[
https://issues.apache.org/jira/browse/SPARK-59347?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Anish Mahto updated SPARK-59347:
--------------------------------
Summary: SCD2 Ignore-null support; union microbatch with target table
during preprocessing (was: SCD2 Ignore-null support; apply version map key
renames after union with target)
> SCD2 Ignore-null support; union microbatch with target table during
> preprocessing
> ---------------------------------------------------------------------------------
>
> Key: SPARK-59347
> URL: https://issues.apache.org/jira/browse/SPARK-59347
> Project: Spark
> Issue Type: Sub-task
> Components: Declarative Pipelines
> Affects Versions: 4.4.0
> Reporter: Anish Mahto
> Priority: Major
>
> Part of the contract for the SCD2 version map is the keys in the version map
> should track exactly with the data column names.
> Complexity arises when case insensitivity is used, and the incoming
> microbatch has a different case-spelling for certain columns than the
> persisted target/aux tables.
> When `Scd2ForeachBatchHandler` does a [union with the incoming microbatch
> rows and the existing target/aux table
> rows|https://github.com/apache/spark/blob/master/sql/pipelines/src/main/scala/org/apache/spark/sql/pipelines/autocdc/Scd2ForeachBatchHandler.scala#L104-L106],
> due to ordering of the union operators, the resulting DF will respect the
> case-spelling on the incoming microbatch.
> But when this dataframe is eventually merged back into the target/aux tables,
> the existing case-spelling on those tables will win. This is no longer
> acceptable since the version map is constructed by referencing column names
> in the dataframe. If those column names change, the version map's keys need
> to also change for consistency.
> The proposal is two-fold:
> 1. The 3-way union should setup target table's df as the first operator, so
> that from this point forward rows always respect the persisted target table's
> case-spellings
> 2. Rename the in-memory version map keys on `preprocessedBatchDf` to use the
> same case-spelling as the target table, right before the union operation
> Then both before and after the union, the version map keys match 1:1 with
> data column names, under whatever case-sensitivity scheme is configured.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]