[
https://issues.apache.org/jira/browse/SPARK-59434?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Cheng Pan updated SPARK-59434:
------------------------------
Affects Version/s: 3.4.0
(was: 5.0.0)
> Non-deterministic predicates should not be pushed down to CTE definitions
> -------------------------------------------------------------------------
>
> Key: SPARK-59434
> URL: https://issues.apache.org/jira/browse/SPARK-59434
> Project: Spark
> Issue Type: Bug
> Components: SQL
> Affects Versions: 3.4.0
> Reporter: Cheng Pan
> Priority: Major
>
> {{PushdownPredicatesAndPruneColumnsForCTEDef}} collects the predicates at
> each CTE reference and pushes their OR-merged combination into the shared CTE
> definition. Each reference keeps its own predicates, so a non-deterministic
> predicate pushed into the definition as well is evaluated twice, which can
> produce wrong results.
> {code:sql}
> with v as (select c1, rand(1) r from t)
> select c1 from v where rand(2) < 0.5
> union all
> select c1 from v where rand(3) < 0.5
> {code}
> The definition is non-deterministic and referenced twice, so it survives
> {{InlineCTE}}. {{(rand(2) < 0.5) OR (rand(3) < 0.5)}} lands in the definition
> while both references keep their own filter, so the optimized plan holds four
> {{rand}} filters instead of two, and rows are dropped twice.
> No {{MATERIALIZED}} option is needed to reach this: any definition that
> survives {{InlineCTE}} has the same shape. The rule's scaladoc claims
> determinism is taken care of by {{ScanOperation}}, but SPARK-39764 (3.4.0)
> replaced that with {{PhysicalOperation}}, which hands back a single filter
> even when it is non-deterministic.
> Fix: only push deterministic predicates into the CTE definition.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]