namanjain24-sudo opened a new pull request, #25781:
URL: https://github.com/apache/datafusion/pull/25781
## Which issue does this PR close?
- Closes #25708.
## Rationale for this change
#25529 fixed a wrong-results bug where a correlated filter under a
grouping-set `Aggregate` was pulled up into the aggregate's grouping sets, but
the guard it added is more conservative than it needs to be: it rejects the
pull up whenever *any* grouping set omits the correlated column, even in cases
like
```sql
SELECT o.k FROM o
WHERE EXISTS (
SELECT 1 FROM i WHERE i.k = o.k
GROUP BY GROUPING SETS ((i.k), (i.j))
)
```
where nothing above the aggregate ever reads `i.k`, so filling it in changes
nothing the query can observe. These queries now fail to plan (`This feature is
not implemented: ... does not support logical expression Exists`) instead of
decorrelating.
## What changes are included in this PR?
- `datafusion/optimizer/src/decorrelate.rs`:
`grouping_sets_cover_pull_up_cols` becomes `grouping_sets_pull_up_outcome`,
returning a three-way `GroupingSetsPullUp` (`NoColumnsToAdd` / `SafeToExtend` /
`Unsafe`) instead of a bool. A set that omits a required column is now only
`Unsafe` when:
- some grouping set is empty (`ROLLUP`/`CUBE` always hold one), or
- the aggregate itself computes `GROUPING`/`GROUPING_ID`, or
- a node above the aggregate reads the column, per a new
`column_refs_above_aggregate` set.
- Since `PullUpCorrelatedExpr::f_up` runs bottom-up, it hasn't visited
whatever sits above the `Aggregate` (a `HAVING`, a `Projection`) by the time it
needs to decide. `column_refs_above_aggregate` is instead computed once,
top-down, over the original (not yet rewritten) subquery plan by a new
`columns_read_above_aggregate` helper, and threaded in via a new
`with_column_refs_above_aggregate` builder call at all three call sites
(`scalar_subquery_to_join.rs`, `decorrelate_predicate_subquery.rs`,
`decorrelate_lateral_join.rs`).
- `datafusion/sqllogictest/test_files/subquery.slt`: the `EXISTS` case above
now decorrelates and asserts correct results; added a case showing a
`GROUPING(k)` read still keeps the subquery correlated (correctly), alongside
the existing `HAVING`/empty-set cases that stay correlated.
## What is the testing strategy for this PR?
Extended `datafusion/sqllogictest/test_files/subquery.slt`'s existing
`#25519` regression block:
- The previously-rejected `EXISTS ... GROUPING SETS ((i.k), (i.j))` case is
now a `query IB` asserting the correct results.
- Added a case with `GROUPING(k)` read out of the aggregate, asserting it
still correctly stays correlated.
- The `HAVING i.k IS NULL` and every empty-set (`ROLLUP`/`CUBE`/explicit
`()`) case are unchanged and still assert they stay correlated.
Ran the full `datafusion-optimizer` unit + integration suite (905 + 26
tests) and the full sqllogictest suite (523 files), all green.
## Are there any user-facing changes?
More `EXISTS`/`IN`/lateral-join subqueries with a correlated filter under a
grouping-set aggregate now decorrelate into a join instead of failing to plan.
No public API changes.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]