adriangb opened a new pull request, #25780:
URL: https://github.com/apache/datafusion/pull/25780

   ## Which issue does this PR close?
   
   - Closes https://github.com/apache/datafusion/issues/25779.
   
   ## Rationale for this change
   
   With `datafusion.execution.parquet.pushdown_filters = false` (the default), 
a Parquet scan claimed equivalence properties from a predicate that it uses 
only for pruning. An order-preserving repartition directly on the scan then 
merged on the wrong keys, and `ORDER BY` and `ORDER BY ... LIMIT` returned 
wrong results. The issue has the reproducer.
   
   ## What changes are included in this PR?
   
   | Change | File |
   | --- | --- |
   | New `FileSource::exact_filter()`: the part of `filter()` that every output 
row satisfies. Default `None`. | `datasource/src/file.rs` |
   | `FileScanConfig::eq_properties` uses `exact_filter()` instead of 
`filter()`. `filter()` is unchanged, because `statistics()` still needs it to 
mark row counts inexact after pruning. | 
`datasource/src/file_scan_config/mod.rs` |
   | `ParquetSource::exact_filter()` returns `None` when `pushdown_filters` is 
off. When it is on, it returns only the conjuncts that pass 
`can_expr_be_pushed_down_with_schemas`, the same test `try_pushdown_filters` 
uses before it replies `PushedDown::Yes`. | `datasource-parquet/src/source.rs` |
   | Upgrade guide entry for custom `FileSource` implementations. | 
`upgrading/56.0.0.md` |
   
   The behavior from https://github.com/apache/datafusion/pull/16686 / 
https://github.com/apache/datafusion/issues/16563 is kept: with 
`pushdown_filters = true` the scan still reports the equivalences, and the 
existing plans do not change.
   
   ## What is the testing strategy for this PR?
   
   - New `sqllogictest` cases in `push_down_filter_parquet.slt`: one `EXPLAIN` 
(the repartition now merges on `a, b`) and four result checks (`ORDER BY`, 
`ORDER BY ... LIMIT`, `GROUP BY ... ORDER BY`, and the `a = b` form). Without 
the fix, all 5 fail.
   - `cte.slt`: one expected plan changes. The scan no longer claims that `k` 
is constant from its pruning-only `k = 2` predicate, so it reports its real 
`[k]` ordering, and the single-input `RepartitionExec` now shows 
`maintains_sort_order=true`.
   - Differential check against DuckDB: 14 query shapes (order by, top-k, group 
by, window, hash join with dynamic filter, sort-merge join, `DISTINCT ON`, 
`first_value`, `a = b`) × 18 configurations (`pushdown_filters`, 
`prefer_existing_sort`, `repartition_file_scans`, `target_partitions`, 
`prefer_hash_join`). 15 failures before this change, 0 after.
   
   ## Are there any user-facing changes?
   
   Queries return correct results. A custom `FileSource` that applies its 
filter to every row must implement `exact_filter` to keep the equivalence 
properties of its scan (see the upgrade guide). This is a new trait method with 
a default, so it is not a breaking change at compile time.
   
   🤖 Generated with [Claude Code](https://claude.com/claude-code)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to