adriangb opened a new pull request, #25780: URL: https://github.com/apache/datafusion/pull/25780
## Which issue does this PR close? - Closes https://github.com/apache/datafusion/issues/25779. ## Rationale for this change With `datafusion.execution.parquet.pushdown_filters = false` (the default), a Parquet scan claimed equivalence properties from a predicate that it uses only for pruning. An order-preserving repartition directly on the scan then merged on the wrong keys, and `ORDER BY` and `ORDER BY ... LIMIT` returned wrong results. The issue has the reproducer. ## What changes are included in this PR? | Change | File | | --- | --- | | New `FileSource::exact_filter()`: the part of `filter()` that every output row satisfies. Default `None`. | `datasource/src/file.rs` | | `FileScanConfig::eq_properties` uses `exact_filter()` instead of `filter()`. `filter()` is unchanged, because `statistics()` still needs it to mark row counts inexact after pruning. | `datasource/src/file_scan_config/mod.rs` | | `ParquetSource::exact_filter()` returns `None` when `pushdown_filters` is off. When it is on, it returns only the conjuncts that pass `can_expr_be_pushed_down_with_schemas`, the same test `try_pushdown_filters` uses before it replies `PushedDown::Yes`. | `datasource-parquet/src/source.rs` | | Upgrade guide entry for custom `FileSource` implementations. | `upgrading/56.0.0.md` | The behavior from https://github.com/apache/datafusion/pull/16686 / https://github.com/apache/datafusion/issues/16563 is kept: with `pushdown_filters = true` the scan still reports the equivalences, and the existing plans do not change. ## What is the testing strategy for this PR? - New `sqllogictest` cases in `push_down_filter_parquet.slt`: one `EXPLAIN` (the repartition now merges on `a, b`) and four result checks (`ORDER BY`, `ORDER BY ... LIMIT`, `GROUP BY ... ORDER BY`, and the `a = b` form). Without the fix, all 5 fail. - `cte.slt`: one expected plan changes. The scan no longer claims that `k` is constant from its pruning-only `k = 2` predicate, so it reports its real `[k]` ordering, and the single-input `RepartitionExec` now shows `maintains_sort_order=true`. - Differential check against DuckDB: 14 query shapes (order by, top-k, group by, window, hash join with dynamic filter, sort-merge join, `DISTINCT ON`, `first_value`, `a = b`) × 18 configurations (`pushdown_filters`, `prefer_existing_sort`, `repartition_file_scans`, `target_partitions`, `prefer_hash_join`). 15 failures before this change, 0 after. ## Are there any user-facing changes? Queries return correct results. A custom `FileSource` that applies its filter to every row must implement `exact_filter` to keep the equivalence properties of its scan (see the upgrade guide). This is a new trait method with a default, so it is not a breaking change at compile time. 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
