andygrove opened a new pull request, #25842: URL: https://github.com/apache/datafusion/pull/25842
## Which issue does this PR close? - Part of #25758 - Related to #24769 ## Rationale for this change On `branch-55`, a Parquet row group is no longer pruned from statistics when the predicate is `col = <literal>` and the statistics show `col` is entirely NULL, so the row group is read instead of skipped (#24769). Results are unchanged, but filters on sparsely populated columns read row groups that 54 skipped. The regression was bisected to #22969, which shipped in 55.0.0. ## What changes are included in this PR? This PR backports #24770 from @jensholdgaard to the `branch-55` line. The cherry-pick applied cleanly, with no changes. ## Are these changes tested? Yes. The backport includes the tests from #24770. On this branch I ran: - `cargo test -p datafusion-datasource-parquet` - `cargo test -p datafusion --test parquet_integration` - `cargo test -p datafusion --test core_integration datasource` - the full sqllogictest suite - `./dev/rust_lint.sh` ## Are there any user-facing changes? Row groups that file statistics prove cannot match are pruned again, as in 54. `RowGroupAccessPlanFilter` gains a public `skip_all` method. No other public API changes. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
