andygrove opened a new pull request, #25842:
URL: https://github.com/apache/datafusion/pull/25842

   ## Which issue does this PR close?
   
   - Part of #25758
   - Related to #24769
   
   ## Rationale for this change
   
   On `branch-55`, a Parquet row group is no longer pruned from statistics when 
the predicate is `col = <literal>` and the statistics show `col` is entirely 
NULL, so the row group is read instead of skipped (#24769). Results are 
unchanged, but filters on sparsely populated columns read row groups that 54 
skipped. The regression was bisected to #22969, which shipped in 55.0.0.
   
   ## What changes are included in this PR?
   
   This PR backports #24770 from @jensholdgaard to the `branch-55` line. The 
cherry-pick applied cleanly, with no changes.
   
   ## Are these changes tested?
   
   Yes. The backport includes the tests from #24770. On this branch I ran:
   
   - `cargo test -p datafusion-datasource-parquet`
   - `cargo test -p datafusion --test parquet_integration`
   - `cargo test -p datafusion --test core_integration datasource`
   - the full sqllogictest suite
   - `./dev/rust_lint.sh`
   
   ## Are there any user-facing changes?
   
   Row groups that file statistics prove cannot match are pruned again, as in 
54. `RowGroupAccessPlanFilter` gains a public `skip_all` method. No other 
public API changes.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to