haohuaijin opened a new issue, #25797:
URL: https://github.com/apache/datafusion/issues/25797
### Is your feature request related to a problem or challenge?
Parquet Bloom filter pruning currently loads filters for the relevant
predicate columns before evaluating whether a row group can be excluded.
For queries with multiple filtering conditions, some of these reads may be
unnecessary. For example:
```sql
SELECT COUNT(*)
FROM traces
WHERE trace_id = 'abc123'
AND service_name = 'checkout'
AND host = 'host-42';
```
If a row group survives statistics pruning, but its `trace_id` Bloom filter
proves that `'abc123'` is absent, the entire row group can already be excluded.
Reading the `service_name` and `host` Bloom filters cannot change that decision.
This is particularly relevant to observability workloads that filter across
several columns, especially when Parquet files reside in remote object storage.
### Describe the solution you'd like
Evaluate Bloom filters incrementally and stop reading additional filters for
a row group as soon as a necessary condition proves that the group cannot match.
The desired behavior is:
1. Read a relevant Bloom filter.
2. Check whether it rules out the row group.
3. If it does, skip the remaining Bloom filter reads for that group.
4. Otherwise, continue with the remaining relevant columns.
Existing literal guarantees could provide the necessary conditions. For
example, an `IN` condition can exclude a group only when all candidate values
are definitely absent. A possible Bloom hit must remain inconclusive.
This would also allow filters to be released after evaluation instead of
retaining all filters until a separate pruning pass. The implementation should
preserve conservative handling of compound predicates, NULLs, missing filters,
and read failures.
### Describe alternatives you've considered
Optimizing evaluation after loading all filters could reduce CPU overhead,
but would not avoid unnecessary Bloom filter reads.
### Additional context
The relevant loading and pruning flow is in
`datafusion/datasource-parquet/src/opener/mod.rs`, particularly
`load_bloom_filters()` and `prune_bloom_filters()`.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]