andygrove commented on PR #5533: URL: https://github.com/apache/datafusion-comet/pull/5533#issuecomment-5876573343
This is a light fully automated review since there are so many PRs open. When a filter that decodes sits under a limit, this also takes the Parquet scan out of native execution. Spark's `FileSourceStrategy` copies every deterministic data-column predicate into `dataFilters`, including ones it can't push into Parquet, so for the fixture query at `unbase64_operator_masks.sql:35` the scan's own `dataFilters` hold `unbase64(bad) <=> X'616263'`. `originalPlan` at `CometExecRule.scala:892` exposes them, `limitName` at line 968 matches on the scan, and line 1010 rebuilds it as a Spark `FileSourceScanExec`. As far as I can tell the native reader never decodes there by default. Per the comment at `native/core/src/parquet/parquet_exec.rs:201`, data filters are only evaluated per row when `spark.comet.parquet.rowFilterPushdown.enabled` is on, and that defaults to `false`. Otherwise they only feed row-group, page-index and bloom-filter pruning, and DataFusion's pruning rewrite rejects a function call over a column. The Spark `Filter` above, which already stays in Spark, is what decodes row by row. Could the scan only count as an evaluation site when row filter pushdown is enabled? `checkSparkAnswerAndFallbackReason` passes either way, so a `nativeScans(plan) == 1` check on the same query at `CometEvaluationMaskSuite.scala:192`, like the one at line 235, would pin it. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
