pingzh opened a new pull request, #6496: URL: https://github.com/apache/datafusion-comet/pull/6496
## Which issue does this PR close? Closes #6123. ## Rationale for this change The runtime-filter schema guard disables reader pruning for permitted INT32-to-BIGINT promotion and struct subset projection, and rewrites every required column for every file even when adaptation is a direct mapping. Preserve the conversion-error protections from #6067 while restoring pruning for proven-infallible adaptations and reducing eligibility work for matching fields. ## What changes are included in this PR? 1. Centralize an explicit read-adaptation safety contract beside the Spark schema adapter. Accept permitted INT32-to-BIGINT promotion and validated structural subsets, retain the DataFusion casts needed for leaf clipping, and keep unknown/fallible adaptations ineligible. Add pruning/error regressions and a repeatable five-fixture benchmark. 2. Retain the concrete Spark adapter factory in type-keyed file extensions, verify it is the scan's active factory, and resolve matching required fields without eligibility rewrites. Keep normal handling for unresolved names, defaults, duplicates, Variant normalization, and other adaptations. Preserve unrelated extensions and generic adapters. Static predicates, partition exclusions, supplied-file-statistics fallback, and decoded-batch filtering keep their existing behavior. No new configuration or production metrics. ## How are these changes tested? - Native release tests: 53 dynamic-filter tests and 99 schema-adapter tests pass. New coverage includes pruning for promotion/struct subsets, unrelated unsafe payloads, retained nested timestamp overflow, static-predicate dependencies, zero eligibility rewrites for a 256-column scan, case/field-ID mappings, defaults, unresolved names, Variant, and replaced factories. - Both new pruning regressions fail on unchanged upstream: zero groups pruned instead of seven. The first commit restores reader pruning while keeping the existing conversion-error regressions passing. - Spark 4.1.3 / JDK 21: all 14 `CometJoinSuite` tests matching `join dynamic filter` pass. The packaged JNI library matches the rebuilt library's SHA-256. - Clippy for native library/test targets with warnings denied, Cargo formatting, Maven Spotless/Scalastyle, and diff whitespace checks pass. ### Release benchmark Compare upstream `62ed9d16d` with this branch using independently archived native test binaries and the same generated Parquet files. Standard release profile (optimization level 3, thin LTO, one codegen unit), Rust 1.97.1, DataFusion 55.1.0; one scan partition, batch size 8192, Snappy, dictionaries disabled, warm OS/metadata caches, and page-index/row-filter pushdown disabled to isolate row-group pruning. Fresh planning and execution are timed; fixture generation and session creation are excluded. The direct/promotion/struct cases use 16 files × 65,536 rows with 8,192-row groups. Wide cases use 128 files × 2,048 rows × 64 columns with 1,024-row groups. The build key is 0; selective probe keys are ordered and unique, and the non-pruning case alternates 0/1 in every group. Run order is baseline → final → final → baseline, with two warmups and seven measured ON/OFF samples per process, yielding fourteen samples per revision/case/mode. An additional seven-sample intermediate run shows restored pruning before the matching-field fast path; intermediate timing differences are indicative. | Case | Filter | Baseline ms | First commit ms¹ | Final ms | Change | Data bytes baseline → final | Groups read baseline → final | |---|---|---:|---:|---:|---:|---:|---:| | direct | OFF | 17.08 | 11.29 | 11.73 | -31.3% | 8,395,970 → 8,395,970 | 128 → 128 | | direct | ON | 1.15 | 1.05 | 1.09 | -4.5% | 65,592 → 65,592 | 1 → 1 | | promotion | OFF | 17.06 | 12.40 | 12.53 | -26.5% | 8,395,970 → 8,395,970 | 128 → 128 | | promotion | ON | 17.27 | 1.13 | 1.16 | -93.3% | 8,395,970 → 65,592 | 128 → 1 | | struct_subset | OFF | 21.03 | 15.01 | 17.43 | -17.1% | 12,593,955 → 12,593,955 | 128 → 128 | | struct_subset | ON | 20.07 | 1.50 | 1.46 | -92.7% | 12,593,955 → 98,388 | 128 → 1 | | wide_selective | OFF | 149.12 | 141.67 | 146.90 | -1.5% | 67,518,016 → 67,518,016 | 256 → 256 | | wide_selective | ON | 34.11 | 34.70 | 29.91 | -12.3% | 263,680 → 263,680 | 1 → 1 | | wide_non_pruning | OFF | 151.88 | 153.67 | 155.42 | +2.3% | 66,519,879 → 66,519,879 | 256 → 256 | | wide_non_pruning | ON | 191.95 | 193.52 | 186.81 | -2.7% | 66,519,879 → 66,519,879 | 256 → 256 | ¹ The intermediate commit has seven samples per case/mode; baseline and final each have fourteen. Requested data bytes are `scan_io_data_bytes`, excluding metadata; measured metadata bytes are zero. Output counts, filter attachment, and deterministic reader work are checked. Timings are local microbenchmark observations; the row-group and byte reductions are deterministic. The ignored `schema_guard_benchmark` test documents how to reproduce these fixtures and compare archived binaries. The wide selective ON median improves 12.3%, with its OFF control changing -1.5%; wide non-pruning ON improves 2.7%, with OFF changing +2.3%. Narrow-scan OFF controls vary by 17–31%, so their timing differences should not be attributed to this change. The firm result for promotion and struct subsets is restoring pruning from 128 groups read to 1. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
