pingzh opened a new pull request, #6496:
URL: https://github.com/apache/datafusion-comet/pull/6496

   ## Which issue does this PR close?
   
   Closes #6123.
   
   ## Rationale for this change
   
   The runtime-filter schema guard disables reader pruning for permitted 
INT32-to-BIGINT promotion and struct subset projection, and rewrites every 
required column for every file even when adaptation is a direct mapping. 
Preserve the conversion-error protections from #6067 while restoring pruning 
for proven-infallible adaptations and reducing eligibility work for matching 
fields.
   
   ## What changes are included in this PR?
   
   1. Centralize an explicit read-adaptation safety contract beside the Spark 
schema adapter. Accept permitted INT32-to-BIGINT promotion and validated 
structural subsets, retain the DataFusion casts needed for leaf clipping, and 
keep unknown/fallible adaptations ineligible. Add pruning/error regressions and 
a repeatable five-fixture benchmark.
   2. Retain the concrete Spark adapter factory in type-keyed file extensions, 
verify it is the scan's active factory, and resolve matching required fields 
without eligibility rewrites. Keep normal handling for unresolved names, 
defaults, duplicates, Variant normalization, and other adaptations. Preserve 
unrelated extensions and generic adapters.
   
   Static predicates, partition exclusions, supplied-file-statistics fallback, 
and decoded-batch filtering keep their existing behavior. No new configuration 
or production metrics.
   
   ## How are these changes tested?
   
   - Native release tests: 53 dynamic-filter tests and 99 schema-adapter tests 
pass. New coverage includes pruning for promotion/struct subsets, unrelated 
unsafe payloads, retained nested timestamp overflow, static-predicate 
dependencies, zero eligibility rewrites for a 256-column scan, case/field-ID 
mappings, defaults, unresolved names, Variant, and replaced factories.
   - Both new pruning regressions fail on unchanged upstream: zero groups 
pruned instead of seven. The first commit restores reader pruning while keeping 
the existing conversion-error regressions passing.
   - Spark 4.1.3 / JDK 21: all 14 `CometJoinSuite` tests matching `join dynamic 
filter` pass. The packaged JNI library matches the rebuilt library's SHA-256.
   - Clippy for native library/test targets with warnings denied, Cargo 
formatting, Maven Spotless/Scalastyle, and diff whitespace checks pass.
   
   ### Release benchmark
   
   Compare upstream `62ed9d16d` with this branch using independently archived 
native test binaries and the same generated Parquet files. Standard release 
profile (optimization level 3, thin LTO, one codegen unit), Rust 1.97.1, 
DataFusion 55.1.0; one scan partition, batch size 8192, Snappy, dictionaries 
disabled, warm OS/metadata caches, and page-index/row-filter pushdown disabled 
to isolate row-group pruning. Fresh planning and execution are timed; fixture 
generation and session creation are excluded.
   
   The direct/promotion/struct cases use 16 files × 65,536 rows with 8,192-row 
groups. Wide cases use 128 files × 2,048 rows × 64 columns with 1,024-row 
groups. The build key is 0; selective probe keys are ordered and unique, and 
the non-pruning case alternates 0/1 in every group. Run order is baseline → 
final → final → baseline, with two warmups and seven measured ON/OFF samples 
per process, yielding fourteen samples per revision/case/mode. An additional 
seven-sample intermediate run shows restored pruning before the matching-field 
fast path; intermediate timing differences are indicative.
   
   | Case | Filter | Baseline ms | First commit ms¹ | Final ms | Change | Data 
bytes baseline → final | Groups read baseline → final |
   |---|---|---:|---:|---:|---:|---:|---:|
   | direct | OFF | 17.08 | 11.29 | 11.73 | -31.3% | 8,395,970 → 8,395,970 | 
128 → 128 |
   | direct | ON | 1.15 | 1.05 | 1.09 | -4.5% | 65,592 → 65,592 | 1 → 1 |
   | promotion | OFF | 17.06 | 12.40 | 12.53 | -26.5% | 8,395,970 → 8,395,970 | 
128 → 128 |
   | promotion | ON | 17.27 | 1.13 | 1.16 | -93.3% | 8,395,970 → 65,592 | 128 → 
1 |
   | struct_subset | OFF | 21.03 | 15.01 | 17.43 | -17.1% | 12,593,955 → 
12,593,955 | 128 → 128 |
   | struct_subset | ON | 20.07 | 1.50 | 1.46 | -92.7% | 12,593,955 → 98,388 | 
128 → 1 |
   | wide_selective | OFF | 149.12 | 141.67 | 146.90 | -1.5% | 67,518,016 → 
67,518,016 | 256 → 256 |
   | wide_selective | ON | 34.11 | 34.70 | 29.91 | -12.3% | 263,680 → 263,680 | 
1 → 1 |
   | wide_non_pruning | OFF | 151.88 | 153.67 | 155.42 | +2.3% | 66,519,879 → 
66,519,879 | 256 → 256 |
   | wide_non_pruning | ON | 191.95 | 193.52 | 186.81 | -2.7% | 66,519,879 → 
66,519,879 | 256 → 256 |
   
   ¹ The intermediate commit has seven samples per case/mode; baseline and 
final each have fourteen.
   
   Requested data bytes are `scan_io_data_bytes`, excluding metadata; measured 
metadata bytes are zero. Output counts, filter attachment, and deterministic 
reader work are checked. Timings are local microbenchmark observations; the 
row-group and byte reductions are deterministic. The ignored 
`schema_guard_benchmark` test documents how to reproduce these fixtures and 
compare archived binaries.
   
   The wide selective ON median improves 12.3%, with its OFF control changing 
-1.5%; wide non-pruning ON improves 2.7%, with OFF changing +2.3%. Narrow-scan 
OFF controls vary by 17–31%, so their timing differences should not be 
attributed to this change. The firm result for promotion and struct subsets is 
restoring pruning from 128 groups read to 1.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to