rich7420 opened a new pull request, #6548:
URL: https://github.com/apache/datafusion-comet/pull/6548

   ## Which issue does this PR close?
   
   Closes #5994.
   
   ## Rationale for this change
   
   Spark hashes decimals with precision above 18 using the minimal signed 
big-endian bytes of the unscaled value. Comet currently falls back for SQL 
`hash` and `xxhash64`, while native shuffle hashes wide decimal keys as 
fixed-width little-endian bytes. The shuffle mismatch can lose matching rows 
when one join input uses native shuffle and the other uses Spark-compatible 
shuffle.
   
   This fixes the shared native hash encoding. #6005 addresses the shuffle 
mismatch by avoiding native shuffle for wide decimal keys. This PR instead 
keeps native shuffle and makes its partition assignment match Spark. If #6005 
lands first, its guard will need to be reconciled with this implementation.
   
   ## What changes are included in this PR?
   
   - Encode decimal precision 19–38 using a 16-byte stack buffer and trim 
redundant sign bytes. Precision at most 18 retains unscaled-long hashing.
   - Enable SQL `hash` and `xxhash64` for wide decimals, including arrays, 
structs and maps. Keep wide decimals on Comet's xxhash64 kernel because the 
pinned DataFusion kernel uses a different encoding.
   - Use typed decimal buffers for List, LargeList and FixedSizeList hashing, 
avoiding per-element Arrow slices and recursive dispatch.
   - Add independent signed-byte reference tests, native execution assertions, 
Spark partition-ID parity, and a join between native Parquet and Spark JSON 
inputs. The join covers native/Spark and native/JVM-columnar shuffle with AQE 
on and off, and checks the executed plan.
   - Add Criterion cases, a whole-query benchmark and update hash support/audit 
documentation.
   
   ## How are these changes tested?
   
   Rebased onto upstream main `ca1c143af`. The original two feature commits 
retain identical patches, and the additional commit only strengthens 
mixed-shuffle regression coverage.
   
   Fresh local verification on the rebased tree (Apple M3 Pro, JDK 21, Spark 
4.1.3):
   
   - Release native build and root-reactor JVM/test compilation passed.
   - 42 hash tests and 2 wide-decimal shuffle tests passed. The final 
executed-plan assertion in the mixed-shuffle test was rerun separately and 
passed. No failed, canceled, ignored or pending tests, and no aborted suites.
   - 86 Rust hash tests passed, including signed-byte reference and decimal 
list/routing regressions.
   - Rustfmt, Maven Spotless/scalastyle and Apache RAT passed.
   
   Performance evidence was collected on pre-rebase feature head `27faa50ed`; 
this rebase does not change the decimal kernels. All 36 Criterion array cases 
improved against the Spark-compatible implementation before the typed-list 
optimization: Murmur3 5.0–14.2x and xxhash64 2.1–15.8x across the three list 
types, 2/32 elements, sliced inputs and no/sparse/dense nulls. Paired scalar 
remeasurement found no significant slowdown.
   
   Two whole-query runs consumed every result through a checksum and verified 
Spark parity, native expression routing and no JVM codegen dispatch. With 
1,048,576 Parquet rows, native array Murmur3 averaged 85/86 ms versus 159/176 
ms with Spark hash fallback; native array xxhash64 averaged 89/83 ms versus 
167/168 ms with fallback. These are local workload measurements.
   
   The pre-rebase head passed [fork 
CI](https://github.com/rich7420/datafusion-comet/actions/runs/36451221737), 
including all Comet Spark profiles and Spark 4.1 SQL shards. CI for the new 
head must run separately; keep this PR draft until that verdict. The fork draft 
retains `run-all-spark-profiles` and `run-spark-4.1-tests`.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to