rich7420 opened a new pull request, #6548: URL: https://github.com/apache/datafusion-comet/pull/6548
## Which issue does this PR close? Closes #5994. ## Rationale for this change Spark hashes decimals with precision above 18 using the minimal signed big-endian bytes of the unscaled value. Comet currently falls back for SQL `hash` and `xxhash64`, while native shuffle hashes wide decimal keys as fixed-width little-endian bytes. The shuffle mismatch can lose matching rows when one join input uses native shuffle and the other uses Spark-compatible shuffle. This fixes the shared native hash encoding. #6005 addresses the shuffle mismatch by avoiding native shuffle for wide decimal keys. This PR instead keeps native shuffle and makes its partition assignment match Spark. If #6005 lands first, its guard will need to be reconciled with this implementation. ## What changes are included in this PR? - Encode decimal precision 19–38 using a 16-byte stack buffer and trim redundant sign bytes. Precision at most 18 retains unscaled-long hashing. - Enable SQL `hash` and `xxhash64` for wide decimals, including arrays, structs and maps. Keep wide decimals on Comet's xxhash64 kernel because the pinned DataFusion kernel uses a different encoding. - Use typed decimal buffers for List, LargeList and FixedSizeList hashing, avoiding per-element Arrow slices and recursive dispatch. - Add independent signed-byte reference tests, native execution assertions, Spark partition-ID parity, and a join between native Parquet and Spark JSON inputs. The join covers native/Spark and native/JVM-columnar shuffle with AQE on and off, and checks the executed plan. - Add Criterion cases, a whole-query benchmark and update hash support/audit documentation. ## How are these changes tested? Rebased onto upstream main `ca1c143af`. The original two feature commits retain identical patches, and the additional commit only strengthens mixed-shuffle regression coverage. Fresh local verification on the rebased tree (Apple M3 Pro, JDK 21, Spark 4.1.3): - Release native build and root-reactor JVM/test compilation passed. - 42 hash tests and 2 wide-decimal shuffle tests passed. The final executed-plan assertion in the mixed-shuffle test was rerun separately and passed. No failed, canceled, ignored or pending tests, and no aborted suites. - 86 Rust hash tests passed, including signed-byte reference and decimal list/routing regressions. - Rustfmt, Maven Spotless/scalastyle and Apache RAT passed. Performance evidence was collected on pre-rebase feature head `27faa50ed`; this rebase does not change the decimal kernels. All 36 Criterion array cases improved against the Spark-compatible implementation before the typed-list optimization: Murmur3 5.0–14.2x and xxhash64 2.1–15.8x across the three list types, 2/32 elements, sliced inputs and no/sparse/dense nulls. Paired scalar remeasurement found no significant slowdown. Two whole-query runs consumed every result through a checksum and verified Spark parity, native expression routing and no JVM codegen dispatch. With 1,048,576 Parquet rows, native array Murmur3 averaged 85/86 ms versus 159/176 ms with Spark hash fallback; native array xxhash64 averaged 89/83 ms versus 167/168 ms with fallback. These are local workload measurements. The pre-rebase head passed [fork CI](https://github.com/rich7420/datafusion-comet/actions/runs/36451221737), including all Comet Spark profiles and Spark 4.1 SQL shards. CI for the new head must run separately; keep this PR draft until that verdict. The fork draft retains `run-all-spark-profiles` and `run-spark-4.1-tests`. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
