sunchao opened a new pull request, #25889:
URL: https://github.com/apache/datafusion/pull/25889

   ## Which issue does this PR close?
   
   Part of #23031. Related to #23032; this is a hash-join-specific alternative 
for comparing the implementation and performance tradeoffs, not a replacement 
for its nested-loop or piecewise-merge join work.
   
   ## Rationale for this change
   
   Hash joins currently concatenate the entire build relation into one Arrow 
batch. For a large build this requires admitting a second payload allocation 
while the input batches are still retained, and combines individually valid 
variable-width arrays into one offset-limited array.
   
   This PR keeps the small-build path and retains large builds as batches. The 
scope includes the full build/probe/output path and a reproducible benchmark 
suite, rather than only exposing a storage helper.
   
   ## What changes are included in this PR?
   
   - Select compact versus batched storage using both retained memory and 
estimated logical copy size, keeping tiny slices of large parents compact.
   - Coalesce independent flat inputs within 8 MiB / 8,192-row targets, account 
for copies before allocation, and release unused backing buffers and 
empty-batch charges.
   - Keep batch-local evaluated keys and a sparse logical-row directory; 
compare candidates without concatenating the keys or payload.
   - Reserve the actual retained metadata capacities and share probe 
preprocessing; keep only one source comparator alive at a time, including 
Arrow's dictionary-null scratch buffers.
   - Gather only referenced source batches, with null-safe nested gathering and 
the existing Arrow kernels for other encodings.
   - Support computed and composite keys, dictionary keys/payloads, nested 
payloads, ordinary mark/semi/anti/outer joins, and batched perfect hashing.
   - Preserve the current bounded final-build-row emission, dynamic-filter 
coordination, and reservation lifetime. Optional IN-list construction falls 
back to the hash-table predicate when it cannot be admitted or represented.
   - Add a 20-case physical-plan benchmark covering small builds, large 
payloads, computed/dictionary keys, dictionary/list payloads, tiny batches, 
shared slices, excess backing, and low-match outer joins, with perfect hashing 
enabled and disabled.
   
   The batched path currently applies to ordinary hash joins above the 64 MiB 
compact-build threshold. Prepared reusable builds and null-aware joins retain 
their existing contiguous state contracts. This does not add spilling, change 
other join operators, or wire the representation into Comet.
   
   ## Are these changes tested?
   
   - 1,489 physical join unit tests passed, including the new computed-key, 
encoding, all-join-type, fetch, memory-limit, slice-retention, 
metadata-capacity, and probe-preprocessing regressions.
   - `cargo clippy --workspace --all-targets --all-features -- -D warnings` 
passed.
   - Formatting, TOML formatting, license headers, and spelling checks passed.
   - The extended workspace test suite is running.
   
   Local validation caveat: the dependency mirror does not yet serve the locked 
`thiserror 2.0.21`. Validation copies use `thiserror` and `thiserror-impl` at 
`2.0.20`; the PR's `Cargo.lock` is unchanged. Both benchmark revisions use the 
same temporary dependency lockfile.
   
   ### Performance
   
   Matched measurements are in progress against base 
`1d9be2e10794994fc1495c35363658dc7db0adf7` and head 
`5e3418baf95f1e7f9e951b9c95f50e02711fcd00`. No speedup claim yet. The benchmark 
harness checks output row counts/checksums and separately records peak 
reservations, which are not process RSS.
   
   ## Are there any user-facing changes?
   
   No public API or configuration changes. Large ordinary hash joins may use 
less build memory by avoiding the full payload copy. Query results and the 
existing Arrow key-coercion requirements are unchanged.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to