dwsmith1983 opened a new pull request, #25933:
URL: https://github.com/apache/datafusion/pull/25933

   ## Which issue does this PR close?
   
   - Part of #25923 (the memory accounting half; the per-row filter cost is 
separate).
   
   ## Rationale for this change
   
   A filtered LeftSemi/LeftAnti sort-merge join under a tight memory pool 
spills once per matched key group, even when each group holds only a few rows. 
Each buffered inner key group is charged its parent batch's full size, so a 
group of at most 7 rows is charged 131,264 bytes (a whole 8192-row batch). On 
the benchmark from the issue, with a `GreedyMemoryPool` of 64 KB over 20K keys, 
every one of the 18,128 matched groups spilled to its own file. In a real plan 
the join shares the pool with the sorts in the same task, so this can start 
well before the data needs to spill.
   
   ## What changes are included in this PR?
   
   `buffer_inner_key_group` in `bitwise_stream.rs` now reserves 
`get_sliced_size()` for each buffered slice instead of 
`get_array_memory_size()`, the same measure sort uses. Only the reservation 
(and so the spill decision and `peak_mem_used`) was inflated; `spilled_bytes` 
counts serialized bytes and is unaffected.
   
   Two cases still count more than the rows: view arrays keep their parent's 
data buffers in the measure, and a group that spans an inner batch boundary 
keeps the earlier batch alive while only its rows are charged, so the 
reservation can be up to one batch low there. Both are noted at the call site.
   
   ## What is the testing strategy for this PR?
   
   `bitwise_small_key_groups_charged_by_sliced_size` joins inputs with 2 and 3 
row key groups under a pool of half a batch, for LeftSemi, LeftAnti, RightSemi 
and RightAnti. It checks that nothing spills and that the results match an 
unbounded run. On main it spills 683 times.
   
   Five existing spill tests used a 100-byte pool to force a spill. A one-row 
slice is now charged about 12 bytes, so their pool is 1 byte; their assertions 
are unchanged.
   
   With this change on 55.1.0, the issue's 64 KB pool case goes from 18,128 
spills to none, and the join runs as fast as with an unbounded pool (11.9 ms 
both, where the 64 KB pool previously took 1,088 ms against 9.3 ms unbounded).
   
   ## Are there any user-facing changes?
   
   No API changes. Filtered semi and anti sort-merge joins spill less under 
memory pressure.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to