kosiew opened a new issue, #25047:
URL: https://github.com/apache/datafusion/issues/25047

   ## Related PR
   #24757 
   
   
   ## Problem
   `ordered_aggregate_spill.slt:166` intermittently fails with 
`ResourcesExhausted` under its fixed 600 KiB memory limit. The failure involves 
two concurrent `OrderedFinalAggregateStream` consumers sharing a greedy pool: 
one holds about 447.5 KiB, the other about 145.4 KiB, then needs another 8.1 
KiB while only 7.0 KiB remains.
   
   ### Flaky runs as of 
https://github.com/apache/datafusion/tree/8f40f33c211937e9b1724925a9ef152f049fccc5
   
   fail - 
https://github.com/apache/datafusion/actions/runs/34140768577/job/101802015716?pr=24757
   pass - https://github.com/apache/datafusion/actions/runs/34140768577?pr=24757
   
   ## Why it matters
   A flaky SLT obscures real regressions and makes CI unreliable. It may also 
expose an ordered-final spill lifecycle/accounting defect rather than test-only 
nondeterminism.
   
   ## Invariant / desired behavior
   The test must have a reproducible outcome under a fixed environment: either 
it consistently completes at 600 KiB, or it reliably exposes a defined 
resource-accounting failure. Test scheduling must not determine the result.
   
   ## Proposed direction
   First make the failure classification reproducible before changing 
production behavior:
   
   - Run the exact SLT repeatedly with identical Rust version, target 
directory, runner parallelism, pool implementation, and partition count.
   - Repeat enough times to capture the interleaving-dependent result; record 
plan stream types and per-consumer peak/current reservations.
   - Add a focused physical-plan reproduction with two ordered-final partitions 
sharing a finite greedy pool. Force one partition to retain a run while the 
other reaches its spill boundary.
   - Assert the intended result, grouped output, spill metrics, and reservation 
cleanup. Avoid timing-dependent orchestration.
   
   If the focused test proves the spill transition allocates an unreserved 
materialized state batch or sort/write workspace, file/fix that production 
accounting defect separately. Do not weaken this test by raising its memory 
limit.
   
   ## Scope
   ### In
   - Deterministic reproduction and classification of the ordered aggregate 
spill failure.
   - Focused finite-pool, multi-partition regression test coverage.
   - Repeated verification under failure-like parallelism.
   
   ### Out
   - Raising `datafusion.runtime.memory_limit` above 600 KiB.
   - Changing aggregate metric docs, labels/timers, or `EXPLAIN ANALYZE` 
formatting.
   - Attributing the flake to the docs PR without a controlled 
parent-versus-HEAD reproduction.
   
   ## Acceptance criteria
   - [ ] A focused test deterministically exercises two ordered-final consumers 
at the spill boundary in a shared finite greedy pool.
   - [ ] The test's result is stable across repeated runs with fixed 
configuration.
   - [ ] The test asserts correct grouped output, spill behavior, and zero 
reservations after stream drop.
   - [ ] The existing SLT remains at 600 KiB and is repeatedly stable under 
relevant parallelism.
   - [ ] Evidence identifies whether the cause is test scheduling, runner 
configuration, or a production spill-accounting defect.
   
   ## Tests / verification
   - Repeat:
     `cargo test -q -p datafusion-sqllogictest --test sqllogictests -- 
ordered_aggregate_spill --nocapture`
   - Run the new focused `datafusion-physical-plan` aggregate test repeatedly 
with its shared-pool configuration.
   - Run relevant `datafusion-physical-plan` aggregate tests.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to