andygrove commented on issue #5997:
URL: 
https://github.com/apache/datafusion-comet/issues/5997#issuecomment-5849554596

   Closing this. #5998 took on the first item, and I've abandoned that 
approach. #6250 replaces it with observability.
   
   #5998 charged each task's JVM Arrow allocations to Spark's 
`TaskMemoryManager`, and refused an allocation the pool couldn't cover. 
Backpressure only runs one way in that design. Spark can't reclaim anything 
from Comet's native consumer, because `NativeMemoryConsumer.spill` returns 0 
(item 2 here, tracked as #3873). The JVM Arrow buffers can't be spilled either.
   
   Native operators keep reserving until they are refused, and only then spill, 
so they fill the task's share of the pool. The JVM allocation that decodes 
their next input batch asks for memory just before native would have been 
refused, so it is the one that gets refused. The task fails where it used to 
spill.
   
   I reproduced this with a sort within partitions, and with a repartition, 
directly over a Comet-cached table: 4 tasks and a 128 MiB off-heap pool. With 
the charge on, both failed with `Unable to reserve ... for a JVM Arrow 
allocation ... got 0`. Spark's memory dump showed the native sort holding 31.7 
to 31.8 MiB of each task's 32 MiB share. With the charge off, both succeeded, 
and the sort spilled 24 times.
   
   Dropping the refusal doesn't rescue the approach. The charge would then 
bound nothing when the pool is full, which is the only time a bound matters. 
With default settings it mostly charged broadcast decodes, which native 
reserves right afterwards anyway. The paths where the charge would have 
mattered are the Comet cache, `sparkToColumnar` under a native shuffle, and 
PyArrow UDF output, and those are exactly the paths where it fails. For 
comparison, Spark doesn't charge its own JVM Arrow buffers either. Its Python 
runners allocate from `ArrowUtils.rootAllocator`, a plain 
`RootAllocator(Long.MaxValue)`.
   
   Bounding this memory safely needs native reclaim first. Until then, #6250 
reports the JVM Arrow totals in the executor's memory usage log and counts them 
in the container warning, so the overhead can be sized for them.
   
   The other items here:
   
   - Native reclaim is tracked by #3873.
   - Validating the off-heap settings and sizing the overhead overlaps #6050.
   - Rounding native reservations to blocks has no issue of its own yet.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to