andygrove commented on issue #5997: URL: https://github.com/apache/datafusion-comet/issues/5997#issuecomment-5849554596
Closing this. #5998 took on the first item, and I've abandoned that approach. #6250 replaces it with observability. #5998 charged each task's JVM Arrow allocations to Spark's `TaskMemoryManager`, and refused an allocation the pool couldn't cover. Backpressure only runs one way in that design. Spark can't reclaim anything from Comet's native consumer, because `NativeMemoryConsumer.spill` returns 0 (item 2 here, tracked as #3873). The JVM Arrow buffers can't be spilled either. Native operators keep reserving until they are refused, and only then spill, so they fill the task's share of the pool. The JVM allocation that decodes their next input batch asks for memory just before native would have been refused, so it is the one that gets refused. The task fails where it used to spill. I reproduced this with a sort within partitions, and with a repartition, directly over a Comet-cached table: 4 tasks and a 128 MiB off-heap pool. With the charge on, both failed with `Unable to reserve ... for a JVM Arrow allocation ... got 0`. Spark's memory dump showed the native sort holding 31.7 to 31.8 MiB of each task's 32 MiB share. With the charge off, both succeeded, and the sort spilled 24 times. Dropping the refusal doesn't rescue the approach. The charge would then bound nothing when the pool is full, which is the only time a bound matters. With default settings it mostly charged broadcast decodes, which native reserves right afterwards anyway. The paths where the charge would have mattered are the Comet cache, `sparkToColumnar` under a native shuffle, and PyArrow UDF output, and those are exactly the paths where it fails. For comparison, Spark doesn't charge its own JVM Arrow buffers either. Its Python runners allocate from `ArrowUtils.rootAllocator`, a plain `RootAllocator(Long.MaxValue)`. Bounding this memory safely needs native reclaim first. Until then, #6250 reports the JVM Arrow totals in the executor's memory usage log and counts them in the container warning, so the overhead can be sized for them. The other items here: - Native reclaim is tracked by #3873. - Validating the off-heap settings and sizing the overhead overlaps #6050. - Rounding native reservations to blocks has no issue of its own yet. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
