comphead commented on issue #3532: URL: https://github.com/apache/datafusion-comet/issues/3532#issuecomment-5898618614
Follow-up on scope: the three leaks above are specific to native C2R. I went through every production call site that passes Arrow data between the JVM and native code, on `main` at `c4dd525`. - **Scan input (JVM → native):** native moves the whole stream with `FFI_ArrowArrayStream::from_raw` ([aligned_stream_reader.rs#L44](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/native/core/src/execution/operators/aligned_stream_reader.rs#L44)). The JVM reader closes each source batch in a `finally` ([ColumnarBatchArrowReader.scala#L71-L72](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/spark/src/main/scala/org/apache/spark/sql/comet/execution/arrow/ColumnarBatchArrowReader.scala#L71-L72)), and the stream and its per-task allocator close at task end ([CometNativeArrowSource.scala#L251-L253](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/spark/src/main/scala/org/apache/spark/sql/comet/execution/arrow/CometNativeArrowSource.scala#L251-L253)). This path can't carry a `ConstantColumnVector`, because the reader casts every column to `CometVector` ([#L62](htt ps://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/spark/src/main/scala/org/apache/spark/sql/comet/execution/arrow/ColumnarBatchArrowReader.scala#L62)). - **Exec output and shuffle decode (native → JVM):** `move_to_spark` builds owned structs and moves them into JVM memory as its last step ([utils.rs#L42-L68](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/native/core/src/execution/utils.rs#L42-L68)). `getNextBatch` releases and closes every struct on failure and at EOF ([NativeUtil.scala#L192-L209](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/spark/src/main/scala/org/apache/comet/vector/NativeUtil.scala#L192-L209)). - **JVM UDF bridge (native → JVM → native):** native keeps every struct in a `Box` that releases on drop unless the JVM consumed it ([jvm_udf/mod.rs#L145-L181](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/native/spark-expr/src/jvm_udf/mod.rs#L145-L181)). The JVM closes its result in a `finally` after exporting it ([CometUdfBridge.java#L266-L290](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/spark/src/main/java/org/apache/comet/udf/CometUdfBridge.java#L266-L290)). Native C2R is the only path that reads structs with `ptr::read`, leaves JVM-allocated structs open, or calls `exportBatch`. It's off by default, so none of these leaks affect a default configuration. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
