comphead commented on issue #3532:
URL: 
https://github.com/apache/datafusion-comet/issues/3532#issuecomment-5898618614

   Follow-up on scope: the three leaks above are specific to native C2R. I went 
through every production call site that passes Arrow data between the JVM and 
native code, on `main` at `c4dd525`.
   
   - **Scan input (JVM → native):** native moves the whole stream with 
`FFI_ArrowArrayStream::from_raw` 
([aligned_stream_reader.rs#L44](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/native/core/src/execution/operators/aligned_stream_reader.rs#L44)).
 The JVM reader closes each source batch in a `finally` 
([ColumnarBatchArrowReader.scala#L71-L72](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/spark/src/main/scala/org/apache/spark/sql/comet/execution/arrow/ColumnarBatchArrowReader.scala#L71-L72)),
 and the stream and its per-task allocator close at task end 
([CometNativeArrowSource.scala#L251-L253](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/spark/src/main/scala/org/apache/spark/sql/comet/execution/arrow/CometNativeArrowSource.scala#L251-L253)).
 This path can't carry a `ConstantColumnVector`, because the reader casts every 
column to `CometVector` ([#L62](htt
 
ps://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/spark/src/main/scala/org/apache/spark/sql/comet/execution/arrow/ColumnarBatchArrowReader.scala#L62)).
   - **Exec output and shuffle decode (native → JVM):** `move_to_spark` builds 
owned structs and moves them into JVM memory as its last step 
([utils.rs#L42-L68](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/native/core/src/execution/utils.rs#L42-L68)).
 `getNextBatch` releases and closes every struct on failure and at EOF 
([NativeUtil.scala#L192-L209](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/spark/src/main/scala/org/apache/comet/vector/NativeUtil.scala#L192-L209)).
   - **JVM UDF bridge (native → JVM → native):** native keeps every struct in a 
`Box` that releases on drop unless the JVM consumed it 
([jvm_udf/mod.rs#L145-L181](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/native/spark-expr/src/jvm_udf/mod.rs#L145-L181)).
 The JVM closes its result in a `finally` after exporting it 
([CometUdfBridge.java#L266-L290](https://github.com/apache/datafusion-comet/blob/c4dd52503a59567db667810afde0f277588df5d1/spark/src/main/java/org/apache/comet/udf/CometUdfBridge.java#L266-L290)).
   
   Native C2R is the only path that reads structs with `ptr::read`, leaves 
JVM-allocated structs open, or calls `exportBatch`. It's off by default, so 
none of these leaks affect a default configuration.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to