andygrove commented on issue #3532:
URL: 
https://github.com/apache/datafusion-comet/issues/3532#issuecomment-5860285801

   Two more leaks on the same path, and these fire on every batch, not only on 
failure. Both were measured on `main` at `31b38196d` with 
`spark.comet.exec.columnarToRow.native.enabled=true`:
   
   - `NativeUtil.exportBatchToAddresses` allocates an 
`ArrowArray`/`ArrowSchema` pair per column from `CometArrowAllocator`, and 
nothing ever closes them. Native copies their contents out and releases the 
exported data, but the two structs stay allocated, at 128 bytes each after 
Arrow's rounding. Collecting 200k rows × 3 columns grows 
`CometArrowAllocator.getAllocatedMemory` by 19,200 bytes on every run: 25 
batches × 3 columns × 256 bytes.
   - The `ConstantColumnVector` arm of `NativeUtil.exportBatch` never closes 
the vector it materializes. `Data.exportVector` retains each buffer it exports, 
so the JVM's own reference outlives the export. A round trip of one `bigint` 
constant column leaked 32 KiB.
   
   Fixing the first one needs care. `columnarToRowConvert` takes the structs 
with `std::ptr::read`, so the JVM copies still carry a live release callback, 
and releasing them from the JVM, as `NativeUtil.releaseArrowStructs` does, 
would run that callback twice. If native used `FFI_ArrowArray::from_raw` / 
`FFI_ArrowSchema::from_raw` instead, which leave the source released, the JVM 
could release and close every struct in a `finally`. That also covers the error 
path above, and the JVM-side gap where `exportBatch` throws after it has 
exported some of the columns.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to