peterxcli opened a new pull request, #6342:
URL: https://github.com/apache/datafusion-comet/pull/6342

   ## Which issue does this PR close?
   
   Related to #5978. This is an allocation improvement; #5978 remains open for 
single-pass reconstruction and its performance target.
   
   ## Rationale for this change
   
   Shredded Variant reconstruction creates value and metadata builders for 
every row, then copies their output into batch buffers. Reuse those buffers 
across rows to reduce allocation traffic with the current Arrow dependency.
   
   ## What changes are included in this PR?
   
   Append reconstructed values and metadata directly to batch buffers and build 
the final Binary arrays from their offsets. Reset dictionary state between 
rows, preserve parent nulls, and avoid intermediate Binary casts before 
reconstruction.
   
   ## How are these changes tested?
   
   Exact value and metadata comparisons cover changing dictionaries, large 
values, null rows and sliced arrays. The 16 native Variant tests, all 13 
`CometVariantProjectionSuite` tests, and six targeted upstream Spark 4.1.3 
Parquet Variant tests pass. The upstream assertions are unchanged, with an 
additional check confirming native scans are active.
   
   Matched builds against `e897f8ab4`, on Apple M4 with Rust 1.95, the 
optimized `ci` profile and jemalloc. Five alternating native runs, 4,096 rows 
with 4 KiB payloads and 30 measured batches per case, gave identical allocation 
counts in every run:
   
   | Input | Main bytes/row | PR bytes/row | Reduction |
   | --- | ---: | ---: | ---: |
   | Canonical | 10,332 | 10,332 | 0% |
   | Partially shredded | 57,402 | 46,964 | 18.2% |
   | Fully shredded | 52,320 | 41,882 | 20.0% |
   | Empty key | 88,760 | 78,375 | 11.7% |
   
   These are allocator bytes requested, not retained memory. Both builds use 
the same benchmark fixtures, including the updated partially shredded empty-key 
fixture.
   
   `CometVariantReadBenchmark 100000 1024` consumes both value and metadata 
bytes from matched Parquet inputs. Four JVM runs per build alternate build 
order and run both reader orders twice, using Spark 4.1.3 / JDK 17. Median best 
scan times in milliseconds:
   
   | Input | Main | PR | Change | Vanilla Spark |
   | --- | ---: | ---: | ---: | ---: |
   | Canonical | 143 | 146.5 | +2.4% | 143.5 |
   | Partially shredded | 298 | 271.5 | -8.9% | 165 |
   | Fully shredded | 255 | 268.5 | +5.3% | 166.5 |
   | Empty key | 385 | 373.5 | -3.0% | 159.5 |
   
   Spark medians pool the eight control runs. Concurrent machine activity makes 
timing inconclusive: even Spark's fully shredded control ranges from 150 to 265 
ms. The allocation reduction is reproducible, but this does not establish a 
scan speedup or rule out a regression. Shredded reads remain slower than Spark. 
Keeping this PR in draft until a quieter run resolves the scan regression 
question.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to