kszucs opened a new issue, #51684:
URL: https://github.com/apache/arrow/issues/51684

   ### Describe the bug, including details regarding any error messages, 
version, and platform.
   
   While implementing content-defined chunking for parquet-java 
(https://github.com/apache/parquet-java/pull/3818), I discovered that the C++ 
chunker's state isn't reused between row groups, despite the original 
intention. The arrow-rs implementation does this correctly.
   
   The chunker is owned by the column writer, which only lives for one column 
chunk:
   
   ```
   ParquetFileWriter / FileSerializer           one per file
   └── RowGroupWriter / RowGroupSerializer      new for every row group
       └── ColumnWriterImpl                     new for every column chunk
           └── ContentDefinedChunker            new, with an empty rolling hash 
state
   ```
   
   So every `AppendRowGroup()` restarts the rolling hash, and page boundaries 
depend on where a row group starts, not only on the data. An insertion or 
deletion that shifts the row group boundaries changes the first pages of every 
following row group, which reduces deduplication.
   
   The tests didn't catch this: they diff the page lengths and only check the 
number and size of the changed hunks. The pages changed by the reset merge into 
the hunk a shifted row group is expected to have at its start, while the shift 
alone should change a single page on each side.
   
   Expected: the chunker state carries over between row groups, so page 
boundaries depend only on the data.
   
   ### Component(s)
   
   C++, Parquet
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to