kszucs opened a new issue, #51684: URL: https://github.com/apache/arrow/issues/51684
### Describe the bug, including details regarding any error messages, version, and platform. While implementing content-defined chunking for parquet-java (https://github.com/apache/parquet-java/pull/3818), I discovered that the C++ chunker's state isn't reused between row groups, despite the original intention. The arrow-rs implementation does this correctly. The chunker is owned by the column writer, which only lives for one column chunk: ``` ParquetFileWriter / FileSerializer one per file └── RowGroupWriter / RowGroupSerializer new for every row group └── ColumnWriterImpl new for every column chunk └── ContentDefinedChunker new, with an empty rolling hash state ``` So every `AppendRowGroup()` restarts the rolling hash, and page boundaries depend on where a row group starts, not only on the data. An insertion or deletion that shifts the row group boundaries changes the first pages of every following row group, which reduces deduplication. The tests didn't catch this: they diff the page lengths and only check the number and size of the changed hunks. The pages changed by the reset merge into the hunk a shifted row group is expected to have at its start, while the shift alone should change a single page on each side. Expected: the chunker state carries over between row groups, so page boundaries depend only on the data. ### Component(s) C++, Parquet -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
