zhuqi-lucas commented on issue #627: URL: https://github.com/apache/parquet-format/issues/627#issuecomment-5893156840
Thanks — that is a much better read of the history than I had, and checking the encoders shows you are right in a way I had backwards. The dictionary page is `PlainEncoder` over the deduplicated values; the data pages are `RleEncoder` over bit-packed indices. The high-entropy argument applies to the data pages, not the dictionary — a dictionary of strings has plenty of shared structure for a codec to work with. So I measured, since my premise deserved checking. On ClickBench `hits.parquet` (105 columns, 226 row groups, SNAPPY), dictionary pages are **14.7% of all compressed bytes**, and for high-cardinality columns they dominate their own chunk: `UserID` 54.5%, `FUniqID` 53.6%, `RefererHash` 45.7%, `WatchID` 27.5%. Leaving those uncompressed is not a cheap trade. The one angle that survives the measurement is that a dictionary is decompressed once per column chunk regardless of how many rows are read, while data page cost scales with what is actually touched. Its share of *decompressed* bytes therefore climbs steeply as a scan gets selective: | rows actually decoded | dictionary share of decompressed bytes | | --- | --- | | 100% | 14.7% | | 10% | 63.4% | | 1% | 94.5% | That is the shape of the workload behind the 13.6–17.9% profile figure. But your two suggestions cover it without a format change, and they cover it on the first scan rather than only on repeats. Where the dictionary dominates, those are the high-cardinality columns where dictionary encoding is barely earning its place to begin with, so turning it off per column is a better answer than making it uncompressed. For the rest, per-column `UNCOMPRESSED` is one line of writer config. We control our writers, so that is the direction we will take. That leaves the format change wanted only where someone needs compressed data pages *and* an uncompressed dictionary in the same column. Narrower than I framed it. And thank you for the versioning pointer — reading that vote, the gating question I raised in the issue is already being answered: this would be an in-preview Version 3 feature behind a writer feature flag, with the two-implementations-and-cross-testing bar to clear. That bar is the right one, and I do not think this case earns it on what I can show today. Happy to leave this open for the record in case others want the middle case — the selectivity numbers may save someone the measurement — but I will not push it further. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
