zhuqi-lucas commented on issue #627:
URL: https://github.com/apache/parquet-format/issues/627#issuecomment-5893156840

   Thanks — that is a much better read of the history than I had, and checking 
the encoders shows you are right in a way I had backwards. The dictionary page 
is `PlainEncoder` over the deduplicated values; the data pages are `RleEncoder` 
over bit-packed indices. The high-entropy argument applies to the data pages, 
not the dictionary — a dictionary of strings has plenty of shared structure for 
a codec to work with.
   
   So I measured, since my premise deserved checking. On ClickBench 
`hits.parquet` (105 columns, 226 row groups, SNAPPY), dictionary pages are 
**14.7% of all compressed bytes**, and for high-cardinality columns they 
dominate their own chunk: `UserID` 54.5%, `FUniqID` 53.6%, `RefererHash` 45.7%, 
`WatchID` 27.5%. Leaving those uncompressed is not a cheap trade.
   
   The one angle that survives the measurement is that a dictionary is 
decompressed once per column chunk regardless of how many rows are read, while 
data page cost scales with what is actually touched. Its share of 
*decompressed* bytes therefore climbs steeply as a scan gets selective:
   
   | rows actually decoded | dictionary share of decompressed bytes |
   | --- | --- |
   | 100% | 14.7% |
   | 10% | 63.4% |
   | 1% | 94.5% |
   
   That is the shape of the workload behind the 13.6–17.9% profile figure.
   
   But your two suggestions cover it without a format change, and they cover it 
on the first scan rather than only on repeats. Where the dictionary dominates, 
those are the high-cardinality columns where dictionary encoding is barely 
earning its place to begin with, so turning it off per column is a better 
answer than making it uncompressed. For the rest, per-column `UNCOMPRESSED` is 
one line of writer config. We control our writers, so that is the direction we 
will take.
   
   That leaves the format change wanted only where someone needs compressed 
data pages *and* an uncompressed dictionary in the same column. Narrower than I 
framed it. And thank you for the versioning pointer — reading that vote, the 
gating question I raised in the issue is already being answered: this would be 
an in-preview Version 3 feature behind a writer feature flag, with the 
two-implementations-and-cross-testing bar to clear. That bar is the right one, 
and I do not think this case earns it on what I can show today.
   
   Happy to leave this open for the record in case others want the middle case 
— the selectivity numbers may save someone the measurement — but I will not 
push it further.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to