zhuqi-lucas commented on issue #627: URL: https://github.com/apache/parquet-format/issues/627#issuecomment-5902744922
Thanks @alamb — "read-heavy" was the wrong word. The contrast I meant is not heavy versus light but **a chunk that decodes nothing versus a chunk that decodes at least one value**. Deferral removes the dictionary cost entirely for the first and cannot help the second, since by then the dictionary is genuinely needed. Not single-row reads specifically — selective scans. The dictionary is decompressed once per column chunk no matter how many rows come out of it, while data page cost scales with what is touched, so the fixed part dominates as selectivity rises. On ClickBench `hits.parquet` the dictionary is 14.7% of compressed bytes, but 63% of *decompressed* bytes when 10% of rows are decoded and 94.5% at 1%. That said, my comment above already walks this back: @etseidl pointed out the high-entropy argument applies to the data pages (bit-packed indices), not the dictionary (PLAIN raw values), and the encoders confirm it. His two existing suggestions — per-column `UNCOMPRESSED`, or turning dictionary encoding off where the dictionary dominates — cover our case today with no format change. I do not think this clears the two-implementations bar on what I can show, so it is not worth more of your time. Left open only so the selectivity numbers are on the record. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
