zhuqi-lucas opened a new issue, #25896: URL: https://github.com/apache/datafusion/issues/25896
## Is your feature request related to a problem or challenge? `EXPLAIN ANALYZE` on a parquet scan reports `time_elapsed_processing`, which covers decompression and decoding together. There is no way to tell a scan that is slow because the codec is expensive from one that is slow because there is a lot to decode, and no way to see either in production. That gap has a concrete cost. On our data servers, dictionary-page decompression turned out to be 14-18% of CPU on the skip path, for column chunks whose values were never decoded (apache/arrow-rs#11154, fixed by apache/arrow-rs#11168). Finding that needed a CPU profiler on a production pod, because `EXPLAIN ANALYZE` could not show it. Anyone hitting the same shape on their own data has no way to notice. It also blocks ranking the remaining work. Three proposals are open for the dictionary decompression cost: - apache/parquet-format#627, let dictionary pages opt out of compression the way `DataPageV2` already can - apache/arrow-rs#11303, stop copying and validating the whole dictionary when few values are referenced - writing a decompressed dictionary sidecar at ingest time Each targets a selective read that needs a few rows and still pays for the whole dictionary. None of them can be compared today, because the only measurement anyone has is a profile of a different path. ## Describe the solution you'd like Surface decompression time as a parquet scan metric, split between dictionary pages and data pages, next to the existing `time_elapsed_*` metrics in `ParquetFileMetrics`. Bytes as well as time would help: a large dictionary that decompresses quickly and a small one that does not are different problems. Most of the plumbing already exists. `ArrowReaderMetrics` is created in `datasource-parquet/src/opener/mod.rs`, attached in `push_decoder.rs`, and copied into `ParquetFileMetrics`, which is how `predicate_cache_records` and `predicate_cache_inner_records` reach `EXPLAIN ANALYZE` today. A decompression timer would follow exactly that path. ## Describe alternatives you've considered **Leave it inside `time_elapsed_processing`.** This is the status quo and it is what made the 14-18% finding invisible until someone attached a profiler. **Profile out of band.** It works, which is how we found the original problem, but it is not per query, not per column, and not something you can leave running in production to see which tables actually suffer. ## Additional context This needs a companion change in arrow-rs, and that is where the real work is. The only `decompress` call site in the whole parquet crate is in `decode_page` (`parquet/src/file/serialized_reader.rs`), which is a free function rather than a method on `SerializedPageReader`, so the metrics handle has to be threaded into it and its signature changed. The page type is already in scope there, so separating dictionary from data pages is nearly free once the handle arrives. One thing to settle with a benchmark rather than by assertion: an `Instant::now()` per page is not free when pages are small and numerous, and this is a hotter path than most instrumented code. Coarse accumulation is one answer; the other is that `ArrowReaderMetrics` is already opt-in and disabled by default, which may make the question moot for anyone who does not ask for it. Happy to do the work in both repos if this looks reasonable. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
