zhuqi-lucas opened a new issue, #25896:
URL: https://github.com/apache/datafusion/issues/25896

   ## Is your feature request related to a problem or challenge?
   
   `EXPLAIN ANALYZE` on a parquet scan reports `time_elapsed_processing`, which 
covers decompression and decoding together. There is no way to tell a scan that 
is slow because the codec is expensive from one that is slow because there is a 
lot to decode, and no way to see either in production.
   
   That gap has a concrete cost. On our data servers, dictionary-page 
decompression turned out to be 14-18% of CPU on the skip path, for column 
chunks whose values were never decoded (apache/arrow-rs#11154, fixed by 
apache/arrow-rs#11168). Finding that needed a CPU profiler on a production pod, 
because `EXPLAIN ANALYZE` could not show it. Anyone hitting the same shape on 
their own data has no way to notice.
   
   It also blocks ranking the remaining work. Three proposals are open for the 
dictionary decompression cost:
   
   - apache/parquet-format#627, let dictionary pages opt out of compression the 
way `DataPageV2` already can
   - apache/arrow-rs#11303, stop copying and validating the whole dictionary 
when few values are referenced
   - writing a decompressed dictionary sidecar at ingest time
   
   Each targets a selective read that needs a few rows and still pays for the 
whole dictionary. None of them can be compared today, because the only 
measurement anyone has is a profile of a different path.
   
   ## Describe the solution you'd like
   
   Surface decompression time as a parquet scan metric, split between 
dictionary pages and data pages, next to the existing `time_elapsed_*` metrics 
in `ParquetFileMetrics`. Bytes as well as time would help: a large dictionary 
that decompresses quickly and a small one that does not are different problems.
   
   Most of the plumbing already exists. `ArrowReaderMetrics` is created in 
`datasource-parquet/src/opener/mod.rs`, attached in `push_decoder.rs`, and 
copied into `ParquetFileMetrics`, which is how `predicate_cache_records` and 
`predicate_cache_inner_records` reach `EXPLAIN ANALYZE` today. A decompression 
timer would follow exactly that path.
   
   ## Describe alternatives you've considered
   
   **Leave it inside `time_elapsed_processing`.** This is the status quo and it 
is what made the 14-18% finding invisible until someone attached a profiler.
   
   **Profile out of band.** It works, which is how we found the original 
problem, but it is not per query, not per column, and not something you can 
leave running in production to see which tables actually suffer.
   
   ## Additional context
   
   This needs a companion change in arrow-rs, and that is where the real work 
is. The only `decompress` call site in the whole parquet crate is in 
`decode_page` (`parquet/src/file/serialized_reader.rs`), which is a free 
function rather than a method on `SerializedPageReader`, so the metrics handle 
has to be threaded into it and its signature changed. The page type is already 
in scope there, so separating dictionary from data pages is nearly free once 
the handle arrives.
   
   One thing to settle with a benchmark rather than by assertion: an 
`Instant::now()` per page is not free when pages are small and numerous, and 
this is a hotter path than most instrumented code. Coarse accumulation is one 
answer; the other is that `ArrowReaderMetrics` is already opt-in and disabled 
by default, which may make the question moot for anyone who does not ask for it.
   
   Happy to do the work in both repos if this looks reasonable.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to