yihua opened a new issue, #20109:
URL: https://github.com/apache/hudi/issues/20109

   Spark reads of Hudi tables generate an `UnsafeProjection` for every file on 
the row-based parquet path, and the file group reader format generates its 
output projection per file or file slice. Spark caches the compiled class, but 
it regenerates the source and splits the expressions for every projection 
before the lookup. On wide nested schemas this is several milliseconds per file 
and a large share of executor CPU, although the projection is the same for most 
files of a scan.
   
   Proposal: reuse the generated projection across files that need the same 
one, keyed by everything the generation reads, without sharing a projection 
between iterators or threads.
   
   part of #20064
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to