yihua opened a new issue, #20109: URL: https://github.com/apache/hudi/issues/20109
Spark reads of Hudi tables generate an `UnsafeProjection` for every file on the row-based parquet path, and the file group reader format generates its output projection per file or file slice. Spark caches the compiled class, but it regenerates the source and splits the expressions for every projection before the lookup. On wide nested schemas this is several milliseconds per file and a large share of executor CPU, although the projection is the same for most files of a scan. Proposal: reuse the generated projection across files that need the same one, keyed by everything the generation reads, without sharing a projection between iterators or threads. part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
