yihua opened a new issue, #20073:
URL: https://github.com/apache/hudi/issues/20073

   The per-file read function built by `HoodieFileGroupReaderBasedFileFormat` 
captures the whole file format and a driver-built `HoodieTableMetaClient` that 
carries a full Hadoop configuration. Every Spark task Java-deserializes both 
and re-parses the schemas, which dominates per-task CPU on scans with many 
small tasks. The function also writes read options into the live table config 
for every file, so a read option named like a table config key can change merge 
results, and with several files per task the result can depend on which file a 
task read first.
   
   Proposal: broadcast the scan state once per executor, keep only small 
handles in the read function, and merge read options into a copy of the table 
properties on the driver.
   
   part of #20064
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to