yihua opened a new issue, #20073: URL: https://github.com/apache/hudi/issues/20073
The per-file read function built by `HoodieFileGroupReaderBasedFileFormat` captures the whole file format and a driver-built `HoodieTableMetaClient` that carries a full Hadoop configuration. Every Spark task Java-deserializes both and re-parses the schemas, which dominates per-task CPU on scans with many small tasks. The function also writes read options into the live table config for every file, so a read option named like a table config key can change merge results, and with several files per task the result can depend on which file a task read first. Proposal: broadcast the scan state once per executor, keep only small handles in the read function, and merge read options into a copy of the table properties on the driver. part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
