yihua opened a new pull request, #20077: URL: https://github.com/apache/hudi/pull/20077
### Describe the issue this Pull Request addresses closes #20073 part of #20064 The Spark file group reader function captured the whole file format and a driver-built meta client with a full Hadoop configuration, so every task deserialized both and re-parsed the schemas. It also wrote read options into the live table config for every file. ### Summary and Changelog The reader function is now `HoodieFileGroupReaderFunction`, which holds only broadcast handles. The scan state (the relation's meta client with the timeline the scan was planned against, schemas, reader properties) is broadcast as Java-serialized bytes and deserialized once per executor. Read options are merged into a copy of the table properties on the driver, and per-file properties are copied before use, so nothing shared is mutated on executors. The relation factory passes its meta client to the file format instead of the format building a second one. Tests: `TestHoodieFileGroupReaderFunction` serializes the function for COW and MOR tables and asserts no meta client, file format or Hadoop configuration is reachable, within a size budget. `TestReadOptionsDoNotOverrideTableConfig` checks that read options named like table config keys change neither merge results nor the relation's table config. This PR is independent of the parquet per-file configuration change (#20075) and can be cherry-picked on its own. The table state change (#20074) stacks on it. ### Impact Per-task deserialization drops from the file format plus meta client to a few broadcast handles. Behavior change: a read option whose key is a table config key (`hoodie.record.merge.mode`, `hoodie.table.ordering.fields`, `hoodie.compaction.payload.class`, `hoodie.table.partial.update.mode`, and similar) no longer overrides the persisted table config on the read path. Before, such an option could change merge results, and with several files per task the result could depend on which file a task read first. The write-side keys (`hoodie.datasource.write.payload.class`, `hoodie.write.record.merge.mode`, `hoodie.datasource.write.precombine.field`) still reach the reader through the reader properties. ### Risk Level medium. The per-file read logic is unchanged (the same reader builders and projections move into the new class), but the payload of every Spark scan task changes, and the scan state is shared by concurrent tasks on an executor. Covered by the new tests plus the existing COW/MOR data source and file group reader suites. ### Documentation Update none. The read option behavior change goes in the release notes. ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
