yihua opened a new pull request, #20098: URL: https://github.com/apache/hudi/pull/20098
### Describe the issue this Pull Request addresses closes #20094 part of #20064 Stacked on #20069 (`HoodieEngineContext#broadcast`, `HoodieBroadcast#destroy`), review that first. Clean, marker reconciliation, clustering and compaction planning and metadata table file group initialization capture the object running the service (with the table, write config and meta client, 80 to 150 KB per task). Marker reconciliation always runs `hoodie.finalize.write.parallelism` tasks, clustering planning reads the pending table services of the whole table per partition, listing-based rollback reloads the timeline per partition, and `WriteMarkersFactory` copies the Hadoop configuration per written file. Separately, the bootstrap source listing runs on executors with a default Hadoop configuration, ignoring the job's settings, and fails for a file system only the job configures. ### Summary and Changelog - Clean, marker reconciliation, clustering planning, compaction planning and metadata table file group initialization: task functions capture small values and a broadcast of the state they read; every broadcast is destroyed once its tasks finish. The task payload drops to about 2 to 3 KB. - Only plan strategies and generators shipped with Hudi are shared by the tasks of an executor; custom classes keep one copy per task. - Marker reconciliation runs at most one task per file or partition. Clustering planning reads pending file groups once per plan (new protected `ClusteringPlanStrategy#loadFileGroupsInPendingTableServices`). - Listing-based rollback reloads the timeline only before the table version 6 file slice lookup. `WriteMarkersFactory` reads the scheme from the meta client's raw storage. - `BootstrapUtils#getAllLeafFoldersWithFiles` lists with the source storage's configuration. - Tests: `TestTableServiceTaskPayload`, `TestListingBasedRollbackStrategy`, new cases in `TestWriteMarkersFactory` and `TestBootstrapUtils`; the payload, rollback, marker and bootstrap tests fail on the base. `TaskFileAccessRecordingFileSystem` is byte-identical to the copy in #20097. ### Impact Less deserialization per task for table service planning and cleaning, fewer marker reconciliation tasks, fewer timeline listings in MOR rollback planning, and no configuration copy per written file with timeline-server markers. The bootstrap source listing honors the job's Hadoop configuration. No public API removed. ### Risk Level medium. Table service objects are now shared by the tasks of an executor instead of copied per task. The local engines already share them across threads, the shared objects were checked for mutable state, and the existing clean, clustering, compaction, rollback, marker and bootstrap suites pass. A one-line import conflict with #20097 in `HoodieBackedTableMetadataWriter` keeps both imports. ### Documentation Update none ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
