yihua opened a new pull request, #20098:
URL: https://github.com/apache/hudi/pull/20098

   ### Describe the issue this Pull Request addresses
   
   closes #20094
   part of #20064
   
   Stacked on #20069 (`HoodieEngineContext#broadcast`, 
`HoodieBroadcast#destroy`), review that first.
   
   Clean, marker reconciliation, clustering and compaction planning and 
metadata table file group initialization capture the object running the service 
(with the table, write config and meta client, 80 to 150 KB per task). Marker 
reconciliation always runs `hoodie.finalize.write.parallelism` tasks, 
clustering planning reads the pending table services of the whole table per 
partition, listing-based rollback reloads the timeline per partition, and 
`WriteMarkersFactory` copies the Hadoop configuration per written file. 
Separately, the bootstrap source listing runs on executors with a default 
Hadoop configuration, ignoring the job's settings, and fails for a file system 
only the job configures.
   
   ### Summary and Changelog
   
   - Clean, marker reconciliation, clustering planning, compaction planning and 
metadata table file group initialization: task functions capture small values 
and a broadcast of the state they read; every broadcast is destroyed once its 
tasks finish. The task payload drops to about 2 to 3 KB.
   - Only plan strategies and generators shipped with Hudi are shared by the 
tasks of an executor; custom classes keep one copy per task.
   - Marker reconciliation runs at most one task per file or partition. 
Clustering planning reads pending file groups once per plan (new protected 
`ClusteringPlanStrategy#loadFileGroupsInPendingTableServices`).
   - Listing-based rollback reloads the timeline only before the table version 
6 file slice lookup. `WriteMarkersFactory` reads the scheme from the meta 
client's raw storage.
   - `BootstrapUtils#getAllLeafFoldersWithFiles` lists with the source 
storage's configuration.
   - Tests: `TestTableServiceTaskPayload`, `TestListingBasedRollbackStrategy`, 
new cases in `TestWriteMarkersFactory` and `TestBootstrapUtils`; the payload, 
rollback, marker and bootstrap tests fail on the base. 
`TaskFileAccessRecordingFileSystem` is byte-identical to the copy in #20097.
   
   ### Impact
   
   Less deserialization per task for table service planning and cleaning, fewer 
marker reconciliation tasks, fewer timeline listings in MOR rollback planning, 
and no configuration copy per written file with timeline-server markers. The 
bootstrap source listing honors the job's Hadoop configuration. No public API 
removed.
   
   ### Risk Level
   
   medium. Table service objects are now shared by the tasks of an executor 
instead of copied per task. The local engines already share them across 
threads, the shared objects were checked for mutable state, and the existing 
clean, clustering, compaction, rollback, marker and bootstrap suites pass. A 
one-line import conflict with #20097 in `HoodieBackedTableMetadataWriter` keeps 
both imports.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to