yihua opened a new issue, #20091:
URL: https://github.com/apache/hudi/issues/20091

   Index lookups on the write path run task functions that capture the whole 
`HoodieTable` and the index with its write config, so each Spark task 
deserializes them and rebuilds the table's transient state. The simple and 
global simple indexes ship the table in one task per base file only to read 
record keys from that file. 
`HoodieIndexUtils#getLatestBaseFilesForAllPartitions` runs one task per 
partition, each building its own file system view. The simple bucket index 
calls `reloadActiveTimeline()` in tasks for every partition, which lists the 
timeline on executors and lets tasks disagree on pending instants. Global index 
partition updates fetch every merged file slice of a partition per file group, 
and the consistent bucket row writer builds two listing-based views per task.
   
   Proposal: capture only the small values these tasks read, share the rest 
through the engine context broadcast, list base files with one metadata table 
lookup on the driver, and tag the bucket index from the timeline the write 
started with.
   
   part of #20064
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to