yihua opened a new issue, #20093:
URL: https://github.com/apache/hudi/issues/20093

   The metadata table writer generates index records for every commit with 
tasks that carry table state and repeat driver-level work. The secondary index 
update builds a file system view manager (and, without the timeline server, a 
metadata table reader) in every task and never closes them. The partition stats 
update ships the table metadata and write config per written partition, builds 
views per task, reads each written base file footer twice and runs its own 
column stats lookup, so each column stats file is read once per written 
partition. Column stats, bloom filter and record index updates capture the data 
meta client, the record tagger ships whole file slices, and the record index 
bootstrap resolves the schema per file slice, reading the timeline when the 
write config has no schema.
   
   Proposal: resolve these inputs once per commit on the driver, broadcast the 
meta client, ship only what the tasks read, and release the broadcasts after 
the metadata table write.
   
   part of #20064
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to