yihua opened a new issue, #20068: URL: https://github.com/apache/hudi/issues/20068
A Flink write task flushes its buffer one bucket at a time, and each bucket builds a new meta client (reading `hoodie.properties`) and validates the write schema (listing the timeline and reading the latest commit). A checkpoint that flushes N buckets pays this N times per task, and on object stores each is a request. The compaction task also re-resolves the table schema for every file group, and the per-checkpoint refresh of the bucket assigner, index bootstrap and record level index builds extra meta clients or lists the timeline without using it. Proposal: initialize the table once per instant on write tasks and drop the redundant reloads. part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
