yihua opened a new pull request, #20072: URL: https://github.com/apache/hudi/pull/20072
### Describe the issue this Pull Request addresses closes #20068 part of #20064 A Flink write task flushes one bucket at a time, and each bucket built a new meta client and revalidated the write schema (timeline listing plus a commit read), so a checkpoint flushing N buckets paid that N times per task. Compaction also re-resolved the table schema per file group, and the per-checkpoint refresh built extra meta clients or reloaded the timeline without using it. ### Summary and Changelog `HoodieFlinkWriteClient` keeps the meta client of the instant being written: the first bucket initializes the table and validates the schema, later buckets reuse the table config and reload only the timeline, and the schema is revalidated only when a commit has completed since the last check. `DataTableCompactHandler` resolves the table schema once per compaction instant. `WriteProfile`, `BootstrapOperator` and the record level index backends drop redundant meta client builds and timeline reloads. Tests count task-side `.hoodie` accesses with a recording file system (4 buckets: `hoodie.properties` reads 4 to 1, commit reads 4 to 1) and cover revalidation after a concurrent commit. ### Impact Per task, a flush of N buckets goes from about 3N `.hoodie` requests to N + 2, and compaction reads commit metadata once per instant instead of once per file group. A schema committed concurrently mid-instant is still caught before the next bucket is written. No public API or config change. ### Risk Level low. Failover creates a new write client, so cached state never outlives a task attempt, and the timeline is still reloaded per bucket. Covered by the new tests and the existing Flink write, compaction and index suites. ### Documentation Update none ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
