yihua opened a new pull request, #20086:
URL: https://github.com/apache/hudi/pull/20086

   ### Describe the issue this Pull Request addresses
   
   closes #20082
   part of #20064
   
   Stacked on #20078 (and #20077), review those first. Until they merge this PR 
shows their commits plus one commit of this change.
   
   In the Hudi Trino connector, every split with log files built a 
`HoodieTableMetaClient` on the worker. That read `hoodie.properties` and the 
index definitions (neither goes through the Trino file system cache), resolved 
the table schema from the timeline and the latest commit metadata unless column 
name casing resolution was on, and for tables before version 8 loaded the 
timeline again to check log blocks against the committed instants. A query over 
N such file groups did this N times.
   
   ### Summary and Changelog
   
   Workers read file groups from table state that the coordinator ships on the 
table handle, so they no longer read table metadata from storage. 
`HudiTableHandle` carries the table config (without the create schema) and, for 
merge-on-read tables before version 8, the committed instants up to the query's 
latest commit (new `HudiCommittedInstants`). `HudiMetadata` ships the table 
schema for every merge-on-read handle. `HudiPageSourceProvider` builds a 
`FileGroupReaderTableState` from the handle and passes it with the split's 
storage to the file group reader.
   
   Tests: `TestHudiWorkerTableMetadataAccess` asserts through file system 
tracing spans that workers read log files but nothing under `.hoodie`, for 
table versions 6 and 8 (it fails on the parent commit), and checks what the 
coordinator puts on the handle. `TestHudiTableHandle` covers the JSON round 
trip and the bounded capture of committed instants.
   
   ### Impact
   
   No worker-side `.hoodie` I/O for splits with log files, down from 3 to 4 
metadata reads per split. For merge-on-read handles the coordinator resolves 
the table schema once per query, even when no split has log files. The handle 
grows to about 2.5 to 3.2 KB for typical merge-on-read tables (about 9 KB for a 
60-column table), plus 20 to 40 bytes per active commit instant for version 6 
and 7 merge-on-read tables.
   
   ### Risk Level
   
   low. The read path gets the same table config, schema and committed-instant 
semantics as before, captured once on the coordinator. The full hudi-trino 
suite passes, including the version 6 and 8 merge-on-read smoke and 
merge-semantics tests.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to