yihua opened a new issue, #20082: URL: https://github.com/apache/hudi/issues/20082
In the Hudi Trino connector, every split with log files builds a `HoodieTableMetaClient` on the worker. It reads `hoodie.properties` and the index definitions (neither goes through the Trino file system cache), resolves the table schema from the timeline and the latest commit metadata unless column name casing resolution is on, and for tables before version 8 loads the timeline again to check log blocks against the committed instants. A query over N such file groups does this N times. Proposal: have the coordinator ship the table config, schema and committed instants on the table handle, and read file groups on workers from a table state built from it. part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
