yihua opened a new pull request, #20078: URL: https://github.com/apache/hudi/pull/20078
### Describe the issue this Pull Request addresses closes #20074 part of #20064 Stacked on #20077, review that first. Until it merges this PR shows its commit plus one commit of this change. The file group reader and the log and buffer code it reaches take a `HoodieTableMetaClient` but use only the table config, the base path, a storage factory and, in two places, the timeline. Those two places (the committed-instant check on log blocks of tables before version 8, and the schema-history lookup for schema-on-read) load the timeline on the executor, so every task repeats that I/O and tasks of one query can see different timelines. ### Summary and Changelog New `FileGroupReaderTableState` holds the table config and base path, plus `CommittedInstants` and the schema history captured where the timeline is already loaded (`snapshotOf`). `fromMetaClient` keeps the lazy meta client behavior for existing `withHoodieTableMetaClient` callers. The file group reader, record buffers, buffer loaders, log record readers, LSM reader and Spark CDC iterator take the table state; the Spark file format and CDC path build it once on the driver. `InternalSchema` and `Types.RecordType` lazy lookup maps are published through volatile fields, since one query schema is shared by concurrent tasks. Public constructors that took a meta client keep deprecated bridges. Tests: `TestFileGroupReaderTableState` checks `CommittedInstants` against timeline semantics on layout v1 and v2 instants and serialization of a captured state. `TestMergeOnReadSkipsUncommittedLogBlocks` reads a version 6 MOR table and checks that log blocks of an inflight delta commit older than the latest completed one are skipped. ### Impact No executor timeline loads on the Spark read path for tables before version 8 or for schema-on-read. Behavior change for tables before version 8: the commit check on log blocks uses the committed instants of the timeline the scan was planned against, not a timeline loaded per task, so every task of a scan sees one snapshot. Before, if an inflight instant older than the planned one completed while the scan ran, late tasks included its log blocks and earlier tasks did not. The native log reader now gets a copy of the table properties, so read options no longer reach native delete block ordering values. SPI changes without a bridge: `FileGroupRecordBufferLoader.getRecordBuffer` takes a `FileGroupReaderTableState` instead of a meta client, and `BaseHoodieLogRecordReader`'s protected constructor takes the table state and its protected `hoodieTableMetaClient` field is removed. ### Risk Level medium. `CommittedInstants.isCommitted` mirrors the timeline check it replaces (inflight exclusion, completed set, second-granularity fallback, archival boundary), and meta client callers keep the old lazy loading. The SPI changes only affect external implementations of the buffer loader or subclasses of the log record reader; none exist in the repo outside `hudi-common`. ### Documentation Update none. The SPI changes and the snapshot behavior for tables before version 8 go in the release notes. ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
