yihua opened a new pull request, #20078:
URL: https://github.com/apache/hudi/pull/20078

   ### Describe the issue this Pull Request addresses
   
   closes #20074
   part of #20064
   
   Stacked on #20077, review that first. Until it merges this PR shows its 
commit plus one commit of this change.
   
   The file group reader and the log and buffer code it reaches take a 
`HoodieTableMetaClient` but use only the table config, the base path, a storage 
factory and, in two places, the timeline. Those two places (the 
committed-instant check on log blocks of tables before version 8, and the 
schema-history lookup for schema-on-read) load the timeline on the executor, so 
every task repeats that I/O and tasks of one query can see different timelines.
   
   ### Summary and Changelog
   
   New `FileGroupReaderTableState` holds the table config and base path, plus 
`CommittedInstants` and the schema history captured where the timeline is 
already loaded (`snapshotOf`). `fromMetaClient` keeps the lazy meta client 
behavior for existing `withHoodieTableMetaClient` callers. The file group 
reader, record buffers, buffer loaders, log record readers, LSM reader and 
Spark CDC iterator take the table state; the Spark file format and CDC path 
build it once on the driver. `InternalSchema` and `Types.RecordType` lazy 
lookup maps are published through volatile fields, since one query schema is 
shared by concurrent tasks. Public constructors that took a meta client keep 
deprecated bridges.
   
   Tests: `TestFileGroupReaderTableState` checks `CommittedInstants` against 
timeline semantics on layout v1 and v2 instants and serialization of a captured 
state. `TestMergeOnReadSkipsUncommittedLogBlocks` reads a version 6 MOR table 
and checks that log blocks of an inflight delta commit older than the latest 
completed one are skipped.
   
   ### Impact
   
   No executor timeline loads on the Spark read path for tables before version 
8 or for schema-on-read.
   
   Behavior change for tables before version 8: the commit check on log blocks 
uses the committed instants of the timeline the scan was planned against, not a 
timeline loaded per task, so every task of a scan sees one snapshot. Before, if 
an inflight instant older than the planned one completed while the scan ran, 
late tasks included its log blocks and earlier tasks did not. The native log 
reader now gets a copy of the table properties, so read options no longer reach 
native delete block ordering values.
   
   SPI changes without a bridge: `FileGroupRecordBufferLoader.getRecordBuffer` 
takes a `FileGroupReaderTableState` instead of a meta client, and 
`BaseHoodieLogRecordReader`'s protected constructor takes the table state and 
its protected `hoodieTableMetaClient` field is removed.
   
   ### Risk Level
   
   medium. `CommittedInstants.isCommitted` mirrors the timeline check it 
replaces (inflight exclusion, completed set, second-granularity fallback, 
archival boundary), and meta client callers keep the old lazy loading. The SPI 
changes only affect external implementations of the buffer loader or subclasses 
of the log record reader; none exist in the repo outside `hudi-common`.
   
   ### Documentation Update
   
   none. The SPI changes and the snapshot behavior for tables before version 8 
go in the release notes.
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to