yihua opened a new pull request, #20085:
URL: https://github.com/apache/hudi/pull/20085

   ### Describe the issue this Pull Request addresses
   
   closes #20081
   part of #20064
   
   Under schema-on-read, every base file read resolves its file schema from the 
executor or task manager: `hoodie.properties`, the file's commit file and, when 
the commit has no internal schema or is not in the valid commits (every new 
file of a Flink streaming read), a `.hoodie/.schema` listing and read plus a 
full meta client build. A two-file query issues 78 `.hoodie` accesses from 
Spark tasks on table version 10. On table version 6 the Spark lookup used the 
version 2 timeline layout, never found the commit files, and returned an empty 
file schema for files older than the first schema history version. The Spark 
incremental relations also wrote schema-on-read settings and the full valid 
commits list into `sparkContext.hadoopConfiguration` on every streaming batch.
   
   Merge order: this PR textually conflicts with #20077 in 
`HoodieFileGroupReaderBasedFileFormat.scala` and with #20079 in 
`ParquetSchemaEvolutionUtils.scala`. Whichever merges second gets rebased; the 
resolution is known.
   
   ### Summary and Changelog
   
   New `InternalSchemaHistory` (hudi-common) holds the schema history 
restricted to the query's valid commits plus each commit's own `latest_schema`, 
read in parallel with the table's timeline layout and cached per commit file, 
since completed commit files never change. It is loaded once where the query is 
planned and resolves a version id without I/O, exactly like the timeline 
lookup. The Spark file group reader format, legacy relations and 
`SparkReaderContextFactory` ship it in the reader conf; 
`ParquetSchemaEvolutionUtils` and the Spark 3 and 4 legacy parquet formats 
resolve from it. Flink's `InternalSchemaManager` holds it instead of a storage 
conf and table config. Incremental relations pass schema-on-read settings as 
read options instead of writing the session conf.
   
   Tests: a Spark functional test (COW and MOR, table versions 6 and current, 
rename, type change, add column) checks results and that no Spark task touches 
`.hoodie` when reading base files; unit tests check equivalence with the 
timeline lookup, including a commit that completed before an earlier-requested 
schema change; a Flink test resolves merge schemas with `.hoodie` deleted.
   
   ### Impact
   
   Schema-on-read base file reads do no per-file metadata I/O on executors and 
task managers. The driver or job manager reads `.schema` once per query and 
each commit file once per JVM. On table version 6, files older than the first 
schema history version now get their commit's schema instead of an empty one. A 
commit file that exists but cannot be read now fails the query instead of 
silently falling back. `InternalSchemaManager`'s public constructor changes. 
MOR log blocks still resolve schemas through the meta client; that moves to the 
table state in #20078.
   
   ### Risk Level
   
   medium. It touches the schema-on-read read path of all Spark readers and 
Flink. Covered by the new tests plus the existing schema evolution suites 
(Spark DDL, time travel, streaming source, file group reader, Flink 
`ITTestSchemaEvolution`). The Spark 4 legacy formats get the same edit as Spark 
3; CI covers their compile and tests.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to