yihua opened a new pull request, #20071:
URL: https://github.com/apache/hudi/pull/20071

   ### Describe the issue this Pull Request addresses
   
   closes #20067
   part of #20064
   
   CDC queries on MERGE_ON_READ tables of table version 6 are broken. Spark and 
Flink fail on the driver with `InvalidAvroMagicException` because layout 1 
delta commit metadata (JSON) is always parsed as Avro. Flink also bounds log 
reads by the instant in the log file name, which before table version 8 is the 
file slice's base instant, so changes of later delta commits are dropped and 
replace commit before images are stale.
   
   ### Summary and Changelog
   
   `CommitMetadataSerDeV1` reads JSON commit metadata into the Avro model, and 
`HoodieCDCExtractor` reads delta commit metadata through 
`readCommitMetadataToAvro` once per instant. 
`HoodieCommitMetadata#getDependentFileSliceForFileGroupFromDeltaCommit` takes 
the Avro metadata, orders log files by log version, and fails with a clear 
message when a write stat has no log files. Flink `CdcIterators` bounds log 
reads by the CDC split instant, as Spark does (replace commits use the later of 
the replace instant and the slice's latest log instant). Tests cover both 
timeline layouts, Spark and Flink MOR CDC on table versions 6 and 10, and the 
split bounds.
   
   ### Impact
   
   CDC queries on table version 6 MOR tables work in Spark and Flink and return 
the changes of all delta commits; no change for table version 8 and above. 
Public signature change: 
`getDependentFileSliceForFileGroupFromDeltaCommit(InputStream, ...)` becomes 
`(org.apache.hudi.avro.model.HoodieCommitMetadata, ...)`; the old form only 
worked on Avro timelines. Delta commits whose metadata was rewritten by an 
upgrade or downgrade still cannot be inferred and now fail with a clear message 
instead of an NPE.
   
   ### Risk Level
   
   low. Limited to the CDC read path and to reading layout 1 metadata into the 
Avro model, which always failed before. Covered by red/green tests on table 
versions 6 and 10 plus the existing Spark and Flink CDC suites.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to