yihua opened a new issue, #20083:
URL: https://github.com/apache/hudi/issues/20083

   Flink's `MergeOnReadInputFormat`, `CdcInputFormat` and FLIP-27 split reader 
functions, and Hive's file group record reader, build a meta client for every 
split. Hive also walks up the path to find the table, loads the timeline and 
reads commit metadata to get the schema, per split.
   
   Two correctness issues sit next to it: Flink lookup joins never see commits 
made after their first cache load, and Hive copies job settings into the live 
table config, so a session merge mode can override the persisted one.
   
   Proposal: capture the table state where the read is planned and read splits 
from it, plan lookup join reloads against a reloaded timeline, and keep Hive 
read options out of the persisted table config.
   
   part of #20064
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to