yihua opened a new issue, #20081:
URL: https://github.com/apache/hudi/issues/20081

   Under schema-on-read, every base file read resolves the schema the file was 
written with from the executor or task manager: it reads `hoodie.properties` 
and the file's commit file and, when the commit has no internal schema or is 
not in the valid commits (every new file of a Flink streaming read), lists and 
reads `.hoodie/.schema` and builds a full meta client. That is dozens of 
`.hoodie` requests per query from tasks.
   
   On table version 6 the Spark lookup also uses the wrong timeline layout, 
never finds the commit files, and returns an empty file schema for files older 
than the first schema history version, so a renamed or retyped column reads as 
null. Separately, the Spark incremental relations write schema-on-read settings 
and the full valid commits list into the session Hadoop configuration on every 
streaming batch.
   
   Proposal: load the schema history restricted to the query's valid commits 
once where the query is planned, ship it to the readers, and resolve file 
schemas from it without I/O.
   
   part of #20064
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to