yihua opened a new issue, #20081: URL: https://github.com/apache/hudi/issues/20081
Under schema-on-read, every base file read resolves the schema the file was written with from the executor or task manager: it reads `hoodie.properties` and the file's commit file and, when the commit has no internal schema or is not in the valid commits (every new file of a Flink streaming read), lists and reads `.hoodie/.schema` and builds a full meta client. That is dozens of `.hoodie` requests per query from tasks. On table version 6 the Spark lookup also uses the wrong timeline layout, never finds the commit files, and returns an empty file schema for files older than the first schema history version, so a renamed or retyped column reads as null. Separately, the Spark incremental relations write schema-on-read settings and the full valid commits list into the session Hadoop configuration on every streaming batch. Proposal: load the schema history restricted to the query's valid commits once where the query is planned, ship it to the readers, and resolve file schemas from it without I/O. part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
