yihua opened a new issue, #20101: URL: https://github.com/apache/hudi/issues/20101
For every parquet base file, `HoodieParquetFileFormatHelper.buildImplicitSchemaChangeInfo` converts the whole footer schema to a Spark schema and compares it with the requested schema, matching nested fields by name at every level, to find implicit type changes. The conversion covers every column of the file even when the query reads a few. The files of a table share a few schemas, so scans over many files repeat the same work. Proposal: cache the result per executor, keyed by everything the computation reads (the file schema, the requested schema and the relevant conf values). part of #20064 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
