yihua opened a new pull request, #20103:
URL: https://github.com/apache/hudi/pull/20103

   ### Describe the issue this Pull Request addresses
   
   closes #20101
   part of #20064
   
   For every parquet base file, 
`HoodieParquetFileFormatHelper.buildImplicitSchemaChangeInfo` converts the 
whole footer schema to a Spark schema and compares it with the requested schema 
to find implicit type changes, even for a query that reads a few columns. The 
files of a table share a few schemas, so scans with many files repeat the same 
work.
   
   ### Summary and Changelog
   
   `buildImplicitSchemaChangeInfo` caches its result per JVM, keyed by the 
parquet file schema, the requested schema, the Hadoop conf values 
`ParquetToSparkSchemaConverter` reads and, on Spark 4.1+, 
`spark.sql.variant.allowReadingShredded`. The cache is bounded by weight (file 
columns plus requested leaf fields) and at 256 entries, results that relied on 
a failed Spark adapter check are not cached, and the returned type change map 
is read-only because files share it. The computation itself is unchanged. 
`TestHoodieParquetFileFormatHelper` checks that equal schemas reuse one result, 
that cached results equal the uncached computation across type, nested, 
nullability and conf changes, and that the key covers every Hadoop conf key the 
running Spark version's converter reads.
   
   ### Impact
   
   Lower per-file CPU on Spark parquet base file reads of tables without schema 
on read, most for narrow queries on wide files. No API, config or output change.
   
   ### Risk Level
   
   low. The cached value is the previous result for an equal key that covers 
all of its inputs, and a test fails if a Spark version's converter reads a 
Hadoop conf key the key misses. Worst-case memory is a few MB. Merges cleanly 
with #20077, #20079, #20090 and #20102.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to