yihua opened a new pull request, #20079:
URL: https://github.com/apache/hudi/pull/20079

   ### Describe the issue this Pull Request addresses
   
   closes #20075
   part of #20064
   
   `HoodieSparkParquetReader` (native parquet log files and parquet log blocks) 
and `HoodieAvroParquetReader` built their `ParquetReader` from a path, which 
creates a new `Configuration` and parses the Hadoop default resources for every 
file; the later `withConf` replaced the configuration but not the parse. Each 
parquet base file read also copied the full Hadoop configuration three times.
   
   ### Summary and Changelog
   
   - Both readers are built from a `HadoopInputFile` over the storage 
configuration, through the protected `ParquetReader.Builder(InputFile)` 
constructor (present in parquet 1.12.2 through 1.17.0). They keep using the 
same configuration object with the per-read keys set.
   - `ParquetUtils.withHadoopReadOptions` gives both readers Hadoop read 
options and resolves `parquet.crypto.factory.class` decryption with the file 
path. On parquet 1.15+ (Spark 4.x), `Builder(InputFile)` alone builds plain 
read options that skip the decryption factory, which would break encrypted log 
files.
   - The Hadoop configuration is copied once per parquet base file: the 
per-read copy is a `JobConf`, and the requested schema is set on it in place 
(`ParquetSchemaEvolutionUtils.getHadoopConfClone` is renamed 
`getHadoopAttemptConf`).
   - `HoodieAvroReadSupport` (parquet 1.15+ entry point) no longer loads the 
Hadoop defaults per file to convert the requested projection.
   - Tests: `TestHoodieSparkParquetReader` and `TestHoodieAvroParquetReader` 
fail on master (two default resource loads per read) and pass with the change; 
encrypted-file read tests for both readers pass on Spark 3.5 and Spark 4.0. A 
manual `ParquetPerFileReadBenchmark` is included.
   
   This PR is independent of #20077 and can be cherry-picked on its own.
   
   ### Impact
   
   Lower executor CPU per file for MOR reads with parquet log files or blocks 
and for every parquet base file read through the Spark file group reader. Local 
benchmark, 400 small files, median per-file CPU: native parquet log file read 
6.4 ms to 0.65 ms, parquet base file read 0.61 ms to 0.44 ms. No change to 
results or public APIs. On parquet 1.12.x (Spark 3.3/3.4) the decryption 
factory now receives the file path instead of null, matching Spark's own reader.
   
   ### Risk Level
   
   low. The readers see the same configuration object as before, and the 
in-place update applies to a configuration created for a single read. Covered 
by the new tests and the existing COW/MOR, schema evolution and file group 
reader suites.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to