yihua opened a new pull request, #20079: URL: https://github.com/apache/hudi/pull/20079
### Describe the issue this Pull Request addresses closes #20075 part of #20064 `HoodieSparkParquetReader` (native parquet log files and parquet log blocks) and `HoodieAvroParquetReader` built their `ParquetReader` from a path, which creates a new `Configuration` and parses the Hadoop default resources for every file; the later `withConf` replaced the configuration but not the parse. Each parquet base file read also copied the full Hadoop configuration three times. ### Summary and Changelog - Both readers are built from a `HadoopInputFile` over the storage configuration, through the protected `ParquetReader.Builder(InputFile)` constructor (present in parquet 1.12.2 through 1.17.0). They keep using the same configuration object with the per-read keys set. - `ParquetUtils.withHadoopReadOptions` gives both readers Hadoop read options and resolves `parquet.crypto.factory.class` decryption with the file path. On parquet 1.15+ (Spark 4.x), `Builder(InputFile)` alone builds plain read options that skip the decryption factory, which would break encrypted log files. - The Hadoop configuration is copied once per parquet base file: the per-read copy is a `JobConf`, and the requested schema is set on it in place (`ParquetSchemaEvolutionUtils.getHadoopConfClone` is renamed `getHadoopAttemptConf`). - `HoodieAvroReadSupport` (parquet 1.15+ entry point) no longer loads the Hadoop defaults per file to convert the requested projection. - Tests: `TestHoodieSparkParquetReader` and `TestHoodieAvroParquetReader` fail on master (two default resource loads per read) and pass with the change; encrypted-file read tests for both readers pass on Spark 3.5 and Spark 4.0. A manual `ParquetPerFileReadBenchmark` is included. This PR is independent of #20077 and can be cherry-picked on its own. ### Impact Lower executor CPU per file for MOR reads with parquet log files or blocks and for every parquet base file read through the Spark file group reader. Local benchmark, 400 small files, median per-file CPU: native parquet log file read 6.4 ms to 0.65 ms, parquet base file read 0.61 ms to 0.44 ms. No change to results or public APIs. On parquet 1.12.x (Spark 3.3/3.4) the decryption factory now receives the file path instead of null, matching Spark's own reader. ### Risk Level low. The readers see the same configuration object as before, and the in-place update applies to a configuration created for a single read. Covered by the new tests and the existing COW/MOR, schema evolution and file group reader suites. ### Documentation Update none ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
