yihua opened a new pull request, #20102: URL: https://github.com/apache/hudi/pull/20102
### Describe the issue this Pull Request addresses closes #20100 part of #20064 For every parquet base file, the Spark readers copy the Hadoop conf three times and render the requested schema to JSON three times, although those keys depend only on the requested schema. Every task also converts the table schema for timestamp repair even when repair is off, the Spark 3.x readers walk every footer for shredded variants, and the vectorized reader allocates a full batch even for small files. ### Summary and Changelog `SparkParquetReaderBase` prepares a `ParquetReadConf` (a `JobConf` with the requested schema keys) once per scan and requested schema when the caller passes a `SharedScanStorageConfiguration`, as the file group reader format now does for base files; other callers get one copy per file instead of three. `ParquetSchemaEvolutionUtils.getHadoopAttemptConf` (was `getHadoopConfClone`) copies the conf only when the file needs its own requested schema or a pushed filter, so the shared conf is never modified. The Spark 3.x readers skip the variant walk when no requested struct can be an unshredded variant, vectorized batches on Spark 3.5+ are sized to the split's rows, and the timestamp repair schema is converted only for tables with a timestamp-millis field. New tests check that a COW scan reads all base files with one conf and that shared and per-file confs return the same rows, including with pushed filters, type changes and 8 concurrent threads. ### Impact Lower per-file CPU on Spark parquet base file reads. No API, config or output change. ### Risk Level low. The files of a scan share a read-only `JobConf`, as vanilla Spark already does for footer reads; per-file keys go on a copy, and a concurrent test checks the shared conf stays unchanged. Merging with #20079 (conflicts in `ParquetSchemaEvolutionUtils`, `SparkParquetReaderBase` and the six readers): take this PR's `getHadoopAttemptConf` and drop #20079's in-place `val hadoopAttemptConf = readConf` (`ParquetSchemaEvolutionUtils.scala:89`). With a shared conf, that write leaks one file's pushed filter to other files and silently drops their rows. Keep `TestSparkParquetReaderBase#testAttemptConfIsTheSharedConfUnlessTheFileNeedsItsOwnKeys` and `TestSharedScanParquetReadConf` through the merge; they fail on the wrong resolution. This PR also conflicts with #20090 (`ParquetSchemaEvolutionUtils`, the readers, `TestBasicSchemaEvolution`) and #20077 (`HoodieFileGroupReaderBasedFileFormat`, whose two edits move to `HoodieFileGroupReaderFunction`); whichever merges second rebases. ### Documentation Update none ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
