yihua opened a new pull request, #20110: URL: https://github.com/apache/hudi/pull/20110
### Describe the issue this Pull Request addresses closes #20109 part of #20064 The row-based parquet read path (`Spark3xParquetReader` / `Spark4xParquetReader` through `ParquetSchemaEvolutionUtils`) generates an `UnsafeProjection` for every file, and `HoodieFileGroupReaderBasedFileFormat` generates its output projection per file or file slice. Spark caches the compiled class but regenerates the source and splits the expressions every time. On a 100-column schema nested to depth 6 that is about 9 ms per file and about 25% of the executor CPU of a COW snapshot scan (JFR), although most files of a scan need the same projection. ### Summary and Changelog New `UnsafeProjectionPool` (hudi-spark-client) keeps up to 16 idle projections per executor thread, keyed by every input of the generation: requested and partition schemas, the file's type changes, the session time zone and, when there is a type change, the task's registered SQL configs (task local properties such as `spark.sql.execution.id` are left out, so reuse works across queries). A lease takes the projection when the first row needs it and returns it only once its iterator has no more rows, so concurrent iterators and threads never share one. A projection that wrote a row over 1 MB is dropped instead of pooled, and idle projections are held through soft references. The Spark 3.3 to 4.2 parquet readers and the file group reader format's output and `projectIter` projections lease instead of generating; the projection is generated lazily, so a file read as batches generates none. Tests compare reused and freshly generated projections over type-changed files, schema on read, partition columns, pushed filters and MOR log files, hold uncopied rows of interleaved iterators through the real reader, and run union, joins, cache, sorts and limits over reused rows against runs without the pool. ### Impact Lower executor CPU for row-based scans: the per-file projection generation (about 7.6 ms per file in a micro-benchmark, about 9 ms in JFR on the wide schema) is paid once per thread and schema. No API, config or output change. ### Risk Level low. The reused projection is what the same inputs would generate, and rows keep the existing contract (valid until the next `hasNext`/`next` on their iterator). Memory per thread is bounded as above. Conflicts with open PRs, whichever merges second rebases: - #20090: in `ParquetSchemaEvolutionUtils` keep both sides (its `canReadVectorized`/`hasTypeChange`, then `leaseRowProjection`); in `HoodieFileGroupReaderBasedFileFormat` lease the output projection and pass `lease.asProjection` as the by-name projection of `FileGroupOutputProjection.create`, releasing the lease when the reader is exhausted. - #20077 (and #20078 on it) moves `projectIter`, `projectSchema` and `appendPartitionAndProject` into `HoodieFileGroupReaderFunction`; the leases move with them. #20085 conflicts on one import line. - After #20079 or #20102, the `[[getHadoopConfClone]]` scaladoc reference in `ParquetSchemaEvolutionUtils` must become `[[getHadoopAttemptConf]]`. ### Documentation Update none ### Contributor's checklist - [ ] Read through [contributor's guide](https://hudi.apache.org/contribute/how-to-contribute) - [ ] Enough context is provided in the sections above - [ ] Adequate tests were added if applicable -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
