yihua opened a new pull request, #20110:
URL: https://github.com/apache/hudi/pull/20110

   ### Describe the issue this Pull Request addresses
   
   closes #20109
   part of #20064
   
   The row-based parquet read path (`Spark3xParquetReader` / 
`Spark4xParquetReader` through `ParquetSchemaEvolutionUtils`) generates an 
`UnsafeProjection` for every file, and `HoodieFileGroupReaderBasedFileFormat` 
generates its output projection per file or file slice. Spark caches the 
compiled class but regenerates the source and splits the expressions every 
time. On a 100-column schema nested to depth 6 that is about 9 ms per file and 
about 25% of the executor CPU of a COW snapshot scan (JFR), although most files 
of a scan need the same projection.
   
   ### Summary and Changelog
   
   New `UnsafeProjectionPool` (hudi-spark-client) keeps up to 16 idle 
projections per executor thread, keyed by every input of the generation: 
requested and partition schemas, the file's type changes, the session time zone 
and, when there is a type change, the task's registered SQL configs (task local 
properties such as `spark.sql.execution.id` are left out, so reuse works across 
queries). A lease takes the projection when the first row needs it and returns 
it only once its iterator has no more rows, so concurrent iterators and threads 
never share one. A projection that wrote a row over 1 MB is dropped instead of 
pooled, and idle projections are held through soft references. The Spark 3.3 to 
4.2 parquet readers and the file group reader format's output and `projectIter` 
projections lease instead of generating; the projection is generated lazily, so 
a file read as batches generates none.
   
   Tests compare reused and freshly generated projections over type-changed 
files, schema on read, partition columns, pushed filters and MOR log files, 
hold uncopied rows of interleaved iterators through the real reader, and run 
union, joins, cache, sorts and limits over reused rows against runs without the 
pool.
   
   ### Impact
   
   Lower executor CPU for row-based scans: the per-file projection generation 
(about 7.6 ms per file in a micro-benchmark, about 9 ms in JFR on the wide 
schema) is paid once per thread and schema. No API, config or output change.
   
   ### Risk Level
   
   low. The reused projection is what the same inputs would generate, and rows 
keep the existing contract (valid until the next `hasNext`/`next` on their 
iterator). Memory per thread is bounded as above.
   
   Conflicts with open PRs, whichever merges second rebases:
   - #20090: in `ParquetSchemaEvolutionUtils` keep both sides (its 
`canReadVectorized`/`hasTypeChange`, then `leaseRowProjection`); in 
`HoodieFileGroupReaderBasedFileFormat` lease the output projection and pass 
`lease.asProjection` as the by-name projection of 
`FileGroupOutputProjection.create`, releasing the lease when the reader is 
exhausted.
   - #20077 (and #20078 on it) moves `projectIter`, `projectSchema` and 
`appendPartitionAndProject` into `HoodieFileGroupReaderFunction`; the leases 
move with them. #20085 conflicts on one import line.
   - After #20079 or #20102, the `[[getHadoopConfClone]]` scaladoc reference in 
`ParquetSchemaEvolutionUtils` must become `[[getHadoopAttemptConf]]`.
   
   ### Documentation Update
   
   none
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to