yihua opened a new pull request, #20090:
URL: https://github.com/apache/hudi/pull/20090

   ### Describe the issue this Pull Request addresses
   
   closes #20089
   part of #20064
   
   `HoodieFileGroupReaderBasedFileFormat` writes 
`spark.sql.parquet.enableVectorizedReader` into the session conf on every scan. 
After one row-based Hudi scan (such as a scan wider than 
`spark.sql.codegen.maxFields`), every later scan in the session reads 
row-based, including plain Parquet tables: a plain Parquet read after a wide 
Hudi scan used 20 to 25% more executor CPU. The wide scans themselves also 
decode with parquet-mr, where vanilla Spark decodes them vectorized.
   
   ### Summary and Changelog
   
   The format decides vectorized decoding per scan from the scan's output 
schema, like Spark's `ParquetFileFormat`, and never writes the session conf. 
Whether to return batches comes from the plan-time 
`FileFormat.OPTION_RETURNING_BATCH`, which the Parquet and ORC reader builders 
now follow instead of rechecking the conf. `supportBatch` returns the same 
answer as before and no longer has side effects.
   
   Behavior changes: wide scans (and base-only slices of wide MOR scans) decode 
vectorized and return rows; on the row path, a file with a type change is read 
row-based with Cast, so no values change and a nested type change no longer 
fails; a scan planned for batches gets a vectorized reader even if the conf 
changes before it runs. MOR file-group merges and every existing exclusion stay 
row-based. New tests cover the session conf, per-file reader choice, conf 
changes after planning and the vector/variant exclusions.
   
   ### Impact
   
   Performance only, no API or config change. Later queries in a session keep 
their own vectorized decision. Wide scans now hold a vectorized batch 
(`spark.sql.parquet.columnarReaderBatchSize` rows across the requested columns) 
per task, as vanilla Spark does for the same schema; lowering that size or 
disabling the vectorized reader limits it.
   
   ### Risk Level
   
   Medium. More reads go through Spark's nested vectorized reader and wide 
scans use more memory per task, matching vanilla Spark. Type-changed files, 
batch output and MOR merges keep their current path. This textually conflicts 
with #20079 (Parquet reader and `ParquetSchemaEvolutionUtils` lines) and in one 
import with #20077; whichever merges second rebases. A pre-existing wrong-value 
case on narrow batch reads after a schema-on-read type change is tracked 
separately.
   
   ### Documentation Update
   
   None.
   
   ### Contributor's checklist
   
   - [ ] Read through [contributor's 
guide](https://hudi.apache.org/contribute/how-to-contribute)
   - [ ] Enough context is provided in the sections above
   - [ ] Adequate tests were added if applicable
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to