yihua opened a new issue, #20100:
URL: https://github.com/apache/hudi/issues/20100

   For every parquet base file, the Spark readers copy the Hadoop conf three 
times (in `SparkParquetReaderBase.read`, in `ParquetSchemaEvolutionUtils`, and 
in `TaskAttemptContextImpl`, because the copy is not a `JobConf`) and render 
the requested schema to JSON three times, although the keys they set depend 
only on the requested schema. Every task also converts the table schema to 
parquet with a new default-loaded `Configuration` for logical timestamp repair, 
even for tables without a timestamp-millis field, where repair is off. With a 
large Hadoop conf and a wide schema this is a noticeable share of the CPU of 
small or selective file reads.
   
   Proposal: prepare the requested-schema read conf once per scan, copy it per 
file only when the file needs keys of its own, and convert the table schema 
only when timestamp repair is on.
   
   part of #20064
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to