yihua opened a new issue, #20108:
URL: https://github.com/apache/hudi/issues/20108

   With `hoodie.schema.on.read.enable=true`, a read that returns columnar 
batches decodes files written before a column type change with 
`HoodieVectorizedParquetRecordReader`. Its converter 
(`SparkInternalSchemaConverter.convertColumnVectorType`) leaves some old types 
unconverted (Boolean, Byte, Short, Binary, Timestamp) and ignores the result of 
`changePrecision`, so an overflowing decimal is returned as the rescaled old 
value instead of null. The same query returns different values depending on 
`spark.sql.parquet.enableVectorizedReader`.
   
   Repro (copy-on-write, ANSI off):
   ```sql
   set hoodie.schema.on.read.enable=true; set spark.sql.ansi.enabled=false;
   create table t (id int, i2d int, ts long) using hudi tblproperties 
(type='cow', primaryKey='id', orderingFields='ts');
   insert into t values (1, 12345, 1000);
   alter table t alter column i2d type decimal(4,2);
   set spark.sql.parquet.enableVectorizedReader=true;  select id, i2d from t;  
-- [1,123.45]
   set spark.sql.parquet.enableVectorizedReader=false; select id, i2d from t;  
-- [1,null]
   ```
   
   Proposed fix: make the converter report types it cannot convert and honor 
`changePrecision` (null on overflow), and route such files to the row-based 
reader (or return false from `supportBatch` when the internal schema has such 
changes).
   
   Related to #20064 (found while working on #20090).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to