yihua opened a new issue, #20108: URL: https://github.com/apache/hudi/issues/20108
With `hoodie.schema.on.read.enable=true`, a read that returns columnar batches decodes files written before a column type change with `HoodieVectorizedParquetRecordReader`. Its converter (`SparkInternalSchemaConverter.convertColumnVectorType`) leaves some old types unconverted (Boolean, Byte, Short, Binary, Timestamp) and ignores the result of `changePrecision`, so an overflowing decimal is returned as the rescaled old value instead of null. The same query returns different values depending on `spark.sql.parquet.enableVectorizedReader`. Repro (copy-on-write, ANSI off): ```sql set hoodie.schema.on.read.enable=true; set spark.sql.ansi.enabled=false; create table t (id int, i2d int, ts long) using hudi tblproperties (type='cow', primaryKey='id', orderingFields='ts'); insert into t values (1, 12345, 1000); alter table t alter column i2d type decimal(4,2); set spark.sql.parquet.enableVectorizedReader=true; select id, i2d from t; -- [1,123.45] set spark.sql.parquet.enableVectorizedReader=false; select id, i2d from t; -- [1,null] ``` Proposed fix: make the converter report types it cannot convert and honor `changePrecision` (null on overflow), and route such files to the row-based reader (or return false from `supportBatch` when the internal schema has such changes). Related to #20064 (found while working on #20090). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
