andygrove opened a new issue, #6506:
URL: https://github.com/apache/datafusion-comet/issues/6506
### Describe the bug
Since #5681, the native Parquet scan checks the type conversion for every
nested field when it opens a file, while Spark checks it only when it decodes a
value. So a file with nothing to decode, either empty or with every row group
pruned, whose nested field has a type that the read schema can't convert, now
fails the query with `SchemaColumnConvertNotSupportedException`. Spark returns
the rows from the other files, and so did 1.0.0.
When the file does have rows to decode, 1.1.0 now fails the way Spark does.
That part is a fix: 1.0.0 returned null for the value it couldn't convert.
### Steps to reproduce
```scala
withTempPath { dir =>
val path = dir.getCanonicalPath
spark.sql("select named_struct('x', 1) as s").write.parquet(path)
// An empty file whose nested field is a string
spark.sql("select named_struct('x', 'a') as s where
false").write.mode("append").parquet(path)
checkSparkAnswer(spark.read.schema("s struct<x:int>").parquet(path))
}
```
Spark and 1.0.0 return `[[1]]`. 1.1.0-rc1 fails when it opens the empty
file, for column `[s, x]`, BINARY to int. A file whose only row group a filter
prunes fails the same way, for example a nested `decimal(10,4)` read as
`decimal(10,2)` with `WHERE id = 100`.
### Expected behavior
The rows Spark returns.
### Additional context
Found by the 1.1.0 regression audit (#6399) and tracked in #6402.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]