andygrove opened a new issue, #6251:
URL: https://github.com/apache/datafusion-comet/issues/6251
### Describe the bug
`arrays_zip` over two inputs with the same name fails the task when Comet
runs it natively. Spark names each struct field after its input, so
`arrays_zip(a, a)` has type `array<struct<a: string, a: string>>`, and Comet's
native projection returns that struct fine. Importing the result on the JVM
then fails, because Java Arrow keys struct children by name and the two `a`
children collapse into one:
```
java.lang.IllegalStateException: ArrowArray struct has 2 children (expected
1)
at
org.apache.arrow.util.Preconditions.checkState(Preconditions.java:562)
at org.apache.arrow.c.ArrayImporter.doImport(ArrayImporter.java:92)
at org.apache.arrow.c.ArrayImporter.importChild(ArrayImporter.java:83)
at org.apache.arrow.c.ArrayImporter.doImport(ArrayImporter.java:101)
at org.apache.arrow.c.ArrayImporter.importArray(ArrayImporter.java:68)
at org.apache.arrow.c.ArrowImporter.importVector(ArrowImporter.java:62)
at org.apache.comet.vector.NativeUtil.importVector(NativeUtil.scala:264)
at org.apache.comet.vector.NativeUtil.getNextBatch(NativeUtil.scala:211)
at
org.apache.comet.CometExecIterator.getNextBatch(CometExecIterator.scala:237)
```
`arrays_zip(a, b)` over the same table works. Spark returns the zipped rows
for all three queries below.
This is the same Java Arrow limitation as #1015 and #5605.
`CometCreateNamedStruct.getSupportLevel` already returns `Unsupported` when
`names` has duplicates, and `DataTypeSupport` rejects such structs in operator
schemas, but `CometArraysZip.getSupportLevel` only checks the input types and
never looks at `expr.names`. Two same-named inputs come up naturally after a
join, for example `arrays_zip(t1.tags, t2.tags)`.
### Steps to reproduce
Default configs, reproduced on `main` at `f7952de73` with Spark 4.1.3:
```scala
spark.range(4)
.selectExpr("id", "array(cast(id as string), 'x') as a", "array(id, id +
1) as b")
.write.parquet(path)
val df = spark.read.parquet(path)
df.selectExpr("id", "arrays_zip(a, a) AS r").collect() //
IllegalStateException
df.selectExpr("id", "arrays_zip(b, b) AS r").collect() //
IllegalStateException
df.selectExpr("id", "arrays_zip(a, b) AS r").collect() // matches Spark
```
The plan is `CometProject` over `CometNativeScan`.
### Expected behavior
Either match Spark or fall back. The smallest fix is probably for
`CometArraysZip.getSupportLevel` to return `Unsupported` when `expr.names` has
duplicates, the same way `CometCreateNamedStruct` does.
### Additional context
Found while reviewing #6036, but unrelated to that PR's change.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]