andygrove opened a new issue, #6251:
URL: https://github.com/apache/datafusion-comet/issues/6251

   ### Describe the bug
   
   `arrays_zip` over two inputs with the same name fails the task when Comet 
runs it natively. Spark names each struct field after its input, so 
`arrays_zip(a, a)` has type `array<struct<a: string, a: string>>`, and Comet's 
native projection returns that struct fine. Importing the result on the JVM 
then fails, because Java Arrow keys struct children by name and the two `a` 
children collapse into one:
   
   ```
   java.lang.IllegalStateException: ArrowArray struct has 2 children (expected 
1)
        at 
org.apache.arrow.util.Preconditions.checkState(Preconditions.java:562)
        at org.apache.arrow.c.ArrayImporter.doImport(ArrayImporter.java:92)
        at org.apache.arrow.c.ArrayImporter.importChild(ArrayImporter.java:83)
        at org.apache.arrow.c.ArrayImporter.doImport(ArrayImporter.java:101)
        at org.apache.arrow.c.ArrayImporter.importArray(ArrayImporter.java:68)
        at org.apache.arrow.c.ArrowImporter.importVector(ArrowImporter.java:62)
        at org.apache.comet.vector.NativeUtil.importVector(NativeUtil.scala:264)
        at org.apache.comet.vector.NativeUtil.getNextBatch(NativeUtil.scala:211)
        at 
org.apache.comet.CometExecIterator.getNextBatch(CometExecIterator.scala:237)
   ```
   
   `arrays_zip(a, b)` over the same table works. Spark returns the zipped rows 
for all three queries below.
   
   This is the same Java Arrow limitation as #1015 and #5605. 
`CometCreateNamedStruct.getSupportLevel` already returns `Unsupported` when 
`names` has duplicates, and `DataTypeSupport` rejects such structs in operator 
schemas, but `CometArraysZip.getSupportLevel` only checks the input types and 
never looks at `expr.names`. Two same-named inputs come up naturally after a 
join, for example `arrays_zip(t1.tags, t2.tags)`.
   
   ### Steps to reproduce
   
   Default configs, reproduced on `main` at `f7952de73` with Spark 4.1.3:
   
   ```scala
   spark.range(4)
     .selectExpr("id", "array(cast(id as string), 'x') as a", "array(id, id + 
1) as b")
     .write.parquet(path)
   val df = spark.read.parquet(path)
   df.selectExpr("id", "arrays_zip(a, a) AS r").collect() // 
IllegalStateException
   df.selectExpr("id", "arrays_zip(b, b) AS r").collect() // 
IllegalStateException
   df.selectExpr("id", "arrays_zip(a, b) AS r").collect() // matches Spark
   ```
   
   The plan is `CometProject` over `CometNativeScan`.
   
   ### Expected behavior
   
   Either match Spark or fall back. The smallest fix is probably for 
`CometArraysZip.getSupportLevel` to return `Unsupported` when `expr.names` has 
duplicates, the same way `CometCreateNamedStruct` does.
   
   ### Additional context
   
   Found while reviewing #6036, but unrelated to that PR's change.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to