andygrove opened a new pull request, #6473:
URL: https://github.com/apache/datafusion-comet/pull/6473

   ## Which issue does this PR close?
   
   Closes #6464.
   
   ## Rationale for this change
   
   #6464 was found in 1.1.0-rc1. Since #5362 and #5667, the native explode 
slices each output batch out of the exploded child instead of gathering it with 
`take`. So the booleans in an exploded struct leave native at a non-zero bit 
offset, and Arrow Java ignores that offset when it imports the batch (#6288). 
From the second output batch on, the JVM reads those booleans and their nulls 
from the wrong bits. An exploded top-level boolean goes wrong the same way when 
a native `named_struct` wraps it or a boolean Scala UDF reads it.
   
   #6339 already fixed this on `main` by zeroing boolean offsets at every level 
before export, and the reproducer in #6464 passes on `main`. None of #6339's 
tests go through explode, though, so this PR adds that coverage. The fix for 
`branch-1.1` is #6449, the backport of #6339.
   
   ## What changes are included in this PR?
   
   Tests only. `CometGenerateExecSuite` gets seven tests, each over a single 
Parquet file so that one input batch explodes into several output batches:
   
   - `explode`, `explode_outer`, `posexplode` and `posexplode_outer` of 3000 
rows of 13 structs, each with a non-null and a nullable boolean field, at the 
default batch size. At 13 elements a row the output batches are 8190 rows, so 
the later ones start both on and off a byte boundary.
   - One row of 250 structs with `spark.comet.batchSize=100`. The row is 
unnested in one build, which is then split at rows 100 and 200.
   - A native `named_struct` over an exploded boolean.
   - A boolean Scala UDF over an exploded boolean. Spark wraps a 
primitive-argument UDF in `IF(isnull(v), NULL, ...)`, and native evaluates that 
branch on a filtered copy whenever a batch has a NULL, which would hide the 
slice. So this array has no NULLs, and a comment on the fixture says why.
   
   ## How are these changes tested?
   
   All seven pass on `main` on Spark 4.1 and on Spark 3.4 with Scala 2.12, 
along with the rest of `CometGenerateExecSuite`. To check that they catch 
#6464, I rebuilt native with `zero_offsets` limited to the top level, which is 
all that rc1's `take` covered, and with the UDF bridge inputs left 
unnormalized. All seven fail there with wrong results, as do #6339's three 
`CometCodegenSuite` bridge tests.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to