0lai0 commented on code in PR #6473:
URL: https://github.com/apache/datafusion-comet/pull/6473#discussion_r4149612392
##########
spark/src/test/scala/org/apache/comet/exec/CometGenerateExecSuite.scala:
##########
@@ -659,4 +659,67 @@ class CometGenerateExecSuite extends CometTestBase {
}
}
+ // The native explode slices each output batch out of the exploded child
instead of gathering
+ // it, so an exploded boolean, or a boolean field of an exploded struct,
leaves native at a
+ // non-zero bit offset. Arrow Java ignores that offset on import, so native
has to zero it at
+ // every level before export, including in a struct built over the booleans
and in the input to
+ // a Scala UDF, or they come back wrong after the first output batch.
+ // https://github.com/apache/datafusion-comet/issues/6464
+ private def withBooleanArrays(numRows: Int, arrayLength: Int)(f: => Unit):
Unit = {
+ withTempPath { dir =>
+ // One file, so a single input batch explodes into several output
batches.
+ withSQLConf(CometConf.COMET_ENABLED.key -> "false") {
+ spark
+ .range(0, numRows, 1, 1)
Review Comment:
Strict Scala warnings (Spark 3.5, JDK 17) fails to compile this line with
implicit numeric widening: `Dataset.range` widens the first three `Int`
arguments to `Long`. Could you change the fixture to
`withBooleanArrays(numRows: Long, arrayLength: Long)` and call `spark.range(0L,
numRows, 1L, 1)`?
The three call sites already pass integer literals.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]