andygrove opened a new issue, #6464:
URL: https://github.com/apache/datafusion-comet/issues/6464

   ### Describe the bug
   
   In 1.1.0-rc1, `explode`, `posexplode` and `explode_outer` of an array of 
structs with a boolean field return wrong booleans and misplaced NULLs once an 
input batch explodes to more than `spark.comet.batchSize` rows. This happens 
with default configs. 1.0.0 returns the right answer for the same queries.
   
   #5362 split the native explode output into chunks of at most 
`spark.comet.batchSize` rows, and #5667 slices each chunk out of the child 
arrays instead of gathering it with `take`. A sliced struct keeps a bit offset 
on its boolean children. Arrow Java ignores that offset when it imports the 
batch (#6288), so the JVM reads those booleans from the start of the buffer. An 
exploded top-level boolean goes wrong the same way when a native `named_struct` 
wraps it or a boolean Scala UDF reads it.
   
   ### Steps to reproduce
   
   ```scala
   withTempPath { dir =>
     val path = dir.getCanonicalPath
     withSQLConf(CometConf.COMET_ENABLED.key -> "false") {
       spark
         .range(0, 3000, 1, 1)
         .selectExpr(
           "id",
           "transform(sequence(0, 11), i -> named_struct(" +
             "'b', hash(id, i) % 2 = 0, " +
             "'bn', IF(hash(id, i, 7) % 5 = 0, NULL, hash(id, i, 3) % 2 = 0), " 
+
             "'n', id * 12 + i)) AS arr")
         .write
         .parquet(path)
     }
     spark.read.parquet(path).createOrReplaceTempView("t")
     checkSparkAnswer(sql("SELECT id, s FROM t LATERAL VIEW explode(arr) x AS 
s"))
   }
   ```
   
   The native explode runs at both 1.0.0 and rc1. On rc1, 22,824 of the 36,000 
rows differ from Spark, starting at output row 8,184, the first row of the 
second chunk. `n` is right everywhere, and `b` and `bn` are wrong.
   
   ### Expected behavior
   
   The same rows as Spark, as in 1.0.0.
   
   ### Additional context
   
   #6339 fixes #6288 on `main` by zeroing boolean offsets at every level before 
export, and its `branch-1.1` backport, #6449, fixes all of these queries on 
rc1. This was found by the 1.1.0 regression audit in #6399 and is tracked in 
#6402.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to