andygrove opened a new issue, #6505:
URL: https://github.com/apache/datafusion-comet/issues/6505

   ### Describe the bug
   
   Since #5237, the native Parquet scan fills in the `_metadata` constant 
columns itself instead of falling back to Spark. `file_block_start` and 
`file_block_length` describe the split that read a row. When Spark splits one 
file into several partitions, Comet and Spark disagree about which split reads 
a given row group. DataFusion keeps a row group in the split that holds its 
first page (`prune_by_range` in `row_group_filter.rs`). Spark's parquet-mr 
reader keeps it in the split that holds its midpoint 
(`filterFileMetaDataByMidpoint`). Every row is still read exactly once, so data 
columns and the other `_metadata` fields are right, but rows report the start 
and length of the wrong split.
   
   In 1.0.0 any metadata column made the scan fall back to Spark, which 
returned the right values.
   
   ### Steps to reproduce
   
   ```scala
   withSQLConf(SQLConf.FILES_MAX_PARTITION_BYTES.key -> "4096") {
     withTempPath { dir =>
       val path = dir.getCanonicalPath
       spark
         .range(0, 5000)
         .selectExpr("id", "concat('value_', cast(id as string)) as s")
         .coalesce(1)
         .write
         .parquet(path)
       checkSparkAnswer(
         spark.read
           .parquet(path)
           .selectExpr("id", "_metadata.file_block_start", 
"_metadata.file_block_length"))
     }
   }
   ```
   
   The file has one row group, which Spark reads in the split that holds its 
midpoint. Spark returns `file_block_start` = 20480 for all 5,000 rows, and 
1.1.0-rc1 returns 0. With seven row groups (`parquet.block.size` = 16384), 
2,987 of the 5,000 rows differ.
   
   ### Expected behavior
   
   The values Spark returns, as in 1.0.0.
   
   ### Workaround
   
   No setting covers this. Adding `_metadata.row_index` to the query makes 
Comet read that scan with Spark's reader, and so does reading the values with 
`input_file_block_start()` and `input_file_block_length()`. Raising 
`spark.sql.files.maxPartitionBytes` and `spark.sql.files.openCostInBytes` above 
the largest file also avoids it, because Spark then doesn't split files.
   
   ### Additional context
   
   The simplest fix is to fall back to Spark when a query reads 
`file_block_start` or `file_block_length`.
   
   Found by the 1.1.0 regression audit (#6399) and tracked in #6402.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to