andygrove opened a new issue, #6505:
URL: https://github.com/apache/datafusion-comet/issues/6505
### Describe the bug
Since #5237, the native Parquet scan fills in the `_metadata` constant
columns itself instead of falling back to Spark. `file_block_start` and
`file_block_length` describe the split that read a row. When Spark splits one
file into several partitions, Comet and Spark disagree about which split reads
a given row group. DataFusion keeps a row group in the split that holds its
first page (`prune_by_range` in `row_group_filter.rs`). Spark's parquet-mr
reader keeps it in the split that holds its midpoint
(`filterFileMetaDataByMidpoint`). Every row is still read exactly once, so data
columns and the other `_metadata` fields are right, but rows report the start
and length of the wrong split.
In 1.0.0 any metadata column made the scan fall back to Spark, which
returned the right values.
### Steps to reproduce
```scala
withSQLConf(SQLConf.FILES_MAX_PARTITION_BYTES.key -> "4096") {
withTempPath { dir =>
val path = dir.getCanonicalPath
spark
.range(0, 5000)
.selectExpr("id", "concat('value_', cast(id as string)) as s")
.coalesce(1)
.write
.parquet(path)
checkSparkAnswer(
spark.read
.parquet(path)
.selectExpr("id", "_metadata.file_block_start",
"_metadata.file_block_length"))
}
}
```
The file has one row group, which Spark reads in the split that holds its
midpoint. Spark returns `file_block_start` = 20480 for all 5,000 rows, and
1.1.0-rc1 returns 0. With seven row groups (`parquet.block.size` = 16384),
2,987 of the 5,000 rows differ.
### Expected behavior
The values Spark returns, as in 1.0.0.
### Workaround
No setting covers this. Adding `_metadata.row_index` to the query makes
Comet read that scan with Spark's reader, and so does reading the values with
`input_file_block_start()` and `input_file_block_length()`. Raising
`spark.sql.files.maxPartitionBytes` and `spark.sql.files.openCostInBytes` above
the largest file also avoids it, because Spark then doesn't split files.
### Additional context
The simplest fix is to fall back to Spark when a query reads
`file_block_start` or `file_block_length`.
Found by the 1.1.0 regression audit (#6399) and tracked in #6402.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]