andygrove opened a new pull request, #6510: URL: https://github.com/apache/datafusion-comet/pull/6510
## Which issue does this PR close? Closes #6505. Found by the 1.1.0 regression audit (#6399) and tracked in #6402. ## Rationale for this change #5237 made the native Parquet scan fill in the `_metadata` constant columns from each split's `PartitionedFile` instead of falling back to Spark. That's right for the per-file columns, but `file_block_start` and `file_block_length` describe the split that reads a row. When Spark splits a file, DataFusion and Spark don't agree on which split reads a row group. DataFusion keeps a row group in the split that holds its first page (`prune_by_range` in `row_group_filter.rs`), and Spark's parquet-mr reader keeps it in the split that holds its midpoint (`filterFileMetaDataByMidpoint`). Every row is still read once, but rows can report the wrong split. With 4 KB splits, a 5,000-row file returned `file_block_start` = 0 for every row on 1.1.0-rc1, where Spark returns 20480. With seven row groups, 2,987 of the 5,000 rows differed. 1.0.0 fell back to Spark for any metadata column, so this is a regression in 1.1.0. ## What changes are included in this PR? `CometScanRule` no longer counts `file_block_start` and `file_block_length` as supported constant metadata columns, so a scan that reads either one falls back to Spark, as it already does for `_metadata.row_index`. `file_path`, `file_name`, `file_size` and `file_modification_time` don't depend on the split, so they stay native. The serde still knows how to fill in the block columns, so they can come back once the native scan assigns row groups to splits the way Spark does. The compatibility guide now lists the two columns as falling back. ## How are these changes tested? A new `ParquetReadSuite` test writes a 5,000-row file and reads it with `spark.sql.files.maxPartitionBytes` = 4096, so that Spark splits it. It checks that each block column matches Spark and falls back with the expected reason, and that the per-file columns still run natively on the split file. Without the change, the test fails with a result mismatch. The existing native `_metadata` test no longer selects the two block columns. `ParquetReadV1Suite` and `CometScanRuleSuite` pass on the Spark 4.1 profile. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
