comphead opened a new issue, #25938:
URL: https://github.com/apache/datafusion/issues/25938

   ### Is your feature request related to a problem or challenge?
   
   When a Parquet file is split into byte ranges (`FileRange`), each row group 
is read by exactly one range. DataFusion gives a row group to the range that 
contains the start of its first column chunk (`row_group_in_range` in 
`datafusion/datasource-parquet/src/row_group_filter.rs`). Spark, through 
parquet-java's `ParquetMetadataConverter.filterFileMetaDataByMidpoint`, gives 
it to the range that contains its midpoint: that start plus half of the row 
group's compressed size.
   
   Both rules read every row once, but they disagree on row groups that cross a 
range boundary, and the start rule piles large row groups onto earlier ranges. 
For a ~190 MB file with row groups at `[4, ~122 MB)` and `[~122 MB, ~190 MB)` 
split at 128 MB, the first range reads both row groups and the second reads 
nothing. Spark gives each range one row group.
   
   This hurts engines that pass Spark's splits to DataFusion, such as Comet:
   
   - apache/datafusion-comet#3817: 600 of 1800 planned scan tasks read nothing.
   - apache/datafusion-comet#6512: rows are read by a different split than in 
Spark, so `_metadata.file_block_start` and `file_block_length` are wrong 
(apache/datafusion-comet#6505).
   
   The `FileRange` rustdoc still says row groups are selected by their 
"midpoint", which stopped being true after #2677 and #5997.
   
   ### Describe the solution you'd like
   
   A reader option `datafusion.execution.parquet.row_group_range_assignment`:
   
   - `start_offset` (default): the current rule.
   - `midpoint`: parquet-java's rule. The start is the smaller of column 0's 
dictionary and data page offsets, and the midpoint adds half of the row group's 
compressed size.
   
   Comet can then set `midpoint` to match Spark.
   
   ### Describe alternatives you've considered
   
   - Make `midpoint` the default, as apache/datafusion-comet#6512 suggests. It 
would also balance the ranges of DataFusion's own `repartition_file_scans` 
better. Starting with an opt-in keeps existing behavior unchanged, and the 
default can switch later.
   - Do it in Comet only. `ParquetAccessPlan` needs one entry per row group, so 
Comet would have to read every footer before planning or wrap `ParquetSource`.
   
   ### Additional context
   
   Comet's native Iceberg scan already uses the midpoint rule 
(apache/iceberg-rust#2615).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to