comphead opened a new pull request, #25940:
URL: https://github.com/apache/datafusion/pull/25940

   ## Which issue does this PR close?
   
   - Closes #25938.
   
   ## Rationale for this change
   
   When a Parquet file is split into byte ranges, DataFusion reads a row group 
in the range that holds its first page, while Spark reads it in the range that 
holds its midpoint. Engines that hand Spark's splits to DataFusion, like Comet, 
therefore read row groups in different partitions than Spark planned. That 
leaves scan tasks idle (apache/datafusion-comet#3817) and reports the wrong 
split in `_metadata.file_block_*` (apache/datafusion-comet#6512). This PR adds 
an opt-in rule that matches Spark and keeps the current rule as the default.
   
   ## What changes are included in this PR?
   
   - New option `datafusion.execution.parquet.row_group_range_assignment` with 
values `start_offset` (default, current behavior) and `midpoint`. It is backed 
by the `RowGroupRangeAssignment` enum in `datafusion_common::parquet_config` 
and can be set with `SET`, as a `format.` table option, or with 
`ParquetSource::with_row_group_range_assignment`.
   - `midpoint` mirrors parquet-java's 
`ParquetMetadataConverter.filterFileMetaDataByMidpoint`: the smaller of column 
0's dictionary and data page offsets, plus half of 
`RowGroupMetaData::compressed_size()`.
   - The opener applies the setting to range pruning and to the 
`bytes_processed` credit, which must agree on which range owns a row group.
   - `RowGroupAccessPlanFilter::prune_by_range` keeps its signature and 
behavior. The new `prune_by_range_with_assignment` takes the assignment.
   - Proto: new `row_group_range_assignment` field in `ParquetOptions`. An 
empty value decodes as the default, so plans encoded before this change still 
load.
   - Docs: `configs.md`, and the `FileRange` rustdoc, which still said row 
groups are picked by midpoint.
   
   ## What is the testing strategy for this PR?
   
   - `row_group_filter.rs`: the layout from apache/datafusion-comet#3817 under 
both rules, a sweep that checks every row group lands in exactly one range for 
every split size, and the dictionary offset case where parquet-java and 
DataFusion's start offsets differ.
   - `opener/mod.rs`: `each_range_reads_the_row_groups_its_assignment_gives_it` 
checks the rows and `bytes_processed` of each range under both rules, on a 
split where the rules disagree. It fails if the byte credit ignores the setting.
   - `datafusion/core/tests/parquet/file_range_pruning.rs`: end-to-end reads of 
a split two-row-group file under both rules, including a sweep of split points.
   - Config parsing, proto round trips (including the COPY TO plan round trip), 
`information_schema.slt`, and a `repartition_scan.slt` case that reads a 
multi-row-group file split into 4 ranges under both rules.
   
   ## Are there any user-facing changes?
   
   A new config option. Nothing changes unless it is set to `midpoint`. 
`ParquetOptions` gains a public field.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to