comphead opened a new pull request, #25940: URL: https://github.com/apache/datafusion/pull/25940
## Which issue does this PR close? - Closes #25938. ## Rationale for this change When a Parquet file is split into byte ranges, DataFusion reads a row group in the range that holds its first page, while Spark reads it in the range that holds its midpoint. Engines that hand Spark's splits to DataFusion, like Comet, therefore read row groups in different partitions than Spark planned. That leaves scan tasks idle (apache/datafusion-comet#3817) and reports the wrong split in `_metadata.file_block_*` (apache/datafusion-comet#6512). This PR adds an opt-in rule that matches Spark and keeps the current rule as the default. ## What changes are included in this PR? - New option `datafusion.execution.parquet.row_group_range_assignment` with values `start_offset` (default, current behavior) and `midpoint`. It is backed by the `RowGroupRangeAssignment` enum in `datafusion_common::parquet_config` and can be set with `SET`, as a `format.` table option, or with `ParquetSource::with_row_group_range_assignment`. - `midpoint` mirrors parquet-java's `ParquetMetadataConverter.filterFileMetaDataByMidpoint`: the smaller of column 0's dictionary and data page offsets, plus half of `RowGroupMetaData::compressed_size()`. - The opener applies the setting to range pruning and to the `bytes_processed` credit, which must agree on which range owns a row group. - `RowGroupAccessPlanFilter::prune_by_range` keeps its signature and behavior. The new `prune_by_range_with_assignment` takes the assignment. - Proto: new `row_group_range_assignment` field in `ParquetOptions`. An empty value decodes as the default, so plans encoded before this change still load. - Docs: `configs.md`, and the `FileRange` rustdoc, which still said row groups are picked by midpoint. ## What is the testing strategy for this PR? - `row_group_filter.rs`: the layout from apache/datafusion-comet#3817 under both rules, a sweep that checks every row group lands in exactly one range for every split size, and the dictionary offset case where parquet-java and DataFusion's start offsets differ. - `opener/mod.rs`: `each_range_reads_the_row_groups_its_assignment_gives_it` checks the rows and `bytes_processed` of each range under both rules, on a split where the rules disagree. It fails if the byte credit ignores the setting. - `datafusion/core/tests/parquet/file_range_pruning.rs`: end-to-end reads of a split two-row-group file under both rules, including a sweep of split points. - Config parsing, proto round trips (including the COPY TO plan round trip), `information_schema.slt`, and a `repartition_scan.slt` case that reads a multi-row-group file split into 4 ranges under both rules. ## Are there any user-facing changes? A new config option. Nothing changes unless it is set to `midpoint`. `ParquetOptions` gains a public field. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
