Gabriel39 opened a new pull request, #68780:
URL: https://github.com/apache/doris/pull/68780
### What problem does this PR solve?
Parquet V2 scans with a direct fixed-width predicate can fail when a page
contributes no candidate rows. The native reader advances past the page and
reports no predicate execution kind, but the adapter merges that kind with the
preceding data page and fails its execution-kind check.
A stale optional OffsetIndex also exposes this path. After the reader
discards the index, subsequent pages still need their headers parsed before
computing row ranges; otherwise their bounds remain stale, causing empty
fragments or missing rows.
Ignore zero-row fragments when merging predicate execution kinds, retain the
no-progress and final row-count checks, and refresh page bounds after an
OffsetIndex fallback. Active, reconciled indexes retain lazy page skipping.
### Release note
Fix Parquet V2 scan failures on empty predicate page fragments and incorrect
row accounting after falling back from a stale OffsetIndex.
### Check List (For Author)
- Test
- [x] Unit Test
The tests cover a skipped middle page, multi-page direct filtering and
ordinary reads after index fallback, and an Arrow-written Parquet file with a
stale first-page index size. File scans compare every filtered value and
projected payload with page indexes enabled and disabled, using both cross-page
and multi-batch reads.
Validation:
- Reproduced the execution-kind failure on the unmodified branch with the
native-reader tests and the complete-file scan test. The extended ordinary-read
test returned zero rows instead of the expected next-page row.
- ASAN BE tests: all 730 tests selected by
`*Parquet*:FileScannerV2Test.*:TableReaderTest.*` passed, including all four
targeted regression tests.
- clang-format 16 and `git diff --check` passed.
- Behavior changed:
- [x] Yes. Valid page data remains readable after discarding an
inconsistent optional index.
- Does this need documentation?
- [x] No.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]