yujun777 opened a new pull request, #68648: URL: https://github.com/apache/doris/pull/68648
### What problem does this PR solve? Issue Number: N/A Related PR: #68646 Trace issue: https://github.com/apache/doris/issues/65418 Problem Summary: A refresh of one MV partition reads each base table through the key range of that MV partition, while the snapshot it records next to the read names the base partitions the MV partition's mapping keeps. The two sets are not the same one: the mapping is what `MTMV.calculatePartitionMappings` answers, and `partition_sync_limit` filters it, so a refresh reads base partitions the snapshot does not describe -- and the rows they put in the MV partition stay there unaccounted for. That difference is what a silent base partition change hides behind. `DROP PARTITION`, `TRUNCATE PARTITION`, `REPLACE PARTITION` and `RECOVER PARTITION` remove visible rows through metadata rather than through row binlog entries, so a refresh has to be told about them. An IVM MV is: its per-partition baseline is invalidated, which reaches every MV partition that reads the changed one. A plain MV is judged by its snapshot, and when the changed partition is one the snapshot never named, the base partition set is back to what the snapshot does name once the change is done: the MV partition is judged synchronized, the transparent rewrite serves the rows of a partition the base table no longer has, and no refresh plans that partition again. This PR makes the read and the record the same set. `MTMVTask#mappedBasePartitions` decides which base partitions a batch may read, from the same mapping the snapshots taken next to that read are generated from, and `UpdateMvByPartitionCommand` builds each base table's read predicate from exactly those partitions: * A base table the mapping gives no partition of is read as nothing rather than in full, which is the reading the IVM rewrite already gives such a scope. * A table the caller does not scope keeps the MV partition's own key range, which is what an external base table and the full refresh of an IVM MV use. * The predicate is built from the base partition's key at the position the MV's partition column has in that base table: a partition of a list partitioned base table holds one key per partition column, and the MV's column is not necessarily the first of them. * The partitions are named by `BaseTableInfo`, so a table is identified by the table it names rather than by which of the MV's partition info and the mapping's own two objects it was read from. Keying by the object would have made "the lookup missed" read as "this table may not be read at all", which would empty the MV quietly. Only the olap base tables of a non-IVM MV are scoped: the changes this answers are olap DDL, and an IVM MV already invalidates its baseline for them and scopes the reads of its delta in the IVM rewrite. With the read equal to the record, a base partition that was read is one the snapshot names, so a later silent change to it is what the sync check compares against, and the partition is planned for a rebuild instead of being served stale. `EXPLAIN REFRESH ... COMPLETE` still shows the MV partitions' own key ranges. The scope is decided by the refresh's own context, which a plan built for an `EXPLAIN` does not have; the partitions pruning picks are the same in the common case, only the predicate differs. Behaviour changed: Yes * With `partition_sync_limit`, an MV partition no longer keeps the rows of base partitions the window no longer keeps: the refresh reads the partitions the window keeps rather than the whole key range of the MV partition, which can be wider. Queries that also read a base partition the window has expired are answered from the base table by the union rewrite, which is what that path is for and is on by default (`enable_materialized_view_rewrite`). ### Release note With `partition_sync_limit`, a materialized view partition no longer keeps the rows of base partitions the sync window no longer keeps. A refresh reads exactly the base partitions the view's partition is recorded with, so what the partition holds and what it is reported as synchronized with agree; a query that also reads expired base partitions is answered from the base table by the union rewrite, as before. ### Check List (For Author) - Test: Unit Test / Regression test - FE unit tests: `UpdateMvByPartitionCommandTest` (10 -> 13: a scoped table is read from the partitions named, a table named with none is read as nothing, and a partition predicate follows the key at the position the MV's partition column has), `MTMVTaskTest`, `MTMVTest`, `MTMVPartitionUtilTest`, `ExplainRefreshMTMVCommandTest` -- all pass. Each new case was checked to fail with the change it covers reverted: the read goes back to the MV partition's own range, an empty scope goes back to reading everything, and the key position is ignored. - Regression: `test_mtmv_base_partition_read_scope` (new: daily base partitions, a year MV partition and a `partition_sync_limit` window keeping two of three days -- the MV holds those two, where it held the whole year before; fails with the change reverted), the `mtmv_p0` suites around roll-up, partition limits, recreate and column changes, and `mtmv_p0/ivm`'s epoch and baseline suites -- all pass. - Behavior changed: Yes (see above) - Does this need documentation: Yes, in a follow-up (no doc PR yet): the pages for `partition_sync_limit` and for the incremental materialized view's refresh modes are tracked in the design document-change list, which records the exact locations and wording. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
