yujun777 commented on code in PR #68646:
URL: https://github.com/apache/doris/pull/68646#discussion_r4143670230
##########
fe/fe-core/src/main/java/org/apache/doris/alter/Alter.java:
##########
@@ -373,8 +374,30 @@ private boolean
processAlterOlapTableInternal(List<AlterOp> alterOps, OlapTable
throw new DdlException("Invalid alter operations: " +
currentAlterOps);
}
if (needChangeMTMVState(alterOps)) {
- Env.getCurrentEnv().getMtmvService()
- .alterTable(oldBaseTableInfo, newBaseTableInfo,
currentAlterOps.hasReplaceTableOp());
+ // Which columns an operation's effect on a view turns on is the
operation's to say, see
+ // AlterOp#queryJudgedColumnNames, and every clause of the alter
has to name them: a batch that
+ // mixes a dropped column with a type change is decided by neither
-- no query says anything
+ // about a type change -- and stays invalidated the way it was
before the queries were asked at
+ // all. Each of them also has to have reached the table. A schema
change that is not a light one
+ // is applied by a job, which may not have run where this hook
runs: the table still holds the
+ // column the change is about, every query still analyses against
it, and an invalidation
+ // decided on that answer would be about the table from before the
change. What is asked is
+ // whether the change has reached the table, which is the same
fact the re-analysis reads, so
+ // the two answers cannot disagree.
+ boolean judgedByQuery = alterOps.stream().allMatch(op ->
!op.queryJudgedColumnNames().isEmpty()
+ && op.hasReachedTheTable(olapTable));
+ // The names of those columns go to the hook rather than a
verdict: what the judgement is about
+ // is the column, and the hook holds the query's answer against it
-- both while asking, in case
+ // the query can reach the name some other way now, and once it
has answered, in case the table
+ // is no longer the one that answered. See MTMVRelationManager.
+ MTMVHookService.QueryJudgedChange queryJudgedChange = judgedByQuery
+ ? new MTMVHookService.QueryJudgedChange(
+
alterOps.stream().map(AlterOp::queryJudgedColumnNames).flatMap(Set::stream)
+ .collect(Collectors.toSet()),
+ () -> alterOps.stream().allMatch(op ->
op.hasReachedTheTable(olapTable)))
+ : null;
+ Env.getCurrentEnv().getMtmvService().alterTable(oldBaseTableInfo,
newBaseTableInfo,
Review Comment:
Agreed, and this one is not fixed here: it is not a precision question like
the others but the window between `SchemaChangeHandler.process` publishing the
new column and this hook invalidating the views that read it, and closing it
means the invalidation has to be established before the new schema is visible
to a planner -- either by moving the whole-MV decision to where the write lock
is still held, or by a barrier the rewrite's plan building respects. It needs a
decision on which of the two, so I have not guessed at it in this commit.
##########
fe/fe-core/src/main/java/org/apache/doris/mtmv/MTMVRelationManager.java:
##########
@@ -349,34 +359,167 @@ public void dropTable(Table table) {
// because a dropped table is the one change whose query is gone
beyond doubt. What the two record
// is the same state either way. Unlike a rename it stays an
invalidation: the table is gone for
// good, so the state is not something a later alter can make obsolete.
- processBaseTableChange(new BaseTableInfo(table), "The base table has
been deleted:", false);
+ processBaseTableChange(new BaseTableInfo(table), "The base table has
been deleted:", null);
}
/**
* update mtmv status to `SCHEMA_CHANGE`.
*
* @param isReplace
+ * @param queryJudgedColumns the names the alter gives the table or takes
away from it, which leave the
+ * judgement about each MV's state to that MV's
own query, or null when the
+ * alter is not one a query decides. The names
are carried rather than judged
+ * before the call because the judgement is
about them; see
+ * {@code AlterOp#queryJudgedColumnNames} for
which operations name one, and
+ * {@link #invalidateMvUnlessQueryHolds} for
what is asked about it. A rename
+ * of the base table names no column: it is left
to the record below, which
+ * says what the MV that keeps spelling the old
name needs to hear
*/
@Override
- public void alterTable(BaseTableInfo oldTableInfo, Optional<BaseTableInfo>
newTableInfo, boolean isReplace) {
+ public void alterTable(BaseTableInfo oldTableInfo, Optional<BaseTableInfo>
newTableInfo, boolean isReplace,
+ QueryJudgedChange queryJudgedChange) {
// when replace, need deal two table
if (isReplace) {
// REPLACE TABLE already invalidates the IVM baseline explicitly,
see Alter#processReplaceTable
- processBaseTableChange(newTableInfo.get(), "The base table has
been updated:", false);
+ processBaseTableChange(newTableInfo.get(), "The base table has
been updated:", null);
}
- boolean renamed = !isReplace && newTableInfo.isPresent()
- && !Objects.equals(oldTableInfo.getTableName(),
newTableInfo.get().getTableName());
- // A rename is the one change whose query check is skipped: the MV
query keeps spelling the old
- // name, so it is unanalyzable by construction, and the reason it
would be invalidated with --
- // "the query is no longer analyzable" -- says less than the message
this call records anyway.
- boolean checkQueryUsable = !renamed;
- processBaseTableChange(oldTableInfo, "The base table has been
updated:", checkQueryUsable);
+ processBaseTableChange(oldTableInfo, "The base table has been
updated:", queryJudgedChange);
}
/**
- * An MV's query is only as good as the base table schema it was analyzed
against. Re-analyzing the
- * MV query here (right after the alter was applied) is what detects a
changed column identity:
+ * Whether the query, as it is analysed now, reads a column of any of
these names, and reads it where
+ * the change can reach it.
+ *
+ * <p>There are two places a name is the change's to answer for. One is a
column of the table the change
+ * is about: that is the column this view's rows were computed from, and
the names are matched
+ * case-insensitively because a name is what moves. The other is a column
the query reaches across a
+ * scope boundary -- the plan records those on the Apply that stands for
the subquery, whose correlation
+ * slots are the outer columns its right side reads -- because such a name
is the scopes' to answer for
+ * rather than the query's: the nearest column to the reference answers
for it, so a column the change
+ * takes away from a scope inside leaves the name to one outside, and a
column it gives to a scope inside
+ * takes the name over. A name reached with the qualifier of another table
inside the query's own scope
+ * is neither: no later change can move it, so one to a column it does not
name is one this view's rows
+ * do not depend on.
+ */
+ private static boolean reachesAnyColumnOf(Plan plan, BaseTableInfo
baseTableInfo, Set<String> columnNames) {
+ if (plan == null) {
+ // A query whose plan was not kept is one this cannot be answered
about, and "it does" is the
+ // answer that keeps the view safe.
+ return true;
+ }
+ Set<String> names = Sets.newTreeSet(String.CASE_INSENSITIVE_ORDER);
+ names.addAll(columnNames);
+ LineageInfo lineage = LineageInfoExtractor.extractLineageInfo(plan);
+ for (SetMultimap<?, Expression> byType :
lineage.getDirectLineageMap().values()) {
+ if (reachesAnyColumn(byType.values(), names, baseTableInfo)) {
+ return true;
+ }
+ }
+ // The dataset predicates once, not once per output column: the
per-output copy of them the lineage
+ // also offers holds the same expressions for every column the query
produces, and scanning it would
+ // visit each of them once per column.
+ if (reachesAnyColumn(lineage.getDatasetIndirectLineageMap().values(),
names, baseTableInfo)) {
Review Comment:
Agreed, and this one is not fixed here either: the dataset lineage is what
the view's own columns are read from, and a CTE producer the result never
consumes is read by the same scan because the analysed plan keeps it under the
anchor. Restricting that scan means walking the plan for what the result
reaches -- skipping the producers no consumer reads -- which is a change to how
this check traverses the plan rather than a name it matches, so it should land
on its own.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]