[
https://issues.apache.org/jira/browse/HDDS-16275?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Sadanand Shenoy reassigned HDDS-16275:
--------------------------------------
Assignee: Sadanand Shenoy
> OM checkpoint pause() stops flush thread but not apply executor, risking
> divergence on install failure paths
> ------------------------------------------------------------------------------------------------------------
>
> Key: HDDS-16275
> URL: https://issues.apache.org/jira/browse/HDDS-16275
> Project: Apache Ozone
> Issue Type: Bug
> Reporter: Wei-Chiu Chuang
> Assignee: Sadanand Shenoy
> Priority: Minor
> Labels: checkpoint-install, double-buffer, om-ha, ratis,
> state-machine
>
> h3. Summary
> During checkpoint installation, {{OzoneManagerStateMachine.pause()}} stops
> the double-buffer flush thread ({{OzoneManagerDoubleBuffer.stop()}}) but does
> *not* stop or block the apply executor
> ({{OMStateMachineApplyTransactionThread}}). Ratis can continue calling
> {{applyTransaction()}} while the state machine is paused and flush is
> disabled.
> This is acceptable on the *successful install path* (DB replaced with leader
> checkpoint; old buffer/state intentionally discarded). It is a potential
> correctness issue on *failure or skip paths* where the old DB is kept and the
> state machine is unpaused without replacing metadata.
> h3. Background
> {{pause()}} is invoked from {{OzoneManager.installCheckpoint()}} before
> replacing OM metadata. It transitions the state machine to {{PAUSED}} and
> stops the flush daemon, but {{executorService}} is not stopped (contrast with
> {{stop()}}, which shuts down both).
> {{applyTransaction()}} continues to submit work:
> {code}
> ozoneManagerDoubleBuffer.acquireUnFlushedTransactions(1);
> return CompletableFuture.supplyAsync(() -> runCommand(request, termIndex),
> executorService)
> {code}
> Ratis {{StateMachineUpdater}} may keep applying committed entries while its
> own state is {{RUNNING}}; it does not check OM lifecycle {{PAUSED}}. Ratis
> also documents that append can continue during snapshot install
> ({{SnapshotInstallationHandler}}).
> Writes from the apply path go to {{doubleBuffer.add()}} but are not flushed
> while the flush thread is stopped. On {{unpause()}}, a new double buffer is
> created and the old one (with any unflushed entries) is discarded.
> h3. Impact by scenario
> || Scenario || Severity || Notes ||
> | Successful install, DB replaced | Low / by design | Old buffer and
> unflushed applies are intentionally discarded; leader checkpoint is
> authoritative. |
> | {{canProceed == false}} (checkpoint rejected, old DB kept) | Medium | Ratis
> may advance applied index during pause; unflushed buffer entries discarded on
> {{unpause()}} -> possible Ratis vs OM DB divergence. |
> | Failed install with rollback to old DB | Medium | Same as above if applies
> occurred during the pause window. |
> | Long install window | Low (liveness) | {{acquireUnFlushedTransactions()}}
> may block when flush is stopped and permits are exhausted -> Ratis apply
> pipeline stalls until install completes. |
> h3. Root cause
> {{pause()}} and {{stop()}} are asymmetric:
> * {{stop()}} — stops flush thread *and* shuts down apply executor
> * {{pause()}} — stops flush thread only
> There is no guard in {{applyTransaction()}} for {{LifeCycle.State.PAUSED}},
> and Ratis is not coordinated to stop applying before OM {{pause()}}.
> h3. Steps to reproduce (conceptual)
> # HA cluster; follower triggers checkpoint install.
> # Fault-inject or force {{canProceed == false}} after {{pause()}} (or fail
> install before DB replace).
> # While paused, allow Ratis to commit/apply one or more write entries.
> # Complete install path that calls {{unpause()}} without replacing DB.
> # Observe: Ratis applied index may be ahead of OM persisted
> {{lastAppliedTermIndex}} / DB contents for those entries.
> h3. Expected behavior
> While the state machine is paused for checkpoint install, no new Ratis write
> applications should be accepted, or any accepted work should be guaranteed
> reconciled before returning to normal operation on paths that keep the old DB.
> h3. Proposed directions (for discussion)
> # Gate {{applyTransaction()}} when lifecycle is {{PAUSED}} / {{PAUSING}}
> (needs Ratis interaction review).
> # Drain/shutdown apply executor in {{pause()}}, restart in {{unpause()}}
> (mirror {{stop()}} for apply path only).
> # On failure/skip paths only: reconcile Ratis applied index with OM
> {{lastAppliedTermIndex}} before {{unpause()}}.
> # Document that {{pause()}} intentionally allows in-flight apply only when DB
> will be replaced.
> h3. Related work
> * HDDS-6685 — checkpoint lifecycle {{pause()}} / {{unpause()}} /
> {{reinitialize()}}
> * HDDS-16155 — flush-thread deadlock during {{pause()}} (PR #10995)
> * HDDS-16057 — bootstrap install during {{BOOTSTRAPPING}} (different issue)
> h3. Files
> *
> {{hadoop-ozone/ozone-manager/src/main/java/org/apache/hadoop/ozone/om/ratis/OzoneManagerStateMachine.java}}
> *
> {{hadoop-ozone/ozone-manager/src/main/java/org/apache/hadoop/ozone/om/ratis/OzoneManagerDoubleBuffer.java}}
> *
> {{hadoop-ozone/ozone-manager/src/main/java/org/apache/hadoop/ozone/om/OzoneManager.java}}
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]