Siyao Meng created HDDS-16841:
---------------------------------
Summary: A repeated OM bootstrap under the other Raft role blocks
membership changes and leader transfer on the leader
Key: HDDS-16841
URL: https://issues.apache.org/jira/browse/HDDS-16841
Project: Apache Ozone
Issue Type: Bug
Reporter: Siyao Meng
Attachments: MC-5-reject-bootstrap-of-existing-ratis-member.patch,
TestBugMC5RepeatedBootstrapBlocksMembershipChanges.java
h3. Mechanism
{{OzoneManagerRatisServer.addOMToRatisRing}} does not check whether the node is
already a member of the Raft group. It takes the leader's voters and listeners
and appends the node to the list for the role named in the bootstrap request.
When the node is already a member under the other role, its id is in both lists
of the {{SetConfigurationRequest}}.
Ratis notices the duplicate too late. {{LeaderStateImpl.startSetConfiguration}}
registers it as the pending configuration change
({{PendingRequests.addConfRequest}}) and only then builds the new
configuration, which throws {{IllegalArgumentException}} ("Failed to initialize
listeners: Found omNode-bootstrap-1 in existing peers"). Nothing clears the
pending request. Every later {{setConfiguration}} on that leader fails the
precondition in {{addConfRequest}} with {{IllegalStateException}} for as long
as it stays leader. This is Ratis 3.3.1, the version Ozone uses. The current
Ratis master has the same order (code reading, not run). The Ratis side is not
changed here.
h3. Trigger
An OM that is already a voter is started with {{ozone om --bootstrap}} again
while its own config now lists it in {{ozone.om.listener.nodes.<serviceId>}},
for example to turn a voter into a listener without a decommission first. The
config of the running OMs does not have to change: the checks on the joining OM
pass and the request reaches the leader.
The mirror case, a listener that is added again as a voter, fails in the same
way at the Ratis server level (the unit test of the patch). It was not run on a
mini cluster.
A repeat under the same role does not have this effect: the id is then twice in
one list and the request is rejected when it is built (code reading).
h3. Impact
Reproduced on unmodified source with real OMs in a mini cluster. On the same
leader, in the same term:
* A decommission of that OM ({{ozone admin om decommission}}) fails with
{{IllegalStateException}}.
* A leader transfer ({{ozone admin om transfer}}) fails with
{{RaftRetryFailureException}} after about a minute of retries, because it first
changes priorities with {{setConfiguration}}.
* The bootstrap of any other OM goes through the same call and would fail too
(code reading, not run).
* The second start of the OM does not fail quickly. Its client retries up to
{{ozone.client.failover.max.attempts}} times (default 500), and every attempt
that reaches the leader is another failed request. A run without a limit was at
71 attempts after about 40 seconds when it was stopped. The reproduction sets
the limit to 5.
The Raft configuration is not changed and no data is lost. Client writes were
not exercised. The state is gone once another OM is leader: in the reproduction
the same decommission succeeds after Ratis has moved the leadership. The admin
transfer command is itself blocked, so in practice the operator has to restart
the leader OM (argued, the attached class moves the leadership through the
Ratis server API).
Affects 2.1.0 and later, where listeners were added (read from the release
tags, not run there).
h3. Reproduction
[^TestBugMC5RepeatedBootstrapBlocksMembershipChanges.java] is self contained,
goes into {{hadoop-ozone/integration-test}} (target path and Maven command are
in the class comment) and passes on unmodified source when the defect is
present. It starts three OMs in a mini cluster and bootstraps a fourth in the
normal way. The fourth is stopped, its own config gets the listener key naming
itself, and it is started again in bootstrap mode on its existing storage
({{OzoneManager.createOm}} with {{StartupOption.BOOTSTRAP}}). Then the class
sends a decommission through {{OMAdminProtocol.decommission}} and calls
{{OzoneManager.transferLeadership}}, checks that leader, term and Raft
configuration are unchanged, asks the Ratis server to transfer the leadership
(no configuration change is involved) and sends the same decommission to the
new leader. The CLI itself is not run. No state is injected and nothing is
timed. The only setting that differs from the defaults is the retry limit of
the second start.
h3. Patch
[^MC-5-reject-bootstrap-of-existing-ratis-member.patch], against
7fcf31294859d5d31163e016b5a9a63f21c6edd2. It also applies to master at
1cc6423590f (not built there).
{{addOMToRatisRing}} refuses a node that the Raft configuration already has as
a voter or as a listener, before a request is built. The joining OM gets "OM
<id> is already a member of the Ratis group <group>. Start it without
--bootstrap." as its bootstrap error and does not retry.
What changes for a repeated bootstrap:
* Under the other role it is refused with that message instead of blocking the
leader.
* Under the same role it failed before as well, when the request was built, and
the joining OM kept retrying. Now it is refused with that message (code reading
for the old behaviour).
* In two narrow cases a repeat that used to be answered with success and no
change is now refused (code reading, not run). One is a member the leader's own
peer list does not have, such as an OM that joined with {{--bootstrap --force}}
without config on the running OMs and is started that way again. The other is a
retried bootstrap request that reaches the leader after the first one has
committed and before the leader has applied it. In both the unmodified code
sent a list equal to the current configuration. The hint in the message covers
them: the OM is a member and starts without {{--bootstrap}}.
A member cannot change its role by bootstrapping again, with or without the
patch. That still needs a decommission and a new bootstrap.
Covered by {{testAddRejectsExistingMember}} in the existing
{{TestOzoneManagerRatisServer}}: a group with one listener, the same id is
added again as a voter and as a listener, and a later removal must still work.
On unmodified source it fails with the Ratis {{IllegalArgumentException}}
quoted above. With the patch {{TestOzoneManagerRatisServer}} (8 tests),
{{TestOMAdminProtocolServerSideImpl}} (1 test) and
{{TestAddRemoveOzoneManager}} (8 tests) pass and checkstyle is clean. With the
patch the attached class fails at its assertion on the request in the Ratis
log, because the leader refuses the second bootstrap with the message above and
no request reaches Ratis.
The patch of HDDS-16840 changes the same method and the same test class. The
two patches do not apply together as they are. Whichever lands second needs a
small rebase: the check goes in front of the place where the lists are built,
the test helpers both patches add exist once, and the removal at the end of
{{testAddRejectsExistingMember}} takes the node id. With both changes in one
tree {{TestOzoneManagerRatisServer}} (12 tests) passes.
Found by TLA+ model checking and code review of the OM HA membership change
paths (bootstrap, decommission and leader transfer) under HDDS-15926, on commit
7fcf31294859d5d31163e016b5a9a63f21c6edd2. Checked against HDDS and RATIS issues
and apache/ozone pull requests for duplicates before filing. No existing issue
was found for the order in {{startSetConfiguration}} on the Ratis side. The
attached patch is a proposal for review. Generated with Specula (Claude Opus
5.5).
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]