Siyao Meng created HDDS-16841:
---------------------------------

             Summary: A repeated OM bootstrap under the other Raft role blocks 
membership changes and leader transfer on the leader
                 Key: HDDS-16841
                 URL: https://issues.apache.org/jira/browse/HDDS-16841
             Project: Apache Ozone
          Issue Type: Bug
            Reporter: Siyao Meng
         Attachments: MC-5-reject-bootstrap-of-existing-ratis-member.patch, 
TestBugMC5RepeatedBootstrapBlocksMembershipChanges.java

h3. Mechanism
{{OzoneManagerRatisServer.addOMToRatisRing}} does not check whether the node is 
already a member of the Raft group. It takes the leader's voters and listeners 
and appends the node to the list for the role named in the bootstrap request. 
When the node is already a member under the other role, its id is in both lists 
of the {{SetConfigurationRequest}}.

Ratis notices the duplicate too late. {{LeaderStateImpl.startSetConfiguration}} 
registers it as the pending configuration change 
({{PendingRequests.addConfRequest}}) and only then builds the new 
configuration, which throws {{IllegalArgumentException}} ("Failed to initialize 
listeners: Found omNode-bootstrap-1 in existing peers"). Nothing clears the 
pending request. Every later {{setConfiguration}} on that leader fails the 
precondition in {{addConfRequest}} with {{IllegalStateException}} for as long 
as it stays leader. This is Ratis 3.3.1, the version Ozone uses. The current 
Ratis master has the same order (code reading, not run). The Ratis side is not 
changed here.

h3. Trigger
An OM that is already a voter is started with {{ozone om --bootstrap}} again 
while its own config now lists it in {{ozone.om.listener.nodes.<serviceId>}}, 
for example to turn a voter into a listener without a decommission first. The 
config of the running OMs does not have to change: the checks on the joining OM 
pass and the request reaches the leader.

The mirror case, a listener that is added again as a voter, fails in the same 
way at the Ratis server level (the unit test of the patch). It was not run on a 
mini cluster.

A repeat under the same role does not have this effect: the id is then twice in 
one list and the request is rejected when it is built (code reading).

h3. Impact
Reproduced on unmodified source with real OMs in a mini cluster. On the same 
leader, in the same term:
* A decommission of that OM ({{ozone admin om decommission}}) fails with 
{{IllegalStateException}}.
* A leader transfer ({{ozone admin om transfer}}) fails with 
{{RaftRetryFailureException}} after about a minute of retries, because it first 
changes priorities with {{setConfiguration}}.
* The bootstrap of any other OM goes through the same call and would fail too 
(code reading, not run).
* The second start of the OM does not fail quickly. Its client retries up to 
{{ozone.client.failover.max.attempts}} times (default 500), and every attempt 
that reaches the leader is another failed request. A run without a limit was at 
71 attempts after about 40 seconds when it was stopped. The reproduction sets 
the limit to 5.

The Raft configuration is not changed and no data is lost. Client writes were 
not exercised. The state is gone once another OM is leader: in the reproduction 
the same decommission succeeds after Ratis has moved the leadership. The admin 
transfer command is itself blocked, so in practice the operator has to restart 
the leader OM (argued, the attached class moves the leadership through the 
Ratis server API).

Affects 2.1.0 and later, where listeners were added (read from the release 
tags, not run there).

h3. Reproduction
[^TestBugMC5RepeatedBootstrapBlocksMembershipChanges.java] is self contained, 
goes into {{hadoop-ozone/integration-test}} (target path and Maven command are 
in the class comment) and passes on unmodified source when the defect is 
present. It starts three OMs in a mini cluster and bootstraps a fourth in the 
normal way. The fourth is stopped, its own config gets the listener key naming 
itself, and it is started again in bootstrap mode on its existing storage 
({{OzoneManager.createOm}} with {{StartupOption.BOOTSTRAP}}). Then the class 
sends a decommission through {{OMAdminProtocol.decommission}} and calls 
{{OzoneManager.transferLeadership}}, checks that leader, term and Raft 
configuration are unchanged, asks the Ratis server to transfer the leadership 
(no configuration change is involved) and sends the same decommission to the 
new leader. The CLI itself is not run. No state is injected and nothing is 
timed. The only setting that differs from the defaults is the retry limit of 
the second start.

h3. Patch
[^MC-5-reject-bootstrap-of-existing-ratis-member.patch], against 
7fcf31294859d5d31163e016b5a9a63f21c6edd2. It also applies to master at 
1cc6423590f (not built there).

{{addOMToRatisRing}} refuses a node that the Raft configuration already has as 
a voter or as a listener, before a request is built. The joining OM gets "OM 
<id> is already a member of the Ratis group <group>. Start it without 
--bootstrap." as its bootstrap error and does not retry.

What changes for a repeated bootstrap:
* Under the other role it is refused with that message instead of blocking the 
leader.
* Under the same role it failed before as well, when the request was built, and 
the joining OM kept retrying. Now it is refused with that message (code reading 
for the old behaviour).
* In two narrow cases a repeat that used to be answered with success and no 
change is now refused (code reading, not run). One is a member the leader's own 
peer list does not have, such as an OM that joined with {{--bootstrap --force}} 
without config on the running OMs and is started that way again. The other is a 
retried bootstrap request that reaches the leader after the first one has 
committed and before the leader has applied it. In both the unmodified code 
sent a list equal to the current configuration. The hint in the message covers 
them: the OM is a member and starts without {{--bootstrap}}.

A member cannot change its role by bootstrapping again, with or without the 
patch. That still needs a decommission and a new bootstrap.

Covered by {{testAddRejectsExistingMember}} in the existing 
{{TestOzoneManagerRatisServer}}: a group with one listener, the same id is 
added again as a voter and as a listener, and a later removal must still work. 
On unmodified source it fails with the Ratis {{IllegalArgumentException}} 
quoted above. With the patch {{TestOzoneManagerRatisServer}} (8 tests), 
{{TestOMAdminProtocolServerSideImpl}} (1 test) and 
{{TestAddRemoveOzoneManager}} (8 tests) pass and checkstyle is clean. With the 
patch the attached class fails at its assertion on the request in the Ratis 
log, because the leader refuses the second bootstrap with the message above and 
no request reaches Ratis.

The patch of HDDS-16840 changes the same method and the same test class. The 
two patches do not apply together as they are. Whichever lands second needs a 
small rebase: the check goes in front of the place where the lists are built, 
the test helpers both patches add exist once, and the removal at the end of 
{{testAddRejectsExistingMember}} takes the node id. With both changes in one 
tree {{TestOzoneManagerRatisServer}} (12 tests) passes.

Found by TLA+ model checking and code review of the OM HA membership change 
paths (bootstrap, decommission and leader transfer) under HDDS-15926, on commit 
7fcf31294859d5d31163e016b5a9a63f21c6edd2. Checked against HDDS and RATIS issues 
and apache/ozone pull requests for duplicates before filing. No existing issue 
was found for the order in {{startSetConfiguration}} on the Ratis side. The 
attached patch is a proposal for review. Generated with Specula (Claude Opus 
5.5).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to