Siyao Meng created HDDS-16836:
---------------------------------
Summary: ozone admin om roles reports LISTENER or FOLLOWER from
the local config of the answering OM, not from the Raft configuration
Key: HDDS-16836
URL: https://issues.apache.org/jira/browse/HDDS-16836
Project: Apache Ozone
Issue Type: Bug
Reporter: Siyao Meng
Attachments:
CR-2-take-listener-role-in-service-list-from-raft-conf.patch,
TestBugCR2ServiceListFromLocalConfig.java
h3. Mechanism
{{OzoneManager.getServiceList}} takes LISTENER or FOLLOWER for each OM from
{{OMNodeDetails.isRatisListener}}, which is what the config of the answering OM
says. Only the leader comes from Ratis. {{OzoneManager.getRatisRoles}} (the OM
web UI and JMX) formats the same list. The reported role can therefore differ
from the committed Raft configuration.
The Raft role of a new OM is decided by the config of the new OM
({{BootstrapOMRequest.isListener}}), while every existing OM reads the flag for
it from its own config. {{OzoneManager.checkRemoteOMConfig}} compares the
address and the decommissioned nodes list of the two configs, not
{{ozone.om.listener.nodes}}.
Present since listener OMs were added in 2.1.0 (HDDS-11523).
h3. Trigger
A normal {{ozone om --bootstrap}} of an OM whose
{{ozone.om.listener.nodes.<serviceId>}} differs from the one on the existing
OMs. The config check passes in both directions.
* The new OM is in the key only in its own config (run): Ratis adds a listener,
and every existing OM reports it as FOLLOWER.
* The new OM is in the key on the existing OMs and not in its own config (from
the code, not run through a bootstrap; the regression test produces the same
view by setting the flag on the leader): Ratis adds a voter, and every existing
OM reports it as LISTENER. This is the direction in which the report hides a
change of the quorum size.
h3. Impact
Reporting only. {{ozone admin om roles}}, the OM web UI, JMX and
{{RpcClient.getOmRoleInfos}} show a role that differs from the Raft
configuration. I found no reader of this list that makes a membership, quorum
or failover decision from the role: the client failover proxy providers and the
decommission and transfer commands use the config, Recon reads only the leader
entry, the S3 gateway and the client read versions and certificates. The
exception is the Freon OM metadata generator, a benchmark tool, which picks
FOLLOWER entries for follower reads.
The same difference between the local view and the Raft configuration is
harmful where the leader builds a new Raft configuration from it. That is a
separate defect and not part of this issue.
h3. Reproduction
PASS, no fault injection and no timing hooks, unmodified source.
[^TestBugCR2ServiceListFromLocalConfig.java]
({{hadoop-ozone/integration-test}}) uses the real OMs of a mini cluster. The
first test shows the listener reported as FOLLOWER by all three existing OMs in
{{getServiceList}} and {{getRatisRoles}} while their Raft configuration has it
as listener. The second test shows the related gap below. A passing test means
the defect is present.
h3. Suggested fix
Take the role from the Raft configuration (the attached patch). The bootstrap
config check could also compare {{ozone.om.listener.nodes}}, so that the
difference cannot arise through a normal bootstrap; that does not replace the
patch for a forced bootstrap or for a ring that is already in that state.
h3. Patch
[^CR-2-take-listener-role-in-service-list-from-raft-conf.patch], against
7fcf31294859d5d31163e016b5a9a63f21c6edd2. It also applies to master at
1cc6423590f (not built there).
{{OzoneManager.getServiceList}} now reports an OM as LISTENER if the Raft
configuration of the answering OM has it as a listener, for the answering OM
and for its peers. The listeners are read through the Ratis division the method
already uses for the leader ID, so no new exception can come out of the call.
Nothing else in the list changes.
Covered by {{testServiceListHasListenersOfRaftConf}} in the existing
{{TestAddRemoveOzoneManager}}. The test gives a voter the listener flag in the
local view of the leader and expects the list to follow the Raft configuration.
Without the change it fails because that voter is reported as LISTENER. With it
the nine tests of {{TestAddRemoveOzoneManager}} and {{TestServiceInfoProvider}}
pass and checkstyle is clean. With the patch the first test of the reproduction
fails because the new OM is reported as LISTENER, and the second still passes.
h3. Related, not fixed by this patch
The set of OMs in the list also comes from the local view ({{peerNodesMap}}).
{{OzoneManager.addOMNodeToPeers}} logs an error and returns when the local
config has no address for a new member, after the change is already committed.
After {{ozone om --bootstrap --force}} against existing OMs whose config does
not have the new OM, Ratis has a fourth voter and the existing OMs keep
reporting three OMs (run, second test of the reproduction). Such a member
cannot be listed completely, because each entry carries the host name and RPC
port from the local config and the Raft configuration only has the Ratis
address.
Found by code review of the OM HA membership change paths (bootstrap,
decommission and leader transfer), as part of the TLA+ verification effort
under HDDS-15926, on commit 7fcf31294859d5d31163e016b5a9a63f21c6edd2. Checked
against HDDS issues and apache/ozone pull requests for duplicates before
filing. The attached patch is a proposal for review. Generated with Specula
(Claude Opus 5.5).
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]