Siyao Meng created HDDS-16836:
---------------------------------

             Summary: ozone admin om roles reports LISTENER or FOLLOWER from 
the local config of the answering OM, not from the Raft configuration
                 Key: HDDS-16836
                 URL: https://issues.apache.org/jira/browse/HDDS-16836
             Project: Apache Ozone
          Issue Type: Bug
            Reporter: Siyao Meng
         Attachments: 
CR-2-take-listener-role-in-service-list-from-raft-conf.patch, 
TestBugCR2ServiceListFromLocalConfig.java

h3. Mechanism
{{OzoneManager.getServiceList}} takes LISTENER or FOLLOWER for each OM from 
{{OMNodeDetails.isRatisListener}}, which is what the config of the answering OM 
says. Only the leader comes from Ratis. {{OzoneManager.getRatisRoles}} (the OM 
web UI and JMX) formats the same list. The reported role can therefore differ 
from the committed Raft configuration.

The Raft role of a new OM is decided by the config of the new OM 
({{BootstrapOMRequest.isListener}}), while every existing OM reads the flag for 
it from its own config. {{OzoneManager.checkRemoteOMConfig}} compares the 
address and the decommissioned nodes list of the two configs, not 
{{ozone.om.listener.nodes}}.

Present since listener OMs were added in 2.1.0 (HDDS-11523).

h3. Trigger
A normal {{ozone om --bootstrap}} of an OM whose 
{{ozone.om.listener.nodes.<serviceId>}} differs from the one on the existing 
OMs. The config check passes in both directions.
* The new OM is in the key only in its own config (run): Ratis adds a listener, 
and every existing OM reports it as FOLLOWER.
* The new OM is in the key on the existing OMs and not in its own config (from 
the code, not run through a bootstrap; the regression test produces the same 
view by setting the flag on the leader): Ratis adds a voter, and every existing 
OM reports it as LISTENER. This is the direction in which the report hides a 
change of the quorum size.

h3. Impact
Reporting only. {{ozone admin om roles}}, the OM web UI, JMX and 
{{RpcClient.getOmRoleInfos}} show a role that differs from the Raft 
configuration. I found no reader of this list that makes a membership, quorum 
or failover decision from the role: the client failover proxy providers and the 
decommission and transfer commands use the config, Recon reads only the leader 
entry, the S3 gateway and the client read versions and certificates. The 
exception is the Freon OM metadata generator, a benchmark tool, which picks 
FOLLOWER entries for follower reads.

The same difference between the local view and the Raft configuration is 
harmful where the leader builds a new Raft configuration from it. That is a 
separate defect and not part of this issue.

h3. Reproduction
PASS, no fault injection and no timing hooks, unmodified source. 
[^TestBugCR2ServiceListFromLocalConfig.java] 
({{hadoop-ozone/integration-test}}) uses the real OMs of a mini cluster. The 
first test shows the listener reported as FOLLOWER by all three existing OMs in 
{{getServiceList}} and {{getRatisRoles}} while their Raft configuration has it 
as listener. The second test shows the related gap below. A passing test means 
the defect is present.

h3. Suggested fix
Take the role from the Raft configuration (the attached patch). The bootstrap 
config check could also compare {{ozone.om.listener.nodes}}, so that the 
difference cannot arise through a normal bootstrap; that does not replace the 
patch for a forced bootstrap or for a ring that is already in that state.

h3. Patch
[^CR-2-take-listener-role-in-service-list-from-raft-conf.patch], against 
7fcf31294859d5d31163e016b5a9a63f21c6edd2. It also applies to master at 
1cc6423590f (not built there).

{{OzoneManager.getServiceList}} now reports an OM as LISTENER if the Raft 
configuration of the answering OM has it as a listener, for the answering OM 
and for its peers. The listeners are read through the Ratis division the method 
already uses for the leader ID, so no new exception can come out of the call. 
Nothing else in the list changes.

Covered by {{testServiceListHasListenersOfRaftConf}} in the existing 
{{TestAddRemoveOzoneManager}}. The test gives a voter the listener flag in the 
local view of the leader and expects the list to follow the Raft configuration. 
Without the change it fails because that voter is reported as LISTENER. With it 
the nine tests of {{TestAddRemoveOzoneManager}} and {{TestServiceInfoProvider}} 
pass and checkstyle is clean. With the patch the first test of the reproduction 
fails because the new OM is reported as LISTENER, and the second still passes.

h3. Related, not fixed by this patch
The set of OMs in the list also comes from the local view ({{peerNodesMap}}). 
{{OzoneManager.addOMNodeToPeers}} logs an error and returns when the local 
config has no address for a new member, after the change is already committed. 
After {{ozone om --bootstrap --force}} against existing OMs whose config does 
not have the new OM, Ratis has a fourth voter and the existing OMs keep 
reporting three OMs (run, second test of the reproduction). Such a member 
cannot be listed completely, because each entry carries the host name and RPC 
port from the local config and the Raft configuration only has the Ratis 
address.

Found by code review of the OM HA membership change paths (bootstrap, 
decommission and leader transfer), as part of the TLA+ verification effort 
under HDDS-15926, on commit 7fcf31294859d5d31163e016b5a9a63f21c6edd2. Checked 
against HDDS issues and apache/ozone pull requests for duplicates before 
filing. The attached patch is a proposal for review. Generated with Specula 
(Claude Opus 5.5).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to