Siyao Meng created HDDS-16837:
---------------------------------
Summary: ozone.om.listener.nodes without the service ID suffix, as
documented, is silently ignored and the intended listener OM joins as a voter
Key: HDDS-16837
URL: https://issues.apache.org/jira/browse/HDDS-16837
Project: Apache Ozone
Issue Type: Bug
Reporter: Siyao Meng
Attachments:
CR-5-warn-when-listener-nodes-key-lacks-service-id-suffix.patch,
TestBugCR5UnsuffixedListenerNodesKey.java
h3. Mechanism
{{OmUtils.getListenerOMNodeIds}} and {{OmUtils.getAllOMHAAddresses}} read only
{{ozone.om.listener.nodes.<serviceId>}}. The key without the suffix is never
read and no warning is logged about it. Both documents that describe the
feature show the key without the suffix ({{ozone.om.listener.nodes}} with the
value {{om4,om5}} and one service ID):
* the Listener OM page of the site,
{{docs/05-administrator-guide/02-configuration/05-high-availability/04-listener-om.md}}
in apache/ozone-site. It is published under {{/docs/next/}} only, no released
docs version has the page.
* the design document {{hadoop-hdds/docs/content/design/listener-om.md}} in the
main repository.
The key has no entry in {{ozone-default.xml}}, so the documents are the only
place an operator can take it from.
With the key ignored, {{OMNodeDetails.isRatisListener}} is false for the node
on every OM, including the node itself, so its bootstrap request asks for a
voting member.
The readers and the design document are in the release tags 2.1.0 to 2.2.1
(HDDS-11523).
h3. Trigger
# One OM service ID.
# {{ozone.om.listener.nodes=om4}} in {{ozone-site.xml}} on all OMs, as
documented.
# {{ozone om --bootstrap}} on om4.
No fault is needed.
h3. Impact
The node the operator declared a listener is committed to the Raft
configuration as a voter. It counts towards the majority, its loss costs fault
tolerance, and it can be elected leader, which is what a listener is meant to
exclude. No warning is logged about the ignored key. The only signs are
{{isListener: false}} in the startup log line of the node (seen in the run) and
FOLLOWER instead of LISTENER in {{ozone admin om roles}} (from the code).
Run on a mini cluster: one OM plus one OM named in the key without suffix.
After the bootstrap the Raft configuration has two voters and no listener, and
after the new OM is stopped the first OM no longer stays leader. An election of
the new OM as leader was not run.
h3. Reproduction
PASS, no fault injection and no timing hooks, unmodified source.
[^TestBugCR5UnsuffixedListenerNodesKey.java]
({{hadoop-ozone/integration-test}}) uses the real OMs of a mini cluster and its
normal bootstrap. A passing test means the defect is present.
h3. Suggested fix
# Correct the two documents to {{ozone.om.listener.nodes.<serviceId>}}. The
site page needs a separate pull request to apache/ozone-site, the design
document is in the main repository. The attached patch touches neither.
# Log a warning at OM start when the key is set without the suffix (the
attached patch).
Reading the key without suffix when one service ID is configured, as HDDS-10942
did for {{ozone.om.decommissioned.nodes}}, is not proposed, for compatibility.
A cluster that already has the key without suffix runs that node as a voter,
and it would keep voting after an upgrade because Ratis starts from its
persisted configuration (from the code, not run). Every process that reads the
config would then treat it as a listener, and the client failover proxy
providers leave listeners out of their OM list. Run with such a fallback
applied: three voters, a client config that names the current leader in the key
without suffix and 5 failover attempts. The client failed after about 4 seconds
with {{OMNotLeaderException}} and "Failed to connect to OMs: [omNode-3,
omNode-2]. Attempted 5 failovers." Without a fallback the same client config
writes normally (run with the attached patch).
h3. Patch
[^CR-5-warn-when-listener-nodes-key-lacks-service-id-suffix.patch], against
7fcf31294859d5d31163e016b5a9a63f21c6edd2. It also applies to master at
1cc6423590f (not built there).
{{OMHANodeDetails.loadOMHAConfig}} logs one warning when
{{ozone.om.listener.nodes}} has a value: "ozone.om.listener.nodes is set
without the OM service ID suffix and is not used. Listener OMs are read from
ozone.om.listener.nodes.<serviceId> only.", with the service ID of the OM
filled in. Nothing else changes: the key without suffix is still not read and
the OM starts as before.
Covered by {{testListenerNodesKeyWithoutServiceIdSuffix}} in the existing
{{TestOzoneManagerConfiguration}}: with the suffixed key the peer is a listener
and nothing is logged, with the key without suffix the peer is not a listener
and the warning names the suffixed key. Without the change it fails because the
warning is missing. With it the ten tests of {{TestOzoneManagerConfiguration}}
pass and checkstyle is clean. The reproduction still passes with the patch,
because the behaviour is unchanged: the new OM joins as a voter, and its log
now has the warning.
Found by code review of the OM HA membership change paths (bootstrap,
decommission and leader transfer), as part of the TLA+ verification effort
under HDDS-15926, on commit 7fcf31294859d5d31163e016b5a9a63f21c6edd2. Checked
against HDDS issues and apache/ozone pull requests for duplicates before
filing. The attached patch is a proposal for review. Generated with Specula
(Claude Opus 5.5).
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]