[
https://issues.apache.org/jira/browse/HDDS-16300?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated HDDS-16300:
----------------------------------
Labels: pull-request-available (was: )
> Allow the Ozone Manager to dynamically reconfigure its SCM node list
> (ozone.scm.nodes / ozone.scm.address) without a restart
> ----------------------------------------------------------------------------------------------------------------------------
>
> Key: HDDS-16300
> URL: https://issues.apache.org/jira/browse/HDDS-16300
> Project: Apache Ozone
> Issue Type: New Feature
> Reporter: Mohammadhani Fouladgar
> Assignee: Mohammadhani Fouladgar
> Priority: Major
> Labels: pull-request-available
>
> HDDS-13890 made DataNodes refresh their SCM connections without a restart,
> but it only touched the DataNode heartbeat/registration path
> ({{{}SCMConnectionManager{}}} + {{{}StateContext{}}} endpoints). The OM
> reaches SCM through a different mechanism: the {{SCMFailoverProxyProvider}}
> family, which loads the SCM list once at startup via
> {{SCMNodeInfo.buildNodeInfo(conf)}} ({{SCMFailoverProxyProviderBase}}) and
> never reloads. The OM does not register {{ozone.scm.nodes}} as
> reconfigurable, and the proxy providers have no reload hook.
> Consequence: when a new SCM is added, the OM does not learn about it. While
> the new SCM stays a follower this is harmless -- the OM still reaches the
> leader via the existing configured SCMs. But once leadership moves to the new
> SCM, the OM cannot route to a leader that is not in its static config and
> requests fail (the same failover-to-unconfigured-leader outage as the
> leadership-transfer scenario). So the OM must eventually learn the new SCM --
> and the only alternative today is an OM restart, which violates the
> restart-free requirement.
> *Proposed change:*
> * Register {{ozone.scm.nodes.<svc>}} (and the address keys) as reconfigurable
> on the OM, implemented in the {{SCMFailoverProxyProviderBase}} reload path.
> * On reconfiguration, reload the SCM list from config and rebuild the SCM
> failover proxy set (add SCM proxies) -- the client-side counterpart to
> HDDS-13890.
> *Acceptance criteria:*
> * With the OM running, updating the SCM node config and invoking
> reconfiguration makes it aware of a newly added SCM without a restart.
> * After reconfiguration, the OM can reach the new SCM when it becomes leader.
> * Covers the SCM proxy providers the OM actually uses: block-location and
> container-location.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]