[
https://issues.apache.org/jira/browse/HDDS-16382?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated HDDS-16382:
----------------------------------
Labels: pull-request-available (was: )
> Dedicated SCM client RPC timeout and retry for OM request critical path
> -----------------------------------------------------------------------
>
> Key: HDDS-16382
> URL: https://issues.apache.org/jira/browse/HDDS-16382
> Project: Apache Ozone
> Issue Type: Improvement
> Reporter: Ivan Andika
> Assignee: Ivan Andika
> Priority: Major
> Labels: pull-request-available
>
> We need reduce the following SCM client configuration so that OM read and
> write critical path is not blocked for too long when SCM is unresponsive
> * hdds.scmclient.rpc.timeout
> * ipc.client.connect.timeout
> * ipc.client.connect.max.retries.on.timeouts
> However, OzoneManager instantiates quite a lot of SCM client and each of them
> share a single configuration although they have different latency and retry
> requirement
> * Critical paths (low latency and fail fast)
> ** scmContainerClient (used in read critical path)
> ** scmBlockClient (used in write critical path)
> *** scmTopologyClient is also part of the scmBlockClient
> * SCM security clients (tolerates high latency and unbounded timeout)
> ** scmSecurityClient
> ** secretKeyClient
> Additionally, they are all stored in the same ClientCache entry. ClientCache
> is a Map<SocketFactory, Client> and the standard implementation SocketFactory
> returned in NetUtils#getDefaultSocketFactory in StandardSocketFactory that
> has the following hashCode
> {code:java}
> @Override
> public int hashCode() {
> return this.getClass().hashCode();
> } {code}
> So this means that even if we use different two configurations, only the
> first configured client will be stored in the ClientCache and the subsequent
> one will reuse the cache entry. Therefore, we cannot simply create a separate
> configuration with different timeout.
> The ClientCache#getClient also documented this decision
> {code:java}
> // Construct & cache client. The configuration is only used for timeout,
> // and Clients have connection pools. So we can either (a) lose some
> // connection pooling and leak sockets, or (b) use the same timeout for all
> // configurations. Since the IPC is usually intended globally, not
> // per-job, we choose (a). {code}
> Since we need two different timeout, it conflicts with the ClientCache
> mechanism.
> Therefore, we need another solution to split the two different clients into
> two different SocketFactory.
> In another note, the Hadoop client design decisions seems to be suited only
> for user Hadoop client (15 minutes might make sense for a single Hadoop
> client). However, Ozone decided to reuse the Hadoop client into our critical
> path without revisiting whether the default timeout makes sense in this
> context. So in the future, we need to check whether OM should use Hadoop RPC
> client for its critical path or whether we need to implement a high
> performance for OM and SCM.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]