[
https://issues.apache.org/jira/browse/HDDS-16843?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated HDDS-16843:
----------------------------------
Labels: pull-request-available (was: )
> Intermittent failure in
> TestOzoneManagerHAFollowerReadWithStoppedNodes#testLeaderOmProxyProviderFailoverOnConnectionFailure
> ---------------------------------------------------------------------------------------------------------------------------
>
> Key: HDDS-16843
> URL: https://issues.apache.org/jira/browse/HDDS-16843
> Project: Apache Ozone
> Issue Type: Sub-task
> Components: OM HA, test
> Reporter: Siyao Meng
> Assignee: Siyao Meng
> Priority: Major
> Labels: pull-request-available
>
> {code}
> Tests run: 8, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 156.3 s <<<
> FAILURE! -- in
> org.apache.hadoop.ozone.om.TestOzoneManagerHAFollowerReadWithStoppedNodes
> org.apache.hadoop.ozone.om.TestOzoneManagerHAFollowerReadWithStoppedNodes.testLeaderOmProxyProviderFailoverOnConnectionFailure
> -- Time elapsed: 15.41 s <<< ERROR!
> java.io.IOException: Could not determine or connect to OM Leader.
> at
> org.apache.hadoop.ozone.om.protocolPB.Hadoop3OmTransport.submitRequest(Hadoop3OmTransport.java:107)
> at
> org.apache.hadoop.ozone.om.protocolPB.OzoneManagerProtocolClientSideTranslatorPB.submitRequest(OzoneManagerProtocolClientSideTranslatorPB.java:385)
> at
> org.apache.hadoop.ozone.om.protocolPB.OzoneManagerProtocolClientSideTranslatorPB.createVolume(OzoneManagerProtocolClientSideTranslatorPB.java:407)
> at
> org.apache.hadoop.ozone.client.rpc.RpcClient.createVolume(RpcClient.java:473)
> at
> org.apache.hadoop.ozone.client.ObjectStore.createVolume(ObjectStore.java:125)
> at
> org.apache.hadoop.ozone.om.AbstractOzoneManagerHATest.createVolumeTest(AbstractOzoneManagerHATest.java:321)
> at
> org.apache.hadoop.ozone.om.TestOzoneManagerHAFollowerReadWithStoppedNodes.testLeaderOmProxyProviderFailoverOnConnectionFailure(TestOzoneManagerHAFollowerReadWithStoppedNodes.java:224)
> {code}
> h3. Root cause
> Same as HDDS-16589, in the follower read variant of the test. The test stops
> the current leader OM and immediately sends a write. The client gives up
> after 5 failovers while the remaining OMs are still electing a new leader,
> failing with "Could not determine or connect to OM Leader". The test also
> reads the current proxy node ID before the first request, so it may not point
> at the leader yet.
> h3. Fix
> Apply the HDDS-16589 change to
> TestOzoneManagerHAFollowerReadWithStoppedNodes: read the proxy node ID after
> the first createVolumeTest and call waitForLeaderToBeReady() after stopping
> the OM.
> - https://github.com/apache/ozone/actions/runs/38025121510/job/114136161996
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]