[ 
https://issues.apache.org/jira/browse/HDDS-16780?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Siyao Meng updated HDDS-16780:
------------------------------
    Description: 
{code}
[ERROR] Tests run: 3, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 82.46 s 
<<< FAILURE! -- in 
org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication
[ERROR] 
org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication.testDecommissionWithPerVolumePools
 -- Time elapsed: 31.25 s <<< ERROR!
java.util.concurrent.TimeoutException:
Timed out waiting for condition.
        at 
org.apache.ozone.test.GenericTestUtils.waitFor(GenericTestUtils.java:137)
        at 
org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication.testDecommissionWithPerVolumePools(TestPerVolumePushReplication.java:271)
{code}

SCM log during the decommission (197 occurrences):

{code}
ERROR replication.UnhealthyReplicationProcessor 
(UnhealthyReplicationProcessor.java:processAll(125)) - Error processing Health 
result of class: class 
org.apache.hadoop.hdds.scm.container.replication.ContainerHealthResult$UnderReplicatedHealthResult
 ...
java.lang.IllegalArgumentException: Capacity cannot be negative.
        at 
org.apache.hadoop.hdds.scm.container.placement.metrics.SCMNodeStat.<init>(SCMNodeStat.java:45)
        at 
org.apache.hadoop.hdds.scm.node.SCMNodeManager.getNodeStatInternal(SCMNodeManager.java:1172)
        at 
org.apache.hadoop.hdds.scm.node.SCMNodeManager.getNodeStat(SCMNodeManager.java:1147)
        at 
org.apache.hadoop.hdds.scm.container.replication.ReplicationManagerUtil.excludeFullNodes(ReplicationManagerUtil.java:245)
        at 
org.apache.hadoop.hdds.scm.container.replication.ReplicationManagerUtil.getExcludedAndUsedNodes(ReplicationManagerUtil.java:223)
        at 
org.apache.hadoop.hdds.scm.container.replication.ECUnderReplicationHandler.processAndSendCommands(ECUnderReplicationHandler.java:130)
{code}

h3. Root cause

The integration test {{ozone-site.xml}} sets 
{{hdds.datanode.du.factory.classname}} to {{MockSpaceUsageCheckFactory$None}}, 
which reports {{Long.MAX_VALUE}} capacity for every volume. 
{{TestPerVolumePushReplication}} starts datanodes with 2 data volumes, so 
{{SCMNodeManager#getNodeStatInternal}} overflows when it sums the storage 
reports of a node, and {{SCMNodeStat}} throws {{Capacity cannot be negative}}.

{{ReplicationManagerUtil#excludeFullNodes}} calls {{getNodeStat}} for every 
datanode with a pending replica ADD. So while any replication is in flight, 
processing of every other under replicated container fails, and the 
decommission effectively replicates one container per ReplicationManager cycle 
(about 3s). The decommissioning node held 12 containers; it had reached 9 of 
them when the 30s {{waitForDnToReachOpState}} wait expired.

h3. Fix

The same overflow affects any mini cluster test with more than one volume per 
datanode, and {{SCMNodeManager#getStats}}, which sums over all datanodes. Lower 
{{MockSpaceUsageSource.unlimited()}} to {{Long.MAX_VALUE / 1024}} (about 8 PB 
per volume), which leaves room to sum 1024 volumes without overflow.

h3. Second failure: testHealthyVolumeReplicationAfterVolumeFailure

{code}
[ERROR] 
org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication.testHealthyVolumeReplicationAfterVolumeFailure
 -- Time elapsed: 35.37 s <<< ERROR!
java.util.concurrent.TimeoutException:
        at 
org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication.queuePushAndWaitForContainer(TestPerVolumePushReplication.java:387)
        at 
org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication.testHealthyVolumeReplicationAfterVolumeFailure(TestPerVolumePushReplication.java:220)
{code}

{code}
WARN  statemachine.StateContext (StateContext.java:getNextCommand(767)) - 
Detect and drop a SCMCommand replicateContainerCommand: ... from stale leader 
SCM, stale term 0, latest term 2.
{code}

The test can pick the datanode that the previous test just restarted as the 
push source. That datanode has not learned the SCM leader term yet, so 
{{queueReplicationCommand}} left the command at term 0, and {{StateContext}} 
dropped it as stale once the first real SCM command arrived. Fix: set the term 
from {{SCMContext#getTermOfLeader}}, as {{TestCloseContainerByPipeline}} does.

- https://github.com/apache/ozone/actions/runs/37686266652/job/113019002301
- https://github.com/apache/ozone/actions/runs/37927673352/job/113815428728

  was:
{code}
[ERROR] Tests run: 3, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 82.46 s 
<<< FAILURE! -- in 
org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication
[ERROR] 
org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication.testDecommissionWithPerVolumePools
 -- Time elapsed: 31.25 s <<< ERROR!
java.util.concurrent.TimeoutException:
Timed out waiting for condition.
        at 
org.apache.ozone.test.GenericTestUtils.waitFor(GenericTestUtils.java:137)
        at 
org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication.testDecommissionWithPerVolumePools(TestPerVolumePushReplication.java:271)
{code}

SCM log during the decommission (197 occurrences):

{code}
ERROR replication.UnhealthyReplicationProcessor 
(UnhealthyReplicationProcessor.java:processAll(125)) - Error processing Health 
result of class: class 
org.apache.hadoop.hdds.scm.container.replication.ContainerHealthResult$UnderReplicatedHealthResult
 ...
java.lang.IllegalArgumentException: Capacity cannot be negative.
        at 
org.apache.hadoop.hdds.scm.container.placement.metrics.SCMNodeStat.<init>(SCMNodeStat.java:45)
        at 
org.apache.hadoop.hdds.scm.node.SCMNodeManager.getNodeStatInternal(SCMNodeManager.java:1172)
        at 
org.apache.hadoop.hdds.scm.node.SCMNodeManager.getNodeStat(SCMNodeManager.java:1147)
        at 
org.apache.hadoop.hdds.scm.container.replication.ReplicationManagerUtil.excludeFullNodes(ReplicationManagerUtil.java:245)
        at 
org.apache.hadoop.hdds.scm.container.replication.ReplicationManagerUtil.getExcludedAndUsedNodes(ReplicationManagerUtil.java:223)
        at 
org.apache.hadoop.hdds.scm.container.replication.ECUnderReplicationHandler.processAndSendCommands(ECUnderReplicationHandler.java:130)
{code}

h3. Root cause

The integration test {{ozone-site.xml}} sets 
{{hdds.datanode.du.factory.classname}} to {{MockSpaceUsageCheckFactory$None}}, 
which reports {{Long.MAX_VALUE}} capacity for every volume. 
{{TestPerVolumePushReplication}} starts datanodes with 2 data volumes, so 
{{SCMNodeManager#getNodeStatInternal}} overflows when it sums the storage 
reports of a node, and {{SCMNodeStat}} throws {{Capacity cannot be negative}}.

{{ReplicationManagerUtil#excludeFullNodes}} calls {{getNodeStat}} for every 
datanode with a pending replica ADD. So while any replication is in flight, 
processing of every other under replicated container fails, and the 
decommission effectively replicates one container per ReplicationManager cycle 
(about 3s). The decommissioning node held 12 containers; it had reached 9 of 
them when the 30s {{waitForDnToReachOpState}} wait expired.

h3. Fix

The same overflow affects any mini cluster test with more than one volume per 
datanode, and {{SCMNodeManager#getStats}}, which sums over all datanodes. Lower 
{{MockSpaceUsageSource.unlimited()}} to {{Long.MAX_VALUE / 1024}} (about 8 PB 
per volume), which leaves room to sum 1024 volumes without overflow.

- https://github.com/apache/ozone/actions/runs/37686266652/job/113019002301

        Summary: Intermittent failures in TestPerVolumePushReplication  (was: 
Intermittent failure in 
TestPerVolumePushReplication#testDecommissionWithPerVolumePools)

> Intermittent failures in TestPerVolumePushReplication
> -----------------------------------------------------
>
>                 Key: HDDS-16780
>                 URL: https://issues.apache.org/jira/browse/HDDS-16780
>             Project: Apache Ozone
>          Issue Type: Sub-task
>          Components: SCM, test
>            Reporter: Siyao Meng
>            Assignee: Siyao Meng
>            Priority: Major
>              Labels: pull-request-available
>
> {code}
> [ERROR] Tests run: 3, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 82.46 
> s <<< FAILURE! -- in 
> org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication
> [ERROR] 
> org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication.testDecommissionWithPerVolumePools
>  -- Time elapsed: 31.25 s <<< ERROR!
> java.util.concurrent.TimeoutException:
> Timed out waiting for condition.
>       at 
> org.apache.ozone.test.GenericTestUtils.waitFor(GenericTestUtils.java:137)
>       at 
> org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication.testDecommissionWithPerVolumePools(TestPerVolumePushReplication.java:271)
> {code}
> SCM log during the decommission (197 occurrences):
> {code}
> ERROR replication.UnhealthyReplicationProcessor 
> (UnhealthyReplicationProcessor.java:processAll(125)) - Error processing 
> Health result of class: class 
> org.apache.hadoop.hdds.scm.container.replication.ContainerHealthResult$UnderReplicatedHealthResult
>  ...
> java.lang.IllegalArgumentException: Capacity cannot be negative.
>       at 
> org.apache.hadoop.hdds.scm.container.placement.metrics.SCMNodeStat.<init>(SCMNodeStat.java:45)
>       at 
> org.apache.hadoop.hdds.scm.node.SCMNodeManager.getNodeStatInternal(SCMNodeManager.java:1172)
>       at 
> org.apache.hadoop.hdds.scm.node.SCMNodeManager.getNodeStat(SCMNodeManager.java:1147)
>       at 
> org.apache.hadoop.hdds.scm.container.replication.ReplicationManagerUtil.excludeFullNodes(ReplicationManagerUtil.java:245)
>       at 
> org.apache.hadoop.hdds.scm.container.replication.ReplicationManagerUtil.getExcludedAndUsedNodes(ReplicationManagerUtil.java:223)
>       at 
> org.apache.hadoop.hdds.scm.container.replication.ECUnderReplicationHandler.processAndSendCommands(ECUnderReplicationHandler.java:130)
> {code}
> h3. Root cause
> The integration test {{ozone-site.xml}} sets 
> {{hdds.datanode.du.factory.classname}} to 
> {{MockSpaceUsageCheckFactory$None}}, which reports {{Long.MAX_VALUE}} 
> capacity for every volume. {{TestPerVolumePushReplication}} starts datanodes 
> with 2 data volumes, so {{SCMNodeManager#getNodeStatInternal}} overflows when 
> it sums the storage reports of a node, and {{SCMNodeStat}} throws {{Capacity 
> cannot be negative}}.
> {{ReplicationManagerUtil#excludeFullNodes}} calls {{getNodeStat}} for every 
> datanode with a pending replica ADD. So while any replication is in flight, 
> processing of every other under replicated container fails, and the 
> decommission effectively replicates one container per ReplicationManager 
> cycle (about 3s). The decommissioning node held 12 containers; it had reached 
> 9 of them when the 30s {{waitForDnToReachOpState}} wait expired.
> h3. Fix
> The same overflow affects any mini cluster test with more than one volume per 
> datanode, and {{SCMNodeManager#getStats}}, which sums over all datanodes. 
> Lower {{MockSpaceUsageSource.unlimited()}} to {{Long.MAX_VALUE / 1024}} 
> (about 8 PB per volume), which leaves room to sum 1024 volumes without 
> overflow.
> h3. Second failure: testHealthyVolumeReplicationAfterVolumeFailure
> {code}
> [ERROR] 
> org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication.testHealthyVolumeReplicationAfterVolumeFailure
>  -- Time elapsed: 35.37 s <<< ERROR!
> java.util.concurrent.TimeoutException:
>       at 
> org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication.queuePushAndWaitForContainer(TestPerVolumePushReplication.java:387)
>       at 
> org.apache.hadoop.ozone.container.replication.TestPerVolumePushReplication.testHealthyVolumeReplicationAfterVolumeFailure(TestPerVolumePushReplication.java:220)
> {code}
> {code}
> WARN  statemachine.StateContext (StateContext.java:getNextCommand(767)) - 
> Detect and drop a SCMCommand replicateContainerCommand: ... from stale leader 
> SCM, stale term 0, latest term 2.
> {code}
> The test can pick the datanode that the previous test just restarted as the 
> push source. That datanode has not learned the SCM leader term yet, so 
> {{queueReplicationCommand}} left the command at term 0, and {{StateContext}} 
> dropped it as stale once the first real SCM command arrived. Fix: set the 
> term from {{SCMContext#getTermOfLeader}}, as {{TestCloseContainerByPipeline}} 
> does.
> - https://github.com/apache/ozone/actions/runs/37686266652/job/113019002301
> - https://github.com/apache/ozone/actions/runs/37927673352/job/113815428728



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to