Maksim Davydov created IGNITE-29092:
---------------------------------------

             Summary: Fix flaky 
IgniteClusterSnapshotStreamerTest.testMetaWarningRestoredByOnlyOneNode
                 Key: IGNITE-29092
                 URL: https://issues.apache.org/jira/browse/IGNITE-29092
             Project: Ignite
          Issue Type: Test
            Reporter: Maksim Davydov
            Assignee: Maksim Davydov


IgniteClusterSnapshotStreamerTest.testMetaWarningRestoredByOnlyOneNode fails 
from time to time in the Snapshots 3 suite with:

ClusterTopologyCheckedException: Snapshot validation stopped. Required node has 
left the cluster [nodeId=[...0001]]

It fails with encryption on and off, for example:
- 
https://ci2.ignite.apache.org/buildConfiguration/IgniteTests24Java8_Snapshots3/9375047
- 
https://ci2.ignite.apache.org/buildConfiguration/IgniteTests24Java8_Snapshots3/9373027

The test stops servers 0 and 1 and then checks the snapshot from the client. 
SnapshotCheckProcess#start takes the required nodes from the local discovery 
cache of the node that starts the check, here the client 
(aliveBaselineNodes()). stopGrid() waits until each node's discovery and 
exchange versions match, but not until other nodes have processed the leave. If 
the client hasn't processed NODE_LEFT for node 1 yet, the check lists node 1 as 
required. Node 1 never answers, and the check fails. In both builds the failure 
names the server stopped last.

The race reproduces every time if the client's processing of that NODE_LEFT is 
delayed by 2 seconds: 12 of 12 executions failed with the same error.

The fix: before the check, wait until every node, the client included, sees the 
new topology (waitForTopology(3)). With the same 2-second delay, 11 of 11 
executions pass.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to