Maksim Davydov created IGNITE-29088:
---------------------------------------

             Summary: Coordinator keeps a lost partition MOVING after its new 
primary owns it (IGNORE loss policy)
                 Key: IGNITE-29088
                 URL: https://issues.apache.org/jira/browse/IGNITE-29088
             Project: Ignite
          Issue Type: Bug
            Reporter: Maksim Davydov
            Assignee: Maksim Davydov


*Problem*

In an in-memory cluster with the IGNORE partition loss policy (baseline 
auto-adjust on, timeout 0), a node leaves and takes the only copy of some 
partitions with it (backups = 0). Each lost partition is recreated empty on its 
new primary and owned there. Sometimes the coordinator never learns that: its 
partition map keeps the partition MOVING until the next exchange.

*Impact*

Until the next topology change, which in a stable cluster can be hours away, 
the coordinator and every node that takes its full map see no owner of the 
partition:
- SQL queries from those nodes fail after the retry timeout:
  "Failed to map SQL query to topology during timeout: 30000ms". With a MOVING 
partition in the map, ReducePartitionMapper maps partitions by their owners and 
finds none. On the new primary, whose own map is right, the same query succeeds.
- awaitPartitionMapExchange() in tests times out. Key operations and scan 
queries still work.

This is not what IGNORE is meant to do: the policy resets a lost partition 
silently, no
partition is LOST, and there is nothing to reset by hand.

*Cause*

With the exchange merge protocol, the new primary creates the lost partition as 
MOVING when it gets the coordinator's full message, and owns it a moment later 
in detectLostPartitions. The coordinator, in its own detectLostPartitions, 
already marks the partition OWNING for that node.

If the new primary sends its partition map between these two steps, the map 
carries MOVING with a newer update sequence, and the coordinator takes it. A 
typical sender is the scheduled resend (scheduleResendPartitions(), 1.5 s after 
an eviction or a map change). Nothing sends the node's map again after it owns 
the partition: the result of detectLostPartitions is dropped in 
GridDhtPartitionsExchangeFuture#detectLostPartitions. 

Also, inside GridDhtPartitionTopologyImpl#detectLostPartitions the result 
reflects only the last lost partition.

*How it shows*

IgniteTopologyValidatorGridSplitCacheTest (32 nodes, 50 caches without backups) 
times out in awaitPartitionMapExchange() after stopping its configless node in 
9 of 15 local runs. Small clusters rarely send a map inside that window, so 
IgniteCachePartitionLossPolicySelfTest (3 nodes) doesn't hit it.

*Fix*

When detectLostPartitions changes a local partition, a non-coordinator node 
sends its single map again. A new test, CachePartitionLossIgnorePolicyMapTest, 
makes the new primary send its map inside the window every time and checks that 
the coordinator's map matches the nodes.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to