[
https://issues.apache.org/jira/browse/CASSANDRA-16213?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17241853#comment-17241853
]
David Capwell commented on CASSANDRA-16213:
-------------------------------------------
[~samt] pushed changes based off your version. I made a few changes (we spoke
in slack about them)
1) rather than modify the block which checks the downed host is "old enough",
we inject the node while doing prepare; this keeps the logic centralized
2) keep the isEmpty logic and flag to enable/disable this feature. If feature
is disabled then this patch has no affect (outside of shadow round sending
empty).
[~paulo] would be good to get your eyes as well!
> Cannot replace_address /X because it doesn't exist in gossip
> ------------------------------------------------------------
>
> Key: CASSANDRA-16213
> URL: https://issues.apache.org/jira/browse/CASSANDRA-16213
> Project: Cassandra
> Issue Type: Bug
> Components: Cluster/Gossip, Cluster/Membership
> Reporter: David Capwell
> Assignee: David Capwell
> Priority: Normal
> Fix For: 4.0-beta
>
>
> We see this exception around nodes crashing and trying to do a host
> replacement; this error appears to be correlated around multiple node
> failures.
> A simplified case to trigger this is the following
> *) Have a N node cluster
> *) Shutdown all N nodes
> *) Bring up N-1 nodes (at least 1 seed, else replace seed)
> *) Host replace the N-1th node -> this will fail with the above
> The reason this happens is that the N-1th node isn’t gossiping anymore, and
> the existing nodes do not have its details in gossip (but have the details in
> the peers table), so the host replacement fails as the node isn’t known in
> gossip.
> This affects all versions (tested 3.0 and trunk, assume 2.2 as well)
--
This message was sent by Atlassian Jira
(v8.3.4#803005)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]