hi,

I have a problem with a heartbeat installation. After a node getting a failure 
from the LSB start script of one resource of a resource group, the node stops 
trying to start the resource group. At this point the other node will take the 
resource group as expected. But after this situation, if the other server also 
fails to start the resource, the first node doesn't ever retries to get the 
resource working.

I don't know a way to detect such a situation. If a administrator looks at the 
cluster status via crm_mon, it look like one node has the resource and the 
other stays there and wait for a failover, but when this failover actually is 
necessary it won't work because the standby node is in a state, where it won't 
take the resource group.

Is this the expected behaviour of heartbeat? Is there any way to detect that 
one node currently is not able to handle a failover?

thanks in advance
lonavera

-- 
Der GMX SmartSurfer hilft bis zu 70% Ihrer Onlinekosten zu sparen! 
Ideal für Modem und ISDN: http://www.gmx.net/de/go/smartsurfer
_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems

Reply via email to