hi, I have a problem with a heartbeat installation. After a node getting a failure from the LSB start script of one resource of a resource group, the node stops trying to start the resource group. At this point the other node will take the resource group as expected. But after this situation, if the other server also fails to start the resource, the first node doesn't ever retries to get the resource working.
I don't know a way to detect such a situation. If a administrator looks at the cluster status via crm_mon, it look like one node has the resource and the other stays there and wait for a failover, but when this failover actually is necessary it won't work because the standby node is in a state, where it won't take the resource group. Is this the expected behaviour of heartbeat? Is there any way to detect that one node currently is not able to handle a failover? thanks in advance lonavera -- Der GMX SmartSurfer hilft bis zu 70% Ihrer Onlinekosten zu sparen! Ideal für Modem und ISDN: http://www.gmx.net/de/go/smartsurfer _______________________________________________ Linux-HA mailing list [email protected] http://lists.linux-ha.org/mailman/listinfo/linux-ha See also: http://linux-ha.org/ReportingProblems
