Hi, On Fri, Jun 05, 2009 at 04:21:59PM +0200, Husemann, Harald wrote: > Hi Dejan, > > thanks for your answer, after playin' a bit more with stonith I figured > out the following: > > - the riloe script only distinguishes between "button" and other > methods, so, it doesn't make a difference if the method is set to "off", > "power", or anything else except "button" > - All methods (including "button") seem to rely on a running acpid on > the stonith'ed machine to switch it off
Yes, button is sort of "quick power button press" which is handled by acpi. "power" should be more like pulling the power plug. > - If I call stonith on the command line with "-T off" the machine is > switched off immediately when acpid is not running, but the cluster > seems to use another method which relies on acpid stonith on command line should behave in exactly the same way like stonithd. > - A reset works regardless of acpid Oh, you're still running 2.1.3. external/riloe was completely rewritten in the meantime. Definitely upgrade to 2.1.4. Or at least update the plugin. > So, my solution is now to stop acpid and set the stonith_method to > "reboot" instead of "poweroff". > With this, the fenced node is rebooted by the DC when the communication > is disturbed. > > Hm... I'd like to have it switched off instead of rebooting to prevent > resources from falling back to a probably defect node, but it seems that > this is impossible with riloe. It should work with the new version. Thanks, Dejan > Now, I'd take a look at the link and configure the second stonith resource > > Thanks, and regards, > > Harald > > Dejan Muhamedagic schrieb: > > Hi, > > > > On Thu, Jun 04, 2009 at 06:04:04PM +0200, Husemann, Harald wrote: > >> Hi list, > >> > >> I have a Heartbeat v2 cluster build out of two HP DL380, each equipped > >> with an iLO. Both nodes can see each other's iLO interface, and STONITH > >> works almost as expected - when I cut off the communication of the nodes > >> (by blocking port 629 on the DC with iptables), it sends a powerdown > >> request to the other node's iLO. > >> The other node starts shutting down, and now things get funny: When it > >> comes to shutdown heartbeat itself, the node wants to tell the DC that > >> it's going down - but it can't reach the DC, so, the shutdown script > >> goes in an endless loop. > > > > Perhaps try with another method. Can't recall which one, but one > > of the methods depends on the system OS to do an orderly > > shutdown. You should use another ilo_powerdown_method (power is > > the default, I guess that that should be OK). BTW, it's not > > really good that fencing depends on the OS cooperation. > > > >> When I re-enable the internal communication (i. e., drop the iptables > >> rule), the node deregisters itself at the DC, continues the shutdown > >> process and finally switches off the power. > >> > >> Hmmm... Any ideas how I can force the DC to *really* powering off the > >> other node, or prevent this one from trying to deregister itself?? > >> > >> Another (maybe stupid) question: I learned that I can use clone > >> resources for stonith, and that I should do this to ease things and make > >> the config better readable. Okay, but how to do this?? All examples I've > >> found deal with ibmhc which has only one IP address for the Stonith > >> device, but my iLO has different adresses for each node. Is it possible > >> to use cloned resources here? Maybe someone can give me an example for > >> this? > > > > No, since you have just two nodes (right?) and the configuration > > is different for each node. You can check this for more info: > > > > http://clusterlabs.org/mediawiki/images/f/f2/Crm_fencing.pdf > > > > Thanks, > > > > Dejan > > > > > >> Thanks + have a nice hackin', > >> > >> Harald > >> > >> Some infos of the hardware and software in use: > >> > >> HP DL380, iLO-version 1.84 > >> HA Version 2.1.3, CRM Version 2.0 (CIB feature set 2.0) > >> OS CentOS 5.3 (final) > >> > >> The "interesting" parts of my CIB: > >> > >> ===================/snip/========================== > >> <cluster_property_set id="cib-bootstrap-options"> > >> <attributes> > >> <nvpair id="cib-bootstrap-options-dc-version" > >> name="dc-version" value="2.1.3-node: > >> 552305612591183b1628baa5bc6e903e0f1e26a3"/> > >> <nvpair id="cib-bootstrap-options-last-lrm-refresh" > >> name="last-lrm-refresh" value="1244115187"/> > >> <nvpair id="cib-bootstrap-options-stonith-enabled" > >> name="stonith-enabled" value="true"/> > >> <nvpair id="cib-bootstrap-options-stonith-action" > >> name="stonith-action" value="poweroff"/> > >> </attributes> > >> </cluster_property_set> > >> (...) > >> <primitive id="rs_stonith-db1" class="stonith" type="external/riloe" > >> provider="heartbeat"> > >> <meta_attributes id="rs_stonith-db1_meta_attrs"> > >> <attributes> > >> <nvpair id="rs_stonith-db1_metaattr_target_role" > >> name="target_role" value="started"/> > >> </attributes> > >> </meta_attributes> > >> <instance_attributes id="rs_stonith-db1_instance_attrs"> > >> <attributes> > >> <nvpair id="a83a7329-5cda-4dbd-9826-7e20bbda3835" > >> name="hostlist" value="mat-db-1.***"/> > >> <nvpair id="052b8d9c-309b-42ce-8d07-d17e612172f9" > >> name="ilo_hostname" value="mat-db-1.***"/> > >> <nvpair id="6da9867d-8d45-4344-8fff-4cfda8cce93d" > >> name="ilo_user" value="stonith"/> > >> <nvpair id="65dfb20e-0c60-41fb-bf76-ce207c8efd47" > >> name="ilo_password" value="******"/> > >> <nvpair id="831ff8dc-55ac-46ea-b160-f4a270dd9452" > >> name="ilo_can_reset" value="1"/> > >> <nvpair id="46649c31-7fbb-4213-a290-ef71e87b970b" > >> name="ilo_protocol" value="2.0"/> > >> <nvpair id="256ccbd8-1a7f-495e-af44-ebf4be4edca3" > >> name="ilo_powerdown_method" value="button"/> > >> </attributes> > >> </instance_attributes> > >> </primitive> > >> (...) > >> <constraints> > >> <rsc_location id="location_stonith-db1" rsc="rs_stonith-db1"> > >> <rule id="prefered_location_stonith-db1" score="INFINITY" > >> boolean_op="and"> > >> <expression attribute="#uname" > >> id="b2bb572c-4e9c-499d-8cb4-9d71cc924268" operation="eq" > >> value="mat-db-2.materna-com.de"/> > >> </rule> > >> </rsc_location> > >> </constraints> > >> ============/snap/================================================== > >> > >> > >> -- > >> Harald Husemann > >> Netzwerk- und Systemadministrator > >> Operation Management Center (OMC) > >> MATERNA GmbH > >> Information & Communications > >> > >> Westfalendamm 98 > >> 44141 Dortmund > >> > >> Gesch?ftsf?hrer: Dr. Winfried Materna, Helmut an de Meulen, Ralph Hartwig > >> Amtsgericht Dortmund HRB 5839 > >> > >> Tel: +49 231 9505 222 > >> Fax: +49 231 9505 100 > >> www.annyway.com <http://www.annyway.com/> > >> www.materna.com <http://www.materna.com/> > >> _______________________________________________ > >> Linux-HA mailing list > >> [email protected] > >> http://lists.linux-ha.org/mailman/listinfo/linux-ha > >> See also: http://linux-ha.org/ReportingProblems > > _______________________________________________ > > Linux-HA mailing list > > [email protected] > > http://lists.linux-ha.org/mailman/listinfo/linux-ha > > See also: http://linux-ha.org/ReportingProblems > > > > -- > Harald Husemann > Netzwerk- und Systemadministrator > Operation Management Center (OMC) > MATERNA GmbH > Information & Communications > > Westfalendamm 98 > 44141 Dortmund > > Gesch?ftsf?hrer: Dr. Winfried Materna, Helmut an de Meulen, Ralph Hartwig > Amtsgericht Dortmund HRB 5839 > > Tel: +49 231 9505 222 > Fax: +49 231 9505 100 > www.annyway.com <http://www.annyway.com/> > www.materna.com <http://www.materna.com/> > _______________________________________________ > Linux-HA mailing list > [email protected] > http://lists.linux-ha.org/mailman/listinfo/linux-ha > See also: http://linux-ha.org/ReportingProblems _______________________________________________ Linux-HA mailing list [email protected] http://lists.linux-ha.org/mailman/listinfo/linux-ha See also: http://linux-ha.org/ReportingProblems
