Hi Dejan,

thanks for your answer, after playin' a bit more with stonith I figured 
out the following:

- the riloe script only distinguishes between "button" and other 
methods, so, it doesn't make a difference if the method is set to "off", 
"power", or anything else except "button"
- All methods (including "button") seem to rely on a running acpid on 
the stonith'ed machine to switch it off
- If I call stonith on the command line with "-T off" the machine is 
switched off immediately when acpid is not running, but the cluster 
seems to use another method which relies on acpid
- A reset works regardless of acpid

So, my solution is now to stop acpid and set the stonith_method to 
"reboot" instead of "poweroff".
With this, the fenced node is rebooted by the DC when the communication 
is disturbed.

Hm... I'd like to have it switched off instead of rebooting to prevent 
resources from falling back to a probably defect node, but it seems that 
this is impossible with riloe.

Now, I'd take a look at the link and configure the second stonith resource

Thanks, and regards,

Harald

Dejan Muhamedagic schrieb:
> Hi,
> 
> On Thu, Jun 04, 2009 at 06:04:04PM +0200, Husemann, Harald wrote:
>> Hi list,
>>
>> I have a Heartbeat v2 cluster build out of two HP DL380, each equipped 
>> with an iLO. Both nodes can see each other's iLO interface, and STONITH 
>> works almost as expected - when I cut off the communication of the nodes 
>> (by blocking port 629 on the DC with iptables), it sends a powerdown 
>> request to the other node's iLO.
>> The other node starts shutting down, and now things get funny: When it 
>> comes to shutdown heartbeat itself, the node wants to tell the DC that 
>> it's going down - but it can't reach the DC, so, the shutdown script 
>> goes in an endless loop.
> 
> Perhaps try with another method. Can't recall which one, but one
> of the methods depends on the system OS to do an orderly
> shutdown. You should use another ilo_powerdown_method (power is
> the default, I guess that that should be OK). BTW, it's not
> really good that fencing depends on the OS cooperation.
> 
>> When I re-enable the internal communication (i. e., drop the iptables 
>> rule), the node deregisters itself at the DC, continues the shutdown 
>> process and finally switches off the power.
>>
>> Hmmm... Any ideas how I can force the DC to *really* powering off the 
>> other node, or prevent this one from trying to deregister itself??
>>
>> Another (maybe stupid) question: I learned that I can use clone 
>> resources for stonith, and that I should do this to ease things and make 
>> the config better readable. Okay, but how to do this?? All examples I've 
>> found deal with ibmhc which has only one IP address for the Stonith 
>> device, but my iLO has different adresses for each node. Is it possible 
>> to use cloned resources here? Maybe someone can give me an example for this?
> 
> No, since you have just two nodes (right?) and the configuration
> is different for each node. You can check this for more info:
> 
> http://clusterlabs.org/mediawiki/images/f/f2/Crm_fencing.pdf
> 
> Thanks,
> 
> Dejan
> 
> 
>> Thanks + have a nice hackin',
>>
>> Harald
>>
>> Some infos of the hardware and software in use:
>>
>> HP DL380, iLO-version 1.84
>> HA Version 2.1.3, CRM Version 2.0 (CIB feature set 2.0)
>> OS CentOS 5.3 (final)
>>
>> The "interesting" parts of my CIB:
>>
>> ===================/snip/==========================
>>   <cluster_property_set id="cib-bootstrap-options">
>>           <attributes>
>>             <nvpair id="cib-bootstrap-options-dc-version" 
>> name="dc-version" value="2.1.3-node: 
>> 552305612591183b1628baa5bc6e903e0f1e26a3"/>
>>             <nvpair id="cib-bootstrap-options-last-lrm-refresh" 
>> name="last-lrm-refresh" value="1244115187"/>
>>             <nvpair id="cib-bootstrap-options-stonith-enabled" 
>> name="stonith-enabled" value="true"/>
>>             <nvpair id="cib-bootstrap-options-stonith-action" 
>> name="stonith-action" value="poweroff"/>
>>           </attributes>
>>         </cluster_property_set>
>> (...)
>>   <primitive id="rs_stonith-db1" class="stonith" type="external/riloe" 
>> provider="heartbeat">
>>           <meta_attributes id="rs_stonith-db1_meta_attrs">
>>             <attributes>
>>               <nvpair id="rs_stonith-db1_metaattr_target_role" 
>> name="target_role" value="started"/>
>>             </attributes>
>>           </meta_attributes>
>>           <instance_attributes id="rs_stonith-db1_instance_attrs">
>>             <attributes>
>>               <nvpair id="a83a7329-5cda-4dbd-9826-7e20bbda3835" 
>> name="hostlist" value="mat-db-1.***"/>
>>               <nvpair id="052b8d9c-309b-42ce-8d07-d17e612172f9" 
>> name="ilo_hostname" value="mat-db-1.***"/>
>>               <nvpair id="6da9867d-8d45-4344-8fff-4cfda8cce93d" 
>> name="ilo_user" value="stonith"/>
>>               <nvpair id="65dfb20e-0c60-41fb-bf76-ce207c8efd47" 
>> name="ilo_password" value="******"/>
>>               <nvpair id="831ff8dc-55ac-46ea-b160-f4a270dd9452" 
>> name="ilo_can_reset" value="1"/>
>>               <nvpair id="46649c31-7fbb-4213-a290-ef71e87b970b" 
>> name="ilo_protocol" value="2.0"/>
>>               <nvpair id="256ccbd8-1a7f-495e-af44-ebf4be4edca3" 
>> name="ilo_powerdown_method" value="button"/>
>>             </attributes>
>>           </instance_attributes>
>>         </primitive>
>> (...)
>> <constraints>
>>         <rsc_location id="location_stonith-db1" rsc="rs_stonith-db1">
>>           <rule id="prefered_location_stonith-db1" score="INFINITY" 
>> boolean_op="and">
>>             <expression attribute="#uname" 
>> id="b2bb572c-4e9c-499d-8cb4-9d71cc924268" operation="eq" 
>> value="mat-db-2.materna-com.de"/>
>>           </rule>
>>         </rsc_location>
>>       </constraints>
>> ============/snap/==================================================
>>
>>
>> -- 
>> Harald Husemann
>> Netzwerk- und Systemadministrator
>> Operation Management Center (OMC)
>> MATERNA GmbH
>> Information & Communications
>>
>> Westfalendamm 98
>> 44141 Dortmund
>>
>> Gesch?ftsf?hrer: Dr. Winfried Materna, Helmut an de Meulen, Ralph Hartwig
>> Amtsgericht Dortmund HRB 5839
>>
>> Tel: +49 231 9505 222
>> Fax: +49 231 9505 100
>> www.annyway.com <http://www.annyway.com/>
>> www.materna.com <http://www.materna.com/>
>> _______________________________________________
>> Linux-HA mailing list
>> [email protected]
>> http://lists.linux-ha.org/mailman/listinfo/linux-ha
>> See also: http://linux-ha.org/ReportingProblems
> _______________________________________________
> Linux-HA mailing list
> [email protected]
> http://lists.linux-ha.org/mailman/listinfo/linux-ha
> See also: http://linux-ha.org/ReportingProblems
> 

-- 
Harald Husemann
Netzwerk- und Systemadministrator
Operation Management Center (OMC)
MATERNA GmbH
Information & Communications

Westfalendamm 98
44141 Dortmund

Geschäftsführer: Dr. Winfried Materna, Helmut an de Meulen, Ralph Hartwig
Amtsgericht Dortmund HRB 5839

Tel: +49 231 9505 222
Fax: +49 231 9505 100
www.annyway.com <http://www.annyway.com/>
www.materna.com <http://www.materna.com/>
_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems

Reply via email to