Hello,
this is a little bit off topic but maybe anybody can help me. The physical 
setup of our Debian etch heartbeat cluster looks like this:

Cluster of two nodes connected with bonding to the network, with a crossover 
connection to each other (interconnect) and with dual port Qlogic HBAs to our 
SAN. The SAN consists of two independent (physical) SANs with different VSANs 
(just to be more human readable) connected to some EMC clariions CX3-20f.

We use multipathd, LVM2 and ocfs2. This is the multipath.conf:

defaults {
        rr_min_io                       1000
        polling_interval                2
        no_path_retry                   5
        user_friendly_names             yes
}

blacklist {
        devnode "^(ram|raw|loop|fd|md|dm-|sr|scd|st)[0-9]*"
        devnode "^hd[a-z][[0-9]*]"
        devnode "^cciss!c[0-9]d[0-9]*[p[0-9]*]"
        device {
                vendor "DGC"
                product "LUNZ" #EMC Clariion management LUN
        }
        device {
                vendor "ATA"  #We do not need mutlipathing for local drives
                product "*"
        }
        device {
                vendor "AMI" # No multipathing for SUN Virtual devices
                product "*"
        }
        device {
                vendor "HITACHI" # No multipathing for local scsi disks
                product "H101414SCSUN146G"
        }
}

devices {
        ## Device attributes for EMC CLARiiON
        device {
                vendor                  "DGC"
                product                 "*"
                path_grouping_policy    group_by_prio
                getuid_callout          "/sbin/scsi_id -g -u -s /block/%n"
                prio_callout            "/sbin/mpath_prio_emc /dev/%n"
                hardware_handler        "1 emc"
                features                "1 queue_if_no_path"
                no_path_retry           fail
                path_checker            emc_clariion
                path_selector           "round-robin 0"
                failback                immediate
        }
}

multipaths {
        multipath {
                wwid                    3600601603ac511001c7c92fec775dd11
                alias                   stosan01_lun070
        }
}


And we changed the "filter" of lvm2 to not scan /dev/sd* devices. After booting 
up the system multipath -ll looks like this:

stosan01_lun070 (3600601603ac511001c7c92fec775dd11) dm-7 DGC,RAID 5
[size=133G][features=0][hwhandler=1 emc]
\_ round-robin 0 [prio=2][active]
 \_ 0:0:1:1 sdd 8:48  [active][ready]
 \_ 3:0:1:1 sdh 8:112 [active][ready]
\_ round-robin 0 [prio=0][enabled]
 \_ 0:0:0:1 sdb 8:16  [active][ready]
 \_ 3:0:0:1 sdf 8:80  [active][ready]

If we now kill one of the connections by disable the zoning on one of our SANs 
this changes to:

stosan01_lun070 (3600601603ac511001c7c92fec775dd11) dm-7 DGC,RAID 5
[size=133G][features=0][hwhandler=1 emc]
\_ round-robin 0 [prio=1][active]
 \_ 0:0:1:1 sdd 8:48  [active][ready]
\_ round-robin 0 [prio=0][enabled]
 \_ 0:0:0:1 sdb 8:16  [active][ready]

And /var/log/messages

Jun  4 14:53:58 defr1elcbtd02 kernel: o2net: no longer connected to node 
defr1elcbtd01 (num 0) at 192.168.0.101:7777
Jun  4 14:54:01 defr1elcbtd02 kernel: (4741,0):dlm_wait_for_node_death:374 
5EE89BC01EFC405E9197C198DEEAE678: waiting 5000ms for notification of death of 
node 0
Jun  4 14:54:16 defr1elcbtd02 last message repeated 3 times
Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: multipath: Failing path 
8:80.
Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: multipath: Failing path 
8:112.
Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: multipath emc: long 
trespass command will be send
Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: multipath emc: honor 
reservation bit will not be set (default)
Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: ioctl: error adding target 
to table
Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: multipath emc: long 
trespass command will be send
Jun  4 14:54:21 defr1elcbtd02 kernel: device-mapper: multipath emc: honor 
reservation bit will not be set (default)
Jun  4 14:54:21 defr1elcbtd02 kernel: (4741,0):dlm_wait_for_node_death:374 
5EE89BC01EFC405E9197C198DEEAE678: waiting 5000ms for notification of death of 
node 0
Jun  4 14:54:26 defr1elcbtd02 kernel: scsi 1:0:1:1: SCSI error: return code = 
0x00020000
Jun  4 14:54:26 defr1elcbtd02 kernel: end_request: I/O error, dev sdh, sector 
1672
Jun  4 14:54:26 defr1elcbtd02 kernel: device-mapper: multipath emc: 
emc_pg_init: sending switch-over command
Jun  4 14:54:26 defr1elcbtd02 kernel: (4741,0):dlm_wait_for_node_death:374 
5EE89BC01EFC405E9197C198DEEAE678: waiting 5000ms for notification of death of 
node 0
Jun  4 14:54:57 defr1elcbtd02 last message repeated 6 times

After 60 seconds the kernel panics because of ocfs2 is fencing this node:

o2quo_make_decision:143 ERROR: fencing this node because it is connected to a 
half-quorum of 1 out of 2 nodes which doesn't include the lowest active node 0

The o2cb settings are:
O2CB_HEARTBEAT_THRESHOLD=31
O2CB_IDLE_TIMEOUT_MS=30000
O2CB_KEEPALIVE_DELAY_MS=2000
O2CB_RECONNECT_DELAY_MS=2000

If we just set the FC Port on one of the switches to shutdown everything works 
fine. Qla2xxx detects port down, multipathd detects the missing path and ocfs2 
is not recognizing anything - just working fine like it should...

Whats going on? Is this a bug?

Regards Florian


Florian Engelmann | Unix System and Storage Administrator | IT Services Germany 
| BT (Germany) GmbH & Co. oHG | Tel:+49 (0)89 2600 8173 | Mob:+49 (0)176 1073 
1081 | Fax:+49 (0)89 2600 4955 | Email: [email protected] | Sitz und 
Registergericht der oHG: München, HRA 77639 | Persönlich haftende 
Gesellschafterin der oHG: BT Deutschland GmbH | Sitz und Registergericht: 
München, HRB 132307 | Geschäftsführung: Karsten Lereuth, Joachim Piroth, Dr. 
Sebastian Brandis, Martin Glock, Neil Smith, Judith Sasse | 
www.bt.com/globalservices

Diese E-Mail enthält Informationen von BT (Germany) GmbH & Co. oHG. Diese sind 
möglicherweise vertraulich und ausschließlich für den Adressaten bestimmt. 
Jeglicher Zugriff auf diese E-Mail durch andere Personen als den Adressaten ist 
untersagt. Sollten Sie nicht der für diese E-Mail bestimmte Adressat sein, ist 
Ihnen jede Veröffentlichung, Vervielfältigung oder Weitergabe untersagt. 
Sollten Sie diese elektronische Nachricht irrtümlicherweise erhalten haben, so 
informieren Sie uns bitte unverzüglich telefonisch oder per E-Mail 
(Kontaktdaten obenstehend). Das E-Mail-System wird von BT (Germany) GmbH & Co. 
oHG im Rahmen der gesetzlichen Bestimmungen laufend überwacht, insbesondere um 
einen störungsfreien Betrieb zu gewährleisten. Auch die vom System übertragenen 
Nachrichten werden im Rahmen der gesetzlichen Bestimmungen durch BT (Germany) 
GmbH & Co. oHG überwacht und können aufgezeichnet oder ausgewertet werden.
This electronic message contains information from BT (Germany) GmbH & Co. oHG 
which may be privileged or confidential. The information is intended to be for 
the use of the individual(s) or entity named above. If you are not the intended 
recipient be aware that any disclosure, copying, distribution or use of the 
contents of this information is prohibited. If you have received this 
electronic message in error, please notify us by telephone or e-mail (to the 
numbers or address above) immediately. Activity and use of the BT (Germany) 
GmbH & Co. oHG e-mail system is monitored to secure its effective operation and 
for other lawful business purposes. Communications using this system will also 
be monitored and may be recorded to secure effective operation and for other 
lawful business purposes.


_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems

Reply via email to