Hi all,
I'm setting up for a friend a 2-nodes cluster using
heartbeat-2.1.3-3.el5.centos.i386.rpm on Centos 5.1 32bit.
It's serving drbd, mysql, and tomcat.
Recently he notices that the application (tomcat) is not responding on the
main node, and he tries to migrate the resources to the backup node by
unplugging the pingd network card.
Heartbeat then tries to move the resources to the backup node, but when it
tries to stop tomcat, it fails (time out):
Sep 7 09:42:55 ibmsrv crmd: [2611]: info: do_lrm_rsc_op: Performing
op=r_tomcat0_stop_0 key=13:11:6b49e2d2-a0ad-492b-beff-5325bacb3e67)
Sep 7 09:42:55 ibmsrv lrmd: [2608]: info: rsc:r_tomcat0: stop
Sep 7 09:43:15 ibmsrv lrmd: [2608]: WARN: r_tomcat0:stop process (PID
21972) timed out (try 1). Killing with signal SIGTERM (15).
Sep 7 09:43:15 ibmsrv lrmd: [2608]: WARN: operation stop[12] on
ocf::tomcat::r_tomcat0 for client 2611, its parameters:
catalina_home=[/opt
/tomcat] CRM_meta_role=[Started] CRM_meta_start_delay=[0]
CATALINA_PID=[/var/run/catalina.pid]
CRM_meta_id=[9cfae9ce-dd9f-42be-9a8b-b076b417
ce03] CRM_meta_timeout=[20000] CRM_meta_on_fail=[restart]
crm_feature_set=[2.0] java_home=[/usr/java/jre1.6.0_10]
CRM_meta_disabled=[false]
CRM_meta_name=[start] : pid [21972] timed out
Sep 7 09:43:15 ibmsrv crmd: [2611]: ERROR: process_lrm_event: LRM
operation r_tomcat0_stop_0 (12) Timed Out (timeout=20000ms)
Then the backup node tries to startup tomcat, but fails:
Sep 7 09:43:16 srvbk pengine: [3860]: notice: StartRsc: srvbk.baru.com
Start r_tomcat0
Sep 7 09:43:16 srvbk pengine: [3860]: WARN: should_dump_action: action 13
(r_tomcat0_stop_0) was for an unmanaged resource (r_tomcat0)
Sep 7 09:43:16 srvbk pengine: [3860]: WARN: should_dump_action: action 14
(r_tomcat0_start_0) was for an unmanaged resource (r_tomcat0)
Sep 7 09:43:16 srvbk pengine: [3860]: WARN: should_dump_action: action 13
(r_tomcat0_stop_0) was for an unmanaged resource (r_tomcat0)
Sep 7 09:43:16 srvbk pengine: [3860]: WARN: should_dump_action: action 14
(r_tomcat0_start_0) was for an unmanaged resource (r_tomcat0)
Sep 7 09:43:16 srvbk pengine: [3860]: WARN: should_dump_action: action 13
(r_tomcat0_stop_0) was for an unmanaged resource (r_tomcat0)
Sep 7 09:43:16 srvbk pengine: [3860]: WARN: should_dump_action: action 14
(r_tomcat0_start_0) was for an unmanaged resource (r_tomcat0)
He then rebooted the main node, after this then heartbeat manages to start
all resources including tomcat in main node.
Is there anything I can do to handle this case?
Thank you,
--
Fajar Priyanto | Reg'd Linux User #327841 | Linux tutorial
http://linux2.arinet.org
13:10:54 up 5:02, 2.6.24-18-generic GNU/Linux
Let's use OpenOffice. http://www.openoffice.org
The real challenge of teaching is getting your students motivated to learn.
_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems