On Wed, May 12, 2010 at 10:59:43AM -0700, Mike Sweetser wrote: > > What is before this? > > Below is "MCP dead" (Master Control Process)... > > it should log why it died. > > Or there should be some core file below > > find /var/lib/heartbeat/cores/ > > Or both. > > > > > May 11 17:38:33 mysql1 crmd: [904]: notice: run_graph: Transition 1 > (Complete=0, Pending=0, Fired=0, Skipped=0, Incomplete=0, > Source=/var/lib/pengine/pe-input-2440.bz2): Complete > May 11 17:38:33 mysql1 crmd: [904]: info: te_graph_trigger: Transition 1 is > now complete > May 11 17:38:33 mysql1 crmd: [904]: info: notify_crmd: Transition 1 status: > done - <null> > May 11 17:38:33 mysql1 crmd: [904]: info: do_state_transition: State > transition S_TRANSITION_ENGINE -> S_IDLE [ input=I_TE_SUCCESS > cause=C_FSA_INTERNAL origin=notify_crmd ] > May 11 17:38:33 mysql1 crmd: [904]: info: do_state_transition: Starting > PEngine Recheck Timer > May 11 17:38:33 mysql1 pengine: [23821]: info: process_pe_message: > Transition 1: PEngine Input stored in: /var/lib/pengine/pe-input-2440.bz2 > May 11 17:45:28 mysql1 cib: [900]: info: cib_stats: Processed 1 operations > (0.00us average, 0% utilization) in the last 10min > > That's before all those messages. Right before that, it actually said it
and "those messages" are those that are interessting. above is boring ;-) > lost a connection with the other server, but it came back right away. > > > > > May 08 05:33:19 mysql1 heartbeat: [5536]: CRIT: Killing pid 5533 with Messages before _this_ one, indicating _why_ the MCP may have crashed would be interessting. > > > SIGTERM > > > May 08 05:33:19 mysql1 heartbeat: [5536]: CRIT: Killing pid 5537 with > > > SIGTERM > > > May 08 05:33:19 mysql1 heartbeat: [5536]: CRIT: Killing pid 5538 with > > > SIGTERM > > > May 08 05:33:19 mysql1 heartbeat: [5536]: CRIT: Killing pid 5539 with > > > SIGTERM > > > May 08 05:33:19 mysql1 heartbeat: [5536]: CRIT: Killing pid 5540 with > > > SIGTERM > > > May 08 05:33:19 mysql1 heartbeat: [5536]: CRIT: Emergency Shutdown(MCP > > > dead): Killing ourselves. probably best to bundle up a hb_report of the time frame from good before "those messages" have started, until this "Emergency Shutdown" incident. And send that to your support contact. You do have a support contact, right? ;-) -- : Lars Ellenberg : LINBIT | Your Way to High Availability : DRBD/HA support and consulting http://www.linbit.com DRBD® and LINBIT® are registered trademarks of LINBIT, Austria. _______________________________________________ Linux-HA mailing list [email protected] http://lists.linux-ha.org/mailman/listinfo/linux-ha See also: http://linux-ha.org/ReportingProblems
