On Wed, May 12, 2010 at 10:59:43AM -0700, Mike Sweetser wrote:
> > What is before this?
> > Below is "MCP dead" (Master Control Process)...
> > it should log why it died.
> > Or there should be some core file below
> >        find /var/lib/heartbeat/cores/
> > Or both.
> >
> >
> May 11 17:38:33 mysql1 crmd: [904]: notice: run_graph: Transition 1
> (Complete=0, Pending=0, Fired=0, Skipped=0, Incomplete=0,
> Source=/var/lib/pengine/pe-input-2440.bz2): Complete
> May 11 17:38:33 mysql1 crmd: [904]: info: te_graph_trigger: Transition 1 is
> now complete
> May 11 17:38:33 mysql1 crmd: [904]: info: notify_crmd: Transition 1 status:
> done - <null>
> May 11 17:38:33 mysql1 crmd: [904]: info: do_state_transition: State
> transition S_TRANSITION_ENGINE -> S_IDLE [ input=I_TE_SUCCESS
> cause=C_FSA_INTERNAL origin=notify_crmd ]
> May 11 17:38:33 mysql1 crmd: [904]: info: do_state_transition: Starting
> PEngine Recheck Timer
> May 11 17:38:33 mysql1 pengine: [23821]: info: process_pe_message:
> Transition 1: PEngine Input stored in: /var/lib/pengine/pe-input-2440.bz2
> May 11 17:45:28 mysql1 cib: [900]: info: cib_stats: Processed 1 operations
> (0.00us average, 0% utilization) in the last 10min
> 
> That's before all those messages. Right before that, it actually said it

and "those messages" are those that are interessting.  above is boring ;-)

> lost a connection with the other server, but it came back right away.
> 
> 
> > > May 08 05:33:19 mysql1 heartbeat: [5536]: CRIT: Killing pid 5533 with

Messages before _this_ one, indicating _why_ the MCP may have crashed
would be interessting.

> > > SIGTERM
> > > May 08 05:33:19 mysql1 heartbeat: [5536]: CRIT: Killing pid 5537 with
> > > SIGTERM
> > > May 08 05:33:19 mysql1 heartbeat: [5536]: CRIT: Killing pid 5538 with
> > > SIGTERM
> > > May 08 05:33:19 mysql1 heartbeat: [5536]: CRIT: Killing pid 5539 with
> > > SIGTERM
> > > May 08 05:33:19 mysql1 heartbeat: [5536]: CRIT: Killing pid 5540 with
> > > SIGTERM
> > > May 08 05:33:19 mysql1 heartbeat: [5536]: CRIT: Emergency Shutdown(MCP
> > > dead): Killing ourselves.

probably best to bundle up a hb_report of the time frame from good
before "those messages" have started, until this "Emergency Shutdown"
incident.  And send that to your support contact.
You do have a support contact, right?

 ;-)

-- 
: Lars Ellenberg
: LINBIT | Your Way to High Availability
: DRBD/HA support and consulting http://www.linbit.com

DRBD® and LINBIT® are registered trademarks of LINBIT, Austria.
_______________________________________________
Linux-HA mailing list
[email protected]
http://lists.linux-ha.org/mailman/listinfo/linux-ha
See also: http://linux-ha.org/ReportingProblems

Reply via email to