public inbox for linux-kernel@vger.kernel.org
 help / color / mirror / Atom feed
From: Roger Heflin <rheflin@atipa.com>
To: Attila Nagy <bra@fsn.hu>
Cc: Alan Cox <alan@lxorguk.ukuu.org.uk>, linux-kernel@vger.kernel.org
Subject: Re: Hangs and reboots under high loads, oops with DEBUG_SHIRQ
Date: Tue, 31 Jul 2007 17:08:11 -0500	[thread overview]
Message-ID: <46AFB2CB.6040906@atipa.com> (raw)
In-Reply-To: <46AF4570.7070606@fsn.hu>

Attila Nagy wrote:
> On 2007.07.30. 18:19, Alan Cox wrote:
>> O> MCE:
>>  
>>> [153103.918654] HARDWARE ERROR
>>> [153103.918655] CPU 1: Machine Check Exception:                5 Bank 
>>> 0: b200004010000400
>>> [153104.066037] RIP !INEXACT! 10:<ffffffff802569e6> 
>>> {mwait_idle+0x46/0x60}
>>> [153104.145699] TSC 1167e915e93ce
>>> [153104.183554] This is not a software problem!
>>> [153104.234724] Run through mcelog --ascii to decode and contact your 
>>> hardware vendor
>>>     
>>
>> If you it through mcelog as it suggests it wil decode the meaning of the
>> MCE data and that should give you some idea. Generally speaking MCE
>> errors are real hardware errors but can certainly be caused by external
>> factors (power supply glitches, heat etc)
>>   
> Sorry, of course I ran that through mcelog, but inadvertently attached 
> the original version.
> 
> I've tried the machines with two types of power sources (different 
> UPSes, line filtering, etc,
> and the chassis have redundant PSes), monitoring the temperatures (seems 
> to be OK,
> the CPUs don't go over 30 °C even under load). I have the latest BIOS 
> for the
> motherboard.
> But I will recheck everything.
> 
> BTW, here's the output from mcelog, I see this occasionally on all four 
> machines:
> 
> HARDWARE ERROR
> HARDWARE ERROR. This is *NOT* a software problem!
> Please contact your hardware vendor
> CPU 1 BANK 0 TSC 1167e915e93ce
> MCG status:RIPV MCIP
> MCi status:
> Uncorrected error
> Error enabled
> Processor context corrupt
> MCA: Internal Timer error
> STATUS b200004010000400 MCGSTATUS 5
> This is not a software problem!
> Run through mcelog --ascii to decode and contact your hardware vendor
> 
> HARDWARE ERROR
> HARDWARE ERROR. This is *NOT* a software problem!
> Please contact your hardware vendor
> CPU 1 BANK 5 TSC 1167e915e9ea8
> MCG status:RIPV MCIP
> MCi status:
> Uncorrected error
> Error enabled
> Processor context corrupt
> MCA: Internal Timer error
> STATUS b200221024080400 MCGSTATUS 5
> This is not a software problem!
> Run through mcelog --ascii to decode and contact your hardware vendor
> 
> Thanks,
> 
> -
> To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html
> Please read the FAQ at  http://www.tux.org/lkml/
> 

Attila,

We had some issues with very similar boards all of the problems
seem to be around the PCIX bus area of the machine, setting the
PCIX buses to 66 mhz in the bios made things stable (but slow).   Not using
the PCIX bus also seemed to make things work.   We got MCE's and
other odd crashes under heavy IO loads.   I believe turning things
down to 100mhz made things more stable, but things still crashed.

Supermicro reported being able to fix the issue with:
setting the PCI Configuration -> PCI-e I/O performance
setting to Colasce 128B.

I am not exactly sure where to set it as we did not try it
as we had already changed to a different motherboard that did not
have the issue.

If this works please tell me.

                      Roger






  reply	other threads:[~2007-07-31 22:13 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2007-07-30 15:30 Hangs and reboots under high loads, oops with DEBUG_SHIRQ Attila Nagy
2007-07-30 16:18 ` Kok, Auke
2007-07-30 16:19 ` Alan Cox
2007-07-31 14:21   ` Attila Nagy
2007-07-31 22:08     ` Roger Heflin [this message]
2007-08-02  8:29       ` Attila Nagy
2007-08-04 22:56         ` Mr. James W. Laferriere

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=46AFB2CB.6040906@atipa.com \
    --to=rheflin@atipa.com \
    --cc=alan@lxorguk.ukuu.org.uk \
    --cc=bra@fsn.hu \
    --cc=linux-kernel@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox