From: Breno Leitao <leitao@debian.org>
To: Borislav Petkov <bp@alien8.de>
Cc: "Luck, Tony" <tony.luck@intel.com>,
Thomas Gleixner <tglx@kernel.org>,
Ingo Molnar <mingo@redhat.com>,
Dave Hansen <dave.hansen@linux.intel.com>,
"x86@kernel.org" <x86@kernel.org>,
"H. Peter Anvin" <hpa@zytor.com>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Randy Dunlap <rdunlap@infradead.org>,
"linux-edac@vger.kernel.org" <linux-edac@vger.kernel.org>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
"linux-doc@vger.kernel.org" <linux-doc@vger.kernel.org>,
"kernel-team@meta.com" <kernel-team@meta.com>
Subject: Re: [PATCH 2/2] x86/mce: Add mce=panic_on_ce_count to panic on a corrected error flood
Date: Thu, 27 Aug 2026 02:47:11 -0700 [thread overview]
Message-ID: <apAEy6c6xuXJXSEQ@gmail.com> (raw)
In-Reply-To: <20260826160934.GBao8PvsM8M-U7XuFR@fat_crate.local>
Hello Borislav,
On Wed, Aug 26, 2026 at 09:09:34AM -0700, Borislav Petkov wrote:
> On Wed, Aug 26, 2026 at 06:22:27AM -0700, Breno Leitao wrote:
> > In other words, I'm more interested in flagging that a host is a lemon
> > than in keeping it limping along half-alive, thus my approach in this
> > patchset.
>
> This is your use case but I'm sure others would prefer the limping and
> gracefully saving work the system has done until now...
Right, and I'm making this configurable: users who prefer a clean
panic over a limping machine can set the kernel up that way.
> Putting my user hat on:
>
> + Panic once a machine check bank has logged this many corrected
> + errors. 0, the default, disables it.
> +
> + The setting is global rather than per-CPU, and it is compared
> + against a running total kept per CPU and per bank. Lowering it
> + below a total a bank has already reached takes the machine
> + down on that bank's next corrected error.
>
> Ah, so how many CEs is "many"?
This is a user-configured option; each user picks whatever number
works best for them.
From a quick look at the Meta fleet, this follows a bi-modal curve:
a bunch of hosts logging a dozen or so, and a few with thousands. So
doing some napkin math, I'd say "many" is around 100ish for Meta
fleet.
> Or you simply say, after X CEs, we simply panic the machine - instead of
> gracefully powering it off after leaving a helpful message in dmesg - and have
> lab technicians inspect it?
Correct, that's exactly my approach here.
After X CE errors, we acknowledge that the machine is no longer
stable enough to keep carrying the workload, and simply panic. It's
the last resort: it means lab technicians haven't fixed the problem,
and nobody has gracefully powered off the host either.
Thanks for the review,
--breno
next prev parent reply other threads:[~2026-08-27 9:47 UTC|newest]
Thread overview: 23+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-21 10:24 [PATCH 0/2] x86/mce: Rate-limit storm notices and add a corrected error ceiling Breno Leitao
2026-08-21 10:24 ` [PATCH 1/2] x86/mce: Rate-limit the CMCI storm transition notices Breno Leitao
2026-08-21 16:18 ` Luck, Tony
2026-08-21 16:35 ` Breno Leitao
2026-08-27 9:49 ` Breno Leitao
2026-08-21 10:24 ` [PATCH 2/2] x86/mce: Add mce=panic_on_ce_count to panic on a corrected error flood Breno Leitao
2026-08-21 16:50 ` Luck, Tony
2026-08-24 8:29 ` Breno Leitao
2026-08-24 16:31 ` Luck, Tony
2026-08-24 20:20 ` Luck, Tony
2026-08-25 13:55 ` Breno Leitao
2026-08-25 15:27 ` Luck, Tony
2026-08-25 16:16 ` Borislav Petkov
2026-08-25 16:26 ` Luck, Tony
2026-08-25 19:15 ` Borislav Petkov
2026-08-25 20:01 ` Luck, Tony
2026-08-25 22:52 ` Borislav Petkov
2026-08-26 13:22 ` Breno Leitao
2026-08-26 16:09 ` Borislav Petkov
2026-08-26 17:32 ` Luck, Tony
2026-08-27 9:47 ` Breno Leitao [this message]
2026-08-27 14:28 ` Borislav Petkov
2026-08-24 20:47 ` Luck, Tony
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=apAEy6c6xuXJXSEQ@gmail.com \
--to=leitao@debian.org \
--cc=bp@alien8.de \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=hpa@zytor.com \
--cc=kernel-team@meta.com \
--cc=linux-doc@vger.kernel.org \
--cc=linux-edac@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=rdunlap@infradead.org \
--cc=skhan@linuxfoundation.org \
--cc=tglx@kernel.org \
--cc=tony.luck@intel.com \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox