From: Breno Leitao <leitao@debian.org>
To: Kiryl Shutsemau <kas@kernel.org>
Cc: Ard Biesheuvel <ardb@kernel.org>,
Ilias Apalodimas <ilias.apalodimas@linaro.org>,
Miaohe Lin <linmiaohe@huawei.com>,
Naoya Horiguchi <nao.horiguchi@gmail.com>,
Andrew Morton <akpm@linux-foundation.org>,
linux-efi@vger.kernel.org, linux-kernel@vger.kernel.org,
kexec@lists.infradead.org, rneu@meta.com, riel@surriel.com,
caggio@meta.com, anilagrawal@meta.com, rmikey@meta.com,
linux-mm@kvack.org, kernel-team@meta.com
Subject: Re: [PATCH RFC 0/3] efi: mm/memory-failure: keep hardware-poisoned pages out of the next kexec
Date: Mon, 27 Jul 2026 06:22:14 -0700 [thread overview]
Message-ID: <amdVuOO_6aCxMkA2@gmail.com> (raw)
In-Reply-To: <amDHmNWfQ8eid9jH@thinkstation>
On Wed, Jul 22, 2026 at 02:50:18PM +0100, Kiryl Shutsemau wrote:
> On Fri, Jul 17, 2026 at 07:03:02AM -0700, Breno Leitao wrote:
> > Problem:
> > ========
> >
> > When a page is hard-offlined due to an uncorrectable memory error (multi
> > bit ECC), memory_failure() sets PG_hwpoison, unmaps it, removes it from
> > the buddy allocator. This information is not carried to the next kernel
> > that is kexeced. The new kernel kexecs and trip over that bad memory
> > bank _again_.
> >
> > Why now:
> > ========
> >
> > Several industry trends make this increasingly important:
> >
> > 1) DRAM is getting more expensive
> > 2) soldered / on-package memory (LPDDR, HBM) is becoming more common, so a
> > failing part can no longer simply be swapped;
> > 3) memory is kept in service far longer (at Meta, DRAM lifetime is being
> > drastically extended).
> > 4) It is more and more common to kexec instead of full reboot
> > 5) Increase of memory per system with CXL
> >
> > Proposed Solution:
> > ==================
> >
> > Carry the poisoned frames to the next kernel in a new EFI configuration
> > table, LINUX_EFI_POISONED_MEMORY, modeled on the existing
> > LINUX_EFI_MEMRESERVE table.
> >
> > EFI configuration tables already survive kexec: firmware hands the EFI
> > system table to every kernel in the chain, so a table installed once is
> > seen by all successors without a new handover channel.
> >
> > The mechanism is architecture independent, so x86 and arm64 use the same
> > code.
> >
> > The EFI stub installs an empty table while EFI boot service is up. A
> > configuration table can only be created there; the running kernel can only
> > append to it.
> >
> > Each hard-offlined frame is appended; an unpoison "removes" its entry
> > so a frame that is good again is not carried forward.
> >
> > The next kernel walks the inherited table early in
> > efi_config_parse_tables(), before memblock and the buddy allocator are
> > up, and memblock_reserve()s every recorded frame. The bad RAM is never
> > handed out.
>
> The allocator is only half of the problem. kexec segment placement
> doesn't know about any of this: kexec_file picks destinations from
> System RAM resources (memblock on arm64), and poisoned frames are only
> ever removed from buddy. So the next kernel image, initrd or purgatory
> can be placed on top of a poisoned frame -- the relocation memcpy then
> consumes the poison and the machine checks during the very kexec this
> series is supposed to protect. Same for the table's own list pages:
> overwrite one at placement time and the next kernel parses garbage.
Agreed, and thanks. This is real and, I think, largely separable from
the cross-kexec table -- the running kernel already knows its own
poisoned frames via PG_hwpoison, so excluding them at load time protects
the immediate kexec without depending on anything carried across the
boot. I'd like to tackle it as an independent change, let me do a PoC
and see how much change it is. I suspect it is less than this one (and
probably nice to have even if this series doesn't end up anywhere)
> Hooking num_poisoned_pages_inc() also means soft-offlined pages are
> recorded. Those are functional pages, offlined predictively. Turning
> them into permanent losses for every kernel down the kexec chain does
> not seem right. I would record hard failures only.
Yea, it seems that action_result() is what we want instead of
num_poisoned_pages_inc()
> On the data structure: I am not an expert on DRAM failure modes, so I
> went reading. The field studies [1] say roughly a quarter of DRAM
> faults are multi-bit structures (row/column/bank), and the address
> interleaving means one such fault shows up to the OS as many separate
> 4K pages. A row fault lands in a window of tens of KB to about a MB; a
> column fault is one bad line per row, strided across the bank's entire
> footprint -- potentially thousands of pages scattered over gigabytes,
> reported one MCE at a time as they get touched. DDR5 on-die ECC hides
> most single-bit faults from the host, so the visible mix shifts toward
> these multi-bit modes over time.
>
> If that is accurate, per-4K entries have no natural bound, and every
> entry becomes a separate scattered memblock_reserve() in every future
> boot.
>
> I would consider a bitmap with one bit per 2M instead, modeled on
> struct efi_unaccepted_memory: the stub sizes it from the EFI memory map
> at cold boot, the running kernel sets a bit on hard offline, clears it
> on unpoison, and the next kernel reserves set units. It is one
> fixed-size allocation, so the whole grow-and-link machinery and the
> chain-parsing trust problem go away, and a row fault collapses into one
> or two bits.
Thanks for digging into the failure modes -- that matches my
understanding, and the bounded, chain-free bitmap is appealing for
exactly the reasons you give: a row fault collapsing to a bit or two,
one fixed allocation, and no cross-kernel chain to trust.
Let me hack a PoC with a bitmap and check how it looks like, let's see
if that looks like smoother.
--breno
prev parent reply other threads:[~2026-07-27 13:22 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-17 14:03 [PATCH RFC 0/3] efi: mm/memory-failure: keep hardware-poisoned pages out of the next kexec Breno Leitao
2026-07-17 14:03 ` [PATCH RFC 1/3] efi: add the LINUX_EFI_POISONED_MEMORY configuration table Breno Leitao
2026-07-17 14:03 ` [PATCH RFC 2/3] efi: record hardware-poisoned frames into the poisoned-memory table Breno Leitao
2026-07-17 14:03 ` [PATCH RFC 3/3] efi: reserve inherited poisoned frames before the allocator comes up Breno Leitao
2026-07-22 13:50 ` [PATCH RFC 0/3] efi: mm/memory-failure: keep hardware-poisoned pages out of the next kexec Kiryl Shutsemau
2026-07-27 13:22 ` Breno Leitao [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=amdVuOO_6aCxMkA2@gmail.com \
--to=leitao@debian.org \
--cc=akpm@linux-foundation.org \
--cc=anilagrawal@meta.com \
--cc=ardb@kernel.org \
--cc=caggio@meta.com \
--cc=ilias.apalodimas@linaro.org \
--cc=kas@kernel.org \
--cc=kernel-team@meta.com \
--cc=kexec@lists.infradead.org \
--cc=linmiaohe@huawei.com \
--cc=linux-efi@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=nao.horiguchi@gmail.com \
--cc=riel@surriel.com \
--cc=rmikey@meta.com \
--cc=rneu@meta.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.