Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Kiryl Shutsemau <kas@kernel.org>
To: Breno Leitao <leitao@debian.org>
Cc: Ard Biesheuvel <ardb@kernel.org>,
	 Ilias Apalodimas <ilias.apalodimas@linaro.org>,
	Miaohe Lin <linmiaohe@huawei.com>,
	 Naoya Horiguchi <nao.horiguchi@gmail.com>,
	Andrew Morton <akpm@linux-foundation.org>,
	 linux-efi@vger.kernel.org, linux-kernel@vger.kernel.org,
	kexec@lists.infradead.org,  rneu@meta.com, riel@surriel.com,
	caggio@meta.com, anilagrawal@meta.com,  rmikey@meta.com,
	linux-mm@kvack.org, kernel-team@meta.com
Subject: Re: [PATCH RFC 0/3] efi: mm/memory-failure: keep hardware-poisoned pages out of the next kexec
Date: Wed, 22 Jul 2026 14:50:18 +0100	[thread overview]
Message-ID: <amDHmNWfQ8eid9jH@thinkstation> (raw)
In-Reply-To: <20260717-hwpoison-kho-v1-0-9c5eda551998@debian.org>

On Fri, Jul 17, 2026 at 07:03:02AM -0700, Breno Leitao wrote:
> Problem:
> ========
> 
> When a page is hard-offlined due to an uncorrectable memory error (multi
> bit ECC), memory_failure() sets PG_hwpoison, unmaps it, removes it from
> the buddy allocator. This information is not carried to the next kernel
> that is kexeced. The new kernel kexecs and trip over that bad memory
> bank _again_.
> 
> Why now:
> ========
> 
> Several industry trends make this increasingly important:
> 
>     1) DRAM is getting more expensive
>     2) soldered / on-package memory (LPDDR, HBM) is becoming more common, so a
>        failing part can no longer simply be swapped;
>     3) memory is kept in service far longer (at Meta, DRAM lifetime is being
>        drastically extended).
>     4) It is more and more common to kexec instead of full reboot
>     5) Increase of memory per system with CXL
> 
> Proposed Solution:
> ==================
> 
> Carry the poisoned frames to the next kernel in a new EFI configuration
> table, LINUX_EFI_POISONED_MEMORY, modeled on the existing
> LINUX_EFI_MEMRESERVE table.
> 
> EFI configuration tables already survive kexec: firmware hands the EFI
> system table to every kernel in the chain, so a table installed once is
> seen by all successors without a new handover channel.
> 
> The mechanism is architecture independent, so x86 and arm64 use the same
> code.
> 
> The EFI stub installs an empty table while EFI boot service is up. A
> configuration table can only be created there; the running kernel can only
> append to it.
> 
> Each hard-offlined frame is appended; an unpoison "removes" its entry
> so a frame that is good again is not carried forward.
> 
> The next kernel walks the inherited table early in
> efi_config_parse_tables(), before memblock and the buddy allocator are
> up, and memblock_reserve()s every recorded frame. The bad RAM is never
> handed out.

The allocator is only half of the problem. kexec segment placement
doesn't know about any of this: kexec_file picks destinations from
System RAM resources (memblock on arm64), and poisoned frames are only
ever removed from buddy. So the next kernel image, initrd or purgatory
can be placed on top of a poisoned frame -- the relocation memcpy then
consumes the poison and the machine checks during the very kexec this
series is supposed to protect. Same for the table's own list pages:
overwrite one at placement time and the next kernel parses garbage.

Hooking num_poisoned_pages_inc() also means soft-offlined pages are
recorded. Those are functional pages, offlined predictively. Turning
them into permanent losses for every kernel down the kexec chain does
not seem right. I would record hard failures only.

On the data structure: I am not an expert on DRAM failure modes, so I
went reading. The field studies [1] say roughly a quarter of DRAM
faults are multi-bit structures (row/column/bank), and the address
interleaving means one such fault shows up to the OS as many separate
4K pages. A row fault lands in a window of tens of KB to about a MB; a
column fault is one bad line per row, strided across the bank's entire
footprint -- potentially thousands of pages scattered over gigabytes,
reported one MCE at a time as they get touched. DDR5 on-die ECC hides
most single-bit faults from the host, so the visible mix shifts toward
these multi-bit modes over time.

If that is accurate, per-4K entries have no natural bound, and every
entry becomes a separate scattered memblock_reserve() in every future
boot.

I would consider a bitmap with one bit per 2M instead, modeled on
struct efi_unaccepted_memory: the stub sizes it from the EFI memory map
at cold boot, the running kernel sets a bit on hard offline, clears it
on unpoison, and the next kernel reserves set units. It is one
fixed-size allocation, so the whole grow-and-link machinery and the
chain-parsing trust problem go away, and a row fault collapses into one
or two bits.

[1] https://arxiv.org/abs/2408.15302

-- 
  Kiryl Shutsemau / Kirill A. Shutemov


      parent reply	other threads:[~2026-07-22 13:50 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-17 14:03 [PATCH RFC 0/3] efi: mm/memory-failure: keep hardware-poisoned pages out of the next kexec Breno Leitao
2026-07-17 14:03 ` [PATCH RFC 1/3] efi: add the LINUX_EFI_POISONED_MEMORY configuration table Breno Leitao
2026-07-17 14:03 ` [PATCH RFC 2/3] efi: record hardware-poisoned frames into the poisoned-memory table Breno Leitao
2026-07-17 14:03 ` [PATCH RFC 3/3] efi: reserve inherited poisoned frames before the allocator comes up Breno Leitao
2026-07-22 13:50 ` Kiryl Shutsemau [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=amDHmNWfQ8eid9jH@thinkstation \
    --to=kas@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=anilagrawal@meta.com \
    --cc=ardb@kernel.org \
    --cc=caggio@meta.com \
    --cc=ilias.apalodimas@linaro.org \
    --cc=kernel-team@meta.com \
    --cc=kexec@lists.infradead.org \
    --cc=leitao@debian.org \
    --cc=linmiaohe@huawei.com \
    --cc=linux-efi@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=nao.horiguchi@gmail.com \
    --cc=riel@surriel.com \
    --cc=rmikey@meta.com \
    --cc=rneu@meta.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox