From: Miaohe Lin <linmiaohe@huawei.com>
To: Breno Leitao <leitao@debian.org>
Cc: <linux-efi@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
<linux-mm@kvack.org>, <rmikey@meta.com>, <riel@surriel.com>,
<harry@kernel.org>, <linux-cxl@vger.kernel.org>,
<driver-core@lists.linux.dev>, <kernel-team@meta.com>,
Ard Biesheuvel <ardb@kernel.org>,
Ilias Apalodimas <ilias.apalodimas@linaro.org>,
Naoya Horiguchi <nao.horiguchi@gmail.com>,
Andrew Morton <akpm@linux-foundation.org>, <kas@kernel.org>,
<kexec@lists.infradead.org>, David Hildenbrand <david@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>,
"Liam R. Howlett" <liam@infradead.org>,
Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>,
Suren Baghdasaryan <surenb@google.com>,
Michal Hocko <mhocko@suse.com>, Thomas Gleixner <tglx@kernel.org>,
Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>, <x86@kernel.org>,
"H. Peter Anvin" <hpa@zytor.com>,
Brendan Jackman <brendan.jackman@linux.dev>,
Johannes Weiner <hannes@cmpxchg.org>, Zi Yan <ziy@nvidia.com>,
Oscar Salvador <osalvador@suse.de>,
Greg Kroah-Hartman <gregkh@linuxfoundation.org>,
"Rafael J. Wysocki" <rafael@kernel.org>,
Danilo Krummrich <dakr@kernel.org>, <hannes@cmpxchg.or>,
<shakeel.butt@linux.dev>
Subject: Re: [PATCH v5 0/9] mm/memory-failure: keep hardware-poisoned pages out of the next kexec
Date: Tue, 22 Sep 2026 14:27:59 +0800 [thread overview]
Message-ID: <d6e35901-f74d-bf81-75d3-bac4b26d9d72@huawei.com> (raw)
In-Reply-To: <20260915-hwpoison-kho-v5-0-3bc7a57bd503@debian.org>
On 2026/9/15 20:53, Breno Leitao wrote:
> Problem:
> ========
>
> When a page is hard-offlined due to an uncorrectable memory error (multi
> bit ECC), memory_failure() sets PG_hwpoison, unmaps it, removes it from
> the buddy allocator. This information is not carried to the next kernel
> that is kexeced. The new kernel kexecs and trip over that bad memory
> bank _again_.
>
> Why now:
> ========
>
> Several industry trends make this increasingly important:
>
> 1) DRAM is getting more expensive
> 2) soldered / on-package memory (LPDDR, HBM) is becoming more common, so a
> failing part can no longer simply be swapped;
> 3) memory is kept in service far longer (at Meta, DRAM lifetime is being
> drastically extended).
> 4) It is more and more common to kexec instead of full reboot
> 5) Increase of memory per system with CXL
>
> What was done already:
> ======================
>
> In order make linux deal better with the problem above, I've done
> already fixed a bunch of stuff in this area, such as:
>
> 1) Panic on unrecoverable errors, instead of "printk and carry over":
> https://lore.kernel.org/all/20260630-ecc_panic-v10-0-c6ed5b62eea2@debian.org/
>
> 2) Respect poisoned memory at kexec time
> https://lore.kernel.org/all/20260812-kexec_posioned-v6-0-e477887086f0@debian.org/
>
> Now, the natural follow up is to carry the poisoned memory information
> to the next kexec kernel, avoiding tripping over a known "bad page".
>
> Proposed Solution:
> ==================
>
> Carry the poisoned frames to the next kernel in a new EFI configuration
> table, LINUX_EFI_POISONED_MEMORY.
>
> The table is a bitmap with one bit per 2MB of physical memory, modeled
> on LINUX_EFI_UNACCEPTED_MEMORY. The stub sizes it from the EFI memory
> map and installs it empty while boot services are up -- a running kernel
> cannot install a configuration table, it can only flip bits -- and a
> table inherited from an earlier boot is reused as-is. That is one
> fixed-size allocation, 64KB per TiB of the span the memory map describes,
> with no list to grow at runtime and no chain to trust at parse time.
>
> The mechanism is architecture independent, so x86 and arm64 use the same
> code.
>
> Each hard offline sets the bit for its unit. Soft-offlined pages are not
> recorded: they are still functional, and were offlined predictively.
>
> The next kernel takes the table into use from efi_config_parse_tables(),
> which vets the inherited header and hands the table's pages to memblock.
> It is EFI ACPI reclaim memory, which x86 leaves out of memblock and so
> out of the direct map; the unaccepted memory table is handled the same
> way for the same reason.
>
> The frames themselves are poisoned in __free_pages_core(), as each block
> reaches the buddy allocator. That is the one point every producer passes
> through -- memblock_free_pages() early, deferred_free_pages() once the
> deferred struct pages are up, and generic_online_page() at any later
> hotplug -- and it is already where the unaccepted memory table is
> consulted. The recorded frames are held back and the rest of the block is
> freed, so they never enter the allocator instead of being taken back out
> of it. They end up in the state a frame poisoned by this kernel would be
> in, so everything that already understands PG_hwpoison covers them --
> including the kexec segment placement check from the series linked above,
> which keeps the segments a kexec places off these frames.
>
> Granularity is the trade-off: one bad 4KB frame costs a whole 2MB unit in
> every later kernel of the chain. In exchange, a row or column fault --
> roughly a quarter of the DRAM faults reported in [1], and potentially
> thousands of 4KB pages scattered over gigabytes -- collapses into a bit
> or two.
>
> A bit is never cleared, which is a known limitation: it stands for a
> whole unit, so an unpoison of one frame cannot tell whether the unit as a
> whole is good again.
Thanks for your patches. I have a question about unpoison: If a bit is never
cleared, after we do some memory-failure+unpoison tests, kexec will lose the
tested memory without reboot?
Thanks.
.
next prev parent reply other threads:[~2026-09-22 6:28 UTC|newest]
Thread overview: 48+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-15 12:53 [PATCH v5 0/9] mm/memory-failure: keep hardware-poisoned pages out of the next kexec Breno Leitao
2026-09-15 12:53 ` [PATCH v5 1/9] mm/page_alloc: factor out the accept-and-free tail of __free_pages_core() Breno Leitao
2026-09-15 12:57 ` sashiko-bot
2026-09-15 12:53 ` [PATCH v5 2/9] mm/memory-failure: efi: add the LINUX_EFI_POISONED_MEMORY configuration table Breno Leitao
2026-09-15 13:05 ` sashiko-bot
2026-09-15 12:53 ` [PATCH v5 3/9] mm/memory-failure: libstub: install the poisoned-memory EFI table Breno Leitao
2026-09-15 13:15 ` sashiko-bot
2026-09-15 14:00 ` Breno Leitao
2026-09-16 15:39 ` Usama Arif
2026-09-17 10:53 ` Breno Leitao
2026-09-15 12:53 ` [PATCH v5 4/9] mm/memory-failure: efi: adopt the inherited poisoned-memory table Breno Leitao
2026-09-15 13:24 ` sashiko-bot
2026-09-15 12:53 ` [PATCH v5 5/9] mm/memory-failure: efi: record hardware-poisoned frames into the " Breno Leitao
2026-09-15 13:36 ` sashiko-bot
2026-09-15 14:33 ` Breno Leitao
2026-09-15 12:53 ` [PATCH v5 6/9] mm/memory-failure: efi: answer whether a range is poisoned Breno Leitao
2026-09-15 13:45 ` sashiko-bot
2026-09-16 6:44 ` David Hildenbrand (Arm)
2026-09-18 15:27 ` Breno Leitao
2026-09-18 15:53 ` Harry Yoo
2026-09-18 20:02 ` David Hildenbrand (Arm)
2026-09-15 12:53 ` [PATCH v5 7/9] drivers/base/memory: count inherited poisoned frames into the block Breno Leitao
2026-09-15 13:59 ` sashiko-bot
2026-09-16 6:48 ` David Hildenbrand (Arm)
2026-09-16 9:35 ` Breno Leitao
2026-09-16 14:51 ` David Hildenbrand (Arm)
2026-09-17 13:01 ` Breno Leitao
2026-09-18 12:18 ` David Hildenbrand (Arm)
2026-09-18 15:22 ` Breno Leitao
2026-09-18 20:16 ` David Hildenbrand (Arm)
2026-09-21 14:31 ` Breno Leitao
2026-09-22 11:33 ` David Hildenbrand (Arm)
2026-09-22 12:34 ` Breno Leitao
2026-09-24 11:22 ` David Hildenbrand (Arm)
2026-09-22 12:51 ` Kiryl Shutsemau
2026-09-22 13:45 ` Harry Yoo
2026-09-22 13:53 ` David Hildenbrand (Arm)
2026-09-15 12:53 ` [PATCH v5 8/9] mm/memory-failure: add hwpoison_boot_page() to flag an inherited frame Breno Leitao
2026-09-15 14:11 ` sashiko-bot
2026-09-19 10:28 ` Shaikh Kamaluddin
2026-09-21 13:41 ` Breno Leitao
2026-09-15 12:53 ` [PATCH v5 9/9] mm/memory-failure: keep inherited poisoned frames out of the buddy allocator Breno Leitao
2026-09-15 14:25 ` sashiko-bot
2026-09-16 8:32 ` Vlastimil Babka (SUSE)
2026-09-22 6:27 ` Miaohe Lin [this message]
2026-09-22 9:14 ` [PATCH v5 0/9] mm/memory-failure: keep hardware-poisoned pages out of the next kexec Breno Leitao
2026-09-22 9:29 ` Miaohe Lin
2026-09-22 10:39 ` Breno Leitao
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=d6e35901-f74d-bf81-75d3-bac4b26d9d72@huawei.com \
--to=linmiaohe@huawei.com \
--cc=akpm@linux-foundation.org \
--cc=ardb@kernel.org \
--cc=bp@alien8.de \
--cc=brendan.jackman@linux.dev \
--cc=dakr@kernel.org \
--cc=dave.hansen@linux.intel.com \
--cc=david@kernel.org \
--cc=driver-core@lists.linux.dev \
--cc=gregkh@linuxfoundation.org \
--cc=hannes@cmpxchg.or \
--cc=hannes@cmpxchg.org \
--cc=harry@kernel.org \
--cc=hpa@zytor.com \
--cc=ilias.apalodimas@linaro.org \
--cc=kas@kernel.org \
--cc=kernel-team@meta.com \
--cc=kexec@lists.infradead.org \
--cc=leitao@debian.org \
--cc=liam@infradead.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-efi@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=mingo@redhat.com \
--cc=nao.horiguchi@gmail.com \
--cc=osalvador@suse.de \
--cc=rafael@kernel.org \
--cc=riel@surriel.com \
--cc=rmikey@meta.com \
--cc=rppt@kernel.org \
--cc=shakeel.butt@linux.dev \
--cc=surenb@google.com \
--cc=tglx@kernel.org \
--cc=vbabka@kernel.org \
--cc=x86@kernel.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox