linux-mm.kvack.org archive mirror
 help / color / mirror / Atom feed
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Harry Yoo <harry@kernel.org>
Cc: David Rientjes <rientjes@google.com>,
	Amit Shah <amit@infradead.org>,
	Andrew Morton <akpm@linux-foundation.org>,
	Aneesh Kumar <AneeshKumar.KizhakeVeetil@arm.com>,
	Christoph Lameter <christoph@gentwo.org>,
	Dave Hansen <dave.hansen@intel.com>,
	Davidlohr Bueso <dave@stgolabs.net>,
	Hugh Dickins <hughd@google.com>,
	Johannes Weiner <hannes@cmpxchg.org>,
	John Hubbard <jhubbard@nvidia.com>,
	Kirill Shutemov <k.shutemov@gmail.com>,
	Matthew Wilcox <willy@infradead.org>,
	Mel Gorman <mel.gorman@gmail.com>, Michal Hocko <mhocko@suse.com>,
	Mike Rapoport <mike.rapoport@gmail.com>,
	Peter Xu <peterx@redhat.com>, Raghavendra K T <rkodsara@amd.com>,
	"Rao, Bharata Bhasker" <bharata@amd.com>,
	Rik van Riel <riel@surriel.com>,
	Roman Gushchin <roman.gushchin@linux.dev>,
	Shakeel Butt <shakeel.butt@linux.dev>,
	Shivank Garg <shivankg@amd.com>,
	Sterling Alexander <stalexan@redhat.com>,
	Suren Baghdasaryan <surenb@google.com>, Tejun Heo <tj@kernel.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	Yang Shi <shy828301@gmail.com>, Zi Yan <ziy@nvidia.com>,
	William Roche <william.roche@oracle.com>,
	linmiaohe@huawei.com, ljs@kernel.org, osalvador@kernel.org,
	nao.horiguchi@gmail.com, tony.luck@intel.com,
	wangkefeng.wang@huawei.com, jane.chu@oracle.com,
	muchun.song@linux.dev, liam@infradead.org, shuah@kernel.org,
	boudewijn@delta-utec.com, linux-mm@kvack.org,
	Breno Leitao <leitao@debian.org>,
	kas@kernel.org
Subject: Re: [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday
Date: Thu, 24 Sep 2026 10:32:29 +0200	[thread overview]
Message-ID: <ce6659a6-a658-4311-8c8d-193d309bf85b@kernel.org> (raw)
In-Reply-To: <arPdnxhX9VHnDx15@thinkstation>

On 9/23/26 16:36, Harry Yoo wrote:
> On Wed, Sep 23, 2026 at 03:37:14PM +0200, David Hildenbrand (Arm) wrote:
>> On 9/23/26 14:57, Harry Yoo wrote:
>>>
>>> [+Cc Breno]
>>>
>>> This might be worth some attention:
>>>
>>>   [PATCH v5 0/9] mm/memory-failure: keep hardware-poisoned pages out of the next kexec
>>>   https://lore.kernel.org/linux-mm/20260915-hwpoison-kho-v5-0-3bc7a57bd503@debian.org/
>>>
>>>   - The question: How best can we pass information about poisoned pages
>>>     across kexec when the kernel accesses a poisoned page, panics and
>>>     calls kexec, so that the next kernel avoids hitting the poisoned
>>>     pages again?
>>>
>>>     One challenge here is introducing another source (a new EFI bitmap
>>>     that represents poisoned memory, in addition to existing
>>>     PG_hwpoison) of poisoned pages adds complexity and confusion.
>>>
>>>     David Hildenbrand suggested that we should simplify this as much
>>>     as possible in a way that we don't have two sources of information
>>>     w/ inconsistency between them and using memmap as the only source.
>>>
>>>     e.g.) By consuming the EFI bitmap only once when initializing
>>>     memmap (to propagate poison information to not just free pages
>>>     through __free_pages_core(), but to the all pages on memmap).
>>>
>>>     Another challenge here is that we don't have functionality
>>>     to poison pages early in the boot process before memmap and buddy
>>>     are ready. So any allocation before initializing them might still
>>>     allocate poisoned pages.
>>
>> I've been thinking some more (and will reply in detail to the series), but I do
>> wonder whether memblock should actually be thought about poisoned ranges and
>> refuse to hand them out (reserved/allocated).
> 
> I believe this is what the patchset actually did in v2.
> 
> That way we'll have to either
> 1) make any architecture that supports kexec select ARCH_KEEP_MEMBLOCK,
> or 2) make kexec scan the bitmap when allocating the memory.

My naive design would be:

1) Teach memblock early about poisoned memory ranges. Don't let it hand them out.

2) Memblock will naturally skip them when freeing pages to the buddy. No need to
hook bitmap scans into other code path.

3) Teach memblock to hwpoison the memmap.

4) Don't update memblock (ARCH_KEEP_MEMBLOCK) when poisoning occurs later during
boot. Any hwpoison happening after boot will be reflected in the memmap (+
synced to the bitmap).


What you are describing is "how to keep kexec from placing data on hwpoisoned
ranges". If we want to avoid scanning a bitmap (which I really think we want to
avoid), the memmap is our ultimate source of truth for hwpoisoned. I would just
consult that. (that's likely what v2 did?)

One complication might be that a single hwpoisoned page is represented as a
bigger chunk in the bitmap. So we might place kexec data on healthy pages within
a poisoned (per bitmap) block. That's fine (pages are healthy after all), we
just have to teach the new kernel to deal with that.

> 
>> So we'd use the bitmap only to feed memblock, but not for anything else later
>> (if possible).
> 
> We can try not to use it for anything later, but still we'd have to
> pass a proper bitmap (through the EFI table) to the next kernel.
> At least the kernel should update the bitmap for newly poisoned pages
> at some point.
Yes. I think it's fine to update the bitmap passively through some callback
mechanism from hwpoison code.

-- 
Cheers,

David


  reply	other threads:[~2026-09-24  8:32 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-21 16:26 [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday David Rientjes
2026-09-23 12:57 ` Harry Yoo
2026-09-23 13:37   ` David Hildenbrand (Arm)
2026-09-23 14:36     ` Harry Yoo
2026-09-24  8:32       ` David Hildenbrand (Arm) [this message]
2026-09-24 10:46         ` Kiryl Shutsemau
2026-09-24 10:55           ` David Hildenbrand (Arm)
2026-09-24 11:19             ` Kiryl Shutsemau
2026-09-24 11:23               ` David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ce6659a6-a658-4311-8c8d-193d309bf85b@kernel.org \
    --to=david@kernel.org \
    --cc=AneeshKumar.KizhakeVeetil@arm.com \
    --cc=akpm@linux-foundation.org \
    --cc=amit@infradead.org \
    --cc=bharata@amd.com \
    --cc=boudewijn@delta-utec.com \
    --cc=christoph@gentwo.org \
    --cc=dave.hansen@intel.com \
    --cc=dave@stgolabs.net \
    --cc=hannes@cmpxchg.org \
    --cc=harry@kernel.org \
    --cc=hughd@google.com \
    --cc=jane.chu@oracle.com \
    --cc=jhubbard@nvidia.com \
    --cc=k.shutemov@gmail.com \
    --cc=kas@kernel.org \
    --cc=leitao@debian.org \
    --cc=liam@infradead.org \
    --cc=linmiaohe@huawei.com \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mel.gorman@gmail.com \
    --cc=mhocko@suse.com \
    --cc=mike.rapoport@gmail.com \
    --cc=muchun.song@linux.dev \
    --cc=nao.horiguchi@gmail.com \
    --cc=osalvador@kernel.org \
    --cc=peterx@redhat.com \
    --cc=riel@surriel.com \
    --cc=rientjes@google.com \
    --cc=rkodsara@amd.com \
    --cc=roman.gushchin@linux.dev \
    --cc=shakeel.butt@linux.dev \
    --cc=shivankg@amd.com \
    --cc=shuah@kernel.org \
    --cc=shy828301@gmail.com \
    --cc=stalexan@redhat.com \
    --cc=surenb@google.com \
    --cc=tj@kernel.org \
    --cc=tony.luck@intel.com \
    --cc=vbabka@kernel.org \
    --cc=wangkefeng.wang@huawei.com \
    --cc=william.roche@oracle.com \
    --cc=willy@infradead.org \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).