Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday
@ 2026-09-21 16:26 David Rientjes
  2026-09-23 12:57 ` Harry Yoo
  0 siblings, 1 reply; 9+ messages in thread
From: David Rientjes @ 2026-09-21 16:26 UTC (permalink / raw)
  To: Amit Shah, Andrew Morton, Aneesh Kumar, Christoph Lameter,
	Dave Hansen, David Hildenbrand, Davidlohr Bueso, Hugh Dickins,
	Johannes Weiner, John Hubbard, Kirill Shutemov, Matthew Wilcox,
	Mel Gorman, Michal Hocko, Mike Rapoport, Peter Xu,
	Raghavendra K T, Rao, Bharata Bhasker, Rik van Riel,
	Roman Gushchin, Shakeel Butt, Shivank Garg, Sterling Alexander,
	Suren Baghdasaryan, Tejun Heo, Vlastimil Babka, Yang Shi, Zi Yan,
	William Roche, linmiaohe, ljs, osalvador, harry.yoo,
	nao.horiguchi, tony.luck, wangkefeng.wang, jane.chu, muchun.song,
	liam, shuah, boudewijn
  Cc: linux-mm

Hi everybody,

We host a biweekly series, the Linux MM Alignment Session, on Wednesdays.
We'd like to invite MM developers to attend and will announce the topic
for the next instance on the Monday prior to the next meeting.

Our next Linux MM Alignment Session is scheduled for Wednesday.  The
details:

Wednesday, September 23 * 9:00 - 10:00am PDT (UTC-7)
https://meet.google.com/csb-wcds-xya
backup: https://tel.meet/csb-wcds-xya?pin=1301132214803
doc:
https://docs.google.com/document/d/1QQfZFPHa-pGEGf4A06cSqS0jdJwuZ8ro9JU5H0YSUQ0/view

This week's topic will be Hwpoison, see the discussion in 
https://marc.info/?l=linux-kernel&m=178796007385067&w=2

No specific agenda other than to gather core MM developers with hwpoison
developers and discuss:
 - where validation should live, esp with regard to the page allocator, or
   strict invariants in the page allocator
   + implications for the fastpath
 - handling for folios of arbitary orders
 - handling partial folio poisoning and split failures
 - race conditions between memory_failure() and concurrent
   alloc/free/compaction
 - soft offline for pages in the pcp lists

We might even want to be bold and discuss a holistic design overhaul if
it's warranted.

If anybody has ideas for future topics, please let me know and I'll try to 
organize them.  We'd love to have volunteers to lead future topics as well 
as requests for MM topics to be presented.

Looking forward to seeing all of you on Wednesday!

Time zones

PDT (UTC-7)		9:00am
MDT (UTC-6)		10:00am
CDT (UTC-5)		11:00am
EDT (UTC-4)		12:00pm
Rio de Janeiro (UTC-3)	1:00pm
London (UTC+1)		5:00pm
Berlin (UTC+2)		6:00pm
Moscow (UTC+3)		7:00pm
Dubai (UTC+4)		8:00pm
Mumbai (UTC+5:30)	9:30pm
Singapore (UTC+8)	12:00am Thursday
Beijing (UTC+8)		12:00am Thursday
Tokyo (UTC+9)		1:00am Thursday
Sydney (UTC+10)		2:00am Thursday
Auckland (UTC+12)	4:00am Thursday


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday
  2026-09-21 16:26 [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday David Rientjes
@ 2026-09-23 12:57 ` Harry Yoo
  2026-09-23 13:37   ` David Hildenbrand (Arm)
  0 siblings, 1 reply; 9+ messages in thread
From: Harry Yoo @ 2026-09-23 12:57 UTC (permalink / raw)
  To: David Rientjes
  Cc: Amit Shah, Andrew Morton, Aneesh Kumar, Christoph Lameter,
	Dave Hansen, David Hildenbrand, Davidlohr Bueso, Hugh Dickins,
	Johannes Weiner, John Hubbard, Kirill Shutemov, Matthew Wilcox,
	Mel Gorman, Michal Hocko, Mike Rapoport, Peter Xu,
	Raghavendra K T, Rao, Bharata Bhasker, Rik van Riel,
	Roman Gushchin, Shakeel Butt, Shivank Garg, Sterling Alexander,
	Suren Baghdasaryan, Tejun Heo, Vlastimil Babka, Yang Shi, Zi Yan,
	William Roche, linmiaohe, ljs, osalvador, nao.horiguchi,
	tony.luck, wangkefeng.wang, jane.chu, muchun.song, liam, shuah,
	boudewijn, linux-mm, Breno Leitao

On Mon, Sep 21, 2026 at 09:26:22AM -0700, David Rientjes wrote:
> Hi everybody,
> 
> We host a biweekly series, the Linux MM Alignment Session, on Wednesdays.
> We'd like to invite MM developers to attend and will announce the topic
> for the next instance on the Monday prior to the next meeting.
> 
> Our next Linux MM Alignment Session is scheduled for Wednesday.  The
> details:
> 
> Wednesday, September 23 * 9:00 - 10:00am PDT (UTC-7)
> https://meet.google.com/csb-wcds-xya
> backup: https://tel.meet/csb-wcds-xya?pin=1301132214803
> doc:
> https://docs.google.com/document/d/1QQfZFPHa-pGEGf4A06cSqS0jdJwuZ8ro9JU5H0YSUQ0/view
> 
> This week's topic will be Hwpoison, see the discussion in 
> https://marc.info/?l=linux-kernel&m=178796007385067&w=2
> 
> No specific agenda other than to gather core MM developers with hwpoison
> developers and discuss:
>  - where validation should live, esp with regard to the page allocator, or
>    strict invariants in the page allocator
>    + implications for the fastpath
>  - handling for folios of arbitary orders
>  - handling partial folio poisoning and split failures
>  - race conditions between memory_failure() and concurrent
>    alloc/free/compaction
>  - soft offline for pages in the pcp lists

[+Cc Breno]

This might be worth some attention:

  [PATCH v5 0/9] mm/memory-failure: keep hardware-poisoned pages out of the next kexec
  https://lore.kernel.org/linux-mm/20260915-hwpoison-kho-v5-0-3bc7a57bd503@debian.org/

  - The question: How best can we pass information about poisoned pages
    across kexec when the kernel accesses a poisoned page, panics and
    calls kexec, so that the next kernel avoids hitting the poisoned
    pages again?

    One challenge here is introducing another source (a new EFI bitmap
    that represents poisoned memory, in addition to existing
    PG_hwpoison) of poisoned pages adds complexity and confusion.

    David Hildenbrand suggested that we should simplify this as much
    as possible in a way that we don't have two sources of information
    w/ inconsistency between them and using memmap as the only source.

    e.g.) By consuming the EFI bitmap only once when initializing
    memmap (to propagate poison information to not just free pages
    through __free_pages_core(), but to the all pages on memmap).

    Another challenge here is that we don't have functionality
    to poison pages early in the boot process before memmap and buddy
    are ready. So any allocation before initializing them might still
    allocate poisoned pages.

> We might even want to be bold and discuss a holistic design overhaul if
> it's warranted.
> 
> If anybody has ideas for future topics, please let me know and I'll try to 
> organize them.  We'd love to have volunteers to lead future topics as well 
> as requests for MM topics to be presented.
> 
> Looking forward to seeing all of you on Wednesday!
> 
> Time zones
> 
> PDT (UTC-7)		9:00am
> MDT (UTC-6)		10:00am
> CDT (UTC-5)		11:00am
> EDT (UTC-4)		12:00pm
> Rio de Janeiro (UTC-3)	1:00pm
> London (UTC+1)		5:00pm
> Berlin (UTC+2)		6:00pm
> Moscow (UTC+3)		7:00pm
> Dubai (UTC+4)		8:00pm
> Mumbai (UTC+5:30)	9:30pm
> Singapore (UTC+8)	12:00am Thursday
> Beijing (UTC+8)		12:00am Thursday
> Tokyo (UTC+9)		1:00am Thursday
> Sydney (UTC+10)		2:00am Thursday
> Auckland (UTC+12)	4:00am Thursday

-- 
Cheers,
Harry / Hyeonggon


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday
  2026-09-23 12:57 ` Harry Yoo
@ 2026-09-23 13:37   ` David Hildenbrand (Arm)
  2026-09-23 14:36     ` Harry Yoo
  0 siblings, 1 reply; 9+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-23 13:37 UTC (permalink / raw)
  To: Harry Yoo, David Rientjes
  Cc: Amit Shah, Andrew Morton, Aneesh Kumar, Christoph Lameter,
	Dave Hansen, Davidlohr Bueso, Hugh Dickins, Johannes Weiner,
	John Hubbard, Kirill Shutemov, Matthew Wilcox, Mel Gorman,
	Michal Hocko, Mike Rapoport, Peter Xu, Raghavendra K T,
	Rao, Bharata Bhasker, Rik van Riel, Roman Gushchin, Shakeel Butt,
	Shivank Garg, Sterling Alexander, Suren Baghdasaryan, Tejun Heo,
	Vlastimil Babka, Yang Shi, Zi Yan, William Roche, linmiaohe, ljs,
	osalvador, nao.horiguchi, tony.luck, wangkefeng.wang, jane.chu,
	muchun.song, liam, shuah, boudewijn, linux-mm, Breno Leitao

On 9/23/26 14:57, Harry Yoo wrote:
> On Mon, Sep 21, 2026 at 09:26:22AM -0700, David Rientjes wrote:
>> Hi everybody,
>>
>> We host a biweekly series, the Linux MM Alignment Session, on Wednesdays.
>> We'd like to invite MM developers to attend and will announce the topic
>> for the next instance on the Monday prior to the next meeting.
>>
>> Our next Linux MM Alignment Session is scheduled for Wednesday.  The
>> details:
>>
>> Wednesday, September 23 * 9:00 - 10:00am PDT (UTC-7)
>> https://meet.google.com/csb-wcds-xya
>> backup: https://tel.meet/csb-wcds-xya?pin=1301132214803
>> doc:
>> https://docs.google.com/document/d/1QQfZFPHa-pGEGf4A06cSqS0jdJwuZ8ro9JU5H0YSUQ0/view
>>
>> This week's topic will be Hwpoison, see the discussion in 
>> https://marc.info/?l=linux-kernel&m=178796007385067&w=2
>>
>> No specific agenda other than to gather core MM developers with hwpoison
>> developers and discuss:
>>  - where validation should live, esp with regard to the page allocator, or
>>    strict invariants in the page allocator
>>    + implications for the fastpath
>>  - handling for folios of arbitary orders
>>  - handling partial folio poisoning and split failures
>>  - race conditions between memory_failure() and concurrent
>>    alloc/free/compaction
>>  - soft offline for pages in the pcp lists
> 
> [+Cc Breno]
> 
> This might be worth some attention:
> 
>   [PATCH v5 0/9] mm/memory-failure: keep hardware-poisoned pages out of the next kexec
>   https://lore.kernel.org/linux-mm/20260915-hwpoison-kho-v5-0-3bc7a57bd503@debian.org/
> 
>   - The question: How best can we pass information about poisoned pages
>     across kexec when the kernel accesses a poisoned page, panics and
>     calls kexec, so that the next kernel avoids hitting the poisoned
>     pages again?
> 
>     One challenge here is introducing another source (a new EFI bitmap
>     that represents poisoned memory, in addition to existing
>     PG_hwpoison) of poisoned pages adds complexity and confusion.
> 
>     David Hildenbrand suggested that we should simplify this as much
>     as possible in a way that we don't have two sources of information
>     w/ inconsistency between them and using memmap as the only source.
> 
>     e.g.) By consuming the EFI bitmap only once when initializing
>     memmap (to propagate poison information to not just free pages
>     through __free_pages_core(), but to the all pages on memmap).
> 
>     Another challenge here is that we don't have functionality
>     to poison pages early in the boot process before memmap and buddy
>     are ready. So any allocation before initializing them might still
>     allocate poisoned pages.

I've been thinking some more (and will reply in detail to the series), but I do
wonder whether memblock should actually be thought about poisoned ranges and
refuse to hand them out (reserved/allocated).

So we'd use the bitmap only to feed memblock, but not for anything else later
(if possible).

-- 
Cheers,

David


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday
  2026-09-23 13:37   ` David Hildenbrand (Arm)
@ 2026-09-23 14:36     ` Harry Yoo
  2026-09-24  8:32       ` David Hildenbrand (Arm)
  0 siblings, 1 reply; 9+ messages in thread
From: Harry Yoo @ 2026-09-23 14:36 UTC (permalink / raw)
  To: David Hildenbrand (Arm)
  Cc: David Rientjes, Amit Shah, Andrew Morton, Aneesh Kumar,
	Christoph Lameter, Dave Hansen, Davidlohr Bueso, Hugh Dickins,
	Johannes Weiner, John Hubbard, Kirill Shutemov, Matthew Wilcox,
	Mel Gorman, Michal Hocko, Mike Rapoport, Peter Xu,
	Raghavendra K T, Rao, Bharata Bhasker, Rik van Riel,
	Roman Gushchin, Shakeel Butt, Shivank Garg, Sterling Alexander,
	Suren Baghdasaryan, Tejun Heo, Vlastimil Babka, Yang Shi, Zi Yan,
	William Roche, linmiaohe, ljs, osalvador, nao.horiguchi,
	tony.luck, wangkefeng.wang, jane.chu, muchun.song, liam, shuah,
	boudewijn, linux-mm, Breno Leitao, kas

On Wed, Sep 23, 2026 at 03:37:14PM +0200, David Hildenbrand (Arm) wrote:
> On 9/23/26 14:57, Harry Yoo wrote:
> > On Mon, Sep 21, 2026 at 09:26:22AM -0700, David Rientjes wrote:
> >> Hi everybody,
> >>
> >> We host a biweekly series, the Linux MM Alignment Session, on Wednesdays.
> >> We'd like to invite MM developers to attend and will announce the topic
> >> for the next instance on the Monday prior to the next meeting.
> >>
> >> Our next Linux MM Alignment Session is scheduled for Wednesday.  The
> >> details:
> >>
> >> Wednesday, September 23 * 9:00 - 10:00am PDT (UTC-7)
> >> https://meet.google.com/csb-wcds-xya
> >> backup: https://tel.meet/csb-wcds-xya?pin=1301132214803
> >> doc:
> >> https://docs.google.com/document/d/1QQfZFPHa-pGEGf4A06cSqS0jdJwuZ8ro9JU5H0YSUQ0/view
> >>
> >> This week's topic will be Hwpoison, see the discussion in 
> >> https://marc.info/?l=linux-kernel&m=178796007385067&w=2
> >>
> >> No specific agenda other than to gather core MM developers with hwpoison
> >> developers and discuss:
> >>  - where validation should live, esp with regard to the page allocator, or
> >>    strict invariants in the page allocator
> >>    + implications for the fastpath
> >>  - handling for folios of arbitary orders
> >>  - handling partial folio poisoning and split failures
> >>  - race conditions between memory_failure() and concurrent
> >>    alloc/free/compaction
> >>  - soft offline for pages in the pcp lists
> > 
> > [+Cc Breno]
> > 
> > This might be worth some attention:
> > 
> >   [PATCH v5 0/9] mm/memory-failure: keep hardware-poisoned pages out of the next kexec
> >   https://lore.kernel.org/linux-mm/20260915-hwpoison-kho-v5-0-3bc7a57bd503@debian.org/
> > 
> >   - The question: How best can we pass information about poisoned pages
> >     across kexec when the kernel accesses a poisoned page, panics and
> >     calls kexec, so that the next kernel avoids hitting the poisoned
> >     pages again?
> > 
> >     One challenge here is introducing another source (a new EFI bitmap
> >     that represents poisoned memory, in addition to existing
> >     PG_hwpoison) of poisoned pages adds complexity and confusion.
> > 
> >     David Hildenbrand suggested that we should simplify this as much
> >     as possible in a way that we don't have two sources of information
> >     w/ inconsistency between them and using memmap as the only source.
> > 
> >     e.g.) By consuming the EFI bitmap only once when initializing
> >     memmap (to propagate poison information to not just free pages
> >     through __free_pages_core(), but to the all pages on memmap).
> > 
> >     Another challenge here is that we don't have functionality
> >     to poison pages early in the boot process before memmap and buddy
> >     are ready. So any allocation before initializing them might still
> >     allocate poisoned pages.
> 
> I've been thinking some more (and will reply in detail to the series), but I do
> wonder whether memblock should actually be thought about poisoned ranges and
> refuse to hand them out (reserved/allocated).

I believe this is what the patchset actually did in v2.

That way we'll have to either
1) make any architecture that supports kexec select ARCH_KEEP_MEMBLOCK,
or 2) make kexec scan the bitmap when allocating the memory.

> So we'd use the bitmap only to feed memblock, but not for anything else later
> (if possible).

We can try not to use it for anything later, but still we'd have to
pass a proper bitmap (through the EFI table) to the next kernel.
At least the kernel should update the bitmap for newly poisoned pages
at some point.

-- 
Cheers,
Harry / Hyeonggon


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday
  2026-09-23 14:36     ` Harry Yoo
@ 2026-09-24  8:32       ` David Hildenbrand (Arm)
  2026-09-24 10:46         ` Kiryl Shutsemau
  0 siblings, 1 reply; 9+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-24  8:32 UTC (permalink / raw)
  To: Harry Yoo
  Cc: David Rientjes, Amit Shah, Andrew Morton, Aneesh Kumar,
	Christoph Lameter, Dave Hansen, Davidlohr Bueso, Hugh Dickins,
	Johannes Weiner, John Hubbard, Kirill Shutemov, Matthew Wilcox,
	Mel Gorman, Michal Hocko, Mike Rapoport, Peter Xu,
	Raghavendra K T, Rao, Bharata Bhasker, Rik van Riel,
	Roman Gushchin, Shakeel Butt, Shivank Garg, Sterling Alexander,
	Suren Baghdasaryan, Tejun Heo, Vlastimil Babka, Yang Shi, Zi Yan,
	William Roche, linmiaohe, ljs, osalvador, nao.horiguchi,
	tony.luck, wangkefeng.wang, jane.chu, muchun.song, liam, shuah,
	boudewijn, linux-mm, Breno Leitao, kas

On 9/23/26 16:36, Harry Yoo wrote:
> On Wed, Sep 23, 2026 at 03:37:14PM +0200, David Hildenbrand (Arm) wrote:
>> On 9/23/26 14:57, Harry Yoo wrote:
>>>
>>> [+Cc Breno]
>>>
>>> This might be worth some attention:
>>>
>>>   [PATCH v5 0/9] mm/memory-failure: keep hardware-poisoned pages out of the next kexec
>>>   https://lore.kernel.org/linux-mm/20260915-hwpoison-kho-v5-0-3bc7a57bd503@debian.org/
>>>
>>>   - The question: How best can we pass information about poisoned pages
>>>     across kexec when the kernel accesses a poisoned page, panics and
>>>     calls kexec, so that the next kernel avoids hitting the poisoned
>>>     pages again?
>>>
>>>     One challenge here is introducing another source (a new EFI bitmap
>>>     that represents poisoned memory, in addition to existing
>>>     PG_hwpoison) of poisoned pages adds complexity and confusion.
>>>
>>>     David Hildenbrand suggested that we should simplify this as much
>>>     as possible in a way that we don't have two sources of information
>>>     w/ inconsistency between them and using memmap as the only source.
>>>
>>>     e.g.) By consuming the EFI bitmap only once when initializing
>>>     memmap (to propagate poison information to not just free pages
>>>     through __free_pages_core(), but to the all pages on memmap).
>>>
>>>     Another challenge here is that we don't have functionality
>>>     to poison pages early in the boot process before memmap and buddy
>>>     are ready. So any allocation before initializing them might still
>>>     allocate poisoned pages.
>>
>> I've been thinking some more (and will reply in detail to the series), but I do
>> wonder whether memblock should actually be thought about poisoned ranges and
>> refuse to hand them out (reserved/allocated).
> 
> I believe this is what the patchset actually did in v2.
> 
> That way we'll have to either
> 1) make any architecture that supports kexec select ARCH_KEEP_MEMBLOCK,
> or 2) make kexec scan the bitmap when allocating the memory.

My naive design would be:

1) Teach memblock early about poisoned memory ranges. Don't let it hand them out.

2) Memblock will naturally skip them when freeing pages to the buddy. No need to
hook bitmap scans into other code path.

3) Teach memblock to hwpoison the memmap.

4) Don't update memblock (ARCH_KEEP_MEMBLOCK) when poisoning occurs later during
boot. Any hwpoison happening after boot will be reflected in the memmap (+
synced to the bitmap).


What you are describing is "how to keep kexec from placing data on hwpoisoned
ranges". If we want to avoid scanning a bitmap (which I really think we want to
avoid), the memmap is our ultimate source of truth for hwpoisoned. I would just
consult that. (that's likely what v2 did?)

One complication might be that a single hwpoisoned page is represented as a
bigger chunk in the bitmap. So we might place kexec data on healthy pages within
a poisoned (per bitmap) block. That's fine (pages are healthy after all), we
just have to teach the new kernel to deal with that.

> 
>> So we'd use the bitmap only to feed memblock, but not for anything else later
>> (if possible).
> 
> We can try not to use it for anything later, but still we'd have to
> pass a proper bitmap (through the EFI table) to the next kernel.
> At least the kernel should update the bitmap for newly poisoned pages
> at some point.
Yes. I think it's fine to update the bitmap passively through some callback
mechanism from hwpoison code.

-- 
Cheers,

David


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday
  2026-09-24  8:32       ` David Hildenbrand (Arm)
@ 2026-09-24 10:46         ` Kiryl Shutsemau
  2026-09-24 10:55           ` David Hildenbrand (Arm)
  0 siblings, 1 reply; 9+ messages in thread
From: Kiryl Shutsemau @ 2026-09-24 10:46 UTC (permalink / raw)
  To: David Hildenbrand (Arm)
  Cc: Harry Yoo, David Rientjes, Amit Shah, Andrew Morton, Aneesh Kumar,
	Christoph Lameter, Dave Hansen, Davidlohr Bueso, Hugh Dickins,
	Johannes Weiner, John Hubbard, Matthew Wilcox, Mel Gorman,
	Michal Hocko, Mike Rapoport, Peter Xu, Raghavendra K T,
	Rao, Bharata Bhasker, Rik van Riel, Roman Gushchin, Shakeel Butt,
	Shivank Garg, Sterling Alexander, Suren Baghdasaryan, Tejun Heo,
	Vlastimil Babka, Yang Shi, Zi Yan, William Roche, linmiaohe, ljs,
	osalvador, nao.horiguchi, tony.luck, wangkefeng.wang, jane.chu,
	muchun.song, liam, shuah, boudewijn, linux-mm, Breno Leitao

On Thu, Sep 24, 2026 at 10:32:29AM +0200, David Hildenbrand (Arm) wrote:
> On 9/23/26 16:36, Harry Yoo wrote:
> > On Wed, Sep 23, 2026 at 03:37:14PM +0200, David Hildenbrand (Arm) wrote:
> >> On 9/23/26 14:57, Harry Yoo wrote:
> >>>
> >>> [+Cc Breno]
> >>>
> >>> This might be worth some attention:
> >>>
> >>>   [PATCH v5 0/9] mm/memory-failure: keep hardware-poisoned pages out of the next kexec
> >>>   https://lore.kernel.org/linux-mm/20260915-hwpoison-kho-v5-0-3bc7a57bd503@debian.org/
> >>>
> >>>   - The question: How best can we pass information about poisoned pages
> >>>     across kexec when the kernel accesses a poisoned page, panics and
> >>>     calls kexec, so that the next kernel avoids hitting the poisoned
> >>>     pages again?
> >>>
> >>>     One challenge here is introducing another source (a new EFI bitmap
> >>>     that represents poisoned memory, in addition to existing
> >>>     PG_hwpoison) of poisoned pages adds complexity and confusion.
> >>>
> >>>     David Hildenbrand suggested that we should simplify this as much
> >>>     as possible in a way that we don't have two sources of information
> >>>     w/ inconsistency between them and using memmap as the only source.
> >>>
> >>>     e.g.) By consuming the EFI bitmap only once when initializing
> >>>     memmap (to propagate poison information to not just free pages
> >>>     through __free_pages_core(), but to the all pages on memmap).
> >>>
> >>>     Another challenge here is that we don't have functionality
> >>>     to poison pages early in the boot process before memmap and buddy
> >>>     are ready. So any allocation before initializing them might still
> >>>     allocate poisoned pages.
> >>
> >> I've been thinking some more (and will reply in detail to the series), but I do
> >> wonder whether memblock should actually be thought about poisoned ranges and
> >> refuse to hand them out (reserved/allocated).
> > 
> > I believe this is what the patchset actually did in v2.
> > 
> > That way we'll have to either
> > 1) make any architecture that supports kexec select ARCH_KEEP_MEMBLOCK,
> > or 2) make kexec scan the bitmap when allocating the memory.
> 
> My naive design would be:
> 
> 1) Teach memblock early about poisoned memory ranges. Don't let it hand them out.

I suggested using a bitmap because ranges are not scalable. Some memory
failure modes generate errors repeated across the physical address space
(think of a column failure, for instance). It would produce too many
ranges to be viable.

We already consult a bitmap in memblock for unaccepted memory. We *can*
do the same for poisoned memory.

But I am not convinced it is needed for the initial implementation.
Memblock is a small portion of kernel allocations. If we step on broken
memory there, the machine is dead and requires repair. It can be
improved later if the rate is too high.

-- 
  Kiryl Shutsemau / Kirill A. Shutemov


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday
  2026-09-24 10:46         ` Kiryl Shutsemau
@ 2026-09-24 10:55           ` David Hildenbrand (Arm)
  2026-09-24 11:19             ` Kiryl Shutsemau
  0 siblings, 1 reply; 9+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-24 10:55 UTC (permalink / raw)
  To: Kiryl Shutsemau
  Cc: Harry Yoo, David Rientjes, Amit Shah, Andrew Morton, Aneesh Kumar,
	Christoph Lameter, Dave Hansen, Davidlohr Bueso, Hugh Dickins,
	Johannes Weiner, John Hubbard, Matthew Wilcox, Mel Gorman,
	Michal Hocko, Mike Rapoport, Peter Xu, Raghavendra K T,
	Rao, Bharata Bhasker, Rik van Riel, Roman Gushchin, Shakeel Butt,
	Shivank Garg, Sterling Alexander, Suren Baghdasaryan, Tejun Heo,
	Vlastimil Babka, Yang Shi, Zi Yan, William Roche, linmiaohe, ljs,
	osalvador, nao.horiguchi, tony.luck, wangkefeng.wang, jane.chu,
	muchun.song, liam, shuah, boudewijn, linux-mm, Breno Leitao

On 9/24/26 12:46, Kiryl Shutsemau wrote:
> On Thu, Sep 24, 2026 at 10:32:29AM +0200, David Hildenbrand (Arm) wrote:
>> On 9/23/26 16:36, Harry Yoo wrote:
>>>
>>> I believe this is what the patchset actually did in v2.
>>>
>>> That way we'll have to either
>>> 1) make any architecture that supports kexec select ARCH_KEEP_MEMBLOCK,
>>> or 2) make kexec scan the bitmap when allocating the memory.
>>
>> My naive design would be:
>>
>> 1) Teach memblock early about poisoned memory ranges. Don't let it hand them out.
> 
> I suggested using a bitmap because ranges are not scalable. Some memory
> failure modes generate errors repeated across the physical address space
> (think of a column failure, for instance). It would produce too many
> ranges to be viable.

These are in 2M granularity. Is that a real problem?

> 
> We already consult a bitmap in memblock for unaccepted memory. We *can*
> do the same for poisoned memory.
> 

We can do many things. I don't want it.

> But I am not convinced it is needed for the initial implementation.

As I expressed yesterday, I won't tolerate anything in hwpoison code that looks
like a hack.

-- 
Cheers,

David


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday
  2026-09-24 10:55           ` David Hildenbrand (Arm)
@ 2026-09-24 11:19             ` Kiryl Shutsemau
  2026-09-24 11:23               ` David Hildenbrand (Arm)
  0 siblings, 1 reply; 9+ messages in thread
From: Kiryl Shutsemau @ 2026-09-24 11:19 UTC (permalink / raw)
  To: David Hildenbrand (Arm)
  Cc: Harry Yoo, David Rientjes, Amit Shah, Andrew Morton, Aneesh Kumar,
	Christoph Lameter, Dave Hansen, Davidlohr Bueso, Hugh Dickins,
	Johannes Weiner, John Hubbard, Matthew Wilcox, Mel Gorman,
	Michal Hocko, Mike Rapoport, Peter Xu, Raghavendra K T,
	Rao, Bharata Bhasker, Rik van Riel, Roman Gushchin, Shakeel Butt,
	Shivank Garg, Sterling Alexander, Suren Baghdasaryan, Tejun Heo,
	Vlastimil Babka, Yang Shi, Zi Yan, William Roche, linmiaohe, ljs,
	osalvador, nao.horiguchi, tony.luck, wangkefeng.wang, jane.chu,
	muchun.song, liam, shuah, boudewijn, linux-mm, Breno Leitao

On Thu, Sep 24, 2026 at 12:55:10PM +0200, David Hildenbrand (Arm) wrote:
> On 9/24/26 12:46, Kiryl Shutsemau wrote:
> > On Thu, Sep 24, 2026 at 10:32:29AM +0200, David Hildenbrand (Arm) wrote:
> >> On 9/23/26 16:36, Harry Yoo wrote:
> >>>
> >>> I believe this is what the patchset actually did in v2.
> >>>
> >>> That way we'll have to either
> >>> 1) make any architecture that supports kexec select ARCH_KEEP_MEMBLOCK,
> >>> or 2) make kexec scan the bitmap when allocating the memory.
> >>
> >> My naive design would be:
> >>
> >> 1) Teach memblock early about poisoned memory ranges. Don't let it hand them out.
> > 
> > I suggested using a bitmap because ranges are not scalable. Some memory
> > failure modes generate errors repeated across the physical address space
> > (think of a column failure, for instance). It would produce too many
> > ranges to be viable.
> 
> These are in 2M granularity. Is that a real problem?

2M is arbitrary. We can trade finer granularity for the size of the bitmap.

> > We already consult a bitmap in memblock for unaccepted memory. We *can*
> > do the same for poisoned memory.
> > 
> 
> We can do many things. I don't want it.

What's wrong with it? The last step on allocation from memblock check if
the memory is poisoned in bitmap. If it is, try again next range,
starting from the next non-poisoned address.

-- 
  Kiryl Shutsemau / Kirill A. Shutemov


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday
  2026-09-24 11:19             ` Kiryl Shutsemau
@ 2026-09-24 11:23               ` David Hildenbrand (Arm)
  0 siblings, 0 replies; 9+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-24 11:23 UTC (permalink / raw)
  To: Kiryl Shutsemau
  Cc: Harry Yoo, David Rientjes, Amit Shah, Andrew Morton, Aneesh Kumar,
	Christoph Lameter, Dave Hansen, Davidlohr Bueso, Hugh Dickins,
	Johannes Weiner, John Hubbard, Matthew Wilcox, Mel Gorman,
	Michal Hocko, Mike Rapoport, Peter Xu, Raghavendra K T,
	Rao, Bharata Bhasker, Rik van Riel, Roman Gushchin, Shakeel Butt,
	Shivank Garg, Sterling Alexander, Suren Baghdasaryan, Tejun Heo,
	Vlastimil Babka, Yang Shi, Zi Yan, William Roche, linmiaohe, ljs,
	osalvador, nao.horiguchi, tony.luck, wangkefeng.wang, jane.chu,
	muchun.song, liam, shuah, boudewijn, linux-mm, Breno Leitao

On 9/24/26 13:19, Kiryl Shutsemau wrote:
> On Thu, Sep 24, 2026 at 12:55:10PM +0200, David Hildenbrand (Arm) wrote:
>> On 9/24/26 12:46, Kiryl Shutsemau wrote:
>>>
>>> I suggested using a bitmap because ranges are not scalable. Some memory
>>> failure modes generate errors repeated across the physical address space
>>> (think of a column failure, for instance). It would produce too many
>>> ranges to be viable.
>>
>> These are in 2M granularity. Is that a real problem?
> 
> 2M is arbitrary. We can trade finer granularity for the size of the bitmap.

Sure. And if you run into scalability problems you can go to coarse granularity.

> 
>>> We already consult a bitmap in memblock for unaccepted memory. We *can*
>>> do the same for poisoned memory.
>>>
>>
>> We can do many things. I don't want it.
> 
> What's wrong with it? The last step on allocation from memblock check if
> the memory is poisoned in bitmap. If it is, try again next range,
> starting from the next non-poisoned address.

See my reply to Breno on the other thread.

-- 
Cheers,

David


^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-09-24 11:23 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-21 16:26 [Invitation] Linux MM Alignment Session on Hwpoison on Wednesday David Rientjes
2026-09-23 12:57 ` Harry Yoo
2026-09-23 13:37   ` David Hildenbrand (Arm)
2026-09-23 14:36     ` Harry Yoo
2026-09-24  8:32       ` David Hildenbrand (Arm)
2026-09-24 10:46         ` Kiryl Shutsemau
2026-09-24 10:55           ` David Hildenbrand (Arm)
2026-09-24 11:19             ` Kiryl Shutsemau
2026-09-24 11:23               ` David Hildenbrand (Arm)

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox