* [PATCH] mm/mm_init: fix out-of-range first_deferred_pfn
@ 2026-08-05 22:44 Alexander Graf
2026-08-06 5:47 ` Mike Rapoport
0 siblings, 1 reply; 2+ messages in thread
From: Alexander Graf @ 2026-08-05 22:44 UTC (permalink / raw)
To: Andrew Morton, Mike Rapoport
Cc: David Hildenbrand, Wei Yang, linux-mm, linux-kernel,
nh-open-source
deferred_grow_zone() initializes deferred struct pages a section at a
time until the allocation that called into it can be satisfied, and
records where to resume in pgdat->first_deferred_pfn. Reserve most of
the top zone early (a large CMA reservation is the easy way) and a single
early allocation has to walk the whole zone instead of stopping in its
first free section. If that zone does not end on a section boundary, the
pgdatinit kthread then dies:
kernel BUG at mm/mm_init.c:2131!
Oops: invalid opcode: 0000 [#1] SMP NOPTI
CPU: 3 UID: 0 PID: 36 Comm: pgdatinit0 Not tainted 7.2.0-rc6 #1
RIP: 0010:deferred_init_memmap+0x1b8/0x1c0
RAX: 0000000000236000 R13: 0000000000238000
Call Trace:
kthread+0xdf/0x120
ret_from_fork+0x187/0x250
RAX is pgdat_end_pfn(), R13 the pfn that was stored. The loop advances
spfn in whole PAGES_PER_SECTION steps and only tests it before entering
an iteration, so once the walk reaches the end of a zone that ends
mid-section the escaping spfn is SECTION_ALIGN_UP(zone_end_pfn()).
Commit 3acb913c9d5b ("mm/mm_init: use deferred_init_memmap_chunk() in
deferred_grow_zone()") dropped the clamp that used to prevent that: epfn
came from __next_mem_pfn_range_in_zone(), since removed, which capped it
with min(zone_end_pfn(zone), epfn), so spfn could reach zone_end_pfn but
never pass it. The assert is fatal either way, panicking under
panic_on_oops and otherwise leaving page_alloc_init_late() waiting
forever for a completion the dead kthread never reports.
Store ULONG_MAX once spfn has left the zone. Nothing is left
uninitialized: the loop covers a single gap-free interval, and because it
only enters with spfn < zone_end_pfn() the escaping value is exactly
SECTION_ALIGN_UP(zone_end_pfn()), which is the last_pfn that
deferred_init_memmap() would have used for the same pfn range. A zone
that does end section-aligned now takes this path too and loses its
zero-work padata job along with that node's pr_info() and the WARN_ON()
on the next zone.
To reproduce with CONFIG_DEFERRED_STRUCT_PAGE_INIT=y and CONFIG_CMA=y:
qemu-system-x86_64 -enable-kvm -m 8032M -kernel bzImage \
-append "nokaslr cma=4768M@0x100000000"
Top of RAM is then 0x235ffffff, so ZONE_NORMAL ends 96 MiB into its last
section, and less than a section stays free above the reservation once
the early memblock allocations are done. The walk therefore runs off the
end of the zone and stores 0x238000. Sweeping that free remainder from
96M to 288M in 16M steps, an unpatched kernel dies on 7 of the 13 boots
and a patched one on none. On an 8 GiB cloud instance that reserves most
of its top zone for a memory pool, roughly one boot in three panicked
before reaching userspace.
Fixes: 3acb913c9d5b ("mm/mm_init: use deferred_init_memmap_chunk() in deferred_grow_zone()")
Cc: stable@vger.kernel.org
Assisted-by: Kiro:claude-opus-5
Signed-off-by: Alexander Graf <graf@amazon.com>
---
Notes:
Applies unchanged to 6.18.y, 6.19.y, 7.0.y and 7.1.y (checked against
v6.18.39, v6.19.14, v7.0.14 and v7.1.4); the deferred_init_memmap_chunk()
signature change in cbbbf7795fc3 sits outside the hunk context, so stable
needs no separate backport.
First seen on 6.18.y and 6.19-rc distribution kernels. There is no public
report to link, hence no Closes: tag.
mm/mm_init.c | 9 ++++++---
1 file changed, 6 insertions(+), 3 deletions(-)
diff --git a/mm/mm_init.c b/mm/mm_init.c
index 498d62c4ece3..91177be58a00 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -2214,10 +2214,13 @@ bool __init deferred_grow_zone(struct zone *zone, unsigned int order)
}
/*
- * There were no pages to initialize and free which means the zone's
- * memory map is completely initialized.
+ * The loop only tests spfn before entering an iteration, so on exit it
+ * may point up to a section past the end of the zone. When it does,
+ * the rest of the zone has already been handed to
+ * deferred_init_memmap_chunk() and nothing is left to initialize.
*/
- pgdat->first_deferred_pfn = nr_pages ? spfn : ULONG_MAX;
+ pgdat->first_deferred_pfn =
+ spfn < zone_end_pfn(zone) ? spfn : ULONG_MAX;
pgdat_resize_unlock(pgdat, &flags);
base-commit: 0d839570765118029aa8bf4a95444c6a11aacf85
--
2.47.1
^ permalink raw reply related [flat|nested] 2+ messages in thread
* Re: [PATCH] mm/mm_init: fix out-of-range first_deferred_pfn
2026-08-05 22:44 [PATCH] mm/mm_init: fix out-of-range first_deferred_pfn Alexander Graf
@ 2026-08-06 5:47 ` Mike Rapoport
0 siblings, 0 replies; 2+ messages in thread
From: Mike Rapoport @ 2026-08-06 5:47 UTC (permalink / raw)
To: Alexander Graf
Cc: Andrew Morton, David Hildenbrand, Wei Yang, linux-mm,
linux-kernel, nh-open-source
Hi Alex,
On Wed, Aug 05, 2026 at 10:44:21PM +0000, Alexander Graf wrote:
> deferred_grow_zone() initializes deferred struct pages a section at a
> time until the allocation that called into it can be satisfied, and
> records where to resume in pgdat->first_deferred_pfn. Reserve most of
> the top zone early (a large CMA reservation is the easy way) and a single
> early allocation has to walk the whole zone instead of stopping in its
> first free section. If that zone does not end on a section boundary, the
> pgdatinit kthread then dies:
>
> kernel BUG at mm/mm_init.c:2131!
> Oops: invalid opcode: 0000 [#1] SMP NOPTI
> CPU: 3 UID: 0 PID: 36 Comm: pgdatinit0 Not tainted 7.2.0-rc6 #1
> RIP: 0010:deferred_init_memmap+0x1b8/0x1c0
> RAX: 0000000000236000 R13: 0000000000238000
> Call Trace:
> kthread+0xdf/0x120
> ret_from_fork+0x187/0x250
>
> RAX is pgdat_end_pfn(), R13 the pfn that was stored. The loop advances
> spfn in whole PAGES_PER_SECTION steps and only tests it before entering
> an iteration, so once the walk reaches the end of a zone that ends
> mid-section the escaping spfn is SECTION_ALIGN_UP(zone_end_pfn()).
> Commit 3acb913c9d5b ("mm/mm_init: use deferred_init_memmap_chunk() in
> deferred_grow_zone()") dropped the clamp that used to prevent that: epfn
> came from __next_mem_pfn_range_in_zone(), since removed, which capped it
> with min(zone_end_pfn(zone), epfn), so spfn could reach zone_end_pfn but
> never pass it. The assert is fatal either way, panicking under
> panic_on_oops and otherwise leaving page_alloc_init_late() waiting
> forever for a completion the dead kthread never reports.
>
> Store ULONG_MAX once spfn has left the zone. Nothing is left
> uninitialized: the loop covers a single gap-free interval, and because it
> only enters with spfn < zone_end_pfn() the escaping value is exactly
> SECTION_ALIGN_UP(zone_end_pfn()), which is the last_pfn that
> deferred_init_memmap() would have used for the same pfn range. A zone
> that does end section-aligned now takes this path too and loses its
> zero-work padata job along with that node's pr_info() and the WARN_ON()
> on the next zone.
>
> To reproduce with CONFIG_DEFERRED_STRUCT_PAGE_INIT=y and CONFIG_CMA=y:
>
> qemu-system-x86_64 -enable-kvm -m 8032M -kernel bzImage \
> -append "nokaslr cma=4768M@0x100000000"
The only two paragraphs I understood is this and the BUG splat ;-P
Can we please have a lot more of human touch on the changelog?
> Top of RAM is then 0x235ffffff, so ZONE_NORMAL ends 96 MiB into its last
> section, and less than a section stays free above the reservation once
> the early memblock allocations are done. The walk therefore runs off the
> end of the zone and stores 0x238000. Sweeping that free remainder from
> 96M to 288M in 16M steps, an unpatched kernel dies on 7 of the 13 boots
> and a patched one on none. On an 8 GiB cloud instance that reserves most
> of its top zone for a memory pool, roughly one boot in three panicked
> before reaching userspace.
>
> Fixes: 3acb913c9d5b ("mm/mm_init: use deferred_init_memmap_chunk() in deferred_grow_zone()")
> Cc: stable@vger.kernel.org
> Assisted-by: Kiro:claude-opus-5
> Signed-off-by: Alexander Graf <graf@amazon.com>
> ---
>
> Notes:
> Applies unchanged to 6.18.y, 6.19.y, 7.0.y and 7.1.y (checked against
> v6.18.39, v6.19.14, v7.0.14 and v7.1.4); the deferred_init_memmap_chunk()
> signature change in cbbbf7795fc3 sits outside the hunk context, so stable
> needs no separate backport.
>
> First seen on 6.18.y and 6.19-rc distribution kernels. There is no public
> report to link, hence no Closes: tag.
>
> mm/mm_init.c | 9 ++++++---
> 1 file changed, 6 insertions(+), 3 deletions(-)
>
> diff --git a/mm/mm_init.c b/mm/mm_init.c
> index 498d62c4ece3..91177be58a00 100644
> --- a/mm/mm_init.c
> +++ b/mm/mm_init.c
> @@ -2214,10 +2214,13 @@ bool __init deferred_grow_zone(struct zone *zone, unsigned int order)
> }
>
> /*
> - * There were no pages to initialize and free which means the zone's
> - * memory map is completely initialized.
> + * The loop only tests spfn before entering an iteration, so on exit it
> + * may point up to a section past the end of the zone. When it does,
> + * the rest of the zone has already been handed to
> + * deferred_init_memmap_chunk() and nothing is left to initialize.
> */
> - pgdat->first_deferred_pfn = nr_pages ? spfn : ULONG_MAX;
> + pgdat->first_deferred_pfn =
> + spfn < zone_end_pfn(zone) ? spfn : ULONG_MAX;
>
> pgdat_resize_unlock(pgdat, &flags);
>
>
> base-commit: 0d839570765118029aa8bf4a95444c6a11aacf85
> --
> 2.47.1
>
--
Sincerely yours,
Mike.
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-08-06 5:47 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-05 22:44 [PATCH] mm/mm_init: fix out-of-range first_deferred_pfn Alexander Graf
2026-08-06 5:47 ` Mike Rapoport
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox