linux-mm.kvack.org archive mirror
 help / color / mirror / Atom feed
From: Mike Rapoport <rppt@kernel.org>
To: Alexander Graf <graf@amazon.com>
Cc: Andrew Morton <akpm@linux-foundation.org>,
	David Hildenbrand <david@kernel.org>,
	Wei Yang <richard.weiyang@gmail.com>,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	nh-open-source@amazon.com
Subject: Re: [PATCH] mm/mm_init: fix out-of-range first_deferred_pfn
Date: Thu, 6 Aug 2026 08:47:06 +0300	[thread overview]
Message-ID: <anQf2kFzA63_3ayh@kernel.org> (raw)
In-Reply-To: <20260805224421.15794-1-graf@amazon.com>

Hi Alex,

On Wed, Aug 05, 2026 at 10:44:21PM +0000, Alexander Graf wrote:
> deferred_grow_zone() initializes deferred struct pages a section at a
> time until the allocation that called into it can be satisfied, and
> records where to resume in pgdat->first_deferred_pfn.  Reserve most of
> the top zone early (a large CMA reservation is the easy way) and a single
> early allocation has to walk the whole zone instead of stopping in its
> first free section.  If that zone does not end on a section boundary, the
> pgdatinit kthread then dies:
> 
>   kernel BUG at mm/mm_init.c:2131!
>   Oops: invalid opcode: 0000 [#1] SMP NOPTI
>   CPU: 3 UID: 0 PID: 36 Comm: pgdatinit0 Not tainted 7.2.0-rc6 #1
>   RIP: 0010:deferred_init_memmap+0x1b8/0x1c0
>   RAX: 0000000000236000 R13: 0000000000238000
>   Call Trace:
>    kthread+0xdf/0x120
>    ret_from_fork+0x187/0x250
> 
> RAX is pgdat_end_pfn(), R13 the pfn that was stored.  The loop advances
> spfn in whole PAGES_PER_SECTION steps and only tests it before entering
> an iteration, so once the walk reaches the end of a zone that ends
> mid-section the escaping spfn is SECTION_ALIGN_UP(zone_end_pfn()).
> Commit 3acb913c9d5b ("mm/mm_init: use deferred_init_memmap_chunk() in
> deferred_grow_zone()") dropped the clamp that used to prevent that: epfn
> came from __next_mem_pfn_range_in_zone(), since removed, which capped it
> with min(zone_end_pfn(zone), epfn), so spfn could reach zone_end_pfn but
> never pass it.  The assert is fatal either way, panicking under
> panic_on_oops and otherwise leaving page_alloc_init_late() waiting
> forever for a completion the dead kthread never reports.
> 
> Store ULONG_MAX once spfn has left the zone.  Nothing is left
> uninitialized: the loop covers a single gap-free interval, and because it
> only enters with spfn < zone_end_pfn() the escaping value is exactly
> SECTION_ALIGN_UP(zone_end_pfn()), which is the last_pfn that
> deferred_init_memmap() would have used for the same pfn range.  A zone
> that does end section-aligned now takes this path too and loses its
> zero-work padata job along with that node's pr_info() and the WARN_ON()
> on the next zone.
> 
> To reproduce with CONFIG_DEFERRED_STRUCT_PAGE_INIT=y and CONFIG_CMA=y:
> 
>   qemu-system-x86_64 -enable-kvm -m 8032M -kernel bzImage \
>       -append "nokaslr cma=4768M@0x100000000"

The only two paragraphs I understood is this and the BUG splat ;-P

Can we please have a lot more of human touch on the changelog?
 
> Top of RAM is then 0x235ffffff, so ZONE_NORMAL ends 96 MiB into its last
> section, and less than a section stays free above the reservation once
> the early memblock allocations are done.  The walk therefore runs off the
> end of the zone and stores 0x238000.  Sweeping that free remainder from
> 96M to 288M in 16M steps, an unpatched kernel dies on 7 of the 13 boots
> and a patched one on none.  On an 8 GiB cloud instance that reserves most
> of its top zone for a memory pool, roughly one boot in three panicked
> before reaching userspace.
> 
> Fixes: 3acb913c9d5b ("mm/mm_init: use deferred_init_memmap_chunk() in deferred_grow_zone()")
> Cc: stable@vger.kernel.org
> Assisted-by: Kiro:claude-opus-5
> Signed-off-by: Alexander Graf <graf@amazon.com>
> ---
> 
> Notes:
>     Applies unchanged to 6.18.y, 6.19.y, 7.0.y and 7.1.y (checked against
>     v6.18.39, v6.19.14, v7.0.14 and v7.1.4); the deferred_init_memmap_chunk()
>     signature change in cbbbf7795fc3 sits outside the hunk context, so stable
>     needs no separate backport.
>     
>     First seen on 6.18.y and 6.19-rc distribution kernels.  There is no public
>     report to link, hence no Closes: tag.
> 
>  mm/mm_init.c | 9 ++++++---
>  1 file changed, 6 insertions(+), 3 deletions(-)
> 
> diff --git a/mm/mm_init.c b/mm/mm_init.c
> index 498d62c4ece3..91177be58a00 100644
> --- a/mm/mm_init.c
> +++ b/mm/mm_init.c
> @@ -2214,10 +2214,13 @@ bool __init deferred_grow_zone(struct zone *zone, unsigned int order)
>  	}
>  
>  	/*
> -	 * There were no pages to initialize and free which means the zone's
> -	 * memory map is completely initialized.
> +	 * The loop only tests spfn before entering an iteration, so on exit it
> +	 * may point up to a section past the end of the zone.  When it does,
> +	 * the rest of the zone has already been handed to
> +	 * deferred_init_memmap_chunk() and nothing is left to initialize.
>  	 */
> -	pgdat->first_deferred_pfn = nr_pages ? spfn : ULONG_MAX;
> +	pgdat->first_deferred_pfn =
> +		spfn < zone_end_pfn(zone) ? spfn : ULONG_MAX;
>  
>  	pgdat_resize_unlock(pgdat, &flags);
>  
> 
> base-commit: 0d839570765118029aa8bf4a95444c6a11aacf85
> -- 
> 2.47.1
> 

-- 
Sincerely yours,
Mike.


  reply	other threads:[~2026-08-06  5:47 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-05 22:44 [PATCH] mm/mm_init: fix out-of-range first_deferred_pfn Alexander Graf
2026-08-06  5:47 ` Mike Rapoport [this message]
2026-08-07  2:43   ` Graf (AWS), Alexander

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anQf2kFzA63_3ayh@kernel.org \
    --to=rppt@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=david@kernel.org \
    --cc=graf@amazon.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=nh-open-source@amazon.com \
    --cc=richard.weiyang@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).