All of lore.kernel.org
 help / color / mirror / Atom feed
From: Muchun Song <muchun.song@linux.dev>
To: Mike Rapoport <rppt@kernel.org>
Cc: Muchun Song <songmuchun@bytedance.com>,
	Andrew Morton <akpm@linux-foundation.org>,
	Oscar Salvador <osalvador@suse.de>,
	David Hildenbrand <david@kernel.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	Lorenzo Stoakes <ljs@kernel.org>, Michal Hocko <mhocko@suse.com>,
	David Laight <david.laight.linux@gmail.com>,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v2 03/17] mm/mm_init: skip initializing shared vmemmap tail pages
Date: Thu, 30 Jul 2026 21:09:00 +0800	[thread overview]
Message-ID: <3FCC8989-95DB-450F-95C3-183787DF4D08@linux.dev> (raw)
In-Reply-To: <178540851830.2154539.6359871825839331816.b4-reply@b4>



> On Jul 30, 2026, at 18:48, Mike Rapoport <rppt@kernel.org> wrote:
> 
> Hi Muchun,

Hi,

> 
> On 2026-07-26 12:30:17+08:00, Muchun Song wrote:
>>> On Jul 20, 2026, at 17:31, Muchun Song <songmuchun@bytedance.com> wrote:
>>> 
>>> memmap_init_range() initializes every struct page in the target range.
>>> For compound pages with vmemmap optimization, the tail struct pages are
>>> backed by a shared vmemmap page.
>>> 
>>> Initializing those tail struct pages would overwrite the shared
>>> vmemmap page contents, requiring users such as HugeTLB to restore the
>>> metadata afterwards.
>>> 
>>> Track the compound order for HVO-backed sections and use that metadata
>>> to detect struct pages that fall into the shared tail vmemmap range.
>>> Skip those shared tail pages in memmap_init_range(), then initialize
>>> pageblock migratetypes for the processed range with a helper after the
>>> per-page initialization loop.
>>> 
>>> The !SPARSEMEM __pfn_to_section() stub is needed only for the build:
>>> memmap_init_range() references __pfn_to_section() after checking
>>> pfn_vmemmap_optimizable(), and !SPARSEMEM builds still have to compile
>>> that code even though pfn_vmemmap_optimizable() folds to false.
>>> 
>>> This is a preparatory change for consolidating handling across users of
>>> vmemmap optimization, and it also avoids redundant initialization of
>>> shared tail vmemmap pages during early boot.
>>> 
>>> Signed-off-by: Muchun Song <songmuchun@bytedance.com>
>>> ---
>>> v2:
>>> - Fold section order tracking into the first user instead of keeping a
>>> standalone API-only patch (suggested by Mike Rapoport)
>>> - Rename page_vmemmap_optimizable() to pfn_vmemmap_optimizable() and
>>> pass a PFN directly (suggested by Mike Rapoport)
>>> - Initialize pageblock migratetypes from a helper after the per-page
>>> loop (suggested by Mike Rapoport)
>>> - Use a 1G PFN chunk for cond_resched() in the pageblock helper
>>> (suggested by Mike Rapoport)
>>> - Guard section_order() with CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP so
>>> it returns 0 when HVO is disabled and lets the compiler optimize the
>>> code as much as possible (suggested by Mike Rapoport)
>>> - Explain why the !SPARSEMEM __pfn_to_section() stub belongs here
>>> (suggested by Mike Rapoport)
>>> ---
>>> include/linux/mmzone.h | 14 ++++++++++++++
>>> mm/mm_init.c           | 33 +++++++++++++++++----------------
>>> mm/sparse.h            | 23 +++++++++++++++++++++++
>>> 3 files changed, 54 insertions(+), 16 deletions(-)
>>> 
>>> diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
>>> index 82b0155d886f..2a32101d55e6 100644
>>> --- a/include/linux/mmzone.h
>>> +++ b/include/linux/mmzone.h
>>> @@ -2011,6 +2011,14 @@ struct mem_section {
>>> unsigned long section_mem_map;
>>> 
>>> struct mem_section_usage *usage;
>>> +#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP
>>> +  	/*
>>> + 	 * Normally, sections hold regular (order-0) pages. However, for
>>> + 	 * sections with HVO enabled, this tracks the compound page order
>>> + 	 * to enable deduplication of redundant vmemmap pages.
>>> + 	 */
>>> +  	unsigned int order;
>>> +#endif
>>> #ifdef CONFIG_PAGE_EXTENSION
>>> /*
>>>  * If SPARSEMEM, pgdat doesn't have page_ext pointer. We use
>>> @@ -2365,8 +2373,14 @@ static inline unsigned long next_present_section_nr(unsigned long section_nr)
>>> #endif
>>> 
>>> #else
>>> +struct mem_section;
>>> +
>>> #define sparse_vmemmap_init_nid_early(_nid) do {} while (0)
>>> #define pfn_in_present_section pfn_valid
>>> +static inline struct mem_section *__pfn_to_section(unsigned long pfn)
>>> +{
>>> +  	return NULL;
>>> +}
>> 
>> I'd like to propose an alternative implementation that doesn't require
>> exposing the mem_section. The idea is to add a new helper function,
>> pfn_to_section_order(), so that for non-sparse-memory configurations,
>> the mem_section concept stays hidden internally. I'd really appreciate
>> any thoughts or concerns — if everyone is comfortable with it, I can go
>> ahead and implement this in the next version.
> 
> A helper that keeps mem_section hidden from !SPARSMEM makes perfect
> sense to me.
> 
> I'd even take it one step further and make it return how many pfns
> should be skipped in pfn_vmemmap_optimizable case.

To make sure we're on the same page, let me walk you through the specific
changes I have in mind. My initial plan is to introduce pfn_to_section_order,
and the expected diff changes are as follow to keep mem_sectionhidden from
!SPARSEMEM.

diff --git a/mm/mm_init.c b/mm/mm_init.c
index dcb757b36902..0b0c2996d080 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -884,7 +884,7 @@ void __meminit memmap_init_range(unsigned long size, int nid, unsigned long zone
                }

                if (pfn_vmemmap_optimizable(pfn)) {
-                       unsigned int order = section_order(__pfn_to_section(pfn));
+                       unsigned int order = pfn_to_section_order(pfn);

                        pfn = min(ALIGN(pfn, 1UL << order), end_pfn);
                        continue;
diff --git a/mm/sparse.h b/mm/sparse.h
index 030248030dc7..c5fbcdde3cee 100644
--- a/mm/sparse.h
+++ b/mm/sparse.h
@@ -47,6 +47,11 @@ static inline void __section_mark_present(struct mem_section *ms,

        ms->section_mem_map |= SECTION_MARKED_PRESENT;
 }
+
+static inline unsigned int pfn_to_section_order(unsigned long pfn)
+{
+       return section_order(__pfn_to_section(pfn));
+}
 #else
 static inline void sparse_init(void) {}
 #endif /* CONFIG_SPARSEMEM */

Since we also use __pfn_to_section in the patch 14 in this series for
!SPARSEMEM, we need to make corresponding adjustments—specifically, by using
pfn_to_section_order to determine whether the vmemmap of a given section is
optimizable. This new helper will be called from several places, so I'm afraid
its introduction is unavoidable.

That said, I've also considered an alternative: introducing another helper that
returns the exact number of PFNs to skip, and using it solely within
memmap_init_range(). However, that approach doesn't seem to offer much in terms
of code simplification. If I'm missing something or if my reasoning doesn't align
with your expectations, I would really appreciate your guidance. Thank you for
your patience!

Thanks,
Muchun





  reply	other threads:[~2026-07-30 13:09 UTC|newest]

Thread overview: 27+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-20  9:31 [PATCH v2 00/17] mm: Introduce section-based vmemmap optimization for HugeTLB Muchun Song
2026-07-20  9:31 ` [PATCH v2 01/17] mm/sparse: relax struct mem_section size constraints Muchun Song
2026-07-20  9:31 ` [PATCH v2 02/17] mm/sparse-vmemmap: rename HVO order macros Muchun Song
2026-07-30 10:57   ` Mike Rapoport
2026-07-20  9:31 ` [PATCH v2 03/17] mm/mm_init: skip initializing shared vmemmap tail pages Muchun Song
2026-07-26  4:30   ` Muchun Song
2026-07-30 10:48     ` Mike Rapoport
2026-07-30 13:09       ` Muchun Song [this message]
2026-07-30 14:32         ` Mike Rapoport
2026-07-20  9:31 ` [PATCH v2 04/17] mm/sparse-vmemmap: initialize shared tail vmemmap pages on allocation Muchun Song
2026-07-20  9:31 ` [PATCH v2 05/17] mm/sparse-vmemmap: support section-based vmemmap accounting Muchun Song
2026-07-20  9:31 ` [PATCH v2 06/17] mm/mm_init: factor out pfn_to_zone() Muchun Song
2026-07-20 11:04   ` Muchun Song
2026-07-20 11:05   ` Muchun Song
2026-07-30 10:57   ` Mike Rapoport
2026-07-20  9:31 ` [PATCH v2 07/17] mm/sparse-vmemmap: move vmemmap_get_tail() before PTE population Muchun Song
2026-07-20  9:31 ` [PATCH v2 08/17] mm/sparse-vmemmap: support section-based vmemmap optimization Muchun Song
2026-07-20 11:10   ` Muchun Song
2026-07-20  9:31 ` [PATCH v2 09/17] mm/sparse: initialize memory sections earlier Muchun Song
2026-07-20  9:31 ` [PATCH v2 10/17] mm/hugetlb: switch HugeTLB to section-based vmemmap optimization Muchun Song
2026-07-20  9:31 ` [PATCH v2 11/17] mm/sparse-vmemmap: remove SPARSEMEM_VMEMMAP_PREINIT support Muchun Song
2026-07-20  9:31 ` [PATCH v2 12/17] mm/sparse: inline usemap allocation into sparse_init_nid() Muchun Song
2026-07-20  9:31 ` [PATCH v2 13/17] mm/sparse: remove section_map_size() Muchun Song
2026-07-20  9:31 ` [PATCH v2 14/17] mm/hugetlb: remove HUGE_BOOTMEM_HVO Muchun Song
2026-07-20  9:31 ` [PATCH v2 15/17] mm/hugetlb: remove HUGE_BOOTMEM_CMA Muchun Song
2026-07-20  9:31 ` [PATCH v2 16/17] mm/hugetlb: localize struct huge_bootmem_page Muchun Song
2026-07-20  9:31 ` [PATCH v2 17/17] mm/hugetlb: localize HUGE_BOOTMEM_ZONES_VALID Muchun Song

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=3FCC8989-95DB-450F-95C3-183787DF4D08@linux.dev \
    --to=muchun.song@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=david.laight.linux@gmail.com \
    --cc=david@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=osalvador@suse.de \
    --cc=rppt@kernel.org \
    --cc=songmuchun@bytedance.com \
    --cc=vbabka@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.