From: Muchun Song <muchun.song@linux.dev>
To: Mike Rapoport <rppt@kernel.org>
Cc: Muchun Song <songmuchun@bytedance.com>,
Andrew Morton <akpm@linux-foundation.org>,
Oscar Salvador <osalvador@suse.de>,
David Hildenbrand <david@kernel.org>,
Vlastimil Babka <vbabka@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>, Michal Hocko <mhocko@suse.com>,
David Laight <david.laight.linux@gmail.com>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v2 03/17] mm/mm_init: skip initializing shared vmemmap tail pages
Date: Thu, 30 Jul 2026 21:09:00 +0800 [thread overview]
Message-ID: <3FCC8989-95DB-450F-95C3-183787DF4D08@linux.dev> (raw)
In-Reply-To: <178540851830.2154539.6359871825839331816.b4-reply@b4>
> On Jul 30, 2026, at 18:48, Mike Rapoport <rppt@kernel.org> wrote:
>
> Hi Muchun,
Hi,
>
> On 2026-07-26 12:30:17+08:00, Muchun Song wrote:
>>> On Jul 20, 2026, at 17:31, Muchun Song <songmuchun@bytedance.com> wrote:
>>>
>>> memmap_init_range() initializes every struct page in the target range.
>>> For compound pages with vmemmap optimization, the tail struct pages are
>>> backed by a shared vmemmap page.
>>>
>>> Initializing those tail struct pages would overwrite the shared
>>> vmemmap page contents, requiring users such as HugeTLB to restore the
>>> metadata afterwards.
>>>
>>> Track the compound order for HVO-backed sections and use that metadata
>>> to detect struct pages that fall into the shared tail vmemmap range.
>>> Skip those shared tail pages in memmap_init_range(), then initialize
>>> pageblock migratetypes for the processed range with a helper after the
>>> per-page initialization loop.
>>>
>>> The !SPARSEMEM __pfn_to_section() stub is needed only for the build:
>>> memmap_init_range() references __pfn_to_section() after checking
>>> pfn_vmemmap_optimizable(), and !SPARSEMEM builds still have to compile
>>> that code even though pfn_vmemmap_optimizable() folds to false.
>>>
>>> This is a preparatory change for consolidating handling across users of
>>> vmemmap optimization, and it also avoids redundant initialization of
>>> shared tail vmemmap pages during early boot.
>>>
>>> Signed-off-by: Muchun Song <songmuchun@bytedance.com>
>>> ---
>>> v2:
>>> - Fold section order tracking into the first user instead of keeping a
>>> standalone API-only patch (suggested by Mike Rapoport)
>>> - Rename page_vmemmap_optimizable() to pfn_vmemmap_optimizable() and
>>> pass a PFN directly (suggested by Mike Rapoport)
>>> - Initialize pageblock migratetypes from a helper after the per-page
>>> loop (suggested by Mike Rapoport)
>>> - Use a 1G PFN chunk for cond_resched() in the pageblock helper
>>> (suggested by Mike Rapoport)
>>> - Guard section_order() with CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP so
>>> it returns 0 when HVO is disabled and lets the compiler optimize the
>>> code as much as possible (suggested by Mike Rapoport)
>>> - Explain why the !SPARSEMEM __pfn_to_section() stub belongs here
>>> (suggested by Mike Rapoport)
>>> ---
>>> include/linux/mmzone.h | 14 ++++++++++++++
>>> mm/mm_init.c | 33 +++++++++++++++++----------------
>>> mm/sparse.h | 23 +++++++++++++++++++++++
>>> 3 files changed, 54 insertions(+), 16 deletions(-)
>>>
>>> diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
>>> index 82b0155d886f..2a32101d55e6 100644
>>> --- a/include/linux/mmzone.h
>>> +++ b/include/linux/mmzone.h
>>> @@ -2011,6 +2011,14 @@ struct mem_section {
>>> unsigned long section_mem_map;
>>>
>>> struct mem_section_usage *usage;
>>> +#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP
>>> + /*
>>> + * Normally, sections hold regular (order-0) pages. However, for
>>> + * sections with HVO enabled, this tracks the compound page order
>>> + * to enable deduplication of redundant vmemmap pages.
>>> + */
>>> + unsigned int order;
>>> +#endif
>>> #ifdef CONFIG_PAGE_EXTENSION
>>> /*
>>> * If SPARSEMEM, pgdat doesn't have page_ext pointer. We use
>>> @@ -2365,8 +2373,14 @@ static inline unsigned long next_present_section_nr(unsigned long section_nr)
>>> #endif
>>>
>>> #else
>>> +struct mem_section;
>>> +
>>> #define sparse_vmemmap_init_nid_early(_nid) do {} while (0)
>>> #define pfn_in_present_section pfn_valid
>>> +static inline struct mem_section *__pfn_to_section(unsigned long pfn)
>>> +{
>>> + return NULL;
>>> +}
>>
>> I'd like to propose an alternative implementation that doesn't require
>> exposing the mem_section. The idea is to add a new helper function,
>> pfn_to_section_order(), so that for non-sparse-memory configurations,
>> the mem_section concept stays hidden internally. I'd really appreciate
>> any thoughts or concerns — if everyone is comfortable with it, I can go
>> ahead and implement this in the next version.
>
> A helper that keeps mem_section hidden from !SPARSMEM makes perfect
> sense to me.
>
> I'd even take it one step further and make it return how many pfns
> should be skipped in pfn_vmemmap_optimizable case.
To make sure we're on the same page, let me walk you through the specific
changes I have in mind. My initial plan is to introduce pfn_to_section_order,
and the expected diff changes are as follow to keep mem_sectionhidden from
!SPARSEMEM.
diff --git a/mm/mm_init.c b/mm/mm_init.c
index dcb757b36902..0b0c2996d080 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -884,7 +884,7 @@ void __meminit memmap_init_range(unsigned long size, int nid, unsigned long zone
}
if (pfn_vmemmap_optimizable(pfn)) {
- unsigned int order = section_order(__pfn_to_section(pfn));
+ unsigned int order = pfn_to_section_order(pfn);
pfn = min(ALIGN(pfn, 1UL << order), end_pfn);
continue;
diff --git a/mm/sparse.h b/mm/sparse.h
index 030248030dc7..c5fbcdde3cee 100644
--- a/mm/sparse.h
+++ b/mm/sparse.h
@@ -47,6 +47,11 @@ static inline void __section_mark_present(struct mem_section *ms,
ms->section_mem_map |= SECTION_MARKED_PRESENT;
}
+
+static inline unsigned int pfn_to_section_order(unsigned long pfn)
+{
+ return section_order(__pfn_to_section(pfn));
+}
#else
static inline void sparse_init(void) {}
#endif /* CONFIG_SPARSEMEM */
Since we also use __pfn_to_section in the patch 14 in this series for
!SPARSEMEM, we need to make corresponding adjustments—specifically, by using
pfn_to_section_order to determine whether the vmemmap of a given section is
optimizable. This new helper will be called from several places, so I'm afraid
its introduction is unavoidable.
That said, I've also considered an alternative: introducing another helper that
returns the exact number of PFNs to skip, and using it solely within
memmap_init_range(). However, that approach doesn't seem to offer much in terms
of code simplification. If I'm missing something or if my reasoning doesn't align
with your expectations, I would really appreciate your guidance. Thank you for
your patience!
Thanks,
Muchun
next prev parent reply other threads:[~2026-07-30 13:09 UTC|newest]
Thread overview: 27+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-20 9:31 [PATCH v2 00/17] mm: Introduce section-based vmemmap optimization for HugeTLB Muchun Song
2026-07-20 9:31 ` [PATCH v2 01/17] mm/sparse: relax struct mem_section size constraints Muchun Song
2026-07-20 9:31 ` [PATCH v2 02/17] mm/sparse-vmemmap: rename HVO order macros Muchun Song
2026-07-30 10:57 ` Mike Rapoport
2026-07-20 9:31 ` [PATCH v2 03/17] mm/mm_init: skip initializing shared vmemmap tail pages Muchun Song
2026-07-26 4:30 ` Muchun Song
2026-07-30 10:48 ` Mike Rapoport
2026-07-30 13:09 ` Muchun Song [this message]
2026-07-30 14:32 ` Mike Rapoport
2026-07-20 9:31 ` [PATCH v2 04/17] mm/sparse-vmemmap: initialize shared tail vmemmap pages on allocation Muchun Song
2026-07-20 9:31 ` [PATCH v2 05/17] mm/sparse-vmemmap: support section-based vmemmap accounting Muchun Song
2026-07-20 9:31 ` [PATCH v2 06/17] mm/mm_init: factor out pfn_to_zone() Muchun Song
2026-07-20 11:04 ` Muchun Song
2026-07-20 11:05 ` Muchun Song
2026-07-30 10:57 ` Mike Rapoport
2026-07-20 9:31 ` [PATCH v2 07/17] mm/sparse-vmemmap: move vmemmap_get_tail() before PTE population Muchun Song
2026-07-20 9:31 ` [PATCH v2 08/17] mm/sparse-vmemmap: support section-based vmemmap optimization Muchun Song
2026-07-20 11:10 ` Muchun Song
2026-07-20 9:31 ` [PATCH v2 09/17] mm/sparse: initialize memory sections earlier Muchun Song
2026-07-20 9:31 ` [PATCH v2 10/17] mm/hugetlb: switch HugeTLB to section-based vmemmap optimization Muchun Song
2026-07-20 9:31 ` [PATCH v2 11/17] mm/sparse-vmemmap: remove SPARSEMEM_VMEMMAP_PREINIT support Muchun Song
2026-07-20 9:31 ` [PATCH v2 12/17] mm/sparse: inline usemap allocation into sparse_init_nid() Muchun Song
2026-07-20 9:31 ` [PATCH v2 13/17] mm/sparse: remove section_map_size() Muchun Song
2026-07-20 9:31 ` [PATCH v2 14/17] mm/hugetlb: remove HUGE_BOOTMEM_HVO Muchun Song
2026-07-20 9:31 ` [PATCH v2 15/17] mm/hugetlb: remove HUGE_BOOTMEM_CMA Muchun Song
2026-07-20 9:31 ` [PATCH v2 16/17] mm/hugetlb: localize struct huge_bootmem_page Muchun Song
2026-07-20 9:31 ` [PATCH v2 17/17] mm/hugetlb: localize HUGE_BOOTMEM_ZONES_VALID Muchun Song
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=3FCC8989-95DB-450F-95C3-183787DF4D08@linux.dev \
--to=muchun.song@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=david.laight.linux@gmail.com \
--cc=david@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=osalvador@suse.de \
--cc=rppt@kernel.org \
--cc=songmuchun@bytedance.com \
--cc=vbabka@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox