From: Muchun Song <muchun.song@linux.dev>
To: Muchun Song <songmuchun@bytedance.com>
Cc: Mike Rapoport <rppt@kernel.org>,
Vlastimil Babka <vbabka@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>, Michal Hocko <mhocko@suse.com>,
David Laight <david.laight.linux@gmail.com>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
Andrew Morton <akpm@linux-foundation.org>,
Oscar Salvador <osalvador@suse.de>,
David Hildenbrand <david@kernel.org>
Subject: Re: [PATCH v3 08/17] mm/sparse-vmemmap: support section-based vmemmap optimization
Date: Thu, 6 Aug 2026 11:04:35 +0800 [thread overview]
Message-ID: <978060e0-a202-43b2-b84d-b1950b2d28c1@linux.dev> (raw)
In-Reply-To: <20260804035535.2846016-9-songmuchun@bytedance.com>
On 2026/8/4 11:55, Muchun Song wrote:
> Teach sparse-vmemmap population code to use the compound page order
> when deciding whether a vmemmap page can be optimized.
>
> With this information, the common sparse-vmemmap population path can
> allocate or reuse shared tail vmemmap pages directly instead of relying
> on HugeTLB-specific handling.
>
> This centralizes vmemmap optimization logic in the sparse-vmemmap code,
> based on section metadata, and prepares for sharing the same mechanism
> across different users of vmemmap optimization, including HugeTLB and
> DAX.
>
> Signed-off-by: Muchun Song <songmuchun@bytedance.com>
> ---
> v2:
> - Keep vmemmap accounting and population logic in sparse-vmemmap.c
> (suggested by Mike Rapoport)
> - Move vmemmap_get_tail() before its first use instead of adding only a
> forward declaration in the previous patch (suggested by Mike Rapoport)
> - Simplify the PMD path handling for HVO-covered sections
> ---
> mm/sparse-vmemmap.c | 36 ++++++++++++++++++++++++++++++------
> mm/sparse.c | 4 ++--
> mm/sparse.h | 7 +++++++
> 3 files changed, 39 insertions(+), 8 deletions(-)
>
> diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c
> index b770fe2428fd..b69a7af76858 100644
> --- a/mm/sparse-vmemmap.c
> +++ b/mm/sparse-vmemmap.c
> @@ -186,6 +186,11 @@ static __meminit struct page *vmemmap_get_tail(unsigned int order, struct zone *
>
> return tail;
> }
> +#else
> +static inline struct page *vmemmap_get_tail(unsigned int order, struct zone *zone)
> +{
> + return NULL;
> +}
> #endif
>
> static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, int node,
> @@ -193,12 +198,24 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
> unsigned long ptpfn, unsigned long flags)
> {
> pte_t *pte = pte_offset_kernel(pmd, addr);
> + unsigned long pfn = page_to_pfn((struct page *)addr);
> +
> if (pte_none(ptep_get(pte))) {
> pte_t entry;
> - void *p;
> +
> + if (pfn_vmemmap_optimizable(pfn) && ptpfn == (unsigned long)-1) {
> + unsigned int order = pfn_to_section_order(pfn);
> + struct zone *zone = pfn_to_zone(pfn, node);
> + struct page *page = vmemmap_get_tail(order, zone);
> +
> + if (!page)
> + return NULL;
> + ptpfn = page_to_pfn(page);
> + }
>
> if (ptpfn == (unsigned long)-1) {
> - p = vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
> + void *p = vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap);
> +
> if (!p)
> return NULL;
> ptpfn = PHYS_PFN(__pa(p));
> @@ -217,7 +234,8 @@ static pte_t * __meminit vmemmap_pte_populate(pmd_t *pmd, unsigned long addr, in
> }
> entry = pfn_pte(ptpfn, PAGE_KERNEL);
> set_pte_at(&init_mm, addr, pte, entry);
> - }
> + } else if (WARN_ON_ONCE(pfn_vmemmap_optimizable(pfn)))
> + return NULL;
> return pte;
> }
>
> @@ -406,6 +424,9 @@ int __meminit vmemmap_populate_hugepages(unsigned long start, unsigned long end,
> pmd_t *pmd;
>
> for (addr = start; addr < end; addr = next) {
> + unsigned long pfn = page_to_pfn((struct page *)addr);
> + const struct mem_section *ms = __pfn_to_section(pfn);
> +
> next = pmd_addr_end(addr, end);
>
> pgd = vmemmap_pgd_populate(addr, node);
> @@ -421,7 +442,7 @@ int __meminit vmemmap_populate_hugepages(unsigned long start, unsigned long end,
> return -ENOMEM;
>
> pmd = pmd_offset(pud, addr);
> - if (pmd_none(pmdp_get(pmd))) {
> + if (pmd_none(pmdp_get(pmd)) && !section_vmemmap_optimizable(ms)) {
> void *p;
>
> p = vmemmap_alloc_block_buf(PMD_SIZE, node, altmap);
> @@ -439,8 +460,11 @@ int __meminit vmemmap_populate_hugepages(unsigned long start, unsigned long end,
> */
> return -ENOMEM;
> }
> - } else if (vmemmap_check_pmd(pmd, node, addr, next))
> + } else if (vmemmap_check_pmd(pmd, node, addr, next)) {
> + if (WARN_ON_ONCE(section_vmemmap_optimizable(ms)))
> + return -ENOTSUPP;
> continue;
> + }
> if (vmemmap_populate_basepages(addr, next, node, altmap))
> return -ENOMEM;
> }
> @@ -648,7 +672,7 @@ void offline_mem_sections(unsigned long start_pfn, unsigned long end_pfn)
> }
> }
>
> -static int __meminit section_nr_vmemmap_pages(unsigned long pfn, unsigned long nr_pages,
> +int __meminit section_nr_vmemmap_pages(unsigned long pfn, unsigned long nr_pages,
The kernel test robot reported a compilation issue: when
CONFIG_MEMORY_HOTPLUG = n && CONFIG_SPARSEMEM_VMEMMAP=y,
section_nr_vmemmap_pages is undefined. This problem is easy to fix, and
I will move the entire function outside the CONFIG_MEMORY_HOTPLUG guard
in the next version.
Thanks.
> struct vmem_altmap *altmap, struct dev_pagemap *pgmap)
> {
> const struct mem_section *ms = __pfn_to_section(pfn);
> diff --git a/mm/sparse.c b/mm/sparse.c
> index ca9875f568d3..24555a32a5d9 100644
> --- a/mm/sparse.c
> +++ b/mm/sparse.c
> @@ -315,8 +315,8 @@ static void __init sparse_init_nid(int nid, unsigned long pnum_begin,
> nid, NULL, NULL);
> if (!map)
> panic("Failed to allocate memmap for section %lu\n", pnum);
> - memmap_boot_pages_add(DIV_ROUND_UP(PAGES_PER_SECTION * sizeof(struct page),
> - PAGE_SIZE));
> + memmap_boot_pages_add(section_nr_vmemmap_pages(pfn, PAGES_PER_SECTION,
> + NULL, NULL));
> sparse_init_early_section(nid, map, pnum, 0);
> }
> }
> diff --git a/mm/sparse.h b/mm/sparse.h
> index 6ad190ec48cf..f8f852f9f8a2 100644
> --- a/mm/sparse.h
> +++ b/mm/sparse.h
> @@ -105,8 +105,15 @@ static inline void sparse_init(void) {}
> */
> #ifdef CONFIG_SPARSEMEM_VMEMMAP
> void sparse_init_subsection_map(void);
> +int __meminit section_nr_vmemmap_pages(unsigned long pfn, unsigned long nr_pages,
> + struct vmem_altmap *altmap, struct dev_pagemap *pgmap);
> #else
> static inline void sparse_init_subsection_map(void) {}
> +static inline int section_nr_vmemmap_pages(unsigned long pfn, unsigned long nr_pages,
> + struct vmem_altmap *altmap, struct dev_pagemap *pgmap)
> +{
> + return DIV_ROUND_UP(nr_pages * sizeof(struct page), PAGE_SIZE);
> +}
> #endif /* CONFIG_SPARSEMEM_VMEMMAP */
>
> #endif /* __MM_SPARSE_H */
next prev parent reply other threads:[~2026-08-06 3:04 UTC|newest]
Thread overview: 23+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-04 3:55 [PATCH v3 00/17] mm: Introduce section-based vmemmap optimization for HugeTLB Muchun Song
2026-08-04 3:55 ` [PATCH v3 01/17] mm/sparse: relax struct mem_section size constraints Muchun Song
2026-08-04 3:55 ` [PATCH v3 02/17] mm/sparse-vmemmap: rename HVO order macros Muchun Song
2026-08-04 3:55 ` [PATCH v3 03/17] mm/mm_init: skip initializing shared vmemmap tail pages Muchun Song
2026-08-11 3:40 ` Muchun Song
2026-08-04 3:55 ` [PATCH v3 04/17] mm/sparse-vmemmap: initialize shared tail vmemmap pages on allocation Muchun Song
2026-08-04 3:55 ` [PATCH v3 05/17] mm/sparse-vmemmap: support section-based vmemmap accounting Muchun Song
2026-08-11 3:45 ` Muchun Song
2026-08-04 3:55 ` [PATCH v3 06/17] mm/mm_init: factor out pfn_to_zone() Muchun Song
2026-08-04 3:55 ` [PATCH v3 07/17] mm/sparse-vmemmap: move vmemmap_get_tail() before PTE population Muchun Song
2026-08-04 3:55 ` [PATCH v3 08/17] mm/sparse-vmemmap: support section-based vmemmap optimization Muchun Song
2026-08-06 3:04 ` Muchun Song [this message]
2026-08-11 3:54 ` Muchun Song
2026-08-04 3:55 ` [PATCH v3 09/17] mm/sparse: initialize memory sections earlier Muchun Song
2026-08-04 3:55 ` [PATCH v3 10/17] mm/hugetlb: switch HugeTLB to section-based vmemmap optimization Muchun Song
2026-08-11 3:58 ` Muchun Song
2026-08-04 3:55 ` [PATCH v3 11/17] mm/sparse-vmemmap: remove SPARSEMEM_VMEMMAP_PREINIT support Muchun Song
2026-08-04 3:55 ` [PATCH v3 12/17] mm/sparse: inline usemap allocation into sparse_init_nid() Muchun Song
2026-08-04 3:55 ` [PATCH v3 13/17] mm/sparse: remove section_map_size() Muchun Song
2026-08-04 3:55 ` [PATCH v3 14/17] mm/hugetlb: remove HUGE_BOOTMEM_HVO Muchun Song
2026-08-04 3:55 ` [PATCH v3 15/17] mm/hugetlb: remove HUGE_BOOTMEM_CMA Muchun Song
2026-08-04 3:55 ` [PATCH v3 16/17] mm/hugetlb: localize struct huge_bootmem_page Muchun Song
2026-08-04 3:55 ` [PATCH v3 17/17] mm/hugetlb: localize HUGE_BOOTMEM_ZONES_VALID Muchun Song
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=978060e0-a202-43b2-b84d-b1950b2d28c1@linux.dev \
--to=muchun.song@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=david.laight.linux@gmail.com \
--cc=david@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=osalvador@suse.de \
--cc=rppt@kernel.org \
--cc=songmuchun@bytedance.com \
--cc=vbabka@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.