Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Qi Zheng <qi.zheng@linux.dev>
To: Muchun Song <muchun.song@linux.dev>
Cc: Muchun Song <songmuchun@bytedance.com>,
	Andrew Morton <akpm@linux-foundation.org>,
	Oscar Salvador <osalvador@suse.de>,
	David Hildenbrand <david@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	Lorenzo Stoakes <ljs@kernel.org>, Michal Hocko <mhocko@suse.com>,
	David Laight <david.laight.linux@gmail.com>,
	"Liam R . Howlett" <liam@infradead.org>,
	Suren Baghdasaryan <surenb@google.com>,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v5 10/17] mm/hugetlb: switch HugeTLB to section-based vmemmap optimization
Date: Wed, 26 Aug 2026 11:24:20 +0800	[thread overview]
Message-ID: <976369f1-7806-4c0d-b2b1-f91c3ac84b8b@linux.dev> (raw)
In-Reply-To: <C06EECC2-9025-49D8-91ED-EC58685398C4@linux.dev>



On 8/26/26 10:38 AM, Muchun Song wrote:
> 
> 
>> On Aug 25, 2026, at 21:12, Qi Zheng <qi.zheng@linux.dev> wrote:
>>
>>
>>
>> On 8/25/26 4:46 PM, Muchun Song wrote:
>>> HugeTLB bootmem vmemmap optimization still carries its own early setup
>>> path, including pre-populating optimized mappings before the generic
>>> sparse-vmemmap code runs.
>>> Now that the section-based vmemmap optimization can derive HugeTLB
>>> vmemmap deduplication from section metadata, HugeTLB only needs to mark
>>> the bootmem huge page range with the appropriate order. The generic
>>> sparse-vmemmap population path can then allocate and map the shared tail
>>> vmemmap pages without any HugeTLB-specific early population code.
>>> Do that by setting the section order when a bootmem huge page is
>>> allocated and dropping the dedicated pre-HVO helpers and related
>>> special-casing.
>>> This removes duplicate early setup logic and switches HugeTLB to the
>>> section-based vmemmap optimization path.
>>> Signed-off-by: Muchun Song <songmuchun@bytedance.com>
>>> Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
>>> ---
>>> v3:
>>> - Use the order-based helper for the bootmem vmemmap-optimized check
>>> v2:
>>> - Collect Acked-by from Mike Rapoport
>>> ---
>>>   include/linux/hugetlb.h |  1 -
>>>   include/linux/mm.h      |  3 --
>>>   mm/hugetlb.c            | 30 ++------------
>>>   mm/hugetlb_vmemmap.c    | 90 +++--------------------------------------
>>>   mm/hugetlb_vmemmap.h    | 14 +++----
>>>   mm/sparse-vmemmap.c     | 31 --------------
>>>   mm/sparse.h             | 27 +++++++++++++
>>>   7 files changed, 42 insertions(+), 154 deletions(-)
>>> diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h
>>> index 16c4c4caa126..fe28f98e1b22 100644
>>> --- a/include/linux/hugetlb.h
>>> +++ b/include/linux/hugetlb.h
>>> @@ -171,7 +171,6 @@ struct address_space *hugetlb_folio_mapping_lock_write(struct folio *folio);
>>>     extern int movable_gigantic_pages __read_mostly;
>>>   extern int sysctl_hugetlb_shm_group __read_mostly;
>>> -extern struct list_head huge_boot_pages[MAX_NUMNODES];
>>>     void hugetlb_bootmem_struct_page_init(void);
>>>   void hugetlb_bootmem_alloc(void);
>>> diff --git a/include/linux/mm.h b/include/linux/mm.h
>>> index dd09c438fa23..441bd39eab73 100644
>>> --- a/include/linux/mm.h
>>> +++ b/include/linux/mm.h
>>> @@ -5159,9 +5159,6 @@ int vmemmap_populate_hugepages(unsigned long start, unsigned long end,
>>>           int node, struct vmem_altmap *altmap);
>>>   int vmemmap_populate(unsigned long start, unsigned long end, int node,
>>>    struct vmem_altmap *altmap);
>>> -int vmemmap_populate_hvo(unsigned long start, unsigned long end,
>>> -  unsigned int order, struct zone *zone,
>>> -  unsigned long headsize);
>>>   void vmemmap_wrprotect_hvo(unsigned long start, unsigned long end, int node,
>>>      unsigned long headsize);
>>>   void vmemmap_populate_print_last(void);
>>> diff --git a/mm/hugetlb.c b/mm/hugetlb.c
>>> index 04e6c4244cd6..fbb0c83bea79 100644
>>> --- a/mm/hugetlb.c
>>> +++ b/mm/hugetlb.c
>>> @@ -52,6 +52,7 @@
>>>   #include "hugetlb_cma.h"
>>>   #include "hugetlb_internal.h"
>>>   #include "mm_init.h"
>>> +#include "sparse.h"
>>>   #include <linux/page-isolation.h>
>>>     int hugetlb_max_hstate __read_mostly;
>>> @@ -59,7 +60,7 @@ unsigned int default_hstate_idx;
>>>   struct hstate hstates[HUGE_MAX_HSTATE];
>>>     __initdata nodemask_t hugetlb_bootmem_nodes;
>>> -__initdata struct list_head huge_boot_pages[MAX_NUMNODES];
>>> +static struct list_head huge_boot_pages[MAX_NUMNODES] __initdata;
>>> /*
>>>   * Due to ordering constraints across the init code for various
>>> @@ -3139,6 +3140,7 @@ static bool __init alloc_bootmem_huge_page(struct hstate *h, int nid)
>>>    	} else {
>>>    		list_add_tail(&m->list, &huge_boot_pages[nid]);
>>>    		m->flags |= HUGE_BOOTMEM_ZONES_VALID;
>>> + 		hugetlb_vmemmap_optimize_bootmem_page(m);
>>>    		/*
>>>    		 * Only initialize the head struct page in memmap_init_reserved_pages,
>>>    		 * rest of the struct pages will be initialized by the HugeTLB
>>> @@ -3299,6 +3301,7 @@ static void __init gather_bootmem_prealloc_node(unsigned long nid)
>>>    	 * this folio.
>>>    	 */
>>>    	folio_set_hugetlb_vmemmap_optimized(folio);
>>> + 	section_set_order_range(folio_pfn(folio), folio_nr_pages(folio), 0);
>>
>> So section->order is only used during initialization. Is it ever
>> accessed later at runtime? If not, can we just skip zeroing it out?
> 
> You have keenly noticed a detail. This was actually done deliberately,
> because in patch 14, HUGE_BOOTMEM_HVO was removed and replaced with
> section->order. The ->order field may store a value that is less than
> VMEMMAP_OPTIMIZATION_MIN_ORDER, so clearing it to 0 here is to prevent
> potential issues with this memory region during the hotplug/hotremove
> process in the future (in my future series).
> 
> However, is there also a way to avoid clearing it to 0? There is. We
> could add an extra check in hugetlb_vmemmap_optimize_bootmem_page to
> only call section_set_order_range when the current hstate->order is
> greater than or equal to VMEMMAP_OPTIMIZATION_MIN_ORDER. I just feel
> that this would add a bit more code. And then during the boot phase,
> doing one extra zeroing-out doesn't introduce much overhead anyway.
> Therefore, I chose the simpler implementation approach.

Got it. Thanks for the detailed explanation!





  reply	other threads:[~2026-08-26  3:24 UTC|newest]

Thread overview: 37+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-25  8:45 [PATCH v5 00/17] mm: Introduce section-based vmemmap optimization for HugeTLB Muchun Song
2026-08-25  8:45 ` [PATCH v5 01/17] mm/sparse: relax struct mem_section size constraints Muchun Song
2026-08-25  8:45 ` [PATCH v5 02/17] mm/sparse-vmemmap: rename HVO order macros Muchun Song
2026-08-25 10:05   ` Qi Zheng
2026-08-25  8:45 ` [PATCH v5 03/17] mm/mm_init: skip initializing shared vmemmap tail pages Muchun Song
2026-08-25  9:30   ` Mike Rapoport
2026-08-25 10:15   ` Qi Zheng
2026-08-25  8:45 ` [PATCH v5 04/17] mm/sparse-vmemmap: initialize shared tail vmemmap pages on allocation Muchun Song
2026-08-25  9:30   ` Mike Rapoport
2026-08-25 10:40   ` Qi Zheng
2026-08-25  8:45 ` [PATCH v5 05/17] mm/sparse-vmemmap: support section-based vmemmap accounting Muchun Song
2026-08-25 10:47   ` Qi Zheng
2026-08-25  8:45 ` [PATCH v5 06/17] mm/mm_init: factor out pfn_to_zone() Muchun Song
2026-08-25 10:51   ` Qi Zheng
2026-08-25  8:45 ` [PATCH v5 07/17] mm/sparse-vmemmap: move helpers ahead of future callers Muchun Song
2026-08-25 10:54   ` Qi Zheng
2026-08-25  8:45 ` [PATCH v5 08/17] mm/sparse-vmemmap: support section-based vmemmap optimization Muchun Song
2026-08-25  9:30   ` Mike Rapoport
2026-08-25 12:38   ` Qi Zheng
2026-08-25  8:46 ` [PATCH v5 09/17] mm/sparse: initialize memory sections earlier Muchun Song
2026-08-25 12:49   ` Qi Zheng
2026-08-25  8:46 ` [PATCH v5 10/17] mm/hugetlb: switch HugeTLB to section-based vmemmap optimization Muchun Song
2026-08-25 13:12   ` Qi Zheng
2026-08-26  2:38     ` Muchun Song
2026-08-26  3:24       ` Qi Zheng [this message]
2026-08-25  8:46 ` [PATCH v5 11/17] mm/sparse-vmemmap: remove SPARSEMEM_VMEMMAP_PREINIT support Muchun Song
2026-08-25 13:16   ` Qi Zheng
2026-08-25  8:46 ` [PATCH v5 12/17] mm/sparse: inline usemap allocation into sparse_init_nid() Muchun Song
2026-08-25  8:46 ` [PATCH v5 13/17] mm/sparse: remove section_map_size() Muchun Song
2026-08-25  8:46 ` [PATCH v5 14/17] mm/hugetlb: remove HUGE_BOOTMEM_HVO Muchun Song
2026-08-25 13:28   ` Qi Zheng
2026-08-25  8:46 ` [PATCH v5 15/17] mm/hugetlb: remove HUGE_BOOTMEM_CMA Muchun Song
2026-08-25 13:42   ` Qi Zheng
2026-08-25  8:46 ` [PATCH v5 16/17] mm/hugetlb: localize struct huge_bootmem_page Muchun Song
2026-08-25 13:38   ` Qi Zheng
2026-08-25  8:46 ` [PATCH v5 17/17] mm/hugetlb: localize HUGE_BOOTMEM_ZONES_VALID Muchun Song
2026-08-25 13:40   ` Qi Zheng

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=976369f1-7806-4c0d-b2b1-f91c3ac84b8b@linux.dev \
    --to=qi.zheng@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=david.laight.linux@gmail.com \
    --cc=david@kernel.org \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=osalvador@suse.de \
    --cc=rppt@kernel.org \
    --cc=songmuchun@bytedance.com \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox