All of lore.kernel.org
 help / color / mirror / Atom feed
From: Qi Zheng <qi.zheng@linux.dev>
To: Muchun Song <muchun.song@linux.dev>
Cc: Muchun Song <songmuchun@bytedance.com>,
	Andrew Morton <akpm@linux-foundation.org>,
	Oscar Salvador <osalvador@suse.de>,
	David Hildenbrand <david@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	Lorenzo Stoakes <ljs@kernel.org>, Michal Hocko <mhocko@suse.com>,
	David Laight <david.laight.linux@gmail.com>,
	"Liam R . Howlett" <liam@infradead.org>,
	Suren Baghdasaryan <surenb@google.com>,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v5 10/17] mm/hugetlb: switch HugeTLB to section-based vmemmap optimization
Date: Wed, 26 Aug 2026 11:24:20 +0800	[thread overview]
Message-ID: <976369f1-7806-4c0d-b2b1-f91c3ac84b8b@linux.dev> (raw)
In-Reply-To: <C06EECC2-9025-49D8-91ED-EC58685398C4@linux.dev>



On 8/26/26 10:38 AM, Muchun Song wrote:
> 
> 
>> On Aug 25, 2026, at 21:12, Qi Zheng <qi.zheng@linux.dev> wrote:
>>
>>
>>
>> On 8/25/26 4:46 PM, Muchun Song wrote:
>>> HugeTLB bootmem vmemmap optimization still carries its own early setup
>>> path, including pre-populating optimized mappings before the generic
>>> sparse-vmemmap code runs.
>>> Now that the section-based vmemmap optimization can derive HugeTLB
>>> vmemmap deduplication from section metadata, HugeTLB only needs to mark
>>> the bootmem huge page range with the appropriate order. The generic
>>> sparse-vmemmap population path can then allocate and map the shared tail
>>> vmemmap pages without any HugeTLB-specific early population code.
>>> Do that by setting the section order when a bootmem huge page is
>>> allocated and dropping the dedicated pre-HVO helpers and related
>>> special-casing.
>>> This removes duplicate early setup logic and switches HugeTLB to the
>>> section-based vmemmap optimization path.
>>> Signed-off-by: Muchun Song <songmuchun@bytedance.com>
>>> Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
>>> ---
>>> v3:
>>> - Use the order-based helper for the bootmem vmemmap-optimized check
>>> v2:
>>> - Collect Acked-by from Mike Rapoport
>>> ---
>>>   include/linux/hugetlb.h |  1 -
>>>   include/linux/mm.h      |  3 --
>>>   mm/hugetlb.c            | 30 ++------------
>>>   mm/hugetlb_vmemmap.c    | 90 +++--------------------------------------
>>>   mm/hugetlb_vmemmap.h    | 14 +++----
>>>   mm/sparse-vmemmap.c     | 31 --------------
>>>   mm/sparse.h             | 27 +++++++++++++
>>>   7 files changed, 42 insertions(+), 154 deletions(-)
>>> diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h
>>> index 16c4c4caa126..fe28f98e1b22 100644
>>> --- a/include/linux/hugetlb.h
>>> +++ b/include/linux/hugetlb.h
>>> @@ -171,7 +171,6 @@ struct address_space *hugetlb_folio_mapping_lock_write(struct folio *folio);
>>>     extern int movable_gigantic_pages __read_mostly;
>>>   extern int sysctl_hugetlb_shm_group __read_mostly;
>>> -extern struct list_head huge_boot_pages[MAX_NUMNODES];
>>>     void hugetlb_bootmem_struct_page_init(void);
>>>   void hugetlb_bootmem_alloc(void);
>>> diff --git a/include/linux/mm.h b/include/linux/mm.h
>>> index dd09c438fa23..441bd39eab73 100644
>>> --- a/include/linux/mm.h
>>> +++ b/include/linux/mm.h
>>> @@ -5159,9 +5159,6 @@ int vmemmap_populate_hugepages(unsigned long start, unsigned long end,
>>>           int node, struct vmem_altmap *altmap);
>>>   int vmemmap_populate(unsigned long start, unsigned long end, int node,
>>>    struct vmem_altmap *altmap);
>>> -int vmemmap_populate_hvo(unsigned long start, unsigned long end,
>>> -  unsigned int order, struct zone *zone,
>>> -  unsigned long headsize);
>>>   void vmemmap_wrprotect_hvo(unsigned long start, unsigned long end, int node,
>>>      unsigned long headsize);
>>>   void vmemmap_populate_print_last(void);
>>> diff --git a/mm/hugetlb.c b/mm/hugetlb.c
>>> index 04e6c4244cd6..fbb0c83bea79 100644
>>> --- a/mm/hugetlb.c
>>> +++ b/mm/hugetlb.c
>>> @@ -52,6 +52,7 @@
>>>   #include "hugetlb_cma.h"
>>>   #include "hugetlb_internal.h"
>>>   #include "mm_init.h"
>>> +#include "sparse.h"
>>>   #include <linux/page-isolation.h>
>>>     int hugetlb_max_hstate __read_mostly;
>>> @@ -59,7 +60,7 @@ unsigned int default_hstate_idx;
>>>   struct hstate hstates[HUGE_MAX_HSTATE];
>>>     __initdata nodemask_t hugetlb_bootmem_nodes;
>>> -__initdata struct list_head huge_boot_pages[MAX_NUMNODES];
>>> +static struct list_head huge_boot_pages[MAX_NUMNODES] __initdata;
>>> /*
>>>   * Due to ordering constraints across the init code for various
>>> @@ -3139,6 +3140,7 @@ static bool __init alloc_bootmem_huge_page(struct hstate *h, int nid)
>>>    	} else {
>>>    		list_add_tail(&m->list, &huge_boot_pages[nid]);
>>>    		m->flags |= HUGE_BOOTMEM_ZONES_VALID;
>>> + 		hugetlb_vmemmap_optimize_bootmem_page(m);
>>>    		/*
>>>    		 * Only initialize the head struct page in memmap_init_reserved_pages,
>>>    		 * rest of the struct pages will be initialized by the HugeTLB
>>> @@ -3299,6 +3301,7 @@ static void __init gather_bootmem_prealloc_node(unsigned long nid)
>>>    	 * this folio.
>>>    	 */
>>>    	folio_set_hugetlb_vmemmap_optimized(folio);
>>> + 	section_set_order_range(folio_pfn(folio), folio_nr_pages(folio), 0);
>>
>> So section->order is only used during initialization. Is it ever
>> accessed later at runtime? If not, can we just skip zeroing it out?
> 
> You have keenly noticed a detail. This was actually done deliberately,
> because in patch 14, HUGE_BOOTMEM_HVO was removed and replaced with
> section->order. The ->order field may store a value that is less than
> VMEMMAP_OPTIMIZATION_MIN_ORDER, so clearing it to 0 here is to prevent
> potential issues with this memory region during the hotplug/hotremove
> process in the future (in my future series).
> 
> However, is there also a way to avoid clearing it to 0? There is. We
> could add an extra check in hugetlb_vmemmap_optimize_bootmem_page to
> only call section_set_order_range when the current hstate->order is
> greater than or equal to VMEMMAP_OPTIMIZATION_MIN_ORDER. I just feel
> that this would add a bit more code. And then during the boot phase,
> doing one extra zeroing-out doesn't introduce much overhead anyway.
> Therefore, I chose the simpler implementation approach.

Got it. Thanks for the detailed explanation!





  reply	other threads:[~2026-08-26  3:24 UTC|newest]

Thread overview: 54+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-25  8:45 [PATCH v5 00/17] mm: Introduce section-based vmemmap optimization for HugeTLB Muchun Song
2026-08-25  8:45 ` [PATCH v5 01/17] mm/sparse: relax struct mem_section size constraints Muchun Song
2026-09-09  9:33   ` David Hildenbrand (Arm)
2026-08-25  8:45 ` [PATCH v5 02/17] mm/sparse-vmemmap: rename HVO order macros Muchun Song
2026-08-25 10:05   ` Qi Zheng
2026-08-25  8:45 ` [PATCH v5 03/17] mm/mm_init: skip initializing shared vmemmap tail pages Muchun Song
2026-08-25  9:30   ` Mike Rapoport
2026-08-25 10:15   ` Qi Zheng
2026-09-09  9:36   ` David Hildenbrand (Arm)
2026-09-09 11:30     ` Muchun Song
2026-08-25  8:45 ` [PATCH v5 04/17] mm/sparse-vmemmap: initialize shared tail vmemmap pages on allocation Muchun Song
2026-08-25  9:30   ` Mike Rapoport
2026-08-25 10:40   ` Qi Zheng
2026-09-09  9:39   ` David Hildenbrand (Arm)
2026-09-09 12:49     ` Muchun Song
2026-08-25  8:45 ` [PATCH v5 05/17] mm/sparse-vmemmap: support section-based vmemmap accounting Muchun Song
2026-08-25 10:47   ` Qi Zheng
2026-08-25  8:45 ` [PATCH v5 06/17] mm/mm_init: factor out pfn_to_zone() Muchun Song
2026-08-25 10:51   ` Qi Zheng
2026-08-25  8:45 ` [PATCH v5 07/17] mm/sparse-vmemmap: move helpers ahead of future callers Muchun Song
2026-08-25 10:54   ` Qi Zheng
2026-08-25  8:45 ` [PATCH v5 08/17] mm/sparse-vmemmap: support section-based vmemmap optimization Muchun Song
2026-08-25  9:30   ` Mike Rapoport
2026-08-25 12:38   ` Qi Zheng
2026-08-25  8:46 ` [PATCH v5 09/17] mm/sparse: initialize memory sections earlier Muchun Song
2026-08-25 12:49   ` Qi Zheng
2026-09-09  9:56   ` David Hildenbrand (Arm)
2026-09-09 12:29     ` Muchun Song
2026-09-09 12:48       ` David Hildenbrand (Arm)
2026-09-09 12:58         ` Muchun Song
2026-08-25  8:46 ` [PATCH v5 10/17] mm/hugetlb: switch HugeTLB to section-based vmemmap optimization Muchun Song
2026-08-25 13:12   ` Qi Zheng
2026-08-26  2:38     ` Muchun Song
2026-08-26  3:24       ` Qi Zheng [this message]
2026-09-09  9:58   ` David Hildenbrand (Arm)
2026-09-09 13:08     ` Muchun Song
2026-09-09 13:39       ` David Hildenbrand (Arm)
2026-08-25  8:46 ` [PATCH v5 11/17] mm/sparse-vmemmap: remove SPARSEMEM_VMEMMAP_PREINIT support Muchun Song
2026-08-25 13:16   ` Qi Zheng
2026-09-09  9:59   ` David Hildenbrand (Arm)
2026-08-25  8:46 ` [PATCH v5 12/17] mm/sparse: inline usemap allocation into sparse_init_nid() Muchun Song
2026-09-09 10:01   ` David Hildenbrand (Arm)
2026-08-25  8:46 ` [PATCH v5 13/17] mm/sparse: remove section_map_size() Muchun Song
2026-09-09 10:02   ` David Hildenbrand (Arm)
2026-08-25  8:46 ` [PATCH v5 14/17] mm/hugetlb: remove HUGE_BOOTMEM_HVO Muchun Song
2026-08-25 13:28   ` Qi Zheng
2026-08-25  8:46 ` [PATCH v5 15/17] mm/hugetlb: remove HUGE_BOOTMEM_CMA Muchun Song
2026-08-25 13:42   ` Qi Zheng
2026-08-25  8:46 ` [PATCH v5 16/17] mm/hugetlb: localize struct huge_bootmem_page Muchun Song
2026-08-25 13:38   ` Qi Zheng
2026-08-25  8:46 ` [PATCH v5 17/17] mm/hugetlb: localize HUGE_BOOTMEM_ZONES_VALID Muchun Song
2026-08-25 13:40   ` Qi Zheng
2026-08-28  4:20 ` [PATCH v5 00/17] mm: Introduce section-based vmemmap optimization for HugeTLB Andrew Morton
2026-09-09  9:31 ` David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=976369f1-7806-4c0d-b2b1-f91c3ac84b8b@linux.dev \
    --to=qi.zheng@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=david.laight.linux@gmail.com \
    --cc=david@kernel.org \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=osalvador@suse.de \
    --cc=rppt@kernel.org \
    --cc=songmuchun@bytedance.com \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.