From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mta1.migadu.com (out-112.mta1.migadu.com [95.215.58.112]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3CBC81DD0EF for ; Wed, 26 Aug 2026 02:38:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.112 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787711912; cv=none; b=ajH16A+3vQTdrikIGexF/0OcqfgFUN6sOeCw5C2MCXuUAtkO7dT80BgaQljuosPO0Bq1NByas8AkdpkkbaQCyd58C5Kqi0XOsvze1g5ZrH9D6cgtnxPnj89FQiG5/op1clfUK1KsePtEvsNlpF5+ITSD9yCke+J8AlX1kQeB99E= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787711912; c=relaxed/simple; bh=iaJbOqWwU7swmIOu9Nbpfo0jDj4KZx3LS8pZAlWsWV8=; h=Content-Type:Mime-Version:Subject:From:In-Reply-To:Date:Cc: Message-Id:References:To; b=FeOzl0BIRbYgR2i9dVMjJZdoBqVuEOpfpK/OaNgH3YMtz+qHRw+1s5oITc73xrLI1++jazVImCnbk3/vehJVDK96OKetSEnukcnqmmFHeP5kfEEZQNEUvWUkzmdFfww36Jf7qmSyo4oR+qwAEXoR1KyPBdE15DP/fmKCypo3rRE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=BKNkuZZ2; arc=none smtp.client-ip=95.215.58.112 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="BKNkuZZ2" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=iaJbOqWwU7swmIOu9Nbpfo0jDj4KZx3LS8pZAlWsWV8=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787711907; v=1; x=1788316707; b=BKNkuZZ2uzuJpmj4ptcyqwXtJcV0xmF9zKMoW8cwAtdbydxMNvciQDNIwfLCvAv1vwVBX/lb nUDdzB9XpsinJH+ek/jgLxZOBFfIf8HUdGHMIaU1oT1kA40oylWr4mom46JkJVsMPDF3DzjpdmF 75N+GDljm3vHFKFzmNNSO1rg= X-Envelope-To: linux-kernel@vger.kernel.org Received: from smtpclient.apple (114.251.196.90) by mta10.migadu.com with ESMTPS id 56a300c3813559b4; Wed, 26 Aug 2026 02:38:27 +0000 X-Mizu-Trace-ID: 56a300c3813559b4 X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset=us-ascii Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 (Mac OS X Mail 16.0 \(3864.600.51.1.1\)) Subject: Re: [PATCH v5 10/17] mm/hugetlb: switch HugeTLB to section-based vmemmap optimization From: Muchun Song In-Reply-To: <83d8d1f2-6ca2-4c72-a494-23ce07bb4181@linux.dev> Date: Wed, 26 Aug 2026 10:38:08 +0800 Cc: Muchun Song , Andrew Morton , Oscar Salvador , David Hildenbrand , Mike Rapoport , Vlastimil Babka , Lorenzo Stoakes , Michal Hocko , David Laight , "Liam R . Howlett" , Suren Baghdasaryan , linux-mm@kvack.org, linux-kernel@vger.kernel.org Content-Transfer-Encoding: quoted-printable Message-Id: References: <20260825084608.47437-1-songmuchun@bytedance.com> <20260825084608.47437-11-songmuchun@bytedance.com> <83d8d1f2-6ca2-4c72-a494-23ce07bb4181@linux.dev> To: Qi Zheng X-Mailer: Apple Mail (2.3864.600.51.1.1) > On Aug 25, 2026, at 21:12, Qi Zheng wrote: >=20 >=20 >=20 > On 8/25/26 4:46 PM, Muchun Song wrote: >> HugeTLB bootmem vmemmap optimization still carries its own early = setup >> path, including pre-populating optimized mappings before the generic >> sparse-vmemmap code runs. >> Now that the section-based vmemmap optimization can derive HugeTLB >> vmemmap deduplication from section metadata, HugeTLB only needs to = mark >> the bootmem huge page range with the appropriate order. The generic >> sparse-vmemmap population path can then allocate and map the shared = tail >> vmemmap pages without any HugeTLB-specific early population code. >> Do that by setting the section order when a bootmem huge page is >> allocated and dropping the dedicated pre-HVO helpers and related >> special-casing. >> This removes duplicate early setup logic and switches HugeTLB to the >> section-based vmemmap optimization path. >> Signed-off-by: Muchun Song >> Acked-by: Mike Rapoport (Microsoft) >> --- >> v3: >> - Use the order-based helper for the bootmem vmemmap-optimized check >> v2: >> - Collect Acked-by from Mike Rapoport >> --- >> include/linux/hugetlb.h | 1 - >> include/linux/mm.h | 3 -- >> mm/hugetlb.c | 30 ++------------ >> mm/hugetlb_vmemmap.c | 90 = +++-------------------------------------- >> mm/hugetlb_vmemmap.h | 14 +++---- >> mm/sparse-vmemmap.c | 31 -------------- >> mm/sparse.h | 27 +++++++++++++ >> 7 files changed, 42 insertions(+), 154 deletions(-) >> diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h >> index 16c4c4caa126..fe28f98e1b22 100644 >> --- a/include/linux/hugetlb.h >> +++ b/include/linux/hugetlb.h >> @@ -171,7 +171,6 @@ struct address_space = *hugetlb_folio_mapping_lock_write(struct folio *folio); >> extern int movable_gigantic_pages __read_mostly; >> extern int sysctl_hugetlb_shm_group __read_mostly; >> -extern struct list_head huge_boot_pages[MAX_NUMNODES]; >> void hugetlb_bootmem_struct_page_init(void); >> void hugetlb_bootmem_alloc(void); >> diff --git a/include/linux/mm.h b/include/linux/mm.h >> index dd09c438fa23..441bd39eab73 100644 >> --- a/include/linux/mm.h >> +++ b/include/linux/mm.h >> @@ -5159,9 +5159,6 @@ int vmemmap_populate_hugepages(unsigned long = start, unsigned long end, >> int node, struct vmem_altmap *altmap); >> int vmemmap_populate(unsigned long start, unsigned long end, int = node, >> struct vmem_altmap *altmap); >> -int vmemmap_populate_hvo(unsigned long start, unsigned long end, >> - unsigned int order, struct zone *zone, >> - unsigned long headsize); >> void vmemmap_wrprotect_hvo(unsigned long start, unsigned long end, = int node, >> unsigned long headsize); >> void vmemmap_populate_print_last(void); >> diff --git a/mm/hugetlb.c b/mm/hugetlb.c >> index 04e6c4244cd6..fbb0c83bea79 100644 >> --- a/mm/hugetlb.c >> +++ b/mm/hugetlb.c >> @@ -52,6 +52,7 @@ >> #include "hugetlb_cma.h" >> #include "hugetlb_internal.h" >> #include "mm_init.h" >> +#include "sparse.h" >> #include >> int hugetlb_max_hstate __read_mostly; >> @@ -59,7 +60,7 @@ unsigned int default_hstate_idx; >> struct hstate hstates[HUGE_MAX_HSTATE]; >> __initdata nodemask_t hugetlb_bootmem_nodes; >> -__initdata struct list_head huge_boot_pages[MAX_NUMNODES]; >> +static struct list_head huge_boot_pages[MAX_NUMNODES] __initdata; >> /* >> * Due to ordering constraints across the init code for various >> @@ -3139,6 +3140,7 @@ static bool __init = alloc_bootmem_huge_page(struct hstate *h, int nid) >> } else { >> list_add_tail(&m->list, &huge_boot_pages[nid]); >> m->flags |=3D HUGE_BOOTMEM_ZONES_VALID; >> + hugetlb_vmemmap_optimize_bootmem_page(m); >> /* >> * Only initialize the head struct page in = memmap_init_reserved_pages, >> * rest of the struct pages will be initialized by the = HugeTLB >> @@ -3299,6 +3301,7 @@ static void __init = gather_bootmem_prealloc_node(unsigned long nid) >> * this folio. >> */ >> folio_set_hugetlb_vmemmap_optimized(folio); >> + section_set_order_range(folio_pfn(folio), folio_nr_pages(folio), = 0); >=20 > So section->order is only used during initialization. Is it ever > accessed later at runtime? If not, can we just skip zeroing it out? You have keenly noticed a detail. This was actually done deliberately, because in patch 14, HUGE_BOOTMEM_HVO was removed and replaced with=20 section->order. The ->order field may store a value that is less than VMEMMAP_OPTIMIZATION_MIN_ORDER, so clearing it to 0 here is to prevent potential issues with this memory region during the hotplug/hotremove process in the future (in my future series). However, is there also a way to avoid clearing it to 0? There is. We could add an extra check in hugetlb_vmemmap_optimize_bootmem_page to only call section_set_order_range when the current hstate->order is greater than or equal to VMEMMAP_OPTIMIZATION_MIN_ORDER. I just feel that this would add a bit more code. And then during the boot phase, doing one extra zeroing-out doesn't introduce much overhead anyway. Therefore, I chose the simpler implementation approach. Muchun, Thanks.=