From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj2-f12.google.com (mail-pj2-f12.google.com [74.125.227.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3A26D39D6E5 for ; Fri, 11 Sep 2026 05:03:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.140 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789103020; cv=none; b=fD8YtIJj4gKK0jBOIHcj6fIvAmaO+qfMwYzx3T28uIkaaanELw/TYNIkMAgBkYnMJOU2TQjMaq5oOQ34sulslHNfZrQ9BR8KGNT4fBnS+0d9radvt0/pRgLm77aT4lJhuYAzQ9Xo32vv1K/bJwCYsOdN1PQnnzG61kfOidS+xdc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789103020; c=relaxed/simple; bh=fil1wyB3n2qakOtRPZGMWnUUGnhepk7poZ7+sHybDhs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=K6dg27QwQyIYiU80SjsidPJDLmQsr0ogFV7u3PyaE7uQudoe936tOD1cyNVtKLQGDKJDzf+cob52iRZe8ugZ0hnR8+euoIX/niMcJjeWFCZ5uW3INlH8gS8ZOZOGp1ZzYRoFoB2G6ZOLGQeIguaOrI0Rf/sXL8x7znTl6tAM/Lg= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=MABlb6Wm; arc=none smtp.client-ip=74.125.227.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="MABlb6Wm" Received: by mail-pj2-f12.google.com with SMTP id 98e67ed59e1d1-396ccafb751so422272a91.2 for ; Thu, 10 Sep 2026 22:03:37 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1789103016; x=1789707816; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=6bR1GMYxjFGrp68HZ8kexJBvgYVBJednZ0YGt4vUS+s=; b=MABlb6WmqgWHq8e1CKuMVOfOGwIqM7Y0uhz3rx/ZUqGvnaymUVG/EB6uwUkC1kTWlR 67DsjyIEYpnvfFnRutlcZ3W57kMVaHpxAB3fS8PLfpOPdJ8I67Zhjl7D/WocCcLSxdN1 zDt0H6MwRHVD5SodswUIWi41stFDSpCuYWrQpVfTfImFO4ySAyeg8MKEYlMBTJuWYJNa 5lKEx+1rxeSiargrHeBAn2+v4dymxXi/2VvuzLehtdewufz22ssofVP2qnGQjHW068N5 w3UGY66jPf6CxRCcLG22vovCi4WbxBgCUn5/p5mq2qK6Qlm6tY/1wnzLjT+51lOkmzdS P4lQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789103016; x=1789707816; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=6bR1GMYxjFGrp68HZ8kexJBvgYVBJednZ0YGt4vUS+s=; b=BswBpfb1jYkX02bpJA/1rP4A59bdie9acmgE77tFB5yyWpCcWuw4OEekzVEaQ6uinT elCIDLrTKcdDOYtjHAPybIzGSoeP/LanRaxcSjsF3Deg7FqPFFcTH6QzEn9tqSJlymA6 qLru0c5kXRcUcX/gpUr442ocGMiID91XczqcZN46vpb9gzbhVko27D0W83aOKAPEggqi DhdVvMNojgoEKS7ja0pmlHvrIN7tkWf+q4AZzGHV7hqnEix0y/OaqVJ9hyfWDt7CMv42 uu+L5LQjkIRH4knRWO2ISZ2YMgmVrnmxPbKOcrMG8+MMV6Tmqo0gH7+TD90XFzjIcKuQ py7Q== X-Forwarded-Encrypted: i=1; AKwUvBwHF1vNsnO1tPQVgLCzPWeTT8e+mxDsbC0DYMaj/cmirdUdu4Rx06vRKbP5Fh8pJN762rcESJUHxAE=@vger.kernel.org X-Gm-Message-State: AFuF++l1tfUxaHeFomVDsDR3y2FDSDP1SfZP6d/Ezf4ZyeH/ggLt+Neb +otVUsWnSoNHkDCHLm1SlDskoCcDeWo3TmqclSvC2k4SbixCQloiEY7iIIBg7SbmjOk= X-Gm-Gg: AYBFou36LP/vW4kGjpXe2tB39BStOHqPuWF3wFXiKtD3TQBEZYbXat32pQXr8qLIQUP VXjn/h6bJYWl9D1oRMGapNlxunZO8Nc7ifz8xxtiNBLeAVkcmMLRwW3bb7fve/T4Tv4jZy20rZV Q8EjqmA12kpS8HtfIP+psFTVLAoGb6sVnw3Nt2Zwop+Z1gN8WAqGNMKyVWmljMkZ6xO7sRDBrDf 7qojwn9Grp82Li/jU0volNKiSGdr7eNeItb/fAV2d4ZndaQwwfV25Eiw25MGHuEA8XIxXtKW7Oy JZkk83Ew6S3IRRXPLLupcF37j1BL6OaysfVSUKLy0dd5QlyX/2qaDpYfYmSOXJ0QPuVCHD6tC89 sLhAiYA0wLOJpTe9/NwFMlWv6V97RMVLU5UR6VrMHyMf/NzcMrtv6Pa6FwteXI4RckjxlMW9Oab N/ujKMKomxo7Skp7VYR8AfU3ka57eWqDnz0Ro4u3SGmD0Cg52kICoQqraOt9x85cRkwK4RuFwvM SZMXMredu+rcgFwG90IPzKIQFvFAnhdCKY22Q== X-Received: by 2002:a17:90b:48c7:b0:398:9be8:ea66 with SMTP id 98e67ed59e1d1-39d9c226c4dmr4442740a91.19.1789103016270; Thu, 10 Sep 2026 22:03:36 -0700 (PDT) Received: from G6L4RL2QG9.bytedance.net ([61.213.176.11]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-39d99531b16sm2651561a91.14.2026.09.10.22.03.31 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Thu, 10 Sep 2026 22:03:35 -0700 (PDT) From: Muchun Song To: Andrew Morton , David Hildenbrand , Oscar Salvador , Madhavan Srinivasan , Michael Ellerman , Jonathan Corbet Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linuxppc-dev@lists.ozlabs.org, linux-doc@vger.kernel.org, Muchun Song , Lorenzo Stoakes , Mike Rapoport , Qi Zheng , Nicholas Piggin , Christophe Leroy , Randy Dunlap , Muchun Song Subject: [PATCH v3 02/11] mm/sparse-vmemmap: factor out shared vmemmap tail page allocation Date: Fri, 11 Sep 2026 13:02:19 +0800 Message-ID: <20260911050228.58884-3-songmuchun@bytedance.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260911050228.58884-1-songmuchun@bytedance.com> References: <20260911050228.58884-1-songmuchun@bytedance.com> Precedence: bulk X-Mailing-List: linux-doc@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit HugeTLB and sparse-vmemmap each have their own helper to allocate the shared vmemmap tail page used by vmemmap optimization. Factor that logic into a common vmemmap_shared_tail_page() helper. It allocates the page through vmemmap_alloc_block(), and uses cmpxchg() to install the per-zone shared page. Expose zone->vmemmap_tails under CONFIG_SPARSEMEM_VMEMMAP to match the shared helper's build condition. This avoids a !CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION stub; when optimization is disabled, the array has no entries and the compiler folds away the unused paths, so no storage or runtime overhead is added. This removes duplicate allocation logic while still handling both the early boot and runtime paths through the same helper. Signed-off-by: Muchun Song Acked-by: Qi Zheng --- v2: - Collect Acked-by from Qi Zheng --- include/linux/mmzone.h | 2 +- mm/hugetlb_vmemmap.c | 28 +--------------- mm/sparse-vmemmap.c | 74 +++++++++++++++++------------------------- mm/sparse.h | 1 + 4 files changed, 33 insertions(+), 72 deletions(-) diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h index 97511f651ebc..1a18c1c40c8c 100644 --- a/include/linux/mmzone.h +++ b/include/linux/mmzone.h @@ -1156,7 +1156,7 @@ struct zone { /* Zone statistics */ atomic_long_t vm_stat[NR_VM_ZONE_STAT_ITEMS]; atomic_long_t vm_numa_event[NR_VM_NUMA_EVENT_ITEMS]; -#ifdef CONFIG_SPARSEMEM_VMEMMAP_OPTIMIZATION +#ifdef CONFIG_SPARSEMEM_VMEMMAP struct page *vmemmap_tails[VMEMMAP_OPTIMIZATION_NR_ORDERS]; #endif } ____cacheline_internodealigned_in_smp; diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c index f977d0a7e002..5ddf06b83c96 100644 --- a/mm/hugetlb_vmemmap.c +++ b/mm/hugetlb_vmemmap.c @@ -493,32 +493,6 @@ static bool vmemmap_should_optimize_folio(const struct hstate *h, struct folio * return true; } -static struct page *vmemmap_get_tail(unsigned int order, struct zone *zone) -{ - const unsigned int idx = order - VMEMMAP_OPTIMIZATION_MIN_ORDER; - struct page *tail, *p; - int node = zone_to_nid(zone); - - tail = READ_ONCE(zone->vmemmap_tails[idx]); - if (likely(tail)) - return tail; - - tail = alloc_pages_node(node, GFP_KERNEL | __GFP_ZERO, 0); - if (!tail) - return NULL; - - p = page_to_virt(tail); - for (int i = 0; i < PAGE_SIZE / sizeof(struct page); i++) - init_compound_tail(p + i, NULL, order, zone); - - if (cmpxchg(&zone->vmemmap_tails[idx], NULL, tail)) { - __free_page(tail); - tail = READ_ONCE(zone->vmemmap_tails[idx]); - } - - return tail; -} - static int __hugetlb_vmemmap_optimize_folio(const struct hstate *h, struct folio *folio, struct list_head *vmemmap_pages, @@ -535,7 +509,7 @@ static int __hugetlb_vmemmap_optimize_folio(const struct hstate *h, return ret; nid = folio_nid(folio); - vmemmap_tail = vmemmap_get_tail(h->order, folio_zone(folio)); + vmemmap_tail = vmemmap_shared_tail_page(h->order, folio_zone(folio)); if (!vmemmap_tail) return -ENOMEM; diff --git a/mm/sparse-vmemmap.c b/mm/sparse-vmemmap.c index f22d815d7af0..7388a5b5cce3 100644 --- a/mm/sparse-vmemmap.c +++ b/mm/sparse-vmemmap.c @@ -42,27 +42,13 @@ #include "mm_init.h" #include "sparse.h" -/* - * Allocate a block of memory to be used to back the virtual memory map - * or to back the page tables that are used to create the mapping. - * Uses the main allocators if they are available, else bootmem. - */ - -static void * __ref __earlyonly_bootmem_alloc(int node, - unsigned long size, - unsigned long align, - unsigned long goal) -{ - return memmap_alloc(size, align, goal, node, false); -} - -void * __meminit vmemmap_alloc_block(unsigned long size, int node) +void __ref *vmemmap_alloc_block(unsigned long size, int node) { /* If the main allocator is up use that, fallback to bootmem. */ if (slab_is_available()) { gfp_t gfp_mask = GFP_KERNEL|__GFP_RETRY_MAYFAIL|__GFP_NOWARN; int order = get_order(size); - static bool warned __meminitdata; + static bool warned; struct page *page; page = alloc_pages_node(node, gfp_mask, order); @@ -76,8 +62,7 @@ void * __meminit vmemmap_alloc_block(unsigned long size, int node) } return NULL; } else - return __earlyonly_bootmem_alloc(node, size, size, - __pa(MAX_DMA_ADDRESS)); + return memmap_alloc(size, size, __pa(MAX_DMA_ADDRESS), node, false); } static void * __meminit altmap_alloc_block_buf(unsigned long size, @@ -184,39 +169,40 @@ static void * __meminit vmemmap_alloc_block_zero(unsigned long size, int node) return p; } -#ifdef CONFIG_HUGETLB_PAGE_OPTIMIZE_VMEMMAP -static __meminit struct page *vmemmap_get_tail(unsigned int order, struct zone *zone) +struct page __ref *vmemmap_shared_tail_page(unsigned int order, struct zone *zone) { - struct page *p, *tail; - unsigned int idx; - int node = zone_to_nid(zone); + void *addr; + struct page *page; + const unsigned int idx = order - VMEMMAP_OPTIMIZATION_MIN_ORDER; - if (WARN_ON_ONCE(order < VMEMMAP_OPTIMIZATION_MIN_ORDER)) - return NULL; - if (WARN_ON_ONCE(order > MAX_FOLIO_ORDER)) + if (WARN_ON_ONCE(idx >= ARRAY_SIZE(zone->vmemmap_tails))) return NULL; - idx = order - VMEMMAP_OPTIMIZATION_MIN_ORDER; - tail = zone->vmemmap_tails[idx]; - if (tail) - return tail; - p = vmemmap_alloc_block_zero(PAGE_SIZE, node); - if (!p) + page = READ_ONCE(zone->vmemmap_tails[idx]); + if (likely(page)) + return page; + + addr = vmemmap_alloc_block(PAGE_SIZE, zone_to_nid(zone)); + if (!addr) return NULL; - for (int i = 0; i < PAGE_SIZE / sizeof(struct page); i++) - init_compound_tail(p + i, NULL, order, zone); - tail = virt_to_page(p); - zone->vmemmap_tails[idx] = tail; + for (int i = 0; i < PAGE_SIZE / sizeof(struct page); i++) { + page = (struct page *)addr + i; + mm_zero_struct_page(page); + init_compound_tail(page, NULL, order, zone); + } - return tail; -} -#else -static inline struct page *vmemmap_get_tail(unsigned int order, struct zone *zone) -{ - return NULL; + page = virt_to_page(addr); + if (cmpxchg(&zone->vmemmap_tails[idx], NULL, page) != NULL) { + if (slab_is_available()) + __free_page(page); + else + memblock_free(addr, PAGE_SIZE); + page = READ_ONCE(zone->vmemmap_tails[idx]); + } + + return page; } -#endif static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node, struct vmem_altmap *altmap) @@ -229,7 +215,7 @@ static __meminit void *vmemmap_alloc_pte(unsigned long pfn, int node, return vmemmap_alloc_block_buf(PAGE_SIZE, node, altmap); zone = pfn_to_zone(pfn, node); - page = vmemmap_get_tail(order, zone); + page = vmemmap_shared_tail_page(order, zone); if (!page) return NULL; diff --git a/mm/sparse.h b/mm/sparse.h index 3151d4db7575..a250bea4088e 100644 --- a/mm/sparse.h +++ b/mm/sparse.h @@ -142,6 +142,7 @@ static inline void sparse_sections_init(void) {} * mm/sparse-vmemmap.c */ #ifdef CONFIG_SPARSEMEM_VMEMMAP +struct page *vmemmap_shared_tail_page(unsigned int order, struct zone *zone); void sparse_init_subsection_map(void); int section_nr_vmemmap_pages(unsigned long pfn, unsigned long nr_pages, struct vmem_altmap *altmap, struct dev_pagemap *pgmap); -- 2.54.0