From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-ej1-f71.google.com (mail-ej1-f71.google.com [209.85.218.71]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A734848F01E for ; Sat, 3 Oct 2026 21:22:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.71 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791062575; cv=none; b=dnsUPHNKspS+k3BjXpltDL6YsYlxKGW+tDaK2KcR/nCD13OSoP3aZG2+gQLyh/sJUYPwNM0Jezz6cBvWuRVA245nRzjeNxKNsz6uTncQfJGiMU1dzkKomqbG2/PrzNMotVsB3iVIKhisbwosTnXzEMJ2z35N5pJRSeOdRW1cs9g= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791062575; c=relaxed/simple; bh=C7ufT8KVq72MWzFJIoRk0UJZVOPIfdSL5MWcPH3Ufkc=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=egZDxnH15SsegehgLU3mC6Iz+oGPNMaYvhj624lfnPYrIxF7AAysozXHpNVp+Mg/E/7f5/ifAf+Z7sNzhRPqNLrp0BKOzcCBz5xon3PaDs8s+XRAlXu2kh4nYcXIMY1Q9nX7Nu0kqJIo4m4MoEhlR8R4IcCB3XlZ8n5B8ynaFjY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=SbGdnTZa; arc=none smtp.client-ip=209.85.218.71 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="SbGdnTZa" Received: by mail-ej1-f71.google.com with SMTP id a640c23a62f3a-c252a2ae8e1so90249266b.1 for ; Sat, 03 Oct 2026 14:22:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1791062570; x=1791667370; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=sxA7oyLrSoVNihu+TTEed7Cv84XwkDi44oXKTt9irD0=; b=SbGdnTZaXKkXlmXqvIAGbqXfLC0KcAHbZbn83UfTiXGKY7xv6NIm11BJ3CbV100+2o po3MxGwXiyWlh9epmtDPb0amnWHSctqEPU3zwNHYMZ2Rj4FEiqYSvm7b7/UhMY2dI4bM b6I39X2yOfed0KTQrabJ45HTMjLyFZM8hEf2BiT6N5UNoekqewNFXpn+kh5ilUOXH0VY pqIcTrdkWDp8iOiYE/BDkf9xKnSkwNQqEDaRsbuHTH80vas8mttocT4Nkh4lYwYInr4H kM45h1qJM4ThUZh6jIdJ0VDn8V34KhdAIsoYEPJaYwQucAhvUK6fber4PF7UGlYHNUIq l6ZA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1791062570; x=1791667370; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=sxA7oyLrSoVNihu+TTEed7Cv84XwkDi44oXKTt9irD0=; b=uqpswyb/KffjarCnLzqxgFbrbrzuH0lnZdADlQdoNYxtYEOsDWG6rjCsYhFKTpzv5x 4m+mty7lEVprDQEJzW4Uv0Kj8aKX1Qq2PinoG9E84AXTQLfIlTMncBTbgw9D2y3LsTdM glDZk+Uy04t1onVAeHSPGFDHaegV8GFYVF1E/P7OZdZ5Smy4jbe/uFALeB09xPmf5pgG 2iYBfDLThUHU8MjhcT20nJFuyJSZilDS+ZkQUZyzSqQnMCpCsRseXUo3zEvicocMOjom zYcqElCaPGJdhvClHYdmMG5/tc7nnISe83tYTKigyRCtV085BsTTuRtsGnssBpGv6+u/ LHkw== X-Forwarded-Encrypted: i=1; AKwUvBzsrWh0HV2bI9BitSmql+Li/FIrjgWC8OEVLA6j0xm5I8V2MJrAzp/B9oBsskBuymBD5FUThq8=@vger.kernel.org X-Gm-Message-State: AFuF++nVcF+N7A/f5RtMkvPkU37Iph3zGXgZxA6oMbVpzPkf0RHN94XL EuIS3kNyvC0duvBR58J7Og7kelMd6GhUE04cdFg4SUycNQT00xS8DyJqNp0SPY4j1BlW0nnsCxT +CRVRmA== X-Received: from ejcay23.prod.google.com ([2002:a17:907:9017:b0:c2d:d3a7:21fe]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a17:907:97d3:b0:c2e:3ed7:7db4 with SMTP id a640c23a62f3a-c2e6ed3f0fcmr244265366b.23.1791062569971; Sat, 03 Oct 2026 14:22:49 -0700 (PDT) Date: Sat, 3 Oct 2026 21:22:22 +0000 In-Reply-To: <20261003212241.3432303-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: netdev@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20261003212241.3432303-1-lrizzo@google.com> X-Mailer: git-send-email 2.56.0.rc1.315.gc6ed9934b7-goog Message-ID: <20261003212241.3432303-4-lrizzo@google.com> Subject: [RFC: DMA_PMD 03/22] mm: Add split_page_compound() From: Luigi Rizzo To: Luigi Rizzo , Joerg Roedel , Will Deacon , Robin Murphy , Christoph Hellwig , Marek Szyprowski , Andrew Morton , Vlastimil Babka , David Hildenbrand , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni Cc: Greg Kroah-Hartman , "Rafael J . Wysocki" , Danilo Krummrich , Jonathan Corbet , Jesper Dangaard Brouer , Ilias Apalodimas , Willem de Bruijn , Kuniyuki Iwashima , Joshua Washington , Harshitha Ramamurthy , Saeed Mahameed , Tariq Toukan , Tony Nguyen , Przemek Kitszel , Alexander Lobakin , Michael Chan , Pavan Chebbi , iommu@lists.linux.dev, netdev@vger.kernel.org, linux-mm@kvack.org, driver-core@lists.linux.dev, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, Luigi Rizzo Content-Type: text/plain; charset="UTF-8" Add split_page_compound(), which splits an order-@old_order page into independently refcounted order-@new_order compound pages (or order-0 pages when @new_order == 0). This is the high-order counterpart of split_page(), used by DMA_PMD pools to carve PMD pages into smaller compound blocks. All resulting pieces are returned frozen (refcount 0) so the caller can park them in a pool and publish each piece with page_ref_unfreeze(piece, 1) when handed out. Also add a KUnit test suite (CONFIG_SPLIT_PAGE_COMPOUND_KUNIT_TEST) in mm/split_page_compound_kunit.c. Signed-off-by: Luigi Rizzo --- include/linux/mm.h | 2 + mm/Kconfig.debug | 11 ++++ mm/Makefile | 1 + mm/page_alloc.c | 85 +++++++++++++++++++++++++ mm/split_page_compound_kunit.c | 113 +++++++++++++++++++++++++++++++++ 5 files changed, 212 insertions(+) create mode 100644 mm/split_page_compound_kunit.c diff --git a/include/linux/mm.h b/include/linux/mm.h index dd09c438fa23e..799a42005627f 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -1985,6 +1985,8 @@ static inline struct folio *virt_to_folio(const void *x) void __folio_put(struct folio *folio); void split_page(struct page *page, unsigned int order); +int split_page_compound(struct page *page, unsigned int old_order, + unsigned int new_order); void folio_copy(struct folio *dst, struct folio *src); int folio_mc_copy(struct folio *dst, struct folio *src); diff --git a/mm/Kconfig.debug b/mm/Kconfig.debug index 15dca19dd07da..e1a04671c3b0d 100644 --- a/mm/Kconfig.debug +++ b/mm/Kconfig.debug @@ -347,3 +347,14 @@ config MEM_ALLOC_PROFILING_DEBUG help Adds warnings with helpful error messages for memory allocation profiling. + +config SPLIT_PAGE_COMPOUND_KUNIT_TEST + bool "KUnit test for split_page_compound() (built-in)" if !KUNIT_ALL_TESTS + depends on KUNIT=y + default KUNIT_ALL_TESTS + help + Builds KUnit unit tests for split_page_compound() in mm/page_alloc.c, + verifying compound splitting, sub-block freezing, and speculative + reference handling. + + If unsure, say N. diff --git a/mm/Makefile b/mm/Makefile index e7245cb88c665..448c0f5f783f3 100644 --- a/mm/Makefile +++ b/mm/Makefile @@ -148,3 +148,4 @@ obj-$(CONFIG_EXECMEM) += execmem.o obj-$(CONFIG_TMPFS_QUOTA) += shmem_quota.o obj-$(CONFIG_LAZY_MMU_MODE_KUNIT_TEST) += tests/lazy_mmu_mode_kunit.o obj-$(CONFIG_MEM_ALLOC_PROFILING) += alloc_tag.o +obj-$(CONFIG_SPLIT_PAGE_COMPOUND_KUNIT_TEST) += split_page_compound_kunit.o diff --git a/mm/page_alloc.c b/mm/page_alloc.c index e8e9905cdea5b..a663ec81a588b 100644 --- a/mm/page_alloc.c +++ b/mm/page_alloc.c @@ -3133,6 +3133,91 @@ void split_page(struct page *page, unsigned int order) } EXPORT_SYMBOL_GPL(split_page); +/* + * split_page_compound() - split an order-@old_order page into compound pages + * of order @new_order + * @page: the page to split + * @old_order: the current order of @page + * @new_order: the order of the resulting pages + * + * Unlike split_page(), which produces order-0 pages, this produces + * 1 << (@old_order - @new_order) compound pages, each with its own head. + * @page must not be compound, and the caller must hold the only reference to + * it: the block is frozen while the new heads are built. + * + * All pieces (including @page at index 0) are returned frozen, with a zero + * refcount. Publish each piece with + * page_ref_unfreeze(piece, 1) before handing it out: its release store orders + * the compound layout built here against anyone who then observes the + * refcount. A relaxed set_page_count() would not, and on a weakly ordered + * machine a PFN walker could see a live refcount on a head whose layout is + * not visible yet. + * + * Memcg-charged pages are rejected: the accounting helpers assume a split + * produces order-0 pages and would mis-attribute the new compound heads. + * + * Return: 0 on success, -EINVAL if @page cannot be split or if @new_order is + * larger than @old_order, -EBUSY if anyone but the caller holds a reference + * to it. + */ +int split_page_compound(struct page *page, unsigned int old_order, + unsigned int new_order) +{ + unsigned int i, step = 1U << new_order, nr = 1U << old_order; + + if (WARN_ON_ONCE(PageCompound(page) || !page_count(page))) + return -EINVAL; + + if (WARN_ON_ONCE(new_order > old_order)) + return -EINVAL; + + if (WARN_ON_ONCE(memcg_kmem_online() && PageMemcgKmem(page))) + return -EINVAL; + + /* + * Shape the new compound pages while the block is frozen, which is + * the order __alloc_pages_noprof() itself uses: prep_new_page() + * builds the page with a zero refcount and the set_page_refcounted() + * that follows is what makes it visible. + * + * Freezing stops a speculative PFN walker from resolving a compound + * page that is only half built: get_page_unless_zero() fails for as + * long as the layout below is being written. It does not order + * against memory_failure(), which sets PG_hwpoison before taking any + * reference, so the non-atomic __SetPageHead() below can still lose a + * concurrent poison bit - exactly as it can for every other + * prep_compound_page() caller, the page allocator included. + * + * A failed freeze is not a bug, it means someone holds a speculative + * reference. Every PFN walker in the tree gates + * get_page_unless_zero() on PageLRU, which a freshly allocated block + * is not, so in practice only the hwpoison machinery gets here. Let + * the caller fall back rather than warn: a machine running + * panic_on_warn should not die of a condition the caller handles. + */ + if (!page_ref_freeze(page, 1)) + return -EBUSY; + + if (!new_order) { + /* + * The subpages of a non-compound block are already + * independent frozen pages, so only the bookkeeping applies. + */ + split_page_owner(page, old_order, 0); + pgalloc_tag_split(page_folio(page), old_order, 0); + split_page_memcg(page, old_order); + return 0; + } + + split_page_owner(page, old_order, new_order); + pgalloc_tag_split(page_folio(page), old_order, new_order); + + for (i = 0; i < nr; i += step) + prep_compound_page(page + i, new_order); + + return 0; +} + int __isolate_free_page(struct page *page, unsigned int order) { struct zone *zone = page_zone(page); diff --git a/mm/split_page_compound_kunit.c b/mm/split_page_compound_kunit.c new file mode 100644 index 0000000000000..bb0ac908bbc1e --- /dev/null +++ b/mm/split_page_compound_kunit.c @@ -0,0 +1,113 @@ +// SPDX-License-Identifier: GPL-2.0 OR BSD-3-Clause +/* + * KUnit tests for split_page_compound(). + */ +#include +#include +#include + +static void test_split_compound_pieces(struct kunit *test) +{ + const unsigned int old_order = 5, new_order = 2; + const unsigned int step = 1U << new_order; + const unsigned int nr = 1U << old_order; + struct page *page; + unsigned int i, j; + int ret; + + page = alloc_pages(GFP_KERNEL, old_order); + KUNIT_ASSERT_NOT_NULL(test, page); + KUNIT_EXPECT_FALSE(test, PageCompound(page)); + + ret = split_page_compound(page, old_order, new_order); + if (ret) + __free_pages(page, old_order); + KUNIT_ASSERT_EQ(test, ret, 0); + + for (i = 0; i < nr; i += step) { + struct page *sub = page + i; + + /* All heads 0..N-1 are returned frozen (refcount 0). */ + KUNIT_EXPECT_EQ(test, page_count(sub), 0); + KUNIT_EXPECT_TRUE(test, PageHead(sub)); + KUNIT_EXPECT_EQ(test, compound_order(sub), new_order); + + for (j = 1; j < step; j++) { + KUNIT_EXPECT_TRUE(test, PageTail(sub + j)); + KUNIT_EXPECT_PTR_EQ(test, compound_head(sub + j), sub); + } + + page_ref_unfreeze(sub, 1); + + /* Tail get_page()/put_page() must operate on this piece's head. */ + get_page(sub + 1); + KUNIT_EXPECT_EQ(test, page_count(sub), 2); + put_page(sub + 1); + KUNIT_EXPECT_EQ(test, page_count(sub), 1); + } + + /* Each compound piece frees independently at new_order. */ + for (i = 0; i < nr; i += step) + __free_pages(page + i, new_order); +} + +static void test_split_order0_pieces(struct kunit *test) +{ + const unsigned int old_order = 3, nr = 1U << old_order; + struct page *page; + unsigned int i; + int ret; + + page = alloc_pages(GFP_KERNEL, old_order); + KUNIT_ASSERT_NOT_NULL(test, page); + + ret = split_page_compound(page, old_order, 0); + if (ret) + __free_pages(page, old_order); + KUNIT_ASSERT_EQ(test, ret, 0); + + for (i = 0; i < nr; i++) { + struct page *sub = page + i; + + KUNIT_EXPECT_FALSE(test, PageCompound(sub)); + KUNIT_EXPECT_EQ(test, page_count(sub), 0); + page_ref_unfreeze(sub, 1); + __free_pages(sub, 0); + } +} + +static void test_split_ebusy_extra_ref(struct kunit *test) +{ + const unsigned int old_order = 4; + const unsigned int new_order = 2; + struct page *page; + int ret; + + page = alloc_pages(GFP_KERNEL, old_order); + KUNIT_ASSERT_NOT_NULL(test, page); + + /* Simulate a concurrent speculative reference. */ + get_page(page); + ret = split_page_compound(page, old_order, new_order); + KUNIT_EXPECT_EQ(test, ret, -EBUSY); + KUNIT_EXPECT_FALSE(test, PageCompound(page)); + + put_page(page); + __free_pages(page, old_order); +} + +static struct kunit_case split_page_compound_test_cases[] = { + KUNIT_CASE(test_split_compound_pieces), + KUNIT_CASE(test_split_order0_pieces), + KUNIT_CASE(test_split_ebusy_extra_ref), + {} +}; + +static struct kunit_suite split_page_compound_test_suite = { + .name = "split_page_compound", + .test_cases = split_page_compound_test_cases, +}; + +kunit_test_suite(split_page_compound_test_suite); +MODULE_DESCRIPTION("KUnit tests for split_page_compound()"); +MODULE_LICENSE("Dual BSD/GPL"); -- 2.56.0.rc1.315.gc6ed9934b7-goog