From: "Li Zhe" <lizhe.67@bytedance.com>
To: <akpm@linux-foundation.org>, <apopple@nvidia.com>,
<arnd@arndb.de>, <balbirs@nvidia.com>, <bp@alien8.de>,
<dave.hansen@linux.intel.com>, <david@kernel.org>,
<kees@kernel.org>, <mingo@redhat.com>, <muchun.song@linux.dev>,
<rppt@kernel.org>, <tglx@kernel.org>
Cc: <linux-arch@vger.kernel.org>, <linux-hardening@vger.kernel.org>,
<linux-kernel@vger.kernel.org>, <linux-mm@kvack.org>,
<x86@kernel.org>, <lizhe.67@bytedance.com>
Subject: [PATCH v9 7/8] mm: use memcpy_nontemporal() in zone-device template copies
Date: Mon, 3 Aug 2026 15:09:28 +0800 [thread overview]
Message-ID: <20260803070929.86075-8-lizhe.67@bytedance.com> (raw)
In-Reply-To: <20260803070929.86075-1-lizhe.67@bytedance.com>
The template fast path currently uses memcpy() for the actual struct
page copy. Switch zone_device_page_init_from_template() to
memcpy_nontemporal().
ZONE_DEVICE memmap initialization is largely write-once: each struct
page is populated once, and most destination cachelines are not expected
to be reused immediately afterwards. On x86, a regular cached memcpy()
can therefore incur write-allocate traffic by pulling destination
cachelines into the cache before writeback, and can populate the cache
with data that has little near-term reuse. Using memcpy_nontemporal()
lets this path request nontemporal stores for that copy pattern, which
can reduce cache pollution and avoid part of the associated
write-allocate overhead, while architectures without a specialized
backend still fall back to memcpy().
Do not add a KASAN/KMSAN-specific fallback around this call site. As
Muchun pointed out, special KASAN handling for memcpy_flushcache() or
memcpy_nontemporal(), if needed, belongs in the low-level helper rather
than in this ZONE_DEVICE caller.
No separate drain is added here. memcpy_nontemporal() is used only as
the copy primitive while memmap_init_zone_device() is still initializing
the struct page array. The ordinary stores that follow in this path,
such as compound-page setup, are part of the same CPU's initialization
sequence; they are not used as a publication store that tells another CPU
or device to consume data written by the non-temporal copy.
Therefore this call site does not need a helper-level drain for
correctness. Callers that use memcpy_nontemporal() as part of a
producer-consumer or device-visible handoff must add the required
ordering themselves.
Tested in a VM with a 100 GB fsdax namespace device configured with
map=dev and a 100 GB devdax namespace (align=2097152) on Intel Ice Lake
server.
Test procedure:
Rebind the nd_pmem and dax_pmem driver 30 times and collect the memmap
initialization time from the pr_debug() output of
memmap_init_zone_device().
Base(v7.2-rc1):
Average of rebinds for nd_pmem driver: 244.28 ms
Average of rebinds for dax_pmem driver: 273.31 ms
With this patch and its prerequisites applied:
Average of rebinds for nd_pmem driver: 150.83 ms
Average of rebinds for dax_pmem driver: 153.55 ms
This reduces the average memmap initialization time measured during rebind
by about 38.3% for nd_pmem and 43.8% for dax_pmem.
Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
---
mm/mm_init.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/mm/mm_init.c b/mm/mm_init.c
index 9691fa2a060d..bb2007806a28 100644
--- a/mm/mm_init.c
+++ b/mm/mm_init.c
@@ -1101,7 +1101,7 @@ static void zone_device_page_init_from_template(struct page *page,
* to the destination page.
*/
zone_device_page_update_template(template, pfn);
- memcpy(page, template, sizeof(*page));
+ memcpy_nontemporal(page, template, sizeof(*page));
}
/*
--
2.20.1
next prev parent reply other threads:[~2026-08-03 7:13 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-03 7:09 [PATCH v9 0/8] mm: optimize zone-device memmap initialization Li Zhe
2026-08-03 7:09 ` [PATCH v9 1/8] mm: fix stale ZONE_DEVICE refcount comment Li Zhe
2026-08-03 7:09 ` [PATCH v9 2/8] mm: factor zone-device page init helpers out of __init_zone_device_page Li Zhe
2026-08-03 7:09 ` [PATCH v9 3/8] mm: add a set_page_section_from_pfn() helper Li Zhe
2026-08-03 7:09 ` [PATCH v9 4/8] mm: add a template-based fast path for zone-device page init Li Zhe
2026-08-03 8:39 ` Muchun Song
2026-08-05 9:50 ` Li Zhe
2026-08-03 7:09 ` [PATCH v9 5/8] mm: extend the template fast path to zone-device compound tails Li Zhe
2026-08-03 7:09 ` [PATCH v9 6/8] string: introduce memcpy_nontemporal() Li Zhe
2026-08-03 7:09 ` Li Zhe [this message]
2026-08-03 7:09 ` [PATCH v9 8/8] x86/string: extend memcpy_flushcache() fixed-size fastpaths Li Zhe
2026-08-04 20:35 ` Borislav Petkov
2026-08-05 11:04 ` Li Zhe
2026-08-03 21:40 ` [PATCH v9 0/8] mm: optimize zone-device memmap initialization Andrew Morton
2026-08-05 9:49 ` Li Zhe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260803070929.86075-8-lizhe.67@bytedance.com \
--to=lizhe.67@bytedance.com \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=arnd@arndb.de \
--cc=balbirs@nvidia.com \
--cc=bp@alien8.de \
--cc=dave.hansen@linux.intel.com \
--cc=david@kernel.org \
--cc=kees@kernel.org \
--cc=linux-arch@vger.kernel.org \
--cc=linux-hardening@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mingo@redhat.com \
--cc=muchun.song@linux.dev \
--cc=rppt@kernel.org \
--cc=tglx@kernel.org \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox