* + mm-extend-the-template-fast-path-to-zone-device-compound-tails.patch added to mm-new branch
@ 2026-08-31 23:46 Andrew Morton
0 siblings, 0 replies; 2+ messages in thread
From: Andrew Morton @ 2026-08-31 23:46 UTC (permalink / raw)
To: mm-commits, rppt, muchun.song, mingo, kees, david, dave.hansen,
bp, balbirs, arnd, apopple, lizhe.67, akpm
The patch titled
Subject: mm: extend the template fast path to zone-device compound tails
has been added to the -mm mm-new branch. Its filename is
mm-extend-the-template-fast-path-to-zone-device-compound-tails.patch
This patch will shortly appear at
https://git.kernel.org/pub/scm/linux/kernel/git/akpm/25-new.git/tree/patches/mm-extend-the-template-fast-path-to-zone-device-compound-tails.patch
This patch will later appear in the mm-new branch at
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Note, mm-new is a provisional staging ground for work-in-progress
patches, and acceptance into mm-new is a notification for others take
notice and to finish up reviews. Please do not hesitate to respond to
review feedback and post updated versions to replace or incrementally
fixup patches in mm-new.
The mm-new branch of mm.git is not included in linux-next
If a few days of testing in mm-new is successful, the patch will me moved
into mm.git's mm-unstable branch, which is included in linux-next
Before you just go and hit "reply", please:
a) Consider who else should be cc'ed
b) Prefer to cc a suitable mailing list as well
c) Ideally: find the original patch on the mailing list and do a
reply-to-all to that, adding suitable additional cc's
*** Remember to use Documentation/process/submit-checklist.rst when testing your code ***
The -mm tree is included into linux-next via various
branches at git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
and is updated there most days
------------------------------------------------------
From: "Li Zhe" <lizhe.67@bytedance.com>
Subject: mm: extend the template fast path to zone-device compound tails
Date: Mon, 31 Aug 2026 19:16:35 +0800
The template fast path from the previous patch only accelerates head
pages. Compound tails in memmap_init_compound() still go through the
normal initialization path one by one.
Build separate head and tail templates and reuse one prepared tail
template across the tail pages in a compound range. Head pages preserve
the existing refcount policy, while compound tails always start with a
refcount of 0 after prep_compound_tail().
This extends the template-copy fast path to pfns_per_compound > 1.
Tail-page PFN-dependent fields are refreshed in the reusable tail template
before each copy.
Do not keep a separate non-template fallback for compound tails either.
These pages are still under memmap initialization, and the
initialization-time refcount updates are not part of the observable
lifetime of pages handed out later.
The impact is controlled for the same reason as for head pages. The first
tail page still seeds the reusable tail template through the normal tail
initialization sequence, and the copied tail pages have the same final
initialized state except for the PFN-dependent fields refreshed before
each copy.
Tested in a VM with a 100 GB devdax namespace (align=2097152) on Intel Ice
Lake server. This test exercises the dax_pmem rebind path and measures
memmap initialization latency.
Test procedure: Unbind and rebind the dax_pmem driver 30 times, collect
memmap initialization time from the pr_debug() output of
memmap_init_zone_device().
Base(v7.3-rc1):
Average of rebinds for dax_pmem driver: 191.20 ms
With this patch and its prerequisites applied:
Average of rebinds for dax_pmem driver: 176.87 ms
This reduces the average memmap initialization time measured during
rebind from 191.20 ms to 176.87 ms, or about 7.5%.
Link: https://lore.kernel.org/20260831111638.76012-5-lizhe.67@bytedance.com
Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
Cc: Alistair Popple <apopple@nvidia.com>
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: Balbir Singh <balbirs@nvidia.com>
Cc: "Borislav Petkov (AMD)" <bp@alien8.de>
Cc: Dave Hansen <dave.hansen@linux.intel.com>
Cc: David Hildenbrand (Arm) <david@kernel.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Kees Cook <kees@kernel.org>
Cc: Mike Rapoport (Microsoft) <rppt@kernel.org>
Cc: Muchun Song <muchun.song@linux.dev>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
---
mm/mm_init.c | 24 ++++++++++++++++++------
1 file changed, 18 insertions(+), 6 deletions(-)
--- a/mm/mm_init.c~mm-extend-the-template-fast-path-to-zone-device-compound-tails
+++ a/mm/mm_init.c
@@ -1072,6 +1072,8 @@ static void __ref memmap_init_compound(s
{
unsigned long pfn, end_pfn = head_pfn + nr_pages;
unsigned int order = pgmap->vmemmap_shift;
+ struct page template;
+ struct page *page;
/*
* We have to initialize the pages, including setting up page links.
@@ -1080,13 +1082,23 @@ static void __ref memmap_init_compound(s
* the pages in the same go.
*/
__SetPageHead(head);
- for (pfn = head_pfn + 1; pfn < end_pfn; pfn++) {
- struct page *page = pfn_to_page(pfn);
- __init_zone_device_page(page, pfn, zone_idx, nid, pgmap);
- prep_compound_tail(page, head, order);
- set_page_count(page, 0);
- }
+ /*
+ * All tails of the same compound page share the state established by
+ * prep_compound_tail(). Reuse one tail template for the whole range and
+ * refresh only the PFN-dependent fields in that template before each copy.
+ */
+ pfn = head_pfn + 1;
+ page = pfn_to_page(pfn);
+ __init_zone_device_page(page, pfn, zone_idx, nid, pgmap);
+ prep_compound_tail(page, head, order);
+ set_page_count(page, 0);
+ memcpy(&template, page, sizeof(*page));
+
+ /* Initialize the remaining tail pages from template. */
+ for (pfn = head_pfn + 2; pfn < end_pfn; pfn++)
+ zone_device_page_init_from_template(pfn_to_page(pfn), pfn,
+ &template);
prep_compound_head(head, order);
}
_
Patches currently in -mm which might be from lizhe.67@bytedance.com are
mm-fix-stale-zone_device-refcount-comment.patch
mm-add-a-set_page_section_from_pfn-helper.patch
mm-add-a-template-based-fast-path-for-zone-device-page-init.patch
mm-extend-the-template-fast-path-to-zone-device-compound-tails.patch
string-introduce-memcpy_nontemporal.patch
mm-use-memcpy_nontemporal-in-zone-device-template-copies.patch
x86-string-extend-memcpy_flushcache-fixed-size-fastpaths.patch
^ permalink raw reply [flat|nested] 2+ messages in thread* + mm-extend-the-template-fast-path-to-zone-device-compound-tails.patch added to mm-new branch
@ 2026-07-01 23:28 Andrew Morton
0 siblings, 0 replies; 2+ messages in thread
From: Andrew Morton @ 2026-07-01 23:28 UTC (permalink / raw)
To: mm-commits, lizhe.67, akpm
The patch titled
Subject: mm: extend the template fast path to zone-device compound tails
has been added to the -mm mm-new branch. Its filename is
mm-extend-the-template-fast-path-to-zone-device-compound-tails.patch
This patch will shortly appear at
https://git.kernel.org/pub/scm/linux/kernel/git/akpm/25-new.git/tree/patches/mm-extend-the-template-fast-path-to-zone-device-compound-tails.patch
This patch will later appear in the mm-new branch at
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Note, mm-new is a provisional staging ground for work-in-progress
patches, and acceptance into mm-new is a notification for others take
notice and to finish up reviews. Please do not hesitate to respond to
review feedback and post updated versions to replace or incrementally
fixup patches in mm-new.
The mm-new branch of mm.git is not included in linux-next
If a few days of testing in mm-new is successful, the patch will me moved
into mm.git's mm-unstable branch, which is included in linux-next
Before you just go and hit "reply", please:
a) Consider who else should be cc'ed
b) Prefer to cc a suitable mailing list as well
c) Ideally: find the original patch on the mailing list and do a
reply-to-all to that, adding suitable additional cc's
*** Remember to use Documentation/process/submit-checklist.rst when testing your code ***
The -mm tree is included into linux-next via various
branches at git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
and is updated there most days
------------------------------------------------------
From: "Li Zhe" <lizhe.67@bytedance.com>
Subject: mm: extend the template fast path to zone-device compound tails
Date: Wed, 1 Jul 2026 17:05:50 +0800
The template fast path from the previous patch only accelerates head
pages. Compound tails in memmap_init_compound() still go through the slow
path one by one.
Build separate head and tail templates and reuse one prepared tail
template across the tail pages in a compound range. Head pages preserve
the existing refcount policy, while compound tails always start with a
refcount of 0 after prep_compound_tail().
This extends the template-copy fast path to pfns_per_compound > 1 without
changing the existing slow path. Tail-page PFN-dependent fields are
refreshed in the reusable tail template before each copy.
Tested in a VM with a 100 GB devdax namespace (align=2097152) on Intel Ice
Lake server. This test exercises the dax_pmem rebind path and measures
memmap initialization latency.
Test procedure:
Unbind and rebind the dax_pmem driver 30 times, collect memmap
initialization time from the pr_debug() output of memmap_init_zone_device().
Base(v7.2-rc1):
First binding: 1462 ms
Average of subsequent rebinds: 273.31 ms
With this patch and its prerequisites applied:
First binding: 1403 ms
Average of subsequent rebinds: 244.37 ms
This reduces the average rebind time from 273.31 ms to 244.37 ms, or
about 10.6%.
Link: https://lore.kernel.org/20260701090553.62691-6-lizhe.67@bytedance.com
Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
Cc: Alistair Popple <apopple@nvidia.com>
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: Balbir Singh <balbirs@nvidia.com>
Cc: "Borislav Petkov (AMD)" <bp@alien8.de>
Cc: David Hildenbrand <david@kernel.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Kees Cook <kees@kernel.org>
Cc: Mike Rapoport (Microsoft) <rppt@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
---
mm/mm_init.c | 47 ++++++++++++++++++++++++++++++++++++++++-------
1 file changed, 40 insertions(+), 7 deletions(-)
--- a/mm/mm_init.c~mm-extend-the-template-fast-path-to-zone-device-compound-tails
+++ a/mm/mm_init.c
@@ -1071,6 +1071,16 @@ static inline void zone_device_template_
memcpy(template, src, sizeof(*template));
}
+static inline void zone_device_tail_page_init(struct page *page,
+ unsigned long pfn, unsigned long zone_idx, int nid,
+ struct dev_pagemap *pgmap, const struct page *head,
+ unsigned int order)
+{
+ zone_device_page_init_slow(page, pfn, zone_idx, nid, pgmap);
+ prep_compound_tail(page, head, order);
+ set_page_count(page, 0);
+}
+
/*
* 'template' is a reusable page prototype rather than a strictly immutable
* object. Most ZONE_DEVICE fields stay constant across the pages covered by
@@ -1128,10 +1138,12 @@ static void __ref memmap_init_compound(s
unsigned long head_pfn,
unsigned long zone_idx, int nid,
struct dev_pagemap *pgmap,
- unsigned long nr_pages)
+ unsigned long nr_pages,
+ bool use_template)
{
unsigned long pfn, end_pfn = head_pfn + nr_pages;
unsigned int order = pgmap->vmemmap_shift;
+ struct page template;
/*
* We have to initialize the pages, including setting up page links.
@@ -1140,12 +1152,31 @@ static void __ref memmap_init_compound(s
* the pages in the same go.
*/
__SetPageHead(head);
- for (pfn = head_pfn + 1; pfn < end_pfn; pfn++) {
+
+ pfn = head_pfn + 1;
+ /*
+ * All tails of the same compound page share the state established by
+ * prep_compound_tail(). Reuse one tail template for the whole range and
+ * refresh only the PFN-dependent fields in that template before each copy.
+ */
+ if (use_template) {
struct page *page = pfn_to_page(pfn);
- zone_device_page_init_slow(page, pfn, zone_idx, nid, pgmap);
- prep_compound_tail(page, head, order);
- set_page_count(page, 0);
+ zone_device_tail_page_init(page, pfn, zone_idx, nid,
+ pgmap, head, order);
+ zone_device_template_page_init(&template, page);
+ pfn++;
+ }
+
+ for (; pfn < end_pfn; pfn++) {
+ struct page *page = pfn_to_page(pfn);
+
+ if (use_template)
+ zone_device_page_init_from_template(page, pfn,
+ &template);
+ else
+ zone_device_tail_page_init(page, pfn, zone_idx, nid,
+ pgmap, head, order);
}
prep_compound_head(head, order);
}
@@ -1195,7 +1226,8 @@ void __ref memmap_init_zone_device(struc
zone_device_template_page_init(&template, page);
if (pfns_per_compound != 1)
memmap_init_compound(page, pfn, zone_idx, nid, pgmap,
- compound_nr_pages(start_pfn, altmap, pgmap));
+ compound_nr_pages(start_pfn, altmap, pgmap),
+ use_template);
pfn += pfns_per_compound;
}
@@ -1216,7 +1248,8 @@ void __ref memmap_init_zone_device(struc
continue;
memmap_init_compound(page, pfn, zone_idx, nid, pgmap,
- compound_nr_pages(pfn, altmap, pgmap));
+ compound_nr_pages(pfn, altmap, pgmap),
+ use_template);
}
pageblock_migratetype_init_range(start_pfn, nr_pages, MIGRATE_MOVABLE, false);
_
Patches currently in -mm which might be from lizhe.67@bytedance.com are
mm-fix-stale-zone_device-refcount-comment.patch
mm-factor-zone-device-page-init-helpers-out-of-__init_zone_device_page.patch
mm-add-a-set_page_section_from_pfn-helper.patch
mm-add-a-template-based-fast-path-for-zone-device-page-init.patch
mm-extend-the-template-fast-path-to-zone-device-compound-tails.patch
string-introduce-memcpy_nt-helpers.patch
x86-string-extend-memcpy_flushcache-fixed-size-fastpaths.patch
mm-use-memcpy_nt-in-zone-device-template-copies.patch
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-08-31 23:46 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-31 23:46 + mm-extend-the-template-fast-path-to-zone-device-compound-tails.patch added to mm-new branch Andrew Morton
-- strict thread matches above, loose matches on Subject: below --
2026-07-01 23:28 Andrew Morton
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.