From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6155439A7F6 for ; Wed, 1 Jul 2026 23:28:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782948516; cv=none; b=gJZqasa6ZuRh5kYSoV8wVLIX2ncHVPCMEWHi+wAUm6Vaq9lp8OX12jleVme+ZPuD4zRt0nOBruLMC+DcplO4ZIZKKLyOkqCm3tTb60R2iZYQdhW9I5wapbHrfd7GsRz+b6AFBakByais95UJXUGSVsx2VlXrjC8XnZhzimfocIo= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782948516; c=relaxed/simple; bh=rXW1XM4kR7nvsr/Fr7GT6kq/PraGGjsFH8RR1FuE6q4=; h=Date:To:From:Subject:Message-Id; b=Z6YNn4cGqC48KMCjLDIEB8A1hK5hjg7RSCPHFBh7ViY97JEHU1ihVjGw8knSCuK4fI6LnKgOR0X7eVn++hKsOytgYvCj1M3DeeimqH4kl4CQf5al1AxaZSJNRVoPlb7CsVx8eqD0lrKwYMsBHgObptn4Qwmz3L4d2R9pJT8MVqM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b=wIzFrD7P; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b="wIzFrD7P" Received: by smtp.kernel.org (Postfix) with ESMTPSA id A77991F000E9; Wed, 1 Jul 2026 23:28:34 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux-foundation.org; s=korg; t=1782948514; bh=xrcLBEV6BiVlMuXhdcVJT3AKbkPSBnYuPRivqcfQPKY=; h=Date:To:From:Subject; b=wIzFrD7P8zkWo9/btSScbfhZ67L+OvKu3BZJuq6dW02RYEENoxKzSqZJGUnDsLgl5 0wHqK/zBk16pACtwLx1zZu/k/mSTWf8zAGd0ypR/Ng5mC4xShov40CIR5I/wITbTNj Z0/9gYVsLVnt8SyR4TIw8YhVV7B7Aa23iMvXhiUg= Date: Wed, 01 Jul 2026 16:28:34 -0700 To: mm-commits@vger.kernel.org,lizhe.67@bytedance.com,akpm@linux-foundation.org From: Andrew Morton Subject: + mm-add-a-template-based-fast-path-for-zone-device-page-init.patch added to mm-new branch Message-Id: <20260701232834.A77991F000E9@smtp.kernel.org> Precedence: bulk X-Mailing-List: mm-commits@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: The patch titled Subject: mm: add a template-based fast path for zone-device page init has been added to the -mm mm-new branch. Its filename is mm-add-a-template-based-fast-path-for-zone-device-page-init.patch This patch will shortly appear at https://git.kernel.org/pub/scm/linux/kernel/git/akpm/25-new.git/tree/patches/mm-add-a-template-based-fast-path-for-zone-device-page-init.patch This patch will later appear in the mm-new branch at git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm Note, mm-new is a provisional staging ground for work-in-progress patches, and acceptance into mm-new is a notification for others take notice and to finish up reviews. Please do not hesitate to respond to review feedback and post updated versions to replace or incrementally fixup patches in mm-new. The mm-new branch of mm.git is not included in linux-next If a few days of testing in mm-new is successful, the patch will me moved into mm.git's mm-unstable branch, which is included in linux-next Before you just go and hit "reply", please: a) Consider who else should be cc'ed b) Prefer to cc a suitable mailing list as well c) Ideally: find the original patch on the mailing list and do a reply-to-all to that, adding suitable additional cc's *** Remember to use Documentation/process/submit-checklist.rst when testing your code *** The -mm tree is included into linux-next via various branches at git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm and is updated there most days ------------------------------------------------------ From: "Li Zhe" Subject: mm: add a template-based fast path for zone-device page init Date: Wed, 1 Jul 2026 17:05:49 +0800 memmap_init_zone_device() repeats nearly identical head-page initialization for each PFN. Prepare one reusable ZONE_DEVICE head-page template through the existing slow path, refresh the PFN-dependent fields in that template before each copy, and memcpy it into each destination page. The optimized path assigns _refcount through the copied template, so keep it disabled when the page_ref_set tracepoint is enabled. This patch accelerates head-page initialization. The pfns_per_compound == 1 case gets the full benefit here, compound tails are handled in the next patch. Tested in a VM with a 100 GB fsdax namespace device configured with map=dev on Intel Ice Lake server. This test exercises the nd_pmem rebind path (pfns_per_compound == 1). Test procedure: Rebind the nd_pmem driver 30 times and collect the memmap initialization time from the pr_debug() output of memmap_init_zone_device(). Base(v7.2-rc1): First binding: 1456 ms Average of subsequent rebinds: 244.28 ms With this patch and its prerequisites applied: First binding: 1440 ms Average of subsequent rebinds: 217.19 ms This reduces the average rebind time from 244.28 ms to 217.19 ms, or about 11%. Link: https://lore.kernel.org/20260701090553.62691-5-lizhe.67@bytedance.com Signed-off-by: Li Zhe Cc: Alistair Popple Cc: Arnd Bergmann Cc: Balbir Singh Cc: "Borislav Petkov (AMD)" Cc: David Hildenbrand Cc: Ingo Molnar Cc: Kees Cook Cc: Mike Rapoport (Microsoft) Signed-off-by: Andrew Morton --- mm/mm_init.c | 76 +++++++++++++++++++++++++++++++++++++++++++++++-- 1 file changed, 74 insertions(+), 2 deletions(-) --- a/mm/mm_init.c~mm-add-a-template-based-fast-path-for-zone-device-page-init +++ a/mm/mm_init.c @@ -1056,6 +1056,50 @@ static void __ref zone_device_page_init_ set_page_count(page, 0); } +static inline bool zone_device_page_init_optimization_enabled(void) +{ + /* + * The template fast path copies a preinitialized struct page image. + * Skip it when the page_ref_set tracepoint is enabled. + */ + return !page_ref_tracepoint_active(page_ref_set); +} + +static inline void zone_device_template_page_init(struct page *template, + struct page *src) +{ + memcpy(template, src, sizeof(*template)); +} + +/* + * 'template' is a reusable page prototype rather than a strictly immutable + * object. Most ZONE_DEVICE fields stay constant across the pages covered by + * the current template, but section bits and page->virtual may still depend + * on the PFN. Refresh those PFN-dependent fields in the template before + * copying it into @page. + */ +static inline void zone_device_page_update_template(struct page *template, + unsigned long pfn) +{ + set_page_section_from_pfn(template, pfn); +#ifdef WANT_PAGE_VIRTUAL + if (!is_highmem_idx(ZONE_DEVICE)) + set_page_address(template, __va(pfn << PAGE_SHIFT)); +#endif +} + +static void zone_device_page_init_from_template(struct page *page, + unsigned long pfn, struct page *template) +{ + /* + * 'template' carries the invariant portion of a ZONE_DEVICE struct + * page. Update the PFN-dependent fields in place before copying it + * to the destination page. + */ + zone_device_page_update_template(template, pfn); + memcpy(page, template, sizeof(*page)); +} + /* * With compound page geometry and when struct pages are stored in ram most * tail pages are reused. Consequently, the amount of unique struct pages to @@ -1111,6 +1155,7 @@ void __ref memmap_init_zone_device(struc unsigned long nr_pages, struct dev_pagemap *pgmap) { + bool use_template = zone_device_page_init_optimization_enabled(); unsigned long pfn, end_pfn = start_pfn + nr_pages; struct pglist_data *pgdat = zone->zone_pgdat; struct vmem_altmap *altmap = pgmap_altmap(pgmap); @@ -1118,6 +1163,7 @@ void __ref memmap_init_zone_device(struc unsigned long zone_idx = zone_idx(zone); unsigned long start = jiffies; int nid = pgdat->node_id; + struct page template; if (WARN_ON_ONCE(!pgmap || zone_idx != ZONE_DEVICE)) return; @@ -1132,10 +1178,36 @@ void __ref memmap_init_zone_device(struc nr_pages = end_pfn - start_pfn; } - for (pfn = start_pfn; pfn < end_pfn; pfn += pfns_per_compound) { - struct page *page = pfn_to_page(pfn); + + if (!nr_pages) + return; + + pfn = start_pfn; + /* + * Seed the reusable head-page template from the first real struct + * page, because the existing page-init and pageblock helpers expect + * a real memmap entry rather than a stack object. + */ + if (use_template) { + struct page *page = pfn_to_page(start_pfn); zone_device_page_init_slow(page, pfn, zone_idx, nid, pgmap); + zone_device_template_page_init(&template, page); + if (pfns_per_compound != 1) + memmap_init_compound(page, pfn, zone_idx, nid, pgmap, + compound_nr_pages(start_pfn, altmap, pgmap)); + pfn += pfns_per_compound; + } + + for (; pfn < end_pfn; pfn += pfns_per_compound) { + struct page *page = pfn_to_page(pfn); + + if (use_template) + zone_device_page_init_from_template(page, pfn, + &template); + else + zone_device_page_init_slow(page, pfn, zone_idx, + nid, pgmap); if (IS_ALIGNED(pfn, PAGES_PER_SECTION)) cond_resched(); _ Patches currently in -mm which might be from lizhe.67@bytedance.com are mm-fix-stale-zone_device-refcount-comment.patch mm-factor-zone-device-page-init-helpers-out-of-__init_zone_device_page.patch mm-add-a-set_page_section_from_pfn-helper.patch mm-add-a-template-based-fast-path-for-zone-device-page-init.patch mm-extend-the-template-fast-path-to-zone-device-compound-tails.patch string-introduce-memcpy_nt-helpers.patch x86-string-extend-memcpy_flushcache-fixed-size-fastpaths.patch mm-use-memcpy_nt-in-zone-device-template-copies.patch