From: Muchun Song <songmuchun@bytedance.com>
To: Andrew Morton <akpm@linux-foundation.org>,
Dan Williams <djbw@kernel.org>,
David Hildenbrand <david@kernel.org>
Cc: linux-mm@kvack.org, nvdimm@lists.linux.dev,
linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org,
linux-cxl@vger.kernel.org,
Vishal Verma <vishal.l.verma@intel.com>,
Dave Jiang <dave.jiang@intel.com>,
Alison Schofield <alison.schofield@intel.com>,
Mike Rapoport <rppt@kernel.org>,
Oscar Salvador <osalvador@suse.de>, Ira Weiny <iweiny@kernel.org>,
Jan Kara <jack@suse.cz>, Matthew Wilcox <willy@infradead.org>,
Lorenzo Stoakes <ljs@kernel.org>,
Vlastimil Babka <vbabka@kernel.org>,
Michal Hocko <mhocko@suse.com>, Qi Zheng <qi.zheng@linux.dev>,
Muchun Song <songmuchun@bytedance.com>,
muchun.song@linux.dev
Subject: [PATCH 4/4] fsdax: materialize pmem vmemmap metadata on faults
Date: Thu, 3 Sep 2026 20:21:27 +0800 [thread overview]
Message-ID: <20260903122128.12264-5-songmuchun@bytedance.com> (raw)
In-Reply-To: <20260903122128.12264-1-songmuchun@bytedance.com>
Pmem FS-DAX registers the device range as ZONE_DEVICE memory, and the
kernel normally allocates and initializes vmemmap storage for every
advertised PFN up front. Sparse backing storage and workloads that only use
the DAX direct-access path may never need writable per-PFN state for most
of that range, but still pay the memory and initialization cost.
Opt pmem FS-DAX into the shared read-only vmemmap mode and materialize
private metadata before inserting a PFN into a userspace mapping. PMD
faults materialize the whole PMD-sized metadata range. PFNs that are never
faulted continue to use the shared metadata.
This shifts private vmemmap allocation from device registration to the
first DAX fault for each metadata page, so the struct page overhead tracks
the faulted DAX working set instead of the full advertised pmem device
size.
Signed-off-by: Muchun Song <songmuchun@bytedance.com>
---
drivers/nvdimm/pmem.c | 1 +
fs/dax.c | 5 +++++
2 files changed, 6 insertions(+)
diff --git a/drivers/nvdimm/pmem.c b/drivers/nvdimm/pmem.c
index b14f75daeda7..5530b32163d3 100644
--- a/drivers/nvdimm/pmem.c
+++ b/drivers/nvdimm/pmem.c
@@ -543,6 +543,7 @@ static int pmem_attach_disk(struct device *dev,
pmem->pgmap.nr_range = 1;
pmem->pgmap.type = MEMORY_DEVICE_FS_DAX;
pmem->pgmap.ops = &fsdax_pagemap_ops;
+ pmem->pgmap.flags |= PGMAP_VMEMMAP_OPTIMIZATION;
addr = devm_memremap_pages(dev, &pmem->pgmap);
bb_range = pmem->pgmap.range;
} else {
diff --git a/fs/dax.c b/fs/dax.c
index 1fbba0d21c13..89377301d51b 100644
--- a/fs/dax.c
+++ b/fs/dax.c
@@ -13,6 +13,7 @@
#include <linux/fs.h>
#include <linux/highmem.h>
#include <linux/memcontrol.h>
+#include <linux/memremap.h>
#include <linux/mm.h>
#include <linux/mutex.h>
#include <linux/sched.h>
@@ -1877,6 +1878,10 @@ static vm_fault_t dax_fault_iter(struct vm_fault *vmf,
if (err)
return pmd ? VM_FAULT_FALLBACK : dax_fault_return(err);
+ err = vmemmap_materialize_page(pfn_to_page(pfn), pmd ? PMD_ORDER : 0);
+ if (err)
+ return dax_fault_return(err);
+
*entry = dax_insert_entry(xas, vmf, iter, *entry, pfn, entry_flags);
if (write && iomap->flags & IOMAP_F_SHARED) {
--
2.54.0
next prev parent reply other threads:[~2026-09-03 12:22 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 12:21 [PATCH 0/4] mm: Reduce struct page overhead for FS-DAX pmem Muchun Song
2026-09-03 12:21 ` [PATCH 1/4] mm: generalize vmemmap remap architecture support Muchun Song
2026-09-21 12:25 ` Oscar Salvador (SUSE)
2026-09-03 12:21 ` [PATCH 2/4] nvdimm/pmem: avoid HWPoison flag updates for clean pages Muchun Song
2026-09-03 12:41 ` sashiko-bot
2026-09-21 7:33 ` Gupta, Pankaj
2026-09-21 12:53 ` Oscar Salvador (SUSE)
2026-09-22 2:33 ` Muchun Song
2026-09-03 12:21 ` [PATCH 3/4] mm: add shared read-only vmemmap support for FS-DAX Muchun Song
2026-09-21 6:42 ` Gupta, Pankaj
2026-09-21 9:33 ` Muchun Song
2026-09-21 11:05 ` Gupta, Pankaj
2026-09-21 13:09 ` Oscar Salvador (SUSE)
2026-09-22 3:17 ` Muchun Song
2026-09-03 12:21 ` Muchun Song [this message]
2026-09-03 12:53 ` [PATCH 4/4] fsdax: materialize pmem vmemmap metadata on faults sashiko-bot
2026-09-05 3:26 ` Muchun Song
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260903122128.12264-5-songmuchun@bytedance.com \
--to=songmuchun@bytedance.com \
--cc=akpm@linux-foundation.org \
--cc=alison.schofield@intel.com \
--cc=dave.jiang@intel.com \
--cc=david@kernel.org \
--cc=djbw@kernel.org \
--cc=iweiny@kernel.org \
--cc=jack@suse.cz \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=muchun.song@linux.dev \
--cc=nvdimm@lists.linux.dev \
--cc=osalvador@suse.de \
--cc=qi.zheng@linux.dev \
--cc=rppt@kernel.org \
--cc=vbabka@kernel.org \
--cc=vishal.l.verma@intel.com \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.