From: Gregory Price <gourry@gourry.net>
To: linux-mm@kvack.org
Cc: Zhigang.Luo@amd.com, arun.george@samsung.com, balbirs@nvidia.com,
brendan.jackman@linux.dev, yuzenghui@huawei.com,
apopple@nvidia.com, alucerop@amd.com, matthew.brost@intel.com,
akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org,
liam@infradead.org, vbabka@kernel.org, rppt@kernel.org,
surenb@google.com, mhocko@suse.com, corbet@lwn.net,
skhan@linuxfoundation.org, gregkh@linuxfoundation.org,
rafael@kernel.org, dakr@kernel.org, djbw@kernel.org,
vishal.l.verma@intel.com, dave.jiang@intel.com,
alison.schofield@intel.com, osandov@osandov.com,
jannh@google.com, pfalcato@suse.de, jackmanb@google.com,
hannes@cmpxchg.org, ziy@nvidia.com, pbonzini@redhat.com,
osalvador@suse.de, joshua.hahnjy@gmail.com, rakie.kim@sk.com,
byungchul@sk.com, ying.huang@linux.alibaba.com,
kasong@tencent.com, qi.zheng@linux.dev, shakeel.butt@linux.dev,
baohua@kernel.org, axelrasmussen@google.com, yuanchu@google.com,
weixugc@google.com, yury.norov@gmail.com,
linux@rasmusvillemoes.dk, longman@redhat.com,
ridong.chen@linux.dev, tj@kernel.org, mkoutny@suse.com,
sj@kernel.org, jgg@ziepe.ca, jhubbard@nvidia.com,
peterx@redhat.com, baolin.wang@linux.alibaba.com,
npache@redhat.com, ryan.roberts@arm.com, dev.jain@arm.com,
lance.yang@linux.dev, usama.arif@linux.dev, xu.xin16@zte.com.cn,
chengming.zhou@linux.dev, roman.gushchin@linux.dev,
muchun.song@linux.dev, linux-kernel@vger.kernel.org,
linux-doc@vger.kernel.org, driver-core@lists.linux.dev,
nvdimm@lists.linux.dev, linux-cxl@vger.kernel.org,
linux-debuggers@vger.kernel.org, linux-fsdevel@vger.kernel.org,
kvm@vger.kernel.org, cgroups@vger.kernel.org,
damon@lists.linux.dev, linux-kselftest@vger.kernel.org,
kernel-team@meta.com
Subject: Re: [PATCH v5 00/36] Private Memory NUMA Nodes
Date: Fri, 25 Sep 2026 19:02:59 -0400 [thread overview]
Message-ID: <arb0mV23QmRVDtB3@gourry-fedora-PF4VCD3F> (raw)
In-Reply-To: <20260720193431.3841992-1-gourry@gourry.net>
On Mon, Jul 20, 2026 at 03:33:54PM -0400, Gregory Price wrote:
> This series introduces the concept of "Private Memory Nodes", which
> are opted into two basic functionalities by default:
> - page allocation (mm/page_alloc.c)
> - OOM killing
>
All, I wanted to preface v6 with an update, I probably won't be sending
it out on-list before LPC (i've been a bit noisey as it is, I'll give
it a rest).
But I would like to give you the highlights for those interested.
The latest working branch can be found here:
https://github.com/gourryinverse/linux/tree/scratch/gourry/managed_nodes/rfc6-nonuma
You will find the following:
- There is no depedency on exported alloc_flags (willy, i found a way),
but there is a need for `folio_alloc_private()` to explicitly ask
for the ALLOC_ZONELIST_PRIVATE.
- feature bits have been converted to node_state[] bits
- N_MEMORY_COMMON was added to mean "common memory" - i.e. the existing
N_MEMORY set. A private node is now defined as !N_MEMORY_COMMON.
- opt-out sites no longer check for !N_MEMORY_PRIVATE, they instead
iterate over node_state[N_MEMORY_COMMON]. This is *much* cleaner.
- opt-in sites get to define bits in terms of their own service, e.g.
/* for each reclaimable node */
for_each_node_mask_state(..., N_MEMORY_RECLAIM) {
...
}
Or maybe even:
for_each_reclaimable_node() :)
- There are only 2 node features in the base set:
N_MEMORY_RECLAIM - generic reclaim runs on the node
N_MEMORY_COMPACTION - compaction will run on the node
You'll notice there's no userland numa control support, more on this
in a moment.
This is all you need to have functional device-private coherent
memory that uses the page allocator.
Disabling N_MEMORY_RECLAIM does not necessarily mean you're prevented
from using vmscan.c and the LRU - it just means we'd need to expose an
interface for a !N_MEMORY_RECLAIM node to explicitly ask for reclaim
to operate on it.
Tiering (Demotion) *to* a node is not supported. We can probably
debate whether this should mean tiering *from* a node should be
supported or not as well. This is easy to change.
- The working branch also includes a minimal compressed-ram service
that only supports anonymous memory.
Supporting file-backed memory (and shmem) is a whole different can
of worms that deserves its own discussion.
I will post this as a separate RFC from the base set, but it is
available to play with on my working branch.
- I ship a basic dax driver for the base series, `anondax` which allows
existing devices to online a private node with both compaction and
reclaim support for testing. This enables this use case:
fd = open(/dev/my_device, ..., O_DIRECT);
buf = mmap(fd, ..., MAP_SHARED)
my_file = open(myfile, ...);
read(buf, myfile); /* fault directly onto device memory */
With swap (and no tiering), this gives you a private over-commitable
node for which aspiring drivers can use to implement their own
mmu_notifier callback stream to manage device page tables and such.
- The reason mempolicy (and user numa in general) is not supported is a
critical relationship between the OOM killer and page allocator.
In trivial scenarios, if a consumer of a private node overshoots and
causes an OOM - it will be the largest consumer and be chosen as the
victim.
If there are many consumers - what actually happens in practice is the
oom killer attempts to select a victim based on *task* policy... which
makes every task a candidate. This leads to, among other things,
either an OOM storm and/or a full blown panic because a victim cannot
be located (depending on the constraints).
I've concluded that as-is, this is not tractable. But also, it is not
actually needed if the devices provide their own chardev to provide
a basic wrapper around folio_alloc_private().
This unfortunately means potential users like guest_memfd are unlikely
to be able to use the nodes without hard-coding the allocator
interface, as opposed to a mempolicy.
I think it's possible instead to *maybe* add N_MEMORY_USER_MIGRATION
and allow for explicit movement between nodes, but not full blown
mempolicies.
The relationship between page_alloc, cpuset, mempolicy, and the oom
killer is just too tight to allow allocations without the use of
__GFP_THISNODE - which the private allocation interface will enforce.
As it stands, the series has been heavily minimized in both line count
and patch number.
The base series is ~22 commits with the scary diffstat below, but many
of the core mm/ patches are similar to the patch below - with nearly
all of the real complexity happening in reclaim and memory hotplug.
despite the branch name, I intend to drop RFC from here-on, everything
else required for the base series is in mm-new.
See you at LPC
~Gregory
-------
diff --git a/mm/migrate.c b/mm/migrate.c
index 7bdcdb57652f8..cd61566b44fa1 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -2266,7 +2266,7 @@ static int __add_folio_for_migration(struct folio *folio, int node,
if (is_zero_folio(folio) || is_huge_zero_folio(folio))
return -EFAULT;
- if (folio_is_zone_device(folio))
+ if (!folio_is_common_memory(folio))
return -ENOENT;
if (folio_nid(folio) == node)
@@ -2390,7 +2390,7 @@ static int do_pages_move(struct mm_struct *mm, nodemask_t task_nodes,
err = -ENODEV;
if (node < 0 || node >= MAX_NUMNODES)
goto out_flush;
- if (!node_state(node, N_MEMORY))
+ if (!node_state(node, N_MEMORY_COMMON))
goto out_flush;
err = -EACCES;
@@ -2475,7 +2475,7 @@ static void do_pages_stat_array(struct mm_struct *mm, unsigned long nr_pages,
if (folio) {
if (is_zero_folio(folio) || is_huge_zero_folio(folio))
err = -EFAULT;
- else if (folio_is_zone_device(folio))
+ else if (!folio_is_common_memory(folio))
err = -ENOENT;
else
err = folio_nid(folio);
Documentation/ABI/stable/sysfs-devices-node | 21 +++++++++++++++++
Documentation/admin-guide/cgroup-v2.rst | 7 ++++++
Documentation/mm/physical_memory.rst | 15 ++++++++++++-
drivers/base/node.c | 67 +++++++++++++++++++++++++++++++++++++++++++++++++++++--
drivers/dax/kmem.c | 2 +-
include/linux/gfp.h | 6 +++++
include/linux/memory_hotplug.h | 2 +-
include/linux/mmzone.h | 45 +++++++++++++++++++++++++++++++++++++
include/linux/node.h | 12 ++++++++++
include/linux/nodemask.h | 53 ++++++++++++++++++++++++++++++++++++++++++-
kernel/cgroup/cpuset.c | 39 ++++++++++++++++++++++++++------
kernel/sched/fair.c | 6 ++---
mm/compaction.c | 21 +++++++++++------
mm/damon/ops-common.c | 15 ++++++++-----
mm/damon/vaddr.c | 20 +++++++++++++----
mm/huge_memory.c | 2 +-
mm/internal.h | 22 ++++++++++++++++++
mm/khugepaged.c | 9 ++++++--
mm/ksm.c | 8 ++++---
mm/madvise.c | 6 ++---
mm/memory-tiers.c | 25 +++++++++++----------
mm/memory_hotplug.c | 109 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------------------------------
mm/mempolicy.c | 51 +++++++++++++++++++++++++-----------------
mm/migrate.c | 6 ++---
mm/mm_init.c | 22 +++++++++---------
mm/page_alloc.c | 115 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---------------
mm/page_alloc.h | 5 +++++
mm/vmscan.c | 31 +++++++++++++++-----------
28 files changed, 585 insertions(+), 157 deletions(-)
prev parent reply other threads:[~2026-09-25 23:03 UTC|newest]
Thread overview: 97+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <CGME20260720193443epcas5p167fa7d8490edfb4d8c0aa6a259db058f@epcas5p1.samsung.com>
2026-07-20 19:33 ` [PATCH v5 00/36] Private Memory NUMA Nodes Gregory Price
2026-07-20 19:33 ` [PATCH v5 01/36] mm: refactor find_next_best_node to find_next_best_node_in Gregory Price
2026-07-21 5:46 ` Balbir Singh
2026-07-20 19:33 ` [PATCH v5 02/36] mm/page_alloc: refactor build_node_zonelist() out of build_zonelists() Gregory Price
2026-07-20 19:33 ` [PATCH v5 03/36] mm/page_alloc: let the bulk and folio allocators carry alloc_flags Gregory Price
2026-07-21 18:26 ` sashiko-bot
2026-07-20 19:33 ` [PATCH v5 04/36] numa: introduce N_MEMORY_PRIVATE Gregory Price
2026-07-21 18:18 ` sashiko-bot
2026-07-20 19:33 ` [PATCH v5 05/36] mm: add ZONELIST_PRIVATE(_NOFALLBACK) for N_MEMORY_PRIVATE nodes Gregory Price
2026-07-22 13:00 ` Richard Cheng
2026-07-22 13:19 ` Gregory Price
2026-07-20 19:34 ` [PATCH v5 06/36] cpuset: exclude private nodes from cpuset.mems (default-open) Gregory Price
2026-07-21 18:22 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 07/36] mm/memory_hotplug: disallow migration-driven private node hotunplug Gregory Price
2026-07-20 19:34 ` [PATCH v5 08/36] mm/mempolicy: skip private node folios when queueing for migration Gregory Price
2026-07-21 18:22 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 09/36] mm/migrate: disallow userland driven migration for private nodes Gregory Price
2026-07-21 18:32 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 10/36] mm/madvise: disallow madvise operations on private node folios Gregory Price
2026-07-21 18:30 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 11/36] mm/compaction: disallow compaction on private nodes Gregory Price
2026-07-21 18:22 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 12/36] mm/page_alloc: clear private node watermarks and system reserves Gregory Price
2026-07-20 19:34 ` [PATCH v5 13/36] mm/mempolicy: disallow NUMA Balancing prot_none on private nodes Gregory Price
2026-07-20 19:34 ` [PATCH v5 14/36] mm/damon: skip private node memory in DAMON migration and pageout Gregory Price
2026-07-21 18:18 ` sashiko-bot
2026-07-21 23:46 ` SJ Park
2026-07-22 12:16 ` Gregory Price
2026-07-23 0:19 ` SJ Park
2026-07-23 3:25 ` Gregory Price
2026-07-23 13:36 ` SJ Park
2026-09-10 8:58 ` David Hildenbrand (Arm)
2026-09-10 14:06 ` SJ Park
2026-07-20 19:34 ` [PATCH v5 15/36] mm/ksm: skip KSM for managed-memory folios Gregory Price
2026-07-20 19:34 ` [PATCH v5 16/36] mm/khugepaged: skip private node folios when trying to collapse Gregory Price
2026-07-21 18:36 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 17/36] mm/vmscan: disallow reclaim of private node memory Gregory Price
2026-07-21 18:35 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 18/36] mm/gup: disallow longterm pin of private node folios Gregory Price
2026-07-20 19:34 ` [PATCH v5 19/36] proc: include N_MEMORY_PRIVATE nodes in numa_maps output Gregory Price
2026-07-21 18:30 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 20/36] mm/memcontrol: account private-node memory in per-node stats Gregory Price
2026-07-21 18:33 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 21/36] proc/kcore: include private-node RAM in the kcore RAM map Gregory Price
2026-07-20 20:05 ` Omar Sandoval
2026-07-20 19:34 ` [PATCH v5 22/36] mm/mempolicy: add MPOL_F_PRIVATE and zonelist selection Gregory Price
2026-07-21 19:28 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 23/36] mm/mempolicy: apply policy at the kernel zone for private-node binds Gregory Price
2026-07-21 19:39 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 24/36] mm/mempolicy: add in-kernel MPOL_BIND interfaces for drivers/services Gregory Price
2026-07-21 18:36 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 25/36] mm/memory_hotplug: support N_MEMORY_PRIVATE node hotplug Gregory Price
2026-07-21 18:33 ` sashiko-bot
2026-08-06 3:30 ` Qiqi Li
2026-08-17 0:49 ` Gregory Price
2026-09-04 1:05 ` Qiqi Li
2026-09-04 3:18 ` Gregory Price
2026-07-20 19:34 ` [PATCH v5 26/36] mm: add NODE_PRIVATE_CAP_RECLAIM for opted-in private node reclaim Gregory Price
2026-07-21 20:02 ` sashiko-bot
2026-07-22 13:36 ` Richard Cheng
2026-07-22 13:48 ` Gregory Price
2026-07-20 19:34 ` [PATCH v5 27/36] mm: add NODE_PRIVATE_CAP_USER_NUMA for userland numa controls Gregory Price
2026-07-21 20:13 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 28/36] mm: add NODE_PRIVATE_CAP_HOTUNPLUG for opted-in private nodes Gregory Price
2026-07-22 14:04 ` Gregory Price
2026-07-20 19:34 ` [PATCH v5 29/36] mm: add NODE_PRIVATE_CAP_DEMOTION for private-node tiering demotion Gregory Price
2026-07-21 20:28 ` sashiko-bot
2026-07-20 19:34 ` [PATCH v5 30/36] mm: add NODE_PRIVATE_CAP_NUMA_BALANCING for private-node NUMA balancing Gregory Price
2026-07-20 19:34 ` [PATCH v5 31/36] mm: add NODE_PRIVATE_CAP_LTPIN for private node folio pinning Gregory Price
2026-07-20 19:34 ` [PATCH v5 32/36] mm/khugepaged: base private node collapse eligiblity on actor/cap bits Gregory Price
2026-07-21 20:49 ` sashiko-bot
2026-07-22 13:24 ` Richard Cheng
2026-07-22 13:43 ` Gregory Price
2026-07-20 19:34 ` [PATCH v5 33/36] Documentation/mm: describe private (N_MEMORY_PRIVATE) memory nodes Gregory Price
2026-07-20 19:34 ` [PATCH v5 34/36] mm/mempolicy: add mpol_set_shared_policy_range() Gregory Price
2026-07-20 19:34 ` [PATCH v5 35/36] KVM: guest_memfd: bind backing memory to a NUMA node at creation Gregory Price
2026-07-21 21:11 ` sashiko-bot
2026-07-21 3:46 ` [PATCH v5 00/36] Private Memory NUMA Nodes Balbir Singh
2026-07-21 18:16 ` Gregory Price
2026-07-22 8:29 ` Balbir Singh
2026-07-22 12:28 ` Gregory Price
2026-07-21 13:26 ` Zenghui Yu
2026-07-21 17:18 ` Gregory Price
2026-07-21 18:06 ` [PATCH v5 36/36] KVM: selftests: add a guest_memfd FLAG_BIND_NODE test Gregory Price
2026-07-22 14:20 ` [PATCH v5 00/36] Private Memory NUMA Nodes Richard Cheng
2026-07-22 15:40 ` Gregory Price
2026-07-22 20:56 ` Gregory Price
2026-07-24 6:29 ` Richard Cheng
2026-07-23 8:38 ` Arun George/Arun George
2026-07-23 16:53 ` Gregory Price
2026-09-22 9:57 ` Arun George/Arun George
2026-09-22 22:04 ` Gregory Price
2026-08-12 12:01 ` Pankaj Gupta
2026-08-13 1:21 ` Gregory Price
2026-08-13 8:55 ` Pankaj Gupta
2026-08-14 13:27 ` Gregory Price
2026-09-25 23:02 ` Gregory Price [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arb0mV23QmRVDtB3@gourry-fedora-PF4VCD3F \
--to=gourry@gourry.net \
--cc=Zhigang.Luo@amd.com \
--cc=akpm@linux-foundation.org \
--cc=alison.schofield@intel.com \
--cc=alucerop@amd.com \
--cc=apopple@nvidia.com \
--cc=arun.george@samsung.com \
--cc=axelrasmussen@google.com \
--cc=balbirs@nvidia.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=brendan.jackman@linux.dev \
--cc=byungchul@sk.com \
--cc=cgroups@vger.kernel.org \
--cc=chengming.zhou@linux.dev \
--cc=corbet@lwn.net \
--cc=dakr@kernel.org \
--cc=damon@lists.linux.dev \
--cc=dave.jiang@intel.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=djbw@kernel.org \
--cc=driver-core@lists.linux.dev \
--cc=gregkh@linuxfoundation.org \
--cc=hannes@cmpxchg.org \
--cc=jackmanb@google.com \
--cc=jannh@google.com \
--cc=jgg@ziepe.ca \
--cc=jhubbard@nvidia.com \
--cc=joshua.hahnjy@gmail.com \
--cc=kasong@tencent.com \
--cc=kernel-team@meta.com \
--cc=kvm@vger.kernel.org \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-debuggers@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux@rasmusvillemoes.dk \
--cc=ljs@kernel.org \
--cc=longman@redhat.com \
--cc=matthew.brost@intel.com \
--cc=mhocko@suse.com \
--cc=mkoutny@suse.com \
--cc=muchun.song@linux.dev \
--cc=npache@redhat.com \
--cc=nvdimm@lists.linux.dev \
--cc=osalvador@suse.de \
--cc=osandov@osandov.com \
--cc=pbonzini@redhat.com \
--cc=peterx@redhat.com \
--cc=pfalcato@suse.de \
--cc=qi.zheng@linux.dev \
--cc=rafael@kernel.org \
--cc=rakie.kim@sk.com \
--cc=ridong.chen@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shakeel.butt@linux.dev \
--cc=sj@kernel.org \
--cc=skhan@linuxfoundation.org \
--cc=surenb@google.com \
--cc=tj@kernel.org \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=vishal.l.verma@intel.com \
--cc=weixugc@google.com \
--cc=xu.xin16@zte.com.cn \
--cc=ying.huang@linux.alibaba.com \
--cc=yuanchu@google.com \
--cc=yury.norov@gmail.com \
--cc=yuzenghui@huawei.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.