Linux debuggers
 help / color / mirror / Atom feed
From: Gregory Price <gourry@gourry.net>
To: Richard Cheng <icheng@nvidia.com>
Cc: linux-mm@kvack.org, Zhigang.Luo@amd.com, arun.george@samsung.com,
	 balbirs@nvidia.com, brendan.jackman@linux.dev,
	yuzenghui@huawei.com,  apopple@nvidia.com, alucerop@amd.com,
	matthew.brost@intel.com,  akpm@linux-foundation.org,
	david@kernel.org, ljs@kernel.org, liam@infradead.org,
	 vbabka@kernel.org, rppt@kernel.org, surenb@google.com,
	mhocko@suse.com,  corbet@lwn.net, skhan@linuxfoundation.org,
	gregkh@linuxfoundation.org,  rafael@kernel.org, dakr@kernel.org,
	djbw@kernel.org, vishal.l.verma@intel.com,  dave.jiang@intel.com,
	alison.schofield@intel.com, osandov@osandov.com,
	 jannh@google.com, pfalcato@suse.de, jackmanb@google.com,
	hannes@cmpxchg.org,  ziy@nvidia.com, pbonzini@redhat.com,
	osalvador@suse.de, joshua.hahnjy@gmail.com,  rakie.kim@sk.com,
	byungchul@sk.com, ying.huang@linux.alibaba.com,
	 kasong@tencent.com, qi.zheng@linux.dev, shakeel.butt@linux.dev,
	baohua@kernel.org,  axelrasmussen@google.com, yuanchu@google.com,
	weixugc@google.com, yury.norov@gmail.com,
	 linux@rasmusvillemoes.dk, longman@redhat.com,
	ridong.chen@linux.dev, tj@kernel.org,  mkoutny@suse.com,
	sj@kernel.org, jgg@ziepe.ca, jhubbard@nvidia.com,
	 peterx@redhat.com, baolin.wang@linux.alibaba.com,
	npache@redhat.com,  ryan.roberts@arm.com, dev.jain@arm.com,
	lance.yang@linux.dev, usama.arif@linux.dev,  xu.xin16@zte.com.cn,
	chengming.zhou@linux.dev, roman.gushchin@linux.dev,
	 muchun.song@linux.dev, linux-kernel@vger.kernel.org,
	linux-doc@vger.kernel.org,  driver-core@lists.linux.dev,
	nvdimm@lists.linux.dev, linux-cxl@vger.kernel.org,
	 linux-debuggers@vger.kernel.org, linux-fsdevel@vger.kernel.org,
	kvm@vger.kernel.org,  cgroups@vger.kernel.org,
	damon@lists.linux.dev, linux-kselftest@vger.kernel.org,
	 kernel-team@meta.com
Subject: Re: [PATCH v5 00/36] Private Memory NUMA Nodes
Date: Wed, 22 Jul 2026 11:40:06 -0400	[thread overview]
Message-ID: <amDVQopyvolxfX8K@gourry-fedora-PF4VCD3F> (raw)
In-Reply-To: <amDL80BSEP6u9XfE@MWDK4CY14F>

On Wed, Jul 22, 2026 at 10:20:00PM +0800, Richard Cheng wrote:
> On Mon, Jul 20, 2026 at 03:33:54PM +0800, Gregory Price wrote:
> Hi Gregory,
> 
> I applied your series on mm-next and give a quick review,
> not thoroughly, still I have some questions below regarding the design.
>  

Hi Richard,

Thank you for the read.

Just a heads up, this was based on mm-new because of some recent work
on Brendan's page_alloc.h cleanup, but if you managed to get it applied
then maybe that's made its way forward already.

> > The goal here is to flip that dynamic, isolate by default and then
> > opt-in to specific services that the device says is safe.
> > 
> > Isolation at the NUMA/Zonelist layer provides a powerful mechanism for
> > memory hosted on accelerators - re-use of the kernel mm/ code.
> > 
> >  - Accelerators (GPUs) can use demotion, numactl, and reclaim.
> >  - Special memory devices (Compressed RAM) with special access controls
> >    (promote-on-write) can have generic services written for them.
> >  - Network devices with large memory regions intended for ring buffers
> >    can use the buddy and standard networking stack.
> >  - Slow, disaggregated memory pools which aren't suitable as general
> >    purpose memory get cleaner interfaces (no need to re-write the buddy
> >    in userland, can use migration interface, etc).
> >  - Per-workload dedicated memory nodes (disaggregated VM memory)
> > 
> 
> For accelerator part, take CXL Type-2 as example, it has its own protocol
> CXL.cache, CXL.mem which is the rule they need to obey during memory transaction,
> Adding more rules for them since they are NUMA node confused me, I wonder the reason ?
> 

CXL protocols simply state how to do the memory transaction at a
physical / transport level.  It does not make any statement on how the
operating systems are to make sense of these devices or what constructs
/ abstractions to build to actually manage the devices themselves.

This series is not necessarily attached to CXL in particular, you could
just as easily carve out memory from the general pool onto a private
node to ensure it only gets used for a particular use.

In fact, that is how I have been testing this with dax/kmem:
https://github.com/gourryinverse/linux/commit/f279e741d9c3643597525bfab916b29b95cb635b

> Because NUMA node, at least for me, representing topology/locality rather than
> something with ownership or capability or rules.
> 

The NUMA abstraction representing topology/locality is a construct we
(the OS developers) have decided on historically - but there's nothing
that dictates we can never create new useful abstractions with it.

Consider:
    N_NORMAL_MEMORY,        /* The node has regular memory */
    N_HIGH_MEMORY,          /* The node has regular or high memory */

These have nothing to do with topology or locality, they are node states
that only have meaning in the context of linux mm/.


To make something clear - there is no *requirement* for any particular
device to use this abstraction.  It simply enables a cleaner way for
devices carrying memory to re-utilize mm/ services while ensuring their
memory does not silently get used under system pressure (or some vagrant
in userspace doing `numactl --interleave all`).

Ignoring all the CAP bits entirely, if you just took the base series
your driver could re-use the buddy allocator without any special logic
AND have confidence that your driver is the only possible user of that
memory (barring some truly obscene bug).

> >  - isolation via a dedicated zonelist:
> >      - private nodes are omitted from FALLBACK/NOFALLBACK
> >      - added ZONELIST_PRIVATE(_NOFALLBACK)
> 
> I saw the reply in patch 5, so in fact there's not only one
> dedicate zonelist, but numerous ?
> I raise the question because zonelist was supposed to be a
> global, unbypasssable thing in MM design, but now what you
> are trying to do is to seperate the whole global list into several
> parts ?
> 
> Can you explain why in current design you don't consider to support
> something like ZONELIST_PRIVATE[n]={0, .. ,n-1} ? that's my imagination
> of what a global zonelist should look like, no matter private or non.
> 

First let me say that there's nothing that prevents us from doing this,
and we *could* make this the default case - but this decision was
intentional by me.

Having them all present in each-other's zonelists by default creates a
number of implicit opt-ins that are unclear:

  - any direct zonelist iterator now iterates all zones on all private
    nodes, even those nodes do not opt into the same services

    e.g. zonelist iteration in reclaim that targets Private Node A would
    attempt to reclaim Private Node B as well.  That would require and
    extra explicit filter.

    I ask: Why do this? Just isolate in the zonelist, and if there is
           a desire for intersections - make it explicit, not implicit.

  - fallback allocations can now occur across private nodes, even
    if those nodes are intended for different purposes.

    Obviously you can use nodemask to tighten the allocation target, but
    I use `numactl --interleave --all` as an example of a clear case
    where intersected nodelists may not give you the behavior you want.


So for consistency - everything including zonelists have full isolation.
This way all interactions with a private node must be explicit - always.


This is actually one of the problems with ZONE_DEVICE - and you can see
it in this patch set.  Some of the hooks in mm/ for zone_device only
apply to PTE cases, and are absent from PMD cases - only because PMD
mappings in ZONE_DEVICE aren't supported.

That kind of implicit behavior is quite bad.

That said, future improvements could include something like:

for_reclaimable_zone(ZONELIST_PRIVATE) {}

where we loosen this isolation, and formalize a filter, but I would
like to see the usecase for it first.  Loosening the isolation defeats
the entire purpose of the series, so there should be a strong reason to
do so.

> > Allocation Isolation
> > ====================
> > page_alloc presently controls whether a node's memory can be allocated
> > on a given call by 4 things (in order of authority)
> > 
> >   1) ZONELIST membership
> >      If a node is not in the walked zonelist, it's unreachable.
> >
> 
> So a device gets hotplugged in the system will get a dedicated zonelist here ?

Yes.

> And make sure it obeys the device's own protocol if it has one ?
> 

I'm not sure i follow this question, can you help me understand?

If by protocol you mean the CAP bits (opting into reclaim, demotion,
etc), then yes.  If you mean something else (CXL) then I think that's
orthogonal and unrelated.

> > Private nodes:
> >   1) Never appear in any ZONELIST_FALLBACK
> >   2) Have an empty ZONELIST_NOFALLBACK
> >   3) Only appear in their own ZONELIST_PRIVATE(_NOFALLBACK)
> > 
> > 1 & 2 mean all existing in-tree callers to page_alloc can NEVER
> > accidentally allocate from a private node.
> > 
> > An allocation must explicitly ask via a zonelist and a nodemask.
> > 
> >   alloc_flags |= ALLOC_ZONELIST_PRIVATE; /* use ZONELIST_PRIVATE */
> >   __alloc_pages(..., nodemask);     /* with the private node set */
> > 
> 
> As I stated above, NUMA concept was quite naive at first glance for my
> limited knowledge.
> 
> This is quite alot to add for NUMA node concept, I'll want to see
> more explanation in v6 and learn from it, thanks.
> 

Sure, I can expand on it.  I think it's not as much to add as you think
though - it's simply adding the concept of isolation to a NUMA node.

Some more explanation below, but if you think there is something i
should explicit spell out in v6 cover, please let me know.

---

I agree that NUMA as a concept was best-effort for, comically enough,
somewhat *Uniform* memory access - instead of *Non*-Uniform memory access.

The current abstraction quite nicely handles the case where all memory
on the system is roughly of the same calibre (DDR4, DDR5, etc) and
roughly for the same purpose (general system memory).

But it is quite incapable of handling truly heterogeneous memory systems
(precious HBM attached to the CPU, GPU's with HBM over a coherent link,
hardware-compressed memory expansion, network devices w/ memory, etc).

If you look at the history of ZONE_DEVICE, what it fundamentally does is
slaps an isolation mechanism on top of NUMA nodes because the NUMA
abstraction doesn't provide one. 

The problem with that approach is now you have to reason about a node
having both fungible and non-fungible memory.  It creates the need for
something like migrate_device.c when a properly isolated NUMA node could
just use migrate.c directly (with a coherent link).


One example of what isolation on the node enables:

A Private Node can hotplug memory in ZONE_NORMAL - which means it
can be GUP pinned.  That means driver support for GPU direct storage
is simply `alloc_pages_node() + pin()`. The driver doesn't have to
worry about something like SLAB accidentally using that same memory.

All without having to rewrite a bunch of mm/ in a driver, and with
having to do some kind of heroics with ZONE_DEVICE that causes even
more special mm/ interactions.

That only comes from adding an isolation primitive to a NUMA node.


> > Isolating private node folios from kernel services
> > ==================================================
> > We implement filter points in mm/ to prevent operations on
> > private node memory.  Where possible, we even re-use existing
> > filter points from ZONE_DEVICE.
> > 
> > Most filter points are one or two lines of code:
> > 
> > Combining ZONE_DEVICE and N_MEMORY_PRIVATE opt-out spots:
> >   -     if (folio_is_zone_device(folio))
> >   +     if (unlikely(folio_is_private_managed(folio)))
> > 
> > Disabling a service:
> >   +  if (!node_is_private(nid)) {
> >   +      kswapd_run(nid);
> >   +      kcompactd_run(nid);
> >   +  }
> > 
> > Disallowing a uapi interaction:
> >   +  if (node_state(nid, N_MEMORY_PRIVATE))
> >   +      return -EINVAL;
> > 
> 
> I'm not sure of why do we re-implement more filter and basically
> doing the same thing ? Any unavoidable scenario ?
> 
> re-using the exisintg filter would be nice if that's possible.
> 

We re-use (combine) existing filters where possible.

We add new ones where ZONE_DEVICE did not implement filters due to some
*implicit* filter already existing.

Two clear examples:
   - reclaim does not target ZONE_DEVICE
   - ZONE_DEVICE does not support PMD

Both cases result in private-node filters that otherwise would have to
exist for ZONE_DEVICE (and if ZONE_DEVICE ever grows PMD support, those
filters will have to be added).

> > Bonus Configuration: HBM device memory tiering
> > ==============================================
> > echo 1 > dax0.0/private       # make the node private
> > echo 0 > dax0.0/adistance     # highest tier
> > echo 1 > dax0.0/reclaim       # reclaim active
> > echo 1 > dax0.0/demotion      # may demote from the node
> > echo 1 > dax0.0/user_numa     # mbind()
> > echo online_movable > dax0.0/state
> > echo 1 > numa/demotion_enabled
> > 
> > This is an HBM device which is treated as the top-tier in the
> > system but for which memory can only enter via explicit mbind().
> > 
> > It can be overcommitted because it can be reclaimed (demotions
> > go to CPU DRAM, and reclaim can swap from it).
> > 
> > If the HBM is managed by an accelerator (GPU), the mmu_notifier
> > allows it to know when reclaim is moving memory out to do
> > device-mmu invalidation prior to migration.
> > 
> > Prereqs, base commit, references
> > ================================
> > akpm/mm-new - for Brendan Jackman's mm/page_alloc.h work[3]
> > 
> 
> Still thanks for the work, I learn alot from your work as well, thanks.
> 

Of course, and thank you for reading.

If nothing else I hope this series helps folks understand the page
allocator better (it certainly has helped me).  If there is anything
you think I can improve or better explain, I am happy to discuss.

I am planning a larger publication of all my research sometime in the
future, but I think code is more impactful and useful - so I am
prioritizing that for now :].

~Gregory

      reply	other threads:[~2026-07-22 15:40 UTC|newest]

Thread overview: 55+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-20 19:33 [PATCH v5 00/36] Private Memory NUMA Nodes Gregory Price
2026-07-20 19:33 ` [PATCH v5 01/36] mm: refactor find_next_best_node to find_next_best_node_in Gregory Price
2026-07-21  5:46   ` Balbir Singh
2026-07-20 19:33 ` [PATCH v5 02/36] mm/page_alloc: refactor build_node_zonelist() out of build_zonelists() Gregory Price
2026-07-20 19:33 ` [PATCH v5 03/36] mm/page_alloc: let the bulk and folio allocators carry alloc_flags Gregory Price
2026-07-20 19:33 ` [PATCH v5 04/36] numa: introduce N_MEMORY_PRIVATE Gregory Price
2026-07-20 19:33 ` [PATCH v5 05/36] mm: add ZONELIST_PRIVATE(_NOFALLBACK) for N_MEMORY_PRIVATE nodes Gregory Price
2026-07-22 13:00   ` Richard Cheng
2026-07-22 13:19     ` Gregory Price
2026-07-20 19:34 ` [PATCH v5 06/36] cpuset: exclude private nodes from cpuset.mems (default-open) Gregory Price
2026-07-20 19:34 ` [PATCH v5 07/36] mm/memory_hotplug: disallow migration-driven private node hotunplug Gregory Price
2026-07-20 19:34 ` [PATCH v5 08/36] mm/mempolicy: skip private node folios when queueing for migration Gregory Price
2026-07-20 19:34 ` [PATCH v5 09/36] mm/migrate: disallow userland driven migration for private nodes Gregory Price
2026-07-20 19:34 ` [PATCH v5 10/36] mm/madvise: disallow madvise operations on private node folios Gregory Price
2026-07-20 19:34 ` [PATCH v5 11/36] mm/compaction: disallow compaction on private nodes Gregory Price
2026-07-20 19:34 ` [PATCH v5 12/36] mm/page_alloc: clear private node watermarks and system reserves Gregory Price
2026-07-20 19:34 ` [PATCH v5 13/36] mm/mempolicy: disallow NUMA Balancing prot_none on private nodes Gregory Price
2026-07-20 19:34 ` [PATCH v5 14/36] mm/damon: skip private node memory in DAMON migration and pageout Gregory Price
2026-07-21 23:46   ` SJ Park
2026-07-22 12:16     ` Gregory Price
2026-07-20 19:34 ` [PATCH v5 15/36] mm/ksm: skip KSM for managed-memory folios Gregory Price
2026-07-20 19:34 ` [PATCH v5 16/36] mm/khugepaged: skip private node folios when trying to collapse Gregory Price
2026-07-20 19:34 ` [PATCH v5 17/36] mm/vmscan: disallow reclaim of private node memory Gregory Price
2026-07-20 19:34 ` [PATCH v5 18/36] mm/gup: disallow longterm pin of private node folios Gregory Price
2026-07-20 19:34 ` [PATCH v5 19/36] proc: include N_MEMORY_PRIVATE nodes in numa_maps output Gregory Price
2026-07-20 19:34 ` [PATCH v5 20/36] mm/memcontrol: account private-node memory in per-node stats Gregory Price
2026-07-20 19:34 ` [PATCH v5 21/36] proc/kcore: include private-node RAM in the kcore RAM map Gregory Price
2026-07-20 20:05   ` Omar Sandoval
2026-07-20 19:34 ` [PATCH v5 22/36] mm/mempolicy: add MPOL_F_PRIVATE and zonelist selection Gregory Price
2026-07-20 19:34 ` [PATCH v5 23/36] mm/mempolicy: apply policy at the kernel zone for private-node binds Gregory Price
2026-07-20 19:34 ` [PATCH v5 24/36] mm/mempolicy: add in-kernel MPOL_BIND interfaces for drivers/services Gregory Price
2026-07-20 19:34 ` [PATCH v5 25/36] mm/memory_hotplug: support N_MEMORY_PRIVATE node hotplug Gregory Price
2026-07-20 19:34 ` [PATCH v5 26/36] mm: add NODE_PRIVATE_CAP_RECLAIM for opted-in private node reclaim Gregory Price
2026-07-22 13:36   ` Richard Cheng
2026-07-22 13:48     ` Gregory Price
2026-07-20 19:34 ` [PATCH v5 27/36] mm: add NODE_PRIVATE_CAP_USER_NUMA for userland numa controls Gregory Price
2026-07-20 19:34 ` [PATCH v5 28/36] mm: add NODE_PRIVATE_CAP_HOTUNPLUG for opted-in private nodes Gregory Price
2026-07-22 14:04   ` Gregory Price
2026-07-20 19:34 ` [PATCH v5 29/36] mm: add NODE_PRIVATE_CAP_DEMOTION for private-node tiering demotion Gregory Price
2026-07-20 19:34 ` [PATCH v5 30/36] mm: add NODE_PRIVATE_CAP_NUMA_BALANCING for private-node NUMA balancing Gregory Price
2026-07-20 19:34 ` [PATCH v5 31/36] mm: add NODE_PRIVATE_CAP_LTPIN for private node folio pinning Gregory Price
2026-07-20 19:34 ` [PATCH v5 32/36] mm/khugepaged: base private node collapse eligiblity on actor/cap bits Gregory Price
2026-07-22 13:24   ` Richard Cheng
2026-07-22 13:43     ` Gregory Price
2026-07-20 19:34 ` [PATCH v5 33/36] Documentation/mm: describe private (N_MEMORY_PRIVATE) memory nodes Gregory Price
2026-07-20 19:34 ` [PATCH v5 34/36] mm/mempolicy: add mpol_set_shared_policy_range() Gregory Price
2026-07-20 19:34 ` [PATCH v5 35/36] KVM: guest_memfd: bind backing memory to a NUMA node at creation Gregory Price
2026-07-21  3:46 ` [PATCH v5 00/36] Private Memory NUMA Nodes Balbir Singh
2026-07-21 18:16   ` Gregory Price
2026-07-22  8:29     ` Balbir Singh
2026-07-22 12:28       ` Gregory Price
2026-07-21 13:26 ` Zenghui Yu
2026-07-21 17:18   ` Gregory Price
2026-07-22 14:20 ` Richard Cheng
2026-07-22 15:40   ` Gregory Price [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=amDVQopyvolxfX8K@gourry-fedora-PF4VCD3F \
    --to=gourry@gourry.net \
    --cc=Zhigang.Luo@amd.com \
    --cc=akpm@linux-foundation.org \
    --cc=alison.schofield@intel.com \
    --cc=alucerop@amd.com \
    --cc=apopple@nvidia.com \
    --cc=arun.george@samsung.com \
    --cc=axelrasmussen@google.com \
    --cc=balbirs@nvidia.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=brendan.jackman@linux.dev \
    --cc=byungchul@sk.com \
    --cc=cgroups@vger.kernel.org \
    --cc=chengming.zhou@linux.dev \
    --cc=corbet@lwn.net \
    --cc=dakr@kernel.org \
    --cc=damon@lists.linux.dev \
    --cc=dave.jiang@intel.com \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=djbw@kernel.org \
    --cc=driver-core@lists.linux.dev \
    --cc=gregkh@linuxfoundation.org \
    --cc=hannes@cmpxchg.org \
    --cc=icheng@nvidia.com \
    --cc=jackmanb@google.com \
    --cc=jannh@google.com \
    --cc=jgg@ziepe.ca \
    --cc=jhubbard@nvidia.com \
    --cc=joshua.hahnjy@gmail.com \
    --cc=kasong@tencent.com \
    --cc=kernel-team@meta.com \
    --cc=kvm@vger.kernel.org \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-debuggers@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux@rasmusvillemoes.dk \
    --cc=ljs@kernel.org \
    --cc=longman@redhat.com \
    --cc=matthew.brost@intel.com \
    --cc=mhocko@suse.com \
    --cc=mkoutny@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=npache@redhat.com \
    --cc=nvdimm@lists.linux.dev \
    --cc=osalvador@suse.de \
    --cc=osandov@osandov.com \
    --cc=pbonzini@redhat.com \
    --cc=peterx@redhat.com \
    --cc=pfalcato@suse.de \
    --cc=qi.zheng@linux.dev \
    --cc=rafael@kernel.org \
    --cc=rakie.kim@sk.com \
    --cc=ridong.chen@linux.dev \
    --cc=roman.gushchin@linux.dev \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=shakeel.butt@linux.dev \
    --cc=sj@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=surenb@google.com \
    --cc=tj@kernel.org \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=vishal.l.verma@intel.com \
    --cc=weixugc@google.com \
    --cc=xu.xin16@zte.com.cn \
    --cc=ying.huang@linux.alibaba.com \
    --cc=yuanchu@google.com \
    --cc=yury.norov@gmail.com \
    --cc=yuzenghui@huawei.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox