Linux Documentation
 help / color / mirror / Atom feed
From: "Li Zhe" <lizhe.67@bytedance.com>
To: <akpm@linux-foundation.org>, <david@kernel.org>, <ljs@kernel.org>,
	 <liam@infradead.org>, <rppt@kernel.org>, <mhocko@suse.com>,
	 <corbet@lwn.net>, <skhan@linuxfoundation.org>, <ziy@nvidia.com>,
	 <joshua.hahnjy@gmail.com>, <gourry@gourry.net>,
	 <ying.huang@linux.alibaba.com>, <apopple@nvidia.com>
Cc: <linux-mm@kvack.org>, <linux-doc@vger.kernel.org>,
	 <linux-kernel@vger.kernel.org>, <lizhe.67@bytedance.com>
Subject: [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy
Date: Wed, 30 Sep 2026 15:26:17 +0800	[thread overview]
Message-ID: <20260930072617.64665-1-lizhe.67@bytedance.com> (raw)

MPOL_WEIGHTED_INTERLEAVE is useful on tiered-memory systems because it
can seed a workload's new allocations across fast memory and slower
capacity memory according to a configured ratio.

That initial placement is useful for workload managers and orchestration
systems.  They can take the amount of fast memory and slower capacity
memory on a machine into account before starting a workload, and choose a
weighted policy that seeds the workload across the tiers at allocation
time.  This avoids starting from an all-fast or all-slow placement and
then relying on promotion or demotion to reshape a large working set.

After those pages have been placed, however, the policy cannot currently
opt in to migrate-on-fault placement.  set_mempolicy() and mbind()
reject MPOL_WEIGHTED_INTERLEAVE when MPOL_F_NUMA_BALANCING is specified,
so memory tiering cannot promote hot pages that were initially placed on
the slower nodes by the weighted policy.

Initial placement is only a starting point.  Pages initially allocated on
fast memory are not necessarily the long-term hot pages, and pages
initially allocated on slower memory may become hot as the workload's hot
set changes.  The policy therefore needs to be able to combine weighted
initial placement with memory tiering's NUMA fault based hot-page
promotion.

Allow MPOL_F_NUMA_BALANCING for MPOL_WEIGHTED_INTERLEAVE.  As with
MPOL_BIND and MPOL_PREFERRED_MANY, keep migration constrained by the
policy nodemask: if the CPU's node is outside the nodemask, do not
migrate the folio there.

This is an opt-in behavior.  MPOL_WEIGHTED_INTERLEAVE without
MPOL_F_NUMA_BALANCING keeps its existing behavior and does not
participate in NUMA balancing.

The weighted interleave weights continue to control new allocations only.
They are not a target residency ratio after migrate-on-fault placement.

Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
---
 Documentation/admin-guide/mm/numa_memory_policy.rst | 10 ++++++++++
 mm/mempolicy.c                                      |  8 +++++++-
 2 files changed, 17 insertions(+), 1 deletion(-)

diff --git a/Documentation/admin-guide/mm/numa_memory_policy.rst b/Documentation/admin-guide/mm/numa_memory_policy.rst
index 90ab26e805a9a..f10a8297fb178 100644
--- a/Documentation/admin-guide/mm/numa_memory_policy.rst
+++ b/Documentation/admin-guide/mm/numa_memory_policy.rst
@@ -259,6 +259,16 @@ MPOL_WEIGHTED_INTERLEAVE
 	weight.  For example if nodes [0,1] are weighted [5,2], 5 pages
 	will be allocated on node0 for every 2 pages allocated on node1.
 
+	When MPOL_F_NUMA_BALANCING is specified, migrate-on-fault
+	placement is allowed within the policy nodemask.  This is an
+	opt-in behavior; without MPOL_F_NUMA_BALANCING,
+	MPOL_WEIGHTED_INTERLEAVE keeps its existing behavior and does
+	not participate in NUMA balancing.  The flag can be used by
+	memory tiering to promote hot pages that were initially
+	allocated on slower nodes.  The weights still affect new
+	allocations only and do not define a target resident ratio after
+	page migration.
+
 NUMA memory policy supports the following optional mode flags:
 
 MPOL_F_STATIC_NODES
diff --git a/mm/mempolicy.c b/mm/mempolicy.c
index 79053ece02cd4..f71cdb488bbdc 100644
--- a/mm/mempolicy.c
+++ b/mm/mempolicy.c
@@ -1731,7 +1731,8 @@ static inline int sanitize_mpol_flags(int *mode, unsigned short *flags)
 	if ((*flags & MPOL_F_STATIC_NODES) && (*flags & MPOL_F_RELATIVE_NODES))
 		return -EINVAL;
 	if (*flags & MPOL_F_NUMA_BALANCING) {
-		if (*mode == MPOL_BIND || *mode == MPOL_PREFERRED_MANY)
+		if (*mode == MPOL_BIND || *mode == MPOL_PREFERRED_MANY ||
+		    *mode == MPOL_WEIGHTED_INTERLEAVE)
 			*flags |= (MPOL_F_MOF | MPOL_F_MORON);
 		else
 			return -EINVAL;
@@ -3005,6 +3006,11 @@ int mpol_misplaced(struct folio *folio, struct vm_fault *vmf,
 		break;
 
 	case MPOL_WEIGHTED_INTERLEAVE:
+		if (pol->flags & MPOL_F_MORON) {
+			if (node_isset(thisnid, pol->nodes))
+				break;
+			goto out;
+		}
 		polnid = weighted_interleave_nid(pol, ilx);
 		break;
 
-- 
2.20.1

             reply	other threads:[~2026-09-30  7:27 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30  7:26 Li Zhe [this message]
2026-09-30  8:43 ` [PATCH] mm/numa_balancing: allow migrate on protnone reference with MPOL_WEIGHTED_INTERLEAVE policy Gregory Price
2026-09-30  8:59   ` Li Zhe
2026-09-30  9:11     ` Gregory Price
2026-09-30 10:59       ` Joshua Hahn
2026-09-30 11:22       ` Li Zhe
2026-09-30 12:25         ` Gregory Price
2026-10-02  7:25           ` Li Zhe
2026-09-30 11:26 ` David Hildenbrand (Arm)
2026-09-30 11:52   ` Li Zhe
2026-09-30 12:11     ` Joshua Hahn
2026-09-30 14:02       ` Li Zhe
2026-09-30 14:50         ` Gregory Price
2026-09-30 15:10           ` Zi Yan
2026-10-01 10:54             ` David Hildenbrand (Arm)
2026-10-01 11:18               ` Zi Yan
2026-10-01 13:26               ` Gregory Price
2026-10-02  7:56                 ` David Hildenbrand (Arm)
2026-10-02  8:15                   ` Li Zhe
2026-10-02 12:05                     ` Gregory Price
2026-09-30 12:03   ` Gregory Price
2026-09-30 12:32     ` David Hildenbrand (Arm)
2026-09-30 13:40       ` Gregory Price
2026-10-01 10:55         ` David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260930072617.64665-1-lizhe.67@bytedance.com \
    --to=lizhe.67@bytedance.com \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=corbet@lwn.net \
    --cc=david@kernel.org \
    --cc=gourry@gourry.net \
    --cc=joshua.hahnjy@gmail.com \
    --cc=liam@infradead.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=rppt@kernel.org \
    --cc=skhan@linuxfoundation.org \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox