From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Gregory Price <gourry@gourry.net>, linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com,
akpm@linux-foundation.org, ljs@kernel.org, liam@infradead.org,
vbabka@kernel.org, rppt@kernel.org, surenb@google.com,
mhocko@suse.com, mingo@redhat.com, peterz@infradead.org,
juri.lelli@redhat.com, vincent.guittot@linaro.org,
dietmar.eggemann@arm.com, rostedt@goodmis.org,
bsegall@google.com, mgorman@suse.de, vschneid@redhat.com,
kprateek.nayak@amd.com, ziy@nvidia.com,
baolin.wang@linux.alibaba.com, nico.pache@linux.dev,
ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org,
lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org,
matthew.brost@intel.com, joshua.hahnjy@gmail.com,
rakie.kim@sk.com, byungchul@sk.com, ying.huang@linux.alibaba.com,
apopple@nvidia.com, jannh@google.com, pfalcato@suse.de,
osalvador@suse.de, hannes@cmpxchg.org, raghavendra.kt@amd.com,
stable@vger.kernel.org
Subject: Re: [PATCH v2 3/4] sched/numa: scan read-only file mappings in tiering mode
Date: Fri, 18 Sep 2026 14:58:36 +0200 [thread overview]
Message-ID: <0ed3ab3a-80b4-492f-867a-0584441722a9@kernel.org> (raw)
In-Reply-To: <20260911001826.2109390-4-gourry@gourry.net>
On 9/11/26 02:18, Gregory Price wrote:
> From: "Gregory Price (Meta)" <gourry@gourry.net>
>
> Commit 4591ce4f2d22 ("sched/numa: Do not trap hinting faults for
> shared libraries") excludes read-only file mappings from NUMA hinting to
> prevent placement bouncing. This also hides hot file folios on slow memory
> from the tiering code.
>
> Scan these mappings when memory tiering is enabled, but make their scans
> promotion-only. Ordinary NUMA placement retains the existing restriction.
>
> On a host with 768 GB of DRAM and 256 GB of CXL memory running a roughly
> 430 GB database service, 169 MB of its 185 MB main binary accumulated on
> CXL before this change. Afterwards its tier residency tracked runtime load.
>
> Fixes: c574bbe91703 ("NUMA balancing: optimize page placement for memory tiering system")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
> ---
> kernel/sched/fair.c | 23 +++++++++++++++++------
> 1 file changed, 17 insertions(+), 6 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 81359b414947..e636e8de53f1 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -4061,6 +4061,16 @@ void task_numa_fault(int last_cpupid, int mem_node, int pages, int flags)
> p->numa_faults_locality[local] += pages;
> }
>
> +/*
> + * Read-only file-backed mappings are expected to be cache replicated between
> + * accessor nodes, so they are not worth sampling for placement. They can
> + * still strand on the slow tier like anything else.
> + */
> +static bool vma_is_ro_file(struct vm_area_struct *vma)
> +{
> + return vma->vm_file && (vma->vm_flags & (VM_READ | VM_WRITE)) == VM_READ;
MAP_PRIVATE can easily map a read-only file with write permissions. So the
function name is a bit misleading.
This smells like a helper that should go next to other vma helpers and have
clear semantics.
> +}
> +
> static void reset_ptenuma_scan(struct task_struct *p)
> {
> /*
> @@ -4220,13 +4230,13 @@ static void task_numa_work(struct callback_head *work)
> }
>
> /*
> - * Shared library pages mapped by multiple processes are not
> - * migrated as it is expected they are cache replicated. Avoid
> - * hinting faults in read-only file-backed mappings or the vDSO
> - * as migrating the pages will be of marginal benefit.
> + * Read-only file-backed folios are poor NUMA placement
> + * candidates, but slow-tier folios still need to be scanned for
> + * promotion.
> */
> if (!vma->vm_mm ||
> - (vma->vm_file && (vma->vm_flags & (VM_READ|VM_WRITE)) == (VM_READ))) {
> + (vma_is_ro_file(vma) &&
> + !(numab_mode & NUMA_BALANCING_MEMORY_TIERING))) {
I'd vote for >80c here and but it into a singe line.
Or just use a magical helper
const bool tiering = numab_mode & NUMA_BALANCING_MEMORY_TIERING;
or sth like that.
> trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_SHARED_RO);
> continue;
> }
> @@ -4306,7 +4316,8 @@ static void task_numa_work(struct callback_head *work)
> continue;
> }
>
> - promo_only = !(numab_mode & NUMA_BALANCING_NORMAL);
> + promo_only = !(numab_mode & NUMA_BALANCING_NORMAL) ||
> + vma_is_ro_file(vma);
As I said, maybe that flag could be voided.
--
Cheers,
David
next prev parent reply other threads:[~2026-09-18 12:58 UTC|newest]
Thread overview: 37+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-11 0:18 [PATCH v2 0/4] sched/numa: stop VMA scan filters from gating promotion Gregory Price
2026-09-11 0:18 ` [PATCH v2 1/4] mm: support promotion-only NUMA hinting scans Gregory Price
2026-09-17 16:03 ` Peter Zijlstra
2026-09-17 16:14 ` Gregory Price
2026-09-18 12:26 ` David Hildenbrand (Arm)
2026-09-18 12:37 ` David Hildenbrand (Arm)
2026-09-18 13:46 ` Gregory Price
2026-09-18 13:56 ` David Hildenbrand (Arm)
2026-09-11 0:18 ` [PATCH v2 2/4] mm: allow shared folios to be promoted to a fast tier Gregory Price
2026-09-17 16:08 ` Peter Zijlstra
2026-09-17 16:18 ` Gregory Price
2026-09-17 16:23 ` Peter Zijlstra
2026-09-17 16:39 ` Gregory Price
2026-09-18 4:14 ` Bharata B Rao
2026-09-17 17:49 ` Zi Yan
2026-09-18 12:54 ` David Hildenbrand (Arm)
2026-09-18 12:53 ` David Hildenbrand (Arm)
2026-09-18 13:54 ` Gregory Price
2026-09-18 13:57 ` David Hildenbrand (Arm)
2026-09-11 0:18 ` [PATCH v2 3/4] sched/numa: scan read-only file mappings in tiering mode Gregory Price
2026-09-18 12:58 ` David Hildenbrand (Arm) [this message]
2026-09-18 13:57 ` Gregory Price
2026-09-18 13:59 ` David Hildenbrand (Arm)
2026-09-18 14:53 ` Lorenzo Stoakes (ARM)
2026-09-18 15:48 ` Gregory Price
2026-09-18 16:19 ` Lorenzo Stoakes (ARM)
2026-09-18 16:38 ` Gregory Price
2026-09-11 0:18 ` [PATCH v2 4/4] sched/numa: do not let VMA PID activity gate promotion Gregory Price
2026-09-17 16:19 ` Peter Zijlstra
2026-09-18 13:01 ` David Hildenbrand (Arm)
2026-09-18 13:59 ` Gregory Price
2026-09-11 5:38 ` [PATCH v2 0/4] sched/numa: stop VMA scan filters from gating promotion Gregory Price
2026-09-17 5:35 ` Andrew Morton
2026-09-17 6:59 ` Gregory Price
2026-09-17 15:53 ` David Hildenbrand (Arm)
2026-09-18 20:56 ` Zi Yan
2026-09-18 21:42 ` Gregory Price
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=0ed3ab3a-80b4-492f-867a-0584441722a9@kernel.org \
--to=david@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=bsegall@google.com \
--cc=byungchul@sk.com \
--cc=dev.jain@arm.com \
--cc=dietmar.eggemann@arm.com \
--cc=gourry@gourry.net \
--cc=hannes@cmpxchg.org \
--cc=jannh@google.com \
--cc=joshua.hahnjy@gmail.com \
--cc=juri.lelli@redhat.com \
--cc=kas@kernel.org \
--cc=kernel-team@meta.com \
--cc=kprateek.nayak@amd.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=matthew.brost@intel.com \
--cc=mgorman@suse.de \
--cc=mhocko@suse.com \
--cc=mingo@redhat.com \
--cc=nico.pache@linux.dev \
--cc=osalvador@suse.de \
--cc=peterz@infradead.org \
--cc=pfalcato@suse.de \
--cc=raghavendra.kt@amd.com \
--cc=rakie.kim@sk.com \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=stable@vger.kernel.org \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=ying.huang@linux.alibaba.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.