From: Peter Zijlstra <peterz@infradead.org>
To: Gregory Price <gourry@gourry.net>
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org,
kernel-team@meta.com, akpm@linux-foundation.org,
david@kernel.org, ljs@kernel.org, liam@infradead.org,
vbabka@kernel.org, rppt@kernel.org, surenb@google.com,
mhocko@suse.com, mingo@redhat.com, juri.lelli@redhat.com,
vincent.guittot@linaro.org, dietmar.eggemann@arm.com,
rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de,
vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com,
baolin.wang@linux.alibaba.com, nico.pache@linux.dev,
ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org,
lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org,
matthew.brost@intel.com, joshua.hahnjy@gmail.com,
rakie.kim@sk.com, byungchul@sk.com, ying.huang@linux.alibaba.com,
apopple@nvidia.com, jannh@google.com, pfalcato@suse.de,
osalvador@suse.de, hannes@cmpxchg.org, raghavendra.kt@amd.com,
stable@vger.kernel.org
Subject: Re: [PATCH v2 4/4] sched/numa: do not let VMA PID activity gate promotion
Date: Thu, 17 Sep 2026 18:19:57 +0200 [thread overview]
Message-ID: <20260917161957.GO4121339@noisy.programming.kicks-ass.net> (raw)
In-Reply-To: <20260911001826.2109390-5-gourry@gourry.net>
On Thu, Sep 10, 2026 at 08:18:26PM -0400, Gregory Price wrote:
> From: "Gregory Price (Meta)" <gourry@gourry.net>
>
> Commit fc137c0ddab2 ("sched/numa: enhance vma scanning logic")
> skips VMAs without recent PID activity. Since only NUMA hint faults record
> that activity, the filter can suppress the fault needed to promote hot
> slow-tier memory.
>
> In tiering mode, scan PID-inactive VMAs using promotion-only scans.
> Reevaluate this choice whenever the scanner visits a VMA and pass it with
> each protection walk.
>
> Record socket-placement scans in prev_placement_scan_seq. Promotion-only
> scans still update prev_scan_seq, but no longer postpone the starvation
> fallback for placement scans.
>
> On a host with 768 GB of DRAM and 256 GB of CXL memory running two roughly
> 430 GB database workloads, a large shmem VMA occupied each scan while 2,537
> other VMAs covering 84 GB were skipped as inactive. A hot 20 GB hash table
> remained entirely on CXL before this change and was split evenly between
> DRAM and CXL afterwards.
>
> Fixes: fc137c0ddab2 ("sched/numa: enhance vma scanning logic")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
> ---
> include/linux/mm_types.h | 7 ++++++
> kernel/sched/fair.c | 48 ++++++++++++++++++++++++++--------------
> 2 files changed, 39 insertions(+), 16 deletions(-)
>
> diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h
> index 5413bd10fff2..9f042d6ad465 100644
> --- a/include/linux/mm_types.h
> +++ b/include/linux/mm_types.h
> @@ -803,6 +803,13 @@ struct vma_numab_state {
> * A VMA is not eligible for scanning if prev_scan_seq == numa_scan_seq
> */
> int prev_scan_seq;
> +
> + /*
> + * MM scan sequence ID when the VMA was last scanned for placement.
> + * The starvation horizon in vma_is_accessed() counts against this, so
> + * promotion-only scans cannot postpone placement indefinitely.
> + */
> + int prev_placement_scan_seq;
> };
>
> #ifdef __HAVE_PFNMAP_TRACKING
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index e636e8de53f1..6d1da13a2ef5 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -4085,6 +4085,10 @@ static void reset_ptenuma_scan(struct task_struct *p)
> p->mm->numa_scan_offset = 0;
> }
>
> +/*
> + * Decide whether this VMA should be sampled for NUMA placement. In addition
> + * to recent accesses, periodically allow a scan to avoid starvation.
> + */
> static bool vma_is_accessed(struct mm_struct *mm, struct vm_area_struct *vma)
> {
> unsigned long pids;
> @@ -4101,22 +4105,13 @@ static bool vma_is_accessed(struct mm_struct *mm, struct vm_area_struct *vma)
> if (test_bit(hash_32(current->pid, ilog2(BITS_PER_LONG)), &pids))
> return true;
>
> - /*
> - * Complete a scan that has already started regardless of PID access, or
> - * some VMAs may never be scanned in multi-threaded applications:
> - */
> - if (mm->numa_scan_offset > vma->vm_start) {
> - trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_IGNORE_PID);
> - return true;
> - }
> -
> /*
> * This vma has not been accessed for a while, and if the number
> * the threads in the same process is low, which means no other
> * threads can help scan this vma, force a vma scan.
> */
> if (READ_ONCE(mm->numa_scan_seq) >
> - (vma->numab_state->prev_scan_seq + get_nr_threads(current)))
> + (vma->numab_state->prev_placement_scan_seq + get_nr_threads(current)))
> return true;
>
> return false;
> @@ -4142,7 +4137,7 @@ static void task_numa_work(struct callback_head *work)
> unsigned int numab_mode = READ_ONCE(sysctl_numa_balancing_mode);
> bool vma_pids_skipped;
> bool vma_pids_forced = false;
> - bool promo_only;
> + bool accessed, scan_started, promo_only;
>
> WARN_ON_ONCE(p != container_of(work, struct task_struct, numa_work));
>
> @@ -4278,6 +4273,7 @@ static void task_numa_work(struct callback_head *work)
> * first scan:
> */
> vma->numab_state->prev_scan_seq = mm->numa_scan_seq - 1;
> + vma->numab_state->prev_placement_scan_seq = mm->numa_scan_seq - 1;
> }
>
> /*
> @@ -4307,17 +4303,30 @@ static void task_numa_work(struct callback_head *work)
> }
>
> /*
> - * Do not scan the VMA if task has not accessed it, unless no other
> - * VMA candidate exists.
> + * The PID filter must not gate promotion. Scan PID-inactive
> + * VMAs in tiering mode using promotion-only scans.
> */
> - if (!vma_pids_forced && !vma_is_accessed(mm, vma)) {
> + accessed = vma_is_accessed(mm, vma);
> + scan_started = mm->numa_scan_offset > vma->vm_start;
> +
> + if (!vma_pids_forced && !accessed && !scan_started &&
> + !(numab_mode & NUMA_BALANCING_MEMORY_TIERING)) {
> vma_pids_skipped = true;
> trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_PID_INACTIVE);
> continue;
> }
> + if (!vma_pids_forced && !accessed &&
> + !(numab_mode & NUMA_BALANCING_MEMORY_TIERING) &&
> + scan_started)
> + trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_IGNORE_PID);
>
> + /*
> + * In combined mode, only VMAs the task uses need placement
> + * samples. Without tiering, every VMA reaching here does.
> + */
> promo_only = !(numab_mode & NUMA_BALANCING_NORMAL) ||
> - vma_is_ro_file(vma);
> + ((numab_mode & NUMA_BALANCING_MEMORY_TIERING) &&
> + !accessed) || vma_is_ro_file(vma);
>
> do {
> start = max(start, vma->vm_start);
> @@ -4345,8 +4354,15 @@ static void task_numa_work(struct callback_head *work)
> cond_resched();
> } while (end != vma->vm_end);
>
> - /* VMA scan is complete, do not scan until next sequence. */
> + /*
> + * VMA scan is complete, do not scan until next sequence. A
> + * promotion-only scan reached the end of the VMA but did not
> + * sample placement, so it does not count towards the starvation
> + * horizon in vma_is_accessed().
> + */
> vma->numab_state->prev_scan_seq = mm->numa_scan_seq;
> + if (!promo_only)
> + vma->numab_state->prev_placement_scan_seq = mm->numa_scan_seq;
>
> /*
> * Only force scan within one VMA at a time, to limit the
Not a fan of what that tiering code is causing :/
But I suppose this will do; Mel?
next prev parent reply other threads:[~2026-09-17 16:20 UTC|newest]
Thread overview: 37+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-11 0:18 [PATCH v2 0/4] sched/numa: stop VMA scan filters from gating promotion Gregory Price
2026-09-11 0:18 ` [PATCH v2 1/4] mm: support promotion-only NUMA hinting scans Gregory Price
2026-09-17 16:03 ` Peter Zijlstra
2026-09-17 16:14 ` Gregory Price
2026-09-18 12:26 ` David Hildenbrand (Arm)
2026-09-18 12:37 ` David Hildenbrand (Arm)
2026-09-18 13:46 ` Gregory Price
2026-09-18 13:56 ` David Hildenbrand (Arm)
2026-09-11 0:18 ` [PATCH v2 2/4] mm: allow shared folios to be promoted to a fast tier Gregory Price
2026-09-17 16:08 ` Peter Zijlstra
2026-09-17 16:18 ` Gregory Price
2026-09-17 16:23 ` Peter Zijlstra
2026-09-17 16:39 ` Gregory Price
2026-09-18 4:14 ` Bharata B Rao
2026-09-17 17:49 ` Zi Yan
2026-09-18 12:54 ` David Hildenbrand (Arm)
2026-09-18 12:53 ` David Hildenbrand (Arm)
2026-09-18 13:54 ` Gregory Price
2026-09-18 13:57 ` David Hildenbrand (Arm)
2026-09-11 0:18 ` [PATCH v2 3/4] sched/numa: scan read-only file mappings in tiering mode Gregory Price
2026-09-18 12:58 ` David Hildenbrand (Arm)
2026-09-18 13:57 ` Gregory Price
2026-09-18 13:59 ` David Hildenbrand (Arm)
2026-09-18 14:53 ` Lorenzo Stoakes (ARM)
2026-09-18 15:48 ` Gregory Price
2026-09-18 16:19 ` Lorenzo Stoakes (ARM)
2026-09-18 16:38 ` Gregory Price
2026-09-11 0:18 ` [PATCH v2 4/4] sched/numa: do not let VMA PID activity gate promotion Gregory Price
2026-09-17 16:19 ` Peter Zijlstra [this message]
2026-09-18 13:01 ` David Hildenbrand (Arm)
2026-09-18 13:59 ` Gregory Price
2026-09-11 5:38 ` [PATCH v2 0/4] sched/numa: stop VMA scan filters from gating promotion Gregory Price
2026-09-17 5:35 ` Andrew Morton
2026-09-17 6:59 ` Gregory Price
2026-09-17 15:53 ` David Hildenbrand (Arm)
2026-09-18 20:56 ` Zi Yan
2026-09-18 21:42 ` Gregory Price
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260917161957.GO4121339@noisy.programming.kicks-ass.net \
--to=peterz@infradead.org \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=bsegall@google.com \
--cc=byungchul@sk.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=dietmar.eggemann@arm.com \
--cc=gourry@gourry.net \
--cc=hannes@cmpxchg.org \
--cc=jannh@google.com \
--cc=joshua.hahnjy@gmail.com \
--cc=juri.lelli@redhat.com \
--cc=kas@kernel.org \
--cc=kernel-team@meta.com \
--cc=kprateek.nayak@amd.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=matthew.brost@intel.com \
--cc=mgorman@suse.de \
--cc=mhocko@suse.com \
--cc=mingo@redhat.com \
--cc=nico.pache@linux.dev \
--cc=osalvador@suse.de \
--cc=pfalcato@suse.de \
--cc=raghavendra.kt@amd.com \
--cc=rakie.kim@sk.com \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=stable@vger.kernel.org \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=ying.huang@linux.alibaba.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.