All of lore.kernel.org
 help / color / mirror / Atom feed
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Gregory Price <gourry@gourry.net>, linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com,
	akpm@linux-foundation.org, ljs@kernel.org, liam@infradead.org,
	vbabka@kernel.org, rppt@kernel.org, surenb@google.com,
	mhocko@suse.com, mingo@redhat.com, peterz@infradead.org,
	juri.lelli@redhat.com, vincent.guittot@linaro.org,
	dietmar.eggemann@arm.com, rostedt@goodmis.org,
	bsegall@google.com, mgorman@suse.de, vschneid@redhat.com,
	kprateek.nayak@amd.com, ziy@nvidia.com,
	baolin.wang@linux.alibaba.com, nico.pache@linux.dev,
	ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org,
	lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org,
	matthew.brost@intel.com, joshua.hahnjy@gmail.com,
	rakie.kim@sk.com, byungchul@sk.com, ying.huang@linux.alibaba.com,
	apopple@nvidia.com, jannh@google.com, pfalcato@suse.de,
	hannes@cmpxchg.org, shy828301@gmail.com, raghavendra.kt@amd.com,
	stable@vger.kernel.org
Subject: Re: [PATCH v3 5/7] sched/numa: scan PID-inactive VMAs for promotion
Date: Fri, 25 Sep 2026 12:55:09 +0200	[thread overview]
Message-ID: <ad99014e-3714-45a5-ad23-aa7fe72745cb@kernel.org> (raw)
In-Reply-To: <20260922182928.2199090-6-gourry@gourry.net>

On 9/22/26 20:29, Gregory Price wrote:
> From: "Gregory Price (Meta)" <gourry@gourry.net>
> 
> Commit fc137c0ddab2 ("sched/numa: enhance vma scanning logic") skips
> VMAs without recent PID activity. Since only NUMA hint faults record
> that activity, the filter can suppress the fault needed to promote hot
> slow-tier memory.
> 
> Let memory tiering bypass the PID scan filter. In combined mode, use the
> placement decision to restrict top-tier sampling to VMAs that need it,
> while inactive VMAs still receive promotion-only scans.
> 
> Track the last completed placement scan separately from scans of any
> kind. Promotion-only scans still update prev_scan_seq, but do not
> advance the placement-starvation horizon.
> 
> On a host with 768 GB of DRAM and 256 GB of CXL memory, one large shmem
> VMA consumed most scanning activity, while 2,537 other VMAs covering
> 84 GB were skipped as inactive. One stand-out result: a hot 20 GB hash
> table ended up trapped entirely on CXL and drove CXL bandwidth
> utilization beyond sustainable levels - resulting in a large regression.
> 
> With this series, the hot hash table ends up split evenly between DRAM
> and CXL, tier residency tracked runtime load, and CXL bandwidth
> utilization drops from 45GB/s (maxed) to 5-10GB/s, while DRAM bandwidth
> utilization increases from ~200GB/s to 250GB/s+, resulting in major
> throughput improvements for the database workload.
> 
> Fixes: fc137c0ddab2 ("sched/numa: enhance vma scanning logic")
> Cc: stable@vger.kernel.org
> Assisted-by: OpenAI:gpt-5
> Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
> ---
>  include/linux/mm_types.h |  7 +++++++
>  kernel/sched/fair.c      | 26 +++++++++++++++++++++-----
>  2 files changed, 28 insertions(+), 5 deletions(-)
> 
> diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h
> index fd35db969bc94..dcca3ead9db59 100644
> --- a/include/linux/mm_types.h
> +++ b/include/linux/mm_types.h
> @@ -804,6 +804,13 @@ struct vma_numab_state {
>  	 */
>  	int prev_scan_seq;
>  
> +	/*
> +	 * MM scan sequence ID when the VMA was last scanned for placement.
> +	 * The starvation horizon in vma_needs_placement_scan() counts against
> +	 * this, so promotion-only scans cannot postpone placement indefinitely.
> +	 */
> +	int prev_placement_scan_seq;
> +
>  	/* Preserve placement-scan eligibility during an in-progress scan. */
>  	bool placement_scan;
>  };
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index a2849e72c4e26..412c72084a63d 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -4097,7 +4097,7 @@ static bool vma_needs_placement_scan(struct mm_struct *mm,
>  	 * threads can help scan this vma, force a vma scan.
>  	 */
>  	if (READ_ONCE(mm->numa_scan_seq) >
> -	   (vma->numab_state->prev_scan_seq + get_nr_threads(current)))
> +	   (vma->numab_state->prev_placement_scan_seq + get_nr_threads(current)))
>  		return true;
>  
>  	return false;
> @@ -4267,7 +4267,8 @@ static void task_numa_work(struct callback_head *work)
>  			 * to prevent VMAs being skipped prematurely on the
>  			 * first scan:
>  			 */
> -			 vma->numab_state->prev_scan_seq = mm->numa_scan_seq - 1;
> +			vma->numab_state->prev_scan_seq = mm->numa_scan_seq - 1;
> +			vma->numab_state->prev_placement_scan_seq = mm->numa_scan_seq - 1;
>  		}
>  
>  		/*
> @@ -4300,10 +4301,13 @@ static void task_numa_work(struct callback_head *work)
>  		 * Do not scan the VMA if a task has not accessed it, unless no other
>  		 * VMA candidate exists. If a scan is already in-progress, finish it,
>  		 * but track continuation separately from starting a new one.
> +		 *
> +		 * The PID filter must not gate promotion. Allow PID-inactive VMAs
> +		 * to proceed when memory tiering is enabled.
>  		 */
>  		placement_due = vma_needs_placement_scan(mm, vma);
>  		scan_started = mm->numa_scan_offset > vma->vm_start;
> -		pid_scan_allowed = vma_pids_forced || placement_due;
> +		pid_scan_allowed = tiering || vma_pids_forced || placement_due;
>  
>  		if (!pid_scan_allowed) {
>  			if (scan_started) {
> @@ -4315,10 +4319,16 @@ static void task_numa_work(struct callback_head *work)
>  			}
>  		}
>  
> -		/* Keep scan policy stable while processing a VMA in chunks.*/
> +		/*
> +		 * Keep scan policy stable while processing a VMA in chunks.
> +		 * A fault in one chunk can make a VMA placement-eligible. Keep a
> +		 * promotion-only decision sticky for the rest of a partial scan.
> +		 */
>  		placement_scan &= numab_mode & NUMA_BALANCING_NORMAL;
>  		if (scan_started)
>  			placement_scan &= vma->numab_state->placement_scan;
> +		else if (tiering)
> +			placement_scan &= placement_due;

Same comment. Apart from that LGTM.

-- 
Cheers,

David


  reply	other threads:[~2026-09-25 10:55 UTC|newest]

Thread overview: 22+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 18:29 [PATCH v3 0/7] sched/numa: stop VMA scan filters from gating promotion Gregory Price
2026-09-22 18:29 ` [PATCH v3 1/7] mm: support promotion-only NUMA hinting scans Gregory Price
2026-09-24 11:45   ` David Hildenbrand (Arm)
2026-09-24 14:08     ` Gregory Price
2026-09-22 18:29 ` [PATCH v3 2/7] mm: allow shared folios to be promoted to a fast tier Gregory Price
2026-09-24 11:59   ` David Hildenbrand (Arm)
2026-09-24 14:07     ` Gregory Price
2026-09-24 15:32       ` David Hildenbrand (Arm)
2026-09-24 15:37         ` Gregory Price
2026-09-24 15:42           ` Zi Yan
2026-09-22 18:29 ` [PATCH v3 3/7] sched/numa: scan read-only file mappings in tiering mode Gregory Price
2026-09-24 20:53   ` David Hildenbrand (Arm)
2026-09-25  0:35     ` Gregory Price
2026-09-25 10:04       ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 4/7] sched/numa: separate VMA placement from scan continuation Gregory Price
2026-09-25 10:53   ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 5/7] sched/numa: scan PID-inactive VMAs for promotion Gregory Price
2026-09-25 10:55   ` David Hildenbrand (Arm) [this message]
2026-09-22 18:29 ` [PATCH v3 6/7] mm: use BIT() for change_protection() flags Gregory Price
2026-09-24 20:38   ` David Hildenbrand (Arm)
2026-09-22 18:29 ` [PATCH v3 7/7] mm: use VMA flag helpers in NUMA balancing Gregory Price
2026-09-24 20:39   ` David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ad99014e-3714-45a5-ad23-aa7fe72745cb@kernel.org \
    --to=david@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=bsegall@google.com \
    --cc=byungchul@sk.com \
    --cc=dev.jain@arm.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=gourry@gourry.net \
    --cc=hannes@cmpxchg.org \
    --cc=jannh@google.com \
    --cc=joshua.hahnjy@gmail.com \
    --cc=juri.lelli@redhat.com \
    --cc=kas@kernel.org \
    --cc=kernel-team@meta.com \
    --cc=kprateek.nayak@amd.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=matthew.brost@intel.com \
    --cc=mgorman@suse.de \
    --cc=mhocko@suse.com \
    --cc=mingo@redhat.com \
    --cc=nico.pache@linux.dev \
    --cc=peterz@infradead.org \
    --cc=pfalcato@suse.de \
    --cc=raghavendra.kt@amd.com \
    --cc=rakie.kim@sk.com \
    --cc=rostedt@goodmis.org \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=shy828301@gmail.com \
    --cc=stable@vger.kernel.org \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.