All of lore.kernel.org
 help / color / mirror / Atom feed
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Gregory Price <gourry@gourry.net>, linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com,
	akpm@linux-foundation.org, ljs@kernel.org, liam@infradead.org,
	vbabka@kernel.org, rppt@kernel.org, surenb@google.com,
	mhocko@suse.com, mingo@redhat.com, peterz@infradead.org,
	juri.lelli@redhat.com, vincent.guittot@linaro.org,
	dietmar.eggemann@arm.com, rostedt@goodmis.org,
	bsegall@google.com, mgorman@suse.de, vschneid@redhat.com,
	kprateek.nayak@amd.com, ziy@nvidia.com,
	baolin.wang@linux.alibaba.com, nico.pache@linux.dev,
	ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org,
	lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org,
	matthew.brost@intel.com, joshua.hahnjy@gmail.com,
	rakie.kim@sk.com, byungchul@sk.com, ying.huang@linux.alibaba.com,
	apopple@nvidia.com, jannh@google.com, pfalcato@suse.de,
	osalvador@suse.de, hannes@cmpxchg.org, raghavendra.kt@amd.com,
	stable@vger.kernel.org
Subject: Re: [PATCH v2 1/4] mm: support promotion-only NUMA hinting scans
Date: Fri, 18 Sep 2026 14:37:52 +0200	[thread overview]
Message-ID: <9b5e1573-6fd8-48a5-a274-a6e7ff83d2fe@kernel.org> (raw)
In-Reply-To: <20260911001826.2109390-2-gourry@gourry.net>

On 9/11/26 02:18, Gregory Price wrote:
> From: "Gregory Price (Meta)" <gourry@gourry.net>
> 
> folio_can_map_prot_numa() derives the eligible memory from the global
> balancing mode. Combined mode needs the scanner to distinguish placement
> scans from promotion-only scans.
> 
> Add MM_CP_PROT_NUMA_PROMO_ONLY and let change_prot_numa() callers request
> a promotion-only protection walk. Carry the choice with each walk so the
> PTE and PMD paths use the same decision.
> 
> Initially derive the value from the normal balancing mode, preserving the
> existing behavior for the following fixes.
> 
> Fixes: c574bbe91703 ("NUMA balancing: optimize page placement for memory tiering system")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
> ---
>  include/linux/mm.h  |  4 +++-
>  kernel/sched/fair.c |  7 ++++++-
>  mm/huge_memory.c    |  3 ++-
>  mm/internal.h       |  5 +++--
>  mm/mempolicy.c      | 22 +++++++++++++---------
>  mm/mprotect.c       |  4 +++-
>  6 files changed, 30 insertions(+), 15 deletions(-)
> 
> diff --git a/include/linux/mm.h b/include/linux/mm.h
> index 969594074fd2..2d1b59a27629 100644
> --- a/include/linux/mm.h
> +++ b/include/linux/mm.h
> @@ -3408,6 +3408,8 @@ int get_cmdline(struct task_struct *task, char *buffer, int buflen);
>  #define  MM_CP_UFFD_RWP_RESOLVE            (1UL << 5) /* resolve rwp */
>  #define  MM_CP_UFFD_RWP_ALL                (MM_CP_UFFD_RWP | \
>  					    MM_CP_UFFD_RWP_RESOLVE)
> +/* Whether a MM_CP_PROT_NUMA change is for promotion only */
> +#define  MM_CP_PROT_NUMA_PROMO_ONLY        (1UL << 6)

BTW, shouldn't we just be using BIT()?

>  
>  bool can_change_pte_writable(struct vm_area_struct *vma, unsigned long addr,
>  			     pte_t pte);
> @@ -4736,7 +4738,7 @@ void vma_set_file(struct vm_area_struct *vma, struct file *file);
>  
>  #ifdef CONFIG_NUMA_BALANCING
>  unsigned long change_prot_numa(struct vm_area_struct *vma,
> -			unsigned long start, unsigned long end);
> +			unsigned long start, unsigned long end, bool promo_only);

While at it use two tabs.

>  #endif
>  
>  struct vm_area_struct *find_extend_vma_locked(struct mm_struct *,
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 8dff37059faf..81359b414947 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -4129,8 +4129,10 @@ static void task_numa_work(struct callback_head *work)
>  	unsigned long nr_pte_updates = 0;
>  	long pages, virtpages;
>  	struct vma_iterator vmi;
> +	unsigned int numab_mode = READ_ONCE(sysctl_numa_balancing_mode);

const and all the way to the top?

>  	bool vma_pids_skipped;
>  	bool vma_pids_forced = false;
> +	bool promo_only;
>  
>  	WARN_ON_ONCE(p != container_of(work, struct task_struct, numa_work));
>  
> @@ -4304,11 +4306,14 @@ static void task_numa_work(struct callback_head *work)
>  			continue;
>  		}
>  
> +		promo_only = !(numab_mode & NUMA_BALANCING_NORMAL);

bool promo_only = !(numab_mode & NUMA_BALANCING_NORMAL);


and in the later patch

if (vma_is_ro_file(vma))
	promo_only = true;

?

But ...

> +
>  		do {
>  			start = max(start, vma->vm_start);
>  			end = ALIGN(start + (pages << PAGE_SHIFT), HPAGE_SIZE);
>  			end = min(end, vma->vm_end);
> -			nr_pte_updates = change_prot_numa(vma, start, end);
> +			nr_pte_updates = change_prot_numa(vma, start, end,
> +							  promo_only);


Stupid question: why can't/shouldn't change_prot_numa() query
sysctl_numa_balancing_mode? Why do we have to query this outside of the function
and forward it?

I mean, change_prot_numa() gets the vma and can query
sysctl_numa_balancing_mode. Why not move that into the function and avoid the
boolean parameter?

We do have a single change_prot_numa() caller in the tree ...

>  
>  			/*
>  			 * Try to scan sysctl_numa_balancing_size worth of
> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
> index 30b7c63b0e35..2f9ada1bbcfc 100644
> --- a/mm/huge_memory.c
> +++ b/mm/huge_memory.c
> @@ -2784,7 +2784,8 @@ int change_huge_pmd(struct mmu_gather *tlb, struct vm_area_struct *vma,
>  			goto unlock;
>  
>  		if (!folio_can_map_prot_numa(pmd_folio(*pmd), vma,
> -					     vma_is_single_threaded_private(vma)))
> +					     vma_is_single_threaded_private(vma),
> +					     cp_flags & MM_CP_PROT_NUMA_PROMO_ONLY))

As raised, maybe just forward cp_flags



-- 
Cheers,

David


  parent reply	other threads:[~2026-09-18 12:38 UTC|newest]

Thread overview: 37+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-11  0:18 [PATCH v2 0/4] sched/numa: stop VMA scan filters from gating promotion Gregory Price
2026-09-11  0:18 ` [PATCH v2 1/4] mm: support promotion-only NUMA hinting scans Gregory Price
2026-09-17 16:03   ` Peter Zijlstra
2026-09-17 16:14     ` Gregory Price
2026-09-18 12:26       ` David Hildenbrand (Arm)
2026-09-18 12:37   ` David Hildenbrand (Arm) [this message]
2026-09-18 13:46     ` Gregory Price
2026-09-18 13:56       ` David Hildenbrand (Arm)
2026-09-11  0:18 ` [PATCH v2 2/4] mm: allow shared folios to be promoted to a fast tier Gregory Price
2026-09-17 16:08   ` Peter Zijlstra
2026-09-17 16:18     ` Gregory Price
2026-09-17 16:23       ` Peter Zijlstra
2026-09-17 16:39         ` Gregory Price
2026-09-18  4:14         ` Bharata B Rao
2026-09-17 17:49   ` Zi Yan
2026-09-18 12:54     ` David Hildenbrand (Arm)
2026-09-18 12:53   ` David Hildenbrand (Arm)
2026-09-18 13:54     ` Gregory Price
2026-09-18 13:57       ` David Hildenbrand (Arm)
2026-09-11  0:18 ` [PATCH v2 3/4] sched/numa: scan read-only file mappings in tiering mode Gregory Price
2026-09-18 12:58   ` David Hildenbrand (Arm)
2026-09-18 13:57     ` Gregory Price
2026-09-18 13:59       ` David Hildenbrand (Arm)
2026-09-18 14:53         ` Lorenzo Stoakes (ARM)
2026-09-18 15:48           ` Gregory Price
2026-09-18 16:19             ` Lorenzo Stoakes (ARM)
2026-09-18 16:38               ` Gregory Price
2026-09-11  0:18 ` [PATCH v2 4/4] sched/numa: do not let VMA PID activity gate promotion Gregory Price
2026-09-17 16:19   ` Peter Zijlstra
2026-09-18 13:01   ` David Hildenbrand (Arm)
2026-09-18 13:59     ` Gregory Price
2026-09-11  5:38 ` [PATCH v2 0/4] sched/numa: stop VMA scan filters from gating promotion Gregory Price
2026-09-17  5:35 ` Andrew Morton
2026-09-17  6:59   ` Gregory Price
2026-09-17 15:53     ` David Hildenbrand (Arm)
2026-09-18 20:56 ` Zi Yan
2026-09-18 21:42   ` Gregory Price

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=9b5e1573-6fd8-48a5-a274-a6e7ff83d2fe@kernel.org \
    --to=david@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=bsegall@google.com \
    --cc=byungchul@sk.com \
    --cc=dev.jain@arm.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=gourry@gourry.net \
    --cc=hannes@cmpxchg.org \
    --cc=jannh@google.com \
    --cc=joshua.hahnjy@gmail.com \
    --cc=juri.lelli@redhat.com \
    --cc=kas@kernel.org \
    --cc=kernel-team@meta.com \
    --cc=kprateek.nayak@amd.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=matthew.brost@intel.com \
    --cc=mgorman@suse.de \
    --cc=mhocko@suse.com \
    --cc=mingo@redhat.com \
    --cc=nico.pache@linux.dev \
    --cc=osalvador@suse.de \
    --cc=peterz@infradead.org \
    --cc=pfalcato@suse.de \
    --cc=raghavendra.kt@amd.com \
    --cc=rakie.kim@sk.com \
    --cc=rostedt@goodmis.org \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=stable@vger.kernel.org \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.