From: Chen Yu <yu.c.chen@intel.com>
To: Jianyong Wu <wujianyong@hygon.cn>
Cc: Tim Chen <tim.c.chen@linux.intel.com>,
Peter Zijlstra <peterz@infradead.org>,
Ingo Molnar <mingo@redhat.com>,
Juri Lelli <juri.lelli@redhat.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
Shrikanth Hegde <sshegde@linux.ibm.com>,
Phil Auld <pauld@redhat.com>,
Andrew Morton <akpm@linux-foundation.org>,
"David Hildenbrand" <david@kernel.org>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
"linux-mm@kvack.org" <linux-mm@kvack.org>,
"jianyong.wu@outlook.com" <jianyong.wu@outlook.com>,
Yuan Zhong <zhongyuan@hygon.cn>, Huangsj <huangsj@hygon.cn>,
Fengyu Wang <wangfengyu@hygon.cn>,
Zhiwei Ying <yingzhiwei@hygon.cn>,
"justin.he@arm.com" <justin.he@arm.com>, <chen.yu@linux.dev>
Subject: Re: [RFC PATCH v2 02/23] sched/topology: Introduce a NUMA distance matrix with unique distance values
Date: Fri, 9 Oct 2026 00:35:25 +0800 [thread overview]
Message-ID: <asfGTaJjCGMLW6fp@chenyu-dev> (raw)
In-Reply-To: <4d2b5dede1f64350ac17ea9d0a5428a1@hygon.cn>
Hi Jianyong,
On Thu, Oct 08, 2026 at 07:35:10AM +0000, Jianyong Wu wrote:
> Hi Tim,
>
> > -----Original Message-----
> > From: Tim Chen <tim.c.chen@linux.intel.com>
> > Sent: Friday, October 2, 2026 6:04 AM
> > To: Jianyong Wu <wujianyong@hygon.cn>; Peter Zijlstra
> > <peterz@infradead.org>
> > Cc: Ingo Molnar <mingo@redhat.com>; Juri Lelli <juri.lelli@redhat.com>;
> > Vincent Guittot <vincent.guittot@linaro.org>; Chen Yu
> > <yu.c.chen@intel.com>; Dietmar Eggemann
> > <dietmar.eggemann@arm.com>; Steven Rostedt <rostedt@goodmis.org>;
> > Ben Segall <bsegall@google.com>; Mel Gorman <mgorman@suse.de>;
> > Valentin Schneider <vschneid@redhat.com>; K Prateek Nayak
> > <kprateek.nayak@amd.com>; Shrikanth Hegde <sshegde@linux.ibm.com>;
> > Phil Auld <pauld@redhat.com>; Andrew Morton
> > <akpm@linux-foundation.org>; David Hildenbrand <david@kernel.org>;
> > linux-kernel@vger.kernel.org; linux-mm@kvack.org;
> > jianyong.wu@outlook.com; Yuan Zhong <zhongyuan@hygon.cn>; Huangsj
> > <huangsj@hygon.cn>; Fengyu Wang <wangfengyu@hygon.cn>; Zhiwei Ying
> > <yingzhiwei@hygon.cn>; justin.he@arm.com
> > Subject: Re: [RFC PATCH v2 02/23] sched/topology: Introduce a NUMA
> > distance matrix with unique distance values
> >
> > On Mon, 2026-09-28 at 09:39 +0000, Jianyong Wu wrote:
> > > Hi Tim,
> > >
> > >
> > > Makes sense. So, what about the following solution?
> > > Given a node affinity sequence, a move from src to dst improves the
> > > affinity of every task whose preferred node i ranks dst better than src.
> > The score
> > > is then
> > >
> > > Di = raw_dist(src, i) - raw_dist(dst, i)
> > > affinity_bias_i = position of src minus position of dst in node i's affinity
> > > sequence, counted among the nodes at the same
> > distance from
> > > i (zero when Di is not zero)
> >
> > Do we really need an affinity bias? I think your intention is to use it for
> > breaking a tie.
> > If there is a tie in affinity score (without injecting bias),
> > just use the position diff to break the tie. Having a bias distorts
> > the affinity score.
> >
> >
> I think there is a misunderstanding about the purpose of the bias.
> It is not intended to break ties between affinity scores. I want the
> score itself to reflect opportunities for aggregation, including moves
> that leave the raw NUMA distance unchanged but follow the fixed node
> affinity order.
>
> That said, I agree with your concern about using artificial node
> distances in the score calculation. The magnitude of an artificial
> distance or rank difference should not determine the weight of those
> aggregation opportunities.
>
Thanks a lot for working on this.
I think there are two level of aggregation:
LLCs aggregation, and Node aggregation.
Your current proposal uses a best-effort strategy for multi-LLC/node
aggregation. It tries to migrate towards a preferred LLC mask, and the
size of the LLC/node mask is the "range" you calculated in task_cache_work().
Take the LLC case, for example. Suppose there are 6 LLCs, and in
task_cache_work(), the range is calculated as 3 for process P, which means
that at most 3 LLCs can hold all the threads of P.
The strategy allows load balancing to move into the preferred LLC mask/range
in random order:
suppose the util of each LLCs are:
LLC0: 45% LLC1: 40% LLC2: 30%
migrate task from LLC4 to LLC1? OK
migrate task from LLC4 to LLC0? OK
migrate task from LLC2 to LLC0? OK
migrate task from LLC0 to LLC1? no
...
My understanding is that the introduction of distance comparison is
because the NUMA nodes are not symmetric to each other, so you need a
realistic sequence to saturate the nodes. However, this might not be
true for LLCs - they are symmetric, and you can start with a pref CPU
and get the next LLC via the LLC id sequence. That is to say,
dist(LLC1, LLC2) = (rank1 + rank2) % k + 1 is not mandatory. Why not
just generate the preferred_mask by adding the "range" siblings of the
preferred LLC, and use preferred_mask directly during load balance?
That seems to be straightforward. So we can get rid of the distance
comparison (at least for LLCs).
thanks,
Chenyu
next prev parent reply other threads:[~2026-10-08 16:48 UTC|newest]
Thread overview: 72+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-27 12:27 [RFC PATCH v2 00/23] sched: Scale cache-aware aggregation at LLC granularity Jianyong Wu
2026-08-27 12:27 ` [RFC PATCH v2 01/23] sched/topology: Add llc_to_node() to translate LLC id to NUMA node Jianyong Wu
2026-08-29 10:31 ` Peter Zijlstra
2026-08-31 9:43 ` Jianyong Wu
2026-08-27 12:27 ` [RFC PATCH v2 02/23] sched/topology: Introduce a NUMA distance matrix with unique distance values Jianyong Wu
2026-08-31 11:50 ` Peter Zijlstra
2026-09-01 6:57 ` Jianyong Wu
2026-09-01 7:13 ` Peter Zijlstra
2026-09-01 7:37 ` Jianyong Wu
2026-09-22 18:38 ` Tim Chen
2026-09-23 3:14 ` Jianyong Wu
2026-09-23 18:29 ` Tim Chen
2026-09-24 5:41 ` Jianyong Wu
2026-09-24 15:46 ` Tim Chen
2026-09-28 9:39 ` Jianyong Wu
2026-10-01 22:04 ` Tim Chen
2026-10-08 7:35 ` Jianyong Wu
2026-10-08 16:35 ` Chen Yu [this message]
2026-09-01 8:48 ` Peter Zijlstra
2026-09-02 7:11 ` Jianyong Wu
2026-08-27 12:27 ` [RFC PATCH v2 03/23] sched/topology: Introduce a macro to traverse node Jianyong Wu
2026-08-27 12:27 ` [RFC PATCH v2 04/23] sched/topology: Introduce a method to calculate the llc distance Jianyong Wu
2026-08-31 13:12 ` Peter Zijlstra
2026-08-27 12:27 ` [RFC PATCH v2 05/23] sched/topology: Introduce a macro to traverse LLC inside node Jianyong Wu
2026-08-27 12:27 ` [RFC PATCH v2 06/23] sched/topology: Add sd_node for the NODE sched domain Jianyong Wu
2026-08-27 12:28 ` [RFC PATCH v2 07/23] sched/cache: Prioritize preferred NUMA node selection over LLC selection Jianyong Wu
2026-08-31 13:16 ` Peter Zijlstra
2026-09-01 7:44 ` Jianyong Wu
2026-08-31 13:22 ` Peter Zijlstra
2026-09-01 8:05 ` Jianyong Wu
2026-08-27 12:28 ` [RFC PATCH v2 08/23] sched/topology: Introduce a per-CPU tasks NUMA preferred counter Jianyong Wu
2026-08-31 13:23 ` Peter Zijlstra
2026-09-01 8:14 ` Jianyong Wu
2026-08-31 13:24 ` Peter Zijlstra
2026-09-01 8:31 ` Jianyong Wu
2026-09-01 10:21 ` Peter Zijlstra
2026-09-01 13:02 ` Jianyong Wu
2026-08-27 12:28 ` [RFC PATCH v2 09/23] sched/cache: Account percpu sd task NUMA preference Jianyong Wu
2026-09-01 7:54 ` Peter Zijlstra
2026-09-01 8:41 ` Jianyong Wu
2026-08-27 12:28 ` [RFC PATCH v2 10/23] sched/topology: Add per-sd scratch for the load balance affinity score Jianyong Wu
2026-09-01 8:02 ` Peter Zijlstra
2026-09-01 11:55 ` Jianyong Wu
2026-08-27 12:28 ` [RFC PATCH v2 11/23] sched/cache: Introduce helpers for task migration decisions Jianyong Wu
2026-09-01 9:08 ` Peter Zijlstra
2026-09-02 5:08 ` Jianyong Wu
2026-09-01 11:32 ` Peter Zijlstra
2026-09-02 5:46 ` Jianyong Wu
2026-09-02 21:11 ` Tim Chen
2026-09-03 2:04 ` Jianyong Wu
2026-08-27 12:28 ` [RFC PATCH v2 12/23] sched/cache: Introduce rq affinity gain calculation Jianyong Wu
2026-09-01 9:58 ` Peter Zijlstra
2026-09-01 12:23 ` Jianyong Wu
2026-09-01 10:16 ` Peter Zijlstra
2026-09-01 12:34 ` Jianyong Wu
2026-08-27 12:28 ` [RFC PATCH v2 13/23] sched/cache: Pick optimal src rq/group using affinity promotion metric Jianyong Wu
2026-08-27 12:28 ` [RFC PATCH v2 14/23] sched/cache: Drop prefer_sibling restriction for llc_balance Jianyong Wu
2026-09-01 10:29 ` Peter Zijlstra
2026-09-01 13:25 ` Jianyong Wu
2026-08-28 1:58 ` [RFC PATCH v2 15/23] sched/cache: Judge migration eligibility in LLC granularity Jianyong Wu
2026-08-28 2:04 ` [RFC PATCH v2 16/23] sched/cache: Allow un-throttled active balance to spread out of a full LLC Jianyong Wu
2026-08-28 2:07 ` [RFC PATCH v2 17/23] sched/fair: Fine-granularity NUMA balancing Jianyong Wu
2026-09-01 12:47 ` Peter Zijlstra
2026-09-02 6:43 ` Jianyong Wu
2026-08-28 2:09 ` [RFC PATCH v2 18/23] sched/cache: Scan all prefer nodes in thread group Jianyong Wu
2026-08-28 2:10 ` [RFC PATCH v2 19/23] sched/cache: Remove preferred LLC/node check no longer needed Jianyong Wu
2026-08-28 2:11 ` [RFC PATCH v2 20/23] sched/cache: Estimate utilization of the whole thread group Jianyong Wu
2026-09-01 14:44 ` Peter Zijlstra
2026-09-08 7:43 ` Jianyong Wu
2026-08-28 2:13 ` [RFC PATCH v2 21/23] sched/cache: Spread workloads within an estimated LLC range Jianyong Wu
2026-08-28 2:14 ` [RFC PATCH v2 22/23] sched/cache: Walk the preferred node from the preferred LLC Jianyong Wu
2026-08-28 2:15 ` [RFC PATCH v2 23/23] sched/debug: Print task preferred LLC for scheduler debugging Jianyong Wu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asfGTaJjCGMLW6fp@chenyu-dev \
--to=yu.c.chen@intel.com \
--cc=akpm@linux-foundation.org \
--cc=bsegall@google.com \
--cc=chen.yu@linux.dev \
--cc=david@kernel.org \
--cc=dietmar.eggemann@arm.com \
--cc=huangsj@hygon.cn \
--cc=jianyong.wu@outlook.com \
--cc=juri.lelli@redhat.com \
--cc=justin.he@arm.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=pauld@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=sshegde@linux.ibm.com \
--cc=tim.c.chen@linux.intel.com \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
--cc=wangfengyu@hygon.cn \
--cc=wujianyong@hygon.cn \
--cc=yingzhiwei@hygon.cn \
--cc=zhongyuan@hygon.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.