All of lore.kernel.org
 help / color / mirror / Atom feed
From: Chen Yu <chen.yu@linux.dev>
To: Peter Zijlstra <peterz@infradead.org>
Cc: Szabina Korbai <szkorbai@linux.ibm.com>,
	K Prateek Nayak <kprateek.nayak@amd.com>,
	mingo@kernel.org, longman@redhat.com, chenridong@huaweicloud.com,
	juri.lelli@redhat.com, vincent.guittot@linaro.org,
	dietmar.eggemann@arm.com, rostedt@goodmis.org,
	bsegall@google.com, mgorman@suse.de, vschneid@redhat.com,
	tj@kernel.org, hannes@cmpxchg.org, mkoutny@suse.com,
	cgroups@vger.kernel.org, linux-kernel@vger.kernel.org,
	jstultz@google.com, qyousef@layalina.io, euan@linux.ibm.com,
	huschle@linux.ibm.com, yu.c.chen@intel.com, tim.c.chen@intel.com,
	philip.li@intel.com
Subject: Re: [PATCH v3 0/7] sched: Flatten the pick
Date: Mon, 24 Aug 2026 22:51:51 +0800	[thread overview]
Message-ID: <aoxah90s0bQ4tcUW@three-body> (raw)
In-Reply-To: <20260818091649.GC1247881@noisy.programming.kicks-ass.net>

On Tue, Aug 18, 2026 at 11:16:49AM +0200, Peter Zijlstra wrote:
> > > 
> > > Are you using tip:sched/core at commit 68e3748781 ("sched/fair: Fix
> > > flat
> > > hierarchy") for the flat_cg numbers or did you checkout at
> > > 85570f10a4c6
> > > ("sched/eevdf: Move to a single runqueue")?
> > > 
> > > There are a couple fixes for vruntime update and Vincent's
> > > optimizations
> > > for preemption bits which might make a difference to the overall
> > > results.
> > 
> > 
> > Hi Prateek,
> > 
> > thank you, that's a good call. I did checkout at "Move to a single
> > runqueue". Let me try it with the fix included, see how the results are
> > affected.
> 
> I've not yet managed to digest your various benchmark results, but also
> double check that patch 6/7 from this series is not to 'blame' for the
> some of the changes.
> 
> The 0day robot fingered that patch for at least one issue.
> 
> In that case the benchmark threads ended up 'heavier' than before, which
> resulted in less preemptions. Probably ksoftirqd getting ran less and
> causing a regression in network throughput for that thing.
> 

The 0day' regression was triggered by running netperf/netserver using loopback,
so ksoftirqd was not really running very frequently IMO.

> I did suggest trying to change the slice of ksoftirqd down, such that it
> might be ran more readily, but I'm not sure that ever got tried.

The 0day regression was triggered by running netperf/netserver over loopback,
so ksoftirqd was not running very frequently IIUC.

I borrowed a similar machine(no-SMT, every 4 Cores share L2, 192 Cores) as 0day used.
I asked AI to generate a simple test script based on 0day's reproducer
(attached at the end of this email), and successfully reproduced the regression.

TL;DR,
The decision made by wake_affine_weight() seems to be affected by the increased
weight of the wakee under concur mode. As a result, wake_affine_weight() became less
likely to select this_cpu. This reduced preference for this_cpu leads to a higher
L2 miss rate on this platform, and consequently lower throughput.

Test summary:

| Config | Throughput (Mbps) | **L2 miss%** | **LLC miss%** |
| smp + WA_WEIGHT | **470249** | **19.37%** | 60.08% |
| smp + NO_WA_WEIGHT | 305360 | 32.67% | 60.46% |
| concur + WA_WEIGHT | 338770 | 31.45% | 62.85% |
| concur + NO_WA_WEIGHT | 259358 | 36.24% | 61.95% |

We can see if WA_WEIGHT is disabled in smp mode, the performance drops to concur. One
possible reason is that task_h_load(p) returns a very small value in smp mode(just the
issue that concur wants to fix), whereas it returns a "normal" value in concur mode.

In smp mode, the extremly small value returned by task_h_load(p) can be effectively ignored
in the comparison in wake_affine_weight(). As a result, the condition for returning this_cpu
becomes:

cpu_load(cpu_rq(this_cpu)) * 100 < cpu_load(cpu_rq(prev_cpu) * 108

As a result, this_cpu has a higher chance of being selected in smp mode. In concur mode,
task_h_load(p) is on par with cpu_load(), so it cannot be simply ignored and provides a fairer
basis for choosing between this_cpu and prev_cpu. Since netperf/netserver is a sync wakeup, it
prefers to be co-located on the same CPU/L2 domain for better L2 cache locality.

I'm not sure if this is the expected behavior before I further digging into the details,
a hack patch could restore part of the performance, because it increases the ratio to
return this_cpu:

thanks,
Chenyu

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 6d881e530f89..8bff8b6698a9 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -8385,6 +8385,7 @@ wake_affine_weight(struct sched_domain *sd, struct task_struct *p,
        unsigned long task_load;
 
        this_eff_load = cpu_load(cpu_rq(this_cpu));
+       task_load = task_h_load(p);
 
        if (sync) {
                unsigned long current_load = task_h_load(current);
@@ -8392,11 +8393,13 @@ wake_affine_weight(struct sched_domain *sd, struct task_struct *p,
                if (current_load > this_eff_load)
                        return this_cpu;
 
+               /* replace current load with wakee load does no harm */
+               if (sched_feat(WA_SYNC_WEIGHT) && current_load >= task_load)
+                       return this_cpu;
+
                this_eff_load -= current_load;
        }
 
-       task_load = task_h_load(p);
-
        this_eff_load += task_load;
        if (sched_feat(WA_BIAS))
                this_eff_load *= 100;
diff --git a/kernel/sched/features.h b/kernel/sched/features.h
index 8f0dee8fc475..99fabe53fb22 100644
--- a/kernel/sched/features.h
+++ b/kernel/sched/features.h
@@ -129,6 +129,7 @@ SCHED_FEAT(ATTACH_AGE_LOAD, true)
 SCHED_FEAT(WA_IDLE, true)
 SCHED_FEAT(WA_WEIGHT, true)
 SCHED_FEAT(WA_BIAS, true)
+SCHED_FEAT(WA_SYNC_WEIGHT, true)


test command:

NR=$(( $(nproc) * 2 ))
run() {
local mode=$1 waw=$2 tag=$3
echo $mode | sudo -n tee /sys/kernel/debug/sched/cgroup_mode >/dev/null
echo $waw  | sudo -n tee /sys/kernel/debug/sched/features    >/dev/null
sudo -n pkill -x netserver; sleep 1
netserver -4 >/dev/null 2>&1; sleep 1
rm -f /tmp/m4.$tag
for i in $(seq $NR); do
netperf -4 -H 127.0.0.1 -t TCP_MAERTS -l 70 -P 0 >> /tmp/m4.$tag 
2>/dev/null &
done
sleep 15
sudo -n ./tools/perf/perf stat -a \
-e 
l2_request.hit,l2_request.miss,longest_lat_cache.reference,longest_lat_cache.miss 
\
-x, -o /tmp/p4.$tag -- sleep 20
wait
} >/dev/null 2>&1
run smp    WA_WEIGHT    smp_waw
run smp    NO_WA_WEIGHT smp_nowaw
run concur WA_WEIGHT    con_waw
run concur NO_WA_WEIGHT con_nowaw
echo WA_WEIGHT | sudo -n tee /sys/kernel/debug/sched/features >/dev/null


printf "%-12s %10s %8s %10s %10s\n" CONFIG Mbps streams L2miss% LLCmiss%
for t in smp_waw smp_nowaw con_waw con_nowaw; do
   read TP NS < <(awk '{s+=$NF;n++} END{print s, n}' /tmp/m4.$t)
   awk -F, -v t=$t -v tp=$TP -v ns=$NS '
     $3=="l2_request.hit"{h=$1} $3=="l2_request.miss"{m=$1}
     $3=="longest_lat_cache.reference"{r=$1} 
$3=="longest_lat_cache.miss"{lm=$1}
     END{printf "%-12s %10.0f %8d %9.2f%% %9.2f%%\n", t, tp, ns, 
100*m/(h+m), 100*lm/r}' /tmp/p4.$t
done

      parent reply	other threads:[~2026-08-23 14:52 UTC|newest]

Thread overview: 33+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-06-05 12:40 [PATCH v3 0/7] sched: Flatten the pick Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 1/7] sched/fair: Add cgroup_mode switch Peter Zijlstra
2026-06-30  9:03   ` [tip: sched/core] " tip-bot2 for Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 2/7] sched/fair: Add cgroup_mode: up Peter Zijlstra
2026-06-05 15:07   ` Peter Zijlstra
2026-06-30  9:03   ` [tip: sched/core] " tip-bot2 for Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 3/7] sched/fair: Add cgroup_mode: max Peter Zijlstra
2026-06-10 15:09   ` Waiman Long
2026-06-10 15:42     ` Waiman Long
2026-06-11 13:49       ` Peter Zijlstra
2026-06-11 13:47     ` Peter Zijlstra
2026-06-11 20:57       ` Waiman Long
2026-06-30  9:03   ` [tip: sched/core] " tip-bot2 for Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 4/7] sched/fair: Add cgroup_mode: concur Peter Zijlstra
2026-06-30  9:03   ` [tip: sched/core] " tip-bot2 for Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 5/7] sched/fair: Add cgroup_mode: tasks Peter Zijlstra
2026-06-30  9:03   ` [tip: sched/core] " tip-bot2 for Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 6/7] sched/fair: Change the default cgroup_mode to concur Peter Zijlstra
2026-06-30  9:03   ` [tip: sched/core] " tip-bot2 for Peter Zijlstra
2026-06-05 12:40 ` [PATCH v3 7/7] sched/eevdf: Move to a single runqueue Peter Zijlstra
2026-06-20  3:54   ` Chen, Yu C
2026-06-26 11:40     ` Peter Zijlstra
2026-06-29 14:02       ` Vincent Guittot
2026-06-30  9:03     ` [tip: sched/core] sched/fair: Fix overflow in update_tg_cfs_runnable() tip-bot2 for Chen, Yu C
2026-06-30  9:03   ` [tip: sched/core] sched/eevdf: Move to a single runqueue tip-bot2 for Peter Zijlstra (Intel)
2026-06-09  5:37 ` [PATCH v3 0/7] sched: Flatten the pick K Prateek Nayak
2026-06-12  2:29 ` Shubhang Kaushik
2026-08-17 16:05 ` Szabina Korbai
2026-08-17 16:35   ` K Prateek Nayak
2026-08-18  9:04     ` Szabina Korbai
2026-08-18  9:16       ` Peter Zijlstra
2026-08-21 10:37         ` Szabina Korbai
2026-08-24 14:51         ` Chen Yu [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aoxah90s0bQ4tcUW@three-body \
    --to=chen.yu@linux.dev \
    --cc=bsegall@google.com \
    --cc=cgroups@vger.kernel.org \
    --cc=chenridong@huaweicloud.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=euan@linux.ibm.com \
    --cc=hannes@cmpxchg.org \
    --cc=huschle@linux.ibm.com \
    --cc=jstultz@google.com \
    --cc=juri.lelli@redhat.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=longman@redhat.com \
    --cc=mgorman@suse.de \
    --cc=mingo@kernel.org \
    --cc=mkoutny@suse.com \
    --cc=peterz@infradead.org \
    --cc=philip.li@intel.com \
    --cc=qyousef@layalina.io \
    --cc=rostedt@goodmis.org \
    --cc=szkorbai@linux.ibm.com \
    --cc=tim.c.chen@intel.com \
    --cc=tj@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    --cc=yu.c.chen@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.