The Linux Kernel Mailing List
 help / color / mirror / Atom feed
From: Mete Durlu <meted@linux.ibm.com>
To: Andrea Righi <arighi@nvidia.com>, Ingo Molnar <mingo@redhat.com>,
	Peter Zijlstra <peterz@infradead.org>,
	Juri Lelli <juri.lelli@redhat.com>,
	Vincent Guittot <vincent.guittot@linaro.org>
Cc: Dietmar Eggemann <dietmar.eggemann@arm.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
	Valentin Schneider <vschneid@redhat.com>,
	K Prateek Nayak <kprateek.nayak@amd.com>,
	Christian Loehle <christian.loehle@arm.com>,
	Shrikanth Hegde <sshegde@linux.ibm.com>,
	Phil Auld <pauld@redhat.com>,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH v3] sched/fair: Prefer fully idle cores for NOHZ balancing
Date: Tue, 4 Aug 2026 14:36:48 +0200	[thread overview]
Message-ID: <533a1615-e0f8-4c7d-b212-b24760f198d3@linux.ibm.com> (raw)
In-Reply-To: <20260731191957.3199642-1-arighi@nvidia.com>

Hi,

> find_new_ilb() selects the first idle housekeeping CPU without
> considering whether another thread is running on the same physical core.
> On an SMT system, the idle load balancer can therefore activate both
> siblings even when another housekeeping CPU has an entirely idle core.
> 
> On most SMT systems, this is not problematic because the idle load
> balancer is a short-lived activity and the transient wakeup of a sibling
> has negligible performance impact.
> 
> However, this can be particularly costly on NVIDIA Olympus cores used in
> Vera. Briefly activating an otherwise idle sibling can reduce the
> performance available to the other sibling and this effect does not
> necessarily end once the activated sibling becomes idle: after the ILB
> finishes and its CPU enters WFI, full single-thread performance is
> restored only after the sibling has remained idle for a qualification
> interval (10 Ki cycles on the tested Vera system). Repeated short
> sibling wakeups can therefore sustain the interference even with little
> actual overlap.
> 
> Prevent this by preferring an idle housekeeping CPU whose entire SMT
> core is idle. Retain the first idle CPU as a fallback when no fully idle
> core is available, so NOHZ balancing continues to make forward progress.
> Once a partially busy core has been examined, skip its remaining SMT
> siblings to avoid repeating the core-idle check on wide SMT systems.
> 
> Tests performed using an ad hoc GEMM benchmark running one CPU-intensive
> task per SMT core within its CPU affinity mask improved from
> approximately 6.2 TFLOP/s to 9.4 TFLOP/s.

Although what you describe above with siblings suffering interference
does not really fit to s390, I'd like to hear more about what sort
of GEMM (general matrix multiplication) tests you did.

I tested this patch with a couple of different tools
- perf bench sched pipe
- hackbench
- uperf
- cyclictest
- stress-ng (3d-matrix and cyclic)

Didn't come across any meaningful difference in any of them on multiple
runs each. So I was curious about the exact sort of benchmark you
mention here.

One minor nit for the diff below;

> 
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 37001c63452e5..574b6b3ee922a 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -13965,28 +13965,66 @@ static inline int on_null_domain(struct rq *rq)
>   static inline int find_new_ilb(void)
>   {
>   	int this_cpu = smp_processor_id();
> -	const struct cpumask *hk_mask;
> -	int ilb_cpu;
> +	struct cpumask *ilb_cpus;
> +	int ilb_cpu, fallback = -1;
> +
> +	lockdep_assert_irqs_disabled();
>   
> -	hk_mask = housekeeping_cpumask(HK_TYPE_KERNEL_NOISE);
> +	/*
> +	 * Reuse the per-CPU select_rq_mask, which is protected from concurrent
> +	 * use on this CPU by having interrupts disabled.
> +	 */
> +	ilb_cpus = this_cpu_cpumask_var_ptr(select_rq_mask);
> +	cpumask_and(ilb_cpus, nohz.idle_cpus_mask,
> +		    housekeeping_cpumask(HK_TYPE_KERNEL_NOISE));
>   
> -	for_each_cpu_and(ilb_cpu, nohz.idle_cpus_mask, hk_mask) {
> +	for_each_cpu(ilb_cpu, ilb_cpus) {
>   		if (ilb_cpu == this_cpu)
>   			continue;
>   
> -		if (idle_cpu(ilb_cpu))
> -			return ilb_cpu;
> +		if (!idle_cpu(ilb_cpu)) {
> +			/*
> +			 * Once an idle fallback exists, a busy CPU proves that
> +			 * this core cannot be fully idle. Skip its siblings.
> +			 */
> +			if (sched_smt_active() && fallback >= 0)
> +				cpumask_andnot(ilb_cpus, ilb_cpus,
> +					       cpu_smt_mask(ilb_cpu));

nit;
With line break this if block is now taking multiple lines and deserves
its own curly braces.

With or without the nit, feel free to add my r-b to v4, I doubt removal
of the "this_cpu" check will change anything as it is a dud.

Reviewed By: Mete Durlu <meted@linux.ibm.com>


  parent reply	other threads:[~2026-08-04 12:37 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-31 19:19 [PATCH v3] sched/fair: Prefer fully idle cores for NOHZ balancing Andrea Righi
2026-08-04  8:42 ` Vincent Guittot
2026-08-04  9:48   ` K Prateek Nayak
2026-08-04 10:30     ` Vincent Guittot
2026-08-04 12:16       ` Andrea Righi
2026-08-04 12:36 ` Mete Durlu [this message]
2026-08-04 14:58   ` Andrea Righi
2026-08-05  8:59 ` K Prateek Nayak

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=533a1615-e0f8-4c7d-b212-b24760f198d3@linux.ibm.com \
    --to=meted@linux.ibm.com \
    --cc=arighi@nvidia.com \
    --cc=bsegall@google.com \
    --cc=christian.loehle@arm.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=juri.lelli@redhat.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=pauld@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=sshegde@linux.ibm.com \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox