All of lore.kernel.org
 help / color / mirror / Atom feed
From: Andrea Righi <arighi@nvidia.com>
To: Shrikanth Hegde <sshegde@linux.ibm.com>
Cc: Ingo Molnar <mingo@redhat.com>,
	Peter Zijlstra <peterz@infradead.org>,
	Juri Lelli <juri.lelli@redhat.com>,
	Vincent Guittot <vincent.guittot@linaro.org>,
	Dietmar Eggemann <dietmar.eggemann@arm.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
	Valentin Schneider <vschneid@redhat.com>,
	K Prateek Nayak <kprateek.nayak@amd.com>,
	Christian Loehle <christian.loehle@arm.com>,
	Phil Auld <pauld@redhat.com>, Mete Durlu <meted@linux.ibm.com>,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH v4] sched/fair: Prefer fully idle cores for NOHZ balancing
Date: Tue, 4 Aug 2026 20:09:46 +0200	[thread overview]
Message-ID: <anIq6pU5KXTTFCDN@gpd4> (raw)
In-Reply-To: <92dae374-14ec-4acb-b451-7ac81d9d262b@linux.ibm.com>

Hi Shrikanth,

On Tue, Aug 04, 2026 at 08:49:01PM +0530, Shrikanth Hegde wrote:
> Hi Andrea,
> 
> On 8/4/26 8:43 PM, Andrea Righi wrote:
> > find_new_ilb() selects the first idle housekeeping CPU without
> > considering whether another thread is running on the same physical core.
> > On an SMT system, the idle load balancer can therefore activate both
> > siblings even when another housekeeping CPU has an entirely idle core.
> > 
> > On most SMT systems, this is not problematic because the idle load
> > balancer is a short-lived activity and the transient wakeup of a sibling
> > has negligible performance impact.
> > 
> > However, this can be particularly costly on NVIDIA Olympus cores used in
> > Vera. Briefly activating an otherwise idle sibling can reduce the
> > performance available to the other sibling and this effect does not
> > necessarily end once the activated sibling becomes idle: after the ILB
> > finishes and its CPU enters WFI, full single-thread performance is
> > restored only after the sibling has remained idle for a qualification
> > interval (10 Ki cycles on the tested Vera system). Repeated short
> > sibling wakeups can therefore sustain the interference even with little
> > actual overlap.
> 
> Is this hardware cost has been accounted in things like cpufreq or cpuidle?
> or have you set it already to max performance.

All the CPUs were using the cppc_cpufreq performance governor during the tests.
The scaling minimum frequency was between 95.5% and 97% of the maximum
frequency, depending on the CPU. The CPUs were not strictly pinned to their
maximum frequency, but the performance governor was in use.

The acpi_idle driver was enabled, with both LPI-0 and LPI-1 available. I didn't
disable the idle states.

> 
> > 
> > Prevent this by preferring an idle housekeeping CPU whose entire SMT
> > core is idle. Retain the first idle CPU as a fallback when no fully idle
> > core is available, so NOHZ balancing continues to make forward progress.
> > Once a partially busy core has been examined, skip its remaining SMT
> > siblings to avoid repeating the core-idle check on wide SMT systems.
> > 
> 
> Just Curious, making nohz_full=<except first core> yeilds similar numbers?

Yes, provided that the housekeeping core is also excluded from the workload.

I tested this booting with:

  nohz_full=1-175,177-351

This leaves CPUs 0 and 176 for housekeeping. I excluded that core from the
workload and ran 87 OpenMP tasks on the remaining node-0 cores (1 task per core,
excluding the housekeeping one):

 $ env OMP_NUM_THREADS=87 \
       OMP_DYNAMIC=false \
       OPENBLAS_LOOPS=10 \
       OPENBLAS_PARAM_M=16384 \
       OPENBLAS_PARAM_N=16384 \
       OPENBLAS_PARAM_K=16384 \
       numactl -C 1-87,177-263 --membind=0 \
       ./benchmark/sgemm.goto 1 1 1

Results:

    unpatched           : 5.246 TFLOP/s
    unpatched+nohz_full : 6.953 TFLOP/s
    patched             : 6.861 TFLOP/s
    patched+nohz_full   : 6.984 TFLOP/s

So, nohz_full seems to prevent the problematic ILB wakeups and can produce
similar results for this CPU-bound workload. However, it shouldn't be considered
a sobstiute for the ILB fix, since it requires explicit partitioning and
reserved housekeeping CPUs. Full dyntick also enables context tracking on the
isolated CPUs, adding kernel entry/exit overhead.

Thanks,
-Andrea

> 
> > Tests performed using an ad hoc GEMM benchmark running one CPU-intensive
> > task per SMT core within its CPU affinity mask improved from
> > approximately 6.2 TFLOP/s to 9.4 TFLOP/s.
> > 
> > Note that this preference may wake a fully idle physical core instead of
> > using an idle sibling of an active core, potentially increasing ILB
> > wakeup latency or energy consumption on some architectures. It may also
> > scan additional CPUs before selecting the one to run the ILB. The
> > selection falls back to the first idle CPU when no fully idle SMT core
> > is available. Non-SMT systems continue to select the first idle
> > housekeeping CPU.
> > 
> > Cc: Vincent Guittot <vincent.guittot@linaro.org>
> > Cc: K Prateek Nayak <kprateek.nayak@amd.com>
> > Cc: Shrikanth Hegde <sshegde@linux.ibm.com>
> > Signed-off-by: Andrea Righi <arighi@nvidia.com>
> > Reviewed-by: Mete Durlu <meted@linux.ibm.com>
> > ---
> > Changes in v4:
> >   - Remove redundant this_cpu check (Prateek Nayak, Vincent Guittot)
> >   - Link to v3: https://lore.kernel.org/all/20260731191957.3199642-1-arighi@nvidia.com/
> > 
> > Changes in v3:
> >   - After finding an idle fallback, skip all siblings when a busy CPU is
> >     encountered, avoiding per-CPU traversal of known-busy cores (Mete Durlu)
> >   - Link to v2: https://lore.kernel.org/all/20260729163225.1987068-1-arighi@nvidia.com/
> > 
> > Changes in v2:
> >   - Avoid repeated is_core_idle() checks on wide SMT systems by pruning
> >     the remaining siblings of a partially busy core (Prateek Nayak)
> >   - Link to v1: https://lore.kernel.org/r/20260728214442.1648483-1-arighi@nvidia.com/
> > 
> >   kernel/sched/fair.c | 55 ++++++++++++++++++++++++++++++++++++---------
> >   1 file changed, 44 insertions(+), 11 deletions(-)
> > 
> > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> > index 37001c63452e5..9a975a684b487 100644
> > --- a/kernel/sched/fair.c
> > +++ b/kernel/sched/fair.c
> > @@ -13964,29 +13964,62 @@ static inline int on_null_domain(struct rq *rq)
> >    */
> >   static inline int find_new_ilb(void)
> >   {
> > -	int this_cpu = smp_processor_id();
> > -	const struct cpumask *hk_mask;
> > -	int ilb_cpu;
> > +	struct cpumask *ilb_cpus;
> > +	int ilb_cpu, fallback = -1;
> > +
> > +	lockdep_assert_irqs_disabled();
> > +
> > +	/*
> > +	 * Reuse the per-CPU select_rq_mask, which is protected from concurrent
> > +	 * use on this CPU by having interrupts disabled.
> > +	 */
> > +	ilb_cpus = this_cpu_cpumask_var_ptr(select_rq_mask);
> > +	cpumask_and(ilb_cpus, nohz.idle_cpus_mask,
> > +		    housekeeping_cpumask(HK_TYPE_KERNEL_NOISE));
> > +
> > +	for_each_cpu(ilb_cpu, ilb_cpus) {
> > +		if (!idle_cpu(ilb_cpu)) {
> > +			/*
> > +			 * Once an idle fallback exists, a busy CPU proves that
> > +			 * this core cannot be fully idle. Skip its siblings.
> > +			 */
> > +			if (sched_smt_active() && fallback >= 0)
> > +				cpumask_andnot(ilb_cpus, ilb_cpus, cpu_smt_mask(ilb_cpu));
> > +			continue;
> > +		}
> > -	hk_mask = housekeeping_cpumask(HK_TYPE_KERNEL_NOISE);
> > +		/*
> > +		 * Running the idle load balancer on an idle sibling of a busy
> > +		 * SMT core can reduce the capacity available to its sibling. Prefer
> > +		 * a CPU whose entire core is idle, but retain the first idle CPU as
> > +		 * a fallback so idle balancing can still make progress when no fully
> > +		 * idle core exists.
> > +		 */
> > +		if (sched_smt_active() && !is_core_idle(ilb_cpu)) {
> > +			if (fallback < 0)
> > +				fallback = ilb_cpu;
> > -	for_each_cpu_and(ilb_cpu, nohz.idle_cpus_mask, hk_mask) {
> > -		if (ilb_cpu == this_cpu)
> > +			/*
> > +			 * The core is not idle, so there is no need to check
> > +			 * any of its other SMT siblings.
> > +			 */
> > +			cpumask_andnot(ilb_cpus, ilb_cpus,
> > +				       cpu_smt_mask(ilb_cpu));
> >   			continue;
> > +		}
> > -		if (idle_cpu(ilb_cpu))
> > -			return ilb_cpu;
> > +		return ilb_cpu;
> >   	}
> > -	return -1;
> > +	return fallback;
> >   }
> >   /*
> >    * Kick a CPU to do the NOHZ balancing, if it is time for it, via a cross-CPU
> >    * SMP function call (IPI).
> >    *
> > - * We pick the first idle CPU in the HK_TYPE_KERNEL_NOISE housekeeping set
> > - * (if there is one).
> > + * Prefer a CPU on a fully idle core in the HK_TYPE_KERNEL_NOISE housekeeping
> > + * set. Fall back to the first idle CPU when no fully idle core exists.
> >    */
> >   static void kick_ilb(unsigned int flags)
> >   {
> 

  reply	other threads:[~2026-08-04 18:10 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-04 15:13 [PATCH v4] sched/fair: Prefer fully idle cores for NOHZ balancing Andrea Righi
2026-08-04 15:19 ` Shrikanth Hegde
2026-08-04 18:09   ` Andrea Righi [this message]
2026-08-04 21:18     ` Shrikanth Hegde
2026-08-05 10:31 ` Vincent Guittot
2026-08-06 10:26 ` Shrikanth Hegde
2026-08-06 13:18   ` Andrea Righi
2026-08-07 12:26     ` Shrikanth Hegde
2026-08-08  9:44 ` [tip: sched/core] " tip-bot2 for Andrea Righi

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anIq6pU5KXTTFCDN@gpd4 \
    --to=arighi@nvidia.com \
    --cc=bsegall@google.com \
    --cc=christian.loehle@arm.com \
    --cc=dietmar.eggemann@arm.com \
    --cc=juri.lelli@redhat.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=meted@linux.ibm.com \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=pauld@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=sshegde@linux.ibm.com \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.