All of lore.kernel.org
 help / color / mirror / Atom feed
From: Yury Norov <ynorov@nvidia.com>
To: Shrikanth Hegde <sshegde@linux.ibm.com>
Cc: linux-kernel@vger.kernel.org, mingo@kernel.org,
	peterz@infradead.org, juri.lelli@redhat.com,
	vincent.guittot@linaro.org, yury.norov@gmail.com,
	kprateek.nayak@amd.com, iii@linux.ibm.com, corbet@lwn.net,
	tglx@kernel.org, gregkh@linuxfoundation.org, pbonzini@redhat.com,
	seanjc@google.com, vschneid@redhat.com, huschle@linux.ibm.com,
	rostedt@goodmis.org, dietmar.eggemann@arm.com,
	maddy@linux.ibm.com, srikar@linux.ibm.com, hdanton@sina.com,
	chleroy@kernel.org, vineeth@bitbyteword.org, frederic@kernel.org,
	arighi@nvidia.com, pauld@redhat.com, christian.loehle@arm.com,
	tj@kernel.org, tommaso.cucinotta@gmail.com, maz@kernel.org,
	rafael@kernel.org, rdunlap@infradead.org, kernellwp@gmail.com,
	linux-doc@vger.kernel.org, jgross@suse.com,
	virtualization@lists.linux.dev
Subject: Re: [PATCH v9 05/11] sched/fair: Load balance only among preferred CPUs
Date: Fri, 24 Jul 2026 17:40:26 -0400	[thread overview]
Message-ID: <amPbythM-IBt6j_e@yury> (raw)
In-Reply-To: <20260724140732.2683314-6-sshegde@linux.ibm.com>

On Fri, Jul 24, 2026 at 07:37:26PM +0530, Shrikanth Hegde wrote:
> When cpu is marked as non preferred, any load pulled towards it is
> pointless since in the next tick task will be pushed out again.
> So, Consider only preferred CPUs for load balance.
> 
> This makes it not fight against the push task mechanism which happens
> at tick. Also, this stops active balance to happen on non-preferred CPU
> pulling the load.
> 
> This means there is no load balancing if the task is pinned only to
> non-preferred CPUs. They will continue to run where they were previously
> running before the CPUs was marked as non-preferred.
> 
> Bailout early for NEWIDLE and IDLE balance as load balancing is done
> only on preferred CPUs.
> 
> Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com>
> ---
>  kernel/sched/fair.c | 11 +++++------
>  1 file changed, 5 insertions(+), 6 deletions(-)
> 
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index df8c9c2c7918..12f5b7de28d2 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -13399,7 +13399,7 @@ static int sched_balance_rq(int this_cpu, struct rq *this_rq,
>  	};
>  	bool need_unlock = false;
>  
> -	cpumask_and(cpus, sched_domain_span(sd), cpu_active_mask);
> +	cpumask_and(cpus, sched_domain_span(sd), cpu_preferred_mask);
>  
>  	schedstat_inc(sd->lb_count[idle]);
>  
> @@ -14337,7 +14337,8 @@ static void _nohz_idle_balance(struct rq *this_rq, unsigned int flags)
>  			update_rq_clock(rq);
>  			rq_unlock_irqrestore(rq, &rf);
>  
> -			if (flags & NOHZ_BALANCE_KICK)
> +			if (flags & NOHZ_BALANCE_KICK &&
> +			    cpu_preferred(balance_cpu))
>  				sched_balance_domains(rq, CPU_IDLE);

Here you skip the sched_balance_domains() for idle non-preferred CPUs,
which means you don't re-calculate the rq->next_balance.

In the following code, we update the global nohz.next_balance
depending on the new rq->next_balance, which doesn't happen if
the sched_balance_domains() is not invoked.

So, you propagate an already-expired per-CPU deadline back into the
global NOHZ deadline. It may lead to unneeded asynchronous IPI in the
nohz_balancer_kick() -> kick_ilb() path.

I think, you need another helper, something like:

        void sched_balance_domains(rq, idle)
        {
                __sched_balance_domains(rq, idle);
                advance_rq_next_balance(rq, idle);
        }

And then in the code above:

        if (flags & NOHZ_BALANCE_KICK) {
                if (cpu_preferred(balance_cpu))
                        __sched_balance_domains(rq, CPU_IDLE);

                advance_rq_next_balance(rq, idle);
        }

> @@ -14481,10 +14482,8 @@ static int sched_balance_newidle(struct rq *this_rq, struct rq_flags *rf)
>  	 */
>  	this_rq->idle_stamp = rq_clock(this_rq);
>  
> -	/*
> -	 * Do not pull tasks towards !active CPUs...
> -	 */
> -	if (!cpu_active(this_cpu))
> +	/* Do not pull tasks towards !preferred CPUs */
> +	if (!cpu_preferred(this_cpu))
>  		return 0;
>  
>  	/*
> -- 
> 2.47.3

  reply	other threads:[~2026-07-24 21:40 UTC|newest]

Thread overview: 18+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-24 14:07 [PATCH v9 00/11] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
2026-07-24 14:07 ` [PATCH v9 01/11] sched/docs: Document cpu_preferred_mask and Preferred CPU concept Shrikanth Hegde
2026-07-24 21:45   ` Yury Norov
2026-07-24 14:07 ` [PATCH v9 02/11] cpumask: Introduce cpu_preferred_mask Shrikanth Hegde
2026-07-24 20:00   ` Yury Norov
2026-07-24 14:07 ` [PATCH v9 03/11] sysfs: Add preferred CPU file Shrikanth Hegde
2026-07-24 14:07 ` [PATCH v9 04/11] sched/core: Try to use a preferred CPU in is_cpu_allowed Shrikanth Hegde
2026-07-24 14:07 ` [PATCH v9 05/11] sched/fair: Load balance only among preferred CPUs Shrikanth Hegde
2026-07-24 21:40   ` Yury Norov [this message]
2026-07-24 14:07 ` [PATCH v9 06/11] sched/core: Push current task from non preferred CPU Shrikanth Hegde
2026-07-24 22:04   ` Yury Norov
2026-07-24 14:07 ` [PATCH v9 07/11] sched/debug: Add migration stats due to non preferred CPUs Shrikanth Hegde
2026-07-24 14:07 ` [PATCH v9 08/11] virt: Introduce steal governor driver Shrikanth Hegde
2026-07-24 14:07 ` [PATCH v9 09/11] virt/steal_governor: Add control knobs for handling steal values Shrikanth Hegde
2026-07-24 14:07 ` [PATCH v9 10/11] virt/steal_governor: Implement steal_governor policy loop Shrikanth Hegde
2026-07-24 21:05   ` Yury Norov
2026-07-24 14:07 ` [PATCH v9 11/11] virt/steal_governor: Enable the driver Shrikanth Hegde
2026-07-24 22:07 ` [PATCH v9 00/11] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Yury Norov

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=amPbythM-IBt6j_e@yury \
    --to=ynorov@nvidia.com \
    --cc=arighi@nvidia.com \
    --cc=chleroy@kernel.org \
    --cc=christian.loehle@arm.com \
    --cc=corbet@lwn.net \
    --cc=dietmar.eggemann@arm.com \
    --cc=frederic@kernel.org \
    --cc=gregkh@linuxfoundation.org \
    --cc=hdanton@sina.com \
    --cc=huschle@linux.ibm.com \
    --cc=iii@linux.ibm.com \
    --cc=jgross@suse.com \
    --cc=juri.lelli@redhat.com \
    --cc=kernellwp@gmail.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=maddy@linux.ibm.com \
    --cc=maz@kernel.org \
    --cc=mingo@kernel.org \
    --cc=pauld@redhat.com \
    --cc=pbonzini@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rafael@kernel.org \
    --cc=rdunlap@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=seanjc@google.com \
    --cc=srikar@linux.ibm.com \
    --cc=sshegde@linux.ibm.com \
    --cc=tglx@kernel.org \
    --cc=tj@kernel.org \
    --cc=tommaso.cucinotta@gmail.com \
    --cc=vincent.guittot@linaro.org \
    --cc=vineeth@bitbyteword.org \
    --cc=virtualization@lists.linux.dev \
    --cc=vschneid@redhat.com \
    --cc=yury.norov@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.