linux-doc.vger.kernel.org archive mirror
 help / color / mirror / Atom feed
From: Shrikanth Hegde <sshegde@linux.ibm.com>
To: Yury Norov <yury.norov@gmail.com>
Cc: linux-kernel@vger.kernel.org, mingo@kernel.org,
	peterz@infradead.org, juri.lelli@redhat.com,
	vincent.guittot@linaro.org, kprateek.nayak@amd.com,
	iii@linux.ibm.com, corbet@lwn.net, meted@linux.ibm.com,
	tglx@kernel.org, gregkh@linuxfoundation.org, pbonzini@redhat.com,
	seanjc@google.com, vschneid@redhat.com, huschle@linux.ibm.com,
	rostedt@goodmis.org, dietmar.eggemann@arm.com,
	maddy@linux.ibm.com, srikar@linux.ibm.com, hdanton@sina.com,
	chleroy@kernel.org, vineeth@bitbyteword.org, frederic@kernel.org,
	arighi@nvidia.com, pauld@redhat.com, christian.loehle@arm.com,
	tj@kernel.org, tommaso.cucinotta@gmail.com, maz@kernel.org,
	rafael@kernel.org, rdunlap@infradead.org, kernellwp@gmail.com,
	linux-doc@vger.kernel.org, jgross@suse.com,
	virtualization@lists.linux.dev, sunlightlinux@gmail.com
Subject: Re: [PATCH v12 08/13] sched/core: Push current task from non preferred CPU
Date: Wed, 9 Sep 2026 08:51:00 +0530	[thread overview]
Message-ID: <ee949902-1158-456e-af4d-b1fe6c42ac61@linux.ibm.com> (raw)
In-Reply-To: <aqCTMREUoCmbtygK@yury>



On 9/9/26 4:28 AM, Yury Norov wrote:
> On Mon, Sep 07, 2026 at 08:53:19AM +0530, Shrikanth Hegde wrote:
>> Hi Yury, thanks for taking a look.
>>

>>>> +	/* This could take rq lock. So call it before rq lock is taken */
>>>> +	cpu = select_fallback_rq(rq->cpu, p);
>>>> +	rq_lock(rq, &rf);
>>>
>>> If select_fallback_rq() grabs the lock, then when it releases the
>>> lock, there's a window for race between the other process and the
>>> subsequent rq_lock(). Or I misunderstand it?
>>>
>>
>> select_fallback_rq taking lock is for any state change that needs to happen
>> such as fallback to possible CPUs etc.
>>
>> Most of the time it won't grab the rq lock. Even if the task got pulled by load balancer
>> before grabbing the lock, Below (task_rq(p) == rq) will catch that, and it bails out.
>>
>> So it is safe.
> 
> OK... Can you please explain it in the comment above?
> 
> ...

Ok.

> 
>>>> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
>>>> index 6c3ad70e58b8..678e44134acf 100644
>>>> --- a/kernel/sched/sched.h
>>>> +++ b/kernel/sched/sched.h
>>>> @@ -1298,6 +1298,8 @@ struct rq {
>>>>    	struct list_head cfs_tasks;
>>>> +	bool			push_task_work_done;
>>>> +
>>>
>>> It should be protected with CONFIG_PREFERRED_CPU. Also, the name
>>> doesn't look correct. You set the variable to 'true' even before
>>> calling the stopper. Maybe need_push_to_npc, or similar?
>>>
>>
>> ok. npc_push_work_pending is probably a better one?
>>> Why did you place it between cfs_tasks and avg_rt? If no specific
>>> reason, maybe place it next to CONFIG_PARAVIRT-guarded fields.
>>>
>>
>> I don't see a common empty space there. I could increase the size.
>>
>>> What about pahole?
>>
>> I did check pahole on powerpc which has 128 byte cachelines.
>>
>> 	int                        online;               /*  4524     4 */
>> 	struct list_head           cfs_tasks;            /*  4528    16 */
>>
>> 	/* XXX 64 bytes hole, try to pack */
>>
>> It was empty space. Now, that i check 64 byte cachelines it may not be the
>> optimal one.
>>
>> I do see, a couple common places for both 64 abd 126 byte cacheline. I believe those
>> are better places. It won't increase the size or cause any existing fields to
>> misalign. It also makes sense to guard it again CONFIG_PREFERRED_CPU. I had not
>> done to avoid ifdefs. But it is used only under it. So i think that makes sense too.
>>
>> 1.
>>
>> 	struct balance_callback *  balance_callback;     /*  3608     8 */
>> 	unsigned char              nohz_idle_balance;    /*  3616     1 */
>> 	unsigned char              idle_balance;         /*  3617     1 */
>>
>> 	/* XXX 6 bytes hole, try to pack */
>> 	long unsigned int          misfit_task_load;     /*  3624     8 */
>>
>>
>> 2.
>> 	unsigned int               ttwu_count;           /*  5276     4 */
>> 	unsigned int               ttwu_local;           /*  5280     4 */
>>
>> 	/* XXX 4 bytes hole, try to pack */
>>
>> 	struct cpuidle_state *     idle_state;           /*  5288     8 */
> 
> The struct rq is highly configurable. Depending on your config,
> the holes will migrate to different places. I'd not rely on just
> 'optimizing holes' problem. Just put the new field next to logically
> related existing fields.
> 
> You've got paravirt-related prev_steal_time and prev_steal_time_rq,
> and you've got the /* For active balancing */ section. Maybe one of
> them?
> 

Ok. Moving it after prev_steal_time_rq.

  #ifdef CONFIG_PARAVIRT_TIME_ACCOUNTING
         u64                     prev_steal_time_rq;
  #endif
+#ifdef CONFIG_PREFERRED_CPU
+       bool                    npc_push_work_pending;
+#endif


I checked on 128 byte cachelines, it didn't increase the number of cachelines
with the change as well. So we are good there.

base:
         /* size: 5632, cachelines: 44, members: 104 */
         /* sum members: 5022, holes: 15, sum holes: 502 */
with above change:
         /* size: 5632, cachelines: 44, members: 105 */
         /* sum members: 5023, holes: 17, sum holes: 533 */


On 64 byte cachelines too, there is 16 bytes hole a bit below. So it should absorb it
as well. So we are fine there as well.

	u64                        prev_steal_time_rq;   /*  3960     8 */
	/* --- cacheline 62 boundary (3968 bytes) --- */
	long unsigned int          calc_load_update;     /*  3968     8 */
	long int                   calc_load_active;     /*  3976     8 */

	/* XXX 16 bytes hole, try to pack */

	call_single_data_t         hrtick_csd __attribute__((__aligned__(32))); /*  4000    32 */

I will make this change and send out v13 today.
> Thanks,
> Yury


  reply	other threads:[~2026-09-09  3:21 UTC|newest]

Thread overview: 19+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-03  6:32 [PATCH v12 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
2026-09-03  6:32 ` [PATCH v12 01/13] sched/cputime: Add kcpustat_field_total helper Shrikanth Hegde
2026-09-03 16:26   ` Frederic Weisbecker
2026-09-03  6:32 ` [PATCH v12 02/13] cpumask: Introduce cpumask_intersects_and Shrikanth Hegde
2026-09-03  6:32 ` [PATCH v12 03/13] sched/docs: Document cpu_preferred_mask and Preferred CPU concept Shrikanth Hegde
2026-09-03  6:32 ` [PATCH v12 04/13] cpumask: Introduce cpu_preferred_mask Shrikanth Hegde
2026-09-03  6:32 ` [PATCH v12 05/13] sysfs: Add preferred CPU file Shrikanth Hegde
2026-09-03  6:32 ` [PATCH v12 06/13] sched/core: Try to use a preferred CPU in is_cpu_allowed Shrikanth Hegde
2026-09-03  6:32 ` [PATCH v12 07/13] sched/fair: Load balance only among preferred CPUs Shrikanth Hegde
2026-09-03  6:32 ` [PATCH v12 08/13] sched/core: Push current task from non preferred CPU Shrikanth Hegde
2026-09-05  0:28   ` Yury Norov
2026-09-07  3:23     ` Shrikanth Hegde
2026-09-08 22:58       ` Yury Norov
2026-09-09  3:21         ` Shrikanth Hegde [this message]
2026-09-03  6:32 ` [PATCH v12 09/13] sched/debug: Add migration stats due to non preferred CPUs Shrikanth Hegde
2026-09-03  6:32 ` [PATCH v12 10/13] virt: Introduce steal governor driver Shrikanth Hegde
2026-09-03  6:32 ` [PATCH v12 11/13] virt/steal_governor: Add control knobs for handling steal values Shrikanth Hegde
2026-09-03  6:32 ` [PATCH v12 12/13] virt/steal_governor: Implement steal_governor policy loop Shrikanth Hegde
2026-09-03  6:32 ` [PATCH v12 13/13] virt/steal_governor: Enable the driver Shrikanth Hegde

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ee949902-1158-456e-af4d-b1fe6c42ac61@linux.ibm.com \
    --to=sshegde@linux.ibm.com \
    --cc=arighi@nvidia.com \
    --cc=chleroy@kernel.org \
    --cc=christian.loehle@arm.com \
    --cc=corbet@lwn.net \
    --cc=dietmar.eggemann@arm.com \
    --cc=frederic@kernel.org \
    --cc=gregkh@linuxfoundation.org \
    --cc=hdanton@sina.com \
    --cc=huschle@linux.ibm.com \
    --cc=iii@linux.ibm.com \
    --cc=jgross@suse.com \
    --cc=juri.lelli@redhat.com \
    --cc=kernellwp@gmail.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=maddy@linux.ibm.com \
    --cc=maz@kernel.org \
    --cc=meted@linux.ibm.com \
    --cc=mingo@kernel.org \
    --cc=pauld@redhat.com \
    --cc=pbonzini@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rafael@kernel.org \
    --cc=rdunlap@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=seanjc@google.com \
    --cc=srikar@linux.ibm.com \
    --cc=sunlightlinux@gmail.com \
    --cc=tglx@kernel.org \
    --cc=tj@kernel.org \
    --cc=tommaso.cucinotta@gmail.com \
    --cc=vincent.guittot@linaro.org \
    --cc=vineeth@bitbyteword.org \
    --cc=virtualization@lists.linux.dev \
    --cc=vschneid@redhat.com \
    --cc=yury.norov@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).