From: Shrikanth Hegde <sshegde@linux.ibm.com>
To: Yury Norov <yury.norov@gmail.com>
Cc: linux-kernel@vger.kernel.org, mingo@kernel.org,
peterz@infradead.org, juri.lelli@redhat.com,
vincent.guittot@linaro.org, kprateek.nayak@amd.com,
iii@linux.ibm.com, corbet@lwn.net, meted@linux.ibm.com,
tglx@kernel.org, gregkh@linuxfoundation.org, pbonzini@redhat.com,
seanjc@google.com, vschneid@redhat.com, huschle@linux.ibm.com,
rostedt@goodmis.org, dietmar.eggemann@arm.com,
maddy@linux.ibm.com, srikar@linux.ibm.com, hdanton@sina.com,
chleroy@kernel.org, vineeth@bitbyteword.org, frederic@kernel.org,
arighi@nvidia.com, pauld@redhat.com, christian.loehle@arm.com,
tj@kernel.org, tommaso.cucinotta@gmail.com, maz@kernel.org,
rafael@kernel.org, rdunlap@infradead.org, kernellwp@gmail.com,
linux-doc@vger.kernel.org, jgross@suse.com,
virtualization@lists.linux.dev, sunlightlinux@gmail.com
Subject: Re: [PATCH v12 08/13] sched/core: Push current task from non preferred CPU
Date: Wed, 9 Sep 2026 08:51:00 +0530 [thread overview]
Message-ID: <ee949902-1158-456e-af4d-b1fe6c42ac61@linux.ibm.com> (raw)
In-Reply-To: <aqCTMREUoCmbtygK@yury>
On 9/9/26 4:28 AM, Yury Norov wrote:
> On Mon, Sep 07, 2026 at 08:53:19AM +0530, Shrikanth Hegde wrote:
>> Hi Yury, thanks for taking a look.
>>
>>>> + /* This could take rq lock. So call it before rq lock is taken */
>>>> + cpu = select_fallback_rq(rq->cpu, p);
>>>> + rq_lock(rq, &rf);
>>>
>>> If select_fallback_rq() grabs the lock, then when it releases the
>>> lock, there's a window for race between the other process and the
>>> subsequent rq_lock(). Or I misunderstand it?
>>>
>>
>> select_fallback_rq taking lock is for any state change that needs to happen
>> such as fallback to possible CPUs etc.
>>
>> Most of the time it won't grab the rq lock. Even if the task got pulled by load balancer
>> before grabbing the lock, Below (task_rq(p) == rq) will catch that, and it bails out.
>>
>> So it is safe.
>
> OK... Can you please explain it in the comment above?
>
> ...
Ok.
>
>>>> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
>>>> index 6c3ad70e58b8..678e44134acf 100644
>>>> --- a/kernel/sched/sched.h
>>>> +++ b/kernel/sched/sched.h
>>>> @@ -1298,6 +1298,8 @@ struct rq {
>>>> struct list_head cfs_tasks;
>>>> + bool push_task_work_done;
>>>> +
>>>
>>> It should be protected with CONFIG_PREFERRED_CPU. Also, the name
>>> doesn't look correct. You set the variable to 'true' even before
>>> calling the stopper. Maybe need_push_to_npc, or similar?
>>>
>>
>> ok. npc_push_work_pending is probably a better one?
>>> Why did you place it between cfs_tasks and avg_rt? If no specific
>>> reason, maybe place it next to CONFIG_PARAVIRT-guarded fields.
>>>
>>
>> I don't see a common empty space there. I could increase the size.
>>
>>> What about pahole?
>>
>> I did check pahole on powerpc which has 128 byte cachelines.
>>
>> int online; /* 4524 4 */
>> struct list_head cfs_tasks; /* 4528 16 */
>>
>> /* XXX 64 bytes hole, try to pack */
>>
>> It was empty space. Now, that i check 64 byte cachelines it may not be the
>> optimal one.
>>
>> I do see, a couple common places for both 64 abd 126 byte cacheline. I believe those
>> are better places. It won't increase the size or cause any existing fields to
>> misalign. It also makes sense to guard it again CONFIG_PREFERRED_CPU. I had not
>> done to avoid ifdefs. But it is used only under it. So i think that makes sense too.
>>
>> 1.
>>
>> struct balance_callback * balance_callback; /* 3608 8 */
>> unsigned char nohz_idle_balance; /* 3616 1 */
>> unsigned char idle_balance; /* 3617 1 */
>>
>> /* XXX 6 bytes hole, try to pack */
>> long unsigned int misfit_task_load; /* 3624 8 */
>>
>>
>> 2.
>> unsigned int ttwu_count; /* 5276 4 */
>> unsigned int ttwu_local; /* 5280 4 */
>>
>> /* XXX 4 bytes hole, try to pack */
>>
>> struct cpuidle_state * idle_state; /* 5288 8 */
>
> The struct rq is highly configurable. Depending on your config,
> the holes will migrate to different places. I'd not rely on just
> 'optimizing holes' problem. Just put the new field next to logically
> related existing fields.
>
> You've got paravirt-related prev_steal_time and prev_steal_time_rq,
> and you've got the /* For active balancing */ section. Maybe one of
> them?
>
Ok. Moving it after prev_steal_time_rq.
#ifdef CONFIG_PARAVIRT_TIME_ACCOUNTING
u64 prev_steal_time_rq;
#endif
+#ifdef CONFIG_PREFERRED_CPU
+ bool npc_push_work_pending;
+#endif
I checked on 128 byte cachelines, it didn't increase the number of cachelines
with the change as well. So we are good there.
base:
/* size: 5632, cachelines: 44, members: 104 */
/* sum members: 5022, holes: 15, sum holes: 502 */
with above change:
/* size: 5632, cachelines: 44, members: 105 */
/* sum members: 5023, holes: 17, sum holes: 533 */
On 64 byte cachelines too, there is 16 bytes hole a bit below. So it should absorb it
as well. So we are fine there as well.
u64 prev_steal_time_rq; /* 3960 8 */
/* --- cacheline 62 boundary (3968 bytes) --- */
long unsigned int calc_load_update; /* 3968 8 */
long int calc_load_active; /* 3976 8 */
/* XXX 16 bytes hole, try to pack */
call_single_data_t hrtick_csd __attribute__((__aligned__(32))); /* 4000 32 */
I will make this change and send out v13 today.
> Thanks,
> Yury
next prev parent reply other threads:[~2026-09-09 3:21 UTC|newest]
Thread overview: 19+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 6:32 [PATCH v12 00/13] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff Shrikanth Hegde
2026-09-03 6:32 ` [PATCH v12 01/13] sched/cputime: Add kcpustat_field_total helper Shrikanth Hegde
2026-09-03 16:26 ` Frederic Weisbecker
2026-09-03 6:32 ` [PATCH v12 02/13] cpumask: Introduce cpumask_intersects_and Shrikanth Hegde
2026-09-03 6:32 ` [PATCH v12 03/13] sched/docs: Document cpu_preferred_mask and Preferred CPU concept Shrikanth Hegde
2026-09-03 6:32 ` [PATCH v12 04/13] cpumask: Introduce cpu_preferred_mask Shrikanth Hegde
2026-09-03 6:32 ` [PATCH v12 05/13] sysfs: Add preferred CPU file Shrikanth Hegde
2026-09-03 6:32 ` [PATCH v12 06/13] sched/core: Try to use a preferred CPU in is_cpu_allowed Shrikanth Hegde
2026-09-03 6:32 ` [PATCH v12 07/13] sched/fair: Load balance only among preferred CPUs Shrikanth Hegde
2026-09-03 6:32 ` [PATCH v12 08/13] sched/core: Push current task from non preferred CPU Shrikanth Hegde
2026-09-05 0:28 ` Yury Norov
2026-09-07 3:23 ` Shrikanth Hegde
2026-09-08 22:58 ` Yury Norov
2026-09-09 3:21 ` Shrikanth Hegde [this message]
2026-09-03 6:32 ` [PATCH v12 09/13] sched/debug: Add migration stats due to non preferred CPUs Shrikanth Hegde
2026-09-03 6:32 ` [PATCH v12 10/13] virt: Introduce steal governor driver Shrikanth Hegde
2026-09-03 6:32 ` [PATCH v12 11/13] virt/steal_governor: Add control knobs for handling steal values Shrikanth Hegde
2026-09-03 6:32 ` [PATCH v12 12/13] virt/steal_governor: Implement steal_governor policy loop Shrikanth Hegde
2026-09-03 6:32 ` [PATCH v12 13/13] virt/steal_governor: Enable the driver Shrikanth Hegde
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ee949902-1158-456e-af4d-b1fe6c42ac61@linux.ibm.com \
--to=sshegde@linux.ibm.com \
--cc=arighi@nvidia.com \
--cc=chleroy@kernel.org \
--cc=christian.loehle@arm.com \
--cc=corbet@lwn.net \
--cc=dietmar.eggemann@arm.com \
--cc=frederic@kernel.org \
--cc=gregkh@linuxfoundation.org \
--cc=hdanton@sina.com \
--cc=huschle@linux.ibm.com \
--cc=iii@linux.ibm.com \
--cc=jgross@suse.com \
--cc=juri.lelli@redhat.com \
--cc=kernellwp@gmail.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=maddy@linux.ibm.com \
--cc=maz@kernel.org \
--cc=meted@linux.ibm.com \
--cc=mingo@kernel.org \
--cc=pauld@redhat.com \
--cc=pbonzini@redhat.com \
--cc=peterz@infradead.org \
--cc=rafael@kernel.org \
--cc=rdunlap@infradead.org \
--cc=rostedt@goodmis.org \
--cc=seanjc@google.com \
--cc=srikar@linux.ibm.com \
--cc=sunlightlinux@gmail.com \
--cc=tglx@kernel.org \
--cc=tj@kernel.org \
--cc=tommaso.cucinotta@gmail.com \
--cc=vincent.guittot@linaro.org \
--cc=vineeth@bitbyteword.org \
--cc=virtualization@lists.linux.dev \
--cc=vschneid@redhat.com \
--cc=yury.norov@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).