From: Andrea Righi <arighi@nvidia.com>
To: K Prateek Nayak <kprateek.nayak@amd.com>
Cc: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
Changwoo Min <changwoo@igalia.com>,
John Stultz <jstultz@google.com>, Ingo Molnar <mingo@redhat.com>,
Peter Zijlstra <peterz@infradead.org>,
Juri Lelli <juri.lelli@redhat.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
Christian Loehle <christian.loehle@arm.com>,
David Dai <david.dai@linux.dev>, Emil Tsalapatis <etsal@meta.com>,
Lee Trager <ltrager@nvidia.com>,
Richard Cheng <icheng@nvidia.com>, Koba Ko <kobak@nvidia.com>,
Aiqun Yu <aiqun.yu@oss.qualcomm.com>,
sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: Re: [PATCH 02/18] sched/core: Dequeue waking proxy donors before reset
Date: Tue, 8 Sep 2026 11:28:00 +0200 [thread overview]
Message-ID: <ap_VIJdG80ZZ_8D0@gpd4> (raw)
In-Reply-To: <3d4116ff-0634-4ab6-be24-0bd2c68f1698@amd.com>
Hi Prateek,
On Tue, Sep 01, 2026 at 10:54:44AM +0530, K Prateek Nayak wrote:
> Hello Andrea,
>
> On 8/31/2026 7:12 PM, Andrea Righi wrote:
> > proxy_needs_return() resets an active donor while holding blocked_lock.
> > proxy_reset_donor() invokes scheduling-class callbacks, adding an
> > unnecessary raw-spinlock nesting. It also presents the waking donor to
> > put_prev_task() as still runnable immediately before block_task()
> > removes it from the runqueue.
>
> Is that an issue for scx?
Yes. When a proxy-migrated donor wakes while it's still rq->donor,
proxy_needs_return() has to reset the rq's donor before clearing the task's
generic on_rq state and returning it through the full wakeup path.
proxy_reset_donor() calls put_prev_set_next_task(), which invokes
put_prev_task_scx() for an EXT donor. If this happens before the
scheduling-class dequeue, SCX_TASK_QUEUED and is_blocked are still set, so
put_prev_task_scx() re-enqueues the donor through scx_do_enqueue_task() and the
following block_task() immediately dequeues it again.
I can clarify this sched_ext-specific ordering better in the patch description.
>
> > Split block_task() so the waking donor can first be dequeued from its
> > scheduling class. Release blocked_lock, dequeue the donor while its
> > generic on_rq state still prevents migration, replace all donor
> > references, and only then complete the generic runqueue removal. This
> > follows the normal sleep ordering and avoids transiently re-enqueuing
> > the waking donor.
> >
> > This is a preparatory change to support proxy execution with sched_ext.
> >
> > Signed-off-by: Andrea Righi <arighi@nvidia.com>
> > ---
> > kernel/sched/core.c | 30 +++++++++++++++++++++++-------
> > 1 file changed, 23 insertions(+), 7 deletions(-)
> >
> > diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> > index 5817d1a4cea2c..237d216382f46 100644
> > --- a/kernel/sched/core.c
> > +++ b/kernel/sched/core.c
> > @@ -2252,7 +2252,8 @@ void deactivate_task(struct rq *rq, struct task_struct *p, int flags)
> > dequeue_task(rq, p, flags);
> > }
> >
> > -static void block_task(struct rq *rq, struct task_struct *p, unsigned long task_state)
> > +static bool dequeue_block_task(struct rq *rq, struct task_struct *p,
> > + unsigned long task_state)
> > {
> > int flags = DEQUEUE_NOCLOCK;
> >
> > @@ -2273,9 +2274,15 @@ static void block_task(struct rq *rq, struct task_struct *p, unsigned long task_
> > *
> > * Where __schedule() and ttwu() have matching control dependencies.
> > *
> > - * After this, schedule() must not care about p->state any more.
> > + * Once the caller invokes __block_task(), schedule() must not care about
> > + * p->state any more.
> > */
> > - if (dequeue_task(rq, p, DEQUEUE_SLEEP | flags))
> > + return dequeue_task(rq, p, DEQUEUE_SLEEP | flags);
> > +}
> > +
> > +static void block_task(struct rq *rq, struct task_struct *p, unsigned long task_state)
> > +{
> > + if (dequeue_block_task(rq, p, task_state))
> > __block_task(rq, p);
> > }
> >
> > @@ -3774,6 +3781,9 @@ static inline void proxy_reset_donor(struct rq *rq)
> > */
> > static inline bool proxy_needs_return(struct rq *rq, struct task_struct *p)
> > {
> > + bool reset_donor = false;
> > + bool dequeued;
> > +
> > /*
> > * Typically per __set_task_cpu(), task_cpu(p) == p->wake_cpu.
> > *
> > @@ -3797,11 +3807,17 @@ static inline bool proxy_needs_return(struct rq *rq, struct task_struct *p)
> > if (task_current(rq, p))
> > return false;
> >
> > - /* If we're return migrating the rq->donor, switch it out for idle */
> > - if (task_current_donor(rq, p))
> > - proxy_reset_donor(rq);
> > + reset_donor = task_current_donor(rq, p);
>
> nit.
>
> Since proxy_needs_return() holds the rq_lock, you can check this outside
> the blocked_lock safely, even after the dequeue. There is no need to
> stash "reset_donor".
Agreed, rq->donor is stable while the rq lock is held, so this can use
task_current_donor(rq, p) directly after dequeue_block_task() and remove
reset_donor.
>
> > }
> > - block_task(rq, p, TASK_WAKING);
> > +
> > + dequeued = dequeue_block_task(rq, p, TASK_WAKING);
> > +
> > + /* Keep on_rq set until all donor references have been replaced. */
> > + if (reset_donor)
> > + proxy_reset_donor(rq);
>
> Since TASK_WAKING is guaranteed to block the task by adding
> DEQUEUE_SPECIAL, you can just move the proxy_reset_donor() bit out the
> blocked lock and keep everything else the same right?
We still need to perform the sched-class dequeue before proxy_reset_donor().
If proxy_reset_donor() runs first, put_prev_task_scx() sees the donor with
SCX_TASK_QUEUED and is_blocked still set and reenqueues it through
scx_do_enqueue_task(). Then the following block_task() dequeues it again.
>
> I'm not sure I understand "transiently re-enqueuing the waking donor"
> bit. How is that possible if we just do:
>
> /* __task_rq_lock is held throughout. */
>
> if (task_current_donor(rq, p))
> proxy_reset_donor(rq, p);
>
> block_task(rq, p, TASK_WAKING);
And the transient reenqueue happens inside proxy_reset_donor():
proxy_reset_donor()
put_prev_set_next_task()
put_prev_task_scx()
scx_do_enqueue_task()
ops.enqueue()
__clear_task_blocked_on() clears the blocked_on relationship, but p->is_blocked
remains set until ttwu_do_wakeup(). As SCX_TASK_QUEUED is also still set,
put_prev_task_scx() takes its retained-blocked-donor path.
The subsequent block_task() would then immediately undo that enqueue.
> Is the split because dequeue_task_scx() needs a correct rq->donor
> reference or does ext requires dequeue_task_scx() to be called
> before doing put_prev_task_scx() always?
Both are aspects of the same ordering requirement here.
dequeue_task_scx() must run with rq->donor == p so it recognizes p as the
current scheduling context, emits the appropriate stopping transition and clears
SCX_TASK_QUEUED. Then put_prev_task_scx() can run as part of proxy_reset_donor()
without placing the blocked donor back into a DSQ or calling ops.enqueue()
again.
And this is specific to resetting a retained donor, it's not a general
requirement that dequeue_task_scx() always precede put_prev_task_scx().
I'll remove the unnecessary reset_donor variable and expand the comment to make
this ordering more explicit.
Thanks!
-Andrea
>
> > +
> > + if (dequeued)
> > + __block_task(rq, p);
> > return true;
> > }
> > #else /* !CONFIG_SCHED_PROXY_EXEC */
>
> --
> Thanks and Regards,
> Prateek
>
next prev parent reply other threads:[~2026-09-08 9:28 UTC|newest]
Thread overview: 42+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-31 13:42 [PATCHSET v13 sched_ext/for-7.4] sched: Make proxy execution compatible with sched_ext Andrea Righi
2026-08-31 13:42 ` [PATCH 01/18] sched/core: Drop mutex locks before proxy rescheduling Andrea Righi
2026-08-31 13:42 ` [PATCH 02/18] sched/core: Dequeue waking proxy donors before reset Andrea Righi
2026-09-01 5:24 ` K Prateek Nayak
2026-09-08 9:28 ` Andrea Righi [this message]
2026-08-31 13:42 ` [PATCH 03/18] sched: Make NOHZ CFS bandwidth checks follow proxy donor Andrea Righi
2026-09-10 9:54 ` Peter Zijlstra
2026-08-31 13:42 ` [PATCH 04/18] sched/core: Avoid false migration warning for proxy donors Andrea Righi
2026-09-10 10:06 ` Peter Zijlstra
2026-08-31 13:42 ` [PATCH 05/18] sched: Pass next class to sched_change_begin() Andrea Righi
2026-09-10 10:12 ` Peter Zijlstra
2026-08-31 13:42 ` [PATCH 06/18] sched: Add helper to block retained proxy donors Andrea Righi
2026-08-31 13:42 ` [PATCH 07/18] sched: Add sched_ext hooks for proxy execution Andrea Righi
2026-09-10 10:38 ` Peter Zijlstra
2026-08-31 13:42 ` [PATCH 08/18] sched: Introduce WF_ON_RQ wake flag Andrea Righi
2026-09-10 10:45 ` Peter Zijlstra
2026-08-31 13:42 ` [PATCH 09/18] sched_ext: Block proxy donors across scheduler transitions Andrea Righi
2026-09-10 10:53 ` Peter Zijlstra
2026-09-10 11:41 ` Peter Zijlstra
2026-08-31 13:42 ` [PATCH 10/18] sched_ext: Fix ops.running/stopping() pairing for proxy-exec donors Andrea Righi
2026-08-31 13:42 ` [PATCH 11/18] sched_ext: Move reject DSQ draining into core Andrea Righi
2026-08-31 13:42 ` [PATCH 12/18] sched_ext: Generalize the reject DSQ reenqueue path Andrea Righi
2026-09-03 22:39 ` Tejun Heo
2026-09-08 9:34 ` Andrea Righi
2026-08-31 13:42 ` [PATCH 13/18] sched_ext: Handle proxy-exec races in remote DSQ transfers Andrea Righi
2026-08-31 13:42 ` [PATCH 14/18] sched_ext: Split curr|donor references properly Andrea Righi
2026-08-31 17:49 ` sashiko-bot
2026-09-08 10:15 ` Andrea Righi
2026-09-10 11:47 ` Peter Zijlstra
2026-08-31 13:42 ` [PATCH 15/18] sched_ext: Delegate proxy donor admission to BPF schedulers Andrea Righi
2026-08-31 18:08 ` sashiko-bot
2026-09-08 10:08 ` Andrea Righi
2026-09-10 13:39 ` Peter Zijlstra
2026-09-10 13:41 ` Peter Zijlstra
2026-08-31 13:42 ` [PATCH 16/18] sched_ext: Add selftest for blocked donor admission Andrea Righi
2026-08-31 13:42 ` [PATCH 17/18] sched_ext: scx_qmap: Add proxy execution support Andrea Righi
2026-08-31 18:33 ` sashiko-bot
2026-09-01 7:52 ` Richard Cheng
2026-09-08 9:42 ` Andrea Righi
2026-08-31 13:42 ` [PATCH 18/18] sched: Allow enabling proxy exec with sched_ext Andrea Righi
2026-09-03 22:51 ` [PATCHSET v13 sched_ext/for-7.4] sched: Make proxy execution compatible " Tejun Heo
2026-09-08 8:02 ` Peter Zijlstra
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ap_VIJdG80ZZ_8D0@gpd4 \
--to=arighi@nvidia.com \
--cc=aiqun.yu@oss.qualcomm.com \
--cc=bsegall@google.com \
--cc=changwoo@igalia.com \
--cc=christian.loehle@arm.com \
--cc=david.dai@linux.dev \
--cc=dietmar.eggemann@arm.com \
--cc=etsal@meta.com \
--cc=icheng@nvidia.com \
--cc=jstultz@google.com \
--cc=juri.lelli@redhat.com \
--cc=kobak@nvidia.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-kernel@vger.kernel.org \
--cc=ltrager@nvidia.com \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=sched-ext@lists.linux.dev \
--cc=tj@kernel.org \
--cc=vincent.guittot@linaro.org \
--cc=void@manifault.com \
--cc=vschneid@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.