From: Andrea Righi <arighi@nvidia.com>
To: Tejun Heo <tj@kernel.org>
Cc: David Vernet <void@manifault.com>,
Changwoo Min <changwoo@igalia.com>,
John Stultz <jstultz@google.com>, Ingo Molnar <mingo@redhat.com>,
Peter Zijlstra <peterz@infradead.org>,
Juri Lelli <juri.lelli@redhat.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
Christian Loehle <christian.loehle@arm.com>,
David Dai <david.dai@linux.dev>, Koba Ko <kobak@nvidia.com>,
Aiqun Yu <aiqun.yu@oss.qualcomm.com>,
Shuah Khan <shuah@kernel.org>,
sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: Re: [PATCH 12/15] sched_ext: Delegate proxy donor admission to BPF schedulers
Date: Thu, 6 Aug 2026 08:10:09 +0200 [thread overview]
Message-ID: <anQlQUOj4PO7XEvy@gpd4> (raw)
In-Reply-To: <anETyb93XUqBau6Y@slm.duckdns.org>
Hi Tejun,
On Mon, Aug 03, 2026 at 12:18:49PM -1000, Tejun Heo wrote:
> On Tue, Jul 28, 2026 at 05:43:30PM +0200, Andrea Righi wrote:
> ...
> > +/*
> > + * Called with @p's pi and rq locks held immediately before
> > + * sched_change_begin(). The caller must pass DEQUEUE_NOCLOCK so the rq clock
> > + * is updated only once.
> > + */
> > +void scx_prepare_task_sched_change(struct task_struct *p, struct scx_sched *sch)
> > +{
> > + lockdep_assert_held(&p->pi_lock);
> > + lockdep_assert_rq_held(task_rq(p));
> > +
> > + update_rq_clock(task_rq(p));
> > +
> > + /* Block retained donors that the incoming scheduler cannot manage. */
> > + if (!(sch->ops.flags & SCX_OPS_ENQ_BLOCKED))
> > + sched_proxy_block_task(task_rq(p), p);
> > }
>
> What are the cases that this one catches that scx_allow_proxy_exec() or
> prepare_switch_scx() doesn't?
scx_allow_proxy_exec() controls whether a task is retained when it first blocks
in __schedule(), it doesn't handle a donor that was already retained before its
scheduler ownership changes.
prepare_switch_scx() handles a scheduling-class transition into EXT, but it
isn't called for an EXT-to-EXT scheduler change. The helper was intended to
cover these same-class transitions, i.e., moving between parent and child
sub-schedulers.
Thinking more about this, retained proxy execution can be terminated on all
class changes centrally in sched_change_begin(). In this way we can remove the
.prepare_switch() class callback and prepare_switch_scx(). We would still need
to terminate retained proxy execution explicitly for EXT-to-EXT scheduler
ownership changes, but that shouldn't be an issue. I'll test this approach, it
should simplify the transition handling considerably.
>
> > @@ -2299,11 +2351,24 @@ static void wakeup_preempt_scx(struct rq *rq, struct task_struct *p, int wake_fl
> > {
> > /*
> > * Preemption between SCX tasks is implemented by resetting the victim
> > - * task's slice to 0 and triggering reschedule on the target CPU.
> > - * Nothing to do.
> > + * task's slice to 0 and triggering reschedule on the target CPU. A
> > + * mutex-blocked task is kept queued for proxy execution, so its wakeup
> > + * doesn't go through enqueue_task_scx(). If the BPF scheduler manages
> > + * blocked donors, reschedule explicitly so that it can reconsider a
> > + * donor it declined to dispatch while blocked.
>
> Can you make this a separate paragraph and is the comment uptodate? I'm
> having a difficulty understanding what "if the BPF scheduler manages blocked
> donors" mean.
"manages blocked donors" means the BPF scheduler sets SCX_OPS_ENQ_BLOCKED. I'll
split the comment and clarify it.
>
> > */
> > - if (p->sched_class == &ext_sched_class)
> > + if (p->sched_class == &ext_sched_class) {
> > + bool enq_wakeup = p->scx.flags & SCX_TASK_ENQ_WAKEUP;
> > +
> > + p->scx.flags &= ~SCX_TASK_ENQ_WAKEUP;
> > + if (!enq_wakeup && p->is_blocked) {
> > + struct scx_sched *sch = scx_task_sched(p);
> > +
> > + if (sch && (sch->ops.flags & SCX_OPS_ENQ_BLOCKED))
> > + resched_curr(rq);
> > + }
> > return;
> > + }
>
> My understanding of what happens here is hazy. I suppose this is for the
> case of an active proxy execution being preempted by another SCX task? I'm
> not following why resched_curr() is needed here.
The relevant case is a mutex waiter receiving a wakeup while it's retained on
the rq as a proxy donor. Although the task is basically blocked and cannot
execute itself, its scheduling context remains on the rq, so that the mutex
owner can execute through it.
There are two wakeup paths when the mutex is released:
1) If the donor was not proxy-migrated and is still on its callback rq,
ttwu_runnable() handles the wakeup while the task remains on the rq. It
calls wakeup_preempt() and then clears p->is_blocked, without calling
enqueue_task_scx(). BPF doesn't receive any new ops.enqueue() notification.
resched_curr() requests another scheduling cycle so that ops.dispatch() can
reconsider the now-unblocked task (BPF scheduler may have kept the blocked
donor in a BPF-managed queue).
2) If the donor was proxy-migrated to the owner's rq, proxy_needs_return()
removes it from that rq and the wakeup proceeds through the full activation
path. That path calls enqueue_task_scx() before wakeup_preempt().
SCX_TASK_ENQ_WAKEUP records that this enqueue already happened, preventing
the additional resched_curr().
>
> > @@ -3198,6 +3279,37 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
> > if (p->scx.flags & SCX_TASK_QUEUED) {
> > set_task_runnable(rq, p);
> >
> > + /*
> > + * The rq lock has remained held since scx_allow_proxy_exec(), so
> > + * @p's scheduler association cannot have changed. An associated
> > + * donor stays queued only when its BPF scheduler enables
> > + * %SCX_OPS_ENQ_BLOCKED; delegate its admission to that scheduler.
> > + *
> > + * If @sch is NULL, @p is transitioning into the root scheduler. The
> > + * root is published before tasks enter EXT and cannot be cleared while
> > + * this rq is locked. Preserve generic proxy execution by placing the
> > + * donor directly on the local DSQ.
> > + */
> > + if (p->is_blocked) {
> > + /*
> > + * If the donor is the same and only the mutex owner
> > + * changes, avoid triggering another ops.enqueue(): the
> > + * BPF scheduler has already admitted the donor, so it
> > + * can continue running.
> > + */
> > + if (next == p)
> > + goto switch_class;
> > +
> > + if (sch) {
> > + WARN_ON_ONCE(!(sch->ops.flags & SCX_OPS_ENQ_BLOCKED));
> > + scx_do_enqueue_task(rq, p, 0, -1);
> > + } else {
> > + scx_dispatch_enqueue(scx_root, rq, &rq->scx.local_dsq,
> > + p, 0);
>
> Does this else arm actually happen? Can you describe the scenario? Oh, maybe
> below is the counterpart.
Right, this was intended for the root-enable transition where a task could
already be on the EXT class but not yet have an associated BPF scheduler. In
that window scx_allow_proxy_exec() permits generic proxy execution and sch can
be NULL.
IF we call sched_proxy_block_task() unconditionally, root enable will block any
retained donor before establishing the new scheduler ownership, so the
NULL-scheduler fallback is then unnecessary and we can remove it.
>
> > @@ -7758,6 +7875,10 @@ static void scx_root_enable_workfn(struct kthread_work *work)
> >
> > if (old_class != new_class)
> > queue_flags |= DEQUEUE_CLASS;
> > + if (old_class == new_class && new_class == &ext_sched_class) {
> > + scx_prepare_task_sched_change(p, sch);
> > + queue_flags |= DEQUEUE_NOCLOCK;
> > + }
>
> I'd appreciate if there's more explanation of what happens during enable.
> Wouldn't it be simpler if we just do sched_proxy_block_task() on all
> transitions and start with a clean slate?
Agreed. I'll change the logic so that a proxy session never survives a
scheduling-class or BPF-scheduler ownership transition.
For scheduling-class changes, we can call sched_proxy_block_task() centrally
from sched_change_begin(), this makes the new .prepare_switch() sched-class
callback and prepare_switch_scx() unnecessary.
Thanks,
-Andrea
next prev parent reply other threads:[~2026-08-06 6:10 UTC|newest]
Thread overview: 29+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-28 15:43 [PATCHSET v10 sched_ext/for-7.3] sched: Make proxy execution compatible with sched_ext Andrea Righi
2026-07-28 15:43 ` [PATCH 01/15] sched/core: Avoid false migration warning for proxy donors Andrea Righi
2026-07-28 15:43 ` [PATCH 02/15] sched: Make NOHZ CFS bandwidth checks follow proxy donor Andrea Righi
2026-07-28 15:43 ` [PATCH 03/15] sched: Add helper to block retained proxy donors Andrea Righi
2026-07-28 15:43 ` [PATCH 04/15] sched: Skip class callbacks with SCHED_FLAG_KEEP_PARAMS Andrea Righi
2026-07-28 15:43 ` [PATCH 05/15] sched: Add prepare_switch() class callback Andrea Righi
2026-07-28 15:43 ` [PATCH 06/15] sched: Add sched_ext hooks for proxy execution Andrea Righi
2026-07-28 15:43 ` [PATCH 07/15] sched_ext: Block proxy donors across scheduler transitions Andrea Righi
2026-08-03 20:36 ` Tejun Heo
2026-08-05 7:02 ` Andrea Righi
2026-08-06 9:07 ` Andrea Righi
2026-07-28 15:43 ` [PATCH 08/15] sched_ext: Fix ops.running/stopping() pairing for proxy-exec donors Andrea Righi
2026-07-28 15:43 ` [PATCH 09/15] sched_ext: Generalize the reject DSQ reenqueue path Andrea Righi
2026-08-03 20:35 ` Tejun Heo
2026-08-03 20:38 ` Tejun Heo
2026-08-05 8:50 ` Andrea Righi
2026-07-28 15:43 ` [PATCH 10/15] sched_ext: Handle proxy-exec races in remote DSQ transfers Andrea Righi
2026-08-03 21:36 ` Tejun Heo
2026-08-05 16:44 ` Andrea Righi
2026-07-28 15:43 ` [PATCH 11/15] sched_ext: Split curr|donor references properly Andrea Righi
2026-07-28 15:43 ` [PATCH 12/15] sched_ext: Delegate proxy donor admission to BPF schedulers Andrea Righi
2026-08-03 22:18 ` Tejun Heo
2026-08-06 6:10 ` Andrea Righi [this message]
2026-07-28 15:43 ` [PATCH 13/15] sched_ext: Add selftest for blocked donor admission Andrea Righi
2026-07-28 15:43 ` [PATCH 14/15] sched_ext: scx_qmap: Add proxy execution support Andrea Righi
2026-08-03 22:22 ` Tejun Heo
2026-08-06 7:28 ` Andrea Righi
2026-07-28 15:43 ` [PATCH 15/15] sched: Allow enabling proxy exec with sched_ext Andrea Righi
2026-08-03 22:19 ` Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anQlQUOj4PO7XEvy@gpd4 \
--to=arighi@nvidia.com \
--cc=aiqun.yu@oss.qualcomm.com \
--cc=bsegall@google.com \
--cc=changwoo@igalia.com \
--cc=christian.loehle@arm.com \
--cc=david.dai@linux.dev \
--cc=dietmar.eggemann@arm.com \
--cc=jstultz@google.com \
--cc=juri.lelli@redhat.com \
--cc=kobak@nvidia.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=sched-ext@lists.linux.dev \
--cc=shuah@kernel.org \
--cc=tj@kernel.org \
--cc=vincent.guittot@linaro.org \
--cc=void@manifault.com \
--cc=vschneid@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox