From: Andrea Righi <arighi@nvidia.com>
To: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
Changwoo Min <changwoo@igalia.com>
Cc: John Stultz <jstultz@google.com>,
sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: [PATCH v3 sched_ext/for-7.4] sched_ext: Keep proxy donors with slice left on the local DSQ
Date: Fri, 9 Oct 2026 22:08:26 +0200 [thread overview]
Message-ID: <20261009200827.4026499-1-arighi@nvidia.com> (raw)
Commit ee172227d0dc ("sched_ext: Delegate proxy donor admission to BPF
schedulers") makes put_prev_task_scx() pass a retained proxy donor to
ops.enqueue() with SCX_ENQ_BLOCKED.
Some of these puts are only proxy bookkeeping. proxy_resched_idle()
drops the rq's donor reference before switching to idle when:
- find_proxy_task() finds a remote mutex owner before the donor
switches out,
- proxy_migrate_task() detaches the donor before migration,
- proxy_deactivate() blocks a donor whose owner cannot run.
In the first case, BPF has already selected a donor with slice left.
Returning it to ops.enqueue() forces BPF to dispatch it again before
proxy resolution can continue. In the other two cases, the caller
deactivates the donor immediately after ops.enqueue(), undoing any
placement BPF makes.
Keep a donor with slice left at the head of the local DSQ so the next
pick can resolve its owner, or deactivation can remove it without an
unnecessary BPF handoff. An IMMED donor can stay there only for a
bookkeeping put during proxy resolution. Mark a blocked pick as
awaiting resolution to distinguish that put from a real preemption.
Return a preempted IMMED donor to BPF and apply the normal IMMED rule
during deferred local reenqueues. Use SCX_ENQ_IMMED on the temporary
local insertion so the base CPU capability is sufficient.
Moreover, allow blocked donors to receive SCX_ENQ_LAST when the
scheduler needs to arrange a follow-up scheduling event. During a
scheduler ownership change, mark the active donor's proxy reset in
sched_ext and skip reenqueueing it: sched_proxy_block_task() will
dequeue it next. This avoids a transient enqueue/dequeue pair and a
false ENQ_LAST warning.
For donors with slice left, this leaves BPF to handle meaningful
placement decisions rather than transient proxy-bookkeeping puts.
Fixes: ee172227d0dc ("sched_ext: Delegate proxy donor admission to BPF schedulers")
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
Changes in v3:
- Restore SCX_RQ_PROXY_PICK_PENDING so only proxy bookkeeping keeps an
IMMED donor local, real preemption and deferred reenqueues follow
the normal IMMED rules (Tejun Heo)
- Allow blocked donors to receive SCX_ENQ_LAST and use
SCX_RQ_PROXY_BLOCKING during scheduler ownership changes to avoid
transient enqueue/dequeue and false ENQ_LAST warning (Tejun Heo)
- Link to v2: https://lore.kernel.org/r/20261002221559.3090900-1-arighi@nvidia.com
kernel/sched/ext/ext.c | 78 +++++++++++++++++++++++--------------
kernel/sched/ext/internal.h | 16 +++++---
kernel/sched/sched.h | 2 +
3 files changed, 62 insertions(+), 34 deletions(-)
diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index 248b39d09ce3c..e168aadb5c032 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -24,6 +24,13 @@
DEFINE_RAW_SPINLOCK(scx_sched_lock);
+static void scx_block_proxy_donor(struct rq *rq, struct task_struct *p)
+{
+ rq->scx.flags |= SCX_RQ_PROXY_BLOCKING;
+ sched_proxy_block_task(rq, p);
+ rq->scx.flags &= ~SCX_RQ_PROXY_BLOCKING;
+}
+
bool __scx_allow_proxy_exec(const struct task_struct *p)
{
struct scx_sched *sch;
@@ -47,7 +54,7 @@ void scx_prepare_task_sched_change(struct task_struct *p)
lockdep_assert_rq_held(task_rq(p));
update_rq_clock(task_rq(p));
- sched_proxy_block_task(task_rq(p), p);
+ scx_block_proxy_donor(task_rq(p), p);
}
/*
@@ -1574,7 +1581,8 @@ static void dsq_inc_nr(struct scx_dispatch_q *dsq, struct task_struct *p, u64 en
/*
* If @rq already had other tasks or the current task is not
- * done yet, @p can't go on the CPU immediately. Re-enqueue.
+ * done yet, @p can't go on the CPU immediately. Check whether
+ * it needs to be reenqueued.
*/
if (unlikely(dsq->nr > 1 || !rq_is_open(rq, enq_flags)))
scx_schedule_reenq_local(rq, 0);
@@ -3318,6 +3326,9 @@ static enum scx_dsp_verdict dispatch_one(struct rq *rq, struct task_struct *prev
*
* - A non-IMMED HEAD task can get queued in front of an IMMED task
* between the IMMED queueing and the subsequent scheduling event.
+ *
+ * A blocked IMMED donor may make this scan a no-op. Skipping it does
+ * not schedule another scan, so avoid separate accounting for it.
*/
if (unlikely(rq->scx.local_dsq.nr > 1 && rq->scx.nr_immed))
scx_schedule_reenq_local(rq, 0);
@@ -3346,6 +3357,11 @@ static void set_next_task_scx(struct rq *rq, struct task_struct *p, enum snt_e t
bool first = type == SNT_PICK;
bool can_stop_tick;
+ rq->scx.flags &= ~SCX_RQ_PROXY_PICK_PENDING;
+ if (sched_proxy_exec() && p->is_blocked &&
+ (type == SNT_PICK || type == SNT_REPICK))
+ rq->scx.flags |= SCX_RQ_PROXY_PICK_PENDING;
+
if (type == SNT_REPICK)
return;
@@ -3425,6 +3441,7 @@ void scx_proxy_donor_start(struct rq *rq)
struct task_struct *donor = rq->donor;
lockdep_assert_rq_held(rq);
+ rq->scx.flags &= ~SCX_RQ_PROXY_PICK_PENDING;
if (donor->sched_class == &ext_sched_class && (donor->scx.flags & SCX_TASK_QUEUED))
scx_start_task_running(rq, donor);
@@ -3485,8 +3502,14 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
struct task_struct *next)
{
struct scx_sched *sch = scx_task_sched(p);
+ bool proxy_put = p->is_blocked && next == rq->idle &&
+ (rq->scx.flags & SCX_RQ_PROXY_PICK_PENDING);
+ bool proxy_block = p->is_blocked && next == rq->curr &&
+ (rq->scx.flags & SCX_RQ_PROXY_BLOCKING);
bool rescue_keep = false;
+ rq->scx.flags &= ~SCX_RQ_PROXY_PICK_PENDING;
+
/* see kick_sync_wait_bal_cb() */
smp_store_release(&rq->scx.kick_sync, rq->scx.kick_sync + 1);
@@ -3510,45 +3533,36 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
if (next != p && (p->scx.flags & SCX_TASK_QUEUED) &&
(p->scx.flags & SCX_TASK_RUN_TRACKED)) {
if (SCX_HAS_OP(sch, stopping))
- SCX_CALL_OP_TASK(sch, stopping, rq, p, true);
+ SCX_CALL_OP_TASK(sch, stopping, rq, p, !proxy_block);
p->scx.flags &= ~SCX_TASK_RUN_TRACKED;
}
if (p->scx.flags & SCX_TASK_QUEUED) {
- set_task_runnable(rq, p);
+ if (proxy_block)
+ goto switch_class;
- /* Delegate retained donor admission to its owning BPF scheduler. */
- if (p->is_blocked) {
- /*
- * If the donor is the same and only the mutex owner
- * changes, avoid triggering another ops.enqueue(): the
- * BPF scheduler has already admitted the donor, so it
- * can continue running.
- */
- if (next == p)
- goto switch_class;
+ set_task_runnable(rq, p);
- if (WARN_ON_ONCE(!sch))
- goto switch_class;
- WARN_ON_ONCE(!(sch->ops.flags & SCX_OPS_ENQ_BLOCKED));
- scx_do_enqueue_task(rq, p, 0, -1);
+ /*
+ * If the donor is the same and only the mutex owner changes,
+ * avoid triggering another ops.enqueue(): the BPF scheduler has
+ * already admitted the donor, so it can continue running.
+ */
+ if (p->is_blocked && next == p)
goto switch_class;
- }
/*
- * If @p has slice left and is being put, @p is getting
- * preempted by a higher priority scheduler class or core-sched
- * forcing a different task. Leave it at the head of the local
- * DSQ unless it was an IMMED task. IMMED tasks should not
- * linger on a busy CPU, reenqueue them to the BPF scheduler.
+ * If @p has slice left, keep it at the head of the local DSQ.
+ * An IMMED task instead returns to BPF if it is preempted, but a
+ * proxy put only drops references before resolving the owner.
*
* An open rescue must keep @p on the local DSQ even if the
* scheduler zeroed the slice in ops.stopping() above.
*/
if ((p->scx.slice || unlikely(p == scx_rescuee(rq))) &&
!scx_bypassing(sch, cpu_of(rq))) {
- if (p->scx.flags & SCX_TASK_IMMED) {
+ if ((p->scx.flags & SCX_TASK_IMMED) && !proxy_put) {
p->scx.flags |= SCX_TASK_REENQ_PREEMPTED;
scx_do_enqueue_task(rq, p, SCX_ENQ_REENQ, -1);
} else {
@@ -3565,6 +3579,9 @@ static void put_prev_task_scx(struct rq *rq, struct task_struct *p,
enq_flags |= SCX_ENQ_HEAD;
} else {
enq_flags |= SCX_ENQ_HEAD;
+ /* SCX_ENQ_IMMED uses SCX_CAP_ENQ_IMMED, the base cap. */
+ if (proxy_put && (p->scx.flags & SCX_TASK_IMMED))
+ enq_flags |= SCX_ENQ_IMMED;
}
scx_dispatch_enqueue(sch, rq, &rq->scx.local_dsq, p, 0, 0,
@@ -3756,6 +3773,9 @@ do_pick_task_scx(struct rq *rq, struct rq_flags *rf, bool force_scx)
enum scx_dsp_verdict verdict;
struct task_struct *p;
+ /* A retry can abandon a provisional blocked-donor pick. */
+ rq->scx.flags &= ~SCX_RQ_PROXY_PICK_PENDING;
+
/* see kick_sync_wait_bal_cb() */
smp_store_release(&rq->scx.kick_sync, rq->scx.kick_sync + 1);
@@ -4776,7 +4796,7 @@ void scx_prepare_setscheduler(struct task_struct *p, int policy)
lockdep_assert_rq_held(task_rq(p));
if (scx_enabled() && p->policy != policy && policy == SCHED_EXT)
- sched_proxy_block_task(task_rq(p), p);
+ scx_block_proxy_donor(task_rq(p), p);
}
static void process_ddsp_deferred_locals(struct rq *rq)
@@ -4823,9 +4843,9 @@ static void process_ddsp_deferred_locals(struct rq *rq)
* rq_is_open() is true.
*
* An IMMED task is kept (returns %false) only if it's the first task in the DSQ
- * AND the current task is done — i.e. it will execute immediately. All other
- * IMMED tasks are reenqueued. This means if a non-IMMED task sits at the head,
- * every IMMED task behind it gets reenqueued.
+ * AND the current task is done, so it will execute immediately. Other IMMED
+ * tasks are reenqueued. If a non-IMMED task sits at the head, every IMMED task
+ * behind it gets reenqueued.
*
* Reenqueued tasks go through ops.enqueue() with %SCX_ENQ_REENQ |
* %SCX_TASK_REENQ_IMMED. If the BPF scheduler dispatches back to the same local
diff --git a/kernel/sched/ext/internal.h b/kernel/sched/ext/internal.h
index 8dcab02a38ab1..d61eb88ea8d74 100644
--- a/kernel/sched/ext/internal.h
+++ b/kernel/sched/ext/internal.h
@@ -1816,14 +1816,14 @@ enum scx_enq_flags {
SCX_ENQ_PREEMPT_LAZY = 1LLU << 35,
/*
- * Only allowed on local DSQs. Guarantees that the task either gets
- * on the CPU immediately and stays on it, or gets reenqueued back
- * to the BPF scheduler. It will never linger on a local DSQ or be
- * silently put back after preemption.
+ * Only allowed on local DSQs. Guarantees that an unblocked task either
+ * gets on the CPU immediately and stays on it, or gets reenqueued back
+ * to the BPF scheduler. A blocked proxy donor can stay on the local DSQ
+ * with slice left because its progress depends on its mutex owner.
*
* The protection persists until the next fresh enqueue - it
* survives SAVE/RESTORE cycles, slice extensions and preemption.
- * If the task can't stay on the CPU for any reason, it gets
+ * If an unblocked task can't stay on the CPU for any reason, it gets
* reenqueued back to the BPF scheduler.
*
* Exiting and migration-disabled tasks bypass ops.enqueue() and
@@ -1865,6 +1865,12 @@ enum scx_enq_flags {
/*
* The task is blocked on a mutex and is being kept runnable as a proxy
* donor. Only passed to ops.enqueue() when %SCX_OPS_ENQ_BLOCKED is set.
+ *
+ * Blocking on the mutex does not enqueue the task by itself. A donor put
+ * with slice left stays at the head of the local DSQ, except when an
+ * IMMED donor is preempted or cannot stay local. It is passed to
+ * ops.enqueue() when its slice runs out, the IMMED placement cannot be
+ * kept, or proxy execution moves it to the owner's CPU.
*/
SCX_ENQ_BLOCKED = 1LLU << 42,
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index fc65296f5c745..d6813eccd7f74 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -793,6 +793,8 @@ enum scx_rq_flags {
SCX_RQ_ROOT_IDLE_RENOTIFY = 1 << 8, /* the root is owed update_idle() */
SCX_RQ_PROXY_RETRY = 1 << 9, /* proxy-rejected tasks need retry */
SCX_RQ_PROXY_TICK = 1 << 10, /* proxy execution requires the tick */
+ SCX_RQ_PROXY_PICK_PENDING = 1 << 11, /* blocked donor awaits proxy resolution */
+ SCX_RQ_PROXY_BLOCKING = 1 << 12, /* donor is being removed for sched change */
SCX_RQ_IN_WAKEUP = 1 << 16,
SCX_RQ_IN_DISPATCH = 1 << 17,
--
2.56.0
next reply other threads:[~2026-10-09 20:08 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-09 20:08 Andrea Righi [this message]
2026-10-09 23:01 ` [PATCH v3 sched_ext/for-7.4] sched_ext: Keep proxy donors with slice left on the local DSQ Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261009200827.4026499-1-arighi@nvidia.com \
--to=arighi@nvidia.com \
--cc=changwoo@igalia.com \
--cc=jstultz@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=sched-ext@lists.linux.dev \
--cc=tj@kernel.org \
--cc=void@manifault.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox