From: Kuba Piecuch <jpiecuch@google.com>
To: Tejun Heo <tj@kernel.org>, Andrea Righi <arighi@nvidia.com>,
Changwoo Min <changwoo@igalia.com>,
David Vernet <void@manifault.com>
Cc: linux-kernel@vger.kernel.org, sched-ext@lists.linux.dev,
Kuba Piecuch <jpiecuch@google.com>
Subject: [PATCH v2 sched_ext/for-7.3-fixes] sched_ext: Generate qseq from a per-task counter
Date: Sat, 3 Oct 2026 11:53:17 +0000 [thread overview]
Message-ID: <20261003115317.43001-1-jpiecuch@google.com> (raw)
finish_dispatch() uses the qseq embedded in p->scx.ops_state to tell
whether the QUEUED instance of a task it's about to claim is the one
scx_bpf_dsq_insert() saw. qseq is generated from rq->scx.ops_qseq, but
the counters of different rqs are independent, so if a task is dequeued
and re-enqueued on a different rq between scx_bpf_dsq_insert() and
finish_dispatch(), the new QUEUED instance can end up with the same
qseq as the old one:
CPU X CPU Z
----- -----
enqueue p on rq A, qseq = N
ops.dispatch()
scx_bpf_dsq_insert(p)
records qseq N
sched_setaffinity(p)
dequeue p from rq A
enqueue p on rq B, qseq = N
finish_dispatch(p, N)
qseq matches, p is claimed
The claim itself is still atomic so the core stays consistent, but an
insert issued for a previous QUEUED instance gets applied to a new one
which the BPF scheduler has just received through ops.enqueue(). This
breaks the guarantee that dispatches targeting a stale instance are
ignored.
Generate qseq from a per-task counter, p->scx.ops_qseq, instead so that
consecutive QUEUED instances of a task never share a qseq regardless of
which rq they're on. The counter is only updated in
scx_do_enqueue_task() with the task's rq locked, so no additional
synchronization is needed, and it fits in an existing hole in struct
sched_ext_entity on 64bit. Remove the now unused rq->scx.ops_qseq.
Never generate qseq 0. NONE and DISPATCHING don't carry a qseq, so
scx_bpf_dsq_insert() on a task in either state records 0. With a
per-task counter, every task's first QUEUED instance would otherwise get
qseq 0 and could be claimed by such an insert. Wrap the counter where
the QSEQ field wraps so that it can't reach a value that shifts to 0 on
32bit either.
Fixes: f0e1a0643a59 ("sched_ext: Implement BPF extensible scheduler class")
Assisted-by: Claude:claude-opus-5.5
Signed-off-by: Kuba Piecuch <jpiecuch@google.com>
---
v2:
- Wrap the counter where the QSEQ field wraps instead of looping to
skip 0 (Tejun).
v1: https://lore.kernel.org/all/20261002205242.3820674-1-jpiecuch@google.com/
include/linux/sched/ext.h | 1 +
kernel/sched/ext/ext.c | 16 ++++++++++++----
kernel/sched/ext/internal.h | 5 +++++
kernel/sched/sched.h | 1 -
4 files changed, 18 insertions(+), 5 deletions(-)
diff --git a/include/linux/sched/ext.h b/include/linux/sched/ext.h
index 582d7cd4a983..36c04797436c 100644
--- a/include/linux/sched/ext.h
+++ b/include/linux/sched/ext.h
@@ -207,6 +207,7 @@ struct sched_ext_entity {
s32 holding_cpu;
s32 selected_cpu;
s32 runnable_cpu; /* cpu @p is runnable on, -1 if not */
+ u32 ops_qseq; /* protected by rq lock */
struct task_struct *kf_tasks[2]; /* see SCX_CALL_OP_TASK() */
struct list_head runnable_node; /* rq->scx.runnable_list */
diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index e56c3c95018f..5acb5c325b5e 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -2057,8 +2057,16 @@ void scx_do_enqueue_task(struct rq *rq, struct task_struct *p, u64 enq_flags,
if (unlikely(!SCX_HAS_OP(sch, enqueue)))
goto global;
- /* DSQ bypass didn't trigger, enqueue on the BPF scheduler */
- qseq = rq->scx.ops_qseq++ << SCX_OPSS_QSEQ_SHIFT;
+ /*
+ * DSQ bypass didn't trigger, enqueue on the BPF scheduler.
+ * Brand this QUEUED instance with a fresh per-task qseq.
+ * Wrap the counter where the QSEQ sub-field of ops_state wraps
+ * and skip 0 as that's what scx_bpf_dsq_insert() records
+ * for a task in NONE or DISPATCHING.
+ */
+ p->scx.ops_qseq = ((p->scx.ops_qseq + 1) &
+ (SCX_OPSS_QSEQ_MASK >> SCX_OPSS_QSEQ_SHIFT)) ?: 1;
+ qseq = (unsigned long)p->scx.ops_qseq << SCX_OPSS_QSEQ_SHIFT;
WARN_ON_ONCE(atomic_long_read(&p->scx.ops_state) != SCX_OPSS_NONE);
atomic_long_set(&p->scx.ops_state, SCX_OPSS_QUEUEING | qseq);
@@ -6947,9 +6955,9 @@ static void scx_dump_cpu(struct scx_sched *sch, struct seq_buf *s,
seq_buf_init(&ns, buf, avail);
dump_newline(&ns);
- scx_dump_line(&ns, "CPU %-4d: nr_run=%u flags=0x%x cpu_rel=%d ops_qseq=%lu ksync=%lu",
+ scx_dump_line(&ns, "CPU %-4d: nr_run=%u flags=0x%x cpu_rel=%d ksync=%lu",
cpu, rq->scx.nr_running, rq->scx.flags, rq->scx.cpu_released,
- rq->scx.ops_qseq, rq->scx.kick_sync);
+ rq->scx.kick_sync);
scx_rescue_dump(&ns, rq);
scx_dump_line(&ns, " curr=%s[%d] class=%ps",
rq->curr->comm, rq->curr->pid, rq->curr->sched_class);
diff --git a/kernel/sched/ext/internal.h b/kernel/sched/ext/internal.h
index 1df8f583b0ec..5b37faa532d6 100644
--- a/kernel/sched/ext/internal.h
+++ b/kernel/sched/ext/internal.h
@@ -1998,6 +1998,11 @@ enum scx_ops_state {
* dequeue/requeue, the dispatcher can tell whether it still has a claim
* on the task being dispatched.
*
+ * QSEQ is generated from the per-task p->scx.ops_qseq counter so that
+ * it doesn't repeat across QUEUED instances of the same task even if
+ * the task moves between rqs. 0 is never used as a valid QSEQ since
+ * NONE and DISPATCHING map to this value.
+ *
* As some 32bit archs can't do 64bit store_release/load_acquire,
* p->scx.ops_state is atomic_long_t which leaves 30 bits for QSEQ on
* 32bit machines. The dispatch race window QSEQ protects is very narrow
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index e656c7059bf8..37df14377516 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -816,7 +816,6 @@ struct scx_rq {
#endif
struct list_head runnable_list; /* runnable tasks on this rq */
struct list_head ddsp_deferred_locals; /* deferred ddsps from enq */
- unsigned long ops_qseq;
/* both stashed across the activate_task() in move_remote_task_to_local_dsq() */
u64 remote_activate_enq_flags;
struct scx_sched *remote_activate_sch;
--
2.56.0.rc1.315.gc6ed9934b7-goog
next reply other threads:[~2026-10-03 11:53 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-03 11:53 Kuba Piecuch [this message]
2026-10-03 23:41 ` [PATCH v2 sched_ext/for-7.3-fixes] sched_ext: Generate qseq from a per-task counter Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261003115317.43001-1-jpiecuch@google.com \
--to=jpiecuch@google.com \
--cc=arighi@nvidia.com \
--cc=changwoo@igalia.com \
--cc=linux-kernel@vger.kernel.org \
--cc=sched-ext@lists.linux.dev \
--cc=tj@kernel.org \
--cc=void@manifault.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.