Sched_ext development
 help / color / mirror / Atom feed
* [PATCH sched_ext/for-7.3-fixes] sched_ext: Fix use-after-free of a destroyed user DSQ by a deferred reenqueue
@ 2026-10-09 22:01 Tejun Heo
  2026-10-09 22:09 ` sashiko-bot
                   ` (2 more replies)
  0 siblings, 3 replies; 5+ messages in thread
From: Tejun Heo @ 2026-10-09 22:01 UTC (permalink / raw)
  To: David Vernet, Andrea Righi, Changwoo Min
  Cc: Hui Su, Emil Tsalapatis, David Dai, Cheng-Yang Chou, sched-ext,
	linux-kernel

A deferred reenqueue can read a freed user DSQ. The request stays on its
CPU's deferred list until run_deferred() takes it, and the RCU callback that
frees the DSQ assumes that the grace period has flushed it, so exit_dsq()
checks the per-CPU node without the deferred list lock. A grace period waits
for readers in flight, not for a pending run_deferred(). A run_deferred()
which starts after the grace period began can detach the request right
before the callback looks and then read the freed DSQ's id.

Cancel the pending requests before the grace period instead. The irq work
that frees the DSQ unlinks every CPU's request under the per-CPU deferred
list lock, and schedule_dsq_reenq() ignores a new request under the same
lock once the id is invalid. Only a run_deferred() that started before the
grace period can hold a request, and it runs with IRQs disabled, so the
grace period waits for it and nothing touches the DSQ after the callback.

Fixes: 84b1a0ea0b7c ("sched_ext: Implement scx_bpf_dsq_reenq() for user DSQs")
Cc: stable@vger.kernel.org # v7.1+
Reported-by: Hui Su <sh_def@163.com>
Link: https://lore.kernel.org/all/20261008024753.4096008-1-sh_def@163.com/
Signed-off-by: Tejun Heo <tj@kernel.org>
---
 kernel/sched/ext/ext.c |   28 +++++++++++++++++++++++++---
 1 file changed, 25 insertions(+), 3 deletions(-)

--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -1149,6 +1149,10 @@ void schedule_dsq_reenq(struct scx_sched
 
 			guard(raw_spinlock_irqsave)(&rq->scx.deferred_reenq_lock);
 
+			/* @dsq is being destroyed, see free_dsq_irq_workfn() */
+			if (unlikely(READ_ONCE(dsq->id) == SCX_DSQ_INVALID))
+				return;
+
 			if (list_empty(&dru->node))
 				list_move_tail(&dru->node, &rq->scx.deferred_reenq_users);
 			WRITE_ONCE(dru->flags, dru->flags | reenq_flags);
@@ -5171,8 +5175,8 @@ static void exit_dsq(struct scx_dispatch
 		struct rq *rq = cpu_rq(cpu);
 
 		/*
-		 * There must have been a RCU grace period since the last
-		 * insertion and @dsq should be off the deferred list by now.
+		 * free_dsq_irq_workfn() cancelled the reenqs before the grace
+		 * period and schedule_dsq_reenq() queues nothing after that.
 		 */
 		if (WARN_ON_ONCE(!list_empty(&dru->node))) {
 			guard(raw_spinlock_irqsave)(&rq->scx.deferred_reenq_lock);
@@ -5196,8 +5200,26 @@ static void free_dsq_irq_workfn(struct i
 	struct llist_node *to_free = llist_del_all(&dsqs_to_free);
 	struct scx_dispatch_q *dsq, *tmp_dsq;
 
-	llist_for_each_entry_safe(dsq, tmp_dsq, to_free, free_node)
+	llist_for_each_entry_safe(dsq, tmp_dsq, to_free, free_node) {
+		s32 cpu;
+
+		/*
+		 * Cancel pending reenqs. Only a run_deferred() that started
+		 * before the grace period can hold one, and it keeps IRQs off,
+		 * so the grace period waits for it. After this sweep,
+		 * schedule_dsq_reenq() sees the invalid id under the same lock
+		 * and queues nothing.
+		 */
+		for_each_possible_cpu(cpu) {
+			struct scx_dsq_pcpu *pcpu = per_cpu_ptr(dsq->pcpu, cpu);
+			struct rq *rq = cpu_rq(cpu);
+
+			guard(raw_spinlock_irqsave)(&rq->scx.deferred_reenq_lock);
+			list_del_init(&pcpu->deferred_reenq_user.node);
+		}
+
 		call_rcu(&dsq->rcu, free_dsq_rcufn);
+	}
 }
 
 static DEFINE_IRQ_WORK(free_dsq_irq_work, free_dsq_irq_workfn);

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-10-10  8:38 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-09 22:01 [PATCH sched_ext/for-7.3-fixes] sched_ext: Fix use-after-free of a destroyed user DSQ by a deferred reenqueue Tejun Heo
2026-10-09 22:09 ` sashiko-bot
2026-10-09 22:40   ` Tejun Heo
2026-10-10  6:23 ` [PATCH] " Hui Su
2026-10-10  8:38 ` [PATCH sched_ext/for-7.3-fixes] " Tejun Heo

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox