* [PATCH v3 1/4] sched/debug: Protect lockless rq->curr access in print_cpu()
2026-08-08 23:55 [PATCH v3 0/4] sched/debug: Introduce per-CPU debugfs files Aaron Tomlin
@ 2026-08-08 23:55 ` Aaron Tomlin
2026-08-08 23:55 ` [PATCH v3 2/4] sched/debug: Protect lockless rq->rd access in print_dl_rq() Aaron Tomlin
` (3 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Aaron Tomlin @ 2026-08-08 23:55 UTC (permalink / raw)
To: mingo, peterz, juri.lelli, vincent.guittot
Cc: dietmar.eggemann, rostedt, bsegall, mgorman, vschneid,
kprateek.nayak, zhanxusheng1024, neelx, chjohnst, mproche, sean,
steve, linux-kernel
In print_cpu(), rq->curr is dereferenced locklessly to print the current
task's PID via task_pid_nr(rq->curr).
While accessing /sys/kernel/debug/sched/debug is inherently best-effort
only; rq->curr is indeed expected to change dynamically while
print_cpu() is executing. However, if the task currently running on the
CPU exits concurrently and its reference count drops to zero,
put_task_struct() calls call_rcu() to schedule
__put_task_struct_rcu_cb(). Because print_cpu() does not hold the RCU
read lock while dereferencing rq->curr, an RCU grace period can complete
concurrently and free the task structure via free_task(), creating a
potential use-after-free race condition.
Resolve this by reading rq->curr using READ_ONCE() inside an RCU
read-side critical section. Holding the RCU read lock delays the
invocation of __put_task_struct_rcu_cb() until after rcu_read_unlock(),
ensuring that the struct task_struct memory remains valid while being
accessed.
Fixes: 02968ccf7b80 ("sched: add /proc/sched_debug file")
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Signed-off-by: Aaron Tomlin <atomlin@atomlin.com>
---
kernel/sched/debug.c | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c
index 40584b27ea0c..78fc02d71710 100644
--- a/kernel/sched/debug.c
+++ b/kernel/sched/debug.c
@@ -1126,7 +1126,10 @@ do { \
P(nr_switches);
P(nr_uninterruptible);
PN(next_balance);
- SEQ_printf(m, " .%-30s: %ld\n", "curr->pid", (long)(task_pid_nr(rq->curr)));
+ rcu_read_lock();
+ SEQ_printf(m, " .%-30s: %ld\n", "curr->pid",
+ (long)(task_pid_nr(READ_ONCE(rq->curr))));
+ rcu_read_unlock();
PN(clock);
PN(clock_task);
#undef P
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* [PATCH v3 2/4] sched/debug: Protect lockless rq->rd access in print_dl_rq()
2026-08-08 23:55 [PATCH v3 0/4] sched/debug: Introduce per-CPU debugfs files Aaron Tomlin
2026-08-08 23:55 ` [PATCH v3 1/4] sched/debug: Protect lockless rq->curr access in print_cpu() Aaron Tomlin
@ 2026-08-08 23:55 ` Aaron Tomlin
2026-08-08 23:55 ` [PATCH v3 3/4] sched/fair: Use list_for_each_entry_rcu() in print_cfs_stats() Aaron Tomlin
` (2 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Aaron Tomlin @ 2026-08-08 23:55 UTC (permalink / raw)
To: mingo, peterz, juri.lelli, vincent.guittot
Cc: dietmar.eggemann, rostedt, bsegall, mgorman, vschneid,
kprateek.nayak, zhanxusheng1024, neelx, chjohnst, mproche, sean,
steve, linux-kernel
In print_dl_rq(), cpu_rq(cpu)->rd is dereferenced locklessly to display
deadline bandwidth statistics.
During CPU hot-unplug or cgroup cpuset repartitioning events,
partition_sched_domains() calls cpu_attach_domain(), which executes
rq_attach_root() to detach the CPU from its root_domain. When the
reference count of the detached root_domain drops to zero,
rq_attach_root() calls call_rcu(&old_rd->rcu, free_rootdomain) to
schedule memory teardown after an RCU grace period.
Because print_dl_rq() does not hold an RCU read lock while dereferencing
cpu_rq(cpu)->rd, an RCU grace period can elapse concurrently while
debugfs is reading the file. This allows free_rootdomain() to execute
kfree(old_rd), introducing a use-after-free race condition when
print_dl_rq() reads dl_bw->bw.
Fix this by fetching rq->rd using READ_ONCE() inside an RCU read-side
critical section. Holding the RCU read lock guarantees that the struct
root_domain memory remains valid while being accessed.
Fixes: 02968ccf7b80 ("sched: add /proc/sched_debug file")
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Signed-off-by: Aaron Tomlin <atomlin@atomlin.com>
---
kernel/sched/debug.c | 12 +++++++++---
1 file changed, 9 insertions(+), 3 deletions(-)
diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c
index 78fc02d71710..e9d40a660346 100644
--- a/kernel/sched/debug.c
+++ b/kernel/sched/debug.c
@@ -1081,6 +1081,7 @@ void print_rt_rq(struct seq_file *m, int cpu, struct rt_rq *rt_rq)
void print_dl_rq(struct seq_file *m, int cpu, struct dl_rq *dl_rq)
{
struct dl_bw *dl_bw;
+ struct root_domain *rd;
SEQ_printf(m, "\n");
SEQ_printf(m, "dl_rq[%d]:\n", cpu);
@@ -1089,9 +1090,14 @@ void print_dl_rq(struct seq_file *m, int cpu, struct dl_rq *dl_rq)
SEQ_printf(m, " .%-30s: %lu\n", #x, (unsigned long)(dl_rq->x))
PU(dl_nr_running);
- dl_bw = &cpu_rq(cpu)->rd->dl_bw;
- SEQ_printf(m, " .%-30s: %lld\n", "dl_bw->bw", dl_bw->bw);
- SEQ_printf(m, " .%-30s: %lld\n", "dl_bw->total_bw", dl_bw->total_bw);
+ rcu_read_lock();
+ rd = READ_ONCE(cpu_rq(cpu)->rd);
+ if (rd) {
+ dl_bw = &rd->dl_bw;
+ SEQ_printf(m, " .%-30s: %lld\n", "dl_bw->bw", dl_bw->bw);
+ SEQ_printf(m, " .%-30s: %lld\n", "dl_bw->total_bw", dl_bw->total_bw);
+ }
+ rcu_read_unlock();
#undef PU
}
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* [PATCH v3 3/4] sched/fair: Use list_for_each_entry_rcu() in print_cfs_stats()
2026-08-08 23:55 [PATCH v3 0/4] sched/debug: Introduce per-CPU debugfs files Aaron Tomlin
2026-08-08 23:55 ` [PATCH v3 1/4] sched/debug: Protect lockless rq->curr access in print_cpu() Aaron Tomlin
2026-08-08 23:55 ` [PATCH v3 2/4] sched/debug: Protect lockless rq->rd access in print_dl_rq() Aaron Tomlin
@ 2026-08-08 23:55 ` Aaron Tomlin
2026-08-08 23:55 ` [PATCH v3 4/4] sched/debug: Introduce per-CPU debugfs files Aaron Tomlin
2026-08-10 3:17 ` [PATCH v3 0/4] " Aaron Tomlin
4 siblings, 0 replies; 6+ messages in thread
From: Aaron Tomlin @ 2026-08-08 23:55 UTC (permalink / raw)
To: mingo, peterz, juri.lelli, vincent.guittot
Cc: dietmar.eggemann, rostedt, bsegall, mgorman, vschneid,
kprateek.nayak, zhanxusheng1024, neelx, chjohnst, mproche, sean,
steve, linux-kernel
In print_cfs_stats(), rq->leaf_cfs_rq_list is traversed using
for_each_leaf_cfs_rq_safe(), which expands to list_for_each_entry_safe().
Although rq->leaf_cfs_rq_list is RCU-protected, list_for_each_entry_safe()
is a non-RCU iteration macro. It dereferences pointer links without
READ_ONCE() and pre-fetches the next pointer.
When a writer concurrently adds a new cfs_rq to the list using
list_add_rcu(), a reader traversing with list_for_each_entry_safe()
lacks READ_ONCE() protection. Without READ_ONCE(), the compiler is free
to re-fetch pointers or reorder instructions. As a result, the reader
can observe a newly inserted cfs_rq's pointer before its internal fields
are fully visible, leading to reading uninitialised data or
dereferencing invalid pointers. This is illustrated below:
CPU 0 CPU 1
Writer (under rq->lock) Reader (print_cfs_stats)
--------------------------- --------------------------
1. Initialise cfs_rq rcu_read_lock()
cfs_rq->tg = tg
cfs_rq->load = 1024
2. list_add_rcu(&cfs_rq->list, ...): 3. for_each_leaf_cfs_rq_safe():
smp_store_release() Reads prev->next (cfs_rq)
prev->next = cfs_rq -----> Missing READ_ONCE()
4. Reads cfs_rq fields; sees
garbage data
cfs_rq->tg
cfs_rq->load
Fix this by introducing for_each_leaf_cfs_rq_rcu(), which expands to
list_for_each_entry_rcu(). This uses READ_ONCE() during list traversal.
Fixes: 0601267ca673 ("sched/fair: Rewrite list_add_leaf_cfs_rq()")
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Signed-off-by: Aaron Tomlin <atomlin@atomlin.com>
---
kernel/sched/fair.c | 11 +++++++++--
1 file changed, 9 insertions(+), 2 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index d78467ec6ee1..25f941ad637a 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -412,6 +412,10 @@ static inline void assert_list_leaf_cfs_rq(struct rq *rq)
list_for_each_entry_safe(cfs_rq, pos, &rq->leaf_cfs_rq_list, \
leaf_cfs_rq_list)
+#define for_each_leaf_cfs_rq_rcu(rq, cfs_rq) \
+ list_for_each_entry_rcu(cfs_rq, &(rq)->leaf_cfs_rq_list, \
+ leaf_cfs_rq_list)
+
/* Do the two (enqueued) entities belong to the same group ? */
static inline struct cfs_rq *
is_same_group(struct sched_entity *se, struct sched_entity *pse)
@@ -497,6 +501,9 @@ static inline void assert_list_leaf_cfs_rq(struct rq *rq)
#define for_each_leaf_cfs_rq_safe(rq, cfs_rq, pos) \
for (cfs_rq = &rq->cfs, pos = NULL; cfs_rq; cfs_rq = pos)
+#define for_each_leaf_cfs_rq_rcu(rq, cfs_rq) \
+ for (cfs_rq = &rq->cfs; cfs_rq; cfs_rq = NULL)
+
static inline struct sched_entity *parent_entity(struct sched_entity *se)
{
return NULL;
@@ -15401,10 +15408,10 @@ DEFINE_SCHED_CLASS(fair) = {
void print_cfs_stats(struct seq_file *m, int cpu)
{
- struct cfs_rq *cfs_rq, *pos;
+ struct cfs_rq *cfs_rq;
rcu_read_lock();
- for_each_leaf_cfs_rq_safe(cpu_rq(cpu), cfs_rq, pos)
+ for_each_leaf_cfs_rq_rcu(cpu_rq(cpu), cfs_rq)
print_cfs_rq(m, cpu, cfs_rq);
rcu_read_unlock();
}
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* [PATCH v3 4/4] sched/debug: Introduce per-CPU debugfs files
2026-08-08 23:55 [PATCH v3 0/4] sched/debug: Introduce per-CPU debugfs files Aaron Tomlin
` (2 preceding siblings ...)
2026-08-08 23:55 ` [PATCH v3 3/4] sched/fair: Use list_for_each_entry_rcu() in print_cfs_stats() Aaron Tomlin
@ 2026-08-08 23:55 ` Aaron Tomlin
2026-08-10 3:17 ` [PATCH v3 0/4] " Aaron Tomlin
4 siblings, 0 replies; 6+ messages in thread
From: Aaron Tomlin @ 2026-08-08 23:55 UTC (permalink / raw)
To: mingo, peterz, juri.lelli, vincent.guittot
Cc: dietmar.eggemann, rostedt, bsegall, mgorman, vschneid,
kprateek.nayak, zhanxusheng1024, neelx, chjohnst, mproche, sean,
steve, linux-kernel
Currently, accessing scheduler debugging details for a specific CPU
requires reading /sys/kernel/debug/sched/debug, which outputs
information for all online CPUs. When investigating a latency anomaly or
scheduling issue isolated to a specific CPU, accessing
/sys/kernel/debug/sched/cpu/cpu<N>/debug provides an immediate, targeted
view of that runqueue.
Add support for per-CPU debug files under:
/sys/kernel/debug/sched/cpu/cpu<N>/debug. Reading
/sys/kernel/debug/sched/cpu/cpu<N>/debug calls print_cpu() specifically
for CPU <N>, exposing CPU-specific runqueue details on demand. If the
target CPU is currently offline, reading its file returns -ENODEV.
Signed-off-by: Aaron Tomlin <atomlin@atomlin.com>
---
kernel/sched/debug.c | 43 +++++++++++++++++++++++++++++++++++++++++++
1 file changed, 43 insertions(+)
diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c
index e9d40a660346..899954b83bf5 100644
--- a/kernel/sched/debug.c
+++ b/kernel/sched/debug.c
@@ -356,6 +356,7 @@ static const struct file_operations sched_verbose_fops = {
};
static const struct seq_operations sched_debug_sops;
+static void print_cpu(struct seq_file *m, int cpu);
static int sched_debug_open(struct inode *inode, struct file *filp)
{
@@ -633,6 +634,47 @@ static void debugfs_fair_server_init(void)
}
}
+static int sched_debug_cpu_show(struct seq_file *m, void *v)
+{
+ unsigned long cpu = (unsigned long) m->private;
+
+ if (!cpu_online(cpu))
+ return -ENODEV;
+
+ print_cpu(m, cpu);
+ return 0;
+}
+
+static int sched_debug_cpu_open(struct inode *inode, struct file *filp)
+{
+ return single_open(filp, sched_debug_cpu_show, inode->i_private);
+}
+
+static const struct file_operations sched_debug_cpu_fops = {
+ .open = sched_debug_cpu_open,
+ .read = seq_read,
+ .llseek = seq_lseek,
+ .release = single_release,
+};
+
+static __init void debugfs_cpu_init(void)
+{
+ struct dentry *d_cpu_dir;
+ unsigned long cpu;
+ char buf[16];
+
+ d_cpu_dir = debugfs_create_dir("cpu", debugfs_sched);
+
+ for_each_possible_cpu(cpu) {
+ struct dentry *d_cpu;
+
+ snprintf(buf, sizeof(buf), "cpu%lu", cpu);
+ d_cpu = debugfs_create_dir(buf, d_cpu_dir);
+
+ debugfs_create_file("debug", 0444, d_cpu, (void *) cpu, &sched_debug_cpu_fops);
+ }
+}
+
static __init int sched_init_debug(void)
{
struct dentry __maybe_unused *numa, *llc;
@@ -690,6 +732,7 @@ static __init int sched_init_debug(void)
#ifdef CONFIG_SCHED_CLASS_EXT
debugfs_ext_server_init();
#endif
+ debugfs_cpu_init();
return 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* Re: [PATCH v3 0/4] sched/debug: Introduce per-CPU debugfs files
2026-08-08 23:55 [PATCH v3 0/4] sched/debug: Introduce per-CPU debugfs files Aaron Tomlin
` (3 preceding siblings ...)
2026-08-08 23:55 ` [PATCH v3 4/4] sched/debug: Introduce per-CPU debugfs files Aaron Tomlin
@ 2026-08-10 3:17 ` Aaron Tomlin
4 siblings, 0 replies; 6+ messages in thread
From: Aaron Tomlin @ 2026-08-10 3:17 UTC (permalink / raw)
To: mingo, peterz, juri.lelli, vincent.guittot
Cc: dietmar.eggemann, rostedt, bsegall, mgorman, vschneid,
kprateek.nayak, zhanxusheng1024, neelx, chjohnst, mproche, sean,
steve, linux-kernel
On Sat, Aug 08, 2026 at 07:55:18PM -0400, Aaron Tomlin wrote:
> Hi Peter, Juri, Ingo, Vincent,
>
> This patch series addresses a few pre-existing memory safety and list
> traversal concurrency issues in scheduler debugfs handlers, and introduces
> per-CPU debugfs files under /sys/kernel/debug/sched/cpu/cpu<N>/debug.
Please ignore.
Sashiko reported additional pre-existing issues [1] which are valid.
[1]: https://sashiko.dev/#/patchset/20260808235522.380038-1-atomlin%40atomlin.com
--
Aaron Tomlin
^ permalink raw reply [flat|nested] 6+ messages in thread