* [PATCH 1/5] selftests/sched_ext: Fix reload_loop error-path use-after-free
2026-10-09 13:40 [PATCH 0/5] sched_ext: selftest infrastructure fixes Tao Cui
@ 2026-10-09 13:41 ` Tao Cui
2026-10-09 13:41 ` [PATCH 2/5] sched_ext: Add SCX_ECODE_RSN_CGROUP_OFFLINE to user_exit_info.h Tao Cui
` (3 subsequent siblings)
4 siblings, 0 replies; 9+ messages in thread
From: Tao Cui @ 2026-10-09 13:41 UTC (permalink / raw)
To: Tejun Heo
Cc: David Vernet, Andrea Righi, Changwoo Min, sched-ext, cui.tao,
Tao Cui
From: Tao Cui <cuitao@kylinos.cn>
If the second pthread_create() fails, run() returns and the runner
calls cleanup, which destroys the skeleton while the first reload
thread is still running do_reload_loop() and dereferencing its maps -
a use-after-free on the error path.
Set force_exit and join the first thread before failing.
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
tools/testing/selftests/sched_ext/reload_loop.c | 6 +++++-
1 file changed, 5 insertions(+), 1 deletion(-)
diff --git a/tools/testing/selftests/sched_ext/reload_loop.c b/tools/testing/selftests/sched_ext/reload_loop.c
index 49297b83d748..ecff86f59e6c 100644
--- a/tools/testing/selftests/sched_ext/reload_loop.c
+++ b/tools/testing/selftests/sched_ext/reload_loop.c
@@ -53,7 +53,11 @@ static enum scx_test_status run(void *ctx)
SCX_FAIL_IF(err, "Failed to create thread 0");
err = pthread_create(&threads[1], NULL, do_reload_loop, NULL);
- SCX_FAIL_IF(err, "Failed to create thread 1");
+ if (err) {
+ force_exit = true;
+ pthread_join(threads[0], NULL);
+ SCX_FAIL_IF(err, "Failed to create thread 1");
+ }
SCX_FAIL_IF(pthread_join(threads[0], &ret), "thread 0 failed");
SCX_FAIL_IF(pthread_join(threads[1], &ret), "thread 1 failed");
--
2.53.0
^ permalink raw reply related [flat|nested] 9+ messages in thread* [PATCH 2/5] sched_ext: Add SCX_ECODE_RSN_CGROUP_OFFLINE to user_exit_info.h
2026-10-09 13:40 [PATCH 0/5] sched_ext: selftest infrastructure fixes Tao Cui
2026-10-09 13:41 ` [PATCH 1/5] selftests/sched_ext: Fix reload_loop error-path use-after-free Tao Cui
@ 2026-10-09 13:41 ` Tao Cui
2026-10-09 13:41 ` [PATCH 3/5] selftests/sched_ext: Fail loudly when EXIT_KIND is missing from BTF Tao Cui
` (2 subsequent siblings)
4 siblings, 0 replies; 9+ messages in thread
From: Tao Cui @ 2026-10-09 13:41 UTC (permalink / raw)
To: Tejun Heo
Cc: David Vernet, Andrea Righi, Changwoo Min, sched-ext, cui.tao,
Tao Cui
From: Tao Cui <cuitao@kylinos.cn>
The kernel-side scx_exit_code enum gained SCX_ECODE_RSN_CGROUP_OFFLINE
when cgroup lifetime notification support was added, but the tools-side
user_exit_info.h was not updated. Userspace cannot decode exit codes
carrying this reason.
Add the missing enum value.
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
tools/sched_ext/include/scx/user_exit_info.h | 1 +
1 file changed, 1 insertion(+)
diff --git a/tools/sched_ext/include/scx/user_exit_info.h b/tools/sched_ext/include/scx/user_exit_info.h
index 56a02b549aef..cf27ac2447ef 100644
--- a/tools/sched_ext/include/scx/user_exit_info.h
+++ b/tools/sched_ext/include/scx/user_exit_info.h
@@ -52,6 +52,7 @@
enum scx_exit_code {
/* Reasons */
SCX_ECODE_RSN_HOTPLUG = 1LLU << 32,
+ SCX_ECODE_RSN_CGROUP_OFFLINE = 2LLU << 32,
/* Actions */
SCX_ECODE_ACT_RESTART = 1LLU << 48,
--
2.53.0
^ permalink raw reply related [flat|nested] 9+ messages in thread* [PATCH 3/5] selftests/sched_ext: Fail loudly when EXIT_KIND is missing from BTF
2026-10-09 13:40 [PATCH 0/5] sched_ext: selftest infrastructure fixes Tao Cui
2026-10-09 13:41 ` [PATCH 1/5] selftests/sched_ext: Fix reload_loop error-path use-after-free Tao Cui
2026-10-09 13:41 ` [PATCH 2/5] sched_ext: Add SCX_ECODE_RSN_CGROUP_OFFLINE to user_exit_info.h Tao Cui
@ 2026-10-09 13:41 ` Tao Cui
2026-10-09 13:54 ` sashiko-bot
2026-10-09 13:41 ` [PATCH 4/5] selftests/sched_ext: Add SCX_ASSERT_ALIVE and adopt it Tao Cui
2026-10-09 13:41 ` [PATCH 5/5] selftests/sched_ext: Bound exit waits and report reload_loop failures Tao Cui
4 siblings, 1 reply; 9+ messages in thread
From: Tao Cui @ 2026-10-09 13:41 UTC (permalink / raw)
To: Tejun Heo
Cc: David Vernet, Andrea Righi, Changwoo Min, sched-ext, cui.tao,
Tao Cui
From: Tao Cui <cuitao@kylinos.cn>
EXIT_KIND() resolves scx_exit_kind values through
__COMPAT_ENUM_OR_ZERO(), which maps a missing enum to 0. 0 is
SCX_EXIT_NONE, so on kernels whose BTF lacks the enum every
EXIT_KIND() comparison is meaningless.
Resolve the value with __COMPAT_read_enum() and fail the running test
when it is missing, matching SCX_ECODE_VAL() and SCX_KIND_VAL(). The
kick and nohz_tick polling helpers switch to UEI_EXITED() so they
keep their semantics without a BTF lookup.
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
tools/testing/selftests/sched_ext/kick.c | 8 +++++---
tools/testing/selftests/sched_ext/nohz_tick.c | 6 +++---
tools/testing/selftests/sched_ext/scx_test.h | 17 ++++++++++++++++-
3 files changed, 24 insertions(+), 7 deletions(-)
diff --git a/tools/testing/selftests/sched_ext/kick.c b/tools/testing/selftests/sched_ext/kick.c
index 7f01602c75b2..591c250d475d 100644
--- a/tools/testing/selftests/sched_ext/kick.c
+++ b/tools/testing/selftests/sched_ext/kick.c
@@ -97,7 +97,7 @@ static bool wait_for_state(struct kick *skel, u32 wanted)
for (i = 0; i < WAIT_LOOPS; i++) {
if (__atomic_load_n(&skel->bss->state, __ATOMIC_ACQUIRE) == wanted)
return true;
- if (skel->data->uei.kind != EXIT_KIND(SCX_EXIT_NONE))
+ if (UEI_EXITED(skel, uei))
return false;
usleep(1000);
}
@@ -111,7 +111,7 @@ static bool wait_for_counter(struct kick *skel, const u64 *counter, u64 wanted)
for (i = 0; i < WAIT_LOOPS; i++) {
if (__atomic_load_n(counter, __ATOMIC_ACQUIRE) >= wanted)
return true;
- if (skel->data->uei.kind != EXIT_KIND(SCX_EXIT_NONE))
+ if (UEI_EXITED(skel, uei))
return false;
usleep(1000);
}
@@ -421,6 +421,7 @@ static enum scx_test_status invalid_one(struct kick_ctx *ctx, u32 scenario)
enum scx_test_status ret = SCX_TEST_FAIL;
int cpu = ctx->target_cpu;
int i;
+ u64 error_kind;
victim = scx_test_spawn_gated_worker(cpu, false);
if (victim.pid < 0)
@@ -435,12 +436,13 @@ static enum scx_test_status invalid_one(struct kick_ctx *ctx, u32 scenario)
bpf_program__set_autoload(skel->progs.kick_wait_callback, false);
if (kick__load(skel))
goto out;
+ error_kind = EXIT_KIND(SCX_EXIT_ERROR);
ops_link = bpf_map__attach_struct_ops(skel->maps.kick_ops);
if (!ops_link || !scx_test_start_gated_worker(&victim))
goto out;
for (i = 0; i < WAIT_LOOPS; i++) {
- if (skel->data->uei.kind == EXIT_KIND(SCX_EXIT_ERROR)) {
+ if (UEI_EXITED(skel, uei) == error_kind) {
ret = SCX_TEST_PASS;
break;
}
diff --git a/tools/testing/selftests/sched_ext/nohz_tick.c b/tools/testing/selftests/sched_ext/nohz_tick.c
index df40976b9166..e19a97034b7f 100644
--- a/tools/testing/selftests/sched_ext/nohz_tick.c
+++ b/tools/testing/selftests/sched_ext/nohz_tick.c
@@ -165,7 +165,7 @@ static bool wait_for_counter(struct nohz_tick *skel, const u64 *counter, u64 val
for (elapsed = 0; elapsed < timeout_ms; elapsed++) {
if (__atomic_load_n(counter, __ATOMIC_RELAXED) >= value)
return true;
- if (skel->data->uei.kind != EXIT_KIND(SCX_EXIT_NONE))
+ if (UEI_EXITED(skel, uei))
return false;
usleep(1000);
}
@@ -182,7 +182,7 @@ static bool wait_for_tick_stop(struct nohz_tick *skel, const u64 *counter, int t
u64 curr;
usleep(1000);
- if (skel->data->uei.kind != EXIT_KIND(SCX_EXIT_NONE))
+ if (UEI_EXITED(skel, uei))
return false;
curr = __atomic_load_n(counter, __ATOMIC_RELAXED);
if (curr == prev) {
@@ -461,7 +461,7 @@ static enum scx_test_status run(void *ctx_ptr)
scx_test_stop_gated_worker(&challenger);
check_exit:
- if (skel->data->uei.kind != EXIT_KIND(SCX_EXIT_NONE)) {
+ if (UEI_EXITED(skel, uei)) {
SCX_ERR("Scheduler exited unexpectedly (kind=%llu code=%lld)",
(unsigned long long)skel->data->uei.kind,
(long long)skel->data->uei.exit_code);
diff --git a/tools/testing/selftests/sched_ext/scx_test.h b/tools/testing/selftests/sched_ext/scx_test.h
index 90b8d6915bb7..fe0e6dabe7a4 100644
--- a/tools/testing/selftests/sched_ext/scx_test.h
+++ b/tools/testing/selftests/sched_ext/scx_test.h
@@ -9,6 +9,7 @@
#define __SCX_TEST_H__
#include <errno.h>
+#include <unistd.h>
#include <scx/common.h>
#include <scx/compat.h>
@@ -18,7 +19,21 @@ enum scx_test_status {
SCX_TEST_FAIL,
};
-#define EXIT_KIND(__ent) __COMPAT_ENUM_OR_ZERO("scx_exit_kind", #__ent)
+/*
+ * EXIT_KIND - Resolve an scx_exit_kind enum value from kernel BTF.
+ *
+ * __COMPAT_ENUM_OR_ZERO() silently maps a missing enum to 0, which for
+ * exit kinds is indistinguishable from SCX_EXIT_NONE and would make
+ * every EXIT_KIND() comparison pass vacuously. Fail the test loudly
+ * instead, mirroring SCX_ECODE_VAL() and SCX_KIND_VAL() below.
+ */
+#define EXIT_KIND(__ent) ({ \
+ u64 __val = 0; \
+ \
+ if (!__COMPAT_read_enum("scx_exit_kind", #__ent, &__val)) \
+ SCX_FAIL("scx_exit_kind.%s not found in kernel BTF", #__ent); \
+ __val; \
+})
struct scx_test {
/**
--
2.53.0
^ permalink raw reply related [flat|nested] 9+ messages in thread* Re: [PATCH 3/5] selftests/sched_ext: Fail loudly when EXIT_KIND is missing from BTF
2026-10-09 13:41 ` [PATCH 3/5] selftests/sched_ext: Fail loudly when EXIT_KIND is missing from BTF Tao Cui
@ 2026-10-09 13:54 ` sashiko-bot
0 siblings, 0 replies; 9+ messages in thread
From: sashiko-bot @ 2026-10-09 13:54 UTC (permalink / raw)
To: Tao Cui; +Cc: sched-ext
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] selftests/sched_ext: EXIT_KIND() hidden return bypasses cleanup and leaks spinning threads
--
commit 083f03de0fcd50d9c9ef1d52267ab98872def3d2
Author: Tao Cui <cuitao@kylinos.cn>
selftests/sched_ext: Fail loudly when EXIT_KIND is missing from BTF
The patch updates EXIT_KIND() in the sched_ext selftests to fail explicitly via
__COMPAT_read_enum() when a BTF enum is missing, avoiding silent failures where
SCX_EXIT_NONE is returned. It changes the polling helpers in kick and nohz_tick
to use UEI_EXITED() to retain their original semantics without requiring a BTF
lookup.
> diff --git a/tools/testing/selftests/sched_ext/kick.c b/tools/testing/selftests/sched_ext/kick.c
> index 7f01602c75b29..591c250d475db 100644
> --- a/tools/testing/selftests/sched_ext/kick.c
> +++ b/tools/testing/selftests/sched_ext/kick.c
[ ... ]
> @@ -421,6 +421,7 @@ static enum scx_test_status invalid_one(struct kick_ctx *ctx, u32 scenario)
> enum scx_test_status ret = SCX_TEST_FAIL;
> int cpu = ctx->target_cpu;
> int i;
> + u64 error_kind;
>
> victim = scx_test_spawn_gated_worker(cpu, false);
> if (victim.pid < 0)
> @@ -435,12 +436,13 @@ static enum scx_test_status invalid_one(struct kick_ctx *ctx, u32 scenario)
> bpf_program__set_autoload(skel->progs.kick_wait_callback, false);
> if (kick__load(skel))
> goto out;
> + error_kind = EXIT_KIND(SCX_EXIT_ERROR);
[Severity: Medium]
Does this code leak the spawned gated worker and skeleton object on BTF lookup
failure?
The new EXIT_KIND() macro embeds SCX_FAIL(), which performs a hidden function
return. If the BTF lookup fails here, invalid_one() returns directly and
bypasses the goto out; cleanup path below.
This leaks the skel object and leaves the victim gated worker spinning
indefinitely in the background, which can corrupt the environment for
subsequent tests in the runner.
This pattern of leaking threads also appears in
tools/testing/selftests/sched_ext/cyclic_kick_wait.c:run() where EXIT_KIND() is
called after background pthreads have been started.
> ops_link = bpf_map__attach_struct_ops(skel->maps.kick_ops);
> if (!ops_link || !scx_test_start_gated_worker(&victim))
> goto out;
>
> for (i = 0; i < WAIT_LOOPS; i++) {
> - if (skel->data->uei.kind == EXIT_KIND(SCX_EXIT_ERROR)) {
> + if (UEI_EXITED(skel, uei) == error_kind) {
> ret = SCX_TEST_PASS;
> break;
> }
[ ... ]
> diff --git a/tools/testing/selftests/sched_ext/scx_test.h b/tools/testing/selftests/sched_ext/scx_test.h
> index 90b8d6915bb7e..fe0e6dabe7a48 100644
> --- a/tools/testing/selftests/sched_ext/scx_test.h
> +++ b/tools/testing/selftests/sched_ext/scx_test.h
[ ... ]
> @@ -18,7 +19,21 @@ enum scx_test_status {
> SCX_TEST_FAIL,
> };
>
> -#define EXIT_KIND(__ent) __COMPAT_ENUM_OR_ZERO("scx_exit_kind", #__ent)
> +/*
> + * EXIT_KIND - Resolve an scx_exit_kind enum value from kernel BTF.
> + *
> + * __COMPAT_ENUM_OR_ZERO() silently maps a missing enum to 0, which for
> + * exit kinds is indistinguishable from SCX_EXIT_NONE and would make
> + * every EXIT_KIND() comparison pass vacuously. Fail the test loudly
> + * instead, mirroring SCX_ECODE_VAL() and SCX_KIND_VAL() below.
> + */
> +#define EXIT_KIND(__ent) ({ \
> + u64 __val = 0; \
> + \
> + if (!__COMPAT_read_enum("scx_exit_kind", #__ent, &__val)) \
> + SCX_FAIL("scx_exit_kind.%s not found in kernel BTF", #__ent); \
[Severity: Medium]
Could this hidden return cause resource leaks?
SCX_FAIL() expands to a direct return SCX_TEST_FAIL;. Because EXIT_KIND() is
used as an expression, this introduces a hidden control flow branch that
silently returns from the caller, bypassing explicit explicit cleanup paths.
> + __val; \
> +})
>
> struct scx_test {
> /**
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261009134104.296891-1-cui.tao@linux.dev?part=3
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH 4/5] selftests/sched_ext: Add SCX_ASSERT_ALIVE and adopt it
2026-10-09 13:40 [PATCH 0/5] sched_ext: selftest infrastructure fixes Tao Cui
` (2 preceding siblings ...)
2026-10-09 13:41 ` [PATCH 3/5] selftests/sched_ext: Fail loudly when EXIT_KIND is missing from BTF Tao Cui
@ 2026-10-09 13:41 ` Tao Cui
2026-10-09 13:41 ` [PATCH 5/5] selftests/sched_ext: Bound exit waits and report reload_loop failures Tao Cui
4 siblings, 0 replies; 9+ messages in thread
From: Tao Cui @ 2026-10-09 13:41 UTC (permalink / raw)
To: Tejun Heo
Cc: David Vernet, Andrea Righi, Changwoo Min, sched-ext, cui.tao,
Tao Cui
From: Tao Cui <cuitao@kylinos.cn>
A scheduler that dies mid-test freezes the state that the test is
about to assert on, so the assertions can pass vacuously. Add
SCX_ASSERT_ALIVE(), which fails the test with the recorded exit kind
and code, and use it in maximal, create_dsq and init_enable_count.
Existing open-coded checks keep their goto-based cleanup and are not
converted. init_enable_count also clears the uei between its two
attach cycles, since the clean unregister of the first cycle would
otherwise trip the check of the second.
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
tools/testing/selftests/sched_ext/create_dsq.bpf.c | 8 ++++++++
tools/testing/selftests/sched_ext/create_dsq.c | 2 ++
.../selftests/sched_ext/init_enable_count.bpf.c | 8 ++++++++
.../selftests/sched_ext/init_enable_count.c | 13 +++++++++++++
tools/testing/selftests/sched_ext/maximal.bpf.c | 6 +++++-
tools/testing/selftests/sched_ext/maximal.c | 2 ++
tools/testing/selftests/sched_ext/scx_test.h | 14 ++++++++++++++
7 files changed, 52 insertions(+), 1 deletion(-)
diff --git a/tools/testing/selftests/sched_ext/create_dsq.bpf.c b/tools/testing/selftests/sched_ext/create_dsq.bpf.c
index 680cc4b6d8c7..24dacd1aab0c 100644
--- a/tools/testing/selftests/sched_ext/create_dsq.bpf.c
+++ b/tools/testing/selftests/sched_ext/create_dsq.bpf.c
@@ -12,6 +12,8 @@ char _license[] SEC("license") = "GPL";
u32 nr_lifecycle_tests;
+UEI_DEFINE(uei);
+
void BPF_STRUCT_OPS(create_dsq_exit_task, struct task_struct *p,
struct scx_exit_task_args *args)
{
@@ -84,10 +86,16 @@ s32 BPF_STRUCT_OPS_SLEEPABLE(create_dsq_init)
return 0;
}
+void BPF_STRUCT_OPS(create_dsq_exit, struct scx_exit_info *ei)
+{
+ UEI_RECORD(uei, ei);
+}
+
SEC(".struct_ops.link")
struct sched_ext_ops create_dsq_ops = {
.init_task = (void *) create_dsq_init_task,
.exit_task = (void *) create_dsq_exit_task,
.init = (void *) create_dsq_init,
+ .exit = (void *) create_dsq_exit,
.name = "create_dsq",
};
diff --git a/tools/testing/selftests/sched_ext/create_dsq.c b/tools/testing/selftests/sched_ext/create_dsq.c
index 422e6532ee71..84143d5710f4 100644
--- a/tools/testing/selftests/sched_ext/create_dsq.c
+++ b/tools/testing/selftests/sched_ext/create_dsq.c
@@ -35,6 +35,8 @@ static enum scx_test_status run(void *ctx)
return SCX_TEST_FAIL;
}
+ SCX_ASSERT_ALIVE(skel, uei);
+
bpf_link__destroy(link);
SCX_EQ(skel->bss->nr_lifecycle_tests, 1024);
diff --git a/tools/testing/selftests/sched_ext/init_enable_count.bpf.c b/tools/testing/selftests/sched_ext/init_enable_count.bpf.c
index 5eb9edb1837d..c0a5dd8fe183 100644
--- a/tools/testing/selftests/sched_ext/init_enable_count.bpf.c
+++ b/tools/testing/selftests/sched_ext/init_enable_count.bpf.c
@@ -15,6 +15,8 @@ char _license[] SEC("license") = "GPL";
u64 init_task_cnt, exit_task_cnt, enable_cnt, disable_cnt;
u64 init_fork_cnt, init_transition_cnt;
+UEI_DEFINE(uei);
+
s32 BPF_STRUCT_OPS_SLEEPABLE(cnt_init_task, struct task_struct *p,
struct scx_init_task_args *args)
{
@@ -43,11 +45,17 @@ void BPF_STRUCT_OPS(cnt_disable, struct task_struct *p)
__sync_fetch_and_add(&disable_cnt, 1);
}
+void BPF_STRUCT_OPS(cnt_exit, struct scx_exit_info *ei)
+{
+ UEI_RECORD(uei, ei);
+}
+
SEC(".struct_ops.link")
struct sched_ext_ops init_enable_count_ops = {
.init_task = (void *) cnt_init_task,
.exit_task = (void *) cnt_exit_task,
.enable = (void *) cnt_enable,
.disable = (void *) cnt_disable,
+ .exit = (void *) cnt_exit,
.name = "init_enable_count",
};
diff --git a/tools/testing/selftests/sched_ext/init_enable_count.c b/tools/testing/selftests/sched_ext/init_enable_count.c
index 44577e30e764..c4a57038a1d6 100644
--- a/tools/testing/selftests/sched_ext/init_enable_count.c
+++ b/tools/testing/selftests/sched_ext/init_enable_count.c
@@ -5,6 +5,7 @@
* Copyright (c) 2023 Tejun Heo <tj@kernel.org>
*/
#include <signal.h>
+#include <string.h>
#include <stdio.h>
#include <unistd.h>
#include <sched.h>
@@ -72,7 +73,17 @@ static enum scx_test_status run_test(bool global)
close(pipe_fds[1]);
signal(SIGCHLD, SIG_DFL);
+ SCX_ASSERT_ALIVE(skel, uei);
+
bpf_link__destroy(link);
+
+ /*
+ * The clean unregister above is recorded in uei and would
+ * otherwise trip the liveness check of the second cycle, which
+ * reuses this skeleton.
+ */
+ memset(&skel->data->uei, 0, sizeof(skel->data->uei));
+
SCX_GE(skel->bss->init_task_cnt, num_pre_forks);
SCX_GE(skel->bss->exit_task_cnt, num_pre_forks);
@@ -124,6 +135,8 @@ static enum scx_test_status run_test(bool global)
status);
}
+ SCX_ASSERT_ALIVE(skel, uei);
+
bpf_link__destroy(link);
SCX_GE(skel->bss->init_task_cnt, 2 * num_children);
diff --git a/tools/testing/selftests/sched_ext/maximal.bpf.c b/tools/testing/selftests/sched_ext/maximal.bpf.c
index 04a369078aac..9ea664e2ff70 100644
--- a/tools/testing/selftests/sched_ext/maximal.bpf.c
+++ b/tools/testing/selftests/sched_ext/maximal.bpf.c
@@ -14,6 +14,8 @@ char _license[] SEC("license") = "GPL";
#define DSQ_ID 0
+UEI_DEFINE(uei);
+
s32 BPF_STRUCT_OPS(maximal_select_cpu, struct task_struct *p, s32 prev_cpu,
u64 wake_flags)
{
@@ -132,7 +134,9 @@ s32 BPF_STRUCT_OPS_SLEEPABLE(maximal_init)
}
void BPF_STRUCT_OPS(maximal_exit, struct scx_exit_info *info)
-{}
+{
+ UEI_RECORD(uei, info);
+}
SEC(".struct_ops.link")
struct sched_ext_ops maximal_ops = {
diff --git a/tools/testing/selftests/sched_ext/maximal.c b/tools/testing/selftests/sched_ext/maximal.c
index 1dc369224670..6c6b80e0446e 100644
--- a/tools/testing/selftests/sched_ext/maximal.c
+++ b/tools/testing/selftests/sched_ext/maximal.c
@@ -35,6 +35,8 @@ static enum scx_test_status run(void *ctx)
link = bpf_map__attach_struct_ops(skel->maps.maximal_ops);
SCX_FAIL_IF(!link, "Failed to attach scheduler");
+ SCX_ASSERT_ALIVE(skel, uei);
+
bpf_link__destroy(link);
return SCX_TEST_PASS;
diff --git a/tools/testing/selftests/sched_ext/scx_test.h b/tools/testing/selftests/sched_ext/scx_test.h
index fe0e6dabe7a4..aa676cd5bb8d 100644
--- a/tools/testing/selftests/sched_ext/scx_test.h
+++ b/tools/testing/selftests/sched_ext/scx_test.h
@@ -125,6 +125,20 @@ void scx_test_register(struct scx_test *test);
#define SCX_ASSERT(_x) SCX_FAIL_IF(!(_x), "Expected %s to be true (%lu)", \
#_x, (u64)(_x))
+/*
+ * SCX_ASSERT_ALIVE - Fail the test if the scheduler under test exited
+ * on its own between attach and this check, e.g. from a runtime error.
+ * A scheduler that dies mid-test freezes whatever state the test is
+ * about to assert on, which can make the assertions pass vacuously.
+ */
+#define SCX_ASSERT_ALIVE(__skel, __uei) \
+ do { \
+ if (UEI_EXITED(__skel, __uei)) \
+ SCX_FAIL("Scheduler exited unexpectedly (kind=%llu code=%lld)", \
+ (unsigned long long)(__skel)->data->__uei.kind, \
+ (long long)(__skel)->data->__uei.exit_code); \
+ } while (0)
+
#define SCX_ECODE_VAL(__ecode) ({ \
u64 __val = 0; \
bool __found = false; \
--
2.53.0
^ permalink raw reply related [flat|nested] 9+ messages in thread* [PATCH 5/5] selftests/sched_ext: Bound exit waits and report reload_loop failures
2026-10-09 13:40 [PATCH 0/5] sched_ext: selftest infrastructure fixes Tao Cui
` (3 preceding siblings ...)
2026-10-09 13:41 ` [PATCH 4/5] selftests/sched_ext: Add SCX_ASSERT_ALIVE and adopt it Tao Cui
@ 2026-10-09 13:41 ` Tao Cui
2026-10-09 13:48 ` sashiko-bot
4 siblings, 1 reply; 9+ messages in thread
From: Tao Cui @ 2026-10-09 13:41 UTC (permalink / raw)
To: Tejun Heo
Cc: David Vernet, Andrea Righi, Changwoo Min, sched-ext, cui.tao,
Tao Cui
From: Tao Cui <cuitao@kylinos.cn>
exit, prog_run and hotplug wait for the scheduler to exit by spinning
on its uei with no time bound; a scheduler that never exits hangs the
runner instead of failing the test. Bound the waits to ~10s
(iteration-based) with the new SCX_TIMED_WAIT_UEI() helper.
reload_loop skips bpf_map__attach_struct_ops() failures silently:
measured on a loaded host, 1125-1233 of 2048 iterations failed (range across
boots), all with EBUSY (the single scheduler slot changing hands between the two
threads). Count the failures per thread and report them; fail the
test only on unexpected errnos or when every attach failed.
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
tools/testing/selftests/sched_ext/exit.c | 4 +-
tools/testing/selftests/sched_ext/hotplug.c | 12 +++--
tools/testing/selftests/sched_ext/prog_run.c | 7 ++-
.../testing/selftests/sched_ext/reload_loop.c | 47 ++++++++++++++++---
tools/testing/selftests/sched_ext/scx_test.h | 17 +++++++
5 files changed, 73 insertions(+), 14 deletions(-)
diff --git a/tools/testing/selftests/sched_ext/exit.c b/tools/testing/selftests/sched_ext/exit.c
index 01b17092d5c8..9270d8e79eb6 100644
--- a/tools/testing/selftests/sched_ext/exit.c
+++ b/tools/testing/selftests/sched_ext/exit.c
@@ -43,8 +43,8 @@ static enum scx_test_status run(void *ctx)
}
/* Assumes uei.kind is written last */
- while (skel->data->uei.kind == EXIT_KIND(SCX_EXIT_NONE))
- sched_yield();
+ SCX_FAIL_IF(!SCX_TIMED_WAIT_UEI(skel, uei, 10000),
+ "Timed out waiting for scheduler to exit");
SCX_EQ(skel->data->uei.kind, EXIT_KIND(SCX_EXIT_UNREG_BPF));
SCX_EQ(skel->data->uei.exit_code, tc);
diff --git a/tools/testing/selftests/sched_ext/hotplug.c b/tools/testing/selftests/sched_ext/hotplug.c
index 10b8d42bd89b..3c07686a93c8 100644
--- a/tools/testing/selftests/sched_ext/hotplug.c
+++ b/tools/testing/selftests/sched_ext/hotplug.c
@@ -85,8 +85,10 @@ static enum scx_test_status test_hotplug(bool onlining, bool cbs_defined)
if (toggle_online_status(onlining ? 1 : 0))
goto out_destroy_link;
- while (!UEI_EXITED(skel, uei))
- sched_yield();
+ if (!SCX_TIMED_WAIT_UEI(skel, uei, 10000)) {
+ SCX_ERR("Timed out waiting for scheduler to exit");
+ goto out_destroy_link;
+ }
SCX_EQ(skel->data->uei.kind, kind);
SCX_EQ(UEI_REPORT(skel, uei), code);
@@ -130,8 +132,10 @@ static enum scx_test_status test_hotplug_attach(void)
goto out_destroy_link;
SCX_ASSERT(link);
- while (!UEI_EXITED(skel, uei))
- sched_yield();
+ if (!SCX_TIMED_WAIT_UEI(skel, uei, 10000)) {
+ SCX_ERR("Timed out waiting for scheduler to exit");
+ goto out_destroy_link;
+ }
kind = SCX_KIND_VAL(SCX_EXIT_UNREG_KERN);
code = SCX_ECODE_VAL(SCX_ECODE_ACT_RESTART) |
diff --git a/tools/testing/selftests/sched_ext/prog_run.c b/tools/testing/selftests/sched_ext/prog_run.c
index 1129ec2aaddc..369c744ac06e 100644
--- a/tools/testing/selftests/sched_ext/prog_run.c
+++ b/tools/testing/selftests/sched_ext/prog_run.c
@@ -55,8 +55,11 @@ static enum scx_test_status run(void *ctx)
}
/* Assumes uei.kind is written last */
- while (skel->data->uei.kind == EXIT_KIND(SCX_EXIT_NONE))
- sched_yield();
+ if (!SCX_TIMED_WAIT_UEI(skel, uei, 10000)) {
+ SCX_ERR("Timed out waiting for scheduler to exit");
+ status = SCX_TEST_FAIL;
+ goto out;
+ }
if (skel->data->uei.kind != EXIT_KIND(SCX_EXIT_UNREG_BPF)) {
SCX_ERR("Unexpected exit kind: %llu",
diff --git a/tools/testing/selftests/sched_ext/reload_loop.c b/tools/testing/selftests/sched_ext/reload_loop.c
index ecff86f59e6c..4ae6fb43b6c5 100644
--- a/tools/testing/selftests/sched_ext/reload_loop.c
+++ b/tools/testing/selftests/sched_ext/reload_loop.c
@@ -4,6 +4,7 @@
* Copyright (c) 2024 David Vernet <dvernet@meta.com>
*/
#include <bpf/bpf.h>
+#include <errno.h>
#include <pthread.h>
#include <scx/common.h>
#include <sys/wait.h>
@@ -11,8 +12,12 @@
#include "maximal.bpf.skel.h"
#include "scx_test.h"
+#define RELOAD_ITERS 1024
+#define RELOAD_THREADS 2
+
static struct maximal *skel;
-static pthread_t threads[2];
+static pthread_t threads[RELOAD_THREADS];
+static int unexpected_errno;
bool force_exit = false;
@@ -29,23 +34,36 @@ static enum scx_test_status setup(void **ctx)
return SCX_TEST_PASS;
}
+
+
static void *do_reload_loop(void *arg)
{
- u32 i;
+ u32 i, failures = 0;
- for (i = 0; i < 1024 && !force_exit; i++) {
+ for (i = 0; i < RELOAD_ITERS && !force_exit; i++) {
struct bpf_link *link;
+ errno = 0;
link = bpf_map__attach_struct_ops(skel->maps.maximal_ops);
- if (link)
- bpf_link__destroy(link);
+ if (!link) {
+ int e = errno;
+
+ if (e != EBUSY)
+ __atomic_store_n(&unexpected_errno, e,
+ __ATOMIC_RELAXED);
+ failures++;
+ continue;
+ }
+
+ bpf_link__destroy(link);
}
- return NULL;
+ return (void *)(unsigned long)failures;
}
static enum scx_test_status run(void *ctx)
{
+ unsigned long failures = 0;
int err;
void *ret;
@@ -60,7 +78,24 @@ static enum scx_test_status run(void *ctx)
}
SCX_FAIL_IF(pthread_join(threads[0], &ret), "thread 0 failed");
+ failures += (unsigned long)ret;
+
SCX_FAIL_IF(pthread_join(threads[1], &ret), "thread 1 failed");
+ failures += (unsigned long)ret;
+
+ SCX_FAIL_IF(unexpected_errno, "unexpected attach errno %d",
+ unexpected_errno);
+
+ if (failures) {
+ printf("reload_loop: %lu attach failures%s\n",
+ (unsigned long)failures,
+ unexpected_errno ? ", unexpected errno" :
+ " (EBUSY: scheduler slot busy)");
+ fflush(stdout);
+ }
+
+ SCX_FAIL_IF(failures == RELOAD_ITERS * RELOAD_THREADS,
+ "every attach failed");
return SCX_TEST_PASS;
}
diff --git a/tools/testing/selftests/sched_ext/scx_test.h b/tools/testing/selftests/sched_ext/scx_test.h
index aa676cd5bb8d..be63447b8898 100644
--- a/tools/testing/selftests/sched_ext/scx_test.h
+++ b/tools/testing/selftests/sched_ext/scx_test.h
@@ -139,6 +139,23 @@ void scx_test_register(struct scx_test *test);
(long long)(__skel)->data->__uei.exit_code); \
} while (0)
+/*
+ * SCX_TIMED_WAIT_UEI - Wait for the scheduler under test to exit, i.e.
+ * its uei to be filled in, for up to @__ms iterations at ~1ms each.
+ * Returns true if the scheduler exited and false on timeout, so a
+ * scheduler that never exits fails the test instead of hanging it.
+ * The bound is iteration-based rather than wall-clock and stretches
+ * on a loaded system instead of failing early.
+ */
+#define SCX_TIMED_WAIT_UEI(__skel, __uei, __ms) \
+ ({ \
+ s64 __left = (__ms); \
+ \
+ while (__left-- > 0 && !UEI_EXITED(__skel, __uei)) \
+ usleep(1000); \
+ UEI_EXITED(__skel, __uei) != 0; \
+ })
+
#define SCX_ECODE_VAL(__ecode) ({ \
u64 __val = 0; \
bool __found = false; \
--
2.53.0
^ permalink raw reply related [flat|nested] 9+ messages in thread* Re: [PATCH 5/5] selftests/sched_ext: Bound exit waits and report reload_loop failures
2026-10-09 13:41 ` [PATCH 5/5] selftests/sched_ext: Bound exit waits and report reload_loop failures Tao Cui
@ 2026-10-09 13:48 ` sashiko-bot
0 siblings, 0 replies; 9+ messages in thread
From: sashiko-bot @ 2026-10-09 13:48 UTC (permalink / raw)
To: Tao Cui; +Cc: sched-ext
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] selftests/sched_ext: leak of BPF link and skeleton on timeout in exit.c
- [Low] selftests/sched_ext: dead code in reload_loop failure reporting
--
commit b49c5aa1d20c7dfda172f19617a72ddbef7da918
Author: Tao Cui <cuitao@kylinos.cn>
selftests/sched_ext: Bound exit waits and report reload_loop failures
This commit replaces unbounded waits for scheduler exit in multiple tests
with an iteration-bound timeout using SCX_TIMED_WAIT_UEI(). It also updates
reload_loop to track and report attach failures, failing the test on
unexpected errors or if every attach attempt fails.
> diff --git a/tools/testing/selftests/sched_ext/exit.c b/tools/testing/selftests/sched_ext/exit.c
> index 01b17092d5c8f..9270d8e79eb63 100644
> --- a/tools/testing/selftests/sched_ext/exit.c
> +++ b/tools/testing/selftests/sched_ext/exit.c
> @@ -43,8 +43,8 @@ static enum scx_test_status run(void *ctx)
> }
>
> /* Assumes uei.kind is written last */
> - while (skel->data->uei.kind == EXIT_KIND(SCX_EXIT_NONE))
> - sched_yield();
> + SCX_FAIL_IF(!SCX_TIMED_WAIT_UEI(skel, uei, 10000),
> + "Timed out waiting for scheduler to exit");
>
[Severity: Medium]
Does this SCX_FAIL_IF() check leak the BPF link and skeleton on timeout?
In run(), skel and link are allocated locally inside the loop. If the wait
times out, SCX_FAIL_IF() will immediately return SCX_TEST_FAIL, bypassing
the bpf_link__destroy() and exit__destroy() calls at the end of the loop block.
Failing to destroy the BPF link leaves the scheduler attached to the kernel,
which breaks test isolation. Subsequent tests in the runner might then fail
with -EBUSY because only one sched_ext scheduler can be attached at a time.
> SCX_EQ(skel->data->uei.kind, EXIT_KIND(SCX_EXIT_UNREG_BPF));
> SCX_EQ(skel->data->uei.exit_code, tc);
[ ... ]
> diff --git a/tools/testing/selftests/sched_ext/reload_loop.c b/tools/testing/selftests/sched_ext/reload_loop.c
> index ecff86f59e6c7..4ae6fb43b6c5c 100644
> --- a/tools/testing/selftests/sched_ext/reload_loop.c
> +++ b/tools/testing/selftests/sched_ext/reload_loop.c
[ ... ]
> @@ -60,7 +78,24 @@ static enum scx_test_status run(void *ctx)
> }
>
> SCX_FAIL_IF(pthread_join(threads[0], &ret), "thread 0 failed");
> + failures += (unsigned long)ret;
> +
> SCX_FAIL_IF(pthread_join(threads[1], &ret), "thread 1 failed");
> + failures += (unsigned long)ret;
> +
> + SCX_FAIL_IF(unexpected_errno, "unexpected attach errno %d",
> + unexpected_errno);
> +
> + if (failures) {
> + printf("reload_loop: %lu attach failures%s\n",
> + (unsigned long)failures,
> + unexpected_errno ? ", unexpected errno" :
> + " (EBUSY: scheduler slot busy)");
[Severity: Low]
Is the true branch of this ternary condition dead code?
The SCX_FAIL_IF(unexpected_errno, ...) check slightly above will immediately
return from the function if unexpected_errno is non-zero. As a result, when
execution reaches this printf statement, unexpected_errno must evaluate to
false.
The true branch (", unexpected errno") can never be executed, preventing the
unexpected errno from being reported in the summary as intended.
> + fflush(stdout);
> + }
> +
> + SCX_FAIL_IF(failures == RELOAD_ITERS * RELOAD_THREADS,
> + "every attach failed");
>
> return SCX_TEST_PASS;
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261009134104.296891-1-cui.tao@linux.dev?part=5
^ permalink raw reply [flat|nested] 9+ messages in thread