From: Andrea Righi <arighi@nvidia.com>
To: Wanwu Li <liwanwu@kylinos.cn>
Cc: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
Changwoo Min <changwoo@igalia.com>,
sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org,
stable@vger.kernel.org
Subject: Re: [PATCH v2] sched_ext: Fix NULL sched deref in kfunc sub-sched error paths
Date: Wed, 2 Sep 2026 20:11:14 +0200 [thread overview]
Message-ID: <aphmwjLkdS4tKpbv@gpd4> (raw)
In-Reply-To: <20260902170751.256434-1-liwanwu@kylinos.cn>
Hi Wanwu,
On Thu, Sep 03, 2026 at 01:07:51AM +0800, Wanwu Li wrote:
> The COMPAT kfunc wrappers scx_bpf_select_cpu_and() and
> scx_bpf_dsq_insert_vtime() error out @p's scheduler when the root
> scheduler has sub-scheds attached:
>
> scx_error(scx_task_sched(p), "... must be used");
>
> scx_task_sched(p) is p->scx.sched, which is NULL for any task that is not
> on an scx scheduler: it is memset() by init_scx_entity() and cleared by
> scx_disable_and_exit_task() -- which sched_ext_dead() runs when a task
> exits. It is also an rcu_dereference_protected() that must be called with
> @p's pi_lock or rq lock held, which neither wrapper does. Passing NULL to
> scx_error() reaches scx_vexit(), which dereferences sch->exit_info
> unconditionally, oopsing the kernel.
>
> Both wrappers are reachable with such a @p. scx_bpf_select_cpu_and() is in
> the select_cpu kfunc group, which scx_kfunc_context_filter() opens to
> BPF_PROG_TYPE_SYSCALL programs. scx_bpf_dsq_insert_vtime() is in the
> enqueue_dispatch group, which ops.enqueue() and ops.dispatch() may call
> with any KF_RCU task: the group has no kf_tasks validation, and
> scx_dsq_insert_preamble() checks task ownership with scx_task_on_sched()
> precisely because @p may be an arbitrary task.
>
> Neither wrapper requires a contrived @p. Tasks that are never enabled --
> kthreads and tasks of other classes under SCX_SWITCH_ALL=n -- keep
> p->scx.sched NULL indefinitely; and a task handed over from
> bpf_task_from_pid() can exit before the call lands, as its zombie stays
> visible to pid lookups until it is reaped.
>
> One concrete trigger exercised for this changelog: a SYSCALL program
> calling the select_cpu_and wrapper on an exited-but-not-reaped task while
> a sub-scheduler is attached (faulting instruction is the scx_vexit()
> prologue "mov r15,[rdi+0x398]" with RDI=NULL and 0x398 the offset of
> sch->exit_info):
>
> sched_ext: BPF scheduler "kfunc_subsched_null" enabled
> sched_ext: BPF sub-scheduler "kfunc_subsched_null" enabled
> sched_ext: Unassociated program run_select_cpu_ (id 76)
> BUG: kernel NULL pointer dereference, address: 0000000000000398
> #PF: supervisor read access in kernel mode
> #PF: error_code(0x0000) - not-present page
> Oops: Oops: 0000 [#1] SMP NOPTI
> CPU: 7 UID: 0 PID: 8201 Comm: kfunc_test_runn Tainted: G W
> RIP: 0010:scx_vexit+0x25/0xa0
> Code: ... <4c> 8b bf 98 03 00 00 ...
> CR2: 0000000000000398
> Call Trace:
> <TASK>
> __scx_exit+0x4f/0x70
> scx_bpf_select_cpu_and+0xab/0xb0
> bpf_prog_430ed61a7b66e03a_run_select_cpu_and+0x9c/0xe7
> ? __x64_sys_bpf+0x2c/0x40
> bpf_prog_test_run_syscall+0x130/0x2f0
> __sys_bpf+0x930/0x10d0
> ? __x64_sys_bpf+0x2c/0x40
> __x64_sys_bpf+0x2c/0x40
> do_syscall_64+0xbc/0x460
> ? rseq_set_ids_get_csaddr+0x81/0x140
> ? __rseq_handle_slowpath+0xd0/0x130
> ? switch_fpu_return+0x51/0xd0
> ? arch_exit_to_user_mode_prepare.constprop.0+0x87/0xb0
> ? do_syscall_64+0xf3/0x460
> ? irqentry_exit+0x48/0x740
> ? clear_bhb_loop+0x40/0x90
> ? do_syscall_64+0x35/0x460
> entry_SYSCALL_64_after_hwframe+0x76/0x7e
> </TASK>
>
> Keep the "error out @p's scheduler" attribution -- it is what every other
> kfunc error path does (select_cpu_from_kfunc()'s cross_task,
> scx_kf_arg_task_ok()) and it is correct for the callers that take this
> path with a live task: p->scx.sched is their scheduler. Just read it
> safely: use scx_task_sched_rcu() (valid under the guard(rcu)() both
> wrappers already hold, no @p lock required) and fall back to @sch -- the
> root scheduler, guaranteed non-NULL here -- when @p is not on an scx
> scheduler, which is precisely the case that used to be NULL.
>
> These COMPAT wrappers are scheduled for eventual removal once the
> deprecation grace period elapses, but until then -- and regardless of
> their removal timeline -- they must not oops the kernel on a task they
> are handed; this fix makes the error path safe.
>
> Cc: <stable@vger.kernel.org>
> Fixes: a5fa0708cbfd ("sched_ext: Enforce scheduling authority in dispatch and select_cpu operations")
> Signed-off-by: Wanwu Li <liwanwu@kylinos.cn>
This looks good to me.
Reviewed-by: Andrea Righi <arighi@nvidia.com>
Thanks,
-Andrea
> ---
>
> Changes v1 -> v2:
> - Extend the same fallback to scx_bpf_dsq_insert_vtime() as suggested
> by Andrea (and flagged by sashiko-bot); rewrite the reachability
> argument in the commit message accordingly.
>
> kernel/sched/ext/ext.c | 9 +++++++--
> kernel/sched/ext/idle.c | 9 +++++++--
> 2 files changed, 14 insertions(+), 4 deletions(-)
>
> diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
> index 10af28a9f2c0..fdfaa7e9c8f5 100644
> --- a/kernel/sched/ext/ext.c
> +++ b/kernel/sched/ext/ext.c
> @@ -8943,10 +8943,15 @@ __bpf_kfunc void scx_bpf_dsq_insert_vtime(struct task_struct *p, u64 dsq_id,
> #ifdef CONFIG_EXT_SUB_SCHED
> /*
> * Disallow if any sub-scheds are attached. There is no way to tell
> - * which scheduler called us, just error out @p's scheduler.
> + * which scheduler called us, so error out @p's scheduler -- but read
> + * it under RCU (@p's locks aren't necessarily held here) and fall
> + * back to @sch if @p isn't on an scx scheduler: enqueue/dispatch
> + * contexts may pass any KF_RCU task, and p->scx.sched is NULL for
> + * one that has exited or is managed by another scheduler.
> */
> if (unlikely(!list_empty(&sch->children))) {
> - scx_error(scx_task_sched(p), "__scx_bpf_dsq_insert_vtime() must be used");
> + scx_error(scx_task_sched_rcu(p) ?: sch,
> + "__scx_bpf_dsq_insert_vtime() must be used");
> return;
> }
> #endif
> diff --git a/kernel/sched/ext/idle.c b/kernel/sched/ext/idle.c
> index d2973fb3af6d..014599d82bb0 100644
> --- a/kernel/sched/ext/idle.c
> +++ b/kernel/sched/ext/idle.c
> @@ -1142,10 +1142,15 @@ __bpf_kfunc s32 scx_bpf_select_cpu_and(struct task_struct *p, s32 prev_cpu, u64
> #ifdef CONFIG_EXT_SUB_SCHED
> /*
> * Disallow if any sub-scheds are attached. There is no way to tell
> - * which scheduler called us, just error out @p's scheduler.
> + * which scheduler called us, so error out @p's scheduler -- but read
> + * it under RCU (@p's locks aren't held here) and fall back to @sch if
> + * @p isn't on an scx scheduler: a BPF_PROG_TYPE_SYSCALL prog can pass
> + * any task and p->scx.sched is NULL for one that has exited or is
> + * managed by another scheduler.
> */
> if (unlikely(!list_empty(&sch->children))) {
> - scx_error(scx_task_sched(p), "__scx_bpf_select_cpu_and() must be used");
> + scx_error(scx_task_sched_rcu(p) ?: sch,
> + "__scx_bpf_select_cpu_and() must be used");
> return -EINVAL;
> }
> #endif
> --
> 2.25.1
>
next prev parent reply other threads:[~2026-09-02 18:11 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 15:36 [PATCH] sched_ext: Fix NULL sched deref in select_cpu_and sub-sched error path Wanwu Li
2026-09-02 15:58 ` sashiko-bot
2026-09-02 16:14 ` Andrea Righi
2026-09-02 16:34 ` liwanwu
2026-09-02 17:07 ` [PATCH v2] sched_ext: Fix NULL sched deref in kfunc sub-sched error paths Wanwu Li
2026-09-02 18:11 ` Andrea Righi [this message]
2026-09-02 19:37 ` Tejun Heo
2026-09-02 19:39 ` Tejun Heo
2026-09-03 6:06 ` [PATCH v3] " Wanwu Li
2026-09-03 17:56 ` Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aphmwjLkdS4tKpbv@gpd4 \
--to=arighi@nvidia.com \
--cc=changwoo@igalia.com \
--cc=linux-kernel@vger.kernel.org \
--cc=liwanwu@kylinos.cn \
--cc=sched-ext@lists.linux.dev \
--cc=stable@vger.kernel.org \
--cc=tj@kernel.org \
--cc=void@manifault.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.