From: Andrea Righi <arighi@nvidia.com>
To: Wanwu Li <liwanwu@kylinos.cn>
Cc: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
Changwoo Min <changwoo@igalia.com>,
sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org,
stable@vger.kernel.org
Subject: Re: [PATCH] sched_ext: Fix NULL sched deref in select_cpu_and sub-sched error path
Date: Wed, 2 Sep 2026 18:14:18 +0200 [thread overview]
Message-ID: <aphLWtCVY1XVME9C@gpd4> (raw)
In-Reply-To: <20260902153640.144791-1-liwanwu@kylinos.cn>
Hi Wanwu,
On Wed, Sep 02, 2026 at 11:36:40PM +0800, Wanwu Li wrote:
> scx_bpf_select_cpu_and() errors out @p's scheduler when the root scheduler
> has sub-scheds attached:
>
> scx_error(scx_task_sched(p), "... must be used");
>
> scx_task_sched(p) is p->scx.sched, which is NULL for any task that is not
> on an scx scheduler: it is memset() by init_scx_entity() and cleared by
> scx_disable_and_exit_task() -- which sched_ext_dead() runs when a task
> exits. It is also an rcu_dereference_protected() that must be called with
> @p's pi_lock or rq lock held -- neither of which a BPF_PROG_TYPE_SYSCALL
> program holds.
>
> The wrapper is reachable from such a program -- scx_kfunc_context_filter()
> allows the select_cpu kfunc group for BPF_PROG_TYPE_SYSCALL -- and the
> program can pass any task, e.g. one obtained with bpf_task_from_pid() that
> exited in between. scx_error() then calls scx_vexit(), which dereferences
> sch->exit_info unconditionally, so passing NULL oopses the kernel.
>
> This was triggered live on a v7.2 based kernel with a
> BPF_PROG_TYPE_SYSCALL program calling the wrapper on an exited-but-not
> reaped task while a sub-scheduler was attached (faulting instruction is
> the scx_vexit() prologue "mov r15,[rdi+0x398]" with RDI=NULL and 0x398
> the offset of sch->exit_info):
>
> sched_ext: BPF scheduler "kfunc_subsched_null" enabled
> sched_ext: BPF sub-scheduler "kfunc_subsched_null" enabled
> sched_ext: Unassociated program run_select_cpu_ (id 76)
> BUG: kernel NULL pointer dereference, address: 0000000000000398
> #PF: supervisor read access in kernel mode
> #PF: error_code(0x0000) - not-present page
> Oops: Oops: 0000 [#1] SMP NOPTI
> CPU: 7 UID: 0 PID: 8201 Comm: kfunc_test_runn Tainted: G W
> RIP: 0010:scx_vexit+0x25/0xa0
> Code: ... <4c> 8b bf 98 03 00 00 ...
> CR2: 0000000000000398
> Call Trace:
> <TASK>
> __scx_exit+0x4f/0x70
> scx_bpf_select_cpu_and+0xab/0xb0
> bpf_prog_430ed61a7b66e03a_run_select_cpu_and+0x9c/0xe7
> ? __x64_sys_bpf+0x2c/0x40
> bpf_prog_test_run_syscall+0x130/0x2f0
> __sys_bpf+0x930/0x10d0
> ? __x64_sys_bpf+0x2c/0x40
> __x64_sys_bpf+0x2c/0x40
> do_syscall_64+0xbc/0x460
> ? rseq_set_ids_get_csaddr+0x81/0x140
> ? __rseq_handle_slowpath+0xd0/0x130
> ? switch_fpu_return+0x51/0xd0
> ? arch_exit_to_user_mode_prepare.constprop.0+0x87/0xb0
> ? do_syscall_64+0xf3/0x460
> ? irqentry_exit+0x48/0x740
> ? clear_bhb_loop+0x40/0x90
> ? do_syscall_64+0x35/0x460
> entry_SYSCALL_64_after_hwframe+0x76/0x7e
> </TASK>
>
> Keep the "error out @p's scheduler" attribution -- it is what every other
> kfunc error path does (select_cpu_from_kfunc()'s cross_task,
> scx_kf_arg_task_ok()) and it is correct for the callers that actually take
> this path: the only other callers are the struct_ops select_cpu/enqueue
> ops, where @p is the caller's own task and p->scx.sched is its scheduler.
> Just read it safely: use scx_task_sched_rcu() (valid under the guard(rcu)()
> the wrapper already holds, no @p lock required) and fall back to @sch -- the
> root scheduler, guaranteed non-NULL here -- when @p is not on an scx
> scheduler, which is precisely the case that used to be NULL.
>
> scx_bpf_dsq_insert_vtime() has the same error path but it is not reachable
> with a NULL @p: SYSCALL programs are rejected for its kfunc set and @p is
> always the calling scheduler's own task in the contexts where it runs, so
> it is left unchanged.
The reported crash looks valid to me, and using scx_task_sched_rcu() with the
root scheduler as fallback also looks correct.
However, as also pointed out by sashiko, the assumption above doesn't hold for
scx_bpf_dsq_insert_vtime(), although SYSCALL programs can't call it, STRUCT_OPS
programs can call it from ops.enqueue() and ops.dispatch().
Can you update scx_bpf_dsq_insert_vtime() as well with the same fallback?
Thanks,
-Andrea
>
> The proper endgame for these COMPAT wrappers is removal once the
> deprecation grace period is announced and elapsed; this fix only keeps
> the window from oopsing the kernel until that happens.
>
> Cc: <stable@vger.kernel.org>
> Fixes: a5fa0708cbfd ("sched_ext: Enforce scheduling authority in dispatch and select_cpu operations")
> Signed-off-by: Wanwu Li <liwanwu@kylinos.cn>
> ---
> kernel/sched/ext/idle.c | 9 +++++++--
> 1 file changed, 7 insertions(+), 2 deletions(-)
>
> diff --git a/kernel/sched/ext/idle.c b/kernel/sched/ext/idle.c
> index d2973fb3af6d..014599d82bb0 100644
> --- a/kernel/sched/ext/idle.c
> +++ b/kernel/sched/ext/idle.c
> @@ -1142,10 +1142,15 @@ __bpf_kfunc s32 scx_bpf_select_cpu_and(struct task_struct *p, s32 prev_cpu, u64
> #ifdef CONFIG_EXT_SUB_SCHED
> /*
> * Disallow if any sub-scheds are attached. There is no way to tell
> - * which scheduler called us, just error out @p's scheduler.
> + * which scheduler called us, so error out @p's scheduler -- but read
> + * it under RCU (@p's locks aren't held here) and fall back to @sch if
> + * @p isn't on an scx scheduler: a BPF_PROG_TYPE_SYSCALL prog can pass
> + * any task and p->scx.sched is NULL for one that has exited or is
> + * managed by another scheduler.
> */
> if (unlikely(!list_empty(&sch->children))) {
> - scx_error(scx_task_sched(p), "__scx_bpf_select_cpu_and() must be used");
> + scx_error(scx_task_sched_rcu(p) ?: sch,
> + "__scx_bpf_select_cpu_and() must be used");
> return -EINVAL;
> }
> #endif
> --
> 2.25.1
>
next prev parent reply other threads:[~2026-09-02 16:14 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 15:36 [PATCH] sched_ext: Fix NULL sched deref in select_cpu_and sub-sched error path Wanwu Li
2026-09-02 15:58 ` sashiko-bot
2026-09-02 16:14 ` Andrea Righi [this message]
2026-09-02 16:34 ` liwanwu
2026-09-02 17:07 ` [PATCH v2] sched_ext: Fix NULL sched deref in kfunc sub-sched error paths Wanwu Li
2026-09-02 18:11 ` Andrea Righi
2026-09-02 19:37 ` Tejun Heo
2026-09-02 19:39 ` Tejun Heo
2026-09-03 6:06 ` [PATCH v3] " Wanwu Li
2026-09-03 17:56 ` Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aphLWtCVY1XVME9C@gpd4 \
--to=arighi@nvidia.com \
--cc=changwoo@igalia.com \
--cc=linux-kernel@vger.kernel.org \
--cc=liwanwu@kylinos.cn \
--cc=sched-ext@lists.linux.dev \
--cc=stable@vger.kernel.org \
--cc=tj@kernel.org \
--cc=void@manifault.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.