From: Tejun Heo <tj@kernel.org>
To: Qiurong Fang <fangqiurong@kylinos.cn>
Cc: arighi@nvidia.com, sched-ext@lists.linux.dev
Subject: Re: [PATCH] sched_ext: Delete sub-scheduler kobjects before releasing scx_enable_mutex
Date: Tue, 1 Sep 2026 20:02:54 -1000 [thread overview]
Message-ID: <ape8DkMkIdWeyAd7@slm.duckdns.org> (raw)
In-Reply-To: <20260902032233.135959-1-fangqiurong@kylinos.cn>
Hello,
On Wed, Sep 02, 2026 at 11:22:33AM +0800, Qiurong Fang wrote:
> Hello Tejun,
>
> On Tue, Sep 01, 2026 at 09:56:51AM -1000, Tejun Heo wrote:
> > Why is this a problem? The subsched hasn't been fully unloaded yet so if you
> > try to attach a new one, it's going to fail.
>
> When the mutex is released, the cgroup has already been given back to the
> parent and the exiting scheduler unlinked - its sysfs name is the only
> thing left, and only kobject_add() can see it:
>
> exiting sched (disable workfn) new sched (enable workfn)
> ----------------------------- ---------------------------
> mutex_lock(&scx_enable_mutex)
> set_cgroup_sched(cgrp, parent)
> scx_unlink_sched(sch)
> mutex_unlock(&scx_enable_mutex)
> mutex_lock(&scx_enable_mutex)
> find_parent_sched() ok
> create/validate/link sch ok
> scx_sched_sysfs_add()
> kobject_add("sub-%llu") -EEXIST
> err_disable: tear down the
> otherwise healthy sched
> ops.sub_detach()/exit() ~100ms
> synchronize_rcu()
> kobject_del("sub-%llu") too late
>
> The new scheduler passes every check and dies only at sysfs_add(); the
> error surfaces via its own ops.exit() while the attach syscall returns 0,
> and there is nothing userspace could wait on. This is the case the root
> path's comment covers by deleting before unlocking.
>
> Reproduced with a slow (~100ms) ops.exit() and a second struct_ops
> instance attaching to the same cgroup in a loop: 40/40 iterations hit
> -EEXIST without the patch, 0/40 with it. Happy to turn the reproducer
> into a selftest as a follow-up if wanted.
How is that distinguishible from new sched being loaded while the previous
one is still there? There are two ways to observe whether a scheduler is in
place - bpf link destruction, which cna propagate through the sched process
exit, and sysfs kobject visibility and the associated uevent notifications.
ops.sub_detach/exit() being called doesn't mean that a subsched is gone -
these are internal cleanup routines. You don't generate userspace visible
indications from those.
You're trying to load a new scheduler when nobody told anybody that the
previous one is gone, declaring that not succeeding a bug and then fixing
that by introducing an actual bug - now you're telling userspace that the
scheduler is gone via uevent while the scheduler is *still* there. Please
stop.
Thanks.
--
tejun
next prev parent reply other threads:[~2026-09-02 6:02 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-01 10:22 [PATCH] sched_ext: Delete sub-scheduler kobjects before releasing scx_enable_mutex Qiurong Fang
2026-09-01 19:56 ` Tejun Heo
2026-09-02 3:22 ` Qiurong Fang
2026-09-02 6:02 ` Tejun Heo [this message]
2026-09-01 20:17 ` Andrea Righi
-- strict thread matches above, loose matches on Subject: below --
2026-09-02 6:20 Qiurong Fang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ape8DkMkIdWeyAd7@slm.duckdns.org \
--to=tj@kernel.org \
--cc=arighi@nvidia.com \
--cc=fangqiurong@kylinos.cn \
--cc=sched-ext@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox