From: Tejun Heo <tj@kernel.org>
To: Qiurong Fang <fangqiurong@kylinos.cn>
Cc: arighi@nvidia.com, sched-ext@lists.linux.dev
Subject: Re: [PATCH] sched_ext: Delete sub-scheduler kobjects before releasing scx_enable_mutex
Date: Tue, 1 Sep 2026 20:02:54 -1000 [thread overview]
Message-ID: <ape8DkMkIdWeyAd7@slm.duckdns.org> (raw)
In-Reply-To: <20260902032233.135959-1-fangqiurong@kylinos.cn>
Hello,
On Wed, Sep 02, 2026 at 11:22:33AM +0800, Qiurong Fang wrote:
> Hello Tejun,
>
> On Tue, Sep 01, 2026 at 09:56:51AM -1000, Tejun Heo wrote:
> > Why is this a problem? The subsched hasn't been fully unloaded yet so if you
> > try to attach a new one, it's going to fail.
>
> When the mutex is released, the cgroup has already been given back to the
> parent and the exiting scheduler unlinked - its sysfs name is the only
> thing left, and only kobject_add() can see it:
>
> exiting sched (disable workfn) new sched (enable workfn)
> ----------------------------- ---------------------------
> mutex_lock(&scx_enable_mutex)
> set_cgroup_sched(cgrp, parent)
> scx_unlink_sched(sch)
> mutex_unlock(&scx_enable_mutex)
> mutex_lock(&scx_enable_mutex)
> find_parent_sched() ok
> create/validate/link sch ok
> scx_sched_sysfs_add()
> kobject_add("sub-%llu") -EEXIST
> err_disable: tear down the
> otherwise healthy sched
> ops.sub_detach()/exit() ~100ms
> synchronize_rcu()
> kobject_del("sub-%llu") too late
>
> The new scheduler passes every check and dies only at sysfs_add(); the
> error surfaces via its own ops.exit() while the attach syscall returns 0,
> and there is nothing userspace could wait on. This is the case the root
> path's comment covers by deleting before unlocking.
>
> Reproduced with a slow (~100ms) ops.exit() and a second struct_ops
> instance attaching to the same cgroup in a loop: 40/40 iterations hit
> -EEXIST without the patch, 0/40 with it. Happy to turn the reproducer
> into a selftest as a follow-up if wanted.
How is that distinguishible from new sched being loaded while the previous
one is still there? There are two ways to observe whether a scheduler is in
place - bpf link destruction, which cna propagate through the sched process
exit, and sysfs kobject visibility and the associated uevent notifications.
ops.sub_detach/exit() being called doesn't mean that a subsched is gone -
these are internal cleanup routines. You don't generate userspace visible
indications from those.
You're trying to load a new scheduler when nobody told anybody that the
previous one is gone, declaring that not succeeding a bug and then fixing
that by introducing an actual bug - now you're telling userspace that the
scheduler is gone via uevent while the scheduler is *still* there. Please
stop.
Thanks.
--
tejun
next prev parent reply other threads:[~2026-09-02 6:02 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-01 10:22 [PATCH] sched_ext: Delete sub-scheduler kobjects before releasing scx_enable_mutex Qiurong Fang
2026-09-01 19:56 ` Tejun Heo
2026-09-02 3:22 ` Qiurong Fang
2026-09-02 6:02 ` Tejun Heo [this message]
2026-09-01 20:17 ` Andrea Righi
-- strict thread matches above, loose matches on Subject: below --
2026-09-02 6:20 Qiurong Fang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ape8DkMkIdWeyAd7@slm.duckdns.org \
--to=tj@kernel.org \
--cc=arighi@nvidia.com \
--cc=fangqiurong@kylinos.cn \
--cc=sched-ext@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.