From: Tao Cui <cui.tao@linux.dev>
To: tj@kernel.org, void@manifault.com
Cc: cui.tao@linux.dev, arighi@nvidia.com, changwoo@igalia.com,
suzhidao@xiaomi.com, sched-ext@lists.linux.dev,
linux-kernel@vger.kernel.org, bpf@vger.kernel.org,
Tao Cui <cuitao@kylinos.cn>
Subject: Re: [PATCH RFC] sched_ext: warn when cpu.max is set but the BPF scheduler doesn't implement bandwidth control
Date: Tue, 18 Aug 2026 21:58:25 +0800 [thread overview]
Message-ID: <ceb9f4e0-5fe3-426a-bdc7-42d918f87fab@linux.dev> (raw)
In-Reply-To: <20260818135328.174152-1-cui.tao@linux.dev>
Hi, tejun
在 2026/8/18 21:53, Tao Cui 写道:
> From: Tao Cui <cuitao@kylinos.cn>
>
> The kernel stores cpu.max bandwidth parameters in the task_group and
> passes them to the BPF scheduler via ops.cgroup_set_bandwidth() and
> scx_cgroup_init_args, but does not enforce the quota itself. If the
> loaded BPF scheduler doesn't implement the callback, cpu.max is
> silently ignored -- the cgroup gets unlimited CPU regardless of the
> configured quota.
>
> Of the example schedulers, only scx_qmap implements the callback --
> and only to bpf_printk() the parameters, so no in-tree scheduler
> actually enforces the quota. Measured with scx_simple: a
> cgroup with cpu.max = "50000 100000" (50% of one CPU) and one
> busy task used 9946ms of CPU in 10 seconds with nr_throttled
> remaining 0.
>
> Print a one-time warning when a finite quota is configured on a
> cgroup while the active scheduler lacks the callback, so users and
> container orchestrators know the quota is not enforced.
>
Some background on how I found this: I was testing how sched_ext
interacts with cgroup CPU controls in a VM, and the cpu.max case stood
out:
cgroup with cpu.max = "50000 100000" (50% of one CPU),
one busy task, scx_simple loaded
sched_ext : 9946ms of CPU in 10s, nr_throttled = 0
CFS : ~5000ms in 10s, throttling as expected
The same happens with scx_flatcg and scx_central -- neither they nor
scx_simple implement ops.cgroup_set_bandwidth(), so the quota is
silently ignored. grep shows scx_qmap is the only in-tree user of the
callback -- and its implementation just bpf_printk()s the parameters, so
even there the quota is not enforced. Outside the
tree, lavd implements its own bandwidth accounting, but as far as I can
tell rusty and bpfland don't, which suggests users on those schedulers
are running containers with cpu.max that does nothing.
That's why I drafted the warning patch -- but I'm not sure a warning is
the right approach. Some alternatives I can think of:
1. pr_warn_once() as in this patch (minimal, but the quota is still
not enforced)
2. refuse to enable sched_ext (or the cgroup support) when a finite
quota exists and the callback is missing
3. kernel-side fallback enforcement, e.g. throttle in
scx_next_task_picked()/dispatch path based on tg->scx.bw_quota_us
Is the missing enforcement considered the BPF scheduler's responsibility
by design (and just under-documented), or would a kernel fallback be
welcome? If it's the former, maybe sched-ext.rst should mention that
cpu.max requires ops.cgroup_set_bandwidth() from the loaded scheduler.
Happy to work on whichever direction you prefer.
Thanks,
Tao
> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
> ---
> kernel/sched/ext/ext.c | 6 ++++++
> 1 file changed, 6 insertions(+)
>
> diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
> index 10af28a9f2c0..de786d0b928a 100644
> --- a/kernel/sched/ext/ext.c
> +++ b/kernel/sched/ext/ext.c
> @@ -4953,6 +4953,12 @@ void scx_group_set_bandwidth(struct task_group *tg,
> tg->scx.bw_burst_us != burst_us))
> SCX_CALL_OP(sch, cgroup_set_bandwidth, NULL,
> tg_cgrp(tg), period_us, quota_us, burst_us);
> + else if (scx_cgroup_enabled && sch &&
> + !SCX_HAS_OP(sch, cgroup_set_bandwidth) &&
> + quota_us != RUNTIME_INF)
> + pr_warn_once("sched_ext: BPF scheduler \"%s\" does not implement "
> + "ops.cgroup_set_bandwidth(); cpu.max will not be enforced\n",
> + sch->ops.name);
>
> tg->scx.bw_period_us = period_us;
> tg->scx.bw_quota_us = quota_us;
next prev parent reply other threads:[~2026-08-18 13:58 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-18 13:53 [PATCH RFC] sched_ext: warn when cpu.max is set but the BPF scheduler doesn't implement bandwidth control Tao Cui
2026-08-18 13:58 ` Tao Cui [this message]
2026-08-18 14:10 ` sashiko-bot
2026-08-18 15:34 ` Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ceb9f4e0-5fe3-426a-bdc7-42d918f87fab@linux.dev \
--to=cui.tao@linux.dev \
--cc=arighi@nvidia.com \
--cc=bpf@vger.kernel.org \
--cc=changwoo@igalia.com \
--cc=cuitao@kylinos.cn \
--cc=linux-kernel@vger.kernel.org \
--cc=sched-ext@lists.linux.dev \
--cc=suzhidao@xiaomi.com \
--cc=tj@kernel.org \
--cc=void@manifault.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox