From: Tao Cui <cui.tao@linux.dev>
To: Andrea Righi <arighi@nvidia.com>
Cc: cui.tao@linux.dev, tj@kernel.org, void@manifault.com,
changwoo@igalia.com, suzhidao@xiaomi.com,
sched-ext@lists.linux.dev, linux-kernel@vger.kernel.org,
bpf@vger.kernel.org, Tao Cui <cuitao@kylinos.cn>
Subject: Re: [PATCH v3] docs/sched_ext: document that cgroup CPU knobs are scheduler-dependent
Date: Mon, 24 Aug 2026 21:08:45 +0800 [thread overview]
Message-ID: <1b2f54bd-97f5-4175-9a92-34871dd6f375@linux.dev> (raw)
In-Reply-To: <aowxJK_KgBLLnyiN@gpd4>
Hello, Andrea,
在 2026/8/24 19:55, Andrea Righi 写道:
> Hi Tao,
>
> On Mon, Aug 24, 2026 at 05:15:01PM +0800, Tao Cui wrote:
>> From: Tao Cui <cuitao@kylinos.cn>
>>
>> The fair class enforces cpu controller knobs such as cpu.max,
>> cpu.weight and cpu.idle in the kernel. sched_ext only passes them to
>> the BPF scheduler through the ops.cgroup_set_*() callbacks. Whether
>> and how a knob takes effect is up to the loaded scheduler: it may
>> implement the corresponding callback partially or not at all. The
>> same applies to other knobs like nice levels.
>>
>> Document this in the basics section so users and container
>> orchestrators know what to expect from a BPF scheduler.
>>
>> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
>> ---
>> v2 -> v3: Drop the scheduler list and the nr_throttled example, keep
>> the section concise and generic, per review.
>>
>> v2: https://lore.kernel.org/r/20260819012157.220932-1-cui.tao@linux.dev
>>
>> Documentation/scheduler/sched-ext.rst | 12 ++++++++++++
>> 1 file changed, 12 insertions(+)
>>
>> diff --git a/Documentation/scheduler/sched-ext.rst b/Documentation/scheduler/sched-ext.rst
>> index 35b550671ca7..b594d93dd6aa 100644
>> --- a/Documentation/scheduler/sched-ext.rst
>> +++ b/Documentation/scheduler/sched-ext.rst
>> @@ -242,6 +242,18 @@ optional. The following modified excerpt is from
>> .name = "simple",
>> };
>>
>> +Scheduler-Dependent Knobs
>> +-------------------------
>> +
>> +The fair class enforces cpu controller knobs such as ``cpu.max``,
>> +``cpu.weight`` and ``cpu.idle`` in the kernel. sched_ext only passes
>> +them to the BPF scheduler through ``ops.cgroup_set_weight()``,
>> +``ops.cgroup_set_idle()``, ``ops.cgroup_set_bandwidth()`` and friends.
>
> Existing cpu.weight and bandwidth values are delivered via ops.cgroup_init(),
> ops.cgroup_set_*() callbacks handle later changes.
>
> Moreover, cpu.idle state is not included in scx_cgroup_init_args, so apparently
> BPF schedulers don't receive an initial value (only subsequent writes). This
> should be probably fixed separately.
>
>> +Whether and how a knob takes effect is up to the loaded scheduler: it
>> +may implement the corresponding callback partially or not at all. The
>> +same applies to other knobs like nice levels. When relying on these
>> +knobs, check the documentation or source of the loaded scheduler.
>> +
>
> "nice levels" can be a bit ambiguous, per-task nice changes are converted to
> weights and reported through ops.set_weight(), writes to the cgroup file
> cpu.weight.nice are reported through ops.cgroup_set_weight().
>
> Maybe we can rephrase the whole paragraph as following, or something along these
> lines:
>
> The fair-class scheduler enforces CPU controller settings such as cpu.max,
> cpu.weight, and cpu.idle. For sched_ext tasks, the scheduler core communicates
> these settings to the BPF scheduler through ops.cgroup_init() and reports
> subsequent changes through the corresponding ops.cgroup_set_*() callbacks.
> Similarly, per-task nice changes are converted to weights and reported through
> ops.set_weight().
>
> Each BPF scheduler is responsible for implementing the scheduling semantics of
> these settings and may choose to ignore them. Consult the loaded scheduler's
> documentation before relying on these controls.
>
Thanks for the review. v4 adopts your wording: ops.cgroup_init()
for the initial values, ops.cgroup_set_*() only for subsequent
changes, and per-task nice disambiguated from cpu.weight.nice via
ops.set_weight().
On the missing cpu.idle initial value in scx_cgroup_init_args: I can
look into a separate fix for that.
Thanks,
Tao
> Thanks,
> -Andrea
prev parent reply other threads:[~2026-08-24 13:09 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-24 9:15 [PATCH v3] docs/sched_ext: document that cgroup CPU knobs are scheduler-dependent Tao Cui
2026-08-24 11:55 ` Andrea Righi
2026-08-24 13:08 ` Tao Cui [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1b2f54bd-97f5-4175-9a92-34871dd6f375@linux.dev \
--to=cui.tao@linux.dev \
--cc=arighi@nvidia.com \
--cc=bpf@vger.kernel.org \
--cc=changwoo@igalia.com \
--cc=cuitao@kylinos.cn \
--cc=linux-kernel@vger.kernel.org \
--cc=sched-ext@lists.linux.dev \
--cc=suzhidao@xiaomi.com \
--cc=tj@kernel.org \
--cc=void@manifault.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.