From: Andrea Righi <arighi@nvidia.com>
To: Tejun Heo <tj@kernel.org>
Cc: sched-ext@lists.linux.dev, David Vernet <void@manifault.com>,
Changwoo Min <changwoo@igalia.com>,
Emil Tsalapatis <emil@etsalapatis.com>,
David Dai <david.dai@linux.dev>,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH 1/2] sched_ext: Add ops.sub_child_ecaps_updated() to report a child's effective cap changes
Date: Fri, 9 Oct 2026 13:17:15 +0200 [thread overview]
Message-ID: <asjNO0ggCPcfIP7Y@gpd4> (raw)
In-Reply-To: <20261008093228.2015427-2-tj@kernel.org>
Hi Tejun,
On Wed, Oct 07, 2026 at 11:32:27PM -1000, Tejun Heo wrote:
> A grant or revoke only records the target caps. They take effect on a cid at
> its next dispatch, and nothing tells the parent when. Without that the
> parent cannot act on a revoke: a child that ran a cid at a low cpuperf
> target leaves the target there when PERF is revoked, as the kernel resets
> targets only at root enable. Nor can the parent tell when it may schedule on
> the cid again.
>
> Add ops.sub_child_ecaps_updated(), delivered to the direct parent right
> after the child's own ops.sub_ecaps_updated() with the child's cgroup id and
> the same before and after caps, in the same dispatch context. Both
> deliveries are suppressed while the child is bypassing and replayed together
> afterwards. A disabled child reports nothing: ops.sub_detach() is where the
> parent restores what it had delegated.
>
> A sub is now bypassed before it is linked, so that the grants queued for it
> from ops.sub_attach() are consumed while it is bypassed and replayed to it
> and the parent once it is enabled. Otherwise a sync consumed before the
> sub's ops are registered would reach the parent but never the sub.
>
> v2: Drop the disable report with its sleep lock and nested-dispatch gate,
> the parent restores from ops.sub_detach() instead.
>
> Signed-off-by: Tejun Heo <tj@kernel.org>
> ---
...
> diff --git a/kernel/sched/ext/sub.c b/kernel/sched/ext/sub.c
> index 51a53467cf98..c7c95661280b 100644
> --- a/kernel/sched/ext/sub.c
> +++ b/kernel/sched/ext/sub.c
> @@ -1121,27 +1121,43 @@ void __scx_process_sync_ecaps(struct rq *rq, struct task_struct *prev)
> lost_all |= lost;
>
> /*
> - * Tell the sched its effective caps on this cid changed. The
> - * invocation is equivalent to the dispatch path and may drop
> - * and re-acquire the rq lock temporarily while the rest of
> - * @batch is held privately, see scx_discard_ecaps_to_sync().
> - * The dispatch kfuncs resolve their context on the executing
> - * cpu, which under core scheduling can differ from @rq's cpu,
> - * so the context is set up there. The rq recorded in it keeps
> - * the dispatches targeting @rq.
> + * Tell the sched and its parent that the sched's effective caps
> + * on this cid changed. The invocations are equivalent to the
> + * dispatch path and may drop and re-acquire the rq lock
> + * temporarily while the rest of @batch is held privately, see
> + * scx_discard_ecaps_to_sync(). The dispatch kfuncs resolve
> + * their context on the executing cpu, which under core
> + * scheduling can differ from @rq's cpu, so the context is set
> + * up there. The rq recorded in it keeps the dispatches
> + * targeting @rq.
> + *
> + * Bypass propagates down the hierarchy, so a sched that isn't
> + * bypassing has no bypassing parent. The child's bypass state
> + * gates both deliveries. Both report the same before value, so
> + * one reported_ecaps covers them.
> */
> - if (ecaps != pcpu->reported_ecaps &&
> - SCX_HAS_OP(pcpu->sch, sub_ecaps_updated) &&
> - !scx_bypassing(pcpu->sch, cpu)) {
> - struct scx_dsp_ctx *dspc = &this_cpu_ptr(pcpu->sch->pcpu)->dsp_ctx;
> + if (ecaps != pcpu->reported_ecaps && !scx_bypassing(pcpu->sch, cpu)) {
> + struct scx_sched *parent = scx_parent(pcpu->sch);
> + struct scx_dsp_ctx *dspc;
Should ops.sub_child_ecaps_updated() still be delivered to the parent when the
child is bypassing but the parent is not?
For example, the shared pool may rotate away from a bypassing child. IIUC, its
PERF revoke takes effect at the next dispatch, but the child's bypass state
suppresses the parent notification too. If the child is being disabled, it never
leaves bypass, so the parent only gets ops.sub_detach() and cannot tell when the
revoke took effect.
Could we notify the parent when the revoke takes effect, even if the child is
bypassing, while continuing to defer the child's own notification until it
leaves bypass?
Thanks,
-Andrea
>
> - dspc->rq = rq;
> /* stash @prev so nested dispatches can access it */
> rq->scx.sub_dispatch_prev = prev;
> - SCX_CALL_OP(pcpu->sch, sub_ecaps_updated, rq, scx_cpu_arg(cpu),
> - pcpu->reported_ecaps, ecaps);
> + if (SCX_HAS_OP(pcpu->sch, sub_ecaps_updated)) {
> + dspc = &this_cpu_ptr(pcpu->sch->pcpu)->dsp_ctx;
> + dspc->rq = rq;
> + SCX_CALL_OP(pcpu->sch, sub_ecaps_updated, rq,
> + scx_cpu_arg(cpu), pcpu->reported_ecaps, ecaps);
> + scx_flush_dispatch_buf(pcpu->sch, rq);
> + }
> + if (SCX_HAS_OP(parent, sub_child_ecaps_updated)) {
> + dspc = &this_cpu_ptr(parent->pcpu)->dsp_ctx;
> + dspc->rq = rq;
> + SCX_CALL_OP(parent, sub_child_ecaps_updated, rq,
> + pcpu->sch->ops.sub_cgroup_id, scx_cpu_arg(cpu),
> + pcpu->reported_ecaps, ecaps);
> + scx_flush_dispatch_buf(parent, rq);
> + }
> rq->scx.sub_dispatch_prev = NULL;
> - scx_flush_dispatch_buf(pcpu->sch, rq);
> pcpu->reported_ecaps = ecaps;
> }
>
> @@ -1940,6 +1956,15 @@ void scx_sub_enable_workfn(struct kthread_work *work)
> if (ret)
> goto err_disable;
>
> + /*
> + * Bypass before @sch is linked and grants can reach it. The parent's
> + * delivery advances reported_ecaps, so a sync consumed before @sch's
> + * ops are registered would reach the parent and never be replayed to
> + * @sch. While bypassed, the syncs are consumed without a delivery and
> + * replayed to both at unbypass.
> + */
> + scx_bypass(sch, true);
> +
> ret = scx_link_sched(sch);
> if (ret)
> goto err_disable;
> @@ -1985,8 +2010,6 @@ void scx_sub_enable_workfn(struct kthread_work *work)
> }
> sch->sub_attached = true;
>
> - scx_bypass(sch, true);
> -
> for (i = SCX_OPI_BEGIN; i < SCX_OPI_END; i++)
> if (((void (**)(void))ops)[i])
> set_bit(i, sch->has_op);
> --
> 2.55.0
>
next prev parent reply other threads:[~2026-10-09 11:17 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-08 9:32 [PATCHSET v2 sched_ext/for-7.4] sched_ext: Add ops.sub_child_ecaps_updated() Tejun Heo
2026-10-08 9:32 ` [PATCH 1/2] sched_ext: Add ops.sub_child_ecaps_updated() to report a child's effective cap changes Tejun Heo
2026-10-09 11:17 ` Andrea Righi [this message]
2026-10-09 20:04 ` Tejun Heo
2026-10-09 20:14 ` Andrea Righi
2026-10-08 9:32 ` [PATCH 2/2] sched_ext: scx_qmap: Reset a child's cpuperf targets once its PERF revoke takes effect Tejun Heo
2026-10-09 11:18 ` Andrea Righi
2026-10-09 20:04 ` Tejun Heo
2026-10-09 20:48 ` [PATCH v3 " Tejun Heo
2026-10-09 21:38 ` Andrea Righi
2026-10-09 21:44 ` [PATCHSET v2 sched_ext/for-7.4] sched_ext: Add ops.sub_child_ecaps_updated() Tejun Heo
-- strict thread matches above, loose matches on Subject: below --
2026-10-08 0:03 [PATCHSET " Tejun Heo
2026-10-08 0:03 ` [PATCH 1/2] sched_ext: Add ops.sub_child_ecaps_updated() to report a child's effective cap changes Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asjNO0ggCPcfIP7Y@gpd4 \
--to=arighi@nvidia.com \
--cc=changwoo@igalia.com \
--cc=david.dai@linux.dev \
--cc=emil@etsalapatis.com \
--cc=linux-kernel@vger.kernel.org \
--cc=sched-ext@lists.linux.dev \
--cc=tj@kernel.org \
--cc=void@manifault.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox