Sched_ext development
 help / color / mirror / Atom feed
* [REPORT] sched_ext: sub-scheduler tasks are severely under-scheduled vs root-owned tasks
@ 2026-10-08  8:36 Tao Cui
  2026-10-08  9:10 ` Tejun Heo
  0 siblings, 1 reply; 4+ messages in thread
From: Tao Cui @ 2026-10-08  8:36 UTC (permalink / raw)
  To: Tejun Heo
  Cc: cui.tao, David Vernet, sched-ext, Andrea Righi, linux-kernel,
	Changwoo Min

Hello Tejun,

During recent testing of the sub-scheduler attach flow, I hit what
looks like a dispatch starvation for claimed tasks and want to check
whether it's a known limitation of the sub-sched dispatch model before
digging further.

Setup:

A kvm guest (4 vCPUs, HZ=1000) running a sched_ext/for-7.4 based tree
with the caps-clear patch applied. A pair of minimal cid-form test
schedulers: a parent that grants SCX_CAP_PERF | SCX_CAP_ENQ_IMMED over
a cgroup to a child, and the child writing cpuperf target 1 from
ops.tick(). A busy-loop task (`taskset 1 sh -c 'while :; do :; done'`)
is placed in the child's cgroup (pinned to cpu0), and
rq->scx.cpuperf_target plus per-task CPU time are sampled every second.

Two attach orders are tested:
(a) the busy-loop task is moved into the cgroup before the child
    attaches;
(b) the task is moved in after the child attaches.

Observation:

Order (a), 8 rounds: the child claims the task in every round
(verified: the enable walk's handover fires, p->scx.sched == child),
but the task then runs at ~0.3% duty cycle - it accumulates 0 or
~20-27 ticks per 10s window (one SCX_SLICE_DFL slice) while an
identical task owned by the parent runs at ~100% (1000 ticks/s).

Order (b), 8 rounds: the child claims the task via the migration path
and the task runs normally (no starvation observed in 16 rounds across
two parent-load conditions).

Based on the current test data, the disparity is ~300x: root-owned
tasks run at full duty while child-owned tasks get one slice and then
starve. The starvation is silent: attach succeeds, the grant succeeds,
the task's cgroup membership is correct - there is no error, warning,
or counter that indicates anything is wrong.

Probes and exclusions:

- ops.dispatch() for the child's cid is never invoked: a counter
  bumped in ops.dispatch() when cid == 0 stays 0 across entire runs,
  including rounds where the task is starved.
- ops.tick() for the child fires only when the task actually runs
  (the tick count matches the starved duty cycle).
- The parent's ops.dispatch() delegation via
  scx_bpf_sub_dispatch(cgroup_id) does not change the behavior.
- An explicit scx_bpf_kick_cid() on the task's cid in ops.enqueue()
  does not change the behavior.
- The enable walk's claim is verified: pass 1 (init) and pass 2
  (handover) both process the task, and a post-handover ownership
  check (scx_task_on_sched()) passes.
- Adding a root-owned busy task on another CPU raises the failure rate
  from ~25% to ~100%, which suggests the claimed task only gets to run
  when some unrelated event runs balance on its CPU.

This is independent of the caps-clear patch (it reproduces with and
without it). The root cause is likely in the sub-sched attach->dispatch
path itself, so any sub-scheduler deployment would be affected.

Questions:

1. Is this a known gap in the sub-sched dispatch model - claimed tasks
   on idle CPUs waiting for a delegation that only happens when the
   parent's dispatch runs for that CPU?
2. Is the intended pattern for parents to kick CPUs on the child's
   behalf (scx_bpf_kick_cid() after grant/claim), or should the core
   route the balance to the child?

The repro is deterministic (8/8 rounds with the starvation in order
(a) under root-side load).

Thanks.
--
Tao

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [REPORT] sched_ext: sub-scheduler tasks are severely under-scheduled vs root-owned tasks
  2026-10-08  8:36 [REPORT] sched_ext: sub-scheduler tasks are severely under-scheduled vs root-owned tasks Tao Cui
@ 2026-10-08  9:10 ` Tejun Heo
  2026-10-09  1:51   ` Tao Cui
  0 siblings, 1 reply; 4+ messages in thread
From: Tejun Heo @ 2026-10-08  9:10 UTC (permalink / raw)
  To: Tao Cui; +Cc: David Vernet, sched-ext, Andrea Righi, linux-kernel, Changwoo Min

Hello, Tao.

On Thu, Oct 08, 2026 at 04:36:49PM +0800, Tao Cui wrote:
> A pair of minimal cid-form test
> schedulers: a parent that grants SCX_CAP_PERF | SCX_CAP_ENQ_IMMED over
> a cgroup to a child, and the child writing cpuperf target 1 from
> ops.tick().

Can you explain exactly what you tested and how, including the scheduler
source and complete reproduction steps? This isn't a useful bug report
without that information.

Thanks.

-- 
tejun

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [REPORT] sched_ext: sub-scheduler tasks are severely under-scheduled vs root-owned tasks
  2026-10-08  9:10 ` Tejun Heo
@ 2026-10-09  1:51   ` Tao Cui
  2026-10-09 23:42     ` Tejun Heo
  0 siblings, 1 reply; 4+ messages in thread
From: Tao Cui @ 2026-10-09  1:51 UTC (permalink / raw)
  To: Tejun Heo
  Cc: cui.tao, David Vernet, sched-ext, Andrea Righi, linux-kernel,
	Changwoo Min

[-- Attachment #1: Type: text/plain, Size: 5628 bytes --]

Hello Tejun,

在 2026/10/8 17:10, Tejun Heo 写道:
> Hello, Tao.
> 
> On Thu, Oct 08, 2026 at 04:36:49PM +0800, Tao Cui wrote:
>> A pair of minimal cid-form test
>> schedulers: a parent that grants SCX_CAP_PERF | SCX_CAP_ENQ_IMMED over
>> a cgroup to a child, and the child writing cpuperf target 1 from
>> ops.tick().
> 
> Can you explain exactly what you tested and how, including the scheduler
> source and complete reproduction steps? This isn't a useful bug report
> without that information.
> 

Sure thing, and thanks for pointing out what was missing. Attached
are the complete scheduler sources together with the exact setup and
reproduction steps.

The scheduler is a single binary (scx_k2) with parent/root and
child/sub modes. The parent grants SCX_CAP_PERF | SCX_CAP_ENQ_IMMED
to the child over a cgroup, while the child only writes cpuperf
target 1 from ops.tick(). Both schedulers print their counters every
2s.

Environment:

- x86_64 KVM guest, QEMU "-enable-kvm -cpu host -M pc -smp 4 -m 4G",
  CONFIG_HZ_1000, default NO_HZ idle, cgroup2, everything as root.
- Kernel: your for-7.4 tree at f55a5ae74ca6 (Merge branch
  'for-7.3-fixes' into for-7.4), plus cda36da11910 (sched_ext: Clear
  a sub-scheduler's caps before ops.sub_detach()).

Reproduce (one round):

  1. Build scx_k2 and scx_perfprobe (recipe below).

  2. Mount cgroup2 if not already mounted:
     mount -t cgroup2 none /sys/fs/cgroup

  3. Enable the CPU controller:
     echo +cpu > /sys/fs/cgroup/cgroup.subtree_control

  4. Create the test cgroup:
     mkdir -p /sys/fs/cgroup/ss/a

  5. Start the parent and let it come up:
     ./scx_k2 -g /sys/fs/cgroup/ss/a > root.log 2>&1 &
     sleep 2

  6. Start the spinner pinned to CPU 0:
     taskset 0x1 sh -c 'while :; do :; done' &

  7. Move it into the cgroup before attaching the child:
     echo $! > /sys/fs/cgroup/ss/a/cgroup.procs

  8. Start a root-owned spinner pinned to CPU 2:
     taskset 0x4 sh -c 'while :; do :; done' &

  9. Attach the child and let it settle:
     ./scx_k2 -s /sys/fs/cgroup/ss/a > sub.log 2>&1 &
     sleep 2

  10. Wait 6s.

  11. Record the spinner's utime+stime from /proc/<pid>/stat before
      and after the window, and inspect the counter lines in
      root.log and sub.log.

The attached repro-subsched-starve.sh automates the setup and
reports the result:

  sh repro-subsched-starve.sh before 1

Expected: the child owns the spinner and it continues running at
full duty cycle, accumulating approximately 600 CPU ticks over 6s
(utime+stime; the guest's USER_HZ is 100). While the task runs,
setcnt increases at roughly 100/s in our healthy rounds.

Actual: the child claims the task, and its ops.tick() initially
services it, but setcnt then stops increasing and the task receives
little or no CPU time - 0-91 ticks per 6s window across our runs,
most rounds 0. In affected rounds, subdisp0 stays at 0 throughout:
the child's ops.dispatch() is never invoked for its CID.

Counters (both logs):

  setcnt    - child ops.tick() invocations
  rootticks - parent ops.tick() on CPU 0
  disp0     - parent's ops.dispatch() invocations for CID 0
  subdisp0  - child's ops.dispatch() invocations for CID 0
  dsqdepth  - parent SHARED DSQ depth, sampled on tick

Attachments:

  scx_k2.bpf.c              parent and child schedulers in one BPF
                            object
  scx_k2.c                  userspace loader and counter printer
  scx_perfprobe.bpf.c       sched_tick kprobe sampling
                            rq->scx.cpuperf_target per CPU
  scx_perfprobe.c           sampler printer
  repro-subsched-starve.sh  one-round runner, accepting
                            [before|after] [hog]

Build (from tools/sched_ext, with the four sources dropped in):

  make $(pwd)/build/include/scx_k2.bpf.skel.h \
       $(pwd)/build/include/scx_perfprobe.bpf.skel.h

  gcc -g -O2 -pthread -Ibuild/include -I../../include/generated \
      -I../../tools/lib -I../../tools/include -I../../tools/include/uapi \
      -Iinclude -o scx_k2 scx_k2.c build/obj/libbpf/libbpf.a \
      -lelf -lz -lzstd -llzma -lpthread

  gcc -g -O2 -pthread -Ibuild/include -I../../include/generated \
      -I../../tools/lib -I../../tools/include -I../../tools/include/uapi \
      -Iinclude -o scx_perfprobe scx_perfprobe.c build/obj/libbpf/libbpf.a \
      -lelf -lz -lzstd -llzma -lpthread

Follow-up testing:

We repeated the tests across 11 boots (89 rounds, covering both
attach orders, with and without the root hog). Every round on 10
boots starved (74 rounds total); on the remaining boot, all 15
rounds ran at full duty cycle with normal child dispatch. The
healthy boot used the same binary build as one of the starving
boots, suggesting that the behavior is boot-dependent rather than
determined by the attach order or binary build. We haven't
identified what accounts for the difference yet. Note that this also
differs from my earlier mail, where order (b) appeared healthy; on
this tree, starvation occurs with both attach orders.

In an A/B test, having the parent kick the task's CPU from
ops.enqueue() with scx_bpf_kick_cid() and hand its dispatch turn to
the child with scx_bpf_sub_dispatch() prevents the starvation.

One possibility is that, on a NO_HZ idle CPU, no balance is
triggered for the child's DSQ, so the child's ops.dispatch() is
never called. This is still a hypothesis; the A/B result suggests
the dispatch/kick path is worth investigating.

I'll keep digging into the per-boot difference. If it doesn't
reproduce on your side, please let me know what differs (kernel
config, topology, or anything else).

Thanks,
Tao

> Thanks.
> 

[-- Attachment #2: repro-subsched-starve.sh --]
[-- Type: application/x-shellscript, Size: 2131 bytes --]

[-- Attachment #3: scx_k2.bpf.c --]
[-- Type: text/x-csrc, Size: 6572 bytes --]

// SPDX-License-Identifier: GPL-2.0
/*
 * scx_k2 - minimal cid-form scheduler pair for sub-scheduler attach
 * testing (single binary, two modes).
 *
 * Root mode (default): private FIFO DSQ; when a sub-scheduler attaches,
 * grants grant_caps (default SCX_CAP_ENQ_IMMED|SCX_CAP_PERF, full cid
 * mask) over the cgroup given with -g. Never writes cpuperf itself.
 * -R <ticks> issues scx_bpf_sub_revoke() of SCX_CAP_PERF on tick N
 * (0 = off).
 *
 * Sub mode (-s): ops.tick() writes perf_target via
 * scx_bpf_cidperf_set(this_cid).
 *
 * cid-form notes: must use SCX_OPS_CID_DEFINE/OPEN; every program must
 * touch the arena (k2_arena_touch) or bpf_prog_arena() comes back NULL
 * and programs get rejected; sub_cgroup_id must be written to both
 * rodata and struct_ops; the minimal cap set is ENQ_IMMED; a private
 * DSQ is required.
 */
#include <scx/common.bpf.h>

char _license[] SEC("license") = "GPL";

const volatile u64 grant_cgid;	/* root: cgid to grant/revoke caps over */
const volatile u64 grant_caps;	/* root: cap bits to grant */
const volatile u64 revoke_at;	/* root: revoke PERF on tick N (0 = off) */
const volatile u32 restore_on_detach; /* root: parent writes neutral values on sub_detach */
const volatile u64 sub_cgroup_id; /* sub mode: dual-written with ops->sub_cgroup_id */
const volatile u32 is_sub;	/* 1: sub mode */
const volatile u32 perf_target;	/* sub: target written every tick */

UEI_DEFINE(uei);

#define SHARED_DSQ 0

enum {
	K2_GRANT_RET = 0,	/* sub_grant return value (errno stored as 8192-errno) */
	K2_REVOKE_DONE = 1,	/* number of revokes performed */
	K2_SETCNT = 2,		/* number of sub tick writes */
	K2_SET_ERR = 3,		/* last cidperf_set return value (errno as 8192-errno) */
	K2_ROOTTICKS = 4,	/* root ticks on cpu0 (ownership discriminator) */
	K2_DSQDEPTH = 7,	/* SHARED_DSQ depth sampled in root tick (starvation probe) */
	K2_DISP0 = 8,		/* ROOT's cid0 dispatch invocations */
	K2_DISP0_MOVE = 9,	/* ROOT's cid0 successful move_to_local */
	K2_SUBDISP0 = 10,	/* SUB's cid0 dispatch invocations (core routing signal) */
	K2_SUBDISP0_MOVE = 11,	/* SUB's cid0 successful move_to_local */
	K2_CNT_NR = 12,
};

struct {
	__uint(type, BPF_MAP_TYPE_ARRAY);
	__uint(max_entries, K2_CNT_NR);
	__type(key, u32);
	__type(value, u64);
} counters SEC(".maps");

/* cid-form requires an arena map (verified by the kernel) */
struct {
	__uint(type, BPF_MAP_TYPE_ARENA);
	__uint(map_flags, BPF_F_MMAPABLE);
	__uint(max_entries, 1 << 16);	/* pages */
} arena SEC(".maps");

struct scx_cmask __arena cap_mask;
u64 __arena k2_ticks;

static void bump(u32 idx)
{
	u64 *cnt = bpf_map_lookup_elem(&counters, &idx);

	if (cnt)
		(*cnt)++;
}

static void setcnt(u32 idx, s64 v)
{
	u64 val = v < 0 ? 8192 + (-v) : v;

	bpf_map_update_elem(&counters, &idx, &val, BPF_ANY);
}

/* every scx program must touch the arena, or bpf_prog_arena() == NULL -> EINVAL */
u64 __arena k2_arena_sink;
static void arena_touch(void)
{
	k2_arena_sink += cap_mask.alloc_words;
}

s32 BPF_STRUCT_OPS(k2_select_cid, struct task_struct *p,
		   s32 prev_cid, u64 wake_flags)
{
	arena_touch();
	return prev_cid;
}

void BPF_STRUCT_OPS(k2_enqueue, struct task_struct *p, u64 enq_flags)
{
	arena_touch();
	scx_bpf_dsq_insert(p, SHARED_DSQ, SCX_SLICE_DFL, enq_flags);
}

void BPF_STRUCT_OPS(k2_dispatch, s32 cid, struct task_struct *prev)
{
	arena_touch();

	if (is_sub) {
		/* the sub's own dispatch turn: consume our SHARED DSQ */
		if (cid == 0)
			bump(K2_SUBDISP0);
		if (scx_bpf_dsq_move_to_local(SHARED_DSQ, 0))
			bump(K2_SUBDISP0_MOVE);
		return;
	}

	/*
	 * Root's dispatch turn: consume our own SHARED DSQ. (A/B tested:
	 * handing the turn down to the sub first via
	 * scx_bpf_sub_dispatch() does not rescue the claimed task.)
	 */

	if (cid == 0) {
		bump(K2_DISP0);
		if (scx_bpf_dsq_move_to_local(SHARED_DSQ, 0))
			bump(K2_DISP0_MOVE);
	} else {
		scx_bpf_dsq_move_to_local(SHARED_DSQ, 0);
	}
}

/* root: grant caps when a sub attaches (full cid mask; test harness
 * deliberately over-grants) */
s32 BPF_STRUCT_OPS(k2_sub_attach, struct scx_sub_attach_args *args)
{
	u32 nr;
	s32 ret;

	arena_touch();
	if (!grant_cgid || !grant_caps)
		return 0;

	nr = scx_bpf_nr_cids();
	cmask_init(&cap_mask, 0, nr);
	cmask_set_range(&cap_mask, 0, nr);
	ret = scx_bpf_sub_grant(grant_cgid, grant_caps, &cap_mask, NULL);
	setcnt(K2_GRANT_RET, ret);
	return 0;
}

/* parent-side restore per the last-writer-wins ownership model: once
 * the sub is inert (caps cleared), the parent writes back neutral values */
void BPF_STRUCT_OPS(k2_sub_detach, struct scx_sub_detach_args *args)
{
	u32 i, nr;

	arena_touch();
	if (!restore_on_detach)
		return;

	nr = scx_bpf_nr_cids();
	bpf_for(i, 0, nr) {
		if (restore_on_detach == 2 || i == 0)
			scx_bpf_cidperf_set(i, SCX_CPUPERF_ONE);
	}
}

/* root: revoke PERF on tick revoke_at; sub: write target every tick */
void BPF_STRUCT_OPS(k2_tick, struct task_struct *p)
{
	s32 cid;

	if (is_sub) {
		cid = scx_bpf_this_cid();
		if (cid >= 0)
			setcnt(K2_SET_ERR,
			       scx_bpf_cidperf_set(cid, perf_target));
		bump(K2_SETCNT);
		return;
	}

	/* count only cpu0 root ticks: discriminates who owns the cpu0
	 * spinner (root ticks mean the spinner runs as a root task) */
	if (bpf_get_smp_processor_id() == 0)
		bump(K2_ROOTTICKS);
	setcnt(K2_DSQDEPTH, scx_bpf_dsq_nr_queued(SHARED_DSQ));
	arena_touch();
	k2_ticks++;
	if (revoke_at && k2_ticks == revoke_at) {
		u32 nr = scx_bpf_nr_cids();

		cmask_init(&cap_mask, 0, nr);
		cmask_set_range(&cap_mask, 0, nr);
		scx_bpf_sub_revoke(grant_cgid, SCX_CAP_PERF, &cap_mask);
		bump(K2_REVOKE_DONE);
	}
}

void BPF_STRUCT_OPS(k2_cid_online, s32 cpu)
{
	arena_touch();
}

void BPF_STRUCT_OPS(k2_cid_offline, s32 cpu)
{
	arena_touch();
}

s32 BPF_STRUCT_OPS(k2_init_task, struct task_struct *p,
		   struct scx_init_task_args *args)
{
	arena_touch();
	return 0;
}

s32 BPF_STRUCT_OPS_SLEEPABLE(k2_init)
{
	asm volatile("" :: "r"(&arena));	/* make the verifier claim the arena map */
	return scx_bpf_create_dsq(SHARED_DSQ, -1);
}

void BPF_STRUCT_OPS(k2_exit, struct scx_exit_info *ei)
{
	UEI_RECORD(uei, ei);
}

SCX_OPS_CID_DEFINE(k2_ops,
		   .select_cid		= (void *)k2_select_cid,
		   .enqueue		= (void *)k2_enqueue,
		   .dispatch		= (void *)k2_dispatch,
		   .sub_attach		= (void *)k2_sub_attach,
		   .sub_detach		= (void *)k2_sub_detach,
		   .tick		= (void *)k2_tick,
		   .cid_online		= (void *)k2_cid_online,
		   .cid_offline		= (void *)k2_cid_offline,
		   .init_task		= (void *)k2_init_task,
		   .init		= (void *)k2_init,
		   .exit		= (void *)k2_exit,
		   .name		= "k2");

[-- Attachment #4: scx_k2.c --]
[-- Type: text/x-csrc, Size: 4006 bytes --]

// SPDX-License-Identifier: GPL-2.0
/* scx_k2 userspace. Usage:
 *   root mode: scx_k2 [-g <sub-cgroup-dir>] [-c <caps>] [-R <ticks>] [-p <perf>]
 *   sub  mode: scx_k2 -s <my-cgroup-dir> [-p <perf>]
 * Prints all counters every 2s; SIGTERM exits (triggers sub disable =>
 * sweep path).
 */
#include <stdio.h>
#include <unistd.h>
#include <signal.h>
#include <fcntl.h>
#include <errno.h>
#include <sys/stat.h>
#include <bpf/bpf.h>
#include <scx/common.h>

/* Mirror of kernel/sched/ext/types.h's arena cap mask: the arena global
 * cap_mask is mirrored to userspace via the skeleton (must be defined
 * before the skel include). The bits[] flexible array is omitted and a
 * __pad field pads the struct to the 16 bytes BTF reports (FAM alignment
 * padding) - userspace never accesses the object, and a FAM would
 * trigger -Wgnu-variable-sized-type-not-at-end. */
struct scx_cmask {
	u32 base;
	u32 nr_cids;
	u32 alloc_words;
	u32 __pad;
};

#include "scx_k2.bpf.skel.h"

static bool verbose;
static volatile int exit_req;

static int libbpf_print_fn(enum libbpf_print_level level, const char *format, va_list args)
{
	if (level == LIBBPF_DEBUG && !verbose)
		return 0;
	return vfprintf(stderr, format, args);
}

static void sig_handler(int sig)
{
	exit_req = 1;
}

static const char *cnt_name(u32 i)
{
	static const char *names[] = { "grant_ret", "revoke_done", "setcnt", "set_err", "rootticks", "detach", "detach_ok", "dsqdepth", "disp0", "moved0", "subdisp0", "submoved0" };

	return names[i];
}

static u64 read_cnt(struct scx_k2 *skel, u32 idx)
{
	u64 v = 0;

	bpf_map_lookup_elem(bpf_map__fd(skel->maps.counters), &idx, &v);
	return v;
}

int main(int argc, char **argv)
{
	struct scx_k2 *skel;
	struct bpf_link *link;
	__s32 opt;
	u64 sub_cgid = 0, grant_cgid = 0, grant_caps = 0, revoke_at = 0;
	u32 restore_on_detach = 0;
	u32 perf_target = 1;

	libbpf_set_print(libbpf_print_fn);
	signal(SIGINT, sig_handler);
	signal(SIGTERM, sig_handler);

	skel = SCX_OPS_CID_OPEN(k2_ops, scx_k2);
	while ((opt = getopt(argc, argv, "s:g:c:R:e:p:v")) != -1) {
		struct stat st;

		switch (opt) {
		case 's':	/* sub mode, arg = own cgroup */
			if (stat(optarg, &st)) { perror(optarg); return 1; }
			sub_cgid = st.st_ino;
			break;
		case 'g':	/* root mode, arg = sub's cgroup */
			if (stat(optarg, &st)) { perror(optarg); return 1; }
			grant_cgid = st.st_ino;
			break;
		case 'c':
			grant_caps = strtoull(optarg, NULL, 0);
			break;
		case 'R':
			revoke_at = strtoull(optarg, NULL, 0);
			break;
		case 'e':	/* on sub_detach parent restores neutral values (1=cpu0 only, 2=all cids) */
			restore_on_detach = (u32)strtoul(optarg, NULL, 0);
			break;
		case 'p':
			perf_target = (u32)strtoul(optarg, NULL, 0);
			break;
		case 'v':
			verbose = true;
			break;
		default:
			fprintf(stderr, "scx_k2: -s <cg>|-g <cg> [-c caps] [-R ticks] [-e 1|2] [-p perf]\n");
			return opt != 'h';
		}
	}

	if (sub_cgid) {
		skel->struct_ops.k2_ops->sub_cgroup_id = sub_cgid;
		skel->rodata->sub_cgroup_id = sub_cgid;
		skel->rodata->is_sub = 1;
	}
	skel->rodata->grant_cgid = grant_cgid;
	/* no SCX_CAP_* enum in the userspace headers (kernel internal.h):
	 * ENQ_IMMED=BIT(0), PERF=BIT(3) */
	skel->rodata->grant_caps = grant_caps ? grant_caps : 0x9;
	skel->rodata->revoke_at = revoke_at;
	skel->rodata->restore_on_detach = restore_on_detach;
	skel->rodata->perf_target = perf_target;

	SCX_OPS_LOAD(skel, k2_ops, scx_k2, uei);
	link = SCX_OPS_ATTACH(skel, k2_ops, scx_k2);
	fprintf(stderr, "[k2] up: %s grant_cgid=%llu caps=0x%llx revoke_at=%llu perf=%u\n",
		sub_cgid ? "sub" : "root",
		(unsigned long long)grant_cgid,
		(unsigned long long)skel->rodata->grant_caps,
		(unsigned long long)revoke_at, perf_target);

	while (!exit_req) {
		u32 i;

		sleep(2);
		for (i = 0; i < 12; i++)
			fprintf(stderr, "[k2] %s=%llu%s", cnt_name(i),
				(unsigned long long)read_cnt(skel, i),
				i == 3 ? "\n" : " ");
	}

	bpf_link__destroy(link);
	scx_k2__destroy(skel);
	fprintf(stderr, "[k2] exited cleanly\n");
	return 0;
}

[-- Attachment #5: scx_perfprobe.bpf.c --]
[-- Type: text/x-csrc, Size: 1058 bytes --]

// SPDX-License-Identifier: GPL-2.0
/*
 * scx_perfprobe - samples per-CPU rq->scx.cpuperf_target (the value
 * schedutil consumes). Kprobes sched_tick and overwrites this CPU's
 * bucket on every tick (~1000/s at HZ=1000).
 */
#include <scx/common.bpf.h>

char _license[] SEC("license") = "GPL";

#define K2_MAX_CPUS 128

struct {
	__uint(type, BPF_MAP_TYPE_ARRAY);
	__uint(max_entries, K2_MAX_CPUS);
	__type(key, u32);
	__type(value, u64);
} targets SEC(".maps");

/* per-CPU sched_tick hits: tells whether a CPU is really running ticks */
struct {
	__uint(type, BPF_MAP_TYPE_ARRAY);
	__uint(max_entries, K2_MAX_CPUS);
	__type(key, u32);
	__type(value, u64);
} tickcnt SEC(".maps");

SEC("kprobe/sched_tick")
int BPF_KPROBE(k2_sample_tick)
{
	struct rq *rq = bpf_this_cpu_ptr(&runqueues);
	/* map value is u64: read the target into a u64 so we don't pass
	 * a u32 address as an 8-byte value */
	u64 t = rq->scx.cpuperf_target;
	u32 cpu = bpf_get_smp_processor_id();

	if (cpu < K2_MAX_CPUS)
		bpf_map_update_elem(&targets, &cpu, &t, BPF_ANY);
	return 0;
}

[-- Attachment #6: scx_perfprobe.c --]
[-- Type: text/x-csrc, Size: 1647 bytes --]

// SPDX-License-Identifier: GPL-2.0
/* scx_perfprobe userspace - prints one line per second with every
 * CPU's rq->scx.cpuperf_target. */
#include <stdio.h>
#include <unistd.h>
#include <signal.h>
#include <errno.h>
#include <bpf/bpf.h>
#include <scx/common.h>
#include "scx_perfprobe.bpf.skel.h"

#define K2_MAX_CPUS 128

static volatile int exit_req;

static void sig_handler(int sig)
{
	exit_req = 1;
}

int main(int argc, char **argv)
{
	struct scx_perfprobe *skel;
	struct bpf_link *link;
	long online = sysconf(_SC_NPROCESSORS_ONLN);
	unsigned long sec = 0;

	signal(SIGINT, sig_handler);
	signal(SIGTERM, sig_handler);

	skel = scx_perfprobe__open();
	if (!skel) {
		fprintf(stderr, "perfprobe: open failed\n");
		return 1;
	}
	if (scx_perfprobe__load(skel)) {
		fprintf(stderr, "perfprobe: load failed\n");
		return 1;
	}
	link = bpf_program__attach(skel->progs.k2_sample_tick);
	if (libbpf_get_error(link)) {
		fprintf(stderr, "perfprobe: attach failed\n");
		return 1;
	}

	while (!exit_req) {
		char line[1024];
		int n = 0, cpu;

		n += snprintf(line + n, sizeof(line) - n, "[T+%lus]", ++sec);
		for (cpu = 0; cpu < online && cpu < K2_MAX_CPUS; cpu++) {
			u64 v = 0;
			u32 key = cpu;

			bpf_map_lookup_elem(bpf_map__fd(skel->maps.targets),
					    &key, &v);
			u64 tk = 0;
			bpf_map_lookup_elem(bpf_map__fd(skel->maps.tickcnt),
					    &key, &tk);
			n += snprintf(line + n, sizeof(line) - n,
				      " c%d=%llu/t%llu", cpu,
				      (unsigned long long)v,
				      (unsigned long long)tk);
		}
		printf("%s\n", line);
		fflush(stdout);
		sleep(1);
	}

	bpf_link__destroy(link);
	scx_perfprobe__destroy(skel);
	return 0;
}

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [REPORT] sched_ext: sub-scheduler tasks are severely under-scheduled vs root-owned tasks
  2026-10-09  1:51   ` Tao Cui
@ 2026-10-09 23:42     ` Tejun Heo
  0 siblings, 0 replies; 4+ messages in thread
From: Tejun Heo @ 2026-10-09 23:42 UTC (permalink / raw)
  To: Tao Cui; +Cc: David Vernet, sched-ext, Andrea Righi, linux-kernel, Changwoo Min

Hello, Tao.

The following is a Claude-generated analysis.

On Fri, Oct 09, 2026 at 09:51:35AM +0800, Tao Cui wrote:
> The scheduler is a single binary (scx_k2) with parent/root and
> child/sub modes. The parent grants SCX_CAP_PERF | SCX_CAP_ENQ_IMMED
> to the child over a cgroup, while the child only writes cpuperf
> target 1 from ops.tick(). Both schedulers print their counters every
> 2s.

Reproduced on for-7.4 with your sources in a 4 cpu VM. The starvation
comes from the two schedulers, not the kernel.

The child's ops.enqueue() inserts every task into its own user DSQ. The
only consumer of that DSQ is the child's ops.dispatch(), and a child's
dispatch runs only when its parent calls scx_bpf_sub_dispatch() from its
own ops.dispatch(). scx_k2's parent never does. See the comment above
enum scx_cap_flags in kernel/sched/ext/internal.h and how scx_qmap
dispatches its children. The SysRq-D dump during the stall shows the
spinner in the child's DSQ 0 with a full slice while cpu0 runs root
tasks.

> In an A/B test, having the parent kick the task's CPU from
> ops.enqueue() with scx_bpf_kick_cid() and hand its dispatch turn to
> the child with scx_bpf_sub_dispatch() prevents the starvation.

scx_bpf_sub_dispatch() is only callable from ops.dispatch(). With the
parent calling it there, the child's dispatch does run, but every
scx_bpf_dsq_move_to_local() is rejected because the grant lacks
SCX_CAP_ENQ. SCX_CAP_ENQ_IMMED only covers IMMED inserts. The task goes
to the reject DSQ and is reenqueued, SCX_EV_REENQ_REPEAT climbs on the
child, and it still starves. With the sub dispatch call and SCX_CAP_ENQ
granted as well, the spinner runs at full duty under the child.

> Actual: the child claims the task, and its ops.tick() initially
> services it, but setcnt then stops increasing and the task receives
> little or no CPU time - 0-91 ticks per 6s window across our runs,
> most rounds 0.

The initial burst is the bypass DSQ drained during attach. In the
move-after-attach order the task keeps running through dispatch's
keep-last path until something else wakes on cpu0 and pushes it through
the child's ops.enqueue(), which is why some windows look healthy. The
stall watchdog catches it past your 6s window:

  sched_ext: BPF sub-scheduler "k2" disabled (runnable task stall)
  sched_ext: k2: sh[1401] failed to run for 30.696s

> One possibility is that, on a NO_HZ idle CPU, no balance is
> triggered for the child's DSQ, so the child's ops.dispatch() is
> never called.

The parent's dispatch on cid 0 ran several times a second throughout
the stall. It only consumed the parent's own DSQ.

Thanks.

-- 
tejun

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-10-09 23:42 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-08  8:36 [REPORT] sched_ext: sub-scheduler tasks are severely under-scheduled vs root-owned tasks Tao Cui
2026-10-08  9:10 ` Tejun Heo
2026-10-09  1:51   ` Tao Cui
2026-10-09 23:42     ` Tejun Heo

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox