Hello Tejun, 在 2026/10/8 17:10, Tejun Heo 写道: > Hello, Tao. > > On Thu, Oct 08, 2026 at 04:36:49PM +0800, Tao Cui wrote: >> A pair of minimal cid-form test >> schedulers: a parent that grants SCX_CAP_PERF | SCX_CAP_ENQ_IMMED over >> a cgroup to a child, and the child writing cpuperf target 1 from >> ops.tick(). > > Can you explain exactly what you tested and how, including the scheduler > source and complete reproduction steps? This isn't a useful bug report > without that information. > Sure thing, and thanks for pointing out what was missing. Attached are the complete scheduler sources together with the exact setup and reproduction steps. The scheduler is a single binary (scx_k2) with parent/root and child/sub modes. The parent grants SCX_CAP_PERF | SCX_CAP_ENQ_IMMED to the child over a cgroup, while the child only writes cpuperf target 1 from ops.tick(). Both schedulers print their counters every 2s. Environment: - x86_64 KVM guest, QEMU "-enable-kvm -cpu host -M pc -smp 4 -m 4G", CONFIG_HZ_1000, default NO_HZ idle, cgroup2, everything as root. - Kernel: your for-7.4 tree at f55a5ae74ca6 (Merge branch 'for-7.3-fixes' into for-7.4), plus cda36da11910 (sched_ext: Clear a sub-scheduler's caps before ops.sub_detach()). Reproduce (one round): 1. Build scx_k2 and scx_perfprobe (recipe below). 2. Mount cgroup2 if not already mounted: mount -t cgroup2 none /sys/fs/cgroup 3. Enable the CPU controller: echo +cpu > /sys/fs/cgroup/cgroup.subtree_control 4. Create the test cgroup: mkdir -p /sys/fs/cgroup/ss/a 5. Start the parent and let it come up: ./scx_k2 -g /sys/fs/cgroup/ss/a > root.log 2>&1 & sleep 2 6. Start the spinner pinned to CPU 0: taskset 0x1 sh -c 'while :; do :; done' & 7. Move it into the cgroup before attaching the child: echo $! > /sys/fs/cgroup/ss/a/cgroup.procs 8. Start a root-owned spinner pinned to CPU 2: taskset 0x4 sh -c 'while :; do :; done' & 9. Attach the child and let it settle: ./scx_k2 -s /sys/fs/cgroup/ss/a > sub.log 2>&1 & sleep 2 10. Wait 6s. 11. Record the spinner's utime+stime from /proc//stat before and after the window, and inspect the counter lines in root.log and sub.log. The attached repro-subsched-starve.sh automates the setup and reports the result: sh repro-subsched-starve.sh before 1 Expected: the child owns the spinner and it continues running at full duty cycle, accumulating approximately 600 CPU ticks over 6s (utime+stime; the guest's USER_HZ is 100). While the task runs, setcnt increases at roughly 100/s in our healthy rounds. Actual: the child claims the task, and its ops.tick() initially services it, but setcnt then stops increasing and the task receives little or no CPU time - 0-91 ticks per 6s window across our runs, most rounds 0. In affected rounds, subdisp0 stays at 0 throughout: the child's ops.dispatch() is never invoked for its CID. Counters (both logs): setcnt - child ops.tick() invocations rootticks - parent ops.tick() on CPU 0 disp0 - parent's ops.dispatch() invocations for CID 0 subdisp0 - child's ops.dispatch() invocations for CID 0 dsqdepth - parent SHARED DSQ depth, sampled on tick Attachments: scx_k2.bpf.c parent and child schedulers in one BPF object scx_k2.c userspace loader and counter printer scx_perfprobe.bpf.c sched_tick kprobe sampling rq->scx.cpuperf_target per CPU scx_perfprobe.c sampler printer repro-subsched-starve.sh one-round runner, accepting [before|after] [hog] Build (from tools/sched_ext, with the four sources dropped in): make $(pwd)/build/include/scx_k2.bpf.skel.h \ $(pwd)/build/include/scx_perfprobe.bpf.skel.h gcc -g -O2 -pthread -Ibuild/include -I../../include/generated \ -I../../tools/lib -I../../tools/include -I../../tools/include/uapi \ -Iinclude -o scx_k2 scx_k2.c build/obj/libbpf/libbpf.a \ -lelf -lz -lzstd -llzma -lpthread gcc -g -O2 -pthread -Ibuild/include -I../../include/generated \ -I../../tools/lib -I../../tools/include -I../../tools/include/uapi \ -Iinclude -o scx_perfprobe scx_perfprobe.c build/obj/libbpf/libbpf.a \ -lelf -lz -lzstd -llzma -lpthread Follow-up testing: We repeated the tests across 11 boots (89 rounds, covering both attach orders, with and without the root hog). Every round on 10 boots starved (74 rounds total); on the remaining boot, all 15 rounds ran at full duty cycle with normal child dispatch. The healthy boot used the same binary build as one of the starving boots, suggesting that the behavior is boot-dependent rather than determined by the attach order or binary build. We haven't identified what accounts for the difference yet. Note that this also differs from my earlier mail, where order (b) appeared healthy; on this tree, starvation occurs with both attach orders. In an A/B test, having the parent kick the task's CPU from ops.enqueue() with scx_bpf_kick_cid() and hand its dispatch turn to the child with scx_bpf_sub_dispatch() prevents the starvation. One possibility is that, on a NO_HZ idle CPU, no balance is triggered for the child's DSQ, so the child's ops.dispatch() is never called. This is still a hypothesis; the A/B result suggests the dispatch/kick path is worth investigating. I'll keep digging into the per-boot difference. If it doesn't reproduce on your side, please let me know what differs (kernel config, topology, or anything else). Thanks, Tao > Thanks. >