From: Guopeng Zhang <guopeng.zhang@linux.dev>
To: Peter Zijlstra <peterz@infradead.org>,
Jake S <j@metarealtyinc.ca>, Waiman Long <longman@redhat.com>
Cc: Ingo Molnar <mingo@redhat.com>,
Juri Lelli <juri.lelli@redhat.com>,
Vincent Guittot <vincent.guittot@linaro.org>,
Dietmar Eggemann <dietmar.eggemann@arm.com>,
Steven Rostedt <rostedt@goodmis.org>,
Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
Valentin Schneider <vschneid@redhat.com>,
K Prateek Nayak <kprateek.nayak@amd.com>,
Tejun Heo <tj@kernel.org>,
linux-kernel@vger.kernel.org, cgroups@vger.kernel.org
Subject: Re: [BUG] sched/fair: divide error in __calc_prop_weight() from the enqueue path (flat-hierarchy series)
Date: Wed, 19 Aug 2026 17:49:56 +0800 [thread overview]
Message-ID: <465d2716-d023-4858-bec6-b688540cafeb@linux.dev> (raw)
In-Reply-To: <20260819073657.GE1247881@noisy.programming.kicks-ass.net>
在 2026/8/19 15:36, Peter Zijlstra 写道:
> On Tue, Aug 18, 2026 at 07:13:31PM -0400, Jake S wrote:
>> Hi,
>>
>> I hit a divide-by-zero panic in __calc_prop_weight(), reached from
>> enqueue_hierarchy() inside enqueue_task_fair(). This is the *enqueue*
>> path, not the task_tick_fair() variant reported in May and addressed by
>> the se->on_rq guard folded into 85570f10a4c6 -- enqueue_hierarchy() and
>> dequeue_hierarchy() carry no equivalent check.
>>
>> The code is from the tip sched/core flat-hierarchy rework; it is not in
>> Linus' tree. I am running it via a distro kernel (CachyOS) that carries
>> the series, on 7.2-rc7 and 7.2.0.
>>
>> I have separated what I verified from what I am guessing. The last link
>> in the causal chain is unexplained and I am asking about it rather than
>> asserting it.
>>
>> === The oops ===
>>
>> Oops: divide error: 0000 [#1] SMP NOPTI
>> CPU: 12 UID: 1000 PID: 312907 Comm: bash
>> Tainted: G U C OE 7.2.0-rc7-2-cachyos-rc #1 PREEMPT(full)
>> Hardware name: Dell Inc. XPS 16 DA16260/0RMV2Y, BIOS 1.5.1 04/01/2026
>> RIP: 0010:enqueue_task_fair.llvm.6536700009857788019+0x422/0x950
>> Code: 0f 84 74 01 00 00 83 bd 68 01 00 00 00 45 0f 4f f4 48 8b 4d 00
>> 4c 89 e8 48 09 c8 48 c1 e8 20 0f 85 53 fd ff ff 44 89 e8 31 d2
>> <f7> f1 41 89 c5 e9 4f fd ff ff 0f 0b e9 1d fe ff ff 4c 89 e6
>> RAX: 0000000000000000 RBX: 0000000000000001 RCX: 0000000000000000
>> RDX: 0000000000000000 RSI: fffff46fbf98e680 RDI: fffff46fbf98ffc0
>> RBP: fffff46fbf98ffc0 R08: ffff8ee25f9b2a80 R09: 0000000000000000
>> R10: 0000000000000000 R11: 0000000000000110 R12: 0000000000000001
>> R13: 0000000000000000 R14: 0000000000000001 R15: fffff46fbf9901c0
>> Call Trace:
>> <TASK>
>> enqueue_task+0x8e/0x250
>> wake_up_new_task+0x148/0x2e0
>> kernel_clone+0x1c6/0x390
>> __x64_sys_clone+0xcc/0x100
>> do_syscall_64+0x147/0x3c0
>> asm_fred_entrypoint_user+0x41/0x41
>> </TASK>
>>
>> Machine was idle, lid closed, 11.66 h into the boot. bash forked, the
>> new task was enqueued, div trapped.
>>
>> It is not survivable in practice. panic_on_oops was 0, so the kernel
>> took the first #DE, printed the oops and continued for 476 ms. It then
>> faulted at the same RIP with byte-identical registers and an identical
>> RSP (ffffd46fff53bbb0):
>>
>> Kernel panic - not syncing: Fatal exception
>> Shutting down cpus with NMI
>>
>> i.e. the oops-recovery path (kill task -> schedule()) re-entered the
>> same enqueue with the rq lock already held mid-enqueue.
>>
>> === Where it divides (confirmed) ===
>>
>> kernel/sched/fair.c, __calc_prop_weight(), inlined into
>> enqueue_hierarchy() -> enqueue_task_fair():
>>
>> weight *= se->load.weight;
>> if (parent_entity(se))
>> weight /= cfs_rq->load.weight; /* <-- #DE */
>>
>> RCX = cfs_rq->load.weight = 0. R13 = 0 means se->load.weight was 0 as
>> well, i.e. a group sched_entity carrying zero weight.
>
> Durr, 0 weight not good, in any scheme. Much of the code strives to
> never let it get below 2 or so.
>
>> Not a miscompile: this is clang 22.1.8 + ThinLTO, hence the .llvm.<hash>
>> suffix. The 32-bit "div %ecx" against 64-bit C operands is clang's
>> BypassSlowDivision -- the preceding "or %rcx,%rax; shr $32,%rax; jne"
>> is its guard. The 64-bit slow path is present in the same function.
>>
>> === How the weight can reach zero (mechanism, partly inferred) ===
>>
>> __calc_smp_shares() ends:
>>
>> return clamp_t(long, shares, MIN_SHARES, shares_max);
>>
>> clamp() yields hi when hi < lo, so shares_max == 0 silently defeats the
>> MIN_SHARES floor and returns 0 -- exactly the case the comment directly
>> above it says must yield MIN_SHARES instead of 0. Note __clamp_once()
>> already carries
>
> Moo..
>
>>
>> BUILD_BUG_ON_MSG(statically_true(ulo > uhi), ...)
>>
>> so lo > hi is considered a bug upstream; it just cannot fire on a
>> runtime-computed shares_max.
>>
>> shares_max arrives from calc_concur_shares() as nr * tg_shares, where
>> nr = min(tg_tasks(tg), tg_cpus(tg)). tg_cpus() returns
>> cpuset_num_cpus(cgrp) unfloored, while its sibling tg_tasks() already
>> floors at 1. That asymmetry is the hole.
>>
>> concur is the live mode here:
>>
>> $ cat /sys/kernel/debug/sched/cgroup_mode
>> up smp (concur) max tasks
>>
>> What I could NOT establish: that tg_cpus() actually returned 0, or what
>> would produce an empty effective cpuset. update_cpumasks_hier()
>> substitutes the parent's effective_cpus before storing; on this machine
>> no cgroup has an empty cpuset.cpus.effective and every
>> cpuset.cpus.partition reads "member". Twelve cgroups here have an empty
>> cpuset.cpus and all report effective = 0-15. I suspected a power daemon
>> that rewrites AllowedCPUs on the top-level systemd slices using an
>> empty-then-set idiom, but I could not make that yield an empty effective
>> mask, so I am not claiming it.
>>
>> The missing floor looks like a hole regardless of what trips it, and I
>> would rather ask than guess: is there a path where cpuset_num_cpus() can
>> legitimately return 0, or should tg_cpus() simply floor at 1 the way
>> tg_tasks() does?
>
> tg_cpus() should probably floor at 1, just to be both symmetric and
> avoid this. But I too was under the impression a cpuset would never have
> an empty set.
>
> However, now that I think about it, IIRC there is a cpu hotplug (and
> suspect I suppose) case where cpuset-v2 allows a cpuset to become empty
> like this, *however* it would then take the parent cgroup until it would
> find one that is non empty.
>
This reminded me of a CPU hotplug / suspend issue I ran into recently.
I had a case where the last online HK_TYPE_DOMAIN CPU could be offlined
during regular hotplug because the check in _cpu_down() did not exclude
the outgoing CPU. That left scheduler-domain rebuilds with an empty span
and eventually caused a crash:
https://lore.kernel.org/all/20260811121307.168471-1-guopeng.zhang@linux.dev/
While testing this with domain isolation, I also saw suspend/freeze reach
a state with no active HK_TYPE_DOMAIN CPUs. I later sent an RFC for
handling that case in the scheduler/cpuset paths:
https://lore.kernel.org/all/20260722115238.351821-1-guopeng.zhang@linux.dev/
The reproducer used:
isolcpus=domain,0,3-31
I have not checked whether this can make cpuset_num_cpus() return 0 in
the path Jake hit, so this may well be unrelated. I just thought it was
worth mentioning since you brought up the hotplug/suspend case.
Thanks,
Guopeng
> Now, cpuset_num_cpus() uses RCU, so perhaps there is a race somewhere.
> Waiman, you know this cpuset stuff beter than me, did I get it wrong?
>
> Anyway, if you send a patch adding the floow to tg_cpus(), I'll apply
> that just on symmetry grounds. If Waiman spots a fail with the cpuset
> bits we can fix that too.
>
> Thanks for the excellent report!
>
next prev parent reply other threads:[~2026-08-19 9:50 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-18 23:13 [BUG] sched/fair: divide error in __calc_prop_weight() from the enqueue path (flat-hierarchy series) Jake S
2026-08-19 5:49 ` K Prateek Nayak
2026-08-19 7:36 ` Peter Zijlstra
2026-08-19 9:49 ` Guopeng Zhang [this message]
2026-08-19 13:20 ` [PATCH] sched/fair: floor tg_cpus() at 1 Jake Steinman
2026-08-19 13:22 ` [BUG] sched/fair: divide error in __calc_prop_weight() from the enqueue path (flat-hierarchy series) Jake Steinman
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=465d2716-d023-4858-bec6-b688540cafeb@linux.dev \
--to=guopeng.zhang@linux.dev \
--cc=bsegall@google.com \
--cc=cgroups@vger.kernel.org \
--cc=dietmar.eggemann@arm.com \
--cc=j@metarealtyinc.ca \
--cc=juri.lelli@redhat.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-kernel@vger.kernel.org \
--cc=longman@redhat.com \
--cc=mgorman@suse.de \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=tj@kernel.org \
--cc=vincent.guittot@linaro.org \
--cc=vschneid@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.