The Linux Kernel Mailing List
 help / color / mirror / Atom feed
From: Guopeng Zhang <guopeng.zhang@linux.dev>
To: Peter Zijlstra <peterz@infradead.org>,
	Jake S <j@metarealtyinc.ca>, Waiman Long <longman@redhat.com>
Cc: Ingo Molnar <mingo@redhat.com>,
	Juri Lelli <juri.lelli@redhat.com>,
	Vincent Guittot <vincent.guittot@linaro.org>,
	Dietmar Eggemann <dietmar.eggemann@arm.com>,
	Steven Rostedt <rostedt@goodmis.org>,
	Ben Segall <bsegall@google.com>, Mel Gorman <mgorman@suse.de>,
	Valentin Schneider <vschneid@redhat.com>,
	K Prateek Nayak <kprateek.nayak@amd.com>,
	Tejun Heo <tj@kernel.org>,
	linux-kernel@vger.kernel.org, cgroups@vger.kernel.org
Subject: Re: [BUG] sched/fair: divide error in __calc_prop_weight() from the enqueue path (flat-hierarchy series)
Date: Wed, 19 Aug 2026 17:49:56 +0800	[thread overview]
Message-ID: <465d2716-d023-4858-bec6-b688540cafeb@linux.dev> (raw)
In-Reply-To: <20260819073657.GE1247881@noisy.programming.kicks-ass.net>



在 2026/8/19 15:36, Peter Zijlstra 写道:
> On Tue, Aug 18, 2026 at 07:13:31PM -0400, Jake S wrote:
>> Hi,
>>
>> I hit a divide-by-zero panic in __calc_prop_weight(), reached from
>> enqueue_hierarchy() inside enqueue_task_fair(). This is the *enqueue*
>> path, not the task_tick_fair() variant reported in May and addressed by
>> the se->on_rq guard folded into 85570f10a4c6 -- enqueue_hierarchy() and
>> dequeue_hierarchy() carry no equivalent check.
>>
>> The code is from the tip sched/core flat-hierarchy rework; it is not in
>> Linus' tree. I am running it via a distro kernel (CachyOS) that carries
>> the series, on 7.2-rc7 and 7.2.0.
>>
>> I have separated what I verified from what I am guessing. The last link
>> in the causal chain is unexplained and I am asking about it rather than
>> asserting it.
>>
>> === The oops ===
>>
>>   Oops: divide error: 0000 [#1] SMP NOPTI
>>   CPU: 12 UID: 1000 PID: 312907 Comm: bash
>>   Tainted: G     U   C OE       7.2.0-rc7-2-cachyos-rc #1 PREEMPT(full)
>>   Hardware name: Dell Inc. XPS 16 DA16260/0RMV2Y, BIOS 1.5.1 04/01/2026
>>   RIP: 0010:enqueue_task_fair.llvm.6536700009857788019+0x422/0x950
>>   Code: 0f 84 74 01 00 00 83 bd 68 01 00 00 00 45 0f 4f f4 48 8b 4d 00
>>         4c 89 e8 48 09 c8 48 c1 e8 20 0f 85 53 fd ff ff 44 89 e8 31 d2
>>         <f7> f1 41 89 c5 e9 4f fd ff ff 0f 0b e9 1d fe ff ff 4c 89 e6
>>   RAX: 0000000000000000 RBX: 0000000000000001 RCX: 0000000000000000
>>   RDX: 0000000000000000 RSI: fffff46fbf98e680 RDI: fffff46fbf98ffc0
>>   RBP: fffff46fbf98ffc0 R08: ffff8ee25f9b2a80 R09: 0000000000000000
>>   R10: 0000000000000000 R11: 0000000000000110 R12: 0000000000000001
>>   R13: 0000000000000000 R14: 0000000000000001 R15: fffff46fbf9901c0
>>   Call Trace:
>>    <TASK>
>>    enqueue_task+0x8e/0x250
>>    wake_up_new_task+0x148/0x2e0
>>    kernel_clone+0x1c6/0x390
>>    __x64_sys_clone+0xcc/0x100
>>    do_syscall_64+0x147/0x3c0
>>    asm_fred_entrypoint_user+0x41/0x41
>>    </TASK>
>>
>> Machine was idle, lid closed, 11.66 h into the boot. bash forked, the
>> new task was enqueued, div trapped.
>>
>> It is not survivable in practice. panic_on_oops was 0, so the kernel
>> took the first #DE, printed the oops and continued for 476 ms. It then
>> faulted at the same RIP with byte-identical registers and an identical
>> RSP (ffffd46fff53bbb0):
>>
>>   Kernel panic - not syncing: Fatal exception
>>   Shutting down cpus with NMI
>>
>> i.e. the oops-recovery path (kill task -> schedule()) re-entered the
>> same enqueue with the rq lock already held mid-enqueue.
>>
>> === Where it divides (confirmed) ===
>>
>> kernel/sched/fair.c, __calc_prop_weight(), inlined into
>> enqueue_hierarchy() -> enqueue_task_fair():
>>
>> 	weight *= se->load.weight;
>> 	if (parent_entity(se))
>> 		weight /= cfs_rq->load.weight;	/* <-- #DE */
>>
>> RCX = cfs_rq->load.weight = 0. R13 = 0 means se->load.weight was 0 as
>> well, i.e. a group sched_entity carrying zero weight.
> 
> Durr, 0 weight not good, in any scheme. Much of the code strives to
> never let it get below 2 or so.
> 
>> Not a miscompile: this is clang 22.1.8 + ThinLTO, hence the .llvm.<hash>
>> suffix. The 32-bit "div %ecx" against 64-bit C operands is clang's
>> BypassSlowDivision -- the preceding "or %rcx,%rax; shr $32,%rax; jne"
>> is its guard. The 64-bit slow path is present in the same function.
>>
>> === How the weight can reach zero (mechanism, partly inferred) ===
>>
>> __calc_smp_shares() ends:
>>
>> 	return clamp_t(long, shares, MIN_SHARES, shares_max);
>>
>> clamp() yields hi when hi < lo, so shares_max == 0 silently defeats the
>> MIN_SHARES floor and returns 0 -- exactly the case the comment directly
>> above it says must yield MIN_SHARES instead of 0. Note __clamp_once()
>> already carries
> 
> Moo..
> 
>>
>> 	BUILD_BUG_ON_MSG(statically_true(ulo > uhi), ...)
>>
>> so lo > hi is considered a bug upstream; it just cannot fire on a
>> runtime-computed shares_max.
>>
>> shares_max arrives from calc_concur_shares() as nr * tg_shares, where
>> nr = min(tg_tasks(tg), tg_cpus(tg)). tg_cpus() returns
>> cpuset_num_cpus(cgrp) unfloored, while its sibling tg_tasks() already
>> floors at 1. That asymmetry is the hole.
>>
>> concur is the live mode here:
>>
>>   $ cat /sys/kernel/debug/sched/cgroup_mode
>>   up smp (concur) max tasks
>>
>> What I could NOT establish: that tg_cpus() actually returned 0, or what
>> would produce an empty effective cpuset. update_cpumasks_hier()
>> substitutes the parent's effective_cpus before storing; on this machine
>> no cgroup has an empty cpuset.cpus.effective and every
>> cpuset.cpus.partition reads "member". Twelve cgroups here have an empty
>> cpuset.cpus and all report effective = 0-15. I suspected a power daemon
>> that rewrites AllowedCPUs on the top-level systemd slices using an
>> empty-then-set idiom, but I could not make that yield an empty effective
>> mask, so I am not claiming it.
>>
>> The missing floor looks like a hole regardless of what trips it, and I
>> would rather ask than guess: is there a path where cpuset_num_cpus() can
>> legitimately return 0, or should tg_cpus() simply floor at 1 the way
>> tg_tasks() does?
> 
> tg_cpus() should probably floor at 1, just to be both symmetric and
> avoid this. But I too was under the impression a cpuset would never have
> an empty set.
> 
> However, now that I think about it, IIRC there is a cpu hotplug (and
> suspect I suppose) case where cpuset-v2 allows a cpuset to become empty
> like this, *however* it would then take the parent cgroup until it would
> find one that is non empty.
> 

This reminded me of a CPU hotplug / suspend issue I ran into recently.

I had a case where the last online HK_TYPE_DOMAIN CPU could be offlined
during regular hotplug because the check in _cpu_down() did not exclude
the outgoing CPU. That left scheduler-domain rebuilds with an empty span
and eventually caused a crash:

https://lore.kernel.org/all/20260811121307.168471-1-guopeng.zhang@linux.dev/

While testing this with domain isolation, I also saw suspend/freeze reach
a state with no active HK_TYPE_DOMAIN CPUs. I later sent an RFC for
handling that case in the scheduler/cpuset paths:

https://lore.kernel.org/all/20260722115238.351821-1-guopeng.zhang@linux.dev/

The reproducer used:

    isolcpus=domain,0,3-31

I have not checked whether this can make cpuset_num_cpus() return 0 in
the path Jake hit, so this may well be unrelated. I just thought it was
worth mentioning since you brought up the hotplug/suspend case.

Thanks,
Guopeng

> Now, cpuset_num_cpus() uses RCU, so perhaps there is a race somewhere.
> Waiman, you know this cpuset stuff beter than me, did I get it wrong?
> 
> Anyway, if you send a patch adding the floow to tg_cpus(), I'll apply
> that just on symmetry grounds. If Waiman spots a fail with the cpuset
> bits we can fix that too.
> 
> Thanks for the excellent report!
> 


  reply	other threads:[~2026-08-19  9:50 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-18 23:13 [BUG] sched/fair: divide error in __calc_prop_weight() from the enqueue path (flat-hierarchy series) Jake S
2026-08-19  5:49 ` K Prateek Nayak
2026-08-19  7:36 ` Peter Zijlstra
2026-08-19  9:49   ` Guopeng Zhang [this message]
2026-08-19 13:20 ` [PATCH] sched/fair: floor tg_cpus() at 1 Jake Steinman
2026-08-19 13:22 ` [BUG] sched/fair: divide error in __calc_prop_weight() from the enqueue path (flat-hierarchy series) Jake Steinman
2026-08-19 18:47 ` Jake Steinman

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=465d2716-d023-4858-bec6-b688540cafeb@linux.dev \
    --to=guopeng.zhang@linux.dev \
    --cc=bsegall@google.com \
    --cc=cgroups@vger.kernel.org \
    --cc=dietmar.eggemann@arm.com \
    --cc=j@metarealtyinc.ca \
    --cc=juri.lelli@redhat.com \
    --cc=kprateek.nayak@amd.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=longman@redhat.com \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=rostedt@goodmis.org \
    --cc=tj@kernel.org \
    --cc=vincent.guittot@linaro.org \
    --cc=vschneid@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox