All of lore.kernel.org
 help / color / mirror / Atom feed
From: Guopeng Zhang <guopeng.zhang@linux.dev>
To: Bradley Morgan <include@grrlz.net>
Cc: frederic@kernel.org, linux-kernel@vger.kernel.org,
	peterz@infradead.org, tglx@kernel.org, zhangguopeng@kylinos.cn
Subject: Re: [PATCH] cpu/hotplug: Prevent offlining the last domain housekeeping CPU
Date: Mon, 20 Jul 2026 16:29:14 +0800	[thread overview]
Message-ID: <5a38dec5-9120-497d-aeb6-aa20ad87b7d5@linux.dev> (raw)
In-Reply-To: <4DE64100-BF83-4591-B554-4E6F7982FDD9@grrlz.net>



在 2026/7/18 22:53, Bradley Morgan 写道:
> Hi Guopeng,
> 
> On Fri, 18 Jul 2026, Guopeng Zhang wrote:
>> _cpu_down() is supposed to leave at least one HK_TYPE_DOMAIN CPU online.
>> However, cpumask_any_and() includes the CPU being offlined. The target
>> therefore satisfies the check when it is the last online CPU in that
>> housekeeping mask.

Hi Bradley,

Thanks for the detailed review. I verified the points you raised.

> 
> The bug is real. Note the code before the revert got this right,
> 38685e2a0476
> ("cpu/hotplug: Don't offline the last non-isolated CPU") had
> 
> 	for_each_cpu_and(cpu, cpu_online_mask, housekeeping_cpumask(HK_TYPE_DOMAIN)) {
> 		if (cpu != work.cpu)
> 
> and the exclusion got lost when de715325cc47 moved the check into
> _cpu_down(). please say that in the changelog, it makes the Fixes: tag
> obvious.
> 

Agreed. I will update the changelog accordingly.

>> Offlining it leaves no online CPU in the HK_TYPE_DOMAIN mask. The next
>> scheduler-domain rebuild passes an empty span to build_sched_domains()
>> and triggers its WARN_ON().
> 
> [...]
> 
>> On the unpatched kernel, offlining the last HK_TYPE_DOMAIN CPU triggered
>> a kernel panic.
> 
> WARN_ON() or panic, pick one. if that was panic_on_warn, say so. if it
> kept going and crashed later, paste the splat. The changelog should say
> what actually happens on a default config.
> 

The panic was not caused by panic_on_warn. After the empty-span warning,
the scheduler-domain rebuild continued and hit a general protection
fault in build_perf_domains(), which was followed by a kernel panic.

The relevant splat is:
[   69.172549] Oops: general protection fault, probably for non-canonical address 0xfc894905e5c15114: 0000 [#1] SMP NOPTI
[   69.173143] CPU: 2 UID: 0 PID: 29 Comm: cpuhp/2 Tainted: G        W   E       7.2.0-rc3-next-20260716-sdaf+ #4 PREEMPT(lazy) 
[   69.173966] Tainted: [W]=WARN, [E]=UNSIGNED_MODULE
[   69.174306] Hardware name: Red Hat KVM, BIOS rel-1.16.3-0-ga6ed6b701f0a-prebuilt.qemu.org 04/01/2014
[   69.174944] RIP: 0010:build_perf_domains+0x40/0x230
[   69.175312] Code: fd 53 e8 83 e1 8e 00 48 c7 c2 80 7c 20 92 48 98 48 8b 04 c5 80 0d 57 90 4c 8b a4 10 48 10 00 00 8b 05 6c 9f 4e 02 85 c0 75 3a <49> 8b 84 24 c0 07 00 00 49 c7 84 24 c0 07 00 00 00 00 00 00 48 85
[   69.176881] RSP: 0018:ffffd1be401efc38 EFLAGS: 00010246
[   69.177496] RAX: 0000000000000000 RBX: 0000000000000000 RCX: 0000000000000000
[   69.178328] RDX: 0000000000000040 RSI: 0000000000000020 RDI: ffff88ed8813ab30
[   69.178821] RBP: ffff88ed8813ab30 R08: 0000000000000001 R09: ffffffff80000000
[   69.179226] R10: 0000000000000000 R11: 0000000000000000 R12: fc894905e5c14954
[   69.179622] R13: 0000000000000008 R14: 0000000000000000 R15: ffff88ed8813a690
[   69.180017] FS:  0000000000000000(0000) GS:ffff88f9bad3e000(0000) knlGS:0000000000000000
[   69.180465] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[   69.180786] CR2: 00007f6cfb06a010 CR3: 0000000115c49001 CR4: 0000000000772ef0
[   69.181182] PKRU: 55555554
[   69.181340] Call Trace:
[   69.181484]  <TASK>
[   69.181608]  partition_sched_domains_locked+0x3d9/0x710
[   69.182013]  ? partition_sched_domains_locked+0x16d/0x710
[   69.182418]  partition_sched_domains+0x30/0x40
[   69.182769]  rebuild_sched_domains_locked+0xe3/0x110
[   69.183151]  cpuset_handle_hotplug+0x578/0xbc0
[   69.183507]  sched_cpu_deactivate+0x1dc/0x300
[   69.183856]  ? cpuhp_thread_fun+0x5a/0x230
[   69.184187]  cpuhp_invoke_callback+0x1a4/0x780
[   69.184542]  ? cpuhp_thread_fun+0x5a/0x230
[   69.184872]  cpuhp_thread_fun+0x1b4/0x230
[   69.185202]  ? __pfx_smpboot_thread_fn+0x10/0x10
[   69.185562]  smpboot_thread_fn+0xf3/0x220
[   69.185889]  kthread+0xe6/0x120
[   69.186168]  ? __pfx_kthread+0x10/0x10
[   69.186482]  ret_from_fork+0x29f/0x330
[   69.186793]  ? __pfx_kthread+0x10/0x10
[   69.187102]  ret_from_fork_asm+0x1a/0x30
[   69.187424]  </TASK>

I will describe this sequence more accurately in the v2 changelog.

>> -	if (cpumask_any_and(cpu_online_mask,
>> -			    housekeeping_cpumask(HK_TYPE_DOMAIN)) >= nr_cpu_ids) {
>> +	if (cpumask_any_and_but(cpu_online_mask,
>> +				housekeeping_cpumask(HK_TYPE_DOMAIN),
>> +				cpu) >= nr_cpu_ids) {
> 
> The check itself is correct for the sysfs path. cpumask_any_and_but()
> returns >= nr_cpu_ids when nothing is left, and excluding a cpu that is
> not in the housekeeping mask is harmless. Placement is fine too,
> de715325cc47 put this under cpus_write_lock() specifically to serialize
> against cpuset partition updates, so it should stay where it is.
> 
> but this breaks suspend. _cpu_down() is also the frozen path,
> freeze_secondary_cpus() does
> 
> 	error = _cpu_down(cpu, 1, CPUHP_OFFLINE);
> 
> for every CPU except the boot CPU. With your own cmdline
> (HK_TYPE_DOMAIN = {1,2}, primary = CPU0) suspend now dies at the last
> housekeeping CPU:
> 
> 	Error taking CPU1 down: -16
> 	Non-boot CPUs are not disabled

Confirmed. I tested suspend with the same isolcpus configuration and
reproduced this failure.

> The old check passed there because it counted the outgoing CPU. I did not
> check whether the frozen path can hit the same empty span WARN today, but
> turning a suspend that used to work into -EBUSY is a regression either
> way. do e.g:
> 
> 	if (!tasks_frozen &&
> 	    cpumask_any_and_but(cpu_online_mask,
> 				housekeeping_cpumask(HK_TYPE_DOMAIN),
> 				cpu) >= nr_cpu_ids) {
> 
> and test suspend/resume with the same isolcpus cmdline for v2, not just
> the sysfs offline.

I will update the patch as suggested and retest both the sysfs CPU
offline path and the suspend/resume path with the same isolcpus command
line before sending v2.

> 
> The Fixes: tag LGTM, the revert is where the exclusion went missing. If
> de715325cc47 is in a released kernel by now, CC stable. iirc cpu hotplug
> is root only and this needs isolcpus on the cmdline, so put that in the
> impact line and let the stable folks judge.

I will add the root-only and isolcpus requirements to the impact
description. I will also check whether de715325cc47 is present in a
released kernel and Cc stable if appropriate.

Thanks again for catching the suspend regression.

Thanks,
Guopeng


      reply	other threads:[~2026-07-20  8:29 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-18  9:52 [PATCH] cpu/hotplug: Prevent offlining the last domain housekeeping CPU Guopeng Zhang
2026-07-18 14:53 ` Bradley Morgan
2026-07-20  8:29   ` Guopeng Zhang [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=5a38dec5-9120-497d-aeb6-aa20ad87b7d5@linux.dev \
    --to=guopeng.zhang@linux.dev \
    --cc=frederic@kernel.org \
    --cc=include@grrlz.net \
    --cc=linux-kernel@vger.kernel.org \
    --cc=peterz@infradead.org \
    --cc=tglx@kernel.org \
    --cc=zhangguopeng@kylinos.cn \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.