From: Jinjie Ruan <ruanjinjie@huawei.com>
To: Will Deacon <will@kernel.org>
Cc: <linux-arm-kernel@lists.infradead.org>,
<linux-kernel@vger.kernel.org>, Thomas Gleixner <tglx@kernel.org>,
Catalin Marinas <catalin.marinas@arm.com>,
Borislav Petkov <bp@alien8.de>,
Lorenzo Pieralisi <lpieralisi@kernel.org>,
Mark Rutland <mark.rutland@arm.com>,
David Woodhouse <dwmw@amazon.co.uk>,
Peter Zijlstra <peterz@infradead.org>,
Marc Zyngier <maz@kernel.org>
Subject: Re: [PATCH 18/19] arm64: smp: Use generic HOTPLUG_PARALLEL machinery for CPU onlining
Date: Wed, 16 Sep 2026 10:34:10 +0800 [thread overview]
Message-ID: <18e3c3d9-562d-42c6-80f6-7744adfad79d@huawei.com> (raw)
In-Reply-To: <aqP6kv9ZfNRWmngj@willie-the-truck>
在 2026/9/11 20:56, Will Deacon 写道:
> On Tue, Sep 08, 2026 at 09:05:18PM +0800, Jinjie Ruan wrote:
>> 在 2026/9/8 0:40, Will Deacon 写道:
>>> diff --git a/arch/arm64/kernel/head.S b/arch/arm64/kernel/head.S
>>> index 17868b497d7c..bec4bc1b12db 100644
>>> --- a/arch/arm64/kernel/head.S
>>> +++ b/arch/arm64/kernel/head.S
>>> @@ -393,7 +393,6 @@ SYM_FUNC_START_LOCAL(__secondary_switched)
>>> mov x0, x20
>>> bl finalise_el2
>>>
>>> - str_l xzr, __early_cpu_boot_status, x3
>>> adr_l x5, vectors
>>> msr vbar_el1, x5
>>> isb
>>> @@ -439,15 +438,15 @@ SYM_FUNC_END(set_cpu_boot_mode_flag)
>>> * with MMU turned off.
>>> *
>>> * update_early_cpu_boot_status tmp, status
>>> - * - Corrupts tmp1, tmp2
>>> - * - Writes 'status' to __early_cpu_boot_status and makes sure
>>> + * - Corrupts tmp1
>>
>> Corrupts tmp1, tmp2 ?
>
> Well spotted, thanks.
>
>>> @@ -125,45 +135,42 @@ int arch_cpuhp_kick_ap_alive(unsigned int cpu, struct task_struct *idle)
>>>
>>> void arch_cpuhp_cleanup_kick_cpu(unsigned int cpu, bool is_alive)
>>> {
>>> - long status;
>>> + union secondary_status status;
>>>
>>> if (is_alive)
>>> return;
>>>
>>> - secondary_data.task = NULL;
>>> - status = READ_ONCE(secondary_data.status);
>>> - if (status == CPU_MMU_OFF)
>>> - status = READ_ONCE(__early_cpu_boot_status);
>>> -
>>> /* A CPU has failed to boot. Try to figure out what happened. */
>>> - switch (status & CPU_BOOT_STATUS_MASK) {
>>> - default:
>>> - pr_err("CPU%u: failed in unknown state : 0x%lx\n",
>>> - cpu, status);
>>> - cpus_stuck_in_kernel++;
>>> - break;
>>> - case CPU_KILL_ME:
>>> - if (cpumask_test_cpu(cpu, &secondary_data.cpu_died_early_mask))
>>> - set_cpu_present(cpu, false);
>>> + if (smp_parallel_bringup)
>>> + pr_warn_once("Parallel CPU bringup failed; consider passing \"cpuhp.parallel=off\" for a more accurate diagnosis.\n");
>>
>> For some production systems, restarting to reproduce the issue may be
>> troublesome.
>
> These errors _really_ shouldn't happen with production systems. They are
> caused by critical, deterministic errors such as the secondary CPU not
> supporting the system page size. If you can't reboot in that situation,
> then you have no system!
That does seem to be the case — this kind of situation is usually more
common with firmware issues or during the testing phase before a chip is
commercially available, and these boot problems should already have been
resolved before production systems.
>
> The parallel bringup code will detect the issue, it just won't be able
> to tell you which CPU caused which issue.
Indeed.
>
>>> + if (cpumask_test_cpu(cpu, &secondary_data.cpu_died_early_mask)) {
>>> + set_cpu_present(cpu, false);
>>> if (!op_cpu_kill(cpu)) {
>>> pr_crit("CPU%u: died during early boot\n", cpu);
>>> - break;
>>> + return;
>>> }
>>> - pr_crit("CPU%u: may not have shut down cleanly\n", cpu);
>>> - fallthrough;
>>> - case CPU_STUCK_IN_KERNEL:
>>> - pr_crit("CPU%u: is stuck in kernel\n", cpu);
>>> - if (status & CPU_STUCK_REASON_52_BIT_VA)
>>> - pr_crit("CPU%u: does not support 52-bit VAs\n", cpu);
>>> - if (status & CPU_STUCK_REASON_NO_GRAN) {
>>> - pr_crit("CPU%u: does not support %luK granule\n",
>>> - cpu, PAGE_SIZE / SZ_1K);
>>> - }
>>> - cpus_stuck_in_kernel++;
>>> - break;
>>> - case CPU_PANIC_KERNEL:
>>> - panic("CPU%u detected unsupported configuration\n", cpu);
>>> }
>>> +
>>> + pr_crit_once("CPUs may be stuck in kernel\n");
>>
>> Need we print the "cpu"
>
> You should already get something like:
>
> CPUn: will not boot
"will not boot" is only for "cpu_died_early_mask" case, I meant the
"unknown state" case in the original default branch of the switch.
>
> so I don't think we need anything extra (as, as above, we don't know
> exactly which CPUs are stuck).
>
>>> static void init_gic_priority_masking(void)
>>> @@ -407,12 +414,8 @@ void __noreturn cpu_die_early(void)
>>>
>>> cpumask_set_cpu(cpu, &secondary_data.cpu_died_early_mask);
>>
>> It seems unsafe for multiple secondary CPUs to update
>> cpu_died_early_mask concurrently.
>
> Why? cpumask_set_cpu() is atomic and each CPU only sets their own bit.
Oh! Sorry, I misread it — the underlying implementation uses atomic LSE
or LDXR/STXR instructions.
>
> Will
next prev parent reply other threads:[~2026-09-16 2:34 UTC|newest]
Thread overview: 61+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-07 16:40 [PATCH 00/19] arm64: Implement parallel CPU onlining with PSCI v0.2+ Will Deacon
2026-09-07 16:40 ` [PATCH 01/19] cpu/hotplug: Clean up cmpxchg() logic in cpuhp_can_boot_ap() Will Deacon
2026-09-08 2:55 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 02/19] cpu/hotplug: Avoid trying to bring up CPUs that are already online Will Deacon
2026-09-08 3:13 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 03/19] cpu/hotplug: Avoid busy-polling on archs where cpu_relax() is a no-op Will Deacon
2026-09-08 4:00 ` Jinjie Ruan
2026-09-11 7:16 ` Jinjie Ruan
2026-09-11 12:57 ` Will Deacon
2026-09-16 1:15 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 04/19] cpu/hotplug: Propagate bring-up status to arch_cpuhp_cleanup_kick_cpu() Will Deacon
2026-09-08 4:05 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 05/19] arm64: smp: Tidy up smp_prepare_cpus() Will Deacon
2026-09-08 7:35 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 06/19] arm64: smp: Tidy up cpuinfo init and cpufeature updates Will Deacon
2026-09-08 7:53 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 07/19] arm64: smp: Defer update of secondary CPU capabilities Will Deacon
2026-09-08 8:20 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 08/19] arm64: smp: Don't bother printing the I-cache policy for each CPU Will Deacon
2026-09-08 8:33 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 09/19] arm64: smp: Defer RCU registration during secondary CPU bringup Will Deacon
2026-09-08 8:55 ` Jinjie Ruan
2026-09-08 10:19 ` Will Deacon
2026-09-08 11:25 ` Jinjie Ruan
2026-09-09 12:36 ` Will Deacon
2026-09-10 2:47 ` Jinjie Ruan
2026-09-11 12:52 ` Will Deacon
2026-09-16 2:40 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 10/19] arm64: smp: Use generic HOTPLUG_CORE_SYNC_FULL machinery for CPU onlining Will Deacon
2026-09-08 9:01 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 11/19] arm64: smp: Use generic HOTPLUG_SPLIT_STARTUP " Will Deacon
2026-09-08 9:10 ` Jinjie Ruan
2026-09-08 11:35 ` Jinjie Ruan
2026-09-11 12:55 ` Will Deacon
2026-09-07 16:40 ` [PATCH 12/19] arm64: cpu_ops: Make 'cpu_operations' pointer global instead of per-cpu Will Deacon
2026-09-08 11:32 ` Jinjie Ruan
2026-09-11 12:55 ` Will Deacon
2026-09-16 1:12 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 13/19] arm64: cpu_ops: Introduce get_secondary_cpu_ops() Will Deacon
2026-09-08 11:56 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 14/19] firmware/psci: Cache PSCI v0.2+ version number to avoid redundant SMCs Will Deacon
2026-09-08 11:57 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 15/19] firmware/psci: Extend ->cpu_on() callback to take an additional argument Will Deacon
2026-09-08 12:05 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 16/19] arm64: cpu_ops: Expose optional argument to target cpu in ->cpu_boot() Will Deacon
2026-09-08 12:12 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 17/19] arm64: smp: Pass secondary CPU boot parameters via firmware if possible Will Deacon
2026-09-08 12:16 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 18/19] arm64: smp: Use generic HOTPLUG_PARALLEL machinery for CPU onlining Will Deacon
2026-09-08 13:05 ` Jinjie Ruan
2026-09-11 12:56 ` Will Deacon
2026-09-16 2:34 ` Jinjie Ruan [this message]
2026-09-18 15:18 ` Will Deacon
2026-09-07 16:40 ` [PATCH 19/19] arm64: smp: Harden parallel CPU bringup against broken PSCI firmware Will Deacon
2026-09-08 13:36 ` Will Deacon
2026-09-29 8:56 ` [PATCH 00/19] arm64: Implement parallel CPU onlining with PSCI v0.2+ Pankaj Patil
2026-10-08 9:38 ` Will Deacon
2026-10-09 10:00 ` David Woodhouse
2026-10-09 11:08 ` Will Deacon
2026-10-09 14:45 ` David Woodhouse
2026-10-09 16:04 ` David Woodhouse
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=18e3c3d9-562d-42c6-80f6-7744adfad79d@huawei.com \
--to=ruanjinjie@huawei.com \
--cc=bp@alien8.de \
--cc=catalin.marinas@arm.com \
--cc=dwmw@amazon.co.uk \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=lpieralisi@kernel.org \
--cc=mark.rutland@arm.com \
--cc=maz@kernel.org \
--cc=peterz@infradead.org \
--cc=tglx@kernel.org \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox