From: Will Deacon <will@kernel.org>
To: Jinjie Ruan <ruanjinjie@huawei.com>
Cc: linux-arm-kernel@lists.infradead.org,
linux-kernel@vger.kernel.org, Thomas Gleixner <tglx@kernel.org>,
Catalin Marinas <catalin.marinas@arm.com>,
Borislav Petkov <bp@alien8.de>,
Lorenzo Pieralisi <lpieralisi@kernel.org>,
Mark Rutland <mark.rutland@arm.com>,
David Woodhouse <dwmw@amazon.co.uk>,
Peter Zijlstra <peterz@infradead.org>,
Marc Zyngier <maz@kernel.org>
Subject: Re: [PATCH 09/19] arm64: smp: Defer RCU registration during secondary CPU bringup
Date: Wed, 9 Sep 2026 13:36:55 +0100 [thread overview]
Message-ID: <aqFS58mJ4GJNrbZ2@willie-the-truck> (raw)
In-Reply-To: <f01fc59e-5254-484d-adf9-8dd1c2e58d91@huawei.com>
On Tue, Sep 08, 2026 at 07:25:54PM +0800, Jinjie Ruan wrote:
> 在 2026/9/8 18:19, Will Deacon 写道:
> > On Tue, Sep 08, 2026 at 04:55:36PM +0800, Jinjie Ruan wrote:
> >> I think we need to handle the printk problem before this patch as we
> >> discussed earlier.
> >>
> >> Otherwise defer the rcutree_report_cpu_starting() will trigger a
> >> false-positive lockdep"suspicious RCU usage" splat during early lock
> >> acquisitions as commit ce3d31ad3cac ("arm64/smp: Move
> >> rcu_cpu_starting() earlier") pointed out.
> >
> > Sorry, I meant to mention this in the cover letter but forgot about it.
> > I'm not sure that ce3d31ad3cac ("arm64/smp: Move rcu_cpu_starting()
> > earlier") is still relevant with the latest printk/console/lockdep code.
> > I tried quite hard to trigger lockdep splats manually, but the only way
> > I could do it was by using the "%pS" specifier to print the name of a
> > symbol in a module, which would cause an RCU walk of the module symbols
> > in the kallsyms code! Manually calling WARN() or even rcu_read_lock() /
> > spin_lock() did _not_ trigger a splat.
>
> Add "dyndbg="+p"" in cmdline, CONFIG_DEBUG_LOCK_ALLOC=y,
> CONFIG_PROVE_RCU_LIST=y, we can reproduce the warning as below:
>
> I believe there is also a problem in the RISC-V code itself here as
> store_cpu_topology() is common for RISC-V.
>
> [ 0.335162] smp: Bringing up secondary CPUs ...
> [ 0.345495]
> [ 0.345513] =============================
> [ 0.345523] WARNING: suspicious RCU usage
> [ 0.345621] 7.3.0-rc2-00010-g2311ba2cd56f #500 Tainted: G W
> [ 0.345637] -----------------------------
> [ 0.345646] kernel/locking/lockdep.c:3845 RCU-list traversed in
> non-reader section!!
> [ 0.345659]
> [ 0.345659] other info that might help us debug this:
> [ 0.345659]
> [ 0.345680]
> [ 0.345680] RCU used illegally from offline CPU!
> [ 0.345680] rcu_scheduler_active = 1, debug_locks = 1
> [ 0.345725] locks held by swapper/1/0: 0, last CPU#1
> [ 0.345743]
> [ 0.345743] stack backtrace:
> [ 0.345834] CPU: 1 UID: 0 PID: 0 Comm: swapper/1 Tainted: G W
> 7.3.0-rc2-00010-g2311ba2cd56f #500 PREEMPT(full)
> [ 0.345885] Tainted: [W]=WARN
> [ 0.346077] Call trace:
> [ 0.346102] show_stack+0x20/0x38 (C)
> [ 0.346153] dump_stack_lvl+0xc4/0x150
> [ 0.346176] dump_stack+0x18/0x28
> [ 0.346194] lockdep_rcu_suspicious+0x170/0x238
> [ 0.346217] __lock_acquire+0xf08/0x1818
> [ 0.346237] lock_acquire+0x1e0/0x450
> [ 0.346256] _raw_spin_lock_irqsave+0x70/0xc0
> [ 0.346277] down_trylock+0x20/0x60
> [ 0.346293] __down_trylock_console_sem+0x4c/0x118
> [ 0.346316] vprintk_emit+0x2d8/0x3f8
> [ 0.346333] vprintk_default+0x40/0x58
> [ 0.346350] vprintk+0x3c/0x80
> [ 0.346366] _printk+0x64/0x98
> [ 0.346386] __dynamic_pr_debug+0x90/0xd8
> [ 0.346406] acpi_get_cache_info+0x140/0x1a0
> [ 0.346430] init_cache_level+0xec/0x110
> [ 0.346450] detect_cache_attributes+0x74/0x7c0
> [ 0.346473] update_siblings_masks+0x30/0x300
> [ 0.346495] store_cpu_topology+0x70/0xf0
> [ 0.346515] secondary_start_kernel+0xe0/0x178
> [ 0.346535] __secondary_switched+0xc0/0xc8
I was about to say "don't do this" but then I realised two things:
1. update_siblings_masks() can trigger lockdep splats outside of
pr_debug() if RCU isn't up and running, e.g.:
[ 0.524042] show_stack+0x18/0x24 (C)
[ 0.524519] __dump_stack+0x28/0x38
[ 0.524546] dump_stack_lvl+0x64/0x84
[ 0.524562] dump_stack+0x18/0x24
[ 0.524576] lockdep_rcu_suspicious+0x134/0x1cc
[ 0.524591] __lock_acquire+0xee8/0x2cb0
[ 0.524606] lock_acquire+0x11c/0x2fc
[ 0.524621] _raw_spin_lock_irqsave+0x64/0x84
[ 0.524641] of_find_property+0x2c/0x8c
[ 0.524659] detect_cache_attributes+0x1c0/0x6d0
[ 0.524676] update_siblings_masks+0x38/0x288
[ 0.524692] store_cpu_topology+0x4c/0x58
[ 0.524706] secondary_start_kernel+0xdc/0x1c8
[ 0.524722] __secondary_switched+0x120/0x124
2. This code is running _after_ cpuhp_ap_sync_alive().
So for the next version, I'll reintroduce the call to
rcutree_report_cpu_starting(), but move it immediately after the call to
cpuhp_ap_sync_alive(). I think that will solve these issues, without
causing issues with the concurrent part of early boot and also without
reintroducing the early call to rcutree_report_cpu_dead().
Cheers,
Will
next prev parent reply other threads:[~2026-09-09 12:37 UTC|newest]
Thread overview: 50+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-07 16:40 [PATCH 00/19] arm64: Implement parallel CPU onlining with PSCI v0.2+ Will Deacon
2026-09-07 16:40 ` [PATCH 01/19] cpu/hotplug: Clean up cmpxchg() logic in cpuhp_can_boot_ap() Will Deacon
2026-09-08 2:55 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 02/19] cpu/hotplug: Avoid trying to bring up CPUs that are already online Will Deacon
2026-09-08 3:13 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 03/19] cpu/hotplug: Avoid busy-polling on archs where cpu_relax() is a no-op Will Deacon
2026-09-08 4:00 ` Jinjie Ruan
2026-09-11 7:16 ` Jinjie Ruan
2026-09-11 12:57 ` Will Deacon
2026-09-07 16:40 ` [PATCH 04/19] cpu/hotplug: Propagate bring-up status to arch_cpuhp_cleanup_kick_cpu() Will Deacon
2026-09-08 4:05 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 05/19] arm64: smp: Tidy up smp_prepare_cpus() Will Deacon
2026-09-08 7:35 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 06/19] arm64: smp: Tidy up cpuinfo init and cpufeature updates Will Deacon
2026-09-08 7:53 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 07/19] arm64: smp: Defer update of secondary CPU capabilities Will Deacon
2026-09-08 8:20 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 08/19] arm64: smp: Don't bother printing the I-cache policy for each CPU Will Deacon
2026-09-08 8:33 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 09/19] arm64: smp: Defer RCU registration during secondary CPU bringup Will Deacon
2026-09-08 8:55 ` Jinjie Ruan
2026-09-08 10:19 ` Will Deacon
2026-09-08 11:25 ` Jinjie Ruan
2026-09-09 12:36 ` Will Deacon [this message]
2026-09-10 2:47 ` Jinjie Ruan
2026-09-11 12:52 ` Will Deacon
2026-09-07 16:40 ` [PATCH 10/19] arm64: smp: Use generic HOTPLUG_CORE_SYNC_FULL machinery for CPU onlining Will Deacon
2026-09-08 9:01 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 11/19] arm64: smp: Use generic HOTPLUG_SPLIT_STARTUP " Will Deacon
2026-09-08 9:10 ` Jinjie Ruan
2026-09-08 11:35 ` Jinjie Ruan
2026-09-11 12:55 ` Will Deacon
2026-09-07 16:40 ` [PATCH 12/19] arm64: cpu_ops: Make 'cpu_operations' pointer global instead of per-cpu Will Deacon
2026-09-08 11:32 ` Jinjie Ruan
2026-09-11 12:55 ` Will Deacon
2026-09-07 16:40 ` [PATCH 13/19] arm64: cpu_ops: Introduce get_secondary_cpu_ops() Will Deacon
2026-09-08 11:56 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 14/19] firmware/psci: Cache PSCI v0.2+ version number to avoid redundant SMCs Will Deacon
2026-09-08 11:57 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 15/19] firmware/psci: Extend ->cpu_on() callback to take an additional argument Will Deacon
2026-09-08 12:05 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 16/19] arm64: cpu_ops: Expose optional argument to target cpu in ->cpu_boot() Will Deacon
2026-09-08 12:12 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 17/19] arm64: smp: Pass secondary CPU boot parameters via firmware if possible Will Deacon
2026-09-08 12:16 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 18/19] arm64: smp: Use generic HOTPLUG_PARALLEL machinery for CPU onlining Will Deacon
2026-09-08 13:05 ` Jinjie Ruan
2026-09-11 12:56 ` Will Deacon
2026-09-07 16:40 ` [PATCH 19/19] arm64: smp: Harden parallel CPU bringup against broken PSCI firmware Will Deacon
2026-09-08 13:36 ` Will Deacon
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqFS58mJ4GJNrbZ2@willie-the-truck \
--to=will@kernel.org \
--cc=bp@alien8.de \
--cc=catalin.marinas@arm.com \
--cc=dwmw@amazon.co.uk \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=lpieralisi@kernel.org \
--cc=mark.rutland@arm.com \
--cc=maz@kernel.org \
--cc=peterz@infradead.org \
--cc=ruanjinjie@huawei.com \
--cc=tglx@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox