From: Will Deacon <will@kernel.org>
To: Jinjie Ruan <ruanjinjie@huawei.com>
Cc: linux-arm-kernel@lists.infradead.org,
linux-kernel@vger.kernel.org, Thomas Gleixner <tglx@kernel.org>,
Catalin Marinas <catalin.marinas@arm.com>,
Borislav Petkov <bp@alien8.de>,
Lorenzo Pieralisi <lpieralisi@kernel.org>,
Mark Rutland <mark.rutland@arm.com>,
David Woodhouse <dwmw@amazon.co.uk>,
Peter Zijlstra <peterz@infradead.org>,
Marc Zyngier <maz@kernel.org>
Subject: Re: [PATCH 09/19] arm64: smp: Defer RCU registration during secondary CPU bringup
Date: Wed, 9 Sep 2026 13:36:55 +0100 [thread overview]
Message-ID: <aqFS58mJ4GJNrbZ2@willie-the-truck> (raw)
In-Reply-To: <f01fc59e-5254-484d-adf9-8dd1c2e58d91@huawei.com>
On Tue, Sep 08, 2026 at 07:25:54PM +0800, Jinjie Ruan wrote:
> 在 2026/9/8 18:19, Will Deacon 写道:
> > On Tue, Sep 08, 2026 at 04:55:36PM +0800, Jinjie Ruan wrote:
> >> I think we need to handle the printk problem before this patch as we
> >> discussed earlier.
> >>
> >> Otherwise defer the rcutree_report_cpu_starting() will trigger a
> >> false-positive lockdep"suspicious RCU usage" splat during early lock
> >> acquisitions as commit ce3d31ad3cac ("arm64/smp: Move
> >> rcu_cpu_starting() earlier") pointed out.
> >
> > Sorry, I meant to mention this in the cover letter but forgot about it.
> > I'm not sure that ce3d31ad3cac ("arm64/smp: Move rcu_cpu_starting()
> > earlier") is still relevant with the latest printk/console/lockdep code.
> > I tried quite hard to trigger lockdep splats manually, but the only way
> > I could do it was by using the "%pS" specifier to print the name of a
> > symbol in a module, which would cause an RCU walk of the module symbols
> > in the kallsyms code! Manually calling WARN() or even rcu_read_lock() /
> > spin_lock() did _not_ trigger a splat.
>
> Add "dyndbg="+p"" in cmdline, CONFIG_DEBUG_LOCK_ALLOC=y,
> CONFIG_PROVE_RCU_LIST=y, we can reproduce the warning as below:
>
> I believe there is also a problem in the RISC-V code itself here as
> store_cpu_topology() is common for RISC-V.
>
> [ 0.335162] smp: Bringing up secondary CPUs ...
> [ 0.345495]
> [ 0.345513] =============================
> [ 0.345523] WARNING: suspicious RCU usage
> [ 0.345621] 7.3.0-rc2-00010-g2311ba2cd56f #500 Tainted: G W
> [ 0.345637] -----------------------------
> [ 0.345646] kernel/locking/lockdep.c:3845 RCU-list traversed in
> non-reader section!!
> [ 0.345659]
> [ 0.345659] other info that might help us debug this:
> [ 0.345659]
> [ 0.345680]
> [ 0.345680] RCU used illegally from offline CPU!
> [ 0.345680] rcu_scheduler_active = 1, debug_locks = 1
> [ 0.345725] locks held by swapper/1/0: 0, last CPU#1
> [ 0.345743]
> [ 0.345743] stack backtrace:
> [ 0.345834] CPU: 1 UID: 0 PID: 0 Comm: swapper/1 Tainted: G W
> 7.3.0-rc2-00010-g2311ba2cd56f #500 PREEMPT(full)
> [ 0.345885] Tainted: [W]=WARN
> [ 0.346077] Call trace:
> [ 0.346102] show_stack+0x20/0x38 (C)
> [ 0.346153] dump_stack_lvl+0xc4/0x150
> [ 0.346176] dump_stack+0x18/0x28
> [ 0.346194] lockdep_rcu_suspicious+0x170/0x238
> [ 0.346217] __lock_acquire+0xf08/0x1818
> [ 0.346237] lock_acquire+0x1e0/0x450
> [ 0.346256] _raw_spin_lock_irqsave+0x70/0xc0
> [ 0.346277] down_trylock+0x20/0x60
> [ 0.346293] __down_trylock_console_sem+0x4c/0x118
> [ 0.346316] vprintk_emit+0x2d8/0x3f8
> [ 0.346333] vprintk_default+0x40/0x58
> [ 0.346350] vprintk+0x3c/0x80
> [ 0.346366] _printk+0x64/0x98
> [ 0.346386] __dynamic_pr_debug+0x90/0xd8
> [ 0.346406] acpi_get_cache_info+0x140/0x1a0
> [ 0.346430] init_cache_level+0xec/0x110
> [ 0.346450] detect_cache_attributes+0x74/0x7c0
> [ 0.346473] update_siblings_masks+0x30/0x300
> [ 0.346495] store_cpu_topology+0x70/0xf0
> [ 0.346515] secondary_start_kernel+0xe0/0x178
> [ 0.346535] __secondary_switched+0xc0/0xc8
I was about to say "don't do this" but then I realised two things:
1. update_siblings_masks() can trigger lockdep splats outside of
pr_debug() if RCU isn't up and running, e.g.:
[ 0.524042] show_stack+0x18/0x24 (C)
[ 0.524519] __dump_stack+0x28/0x38
[ 0.524546] dump_stack_lvl+0x64/0x84
[ 0.524562] dump_stack+0x18/0x24
[ 0.524576] lockdep_rcu_suspicious+0x134/0x1cc
[ 0.524591] __lock_acquire+0xee8/0x2cb0
[ 0.524606] lock_acquire+0x11c/0x2fc
[ 0.524621] _raw_spin_lock_irqsave+0x64/0x84
[ 0.524641] of_find_property+0x2c/0x8c
[ 0.524659] detect_cache_attributes+0x1c0/0x6d0
[ 0.524676] update_siblings_masks+0x38/0x288
[ 0.524692] store_cpu_topology+0x4c/0x58
[ 0.524706] secondary_start_kernel+0xdc/0x1c8
[ 0.524722] __secondary_switched+0x120/0x124
2. This code is running _after_ cpuhp_ap_sync_alive().
So for the next version, I'll reintroduce the call to
rcutree_report_cpu_starting(), but move it immediately after the call to
cpuhp_ap_sync_alive(). I think that will solve these issues, without
causing issues with the concurrent part of early boot and also without
reintroducing the early call to rcutree_report_cpu_dead().
Cheers,
Will
next prev parent reply other threads:[~2026-09-09 12:37 UTC|newest]
Thread overview: 44+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-07 16:40 [PATCH 00/19] arm64: Implement parallel CPU onlining with PSCI v0.2+ Will Deacon
2026-09-07 16:40 ` [PATCH 01/19] cpu/hotplug: Clean up cmpxchg() logic in cpuhp_can_boot_ap() Will Deacon
2026-09-08 2:55 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 02/19] cpu/hotplug: Avoid trying to bring up CPUs that are already online Will Deacon
2026-09-08 3:13 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 03/19] cpu/hotplug: Avoid busy-polling on archs where cpu_relax() is a no-op Will Deacon
2026-09-08 4:00 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 04/19] cpu/hotplug: Propagate bring-up status to arch_cpuhp_cleanup_kick_cpu() Will Deacon
2026-09-08 4:05 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 05/19] arm64: smp: Tidy up smp_prepare_cpus() Will Deacon
2026-09-08 7:35 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 06/19] arm64: smp: Tidy up cpuinfo init and cpufeature updates Will Deacon
2026-09-08 7:53 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 07/19] arm64: smp: Defer update of secondary CPU capabilities Will Deacon
2026-09-08 8:20 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 08/19] arm64: smp: Don't bother printing the I-cache policy for each CPU Will Deacon
2026-09-08 8:33 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 09/19] arm64: smp: Defer RCU registration during secondary CPU bringup Will Deacon
2026-09-08 8:55 ` Jinjie Ruan
2026-09-08 10:19 ` Will Deacon
2026-09-08 11:25 ` Jinjie Ruan
2026-09-09 12:36 ` Will Deacon [this message]
2026-09-10 2:47 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 10/19] arm64: smp: Use generic HOTPLUG_CORE_SYNC_FULL machinery for CPU onlining Will Deacon
2026-09-08 9:01 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 11/19] arm64: smp: Use generic HOTPLUG_SPLIT_STARTUP " Will Deacon
2026-09-08 9:10 ` Jinjie Ruan
2026-09-08 11:35 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 12/19] arm64: cpu_ops: Make 'cpu_operations' pointer global instead of per-cpu Will Deacon
2026-09-08 11:32 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 13/19] arm64: cpu_ops: Introduce get_secondary_cpu_ops() Will Deacon
2026-09-08 11:56 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 14/19] firmware/psci: Cache PSCI v0.2+ version number to avoid redundant SMCs Will Deacon
2026-09-08 11:57 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 15/19] firmware/psci: Extend ->cpu_on() callback to take an additional argument Will Deacon
2026-09-08 12:05 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 16/19] arm64: cpu_ops: Expose optional argument to target cpu in ->cpu_boot() Will Deacon
2026-09-08 12:12 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 17/19] arm64: smp: Pass secondary CPU boot parameters via firmware if possible Will Deacon
2026-09-08 12:16 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 18/19] arm64: smp: Use generic HOTPLUG_PARALLEL machinery for CPU onlining Will Deacon
2026-09-08 13:05 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 19/19] arm64: smp: Harden parallel CPU bringup against broken PSCI firmware Will Deacon
2026-09-08 13:36 ` Will Deacon
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqFS58mJ4GJNrbZ2@willie-the-truck \
--to=will@kernel.org \
--cc=bp@alien8.de \
--cc=catalin.marinas@arm.com \
--cc=dwmw@amazon.co.uk \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=lpieralisi@kernel.org \
--cc=mark.rutland@arm.com \
--cc=maz@kernel.org \
--cc=peterz@infradead.org \
--cc=ruanjinjie@huawei.com \
--cc=tglx@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.