From: Will Deacon <will@kernel.org>
To: David Woodhouse <dwmw2@infradead.org>
Cc: linux-arm-kernel@lists.infradead.org,
Pasha Tatashin <pasha.tatashin@soleen.com>,
Luka Absandze <absandze@amazon.de>,
linux-kernel@vger.kernel.org, Thomas Gleixner <tglx@kernel.org>,
Catalin Marinas <catalin.marinas@arm.com>,
Borislav Petkov <bp@alien8.de>,
Lorenzo Pieralisi <lpieralisi@kernel.org>,
Jinjie Ruan <ruanjinjie@huawei.com>,
Mark Rutland <mark.rutland@arm.com>,
Peter Zijlstra <peterz@infradead.org>,
Marc Zyngier <maz@kernel.org>
Subject: Re: [PATCH 00/19] arm64: Implement parallel CPU onlining with PSCI v0.2+
Date: Fri, 9 Oct 2026 12:08:03 +0100 [thread overview]
Message-ID: <asjLE0VkeVJKA554@willie-the-truck> (raw)
In-Reply-To: <f1ab0c03257d5055ada02877ac92a06f370c72b9.camel@infradead.org>
Hi David,
Thanks for replying with so much data! It looks like I sent my v2 out
at the same time.
On Fri, Oct 09, 2026 at 11:00:08AM +0100, David Woodhouse wrote:
> Thanks both of you for working on this. It was always on my list to
> come back and do it for arm64, but that list is long. Also on my list
> FWIW is parallelising the *later* stages of CPU hotplug¹, not just the
> early path into start_secondary()/secondary_start_kernel(). IIRC there
> was more win to be had, but we had to prove that a lot of once-
> serialized code could safely run concurrently. Ideally without just
> naïvely adding locking and serializing it all again.
That gets _really_ hard because you need RCU up and running pretty early.
> I gave your series a quick spin across a range of EC2 systems, and it
> gives a ~40% win on systems with 192 cores — at least, as a
> microbenchmark of the CPU onlining. For these systems, PSCI is fairly
> fast to bring the CPUs online and it isn't a huge proportion of the
> *overall* kexec time.
>
> However, your code in serial mode (cpuhp.parallel=0) is up to 40%
> *slower* than before. A/B testing results, all values in milliseconds:
Oh, that's unexpected.
> ┌───────────┬──────┬───────────┬────────────────┬─────────────────┐
> │ platform │ ncpu │ A 7.3-rc2 │ A' Will-serial │ B Will-parallel │
> ├───────────┼──────┼───────────┼────────────────┼─────────────────┤
> │ a1-metal │ 16 │ 5.7 │ 6.7 (+16.8%) │ 5.2 (−8.6%) │
> ├───────────┼──────┼───────────┼────────────────┼─────────────────┤
> │ a1-virt │ 16 │ 13.6 │ 15.7 (+15.4%) │ 10.5 (−22.5%) │
> ├───────────┼──────┼───────────┼────────────────┼─────────────────┤
> │ m6g-metal │ 64 │ 21.7 │ 29.8 (+37.3%) │ 15.5 (−28.8%) │
> ├───────────┼──────┼───────────┼────────────────┼─────────────────┤
> │ m6g-virt │ 64 │ 27.0 │ 30.9 (+14.5%) │ 23.0 (−14.6%) │
> ├───────────┼──────┼───────────┼────────────────┼─────────────────┤
> │ c7g-metal │ 64 │ 20.6 │ 26.0 (+26.3%) │ 11.9 (−41.9%) │
> ├───────────┼──────┼───────────┼────────────────┼─────────────────┤
> │ c7g-virt │ 64 │ 26.8 │ 31.0 (+15.4%) │ 21.7 (−19.1%) │
> ├───────────┼──────┼───────────┼────────────────┼─────────────────┤
> │ c8g-metal │ 192 │ 81.7 │ 111.2 (+36.0%) │ 53.0 (−35.2%) │
> ├───────────┼──────┼───────────┼────────────────┼─────────────────┤
> │ c8g-virt │ 192 │ 110.7 │ 120.4 (+8.8%) │ 96.4 (−12.9%) │
> ├───────────┼──────┼───────────┼────────────────┼─────────────────┤
> │ m9g-metal │ 192 │ 69.0 │ 97.2 (+40.9%) │ 43.9 (−36.3%) │
> ├───────────┼──────┼───────────┼────────────────┼─────────────────┤
> │ m9g-virt │ 192 │ 128.6 │ 138.9 (+8.0%) │ 112.2 (−12.7%) │
> └───────────┴──────┴───────────┴────────────────┴─────────────────┘
Would it be possible for you to pick one of these platforms and bisect
the serial regression, please? I'm not really sure where to look, but
serial boot should work all the way through the series so if you can
identify the point at which it regresses (and presumably doesn't recover)
then that would hopefully point me in the right direction. One possibility
is that the extra atomics in the new state machine logic are slowing things
down. Another possibility is the timeout logic in
cpuhp_wait_for_sync_state() ends up sleeping in your tests.
Probably worth using my v2 just in case one of the fixes there helps,
but I'm not hopeful.
> We also tried it on another system where the firmware takes about 9ms
> to bring each CPU up (after CPU_ON returns fairly quickly). Fanning out
> the CPU_ON calls didn't make any difference either; the CPUs came
> online, one at a time, about 9ms apart. The parallel onlining saved
> only about 20ms out of 820ms (→800ms) here.
Damn, I guess the firmware has some serialisation in that case?
> The real answer for such platforms (at least for kexec) is *not* to do
> the CPU_OFF/CPU_ON thing at all. Pasha's Caretaker work² still does so
> for the reclaim at hotplug time in the next kernel; I'm experimenting
> with eliding that, which should give the biggest improvement on such
> platforms. So during kexec the APs just spin and wait to be asked to
> come back, instead of going completely offline.
I was talking to Tarun about that at LPC. I was envisaging a form of
CPU_OFF that would leave the MMU enabled so we could make use of the
existing EFI boot logic, but I hadn't thought about it beyond that.
Will
next prev parent reply other threads:[~2026-10-09 11:08 UTC|newest]
Thread overview: 61+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-07 16:40 [PATCH 00/19] arm64: Implement parallel CPU onlining with PSCI v0.2+ Will Deacon
2026-09-07 16:40 ` [PATCH 01/19] cpu/hotplug: Clean up cmpxchg() logic in cpuhp_can_boot_ap() Will Deacon
2026-09-08 2:55 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 02/19] cpu/hotplug: Avoid trying to bring up CPUs that are already online Will Deacon
2026-09-08 3:13 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 03/19] cpu/hotplug: Avoid busy-polling on archs where cpu_relax() is a no-op Will Deacon
2026-09-08 4:00 ` Jinjie Ruan
2026-09-11 7:16 ` Jinjie Ruan
2026-09-11 12:57 ` Will Deacon
2026-09-16 1:15 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 04/19] cpu/hotplug: Propagate bring-up status to arch_cpuhp_cleanup_kick_cpu() Will Deacon
2026-09-08 4:05 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 05/19] arm64: smp: Tidy up smp_prepare_cpus() Will Deacon
2026-09-08 7:35 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 06/19] arm64: smp: Tidy up cpuinfo init and cpufeature updates Will Deacon
2026-09-08 7:53 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 07/19] arm64: smp: Defer update of secondary CPU capabilities Will Deacon
2026-09-08 8:20 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 08/19] arm64: smp: Don't bother printing the I-cache policy for each CPU Will Deacon
2026-09-08 8:33 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 09/19] arm64: smp: Defer RCU registration during secondary CPU bringup Will Deacon
2026-09-08 8:55 ` Jinjie Ruan
2026-09-08 10:19 ` Will Deacon
2026-09-08 11:25 ` Jinjie Ruan
2026-09-09 12:36 ` Will Deacon
2026-09-10 2:47 ` Jinjie Ruan
2026-09-11 12:52 ` Will Deacon
2026-09-16 2:40 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 10/19] arm64: smp: Use generic HOTPLUG_CORE_SYNC_FULL machinery for CPU onlining Will Deacon
2026-09-08 9:01 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 11/19] arm64: smp: Use generic HOTPLUG_SPLIT_STARTUP " Will Deacon
2026-09-08 9:10 ` Jinjie Ruan
2026-09-08 11:35 ` Jinjie Ruan
2026-09-11 12:55 ` Will Deacon
2026-09-07 16:40 ` [PATCH 12/19] arm64: cpu_ops: Make 'cpu_operations' pointer global instead of per-cpu Will Deacon
2026-09-08 11:32 ` Jinjie Ruan
2026-09-11 12:55 ` Will Deacon
2026-09-16 1:12 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 13/19] arm64: cpu_ops: Introduce get_secondary_cpu_ops() Will Deacon
2026-09-08 11:56 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 14/19] firmware/psci: Cache PSCI v0.2+ version number to avoid redundant SMCs Will Deacon
2026-09-08 11:57 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 15/19] firmware/psci: Extend ->cpu_on() callback to take an additional argument Will Deacon
2026-09-08 12:05 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 16/19] arm64: cpu_ops: Expose optional argument to target cpu in ->cpu_boot() Will Deacon
2026-09-08 12:12 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 17/19] arm64: smp: Pass secondary CPU boot parameters via firmware if possible Will Deacon
2026-09-08 12:16 ` Jinjie Ruan
2026-09-07 16:40 ` [PATCH 18/19] arm64: smp: Use generic HOTPLUG_PARALLEL machinery for CPU onlining Will Deacon
2026-09-08 13:05 ` Jinjie Ruan
2026-09-11 12:56 ` Will Deacon
2026-09-16 2:34 ` Jinjie Ruan
2026-09-18 15:18 ` Will Deacon
2026-09-07 16:40 ` [PATCH 19/19] arm64: smp: Harden parallel CPU bringup against broken PSCI firmware Will Deacon
2026-09-08 13:36 ` Will Deacon
2026-09-29 8:56 ` [PATCH 00/19] arm64: Implement parallel CPU onlining with PSCI v0.2+ Pankaj Patil
2026-10-08 9:38 ` Will Deacon
2026-10-09 10:00 ` David Woodhouse
2026-10-09 11:08 ` Will Deacon [this message]
2026-10-09 14:45 ` David Woodhouse
2026-10-09 16:04 ` David Woodhouse
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asjLE0VkeVJKA554@willie-the-truck \
--to=will@kernel.org \
--cc=absandze@amazon.de \
--cc=bp@alien8.de \
--cc=catalin.marinas@arm.com \
--cc=dwmw2@infradead.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=lpieralisi@kernel.org \
--cc=mark.rutland@arm.com \
--cc=maz@kernel.org \
--cc=pasha.tatashin@soleen.com \
--cc=peterz@infradead.org \
--cc=ruanjinjie@huawei.com \
--cc=tglx@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox