From: David Laight <david.laight.linux@gmail.com>
To: "David Hildenbrand (Arm)" <david@kernel.org>
Cc: "Christoph Lameter (Ampere)" <cl@gentwo.org>,
"Lorenzo Stoakes (ARM)" <ljs@kernel.org>,
Mark Rutland <mark.rutland@arm.com>,
Yang Shi <yang@os.amperecomputing.com>,
Ryan Roberts <ryan.roberts@arm.com>,
dennis@kernel.org, tj@kernel.org, urezki@gmail.com,
catalin.marinas@arm.com, will@kernel.org,
akpm@linux-foundation.org, hca@linux.ibm.com, gor@linux.ibm.com,
agordeev@linux.ibm.com, linux-mm@kvack.org,
linux-arm-kernel@lists.infradead.org,
linux-kernel@vger.kernel.org,
Linus Torvalds <torvalds@linux-foundation.org>,
Jason Gunthorpe <jgg@nvidia.com>
Subject: Re: [RFC v2 PATCH 0/16] Optimize this_cpu_*() ops for non-x86 (ARM64 for this series)
Date: Wed, 5 Aug 2026 12:10:39 +0100 [thread overview]
Message-ID: <20260805121039.471c7340@pumpkin> (raw)
In-Reply-To: <84b31836-1d10-4bec-a469-468b91fd4645@kernel.org>
On Tue, 4 Aug 2026 18:47:26 +0200
"David Hildenbrand (Arm)" <david@kernel.org> wrote:
> On 8/4/26 18:19, Christoph Lameter (Ampere) wrote:
> > On Tue, 4 Aug 2026, Lorenzo Stoakes (ARM) wrote:
> >
> >> Since this work seems to be very much arm64-focused, perhaps it's therefore
> >> worth looking at an alterative solution that's specific to the arch, like the
> >> one suggested by Mark ([1])?
> >>
> >> [0]:https://lore.kernel.org/all/CAHk-=wire3dzhHx=KiL_f5Rj0=1u9ustsa33QoR-F9-v-NU9Ng@mail.gmail.com/
> >> [1]:https://lore.kernel.org/linux-arm-kernel/al_DpFJFcmVhxpvW@J2N7QTR9R3/
> >
> > Mark's solution does replace the preempt_enable/disable sections with a
> > rather hacky restart logic. It relies on a long preemable and postscript
> > to each per cpu operations.
>
> Okay, so 3 simple instructions of preemable is "long preemable"? In which universe?
>
> But I am sure you did you homework and have data to back up your claims. Please
> share that data, because I am very curious.
The proposed sequence is:
> // Prologue. Enable fixups for <off> and <addr>.
> 1 mrs <tsk>, sp_el0
> 2 mov <tmp>, #__VAL_PCPU_GPRS(<pcp>, <off>, <addr>)
> 3 strh <tmp>, [<tsk>, #TSK_TI_PCPU_GPRS]
>
> // Generate cpu-specific address
> 4 mrs <off>, TPIDR_ELx
> 5 add <addr>, <pcp>, <off>
>
> // Perform access sequence
> 6 ldr <val>, [<addr>]
>
> // Epilogue. Disable fixups
> 7 strh wzr, [<tsk>, #TSK_TI_PCPU_GPRS]
Think about how that actually gets execute by a real cpu.
Instructions will be read from the I-cache in 'chunks' (maybe half a cache line).
They are then fed to multiple decoders that generate u-ops for the execution units.
The decoder is unlikely to be a bottleneck.
I've numbered the instructions:
First clock can run instructions 1, 2 and 4.
Assuming the mrs have no extra latency the second runs 3 and 5.
The third will then run 6 and 7.
The cpu then probably has to wait for the result of the ldr.
If the access is a write then there may be a stall waiting for the value
to be written to be available.
The only real effect of the extra instructions is likely to be code size.
David
next prev parent reply other threads:[~2026-08-05 11:10 UTC|newest]
Thread overview: 63+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-15 18:04 [RFC v2 PATCH 0/16] Optimize this_cpu_*() ops for non-x86 (ARM64 for this series) Yang Shi
2026-07-15 18:04 ` [PATCH 01/16] drivers: arch_numa: move percpu set up code to arch Yang Shi
2026-07-15 18:04 ` [PATCH 02/16] arm64: kconfig: make percpu related configs not depend on NUMA Yang Shi
2026-07-15 18:04 ` [PATCH 03/16] mm: pgalloc: introduce {pud|pmd}_populate_sync() Yang Shi
2026-07-15 18:04 ` [PATCH 04/16] vmalloc: pass in pgd pointer for vmap{__vunmap}_range_noflush() Yang Shi
2026-07-29 9:32 ` Lorenzo Stoakes (ARM)
2026-08-03 19:12 ` Yang Shi
2026-08-04 13:47 ` Lorenzo Stoakes (ARM)
2026-07-15 18:04 ` [PATCH 05/16] arm64: mm: enable percpu kernel page table Yang Shi
2026-07-15 18:04 ` [PATCH 06/16] arm64: mm: defined {pud|pmd}_populate_sync() Yang Shi
2026-07-15 18:04 ` [PATCH 07/16] arm64: mm: sync percpu page table for memory hotplug/unplug Yang Shi
2026-07-15 18:04 ` [PATCH 08/16] arm64: kasan: sync up kasan shadow area page table Yang Shi
2026-07-15 18:04 ` [PATCH 09/16] arm64: mm: define percpu virtual space area Yang Shi
2026-07-15 18:04 ` [PATCH 10/16] mm: percpu: prepare to use dedicated percpu area Yang Shi
2026-07-15 18:04 ` [PATCH 11/16] arm64: mm: map local percpu first chunk Yang Shi
2026-07-15 18:04 ` [PATCH 12/16] mm: percpu: set up first chunk and reserve chunk Yang Shi
2026-07-15 18:04 ` [PATCH 13/16] arm64: mm: introduce __per_cpu_local_off Yang Shi
2026-07-15 18:04 ` [PATCH 14/16] mm: percpu: allocate and free local percpu vm area Yang Shi
2026-07-15 18:04 ` [PATCH 15/16] arm64: kconfig: select HAVE_LOCAL_PER_CPU_MAP Yang Shi
2026-07-15 18:04 ` [PATCH 16/16] arm64: percpu: use local percpu for this_cpu_*() APIs Yang Shi
2026-07-16 13:23 ` [RFC v2 PATCH 0/16] Optimize this_cpu_*() ops for non-x86 (ARM64 for this series) Ryan Roberts
2026-07-21 19:08 ` Mark Rutland
2026-07-21 23:20 ` Yang Shi
2026-07-22 9:36 ` Mark Rutland
2026-07-27 21:10 ` Yang Shi
2026-07-27 22:06 ` Christoph Lameter (Ampere)
2026-07-29 9:28 ` David Hildenbrand (Arm)
2026-08-04 14:15 ` Lorenzo Stoakes (ARM)
2026-08-04 14:21 ` Lorenzo Stoakes (ARM)
2026-08-04 14:40 ` Jason Gunthorpe
2026-08-04 18:06 ` Matthew Wilcox
2026-08-04 18:16 ` Jason Gunthorpe
2026-08-04 15:21 ` Linus Torvalds
2026-08-04 16:15 ` Christoph Lameter (Ampere)
2026-08-04 16:30 ` Linus Torvalds
2026-08-04 16:54 ` David Hildenbrand (Arm)
2026-08-04 17:01 ` Linus Torvalds
2026-08-04 17:32 ` Lorenzo Stoakes (ARM)
2026-08-04 21:40 ` Christoph Lameter (Ampere)
2026-08-04 21:48 ` David Hildenbrand (Arm)
2026-08-04 21:56 ` Christoph Lameter (Ampere)
2026-08-04 22:01 ` David Hildenbrand (Arm)
2026-08-05 8:16 ` Lorenzo Stoakes (ARM)
2026-08-04 17:23 ` Lorenzo Stoakes (ARM)
2026-08-04 17:28 ` Linus Torvalds
2026-08-04 21:51 ` Yang Shi
2026-08-04 22:05 ` David Hildenbrand (Arm)
2026-08-04 22:35 ` Christoph Lameter (Ampere)
2026-08-05 6:11 ` David Hildenbrand (Arm)
2026-08-05 14:48 ` Mark Rutland
2026-08-05 7:52 ` Lorenzo Stoakes (ARM)
2026-08-06 17:15 ` Will Deacon
2026-08-06 17:31 ` Lorenzo Stoakes (ARM)
2026-08-04 16:19 ` Christoph Lameter (Ampere)
2026-08-04 16:47 ` David Hildenbrand (Arm)
2026-08-04 21:25 ` Christoph Lameter (Ampere)
2026-08-04 21:47 ` David Hildenbrand (Arm)
2026-08-04 22:01 ` Christoph Lameter (Ampere)
2026-08-05 6:12 ` David Hildenbrand (Arm)
2026-08-05 8:45 ` Heiko Carstens
2026-08-05 11:10 ` David Laight [this message]
2026-08-05 14:53 ` Mark Rutland
2026-07-30 6:06 ` [RFC v2 PATCH 0/16] Optimize this_cpu_*() ops for non-x86 (ARM64 for this series)~ Mete Durlu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260805121039.471c7340@pumpkin \
--to=david.laight.linux@gmail.com \
--cc=agordeev@linux.ibm.com \
--cc=akpm@linux-foundation.org \
--cc=catalin.marinas@arm.com \
--cc=cl@gentwo.org \
--cc=david@kernel.org \
--cc=dennis@kernel.org \
--cc=gor@linux.ibm.com \
--cc=hca@linux.ibm.com \
--cc=jgg@nvidia.com \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mark.rutland@arm.com \
--cc=ryan.roberts@arm.com \
--cc=tj@kernel.org \
--cc=torvalds@linux-foundation.org \
--cc=urezki@gmail.com \
--cc=will@kernel.org \
--cc=yang@os.amperecomputing.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox