Linux-ARM-Kernel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Mark Rutland <mark.rutland@arm.com>
To: "Christoph Lameter (Ampere)" <cl@gentwo.org>
Cc: Yang Shi <yang@os.amperecomputing.com>,
	Ryan Roberts <ryan.roberts@arm.com>,
	dennis@kernel.org, tj@kernel.org, urezki@gmail.com,
	catalin.marinas@arm.com, will@kernel.org, david@kernel.org,
	akpm@linux-foundation.org, hca@linux.ibm.com, gor@linux.ibm.com,
	agordeev@linux.ibm.com, linux-mm@kvack.org,
	linux-arm-kernel@lists.infradead.org,
	linux-kernel@vger.kernel.org
Subject: Re: [RFC v2 PATCH 0/16] Optimize this_cpu_*() ops for non-x86 (ARM64 for this series)
Date: Wed, 5 Aug 2026 15:53:28 +0100	[thread overview]
Message-ID: <anNMvTAvUJkfspSR@J2N7QTR9R3> (raw)
In-Reply-To: <25d1e09b-53e4-7cd5-87db-b58437e4e690@gentwo.org>

On Mon, Jul 27, 2026 at 03:06:27PM -0700, Christoph Lameter (Ampere) wrote:
> On Wed, 22 Jul 2026, Mark Rutland wrote:
> > I expect that should come with a reasonable benefit, but I don't have
> > benchmark figures yet as I haven't finished converting the xchg and
> > cmpxchg implementations.
> >
> > > It sounds like it just moved the cost from one place to the other
> > > place and it also seems hacky TBH.
> 
> Yang Shi's patch has *no* critical section. There is no additional code
> for the RMV instruction. The RMV instruction is executed on the correct
> per cpu area.

None of that was in question.

> One of the reasons for the performance win is the
> eliminattion of these critical sections. Your approach still has some form
> of prologue and posthandling like the current preempt approach and
> therefore will not be able to have the same performance gains.

In absolute terms, yes.

However, I'm fairly confident that the vast majority of the overhead we
have today can be eliminated with simpler alternatives.

There is a trade-off, and there are surprisingly complex interactions
between page tables and other things (e.g. entry code). There is risk
and maintenance burden associated with that. Hence people want to
understand how much of the benefit is attributable to what. So far, the
statements haven't convinced me people actually know what portion of the
overhead come from which factor, e.g.

* How much of that attributable to conditional work when re-enabling
  preemption?

* How much of that is attributable to RMW sequences to modify the
  preempt count itself?
 
* How much of that is attributable to system register accesses (SP_EL0
  and TPIDR_ELx)?

Any of those could easily dominate the other factors and might easily be
avoidable. Most of that should be measurable today. For example you
could restore the preempt_{enable,disable} calls atop Yang Shi's
patches.

> The code is more efficient, there is no restart necessary and the
> technique is already widely used on x86 for a long time.

It's true that the per-cpu page table approach will have fewer
instructions in the fast path.

However, the other statements here are potentially misleading:

(1) There is no restart in the scheme I have proposed, so restarting is
    irrelevant to the comparison.

(2) On x86, this_cpu*() operations use segment relative addressing, NOT
    per-cpu page tables. If arm64 had a similar addressing scheme, I
    expect we would use it.

(3) There are a number of novel problems associated with per-cpu page
    tables (e.g the various unsolved issues Yang has described), which
    do not apply to x86's implementation of this_cpu_*() operations.

> Having the ability in general to map mmemory differently depending on the
> cpu opens up a number of other optimization like

> What we are proposing here is a basic new feature that simplifies code and
> allows addititonal performance and functional features that are so far not
> possible on ARM64.

While this simplifies the this_cpu_*() operations, I don't believe this
is a simplification overall, and IMO, describing it as such is
misleading.

Mark.


  parent reply	other threads:[~2026-08-05 14:53 UTC|newest]

Thread overview: 63+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-15 18:04 [RFC v2 PATCH 0/16] Optimize this_cpu_*() ops for non-x86 (ARM64 for this series) Yang Shi
2026-07-15 18:04 ` [PATCH 01/16] drivers: arch_numa: move percpu set up code to arch Yang Shi
2026-07-15 18:04 ` [PATCH 02/16] arm64: kconfig: make percpu related configs not depend on NUMA Yang Shi
2026-07-15 18:04 ` [PATCH 03/16] mm: pgalloc: introduce {pud|pmd}_populate_sync() Yang Shi
2026-07-15 18:04 ` [PATCH 04/16] vmalloc: pass in pgd pointer for vmap{__vunmap}_range_noflush() Yang Shi
2026-07-29  9:32   ` Lorenzo Stoakes (ARM)
2026-08-03 19:12     ` Yang Shi
2026-08-04 13:47       ` Lorenzo Stoakes (ARM)
2026-07-15 18:04 ` [PATCH 05/16] arm64: mm: enable percpu kernel page table Yang Shi
2026-07-15 18:04 ` [PATCH 06/16] arm64: mm: defined {pud|pmd}_populate_sync() Yang Shi
2026-07-15 18:04 ` [PATCH 07/16] arm64: mm: sync percpu page table for memory hotplug/unplug Yang Shi
2026-07-15 18:04 ` [PATCH 08/16] arm64: kasan: sync up kasan shadow area page table Yang Shi
2026-07-15 18:04 ` [PATCH 09/16] arm64: mm: define percpu virtual space area Yang Shi
2026-07-15 18:04 ` [PATCH 10/16] mm: percpu: prepare to use dedicated percpu area Yang Shi
2026-07-15 18:04 ` [PATCH 11/16] arm64: mm: map local percpu first chunk Yang Shi
2026-07-15 18:04 ` [PATCH 12/16] mm: percpu: set up first chunk and reserve chunk Yang Shi
2026-07-15 18:04 ` [PATCH 13/16] arm64: mm: introduce __per_cpu_local_off Yang Shi
2026-07-15 18:04 ` [PATCH 14/16] mm: percpu: allocate and free local percpu vm area Yang Shi
2026-07-15 18:04 ` [PATCH 15/16] arm64: kconfig: select HAVE_LOCAL_PER_CPU_MAP Yang Shi
2026-07-15 18:04 ` [PATCH 16/16] arm64: percpu: use local percpu for this_cpu_*() APIs Yang Shi
2026-07-16 13:23 ` [RFC v2 PATCH 0/16] Optimize this_cpu_*() ops for non-x86 (ARM64 for this series) Ryan Roberts
2026-07-21 19:08   ` Mark Rutland
2026-07-21 23:20   ` Yang Shi
2026-07-22  9:36     ` Mark Rutland
2026-07-27 21:10       ` Yang Shi
2026-07-27 22:06       ` Christoph Lameter (Ampere)
2026-07-29  9:28         ` David Hildenbrand (Arm)
2026-08-04 14:15           ` Lorenzo Stoakes (ARM)
2026-08-04 14:21             ` Lorenzo Stoakes (ARM)
2026-08-04 14:40               ` Jason Gunthorpe
2026-08-04 18:06                 ` Matthew Wilcox
2026-08-04 18:16                   ` Jason Gunthorpe
2026-08-04 15:21             ` Linus Torvalds
2026-08-04 16:15               ` Christoph Lameter (Ampere)
2026-08-04 16:30                 ` Linus Torvalds
2026-08-04 16:54                   ` David Hildenbrand (Arm)
2026-08-04 17:01                   ` Linus Torvalds
2026-08-04 17:32                     ` Lorenzo Stoakes (ARM)
2026-08-04 21:40                       ` Christoph Lameter (Ampere)
2026-08-04 21:48                         ` David Hildenbrand (Arm)
2026-08-04 21:56                           ` Christoph Lameter (Ampere)
2026-08-04 22:01                             ` David Hildenbrand (Arm)
2026-08-05  8:16                         ` Lorenzo Stoakes (ARM)
2026-08-04 17:23                   ` Lorenzo Stoakes (ARM)
2026-08-04 17:28                     ` Linus Torvalds
2026-08-04 21:51                     ` Yang Shi
2026-08-04 22:05                       ` David Hildenbrand (Arm)
2026-08-04 22:35                         ` Christoph Lameter (Ampere)
2026-08-05  6:11                           ` David Hildenbrand (Arm)
2026-08-05 14:48                           ` Mark Rutland
2026-08-05  7:52                       ` Lorenzo Stoakes (ARM)
2026-08-06 17:15                     ` Will Deacon
2026-08-06 17:31                       ` Lorenzo Stoakes (ARM)
2026-08-04 16:19             ` Christoph Lameter (Ampere)
2026-08-04 16:47               ` David Hildenbrand (Arm)
2026-08-04 21:25                 ` Christoph Lameter (Ampere)
2026-08-04 21:47                   ` David Hildenbrand (Arm)
2026-08-04 22:01                     ` Christoph Lameter (Ampere)
2026-08-05  6:12                       ` David Hildenbrand (Arm)
2026-08-05  8:45                     ` Heiko Carstens
2026-08-05 11:10                 ` David Laight
2026-08-05 14:53         ` Mark Rutland [this message]
2026-07-30  6:06     ` [RFC v2 PATCH 0/16] Optimize this_cpu_*() ops for non-x86 (ARM64 for this series)~ Mete Durlu

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anNMvTAvUJkfspSR@J2N7QTR9R3 \
    --to=mark.rutland@arm.com \
    --cc=agordeev@linux.ibm.com \
    --cc=akpm@linux-foundation.org \
    --cc=catalin.marinas@arm.com \
    --cc=cl@gentwo.org \
    --cc=david@kernel.org \
    --cc=dennis@kernel.org \
    --cc=gor@linux.ibm.com \
    --cc=hca@linux.ibm.com \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ryan.roberts@arm.com \
    --cc=tj@kernel.org \
    --cc=urezki@gmail.com \
    --cc=will@kernel.org \
    --cc=yang@os.amperecomputing.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox