From: Peter Zijlstra <peterz@infradead.org>
To: Mark Rutland <mark.rutland@arm.com>
Cc: vladimir.murzin@arm.com, catalin.marinas@arm.com,
hca@linux.ibm.com, linux-kernel@vger.kernel.org,
ruanjinjie@huawei.com, yang@os.amperecomputing.com,
maz@kernel.org, will@kernel.org, ardb@kernel.org,
linux-arm-kernel@lists.infradead.org
Subject: Re: [RFC PATCH 06/13] arm64: percpu: Add infrastructure for preemptible this_cpu_*() ops
Date: Tue, 28 Jul 2026 16:19:58 +0200 [thread overview]
Message-ID: <20260728141958.GV751831@noisy.programming.kicks-ass.net> (raw)
In-Reply-To: <20260728123859.2911495-7-mark.rutland@arm.com>
On Tue, Jul 28, 2026 at 01:38:52PM +0100, Mark Rutland wrote:
> With the scheme added in this patch, this can be compiled as:
>
> | <outline_this_cpu_add_u64>:
> | mrs x2, sp_el0
> | mov x4, #0xc80
Perhaps add a comment to that like:
| mov x4, #0xc80 # {x0, x4, x3}
to make the literal less magical?
> | strh w4, [x2, #20]
Anyway, interrupt can hit here, at which point it will:
x4 := this_cpu_offset
x3 := x0 + x4
> | mrs x4, tpidr_el1
> | add x3, x0, x4
Which it then recomputes, and so is harmless.
> | 1: ldxr x6, [x3]
> | add x6, x6, x1
> | stxr w5, x6, [x3]
IIRC any interrupt in between LL/SC will cause the stxr to fail anyway,
irrespective of the fixup, so goto 1b we go.
> | cbnz w5, 1b
If we get here and interrupts happens, we recompute pointlessly, no harm
done.
Anyway, per this sequence there is no point in ever doing the fixup. So
perhaps give a better example?
> | strh wzr, [x2, #20]
> | ret
>
> TODO: Save/restore the PCPU GPRs in __sdei_asm_handler(). This will
> require some mechanical rework to the __sdei_asm_handler() assembly.
>
> Signed-off-by: Mark Rutland <mark.rutland@arm.com>
> @@ -51,6 +53,47 @@ static inline unsigned long __kern_my_cpu_offset(void)
> return off;
> }
>
> +#define PCPU_GPR_PCP GENMASK(4, 0)
> +#define PCPU_GPR_OFF GENMASK(9, 5)
> +#define PCPU_GPR_ADDR GENMASK(14, 10)
I hate those macros, I can never remember what they all do. Now, esp.
since you then not use them here:
> +#define __VAL_PCPU_GPRS(pcp, off, addr) \
> + "(" \
> + "(.L__gpr_num_" pcp " << 0) | " \
> + "(.L__gpr_num_" off " << 5) | " \
> + "(.L__gpr_num_" addr " << 10)" \
> + ")"
... continue comment at irqentry_exit_pcpu_adjust().
> +#define ____PCPU_GPRS_BEGIN(gprs, pcp, off, addr) \
> + __DEFINE_ASM_GPR_NUMS \
> + __DEFINE_ASM_GPR_ALIASES \
> + " mov w" off ", #" __VAL_PCPU_GPRS(pcp, off, addr) "\n" \
> + " strh w" off ", " gprs "\n" \
> + __KERN_ASM_CPU_OFFSET(off) "\n"
Can this macro also generate a readable comment for those few of us
building the .i file ?
> +/*
> + * Where the context being returned to had an active percpu GPR critical
> + * section, ensure that the offset and address GPRs are updated to match the
> + * current CPU.
> + *
> + * For simplicity we always update the GPRs when a critical section is active
> + * and preemption was *possible*, regardless of whether preemption actually
> + * occurred. Where preemption did not occur, the updates are redundant but not
> + * harmful.
> + */
> +static __always_inline void irqentry_exit_pcpu_adjust(struct pt_regs *regs)
> +{
> + int reg_pcp, reg_off, reg_addr;
> + unsigned long pcp, off, addr;
> + u16 gprs = regs->pcpu_gprs;
> +
> + /*
> + * Zero means no active PCPU GPRs. As the PCPU GPRs must be distinct,
> + * a PCPU critical section cannot possibly use {x0,x0,x0}.
> + */
> + if (likely(!gprs))
> + return;
> +
> + reg_pcp = FIELD_GET(PCPU_GPR_PCP, gprs);
> + reg_off = FIELD_GET(PCPU_GPR_OFF, gprs);
> + reg_addr = FIELD_GET(PCPU_GPR_ADDR, gprs);
Would it not be more readable to write this like:
reg_pcp = (grps >> 0) & 0x1f;
reg_off = (grps >> 5) & 0x1f;
reg_addr = (grps >> 10) & 0x1f;
To better match __VAL_PCPU_GRPS() ?
> + pcp = pt_regs_read_reg(regs, reg_pcp);
> +
> + off = __kern_my_cpu_offset();
> + pt_regs_write_reg(regs, reg_off, off);
> +
> + addr = pcp + off;
> + pt_regs_write_reg(regs, reg_addr, addr);
> +}
WARNING: multiple messages have this Message-ID (diff)
From: Peter Zijlstra <peterz@infradead.org>
To: Mark Rutland <mark.rutland@arm.com>
Cc: linux-arm-kernel@lists.infradead.org, ada.coupriediaz@arm.com,
ardb@kernel.org, catalin.marinas@arm.com, hca@linux.ibm.com,
linux-kernel@vger.kernel.org, maz@kernel.org,
ruanjinjie@huawei.com, vladimir.murzin@arm.com, will@kernel.org,
yang@os.amperecomputing.com
Subject: Re: [RFC PATCH 06/13] arm64: percpu: Add infrastructure for preemptible this_cpu_*() ops
Date: Tue, 28 Jul 2026 16:19:58 +0200 [thread overview]
Message-ID: <20260728141958.GV751831@noisy.programming.kicks-ass.net> (raw)
In-Reply-To: <20260728123859.2911495-7-mark.rutland@arm.com>
On Tue, Jul 28, 2026 at 01:38:52PM +0100, Mark Rutland wrote:
> With the scheme added in this patch, this can be compiled as:
>
> | <outline_this_cpu_add_u64>:
> | mrs x2, sp_el0
> | mov x4, #0xc80
Perhaps add a comment to that like:
| mov x4, #0xc80 # {x0, x4, x3}
to make the literal less magical?
> | strh w4, [x2, #20]
Anyway, interrupt can hit here, at which point it will:
x4 := this_cpu_offset
x3 := x0 + x4
> | mrs x4, tpidr_el1
> | add x3, x0, x4
Which it then recomputes, and so is harmless.
> | 1: ldxr x6, [x3]
> | add x6, x6, x1
> | stxr w5, x6, [x3]
IIRC any interrupt in between LL/SC will cause the stxr to fail anyway,
irrespective of the fixup, so goto 1b we go.
> | cbnz w5, 1b
If we get here and interrupts happens, we recompute pointlessly, no harm
done.
Anyway, per this sequence there is no point in ever doing the fixup. So
perhaps give a better example?
> | strh wzr, [x2, #20]
> | ret
>
> TODO: Save/restore the PCPU GPRs in __sdei_asm_handler(). This will
> require some mechanical rework to the __sdei_asm_handler() assembly.
>
> Signed-off-by: Mark Rutland <mark.rutland@arm.com>
> @@ -51,6 +53,47 @@ static inline unsigned long __kern_my_cpu_offset(void)
> return off;
> }
>
> +#define PCPU_GPR_PCP GENMASK(4, 0)
> +#define PCPU_GPR_OFF GENMASK(9, 5)
> +#define PCPU_GPR_ADDR GENMASK(14, 10)
I hate those macros, I can never remember what they all do. Now, esp.
since you then not use them here:
> +#define __VAL_PCPU_GPRS(pcp, off, addr) \
> + "(" \
> + "(.L__gpr_num_" pcp " << 0) | " \
> + "(.L__gpr_num_" off " << 5) | " \
> + "(.L__gpr_num_" addr " << 10)" \
> + ")"
... continue comment at irqentry_exit_pcpu_adjust().
> +#define ____PCPU_GPRS_BEGIN(gprs, pcp, off, addr) \
> + __DEFINE_ASM_GPR_NUMS \
> + __DEFINE_ASM_GPR_ALIASES \
> + " mov w" off ", #" __VAL_PCPU_GPRS(pcp, off, addr) "\n" \
> + " strh w" off ", " gprs "\n" \
> + __KERN_ASM_CPU_OFFSET(off) "\n"
Can this macro also generate a readable comment for those few of us
building the .i file ?
> +/*
> + * Where the context being returned to had an active percpu GPR critical
> + * section, ensure that the offset and address GPRs are updated to match the
> + * current CPU.
> + *
> + * For simplicity we always update the GPRs when a critical section is active
> + * and preemption was *possible*, regardless of whether preemption actually
> + * occurred. Where preemption did not occur, the updates are redundant but not
> + * harmful.
> + */
> +static __always_inline void irqentry_exit_pcpu_adjust(struct pt_regs *regs)
> +{
> + int reg_pcp, reg_off, reg_addr;
> + unsigned long pcp, off, addr;
> + u16 gprs = regs->pcpu_gprs;
> +
> + /*
> + * Zero means no active PCPU GPRs. As the PCPU GPRs must be distinct,
> + * a PCPU critical section cannot possibly use {x0,x0,x0}.
> + */
> + if (likely(!gprs))
> + return;
> +
> + reg_pcp = FIELD_GET(PCPU_GPR_PCP, gprs);
> + reg_off = FIELD_GET(PCPU_GPR_OFF, gprs);
> + reg_addr = FIELD_GET(PCPU_GPR_ADDR, gprs);
Would it not be more readable to write this like:
reg_pcp = (grps >> 0) & 0x1f;
reg_off = (grps >> 5) & 0x1f;
reg_addr = (grps >> 10) & 0x1f;
To better match __VAL_PCPU_GRPS() ?
> + pcp = pt_regs_read_reg(regs, reg_pcp);
> +
> + off = __kern_my_cpu_offset();
> + pt_regs_write_reg(regs, reg_off, off);
> +
> + addr = pcp + off;
> + pt_regs_write_reg(regs, reg_addr, addr);
> +}
next prev parent reply other threads:[~2026-07-28 14:20 UTC|newest]
Thread overview: 48+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-28 12:38 [RFC PATCH 00/13] arm64: Preemptible this_cpu_*() operations Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 01/13] arm64: preempt: Simplify and optimize __preempt_count_dec_and_test() Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 02/13] arm64: preempt: Treat should_resched() as unlikely Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 03/13] arm64: ptrace: Always inline pt_regs_[read,write}_reg() Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 04/13] arm64: percpu: Factor out percpu offset asm Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 05/13] arm64: gpr-num: Add wxN aliases for wN registers Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 06/13] arm64: percpu: Add infrastructure for preemptible this_cpu_*() ops Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 14:19 ` Peter Zijlstra [this message]
2026-07-28 14:19 ` Peter Zijlstra
2026-07-28 14:30 ` Peter Zijlstra
2026-07-28 14:30 ` Peter Zijlstra
2026-07-28 16:03 ` Mark Rutland
2026-07-28 16:03 ` Mark Rutland
2026-07-28 16:49 ` Peter Zijlstra
2026-07-28 16:49 ` Peter Zijlstra
2026-07-28 15:53 ` Mark Rutland
2026-07-28 15:53 ` Mark Rutland
2026-07-28 16:46 ` Peter Zijlstra
2026-07-28 16:46 ` Peter Zijlstra
2026-07-28 16:49 ` Peter Zijlstra
2026-07-28 16:49 ` Peter Zijlstra
2026-07-28 16:51 ` Heiko Carstens
2026-07-28 16:51 ` Heiko Carstens
2026-07-28 12:38 ` [RFC PATCH 07/13] arm64: percpu: Implement preemptible read/write ops Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 21:47 ` David Laight
2026-07-28 21:47 ` David Laight
2026-07-28 22:05 ` David Laight
2026-07-28 22:05 ` David Laight
2026-07-28 12:38 ` [RFC PATCH 08/13] arm64: percpu: Implement preemptible void RMW ops Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 09/13] arm64: percpu: Implement preemptible return " Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 10/13] arm64: percpu: Implement preemptible XCHG ops Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 11/13] arm64: percpu: Implement preemptible CMPXCHG ops Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 12/13] arm64: percpu: Implement preemptible CMPXCHG128 ops Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 13/13] arm64: percpu: Remove _pcp_protect*() wrappers Mark Rutland
2026-07-28 12:38 ` Mark Rutland
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260728141958.GV751831@noisy.programming.kicks-ass.net \
--to=peterz@infradead.org \
--cc=ardb@kernel.org \
--cc=catalin.marinas@arm.com \
--cc=hca@linux.ibm.com \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=maz@kernel.org \
--cc=ruanjinjie@huawei.com \
--cc=vladimir.murzin@arm.com \
--cc=will@kernel.org \
--cc=yang@os.amperecomputing.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.