All of lore.kernel.org
 help / color / mirror / Atom feed
From: Peter Zijlstra <peterz@infradead.org>
To: Mark Rutland <mark.rutland@arm.com>
Cc: vladimir.murzin@arm.com, catalin.marinas@arm.com,
	hca@linux.ibm.com, linux-kernel@vger.kernel.org,
	ruanjinjie@huawei.com, yang@os.amperecomputing.com,
	maz@kernel.org, will@kernel.org, ardb@kernel.org,
	linux-arm-kernel@lists.infradead.org
Subject: Re: [RFC PATCH 06/13] arm64: percpu: Add infrastructure for preemptible this_cpu_*() ops
Date: Tue, 28 Jul 2026 16:19:58 +0200	[thread overview]
Message-ID: <20260728141958.GV751831@noisy.programming.kicks-ass.net> (raw)
In-Reply-To: <20260728123859.2911495-7-mark.rutland@arm.com>

On Tue, Jul 28, 2026 at 01:38:52PM +0100, Mark Rutland wrote:

> With the scheme added in this patch, this can be compiled as:
> 
> | <outline_this_cpu_add_u64>:
> |        mrs     x2, sp_el0
> |        mov     x4, #0xc80

Perhaps add a comment to that like:

  |        mov     x4, #0xc80 # {x0, x4, x3}

to make the literal less magical?

> |        strh    w4, [x2, #20]

Anyway, interrupt can hit here, at which point it will:

  x4 := this_cpu_offset
  x3 := x0 + x4

> |        mrs     x4, tpidr_el1
> |        add     x3, x0, x4

Which it then recomputes, and so is harmless.

> | 1:     ldxr    x6, [x3]
> |        add     x6, x6, x1
> |        stxr    w5, x6, [x3]

IIRC any interrupt in between LL/SC will cause the stxr to fail anyway,
irrespective of the fixup, so goto 1b we go.

> |        cbnz    w5, 1b

If we get here and interrupts happens, we recompute pointlessly, no harm
done.

Anyway, per this sequence there is no point in ever doing the fixup. So
perhaps give a better example?

> |        strh    wzr, [x2, #20]
> |        ret
> 
> TODO: Save/restore the PCPU GPRs in __sdei_asm_handler(). This will
> require some mechanical rework to the __sdei_asm_handler() assembly.
> 
> Signed-off-by: Mark Rutland <mark.rutland@arm.com>

> @@ -51,6 +53,47 @@ static inline unsigned long __kern_my_cpu_offset(void)
>  	return off;
>  }
>  
> +#define PCPU_GPR_PCP			GENMASK(4, 0)
> +#define PCPU_GPR_OFF			GENMASK(9, 5)
> +#define PCPU_GPR_ADDR			GENMASK(14, 10)

I hate those macros, I can never remember what they all do. Now, esp.
since you then not use them here:

> +#define __VAL_PCPU_GPRS(pcp, off, addr)				\
> +	"("							\
> +		"(.L__gpr_num_" pcp  " << 0) | "		\
> +		"(.L__gpr_num_" off  " << 5) | "		\
> +		"(.L__gpr_num_" addr " << 10)"			\
> +	")"

... continue comment at irqentry_exit_pcpu_adjust().

> +#define ____PCPU_GPRS_BEGIN(gprs, pcp, off, addr)			\
> +	__DEFINE_ASM_GPR_NUMS						\
> +	__DEFINE_ASM_GPR_ALIASES					\
> +	"	mov w" off ", #" __VAL_PCPU_GPRS(pcp, off, addr) "\n"	\
> +	"	strh	w" off ", " gprs "\n"				\
> +	__KERN_ASM_CPU_OFFSET(off) "\n"

Can this macro also generate a readable comment for those few of us
building the .i file ?

> +/*
> + * Where the context being returned to had an active percpu GPR critical
> + * section, ensure that the offset and address GPRs are updated to match the
> + * current CPU.
> + *
> + * For simplicity we always update the GPRs when a critical section is active
> + * and preemption was *possible*, regardless of whether preemption actually
> + * occurred. Where preemption did not occur, the updates are redundant but not
> + * harmful.
> + */
> +static __always_inline void irqentry_exit_pcpu_adjust(struct pt_regs *regs)
> +{
> +	int reg_pcp, reg_off, reg_addr;
> +	unsigned long pcp, off, addr;
> +	u16 gprs = regs->pcpu_gprs;
> +
> +	/*
> +	 * Zero means no active PCPU GPRs. As the PCPU GPRs must be distinct,
> +	 * a PCPU critical section cannot possibly use {x0,x0,x0}.
> +	 */
> +	if (likely(!gprs))
> +		return;
> +
> +	reg_pcp  = FIELD_GET(PCPU_GPR_PCP,  gprs);
> +	reg_off  = FIELD_GET(PCPU_GPR_OFF,  gprs);
> +	reg_addr = FIELD_GET(PCPU_GPR_ADDR, gprs);

Would it not be more readable to write this like:

	reg_pcp  = (grps >>  0) & 0x1f;
	reg_off  = (grps >>  5) & 0x1f;
	reg_addr = (grps >> 10) & 0x1f;

To better match __VAL_PCPU_GRPS() ?

> +	pcp = pt_regs_read_reg(regs, reg_pcp);
> +
> +	off = __kern_my_cpu_offset();
> +	pt_regs_write_reg(regs, reg_off, off);
> +
> +	addr = pcp + off;
> +	pt_regs_write_reg(regs, reg_addr, addr);
> +}


WARNING: multiple messages have this Message-ID (diff)
From: Peter Zijlstra <peterz@infradead.org>
To: Mark Rutland <mark.rutland@arm.com>
Cc: linux-arm-kernel@lists.infradead.org, ada.coupriediaz@arm.com,
	ardb@kernel.org, catalin.marinas@arm.com, hca@linux.ibm.com,
	linux-kernel@vger.kernel.org, maz@kernel.org,
	ruanjinjie@huawei.com, vladimir.murzin@arm.com, will@kernel.org,
	yang@os.amperecomputing.com
Subject: Re: [RFC PATCH 06/13] arm64: percpu: Add infrastructure for preemptible this_cpu_*() ops
Date: Tue, 28 Jul 2026 16:19:58 +0200	[thread overview]
Message-ID: <20260728141958.GV751831@noisy.programming.kicks-ass.net> (raw)
In-Reply-To: <20260728123859.2911495-7-mark.rutland@arm.com>

On Tue, Jul 28, 2026 at 01:38:52PM +0100, Mark Rutland wrote:

> With the scheme added in this patch, this can be compiled as:
> 
> | <outline_this_cpu_add_u64>:
> |        mrs     x2, sp_el0
> |        mov     x4, #0xc80

Perhaps add a comment to that like:

  |        mov     x4, #0xc80 # {x0, x4, x3}

to make the literal less magical?

> |        strh    w4, [x2, #20]

Anyway, interrupt can hit here, at which point it will:

  x4 := this_cpu_offset
  x3 := x0 + x4

> |        mrs     x4, tpidr_el1
> |        add     x3, x0, x4

Which it then recomputes, and so is harmless.

> | 1:     ldxr    x6, [x3]
> |        add     x6, x6, x1
> |        stxr    w5, x6, [x3]

IIRC any interrupt in between LL/SC will cause the stxr to fail anyway,
irrespective of the fixup, so goto 1b we go.

> |        cbnz    w5, 1b

If we get here and interrupts happens, we recompute pointlessly, no harm
done.

Anyway, per this sequence there is no point in ever doing the fixup. So
perhaps give a better example?

> |        strh    wzr, [x2, #20]
> |        ret
> 
> TODO: Save/restore the PCPU GPRs in __sdei_asm_handler(). This will
> require some mechanical rework to the __sdei_asm_handler() assembly.
> 
> Signed-off-by: Mark Rutland <mark.rutland@arm.com>

> @@ -51,6 +53,47 @@ static inline unsigned long __kern_my_cpu_offset(void)
>  	return off;
>  }
>  
> +#define PCPU_GPR_PCP			GENMASK(4, 0)
> +#define PCPU_GPR_OFF			GENMASK(9, 5)
> +#define PCPU_GPR_ADDR			GENMASK(14, 10)

I hate those macros, I can never remember what they all do. Now, esp.
since you then not use them here:

> +#define __VAL_PCPU_GPRS(pcp, off, addr)				\
> +	"("							\
> +		"(.L__gpr_num_" pcp  " << 0) | "		\
> +		"(.L__gpr_num_" off  " << 5) | "		\
> +		"(.L__gpr_num_" addr " << 10)"			\
> +	")"

... continue comment at irqentry_exit_pcpu_adjust().

> +#define ____PCPU_GPRS_BEGIN(gprs, pcp, off, addr)			\
> +	__DEFINE_ASM_GPR_NUMS						\
> +	__DEFINE_ASM_GPR_ALIASES					\
> +	"	mov w" off ", #" __VAL_PCPU_GPRS(pcp, off, addr) "\n"	\
> +	"	strh	w" off ", " gprs "\n"				\
> +	__KERN_ASM_CPU_OFFSET(off) "\n"

Can this macro also generate a readable comment for those few of us
building the .i file ?

> +/*
> + * Where the context being returned to had an active percpu GPR critical
> + * section, ensure that the offset and address GPRs are updated to match the
> + * current CPU.
> + *
> + * For simplicity we always update the GPRs when a critical section is active
> + * and preemption was *possible*, regardless of whether preemption actually
> + * occurred. Where preemption did not occur, the updates are redundant but not
> + * harmful.
> + */
> +static __always_inline void irqentry_exit_pcpu_adjust(struct pt_regs *regs)
> +{
> +	int reg_pcp, reg_off, reg_addr;
> +	unsigned long pcp, off, addr;
> +	u16 gprs = regs->pcpu_gprs;
> +
> +	/*
> +	 * Zero means no active PCPU GPRs. As the PCPU GPRs must be distinct,
> +	 * a PCPU critical section cannot possibly use {x0,x0,x0}.
> +	 */
> +	if (likely(!gprs))
> +		return;
> +
> +	reg_pcp  = FIELD_GET(PCPU_GPR_PCP,  gprs);
> +	reg_off  = FIELD_GET(PCPU_GPR_OFF,  gprs);
> +	reg_addr = FIELD_GET(PCPU_GPR_ADDR, gprs);

Would it not be more readable to write this like:

	reg_pcp  = (grps >>  0) & 0x1f;
	reg_off  = (grps >>  5) & 0x1f;
	reg_addr = (grps >> 10) & 0x1f;

To better match __VAL_PCPU_GRPS() ?

> +	pcp = pt_regs_read_reg(regs, reg_pcp);
> +
> +	off = __kern_my_cpu_offset();
> +	pt_regs_write_reg(regs, reg_off, off);
> +
> +	addr = pcp + off;
> +	pt_regs_write_reg(regs, reg_addr, addr);
> +}

  reply	other threads:[~2026-07-28 14:20 UTC|newest]

Thread overview: 48+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-28 12:38 [RFC PATCH 00/13] arm64: Preemptible this_cpu_*() operations Mark Rutland
2026-07-28 12:38 ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 01/13] arm64: preempt: Simplify and optimize __preempt_count_dec_and_test() Mark Rutland
2026-07-28 12:38   ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 02/13] arm64: preempt: Treat should_resched() as unlikely Mark Rutland
2026-07-28 12:38   ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 03/13] arm64: ptrace: Always inline pt_regs_[read,write}_reg() Mark Rutland
2026-07-28 12:38   ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 04/13] arm64: percpu: Factor out percpu offset asm Mark Rutland
2026-07-28 12:38   ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 05/13] arm64: gpr-num: Add wxN aliases for wN registers Mark Rutland
2026-07-28 12:38   ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 06/13] arm64: percpu: Add infrastructure for preemptible this_cpu_*() ops Mark Rutland
2026-07-28 12:38   ` Mark Rutland
2026-07-28 14:19   ` Peter Zijlstra [this message]
2026-07-28 14:19     ` Peter Zijlstra
2026-07-28 14:30     ` Peter Zijlstra
2026-07-28 14:30       ` Peter Zijlstra
2026-07-28 16:03       ` Mark Rutland
2026-07-28 16:03         ` Mark Rutland
2026-07-28 16:49         ` Peter Zijlstra
2026-07-28 16:49           ` Peter Zijlstra
2026-07-28 15:53     ` Mark Rutland
2026-07-28 15:53       ` Mark Rutland
2026-07-28 16:46       ` Peter Zijlstra
2026-07-28 16:46         ` Peter Zijlstra
2026-07-28 16:49       ` Peter Zijlstra
2026-07-28 16:49         ` Peter Zijlstra
2026-07-28 16:51   ` Heiko Carstens
2026-07-28 16:51     ` Heiko Carstens
2026-07-28 12:38 ` [RFC PATCH 07/13] arm64: percpu: Implement preemptible read/write ops Mark Rutland
2026-07-28 12:38   ` Mark Rutland
2026-07-28 21:47   ` David Laight
2026-07-28 21:47     ` David Laight
2026-07-28 22:05   ` David Laight
2026-07-28 22:05     ` David Laight
2026-07-28 12:38 ` [RFC PATCH 08/13] arm64: percpu: Implement preemptible void RMW ops Mark Rutland
2026-07-28 12:38   ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 09/13] arm64: percpu: Implement preemptible return " Mark Rutland
2026-07-28 12:38   ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 10/13] arm64: percpu: Implement preemptible XCHG ops Mark Rutland
2026-07-28 12:38   ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 11/13] arm64: percpu: Implement preemptible CMPXCHG ops Mark Rutland
2026-07-28 12:38   ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 12/13] arm64: percpu: Implement preemptible CMPXCHG128 ops Mark Rutland
2026-07-28 12:38   ` Mark Rutland
2026-07-28 12:38 ` [RFC PATCH 13/13] arm64: percpu: Remove _pcp_protect*() wrappers Mark Rutland
2026-07-28 12:38   ` Mark Rutland

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260728141958.GV751831@noisy.programming.kicks-ass.net \
    --to=peterz@infradead.org \
    --cc=ardb@kernel.org \
    --cc=catalin.marinas@arm.com \
    --cc=hca@linux.ibm.com \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mark.rutland@arm.com \
    --cc=maz@kernel.org \
    --cc=ruanjinjie@huawei.com \
    --cc=vladimir.murzin@arm.com \
    --cc=will@kernel.org \
    --cc=yang@os.amperecomputing.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.