Linux-ARM-Kernel Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v2] arm64/sve: Don't zero the SVE state buffer when the SVE state is live
@ 2026-09-15 10:51 Breno Leitao
  2026-10-02  9:43 ` Mark Rutland
  2026-10-02 18:58 ` Catalin Marinas
  0 siblings, 2 replies; 3+ messages in thread
From: Breno Leitao @ 2026-09-15 10:51 UTC (permalink / raw)
  To: Catalin Marinas, Will Deacon, Mark Rutland
  Cc: rmikey, kas, usama.arif, linux-arm-kernel, linux-kernel,
	kernel-team, Breno Leitao

Currently do_sve_acc() always zeroes current->thread.sve_state. This is
not necessary in the common case, and avoiding the zeroing has a
measurable impact on some benchmarks.

In the common case where the task is not preempted and its state is not
altered by a tracer, do_sve_acc() will observe that TIF_FOREIGN_FPSTATE
is clear. In such cases, only the live register values matter, and the
in-memory copy is stale regardless of whether it is saved in
FP_STATE_FPSIMD format or FP_STATE_SVE format.

It is worth skipping the zeroing because the SVE state is discarded on
syscall entry, so userspace that mixes SVE and syscalls re-traps
constantly. A fleet profile of arm64 hosts running services whose
memset() is SVE shows the memset under do_sve_acc() accounting for 29%
of the trap handling cost.

Measured on a 72-core Neoverse V2 (SVE VL 128, sve_state_size 546,
performance governor) with perf bench sched pipe pinned to one CPU, and
SVE operation on write, so that each loop also takes an SVE access trap.

	* -0.99% kernel instructions
	* -1.38% kernel cycles
	* -1.12% wall clock

Signed-off-by: Breno Leitao <leitao@debian.org>
---
Changes in v2:
- Rewrote the commit message using Mark's suggested wording: what
  matters is that the in-memory copy is stale whenever the state is
  live, not the format it was last saved in
- Link to v1:
  https://patch.msgid.link/20260914-b4-arm64-sve-acc-memset-v1-1-67866e442393@debian.org
---
 arch/arm64/kernel/fpsimd.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/arch/arm64/kernel/fpsimd.c b/arch/arm64/kernel/fpsimd.c
index e7f1682a3059b..324c9799b0511 100644
--- a/arch/arm64/kernel/fpsimd.c
+++ b/arch/arm64/kernel/fpsimd.c
@@ -1316,7 +1316,7 @@ void do_sve_acc(unsigned long esr, struct pt_regs *regs)
 		return;
 	}
 
-	sve_alloc(current, true);
+	sve_alloc(current, false);
 	if (!current->thread.sve_state) {
 		force_sig(SIGKILL);
 		return;
@@ -1341,6 +1341,7 @@ void do_sve_acc(unsigned long esr, struct pt_regs *regs)
 		sve_flush_live();
 		fpsimd_bind_task_to_cpu();
 	} else {
+		memset(current->thread.sve_state, 0, sve_state_size(current));
 		fpsimd_to_sve(current);
 		current->thread.fp_type = FP_STATE_SVE;
 		fpsimd_flush_task_state(current);

---
base-commit: f2bfbc3554ca6919484030729424b9dee2942d24
change-id: 20260911-b4-arm64-sve-acc-memset-3425ad567857

Best regards,
--  
Breno Leitao <leitao@debian.org>



^ permalink raw reply related	[flat|nested] 3+ messages in thread

* Re: [PATCH v2] arm64/sve: Don't zero the SVE state buffer when the SVE state is live
  2026-09-15 10:51 [PATCH v2] arm64/sve: Don't zero the SVE state buffer when the SVE state is live Breno Leitao
@ 2026-10-02  9:43 ` Mark Rutland
  2026-10-02 18:58 ` Catalin Marinas
  1 sibling, 0 replies; 3+ messages in thread
From: Mark Rutland @ 2026-10-02  9:43 UTC (permalink / raw)
  To: Breno Leitao
  Cc: Catalin Marinas, Will Deacon, rmikey, kas, usama.arif,
	linux-arm-kernel, linux-kernel, kernel-team

On Tue, Sep 15, 2026 at 03:51:05AM -0700, Breno Leitao wrote:
> Currently do_sve_acc() always zeroes current->thread.sve_state. This is
> not necessary in the common case, and avoiding the zeroing has a
> measurable impact on some benchmarks.
> 
> In the common case where the task is not preempted and its state is not
> altered by a tracer, do_sve_acc() will observe that TIF_FOREIGN_FPSTATE
> is clear. In such cases, only the live register values matter, and the
> in-memory copy is stale regardless of whether it is saved in
> FP_STATE_FPSIMD format or FP_STATE_SVE format.
> 
> It is worth skipping the zeroing because the SVE state is discarded on
> syscall entry, so userspace that mixes SVE and syscalls re-traps
> constantly. A fleet profile of arm64 hosts running services whose
> memset() is SVE shows the memset under do_sve_acc() accounting for 29%
> of the trap handling cost.
> 
> Measured on a 72-core Neoverse V2 (SVE VL 128, sve_state_size 546,
> performance governor) with perf bench sched pipe pinned to one CPU, and
> SVE operation on write, so that each loop also takes an SVE access trap.
> 
> 	* -0.99% kernel instructions
> 	* -1.38% kernel cycles
> 	* -1.12% wall clock
> 
> Signed-off-by: Breno Leitao <leitao@debian.org>

Avoiding the memset in the fast path makes sense to me, and I believe
this is sound, so:

Acked-by: Mark Rutlame <mark.rutland@arm.com>

Mark.

> ---
> Changes in v2:
> - Rewrote the commit message using Mark's suggested wording: what
>   matters is that the in-memory copy is stale whenever the state is
>   live, not the format it was last saved in
> - Link to v1:
>   https://patch.msgid.link/20260914-b4-arm64-sve-acc-memset-v1-1-67866e442393@debian.org
> ---
>  arch/arm64/kernel/fpsimd.c | 3 ++-
>  1 file changed, 2 insertions(+), 1 deletion(-)
> 
> diff --git a/arch/arm64/kernel/fpsimd.c b/arch/arm64/kernel/fpsimd.c
> index e7f1682a3059b..324c9799b0511 100644
> --- a/arch/arm64/kernel/fpsimd.c
> +++ b/arch/arm64/kernel/fpsimd.c
> @@ -1316,7 +1316,7 @@ void do_sve_acc(unsigned long esr, struct pt_regs *regs)
>  		return;
>  	}
>  
> -	sve_alloc(current, true);
> +	sve_alloc(current, false);
>  	if (!current->thread.sve_state) {
>  		force_sig(SIGKILL);
>  		return;
> @@ -1341,6 +1341,7 @@ void do_sve_acc(unsigned long esr, struct pt_regs *regs)
>  		sve_flush_live();
>  		fpsimd_bind_task_to_cpu();
>  	} else {
> +		memset(current->thread.sve_state, 0, sve_state_size(current));
>  		fpsimd_to_sve(current);
>  		current->thread.fp_type = FP_STATE_SVE;
>  		fpsimd_flush_task_state(current);
> 
> ---
> base-commit: f2bfbc3554ca6919484030729424b9dee2942d24
> change-id: 20260911-b4-arm64-sve-acc-memset-3425ad567857
> 
> Best regards,
> --  
> Breno Leitao <leitao@debian.org>
> 


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH v2] arm64/sve: Don't zero the SVE state buffer when the SVE state is live
  2026-09-15 10:51 [PATCH v2] arm64/sve: Don't zero the SVE state buffer when the SVE state is live Breno Leitao
  2026-10-02  9:43 ` Mark Rutland
@ 2026-10-02 18:58 ` Catalin Marinas
  1 sibling, 0 replies; 3+ messages in thread
From: Catalin Marinas @ 2026-10-02 18:58 UTC (permalink / raw)
  To: Will Deacon, Mark Rutland, Breno Leitao
  Cc: rmikey, kas, usama.arif, linux-arm-kernel, linux-kernel,
	kernel-team

On Tue, 15 Sep 2026 03:51:05 -0700, Breno Leitao wrote:
> Currently do_sve_acc() always zeroes current->thread.sve_state. This is
> not necessary in the common case, and avoiding the zeroing has a
> measurable impact on some benchmarks.
> 
> In the common case where the task is not preempted and its state is not
> altered by a tracer, do_sve_acc() will observe that TIF_FOREIGN_FPSTATE
> is clear. In such cases, only the live register values matter, and the
> in-memory copy is stale regardless of whether it is saved in
> FP_STATE_FPSIMD format or FP_STATE_SVE format.
> 
> [...]

Applied to arm64 (for-next/misc), thanks!

[1/1] arm64/sve: Don't zero the SVE state buffer when the SVE state is live
      https://git.kernel.org/arm64/c/d6ae12c900f2


^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-10-02 18:58 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-15 10:51 [PATCH v2] arm64/sve: Don't zero the SVE state buffer when the SVE state is live Breno Leitao
2026-10-02  9:43 ` Mark Rutland
2026-10-02 18:58 ` Catalin Marinas

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox