* [PATCH v2] arm64/sve: Don't zero the SVE state buffer when the SVE state is live
@ 2026-09-15 10:51 Breno Leitao
2026-10-02 9:43 ` Mark Rutland
2026-10-02 18:58 ` Catalin Marinas
0 siblings, 2 replies; 3+ messages in thread
From: Breno Leitao @ 2026-09-15 10:51 UTC (permalink / raw)
To: Catalin Marinas, Will Deacon, Mark Rutland
Cc: rmikey, kas, usama.arif, linux-arm-kernel, linux-kernel,
kernel-team, Breno Leitao
Currently do_sve_acc() always zeroes current->thread.sve_state. This is
not necessary in the common case, and avoiding the zeroing has a
measurable impact on some benchmarks.
In the common case where the task is not preempted and its state is not
altered by a tracer, do_sve_acc() will observe that TIF_FOREIGN_FPSTATE
is clear. In such cases, only the live register values matter, and the
in-memory copy is stale regardless of whether it is saved in
FP_STATE_FPSIMD format or FP_STATE_SVE format.
It is worth skipping the zeroing because the SVE state is discarded on
syscall entry, so userspace that mixes SVE and syscalls re-traps
constantly. A fleet profile of arm64 hosts running services whose
memset() is SVE shows the memset under do_sve_acc() accounting for 29%
of the trap handling cost.
Measured on a 72-core Neoverse V2 (SVE VL 128, sve_state_size 546,
performance governor) with perf bench sched pipe pinned to one CPU, and
SVE operation on write, so that each loop also takes an SVE access trap.
* -0.99% kernel instructions
* -1.38% kernel cycles
* -1.12% wall clock
Signed-off-by: Breno Leitao <leitao@debian.org>
---
Changes in v2:
- Rewrote the commit message using Mark's suggested wording: what
matters is that the in-memory copy is stale whenever the state is
live, not the format it was last saved in
- Link to v1:
https://patch.msgid.link/20260914-b4-arm64-sve-acc-memset-v1-1-67866e442393@debian.org
---
arch/arm64/kernel/fpsimd.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/arch/arm64/kernel/fpsimd.c b/arch/arm64/kernel/fpsimd.c
index e7f1682a3059b..324c9799b0511 100644
--- a/arch/arm64/kernel/fpsimd.c
+++ b/arch/arm64/kernel/fpsimd.c
@@ -1316,7 +1316,7 @@ void do_sve_acc(unsigned long esr, struct pt_regs *regs)
return;
}
- sve_alloc(current, true);
+ sve_alloc(current, false);
if (!current->thread.sve_state) {
force_sig(SIGKILL);
return;
@@ -1341,6 +1341,7 @@ void do_sve_acc(unsigned long esr, struct pt_regs *regs)
sve_flush_live();
fpsimd_bind_task_to_cpu();
} else {
+ memset(current->thread.sve_state, 0, sve_state_size(current));
fpsimd_to_sve(current);
current->thread.fp_type = FP_STATE_SVE;
fpsimd_flush_task_state(current);
---
base-commit: f2bfbc3554ca6919484030729424b9dee2942d24
change-id: 20260911-b4-arm64-sve-acc-memset-3425ad567857
Best regards,
--
Breno Leitao <leitao@debian.org>
^ permalink raw reply related [flat|nested] 3+ messages in thread
* Re: [PATCH v2] arm64/sve: Don't zero the SVE state buffer when the SVE state is live
2026-09-15 10:51 [PATCH v2] arm64/sve: Don't zero the SVE state buffer when the SVE state is live Breno Leitao
@ 2026-10-02 9:43 ` Mark Rutland
2026-10-02 18:58 ` Catalin Marinas
1 sibling, 0 replies; 3+ messages in thread
From: Mark Rutland @ 2026-10-02 9:43 UTC (permalink / raw)
To: Breno Leitao
Cc: Catalin Marinas, Will Deacon, rmikey, kas, usama.arif,
linux-arm-kernel, linux-kernel, kernel-team
On Tue, Sep 15, 2026 at 03:51:05AM -0700, Breno Leitao wrote:
> Currently do_sve_acc() always zeroes current->thread.sve_state. This is
> not necessary in the common case, and avoiding the zeroing has a
> measurable impact on some benchmarks.
>
> In the common case where the task is not preempted and its state is not
> altered by a tracer, do_sve_acc() will observe that TIF_FOREIGN_FPSTATE
> is clear. In such cases, only the live register values matter, and the
> in-memory copy is stale regardless of whether it is saved in
> FP_STATE_FPSIMD format or FP_STATE_SVE format.
>
> It is worth skipping the zeroing because the SVE state is discarded on
> syscall entry, so userspace that mixes SVE and syscalls re-traps
> constantly. A fleet profile of arm64 hosts running services whose
> memset() is SVE shows the memset under do_sve_acc() accounting for 29%
> of the trap handling cost.
>
> Measured on a 72-core Neoverse V2 (SVE VL 128, sve_state_size 546,
> performance governor) with perf bench sched pipe pinned to one CPU, and
> SVE operation on write, so that each loop also takes an SVE access trap.
>
> * -0.99% kernel instructions
> * -1.38% kernel cycles
> * -1.12% wall clock
>
> Signed-off-by: Breno Leitao <leitao@debian.org>
Avoiding the memset in the fast path makes sense to me, and I believe
this is sound, so:
Acked-by: Mark Rutlame <mark.rutland@arm.com>
Mark.
> ---
> Changes in v2:
> - Rewrote the commit message using Mark's suggested wording: what
> matters is that the in-memory copy is stale whenever the state is
> live, not the format it was last saved in
> - Link to v1:
> https://patch.msgid.link/20260914-b4-arm64-sve-acc-memset-v1-1-67866e442393@debian.org
> ---
> arch/arm64/kernel/fpsimd.c | 3 ++-
> 1 file changed, 2 insertions(+), 1 deletion(-)
>
> diff --git a/arch/arm64/kernel/fpsimd.c b/arch/arm64/kernel/fpsimd.c
> index e7f1682a3059b..324c9799b0511 100644
> --- a/arch/arm64/kernel/fpsimd.c
> +++ b/arch/arm64/kernel/fpsimd.c
> @@ -1316,7 +1316,7 @@ void do_sve_acc(unsigned long esr, struct pt_regs *regs)
> return;
> }
>
> - sve_alloc(current, true);
> + sve_alloc(current, false);
> if (!current->thread.sve_state) {
> force_sig(SIGKILL);
> return;
> @@ -1341,6 +1341,7 @@ void do_sve_acc(unsigned long esr, struct pt_regs *regs)
> sve_flush_live();
> fpsimd_bind_task_to_cpu();
> } else {
> + memset(current->thread.sve_state, 0, sve_state_size(current));
> fpsimd_to_sve(current);
> current->thread.fp_type = FP_STATE_SVE;
> fpsimd_flush_task_state(current);
>
> ---
> base-commit: f2bfbc3554ca6919484030729424b9dee2942d24
> change-id: 20260911-b4-arm64-sve-acc-memset-3425ad567857
>
> Best regards,
> --
> Breno Leitao <leitao@debian.org>
>
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH v2] arm64/sve: Don't zero the SVE state buffer when the SVE state is live
2026-09-15 10:51 [PATCH v2] arm64/sve: Don't zero the SVE state buffer when the SVE state is live Breno Leitao
2026-10-02 9:43 ` Mark Rutland
@ 2026-10-02 18:58 ` Catalin Marinas
1 sibling, 0 replies; 3+ messages in thread
From: Catalin Marinas @ 2026-10-02 18:58 UTC (permalink / raw)
To: Will Deacon, Mark Rutland, Breno Leitao
Cc: rmikey, kas, usama.arif, linux-arm-kernel, linux-kernel,
kernel-team
On Tue, 15 Sep 2026 03:51:05 -0700, Breno Leitao wrote:
> Currently do_sve_acc() always zeroes current->thread.sve_state. This is
> not necessary in the common case, and avoiding the zeroing has a
> measurable impact on some benchmarks.
>
> In the common case where the task is not preempted and its state is not
> altered by a tracer, do_sve_acc() will observe that TIF_FOREIGN_FPSTATE
> is clear. In such cases, only the live register values matter, and the
> in-memory copy is stale regardless of whether it is saved in
> FP_STATE_FPSIMD format or FP_STATE_SVE format.
>
> [...]
Applied to arm64 (for-next/misc), thanks!
[1/1] arm64/sve: Don't zero the SVE state buffer when the SVE state is live
https://git.kernel.org/arm64/c/d6ae12c900f2
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-02 18:58 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-15 10:51 [PATCH v2] arm64/sve: Don't zero the SVE state buffer when the SVE state is live Breno Leitao
2026-10-02 9:43 ` Mark Rutland
2026-10-02 18:58 ` Catalin Marinas
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox