From: Dapeng Mi <dapeng1.mi@linux.intel.com>
To: Peter Zijlstra <peterz@infradead.org>,
Ingo Molnar <mingo@redhat.com>,
Arnaldo Carvalho de Melo <acme@kernel.org>,
Namhyung Kim <namhyung@kernel.org>,
Thomas Gleixner <tglx@linutronix.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
Ian Rogers <irogers@google.com>,
Adrian Hunter <adrian.hunter@intel.com>,
Jiri Olsa <jolsa@kernel.org>,
Alexander Shishkin <alexander.shishkin@linux.intel.com>,
Andi Kleen <ak@linux.intel.com>,
Eranian Stephane <eranian@google.com>
Cc: Mark Rutland <mark.rutland@arm.com>,
broonie@kernel.org, Ravi Bangoria <ravi.bangoria@amd.com>,
linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org,
Zide Chen <zide.chen@intel.com>,
Falcon Thomas <thomas.falcon@intel.com>,
Dapeng Mi <dapeng1.mi@intel.com>,
Xudong Hao <xudong.hao@intel.com>,
Dapeng Mi <dapeng1.mi@linux.intel.com>
Subject: [Patch v10 07/23] x86/fpu: Add update_fpu_state_and_flag() helper
Date: Tue, 21 Jul 2026 14:24:50 +0800 [thread overview]
Message-ID: <20260721062506.3745816-8-dapeng1.mi@linux.intel.com> (raw)
In-Reply-To: <20260721062506.3745816-1-dapeng1.mi@linux.intel.com>
Add update_fpu_state_and_flag() as suggested by Peter and Dave.
The helper saves user FPU state and then sets TIF_NEED_FPU_LOAD,
ensuring the task FPU state is saved whenever the flag is set.
Subsequent patches will use this guarantee in NMI context by checking
TIF_NEED_FPU_LOAD before retrieving user FPU state from the saved
task FPU state.
Also add barrier() in the host/guest FPU state switch path and move
clearing of fpu->__task_fpstate after switching back to host state. So
fpu->__task_fpstate is always observed as host FPU state when non-NULL.
Link: https://lore.kernel.org/all/20251204154721.GB2619703@noisy.programming.kicks-ass.net/
Signed-off-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
arch/x86/include/asm/fpu/sched.h | 5 +++--
arch/x86/kernel/fpu/core.c | 33 ++++++++++++++++++++++++++------
2 files changed, 30 insertions(+), 8 deletions(-)
diff --git a/arch/x86/include/asm/fpu/sched.h b/arch/x86/include/asm/fpu/sched.h
index 89004f4ca208..dcb2fa5f06d6 100644
--- a/arch/x86/include/asm/fpu/sched.h
+++ b/arch/x86/include/asm/fpu/sched.h
@@ -10,6 +10,8 @@
#include <asm/trace/fpu.h>
extern void save_fpregs_to_fpstate(struct fpu *fpu);
+extern void update_fpu_state_and_flag(struct fpu *fpu,
+ struct task_struct *task);
extern void fpu__drop(struct task_struct *tsk);
extern int fpu_clone(struct task_struct *dst, u64 clone_flags, bool minimal,
unsigned long shstk_addr);
@@ -36,8 +38,7 @@ static inline void switch_fpu(struct task_struct *old, int cpu)
!(old->flags & (PF_KTHREAD | PF_USER_WORKER))) {
struct fpu *old_fpu = x86_task_fpu(old);
- set_tsk_thread_flag(old, TIF_NEED_FPU_LOAD);
- save_fpregs_to_fpstate(old_fpu);
+ update_fpu_state_and_flag(old_fpu, old);
/*
* The save operation preserved register state, so the
* fpu_fpregs_owner_ctx is still @old_fpu. Store the
diff --git a/arch/x86/kernel/fpu/core.c b/arch/x86/kernel/fpu/core.c
index 584fb9913be4..9e029b2f8937 100644
--- a/arch/x86/kernel/fpu/core.c
+++ b/arch/x86/kernel/fpu/core.c
@@ -213,6 +213,19 @@ void restore_fpregs_from_fpstate(struct fpstate *fpstate, u64 mask)
}
}
+/*
+ * Save the FPU register state in fpu->fpstate->regs and set
+ * TIF_NEED_FPU_LOAD subsequently.
+ *
+ * Must be called with fpregs_lock() held, ensuring flag
+ * TIF_NEED_FPU_LOAD is set last.
+ */
+void update_fpu_state_and_flag(struct fpu *fpu, struct task_struct *task)
+{
+ save_fpregs_to_fpstate(fpu);
+ set_tsk_thread_flag(task, TIF_NEED_FPU_LOAD);
+}
+
void fpu_reset_from_exception_fixup(void)
{
restore_fpregs_from_fpstate(&init_fpstate, XFEATURE_MASK_FPSTATE);
@@ -383,13 +396,13 @@ int fpu_swap_kvm_fpstate(struct fpu_guest *guest_fpu, bool enter_guest)
/* Swap fpstate */
if (enter_guest) {
- fpu->__task_fpstate = cur_fps;
+ WRITE_ONCE(fpu->__task_fpstate, cur_fps);
+ barrier();
fpu->fpstate = guest_fps;
guest_fps->in_use = true;
} else {
guest_fps->in_use = false;
fpu->fpstate = fpu->__task_fpstate;
- fpu->__task_fpstate = NULL;
}
cur_fps = fpu->fpstate;
@@ -406,6 +419,16 @@ int fpu_swap_kvm_fpstate(struct fpu_guest *guest_fpu, bool enter_guest)
xfd_update_state(cur_fps);
}
+ /*
+ * Clear fpu->__task_fpstate after switching back to host state.
+ * A non-NULL __task_fpstate means guest state is still resident in
+ * hardware; reset it only once host state has been restored.
+ */
+ if (!enter_guest) {
+ barrier();
+ WRITE_ONCE(fpu->__task_fpstate, NULL);
+ }
+
fpregs_mark_activate();
fpregs_unlock();
return 0;
@@ -481,10 +504,8 @@ void kernel_fpu_begin_mask(unsigned int kfpu_mask)
this_cpu_write(kernel_fpu_allowed, false);
if (!(current->flags & (PF_KTHREAD | PF_USER_WORKER)) &&
- !test_thread_flag(TIF_NEED_FPU_LOAD)) {
- set_thread_flag(TIF_NEED_FPU_LOAD);
- save_fpregs_to_fpstate(x86_task_fpu(current));
- }
+ !test_thread_flag(TIF_NEED_FPU_LOAD))
+ update_fpu_state_and_flag(x86_task_fpu(current), current);
__cpu_invalidate_fpregs_state();
/* Put sane initial values into the control registers. */
--
2.34.1
next prev parent reply other threads:[~2026-07-21 6:32 UTC|newest]
Thread overview: 26+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-21 6:24 [Patch v10 00/23] Support SIMD/eGPRs/SSP registers sampling for perf Dapeng Mi
2026-07-21 6:24 ` [Patch v10 01/23] perf/x86: Move hybrid PMU initialization before x86_pmu_starting_cpu() Dapeng Mi
2026-07-21 6:24 ` [Patch v10 02/23] perf/x86/intel: Enable large PEBS sampling for XMMs Dapeng Mi
2026-07-21 6:24 ` [Patch v10 03/23] perf/x86/intel: Convert x86_perf_regs to per-cpu variables Dapeng Mi
2026-07-21 6:24 ` [Patch v10 04/23] perf: Eliminate duplicate arch-specific function definitions Dapeng Mi
2026-07-21 6:24 ` [Patch v10 05/23] perf/x86: Use x86_perf_regs in NMI handlers Dapeng Mi
2026-07-21 6:24 ` [Patch v10 06/23] x86/fpu/xstate: Add xsaves_nmi() helper Dapeng Mi
2026-07-21 6:24 ` Dapeng Mi [this message]
2026-07-21 6:24 ` [Patch v10 08/23] perf: Move and enhance has_extended_regs() for arch-specific use Dapeng Mi
2026-07-21 6:24 ` [Patch v10 09/23] perf/x86/intel: Centralize PERF_PMU_CAP_EXTENDED_REGS updates Dapeng Mi
2026-07-21 6:24 ` [Patch v10 10/23] perf/x86: Enable XMM register sampling for non-PEBS events Dapeng Mi
2026-07-21 6:24 ` [Patch v10 11/23] perf/x86: Enable XMM register sampling for REGS_USER case Dapeng Mi
2026-07-21 6:24 ` [Patch v10 12/23] perf: Add sampling support for SIMD registers Dapeng Mi
2026-07-21 6:24 ` [Patch v10 13/23] perf/x86: Support XMM sampling using sample_simd_vec_reg_* fields Dapeng Mi
2026-07-21 6:24 ` [Patch v10 14/23] perf/x86: Support YMM " Dapeng Mi
2026-07-21 6:24 ` [Patch v10 15/23] perf/x86: Support ZMM " Dapeng Mi
2026-07-21 6:24 ` [Patch v10 16/23] perf/x86: Support OPMASK sampling using sample_simd_pred_reg_* fields Dapeng Mi
2026-07-21 6:25 ` [Patch v10 17/23] perf: Enhance perf_reg_validate() with simd_enabled argument Dapeng Mi
2026-07-21 6:25 ` [Patch v10 18/23] perf/x86: Support eGPRs sampling using sample_regs_* fields Dapeng Mi
2026-07-21 6:25 ` [Patch v10 19/23] perf/x86: Support SSP " Dapeng Mi
2026-07-21 6:25 ` [Patch v10 20/23] perf/x86/intel: Support arch-PEBS based SIMD/eGPRs sampling Dapeng Mi
2026-07-21 6:25 ` [Patch v10 21/23] perf/x86/intel: Advertise PERF_PMU_CAP_SIMD_REGS capability Dapeng Mi
2026-07-21 6:25 ` [Patch v10 22/23] perf/x86: Activate back-to-back NMI detection for arch-PEBS induced NMIs Dapeng Mi
2026-07-21 6:25 ` [Patch v10 23/23] perf/x86/intel: Add sanity check for PEBS record/fragment size Dapeng Mi
2026-07-21 17:42 ` [Patch v10 00/23] Support SIMD/eGPRs/SSP registers sampling for perf Ian Rogers
2026-07-22 0:17 ` Mi, Dapeng
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260721062506.3745816-8-dapeng1.mi@linux.intel.com \
--to=dapeng1.mi@linux.intel.com \
--cc=acme@kernel.org \
--cc=adrian.hunter@intel.com \
--cc=ak@linux.intel.com \
--cc=alexander.shishkin@linux.intel.com \
--cc=broonie@kernel.org \
--cc=dapeng1.mi@intel.com \
--cc=dave.hansen@linux.intel.com \
--cc=eranian@google.com \
--cc=irogers@google.com \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=mingo@redhat.com \
--cc=namhyung@kernel.org \
--cc=peterz@infradead.org \
--cc=ravi.bangoria@amd.com \
--cc=tglx@linutronix.de \
--cc=thomas.falcon@intel.com \
--cc=xudong.hao@intel.com \
--cc=zide.chen@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.