All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested)
@ 2026-07-28 17:42 Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 01/13] KVM: selftests: Use __stringify() instead of custom XSTR() macros Yosry Ahmed
                   ` (12 more replies)
  0 siblings, 13 replies; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

Add a stress test for save+restore while the guest is triggering and
handling #PFs, in both L1 and L2. The goal was to create a generic
selftest that would catch bugs like the one fixed by commit 5c247d08bc81
("KVM: nSVM: Use vcpu->arch.cr2 when updating vmcb12 on nested
#VMEXIT"), instead of relying on high-level testing (e.g. building GCC
in L2) to catch it.

The test tries to be as generic as possible by triggering #PFs in a
guest and installing a proper #PF handler, while the host is
continuously doing save+restore cycles. Exiting to userspace is randomly
triggered by a second thread that constantly signals the vCPU thread.

Patches (1-6) are prep patches, fixing GPR and RFLAGS switching for nSVM
and generalizing GPR switching to cover nVMX, which is needed for the
test to run properly with nVMX. Patch 6 removes
HORRIFIC_L2_UCALL_CLOBBER_HACK, as it is no longer needed. While this
series does not have the "complete" fix added by commit 6783ca4105a7
("KVM: selftests: Add a shameful hack to preserve/clobber GPRs across
ucall"), it's a good step in the right direction.

Patches (7-9) add minor infrastructure improvements (assertion failure
formatting and exposing MMU masks to the guest).

Patches (10-13) add the actual test. The test is first introduced as a
simple (read: dummy) stress test that just explicitly syncs to userspace
after each #PF handling to do save+restore, then gradually evolves to
add the random signaling, nested support, and triggering L2->L1 exits.
After the last patch, the test reliably reproduces the CR2 bug.

v4 -> v5:
- Only set KVM_CAP_EXCEPTION_PAYLOAD when running nested [Sashiko].
- Do not intercept #PF by default for nVMX tests [Sean].
- Nits and cleanups [Sean].

v3 -> v4:
- Switch guest_regs back to a struct and provide struct offsets as
  ASM constraints similar to kvm-unit-tests [Sean].
- Expose PTE masks to guests as part of a guest MMU [Sean].
- Use __stringify() from <linux/stringify.h> instead of custom STR() and
  XSTR() macros [Sean].
- Handle guest RFLAGS save/restore as part of guest_regs.
- Run both L0 and L2 tests sequentially by default if nested is supported,
  instead of requiring a '-n' command line argument [Sean].
- Minor fixups and cleanups [Sean].

v2 -> v3:
- Rebased on top of L2 stack rework in selftests.
- Fix GPR array size (i.e. NR_GUEST_REGS) [Sashiko].
- Handle evmcs_vmlaunch() and evmcs_vmresume() [Sashiko].
- Fix off-by-one assertion error in intermediate patches.
- Increase inter-signal delay from 100us to 1msec.

v1 -> v2:
- Switch guest_regs to an array, which simplifies the offsets
  calculation and forgoes the dependency on using OFFSET() or defining
  the struct offsets for assembly otherwise.
- Move page table mapping to the test (instead of a generic helper), as
  the helper mistakenly tried to map the entire memslot, not just page
  tables.
- Do not use identity mappings for page tables as it collisions with
  GVAs used for ELF in some cases.
- Simplify page table walking by using loops.
- Make sure the signals are ignored before creating the signaling
  thread [Sashiko]
- Assert that the guest actually ran and had page faults [Sashiko]
- Add a patch to fix RAX and RFLAGS offsets in run_guest() [Sashiko]
- Initialize exception_has_payload when injecting a #UD [Sashiko]
- Only check KVM_STATE_NESTED_GUEST_MODE when running in nested mode
  [Sashiko]

v1: https://lore.kernel.org/all/20260518202514.2037078-1-yosry@kernel.org/
v2: https://lore.kernel.org/kvm/20260604203546.365658-1-yosry@kernel.org/
v3: https://lore.kernel.org/kvm/20260629183746.699840-1-yosry@kernel.org/
v4: https://lore.kernel.org/kvm/20260727235228.1007324-1-yosry@kernel.org/

Yosry Ahmed (13):
  KVM: selftests: Use __stringify() instead of custom XSTR() macros
  KVM: selftests: Fix RAX and RFLAGS VMCB offsets when running L2
  KVM: selftests: Rework GPR registers switching for SVM (and fix
    offsets)
  KVM: selftests: Handle rflags save/restore for SVM in guest_regs
  KVM: selftests: Reuse GPR switching logic for nVMX
  KVM: selftests: Drop HORRIFIC_L2_UCALL_CLOBBER_HACK
  KVM: selftests: Add a blank line before logging assertion failures
  KVM: selftests: Expose PTE masks to guests as part of an MMU
  KVM: selftests: Do not intercept #PF by default in nVMX tests
  KVM: selftests: Add basic stress test for save+restore and #PF
    handling
  KVM: selftests: Trigger save+restore randomly in the #PF stress test
  KVM: selftests: Support running stress save+restore and #PF test in L2
  KVM: selftests: Trigger L2->L1 exits stress save+restore and #PF test

 tools/testing/selftests/kvm/Makefile.kvm      |   1 +
 .../testing/selftests/kvm/include/test_util.h |   1 +
 .../testing/selftests/kvm/include/x86/evmcs.h |  46 +--
 .../selftests/kvm/include/x86/processor.h     |  52 +++-
 tools/testing/selftests/kvm/include/x86/vmx.h |  69 ++---
 tools/testing/selftests/kvm/lib/assert.c      |   2 +-
 .../testing/selftests/kvm/lib/x86/processor.c |  14 +
 tools/testing/selftests/kvm/lib/x86/svm.c     |  62 ++--
 tools/testing/selftests/kvm/lib/x86/ucall.c   |  32 +-
 tools/testing/selftests/kvm/lib/x86/vmx.c     |   2 +-
 .../kvm/x86/evmcs_smm_controls_test.c         |   5 +-
 .../selftests/kvm/x86/fix_hypercall_test.c    |   1 -
 .../kvm/x86/save_restore_pf_stress_test.c     | 288 ++++++++++++++++++
 tools/testing/selftests/kvm/x86/smm_test.c    |   5 +-
 14 files changed, 439 insertions(+), 141 deletions(-)
 create mode 100644 tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c


base-commit: 271255273d5ff348fe29d89fe4712b2f7f7907c3
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [PATCH v5 01/13] KVM: selftests: Use __stringify() instead of custom XSTR() macros
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 02/13] KVM: selftests: Fix RAX and RFLAGS VMCB offsets when running L2 Yosry Ahmed
                   ` (11 subsequent siblings)
  12 siblings, 0 replies; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

Drop the custom defined XSTR() macros in KVM selftests and use
__stringify() instead. Include stringify.h in test_util.h to make it
available for all tests instead of including it in all the tests that
need it, as more tests will start using it.

No functional change intended.

Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 tools/testing/selftests/kvm/include/test_util.h           | 1 +
 tools/testing/selftests/kvm/x86/evmcs_smm_controls_test.c | 5 +----
 tools/testing/selftests/kvm/x86/fix_hypercall_test.c      | 1 -
 tools/testing/selftests/kvm/x86/smm_test.c                | 5 +----
 4 files changed, 3 insertions(+), 9 deletions(-)

diff --git a/tools/testing/selftests/kvm/include/test_util.h b/tools/testing/selftests/kvm/include/test_util.h
index d64c8a228207a..e8356ee54d7b9 100644
--- a/tools/testing/selftests/kvm/include/test_util.h
+++ b/tools/testing/selftests/kvm/include/test_util.h
@@ -23,6 +23,7 @@
 
 #include <linux/mman.h>
 #include <linux/types.h>
+#include <linux/stringify.h>
 
 #define msecs_to_usecs(msec)    ((msec) * 1000ULL)
 
diff --git a/tools/testing/selftests/kvm/x86/evmcs_smm_controls_test.c b/tools/testing/selftests/kvm/x86/evmcs_smm_controls_test.c
index 77ce87c41a868..aa7f3b405fd32 100644
--- a/tools/testing/selftests/kvm/x86/evmcs_smm_controls_test.c
+++ b/tools/testing/selftests/kvm/x86/evmcs_smm_controls_test.c
@@ -22,9 +22,6 @@
 
 #define SYNC_PORT	0xe
 
-#define STR(x) #x
-#define XSTR(s) STR(s)
-
 /*
  * SMI handler: runs in real-address mode.
  * Reports SMRAM_STAGE via port IO, then does RSM.
@@ -37,7 +34,7 @@ static u8 smi_handler[] = {
 
 static inline void sync_with_host(u64 phase)
 {
-	asm volatile("in $" XSTR(SYNC_PORT) ", %%al \n"
+	asm volatile("in $" __stringify(SYNC_PORT) ", %%al \n"
 		     : "+a" (phase));
 }
 
diff --git a/tools/testing/selftests/kvm/x86/fix_hypercall_test.c b/tools/testing/selftests/kvm/x86/fix_hypercall_test.c
index 753a0e730ea8d..4931ec22768ee 100644
--- a/tools/testing/selftests/kvm/x86/fix_hypercall_test.c
+++ b/tools/testing/selftests/kvm/x86/fix_hypercall_test.c
@@ -6,7 +6,6 @@
  */
 #include <asm/kvm_para.h>
 #include <linux/kvm_para.h>
-#include <linux/stringify.h>
 #include <stdint.h>
 
 #include "kvm_test_harness.h"
diff --git a/tools/testing/selftests/kvm/x86/smm_test.c b/tools/testing/selftests/kvm/x86/smm_test.c
index e2542f4ced605..d1edafd5af755 100644
--- a/tools/testing/selftests/kvm/x86/smm_test.c
+++ b/tools/testing/selftests/kvm/x86/smm_test.c
@@ -22,9 +22,6 @@
 #define SMRAM_GPA 0x1000000
 #define SMRAM_STAGE 0xfe
 
-#define STR(x) #x
-#define XSTR(s) STR(s)
-
 #define SYNC_PORT 0xe
 #define DONE 0xff
 
@@ -42,7 +39,7 @@ u8 smi_handler[] = {
 
 static inline void sync_with_host(u64 phase)
 {
-	asm volatile("in $" XSTR(SYNC_PORT)", %%al \n"
+	asm volatile("in $" __stringify(SYNC_PORT)", %%al \n"
 		     : "+a" (phase));
 }
 
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [PATCH v5 02/13] KVM: selftests: Fix RAX and RFLAGS VMCB offsets when running L2
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 01/13] KVM: selftests: Use __stringify() instead of custom XSTR() macros Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  2026-07-28 17:55   ` sashiko-bot
  2026-07-28 17:42 ` [PATCH v5 03/13] KVM: selftests: Rework GPR registers switching for SVM (and fix offsets) Yosry Ahmed
                   ` (10 subsequent siblings)
  12 siblings, 1 reply; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson
  Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed, Sashiko

The offsets used (0x170 and 0x1f8) are offsets within vmcb_save_area,
not vmcb. The correct offsets should include the base of vmcb_save_area
within vmcb (which is 0x400 -- so 0x570 and 0x5f8).

Instead of just correcting the offsets, use vmcb->save.rax and
vmcb->save.rflags as parameters to the asm block and avoid hardcoding
offsets completely. While at it, also use guest_regs.rax directly
instead of assuming it's at offset 0 of guest_regs.

Note: "+m" must be used for vmcb_rax and vmcb_rflags, as caching those
fields in registers would be wrong as the underlying KVM will update
them in memory.

The same problem was recently fixed (differently) for kvm-unit-tests
[1].

[1]https://lore.kernel.org/all/20260521092311.86030-1-pbonzini@redhat.com/

Reported-by: Sashiko <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260518202514.2037078-1-yosry%40kernel.org?part=1
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 tools/testing/selftests/kvm/lib/x86/svm.c | 19 +++++++++++--------
 1 file changed, 11 insertions(+), 8 deletions(-)

diff --git a/tools/testing/selftests/kvm/lib/x86/svm.c b/tools/testing/selftests/kvm/lib/x86/svm.c
index 1445b890986fd..766d15f1d534a 100644
--- a/tools/testing/selftests/kvm/lib/x86/svm.c
+++ b/tools/testing/selftests/kvm/lib/x86/svm.c
@@ -164,19 +164,22 @@ void run_guest(struct vmcb *vmcb, u64 vmcb_gpa)
 {
 	asm volatile (
 		"vmload %[vmcb_gpa]\n\t"
-		"mov rflags, %%r15\n\t"	// rflags
-		"mov %%r15, 0x170(%[vmcb])\n\t"
-		"mov guest_regs, %%r15\n\t"	// rax
-		"mov %%r15, 0x1f8(%[vmcb])\n\t"
+		"mov rflags, %%r15\n\t"
+		"mov %%r15, %[vmcb_rflags]\n\t"
+		"mov %[guest_regs_rax], %%r15\n\t"
+		"mov %%r15, %[vmcb_rax]\n\t"
 		LOAD_GPR_C
 		"vmrun %[vmcb_gpa]\n\t"
 		SAVE_GPR_C
-		"mov 0x170(%[vmcb]), %%r15\n\t"	// rflags
+		"mov %[vmcb_rflags], %%r15\n\t"
 		"mov %%r15, rflags\n\t"
-		"mov 0x1f8(%[vmcb]), %%r15\n\t"	// rax
-		"mov %%r15, guest_regs\n\t"
+		"mov %[vmcb_rax], %%r15\n\t"	// rax
+		"mov %%r15, %[guest_regs_rax]\n\t"
 		"vmsave %[vmcb_gpa]\n\t"
-		: : [vmcb] "r" (vmcb), [vmcb_gpa] "a" (vmcb_gpa)
+		: [vmcb_rflags] "+m" (vmcb->save.rflags),
+		  [vmcb_rax] "+m" (vmcb->save.rax),
+		  [guest_regs_rax] "+rm" (guest_regs.rax)
+		: [vmcb_gpa] "a" (vmcb_gpa)
 		: "r15", "memory");
 }
 
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [PATCH v5 03/13] KVM: selftests: Rework GPR registers switching for SVM (and fix offsets)
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 01/13] KVM: selftests: Use __stringify() instead of custom XSTR() macros Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 02/13] KVM: selftests: Fix RAX and RFLAGS VMCB offsets when running L2 Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 04/13] KVM: selftests: Handle rflags save/restore for SVM in guest_regs Yosry Ahmed
                   ` (9 subsequent siblings)
  12 siblings, 0 replies; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

The assembly code defined by SAVE_GPR_C uses the wrong offsets for some
registers in guest_regs. For example, the offset of RCX should be 0x08
not 0x10. Also, the last offset in the struct (R15) is 0x78, not 0x80,
so the code actually saves and restore beyond the end of gpr64_regs.

Eliminate hardcoded offsets by dynamically generating offsets using
offset_of() and using macros to pass the offsets to assembly as asm
constraints.

To avoid register conflicts in inline assembly (since almost all GPRs are
context-switched), access guest_regs via absolute symbol addressing
(guest_regs + offset) rather than using a base register which could
get overwritten mid-assembly.

While at it, rename SAVE_GPR_C and LOAD_GPR_C to a single macro,
SVM_SWITCH_GPRS_ASM, rename gpr64_regs to guest_regs, and expose it in
processor.h (in preparation for reusing it for VMX).

Assisted-by: Gemini:Gemini-Next
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 .../selftests/kvm/include/x86/processor.h     | 31 +++++++++++-
 .../testing/selftests/kvm/lib/x86/processor.c |  2 +
 tools/testing/selftests/kvm/lib/x86/svm.c     | 49 +++++++++----------
 3 files changed, 54 insertions(+), 28 deletions(-)

diff --git a/tools/testing/selftests/kvm/include/x86/processor.h b/tools/testing/selftests/kvm/include/x86/processor.h
index b161174ece453..d7f1acaddd4ee 100644
--- a/tools/testing/selftests/kvm/include/x86/processor.h
+++ b/tools/testing/selftests/kvm/include/x86/processor.h
@@ -398,8 +398,7 @@ static inline unsigned int x86_model(unsigned int eax)
 #define PTE_GET_PA(pte)		((pte) & PHYSICAL_PAGE_MASK)
 #define PTE_GET_PFN(pte)        (PTE_GET_PA(pte) >> PAGE_SHIFT)
 
-/* General Registers in 64-Bit Mode */
-struct gpr64_regs {
+struct guest_regs {
 	u64 rax;
 	u64 rcx;
 	u64 rdx;
@@ -418,6 +417,34 @@ struct gpr64_regs {
 	u64 r15;
 };
 
+extern struct guest_regs guest_regs;
+
+#define GUEST_REG_OFFSET(name) \
+	[off_##name] "i" (offsetof(struct guest_regs, name))
+
+#define GUEST_REGS_OFFSETS	\
+	GUEST_REG_OFFSET(rax),	\
+	GUEST_REG_OFFSET(rcx),	\
+	GUEST_REG_OFFSET(rdx),	\
+	GUEST_REG_OFFSET(rbx),	\
+	GUEST_REG_OFFSET(rsp),	\
+	GUEST_REG_OFFSET(rbp),	\
+	GUEST_REG_OFFSET(rsi),	\
+	GUEST_REG_OFFSET(rdi),	\
+	GUEST_REG_OFFSET(r8),	\
+	GUEST_REG_OFFSET(r9),	\
+	GUEST_REG_OFFSET(r10),	\
+	GUEST_REG_OFFSET(r11),	\
+	GUEST_REG_OFFSET(r12),	\
+	GUEST_REG_OFFSET(r13),	\
+	GUEST_REG_OFFSET(r14),	\
+	GUEST_REG_OFFSET(r15)
+
+#define GUEST_REG(name) "guest_regs + %c[off_" #name "]"
+
+#define GUEST_SWITCH_GPR_ASM(name) \
+	"xchg %%" #name ", " GUEST_REG(name) "\n\t"
+
 struct desc64 {
 	u16 limit0;
 	u16 base0;
diff --git a/tools/testing/selftests/kvm/lib/x86/processor.c b/tools/testing/selftests/kvm/lib/x86/processor.c
index ef56dcefe0119..1f9201590f5b3 100644
--- a/tools/testing/selftests/kvm/lib/x86/processor.c
+++ b/tools/testing/selftests/kvm/lib/x86/processor.c
@@ -29,6 +29,8 @@ bool host_cpu_is_amd_compatible;
 bool is_forced_emulation_enabled;
 u64 guest_tsc_khz;
 
+struct guest_regs guest_regs;
+
 const char *ex_str(int vector)
 {
 	switch (vector) {
diff --git a/tools/testing/selftests/kvm/lib/x86/svm.c b/tools/testing/selftests/kvm/lib/x86/svm.c
index 766d15f1d534a..7db38faded89e 100644
--- a/tools/testing/selftests/kvm/lib/x86/svm.c
+++ b/tools/testing/selftests/kvm/lib/x86/svm.c
@@ -13,7 +13,6 @@
 
 #define SEV_DEV_PATH "/dev/sev"
 
-struct gpr64_regs guest_regs;
 u64 rflags;
 
 /* Allocate memory regions for nested SVM tests.
@@ -137,27 +136,25 @@ void generic_svm_setup(struct svm_test_data *svm, void *guest_rip)
  * save/restore 64-bit general registers except rax, rip, rsp
  * which are directly handed through the VMCB guest processor state
  */
-#define SAVE_GPR_C				\
-	"xchg %%rbx, guest_regs+0x20\n\t"	\
-	"xchg %%rcx, guest_regs+0x10\n\t"	\
-	"xchg %%rdx, guest_regs+0x18\n\t"	\
-	"xchg %%rbp, guest_regs+0x30\n\t"	\
-	"xchg %%rsi, guest_regs+0x38\n\t"	\
-	"xchg %%rdi, guest_regs+0x40\n\t"	\
-	"xchg %%r8,  guest_regs+0x48\n\t"	\
-	"xchg %%r9,  guest_regs+0x50\n\t"	\
-	"xchg %%r10, guest_regs+0x58\n\t"	\
-	"xchg %%r11, guest_regs+0x60\n\t"	\
-	"xchg %%r12, guest_regs+0x68\n\t"	\
-	"xchg %%r13, guest_regs+0x70\n\t"	\
-	"xchg %%r14, guest_regs+0x78\n\t"	\
-	"xchg %%r15, guest_regs+0x80\n\t"
-
-#define LOAD_GPR_C      SAVE_GPR_C
+#define SVM_SWITCH_GPRS_ASM \
+	GUEST_SWITCH_GPR_ASM(rbx) \
+	GUEST_SWITCH_GPR_ASM(rcx) \
+	GUEST_SWITCH_GPR_ASM(rdx) \
+	GUEST_SWITCH_GPR_ASM(rbp) \
+	GUEST_SWITCH_GPR_ASM(rsi) \
+	GUEST_SWITCH_GPR_ASM(rdi) \
+	GUEST_SWITCH_GPR_ASM(r8)  \
+	GUEST_SWITCH_GPR_ASM(r9)  \
+	GUEST_SWITCH_GPR_ASM(r10) \
+	GUEST_SWITCH_GPR_ASM(r11) \
+	GUEST_SWITCH_GPR_ASM(r12) \
+	GUEST_SWITCH_GPR_ASM(r13) \
+	GUEST_SWITCH_GPR_ASM(r14) \
+	GUEST_SWITCH_GPR_ASM(r15)
 
 /*
  * selftests do not use interrupts so we dropped clgi/sti/cli/stgi
- * for now. registers involved in LOAD/SAVE_GPR_C are eventually
+ * for now. Registers involved in SVM_SWITCH_GPRS_ASM are eventually
  * unmodified so they do not need to be in the clobber list.
  */
 void run_guest(struct vmcb *vmcb, u64 vmcb_gpa)
@@ -166,20 +163,20 @@ void run_guest(struct vmcb *vmcb, u64 vmcb_gpa)
 		"vmload %[vmcb_gpa]\n\t"
 		"mov rflags, %%r15\n\t"
 		"mov %%r15, %[vmcb_rflags]\n\t"
-		"mov %[guest_regs_rax], %%r15\n\t"
+		"mov " GUEST_REG(rax) ", %%r15\n\t"
 		"mov %%r15, %[vmcb_rax]\n\t"
-		LOAD_GPR_C
+		SVM_SWITCH_GPRS_ASM
 		"vmrun %[vmcb_gpa]\n\t"
-		SAVE_GPR_C
+		SVM_SWITCH_GPRS_ASM
 		"mov %[vmcb_rflags], %%r15\n\t"
 		"mov %%r15, rflags\n\t"
 		"mov %[vmcb_rax], %%r15\n\t"	// rax
-		"mov %%r15, %[guest_regs_rax]\n\t"
+		"mov %%r15, " GUEST_REG(rax) "\n\t"
 		"vmsave %[vmcb_gpa]\n\t"
 		: [vmcb_rflags] "+m" (vmcb->save.rflags),
-		  [vmcb_rax] "+m" (vmcb->save.rax),
-		  [guest_regs_rax] "+rm" (guest_regs.rax)
-		: [vmcb_gpa] "a" (vmcb_gpa)
+		  [vmcb_rax] "+m" (vmcb->save.rax)
+		: [vmcb_gpa] "a" (vmcb_gpa),
+		  GUEST_REGS_OFFSETS
 		: "r15", "memory");
 }
 
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [PATCH v5 04/13] KVM: selftests: Handle rflags save/restore for SVM in guest_regs
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
                   ` (2 preceding siblings ...)
  2026-07-28 17:42 ` [PATCH v5 03/13] KVM: selftests: Rework GPR registers switching for SVM (and fix offsets) Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 05/13] KVM: selftests: Reuse GPR switching logic for nVMX Yosry Ahmed
                   ` (8 subsequent siblings)
  12 siblings, 0 replies; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

Instead of handling rflags separately, add it to guest_regs. No
functional change intended.

Assisted-by: Gemini:Gemini-Next
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 tools/testing/selftests/kvm/include/x86/processor.h | 4 +++-
 tools/testing/selftests/kvm/lib/x86/svm.c           | 6 ++----
 2 files changed, 5 insertions(+), 5 deletions(-)

diff --git a/tools/testing/selftests/kvm/include/x86/processor.h b/tools/testing/selftests/kvm/include/x86/processor.h
index d7f1acaddd4ee..dab64688e8eae 100644
--- a/tools/testing/selftests/kvm/include/x86/processor.h
+++ b/tools/testing/selftests/kvm/include/x86/processor.h
@@ -415,6 +415,7 @@ struct guest_regs {
 	u64 r13;
 	u64 r14;
 	u64 r15;
+	u64 rflags;
 };
 
 extern struct guest_regs guest_regs;
@@ -438,7 +439,8 @@ extern struct guest_regs guest_regs;
 	GUEST_REG_OFFSET(r12),	\
 	GUEST_REG_OFFSET(r13),	\
 	GUEST_REG_OFFSET(r14),	\
-	GUEST_REG_OFFSET(r15)
+	GUEST_REG_OFFSET(r15),	\
+	GUEST_REG_OFFSET(rflags)
 
 #define GUEST_REG(name) "guest_regs + %c[off_" #name "]"
 
diff --git a/tools/testing/selftests/kvm/lib/x86/svm.c b/tools/testing/selftests/kvm/lib/x86/svm.c
index 7db38faded89e..b05be50f075d6 100644
--- a/tools/testing/selftests/kvm/lib/x86/svm.c
+++ b/tools/testing/selftests/kvm/lib/x86/svm.c
@@ -13,8 +13,6 @@
 
 #define SEV_DEV_PATH "/dev/sev"
 
-u64 rflags;
-
 /* Allocate memory regions for nested SVM tests.
  *
  * Input Args:
@@ -161,7 +159,7 @@ void run_guest(struct vmcb *vmcb, u64 vmcb_gpa)
 {
 	asm volatile (
 		"vmload %[vmcb_gpa]\n\t"
-		"mov rflags, %%r15\n\t"
+		"mov " GUEST_REG(rflags) ", %%r15\n\t"
 		"mov %%r15, %[vmcb_rflags]\n\t"
 		"mov " GUEST_REG(rax) ", %%r15\n\t"
 		"mov %%r15, %[vmcb_rax]\n\t"
@@ -169,7 +167,7 @@ void run_guest(struct vmcb *vmcb, u64 vmcb_gpa)
 		"vmrun %[vmcb_gpa]\n\t"
 		SVM_SWITCH_GPRS_ASM
 		"mov %[vmcb_rflags], %%r15\n\t"
-		"mov %%r15, rflags\n\t"
+		"mov %%r15, " GUEST_REG(rflags) "\n\t"
 		"mov %[vmcb_rax], %%r15\n\t"	// rax
 		"mov %%r15, " GUEST_REG(rax) "\n\t"
 		"vmsave %[vmcb_gpa]\n\t"
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [PATCH v5 05/13] KVM: selftests: Reuse GPR switching logic for nVMX
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
                   ` (3 preceding siblings ...)
  2026-07-28 17:42 ` [PATCH v5 04/13] KVM: selftests: Handle rflags save/restore for SVM in guest_regs Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  2026-07-28 17:54   ` sashiko-bot
  2026-07-28 17:42 ` [PATCH v5 06/13] KVM: selftests: Drop HORRIFIC_L2_UCALL_CLOBBER_HACK Yosry Ahmed
                   ` (7 subsequent siblings)
  12 siblings, 1 reply; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

Reuse the GPR switching logic for nVMX by defining VMX_SWITCH_GPRS_ASM,
which is essentially the same as SVM_SWITCH_GPRS_ASM but also switches
RAX and doesn't switch RFLAGS, replacing the push/pop of a subset of the
registers.

The long clobber list of registers is no longer needed as registers are
saved and restored appropriately (and not clobbered by L2).

Define VMX_SWITCH_GPRS_ASM before including evmcs.h, such that it can be
used by evmcs_vmlaunch() and evmcs_vmresume().

This replaces the apparently thread-safe push/pop sequence with the
global GPR switching logic used by SVM, which isn't thread-safe at all.

However this is still an improvement because:
- The VMX logic is half-baked and prompts the UCALL clobber hack as it
  doesn't properly save/restore everything. Reusing the GPR switching
  logic used by SVM allows for dropping that hack.

- Hitting a problem due to half-baked GPR save/restore logic is arguably
  more likely than thread-safety. Evidently, adding more involved stress
  tests fails on VMX with the existing push/pop sequence. OTOH, there
  are no known failures on SVM due to lack of thread-safety fo
  save/restore. Only one test currently uses more than one vCPU with
  nested (the memstress test).

The logical next step is to move the guest_regs to be per-vCPU,
making it thread-safe for both VMX and SVM in a proper way.

Assisted-by: Gemini:gemini-3.1-pro
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 .../testing/selftests/kvm/include/x86/evmcs.h | 46 +++++--------
 tools/testing/selftests/kvm/include/x86/vmx.h | 69 +++++++++----------
 2 files changed, 49 insertions(+), 66 deletions(-)

diff --git a/tools/testing/selftests/kvm/include/x86/evmcs.h b/tools/testing/selftests/kvm/include/x86/evmcs.h
index be79bda024bf1..82a8ea6b661f0 100644
--- a/tools/testing/selftests/kvm/include/x86/evmcs.h
+++ b/tools/testing/selftests/kvm/include/x86/evmcs.h
@@ -1207,30 +1207,23 @@ static inline int evmcs_vmlaunch(void)
 
 	current_evmcs->hv_clean_fields = 0;
 
-	__asm__ __volatile__("push %%rbp;"
-			     "push %%rcx;"
-			     "push %%rdx;"
-			     "push %%rsi;"
-			     "push %%rdi;"
-			     "push $0;"
+	__asm__ __volatile__("push $0;"
 			     "mov %%rsp, (%[host_rsp]);"
 			     "lea 1f(%%rip), %%rax;"
 			     "mov %%rax, (%[host_rip]);"
+			     VMX_SWITCH_GPRS_ASM
 			     "vmlaunch;"
 			     "incq (%%rsp);"
-			     "1: pop %%rax;"
-			     "pop %%rdi;"
-			     "pop %%rsi;"
-			     "pop %%rdx;"
-			     "pop %%rcx;"
-			     "pop %%rbp;"
+			     "1: ;"
+			     VMX_SWITCH_GPRS_ASM
+			     "pop %%rax;"
 			     : [ret]"=&a"(ret)
 			     : [host_rsp]"r"
 			       ((u64)&current_evmcs->host_rsp),
 			       [host_rip]"r"
-			       ((u64)&current_evmcs->host_rip)
-			     : "memory", "cc", "rbx", "r8", "r9", "r10",
-			       "r11", "r12", "r13", "r14", "r15");
+			       ((u64)&current_evmcs->host_rip),
+			       GUEST_REGS_OFFSETS
+			     : "memory", "cc");
 	return ret;
 }
 
@@ -1246,30 +1239,23 @@ static inline int evmcs_vmresume(void)
 	/* HOST_RSP */
 	current_evmcs->hv_clean_fields &= ~HV_VMX_ENLIGHTENED_CLEAN_FIELD_HOST_POINTER;
 
-	__asm__ __volatile__("push %%rbp;"
-			     "push %%rcx;"
-			     "push %%rdx;"
-			     "push %%rsi;"
-			     "push %%rdi;"
-			     "push $0;"
+	__asm__ __volatile__("push $0;"
 			     "mov %%rsp, (%[host_rsp]);"
 			     "lea 1f(%%rip), %%rax;"
 			     "mov %%rax, (%[host_rip]);"
+			     VMX_SWITCH_GPRS_ASM
 			     "vmresume;"
 			     "incq (%%rsp);"
-			     "1: pop %%rax;"
-			     "pop %%rdi;"
-			     "pop %%rsi;"
-			     "pop %%rdx;"
-			     "pop %%rcx;"
-			     "pop %%rbp;"
+			     "1: ;"
+			     VMX_SWITCH_GPRS_ASM
+			     "pop %%rax;"
 			     : [ret]"=&a"(ret)
 			     : [host_rsp]"r"
 			       ((u64)&current_evmcs->host_rsp),
 			       [host_rip]"r"
-			       ((u64)&current_evmcs->host_rip)
-			     : "memory", "cc", "rbx", "r8", "r9", "r10",
-			       "r11", "r12", "r13", "r14", "r15");
+			       ((u64)&current_evmcs->host_rip),
+			       GUEST_REGS_OFFSETS
+			     : "memory", "cc");
 	return ret;
 }
 
diff --git a/tools/testing/selftests/kvm/include/x86/vmx.h b/tools/testing/selftests/kvm/include/x86/vmx.h
index 4bcfd60e3aecb..04f5e34dea3ae 100644
--- a/tools/testing/selftests/kvm/include/x86/vmx.h
+++ b/tools/testing/selftests/kvm/include/x86/vmx.h
@@ -290,6 +290,23 @@ struct vmx_msr_entry {
 	u64 value;
 } __attribute__ ((aligned(16)));
 
+#define VMX_SWITCH_GPRS_ASM \
+	GUEST_SWITCH_GPR_ASM(rax) \
+	GUEST_SWITCH_GPR_ASM(rbx) \
+	GUEST_SWITCH_GPR_ASM(rcx) \
+	GUEST_SWITCH_GPR_ASM(rdx) \
+	GUEST_SWITCH_GPR_ASM(rbp) \
+	GUEST_SWITCH_GPR_ASM(rsi) \
+	GUEST_SWITCH_GPR_ASM(rdi) \
+	GUEST_SWITCH_GPR_ASM(r8)  \
+	GUEST_SWITCH_GPR_ASM(r9)  \
+	GUEST_SWITCH_GPR_ASM(r10) \
+	GUEST_SWITCH_GPR_ASM(r11) \
+	GUEST_SWITCH_GPR_ASM(r12) \
+	GUEST_SWITCH_GPR_ASM(r13) \
+	GUEST_SWITCH_GPR_ASM(r14) \
+	GUEST_SWITCH_GPR_ASM(r15)
+
 #include "evmcs.h"
 
 static inline int vmxon(u64 phys)
@@ -363,9 +380,6 @@ static inline u64 vmptrstz(void)
 	return value;
 }
 
-/*
- * No guest state (e.g. GPRs) is established by this vmlaunch.
- */
 static inline int vmlaunch(void)
 {
 	int ret;
@@ -373,34 +387,24 @@ static inline int vmlaunch(void)
 	if (enable_evmcs)
 		return evmcs_vmlaunch();
 
-	__asm__ __volatile__("push %%rbp;"
-			     "push %%rcx;"
-			     "push %%rdx;"
-			     "push %%rsi;"
-			     "push %%rdi;"
-			     "push $0;"
+	__asm__ __volatile__("push $0;"
 			     "vmwrite %%rsp, %[host_rsp];"
 			     "lea 1f(%%rip), %%rax;"
 			     "vmwrite %%rax, %[host_rip];"
+			     VMX_SWITCH_GPRS_ASM
 			     "vmlaunch;"
 			     "incq (%%rsp);"
-			     "1: pop %%rax;"
-			     "pop %%rdi;"
-			     "pop %%rsi;"
-			     "pop %%rdx;"
-			     "pop %%rcx;"
-			     "pop %%rbp;"
+			     "1: ;"
+			     VMX_SWITCH_GPRS_ASM
+			     "pop %%rax;"
 			     : [ret]"=&a"(ret)
 			     : [host_rsp]"r"((u64)HOST_RSP),
-			       [host_rip]"r"((u64)HOST_RIP)
-			     : "memory", "cc", "rbx", "r8", "r9", "r10",
-			       "r11", "r12", "r13", "r14", "r15");
+			       [host_rip]"r"((u64)HOST_RIP),
+			       GUEST_REGS_OFFSETS
+			     : "memory", "cc");
 	return ret;
 }
 
-/*
- * No guest state (e.g. GPRs) is established by this vmresume.
- */
 static inline int vmresume(void)
 {
 	int ret;
@@ -408,28 +412,21 @@ static inline int vmresume(void)
 	if (enable_evmcs)
 		return evmcs_vmresume();
 
-	__asm__ __volatile__("push %%rbp;"
-			     "push %%rcx;"
-			     "push %%rdx;"
-			     "push %%rsi;"
-			     "push %%rdi;"
-			     "push $0;"
+	__asm__ __volatile__("push $0;"
 			     "vmwrite %%rsp, %[host_rsp];"
 			     "lea 1f(%%rip), %%rax;"
 			     "vmwrite %%rax, %[host_rip];"
+			     VMX_SWITCH_GPRS_ASM
 			     "vmresume;"
 			     "incq (%%rsp);"
-			     "1: pop %%rax;"
-			     "pop %%rdi;"
-			     "pop %%rsi;"
-			     "pop %%rdx;"
-			     "pop %%rcx;"
-			     "pop %%rbp;"
+			     "1: ;"
+			     VMX_SWITCH_GPRS_ASM
+			     "pop %%rax;"
 			     : [ret]"=&a"(ret)
 			     : [host_rsp]"r"((u64)HOST_RSP),
-			       [host_rip]"r"((u64)HOST_RIP)
-			     : "memory", "cc", "rbx", "r8", "r9", "r10",
-			       "r11", "r12", "r13", "r14", "r15");
+			       [host_rip]"r"((u64)HOST_RIP),
+			       GUEST_REGS_OFFSETS
+			     : "memory", "cc");
 	return ret;
 }
 
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [PATCH v5 06/13] KVM: selftests: Drop HORRIFIC_L2_UCALL_CLOBBER_HACK
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
                   ` (4 preceding siblings ...)
  2026-07-28 17:42 ` [PATCH v5 05/13] KVM: selftests: Reuse GPR switching logic for nVMX Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  2026-07-28 17:55   ` sashiko-bot
  2026-07-28 17:42 ` [PATCH v5 07/13] KVM: selftests: Add a blank line before logging assertion failures Yosry Ahmed
                   ` (6 subsequent siblings)
  12 siblings, 1 reply; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

Now that nVMX test codes preserves GPRs across nested VM-Exits
(specifically RBP, RDX, and RDI among others), drop the ucall-specific
hack to avoid clobbering these registers.

Assisted-by: Gemini:gemini-3.1-pro
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 tools/testing/selftests/kvm/lib/x86/ucall.c | 32 ++-------------------
 1 file changed, 2 insertions(+), 30 deletions(-)

diff --git a/tools/testing/selftests/kvm/lib/x86/ucall.c b/tools/testing/selftests/kvm/lib/x86/ucall.c
index e7dd5791959ba..38050c60a0670 100644
--- a/tools/testing/selftests/kvm/lib/x86/ucall.c
+++ b/tools/testing/selftests/kvm/lib/x86/ucall.c
@@ -10,36 +10,8 @@
 
 void ucall_arch_do_ucall(gva_t uc)
 {
-	/*
-	 * FIXME: Revert this hack (the entire commit that added it) once nVMX
-	 * preserves L2 GPRs across a nested VM-Exit.  If a ucall from L2, e.g.
-	 * to do a GUEST_SYNC(), lands the vCPU in L1, any and all GPRs can be
-	 * clobbered by L1.  Save and restore non-volatile GPRs (clobbering RBP
-	 * in particular is problematic) along with RDX and RDI (which are
-	 * inputs), and clobber volatile GPRs. *sigh*
-	 */
-#define HORRIFIC_L2_UCALL_CLOBBER_HACK	\
-	"rcx", "rsi", "r8", "r9", "r10", "r11"
-
-	asm volatile("push %%rbp\n\t"
-		     "push %%r15\n\t"
-		     "push %%r14\n\t"
-		     "push %%r13\n\t"
-		     "push %%r12\n\t"
-		     "push %%rbx\n\t"
-		     "push %%rdx\n\t"
-		     "push %%rdi\n\t"
-		     "in %[port], %%al\n\t"
-		     "pop %%rdi\n\t"
-		     "pop %%rdx\n\t"
-		     "pop %%rbx\n\t"
-		     "pop %%r12\n\t"
-		     "pop %%r13\n\t"
-		     "pop %%r14\n\t"
-		     "pop %%r15\n\t"
-		     "pop %%rbp\n\t"
-		: : [port] "d" (UCALL_PIO_PORT), "D" (uc) : "rax", "memory",
-		     HORRIFIC_L2_UCALL_CLOBBER_HACK);
+	asm volatile("in %[port], %%al"
+		: : [port] "d" (UCALL_PIO_PORT), "D" (uc) : "rax", "memory");
 }
 
 void *ucall_arch_get_ucall(struct kvm_vcpu *vcpu)
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [PATCH v5 07/13] KVM: selftests: Add a blank line before logging assertion failures
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
                   ` (5 preceding siblings ...)
  2026-07-28 17:42 ` [PATCH v5 06/13] KVM: selftests: Drop HORRIFIC_L2_UCALL_CLOBBER_HACK Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 08/13] KVM: selftests: Expose PTE masks to guests as part of an MMU Yosry Ahmed
                   ` (5 subsequent siblings)
  12 siblings, 0 replies; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

Add a blank line for visual separation, which is especially useful for
tests that use carriage returns for same-line printing, in case a
failure occurs before a blank line is ever logged by the test.

Suggested-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 tools/testing/selftests/kvm/lib/assert.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/tools/testing/selftests/kvm/lib/assert.c b/tools/testing/selftests/kvm/lib/assert.c
index 1d72dcdfce3b6..3e353ac39eeb7 100644
--- a/tools/testing/selftests/kvm/lib/assert.c
+++ b/tools/testing/selftests/kvm/lib/assert.c
@@ -74,7 +74,7 @@ test_assert(bool exp, const char *exp_str,
 	if (!(exp)) {
 		va_start(ap, fmt);
 
-		fprintf(stderr, "==== Test Assertion Failure ====\n"
+		fprintf(stderr, "\n==== Test Assertion Failure ====\n"
 			"  %s:%u: %s\n"
 			"  pid=%d tid=%d errno=%d - %s\n",
 			file, line, exp_str, getpid(), kvm_gettid(),
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [PATCH v5 08/13] KVM: selftests: Expose PTE masks to guests as part of an MMU
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
                   ` (6 preceding siblings ...)
  2026-07-28 17:42 ` [PATCH v5 07/13] KVM: selftests: Add a blank line before logging assertion failures Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 09/13] KVM: selftests: Do not intercept #PF by default in nVMX tests Yosry Ahmed
                   ` (4 subsequent siblings)
  12 siblings, 0 replies; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

Expose a guest_mmu to the guest to allow guest code to use the PTE masks
for page table manipulation. Since guest page tables are not mapped in
the guest by default, zero the PGD in guest_mmu in an attempt to make it
more difficult for new tests to shoot themselves in the foot and assume
that page tables can be immediately used by guest code.

Ultimately, guest code can read CR3 any way, so guest_mmu.pgd doesn't
add a lot of value.

Suggested-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 tools/testing/selftests/kvm/include/x86/processor.h |  1 +
 tools/testing/selftests/kvm/lib/x86/processor.c     | 12 ++++++++++++
 2 files changed, 13 insertions(+)

diff --git a/tools/testing/selftests/kvm/include/x86/processor.h b/tools/testing/selftests/kvm/include/x86/processor.h
index dab64688e8eae..06d9ba1c4df33 100644
--- a/tools/testing/selftests/kvm/include/x86/processor.h
+++ b/tools/testing/selftests/kvm/include/x86/processor.h
@@ -24,6 +24,7 @@ extern bool host_cpu_is_amd;
 extern bool host_cpu_is_hygon;
 extern bool host_cpu_is_amd_compatible;
 extern u64 guest_tsc_khz;
+extern struct kvm_mmu guest_mmu;
 
 #ifndef MAX_NR_CPUID_ENTRIES
 #define MAX_NR_CPUID_ENTRIES 100
diff --git a/tools/testing/selftests/kvm/lib/x86/processor.c b/tools/testing/selftests/kvm/lib/x86/processor.c
index 1f9201590f5b3..d31fa81ea0756 100644
--- a/tools/testing/selftests/kvm/lib/x86/processor.c
+++ b/tools/testing/selftests/kvm/lib/x86/processor.c
@@ -28,6 +28,7 @@ bool host_cpu_is_hygon;
 bool host_cpu_is_amd_compatible;
 bool is_forced_emulation_enabled;
 u64 guest_tsc_khz;
+struct kvm_mmu guest_mmu;
 
 struct guest_regs guest_regs;
 
@@ -831,6 +832,17 @@ void kvm_arch_vm_post_create(struct kvm_vm *vm, unsigned int nr_vcpus)
 	TEST_ASSERT(r > 0, "KVM_GET_TSC_KHZ did not provide a valid TSC frequency.");
 	guest_tsc_khz = r;
 	sync_global_to_guest(vm, guest_tsc_khz);
+
+	/*
+	 * The guest MMU is just a placeholder to provide access to PTE masks
+	 * (for now). The guest does not have mappings for its own page tables
+	 * by default, so any meaningful use of guest page tables requires
+	 * explicit setup by the test. Zero the PGD to make it obvious the guest
+	 * page tables are not immediately usable by guest code.
+	 */
+	guest_mmu = vm->mmu;
+	guest_mmu.pgd = 0;
+	sync_global_to_guest(vm, guest_mmu);
 }
 
 void vcpu_arch_set_entry_point(struct kvm_vcpu *vcpu, void *guest_code)
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [PATCH v5 09/13] KVM: selftests: Do not intercept #PF by default in nVMX tests
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
                   ` (7 preceding siblings ...)
  2026-07-28 17:42 ` [PATCH v5 08/13] KVM: selftests: Expose PTE masks to guests as part of an MMU Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  2026-07-28 17:45   ` Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 10/13] KVM: selftests: Add basic stress test for save+restore and #PF handling Yosry Ahmed
                   ` (3 subsequent siblings)
  12 siblings, 1 reply; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

init_vmcs_control_fields() sets PFEC_MASK and PFEC_MATCH so that they
never match, which reverses the meaning of the PF_VECTOR bit in
EXCEPTION_BITMAP (which is zeroed), effectively enabling #PF
interception by default.

The relevant part of the SDM describes this:
  When a page fault occurs, a processor consults (1) bit 14
  of the exception bitmap; (2) the error code produced with the page
  fault [PFEC]; (3) the page-fault error-code mask field [PFEC_MASK];
  and (4) the page-fault error-code match field [PFEC_MATCH]. It checks
  if PFEC & PFEC_MASK = PFEC_MATCH. If there is equality, the
  specification of bit 14 in the exception bitmap is followed (for
  example, a VM exit occurs if that bit is set). If there is inequality,
  the meaning of that bit is reversed (for example, a VM exit occurs if
  that bit is clear).

Clear PFEC_MATCH such that there is equality, and the #PF bit in the
exception bitmap is followed.

Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 tools/testing/selftests/kvm/lib/x86/vmx.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/tools/testing/selftests/kvm/lib/x86/vmx.c b/tools/testing/selftests/kvm/lib/x86/vmx.c
index cd09c9de4485a..089e1a8af53fc 100644
--- a/tools/testing/selftests/kvm/lib/x86/vmx.c
+++ b/tools/testing/selftests/kvm/lib/x86/vmx.c
@@ -232,7 +232,7 @@ static inline void init_vmcs_control_fields(struct vmx_pages *vmx)
 
 	vmwrite(EXCEPTION_BITMAP, 0);
 	vmwrite(PAGE_FAULT_ERROR_CODE_MASK, 0);
-	vmwrite(PAGE_FAULT_ERROR_CODE_MATCH, -1); /* Never match */
+	vmwrite(PAGE_FAULT_ERROR_CODE_MATCH, 0);
 	vmwrite(CR3_TARGET_COUNT, 0);
 	vmwrite(VM_EXIT_CONTROLS, rdmsr(MSR_IA32_VMX_EXIT_CTLS) |
 		VM_EXIT_HOST_ADDR_SPACE_SIZE);	  /* 64-bit host */
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [PATCH v5 10/13] KVM: selftests: Add basic stress test for save+restore and #PF handling
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
                   ` (8 preceding siblings ...)
  2026-07-28 17:42 ` [PATCH v5 09/13] KVM: selftests: Do not intercept #PF by default in nVMX tests Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 11/13] KVM: selftests: Trigger save+restore randomly in the #PF stress test Yosry Ahmed
                   ` (2 subsequent siblings)
  12 siblings, 0 replies; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

Add a basic stress test for handling #PFs in a guest while the host is
doing save+restore cycles. The guest periodically accesses non-present
memory causing a #PF, and the #PF handler walks the page tables and
updates the PTE to be present, like a proper #PF handler.

After every access (and #PF), the guest triggers a sync and the test
performs save+restore of the VM. This is not very meaningful as
save+restore are performed after the access and #PF handling complete,
but following changes will change that.

Assisted-by: Gemini:gemini-3.1-pro
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 tools/testing/selftests/kvm/Makefile.kvm      |   1 +
 .../selftests/kvm/include/x86/processor.h     |  13 ++
 .../kvm/x86/save_restore_pf_stress_test.c     | 160 ++++++++++++++++++
 3 files changed, 174 insertions(+)
 create mode 100644 tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c

diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm
index 88c6c8046ddec..d94f56c630fd5 100644
--- a/tools/testing/selftests/kvm/Makefile.kvm
+++ b/tools/testing/selftests/kvm/Makefile.kvm
@@ -107,6 +107,7 @@ TEST_GEN_PROGS_x86 += x86/pmu_counters_test
 TEST_GEN_PROGS_x86 += x86/pmu_event_filter_test
 TEST_GEN_PROGS_x86 += x86/private_mem_conversions_test
 TEST_GEN_PROGS_x86 += x86/private_mem_kvm_exits_test
+TEST_GEN_PROGS_x86 += x86/save_restore_pf_stress_test
 TEST_GEN_PROGS_x86 += x86/set_boot_cpu_id
 TEST_GEN_PROGS_x86 += x86/set_sregs_test
 TEST_GEN_PROGS_x86 += x86/smaller_maxphyaddr_emulation_test
diff --git a/tools/testing/selftests/kvm/include/x86/processor.h b/tools/testing/selftests/kvm/include/x86/processor.h
index 06d9ba1c4df33..5979d70f49147 100644
--- a/tools/testing/selftests/kvm/include/x86/processor.h
+++ b/tools/testing/selftests/kvm/include/x86/processor.h
@@ -614,6 +614,14 @@ static inline void set_cr0(u64 val)
 	__asm__ __volatile__("mov %0, %%cr0" : : "r" (val) : "memory");
 }
 
+static inline u64 get_cr2(void)
+{
+	u64 cr2;
+
+	__asm__ __volatile__("mov %%cr2, %[cr2]" : [cr2]"=r"(cr2));
+	return cr2;
+}
+
 static inline u64 get_cr3(void)
 {
 	u64 cr3;
@@ -909,6 +917,11 @@ static inline void write_sse_reg(int reg, const sse128_t *data)
 	}
 }
 
+static inline void invlpg(u64 addr)
+{
+	__asm__ __volatile__("invlpg (%0)" : : "r"(addr) : "memory");
+}
+
 static inline void cpu_relax(void)
 {
 	asm volatile("rep; nop" ::: "memory");
diff --git a/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c b/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c
new file mode 100644
index 0000000000000..1f7e142e5cfec
--- /dev/null
+++ b/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c
@@ -0,0 +1,160 @@
+// SPDX-License-Identifier: GPL-2.0-only
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <errno.h>
+#include <sys/types.h>
+#include <time.h>
+#include <unistd.h>
+
+#include "test_util.h"
+#include "kvm_util.h"
+#include "processor.h"
+
+#define NR_ITERATIONS		500
+
+#define PTRS_PER_PTE		512
+#define PXD_INDEX(vaddr, level)	(((vaddr) >> PG_LEVEL_SHIFT(level)) & (PTRS_PER_PTE - 1))
+
+#define TEST_MEM_BASE_GVA	0xc0000000ULL
+#define TEST_PGTABLE_GVA_OFFSET	0xd0000000ULL
+#define PATTERN			0xabcdefabcdefabcdULL
+
+static u64 expected_vaddr;
+static u64 guest_faults;
+
+static u64 *guest_get_pte(u64 vaddr)
+{
+	u64 pgtable_pa, pte;
+	u64 *pgtable;
+	int level;
+
+	level = (get_cr4() & X86_CR4_LA57) ? PG_LEVEL_256T : PG_LEVEL_512G;
+
+	pgtable_pa = get_cr3() & PHYSICAL_PAGE_MASK;
+	for (; level > PG_LEVEL_4K; level--) {
+		pgtable = (u64 *)(pgtable_pa + TEST_PGTABLE_GVA_OFFSET);
+		pte = pgtable[PXD_INDEX(vaddr, level)];
+		GUEST_ASSERT(pte & PTE_PRESENT_MASK(&guest_mmu));
+		GUEST_ASSERT(!(pte & PTE_HUGE_MASK(&guest_mmu)));
+		pgtable_pa = PTE_GET_PA(pte);
+	}
+
+	pgtable = (u64 *)(pgtable_pa + TEST_PGTABLE_GVA_OFFSET);
+	return &pgtable[PXD_INDEX(vaddr, PG_LEVEL_4K)];
+}
+
+static void guest_pf_handler(struct ex_regs *regs)
+{
+	u64 fault_addr;
+	u64 *ptep;
+
+	fault_addr = get_cr2();
+	GUEST_ASSERT_EQ(fault_addr, READ_ONCE(expected_vaddr));
+
+	ptep = guest_get_pte(fault_addr);
+	GUEST_ASSERT(ptep);
+	GUEST_ASSERT(!(*ptep & PTE_PRESENT_MASK(&guest_mmu)));
+
+	*ptep |= PTE_PRESENT_MASK(&guest_mmu);
+	guest_faults++;
+}
+
+static void guest_access_memory(void *arg)
+{
+	u64 vaddr, val;
+	int i;
+
+	for (i = 0; ; i++) {
+		vaddr = TEST_MEM_BASE_GVA + (i % PTRS_PER_PTE) * PAGE_SIZE;
+		WRITE_ONCE(expected_vaddr, vaddr);
+
+		/* Read to trigger #PF */
+		val = READ_ONCE(*(u64 *)vaddr);
+		GUEST_ASSERT_EQ(val, PATTERN);
+
+		/* Clear the present bit again so it faults next time */
+		*guest_get_pte(vaddr) &= ~PTE_PRESENT_MASK(&guest_mmu);
+		invlpg(vaddr);
+
+		GUEST_SYNC(guest_faults);
+	}
+}
+
+int main(int argc, char *argv[])
+{
+	struct kvm_x86_state *state;
+	int r, i, level;
+	gpa_t gpa, pgtable_gpa;
+	struct kvm_vcpu *vcpu;
+	struct kvm_vm *vm;
+	struct ucall uc;
+	u64 *pgtable;
+	gva_t gva;
+	u64 pte;
+
+	vm = vm_create_with_one_vcpu(&vcpu, guest_access_memory);
+	vm_install_exception_handler(vm, PF_VECTOR, guest_pf_handler);
+
+	/* Allocate a page and write the pattern to it */
+	gva = vm_alloc_page(vm);
+	*(u64 *)addr_gva2hva(vm, gva) = PATTERN;
+	gpa = addr_gva2gpa(vm, gva);
+
+	/*
+	 * Map all virtual addresses to the pattern page and clear the present
+	 * bit such that guest accesses will cause a #PF.
+	 */
+	for (i = 0; i < PTRS_PER_PTE; i++) {
+		gva = TEST_MEM_BASE_GVA + i * getpagesize();
+		virt_pg_map(vm, gva, gpa);
+		*vm_get_pte(vm, gva) &= ~PTE_PRESENT_MASK(&vm->mmu);
+	}
+
+	/*
+	 * Now create mappings for the page tables created above so that the
+	 * guest #PF handler can walk them. All PTEs for test virtual addresses
+	 * should lie on the same PTE page, so one page is mapped for each page
+	 * table level.
+	 *
+	 * Use an offset for the GVA instead of creating identity mappings to
+	 * avoid collision with existing mappings at low GVAs (e.g. ELF).
+	 */
+	pgtable_gpa = vm->mmu.pgd;
+	for (level = vm->mmu.pgtable_levels; level >= PG_LEVEL_4K; level--) {
+		virt_map(vm, pgtable_gpa + TEST_PGTABLE_GVA_OFFSET, pgtable_gpa, 1);
+		pgtable = addr_gpa2hva(vm, pgtable_gpa);
+		pte = pgtable[PXD_INDEX(TEST_MEM_BASE_GVA, level)];
+		pgtable_gpa = PTE_GET_PA(pte);
+	}
+
+	for (i = 1; i <= NR_ITERATIONS; i++) {
+		r = __vcpu_run(vcpu);
+		TEST_ASSERT(!r, "vcpu_run failed");
+		TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_IO);
+
+		get_ucall(vcpu, &uc);
+		if (uc.cmd == UCALL_ABORT) {
+			REPORT_GUEST_ASSERT(uc);
+			break;
+		}
+		TEST_ASSERT_EQ(uc.cmd, UCALL_SYNC);
+		TEST_ASSERT_EQ(uc.args[1], i);
+
+		state = vcpu_save_state(vcpu);
+
+		kvm_vm_release(vm);
+		vcpu = vm_recreate_with_one_vcpu(vm);
+		vcpu_load_state(vcpu, state);
+		kvm_x86_state_cleanup(state);
+
+		pr_info("\rSave+restore iterations: %d", i);
+	}
+	pr_info("\n");
+
+	sync_global_from_guest(vm, guest_faults);
+	pr_info("Guest page faults: %lu\n", guest_faults);
+
+	kvm_vm_free(vm);
+	return 0;
+}
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [PATCH v5 11/13] KVM: selftests: Trigger save+restore randomly in the #PF stress test
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
                   ` (9 preceding siblings ...)
  2026-07-28 17:42 ` [PATCH v5 10/13] KVM: selftests: Add basic stress test for save+restore and #PF handling Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 12/13] KVM: selftests: Support running stress save+restore and #PF test in L2 Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 13/13] KVM: selftests: Trigger L2->L1 exits stress save+restore and #PF test Yosry Ahmed
  12 siblings, 0 replies; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

Instead of an explicit GUEST_SYNC() after each access+#PF, run another
thread that keeps sending SIGUSR to the vCPU thread, essentially
triggering exits to userspace and save+restore on random points in guest
execution. This makes the test a lot more meaningful as it opens the
door to exercising race conditions between #PF handling in the guest
and save+restore in the host.

The signals are ignored using SIG_IGN outside of __vcpu_run() to avoid
interrupting other ioctls/sysctls performed by the test.

Assisted-by: Gemini:gemini-3.1-pro
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 .../kvm/x86/save_restore_pf_stress_test.c     | 56 ++++++++++++++++---
 1 file changed, 49 insertions(+), 7 deletions(-)

diff --git a/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c b/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c
index 1f7e142e5cfec..4b5accec642c0 100644
--- a/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c
+++ b/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c
@@ -5,6 +5,8 @@
 #include <errno.h>
 #include <sys/types.h>
 #include <time.h>
+#include <pthread.h>
+#include <signal.h>
 #include <unistd.h>
 
 #include "test_util.h"
@@ -76,15 +78,41 @@ static void guest_access_memory(void *arg)
 		/* Clear the present bit again so it faults next time */
 		*guest_get_pte(vaddr) &= ~PTE_PRESENT_MASK(&guest_mmu);
 		invlpg(vaddr);
+	}
+}
+
+static void *sigusr_thread_fn(void *arg)
+{
+	pthread_t vcpu_thread = (pthread_t)arg;
 
-		GUEST_SYNC(guest_faults);
+	for (;;) {
+		pthread_testcancel();
+		pthread_kill(vcpu_thread, SIGUSR1);
+		usleep(msecs_to_usecs(1));
 	}
+	return NULL;
+}
+
+static void dummy_signal_handler(int signo) {}
+static struct sigaction sa;
+
+static void vcpu_sigusr_listen(void)
+{
+	sa.sa_handler = dummy_signal_handler;
+	sigaction(SIGUSR1, &sa, NULL);
+}
+
+static void vcpu_sigusr_ignore(void)
+{
+	sa.sa_handler = SIG_IGN;
+	sigaction(SIGUSR1, &sa, NULL);
 }
 
 int main(int argc, char *argv[])
 {
 	struct kvm_x86_state *state;
 	int r, i, level;
+	pthread_t sigusr_thread;
 	gpa_t gpa, pgtable_gpa;
 	struct kvm_vcpu *vcpu;
 	struct kvm_vm *vm;
@@ -128,18 +156,29 @@ int main(int argc, char *argv[])
 		pgtable_gpa = PTE_GET_PA(pte);
 	}
 
+	/* Initialize the thread sending SIGUSR and install the handler */
+	vcpu_sigusr_ignore();
+	r = pthread_create(&sigusr_thread, NULL, sigusr_thread_fn,
+			   (void *)pthread_self());
+	TEST_ASSERT(!r, "pthread_create() failed: %d", r);
+
 	for (i = 1; i <= NR_ITERATIONS; i++) {
+		/*
+		 * Only handle SIGUSR while the vCPU is running, otherwise
+		 * ignore it to avoid interrupting other ioctls/syscalls.
+		 */
+		vcpu_sigusr_listen();
 		r = __vcpu_run(vcpu);
-		TEST_ASSERT(!r, "vcpu_run failed");
-		TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_IO);
+		TEST_ASSERT(!r || errno == EINTR, "Expected success or SIGUSR1");
+		vcpu_sigusr_ignore();
 
-		get_ucall(vcpu, &uc);
-		if (uc.cmd == UCALL_ABORT) {
+		/* The guest only exits due to a signal or failed assertion */
+		if (!r) {
+			TEST_ASSERT_KVM_EXIT_REASON(vcpu, KVM_EXIT_IO);
+			TEST_ASSERT_EQ(get_ucall(vcpu, &uc), UCALL_ABORT);
 			REPORT_GUEST_ASSERT(uc);
 			break;
 		}
-		TEST_ASSERT_EQ(uc.cmd, UCALL_SYNC);
-		TEST_ASSERT_EQ(uc.args[1], i);
 
 		state = vcpu_save_state(vcpu);
 
@@ -153,8 +192,11 @@ int main(int argc, char *argv[])
 	pr_info("\n");
 
 	sync_global_from_guest(vm, guest_faults);
+	TEST_ASSERT(guest_faults, "No guest page faults triggered");
 	pr_info("Guest page faults: %lu\n", guest_faults);
 
+	pthread_cancel(sigusr_thread);
+	pthread_join(sigusr_thread, NULL);
 	kvm_vm_free(vm);
 	return 0;
 }
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [PATCH v5 12/13] KVM: selftests: Support running stress save+restore and #PF test in L2
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
                   ` (10 preceding siblings ...)
  2026-07-28 17:42 ` [PATCH v5 11/13] KVM: selftests: Trigger save+restore randomly in the #PF stress test Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  2026-07-28 17:42 ` [PATCH v5 13/13] KVM: selftests: Trigger L2->L1 exits stress save+restore and #PF test Yosry Ahmed
  12 siblings, 0 replies; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

Extend the stress test to allow running the access+#PF code in L2
instead of L1 by adding proper L1 guest code to bootstrap L2. By
default, the test runs in L2 after running in L1 if nested is supported.

Assisted-by: Gemini:gemini-3.1-pro
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 .../kvm/x86/save_restore_pf_stress_test.c     | 56 ++++++++++++++++++-
 1 file changed, 53 insertions(+), 3 deletions(-)

diff --git a/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c b/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c
index 4b5accec642c0..7418e5fb879bd 100644
--- a/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c
+++ b/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c
@@ -8,10 +8,13 @@
 #include <pthread.h>
 #include <signal.h>
 #include <unistd.h>
+#include <getopt.h>
 
 #include "test_util.h"
 #include "kvm_util.h"
 #include "processor.h"
+#include "svm_util.h"
+#include "vmx.h"
 
 #define NR_ITERATIONS		500
 
@@ -81,6 +84,31 @@ static void guest_access_memory(void *arg)
 	}
 }
 
+static void l1_svm_code(struct svm_test_data *svm)
+{
+	generic_svm_setup(svm, guest_access_memory);
+	run_guest(svm->vmcb, svm->vmcb_gpa);
+	GUEST_ASSERT(false);
+}
+
+static void l1_vmx_code(struct vmx_pages *vmx)
+{
+	GUEST_ASSERT(prepare_for_vmx_operation(vmx));
+	GUEST_ASSERT(load_vmcs(vmx));
+	prepare_vmcs(vmx, guest_access_memory);
+
+	GUEST_ASSERT(!vmlaunch());
+	GUEST_ASSERT(false);
+}
+
+static void l1_guest_code(void *test_data)
+{
+	if (this_cpu_has(X86_FEATURE_SVM))
+		l1_svm_code(test_data);
+	else
+		l1_vmx_code(test_data);
+}
+
 static void *sigusr_thread_fn(void *arg)
 {
 	pthread_t vcpu_thread = (pthread_t)arg;
@@ -108,7 +136,7 @@ static void vcpu_sigusr_ignore(void)
 	sigaction(SIGUSR1, &sa, NULL);
 }
 
-int main(int argc, char *argv[])
+static void run_test(bool nested)
 {
 	struct kvm_x86_state *state;
 	int r, i, level;
@@ -121,9 +149,17 @@ int main(int argc, char *argv[])
 	gva_t gva;
 	u64 pte;
 
-	vm = vm_create_with_one_vcpu(&vcpu, guest_access_memory);
+	vm = vm_create_with_one_vcpu(&vcpu, nested ? l1_guest_code : guest_access_memory);
 	vm_install_exception_handler(vm, PF_VECTOR, guest_pf_handler);
 
+	if (nested) {
+		if (kvm_cpu_has(X86_FEATURE_SVM))
+			vcpu_alloc_svm(vm, &gva);
+		else
+			vcpu_alloc_vmx(vm, &gva);
+		vcpu_args_set(vcpu, 1, gva);
+	}
+
 	/* Allocate a page and write the pattern to it */
 	gva = vm_alloc_page(vm);
 	*(u64 *)addr_gva2hva(vm, gva) = PATTERN;
@@ -193,10 +229,24 @@ int main(int argc, char *argv[])
 
 	sync_global_from_guest(vm, guest_faults);
 	TEST_ASSERT(guest_faults, "No guest page faults triggered");
-	pr_info("Guest page faults: %lu\n", guest_faults);
+	pr_info("Guest page faults%s: %lu\n", nested ? " (in L2)" : "", guest_faults);
 
 	pthread_cancel(sigusr_thread);
 	pthread_join(sigusr_thread, NULL);
 	kvm_vm_free(vm);
+}
+
+int main(int argc, char *argv[])
+{
+	pr_info("Running save+restore stress test...\n");
+	run_test(/*nested=*/false);
+
+	if (!kvm_cpu_has(X86_FEATURE_SVM) && !kvm_cpu_has(X86_FEATURE_VMX)) {
+		pr_info("Nested virtualization not supported, skipping nested test\n");
+		return 0;
+	}
+
+	pr_info("Running save+restore stress test with a nested guest...\n");
+	run_test(/*nested=*/true);
 	return 0;
 }
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [PATCH v5 13/13] KVM: selftests: Trigger L2->L1 exits stress save+restore and #PF test
  2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
                   ` (11 preceding siblings ...)
  2026-07-28 17:42 ` [PATCH v5 12/13] KVM: selftests: Support running stress save+restore and #PF test in L2 Yosry Ahmed
@ 2026-07-28 17:42 ` Yosry Ahmed
  12 siblings, 0 replies; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:42 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel, Yosry Ahmed

Extend the testing coverage in L2 by forcing a nested VM-Exit from L2 to
L1 right after restore on every other iteration. Forcing a nested
VM-Exit while L0 has control (e.g. without explicitly running L2 and
making a hypercall) is valuable, as it often happens during live
migration (e.g. L1 timer interrupt fires by the time the VM lands on the
destination).

To force the nested VM-Exit inject a #UD in to the saved vCPU state, and
intercept #UD from L1.

With this change, the test reliably reproduces the CR2 bug fixed by
commit 5c247d08bc81 ("KVM: nSVM: Use vcpu->arch.cr2 when updating vmcb12
on nested #VMEXIT") -- at least on Milan, Genoa, and Turin CPUs.

Assisted-by: Gemini:gemini-3.1-pro
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
---
 .../selftests/kvm/include/x86/processor.h     |  5 +++
 .../kvm/x86/save_restore_pf_stress_test.c     | 44 +++++++++++++++++--
 2 files changed, 45 insertions(+), 4 deletions(-)

diff --git a/tools/testing/selftests/kvm/include/x86/processor.h b/tools/testing/selftests/kvm/include/x86/processor.h
index 5979d70f49147..6e6f70035508a 100644
--- a/tools/testing/selftests/kvm/include/x86/processor.h
+++ b/tools/testing/selftests/kvm/include/x86/processor.h
@@ -958,6 +958,11 @@ struct kvm_x86_state *vcpu_save_state(struct kvm_vcpu *vcpu);
 void vcpu_load_state(struct kvm_vcpu *vcpu, struct kvm_x86_state *state);
 void kvm_x86_state_cleanup(struct kvm_x86_state *state);
 
+static inline bool kvm_x86_state_is_guest_mode(struct kvm_x86_state *state)
+{
+	return state->nested.size && (state->nested.flags & KVM_STATE_NESTED_GUEST_MODE);
+}
+
 const struct kvm_msr_list *kvm_get_msr_index_list(void);
 const struct kvm_msr_list *kvm_get_feature_msr_index_list(void);
 bool kvm_msr_is_in_save_restore_list(u32 msr_index);
diff --git a/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c b/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c
index 7418e5fb879bd..507391ab2c930 100644
--- a/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c
+++ b/tools/testing/selftests/kvm/x86/save_restore_pf_stress_test.c
@@ -87,8 +87,13 @@ static void guest_access_memory(void *arg)
 static void l1_svm_code(struct svm_test_data *svm)
 {
 	generic_svm_setup(svm, guest_access_memory);
-	run_guest(svm->vmcb, svm->vmcb_gpa);
-	GUEST_ASSERT(false);
+	svm->vmcb->control.intercept_exceptions |= BIT(UD_VECTOR);
+
+	while (1) {
+		run_guest(svm->vmcb, svm->vmcb_gpa);
+		GUEST_ASSERT_EQ(svm->vmcb->control.exit_code,
+				(SVM_EXIT_EXCP_BASE + UD_VECTOR));
+	}
 }
 
 static void l1_vmx_code(struct vmx_pages *vmx)
@@ -97,8 +102,14 @@ static void l1_vmx_code(struct vmx_pages *vmx)
 	GUEST_ASSERT(load_vmcs(vmx));
 	prepare_vmcs(vmx, guest_access_memory);
 
+	GUEST_ASSERT(!vmwrite(EXCEPTION_BITMAP, BIT(UD_VECTOR)));
+
 	GUEST_ASSERT(!vmlaunch());
-	GUEST_ASSERT(false);
+	while (1) {
+		GUEST_ASSERT_EQ(vmreadz(VM_EXIT_REASON), EXIT_REASON_EXCEPTION_NMI);
+		GUEST_ASSERT_EQ(vmreadz(VM_EXIT_INTR_INFO) & 0xff, UD_VECTOR);
+		GUEST_ASSERT(!vmresume());
+	}
 }
 
 static void l1_guest_code(void *test_data)
@@ -136,6 +147,19 @@ static void vcpu_sigusr_ignore(void)
 	sigaction(SIGUSR1, &sa, NULL);
 }
 
+static void kvm_x86_state_queue_ud(struct kvm_x86_state *state)
+{
+	if (state->events.exception.pending || state->events.exception.injected)
+		return;
+
+	state->events.flags |= KVM_VCPUEVENT_VALID_PAYLOAD;
+	state->events.exception.pending = true;
+	state->events.exception.injected = false;
+	state->events.exception.nr = UD_VECTOR;
+	state->events.exception.has_error_code = false;
+	state->events.exception_has_payload = false;
+}
+
 static void run_test(bool nested)
 {
 	struct kvm_x86_state *state;
@@ -153,6 +177,7 @@ static void run_test(bool nested)
 	vm_install_exception_handler(vm, PF_VECTOR, guest_pf_handler);
 
 	if (nested) {
+		vm_enable_cap(vm, KVM_CAP_EXCEPTION_PAYLOAD, -2ul);
 		if (kvm_cpu_has(X86_FEATURE_SVM))
 			vcpu_alloc_svm(vm, &gva);
 		else
@@ -218,8 +243,17 @@ static void run_test(bool nested)
 
 		state = vcpu_save_state(vcpu);
 
+		/*
+		 * If the vCPU is in guest mode, inject a #UD to trigger an
+		 * L2->L1 VM-Exit every other iteration.
+		 */
+		if (kvm_x86_state_is_guest_mode(state) && i % 2 == 0)
+			kvm_x86_state_queue_ud(state);
+
 		kvm_vm_release(vm);
 		vcpu = vm_recreate_with_one_vcpu(vm);
+		if (nested)
+			vm_enable_cap(vm, KVM_CAP_EXCEPTION_PAYLOAD, -2ul);
 		vcpu_load_state(vcpu, state);
 		kvm_x86_state_cleanup(state);
 
@@ -241,7 +275,9 @@ int main(int argc, char *argv[])
 	pr_info("Running save+restore stress test...\n");
 	run_test(/*nested=*/false);
 
-	if (!kvm_cpu_has(X86_FEATURE_SVM) && !kvm_cpu_has(X86_FEATURE_VMX)) {
+	if (!kvm_has_cap(KVM_CAP_EXCEPTION_PAYLOAD) ||
+	    !kvm_has_cap(KVM_CAP_NESTED_STATE) ||
+	    (!kvm_cpu_has(X86_FEATURE_SVM) && !kvm_cpu_has(X86_FEATURE_VMX))) {
 		pr_info("Nested virtualization not supported, skipping nested test\n");
 		return 0;
 	}
-- 
2.55.0.487.gaf234c4eb3-goog


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* Re: [PATCH v5 09/13] KVM: selftests: Do not intercept #PF by default in nVMX tests
  2026-07-28 17:42 ` [PATCH v5 09/13] KVM: selftests: Do not intercept #PF by default in nVMX tests Yosry Ahmed
@ 2026-07-28 17:45   ` Yosry Ahmed
  0 siblings, 0 replies; 18+ messages in thread
From: Yosry Ahmed @ 2026-07-28 17:45 UTC (permalink / raw)
  To: Sean Christopherson; +Cc: Paolo Bonzini, kvm, linux-kernel

On Tue, Jul 28, 2026 at 10:42 AM Yosry Ahmed <yosry@kernel.org> wrote:
>
> init_vmcs_control_fields() sets PFEC_MASK and PFEC_MATCH so that they
> never match, which reverses the meaning of the PF_VECTOR bit in
> EXCEPTION_BITMAP (which is zeroed), effectively enabling #PF
> interception by default.
>
> The relevant part of the SDM describes this:
>   When a page fault occurs, a processor consults (1) bit 14
>   of the exception bitmap; (2) the error code produced with the page
>   fault [PFEC]; (3) the page-fault error-code mask field [PFEC_MASK];
>   and (4) the page-fault error-code match field [PFEC_MATCH]. It checks
>   if PFEC & PFEC_MASK = PFEC_MATCH. If there is equality, the
>   specification of bit 14 in the exception bitmap is followed (for
>   example, a VM exit occurs if that bit is set). If there is inequality,
>   the meaning of that bit is reversed (for example, a VM exit occurs if
>   that bit is clear).
>
> Clear PFEC_MATCH such that there is equality, and the #PF bit in the
> exception bitmap is followed.
>
> Signed-off-by: Yosry Ahmed <yosry@kernel.org>

Oh this needs either:

Suggested-by: Sean Christopherson <seanjc@google.com>

or

Co-developed-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>

^ permalink raw reply	[flat|nested] 18+ messages in thread

* Re: [PATCH v5 05/13] KVM: selftests: Reuse GPR switching logic for nVMX
  2026-07-28 17:42 ` [PATCH v5 05/13] KVM: selftests: Reuse GPR switching logic for nVMX Yosry Ahmed
@ 2026-07-28 17:54   ` sashiko-bot
  0 siblings, 0 replies; 18+ messages in thread
From: sashiko-bot @ 2026-07-28 17:54 UTC (permalink / raw)
  To: Yosry Ahmed; +Cc: kvm

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] The newly introduced VMX_SWITCH_GPRS_ASM macro uses a shared global variable `guest_regs` for swapping host and guest general-purpose registers, introducing a race condition when multiple vCPUs execute nested VMX instructions concurrently.
--

--- Patch [5]: [PATCH v5 05/13] KVM: selftests: Reuse GPR switching logic for nVMX ---
commit a0021d28e505a4c708c955c20d96663fad9fc879
Author: Yosry Ahmed <yosry@kernel.org>

KVM: selftests: Reuse GPR switching logic for nVMX

This commit reuses the SVM GPR switching logic for nVMX by defining
VMX_SWITCH_GPRS_ASM. It replaces the previous push/pop sequence with
macro calls that swap a broader set of registers using GUEST_SWITCH_GPR_ASM.

> diff --git a/tools/testing/selftests/kvm/include/x86/vmx.h b/tools/testing/selftests/kvm/include/x86/vmx.h
> --- a/tools/testing/selftests/kvm/include/x86/vmx.h
> +++ b/tools/testing/selftests/kvm/include/x86/vmx.h
> @@ -290,6 +290,23 @@ struct vmx_msr_entry {
>  	u64 value;
>  } __attribute__ ((aligned(16)));
>  
> +#define VMX_SWITCH_GPRS_ASM \
> +	GUEST_SWITCH_GPR_ASM(rax) \
> +	GUEST_SWITCH_GPR_ASM(rbx) \

[Severity: Medium]
Since GUEST_SWITCH_GPR_ASM() uses the global guest_regs structure, will this
cause a data race when multiple vCPUs execute nested VMX instructions
concurrently?

When vmlaunch or vmresume is executed concurrently by multiple threads (such
as in the memstress test), they might race to read and write to this shared
variable, potentially corrupting host callee-saved registers like rbx or r12.
Could this lead to host test code crashes when one vCPU restores another
vCPU's parked host registers?

> +	GUEST_SWITCH_GPR_ASM(rcx) \
> +	GUEST_SWITCH_GPR_ASM(rdx) \
[ ... ]

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260728174232.2423257-1-yosry@kernel.org?part=5

^ permalink raw reply	[flat|nested] 18+ messages in thread

* Re: [PATCH v5 02/13] KVM: selftests: Fix RAX and RFLAGS VMCB offsets when running L2
  2026-07-28 17:42 ` [PATCH v5 02/13] KVM: selftests: Fix RAX and RFLAGS VMCB offsets when running L2 Yosry Ahmed
@ 2026-07-28 17:55   ` sashiko-bot
  0 siblings, 0 replies; 18+ messages in thread
From: sashiko-bot @ 2026-07-28 17:55 UTC (permalink / raw)
  To: Yosry Ahmed; +Cc: kvm

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:

Pre-existing issues:
- [High] The inline assembly macros LOAD_GPR_C and SAVE_GPR_C use incorrect offsets to access guest_regs, shifted by 8 bytes. This corrupts guest registers and causes an out-of-bounds access that clobbers the rflags variable.
--

--- Patch [2]: [PATCH v5 02/13] KVM: selftests: Fix RAX and RFLAGS VMCB offsets when running L2 ---
commit b67ef5c1f8e53b1e2adf72da7c94eb12e46b483e
Author: Yosry Ahmed <yosry@kernel.org>

KVM: selftests: Fix RAX and RFLAGS VMCB offsets when running L2
    
This commit corrects the offsets for rax and rflags within the VMCB save
area. It uses the vmcb->save.rax and vmcb->save.rflags C variables directly
as parameters in the inline assembly, avoiding hardcoded miscalculated
offsets.
    
Link: https://lore.kernel.org/all/20260521092311.86030-1-pbonzini@redhat.com/
Closes: https://sashiko.dev/#/patchset/20260518202514.2037078-1-yosry%40kernel.org?part=1

> diff --git a/tools/testing/selftests/kvm/lib/x86/svm.c b/tools/testing/selftests/kvm/lib/x86/svm.c
> index 1445b890986fd..766d15f1d534a 100644
> --- a/tools/testing/selftests/kvm/lib/x86/svm.c
> +++ b/tools/testing/selftests/kvm/lib/x86/svm.c
> @@ -164,19 +164,22 @@ void run_guest(struct vmcb *vmcb, u64 vmcb_gpa)
>  {
>  	asm volatile (
[ ... ]
>  		LOAD_GPR_C
>  		"vmrun %[vmcb_gpa]\n\t"
>  		SAVE_GPR_C

[Severity: High]
This is a pre-existing issue, but I noticed an out-of-bounds memory access
in tools/testing/selftests/kvm/lib/x86/svm.c related to the SAVE_GPR_C macro 
used here in run_guest(). 

The macro at line 135 contains the instruction "xchg %%r15, guest_regs+0x80".

Since struct gpr64_regs is exactly 128 bytes (0x80) long, doesn't this
instruction write past the end of the guest_regs structure and overwrite the
adjacent rflags variable?

Could the offsets in LOAD_GPR_C and SAVE_GPR_C be shifted by 8 bytes, causing
them to initialize guest registers from the wrong offsets and corrupt memory?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260728174232.2423257-1-yosry@kernel.org?part=2

^ permalink raw reply	[flat|nested] 18+ messages in thread

* Re: [PATCH v5 06/13] KVM: selftests: Drop HORRIFIC_L2_UCALL_CLOBBER_HACK
  2026-07-28 17:42 ` [PATCH v5 06/13] KVM: selftests: Drop HORRIFIC_L2_UCALL_CLOBBER_HACK Yosry Ahmed
@ 2026-07-28 17:55   ` sashiko-bot
  0 siblings, 0 replies; 18+ messages in thread
From: sashiko-bot @ 2026-07-28 17:55 UTC (permalink / raw)
  To: Yosry Ahmed; +Cc: kvm

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] Data race causing register corruption in multi-vCPU nested tests due to globally shared guest_regs structure.
--

--- Patch [6]: [PATCH v5 06/13] KVM: selftests: Drop HORRIFIC_L2_UCALL_CLOBBER_HACK ---
commit 9a9a1bcc6fc0c85786cd983d4fae11f1dfafdf51
Author: Yosry Ahmed <yosry@kernel.org>

KVM: selftests: Drop HORRIFIC_L2_UCALL_CLOBBER_HACK

This commit drops the hack in ucall_arch_do_ucall(), as the nVMX test
code now preserves GPRs across nested VM-Exits.

> diff --git a/tools/testing/selftests/kvm/lib/x86/ucall.c b/tools/testing/selftests/kvm/lib/x86/ucall.c
> index e7dd5791959ba..38050c60a0670 100644
> --- a/tools/testing/selftests/kvm/lib/x86/ucall.c
> +++ b/tools/testing/selftests/kvm/lib/x86/ucall.c
> @@ -10,36 +10,8 @@
>  
>  void ucall_arch_do_ucall(gva_t uc)
>  {
[ ... ]
> -		     "pop %%r15\n\t"
> -		     "pop %%rbp\n\t"
> -		: : [port] "d" (UCALL_PIO_PORT), "D" (uc) : "rax", "memory",
> -		     HORRIFIC_L2_UCALL_CLOBBER_HACK);
> +	asm volatile("in %[port], %%al"
> +		: : [port] "d" (UCALL_PIO_PORT), "D" (uc) : "rax", "memory");
>  }

[Severity: Medium]
Does dropping this hack introduce a regression for multi-vCPU nested tests
where L2 registers could get clobbered?

The nVMX code that now preserves GPRs appears to use a globally shared
structure:

tools/testing/selftests/kvm/lib/x86/processor.c:
struct guest_regs guest_regs;

The VMX_SWITCH_GPRS_ASM macro uses this struct during nested VM entry and
exit:

tools/testing/selftests/kvm/include/x86/vmx.h:
#define VMX_SWITCH_GPRS_ASM \
	GUEST_SWITCH_GPR_ASM(rax) \
	GUEST_SWITCH_GPR_ASM(rbx) \
...

Because guest_regs is not thread-local, wouldn't concurrent vCPU threads
running vmlaunch or vmresume clobber each other's register state when
accessing this shared struct?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260728174232.2423257-1-yosry@kernel.org?part=6

^ permalink raw reply	[flat|nested] 18+ messages in thread

end of thread, other threads:[~2026-07-28 17:55 UTC | newest]

Thread overview: 18+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-28 17:42 [PATCH v5 00/13] KVM: selftests: Stress save+restore and #PF (ft. nested) Yosry Ahmed
2026-07-28 17:42 ` [PATCH v5 01/13] KVM: selftests: Use __stringify() instead of custom XSTR() macros Yosry Ahmed
2026-07-28 17:42 ` [PATCH v5 02/13] KVM: selftests: Fix RAX and RFLAGS VMCB offsets when running L2 Yosry Ahmed
2026-07-28 17:55   ` sashiko-bot
2026-07-28 17:42 ` [PATCH v5 03/13] KVM: selftests: Rework GPR registers switching for SVM (and fix offsets) Yosry Ahmed
2026-07-28 17:42 ` [PATCH v5 04/13] KVM: selftests: Handle rflags save/restore for SVM in guest_regs Yosry Ahmed
2026-07-28 17:42 ` [PATCH v5 05/13] KVM: selftests: Reuse GPR switching logic for nVMX Yosry Ahmed
2026-07-28 17:54   ` sashiko-bot
2026-07-28 17:42 ` [PATCH v5 06/13] KVM: selftests: Drop HORRIFIC_L2_UCALL_CLOBBER_HACK Yosry Ahmed
2026-07-28 17:55   ` sashiko-bot
2026-07-28 17:42 ` [PATCH v5 07/13] KVM: selftests: Add a blank line before logging assertion failures Yosry Ahmed
2026-07-28 17:42 ` [PATCH v5 08/13] KVM: selftests: Expose PTE masks to guests as part of an MMU Yosry Ahmed
2026-07-28 17:42 ` [PATCH v5 09/13] KVM: selftests: Do not intercept #PF by default in nVMX tests Yosry Ahmed
2026-07-28 17:45   ` Yosry Ahmed
2026-07-28 17:42 ` [PATCH v5 10/13] KVM: selftests: Add basic stress test for save+restore and #PF handling Yosry Ahmed
2026-07-28 17:42 ` [PATCH v5 11/13] KVM: selftests: Trigger save+restore randomly in the #PF stress test Yosry Ahmed
2026-07-28 17:42 ` [PATCH v5 12/13] KVM: selftests: Support running stress save+restore and #PF test in L2 Yosry Ahmed
2026-07-28 17:42 ` [PATCH v5 13/13] KVM: selftests: Trigger L2->L1 exits stress save+restore and #PF test Yosry Ahmed

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.