Kernel KVM virtualization development
 help / color / mirror / Atom feed
* [PATCH v3 0/9] Implement support for IBS virtualization
@ 2026-03-10  6:00 Manali Shukla
  2026-03-10  6:00 ` [PATCH v3 1/9] perf/amd/ibs: Fix race condition in IBS Manali Shukla
                   ` (8 more replies)
  0 siblings, 9 replies; 20+ messages in thread
From: Manali Shukla @ 2026-03-10  6:00 UTC (permalink / raw)
  To: seanjc, pbonzini
  Cc: mingo, bp, kvm, x86, santosh.shukla, nikunj.dadhania, Naveen.Rao,
	dapeng1.mi, ravi.bangoria, peterz, Sandipan.Das

Add support for IBS virtualization (VIBS). VIBS feature allows the
guest to collect IBS samples without exiting the guest.  There are
2 parts to it [1].
 - Virtualizing the IBS register state.
 - Ensuring the IBS interrupt is handled in the guest without exiting
   the hypervisor.

To deliver virtualized IBS interrupts to the guest, VIBS requires either
AVIC or Virtual NMI (VNMI) support [1]. During IBS sampling, the
hardware signals a VNMI. The source of this VNMI depends on the AVIC
configuration:

 - With AVIC disabled, the virtual NMI is hardware-accelerated.
 - With AVIC enabled, the virtual NMI is delivered via AVIC using Extended
   LVT.

The local interrupts are extended to include more LVT registers, to
allow additional interrupt sources, like instruction based sampling
etc. [3].

Although IBS virtualization requires either AVIC or VNMI to be enabled
in order to successfully deliver IBS NMIs to the guest, VNMI must be
enabled to ensure reliable delivery. This requirement stems from the
dynamic behavior of AVIC (This is needed because AVIC can change its
state while the guest is running). While a guest is launched with AVIC
enabled, AVIC can be inhibited at runtime. When AVIC is inhibited and
VNMI is disabled, there is no mechanism to deliver IBS NMIs to the
guest. Therefore, enabling VNMI is necessary to support IBS
virtualization reliably.

Note that, since IBS registers are swap type C [2], the hypervisor is
responsible for saving and restoring of IBS host state. Hypervisor needs
to disable host IBS before saving the state and enter the guest. After a
guest exit, the hypervisor needs to restore host IBS state and re-enable
IBS.

The mediated PMU has the capability to save the host context when
entering the guest by scheduling out all exclude_guest events, and to
restore the host context when exiting the guest by scheduling in the
previously scheduled-out events. This behavior aligns with the
requirement for IBS registers being of swap type C. Therefore, the
mediated PMU design can be leveraged to implement IBS virtualization.

How to enable VIBS?
sudo echo 0 | sudo tee /proc/sys/kernel/nmi_watchdog
sudo modprobe -r kvm_amd
sudo modprobe kvm_amd enable_mediated_pmu=1 vnmi=1

Testing done:
- Following tests were executed on guest
  - Basic IBS testing was done on the guest from IBS cheatsheet [5].
  - perf_fuzzer was run for 12hrs, no softlockups or unknown NMIs were
    seen.

TO-DO:
Enable IBS virtualization on SEV-ES and SEV-SNP guests.
Enable IBS virtualization on nested guests.

Patches are rebased on top of
https://github.com/kvm-x86/linux next
commit 7d14d027e733 ("Documentation: KVM: Formalizing taking vcpu->mutex *outside* of kvm->slots_lock") + [5] + [6] + [7].

[1]: https://bugzilla.kernel.org/attachment.cgi?id=306250
     AMD64 Architecture Programmer’s Manual, Vol 2, Section 15.38
     Instruction-Based Sampling Virtualization.

[2]: https://bugzilla.kernel.org/attachment.cgi?id=306250
     AMD64 Architecture Programmer’s Manual, Vol 2, Appendix B Layout
     of VMCB, Table B-3 Swap Types.

[3]: https://bugzilla.kernel.org/attachment.cgi?id=306250
     AMD64 Architecture Programmer’s Manual, Vol 2, Section 16.4.5
     Extended Interrupts.

[4]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/tools/perf/Documentation/perf-amd-ibs.txt

[5]: https://lore.kernel.org/all/20260216042216.1440-1-ravi.bangoria@amd.com/

[6]: https://lore.kernel.org/all/20260216042530.1546-1-ravi.bangoria@amd.com/

[7]: https://lore.kernel.org/kvm/20260204074452.55453-1-manali.shukla@amd.com/

v2->v3
- Moved EXTLVT implementation to a different series
- Added virtualization for newly added IBS MSRs for future hardware
- Miscellaneous changes

v1->v2
- Incorporated review comments from Mi Dapeng
  - Change the name of kvm_lapic_state_w_extapic to kvm_ext_lapic_state.
  - Refactor APIC register mask handling in order to support extended
    APIC registers.
  - cpuid new leaf(CPUID_8000_001B) for IBS capabilities related code changes. 
  - Miscellaneous changes

---
v2: https://lore.kernel.org/kvm/20250901051656.209083-1-manali.shukla@amd.com/T/
v1: https://lore.kernel.org/kvm/afafc865-b42f-4a9d-82d7-a72de16bb47b@amd.com/T/

Manali Shukla (6):
  perf/amd/ibs: Fix race condition in IBS
  KVM: x86/cpuid: Add a KVM-only leaf for IBS capabilities
  KVM: x86: Extend CPUID range to include new leaf
  perf/x86/amd: Enable VPMU passthrough capability for IBS PMU
  perf/x86/amd: Remove exclude_guest check from perf_ibs_init()
  KVM: SVM: Add newly added IBS capabilities and MSRs

Santosh Shukla (3):
  x86/cpufeatures: Add CPUID feature bit for VIBS in SVM/SEV guests
  KVM: SVM: Extend VMCB area for virtualized IBS registers
  KVM: SVM: Add support for IBS Virtualization

 arch/x86/events/amd/ibs.c          |  8 +--
 arch/x86/include/asm/cpufeatures.h |  1 +
 arch/x86/include/asm/kvm_host.h    |  1 +
 arch/x86/include/asm/svm.h         | 18 ++++++-
 arch/x86/kvm/cpuid.c               | 31 ++++++++++++
 arch/x86/kvm/reverse_cpuid.h       | 22 ++++++++
 arch/x86/kvm/svm/svm.c             | 80 +++++++++++++++++++++++++++++-
 7 files changed, 156 insertions(+), 5 deletions(-)

-- 
2.43.0


^ permalink raw reply	[flat|nested] 20+ messages in thread

* [PATCH v3 1/9] perf/amd/ibs: Fix race condition in IBS
  2026-03-10  6:00 [PATCH v3 0/9] Implement support for IBS virtualization Manali Shukla
@ 2026-03-10  6:00 ` Manali Shukla
  2026-10-08 17:19   ` Jim Mattson
  2026-03-10  6:00 ` [PATCH v3 2/9] x86/cpufeatures: Add CPUID feature bit for VIBS in SVM/SEV guests Manali Shukla
                   ` (7 subsequent siblings)
  8 siblings, 1 reply; 20+ messages in thread
From: Manali Shukla @ 2026-03-10  6:00 UTC (permalink / raw)
  To: seanjc, pbonzini
  Cc: mingo, bp, kvm, x86, santosh.shukla, nikunj.dadhania, Naveen.Rao,
	dapeng1.mi, ravi.bangoria, peterz, Sandipan.Das

Consider the following scenario,

While scheduling out an IBS event from perf's core scheduling path,
event_sched_out() disables the IBS event by clearing the IBS enable
bit in perf_ibs_disable_event(). However, if a delayed IBS NMI is
delivered after the IBS enable bit is cleared, the IBS NMI handler
may still observe the valid bit set and incorrectly treat the sample
as valid. As a result, it re-enables IBS by setting the enable bit,
even though the event has already been scheduled out.

This leads to a situation where IBS is re-enabled after being
explicitly disabled, which is incorrect. Although this race does not
have visible side effects, it violates the expected behavior of the
perf subsystem.

The race is particularly noticeable when userspace repeatedly disables
and re-enables IBS using PERF_EVENT_IOC_DISABLE and
PERF_EVENT_IOC_ENABLE ioctls in a loop.

Fix this by checking the IBS_STOPPING bit in the IBS NMI handler before
re-enabling the IBS event. If the IBS_STOPPING bit is set, it indicates
that the event is either disabled or in the process of being disabled,
and the NMI handler should not re-enable it.

Signed-off-by: Manali Shukla <manali.shukla@amd.com>
---
 arch/x86/events/amd/ibs.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/arch/x86/events/amd/ibs.c b/arch/x86/events/amd/ibs.c
index eeb607b84dda..09b56bab510a 100644
--- a/arch/x86/events/amd/ibs.c
+++ b/arch/x86/events/amd/ibs.c
@@ -1582,7 +1582,8 @@ static int perf_ibs_handle_irq(struct perf_ibs *perf_ibs, struct pt_regs *iregs)
 		}
 		new_config |= period >> 4;
 
-		perf_ibs_enable_event(perf_ibs, hwc, new_config);
+		if (!test_bit(IBS_STOPPING, pcpu->state))
+			perf_ibs_enable_event(perf_ibs, hwc, new_config);
 	}
 
 	perf_event_update_userpage(event);
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v3 2/9] x86/cpufeatures: Add CPUID feature bit for VIBS in SVM/SEV guests
  2026-03-10  6:00 [PATCH v3 0/9] Implement support for IBS virtualization Manali Shukla
  2026-03-10  6:00 ` [PATCH v3 1/9] perf/amd/ibs: Fix race condition in IBS Manali Shukla
@ 2026-03-10  6:00 ` Manali Shukla
  2026-10-08 17:28   ` Jim Mattson
  2026-03-10  6:00 ` [PATCH v3 3/9] KVM: x86/cpuid: Add a KVM-only leaf for IBS capabilities Manali Shukla
                   ` (6 subsequent siblings)
  8 siblings, 1 reply; 20+ messages in thread
From: Manali Shukla @ 2026-03-10  6:00 UTC (permalink / raw)
  To: seanjc, pbonzini
  Cc: mingo, bp, kvm, x86, santosh.shukla, nikunj.dadhania, Naveen.Rao,
	dapeng1.mi, ravi.bangoria, peterz, Sandipan.Das

From: Santosh Shukla <santosh.shukla@amd.com>

The virtualized IBS (VIBS) feature allows the guest to collect IBS
samples without exiting the guest.

Presence of the VIBS feature is indicated via CPUID function
0x8000000A_EDX[26].

Signed-off-by: Santosh Shukla <santosh.shukla@amd.com>
Signed-off-by: Manali Shukla <manali.shukla@amd.com>
---
 arch/x86/include/asm/cpufeatures.h | 1 +
 1 file changed, 1 insertion(+)

diff --git a/arch/x86/include/asm/cpufeatures.h b/arch/x86/include/asm/cpufeatures.h
index b1631eb15e74..a1cd6437a052 100644
--- a/arch/x86/include/asm/cpufeatures.h
+++ b/arch/x86/include/asm/cpufeatures.h
@@ -382,6 +382,7 @@
 #define X86_FEATURE_X2AVIC		(15*32+18) /* "x2avic" Virtual x2apic */
 #define X86_FEATURE_V_SPEC_CTRL		(15*32+20) /* "v_spec_ctrl" Virtual SPEC_CTRL */
 #define X86_FEATURE_VNMI		(15*32+25) /* "vnmi" Virtual NMI */
+#define X86_FEATURE_VIBS		(15*32+26) /* Virtual IBS */
 #define X86_FEATURE_AVIC_EXTLVT		(15*32+27) /* Extended LVT AVIC acceleration support */
 #define X86_FEATURE_SVME_ADDR_CHK	(15*32+28) /* SVME addr check */
 #define X86_FEATURE_BUS_LOCK_THRESHOLD	(15*32+29) /* Bus lock threshold */
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v3 3/9] KVM: x86/cpuid: Add a KVM-only leaf for IBS capabilities
  2026-03-10  6:00 [PATCH v3 0/9] Implement support for IBS virtualization Manali Shukla
  2026-03-10  6:00 ` [PATCH v3 1/9] perf/amd/ibs: Fix race condition in IBS Manali Shukla
  2026-03-10  6:00 ` [PATCH v3 2/9] x86/cpufeatures: Add CPUID feature bit for VIBS in SVM/SEV guests Manali Shukla
@ 2026-03-10  6:00 ` Manali Shukla
  2026-10-08 17:44   ` Jim Mattson
  2026-03-10  6:00 ` [PATCH v3 4/9] KVM: x86: Extend CPUID range to include new leaf Manali Shukla
                   ` (5 subsequent siblings)
  8 siblings, 1 reply; 20+ messages in thread
From: Manali Shukla @ 2026-03-10  6:00 UTC (permalink / raw)
  To: seanjc, pbonzini
  Cc: mingo, bp, kvm, x86, santosh.shukla, nikunj.dadhania, Naveen.Rao,
	dapeng1.mi, ravi.bangoria, peterz, Sandipan.Das

Add a KVM-only leaf for AMD's Instruction Based Sampling capabilities.
Multiple IBS related capabilities are added to KVM-only leaf, so that KVM
can set these capabilities for the guest, when IBS feature bit is
enabled on the guest.

Signed-off-by: Manali Shukla <manali.shukla@amd.com>
---
 arch/x86/include/asm/kvm_host.h |  1 +
 arch/x86/kvm/reverse_cpuid.h    | 16 ++++++++++++++++
 2 files changed, 17 insertions(+)

diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
index 32dd2d55e6f0..01abdf7f112b 100644
--- a/arch/x86/include/asm/kvm_host.h
+++ b/arch/x86/include/asm/kvm_host.h
@@ -782,6 +782,7 @@ enum kvm_only_cpuid_leafs {
 	CPUID_12_EAX	 = NCAPINTS,
 	CPUID_7_1_EDX,
 	CPUID_8000_0007_EDX,
+	CPUID_8000_001B_EAX,
 	CPUID_8000_0022_EAX,
 	CPUID_7_2_EDX,
 	CPUID_24_0_EBX,
diff --git a/arch/x86/kvm/reverse_cpuid.h b/arch/x86/kvm/reverse_cpuid.h
index 657f5f743ed9..22cfdb331e9e 100644
--- a/arch/x86/kvm/reverse_cpuid.h
+++ b/arch/x86/kvm/reverse_cpuid.h
@@ -76,6 +76,21 @@
 #define KVM_X86_FEATURE_TSA_SQ_NO	KVM_X86_FEATURE(CPUID_8000_0021_ECX, 1)
 #define KVM_X86_FEATURE_TSA_L1_NO	KVM_X86_FEATURE(CPUID_8000_0021_ECX, 2)
 
+/* AMD defined Instruction-base Sampling capabilities. CPUID level 0x8000001B (EAX). */
+#define X86_FEATURE_IBS_AVAIL			KVM_X86_FEATURE(CPUID_8000_001B_EAX, 0)
+#define X86_FEATURE_IBS_FETCHSAM		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 1)
+#define X86_FEATURE_IBS_OPSAM			KVM_X86_FEATURE(CPUID_8000_001B_EAX, 2)
+#define X86_FEATURE_IBS_RDWROPCNT		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 3)
+#define X86_FEATURE_IBS_OPCNT			KVM_X86_FEATURE(CPUID_8000_001B_EAX, 4)
+#define X86_FEATURE_IBS_BRNTRGT			KVM_X86_FEATURE(CPUID_8000_001B_EAX, 5)
+#define X86_FEATURE_IBS_OPCNTEXT		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 6)
+#define X86_FEATURE_IBS_RIPINVALIDCHK		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 7)
+#define X86_FEATURE_IBS_OPBRNFUSE		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 8)
+#define X86_FEATURE_IBS_FETCHCTLEXTD		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 9)
+#define X86_FEATURE_IBS_ZEN4_EXT		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 11)
+#define X86_FEATURE_IBS_LOADLATFIL		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 12)
+#define X86_FEATURE_IBS_ZEN4_DTLBSTAT		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 19)
+
 struct cpuid_reg {
 	u32 function;
 	u32 index;
@@ -105,6 +120,7 @@ static const struct cpuid_reg reverse_cpuid[] = {
 	[CPUID_8000_0022_EAX] = {0x80000022, 0, CPUID_EAX},
 	[CPUID_7_2_EDX]       = {         7, 2, CPUID_EDX},
 	[CPUID_24_0_EBX]      = {      0x24, 0, CPUID_EBX},
+	[CPUID_8000_001B_EAX] = {0x8000001b, 0, CPUID_EAX},
 	[CPUID_8000_0021_ECX] = {0x80000021, 0, CPUID_ECX},
 	[CPUID_7_1_ECX]       = {         7, 1, CPUID_ECX},
 	[CPUID_1E_1_EAX]      = {      0x1e, 1, CPUID_EAX},
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v3 4/9] KVM: x86: Extend CPUID range to include new leaf
  2026-03-10  6:00 [PATCH v3 0/9] Implement support for IBS virtualization Manali Shukla
                   ` (2 preceding siblings ...)
  2026-03-10  6:00 ` [PATCH v3 3/9] KVM: x86/cpuid: Add a KVM-only leaf for IBS capabilities Manali Shukla
@ 2026-03-10  6:00 ` Manali Shukla
  2026-10-08 17:59   ` Jim Mattson
  2026-03-10  6:00 ` [PATCH v3 5/9] KVM: SVM: Extend VMCB area for virtualized IBS registers Manali Shukla
                   ` (4 subsequent siblings)
  8 siblings, 1 reply; 20+ messages in thread
From: Manali Shukla @ 2026-03-10  6:00 UTC (permalink / raw)
  To: seanjc, pbonzini
  Cc: mingo, bp, kvm, x86, santosh.shukla, nikunj.dadhania, Naveen.Rao,
	dapeng1.mi, ravi.bangoria, peterz, Sandipan.Das

CPUID leaf 0x8000001b (EAX) provides information about Instruction-Based
sampling capabilities on AMD Platforms. Add the new leaf to
kvm_cpu_cap_init() using F() macros, which automatically gate each
capability bits against raw hardware CPUID via raw_cpuid_get().

This allows vendor code to simply clear entire leaf when vibs is not
enabled, rather than reading hardware CPUID and calling
kvm_cpu_cap_set() for each capability bits inidividually in later
patches.

Signed-off-by: Manali Shukla <manali.shukla@amd.com>
---
 arch/x86/kvm/cpuid.c | 25 +++++++++++++++++++++++++
 1 file changed, 25 insertions(+)

diff --git a/arch/x86/kvm/cpuid.c b/arch/x86/kvm/cpuid.c
index 96a08a556543..4e626e77e6a6 100644
--- a/arch/x86/kvm/cpuid.c
+++ b/arch/x86/kvm/cpuid.c
@@ -1226,6 +1226,22 @@ void kvm_initialize_cpu_caps(void)
 		VENDOR_F(SVME_ADDR_CHK),
 	);
 
+	kvm_cpu_cap_init(CPUID_8000_001B_EAX,
+		F(IBS_AVAIL),
+		F(IBS_FETCHSAM),
+		F(IBS_OPSAM),
+		F(IBS_RDWROPCNT),
+		F(IBS_OPCNT),
+		F(IBS_BRNTRGT),
+		F(IBS_OPCNTEXT),
+		F(IBS_RIPINVALIDCHK),
+		F(IBS_OPBRNFUSE),
+		F(IBS_FETCHCTLEXTD),
+		F(IBS_ZEN4_EXT),
+		F(IBS_LOADLATFIL),
+		F(IBS_ZEN4_DTLBSTAT),
+	);
+
 	kvm_cpu_cap_init(CPUID_8000_001F_EAX,
 		VENDOR_F(SME),
 		VENDOR_F(SEV),
@@ -1848,6 +1864,15 @@ static inline int __do_cpuid_func(struct kvm_cpuid_array *array, u32 function)
 		entry->eax = entry->ebx = entry->ecx = 0;
 		entry->edx = 0; /* reserved */
 		break;
+	/* AMD IBS capability */
+	case 0x8000001B:
+		if (!kvm_cpu_cap_has(X86_FEATURE_IBS))
+			entry->eax = 0;
+		else
+			cpuid_entry_override(entry, CPUID_8000_001B_EAX);
+
+		entry->ebx = entry->ecx = entry->edx = 0;
+		break;
 	case 0x8000001F:
 		if (!kvm_cpu_cap_has(X86_FEATURE_SEV)) {
 			entry->eax = entry->ebx = entry->ecx = entry->edx = 0;
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v3 5/9] KVM: SVM: Extend VMCB area for virtualized IBS registers
  2026-03-10  6:00 [PATCH v3 0/9] Implement support for IBS virtualization Manali Shukla
                   ` (3 preceding siblings ...)
  2026-03-10  6:00 ` [PATCH v3 4/9] KVM: x86: Extend CPUID range to include new leaf Manali Shukla
@ 2026-03-10  6:00 ` Manali Shukla
  2026-10-08 18:01   ` Jim Mattson
  2026-03-10  6:00 ` [PATCH v3 6/9] KVM: SVM: Add support for IBS Virtualization Manali Shukla
                   ` (3 subsequent siblings)
  8 siblings, 1 reply; 20+ messages in thread
From: Manali Shukla @ 2026-03-10  6:00 UTC (permalink / raw)
  To: seanjc, pbonzini
  Cc: mingo, bp, kvm, x86, santosh.shukla, nikunj.dadhania, Naveen.Rao,
	dapeng1.mi, ravi.bangoria, peterz, Sandipan.Das

From: Santosh Shukla <santosh.shukla@amd.com>

Define the new VMCB fields that will be used to save and restore the
state of the following fetch and op IBS related MSRs.

  * MSRC001_1030 [IBS Fetch Control]
  * MSRC001_1031 [IBS Fetch Linear Address]
  * MSRC001_1033 [IBS Execution Control]
  * MSRC001_1034 [IBS Op Logical Address]
  * MSRC001_1035 [IBS Op Data]
  * MSRC001_1036 [IBS Op Data 2]
  * MSRC001_1037 [IBS Op Data 3]
  * MSRC001_1038 [IBS DC Linear Address]
  * MSRC001_103B [IBS Branch Target Address]
  * MSRC001_103C [IBS Fetch Control Extended]

Signed-off-by: Santosh Shukla <santosh.shukla@amd.com>
Signed-off-by: Manali Shukla <manali.shukla@amd.com>
---
 arch/x86/include/asm/svm.h | 14 +++++++++++++-
 1 file changed, 13 insertions(+), 1 deletion(-)

diff --git a/arch/x86/include/asm/svm.h b/arch/x86/include/asm/svm.h
index bcfeb5e7c0ed..4296efc1dafe 100644
--- a/arch/x86/include/asm/svm.h
+++ b/arch/x86/include/asm/svm.h
@@ -369,6 +369,17 @@ struct vmcb_save_area {
 	u64 last_excp_to;
 	u8 reserved_0x298[72];
 	u64 spec_ctrl;		/* Guest version of SPEC_CTRL at 0x2E0 */
+	u8 reserved_0x2e8[1168];
+	u64 ibs_fetch_ctl;
+	u64 ibs_fetch_linear_addr;
+	u64 ibs_op_ctl;
+	u64 ibs_op_rip;
+	u64 ibs_op_data;
+	u64 ibs_op_data2;
+	u64 ibs_op_data3;
+	u64 ibs_dc_linear_addr;
+	u64 ibs_br_target;
+	u64 ibs_fetch_extd_ctl;
 } __packed;
 
 /* Save area definition for SEV-ES and SEV-SNP guests */
@@ -551,7 +562,7 @@ struct vmcb {
 	};
 } __packed;
 
-#define EXPECTED_VMCB_SAVE_AREA_SIZE		744
+#define EXPECTED_VMCB_SAVE_AREA_SIZE		1992
 #define EXPECTED_GHCB_SAVE_AREA_SIZE		1032
 #define EXPECTED_SEV_ES_SAVE_AREA_SIZE		1648
 #define EXPECTED_VMCB_CONTROL_AREA_SIZE		1024
@@ -577,6 +588,7 @@ static inline void __unused_size_checks(void)
 	BUILD_BUG_RESERVED_OFFSET(vmcb_save_area, 0x180);
 	BUILD_BUG_RESERVED_OFFSET(vmcb_save_area, 0x248);
 	BUILD_BUG_RESERVED_OFFSET(vmcb_save_area, 0x298);
+	BUILD_BUG_RESERVED_OFFSET(vmcb_save_area, 0x2e8);
 
 	BUILD_BUG_RESERVED_OFFSET(sev_es_save_area, 0xc8);
 	BUILD_BUG_RESERVED_OFFSET(sev_es_save_area, 0xcc);
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v3 6/9] KVM: SVM: Add support for IBS Virtualization
  2026-03-10  6:00 [PATCH v3 0/9] Implement support for IBS virtualization Manali Shukla
                   ` (4 preceding siblings ...)
  2026-03-10  6:00 ` [PATCH v3 5/9] KVM: SVM: Extend VMCB area for virtualized IBS registers Manali Shukla
@ 2026-03-10  6:00 ` Manali Shukla
  2026-10-08 19:19   ` Jim Mattson
  2026-03-10  6:00 ` [PATCH v3 7/9] perf/x86/amd: Enable VPMU passthrough capability for IBS PMU Manali Shukla
                   ` (2 subsequent siblings)
  8 siblings, 1 reply; 20+ messages in thread
From: Manali Shukla @ 2026-03-10  6:00 UTC (permalink / raw)
  To: seanjc, pbonzini
  Cc: mingo, bp, kvm, x86, santosh.shukla, nikunj.dadhania, Naveen.Rao,
	dapeng1.mi, ravi.bangoria, peterz, Sandipan.Das

From: Santosh Shukla <santosh.shukla@amd.com>

IBS virtualization (VIBS) allows a guest to collect Instruction-Based
Sampling (IBS) data using hardware-assisted virtualization. With VIBS
enabled, the hardware automatically saves and restores guest IBS state
during VM-Entry and VM-Exit via the VMCB State Save Area.

IBS-generated interrupts are delivered directly to the guest without
causing a VMEXIT.

VIBS depends on mediated PMU mode and requires either AVIC or NMI
virtualization for interrupt delivery. However, since AVIC can be
dynamically inhibited, VIBS requires VNMI to be enabled to ensure
reliable interrupt delivery. If AVIC is inhibited and VNMI is
disabled, the guest can encounter a VMEXIT_INVALID when IBS
virtualization is enabled for the guest.

Because IBS state is classified as swap type C, the hypervisor must
save its own IBS state before VMRUN and restore it after VMEXIT. It
must also disable IBS before VMRUN and re-enable it afterward. This
will be handled using mediated PMU support in subsequent patches by
enabling mediated PMU capability for IBS PMUs.

More details about IBS virtualization can be found at [1].

[1]: https://bugzilla.kernel.org/attachment.cgi?id=306250
     AMD64 Architecture Programmer’s Manual, Vol 2, Section 15.38
     Instruction-Based Sampling Virtualization.

Signed-off-by: Santosh Shukla <santosh.shukla@amd.com>
Co-developed-by: Manali Shukla <manali.shukla@amd.com>
Signed-off-by: Manali Shukla <manali.shukla@amd.com>
---
 arch/x86/include/asm/svm.h |  2 ++
 arch/x86/kvm/svm/svm.c     | 73 +++++++++++++++++++++++++++++++++++++-
 2 files changed, 74 insertions(+), 1 deletion(-)

diff --git a/arch/x86/include/asm/svm.h b/arch/x86/include/asm/svm.h
index 4296efc1dafe..17aa6bf76bce 100644
--- a/arch/x86/include/asm/svm.h
+++ b/arch/x86/include/asm/svm.h
@@ -226,6 +226,8 @@ struct __attribute__ ((__packed__)) vmcb_control_area {
 
 #define SVM_INT_VECTOR_MASK GENMASK(7, 0)
 
+#define SVM_MISC_ENABLE_V_IBS BIT_ULL(2)
+
 #define SVM_INTERRUPT_SHADOW_MASK	BIT_ULL(0)
 #define SVM_GUEST_INTERRUPT_MASK	BIT_ULL(1)
 
diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c
index 5af3479cd264..421a929398da 100644
--- a/arch/x86/kvm/svm/svm.c
+++ b/arch/x86/kvm/svm/svm.c
@@ -161,12 +161,15 @@ module_param(lbrv, int, 0444);
 static int __ro_after_init tsc_scaling = true;
 module_param(tsc_scaling, int, 0444);
 
+/* enable/disable IBS virtualization */
+static bool __ro_after_init vibs = true;
+module_param(vibs, bool, 0444);
+
 module_param(enable_device_posted_irqs, bool, 0444);
 
 bool __read_mostly dump_invalid_vmcb;
 module_param(dump_invalid_vmcb, bool, 0644);
 
-
 bool __ro_after_init intercept_smi = true;
 module_param(intercept_smi, bool, 0444);
 
@@ -779,6 +782,26 @@ static void svm_recalc_pmu_msr_intercepts(struct kvm_vcpu *vcpu)
 				  MSR_TYPE_RW, intercept);
 }
 
+static void svm_recalc_ibs_msr_intercepts(struct kvm_vcpu *vcpu)
+{
+	bool intercept = !(guest_cpu_cap_has(vcpu, X86_FEATURE_IBS) &&
+			   kvm_vcpu_has_mediated_pmu(vcpu));
+
+	if (!enable_mediated_pmu || !vibs)
+		return;
+
+	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSFETCHCTL, MSR_TYPE_RW, intercept);
+	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSFETCHLINAD, MSR_TYPE_RW, intercept);
+	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSOPCTL, MSR_TYPE_RW, intercept);
+	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSOPRIP, MSR_TYPE_RW, intercept);
+	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSOPDATA, MSR_TYPE_RW, intercept);
+	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSOPDATA2, MSR_TYPE_RW, intercept);
+	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSOPDATA3, MSR_TYPE_RW, intercept);
+	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSDCLINAD, MSR_TYPE_RW, intercept);
+	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSBRTARGET, MSR_TYPE_RW, intercept);
+	svm_set_intercept_for_msr(vcpu, MSR_AMD64_ICIBSEXTDCTL, MSR_TYPE_RW, intercept);
+}
+
 static void svm_recalc_msr_intercepts(struct kvm_vcpu *vcpu)
 {
 	struct vcpu_svm *svm = to_svm(vcpu);
@@ -848,6 +871,7 @@ static void svm_recalc_msr_intercepts(struct kvm_vcpu *vcpu)
 		sev_es_recalc_msr_intercepts(vcpu);
 
 	svm_recalc_pmu_msr_intercepts(vcpu);
+	svm_recalc_ibs_msr_intercepts(vcpu);
 
 	/*
 	 * x2APIC intercepts are modified on-demand and cannot be filtered by
@@ -2880,6 +2904,27 @@ static int svm_get_msr(struct kvm_vcpu *vcpu, struct msr_data *msr_info)
 	case MSR_AMD64_DE_CFG:
 		msr_info->data = svm->msr_decfg;
 		break;
+
+	case MSR_AMD64_IBSCTL:
+		if (guest_cpu_cap_has(vcpu, X86_FEATURE_IBS))
+			msr_info->data = IBSCTL_LVT_OFFSET_VALID;
+		else
+			msr_info->data = 0;
+		break;
+
+
+	/*
+	 * When IBS virtualization is enabled, guest reads from
+	 * MSR_AMD64_IBSFETCHPHYSAD and MSR_AMD64_IBSDCPHYSAD must return 0.
+	 * This is done for security reasons, as guests should not be allowed to
+	 * access or infer any information about the system's physical
+	 * addresses.
+	 */
+	case MSR_AMD64_IBSDCPHYSAD:
+	case MSR_AMD64_IBSFETCHPHYSAD:
+		msr_info->data = 0;
+		break;
+
 	default:
 		return kvm_get_msr_common(vcpu, msr_info);
 	}
@@ -3171,6 +3216,16 @@ static int svm_set_msr(struct kvm_vcpu *vcpu, struct msr_data *msr)
 		svm->msr_decfg = data;
 		break;
 	}
+	/*
+	 * When IBS virtualization is enabled, guest writes to
+	 * MSR_AMD64_IBSFETCHPHYSAD and MSR_AMD64_IBSDCPHYSAD must be ignored.
+	 * This is done for security reasons, as guests should not be allowed to
+	 * access or infer any information about the system's physical
+	 * addresses.
+	 */
+	case MSR_AMD64_IBSDCPHYSAD:
+	case MSR_AMD64_IBSFETCHPHYSAD:
+		return 1;
 	default:
 		return kvm_set_msr_common(vcpu, msr);
 	}
@@ -4678,6 +4733,11 @@ static void svm_vcpu_after_set_cpuid(struct kvm_vcpu *vcpu)
 	if (guest_cpuid_is_intel_compatible(vcpu))
 		guest_cpu_cap_clear(vcpu, X86_FEATURE_V_VMSAVE_VMLOAD);
 
+	if (guest_cpu_cap_has(vcpu, X86_FEATURE_IBS))
+		svm->vmcb->control.misc_ctl2 |= SVM_MISC_ENABLE_V_IBS;
+	else
+		svm->vmcb->control.misc_ctl2 &= ~SVM_MISC_ENABLE_V_IBS;
+
 	if (sev_guest(vcpu->kvm))
 		sev_vcpu_after_set_cpuid(svm);
 }
@@ -5510,6 +5570,11 @@ static __init void svm_set_cpu_caps(void)
 	if (cpu_feature_enabled(X86_FEATURE_EXTAPIC))
 		kvm_caps.has_extapic = true;
 
+	if (vibs)
+		kvm_cpu_cap_check_and_set(X86_FEATURE_IBS);
+	else
+		kvm_cpu_caps[CPUID_8000_001B_EAX] = 0;
+
 	/* CPUID 0x80000008 */
 	if (boot_cpu_has(X86_FEATURE_LS_CFG_SSBD) ||
 	    boot_cpu_has(X86_FEATURE_AMD_SSBD))
@@ -5698,6 +5763,12 @@ static __init int svm_hardware_setup(void)
 		svm_x86_ops.set_vnmi_pending = NULL;
 	}
 
+	vibs = enable_mediated_pmu && vnmi && vibs
+		&& boot_cpu_has(X86_FEATURE_VIBS);
+
+	if (vibs)
+		pr_info("IBS virtualization supported\n");
+
 	if (!enable_pmu)
 		pr_info("PMU virtualization is disabled\n");
 
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v3 7/9] perf/x86/amd: Enable VPMU passthrough capability for IBS PMU
  2026-03-10  6:00 [PATCH v3 0/9] Implement support for IBS virtualization Manali Shukla
                   ` (5 preceding siblings ...)
  2026-03-10  6:00 ` [PATCH v3 6/9] KVM: SVM: Add support for IBS Virtualization Manali Shukla
@ 2026-03-10  6:00 ` Manali Shukla
  2026-10-08 19:32   ` Jim Mattson
  2026-03-10  6:00 ` [PATCH v3 8/9] perf/x86/amd: Remove exclude_guest check from perf_ibs_init() Manali Shukla
  2026-03-10  6:00 ` [PATCH v3 9/9] KVM: SVM: Add newly added IBS capabilities and MSRs Manali Shukla
  8 siblings, 1 reply; 20+ messages in thread
From: Manali Shukla @ 2026-03-10  6:00 UTC (permalink / raw)
  To: seanjc, pbonzini
  Cc: mingo, bp, kvm, x86, santosh.shukla, nikunj.dadhania, Naveen.Rao,
	dapeng1.mi, ravi.bangoria, peterz, Sandipan.Das

IBS MSRs are classified as Swap Type C, which requires the hypervisor
to save and restore its own IBS state before VMENTRY and after VMEXIT.

To support this, set the ibs_op and ibs_fetch PMUs with the
PERF_PMU_CAP_MEDIATED_VPMU capability. This ensures that these PMUs are
exclusively owned by the guest while it is running, allowing the
hypervisor to manage IBS state transitions correctly.

Signed-off-by: Manali Shukla <manali.shukla@amd.com>
---
 arch/x86/events/amd/ibs.c | 2 ++
 1 file changed, 2 insertions(+)

diff --git a/arch/x86/events/amd/ibs.c b/arch/x86/events/amd/ibs.c
index 09b56bab510a..034a992bbfe3 100644
--- a/arch/x86/events/amd/ibs.c
+++ b/arch/x86/events/amd/ibs.c
@@ -971,6 +971,7 @@ static struct perf_ibs perf_ibs_fetch = {
 		.stop		= perf_ibs_stop,
 		.read		= perf_ibs_read,
 		.check_period	= perf_ibs_check_period,
+		.capabilities	= PERF_PMU_CAP_MEDIATED_VPMU,
 	},
 	.msr			= MSR_AMD64_IBSFETCHCTL,
 	.msr2			= MSR_AMD64_IBSFETCHCTL2,
@@ -997,6 +998,7 @@ static struct perf_ibs perf_ibs_op = {
 		.stop		= perf_ibs_stop,
 		.read		= perf_ibs_read,
 		.check_period	= perf_ibs_check_period,
+		.capabilities	= PERF_PMU_CAP_MEDIATED_VPMU,
 	},
 	.msr			= MSR_AMD64_IBSOPCTL,
 	.msr2			= MSR_AMD64_IBSOPCTL2,
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v3 8/9] perf/x86/amd: Remove exclude_guest check from perf_ibs_init()
  2026-03-10  6:00 [PATCH v3 0/9] Implement support for IBS virtualization Manali Shukla
                   ` (6 preceding siblings ...)
  2026-03-10  6:00 ` [PATCH v3 7/9] perf/x86/amd: Enable VPMU passthrough capability for IBS PMU Manali Shukla
@ 2026-03-10  6:00 ` Manali Shukla
  2026-10-08 19:38   ` Jim Mattson
  2026-03-10  6:00 ` [PATCH v3 9/9] KVM: SVM: Add newly added IBS capabilities and MSRs Manali Shukla
  8 siblings, 1 reply; 20+ messages in thread
From: Manali Shukla @ 2026-03-10  6:00 UTC (permalink / raw)
  To: seanjc, pbonzini
  Cc: mingo, bp, kvm, x86, santosh.shukla, nikunj.dadhania, Naveen.Rao,
	dapeng1.mi, ravi.bangoria, peterz, Sandipan.Das

Currently IBS driver doesn't allow the creation of IBS event with
exclude_guest set. As a result, amd_ibs_init() returns -EINVAL if
IBS event is created with exclude_guest set.

With the introduction of mediated PMU support, software-based handling
of exclude_guest is permitted for PMUs that have the
PERF_PMU_CAP_MEDIATED_VPMU capability.

Since ibs_op and ibs_fetch pmus has PERF_PMU_CAP_MEDIATED_VPMU
capability set, update perf_ibs_init() to remove exclude_guest check.

Signed-off-by: Manali Shukla <manali.shukla@amd.com>
---
 arch/x86/events/amd/ibs.c | 3 +--
 1 file changed, 1 insertion(+), 2 deletions(-)

diff --git a/arch/x86/events/amd/ibs.c b/arch/x86/events/amd/ibs.c
index 034a992bbfe3..7da06c143b32 100644
--- a/arch/x86/events/amd/ibs.c
+++ b/arch/x86/events/amd/ibs.c
@@ -327,8 +327,7 @@ static int perf_ibs_init(struct perf_event *event)
 		return -EOPNOTSUPP;
 
 	/* handle exclude_{user,kernel} in the IRQ handler */
-	if (event->attr.exclude_host || event->attr.exclude_guest ||
-	    event->attr.exclude_idle)
+	if (event->attr.exclude_host || event->attr.exclude_idle)
 		return -EINVAL;
 
 	ret = validate_group(event);
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v3 9/9] KVM: SVM: Add newly added IBS capabilities and MSRs
  2026-03-10  6:00 [PATCH v3 0/9] Implement support for IBS virtualization Manali Shukla
                   ` (7 preceding siblings ...)
  2026-03-10  6:00 ` [PATCH v3 8/9] perf/x86/amd: Remove exclude_guest check from perf_ibs_init() Manali Shukla
@ 2026-03-10  6:00 ` Manali Shukla
  2026-10-08 21:03   ` Jim Mattson
  8 siblings, 1 reply; 20+ messages in thread
From: Manali Shukla @ 2026-03-10  6:00 UTC (permalink / raw)
  To: seanjc, pbonzini
  Cc: mingo, bp, kvm, x86, santosh.shukla, nikunj.dadhania, Naveen.Rao,
	dapeng1.mi, ravi.bangoria, peterz, Sandipan.Das

IBS on upcoming microarch introduced two new control MSRs and a couple of
new features. Define macros for them. Add these newly added IBS
capabilities to KVM-only leaf 0x8000001b, so that when IBS feature bit
is enabled on the guest, these newly added features can be used by
guests if the hardware and guest os supports it.

 - X86_FEATURE_IBS_DISABLE: Independent IBS disable capability to
   avoid RMW race
 - X86_FEATURE_IBS_FETCHLATFIL: Fetch Latency filtering
 - X86_FEATURE_IBS_ADDRFILTER: Address Bit 63 based filtering
 - X86_FEATURE_IBS_STRMST_RMTSOCKET: Streaming store filter and
   indicator. Remote socket indicator.
 - X86_FEATURE_IBS_BUFFER1: IBS buffering v1
 - X86_FEATURE_IBS_MEMPROFILER: IBS memory profiler

Extend VMCB save area to include to the newly added MSRs:
MSR_AMD64_IBSFETCHCTL2 and MSR_AMD64_IBSOPCTL2.

Signed-off-by: Manali Shukla <manali.shukla@amd.com>
---
 arch/x86/include/asm/svm.h   | 4 +++-
 arch/x86/kvm/cpuid.c         | 6 ++++++
 arch/x86/kvm/reverse_cpuid.h | 6 ++++++
 arch/x86/kvm/svm/svm.c       | 7 +++++++
 4 files changed, 22 insertions(+), 1 deletion(-)

diff --git a/arch/x86/include/asm/svm.h b/arch/x86/include/asm/svm.h
index 17aa6bf76bce..88833db2e739 100644
--- a/arch/x86/include/asm/svm.h
+++ b/arch/x86/include/asm/svm.h
@@ -382,6 +382,8 @@ struct vmcb_save_area {
 	u64 ibs_dc_linear_addr;
 	u64 ibs_br_target;
 	u64 ibs_fetch_extd_ctl;
+	u64 ibs_fetch_ctl2;
+	u64 ibs_op_ctl2;
 } __packed;
 
 /* Save area definition for SEV-ES and SEV-SNP guests */
@@ -564,7 +566,7 @@ struct vmcb {
 	};
 } __packed;
 
-#define EXPECTED_VMCB_SAVE_AREA_SIZE		1992
+#define EXPECTED_VMCB_SAVE_AREA_SIZE		2008
 #define EXPECTED_GHCB_SAVE_AREA_SIZE		1032
 #define EXPECTED_SEV_ES_SAVE_AREA_SIZE		1648
 #define EXPECTED_VMCB_CONTROL_AREA_SIZE		1024
diff --git a/arch/x86/kvm/cpuid.c b/arch/x86/kvm/cpuid.c
index 4e626e77e6a6..e8a664cb0bb8 100644
--- a/arch/x86/kvm/cpuid.c
+++ b/arch/x86/kvm/cpuid.c
@@ -1239,6 +1239,12 @@ void kvm_initialize_cpu_caps(void)
 		F(IBS_FETCHCTLEXTD),
 		F(IBS_ZEN4_EXT),
 		F(IBS_LOADLATFIL),
+		F(IBS_DISABLE),
+		F(IBS_FETCHLATFIL),
+		F(IBS_ADDRFILTER),
+		F(IBS_STRMST_RMTSOCKET),
+		F(IBS_BUFFER1),
+		F(IBS_MEMPROFILER),
 		F(IBS_ZEN4_DTLBSTAT),
 	);
 
diff --git a/arch/x86/kvm/reverse_cpuid.h b/arch/x86/kvm/reverse_cpuid.h
index 22cfdb331e9e..1af2ba207b8a 100644
--- a/arch/x86/kvm/reverse_cpuid.h
+++ b/arch/x86/kvm/reverse_cpuid.h
@@ -89,6 +89,12 @@
 #define X86_FEATURE_IBS_FETCHCTLEXTD		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 9)
 #define X86_FEATURE_IBS_ZEN4_EXT		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 11)
 #define X86_FEATURE_IBS_LOADLATFIL		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 12)
+#define X86_FEATURE_IBS_DISABLE			KVM_X86_FEATURE(CPUID_8000_001B_EAX, 13)
+#define X86_FEATURE_IBS_FETCHLATFIL		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 14)
+#define X86_FEATURE_IBS_ADDRFILTER		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 15)
+#define X86_FEATURE_IBS_STRMST_RMTSOCKET	KVM_X86_FEATURE(CPUID_8000_001B_EAX, 16)
+#define X86_FEATURE_IBS_BUFFER1			KVM_X86_FEATURE(CPUID_8000_001B_EAX, 17)
+#define X86_FEATURE_IBS_MEMPROFILER		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 18)
 #define X86_FEATURE_IBS_ZEN4_DTLBSTAT		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 19)
 
 struct cpuid_reg {
diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c
index 421a929398da..9bf0d5f66239 100644
--- a/arch/x86/kvm/svm/svm.c
+++ b/arch/x86/kvm/svm/svm.c
@@ -800,6 +800,13 @@ static void svm_recalc_ibs_msr_intercepts(struct kvm_vcpu *vcpu)
 	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSDCLINAD, MSR_TYPE_RW, intercept);
 	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSBRTARGET, MSR_TYPE_RW, intercept);
 	svm_set_intercept_for_msr(vcpu, MSR_AMD64_ICIBSEXTDCTL, MSR_TYPE_RW, intercept);
+
+	if (guest_cpu_cap_has(vcpu, X86_FEATURE_IBS_DISABLE)) {
+		svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSFETCHCTL2, MSR_TYPE_RW,
+				intercept);
+		svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSOPCTL2, MSR_TYPE_RW,
+				intercept);
+	}
 }
 
 static void svm_recalc_msr_intercepts(struct kvm_vcpu *vcpu)
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* Re: [PATCH v3 1/9] perf/amd/ibs: Fix race condition in IBS
  2026-03-10  6:00 ` [PATCH v3 1/9] perf/amd/ibs: Fix race condition in IBS Manali Shukla
@ 2026-10-08 17:19   ` Jim Mattson
  0 siblings, 0 replies; 20+ messages in thread
From: Jim Mattson @ 2026-10-08 17:19 UTC (permalink / raw)
  To: Manali Shukla
  Cc: Jim Mattson, seanjc, pbonzini, mingo, bp, kvm, x86,
	santosh.shukla, nikunj.dadhania, Naveen.Rao, dapeng1.mi,
	ravi.bangoria, peterz, Sandipan.Das, Yosry Ahmed,
	linux-perf-users

On Tue, Mar 10, 2026 at 06:00:13AM +0000, Manali Shukla wrote:
> Consider the following scenario,
>
> While scheduling out an IBS event from perf's core scheduling path,
> event_sched_out() disables the IBS event by clearing the IBS enable
> bit in perf_ibs_disable_event(). However, if a delayed IBS NMI is
> delivered after the IBS enable bit is cleared, the IBS NMI handler
> may still observe the valid bit set and incorrectly treat the sample
> as valid.

The sample is valid. It was collected while the event was active, and
it is correct to record it. The bug is only that the handler re-arms
the hardware after perf_ibs_stop() disables the event.

> As a result, it re-enables IBS by setting the enable bit,
> even though the event has already been scheduled out.
>
> This leads to a situation where IBS is re-enabled after being
> explicitly disabled, which is incorrect. Although this race does not
> have visible side effects, it violates the expected behavior of the
> perf subsystem.

This race does have visible side effects:

  1. When the delayed NMI arrives before perf_ibs_stop() clears
     IBS_STARTED, the handler takes the normal path (not the fail:
     path), leaves IBS_STOPPED set, and re-arms the hardware for one
     more period. When that extra period overflows after
     perf_ibs_stop() clears IBS_STARTED, a second NMI arrives. If an
     unrelated NMI arrives first, the IBS handler takes the fail:
     path, clears IBS_STOPPED, and claims that NMI. The second IBS
     NMI is then unhandled ("Uhhuh. NMI received for unknown
     reason").
  2. With VIBS enabled, on hardware without IBS_CAPS_DIS, if this
     race happens when perf schedules out a host IBS event before
     VMRUN, IbsFetchEn or IbsOpEn is 1 at VMRUN. APM vol. 2, section
     15.38, says that these bits must be 0 at VMRUN of an SEV-ES or
     SEV-SNP guest with IBS virtualization enabled.

> The race is particularly noticeable when userspace repeatedly disables
> and re-enables IBS using PERF_EVENT_IOC_DISABLE and
> PERF_EVENT_IOC_ENABLE ioctls in a loop.
>
> Fix this by checking the IBS_STOPPING bit in the IBS NMI handler before
> re-enabling the IBS event. If the IBS_STOPPING bit is set, it indicates
> that the event is either disabled or in the process of being disabled,
> and the NMI handler should not re-enable it.
>
> Signed-off-by: Manali Shukla <manali.shukla@amd.com>

I think this warrants a Fixes tag:

Fixes: 85dc600263c2 ("perf/x86/amd/ibs: Fix pmu::stop() nesting")

This fix does not depend on VIBS. It is probably better to send it
separately through tip/perf:core, so that it can go in before the
rest of this series.

> ---
>  arch/x86/events/amd/ibs.c | 3 ++-
>  1 file changed, 2 insertions(+), 1 deletion(-)
>
> diff --git a/arch/x86/events/amd/ibs.c b/arch/x86/events/amd/ibs.c
> index eeb607b84dda..09b56bab510a 100644
> --- a/arch/x86/events/amd/ibs.c
> +++ b/arch/x86/events/amd/ibs.c
> @@ -1582,7 +1582,8 @@ static int perf_ibs_handle_irq(struct perf_ibs *perf_ibs, struct pt_regs *iregs)
>  		}
>  		new_config |= period >> 4;
>
> -		perf_ibs_enable_event(perf_ibs, hwc, new_config);
> +		if (!test_bit(IBS_STOPPING, pcpu->state))
> +			perf_ibs_enable_event(perf_ibs, hwc, new_config);

This stops the late re-arm. perf_ibs_stop() sets IBS_STOPPING first,
with test_and_set_bit(). An NMI before that point can re-arm the
hardware, but perf_ibs_stop() then disables the hardware. An NMI
after that point does not re-arm. Both sides run on the same CPU, so
a plain test_bit() in NMI context is sufficient.

However, the skipped re-arm causes three problems when an NMI arrives
after perf_ibs_stop() sets IBS_STOPPING and before it clears
IBS_STARTED.

First, the event count can increase twice for the same sample:

  1. The handler calls perf_ibs_event_update() for the sample and
     adds a full period. perf_ibs_set_period() sets prev_count to 0.
  2. Because of this patch, the handler does not re-arm. CTL (and
     perf_ibs_stop()'s local config copy, if already read) still
     holds the old sample with Val=1, and the handler does not set
     PERF_HES_UPTODATE.
  3. perf_ibs_stop() clears Val in its config copy and calls
     perf_ibs_event_update() again, because PERF_HES_UPTODATE is
     clear. Because prev_count is now 0, the delta is the whole
     count field: CurCnt (op) or FetchCnt (fetch) from the old
     sample, as if it were progress in a new period.

The amount added in step 3 depends on what the hardware leaves in the
count fields after a sample. For op, the comment in
get_ibs_op_count() says that the lower 7 bits of CurCnt are
randomized after a rollover, so the amount is in general not zero.

Second, IBS_STOPPED can stay set in pcpu->state. perf_ibs_stop() sets
IBS_STOPPED so that a late NMI can clear it at the fail: label. When
the NMI instead arrives before perf_ibs_stop() clears IBS_STARTED,
the handler takes the normal path and does not clear IBS_STOPPED.
Before this patch, the re-armed period gave a second NMI that cleared
it. Now that the handler does not re-arm, IBS_STOPPED can stay set
and falsely claim a later unrelated NMI.

Third, the same sample can be recorded twice. After the handler skips
the re-arm, CTL still holds the old sample with Val=1, and
IBS_STARTED is still set. This is true at least until
perf_ibs_stop() calls perf_ibs_disable_event(). With IBS_CAPS_DIS,
that call writes only CTL2, so it stays true until perf_ibs_stop()
clears IBS_STARTED. If an unrelated NMI arrives in this window, the
IBS handler takes the normal path again, records the same sample a
second time, and adds another full period, because prev_count is 0.
Before this patch, the re-arm cleared Val, so this could not happen.

The throttle path (throttle != 0, so the handler does not re-arm) has
the first two problems already, when perf_event_overflow() calls
pmu::stop() from the NMI. They are not caused by this patch, but they
may be worth a look at the same time.

>  	}
>
>  	perf_event_update_userpage(event);

^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH v3 2/9] x86/cpufeatures: Add CPUID feature bit for VIBS in SVM/SEV guests
  2026-03-10  6:00 ` [PATCH v3 2/9] x86/cpufeatures: Add CPUID feature bit for VIBS in SVM/SEV guests Manali Shukla
@ 2026-10-08 17:28   ` Jim Mattson
  0 siblings, 0 replies; 20+ messages in thread
From: Jim Mattson @ 2026-10-08 17:28 UTC (permalink / raw)
  To: Manali Shukla
  Cc: Jim Mattson, seanjc, pbonzini, mingo, bp, kvm, x86,
	santosh.shukla, nikunj.dadhania, Naveen.Rao, dapeng1.mi,
	ravi.bangoria, peterz, Sandipan.Das, Yosry Ahmed

On Tue, Mar 10, 2026 at 06:00:14AM +0000, Manali Shukla wrote:
> From: Santosh Shukla <santosh.shukla@amd.com>
>
> The virtualized IBS (VIBS) feature allows the guest to collect IBS
> samples without exiting the guest.
>
> Presence of the VIBS feature is indicated via CPUID function
> 0x8000000A_EDX[26].
>
> Signed-off-by: Santosh Shukla <santosh.shukla@amd.com>
> Signed-off-by: Manali Shukla <manali.shukla@amd.com>

Reviewed-by: Jim Mattson <jmattson@google.com>

^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH v3 3/9] KVM: x86/cpuid: Add a KVM-only leaf for IBS capabilities
  2026-03-10  6:00 ` [PATCH v3 3/9] KVM: x86/cpuid: Add a KVM-only leaf for IBS capabilities Manali Shukla
@ 2026-10-08 17:44   ` Jim Mattson
  0 siblings, 0 replies; 20+ messages in thread
From: Jim Mattson @ 2026-10-08 17:44 UTC (permalink / raw)
  To: Manali Shukla
  Cc: Jim Mattson, seanjc, pbonzini, mingo, bp, kvm, x86,
	santosh.shukla, nikunj.dadhania, Naveen.Rao, dapeng1.mi,
	ravi.bangoria, peterz, Sandipan.Das, Yosry Ahmed

On Tue, Mar 10, 2026 at 06:00:15AM +0000, Manali Shukla wrote:
> Add a KVM-only leaf for AMD's Instruction Based Sampling capabilities.
> Multiple IBS related capabilities are added to KVM-only leaf, so that KVM
> can set these capabilities for the guest, when IBS feature bit is
> enabled on the guest.
>
> Signed-off-by: Manali Shukla <manali.shukla@amd.com>
> ---
>  arch/x86/include/asm/kvm_host.h |  1 +
>  arch/x86/kvm/reverse_cpuid.h    | 16 ++++++++++++++++
>  2 files changed, 17 insertions(+)
>
> diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
> index 32dd2d55e6f0..01abdf7f112b 100644
> --- a/arch/x86/include/asm/kvm_host.h
> +++ b/arch/x86/include/asm/kvm_host.h
> @@ -782,6 +782,7 @@ enum kvm_only_cpuid_leafs {
>  	CPUID_12_EAX	 = NCAPINTS,
>  	CPUID_7_1_EDX,
>  	CPUID_8000_0007_EDX,
> +	CPUID_8000_001B_EAX,
>  	CPUID_8000_0022_EAX,
>  	CPUID_7_2_EDX,
>  	CPUID_24_0_EBX,

The other KVM-only leaves are in the order in which they were added, and
reverse_cpuid[] uses the same order. This patch puts the new leaf in the
middle of the enum, but near the end of reverse_cpuid[].  Please add it at
the end of the enum, after the last existing leaf, and in the same place in
reverse_cpuid[].

> diff --git a/arch/x86/kvm/reverse_cpuid.h b/arch/x86/kvm/reverse_cpuid.h
> index 657f5f743ed9..22cfdb331e9e 100644
> --- a/arch/x86/kvm/reverse_cpuid.h
> +++ b/arch/x86/kvm/reverse_cpuid.h
> @@ -76,6 +76,21 @@
>  #define KVM_X86_FEATURE_TSA_SQ_NO	KVM_X86_FEATURE(CPUID_8000_0021_ECX, 1)
>  #define KVM_X86_FEATURE_TSA_L1_NO	KVM_X86_FEATURE(CPUID_8000_0021_ECX, 2)
>
> +/* AMD defined Instruction-base Sampling capabilities. CPUID level 0x8000001B (EAX). */

Nit: "Instruction-base" should be "Instruction-Based."

> +#define X86_FEATURE_IBS_AVAIL			KVM_X86_FEATURE(CPUID_8000_001B_EAX, 0)
> +#define X86_FEATURE_IBS_FETCHSAM		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 1)
> +#define X86_FEATURE_IBS_OPSAM			KVM_X86_FEATURE(CPUID_8000_001B_EAX, 2)
> +#define X86_FEATURE_IBS_RDWROPCNT		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 3)
> +#define X86_FEATURE_IBS_OPCNT			KVM_X86_FEATURE(CPUID_8000_001B_EAX, 4)
> +#define X86_FEATURE_IBS_BRNTRGT			KVM_X86_FEATURE(CPUID_8000_001B_EAX, 5)
> +#define X86_FEATURE_IBS_OPCNTEXT		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 6)
> +#define X86_FEATURE_IBS_RIPINVALIDCHK		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 7)
> +#define X86_FEATURE_IBS_OPBRNFUSE		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 8)
> +#define X86_FEATURE_IBS_FETCHCTLEXTD		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 9)
> +#define X86_FEATURE_IBS_ZEN4_EXT		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 11)
> +#define X86_FEATURE_IBS_LOADLATFIL		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 12)
> +#define X86_FEATURE_IBS_ZEN4_DTLBSTAT		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 19)

Bit 10 (IBS_CAPS_OPDATA4 in asm/perf_event.h) is not defined here or later
in the series. Is that an oversight?

Three of these names are different from the IBS_CAPS_* names for the same
bits in asm/perf_event.h: ZEN4_EXT (IBS_CAPS_ZEN4), LOADLATFIL
(IBS_CAPS_OPLDLAT), and ZEN4_DTLBSTAT (IBS_CAPS_OPDTLBPGSIZE). Please
standardize on the existing names.

> +
>  struct cpuid_reg {
>  	u32 function;
>  	u32 index;
> @@ -105,6 +120,7 @@ static const struct cpuid_reg reverse_cpuid[] = {
>  	[CPUID_8000_0022_EAX] = {0x80000022, 0, CPUID_EAX},
>  	[CPUID_7_2_EDX]       = {         7, 2, CPUID_EDX},
>  	[CPUID_24_0_EBX]      = {      0x24, 0, CPUID_EBX},
> +	[CPUID_8000_001B_EAX] = {0x8000001b, 0, CPUID_EAX},
>  	[CPUID_8000_0021_ECX] = {0x80000021, 0, CPUID_ECX},
>  	[CPUID_7_1_ECX]       = {         7, 1, CPUID_ECX},
>  	[CPUID_1E_1_EAX]      = {      0x1e, 1, CPUID_EAX},

^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH v3 4/9] KVM: x86: Extend CPUID range to include new leaf
  2026-03-10  6:00 ` [PATCH v3 4/9] KVM: x86: Extend CPUID range to include new leaf Manali Shukla
@ 2026-10-08 17:59   ` Jim Mattson
  0 siblings, 0 replies; 20+ messages in thread
From: Jim Mattson @ 2026-10-08 17:59 UTC (permalink / raw)
  To: Manali Shukla
  Cc: Jim Mattson, seanjc, pbonzini, mingo, bp, kvm, x86,
	santosh.shukla, nikunj.dadhania, Naveen.Rao, dapeng1.mi,
	ravi.bangoria, peterz, Sandipan.Das, Yosry Ahmed

On Tue, Mar 10, 2026 at 06:00:16AM +0000, Manali Shukla wrote:
> CPUID leaf 0x8000001b (EAX) provides information about Instruction-Based
> sampling capabilities on AMD Platforms. Add the new leaf to
> kvm_cpu_cap_init() using F() macros, which automatically gate each
> capability bits against raw hardware CPUID via raw_cpuid_get().
>
> This allows vendor code to simply clear entire leaf when vibs is not
> enabled, rather than reading hardware CPUID and calling
> kvm_cpu_cap_set() for each capability bits inidividually in later
> patches.
>
> Signed-off-by: Manali Shukla <manali.shukla@amd.com>

The shortlog is not accurate. The CPUID range does not change; case
0x80000000 already limits the maximum extended leaf to 0x80000022, which
includes 0x8000001B. Before this patch, leaf 0x8000001B went to the default
case, which returns all zeros. This patch adds KVM's supported bits for the
leaf and a case that reports them when X86_FEATURE_IBS is supported. Please
change the shortlog to say that.

Also, there are some typos in the changelog:
  - "Instruction-Based sampling" should be "Instruction-Based Sampling"
  - "capability bits" should be "capability bit" (twice)
  - "inidividually" should be "individually"
  - "vibs" should be "VIBS"

> ---
>  arch/x86/kvm/cpuid.c | 25 +++++++++++++++++++++++++
>  1 file changed, 25 insertions(+)
>
> diff --git a/arch/x86/kvm/cpuid.c b/arch/x86/kvm/cpuid.c
> index 96a08a556543..4e626e77e6a6 100644
> --- a/arch/x86/kvm/cpuid.c
> +++ b/arch/x86/kvm/cpuid.c

[...]

> @@ -1848,6 +1864,15 @@ static inline int __do_cpuid_func(struct kvm_cpuid_array *array, u32 function)
>  		entry->eax = entry->ebx = entry->ecx = 0;
>  		entry->edx = 0; /* reserved */
>  		break;
> +	/* AMD IBS capability */
> +	case 0x8000001B:
> +		if (!kvm_cpu_cap_has(X86_FEATURE_IBS))
> +			entry->eax = 0;
> +		else
> +			cpuid_entry_override(entry, CPUID_8000_001B_EAX);
> +
> +		entry->ebx = entry->ecx = entry->edx = 0;
> +		break;

Please put case 0x8000001b between case 0x8000001a and case 0x8000001e.

>  	case 0x8000001F:
>  		if (!kvm_cpu_cap_has(X86_FEATURE_SEV)) {
>  			entry->eax = entry->ebx = entry->ecx = entry->edx = 0;

^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH v3 5/9] KVM: SVM: Extend VMCB area for virtualized IBS registers
  2026-03-10  6:00 ` [PATCH v3 5/9] KVM: SVM: Extend VMCB area for virtualized IBS registers Manali Shukla
@ 2026-10-08 18:01   ` Jim Mattson
  0 siblings, 0 replies; 20+ messages in thread
From: Jim Mattson @ 2026-10-08 18:01 UTC (permalink / raw)
  To: Manali Shukla
  Cc: Jim Mattson, seanjc, pbonzini, mingo, bp, kvm, x86,
	santosh.shukla, nikunj.dadhania, Naveen.Rao, dapeng1.mi,
	ravi.bangoria, peterz, Sandipan.Das, Yosry Ahmed

On Tue, Mar 10, 2026 at 06:00:17AM +0000, Manali Shukla wrote:
> From: Santosh Shukla <santosh.shukla@amd.com>
>
> Define the new VMCB fields that will be used to save and restore the
> state of the following fetch and op IBS related MSRs.
>
>   * MSRC001_1030 [IBS Fetch Control]
>   * MSRC001_1031 [IBS Fetch Linear Address]
>   * MSRC001_1033 [IBS Execution Control]
>   * MSRC001_1034 [IBS Op Logical Address]
>   * MSRC001_1035 [IBS Op Data]
>   * MSRC001_1036 [IBS Op Data 2]
>   * MSRC001_1037 [IBS Op Data 3]
>   * MSRC001_1038 [IBS DC Linear Address]
>   * MSRC001_103B [IBS Branch Target Address]
>   * MSRC001_103C [IBS Fetch Control Extended]
>
> Signed-off-by: Santosh Shukla <santosh.shukla@amd.com>
> Signed-off-by: Manali Shukla <manali.shukla@amd.com>

Reviewed-by: Jim Mattson <jmattson@google.com>

^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH v3 6/9] KVM: SVM: Add support for IBS Virtualization
  2026-03-10  6:00 ` [PATCH v3 6/9] KVM: SVM: Add support for IBS Virtualization Manali Shukla
@ 2026-10-08 19:19   ` Jim Mattson
  2026-10-08 21:28     ` Jim Mattson
  0 siblings, 1 reply; 20+ messages in thread
From: Jim Mattson @ 2026-10-08 19:19 UTC (permalink / raw)
  To: Manali Shukla
  Cc: Jim Mattson, seanjc, pbonzini, mingo, bp, kvm, x86,
	santosh.shukla, nikunj.dadhania, Naveen.Rao, dapeng1.mi,
	ravi.bangoria, peterz, Sandipan.Das, Yosry Ahmed

On Tue, Mar 10, 2026 at 06:00:18AM +0000, Manali Shukla wrote:
> From: Santosh Shukla <santosh.shukla@amd.com>
>
> IBS virtualization (VIBS) allows a guest to collect Instruction-Based
> Sampling (IBS) data using hardware-assisted virtualization. With VIBS
> enabled, the hardware automatically saves and restores guest IBS state
> during VM-Entry and VM-Exit via the VMCB State Save Area.
>
> IBS-generated interrupts are delivered directly to the guest without
> causing a VMEXIT.
>
> VIBS depends on mediated PMU mode and requires either AVIC or NMI
> virtualization for interrupt delivery. However, since AVIC can be
> dynamically inhibited, VIBS requires VNMI to be enabled to ensure
> reliable interrupt delivery. If AVIC is inhibited and VNMI is
> disabled, the guest can encounter a VMEXIT_INVALID when IBS
> virtualization is enabled for the guest.

APM vol. 2, section 15.38, disagrees. Without virtualized interrupt
delivery, "an IBS interrupt occurring in the guest will not be delivered to
either the guest or the hypervisor."  The only VMEXIT_INVALID that section
15.38 describes is for VMRUN of an SEV-ES or SEV-SNP guest with IbsFetchEn
or IbsOpEn set.

> Because IBS state is classified as swap type C, the hypervisor must
> save its own IBS state before VMRUN and restore it after VMEXIT. It
> must also disable IBS before VMRUN and re-enable it afterward. This
> will be handled using mediated PMU support in subsequent patches by
> enabling mediated PMU capability for IBS PMUs.

On hardware with IBS_CAPS_DIS, perf_ibs_stop() does not clear IbsFetchEn or
IbsOpEn. perf_ibs_disable_event() only sets CTL2[Dis], and
IBS_{FETCH|OP}_CTL[En] stays 1. Section 15.38 says that these bits must be
0 at VMRUN of an SEV-ES or SEV-SNP guest with IBS virtualization enabled
(otherwise VMRUN fails with VMEXIT_INVALID), and that they should be 0 for
other guests, "to prevent host IBS interrupts from leaking across a world
switch."  Section 15.38 does not say that CTL2[Dis]=1 is sufficient. Is it?
If not, the host must also clear En before VMRUN on hardware with
IBS_CAPS_DIS.

> More details about IBS virtualization can be found at [1].
>
> [1]: https://bugzilla.kernel.org/attachment.cgi?id=306250
>      AMD64 Architecture Programmer’s Manual, Vol 2, Section 15.38
>      Instruction-Based Sampling Virtualization.
>
> Signed-off-by: Santosh Shukla <santosh.shukla@amd.com>
> Co-developed-by: Manali Shukla <manali.shukla@amd.com>
> Signed-off-by: Manali Shukla <manali.shukla@amd.com>
> ---
>  arch/x86/include/asm/svm.h |  2 ++
>  arch/x86/kvm/svm/svm.c     | 73 +++++++++++++++++++++++++++++++++++++-
>  2 files changed, 74 insertions(+), 1 deletion(-)
>
> diff --git a/arch/x86/include/asm/svm.h b/arch/x86/include/asm/svm.h
> index 4296efc1dafe..17aa6bf76bce 100644
> --- a/arch/x86/include/asm/svm.h
> +++ b/arch/x86/include/asm/svm.h
> @@ -226,6 +226,8 @@ struct __attribute__ ((__packed__)) vmcb_control_area {
>
>  #define SVM_INT_VECTOR_MASK GENMASK(7, 0)
>
> +#define SVM_MISC_ENABLE_V_IBS BIT_ULL(2)
> +

This bit is in misc_ctl2, not misc_ctl. Please name it SVM_MISC2_ENABLE_V_IBS
and place it after SVM_MISC2_ENABLE_V_VMLOAD_VMSAVE.

[...]

> @@ -779,6 +782,26 @@ static void svm_recalc_pmu_msr_intercepts(struct kvm_vcpu *vcpu)
>  				  MSR_TYPE_RW, intercept);
>  }
>
> +static void svm_recalc_ibs_msr_intercepts(struct kvm_vcpu *vcpu)
> +{
> +	bool intercept = !(guest_cpu_cap_has(vcpu, X86_FEATURE_IBS) &&
> +			   kvm_vcpu_has_mediated_pmu(vcpu));
> +
> +	if (!enable_mediated_pmu || !vibs)
> +		return;
> +
> +	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSFETCHCTL, MSR_TYPE_RW, intercept);

For SEV-ES and SEV-SNP guests, section 15.38 says that IBS virtualization
is enabled by bit 12 of SEV_FEATURES in the VMSA, not by bit 2 at offset
B8h in the VMCB. This patch sets only the VMCB bit, but it disables these
intercepts for every vCPU that has X86_FEATURE_IBS and a mediated PMU,
*including* SEV-ES guests.  IBS virtualization is not enabled for SEV-ES
guests, so they will be able to read and write the host's IBS MSRs!

Since nested_vmcb02_prepare_control() does not set SVM_MISC_ENABLE_V_IBS in
vmcb02->control.misc_ctl2, nested SVM has a similar problem. If L1 does not
set INTERCEPT_MSR_PROT in vmcb12, vmcb02 uses vmcb01's msrpm, which
disables these MSR intercepts. Because VIBS is disabled in vmcb02, L2 can
read and write the host's physical IBS MSRs.

OTOH, if L1 sets INTERCEPT_MSR_PROT in vmcb12 and clears the IBS MSR
intercept bits in msrpm12, nested_svm_init_msrpm_merge_offsets() does not
include the IBS MSRs in merge_msrs[], so msrpm02 still intercepts them, and
L0 injects #GP into L2 on accesses to IBS_FETCH_CTL, IBS_OP_CTL, etc.

Furthermore, the guest's IBS state is in the VMCB save area, but userspace
has no way to save or restore it. The IBS MSRs are not in msrs_to_save[],
and svm_get_msr() and svm_set_msr() reject them, so KVM_{GET,SET}_MSRS will
fail. Hence, the IBS state is lost on suspend/resume.

[...]

> @@ -2880,6 +2904,27 @@ static int svm_get_msr(struct kvm_vcpu *vcpu, struct msr_data *msr_info)
>  	case MSR_AMD64_DE_CFG:
>  		msr_info->data = svm->msr_decfg;
>  		break;
> +
> +	case MSR_AMD64_IBSCTL:
> +		if (guest_cpu_cap_has(vcpu, X86_FEATURE_IBS))
> +			msr_info->data = IBSCTL_LVT_OFFSET_VALID;
> +		else
> +			msr_info->data = 0;
> +		break;

Before this patch, KVM synthesized a #GP when a guest without
X86_FEATURE_IBS read MSR_AMD64_IBSCTL. Now it returns 0.

The condition should also test for kvm_vcpu_has_mediated_pmu(vcpu).

> +
> +

Nit: there is an extra blank line above.

> +	/*
> +	 * When IBS virtualization is enabled, guest reads from
> +	 * MSR_AMD64_IBSFETCHPHYSAD and MSR_AMD64_IBSDCPHYSAD must return 0.
> +	 * This is done for security reasons, as guests should not be allowed to
> +	 * access or infer any information about the system's physical
> +	 * addresses.
> +	 */
> +	case MSR_AMD64_IBSDCPHYSAD:
> +	case MSR_AMD64_IBSFETCHPHYSAD:
> +		msr_info->data = 0;
> +		break;
> +

These new cases should be conditional. Before this patch, KVM synthesized
a #GP for a guest without X86_FEATURE_IBS. Now it returns 0.

But, is this even necessary? APM vol. 2, section 15.38, says that when IBS
virtualization is enabled in the VMCB, the processor returns 0 on guest
reads from IbsFetchPhysAd and IbsDcPhysAd and ignores guest writes to
them. Can't we just disable interception of these MSRs when VIBS is enabled
for the vCPU, and let the hardware handle them?

>  	default:
>  		return kvm_get_msr_common(vcpu, msr_info);
>  	}
> @@ -3171,6 +3216,16 @@ static int svm_set_msr(struct kvm_vcpu *vcpu, struct msr_data *msr)
>  		svm->msr_decfg = data;
>  		break;
>  	}
> +	/*
> +	 * When IBS virtualization is enabled, guest writes to
> +	 * MSR_AMD64_IBSFETCHPHYSAD and MSR_AMD64_IBSDCPHYSAD must be ignored.
> +	 * This is done for security reasons, as guests should not be allowed to
> +	 * access or infer any information about the system's physical
> +	 * addresses.
> +	 */
> +	case MSR_AMD64_IBSDCPHYSAD:
> +	case MSR_AMD64_IBSFETCHPHYSAD:
> +		return 1;

Returning 1 does not "ignore" the write; it induces a #GP. But, see the
point I made above.

[...]

> @@ -4678,6 +4733,11 @@ static void svm_vcpu_after_set_cpuid(struct kvm_vcpu *vcpu)
>  	if (guest_cpuid_is_intel_compatible(vcpu))
>  		guest_cpu_cap_clear(vcpu, X86_FEATURE_V_VMSAVE_VMLOAD);
>
> +	if (guest_cpu_cap_has(vcpu, X86_FEATURE_IBS))
> +		svm->vmcb->control.misc_ctl2 |= SVM_MISC_ENABLE_V_IBS;
> +	else
> +		svm->vmcb->control.misc_ctl2 &= ~SVM_MISC_ENABLE_V_IBS;
> +

The condition should also check kvm_vcpu_has_mediated_pmu(vcpu), as in
svm_recalc_ibs_msr_intercepts().

^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH v3 7/9] perf/x86/amd: Enable VPMU passthrough capability for IBS PMU
  2026-03-10  6:00 ` [PATCH v3 7/9] perf/x86/amd: Enable VPMU passthrough capability for IBS PMU Manali Shukla
@ 2026-10-08 19:32   ` Jim Mattson
  0 siblings, 0 replies; 20+ messages in thread
From: Jim Mattson @ 2026-10-08 19:32 UTC (permalink / raw)
  To: Manali Shukla
  Cc: Jim Mattson, seanjc, pbonzini, mingo, bp, kvm, x86,
	santosh.shukla, nikunj.dadhania, Naveen.Rao, dapeng1.mi,
	ravi.bangoria, peterz, Sandipan.Das, Yosry Ahmed

On Tue, Mar 10, 2026 at 06:00:19AM +0000, Manali Shukla wrote:
> IBS MSRs are classified as Swap Type C, which requires the hypervisor
> to save and restore its own IBS state before VMENTRY and after VMEXIT.
>
> To support this, set the ibs_op and ibs_fetch PMUs with the
> PERF_PMU_CAP_MEDIATED_VPMU capability. This ensures that these PMUs are
> exclusively owned by the guest while it is running, allowing the
> hypervisor to manage IBS state transitions correctly.
>
> Signed-off-by: Manali Shukla <manali.shukla@amd.com>
> ---
>  arch/x86/events/amd/ibs.c | 2 ++
>  1 file changed, 2 insertions(+)
>
> diff --git a/arch/x86/events/amd/ibs.c b/arch/x86/events/amd/ibs.c
> index 6d230e417b61..e075c5ed136c 100644
> --- a/arch/x86/events/amd/ibs.c
> +++ b/arch/x86/events/amd/ibs.c
> @@ -978,6 +978,7 @@ static struct perf_ibs perf_ibs_fetch = {
>  		.stop		= perf_ibs_stop,
>  		.read		= perf_ibs_read,
>  		.check_period	= perf_ibs_check_period,
> +		.capabilities	= PERF_PMU_CAP_MEDIATED_VPMU,

At this point in the series, perf_ibs_init() still rejects events with
exclude_guest set, so every IBS event is an "include guest" event, as
defined by is_include_guest_event(). With this capability:

- If any IBS event exists, perf_create_mediated_pmu() returns -EBUSY, so
  KVM cannot create a vCPU in a VM with a mediated PMU.

- If any VM with a mediated PMU exists, mediated_pmu_account_event()
  returns -EOPNOTSUPP, so no IBS event can be created.

Patch 8/9 removes the exclude_guest check. Please squash patches 7/9 and
8/9 together, so that the series does not have this regression at any
commit.

Do we really want to set this capability when the CPU does not support IBS
virtualization (or when kvm_amd.vibs=0)? In that case, the guest cannot use
IBS, but host IBS events that do not set exclude_guest still block mediated
PMU VMs, and vice versa.

When the guest context is loaded, perf_load_guest_context() schedules
out the exclude_guest IBS events through perf_ibs_stop(). On hardware
with IBS_CAPS_DIS, perf_ibs_stop() does not clear IbsFetchEn or IbsOpEn.
See my comment on patch 6/9.

>  	},
>  	.msr			= MSR_AMD64_IBSFETCHCTL,
>  	.msr2			= MSR_AMD64_IBSFETCHCTL2,

^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH v3 8/9] perf/x86/amd: Remove exclude_guest check from perf_ibs_init()
  2026-03-10  6:00 ` [PATCH v3 8/9] perf/x86/amd: Remove exclude_guest check from perf_ibs_init() Manali Shukla
@ 2026-10-08 19:38   ` Jim Mattson
  0 siblings, 0 replies; 20+ messages in thread
From: Jim Mattson @ 2026-10-08 19:38 UTC (permalink / raw)
  To: Manali Shukla
  Cc: Jim Mattson, seanjc, pbonzini, mingo, bp, kvm, x86,
	santosh.shukla, nikunj.dadhania, Naveen.Rao, dapeng1.mi,
	ravi.bangoria, peterz, Sandipan.Das, Yosry Ahmed

On Tue, Mar 10, 2026 at 06:00:20AM +0000, Manali Shukla wrote:
> Currently IBS driver doesn't allow the creation of IBS event with
> exclude_guest set. As a result, amd_ibs_init() returns -EINVAL if
> IBS event is created with exclude_guest set.

s/amd_ibs_init()/perf_ibs_init()/

> With the introduction of mediated PMU support, software-based handling
> of exclude_guest is permitted for PMUs that have the
> PERF_PMU_CAP_MEDIATED_VPMU capability.

This software-based handling applies only while a vCPU with a mediated
PMU is loaded (perf_load_guest_context() through perf_put_guest_context()).
The core PMU also sets AMD64_EVENTSEL_HOSTONLY for exclude_guest events,
so the hardware excludes guest execution in all other cases. IBS has no
equivalent, and nothing in the IBS driver or in KVM stops an IBS event
while a guest without a mediated PMU runs (for example, when
kvm_amd.enable_mediated_pmu=0, or when the VM is created with
KVM_PMU_CAP_DISABLE). After this patch, perf_ibs_init() accepts
exclude_guest=1 in those cases, but the event does not exclude the guest.
This seems like a problem.

> Since ibs_op and ibs_fetch pmus has PERF_PMU_CAP_MEDIATED_VPMU
> capability set, update perf_ibs_init() to remove exclude_guest check.

s/pmus has/PMUs have/

As I said on patch 7/9, please squash this patch with patch 7/9.

> Signed-off-by: Manali Shukla <manali.shukla@amd.com>
> ---
>  arch/x86/events/amd/ibs.c | 3 +--
>  1 file changed, 1 insertion(+), 2 deletions(-)
>
> diff --git a/arch/x86/events/amd/ibs.c b/arch/x86/events/amd/ibs.c
> index e075c5ed136c..47f4dc0b6341 100644
> --- a/arch/x86/events/amd/ibs.c
> +++ b/arch/x86/events/amd/ibs.c
> @@ -327,8 +327,7 @@ static int perf_ibs_init(struct perf_event *event)
>  		return -EOPNOTSUPP;
>
>  	/* handle exclude_{user,kernel} in the IRQ handler */
> -	if (event->attr.exclude_host || event->attr.exclude_guest ||
> -	    event->attr.exclude_idle)
> +	if (event->attr.exclude_host || event->attr.exclude_idle)
>  		return -EINVAL;
>
>  	ret = validate_group(event);

^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH v3 9/9] KVM: SVM: Add newly added IBS capabilities and MSRs
  2026-03-10  6:00 ` [PATCH v3 9/9] KVM: SVM: Add newly added IBS capabilities and MSRs Manali Shukla
@ 2026-10-08 21:03   ` Jim Mattson
  0 siblings, 0 replies; 20+ messages in thread
From: Jim Mattson @ 2026-10-08 21:03 UTC (permalink / raw)
  To: Manali Shukla
  Cc: Jim Mattson, seanjc, pbonzini, mingo, bp, kvm, x86,
	santosh.shukla, nikunj.dadhania, Naveen.Rao, dapeng1.mi,
	ravi.bangoria, peterz, Sandipan.Das, Yosry Ahmed

On Tue, Mar 10, 2026 at 06:00:21AM +0000, Manali Shukla wrote:
> IBS on upcoming microarch introduced two new control MSRs and a couple of
> new features. Define macros for them. Add these newly added IBS
> capabilities to KVM-only leaf 0x8000001b, so that when IBS feature bit
> is enabled on the guest, these newly added features can be used by
> guests if the hardware and guest os supports it.
>
>  - X86_FEATURE_IBS_DISABLE: Independent IBS disable capability to
>    avoid RMW race
>  - X86_FEATURE_IBS_FETCHLATFIL: Fetch Latency filtering
>  - X86_FEATURE_IBS_ADDRFILTER: Address Bit 63 based filtering
>  - X86_FEATURE_IBS_STRMST_RMTSOCKET: Streaming store filter and
>    indicator. Remote socket indicator.
>  - X86_FEATURE_IBS_BUFFER1: IBS buffering v1
>  - X86_FEATURE_IBS_MEMPROFILER: IBS memory profiler

APM vol. 2 (rev. 3.45), section 13.3.7, says, "Hardware virtualization and
IBS Buffering are not supported for IBS Memory Profiler Version 1."
Moreover, the memory profiler requires additional MSRs
(C001_0380h-C001_0386h), which are not in the list of IBS virtualization
state in section 15.38.

KVM should not advertise X86_FEATURE_IBS_MEMPROFILER.

IBS buffering *is* virtualized (IBS_BUFFER_BASE, IBS_BUFFER_SIZE, and
IBS_BUFFER_TAIL at offsets 7D8h-7E4h in Table B-2). However, this series
does not add these fields to struct vmcb_save_area, and it does not disable
interception of MSRs C001_0390h-C001_0392h. svm_get_msr() and svm_set_msr()
do not emulate these MSRs, so a guest cannot really use IBS buffering.

KVM should not advertise X86_FEATURE_IBS_BUFFER1.

> Extend VMCB save area to include to the newly added MSRs:
> MSR_AMD64_IBSFETCHCTL2 and MSR_AMD64_IBSOPCTL2.

s/include to/include/

As with the other IBS MSRs (see my comment on patch 6/9), svm_get_msr() and
svm_set_msr() do not handle IBS_FETCH_CTL2 and IBS_OP_CTL2, which breaks
KVM_{GET,SET}_MSRS and KVM_FEP. (And to correct what I wrote on patch 6/9:
even if we disable interception of MSR_AMD64_IBSFETCHPHYSAD and
MSR_AMD64_IBSDCPHYSAD so that hardware handles native guest accesses,
svm_get_msr() and svm_set_msr() still must handle those two MSRs as well,
for KVM_FEP if nothing else.)

> Signed-off-by: Manali Shukla <manali.shukla@amd.com>
> ---

[...]

> diff --git a/arch/x86/kvm/reverse_cpuid.h b/arch/x86/kvm/reverse_cpuid.h
> index 22cfdb331e9e..1af2ba207b8a 100644
> --- a/arch/x86/kvm/reverse_cpuid.h
> +++ b/arch/x86/kvm/reverse_cpuid.h
> @@ -89,6 +89,12 @@
>  #define X86_FEATURE_IBS_FETCHCTLEXTD		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 9)
>  #define X86_FEATURE_IBS_ZEN4_EXT		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 11)
>  #define X86_FEATURE_IBS_LOADLATFIL		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 12)
> +#define X86_FEATURE_IBS_DISABLE			KVM_X86_FEATURE(CPUID_8000_001B_EAX, 13)
> +#define X86_FEATURE_IBS_FETCHLATFIL		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 14)
> +#define X86_FEATURE_IBS_ADDRFILTER		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 15)
> +#define X86_FEATURE_IBS_STRMST_RMTSOCKET	KVM_X86_FEATURE(CPUID_8000_001B_EAX, 16)
> +#define X86_FEATURE_IBS_BUFFER1			KVM_X86_FEATURE(CPUID_8000_001B_EAX, 17)
> +#define X86_FEATURE_IBS_MEMPROFILER		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 18)

Three of these names are different from the IBS_CAPS_* names for the same
bits in asm/perf_event.h: IBS_DISABLE (IBS_CAPS_DIS), IBS_FETCHLATFIL
(IBS_CAPS_FETCHLAT), and IBS_ADDRFILTER (IBS_CAPS_BIT63_FILTER). Please
standardize on the existing names.

>  #define X86_FEATURE_IBS_ZEN4_DTLBSTAT		KVM_X86_FEATURE(CPUID_8000_001B_EAX, 19)
>
>  struct cpuid_reg {
> diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c
> index 421a929398da..9bf0d5f66239 100644
> --- a/arch/x86/kvm/svm/svm.c
> +++ b/arch/x86/kvm/svm/svm.c
> @@ -800,6 +800,13 @@ static void svm_recalc_ibs_msr_intercepts(struct kvm_vcpu *vcpu)
>  	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSDCLINAD, MSR_TYPE_RW, intercept);
>  	svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSBRTARGET, MSR_TYPE_RW, intercept);
>  	svm_set_intercept_for_msr(vcpu, MSR_AMD64_ICIBSEXTDCTL, MSR_TYPE_RW, intercept);
> +
> +	if (guest_cpu_cap_has(vcpu, X86_FEATURE_IBS_DISABLE)) {
> +		svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSFETCHCTL2, MSR_TYPE_RW,
> +				intercept);
> +		svm_set_intercept_for_msr(vcpu, MSR_AMD64_IBSOPCTL2, MSR_TYPE_RW,
> +				intercept);
> +	}

First, this function never re-enables the intercepts for
MSR_AMD64_IBSFETCHCTL2 and MSR_AMD64_IBSOPCTL2. If KVM_RUN recalculates the
intercepts but returns to userspace before VMRUN (e.g. due to a pending
signal), userspace can still clear X86_FEATURE_IBS_DISABLE with
KVM_SET_CPUID2, and the intercepts stay disabled. Call
svm_set_intercept_for_msr() unconditionally with the computed intercept
state, as svm_recalc_pmu_msr_intercepts() does.

Second, IbsDis (bit 13) is not the only CPUID bit that enumerates fields in
these MSRs. Per APM vol. 2 (rev. 3.45), section 13.3:

- IbsFetchCtl2[IbsFetchLatFilter] depends on IbsFetchLatencyFiltering
  (bit 14).

- IbsFetchCtl2[IbsFetchExclAddr63Eq{0,1}] and IbsOpCtl2[IbsOpExclRip63Eq{0,1}]
  depend on IbsAddrBit63Filtering (bit 15).

- IbsOpCtl2[IbsOpStrmStFilter] depends on IbsStrmStAndRmtSocket (bit 16).

The Linux IBS driver already uses these MSRs for IBS_CAPS_BIT63_FILTER
without IBS_CAPS_DIS (see perf_ibs_init()). If userspace gives the guest
any of bits 14-16 without bit 13, these MSRs stay intercepted,
svm_set_msr() rejects the guest's writes, and KVM synthesizes a #GP.

^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH v3 6/9] KVM: SVM: Add support for IBS Virtualization
  2026-10-08 19:19   ` Jim Mattson
@ 2026-10-08 21:28     ` Jim Mattson
  0 siblings, 0 replies; 20+ messages in thread
From: Jim Mattson @ 2026-10-08 21:28 UTC (permalink / raw)
  To: Manali Shukla
  Cc: seanjc, pbonzini, mingo, bp, kvm, x86, santosh.shukla,
	nikunj.dadhania, Naveen.Rao, dapeng1.mi, ravi.bangoria, peterz,
	Sandipan.Das, Yosry Ahmed

On Thu, Oct 8, 2026 at 12:19 PM Jim Mattson <jmattson@google.com> wrote:
>
> On Tue, Mar 10, 2026 at 06:00:18AM +0000, Manali Shukla wrote:
> > +     /*
> > +      * When IBS virtualization is enabled, guest reads from
> > +      * MSR_AMD64_IBSFETCHPHYSAD and MSR_AMD64_IBSDCPHYSAD must return 0.
> > +      * This is done for security reasons, as guests should not be allowed to
> > +      * access or infer any information about the system's physical
> > +      * addresses.
> > +      */
> > +     case MSR_AMD64_IBSDCPHYSAD:
> > +     case MSR_AMD64_IBSFETCHPHYSAD:
> > +             msr_info->data = 0;
> > +             break;
> > +
>
> These new cases should be conditional. Before this patch, KVM synthesized
> a #GP for a guest without X86_FEATURE_IBS. Now it returns 0.
>
> But, is this even necessary? APM vol. 2, section 15.38, says that when IBS
> virtualization is enabled in the VMCB, the processor returns 0 on guest
> reads from IbsFetchPhysAd and IbsDcPhysAd and ignores guest writes to
> them. Can't we just disable interception of these MSRs when VIBS is enabled
> for the vCPU, and let the hardware handle them?

Doh! As I realized later in the series, KVM_FEP requires that KVM must
be able to emulate reads and writes of every virtualized MSR...even if
virtualization just means "ignore."

^ permalink raw reply	[flat|nested] 20+ messages in thread

end of thread, other threads:[~2026-10-08 21:28 UTC | newest]

Thread overview: 20+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-03-10  6:00 [PATCH v3 0/9] Implement support for IBS virtualization Manali Shukla
2026-03-10  6:00 ` [PATCH v3 1/9] perf/amd/ibs: Fix race condition in IBS Manali Shukla
2026-10-08 17:19   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 2/9] x86/cpufeatures: Add CPUID feature bit for VIBS in SVM/SEV guests Manali Shukla
2026-10-08 17:28   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 3/9] KVM: x86/cpuid: Add a KVM-only leaf for IBS capabilities Manali Shukla
2026-10-08 17:44   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 4/9] KVM: x86: Extend CPUID range to include new leaf Manali Shukla
2026-10-08 17:59   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 5/9] KVM: SVM: Extend VMCB area for virtualized IBS registers Manali Shukla
2026-10-08 18:01   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 6/9] KVM: SVM: Add support for IBS Virtualization Manali Shukla
2026-10-08 19:19   ` Jim Mattson
2026-10-08 21:28     ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 7/9] perf/x86/amd: Enable VPMU passthrough capability for IBS PMU Manali Shukla
2026-10-08 19:32   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 8/9] perf/x86/amd: Remove exclude_guest check from perf_ibs_init() Manali Shukla
2026-10-08 19:38   ` Jim Mattson
2026-03-10  6:00 ` [PATCH v3 9/9] KVM: SVM: Add newly added IBS capabilities and MSRs Manali Shukla
2026-10-08 21:03   ` Jim Mattson

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox