Kernel KVM virtualization development
 help / color / mirror / Atom feed
* [kvm-unit-tests PATCH] x86/pmu: Relax precise count check for emulated instructions on AMD
@ 2026-07-15  6:28 Sandipan Das
  2026-09-30 20:18 ` Sean Christopherson
  0 siblings, 1 reply; 2+ messages in thread
From: Sandipan Das @ 2026-07-15  6:28 UTC (permalink / raw)
  To: kvm
  Cc: Sean Christopherson, Paolo Bonzini, Dapeng Mi,
	Nikunj A . Dadhania, Manali Shukla, Sandipan Das

check_emulated_instr() expects the retired instruction and branch
counts to match the expected values exactly when a global control MSR
is available. This does not hold on AMD processors because VMRUN is
counted as a retired instruction and branch in guest context. The exact
comparison therefore fails, with the surplus varying with how many
asynchronous #VMEXITs occur while the measured code runs.

Hence, gate the precise comparison on pmu.is_intel and fall back to the
lower-bound check, mirroring how adjust_events_range() already handles
the same VMRUN behaviour.

Signed-off-by: Sandipan Das <sandipan.das@amd.com>
---
 x86/pmu.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/x86/pmu.c b/x86/pmu.c
index b262ea59..2b8d1b5b 100644
--- a/x86/pmu.c
+++ b/x86/pmu.c
@@ -786,12 +786,12 @@ static void check_emulated_instr(void)
 
 	// Check that the end count - start count is at least the expected
 	// number of instructions and branches.
-	if (has_perf_global_ctrl && !pmu.errata.instructions_retired_overcount)
+	if (pmu.is_intel && has_perf_global_ctrl && !pmu.errata.instructions_retired_overcount)
 		report(instr_cnt.count - instr_start == KVM_FEP_INSNS, "instruction count");
 	else
 		report(instr_cnt.count - instr_start >= KVM_FEP_INSNS, "instruction count");
 
-	if (has_perf_global_ctrl && !pmu.errata.branches_retired_overcount)
+	if (pmu.is_intel && has_perf_global_ctrl && !pmu.errata.branches_retired_overcount)
 		report(brnch_cnt.count - brnch_start == KVM_FEP_BRANCHES, "branch count");
 	else
 		report(brnch_cnt.count - brnch_start >= KVM_FEP_BRANCHES, "branch count");
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 2+ messages in thread

* Re: [kvm-unit-tests PATCH] x86/pmu: Relax precise count check for emulated instructions on AMD
  2026-07-15  6:28 [kvm-unit-tests PATCH] x86/pmu: Relax precise count check for emulated instructions on AMD Sandipan Das
@ 2026-09-30 20:18 ` Sean Christopherson
  0 siblings, 0 replies; 2+ messages in thread
From: Sean Christopherson @ 2026-09-30 20:18 UTC (permalink / raw)
  To: Sandipan Das
  Cc: kvm, Paolo Bonzini, Dapeng Mi, Nikunj A . Dadhania, Manali Shukla

On Wed, Jul 15, 2026, Sandipan Das wrote:
> check_emulated_instr() expects the retired instruction and branch
> counts to match the expected values exactly when a global control MSR
> is available. This does not hold on AMD processors because VMRUN is
> counted as a retired instruction and branch in guest context. The exact
> comparison therefore fails, with the surplus varying with how many
> asynchronous #VMEXITs occur while the measured code runs.
> 
> Hence, gate the precise comparison on pmu.is_intel and fall back to the
> lower-bound check, mirroring how adjust_events_range() already handles
> the same VMRUN behaviour.
> 
> Signed-off-by: Sandipan Das <sandipan.das@amd.com>
> ---
>  x86/pmu.c | 4 ++--
>  1 file changed, 2 insertions(+), 2 deletions(-)
> 
> diff --git a/x86/pmu.c b/x86/pmu.c
> index b262ea59..2b8d1b5b 100644
> --- a/x86/pmu.c
> +++ b/x86/pmu.c
> @@ -786,12 +786,12 @@ static void check_emulated_instr(void)
>  
>  	// Check that the end count - start count is at least the expected
>  	// number of instructions and branches.
> -	if (has_perf_global_ctrl && !pmu.errata.instructions_retired_overcount)
> +	if (pmu.is_intel && has_perf_global_ctrl && !pmu.errata.instructions_retired_overcount)
>  		report(instr_cnt.count - instr_start == KVM_FEP_INSNS, "instruction count");
>  	else
>  		report(instr_cnt.count - instr_start >= KVM_FEP_INSNS, "instruction count");
>  
> -	if (has_perf_global_ctrl && !pmu.errata.branches_retired_overcount)
> +	if (pmu.is_intel && has_perf_global_ctrl && !pmu.errata.branches_retired_overcount)
>  		report(brnch_cnt.count - brnch_start == KVM_FEP_BRANCHES, "branch count");
>  	else
>  		report(brnch_cnt.count - brnch_start >= KVM_FEP_BRANCHES, "branch count");

I'd rather treat the AMD behavior as errata, e.g. so that we don't play whack-a-mole
with thing like measure_for_overflow().  Even with the below, I still see random
one-off failures in the overflow test (I've been ignoring AMD PMU failures for a
very long time).  I haven't debugged why, but AFAICT this is still a strict
improvement:

From: Sean Christopherson <seanjc@google.com>
Date: Wed, 30 Sep 2026 13:15:18 -0700
Subject: [PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch
 overcount errata

Apply the instructions and branches overcount errata to all AMD CPUs, which
count VMRUN as a retired branch instruction in guest context.  I.e. any
asynchronous #VMEXITs during any measurement will result in an overcount.

Reported-by: Sandipan Das <sandipan.das@amd.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
 lib/x86/pmu.c | 7 +++++++
 x86/pmu.c     | 7 +------
 2 files changed, 8 insertions(+), 6 deletions(-)

diff --git a/lib/x86/pmu.c b/lib/x86/pmu.c
index 67f3b23e..3231ce2d 100644
--- a/lib/x86/pmu.c
+++ b/lib/x86/pmu.c
@@ -72,6 +72,13 @@ void pmu_init(void)
 			pmu.msr_global_status_clr = MSR_CORE_PERF_GLOBAL_OVF_CTRL;
 		}
 	} else {
+		/*
+		 * All AMD CPUs overcount instructions and branches retired, as
+		 * they count VMRUN as a branch instruction in guest context.
+		 */
+		pmu.errata.instructions_retired_overcount = true;
+		pmu.errata.branches_retired_overcount = true;
+
 		if (this_cpu_has(X86_FEATURE_PERFCTR_CORE)) {
 			/* Performance Monitoring Version 2 Supported */
 			if (this_cpu_has(X86_FEATURE_AMD_PMU_V2)) {
diff --git a/x86/pmu.c b/x86/pmu.c
index b262ea59..35f8a818 100644
--- a/x86/pmu.c
+++ b/x86/pmu.c
@@ -222,13 +222,8 @@ static void adjust_events_range(struct pmu_event *gp_events,
 	 * If HW supports GLOBAL_CTRL MSR, enabling and disabling PMCs are
 	 * moved in __precise_loop(). Thus, instructions and branches events
 	 * can be verified against a precise count instead of a rough range.
-	 *
-	 * Skip the precise checks on AMD, as AMD CPUs count VMRUN as a branch
-	 * instruction in guest context, which* leads to intermittent failures
-	 * as the counts will vary depending on how many asynchronous VM-Exits
-	 * occur while running the measured code, e.g. if the host takes IRQs.
 	 */
-	if (pmu.is_intel && this_cpu_has_perf_global_ctrl()) {
+	if (this_cpu_has_perf_global_ctrl()) {
 		if (!pmu.errata.instructions_retired_overcount) {
 			gp_events[instruction_idx].min = LOOP_INSNS;
 			gp_events[instruction_idx].max = LOOP_INSNS;

base-commit: eae36be65d135b609f22ca72c6ad836f5c2e2699
-- 

^ permalink raw reply related	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-30 20:18 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-15  6:28 [kvm-unit-tests PATCH] x86/pmu: Relax precise count check for emulated instructions on AMD Sandipan Das
2026-09-30 20:18 ` Sean Christopherson

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox