* [kvm-unit-tests PATCH] x86/pmu: Relax precise count check for emulated instructions on AMD
@ 2026-07-15 6:28 Sandipan Das
2026-09-30 20:18 ` Sean Christopherson
0 siblings, 1 reply; 2+ messages in thread
From: Sandipan Das @ 2026-07-15 6:28 UTC (permalink / raw)
To: kvm
Cc: Sean Christopherson, Paolo Bonzini, Dapeng Mi,
Nikunj A . Dadhania, Manali Shukla, Sandipan Das
check_emulated_instr() expects the retired instruction and branch
counts to match the expected values exactly when a global control MSR
is available. This does not hold on AMD processors because VMRUN is
counted as a retired instruction and branch in guest context. The exact
comparison therefore fails, with the surplus varying with how many
asynchronous #VMEXITs occur while the measured code runs.
Hence, gate the precise comparison on pmu.is_intel and fall back to the
lower-bound check, mirroring how adjust_events_range() already handles
the same VMRUN behaviour.
Signed-off-by: Sandipan Das <sandipan.das@amd.com>
---
x86/pmu.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/x86/pmu.c b/x86/pmu.c
index b262ea59..2b8d1b5b 100644
--- a/x86/pmu.c
+++ b/x86/pmu.c
@@ -786,12 +786,12 @@ static void check_emulated_instr(void)
// Check that the end count - start count is at least the expected
// number of instructions and branches.
- if (has_perf_global_ctrl && !pmu.errata.instructions_retired_overcount)
+ if (pmu.is_intel && has_perf_global_ctrl && !pmu.errata.instructions_retired_overcount)
report(instr_cnt.count - instr_start == KVM_FEP_INSNS, "instruction count");
else
report(instr_cnt.count - instr_start >= KVM_FEP_INSNS, "instruction count");
- if (has_perf_global_ctrl && !pmu.errata.branches_retired_overcount)
+ if (pmu.is_intel && has_perf_global_ctrl && !pmu.errata.branches_retired_overcount)
report(brnch_cnt.count - brnch_start == KVM_FEP_BRANCHES, "branch count");
else
report(brnch_cnt.count - brnch_start >= KVM_FEP_BRANCHES, "branch count");
--
2.53.0
^ permalink raw reply related [flat|nested] 2+ messages in thread
* Re: [kvm-unit-tests PATCH] x86/pmu: Relax precise count check for emulated instructions on AMD
2026-07-15 6:28 [kvm-unit-tests PATCH] x86/pmu: Relax precise count check for emulated instructions on AMD Sandipan Das
@ 2026-09-30 20:18 ` Sean Christopherson
0 siblings, 0 replies; 2+ messages in thread
From: Sean Christopherson @ 2026-09-30 20:18 UTC (permalink / raw)
To: Sandipan Das
Cc: kvm, Paolo Bonzini, Dapeng Mi, Nikunj A . Dadhania, Manali Shukla
On Wed, Jul 15, 2026, Sandipan Das wrote:
> check_emulated_instr() expects the retired instruction and branch
> counts to match the expected values exactly when a global control MSR
> is available. This does not hold on AMD processors because VMRUN is
> counted as a retired instruction and branch in guest context. The exact
> comparison therefore fails, with the surplus varying with how many
> asynchronous #VMEXITs occur while the measured code runs.
>
> Hence, gate the precise comparison on pmu.is_intel and fall back to the
> lower-bound check, mirroring how adjust_events_range() already handles
> the same VMRUN behaviour.
>
> Signed-off-by: Sandipan Das <sandipan.das@amd.com>
> ---
> x86/pmu.c | 4 ++--
> 1 file changed, 2 insertions(+), 2 deletions(-)
>
> diff --git a/x86/pmu.c b/x86/pmu.c
> index b262ea59..2b8d1b5b 100644
> --- a/x86/pmu.c
> +++ b/x86/pmu.c
> @@ -786,12 +786,12 @@ static void check_emulated_instr(void)
>
> // Check that the end count - start count is at least the expected
> // number of instructions and branches.
> - if (has_perf_global_ctrl && !pmu.errata.instructions_retired_overcount)
> + if (pmu.is_intel && has_perf_global_ctrl && !pmu.errata.instructions_retired_overcount)
> report(instr_cnt.count - instr_start == KVM_FEP_INSNS, "instruction count");
> else
> report(instr_cnt.count - instr_start >= KVM_FEP_INSNS, "instruction count");
>
> - if (has_perf_global_ctrl && !pmu.errata.branches_retired_overcount)
> + if (pmu.is_intel && has_perf_global_ctrl && !pmu.errata.branches_retired_overcount)
> report(brnch_cnt.count - brnch_start == KVM_FEP_BRANCHES, "branch count");
> else
> report(brnch_cnt.count - brnch_start >= KVM_FEP_BRANCHES, "branch count");
I'd rather treat the AMD behavior as errata, e.g. so that we don't play whack-a-mole
with thing like measure_for_overflow(). Even with the below, I still see random
one-off failures in the overflow test (I've been ignoring AMD PMU failures for a
very long time). I haven't debugged why, but AFAICT this is still a strict
improvement:
From: Sean Christopherson <seanjc@google.com>
Date: Wed, 30 Sep 2026 13:15:18 -0700
Subject: [PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch
overcount errata
Apply the instructions and branches overcount errata to all AMD CPUs, which
count VMRUN as a retired branch instruction in guest context. I.e. any
asynchronous #VMEXITs during any measurement will result in an overcount.
Reported-by: Sandipan Das <sandipan.das@amd.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
lib/x86/pmu.c | 7 +++++++
x86/pmu.c | 7 +------
2 files changed, 8 insertions(+), 6 deletions(-)
diff --git a/lib/x86/pmu.c b/lib/x86/pmu.c
index 67f3b23e..3231ce2d 100644
--- a/lib/x86/pmu.c
+++ b/lib/x86/pmu.c
@@ -72,6 +72,13 @@ void pmu_init(void)
pmu.msr_global_status_clr = MSR_CORE_PERF_GLOBAL_OVF_CTRL;
}
} else {
+ /*
+ * All AMD CPUs overcount instructions and branches retired, as
+ * they count VMRUN as a branch instruction in guest context.
+ */
+ pmu.errata.instructions_retired_overcount = true;
+ pmu.errata.branches_retired_overcount = true;
+
if (this_cpu_has(X86_FEATURE_PERFCTR_CORE)) {
/* Performance Monitoring Version 2 Supported */
if (this_cpu_has(X86_FEATURE_AMD_PMU_V2)) {
diff --git a/x86/pmu.c b/x86/pmu.c
index b262ea59..35f8a818 100644
--- a/x86/pmu.c
+++ b/x86/pmu.c
@@ -222,13 +222,8 @@ static void adjust_events_range(struct pmu_event *gp_events,
* If HW supports GLOBAL_CTRL MSR, enabling and disabling PMCs are
* moved in __precise_loop(). Thus, instructions and branches events
* can be verified against a precise count instead of a rough range.
- *
- * Skip the precise checks on AMD, as AMD CPUs count VMRUN as a branch
- * instruction in guest context, which* leads to intermittent failures
- * as the counts will vary depending on how many asynchronous VM-Exits
- * occur while running the measured code, e.g. if the host takes IRQs.
*/
- if (pmu.is_intel && this_cpu_has_perf_global_ctrl()) {
+ if (this_cpu_has_perf_global_ctrl()) {
if (!pmu.errata.instructions_retired_overcount) {
gp_events[instruction_idx].min = LOOP_INSNS;
gp_events[instruction_idx].max = LOOP_INSNS;
base-commit: eae36be65d135b609f22ca72c6ad836f5c2e2699
--
^ permalink raw reply related [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-09-30 20:18 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-15 6:28 [kvm-unit-tests PATCH] x86/pmu: Relax precise count check for emulated instructions on AMD Sandipan Das
2026-09-30 20:18 ` Sean Christopherson
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox