* [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata @ 2026-09-30 21:02 Sean Christopherson 2026-10-01 7:48 ` Sandipan Das 2026-10-02 21:07 ` Sean Christopherson 0 siblings, 2 replies; 4+ messages in thread From: Sean Christopherson @ 2026-09-30 21:02 UTC (permalink / raw) To: Paolo Bonzini; +Cc: kvm, Sandipan Das, Sean Christopherson Apply the instructions and branches overcount errata to all AMD CPUs, which count VMRUN as a retired branch instruction in guest context. I.e. any asynchronous #VMEXITs during any measurement will result in an overcount. Reported-by: Sandipan Das <sandipan.das@amd.com> Signed-off-by: Sean Christopherson <seanjc@google.com> --- lib/x86/pmu.c | 7 +++++++ x86/pmu.c | 7 +------ 2 files changed, 8 insertions(+), 6 deletions(-) diff --git a/lib/x86/pmu.c b/lib/x86/pmu.c index 67f3b23e..3231ce2d 100644 --- a/lib/x86/pmu.c +++ b/lib/x86/pmu.c @@ -72,6 +72,13 @@ void pmu_init(void) pmu.msr_global_status_clr = MSR_CORE_PERF_GLOBAL_OVF_CTRL; } } else { + /* + * All AMD CPUs overcount instructions and branches retired, as + * they count VMRUN as a branch instruction in guest context. + */ + pmu.errata.instructions_retired_overcount = true; + pmu.errata.branches_retired_overcount = true; + if (this_cpu_has(X86_FEATURE_PERFCTR_CORE)) { /* Performance Monitoring Version 2 Supported */ if (this_cpu_has(X86_FEATURE_AMD_PMU_V2)) { diff --git a/x86/pmu.c b/x86/pmu.c index b262ea59..35f8a818 100644 --- a/x86/pmu.c +++ b/x86/pmu.c @@ -222,13 +222,8 @@ static void adjust_events_range(struct pmu_event *gp_events, * If HW supports GLOBAL_CTRL MSR, enabling and disabling PMCs are * moved in __precise_loop(). Thus, instructions and branches events * can be verified against a precise count instead of a rough range. - * - * Skip the precise checks on AMD, as AMD CPUs count VMRUN as a branch - * instruction in guest context, which* leads to intermittent failures - * as the counts will vary depending on how many asynchronous VM-Exits - * occur while running the measured code, e.g. if the host takes IRQs. */ - if (pmu.is_intel && this_cpu_has_perf_global_ctrl()) { + if (this_cpu_has_perf_global_ctrl()) { if (!pmu.errata.instructions_retired_overcount) { gp_events[instruction_idx].min = LOOP_INSNS; gp_events[instruction_idx].max = LOOP_INSNS; base-commit: 5a221342a025948fa7d84f0c404748f6ba57f3ad -- 2.56.0.rc1.315.gc6ed9934b7-goog ^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata 2026-09-30 21:02 [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata Sean Christopherson @ 2026-10-01 7:48 ` Sandipan Das 2026-10-01 23:38 ` Sean Christopherson 2026-10-02 21:07 ` Sean Christopherson 1 sibling, 1 reply; 4+ messages in thread From: Sandipan Das @ 2026-10-01 7:48 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini; +Cc: kvm On 01-10-2026 02:32, Sean Christopherson wrote: > Apply the instructions and branches overcount errata to all AMD CPUs, which > count VMRUN as a retired branch instruction in guest context. I.e. any > asynchronous #VMEXITs during any measurement will result in an overcount. > > Reported-by: Sandipan Das <sandipan.das@amd.com> > Signed-off-by: Sean Christopherson <seanjc@google.com> > --- > lib/x86/pmu.c | 7 +++++++ > x86/pmu.c | 7 +------ > 2 files changed, 8 insertions(+), 6 deletions(-) > > diff --git a/lib/x86/pmu.c b/lib/x86/pmu.c > index 67f3b23e..3231ce2d 100644 > --- a/lib/x86/pmu.c > +++ b/lib/x86/pmu.c > @@ -72,6 +72,13 @@ void pmu_init(void) > pmu.msr_global_status_clr = MSR_CORE_PERF_GLOBAL_OVF_CTRL; > } > } else { > + /* > + * All AMD CPUs overcount instructions and branches retired, as > + * they count VMRUN as a branch instruction in guest context. > + */ > + pmu.errata.instructions_retired_overcount = true; > + pmu.errata.branches_retired_overcount = true; > + > if (this_cpu_has(X86_FEATURE_PERFCTR_CORE)) { > /* Performance Monitoring Version 2 Supported */ > if (this_cpu_has(X86_FEATURE_AMD_PMU_V2)) { > diff --git a/x86/pmu.c b/x86/pmu.c > index b262ea59..35f8a818 100644 > --- a/x86/pmu.c > +++ b/x86/pmu.c > @@ -222,13 +222,8 @@ static void adjust_events_range(struct pmu_event *gp_events, > * If HW supports GLOBAL_CTRL MSR, enabling and disabling PMCs are > * moved in __precise_loop(). Thus, instructions and branches events > * can be verified against a precise count instead of a rough range. > - * > - * Skip the precise checks on AMD, as AMD CPUs count VMRUN as a branch > - * instruction in guest context, which* leads to intermittent failures > - * as the counts will vary depending on how many asynchronous VM-Exits > - * occur while running the measured code, e.g. if the host takes IRQs. > */ > - if (pmu.is_intel && this_cpu_has_perf_global_ctrl()) { > + if (this_cpu_has_perf_global_ctrl()) { > if (!pmu.errata.instructions_retired_overcount) { > gp_events[instruction_idx].min = LOOP_INSNS; > gp_events[instruction_idx].max = LOOP_INSNS; > Commit 61a7ee521d ("x86/pmu: Handle instruction overcount issue in overflow test") makes measure_for_overflow() always return 1 - LOOP_INSNS (-10000016) when pmu.errata.instructions_retired_overcount is true. For runs without "perfmon-v2" on a Zen 5 system, I see that the overshoot from this preset value is between 125 and 135, which makes the following condition in check_counter_overflow() fail. report(cnt.count == 0xffffffffffff || cnt.count < 7, "cntr-%d", i); How should we handle this? Restrict the condition in measure_for_overflow() to return 1 - LOOP_INSNS if (pmu.is_intel && pmu.errata.instructions_retired_overcount) or, accept a larger overshoot. diff --git a/x86/pmu.c b/x86/pmu.c index 35f8a8183a..74e30b149d 100644 --- a/x86/pmu.c +++ b/x86/pmu.c @@ -584,6 +584,8 @@ static void check_counter_overflow(void) else report(cnt.count == 1, "cntr-%d", i); } + else if (pmu.errata.instructions_retired_overcount) + report(cnt.count == 0xffffffffffff || cnt.count < 150, "cntr-%d", i); else report(cnt.count == 0xffffffffffff || cnt.count < 7, "cntr-%d", i); ^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata 2026-10-01 7:48 ` Sandipan Das @ 2026-10-01 23:38 ` Sean Christopherson 0 siblings, 0 replies; 4+ messages in thread From: Sean Christopherson @ 2026-10-01 23:38 UTC (permalink / raw) To: Sandipan Das; +Cc: Paolo Bonzini, kvm On Thu, Oct 01, 2026, Sandipan Das wrote: > On 01-10-2026 02:32, Sean Christopherson wrote: > > diff --git a/x86/pmu.c b/x86/pmu.c > > index b262ea59..35f8a818 100644 > > --- a/x86/pmu.c > > +++ b/x86/pmu.c > > @@ -222,13 +222,8 @@ static void adjust_events_range(struct pmu_event *gp_events, > > * If HW supports GLOBAL_CTRL MSR, enabling and disabling PMCs are > > * moved in __precise_loop(). Thus, instructions and branches events > > * can be verified against a precise count instead of a rough range. > > - * > > - * Skip the precise checks on AMD, as AMD CPUs count VMRUN as a branch > > - * instruction in guest context, which* leads to intermittent failures > > - * as the counts will vary depending on how many asynchronous VM-Exits > > - * occur while running the measured code, e.g. if the host takes IRQs. > > */ > > - if (pmu.is_intel && this_cpu_has_perf_global_ctrl()) { > > + if (this_cpu_has_perf_global_ctrl()) { > > if (!pmu.errata.instructions_retired_overcount) { > > gp_events[instruction_idx].min = LOOP_INSNS; > > gp_events[instruction_idx].max = LOOP_INSNS; > > > > Commit 61a7ee521d ("x86/pmu: Handle instruction overcount issue in overflow > test") makes measure_for_overflow() always return 1 - LOOP_INSNS (-10000016) > when pmu.errata.instructions_retired_overcount is true. For runs without > "perfmon-v2" on a Zen 5 system, I see that the overshoot from this preset value > is between 125 and 135, Heh, I was going to say "woah, that's a lot of exits!", then I realized the loop does a million runs... > which makes the following condition in check_counter_overflow() fail. > > report(cnt.count == 0xffffffffffff || cnt.count < 7, "cntr-%d", i); > > How should we handle this? > > Restrict the condition in measure_for_overflow() to return 1 - LOOP_INSNS > if (pmu.is_intel && pmu.errata.instructions_retired_overcount) > > or, accept a larger overshoot. Doesn't it have to be the latter? Because won't the test fail if the first run of __measure() from measure_for_overflow() takes more exits than the second run of __measure()? > diff --git a/x86/pmu.c b/x86/pmu.c > index 35f8a8183a..74e30b149d 100644 > --- a/x86/pmu.c > +++ b/x86/pmu.c > @@ -584,6 +584,8 @@ static void check_counter_overflow(void) > else > report(cnt.count == 1, "cntr-%d", i); > } > + else if (pmu.errata.instructions_retired_overcount) > + report(cnt.count == 0xffffffffffff || cnt.count < 150, "cntr-%d", i); Do you have any idea where the 0xffffffffffff check comes from? Commit b883751a ("x86/pmu: Update testcases to cover AMD PMU") doesn't explain it at all. Depending on what's up with the 0xffffffffffff thing, my vote would be to go big hammer, and express the overflow as a fraction of loop instructions, e.g. allow for an extra 200 exits: if (pmu.errata.instructions_retired_overcount) report(cnt.count < N / 5000, "cntr-%d", i); else report(cnt.count == 1, "cntr-%d", i); > else > report(cnt.count == 0xffffffffffff || cnt.count < 7, "cntr-%d", i); > ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata 2026-09-30 21:02 [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata Sean Christopherson 2026-10-01 7:48 ` Sandipan Das @ 2026-10-02 21:07 ` Sean Christopherson 1 sibling, 0 replies; 4+ messages in thread From: Sean Christopherson @ 2026-10-02 21:07 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini; +Cc: kvm, Sandipan Das On Wed, 30 Sep 2026 14:02:01 -0700, Sean Christopherson wrote: > Apply the instructions and branches overcount errata to all AMD CPUs, which > count VMRUN as a retired branch instruction in guest context. I.e. any > asynchronous #VMEXITs during any measurement will result in an overcount. Applied to kvm-x86 next, thanks! [1/1] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata https://github.com/kvm-x86/kvm-unit-tests/commit/7999ff1d28b0 -- https://github.com/kvm-x86/kvm-unit-tests/tree/next ^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-10-02 21:10 UTC | newest] Thread overview: 4+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-09-30 21:02 [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata Sean Christopherson 2026-10-01 7:48 ` Sandipan Das 2026-10-01 23:38 ` Sean Christopherson 2026-10-02 21:07 ` Sean Christopherson
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox