* [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata
@ 2026-09-30 21:02 Sean Christopherson
2026-10-01 7:48 ` Sandipan Das
2026-10-02 21:07 ` Sean Christopherson
0 siblings, 2 replies; 4+ messages in thread
From: Sean Christopherson @ 2026-09-30 21:02 UTC (permalink / raw)
To: Paolo Bonzini; +Cc: kvm, Sandipan Das, Sean Christopherson
Apply the instructions and branches overcount errata to all AMD CPUs, which
count VMRUN as a retired branch instruction in guest context. I.e. any
asynchronous #VMEXITs during any measurement will result in an overcount.
Reported-by: Sandipan Das <sandipan.das@amd.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
lib/x86/pmu.c | 7 +++++++
x86/pmu.c | 7 +------
2 files changed, 8 insertions(+), 6 deletions(-)
diff --git a/lib/x86/pmu.c b/lib/x86/pmu.c
index 67f3b23e..3231ce2d 100644
--- a/lib/x86/pmu.c
+++ b/lib/x86/pmu.c
@@ -72,6 +72,13 @@ void pmu_init(void)
pmu.msr_global_status_clr = MSR_CORE_PERF_GLOBAL_OVF_CTRL;
}
} else {
+ /*
+ * All AMD CPUs overcount instructions and branches retired, as
+ * they count VMRUN as a branch instruction in guest context.
+ */
+ pmu.errata.instructions_retired_overcount = true;
+ pmu.errata.branches_retired_overcount = true;
+
if (this_cpu_has(X86_FEATURE_PERFCTR_CORE)) {
/* Performance Monitoring Version 2 Supported */
if (this_cpu_has(X86_FEATURE_AMD_PMU_V2)) {
diff --git a/x86/pmu.c b/x86/pmu.c
index b262ea59..35f8a818 100644
--- a/x86/pmu.c
+++ b/x86/pmu.c
@@ -222,13 +222,8 @@ static void adjust_events_range(struct pmu_event *gp_events,
* If HW supports GLOBAL_CTRL MSR, enabling and disabling PMCs are
* moved in __precise_loop(). Thus, instructions and branches events
* can be verified against a precise count instead of a rough range.
- *
- * Skip the precise checks on AMD, as AMD CPUs count VMRUN as a branch
- * instruction in guest context, which* leads to intermittent failures
- * as the counts will vary depending on how many asynchronous VM-Exits
- * occur while running the measured code, e.g. if the host takes IRQs.
*/
- if (pmu.is_intel && this_cpu_has_perf_global_ctrl()) {
+ if (this_cpu_has_perf_global_ctrl()) {
if (!pmu.errata.instructions_retired_overcount) {
gp_events[instruction_idx].min = LOOP_INSNS;
gp_events[instruction_idx].max = LOOP_INSNS;
base-commit: 5a221342a025948fa7d84f0c404748f6ba57f3ad
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata
2026-09-30 21:02 [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata Sean Christopherson
@ 2026-10-01 7:48 ` Sandipan Das
2026-10-01 23:38 ` Sean Christopherson
2026-10-02 21:07 ` Sean Christopherson
1 sibling, 1 reply; 4+ messages in thread
From: Sandipan Das @ 2026-10-01 7:48 UTC (permalink / raw)
To: Sean Christopherson, Paolo Bonzini; +Cc: kvm
On 01-10-2026 02:32, Sean Christopherson wrote:
> Apply the instructions and branches overcount errata to all AMD CPUs, which
> count VMRUN as a retired branch instruction in guest context. I.e. any
> asynchronous #VMEXITs during any measurement will result in an overcount.
>
> Reported-by: Sandipan Das <sandipan.das@amd.com>
> Signed-off-by: Sean Christopherson <seanjc@google.com>
> ---
> lib/x86/pmu.c | 7 +++++++
> x86/pmu.c | 7 +------
> 2 files changed, 8 insertions(+), 6 deletions(-)
>
> diff --git a/lib/x86/pmu.c b/lib/x86/pmu.c
> index 67f3b23e..3231ce2d 100644
> --- a/lib/x86/pmu.c
> +++ b/lib/x86/pmu.c
> @@ -72,6 +72,13 @@ void pmu_init(void)
> pmu.msr_global_status_clr = MSR_CORE_PERF_GLOBAL_OVF_CTRL;
> }
> } else {
> + /*
> + * All AMD CPUs overcount instructions and branches retired, as
> + * they count VMRUN as a branch instruction in guest context.
> + */
> + pmu.errata.instructions_retired_overcount = true;
> + pmu.errata.branches_retired_overcount = true;
> +
> if (this_cpu_has(X86_FEATURE_PERFCTR_CORE)) {
> /* Performance Monitoring Version 2 Supported */
> if (this_cpu_has(X86_FEATURE_AMD_PMU_V2)) {
> diff --git a/x86/pmu.c b/x86/pmu.c
> index b262ea59..35f8a818 100644
> --- a/x86/pmu.c
> +++ b/x86/pmu.c
> @@ -222,13 +222,8 @@ static void adjust_events_range(struct pmu_event *gp_events,
> * If HW supports GLOBAL_CTRL MSR, enabling and disabling PMCs are
> * moved in __precise_loop(). Thus, instructions and branches events
> * can be verified against a precise count instead of a rough range.
> - *
> - * Skip the precise checks on AMD, as AMD CPUs count VMRUN as a branch
> - * instruction in guest context, which* leads to intermittent failures
> - * as the counts will vary depending on how many asynchronous VM-Exits
> - * occur while running the measured code, e.g. if the host takes IRQs.
> */
> - if (pmu.is_intel && this_cpu_has_perf_global_ctrl()) {
> + if (this_cpu_has_perf_global_ctrl()) {
> if (!pmu.errata.instructions_retired_overcount) {
> gp_events[instruction_idx].min = LOOP_INSNS;
> gp_events[instruction_idx].max = LOOP_INSNS;
>
Commit 61a7ee521d ("x86/pmu: Handle instruction overcount issue in overflow
test") makes measure_for_overflow() always return 1 - LOOP_INSNS (-10000016)
when pmu.errata.instructions_retired_overcount is true. For runs without
"perfmon-v2" on a Zen 5 system, I see that the overshoot from this preset value
is between 125 and 135, which makes the following condition in
check_counter_overflow() fail.
report(cnt.count == 0xffffffffffff || cnt.count < 7, "cntr-%d", i);
How should we handle this?
Restrict the condition in measure_for_overflow() to return 1 - LOOP_INSNS
if (pmu.is_intel && pmu.errata.instructions_retired_overcount)
or, accept a larger overshoot.
diff --git a/x86/pmu.c b/x86/pmu.c
index 35f8a8183a..74e30b149d 100644
--- a/x86/pmu.c
+++ b/x86/pmu.c
@@ -584,6 +584,8 @@ static void check_counter_overflow(void)
else
report(cnt.count == 1, "cntr-%d", i);
}
+ else if (pmu.errata.instructions_retired_overcount)
+ report(cnt.count == 0xffffffffffff || cnt.count < 150, "cntr-%d", i);
else
report(cnt.count == 0xffffffffffff || cnt.count < 7, "cntr-%d", i);
^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata
2026-10-01 7:48 ` Sandipan Das
@ 2026-10-01 23:38 ` Sean Christopherson
0 siblings, 0 replies; 4+ messages in thread
From: Sean Christopherson @ 2026-10-01 23:38 UTC (permalink / raw)
To: Sandipan Das; +Cc: Paolo Bonzini, kvm
On Thu, Oct 01, 2026, Sandipan Das wrote:
> On 01-10-2026 02:32, Sean Christopherson wrote:
> > diff --git a/x86/pmu.c b/x86/pmu.c
> > index b262ea59..35f8a818 100644
> > --- a/x86/pmu.c
> > +++ b/x86/pmu.c
> > @@ -222,13 +222,8 @@ static void adjust_events_range(struct pmu_event *gp_events,
> > * If HW supports GLOBAL_CTRL MSR, enabling and disabling PMCs are
> > * moved in __precise_loop(). Thus, instructions and branches events
> > * can be verified against a precise count instead of a rough range.
> > - *
> > - * Skip the precise checks on AMD, as AMD CPUs count VMRUN as a branch
> > - * instruction in guest context, which* leads to intermittent failures
> > - * as the counts will vary depending on how many asynchronous VM-Exits
> > - * occur while running the measured code, e.g. if the host takes IRQs.
> > */
> > - if (pmu.is_intel && this_cpu_has_perf_global_ctrl()) {
> > + if (this_cpu_has_perf_global_ctrl()) {
> > if (!pmu.errata.instructions_retired_overcount) {
> > gp_events[instruction_idx].min = LOOP_INSNS;
> > gp_events[instruction_idx].max = LOOP_INSNS;
> >
>
> Commit 61a7ee521d ("x86/pmu: Handle instruction overcount issue in overflow
> test") makes measure_for_overflow() always return 1 - LOOP_INSNS (-10000016)
> when pmu.errata.instructions_retired_overcount is true. For runs without
> "perfmon-v2" on a Zen 5 system, I see that the overshoot from this preset value
> is between 125 and 135,
Heh, I was going to say "woah, that's a lot of exits!", then I realized the
loop does a million runs...
> which makes the following condition in check_counter_overflow() fail.
>
> report(cnt.count == 0xffffffffffff || cnt.count < 7, "cntr-%d", i);
>
> How should we handle this?
>
> Restrict the condition in measure_for_overflow() to return 1 - LOOP_INSNS
> if (pmu.is_intel && pmu.errata.instructions_retired_overcount)
>
> or, accept a larger overshoot.
Doesn't it have to be the latter? Because won't the test fail if the first run
of __measure() from measure_for_overflow() takes more exits than the second run
of __measure()?
> diff --git a/x86/pmu.c b/x86/pmu.c
> index 35f8a8183a..74e30b149d 100644
> --- a/x86/pmu.c
> +++ b/x86/pmu.c
> @@ -584,6 +584,8 @@ static void check_counter_overflow(void)
> else
> report(cnt.count == 1, "cntr-%d", i);
> }
> + else if (pmu.errata.instructions_retired_overcount)
> + report(cnt.count == 0xffffffffffff || cnt.count < 150, "cntr-%d", i);
Do you have any idea where the 0xffffffffffff check comes from? Commit b883751a
("x86/pmu: Update testcases to cover AMD PMU") doesn't explain it at all.
Depending on what's up with the 0xffffffffffff thing, my vote would be to go big
hammer, and express the overflow as a fraction of loop instructions, e.g. allow
for an extra 200 exits:
if (pmu.errata.instructions_retired_overcount)
report(cnt.count < N / 5000, "cntr-%d", i);
else
report(cnt.count == 1, "cntr-%d", i);
> else
> report(cnt.count == 0xffffffffffff || cnt.count < 7, "cntr-%d", i);
>
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata
2026-09-30 21:02 [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata Sean Christopherson
2026-10-01 7:48 ` Sandipan Das
@ 2026-10-02 21:07 ` Sean Christopherson
1 sibling, 0 replies; 4+ messages in thread
From: Sean Christopherson @ 2026-10-02 21:07 UTC (permalink / raw)
To: Sean Christopherson, Paolo Bonzini; +Cc: kvm, Sandipan Das
On Wed, 30 Sep 2026 14:02:01 -0700, Sean Christopherson wrote:
> Apply the instructions and branches overcount errata to all AMD CPUs, which
> count VMRUN as a retired branch instruction in guest context. I.e. any
> asynchronous #VMEXITs during any measurement will result in an overcount.
Applied to kvm-x86 next, thanks!
[1/1] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata
https://github.com/kvm-x86/kvm-unit-tests/commit/7999ff1d28b0
--
https://github.com/kvm-x86/kvm-unit-tests/tree/next
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-10-02 21:10 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30 21:02 [kvm-unit-tests PATCH] x86/pmu: Treat all AMD CPUs as having instruction/branch overcount errata Sean Christopherson
2026-10-01 7:48 ` Sandipan Das
2026-10-01 23:38 ` Sean Christopherson
2026-10-02 21:07 ` Sean Christopherson
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox