* [PATCH v8 1/7] target/i386/kvm: set KVM_PMU_CAP_DISABLE if "-pmu" is configured
2025-12-30 7:42 [PATCH v8 0/7] target/i386/kvm/pmu: PMU Enhancement, Bugfix and Cleanup Dongli Zhang
@ 2025-12-30 7:42 ` Dongli Zhang
2026-10-01 8:32 ` Michael Tokarev
2025-12-30 7:42 ` [PATCH v8 2/7] target/i386/kvm: extract unrelated code out of kvm_x86_build_cpuid() Dongli Zhang
` (5 subsequent siblings)
6 siblings, 1 reply; 25+ messages in thread
From: Dongli Zhang @ 2025-12-30 7:42 UTC (permalink / raw)
To: qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
Although AMD PERFCORE and PerfMonV2 are removed when "-pmu" is configured,
there is no way to fully disable KVM AMD PMU virtualization. Neither
"-cpu host,-pmu" nor "-cpu EPYC" achieves this.
As a result, the following message still appears in the VM dmesg:
[ 0.263615] Performance Events: AMD PMU driver.
However, the expected output should be:
[ 0.596381] Performance Events: PMU not available due to virtualization, using software events only.
[ 0.600972] NMI watchdog: Perf NMI watchdog permanently disabled
This occurs because AMD does not use any CPUID bit to indicate PMU
availability.
To address this, KVM_CAP_PMU_CAPABILITY is used to set KVM_PMU_CAP_DISABLE
when "-pmu" is configured.
Signed-off-by: Dongli Zhang <dongli.zhang@oracle.com>
Reviewed-by: Xiaoyao Li <xiaoyao.li@intel.com>
Reviewed-by: Zhao Liu <zhao1.liu@intel.com>
Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
Changed since v1:
- Switch back to the initial implementation with "-pmu".
https://lore.kernel.org/all/20221119122901.2469-3-dongli.zhang@oracle.com
- Mention that "KVM_PMU_CAP_DISABLE doesn't change the PMU behavior on
Intel platform because current "pmu" property works as expected."
Changed since v2:
- Change has_pmu_cap to pmu_cap.
- Use (pmu_cap & KVM_PMU_CAP_DISABLE) instead of only pmu_cap in if
statement.
- Add Reviewed-by from Xiaoyao and Zhao as the change is minor.
Changed since v5:
- Re-base on top of most recent mainline QEMU.
- To resolve conflicts, move the PMU related code before the
call site of is_tdx_vm().
Changed since v6:
- Add Reviewed-by from Dapeng Mi.
target/i386/kvm/kvm.c | 31 +++++++++++++++++++++++++++++++
1 file changed, 31 insertions(+)
diff --git a/target/i386/kvm/kvm.c b/target/i386/kvm/kvm.c
index 7b9b740a8e..c98832f423 100644
--- a/target/i386/kvm/kvm.c
+++ b/target/i386/kvm/kvm.c
@@ -179,6 +179,8 @@ static int has_triple_fault_event;
static bool has_msr_mcg_ext_ctl;
+static int pmu_cap;
+
static struct kvm_cpuid2 *cpuid_cache;
static struct kvm_cpuid2 *hv_cpuid_cache;
static struct kvm_msr_list *kvm_feature_msrs;
@@ -2080,6 +2082,33 @@ full:
int kvm_arch_pre_create_vcpu(CPUState *cpu, Error **errp)
{
+ static bool first = true;
+ int ret;
+
+ if (first) {
+ first = false;
+
+ /*
+ * Since Linux v5.18, KVM provides a VM-level capability to easily
+ * disable PMUs; however, QEMU has been providing PMU property per
+ * CPU since v1.6. In order to accommodate both, have to configure
+ * the VM-level capability here.
+ *
+ * KVM_PMU_CAP_DISABLE doesn't change the PMU
+ * behavior on Intel platform because current "pmu" property works
+ * as expected.
+ */
+ if ((pmu_cap & KVM_PMU_CAP_DISABLE) && !X86_CPU(cpu)->enable_pmu) {
+ ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
+ KVM_PMU_CAP_DISABLE);
+ if (ret < 0) {
+ error_setg_errno(errp, -ret,
+ "Failed to set KVM_PMU_CAP_DISABLE");
+ return ret;
+ }
+ }
+ }
+
if (is_tdx_vm()) {
return tdx_pre_create_vcpu(cpu, errp);
}
@@ -3391,6 +3420,8 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
}
}
+ pmu_cap = kvm_check_extension(s, KVM_CAP_PMU_CAPABILITY);
+
return 0;
}
--
2.39.3
^ permalink raw reply related [flat|nested] 25+ messages in thread* Re: [PATCH v8 1/7] target/i386/kvm: set KVM_PMU_CAP_DISABLE if "-pmu" is configured
2025-12-30 7:42 ` [PATCH v8 1/7] target/i386/kvm: set KVM_PMU_CAP_DISABLE if "-pmu" is configured Dongli Zhang
@ 2026-10-01 8:32 ` Michael Tokarev
2026-10-01 8:54 ` Sandipan Das
0 siblings, 1 reply; 25+ messages in thread
From: Michael Tokarev @ 2026-10-01 8:32 UTC (permalink / raw)
To: Dongli Zhang, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
On 12/30/25 10:42, Dongli Zhang wrote:
> Although AMD PERFCORE and PerfMonV2 are removed when "-pmu" is configured,
> there is no way to fully disable KVM AMD PMU virtualization. Neither
> "-cpu host,-pmu" nor "-cpu EPYC" achieves this.
>
> As a result, the following message still appears in the VM dmesg:
>
> [ 0.263615] Performance Events: AMD PMU driver.
>
> However, the expected output should be:
>
> [ 0.596381] Performance Events: PMU not available due to virtualization, using software events only.
> [ 0.600972] NMI watchdog: Perf NMI watchdog permanently disabled
>
> This occurs because AMD does not use any CPUID bit to indicate PMU
> availability.
>
> To address this, KVM_CAP_PMU_CAPABILITY is used to set KVM_PMU_CAP_DISABLE
> when "-pmu" is configured.
>
> Signed-off-by: Dongli Zhang <dongli.zhang@oracle.com>
> Reviewed-by: Xiaoyao Li <xiaoyao.li@intel.com>
> Reviewed-by: Zhao Liu <zhao1.liu@intel.com>
> Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
This is commit 4f0dd09e61b080c27589fe35b575503783c5a256 (v10.2.0-1114-g4f0dd09e61b)
in the master branch now.
This change seem to broke booting windows 10/11 on an AMD host, in paticular
with -cpu EPYC-Rome-v4. It were initially reported in debian qemu bug tracker,
and I verified and bisected the issue on my AMD desktop machine (which is AMD
Ryzen 7 5700G, Zen 3.
When booting windows11 VM with -cpu EPYC-Rome-v4 -enable-kvm, the
guest enters tight CPU loop while displaying a boot screen, before
it starts showing up the rotating circle (which shows boot progress).
Before this patch, it boots ok.
There's no other fancy options, just
qemu-system-x86_64 -m 2G -enable-kvm -cpu EPYC-Rome-v4 -smp 2
(I'm replying to the original email message with all original recipients,
hope it's okay).
Can someone please take a look?
Thanks,
/mjt
> target/i386/kvm/kvm.c | 31 +++++++++++++++++++++++++++++++
> 1 file changed, 31 insertions(+)
>
> diff --git a/target/i386/kvm/kvm.c b/target/i386/kvm/kvm.c
> index 7b9b740a8e..c98832f423 100644
> --- a/target/i386/kvm/kvm.c
> +++ b/target/i386/kvm/kvm.c
> @@ -179,6 +179,8 @@ static int has_triple_fault_event;
>
> static bool has_msr_mcg_ext_ctl;
>
> +static int pmu_cap;
> +
> static struct kvm_cpuid2 *cpuid_cache;
> static struct kvm_cpuid2 *hv_cpuid_cache;
> static struct kvm_msr_list *kvm_feature_msrs;
> @@ -2080,6 +2082,33 @@ full:
>
> int kvm_arch_pre_create_vcpu(CPUState *cpu, Error **errp)
> {
> + static bool first = true;
> + int ret;
> +
> + if (first) {
> + first = false;
> +
> + /*
> + * Since Linux v5.18, KVM provides a VM-level capability to easily
> + * disable PMUs; however, QEMU has been providing PMU property per
> + * CPU since v1.6. In order to accommodate both, have to configure
> + * the VM-level capability here.
> + *
> + * KVM_PMU_CAP_DISABLE doesn't change the PMU
> + * behavior on Intel platform because current "pmu" property works
> + * as expected.
> + */
> + if ((pmu_cap & KVM_PMU_CAP_DISABLE) && !X86_CPU(cpu)->enable_pmu) {
> + ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
> + KVM_PMU_CAP_DISABLE);
> + if (ret < 0) {
> + error_setg_errno(errp, -ret,
> + "Failed to set KVM_PMU_CAP_DISABLE");
> + return ret;
> + }
> + }
> + }
> +
> if (is_tdx_vm()) {
> return tdx_pre_create_vcpu(cpu, errp);
> }
> @@ -3391,6 +3420,8 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
> }
> }
>
> + pmu_cap = kvm_check_extension(s, KVM_CAP_PMU_CAPABILITY);
> +
> return 0;
> }
>
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v8 1/7] target/i386/kvm: set KVM_PMU_CAP_DISABLE if "-pmu" is configured
2026-10-01 8:32 ` Michael Tokarev
@ 2026-10-01 8:54 ` Sandipan Das
2026-10-01 10:00 ` Michael Tokarev
0 siblings, 1 reply; 25+ messages in thread
From: Sandipan Das @ 2026-10-01 8:54 UTC (permalink / raw)
To: Michael Tokarev, Dongli Zhang, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, babu.moger, likexu, like.xu.linux,
groug, khorenko, alexander.ivanov, den, davydov-max, xiaoyao.li,
dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
On 01-10-2026 14:02, Michael Tokarev wrote:
> On 12/30/25 10:42, Dongli Zhang wrote:
>> Although AMD PERFCORE and PerfMonV2 are removed when "-pmu" is configured,
>> there is no way to fully disable KVM AMD PMU virtualization. Neither
>> "-cpu host,-pmu" nor "-cpu EPYC" achieves this.
>>
>> As a result, the following message still appears in the VM dmesg:
>>
>> [ 0.263615] Performance Events: AMD PMU driver.
>>
>> However, the expected output should be:
>>
>> [ 0.596381] Performance Events: PMU not available due to virtualization, using software events only.
>> [ 0.600972] NMI watchdog: Perf NMI watchdog permanently disabled
>>
>> This occurs because AMD does not use any CPUID bit to indicate PMU
>> availability.
>>
>> To address this, KVM_CAP_PMU_CAPABILITY is used to set KVM_PMU_CAP_DISABLE
>> when "-pmu" is configured.
>>
>> Signed-off-by: Dongli Zhang <dongli.zhang@oracle.com>
>> Reviewed-by: Xiaoyao Li <xiaoyao.li@intel.com>
>> Reviewed-by: Zhao Liu <zhao1.liu@intel.com>
>> Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
>
> This is commit 4f0dd09e61b080c27589fe35b575503783c5a256 (v10.2.0-1114-g4f0dd09e61b)
> in the master branch now.
>
> This change seem to broke booting windows 10/11 on an AMD host, in paticular
> with -cpu EPYC-Rome-v4. It were initially reported in debian qemu bug tracker,
> and I verified and bisected the issue on my AMD desktop machine (which is AMD
> Ryzen 7 5700G, Zen 3.
>
> When booting windows11 VM with -cpu EPYC-Rome-v4 -enable-kvm, the
> guest enters tight CPU loop while displaying a boot screen, before
> it starts showing up the rotating circle (which shows boot progress).
>
> Before this patch, it boots ok.
>
> There's no other fancy options, just
>
> qemu-system-x86_64 -m 2G -enable-kvm -cpu EPYC-Rome-v4 -smp 2
>
> (I'm replying to the original email message with all original recipients,
> hope it's okay).
>
> Can someone please take a look?
>
What happens if you use "-cpu EPYC-Rome-v4,+pmu"?
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [PATCH v8 1/7] target/i386/kvm: set KVM_PMU_CAP_DISABLE if "-pmu" is configured
2026-10-01 8:54 ` Sandipan Das
@ 2026-10-01 10:00 ` Michael Tokarev
2026-10-01 10:43 ` Sandipan Das
0 siblings, 1 reply; 25+ messages in thread
From: Michael Tokarev @ 2026-10-01 10:00 UTC (permalink / raw)
To: Sandipan Das, Dongli Zhang, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, babu.moger, likexu, like.xu.linux,
groug, khorenko, alexander.ivanov, den, davydov-max, xiaoyao.li,
dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
On 10/1/26 11:54, Sandipan Das wrote:
> On 01-10-2026 14:02, Michael Tokarev wrote:
...
>> When booting windows11 VM with -cpu EPYC-Rome-v4 -enable-kvm, the
>> guest enters tight CPU loop while displaying a boot screen, before
>> it starts showing up the rotating circle (which shows boot progress).
> What happens if you use "-cpu EPYC-Rome-v4,+pmu"?
With that, it boots fine, just like before this patch.
I checked current master too - same, +pmu fixes it.
Thanks,
/mjt
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [PATCH v8 1/7] target/i386/kvm: set KVM_PMU_CAP_DISABLE if "-pmu" is configured
2026-10-01 10:00 ` Michael Tokarev
@ 2026-10-01 10:43 ` Sandipan Das
2026-10-01 11:05 ` Michael Tokarev
0 siblings, 1 reply; 25+ messages in thread
From: Sandipan Das @ 2026-10-01 10:43 UTC (permalink / raw)
To: Michael Tokarev, Dongli Zhang, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, babu.moger, likexu, like.xu.linux,
groug, khorenko, alexander.ivanov, den, davydov-max, xiaoyao.li,
dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
On 01-10-2026 15:30, Michael Tokarev wrote:
> On 10/1/26 11:54, Sandipan Das wrote:
>> On 01-10-2026 14:02, Michael Tokarev wrote:
> ...
>
>>> When booting windows11 VM with -cpu EPYC-Rome-v4 -enable-kvm, the
>>> guest enters tight CPU loop while displaying a boot screen, before
>>> it starts showing up the rotating circle (which shows boot progress).
>
>
>> What happens if you use "-cpu EPYC-Rome-v4,+pmu"?
>
> With that, it boots fine, just like before this patch.
> I checked current master too - same, +pmu fixes it.
>
A guest OS cannot expect a functioning PMU if the hypervisor has disabled
it. Linux runs a basic sanity test to determine this. See:
https://github.com/torvalds/linux/blob/master/arch/x86/events/core.c#L268
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [PATCH v8 1/7] target/i386/kvm: set KVM_PMU_CAP_DISABLE if "-pmu" is configured
2026-10-01 10:43 ` Sandipan Das
@ 2026-10-01 11:05 ` Michael Tokarev
2026-10-01 11:12 ` Sandipan Das
0 siblings, 1 reply; 25+ messages in thread
From: Michael Tokarev @ 2026-10-01 11:05 UTC (permalink / raw)
To: Sandipan Das, Dongli Zhang, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, babu.moger, likexu, like.xu.linux,
groug, khorenko, alexander.ivanov, den, davydov-max, xiaoyao.li,
dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
On 10/1/26 13:43, Sandipan Das wrote:
> On 01-10-2026 15:30, Michael Tokarev wrote:
>> On 10/1/26 11:54, Sandipan Das wrote:
>>> On 01-10-2026 14:02, Michael Tokarev wrote:
>> ...
>>
>>>> When booting windows11 VM with -cpu EPYC-Rome-v4 -enable-kvm, the
>>>> guest enters tight CPU loop while displaying a boot screen, before
>>>> it starts showing up the rotating circle (which shows boot progress).
>>
>>
>>> What happens if you use "-cpu EPYC-Rome-v4,+pmu"?
>>
>> With that, it boots fine, just like before this patch.
>> I checked current master too - same, +pmu fixes it.
>>
> A guest OS cannot expect a functioning PMU if the hypervisor has disabled
> it. Linux runs a basic sanity test to determine this. See:
> https://github.com/torvalds/linux/blob/master/arch/x86/events/core.c#L268
So.. what do we do with all this? :)
I kinda see the intention of the original change, but it results in
non-bootable systems (linux too, as I think you demonstrate above).
hmm.. :)
Thanks,
/mjt
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [PATCH v8 1/7] target/i386/kvm: set KVM_PMU_CAP_DISABLE if "-pmu" is configured
2026-10-01 11:05 ` Michael Tokarev
@ 2026-10-01 11:12 ` Sandipan Das
2026-10-06 8:41 ` Sandipan Das
0 siblings, 1 reply; 25+ messages in thread
From: Sandipan Das @ 2026-10-01 11:12 UTC (permalink / raw)
To: Michael Tokarev, Dongli Zhang, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, babu.moger, likexu, like.xu.linux,
groug, khorenko, alexander.ivanov, den, davydov-max, xiaoyao.li,
dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
On 01-10-2026 16:35, Michael Tokarev wrote:
> On 10/1/26 13:43, Sandipan Das wrote:
>> On 01-10-2026 15:30, Michael Tokarev wrote:
>>> On 10/1/26 11:54, Sandipan Das wrote:
>>>> On 01-10-2026 14:02, Michael Tokarev wrote:
>>> ...
>>>
>>>>> When booting windows11 VM with -cpu EPYC-Rome-v4 -enable-kvm, the
>>>>> guest enters tight CPU loop while displaying a boot screen, before
>>>>> it starts showing up the rotating circle (which shows boot progress).
>>>
>>>
>>>> What happens if you use "-cpu EPYC-Rome-v4,+pmu"?
>>>
>>> With that, it boots fine, just like before this patch.
>>> I checked current master too - same, +pmu fixes it.
>>>
>> A guest OS cannot expect a functioning PMU if the hypervisor has disabled
>> it. Linux runs a basic sanity test to determine this. See:
>> https://github.com/torvalds/linux/blob/master/arch/x86/events/core.c#L268
>
> So.. what do we do with all this? :)
>
> I kinda see the intention of the original change, but it results in
> non-bootable systems (linux too, as I think you demonstrate above).
>
Because of the sanity test, Linux will boot but not register the "cpu" PMU.
Perhaps others should also run a similar sanity test on PMU hardware rather
than making assumptions and failing to boot.
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [PATCH v8 1/7] target/i386/kvm: set KVM_PMU_CAP_DISABLE if "-pmu" is configured
2026-10-01 11:12 ` Sandipan Das
@ 2026-10-06 8:41 ` Sandipan Das
2026-10-07 0:51 ` Dongli Zhang
0 siblings, 1 reply; 25+ messages in thread
From: Sandipan Das @ 2026-10-06 8:41 UTC (permalink / raw)
To: zhao1.liu, Dongli Zhang
Cc: pbonzini, mtosatti, babu.moger, likexu, like.xu.linux, groug,
khorenko, alexander.ivanov, den, davydov-max, xiaoyao.li,
dapeng1.mi, joe.jin, ewanhai-oc, ewanhai, Michael Tokarev,
qemu-devel, kvm
On 01-10-2026 16:42, Sandipan Das wrote:
> On 01-10-2026 16:35, Michael Tokarev wrote:
>> On 10/1/26 13:43, Sandipan Das wrote:
>>> On 01-10-2026 15:30, Michael Tokarev wrote:
>>>> On 10/1/26 11:54, Sandipan Das wrote:
>>>>> On 01-10-2026 14:02, Michael Tokarev wrote:
>>>> ...
>>>>
>>>>>> When booting windows11 VM with -cpu EPYC-Rome-v4 -enable-kvm, the
>>>>>> guest enters tight CPU loop while displaying a boot screen, before
>>>>>> it starts showing up the rotating circle (which shows boot progress).
>>>>
>>>>
>>>>> What happens if you use "-cpu EPYC-Rome-v4,+pmu"?
>>>>
>>>> With that, it boots fine, just like before this patch.
>>>> I checked current master too - same, +pmu fixes it.
>>>>
>>> A guest OS cannot expect a functioning PMU if the hypervisor has disabled
>>> it. Linux runs a basic sanity test to determine this. See:
>>> https://github.com/torvalds/linux/blob/master/arch/x86/events/core.c#L268
>>
>> So.. what do we do with all this? :)
>>
>> I kinda see the intention of the original change, but it results in
>> non-bootable systems (linux too, as I think you demonstrate above).
>>
>
> Because of the sanity test, Linux will boot but not register the "cpu" PMU.
> Perhaps others should also run a similar sanity test on PMU hardware rather
> than making assumptions and failing to boot.
After looking at kvm_msr and kvm_cpuid traces, I see that Windows is indeed
getting #GPs upon accessing MSRs 0xc0010200 and 0xc0010201 but it is not at
fault here since it expects these MSRs to be available as guest CPUID
0x80000001[ECX].PerfCtrExtCore is set. This feature bit should have been
cleared by QEMU when "pmu=off".
It boots up when "pmu=off" is combined with "perfctr-core=off" and
"perfmon-v2=off". The missing pieces here are the following:
https://lore.kernel.org/qemu-devel/20251111061532.36702-2-dongli.zhang@oracle.com/
https://lore.kernel.org/qemu-devel/20251111061532.36702-3-dongli.zhang@oracle.com/
Dongli, Zhao,
I see a discussion about compat array at
https://lore.kernel.org/qemu-devel/aUEXkDDOba+oZ4v+@intel.com/
Is this ready for the two patches above to be pulled in?
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [PATCH v8 1/7] target/i386/kvm: set KVM_PMU_CAP_DISABLE if "-pmu" is configured
2026-10-06 8:41 ` Sandipan Das
@ 2026-10-07 0:51 ` Dongli Zhang
0 siblings, 0 replies; 25+ messages in thread
From: Dongli Zhang @ 2026-10-07 0:51 UTC (permalink / raw)
To: Sandipan Das, zhao1.liu
Cc: pbonzini, mtosatti, babu.moger, likexu, like.xu.linux, groug,
khorenko, alexander.ivanov, den, davydov-max, xiaoyao.li,
dapeng1.mi, joe.jin, ewanhai-oc, ewanhai, Michael Tokarev,
qemu-devel, kvm
On Tue, Oct 6, 2026 1:41:05AM -0700, Sandipan Das wrote:
> On 01-10-2026 16: 42, Sandipan Das wrote: > On 01-10-2026 16: 35, Michael
> Tokarev wrote: >> On 10/1/26 13: 43, Sandipan Das wrote: >>> On 01-10-2026 15:
> 30, Michael Tokarev wrote: >>>> On 10/1/26 11: 54, Sandipan Das
>
> On 01-10-2026 16:42, Sandipan Das wrote:
>> On 01-10-2026 16:35, Michael Tokarev wrote:
>>> On 10/1/26 13:43, Sandipan Das wrote:
>>>> On 01-10-2026 15:30, Michael Tokarev wrote:
>>>>> On 10/1/26 11:54, Sandipan Das wrote:
>>>>>> On 01-10-2026 14:02, Michael Tokarev wrote:
>>>>> ...
>>>>>
>>>>>>> When booting windows11 VM with -cpu EPYC-Rome-v4 -enable-kvm, the
>>>>>>> guest enters tight CPU loop while displaying a boot screen, before
>>>>>>> it starts showing up the rotating circle (which shows boot progress).
>>>>>
>>>>>
>>>>>> What happens if you use "-cpu EPYC-Rome-v4,+pmu"?
>>>>>
>>>>> With that, it boots fine, just like before this patch.
>>>>> I checked current master too - same, +pmu fixes it.
>>>>>
>>>> A guest OS cannot expect a functioning PMU if the hypervisor has disabled
>>>> it. Linux runs a basic sanity test to determine this. See:
>>>> https://urldefense.com/v3/__https://github.com/torvalds/linux/blob/master/arch/
> x86/events/core.c*L268__;Iw!!ACWV5N9M2RV99hQ!
> PdJ7r8w0UUuD5iOlCxu7931wevMkDprX9yCV5mkfktbFkpOobtIWrkMt0Gmb9Ztnj3iZHwTK4UQHh85raOPEmgo$ <https://urldefense.com/v3/__https://github.com/torvalds/linux/blob/master/arch/x86/events/core.c*L268__;Iw!!ACWV5N9M2RV99hQ!PdJ7r8w0UUuD5iOlCxu7931wevMkDprX9yCV5mkfktbFkpOobtIWrkMt0Gmb9Ztnj3iZHwTK4UQHh85raOPEmgo$>
>>>
>>> So.. what do we do with all this? :)
>>>
>>> I kinda see the intention of the original change, but it results in
>>> non-bootable systems (linux too, as I think you demonstrate above).
>>>
>>
>> Because of the sanity test, Linux will boot but not register the "cpu" PMU.
>> Perhaps others should also run a similar sanity test on PMU hardware rather
>> than making assumptions and failing to boot.
>
> After looking at kvm_msr and kvm_cpuid traces, I see that Windows is indeed
> getting #GPs upon accessing MSRs 0xc0010200 and 0xc0010201 but it is not at
> fault here since it expects these MSRs to be available as guest CPUID
> 0x80000001[ECX].PerfCtrExtCore is set. This feature bit should have been
> cleared by QEMU when "pmu=off".
>
> It boots up when "pmu=off" is combined with "perfctr-core=off" and
> "perfmon-v2=off". The missing pieces here are the following:
>
> https://lore.kernel.org/qemu-devel/20251111061532.36702-2-dongli.zhang@oracle.com/
> https://lore.kernel.org/qemu-devel/20251111061532.36702-3-dongli.zhang@oracle.com/
>
> Dongli, Zhao,
>
> I see a discussion about compat array at
> https://lore.kernel.org/qemu-devel/aUEXkDDOba+oZ4v+@intel.com/
>
> Is this ready for the two patches above to be pulled in?
>
Thanks Sandipan for analyzing the issue!
Hi Zhao,
I noticed that the machine compat arrays are now available. Would you be able to
send updated versions of patches 1 and 2 with the compatibility property you
suggested?
Thank you very much!
Dongli Zhang
^ permalink raw reply [flat|nested] 25+ messages in thread
* [PATCH v8 2/7] target/i386/kvm: extract unrelated code out of kvm_x86_build_cpuid()
2025-12-30 7:42 [PATCH v8 0/7] target/i386/kvm/pmu: PMU Enhancement, Bugfix and Cleanup Dongli Zhang
2025-12-30 7:42 ` [PATCH v8 1/7] target/i386/kvm: set KVM_PMU_CAP_DISABLE if "-pmu" is configured Dongli Zhang
@ 2025-12-30 7:42 ` Dongli Zhang
2025-12-30 7:42 ` [PATCH v8 3/7] target/i386/kvm: rename architectural PMU variables Dongli Zhang
` (4 subsequent siblings)
6 siblings, 0 replies; 25+ messages in thread
From: Dongli Zhang @ 2025-12-30 7:42 UTC (permalink / raw)
To: qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
The initialization of 'has_architectural_pmu_version',
'num_architectural_pmu_gp_counters', and
'num_architectural_pmu_fixed_counters' is unrelated to the process of
building the CPUID.
Extract them out of kvm_x86_build_cpuid().
In addition, use cpuid_find_entry() instead of cpu_x86_cpuid(), because
CPUID has already been filled at this stage.
Signed-off-by: Dongli Zhang <dongli.zhang@oracle.com>
Reviewed-by: Zhao Liu <zhao1.liu@intel.com>
Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
Changed since v1:
- Still extract the code, but call them for all CPUs.
Changed since v2:
- Use cpuid_find_entry() instead of cpu_x86_cpuid().
- Didn't add Reviewed-by from Dapeng as the change isn't minor.
Changed since v6:
- Add Reviewed-by from Dapeng Mi.
target/i386/kvm/kvm.c | 62 ++++++++++++++++++++++++-------------------
1 file changed, 35 insertions(+), 27 deletions(-)
diff --git a/target/i386/kvm/kvm.c b/target/i386/kvm/kvm.c
index c98832f423..08d80ff677 100644
--- a/target/i386/kvm/kvm.c
+++ b/target/i386/kvm/kvm.c
@@ -1986,33 +1986,6 @@ uint32_t kvm_x86_build_cpuid(CPUX86State *env, struct kvm_cpuid_entry2 *entries,
}
}
- if (limit >= 0x0a) {
- uint32_t eax, edx;
-
- cpu_x86_cpuid(env, 0x0a, 0, &eax, &unused, &unused, &edx);
-
- has_architectural_pmu_version = eax & 0xff;
- if (has_architectural_pmu_version > 0) {
- num_architectural_pmu_gp_counters = (eax & 0xff00) >> 8;
-
- /* Shouldn't be more than 32, since that's the number of bits
- * available in EBX to tell us _which_ counters are available.
- * Play it safe.
- */
- if (num_architectural_pmu_gp_counters > MAX_GP_COUNTERS) {
- num_architectural_pmu_gp_counters = MAX_GP_COUNTERS;
- }
-
- if (has_architectural_pmu_version > 1) {
- num_architectural_pmu_fixed_counters = edx & 0x1f;
-
- if (num_architectural_pmu_fixed_counters > MAX_FIXED_COUNTERS) {
- num_architectural_pmu_fixed_counters = MAX_FIXED_COUNTERS;
- }
- }
- }
- }
-
cpu_x86_cpuid(env, 0x80000000, 0, &limit, &unused, &unused, &unused);
for (i = 0x80000000; i <= limit; i++) {
@@ -2116,6 +2089,39 @@ int kvm_arch_pre_create_vcpu(CPUState *cpu, Error **errp)
return 0;
}
+static void kvm_init_pmu_info(struct kvm_cpuid2 *cpuid)
+{
+ struct kvm_cpuid_entry2 *c;
+
+ c = cpuid_find_entry(cpuid, 0xa, 0);
+
+ if (!c) {
+ return;
+ }
+
+ has_architectural_pmu_version = c->eax & 0xff;
+ if (has_architectural_pmu_version > 0) {
+ num_architectural_pmu_gp_counters = (c->eax & 0xff00) >> 8;
+
+ /*
+ * Shouldn't be more than 32, since that's the number of bits
+ * available in EBX to tell us _which_ counters are available.
+ * Play it safe.
+ */
+ if (num_architectural_pmu_gp_counters > MAX_GP_COUNTERS) {
+ num_architectural_pmu_gp_counters = MAX_GP_COUNTERS;
+ }
+
+ if (has_architectural_pmu_version > 1) {
+ num_architectural_pmu_fixed_counters = c->edx & 0x1f;
+
+ if (num_architectural_pmu_fixed_counters > MAX_FIXED_COUNTERS) {
+ num_architectural_pmu_fixed_counters = MAX_FIXED_COUNTERS;
+ }
+ }
+ }
+}
+
int kvm_arch_init_vcpu(CPUState *cs)
{
struct {
@@ -2306,6 +2312,8 @@ int kvm_arch_init_vcpu(CPUState *cs)
cpuid_i = kvm_x86_build_cpuid(env, cpuid_data.entries, cpuid_i);
cpuid_data.cpuid.nent = cpuid_i;
+ kvm_init_pmu_info(&cpuid_data.cpuid);
+
if (x86_cpu_family(env->cpuid_version) >= 6
&& (env->features[FEAT_1_EDX] & (CPUID_MCE | CPUID_MCA)) ==
(CPUID_MCE | CPUID_MCA)) {
--
2.39.3
^ permalink raw reply related [flat|nested] 25+ messages in thread* [PATCH v8 3/7] target/i386/kvm: rename architectural PMU variables
2025-12-30 7:42 [PATCH v8 0/7] target/i386/kvm/pmu: PMU Enhancement, Bugfix and Cleanup Dongli Zhang
2025-12-30 7:42 ` [PATCH v8 1/7] target/i386/kvm: set KVM_PMU_CAP_DISABLE if "-pmu" is configured Dongli Zhang
2025-12-30 7:42 ` [PATCH v8 2/7] target/i386/kvm: extract unrelated code out of kvm_x86_build_cpuid() Dongli Zhang
@ 2025-12-30 7:42 ` Dongli Zhang
2025-12-30 7:42 ` [PATCH v8 4/7] target/i386/kvm: query kvm.enable_pmu parameter Dongli Zhang
` (3 subsequent siblings)
6 siblings, 0 replies; 25+ messages in thread
From: Dongli Zhang @ 2025-12-30 7:42 UTC (permalink / raw)
To: qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
AMD does not have what is commonly referred to as an architectural PMU.
Therefore, we need to rename the following variables to be applicable for
both Intel and AMD:
- has_architectural_pmu_version
- num_architectural_pmu_gp_counters
- num_architectural_pmu_fixed_counters
For Intel processors, the meaning of pmu_version remains unchanged.
For AMD processors:
pmu_version == 1 corresponds to versions before AMD PerfMonV2.
pmu_version == 2 corresponds to AMD PerfMonV2.
Signed-off-by: Dongli Zhang <dongli.zhang@oracle.com>
Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
Reviewed-by: Zhao Liu <zhao1.liu@intel.com>
Reviewed-by: Sandipan Das <sandipan.das@amd.com>
---
Changed since v2:
- Change has_pmu_version to pmu_version.
- Add Reviewed-by since the change is minor.
- As a reminder, there are some contextual change due to PATCH 05,
i.e., c->edx vs. edx.
Changed since v6:
- Add Reviewed-by from Sandipan.
target/i386/kvm/kvm.c | 49 ++++++++++++++++++++++++-------------------
1 file changed, 28 insertions(+), 21 deletions(-)
diff --git a/target/i386/kvm/kvm.c b/target/i386/kvm/kvm.c
index 08d80ff677..3b803c662d 100644
--- a/target/i386/kvm/kvm.c
+++ b/target/i386/kvm/kvm.c
@@ -167,9 +167,16 @@ static bool has_msr_perf_capabs;
static bool has_msr_pkrs;
static bool has_msr_hwcr;
-static uint32_t has_architectural_pmu_version;
-static uint32_t num_architectural_pmu_gp_counters;
-static uint32_t num_architectural_pmu_fixed_counters;
+/*
+ * For Intel processors, the meaning is the architectural PMU version
+ * number.
+ *
+ * For AMD processors: 1 corresponds to the prior versions, and 2
+ * corresponds to AMD PerfMonV2.
+ */
+static uint32_t pmu_version;
+static uint32_t num_pmu_gp_counters;
+static uint32_t num_pmu_fixed_counters;
static int has_xsave2;
static int has_xcrs;
@@ -2099,24 +2106,24 @@ static void kvm_init_pmu_info(struct kvm_cpuid2 *cpuid)
return;
}
- has_architectural_pmu_version = c->eax & 0xff;
- if (has_architectural_pmu_version > 0) {
- num_architectural_pmu_gp_counters = (c->eax & 0xff00) >> 8;
+ pmu_version = c->eax & 0xff;
+ if (pmu_version > 0) {
+ num_pmu_gp_counters = (c->eax & 0xff00) >> 8;
/*
* Shouldn't be more than 32, since that's the number of bits
* available in EBX to tell us _which_ counters are available.
* Play it safe.
*/
- if (num_architectural_pmu_gp_counters > MAX_GP_COUNTERS) {
- num_architectural_pmu_gp_counters = MAX_GP_COUNTERS;
+ if (num_pmu_gp_counters > MAX_GP_COUNTERS) {
+ num_pmu_gp_counters = MAX_GP_COUNTERS;
}
- if (has_architectural_pmu_version > 1) {
- num_architectural_pmu_fixed_counters = c->edx & 0x1f;
+ if (pmu_version > 1) {
+ num_pmu_fixed_counters = c->edx & 0x1f;
- if (num_architectural_pmu_fixed_counters > MAX_FIXED_COUNTERS) {
- num_architectural_pmu_fixed_counters = MAX_FIXED_COUNTERS;
+ if (num_pmu_fixed_counters > MAX_FIXED_COUNTERS) {
+ num_pmu_fixed_counters = MAX_FIXED_COUNTERS;
}
}
}
@@ -4087,25 +4094,25 @@ static int kvm_put_msrs(X86CPU *cpu, KvmPutState level)
kvm_msr_entry_add(cpu, MSR_KVM_POLL_CONTROL, env->poll_control_msr);
}
- if (has_architectural_pmu_version > 0) {
- if (has_architectural_pmu_version > 1) {
+ if (pmu_version > 0) {
+ if (pmu_version > 1) {
/* Stop the counter. */
kvm_msr_entry_add(cpu, MSR_CORE_PERF_FIXED_CTR_CTRL, 0);
kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_CTRL, 0);
}
/* Set the counter values. */
- for (i = 0; i < num_architectural_pmu_fixed_counters; i++) {
+ for (i = 0; i < num_pmu_fixed_counters; i++) {
kvm_msr_entry_add(cpu, MSR_CORE_PERF_FIXED_CTR0 + i,
env->msr_fixed_counters[i]);
}
- for (i = 0; i < num_architectural_pmu_gp_counters; i++) {
+ for (i = 0; i < num_pmu_gp_counters; i++) {
kvm_msr_entry_add(cpu, MSR_P6_PERFCTR0 + i,
env->msr_gp_counters[i]);
kvm_msr_entry_add(cpu, MSR_P6_EVNTSEL0 + i,
env->msr_gp_evtsel[i]);
}
- if (has_architectural_pmu_version > 1) {
+ if (pmu_version > 1) {
kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_STATUS,
env->msr_global_status);
kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_OVF_CTRL,
@@ -4622,17 +4629,17 @@ static int kvm_get_msrs(X86CPU *cpu)
if (env->features[FEAT_KVM] & CPUID_KVM_POLL_CONTROL) {
kvm_msr_entry_add(cpu, MSR_KVM_POLL_CONTROL, 1);
}
- if (has_architectural_pmu_version > 0) {
- if (has_architectural_pmu_version > 1) {
+ if (pmu_version > 0) {
+ if (pmu_version > 1) {
kvm_msr_entry_add(cpu, MSR_CORE_PERF_FIXED_CTR_CTRL, 0);
kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_CTRL, 0);
kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_STATUS, 0);
kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_OVF_CTRL, 0);
}
- for (i = 0; i < num_architectural_pmu_fixed_counters; i++) {
+ for (i = 0; i < num_pmu_fixed_counters; i++) {
kvm_msr_entry_add(cpu, MSR_CORE_PERF_FIXED_CTR0 + i, 0);
}
- for (i = 0; i < num_architectural_pmu_gp_counters; i++) {
+ for (i = 0; i < num_pmu_gp_counters; i++) {
kvm_msr_entry_add(cpu, MSR_P6_PERFCTR0 + i, 0);
kvm_msr_entry_add(cpu, MSR_P6_EVNTSEL0 + i, 0);
}
--
2.39.3
^ permalink raw reply related [flat|nested] 25+ messages in thread* [PATCH v8 4/7] target/i386/kvm: query kvm.enable_pmu parameter
2025-12-30 7:42 [PATCH v8 0/7] target/i386/kvm/pmu: PMU Enhancement, Bugfix and Cleanup Dongli Zhang
` (2 preceding siblings ...)
2025-12-30 7:42 ` [PATCH v8 3/7] target/i386/kvm: rename architectural PMU variables Dongli Zhang
@ 2025-12-30 7:42 ` Dongli Zhang
2026-01-02 22:59 ` Chen, Zide
2025-12-30 7:42 ` [PATCH v8 5/7] target/i386/kvm: reset AMD PMU registers during VM reset Dongli Zhang
` (2 subsequent siblings)
6 siblings, 1 reply; 25+ messages in thread
From: Dongli Zhang @ 2025-12-30 7:42 UTC (permalink / raw)
To: qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
When PMU is enabled in QEMU, there is a chance that PMU virtualization is
completely disabled by the KVM module parameter kvm.enable_pmu=N.
The kvm.enable_pmu parameter is introduced since Linux v5.17.
Its permission is 0444. It does not change until a reload of the KVM
module.
Read the kvm.enable_pmu value from the module sysfs to give a chance to
provide more information about vPMU enablement.
Signed-off-by: Dongli Zhang <dongli.zhang@oracle.com>
Reviewed-by: Zhao Liu <zhao1.liu@intel.com>
Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
Changed since v2:
- Rework the code flow following Zhao's suggestion.
- Return error when:
(*kvm_enable_pmu == 'N' && X86_CPU(cpu)->enable_pmu)
Changed since v3:
- Re-split the cases into enable_pmu and !enable_pmu, following Zhao's
suggestion.
- Rework the commit messages.
- Bring back global static variable 'kvm_pmu_disabled' from v2.
Changed since v4:
- Add Reviewed-by from Zhao.
Changed since v5:
- Rebase on top of most recent QEMU.
Changed since v6:
- Add Reviewed-by from Dapeng Mi.
target/i386/kvm/kvm.c | 61 +++++++++++++++++++++++++++++++------------
1 file changed, 44 insertions(+), 17 deletions(-)
diff --git a/target/i386/kvm/kvm.c b/target/i386/kvm/kvm.c
index 3b803c662d..338b9558e4 100644
--- a/target/i386/kvm/kvm.c
+++ b/target/i386/kvm/kvm.c
@@ -187,6 +187,10 @@ static int has_triple_fault_event;
static bool has_msr_mcg_ext_ctl;
static int pmu_cap;
+/*
+ * Read from /sys/module/kvm/parameters/enable_pmu.
+ */
+static bool kvm_pmu_disabled;
static struct kvm_cpuid2 *cpuid_cache;
static struct kvm_cpuid2 *hv_cpuid_cache;
@@ -2068,23 +2072,30 @@ int kvm_arch_pre_create_vcpu(CPUState *cpu, Error **errp)
if (first) {
first = false;
- /*
- * Since Linux v5.18, KVM provides a VM-level capability to easily
- * disable PMUs; however, QEMU has been providing PMU property per
- * CPU since v1.6. In order to accommodate both, have to configure
- * the VM-level capability here.
- *
- * KVM_PMU_CAP_DISABLE doesn't change the PMU
- * behavior on Intel platform because current "pmu" property works
- * as expected.
- */
- if ((pmu_cap & KVM_PMU_CAP_DISABLE) && !X86_CPU(cpu)->enable_pmu) {
- ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
- KVM_PMU_CAP_DISABLE);
- if (ret < 0) {
- error_setg_errno(errp, -ret,
- "Failed to set KVM_PMU_CAP_DISABLE");
- return ret;
+ if (X86_CPU(cpu)->enable_pmu) {
+ if (kvm_pmu_disabled) {
+ warn_report("Failed to enable PMU since "
+ "KVM's enable_pmu parameter is disabled");
+ }
+ } else {
+ /*
+ * Since Linux v5.18, KVM provides a VM-level capability to easily
+ * disable PMUs; however, QEMU has been providing PMU property per
+ * CPU since v1.6. In order to accommodate both, have to configure
+ * the VM-level capability here.
+ *
+ * KVM_PMU_CAP_DISABLE doesn't change the PMU
+ * behavior on Intel platform because current "pmu" property works
+ * as expected.
+ */
+ if (pmu_cap & KVM_PMU_CAP_DISABLE) {
+ ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
+ KVM_PMU_CAP_DISABLE);
+ if (ret < 0) {
+ error_setg_errno(errp, -ret,
+ "Failed to set KVM_PMU_CAP_DISABLE");
+ return ret;
+ }
}
}
}
@@ -3302,6 +3313,7 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
int ret;
struct utsname utsname;
Error *local_err = NULL;
+ g_autofree char *kvm_enable_pmu;
/*
* Initialize confidential guest (SEV/TDX) context, if required
@@ -3437,6 +3449,21 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
pmu_cap = kvm_check_extension(s, KVM_CAP_PMU_CAPABILITY);
+ /*
+ * The enable_pmu parameter is introduced since Linux v5.17,
+ * give a chance to provide more information about vPMU
+ * enablement.
+ *
+ * The kvm.enable_pmu's permission is 0444. It does not change
+ * until a reload of the KVM module.
+ */
+ if (g_file_get_contents("/sys/module/kvm/parameters/enable_pmu",
+ &kvm_enable_pmu, NULL, NULL)) {
+ if (*kvm_enable_pmu == 'N') {
+ kvm_pmu_disabled = true;
+ }
+ }
+
return 0;
}
--
2.39.3
^ permalink raw reply related [flat|nested] 25+ messages in thread* Re: [PATCH v8 4/7] target/i386/kvm: query kvm.enable_pmu parameter
2025-12-30 7:42 ` [PATCH v8 4/7] target/i386/kvm: query kvm.enable_pmu parameter Dongli Zhang
@ 2026-01-02 22:59 ` Chen, Zide
2026-01-05 20:21 ` Dongli Zhang
0 siblings, 1 reply; 25+ messages in thread
From: Chen, Zide @ 2026-01-02 22:59 UTC (permalink / raw)
To: Dongli Zhang, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
On 12/29/2025 11:42 PM, Dongli Zhang wrote:
> When PMU is enabled in QEMU, there is a chance that PMU virtualization is
> completely disabled by the KVM module parameter kvm.enable_pmu=N.
>
> The kvm.enable_pmu parameter is introduced since Linux v5.17.
> Its permission is 0444. It does not change until a reload of the KVM
> module.
>
> Read the kvm.enable_pmu value from the module sysfs to give a chance to
> provide more information about vPMU enablement.> Signed-off-by: Dongli Zhang <dongli.zhang@oracle.com>
> Reviewed-by: Zhao Liu <zhao1.liu@intel.com>
> Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
> ---
> Changed since v2:
> - Rework the code flow following Zhao's suggestion.
> - Return error when:
> (*kvm_enable_pmu == 'N' && X86_CPU(cpu)->enable_pmu)
> Changed since v3:
> - Re-split the cases into enable_pmu and !enable_pmu, following Zhao's
> suggestion.
> - Rework the commit messages.
> - Bring back global static variable 'kvm_pmu_disabled' from v2.
> Changed since v4:
> - Add Reviewed-by from Zhao.
> Changed since v5:
> - Rebase on top of most recent QEMU.
> Changed since v6:
> - Add Reviewed-by from Dapeng Mi.
>
> target/i386/kvm/kvm.c | 61 +++++++++++++++++++++++++++++++------------
> 1 file changed, 44 insertions(+), 17 deletions(-)
>
> diff --git a/target/i386/kvm/kvm.c b/target/i386/kvm/kvm.c
> index 3b803c662d..338b9558e4 100644
> --- a/target/i386/kvm/kvm.c
> +++ b/target/i386/kvm/kvm.c
> @@ -187,6 +187,10 @@ static int has_triple_fault_event;
> static bool has_msr_mcg_ext_ctl;
>
> static int pmu_cap;
> +/*
> + * Read from /sys/module/kvm/parameters/enable_pmu.
> + */
> +static bool kvm_pmu_disabled;
>
> static struct kvm_cpuid2 *cpuid_cache;
> static struct kvm_cpuid2 *hv_cpuid_cache;
> @@ -2068,23 +2072,30 @@ int kvm_arch_pre_create_vcpu(CPUState *cpu, Error **errp)
> if (first) {
> first = false;
>
> - /*
> - * Since Linux v5.18, KVM provides a VM-level capability to easily
> - * disable PMUs; however, QEMU has been providing PMU property per
> - * CPU since v1.6. In order to accommodate both, have to configure
> - * the VM-level capability here.
> - *
> - * KVM_PMU_CAP_DISABLE doesn't change the PMU
> - * behavior on Intel platform because current "pmu" property works
> - * as expected.
> - */
> - if ((pmu_cap & KVM_PMU_CAP_DISABLE) && !X86_CPU(cpu)->enable_pmu) {
> - ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
> - KVM_PMU_CAP_DISABLE);
> - if (ret < 0) {
> - error_setg_errno(errp, -ret,
> - "Failed to set KVM_PMU_CAP_DISABLE");
> - return ret;
> + if (X86_CPU(cpu)->enable_pmu) {
> + if (kvm_pmu_disabled) {
> + warn_report("Failed to enable PMU since "
> + "KVM's enable_pmu parameter is disabled");
I'm wondering about the intended value of this patch?
If enable_pmu is true in QEMU but the corresponding KVM parameter is
false, then KVM_GET_SUPPORTED_CPUID or KVM_GET_MSRS should be able to
tell that the PMU feature is not supported by host.
The logic implemented in this patch seems somewhat redundant.
Additionally, in this scenario — where the user intends to enable a
feature but the host cannot support it — normally no warning is emitted
by QEMU.
> + }
> + } else {
> + /*
> + * Since Linux v5.18, KVM provides a VM-level capability to easily
> + * disable PMUs; however, QEMU has been providing PMU property per
> + * CPU since v1.6. In order to accommodate both, have to configure
> + * the VM-level capability here.
> + *
> + * KVM_PMU_CAP_DISABLE doesn't change the PMU
> + * behavior on Intel platform because current "pmu" property works
> + * as expected.
> + */
> + if (pmu_cap & KVM_PMU_CAP_DISABLE) {
> + ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
> + KVM_PMU_CAP_DISABLE);
> + if (ret < 0) {
> + error_setg_errno(errp, -ret,
> + "Failed to set KVM_PMU_CAP_DISABLE");
> + return ret;
> + }
> }
> }
> }
> @@ -3302,6 +3313,7 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
> int ret;
> struct utsname utsname;
> Error *local_err = NULL;
> + g_autofree char *kvm_enable_pmu;
>
> /*
> * Initialize confidential guest (SEV/TDX) context, if required
> @@ -3437,6 +3449,21 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
>
> pmu_cap = kvm_check_extension(s, KVM_CAP_PMU_CAPABILITY);
>
> + /*
> + * The enable_pmu parameter is introduced since Linux v5.17,
> + * give a chance to provide more information about vPMU
> + * enablement.
> + *
> + * The kvm.enable_pmu's permission is 0444. It does not change
> + * until a reload of the KVM module.
> + */
> + if (g_file_get_contents("/sys/module/kvm/parameters/enable_pmu",
> + &kvm_enable_pmu, NULL, NULL)) {
> + if (*kvm_enable_pmu == 'N') {
> + kvm_pmu_disabled = true;
It’s generally better not to rely on KVM’s internal implementation
unless really necessary.
For example, in the new mediated vPMU framework, even if the KVM module
parameter enable_pmu is set, the per-guest kvm->arch.enable_pmu could
still be cleared.
In such a case, the logic here might not be correct.
> + }
> + }
> +
> return 0;
> }
>
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v8 4/7] target/i386/kvm: query kvm.enable_pmu parameter
2026-01-02 22:59 ` Chen, Zide
@ 2026-01-05 20:21 ` Dongli Zhang
2026-01-06 21:03 ` Chen, Zide
0 siblings, 1 reply; 25+ messages in thread
From: Dongli Zhang @ 2026-01-05 20:21 UTC (permalink / raw)
To: Chen, Zide, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
Hi Zide,
On 1/2/26 2:59 PM, Chen, Zide wrote:
>
>
> On 12/29/2025 11:42 PM, Dongli Zhang wrote:
[snip]
>>
>> static struct kvm_cpuid2 *cpuid_cache;
>> static struct kvm_cpuid2 *hv_cpuid_cache;
>> @@ -2068,23 +2072,30 @@ int kvm_arch_pre_create_vcpu(CPUState *cpu, Error **errp)
>> if (first) {
>> first = false;
>>
>> - /*
>> - * Since Linux v5.18, KVM provides a VM-level capability to easily
>> - * disable PMUs; however, QEMU has been providing PMU property per
>> - * CPU since v1.6. In order to accommodate both, have to configure
>> - * the VM-level capability here.
>> - *
>> - * KVM_PMU_CAP_DISABLE doesn't change the PMU
>> - * behavior on Intel platform because current "pmu" property works
>> - * as expected.
>> - */
>> - if ((pmu_cap & KVM_PMU_CAP_DISABLE) && !X86_CPU(cpu)->enable_pmu) {
>> - ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
>> - KVM_PMU_CAP_DISABLE);
>> - if (ret < 0) {
>> - error_setg_errno(errp, -ret,
>> - "Failed to set KVM_PMU_CAP_DISABLE");
>> - return ret;
>> + if (X86_CPU(cpu)->enable_pmu) {
>> + if (kvm_pmu_disabled) {
>> + warn_report("Failed to enable PMU since "
>> + "KVM's enable_pmu parameter is disabled");
>
> I'm wondering about the intended value of this patch?
>
> If enable_pmu is true in QEMU but the corresponding KVM parameter is
> false, then KVM_GET_SUPPORTED_CPUID or KVM_GET_MSRS should be able to
> tell that the PMU feature is not supported by host.
>
> The logic implemented in this patch seems somewhat redundant.
For Intel, the QEMU userspace can determine if the vPMU is disabled by KVM
through the use of KVM_GET_SUPPORTED_CPUID.
However, this approach does not apply to AMD. Unlike Intel, AMD does not rely on
CPUID to detect whether PMU is supported. By default, we can assume that PMU is
always available, except for the recent PerfMonV2 feature.
The main objective of this PATCH 4/7 is to introduce the variable
'kvm_pmu_disabled', which will be reused in PATCH 5/7 to skip any PMU
initialization if the parameter is set to 'N'.
+static void kvm_init_pmu_info(struct kvm_cpuid2 *cpuid, X86CPU *cpu)
+{
+ CPUX86State *env = &cpu->env;
+
+ /*
+ * The PMU virtualization is disabled by kvm.enable_pmu=N.
+ */
+ if (kvm_pmu_disabled) {
+ return;
+ }
The 'kvm_pmu_disabled' variable is used to differentiate between the following
two scenarios on AMD:
(1) A newer KVM with KVM_PMU_CAP_DISABLE support, but explicitly disabled via
the KVM parameter ('N').
(2) An older KVM without KVM_CAP_PMU_CAPABILITY support.
In both cases, the call to KVM_CAP_PMU_CAPABILITY extension support check may
return 0.
By reading the file "/sys/module/kvm/parameters/enable_pmu", we can distinguish
between these two scenarios.
As you mentioned, another approach would be to use KVM_GET_MSRS to specifically
probe for AMD during QEMU initialization. In this case, we can set
'kvm_pmu_disabled' to true if reading the AMD PMU MSR registers fails.
To implement this, we may need to:
1. Turn this patch to be AMD specific by probing the AMD PMU registers during
initialization. We may need go create a new function in QEMU to use KVM_GET_MSRS
for probing only, or we may re-use kvm_arch_get_supported_msr_feature() or
kvm_get_one_msr(). I may change in the next version.
2. Limit the usage of 'kvm_pmu_disabled' to be AMD specific in PATCH 5/7.
>
> Additionally, in this scenario — where the user intends to enable a
> feature but the host cannot support it — normally no warning is emitted
> by QEMU.
According to the usage of QEMU, may I assume QEMU already prints warning logs
for unsupported features? The below is an example.
QEMU 10.2.50 monitor - type 'help' for more information
qemu-system-x86_64: warning: host doesn't support requested feature:
CPUID[eax=07h,ecx=00h].EBX.hle [bit 4]
qemu-system-x86_64: warning: host doesn't support requested feature:
CPUID[eax=07h,ecx=00h].EBX.rtm [bit 11]
>
>> + }
>> + } else {
>> + /*
>> + * Since Linux v5.18, KVM provides a VM-level capability to easily
>> + * disable PMUs; however, QEMU has been providing PMU property per
>> + * CPU since v1.6. In order to accommodate both, have to configure
>> + * the VM-level capability here.
>> + *
>> + * KVM_PMU_CAP_DISABLE doesn't change the PMU
>> + * behavior on Intel platform because current "pmu" property works
>> + * as expected.
>> + */
>> + if (pmu_cap & KVM_PMU_CAP_DISABLE) {
>> + ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
>> + KVM_PMU_CAP_DISABLE);
>> + if (ret < 0) {
>> + error_setg_errno(errp, -ret,
>> + "Failed to set KVM_PMU_CAP_DISABLE");
>> + return ret;
>> + }
>> }
>> }
>> }
>> @@ -3302,6 +3313,7 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
>> int ret;
>> struct utsname utsname;
>> Error *local_err = NULL;
>> + g_autofree char *kvm_enable_pmu;
>>
>> /*
>> * Initialize confidential guest (SEV/TDX) context, if required
>> @@ -3437,6 +3449,21 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
>>
>> pmu_cap = kvm_check_extension(s, KVM_CAP_PMU_CAPABILITY);
>>
>> + /*
>> + * The enable_pmu parameter is introduced since Linux v5.17,
>> + * give a chance to provide more information about vPMU
>> + * enablement.
>> + *
>> + * The kvm.enable_pmu's permission is 0444. It does not change
>> + * until a reload of the KVM module.
>> + */
>> + if (g_file_get_contents("/sys/module/kvm/parameters/enable_pmu",
>> + &kvm_enable_pmu, NULL, NULL)) {
>> + if (*kvm_enable_pmu == 'N') {
>> + kvm_pmu_disabled = true;
>
> It’s generally better not to rely on KVM’s internal implementation
> unless really necessary.
>
> For example, in the new mediated vPMU framework, even if the KVM module
> parameter enable_pmu is set, the per-guest kvm->arch.enable_pmu could
> still be cleared.
>
> In such a case, the logic here might not be correct.
Would the Mediated vPMU set KVM_PMU_CAP_DISABLE to clear per-VM enable_pmu even
when the global KVM parameter enable_pmu=N is set?
In this scenario, we plan to rely on KVM_PMU_CAP_DISABLE only when the value of
"/sys/module/kvm/parameters/enable_pmu" is not equal to N.
Can I assume that this will work with Mediated vPMU?
Is there any possibility to follow the current approach before Mediated vPMU is
finalized for mainline, and later introduce an incremental change using
KVM_GET_MSRS probing? The current approach is straightforward and can work with
existing Linux kernel source code.
For quite some time, QEMU has lacked support for disabling or resetting AMD PMU
registers. If we could add this feature before Mediated vPMU is finalized, it
would benefit many existing kernel versions. This patchset solves production bugs.
Feel free to let me know your thought, while I would starting working on next
version now.
Thank you very much!
Dongli Zhang
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v8 4/7] target/i386/kvm: query kvm.enable_pmu parameter
2026-01-05 20:21 ` Dongli Zhang
@ 2026-01-06 21:03 ` Chen, Zide
2026-01-07 8:05 ` Dongli Zhang
0 siblings, 1 reply; 25+ messages in thread
From: Chen, Zide @ 2026-01-06 21:03 UTC (permalink / raw)
To: Dongli Zhang, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
On 1/5/2026 12:21 PM, Dongli Zhang wrote:
> Hi Zide,
>
> On 1/2/26 2:59 PM, Chen, Zide wrote:
>>
>>
>> On 12/29/2025 11:42 PM, Dongli Zhang wrote:
>
> [snip]
>
>>>
>>> static struct kvm_cpuid2 *cpuid_cache;
>>> static struct kvm_cpuid2 *hv_cpuid_cache;
>>> @@ -2068,23 +2072,30 @@ int kvm_arch_pre_create_vcpu(CPUState *cpu, Error **errp)
>>> if (first) {
>>> first = false;
>>>
>>> - /*
>>> - * Since Linux v5.18, KVM provides a VM-level capability to easily
>>> - * disable PMUs; however, QEMU has been providing PMU property per
>>> - * CPU since v1.6. In order to accommodate both, have to configure
>>> - * the VM-level capability here.
>>> - *
>>> - * KVM_PMU_CAP_DISABLE doesn't change the PMU
>>> - * behavior on Intel platform because current "pmu" property works
>>> - * as expected.
>>> - */
>>> - if ((pmu_cap & KVM_PMU_CAP_DISABLE) && !X86_CPU(cpu)->enable_pmu) {
>>> - ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
>>> - KVM_PMU_CAP_DISABLE);
>>> - if (ret < 0) {
>>> - error_setg_errno(errp, -ret,
>>> - "Failed to set KVM_PMU_CAP_DISABLE");
>>> - return ret;
>>> + if (X86_CPU(cpu)->enable_pmu) {
>>> + if (kvm_pmu_disabled) {
>>> + warn_report("Failed to enable PMU since "
>>> + "KVM's enable_pmu parameter is disabled");
>>
>> I'm wondering about the intended value of this patch?
>>
>> If enable_pmu is true in QEMU but the corresponding KVM parameter is
>> false, then KVM_GET_SUPPORTED_CPUID or KVM_GET_MSRS should be able to
>> tell that the PMU feature is not supported by host.
>>
>> The logic implemented in this patch seems somewhat redundant.
>
> For Intel, the QEMU userspace can determine if the vPMU is disabled by KVM
> through the use of KVM_GET_SUPPORTED_CPUID.
>
> However, this approach does not apply to AMD. Unlike Intel, AMD does not rely on
> CPUID to detect whether PMU is supported. By default, we can assume that PMU is
> always available, except for the recent PerfMonV2 feature.
>
> The main objective of this PATCH 4/7 is to introduce the variable
> 'kvm_pmu_disabled', which will be reused in PATCH 5/7 to skip any PMU
> initialization if the parameter is set to 'N'.
>
> +static void kvm_init_pmu_info(struct kvm_cpuid2 *cpuid, X86CPU *cpu)
> +{
> + CPUX86State *env = &cpu->env;
> +
> + /*
> + * The PMU virtualization is disabled by kvm.enable_pmu=N.
> + */
> + if (kvm_pmu_disabled) {
> + return;
> + }
Thanks for explanation.
> The 'kvm_pmu_disabled' variable is used to differentiate between the following
> two scenarios on AMD:
>
> (1) A newer KVM with KVM_PMU_CAP_DISABLE support, but explicitly disabled via
> the KVM parameter ('N').
>
> (2) An older KVM without KVM_CAP_PMU_CAPABILITY support.
>
> In both cases, the call to KVM_CAP_PMU_CAPABILITY extension support check may
> return 0.
>
> By reading the file "/sys/module/kvm/parameters/enable_pmu", we can distinguish
> between these two scenarios.
As described in PATCH 1/7, without issuing KVM_PMU_CAP_DISABLE, KVM has
no way to know that userspace does not intend to enable vPMU in AMD
platforms, and therefore does not fault guest accesses to PMU MSRs.
My understanding is that the issue being addressed here is basically the
opposite: QEMU does not know that vPMU is disabled by KVM.
IIUC, one difference between Intel and AMD is that AMD lacks a CPUID
leaf to indicate the availability of PMU version 1. But Intel
potentially could be in the same situation that KVM advertises PMU
availability but it's not actually supported. (e.g. kvm->arch.enable_pmu
is false while modules parameter enable_pmu is true).
From the guest’s point of view, it probes PMU MSRs to determine whether
PMU support is present and it's fine in this situation.
In userspace, QEMU may issue KVM_SET_MSRS / KVM_GET_MSRS to KVM without
knowing that vPMU has been disabled by KVM. I think these IOCTLs should
not fail, since KVM states that “Userspace is allowed to read MSRs, and
write ‘0’ to MSRs, that KVM advertises to userspace, even if an MSR
isn’t fully supported.”
My current understanding is that AMD should be fine even without
kvm_pmu_disabled, but I may be missing some context here.
The bottom line is this patch doesn't handle the cases that KVM still
could disable vPMU support even if enable_pmu is true.
> As you mentioned, another approach would be to use KVM_GET_MSRS to specifically
> probe for AMD during QEMU initialization. In this case, we can set
> 'kvm_pmu_disabled' to true if reading the AMD PMU MSR registers fails.
>
> To implement this, we may need to:
>
> 1. Turn this patch to be AMD specific by probing the AMD PMU registers during
> initialization. We may need go create a new function in QEMU to use KVM_GET_MSRS
> for probing only, or we may re-use kvm_arch_get_supported_msr_feature() or
> kvm_get_one_msr(). I may change in the next version.
>
> 2. Limit the usage of 'kvm_pmu_disabled' to be AMD specific in PATCH 5/7.
I guess this might make things more complicated.
>>
>> Additionally, in this scenario — where the user intends to enable a
>> feature but the host cannot support it — normally no warning is emitted
>> by QEMU.
>
> According to the usage of QEMU, may I assume QEMU already prints warning logs
> for unsupported features? The below is an example.
>
> QEMU 10.2.50 monitor - type 'help' for more information
> qemu-system-x86_64: warning: host doesn't support requested feature:
> CPUID[eax=07h,ecx=00h].EBX.hle [bit 4]
> qemu-system-x86_64: warning: host doesn't support requested feature:
> CPUID[eax=07h,ecx=00h].EBX.rtm [bit 11]
>
>>
>>> + }
>>> + } else {
>>> + /*
>>> + * Since Linux v5.18, KVM provides a VM-level capability to easily
>>> + * disable PMUs; however, QEMU has been providing PMU property per
>>> + * CPU since v1.6. In order to accommodate both, have to configure
>>> + * the VM-level capability here.
>>> + *
>>> + * KVM_PMU_CAP_DISABLE doesn't change the PMU
>>> + * behavior on Intel platform because current "pmu" property works
>>> + * as expected.
>>> + */
>>> + if (pmu_cap & KVM_PMU_CAP_DISABLE) {
>>> + ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
>>> + KVM_PMU_CAP_DISABLE);
>>> + if (ret < 0) {
>>> + error_setg_errno(errp, -ret,
>>> + "Failed to set KVM_PMU_CAP_DISABLE");
>>> + return ret;
>>> + }
>>> }
>>> }
>>> }
>>> @@ -3302,6 +3313,7 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
>>> int ret;
>>> struct utsname utsname;
>>> Error *local_err = NULL;
>>> + g_autofree char *kvm_enable_pmu;
>>>
>>> /*
>>> * Initialize confidential guest (SEV/TDX) context, if required
>>> @@ -3437,6 +3449,21 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
>>>
>>> pmu_cap = kvm_check_extension(s, KVM_CAP_PMU_CAPABILITY);
>>>
>>> + /*
>>> + * The enable_pmu parameter is introduced since Linux v5.17,
>>> + * give a chance to provide more information about vPMU
>>> + * enablement.
>>> + *
>>> + * The kvm.enable_pmu's permission is 0444. It does not change
>>> + * until a reload of the KVM module.
>>> + */
>>> + if (g_file_get_contents("/sys/module/kvm/parameters/enable_pmu",
>>> + &kvm_enable_pmu, NULL, NULL)) {
>>> + if (*kvm_enable_pmu == 'N') {
>>> + kvm_pmu_disabled = true;
>>
>> It’s generally better not to rely on KVM’s internal implementation
>> unless really necessary.
>>
>> For example, in the new mediated vPMU framework, even if the KVM module
>> parameter enable_pmu is set, the per-guest kvm->arch.enable_pmu could
>> still be cleared.
>>
>> In such a case, the logic here might not be correct.
>
> Would the Mediated vPMU set KVM_PMU_CAP_DISABLE to clear per-VM enable_pmu even
> when the global KVM parameter enable_pmu=N is set?
>
> In this scenario, we plan to rely on KVM_PMU_CAP_DISABLE only when the value of
> "/sys/module/kvm/parameters/enable_pmu" is not equal to N.
>
> Can I assume that this will work with Mediated vPMU?
>
>
> Is there any possibility to follow the current approach before Mediated vPMU is
> finalized for mainline, and later introduce an incremental change using
> KVM_GET_MSRS probing? The current approach is straightforward and can work with
> existing Linux kernel source code.
Apologies for the incorrect statement I made earlier regarding mediated
vPMU.
According to the mediated vPMU v6, the only behavior specific to
mediated vPMU is that kvm->arch.enable_pmu may be cleared when
irqchip_in_kernel() is not true:
https://lore.kernel.org/all/20251206001720.468579-17-seanjc@google.com/
However, this does not imply that mediated vPMU requires any special
handling here. In theory, KVM could clear kvm->arch.enable_pmu in the
future for other reasons.
> For quite some time, QEMU has lacked support for disabling or resetting AMD PMU
> registers. If we could add this feature before Mediated vPMU is finalized, it
> would benefit many existing kernel versions. This patchset solves production bugs.
>
>
> Feel free to let me know your thought, while I would starting working on next
> version now.
>
> Thank you very much!
>
> Dongli Zhang
>
>
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v8 4/7] target/i386/kvm: query kvm.enable_pmu parameter
2026-01-06 21:03 ` Chen, Zide
@ 2026-01-07 8:05 ` Dongli Zhang
2026-01-07 18:09 ` Chen, Zide
0 siblings, 1 reply; 25+ messages in thread
From: Dongli Zhang @ 2026-01-07 8:05 UTC (permalink / raw)
To: Chen, Zide, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
Hi Zide,
On 1/6/26 1:03 PM, Chen, Zide wrote:
>
>
> On 1/5/2026 12:21 PM, Dongli Zhang wrote:
>> Hi Zide,
>>
>> On 1/2/26 2:59 PM, Chen, Zide wrote:
>>>
>>>
>>> On 12/29/2025 11:42 PM, Dongli Zhang wrote:
>>
>> [snip]
>>
>>>>
>>>> static struct kvm_cpuid2 *cpuid_cache;
>>>> static struct kvm_cpuid2 *hv_cpuid_cache;
>>>> @@ -2068,23 +2072,30 @@ int kvm_arch_pre_create_vcpu(CPUState *cpu, Error **errp)
>>>> if (first) {
>>>> first = false;
>>>>
>>>> - /*
>>>> - * Since Linux v5.18, KVM provides a VM-level capability to easily
>>>> - * disable PMUs; however, QEMU has been providing PMU property per
>>>> - * CPU since v1.6. In order to accommodate both, have to configure
>>>> - * the VM-level capability here.
>>>> - *
>>>> - * KVM_PMU_CAP_DISABLE doesn't change the PMU
>>>> - * behavior on Intel platform because current "pmu" property works
>>>> - * as expected.
>>>> - */
>>>> - if ((pmu_cap & KVM_PMU_CAP_DISABLE) && !X86_CPU(cpu)->enable_pmu) {
>>>> - ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
>>>> - KVM_PMU_CAP_DISABLE);
>>>> - if (ret < 0) {
>>>> - error_setg_errno(errp, -ret,
>>>> - "Failed to set KVM_PMU_CAP_DISABLE");
>>>> - return ret;
>>>> + if (X86_CPU(cpu)->enable_pmu) {
>>>> + if (kvm_pmu_disabled) {
>>>> + warn_report("Failed to enable PMU since "
>>>> + "KVM's enable_pmu parameter is disabled");
>>>
>>> I'm wondering about the intended value of this patch?
>>>
>>> If enable_pmu is true in QEMU but the corresponding KVM parameter is
>>> false, then KVM_GET_SUPPORTED_CPUID or KVM_GET_MSRS should be able to
>>> tell that the PMU feature is not supported by host.
>>>
>>> The logic implemented in this patch seems somewhat redundant.
>>
>> For Intel, the QEMU userspace can determine if the vPMU is disabled by KVM
>> through the use of KVM_GET_SUPPORTED_CPUID.
>>
>> However, this approach does not apply to AMD. Unlike Intel, AMD does not rely on
>> CPUID to detect whether PMU is supported. By default, we can assume that PMU is
>> always available, except for the recent PerfMonV2 feature.
>>
>> The main objective of this PATCH 4/7 is to introduce the variable
>> 'kvm_pmu_disabled', which will be reused in PATCH 5/7 to skip any PMU
>> initialization if the parameter is set to 'N'.
>>
>> +static void kvm_init_pmu_info(struct kvm_cpuid2 *cpuid, X86CPU *cpu)
>> +{
>> + CPUX86State *env = &cpu->env;
>> +
>> + /*
>> + * The PMU virtualization is disabled by kvm.enable_pmu=N.
>> + */
>> + if (kvm_pmu_disabled) {
>> + return;
>> + }
>
> Thanks for explanation.
>
>> The 'kvm_pmu_disabled' variable is used to differentiate between the following
>> two scenarios on AMD:
>>
>> (1) A newer KVM with KVM_PMU_CAP_DISABLE support, but explicitly disabled via
>> the KVM parameter ('N').
>>
>> (2) An older KVM without KVM_CAP_PMU_CAPABILITY support.
>>
>> In both cases, the call to KVM_CAP_PMU_CAPABILITY extension support check may
>> return 0.
>>
>> By reading the file "/sys/module/kvm/parameters/enable_pmu", we can distinguish
>> between these two scenarios.
>
> As described in PATCH 1/7, without issuing KVM_PMU_CAP_DISABLE, KVM has
> no way to know that userspace does not intend to enable vPMU in AMD
> platforms, and therefore does not fault guest accesses to PMU MSRs.
>
> My understanding is that the issue being addressed here is basically the
> opposite: QEMU does not know that vPMU is disabled by KVM.
Exactly.
Otherwise, QEMU issues unwanted MSR writes for every vCPU during QEMU reset.
>
> IIUC, one difference between Intel and AMD is that AMD lacks a CPUID
> leaf to indicate the availability of PMU version 1. But Intel
> potentially could be in the same situation that KVM advertises PMU
> availability but it's not actually supported. (e.g. kvm->arch.enable_pmu
> is false while modules parameter enable_pmu is true).
>
> From the guest’s point of view, it probes PMU MSRs to determine whether
> PMU support is present and it's fine in this situation.
>
> In userspace, QEMU may issue KVM_SET_MSRS / KVM_GET_MSRS to KVM without
> knowing that vPMU has been disabled by KVM. I think these IOCTLs should
> not fail, since KVM states that “Userspace is allowed to read MSRs, and
> write ‘0’ to MSRs, that KVM advertises to userspace, even if an MSR
> isn’t fully supported.”
>
> My current understanding is that AMD should be fine even without
> kvm_pmu_disabled, but I may be missing some context here.
>
> The bottom line is this patch doesn't handle the cases that KVM still
> could disable vPMU support even if enable_pmu is true.
Yes. There are still unwanted PMU MSR writes from QEMU. This just seems odd.
The concern with unwanted MSR writes was initially raised by Maksim Davydov:
https://lore.kernel.org/qemu-devel/a7f9c3c9-09af-4941-b137-2cb83ef8ceb3@yandex-team.ru/
As shown below on the v6.0 KVM hypervisor (AMD), while there are no errors from
QEMU, numerous annoying warnings are generated. (If I recall correctly, this can
also be triggered from the VM itself.)
However, here the logs are not only due to vcpu0, but indeed every vcpu.
[ 280.802976] kvm_set_msr_common: 1910 callbacks suppressed
[ 280.802981] kvm [18411]: vcpu0, guest rIP: 0xffffffffa4c97844 disabled
perfctr wrmsr: 0xc0010007 data 0xffff
[ 295.345747] kvm [18411]: vcpu0, guest rIP: 0xfff0 disabled perfctr wrmsr:
0xc0010004 data 0x0
[ 295.355379] kvm [18411]: vcpu0, guest rIP: 0xfff0 disabled perfctr wrmsr:
0xc0010005 data 0x0
[ 295.364997] kvm [18411]: vcpu0, guest rIP: 0xfff0 disabled perfctr wrmsr:
0xc0010006 data 0x0
[ 295.374618] kvm [18411]: vcpu0, guest rIP: 0xfff0 disabled perfctr wrmsr:
0xc0010007 data 0x0
[ 295.385048] kvm [18411]: vcpu1, guest rIP: 0xfff0 disabled perfctr wrmsr:
0xc0010004 data 0x0
[ 295.394694] kvm [18411]: vcpu1, guest rIP: 0xfff0 disabled perfctr wrmsr:
0xc0010005 data 0x0
[ 295.404317] kvm [18411]: vcpu1, guest rIP: 0xfff0 disabled perfctr wrmsr:
0xc0010006 data 0x0
[ 295.413928] kvm [18411]: vcpu1, guest rIP: 0xfff0 disabled perfctr wrmsr:
0xc0010007 data 0x0
[ 295.424319] kvm [18411]: vcpu2, guest rIP: 0xfff0 disabled perfctr wrmsr:
0xc0010004 data 0x0
[ 295.433963] kvm [18411]: vcpu2, guest rIP: 0xfff0 disabled perfctr wrmsr:
0xc0010005 data 0x0
[ 301.966571] kvm_set_msr_common: 1910 callbacks suppressed
[ 301.966577] kvm [18411]: vcpu0, guest rIP: 0xffffffff8ac97844 disabled
perfctr wrmsr: 0xc0010007 data 0xffff
>
>
>> As you mentioned, another approach would be to use KVM_GET_MSRS to specifically
>> probe for AMD during QEMU initialization. In this case, we can set
>> 'kvm_pmu_disabled' to true if reading the AMD PMU MSR registers fails.
>>
>> To implement this, we may need to:
>>
>> 1. Turn this patch to be AMD specific by probing the AMD PMU registers during
>> initialization. We may need go create a new function in QEMU to use KVM_GET_MSRS
>> for probing only, or we may re-use kvm_arch_get_supported_msr_feature() or
>> kvm_get_one_msr(). I may change in the next version.
>>
>> 2. Limit the usage of 'kvm_pmu_disabled' to be AMD specific in PATCH 5/7.
>
> I guess this might make things more complicated.
>
>>>
>>> Additionally, in this scenario — where the user intends to enable a
>>> feature but the host cannot support it — normally no warning is emitted
>>> by QEMU.
>>
>> According to the usage of QEMU, may I assume QEMU already prints warning logs
>> for unsupported features? The below is an example.
>>
>> QEMU 10.2.50 monitor - type 'help' for more information
>> qemu-system-x86_64: warning: host doesn't support requested feature:
>> CPUID[eax=07h,ecx=00h].EBX.hle [bit 4]
>> qemu-system-x86_64: warning: host doesn't support requested feature:
>> CPUID[eax=07h,ecx=00h].EBX.rtm [bit 11]
>>
>>>
>>>> + }
>>>> + } else {
>>>> + /*
>>>> + * Since Linux v5.18, KVM provides a VM-level capability to easily
>>>> + * disable PMUs; however, QEMU has been providing PMU property per
>>>> + * CPU since v1.6. In order to accommodate both, have to configure
>>>> + * the VM-level capability here.
>>>> + *
>>>> + * KVM_PMU_CAP_DISABLE doesn't change the PMU
>>>> + * behavior on Intel platform because current "pmu" property works
>>>> + * as expected.
>>>> + */
>>>> + if (pmu_cap & KVM_PMU_CAP_DISABLE) {
>>>> + ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
>>>> + KVM_PMU_CAP_DISABLE);
>>>> + if (ret < 0) {
>>>> + error_setg_errno(errp, -ret,
>>>> + "Failed to set KVM_PMU_CAP_DISABLE");
>>>> + return ret;
>>>> + }
>>>> }
>>>> }
>>>> }
>>>> @@ -3302,6 +3313,7 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
>>>> int ret;
>>>> struct utsname utsname;
>>>> Error *local_err = NULL;
>>>> + g_autofree char *kvm_enable_pmu;
>>>>
>>>> /*
>>>> * Initialize confidential guest (SEV/TDX) context, if required
>>>> @@ -3437,6 +3449,21 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
>>>>
>>>> pmu_cap = kvm_check_extension(s, KVM_CAP_PMU_CAPABILITY);
>>>>
>>>> + /*
>>>> + * The enable_pmu parameter is introduced since Linux v5.17,
>>>> + * give a chance to provide more information about vPMU
>>>> + * enablement.
>>>> + *
>>>> + * The kvm.enable_pmu's permission is 0444. It does not change
>>>> + * until a reload of the KVM module.
>>>> + */
>>>> + if (g_file_get_contents("/sys/module/kvm/parameters/enable_pmu",
>>>> + &kvm_enable_pmu, NULL, NULL)) {
>>>> + if (*kvm_enable_pmu == 'N') {
>>>> + kvm_pmu_disabled = true;
>>>
>>> It’s generally better not to rely on KVM’s internal implementation
>>> unless really necessary.
>>>
>>> For example, in the new mediated vPMU framework, even if the KVM module
>>> parameter enable_pmu is set, the per-guest kvm->arch.enable_pmu could
>>> still be cleared.
>>>
>>> In such a case, the logic here might not be correct.
>>
>> Would the Mediated vPMU set KVM_PMU_CAP_DISABLE to clear per-VM enable_pmu even
>> when the global KVM parameter enable_pmu=N is set?
>>
>> In this scenario, we plan to rely on KVM_PMU_CAP_DISABLE only when the value of
>> "/sys/module/kvm/parameters/enable_pmu" is not equal to N.
>>
>> Can I assume that this will work with Mediated vPMU?
>>
>>
>> Is there any possibility to follow the current approach before Mediated vPMU is
>> finalized for mainline, and later introduce an incremental change using
>> KVM_GET_MSRS probing? The current approach is straightforward and can work with
>> existing Linux kernel source code.
>
> Apologies for the incorrect statement I made earlier regarding mediated
> vPMU.
>
> According to the mediated vPMU v6, the only behavior specific to
> mediated vPMU is that kvm->arch.enable_pmu may be cleared when
> irqchip_in_kernel() is not true:
> https://urldefense.com/v3/__https://lore.kernel.org/all/20251206001720.468579-17-seanjc@google.com/__;!!ACWV5N9M2RV99hQ!IS8XQG3Zx84utP2QScNlxp-0H5JAgr89lBb1j2oGVJJop3WMyK6X2I5mlerMPA06wkJy8VFd1x6XGEHc6kWn$
>
> However, this does not imply that mediated vPMU requires any special
> handling here. In theory, KVM could clear kvm->arch.enable_pmu in the
> future for other reasons.
>
While "KVM could clear kvm->arch.enable_pmu in the future," I don't think KVM
may set kvm->arch.enable_pmu if the global enable_pmu is set to 'N'.
Taking Intel VMX EPT as an example, once "/sys/module/kvm_intel/parameters/ept"
is globally disabled, there's no way within KVM software to enable it for any
guest VM **after** 'ept' is set to N.
Similarly, "/sys/module/kvm/parameters/enable_pmu=N" indicates that this KVM
host will not support PMU virtualization in any way. Therefore, there should be
no way to enable vPMU for any guest VM if the global parameter is set to 'N'.
Here we read from ths parameter only during QEMU initialization.
That's why I believe it's reliable to trust the setting when
"/sys/module/kvm/parameters/enable_pmu=N".
In this way, we can avoid many unnecessary MSR writes, especially in cases where
a VM has 300+ vCPUs, even though these may be equivalent to NOPs with
optimizations in more recent KVM versions.
The objective isn't to improve performance. Minimizing the number of unwanted
MSR writes from QEMU reduces the chances of failure (e.g., due to any QEMU
software bug). We can simply avoid those unwanted MSR/NOPs by reading from a KVM
parameter.
From a user's perspective, this just seems odd.
"/sys/module/kvm/parameters/enable_pmu=N" is a reliable setting. If there's a
configuration mismatch between QEMU and KVM, a warning could alert the user.
I can remove this patch, along with the 'kvm_pmu_disabled' variable.
Thank you very much!
Dongli Zhang
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v8 4/7] target/i386/kvm: query kvm.enable_pmu parameter
2026-01-07 8:05 ` Dongli Zhang
@ 2026-01-07 18:09 ` Chen, Zide
0 siblings, 0 replies; 25+ messages in thread
From: Chen, Zide @ 2026-01-07 18:09 UTC (permalink / raw)
To: Dongli Zhang, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
On 1/7/2026 12:05 AM, Dongli Zhang wrote:
> Hi Zide,
>
> On 1/6/26 1:03 PM, Chen, Zide wrote:
>>
>>
>> On 1/5/2026 12:21 PM, Dongli Zhang wrote:
>>> Hi Zide,
>>>
>>> On 1/2/26 2:59 PM, Chen, Zide wrote:
>>>>
>>>>
>>>> On 12/29/2025 11:42 PM, Dongli Zhang wrote:
>>>
>>> [snip]
>>>
>>>>>
>>>>> static struct kvm_cpuid2 *cpuid_cache;
>>>>> static struct kvm_cpuid2 *hv_cpuid_cache;
>>>>> @@ -2068,23 +2072,30 @@ int kvm_arch_pre_create_vcpu(CPUState *cpu, Error **errp)
>>>>> if (first) {
>>>>> first = false;
>>>>>
>>>>> - /*
>>>>> - * Since Linux v5.18, KVM provides a VM-level capability to easily
>>>>> - * disable PMUs; however, QEMU has been providing PMU property per
>>>>> - * CPU since v1.6. In order to accommodate both, have to configure
>>>>> - * the VM-level capability here.
>>>>> - *
>>>>> - * KVM_PMU_CAP_DISABLE doesn't change the PMU
>>>>> - * behavior on Intel platform because current "pmu" property works
>>>>> - * as expected.
>>>>> - */
>>>>> - if ((pmu_cap & KVM_PMU_CAP_DISABLE) && !X86_CPU(cpu)->enable_pmu) {
>>>>> - ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
>>>>> - KVM_PMU_CAP_DISABLE);
>>>>> - if (ret < 0) {
>>>>> - error_setg_errno(errp, -ret,
>>>>> - "Failed to set KVM_PMU_CAP_DISABLE");
>>>>> - return ret;
>>>>> + if (X86_CPU(cpu)->enable_pmu) {
>>>>> + if (kvm_pmu_disabled) {
>>>>> + warn_report("Failed to enable PMU since "
>>>>> + "KVM's enable_pmu parameter is disabled");
>>>>
>>>> I'm wondering about the intended value of this patch?
>>>>
>>>> If enable_pmu is true in QEMU but the corresponding KVM parameter is
>>>> false, then KVM_GET_SUPPORTED_CPUID or KVM_GET_MSRS should be able to
>>>> tell that the PMU feature is not supported by host.
>>>>
>>>> The logic implemented in this patch seems somewhat redundant.
>>>
>>> For Intel, the QEMU userspace can determine if the vPMU is disabled by KVM
>>> through the use of KVM_GET_SUPPORTED_CPUID.
>>>
>>> However, this approach does not apply to AMD. Unlike Intel, AMD does not rely on
>>> CPUID to detect whether PMU is supported. By default, we can assume that PMU is
>>> always available, except for the recent PerfMonV2 feature.
>>>
>>> The main objective of this PATCH 4/7 is to introduce the variable
>>> 'kvm_pmu_disabled', which will be reused in PATCH 5/7 to skip any PMU
>>> initialization if the parameter is set to 'N'.
>>>
>>> +static void kvm_init_pmu_info(struct kvm_cpuid2 *cpuid, X86CPU *cpu)
>>> +{
>>> + CPUX86State *env = &cpu->env;
>>> +
>>> + /*
>>> + * The PMU virtualization is disabled by kvm.enable_pmu=N.
>>> + */
>>> + if (kvm_pmu_disabled) {
>>> + return;
>>> + }
>>
>> Thanks for explanation.
>>
>>> The 'kvm_pmu_disabled' variable is used to differentiate between the following
>>> two scenarios on AMD:
>>>
>>> (1) A newer KVM with KVM_PMU_CAP_DISABLE support, but explicitly disabled via
>>> the KVM parameter ('N').
>>>
>>> (2) An older KVM without KVM_CAP_PMU_CAPABILITY support.
>>>
>>> In both cases, the call to KVM_CAP_PMU_CAPABILITY extension support check may
>>> return 0.
>>>
>>> By reading the file "/sys/module/kvm/parameters/enable_pmu", we can distinguish
>>> between these two scenarios.
>>
>> As described in PATCH 1/7, without issuing KVM_PMU_CAP_DISABLE, KVM has
>> no way to know that userspace does not intend to enable vPMU in AMD
>> platforms, and therefore does not fault guest accesses to PMU MSRs.
>>
>> My understanding is that the issue being addressed here is basically the
>> opposite: QEMU does not know that vPMU is disabled by KVM.
>
> Exactly.
>
> Otherwise, QEMU issues unwanted MSR writes for every vCPU during QEMU reset.
>
>>
>> IIUC, one difference between Intel and AMD is that AMD lacks a CPUID
>> leaf to indicate the availability of PMU version 1. But Intel
>> potentially could be in the same situation that KVM advertises PMU
>> availability but it's not actually supported. (e.g. kvm->arch.enable_pmu
>> is false while modules parameter enable_pmu is true).
>>
>> From the guest’s point of view, it probes PMU MSRs to determine whether
>> PMU support is present and it's fine in this situation.
>>
>> In userspace, QEMU may issue KVM_SET_MSRS / KVM_GET_MSRS to KVM without
>> knowing that vPMU has been disabled by KVM. I think these IOCTLs should
>> not fail, since KVM states that “Userspace is allowed to read MSRs, and
>> write ‘0’ to MSRs, that KVM advertises to userspace, even if an MSR
>> isn’t fully supported.”
>>
>> My current understanding is that AMD should be fine even without
>> kvm_pmu_disabled, but I may be missing some context here.
>>
>> The bottom line is this patch doesn't handle the cases that KVM still
>> could disable vPMU support even if enable_pmu is true.
>
> Yes. There are still unwanted PMU MSR writes from QEMU. This just seems odd.
>
> The concern with unwanted MSR writes was initially raised by Maksim Davydov:
>
> https://lore.kernel.org/qemu-devel/a7f9c3c9-09af-4941-b137-2cb83ef8ceb3@yandex-team.ru/
>
> As shown below on the v6.0 KVM hypervisor (AMD), while there are no errors from
> QEMU, numerous annoying warnings are generated. (If I recall correctly, this can
> also be triggered from the VM itself.)
>
> However, here the logs are not only due to vcpu0, but indeed every vcpu.
>
> [ 280.802976] kvm_set_msr_common: 1910 callbacks suppressed
> [ 280.802981] kvm [18411]: vcpu0, guest rIP: 0xffffffffa4c97844 disabled
> perfctr wrmsr: 0xc0010007 data 0xffff
> [ 295.345747] kvm [18411]: vcpu0, guest rIP: 0xfff0 disabled perfctr wrmsr:
> 0xc0010004 data 0x0
> [ 295.355379] kvm [18411]: vcpu0, guest rIP: 0xfff0 disabled perfctr wrmsr:
> 0xc0010005 data 0x0
> [ 295.364997] kvm [18411]: vcpu0, guest rIP: 0xfff0 disabled perfctr wrmsr:
> 0xc0010006 data 0x0
> [ 295.374618] kvm [18411]: vcpu0, guest rIP: 0xfff0 disabled perfctr wrmsr:
> 0xc0010007 data 0x0
> [ 295.385048] kvm [18411]: vcpu1, guest rIP: 0xfff0 disabled perfctr wrmsr:
> 0xc0010004 data 0x0
> [ 295.394694] kvm [18411]: vcpu1, guest rIP: 0xfff0 disabled perfctr wrmsr:
> 0xc0010005 data 0x0
> [ 295.404317] kvm [18411]: vcpu1, guest rIP: 0xfff0 disabled perfctr wrmsr:
> 0xc0010006 data 0x0
> [ 295.413928] kvm [18411]: vcpu1, guest rIP: 0xfff0 disabled perfctr wrmsr:
> 0xc0010007 data 0x0
> [ 295.424319] kvm [18411]: vcpu2, guest rIP: 0xfff0 disabled perfctr wrmsr:
> 0xc0010004 data 0x0
> [ 295.433963] kvm [18411]: vcpu2, guest rIP: 0xfff0 disabled perfctr wrmsr:
> 0xc0010005 data 0x0
> [ 301.966571] kvm_set_msr_common: 1910 callbacks suppressed
> [ 301.966577] kvm [18411]: vcpu0, guest rIP: 0xffffffff8ac97844 disabled
> perfctr wrmsr: 0xc0010007 data 0xffff
In e76ae52747a8 ("KVM: x86/pmu: Gate all "unimplemented MSR" prints on
report_ignored_msrs"), in "disabled perfctr wrmsr" case, vcpu_unimpl()
is no longer forced for counter MSRs, so most of the above warnings go away.
For the remaining warnings, vcpu_unimpl() is ratelimitedm and plus all
these logs can be removed by setting report_ignored_msrs=false.
So, it should be not that bad now.
>>
>>
>>> As you mentioned, another approach would be to use KVM_GET_MSRS to specifically
>>> probe for AMD during QEMU initialization. In this case, we can set
>>> 'kvm_pmu_disabled' to true if reading the AMD PMU MSR registers fails.
>>>
>>> To implement this, we may need to:
>>>
>>> 1. Turn this patch to be AMD specific by probing the AMD PMU registers during
>>> initialization. We may need go create a new function in QEMU to use KVM_GET_MSRS
>>> for probing only, or we may re-use kvm_arch_get_supported_msr_feature() or
>>> kvm_get_one_msr(). I may change in the next version.
>>>
>>> 2. Limit the usage of 'kvm_pmu_disabled' to be AMD specific in PATCH 5/7.
>>
>> I guess this might make things more complicated.
>>
>>>>
>>>> Additionally, in this scenario — where the user intends to enable a
>>>> feature but the host cannot support it — normally no warning is emitted
>>>> by QEMU.
>>>
>>> According to the usage of QEMU, may I assume QEMU already prints warning logs
>>> for unsupported features? The below is an example.
>>>
>>> QEMU 10.2.50 monitor - type 'help' for more information
>>> qemu-system-x86_64: warning: host doesn't support requested feature:
>>> CPUID[eax=07h,ecx=00h].EBX.hle [bit 4]
>>> qemu-system-x86_64: warning: host doesn't support requested feature:
>>> CPUID[eax=07h,ecx=00h].EBX.rtm [bit 11]
>>>
>>>>
>>>>> + }
>>>>> + } else {
>>>>> + /*
>>>>> + * Since Linux v5.18, KVM provides a VM-level capability to easily
>>>>> + * disable PMUs; however, QEMU has been providing PMU property per
>>>>> + * CPU since v1.6. In order to accommodate both, have to configure
>>>>> + * the VM-level capability here.
>>>>> + *
>>>>> + * KVM_PMU_CAP_DISABLE doesn't change the PMU
>>>>> + * behavior on Intel platform because current "pmu" property works
>>>>> + * as expected.
>>>>> + */
>>>>> + if (pmu_cap & KVM_PMU_CAP_DISABLE) {
>>>>> + ret = kvm_vm_enable_cap(kvm_state, KVM_CAP_PMU_CAPABILITY, 0,
>>>>> + KVM_PMU_CAP_DISABLE);
>>>>> + if (ret < 0) {
>>>>> + error_setg_errno(errp, -ret,
>>>>> + "Failed to set KVM_PMU_CAP_DISABLE");
>>>>> + return ret;
>>>>> + }
>>>>> }
>>>>> }
>>>>> }
>>>>> @@ -3302,6 +3313,7 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
>>>>> int ret;
>>>>> struct utsname utsname;
>>>>> Error *local_err = NULL;
>>>>> + g_autofree char *kvm_enable_pmu;
>>>>>
>>>>> /*
>>>>> * Initialize confidential guest (SEV/TDX) context, if required
>>>>> @@ -3437,6 +3449,21 @@ int kvm_arch_init(MachineState *ms, KVMState *s)
>>>>>
>>>>> pmu_cap = kvm_check_extension(s, KVM_CAP_PMU_CAPABILITY);
>>>>>
>>>>> + /*
>>>>> + * The enable_pmu parameter is introduced since Linux v5.17,
>>>>> + * give a chance to provide more information about vPMU
>>>>> + * enablement.
>>>>> + *
>>>>> + * The kvm.enable_pmu's permission is 0444. It does not change
>>>>> + * until a reload of the KVM module.
>>>>> + */
>>>>> + if (g_file_get_contents("/sys/module/kvm/parameters/enable_pmu",
>>>>> + &kvm_enable_pmu, NULL, NULL)) {
>>>>> + if (*kvm_enable_pmu == 'N') {
>>>>> + kvm_pmu_disabled = true;
>>>>
>>>> It’s generally better not to rely on KVM’s internal implementation
>>>> unless really necessary.
>>>>
>>>> For example, in the new mediated vPMU framework, even if the KVM module
>>>> parameter enable_pmu is set, the per-guest kvm->arch.enable_pmu could
>>>> still be cleared.
>>>>
>>>> In such a case, the logic here might not be correct.
>>>
>>> Would the Mediated vPMU set KVM_PMU_CAP_DISABLE to clear per-VM enable_pmu even
>>> when the global KVM parameter enable_pmu=N is set?
>>>
>>> In this scenario, we plan to rely on KVM_PMU_CAP_DISABLE only when the value of
>>> "/sys/module/kvm/parameters/enable_pmu" is not equal to N.
>>>
>>> Can I assume that this will work with Mediated vPMU?
>>>
>>>
>>> Is there any possibility to follow the current approach before Mediated vPMU is
>>> finalized for mainline, and later introduce an incremental change using
>>> KVM_GET_MSRS probing? The current approach is straightforward and can work with
>>> existing Linux kernel source code.
>>
>> Apologies for the incorrect statement I made earlier regarding mediated
>> vPMU.
>>
>> According to the mediated vPMU v6, the only behavior specific to
>> mediated vPMU is that kvm->arch.enable_pmu may be cleared when
>> irqchip_in_kernel() is not true:
>> https://urldefense.com/v3/__https://lore.kernel.org/all/20251206001720.468579-17-seanjc@google.com/__;!!ACWV5N9M2RV99hQ!IS8XQG3Zx84utP2QScNlxp-0H5JAgr89lBb1j2oGVJJop3WMyK6X2I5mlerMPA06wkJy8VFd1x6XGEHc6kWn$
>>
>> However, this does not imply that mediated vPMU requires any special
>> handling here. In theory, KVM could clear kvm->arch.enable_pmu in the
>> future for other reasons.
>>
>
> While "KVM could clear kvm->arch.enable_pmu in the future," I don't think KVM
> may set kvm->arch.enable_pmu if the global enable_pmu is set to 'N'.
>
> Taking Intel VMX EPT as an example, once "/sys/module/kvm_intel/parameters/ept"
> is globally disabled, there's no way within KVM software to enable it for any
> guest VM **after** 'ept' is set to N.
>
> Similarly, "/sys/module/kvm/parameters/enable_pmu=N" indicates that this KVM
> host will not support PMU virtualization in any way. Therefore, there should be
> no way to enable vPMU for any guest VM if the global parameter is set to 'N'.
> Here we read from ths parameter only during QEMU initialization.
>
> That's why I believe it's reliable to trust the setting when
> "/sys/module/kvm/parameters/enable_pmu=N".
>
> In this way, we can avoid many unnecessary MSR writes, especially in cases where
> a VM has 300+ vCPUs, even though these may be equivalent to NOPs with
> optimizations in more recent KVM versions.
>
> The objective isn't to improve performance. Minimizing the number of unwanted
> MSR writes from QEMU reduces the chances of failure (e.g., due to any QEMU
> software bug). We can simply avoid those unwanted MSR/NOPs by reading from a KVM
> parameter.
>
> From a user's perspective, this just seems odd.
> "/sys/module/kvm/parameters/enable_pmu=N" is a reliable setting. If there's a
> configuration mismatch between QEMU and KVM, a warning could alert the user.
>
> I can remove this patch, along with the 'kvm_pmu_disabled' variable.
Even with /sys/module/kvm/parameters/enable_pmu=Y, theoretically it's
possible for kvm->arch.enable_pmu to be false. In such a case, vPMU
could still be advertised, and QEMU doens't know that vPMU is not
supported by KVM, on either Intel or AMD platforms.
Anyway, this is likely only a theoretical scenario and may not actually
happens in practice.
> Thank you very much!
>
> Dongli Zhang
^ permalink raw reply [flat|nested] 25+ messages in thread
* [PATCH v8 5/7] target/i386/kvm: reset AMD PMU registers during VM reset
2025-12-30 7:42 [PATCH v8 0/7] target/i386/kvm/pmu: PMU Enhancement, Bugfix and Cleanup Dongli Zhang
` (3 preceding siblings ...)
2025-12-30 7:42 ` [PATCH v8 4/7] target/i386/kvm: query kvm.enable_pmu parameter Dongli Zhang
@ 2025-12-30 7:42 ` Dongli Zhang
2025-12-30 7:42 ` [PATCH v8 6/7] target/i386/kvm: support perfmon-v2 for reset Dongli Zhang
2025-12-30 7:42 ` [PATCH v8 7/7] target/i386/kvm: don't stop Intel PMU counters Dongli Zhang
6 siblings, 0 replies; 25+ messages in thread
From: Dongli Zhang @ 2025-12-30 7:42 UTC (permalink / raw)
To: qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
QEMU uses the kvm_get_msrs() function to save Intel PMU registers from KVM
and kvm_put_msrs() to restore them to KVM. However, there is no support for
AMD PMU registers. Currently, pmu_version and num_pmu_gp_counters are
initialized based on cpuid(0xa), which does not apply to AMD processors.
For AMD CPUs, prior to PerfMonV2, the number of general-purpose registers
is determined based on the CPU version.
To address this issue, we need to add support for AMD PMU registers.
Without this support, the following problems can arise:
1. If the VM is reset (e.g., via QEMU system_reset or VM kdump/kexec) while
running "perf top", the PMU registers are not disabled properly.
2. Despite x86_cpu_reset() resetting many registers to zero, kvm_put_msrs()
does not handle AMD PMU registers, causing some PMU events to remain
enabled in KVM.
3. The KVM kvm_pmc_speculative_in_use() function consistently returns true,
preventing the reclamation of these events. Consequently, the
kvm_pmc->perf_event remains active.
4. After a reboot, the VM kernel may report the following error:
[ 0.092011] Performance Events: Fam17h+ core perfctr, Broken BIOS detected, complain to your hardware vendor.
[ 0.092023] [Firmware Bug]: the BIOS has corrupted hw-PMU resources (MSR c0010200 is 530076)
5. In the worst case, the active kvm_pmc->perf_event may inject unknown
NMIs randomly into the VM kernel:
[...] Uhhuh. NMI received for unknown reason 30 on CPU 0.
To resolve these issues, we propose resetting AMD PMU registers during the
VM reset process.
Signed-off-by: Dongli Zhang <dongli.zhang@oracle.com>
Reviewed-by: Zhao Liu <zhao1.liu@intel.com>
Reviewed-by: Sandipan Das <sandipan.das@amd.com>
Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
Changed since v1:
- Modify "MSR_K7_EVNTSEL0 + 3" and "MSR_K7_PERFCTR0 + 3" by using
AMD64_NUM_COUNTERS (suggested by Sandipan Das).
- Use "AMD64_NUM_COUNTERS_CORE * 2 - 1", not "MSR_F15H_PERF_CTL0 + 0xb".
(suggested by Sandipan Das).
- Switch back to "-pmu" instead of using a global "pmu-cap-disabled".
- Don't initialize PMU info if kvm.enable_pmu=N.
Changed since v2:
- Remove 'static' from host_cpuid_vendorX.
- Change has_pmu_version to pmu_version.
- Use object_property_get_int() to get CPU family.
- Use cpuid_find_entry() instead of cpu_x86_cpuid().
- Send error log when host and guest are from different vendors.
- Move "if (!cpu->enable_pmu)" to begin of function. Add comments to
reminder developers.
- Add support to Zhaoxin. Change is_same_vendor() to
is_host_compat_vendor().
- Didn't add Reviewed-by from Sandipan because the change isn't minor.
Changed since v3:
- Use host_cpu_vendor_fms() from Zhao's patch.
- Check AMD directly makes the "compat" rule clear.
- Add comment to MAX_GP_COUNTERS.
- Skip PMU info initialization if !kvm_pmu_disabled.
Changed since v4:
- Add Reviewed-by from Zhao and Sandipan.
Changed since v6:
- Add Reviewed-by from Dapeng Mi.
target/i386/cpu.h | 12 +++
target/i386/kvm/kvm.c | 175 +++++++++++++++++++++++++++++++++++++++++-
2 files changed, 183 insertions(+), 4 deletions(-)
diff --git a/target/i386/cpu.h b/target/i386/cpu.h
index 41ea04099b..c1649d1247 100644
--- a/target/i386/cpu.h
+++ b/target/i386/cpu.h
@@ -506,6 +506,14 @@ typedef enum X86Seg {
#define MSR_CORE_PERF_GLOBAL_CTRL 0x38f
#define MSR_CORE_PERF_GLOBAL_OVF_CTRL 0x390
+#define MSR_K7_EVNTSEL0 0xc0010000
+#define MSR_K7_PERFCTR0 0xc0010004
+#define MSR_F15H_PERF_CTL0 0xc0010200
+#define MSR_F15H_PERF_CTR0 0xc0010201
+
+#define AMD64_NUM_COUNTERS 4
+#define AMD64_NUM_COUNTERS_CORE 6
+
#define MSR_MC0_CTL 0x400
#define MSR_MC0_STATUS 0x401
#define MSR_MC0_ADDR 0x402
@@ -1737,6 +1745,10 @@ typedef struct {
#endif
#define MAX_FIXED_COUNTERS 3
+/*
+ * This formula is based on Intel's MSR. The current size also meets AMD's
+ * needs.
+ */
#define MAX_GP_COUNTERS (MSR_IA32_PERF_STATUS - MSR_P6_EVNTSEL0)
#define NB_OPMASK_REGS 8
diff --git a/target/i386/kvm/kvm.c b/target/i386/kvm/kvm.c
index 338b9558e4..b8e9c9192b 100644
--- a/target/i386/kvm/kvm.c
+++ b/target/i386/kvm/kvm.c
@@ -2107,7 +2107,7 @@ int kvm_arch_pre_create_vcpu(CPUState *cpu, Error **errp)
return 0;
}
-static void kvm_init_pmu_info(struct kvm_cpuid2 *cpuid)
+static void kvm_init_pmu_info_intel(struct kvm_cpuid2 *cpuid)
{
struct kvm_cpuid_entry2 *c;
@@ -2140,6 +2140,96 @@ static void kvm_init_pmu_info(struct kvm_cpuid2 *cpuid)
}
}
+static void kvm_init_pmu_info_amd(struct kvm_cpuid2 *cpuid, X86CPU *cpu)
+{
+ struct kvm_cpuid_entry2 *c;
+ int64_t family;
+
+ family = object_property_get_int(OBJECT(cpu), "family", NULL);
+ if (family < 0) {
+ return;
+ }
+
+ if (family < 6) {
+ error_report("AMD performance-monitoring is supported from "
+ "K7 and later");
+ return;
+ }
+
+ pmu_version = 1;
+ num_pmu_gp_counters = AMD64_NUM_COUNTERS;
+
+ c = cpuid_find_entry(cpuid, 0x80000001, 0);
+ if (!c) {
+ return;
+ }
+
+ if (!(c->ecx & CPUID_EXT3_PERFCORE)) {
+ return;
+ }
+
+ num_pmu_gp_counters = AMD64_NUM_COUNTERS_CORE;
+}
+
+static bool is_host_compat_vendor(CPUX86State *env)
+{
+ char host_vendor[CPUID_VENDOR_SZ + 1];
+
+ host_cpu_vendor_fms(host_vendor, NULL, NULL, NULL);
+
+ /*
+ * Intel and Zhaoxin are compatible.
+ */
+ if ((g_str_equal(host_vendor, CPUID_VENDOR_INTEL) ||
+ g_str_equal(host_vendor, CPUID_VENDOR_ZHAOXIN1) ||
+ g_str_equal(host_vendor, CPUID_VENDOR_ZHAOXIN2)) &&
+ (IS_INTEL_CPU(env) || IS_ZHAOXIN_CPU(env))) {
+ return true;
+ }
+
+ return g_str_equal(host_vendor, CPUID_VENDOR_AMD) &&
+ IS_AMD_CPU(env);
+}
+
+static void kvm_init_pmu_info(struct kvm_cpuid2 *cpuid, X86CPU *cpu)
+{
+ CPUX86State *env = &cpu->env;
+
+ /*
+ * The PMU virtualization is disabled by kvm.enable_pmu=N.
+ */
+ if (kvm_pmu_disabled) {
+ return;
+ }
+
+ /*
+ * If KVM_CAP_PMU_CAPABILITY is not supported, there is no way to
+ * disable the AMD PMU virtualization.
+ *
+ * Assume the user is aware of this when !cpu->enable_pmu. AMD PMU
+ * registers are not going to reset, even they are still available to
+ * guest VM.
+ */
+ if (!cpu->enable_pmu) {
+ return;
+ }
+
+ /*
+ * It is not supported to virtualize AMD PMU registers on Intel
+ * processors, nor to virtualize Intel PMU registers on AMD processors.
+ */
+ if (!is_host_compat_vendor(env)) {
+ error_report("host doesn't support requested feature: vPMU");
+ return;
+ }
+
+ if (IS_INTEL_CPU(env) || IS_ZHAOXIN_CPU(env)) {
+ kvm_init_pmu_info_intel(cpuid);
+ } else if (IS_AMD_CPU(env)) {
+ kvm_init_pmu_info_amd(cpuid, cpu);
+ }
+}
+
int kvm_arch_init_vcpu(CPUState *cs)
{
struct {
@@ -2330,7 +2420,7 @@ int kvm_arch_init_vcpu(CPUState *cs)
cpuid_i = kvm_x86_build_cpuid(env, cpuid_data.entries, cpuid_i);
cpuid_data.cpuid.nent = cpuid_i;
- kvm_init_pmu_info(&cpuid_data.cpuid);
+ kvm_init_pmu_info(&cpuid_data.cpuid, cpu);
if (x86_cpu_family(env->cpuid_version) >= 6
&& (env->features[FEAT_1_EDX] & (CPUID_MCE | CPUID_MCA)) ==
@@ -4121,7 +4211,7 @@ static int kvm_put_msrs(X86CPU *cpu, KvmPutState level)
kvm_msr_entry_add(cpu, MSR_KVM_POLL_CONTROL, env->poll_control_msr);
}
- if (pmu_version > 0) {
+ if ((IS_INTEL_CPU(env) || IS_ZHAOXIN_CPU(env)) && pmu_version > 0) {
if (pmu_version > 1) {
/* Stop the counter. */
kvm_msr_entry_add(cpu, MSR_CORE_PERF_FIXED_CTR_CTRL, 0);
@@ -4152,6 +4242,38 @@ static int kvm_put_msrs(X86CPU *cpu, KvmPutState level)
env->msr_global_ctrl);
}
}
+
+ if (IS_AMD_CPU(env) && pmu_version > 0) {
+ uint32_t sel_base = MSR_K7_EVNTSEL0;
+ uint32_t ctr_base = MSR_K7_PERFCTR0;
+ /*
+ * The address of the next selector or counter register is
+ * obtained by incrementing the address of the current selector
+ * or counter register by one.
+ */
+ uint32_t step = 1;
+
+ /*
+ * When PERFCORE is enabled, AMD PMU uses a separate set of
+ * addresses for the selector and counter registers.
+ * Additionally, the address of the next selector or counter
+ * register is determined by incrementing the address of the
+ * current register by two.
+ */
+ if (num_pmu_gp_counters == AMD64_NUM_COUNTERS_CORE) {
+ sel_base = MSR_F15H_PERF_CTL0;
+ ctr_base = MSR_F15H_PERF_CTR0;
+ step = 2;
+ }
+
+ for (i = 0; i < num_pmu_gp_counters; i++) {
+ kvm_msr_entry_add(cpu, ctr_base + i * step,
+ env->msr_gp_counters[i]);
+ kvm_msr_entry_add(cpu, sel_base + i * step,
+ env->msr_gp_evtsel[i]);
+ }
+ }
+
/*
* Hyper-V partition-wide MSRs: to avoid clearing them on cpu hot-add,
* only sync them to KVM on the first cpu
@@ -4656,7 +4778,8 @@ static int kvm_get_msrs(X86CPU *cpu)
if (env->features[FEAT_KVM] & CPUID_KVM_POLL_CONTROL) {
kvm_msr_entry_add(cpu, MSR_KVM_POLL_CONTROL, 1);
}
- if (pmu_version > 0) {
+
+ if ((IS_INTEL_CPU(env) || IS_ZHAOXIN_CPU(env)) && pmu_version > 0) {
if (pmu_version > 1) {
kvm_msr_entry_add(cpu, MSR_CORE_PERF_FIXED_CTR_CTRL, 0);
kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_CTRL, 0);
@@ -4672,6 +4795,35 @@ static int kvm_get_msrs(X86CPU *cpu)
}
}
+ if (IS_AMD_CPU(env) && pmu_version > 0) {
+ uint32_t sel_base = MSR_K7_EVNTSEL0;
+ uint32_t ctr_base = MSR_K7_PERFCTR0;
+ /*
+ * The address of the next selector or counter register is
+ * obtained by incrementing the address of the current selector
+ * or counter register by one.
+ */
+ uint32_t step = 1;
+
+ /*
+ * When PERFCORE is enabled, AMD PMU uses a separate set of
+ * addresses for the selector and counter registers.
+ * Additionally, the address of the next selector or counter
+ * register is determined by incrementing the address of the
+ * current register by two.
+ */
+ if (num_pmu_gp_counters == AMD64_NUM_COUNTERS_CORE) {
+ sel_base = MSR_F15H_PERF_CTL0;
+ ctr_base = MSR_F15H_PERF_CTR0;
+ step = 2;
+ }
+
+ for (i = 0; i < num_pmu_gp_counters; i++) {
+ kvm_msr_entry_add(cpu, ctr_base + i * step, 0);
+ kvm_msr_entry_add(cpu, sel_base + i * step, 0);
+ }
+ }
+
if (env->mcg_cap) {
kvm_msr_entry_add(cpu, MSR_MCG_STATUS, 0);
kvm_msr_entry_add(cpu, MSR_MCG_CTL, 0);
@@ -5002,6 +5154,21 @@ static int kvm_get_msrs(X86CPU *cpu)
case MSR_P6_EVNTSEL0 ... MSR_P6_EVNTSEL0 + MAX_GP_COUNTERS - 1:
env->msr_gp_evtsel[index - MSR_P6_EVNTSEL0] = msrs[i].data;
break;
+ case MSR_K7_EVNTSEL0 ... MSR_K7_EVNTSEL0 + AMD64_NUM_COUNTERS - 1:
+ env->msr_gp_evtsel[index - MSR_K7_EVNTSEL0] = msrs[i].data;
+ break;
+ case MSR_K7_PERFCTR0 ... MSR_K7_PERFCTR0 + AMD64_NUM_COUNTERS - 1:
+ env->msr_gp_counters[index - MSR_K7_PERFCTR0] = msrs[i].data;
+ break;
+ case MSR_F15H_PERF_CTL0 ...
+ MSR_F15H_PERF_CTL0 + AMD64_NUM_COUNTERS_CORE * 2 - 1:
+ index = index - MSR_F15H_PERF_CTL0;
+ if (index & 0x1) {
+ env->msr_gp_counters[index] = msrs[i].data;
+ } else {
+ env->msr_gp_evtsel[index] = msrs[i].data;
+ }
+ break;
case HV_X64_MSR_HYPERCALL:
env->msr_hv_hypercall = msrs[i].data;
break;
--
2.39.3
^ permalink raw reply related [flat|nested] 25+ messages in thread* [PATCH v8 6/7] target/i386/kvm: support perfmon-v2 for reset
2025-12-30 7:42 [PATCH v8 0/7] target/i386/kvm/pmu: PMU Enhancement, Bugfix and Cleanup Dongli Zhang
` (4 preceding siblings ...)
2025-12-30 7:42 ` [PATCH v8 5/7] target/i386/kvm: reset AMD PMU registers during VM reset Dongli Zhang
@ 2025-12-30 7:42 ` Dongli Zhang
2025-12-30 7:42 ` [PATCH v8 7/7] target/i386/kvm: don't stop Intel PMU counters Dongli Zhang
6 siblings, 0 replies; 25+ messages in thread
From: Dongli Zhang @ 2025-12-30 7:42 UTC (permalink / raw)
To: qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
Since perfmon-v2, the AMD PMU supports additional registers. This update
includes get/put functionality for these extra registers.
Similar to the implementation in KVM:
- MSR_CORE_PERF_GLOBAL_STATUS and MSR_AMD64_PERF_CNTR_GLOBAL_STATUS both
use env->msr_global_status.
- MSR_CORE_PERF_GLOBAL_CTRL and MSR_AMD64_PERF_CNTR_GLOBAL_CTL both use
env->msr_global_ctrl.
- MSR_CORE_PERF_GLOBAL_OVF_CTRL and MSR_AMD64_PERF_CNTR_GLOBAL_STATUS_CLR
both use env->msr_global_ovf_ctrl.
No changes are needed for vmstate_msr_architectural_pmu or
pmu_enable_needed().
Signed-off-by: Dongli Zhang <dongli.zhang@oracle.com>
Reviewed-by: Zhao Liu <zhao1.liu@intel.com>
Reviewed-by: Sandipan Das <sandipan.das@amd.com>
---
Changed since v1:
- Use "has_pmu_version > 1", not "has_pmu_version == 2".
Changed since v2:
- Use cpuid_find_entry() instead of cpu_x86_cpuid().
- Change has_pmu_version to pmu_version.
- Cap num_pmu_gp_counters with MAX_GP_COUNTERS.
Changed since v4:
- Add Reviewed-by from Sandipan.
target/i386/cpu.h | 4 ++++
target/i386/kvm/kvm.c | 48 +++++++++++++++++++++++++++++++++++--------
2 files changed, 43 insertions(+), 9 deletions(-)
diff --git a/target/i386/cpu.h b/target/i386/cpu.h
index c1649d1247..c6dab03a92 100644
--- a/target/i386/cpu.h
+++ b/target/i386/cpu.h
@@ -506,6 +506,10 @@ typedef enum X86Seg {
#define MSR_CORE_PERF_GLOBAL_CTRL 0x38f
#define MSR_CORE_PERF_GLOBAL_OVF_CTRL 0x390
+#define MSR_AMD64_PERF_CNTR_GLOBAL_STATUS 0xc0000300
+#define MSR_AMD64_PERF_CNTR_GLOBAL_CTL 0xc0000301
+#define MSR_AMD64_PERF_CNTR_GLOBAL_STATUS_CLR 0xc0000302
+
#define MSR_K7_EVNTSEL0 0xc0010000
#define MSR_K7_PERFCTR0 0xc0010004
#define MSR_F15H_PERF_CTL0 0xc0010200
diff --git a/target/i386/kvm/kvm.c b/target/i386/kvm/kvm.c
index b8e9c9192b..99837048b8 100644
--- a/target/i386/kvm/kvm.c
+++ b/target/i386/kvm/kvm.c
@@ -2169,6 +2169,16 @@ static void kvm_init_pmu_info_amd(struct kvm_cpuid2 *cpuid, X86CPU *cpu)
}
num_pmu_gp_counters = AMD64_NUM_COUNTERS_CORE;
+
+ c = cpuid_find_entry(cpuid, 0x80000022, 0);
+ if (c && (c->eax & CPUID_8000_0022_EAX_PERFMON_V2)) {
+ pmu_version = 2;
+ num_pmu_gp_counters = c->ebx & 0xf;
+
+ if (num_pmu_gp_counters > MAX_GP_COUNTERS) {
+ num_pmu_gp_counters = MAX_GP_COUNTERS;
+ }
+ }
}
static bool is_host_compat_vendor(CPUX86State *env)
@@ -4254,13 +4264,14 @@ static int kvm_put_msrs(X86CPU *cpu, KvmPutState level)
uint32_t step = 1;
/*
- * When PERFCORE is enabled, AMD PMU uses a separate set of
- * addresses for the selector and counter registers.
- * Additionally, the address of the next selector or counter
- * register is determined by incrementing the address of the
- * current register by two.
+ * When PERFCORE or PerfMonV2 is enabled, AMD PMU uses a
+ * separate set of addresses for the selector and counter
+ * registers. Additionally, the address of the next selector or
+ * counter register is determined by incrementing the address
+ * of the current register by two.
*/
- if (num_pmu_gp_counters == AMD64_NUM_COUNTERS_CORE) {
+ if (num_pmu_gp_counters == AMD64_NUM_COUNTERS_CORE ||
+ pmu_version > 1) {
sel_base = MSR_F15H_PERF_CTL0;
ctr_base = MSR_F15H_PERF_CTR0;
step = 2;
@@ -4272,6 +4283,15 @@ static int kvm_put_msrs(X86CPU *cpu, KvmPutState level)
kvm_msr_entry_add(cpu, sel_base + i * step,
env->msr_gp_evtsel[i]);
}
+
+ if (pmu_version > 1) {
+ kvm_msr_entry_add(cpu, MSR_AMD64_PERF_CNTR_GLOBAL_STATUS,
+ env->msr_global_status);
+ kvm_msr_entry_add(cpu, MSR_AMD64_PERF_CNTR_GLOBAL_STATUS_CLR,
+ env->msr_global_ovf_ctrl);
+ kvm_msr_entry_add(cpu, MSR_AMD64_PERF_CNTR_GLOBAL_CTL,
+ env->msr_global_ctrl);
+ }
}
/*
@@ -4806,13 +4826,14 @@ static int kvm_get_msrs(X86CPU *cpu)
uint32_t step = 1;
/*
- * When PERFCORE is enabled, AMD PMU uses a separate set of
- * addresses for the selector and counter registers.
+ * When PERFCORE or PerfMonV2 is enabled, AMD PMU uses a separate
+ * set of addresses for the selector and counter registers.
* Additionally, the address of the next selector or counter
* register is determined by incrementing the address of the
* current register by two.
*/
- if (num_pmu_gp_counters == AMD64_NUM_COUNTERS_CORE) {
+ if (num_pmu_gp_counters == AMD64_NUM_COUNTERS_CORE ||
+ pmu_version > 1) {
sel_base = MSR_F15H_PERF_CTL0;
ctr_base = MSR_F15H_PERF_CTR0;
step = 2;
@@ -4822,6 +4843,12 @@ static int kvm_get_msrs(X86CPU *cpu)
kvm_msr_entry_add(cpu, ctr_base + i * step, 0);
kvm_msr_entry_add(cpu, sel_base + i * step, 0);
}
+
+ if (pmu_version > 1) {
+ kvm_msr_entry_add(cpu, MSR_AMD64_PERF_CNTR_GLOBAL_CTL, 0);
+ kvm_msr_entry_add(cpu, MSR_AMD64_PERF_CNTR_GLOBAL_STATUS, 0);
+ kvm_msr_entry_add(cpu, MSR_AMD64_PERF_CNTR_GLOBAL_STATUS_CLR, 0);
+ }
}
if (env->mcg_cap) {
@@ -5137,12 +5164,15 @@ static int kvm_get_msrs(X86CPU *cpu)
env->msr_fixed_ctr_ctrl = msrs[i].data;
break;
case MSR_CORE_PERF_GLOBAL_CTRL:
+ case MSR_AMD64_PERF_CNTR_GLOBAL_CTL:
env->msr_global_ctrl = msrs[i].data;
break;
case MSR_CORE_PERF_GLOBAL_STATUS:
+ case MSR_AMD64_PERF_CNTR_GLOBAL_STATUS:
env->msr_global_status = msrs[i].data;
break;
case MSR_CORE_PERF_GLOBAL_OVF_CTRL:
+ case MSR_AMD64_PERF_CNTR_GLOBAL_STATUS_CLR:
env->msr_global_ovf_ctrl = msrs[i].data;
break;
case MSR_CORE_PERF_FIXED_CTR0 ... MSR_CORE_PERF_FIXED_CTR0 + MAX_FIXED_COUNTERS - 1:
--
2.39.3
^ permalink raw reply related [flat|nested] 25+ messages in thread* [PATCH v8 7/7] target/i386/kvm: don't stop Intel PMU counters
2025-12-30 7:42 [PATCH v8 0/7] target/i386/kvm/pmu: PMU Enhancement, Bugfix and Cleanup Dongli Zhang
` (5 preceding siblings ...)
2025-12-30 7:42 ` [PATCH v8 6/7] target/i386/kvm: support perfmon-v2 for reset Dongli Zhang
@ 2025-12-30 7:42 ` Dongli Zhang
2026-01-03 0:27 ` Chen, Zide
6 siblings, 1 reply; 25+ messages in thread
From: Dongli Zhang @ 2025-12-30 7:42 UTC (permalink / raw)
To: qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
PMU MSRs are set by QEMU only at levels >= KVM_PUT_RESET_STATE,
excluding runtime. Therefore, updating these MSRs without stopping events
should be acceptable.
In addition, KVM creates kernel perf events with host mode excluded
(exclude_host = 1). While the events remain active, they don't increment
the counter during QEMU vCPU userspace mode.
Finally, The kvm_put_msrs() sets the MSRs using KVM_SET_MSRS. The x86 KVM
processes these MSRs one by one in a loop, only saving the config and
triggering the KVM_REQ_PMU request. This approach does not immediately stop
the event before updating PMC. This approach is true since Linux kernel
commit 68fb4757e867 ("KVM: x86/pmu: Defer reprogram_counter() to
kvm_pmu_handle_event"), that is, v6.2.
No Fixed tag is going to be added for the commit 0d89436786b0 ("kvm:
migrate vPMU state"), because this isn't a bugfix.
Signed-off-by: Dongli Zhang <dongli.zhang@oracle.com>
Reviewed-by: Zhao Liu <zhao1.liu@intel.com>
Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
---
Changed since v3:
- Re-order reasons in commit messages.
- Mention KVM's commit 68fb4757e867 (v6.2).
- Keep Zhao's review as there isn't code change.
Changed since v6:
- Add Reviewed-by from Dapeng Mi.
target/i386/kvm/kvm.c | 9 ---------
1 file changed, 9 deletions(-)
diff --git a/target/i386/kvm/kvm.c b/target/i386/kvm/kvm.c
index 99837048b8..742dc6ac0d 100644
--- a/target/i386/kvm/kvm.c
+++ b/target/i386/kvm/kvm.c
@@ -4222,13 +4222,6 @@ static int kvm_put_msrs(X86CPU *cpu, KvmPutState level)
}
if ((IS_INTEL_CPU(env) || IS_ZHAOXIN_CPU(env)) && pmu_version > 0) {
- if (pmu_version > 1) {
- /* Stop the counter. */
- kvm_msr_entry_add(cpu, MSR_CORE_PERF_FIXED_CTR_CTRL, 0);
- kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_CTRL, 0);
- }
-
- /* Set the counter values. */
for (i = 0; i < num_pmu_fixed_counters; i++) {
kvm_msr_entry_add(cpu, MSR_CORE_PERF_FIXED_CTR0 + i,
env->msr_fixed_counters[i]);
@@ -4244,8 +4237,6 @@ static int kvm_put_msrs(X86CPU *cpu, KvmPutState level)
env->msr_global_status);
kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_OVF_CTRL,
env->msr_global_ovf_ctrl);
-
- /* Now start the PMU. */
kvm_msr_entry_add(cpu, MSR_CORE_PERF_FIXED_CTR_CTRL,
env->msr_fixed_ctr_ctrl);
kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_CTRL,
--
2.39.3
^ permalink raw reply related [flat|nested] 25+ messages in thread* Re: [PATCH v8 7/7] target/i386/kvm: don't stop Intel PMU counters
2025-12-30 7:42 ` [PATCH v8 7/7] target/i386/kvm: don't stop Intel PMU counters Dongli Zhang
@ 2026-01-03 0:27 ` Chen, Zide
2026-01-05 20:24 ` Dongli Zhang
0 siblings, 1 reply; 25+ messages in thread
From: Chen, Zide @ 2026-01-03 0:27 UTC (permalink / raw)
To: Dongli Zhang, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
On 12/29/2025 11:42 PM, Dongli Zhang wrote:
> PMU MSRs are set by QEMU only at levels >= KVM_PUT_RESET_STATE,
> excluding runtime. Therefore, updating these MSRs without stopping events
> should be acceptable.
It seems preferable to keep the existing logic. The sequence of
disabling -> setting new counters -> re-enabling is complete and
reasonable. Re-enabling the PMU implicitly tell KVM to do whatever
actions are needed to make the new counters take effect.
If the purpose of this patch to improve performance, given that this is
a non-critical path, trading this clear and robust logic for a minor
performance gain does not seem necessary.
> In addition, KVM creates kernel perf events with host mode excluded
> (exclude_host = 1). While the events remain active, they don't increment
> the counter during QEMU vCPU userspace mode.
>
> Finally, The kvm_put_msrs() sets the MSRs using KVM_SET_MSRS. The x86 KVM
> processes these MSRs one by one in a loop, only saving the config and
> triggering the KVM_REQ_PMU request. This approach does not immediately stop
> the event before updating PMC. This approach is true since Linux kernel
> commit 68fb4757e867 ("KVM: x86/pmu: Defer reprogram_counter() to
> kvm_pmu_handle_event"), that is, v6.2.
This seems to assume KVM's internal behavior. While that is true today
(and possibly in the future), it’s not necessary for QEMU to make such
assumptions, as that could unnecessarily limit KVM’s flexibility to
change its behavior later.
> No Fixed tag is going to be added for the commit 0d89436786b0 ("kvm:
> migrate vPMU state"), because this isn't a bugfix.
>
> Signed-off-by: Dongli Zhang <dongli.zhang@oracle.com>
> Reviewed-by: Zhao Liu <zhao1.liu@intel.com>
> Reviewed-by: Dapeng Mi <dapeng1.mi@linux.intel.com>
> ---
> Changed since v3:
> - Re-order reasons in commit messages.
> - Mention KVM's commit 68fb4757e867 (v6.2).
> - Keep Zhao's review as there isn't code change.
> Changed since v6:
> - Add Reviewed-by from Dapeng Mi.
>
> target/i386/kvm/kvm.c | 9 ---------
> 1 file changed, 9 deletions(-)
>
> diff --git a/target/i386/kvm/kvm.c b/target/i386/kvm/kvm.c
> index 99837048b8..742dc6ac0d 100644
> --- a/target/i386/kvm/kvm.c
> +++ b/target/i386/kvm/kvm.c
> @@ -4222,13 +4222,6 @@ static int kvm_put_msrs(X86CPU *cpu, KvmPutState level)
> }
>
> if ((IS_INTEL_CPU(env) || IS_ZHAOXIN_CPU(env)) && pmu_version > 0) {
> - if (pmu_version > 1) {
> - /* Stop the counter. */
> - kvm_msr_entry_add(cpu, MSR_CORE_PERF_FIXED_CTR_CTRL, 0);
> - kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_CTRL, 0);
> - }
> -
> - /* Set the counter values. */
> for (i = 0; i < num_pmu_fixed_counters; i++) {
> kvm_msr_entry_add(cpu, MSR_CORE_PERF_FIXED_CTR0 + i,
> env->msr_fixed_counters[i]);
> @@ -4244,8 +4237,6 @@ static int kvm_put_msrs(X86CPU *cpu, KvmPutState level)
> env->msr_global_status);
> kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_OVF_CTRL,
> env->msr_global_ovf_ctrl);
> -
> - /* Now start the PMU. */
> kvm_msr_entry_add(cpu, MSR_CORE_PERF_FIXED_CTR_CTRL,
> env->msr_fixed_ctr_ctrl);
> kvm_msr_entry_add(cpu, MSR_CORE_PERF_GLOBAL_CTRL,
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v8 7/7] target/i386/kvm: don't stop Intel PMU counters
2026-01-03 0:27 ` Chen, Zide
@ 2026-01-05 20:24 ` Dongli Zhang
2026-01-06 21:04 ` Chen, Zide
0 siblings, 1 reply; 25+ messages in thread
From: Dongli Zhang @ 2026-01-05 20:24 UTC (permalink / raw)
To: Chen, Zide, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
Hi Zide,
On 1/2/26 4:27 PM, Chen, Zide wrote:
>
>
> On 12/29/2025 11:42 PM, Dongli Zhang wrote:
>> PMU MSRs are set by QEMU only at levels >= KVM_PUT_RESET_STATE,
>> excluding runtime. Therefore, updating these MSRs without stopping events
>> should be acceptable.
>
> It seems preferable to keep the existing logic. The sequence of
> disabling -> setting new counters -> re-enabling is complete and
> reasonable. Re-enabling the PMU implicitly tell KVM to do whatever
> actions are needed to make the new counters take effect.
>
> If the purpose of this patch to improve performance, given that this is
> a non-critical path, trading this clear and robust logic for a minor
> performance gain does not seem necessary.
>
>
>> In addition, KVM creates kernel perf events with host mode excluded
>> (exclude_host = 1). While the events remain active, they don't increment
>> the counter during QEMU vCPU userspace mode.
>>
>> Finally, The kvm_put_msrs() sets the MSRs using KVM_SET_MSRS. The x86 KVM
>> processes these MSRs one by one in a loop, only saving the config and
>> triggering the KVM_REQ_PMU request. This approach does not immediately stop
>> the event before updating PMC. This approach is true since Linux kernel
>> commit 68fb4757e867 ("KVM: x86/pmu: Defer reprogram_counter() to
>> kvm_pmu_handle_event"), that is, v6.2.
>
> This seems to assume KVM's internal behavior. While that is true today
> (and possibly in the future), it’s not necessary for QEMU to make such
> assumptions, as that could unnecessarily limit KVM’s flexibility to
> change its behavior later.
>
To "assume KVM's internal behavior" is only one of the two reasons. The first
reason is that QEMU controls the state of the vCPU to ensure this action only
occurs when "levels >= KVM_PUT_RESET_STATE."
Thanks to "(level >= KVM_PUT_RESET_STATE)", QEMU is able to avoid unnecessary
updates to many MSR registers during runtime.
The main objective is to sync the implementation for Intel and AMD.
Both MSR_CORE_PERF_FIXED_CTR_CTRL and MSR_CORE_PERF_GLOBAL_CTRL are reset to
zero only in the case where "has_pmu_version > 1." Otherwise, we may need to
reset the MSR_P6_PERFCTR_N registers before writing to the counter registers.
Without PATCH 7/7, an additional patch will be required to fix the workflow for
handling PMU registers to reset control registers before counter registers.
If the plan is to maintain the current logic, we may need to adjust the logic
for the AMD registers as well. In PATCH 6/7, we never reset global registers
before writing to control and counter registers.
Would you mine sharing your thoughts on it?
Thank you very much!
Dongli Zhang
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v8 7/7] target/i386/kvm: don't stop Intel PMU counters
2026-01-05 20:24 ` Dongli Zhang
@ 2026-01-06 21:04 ` Chen, Zide
2026-01-07 8:36 ` Dongli Zhang
0 siblings, 1 reply; 25+ messages in thread
From: Chen, Zide @ 2026-01-06 21:04 UTC (permalink / raw)
To: Dongli Zhang, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
On 1/5/2026 12:24 PM, Dongli Zhang wrote:
> Hi Zide,
>
> On 1/2/26 4:27 PM, Chen, Zide wrote:
>>
>>
>> On 12/29/2025 11:42 PM, Dongli Zhang wrote:
>>> PMU MSRs are set by QEMU only at levels >= KVM_PUT_RESET_STATE,
>>> excluding runtime. Therefore, updating these MSRs without stopping events
>>> should be acceptable.
>>
>> It seems preferable to keep the existing logic. The sequence of
>> disabling -> setting new counters -> re-enabling is complete and
>> reasonable. Re-enabling the PMU implicitly tell KVM to do whatever
>> actions are needed to make the new counters take effect.
>>
>> If the purpose of this patch to improve performance, given that this is
>> a non-critical path, trading this clear and robust logic for a minor
>> performance gain does not seem necessary.
>>
>>
>>> In addition, KVM creates kernel perf events with host mode excluded
>>> (exclude_host = 1). While the events remain active, they don't increment
>>> the counter during QEMU vCPU userspace mode.
>>>
>>> Finally, The kvm_put_msrs() sets the MSRs using KVM_SET_MSRS. The x86 KVM
>>> processes these MSRs one by one in a loop, only saving the config and
>>> triggering the KVM_REQ_PMU request. This approach does not immediately stop
>>> the event before updating PMC. This approach is true since Linux kernel
>>> commit 68fb4757e867 ("KVM: x86/pmu: Defer reprogram_counter() to
>>> kvm_pmu_handle_event"), that is, v6.2.
>>
>> This seems to assume KVM's internal behavior. While that is true today
>> (and possibly in the future), it’s not necessary for QEMU to make such
>> assumptions, as that could unnecessarily limit KVM’s flexibility to
>> change its behavior later.
>>
>
> To "assume KVM's internal behavior" is only one of the two reasons. The first
> reason is that QEMU controls the state of the vCPU to ensure this action only
> occurs when "levels >= KVM_PUT_RESET_STATE."
>
> Thanks to "(level >= KVM_PUT_RESET_STATE)", QEMU is able to avoid unnecessary
> updates to many MSR registers during runtime.
>
>
> The main objective is to sync the implementation for Intel and AMD.
>
> Both MSR_CORE_PERF_FIXED_CTR_CTRL and MSR_CORE_PERF_GLOBAL_CTRL are reset to
> zero only in the case where "has_pmu_version > 1." Otherwise, we may need to
> reset the MSR_P6_PERFCTR_N registers before writing to the counter registers.
> Without PATCH 7/7, an additional patch will be required to fix the workflow for
> handling PMU registers to reset control registers before counter registers.
I might be missing something here, but since this is not for runtime,
I don’t quite understand the need to reset the control registers.
> If the plan is to maintain the current logic, we may need to adjust the logic
> for the AMD registers as well. In PATCH 6/7, we never reset global registers
> before writing to control and counter registers.
>
> Would you mine sharing your thoughts on it?
Personally, I would lean towards keeping the current logic and instead
adjusting patch 6/7 to reset the global registers. This is just my
view, and please don’t feel obligated to follow it.
> Thank you very much!
>
> Dongli Zhang
^ permalink raw reply [flat|nested] 25+ messages in thread* Re: [PATCH v8 7/7] target/i386/kvm: don't stop Intel PMU counters
2026-01-06 21:04 ` Chen, Zide
@ 2026-01-07 8:36 ` Dongli Zhang
0 siblings, 0 replies; 25+ messages in thread
From: Dongli Zhang @ 2026-01-07 8:36 UTC (permalink / raw)
To: Chen, Zide, qemu-devel, kvm
Cc: pbonzini, zhao1.liu, mtosatti, sandipan.das, babu.moger, likexu,
like.xu.linux, groug, khorenko, alexander.ivanov, den,
davydov-max, xiaoyao.li, dapeng1.mi, joe.jin, ewanhai-oc, ewanhai
Hi Zide,
On 1/6/26 1:04 PM, Chen, Zide wrote:
>
>
> On 1/5/2026 12:24 PM, Dongli Zhang wrote:
>> Hi Zide,
>>
>> On 1/2/26 4:27 PM, Chen, Zide wrote:
>>>
>>>
>>> On 12/29/2025 11:42 PM, Dongli Zhang wrote:
>>>> PMU MSRs are set by QEMU only at levels >= KVM_PUT_RESET_STATE,
>>>> excluding runtime. Therefore, updating these MSRs without stopping events
>>>> should be acceptable.
>>>
>>> It seems preferable to keep the existing logic. The sequence of
>>> disabling -> setting new counters -> re-enabling is complete and
>>> reasonable. Re-enabling the PMU implicitly tell KVM to do whatever
>>> actions are needed to make the new counters take effect.
>>>
>>> If the purpose of this patch to improve performance, given that this is
>>> a non-critical path, trading this clear and robust logic for a minor
>>> performance gain does not seem necessary.
>>>
>>>
>>>> In addition, KVM creates kernel perf events with host mode excluded
>>>> (exclude_host = 1). While the events remain active, they don't increment
>>>> the counter during QEMU vCPU userspace mode.
>>>>
>>>> Finally, The kvm_put_msrs() sets the MSRs using KVM_SET_MSRS. The x86 KVM
>>>> processes these MSRs one by one in a loop, only saving the config and
>>>> triggering the KVM_REQ_PMU request. This approach does not immediately stop
>>>> the event before updating PMC. This approach is true since Linux kernel
>>>> commit 68fb4757e867 ("KVM: x86/pmu: Defer reprogram_counter() to
>>>> kvm_pmu_handle_event"), that is, v6.2.
>>>
>>> This seems to assume KVM's internal behavior. While that is true today
>>> (and possibly in the future), it’s not necessary for QEMU to make such
>>> assumptions, as that could unnecessarily limit KVM’s flexibility to
>>> change its behavior later.
>>>
>>
>> To "assume KVM's internal behavior" is only one of the two reasons. The first
>> reason is that QEMU controls the state of the vCPU to ensure this action only
>> occurs when "levels >= KVM_PUT_RESET_STATE."
>>
>> Thanks to "(level >= KVM_PUT_RESET_STATE)", QEMU is able to avoid unnecessary
>> updates to many MSR registers during runtime.
>>
>>
>> The main objective is to sync the implementation for Intel and AMD.
>>
>> Both MSR_CORE_PERF_FIXED_CTR_CTRL and MSR_CORE_PERF_GLOBAL_CTRL are reset to
>> zero only in the case where "has_pmu_version > 1." Otherwise, we may need to
>> reset the MSR_P6_PERFCTR_N registers before writing to the counter registers.
>> Without PATCH 7/7, an additional patch will be required to fix the workflow for
>> handling PMU registers to reset control registers before counter registers.
>
> I might be missing something here, but since this is not for runtime,
> I don’t quite understand the need to reset the control registers.
Global control registers take priority over per-counter control registers. If a
counter is disabled by the global register, enabling it in the per-counter
control register will have no effect.
However, assuming !(has_architectural_pmu_version > 1), which is highly unlikely
in modern systems, there will be no support for the global register.
As a result, QEMU should:
1. Set the control register to 0 for each counter.
2. Write to the counter.
3. Write to the control register.
From the KVM source code, the minimum version is 2, although QEMU assumes any
version can be utilized.
>
>> If the plan is to maintain the current logic, we may need to adjust the logic
>> for the AMD registers as well. In PATCH 6/7, we never reset global registers
>> before writing to control and counter registers.
>>
>> Would you mine sharing your thoughts on it?
>
> Personally, I would lean towards keeping the current logic and instead
> adjusting patch 6/7 to reset the global registers. This is just my
> view, and please don’t feel obligated to follow it.
>
It just seems odd to me that we would need to emulate PMU MSR usage in a VM,
i.e., stop all counters in the global registers at the beginning, when they
won't actually take effect.
Perhaps other reviewers have some insights?
Thank you very much!
Dongli Zhang
^ permalink raw reply [flat|nested] 25+ messages in thread