* [PATCH v6 0/2] KVM: VMX: Workaround VMX preemption timer erratum
@ 2026-07-29 13:42 Sean Christopherson
2026-07-29 13:42 ` [PATCH v6 1/2] KVM: VMX: Bury all of the VMX preemption timer code under CONFIG_X86_64=y Sean Christopherson
2026-07-29 13:42 ` [PATCH v6 2/2] KVM: VMX: Cap VMX preemption timer to work around Intel erratum Sean Christopherson
0 siblings, 2 replies; 5+ messages in thread
From: Sean Christopherson @ 2026-07-29 13:42 UTC (permalink / raw)
To: Sean Christopherson, Paolo Bonzini
Cc: kvm, linux-kernel, Binbin Wu, Chao Gao, Jim Mattson
Due to a widespread Intel erratum (e.g. EMR158), programming the VMX
preemption timer with certain large values may cause the timer to expire
earlier than expected. The recommended workaround is to cap the timer value
to strictly less than 2^25 * CPUID.15H:EBX[31:0] / CPUID.15H:EAX[31:0].
v6:
- Collect reviews. [Chao, Binbin]
- Massage the wording in patch 1's changelog. [Binbin]
- Use an inclusive "max" instead of an exclusive "limit".
- Define a local ARCHITECTURAL_MAX_VALUE instead of open coding it in multiple
locations.
- Don't apply the workaround when running as a guest.
- WARN if the CPUID.0x15 information would result in max value of 0.
v5:
- https://lore.kernel.org/all/20260724234914.987987-1-jmattson@google.com
- Include Sean's preparatory patch burying VMX preemption timer code under
CONFIG_X86_64=y [Sean]
- Replace div_u64() with direct division (/)
v4: https://lore.kernel.org/all/20260722143024.3938899-1-jmattson@google.com
v3: https://lore.kernel.org/all/20260722040311.3369898-1-jmattson@google.com
v2: https://lore.kernel.org/all/20260720231639.1592848-1-jmattson@google.com
v1: https://lore.kernel.org/all/20260720205230.1457146-1-jmattson@google.com
Jim Mattson (1):
KVM: VMX: Cap VMX preemption timer to work around Intel erratum
Sean Christopherson (1):
KVM: VMX: Bury all of the VMX preemption timer code under
CONFIG_X86_64=y
arch/x86/kvm/vmx/vmx.c | 158 ++++++++++++++++++++++++++---------------
1 file changed, 102 insertions(+), 56 deletions(-)
base-commit: 3c7d7f908d574277a845423ec32250a8d8df44c8
--
2.55.0.487.gaf234c4eb3-goog
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v6 1/2] KVM: VMX: Bury all of the VMX preemption timer code under CONFIG_X86_64=y
2026-07-29 13:42 [PATCH v6 0/2] KVM: VMX: Workaround VMX preemption timer erratum Sean Christopherson
@ 2026-07-29 13:42 ` Sean Christopherson
2026-07-29 13:42 ` [PATCH v6 2/2] KVM: VMX: Cap VMX preemption timer to work around Intel erratum Sean Christopherson
1 sibling, 0 replies; 5+ messages in thread
From: Sean Christopherson @ 2026-07-29 13:42 UTC (permalink / raw)
To: Sean Christopherson, Paolo Bonzini
Cc: kvm, linux-kernel, Binbin Wu, Chao Gao, Jim Mattson
Double down on using the VMX preemption timer only for 64-bit kernels, and
bury the setup and runtime adjustment code, and all global variables, under
CONFIG_X86_64=y. This will allow addressing a widespread Intel erratum
without running afoul of unused-but-set-variable and __udivdi3() errors on
32-bit kernels.
No functional change intended.
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Reviewed-by: Chao Gao <chao.gao@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
arch/x86/kvm/vmx/vmx.c | 122 +++++++++++++++++++++++------------------
1 file changed, 69 insertions(+), 53 deletions(-)
diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c
index e4b9ac7fed9f..a07faa066ef0 100644
--- a/arch/x86/kvm/vmx/vmx.c
+++ b/arch/x86/kvm/vmx/vmx.c
@@ -150,10 +150,12 @@ module_param(dump_invalid_vmcs, bool, 0644);
#define KVM_VMX_TSC_MULTIPLIER_MAX 0xffffffffffffffffULL
/* Guest_tsc -> host_tsc conversion requires 64-bit division. */
+#ifdef CONFIG_X86_64
static int __read_mostly cpu_preemption_timer_multi;
static bool __read_mostly enable_preemption_timer = 1;
-#ifdef CONFIG_X86_64
module_param_named(preemption_timer, enable_preemption_timer, bool, S_IRUGO);
+#else
+#define enable_preemption_timer false
#endif
extern bool __read_mostly allow_smaller_maxphyaddr;
@@ -7408,32 +7410,6 @@ static void vmx_refresh_guest_perf_global_control(struct kvm_vcpu *vcpu)
pmu->global_ctrl = vmcs_read64(GUEST_IA32_PERF_GLOBAL_CTRL);
}
-static void vmx_update_hv_timer(struct kvm_vcpu *vcpu, bool force_immediate_exit)
-{
- struct vcpu_vmx *vmx = to_vmx(vcpu);
- u64 tscl;
- u32 delta_tsc;
-
- if (force_immediate_exit) {
- vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, 0);
- vmx->loaded_vmcs->hv_timer_soft_disabled = false;
- } else if (vmx->hv_deadline_tsc != -1) {
- tscl = rdtsc();
- if (vmx->hv_deadline_tsc > tscl)
- /* set_hv_timer ensures the delta fits in 32-bits */
- delta_tsc = (u32)((vmx->hv_deadline_tsc - tscl) >>
- cpu_preemption_timer_multi);
- else
- delta_tsc = 0;
-
- vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, delta_tsc);
- vmx->loaded_vmcs->hv_timer_soft_disabled = false;
- } else if (!vmx->loaded_vmcs->hv_timer_soft_disabled) {
- vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, -1);
- vmx->loaded_vmcs->hv_timer_soft_disabled = true;
- }
-}
-
void noinstr vmx_update_host_rsp(struct vcpu_vmx *vmx, unsigned long host_rsp)
{
if (unlikely(host_rsp != vmx->loaded_vmcs->host_state.rsp)) {
@@ -7519,6 +7495,8 @@ static noinstr void vmx_vcpu_enter_exit(struct kvm_vcpu *vcpu,
guest_state_exit_irqoff();
}
+static void vmx_update_hv_timer(struct kvm_vcpu *vcpu, bool force_immediate_exit);
+
fastpath_t vmx_vcpu_run(struct kvm_vcpu *vcpu, u64 run_flags)
{
bool force_immediate_exit = run_flags & KVM_RUN_FORCE_IMMEDIATE_EXIT;
@@ -8328,6 +8306,36 @@ static inline int u64_shl_div_u64(u64 a, unsigned int shift,
return 0;
}
+static __init void vmx_setup_preemption_timer(void)
+{
+ if (!cpu_has_vmx_preemption_timer())
+ enable_preemption_timer = false;
+
+ if (enable_preemption_timer) {
+ u64 use_timer_freq = 5000ULL * 1000 * 1000;
+
+ cpu_preemption_timer_multi =
+ vmx_misc_preemption_timer_rate(vmcs_config.misc);
+
+ if (tsc_khz)
+ use_timer_freq = (u64)tsc_khz * 1000;
+ use_timer_freq >>= cpu_preemption_timer_multi;
+
+ /*
+ * KVM "disables" the preemption timer by setting it to its max
+ * value. Don't use the timer if it might cause spurious exits
+ * at a rate faster than 0.1 Hz (of uninterrupted guest time).
+ */
+ if (use_timer_freq > 0xffffffffu / 10)
+ enable_preemption_timer = false;
+ }
+
+ if (!enable_preemption_timer) {
+ vt_x86_ops.set_hv_timer = NULL;
+ vt_x86_ops.cancel_hv_timer = NULL;
+ }
+}
+
int vmx_set_hv_timer(struct kvm_vcpu *vcpu, u64 guest_deadline_tsc,
bool *expired)
{
@@ -8372,6 +8380,39 @@ void vmx_cancel_hv_timer(struct kvm_vcpu *vcpu)
{
to_vmx(vcpu)->hv_deadline_tsc = -1;
}
+
+static void vmx_update_hv_timer(struct kvm_vcpu *vcpu, bool force_immediate_exit)
+{
+ struct vcpu_vmx *vmx = to_vmx(vcpu);
+ u64 tscl;
+ u32 delta_tsc;
+
+ if (force_immediate_exit) {
+ vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, 0);
+ vmx->loaded_vmcs->hv_timer_soft_disabled = false;
+ } else if (vmx->hv_deadline_tsc != -1) {
+ tscl = rdtsc();
+ if (vmx->hv_deadline_tsc > tscl)
+ /* set_hv_timer ensures the delta fits in 32-bits */
+ delta_tsc = (u32)((vmx->hv_deadline_tsc - tscl) >>
+ cpu_preemption_timer_multi);
+ else
+ delta_tsc = 0;
+
+ vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, delta_tsc);
+ vmx->loaded_vmcs->hv_timer_soft_disabled = false;
+ } else if (!vmx->loaded_vmcs->hv_timer_soft_disabled) {
+ vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, -1);
+ vmx->loaded_vmcs->hv_timer_soft_disabled = true;
+ }
+}
+#else
+static __init void vmx_setup_preemption_timer(void) { }
+
+static void vmx_update_hv_timer(struct kvm_vcpu *vcpu, bool force_immediate_exit)
+{
+ BUILD_BUG_ON(1);
+}
#endif
void vmx_update_cpu_dirty_logging(struct kvm_vcpu *vcpu)
@@ -8734,32 +8775,7 @@ __init int vmx_hardware_setup(void)
if (!enable_ept || !enable_ept_ad_bits || !cpu_has_vmx_pml())
enable_pml = 0;
- if (!cpu_has_vmx_preemption_timer())
- enable_preemption_timer = false;
-
- if (enable_preemption_timer) {
- u64 use_timer_freq = 5000ULL * 1000 * 1000;
-
- cpu_preemption_timer_multi =
- vmx_misc_preemption_timer_rate(vmcs_config.misc);
-
- if (tsc_khz)
- use_timer_freq = (u64)tsc_khz * 1000;
- use_timer_freq >>= cpu_preemption_timer_multi;
-
- /*
- * KVM "disables" the preemption timer by setting it to its max
- * value. Don't use the timer if it might cause spurious exits
- * at a rate faster than 0.1 Hz (of uninterrupted guest time).
- */
- if (use_timer_freq > 0xffffffffu / 10)
- enable_preemption_timer = false;
- }
-
- if (!enable_preemption_timer) {
- vt_x86_ops.set_hv_timer = NULL;
- vt_x86_ops.cancel_hv_timer = NULL;
- }
+ vmx_setup_preemption_timer();
kvm_caps.supported_mce_cap |= MCG_LMCE_P;
kvm_caps.supported_mce_cap |= MCG_CMCI_P;
--
2.55.0.487.gaf234c4eb3-goog
^ permalink raw reply related [flat|nested] 5+ messages in thread
* [PATCH v6 2/2] KVM: VMX: Cap VMX preemption timer to work around Intel erratum
2026-07-29 13:42 [PATCH v6 0/2] KVM: VMX: Workaround VMX preemption timer erratum Sean Christopherson
2026-07-29 13:42 ` [PATCH v6 1/2] KVM: VMX: Bury all of the VMX preemption timer code under CONFIG_X86_64=y Sean Christopherson
@ 2026-07-29 13:42 ` Sean Christopherson
2026-07-29 13:50 ` sashiko-bot
1 sibling, 1 reply; 5+ messages in thread
From: Sean Christopherson @ 2026-07-29 13:42 UTC (permalink / raw)
To: Sean Christopherson, Paolo Bonzini
Cc: kvm, linux-kernel, Binbin Wu, Chao Gao, Jim Mattson
From: Jim Mattson <jmattson@google.com>
Due to a widespread Intel erratum (e.g. EMR158), programming the
VMX-preemption timer with certain large values may cause the timer to
expire earlier than expected. The recommended workaround is to cap the
VMX-preemption timer value to strictly less than:
2^25 * CPUID.15H:EBX[31:0] / CPUID.15H:EAX[31:0].
Calculate the maximum "safe" preemption timer value during hardware setup
based on CPUID 15H when available, and use the adjusted max value in all
locations where KVM currently hardcodes the max architectural value,
including in the subtle case where KVM soft-disables the timer.
Don't apply the workaround when running as a VM, because absent explicit
enumeration to state the bug is present (or not), it's L0's responsibility
to faithfully emulate/virtualize the VMX preemption timer.
WARN if the above logic would result in a max value of zero and fall back
to the maximum architectural value, as the expectation is that real
hardware will never provide problematic EAX/EBX values (which is another
reason to ignore the erratum when running as a VM; there's less chance of
a false positive on the WARN due to L0 providing an unanticipated ratio).
Reported-by: Sean Christopherson <seanjc@google.com>
Closes: https://lore.kernel.org/all/Zn9X0yFxZi_Mrlnt@google.com/
Suggested-by: Chao Gao <chao.gao@intel.com>
Assisted-by: Gemini:Gemini-Next
Reviewed-by: Chao Gao <chao.gao@intel.com>
Signed-off-by: Jim Mattson <jmattson@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
[sean: track inclusive max instead of exclusive limit, massage changelog]
Signed-off-by: Sean Christopherson <seanjc@google.com>
---
arch/x86/kvm/vmx/vmx.c | 40 +++++++++++++++++++++++++++++++++++-----
1 file changed, 35 insertions(+), 5 deletions(-)
diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c
index a07faa066ef0..dae41842a9ce 100644
--- a/arch/x86/kvm/vmx/vmx.c
+++ b/arch/x86/kvm/vmx/vmx.c
@@ -153,6 +153,7 @@ module_param(dump_invalid_vmcs, bool, 0644);
#ifdef CONFIG_X86_64
static int __read_mostly cpu_preemption_timer_multi;
static bool __read_mostly enable_preemption_timer = 1;
+static u64 __ro_after_init preemption_timer_max_value;
module_param_named(preemption_timer, enable_preemption_timer, bool, S_IRUGO);
#else
#define enable_preemption_timer false
@@ -8306,6 +8307,33 @@ static inline int u64_shl_div_u64(u64 a, unsigned int shift,
return 0;
}
+/*
+ * Workaround for a widespread Intel erratum (e.g. EMR158) where the
+ * VMX-preemption timer may expire earlier than expected when programmed
+ * with large values. The workaround is to cap the timer value to strictly
+ * less than 2^25 * CPUID.15H:EBX / CPUID.15H:EAX.
+ */
+static __init u64 calc_preemption_timer_max_value(void)
+{
+ const u64 ARCHITECTURAL_MAX_VALUE = UINT_MAX;
+ u32 eax, ebx, ecx, edx;
+
+ if (cpu_feature_enabled(X86_FEATURE_HYPERVISOR))
+ return ARCHITECTURAL_MAX_VALUE;
+
+ if (cpuid_eax(0) < 0x15)
+ return ARCHITECTURAL_MAX_VALUE;
+
+ cpuid(0x15, &eax, &ebx, &ecx, &edx);
+ if (!eax || !ebx)
+ return ARCHITECTURAL_MAX_VALUE;
+
+ if (WARN_ON_ONCE(!(((u64)ebx << 25) / eax)))
+ return ARCHITECTURAL_MAX_VALUE;
+
+ return (((u64)ebx << 25) / eax) - 1;
+}
+
static __init void vmx_setup_preemption_timer(void)
{
if (!cpu_has_vmx_preemption_timer())
@@ -8317,6 +8345,8 @@ static __init void vmx_setup_preemption_timer(void)
cpu_preemption_timer_multi =
vmx_misc_preemption_timer_rate(vmcs_config.misc);
+ preemption_timer_max_value = calc_preemption_timer_max_value();
+
if (tsc_khz)
use_timer_freq = (u64)tsc_khz * 1000;
use_timer_freq >>= cpu_preemption_timer_multi;
@@ -8326,7 +8356,7 @@ static __init void vmx_setup_preemption_timer(void)
* value. Don't use the timer if it might cause spurious exits
* at a rate faster than 0.1 Hz (of uninterrupted guest time).
*/
- if (use_timer_freq > 0xffffffffu / 10)
+ if (use_timer_freq > preemption_timer_max_value / 10)
enable_preemption_timer = false;
}
@@ -8363,12 +8393,12 @@ int vmx_set_hv_timer(struct kvm_vcpu *vcpu, u64 guest_deadline_tsc,
return -ERANGE;
/*
- * If the delta tsc can't fit in the 32 bit after the multi shift,
- * we can't use the preemption timer.
+ * If the delta tsc exceeds the preemption timer limit after the
+ * multi shift, we can't use the preemption timer.
* It's possible that it fits on later vmentries, but checking
* on every vmentry is costly so we just use an hrtimer.
*/
- if (delta_tsc >> (cpu_preemption_timer_multi + 32))
+ if ((delta_tsc >> cpu_preemption_timer_multi) > preemption_timer_max_value)
return -ERANGE;
vmx->hv_deadline_tsc = tscl + delta_tsc;
@@ -8402,7 +8432,7 @@ static void vmx_update_hv_timer(struct kvm_vcpu *vcpu, bool force_immediate_exit
vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, delta_tsc);
vmx->loaded_vmcs->hv_timer_soft_disabled = false;
} else if (!vmx->loaded_vmcs->hv_timer_soft_disabled) {
- vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, -1);
+ vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, preemption_timer_max_value);
vmx->loaded_vmcs->hv_timer_soft_disabled = true;
}
}
--
2.55.0.487.gaf234c4eb3-goog
^ permalink raw reply related [flat|nested] 5+ messages in thread
* Re: [PATCH v6 2/2] KVM: VMX: Cap VMX preemption timer to work around Intel erratum
2026-07-29 13:42 ` [PATCH v6 2/2] KVM: VMX: Cap VMX preemption timer to work around Intel erratum Sean Christopherson
@ 2026-07-29 13:50 ` sashiko-bot
2026-07-29 13:57 ` Sean Christopherson
0 siblings, 1 reply; 5+ messages in thread
From: sashiko-bot @ 2026-07-29 13:50 UTC (permalink / raw)
To: Sean Christopherson; +Cc: kvm
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] The computed VMX preemption timer maximum value can exceed 32 bits on modern CPUs because it is not clamped to the architectural limit (UINT_MAX). This leads to silent 32-bit truncation when written to the VMCS, breaking the timer's soft-disable mechanism and causing premature guest timer expirations.
--
--- Patch [2]: [PATCH v6 2/2] KVM: VMX: Cap VMX preemption timer to work around Intel erratum ---
commit 021d423821c3a8ec582691801fee4d57b9a316fc
Author: Jim Mattson <jmattson@google.com>
KVM: VMX: Cap VMX preemption timer to work around Intel erratum
This commit introduces a workaround for an Intel erratum where the VMX
preemption timer can expire early if programmed with large values. It
calculates a safe maximum value based on CPUID.15H. This maximum is then
used when the timer is soft-disabled and to bound delta_tsc.
Closes: https://lore.kernel.org/all/Zn9X0yFxZi_Mrlnt@google.com/
> diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c
> index 7d2cc36a04e51..67c8dc1eb2c4f 100644
> --- a/arch/x86/kvm/vmx/vmx.c
> +++ b/arch/x86/kvm/vmx/vmx.c
[ ... ]
> @@ -8305,6 +8306,33 @@ static inline int u64_shl_div_u64(u64 a, unsigned int shift,
> return 0;
> }
>
> +/*
> + * Workaround for a widespread Intel erratum (e.g. EMR158) where the
> + * VMX-preemption timer may expire earlier than expected when programmed
> + * with large values. The workaround is to cap the timer value to strictly
> + * less than 2^25 * CPUID.15H:EBX / CPUID.15H:EAX.
> + */
> +static __init u64 calc_preemption_timer_max_value(void)
> +{
> + const u64 ARCHITECTURAL_MAX_VALUE = UINT_MAX;
> + u32 eax, ebx, ecx, edx;
[ ... ]
> + if (WARN_ON_ONCE(!(((u64)ebx << 25) / eax)))
> + return ARCHITECTURAL_MAX_VALUE;
> +
> + return (((u64)ebx << 25) / eax) - 1;
[Severity: High]
If a CPU has a high TSC-to-crystal clock ratio (for example, EBX/EAX >= 128),
could this calculation exceed UINT_MAX? Because the returned u64 is not clamped
to ARCHITECTURAL_MAX_VALUE, preemption_timer_max_value could end up storing a
value larger than 32 bits.
> +}
[ ... ]
> @@ -8362,12 +8392,12 @@ int vmx_set_hv_timer(struct kvm_vcpu *vcpu, u64 guest_deadline_tsc,
> return -ERANGE;
>
> /*
> - * If the delta tsc can't fit in the 32 bit after the multi shift,
> - * we can't use the preemption timer.
> + * If the delta tsc exceeds the preemption timer limit after the
> + * multi shift, we can't use the preemption timer.
> * It's possible that it fits on later vmentries, but checking
> * on every vmentry is costly so we just use an hrtimer.
> */
> - if (delta_tsc >> (cpu_preemption_timer_multi + 32))
> + if ((delta_tsc >> cpu_preemption_timer_multi) > preemption_timer_max_value)
> return -ERANGE;
[Severity: High]
If preemption_timer_max_value is greater than UINT_MAX, does this condition fail
to catch delta_tsc limits that exceed 32 bits? It appears this might bypass the
required hrtimer fallback for large guest deadlines, as the value will later be
silently truncated.
>
> vmx->hv_deadline_tsc = tscl + delta_tsc;
[ ... ]
> @@ -8401,7 +8431,7 @@ static void vmx_update_hv_timer(struct kvm_vcpu *vcpu, bool force_immediate_exit
> vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, delta_tsc);
> vmx->loaded_vmcs->hv_timer_soft_disabled = false;
> } else if (!vmx->loaded_vmcs->hv_timer_soft_disabled) {
> - vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, -1);
> + vmcs_write32(VMX_PREEMPTION_TIMER_VALUE, preemption_timer_max_value);
> vmx->loaded_vmcs->hv_timer_soft_disabled = true;
> }
> }
[Severity: High]
Will writing an un-clamped u64 to vmcs_write32() cause it to be truncated? If
truncated, the preemption timer could fire prematurely when soft-disabled
instead of waiting for the maximum duration.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260729134203.1377606-1-seanjc@google.com?part=2
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v6 2/2] KVM: VMX: Cap VMX preemption timer to work around Intel erratum
2026-07-29 13:50 ` sashiko-bot
@ 2026-07-29 13:57 ` Sean Christopherson
0 siblings, 0 replies; 5+ messages in thread
From: Sean Christopherson @ 2026-07-29 13:57 UTC (permalink / raw)
To: sashiko-reviews; +Cc: kvm
On Wed, Jul 29, 2026, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> > diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c
> > index 7d2cc36a04e51..67c8dc1eb2c4f 100644
> > --- a/arch/x86/kvm/vmx/vmx.c
> > +++ b/arch/x86/kvm/vmx/vmx.c
>
> [ ... ]
>
> > @@ -8305,6 +8306,33 @@ static inline int u64_shl_div_u64(u64 a, unsigned int shift,
> > return 0;
> > }
> >
> > +/*
> > + * Workaround for a widespread Intel erratum (e.g. EMR158) where the
> > + * VMX-preemption timer may expire earlier than expected when programmed
> > + * with large values. The workaround is to cap the timer value to strictly
> > + * less than 2^25 * CPUID.15H:EBX / CPUID.15H:EAX.
> > + */
> > +static __init u64 calc_preemption_timer_max_value(void)
> > +{
> > + const u64 ARCHITECTURAL_MAX_VALUE = UINT_MAX;
> > + u32 eax, ebx, ecx, edx;
>
> [ ... ]
>
> > + if (WARN_ON_ONCE(!(((u64)ebx << 25) / eax)))
> > + return ARCHITECTURAL_MAX_VALUE;
> > +
> > + return (((u64)ebx << 25) / eax) - 1;
>
> [Severity: High]
> If a CPU has a high TSC-to-crystal clock ratio (for example, EBX/EAX >= 128),
> could this calculation exceed UINT_MAX? Because the returned u64 is not clamped
> to ARCHITECTURAL_MAX_VALUE, preemption_timer_max_value could end up storing a
> value larger than 32 bits.
/facepalm
I completely misread this:
return min_t(u64, 1ULL << 32, ((u64)ebx << 25) / eax);
as max_t. I'll add it back...
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-07-29 13:58 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-29 13:42 [PATCH v6 0/2] KVM: VMX: Workaround VMX preemption timer erratum Sean Christopherson
2026-07-29 13:42 ` [PATCH v6 1/2] KVM: VMX: Bury all of the VMX preemption timer code under CONFIG_X86_64=y Sean Christopherson
2026-07-29 13:42 ` [PATCH v6 2/2] KVM: VMX: Cap VMX preemption timer to work around Intel erratum Sean Christopherson
2026-07-29 13:50 ` sashiko-bot
2026-07-29 13:57 ` Sean Christopherson
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.