Kernel KVM virtualization development
 help / color / mirror / Atom feed
* [PATCH] KVM: x86: Guard against division by zero in adjust_lapic_timer_advance()
@ 2026-10-08  9:00 Peng Hao
  2026-10-08  9:15 ` sashiko-bot
  2026-10-09 20:35 ` Sean Christopherson
  0 siblings, 2 replies; 3+ messages in thread
From: Peng Hao @ 2026-10-08  9:00 UTC (permalink / raw)
  To: pbonzini, seanjc; +Cc: kvm

When TSC calibration fails and userspace never programs a tsc khz,
vcpu->arch.virtual_tsc_khz stays zero, but the lapic timer's
software-mode expiry path still calls adjust_lapic_timer_advance()
unconditionally on every timer expiration (unlike
__wait_lapic_expire(), this path is not gated on tsc_scaling_ratio).
Both the "too early" and "too late" branches do_div() by
virtual_tsc_khz, so a zero khz crashes the host with #DE as soon as
the timer advance drifts outside the adjust window.

Bail out before either division.

Signed-off-by: Peng Hao <flyingpeng@tencent.com>
---
 arch/x86/kvm/lapic.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/arch/x86/kvm/lapic.c b/arch/x86/kvm/lapic.c
index 13e32c4915e2..4dcbc66c8e07 100644
--- a/arch/x86/kvm/lapic.c
+++ b/arch/x86/kvm/lapic.c
@@ -1991,6 +1991,9 @@ static inline void adjust_lapic_timer_advance(struct kvm_vcpu *vcpu,
 	    abs(advance_expire_delta) < LAPIC_TIMER_ADVANCE_ADJUST_MIN)
 		return;
 
+	if (unlikely(!vcpu->arch.virtual_tsc_khz))
+		return;
+
 	/* too early */
 	if (advance_expire_delta < 0) {
 		ns = -advance_expire_delta * 1000000ULL;
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 3+ messages in thread

* Re: [PATCH] KVM: x86: Guard against division by zero in adjust_lapic_timer_advance()
  2026-10-08  9:00 [PATCH] KVM: x86: Guard against division by zero in adjust_lapic_timer_advance() Peng Hao
@ 2026-10-08  9:15 ` sashiko-bot
  2026-10-09 20:35 ` Sean Christopherson
  1 sibling, 0 replies; 3+ messages in thread
From: sashiko-bot @ 2026-10-08  9:15 UTC (permalink / raw)
  To: Peng Hao; +Cc: kvm

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] kvm: x86: divide by zero in __wait_lapic_expire() with nested TSC scaling
--

commit 615a870f84e9397c4bfba8db0f187034794e3d9e
Author: Peng Hao <flyingpenghao@gmail.com>

KVM: x86: Guard against division by zero in adjust_lapic_timer_advance()

When TSC calibration fails, vcpu->arch.virtual_tsc_khz is zero. This patch
adds a check to avoid division by zero in adjust_lapic_timer_advance().

>     When TSC calibration fails and userspace never programs a tsc khz,
>     vcpu->arch.virtual_tsc_khz stays zero, but the lapic timer's
>     software-mode expiry path still calls adjust_lapic_timer_advance()
>     unconditionally on every timer expiration (unlike
>     __wait_lapic_expire(), this path is not gated on tsc_scaling_ratio).

[ ... ]

> diff --git a/arch/x86/kvm/lapic.c b/arch/x86/kvm/lapic.c
> index 48b019114c196..dc4a669d1d460 100644
> --- a/arch/x86/kvm/lapic.c
> +++ b/arch/x86/kvm/lapic.c
> @@ -1978,6 +1978,9 @@ static inline void adjust_lapic_timer_advance(struct kvm_vcpu *vcpu,
>  	    abs(advance_expire_delta) < LAPIC_TIMER_ADVANCE_ADJUST_MIN)
>  		return;
>  
> +	if (unlikely(!vcpu->arch.virtual_tsc_khz))
> +		return;
> +

[Severity: High]
Does __wait_lapic_expire() need this same check?

The commit message mentions __wait_lapic_expire() is safely gated by
tsc_scaling_ratio, but is that description accurate when nested
virtualization is enabled?

An L1 guest can enable nested virtualization and set a custom tsc_multiplier
in VMCS12 for its L2 guest. KVM computes vcpu->arch.tsc_scaling_ratio to this
non-default value:

arch/x86/kvm/vmx/nested.c:prepare_vmcs02_early() {
    ...
	vcpu->arch.tsc_scaling_ratio = kvm_calc_nested_tsc_multiplier(
			vcpu->arch.l1_tsc_scaling_ratio,
			vmx_get_l2_tsc_multiplier(vcpu));
    ...
}

If host TSC calibration failed and virtual_tsc_khz is 0, when the LAPIC timer
expires during L2 execution, __wait_lapic_expire() takes the else branch
because tsc_scaling_ratio != default_tsc_scaling_ratio:

arch/x86/kvm/lapic.c:__wait_lapic_expire() {
    ...
	if (vcpu->arch.tsc_scaling_ratio == kvm_caps.default_tsc_scaling_ratio) {
		__delay(min(guest_cycles,
			nsec_to_cycles(vcpu, timer_advance_ns)));
	} else {
		u64 delay_ns = guest_cycles * 1000000ULL;
		do_div(delay_ns, vcpu->arch.virtual_tsc_khz);
		ndelay(min_t(u32, delay_ns, timer_advance_ns));
	}
    ...
}

Can an untrusted L1 guest use this path to reliably crash the host with a
divide-by-zero panic?

>  	/* too early */
>  	if (advance_expire_delta < 0) {
>  		ns = -advance_expire_delta * 1000000ULL;

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261008090037.86931-1-flyingpeng@tencent.com?part=1

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] KVM: x86: Guard against division by zero in adjust_lapic_timer_advance()
  2026-10-08  9:00 [PATCH] KVM: x86: Guard against division by zero in adjust_lapic_timer_advance() Peng Hao
  2026-10-08  9:15 ` sashiko-bot
@ 2026-10-09 20:35 ` Sean Christopherson
  1 sibling, 0 replies; 3+ messages in thread
From: Sean Christopherson @ 2026-10-09 20:35 UTC (permalink / raw)
  To: Peng Hao; +Cc: pbonzini, kvm

On Thu, Oct 08, 2026, Peng Hao wrote:
> When TSC calibration fails and userspace never programs a tsc khz,
> vcpu->arch.virtual_tsc_khz stays zero, but the lapic timer's
> software-mode expiry path still calls adjust_lapic_timer_advance()
> unconditionally on every timer expiration (unlike
> __wait_lapic_expire(), this path is not gated on tsc_scaling_ratio).
> Both the "too early" and "too late" branches do_div() by
> virtual_tsc_khz, so a zero khz crashes the host with #DE as soon as
> the timer advance drifts outside the adjust window.
> 
> Bail out before either division.
> 
> Signed-off-by: Peng Hao <flyingpeng@tencent.com>
> ---
>  arch/x86/kvm/lapic.c | 3 +++
>  1 file changed, 3 insertions(+)
> 
> diff --git a/arch/x86/kvm/lapic.c b/arch/x86/kvm/lapic.c
> index 13e32c4915e2..4dcbc66c8e07 100644
> --- a/arch/x86/kvm/lapic.c
> +++ b/arch/x86/kvm/lapic.c
> @@ -1991,6 +1991,9 @@ static inline void adjust_lapic_timer_advance(struct kvm_vcpu *vcpu,
>  	    abs(advance_expire_delta) < LAPIC_TIMER_ADVANCE_ADJUST_MIN)
>  		return;
>  
> +	if (unlikely(!vcpu->arch.virtual_tsc_khz))
> +		return;

Ugh.  I'd rather reject KVM_CREATE_VCPU, even though there is a rather surprising
amount of code in KVM that does indeed play nice with tsc_khz == 0.  I don't see
how the guest can possibly function with virtual_tsc_khz==0.

vcpu->arch.virtual_tsc_{shift,mult} will also be left as zero (to avoid #DE there
as well):

	/* tsc_khz can be zero if TSC calibration fails */
	if (user_tsc_khz == 0) {
		/* set tsc_scaling_ratio to a safe value */
		kvm_vcpu_write_tsc_multiplier(vcpu, kvm_caps.default_tsc_scaling_ratio);
		return -1;
	}

	/* Compute a scale to convert nanoseconds in TSC cycles */
	kvm_get_time_scale(user_tsc_khz * 1000LL, NSEC_PER_SEC,
			   &vcpu->arch.virtual_tsc_shift,
			   &vcpu->arch.virtual_tsc_mult);

and so nsec_to_cycles() will always return zero due to multiplying by zero.  That
means things like the APIC timer will always fire immediately:

	apic->lapic_timer.tscdeadline = kvm_read_l1_tsc(apic->vcpu, tscl) +
		nsec_to_cycles(apic->vcpu, deadline);

The first instance of this goes back to 03ba32cae66e ("VMX: x86: handle host TSC
calibration failure"), and every "fix" since then has been to avoid #DE.

> +
>  	/* too early */
>  	if (advance_expire_delta < 0) {
>  		ns = -advance_expire_delta * 1000000ULL;
> -- 
> 2.43.7
> 

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-10-09 20:35 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-08  9:00 [PATCH] KVM: x86: Guard against division by zero in adjust_lapic_timer_advance() Peng Hao
2026-10-08  9:15 ` sashiko-bot
2026-10-09 20:35 ` Sean Christopherson

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox