* [PATCH 0/3] KVM: x86: Support hibernation of a nested Hyper-V
@ 2026-10-05 19:24 Mushahid Hussain
2026-10-05 19:24 ` [PATCH 1/3] KVM: nVMX: Clear stale vmcs02 sync flag in free_nested() Mushahid Hussain
` (3 more replies)
0 siblings, 4 replies; 7+ messages in thread
From: Mushahid Hussain @ 2026-10-05 19:24 UTC (permalink / raw)
To: kvm
Cc: Sean Christopherson, Paolo Bonzini, Vitaly Kuznetsov,
K . Y . Srinivasan, Haiyang Zhang, Wei Liu, Dexuan Cui, Long Li,
linux-hyperv, linux-kernel, nh-open-source, mushi.shar
A Windows guest with the Hyper-V role enabled runs its own hypervisor
(hvix64) nested under KVM. When KVM advertises the Hyper-V time and
frequency MSRs, hvix64 lets KVM own the partition reference time. It
then offers hibernation to Windows only if KVM also advertises CPUID
0x40000004 EAX bit 20 (RestoreTimeOnResume), and on resume restores
the saved TSC and reference time through HvCallRestorePartitionTime
(0x103). Both are reserved in the TLFS and defined in Microsoft's
hvdef crate.
Patch 1 fixes a stale nVMX sync flag that zeroed the vmcs12 segment
state hvix64 reloads on resume. It stands alone and is tagged for
stable. Patch 2 is a refactor with no functional change. Patch 3
advertises the bit and implements the hypercall, restoring kvmclock
and every vCPU's TSC inside one pvclock update.
Tested on Windows Server 2025 with Hyper-V running, guest initiated
hibernate and resume, under QEMU with a matching CPU property:
https://github.com/hmushi/qemu/tree/hv-restore-time-on-resume
Mushahid Hussain (3):
KVM: nVMX: Clear stale vmcs02 sync flag in free_nested()
KVM: x86: Extract __kvm_set_clock() from kvm_vm_ioctl_set_clock()
KVM: x86: hyper-v: Implement HvCallRestorePartitionTime
arch/x86/include/asm/kvm_host.h | 5 ++
arch/x86/kvm/hyperv.c | 36 +++++++++
arch/x86/kvm/vmx/nested.c | 1 +
arch/x86/kvm/x86.c | 131 +++++++++++++++++++++++++++-----
arch/x86/kvm/x86.h | 1 +
include/hyperv/hvgdk_mini.h | 12 +++
6 files changed, 168 insertions(+), 18 deletions(-)
base-commit: 6bd2905303c58679e303941b7d8ae8c074cd95ce
--
2.47.3
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH 1/3] KVM: nVMX: Clear stale vmcs02 sync flag in free_nested()
2026-10-05 19:24 [PATCH 0/3] KVM: x86: Support hibernation of a nested Hyper-V Mushahid Hussain
@ 2026-10-05 19:24 ` Mushahid Hussain
2026-10-05 19:24 ` [PATCH 2/3] KVM: x86: Extract __kvm_set_clock() from kvm_vm_ioctl_set_clock() Mushahid Hussain
` (2 subsequent siblings)
3 siblings, 0 replies; 7+ messages in thread
From: Mushahid Hussain @ 2026-10-05 19:24 UTC (permalink / raw)
To: kvm
Cc: Sean Christopherson, Paolo Bonzini, Vitaly Kuznetsov,
K . Y . Srinivasan, Haiyang Zhang, Wei Liu, Dexuan Cui, Long Li,
linux-hyperv, linux-kernel, nh-open-source, mushi.shar, stable
free_nested() frees vmcs02 but leaves need_sync_vmcs02_to_vmcs12_rare
set. The flag means that the rare guest fields of vmcs12 are valid
only in vmcs02. After vmcs02 is freed the flag is wrong.
L1 then executes VMXON and loads a vmcs12 with VMPTRLD. The flag is
still set, so the first sync copies the rare fields from the new,
never launched vmcs02 into vmcs12. This sets TR, LDTR, GDTR, IDTR,
the segment registers and the FS/GS bases to zero in guest memory.
set_current_vmptr() re-arms the other lazy flags at VMPTRLD. This
flag has no such point, so VMXOFF is the place to clear it.
A nested Hyper-V follows this sequence when it resumes from
hibernation. It executes VMXOFF before hibernation. On resume, it
reloads the VMCS images that winresume restored from the hiberfile.
VM-entry fails with exit reason 0x80000021 (invalid guest state) and
the guest hypervisor resets the machine. vmx_leave_nested() also goes
through free_nested(), so KVM_SET_NESTED_STATE takes the same path.
Clear the flag together with the other nested state.
Fixes: 7952d769c29c ("KVM: nVMX: Sync rarely accessed guest fields only when needed")
Cc: stable@vger.kernel.org
Assisted-by: Claude:claude-fable-5.1
Signed-off-by: Mushahid Hussain <hmushi@amazon.co.uk>
---
arch/x86/kvm/vmx/nested.c | 1 +
1 file changed, 1 insertion(+)
diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c
index 9b0bfa2f854cf..6c3723ddd40b8 100644
--- a/arch/x86/kvm/vmx/nested.c
+++ b/arch/x86/kvm/vmx/nested.c
@@ -350,6 +350,7 @@ static void free_nested(struct kvm_vcpu *vcpu)
vmx->nested.vmxon = false;
vmx->nested.smm.vmxon = false;
vmx->nested.vmxon_ptr = INVALID_GPA;
+ vmx->nested.need_sync_vmcs02_to_vmcs12_rare = false;
free_vpid(vmx->nested.vpid02);
vmx->nested.posted_intr_nv = -1;
vmx->nested.current_vmptr = INVALID_GPA;
--
2.47.3
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH 2/3] KVM: x86: Extract __kvm_set_clock() from kvm_vm_ioctl_set_clock()
2026-10-05 19:24 [PATCH 0/3] KVM: x86: Support hibernation of a nested Hyper-V Mushahid Hussain
2026-10-05 19:24 ` [PATCH 1/3] KVM: nVMX: Clear stale vmcs02 sync flag in free_nested() Mushahid Hussain
@ 2026-10-05 19:24 ` Mushahid Hussain
2026-10-05 19:24 ` [PATCH 3/3] KVM: x86: hyper-v: Implement HvCallRestorePartitionTime Mushahid Hussain
2026-10-06 5:08 ` [PATCH 0/3] KVM: x86: Support hibernation of a nested Hyper-V David Woodhouse
3 siblings, 0 replies; 7+ messages in thread
From: Mushahid Hussain @ 2026-10-05 19:24 UTC (permalink / raw)
To: kvm
Cc: Sean Christopherson, Paolo Bonzini, Vitaly Kuznetsov,
K . Y . Srinivasan, Haiyang Zhang, Wei Liu, Dexuan Cui, Long Li,
linux-hyperv, linux-kernel, nh-open-source, mushi.shar
Move the body of KVM_SET_CLOCK into __kvm_set_clock() and leave the
copy from userspace and the flags check in the ioctl handler. A
following patch steps kvmclock from inside KVM and needs the worker.
No functional change intended.
Assisted-by: Claude:claude-fable-5.1
Signed-off-by: Mushahid Hussain <hmushi@amazon.co.uk>
---
arch/x86/kvm/x86.c | 40 +++++++++++++++++++++++-----------------
1 file changed, 23 insertions(+), 17 deletions(-)
diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index 32563a91fba99..19666a80240a5 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -4354,22 +4354,11 @@ static int kvm_vm_ioctl_get_clock(struct kvm *kvm, void __user *argp)
return 0;
}
-static int kvm_vm_ioctl_set_clock(struct kvm *kvm, void __user *argp)
+static void __kvm_set_clock(struct kvm *kvm, struct kvm_clock_data *data)
{
struct kvm_arch *ka = &kvm->arch;
- struct kvm_clock_data data;
u64 now_raw_ns;
- if (copy_from_user(&data, argp, sizeof(data)))
- return -EFAULT;
-
- /*
- * Only KVM_CLOCK_REALTIME is used, but allow passing the
- * result of KVM_GET_CLOCK back to KVM_SET_CLOCK.
- */
- if (data.flags & ~KVM_CLOCK_VALID_FLAGS)
- return -EINVAL;
-
kvm_hv_request_tsc_page_update(kvm);
kvm_start_pvclock_update(kvm);
pvclock_update_vm_gtod_copy(kvm);
@@ -4379,24 +4368,41 @@ static int kvm_vm_ioctl_set_clock(struct kvm *kvm, void __user *argp)
* in use, we use master_kernel_ns + kvmclock_offset to set
* unsigned 'system_time' so if we use get_kvmclock_ns() (which
* is slightly ahead) here we risk going negative on unsigned
- * 'system_time' when 'data.clock' is very small.
+ * 'system_time' when 'data->clock' is very small.
*/
- if (data.flags & KVM_CLOCK_REALTIME) {
+ if (data->flags & KVM_CLOCK_REALTIME) {
u64 now_real_ns = ktime_get_real_ns();
/*
* Avoid stepping the kvmclock backwards.
*/
- if (now_real_ns > data.realtime)
- data.clock += now_real_ns - data.realtime;
+ if (now_real_ns > data->realtime)
+ data->clock += now_real_ns - data->realtime;
}
if (ka->use_master_clock)
now_raw_ns = ka->master_kernel_ns;
else
now_raw_ns = get_kvmclock_base_ns();
- ka->kvmclock_offset = data.clock - now_raw_ns;
+ ka->kvmclock_offset = data->clock - now_raw_ns;
kvm_end_pvclock_update(kvm);
+}
+
+static int kvm_vm_ioctl_set_clock(struct kvm *kvm, void __user *argp)
+{
+ struct kvm_clock_data data;
+
+ if (copy_from_user(&data, argp, sizeof(data)))
+ return -EFAULT;
+
+ /*
+ * Only KVM_CLOCK_REALTIME is used, but allow passing the
+ * result of KVM_GET_CLOCK back to KVM_SET_CLOCK.
+ */
+ if (data.flags & ~KVM_CLOCK_VALID_FLAGS)
+ return -EINVAL;
+
+ __kvm_set_clock(kvm, &data);
return 0;
}
--
2.47.3
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH 3/3] KVM: x86: hyper-v: Implement HvCallRestorePartitionTime
2026-10-05 19:24 [PATCH 0/3] KVM: x86: Support hibernation of a nested Hyper-V Mushahid Hussain
2026-10-05 19:24 ` [PATCH 1/3] KVM: nVMX: Clear stale vmcs02 sync flag in free_nested() Mushahid Hussain
2026-10-05 19:24 ` [PATCH 2/3] KVM: x86: Extract __kvm_set_clock() from kvm_vm_ioctl_set_clock() Mushahid Hussain
@ 2026-10-05 19:24 ` Mushahid Hussain
2026-10-05 19:38 ` sashiko-bot
2026-10-05 22:30 ` David Woodhouse
2026-10-06 5:08 ` [PATCH 0/3] KVM: x86: Support hibernation of a nested Hyper-V David Woodhouse
3 siblings, 2 replies; 7+ messages in thread
From: Mushahid Hussain @ 2026-10-05 19:24 UTC (permalink / raw)
To: kvm
Cc: Sean Christopherson, Paolo Bonzini, Vitaly Kuznetsov,
K . Y . Srinivasan, Haiyang Zhang, Wei Liu, Dexuan Cui, Long Li,
linux-hyperv, linux-kernel, nh-open-source, mushi.shar
A nested Hyper-V that sees the frequency MSRs lets its parent own
the partition reference time. After a resume from hibernation it
issues HvCallRestorePartitionTime (0x103) with the TSC and the
reference counter it saved, and expects both clocks to continue
from there. KVM rejected the hypercall. The guest time base stayed
inconsistent and the root partition did not run.
Restore both clocks in one pvclock update, so no vCPU runs with the
new TSC and the old reference time. Each vCPU computes and writes its
own TSC offset from the recorded step, so no other vCPU writes its
TSC state. kvm_synchronize_tsc() writes the offset of the calling
vCPU only. Mark the TSC page as changed by the guest, so it is
recomputed even when TSC emulation control is set.
The TLFS marks the hypercall, its input and CPUID 0x40000004 EAX
bit 20 (RestoreTimeOnResume) as reserved. Microsoft's hvdef crate
defines them (hypercall::RestorePartitionTime), and Microsoft's
OHCL kernel carries the same C definitions in hvgdk_mini.h.
Advertise bit 20 and gate the hypercall on it under
KVM_CAP_HYPERV_ENFORCE_CPUID. Hyper-V offers hibernation to its root
partition only when the parent sets this bit.
Assisted-by: Claude:claude-fable-5.1
Signed-off-by: Mushahid Hussain <hmushi@amazon.co.uk>
---
arch/x86/include/asm/kvm_host.h | 5 ++
arch/x86/kvm/hyperv.c | 36 ++++++++++++
arch/x86/kvm/x86.c | 99 +++++++++++++++++++++++++++++++--
arch/x86/kvm/x86.h | 1 +
include/hyperv/hvgdk_mini.h | 12 ++++
5 files changed, 148 insertions(+), 5 deletions(-)
diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h
index e2eef7057bd35..e97ff631196f8 100644
--- a/arch/x86/include/asm/kvm_host.h
+++ b/arch/x86/include/asm/kvm_host.h
@@ -127,6 +127,7 @@
KVM_ARCH_REQ_FLAGS(33, KVM_REQUEST_WAIT | KVM_REQUEST_NO_WAKEUP)
#define KVM_REQ_UPDATE_PROTECTED_GUEST_STATE \
KVM_ARCH_REQ_FLAGS(34, KVM_REQUEST_WAIT)
+#define KVM_REQ_WRITE_TSC_OFFSET KVM_ARCH_REQ(35)
#define INVALID_PAGE (~(hpa_t)0)
#define VALID_PAGE(x) ((x) != INVALID_PAGE)
@@ -1246,6 +1247,10 @@ struct kvm_arch {
u64 cur_tsc_offset;
u64 cur_tsc_generation;
int nr_vcpus_matched_tsc;
+ /* TSC step requested by the guest; each vCPU applies it. */
+ u64 restore_host_tsc;
+ u64 restore_guest_tsc;
+ u64 restore_tsc_nsec;
u32 default_tsc_khz;
bool user_set_tsc;
diff --git a/arch/x86/kvm/hyperv.c b/arch/x86/kvm/hyperv.c
index 5d2e3617c9520..923373dee0e4f 100644
--- a/arch/x86/kvm/hyperv.c
+++ b/arch/x86/kvm/hyperv.c
@@ -2538,6 +2538,9 @@ static bool hv_check_hypercall_access(struct kvm_vcpu_hv *hv_vcpu, u16 code)
case HVCALL_SEND_IPI:
return hv_vcpu->cpuid_cache.enlightenments_eax &
HV_X64_CLUSTER_IPI_RECOMMENDED;
+ case HVCALL_RESTORE_PARTITION_TIME:
+ return hv_vcpu->cpuid_cache.enlightenments_eax &
+ HV_X64_RESTORE_TIME_ON_RESUME;
case HV_EXT_CALL_QUERY_CAPABILITIES ... HV_EXT_CALL_MAX:
return hv_vcpu->cpuid_cache.features_ebx &
HV_ENABLE_EXTENDED_HYPERCALLS;
@@ -2548,6 +2551,31 @@ static bool hv_check_hypercall_access(struct kvm_vcpu_hv *hv_vcpu, u16 code)
return true;
}
+static u64 kvm_hv_restore_partition_time(struct kvm_vcpu *vcpu,
+ struct kvm_hv_hcall *hc)
+{
+ struct kvm_hv *hv = to_kvm_hv(vcpu->kvm);
+ struct hv_input_restore_partition_time in;
+
+ BUILD_BUG_ON(sizeof(in) != 32);
+
+ if (kvm_vcpu_read_guest(vcpu, hc->ingpa, &in, sizeof(in)))
+ return HV_STATUS_INVALID_HYPERCALL_INPUT;
+ if (in.reserved)
+ return HV_STATUS_INVALID_HYPERCALL_INPUT;
+ if (in.partition_id != HV_PARTITION_ID_SELF)
+ return HV_STATUS_INVALID_PARTITION_ID;
+
+ /* The guest asked for the step; recompute the TSC page as for its own write. */
+ mutex_lock(&hv->hv_lock);
+ if (hv->hv_tsc_page & HV_X64_MSR_TSC_REFERENCE_ENABLE)
+ hv->hv_tsc_page_status = HV_TSC_PAGE_GUEST_CHANGED;
+ mutex_unlock(&hv->hv_lock);
+
+ kvm_set_clock_and_tsc(vcpu->kvm, in.reference_time_in_100ns * 100, in.tsc);
+ return HV_STATUS_SUCCESS;
+}
+
int kvm_hv_hypercall(struct kvm_vcpu *vcpu)
{
struct kvm_vcpu_hv *hv_vcpu = to_hv_vcpu(vcpu);
@@ -2670,6 +2698,13 @@ int kvm_hv_hypercall(struct kvm_vcpu *vcpu)
}
ret = kvm_hv_send_ipi(vcpu, &hc);
break;
+ case HVCALL_RESTORE_PARTITION_TIME:
+ if (unlikely(hc.fast || hc.rep || hc.var_cnt)) {
+ ret = HV_STATUS_INVALID_HYPERCALL_INPUT;
+ break;
+ }
+ ret = kvm_hv_restore_partition_time(vcpu, &hc);
+ break;
case HVCALL_POST_DEBUG_DATA:
case HVCALL_RETRIEVE_DEBUG_DATA:
if (unlikely(hc.fast)) {
@@ -2889,6 +2924,7 @@ int kvm_get_hv_cpuid(struct kvm_vcpu *vcpu, struct kvm_cpuid2 *cpuid,
ent->eax |= HV_X64_NO_NONARCH_CORESHARING;
ent->eax |= HV_DEPRECATING_AEOI_RECOMMENDED;
+ ent->eax |= HV_X64_RESTORE_TIME_ON_RESUME;
/*
* Default number of spinlock retry attempts, matches
* HyperV 2016.
diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index 19666a80240a5..c0865468b02e7 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -4354,13 +4354,85 @@ static int kvm_vm_ioctl_get_clock(struct kvm *kvm, void __user *argp)
return 0;
}
-static void __kvm_set_clock(struct kvm *kvm, struct kvm_clock_data *data)
+/* Must precede pvclock_update_vm_gtod_copy(), which reads the matched count. */
+static void kvm_open_tsc_generation(struct kvm *kvm, u64 guest_tsc)
{
struct kvm_arch *ka = &kvm->arch;
- u64 now_raw_ns;
+ struct kvm_vcpu *vcpu;
+ unsigned long i;
+
+ lockdep_assert_held(&ka->tsc_write_lock);
+
+ ka->cur_tsc_generation++;
+ ka->cur_tsc_write = guest_tsc;
+ ka->last_tsc_write = guest_tsc;
+ ka->nr_vcpus_matched_tsc = atomic_read(&kvm->online_vcpus) - 1;
+
+ kvm_for_each_vcpu(i, vcpu, kvm) {
+ if (vcpu->arch.guest_tsc_protected)
+ continue;
+ ka->last_tsc_khz = vcpu->arch.virtual_tsc_khz;
+ ka->last_tsc_scaling_ratio = vcpu->arch.l1_tsc_scaling_ratio;
+ break;
+ }
+}
+
+/* Every vCPU reads @guest_tsc at host TSC @host_tsc. */
+static void kvm_set_tsc_generation(struct kvm *kvm, u64 host_tsc,
+ u64 guest_tsc, u64 ns)
+{
+ struct kvm_arch *ka = &kvm->arch;
+ struct kvm_vcpu *vcpu;
+ unsigned long i;
+
+ lockdep_assert_held(&ka->tsc_write_lock);
+
+ ka->cur_tsc_nsec = ns;
+ ka->last_tsc_nsec = ns;
+ ka->restore_host_tsc = host_tsc;
+ ka->restore_guest_tsc = guest_tsc;
+ ka->restore_tsc_nsec = ns;
+
+ kvm_for_each_vcpu(i, vcpu, kvm) {
+ if (vcpu->arch.guest_tsc_protected)
+ continue;
+
+ ka->cur_tsc_offset = kvm_compute_l1_tsc_offset(vcpu, host_tsc,
+ guest_tsc);
+ ka->last_tsc_offset = ka->cur_tsc_offset;
+ kvm_make_request(KVM_REQ_WRITE_TSC_OFFSET, vcpu);
+ }
+}
+
+/* Runs on the vCPU; only the owner writes its TSC offset. */
+static void kvm_vcpu_apply_tsc_generation(struct kvm_vcpu *vcpu)
+{
+ struct kvm_arch *ka = &vcpu->kvm->arch;
+ unsigned long flags;
+ u64 offset;
+
+ raw_spin_lock_irqsave(&ka->tsc_write_lock, flags);
+ offset = kvm_compute_l1_tsc_offset(vcpu, ka->restore_host_tsc,
+ ka->restore_guest_tsc);
+ vcpu->arch.last_guest_tsc = ka->restore_guest_tsc;
+ vcpu->arch.this_tsc_generation = ka->cur_tsc_generation;
+ vcpu->arch.this_tsc_nsec = ka->restore_tsc_nsec;
+ vcpu->arch.this_tsc_write = ka->restore_guest_tsc;
+ raw_spin_unlock_irqrestore(&ka->tsc_write_lock, flags);
+
+ kvm_vcpu_write_tsc_offset(vcpu, offset);
+}
+
+static void __kvm_set_clock(struct kvm *kvm, struct kvm_clock_data *data,
+ const u64 *guest_tsc)
+{
+ struct kvm_arch *ka = &kvm->arch;
+ u64 host_tsc, now_raw_ns;
kvm_hv_request_tsc_page_update(kvm);
kvm_start_pvclock_update(kvm);
+ if (guest_tsc)
+ kvm_open_tsc_generation(kvm, *guest_tsc);
pvclock_update_vm_gtod_copy(kvm);
/*
@@ -4380,14 +4452,29 @@ static void __kvm_set_clock(struct kvm *kvm, struct kvm_clock_data *data)
data->clock += now_real_ns - data->realtime;
}
- if (ka->use_master_clock)
+ if (ka->use_master_clock) {
now_raw_ns = ka->master_kernel_ns;
- else
+ host_tsc = ka->master_cycle_now;
+ } else {
+ host_tsc = rdtsc();
now_raw_ns = get_kvmclock_base_ns();
+ }
ka->kvmclock_offset = data->clock - now_raw_ns;
+
+ if (guest_tsc)
+ kvm_set_tsc_generation(kvm, host_tsc, *guest_tsc, now_raw_ns);
+
kvm_end_pvclock_update(kvm);
}
+/* Step kvmclock to @ns and the guest TSC on every vCPU to @guest_tsc. */
+void kvm_set_clock_and_tsc(struct kvm *kvm, u64 ns, u64 guest_tsc)
+{
+ struct kvm_clock_data data = { .clock = ns };
+
+ __kvm_set_clock(kvm, &data, &guest_tsc);
+}
+
static int kvm_vm_ioctl_set_clock(struct kvm *kvm, void __user *argp)
{
struct kvm_clock_data data;
@@ -4402,7 +4489,7 @@ static int kvm_vm_ioctl_set_clock(struct kvm *kvm, void __user *argp)
if (data.flags & ~KVM_CLOCK_VALID_FLAGS)
return -EINVAL;
- __kvm_set_clock(kvm, &data);
+ __kvm_set_clock(kvm, &data, NULL);
return 0;
}
@@ -8168,6 +8255,8 @@ static int vcpu_enter_guest(struct kvm_vcpu *vcpu)
kvm_update_masterclock(vcpu->kvm);
if (kvm_check_request(KVM_REQ_GLOBAL_CLOCK_UPDATE, vcpu))
kvm_gen_kvmclock_update(vcpu);
+ if (kvm_check_request(KVM_REQ_WRITE_TSC_OFFSET, vcpu))
+ kvm_vcpu_apply_tsc_generation(vcpu);
if (kvm_check_request(KVM_REQ_CLOCK_UPDATE, vcpu)) {
r = kvm_guest_time_update(vcpu);
if (unlikely(r))
diff --git a/arch/x86/kvm/x86.h b/arch/x86/kvm/x86.h
index 436e2fb1396c4..e26ca67ac7219 100644
--- a/arch/x86/kvm/x86.h
+++ b/arch/x86/kvm/x86.h
@@ -326,6 +326,7 @@ void kvm_vcpu_reset(struct kvm_vcpu *vcpu, bool init_event);
void kvm_inject_realmode_interrupt(struct kvm_vcpu *vcpu, int irq, int inc_eip);
u64 get_kvmclock_ns(struct kvm *kvm);
+void kvm_set_clock_and_tsc(struct kvm *kvm, u64 ns, u64 guest_tsc);
uint64_t kvm_get_wall_clock_epoch(struct kvm *kvm);
bool kvm_get_monotonic_and_clockread(s64 *kernel_ns, u64 *tsc_timestamp);
int kvm_guest_time_update(struct kvm_vcpu *v);
diff --git a/include/hyperv/hvgdk_mini.h b/include/hyperv/hvgdk_mini.h
index 6a4e8b9d570fd..87dfb0aed827e 100644
--- a/include/hyperv/hvgdk_mini.h
+++ b/include/hyperv/hvgdk_mini.h
@@ -343,6 +343,8 @@ union hv_hypervisor_version_info {
#define HV_X64_EX_PROCESSOR_MASKS_RECOMMENDED BIT(11)
#define HV_X64_HYPERV_NESTED BIT(12)
#define HV_X64_ENLIGHTENED_VMCS_RECOMMENDED BIT(14)
+/* Reserved in the TLFS; from Microsoft's hvdef. */
+#define HV_X64_RESTORE_TIME_ON_RESUME BIT(20)
#define HV_X64_USE_MMIO_HYPERCALLS BIT(21)
/*
@@ -499,6 +501,7 @@ union hv_vp_assist_msr_contents { /* HV_REGISTER_VP_ASSIST_PAGE */
#define HVCALL_SET_VP_STATE 0x00e4
#define HVCALL_GET_VP_CPUID_VALUES 0x00f4
#define HVCALL_GET_PARTITION_PROPERTY_EX 0x0101
+#define HVCALL_RESTORE_PARTITION_TIME 0x0103
#define HVCALL_MMIO_READ 0x0106
#define HVCALL_MMIO_WRITE 0x0107
#define HVCALL_DISABLE_HYP_EX 0x010f
@@ -1514,6 +1517,15 @@ union hv_intercept_parameters {
/* N.B. Other intercept types do not have any parameters. */
};
+/* HVCALL_RESTORE_PARTITION_TIME input. Layout from Microsoft's hvdef. */
+struct hv_input_restore_partition_time {
+ u64 partition_id;
+ u32 tsc_sequence;
+ u32 reserved;
+ u64 reference_time_in_100ns;
+ u64 tsc;
+} __packed;
+
/* Data structures for HVCALL_MMIO_READ and HVCALL_MMIO_WRITE */
#define HV_HYPERCALL_MMIO_MAX_DATA_LENGTH 64
--
2.47.3
^ permalink raw reply related [flat|nested] 7+ messages in thread
* Re: [PATCH 3/3] KVM: x86: hyper-v: Implement HvCallRestorePartitionTime
2026-10-05 19:24 ` [PATCH 3/3] KVM: x86: hyper-v: Implement HvCallRestorePartitionTime Mushahid Hussain
@ 2026-10-05 19:38 ` sashiko-bot
2026-10-05 22:30 ` David Woodhouse
1 sibling, 0 replies; 7+ messages in thread
From: sashiko-bot @ 2026-10-05 19:38 UTC (permalink / raw)
To: Mushahid Hussain; +Cc: kvm, linux-hyperv
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] x86/kvm: Race condition in kvm_vcpu_apply_tsc_generation() causes state corruption
- [High] x86/kvm: Pre-mature matched TSC count increment permits double-counting
- [Medium] x86/kvm: VM-wide cur_tsc_offset overwritten in per-vCPU loop
--
commit 3619ce01761243ff0d6db55e73b6eefb9f3cc10b
Author: Mushahid Hussain <hmushi@amazon.co.uk>
Subject: KVM: x86: hyper-v: Implement HvCallRestorePartitionTime
A nested Hyper-V guest can resume from hibernation and use
HvCallRestorePartitionTime to sync its TSC and reference counter. This commit
implements handling for the hypercall in KVM, stepping both clocks in a
single update. It updates kvm_set_clock_and_tsc() and issues requests to
write the TSC offset for each vCPU.
> diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
> index 19666a80240a5..c0865468b02e7 100644
> --- a/arch/x86/kvm/x86.c
> +++ b/arch/x86/kvm/x86.c
[ ... ]
> @@ -4354,13 +4354,85 @@ static int kvm_vm_ioctl_get_clock(struct kvm *kvm, void __user *argp)
> return 0;
> }
>
> -static void __kvm_set_clock(struct kvm *kvm, struct kvm_clock_data *data)
> +/* Must precede pvclock_update_vm_gtod_copy(), which reads the matched count. */
> +static void kvm_open_tsc_generation(struct kvm *kvm, u64 guest_tsc)
> +{
> + struct kvm_arch *ka = &kvm->arch;
> + struct kvm_vcpu *vcpu;
> + unsigned long i;
> +
> + lockdep_assert_held(&ka->tsc_write_lock);
> +
> + ka->cur_tsc_generation++;
> + ka->cur_tsc_write = guest_tsc;
> + ka->last_tsc_write = guest_tsc;
> + ka->nr_vcpus_matched_tsc = atomic_read(&kvm->online_vcpus) - 1;
[Severity: High]
Is it safe to pre-increment the matched count globally before the per-vCPU
updates complete?
If a vCPU executes kvm_synchronize_tsc() before processing the asynchronous
KVM_REQ_WRITE_TSC_OFFSET, it might see a stale generation, match it, and
increment ka->nr_vcpus_matched_tsc.
Since the count is pre-incremented here, double-counting could cause
the total to reach or exceed online_vcpus, permanently breaking the KVM
masterclock validation (which checks if matched_tsc + 1 == online_vcpus).
> +
> + kvm_for_each_vcpu(i, vcpu, kvm) {
> + if (vcpu->arch.guest_tsc_protected)
> + continue;
> + ka->last_tsc_khz = vcpu->arch.virtual_tsc_khz;
> + ka->last_tsc_scaling_ratio = vcpu->arch.l1_tsc_scaling_ratio;
> + break;
> + }
> +}
> +
> +/* Every vCPU reads @guest_tsc at host TSC @host_tsc. */
> +static void kvm_set_tsc_generation(struct kvm *kvm, u64 host_tsc,
> + u64 guest_tsc, u64 ns)
> +{
> + struct kvm_arch *ka = &kvm->arch;
> + struct kvm_vcpu *vcpu;
> + unsigned long i;
> +
> + lockdep_assert_held(&ka->tsc_write_lock);
> +
> + ka->cur_tsc_nsec = ns;
> + ka->last_tsc_nsec = ns;
> + ka->restore_host_tsc = host_tsc;
> + ka->restore_guest_tsc = guest_tsc;
> + ka->restore_tsc_nsec = ns;
> +
> + kvm_for_each_vcpu(i, vcpu, kvm) {
> + if (vcpu->arch.guest_tsc_protected)
> + continue;
> +
> + ka->cur_tsc_offset = kvm_compute_l1_tsc_offset(vcpu, host_tsc,
> + guest_tsc);
> + ka->last_tsc_offset = ka->cur_tsc_offset;
[Severity: Medium]
Could this loop repeatedly overwrite the VM-wide offset, leaving it configured
for the last vCPU?
If vCPUs have different TSC scaling ratios (which can be set via
KVM_SET_TSC_KHZ), the reference frequency is bound to the first valid vCPU in
kvm_open_tsc_generation() above, while the offset here ends up bound to the
last vCPU, breaking consistency.
> + kvm_make_request(KVM_REQ_WRITE_TSC_OFFSET, vcpu);
> + }
> +}
> +
> +/* Runs on the vCPU; only the owner writes its TSC offset. */
> +static void kvm_vcpu_apply_tsc_generation(struct kvm_vcpu *vcpu)
> +{
> + struct kvm_arch *ka = &vcpu->kvm->arch;
> + unsigned long flags;
> + u64 offset;
> +
> + raw_spin_lock_irqsave(&ka->tsc_write_lock, flags);
> + offset = kvm_compute_l1_tsc_offset(vcpu, ka->restore_host_tsc,
> + ka->restore_guest_tsc);
> + vcpu->arch.last_guest_tsc = ka->restore_guest_tsc;
> + vcpu->arch.this_tsc_generation = ka->cur_tsc_generation;
[Severity: High]
Could a race condition corrupt state tracking here?
If a concurrent write to MSR_IA32_TSC bumps ka->cur_tsc_generation while
KVM_REQ_WRITE_TSC_OFFSET is queued, the vCPU processing the request will copy
the newer generation counter here while applying the stale ka->restore_*
offset data.
This would leave the vCPU claiming synchronization with the new generation
while actually holding old data, breaking KVM masterclock and TSC
synchronization.
> + vcpu->arch.this_tsc_nsec = ka->restore_tsc_nsec;
> + vcpu->arch.this_tsc_write = ka->restore_guest_tsc;
> + raw_spin_unlock_irqrestore(&ka->tsc_write_lock, flags);
> +
> + kvm_vcpu_write_tsc_offset(vcpu, offset);
> +}
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261005192431.87317-1-hmushi@amazon.co.uk?part=3
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 3/3] KVM: x86: hyper-v: Implement HvCallRestorePartitionTime
2026-10-05 19:24 ` [PATCH 3/3] KVM: x86: hyper-v: Implement HvCallRestorePartitionTime Mushahid Hussain
2026-10-05 19:38 ` sashiko-bot
@ 2026-10-05 22:30 ` David Woodhouse
1 sibling, 0 replies; 7+ messages in thread
From: David Woodhouse @ 2026-10-05 22:30 UTC (permalink / raw)
To: Mushahid Hussain, kvm
Cc: Sean Christopherson, Paolo Bonzini, Vitaly Kuznetsov,
K . Y . Srinivasan, Haiyang Zhang, Wei Liu, Dexuan Cui, Long Li,
linux-hyperv, linux-kernel, nh-open-source, mushi.shar
[-- Attachment #1: Type: text/plain, Size: 1618 bytes --]
On Mon, 2026-10-05 at 19:24 +0000, Mushahid Hussain wrote:
>
> -static void __kvm_set_clock(struct kvm *kvm, struct kvm_clock_data
> *data)
> +/* Must precede pvclock_update_vm_gtod_copy(), which reads the
> matched count. */
> +static void kvm_open_tsc_generation(struct kvm *kvm, u64 guest_tsc)
> {
> struct kvm_arch *ka = &kvm->arch;
> - u64 now_raw_ns;
> + struct kvm_vcpu *vcpu;
> + unsigned long i;
> +
> + lockdep_assert_held(&ka->tsc_write_lock);
> +
> + ka->cur_tsc_generation++;
> + ka->cur_tsc_write = guest_tsc;
> + ka->last_tsc_write = guest_tsc;
> + ka->nr_vcpus_matched_tsc = atomic_read(&kvm->online_vcpus) -
> 1;
> +
> + kvm_for_each_vcpu(i, vcpu, kvm) {
> + if (vcpu->arch.guest_tsc_protected)
> + continue;
> + ka->last_tsc_khz = vcpu->arch.virtual_tsc_khz;
> + ka->last_tsc_scaling_ratio = vcpu-
> >arch.l1_tsc_scaling_ratio;
> + break;
> + }
> +}
> +
> +/* Every vCPU reads @guest_tsc at host TSC @host_tsc. */
> +static void kvm_set_tsc_generation(struct kvm *kvm, u64 host_tsc,
> + u64 guest_tsc, u64 ns)
This series doesn't add a KVM selftest. Please do.
And *if* you're going to do the rework I'm citing here, please do it in
a preliminary patch, called out separately from the functional
addition.
And... *don't* do the rework I'm citing here. Or at least if you must,
please do it on top of
https://git.infradead.org/?p=users/dwmw2/linux.git;a=shortlog;h=refs/heads/kvmclock10-part2
But actually, I have a strong suspicion you just want to set the TSC
and then use a version of KVM_SET_CLOCK_GUEST, *not* KVM_SET_CLOCK?
[-- Attachment #2: smime.p7s --]
[-- Type: application/pkcs7-signature, Size: 6179 bytes --]
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 0/3] KVM: x86: Support hibernation of a nested Hyper-V
2026-10-05 19:24 [PATCH 0/3] KVM: x86: Support hibernation of a nested Hyper-V Mushahid Hussain
` (2 preceding siblings ...)
2026-10-05 19:24 ` [PATCH 3/3] KVM: x86: hyper-v: Implement HvCallRestorePartitionTime Mushahid Hussain
@ 2026-10-06 5:08 ` David Woodhouse
3 siblings, 0 replies; 7+ messages in thread
From: David Woodhouse @ 2026-10-06 5:08 UTC (permalink / raw)
To: Mushahid Hussain, kvm
Cc: Sean Christopherson, Paolo Bonzini, Vitaly Kuznetsov,
K . Y . Srinivasan, Haiyang Zhang, Wei Liu, Dexuan Cui, Long Li,
linux-hyperv, linux-kernel, nh-open-source, mushi.shar
[-- Attachment #1: Type: text/plain, Size: 2299 bytes --]
On Mon, 2026-10-05 at 19:24 +0000, Mushahid Hussain wrote:
> A Windows guest with the Hyper-V role enabled runs its own hypervisor
> (hvix64) nested under KVM. When KVM advertises the Hyper-V time and
> frequency MSRs, hvix64 lets KVM own the partition reference time. It
> then offers hibernation to Windows only if KVM also advertises CPUID
> 0x40000004 EAX bit 20 (RestoreTimeOnResume), and on resume restores
> the saved TSC and reference time through HvCallRestorePartitionTime
> (0x103). Both are reserved in the TLFS and defined in Microsoft's
> hvdef crate.
>
> Patch 1 fixes a stale nVMX sync flag that zeroed the vmcs12 segment
> state hvix64 reloads on resume. It stands alone and is tagged for
> stable. Patch 2 is a refactor with no functional change. Patch 3
> advertises the bit and implements the hypercall, restoring kvmclock
> and every vCPU's TSC inside one pvclock update.
>
> Tested on Windows Server 2025 with Hyper-V running, guest initiated
> hibernate and resume, under QEMU with a matching CPU property:
> https://github.com/hmushi/qemu/tree/hv-restore-time-on-resume
>
> Mushahid Hussain (3):
> KVM: nVMX: Clear stale vmcs02 sync flag in free_nested()
> KVM: x86: Extract __kvm_set_clock() from kvm_vm_ioctl_set_clock()
> KVM: x86: hyper-v: Implement HvCallRestorePartitionTime
>
> arch/x86/include/asm/kvm_host.h | 5 ++
> arch/x86/kvm/hyperv.c | 36 +++++++++
> arch/x86/kvm/vmx/nested.c | 1 +
> arch/x86/kvm/x86.c | 131 +++++++++++++++++++++++++++-----
> arch/x86/kvm/x86.h | 1 +
> include/hyperv/hvgdk_mini.h | 12 +++
> 6 files changed, 168 insertions(+), 18 deletions(-)
>
>
> base-commit: 6bd2905303c58679e303941b7d8ae8c074cd95ce
Please test
https://git.infradead.org/?p=users/dwmw2/linux.git;a=shortlog;h=refs/heads/kvmclock10-hmushi
I kind of wish we didn't have to do this in the kernel at all, and
could punt to userspace. But we don't have that plumbing for Hv
hypercalls, do we? So we'd have to invent that, and then also a way for
userspace to set the sequence in the TSC page (which I wired up to be
restored as I think that's probably what should happen; can we
confirm?).
[-- Attachment #2: smime.p7s --]
[-- Type: application/pkcs7-signature, Size: 6179 bytes --]
^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-10-06 5:08 UTC | newest]
Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-05 19:24 [PATCH 0/3] KVM: x86: Support hibernation of a nested Hyper-V Mushahid Hussain
2026-10-05 19:24 ` [PATCH 1/3] KVM: nVMX: Clear stale vmcs02 sync flag in free_nested() Mushahid Hussain
2026-10-05 19:24 ` [PATCH 2/3] KVM: x86: Extract __kvm_set_clock() from kvm_vm_ioctl_set_clock() Mushahid Hussain
2026-10-05 19:24 ` [PATCH 3/3] KVM: x86: hyper-v: Implement HvCallRestorePartitionTime Mushahid Hussain
2026-10-05 19:38 ` sashiko-bot
2026-10-05 22:30 ` David Woodhouse
2026-10-06 5:08 ` [PATCH 0/3] KVM: x86: Support hibernation of a nested Hyper-V David Woodhouse
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox