From: Sean Christopherson <seanjc@google.com>
To: David Woodhouse <dwmw2@infradead.org>
Cc: Paolo Bonzini <pbonzini@redhat.com>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Thomas Gleixner <tglx@kernel.org>,
Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
x86@kernel.org, "H. Peter Anvin" <hpa@zytor.com>,
Vitaly Kuznetsov <vkuznets@redhat.com>,
Juergen Gross <jgross@suse.com>,
Boris Ostrovsky <boris.ostrovsky@oracle.com>,
Paul Durrant <paul@xen.org>, Jonathan Cameron <jic23@kernel.org>,
Sascha Bischoff <Sascha.Bischoff@arm.com>,
Marc Zyngier <maz@kernel.org>, Joey Gouly <joey.gouly@arm.com>,
Jack Allister <jalliste@amazon.com>,
Dongli Zhang <dongli.zhang@oracle.com>,
joe.jin@oracle.com, kvm@vger.kernel.org,
linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
xen-devel@lists.xenproject.org, linux-kselftest@vger.kernel.org
Subject: Re: [PATCH v6 05/36] KVM: x86: Add KVM_[GS]ET_CLOCK_GUEST for accurate KVM clock migration
Date: Fri, 24 Jul 2026 14:17:19 -0700 [thread overview]
Message-ID: <amPWX8lKXw2KS_kZ@google.com> (raw)
In-Reply-To: <20260703212145.343527-6-dwmw2@infradead.org>
On Fri, Jul 03, 2026, David Woodhouse wrote:
> From: Jack Allister <jalliste@amazon.com>
>
> In the common case (where kvm->arch.use_master_clock is true), the KVM
> clock is defined as a simple arithmetic function of the guest TSC, based
> on a reference point stored in kvm->arch.master_kernel_ns and
> kvm->arch.master_cycle_now.
>
> The existing KVM_[GS]ET_CLOCK functionality does not allow for this
> relationship to be precisely saved and restored by userspace. All it can
> currently do is set the KVM clock at a given UTC reference time, which
> is necessarily imprecise.
>
> So on live update, the guest TSC can remain cycle accurate at precisely
> the same offset from the host TSC, but there is no way for userspace to
> restore the KVM clock accurately.
>
> Even on live migration to a new host, where the accuracy of the guest
> time-keeping is fundamentally limited by the accuracy of wallclock
> synchronization between the source and destination hosts, the clock jump
> experienced by the guest's TSC and its KVM clock should at least be
> *consistent*. Even when the guest TSC suffers a discontinuity, its KVM
> clock should still remain the *same* arithmetic function of the guest
> TSC, and not suffer an *additional* discontinuity.
>
> To allow for accurate migration of the KVM clock, add per-vCPU ioctls
> which save and restore the actual PV clock info in
> pvclock_vcpu_time_info.
>
> The restoration in KVM_SET_CLOCK_GUEST works by creating a new reference
> point in time just as kvm_update_masterclock() does, and calculating the
> corresponding guest TSC value. This guest TSC value is then passed
> through the user-provided pvclock structure to generate the *intended*
> KVM clock value at that point in time, and through the *actual* KVM
> clock calculation. Then kvm->arch.kvmclock_offset is adjusted to
> eliminate the difference.
>
> Where kvm->arch.use_master_clock is false (because the host TSC is
> unreliable, or the guest TSCs are configured strangely), the KVM clock
> is *not* defined as a function of the guest TSC so KVM_GET_CLOCK_GUEST
> returns an error. In this case, as documented, userspace shall use the
> legacy KVM_GET_CLOCK ioctl. The loss of precision is acceptable in this
> case since the clocks are imprecise in this mode anyway.
>
> On *restoration*, if kvm->arch.use_master_clock is false, an error is
> returned for similar reasons and userspace shall fall back to using
> KVM_SET_CLOCK. This does mean that, as documented, userspace needs to
> use *both* KVM_GET_CLOCK_GUEST and KVM_GET_CLOCK and send both results
> with the migration data (unless the intent is to refuse to resume on a
> host with bad TSC).
Please post this as a standalone mini-series. AFAICT, the only dependency of
any kind is a minor conflict with the s/hw_tsc_khz/hw_tsc_hz change, and that's
trivial to sort out later on.
I'm comfortable stumbling my way through the clock fixes, but I want Paolo (and
others) eyeballs on new uAPI like this. And because this series is plenty big
without this one :-)
> Co-developed-by: David Woodhouse <dwmw@amazon.co.uk>
> Signed-off-by: David Woodhouse <dwmw@amazon.co.uk>
> Signed-off-by: Jack Allister <jalliste@amazon.com>
> Reviewed-by: Paul Durrant <paul@xen.org>
> Cc: Dongli Zhang <dongli.zhang@oracle.com>
> Tested-by: Dongli Zhang <dongli.zhang@oracle.com>
> ---
> Documentation/virt/kvm/api.rst | 37 +++++++
> arch/x86/include/uapi/asm/kvm.h | 1 +
> arch/x86/kvm/x86.c | 171 ++++++++++++++++++++++++++++++++
> include/uapi/linux/kvm.h | 3 +
> 4 files changed, 212 insertions(+)
>
> diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
> index 52bbbb553ce1..2268b4442df6 100644
> --- a/Documentation/virt/kvm/api.rst
> +++ b/Documentation/virt/kvm/api.rst
> @@ -6553,6 +6553,43 @@ KVM_S390_KEYOP_SSKE
> Sets the storage key for the guest address ``guest_addr`` to the key
> specified in ``key``, returning the previous value in ``key``.
>
> +4.145 KVM_GET_CLOCK_GUEST
> +----------------------------
> +
> +:Capability: none
Why not add a CAP? The check in the subsequent selftest is quite gross.
next prev parent reply other threads:[~2026-07-24 21:17 UTC|newest]
Thread overview: 42+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-03 21:17 [PATCH v6 00/36] Cleaning up the KVM clock mess David Woodhouse
2026-07-03 21:17 ` [PATCH v6 01/36] KVM: x86/xen: Do not corrupt KVM clock in kvm_xen_shared_info_init() David Woodhouse
2026-07-03 21:17 ` [PATCH v6 02/36] KVM: x86: Improve accuracy of KVM clock when TSC scaling is in force David Woodhouse
2026-07-24 21:10 ` Sean Christopherson
2026-07-03 21:17 ` [PATCH v6 03/36] UAPI: x86: Move pvclock-abi to UAPI for x86 platforms David Woodhouse
2026-07-03 21:17 ` [PATCH v6 04/36] KVM: selftests: Use UAPI pvclock-abi.h in xen_shinfo_test David Woodhouse
2026-07-03 21:17 ` [PATCH v6 05/36] KVM: x86: Add KVM_[GS]ET_CLOCK_GUEST for accurate KVM clock migration David Woodhouse
2026-07-24 21:17 ` Sean Christopherson [this message]
2026-07-03 21:17 ` [PATCH v6 06/36] KVM: selftests: Add KVM/PV clock selftest to prove timer correction David Woodhouse
2026-07-03 21:17 ` [PATCH v6 07/36] KVM: x86: Explicitly disable TSC scaling without CONSTANT_TSC David Woodhouse
2026-07-03 21:17 ` [PATCH v6 08/36] KVM: x86: Activate master clock immediately on vCPU creation David Woodhouse
2026-07-24 21:19 ` Sean Christopherson
2026-07-03 21:17 ` [PATCH v6 09/36] KVM: x86: Add KVM_VCPU_TSC_SCALE and fix the documentation on TSC migration David Woodhouse
2026-07-03 21:17 ` [PATCH v6 10/36] KVM: x86: Avoid NTP frequency skew for KVM clock on 32-bit host David Woodhouse
2026-07-03 21:17 ` [PATCH v6 11/36] KVM: x86: Fold __get_kvmclock() into get_kvmclock() David Woodhouse
2026-07-03 21:17 ` [PATCH v6 12/36] KVM: x86: Restructure get_kvmclock() David Woodhouse
2026-07-24 21:29 ` Sean Christopherson
2026-07-03 21:17 ` [PATCH v6 13/36] KVM: x86: Fix KVM clock precision in get_kvmclock() with TSC scaling David Woodhouse
2026-07-03 21:17 ` [PATCH v6 14/36] KVM: x86: Use get_kvmclock() in kvm_get_wall_clock_epoch() David Woodhouse
2026-07-03 21:17 ` [PATCH v6 15/36] KVM: x86: Fix compute_guest_tsc() to handle negative time deltas David Woodhouse
2026-07-24 21:27 ` Sean Christopherson
2026-07-03 21:17 ` [PATCH v6 16/36] KVM: x86: Restructure kvm_guest_time_update() for TSC upscaling David Woodhouse
2026-07-03 21:17 ` [PATCH v6 17/36] KVM: x86: Simplify and comment kvm_get_time_scale() David Woodhouse
2026-07-03 21:17 ` [PATCH v6 18/36] KVM: x86: Remove implicit rdtsc() from kvm_compute_l1_tsc_offset() David Woodhouse
2026-07-03 21:17 ` [PATCH v6 19/36] KVM: x86: Improve synchronization in kvm_synchronize_tsc() David Woodhouse
2026-07-03 21:17 ` [PATCH v6 20/36] KVM: x86: Kill last_tsc_{nsec,write,offset} fields David Woodhouse
2026-07-03 21:18 ` [PATCH v6 21/36] KVM: x86: Replace nr_vcpus_matched_tsc count with all_vcpus_matched_tsc bool David Woodhouse
2026-07-03 21:18 ` [PATCH v6 22/36] KVM: x86: Allow KVM master clock mode when TSCs are offset from each other David Woodhouse
2026-07-03 21:18 ` [PATCH v6 23/36] KVM: selftests: Add master clock offset test David Woodhouse
2026-07-03 21:18 ` [PATCH v6 24/36] KVM: x86: Factor out kvm_use_master_clock() David Woodhouse
2026-07-03 21:18 ` [PATCH v6 25/36] KVM: x86: Avoid gratuitous global clock updates David Woodhouse
2026-07-03 21:18 ` [PATCH v6 26/36] KVM: x86/xen: Prevent runstate times from becoming negative David Woodhouse
2026-07-03 21:18 ` [PATCH v6 27/36] KVM: x86: Avoid redundant masterclock updates from multiple vCPUs David Woodhouse
2026-07-03 21:18 ` [PATCH v6 28/36] KVM: x86: Remove runtime Xen TSC frequency CPUID update David Woodhouse
2026-07-03 21:18 ` [PATCH v6 29/36] KVM: selftests: Add Xen/generic CPUID timing leaf test David Woodhouse
2026-07-03 21:18 ` [PATCH v6 30/36] KVM: x86: Re-synchronize TSC after KVM_SET_TSC_KHZ David Woodhouse
2026-07-03 21:18 ` [PATCH v6 31/36] KVM: selftests: Add Xen runstate migration test David Woodhouse
2026-07-03 21:18 ` [PATCH v6 32/36] KVM: x86: Use ktime_get_snapshot_id() for master clock David Woodhouse
2026-07-03 21:18 ` [PATCH v6 33/36] KVM: x86: Compute kvmclock base without pvclock_gtod_data David Woodhouse
2026-07-03 21:18 ` [PATCH v6 34/36] KVM: x86: Cache host vclock_mode for masterclock eligibility checks David Woodhouse
2026-07-03 21:18 ` [PATCH v6 35/36] KVM: x86: Remove pvclock_gtod_data and private timekeeping code David Woodhouse
2026-07-03 21:18 ` [PATCH v6 36/36] KVM: x86: Activate master clock from kvm_arch_init_vm() David Woodhouse
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=amPWX8lKXw2KS_kZ@google.com \
--to=seanjc@google.com \
--cc=Sascha.Bischoff@arm.com \
--cc=boris.ostrovsky@oracle.com \
--cc=bp@alien8.de \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=dongli.zhang@oracle.com \
--cc=dwmw2@infradead.org \
--cc=hpa@zytor.com \
--cc=jalliste@amazon.com \
--cc=jgross@suse.com \
--cc=jic23@kernel.org \
--cc=joe.jin@oracle.com \
--cc=joey.gouly@arm.com \
--cc=kvm@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=maz@kernel.org \
--cc=mingo@redhat.com \
--cc=paul@xen.org \
--cc=pbonzini@redhat.com \
--cc=skhan@linuxfoundation.org \
--cc=tglx@kernel.org \
--cc=vkuznets@redhat.com \
--cc=x86@kernel.org \
--cc=xen-devel@lists.xenproject.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.