The Linux Kernel Mailing List
 help / color / mirror / Atom feed
From: Sean Christopherson <seanjc@google.com>
To: David Woodhouse <dwmw2@infradead.org>
Cc: Paolo Bonzini <pbonzini@redhat.com>,
	Jonathan Corbet <corbet@lwn.net>,
	 Shuah Khan <skhan@linuxfoundation.org>,
	Thomas Gleixner <tglx@kernel.org>,
	 Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
	 Dave Hansen <dave.hansen@linux.intel.com>,
	x86@kernel.org,  "H. Peter Anvin" <hpa@zytor.com>,
	Vitaly Kuznetsov <vkuznets@redhat.com>,
	Juergen Gross <jgross@suse.com>,
	 Boris Ostrovsky <boris.ostrovsky@oracle.com>,
	Paul Durrant <paul@xen.org>,  Jonathan Cameron <jic23@kernel.org>,
	Sascha Bischoff <Sascha.Bischoff@arm.com>,
	 Marc Zyngier <maz@kernel.org>, Joey Gouly <joey.gouly@arm.com>,
	Jack Allister <jalliste@amazon.com>,
	 Dongli Zhang <dongli.zhang@oracle.com>,
	joe.jin@oracle.com, kvm@vger.kernel.org,
	 linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
	 xen-devel@lists.xenproject.org, linux-kselftest@vger.kernel.org
Subject: Re: [PATCH v7 17/36] KVM: x86: Allow KVM master clock mode when TSCs are offset from each other
Date: Tue, 11 Aug 2026 14:46:31 -0700	[thread overview]
Message-ID: <anuYNxD81OE3MrQw@google.com> (raw)
In-Reply-To: <3dc73f745b6abfd1eb53e7d3fce2067eaa3b3c92.camel@infradead.org>

On Tue, Aug 11, 2026, David Woodhouse wrote:
> On Tue, 2026-08-11 at 11:41 -0700, Sean Christopherson wrote:
> > OMG, I hate this code.  After literally hours of staring at this, and even typing
> > up a lengthy example of why guest time would go off the rails, I finally spotted
> > that l1_tsc_offset is accounted for by the call to kvm_read_l1_tsc().  FML.
> > 
> > Thanks for being patient and not flaming me too much :-)
> 
> Haha, no judgement. We *all* hate this code. That's why I threw my toys
> out of the pram and went on this crusade to clean it up a bit.
> 
> > > > > And your variant just added a dependency on wallclock time back into it
> > > > 
> > > > Can you elaborate?  I'm guessing I don't entirely understand what you mean by
> > > > wallclock time.
> > > 
> > > The system_time field? The unspecified might-be-UTC-might-have-leap-seconds one :)
> > 
> > Ok, I think I finally understand the goal.  I got turned around by the combination
> > of the name SET_CLOCK_GUEST and the full pvclock structure being passed to the
> > guest.  I was expecting SET_CLOCK_GUEST to literally set the entire clock, e.g.
> > mul+shift, timestamp, etc.
> 
> That's an implementation detail. 

Yes and no.  If the payload didn't literally have all the assets needed to set
the kvmclock fields, then I wouldn't care.  But I don't think I'd be the only
person to see a GET+SET pair and expect GET to return exactly what was written
via SET.

> It is literally getting the clock as the guest sees it, and setting it
> again on the destination from the same guest-ABI pvclock structure.
> From the userspace point of view those actions *are* symmetrical.

Only if userspace holds it just so.

> I'd actually *like* it to just be a memcpy at both ends, even on the
> SET side, just copying what userspace provides into what we offer to
> the guest as its pvclock.
> 
> But as well as wanting to do some sanity checking, we also live in a
> world where we might have to switch to the non-masterclock mode at any
> time, and we have to ingest the information into the per-VM kvmclock
> setup in a way that the kernel "understands", and that's why it ends up
> implemented the way it is.
> 
> I looked at rewriting the masterclock base information from what
> userspace is providing, but there are *host* TSC values in there, and
> it ended up in some cases wanting to set ka->master_cycle_now to a
> value which is *negative* on the new host, and I didn't want to
> exercise that wrap-around path. So instead we just adjust
> ka->kvmclock_offset to give appropriate results based on the existing
> masterclock base.
> 
> Each vCPU's pvclock is then *regenerated* from the VM-side kvmclock
> data, giving rise to that annoying ±1ns discrepancy that I whined about
> a while back, but didn't give in to my perfectionism and eliminate...
> yet.
> 
> As far as userspace is concerned, KVM_SET_CLOCK_GUEST *does* set the
> entire clock (at least the relationship between guest TSC and kvmclock
> nanoseconds, which is what it's for). It's just that the kernel then
> "tweaks" it a little bit to give a slightly different y=m(x-x')+c
> equation which is still within the noise of the original.

Again, if and only if the fields match what has been written previoiusly.  To
me, that's not a SET operation given the full inputs.

I completely agree that conceptually this is intended to SET the entire clock,
but as you note above, reality doesn't allow for that.

> And the kernel actually does that kind of 'tweak' all the time.
> Although we're working on narrowing them down because "within the
> noise" is in the eye of the beholder; Dongli had some patches for that
> which I think I rounded up and included?
> 
> > But all of that metadata is just a means to an end: the one and only goal is to
> > calculate the per-VM kvmclock_offset for the "new" host's TSC+time snapshot, by
> > computing the nanoseconds delta for the new snapshot as if it the guest observed
> > the TSC while running on the old host.
> 
> That is currently how it is implemented. It isn't the API contract.
> 
> > And that is done in the kernel instead of in userspace to minimize the amount of
> > slop introduced due to delay between taking the snapshot and computing the offset.
> 
> Huh? There should be no delays here. If *anything* in this new code is
> done with something other than a *simultaneous* reading of TSC and
> ktime that I sweated blood and tears and got shouted at by Thomas for,
> then that *is* something I care about...

Sorry, I didn't mean to imply there would be delay on the SET side.  What I was
trying to say is that if this were punted to userspace, then there _would_ be a
ton of slop because it would be practically impossible for userspace to provide
the correct offset. (I was walking myself through why KVM needed to provide uAPI
to compute the offset, as opposed to provide uAPI to let userspace jam in whatever
value it wanted).

  reply	other threads:[~2026-08-11 21:46 UTC|newest]

Thread overview: 68+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-28 14:39 [PATCH v7 00/36] Cleaning up the KVM clock mess David Woodhouse
2026-07-28 14:39 ` [PATCH v7 01/36] KVM: x86: Improve accuracy of KVM clock when TSC scaling is in force David Woodhouse
2026-07-28 14:39 ` [PATCH v7 02/36] KVM: x86: Explicitly disable TSC scaling without CONSTANT_TSC David Woodhouse
2026-07-28 14:39 ` [PATCH v7 03/36] KVM: x86: Activate master clock immediately on vCPU creation David Woodhouse
2026-07-28 14:39 ` [PATCH v7 04/36] KVM: x86: Avoid NTP frequency skew for KVM clock on 32-bit host David Woodhouse
2026-07-28 14:39 ` [PATCH v7 05/36] KVM: x86: Fold __get_kvmclock() into get_kvmclock() David Woodhouse
2026-07-28 23:06   ` Sean Christopherson
2026-07-28 14:39 ` [PATCH v7 06/36] KVM: x86: Drop CPU pinning in get_kvmclock() David Woodhouse
2026-07-28 14:39 ` [PATCH v7 07/36] KVM: x86: Restructure get_kvmclock() David Woodhouse
2026-07-28 23:04   ` Sean Christopherson
2026-07-28 14:39 ` [PATCH v7 08/36] KVM: x86: Fix KVM clock precision in get_kvmclock() with TSC scaling David Woodhouse
2026-07-28 14:39 ` [PATCH v7 09/36] KVM: x86: Use get_kvmclock() in kvm_get_wall_clock_epoch() David Woodhouse
2026-07-28 14:39 ` [PATCH v7 10/36] KVM: x86: Fix compute_guest_tsc() to handle negative time deltas David Woodhouse
2026-07-28 14:39 ` [PATCH v7 11/36] KVM: x86: Restructure kvm_guest_time_update() for TSC upscaling David Woodhouse
2026-08-04  1:26   ` Sean Christopherson
2026-07-28 14:39 ` [PATCH v7 12/36] KVM: x86: Simplify and comment kvm_get_time_scale() David Woodhouse
2026-07-28 14:39 ` [PATCH v7 13/36] KVM: x86: Remove implicit rdtsc() from kvm_compute_l1_tsc_offset() David Woodhouse
2026-07-28 14:39 ` [PATCH v7 14/36] KVM: x86: Improve synchronization in kvm_synchronize_tsc() David Woodhouse
2026-07-28 14:39 ` [PATCH v7 15/36] KVM: x86: Kill last_tsc_{nsec,write,offset} fields David Woodhouse
2026-07-28 14:39 ` [PATCH v7 16/36] KVM: x86: Replace nr_vcpus_matched_tsc count with all_vcpus_matched_tsc bool David Woodhouse
2026-07-28 14:39 ` [PATCH v7 17/36] KVM: x86: Allow KVM master clock mode when TSCs are offset from each other David Woodhouse
2026-08-10 17:47   ` Sean Christopherson
2026-08-10 18:20     ` David Woodhouse
2026-08-10 20:56       ` Sean Christopherson
2026-08-10 21:05         ` David Woodhouse
2026-08-11 14:16           ` Sean Christopherson
2026-08-11 14:33             ` Sean Christopherson
2026-08-11 15:05               ` David Woodhouse
2026-08-11 16:40                 ` Sean Christopherson
2026-08-11 17:18                   ` David Woodhouse
2026-08-11 17:28                     ` Sean Christopherson
2026-08-11 17:32                       ` David Woodhouse
2026-08-11 18:41                         ` Sean Christopherson
2026-08-11 21:00                           ` David Woodhouse
2026-08-11 21:46                             ` Sean Christopherson [this message]
2026-08-11 21:56                               ` David Woodhouse
2026-08-11 23:14                                 ` Sean Christopherson
2026-07-28 14:39 ` [PATCH v7 18/36] KVM: x86: Factor out kvm_use_master_clock() David Woodhouse
2026-07-28 14:39 ` [PATCH v7 19/36] KVM: x86: Avoid gratuitous global clock updates David Woodhouse
2026-07-28 14:40 ` [PATCH v7 20/36] KVM: x86/xen: Prevent runstate times from becoming negative David Woodhouse
2026-07-28 14:40 ` [PATCH v7 21/36] KVM: x86: Avoid redundant masterclock updates from multiple vCPUs David Woodhouse
2026-07-28 14:40 ` [PATCH v7 22/36] KVM: x86: Remove runtime Xen TSC frequency CPUID update David Woodhouse
2026-07-28 14:40 ` [PATCH v7 23/36] KVM: x86: Re-synchronize TSC after KVM_SET_TSC_KHZ David Woodhouse
2026-07-28 14:40 ` [PATCH v7 24/36] KVM: x86: Use ktime_get_snapshot_id() for master clock David Woodhouse
2026-07-28 14:40 ` [PATCH v7 25/36] KVM: x86: Compute kvmclock base without pvclock_gtod_data David Woodhouse
2026-07-28 14:40 ` [PATCH v7 26/36] KVM: x86: Cache host vclock_mode for masterclock eligibility checks David Woodhouse
2026-07-28 14:40 ` [PATCH v7 27/36] KVM: x86: Remove pvclock_gtod_data and private timekeeping code David Woodhouse
2026-07-28 14:40 ` [PATCH v7 28/36] KVM: x86: Activate master clock from kvm_arch_init_vm() David Woodhouse
2026-07-28 14:40 ` [PATCH v7 29/36] UAPI: x86: Move pvclock-abi to UAPI for x86 platforms David Woodhouse
2026-07-28 14:40 ` [PATCH v7 30/36] KVM: selftests: Use UAPI pvclock-abi.h in xen_shinfo_test David Woodhouse
2026-07-28 14:40 ` [PATCH v7 31/36] KVM: x86: Add KVM_[GS]ET_CLOCK_GUEST for accurate KVM clock migration David Woodhouse
2026-07-31 23:24   ` Sean Christopherson
2026-08-01  8:07     ` David Woodhouse
2026-08-04 23:38       ` Sean Christopherson
2026-08-05  9:26         ` David Woodhouse
2026-08-11 23:40   ` Sean Christopherson
2026-08-11 23:47     ` David Woodhouse
2026-07-28 14:40 ` [PATCH v7 32/36] KVM: x86: Add KVM_VCPU_TSC_SCALE and fix the documentation on TSC migration David Woodhouse
2026-07-28 14:40 ` [PATCH v7 33/36] KVM: selftests: Add KVM/PV clock selftest to prove timer correction David Woodhouse
2026-07-31 23:32   ` Sean Christopherson
2026-07-28 14:40 ` [PATCH v7 34/36] KVM: selftests: Add master clock offset test David Woodhouse
2026-07-31 23:38   ` Sean Christopherson
2026-07-28 14:40 ` [PATCH v7 35/36] KVM: selftests: Add Xen/generic CPUID timing leaf test David Woodhouse
2026-07-28 14:40 ` [PATCH v7 36/36] KVM: selftests: Add Xen runstate migration test David Woodhouse
2026-07-28 23:18 ` [PATCH v7 00/36] Cleaning up the KVM clock mess Sean Christopherson
2026-07-29 10:42   ` David Woodhouse
2026-08-10 16:42 ` Sean Christopherson
2026-08-10 16:56   ` David Woodhouse

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anuYNxD81OE3MrQw@google.com \
    --to=seanjc@google.com \
    --cc=Sascha.Bischoff@arm.com \
    --cc=boris.ostrovsky@oracle.com \
    --cc=bp@alien8.de \
    --cc=corbet@lwn.net \
    --cc=dave.hansen@linux.intel.com \
    --cc=dongli.zhang@oracle.com \
    --cc=dwmw2@infradead.org \
    --cc=hpa@zytor.com \
    --cc=jalliste@amazon.com \
    --cc=jgross@suse.com \
    --cc=jic23@kernel.org \
    --cc=joe.jin@oracle.com \
    --cc=joey.gouly@arm.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=maz@kernel.org \
    --cc=mingo@redhat.com \
    --cc=paul@xen.org \
    --cc=pbonzini@redhat.com \
    --cc=skhan@linuxfoundation.org \
    --cc=tglx@kernel.org \
    --cc=vkuznets@redhat.com \
    --cc=x86@kernel.org \
    --cc=xen-devel@lists.xenproject.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox