Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Dongli Zhang <dongli.zhang@oracle.com>
To: Marc Zyngier <maz@kernel.org>
Cc: kvm@vger.kernel.org, kvmarm@lists.linux.dev,
	linux-kselftest@vger.kernel.org, seanjc@google.com,
	pbonzini@redhat.com, oupton@kernel.org, fuad.tabba@linux.dev,
	joey.gouly@arm.com, seiden@linux.ibm.com, suzuki.poulose@arm.com,
	yuzenghui@huawei.com, dwmw2@infradead.org, joe.jin@oracle.com
Subject: Re: [PATCH 0/4] KVM: Reset steal time accounting on vCPU pid change (x86 and arm64)
Date: Mon, 17 Aug 2026 14:05:10 -0700	[thread overview]
Message-ID: <d71a5dbd-cabe-476d-8837-9519b88ee5cf@oracle.com> (raw)
In-Reply-To: <86tsotyxh5.wl-maz@kernel.org>



On Mon, Aug 17, 2026 1:15:02AM -0700, Marc Zyngier wrote:
> On Sun, 16 Aug 2026 06:33:01 +0100,
> Dongli Zhang <dongli.zhang@oracle.com> wrote:
> 
> [...]
> 
>> 4. Guest kernel changes are not included. I may send it separately to keep
>> this series limited to the KVM hypervisor.
> 
> Hold on. Do you mean you are changing the guest visible behaviour of a
> PV interface? That's an ABI. It *cannot* change unilaterally.

No, this does not change the ABI.

Taking x86 as an example, the Linux guest currently does not reset the stealtime
accounting metadata, i.e. rq->prev_steal_time and rq->prev_steal_time_rq, when a
vCPU is brought online. I meant to reset that metadata before enabling KVM x86
steal time via MSR_KVM_STEAL_TIME, so that the guest is not affected by a stale
steal-time value from before the vCPU was offlined.

For example, in the code below, the guest resets rq->prev_steal_time and
rq->prev_steal_time_rq after enabling KVM steal time via MSR_KVM_STEAL_TIME.

[PATCH 1/5] x86/kvm: Reset prev_steal_time and prev_steal_time_rq when enabling
steal time
https://lore.kernel.org/all/20260505003044.78693-2-dongli.zhang@oracle.com

+void sched_steal_time_cpu_init(int cpu, u64 steal)
+{
+	struct rq *rq = cpu_rq(cpu);
+
+	rq->prev_steal_time = steal;
+#ifdef CONFIG_PARAVIRT_TIME_ACCOUNTING
+	rq->prev_steal_time_rq = steal;
+#endif
+}
+
 #ifdef CONFIG_HAVE_PV_STEAL_CLOCK_GEN
 static u64 native_steal_clock(int cpu)
 {


@@ -337,6 +354,12 @@ static void kvm_register_steal_time(void)
 		return;

 	wrmsrq(MSR_KVM_STEAL_TIME, (slow_virt_to_phys(st) | KVM_MSR_ENABLED));
+
+	/*
+	 * This CPU is not ready to be scheduled yet.
+	 */
+	sched_steal_time_cpu_init(cpu, kvm_steal_clock(cpu));
+
 	pr_debug("stealtime: cpu %d, msr %llx\n", cpu,
 		(unsigned long long) slow_virt_to_phys(st));
 }

> 
>>
>> [PATCH 1/5] x86/kvm: Reset prev_steal_time and prev_steal_time_rq when enabling steal time
>> https://urldefense.com/v3/__https://lore.kernel.org/all/20260505003044.78693-2-dongli.zhang@oracle.com__;!!ACWV5N9M2RV99hQ!Kt5J7E63dvSlZsI1Ovt02JMZuh0BwYZZupi0BMQXDxnYkQ6iPNwB7tuWgd7wSVOEVHCF40thrD8Apg$ 
>>
>> 5. There is one remaining corner case: this series resets last_steal when
>> the vCPU run PID changes, but not when steal time is enabled. If additional
>> host run_delay is accumulated after the PID changes but before the guest
>> enables steal time, that delta could be unexpectedly accounted to guest
>> vCPU steal time. In practice, this should not happen for Linux guests.
> 
> Why is Linux immune to this? Also, KVM does not cater for Linux guests
> only.
I should have explained this more clearly.

I did not mean that Linux is immune in the architectural sense, or that KVM
should rely on Linux guest behavior. The corner case I had in mind is:

1. The vCPU is run by a new host PID, so KVM resets last_steal to the new PID's
current run_delay.

2. The guest has not enabled steal time yet.

3. The new host task accumulates additional run_delay before the guest enables
steal time.

4. When steal time is later enabled, that pre-enable run_delay can be included
in the first stealtime update.

For the Linux guest on QEMU/KVM x86 path, this window should normally be very
small. For example, on x86, a Linux guest enables KVM steal time immediately
from the CPU online path.  Also, QEMU does not enter KVM_RUN for an offline
vCPU, so the normal QEMU/KVM hotplug path leaves little opportunity to
accumulate run_delay before steal time is enabled.

Thank you very much!

Dongli Zhang

  reply	other threads:[~2026-08-17 21:07 UTC|newest]

Thread overview: 16+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-16  5:33 [PATCH 0/4] KVM: Reset steal time accounting on vCPU pid change (x86 and arm64) Dongli Zhang
2026-08-16  5:33 ` [PATCH 1/4] KVM: x86: Reset last_steal on vCPU pid change Dongli Zhang
2026-08-16  5:55   ` sashiko-bot
2026-08-17 20:21     ` Dongli Zhang
2026-08-16  5:33 ` [PATCH 2/4] KVM: arm64: " Dongli Zhang
2026-08-17  8:42   ` Marc Zyngier
2026-08-17 21:29     ` Dongli Zhang
2026-08-18 13:24       ` Marc Zyngier
2026-08-16  5:33 ` [PATCH 3/4] KVM: selftests: Test steal time across vCPU pid changes on x86 Dongli Zhang
2026-08-16  5:55   ` sashiko-bot
2026-08-17 20:28     ` Dongli Zhang
2026-08-17  8:15 ` [PATCH 0/4] KVM: Reset steal time accounting on vCPU pid change (x86 and arm64) Marc Zyngier
2026-08-17 21:05   ` Dongli Zhang [this message]
2026-08-17 11:51 ` David Woodhouse
2026-08-17 22:04   ` Dongli Zhang
2026-08-17 22:23     ` David Woodhouse

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=d71a5dbd-cabe-476d-8837-9519b88ee5cf@oracle.com \
    --to=dongli.zhang@oracle.com \
    --cc=dwmw2@infradead.org \
    --cc=fuad.tabba@linux.dev \
    --cc=joe.jin@oracle.com \
    --cc=joey.gouly@arm.com \
    --cc=kvm@vger.kernel.org \
    --cc=kvmarm@lists.linux.dev \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=maz@kernel.org \
    --cc=oupton@kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=seanjc@google.com \
    --cc=seiden@linux.ibm.com \
    --cc=suzuki.poulose@arm.com \
    --cc=yuzenghui@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox