Linux Kernel Selftest development
 help / color / mirror / Atom feed
From: Pratyush Yadav <pratyush@kernel.org>
To: Sean Christopherson <seanjc@google.com>
Cc: Pratyush Yadav <pratyush@kernel.org>,
	 Tarun Sahu <tarunsahu@google.com>,
	 ackerleytng@google.com,  fuad.tabba@linux.dev,
	Andrew Morton <akpm@linux-foundation.org>,
	 dmatlack@google.com,  Shuah Khan <skhan@linuxfoundation.org>,
	 Jonathan Corbet <corbet@lwn.net>,
	david@redhat.com,  Pasha Tatashin <pasha.tatashin@soleen.com>,
	sagis@google.com,  Paolo Bonzini <pbonzini@redhat.com>,
	 Mike Rapoport <rppt@kernel.org>,
	 Alexander Graf <graf@amazon.com>,
	linux-kselftest@vger.kernel.org,  andre.przywara@arm.com,
	michael.roth@amd.com,  linux-kernel@vger.kernel.org,
	 linux-mm@kvack.org, will@kernel.org,  vannapurve@google.com,
	 maz@kernel.org, fvdl@google.com,  kvm@vger.kernel.org,
	 oliver.upton@linux.dev, kvmarm@lists.linux.dev,
	 alexandru.elisei@arm.com,  skhawaja@google.com,
	aneesh.kumar@kernel.org,  linux-doc@vger.kernel.org,
	 David Hildenbrand <david@kernel.org>,
	 yan.y.zhao@intel.com,  kexec@lists.infradead.org,
	suzuki.poulose@arm.com
Subject: Re: [PATCH v4 05/11] KVM: LUO: Support VM preservation across live updates
Date: Sat, 15 Aug 2026 12:43:56 +0200	[thread overview]
Message-ID: <2vxzqzjz1x5f.fsf@kernel.org> (raw)
In-Reply-To: <anyObKiZIhfok5mk@google.com> (Sean Christopherson's message of "Wed, 12 Aug 2026 08:17:00 -0700")

On Wed, Aug 12 2026, Sean Christopherson wrote:

> On Wed, Aug 12, 2026, Pratyush Yadav wrote:
>> On Tue, Aug 11 2026, Sean Christopherson wrote:
[...]
> But that's all beside the point.  What I'm saying is that SNP and TDX *must* use
> guest_memfd, whereas guest_memfd is optional for all other VM types.  And so adding
> LUO support for guest_memfd without even sketching out a plan for CoCo VMs feels
> backwards.
>
> I'm not necessarily opposed to gradual guest_memfd support, but there needs to be
> a clear plan of how all of this is going to fit together.  It doesn't need to be
> perfect, and I'm sure we'll make mistakes along the way, but I want to at least
> try not to paint ourselves into a corner, especially with respect to the ABI.

Right. I think landing shared guest_memfd is the first step. It is
simple enough and doesn't need to preserve the more complicated
architecture specific things like secure EPT, etc. The CoCo support
can build on top of this.

There is work underway already for an RFC of TDX live update, but Tarun
knows the details better than me so I'll let him fill in the details for
the next version of this series.

>
>> >> > E.g. the LUO APIs seem pretty straightforward; I assume the bulk of the complexity
>> >> > is going to be in knowing what to save/restore, and how, which is much more about
>> >> > KVM than it is about liveupdate.
>> >> 
>> >> I think it is fine if you want to take these changes through the KVM
>> >> tree, but I would like live update maintainers to be listed as reviewers
>> >> at least.
>> >
>> > Why not simply add a file pattern match to the LIVE UPDATE entry?
>> >
>> > diff --git MAINTAINERS MAINTAINERS
>> > index 8014b9f8253e..2eb57b22c37f 100644
>> > --- MAINTAINERS
>> > +++ MAINTAINERS
>> > @@ -15052,8 +15052,8 @@ F:      include/linux/liveupdate.h
>> >  F:     include/uapi/linux/liveupdate.h
>> >  F:     kernel/liveupdate/
>> >  F:     lib/tests/liveupdate.c
>> > -F:     mm/memfd_luo.c
>> >  F:     tools/testing/selftests/liveupdate/
>> > +N:     [^a-z]luo
>> >  
>> >  LLC (802.2)
>> >  L:     netdev@vger.kernel.org
>> 
>> This would list us as maintainers of kvm_luo.c and the tree as liveupdate.git, 
>
> It would list both "LIVE UPDATE" and "KERNEL VIRTUAL MACHINE (KVM)", as KVM would
> still cover the files via "F:	virt/kvm/*".  And my read of MAINTAINERS is that
> N: and K: entries are "secondary" if the files/scope is covered by an F: entry.
>
> E.g. arch/x86/kvm/vmx/tdx.c is covered by "KERNEL VIRTUAL MACHINE FOR X86 (KVM/x86)"
> via "F:	arch/x86/kvm/*/", and also by "X86 TRUST DOMAIN EXTENSIONS (TDX)" via
> "N:	tdx" (or maybe "K:	\b(tdx)"?  I haven't bothered to check which one
> triggers).  And while I don't think there was ever any formal discussion, AFAIK
> everyone reads the situation as KVM still being the primary maintainer/tree for
> that code.
>
>> both of which is something you're saying you _don't_ want.
>
> No, what I don't want is a dedicated "KVM LIVE UPDATE" entry, because I think most
> people would read "F:     virt/kvm/kvm_luo.c" and "F:     virt/kvm/guest_memfd_luo.c"
> as being more precise than KVM's "F:   virt/kvm/*" and thus would read things as
> "KVM LIVE UPDATE" being the primary maintainer.
>
> In other words, I'm more than ok with LUO being looped in on KVM LUO changes and
> having the authority to object to problematic changes, but I'm not ok with LUO
> taking primary ownership of KVM code.

Sure, that sounds good to me.

>
> I realize there's more than a bit of nuance in my interpretation of N: and K:,
> but again my experience with TDX is that so long as the maintainers are aligned
> on expectations, it's a non-issue in practice.

Yep, as long as we agree among ourselves I don't think the red tape
around which entry has higher priority matters much.

>
>> >> At the same time, I also keep being (pleasantly)
>> >> surprised at preservation being relatively simple. For example, the code
>> >> to preserve a shmem file (via memfd) is roughly 600 lines, a big chunk
>> >> of which is comments. The code of course has some limitations, but it is
>> >> good enough for use in production.
>> >> 
>> >> For one, we care about ABI breakages and versioning. 
>> >
>> > Which is amusing to me because that implies KVM does not, and I would hazard to
>> 
>> No, it doesn't. What I'm saying is I care about changes to the _live
>> update ABI_. Just like you probably care about changes to KVM ABI but
>> not so much about BPF for example.
>> 
>> > guess that KVM has the biggest ABI surface of any subsystem in the kernel by a
>> > country mile (though I'm probably wildly underestimating the effective ABI surface
>> > of filesystems).
>> >
>> >> The serialized state is a part of live update ABI and changes to it should be
>> >> ACKed by us. 
>> >
>> > Meh, "Don't break userspace" is a universal rule in the kernel, I genuinely don't
>> > see why liveupdate needs special treatment.
>> 
>> Ironically enough, you miss my point. I'm not talking about userspace
>> ABI. We all know not to break that. I am talking about live update
>> _serialization ABI_. See the stuff under include/linux/kho/abi. This
>> series also adds things there.
>
> No, I understand exactly what ABI you're talking about.
>
>> This is ABI between kernels. It needs to be stable-ish so you can move
>> from one kernel version to another. At the same time, unlike userspace
>> ABI, it can change.
>
> Uh, yeah, so KVM has been managing such immutable ABI for practically its entire
> existence.  KVM's save/restore uAPI has exactly what you're describing: serialization
> ABI that needs to be backwards and forwards compatible between different kernels
> in order to support both upgrade and rollback scenarios via live migration.

I think there is a slight difference between KVM's save/resture uAPI and
live update's ABI. With KVM's uAPI, you need to maintain strict
backwards compatibility because userspace reads what you output. So if
you change the layout, userspace will interpret it wrong and might
break.

With live update, the ABI never gets to userspace. It is used to talk
between kernels directly. So if you do break that, your userspace keeps
working fine, you just might not be able to live update to the
incompatible kernel.

So with live update, we don't need to keep backwards compatibility in
the ABI forever. Of course, it is good to minimize changes, but we have
more freedom to change it.

The current idea is we change the ABI whenever needed. This is because
live update is still in development so we are likely to see a lot more
changes before we settle down into something that mostly works. At some
point, we want to start being more strict between incompatible ABI
changes. The current idea is that we would like to keep backwards
compatibility across LTS kernels. But none of that is set in stone yet.

That's why I would like to have an explicit ACK from the live update
group for ABI changes. I would like to enforce these compatibility
windows.

>
> KVM also has immutable ABI between itself and guest kernels, including implicit
> "ABI" in the form of not changing guest-visible behavior.
>
>> Today we don't have any rules and let you change things freely as long as you
>> do a version bump. But at a later point, the plan is to add some stability
>> requirements to the ABI so you can actually upgrade the kernel across major
>> versions.
>> 
>> So at least for ABI changes, there should be an explicit ACK from the
>> live update group.
>
> I don't entirely agree.  I get where you're coming from, and I 100% agree that
> the more eyeballs on changes that may affect ABI, the better.
>
> Where I have problems with the above is that it doesn't account for the nuances
> of save/restore across different kernels.  The literal format of the serialized
> data is the most obvious form of ABI, but it's absolutely possible to change ABI
> without changing the data format.
>
> E.g. say there's a flags field in some serialization structure, and flags X and Y
> are mutually exclusive in current kernels.  If a future kernel relaxes that
> restriction for whatever reason, then the effective ABI has been broken because
> state created and saved on new kernels can't be restored on old kernels.  And
> this is not a theoretical concern, KVM has run afoul of this a few times (though
> thankfully very rarely).

That's a bug, and should be fixed.

I am talking here about intentional ABI changes. As I mentioned above we
_can_ change the ABI, we just need to be careful about it.

>
> I don't think it's reasonable to expect LUO maintainers to gain enough expertise
> in each subsystem to be able to ensure changes are forwards and backwards
> compatible.  The only way I see LUO being successful in the long term is to get
> subsystem maintainers/contributors to understand *and buy-in* to the LUO model
> and rules, so that each subsystem can largely be self-sustatining.  I.e. setting
> yourselves up as literal gatekeepers will help prevent blatant breakage, but it's
> less likely to help guard against more subtle breakage, and in my experience,
> subtle breakage is by far harder to detect and more painful to deal with.

Outside of enforcing the ABI compatibility windows, I think this makes
sense.

In principle at least, though from my experience the maintainer appetite
to care about live update varies. The feeling I've got from MM for
example is more along the lines of "you break it, you buy it". Which is
also a valid model IMO. The person who wrote the live update code for a
subsystem is well equipped to understand these nuances.

>
> Somewhat of a side topic: in my experience, using monotonically increasing version
> numbers is a horrible way to enumerate features/content.  So for me, allowing
> KVM's LUO ABI to change with a verson bump is probably a non-starter.  I.e.
> whatever gets merged needs to be more future-proof than "we'll deal with it later".

As I wrote above, the compatibility model for live update allows
breaking ABI with version bumps. Userspace keeps working fine, it only
restricts the set of kernels that you can go to. So the "we'll deal with
it later" is not as harmful because our choices here are not permanent.

Also, there is work underway to allow transitions between versions [0].

But still, I am not opposed to a more flexible guest_memfd ABI. I think
we can reserve some space for feature flags to let us add things in the
future.

[0] https://lore.kernel.org/kexec/20260731215224.831696-1-loganodell@google.com/

-- 
Regards,
Pratyush Yadav

  reply	other threads:[~2026-08-15 10:44 UTC|newest]

Thread overview: 37+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-28 12:11 [PATCH v4 00/11] liveupdate: kvm: Guest_memfd preservation Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 01/11] liveupdate: Add LIVEUPDATE_GUEST_MEMFD config option Tarun Sahu
2026-08-10 22:58   ` Sean Christopherson
2026-08-11 13:26     ` tarunsahu
2026-08-11 14:48       ` Sean Christopherson
2026-07-28 12:11 ` [PATCH v4 02/11] KVM: Introduce kvm_create_vm_file() helper Tarun Sahu
2026-07-30 17:36   ` Ackerley Tng
2026-08-10 10:14     ` tarunsahu
2026-08-10 23:05   ` Sean Christopherson
2026-08-11 13:27     ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 03/11] KVM: Export kvm_uevent_notify_vm_create() Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 04/11] KVM: Track weak reference to vm_file in struct kvm Tarun Sahu
2026-08-10 23:23   ` Sean Christopherson
2026-07-28 12:11 ` [PATCH v4 05/11] KVM: LUO: Support VM preservation across live updates Tarun Sahu
2026-08-10 23:42   ` Sean Christopherson
2026-08-11 11:31     ` Pratyush Yadav
2026-08-11 14:05       ` Sean Christopherson
2026-08-12 13:45         ` Pratyush Yadav
2026-08-12 15:17           ` Sean Christopherson
2026-08-15 10:43             ` Pratyush Yadav [this message]
2026-07-28 12:11 ` [PATCH v4 06/11] KVM: guest_memfd: Move internal definitions to internal header Tarun Sahu
2026-07-30 18:12   ` Ackerley Tng
2026-08-11 10:31     ` Pratyush Yadav
2026-07-28 12:11 ` [PATCH v4 07/11] KVM: guest_memfd: Add support for freezing mappings Tarun Sahu
2026-07-30 17:46   ` Ackerley Tng
2026-08-10 13:15     ` tarunsahu
2026-07-30 18:12   ` Ackerley Tng
2026-08-10 13:08     ` tarunsahu
2026-08-10 23:44   ` Sean Christopherson
2026-07-28 12:11 ` [PATCH v4 08/11] KVM: guest_memfd: Add support for preservation via LUO Tarun Sahu
2026-07-30 18:16   ` Ackerley Tng
2026-08-10 13:20     ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 09/11] docs: liveupdate: Add documentation for VM and guest_memfd preservation Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 10/11] KVM: selftests: Split ____vm_create() and add vm_create_from_fd() Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 11/11] KVM: selftests: Add guest_memfd_preservation_test Tarun Sahu
2026-07-30 18:18   ` Ackerley Tng
2026-08-10 13:22     ` tarunsahu

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=2vxzqzjz1x5f.fsf@kernel.org \
    --to=pratyush@kernel.org \
    --cc=ackerleytng@google.com \
    --cc=akpm@linux-foundation.org \
    --cc=alexandru.elisei@arm.com \
    --cc=andre.przywara@arm.com \
    --cc=aneesh.kumar@kernel.org \
    --cc=corbet@lwn.net \
    --cc=david@kernel.org \
    --cc=david@redhat.com \
    --cc=dmatlack@google.com \
    --cc=fuad.tabba@linux.dev \
    --cc=fvdl@google.com \
    --cc=graf@amazon.com \
    --cc=kexec@lists.infradead.org \
    --cc=kvm@vger.kernel.org \
    --cc=kvmarm@lists.linux.dev \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=maz@kernel.org \
    --cc=michael.roth@amd.com \
    --cc=oliver.upton@linux.dev \
    --cc=pasha.tatashin@soleen.com \
    --cc=pbonzini@redhat.com \
    --cc=rppt@kernel.org \
    --cc=sagis@google.com \
    --cc=seanjc@google.com \
    --cc=skhan@linuxfoundation.org \
    --cc=skhawaja@google.com \
    --cc=suzuki.poulose@arm.com \
    --cc=tarunsahu@google.com \
    --cc=vannapurve@google.com \
    --cc=will@kernel.org \
    --cc=yan.y.zhao@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox