From: tarunsahu@google.com
To: Sean Christopherson <seanjc@google.com>,
Pratyush Yadav <pratyush@kernel.org>
Cc: ackerleytng@google.com, fuad.tabba@linux.dev,
Andrew Morton <akpm@linux-foundation.org>,
dmatlack@google.com, Shuah Khan <skhan@linuxfoundation.org>,
Jonathan Corbet <corbet@lwn.net>,
david@redhat.com, Pasha Tatashin <pasha.tatashin@soleen.com>,
sagis@google.com, Paolo Bonzini <pbonzini@redhat.com>,
Mike Rapoport <rppt@kernel.org>, Alexander Graf <graf@amazon.com>,
linux-kselftest@vger.kernel.org, andre.przywara@arm.com,
michael.roth@amd.com, linux-kernel@vger.kernel.org,
linux-mm@kvack.org, will@kernel.org, vannapurve@google.com,
maz@kernel.org, fvdl@google.com, kvm@vger.kernel.org,
oliver.upton@linux.dev, kvmarm@lists.linux.dev,
alexandru.elisei@arm.com, skhawaja@google.com,
aneesh.kumar@kernel.org, linux-doc@vger.kernel.org,
David Hildenbrand <david@kernel.org>,
yan.y.zhao@intel.com, kexec@lists.infradead.org,
suzuki.poulose@arm.com
Subject: Re: [PATCH v4 05/11] KVM: LUO: Support VM preservation across live updates
Date: Tue, 18 Aug 2026 16:10:31 +0000 [thread overview]
Message-ID: <9huzo6ezv288.fsf@tarunix.c.googlers.com> (raw)
In-Reply-To: <anssOQFNZ2LfihbU@google.com>
Sean Christopherson <seanjc@google.com> writes:
> On Tue, Aug 11, 2026, Pratyush Yadav wrote:
>> On Mon, Aug 10 2026, Sean Christopherson wrote:
>>
>> > On Tue, Jul 28, 2026, Tarun Sahu wrote:
>> >> Register a Live Update Orchestrator (LUO) file handler for KVM VM files
>> >> to serialize and deserialize VM state across kexec live updates.
>> >>
>> >> Currently, Only VM type (e.g. arch.vm_type on x86) is preserved as part
>> >> of VM preservation.
>> >
>> > Why?
>> >
>> >> On retrieval, kvm_luo_retrieve() recreates the KVM VM file via
>> >> kvm_create_vm_file() and use an atomically incremented ID for the internal
>> >> fdname, as the final fdname assigned by userspace is not yet known during
>> >> retrieval. As this fdname is only used in debugfs infra, This will not break
>> >> any UAPI.
>> >>
>> >> This infrastructure establishes the foundation for preserving guest_memfd
>> >> instances across live updates, and can be expanded in the future to
>> >> preserve additional VM state.
>> >
>> > Uh, why guest_memfd? As much as I want to push guest_memfd adoption, it seems
>> > guest_memfd should be the _last_ thing we support, not the first. As evidenced
>> > by the last two decades, it's very doable to have KVM VMs without guest_memfd,
>> > but it's rather hard to have VMs without vCPUs.
>>
>> You _can_ preserve vCPUs today using KVM_{GET,SET}_REGS, they just won't
>> run in the background during the reboot.
>
> What about x86 CoCo VMs? Which are quite literally _the_ reason guest_memfd was
> created in the first place.
>
>> This series can save you from dumping VM memory to disk if it is backed by
>> guest_memfd.
>
> Or to word it another way, one _can_ save guest_memfd, it's just
> slower.
>
>
> My point is that this series needs to provide a _lot_ more information about the
> bigger KVM picture.
Yes, I agree. I will try to layout the plan.
End goal of this series is to preserve VM memory which is backed by
guest_memfd. Not all VM are backed by guest_memfd. So preservation of
KVM (vm_file) is independent of preservation of guest_memfd.
But guest_memefd can not be preserved without preserving the KVM (vm_file).
So I agree to your suggestion:
diff --git virt/kvm/Makefile.kvm virt/kvm/Makefile.kvm
index d047d4cf58c9..e6f098498795 100644
--- virt/kvm/Makefile.kvm
+++ virt/kvm/Makefile.kvm
@@ -13,3 +13,8 @@ kvm-$(CONFIG_HAVE_KVM_IRQ_ROUTING) += $(KVM)/irqchip.o
kvm-$(CONFIG_HAVE_KVM_DIRTY_RING) += $(KVM)/dirty_ring.o
kvm-$(CONFIG_HAVE_KVM_PFNCACHE) += $(KVM)/pfncache.o
kvm-$(CONFIG_KVM_GUEST_MEMFD) += $(KVM)/guest_memfd.o
+
+ifdef CONFIG_LIVEUPDATE
+kvm-y += $(KVM)/kvm_luo.o
+kvm-$(CONFIG_KVM_GUEST_MEMFD) += $(KVM)/guest_memfd_luo.o
+endif
Guest_memfd can't be created alone without struct kvm
(kvm_gmem_create() and KVM_GMEM_CREATE ioctl). This is how
guest_memfd has been designed. I don't want to break this design which
has been accepted upstream after lot of discussion.
So while LUO preserve guest_memfd's data, It just preserve the PFNs
value, few flags that belongs to the guest_memfd.
On restore side in new kernel, These PFNs, flags will be populated
to a _newly_ created guest_memfd. that is how preservation and
retrieval works. So To create this new guest_memfd, We need struct
kvm (vm_file). this vm_file must be the same VM which had this
guest_memfd in old kernel, So we preserve the vm_file, get its TOKEN
preserve with guest_memfd and during restore, we create the guest_memfd
with the same VM (vm_file/struct kvm).
I agree, I did not do good job explaining things in commit message. I
will make sure to update them in next revision.
> For those of us that are on the very fringes of live update,
> it's practically impossible to review because, to us, it seems very arbitrary.
>
> The part that's especially confusing is the saving of the VM type. That comes
> straight from userspace, so it's super bizarre to automatically save/restore that,
> but nothing else.
Like, guest_memfd needs struct kvm to create itself. struct kvm (vm_file) needs
vm_type to create itself (kvm_create_vm() or KVM_CREATE_VM IOCTL). LUO
does not provde functionality to pass any subsystem specific arguments. So
vm_type needs to be preserved even though, userspace is aware about it.
So Why do we preserve only vm_type: To keep things simple for this
series, As target is guest_memfd.
Currently I dont have discreet plan on what else will
,in future, be needed to be preserved. Which I agree not a absoulute right
way to approach. I will layout a rough plan on KVM side preservation.
Having this need for backward compatiblity: I responded here:
https://lore.kernel.org/all/9huzv797v45k.fsf@tarunix.c.googlers.com/
>
>> >> +KVM LIVE UPDATE
>> >> +M: Pasha Tatashin <pasha.tatashin@soleen.com>
>> >> +M: Mike Rapoport <rppt@kernel.org>
>> >> +M: Pratyush Yadav <pratyush@kernel.org>
>> >> +R: Tarun Sahu <tarunsahu@google.com>
>> >> +L: kexec@lists.infradead.org
>> >> +L: kvm@vger.kernel.org
>> >> +S: Maintained
>> >> +T: git git://git.kernel.org/pub/scm/linux/kernel/git/liveupdate/linux.git
>> >
>> > NAK on taking changes through a different tree. This is KVM code, period.
>> >
>> > In general, I'm skeptical of the dedicated MAINTAINERS entry. It's extremely
>> > difficult to tell since this series is little more than a skeleton (either that
>> > or liveupdate is way simpler that I was expecting), but I suspect that maintaining
>
> ...
>
>> > E.g. the LUO APIs seem pretty straightforward; I assume the bulk of the complexity
>> > is going to be in knowing what to save/restore, and how, which is much more about
>> > KVM than it is about liveupdate.
>>
>> I think it is fine if you want to take these changes through the KVM
>> tree, but I would like live update maintainers to be listed as reviewers
>> at least.
>
> Why not simply add a file pattern match to the LIVE UPDATE entry?
>
> diff --git MAINTAINERS MAINTAINERS
> index 8014b9f8253e..2eb57b22c37f 100644
> --- MAINTAINERS
> +++ MAINTAINERS
> @@ -15052,8 +15052,8 @@ F: include/linux/liveupdate.h
> F: include/uapi/linux/liveupdate.h
> F: kernel/liveupdate/
> F: lib/tests/liveupdate.c
> -F: mm/memfd_luo.c
> F: tools/testing/selftests/liveupdate/
> +N: [^a-z]luo
>
> LLC (802.2)
> L: netdev@vger.kernel.org
>
>> At the same time, I also keep being (pleasantly)
>> surprised at preservation being relatively simple. For example, the code
>> to preserve a shmem file (via memfd) is roughly 600 lines, a big chunk
>> of which is comments. The code of course has some limitations, but it is
>> good enough for use in production.
>>
>> For one, we care about ABI breakages and versioning.
>
> Which is amusing to me because that implies KVM does not, and I would hazard to
> guess that KVM has the biggest ABI surface of any subsystem in the kernel by a
> country mile (though I'm probably wildly underestimating the effective ABI surface
> of filesystems).
>
>> The serialized state is a part of live update ABI and changes to it should be
>> ACKed by us.
>
> Meh, "Don't break userspace" is a universal rule in the kernel, I genuinely don't
> see why liveupdate needs special treatment.
>
>> For another, how the file handlers interact with their dependencies can
>> affect the behaviour that VMMs observe. Those changes should also pass by
>> some live update eyes.
>
> Perhaps in the short term, but IMO, that's not a winning strategy in the long
> term. From my perspective, that like saying the PAGE CACHE maintainers should
> review every usage of the filemap APIs, because how the APIs are used impacts
> the page cache and affects userspace-visible behavior. There are myriad analogies
> like that throughout the kernel.
>
> Yes, liveupdate is new and shiny, but IMO for it to be successful and maintainable,
> it needs to be treated like any other core infrastructure in the kernel, not a
> special snowflake whose details are known only by a handful of people. Because
> I think it's likely liveupdate goes one of two ways: either liveupdate becomes a
> very niche thing that is used sparingly throughout the kernel, or it becomes a
> broadly used feature that is supported by many filesystems and subsystems.
>
> If liveupdate is relegated to niche status, then it probably isn't going to see
> a significant amount of ongoing development, at which point the folks working on
> liveupdate will naturally migrate to other projects, and maintenance will largely
> be left to subsystem maintainers.
>
> If liveupdate is broadly used, then having a single group of people maintain
> every subsystem's usage won't scale, and maintenance will again largely fall on
> the shoulder of subsystem maintainers. Which is totally fine and working as
> intended, because that's exactly what subystem maintainers are signing up for
> by merging support for liveupdate.
next prev parent reply other threads:[~2026-08-18 16:10 UTC|newest]
Thread overview: 47+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-28 12:11 [PATCH v4 00/11] liveupdate: kvm: Guest_memfd preservation Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 01/11] liveupdate: Add LIVEUPDATE_GUEST_MEMFD config option Tarun Sahu
2026-08-10 22:58 ` Sean Christopherson
2026-08-11 13:26 ` tarunsahu
2026-08-11 14:48 ` Sean Christopherson
2026-08-18 16:11 ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 02/11] KVM: Introduce kvm_create_vm_file() helper Tarun Sahu
2026-07-30 17:36 ` Ackerley Tng
2026-08-10 10:14 ` tarunsahu
2026-08-10 23:05 ` Sean Christopherson
2026-07-28 12:11 ` [PATCH v4 03/11] KVM: Export kvm_uevent_notify_vm_create() Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 04/11] KVM: Track weak reference to vm_file in struct kvm Tarun Sahu
2026-08-10 23:23 ` Sean Christopherson
2026-08-18 16:29 ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 05/11] KVM: LUO: Support VM preservation across live updates Tarun Sahu
2026-08-10 23:42 ` Sean Christopherson
2026-08-11 11:31 ` Pratyush Yadav
2026-08-11 14:05 ` Sean Christopherson
2026-08-12 13:45 ` Pratyush Yadav
2026-08-12 15:17 ` Sean Christopherson
[not found] ` <2vxzqzjz1x5f.fsf@kernel.org>
2026-08-17 14:37 ` Sean Christopherson
2026-08-18 13:43 ` Pratyush Yadav
2026-08-18 16:02 ` Sean Christopherson
2026-08-21 13:43 ` Pratyush Yadav
2026-08-21 15:34 ` Sean Christopherson
2026-08-18 15:28 ` tarunsahu
2026-08-18 16:10 ` tarunsahu [this message]
2026-07-28 12:11 ` [PATCH v4 06/11] KVM: guest_memfd: Move internal definitions to internal header Tarun Sahu
2026-07-30 18:12 ` Ackerley Tng
2026-08-11 10:31 ` Pratyush Yadav
2026-07-28 12:11 ` [PATCH v4 07/11] KVM: guest_memfd: Add support for freezing mappings Tarun Sahu
2026-07-30 17:46 ` Ackerley Tng
2026-08-10 13:15 ` tarunsahu
2026-07-30 18:12 ` Ackerley Tng
2026-08-10 13:08 ` tarunsahu
2026-08-10 23:44 ` Sean Christopherson
2026-08-18 16:33 ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 08/11] KVM: guest_memfd: Add support for preservation via LUO Tarun Sahu
2026-07-30 18:16 ` Ackerley Tng
2026-08-10 13:20 ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 09/11] docs: liveupdate: Add documentation for VM and guest_memfd preservation Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 10/11] KVM: selftests: Split ____vm_create() and add vm_create_from_fd() Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 11/11] KVM: selftests: Add guest_memfd_preservation_test Tarun Sahu
2026-07-30 18:18 ` Ackerley Tng
2026-08-10 13:22 ` tarunsahu
2026-08-18 16:35 ` tarunsahu
2026-08-11 10:06 ` Pratyush Yadav
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=9huzo6ezv288.fsf@tarunix.c.googlers.com \
--to=tarunsahu@google.com \
--cc=ackerleytng@google.com \
--cc=akpm@linux-foundation.org \
--cc=alexandru.elisei@arm.com \
--cc=andre.przywara@arm.com \
--cc=aneesh.kumar@kernel.org \
--cc=corbet@lwn.net \
--cc=david@kernel.org \
--cc=david@redhat.com \
--cc=dmatlack@google.com \
--cc=fuad.tabba@linux.dev \
--cc=fvdl@google.com \
--cc=graf@amazon.com \
--cc=kexec@lists.infradead.org \
--cc=kvm@vger.kernel.org \
--cc=kvmarm@lists.linux.dev \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=maz@kernel.org \
--cc=michael.roth@amd.com \
--cc=oliver.upton@linux.dev \
--cc=pasha.tatashin@soleen.com \
--cc=pbonzini@redhat.com \
--cc=pratyush@kernel.org \
--cc=rppt@kernel.org \
--cc=sagis@google.com \
--cc=seanjc@google.com \
--cc=skhan@linuxfoundation.org \
--cc=skhawaja@google.com \
--cc=suzuki.poulose@arm.com \
--cc=vannapurve@google.com \
--cc=will@kernel.org \
--cc=yan.y.zhao@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox