Linux Documentation
 help / color / mirror / Atom feed
From: tarunsahu@google.com
To: Sean Christopherson <seanjc@google.com>,
	Pratyush Yadav <pratyush@kernel.org>
Cc: ackerleytng@google.com, fuad.tabba@linux.dev,
	 Andrew Morton <akpm@linux-foundation.org>,
	dmatlack@google.com,  Shuah Khan <skhan@linuxfoundation.org>,
	Jonathan Corbet <corbet@lwn.net>,
	david@redhat.com,  Pasha Tatashin <pasha.tatashin@soleen.com>,
	sagis@google.com,  Paolo Bonzini <pbonzini@redhat.com>,
	Mike Rapoport <rppt@kernel.org>, Alexander Graf <graf@amazon.com>,
	 linux-kselftest@vger.kernel.org, andre.przywara@arm.com,
	michael.roth@amd.com,  linux-kernel@vger.kernel.org,
	linux-mm@kvack.org, will@kernel.org,  vannapurve@google.com,
	maz@kernel.org, fvdl@google.com, kvm@vger.kernel.org,
	 oliver.upton@linux.dev, kvmarm@lists.linux.dev,
	alexandru.elisei@arm.com,  skhawaja@google.com,
	aneesh.kumar@kernel.org, linux-doc@vger.kernel.org,
	 David Hildenbrand <david@kernel.org>,
	yan.y.zhao@intel.com, kexec@lists.infradead.org,
	 suzuki.poulose@arm.com
Subject: Re: [PATCH v4 05/11] KVM: LUO: Support VM preservation across live updates
Date: Tue, 18 Aug 2026 16:10:31 +0000	[thread overview]
Message-ID: <9huzo6ezv288.fsf@tarunix.c.googlers.com> (raw)
In-Reply-To: <anssOQFNZ2LfihbU@google.com>

Sean Christopherson <seanjc@google.com> writes:

> On Tue, Aug 11, 2026, Pratyush Yadav wrote:
>> On Mon, Aug 10 2026, Sean Christopherson wrote:
>> 
>> > On Tue, Jul 28, 2026, Tarun Sahu wrote:
>> >> Register a Live Update Orchestrator (LUO) file handler for KVM VM files
>> >> to serialize and deserialize VM state across kexec live updates.
>> >> 
>> >> Currently, Only VM type (e.g. arch.vm_type on x86) is preserved as part
>> >> of VM preservation.
>> >
>> > Why?
>> >
>> >> On retrieval, kvm_luo_retrieve() recreates the KVM VM file via
>> >> kvm_create_vm_file() and use an atomically incremented ID for the internal
>> >> fdname, as the final fdname assigned by userspace is not yet known during
>> >> retrieval. As this fdname is only used in debugfs infra, This will not break
>> >> any UAPI.
>> >> 
>> >> This infrastructure establishes the foundation for preserving guest_memfd
>> >> instances across live updates, and can be expanded in the future to
>> >> preserve additional VM state.
>> >
>> > Uh, why guest_memfd?  As much as I want to push guest_memfd adoption, it seems
>> > guest_memfd should be the _last_ thing we support, not the first.  As evidenced
>> > by the last two decades, it's very doable to have KVM VMs without guest_memfd,
>> > but it's rather hard to have VMs without vCPUs.
>> 
>> You _can_ preserve vCPUs today using KVM_{GET,SET}_REGS, they just won't
>> run in the background during the reboot.
>
> What about x86 CoCo VMs?  Which are quite literally _the_ reason guest_memfd was
> created in the first place.
>
>> This series can save you from dumping VM memory to disk if it is backed by
>> guest_memfd.
>
> Or to word it another way, one _can_ save guest_memfd, it's just
> slower.
>
>
> My point is that this series needs to provide a _lot_ more information about the
> bigger KVM picture.

Yes, I agree. I will try to layout the plan.

End goal of this series is to preserve VM memory which is backed by
guest_memfd. Not all VM are backed by guest_memfd. So preservation of
KVM (vm_file) is independent of preservation of guest_memfd.
But guest_memefd can not be preserved without preserving the KVM (vm_file).
So I agree to your suggestion:

diff --git virt/kvm/Makefile.kvm virt/kvm/Makefile.kvm
index d047d4cf58c9..e6f098498795 100644
--- virt/kvm/Makefile.kvm
+++ virt/kvm/Makefile.kvm
@@ -13,3 +13,8 @@ kvm-$(CONFIG_HAVE_KVM_IRQ_ROUTING) += $(KVM)/irqchip.o
 kvm-$(CONFIG_HAVE_KVM_DIRTY_RING) += $(KVM)/dirty_ring.o
 kvm-$(CONFIG_HAVE_KVM_PFNCACHE) += $(KVM)/pfncache.o
 kvm-$(CONFIG_KVM_GUEST_MEMFD) += $(KVM)/guest_memfd.o
+
+ifdef CONFIG_LIVEUPDATE
+kvm-y += $(KVM)/kvm_luo.o
+kvm-$(CONFIG_KVM_GUEST_MEMFD) += $(KVM)/guest_memfd_luo.o
+endif

Guest_memfd can't be created alone without struct kvm
(kvm_gmem_create() and KVM_GMEM_CREATE ioctl). This is how
guest_memfd has been designed. I don't want to break this design which
has been accepted upstream after lot of discussion.

So while LUO preserve guest_memfd's data, It just preserve the PFNs
value, few flags that belongs to the guest_memfd.

On restore side in new kernel, These PFNs, flags will be populated
to a _newly_ created guest_memfd. that is how preservation and
retrieval works. So To create this new guest_memfd, We need struct
kvm (vm_file). this vm_file must be the same VM which had this
guest_memfd in old kernel, So we preserve the vm_file, get its TOKEN
preserve with guest_memfd and during restore, we create the guest_memfd
with the same VM (vm_file/struct kvm).

I agree, I did not do good job explaining things in commit message. I
will make sure to update them in next revision.

> For those of us that are on the very fringes of live update,
> it's practically impossible to review because, to us, it seems very arbitrary.
>
> The part that's especially confusing is the saving of the VM type.  That comes
> straight from userspace, so it's super bizarre to automatically save/restore that,
> but nothing else.

Like, guest_memfd needs struct kvm to create itself. struct kvm (vm_file) needs
vm_type to create itself (kvm_create_vm() or KVM_CREATE_VM IOCTL). LUO
does not provde functionality to pass any subsystem specific arguments. So
vm_type needs to be preserved even though, userspace is aware about it.

So Why do we preserve only vm_type:  To keep things simple for this
series, As target is guest_memfd.

Currently I dont have discreet plan on what else will
,in future, be needed to be preserved. Which I agree not a absoulute right
way to approach. I will layout a rough plan on KVM side preservation.

Having this need for backward compatiblity: I responded here:
https://lore.kernel.org/all/9huzv797v45k.fsf@tarunix.c.googlers.com/



>
>> >> +KVM LIVE UPDATE
>> >> +M:	Pasha Tatashin <pasha.tatashin@soleen.com>
>> >> +M:	Mike Rapoport <rppt@kernel.org>
>> >> +M:	Pratyush Yadav <pratyush@kernel.org>
>> >> +R:	Tarun Sahu <tarunsahu@google.com>
>> >> +L:	kexec@lists.infradead.org
>> >> +L:	kvm@vger.kernel.org
>> >> +S:	Maintained
>> >> +T:	git git://git.kernel.org/pub/scm/linux/kernel/git/liveupdate/linux.git
>> >
>> > NAK on taking changes through a different tree.  This is KVM code, period.
>> >
>> > In general, I'm skeptical of the dedicated MAINTAINERS entry.  It's extremely
>> > difficult to tell since this series is little more than a skeleton (either that
>> > or liveupdate is way simpler that I was expecting), but I suspect that maintaining
>
> ...
>
>> > E.g. the LUO APIs seem pretty straightforward; I assume the bulk of the complexity
>> > is going to be in knowing what to save/restore, and how, which is much more about
>> > KVM than it is about liveupdate.
>> 
>> I think it is fine if you want to take these changes through the KVM
>> tree, but I would like live update maintainers to be listed as reviewers
>> at least.
>
> Why not simply add a file pattern match to the LIVE UPDATE entry?
>
> diff --git MAINTAINERS MAINTAINERS
> index 8014b9f8253e..2eb57b22c37f 100644
> --- MAINTAINERS
> +++ MAINTAINERS
> @@ -15052,8 +15052,8 @@ F:      include/linux/liveupdate.h
>  F:     include/uapi/linux/liveupdate.h
>  F:     kernel/liveupdate/
>  F:     lib/tests/liveupdate.c
> -F:     mm/memfd_luo.c
>  F:     tools/testing/selftests/liveupdate/
> +N:     [^a-z]luo
>  
>  LLC (802.2)
>  L:     netdev@vger.kernel.org
>
>> At the same time, I also keep being (pleasantly)
>> surprised at preservation being relatively simple. For example, the code
>> to preserve a shmem file (via memfd) is roughly 600 lines, a big chunk
>> of which is comments. The code of course has some limitations, but it is
>> good enough for use in production.
>> 
>> For one, we care about ABI breakages and versioning. 
>
> Which is amusing to me because that implies KVM does not, and I would hazard to
> guess that KVM has the biggest ABI surface of any subsystem in the kernel by a
> country mile (though I'm probably wildly underestimating the effective ABI surface
> of filesystems).
>
>> The serialized state is a part of live update ABI and changes to it should be
>> ACKed by us. 
>
> Meh, "Don't break userspace" is a universal rule in the kernel, I genuinely don't
> see why liveupdate needs special treatment.
>
>> For another, how the file handlers interact with their dependencies can
>> affect the behaviour that VMMs observe. Those changes should also pass by
>> some live update eyes.
>
> Perhaps in the short term, but IMO, that's not a winning strategy in the long
> term.  From my perspective, that like saying the PAGE CACHE maintainers should
> review every usage of the filemap APIs, because how the APIs are used impacts
> the page cache and affects userspace-visible behavior.  There are myriad analogies
> like that throughout the kernel.
>
> Yes, liveupdate is new and shiny, but IMO for it to be successful and maintainable,
> it needs to be treated like any other core infrastructure in the kernel, not a
> special snowflake whose details are known only by a handful of people.  Because
> I think it's likely liveupdate goes one of two ways: either liveupdate becomes a
> very niche thing that is used sparingly throughout the kernel, or it becomes a
> broadly used feature that is supported by many filesystems and subsystems.
>
> If liveupdate is relegated to niche status, then it probably isn't going to see
> a significant amount of ongoing development, at which point the folks working on
> liveupdate will naturally migrate to other projects, and maintenance will largely
> be left to subsystem maintainers.
>
> If liveupdate is broadly used, then having a single group of people maintain
> every subsystem's usage won't scale, and maintenance will again largely fall on
> the shoulder of subsystem maintainers.  Which is totally fine and working as
> intended, because that's exactly what subystem maintainers are signing up for
> by merging support for liveupdate.

  parent reply	other threads:[~2026-08-18 16:10 UTC|newest]

Thread overview: 47+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-28 12:11 [PATCH v4 00/11] liveupdate: kvm: Guest_memfd preservation Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 01/11] liveupdate: Add LIVEUPDATE_GUEST_MEMFD config option Tarun Sahu
2026-08-10 22:58   ` Sean Christopherson
2026-08-11 13:26     ` tarunsahu
2026-08-11 14:48       ` Sean Christopherson
2026-08-18 16:11         ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 02/11] KVM: Introduce kvm_create_vm_file() helper Tarun Sahu
2026-07-30 17:36   ` Ackerley Tng
2026-08-10 10:14     ` tarunsahu
2026-08-10 23:05   ` Sean Christopherson
2026-07-28 12:11 ` [PATCH v4 03/11] KVM: Export kvm_uevent_notify_vm_create() Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 04/11] KVM: Track weak reference to vm_file in struct kvm Tarun Sahu
2026-08-10 23:23   ` Sean Christopherson
2026-08-18 16:29     ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 05/11] KVM: LUO: Support VM preservation across live updates Tarun Sahu
2026-08-10 23:42   ` Sean Christopherson
2026-08-11 11:31     ` Pratyush Yadav
2026-08-11 14:05       ` Sean Christopherson
2026-08-12 13:45         ` Pratyush Yadav
2026-08-12 15:17           ` Sean Christopherson
     [not found]             ` <2vxzqzjz1x5f.fsf@kernel.org>
2026-08-17 14:37               ` Sean Christopherson
2026-08-18 13:43                 ` Pratyush Yadav
2026-08-18 16:02                   ` Sean Christopherson
2026-08-21 13:43                     ` Pratyush Yadav
2026-08-21 15:34                       ` Sean Christopherson
2026-08-18 15:28               ` tarunsahu
2026-08-18 16:10         ` tarunsahu [this message]
2026-07-28 12:11 ` [PATCH v4 06/11] KVM: guest_memfd: Move internal definitions to internal header Tarun Sahu
2026-07-30 18:12   ` Ackerley Tng
2026-08-11 10:31     ` Pratyush Yadav
2026-07-28 12:11 ` [PATCH v4 07/11] KVM: guest_memfd: Add support for freezing mappings Tarun Sahu
2026-07-30 17:46   ` Ackerley Tng
2026-08-10 13:15     ` tarunsahu
2026-07-30 18:12   ` Ackerley Tng
2026-08-10 13:08     ` tarunsahu
2026-08-10 23:44   ` Sean Christopherson
2026-08-18 16:33     ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 08/11] KVM: guest_memfd: Add support for preservation via LUO Tarun Sahu
2026-07-30 18:16   ` Ackerley Tng
2026-08-10 13:20     ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 09/11] docs: liveupdate: Add documentation for VM and guest_memfd preservation Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 10/11] KVM: selftests: Split ____vm_create() and add vm_create_from_fd() Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 11/11] KVM: selftests: Add guest_memfd_preservation_test Tarun Sahu
2026-07-30 18:18   ` Ackerley Tng
2026-08-10 13:22     ` tarunsahu
2026-08-18 16:35       ` tarunsahu
2026-08-11 10:06     ` Pratyush Yadav

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=9huzo6ezv288.fsf@tarunix.c.googlers.com \
    --to=tarunsahu@google.com \
    --cc=ackerleytng@google.com \
    --cc=akpm@linux-foundation.org \
    --cc=alexandru.elisei@arm.com \
    --cc=andre.przywara@arm.com \
    --cc=aneesh.kumar@kernel.org \
    --cc=corbet@lwn.net \
    --cc=david@kernel.org \
    --cc=david@redhat.com \
    --cc=dmatlack@google.com \
    --cc=fuad.tabba@linux.dev \
    --cc=fvdl@google.com \
    --cc=graf@amazon.com \
    --cc=kexec@lists.infradead.org \
    --cc=kvm@vger.kernel.org \
    --cc=kvmarm@lists.linux.dev \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=maz@kernel.org \
    --cc=michael.roth@amd.com \
    --cc=oliver.upton@linux.dev \
    --cc=pasha.tatashin@soleen.com \
    --cc=pbonzini@redhat.com \
    --cc=pratyush@kernel.org \
    --cc=rppt@kernel.org \
    --cc=sagis@google.com \
    --cc=seanjc@google.com \
    --cc=skhan@linuxfoundation.org \
    --cc=skhawaja@google.com \
    --cc=suzuki.poulose@arm.com \
    --cc=vannapurve@google.com \
    --cc=will@kernel.org \
    --cc=yan.y.zhao@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox