From: Jason Gunthorpe <jgg@nvidia.com>
To: Sean Christopherson <seanjc@google.com>
Cc: Pratyush Yadav <pratyush@kernel.org>,
Tarun Sahu <tarunsahu@google.com>,
ackerleytng@google.com, fuad.tabba@linux.dev,
Andrew Morton <akpm@linux-foundation.org>,
dmatlack@google.com, Shuah Khan <skhan@linuxfoundation.org>,
Jonathan Corbet <corbet@lwn.net>,
david@redhat.com, Pasha Tatashin <pasha.tatashin@soleen.com>,
sagis@google.com, Paolo Bonzini <pbonzini@redhat.com>,
Mike Rapoport <rppt@kernel.org>, Alexander Graf <graf@amazon.com>,
linux-kselftest@vger.kernel.org, andre.przywara@arm.com,
michael.roth@amd.com, linux-kernel@vger.kernel.org,
linux-mm@kvack.org, will@kernel.org, vannapurve@google.com,
maz@kernel.org, fvdl@google.com, kvm@vger.kernel.org,
oliver.upton@linux.dev, kvmarm@lists.linux.dev,
alexandru.elisei@arm.com, skhawaja@google.com,
aneesh.kumar@kernel.org, linux-doc@vger.kernel.org,
David Hildenbrand <david@kernel.org>,
yan.y.zhao@intel.com, kexec@lists.infradead.org,
suzuki.poulose@arm.com
Subject: Re: [PATCH v4 05/11] KVM: LUO: Support VM preservation across live updates
Date: Thu, 10 Sep 2026 14:47:41 -0300 [thread overview]
Message-ID: <20260910174741.GG3968357@nvidia.com> (raw)
In-Reply-To: <aqLU653FXjJjhbop@google.com>
On Thu, Sep 10, 2026 at 09:03:55AM -0700, Sean Christopherson wrote:
> From that perspective, what I am proposing actually goes a step further. I'm
> saying don't commit to supporting *any* specific versions in upstream. Express
> the feature requirements in the serialization payloads, and let userspace sort
> out what kernels are compatible based on their actual usage. The kernel may need
> to provide additional discovery mechanisms, e.g. so that userspace can probe to
> see what is supported, but discovery is usually fairly simple to implement and
> maintain.
I'm not so fussed about the version number scheme itself. Alot of
different schemes have been proposed over the last year and half, and
some were very complicated. Nobody came with a more granular proposal
that also didn't have alot of complexity attached to it.
FWWI this started out with creating device tree fragments with a full
YAML schema, and the same kind of perfect ABI like you are talking
about. Yet it did not figure out how to make it discoverable, it was
super inefficient and very hard to code for.
It has to be discoverable from the ELF, restrictable on the export
side, easy on maintainers, and not complex to implement.
> No, the subsystem just needs to make sure that it serializes its data using the
> defined ABI (where ABI here means the format of the payload and the meaning of
> any flags in the header). That should be *easier* for maintainers to handle than
> trying to support arbitrary versions, because it eliminates subjectivity and
> having to make judgment calls or remember magic version numbers.
I'm pretty sure I have something like PTSD from ABI definitions
adventures on the uAPI side. :( Please don't call it easier!
I really don't want more of those in the kernel process. I think
Linus's non-stable-api-nonsense is really a good thing for community
health.
With the simplification that upstream supports only one version at
once, a simple "id" to represent that ABI was the simplest, easiest on
maintainers thing. There is never a fight. Someone has a new idea,
great no worries, change the ID. Done.
I guess I should say my perspective is to prioritize not burdening the
maintainers.
> > I was told KVM had the smallest luo footprint of everything, so
> > perhaps your perspective is different.
>
> Only because KVM already has a massive ABI surface for save/restore.
Yeah, you are lucky, other subsystems haven't done that. Honestly, was
it easy to create?
> If you want to convince me that magic version numbers are the right
> approach, then show me how KVM's existing save/restore support would
> be made "better" and easier to maintain by throwing away all of
> KVM's save/restore uAPI and replacing it with a versioning scheme.
I don't know about KVM, but how do I manage something like serializing
an iommu page table?
I'm replacing all the iommu page table code. It behaves
differently. It supports different things that old kernels don't
understand. iommu page tables have to be under continuous active DMA
during kexec, so they cannot be serialized.
Setting the old stuff as V1 and the new suff as V2 is so easy and is
unconditionally correct. To my point about supporting only stable
branches this achieves it with minimal maintainer effort.
Yes, we could do some comprehensive analysis and try to determine
everything that changed and make micro feature bits, and some thing to
map the current layout to those bits and then hope all of this is
correct.. But that's *a lot* of work, probably will have gaps since it
will never be tested. Seriously, why bother? Explain to me why I
should spend my time on this and who benifits?
IMHO your KVM example has a much clearer answer to that question. It
is UAPI so you have to do it anyhow, and you have a robust open source
ecosystem with alot of VMMs that will consume it. Yeah, I'm on board,
makes sense.
> Have to, or choose to? I have a very, very hard time believing that
> it's infeasible to define a serialization format that is decoupled
> from kernel internals.
Reduced perfectly everything should become grounded in either HW or
kernel UAPI definitions. Things are not perfect, the kernel has
limitations, it doesn't support every HW feature, stuff leaks
in.. Like my page table example, the new stuff supports ARM CONT, the
old stuff does not. That is kexec ABI breaking and is only happening
because of kernel internal details. In the page table work alone there
are probably ~10-20 micro features like that that would have to be
identified and delt with. Frankly I'm not confident I could even
capture them all.
Also there is the push to make kexec downtime lower. That puts
pressure to just retain kernel things exactly as is, and maybe even
keep kernel memory as-is. ie a sloppy job serializing to get better
performance.
> Only if the relevant subsystem defined a poor save/restore ABI in the first place.
I think KVM is the only subsystem that has save/restore :)
Jason
next prev parent reply other threads:[~2026-09-10 17:48 UTC|newest]
Thread overview: 65+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-28 12:11 [PATCH v4 00/11] liveupdate: kvm: Guest_memfd preservation Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 01/11] liveupdate: Add LIVEUPDATE_GUEST_MEMFD config option Tarun Sahu
2026-07-28 12:22 ` sashiko-bot
2026-07-30 18:06 ` Ackerley Tng
2026-08-10 10:13 ` tarunsahu
2026-08-10 22:58 ` Sean Christopherson
2026-08-11 13:26 ` tarunsahu
2026-08-11 14:48 ` Sean Christopherson
2026-08-18 16:11 ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 02/11] KVM: Introduce kvm_create_vm_file() helper Tarun Sahu
2026-07-30 17:36 ` Ackerley Tng
2026-08-10 10:14 ` tarunsahu
2026-08-10 23:05 ` Sean Christopherson
2026-08-11 13:27 ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 03/11] KVM: Export kvm_uevent_notify_vm_create() Tarun Sahu
2026-07-28 12:26 ` sashiko-bot
2026-07-30 17:43 ` Ackerley Tng
2026-08-06 1:14 ` Sean Christopherson
2026-08-10 12:55 ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 04/11] KVM: Track weak reference to vm_file in struct kvm Tarun Sahu
2026-08-10 23:23 ` Sean Christopherson
2026-08-18 16:29 ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 05/11] KVM: LUO: Support VM preservation across live updates Tarun Sahu
2026-07-28 12:27 ` sashiko-bot
2026-08-10 23:42 ` Sean Christopherson
2026-08-11 11:31 ` Pratyush Yadav
2026-08-11 14:05 ` Sean Christopherson
2026-08-12 13:45 ` Pratyush Yadav
2026-08-12 15:17 ` Sean Christopherson
2026-08-15 10:43 ` Pratyush Yadav
2026-08-17 14:37 ` Sean Christopherson
2026-08-18 13:43 ` Pratyush Yadav
2026-08-18 16:02 ` Sean Christopherson
2026-08-21 13:43 ` Pratyush Yadav
2026-08-21 15:34 ` Sean Christopherson
2026-09-03 15:23 ` tarunsahu
2026-09-05 1:19 ` Jason Gunthorpe
2026-09-10 16:03 ` Sean Christopherson
2026-09-10 17:47 ` Jason Gunthorpe [this message]
2026-08-18 15:28 ` tarunsahu
2026-08-18 16:10 ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 06/11] KVM: guest_memfd: Move internal definitions to internal header Tarun Sahu
2026-07-30 18:12 ` Ackerley Tng
2026-08-11 10:31 ` Pratyush Yadav
2026-07-28 12:11 ` [PATCH v4 07/11] KVM: guest_memfd: Add support for freezing mappings Tarun Sahu
2026-07-28 12:20 ` sashiko-bot
2026-07-30 17:46 ` Ackerley Tng
2026-08-10 13:15 ` tarunsahu
2026-07-30 18:12 ` Ackerley Tng
2026-08-10 13:08 ` tarunsahu
2026-08-10 23:44 ` Sean Christopherson
2026-08-18 16:33 ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 08/11] KVM: guest_memfd: Add support for preservation via LUO Tarun Sahu
2026-07-28 12:23 ` sashiko-bot
2026-07-30 18:16 ` Ackerley Tng
2026-08-10 13:20 ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 09/11] docs: liveupdate: Add documentation for VM and guest_memfd preservation Tarun Sahu
2026-07-28 12:21 ` sashiko-bot
2026-08-10 13:21 ` tarunsahu
2026-07-28 12:11 ` [PATCH v4 10/11] KVM: selftests: Split ____vm_create() and add vm_create_from_fd() Tarun Sahu
2026-07-28 12:11 ` [PATCH v4 11/11] KVM: selftests: Add guest_memfd_preservation_test Tarun Sahu
2026-07-30 18:18 ` Ackerley Tng
2026-08-10 13:22 ` tarunsahu
2026-08-18 16:35 ` tarunsahu
2026-08-11 10:06 ` Pratyush Yadav
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260910174741.GG3968357@nvidia.com \
--to=jgg@nvidia.com \
--cc=ackerleytng@google.com \
--cc=akpm@linux-foundation.org \
--cc=alexandru.elisei@arm.com \
--cc=andre.przywara@arm.com \
--cc=aneesh.kumar@kernel.org \
--cc=corbet@lwn.net \
--cc=david@kernel.org \
--cc=david@redhat.com \
--cc=dmatlack@google.com \
--cc=fuad.tabba@linux.dev \
--cc=fvdl@google.com \
--cc=graf@amazon.com \
--cc=kexec@lists.infradead.org \
--cc=kvm@vger.kernel.org \
--cc=kvmarm@lists.linux.dev \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=maz@kernel.org \
--cc=michael.roth@amd.com \
--cc=oliver.upton@linux.dev \
--cc=pasha.tatashin@soleen.com \
--cc=pbonzini@redhat.com \
--cc=pratyush@kernel.org \
--cc=rppt@kernel.org \
--cc=sagis@google.com \
--cc=seanjc@google.com \
--cc=skhan@linuxfoundation.org \
--cc=skhawaja@google.com \
--cc=suzuki.poulose@arm.com \
--cc=tarunsahu@google.com \
--cc=vannapurve@google.com \
--cc=will@kernel.org \
--cc=yan.y.zhao@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox