From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) (using TLSv1 with cipher DHE-RSA-AES256-SHA (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 8D172C61DD3 for ; Thu, 3 Sep 2026 15:23:08 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 9A7586B008A; Thu, 3 Sep 2026 11:23:06 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 97EDB6B008C; Thu, 3 Sep 2026 11:23:06 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 895F96B0092; Thu, 3 Sep 2026 11:23:06 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0016.hostedemail.com [216.40.44.16]) by kanga.kvack.org (Postfix) with ESMTP id 570976B008A for ; Thu, 3 Sep 2026 11:23:05 -0400 (EDT) Received: from smtpin23.hostedemail.com (lb01a-stub [10.200.18.249]) by unirelay06.hostedemail.com (Postfix) with ESMTP id C64FCA4743 for ; Thu, 3 Sep 2026 15:23:04 +0000 (UTC) X-FDA: 85172819088.23.492A951 Received: from mail-ej1-f70.google.com (mail-ej1-f70.google.com [209.85.218.70]) by imf04.hostedemail.com (Postfix) with ESMTP id 1467740004 for ; Thu, 3 Sep 2026 15:23:02 +0000 (UTC) Authentication-Results: imf04.hostedemail.com; dkim=pass header.d=google.com header.s=20251104 header.b=tNiHO54A; spf=pass (imf04.hostedemail.com: domain of 31ZCZagkKCLgrYpslqYfsemmejc.amkjglsv-kkitYai.mpe@flex--tarunsahu.bounces.google.com designates 209.85.218.70 as permitted sender) smtp.mailfrom=31ZCZagkKCLgrYpslqYfsemmejc.amkjglsv-kkitYai.mpe@flex--tarunsahu.bounces.google.com; dmarc=pass (policy=reject) header.from=google.com ARC-Seal: i=1; a=rsa-sha256; d=hostedemail.com; s=arc-20220608; cv=none; t=1788448983; b=HTegv+nRBUHf41RJ2wW5MHMCuwKooQMKeHTXaEzaU0orXSccuVq9JfpeHtNZSXiACPkIPd 2tcep53JwOIDcmqiupUCvGnHTugwuEvGBqEDauDrdUjnOElPJNyNcml4I3liavZP5v3+PM XZKWKzj6p17tNLVQf0tw+2AXTzP+EI0= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1788448983; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=GmFek9macUO7lb98mH13+DwSxNV0cpAjOXbf35SyPVg=; b=2XB28JAlbWueZWh1JUpl8L/L+ePLCKsXuDuR8lhjjtXlBUqYeedJr8Uhz1PMRx7hfrEOev 3jhOCNDHQuVDa2IxXDp1/fgb365/VktynQT+omluqmwMMHoTTL/xh9rYc6CYWkQ69qhkui zIkxT0GV3bBkEtDVpG6hrf82wVSCkJ4= ARC-Authentication-Results: i=1; imf04.hostedemail.com; dkim=pass header.d=google.com header.s=20251104 header.b=tNiHO54A; spf=pass (imf04.hostedemail.com: domain of 31ZCZagkKCLgrYpslqYfsemmejc.amkjglsv-kkitYai.mpe@flex--tarunsahu.bounces.google.com designates 209.85.218.70 as permitted sender) smtp.mailfrom=31ZCZagkKCLgrYpslqYfsemmejc.amkjglsv-kkitYai.mpe@flex--tarunsahu.bounces.google.com; dmarc=pass (policy=reject) header.from=google.com Received: by mail-ej1-f70.google.com with SMTP id a640c23a62f3a-c1f3117bc6bso190778166b.0 for ; Thu, 03 Sep 2026 08:23:02 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1788448981; x=1789053781; darn=kvack.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=GmFek9macUO7lb98mH13+DwSxNV0cpAjOXbf35SyPVg=; b=tNiHO54AlFDA6VaevMvH5UaoD0Fl2Cq6myKexIfbnJxQiZ2hEeC17X17RcvD/7dQ6n aLzbCePFjtJTYJfiJklKZSNL4Zgd3psnHtJTQooP8ozS/ZUU6JtPKOM64pStnGKQuCep dUtCe+8OzTVZiRmugn0+Lyq4YHdfsGlAmkzG12/fJ0Bbcr8cchXMzn5wyXd7lY7h7mTH XB4sE9KxbZ/W+oLnYpLL+QSAZoDBO1BizBEU+nkCXmDVin5/js1ZQFUKmqtXm274mrYJ 9lWH1mgFz3lY+dBWI2076luqGwPKA2Ekh+O6UK5Gkpc0OycUAhYuKcnRxy7WGJvEKbPB /jaw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788448981; x=1789053781; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=GmFek9macUO7lb98mH13+DwSxNV0cpAjOXbf35SyPVg=; b=J++7t7zn6Jbkb9FpWiMlDB640/audLE7IqUdbXz4YZ59F4psv4iz9P1r98GojfpKcr NxQyyY464KI44wb/cNcu+6esYEqvg5eFUt5ZSzZ8uNB67MQePTWEpG5azPVCk3f6Y2Oe /p+z1zBeHfavApqwa+xwUrfO9JqCv5FqyB4MgX4xWYcsJseio0jhu4fBUGAycgU8dYDd KzvYQqZ/lmZ8B0/9FI98BeUtNXvJwLZjiltVINdreTyCA6cYwgw4jyV5sLyzn4AuRdYL iJper/LDh440qOtonecBxPah4IprXRHlU65lHVZoeBsbwzzcHGcCo9VIQbTLzPtBu/8C 4RZQ== X-Forwarded-Encrypted: i=1; AKwUvByuFe5fYe/jWYE6JApFrR6M55XEoXcoG5CH71O53uu/yzBnCMovWGbDLCFSTVCGxMwvCb+aZfU8aA==@kvack.org X-Gm-Message-State: AFuF++kjAgJlEpTGF00p7Vyth39b0/VLynRmOrb9HXY9fxR4On+O8JcX OkxENeW0Yuz1tqj26q5FGW6T32doAtxbCoXvK2kSf2pUuvSrBhnw6KJaIjjq5YixnWoDDoVi3Ud vtEEbzVzSNZgdp7jY1w== X-Received: from ejctj4.prod.google.com ([2002:a17:907:c244:b0:c25:127e:8e75]) (user=tarunsahu job=prod-delivery.src-stubby-dispatcher) by 2002:a17:907:d508:b0:c16:13e7:fd63 with SMTP id a640c23a62f3a-c260436b4e0mr108793366b.0.1788448981126; Thu, 03 Sep 2026 08:23:01 -0700 (PDT) Date: Thu, 03 Sep 2026 15:23:00 +0000 In-Reply-To: Mime-Version: 1.0 References: <2vxzpkzo51wg.fsf@kernel.org> <2vxzik5f311e.fsf@kernel.org> <2vxzqzjz1x5f.fsf@kernel.org> <2vxzy0e3zgqz.fsf@kernel.org> <2vxzy0dzy4gn.fsf@kernel.org> Message-ID: <9huzv78m2wbv.fsf@tarunix.c.googlers.com> Subject: Re: [PATCH v4 05/11] KVM: LUO: Support VM preservation across live updates From: tarunsahu@google.com To: Sean Christopherson , Pratyush Yadav Cc: ackerleytng@google.com, fuad.tabba@linux.dev, Andrew Morton , dmatlack@google.com, Shuah Khan , Jonathan Corbet , david@redhat.com, Pasha Tatashin , sagis@google.com, Paolo Bonzini , Mike Rapoport , Alexander Graf , linux-kselftest@vger.kernel.org, andre.przywara@arm.com, michael.roth@amd.com, linux-kernel@vger.kernel.org, linux-mm@kvack.org, will@kernel.org, vannapurve@google.com, maz@kernel.org, fvdl@google.com, kvm@vger.kernel.org, oliver.upton@linux.dev, kvmarm@lists.linux.dev, alexandru.elisei@arm.com, skhawaja@google.com, aneesh.kumar@kernel.org, linux-doc@vger.kernel.org, David Hildenbrand , yan.y.zhao@intel.com, kexec@lists.infradead.org, suzuki.poulose@arm.com Content-Type: text/plain; charset="UTF-8" X-Rspam-User: X-Rspamd-Server: rspam07 X-Rspamd-Queue-Id: 1467740004 X-Stat-Signature: o6ftsdqkszfjbo48eytyiposu3scjypi X-HE-Tag: 1788448982-482517 X-HE-Meta: U2FsdGVkX1/5juT/VVR+4lcbZF5GYc/jeoRspCUu0sEyZtiNJDeoBET1DrXsIz0iMdlI9RA2geAe5R4TvgJnhyqTh4AWud6ZzLIYBO+XuegZTjPfoJiUwpy9FyfX4nmNJeZQIh/f488qR6OHAgiHZJzeIUia1Zrb1K34NJA24rrAYggmDPHFC1I9LRjZ/kMNmv3z2jTTvQHncQ7PglswesiQ2BwaNpbyqIroCK5aRH3h5n/uUD3/Z8LJVaz5PDlIaGvDzjrbG/H7srLKY3OaRX/210C9/XJTJCw1uCpGwEle54cseVaHk/XW/bQiM5Qg2TmCys94o1OuO9LdWZxCMQFugp9X1RvD1YzWEB6Gq09KqKM3/pEXbn/9DcSSHMdr7j7RoAeQiqsAAVn/iscuaU7LRI5zS69/1Nzld959Ys7gqVur1gm4MUTKqaP/ATv1hGEEmOeRSDpFpPyRt0AiySozX0P1xSs9LQx1YHMpfmneWHGMXAWiHm4udV9+G+zi1liw7t1cjnbTuWl390aXk3Cnpm5fHGS0YesBe3svG0G9uqA+ZtPkvltlaE9957hxq4MpPlTw50it6uLXcNYmfw2gauhAf+QfaqTNUk4g9dhnHtYS9M+8Ym8+6Su61TQleWrpt7wbGI4HWdD65S7bCWICR8mrqgeAa1+NpBSOkFaJ8w5Rf5J3iDx4FpALipzgImUcFteFDce2MtpQJ+mKamwG1WZti3/b5sdybetswMmSHfFduleT4VSQ/zS0gq0lMF8oiL5cnXb7qQORF92fZxIaqBKIdfcnk8EnP1E6kX1+/Cj1yod8HJSuo/l7plpid8GxGZqtHt+aashnmtBt6/CO9pZblnpCfWAghIrxdl9L3K80/bckkkNBo5JVPVF+c+wMMd9HC+D84mfSndj0vfzK+yMHXWdAd6YOZCDXvKpJo7FP+tSUUqmkEdOb2+HNSjm9WDE3PMYYWZpjc8J qHyBT5mo AbTwFBuyfpJhTUld/EjsEPG/A4RriwNFZKwXO9UVaZA5hO6KeGCKqD7c5+y88znIWq1ttmD9TYZhHoB050I6D3mLh8STflv/gHQ3XeDfujeRnTFWcV5TjP59XGZA6AUdDcXG9g2tTPkHYCWtuVlxsLwmj+yY9Ao5uWaWFQnSu+FRiPhHqTkAjFc/KrCN/leqUp6FICf3Z0EAKbDL2m6llyXRysk9FQ6dAQteQYUL5EdmN6x+CbOaR6eAwA5wwRWGC/V0ALYnjYoi2i8e/vJsuxg5OI/vQgYyyGTyuryUSouHhZ0PTb82oFANvDG9F+911cDhHSZp/SbASi5CaNLeltVcGmURY+a6paZqGfdw1rCs07DGX3WkRsXTwhfkfeXQx0/bc7vpf+AA3f237wHNYnjyTTvVmscs5UHAyYIa9/rYGpXrDqWZZUkEDP4KGOq6EmZ6F9sH4gJs6ziOHbXgcCkypkLLbMo3GFHe5G0rvUPuk9R67sGos7f9c6TijgNAiufSgmZbXht8ioqC4Xdz+xlGMPTqRSbDr86Z7WNz/71pBkbMj5CGN+UwSrBHxEO5Vw7Xh/3fSEREZknky5CCoEwHcEk8cJPZSTKU1GDYnjLk+buk= Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: Sean Christopherson writes: > On Fri, Aug 21, 2026, Pratyush Yadav wrote: >> Hi Sean, >> >> On Tue, Aug 18 2026, Sean Christopherson wrote: >> >> Does this idea of "backwards compatibility" sound acceptable to you, at >> >> least at a high level? >> > >> > No. >> > >> > It's probably fine for Google and other large companies that tightly control their >> > kernels and use cases, and have the resources to juggle the resulting complexity, >> > e.g. have kernel engineers on staff to track feature and dependencies, coordinate >> > and plan kernel upgrades, etc. >> > >> > It's not acceptable for upstream, where downstream consumers often run a distro >> > kernel, have much more varied use cases, and don't always have a horde of kernel >> > engineers on staff to help them thread the needle you describe above. And if >> > supporting live update as a general feature for all users of the kernel isn't >> > being factored into design considerations, then that needs to change, otherwise >> > this is all dead in the water. >> > >> > I also don't see the point. Maintaining a rigid save/restore ABI is annoying, >> > but it's not _hard_ (or at least, not _that_ hard), especially if there's a set >> > of well-documented best known practices that subsystems can follow, e.g. so that >> > individual subsystems don't need to learn painful lessons first-hand. I genuinely >> > believe that maintaining the version hell you describe above would be more costly >> > in the long run than simply committing to full backwards compatibility within a >> > given subsystem. I can imagine that enumerating what subsystems' information is >> > in the payload will require a different scheme, but for a given subsystem, I don't >> > see any reason to aim for anything less than full backwards compatibility. >> >> Let's say for argument's sake that we commit for a fully stable >> backwards compatible ABI. Even then, you have to deal with multiple ABI >> versions. >> >> Live update's ABI is more complex compared to KVM's save/restore ABI. >> For the KVM save/restore uAPI, you are largely describing architectural >> state like CPU registers, etc. These things don't evolve as fast and >> more or less stay the same. > > Right, because nothing meaningful has changed in any architecture in the 20+ > years since KVM has provided save/restore support, whereas guest_memfd looks > nothing like it did when it was introduced three years ago. > >> Live update needs to describe the state of kernel objects. These are >> more complex > > LOL, you might be the first to claim x86 virtualization isn't all that complex. > >> and evolve faster. > > The speed at which things change doesn't automatically mean we shouldn't strive > for backwards compatibility. Yes, providing backwards compatibility requires > additional care and planning, and over time *might* lead to an ABI that is > difficult to maintain. But IMO, that just makes it all the more important to > get the design right the first time, not that we shouldn't even try because it's > hard. > >> For example, say you merge guest_memfd preservation today. Some time >> later, someone comes up with a more efficient data structure to track >> the folios in the file. You _have_ to make a backwards-incompatible ABI >> change to use this data structure. > > Only if those details bleed into the ABI/contract. I actually have a concrete > KVM (well, virtualization) example for this. > > Intel's VMX architecture disallows direct memory accesses to the VMCS, and instead > requires software to access the VMCS via dedicated ISA, using architectural encoding > numbers to reference VMCS fields. I.e. VMX decouples how data is stored in memory > (the data structures) from the ABI/contract with software (VMCS field encodings). > > This allows Intel to optimize the data structures to be more efficient and performant > for each microarchitecture based on the features and properties of each uarch, all > without breaking backwards/forwards compatibility with software. > > My favorite esoteric example is AR_BYTES packing. For Haswell, Intel added an > optimization in ucode to allow saving/loading segment register state in a single > uop (IIRC). The optimization was especially valuable for virtualization as it > shaved cycles off the VM-Enter/VM-Exit hot paths. A key piece of the optimization > was it required the AR_BYTES metadata to be stored in 16 bits, but existing CPUs > stored AR_BYTES using 32 bits in an "unpacked" format. > > Fortunately, because the in-memory representation was decoupled from the contract > with software, Intel could pack AR_BYTES into 16 bits for Haswell+ and pack/unpack > the data on VMWRITE/VMREAD, so that the format presented to software remained > unchanged. > > Does VMX's decoupling of the in-memory represntation of a VMCS have downsides? > Absolutely. Most notably, it incurs extra complexity (in software and hardware) > to achieve comparable performance to directly accessible data structures (AMD's > VMCB and Hyper-V's eVMCS) for nested virtualization. And I'm sure it has placed > contraints on Intel's designs, and obviously introduces complexity into the > overall system by adding a layer of indirection. But IMO the VMX architecture > has been a huge win overall for Intel. > >> Or say you add a new memory backend (like the HugeTLB patches in >> flight). That likely will need a different ABI to describe the state of >> the guest_memfd. >> >> So you will end up with multiple ABI versions that aren't always >> backwards compatible. > > No, you end up with *features* that aren't backwards compatible. I can't imagine > anyone will argue that we should never add new features because then we can't > rollback to an older kernel. > > But adding a new feature shouldn't break the existing ABI. E.g. adding support > for HugeTLB in guest_memfd shouldn't prevent rolling back to an older kernel when > the HugeTLB functionality isn't being used. > > Using AMD's VMCB and Intel's VMCS as examples, literally every major new AMD/Intel > uarch extends the VMC{B,S} in some way, but without fail it's always done in a way > that is backwards compatible with existing software. I.e. AMD and Intel ship new > features, but existing software continues to work, and VMs continue to be migratable > across CPU generations[*], with the obvious restriction that migrating a VM using > a feature introduced on generation N to a generation N-1 CPU isn't a smart idea. > > [*] There are exceptions. E.g. Intel removed MPX, and so VMs with MPX can't be > migrated to newer CPUs. Migrating between CPUs with different MAXPHYADDR is > sketchy (and simply not done by some CSPs) because neither AMD nor Intel > virtualizes MAXPHYADDR. But those exceptions are absolutely Big Deals that > undergo significant scrutiny, from all parties involved. > >> If you refuse that idea too, then KVM live update will be dead in the >> water for a different reason. It will be damn near useless because it >> can't keep up with an evolving subsystem. >> >> Now once you get multiple ABI versions and you can seamlessly go from >> old to new one, say you have a version that was superseded 5 years ago. >> It would be completely reasonable to say that this version is old enough >> and no one should be going from a 5 year old kernel to a modern one. > > LOL, Google literally does this. Granted, the extreme cases only happen for > stragglers, and I think we do force VMs to bounce through a "middle" kernel in > those cases, but I doubt Google is the only company that runs frankenkernels > for an absurd number of years for a variety of reasons. E.g. 4.4 LTS was officially > supported for 6 years, and I'll bet dollars to donuts people ran it for much longer > than that. > >> So you deprecate this ABI. Deprecating old unused uAPIs is not unprecedented. > > When there are provably no users, or we can convince existing userspace to migrate > to an alternative. > >> I think we are better off formalizing this deprecation period from the >> get go. > > Why? What does it buy us? Because all I see is potential abuse and an excuse > for not spending time getting the designs right. > >> A somewhat tangential example is BPF kfuncs. My BPF program that works >> in kernel X might not work in kernel Y because the kfunc has changed or >> been removed. >> >> The argument they make in kfuncs.rst is that kfuncs "provide a kernel >> <-> kernel API, and thus are not bound by any of the strict stability >> restrictions associated with kernel <-> user UAPIs". > > BPF's documentation isn't arguing anything, it's merely reiterating Linux's > long-standing policy that there is no such thing as a stable kernel ABI (in upstream). > >> For LUO as well, this is a kernel -> kernel API. > > Stating the obvious, I disagree with this. As I said before, if this is the > stance LUO wants to take, then so be it, but my NAK stands. > >> Users can also still do a regular kexec or reboot. They just won't get the >> performance optimization of LUO. > > The amount of time, energy, and money poured into minimizing VM downtime on live > migration suggests the overwhelming majority of LUO's targeted users aren't going > to take kindly to this stance. > >> Regardless of if you agree with the last bit about deprecating old >> versions, ABIs evolving with the subsystem is a ground reality of live >> update and it would be foolish to think we can do with only >> backwards-compatible ABI changes forever. > > I never said the ABI is immutable, I said it needs to be backwards/forwards > compatible. Thanks Sean, Pratyush. @loganodell has recently sent the RFC[1] on backward compatibility for liveupdate. This discussion is important so I propose to move the backward compatibility discussion on [1]. For guest_memfd preservation, I will sent v5 soon with the following changes: 1. Suggestions from across the patches 2. tentative Plan for future Guest_memfd preservation 3. tentative Plan for VM preservation. 4. Backward compatibility built on top of RFC [1] What do you think? [1]: https://lore.kernel.org/all/20260903023452.721732-1-loganodell@google.com/