From: Michael Roth <michael.roth@amd.com>
To: <qemu-devel@nongnu.org>
Cc: <kvm@vger.kernel.org>, <pbonzini@redhat.com>,
<berrange@redhat.com>, <armbru@redhat.com>,
<pankaj.gupta@amd.com>, <xiaoyao.li@intel.com>,
<chao.p.peng@linux.intel.com>, <david@kernel.org>,
<ashish.kalra@amd.com>, <ackerleytng@google.com>,
<lpieralisi@kernel.org>
Subject: [PATCH v3 00/19] guest_memfd: support in-place memory conversion
Date: Wed, 7 Oct 2026 08:41:24 -0500 [thread overview]
Message-ID: <20261007134323.1606088-1-michael.roth@amd.com> (raw)
v1: https://lore.kernel.org/qemu-devel/20260528000416.8161-1-michael.roth@amd.com/
v2: https://lore.kernel.org/qemu-devel/20260908205236.838281-1-michael.roth@amd.com/
v3:
- rebase on top of master, which now includes prerequisite guest-memfd support
for non-confidential VMs
- make CGS option implementation-specific (Markus)
- beef up conf/guest-memfd rst docs and reference in qapi changes (Markus)
- rename _SHARED flag to _SHAREABLE (David)
- refactor/reorder patches so that SNP QAPI definitions to enable in-place
conversion are deferred until all the required code is in place rather
than relying on the new removed 'allow_convert_in_place' CGS flag
This patchset is also available at:
https://github.com/amdese/qemu/commits/snp-inplace-v3
The first 5 patches are bug-fixes for code that was added for TDX to deal with
private MMIO ranges. Since much of that needed to be refactored as part of this
series, they are included here for context/review. Testing of those patches
would also be appreciated since I don't have access to TDX hardware and am
unable to check for regressions.
OVERVIEW
--------
This series adds guest_memfd support for in-place conversion of memory
between private/shared, and enables it for SEV-SNP guests. It is based
on recently-added kernel support for mmap()-able guest_memfd
instances[1], which allow it to be used for shared memory, and the
following patchset[2], which adds additional guest_memfd interfaces to
allow it to be used to perform in-place conversion:
[PATCH v13 00/44] guest_memfd: In-place conversion support
https://lore.kernel.org/kvm/20260910-gmem-inplace-conversion-v13-0-dd6fbf94f4e1@google.com/
That series also introduces a new 'gmem_in_place_conversion' KVM
module option that controls setting this mode vs. non-in-place
behavior. Currently in-place mode is planned to eventually deprecate
non-in-place mode.
MOTIVATION
----------
Today, SEV-SNP guests (and other CoCo VM types using guest_memfd) keep
shared and private memory on separate physical backings: a userspace
memory-backend object for shared pages, and a kernel-allocated
guest_memfd file descriptor for private pages. KVM_SET_MEMORY_ATTRIBUTES
flips which backing the guest sees for a given GPA range, and the old
backing is typically discarded / hole-punched on conversion to avoid
doubled memory usage.
That model works, but has a number of downsides that impact certain
use-cases:
- Each conversion involves discarding pages on one side and faulting
them in on the other, which incurs allocation overheads in the
host kernel for every conversion.
- Some use-cases, like pKVM[3], rely on memory isolation rather than
encryption and rely on in-place conversion to pass through things
like secured framebuffer memory without needing to bounce data
through separate shared/private HPAs, which would introduce
unacceptable latency for that sort of workload.
- Hugetlb support[4] for guest_memfd will rely on it, since things like
1GB hugepages with a mix of shared/private sub-ranges would generally
require 2 1GB hugetlb pages to remain available to handle shared vs.
private accesses, which quickly causes doubling of guest memory usage.
Recent kernel work[2] makes guest_memfd mmap()-able and lets the *same*
physical pages be used for both shared and private states for a given
GPA range, allowing the above pitfalls to be naturally avoided.
This series wires that support up in QEMU.
DESIGN
------
For confidential VMs, a new 'convert-in-place' flag is added to switch
on in-place conversion support. When running in this mode, the user
*MUST* use memory-backend-guest-memfd for backing guest RAM. A new
RAM_GUEST_MEMFD_SHARED RAMBlock flag is added to track/enforce the
dependency. Additionally, QEMU is modified to use mmap()-able
guest_memfd and set this flag for other cases where it allocates RAM
internally. As a result, block->fd will generally always be a
guest_memfd, and when RAM_GUEST_MEMFD_SHARED is set then that block->fd
will be qemu_dup()'d as the FD handle for private memory as well. This
allows the prior non-in-place handling around block->guest_memfd_private
to be kept mostly unchanged for the in-place case.
When running with convert-in-place=true, shared/private conversions
are no longer handled directly by KVM, but instead by a new guest_memfd
ioctl, KVM_SET_MEMORY_ATTRIBUTES2, which purposely provides similar
naming/implementation to the KVM_SET_MEMORY_ATTRIBUTES KVM ioctl that
it replaces. This series adds handling to route conversion requests to
the appropriate ioctls based on whether or not in-place conversion is
enabled.
This support also relies on the memory-backend-memfd,guest-memfd=on
option, since it is necessary to use guest_memfd for both the shared
and private memory. This is set automatically based on whether or
not in-place conversion is enabled.
Since guest_memfd ioctls need to be called against the specific
guest_memfd inode associated with each memory slot/region, some
refactoring is needed to handle conversions on a per-region basis. Much
of that is inherited from the bugfix series this patchset is based on
top of, which adds the initial logic for handling multiple sections
within a range.
USAGE
-----
After applying this series against a kernel with the RFC patches above
present, an SEV-SNP guest can be started with in-place conversion via:
qemu-system-x86_64 \
-machine q35,confidential-guest-support=sev0,memory-backend=ram0 \
-object memory-backend-memfd,id=ram0,size=8G,share=on \
-object sev-snp-guest,id=sev0,cbitpos=51,reduced-phys-bits=1,\
convert-in-place=on \
...
NOTES/TODO
----------
- TDX testing would be great, in theory it can be enabled with this
series (similarly to patch #18) but I'm not sure if there are
other special requirements before we can switch it on.
- kernel patches are still in-flight, but are now in kvm-x86/next and
nearing upstream
REFERENCES
----------
[1] https://lore.kernel.org/kvm/20250729225455.670324-1-seanjc@google.com/
[2] https://lore.kernel.org/kvm/20260910-gmem-inplace-conversion-v13-0-dd6fbf94f4e1@google.com/
[3] https://www.youtube.com/watch?v=MMfAGNW9RVg
[4] https://lore.kernel.org/kvm/cover.1747264138.git.ackerleytng@google.com/
Thoughts, feedback, and testing are very much appreciated.
Thanks,
Mike
----------------------------------------------------------------
Ashish Kalra (1):
accel/kvm: Fix kvm_convert_memory() calls crossing memory regions
Michael Roth (18):
accel/kvm: Add helper for handling conversions of MMIO holes
accel/kvm: Fix handling of MMIO holes at start of conversion ranges
accel/kvm: Fix handling of conversion ranges with multiple MMIO holes
accel/kvm: Use dedicated helper for creating private-only gmem instances
[NOT-FOR-MERGE] linux-headers: Update headers for v13 of in-place conversion kernel support
accel/kvm: Add CGS flag to control in-place conversion support
system/memory: Re-use memory-backend-guest-memfd inode for private memory
accel/kvm: Handle guest_memfd flags internally when creating instances
system/memory: Default to guest_memfd for RAM for in-place conversion
accel/kvm: Move post-conversion updates to a separate helper
accel/kvm: Re-order attribute notifications for in-place conversion
accel/kvm: Support shared/private conversions via guest_memfd ioctls
accel/kvm: Don't default to private attributes for in-place conversion
i386/sev: Update SNP_LAUNCH_UPDATE for in-place conversion
i386/sev: Update CPUID failure handling for in-place conversion
accel/kvm: Disable discard for in-place conversion
i386/sev: Add sev-snp-guest parameter to enable in-place conversion
hostmem: Automatically select set guest-memfd=on for in-place conversion
accel/kvm/kvm-all.c | 449 ++++++++++++++++++++++++----
accel/stubs/kvm-stub.c | 8 +-
backends/hostmem-memfd.c | 19 +-
docs/system/guest-memfd.rst | 21 ++
hw/core/machine.c | 5 +
include/hw/core/boards.h | 1 +
include/system/confidential-guest-support.h | 12 +
include/system/kvm.h | 3 +-
include/system/memory.h | 7 +
linux-headers/linux/kvm.h | 16 +
qapi/qom.json | 13 +-
system/memory.c | 23 +-
system/physmem.c | 54 +++-
target/i386/sev.c | 36 ++-
14 files changed, 585 insertions(+), 82 deletions(-)
next reply other threads:[~2026-10-07 13:44 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-07 13:41 Michael Roth [this message]
2026-10-07 13:41 ` [PATCH v3 01/19] accel/kvm: Add helper for handling conversions of MMIO holes Michael Roth
2026-10-07 13:41 ` [PATCH v3 02/19] accel/kvm: Fix kvm_convert_memory() calls crossing memory regions Michael Roth
2026-10-07 13:41 ` [PATCH v3 03/19] accel/kvm: Fix handling of MMIO holes at start of conversion ranges Michael Roth
2026-10-07 13:41 ` [PATCH v3 04/19] accel/kvm: Fix handling of conversion ranges with multiple MMIO holes Michael Roth
2026-10-07 13:41 ` [PATCH v3 05/19] accel/kvm: Use dedicated helper for creating private-only gmem instances Michael Roth
2026-10-07 13:41 ` [PATCH v3 06/19] [NOT-FOR-MERGE] linux-headers: Update headers for v13 of in-place conversion kernel support Michael Roth
2026-10-07 13:41 ` [PATCH v3 07/19] accel/kvm: Add CGS flag to control in-place conversion support Michael Roth
2026-10-07 13:41 ` [PATCH v3 08/19] system/memory: Re-use memory-backend-guest-memfd inode for private memory Michael Roth
2026-10-07 13:41 ` [PATCH v3 09/19] accel/kvm: Handle guest_memfd flags internally when creating instances Michael Roth
2026-10-07 13:41 ` [PATCH v3 10/19] system/memory: Default to guest_memfd for RAM for in-place conversion Michael Roth
2026-10-07 13:41 ` [PATCH v3 11/19] accel/kvm: Move post-conversion updates to a separate helper Michael Roth
2026-10-07 13:41 ` [PATCH v3 12/19] accel/kvm: Re-order attribute notifications for in-place conversion Michael Roth
2026-10-07 13:41 ` [PATCH v3 13/19] accel/kvm: Support shared/private conversions via guest_memfd ioctls Michael Roth
2026-10-07 13:41 ` [PATCH v3 14/19] accel/kvm: Don't default to private attributes for in-place conversion Michael Roth
2026-10-07 13:41 ` [PATCH v3 15/19] i386/sev: Update SNP_LAUNCH_UPDATE " Michael Roth
2026-10-07 13:41 ` [PATCH v3 16/19] i386/sev: Update CPUID failure handling " Michael Roth
2026-10-07 13:41 ` [PATCH v3 17/19] accel/kvm: Disable discard " Michael Roth
2026-10-07 13:41 ` [PATCH v3 18/19] i386/sev: Add sev-snp-guest parameter to enable " Michael Roth
2026-10-08 6:11 ` Markus Armbruster
2026-10-07 13:41 ` [PATCH v3 19/19] hostmem: Automatically select set guest-memfd=on for " Michael Roth
2026-10-08 6:14 ` Markus Armbruster
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261007134323.1606088-1-michael.roth@amd.com \
--to=michael.roth@amd.com \
--cc=ackerleytng@google.com \
--cc=armbru@redhat.com \
--cc=ashish.kalra@amd.com \
--cc=berrange@redhat.com \
--cc=chao.p.peng@linux.intel.com \
--cc=david@kernel.org \
--cc=kvm@vger.kernel.org \
--cc=lpieralisi@kernel.org \
--cc=pankaj.gupta@amd.com \
--cc=pbonzini@redhat.com \
--cc=qemu-devel@nongnu.org \
--cc=xiaoyao.li@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox