From: Michael Roth <michael.roth@amd.com>
To: <qemu-devel@nongnu.org>
Cc: <kvm@vger.kernel.org>, <pbonzini@redhat.com>,
<berrange@redhat.com>, <armbru@redhat.com>,
<pankaj.gupta@amd.com>, <isaku.yamahata@intel.com>,
<xiaoyao.li@intel.com>, <chao.p.peng@linux.intel.com>,
<david@kernel.org>, <ashish.kalra@amd.com>,
<ackerleytng@google.com>, <lpieralisi@kernel.org>
Subject: [PATCH v2 13/19] accel/kvm: Support shared/private conversions via guest_memfd ioctls
Date: Tue, 8 Sep 2026 15:48:33 -0500 [thread overview]
Message-ID: <20260908205236.838281-14-michael.roth@amd.com> (raw)
In-Reply-To: <20260908205236.838281-1-michael.roth@amd.com>
When using guest_memfd with support for shared memory / in-place
conversion, it is necessary to use the guest_memfd ioctls to handle
conversions instead of KVM ioctls. Implement support for this by looping
through all the sections within a converison range. Implement everything
in terms of the kvm_convert_memory() loop, which already deals with some
special considerations regarding various holes / region types that might
be encountered.
Also update kvm_set_memory_attributes_*() to use the same common path
when convert-in-place=false. This potentially results in a small change
in behavior due to the additional MMIO checks/skips now being applied in
that case (generally qemu-triggered during setup) rather than only for
kvm_convert_memory() (generally guest-triggered), but this is arguably
safer, and it provides similar behavior between convert-in-place=false
vs. convert-in-place=true, the latter of which *must* skip MMIO holes
because the regions (and associated guest_memfds) themselves track
shared/private state internally and passing the whole conversion range
through to KVM is not an option in that case.
Signed-off-by: Michael Roth <michael.roth@amd.com>
---
accel/kvm/kvm-all.c | 131 ++++++++++++++++++++++++++++++++++++++------
1 file changed, 114 insertions(+), 17 deletions(-)
diff --git a/accel/kvm/kvm-all.c b/accel/kvm/kvm-all.c
index 17e25f9074..fcb6f3bee9 100644
--- a/accel/kvm/kvm-all.c
+++ b/accel/kvm/kvm-all.c
@@ -1624,14 +1624,78 @@ static int kvm_set_memory_attributes(hwaddr start, uint64_t size, uint64_t attr)
return r;
}
-int kvm_set_memory_attributes_private(hwaddr start, uint64_t size)
+static int kvm_gmem_ioctl(int guest_memfd, unsigned long type, ...)
{
- return kvm_set_memory_attributes(start, size, KVM_MEMORY_ATTRIBUTE_PRIVATE);
+ int ret;
+ void *arg;
+ va_list ap;
+
+ va_start(ap, type);
+ arg = va_arg(ap, void *);
+ va_end(ap);
+
+ ret = ioctl(guest_memfd, type, arg);
+ if (ret == -1) {
+ ret = -errno;
+ }
+ return ret;
}
-int kvm_set_memory_attributes_shared(hwaddr start, uint64_t size)
+static int guest_memfd_set_memory_attributes_fd(int guest_memfd, hwaddr offset,
+ uint64_t size, uint64_t attr)
{
- return kvm_set_memory_attributes(start, size, 0);
+ struct kvm_memory_attributes2 attrs = {0};
+ int r;
+
+ assert((attr & kvm_supported_memory_attributes) == attr);
+ attrs.attributes = attr;
+ attrs.offset = offset;
+ attrs.size = size;
+ attrs.flags = 0;
+
+ /*
+ * guest_memfd may need to delay conversion requests due to
+ * the memory being in-use by the kernel. In most cases these
+ * will be transient uses. In some cases, userspace itself may
+ * be the cause of the memory being considered in-use, though
+ * QEMU currently takes steps to avoid this (e.g. via
+ * RamBlockAttributes). On that basis, this code loops
+ * indefinitely with the assumption that only transient cases
+ * will block, and that those will be for relatively short
+ * periods vs. the overall conversion path.
+ * If those assumptions at some point prove false, most likely
+ * this will manifest as guest-side lockups on their conversion
+ * path, which seems like the appropriate way to surface this
+ * situation to the guest owner rather than some hard timeout.
+ */
+ do {
+ r = kvm_gmem_ioctl(guest_memfd, KVM_SET_MEMORY_ATTRIBUTES2, &attrs);
+ } while (r == -EAGAIN);
+
+ if (r) {
+ error_report("failed to set memory (0x%" HWADDR_PRIx "+0x%" PRIx64 ") "
+ "with attr 0x%" PRIx64 " error '%s'",
+ offset, size, attr, strerror(-r));
+ }
+ return r;
+}
+
+static int guest_memfd_set_memory_section_attributes(MemoryRegionSection *section, uint64_t attr)
+{
+ hwaddr convert_offset, convert_size;
+ MemoryRegion *mr = section->mr;
+ RAMBlock *rb;
+
+ assert(mr);
+ rb = mr->ram_block;
+ assert(rb->guest_memfd_private >= 0);
+ convert_offset = section->offset_within_region;
+ convert_size = int128_get64(section->size);
+
+ return guest_memfd_set_memory_attributes_fd(rb->guest_memfd_private,
+ convert_offset,
+ convert_size,
+ attr);
}
bool kvm_private_memory_attribute_supported(void)
@@ -3439,10 +3503,18 @@ static int kvm_convert_section(MemoryRegionSection *section, bool to_private)
hwaddr size = int128_get64(section->size);
int ret;
- if (to_private) {
- ret = kvm_set_memory_attributes_private(start, size);
+ if (machine_require_guest_memfd_convert_in_place(current_machine)) {
+ ret = guest_memfd_set_memory_section_attributes(section,
+ to_private ? KVM_MEMORY_ATTRIBUTE_PRIVATE
+ : 0);
} else {
- ret = kvm_set_memory_attributes_shared(start, size);
+ /*
+ * Without in-place conversion, attribute-tracking is handled by KVM
+ * across all guest memory rather than on a per-section/slot basis.
+ */
+ ret = kvm_set_memory_attributes(start, size,
+ to_private ? KVM_MEMORY_ATTRIBUTE_PRIVATE
+ : 0);
}
return ret;
@@ -3536,7 +3608,8 @@ static int kvm_post_convert_section(MemoryRegionSection *section, bool to_privat
return ret;
}
-int kvm_convert_memory(hwaddr start, hwaddr size, bool to_private)
+static int kvm_convert_memory_full(hwaddr start, hwaddr size, bool to_private,
+ bool pre_hooks, bool post_hooks)
{
int ret = -EINVAL;
@@ -3580,10 +3653,12 @@ int kvm_convert_memory(hwaddr start, hwaddr size, bool to_private)
continue;
}
- ret = kvm_pre_convert_section(§ion, to_private);
- if (ret) {
- memory_region_unref(section.mr);
- break;
+ if (pre_hooks) {
+ ret = kvm_pre_convert_section(§ion, to_private);
+ if (ret) {
+ memory_region_unref(section.mr);
+ break;
+ }
}
ret = kvm_convert_section(§ion, to_private);
@@ -3592,13 +3667,15 @@ int kvm_convert_memory(hwaddr start, hwaddr size, bool to_private)
break;
}
- ret = kvm_post_convert_section(§ion, to_private);
- memory_region_unref(section.mr);
-
- if (ret) {
- break;
+ if (post_hooks) {
+ ret = kvm_post_convert_section(§ion, to_private);
+ if (ret) {
+ memory_region_unref(section.mr);
+ break;
+ }
}
+ memory_region_unref(section.mr);
size -= section_end - start;
start = section_end;
}
@@ -3606,6 +3683,26 @@ int kvm_convert_memory(hwaddr start, hwaddr size, bool to_private)
return ret;
}
+int kvm_convert_memory(hwaddr start, hwaddr size, bool to_private)
+{
+ return kvm_convert_memory_full(start, size, to_private, true, true);
+}
+
+static int kvm_convert_memory_attributes(hwaddr start, hwaddr size, bool to_private)
+{
+ return kvm_convert_memory_full(start, size, to_private, false, false);
+}
+
+int kvm_set_memory_attributes_private(hwaddr start, uint64_t size)
+{
+ return kvm_convert_memory_attributes(start, size, KVM_MEMORY_ATTRIBUTE_PRIVATE);
+}
+
+int kvm_set_memory_attributes_shared(hwaddr start, uint64_t size)
+{
+ return kvm_convert_memory_attributes(start, size, 0);
+}
+
int kvm_cpu_exec(CPUState *cpu)
{
struct kvm_run *run = cpu->kvm_run;
--
2.43.0
next prev parent reply other threads:[~2026-09-08 20:55 UTC|newest]
Thread overview: 24+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-08 20:48 [PATCH v2 00/19] guest_memfd: support in-place memory conversion Michael Roth
2026-09-08 20:48 ` [PATCH v2 01/19] accel/kvm: Add helper for handling conversions of MMIO holes Michael Roth
2026-09-08 20:48 ` [PATCH v2 02/19] accel/kvm: Fix kvm_convert_memory() calls crossing memory regions Michael Roth
2026-09-08 20:48 ` [PATCH v2 03/19] accel/kvm: Fix handling of MMIO holes at start of conversion ranges Michael Roth
2026-09-08 20:48 ` [PATCH v2 04/19] accel/kvm: Fix handling of conversion ranges with multiple MMIO holes Michael Roth
2026-09-08 20:48 ` [PATCH v2 05/19] accel/kvm: Use dedicated helper for creating private-only gmem instances Michael Roth
2026-09-08 20:48 ` [PATCH v2 06/19] linux-headers: Update headers for v12 of in-place conversion kernel support Michael Roth
2026-09-08 20:48 ` [PATCH v2 07/19] accel/kvm: Add CGS option to control in-place conversion support Michael Roth
2026-09-09 6:20 ` Markus Armbruster
2026-09-08 20:48 ` [PATCH v2 08/19] system/memory: Re-use memory-backend-guest-memfd inode for private memory Michael Roth
2026-09-10 8:39 ` David Hildenbrand
2026-09-08 20:48 ` [PATCH v2 09/19] accel/kvm: Handle guest_memfd flags internally when creating instances Michael Roth
2026-09-08 20:48 ` [PATCH v2 10/19] system/memory: Default to guest_memfd for RAM for in-place conversion Michael Roth
2026-09-08 20:48 ` [PATCH v2 11/19] accel/kvm: Move post-conversion updates to a separate helper Michael Roth
2026-09-08 20:48 ` [PATCH v2 12/19] accel/kvm: Re-order attribute notifications for in-place conversion Michael Roth
2026-09-08 20:48 ` Michael Roth [this message]
2026-09-08 20:48 ` [PATCH v2 14/19] accel/kvm: Don't default to private attributes " Michael Roth
2026-09-08 20:48 ` [PATCH v2 15/19] i386/sev: Update SNP_LAUNCH_UPDATE " Michael Roth
2026-09-08 20:48 ` [PATCH v2 16/19] i386/sev: Allow in-place conversion for SEV-SNP guests Michael Roth
2026-09-08 20:48 ` [PATCH v2 17/19] i386/sev: Update CPUID failure handling for in-place conversion Michael Roth
2026-09-08 20:48 ` [PATCH v2 18/19] accel/kvm: Disable discard " Michael Roth
2026-09-08 22:16 ` Michael Roth
2026-09-08 20:48 ` [PATCH v2 19/19] hostmem: Automatically select set guest-memfd=on " Michael Roth
2026-09-09 6:30 ` Markus Armbruster
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260908205236.838281-14-michael.roth@amd.com \
--to=michael.roth@amd.com \
--cc=ackerleytng@google.com \
--cc=armbru@redhat.com \
--cc=ashish.kalra@amd.com \
--cc=berrange@redhat.com \
--cc=chao.p.peng@linux.intel.com \
--cc=david@kernel.org \
--cc=isaku.yamahata@intel.com \
--cc=kvm@vger.kernel.org \
--cc=lpieralisi@kernel.org \
--cc=pankaj.gupta@amd.com \
--cc=pbonzini@redhat.com \
--cc=qemu-devel@nongnu.org \
--cc=xiaoyao.li@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox