Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Michael Roth <michael.roth@amd.com>
To: <qemu-devel@nongnu.org>
Cc: <kvm@vger.kernel.org>, <pbonzini@redhat.com>,
	<berrange@redhat.com>, <armbru@redhat.com>,
	<pankaj.gupta@amd.com>, <xiaoyao.li@intel.com>,
	<chao.p.peng@linux.intel.com>, <david@kernel.org>,
	<ashish.kalra@amd.com>, <ackerleytng@google.com>,
	<lpieralisi@kernel.org>
Subject: [PATCH v3 12/19] accel/kvm: Re-order attribute notifications for in-place conversion
Date: Wed, 7 Oct 2026 08:41:36 -0500	[thread overview]
Message-ID: <20261007134323.1606088-13-michael.roth@amd.com> (raw)
In-Reply-To: <20261007134323.1606088-1-michael.roth@amd.com>

ram-block-attribute update notifications are currently sent after
conversions from/to private pages to trigger DMA maps/unmaps of shared
GPA ranges (respectively). However, with in-place conversion additional
requirements on the kernel side come into play which require this
behavior to be adjusted.

For shared->private conversions: the attributes need to be set to
private *after* the notification, since when using VFIO it may not be
possible to update the attribute while it remains pinned due to the
IOMMU mapping, so issue the notification first to ensure unmappings are
done in advance.

For private->shared conversions: the attributes need to be set to shared
*before* the notification, since it will possibly result in the page
being mapped into an IOMMU and trigger guest_memfd's fault handler,
which will expect the page to have its attributes set to shared or
otherwise SIGBUS.

Implement this to enable passthrough support for CoCo guests with
in-place conversion support enabled. For non-inplace conversion, pages
mapped into the IOMMU are not the same physical pages as the one used
for private accesses by the guest, so neither order risks DMA accesses
to private memory and that path can be consolidated to use the same
handling as well.

Signed-off-by: Michael Roth <michael.roth@amd.com>
---
 accel/kvm/kvm-all.c | 60 +++++++++++++++++++++++++++++++++++++++++++--
 1 file changed, 58 insertions(+), 2 deletions(-)

diff --git a/accel/kvm/kvm-all.c b/accel/kvm/kvm-all.c
index b824b803cb..60ab5193c9 100644
--- a/accel/kvm/kvm-all.c
+++ b/accel/kvm/kvm-all.c
@@ -3448,7 +3448,7 @@ static int kvm_convert_section(MemoryRegionSection *section, bool to_private)
     return ret;
 }
 
-static int kvm_post_convert_section(MemoryRegionSection *section, bool to_private)
+static int kvm_pre_convert_section(MemoryRegionSection *section, bool to_private)
 {
     hwaddr start = section->offset_within_address_space;
     hwaddr size = int128_get64(section->size);
@@ -3458,16 +3458,66 @@ static int kvm_post_convert_section(MemoryRegionSection *section, bool to_privat
     void *addr;
     int ret;
 
+    if (!to_private)
+        return 0;
+
     addr = memory_region_get_ram_ptr(mr) + section->offset_within_region;
     rb = qemu_ram_block_from_host(addr, false, &offset);
 
+    /*
+     * The attributes need to be set to private *after* the notification
+     * of a shared->private conversion, since when using VFIO it may not
+     * be possible to update the attribute while it remains pinned due
+     * to the IOMMU mapping, so issue the notification first to ensure
+     * unmappings are done in advance.
+     *
+     * There is an asymmetry here in that if the subsequent memory
+     * attribute update fails, this notification is out of sync with the
+     * state as tracked by guest_memfd, which isn't ideal, but memory
+     * attribute failures are not expected to be recoverable any way so
+     * there it would be a waste of time to roll back the notification and
+     * re-trigger things like mapping the page via iommufd.
+     */
     ret = ram_block_attributes_state_change(rb->attributes,
                                             offset, size, to_private);
     if (ret) {
         error_report("Failed to notify the listener the state change of "
                      "(0x%"HWADDR_PRIx" + 0x%"HWADDR_PRIx") to %s, ret %d",
                      start, size, to_private ? "private" : "shared", ret);
-        return ret;
+    }
+
+    return ret;
+}
+
+static int kvm_post_convert_section(MemoryRegionSection *section, bool to_private)
+{
+    hwaddr start = section->offset_within_address_space;
+    hwaddr size = int128_get64(section->size);
+    MemoryRegion *mr = section->mr;
+    ram_addr_t offset;
+    RAMBlock *rb;
+    void *addr;
+    int ret;
+
+    addr = memory_region_get_ram_ptr(mr) + section->offset_within_region;
+    rb = qemu_ram_block_from_host(addr, false, &offset);
+
+    /*
+     * The attributes need to have been set to shared *before* the notification
+     * of a private->shared conversion, since it will possibly result in the
+     * page being mapped into an IOMMU when using VFIO and trigger
+     * guest_memfd's fault handler, which will expect the page to have its
+     * attributes set to shared.
+     */
+    if (!to_private) {
+        ret = ram_block_attributes_state_change(rb->attributes,
+                                                offset, size, to_private);
+        if (ret) {
+            error_report("Failed to notify the listener the state change of "
+                         "(0x%"HWADDR_PRIx" + 0x%"HWADDR_PRIx") to %s, ret %d",
+                         start, size, to_private ? "private" : "shared", ret);
+            return ret;
+        }
     }
 
     if (to_private) {
@@ -3530,6 +3580,12 @@ int kvm_convert_memory(hwaddr start, hwaddr size, bool to_private)
             continue;
         }
 
+        ret = kvm_pre_convert_section(&section, to_private);
+        if (ret) {
+            memory_region_unref(section.mr);
+            break;
+        }
+
         ret = kvm_convert_section(&section, to_private);
         if (ret) {
             memory_region_unref(section.mr);
-- 
2.43.0


  parent reply	other threads:[~2026-10-07 13:45 UTC|newest]

Thread overview: 22+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-07 13:41 [PATCH v3 00/19] guest_memfd: support in-place memory conversion Michael Roth
2026-10-07 13:41 ` [PATCH v3 01/19] accel/kvm: Add helper for handling conversions of MMIO holes Michael Roth
2026-10-07 13:41 ` [PATCH v3 02/19] accel/kvm: Fix kvm_convert_memory() calls crossing memory regions Michael Roth
2026-10-07 13:41 ` [PATCH v3 03/19] accel/kvm: Fix handling of MMIO holes at start of conversion ranges Michael Roth
2026-10-07 13:41 ` [PATCH v3 04/19] accel/kvm: Fix handling of conversion ranges with multiple MMIO holes Michael Roth
2026-10-07 13:41 ` [PATCH v3 05/19] accel/kvm: Use dedicated helper for creating private-only gmem instances Michael Roth
2026-10-07 13:41 ` [PATCH v3 06/19] [NOT-FOR-MERGE] linux-headers: Update headers for v13 of in-place conversion kernel support Michael Roth
2026-10-07 13:41 ` [PATCH v3 07/19] accel/kvm: Add CGS flag to control in-place conversion support Michael Roth
2026-10-07 13:41 ` [PATCH v3 08/19] system/memory: Re-use memory-backend-guest-memfd inode for private memory Michael Roth
2026-10-07 13:41 ` [PATCH v3 09/19] accel/kvm: Handle guest_memfd flags internally when creating instances Michael Roth
2026-10-07 13:41 ` [PATCH v3 10/19] system/memory: Default to guest_memfd for RAM for in-place conversion Michael Roth
2026-10-07 13:41 ` [PATCH v3 11/19] accel/kvm: Move post-conversion updates to a separate helper Michael Roth
2026-10-07 13:41 ` Michael Roth [this message]
2026-10-07 13:41 ` [PATCH v3 13/19] accel/kvm: Support shared/private conversions via guest_memfd ioctls Michael Roth
2026-10-07 13:41 ` [PATCH v3 14/19] accel/kvm: Don't default to private attributes for in-place conversion Michael Roth
2026-10-07 13:41 ` [PATCH v3 15/19] i386/sev: Update SNP_LAUNCH_UPDATE " Michael Roth
2026-10-07 13:41 ` [PATCH v3 16/19] i386/sev: Update CPUID failure handling " Michael Roth
2026-10-07 13:41 ` [PATCH v3 17/19] accel/kvm: Disable discard " Michael Roth
2026-10-07 13:41 ` [PATCH v3 18/19] i386/sev: Add sev-snp-guest parameter to enable " Michael Roth
2026-10-08  6:11   ` Markus Armbruster
2026-10-07 13:41 ` [PATCH v3 19/19] hostmem: Automatically select set guest-memfd=on for " Michael Roth
2026-10-08  6:14   ` Markus Armbruster

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261007134323.1606088-13-michael.roth@amd.com \
    --to=michael.roth@amd.com \
    --cc=ackerleytng@google.com \
    --cc=armbru@redhat.com \
    --cc=ashish.kalra@amd.com \
    --cc=berrange@redhat.com \
    --cc=chao.p.peng@linux.intel.com \
    --cc=david@kernel.org \
    --cc=kvm@vger.kernel.org \
    --cc=lpieralisi@kernel.org \
    --cc=pankaj.gupta@amd.com \
    --cc=pbonzini@redhat.com \
    --cc=qemu-devel@nongnu.org \
    --cc=xiaoyao.li@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox