From: Ackerley Tng <ackerleytng@google.com>
To: aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com,
brauner@kernel.org, chao.p.peng@linux.intel.com,
david@kernel.org, jmattson@google.com, jthoughton@google.com,
michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com,
qperret@google.com, rick.p.edgecombe@intel.com,
rientjes@google.com, shivankg@amd.com, steven.price@arm.com,
willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com,
forkloop@google.com, pratyush@kernel.org,
suzuki.poulose@arm.com, aneesh.kumar@kernel.org,
liam@infradead.org, Paolo Bonzini <pbonzini@redhat.com>,
Sean Christopherson <seanjc@google.com>,
Thomas Gleixner <tglx@kernel.org>,
Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
x86@kernel.org, "H. Peter Anvin" <hpa@zytor.com>,
Steven Rostedt <rostedt@goodmis.org>,
Masami Hiramatsu <mhiramat@kernel.org>,
Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Shuah Khan <shuah@kernel.org>,
Vishal Annapurve <vannapurve@google.com>,
Andrew Morton <akpm@linux-foundation.org>,
Chris Li <chrisl@kernel.org>, Kairui Song <kasong@tencent.com>,
Kemeng Shi <shikemeng@huaweicloud.com>,
Nhat Pham <nphamcs@gmail.com>, Barry Song <baohua@kernel.org>,
Axel Rasmussen <axelrasmussen@google.com>,
Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
Youngjun Park <youngjun.park@lge.com>,
Qi Zheng <qi.zheng@linux.dev>,
Shakeel Butt <shakeel.butt@linux.dev>,
Kiryl Shutsemau <kas@kernel.org>,
Baoquan He <baoquan.he@linux.dev>, Jason Gunthorpe <jgg@ziepe.ca>,
John Hubbard <jhubbard@nvidia.com>, Peter Xu <peterx@redhat.com>,
tarunsahu@google.com, Fuad Tabba <fuad.tabba@linux.dev>,
Vlastimil Babka <vbabka@kernel.org>
Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org,
linux-kselftest@vger.kernel.org, linux-mm@kvack.org,
linux-coco@lists.linux.dev,
Ackerley Tng <ackerleytng@google.com>
Subject: [PATCH v11 14/46] KVM: guest_memfd: Ensure pages are not in use before conversion
Date: Wed, 26 Aug 2026 09:18:12 +0000 [thread overview]
Message-ID: <20260826-gmem-inplace-conversion-v11-14-0a15d8a799aa@google.com> (raw)
In-Reply-To: <20260826-gmem-inplace-conversion-v11-0-0a15d8a799aa@google.com>
When converting memory to private in guest_memfd, it is necessary to ensure
that the pages are not currently being accessed by any other part of the
kernel or userspace to avoid any current user writing to guest private
memory.
guest_memfd checks for any outstanding references to determine whether a
page is still in use. The only expected references after unmapping the
range requested for conversion are those that are held by guest_memfd
itself.
Update the kvm_memory_attributes2 structure to include an error_offset
field. This allows KVM to report the exact offset where a conversion
failed. If the safety check fails, return -EAGAIN and copy the error_offset
back to userspace so that it can potentially retry the operation or handle
the failure gracefully.
Update documentation to document the error_offset field and the possible
-EAGAIN error.
Suggested-by: David Hildenbrand <david@kernel.org>
Co-developed-by: Vishal Annapurve <vannapurve@google.com>
Signed-off-by: Vishal Annapurve <vannapurve@google.com>
Reviewed-by: Fuad Tabba <tabba@google.com>
Tested-by: Shivank Garg <shivankg@amd.com>
Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
Documentation/virt/kvm/api.rst | 19 +++++++++--
include/uapi/linux/kvm.h | 3 +-
virt/kvm/guest_memfd.c | 77 +++++++++++++++++++++++++++++++++++++++---
3 files changed, 91 insertions(+), 8 deletions(-)
diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
index 7aad097ffd392..f88d65b78c504 100644
--- a/Documentation/virt/kvm/api.rst
+++ b/Documentation/virt/kvm/api.rst
@@ -6594,7 +6594,7 @@ KVM_S390_KEYOP_SSKE
:Capability: KVM_CAP_GUEST_MEMFD_MEMORY_ATTRIBUTES
:Architectures: all
:Type: guest_memfd ioctl
-:Parameters: struct kvm_memory_attributes2 (in)
+:Parameters: struct kvm_memory_attributes2 (in/out)
:Returns: 0 on success, <0 on error
Errors:
@@ -6603,6 +6603,8 @@ Errors:
EINVAL The specified `offset` or `size` was invalid (e.g. not
page aligned, causes an overflow, or size is zero).
EFAULT The parameter address was invalid.
+ EAGAIN Some page within requested range had unexpected refcounts. The
+ offset of the page will be returned in `error_offset`.
ENOMEM Ran out of memory trying to track private/shared state
========== ===============================================================
@@ -6616,6 +6618,7 @@ Attribute values are shared with KVM_SET_MEMORY_ATTRIBUTES.
::
struct kvm_memory_attributes2 {
+ /* in */
union {
__u64 address;
__u64 offset;
@@ -6623,7 +6626,9 @@ Attribute values are shared with KVM_SET_MEMORY_ATTRIBUTES.
__u64 size;
__u64 attributes;
__u64 flags;
- __u64 reserved[12];
+ /* out */
+ __u64 error_offset;
+ __u64 reserved[11];
};
#define KVM_MEMORY_ATTRIBUTE_PRIVATE (1ULL << 3)
@@ -6645,6 +6650,16 @@ which includes operations such as unmapping pages from the host or
stage-2 page tables, may result in side effects on memory contents
that vary across different trusted firmware implementations.
+If this ioctl returns -EAGAIN, the offset of the page with unexpected
+refcounts will be returned in `error_offset`. This can occur if there
+are transient refcounts on the pages, taken by other parts of the
+kernel.
+
+Userspace is expected to figure out how to remove all known refcounts
+on the shared pages, such as refcounts taken by get_user_pages(), and
+try the ioctl again. A possible source of these long term refcounts is
+if the guest_memfd memory was pinned in IOMMU page tables.
+
See also: :ref: `KVM_SET_MEMORY_ATTRIBUTES`.
.. _kvm_run:
diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h
index ccbf7c475ec29..0ac9e6a790111 100644
--- a/include/uapi/linux/kvm.h
+++ b/include/uapi/linux/kvm.h
@@ -1662,7 +1662,8 @@ struct kvm_memory_attributes2 {
__u64 size;
__u64 attributes;
__u64 flags;
- __u64 reserved[12];
+ __u64 error_offset;
+ __u64 reserved[11];
};
#define KVM_MEMORY_ATTRIBUTE_PRIVATE (1ULL << 3)
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index 2987611587b18..e4722eb6405b6 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -538,8 +538,46 @@ static int kvm_gmem_mas_preallocate(struct ma_state *mas, u64 attributes,
return mas_preallocate(mas, xa_mk_value(attributes), GFP_KERNEL);
}
+static bool kvm_gmem_has_outstanding_references(struct inode *inode,
+ pgoff_t start, size_t nr_pages,
+ pgoff_t *err_index)
+{
+ struct address_space *mapping = inode->i_mapping;
+ pgoff_t last = start + nr_pages - 1;
+ bool has_outstanding = false;
+ struct folio_batch fbatch;
+ pgoff_t next;
+ int i;
+
+ folio_batch_init(&fbatch);
+
+ next = start;
+ while (has_outstanding && filemap_get_folios(mapping, &next, last, &fbatch)) {
+ for (i = 0; i < folio_batch_count(&fbatch); ++i) {
+ struct folio *folio = fbatch.folios[i];
+
+ /*
+ * Outstanding references are anything other than those
+ * from the page cache, plus 1 temporary reference held
+ * by filemap_get_folios() in the folio batch.
+ */
+ if (folio_ref_count(folio) != folio_nr_pages(folio) + 1) {
+ has_outstanding = true;
+ *err_index = max(start, folio->index);
+ break;
+ }
+ }
+
+ folio_batch_release(&fbatch);
+ cond_resched();
+ }
+
+ return has_outstanding;
+}
+
static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
- size_t nr_pages, uint64_t attrs)
+ size_t nr_pages, uint64_t attrs,
+ pgoff_t *err_index)
{
bool to_private = attrs & KVM_MEMORY_ATTRIBUTE_PRIVATE;
struct address_space *mapping = inode->i_mapping;
@@ -556,8 +594,28 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
mas_init(&mas, mt, start);
r = kvm_gmem_mas_preallocate(&mas, attrs, start, nr_pages);
- if (r)
+ if (r) {
+ *err_index = start;
goto out;
+ }
+
+ if (to_private) {
+ /*
+ * Forcefully unmap the pages from all userspace page tables,
+ * and then verify there are no outstanding references, e.g.
+ * acquired via GUP or similar. Tell userspace to try again if
+ * there are outstanding references and hope that whatever has
+ * pinned the page will put its reference "soon".
+ */
+ unmap_mapping_pages(mapping, start, nr_pages, false);
+
+ if (kvm_gmem_has_outstanding_references(inode, start, nr_pages,
+ err_index)) {
+ mas_destroy(&mas);
+ r = -EAGAIN;
+ goto out;
+ }
+ }
/*
* From this point on guest_memfd has performed necessary
@@ -578,9 +636,10 @@ static long kvm_gmem_set_attributes(struct file *file, void __user *argp)
struct gmem_file *f = file->private_data;
struct inode *inode = file_inode(file);
struct kvm_memory_attributes2 attrs;
+ pgoff_t err_index;
size_t nr_pages;
pgoff_t index;
- int i;
+ int i, r;
if (copy_from_user(&attrs, argp, sizeof(attrs)))
return -EFAULT;
@@ -606,8 +665,16 @@ static long kvm_gmem_set_attributes(struct file *file, void __user *argp)
nr_pages = attrs.size >> PAGE_SHIFT;
index = attrs.offset >> PAGE_SHIFT;
- return __kvm_gmem_set_attributes(inode, index, nr_pages,
- attrs.attributes);
+ r = __kvm_gmem_set_attributes(inode, index, nr_pages, attrs.attributes,
+ &err_index);
+ if (r) {
+ attrs.error_offset = ((uint64_t)err_index) << PAGE_SHIFT;
+
+ if (copy_to_user(argp, &attrs, sizeof(attrs)))
+ return -EFAULT;
+ }
+
+ return r;
}
static long kvm_gmem_ioctl(struct file *file, unsigned int ioctl,
--
2.55.0.887.g758fc8c411-goog
next prev parent reply other threads:[~2026-08-26 9:18 UTC|newest]
Thread overview: 51+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-26 9:17 [PATCH v11 00/46] guest_memfd: In-place conversion support Ackerley Tng
2026-08-26 9:17 ` [PATCH v11 01/46] KVM: guest_memfd: Optimize away conversion overheads via dead-code elimination Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 02/46] KVM: guest_memfd: Use kvm_mem_is_private() when populating guest_memfd memory Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 03/46] KVM: guest_memfd: Introduce per-gmem attributes, use to guard user mappings Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 04/46] KVM: Rename KVM_GENERIC_MEMORY_ATTRIBUTES to KVM_VM_MEMORY_ATTRIBUTES Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 05/46] KVM: Enumerate support for PRIVATE memory iff kvm_arch_has_private_mem is defined Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 06/46] KVM: Rename memory attribute APIs to prepare for in-place gmem conversion Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 07/46] KVM: Rename kvm_mem_is_private() to kvm_is_private_gfn() Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 08/46] KVM: Provide generic interface for checking memory private/shared status Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 09/46] KVM: guest_memfd: Stub in ability to enable in-place shared<=>private conversion Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 10/46] KVM: Consolidate private memory and guest_memfd ifdeffery in kvm_host.h Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 11/46] KVM: guest_memfd: Invalidate both SHARED and PRIVATE mappings for in-place conversions Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 12/46] KVM: guest_memfd: Pass mapping type filter to invalidation helper Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 13/46] KVM: guest_memfd: Add base support for KVM_SET_MEMORY_ATTRIBUTES2 Ackerley Tng
2026-08-26 9:18 ` Ackerley Tng [this message]
2026-08-26 9:18 ` [PATCH v11 15/46] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion Ackerley Tng
2026-08-26 19:44 ` Sean Christopherson
2026-08-26 22:17 ` Michael Roth
2026-08-26 22:33 ` Sean Christopherson
2026-08-26 23:47 ` Michael Roth
2026-08-26 9:18 ` [PATCH v11 16/46] KVM: guest_memfd: Return early if range already has requested attributes Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 17/46] mm/gup: factor out LRU cache draining for folio into lru_cache_drain_for_folio() Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 18/46] KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 19/46] KVM: guest_memfd: Zero page while getting pfn Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 20/46] KVM: SEV: Make 'uaddr' parameter optional for KVM_SEV_SNP_LAUNCH_UPDATE Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 21/46] KVM: TDX: Make source page optional for KVM_TDX_INIT_MEM_REGION Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 22/46] KVM: Move KVM_VM_MEMORY_ATTRIBUTES config definition to x86 Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 23/46] KVM: Let userspace disable per-VM mem attributes, enable per-gmem attributes Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 24/46] KVM: guest_memfd: Enable INIT_SHARED on guest_memfd for x86 Coco VMs Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 25/46] KVM: selftests: Create gmem fd before "regular" fd when adding memslot Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 26/46] KVM: selftests: Rename guest_memfd{,_offset} to gmem_{fd,offset} Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 27/46] KVM: selftests: Add support for mmap() on guest_memfd in core library Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 28/46] KVM: selftests: Add selftests global for guest memory attributes capability Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 29/46] KVM: selftests: Add helpers for calling ioctls on guest_memfd Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 30/46] KVM: selftests: Test basic single-page conversion flow Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 31/46] KVM: selftests: Test conversion flow when INIT_SHARED Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 32/46] KVM: selftests: Test conversion precision in guest_memfd Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 33/46] KVM: selftests: Test conversion before allocation Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 34/46] KVM: selftests: Convert with allocated folios in different layouts Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 35/46] KVM: selftests: Test that truncation does not change shared/private status Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 36/46] KVM: selftests: Test that shared/private status is consistent across processes Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 37/46] KVM: selftests: Add helpers to pin pages with CONFIG_GUP_TEST Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 38/46] KVM: selftests: Test conversion with elevated page refcount Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 39/46] KVM: selftests: Reset shared memory after hole-punching Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 40/46] KVM: selftests: Provide function to look up guest_memfd details from gpa Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 41/46] KVM: selftests: Provide common function to set memory attributes Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 42/46] KVM: selftests: Make TEST_EXPECT_SIGBUS thread-safe Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 43/46] KVM: selftests: Support guest_memfd attributes in private_mem_conversions_test Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 44/46] KVM: selftests: Set up page size and alignment independently for guest_memfd Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 45/46] KVM: selftests: Test in-place conversions in private_mem_conversions_test Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 46/46] KVM: selftests: Update private memory exits test to work with per-gmem attributes Ackerley Tng
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260826-gmem-inplace-conversion-v11-14-0a15d8a799aa@google.com \
--to=ackerleytng@google.com \
--cc=aik@amd.com \
--cc=akpm@linux-foundation.org \
--cc=andrew.jones@linux.dev \
--cc=aneesh.kumar@kernel.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=baoquan.he@linux.dev \
--cc=binbin.wu@linux.intel.com \
--cc=bp@alien8.de \
--cc=brauner@kernel.org \
--cc=chao.p.peng@linux.intel.com \
--cc=chrisl@kernel.org \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=david@kernel.org \
--cc=forkloop@google.com \
--cc=fuad.tabba@linux.dev \
--cc=hpa@zytor.com \
--cc=jgg@ziepe.ca \
--cc=jhubbard@nvidia.com \
--cc=jmattson@google.com \
--cc=jthoughton@google.com \
--cc=kas@kernel.org \
--cc=kasong@tencent.com \
--cc=kvm@vger.kernel.org \
--cc=liam@infradead.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=mathieu.desnoyers@efficios.com \
--cc=mhiramat@kernel.org \
--cc=michael.roth@amd.com \
--cc=mingo@redhat.com \
--cc=nphamcs@gmail.com \
--cc=oupton@kernel.org \
--cc=pankaj.gupta@amd.com \
--cc=pbonzini@redhat.com \
--cc=peterx@redhat.com \
--cc=pratyush@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=qperret@google.com \
--cc=rick.p.edgecombe@intel.com \
--cc=rientjes@google.com \
--cc=rostedt@goodmis.org \
--cc=seanjc@google.com \
--cc=shakeel.butt@linux.dev \
--cc=shikemeng@huaweicloud.com \
--cc=shivankg@amd.com \
--cc=shuah@kernel.org \
--cc=skhan@linuxfoundation.org \
--cc=steven.price@arm.com \
--cc=suzuki.poulose@arm.com \
--cc=tarunsahu@google.com \
--cc=tglx@kernel.org \
--cc=vannapurve@google.com \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=willy@infradead.org \
--cc=wyihan@google.com \
--cc=x86@kernel.org \
--cc=yan.y.zhao@intel.com \
--cc=youngjun.park@lge.com \
--cc=yuanchu@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox