From: Yan Zhao <yan.y.zhao@intel.com>
To: <ackerleytng@google.com>
Cc: <aik@amd.com>, <andrew.jones@linux.dev>,
<binbin.wu@linux.intel.com>, <brauner@kernel.org>,
<chao.p.peng@linux.intel.com>, <david@kernel.org>,
<jmattson@google.com>, <jthoughton@google.com>,
<michael.roth@amd.com>, <oupton@kernel.org>,
<pankaj.gupta@amd.com>, <qperret@google.com>,
<rick.p.edgecombe@intel.com>, <rientjes@google.com>,
<shivankg@amd.com>, <steven.price@arm.com>, <tabba@google.com>,
<willy@infradead.org>, <wyihan@google.com>, <forkloop@google.com>,
<pratyush@kernel.org>, <suzuki.poulose@arm.com>,
<aneesh.kumar@kernel.org>, <liam@infradead.org>,
Paolo Bonzini <pbonzini@redhat.com>,
Sean Christopherson <seanjc@google.com>,
Thomas Gleixner <tglx@kernel.org>, Ingo Molnar <mingo@redhat.com>,
"Borislav Petkov" <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>, <x86@kernel.org>,
"H. Peter Anvin" <hpa@zytor.com>,
Steven Rostedt <rostedt@goodmis.org>,
Masami Hiramatsu <mhiramat@kernel.org>,
"Mathieu Desnoyers" <mathieu.desnoyers@efficios.com>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Shuah Khan <shuah@kernel.org>,
"Vishal Annapurve" <vannapurve@google.com>,
Andrew Morton <akpm@linux-foundation.org>,
Chris Li <chrisl@kernel.org>, Kairui Song <kasong@tencent.com>,
Kemeng Shi <shikemeng@huaweicloud.com>,
Nhat Pham <nphamcs@gmail.com>, Barry Song <baohua@kernel.org>,
Axel Rasmussen <axelrasmussen@google.com>,
Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
Youngjun Park <youngjun.park@lge.com>,
Qi Zheng <qi.zheng@linux.dev>,
Shakeel Butt <shakeel.butt@linux.dev>,
Kiryl Shutsemau <kas@kernel.org>,
Baoquan He <baoquan.he@linux.dev>, Jason Gunthorpe <jgg@ziepe.ca>,
John Hubbard <jhubbard@nvidia.com>, Peter Xu <peterx@redhat.com>,
<tarunsahu@google.com>, Vlastimil Babka <vbabka@kernel.org>,
<kvm@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
<linux-trace-kernel@vger.kernel.org>, <linux-doc@vger.kernel.org>,
<linux-kselftest@vger.kernel.org>, <linux-mm@kvack.org>,
<linux-coco@lists.linux.dev>
Subject: Re: [PATCH v10 11/41] KVM: guest_memfd: Ensure pages are not in use before conversion
Date: Sat, 8 Aug 2026 08:29:15 +0800 [thread overview]
Message-ID: <anZ4W9o5pTWIEgMY@yzhao56-desk.sh.intel.com> (raw)
In-Reply-To: <20260807-gmem-inplace-conversion-v10-11-2fc18ee6d3ba@google.com>
On Fri, Aug 07, 2026 at 02:52:50PM -0700, Ackerley Tng via B4 Relay wrote:
> From: Ackerley Tng <ackerleytng@google.com>
>
> When converting memory to private in guest_memfd, it is necessary to ensure
> that the pages are not currently being accessed by any other part of the
> kernel or userspace to avoid any current user writing to guest private
> memory.
>
> guest_memfd checks for unexpected refcounts to determine whether a page is
> still in use. The only expected refcounts after unmapping the range
> requested for conversion are those that are held by guest_memfd itself.
>
> Update the kvm_memory_attributes2 structure to include an error_offset
> field. This allows KVM to report the exact offset where a conversion
> failed to userspace. If the safety check fails, return -EAGAIN and copy
> the error_offset back to userspace so that it can potentially retry the
> operation or handle the failure gracefully.
>
> Update documentation to document the error_offset field and the possible
> -EAGAIN error.
>
> Suggested-by: David Hildenbrand <david@kernel.org>
> Co-developed-by: Vishal Annapurve <vannapurve@google.com>
> Signed-off-by: Vishal Annapurve <vannapurve@google.com>
> Reviewed-by: Fuad Tabba <tabba@google.com>
> Tested-by: Shivank Garg <shivankg@amd.com>
> Signed-off-by: Ackerley Tng <ackerleytng@google.com>
> ---
> Documentation/virt/kvm/api.rst | 19 ++++++++++--
> include/uapi/linux/kvm.h | 3 +-
> virt/kvm/guest_memfd.c | 66 ++++++++++++++++++++++++++++++++++++++----
> 3 files changed, 80 insertions(+), 8 deletions(-)
>
> diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
> index 1a3f664dbb197..1e64026d7c1e9 100644
> --- a/Documentation/virt/kvm/api.rst
> +++ b/Documentation/virt/kvm/api.rst
> @@ -6583,7 +6583,7 @@ KVM_S390_KEYOP_SSKE
> :Capability: KVM_CAP_GUEST_MEMFD_MEMORY_ATTRIBUTES
> :Architectures: all
> :Type: guest_memfd ioctl
> -:Parameters: struct kvm_memory_attributes2 (in)
> +:Parameters: struct kvm_memory_attributes2 (in/out)
> :Returns: 0 on success, <0 on error
>
> Errors:
> @@ -6592,6 +6592,8 @@ Errors:
> EINVAL The specified `offset` or `size` was invalid (e.g. not
> page aligned, causes an overflow, or size is zero).
> EFAULT The parameter address was invalid.
> + EAGAIN Some page within requested range had unexpected refcounts. The
> + offset of the page will be returned in `error_offset`.
> ENOMEM Ran out of memory trying to track private/shared state
> ========== ===============================================================
>
> @@ -6605,6 +6607,7 @@ Attribute values are shared with KVM_SET_MEMORY_ATTRIBUTES.
> ::
>
> struct kvm_memory_attributes2 {
> + /* in */
> union {
> __u64 address;
> __u64 offset;
> @@ -6612,7 +6615,9 @@ Attribute values are shared with KVM_SET_MEMORY_ATTRIBUTES.
> __u64 size;
> __u64 attributes;
> __u64 flags;
> - __u64 reserved[12];
> + /* out */
> + __u64 error_offset;
> + __u64 reserved[11];
> };
>
> #define KVM_MEMORY_ATTRIBUTE_PRIVATE (1ULL << 3)
> @@ -6634,6 +6639,16 @@ which includes operations such as unmapping pages from the host or
> stage-2 page tables, may result in side effects on memory contents
> that vary across different trusted firmware implementations.
>
> +If this ioctl returns -EAGAIN, the offset of the page with unexpected
> +refcounts will be returned in `error_offset`. This can occur if there
> +are transient refcounts on the pages, taken by other parts of the
> +kernel.
> +
> +Userspace is expected to figure out how to remove all known refcounts
> +on the shared pages, such as refcounts taken by get_user_pages(), and
> +try the ioctl again. A possible source of these long term refcounts is
> +if the guest_memfd memory was pinned in IOMMU page tables.
> +
> See also: :ref: `KVM_SET_MEMORY_ATTRIBUTES`.
>
> .. _kvm_run:
> diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h
> index 80985e28e3b21..129d6f6303251 100644
> --- a/include/uapi/linux/kvm.h
> +++ b/include/uapi/linux/kvm.h
> @@ -1661,7 +1661,8 @@ struct kvm_memory_attributes2 {
> __u64 size;
> __u64 attributes;
> __u64 flags;
> - __u64 reserved[12];
> + __u64 error_offset;
> + __u64 reserved[11];
> };
>
> #define KVM_MEMORY_ATTRIBUTE_PRIVATE (1ULL << 3)
> diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
> index 3783e63476569..13c3989136f67 100644
> --- a/virt/kvm/guest_memfd.c
> +++ b/virt/kvm/guest_memfd.c
> @@ -524,8 +524,42 @@ static int kvm_gmem_mas_preallocate(struct ma_state *mas, u64 attributes,
> return mas_preallocate(mas, xa_mk_value(attributes), GFP_KERNEL);
> }
>
> +static bool kvm_gmem_is_safe_for_conversion(struct inode *inode, pgoff_t start,
> + size_t nr_pages, pgoff_t *err_index)
> +{
> + struct address_space *mapping = inode->i_mapping;
> + const int filemap_get_folios_refcount = 1;
> + pgoff_t last = start + nr_pages - 1;
> + struct folio_batch fbatch;
> + bool safe = true;
> + pgoff_t next;
> + int i;
> +
> + folio_batch_init(&fbatch);
> +
> + next = start;
> + while (safe && filemap_get_folios(mapping, &next, last, &fbatch)) {
> + for (i = 0; i < folio_batch_count(&fbatch); ++i) {
> + struct folio *folio = fbatch.folios[i];
> +
> + if (folio_ref_count(folio) !=
> + folio_nr_pages(folio) + filemap_get_folios_refcount) {
> + safe = false;
> + *err_index = max(start, folio->index);
> + break;
> + }
> + }
> +
> + folio_batch_release(&fbatch);
> + cond_resched();
> + }
> +
> + return safe;
> +}
> +
> static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
> - size_t nr_pages, uint64_t attrs)
> + size_t nr_pages, uint64_t attrs,
> + pgoff_t *err_index)
> {
> bool to_private = attrs & KVM_MEMORY_ATTRIBUTE_PRIVATE;
> struct address_space *mapping = inode->i_mapping;
> @@ -542,8 +576,21 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
>
> mas_init(&mas, mt, start);
> r = kvm_gmem_mas_preallocate(&mas, attrs, start, nr_pages);
> - if (r)
> + if (r) {
> + *err_index = start;
> goto out;
> + }
> +
> + if (to_private) {
> + unmap_mapping_pages(mapping, start, nr_pages, false);
> +
> + if (!kvm_gmem_is_safe_for_conversion(inode, start, nr_pages,
> + err_index)) {
Note: conversion failures could occur if another vCPU is attempting to map a GFN
within this range.
CPU 0 (setting attributes) CPU 1 (attempting to map)
-------------------------- --------------------
A: mmu_invalidate_retry_gfn_unsafe
filemap_invalidate_lock_shared
__kvm_gmem_get_pfn ==> folio refcount++
filemap_invalidate_unlock_shared
filemap_invalidate_lock
filemap_get_folios
check folio_ref_count(folio) ==> Not match !!
filemap_invalidate_unlock
B: read_lock(&vcpu->kvm->mmu_lock);
is_page_fault_stale
kvm_mmu_finish_page_fault ==>folio recount--
read_unlock(&vcpu->kvm->mmu_lock);
Retrying in kvm_gmem_is_safe_for_conversion() or moving the invocation of
kvm_mmu_invalidate_start() + kvm_mmu_invalidate_range_add() to an earlier
position does not help as long as CPU 1 stays at stage A.
So, should we avoid this failure?
e.g., by moving filemap_invalidate_unlock_shared() from stage A to after
stage B?
> + mas_destroy(&mas);
> + r = -EAGAIN;
> + goto out;
> + }
> + }
>
> /*
> * From this point on guest_memfd has performed necessary
> @@ -564,9 +611,10 @@ static long kvm_gmem_set_attributes(struct file *file, void __user *argp)
> struct gmem_file *f = file->private_data;
> struct inode *inode = file_inode(file);
> struct kvm_memory_attributes2 attrs;
> + pgoff_t err_index;
> size_t nr_pages;
> pgoff_t index;
> - int i;
> + int i, r;
>
> if (copy_from_user(&attrs, argp, sizeof(attrs)))
> return -EFAULT;
> @@ -592,8 +640,16 @@ static long kvm_gmem_set_attributes(struct file *file, void __user *argp)
>
> nr_pages = attrs.size >> PAGE_SHIFT;
> index = attrs.offset >> PAGE_SHIFT;
> - return __kvm_gmem_set_attributes(inode, index, nr_pages,
> - attrs.attributes);
> + r = __kvm_gmem_set_attributes(inode, index, nr_pages, attrs.attributes,
> + &err_index);
> + if (r) {
> + attrs.error_offset = ((uint64_t)err_index) << PAGE_SHIFT;
> +
> + if (copy_to_user(argp, &attrs, sizeof(attrs)))
> + return -EFAULT;
> + }
> +
> + return r;
> }
>
> static long kvm_gmem_ioctl(struct file *file, unsigned int ioctl,
>
> --
> 2.55.0.654.g21b8a5bc05-goog
>
>
next prev parent reply other threads:[~2026-08-08 1:10 UTC|newest]
Thread overview: 43+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-07 21:52 [PATCH v10 00/41] guest_memfd: In-place conversion support Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 01/41] KVM: guest_memfd: Use kvm_mem_is_private() when populating guest_memfd memory Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 02/41] KVM: guest_memfd: Introduce per-gmem attributes, use to guard user mappings Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 03/41] KVM: Rename KVM_GENERIC_MEMORY_ATTRIBUTES to KVM_VM_MEMORY_ATTRIBUTES Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 04/41] KVM: Enumerate support for PRIVATE memory iff kvm_arch_has_private_mem is defined Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 05/41] KVM: Rename memory attribute APIs to prepare for in-place gmem conversion Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 06/41] KVM: Provide generic interface for checking memory private/shared status Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 07/41] KVM: guest_memfd: Stub in ability to enable in-place shared<=>private conversion Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 08/41] KVM: Consolidate private memory and guest_memfd ifdeffery in kvm_host.h Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 09/41] KVM: guest_memfd: Filter both shared and private when invalidating Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 10/41] KVM: guest_memfd: Add base support for KVM_SET_MEMORY_ATTRIBUTES2 Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 11/41] KVM: guest_memfd: Ensure pages are not in use before conversion Ackerley Tng via B4 Relay
2026-08-08 0:29 ` Yan Zhao [this message]
2026-08-07 21:52 ` [PATCH v10 12/41] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 13/41] KVM: guest_memfd: Return early if range already has requested attributes Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 14/41] mm/gup: factor out LRU cache draining for folio into lru_cache_drain_for_folio() Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 15/41] KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 16/41] KVM: guest_memfd: Zero page while getting pfn Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 17/41] KVM: SEV: Make 'uaddr' parameter optional for KVM_SEV_SNP_LAUNCH_UPDATE Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 18/41] KVM: TDX: Make source page optional for KVM_TDX_INIT_MEM_REGION Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 19/41] KVM: Move KVM_VM_MEMORY_ATTRIBUTES config definition to x86 Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 20/41] KVM: Let userspace disable per-VM mem attributes, enable per-gmem attributes Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 21/41] KVM: guest_memfd: Enable INIT_SHARED on guest_memfd for x86 Coco VMs Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 22/41] KVM: selftests: Create gmem fd before "regular" fd when adding memslot Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 23/41] KVM: selftests: Rename guest_memfd{,_offset} to gmem_{fd,offset} Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 24/41] KVM: selftests: Add support for mmap() on guest_memfd in core library Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 25/41] KVM: selftests: Add selftests global for guest memory attributes capability Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 26/41] KVM: selftests: Add helpers for calling ioctls on guest_memfd Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 27/41] KVM: selftests: Test basic single-page conversion flow Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 28/41] KVM: selftests: Test conversion flow when INIT_SHARED Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 29/41] KVM: selftests: Test conversion precision in guest_memfd Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 30/41] KVM: selftests: Test conversion before allocation Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 31/41] KVM: selftests: Convert with allocated folios in different layouts Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 32/41] KVM: selftests: Test that truncation does not change shared/private status Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 33/41] KVM: selftests: Test that shared/private status is consistent across processes Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 34/41] KVM: selftests: Add helpers to pin pages with CONFIG_GUP_TEST Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 35/41] KVM: selftests: Test conversion with elevated page refcount Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 36/41] KVM: selftests: Reset shared memory after hole-punching Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 37/41] KVM: selftests: Provide function to look up guest_memfd details from gpa Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 38/41] KVM: selftests: Provide common function to set memory attributes Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 39/41] KVM: selftests: Make TEST_EXPECT_SIGBUS thread-safe Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 40/41] KVM: selftests: Update private_mem_conversions_test to mmap() guest_memfd Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 41/41] KVM: selftests: Update private memory exits test to work with per-gmem attributes Ackerley Tng via B4 Relay
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anZ4W9o5pTWIEgMY@yzhao56-desk.sh.intel.com \
--to=yan.y.zhao@intel.com \
--cc=ackerleytng@google.com \
--cc=aik@amd.com \
--cc=akpm@linux-foundation.org \
--cc=andrew.jones@linux.dev \
--cc=aneesh.kumar@kernel.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=baoquan.he@linux.dev \
--cc=binbin.wu@linux.intel.com \
--cc=bp@alien8.de \
--cc=brauner@kernel.org \
--cc=chao.p.peng@linux.intel.com \
--cc=chrisl@kernel.org \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=david@kernel.org \
--cc=forkloop@google.com \
--cc=hpa@zytor.com \
--cc=jgg@ziepe.ca \
--cc=jhubbard@nvidia.com \
--cc=jmattson@google.com \
--cc=jthoughton@google.com \
--cc=kas@kernel.org \
--cc=kasong@tencent.com \
--cc=kvm@vger.kernel.org \
--cc=liam@infradead.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=mathieu.desnoyers@efficios.com \
--cc=mhiramat@kernel.org \
--cc=michael.roth@amd.com \
--cc=mingo@redhat.com \
--cc=nphamcs@gmail.com \
--cc=oupton@kernel.org \
--cc=pankaj.gupta@amd.com \
--cc=pbonzini@redhat.com \
--cc=peterx@redhat.com \
--cc=pratyush@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=qperret@google.com \
--cc=rick.p.edgecombe@intel.com \
--cc=rientjes@google.com \
--cc=rostedt@goodmis.org \
--cc=seanjc@google.com \
--cc=shakeel.butt@linux.dev \
--cc=shikemeng@huaweicloud.com \
--cc=shivankg@amd.com \
--cc=shuah@kernel.org \
--cc=skhan@linuxfoundation.org \
--cc=steven.price@arm.com \
--cc=suzuki.poulose@arm.com \
--cc=tabba@google.com \
--cc=tarunsahu@google.com \
--cc=tglx@kernel.org \
--cc=vannapurve@google.com \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=willy@infradead.org \
--cc=wyihan@google.com \
--cc=x86@kernel.org \
--cc=youngjun.park@lge.com \
--cc=yuanchu@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox