From: Sean Christopherson <seanjc@google.com>
To: Ackerley Tng <ackerleytng@google.com>
Cc: "David Hildenbrand (Arm)" <david@kernel.org>,
aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com,
brauner@kernel.org, chao.p.peng@linux.intel.com,
jmattson@google.com, jthoughton@google.com,
michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com,
qperret@google.com, rick.p.edgecombe@intel.com,
rientjes@google.com, shivankg@amd.com, steven.price@arm.com,
tabba@google.com, willy@infradead.org, wyihan@google.com,
yan.y.zhao@intel.com, forkloop@google.com, pratyush@kernel.org,
suzuki.poulose@arm.com, aneesh.kumar@kernel.org,
liam@infradead.org, Paolo Bonzini <pbonzini@redhat.com>,
Thomas Gleixner <tglx@kernel.org>,
Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
x86@kernel.org, "H. Peter Anvin" <hpa@zytor.com>,
Steven Rostedt <rostedt@goodmis.org>,
Masami Hiramatsu <mhiramat@kernel.org>,
Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Shuah Khan <shuah@kernel.org>,
Vishal Annapurve <vannapurve@google.com>,
Andrew Morton <akpm@linux-foundation.org>,
Chris Li <chrisl@kernel.org>, Kairui Song <kasong@tencent.com>,
Kemeng Shi <shikemeng@huaweicloud.com>,
Nhat Pham <nphamcs@gmail.com>, Barry Song <baohua@kernel.org>,
Axel Rasmussen <axelrasmussen@google.com>,
Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
Youngjun Park <youngjun.park@lge.com>,
Qi Zheng <qi.zheng@linux.dev>,
Shakeel Butt <shakeel.butt@linux.dev>,
Kiryl Shutsemau <kas@kernel.org>,
Baoquan He <baoquan.he@linux.dev>, Jason Gunthorpe <jgg@ziepe.ca>,
John Hubbard <jhubbard@nvidia.com>, Peter Xu <peterx@redhat.com>,
tarunsahu@google.com, Vlastimil Babka <vbabka@kernel.org>,
kvm@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org,
linux-kselftest@vger.kernel.org, linux-mm@kvack.org,
linux-coco@lists.linux.dev
Subject: Re: [PATCH v10 11/41] KVM: guest_memfd: Ensure pages are not in use before conversion
Date: Mon, 10 Aug 2026 15:26:46 -0700 [thread overview]
Message-ID: <anpQJvFw9s-XP3pW@google.com> (raw)
In-Reply-To: <CAEvNRgHyRdt7rp=5e5FMbTT=fXNfB=np7VWd_GdmJcPOQgps4g@mail.gmail.com>
On Mon, Aug 10, 2026, Ackerley Tng wrote:
> "David Hildenbrand (Arm)" <david@kernel.org> writes:
>
> > On 8/7/26 23:52, Ackerley Tng via B4 Relay wrote:
> >> From: Ackerley Tng <ackerleytng@google.com>
> >>
> >> When converting memory to private in guest_memfd, it is necessary to ensure
> >> that the pages are not currently being accessed by any other part of the
> >> kernel or userspace to avoid any current user writing to guest private
> >> memory.
> >>
> >> guest_memfd checks for unexpected refcounts to determine whether a page is
> >> still in use. The only expected refcounts after unmapping the range
> >> requested for conversion are those that are held by guest_memfd itself.
> >>
> >> Update the kvm_memory_attributes2 structure to include an error_offset
> >> field. This allows KVM to report the exact offset where a conversion
> >> failed to userspace. If the safety check fails, return -EAGAIN and copy
> >> the error_offset back to userspace so that it can potentially retry the
> >> operation or handle the failure gracefully.
> >>
> >> Update documentation to document the error_offset field and the possible
> >> -EAGAIN error.
> >>
> >> Suggested-by: David Hildenbrand <david@kernel.org>
> >> Co-developed-by: Vishal Annapurve <vannapurve@google.com>
> >> Signed-off-by: Vishal Annapurve <vannapurve@google.com>
> >> Reviewed-by: Fuad Tabba <tabba@google.com>
> >> Tested-by: Shivank Garg <shivankg@amd.com>
> >> Signed-off-by: Ackerley Tng <ackerleytng@google.com>
> >> ---
> >
> > [...]
> >
> >> #define KVM_MEMORY_ATTRIBUTE_PRIVATE (1ULL << 3)
> >> diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
> >> index 3783e63476569..13c3989136f67 100644
> >> --- a/virt/kvm/guest_memfd.c
> >> +++ b/virt/kvm/guest_memfd.c
> >> @@ -524,8 +524,42 @@ static int kvm_gmem_mas_preallocate(struct ma_state *mas, u64 attributes,
> >> return mas_preallocate(mas, xa_mk_value(attributes), GFP_KERNEL);
> >> }
> >>
> >> +static bool kvm_gmem_is_safe_for_conversion(struct inode *inode, pgoff_t start,
> >> + size_t nr_pages, pgoff_t *err_index)
> >
> > I would focus on the "to_private" aspect or abstract it to
> > "kvm_gmem_mem_has_unexpected_refs" or sth like that.
+1. Maybe "kvm_gmem_page_has_outstanding_references"?
> Do you mean something like kvm_gmem_is_safe_for_to_private_conversion,
> as in that you want to emphasise that "safe" here refers to a to_private
> and not a to_shared conversion?
I'm obviously not David, but for me, the problem with names like
kvm_gmem_is_safe_for_conversion() is that (a) it conflates what the function is
literally doing with how the function is being used, which often makes the code
harder to understand as it obfuscates things, and (b) can become stale or even
outright broken far too easily. E.g. if KVM adds more checks on whether
or not a conversion is "safe", then the name of the function is a lie because it
doesn't actually check that the target data is safe for conversion, only that
its "safe" for a specific aspect of conversion.
And there is real risk to hiding what a function does. E.g. looking at this code
without diving into the details:
if (to_private) {
unmap_mapping_pages(mapping, start, nr_pages, false);
if (!kvm_gmem_is_safe_for_conversion(inode, start, nr_pages,
err_index)) {
mas_destroy(&mas);
r = -EAGAIN;
goto out;
}
}
and one might thing that it's perfectly find to check for "safety" before
unmapping pages. In fact, looking at the code without a priori knowledge of
the rules, and the above flat out looks wrong. E.g. I could definitely see
someone "fixing" the code to:
if (to_private) {
if (!kvm_gmem_is_safe_for_conversion(inode, start, nr_pages,
err_index)) {
mas_destroy(&mas);
r = -EAGAIN;
goto out;
}
unmap_mapping_pages(mapping, start, nr_pages, false);
}
Whereas this:
if (to_private) {
unmap_mapping_pages(mapping, start, nr_pages, false);
if (kvm_gmem_has_outstanding_references(inode, start, nr_pages,
err_index)) {
mas_destroy(&mas);
r = -EAGAIN;
goto out;
}
}
helps the reader understand what's being checked without having to look at the
details, and also helps communicate the ordering dependency without needing a
comment.
There are definitely times where the usage of a function bleeds into its name,
but usually that's because the name and the usage are on and the same. E.g.
get_user() describes both the usage and the "what". And it's easy/possible to
go too far in the opposite direction, e.g. by giving a play-by-play of what a
function is doing, but that's why we have bikshedding sessions :-)
> Perhaps a little ahead of its time,
Ya.
> but later with restructuring for huge pages, we also need no additional
> refcounts other than gmem's own so that restructuring is safe, hence this
> function name was meant to extend there as well.
Given that I've read that at least five times and still don't understand the
nuance, I think it's safe (ha!) to say we'll need to revisit and review those
changes no matter what. :-)
> In this case "unexpected" (especially since the next patch adds checks
> for maybe dma pinned and unmapping), begs the question "unexpected in
> what way"?
Ya, that's why I like "outstanding", it succinctly captures that one or more
references have been "loaned" but not yet "repaid".
> >> + struct address_space *mapping = inode->i_mapping;
> >> + const int filemap_get_folios_refcount = 1;
> >> + pgoff_t last = start + nr_pages - 1;
> >> + struct folio_batch fbatch;
> >> + bool safe = true;
> >> + pgoff_t next;
> >> + int i;
> >> +
> >> + folio_batch_init(&fbatch);
> >> +
> >> + next = start;
> >> + while (safe && filemap_get_folios(mapping, &next, last, &fbatch)) {
> >> + for (i = 0; i < folio_batch_count(&fbatch); ++i) {
> >> + struct folio *folio = fbatch.folios[i];
> >> +
> >> + if (folio_ref_count(folio) !=
> >> + folio_nr_pages(folio) + filemap_get_folios_refcount) {
> >
> > I'd rather add a comment than have this filemap_get_folios_refcount.
+1, the local variable just made me scratch my head.
> >
> > /*
> > * We expect one reference per folio-page in the pagecache and one
> > * reference from filemap_get_folios().
Nit, please no pronouns in KVM code.
> > */
> > if (folio_ref_count(folio) != folio_nr_pages(folio) + 1)
> >
>
> This comment explains what's "unexpected". I can do this and switch it
> to kvm_gmem_mem_has_unexpected_refs() unless people have other
> suggestions.
>
> I wish there was a folio_pagecache_refs(folio) that
> folio_expected_ref_count() can share with this, and also
> folio_swapcache_refs(), to solidify the definition of refcounts taken by
> the pagecache.
...
> >> @@ -542,8 +576,21 @@ static int __kvm_gmem_set_attributes(struct inode *inode, pgoff_t start,
> >>
> >> mas_init(&mas, mt, start);
> >> r = kvm_gmem_mas_preallocate(&mas, attrs, start, nr_pages);
> >> - if (r)
> >> + if (r) {
> >> + *err_index = start;
> >> goto out;
> >> + }
> >> +
> >> + if (to_private) {
> >
> > I'd add a comment here for the "why are we unmapping".
> >
>
> Does this sound right:
>
> Unmap here to ensure that userspace page tables have no mappings, which
> also ensures refcounts from those mappings are dropped.
How about:
/*
* Forcefully unmap the pages from all userspace page tables,
* and then verify there are no outstanding references, e.g.
* acquired via GUP or similar. Tell userspace to try again if
* there are oustanding references and hope that whatever has
* pinned the page will put its reference "soon".
*/
unmap_mapping_pages(mapping, start, nr_pages, false);
if (!kvm_gmem_is_safe_for_conversion(inode, start, nr_pages,
err_index)) {
mas_destroy(&mas);
r = -EAGAIN;
goto out;
}
>
> > --
> > Cheers,
> >
> > David
next prev parent reply other threads:[~2026-08-10 22:26 UTC|newest]
Thread overview: 67+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-07 21:52 [PATCH v10 00/41] guest_memfd: In-place conversion support Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 01/41] KVM: guest_memfd: Use kvm_mem_is_private() when populating guest_memfd memory Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 02/41] KVM: guest_memfd: Introduce per-gmem attributes, use to guard user mappings Ackerley Tng via B4 Relay
2026-08-10 15:12 ` Sean Christopherson
2026-08-07 21:52 ` [PATCH v10 03/41] KVM: Rename KVM_GENERIC_MEMORY_ATTRIBUTES to KVM_VM_MEMORY_ATTRIBUTES Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 04/41] KVM: Enumerate support for PRIVATE memory iff kvm_arch_has_private_mem is defined Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 05/41] KVM: Rename memory attribute APIs to prepare for in-place gmem conversion Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 06/41] KVM: Provide generic interface for checking memory private/shared status Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 07/41] KVM: guest_memfd: Stub in ability to enable in-place shared<=>private conversion Ackerley Tng via B4 Relay
2026-08-10 8:49 ` David Hildenbrand (Arm)
2026-08-10 15:01 ` Sean Christopherson
2026-08-07 21:52 ` [PATCH v10 08/41] KVM: Consolidate private memory and guest_memfd ifdeffery in kvm_host.h Ackerley Tng via B4 Relay
2026-08-10 8:50 ` David Hildenbrand (Arm)
2026-08-07 21:52 ` [PATCH v10 09/41] KVM: guest_memfd: Filter both shared and private when invalidating Ackerley Tng via B4 Relay
2026-08-10 9:06 ` David Hildenbrand (Arm)
2026-08-10 9:27 ` Suzuki K Poulose
2026-08-10 19:11 ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 10/41] KVM: guest_memfd: Add base support for KVM_SET_MEMORY_ATTRIBUTES2 Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 11/41] KVM: guest_memfd: Ensure pages are not in use before conversion Ackerley Tng via B4 Relay
2026-08-08 0:29 ` Yan Zhao
2026-08-09 21:51 ` Yan Zhao
2026-08-10 21:06 ` Ackerley Tng
2026-08-11 1:04 ` Yan Zhao
2026-08-11 2:17 ` Ackerley Tng
2026-08-11 4:50 ` Yan Zhao
2026-08-10 9:19 ` David Hildenbrand (Arm)
2026-08-10 21:41 ` Ackerley Tng
2026-08-10 22:26 ` Sean Christopherson [this message]
2026-08-07 21:52 ` [PATCH v10 12/41] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 13/41] KVM: guest_memfd: Return early if range already has requested attributes Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 14/41] mm/gup: factor out LRU cache draining for folio into lru_cache_drain_for_folio() Ackerley Tng via B4 Relay
2026-08-10 9:22 ` David Hildenbrand (Arm)
2026-08-07 21:52 ` [PATCH v10 15/41] KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check Ackerley Tng via B4 Relay
2026-08-10 9:25 ` David Hildenbrand (Arm)
2026-08-10 21:29 ` Ackerley Tng
2026-08-07 21:52 ` [PATCH v10 16/41] KVM: guest_memfd: Zero page while getting pfn Ackerley Tng via B4 Relay
2026-08-10 9:30 ` David Hildenbrand (Arm)
2026-08-10 23:41 ` Ackerley Tng
2026-08-11 0:18 ` Sean Christopherson
2026-08-07 21:52 ` [PATCH v10 17/41] KVM: SEV: Make 'uaddr' parameter optional for KVM_SEV_SNP_LAUNCH_UPDATE Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 18/41] KVM: TDX: Make source page optional for KVM_TDX_INIT_MEM_REGION Ackerley Tng via B4 Relay
2026-08-07 21:52 ` [PATCH v10 19/41] KVM: Move KVM_VM_MEMORY_ATTRIBUTES config definition to x86 Ackerley Tng via B4 Relay
2026-08-10 9:32 ` David Hildenbrand (Arm)
2026-08-07 21:52 ` [PATCH v10 20/41] KVM: Let userspace disable per-VM mem attributes, enable per-gmem attributes Ackerley Tng via B4 Relay
2026-08-10 9:36 ` David Hildenbrand (Arm)
2026-08-07 21:53 ` [PATCH v10 21/41] KVM: guest_memfd: Enable INIT_SHARED on guest_memfd for x86 Coco VMs Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 22/41] KVM: selftests: Create gmem fd before "regular" fd when adding memslot Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 23/41] KVM: selftests: Rename guest_memfd{,_offset} to gmem_{fd,offset} Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 24/41] KVM: selftests: Add support for mmap() on guest_memfd in core library Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 25/41] KVM: selftests: Add selftests global for guest memory attributes capability Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 26/41] KVM: selftests: Add helpers for calling ioctls on guest_memfd Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 27/41] KVM: selftests: Test basic single-page conversion flow Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 28/41] KVM: selftests: Test conversion flow when INIT_SHARED Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 29/41] KVM: selftests: Test conversion precision in guest_memfd Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 30/41] KVM: selftests: Test conversion before allocation Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 31/41] KVM: selftests: Convert with allocated folios in different layouts Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 32/41] KVM: selftests: Test that truncation does not change shared/private status Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 33/41] KVM: selftests: Test that shared/private status is consistent across processes Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 34/41] KVM: selftests: Add helpers to pin pages with CONFIG_GUP_TEST Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 35/41] KVM: selftests: Test conversion with elevated page refcount Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 36/41] KVM: selftests: Reset shared memory after hole-punching Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 37/41] KVM: selftests: Provide function to look up guest_memfd details from gpa Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 38/41] KVM: selftests: Provide common function to set memory attributes Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 39/41] KVM: selftests: Make TEST_EXPECT_SIGBUS thread-safe Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 40/41] KVM: selftests: Update private_mem_conversions_test to mmap() guest_memfd Ackerley Tng via B4 Relay
2026-08-07 21:53 ` [PATCH v10 41/41] KVM: selftests: Update private memory exits test to work with per-gmem attributes Ackerley Tng via B4 Relay
2026-08-10 8:43 ` [PATCH v10 00/41] guest_memfd: In-place conversion support David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anpQJvFw9s-XP3pW@google.com \
--to=seanjc@google.com \
--cc=ackerleytng@google.com \
--cc=aik@amd.com \
--cc=akpm@linux-foundation.org \
--cc=andrew.jones@linux.dev \
--cc=aneesh.kumar@kernel.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=baoquan.he@linux.dev \
--cc=binbin.wu@linux.intel.com \
--cc=bp@alien8.de \
--cc=brauner@kernel.org \
--cc=chao.p.peng@linux.intel.com \
--cc=chrisl@kernel.org \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=david@kernel.org \
--cc=forkloop@google.com \
--cc=hpa@zytor.com \
--cc=jgg@ziepe.ca \
--cc=jhubbard@nvidia.com \
--cc=jmattson@google.com \
--cc=jthoughton@google.com \
--cc=kas@kernel.org \
--cc=kasong@tencent.com \
--cc=kvm@vger.kernel.org \
--cc=liam@infradead.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=mathieu.desnoyers@efficios.com \
--cc=mhiramat@kernel.org \
--cc=michael.roth@amd.com \
--cc=mingo@redhat.com \
--cc=nphamcs@gmail.com \
--cc=oupton@kernel.org \
--cc=pankaj.gupta@amd.com \
--cc=pbonzini@redhat.com \
--cc=peterx@redhat.com \
--cc=pratyush@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=qperret@google.com \
--cc=rick.p.edgecombe@intel.com \
--cc=rientjes@google.com \
--cc=rostedt@goodmis.org \
--cc=shakeel.butt@linux.dev \
--cc=shikemeng@huaweicloud.com \
--cc=shivankg@amd.com \
--cc=shuah@kernel.org \
--cc=skhan@linuxfoundation.org \
--cc=steven.price@arm.com \
--cc=suzuki.poulose@arm.com \
--cc=tabba@google.com \
--cc=tarunsahu@google.com \
--cc=tglx@kernel.org \
--cc=vannapurve@google.com \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=willy@infradead.org \
--cc=wyihan@google.com \
--cc=x86@kernel.org \
--cc=yan.y.zhao@intel.com \
--cc=youngjun.park@lge.com \
--cc=yuanchu@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox