From: Ackerley Tng <ackerleytng@google.com>
To: aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com,
brauner@kernel.org, chao.p.peng@linux.intel.com,
david@kernel.org, jmattson@google.com, jthoughton@google.com,
michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com,
qperret@google.com, rick.p.edgecombe@intel.com,
rientjes@google.com, shivankg@amd.com, steven.price@arm.com,
willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com,
forkloop@google.com, pratyush@kernel.org,
suzuki.poulose@arm.com, aneesh.kumar@kernel.org,
liam@infradead.org, Paolo Bonzini <pbonzini@redhat.com>,
Sean Christopherson <seanjc@google.com>,
Thomas Gleixner <tglx@kernel.org>,
Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
x86@kernel.org, "H. Peter Anvin" <hpa@zytor.com>,
Steven Rostedt <rostedt@goodmis.org>,
Masami Hiramatsu <mhiramat@kernel.org>,
Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Shuah Khan <shuah@kernel.org>,
Vishal Annapurve <vannapurve@google.com>,
Andrew Morton <akpm@linux-foundation.org>,
Chris Li <chrisl@kernel.org>, Kairui Song <kasong@tencent.com>,
Kemeng Shi <shikemeng@huaweicloud.com>,
Nhat Pham <nphamcs@gmail.com>, Barry Song <baohua@kernel.org>,
Axel Rasmussen <axelrasmussen@google.com>,
Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
Youngjun Park <youngjun.park@lge.com>,
Qi Zheng <qi.zheng@linux.dev>,
Shakeel Butt <shakeel.butt@linux.dev>,
Kiryl Shutsemau <kas@kernel.org>,
Baoquan He <baoquan.he@linux.dev>, Jason Gunthorpe <jgg@ziepe.ca>,
John Hubbard <jhubbard@nvidia.com>, Peter Xu <peterx@redhat.com>,
tarunsahu@google.com, Fuad Tabba <fuad.tabba@linux.dev>,
Vlastimil Babka <vbabka@kernel.org>
Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org,
linux-kselftest@vger.kernel.org, linux-mm@kvack.org,
linux-coco@lists.linux.dev,
Ackerley Tng <ackerleytng@google.com>,
"Vlastimil Babka (SUSE)" <vbabka@kernel.org>,
Fuad Tabba <fuad.tabba@linux.dev>
Subject: [PATCH v11 18/46] KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check
Date: Wed, 26 Aug 2026 09:18:16 +0000 [thread overview]
Message-ID: <20260826-gmem-inplace-conversion-v11-18-0a15d8a799aa@google.com> (raw)
In-Reply-To: <20260826-gmem-inplace-conversion-v11-0-0a15d8a799aa@google.com>
A guest_memfd folio has no outstanding references if guest_memfd holds the
only references on it. Any other references on the folio may indicate
another user, and guest_memfd cannot convert it to private if there may be
an existing host user.
A folio will have outstanding references if it is present in a per-CPU
lru_add fbatch. guest_memfd does not actually participate in LRU, but
freshly-allocated folios are still added to the lru_add fbatch for batch
LRU statistics processing.
A folio may also have extra refcounts if it is on the mlock fbatch.
These two known "usages" of the folio are handled by calling
lru_cache_drain_for_folio, which drains both the lru_add and mlock
fbatches. After draining, if the refcount is still elevated, then there are
truly outstanding references.
If the page may be dma pinned, DMA is using it and hence there are
outstanding references. folio_maybe_dma_pinned() can have false positives,
but that's only with a significant number of refcounts, at which point
draining LRU is not going to move the needle - it can still be concluded
that the folio has outstanding references.
If the page is still mapped after guest_memfd tried to unmap it earlier in
the conversion process, it also has outstanding references.
Return true and exit early to avoid unnecessary draining in these 2 cases.
Provide a drain status to only drain once ever while processing a batch of
folios.
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Suggested-by: David Hildenbrand <david@kernel.org>
Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Signed-off-by: Ackerley Tng <ackerleytng@google.com>
---
mm/swap.c | 2 ++
virt/kvm/guest_memfd.c | 30 ++++++++++++++++++++++--------
2 files changed, 24 insertions(+), 8 deletions(-)
diff --git a/mm/swap.c b/mm/swap.c
index 8e965c8ce9aa9..9f511b97ab110 100644
--- a/mm/swap.c
+++ b/mm/swap.c
@@ -37,6 +37,7 @@
#include <linux/page_idle.h>
#include <linux/local_lock.h>
#include <linux/buffer_head.h>
+#include <linux/kvm_types.h>
#include "internal.h"
@@ -995,6 +996,7 @@ void lru_cache_drain_for_folio(const struct folio *folio,
*drained = LRU_CACHE_DRAINED_ALL;
}
}
+EXPORT_SYMBOL_FOR_KVM(lru_cache_drain_for_folio);
atomic_t lru_disable_count = ATOMIC_INIT(0);
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index 6dc199be0eb87..4912f90567fe8 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -8,6 +8,7 @@
#include <linux/mempolicy.h>
#include <linux/pseudo_fs.h>
#include <linux/pagemap.h>
+#include <linux/swap.h>
#include "kvm_mm.h"
#include "guest_memfd.h"
@@ -556,10 +557,28 @@ static int kvm_gmem_mas_preallocate(struct ma_state *mas, u64 attributes,
return mas_preallocate(mas, xa_mk_value(attributes), GFP_KERNEL);
}
+static bool __folio_has_outstanding_references(struct folio *folio,
+ enum lru_cache_drained *drained)
+{
+ if (folio_maybe_dma_pinned(folio) || folio_mapped(folio))
+ return true;
+
+ /* 1 reference held by filemap_get_folios() in the folio batch. */
+ lru_cache_drain_for_folio(folio, 1, drained);
+
+ /*
+ * Outstanding references are anything other than those from the page
+ * cache, plus 1 temporary reference held by filemap_get_folios() in the
+ * folio batch.
+ */
+ return folio_ref_count(folio) != folio_nr_pages(folio) + 1;
+}
+
static bool kvm_gmem_has_outstanding_references(struct inode *inode,
pgoff_t start, size_t nr_pages,
pgoff_t *err_index)
{
+ enum lru_cache_drained drained = LRU_CACHE_NOT_DRAINED;
struct address_space *mapping = inode->i_mapping;
pgoff_t last = start + nr_pages - 1;
bool has_outstanding = false;
@@ -570,17 +589,12 @@ static bool kvm_gmem_has_outstanding_references(struct inode *inode,
folio_batch_init(&fbatch);
next = start;
- while (has_outstanding && filemap_get_folios(mapping, &next, last, &fbatch)) {
+ while (!has_outstanding && filemap_get_folios(mapping, &next, last, &fbatch)) {
for (i = 0; i < folio_batch_count(&fbatch); ++i) {
struct folio *folio = fbatch.folios[i];
- /*
- * Outstanding references are anything other than those
- * from the page cache, plus 1 temporary reference held
- * by filemap_get_folios() in the folio batch.
- */
- if (folio_ref_count(folio) != folio_nr_pages(folio) + 1) {
- has_outstanding = true;
+ has_outstanding = __folio_has_outstanding_references(folio, &drained);
+ if (has_outstanding) {
*err_index = max(start, folio->index);
break;
}
--
2.55.0.887.g758fc8c411-goog
next prev parent reply other threads:[~2026-08-26 9:18 UTC|newest]
Thread overview: 53+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-26 9:17 [PATCH v11 00/46] guest_memfd: In-place conversion support Ackerley Tng
2026-08-26 9:17 ` [PATCH v11 01/46] KVM: guest_memfd: Optimize away conversion overheads via dead-code elimination Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 02/46] KVM: guest_memfd: Use kvm_mem_is_private() when populating guest_memfd memory Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 03/46] KVM: guest_memfd: Introduce per-gmem attributes, use to guard user mappings Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 04/46] KVM: Rename KVM_GENERIC_MEMORY_ATTRIBUTES to KVM_VM_MEMORY_ATTRIBUTES Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 05/46] KVM: Enumerate support for PRIVATE memory iff kvm_arch_has_private_mem is defined Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 06/46] KVM: Rename memory attribute APIs to prepare for in-place gmem conversion Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 07/46] KVM: Rename kvm_mem_is_private() to kvm_is_private_gfn() Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 08/46] KVM: Provide generic interface for checking memory private/shared status Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 09/46] KVM: guest_memfd: Stub in ability to enable in-place shared<=>private conversion Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 10/46] KVM: Consolidate private memory and guest_memfd ifdeffery in kvm_host.h Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 11/46] KVM: guest_memfd: Invalidate both SHARED and PRIVATE mappings for in-place conversions Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 12/46] KVM: guest_memfd: Pass mapping type filter to invalidation helper Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 13/46] KVM: guest_memfd: Add base support for KVM_SET_MEMORY_ATTRIBUTES2 Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 14/46] KVM: guest_memfd: Ensure pages are not in use before conversion Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 15/46] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion Ackerley Tng
2026-08-26 19:44 ` Sean Christopherson
2026-08-26 22:17 ` Michael Roth
2026-08-26 22:33 ` Sean Christopherson
2026-08-26 23:47 ` Michael Roth
2026-08-28 4:19 ` Ackerley Tng
2026-08-28 14:57 ` Sean Christopherson
2026-08-26 9:18 ` [PATCH v11 16/46] KVM: guest_memfd: Return early if range already has requested attributes Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 17/46] mm/gup: factor out LRU cache draining for folio into lru_cache_drain_for_folio() Ackerley Tng
2026-08-26 9:18 ` Ackerley Tng [this message]
2026-08-26 9:18 ` [PATCH v11 19/46] KVM: guest_memfd: Zero page while getting pfn Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 20/46] KVM: SEV: Make 'uaddr' parameter optional for KVM_SEV_SNP_LAUNCH_UPDATE Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 21/46] KVM: TDX: Make source page optional for KVM_TDX_INIT_MEM_REGION Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 22/46] KVM: Move KVM_VM_MEMORY_ATTRIBUTES config definition to x86 Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 23/46] KVM: Let userspace disable per-VM mem attributes, enable per-gmem attributes Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 24/46] KVM: guest_memfd: Enable INIT_SHARED on guest_memfd for x86 Coco VMs Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 25/46] KVM: selftests: Create gmem fd before "regular" fd when adding memslot Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 26/46] KVM: selftests: Rename guest_memfd{,_offset} to gmem_{fd,offset} Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 27/46] KVM: selftests: Add support for mmap() on guest_memfd in core library Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 28/46] KVM: selftests: Add selftests global for guest memory attributes capability Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 29/46] KVM: selftests: Add helpers for calling ioctls on guest_memfd Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 30/46] KVM: selftests: Test basic single-page conversion flow Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 31/46] KVM: selftests: Test conversion flow when INIT_SHARED Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 32/46] KVM: selftests: Test conversion precision in guest_memfd Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 33/46] KVM: selftests: Test conversion before allocation Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 34/46] KVM: selftests: Convert with allocated folios in different layouts Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 35/46] KVM: selftests: Test that truncation does not change shared/private status Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 36/46] KVM: selftests: Test that shared/private status is consistent across processes Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 37/46] KVM: selftests: Add helpers to pin pages with CONFIG_GUP_TEST Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 38/46] KVM: selftests: Test conversion with elevated page refcount Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 39/46] KVM: selftests: Reset shared memory after hole-punching Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 40/46] KVM: selftests: Provide function to look up guest_memfd details from gpa Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 41/46] KVM: selftests: Provide common function to set memory attributes Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 42/46] KVM: selftests: Make TEST_EXPECT_SIGBUS thread-safe Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 43/46] KVM: selftests: Support guest_memfd attributes in private_mem_conversions_test Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 44/46] KVM: selftests: Set up page size and alignment independently for guest_memfd Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 45/46] KVM: selftests: Test in-place conversions in private_mem_conversions_test Ackerley Tng
2026-08-26 9:18 ` [PATCH v11 46/46] KVM: selftests: Update private memory exits test to work with per-gmem attributes Ackerley Tng
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260826-gmem-inplace-conversion-v11-18-0a15d8a799aa@google.com \
--to=ackerleytng@google.com \
--cc=aik@amd.com \
--cc=akpm@linux-foundation.org \
--cc=andrew.jones@linux.dev \
--cc=aneesh.kumar@kernel.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=baoquan.he@linux.dev \
--cc=binbin.wu@linux.intel.com \
--cc=bp@alien8.de \
--cc=brauner@kernel.org \
--cc=chao.p.peng@linux.intel.com \
--cc=chrisl@kernel.org \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=david@kernel.org \
--cc=forkloop@google.com \
--cc=fuad.tabba@linux.dev \
--cc=hpa@zytor.com \
--cc=jgg@ziepe.ca \
--cc=jhubbard@nvidia.com \
--cc=jmattson@google.com \
--cc=jthoughton@google.com \
--cc=kas@kernel.org \
--cc=kasong@tencent.com \
--cc=kvm@vger.kernel.org \
--cc=liam@infradead.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=mathieu.desnoyers@efficios.com \
--cc=mhiramat@kernel.org \
--cc=michael.roth@amd.com \
--cc=mingo@redhat.com \
--cc=nphamcs@gmail.com \
--cc=oupton@kernel.org \
--cc=pankaj.gupta@amd.com \
--cc=pbonzini@redhat.com \
--cc=peterx@redhat.com \
--cc=pratyush@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=qperret@google.com \
--cc=rick.p.edgecombe@intel.com \
--cc=rientjes@google.com \
--cc=rostedt@goodmis.org \
--cc=seanjc@google.com \
--cc=shakeel.butt@linux.dev \
--cc=shikemeng@huaweicloud.com \
--cc=shivankg@amd.com \
--cc=shuah@kernel.org \
--cc=skhan@linuxfoundation.org \
--cc=steven.price@arm.com \
--cc=suzuki.poulose@arm.com \
--cc=tarunsahu@google.com \
--cc=tglx@kernel.org \
--cc=vannapurve@google.com \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=willy@infradead.org \
--cc=wyihan@google.com \
--cc=x86@kernel.org \
--cc=yan.y.zhao@intel.com \
--cc=youngjun.park@lge.com \
--cc=yuanchu@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox