From: Hugh Dickins <hughd@google.com>
To: Ackerley Tng <ackerleytng@google.com>,
Andrew Morton <akpm@linux-foundation.org>
Cc: aik@amd.com, andrew.jones@linux.dev, binbin.wu@linux.intel.com,
brauner@kernel.org, chao.p.peng@linux.intel.com,
david@kernel.org, jmattson@google.com, jthoughton@google.com,
michael.roth@amd.com, oupton@kernel.org, pankaj.gupta@amd.com,
qperret@google.com, rick.p.edgecombe@intel.com,
rientjes@google.com, shivankg@amd.com, steven.price@arm.com,
willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com,
forkloop@google.com, pratyush@kernel.org,
suzuki.poulose@arm.com, aneesh.kumar@kernel.org,
liam@infradead.org, Paolo Bonzini <pbonzini@redhat.com>,
Sean Christopherson <seanjc@google.com>,
Thomas Gleixner <tglx@kernel.org>,
Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
x86@kernel.org, "H. Peter Anvin" <hpa@zytor.com>,
Steven Rostedt <rostedt@goodmis.org>,
Masami Hiramatsu <mhiramat@kernel.org>,
Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Shuah Khan <shuah@kernel.org>,
Vishal Annapurve <vannapurve@google.com>,
Andrew Morton <akpm@linux-foundation.org>,
Chris Li <chrisl@kernel.org>, Kairui Song <kasong@tencent.com>,
Kemeng Shi <shikemeng@huaweicloud.com>,
Nhat Pham <nphamcs@gmail.com>, Barry Song <baohua@kernel.org>,
Axel Rasmussen <axelrasmussen@google.com>,
Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
Youngjun Park <youngjun.park@lge.com>,
Qi Zheng <qi.zheng@linux.dev>,
Shakeel Butt <shakeel.butt@linux.dev>,
Kiryl Shutsemau <kas@kernel.org>,
Baoquan He <baoquan.he@linux.dev>,
Jason Gunthorpe <jgg@ziepe.ca>,
John Hubbard <jhubbard@nvidia.com>, Peter Xu <peterx@redhat.com>,
tarunsahu@google.com, Randy Dunlap <rdunlap@infradead.org>,
Lorenzo Stoakes <ljs@kernel.org>,
Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>,
Suren Baghdasaryan <surenb@google.com>,
Michal Hocko <mhocko@suse.com>,
Fuad Tabba <fuad.tabba@linux.dev>,
kvm@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org,
linux-kselftest@vger.kernel.org, linux-mm@kvack.org,
linux-coco@lists.linux.dev
Subject: Re: [PATCH v12 18/45] KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check
Date: Tue, 1 Sep 2026 19:21:05 -0700 (PDT) [thread overview]
Message-ID: <bd6c9c74-e374-a9d3-ba1f-8b6f430894fc@google.com> (raw)
In-Reply-To: <20260830-gmem-inplace-conversion-v12-18-85e5fd25252a@google.com>
On Sun, 30 Aug 2026, Ackerley Tng via B4 Relay wrote:
> From: Ackerley Tng <ackerleytng@google.com>
>
> A guest_memfd folio has no outstanding references if guest_memfd holds the
> only references on it. Any other references on the folio may indicate
> another user, and guest_memfd cannot convert it to private if there may be
> an existing host user.
>
> A folio will have outstanding references if it is present in a per-CPU
> lru_add fbatch. guest_memfd does not actually participate in LRU, but
> freshly-allocated folios are still added to the lru_add fbatch for batch
> LRU statistics processing.
>
> A folio may also have extra refcounts if it is on the mlock fbatch.
>
> These two known "usages" of the folio are handled by calling
> lru_cache_drain_for_folio, which drains both the lru_add and mlock
> fbatches. After draining, if the refcount is still elevated, then there are
> truly outstanding references.
>
> If the page may be dma pinned, DMA is using it and hence there are
> outstanding references. folio_maybe_dma_pinned() can have false positives,
> but that's only with a significant number of refcounts, at which point
> draining LRU is not going to move the needle - it can still be concluded
> that the folio has outstanding references.
>
> If the page is still mapped after guest_memfd tried to unmap it earlier in
> the conversion process, it also has outstanding references.
>
> Return true and exit early to avoid unnecessary draining in these 2 cases.
>
> Provide a drain status to only drain once ever while processing a batch of
> folios.
>
> Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
> Suggested-by: David Hildenbrand <david@kernel.org>
> Reviewed-by: Fuad Tabba <fuad.tabba@linux.dev>
> Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
> Signed-off-by: Ackerley Tng <ackerleytng@google.com>
> ---
> mm/folio.c | 2 ++
> virt/kvm/guest_memfd.c | 30 ++++++++++++++++++++++--------
> 2 files changed, 24 insertions(+), 8 deletions(-)
>
> diff --git a/mm/folio.c b/mm/folio.c
> index c02dcea9c03c2..50a6dbe55998e 100644
> --- a/mm/folio.c
> +++ b/mm/folio.c
> @@ -33,6 +33,7 @@
> #include <linux/page_idle.h>
> #include <linux/local_lock.h>
> #include <linux/buffer_head.h>
> +#include <linux/kvm_types.h>
>
> #include "internal.h"
> #include "page_alloc.h"
> @@ -926,6 +927,7 @@ void lru_cache_drain_for_folio(const struct folio *folio,
> *drained = LRU_CACHE_DRAINED_ALL;
> }
> }
> +EXPORT_SYMBOL_FOR_KVM(lru_cache_drain_for_folio);
>
> atomic_t lru_disable_count = ATOMIC_INIT(0);
>
I don't mind about the virt/kvm/guest_memfd.c part of it, but I'm finding
a KVM patchset modifying mm/folio.c there hard to deal with: and notice
Sean also suggesting to separate this part out.
As you know, I've worked up a patchset "mm/fbatch: drain lru_add_drain()
and _all()" which finally removes the problem lru_cache_drain_for_folio()
works around. In the initial version posted a week ago, there was no
lru_cache_drain_for_folio() in the tree. Now 7.3-rc1 has it, so I
intended a replacement 13/25 in my series, giving you just an empty
inline lru_cache_drain_for_folio() stub (and enum lru_cache_drained)
in linux/swap.h.
But that won't work for you, if you're adding an EXPORT_SYMBOL_FOR_KVM()
in mm/folio.c, and of course conflicts with my removals (in context both
above and below your EXPORT line). It's easy for me to remove what's in
mm/gup.c and mm/folio.c, but I cannot remove what is not yet there.
I've wasted hours on this, hoping not to trouble either of you; but
seeing now that I shall have to rebase anyway (an unrelated mlock fix),
I'm electing to take the only clean way out: I'm going to submit this
mm/folio.c part of your patch to Andrew tonight (with a shorter Cc list!),
in the hope that it can be accelerated into 7.3-rc2 (or at least get an
mm-stable stable base-commit id) which we can both work off independently.
Whether that's acceptable to Ackerley and to Andrew, I don't know
(just as we don't know when either of our patchsets will go further),
but let me try.
Thanks,
Hugh
next prev parent reply other threads:[~2026-09-02 2:21 UTC|newest]
Thread overview: 71+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-31 0:25 [PATCH v12 00/45] guest_memfd: In-place conversion support Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 01/45] KVM: guest_memfd: Optimize away conversion overheads via dead-code elimination Ackerley Tng via B4 Relay
2026-09-01 8:39 ` Fuad Tabba
2026-09-01 9:13 ` Binbin Wu
2026-08-31 0:25 ` [PATCH v12 02/45] KVM: guest_memfd: Use kvm_mem_is_private() when populating guest_memfd memory Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 03/45] KVM: guest_memfd: Introduce per-gmem attributes, use to guard user mappings Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 04/45] KVM: Rename KVM_GENERIC_MEMORY_ATTRIBUTES to KVM_VM_MEMORY_ATTRIBUTES Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 05/45] KVM: Enumerate support for PRIVATE memory iff kvm_arch_has_private_mem is defined Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 06/45] KVM: Rename memory attribute APIs to prepare for in-place gmem conversion Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 07/45] KVM: Rename kvm_mem_is_private() to kvm_is_private_gfn() Ackerley Tng via B4 Relay
2026-09-01 8:42 ` Fuad Tabba
2026-09-01 9:18 ` Binbin Wu
2026-08-31 0:25 ` [PATCH v12 08/45] KVM: Provide generic interface for checking memory private/shared status Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 09/45] KVM: guest_memfd: Stub in ability to enable in-place shared<=>private conversion Ackerley Tng via B4 Relay
2026-09-01 9:10 ` Fuad Tabba
2026-09-01 9:47 ` Binbin Wu
2026-08-31 0:25 ` [PATCH v12 10/45] KVM: Consolidate private memory and guest_memfd ifdeffery in kvm_host.h Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 11/45] KVM: guest_memfd: Invalidate both SHARED and PRIVATE mappings for in-place conversions Ackerley Tng via B4 Relay
2026-09-01 9:44 ` Fuad Tabba
2026-09-02 3:27 ` Binbin Wu
2026-08-31 0:25 ` [PATCH v12 12/45] KVM: guest_memfd: Always fault from guest_memfd if in-place conversion is enabled Ackerley Tng via B4 Relay
2026-09-01 10:13 ` Fuad Tabba
2026-09-02 5:01 ` Yan Zhao
2026-09-02 6:00 ` Binbin Wu
2026-08-31 0:25 ` [PATCH v12 13/45] KVM: guest_memfd: Pass mapping type filter to invalidation helper Ackerley Tng via B4 Relay
2026-09-01 10:33 ` Fuad Tabba
2026-09-02 6:02 ` Binbin Wu
2026-08-31 0:25 ` [PATCH v12 14/45] KVM: guest_memfd: Add base support for KVM_SET_MEMORY_ATTRIBUTES2 Ackerley Tng via B4 Relay
2026-09-01 8:00 ` Fuad Tabba
2026-08-31 0:25 ` [PATCH v12 15/45] KVM: guest_memfd: Ensure pages are not in use before conversion Ackerley Tng via B4 Relay
2026-09-01 7:45 ` Fuad Tabba
2026-09-01 18:35 ` Sean Christopherson
2026-09-01 18:41 ` Sean Christopherson
2026-08-31 0:25 ` [PATCH v12 16/45] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 17/45] KVM: guest_memfd: Return early if range already has requested attributes Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 18/45] KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check Ackerley Tng via B4 Relay
2026-09-02 2:21 ` Hugh Dickins [this message]
2026-08-31 0:25 ` [PATCH v12 19/45] KVM: guest_memfd: Zero page while getting pfn Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 20/45] KVM: SEV: Make 'uaddr' parameter optional for KVM_SEV_SNP_LAUNCH_UPDATE Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 21/45] KVM: TDX: Make source page optional for KVM_TDX_INIT_MEM_REGION Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 22/45] KVM: Move KVM_VM_MEMORY_ATTRIBUTES config definition to x86 Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 23/45] KVM: Let userspace disable per-VM mem attributes, enable per-gmem attributes Ackerley Tng via B4 Relay
2026-09-01 8:19 ` Fuad Tabba
2026-09-01 18:26 ` Sean Christopherson
2026-09-01 22:22 ` Fuad Tabba
2026-09-02 5:56 ` Binbin Wu
2026-09-02 13:30 ` Sean Christopherson
2026-08-31 0:25 ` [PATCH v12 24/45] KVM: guest_memfd: Enable INIT_SHARED on guest_memfd for x86 Coco VMs Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 25/45] KVM: selftests: Create gmem fd before "regular" fd when adding memslot Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 26/45] KVM: selftests: Rename guest_memfd{,_offset} to gmem_{fd,offset} Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 27/45] KVM: selftests: Add support for mmap() on guest_memfd in core library Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 28/45] KVM: selftests: Add selftests global for guest memory attributes capability Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 29/45] KVM: selftests: Add helpers for calling ioctls on guest_memfd Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 30/45] KVM: selftests: Test basic single-page conversion flow Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 31/45] KVM: selftests: Test conversion flow when INIT_SHARED Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 32/45] KVM: selftests: Test conversion precision in guest_memfd Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 33/45] KVM: selftests: Test conversion before allocation Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 34/45] KVM: selftests: Convert with allocated folios in different layouts Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 35/45] KVM: selftests: Test that truncation does not change shared/private status Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 36/45] KVM: selftests: Test that shared/private status is consistent across processes Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 37/45] KVM: selftests: Add helpers to pin pages with CONFIG_GUP_TEST Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 38/45] KVM: selftests: Test conversion with elevated page refcount Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 39/45] KVM: selftests: Reset shared memory after hole-punching Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 40/45] KVM: selftests: Provide function to look up guest_memfd details from gpa Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 41/45] KVM: selftests: Provide common function to set memory attributes Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 42/45] KVM: selftests: Make TEST_EXPECT_SIGBUS thread-safe Ackerley Tng via B4 Relay
2026-08-31 0:25 ` [PATCH v12 43/45] KVM: selftests: Set up page size and alignment independently for guest_memfd Ackerley Tng via B4 Relay
2026-09-01 11:03 ` Fuad Tabba
2026-08-31 0:25 ` [PATCH v12 44/45] KVM: selftests: Update private_mem_conversions_test for in-place conversions Ackerley Tng via B4 Relay
2026-09-01 11:11 ` Fuad Tabba
2026-08-31 0:25 ` [PATCH v12 45/45] KVM: selftests: Update private memory exits test to work with per-gmem attributes Ackerley Tng via B4 Relay
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=bd6c9c74-e374-a9d3-ba1f-8b6f430894fc@google.com \
--to=hughd@google.com \
--cc=ackerleytng@google.com \
--cc=aik@amd.com \
--cc=akpm@linux-foundation.org \
--cc=andrew.jones@linux.dev \
--cc=aneesh.kumar@kernel.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=baoquan.he@linux.dev \
--cc=binbin.wu@linux.intel.com \
--cc=bp@alien8.de \
--cc=brauner@kernel.org \
--cc=chao.p.peng@linux.intel.com \
--cc=chrisl@kernel.org \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=david@kernel.org \
--cc=forkloop@google.com \
--cc=fuad.tabba@linux.dev \
--cc=hpa@zytor.com \
--cc=jgg@ziepe.ca \
--cc=jhubbard@nvidia.com \
--cc=jmattson@google.com \
--cc=jthoughton@google.com \
--cc=kas@kernel.org \
--cc=kasong@tencent.com \
--cc=kvm@vger.kernel.org \
--cc=liam@infradead.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=ljs@kernel.org \
--cc=mathieu.desnoyers@efficios.com \
--cc=mhiramat@kernel.org \
--cc=mhocko@suse.com \
--cc=michael.roth@amd.com \
--cc=mingo@redhat.com \
--cc=nphamcs@gmail.com \
--cc=oupton@kernel.org \
--cc=pankaj.gupta@amd.com \
--cc=pbonzini@redhat.com \
--cc=peterx@redhat.com \
--cc=pratyush@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=qperret@google.com \
--cc=rdunlap@infradead.org \
--cc=rick.p.edgecombe@intel.com \
--cc=rientjes@google.com \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=seanjc@google.com \
--cc=shakeel.butt@linux.dev \
--cc=shikemeng@huaweicloud.com \
--cc=shivankg@amd.com \
--cc=shuah@kernel.org \
--cc=skhan@linuxfoundation.org \
--cc=steven.price@arm.com \
--cc=surenb@google.com \
--cc=suzuki.poulose@arm.com \
--cc=tarunsahu@google.com \
--cc=tglx@kernel.org \
--cc=vannapurve@google.com \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=willy@infradead.org \
--cc=wyihan@google.com \
--cc=x86@kernel.org \
--cc=yan.y.zhao@intel.com \
--cc=youngjun.park@lge.com \
--cc=yuanchu@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox