From: "David Hildenbrand (Arm)" <david@kernel.org>
To: ackerleytng@google.com, aik@amd.com, andrew.jones@linux.dev,
binbin.wu@linux.intel.com, brauner@kernel.org,
chao.p.peng@linux.intel.com, jmattson@google.com,
jthoughton@google.com, michael.roth@amd.com, oupton@kernel.org,
pankaj.gupta@amd.com, qperret@google.com,
rick.p.edgecombe@intel.com, rientjes@google.com,
shivankg@amd.com, steven.price@arm.com, tabba@google.com,
willy@infradead.org, wyihan@google.com, yan.y.zhao@intel.com,
forkloop@google.com, pratyush@kernel.org, suzuki.poulose@arm.com,
aneesh.kumar@kernel.org, liam@infradead.org,
Paolo Bonzini <pbonzini@redhat.com>,
Sean Christopherson <seanjc@google.com>,
Thomas Gleixner <tglx@kernel.org>, Ingo Molnar <mingo@redhat.com>,
Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
x86@kernel.org, "H. Peter Anvin" <hpa@zytor.com>,
Steven Rostedt <rostedt@goodmis.org>,
Masami Hiramatsu <mhiramat@kernel.org>,
Mathieu Desnoyers <mathieu.desnoyers@efficios.com>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Shuah Khan <shuah@kernel.org>,
Vishal Annapurve <vannapurve@google.com>,
Andrew Morton <akpm@linux-foundation.org>,
Chris Li <chrisl@kernel.org>, Kairui Song <kasong@tencent.com>,
Kemeng Shi <shikemeng@huaweicloud.com>,
Nhat Pham <nphamcs@gmail.com>, Barry Song <baohua@kernel.org>,
Axel Rasmussen <axelrasmussen@google.com>,
Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
Youngjun Park <youngjun.park@lge.com>,
Qi Zheng <qi.zheng@linux.dev>,
Shakeel Butt <shakeel.butt@linux.dev>,
Kiryl Shutsemau <kas@kernel.org>,
Baoquan He <baoquan.he@linux.dev>, Jason Gunthorpe <jgg@ziepe.ca>,
John Hubbard <jhubbard@nvidia.com>, Peter Xu <peterx@redhat.com>,
Vlastimil Babka <vbabka@kernel.org>,
"hughd@google.com" <hughd@google.com>
Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org,
linux-trace-kernel@vger.kernel.org, linux-doc@vger.kernel.org,
linux-kselftest@vger.kernel.org, linux-mm@kvack.org,
linux-coco@lists.linux.dev
Subject: Re: [PATCH v9 14/41] mm: swap: Introduce lru_add_drain_progressive()
Date: Thu, 30 Jul 2026 12:21:36 +0200 [thread overview]
Message-ID: <4460e252-b6de-4df3-bea8-0368a9d0f4ff@kernel.org> (raw)
In-Reply-To: <20260728-gmem-inplace-conversion-v9-14-35f9aec2aed2@google.com>
On 7/29/26 02:35, Ackerley Tng via B4 Relay wrote:
> From: Ackerley Tng <ackerleytng@google.com>
>
> Extract the progressive LRU drain retry logic from
> collect_longterm_unpinnable_folios() into a reusable helper,
> lru_add_drain_progressive().
>
> When attempting to isolate folios that may still reside in per-CPU folio
> batches, draining is escalated progressively:
>
> 1. State 0: Call lru_add_drain() to flush local CPU batches.
> 2. State 1: Call lru_add_drain_all() to flush all CPU batches.
> 3. State >= 2: Return false to stop retrying.
>
> Refactor collect_longterm_unpinnable_folios() to use this new helper.
>
> The helper will be used by KVM's guest_memfd in a later patch.
>
> Signed-off-by: Ackerley Tng <ackerleytng@google.com>
> ---
> include/linux/swap.h | 2 ++
> mm/gup.c | 19 ++++++-------------
> mm/swap.c | 15 +++++++++++++++
> 3 files changed, 23 insertions(+), 13 deletions(-)
>
> diff --git a/include/linux/swap.h b/include/linux/swap.h
> index 8f0f68e245baa..cd54f73f34f39 100644
> --- a/include/linux/swap.h
> +++ b/include/linux/swap.h
> @@ -344,6 +344,8 @@ extern void lru_add_drain(void);
> extern void lru_add_drain_cpu(int cpu);
> extern void lru_add_drain_cpu_zone(struct zone *zone);
> extern void lru_add_drain_all(void);
> +bool lru_add_drain_progressive(int *drain_state);
> +
> void folio_deactivate(struct folio *folio);
> void folio_mark_lazyfree(struct folio *folio);
> extern void swap_setup(void);
> diff --git a/mm/gup.c b/mm/gup.c
> index 0692119b79043..5f00435e2c635 100644
> --- a/mm/gup.c
> +++ b/mm/gup.c
> @@ -2268,7 +2268,7 @@ static unsigned long collect_longterm_unpinnable_folios(
> {
> unsigned long collected = 0;
> struct folio *folio;
> - int drained = 0;
> + int drain_state = 0;
> long i = 0;
>
> for (folio = pofs_get_folio(pofs, i); folio;
> @@ -2287,18 +2287,11 @@ static unsigned long collect_longterm_unpinnable_folios(
> continue;
> }
>
> - if (drained == 0 && folio_may_be_lru_cached(folio) &&
> - folio_ref_count(folio) !=
> - folio_expected_ref_count(folio) + 1) {
> - lru_add_drain();
> - drained = 1;
> - }
> - if (drained == 1 && folio_may_be_lru_cached(folio) &&
> - folio_ref_count(folio) !=
> - folio_expected_ref_count(folio) + 1) {
> - lru_add_drain_all();
> - drained = 2;
> - }
> + while (folio_may_be_lru_cached(folio) &&
> + folio_ref_count(folio) !=
> + folio_expected_ref_count(folio) + 1 &&
> + lru_add_drain_progressive(&drain_state))
> + ;
That's rather nasty.
I was hoping that we could embed more logic in a helper. The history [1] of the
refcount check is rather sad:
https://lore.kernel.org/all/c5bac539-fd8a-4db7-c21c-cd3e457eee91@google.com/
... primarily because of mlock() handling.
For guest_memfd(), would mlock() ever apply on a path where you need that check?
Conceptually, I wonder whether we can do the following, and rely on the refcount
check only on the mlock path.
From ae30f7594692b0c47b44922c26b6480dd9584872 Mon Sep 17 00:00:00 2001
From: "David Hildenbrand (Arm)" <david@kernel.org>
Date: Thu, 30 Jul 2026 12:20:14 +0200
Subject: [PATCH] tmp
Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>
---
mm/gup.c | 66 ++++++++++++++++++++++++++++++++++++++++++++------------
1 file changed, 52 insertions(+), 14 deletions(-)
diff --git a/mm/gup.c b/mm/gup.c
index 99902c15703b0..ec43071bd035a 100644
--- a/mm/gup.c
+++ b/mm/gup.c
@@ -2259,6 +2259,56 @@ static struct folio *pofs_next_folio(struct folio *folio,
return pofs_get_folio(pofs, i);
}
+enum lru_cache_drained {
+ LRU_CACHE_NOT_DRAINED,
+ LRU_CACHE_DRAINED,
+ LRU_CACHE_DRAINED_ALL,
+};
+
+static inline void folio_likely_lru_cached(const struct folio *folio)
+{
+ if (!folio_may_be_lru_cached(folio))
+ return false;
+ /*
+ * Having the LRU flag clear either indicates LRU cache references
+ * or LRU isolation.
+ */
+ if (!folio_test_lru(folio))
+ return true;
+ /*
+ * For mlocked folios, we don't have a real indication: we can only
+ * take a guess based on the refcount.
+ */
+ if (!folio_test_mlock(folio))
+ return false;
+ return folio_expected_ref_count(folio) == folio_ref_count(folio) = 1;
+}
+
+/**
+ * lru_cache_drain_for_folio() - progressively try draining the lru cache
+ * @folio: The folio.
+ * @drained: Status initialized to LRU_CACHE_NOT_DRAINED by the caller.
+ *
+ * TODO
+ */
+static void lru_cache_drain_for_folio(const struct folio *folio,
+ enum lru_cache_drained *drained)
+{
+ if (!folio_likely_lru_cached(folio))
+ return false;
+
+ /* Try local draining first, if not already done previously. */
+ if (*drained == LRU_CACHE_NOT_DRAINED) {
+ lru_add_drain();
+ *drained = LRU_CACHE_DRAINED;
+ }
+ /* Try draining all CPUs next if still not an LRU folio. */
+ if (folio_likely_lru_cached(folio) && *drained == LRU_CACHE_DRAINED) {
+ lru_add_drain_all();
+ *drained = LRU_CACHE_DRAINED_ALL;
+ }
+}
+
/*
* Returns the number of collected folios. Return value is always >= 0.
*/
@@ -2266,9 +2316,9 @@ static unsigned long collect_longterm_unpinnable_folios(
struct list_head *movable_folio_list,
struct pages_or_folios *pofs)
{
+ enum lru_cache_drained drained = LRU_CACHE_NOT_DRAINED;
unsigned long collected = 0;
struct folio *folio;
- int drained = 0;
long i = 0;
for (folio = pofs_get_folio(pofs, i); folio;
@@ -2287,19 +2337,7 @@ static unsigned long collect_longterm_unpinnable_folios(
continue;
}
- if (drained == 0 && folio_may_be_lru_cached(folio) &&
- folio_ref_count(folio) !=
- folio_expected_ref_count(folio) + 1) {
- lru_add_drain();
- drained = 1;
- }
- if (drained == 1 && folio_may_be_lru_cached(folio) &&
- folio_ref_count(folio) !=
- folio_expected_ref_count(folio) + 1) {
- lru_add_drain_all();
- drained = 2;
- }
-
+ lru_cache_drain_for_folio(folio, &drained);
if (!folio_isolate_lru(folio))
continue;
--
2.43.0
--
Cheers,
David
next prev parent reply other threads:[~2026-07-30 10:21 UTC|newest]
Thread overview: 47+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-29 0:34 [PATCH v9 00/41] guest_memfd: In-place conversion support Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 01/41] KVM: guest_memfd: Introduce per-gmem attributes, use to guard user mappings Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 02/41] KVM: Rename KVM_GENERIC_MEMORY_ATTRIBUTES to KVM_VM_MEMORY_ATTRIBUTES Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 03/41] KVM: Enumerate support for PRIVATE memory iff kvm_arch_has_private_mem is defined Ackerley Tng via B4 Relay
2026-07-30 9:57 ` Xiaoyao Li
2026-07-29 0:35 ` [PATCH v9 04/41] KVM: Rename memory attribute APIs to prepare for in-place gmem conversion Ackerley Tng via B4 Relay
2026-07-30 8:19 ` Xiaoyao Li
2026-07-29 0:35 ` [PATCH v9 05/41] KVM: Provide generic interface for checking memory private/shared status Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 06/41] KVM: guest_memfd: Introduce function to check GFN " Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 07/41] KVM: guest_memfd: Wire up core private/shared attribute interfaces Ackerley Tng via B4 Relay
2026-07-30 11:05 ` Xiaoyao Li
2026-07-29 0:35 ` [PATCH v9 08/41] KVM: Consolidate private memory and guest_memfd ifdeffery in kvm_host.h Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 09/41] KVM: guest_memfd: Filter both shared and private when invalidating Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 10/41] KVM: guest_memfd: Add base support for KVM_SET_MEMORY_ATTRIBUTES2 Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 11/41] KVM: guest_memfd: Ensure pages are not in use before conversion Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 12/41] KVM: guest_memfd: Call arch make_shared callback for to-shared conversion Ackerley Tng via B4 Relay
2026-07-30 10:35 ` Fuad Tabba
2026-07-29 0:35 ` [PATCH v9 13/41] KVM: guest_memfd: Return early if range already has requested attributes Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 14/41] mm: swap: Introduce lru_add_drain_progressive() Ackerley Tng via B4 Relay
2026-07-30 10:21 ` David Hildenbrand (Arm) [this message]
2026-07-29 0:35 ` [PATCH v9 15/41] KVM: guest_memfd: Handle lru_add fbatch refcounts during conversion safety check Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 16/41] KVM: guest_memfd: Zero page while getting pfn Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 17/41] KVM: SEV: Make 'uaddr' parameter optional for KVM_SEV_SNP_LAUNCH_UPDATE Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 18/41] KVM: TDX: Make source page optional for KVM_TDX_INIT_MEM_REGION Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 19/41] KVM: Move KVM_VM_MEMORY_ATTRIBUTES config definition to x86 Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 20/41] KVM: Let userspace disable per-VM mem attributes, enable per-gmem attributes Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 21/41] KVM: guest_memfd: Enable INIT_SHARED on guest_memfd for x86 Coco VMs Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 22/41] KVM: selftests: Create gmem fd before "regular" fd when adding memslot Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 23/41] KVM: selftests: Rename guest_memfd{,_offset} to gmem_{fd,offset} Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 24/41] KVM: selftests: Add support for mmap() on guest_memfd in core library Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 25/41] KVM: selftests: Add selftests global for guest memory attributes capability Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 26/41] KVM: selftests: Add helpers for calling ioctls on guest_memfd Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 27/41] KVM: selftests: Test basic single-page conversion flow Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 28/41] KVM: selftests: Test conversion flow when INIT_SHARED Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 29/41] KVM: selftests: Test conversion precision in guest_memfd Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 30/41] KVM: selftests: Test conversion before allocation Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 31/41] KVM: selftests: Convert with allocated folios in different layouts Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 32/41] KVM: selftests: Test that truncation does not change shared/private status Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 33/41] KVM: selftests: Test that shared/private status is consistent across processes Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 34/41] KVM: selftests: Add helpers to pin pages with CONFIG_GUP_TEST Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 35/41] KVM: selftests: Test conversion with elevated page refcount Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 36/41] KVM: selftests: Reset shared memory after hole-punching Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 37/41] KVM: selftests: Provide function to look up guest_memfd details from gpa Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 38/41] KVM: selftests: Provide common function to set memory attributes Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 39/41] KVM: selftests: Make TEST_EXPECT_SIGBUS thread-safe Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 40/41] KVM: selftests: Update private_mem_conversions_test to mmap() guest_memfd Ackerley Tng via B4 Relay
2026-07-29 0:35 ` [PATCH v9 41/41] KVM: selftests: Update private memory exits test to work with per-gmem attributes Ackerley Tng via B4 Relay
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=4460e252-b6de-4df3-bea8-0368a9d0f4ff@kernel.org \
--to=david@kernel.org \
--cc=ackerleytng@google.com \
--cc=aik@amd.com \
--cc=akpm@linux-foundation.org \
--cc=andrew.jones@linux.dev \
--cc=aneesh.kumar@kernel.org \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=baoquan.he@linux.dev \
--cc=binbin.wu@linux.intel.com \
--cc=bp@alien8.de \
--cc=brauner@kernel.org \
--cc=chao.p.peng@linux.intel.com \
--cc=chrisl@kernel.org \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=forkloop@google.com \
--cc=hpa@zytor.com \
--cc=hughd@google.com \
--cc=jgg@ziepe.ca \
--cc=jhubbard@nvidia.com \
--cc=jmattson@google.com \
--cc=jthoughton@google.com \
--cc=kas@kernel.org \
--cc=kasong@tencent.com \
--cc=kvm@vger.kernel.org \
--cc=liam@infradead.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=mathieu.desnoyers@efficios.com \
--cc=mhiramat@kernel.org \
--cc=michael.roth@amd.com \
--cc=mingo@redhat.com \
--cc=nphamcs@gmail.com \
--cc=oupton@kernel.org \
--cc=pankaj.gupta@amd.com \
--cc=pbonzini@redhat.com \
--cc=peterx@redhat.com \
--cc=pratyush@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=qperret@google.com \
--cc=rick.p.edgecombe@intel.com \
--cc=rientjes@google.com \
--cc=rostedt@goodmis.org \
--cc=seanjc@google.com \
--cc=shakeel.butt@linux.dev \
--cc=shikemeng@huaweicloud.com \
--cc=shivankg@amd.com \
--cc=shuah@kernel.org \
--cc=skhan@linuxfoundation.org \
--cc=steven.price@arm.com \
--cc=suzuki.poulose@arm.com \
--cc=tabba@google.com \
--cc=tglx@kernel.org \
--cc=vannapurve@google.com \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=willy@infradead.org \
--cc=wyihan@google.com \
--cc=x86@kernel.org \
--cc=yan.y.zhao@intel.com \
--cc=youngjun.park@lge.com \
--cc=yuanchu@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox