From: Lance Yang <lance.yang@linux.dev>
To: kirill@shutemov.name
Cc: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org,
nico.pache@linux.dev, baolin.wang@linux.alibaba.com,
baohua@kernel.org, dev.jain@arm.com, hughd@google.com,
lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com,
rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org,
surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org,
ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com,
linux-mm@kvack.org, linux-kselftest@vger.kernel.org,
linux-kernel@vger.kernel.org, kas@kernel.org, jannh@google.com,
willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org,
mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org,
bpf@vger.kernel.org
Subject: Re: [RFC PATCH 16/57] mm/collapse: freeze the sources behind migration entries
Date: Mon, 24 Aug 2026 21:12:24 +0800 [thread overview]
Message-ID: <20260824131224.73344-1-lance.yang@linux.dev> (raw)
In-Reply-To: <20260816224609.308019-17-kirill@shutemov.name>
On Sun, Aug 16, 2026 at 11:45:28PM +0100, Kiryl Shutsemau wrote:
>From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
[...]
>+static enum scan_result collapse_freeze_candidate(struct mm_struct *mm,
>+ struct collapse_candidate *cand, pte_t *pte)
>+{
>+ const unsigned int nr_pages = candidate_nr_pages(cand);
>+ unsigned int nr_saved = 0, nr_frozen = 0;
>+ enum scan_result result;
>+ struct folio *folio;
>+ unsigned long addr;
>+ unsigned int i;
>+
>+ for (i = 0, addr = cand->addr; i < nr_pages;) {
>+ pte_t ptent = ptep_get(pte + i);
>+ unsigned int nr, nr_max, k;
>+ pte_t rep;
>+
>+ if (pte_none(ptent)) {
>+ /* Hole: nothing to freeze; install verifies it stayed one */
>+ cand->saved_ptes[i] = ptent;
>+ nr_saved = ++i;
>+ addr += PAGE_SIZE;
>+ continue;
>+ }
>+ if (is_zero_pfn(pte_pfn(ptent))) {
>+ /*
>+ * Clear the zeropage mapping now, covered by the round's
>+ * ranged flush: overwriting a live PTE at install would
>+ * be a valid->valid transition, breaking arm64's
>+ * break-before-make. The zeropage has neither rmap nor
>+ * per-map references -- the saved value alone undoes it.
>+ */
>+ cand->saved_ptes[i] =
>+ ptep_get_and_clear(mm, addr, pte + i);
>+ nr_saved = ++i;
>+ addr += PAGE_SIZE;
>+ continue;
>+ }
>+
>+ folio = pte_folio(ptent);
>+
>+ /*
>+ * A folio revisited by a second span of this round is already
>+ * ours and frozen at its first span: folio_get() on a zero count
>+ * is a bug, and try-get fails cleanly. Scrambled layouts
>+ * (mremap) construct this; nothing else can hold a folio frozen
>+ * while its PTE is live under our ptl, so it is not transient.
>+ */
>+ if (!folio_try_get(folio)) {
>+ result = SCAN_PAGE_COUNT;
>+ goto unfreeze;
>+ }
>+ if (!folio_trylock(folio)) {
>+ folio_put(folio);
>+ result = SCAN_PAGE_LOCK;
>+ goto unfreeze;
>+ }
>+
>+ /*
>+ * Never freeze a folio under writeback. PG_writeback holds no
>+ * reference of its own -- the swapcache reference keeps the folio
>+ * alive, and everything that would drop it waits for the flag --
>+ * so folio_end_writeback() plain folio_get()s a folio it may
>+ * assume is alive: a BUG on a frozen one, or with
>+ * CONFIG_DEBUG_VM off, a free under our copy.
>+ *
>+ * Unlike every other hazard here, the freeze does not catch it.
>+ * folio_expected_ref_count() counts the swapcache reference, so
>+ * the count is exactly right and the freeze succeeds. Nor can
>+ * "is it in the swapcache" stand in for this test: that would
>+ * refuse the pages the fault-in pass just swapped in.
>+ *
>+ * Reachable even though writeback starts on an unmapped folio: a
>+ * re-fault from the swapcache maps it back before the bio
>+ * completes, and folio_free_swap() will not drop the cache entry
>+ * under writeback. Testing once is enough -- writeback starts
>+ * only under the folio lock, which we hold from here through
>+ * putback.
>+ */
>+ if (folio_test_writeback(folio)) {
>+ folio_unlock(folio);
>+ folio_put(folio);
>+ result = SCAN_PAGE_DIRTY_OR_WRITEBACK;
>+ goto unfreeze;
>+ }
>+
>+ /* Each slot's own value: a span agrees on the PFN, not the rest */
>+ cand->saved_ptes[i] = ptent;
>+ nr_max = collapse_span_max(ptent, nr_pages - i);
>+ for (nr = 1; nr < nr_max; nr++) {
>+ pte_t tail = ptep_get(pte + i + nr);
>+
>+ if (!pte_present(tail) ||
>+ pte_pfn(tail) != pte_pfn(ptent) + nr)
>+ break;
>+ cand->saved_ptes[i + nr] = tail;
>+ }
>+
>+ /*
>+ * The clear is the GUP-fast linearization point: a grab landing
>+ * before it elevates the refcount and the freeze below fails
>+ * (the candidate unfreezes); one landing after fails its PTE
>+ * re-read and retries. Clear and store sit adjacent under one
>+ * uninterrupted ptl hold, batched per span
>+ * (get_and_clear_full_ptes() unfolds contpte), so the transient
>+ * none window is invisible to installers, which all take the ptl.
>+ */
>+ rep = get_and_clear_full_ptes(mm, addr, pte + i, nr, 0);
>+
>+ /*
>+ * Dirty from the clear -- including any the hardware set since
>+ * the reads above -- goes to the folio, the way unmap does,
>+ * rather than onto PTEs that never had it. Young needs no such
>+ * care: a migration entry drops it either way.
>+ */
>+ if (pte_dirty(rep))
>+ folio_mark_dirty(folio);
>+
>+ for (k = 0; k < nr; k++) {
>+ pte_t saved = cand->saved_ptes[i + k];
>+ swp_entry_t entry;
>+ pte_t swp_pte;
>+
>+ entry = make_readable_migration_entry(pte_pfn(saved));
>+ swp_pte = swp_entry_to_pte(entry);
>+ if (pte_soft_dirty(saved))
>+ swp_pte = pte_swp_mksoft_dirty(swp_pte);
>+ set_pte_at(mm, addr + k * PAGE_SIZE, pte + i + k,
>+ swp_pte);
>+ }
>+ nr_saved = i + nr;
>+
>+ if (!folio_ref_freeze(folio,
>+ folio_expected_ref_count(folio) + 1)) {
>+ result = SCAN_PAGE_COUNT;
>+ goto unfreeze;
>+ }
>+ nr_frozen = nr_saved;
Just one thing I was wondering about ... can deferred_split_isolate()
remove a source folio from deferred_split_lru while its refcount is
frozen by collapse_freeze_candidate()?
Assume an earlier span belongs to an anonymous large folio on the
deferred split queue, then a later span fails folio_trylock().
collapse_freeze_candidate() continues after freezing each source folio
and calls collapse_unfreeze_candidate() on a later failure:
static noinline enum scan_result collapse_freeze_candidate(struct mm_struct *mm,
struct collapse_candidate *cand, pte_t *pte)
{
...
for (i = 0, addr = cand->addr; i < nr_pages;) {
...
if (!folio_trylock(folio)) {
folio_put(folio);
result = SCAN_PAGE_LOCK;
goto unfreeze;
}
...
nr_saved = i + nr;
if (!folio_ref_freeze(folio,
folio_expected_ref_count(folio) + 1)) {
result = SCAN_PAGE_COUNT;
goto unfreeze;
}
nr_frozen = nr_saved;
i += nr;
addr += nr * PAGE_SIZE;
}
...
unfreeze:
collapse_unfreeze_candidate(mm, cand, pte, nr_saved, nr_frozen);
return result;
}
folio_ref_freeze() takes the source folio's refcount to zero:
static inline int folio_ref_freeze(struct folio *folio, int count)
{
return page_ref_freeze(&folio->page, count);
}
static inline int page_ref_freeze(struct page *page, int count)
{
int ret = likely(atomic_cmpxchg(&page->_refcount, count, 0) == count);
...
return ret;
}
While collapse_freeze_candidate() still holds the source folio lock,
deferred_split_scan() can call deferred_split_isolate():
static unsigned long deferred_split_scan(struct shrinker *shrink,
struct shrink_control *sc)
{
LIST_HEAD(dispose);
struct folio *folio, *next;
int split = 0;
unsigned long isolated;
isolated = list_lru_shrink_walk_irq(&deferred_split_lru, sc,
deferred_split_isolate, &dispose);
}
static enum lru_status deferred_split_isolate(struct list_head *item,
struct list_lru_one *lru,
void *cb_arg)
{
struct folio *folio = container_of(item, struct folio, _deferred_list);
struct list_head *freeable = cb_arg;
if (folio_try_get(folio)) {
list_lru_isolate_move(lru, item, freeable);
return LRU_REMOVED;
}
/*
* We lost race with folio_put(). Read folio state before the
* isolate: folio_unqueue_deferred_split() checks list_empty()
* locklessly, so once removed the folio can be freed any time.
*/
if (folio_test_partially_mapped(folio)) {
folio_clear_partially_mapped(folio);
mod_mthp_stat(folio_order(folio),
MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1);
}
list_lru_isolate(lru, item);
return LRU_REMOVED;
}
And folio_try_get() fails because the source folio has a frozen refcount.
deferred_split_isolate() treats the failure as a race with folio_put(),
clears PG_partially_mapped and its MTHP_STAT_NR_ANON_PARTIALLY_MAPPED
accounting when set, then removes the folio from deferred_split_lru ...
>+
>+ i += nr;
>+ addr += nr * PAGE_SIZE;
>+ }
>+
>+ cand->state = CAND_FROZEN;
>+ return SCAN_SUCCEED;
>+
>+unfreeze:
>+ collapse_unfreeze_candidate(mm, cand, pte, nr_saved, nr_frozen);
>+ return result;
And collapse_unfreeze_candidate() restores the source PTEs and refcount,
then unlocks and puts the source folio :)
The source folio is not added back to the deferred split queue, so an
underused or partially mapped folio can remain off the queue.
Should collapse unqueue the source folio after folio_ref_freeze()
succeeds, remember whether it was on the deferred split queue and whether
PG_partially_mapped was set, then requeue it if
collapse_unfreeze_candidate() restores the source folio?
[...]
Cheers, Lance
next prev parent reply other threads:[~2026-08-24 13:12 UTC|newest]
Thread overview: 98+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-16 22:45 [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 01/57] mm: add pte_folio() Kiryl Shutsemau
2026-08-18 16:38 ` Rik van Riel
2026-08-18 18:13 ` David Hildenbrand (Arm)
2026-08-18 20:04 ` Rik van Riel
2026-08-19 7:57 ` David Hildenbrand (Arm)
2026-08-18 17:09 ` David Hildenbrand (Arm)
2026-08-18 18:30 ` Lorenzo Stoakes (ARM)
2026-08-20 10:52 ` Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 02/57] mm: add pte_none_or_zero() Kiryl Shutsemau
2026-08-17 17:57 ` David Hildenbrand (Arm)
2026-08-20 11:03 ` Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 03/57] mm/collapse: add collapse.h for the shared collapse state Kiryl Shutsemau
2026-08-18 10:50 ` Lorenzo Stoakes (ARM)
2026-08-20 11:06 ` Kiryl Shutsemau
2026-08-19 14:19 ` David Hildenbrand (Arm)
2026-08-20 11:11 ` Kiryl Shutsemau
2026-08-24 11:47 ` David Hildenbrand (Arm)
2026-08-24 12:10 ` Kiryl Shutsemau
2026-08-24 12:28 ` David Hildenbrand (Arm)
2026-08-24 12:36 ` Kiryl Shutsemau
2026-08-24 14:09 ` Lorenzo Stoakes (ARM)
2026-08-16 22:45 ` [RFC PATCH 04/57] mm/collapse: rename mthp_present_ptes to eligible_ptes Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 05/57] mm/collapse: state what a collapse may do in the policy Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 06/57] mm/collapse: move the smallest collapse order to collapse.h Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 07/57] mm/collapse: sketch the new anonymous collapse engine Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 08/57] mm/collapse: scan a table for what a collapse could use Kiryl Shutsemau
2026-08-24 8:39 ` Lance Yang
2026-08-24 9:36 ` Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 09/57] mm/collapse: collect candidate windows into a round Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 10/57] mm/collapse: run a round and feed the outcomes back Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 11/57] mm/collapse: sketch the passes of a round Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 12/57] mm/collapse: allocate a destination per candidate Kiryl Shutsemau
2026-08-24 11:20 ` Lance Yang
2026-08-24 12:37 ` Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 13/57] mm/collapse: revalidate a round against the VMA Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 14/57] mm/collapse: fault the sources in before the freeze Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 15/57] mm/collapse: check what a candidate would freeze Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 16/57] mm/collapse: freeze the sources behind migration entries Kiryl Shutsemau
2026-08-24 13:12 ` Lance Yang [this message]
2026-08-16 22:45 ` [RFC PATCH 17/57] mm/collapse: copy the sources into the destinations Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 18/57] mm/collapse: install the destinations at PTE level Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 19/57] mm/collapse: install a PMD leaf as the terminal layer Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 20/57] mm/collapse: put the sources back Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 21/57] mm/collapse: settle whatever the round reached Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 22/57] mm/collapse: walk a table with a selection cursor Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 23/57] mm/collapse: give a refused region a second chance Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 24/57] mm/collapse: report each candidate's outcome to tracing Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 25/57] mm/collapse: collapse anonymous memory with the new engine Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 26/57] mm/collapse: give collapse_single_pmd() the range to work on Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 27/57] mm/collapse: scan the windows a VMA can hold Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 28/57] mm/collapse: remove the mechanism the engine replaces Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 29/57] mm/collapse: move what a collapse is judged on into collapse.c Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 30/57] mm/collapse: name the max_ptes ceiling after collapse Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 31/57] mm/khugepaged: count collapses where khugepaged makes them Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 32/57] mm/collapse: move the file collapse into collapse.c Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 33/57] mm/collapse: split collapse into a scan and a run Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 34/57] mm/collapse: implement MADV_COLLAPSE in madvise.c Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 35/57] mm/madvise: drop MADV_COLLAPSE's redundant mm reference Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 36/57] mm/collapse: report what the scan found Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 37/57] mm/collapse: report what the fault-in pass paid Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 38/57] mm/collapse: report the round, and what it made faulters wait Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 39/57] mm/collapse: name the file collapse's tracepoints after collapse Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 40/57] mm/collapse: remove the tracepoints of the mechanism that is gone Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 41/57] mm/collapse: give collapse its own trace header Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 42/57] mm/collapse: allow error injection into the freeze Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 43/57] mm/khugepaged: check the scan budget before the work, not after Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 44/57] mm/khugepaged: hold the address space open across a scan Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 45/57] mm/collapse: take a per-VMA read lock for the round Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 46/57] mm/khugepaged: scan under a per-VMA read lock Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 47/57] mm/madvise: collapse " Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 48/57] mm/collapse: assert the mm reference the engine relies on Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 49/57] mm/khugepaged: drop the mmap_lock barrier from __khugepaged_exit() Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 50/57] selftests/mm: attribute collapses by candidate event alone Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 51/57] selftests/mm: cover collapse inside a sub-PMD VMA Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 52/57] selftests/mm: cover a hole-y window in " Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 53/57] selftests/mm: cover collapse of mlocked ranges Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 54/57] selftests/mm: cover collapse beside a MADV_FREE'd page Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 55/57] selftests/mm: cover collapse beside a pinned page Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 56/57] selftests/mm: cover the scaled max_ptes_shared limit Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 57/57] MAINTAINERS: add an entry for collapse Kiryl Shutsemau
2026-08-17 8:04 ` Lorenzo Stoakes (ARM)
2026-08-17 8:08 ` David Hildenbrand (Arm)
2026-08-17 10:12 ` Kiryl Shutsemau
2026-08-17 2:02 ` [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives Zi Yan
2026-08-17 10:07 ` Kiryl Shutsemau
2026-08-17 8:52 ` Lorenzo Stoakes (ARM)
2026-08-17 13:38 ` Kiryl Shutsemau
2026-08-18 13:06 ` Lorenzo Stoakes (ARM)
2026-08-18 14:12 ` David Hildenbrand (Arm)
2026-08-18 14:33 ` Lorenzo Stoakes (ARM)
2026-08-19 18:08 ` Kiryl Shutsemau
2026-08-18 14:15 ` David Hildenbrand (Arm)
2026-08-18 14:41 ` Lorenzo Stoakes (ARM)
2026-08-19 18:22 ` Kiryl Shutsemau
2026-08-19 18:14 ` Kiryl Shutsemau
2026-08-18 13:55 ` David Hildenbrand (Arm)
2026-08-19 17:09 ` Kiryl Shutsemau
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260824131224.73344-1-lance.yang@linux.dev \
--to=lance.yang@linux.dev \
--cc=agordeev@linux.ibm.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=bpf@vger.kernel.org \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=hughd@google.com \
--cc=jannh@google.com \
--cc=kas@kernel.org \
--cc=kirill@shutemov.name \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=ljs@kernel.org \
--cc=mhiramat@kernel.org \
--cc=mhocko@suse.com \
--cc=nico.pache@linux.dev \
--cc=pfalcato@suse.de \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shuah@kernel.org \
--cc=surenb@google.com \
--cc=usama.anjum@arm.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox