* [PATCH v2 0/2] Fix deferred_split_isolate() and drop the split workaround
@ 2026-08-31 9:15 Kiryl Shutsemau
2026-08-31 9:15 ` [PATCH v2 1/2] mm/huge_memory: do not touch frozen folios in deferred_split_isolate() Kiryl Shutsemau
2026-08-31 9:15 ` [PATCH v2 2/2] mm/huge_memory: dequeue the deferred split after the split freeze Kiryl Shutsemau
0 siblings, 2 replies; 11+ messages in thread
From: Kiryl Shutsemau @ 2026-08-31 9:15 UTC (permalink / raw)
To: akpm, david, ljs, hannes, usama.arif
Cc: lance.yang, ziy, kasong, hughd, baolin.wang, baohua, liam,
nico.pache, dev.jain, ryan.roberts, balbirs, kas, linux-mm,
linux-kernel
From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
deferred_split_isolate() probes each queued folio with folio_try_get().
folio_try_get() failure is treated as a lost race with folio_put().
It leads to wrong results when !folio_try_get() was not caused by
folio_put(): for a frozen folio, PG_partially_mapped gets wrongfully
cleared and the folio dropped from the queue.
It came up in the review of my collapse RFC series:
https://lore.kernel.org/all/20260824131224.73344-1-lance.yang@linux.dev/
The bug is inert in upstream code:
- __folio_split() works around it;
- __folio_migrate_mapping() freezes a folio it is about to replace;
- reclaim freezes only what try_to_unmap() already unmapped.
No stable@ needed. But my collapse rework steps on it, so it is worth
fixing.
The branch the first patch removes also hid an inert, pre-existing bug in
the zone device path:
https://lore.kernel.org/all/20260827163838.1813081-1-usama.arif@linux.dev/
The first patch fixes deferred_split_isolate().
The second patch removes the workaround for this deferred_split_isolate()
behaviour from __folio_freeze_and_split_unmapped().
Tested in a VM: split_huge_page_test, folio_split_race_test and cow pass.
Also ran a test that leaves 16 partially mapped THPs on the deferred
split queue and drives thp-deferred_split through debugfs, checking
nr_anon_partially_mapped.
v1: https://lore.kernel.org/all/20260826162101.1314941-1-kirill@shutemov.name/
Changes since v1:
- Spell out in the first patch who clears PG_partially_mapped and the
stat once the shrinker stops doing it (David Hildenbrand).
- Add Fixes: and Closes: tags (David Hildenbrand, Zi Yan).
- Drop the three __folio_freeze_and_split_unmapped() cleanups: Kairui
Song's swap THP cleanup series already carries the same changes.
- Collect review tags.
Kiryl Shutsemau (Meta) (2):
mm/huge_memory: do not touch frozen folios in deferred_split_isolate()
mm/huge_memory: dequeue the deferred split after the split freeze
mm/huge_memory.c | 63 +++++++++++++-----------------------------------
1 file changed, 17 insertions(+), 46 deletions(-)
base-commit: 33f61b12d297562321533c048e034b1fb21c1cf3
--
2.54.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v2 1/2] mm/huge_memory: do not touch frozen folios in deferred_split_isolate()
2026-08-31 9:15 [PATCH v2 0/2] Fix deferred_split_isolate() and drop the split workaround Kiryl Shutsemau
@ 2026-08-31 9:15 ` Kiryl Shutsemau
2026-08-31 11:16 ` Lance Yang
` (3 more replies)
2026-08-31 9:15 ` [PATCH v2 2/2] mm/huge_memory: dequeue the deferred split after the split freeze Kiryl Shutsemau
1 sibling, 4 replies; 11+ messages in thread
From: Kiryl Shutsemau @ 2026-08-31 9:15 UTC (permalink / raw)
To: akpm, david, ljs, hannes, usama.arif
Cc: lance.yang, ziy, kasong, hughd, baolin.wang, baohua, liam,
nico.pache, dev.jain, ryan.roberts, balbirs, kas, linux-mm,
linux-kernel
From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
deferred_split_isolate() probes each queued folio with folio_try_get().
folio_try_get() failure is treated as a lost race with folio_put(): clear
PG_partially_mapped, correct MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, take
the folio off the queue.
The folio_put() race is the most common case for !folio_try_get(), but
it is not the only option. Another scenario is folio_ref_freeze().
A zero refcount in such cases does not mean the folio is going away. It
means "don't touch me" and current deferred_split_isolate() doesn't
respect it. It can lead to unqueueing folios from the deferred list for
no reason:
CPU 0 CPU 1
--------------------------- ------------------------------
freeze a mapped folio deferred_split_scan()
folio_ref_freeze() folio_try_get() fails
folio_clear_partially_mapped()
NR_ANON_PARTIALLY_MAPPED--
folio off the queue
give up, put it back
folio_ref_unfreeze()
The folio is still partially mapped, but it is no longer a split candidate.
Nothing queues it again until part of it is unmapped once more.
Skip the folio instead: whoever freezes the folio, owns it and owner is
responsible for its fate. It also covers the folio_put() case:
__folio_put() unqueues the folio via folio_unqueue_deferred_split().
Nothing is lost by skipping. Everything that frees a queued folio
unqueues it first, and folio_unqueue_deferred_split() clears
PG_partially_mapped and brings MTHP_STAT_NR_ANON_PARTIALLY_MAPPED down
on the way:
__folio_put(), folios_put_refs() mm/folio.c
__folio_migrate_mapping() mm/migrate.c
shrink_folio_list() mm/vmscan.c
__folio_freeze_and_split_unmapped() does the same by hand, under the
list_lru lock it holds across the freeze. A freeze that ends in
folio_ref_unfreeze() leaves a folio that is still partially mapped and
still belongs on the queue.
Fixes: 8422acdc97ed ("mm: introduce a pageflag for partially mapped folios")
Reported-by: Lance Yang <lance.yang@linux.dev>
Closes: https://lore.kernel.org/all/20260824131224.73344-1-lance.yang@linux.dev/
Assisted-by: Claude-Code:claude-opus-5
Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
Reviewed-by: Zi Yan <ziy@nvidia.com>
Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
---
mm/huge_memory.c | 19 ++++---------------
1 file changed, 4 insertions(+), 15 deletions(-)
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index ced400f72d43..6281ed993243 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -4590,22 +4590,11 @@ static enum lru_status deferred_split_isolate(struct list_head *item,
struct folio *folio = container_of(item, struct folio, _deferred_list);
struct list_head *freeable = cb_arg;
- if (folio_try_get(folio)) {
- list_lru_isolate_move(lru, item, freeable);
- return LRU_REMOVED;
- }
+ /* Lost race to folio_put() or the folio is under folio_ref_freeze() */
+ if (!folio_try_get(folio))
+ return LRU_SKIP;
- /*
- * We lost race with folio_put(). Read folio state before the
- * isolate: folio_unqueue_deferred_split() checks list_empty()
- * locklessly, so once removed the folio can be freed any time.
- */
- if (folio_test_partially_mapped(folio)) {
- folio_clear_partially_mapped(folio);
- mod_mthp_stat(folio_order(folio),
- MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1);
- }
- list_lru_isolate(lru, item);
+ list_lru_isolate_move(lru, item, freeable);
return LRU_REMOVED;
}
--
2.54.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [PATCH v2 2/2] mm/huge_memory: dequeue the deferred split after the split freeze
2026-08-31 9:15 [PATCH v2 0/2] Fix deferred_split_isolate() and drop the split workaround Kiryl Shutsemau
2026-08-31 9:15 ` [PATCH v2 1/2] mm/huge_memory: do not touch frozen folios in deferred_split_isolate() Kiryl Shutsemau
@ 2026-08-31 9:15 ` Kiryl Shutsemau
2026-08-31 11:34 ` Lance Yang
2026-09-07 13:55 ` David Hildenbrand (Arm)
1 sibling, 2 replies; 11+ messages in thread
From: Kiryl Shutsemau @ 2026-08-31 9:15 UTC (permalink / raw)
To: akpm, david, ljs, hannes, usama.arif
Cc: lance.yang, ziy, kasong, hughd, baolin.wang, baohua, liam,
nico.pache, dev.jain, ryan.roberts, balbirs, kas, linux-mm,
linux-kernel
From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
__folio_freeze_and_split_unmapped() takes the deferred split list_lru lock
across the freeze. It is only there to stop deferred_split_scan() from
touching the folio under split.
With deferred_split_isolate() fixed, the workaround can be dropped.
Unqueue the folio after folio_ref_freeze(), the way
__folio_migrate_mapping() does: folio_unqueue_deferred_split() needs a
zero refcount and a memcg still set, and both hold there.
If the split is called from deferred_split_scan(), the unqueue is a
no-op -- the folio is already removed from the list. But
PG_partially_mapped is still set, so it has to be cleared here or
MTHP_STAT_NR_ANON_PARTIALLY_MAPPED never comes back down.
Assisted-by: Claude-Code:claude-opus-5
Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
Reviewed-by: Zi Yan <ziy@nvidia.com>
Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
---
mm/huge_memory.c | 44 +++++++++++++-------------------------------
1 file changed, 13 insertions(+), 31 deletions(-)
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 6281ed993243..c84e8cbc986d 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -3931,41 +3931,27 @@ static int __folio_freeze_and_split_unmapped(struct folio *folio, unsigned int n
struct folio *end_folio = folio_next(folio);
struct folio *new_folio, *next;
int old_order = folio_order(folio);
- struct list_lru_one *lru;
- bool dequeue_deferred;
int ret = 0;
VM_WARN_ON_ONCE(!mapping && end);
- /*
- * If this folio can be on the deferred split queue, lock out
- * the shrinker before freezing the ref. If the shrinker sees
- * a 0-ref folio, it assumes it beat folio_put() to the list
- * lock and must clean up the LRU state - the same dequeue we
- * will do below as part of the split.
- */
- dequeue_deferred = folio_test_anon(folio) && old_order > 1;
- if (dequeue_deferred) {
- struct mem_cgroup *memcg;
- rcu_read_lock();
- memcg = folio_memcg(folio);
- lru = list_lru_lock(&deferred_split_lru,
- folio_nid(folio), &memcg);
- }
if (folio_ref_freeze(folio, folio_cache_ref_count(folio) + 1)) {
struct swap_cluster_info *ci = NULL;
struct lruvec *lruvec;
- if (dequeue_deferred) {
- __list_lru_del(&deferred_split_lru, lru,
- &folio->_deferred_list, folio_nid(folio));
- if (folio_test_partially_mapped(folio)) {
- folio_clear_partially_mapped(folio);
- mod_mthp_stat(old_order,
- MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1);
- }
- list_lru_unlock(lru);
- rcu_read_unlock();
+ /* Take off the deferred split queue while frozen and memcg set */
+ folio_unqueue_deferred_split(folio);
+
+ /*
+ * deferred_split_scan() takes the folio off the queue before it
+ * splits it, so the unqueue above finds an empty list and
+ * leaves PG_partially_mapped set.
+ * Clear it here: the flag does not survive the split.
+ */
+ if (folio_test_partially_mapped(folio)) {
+ folio_clear_partially_mapped(folio);
+ mod_mthp_stat(old_order,
+ MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1);
}
if (mapping) {
@@ -4067,10 +4053,6 @@ static int __folio_freeze_and_split_unmapped(struct folio *folio, unsigned int n
if (ci)
swap_cluster_unlock(ci);
} else {
- if (dequeue_deferred) {
- list_lru_unlock(lru);
- rcu_read_unlock();
- }
return -EAGAIN;
}
--
2.54.0
^ permalink raw reply related [flat|nested] 11+ messages in thread
* Re: [PATCH v2 1/2] mm/huge_memory: do not touch frozen folios in deferred_split_isolate()
2026-08-31 9:15 ` [PATCH v2 1/2] mm/huge_memory: do not touch frozen folios in deferred_split_isolate() Kiryl Shutsemau
@ 2026-08-31 11:16 ` Lance Yang
2026-08-31 11:23 ` Usama Arif
` (2 subsequent siblings)
3 siblings, 0 replies; 11+ messages in thread
From: Lance Yang @ 2026-08-31 11:16 UTC (permalink / raw)
To: Kiryl Shutsemau
Cc: ziy, kasong, hannes, david, hughd, baolin.wang, baohua, liam,
usama.arif, nico.pache, dev.jain, ryan.roberts, balbirs, kas,
linux-mm, linux-kernel, akpm, ljs
On 2026/8/31 17:15, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> deferred_split_isolate() probes each queued folio with folio_try_get().
> folio_try_get() failure is treated as a lost race with folio_put(): clear
> PG_partially_mapped, correct MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, take
> the folio off the queue.
>
> The folio_put() race is the most common case for !folio_try_get(), but
> it is not the only option. Another scenario is folio_ref_freeze().
>
> A zero refcount in such cases does not mean the folio is going away. It
> means "don't touch me" and current deferred_split_isolate() doesn't
> respect it. It can lead to unqueueing folios from the deferred list for
> no reason:
>
> CPU 0 CPU 1
> --------------------------- ------------------------------
> freeze a mapped folio deferred_split_scan()
> folio_ref_freeze() folio_try_get() fails
> folio_clear_partially_mapped()
> NR_ANON_PARTIALLY_MAPPED--
> folio off the queue
> give up, put it back
> folio_ref_unfreeze()
>
> The folio is still partially mapped, but it is no longer a split candidate.
> Nothing queues it again until part of it is unmapped once more.
>
> Skip the folio instead: whoever freezes the folio, owns it and owner is
> responsible for its fate. It also covers the folio_put() case:
> __folio_put() unqueues the folio via folio_unqueue_deferred_split().
>
> Nothing is lost by skipping. Everything that frees a queued folio
> unqueues it first, and folio_unqueue_deferred_split() clears
> PG_partially_mapped and brings MTHP_STAT_NR_ANON_PARTIALLY_MAPPED down
> on the way:
>
> __folio_put(), folios_put_refs() mm/folio.c
> __folio_migrate_mapping() mm/migrate.c
> shrink_folio_list() mm/vmscan.c
>
> __folio_freeze_and_split_unmapped() does the same by hand, under the
> list_lru lock it holds across the freeze. A freeze that ends in
> folio_ref_unfreeze() leaves a folio that is still partially mapped and
> still belongs on the queue.
>
> Fixes: 8422acdc97ed ("mm: introduce a pageflag for partially mapped folios")
> Reported-by: Lance Yang <lance.yang@linux.dev>
> Closes: https://lore.kernel.org/all/20260824131224.73344-1-lance.yang@linux.dev/
> Assisted-by: Claude-Code:claude-opus-5
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> Reviewed-by: Zi Yan <ziy@nvidia.com>
> Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
> ---
LGTM.
Reviewed-by: Lance Yang <lance.yang@linux.dev>
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 1/2] mm/huge_memory: do not touch frozen folios in deferred_split_isolate()
2026-08-31 9:15 ` [PATCH v2 1/2] mm/huge_memory: do not touch frozen folios in deferred_split_isolate() Kiryl Shutsemau
2026-08-31 11:16 ` Lance Yang
@ 2026-08-31 11:23 ` Usama Arif
2026-09-01 1:38 ` Baolin Wang
2026-09-07 13:48 ` David Hildenbrand (Arm)
3 siblings, 0 replies; 11+ messages in thread
From: Usama Arif @ 2026-08-31 11:23 UTC (permalink / raw)
To: Kiryl Shutsemau
Cc: Usama Arif, akpm, david, ljs, hannes, lance.yang, ziy, kasong,
hughd, baolin.wang, baohua, liam, nico.pache, dev.jain,
ryan.roberts, balbirs, kas, linux-mm, linux-kernel
On Mon, 31 Aug 2026 10:15:13 +0100 Kiryl Shutsemau <kirill@shutemov.name> wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> deferred_split_isolate() probes each queued folio with folio_try_get().
> folio_try_get() failure is treated as a lost race with folio_put(): clear
> PG_partially_mapped, correct MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, take
> the folio off the queue.
>
> The folio_put() race is the most common case for !folio_try_get(), but
> it is not the only option. Another scenario is folio_ref_freeze().
>
> A zero refcount in such cases does not mean the folio is going away. It
> means "don't touch me" and current deferred_split_isolate() doesn't
> respect it. It can lead to unqueueing folios from the deferred list for
> no reason:
>
> CPU 0 CPU 1
> --------------------------- ------------------------------
> freeze a mapped folio deferred_split_scan()
> folio_ref_freeze() folio_try_get() fails
> folio_clear_partially_mapped()
> NR_ANON_PARTIALLY_MAPPED--
> folio off the queue
> give up, put it back
> folio_ref_unfreeze()
>
> The folio is still partially mapped, but it is no longer a split candidate.
> Nothing queues it again until part of it is unmapped once more.
>
> Skip the folio instead: whoever freezes the folio, owns it and owner is
> responsible for its fate. It also covers the folio_put() case:
> __folio_put() unqueues the folio via folio_unqueue_deferred_split().
>
> Nothing is lost by skipping. Everything that frees a queued folio
> unqueues it first, and folio_unqueue_deferred_split() clears
> PG_partially_mapped and brings MTHP_STAT_NR_ANON_PARTIALLY_MAPPED down
> on the way:
>
> __folio_put(), folios_put_refs() mm/folio.c
> __folio_migrate_mapping() mm/migrate.c
> shrink_folio_list() mm/vmscan.c
>
> __folio_freeze_and_split_unmapped() does the same by hand, under the
> list_lru lock it holds across the freeze. A freeze that ends in
> folio_ref_unfreeze() leaves a folio that is still partially mapped and
> still belongs on the queue.
>
> Fixes: 8422acdc97ed ("mm: introduce a pageflag for partially mapped folios")
> Reported-by: Lance Yang <lance.yang@linux.dev>
> Closes: https://lore.kernel.org/all/20260824131224.73344-1-lance.yang@linux.dev/
> Assisted-by: Claude-Code:claude-opus-5
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> Reviewed-by: Zi Yan <ziy@nvidia.com>
> Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
> ---
> mm/huge_memory.c | 19 ++++---------------
> 1 file changed, 4 insertions(+), 15 deletions(-)
>
> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
> index ced400f72d43..6281ed993243 100644
> --- a/mm/huge_memory.c
> +++ b/mm/huge_memory.c
> @@ -4590,22 +4590,11 @@ static enum lru_status deferred_split_isolate(struct list_head *item,
> struct folio *folio = container_of(item, struct folio, _deferred_list);
> struct list_head *freeable = cb_arg;
>
> - if (folio_try_get(folio)) {
> - list_lru_isolate_move(lru, item, freeable);
> - return LRU_REMOVED;
> - }
> + /* Lost race to folio_put() or the folio is under folio_ref_freeze() */
> + if (!folio_try_get(folio))
> + return LRU_SKIP;
>
> - /*
> - * We lost race with folio_put(). Read folio state before the
> - * isolate: folio_unqueue_deferred_split() checks list_empty()
> - * locklessly, so once removed the folio can be freed any time.
> - */
> - if (folio_test_partially_mapped(folio)) {
> - folio_clear_partially_mapped(folio);
> - mod_mthp_stat(folio_order(folio),
> - MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1);
> - }
> - list_lru_isolate(lru, item);
> + list_lru_isolate_move(lru, item, freeable);
> return LRU_REMOVED;
> }
Acked-by: Usama Arif <usama.arif@linux.dev>
>
> --
> 2.54.0
>
>
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 2/2] mm/huge_memory: dequeue the deferred split after the split freeze
2026-08-31 9:15 ` [PATCH v2 2/2] mm/huge_memory: dequeue the deferred split after the split freeze Kiryl Shutsemau
@ 2026-08-31 11:34 ` Lance Yang
2026-09-07 13:55 ` David Hildenbrand (Arm)
1 sibling, 0 replies; 11+ messages in thread
From: Lance Yang @ 2026-08-31 11:34 UTC (permalink / raw)
To: Kiryl Shutsemau
Cc: ziy, kasong, usama.arif, hannes, ljs, hughd, baolin.wang, baohua,
liam, nico.pache, dev.jain, david, akpm, ryan.roberts, balbirs,
kas, linux-mm, linux-kernel
On 2026/8/31 17:15, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> __folio_freeze_and_split_unmapped() takes the deferred split list_lru lock
> across the freeze. It is only there to stop deferred_split_scan() from
> touching the folio under split.
>
> With deferred_split_isolate() fixed, the workaround can be dropped.
>
> Unqueue the folio after folio_ref_freeze(), the way
> __folio_migrate_mapping() does: folio_unqueue_deferred_split() needs a
> zero refcount and a memcg still set, and both hold there.
>
> If the split is called from deferred_split_scan(), the unqueue is a
> no-op -- the folio is already removed from the list. But
> PG_partially_mapped is still set, so it has to be cleared here or
> MTHP_STAT_NR_ANON_PARTIALLY_MAPPED never comes back down.
>
> Assisted-by: Claude-Code:claude-opus-5
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> Reviewed-by: Zi Yan <ziy@nvidia.com>
> Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
> Acked-by: David Hildenbrand (Arm) <david@kernel.org>
> ---
Nothing jumped out at me :)
Reviewed-by: Lance Yang <lance.yang@linux.dev>
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 1/2] mm/huge_memory: do not touch frozen folios in deferred_split_isolate()
2026-08-31 9:15 ` [PATCH v2 1/2] mm/huge_memory: do not touch frozen folios in deferred_split_isolate() Kiryl Shutsemau
2026-08-31 11:16 ` Lance Yang
2026-08-31 11:23 ` Usama Arif
@ 2026-09-01 1:38 ` Baolin Wang
2026-09-07 13:48 ` David Hildenbrand (Arm)
3 siblings, 0 replies; 11+ messages in thread
From: Baolin Wang @ 2026-09-01 1:38 UTC (permalink / raw)
To: Kiryl Shutsemau, akpm, david, ljs, hannes, usama.arif
Cc: lance.yang, ziy, kasong, hughd, baohua, liam, nico.pache,
dev.jain, ryan.roberts, balbirs, kas, linux-mm, linux-kernel
On 8/31/26 5:15 PM, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> deferred_split_isolate() probes each queued folio with folio_try_get().
> folio_try_get() failure is treated as a lost race with folio_put(): clear
> PG_partially_mapped, correct MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, take
> the folio off the queue.
>
> The folio_put() race is the most common case for !folio_try_get(), but
> it is not the only option. Another scenario is folio_ref_freeze().
>
> A zero refcount in such cases does not mean the folio is going away. It
> means "don't touch me" and current deferred_split_isolate() doesn't
> respect it. It can lead to unqueueing folios from the deferred list for
> no reason:
>
> CPU 0 CPU 1
> --------------------------- ------------------------------
> freeze a mapped folio deferred_split_scan()
> folio_ref_freeze() folio_try_get() fails
> folio_clear_partially_mapped()
> NR_ANON_PARTIALLY_MAPPED--
> folio off the queue
> give up, put it back
> folio_ref_unfreeze()
>
> The folio is still partially mapped, but it is no longer a split candidate.
> Nothing queues it again until part of it is unmapped once more.
>
> Skip the folio instead: whoever freezes the folio, owns it and owner is
> responsible for its fate. It also covers the folio_put() case:
> __folio_put() unqueues the folio via folio_unqueue_deferred_split().
>
> Nothing is lost by skipping. Everything that frees a queued folio
> unqueues it first, and folio_unqueue_deferred_split() clears
> PG_partially_mapped and brings MTHP_STAT_NR_ANON_PARTIALLY_MAPPED down
> on the way:
>
> __folio_put(), folios_put_refs() mm/folio.c
> __folio_migrate_mapping() mm/migrate.c
> shrink_folio_list() mm/vmscan.c
>
> __folio_freeze_and_split_unmapped() does the same by hand, under the
> list_lru lock it holds across the freeze. A freeze that ends in
> folio_ref_unfreeze() leaves a folio that is still partially mapped and
> still belongs on the queue.
>
> Fixes: 8422acdc97ed ("mm: introduce a pageflag for partially mapped folios")
> Reported-by: Lance Yang <lance.yang@linux.dev>
> Closes: https://lore.kernel.org/all/20260824131224.73344-1-lance.yang@linux.dev/
> Assisted-by: Claude-Code:claude-opus-5
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> Reviewed-by: Zi Yan <ziy@nvidia.com>
> Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
> ---
LGTM.
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 1/2] mm/huge_memory: do not touch frozen folios in deferred_split_isolate()
2026-08-31 9:15 ` [PATCH v2 1/2] mm/huge_memory: do not touch frozen folios in deferred_split_isolate() Kiryl Shutsemau
` (2 preceding siblings ...)
2026-09-01 1:38 ` Baolin Wang
@ 2026-09-07 13:48 ` David Hildenbrand (Arm)
3 siblings, 0 replies; 11+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-07 13:48 UTC (permalink / raw)
To: Kiryl Shutsemau, akpm, ljs, hannes, usama.arif
Cc: lance.yang, ziy, kasong, hughd, baolin.wang, baohua, liam,
nico.pache, dev.jain, ryan.roberts, balbirs, kas, linux-mm,
linux-kernel
On 8/31/26 11:15, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> deferred_split_isolate() probes each queued folio with folio_try_get().
> folio_try_get() failure is treated as a lost race with folio_put(): clear
> PG_partially_mapped, correct MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, take
> the folio off the queue.
>
> The folio_put() race is the most common case for !folio_try_get(), but
> it is not the only option. Another scenario is folio_ref_freeze().
>
> A zero refcount in such cases does not mean the folio is going away. It
> means "don't touch me" and current deferred_split_isolate() doesn't
> respect it. It can lead to unqueueing folios from the deferred list for
> no reason:
>
> CPU 0 CPU 1
> --------------------------- ------------------------------
> freeze a mapped folio deferred_split_scan()
> folio_ref_freeze() folio_try_get() fails
> folio_clear_partially_mapped()
> NR_ANON_PARTIALLY_MAPPED--
> folio off the queue
> give up, put it back
> folio_ref_unfreeze()
>
> The folio is still partially mapped, but it is no longer a split candidate.
> Nothing queues it again until part of it is unmapped once more.
>
> Skip the folio instead: whoever freezes the folio, owns it and owner is
> responsible for its fate. It also covers the folio_put() case:
> __folio_put() unqueues the folio via folio_unqueue_deferred_split().
>
> Nothing is lost by skipping. Everything that frees a queued folio
> unqueues it first, and folio_unqueue_deferred_split() clears
> PG_partially_mapped and brings MTHP_STAT_NR_ANON_PARTIALLY_MAPPED down
> on the way:
>
> __folio_put(), folios_put_refs() mm/folio.c
> __folio_migrate_mapping() mm/migrate.c
> shrink_folio_list() mm/vmscan.c
>
> __folio_freeze_and_split_unmapped() does the same by hand, under the
> list_lru lock it holds across the freeze. A freeze that ends in
> folio_ref_unfreeze() leaves a folio that is still partially mapped and
> still belongs on the queue.
>
> Fixes: 8422acdc97ed ("mm: introduce a pageflag for partially mapped folios")
> Reported-by: Lance Yang <lance.yang@linux.dev>
> Closes: https://lore.kernel.org/all/20260824131224.73344-1-lance.yang@linux.dev/
> Assisted-by: Claude-Code:claude-opus-5
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> Reviewed-by: Zi Yan <ziy@nvidia.com>
> Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
> ---
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
--
Cheers,
David
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 2/2] mm/huge_memory: dequeue the deferred split after the split freeze
2026-08-31 9:15 ` [PATCH v2 2/2] mm/huge_memory: dequeue the deferred split after the split freeze Kiryl Shutsemau
2026-08-31 11:34 ` Lance Yang
@ 2026-09-07 13:55 ` David Hildenbrand (Arm)
2026-09-07 14:12 ` Kiryl Shutsemau
1 sibling, 1 reply; 11+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-07 13:55 UTC (permalink / raw)
To: Kiryl Shutsemau, akpm, ljs, hannes, usama.arif
Cc: lance.yang, ziy, kasong, hughd, baolin.wang, baohua, liam,
nico.pache, dev.jain, ryan.roberts, balbirs, kas, linux-mm,
linux-kernel
On 8/31/26 11:15, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> __folio_freeze_and_split_unmapped() takes the deferred split list_lru lock
> across the freeze. It is only there to stop deferred_split_scan() from
> touching the folio under split.
>
> With deferred_split_isolate() fixed, the workaround can be dropped.
>
> Unqueue the folio after folio_ref_freeze(), the way
> __folio_migrate_mapping() does: folio_unqueue_deferred_split() needs a
> zero refcount and a memcg still set, and both hold there.
>
> If the split is called from deferred_split_scan(), the unqueue is a
> no-op -- the folio is already removed from the list. But
> PG_partially_mapped is still set, so it has to be cleared here or
> MTHP_STAT_NR_ANON_PARTIALLY_MAPPED never comes back down.
>
> Assisted-by: Claude-Code:claude-opus-5
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> Reviewed-by: Zi Yan <ziy@nvidia.com>
> Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
> Acked-by: David Hildenbrand (Arm) <david@kernel.org>
> ---
> mm/huge_memory.c | 44 +++++++++++++-------------------------------
> 1 file changed, 13 insertions(+), 31 deletions(-)
>
> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
> index 6281ed993243..c84e8cbc986d 100644
> --- a/mm/huge_memory.c
> +++ b/mm/huge_memory.c
> @@ -3931,41 +3931,27 @@ static int __folio_freeze_and_split_unmapped(struct folio *folio, unsigned int n
> struct folio *end_folio = folio_next(folio);
> struct folio *new_folio, *next;
> int old_order = folio_order(folio);
> - struct list_lru_one *lru;
> - bool dequeue_deferred;
> int ret = 0;
>
> VM_WARN_ON_ONCE(!mapping && end);
> - /*
> - * If this folio can be on the deferred split queue, lock out
> - * the shrinker before freezing the ref. If the shrinker sees
> - * a 0-ref folio, it assumes it beat folio_put() to the list
> - * lock and must clean up the LRU state - the same dequeue we
> - * will do below as part of the split.
> - */
> - dequeue_deferred = folio_test_anon(folio) && old_order > 1;
> - if (dequeue_deferred) {
> - struct mem_cgroup *memcg;
>
> - rcu_read_lock();
> - memcg = folio_memcg(folio);
> - lru = list_lru_lock(&deferred_split_lru,
> - folio_nid(folio), &memcg);
> - }
Nice cleanup.
> if (folio_ref_freeze(folio, folio_cache_ref_count(folio) + 1)) {
> struct swap_cluster_info *ci = NULL;
> struct lruvec *lruvec;
>
> - if (dequeue_deferred) {
> - __list_lru_del(&deferred_split_lru, lru,
> - &folio->_deferred_list, folio_nid(folio));
> - if (folio_test_partially_mapped(folio)) {
> - folio_clear_partially_mapped(folio);
> - mod_mthp_stat(old_order,
> - MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1);
> - }
> - list_lru_unlock(lru);
> - rcu_read_unlock();
> + /* Take off the deferred split queue while frozen and memcg set */
> + folio_unqueue_deferred_split(folio);
> +
> + /*
> + * deferred_split_scan() takes the folio off the queue before it
> + * splits it, so the unqueue above finds an empty list and
> + * leaves PG_partially_mapped set.
> + * Clear it here: the flag does not survive the split.
> + */
> + if (folio_test_partially_mapped(folio)) {
> + folio_clear_partially_mapped(folio);
> + mod_mthp_stat(old_order,
> + MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1);
> }
In general,
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
But I do wonder whether this sequence (that also
__folio_unqueue_deferred_split()) performs would deserve a small local helper in
mm/huge_memory.c
Could be done as a separate cleanup.
--
Cheers,
David
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 2/2] mm/huge_memory: dequeue the deferred split after the split freeze
2026-09-07 13:55 ` David Hildenbrand (Arm)
@ 2026-09-07 14:12 ` Kiryl Shutsemau
2026-09-07 14:30 ` David Hildenbrand (Arm)
0 siblings, 1 reply; 11+ messages in thread
From: Kiryl Shutsemau @ 2026-09-07 14:12 UTC (permalink / raw)
To: David Hildenbrand (Arm)
Cc: akpm, ljs, hannes, usama.arif, lance.yang, ziy, kasong, hughd,
baolin.wang, baohua, liam, nico.pache, dev.jain, ryan.roberts,
balbirs, linux-mm, linux-kernel
On Mon, Sep 07, 2026 at 03:55:53PM +0200, David Hildenbrand (Arm) wrote:
> On 8/31/26 11:15, Kiryl Shutsemau wrote:
> > From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> >
> > __folio_freeze_and_split_unmapped() takes the deferred split list_lru lock
> > across the freeze. It is only there to stop deferred_split_scan() from
> > touching the folio under split.
> >
> > With deferred_split_isolate() fixed, the workaround can be dropped.
> >
> > Unqueue the folio after folio_ref_freeze(), the way
> > __folio_migrate_mapping() does: folio_unqueue_deferred_split() needs a
> > zero refcount and a memcg still set, and both hold there.
> >
> > If the split is called from deferred_split_scan(), the unqueue is a
> > no-op -- the folio is already removed from the list. But
> > PG_partially_mapped is still set, so it has to be cleared here or
> > MTHP_STAT_NR_ANON_PARTIALLY_MAPPED never comes back down.
> >
> > Assisted-by: Claude-Code:claude-opus-5
> > Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> > Reviewed-by: Zi Yan <ziy@nvidia.com>
> > Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
> > Acked-by: David Hildenbrand (Arm) <david@kernel.org>
> > ---
> > mm/huge_memory.c | 44 +++++++++++++-------------------------------
> > 1 file changed, 13 insertions(+), 31 deletions(-)
> >
> > diff --git a/mm/huge_memory.c b/mm/huge_memory.c
> > index 6281ed993243..c84e8cbc986d 100644
> > --- a/mm/huge_memory.c
> > +++ b/mm/huge_memory.c
> > @@ -3931,41 +3931,27 @@ static int __folio_freeze_and_split_unmapped(struct folio *folio, unsigned int n
> > struct folio *end_folio = folio_next(folio);
> > struct folio *new_folio, *next;
> > int old_order = folio_order(folio);
> > - struct list_lru_one *lru;
> > - bool dequeue_deferred;
> > int ret = 0;
> >
> > VM_WARN_ON_ONCE(!mapping && end);
> > - /*
> > - * If this folio can be on the deferred split queue, lock out
> > - * the shrinker before freezing the ref. If the shrinker sees
> > - * a 0-ref folio, it assumes it beat folio_put() to the list
> > - * lock and must clean up the LRU state - the same dequeue we
> > - * will do below as part of the split.
> > - */
> > - dequeue_deferred = folio_test_anon(folio) && old_order > 1;
> > - if (dequeue_deferred) {
> > - struct mem_cgroup *memcg;
> >
> > - rcu_read_lock();
> > - memcg = folio_memcg(folio);
> > - lru = list_lru_lock(&deferred_split_lru,
> > - folio_nid(folio), &memcg);
> > - }
>
> Nice cleanup.
>
> > if (folio_ref_freeze(folio, folio_cache_ref_count(folio) + 1)) {
> > struct swap_cluster_info *ci = NULL;
> > struct lruvec *lruvec;
> >
> > - if (dequeue_deferred) {
> > - __list_lru_del(&deferred_split_lru, lru,
> > - &folio->_deferred_list, folio_nid(folio));
> > - if (folio_test_partially_mapped(folio)) {
> > - folio_clear_partially_mapped(folio);
> > - mod_mthp_stat(old_order,
> > - MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1);
> > - }
> > - list_lru_unlock(lru);
> > - rcu_read_unlock();
> > + /* Take off the deferred split queue while frozen and memcg set */
> > + folio_unqueue_deferred_split(folio);
> > +
> > + /*
> > + * deferred_split_scan() takes the folio off the queue before it
> > + * splits it, so the unqueue above finds an empty list and
> > + * leaves PG_partially_mapped set.
> > + * Clear it here: the flag does not survive the split.
> > + */
> > + if (folio_test_partially_mapped(folio)) {
> > + folio_clear_partially_mapped(folio);
> > + mod_mthp_stat(old_order,
> > + MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1);
> > }
>
> In general,
>
> Acked-by: David Hildenbrand (Arm) <david@kernel.org>
>
> But I do wonder whether this sequence (that also
> __folio_unqueue_deferred_split()) performs would deserve a small local helper in
> mm/huge_memory.c
>
> Could be done as a separate cleanup.
Just to be sure, do you want a helper like this:
static void folio_clear_partially_mapped_stat(struct folio *folio)
{
if (!folio_test_partially_mapped(folio))
return;
folio_clear_partially_mapped(folio);
mod_mthp_stat(folio_order(folio),
MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1);
}
?
--
Kiryl Shutsemau / Kirill A. Shutemov
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v2 2/2] mm/huge_memory: dequeue the deferred split after the split freeze
2026-09-07 14:12 ` Kiryl Shutsemau
@ 2026-09-07 14:30 ` David Hildenbrand (Arm)
0 siblings, 0 replies; 11+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-07 14:30 UTC (permalink / raw)
To: Kiryl Shutsemau
Cc: akpm, ljs, hannes, usama.arif, lance.yang, ziy, kasong, hughd,
baolin.wang, baohua, liam, nico.pache, dev.jain, ryan.roberts,
balbirs, linux-mm, linux-kernel
On 9/7/26 16:12, Kiryl Shutsemau wrote:
> On Mon, Sep 07, 2026 at 03:55:53PM +0200, David Hildenbrand (Arm) wrote:
>> On 8/31/26 11:15, Kiryl Shutsemau wrote:
>>> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>>>
>>> __folio_freeze_and_split_unmapped() takes the deferred split list_lru lock
>>> across the freeze. It is only there to stop deferred_split_scan() from
>>> touching the folio under split.
>>>
>>> With deferred_split_isolate() fixed, the workaround can be dropped.
>>>
>>> Unqueue the folio after folio_ref_freeze(), the way
>>> __folio_migrate_mapping() does: folio_unqueue_deferred_split() needs a
>>> zero refcount and a memcg still set, and both hold there.
>>>
>>> If the split is called from deferred_split_scan(), the unqueue is a
>>> no-op -- the folio is already removed from the list. But
>>> PG_partially_mapped is still set, so it has to be cleared here or
>>> MTHP_STAT_NR_ANON_PARTIALLY_MAPPED never comes back down.
>>>
>>> Assisted-by: Claude-Code:claude-opus-5
>>> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
>>> Reviewed-by: Zi Yan <ziy@nvidia.com>
>>> Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
>>> Acked-by: David Hildenbrand (Arm) <david@kernel.org>
>>> ---
>>> mm/huge_memory.c | 44 +++++++++++++-------------------------------
>>> 1 file changed, 13 insertions(+), 31 deletions(-)
>>>
>>> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
>>> index 6281ed993243..c84e8cbc986d 100644
>>> --- a/mm/huge_memory.c
>>> +++ b/mm/huge_memory.c
>>> @@ -3931,41 +3931,27 @@ static int __folio_freeze_and_split_unmapped(struct folio *folio, unsigned int n
>>> struct folio *end_folio = folio_next(folio);
>>> struct folio *new_folio, *next;
>>> int old_order = folio_order(folio);
>>> - struct list_lru_one *lru;
>>> - bool dequeue_deferred;
>>> int ret = 0;
>>>
>>> VM_WARN_ON_ONCE(!mapping && end);
>>> - /*
>>> - * If this folio can be on the deferred split queue, lock out
>>> - * the shrinker before freezing the ref. If the shrinker sees
>>> - * a 0-ref folio, it assumes it beat folio_put() to the list
>>> - * lock and must clean up the LRU state - the same dequeue we
>>> - * will do below as part of the split.
>>> - */
>>> - dequeue_deferred = folio_test_anon(folio) && old_order > 1;
>>> - if (dequeue_deferred) {
>>> - struct mem_cgroup *memcg;
>>>
>>> - rcu_read_lock();
>>> - memcg = folio_memcg(folio);
>>> - lru = list_lru_lock(&deferred_split_lru,
>>> - folio_nid(folio), &memcg);
>>> - }
>>
>> Nice cleanup.
>>
>>> if (folio_ref_freeze(folio, folio_cache_ref_count(folio) + 1)) {
>>> struct swap_cluster_info *ci = NULL;
>>> struct lruvec *lruvec;
>>>
>>> - if (dequeue_deferred) {
>>> - __list_lru_del(&deferred_split_lru, lru,
>>> - &folio->_deferred_list, folio_nid(folio));
>>> - if (folio_test_partially_mapped(folio)) {
>>> - folio_clear_partially_mapped(folio);
>>> - mod_mthp_stat(old_order,
>>> - MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1);
>>> - }
>>> - list_lru_unlock(lru);
>>> - rcu_read_unlock();
>>> + /* Take off the deferred split queue while frozen and memcg set */
>>> + folio_unqueue_deferred_split(folio);
>>> +
>>> + /*
>>> + * deferred_split_scan() takes the folio off the queue before it
>>> + * splits it, so the unqueue above finds an empty list and
>>> + * leaves PG_partially_mapped set.
>>> + * Clear it here: the flag does not survive the split.
>>> + */
>>> + if (folio_test_partially_mapped(folio)) {
>>> + folio_clear_partially_mapped(folio);
>>> + mod_mthp_stat(old_order,
>>> + MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1);
>>> }
>>
>> In general,
>>
>> Acked-by: David Hildenbrand (Arm) <david@kernel.org>
>>
>> But I do wonder whether this sequence (that also
>> __folio_unqueue_deferred_split()) performs would deserve a small local helper in
>> mm/huge_memory.c
>>
>> Could be done as a separate cleanup.
>
> Just to be sure, do you want a helper like this:
Something like that. I'd probably call it folio_reset_partially_mapped()
--
Cheers,
David
^ permalink raw reply [flat|nested] 11+ messages in thread
end of thread, other threads:[~2026-09-07 14:30 UTC | newest]
Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-31 9:15 [PATCH v2 0/2] Fix deferred_split_isolate() and drop the split workaround Kiryl Shutsemau
2026-08-31 9:15 ` [PATCH v2 1/2] mm/huge_memory: do not touch frozen folios in deferred_split_isolate() Kiryl Shutsemau
2026-08-31 11:16 ` Lance Yang
2026-08-31 11:23 ` Usama Arif
2026-09-01 1:38 ` Baolin Wang
2026-09-07 13:48 ` David Hildenbrand (Arm)
2026-08-31 9:15 ` [PATCH v2 2/2] mm/huge_memory: dequeue the deferred split after the split freeze Kiryl Shutsemau
2026-08-31 11:34 ` Lance Yang
2026-09-07 13:55 ` David Hildenbrand (Arm)
2026-09-07 14:12 ` Kiryl Shutsemau
2026-09-07 14:30 ` David Hildenbrand (Arm)
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox