* [PATCH v2 1/3] mm/huge_memory: make thp_underused() work for mTHP folios
2026-09-16 22:44 [PATCH v2 0/3] mm: split underused anonymous mTHP folios Joanne Koong
@ 2026-09-16 22:44 ` Joanne Koong
2026-09-16 22:44 ` [PATCH v2 2/3] mm/huge_memory: don't queue folios that can never be underused Joanne Koong
` (3 subsequent siblings)
4 siblings, 0 replies; 22+ messages in thread
From: Joanne Koong @ 2026-09-16 22:44 UTC (permalink / raw)
To: akpm, david, ljs
Cc: usama.arif, hannes, baohua, alex, ziy, baolin.wang, liam, npache,
ryan.roberts, dev.jain, lance.yang, vbabka, rppt, surenb, mhocko,
willy, linux-mm
thp_underused() decides whether a large folio on the deferred split list
is underused and should be split so that its zero-filled subpages can be
reclaimed.
Today it only ever checks if the folio is underused on PMD-sized folios.
As such, thp_underused() assumes throughout that the folio has
HPAGE_PMD_NR pages. A later patch will be queueing anonymous mTHP folios
on the deferred split list as well, at which point this assumption will
not be correct anymore.
Generalize thp_underused() to be folio-size-aware instead of hardcoded
to PMD-sized folios. The check for whether a folio can ever be underused
is moved into its own function, thp_can_be_underused(), as this will get
reused in the next patch when determining if the folio should get added
to the deferred split queue.
khugepaged_max_ptes_none keeps its meaning as an absolute number of
pages, which is also how khugepaged applies it when collapsing to PMD
order. It is deliberately not scaled down per folio order (collapse
rejected proportional scaling of intermediate values because it either
lets the memory footprint creep or ends up confusing to reason about).
Read as an absolute count, this knob gains a useful second meaning for
mTHP as being the smallest folio size that takes part in underused
splitting at all.
There is no functional change, as for a PMD-sized folio nr_pages is
HPAGE_PMD_NR.
Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
---
mm/huge_memory.c | 31 +++++++++++++++++++++++++------
1 file changed, 25 insertions(+), 6 deletions(-)
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 1e5d68acf62a..c6ca2a541128 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -4541,6 +4541,23 @@ bool __folio_unqueue_deferred_split(struct folio *folio)
return unqueued; /* useful for debug warnings */
}
+static bool thp_can_be_underused(unsigned long nr_pages,
+ unsigned long max_ptes_none)
+{
+ /*
+ * The sysctl maximum means the user tolerates any number of zero-filled
+ * pages, so nothing is ever underused.
+ */
+ if (max_ptes_none == HPAGE_PMD_NR - 1)
+ return false;
+
+ /*
+ * A folio no larger than the number of zero-filled pages the user
+ * tolerates can never exceed it. It can never be underused.
+ */
+ return nr_pages > max_ptes_none;
+}
+
/* partially_mapped=false won't clear PG_partially_mapped folio flag */
void deferred_split_folio(struct folio *folio, bool partially_mapped)
{
@@ -4602,25 +4619,27 @@ static unsigned long deferred_split_count(struct shrinker *shrink,
static bool thp_underused(struct folio *folio)
{
- int num_zero_pages = 0, num_filled_pages = 0;
- int i;
+ const unsigned long max_ptes_none = khugepaged_max_ptes_none;
+ const unsigned long nr_pages = folio_nr_pages(folio);
+ unsigned long num_zero_pages = 0, num_filled_pages = 0;
+ unsigned long i;
- if (khugepaged_max_ptes_none == HPAGE_PMD_NR - 1)
+ if (!thp_can_be_underused(nr_pages, max_ptes_none))
return false;
if (folio_contain_hwpoisoned_page(folio))
return false;
- for (i = 0; i < folio_nr_pages(folio); i++) {
+ for (i = 0; i < nr_pages; i++) {
if (pages_identical(folio_page(folio, i), ZERO_PAGE(0))) {
- if (++num_zero_pages > khugepaged_max_ptes_none)
+ if (++num_zero_pages > max_ptes_none)
return true;
} else {
/*
* Another path for early exit once the number
* of non-zero filled pages exceeds threshold.
*/
- if (++num_filled_pages >= HPAGE_PMD_NR - khugepaged_max_ptes_none)
+ if (++num_filled_pages >= nr_pages - max_ptes_none)
return false;
}
}
--
2.52.0
^ permalink raw reply related [flat|nested] 22+ messages in thread* [PATCH v2 2/3] mm/huge_memory: don't queue folios that can never be underused
2026-09-16 22:44 [PATCH v2 0/3] mm: split underused anonymous mTHP folios Joanne Koong
2026-09-16 22:44 ` [PATCH v2 1/3] mm/huge_memory: make thp_underused() work for " Joanne Koong
@ 2026-09-16 22:44 ` Joanne Koong
2026-09-16 22:44 ` [PATCH v2 3/3] mm/memory: add anonymous mTHP folios to the deferred split list Joanne Koong
` (2 subsequent siblings)
4 siblings, 0 replies; 22+ messages in thread
From: Joanne Koong @ 2026-09-16 22:44 UTC (permalink / raw)
To: akpm, david, ljs
Cc: usama.arif, hannes, baohua, alex, ziy, baolin.wang, liam, npache,
ryan.roberts, dev.jain, lance.yang, vbabka, rppt, surenb, mhocko,
willy, linux-mm
Every anonymous PMD-sized folio is put on the deferred split queue when
it is first mapped, so that the shrinker can find it under memory
pressure and split it if it turns out to be mostly zero-filled.
With the default khugepaged/max_ptes_none this is wasted work. The
default is HPAGE_PMD_NR - 1, which tells thp_underused() that any number
of zero-filled pages is tolerable, which means the shrinker will never
split any of these folios for being underused. They take the list_lru
lock at fault time, inflate the object count the shrinker reports, and
are then walked and dropped when they're first scanned.
Skip the queuing for folios that thp_can_be_underused() says can never
qualify. At the default khugepaged/max_ptes_none that is all of them,
and once the knob is lowered only folios with more pages than it are
queued. This matters for a subsequent patch that adds anonymous mTHP
folios to the deferred split list, as it prevents small mTHP orders from
taking the list_lru lock on every anonymous fault.
Please note that the queue has always been best-effort. A folio that was
queued before khugepaged/max_ptes_none is lowered gets dropped from the
queue by the first scan that finds it's not underused, so changing the
sysctl has never retroactively applied to folios that were already
scanned. Requeueing eligible folios when the sysctl changes will be
addressed in a separate patch.
Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
---
mm/huge_memory.c | 14 ++++++++++++--
1 file changed, 12 insertions(+), 2 deletions(-)
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index c6ca2a541128..e3349f2314cb 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -4573,8 +4573,18 @@ void deferred_split_folio(struct folio *folio, bool partially_mapped)
if (folio_order(folio) <= 1)
return;
- if (!partially_mapped && !split_underused_thp)
- return;
+ if (!partially_mapped) {
+ if (!split_underused_thp)
+ return;
+ /*
+ * Nothing will ever split this folio for being underused, so
+ * keep it off the queue entirely rather than paying for the
+ * list_lru lock here and a shrinker scan later.
+ */
+ if (!thp_can_be_underused(folio_nr_pages(folio),
+ khugepaged_max_ptes_none))
+ return;
+ }
/*
* Exclude swapcache: originally to avoid a corrupt deferred split
--
2.52.0
^ permalink raw reply related [flat|nested] 22+ messages in thread* [PATCH v2 3/3] mm/memory: add anonymous mTHP folios to the deferred split list
2026-09-16 22:44 [PATCH v2 0/3] mm: split underused anonymous mTHP folios Joanne Koong
2026-09-16 22:44 ` [PATCH v2 1/3] mm/huge_memory: make thp_underused() work for " Joanne Koong
2026-09-16 22:44 ` [PATCH v2 2/3] mm/huge_memory: don't queue folios that can never be underused Joanne Koong
@ 2026-09-16 22:44 ` Joanne Koong
2026-09-20 16:20 ` [PATCH v2 0/3] mm: split underused anonymous mTHP folios Lance Yang
2026-09-21 10:23 ` David Hildenbrand (Arm)
4 siblings, 0 replies; 22+ messages in thread
From: Joanne Koong @ 2026-09-16 22:44 UTC (permalink / raw)
To: akpm, david, ljs
Cc: usama.arif, hannes, baohua, alex, ziy, baolin.wang, liam, npache,
ryan.roberts, dev.jain, lance.yang, vbabka, rppt, surenb, mhocko,
willy, linux-mm
Unlike for PMD-sized folios, an anonymous mTHP folio doesn't get added
to the deferred split list at fault or collapse time. As a result, a
fully mapped mTHP folio that is mostly zero-filled doesn't get split by
the deferred split shrinker when the system is under memory pressure.
At Meta we would like to deploy 2M THP=always on arm64 with 64k base
pages, as 2M gives the contpte benefits while the PMD size (512M) is too
big to use. Without underused splitting this causes memory regressions,
as the unused parts of those folios can never be broken down and
reclaimed.
Add anonymous mTHP folios to the deferred split list from
map_anon_folio_pte_nopf(), mirroring what map_anon_folio_pmd_nopf()
already does for PMD-sized folios. This covers both the fault path and
the khugepaged mTHP collapse path. If there is memory pressure, a
zero-filled mTHP can then be split with its zero pages remapped to the
shared zero page and reclaimed.
The preceding patch bounds what folios can get added to the deferred
split list. Nothing gets added at the default khugepaged/max_ptes_none,
and in cases where it is lowered, only folios with more pages than it
are added, so orders that could never be underused are left alone and
systems that enable only small mTHP orders are unaffected. For underused
splitting to happen, khugepaged/max_ptes_none has to be set below the
folio's page count.
To minimize overhead on the common order-0 fault path, the
deferred_split_folio() call is guarded by an inline folio_test_large()
check.
Suggested-by: Usama Arif <usama.arif@linux.dev>
Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
---
mm/memory.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/mm/memory.c b/mm/memory.c
index 8b0c2c735d3d..1fe76f72868d 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -5406,6 +5406,8 @@ void map_anon_folio_pte_nopf(struct folio *folio, pte_t *pte,
folio_add_lru_vma(folio, vma);
set_ptes(vma->vm_mm, addr, pte, entry, nr_pages);
update_mmu_cache_range(NULL, vma, addr, pte, nr_pages);
+ if (folio_test_large(folio))
+ deferred_split_folio(folio, false);
}
static void map_anon_folio_pte_pf(struct folio *folio, pte_t *pte,
--
2.52.0
^ permalink raw reply related [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-16 22:44 [PATCH v2 0/3] mm: split underused anonymous mTHP folios Joanne Koong
` (2 preceding siblings ...)
2026-09-16 22:44 ` [PATCH v2 3/3] mm/memory: add anonymous mTHP folios to the deferred split list Joanne Koong
@ 2026-09-20 16:20 ` Lance Yang
2026-09-21 10:04 ` Barry Song
` (2 more replies)
2026-09-21 10:23 ` David Hildenbrand (Arm)
4 siblings, 3 replies; 22+ messages in thread
From: Lance Yang @ 2026-09-20 16:20 UTC (permalink / raw)
To: joannelkoong
Cc: akpm, david, ljs, usama.arif, hannes, baohua, alex, ziy,
baolin.wang, liam, npache, ryan.roberts, dev.jain, lance.yang,
vbabka, rppt, surenb, mhocko, willy, linux-mm
On Wed, Sep 16, 2026 at 03:44:34PM -0700, Joanne Koong wrote:
>PMD-sized THPs that are mostly zero-filled are reclaimed under memory
>pressure by the deferred split shrinker, but this is not done for mTHP
>folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
>base pages, as 2M provides the contpte benefits while the PMD size there
>(512M) is too big to use. However, 2M THP=always causes memory regressions
>unless the unused portions of those folios can be broken down and
>reclaimed.
>
>v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
>it queued every anonymous mTHP folio. David objected to the scaling since
>collapse had already rejected proportional scaling as either letting the memory
>footprint creep or being confusing to reason about. David and Barry separately
>objected to the queuing, which adds list_lru lock contention on every anonymous
>fault, and hurts lower-order use cases such as Android's order-2-only
>configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
>simply not queuing folios that could never exceed it, which addresses both.
>
>Read as an absolute count, khugepaged_max_ptes_none also implies the
>smallest folio that takes part in underused splitting. A 2M folio on 64k
>base pages is 32 pages, so the knob has to be set below 32 for those to
>be queued at all, and any folio with no more pages than the value stays off
>the queue entirely.
Em ... I'm not sure max_ptes_none should also become the mTHP underused
threshold ...
Say we're on arm64 with 64K pages and set it to 16. Khugepaged PMD
collapse uses 16, while mTHP collapse turns the same value into 0. With
this series, an order-5 (2 MiB) mTHP enters the deferred split queue
because it has 32 pages, and is considered underused once more than 16
pages are zero-filled. So the same knob means 16, 0, and 16 depending on
where it is used ... That's a bit odd ... no?
Would a per-order value make more sense here?
An unset value could preserve the current behavior, while an explicit
value could control both mTHP collapse and underused splitting. That
would also give us a way to leave max_ptes_none as the legacy fallback
instead of adding more meaning to it :)
Cheers, Lance
>Patch 1 makes thp_underused() folio-size aware. Today it only ever sees
>PMD-sized folios, so it assumes HPAGE_PMD_NR pages throughout. Patch 3
>breaks that assumption, so patch 1 generalizes it first. No functional
>changes are introduced.
>
>Patch 2 keeps folios that can never be found underused off the deferred
>split queue, as suggested by Johannes. This changes/optimizes existing
>behavior. With the default khugepaged/max_ptes_none, PMD folios are queued
>today and then dropped again by the first scan without ever having been
>splittable. After this patch they are not queued at all.
>
>Patch 3 queues anonymous mTHP folios from map_anon_folio_pte_nopf(),
>mirroring what map_anon_folio_pmd_nopf() already does for PMD folios. This
>covers both the fault path and the khugepaged mTHP collapse path.
>
>One consequence of keeping the knob absolute is that it is shared with
>collapse. A value low enough to be useful for 2M mTHP is a tiny
>fraction of a 512M PMD, so khugepaged will only collapse to PMD order
>when the region is almost fully populated. That is fine for the deployment
>this series targets, which does not use PMD THP on arm64, but anyone wanting
>both PMD THP and 2M mTHP on 64k pages should be aware of it.
>
>Thanks,
>Joanne
>
>Changelog
>---------
>v1: https://lore.kernel.org/linux-mm/20260707201735.4113107-1-joannelkoong@gmail.com/
>
>Changes since v1:
>* Drop the per-order scaling of khugepaged_max_ptes_none. Keep it
> absolute, matching collapse (David, Johannes)
>* New patch 2 (suggested by Johannes): don't queue folios that can never
> be underused, which both bounds what patch 3 adds and stops queuing PMD
> folios under default settings (Johannes, David, Barry)
>* Fix the thp_underused() early exit to scale to the folio's own size
>
>Joanne Koong (3):
> mm/huge_memory: make thp_underused() work for mTHP folios
> mm/huge_memory: don't queue folios that can never be underused
> mm/memory: add anonymous mTHP folios to the deferred split list
>
> mm/huge_memory.c | 45 +++++++++++++++++++++++++++++++++++++--------
> mm/memory.c | 2 ++
> 2 files changed, 39 insertions(+), 8 deletions(-)
>
>--
>2.52.0
>
>
>
^ permalink raw reply [flat|nested] 22+ messages in thread* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-20 16:20 ` [PATCH v2 0/3] mm: split underused anonymous mTHP folios Lance Yang
@ 2026-09-21 10:04 ` Barry Song
2026-09-21 10:08 ` Usama Arif
2026-09-21 10:12 ` David Hildenbrand (Arm)
2 siblings, 0 replies; 22+ messages in thread
From: Barry Song @ 2026-09-21 10:04 UTC (permalink / raw)
To: Lance Yang
Cc: joannelkoong, akpm, david, ljs, usama.arif, hannes, alex, ziy,
baolin.wang, liam, npache, ryan.roberts, dev.jain, vbabka, rppt,
surenb, mhocko, willy, linux-mm
On Mon, Sep 21, 2026 at 12:20 AM Lance Yang <lance.yang@linux.dev> wrote:
>
>
> On Wed, Sep 16, 2026 at 03:44:34PM -0700, Joanne Koong wrote:
> >PMD-sized THPs that are mostly zero-filled are reclaimed under memory
> >pressure by the deferred split shrinker, but this is not done for mTHP
> >folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
> >base pages, as 2M provides the contpte benefits while the PMD size there
> >(512M) is too big to use. However, 2M THP=always causes memory regressions
> >unless the unused portions of those folios can be broken down and
> >reclaimed.
> >
> >v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
> >it queued every anonymous mTHP folio. David objected to the scaling since
> >collapse had already rejected proportional scaling as either letting the memory
> >footprint creep or being confusing to reason about. David and Barry separately
> >objected to the queuing, which adds list_lru lock contention on every anonymous
> >fault, and hurts lower-order use cases such as Android's order-2-only
> >configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
> >simply not queuing folios that could never exceed it, which addresses both.
> >
> >Read as an absolute count, khugepaged_max_ptes_none also implies the
> >smallest folio that takes part in underused splitting. A 2M folio on 64k
> >base pages is 32 pages, so the knob has to be set below 32 for those to
> >be queued at all, and any folio with no more pages than the value stays off
> >the queue entirely.
>
> Em ... I'm not sure max_ptes_none should also become the mTHP underused
> threshold ...
>
> Say we're on arm64 with 64K pages and set it to 16. Khugepaged PMD
> collapse uses 16, while mTHP collapse turns the same value into 0. With
> this series, an order-5 (2 MiB) mTHP enters the deferred split queue
> because it has 32 pages, and is considered underused once more than 16
> pages are zero-filled. So the same knob means 16, 0, and 16 depending on
> where it is used ... That's a bit odd ... no?
Lance, I actually had a long discussion with Usama, Joanne, and
Johannes:
https://lore.kernel.org/linux-mm/7bb32d25-2c3e-40ab-904a-ba22fff60cca@linux.dev/
https://lore.kernel.org/linux-mm/anHxPU17Ui4MTcCb@cmpxchg.org/
I feel this is correct though it is a bit hard to follow.
Best Regards
Barry
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-20 16:20 ` [PATCH v2 0/3] mm: split underused anonymous mTHP folios Lance Yang
2026-09-21 10:04 ` Barry Song
@ 2026-09-21 10:08 ` Usama Arif
2026-09-21 10:22 ` David Hildenbrand (Arm)
2026-09-22 10:33 ` Kiryl Shutsemau
2026-09-21 10:12 ` David Hildenbrand (Arm)
2 siblings, 2 replies; 22+ messages in thread
From: Usama Arif @ 2026-09-21 10:08 UTC (permalink / raw)
To: Lance Yang
Cc: Usama Arif, joannelkoong, akpm, david, ljs, hannes, baohua, alex,
ziy, baolin.wang, liam, npache, ryan.roberts, dev.jain, vbabka,
rppt, surenb, mhocko, willy, linux-mm, kas
On Mon, 21 Sep 2026 00:20:14 +0800 Lance Yang <lance.yang@linux.dev> wrote:
+Kiryl who is working on khugepaged rework
>
> On Wed, Sep 16, 2026 at 03:44:34PM -0700, Joanne Koong wrote:
> >PMD-sized THPs that are mostly zero-filled are reclaimed under memory
> >pressure by the deferred split shrinker, but this is not done for mTHP
> >folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
> >base pages, as 2M provides the contpte benefits while the PMD size there
> >(512M) is too big to use. However, 2M THP=always causes memory regressions
> >unless the unused portions of those folios can be broken down and
> >reclaimed.
> >
> >v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
> >it queued every anonymous mTHP folio. David objected to the scaling since
> >collapse had already rejected proportional scaling as either letting the memory
> >footprint creep or being confusing to reason about. David and Barry separately
> >objected to the queuing, which adds list_lru lock contention on every anonymous
> >fault, and hurts lower-order use cases such as Android's order-2-only
> >configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
> >simply not queuing folios that could never exceed it, which addresses both.
> >
> >Read as an absolute count, khugepaged_max_ptes_none also implies the
> >smallest folio that takes part in underused splitting. A 2M folio on 64k
> >base pages is 32 pages, so the knob has to be set below 32 for those to
> >be queued at all, and any folio with no more pages than the value stays off
> >the queue entirely.
>
> Em ... I'm not sure max_ptes_none should also become the mTHP underused
> threshold ...
>
> Say we're on arm64 with 64K pages and set it to 16. Khugepaged PMD
> collapse uses 16, while mTHP collapse turns the same value into 0. With
> this series, an order-5 (2 MiB) mTHP enters the deferred split queue
> because it has 32 pages, and is considered underused once more than 16
> pages are zero-filled. So the same knob means 16, 0, and 16 depending on
> where it is used ... That's a bit odd ... no?
>
I feel like the odd part of it is the 0 in the middle used by mTHP collapse,
not the 16 for checking zero-filled pages?
Kiryl, would your rework allow all values for mTHP collapse or would
it stil only allow HPAGE_PMD_NR = 1 and 0?
> Would a per-order value make more sense here?
>
I feel like per-order might be too many knobs? What about a percentage
that might apply to all mTHP orders?
> An unset value could preserve the current behavior, while an explicit
> value could control both mTHP collapse and underused splitting. That
Ah are you proposing that mTHP collapse can then take all values from
0 to mTHP number of pages?
I think one of the inital version of Nicos patches for mTHP collpase
allowed that (hopefully I am not misremembering). Nico, what was the
reason for only allowing 0 and max?
> would also give us a way to leave max_ptes_none as the legacy fallback
> instead of adding more meaning to it :)
>
> Cheers, Lance
[...]
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-21 10:08 ` Usama Arif
@ 2026-09-21 10:22 ` David Hildenbrand (Arm)
2026-09-22 10:33 ` Kiryl Shutsemau
1 sibling, 0 replies; 22+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-21 10:22 UTC (permalink / raw)
To: Usama Arif, Lance Yang
Cc: joannelkoong, akpm, ljs, hannes, baohua, alex, ziy, baolin.wang,
liam, npache, ryan.roberts, dev.jain, vbabka, rppt, surenb,
mhocko, willy, linux-mm, kas
> Kiryl, would your rework allow all values for mTHP collapse or would
> it stil only allow HPAGE_PMD_NR = 1 and 0?
>
>> Would a per-order value make more sense here?
>>
>
> I feel like per-order might be too many knobs? What about a percentage
> that might apply to all mTHP orders?
We discussed all that in the past, and concluded that we will have to support
the existing toggle. I played with using a percentage, but having two toggles
possibly mean the same thing (and some values not being possible) turned ugly.
So the conclusion for me was: no new toggles.
>
>> An unset value could preserve the current behavior, while an explicit
>> value could control both mTHP collapse and underused splitting. That
>
> Ah are you proposing that mTHP collapse can then take all values from
> 0 to mTHP number of pages?
> I think one of the inital version of Nicos patches for mTHP collpase
> allowed that (hopefully I am not misremembering). Nico, what was the
> reason for only allowing 0 and max?
The creep problem was discussed plenty of times. That's when we decided to start
with 0 and max, and only add support for scaling later, once we actually know
which use cases would need them. Without support for deferred shrinker, there
wasn't really the need to support other values.
Some of it can be had in:
commit 5a9c05d86683db5449b7fefd278c461cffb55234
Author: Nico Pache <nico.pache@linux.dev>
Date: Fri Jun 5 10:14:11 2026 -0600
mm/khugepaged: generalize __collapse_huge_page_* for mTHP support
generalize the order of the __collapse_huge_page_* and collapse_max_*
functions to support future mTHP collapse.
The current mechanism for determining collapse with the
khugepaged_max_ptes_none value is not designed with mTHP in mind. This
raises a key design issue: if we support user defined max_pte_none values
(even those scaled by order), a collapse of a lower order can introduces
an feedback loop, or "creep", when max_ptes_none is set to a value greater
than HPAGE_PMD_NR / 2. [1]
With this configuration, a successful collapse to order N will populate
enough pages to satisfy the collapse condition on order N+1 on the next
scan. This leads to unnecessary work and memory churn.
To fix this issue introduce a helper function that will limit mTHP
collapse support to two max_ptes_none values, 0 and HPAGE_PMD_NR - 1.
This effectively supports two modes: [2]
- max_ptes_none=0: never collapses if it encounters an empty PTE or a PTE
that maps the shared zeropage. Consequently, no memory bloat.
- max_ptes_none=511 (on 4k pagesz): Always collapse to the highest
available mTHP order.
This removes the possibility of "creep", and a warning will be emitted if
any non-supported max_ptes_none value is configured with mTHP enabled.
Any intermediate value will default mTHP collapse to max_ptes_none=0.
mTHP collapse will not honor the khugepaged_max_ptes_shared or
khugepaged_max_ptes_swap parameters, and will fail if it encounters a
shared or swapped entry.
Tracing back the discussion on that patch should give you all the details.
--
Cheers,
David
^ permalink raw reply [flat|nested] 22+ messages in thread* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-21 10:08 ` Usama Arif
2026-09-21 10:22 ` David Hildenbrand (Arm)
@ 2026-09-22 10:33 ` Kiryl Shutsemau
1 sibling, 0 replies; 22+ messages in thread
From: Kiryl Shutsemau @ 2026-09-22 10:33 UTC (permalink / raw)
To: Usama Arif
Cc: Lance Yang, joannelkoong, akpm, david, ljs, hannes, baohua, alex,
ziy, baolin.wang, liam, npache, ryan.roberts, dev.jain, vbabka,
rppt, surenb, mhocko, willy, linux-mm
On Mon, Sep 21, 2026 at 03:08:49AM -0700, Usama Arif wrote:
> On Mon, 21 Sep 2026 00:20:14 +0800 Lance Yang <lance.yang@linux.dev> wrote:
>
>
>
> +Kiryl who is working on khugepaged rework
>
> >
> > On Wed, Sep 16, 2026 at 03:44:34PM -0700, Joanne Koong wrote:
> > >PMD-sized THPs that are mostly zero-filled are reclaimed under memory
> > >pressure by the deferred split shrinker, but this is not done for mTHP
> > >folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
> > >base pages, as 2M provides the contpte benefits while the PMD size there
> > >(512M) is too big to use. However, 2M THP=always causes memory regressions
> > >unless the unused portions of those folios can be broken down and
> > >reclaimed.
> > >
> > >v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
> > >it queued every anonymous mTHP folio. David objected to the scaling since
> > >collapse had already rejected proportional scaling as either letting the memory
> > >footprint creep or being confusing to reason about. David and Barry separately
> > >objected to the queuing, which adds list_lru lock contention on every anonymous
> > >fault, and hurts lower-order use cases such as Android's order-2-only
> > >configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
> > >simply not queuing folios that could never exceed it, which addresses both.
> > >
> > >Read as an absolute count, khugepaged_max_ptes_none also implies the
> > >smallest folio that takes part in underused splitting. A 2M folio on 64k
> > >base pages is 32 pages, so the knob has to be set below 32 for those to
> > >be queued at all, and any folio with no more pages than the value stays off
> > >the queue entirely.
> >
> > Em ... I'm not sure max_ptes_none should also become the mTHP underused
> > threshold ...
> >
> > Say we're on arm64 with 64K pages and set it to 16. Khugepaged PMD
> > collapse uses 16, while mTHP collapse turns the same value into 0. With
> > this series, an order-5 (2 MiB) mTHP enters the deferred split queue
> > because it has 32 pages, and is considered underused once more than 16
> > pages are zero-filled. So the same knob means 16, 0, and 16 depending on
> > where it is used ... That's a bit odd ... no?
> >
>
> I feel like the odd part of it is the 0 in the middle used by mTHP collapse,
> not the 16 for checking zero-filled pages?
>
> Kiryl, would your rework allow all values for mTHP collapse or would
> it stil only allow HPAGE_PMD_NR = 1 and 0?
I didn't change the policy here.
What I considered doing it allow proportional treatment of max_ptes_none
for terminal mTHP size -- when there's no larger mTHP or PMD THP allowed
on the system. It can be helpful for 2M mTHPs on ARM machines with 64K
base page size and 512M PMDs.
--
Kiryl Shutsemau / Kirill A. Shutemov
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-20 16:20 ` [PATCH v2 0/3] mm: split underused anonymous mTHP folios Lance Yang
2026-09-21 10:04 ` Barry Song
2026-09-21 10:08 ` Usama Arif
@ 2026-09-21 10:12 ` David Hildenbrand (Arm)
2026-09-23 0:44 ` Joanne Koong
2 siblings, 1 reply; 22+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-21 10:12 UTC (permalink / raw)
To: Lance Yang, joannelkoong
Cc: akpm, ljs, usama.arif, hannes, baohua, alex, ziy, baolin.wang,
liam, npache, ryan.roberts, dev.jain, vbabka, rppt, surenb,
mhocko, willy, linux-mm
On 9/20/26 18:20, Lance Yang wrote:
>
> On Wed, Sep 16, 2026 at 03:44:34PM -0700, Joanne Koong wrote:
>> PMD-sized THPs that are mostly zero-filled are reclaimed under memory
>> pressure by the deferred split shrinker, but this is not done for mTHP
>> folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
>> base pages, as 2M provides the contpte benefits while the PMD size there
>> (512M) is too big to use. However, 2M THP=always causes memory regressions
>> unless the unused portions of those folios can be broken down and
>> reclaimed.
>>
>> v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
>> it queued every anonymous mTHP folio. David objected to the scaling since
"David objected to the scaling", not quite. The tricky bit is the scaling when
collapsing (creep).
When freeing, I'd assume scaling should be alright.
>> collapse had already rejected proportional scaling as either letting the memory
>> footprint creep or being confusing to reason about. David and Barry separately
>> objected to the queuing, which adds list_lru lock contention on every anonymous
>> fault, and hurts lower-order use cases such as Android's order-2-only
>> configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
>> simply not queuing folios that could never exceed it, which addresses both.
>>
>> Read as an absolute count, khugepaged_max_ptes_none also implies the
>> smallest folio that takes part in underused splitting. A 2M folio on 64k
>> base pages is 32 pages, so the knob has to be set below 32 for those to
>> be queued at all, and any folio with no more pages than the value stays off
>> the queue entirely.
>
> Em ... I'm not sure max_ptes_none should also become the mTHP underused
> threshold ...
>
> Say we're on arm64 with 64K pages and set it to 16. Khugepaged PMD
> collapse uses 16, while mTHP collapse turns the same value into 0. With
> this series, an order-5 (2 MiB) mTHP enters the deferred split queue
> because it has 32 pages, and is considered underused once more than 16
> pages are zero-filled. So the same knob means 16, 0, and 16 depending on
> where it is used ... That's a bit odd ... no?
I very much prefer scaling over using absolute numbers as proposed here.
>
> Would a per-order value make more sense here?
>
> An unset value could preserve the current behavior, while an explicit
> value could control both mTHP collapse and underused splitting. That
> would also give us a way to leave max_ptes_none as the legacy fallback
> instead of adding more meaning to it :)
The problem is that: we don't want any new toggles unless unavoidable.
--
Cheers,
David
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-21 10:12 ` David Hildenbrand (Arm)
@ 2026-09-23 0:44 ` Joanne Koong
2026-09-23 6:00 ` Barry Song
2026-09-23 9:44 ` David Hildenbrand (Arm)
0 siblings, 2 replies; 22+ messages in thread
From: Joanne Koong @ 2026-09-23 0:44 UTC (permalink / raw)
To: David Hildenbrand (Arm)
Cc: Lance Yang, akpm, ljs, usama.arif, hannes, baohua, alex, ziy,
baolin.wang, liam, npache, ryan.roberts, dev.jain, vbabka, rppt,
surenb, mhocko, willy, linux-mm
On Mon, Sep 21, 2026 at 3:12 AM David Hildenbrand (Arm)
<david@kernel.org> wrote:
>
> On 9/20/26 18:20, Lance Yang wrote:
> >
> > On Wed, Sep 16, 2026 at 03:44:34PM -0700, Joanne Koong wrote:
> >> PMD-sized THPs that are mostly zero-filled are reclaimed under memory
> >> pressure by the deferred split shrinker, but this is not done for mTHP
> >> folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
> >> base pages, as 2M provides the contpte benefits while the PMD size there
> >> (512M) is too big to use. However, 2M THP=always causes memory regressions
> >> unless the unused portions of those folios can be broken down and
> >> reclaimed.
> >>
> >> v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
> >> it queued every anonymous mTHP folio. David objected to the scaling since
>
> "David objected to the scaling", not quite. The tricky bit is the scaling when
> collapsing (creep).
>
> When freeing, I'd assume scaling should be alright.
Ah, apologies for misinterpreting your comment.
>
> >> collapse had already rejected proportional scaling as either letting the memory
> >> footprint creep or being confusing to reason about. David and Barry separately
> >> objected to the queuing, which adds list_lru lock contention on every anonymous
> >> fault, and hurts lower-order use cases such as Android's order-2-only
> >> configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
> >> simply not queuing folios that could never exceed it, which addresses both.
> >>
> >> Read as an absolute count, khugepaged_max_ptes_none also implies the
> >> smallest folio that takes part in underused splitting. A 2M folio on 64k
> >> base pages is 32 pages, so the knob has to be set below 32 for those to
> >> be queued at all, and any folio with no more pages than the value stays off
> >> the queue entirely.
> >
> > Em ... I'm not sure max_ptes_none should also become the mTHP underused
> > threshold ...
> >
> > Say we're on arm64 with 64K pages and set it to 16. Khugepaged PMD
> > collapse uses 16, while mTHP collapse turns the same value into 0. With
> > this series, an order-5 (2 MiB) mTHP enters the deferred split queue
> > because it has 32 pages, and is considered underused once more than 16
> > pages are zero-filled. So the same knob means 16, 0, and 16 depending on
> > where it is used ... That's a bit odd ... no?
>
> I very much prefer scaling over using absolute numbers as proposed here.
>
With scaling, I don't see a way of avoiding having to queue every
anonymous folio of order >=2 to the deferred split list at fault time
in the non-default max_ptes_none case, since any folio can now
potentially qualify as underused. This drew some objections in v1,
particularly for the Android small folios case. I read that as ruling
out scaling. The objections were a) the list_lru lock contention and
b) the limited upside vs cost of splitting small folios.
For a), I don't think there were any hard numbers showing this is a
problem. If it would be helpful to have some benchmark numbers, I'm
happy to run a synthetic one (patch 3 applied (which adds anon mTHP
folios to the deferred split list), only order-2 enabled,
max_ptes_none set low enough, all threads in one cgroup on one node so
they contend on the same lock, each thread faulting in a large private
anonymous mapping). I'm happy to run something else if anyone has a
workload they would trust more. Johannes also mentioned batching as a
mitigation. If the contention does turn out to be an issue, I'm happy
to prototype that and measure it too.
For b), v1 unconditionally queued every folio of order >= 2. v2 added
an optimization patch for skipping queueing entirely if max_ptes_none
is at its default value, which would also carry over if we went with
scaling. So I think with that patch added, for the objection to b), it
depends on what max_ptes_none Android actually runs. At the default
max_ptes_none, nothing would be queued at all, so a config that sees
no benefit from splitting small folios would pay nothing for it.
Does that make scaling acceptable to you and Barry? If so, I'd prefer
that as well because if/when intermediary values are supported for
mTHP collapse, that support would have to be proportional, and it'd be
easier to reconcile the two with split also scaled. If not, then what
would be the suggestion for where to take v3?
Happy to also discuss some of this at LPC in person, if that would be
more helpful.
> >
> > Would a per-order value make more sense here?
> >
> > An unset value could preserve the current behavior, while an explicit
> > value could control both mTHP collapse and underused splitting. That
> > would also give us a way to leave max_ptes_none as the legacy fallback
> > instead of adding more meaning to it :)
>
> The problem is that: we don't want any new toggles unless unavoidable.
Thanks for your feedback, Lance - fwiw, it was my preference too as it
seemed least confusing :)
But I get the hesitation about adding something that can't be deleted.
Thanks,
Joanne
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-23 0:44 ` Joanne Koong
@ 2026-09-23 6:00 ` Barry Song
2026-09-25 22:40 ` Joanne Koong
2026-09-23 9:44 ` David Hildenbrand (Arm)
1 sibling, 1 reply; 22+ messages in thread
From: Barry Song @ 2026-09-23 6:00 UTC (permalink / raw)
To: Joanne Koong
Cc: David Hildenbrand (Arm), Lance Yang, akpm, ljs, usama.arif,
hannes, alex, ziy, baolin.wang, liam, npache, ryan.roberts,
dev.jain, vbabka, rppt, surenb, mhocko, willy, linux-mm
On Wed, Sep 23, 2026 at 8:44 AM Joanne Koong <joannelkoong@gmail.com> wrote:
>
> On Mon, Sep 21, 2026 at 3:12 AM David Hildenbrand (Arm)
> <david@kernel.org> wrote:
> >
> > On 9/20/26 18:20, Lance Yang wrote:
> > >
> > > On Wed, Sep 16, 2026 at 03:44:34PM -0700, Joanne Koong wrote:
> > >> PMD-sized THPs that are mostly zero-filled are reclaimed under memory
> > >> pressure by the deferred split shrinker, but this is not done for mTHP
> > >> folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
> > >> base pages, as 2M provides the contpte benefits while the PMD size there
> > >> (512M) is too big to use. However, 2M THP=always causes memory regressions
> > >> unless the unused portions of those folios can be broken down and
> > >> reclaimed.
> > >>
> > >> v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
> > >> it queued every anonymous mTHP folio. David objected to the scaling since
> >
> > "David objected to the scaling", not quite. The tricky bit is the scaling when
> > collapsing (creep).
> >
> > When freeing, I'd assume scaling should be alright.
>
> Ah, apologies for misinterpreting your comment.
>
> >
> > >> collapse had already rejected proportional scaling as either letting the memory
> > >> footprint creep or being confusing to reason about. David and Barry separately
> > >> objected to the queuing, which adds list_lru lock contention on every anonymous
> > >> fault, and hurts lower-order use cases such as Android's order-2-only
> > >> configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
> > >> simply not queuing folios that could never exceed it, which addresses both.
> > >>
> > >> Read as an absolute count, khugepaged_max_ptes_none also implies the
> > >> smallest folio that takes part in underused splitting. A 2M folio on 64k
> > >> base pages is 32 pages, so the knob has to be set below 32 for those to
> > >> be queued at all, and any folio with no more pages than the value stays off
> > >> the queue entirely.
> > >
> > > Em ... I'm not sure max_ptes_none should also become the mTHP underused
> > > threshold ...
> > >
> > > Say we're on arm64 with 64K pages and set it to 16. Khugepaged PMD
> > > collapse uses 16, while mTHP collapse turns the same value into 0. With
> > > this series, an order-5 (2 MiB) mTHP enters the deferred split queue
> > > because it has 32 pages, and is considered underused once more than 16
> > > pages are zero-filled. So the same knob means 16, 0, and 16 depending on
> > > where it is used ... That's a bit odd ... no?
> >
> > I very much prefer scaling over using absolute numbers as proposed here.
> >
>
> With scaling, I don't see a way of avoiding having to queue every
> anonymous folio of order >=2 to the deferred split list at fault time
> in the non-default max_ptes_none case, since any folio can now
> potentially qualify as underused. This drew some objections in v1,
> particularly for the Android small folios case. I read that as ruling
> out scaling. The objections were a) the list_lru lock contention and
> b) the limited upside vs cost of splitting small folios.
>
> For a), I don't think there were any hard numbers showing this is a
> problem. If it would be helpful to have some benchmark numbers, I'm
> happy to run a synthetic one (patch 3 applied (which adds anon mTHP
> folios to the deferred split list), only order-2 enabled,
> max_ptes_none set low enough, all threads in one cgroup on one node so
> they contend on the same lock, each thread faulting in a large private
> anonymous mapping). I'm happy to run something else if anyone has a
> workload they would trust more. Johannes also mentioned batching as a
> mitigation. If the contention does turn out to be an issue, I'm happy
> to prototype that and measure it too.
Hi Joanne,
you may take a look at Shivank's report:
https://lore.kernel.org/linux-mm/08f63301a58f0e0c38480553295ad9af95928da6.camel@amd.com/
It's not the same lock, but the problem could be quite similar to the
`lru_cache` case.
The `lru_cache` work for mTHP order-2 not only resolves AMD's
regression after enabling order-2 mTHP, but also improves overall
performance. before that, there was a regression after enabling mTHP
due to lru lock contention:
"
Performance
===========
Base Patched delta
4K 109.60 109.11 -0.4%
mTHP-16K 144.10 99.00 -31.3%
mTHP-64K 104.75 104.25 -0.5%
mTHP-16K was ~31% slower than THP-never, due to lock contentions.
With your series, mTHP-16 performs ~9% faster than THP-never.
"
>
> For b), v1 unconditionally queued every folio of order >= 2. v2 added
> an optimization patch for skipping queueing entirely if max_ptes_none
> is at its default value, which would also carry over if we went with
> scaling. So I think with that patch added, for the objection to b), it
> depends on what max_ptes_none Android actually runs. At the default
> max_ptes_none, nothing would be queued at all, so a config that sees
> no benefit from splitting small folios would pay nothing for it.
I assume your case is that you always use 2MB mTHP and don't use
512MB PMD-sized THP at all, so you don't have the conflict Lance
mentioned? For Android, I don't think we would enable both 16KB mTHP
and THP either.
But thinking about it again, for a general-purpose system, I also find
it a bit strange if we have both mTHP and THP enabled.
>
> Does that make scaling acceptable to you and Barry? If so, I'd prefer
> that as well because if/when intermediary values are supported for
> mTHP collapse, that support would have to be proportional, and it'd be
> easier to reconcile the two with split also scaled. If not, then what
> would be the suggestion for where to take v3?
>
> Happy to also discuss some of this at LPC in person, if that would be
> more helpful.
Thanks
Barry
^ permalink raw reply [flat|nested] 22+ messages in thread* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-23 6:00 ` Barry Song
@ 2026-09-25 22:40 ` Joanne Koong
0 siblings, 0 replies; 22+ messages in thread
From: Joanne Koong @ 2026-09-25 22:40 UTC (permalink / raw)
To: Barry Song
Cc: David Hildenbrand (Arm), Lance Yang, akpm, ljs, usama.arif,
hannes, alex, ziy, baolin.wang, liam, npache, ryan.roberts,
dev.jain, vbabka, rppt, surenb, mhocko, willy, linux-mm
On Tue, Sep 22, 2026 at 11:00 PM Barry Song <baohua@kernel.org> wrote:
>
> On Wed, Sep 23, 2026 at 8:44 AM Joanne Koong <joannelkoong@gmail.com> wrote:
> >
> > On Mon, Sep 21, 2026 at 3:12 AM David Hildenbrand (Arm)
> > <david@kernel.org> wrote:
> > >
> > > On 9/20/26 18:20, Lance Yang wrote:
> > > >
> > > > On Wed, Sep 16, 2026 at 03:44:34PM -0700, Joanne Koong wrote:
> > > >> PMD-sized THPs that are mostly zero-filled are reclaimed under memory
> > > >> pressure by the deferred split shrinker, but this is not done for mTHP
> > > >> folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
> > > >> base pages, as 2M provides the contpte benefits while the PMD size there
> > > >> (512M) is too big to use. However, 2M THP=always causes memory regressions
> > > >> unless the unused portions of those folios can be broken down and
> > > >> reclaimed.
> > > >>
> > > >> v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
> > > >> it queued every anonymous mTHP folio. David objected to the scaling since
> > >
> > > "David objected to the scaling", not quite. The tricky bit is the scaling when
> > > collapsing (creep).
> > >
> > > When freeing, I'd assume scaling should be alright.
> >
> > Ah, apologies for misinterpreting your comment.
> >
> > >
> > > >> collapse had already rejected proportional scaling as either letting the memory
> > > >> footprint creep or being confusing to reason about. David and Barry separately
> > > >> objected to the queuing, which adds list_lru lock contention on every anonymous
> > > >> fault, and hurts lower-order use cases such as Android's order-2-only
> > > >> configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
> > > >> simply not queuing folios that could never exceed it, which addresses both.
> > > >>
> > > >> Read as an absolute count, khugepaged_max_ptes_none also implies the
> > > >> smallest folio that takes part in underused splitting. A 2M folio on 64k
> > > >> base pages is 32 pages, so the knob has to be set below 32 for those to
> > > >> be queued at all, and any folio with no more pages than the value stays off
> > > >> the queue entirely.
> > > >
> > > > Em ... I'm not sure max_ptes_none should also become the mTHP underused
> > > > threshold ...
> > > >
> > > > Say we're on arm64 with 64K pages and set it to 16. Khugepaged PMD
> > > > collapse uses 16, while mTHP collapse turns the same value into 0. With
> > > > this series, an order-5 (2 MiB) mTHP enters the deferred split queue
> > > > because it has 32 pages, and is considered underused once more than 16
> > > > pages are zero-filled. So the same knob means 16, 0, and 16 depending on
> > > > where it is used ... That's a bit odd ... no?
> > >
> > > I very much prefer scaling over using absolute numbers as proposed here.
> > >
> >
> > With scaling, I don't see a way of avoiding having to queue every
> > anonymous folio of order >=2 to the deferred split list at fault time
> > in the non-default max_ptes_none case, since any folio can now
> > potentially qualify as underused. This drew some objections in v1,
> > particularly for the Android small folios case. I read that as ruling
> > out scaling. The objections were a) the list_lru lock contention and
> > b) the limited upside vs cost of splitting small folios.
> >
> > For a), I don't think there were any hard numbers showing this is a
> > problem. If it would be helpful to have some benchmark numbers, I'm
> > happy to run a synthetic one (patch 3 applied (which adds anon mTHP
> > folios to the deferred split list), only order-2 enabled,
> > max_ptes_none set low enough, all threads in one cgroup on one node so
> > they contend on the same lock, each thread faulting in a large private
> > anonymous mapping). I'm happy to run something else if anyone has a
> > workload they would trust more. Johannes also mentioned batching as a
> > mitigation. If the contention does turn out to be an issue, I'm happy
> > to prototype that and measure it too.
>
> Hi Joanne,
>
> you may take a look at Shivank's report:
> https://lore.kernel.org/linux-mm/08f63301a58f0e0c38480553295ad9af95928da6.camel@amd.com/
>
> It's not the same lock, but the problem could be quite similar to the
> `lru_cache` case.
>
> The `lru_cache` work for mTHP order-2 not only resolves AMD's
> regression after enabling order-2 mTHP, but also improves overall
> performance. before that, there was a regression after enabling mTHP
> due to lru lock contention:
>
> "
> Performance
> ===========
>
> Base Patched delta
> 4K 109.60 109.11 -0.4%
> mTHP-16K 144.10 99.00 -31.3%
> mTHP-64K 104.75 104.25 -0.5%
>
> mTHP-16K was ~31% slower than THP-never, due to lock contentions.
> With your series, mTHP-16 performs ~9% faster than THP-never.
> "
Thanks for the link, Barry! Apologies for the late reply, I was out on PTO.
I ran your microbenchmark on my machine to see if I could reproduce
the results above and to see at what mTHP granularity it stops
degrading. My setup is a 2-socket Skylake, 40 cores / 80 threads,
threads free to run on all CPUs with memory bound to node 0 so they
all contend on one lruvec, running 80 threads x 16 MB x 200 loops.
These are the numbers I saw on the base (unpatched) version from
running perf lock contention on folio_lruvec_lock_irqsave ("lock wait"
is all 80 threads summed):
size contentions lock wait folios/MB
4K 1075221 321.4 s 256
mTHP-16K 4203612 392.8 s 64
mTHP-32K 2064073 190.5 s 32
mTHP-64K 212012 1.7 s 16
mTHP-128K 117993 1.8 s 8
mTHP-256K 64532 4.2 s 4
mTHP-512K 32485 3.1 s 2
mTHP-1024K 4313 0.3 s 1
mTHP-2048K 725 0.1 s 0.5
I'm pretty much seeing the same results you mentioned above, where
mTHP-16k results in more lock contentions (though I didn't see a
runtime regression. I saw that 16k comes out a bit faster than 4k in
wall time), but 64k results in significantly less. For mTHP-32k, I see
less time overall spent waiting, but 2x the lock contention.
I re-ran the benchmark with the 3 patches in this series applied, with
the khugepaged/max_ptes_none sysfs value set to 0 so every folio gets
queued (the worst case) and measured the queueing cost by switching
shrink_underused on/off. I saw:
size runtime delta between on/off folios/MB
mTHP-16K +15.8% 64
mTHP-32K +11.5% 32
mTHP-64K +1.1% 16
mTHP-128K +1.2% 8
mTHP-256K -0.6% 4
mTHP-512K +0.5% 2
mTHP-1024K -1.2% 1
mTHP-2048K -0.3% 0.5
which mirrors the same pattern/trend as above (eg regresses at the 16k
and 32k cases, but not for 64k+).
I ran a perf analysis and the delta doesn't come from the deferred
split list's own lock, but actually from folio_lruvec_lock_irqsave:
shrink_underused 0 1
runtime 6.00s 6.86s +14.3%
lruvec contentions 4206376 4234324 +0.7%
lruvec wait 406.4 s 485.8 s +19.5%
avg wait 96.5us 114.5us
list_lru_lock absent absent
The lruvec contentions count doesn't move much (0.7%) but the time
spent waiting goes up 19.5%. I think that's coming from the free path
rather than the fault path. In folios_put_refs(),
__page_cache_release() keeps the lruvec locked across the whole batch,
and from what I'm seeing, folio_unqueue_deferred_split() gets called
from inside that loop. I don't think your lru cache series helps this
directly since it batches on the add side but this cost is coming from
the free side inside a section that's already batched, but I think it
should help indirectly, since there would be less lruvec pressure /
contention overall. I'm not 100% sure about this, but think it might
be possible to move the unqueue from the deferred list to outside the
lruvec lock hold, though I need to look more into this.
For now though, I think it makes sense to have a hard-coded minimum
folio size for the underused shrinker, as you and David suggested.
Based on the results above, I think having 64k as the cut-off makes
sense.
>
> >
> > For b), v1 unconditionally queued every folio of order >= 2. v2 added
> > an optimization patch for skipping queueing entirely if max_ptes_none
> > is at its default value, which would also carry over if we went with
> > scaling. So I think with that patch added, for the objection to b), it
> > depends on what max_ptes_none Android actually runs. At the default
> > max_ptes_none, nothing would be queued at all, so a config that sees
> > no benefit from splitting small folios would pay nothing for it.
>
> I assume your case is that you always use 2MB mTHP and don't use
> 512MB PMD-sized THP at all, so you don't have the conflict Lance
> mentioned? For Android, I don't think we would enable both 16KB mTHP
> and THP either.
Thanks for the context on the Android case. For Meta's case, that is
correct, we are deploying 2M mTHP = always on ARM with 64k base pages,
with the 512M PMD size disabled.
Thanks,
Joanne
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-23 0:44 ` Joanne Koong
2026-09-23 6:00 ` Barry Song
@ 2026-09-23 9:44 ` David Hildenbrand (Arm)
2026-09-25 23:47 ` Joanne Koong
1 sibling, 1 reply; 22+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-23 9:44 UTC (permalink / raw)
To: Joanne Koong
Cc: Lance Yang, akpm, ljs, usama.arif, hannes, baohua, alex, ziy,
baolin.wang, liam, npache, ryan.roberts, dev.jain, vbabka, rppt,
surenb, mhocko, willy, linux-mm
On 9/23/26 02:44, Joanne Koong wrote:
> On Mon, Sep 21, 2026 at 3:12 AM David Hildenbrand (Arm)
> <david@kernel.org> wrote:
>>
>> On 9/20/26 18:20, Lance Yang wrote:
>>>
>>
>> "David objected to the scaling", not quite. The tricky bit is the scaling when
>> collapsing (creep).
>>
>> When freeing, I'd assume scaling should be alright.
>
> Ah, apologies for misinterpreting your comment.
>
>>
>>>
>>> Em ... I'm not sure max_ptes_none should also become the mTHP underused
>>> threshold ...
>>>
>>> Say we're on arm64 with 64K pages and set it to 16. Khugepaged PMD
>>> collapse uses 16, while mTHP collapse turns the same value into 0. With
>>> this series, an order-5 (2 MiB) mTHP enters the deferred split queue
>>> because it has 32 pages, and is considered underused once more than 16
>>> pages are zero-filled. So the same knob means 16, 0, and 16 depending on
>>> where it is used ... That's a bit odd ... no?
>>
>> I very much prefer scaling over using absolute numbers as proposed here.
>>
>
> With scaling, I don't see a way of avoiding having to queue every
> anonymous folio of order >=2 to the deferred split list at fault time
> in the non-default max_ptes_none case, since any folio can now
> potentially qualify as underused.
Exactly.
I was primarily arguing that I don't want this overhead for the majority of
Linux installations out there that ship with
$ cat /sys/kernel/mm/transparent_hugepage/khugepaged/max_ptes_none
511
IOW: we add them (PMD THP folios today!) to the deferred list for
uuderused-scanning but never actually scan them because of:
if (khugepaged_max_ptes_none == HPAGE_PMD_NR - 1)
return false;
Now, what you are saying is we might want less overhead for installations that
modify max_ptes_none, like Meta's fleet. See below.
> This drew some objections in v1,
> particularly for the Android small folios case. I read that as ruling
> out scaling. The objections were a) the list_lru lock contention and
It was the "deferred shrinker list and lock" contention, yes.
> b) the limited upside vs cost of splitting small folios.
>
> For a), I don't think there were any hard numbers showing this is a
> problem. If it would be helpful to have some benchmark numbers, I'm
> happy to run a synthetic one (patch 3 applied (which adds anon mTHP
> folios to the deferred split list), only order-2 enabled,
> max_ptes_none set low enough, all threads in one cgroup on one node so
> they contend on the same lock, each thread faulting in a large private
> anonymous mapping). I'm happy to run something else if anyone has a
> workload they would trust more. Johannes also mentioned batching as a
> mitigation. If the contention does turn out to be an issue, I'm happy
> to prototype that and measure it too.
Right, I raised that. Johannes thinks it could be fixed with batching. I was
concerned that it would become rather ugly. We'd have to see how that could look.
>
> For b), v1 unconditionally queued every folio of order >= 2. v2 added
> an optimization patch for skipping queueing entirely if max_ptes_none
> is at its default value, which would also carry over if we went with
> scaling. So I think with that patch added, for the objection to b), it
> depends on what max_ptes_none Android actually runs. At the default
> max_ptes_none, nothing would be queued at all, so a config that sees
> no benefit from splitting small folios would pay nothing for it.
I would assume everybody except some special users that fine-tune underused
shrinker run with the default. I did some digging in the past regarding public
documentation about how to set the parameter, and most of them just defaulted to
511.
I would be surprised if Android even knows about this parameter ;)
>
> Does that make scaling acceptable to you and Barry? If so, I'd prefer
> that as well because if/when intermediary values are supported for
> mTHP collapse, that support would have to be proportional, and it'd be
> easier to reconcile the two with split also scaled. If not, then what
> would be the suggestion for where to take v3?
My opinion is (open for discussion :) ) that scaling is likely the better
approach. With the following notes:
(1) If we want to enable mTHP collapse with scaling as well, this needs very
good documentation and also another thought on how to work around the problem of
creep (e.g., refusing to collapse if creep would be possible according to the
max_ptes_non setting and warning).
(2) Systems where the underused shrinker is effectively inactive should not add
folios to the deferred list. Including PMD THPs, which we unconditionally add today.
(3) If we want to exclude certain small folio sizes from the underused shrinker
(e.g., order-2? order-3?) it might be better to just hard-code that in the
kernel instead of giving the admin a choice it cannot possibly make easily.
--
Cheers,
David
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-23 9:44 ` David Hildenbrand (Arm)
@ 2026-09-25 23:47 ` Joanne Koong
2026-09-28 8:03 ` Barry Song
2026-09-28 19:20 ` David Hildenbrand (Arm)
0 siblings, 2 replies; 22+ messages in thread
From: Joanne Koong @ 2026-09-25 23:47 UTC (permalink / raw)
To: David Hildenbrand (Arm)
Cc: Lance Yang, akpm, ljs, usama.arif, hannes, baohua, alex, ziy,
baolin.wang, liam, npache, ryan.roberts, dev.jain, vbabka, rppt,
surenb, mhocko, willy, linux-mm
On Wed, Sep 23, 2026 at 2:44 AM David Hildenbrand (Arm)
<david@kernel.org> wrote:
>
> On 9/23/26 02:44, Joanne Koong wrote:
> > On Mon, Sep 21, 2026 at 3:12 AM David Hildenbrand (Arm)
> > <david@kernel.org> wrote:
> >>
> >> On 9/20/26 18:20, Lance Yang wrote:
> >>>
> >>
> >> "David objected to the scaling", not quite. The tricky bit is the scaling when
> >> collapsing (creep).
> >>
> >> When freeing, I'd assume scaling should be alright.
> >
> > Ah, apologies for misinterpreting your comment.
> >
> >>
> >>>
> >>> Em ... I'm not sure max_ptes_none should also become the mTHP underused
> >>> threshold ...
> >>>
> >>> Say we're on arm64 with 64K pages and set it to 16. Khugepaged PMD
> >>> collapse uses 16, while mTHP collapse turns the same value into 0. With
> >>> this series, an order-5 (2 MiB) mTHP enters the deferred split queue
> >>> because it has 32 pages, and is considered underused once more than 16
> >>> pages are zero-filled. So the same knob means 16, 0, and 16 depending on
> >>> where it is used ... That's a bit odd ... no?
> >>
> >> I very much prefer scaling over using absolute numbers as proposed here.
> >>
> >
> > With scaling, I don't see a way of avoiding having to queue every
> > anonymous folio of order >=2 to the deferred split list at fault time
> > in the non-default max_ptes_none case, since any folio can now
> > potentially qualify as underused.
>
> Exactly.
>
> I was primarily arguing that I don't want this overhead for the majority of
> Linux installations out there that ship with
>
> $ cat /sys/kernel/mm/transparent_hugepage/khugepaged/max_ptes_none
> 511
>
> IOW: we add them (PMD THP folios today!) to the deferred list for
> uuderused-scanning but never actually scan them because of:
>
> if (khugepaged_max_ptes_none == HPAGE_PMD_NR - 1)
> return false;
Ah, I missed this nuance in your v1 reply - I thought it was objecting
to the queueing overhead for the non-default max_ptes_none case as
well. Patch 2 in this series does what you mentioned above and it
automatically applies to PMD THPs too. I'll add your Suggested-by: tag
to this in v3.
>
> Now, what you are saying is we might want less overhead for installations that
> modify max_ptes_none, like Meta's fleet. See below.
>
> > This drew some objections in v1,
> > particularly for the Android small folios case. I read that as ruling
> > out scaling. The objections were a) the list_lru lock contention and
>
> It was the "deferred shrinker list and lock" contention, yes.
>
> > b) the limited upside vs cost of splitting small folios.
> >
> > For a), I don't think there were any hard numbers showing this is a
> > problem. If it would be helpful to have some benchmark numbers, I'm
> > happy to run a synthetic one (patch 3 applied (which adds anon mTHP
> > folios to the deferred split list), only order-2 enabled,
> > max_ptes_none set low enough, all threads in one cgroup on one node so
> > they contend on the same lock, each thread faulting in a large private
> > anonymous mapping). I'm happy to run something else if anyone has a
> > workload they would trust more. Johannes also mentioned batching as a
> > mitigation. If the contention does turn out to be an issue, I'm happy
> > to prototype that and measure it too.
>
> Right, I raised that. Johannes thinks it could be fixed with batching. I was
> concerned that it would become rather ugly. We'd have to see how that could look.
I ran Barry's microbenchmark to get a sense of the contention (more of
the details are in my previous reply to Barry [1]). With max_ptes_none
set to 0 so that everything gets queued, and toggling shrink_underused
0 vs 1 with patch 3 applied, I saw roughly a 15.8% and 11.5%
performance hit on runtime at 16k and 32k, but only 1.1% at 64k and
noisy above that.
I didn't see the cost being from contention on the deferred list's
lock though. Perf showed that it was coming from
folio_lruvec_lock_irqsave, and I think that's from the free path where
in folios_put_refs(), __page_cache_release() keeps the lruvec locked
across the whole batch, and folio_unqueue_deferred_split() gets called
from inside that loop with that lruvec lock held. I don't think
batching the enqueues would help in that case, since the cost is
coming from the unqueue side and on a path that's already batched. I
think maybe we could move the unqueue out where the lruvec lock
doesn't need to be held while the unqueueing happens, but I need to
look more into that.
>
> >
> > For b), v1 unconditionally queued every folio of order >= 2. v2 added
> > an optimization patch for skipping queueing entirely if max_ptes_none
> > is at its default value, which would also carry over if we went with
> > scaling. So I think with that patch added, for the objection to b), it
> > depends on what max_ptes_none Android actually runs. At the default
> > max_ptes_none, nothing would be queued at all, so a config that sees
> > no benefit from splitting small folios would pay nothing for it.
>
> I would assume everybody except some special users that fine-tune underused
> shrinker run with the default. I did some digging in the past regarding public
> documentation about how to set the parameter, and most of them just defaulted to
> 511.
>
> I would be surprised if Android even knows about this parameter ;)
>
> >
> > Does that make scaling acceptable to you and Barry? If so, I'd prefer
> > that as well because if/when intermediary values are supported for
> > mTHP collapse, that support would have to be proportional, and it'd be
> > easier to reconcile the two with split also scaled. If not, then what
> > would be the suggestion for where to take v3?
> My opinion is (open for discussion :) ) that scaling is likely the better
> approach. With the following notes:
>
> (1) If we want to enable mTHP collapse with scaling as well, this needs very
> good documentation and also another thought on how to work around the problem of
> creep (e.g., refusing to collapse if creep would be possible according to the
> max_ptes_non setting and warning).
Agreed. And if/when collapse gains proportional support later on,
having split already scaled means one value governing both instead of
two.
>
> (2) Systems where the underused shrinker is effectively inactive should not add
> folios to the deferred list. Including PMD THPs, which we unconditionally add today.
Agreed.
>
> (3) If we want to exclude certain small folio sizes from the underused shrinker
> (e.g., order-2? order-3?) it might be better to just hard-code that in the
> kernel instead of giving the admin a choice it cannot possibly make easily.
Based off the benchmark results in [1] and [2], I think the cutoff
needs to be at 64k where we exclude anything under that. On 4k base
pages, this is order 4, but on 64k base pages, order 4 would mean 1M,
which I think would be overly conservative. I think it makes more
sense to express it as an absolute threshold value instead of by
order. I'll try to find an arm64 machine to double-check this on.
For v3, I'll go back to the scaling approach. I'll be traveling next
week and then will be at LPC after that, so my timeline is to submit
v3 after I come back from LPC.
Thanks,
Joanne
[1] https://lore.kernel.org/linux-mm/CAJnrk1a2PpUGTsCgS7jct2KET4K2zAL+_h+te8G1MaU9bXs99w@mail.gmail.com/
[2] https://lore.kernel.org/linux-mm/CAGsJ_4x_Q1Ku9r6zfC-K-o87juQ9JNQ8=OSNp+UZzW3MPtZXbQ@mail.gmail.com/
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-25 23:47 ` Joanne Koong
@ 2026-09-28 8:03 ` Barry Song
2026-10-01 9:10 ` Joanne Koong
2026-09-28 19:20 ` David Hildenbrand (Arm)
1 sibling, 1 reply; 22+ messages in thread
From: Barry Song @ 2026-09-28 8:03 UTC (permalink / raw)
To: Joanne Koong
Cc: David Hildenbrand (Arm), Lance Yang, akpm, ljs, usama.arif,
hannes, alex, ziy, baolin.wang, liam, npache, ryan.roberts,
dev.jain, vbabka, rppt, surenb, mhocko, willy, linux-mm
On Sat, Sep 26, 2026 at 7:47 AM Joanne Koong <joannelkoong@gmail.com> wrote:
[...]
>
> Based off the benchmark results in [1] and [2], I think the cutoff
> needs to be at 64k where we exclude anything under that. On 4k base
> pages, this is order 4, but on 64k base pages, order 4 would mean 1M,
> which I think would be overly conservative. I think it makes more
> sense to express it as an absolute threshold value instead of by
> order. I'll try to find an arm64 machine to double-check this on.
Hi Joanne,
I once recommended a hard-coded value such as 2 MB[1]. Would that work?
Maybe 1 MB or 512 KB would also be fine. I’m still worried that 64 KB
is too small, especially on machines with 64 KB or 16 KB page sizes.
Also, we could have more contention between lru_lock and the deferred
split lock, since, as you mentioned, folio_unqueue_deferred_split()
is called under lru_lock in many cases.
smaller folio sizes make mTHP collapse and the underused shrinker
more likely to ping-pong?
Just my two cents. I'd also like to hear what David thinks.
[1] https://lore.kernel.org/linux-mm/CAGsJ_4x86EYsMz_n2NKDwTqeiv2pi4AcZb1RuZoBeUBR1iFqOA@mail.gmail.com/
Best Regards
Barry
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-28 8:03 ` Barry Song
@ 2026-10-01 9:10 ` Joanne Koong
0 siblings, 0 replies; 22+ messages in thread
From: Joanne Koong @ 2026-10-01 9:10 UTC (permalink / raw)
To: Barry Song
Cc: David Hildenbrand (Arm), Lance Yang, akpm, ljs, usama.arif,
hannes, alex, ziy, baolin.wang, liam, npache, ryan.roberts,
dev.jain, vbabka, rppt, surenb, mhocko, willy, linux-mm
On Mon, Sep 28, 2026 at 9:03 AM Barry Song <baohua@kernel.org> wrote:
>
> On Sat, Sep 26, 2026 at 7:47 AM Joanne Koong <joannelkoong@gmail.com> wrote:
> [...]
> >
> > Based off the benchmark results in [1] and [2], I think the cutoff
> > needs to be at 64k where we exclude anything under that. On 4k base
> > pages, this is order 4, but on 64k base pages, order 4 would mean 1M,
> > which I think would be overly conservative. I think it makes more
> > sense to express it as an absolute threshold value instead of by
> > order. I'll try to find an arm64 machine to double-check this on.
>
> Hi Joanne,
Hi Barry,
>
> I once recommended a hard-coded value such as 2 MB[1]. Would that work?
>
> Maybe 1 MB or 512 KB would also be fine. I’m still worried that 64 KB
> is too small, especially on machines with 64 KB or 16 KB page sizes.
>
> Also, we could have more contention between lru_lock and the deferred
> split lock, since, as you mentioned, folio_unqueue_deferred_split()
> is called under lru_lock in many cases.
For my use case (64k base pages on arm64 with 2M anon THP set to
always), the >= 2MB threshold works. My one reservation is that on 4k
base pages this is effectively a no-op since 2M is the PMD size there
and those folios are already queued today. I agree with you though
that it's better to err conservative. I'm happy to go with 2M for now
and revisit lowering it later.
>
> smaller folio sizes make mTHP collapse and the underused shrinker
> more likely to ping-pong?
If collapse and split use the same scaled value, I think they're
self-consistent at a given order. Right now there's no ping-ponging
possible because mTHP collapse supports only 0 or HPAGE_PMD_NR - 1,
but I'm keeping an eye out for Kiryl's patches that'll change that, to
make sure that stays true.
Thanks,
Joanne
>
> Just my two cents. I'd also like to hear what David thinks.
>
> [1] https://lore.kernel.org/linux-mm/CAGsJ_4x86EYsMz_n2NKDwTqeiv2pi4AcZb1RuZoBeUBR1iFqOA@mail.gmail.com/
>
> Best Regards
> Barry
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-25 23:47 ` Joanne Koong
2026-09-28 8:03 ` Barry Song
@ 2026-09-28 19:20 ` David Hildenbrand (Arm)
2026-10-01 9:59 ` Joanne Koong
1 sibling, 1 reply; 22+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-28 19:20 UTC (permalink / raw)
To: Joanne Koong
Cc: Lance Yang, akpm, ljs, usama.arif, hannes, baohua, alex, ziy,
baolin.wang, liam, npache, ryan.roberts, dev.jain, vbabka, rppt,
surenb, mhocko, willy, linux-mm
On 9/26/26 01:47, Joanne Koong wrote:
> On Wed, Sep 23, 2026 at 2:44 AM David Hildenbrand (Arm)
> <david@kernel.org> wrote:
>>
>> On 9/23/26 02:44, Joanne Koong wrote:
>>> On Mon, Sep 21, 2026 at 3:12 AM David Hildenbrand (Arm)
>>> <david@kernel.org> wrote:
>>>
>>> Ah, apologies for misinterpreting your comment.
>>>
>>>
>>> With scaling, I don't see a way of avoiding having to queue every
>>> anonymous folio of order >=2 to the deferred split list at fault time
>>> in the non-default max_ptes_none case, since any folio can now
>>> potentially qualify as underused.
>>
>> Exactly.
>>
>> I was primarily arguing that I don't want this overhead for the majority of
>> Linux installations out there that ship with
>>
>> $ cat /sys/kernel/mm/transparent_hugepage/khugepaged/max_ptes_none
>> 511
>>
>> IOW: we add them (PMD THP folios today!) to the deferred list for
>> uuderused-scanning but never actually scan them because of:
>>
>> if (khugepaged_max_ptes_none == HPAGE_PMD_NR - 1)
>> return false;
>
> Ah, I missed this nuance in your v1 reply - I thought it was objecting
> to the queueing overhead for the non-default max_ptes_none case as
> well. Patch 2 in this series does what you mentioned above and it
> automatically applies to PMD THPs too. I'll add your Suggested-by: tag
> to this in v3.
I though Johannes suggested that :)
>> Right, I raised that. Johannes thinks it could be fixed with batching. I was
>> concerned that it would become rather ugly. We'd have to see how that could look.
>
> I ran Barry's microbenchmark to get a sense of the contention (more of
> the details are in my previous reply to Barry [1]). With max_ptes_none
> set to 0 so that everything gets queued, and toggling shrink_underused
> 0 vs 1 with patch 3 applied, I saw roughly a 15.8% and 11.5%
> performance hit on runtime at 16k and 32k, but only 1.1% at 64k and
> noisy above that.
>
> I didn't see the cost being from contention on the deferred list's
> lock though. Perf showed that it was coming from
> folio_lruvec_lock_irqsave, and I think that's from the free path where
> in folios_put_refs(), __page_cache_release() keeps the lruvec locked
> across the whole batch, and folio_unqueue_deferred_split() gets called
> from inside that loop with that lruvec lock held. I don't think
> batching the enqueues would help in that case, since the cost is
> coming from the unqueue side and on a path that's already batched. I
> think maybe we could move the unqueue out where the lruvec lock
> doesn't need to be held while the unqueueing happens, but I need to
> look more into that.
That's an interesting insight. If contention is less of a problem right now with
that reproducer, even better!
I prefer us not doing unnecessary work (adding folios to the deferred split
queue) when we won't really split them. (if it's contention or some other
overhead doesn't really matter)
>>>
>>> Does that make scaling acceptable to you and Barry? If so, I'd prefer
>>> that as well because if/when intermediary values are supported for
>>> mTHP collapse, that support would have to be proportional, and it'd be
>>> easier to reconcile the two with split also scaled. If not, then what
>>> would be the suggestion for where to take v3?
>> My opinion is (open for discussion :) ) that scaling is likely the better
>> approach. With the following notes:
>>
>> (1) If we want to enable mTHP collapse with scaling as well, this needs very
>> good documentation and also another thought on how to work around the problem of
>> creep (e.g., refusing to collapse if creep would be possible according to the
>> max_ptes_non setting and warning).
>
> Agreed. And if/when collapse gains proportional support later on,
> having split already scaled means one value governing both instead of
> two.
>
>>
>> (2) Systems where the underused shrinker is effectively inactive should not add
>> folios to the deferred list. Including PMD THPs, which we unconditionally add today.
>
> Agreed.
>
>>
>> (3) If we want to exclude certain small folio sizes from the underused shrinker
>> (e.g., order-2? order-3?) it might be better to just hard-code that in the
>> kernel instead of giving the admin a choice it cannot possibly make easily.
>
> Based off the benchmark results in [1] and [2], I think the cutoff
> needs to be at 64k where we exclude anything under that. On 4k base
> pages, this is order 4, but on 64k base pages, order 4 would mean 1M,
> which I think would be overly conservative. I think it makes more
> sense to express it as an absolute threshold value instead of by
> order. I'll try to find an arm64 machine to double-check this on.
If we don't add folios to deferred split queue if current underused shrinker
config wouldn't ever split them, I think there is no overhead for Android in any
case (so Barry wouldn't have to worry, at least for now).
So the magic number we chose only applies if the underused shrinker is enabled
and could eventually split them.
There, I think what should drive our decision is not the current runtime
overhead, but instead how realistic it is that we would actually reclaim
"enough" memory on average from these folios.
IOW, if the whole effort of scanning these small things is really worth it.
For an order-2 folio I'd assume "unlikely". Maybe we could collect some data
from some real workloads?
But I guess Meta is mostly focusing on 2M and doesn't really have data for any
other folio size how effective the underused scanner is for them.
>
> For v3, I'll go back to the scaling approach. I'll be traveling next
> week and then will be at LPC after that, so my timeline is to submit
> v3 after I come back from LPC.
Cool, we can continue the discussion at LPC! :)
Yeah, best to wait until after the next merge window.
--
Cheers,
David
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-28 19:20 ` David Hildenbrand (Arm)
@ 2026-10-01 9:59 ` Joanne Koong
0 siblings, 0 replies; 22+ messages in thread
From: Joanne Koong @ 2026-10-01 9:59 UTC (permalink / raw)
To: David Hildenbrand (Arm)
Cc: Lance Yang, akpm, ljs, usama.arif, hannes, baohua, alex, ziy,
baolin.wang, liam, npache, ryan.roberts, dev.jain, vbabka, rppt,
surenb, mhocko, willy, linux-mm
On Mon, Sep 28, 2026 at 8:20 PM David Hildenbrand (Arm)
<david@kernel.org> wrote:
>
> On 9/26/26 01:47, Joanne Koong wrote:
> > On Wed, Sep 23, 2026 at 2:44 AM David Hildenbrand (Arm)
> > <david@kernel.org> wrote:
> >>
> >> On 9/23/26 02:44, Joanne Koong wrote:
> >>> On Mon, Sep 21, 2026 at 3:12 AM David Hildenbrand (Arm)
> >>> <david@kernel.org> wrote:
> >>>
> >>> Ah, apologies for misinterpreting your comment.
> >>>
> >>>
> >>> With scaling, I don't see a way of avoiding having to queue every
> >>> anonymous folio of order >=2 to the deferred split list at fault time
> >>> in the non-default max_ptes_none case, since any folio can now
> >>> potentially qualify as underused.
> >>
> >> Exactly.
> >>
> >> I was primarily arguing that I don't want this overhead for the majority of
> >> Linux installations out there that ship with
> >>
> >> $ cat /sys/kernel/mm/transparent_hugepage/khugepaged/max_ptes_none
> >> 511
> >>
> >> IOW: we add them (PMD THP folios today!) to the deferred list for
> >> uuderused-scanning but never actually scan them because of:
> >>
> >> if (khugepaged_max_ptes_none == HPAGE_PMD_NR - 1)
> >> return false;
> >
> > Ah, I missed this nuance in your v1 reply - I thought it was objecting
> > to the queueing overhead for the non-default max_ptes_none case as
> > well. Patch 2 in this series does what you mentioned above and it
> > automatically applies to PMD THPs too. I'll add your Suggested-by: tag
> > to this in v3.
>
> I though Johannes suggested that :)
Reading back this message in v1 [1], I think both of you had suggested
it, so I'll just tag you both :D
>
> >> Right, I raised that. Johannes thinks it could be fixed with batching. I was
> >> concerned that it would become rather ugly. We'd have to see how that could look.
> >
> > I ran Barry's microbenchmark to get a sense of the contention (more of
> > the details are in my previous reply to Barry [1]). With max_ptes_none
> > set to 0 so that everything gets queued, and toggling shrink_underused
> > 0 vs 1 with patch 3 applied, I saw roughly a 15.8% and 11.5%
> > performance hit on runtime at 16k and 32k, but only 1.1% at 64k and
> > noisy above that.
> >
> > I didn't see the cost being from contention on the deferred list's
> > lock though. Perf showed that it was coming from
> > folio_lruvec_lock_irqsave, and I think that's from the free path where
> > in folios_put_refs(), __page_cache_release() keeps the lruvec locked
> > across the whole batch, and folio_unqueue_deferred_split() gets called
> > from inside that loop with that lruvec lock held. I don't think
> > batching the enqueues would help in that case, since the cost is
> > coming from the unqueue side and on a path that's already batched. I
> > think maybe we could move the unqueue out where the lruvec lock
> > doesn't need to be held while the unqueueing happens, but I need to
> > look more into that.
>
> That's an interesting insight. If contention is less of a problem right now with
> that reproducer, even better!
>
> I prefer us not doing unnecessary work (adding folios to the deferred split
> queue) when we won't really split them. (if it's contention or some other
> overhead doesn't really matter)
>
>
> >>>
> >>> Does that make scaling acceptable to you and Barry? If so, I'd prefer
> >>> that as well because if/when intermediary values are supported for
> >>> mTHP collapse, that support would have to be proportional, and it'd be
> >>> easier to reconcile the two with split also scaled. If not, then what
> >>> would be the suggestion for where to take v3?
> >> My opinion is (open for discussion :) ) that scaling is likely the better
> >> approach. With the following notes:
> >>
> >> (1) If we want to enable mTHP collapse with scaling as well, this needs very
> >> good documentation and also another thought on how to work around the problem of
> >> creep (e.g., refusing to collapse if creep would be possible according to the
> >> max_ptes_non setting and warning).
> >
> > Agreed. And if/when collapse gains proportional support later on,
> > having split already scaled means one value governing both instead of
> > two.
> >
> >>
> >> (2) Systems where the underused shrinker is effectively inactive should not add
> >> folios to the deferred list. Including PMD THPs, which we unconditionally add today.
> >
> > Agreed.
> >
> >>
> >> (3) If we want to exclude certain small folio sizes from the underused shrinker
> >> (e.g., order-2? order-3?) it might be better to just hard-code that in the
> >> kernel instead of giving the admin a choice it cannot possibly make easily.
> >
> > Based off the benchmark results in [1] and [2], I think the cutoff
> > needs to be at 64k where we exclude anything under that. On 4k base
> > pages, this is order 4, but on 64k base pages, order 4 would mean 1M,
> > which I think would be overly conservative. I think it makes more
> > sense to express it as an absolute threshold value instead of by
> > order. I'll try to find an arm64 machine to double-check this on.
>
> If we don't add folios to deferred split queue if current underused shrinker
> config wouldn't ever split them, I think there is no overhead for Android in any
> case (so Barry wouldn't have to worry, at least for now).
>
> So the magic number we chose only applies if the underused shrinker is enabled
> and could eventually split them.
>
> There, I think what should drive our decision is not the current runtime
> overhead, but instead how realistic it is that we would actually reclaim
> "enough" memory on average from these folios.
>
> IOW, if the whole effort of scanning these small things is really worth it.
>
> For an order-2 folio I'd assume "unlikely". Maybe we could collect some data
> from some real workloads?
>
> But I guess Meta is mostly focusing on 2M and doesn't really have data for any
> other folio size how effective the underused scanner is for them.
>
> >
I don't have any measurable data for the smaller sizes unfortunately
as our fleet only runs 2M. On the other subthread, Barry suggested
using 2M [2] as the threshold, which I like. That would cover Meta's
use case (64k base pages with 2M mTHP set to always) while being
conservative enough to keep smaller folios off the queue entirely,
where if we want to lower the threshold later, we could revisit it
when there is more substantial data. For v3, I'll use 2M as the
threshold.
Thanks,
Joanne
[1] https://lore.kernel.org/linux-mm/65b4baa3-77bd-4d2c-a128-9884f858e63d@kernel.org/
[2] https://lore.kernel.org/linux-mm/CAGsJ_4xWC4hg_-MLoDn3Zav7NUM709gu1jPeQpRuODdt8OUpZQ@mail.gmail.com/
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-16 22:44 [PATCH v2 0/3] mm: split underused anonymous mTHP folios Joanne Koong
` (3 preceding siblings ...)
2026-09-20 16:20 ` [PATCH v2 0/3] mm: split underused anonymous mTHP folios Lance Yang
@ 2026-09-21 10:23 ` David Hildenbrand (Arm)
2026-09-21 20:33 ` Joanne Koong
4 siblings, 1 reply; 22+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-21 10:23 UTC (permalink / raw)
To: Joanne Koong, akpm, ljs
Cc: usama.arif, hannes, baohua, alex, ziy, baolin.wang, liam, npache,
ryan.roberts, dev.jain, lance.yang, vbabka, rppt, surenb, mhocko,
willy, linux-mm
On 9/17/26 00:44, Joanne Koong wrote:
> PMD-sized THPs that are mostly zero-filled are reclaimed under memory
> pressure by the deferred split shrinker, but this is not done for mTHP
> folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
> base pages, as 2M provides the contpte benefits while the PMD size there
> (512M) is too big to use. However, 2M THP=always causes memory regressions
> unless the unused portions of those folios can be broken down and
> reclaimed.
>
> v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
> it queued every anonymous mTHP folio. David objected to the scaling since
> collapse had already rejected proportional scaling as either letting the memory
> footprint creep or being confusing to reason about. David and Barry separately
> objected to the queuing, which adds list_lru lock contention on every anonymous
> fault, and hurts lower-order use cases such as Android's order-2-only
> configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
> simply not queuing folios that could never exceed it, which addresses both.
>
> Read as an absolute count, khugepaged_max_ptes_none also implies the
> smallest folio that takes part in underused splitting. A 2M folio on 64k
> base pages is 32 pages, so the knob has to be set below 32 for those to
> be queued at all, and any folio with no more pages than the value stays off
> the queue entirely.
>
> Patch 1 makes thp_underused() folio-size aware. Today it only ever sees
> PMD-sized folios, so it assumes HPAGE_PMD_NR pages throughout. Patch 3
> breaks that assumption, so patch 1 generalizes it first. No functional
> changes are introduced.
>
> Patch 2 keeps folios that can never be found underused off the deferred
> split queue, as suggested by Johannes. This changes/optimizes existing
> behavior. With the default khugepaged/max_ptes_none, PMD folios are queued
> today and then dropped again by the first scan without ever having been
> splittable. After this patch they are not queued at all.
I think we discussed queuing folios once the value changes and you would
actually be able to split some. Was there a good reason why that idea was dropped?
--
Cheers,
David
^ permalink raw reply [flat|nested] 22+ messages in thread* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-21 10:23 ` David Hildenbrand (Arm)
@ 2026-09-21 20:33 ` Joanne Koong
2026-09-22 18:59 ` David Hildenbrand (Arm)
0 siblings, 1 reply; 22+ messages in thread
From: Joanne Koong @ 2026-09-21 20:33 UTC (permalink / raw)
To: David Hildenbrand (Arm)
Cc: akpm, ljs, usama.arif, hannes, baohua, alex, ziy, baolin.wang,
liam, npache, ryan.roberts, dev.jain, lance.yang, vbabka, rppt,
surenb, mhocko, willy, linux-mm
On Mon, Sep 21, 2026 at 3:23 AM David Hildenbrand (Arm)
<david@kernel.org> wrote:
>
> On 9/17/26 00:44, Joanne Koong wrote:
> > PMD-sized THPs that are mostly zero-filled are reclaimed under memory
> > pressure by the deferred split shrinker, but this is not done for mTHP
> > folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
> > base pages, as 2M provides the contpte benefits while the PMD size there
> > (512M) is too big to use. However, 2M THP=always causes memory regressions
> > unless the unused portions of those folios can be broken down and
> > reclaimed.
> >
> > v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
> > it queued every anonymous mTHP folio. David objected to the scaling since
> > collapse had already rejected proportional scaling as either letting the memory
> > footprint creep or being confusing to reason about. David and Barry separately
> > objected to the queuing, which adds list_lru lock contention on every anonymous
> > fault, and hurts lower-order use cases such as Android's order-2-only
> > configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
> > simply not queuing folios that could never exceed it, which addresses both.
> >
> > Read as an absolute count, khugepaged_max_ptes_none also implies the
> > smallest folio that takes part in underused splitting. A 2M folio on 64k
> > base pages is 32 pages, so the knob has to be set below 32 for those to
> > be queued at all, and any folio with no more pages than the value stays off
> > the queue entirely.
> >
> > Patch 1 makes thp_underused() folio-size aware. Today it only ever sees
> > PMD-sized folios, so it assumes HPAGE_PMD_NR pages throughout. Patch 3
> > breaks that assumption, so patch 1 generalizes it first. No functional
> > changes are introduced.
> >
> > Patch 2 keeps folios that can never be found underused off the deferred
> > split queue, as suggested by Johannes. This changes/optimizes existing
> > behavior. With the default khugepaged/max_ptes_none, PMD folios are queued
> > today and then dropped again by the first scan without ever having been
> > splittable. After this patch they are not queued at all.
>
> I think we discussed queuing folios once the value changes and you would
> actually be able to split some. Was there a good reason why that idea was dropped?
>
Hi David,
I have a paragraph about this in the commit message of patch 2:
"Please note that the queue has always been best-effort. A folio that was
queued before khugepaged/max_ptes_none is lowered gets dropped from the
queue by the first scan that finds it's not underused, so changing the
sysctl has never retroactively applied to folios that were already
scanned. Requeueing eligible folios when the sysctl changes will be
addressed in a separate patch."
I will be sending out the change for this as its own separate patch.
Thanks,
Joanne
^ permalink raw reply [flat|nested] 22+ messages in thread
* Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
2026-09-21 20:33 ` Joanne Koong
@ 2026-09-22 18:59 ` David Hildenbrand (Arm)
0 siblings, 0 replies; 22+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-22 18:59 UTC (permalink / raw)
To: Joanne Koong
Cc: akpm, ljs, usama.arif, hannes, baohua, alex, ziy, baolin.wang,
liam, npache, ryan.roberts, dev.jain, lance.yang, vbabka, rppt,
surenb, mhocko, willy, linux-mm
On 9/21/26 22:33, Joanne Koong wrote:
> On Mon, Sep 21, 2026 at 3:23 AM David Hildenbrand (Arm)
> <david@kernel.org> wrote:
>>
>> On 9/17/26 00:44, Joanne Koong wrote:
>>> PMD-sized THPs that are mostly zero-filled are reclaimed under memory
>>> pressure by the deferred split shrinker, but this is not done for mTHP
>>> folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
>>> base pages, as 2M provides the contpte benefits while the PMD size there
>>> (512M) is too big to use. However, 2M THP=always causes memory regressions
>>> unless the unused portions of those folios can be broken down and
>>> reclaimed.
>>>
>>> v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
>>> it queued every anonymous mTHP folio. David objected to the scaling since
>>> collapse had already rejected proportional scaling as either letting the memory
>>> footprint creep or being confusing to reason about. David and Barry separately
>>> objected to the queuing, which adds list_lru lock contention on every anonymous
>>> fault, and hurts lower-order use cases such as Android's order-2-only
>>> configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
>>> simply not queuing folios that could never exceed it, which addresses both.
>>>
>>> Read as an absolute count, khugepaged_max_ptes_none also implies the
>>> smallest folio that takes part in underused splitting. A 2M folio on 64k
>>> base pages is 32 pages, so the knob has to be set below 32 for those to
>>> be queued at all, and any folio with no more pages than the value stays off
>>> the queue entirely.
>>>
>>> Patch 1 makes thp_underused() folio-size aware. Today it only ever sees
>>> PMD-sized folios, so it assumes HPAGE_PMD_NR pages throughout. Patch 3
>>> breaks that assumption, so patch 1 generalizes it first. No functional
>>> changes are introduced.
>>>
>>> Patch 2 keeps folios that can never be found underused off the deferred
>>> split queue, as suggested by Johannes. This changes/optimizes existing
>>> behavior. With the default khugepaged/max_ptes_none, PMD folios are queued
>>> today and then dropped again by the first scan without ever having been
>>> splittable. After this patch they are not queued at all.
>>
>> I think we discussed queuing folios once the value changes and you would
>> actually be able to split some. Was there a good reason why that idea was dropped?
>>
>
> Hi David,
>
> I have a paragraph about this in the commit message of patch 2:
>
Ah, okay, I only skimmed the cover letter so far.
--
Cheers,
David
^ permalink raw reply [flat|nested] 22+ messages in thread