Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v2 0/3] mm: split underused anonymous mTHP folios
@ 2026-09-16 22:44 Joanne Koong
  2026-09-16 22:44 ` [PATCH v2 1/3] mm/huge_memory: make thp_underused() work for " Joanne Koong
                   ` (4 more replies)
  0 siblings, 5 replies; 22+ messages in thread
From: Joanne Koong @ 2026-09-16 22:44 UTC (permalink / raw)
  To: akpm, david, ljs
  Cc: usama.arif, hannes, baohua, alex, ziy, baolin.wang, liam, npache,
	ryan.roberts, dev.jain, lance.yang, vbabka, rppt, surenb, mhocko,
	willy, linux-mm

PMD-sized THPs that are mostly zero-filled are reclaimed under memory
pressure by the deferred split shrinker, but this is not done for mTHP
folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
base pages, as 2M provides the contpte benefits while the PMD size there
(512M) is too big to use. However, 2M THP=always causes memory regressions
unless the unused portions of those folios can be broken down and
reclaimed.

v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
it queued every anonymous mTHP folio. David objected to the scaling since
collapse had already rejected proportional scaling as either letting the memory
footprint creep or being confusing to reason about. David and Barry separately
objected to the queuing, which adds list_lru lock contention on every anonymous
fault, and hurts lower-order use cases such as Android's order-2-only
configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
simply not queuing folios that could never exceed it, which addresses both.

Read as an absolute count, khugepaged_max_ptes_none also implies the
smallest folio that takes part in underused splitting. A 2M folio on 64k
base pages is 32 pages, so the knob has to be set below 32 for those to
be queued at all, and any folio with no more pages than the value stays off
the queue entirely.

Patch 1 makes thp_underused() folio-size aware. Today it only ever sees
PMD-sized folios, so it assumes HPAGE_PMD_NR pages throughout. Patch 3
breaks that assumption, so patch 1 generalizes it first. No functional
changes are introduced.

Patch 2 keeps folios that can never be found underused off the deferred
split queue, as suggested by Johannes. This changes/optimizes existing
behavior. With the default khugepaged/max_ptes_none, PMD folios are queued
today and then dropped again by the first scan without ever having been
splittable. After this patch they are not queued at all.

Patch 3 queues anonymous mTHP folios from map_anon_folio_pte_nopf(),
mirroring what map_anon_folio_pmd_nopf() already does for PMD folios. This
covers both the fault path and the khugepaged mTHP collapse path.

One consequence of keeping the knob absolute is that it is shared with
collapse. A value low enough to be useful for 2M mTHP is a tiny
fraction of a 512M PMD, so khugepaged will only collapse to PMD order
when the region is almost fully populated. That is fine for the deployment
this series targets, which does not use PMD THP on arm64, but anyone wanting
both PMD THP and 2M mTHP on 64k pages should be aware of it.

Thanks,
Joanne

Changelog
---------
v1: https://lore.kernel.org/linux-mm/20260707201735.4113107-1-joannelkoong@gmail.com/

Changes since v1:
* Drop the per-order scaling of khugepaged_max_ptes_none. Keep it
  absolute, matching collapse (David, Johannes)
* New patch 2 (suggested by Johannes): don't queue folios that can never
  be underused, which both bounds what patch 3 adds and stops queuing PMD
  folios under default settings (Johannes, David, Barry)
* Fix the thp_underused() early exit to scale to the folio's own size

Joanne Koong (3):
  mm/huge_memory: make thp_underused() work for mTHP folios
  mm/huge_memory: don't queue folios that can never be underused
  mm/memory: add anonymous mTHP folios to the deferred split list

 mm/huge_memory.c | 45 +++++++++++++++++++++++++++++++++++++--------
 mm/memory.c      |  2 ++
 2 files changed, 39 insertions(+), 8 deletions(-)

-- 
2.52.0



^ permalink raw reply	[flat|nested] 22+ messages in thread

end of thread, other threads:[~2026-10-01 10:00 UTC | newest]

Thread overview: 22+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-16 22:44 [PATCH v2 0/3] mm: split underused anonymous mTHP folios Joanne Koong
2026-09-16 22:44 ` [PATCH v2 1/3] mm/huge_memory: make thp_underused() work for " Joanne Koong
2026-09-16 22:44 ` [PATCH v2 2/3] mm/huge_memory: don't queue folios that can never be underused Joanne Koong
2026-09-16 22:44 ` [PATCH v2 3/3] mm/memory: add anonymous mTHP folios to the deferred split list Joanne Koong
2026-09-20 16:20 ` [PATCH v2 0/3] mm: split underused anonymous mTHP folios Lance Yang
2026-09-21 10:04   ` Barry Song
2026-09-21 10:08   ` Usama Arif
2026-09-21 10:22     ` David Hildenbrand (Arm)
2026-09-22 10:33     ` Kiryl Shutsemau
2026-09-21 10:12   ` David Hildenbrand (Arm)
2026-09-23  0:44     ` Joanne Koong
2026-09-23  6:00       ` Barry Song
2026-09-25 22:40         ` Joanne Koong
2026-09-23  9:44       ` David Hildenbrand (Arm)
2026-09-25 23:47         ` Joanne Koong
2026-09-28  8:03           ` Barry Song
2026-10-01  9:10             ` Joanne Koong
2026-09-28 19:20           ` David Hildenbrand (Arm)
2026-10-01  9:59             ` Joanne Koong
2026-09-21 10:23 ` David Hildenbrand (Arm)
2026-09-21 20:33   ` Joanne Koong
2026-09-22 18:59     ` David Hildenbrand (Arm)

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox