From: Lance Yang <lance.yang@linux.dev>
To: joannelkoong@gmail.com
Cc: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org,
usama.arif@linux.dev, hannes@cmpxchg.org, baohua@kernel.org,
alex@ghiti.fr, ziy@nvidia.com, baolin.wang@linux.alibaba.com,
liam@infradead.org, npache@redhat.com, ryan.roberts@arm.com,
dev.jain@arm.com, lance.yang@linux.dev, vbabka@kernel.org,
rppt@kernel.org, surenb@google.com, mhocko@suse.com,
willy@infradead.org, linux-mm@kvack.org
Subject: Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
Date: Mon, 21 Sep 2026 00:20:14 +0800 [thread overview]
Message-ID: <20260920162014.4479-1-lance.yang@linux.dev> (raw)
In-Reply-To: <20260916224437.1164512-1-joannelkoong@gmail.com>
On Wed, Sep 16, 2026 at 03:44:34PM -0700, Joanne Koong wrote:
>PMD-sized THPs that are mostly zero-filled are reclaimed under memory
>pressure by the deferred split shrinker, but this is not done for mTHP
>folios. At Meta we would like to deploy 2M THP=always on arm64 with 64k
>base pages, as 2M provides the contpte benefits while the PMD size there
>(512M) is too big to use. However, 2M THP=always causes memory regressions
>unless the unused portions of those folios can be broken down and
>reclaimed.
>
>v1 did two things. It scaled khugepaged_max_ptes_none down per folio order and
>it queued every anonymous mTHP folio. David objected to the scaling since
>collapse had already rejected proportional scaling as either letting the memory
>footprint creep or being confusing to reason about. David and Barry separately
>objected to the queuing, which adds list_lru lock contention on every anonymous
>fault, and hurts lower-order use cases such as Android's order-2-only
>configuration. Johannes suggested keeping khugepaged_max_ptes_none absolute and
>simply not queuing folios that could never exceed it, which addresses both.
>
>Read as an absolute count, khugepaged_max_ptes_none also implies the
>smallest folio that takes part in underused splitting. A 2M folio on 64k
>base pages is 32 pages, so the knob has to be set below 32 for those to
>be queued at all, and any folio with no more pages than the value stays off
>the queue entirely.
Em ... I'm not sure max_ptes_none should also become the mTHP underused
threshold ...
Say we're on arm64 with 64K pages and set it to 16. Khugepaged PMD
collapse uses 16, while mTHP collapse turns the same value into 0. With
this series, an order-5 (2 MiB) mTHP enters the deferred split queue
because it has 32 pages, and is considered underused once more than 16
pages are zero-filled. So the same knob means 16, 0, and 16 depending on
where it is used ... That's a bit odd ... no?
Would a per-order value make more sense here?
An unset value could preserve the current behavior, while an explicit
value could control both mTHP collapse and underused splitting. That
would also give us a way to leave max_ptes_none as the legacy fallback
instead of adding more meaning to it :)
Cheers, Lance
>Patch 1 makes thp_underused() folio-size aware. Today it only ever sees
>PMD-sized folios, so it assumes HPAGE_PMD_NR pages throughout. Patch 3
>breaks that assumption, so patch 1 generalizes it first. No functional
>changes are introduced.
>
>Patch 2 keeps folios that can never be found underused off the deferred
>split queue, as suggested by Johannes. This changes/optimizes existing
>behavior. With the default khugepaged/max_ptes_none, PMD folios are queued
>today and then dropped again by the first scan without ever having been
>splittable. After this patch they are not queued at all.
>
>Patch 3 queues anonymous mTHP folios from map_anon_folio_pte_nopf(),
>mirroring what map_anon_folio_pmd_nopf() already does for PMD folios. This
>covers both the fault path and the khugepaged mTHP collapse path.
>
>One consequence of keeping the knob absolute is that it is shared with
>collapse. A value low enough to be useful for 2M mTHP is a tiny
>fraction of a 512M PMD, so khugepaged will only collapse to PMD order
>when the region is almost fully populated. That is fine for the deployment
>this series targets, which does not use PMD THP on arm64, but anyone wanting
>both PMD THP and 2M mTHP on 64k pages should be aware of it.
>
>Thanks,
>Joanne
>
>Changelog
>---------
>v1: https://lore.kernel.org/linux-mm/20260707201735.4113107-1-joannelkoong@gmail.com/
>
>Changes since v1:
>* Drop the per-order scaling of khugepaged_max_ptes_none. Keep it
> absolute, matching collapse (David, Johannes)
>* New patch 2 (suggested by Johannes): don't queue folios that can never
> be underused, which both bounds what patch 3 adds and stops queuing PMD
> folios under default settings (Johannes, David, Barry)
>* Fix the thp_underused() early exit to scale to the folio's own size
>
>Joanne Koong (3):
> mm/huge_memory: make thp_underused() work for mTHP folios
> mm/huge_memory: don't queue folios that can never be underused
> mm/memory: add anonymous mTHP folios to the deferred split list
>
> mm/huge_memory.c | 45 +++++++++++++++++++++++++++++++++++++--------
> mm/memory.c | 2 ++
> 2 files changed, 39 insertions(+), 8 deletions(-)
>
>--
>2.52.0
>
>
>
next prev parent reply other threads:[~2026-09-20 16:20 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 22:44 [PATCH v2 0/3] mm: split underused anonymous mTHP folios Joanne Koong
2026-09-16 22:44 ` [PATCH v2 1/3] mm/huge_memory: make thp_underused() work for " Joanne Koong
2026-09-16 22:44 ` [PATCH v2 2/3] mm/huge_memory: don't queue folios that can never be underused Joanne Koong
2026-09-16 22:44 ` [PATCH v2 3/3] mm/memory: add anonymous mTHP folios to the deferred split list Joanne Koong
2026-09-20 16:20 ` Lance Yang [this message]
2026-09-21 10:04 ` [PATCH v2 0/3] mm: split underused anonymous mTHP folios Barry Song
2026-09-21 10:08 ` Usama Arif
2026-09-21 10:22 ` David Hildenbrand (Arm)
2026-09-22 10:33 ` Kiryl Shutsemau
2026-09-21 10:12 ` David Hildenbrand (Arm)
2026-09-23 0:44 ` Joanne Koong
2026-09-23 6:00 ` Barry Song
2026-09-25 22:40 ` Joanne Koong
2026-09-23 9:44 ` David Hildenbrand (Arm)
2026-09-25 23:47 ` Joanne Koong
2026-09-28 8:03 ` Barry Song
2026-10-01 9:10 ` Joanne Koong
2026-09-28 19:20 ` David Hildenbrand (Arm)
2026-10-01 9:59 ` Joanne Koong
2026-09-21 10:23 ` David Hildenbrand (Arm)
2026-09-21 20:33 ` Joanne Koong
2026-09-22 18:59 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260920162014.4479-1-lance.yang@linux.dev \
--to=lance.yang@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=alex@ghiti.fr \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=hannes@cmpxchg.org \
--cc=joannelkoong@gmail.com \
--cc=liam@infradead.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=npache@redhat.com \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox