From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Joanne Koong <joannelkoong@gmail.com>
Cc: Lance Yang <lance.yang@linux.dev>,
akpm@linux-foundation.org, ljs@kernel.org, usama.arif@linux.dev,
hannes@cmpxchg.org, baohua@kernel.org, alex@ghiti.fr,
ziy@nvidia.com, baolin.wang@linux.alibaba.com,
liam@infradead.org, npache@redhat.com, ryan.roberts@arm.com,
dev.jain@arm.com, vbabka@kernel.org, rppt@kernel.org,
surenb@google.com, mhocko@suse.com, willy@infradead.org,
linux-mm@kvack.org
Subject: Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
Date: Wed, 23 Sep 2026 11:44:27 +0200 [thread overview]
Message-ID: <00d07f4a-f679-486f-9413-6917b5213854@kernel.org> (raw)
In-Reply-To: <CAJnrk1Zk_4awd5OfbvsCLts1YPsB6JjjOF2hiSmg-wAp+_KUDw@mail.gmail.com>
On 9/23/26 02:44, Joanne Koong wrote:
> On Mon, Sep 21, 2026 at 3:12 AM David Hildenbrand (Arm)
> <david@kernel.org> wrote:
>>
>> On 9/20/26 18:20, Lance Yang wrote:
>>>
>>
>> "David objected to the scaling", not quite. The tricky bit is the scaling when
>> collapsing (creep).
>>
>> When freeing, I'd assume scaling should be alright.
>
> Ah, apologies for misinterpreting your comment.
>
>>
>>>
>>> Em ... I'm not sure max_ptes_none should also become the mTHP underused
>>> threshold ...
>>>
>>> Say we're on arm64 with 64K pages and set it to 16. Khugepaged PMD
>>> collapse uses 16, while mTHP collapse turns the same value into 0. With
>>> this series, an order-5 (2 MiB) mTHP enters the deferred split queue
>>> because it has 32 pages, and is considered underused once more than 16
>>> pages are zero-filled. So the same knob means 16, 0, and 16 depending on
>>> where it is used ... That's a bit odd ... no?
>>
>> I very much prefer scaling over using absolute numbers as proposed here.
>>
>
> With scaling, I don't see a way of avoiding having to queue every
> anonymous folio of order >=2 to the deferred split list at fault time
> in the non-default max_ptes_none case, since any folio can now
> potentially qualify as underused.
Exactly.
I was primarily arguing that I don't want this overhead for the majority of
Linux installations out there that ship with
$ cat /sys/kernel/mm/transparent_hugepage/khugepaged/max_ptes_none
511
IOW: we add them (PMD THP folios today!) to the deferred list for
uuderused-scanning but never actually scan them because of:
if (khugepaged_max_ptes_none == HPAGE_PMD_NR - 1)
return false;
Now, what you are saying is we might want less overhead for installations that
modify max_ptes_none, like Meta's fleet. See below.
> This drew some objections in v1,
> particularly for the Android small folios case. I read that as ruling
> out scaling. The objections were a) the list_lru lock contention and
It was the "deferred shrinker list and lock" contention, yes.
> b) the limited upside vs cost of splitting small folios.
>
> For a), I don't think there were any hard numbers showing this is a
> problem. If it would be helpful to have some benchmark numbers, I'm
> happy to run a synthetic one (patch 3 applied (which adds anon mTHP
> folios to the deferred split list), only order-2 enabled,
> max_ptes_none set low enough, all threads in one cgroup on one node so
> they contend on the same lock, each thread faulting in a large private
> anonymous mapping). I'm happy to run something else if anyone has a
> workload they would trust more. Johannes also mentioned batching as a
> mitigation. If the contention does turn out to be an issue, I'm happy
> to prototype that and measure it too.
Right, I raised that. Johannes thinks it could be fixed with batching. I was
concerned that it would become rather ugly. We'd have to see how that could look.
>
> For b), v1 unconditionally queued every folio of order >= 2. v2 added
> an optimization patch for skipping queueing entirely if max_ptes_none
> is at its default value, which would also carry over if we went with
> scaling. So I think with that patch added, for the objection to b), it
> depends on what max_ptes_none Android actually runs. At the default
> max_ptes_none, nothing would be queued at all, so a config that sees
> no benefit from splitting small folios would pay nothing for it.
I would assume everybody except some special users that fine-tune underused
shrinker run with the default. I did some digging in the past regarding public
documentation about how to set the parameter, and most of them just defaulted to
511.
I would be surprised if Android even knows about this parameter ;)
>
> Does that make scaling acceptable to you and Barry? If so, I'd prefer
> that as well because if/when intermediary values are supported for
> mTHP collapse, that support would have to be proportional, and it'd be
> easier to reconcile the two with split also scaled. If not, then what
> would be the suggestion for where to take v3?
My opinion is (open for discussion :) ) that scaling is likely the better
approach. With the following notes:
(1) If we want to enable mTHP collapse with scaling as well, this needs very
good documentation and also another thought on how to work around the problem of
creep (e.g., refusing to collapse if creep would be possible according to the
max_ptes_non setting and warning).
(2) Systems where the underused shrinker is effectively inactive should not add
folios to the deferred list. Including PMD THPs, which we unconditionally add today.
(3) If we want to exclude certain small folio sizes from the underused shrinker
(e.g., order-2? order-3?) it might be better to just hard-code that in the
kernel instead of giving the admin a choice it cannot possibly make easily.
--
Cheers,
David
next prev parent reply other threads:[~2026-09-23 9:44 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 22:44 [PATCH v2 0/3] mm: split underused anonymous mTHP folios Joanne Koong
2026-09-16 22:44 ` [PATCH v2 1/3] mm/huge_memory: make thp_underused() work for " Joanne Koong
2026-09-16 22:44 ` [PATCH v2 2/3] mm/huge_memory: don't queue folios that can never be underused Joanne Koong
2026-09-16 22:44 ` [PATCH v2 3/3] mm/memory: add anonymous mTHP folios to the deferred split list Joanne Koong
2026-09-20 16:20 ` [PATCH v2 0/3] mm: split underused anonymous mTHP folios Lance Yang
2026-09-21 10:04 ` Barry Song
2026-09-21 10:08 ` Usama Arif
2026-09-21 10:22 ` David Hildenbrand (Arm)
2026-09-22 10:33 ` Kiryl Shutsemau
2026-09-21 10:12 ` David Hildenbrand (Arm)
2026-09-23 0:44 ` Joanne Koong
2026-09-23 6:00 ` Barry Song
2026-09-25 22:40 ` Joanne Koong
2026-09-23 9:44 ` David Hildenbrand (Arm) [this message]
2026-09-25 23:47 ` Joanne Koong
2026-09-28 8:03 ` Barry Song
2026-10-01 9:10 ` Joanne Koong
2026-09-28 19:20 ` David Hildenbrand (Arm)
2026-10-01 9:59 ` Joanne Koong
2026-09-21 10:23 ` David Hildenbrand (Arm)
2026-09-21 20:33 ` Joanne Koong
2026-09-22 18:59 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=00d07f4a-f679-486f-9413-6917b5213854@kernel.org \
--to=david@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=alex@ghiti.fr \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=dev.jain@arm.com \
--cc=hannes@cmpxchg.org \
--cc=joannelkoong@gmail.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=npache@redhat.com \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox