Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Joanne Koong <joannelkoong@gmail.com>
Cc: Lance Yang <lance.yang@linux.dev>,
	akpm@linux-foundation.org, ljs@kernel.org, usama.arif@linux.dev,
	hannes@cmpxchg.org, baohua@kernel.org, alex@ghiti.fr,
	ziy@nvidia.com, baolin.wang@linux.alibaba.com,
	liam@infradead.org, npache@redhat.com, ryan.roberts@arm.com,
	dev.jain@arm.com, vbabka@kernel.org, rppt@kernel.org,
	surenb@google.com, mhocko@suse.com, willy@infradead.org,
	linux-mm@kvack.org
Subject: Re: [PATCH v2 0/3] mm: split underused anonymous mTHP folios
Date: Mon, 28 Sep 2026 21:20:28 +0200	[thread overview]
Message-ID: <790584e4-e000-45c4-921d-e55f173466dc@kernel.org> (raw)
In-Reply-To: <CAJnrk1bD+4tndrOWuSENvEb7ody47HL2aNv2n+RzUqDE8mpkrQ@mail.gmail.com>

On 9/26/26 01:47, Joanne Koong wrote:
> On Wed, Sep 23, 2026 at 2:44 AM David Hildenbrand (Arm)
> <david@kernel.org> wrote:
>>
>> On 9/23/26 02:44, Joanne Koong wrote:
>>> On Mon, Sep 21, 2026 at 3:12 AM David Hildenbrand (Arm)
>>> <david@kernel.org> wrote:
>>>
>>> Ah, apologies for misinterpreting your comment.
>>>
>>>
>>> With scaling, I don't see a way of avoiding having to queue every
>>> anonymous folio of order >=2 to the deferred split list at fault time
>>> in the non-default max_ptes_none case, since any folio can now
>>> potentially qualify as underused.
>>
>> Exactly.
>>
>> I was primarily arguing that I don't want this overhead for the majority of
>> Linux installations out there that ship with
>>
>> $ cat /sys/kernel/mm/transparent_hugepage/khugepaged/max_ptes_none
>> 511
>>
>> IOW: we add them (PMD THP folios today!) to the deferred list for
>> uuderused-scanning but never actually scan them because of:
>>
>>         if (khugepaged_max_ptes_none == HPAGE_PMD_NR - 1)
>>                 return false;
> 
> Ah, I missed this nuance in your v1 reply - I thought it was objecting
> to the queueing overhead for the non-default max_ptes_none case as
> well. Patch 2 in this series does what you mentioned above and it
> automatically applies to PMD THPs too. I'll add your Suggested-by: tag
> to this in v3.

I though Johannes suggested that :)

>> Right, I raised that. Johannes thinks it could be fixed with batching. I was
>> concerned that it would become rather ugly. We'd have to see how that could look.
> 
> I ran Barry's microbenchmark to get a sense of the contention (more of
> the details are in my previous reply to Barry [1]). With max_ptes_none
> set to 0 so that everything gets queued, and toggling shrink_underused
> 0 vs 1 with patch 3 applied, I saw roughly a 15.8% and 11.5%
> performance hit on runtime at 16k and 32k, but only 1.1% at 64k and
> noisy above that.
> 
> I didn't see the cost being from contention on the deferred list's
> lock though. Perf showed that it was coming from
> folio_lruvec_lock_irqsave, and I think that's from the free path where
> in folios_put_refs(), __page_cache_release() keeps the lruvec locked
> across the whole batch, and folio_unqueue_deferred_split() gets called
> from inside that loop with that lruvec lock held. I don't think
> batching the enqueues would help in that case, since the cost is
> coming from the unqueue side and on a path that's already batched. I
> think maybe we could move the unqueue out where the lruvec lock
> doesn't need to be held while the unqueueing happens, but I need to
> look more into that.

That's an interesting insight. If contention is less of a problem right now with
that reproducer, even better!

I prefer us not doing unnecessary work (adding folios to the deferred split
queue) when we won't really split them. (if it's contention or some other
overhead doesn't really matter)


>>>
>>> Does that make scaling acceptable to you and Barry? If so, I'd prefer
>>> that as well because if/when intermediary values are supported for
>>> mTHP collapse, that support would have to be proportional, and it'd be
>>> easier to reconcile the two with split also scaled. If not, then what
>>> would be the suggestion for where to take v3?
>> My opinion is (open for discussion :) ) that scaling is likely the better
>> approach. With the following notes:
>>
>> (1) If we want to enable mTHP collapse with scaling as well, this needs very
>> good documentation and also another thought on how to work around the problem of
>> creep (e.g., refusing to collapse if creep would be possible according to the
>> max_ptes_non setting and warning).
> 
> Agreed. And if/when collapse gains proportional support later on,
> having split already scaled means one value governing both instead of
> two.
> 
>>
>> (2) Systems where the underused shrinker is effectively inactive should not add
>> folios to the deferred list. Including PMD THPs, which we unconditionally add today.
> 
> Agreed.
> 
>>
>> (3) If we want to exclude certain small folio sizes from the underused shrinker
>> (e.g., order-2? order-3?) it might be better to just hard-code that in the
>> kernel instead of giving the admin a choice it cannot possibly make easily.
> 
> Based off the benchmark results in [1] and [2], I think the cutoff
> needs to be at 64k where we exclude anything under that. On 4k base
> pages, this is order 4, but on 64k base pages, order 4 would mean 1M,
> which I think would be overly conservative. I think it makes more
> sense to express it as an absolute threshold value instead of by
> order. I'll try to find an arm64 machine to double-check this on.

If we don't add folios to deferred split queue if current underused shrinker
config wouldn't ever split them, I think there is no overhead for Android in any
case (so Barry wouldn't have to worry, at least for now).

So the magic number we chose only applies if the underused shrinker is enabled
and could eventually split them.

There, I think what should drive our decision is not the current runtime
overhead, but instead how realistic it is that we would actually reclaim
"enough" memory on average from these folios.

IOW, if the whole effort of scanning these small things is really worth it.

For an order-2 folio I'd assume "unlikely". Maybe we could collect some data
from some real workloads?

But I guess Meta is mostly focusing on 2M and doesn't really have data for any
other folio size how effective the underused scanner is for them.

> 
> For v3, I'll go back to the scaling approach. I'll be traveling next
> week and then will be at LPC after that, so my timeline is to submit
> v3 after I come back from LPC.
Cool, we can continue the discussion at LPC! :)

Yeah, best to wait until after the next merge window.

-- 
Cheers,

David


  parent reply	other threads:[~2026-09-28 19:20 UTC|newest]

Thread overview: 22+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16 22:44 [PATCH v2 0/3] mm: split underused anonymous mTHP folios Joanne Koong
2026-09-16 22:44 ` [PATCH v2 1/3] mm/huge_memory: make thp_underused() work for " Joanne Koong
2026-09-16 22:44 ` [PATCH v2 2/3] mm/huge_memory: don't queue folios that can never be underused Joanne Koong
2026-09-16 22:44 ` [PATCH v2 3/3] mm/memory: add anonymous mTHP folios to the deferred split list Joanne Koong
2026-09-20 16:20 ` [PATCH v2 0/3] mm: split underused anonymous mTHP folios Lance Yang
2026-09-21 10:04   ` Barry Song
2026-09-21 10:08   ` Usama Arif
2026-09-21 10:22     ` David Hildenbrand (Arm)
2026-09-22 10:33     ` Kiryl Shutsemau
2026-09-21 10:12   ` David Hildenbrand (Arm)
2026-09-23  0:44     ` Joanne Koong
2026-09-23  6:00       ` Barry Song
2026-09-25 22:40         ` Joanne Koong
2026-09-23  9:44       ` David Hildenbrand (Arm)
2026-09-25 23:47         ` Joanne Koong
2026-09-28  8:03           ` Barry Song
2026-10-01  9:10             ` Joanne Koong
2026-09-28 19:20           ` David Hildenbrand (Arm) [this message]
2026-10-01  9:59             ` Joanne Koong
2026-09-21 10:23 ` David Hildenbrand (Arm)
2026-09-21 20:33   ` Joanne Koong
2026-09-22 18:59     ` David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=790584e4-e000-45c4-921d-e55f173466dc@kernel.org \
    --to=david@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=alex@ghiti.fr \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=dev.jain@arm.com \
    --cc=hannes@cmpxchg.org \
    --cc=joannelkoong@gmail.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=npache@redhat.com \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=willy@infradead.org \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox