From: Johannes Weiner <hannes@cmpxchg.org>
To: Barry Song <baohua@kernel.org>
Cc: "David Hildenbrand (Arm)" <david@kernel.org>,
Joanne Koong <joannelkoong@gmail.com>,
akpm@linux-foundation.org, ljs@kernel.org, usama.arif@linux.dev,
alex@ghiti.fr, ziy@nvidia.com, baolin.wang@linux.alibaba.com,
liam@infradead.org, npache@redhat.com, ryan.roberts@arm.com,
dev.jain@arm.com, lance.yang@linux.dev, vbabka@kernel.org,
rppt@kernel.org, surenb@google.com, mhocko@suse.com,
willy@infradead.org, linux-mm@kvack.org
Subject: Re: [PATCH v1 2/2] mm/memory: add anonymous mTHP folios to deferred split list
Date: Mon, 3 Aug 2026 20:48:44 -0400 [thread overview]
Message-ID: <anE27CtvCFx7sAuO@cmpxchg.org> (raw)
In-Reply-To: <CAGsJ_4x86EYsMz_n2NKDwTqeiv2pi4AcZb1RuZoBeUBR1iFqOA@mail.gmail.com>
Hello Barry,
On Tue, Aug 04, 2026 at 05:20:00AM +0800, Barry Song wrote:
> On Mon, Aug 3, 2026 at 10:45 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
> > On Sat, Aug 01, 2026 at 03:33:42PM +0800, Barry Song wrote:
> > > On Fri, Jul 31, 2026 at 11:04 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
> > > > Here is an idea: the THP shrinker will not consider anything unused
> > > > that has <= max_ptes_none zero pages. See thp_underused(). Joanne was
> > > > proposing to scale this knob down relative to the folio size for mTHP
> > > > shrinking. What if instead we kept the meaning absolute?
> > > >
> > > > The knob is an expression of how much waste the user is willing to
> > > > tolerate per folio. If the folio order in question couldn't possibly
> > > > have that much waste in the first place, we don't have to queue it?
> > > >
> > > > Something like this:
> > > >
> > > > diff --git a/mm/huge_memory.c b/mm/huge_memory.c
> > > > index 2bccb0a53a0a..1670e9869bd3 100644
> > > > --- a/mm/huge_memory.c
> > > > +++ b/mm/huge_memory.c
> > > > @@ -4364,6 +4364,9 @@ void deferred_split_folio(struct folio *folio, bool partially_mapped)
> > > > if (!partially_mapped && !split_underused_thp)
> > > > return;
> > > >
> > > > + if (!partially_mapped && folio_nr_pages(folio) <= khugepaged_max_ptes_none)
> > > > + return;
> > > > +
> > > > /*
> > > > * Exclude swapcache: originally to avoid a corrupt deferred split
> > > > * queue. Nowadays that is fully prevented by __memcg1_swapout();
> > > >
> > > > The setting defaults to the PMD-1, so out of the box we wouldn't queue
> > > > any new orders. It would allow that 2MB on 64k ARM usecase, without
> > > > jeopardizing smaller mTHP usecases like Barry's.
^^^
> > > My understanding is that adding smaller folios to the deferred split
> > > list is not the right approach. We could end up with a very large list
> > > where we cannot distinguish partially unmapped folios from fully mapped
> > > ones. For example, the deferred split list could contain 100 fully mapped
> > > folios but only a single partially mapped folio.
> > > Moreover, for smaller large folios, the number of zero subpages
> > > is likely to be small and short-lived.
> > >
> > > However, this is probably fine for larger large folios, since it is
> > > unlikely to significantly increase the size of the deferred split
> > > list. In other words, the deferred split list should remain manageable.
> > > For the same reason, larger large folios may not benefit much from the
> > > LRU cache either.
> > >
> > > So if we have some mechanism to prevent users from doing things that
> > > are not beneficial, such as adding smaller large folios to the list, it
> > > seems reasonable to me.
> >
> > Just to clarify, we're on the same page, right? I was proposing a
> > mechanism to do just that.
>
> For example, with smaller folios, the distribution might look
> something like this:
>
> Smaller large folios
>
> +------------------------------------------------------------+
> | F | F | F | F | F | F | F | F | F | F | F | F | F | P | Z |
> +------------------------------------------------------------+
>
> F: Fully mapped (dominant)
> P: Partially mapped (rare)
> Z: Zero subpages mapped (rare)
>
> So it doesn't make much sense to add them to the list, because we
> would rarely find P, and even the few Z entries we do find would
> soon be filled with non-zero data anyway.
>
> For larger folios, the list becomes much shorter. As folio size
> increases, they are more likely to contain zero-filled subpages:
>
> Larger large folios:
>
> +------------------------------------------------------------+
> | F | P | Z | F | P | Z | P | F | Z | F | P | Z | F | P | Z |
> +------------------------------------------------------------+
>
> F, P and Z become much more evenly distributed.
>
> So if the folio is large enough, this seems acceptable. I know a
> sysctl knob may not be well received, as it adds to the user's
> configuration burden. Perhaps we could just hard-code a
> sufficiently large value instead? for example,
>
> #define LARGE_FOLIO_ZERO_SCAN_MIN_SIZE SZ_2M
>
> if (folio_size(folio) >= LARGE_FOLIO_ZERO_SCAN_MIN_SIZE)
> deferred_split_folio(folio, false);
>
> If, someday, people find that 1 MiB also helps, they can provide
> data to support it.
I am very confused. Did you not see my proposal above?
Why not this?
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 2bccb0a53a0a..1670e9869bd3 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -4364,6 +4364,9 @@ void deferred_split_folio(struct folio *folio, bool partially_mapped)
if (!partially_mapped && !split_underused_thp)
return;
+ if (!partially_mapped && folio_nr_pages(folio) <= khugepaged_max_ptes_none)
+ return;
+
/*
* Exclude swapcache: originally to avoid a corrupt deferred split
* queue. Nowadays that is fully prevented by __memcg1_swapout();
next prev parent reply other threads:[~2026-08-04 0:48 UTC|newest]
Thread overview: 33+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-07 20:17 [PATCH v1 0/2] mm: split underused anonymous mTHP folios Joanne Koong
2026-07-07 20:17 ` [PATCH v1 1/2] mm/huge_memory: extend thp_underused() to " Joanne Koong
2026-07-08 4:03 ` Joanne Koong
2026-07-07 20:17 ` [PATCH v1 2/2] mm/memory: add anonymous mTHP folios to deferred split list Joanne Koong
2026-07-08 7:56 ` David Hildenbrand (Arm)
2026-07-08 9:58 ` Usama Arif
2026-07-29 12:36 ` David Hildenbrand (Arm)
2026-07-29 13:43 ` Usama Arif
2026-07-29 14:02 ` Usama Arif
2026-07-30 13:28 ` David Hildenbrand (Arm)
2026-07-08 17:52 ` Joanne Koong
2026-07-29 12:38 ` David Hildenbrand (Arm)
2026-07-29 15:04 ` Johannes Weiner
2026-07-30 13:46 ` David Hildenbrand (Arm)
2026-07-30 15:37 ` Johannes Weiner
2026-07-30 21:31 ` Barry Song
2026-07-30 23:32 ` Barry Song
2026-07-31 15:04 ` Johannes Weiner
2026-08-01 7:33 ` Barry Song
2026-08-03 14:45 ` Johannes Weiner
2026-08-03 21:20 ` Barry Song
2026-08-04 0:48 ` Johannes Weiner [this message]
2026-08-04 2:03 ` Barry Song
2026-08-04 14:03 ` Johannes Weiner
2026-08-04 14:11 ` David Hildenbrand (Arm)
2026-08-04 18:54 ` Johannes Weiner
2026-08-04 14:39 ` Usama Arif
2026-08-04 22:07 ` Barry Song
2026-08-04 23:26 ` Joanne Koong
2026-07-08 7:46 ` [PATCH v1 0/2] mm: split underused anonymous mTHP folios David Hildenbrand (Arm)
2026-07-08 18:19 ` Joanne Koong
2026-07-30 13:52 ` David Hildenbrand (Arm)
2026-07-30 23:10 ` Joanne Koong
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anE27CtvCFx7sAuO@cmpxchg.org \
--to=hannes@cmpxchg.org \
--cc=akpm@linux-foundation.org \
--cc=alex@ghiti.fr \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=joannelkoong@gmail.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=npache@redhat.com \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox