From: Kiryl Shutsemau <kirill@shutemov.name>
To: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
Cc: akpm@linux-foundation.org, david@kernel.org,
nico.pache@linux.dev, baolin.wang@linux.alibaba.com,
baohua@kernel.org, dev.jain@arm.com, hughd@google.com,
lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com,
rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org,
surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org,
ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com,
linux-mm@kvack.org, linux-kselftest@vger.kernel.org,
linux-kernel@vger.kernel.org, jannh@google.com,
willy@infradead.org, pfalcato@suse.de, rostedt@goodmis.org,
mhiramat@kernel.org, linux-trace-kernel@vger.kernel.org,
bpf@vger.kernel.org
Subject: Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Date: Mon, 17 Aug 2026 14:38:44 +0100 [thread overview]
Message-ID: <aoLkMUtmzRVuv2Hx@thinkstation> (raw)
In-Reply-To: <aoLAly-oL2OAa4n8@lucifer>
On Mon, Aug 17, 2026 at 09:52:13AM +0100, Lorenzo Stoakes (ARM) wrote:
> We have a THP cabal meeting every couple of weeks where it would have been
> useful for you to raise this first.
Fair -- though my invite is on my old @linux.intel.com address. Could
you forward it to kas@kernel.org?
> In any case - this series is not something we'd consider at the moment,
> even broken into parts.
>
> David and I have put THP into feature freeze - until the codebase is
> subtantially improved we're not really interested in seeing significant
> development.
>
> The technical debt is substantial and has to be paid down first.
>
> See [0] for a rough list of TODOs in this regard.
I read the TODO list and I'll pick from it -- though I notice the
technical debt section includes "Literally all of the code in
mm/huge_memory.c and mm/khugepaged.c", which I'd argue this series is a
fairly committed attempt at :)
One clean up I wanted to do is consolidate code by functionality, not by
the THP/non-THP split. Move all page fault handler code into mm/fault.c,
unmap code into mm/zap.c, fork's copying into mm/fork.c -- mirroring
kernel/fork.c, so the mm half of a subsystem sits under the same name.
Large folios are an integral part of mm nowadays and I don't think we
benefit from keeping THP in a separate file. It is also an opportunity to
shift away from mm/memory.c being a kitchen sink.
David and I talked about this at LSF/MM. I can give it a try if it fits
your idea of "feature freeze" -- and if it doesn't collide with the series
you have in flight, in which case I'd rather go after yours than around it.
> > Why
> > ===
> >
> > mTHP collapse landed in khugepaged in 7.2 and I was glad to see it. We
> > at Meta run arm64 with 64K base pages, where a PMD is 512M: PMD-order THP
> > is of limited use at that size, and mTHP is exactly what we want.
>
> Do you have some numbers that indicate to what degree mTHP khugepaged is
> beneficial?
Not from the fleet yet -- that experiment is still ahead of me, so I can't
give you order-by-order numbers.
What I can say is that on x86 we lean on khugepaged heavily to get THPs in
place; it is not a marginal contributor for us. On arm64 with 64K base
pages we get nowhere near the x86 numbers without khugepaged being able to
produce mTHP at all.
I'll grant the other half of it, and more strongly than you put it: for a
64K mTHP on a 4K base page the TLB win is modest, and with today's
mechanism -- which clears and flushes the whole 2M PMD to install it -- I
can believe the disruption exceeds the gain and the net effect on a
workload is negative. That is what I measured: a thread reading and
writing a region while khugepaged collapses it at order-4, read p99 3071ns
against 1023ns. It shows the disruption is real and that it comes down;
whether the collapse pays for itself at that order is a separate question.
> > khugepaged only ever looks at PMD-aligned windows, and it is not an easy
> > limitation to lift.
>
> Yes. This assumption is very much baked in.
>
> I guess this is coming from the perspective of having ranges that are
> neither PMD-aligned nor sized (far harder to achieve with 512 MiB PMD size
> obviously).
Right, and it is the size rather than the alignment that bites. The real
requirement is that a VMA contain a whole PMD-aligned, PMD-sized range,
which at 2M most anonymous mappings of any size manage and at 512M almost
none do. That is why the limitation was easy to miss until the PMD got
big.
> > Fixing the alignment is a one-line change, but what it feeds assumes the
>
> Hmm not so sure about that... especially given how baked in these
> assumptions are.
I think we agree -- that was the setup, not the claim. Dropping the
ALIGN() is the one line; the point of the sentence is that it buys nothing
on its own, because what you then hand a sub-PMD range to still clears the
whole PMD and still demands the VMA span it.
> Which by the way, all speaks to the need for rework.
>
> The first stage in my view would be to improve the code to the point that
> these kinds of assumptions fall out of it, which then lays the foundations
> for future changes to eliminate the assumptions.
For the plumbing, yes -- policy, file layout, the scan/run split all
improve by refactoring in place.
I don't think the locking model gets there that way, though, which is why
I built a second engine rather than morphing the first. The old safety
argument is "hold the address space still"; the new one is "make the
sources inert". They are not two points on a line -- mmap_write cannot go
before something else holds the sources still, and doing that inside
collapse_huge_page() means replacing the copy, the install and the rollback
at the same time, which is the whole function. Every halfway state has
neither argument in full.
What is gradual here is the review rather than the mechanism: the engine
arrives one pass at a time, each reviewable alone, with the old one live
until one patch switches over.
> > PMD everywhere that matters: collapse_huge_page() clears and flushes the
> > whole PMD whatever order it is collapsing, installs a PMD leaf because
> > that is the only thing it can produce, and keeps everyone out with
> > mmap_write_lock, anon_vma_lock_write() and an IPI broadcast while it
> > does.
>
> To be clear - the anon path. I think important to clarify :)
Yes, anon only. The file path moves into collapse.c and picks up the
scan/run split, but its mechanism is untouched.
And you are right that "installs a PMD leaf" is wrong above: only a
PMD-order collapse installs one, a sub-PMD collapse repopulates the table
with PTEs. What is order-blind is everything around it -- the PMD is
still cleared and flushed first.
I would like to bring it into the engine as well, and with it the private
copies in MAP_PRIVATE file mappings, which neither path collapses today --
the anon side requires vma_is_anonymous() and the file side works on the
page cache.
I stopped because I wanted to keep the patch count in double digits. :P
> And yeah it does IPI for any sensible arch (with
> CONFIG_MMU_GATHER_RCU_TABLE_FREE) via tlb_remove_table_sync_one(). The
> other arches IPI anyway on TLB invalidation.
>
> [Though I intend to make all page table freeing RCU relatively soon which
> should? Eliminate the need for this, possibly?]
That would suit this well, and it is worth covering deposited page tables
in it if they are not already in scope. Today PMD collapse has to deposit
a freshly allocated table rather than redepositing the one it detached,
because a deposited table must be safe for zap_deposited_table() to free
immediately. Make that free RCU-deferred and the detached table can go
straight back -- one allocation less per PMD collapse.
> However we have to remember that a lot of the user-visible API assumes PMD
> sizing and so the code has to clearly reflect this and make it clear that
Your sentence got cut off, but if this is about the tunables then let me
flag what the series does with them. max_ptes_none, _swap and _shared are
counts out of a PMD, and a range smaller than one has fewer PTEs than the
budget, so a raw comparison can never refuse it -- a 64K range on 4K pages
is 16 PTEs against a max_ptes_shared default of 256. For swap and shared
the engine therefore compares fractions: count * HPAGE_PMD_NR against
max * nr_scanned.
max_ptes_none stays as mTHP collapse has it, all-or-nothing: 0, or
everything at that order. That is deliberate -- allowing holes at an order
below the largest enabled one lets khugepaged fill them and collapse the
result at the next order up, which is the ratchet max_ptes_none exists to
bound. There is room to scale it at the terminal order, where there is no
larger order to creep into, but I have not done that here.
> > Design
> > ======
> >
> > The old mechanism holds the address space still because it has nothing
> > else stopping the sources from moving under the copy. The new engine
> > makes the sources themselves inert instead, with the two barriers
> > migration already uses, raised in that order:
> >
> > 1. migration entries replace the source PTEs. Faults and GUP-slow
> > now wait on the source folio's lock, which is taken before the
> > first entry becomes visible.
> > 2. the source folio's refcount is frozen to its expected value.
> > GUP-fast, pfn walkers, reclaim, compaction and memory-failure all
> > fail folio_try_get() and back off.
> >
> > Between the two, nothing can reach a source, so the copy runs with no
> > lock held at all -- and the address space is left alone while it does.
>
> Hmm, are migration entries the right mechnanism here? Are you actually
> migrating the pages to a large folio here, or using them to get the
> behaviour you want on fault/GUP?
Both, and I would argue the behaviour is not a side effect: what a
migration entry means to a waiter -- this page is going away, sleep on its
folio lock and look again -- is exactly true of a source under collapse.
Fault, GUP-slow and rmap then all do the right thing with no new code,
which is the case for reusing the entry rather than inventing a marker
every waiter would have to learn.
What is not reused is mm/migrate.c. A migration entry encodes one PFN, so
it cannot name an N:1 destination of a different order; the engine takes
the hold-still half and does the remap itself at install. It is a
migration in substance -- contents move to another folio, the old mappings
are replaced -- but not one migrate.c could drive.
> Same question in general for the freezing.
folio_ref_freeze() means nobody may take a new reference, which is the
property the copy needs, and is why migration and split use it too.
> > What that removes from every collapse path:
> >
> > mmap_write_lock -> mmap_read
> > anon_vma_lock_write() -> nothing: an rmap walk needs the folio
> > locked, and the engine holds that lock
> > from freeze to putback
>
> I do like the idea of eliminating uses of the rmap lock like this, not only
> for contention's sake but also for scalable CoW purposes which introduces
> challenges with regards to holding these.
>
> In fact, migration and huge memory collapse are the really problematic
> areas.
If that is about their rmap complexity, collapse gets easier here rather
than harder.
The engine takes no rmap lock and walks no rmap. A page shared with
another process is unshared before anything is frozen -- the fault-in
pass breaks CoW, exactly as a write would -- so by freeze time a source
is exclusive to this mm and its only live mappings are the ones being
replaced. Migration has to cope with a folio mapped from many mms; this
never sees one.
If what scalable CoW needs is that collapse stops messing with rmap,
that is what this does.
>
> > tlb_remove_table_sync_one() -> nothing: one ranged flush per round
>
> Aren't we reliant upon this synchronisation for correctness?
Not in the new engine.
In collapse_huge_page() the IPI is what makes the *detached* table safe
to use: it pmdp_collapse_flush()es the PMD and then copies out of the
table it just detached, so it has to know no lockless walker is still
inside it.
The new engine never does that -- a sub-PMD window is collapsed in place
under the page table lock, so nothing is detached, and at PMD order the
table is detached, never touched again, and freed with pte_free_defer().
The synchronisation is still there, it is RCU rather than an IPI; and on
the arches without RCU table free the flush itself IPIs, as you say.
> > - A table that cannot become one huge page still yields the largest
> > windows inside it, where before a single disqualified PTE gave up
> > the whole table.
>
> Are you permitting collapse of ranges that straddle PTEs?
No -- a candidate never crosses a page table or a VMA. A round works
within one table, and each candidate is validated against the VMA at its
own order.
> Though in general I'm confused by the single disqualified PTE here -
Taking that literally is how I meant it: collapse_scan_pmd() goto
out_unmap's on the first PTE that fails any of its checks -- uffd, non-anon,
clean lazyfree, off the LRU, unexpected refcount -- and mthp_collapse() only
runs if the verdict came back SCAN_SUCCEED. So one
such PTE anywhere in the table means nothing in that table collapses, at
any order, even at an order whose windows avoid it entirely.
> > There may be a way out -- a PMD migration entry over the table during
> > the window, so the CPU never caches a walk to shoot down -- but that
> > means teaching every pmd-level walker a new kind of entry, and I have
> > not tried it.
>
> Hmm this seems like complexity on top of complexity...
I find it rather elegant, and expect it to be minimally intrusive: let such
a PMD be walkable exactly as a present one is, so software descends through
it as usual while the CPU sees a non-present entry and caches nothing.
Transparent to software, opaque to the CPU.
Out of scope for this patchset either way.
--
Kiryl Shutsemau / Kirill A. Shutemov
prev parent reply other threads:[~2026-08-17 13:38 UTC|newest]
Thread overview: 65+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-16 22:45 [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 01/57] mm: add pte_folio() Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 02/57] mm: add pte_none_or_zero() Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 03/57] mm/collapse: add collapse.h for the shared collapse state Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 04/57] mm/collapse: rename mthp_present_ptes to eligible_ptes Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 05/57] mm/collapse: state what a collapse may do in the policy Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 06/57] mm/collapse: move the smallest collapse order to collapse.h Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 07/57] mm/collapse: sketch the new anonymous collapse engine Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 08/57] mm/collapse: scan a table for what a collapse could use Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 09/57] mm/collapse: collect candidate windows into a round Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 10/57] mm/collapse: run a round and feed the outcomes back Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 11/57] mm/collapse: sketch the passes of a round Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 12/57] mm/collapse: allocate a destination per candidate Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 13/57] mm/collapse: revalidate a round against the VMA Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 14/57] mm/collapse: fault the sources in before the freeze Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 15/57] mm/collapse: check what a candidate would freeze Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 16/57] mm/collapse: freeze the sources behind migration entries Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 17/57] mm/collapse: copy the sources into the destinations Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 18/57] mm/collapse: install the destinations at PTE level Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 19/57] mm/collapse: install a PMD leaf as the terminal layer Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 20/57] mm/collapse: put the sources back Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 21/57] mm/collapse: settle whatever the round reached Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 22/57] mm/collapse: walk a table with a selection cursor Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 23/57] mm/collapse: give a refused region a second chance Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 24/57] mm/collapse: report each candidate's outcome to tracing Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 25/57] mm/collapse: collapse anonymous memory with the new engine Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 26/57] mm/collapse: give collapse_single_pmd() the range to work on Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 27/57] mm/collapse: scan the windows a VMA can hold Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 28/57] mm/collapse: remove the mechanism the engine replaces Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 29/57] mm/collapse: move what a collapse is judged on into collapse.c Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 30/57] mm/collapse: name the max_ptes ceiling after collapse Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 31/57] mm/khugepaged: count collapses where khugepaged makes them Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 32/57] mm/collapse: move the file collapse into collapse.c Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 33/57] mm/collapse: split collapse into a scan and a run Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 34/57] mm/collapse: implement MADV_COLLAPSE in madvise.c Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 35/57] mm/madvise: drop MADV_COLLAPSE's redundant mm reference Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 36/57] mm/collapse: report what the scan found Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 37/57] mm/collapse: report what the fault-in pass paid Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 38/57] mm/collapse: report the round, and what it made faulters wait Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 39/57] mm/collapse: name the file collapse's tracepoints after collapse Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 40/57] mm/collapse: remove the tracepoints of the mechanism that is gone Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 41/57] mm/collapse: give collapse its own trace header Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 42/57] mm/collapse: allow error injection into the freeze Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 43/57] mm/khugepaged: check the scan budget before the work, not after Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 44/57] mm/khugepaged: hold the address space open across a scan Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 45/57] mm/collapse: take a per-VMA read lock for the round Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 46/57] mm/khugepaged: scan under a per-VMA read lock Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 47/57] mm/madvise: collapse " Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 48/57] mm/collapse: assert the mm reference the engine relies on Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 49/57] mm/khugepaged: drop the mmap_lock barrier from __khugepaged_exit() Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 50/57] selftests/mm: attribute collapses by candidate event alone Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 51/57] selftests/mm: cover collapse inside a sub-PMD VMA Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 52/57] selftests/mm: cover a hole-y window in " Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 53/57] selftests/mm: cover collapse of mlocked ranges Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 54/57] selftests/mm: cover collapse beside a MADV_FREE'd page Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 55/57] selftests/mm: cover collapse beside a pinned page Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 56/57] selftests/mm: cover the scaled max_ptes_shared limit Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 57/57] MAINTAINERS: add an entry for collapse Kiryl Shutsemau
2026-08-17 8:04 ` Lorenzo Stoakes (ARM)
2026-08-17 8:08 ` David Hildenbrand (Arm)
2026-08-17 10:12 ` Kiryl Shutsemau
2026-08-17 2:02 ` [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives Zi Yan
2026-08-17 10:07 ` Kiryl Shutsemau
2026-08-17 8:52 ` Lorenzo Stoakes (ARM)
2026-08-17 13:38 ` Kiryl Shutsemau [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aoLkMUtmzRVuv2Hx@thinkstation \
--to=kirill@shutemov.name \
--cc=agordeev@linux.ibm.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=bpf@vger.kernel.org \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=hughd@google.com \
--cc=jannh@google.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=ljs@kernel.org \
--cc=mhiramat@kernel.org \
--cc=mhocko@suse.com \
--cc=nico.pache@linux.dev \
--cc=pfalcato@suse.de \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shuah@kernel.org \
--cc=surenb@google.com \
--cc=usama.anjum@arm.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.