From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Kiryl Shutsemau <kirill@shutemov.name>,
akpm@linux-foundation.org, ljs@kernel.org, nico.pache@linux.dev
Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org,
dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev,
liam@infradead.org, mhocko@suse.com, rppt@kernel.org,
ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com,
usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com,
usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org,
linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org,
kas@kernel.org, jannh@google.com, willy@infradead.org,
pfalcato@suse.de, rostedt@goodmis.org, mhiramat@kernel.org,
linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org
Subject: Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Date: Tue, 18 Aug 2026 15:55:55 +0200 [thread overview]
Message-ID: <9f51ac27-24b2-495b-b397-84865f977d24@kernel.org> (raw)
In-Reply-To: <20260816224609.308019-1-kirill@shutemov.name>
On 8/17/26 00:45, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> Yes, I know, this is a lot of changes. But I'm happy with the overall state
> of the patchset and the only reason I tag it as RFC is that it is tricky
> to get 57 patches upstream.
>
> I wanted to give a view of the end state first. I will suggest a possible
> way to split it below.
>
> I would appreciate any feedback.
Replying here on the overall design first before reading into the other
discussions (and process related topics).
>
> TL;DR
> =====
>
> This replaces khugepaged's anonymous collapse with an engine that
> can collapse sub-PMD ranges. It is built around migration entries and
> frozen folios instead of heavy locking and isolation, aiming for better
> scalability and less disruption to the workload being collapsed.
I recall us discussing something around using some PTE/PMD markers (e.g.,
migration entries) in the past.
One thing that needed care is handling concurrent MADV_DONTNEED + faultin after
dropping relevant locks.
>
> Why
> ===
>
> mTHP collapse landed in khugepaged in 7.2 and I was glad to see it. We
> at Meta run arm64 with 64K base pages, where a PMD is 512M: PMD-order THP
> is of limited use at that size, and mTHP is exactly what we want.
Right, as the first step, we decided to go for the simpler and minimally
intrusive approach of using the existing mechanism that always operates on PMD
ranges, keeping using the existing pmdp_collapse_flush()-based mechanism and
locking in place.
That's why anything more elaborate will require significantly more LOC :)
>
> It turned out not to help us.
>
> khugepaged only ever looks at PMD-aligned windows, and it is not an easy
> limitation to lift.
I recall we discussed some simpler way to make this work with VMAs that don't
fully span PMDs: I think write-locking VMAs (+ rmap) that cover the PMD was
discussed as a low-hanging fruit, such that other page table walkers would not
suddenly stumble over the temporarily removed page table.
So the mmap_write_lock() + vma write-locks prevents concurrent mmap+page faults
and the rmap locks prevent concurrent rmap walks.
There were discussions on the impact when a PMD spans many VMAs, and I think one
conclusion was that such scenarios are likely not worth considering (e.g., 512
VMAs in a single PMD, all with different rmap locks; your workload sucks already).
>
> Fixing the alignment is a one-line change, but what it feeds assumes the
> PMD everywhere that matters: collapse_huge_page() clears and flushes the
> whole PMD whatever order it is collapsing, installs a PMD leaf because
> that is the only thing it can produce, and keeps everyone out with
> mmap_write_lock, anon_vma_lock_write() and an IPI broadcast while it
> does.
>
> Which is why hugepage_vma_revalidate() demands that the VMA span the
> whole PMD even for an mTHP order -- "we'd need to lock all VMAs in the
> PMD range to support this", as the comment there puts it. A PMD-granular
> operation is only safe when one VMA owns the PMD, and that is exactly the
> restriction in the way. The alignment is the symptom; the PMD is the
> design.
I disagree with "A PMD-granular operation is only safe when one VMA owns the
PMD". It's safe when all page table walkers can be stopped (see above).
>
> So both roots have to go.
>
> Design
> ======
>
> The old mechanism holds the address space still because it has nothing
> else stopping the sources from moving under the copy. The new engine
> makes the sources themselves inert instead, with the two barriers
> migration already uses, raised in that order:
>
> 1. migration entries replace the source PTEs. Faults and GUP-slow
> now wait on the source folio's lock, which is taken before the
> first entry becomes visible.
> 2. the source folio's refcount is frozen to its expected value.
> GUP-fast, pfn walkers, reclaim, compaction and memory-failure all
> fail folio_try_get() and back off.
>
> Between the two, nothing can reach a source, so the copy runs with no
> lock held at all -- and the address space is left alone while it does.
Right. Concurrent MADV_DONTNEED can zap migration PTEs and other faults even
re-fault fresh anon folios. So that must be detected before replacing migration
entries again I guess.
>
> What that removes from every collapse path:
>
> mmap_write_lock -> mmap_read
> anon_vma_lock_write() -> nothing: an rmap walk needs the folio
> locked, and the engine holds that lock
> from freeze to putback
> tlb_remove_table_sync_one() -> nothing: one ranged flush per round
> LRU isolation -> nothing: sources are inert in place
>
Not sure how you handle PMD collapse. I recall problems with migration entries
on the PMD level for non-folio things (we discussed something along these lines
also in the past).
> Working in windows rather than whole PMDs takes care of the other root.
> A sub-PMD window is collapsed under the page table lock, so a collapse
> disturbs only the window it collapses, and each candidate is validated
I recall us discussing that holding the PT lock for a longer collapse operation
(especially on 64k) is problematic. But I don't get all the details from your
description here.
> at its own order -- a window need only fit its own VMA. A PMD-order
> candidate still has to own the whole PMD, which is the old rule kept
> where it is still needed.
>
> Candidates are carried through the passes a batch at a time rather than
> one window at a time, so a round pays for its flush and its lock
> acquisitions once.
Now I am starting to feel that there are too many changes packed in a single
series :)
>
> With the barriers holding the sources still, which read lock the engine
> takes stops being part of the design. A round works inside a single
> VMA, so patches 43-49 switch it from mmap_read to per-VMA locking: an
> mmap_write elsewhere in the mm then stops waiting for a collapse that
> has nothing to do with it. That block is the only part of the series
> that needs per-VMA locking to be unconditional, and it is a separate
> dependency (see below); everything before it runs under mmap_read and
> does not care.
>
> Patch 7 sketches the engine as a comment naming every pass, what lock it
> takes and what it may sleep on; the details are there rather than here.
>
> What falls out beyond the lock diet:
>
> - mTHP collapse in VMAs smaller than a PMD, which is the arm64 case
> above: a 2M VMA on an arm64/64K machine collapses nothing today at
> any order, and collapses to mTHP here.
> - Hole and zeropage population at every order, so partially populated
> windows collapse to mTHP under the same max_ptes_none policy as PMD.
> - Sources come in spans -- any stretch of consecutive PTEs mapping
> consecutive pages of one folio -- so partially mapped and scrambled
> compound sources (the PTE-mapped-THP re-collapse class) work at
> every order.
> - A table that cannot become one huge page still yields the largest
> windows inside it, where before a single disqualified PTE gave up
> the whole table.
>
> Reading the series
> ==================
>
> 57 patches is a lot to land on a list. They go in blocks:
>
> 1-6 helpers and shared state: pte_folio(), pte_none_or_zero(),
> mm/collapse.h, and the policy that replaces asking whether
> khugepaged started a collapse
> 7-8 the engine's shape: entry points, a call-tree comment naming
> every pass, and the scan filled in
> 9-23 the collapse half, top down: the round frame, then each pass
> in turn, then selection and the retry store
I fail to parse this sentence.
> 24 per-candidate tracing, before the switch takes the old
> tracepoints away
> 25-28 the switch: point the anon path at the engine, widen coverage
> to sub-PMD VMAs, delete the mechanism it replaces
> 29-35 move what is left of collapse out of khugepaged.c, and
> MADV_COLLAPSE into madvise.c
> 36-42 tracing: the engine's own events and trace header
> 43-49 per-VMA locking, and the mm reference that makes it safe
> 50-56 selftests for what the engine can now do
> 57 MAINTAINERS
>
> The two patches worth reading first if you read nothing else are 7 (the
> design, as a comment naming the whole call tree) and 16 (the freeze,
> which is where the safety argument lives).
>
> A possible split, if that helps:
>
> 1-2 two mm helpers, pte_folio() and pte_none_or_zero(). Both
> convert callers outside collapse and are useful on their own
> 3-27 the engine and the switch-over. This is the smallest unit
> that does anything: stop earlier and the tree carries an
> engine nothing calls
> 28 remove the mechanism the engine replaces
> 29-42 moving what is left of collapse out of khugepaged.c, and the
> engine's own tracepoints
> 43-49 per-VMA locking
> 50-57 selftests and MAINTAINERS
>
> Keeping the removal separate leaves both engines in the tree with only
> the new one reachable, so the switch can be reverted on its own if
> something turns up. The old mechanism is already carried that way for
> three patches inside the series, so this costs nothing but 975 lines of
> unreferenced code until 28 lands. That safety net only lasts until the
> blocks after it land, though: once collapse has moved out of
> khugepaged.c and the locking has changed, reverting the switch no longer
> gives back a working old engine.
[...]
> Performance
> ===========
>
> Measuring khugepaged is awkward. It is a background daemon, so what
> matters is what a workload feels while it runs, not what the daemon
> reports about itself -- and the usual coverage instrument is no help
> below the PMD: smaps AnonHugePages only counts PMD-order folios, so it
> reads zero however much mTHP has been collapsed.
>
> So I wrote "perf bench mem usemem" for this. It touches a region while
> khugepaged works on it and reports the workload's own latency
> percentiles and throughput, against per-size counters that can see
> sub-PMD folios. The branch is above; it is unposted and not a
> dependency.
Something more realistic might be running some workload in a VM whereby the VM
is getting collapsed by khugepaged.
[...]
> There may be a way out -- a PMD migration entry over the table during
> the window, so the CPU never caches a walk to shoot down -- but that
> means teaching every pmd-level walker a new kind of entry, and I have
> not tried it.
I think we discussed that in the past and it's absolutely nasty.
[...]
> Size
> ====
>
> mm/ grows by 1915 lines net: 4475 added against 2560 deleted.
>
That's quite a lot for something that reads like a cleanup at first.
Okay, let me read the other discussions.
--
Cheers,
David
prev parent reply other threads:[~2026-08-18 13:56 UTC|newest]
Thread overview: 78+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-16 22:45 [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 01/57] mm: add pte_folio() Kiryl Shutsemau
2026-08-18 16:38 ` Rik van Riel
2026-08-18 18:13 ` David Hildenbrand (Arm)
2026-08-18 20:04 ` Rik van Riel
2026-08-18 17:09 ` David Hildenbrand (Arm)
2026-08-18 18:30 ` Lorenzo Stoakes (ARM)
2026-08-16 22:45 ` [RFC PATCH 02/57] mm: add pte_none_or_zero() Kiryl Shutsemau
2026-08-17 17:57 ` David Hildenbrand (Arm)
2026-08-16 22:45 ` [RFC PATCH 03/57] mm/collapse: add collapse.h for the shared collapse state Kiryl Shutsemau
2026-08-18 10:50 ` Lorenzo Stoakes (ARM)
2026-08-16 22:45 ` [RFC PATCH 04/57] mm/collapse: rename mthp_present_ptes to eligible_ptes Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 05/57] mm/collapse: state what a collapse may do in the policy Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 06/57] mm/collapse: move the smallest collapse order to collapse.h Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 07/57] mm/collapse: sketch the new anonymous collapse engine Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 08/57] mm/collapse: scan a table for what a collapse could use Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 09/57] mm/collapse: collect candidate windows into a round Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 10/57] mm/collapse: run a round and feed the outcomes back Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 11/57] mm/collapse: sketch the passes of a round Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 12/57] mm/collapse: allocate a destination per candidate Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 13/57] mm/collapse: revalidate a round against the VMA Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 14/57] mm/collapse: fault the sources in before the freeze Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 15/57] mm/collapse: check what a candidate would freeze Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 16/57] mm/collapse: freeze the sources behind migration entries Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 17/57] mm/collapse: copy the sources into the destinations Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 18/57] mm/collapse: install the destinations at PTE level Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 19/57] mm/collapse: install a PMD leaf as the terminal layer Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 20/57] mm/collapse: put the sources back Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 21/57] mm/collapse: settle whatever the round reached Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 22/57] mm/collapse: walk a table with a selection cursor Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 23/57] mm/collapse: give a refused region a second chance Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 24/57] mm/collapse: report each candidate's outcome to tracing Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 25/57] mm/collapse: collapse anonymous memory with the new engine Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 26/57] mm/collapse: give collapse_single_pmd() the range to work on Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 27/57] mm/collapse: scan the windows a VMA can hold Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 28/57] mm/collapse: remove the mechanism the engine replaces Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 29/57] mm/collapse: move what a collapse is judged on into collapse.c Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 30/57] mm/collapse: name the max_ptes ceiling after collapse Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 31/57] mm/khugepaged: count collapses where khugepaged makes them Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 32/57] mm/collapse: move the file collapse into collapse.c Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 33/57] mm/collapse: split collapse into a scan and a run Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 34/57] mm/collapse: implement MADV_COLLAPSE in madvise.c Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 35/57] mm/madvise: drop MADV_COLLAPSE's redundant mm reference Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 36/57] mm/collapse: report what the scan found Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 37/57] mm/collapse: report what the fault-in pass paid Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 38/57] mm/collapse: report the round, and what it made faulters wait Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 39/57] mm/collapse: name the file collapse's tracepoints after collapse Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 40/57] mm/collapse: remove the tracepoints of the mechanism that is gone Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 41/57] mm/collapse: give collapse its own trace header Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 42/57] mm/collapse: allow error injection into the freeze Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 43/57] mm/khugepaged: check the scan budget before the work, not after Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 44/57] mm/khugepaged: hold the address space open across a scan Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 45/57] mm/collapse: take a per-VMA read lock for the round Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 46/57] mm/khugepaged: scan under a per-VMA read lock Kiryl Shutsemau
2026-08-16 22:45 ` [RFC PATCH 47/57] mm/madvise: collapse " Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 48/57] mm/collapse: assert the mm reference the engine relies on Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 49/57] mm/khugepaged: drop the mmap_lock barrier from __khugepaged_exit() Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 50/57] selftests/mm: attribute collapses by candidate event alone Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 51/57] selftests/mm: cover collapse inside a sub-PMD VMA Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 52/57] selftests/mm: cover a hole-y window in " Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 53/57] selftests/mm: cover collapse of mlocked ranges Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 54/57] selftests/mm: cover collapse beside a MADV_FREE'd page Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 55/57] selftests/mm: cover collapse beside a pinned page Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 56/57] selftests/mm: cover the scaled max_ptes_shared limit Kiryl Shutsemau
2026-08-16 22:46 ` [RFC PATCH 57/57] MAINTAINERS: add an entry for collapse Kiryl Shutsemau
2026-08-17 8:04 ` Lorenzo Stoakes (ARM)
2026-08-17 8:08 ` David Hildenbrand (Arm)
2026-08-17 10:12 ` Kiryl Shutsemau
2026-08-17 2:02 ` [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives Zi Yan
2026-08-17 10:07 ` Kiryl Shutsemau
2026-08-17 8:52 ` Lorenzo Stoakes (ARM)
2026-08-17 13:38 ` Kiryl Shutsemau
2026-08-18 13:06 ` Lorenzo Stoakes (ARM)
2026-08-18 14:12 ` David Hildenbrand (Arm)
2026-08-18 14:33 ` Lorenzo Stoakes (ARM)
2026-08-18 14:15 ` David Hildenbrand (Arm)
2026-08-18 14:41 ` Lorenzo Stoakes (ARM)
2026-08-18 13:55 ` David Hildenbrand (Arm) [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=9f51ac27-24b2-495b-b397-84865f977d24@kernel.org \
--to=david@kernel.org \
--cc=agordeev@linux.ibm.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=bpf@vger.kernel.org \
--cc=dev.jain@arm.com \
--cc=hughd@google.com \
--cc=jannh@google.com \
--cc=kas@kernel.org \
--cc=kirill@shutemov.name \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=ljs@kernel.org \
--cc=mhiramat@kernel.org \
--cc=mhocko@suse.com \
--cc=nico.pache@linux.dev \
--cc=pfalcato@suse.de \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shuah@kernel.org \
--cc=surenb@google.com \
--cc=usama.anjum@arm.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox