From: Baolin Wang <baolin.wang@linux.alibaba.com>
To: Kiryl Shutsemau <kirill@shutemov.name>,
Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>, Zi Yan <ziy@nvidia.com>
Cc: "Kiryl Shutsemau (Meta)" <kas@kernel.org>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
kernel-team@meta.com, "Liam R . Howlett" <liam@infradead.org>,
Nico Pache <nico.pache@linux.dev>,
Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
Barry Song <baohua@kernel.org>, Lance Yang <lance.yang@linux.dev>,
Usama Arif <usama.arif@linux.dev>,
Vlastimil Babka <vbabka@kernel.org>, Jann Horn <jannh@google.com>
Subject: Re: [PATCH v2 08/12] mm/collapse: separate scanning a PTE table from collapsing it
Date: Tue, 15 Sep 2026 19:09:06 +0800 [thread overview]
Message-ID: <1e91c48c-cfdf-4a40-bf1c-88d8e0295c5a@linux.alibaba.com> (raw)
In-Reply-To: <20260910120238.2529819-9-kirill@shutemov.name>
On 9/10/26 8:02 PM, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> A collapse is two jobs. One reads a PTE table under mmap_lock and decides
> whether the range is worth collapsing. The other allocates, isolates,
> copies and flushes, and wants the lock given up first.
>
> collapse_single_pmd() did both, so the boundary between them was somewhere
> in the middle of a function.
>
> Give each half its own function:
>
> - collapse_scan_pmd() scans one table and only reads. The anonymous
> scan that used to carry that name keeps its body as
> collapse_scan_anon_pmd(), and collapse_scan_pmd() is now the entry
> that picks the anonymous or the file side.
>
> - collapse_run_pmd() does the collapse the scan asked for.
> SCAN_SUCCEED from the scan means there is something to run; anything
> else is why there is not.
>
> collapse_single_pmd() is now the two of them with the mmap_lock drop in
> between, so its callers see what they saw before.
>
> What the scan found and the run needs travels in collapse_control. For
> an anonymous table that is the orders and the referenced and swapped-out
> counts. For a file it is the file itself, the offset in it, and whether
> the PMD folio is already in the page cache.
>
> The file side moves with the anonymous one. collapse_scan_file() used to
> run with mmap_lock already given up, and called collapse_file() itself
> when the page cache looked worth it. It now runs under the lock like the
> anonymous scan and only judges; the run does the collapse. A file
> collapse works on the page cache and never sees a VMA, so the scan takes
> the file reference while it still has one and the run gives it back.
>
> That changes what a refused file table costs khugepaged. Every file
> table it scanned used to end its pass over that mm, because the lock had
> been dropped to scan it; now only a table it goes on to collapse does.
>
> Two things on the file side stop being rescanned. When the page cache
> already holds the PMD folio, the scan says so and the run goes straight
> to retracting the PTE table. A run that refuses dirty pages and may
> write them back retries collapse_file() alone. The checks the scan makes
> ahead of it are ones collapse_file() repeats under the page cache lock.
>
> Tracing changes with it. mm_khugepaged_scan_pmd and
> mm_khugepaged_scan_file used to fire after the collapse, so for an
> accepted table their status field carried what the collapse made of it.
> They now fire before it and read SCAN_SUCCEED for an accepted table. What
> the collapse then made of it is for mm_collapse_huge_page and
> mm_khugepaged_collapse_file to report.
>
> Assisted-by: LLM
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> ---
> mm/collapse.h | 16 ++++++
> mm/khugepaged.c | 147 ++++++++++++++++++++++++++++++++++++------------
> 2 files changed, 128 insertions(+), 35 deletions(-)
>
> diff --git a/mm/collapse.h b/mm/collapse.h
> index 7044dc71c7c2..346859a2184f 100644
> --- a/mm/collapse.h
> +++ b/mm/collapse.h
> @@ -88,6 +88,22 @@ struct collapse_control {
>
> /* Each bit marks a PTE the scan accepted as a collapse source */
> DECLARE_BITMAP(eligible_ptes, MAX_PTRS_PER_PTE);
> +
> + /*
> + * What a scan found and the run after it needs. Live only between the
> + * two, and read by nobody else.
> + *
> + * The file side takes a reference while it still has the VMA, since a
> + * file collapse works on the page cache and never sees one; the run is
> + * what gives it back. A scan that found the PMD folio already in the
> + * cache leaves only the PTE table to retract.
> + */
> + unsigned long scan_orders;
> + int scan_referenced;
> + int scan_unmapped;
> + struct file *scan_file;
> + pgoff_t scan_pgoff;
> + bool scan_retract_only;
> };
[snip]
> +
> +static enum scan_result collapse_scan_pmd(struct vm_area_struct *vma,
> + unsigned long addr, struct collapse_control *cc)
> {
> - struct mm_struct *mm = vma->vm_mm;
> - bool triggered_wb = false;
> enum scan_result result;
> - struct file *file;
> pgoff_t pgoff;
>
> - mmap_assert_locked(mm);
> + mmap_assert_locked(vma->vm_mm);
> + /* Whatever the last scan found has to have been run by now */
> + if (WARN_ON_ONCE(cc->scan_file)) {
> + fput(cc->scan_file);
> + cc->scan_file = NULL;
> + }
>
> if (vma_is_anonymous(vma))
> - return collapse_scan_pmd(mm, vma, addr, lock_dropped, cc);
> + return collapse_scan_anon_pmd(vma, addr, cc);
>
> - file = get_file(vma->vm_file);
> pgoff = linear_page_index(vma, addr);
> + result = collapse_scan_file(vma->vm_mm, addr, vma->vm_file, pgoff, cc);
> + switch (result) {
> + case SCAN_SUCCEED:
> + cc->scan_retract_only = false;
> + break;
> + case SCAN_PTE_MAPPED_HUGEPAGE:
> + /*
> + * The page cache already holds the PMD folio; what is left is
> + * to retract the PTE table, which is the run's job.
> + */
> + cc->scan_retract_only = true;
> + result = SCAN_SUCCEED;
> + break;
> + default:
> + return result;
> + }
>
> - mmap_read_unlock(mm);
> - *lock_dropped = true;
> + /*
> + * A file collapse works on the page cache and never sees a VMA, so take
> + * what it needs from this one while it is still here.
> + */
> + cc->scan_file = get_file(vma->vm_file);
> + cc->scan_pgoff = pgoff;
> + return result;
> +}
> +
> +static enum scan_result collapse_run_pmd(struct mm_struct *mm,
> + unsigned long addr, struct collapse_control *cc)
> +{
> + struct file *file = cc->scan_file;
> + bool triggered_wb = false;
> + enum scan_result result;
> + pgoff_t pgoff;
> +
> + if (!file)
> + return mthp_collapse(mm, addr, cc->scan_referenced,
> + cc->scan_unmapped, cc, cc->scan_orders);
You already record scan_referenced, scan_unmapped, and scan_orders in
'cc'. It looks like simply passing 'cc' as a parameter should be enough.
next prev parent reply other threads:[~2026-09-15 11:09 UTC|newest]
Thread overview: 34+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-10 12:02 [PATCH v2 00/12] mm/collapse: separate a collapse from its callers Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 01/12] mm/khugepaged: drop redundant mm_struct pin in madvise_collapse() Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 02/12] mm/khugepaged: count collapses where khugepaged makes them Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 03/12] mm/khugepaged: rename mthp_present_ptes bitmap to eligible_ptes Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 04/12] mm/collapse: add collapse.h for the collapse interface Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 05/12] mm/collapse: state what a collapse may do in the policy Kiryl Shutsemau
2026-09-11 2:06 ` Zi Yan
2026-09-14 8:32 ` Baolin Wang
2026-09-10 12:02 ` [PATCH v2 06/12] mm/collapse: drop the collapse_possible() wrapper Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 07/12] mm/collapse: name the per-table scan reset for what it resets Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 08/12] mm/collapse: separate scanning a PTE table from collapsing it Kiryl Shutsemau
2026-09-11 2:38 ` Zi Yan
2026-09-11 13:37 ` Kiryl Shutsemau
2026-09-11 14:40 ` Zi Yan
2026-09-15 11:09 ` Baolin Wang [this message]
2026-09-15 13:46 ` Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 09/12] mm/collapse: open-code collapse_single_pmd() in its two callers Kiryl Shutsemau
2026-09-11 14:57 ` Zi Yan
2026-09-11 15:24 ` Kiryl Shutsemau
2026-09-11 15:26 ` Zi Yan
2026-09-11 22:09 ` Zi Yan
2026-09-14 11:47 ` Kiryl Shutsemau
2026-09-15 15:51 ` Zi Yan
2026-09-10 12:02 ` [PATCH v2 10/12] mm/collapse: work out the orders a VMA allows once per VMA Kiryl Shutsemau
2026-09-11 15:56 ` Zi Yan
2026-09-15 11:10 ` Baolin Wang
2026-09-10 12:02 ` [PATCH v2 11/12] mm/collapse: declare the collapse interface in collapse.h Kiryl Shutsemau
2026-09-11 19:02 ` Zi Yan
2026-09-14 12:54 ` Kiryl Shutsemau
2026-09-10 12:02 ` [PATCH v2 12/12] mm/collapse: implement MADV_COLLAPSE in madvise.c Kiryl Shutsemau
2026-09-11 15:06 ` [PATCH v2 00/12] mm/collapse: separate a collapse from its callers David Hildenbrand (Arm)
2026-09-11 15:56 ` Kiryl Shutsemau
2026-09-11 15:58 ` Kiryl Shutsemau
2026-09-11 18:35 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1e91c48c-cfdf-4a40-bf1c-88d8e0295c5a@linux.alibaba.com \
--to=baolin.wang@linux.alibaba.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=jannh@google.com \
--cc=kas@kernel.org \
--cc=kernel-team@meta.com \
--cc=kirill@shutemov.name \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=nico.pache@linux.dev \
--cc=ryan.roberts@arm.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.