All of lore.kernel.org
 help / color / mirror / Atom feed
From: Kiryl Shutsemau <kirill@shutemov.name>
To: Zi Yan <ziy@nvidia.com>
Cc: Andrew Morton <akpm@linux-foundation.org>,
	 David Hildenbrand <david@kernel.org>,
	Lorenzo Stoakes <ljs@kernel.org>,
	linux-mm@kvack.org,  linux-kernel@vger.kernel.org,
	kernel-team@meta.com,
	 Baolin Wang <baolin.wang@linux.alibaba.com>,
	"Liam R . Howlett" <liam@infradead.org>,
	 Nico Pache <nico.pache@linux.dev>,
	Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
	 Barry Song <baohua@kernel.org>,
	Lance Yang <lance.yang@linux.dev>,
	 Usama Arif <usama.arif@linux.dev>,
	Vlastimil Babka <vbabka@kernel.org>,
	 Jann Horn <jannh@google.com>
Subject: Re: [PATCH 08/12] mm/collapse: separate scanning a PTE table from collapsing it
Date: Mon, 7 Sep 2026 12:34:53 +0100	[thread overview]
Message-ID: <ap6hM1Mt0PFWloM1@thinkstation> (raw)
In-Reply-To: <DL7VP2ZBCB4J.9NUOEM7GWXR6@nvidia.com>

On Sat, Sep 05, 2026 at 10:30:16PM -0400, Zi Yan wrote:
> On Fri Sep 4, 2026 at 11:10 AM EDT, Kiryl Shutsemau wrote:
> > From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> >
> > A collapse is two jobs.  One reads a PTE table under mmap_lock and decides
> > whether the range is worth collapsing.  The other allocates, isolates,
> > copies and flushes, and wants the lock given up first.
> >
> > collapse_single_pmd() did both, so the boundary between them was somewhere
> > in the middle of a function.
> >
> > Give each half its own function:
> >
> >   - collapse_scan_pmd() scans one table and only reads.  The anonymous
> >     scan that used to carry that name keeps its body as
> >     collapse_scan_anon_pmd(), and collapse_scan_pmd() is now the entry
> >     that picks the anonymous or the file side.
> >
> >   - collapse_run_pmd() does the collapse the scan asked for.
> >     SCAN_SUCCEED from the scan means there is something to run; anything
> >     else is why there is not.
> >
> > collapse_single_pmd() is now the two of them with the mmap_lock drop in
> > between, so its callers see what they saw before.
> >
> > Scan results (beyond SCAN_SUCCEED) communicated via collapse_control
> > structure: the orders, the referenced and swapped-out counts, and for a
> > file the file itself and the offset in it.
> >
> > A file collapse works on the page cache and never sees a VMA.  The scan
> > takes the file reference while it still has VMA and the run unpins it
> > when it is done.
> >
> > Tracing changes with it.  mm_khugepaged_scan_pmd now fires before
> > mm_collapse_huge_page instead of after it.  Its status field already reads
> > SCAN_SUCCEED for an accepted table, so what the collapse then made of that
> > table is mm_collapse_huge_page's to report, per order.
> >
> > The two calls to that tracepoint become one.  They differed in what the
> > collapse between them changed; with the collapse no longer here, both
> > carry the same arguments.  failed_pfn is set only where a PTE was refused,
> > so it is -1 exactly when the result is SCAN_SUCCEED.
> >
> > Assisted-by: Claude-Code:claude-opus-5
> > Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> > ---
> >  mm/collapse.h   |  14 ++++++
> >  mm/khugepaged.c | 121 ++++++++++++++++++++++++++++++++++--------------
> >  2 files changed, 100 insertions(+), 35 deletions(-)
> >
> 
> <snip>
> 
> >  
> > -	mmap_read_unlock(mm);
> > -	*lock_dropped = true;
> > +static enum scan_result collapse_run_pmd(struct mm_struct *mm,
> > +		unsigned long addr, struct collapse_control *cc)
> > +{
> > +	struct file *file = cc->scan_file;
> > +	bool triggered_wb = false;
> > +	enum scan_result result;
> > +	pgoff_t pgoff;
> > +
> > +	if (!file)
> > +		return mthp_collapse(mm, addr, cc->scan_referenced,
> > +				     cc->scan_unmapped, cc, cc->scan_orders);
> > +
> > +	cc->scan_file = NULL;
> > +	pgoff = cc->scan_pgoff;
> >  retry:
> >  	result = collapse_scan_file(mm, addr, file, pgoff, cc);
> 
> In the commit message, collapse_run_pmd() is said to do the collapse
> work, but collapse_scan_file() is scanning, right?
> 
> It seems that the code only separate anonymous scan and collapse.
> Why cannot pagecache code be separated in a similar way?

Will fold the patch below into v2:

diff --git a/mm/collapse.h b/mm/collapse.h
index 4baf2228d2c4..1ebbbf63fb25 100644
--- a/mm/collapse.h
+++ b/mm/collapse.h
@@ -95,13 +95,15 @@ struct collapse_control {
 	 *
 	 * The file side takes a reference while it still has the VMA, since a
 	 * file collapse works on the page cache and never sees one; the run is
-	 * what gives it back.
+	 * what gives it back.  A scan that found the PMD folio already in the
+	 * cache leaves only the PTE table to retract.
 	 */
 	unsigned long scan_orders;
 	int scan_referenced;
 	int scan_unmapped;
 	struct file *scan_file;
 	pgoff_t scan_pgoff;
+	bool scan_retract_only;
 };
 
 /* Which orders a VMA may collapse to, zero when it may not collapse at all */
diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index 4ae292a6392e..403e5fee942d 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -2722,20 +2722,13 @@ static enum scan_result collapse_scan_file(struct mm_struct *mm,
 	else
 		cc->progress += HPAGE_PMD_NR;
 
-	if (result == SCAN_SUCCEED) {
-		if (present < HPAGE_PMD_NR - max_ptes_none) {
-			result = SCAN_EXCEED_NONE_PTE;
-			count_vm_event(THP_SCAN_EXCEED_NONE_PTE);
-		} else {
-			result = collapse_file(mm, addr, file, start, cc);
-		}
-		trace_mm_khugepaged_scan_file(mm, -1, file, present, swap,
-					      SCAN_SUCCEED);
-	} else {
-		trace_mm_khugepaged_scan_file(mm, failed_pfn, file, present,
-					      swap, result);
+	if (result == SCAN_SUCCEED && present < HPAGE_PMD_NR - max_ptes_none) {
+		result = SCAN_EXCEED_NONE_PTE;
+		count_vm_event(THP_SCAN_EXCEED_NONE_PTE);
 	}
 
+	trace_mm_khugepaged_scan_file(mm, failed_pfn, file, present, swap,
+				      result);
 	return result;
 }
 
@@ -2758,6 +2751,9 @@ enum scan_result collapse_scan_pmd(struct vm_area_struct *vma,
 		unsigned long addr, struct collapse_control *cc,
 		unsigned long orders)
 {
+	enum scan_result result;
+	pgoff_t pgoff;
+
 	mmap_assert_locked(vma->vm_mm);
 	/* Whatever the last scan found has to have been run by now */
 	if (WARN_ON_ONCE(cc->scan_file)) {
@@ -2768,14 +2764,31 @@ enum scan_result collapse_scan_pmd(struct vm_area_struct *vma,
 	if (vma_is_anonymous(vma))
 		return collapse_scan_anon_pmd(vma, addr, cc, orders);
 
+	pgoff = linear_page_index(vma, addr);
+	result = collapse_scan_file(vma->vm_mm, addr, vma->vm_file, pgoff, cc);
+	switch (result) {
+	case SCAN_SUCCEED:
+		cc->scan_retract_only = false;
+		break;
+	case SCAN_PTE_MAPPED_HUGEPAGE:
+		/*
+		 * The page cache already holds the PMD folio; what is left is
+		 * to retract the PTE table, which is the run's job.
+		 */
+		cc->scan_retract_only = true;
+		result = SCAN_SUCCEED;
+		break;
+	default:
+		return result;
+	}
+
 	/*
 	 * A file collapse works on the page cache and never sees a VMA, so take
-	 * what it needs from this one while it is still here.  Judging the
-	 * range needs the page cache and no lock, so it happens in the run.
+	 * what it needs from this one while it is still here.
 	 */
 	cc->scan_file = get_file(vma->vm_file);
-	cc->scan_pgoff = linear_page_index(vma, addr);
-	return SCAN_SUCCEED;
+	cc->scan_pgoff = pgoff;
+	return result;
 }
 
 enum scan_result collapse_run_pmd(struct mm_struct *mm, unsigned long addr,
@@ -2792,8 +2805,13 @@ enum scan_result collapse_run_pmd(struct mm_struct *mm, unsigned long addr,
 
 	cc->scan_file = NULL;
 	pgoff = cc->scan_pgoff;
+
+	if (cc->scan_retract_only) {
+		result = SCAN_PTE_MAPPED_HUGEPAGE;
+		goto retract;
+	}
 retry:
-	result = collapse_scan_file(mm, addr, file, pgoff, cc);
+	result = collapse_file(mm, addr, file, pgoff, cc);
 
 	/* Dirty pages are worth a writeback and one more try, if asked for */
 	if (cc->policy.writeback_dirty && result == SCAN_PAGE_DIRTY_OR_WRITEBACK &&
@@ -2805,8 +2823,13 @@ enum scan_result collapse_run_pmd(struct mm_struct *mm, unsigned long addr,
 		triggered_wb = true;
 		goto retry;
 	}
+retract:
 	fput(file);
 
+	/*
+	 * A PMD folio is in the page cache, whether the collapse just put it
+	 * there or found it: retract the PTE table, and map the PMD if asked.
+	 */
 	if (result == SCAN_PTE_MAPPED_HUGEPAGE) {
 		mmap_read_lock(mm);
 		if (collapse_test_exit_or_disable(mm))
-- 
  Kiryl Shutsemau / Kirill A. Shutemov


  reply	other threads:[~2026-09-07 11:35 UTC|newest]

Thread overview: 37+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-04 15:10 [PATCH 00/12] mm/collapse: separate a collapse from its callers Kiryl Shutsemau
2026-09-04 15:10 ` [PATCH 01/12] mm/khugepaged: drop redundant mm_struct pin in madvise_collapse() Kiryl Shutsemau
2026-09-04 15:58   ` Zi Yan
2026-09-07  7:33   ` Baolin Wang
2026-09-04 15:10 ` [PATCH 02/12] mm/khugepaged: count collapses where khugepaged makes them Kiryl Shutsemau
2026-09-05  2:25   ` Zi Yan
2026-09-07  7:40   ` Baolin Wang
2026-09-04 15:10 ` [PATCH 03/12] mm/khugepaged: rename mthp_present_ptes bitmap to eligible_ptes Kiryl Shutsemau
2026-09-05  2:28   ` Zi Yan
2026-09-07  7:54   ` Baolin Wang
2026-09-07 10:35     ` Kiryl Shutsemau
2026-09-04 15:10 ` [PATCH 04/12] mm/collapse: add collapse.h for the collapse interface Kiryl Shutsemau
2026-09-05  2:36   ` Zi Yan
2026-09-07 10:41     ` Kiryl Shutsemau
2026-09-07  8:04   ` Baolin Wang
2026-09-04 15:10 ` [PATCH 05/12] mm/collapse: state what a collapse may do in the policy Kiryl Shutsemau
2026-09-05  2:44   ` Zi Yan
2026-09-07 10:49     ` Kiryl Shutsemau
2026-09-07 19:40       ` Zi Yan
2026-09-07  9:05   ` Baolin Wang
2026-09-07 10:56     ` Kiryl Shutsemau
2026-09-08  1:48       ` Baolin Wang
2026-09-04 15:10 ` [PATCH 06/12] mm/collapse: drop the collapse_possible() wrapper Kiryl Shutsemau
2026-09-05  2:45   ` Zi Yan
2026-09-07  8:28   ` Baolin Wang
2026-09-04 15:10 ` [PATCH 07/12] mm/collapse: name the per-table scan reset for what it resets Kiryl Shutsemau
2026-09-05 18:05   ` Zi Yan
2026-09-07  8:31   ` Baolin Wang
2026-09-04 15:10 ` [PATCH 08/12] mm/collapse: separate scanning a PTE table from collapsing it Kiryl Shutsemau
2026-09-06  2:30   ` Zi Yan
2026-09-07 11:34     ` Kiryl Shutsemau [this message]
2026-09-04 15:10 ` [PATCH 09/12] mm/collapse: open-code collapse_single_pmd() in its two callers Kiryl Shutsemau
2026-09-04 15:10 ` [PATCH 10/12] mm/collapse: work out the orders a VMA allows once per VMA Kiryl Shutsemau
2026-09-04 15:10 ` [PATCH 11/12] mm/collapse: declare the collapse interface in collapse.h Kiryl Shutsemau
2026-09-04 15:10 ` [PATCH 12/12] mm/collapse: implement MADV_COLLAPSE in madvise.c Kiryl Shutsemau
2026-09-06  0:23 ` [PATCH 00/12] mm/collapse: separate a collapse from its callers Andrew Morton
2026-09-07 10:28   ` Kiryl Shutsemau

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ap6hM1Mt0PFWloM1@thinkstation \
    --to=kirill@shutemov.name \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=jannh@google.com \
    --cc=kernel-team@meta.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=nico.pache@linux.dev \
    --cc=ryan.roberts@arm.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.