Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Gregory Price <gourry@gourry.net>
To: linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org,
	kernel-team@meta.com, akpm@linux-foundation.org,
	liam@infradead.org, ljs@kernel.org, david@kernel.org,
	vbabka@kernel.org, jannh@google.com, rppt@kernel.org,
	surenb@google.com, mhocko@suse.com, shuah@kernel.org,
	"Gregory Price (Meta)" <gourry@gourry.net>
Subject: [PATCH 09/10] mm/madvise: make cold and pageout PTE lock ownership explicit
Date: Tue, 22 Sep 2026 19:58:29 -0400	[thread overview]
Message-ID: <20260922235830.2350770-10-gourry@gourry.net> (raw)
In-Reply-To: <20260922235830.2350770-1-gourry@gourry.net>

The PMD callback still maps and unmaps the PTE table, drives split retries
and handles rescheduling after dispatching huge PMDs. This leaves PTE lock
ownership mixed with the page-walk callback.

Move that lifecycle into madvise_lru_pte_range(). Its outer loop drops the
lock before splitting or yielding and resumes at the current address,
leaving the PMD callback to select the huge-PMD or PTE path.

No functional change intended.

Assisted-by: LLM
Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
---
 mm/madvise.c | 94 ++++++++++++++++++++++++++--------------------------
 1 file changed, 47 insertions(+), 47 deletions(-)

diff --git a/mm/madvise.c b/mm/madvise.c
index e3c3acfcd9b66..88a4a03dee04c 100644
--- a/mm/madvise.c
+++ b/mm/madvise.c
@@ -553,66 +553,66 @@ madvise_lru_pte_range_locked(pte_t *pte, unsigned long *addr,
 	return NULL;
 }
 
-static int madvise_lru_pmd_entry(pmd_t *pmd, unsigned long addr,
-		unsigned long end, struct mm_walk *walk)
+static void madvise_lru_pte_range(pmd_t *pmd, unsigned long addr,
+		unsigned long end, struct mm_walk *walk,
+		bool pageout_anon_only)
 {
-	struct madvise_walk_private *private = walk->private;
-	struct mmu_gather *tlb = private->tlb;
-	bool pageout = private->pageout;
-	struct mm_struct *mm = tlb->mm;
-	struct vm_area_struct *vma = walk->vma;
+	const struct madvise_walk_private *private = walk->private;
+	struct mm_struct *mm = private->tlb->mm;
+	LIST_HEAD(folio_list);
 	pte_t *start_pte, *pte;
 	spinlock_t *ptl;
-	struct folio *folio = NULL;
-	LIST_HEAD(folio_list);
+	struct folio *folio;
 	unsigned int batch_count = 0;
-	bool pageout_anon_only;
 	int nr;
 
-	if (fatal_signal_pending(current))
-		return -EINTR;
-	pageout_anon_only = pageout && !vma_is_anonymous(vma) &&
-				       !can_do_file_pageout(vma);
+	tlb_change_page_size(private->tlb, PAGE_SIZE);
+	while (addr < end) {
+		start_pte = pte_offset_map_lock(mm, pmd, addr, &ptl);
+		if (!start_pte)
+			break;
+		pte = start_pte;
+		flush_tlb_batched_pending(mm);
+		lazy_mmu_mode_enable();
+		folio = madvise_lru_pte_range_locked(pte, &addr, end, walk,
+				&folio_list, pageout_anon_only, &nr, &batch_count);
 
-	if (pmd_trans_huge(*pmd) &&
-	    madvise_lru_huge_pmd(pmd, addr, end, walk, pageout_anon_only))
-		return 0;
-	tlb_change_page_size(tlb, PAGE_SIZE);
-restart:
-	start_pte = pte = pte_offset_map_lock(vma->vm_mm, pmd, addr, &ptl);
-	if (!start_pte)
-		goto out;
-	flush_tlb_batched_pending(mm);
-	lazy_mmu_mode_enable();
-	folio = madvise_lru_pte_range_locked(pte, &addr, end, walk,
-			&folio_list, pageout_anon_only, &nr, &batch_count);
-	if (!folio && addr < end) {
 		lazy_mmu_mode_disable();
 		pte_unmap_unlock(start_pte, ptl);
-		cond_resched();
-		goto restart;
-	}
-	if (!folio)
-		goto out;
 
-	lazy_mmu_mode_disable();
-	pte_unmap_unlock(start_pte, ptl);
-	start_pte = NULL;
-	if (!split_folio(folio))
-		nr = 0;
-	folio_unlock(folio);
-	folio_put(folio);
-	addr += nr * PAGE_SIZE;
-	goto restart;
-
-out:
-	if (start_pte) {
-		lazy_mmu_mode_disable();
-		pte_unmap_unlock(start_pte, ptl);
+		if (!folio && addr < end) {
+			cond_resched();
+			continue;
+		}
+		if (folio) {
+			if (!split_folio(folio))
+				nr = 0;
+			folio_unlock(folio);
+			folio_put(folio);
+			addr += nr * PAGE_SIZE;
+		}
 	}
-	if (pageout)
+	if (private->pageout)
 		reclaim_pages(&folio_list);
 	cond_resched();
+}
+
+static int madvise_lru_pmd_entry(pmd_t *pmd, unsigned long addr,
+		unsigned long next, struct mm_walk *walk)
+{
+	const struct madvise_walk_private *private = walk->private;
+	struct vm_area_struct *vma = walk->vma;
+	bool pageout_anon_only;
+
+	if (fatal_signal_pending(current))
+		return -EINTR;
+	pageout_anon_only = private->pageout && !vma_is_anonymous(vma) &&
+			    !can_do_file_pageout(vma);
+
+	if (pmd_trans_huge(*pmd) &&
+	    madvise_lru_huge_pmd(pmd, addr, next, walk, pageout_anon_only))
+		return 0;
+	madvise_lru_pte_range(pmd, addr, next, walk, pageout_anon_only);
 
 	return 0;
 }
-- 
2.53.0-Meta



  parent reply	other threads:[~2026-09-22 23:59 UTC|newest]

Thread overview: 23+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 23:58 [PATCH 00/10] mm/madvise: refactor cold and pageout page table walks Gregory Price
2026-09-22 23:58 ` [PATCH 01/10] selftests/mm: exercise MADV_COLD and MADV_PAGEOUT Gregory Price
2026-09-23 14:26   ` Lorenzo Stoakes (ARM)
2026-09-23 14:44     ` Gregory Price
2026-09-23 14:46       ` Lorenzo Stoakes (ARM)
2026-09-24 11:32         ` David Hildenbrand (Arm)
2026-09-24 14:01           ` Gregory Price
2026-09-22 23:58 ` [PATCH 02/10] mm/madvise: name the shared LRU PMD callback Gregory Price
2026-09-23 14:44   ` Lorenzo Stoakes (ARM)
2026-09-22 23:58 ` [PATCH 03/10] mm/madvise: factor shared LRU folio handling Gregory Price
2026-09-23 16:00   ` Lorenzo Stoakes (ARM)
2026-09-22 23:58 ` [PATCH 04/10] mm/madvise: use the PMD softleaf validity helper Gregory Price
2026-09-23 16:02   ` Lorenzo Stoakes (ARM)
2026-09-22 23:58 ` [PATCH 05/10] mm/madvise: factor huge-PMD folio processing Gregory Price
2026-09-23 16:43   ` Lorenzo Stoakes (ARM)
2026-09-23 17:06     ` Gregory Price
2026-09-23 17:14       ` Lorenzo Stoakes (ARM)
2026-09-23 17:26         ` Gregory Price
2026-09-22 23:58 ` [PATCH 06/10] mm/madvise: separate huge PMDs from the PTE walk Gregory Price
2026-09-22 23:58 ` [PATCH 07/10] mm/madvise: separate PTE-batch folio processing Gregory Price
2026-09-22 23:58 ` [PATCH 08/10] mm/madvise: separate the PTL-held PTE scan Gregory Price
2026-09-22 23:58 ` Gregory Price [this message]
2026-09-22 23:58 ` [PATCH 10/10] mm/madvise: share cold and pageout walk setup Gregory Price

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260922235830.2350770-10-gourry@gourry.net \
    --to=gourry@gourry.net \
    --cc=akpm@linux-foundation.org \
    --cc=david@kernel.org \
    --cc=jannh@google.com \
    --cc=kernel-team@meta.com \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=rppt@kernel.org \
    --cc=shuah@kernel.org \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox