From: Gregory Price <gourry@gourry.net>
To: linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org,
kernel-team@meta.com, akpm@linux-foundation.org,
liam@infradead.org, ljs@kernel.org, david@kernel.org,
vbabka@kernel.org, jannh@google.com, rppt@kernel.org,
surenb@google.com, mhocko@suse.com, shuah@kernel.org,
"Gregory Price (Meta)" <gourry@gourry.net>
Subject: [PATCH 09/10] mm/madvise: make cold and pageout PTE lock ownership explicit
Date: Tue, 22 Sep 2026 19:58:29 -0400 [thread overview]
Message-ID: <20260922235830.2350770-10-gourry@gourry.net> (raw)
In-Reply-To: <20260922235830.2350770-1-gourry@gourry.net>
The PMD callback still maps and unmaps the PTE table, drives split retries
and handles rescheduling after dispatching huge PMDs. This leaves PTE lock
ownership mixed with the page-walk callback.
Move that lifecycle into madvise_lru_pte_range(). Its outer loop drops the
lock before splitting or yielding and resumes at the current address,
leaving the PMD callback to select the huge-PMD or PTE path.
No functional change intended.
Assisted-by: LLM
Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
---
mm/madvise.c | 94 ++++++++++++++++++++++++++--------------------------
1 file changed, 47 insertions(+), 47 deletions(-)
diff --git a/mm/madvise.c b/mm/madvise.c
index e3c3acfcd9b66..88a4a03dee04c 100644
--- a/mm/madvise.c
+++ b/mm/madvise.c
@@ -553,66 +553,66 @@ madvise_lru_pte_range_locked(pte_t *pte, unsigned long *addr,
return NULL;
}
-static int madvise_lru_pmd_entry(pmd_t *pmd, unsigned long addr,
- unsigned long end, struct mm_walk *walk)
+static void madvise_lru_pte_range(pmd_t *pmd, unsigned long addr,
+ unsigned long end, struct mm_walk *walk,
+ bool pageout_anon_only)
{
- struct madvise_walk_private *private = walk->private;
- struct mmu_gather *tlb = private->tlb;
- bool pageout = private->pageout;
- struct mm_struct *mm = tlb->mm;
- struct vm_area_struct *vma = walk->vma;
+ const struct madvise_walk_private *private = walk->private;
+ struct mm_struct *mm = private->tlb->mm;
+ LIST_HEAD(folio_list);
pte_t *start_pte, *pte;
spinlock_t *ptl;
- struct folio *folio = NULL;
- LIST_HEAD(folio_list);
+ struct folio *folio;
unsigned int batch_count = 0;
- bool pageout_anon_only;
int nr;
- if (fatal_signal_pending(current))
- return -EINTR;
- pageout_anon_only = pageout && !vma_is_anonymous(vma) &&
- !can_do_file_pageout(vma);
+ tlb_change_page_size(private->tlb, PAGE_SIZE);
+ while (addr < end) {
+ start_pte = pte_offset_map_lock(mm, pmd, addr, &ptl);
+ if (!start_pte)
+ break;
+ pte = start_pte;
+ flush_tlb_batched_pending(mm);
+ lazy_mmu_mode_enable();
+ folio = madvise_lru_pte_range_locked(pte, &addr, end, walk,
+ &folio_list, pageout_anon_only, &nr, &batch_count);
- if (pmd_trans_huge(*pmd) &&
- madvise_lru_huge_pmd(pmd, addr, end, walk, pageout_anon_only))
- return 0;
- tlb_change_page_size(tlb, PAGE_SIZE);
-restart:
- start_pte = pte = pte_offset_map_lock(vma->vm_mm, pmd, addr, &ptl);
- if (!start_pte)
- goto out;
- flush_tlb_batched_pending(mm);
- lazy_mmu_mode_enable();
- folio = madvise_lru_pte_range_locked(pte, &addr, end, walk,
- &folio_list, pageout_anon_only, &nr, &batch_count);
- if (!folio && addr < end) {
lazy_mmu_mode_disable();
pte_unmap_unlock(start_pte, ptl);
- cond_resched();
- goto restart;
- }
- if (!folio)
- goto out;
- lazy_mmu_mode_disable();
- pte_unmap_unlock(start_pte, ptl);
- start_pte = NULL;
- if (!split_folio(folio))
- nr = 0;
- folio_unlock(folio);
- folio_put(folio);
- addr += nr * PAGE_SIZE;
- goto restart;
-
-out:
- if (start_pte) {
- lazy_mmu_mode_disable();
- pte_unmap_unlock(start_pte, ptl);
+ if (!folio && addr < end) {
+ cond_resched();
+ continue;
+ }
+ if (folio) {
+ if (!split_folio(folio))
+ nr = 0;
+ folio_unlock(folio);
+ folio_put(folio);
+ addr += nr * PAGE_SIZE;
+ }
}
- if (pageout)
+ if (private->pageout)
reclaim_pages(&folio_list);
cond_resched();
+}
+
+static int madvise_lru_pmd_entry(pmd_t *pmd, unsigned long addr,
+ unsigned long next, struct mm_walk *walk)
+{
+ const struct madvise_walk_private *private = walk->private;
+ struct vm_area_struct *vma = walk->vma;
+ bool pageout_anon_only;
+
+ if (fatal_signal_pending(current))
+ return -EINTR;
+ pageout_anon_only = private->pageout && !vma_is_anonymous(vma) &&
+ !can_do_file_pageout(vma);
+
+ if (pmd_trans_huge(*pmd) &&
+ madvise_lru_huge_pmd(pmd, addr, next, walk, pageout_anon_only))
+ return 0;
+ madvise_lru_pte_range(pmd, addr, next, walk, pageout_anon_only);
return 0;
}
--
2.53.0-Meta
next prev parent reply other threads:[~2026-09-22 23:59 UTC|newest]
Thread overview: 23+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-22 23:58 [PATCH 00/10] mm/madvise: refactor cold and pageout page table walks Gregory Price
2026-09-22 23:58 ` [PATCH 01/10] selftests/mm: exercise MADV_COLD and MADV_PAGEOUT Gregory Price
2026-09-23 14:26 ` Lorenzo Stoakes (ARM)
2026-09-23 14:44 ` Gregory Price
2026-09-23 14:46 ` Lorenzo Stoakes (ARM)
2026-09-24 11:32 ` David Hildenbrand (Arm)
2026-09-24 14:01 ` Gregory Price
2026-09-22 23:58 ` [PATCH 02/10] mm/madvise: name the shared LRU PMD callback Gregory Price
2026-09-23 14:44 ` Lorenzo Stoakes (ARM)
2026-09-22 23:58 ` [PATCH 03/10] mm/madvise: factor shared LRU folio handling Gregory Price
2026-09-23 16:00 ` Lorenzo Stoakes (ARM)
2026-09-22 23:58 ` [PATCH 04/10] mm/madvise: use the PMD softleaf validity helper Gregory Price
2026-09-23 16:02 ` Lorenzo Stoakes (ARM)
2026-09-22 23:58 ` [PATCH 05/10] mm/madvise: factor huge-PMD folio processing Gregory Price
2026-09-23 16:43 ` Lorenzo Stoakes (ARM)
2026-09-23 17:06 ` Gregory Price
2026-09-23 17:14 ` Lorenzo Stoakes (ARM)
2026-09-23 17:26 ` Gregory Price
2026-09-22 23:58 ` [PATCH 06/10] mm/madvise: separate huge PMDs from the PTE walk Gregory Price
2026-09-22 23:58 ` [PATCH 07/10] mm/madvise: separate PTE-batch folio processing Gregory Price
2026-09-22 23:58 ` [PATCH 08/10] mm/madvise: separate the PTL-held PTE scan Gregory Price
2026-09-22 23:58 ` Gregory Price [this message]
2026-09-22 23:58 ` [PATCH 10/10] mm/madvise: share cold and pageout walk setup Gregory Price
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260922235830.2350770-10-gourry@gourry.net \
--to=gourry@gourry.net \
--cc=akpm@linux-foundation.org \
--cc=david@kernel.org \
--cc=jannh@google.com \
--cc=kernel-team@meta.com \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=rppt@kernel.org \
--cc=shuah@kernel.org \
--cc=surenb@google.com \
--cc=vbabka@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox