All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v2] mm/mglru: Fix young counter undercount for large folios
@ 2026-08-13  6:10 Hui Zhu
  2026-08-13  6:35 ` Barry Song
  0 siblings, 1 reply; 2+ messages in thread
From: Hui Zhu @ 2026-08-13  6:10 UTC (permalink / raw)
  To: Andrew Morton, Johannes Weiner, David Hildenbrand, Michal Hocko,
	Qi Zheng, Shakeel Butt, Lorenzo Stoakes, Kairui Song, Barry Song,
	Axel Rasmussen, Yuanchu Xie, Wei Xu, linux-mm, linux-kernel
  Cc: Hui Zhu, Baolin Wang

From: Hui Zhu <zhuhui@kylinos.cn>

lru_gen_look_around() feeds its local 'young' counter into
suitable_to_scan(), which decides whether the current PMD is added to
the bloom filter and checked again on the next aging round.

The folio triggering the look-around is processed at function entry:
test_and_clear_young_ptes_notify() clears the accessed bits of the nr
PTEs it maps, and the function bails out if none of them is young.  The
loop that follows therefore never recounts this folio, since its
accessed bits are already cleared.  Every other young folio the loop
finds is accounted as a batch (young += nr), where nr is the number of
consecutive PTEs it maps.  The triggering folio, however, still
contributes a fixed young = 1 regardless of its size -- a leftover from
before PTE batching.  A large triggering folio is thus accounted
inconsistently with the rest of the window.

Initialize young to nr so the triggering folio is accounted the same way
as any other young folio batch in the loop.

Note this is a deliberate overestimate, not a measured value.  The
test-and-clear helper only reports whether any of the nr PTEs is young,
not how many were accessed, so the true number of accessed PTEs in a
large folio is unknown and can be smaller than nr.  Counting the full
batch is intentional: the mm core tracks accessed/dirty state per folio,
not per page, so a per-page count is neither obtainable nor meaningful.
The only consumer is suitable_to_scan(), and the bloom filter it feeds
tolerates error.  Overestimating is also the safe direction: at worst a
PMD that saw little access is rescanned, whereas underestimating could
skip rescanning a PMD whose folios are still hot and reclaim them
incorrectly.  (nr here is the PTE batch size, not necessarily
folio_nr_pages().)

Fixes: 56e5b60b2114 ("mm: support batched checking of the young flag for MGLRU")
Signed-off-by: Hui Zhu <zhuhui@kylinos.cn>
Reviewed-by: Baolin Wang <baolin.wang@linux.alibaba.com>
---
Changelog:
v2:
According to the comments of Baolin and Barry, update git commit log.

 mm/vmscan.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/mm/vmscan.c b/mm/vmscan.c
index bc324e37c5f1..264017850a55 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -4192,7 +4192,7 @@ bool lru_gen_look_around(struct page_vma_mapped_walk *pvmw, unsigned int nr)
 	unsigned long end;
 	struct lru_gen_mm_walk *walk;
 	struct folio *last = NULL;
-	int young = 1;
+	int young = nr;
 	pte_t *pte = pvmw->pte;
 	unsigned long addr = pvmw->address;
 	struct vm_area_struct *vma = pvmw->vma;
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-08-13  6:36 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-13  6:10 [PATCH v2] mm/mglru: Fix young counter undercount for large folios Hui Zhu
2026-08-13  6:35 ` Barry Song

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.