* [PATCH v2 1/7] mm/mglru: separate folio generation update from LRU accounting
2026-08-27 23:46 [PATCH v2 0/7] mm/mglru: speed up inc_min_seq() and fix cold/hot inversions Barry Song (Xiaomi)
@ 2026-08-27 23:46 ` Barry Song (Xiaomi)
2026-08-27 23:46 ` [PATCH v2 2/7] mm/mglru: batch update lrugen->nr_pages in inc_min_seq() Barry Song (Xiaomi)
` (5 subsequent siblings)
6 siblings, 0 replies; 9+ messages in thread
From: Barry Song (Xiaomi) @ 2026-08-27 23:46 UTC (permalink / raw)
To: akpm, lianux.mm
Cc: axelrasmussen, baolin.wang, baoquan.he, chenridong, david, hannes,
kasong, linux-kernel, linux-mm, ljs, lyugaofei, mhocko, qi.zheng,
shakeel.butt, stevensd, wangzicheng, weixugc, yuanchu, zhangbo56,
Barry Song (Xiaomi), Xueyuan Chen
folio_inc_gen() currently updates both the folio's generation and the
LRU size accounting. This makes it difficult to batch the LRU size updates
when moving multiple folios.
Extract the generation update into __folio_inc_gen(), which only updates
the folio's generation and reports whether the generation was actually
increased. Keep folio_inc_gen() as the wrapper that performs the LRU size
accounting when needed.
This separates the per-folio generation update from LRU accounting and
allows the latter to be batched by subsequent changes.
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
Tested-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
---
mm/vmscan.c | 26 ++++++++++++++++++++------
1 file changed, 20 insertions(+), 6 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index fdd13299a04a..8f187d296b8e 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -3296,20 +3296,21 @@ static int folio_update_gen(struct folio *folio, int gen, const vma_flags_t *vma
}
/* protect pages accessed multiple times through file descriptors */
-static int folio_inc_gen(struct lruvec *lruvec, struct folio *folio)
+static int __folio_inc_gen(struct folio *folio, int old_gen, bool *increased)
{
- int type = folio_is_file_lru(folio);
- struct lru_gen_folio *lrugen = &lruvec->lrugen;
- int new_gen, old_gen = lru_gen_from_seq(lrugen->min_seq[type]);
unsigned long new_flags, old_flags = READ_ONCE(folio->flags.f);
+ int new_gen;
VM_WARN_ON_ONCE_FOLIO(!(old_flags & LRU_GEN_MASK), folio);
do {
new_gen = ((old_flags & LRU_GEN_MASK) >> LRU_GEN_PGOFF) - 1;
/* folio_update_gen() has promoted this page? */
- if (new_gen >= 0 && new_gen != old_gen)
+ if (new_gen >= 0 && new_gen != old_gen) {
+ if (increased)
+ *increased = false;
return new_gen;
+ }
new_gen = (old_gen + 1) % MAX_NR_GENS;
@@ -3317,8 +3318,21 @@ static int folio_inc_gen(struct lruvec *lruvec, struct folio *folio)
new_flags |= (new_gen + 1UL) << LRU_GEN_PGOFF;
} while (!try_cmpxchg(&folio->flags.f, &old_flags, new_flags));
- lru_gen_update_size(lruvec, folio, old_gen, new_gen);
+ if (increased)
+ *increased = true;
+ return new_gen;
+}
+
+static int folio_inc_gen(struct lruvec *lruvec, struct folio *folio)
+{
+ int type = folio_is_file_lru(folio);
+ struct lru_gen_folio *lrugen = &lruvec->lrugen;
+ int new_gen, old_gen = lru_gen_from_seq(lrugen->min_seq[type]);
+ bool gen_increased;
+ new_gen = __folio_inc_gen(folio, old_gen, &gen_increased);
+ if (gen_increased)
+ lru_gen_update_size(lruvec, folio, old_gen, new_gen);
return new_gen;
}
--
2.34.1
^ permalink raw reply related [flat|nested] 9+ messages in thread* [PATCH v2 2/7] mm/mglru: batch update lrugen->nr_pages in inc_min_seq()
2026-08-27 23:46 [PATCH v2 0/7] mm/mglru: speed up inc_min_seq() and fix cold/hot inversions Barry Song (Xiaomi)
2026-08-27 23:46 ` [PATCH v2 1/7] mm/mglru: separate folio generation update from LRU accounting Barry Song (Xiaomi)
@ 2026-08-27 23:46 ` Barry Song (Xiaomi)
2026-08-27 23:47 ` [PATCH v2 3/7] mm/mglru: enhance cold/hot inversion handling " Barry Song (Xiaomi)
` (4 subsequent siblings)
6 siblings, 0 replies; 9+ messages in thread
From: Barry Song (Xiaomi) @ 2026-08-27 23:46 UTC (permalink / raw)
To: akpm, lianux.mm
Cc: axelrasmussen, baolin.wang, baoquan.he, chenridong, david, hannes,
kasong, linux-kernel, linux-mm, ljs, lyugaofei, mhocko, qi.zheng,
shakeel.butt, stevensd, wangzicheng, weixugc, yuanchu, zhangbo56,
Barry Song (Xiaomi), Xueyuan Chen
Currently, folio_inc_gen() updates lrugen->nr_pages for every folio
as it advances generations. Instead, accumulate the size changes
and update lrugen->nr_pages in a batch after scanning the entire
oldest generation, or when the scan stops because remaining reaches
zero.
Since we only move folios from the oldest generation to the second
oldest generation, the active/inactive state cannot change. We can
therefore skip __lru_update_size().
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
Tested-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
---
mm/vmscan.c | 23 ++++++++++++++++++-----
1 file changed, 18 insertions(+), 5 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 8f187d296b8e..07c22d51debd 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -3918,6 +3918,7 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
struct lru_gen_folio *lrugen = &lruvec->lrugen;
int hist = lru_hist_from_seq(lrugen->min_seq[type]);
int new_gen, old_gen = lru_gen_from_seq(lrugen->min_seq[type]);
+ int target_gen = (old_gen + 1) % MAX_NR_GENS;
/* For file type, skip the check if swappiness is anon only */
if (type && (swappiness == SWAPPINESS_ANON_ONLY))
@@ -3927,35 +3928,47 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
if (!type && !swappiness)
goto done;
+ VM_WARN_ON_ONCE(get_nr_gens(lruvec, type) != MAX_NR_GENS);
+ VM_WARN_ON_ONCE(lru_gen_is_active(lruvec, old_gen) !=
+ lru_gen_is_active(lruvec, target_gen));
/* prevent cold/hot inversion if the type is evictable */
for (zone = 0; zone < MAX_NR_ZONES; zone++) {
struct list_head *head = &lrugen->folios[old_gen][type][zone];
+ long delta = 0;
while (!list_empty(head)) {
struct folio *folio = lru_to_folio(head);
+ long nr_pages = folio_nr_pages(folio);
int refs = folio_lru_refs(folio);
bool workingset = folio_test_workingset(folio);
+ bool gen_increased;
VM_WARN_ON_ONCE_FOLIO(folio_test_unevictable(folio), folio);
VM_WARN_ON_ONCE_FOLIO(folio_test_active(folio), folio);
VM_WARN_ON_ONCE_FOLIO(folio_is_file_lru(folio) != type, folio);
VM_WARN_ON_ONCE_FOLIO(folio_zonenum(folio) != zone, folio);
- new_gen = folio_inc_gen(lruvec, folio);
+ new_gen = __folio_inc_gen(folio, old_gen, &gen_increased);
list_move_tail(&folio->lru, &lrugen->folios[new_gen][type][zone]);
-
+ if (gen_increased)
+ delta += nr_pages;
/* don't count the workingset being lazily promoted */
if (refs + workingset != BIT(LRU_REFS_WIDTH) + 1) {
int tier = lru_tier_from_refs(refs, workingset);
- int delta = folio_nr_pages(folio);
WRITE_ONCE(lrugen->protected[hist][type][tier],
- lrugen->protected[hist][type][tier] + delta);
+ lrugen->protected[hist][type][tier] + nr_pages);
}
if (!--remaining)
- return false;
+ break;
}
+ WRITE_ONCE(lrugen->nr_pages[old_gen][type][zone],
+ lrugen->nr_pages[old_gen][type][zone] - delta);
+ WRITE_ONCE(lrugen->nr_pages[target_gen][type][zone],
+ lrugen->nr_pages[target_gen][type][zone] + delta);
+ if (!remaining)
+ return false;
}
done:
reset_ctrl_pos(lruvec, type, true);
--
2.34.1
^ permalink raw reply related [flat|nested] 9+ messages in thread* [PATCH v2 3/7] mm/mglru: enhance cold/hot inversion handling in inc_min_seq()
2026-08-27 23:46 [PATCH v2 0/7] mm/mglru: speed up inc_min_seq() and fix cold/hot inversions Barry Song (Xiaomi)
2026-08-27 23:46 ` [PATCH v2 1/7] mm/mglru: separate folio generation update from LRU accounting Barry Song (Xiaomi)
2026-08-27 23:46 ` [PATCH v2 2/7] mm/mglru: batch update lrugen->nr_pages in inc_min_seq() Barry Song (Xiaomi)
@ 2026-08-27 23:47 ` Barry Song (Xiaomi)
2026-08-27 23:47 ` [PATCH v2 4/7] mm/mglru: exclude folios promoted by aging from protected " Barry Song (Xiaomi)
` (3 subsequent siblings)
6 siblings, 0 replies; 9+ messages in thread
From: Barry Song (Xiaomi) @ 2026-08-27 23:47 UTC (permalink / raw)
To: akpm, lianux.mm
Cc: axelrasmussen, baolin.wang, baoquan.he, chenridong, david, hannes,
kasong, linux-kernel, linux-mm, ljs, lyugaofei, mhocko, qi.zheng,
shakeel.butt, stevensd, wangzicheng, weixugc, yuanchu, zhangbo56,
Barry Song (Xiaomi), Xueyuan Chen
During aging, a folio's generation may already have been updated by
folio_update_gen(), even though it has not yet been moved to the
corresponding generation list. Such folios are hotter than those
already in that generation.
It makes sense for inc_min_seq() to increment the generation of
folios that were never promoted during aging and move them to the
tail of the new oldest generation. However, folios that were already
promoted should instead be moved to the head of their updated
generation, just as sort_folio() does in scan_folios().
Otherwise, promoted folios could end up behind folios that were
never promoted, effectively inverting their hot/cold ordering.
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
Reviewed-by: Kairui Song <kasong@tencent.com>
Tested-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
---
mm/vmscan.c | 7 +++++--
1 file changed, 5 insertions(+), 2 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 07c22d51debd..b10d1703d907 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -3949,9 +3949,12 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
VM_WARN_ON_ONCE_FOLIO(folio_zonenum(folio) != zone, folio);
new_gen = __folio_inc_gen(folio, old_gen, &gen_increased);
- list_move_tail(&folio->lru, &lrugen->folios[new_gen][type][zone]);
- if (gen_increased)
+ if (gen_increased) {
delta += nr_pages;
+ list_move_tail(&folio->lru, &lrugen->folios[new_gen][type][zone]);
+ } else {
+ list_move(&folio->lru, &lrugen->folios[new_gen][type][zone]);
+ }
/* don't count the workingset being lazily promoted */
if (refs + workingset != BIT(LRU_REFS_WIDTH) + 1) {
int tier = lru_tier_from_refs(refs, workingset);
--
2.34.1
^ permalink raw reply related [flat|nested] 9+ messages in thread* [PATCH v2 4/7] mm/mglru: exclude folios promoted by aging from protected in inc_min_seq()
2026-08-27 23:46 [PATCH v2 0/7] mm/mglru: speed up inc_min_seq() and fix cold/hot inversions Barry Song (Xiaomi)
` (2 preceding siblings ...)
2026-08-27 23:47 ` [PATCH v2 3/7] mm/mglru: enhance cold/hot inversion handling " Barry Song (Xiaomi)
@ 2026-08-27 23:47 ` Barry Song (Xiaomi)
2026-08-27 23:47 ` [PATCH v2 5/7] mm/mglru: make LRU folio prefetch helper an inline function Barry Song (Xiaomi)
` (2 subsequent siblings)
6 siblings, 0 replies; 9+ messages in thread
From: Barry Song (Xiaomi) @ 2026-08-27 23:47 UTC (permalink / raw)
To: akpm, lianux.mm
Cc: axelrasmussen, baolin.wang, baoquan.he, chenridong, david, hannes,
kasong, linux-kernel, linux-mm, ljs, lyugaofei, mhocko, qi.zheng,
shakeel.butt, stevensd, wangzicheng, weixugc, yuanchu, zhangbo56,
Barry Song (Xiaomi), Xueyuan Chen
Some folios may have been promoted during aging, so don't count them
as protected, similar to sort_folio().
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
Reviewed-by: Baoquan He <baoquan.he@linux.dev>
Tested-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
---
mm/vmscan.c | 16 ++++++++--------
1 file changed, 8 insertions(+), 8 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index b10d1703d907..2b33d3f44682 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -3952,17 +3952,17 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
if (gen_increased) {
delta += nr_pages;
list_move_tail(&folio->lru, &lrugen->folios[new_gen][type][zone]);
+
+ /* don't count the workingset being lazily promoted */
+ if (refs + workingset != BIT(LRU_REFS_WIDTH) + 1) {
+ int tier = lru_tier_from_refs(refs, workingset);
+
+ WRITE_ONCE(lrugen->protected[hist][type][tier],
+ lrugen->protected[hist][type][tier] + nr_pages);
+ }
} else {
list_move(&folio->lru, &lrugen->folios[new_gen][type][zone]);
}
- /* don't count the workingset being lazily promoted */
- if (refs + workingset != BIT(LRU_REFS_WIDTH) + 1) {
- int tier = lru_tier_from_refs(refs, workingset);
-
- WRITE_ONCE(lrugen->protected[hist][type][tier],
- lrugen->protected[hist][type][tier] + nr_pages);
- }
-
if (!--remaining)
break;
}
--
2.34.1
^ permalink raw reply related [flat|nested] 9+ messages in thread* [PATCH v2 5/7] mm/mglru: make LRU folio prefetch helper an inline function
2026-08-27 23:46 [PATCH v2 0/7] mm/mglru: speed up inc_min_seq() and fix cold/hot inversions Barry Song (Xiaomi)
` (3 preceding siblings ...)
2026-08-27 23:47 ` [PATCH v2 4/7] mm/mglru: exclude folios promoted by aging from protected " Barry Song (Xiaomi)
@ 2026-08-27 23:47 ` Barry Song (Xiaomi)
2026-08-27 23:47 ` [PATCH v2 6/7] mm/mglru: move folios from oldest gen to second-oldest gen from head to tail Barry Song (Xiaomi)
2026-08-27 23:47 ` [PATCH v2 7/7] mm/mglru: batch move folios to the second-oldest gen's LRU Barry Song (Xiaomi)
6 siblings, 0 replies; 9+ messages in thread
From: Barry Song (Xiaomi) @ 2026-08-27 23:47 UTC (permalink / raw)
To: akpm, lianux.mm
Cc: axelrasmussen, baolin.wang, baoquan.he, chenridong, david, hannes,
kasong, linux-kernel, linux-mm, ljs, lyugaofei, mhocko, qi.zheng,
shakeel.butt, stevensd, wangzicheng, weixugc, yuanchu, zhangbo56,
Barry Song (Xiaomi), Xueyuan Chen
`prefetchw_prev_lru_folio()` is currently implemented as a macro with a
potentially unused argument. This makes the helper harder to read and can
also trigger checkpatch warnings about unused macro arguments.
Make it a `static inline` function and remove the unnecessary `_field`
argument. The helper always prefetches the previous folio's `flags`, so
there is no need to make the field configurable.
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
Tested-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
---
mm/vmscan.c | 26 +++++++++++++++-----------
1 file changed, 15 insertions(+), 11 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 2b33d3f44682..81f95a968447 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -182,17 +182,21 @@ struct scan_control {
};
#ifdef ARCH_HAS_PREFETCHW
-#define prefetchw_prev_lru_folio(_folio, _base, _field) \
- do { \
- if ((_folio)->lru.prev != _base) { \
- struct folio *prev; \
- \
- prev = lru_to_folio(&(_folio->lru)); \
- prefetchw(&prev->_field); \
- } \
- } while (0)
+static inline void prefetchw_prev_lru_folio(struct folio *folio,
+ struct list_head *base)
+{
+ if (folio->lru.prev != base) {
+ struct folio *prev;
+
+ prev = lru_to_folio(&folio->lru);
+ prefetchw(&prev->flags);
+ }
+}
#else
-#define prefetchw_prev_lru_folio(_folio, _base, _field) do { } while (0)
+static inline void prefetchw_prev_lru_folio(struct folio *folio,
+ struct list_head *base)
+{
+}
#endif
/*
@@ -1695,7 +1699,7 @@ static unsigned long isolate_lru_folios(unsigned long nr_to_scan,
struct folio *folio;
folio = lru_to_folio(src);
- prefetchw_prev_lru_folio(folio, src, flags);
+ prefetchw_prev_lru_folio(folio, src);
nr_pages = folio_nr_pages(folio);
total_scan += nr_pages;
--
2.34.1
^ permalink raw reply related [flat|nested] 9+ messages in thread* [PATCH v2 6/7] mm/mglru: move folios from oldest gen to second-oldest gen from head to tail
2026-08-27 23:46 [PATCH v2 0/7] mm/mglru: speed up inc_min_seq() and fix cold/hot inversions Barry Song (Xiaomi)
` (4 preceding siblings ...)
2026-08-27 23:47 ` [PATCH v2 5/7] mm/mglru: make LRU folio prefetch helper an inline function Barry Song (Xiaomi)
@ 2026-08-27 23:47 ` Barry Song (Xiaomi)
2026-08-27 23:47 ` [PATCH v2 7/7] mm/mglru: batch move folios to the second-oldest gen's LRU Barry Song (Xiaomi)
6 siblings, 0 replies; 9+ messages in thread
From: Barry Song (Xiaomi) @ 2026-08-27 23:47 UTC (permalink / raw)
To: akpm, lianux.mm
Cc: axelrasmussen, baolin.wang, baoquan.he, chenridong, david, hannes,
kasong, linux-kernel, linux-mm, ljs, lyugaofei, mhocko, qi.zheng,
shakeel.butt, stevensd, wangzicheng, weixugc, yuanchu, zhangbo56,
Barry Song (Xiaomi), Xueyuan Chen
For reclamation, it makes sense to reclaim folios from tail to
head, as folios near the head are relatively hot. However, when
moving folios from the oldest generation to the second-oldest
generation, using the tail-to-head order would effectively cause
a cold/hot inversion.
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
Reviewed-by: Baoquan He <baoquan.he@linux.dev>
Tested-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
---
mm/vmscan.c | 23 +++++++++++++++++++++--
1 file changed, 21 insertions(+), 2 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 81f95a968447..17524e96fe64 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -192,11 +192,27 @@ static inline void prefetchw_prev_lru_folio(struct folio *folio,
prefetchw(&prev->flags);
}
}
+
+static inline void prefetchw_next_lru_folio(struct folio *folio,
+ struct list_head *base)
+{
+ if (folio->lru.next != base) {
+ struct folio *next;
+
+ next = list_entry(folio->lru.next, struct folio, lru);
+ prefetchw(&next->flags);
+ }
+}
#else
static inline void prefetchw_prev_lru_folio(struct folio *folio,
struct list_head *base)
{
}
+
+static inline void prefetchw_next_lru_folio(struct folio *folio,
+ struct list_head *base)
+{
+}
#endif
/*
@@ -3938,10 +3954,11 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
/* prevent cold/hot inversion if the type is evictable */
for (zone = 0; zone < MAX_NR_ZONES; zone++) {
struct list_head *head = &lrugen->folios[old_gen][type][zone];
+ struct list_head *pos = head->next;
long delta = 0;
- while (!list_empty(head)) {
- struct folio *folio = lru_to_folio(head);
+ while (pos != head) {
+ struct folio *folio = list_entry(pos, struct folio, lru);
long nr_pages = folio_nr_pages(folio);
int refs = folio_lru_refs(folio);
bool workingset = folio_test_workingset(folio);
@@ -3952,6 +3969,8 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
VM_WARN_ON_ONCE_FOLIO(folio_is_file_lru(folio) != type, folio);
VM_WARN_ON_ONCE_FOLIO(folio_zonenum(folio) != zone, folio);
+ prefetchw_next_lru_folio(folio, head);
+ pos = pos->next;
new_gen = __folio_inc_gen(folio, old_gen, &gen_increased);
if (gen_increased) {
delta += nr_pages;
--
2.34.1
^ permalink raw reply related [flat|nested] 9+ messages in thread* [PATCH v2 7/7] mm/mglru: batch move folios to the second-oldest gen's LRU
2026-08-27 23:46 [PATCH v2 0/7] mm/mglru: speed up inc_min_seq() and fix cold/hot inversions Barry Song (Xiaomi)
` (5 preceding siblings ...)
2026-08-27 23:47 ` [PATCH v2 6/7] mm/mglru: move folios from oldest gen to second-oldest gen from head to tail Barry Song (Xiaomi)
@ 2026-08-27 23:47 ` Barry Song (Xiaomi)
2026-08-28 3:29 ` Barry Song
6 siblings, 1 reply; 9+ messages in thread
From: Barry Song (Xiaomi) @ 2026-08-27 23:47 UTC (permalink / raw)
To: akpm, lianux.mm
Cc: axelrasmussen, baolin.wang, baoquan.he, chenridong, david, hannes,
kasong, linux-kernel, linux-mm, ljs, lyugaofei, mhocko, qi.zheng,
shakeel.butt, stevensd, wangzicheng, weixugc, yuanchu, zhangbo56,
Barry Song (Xiaomi), Xueyuan Chen
Detect folios that need to move from the oldest generation to
the second-oldest generation, and batch-move them together.
This can significantly reduce the sys time of inc_min_seq(),
especially when the other type is significantly behind the
preferred type.
Assisted-by: gemini:gemini-3.6-flash
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
Reviewed-by: Baoquan He <baoquan.he@linux.dev>
Tested-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
---
mm/vmscan.c | 20 +++++++++++++++++++-
1 file changed, 19 insertions(+), 1 deletion(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 17524e96fe64..2b6f3f05ce60 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -3931,6 +3931,19 @@ static void clear_mm_walk(void)
kfree(walk);
}
+static inline void flush_lru_batch(struct list_head *head, struct list_head **batch_end,
+ struct list_head *dst)
+{
+ LIST_HEAD(movable);
+
+ if (!*batch_end)
+ return;
+
+ list_cut_position(&movable, head, *batch_end);
+ list_splice_tail_init(&movable, dst);
+ *batch_end = NULL;
+}
+
static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
{
int zone;
@@ -3953,8 +3966,10 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
lru_gen_is_active(lruvec, target_gen));
/* prevent cold/hot inversion if the type is evictable */
for (zone = 0; zone < MAX_NR_ZONES; zone++) {
+ struct list_head *target_list = &lrugen->folios[target_gen][type][zone];
struct list_head *head = &lrugen->folios[old_gen][type][zone];
struct list_head *pos = head->next;
+ struct list_head *batch_end = NULL;
long delta = 0;
while (pos != head) {
@@ -3974,7 +3989,7 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
new_gen = __folio_inc_gen(folio, old_gen, &gen_increased);
if (gen_increased) {
delta += nr_pages;
- list_move_tail(&folio->lru, &lrugen->folios[new_gen][type][zone]);
+ batch_end = &folio->lru;
/* don't count the workingset being lazily promoted */
if (refs + workingset != BIT(LRU_REFS_WIDTH) + 1) {
@@ -3984,11 +3999,14 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
lrugen->protected[hist][type][tier] + nr_pages);
}
} else {
+ flush_lru_batch(head, &batch_end, target_list);
list_move(&folio->lru, &lrugen->folios[new_gen][type][zone]);
}
if (!--remaining)
break;
}
+ flush_lru_batch(head, &batch_end, target_list);
+
WRITE_ONCE(lrugen->nr_pages[old_gen][type][zone],
lrugen->nr_pages[old_gen][type][zone] - delta);
WRITE_ONCE(lrugen->nr_pages[target_gen][type][zone],
--
2.34.1
^ permalink raw reply related [flat|nested] 9+ messages in thread* Re: [PATCH v2 7/7] mm/mglru: batch move folios to the second-oldest gen's LRU
2026-08-27 23:47 ` [PATCH v2 7/7] mm/mglru: batch move folios to the second-oldest gen's LRU Barry Song (Xiaomi)
@ 2026-08-28 3:29 ` Barry Song
0 siblings, 0 replies; 9+ messages in thread
From: Barry Song @ 2026-08-28 3:29 UTC (permalink / raw)
To: akpm, lianux.mm
Cc: axelrasmussen, baolin.wang, baoquan.he, chenridong, david, hannes,
kasong, linux-kernel, linux-mm, ljs, lyugaofei, mhocko, qi.zheng,
shakeel.butt, stevensd, wangzicheng, weixugc, yuanchu, zhangbo56,
Xueyuan Chen
On Fri, Aug 28, 2026 at 7:47 AM Barry Song (Xiaomi) <baohua@kernel.org> wrote:
>
> Detect folios that need to move from the oldest generation to
> the second-oldest generation, and batch-move them together.
> This can significantly reduce the sys time of inc_min_seq(),
> especially when the other type is significantly behind the
> preferred type.
>
> Assisted-by: gemini:gemini-3.6-flash
> Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
> Reviewed-by: Baoquan He <baoquan.he@linux.dev>
> Tested-by: Xueyuan Chen <xueyuan.chen21@gmail.com>
> ---
> mm/vmscan.c | 20 +++++++++++++++++++-
> 1 file changed, 19 insertions(+), 1 deletion(-)
>
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index 17524e96fe64..2b6f3f05ce60 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -3931,6 +3931,19 @@ static void clear_mm_walk(void)
> kfree(walk);
> }
>
[...]
> while (pos != head) {
> @@ -3974,7 +3989,7 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
> new_gen = __folio_inc_gen(folio, old_gen, &gen_increased);
> if (gen_increased) {
> delta += nr_pages;
> - list_move_tail(&folio->lru, &lrugen->folios[new_gen][type][zone]);
> + batch_end = &folio->lru;
>
> /* don't count the workingset being lazily promoted */
> if (refs + workingset != BIT(LRU_REFS_WIDTH) + 1) {
> @@ -3984,11 +3999,14 @@ static bool inc_min_seq(struct lruvec *lruvec, int type, int swappiness)
> lrugen->protected[hist][type][tier] + nr_pages);
> }
> } else {
> + flush_lru_batch(head, &batch_end, target_list);
> list_move(&folio->lru, &lrugen->folios[new_gen][type][zone]);
Sashiko says
"Does this change from list_move_tail() to list_move() unintentionally
reverse the relative LRU ordering of lazily promoted folios?
In the baseline code, list_move_tail() was used for all folios, preserving
their relative order when moving them to the new generation's list.
Because the loop in inc_min_seq() processes folios sequentially from head
to tail, inserting them one by one at the head of the new list using
list_move() will reverse their relative order (for example, folios A, B,
and C will end up as C, B, A in the new list).
Could this subtly distort eviction fairness and page replacement optimality
during MGLRU reclaim under memory pressure?"
I feel this comment is not particularly useful. We are focusing on
resolving two specific issues:
1. Make sure promoted folios are always placed ahead of non-promoted
folios. Right now, mainline can place promoted folios after
non-promoted folios.
2. Preserve the existing order of non-promoted folios instead of
reversing them, as mainline currently does.
Sashiko seems to be concerned about the ordering among promoted folios.
I don't think that ordering is particularly meaningful, because we don't
know which folio was actually accessed more recently. Accessed bits are
accumulated over time and handled together. The only way to establish a
meaningful order among promoted folios would be to make every access
trigger a page fault, which is obviously not feasible.
Best Regards
Barry
^ permalink raw reply [flat|nested] 9+ messages in thread