* [RFC PATCH v2 0/5] mm: mglru: fix swappiness behavior
@ 2026-07-26 12:21 Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 1/5] mm: mglru: avoid scanning empty generations in scan_folios() Barry Song (Xiaomi)
` (4 more replies)
0 siblings, 5 replies; 6+ messages in thread
From: Barry Song (Xiaomi) @ 2026-07-26 12:21 UTC (permalink / raw)
To: akpm, linux-mm
Cc: hannes, david, mhocko, qi.zheng, shakeel.butt, ljs, kasong,
axelrasmussen, yuanchu, weixugc, linux-kernel, lyugaofei,
stevensd, Barry Song (Xiaomi)
RFC v2:
- Quickly address a few issues raised in the Sashiko comments so
reviewers can ignore v1 and review a cleaner version instead.
https://sashiko.dev/#/patchset/20260726012946.18684-1-baohua@kernel.org
Thanks, Sashiko!
The active/inactive LRU respects swappiness well. Anonymous page
scanning and reclamation increase roughly linearly with
swappiness, while file page scanning and reclamation decrease
accordingly.
For example, when swappiness reaches 200, both pgsteal_file and
pgscan_file drop to zero while building the kernel in a 1 GB
memcg. In contrast, MGLRU shows almost no change across different
swappiness values.
pgsteal_file
Swappiness LRU MGLRU
--------------------------------
1 10567455 763612
36 990706 480205
71 688170 415848
106 446294 386164
141 286307 359196
176 201733 351686
200 0 330093
pgsteal_anon
Swappiness LRU MGLRU
--------------------------------
1 4410548 2726362
36 2465268 2762859
71 2677908 2885124
106 2737227 2841796
141 2984276 3035015
176 3381338 2938302
200 13116359 3113499
pgscan_file
Swappiness LRU MGLRU
--------------------------------
1 17997223 923094
36 1325674 539571
71 852345 464222
106 538207 464477
141 357253 412277
176 217536 399446
200 0 375902
pgscan_anon
Swappiness LRU MGLRU
--------------------------------
1 31639423 5987136
36 23441521 5753224
71 26067110 6101780
106 25619448 5782919
141 26842088 6234264
176 29200021 5980292
200 62193924 6413125
This patchset respects the type selected by positive_ctrl_err(),
which uses swappiness as its gain. It does so by:
1. Avoiding premature fallback to the other type. Only fall back
when reclaim is running at high priority.
2. Running aging when the preferred type has few or no
reclaimable folios, so more folios of that type become
reclaimable.
With this patchset, swappiness starts to behave similarly to the
active/inactive LRU.
pgsteal_file
Swappiness LRU MGLRU MGLRU+Patch
-------------------------------------------------
1 10567455 763612 2276977
36 990706 480205 485771
71 688170 415848 425893
106 446294 386164 391528
141 286307 359196 371844
176 201733 351686 322634
200 0 330093 1343
pgsteal_anon
Swappiness LRU MGLRU MGLRU+Patch
-------------------------------------------------
1 4410548 2726362 2212243
36 2465268 2762859 3036395
71 2677908 2885124 3111459
106 2737227 2841796 3021429
141 2984276 3035015 3341110
176 3381338 2938302 3232294
200 13116359 3113499 14144584
pgscan_file
Swappiness LRU MGLRU MGLRU+Patch
-------------------------------------------------
1 17997223 923094 3243420
36 1325674 539571 595233
71 852345 464222 523441
106 538207 464477 457391
141 357253 412277 434981
176 217536 399446 369302
200 0 375902 1344
pgscan_anon
Swappiness LRU MGLRU MGLRU+Patch
-------------------------------------------------
1 31639423 5987136 4136386
36 23441521 5753224 5489636
71 26067110 6101780 5699571
106 25619448 5782919 5910594
141 26842088 6234264 6511215
176 29200021 5980292 6286023
200 62193924 6413125 22582174
Another possible approach is to decouple anonymous and file-backed
aging by maintaining separate max_seq values for each type. This
allows anonymous and file-backed memory to age and be reclaimed
independently, enabling the swappiness-preferred type to be reclaimed
more aggressively while allowing the other type to lag behind. This
approach has already been adopted by projects such as CachyOS [1] and
Chromium [2].
However, this approach requires substantial changes to MGLRU and
fundamentally alters its design by breaking the shared aging timeline
between anonymous and file-backed memory. This timeline is the
foundation for mechanisms such as the PID controller and the
min_ttl_ms thrashing protection.
That is why this patchset aims to fix the swappiness behavior
without fundamentally changing MGLRU's design, with minimal changes.
[1] https://github.com/firelzrd/re-swappiness
[2] https://chromium.googlesource.com/chromiumos/third_party/kernel/+log/929932351492d01f0aee37a0ac3be8c7bd88f80d
Barry Song (Xiaomi) (4):
mm: mglru: avoid scanning empty generations in scan_folios()
mm: mglru: only fall back when reclaim is running at high priority
mm: mglru: prevent min_seq[type] from pointing to an empty generation
mm: mglru: run aging if the preferred type has no folios in
reclaimable gens
lyugaofei (1):
mm: mglru: run aging when pages are severely imbalanced across gens
include/linux/mmzone.h | 6 ++--
mm/vmscan.c | 66 ++++++++++++++++++++++++++++++++----------
2 files changed, 53 insertions(+), 19 deletions(-)
--
2.39.3 (Apple Git-146)
^ permalink raw reply [flat|nested] 6+ messages in thread
* [RFC PATCH v2 1/5] mm: mglru: avoid scanning empty generations in scan_folios()
2026-07-26 12:21 [RFC PATCH v2 0/5] mm: mglru: fix swappiness behavior Barry Song (Xiaomi)
@ 2026-07-26 12:21 ` Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 2/5] mm: mglru: only fall back when reclaim is running at high priority Barry Song (Xiaomi)
` (3 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Barry Song (Xiaomi) @ 2026-07-26 12:21 UTC (permalink / raw)
To: akpm, linux-mm
Cc: hannes, david, mhocko, qi.zheng, shakeel.butt, ljs, kasong,
axelrasmussen, yuanchu, weixugc, linux-kernel, lyugaofei,
stevensd, Barry Song (Xiaomi)
Commit 16b475d2ac3c ("mm/mglru: avoid reclaim type fall back when
isolation makes no progress") only falls back to the other type when
scanned == 0. However, I have frequently observed cases where
scanned > 0, but the older reclaimable generation becomes empty
after the first scan_folios(). As a result, the second
scan_folios() for the same type performs a redundant scan over an
empty generation.
We can avoid this by checking whether the reclaimable generation has
become empty when scanned < nr_to_scan and we still have fewer than
MIN_LRU_BATCH isolated folios after scan_folios().
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
---
mm/vmscan.c | 11 +++++++----
1 file changed, 7 insertions(+), 4 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 566c4e837c7d..babbce4bbfe8 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -4852,11 +4852,14 @@ static int isolate_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
break;
}
/*
- * If scanned > 0 and isolated == 0, avoid falling back to the
- * other type, as this type remains sufficient. Falling back
- * too readily can disrupt the positive_ctrl_err() bias.
+ * If scanned >= nr_to_scan or isolated >= MIN_LRU_BATCH,
+ * avoid falling back to the other type. The preferred
+ * type is still reclaimable; otherwise, it would have
+ * already run out of reclaimable generations. Falling
+ * back too readily can disrupt the positive_ctrl_err()
+ * bias.
*/
- if (!scanned)
+ if (scanned < nr_to_scan && *isolated < MIN_LRU_BATCH)
type = !type;
}
--
2.39.3 (Apple Git-146)
^ permalink raw reply related [flat|nested] 6+ messages in thread
* [RFC PATCH v2 2/5] mm: mglru: only fall back when reclaim is running at high priority
2026-07-26 12:21 [RFC PATCH v2 0/5] mm: mglru: fix swappiness behavior Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 1/5] mm: mglru: avoid scanning empty generations in scan_folios() Barry Song (Xiaomi)
@ 2026-07-26 12:21 ` Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 3/5] mm: mglru: prevent min_seq[type] from pointing to an empty generation Barry Song (Xiaomi)
` (2 subsequent siblings)
4 siblings, 0 replies; 6+ messages in thread
From: Barry Song (Xiaomi) @ 2026-07-26 12:21 UTC (permalink / raw)
To: akpm, linux-mm
Cc: hannes, david, mhocko, qi.zheng, shakeel.butt, ljs, kasong,
axelrasmussen, yuanchu, weixugc, linux-kernel, lyugaofei,
stevensd, Barry Song (Xiaomi)
Respect the type selected by positive_ctrl_err(), where swappiness
controls the relative weights of SP and PV. Falling back too readily
undermines that bias. Only fall back after making a sufficient effort
to reclaim from the preferred type.
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
---
mm/vmscan.c | 8 ++++++--
1 file changed, 6 insertions(+), 2 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index babbce4bbfe8..4a387cc4145a 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -4857,10 +4857,14 @@ static int isolate_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
* type is still reclaimable; otherwise, it would have
* already run out of reclaimable generations. Falling
* back too readily can disrupt the positive_ctrl_err()
- * bias.
+ * bias. Also, only fall back when reclaim is running at
+ * a high priority.
*/
- if (scanned < nr_to_scan && *isolated < MIN_LRU_BATCH)
+ if (scanned < nr_to_scan && *isolated < MIN_LRU_BATCH) {
+ if (sc->priority > 2)
+ break;
type = !type;
+ }
}
return total_scanned;
--
2.39.3 (Apple Git-146)
^ permalink raw reply related [flat|nested] 6+ messages in thread
* [RFC PATCH v2 3/5] mm: mglru: prevent min_seq[type] from pointing to an empty generation
2026-07-26 12:21 [RFC PATCH v2 0/5] mm: mglru: fix swappiness behavior Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 1/5] mm: mglru: avoid scanning empty generations in scan_folios() Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 2/5] mm: mglru: only fall back when reclaim is running at high priority Barry Song (Xiaomi)
@ 2026-07-26 12:21 ` Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 4/5] mm: mglru: run aging when pages are severely imbalanced across gens Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 5/5] mm: mglru: run aging if the preferred type has no folios in reclaimable gens Barry Song (Xiaomi)
4 siblings, 0 replies; 6+ messages in thread
From: Barry Song (Xiaomi) @ 2026-07-26 12:21 UTC (permalink / raw)
To: akpm, linux-mm
Cc: hannes, david, mhocko, qi.zheng, shakeel.butt, ljs, kasong,
axelrasmussen, yuanchu, weixugc, linux-kernel, lyugaofei,
stevensd, Barry Song (Xiaomi)
In try_to_inc_min_seq(), min_seq[LRU_GEN_ANON] and
min_seq[LRU_GEN_FILE] can be fixed up to point to empty generations
in order to keep their gap within one. As a result, scan_folios()
may repeatedly scan empty generations. This is confusing, as I
observed scan_folios() returning 0 even though the following check in
scan_folios() doesn't take effect:
if (get_nr_gens(lruvec, type) == MIN_NR_GENS)
return 0;
There is no need to adjust min_seq[], since the reclaim logic already
triggers aging when the number of generations reaches
MIN_NR_GENS, and reclaim never reduces it below MIN_NR_GENS.
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
---
include/linux/mmzone.h | 6 ++----
mm/vmscan.c | 10 ----------
2 files changed, 2 insertions(+), 14 deletions(-)
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index a26c8b855222..233d2006a541 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -552,10 +552,8 @@ enum {
* The youngest generation number is stored in max_seq for both anon and file
* types as they are aged on an equal footing. The oldest generation numbers are
* stored in min_seq[] separately for anon and file types so that they can be
- * incremented independently. Ideally min_seq[] are kept in sync when both anon
- * and file types are evictable. However, to adapt to situations like extreme
- * swappiness, they are allowed to be out of sync by at most
- * MAX_NR_GENS-MIN_NR_GENS-1.
+ * incremented independently. For both file and anonymous memory, the minimum
+ * generation must be at least MIN_NR_GENS.
*
* The number of pages in each generation is eventually consistent and therefore
* can be transiently negative when reset_batch_size() is pending.
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 4a387cc4145a..d6bd64b4dced 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -3972,16 +3972,6 @@ static void try_to_inc_min_seq(struct lruvec *lruvec, int swappiness)
if (!seq_inc_flag)
return;
- /* see the comment on lru_gen_folio */
- if (swappiness && swappiness <= MAX_SWAPPINESS) {
- unsigned long seq = lrugen->max_seq - MIN_NR_GENS;
-
- if (min_seq[LRU_GEN_ANON] > seq && min_seq[LRU_GEN_FILE] < seq)
- min_seq[LRU_GEN_ANON] = seq;
- else if (min_seq[LRU_GEN_FILE] > seq && min_seq[LRU_GEN_ANON] < seq)
- min_seq[LRU_GEN_FILE] = seq;
- }
-
for_each_evictable_type(type, swappiness) {
if (min_seq[type] <= lrugen->min_seq[type])
continue;
--
2.39.3 (Apple Git-146)
^ permalink raw reply related [flat|nested] 6+ messages in thread
* [RFC PATCH v2 4/5] mm: mglru: run aging when pages are severely imbalanced across gens
2026-07-26 12:21 [RFC PATCH v2 0/5] mm: mglru: fix swappiness behavior Barry Song (Xiaomi)
` (2 preceding siblings ...)
2026-07-26 12:21 ` [RFC PATCH v2 3/5] mm: mglru: prevent min_seq[type] from pointing to an empty generation Barry Song (Xiaomi)
@ 2026-07-26 12:21 ` Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 5/5] mm: mglru: run aging if the preferred type has no folios in reclaimable gens Barry Song (Xiaomi)
4 siblings, 0 replies; 6+ messages in thread
From: Barry Song (Xiaomi) @ 2026-07-26 12:21 UTC (permalink / raw)
To: akpm, linux-mm
Cc: hannes, david, mhocko, qi.zheng, shakeel.butt, ljs, kasong,
axelrasmussen, yuanchu, weixugc, linux-kernel, lyugaofei,
stevensd, Barry Song
From: lyugaofei <lyugaofei@xiaomi.com>
This partially restores the reclaim behavior introduced in Yu
Zhao's initial MGLRU commit, ac35a4902370 ("mm: multi-gen LRU:
minimal implementation"):
/*
* It's also ideal to spread pages out evenly, i.e., 1/(MIN_NR_GENS+1)
* of the total number of pages for each generation. A reasonable range
* for this average portion is [1/MIN_NR_GENS, 1/(MIN_NR_GENS+2)]. The
* aging cares about the upper bound of hot pages, while the eviction
* cares about the lower bound of cold pages.
*/
if (young * MIN_NR_GENS > total)
return true;
if (old * (MIN_NR_GENS + 2) < total)
return true;
But with a stricter condition: the younger generations must contain
at least four times as many folios as the older generations. This
allows aging to keep folios of the preferred type spread across the
reclaimable generations.
Signed-off-by: lyugaofei <lyugaofei@xiaomi.com>
Co-developed-by: Barry Song (Xiaomi) <baohua@kernel.org>
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
---
mm/vmscan.c | 37 ++++++++++++++++++++++++++++++++++++-
1 file changed, 36 insertions(+), 1 deletion(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index d6bd64b4dced..7c13dedb0b1f 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -4956,9 +4956,40 @@ static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
return scanned;
}
+static bool lru_gen_imbalanced(struct lruvec *lruvec, int type,
+ unsigned long max_seq, unsigned long min_seq,
+ int swappiness)
+{
+ struct lru_gen_folio *lrugen = &lruvec->lrugen;
+ unsigned long young = 0, old = 0, seq;
+
+ /*
+ * reclaim is forced to a single type in those cases, so there is
+ * no need to consider the swappiness bias
+ */
+ if (swappiness == MIN_SWAPPINESS || swappiness > MAX_SWAPPINESS)
+ return false;
+
+ for (seq = min_seq; seq <= max_seq; seq++) {
+ int gen = lru_gen_from_seq(seq);
+ unsigned long size = 0;
+ int zone;
+
+ for (zone = 0; zone < MAX_NR_ZONES; zone++)
+ size += max(READ_ONCE(lrugen->nr_pages[gen][type][zone]), 0L);
+
+ if (seq + MIN_NR_GENS > max_seq)
+ young += size;
+ else
+ old += size;
+ }
+ return young > old * 4;
+}
+
static bool should_run_aging(struct lruvec *lruvec, unsigned long max_seq,
struct scan_control *sc, int swappiness)
{
+ int type = get_type_to_scan(lruvec, swappiness);
DEFINE_MIN_SEQ(lruvec);
/* have to run aging, since eviction is not possible anymore */
@@ -4970,7 +5001,11 @@ static bool should_run_aging(struct lruvec *lruvec, unsigned long max_seq,
return false;
/* better to run aging even though eviction is still possible */
- return evictable_min_seq(min_seq, swappiness) + MIN_NR_GENS == max_seq;
+ if (evictable_min_seq(min_seq, swappiness) + MIN_NR_GENS == max_seq)
+ return true;
+
+ /* Run aging if the preferred type is severely imbalanced across gens */
+ return lru_gen_imbalanced(lruvec, type, max_seq, min_seq[type], swappiness);
}
static long get_nr_to_scan(struct lruvec *lruvec, struct scan_control *sc,
--
2.39.3 (Apple Git-146)
^ permalink raw reply related [flat|nested] 6+ messages in thread
* [RFC PATCH v2 5/5] mm: mglru: run aging if the preferred type has no folios in reclaimable gens
2026-07-26 12:21 [RFC PATCH v2 0/5] mm: mglru: fix swappiness behavior Barry Song (Xiaomi)
` (3 preceding siblings ...)
2026-07-26 12:21 ` [RFC PATCH v2 4/5] mm: mglru: run aging when pages are severely imbalanced across gens Barry Song (Xiaomi)
@ 2026-07-26 12:21 ` Barry Song (Xiaomi)
4 siblings, 0 replies; 6+ messages in thread
From: Barry Song (Xiaomi) @ 2026-07-26 12:21 UTC (permalink / raw)
To: akpm, linux-mm
Cc: hannes, david, mhocko, qi.zheng, shakeel.butt, ljs, kasong,
axelrasmussen, yuanchu, weixugc, linux-kernel, lyugaofei,
stevensd, Barry Song (Xiaomi)
Respect the type selected by positive_ctrl_err(). If there are no
reclaimable gens left for that type, run aging.
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
---
mm/vmscan.c | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 7c13dedb0b1f..d63322cdb1bc 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -4996,6 +4996,10 @@ static bool should_run_aging(struct lruvec *lruvec, unsigned long max_seq,
if (evictable_min_seq(min_seq, swappiness) + MIN_NR_GENS > max_seq)
return true;
+ /* run aging if the preferred type is exhausted */
+ if (min_seq[type] + MIN_NR_GENS > max_seq)
+ return true;
+
/* try to avoid aging, do gentle reclaim at the default priority */
if (sc->priority == DEF_PRIORITY)
return false;
--
2.39.3 (Apple Git-146)
^ permalink raw reply related [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-07-26 12:21 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-26 12:21 [RFC PATCH v2 0/5] mm: mglru: fix swappiness behavior Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 1/5] mm: mglru: avoid scanning empty generations in scan_folios() Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 2/5] mm: mglru: only fall back when reclaim is running at high priority Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 3/5] mm: mglru: prevent min_seq[type] from pointing to an empty generation Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 4/5] mm: mglru: run aging when pages are severely imbalanced across gens Barry Song (Xiaomi)
2026-07-26 12:21 ` [RFC PATCH v2 5/5] mm: mglru: run aging if the preferred type has no folios in reclaimable gens Barry Song (Xiaomi)
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox