* [RFC PATCH] mm/mglru: Avoid reclaiming kept folios during retrying missed rotated folios
@ 2026-10-03 5:42 Barry Song (Xiaomi)
2026-10-08 6:52 ` Baolin Wang
0 siblings, 1 reply; 5+ messages in thread
From: Barry Song (Xiaomi) @ 2026-10-03 5:42 UTC (permalink / raw)
To: linux-mm
Cc: akpm, kasong, qi.zheng, shakeel.butt, baohua, axelrasmussen,
yuanchu.xie, weixugc, baoquan.he, baolin.wang, chenridong,
Yu Zhao, Bo Zhang
Since commit 4d5d14a01e2c ("mm/mglru: rework workingset protection"),
`folio_test_referenced()` was accidentally removed when checking for
missed rotated folios. As a result, kept folios without an active set
could be added to the retry list and eventually reclaimed.
The impact should be very small, as we have the `folio_mapped(folio)`
check before retrying. Kept folios won't have `try_to_unmap()` called,
so this is really only applicable to folios that are unmapped by users
after `folio_referenced()` has been called.
I'm not sure if anyone else has seen any issues with this. Sending an
RFC to check whether it is worth fixing.
Fixes: 4d5d14a01e2c ("mm/mglru: rework workingset protection")
Cc: Yu Zhao <yuzhao@google.com>
Reported-by: Bo Zhang <zhangbo56@xiaomi.com>
Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org>
---
mm/vmscan.c | 7 +++++--
1 file changed, 5 insertions(+), 2 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 91295070ca33..6a931b576b7d 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -994,8 +994,10 @@ static enum folio_references folio_check_references(struct folio *folio,
return FOLIOREF_KEEP;
if (lru_gen_enabled() && !lru_gen_switching()) {
- if (!referenced_ptes)
+ if (!referenced_ptes) {
+ folio_clear_referenced(folio);
return FOLIOREF_RECLAIM;
+ }
return lru_gen_set_refs(folio, &vma_flags) ? FOLIOREF_ACTIVATE : FOLIOREF_KEEP;
}
@@ -5110,7 +5112,8 @@ static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec,
continue;
/* retry folios that may have missed folio_rotate_reclaimable() */
- if (!skip_retry && !folio_test_active(folio) && !folio_mapped(folio) &&
+ if (!skip_retry && !folio_test_active(folio) &&
+ !folio_test_referenced(folio) && !folio_mapped(folio) &&
!folio_test_dirty(folio) && !folio_test_writeback(folio)) {
list_move(&folio->lru, &clean);
continue;
--
2.39.3 (Apple Git-146)
^ permalink raw reply related [flat|nested] 5+ messages in thread* Re: [RFC PATCH] mm/mglru: Avoid reclaiming kept folios during retrying missed rotated folios 2026-10-03 5:42 [RFC PATCH] mm/mglru: Avoid reclaiming kept folios during retrying missed rotated folios Barry Song (Xiaomi) @ 2026-10-08 6:52 ` Baolin Wang 2026-10-08 11:43 ` Kairui Song 0 siblings, 1 reply; 5+ messages in thread From: Baolin Wang @ 2026-10-08 6:52 UTC (permalink / raw) To: Barry Song (Xiaomi), linux-mm Cc: akpm, kasong, qi.zheng, shakeel.butt, axelrasmussen, yuanchu.xie, weixugc, baoquan.he, chenridong, Yu Zhao, Bo Zhang Hi Barry, On 10/3/26 1:42 PM, Barry Song (Xiaomi) wrote: > Since commit 4d5d14a01e2c ("mm/mglru: rework workingset protection"), > `folio_test_referenced()` was accidentally removed when checking for > missed rotated folios. As a result, kept folios without an active set > could be added to the retry list and eventually reclaimed. > > The impact should be very small, as we have the `folio_mapped(folio)` > check before retrying. Kept folios won't have `try_to_unmap()` called, > so this is really only applicable to folios that are unmapped by users > after `folio_referenced()` has been called. > > I'm not sure if anyone else has seen any issues with this. Sending an > RFC to check whether it is worth fixing. > > Fixes: 4d5d14a01e2c ("mm/mglru: rework workingset protection") > Cc: Yu Zhao <yuzhao@google.com> > Reported-by: Bo Zhang <zhangbo56@xiaomi.com> > Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org> > --- > mm/vmscan.c | 7 +++++-- > 1 file changed, 5 insertions(+), 2 deletions(-) > > diff --git a/mm/vmscan.c b/mm/vmscan.c > index 91295070ca33..6a931b576b7d 100644 > --- a/mm/vmscan.c > +++ b/mm/vmscan.c > @@ -994,8 +994,10 @@ static enum folio_references folio_check_references(struct folio *folio, > return FOLIOREF_KEEP; > > if (lru_gen_enabled() && !lru_gen_switching()) { > - if (!referenced_ptes) > + if (!referenced_ptes) { > + folio_clear_referenced(folio); > return FOLIOREF_RECLAIM; > + } I've also been investigating the issue of referenced folios being reclaimed recently. The background is that classical LRU will return FOLIOREF_KEEP for folios accessed once via page table, giving the accessed folio another chance. However, for MGLRU, if the folio's pagetable access has already been consumed by aging or lru_gen_look_around(), referenced_ptes may be 0 at that point, leading to reclaiming folios that were accessed via page table, which could cause more refaults (I don't have data yet). But it's not easy to maintain logic consistent with classical LRU, because we cannot distinguish whether this one access count comes from page table or file descriptors (for anonymous pages, it can be distinguished). I'm not sure if Kairui's MGLRU-FG patchset helps with this case (I haven't looked at it carefully yet). Perhaps at least for anonymous pages, if the referenced flag is set, we should return FOLIOREF_KEEP to give it another scan chance. > return lru_gen_set_refs(folio, &vma_flags) ? FOLIOREF_ACTIVATE : FOLIOREF_KEEP; > } > @@ -5110,7 +5112,8 @@ static int evict_folios(unsigned long nr_to_scan, struct lruvec *lruvec, > continue; > > /* retry folios that may have missed folio_rotate_reclaimable() */ > - if (!skip_retry && !folio_test_active(folio) && !folio_mapped(folio) && > + if (!skip_retry && !folio_test_active(folio) && > + !folio_test_referenced(folio) && !folio_mapped(folio) && This part makes sense to me. > !folio_test_dirty(folio) && !folio_test_writeback(folio)) { > list_move(&folio->lru, &clean); > continue; ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [RFC PATCH] mm/mglru: Avoid reclaiming kept folios during retrying missed rotated folios 2026-10-08 6:52 ` Baolin Wang @ 2026-10-08 11:43 ` Kairui Song 2026-10-08 11:52 ` Kairui Song 0 siblings, 1 reply; 5+ messages in thread From: Kairui Song @ 2026-10-08 11:43 UTC (permalink / raw) To: Baolin Wang Cc: Barry Song (Xiaomi), linux-mm, akpm, kasong, qi.zheng, shakeel.butt, axelrasmussen, yuanchu.xie, weixugc, baoquan.he, chenridong, Yu Zhao, Bo Zhang On Thu, Oct 08, 2026 at 02:52:35PM +0100, Baolin Wang wrote: > Hi Barry, > > On 10/3/26 1:42 PM, Barry Song (Xiaomi) wrote: > > Since commit 4d5d14a01e2c ("mm/mglru: rework workingset protection"), > > `folio_test_referenced()` was accidentally removed when checking for > > missed rotated folios. As a result, kept folios without an active set > > could be added to the retry list and eventually reclaimed. > > > > The impact should be very small, as we have the `folio_mapped(folio)` > > check before retrying. Kept folios won't have `try_to_unmap()` called, > > so this is really only applicable to folios that are unmapped by users > > after `folio_referenced()` has been called. > > > > I'm not sure if anyone else has seen any issues with this. Sending an > > RFC to check whether it is worth fixing. > > > > Fixes: 4d5d14a01e2c ("mm/mglru: rework workingset protection") > > Cc: Yu Zhao <yuzhao@google.com> > > Reported-by: Bo Zhang <zhangbo56@xiaomi.com> > > Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org> > > --- > > mm/vmscan.c | 7 +++++-- > > 1 file changed, 5 insertions(+), 2 deletions(-) > > > > diff --git a/mm/vmscan.c b/mm/vmscan.c > > index 91295070ca33..6a931b576b7d 100644 > > --- a/mm/vmscan.c > > +++ b/mm/vmscan.c > > @@ -994,8 +994,10 @@ static enum folio_references folio_check_references(struct folio *folio, > > return FOLIOREF_KEEP; > > if (lru_gen_enabled() && !lru_gen_switching()) { > > - if (!referenced_ptes) > > + if (!referenced_ptes) { > > + folio_clear_referenced(folio); > > return FOLIOREF_RECLAIM; > > + } > > I've also been investigating the issue of referenced folios being reclaimed > recently. The background is that classical LRU will return FOLIOREF_KEEP for > folios accessed once via page table, giving the accessed folio another > chance. Hi, thanks for looking into this. > > However, for MGLRU, if the folio's pagetable access has already been > consumed by aging or lru_gen_look_around(), referenced_ptes may be 0 at that > point, leading to reclaiming folios that were accessed via page table, which > could cause more refaults (I don't have data yet). If we go with MGLRU-FG, aging or lru_gen_look_around may move the folio to newer gen which is identical to the behavior of promoting at eviction time. > But it's not easy to maintain logic consistent with classical LRU, because > we cannot distinguish whether this one access count comes from page table or > file descriptors (for anonymous pages, it can be distinguished). I'm not > sure if Kairui's MGLRU-FG patchset helps with this case (I haven't looked at > it carefully yet). It definatly helps, check this one: https://lore.kernel.org/linux-mm/20261003-mglru-fg-v3-4-cbd4546a5bd9@tencent.com/ folio_inc_lru_refs and variants will promote the folios in a consistent way. See "First access only defers eviction from the oldest gen" part, I believe that is the issue you worries about? BTW promote then to newest gen directly does over protect them, I've see regression with that. > > Perhaps at least for anonymous pages, if the referenced flag is set, we > should return FOLIOREF_KEEP to give it another scan chance. No I think we really shouldn't do that, it will over protect single used folios. Promoting them to second newest gen or second oldest gen seems a good idea and similiar to CLRU, and this is the behavior of MGLRU-FG. ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [RFC PATCH] mm/mglru: Avoid reclaiming kept folios during retrying missed rotated folios 2026-10-08 11:43 ` Kairui Song @ 2026-10-08 11:52 ` Kairui Song 2026-10-08 12:23 ` Kairui Song 0 siblings, 1 reply; 5+ messages in thread From: Kairui Song @ 2026-10-08 11:52 UTC (permalink / raw) To: Baolin Wang Cc: Barry Song (Xiaomi), linux-mm, akpm, qi.zheng, shakeel.butt, axelrasmussen, yuanchu.xie, weixugc, baoquan.he, chenridong, Yu Zhao, Bo Zhang On Thu, Oct 8, 2026 at 1:43 PM Kairui Song <ryncsn@gmail.com> wrote: > > On Thu, Oct 08, 2026 at 02:52:35PM +0100, Baolin Wang wrote: > > Hi Barry, > > > > On 10/3/26 1:42 PM, Barry Song (Xiaomi) wrote: > > > Since commit 4d5d14a01e2c ("mm/mglru: rework workingset protection"), > > > `folio_test_referenced()` was accidentally removed when checking for > > > missed rotated folios. As a result, kept folios without an active set > > > could be added to the retry list and eventually reclaimed. > > > > > > The impact should be very small, as we have the `folio_mapped(folio)` > > > check before retrying. Kept folios won't have `try_to_unmap()` called, > > > so this is really only applicable to folios that are unmapped by users > > > after `folio_referenced()` has been called. > > > > > > I'm not sure if anyone else has seen any issues with this. Sending an > > > RFC to check whether it is worth fixing. > > > > > > Fixes: 4d5d14a01e2c ("mm/mglru: rework workingset protection") > > > Cc: Yu Zhao <yuzhao@google.com> > > > Reported-by: Bo Zhang <zhangbo56@xiaomi.com> > > > Signed-off-by: Barry Song (Xiaomi) <baohua@kernel.org> > > > --- > > > mm/vmscan.c | 7 +++++-- > > > 1 file changed, 5 insertions(+), 2 deletions(-) > > > > > > diff --git a/mm/vmscan.c b/mm/vmscan.c > > > index 91295070ca33..6a931b576b7d 100644 > > > --- a/mm/vmscan.c > > > +++ b/mm/vmscan.c > > > @@ -994,8 +994,10 @@ static enum folio_references folio_check_references(struct folio *folio, > > > return FOLIOREF_KEEP; > > > if (lru_gen_enabled() && !lru_gen_switching()) { > > > - if (!referenced_ptes) > > > + if (!referenced_ptes) { > > > + folio_clear_referenced(folio); > > > return FOLIOREF_RECLAIM; > > > + } > > > > I've also been investigating the issue of referenced folios being reclaimed > > recently. The background is that classical LRU will return FOLIOREF_KEEP for > > folios accessed once via page table, giving the accessed folio another > > chance. > > Hi, thanks for looking into this. > > > > > However, for MGLRU, if the folio's pagetable access has already been > > consumed by aging or lru_gen_look_around(), referenced_ptes may be 0 at that > > point, leading to reclaiming folios that were accessed via page table, which > > could cause more refaults (I don't have data yet). > > If we go with MGLRU-FG, aging or lru_gen_look_around may move the folio > to newer gen which is identical to the behavior of promoting at eviction > time. > > > But it's not easy to maintain logic consistent with classical LRU, because > > we cannot distinguish whether this one access count comes from page table or > > file descriptors (for anonymous pages, it can be distinguished). I'm not > > sure if Kairui's MGLRU-FG patchset helps with this case (I haven't looked at > > it carefully yet). > > It definatly helps, check this one: > https://lore.kernel.org/linux-mm/20261003-mglru-fg-v3-4-cbd4546a5bd9@tencent.com/ > > folio_inc_lru_refs and variants will promote the folios in a consistent way. > See "First access only defers eviction from the oldest gen" part, I believe > that is the issue you worries about? > > BTW promote then to newest gen directly does over protect them, I've see > regression with that. > > > > > Perhaps at least for anonymous pages, if the referenced flag is set, we > > should return FOLIOREF_KEEP to give it another scan chance. > > No I think we really shouldn't do that, it will over protect single used > folios. Promoting them to second newest gen or second oldest gen seems > a good idea and similiar to CLRU, and this is the behavior of MGLRU-FG. Oh I also mean that protecting them by lookaround / walking, and not at eviction seems better? As the rmap here have higher overhead and can see stall hotness info. ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [RFC PATCH] mm/mglru: Avoid reclaiming kept folios during retrying missed rotated folios 2026-10-08 11:52 ` Kairui Song @ 2026-10-08 12:23 ` Kairui Song 0 siblings, 0 replies; 5+ messages in thread From: Kairui Song @ 2026-10-08 12:23 UTC (permalink / raw) To: Baolin Wang Cc: Barry Song (Xiaomi), linux-mm, akpm, qi.zheng, shakeel.butt, axelrasmussen, yuanchu.xie, weixugc, baoquan.he, chenridong, Yu Zhao, Bo Zhang On Thu, Oct 8, 2026 at 1:52 PM Kairui Song <ryncsn@gmail.com> wrote: > > On Thu, Oct 8, 2026 at 1:43 PM Kairui Song <ryncsn@gmail.com> wrote: > > It definatly helps, check this one: > > https://lore.kernel.org/linux-mm/20261003-mglru-fg-v3-4-cbd4546a5bd9@tencent.com/ > > > > folio_inc_lru_refs and variants will promote the folios in a consistent way. > > See "First access only defers eviction from the oldest gen" part, I believe > > that is the issue you worries about? > > > > BTW promote then to newest gen directly does over protect them, I've see > > regression with that. > > > > > > > > Perhaps at least for anonymous pages, if the referenced flag is set, we > > > should return FOLIOREF_KEEP to give it another scan chance. > > > > No I think we really shouldn't do that, it will over protect single used > > folios. Promoting them to second newest gen or second oldest gen seems > > a good idea and similiar to CLRU, and this is the behavior of MGLRU-FG. > > Oh I also mean that protecting them by lookaround / walking, and not > at eviction seems better? As the rmap here have higher overhead and > can see stall hotness info. I also found this seems to work much better too on top of MGLRU-FG v3: Will send a V4 later, the tricky part here is that oldest gen could be very short so it should cover second oldest gen as well. And swap readaheads are mostly unreferenced. diff --git a/include/linux/mm_inline.h b/include/linux/mm_inline.h index e651f60cb914..7fb9c7e85732 100644 --- a/include/linux/mm_inline.h +++ b/include/linux/mm_inline.h @@ -311,6 +311,13 @@ static inline int folio_lru_gen(const struct folio *folio) return lru_get_gen_flags(READ_ONCE(*const_folio_flags(folio, 0))); } +static inline bool lru_gen_is_old(unsigned long max_seq, int gen) +{ + VM_WARN_ON_ONCE(gen >= MAX_NR_GENS); + /* see the comment on MIN_NR_GENS */ + return gen == lru_gen_from_seq(max_seq) || gen == lru_gen_from_seq(max_seq - 1); +} + static inline void lru_gen_update_size(struct lruvec *lruvec, int type, struct folio *folio, int old_gen, int new_gen) { diff --git a/mm/vmscan.c b/mm/vmscan.c index 5892be56b617..8514182b92b2 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -1011,7 +1011,7 @@ void folio_inc_lru_refs(struct folio *folio, unsigned int flags) flags & (LRU_REF_EXEC | LRU_REF_FORCE)) gen = max_gen; /* First access only defers eviction from the oldest gen */ - else if (old_gen == min_gen) + else if (lru_gen_is_old(max_seq, gen)) gen = (old_gen + 1) % MAX_NR_GENS; refs = min(refs, LRU_REFS_PROTECTED); } else if (refs > LRU_REFS_MAX) { @@ -1132,7 +1132,7 @@ static int folio_inc_lru_refs_walk(struct folio *folio, struct lruvec *lruvec, if (refs > LRU_REFS_REFERENCED || is_exec_file_folio(folio, vma_flags)) *new_gen = max_gen; /* First access only defers eviction from the oldest gen */ - else if (old_gen == min_gen) + else if (lru_gen_is_old(max_seq, gen)) *new_gen = (old_gen + 1) % MAX_NR_GENS; else *new_gen = old_gen; ^ permalink raw reply related [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-10-08 12:24 UTC | newest] Thread overview: 5+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-10-03 5:42 [RFC PATCH] mm/mglru: Avoid reclaiming kept folios during retrying missed rotated folios Barry Song (Xiaomi) 2026-10-08 6:52 ` Baolin Wang 2026-10-08 11:43 ` Kairui Song 2026-10-08 11:52 ` Kairui Song 2026-10-08 12:23 ` Kairui Song
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox