Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH] mm/list_lru: don't copy stale shrinker id from non-memcg-aware shrinkers
@ 2026-09-01 11:51 Qinyun Tan
  2026-09-01 17:29 ` Andrew Morton
                   ` (3 more replies)
  0 siblings, 4 replies; 7+ messages in thread
From: Qinyun Tan @ 2026-09-01 11:51 UTC (permalink / raw)
  To: Andrew Morton
  Cc: Johannes Weiner, Michal Koutný, Lance Yang, Qi Zheng,
	Roman Gushchin, Muchun Song, Dave Chinner, Baolin Wang,
	David Hildenbrand, Xunlei Pang, linux-mm, linux-kernel,
	Qinyun Tan

With cgroup.memory=nokmem, shrinker_memcg_alloc() fails with -ENOSYS
for shrinkers without SHRINKER_NONSLAB, and shrinker_alloc() falls
back to a non-memcg-aware shrinker.  On this fallback path,
shrinker->id is never assigned and keeps 0 from kzalloc(), which is a
valid id belonging to whichever memcg-aware shrinker registers first.

__list_lru_init() copies shrinker->id unconditionally, so every
list_lru backed by such a fallback shrinker (thp-deferred_split,
zswap-shrinker, workingset shadow nodes, superblock lrus, ...) ends
up with lru->shrinker_id == 0 instead of -1.

Under nokmem the list_lru collapses to the shared per-node lists, but
__list_lru_add() still calls set_shrinker_bit() against the memcg of
the added object.  Most list_lru users are unaffected because their
objects resolve to a NULL memcg without kmem accounting, but the THP
deferred split queue holds user folios, which are charged regardless
of nokmem.  Since no memcg-aware shrinker can register under nokmem,
shrinker_nr_max stays 0 and every memcg's shrinker_info has
map_nr_max == 0, so the first folio added by khugepaged triggers on
every boot:

  WARNING: mm/shrinker.c:212 at set_shrinker_bit+0x99/0xa0

On systems where a SHRINKER_NONSLAB shrinker (btrfs, xfs) did register
and expand the maps, there is no warning; instead bit 0 is set
spuriously for an unrelated shrinker.

shrinker->id is only meaningful while SHRINKER_MEMCG_AWARE is set,
and all readers inside mm/shrinker.c already check the flag before
using the id.  Make __list_lru_init() do the same and fall back to -1,
so set_shrinker_bit() is never reached with a bogus id.  The stale
shrinker->id itself is left as is; cleaning that up is a separate
topic.

Fixes: 03375203e1da8 ("mm: do not allocate shrinker info with cgroup.memory=nokmem")
Signed-off-by: Qinyun Tan <qinyuntan@linux.alibaba.com>
---

Verified on a machine booting with cgroup.memory=nokmem and
CONFIG_TRANSPARENT_HUGEPAGE=y: the warning fires once per boot from
khugepaged, disappears when nokmem is removed from the command line,
and no longer triggers with this fix applied and nokmem set.

 mm/list_lru.c | 7 ++++++-
 1 file changed, 6 insertions(+), 1 deletion(-)

diff --git a/mm/list_lru.c b/mm/list_lru.c
index 36662d02ff963..96d02ea206a88 100644
--- a/mm/list_lru.c
+++ b/mm/list_lru.c
@@ -666,7 +666,12 @@ int __list_lru_init(struct list_lru *lru, bool memcg_aware, struct shrinker *shr
 	int i;
 
 #ifdef CONFIG_MEMCG
-	if (shrinker)
+	/*
+	 * If the shrinker fell back to being non-memcg-aware (e.g. with
+	 * cgroup.memory=nokmem), its id was never assigned and holds a
+	 * stale 0. Don't let set_shrinker_bit() act on it.
+	 */
+	if (shrinker && (shrinker->flags & SHRINKER_MEMCG_AWARE))
 		lru->shrinker_id = shrinker->id;
 	else
 		lru->shrinker_id = -1;
-- 
2.43.7



^ permalink raw reply related	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-09-03  4:07 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-01 11:51 [PATCH] mm/list_lru: don't copy stale shrinker id from non-memcg-aware shrinkers Qinyun Tan
2026-09-01 17:29 ` Andrew Morton
2026-09-02  3:20   ` Qinyun Tan
2026-09-02  2:25 ` Muchun Song
2026-09-02  5:30 ` Baolin Wang
2026-09-02 17:17 ` Michal Koutný
2026-09-03  4:06   ` Qinyun Tan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox