From: Qinyun Tan <qinyuntan@linux.alibaba.com>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: Dave Chinner <david@fromorbit.com>, Qi Zheng <qi.zheng@linux.dev>,
Roman Gushchin <roman.gushchin@linux.dev>,
Muchun Song <muchun.song@linux.dev>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
Xunlei Pang <xlpang@linux.alibaba.com>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
Qinyun Tan <qinyuntan@linux.alibaba.com>
Subject: [PATCH] mm/list_lru: disable memcg awareness under cgroup_disable=memory
Date: Wed, 2 Sep 2026 11:28:40 +0800 [thread overview]
Message-ID: <20260902032840.30250-1-qinyuntan@linux.alibaba.com> (raw)
__list_lru_init() only collapses a memcg-aware list_lru into plain
per-node lists when kmem accounting is disabled
(cgroup.memory=nokmem). When the memory controller is disabled
entirely (cgroup_disable=memory), mem_cgroup_kmem_disabled() is
false, so the lru stays memcg aware even though no object will ever
be charged to a memcg.
This is more than a semantic inconsistency.
folio_memcg_list_lru_alloc() trusts list_lru_memcg_aware() and
dereferences the folio's memcg, which is always NULL with the
controller disabled. The only mainline caller,
folio_memcg_alloc_deferred(), papers over this with an explicit
mem_cgroup_disabled() check. The shmem unused-huge shrinker
conversion ("mm: shmem: make unused huge shrinker memcg aware") adds
a second caller without such a guard, so booting with
cgroup_disable=memory and writing to a huge=always tmpfs oopses:
BUG: unable to handle page fault for address: 0000000000000488
RIP: 0010:folio_memcg_list_lru_alloc+0x41/0xf0
Call Trace:
<TASK>
shmem_get_folio_gfp+0x1cd/0x7c0
shmem_write_begin+0x5d/0x100
generic_perform_write+0x89/0x2a0
shmem_file_write_iter+0x82/0x90
vfs_write+0x256/0x410
ksys_write+0x61/0xe0
do_syscall_64+0x8d/0x460
entry_SYSCALL_64_after_hwframe+0x76/0x7e
The faulting address is the offset of mem_cgroup->kmemcg_id,
dereferenced on a NULL memcg in memcg_list_lru_allocated():
folio_memcg_list_lru_alloc()
list_lru_memcg_aware() <- true, only nokmem checked
memcg = folio_memcg(folio) <- NULL
memcg_list_lru_allocated(memcg, lru)
memcg->kmemcg_id <- NULL pointer dereference
Check mem_cgroup_disabled() in __list_lru_init() so that all
list_lrus fall back to plain per-node lists when the controller is
disabled, matching what the shrinker side already does
(shrinker_memcg_alloc() bails out on mem_cgroup_disabled()). This
makes the mem_cgroup_disabled() check in callers unnecessary rather
than mandatory.
Signed-off-by: Qinyun Tan <qinyuntan@linux.alibaba.com>
---
Applies on top of mm-new. No Fixes: tag since the commit that makes
the crash reachable ("mm: shmem: make unused huge shrinker memcg
aware") is only in mm-new; no stable backport is needed either.
Reproducer, on mm-new booted with cgroup_disable=memory
(CONFIG_MEMCG=y, CONFIG_TRANSPARENT_HUGEPAGE=y):
# mount -t tmpfs -o huge=always tmpfs /mnt
# echo x > /mnt/f
Without this patch the write oopses immediately as shown above: the
freshly allocated huge folio extends beyond i_size, so
shmem_get_folio_gfp() queues the inode via shmem_unused_huge_add()
-> folio_memcg_list_lru_alloc(), which dereferences the NULL
folio_memcg().
With this patch the same steps run cleanly: the lru falls back to
plain per-node lists and folio_memcg_list_lru_alloc() returns early.
From code inspection the rest of the shmem path handles the NULL
objcg fine (obj_cgroup_memcg() and obj_cgroup_put() are NULL-safe,
and list_lru_add() with a NULL memcg lands on the per-node list),
but I have not exercised the shrinker reclaim itself under
cgroup_disable=memory.
Discussion: https://lore.kernel.org/linux-mm/20260901115104.2944996-1-qinyuntan@linux.alibaba.com/
mm/list_lru.c | 7 ++++++-
1 file changed, 6 insertions(+), 1 deletion(-)
diff --git a/mm/list_lru.c b/mm/list_lru.c
index 36662d02ff963..f8be119351cca 100644
--- a/mm/list_lru.c
+++ b/mm/list_lru.c
@@ -671,7 +671,12 @@ int __list_lru_init(struct list_lru *lru, bool memcg_aware, struct shrinker *shr
else
lru->shrinker_id = -1;
- if (mem_cgroup_kmem_disabled())
+ /*
+ * With the memory controller disabled entirely, no object is ever
+ * charged to a memcg, so collapse to plain per-node lists just
+ * like under nokmem.
+ */
+ if (mem_cgroup_disabled() || mem_cgroup_kmem_disabled())
memcg_aware = false;
#endif
--
2.43.7
next reply other threads:[~2026-09-02 3:28 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 3:28 Qinyun Tan [this message]
2026-09-02 9:16 ` [PATCH] mm/list_lru: disable memcg awareness under cgroup_disable=memory Baolin Wang
2026-09-02 9:30 ` Qinyun Tan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260902032840.30250-1-qinyuntan@linux.alibaba.com \
--to=qinyuntan@linux.alibaba.com \
--cc=akpm@linux-foundation.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@fromorbit.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=muchun.song@linux.dev \
--cc=qi.zheng@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=xlpang@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.