All of lore.kernel.org
 help / color / mirror / Atom feed
From: Qinyun Tan <qinyuntan@linux.alibaba.com>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: Dave Chinner <david@fromorbit.com>, Qi Zheng <qi.zheng@linux.dev>,
	Roman Gushchin <roman.gushchin@linux.dev>,
	Muchun Song <muchun.song@linux.dev>,
	Baolin Wang <baolin.wang@linux.alibaba.com>,
	Xunlei Pang <xlpang@linux.alibaba.com>,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	Qinyun Tan <qinyuntan@linux.alibaba.com>
Subject: [PATCH] mm/list_lru: disable memcg awareness under cgroup_disable=memory
Date: Wed,  2 Sep 2026 11:28:40 +0800	[thread overview]
Message-ID: <20260902032840.30250-1-qinyuntan@linux.alibaba.com> (raw)

__list_lru_init() only collapses a memcg-aware list_lru into plain
per-node lists when kmem accounting is disabled
(cgroup.memory=nokmem).  When the memory controller is disabled
entirely (cgroup_disable=memory), mem_cgroup_kmem_disabled() is
false, so the lru stays memcg aware even though no object will ever
be charged to a memcg.

This is more than a semantic inconsistency.
folio_memcg_list_lru_alloc() trusts list_lru_memcg_aware() and
dereferences the folio's memcg, which is always NULL with the
controller disabled.  The only mainline caller,
folio_memcg_alloc_deferred(), papers over this with an explicit
mem_cgroup_disabled() check.  The shmem unused-huge shrinker
conversion ("mm: shmem: make unused huge shrinker memcg aware") adds
a second caller without such a guard, so booting with
cgroup_disable=memory and writing to a huge=always tmpfs oopses:

  BUG: unable to handle page fault for address: 0000000000000488
  RIP: 0010:folio_memcg_list_lru_alloc+0x41/0xf0
  Call Trace:
   <TASK>
   shmem_get_folio_gfp+0x1cd/0x7c0
   shmem_write_begin+0x5d/0x100
   generic_perform_write+0x89/0x2a0
   shmem_file_write_iter+0x82/0x90
   vfs_write+0x256/0x410
   ksys_write+0x61/0xe0
   do_syscall_64+0x8d/0x460
   entry_SYSCALL_64_after_hwframe+0x76/0x7e

The faulting address is the offset of mem_cgroup->kmemcg_id,
dereferenced on a NULL memcg in memcg_list_lru_allocated():

  folio_memcg_list_lru_alloc()
    list_lru_memcg_aware()               <- true, only nokmem checked
    memcg = folio_memcg(folio)           <- NULL
    memcg_list_lru_allocated(memcg, lru)
      memcg->kmemcg_id                   <- NULL pointer dereference

Check mem_cgroup_disabled() in __list_lru_init() so that all
list_lrus fall back to plain per-node lists when the controller is
disabled, matching what the shrinker side already does
(shrinker_memcg_alloc() bails out on mem_cgroup_disabled()).  This
makes the mem_cgroup_disabled() check in callers unnecessary rather
than mandatory.

Signed-off-by: Qinyun Tan <qinyuntan@linux.alibaba.com>
---
Applies on top of mm-new.  No Fixes: tag since the commit that makes
the crash reachable ("mm: shmem: make unused huge shrinker memcg
aware") is only in mm-new; no stable backport is needed either.

Reproducer, on mm-new booted with cgroup_disable=memory
(CONFIG_MEMCG=y, CONFIG_TRANSPARENT_HUGEPAGE=y):

  # mount -t tmpfs -o huge=always tmpfs /mnt
  # echo x > /mnt/f

Without this patch the write oopses immediately as shown above: the
freshly allocated huge folio extends beyond i_size, so
shmem_get_folio_gfp() queues the inode via shmem_unused_huge_add()
-> folio_memcg_list_lru_alloc(), which dereferences the NULL
folio_memcg().

With this patch the same steps run cleanly: the lru falls back to
plain per-node lists and folio_memcg_list_lru_alloc() returns early.
From code inspection the rest of the shmem path handles the NULL
objcg fine (obj_cgroup_memcg() and obj_cgroup_put() are NULL-safe,
and list_lru_add() with a NULL memcg lands on the per-node list),
but I have not exercised the shrinker reclaim itself under
cgroup_disable=memory.

Discussion: https://lore.kernel.org/linux-mm/20260901115104.2944996-1-qinyuntan@linux.alibaba.com/

 mm/list_lru.c | 7 ++++++-
 1 file changed, 6 insertions(+), 1 deletion(-)

diff --git a/mm/list_lru.c b/mm/list_lru.c
index 36662d02ff963..f8be119351cca 100644
--- a/mm/list_lru.c
+++ b/mm/list_lru.c
@@ -671,7 +671,12 @@ int __list_lru_init(struct list_lru *lru, bool memcg_aware, struct shrinker *shr
 	else
 		lru->shrinker_id = -1;
 
-	if (mem_cgroup_kmem_disabled())
+	/*
+	 * With the memory controller disabled entirely, no object is ever
+	 * charged to a memcg, so collapse to plain per-node lists just
+	 * like under nokmem.
+	 */
+	if (mem_cgroup_disabled() || mem_cgroup_kmem_disabled())
 		memcg_aware = false;
 #endif
 
-- 
2.43.7



             reply	other threads:[~2026-09-02  3:28 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-02  3:28 Qinyun Tan [this message]
2026-09-02  9:16 ` [PATCH] mm/list_lru: disable memcg awareness under cgroup_disable=memory Baolin Wang
2026-09-02  9:30   ` Qinyun Tan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260902032840.30250-1-qinyuntan@linux.alibaba.com \
    --to=qinyuntan@linux.alibaba.com \
    --cc=akpm@linux-foundation.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=david@fromorbit.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=muchun.song@linux.dev \
    --cc=qi.zheng@linux.dev \
    --cc=roman.gushchin@linux.dev \
    --cc=xlpang@linux.alibaba.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.