From: Muchun Song <muchun.song@linux.dev>
To: Prakash Gupta <prakash.gupta@oss.qualcomm.com>
Cc: Johannes Weiner <hannes@cmpxchg.org>,
Michal Hocko <mhocko@kernel.org>,
Roman Gushchin <roman.gushchin@linux.dev>,
Shakeel Butt <shakeel.butt@linux.dev>,
Andrew Morton <akpm@linux-foundation.org>,
Kefeng Wang <wangkefeng.wang@huawei.com>,
"Matthew Wilcox (Oracle)" <willy@infradead.org>,
linux-arm-msm@vger.kernel.org, cgroups@vger.kernel.org,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
stable@vger.kernel.org
Subject: Re: [PATCH] mm/memcg: fix NULL nodeinfo[] dereference on late-onlined nodes
Date: Mon, 10 Aug 2026 15:58:09 +0800 [thread overview]
Message-ID: <5059A780-9C10-4171-A254-38944464F1AC@linux.dev> (raw)
In-Reply-To: <20260810-memcg-nodeinfo-null-guard-v1-1-77a73204d63a@oss.qualcomm.com>
> On Aug 10, 2026, at 13:46, Prakash Gupta <prakash.gupta@oss.qualcomm.com> wrote:
>
> memcg->nodeinfo[] entries are allocated only for nodes present at
> css_alloc time. When a node is onlined after a memcg is created its
> nodeinfo[] slot remains NULL. Two call sites dereference these slots
> unconditionally:
I don't think the premise of this patch is correct.
memcg->nodeinfo[] is not allocated only for nodes that are present or
online at css_alloc time. mem_cgroup_alloc() allocates per-node info with
for_each_node(), and for_each_node() iterates N_POSSIBLE nodes:
for_each_node(node)
alloc_mem_cgroup_per_node_info(memcg, node);
Memory hotplug also rejects memory being added to a node that is not in
node_possible_map. So a node that can be onlined later should already
have a nodeinfo[] entry for all existing memcgs.
If memcg->nodeinfo[nid] is NULL on your system, that looks like a
violation of this invariant, or possibly a downstream-specific change,
bad nid, allocation/lifetime issue, or memory corruption. I don't think
the generic explanation that "the node was onlined after the memcg was
created" is sufficient.
>
> lruvec_stat_mod_folio() calls mem_cgroup_lruvec() which reads
> memcg->nodeinfo[pgdat->node_id] without a NULL check. On a system
> where a node is onlined after the memcg is created, any folio stat
> update for that node crashes with a NULL pointer dereference:
>
> Unable to handle kernel paging request at virtual address ffffffbebf7e1908
> pc : lruvec_stat_mod_folio+0xf0/0x444 [6.18.21-android17-5]
> lr : lruvec_stat_mod_folio+0x88/0x444
> Call trace:
> lruvec_stat_mod_folio+0xf0/0x444
> folio_add_new_anon_rmap+0xac/0x2b8
> do_wp_page+0x768/0xc80
> handle_mm_fault+0x37c/0x8c4
> do_page_fault+0x140/0xa1c
> do_mem_abort+0x54/0x74
> el0_da+0x48/0x8c
> el0t_64_sync_handler+0x20/0x130
> el0t_64_sync+0x1c4/0x1c8
>
> __invalidate_reclaim_iterators() iterates for_each_node() and reads
> from->nodeinfo[nid]->iter without checking for NULL. for_each_node()
> visits all possible nodes, so this is reachable whenever a node is
> onlined after the memcg was created.
>
> Fix lruvec_stat_mod_folio() by checking nodeinfo[pgdat->node_id]
> directly and falling back to mod_node_page_state() when NULL, mirroring
> the existing !memcg early-return path. The fallback must not go through
> mod_lruvec_state() since that calls mod_memcg_lruvec_state() which uses
> container_of() to recover the mem_cgroup_per_node from the lruvec
> pointer; passing &pgdat->__lruvec there produces a garbage pointer.
> When nodeinfo[nid] is NULL the memcg has no per-node accounting
> structure for that node, so node-level accounting is correct.
>
> Fix __invalidate_reclaim_iterators() by skipping NULL nodeinfo[] slots.
>
> Also switch lruvec_stat_mod_folio() to folio_memcg_check() which uses
> READ_ONCE() to safely read folio->memcg_data in an unlocked context.
>
> Fixes: 6c77b607ee26 ("mm: kill lock|unlock_page_memcg()")
The Fixes tag also seems odd. 6c77b607ee26 ("mm: kill
lock|unlock_page_memcg()") only removed/renamed the lock_page_memcg()
wrappers and does not appear to change nodeinfo[] allocation or memory
hotplug handling. Could you explain how that commit introduced the NULL
nodeinfo condition?
> Cc: stable@vger.kernel.org
> Assisted-by: pi:claude-sonnet-4-5
Given the Assisted-by tag, I assume some of the analysis may have been
tool-assisted. That's fine, but the author still needs to validate the
reasoning against the actual code before submission.
Did you confirm that the relevant allocation and hotplug paths were
manually checked against the affected tree? In particular, I wonder
whether this behavior depends on downstream changes around
mem_cgroup_alloc(), for_each_node(), node_possible_map setup, or memory
hotplug nid validation.
Muchun,
Thanks.
prev parent reply other threads:[~2026-08-10 7:58 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-10 5:46 [PATCH] mm/memcg: fix NULL nodeinfo[] dereference on late-onlined nodes Prakash Gupta
2026-08-10 7:58 ` Muchun Song [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=5059A780-9C10-4171-A254-38944464F1AC@linux.dev \
--to=muchun.song@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=cgroups@vger.kernel.org \
--cc=hannes@cmpxchg.org \
--cc=linux-arm-msm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mhocko@kernel.org \
--cc=prakash.gupta@oss.qualcomm.com \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=stable@vger.kernel.org \
--cc=wangkefeng.wang@huawei.com \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox