From: Michal Hocko <mhocko@suse.com>
To: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Andrew Morton <akpm@linux-foundation.org>,
Johannes Weiner <hannes@cmpxchg.org>,
Muchun Song <muchun.song@linux.dev>,
Qi Zheng <qi.zheng@linux.dev>,
Roman Gushchin <roman.gushchin@linux.dev>,
Meta kernel team <kernel-team@meta.com>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org,
stable@vger.kernel.org
Subject: Re: [PATCH] memcg: move LRU size accounting on reparenting instead of copying it
Date: Mon, 24 Aug 2026 13:47:54 +0200 [thread overview]
Message-ID: <aowvaq5RRDAtIKIY@tiehlicka> (raw)
In-Reply-To: <20260822024707.77192-1-shakeel.butt@linux.dev>
On Fri 21-08-26 19:47:07, Shakeel Butt wrote:
> When a memory cgroup is offlined its LRU folios are reparented to the
> parent. lruvec_reparent_lru() splices the child's lists into the
> parent's and credits the parent with the child's per-zone
> lru_zone_size[], but never clears the child's copy, so the size is
> copied rather than moved. lru_gen_reparent_memcg() does the same for
> MGLRU.
>
> The parent is left correct, credited with exactly the folios it took
> over. The stale value sits on the child and nothing will correct it:
> folio->memcg_data now resolves to the parent, so every later
> update_lru_size() for those folios goes there.
>
> Dying cgroups are not freed immediately and mem_cgroup_iter() still
> walks them, so shrink_lruvec() keeps being called on them.
> get_scan_count() reads the phantom counter through lruvec_lru_size() and
> the scan loop then grinds through nr[] in SWAP_CLUSTER_MAX steps against
> an empty list, for as long as the dead cgroup lives. Under MGLRU the
> MGLRU scanner runs instead, but count_shadow_nodes() sums all of
> NR_LRU_LISTS through lruvec_lru_size() and over-budgets the shadow node
> limit just the same.
>
> On one 251 GiB host a sweep of every mz->lru_zone_size[] found 380
> counters describing folios on no list at all: 124777314 pages, 476 GiB,
> 1.89x the machine's RAM, across 57 cgroups. All were on memcgs with
> CSS_DYING set and CSS_ONLINE clear, and parent/child pairs reported
> byte-identical sizes.
>
> LRU_UNEVICTABLE needs its size moved too. Its list is deliberately not
> spliced because lruvec_init() poisons the head - the unevictable LRU is
> imaginary and folios are never threaded on it - but the size is kept by
> lruvec_add_folio()/lruvec_del_folio() and those folios account to the
> parent from here on.
>
> This depends on commit bf4ade7dbd76 ("memcg: keep folio's objcg same as
> its node") and must not be backported ahead of it. Without that
> invariant a folio's objcg can belong to another node, so a folio already
> spliced onto the parent's list can still resolve to the child's lruvec
> until the objcg's node is reparented in a later iteration of
> memcg_reparent_objcgs(); clearing the child's counter early then lets
> lruvec_del_folio() underflow it and trip the WARN_ONCE()/VM_BUG_ON() in
> mem_cgroup_update_lru_size().
>
> Fixes: 07a6e9a2c199 ("mm: vmscan: prepare for reparenting traditional LRU folios")
> Fixes: f304652609ea ("mm: vmscan: prepare for reparenting MGLRU folios")
> Cc: <stable@vger.kernel.org> # After: bf4ade7dbd76: memcg: keep folio's objcg same as its node
> Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
Acked-by: Michal Hocko <mhocko@suse.com>
Thanks!
> ---
> mm/folio.c | 9 +++++++++
> mm/vmscan.c | 5 +++++
> 2 files changed, 14 insertions(+)
>
> diff --git a/mm/folio.c b/mm/folio.c
> index 59c477120b9a..c02dcea9c03c 100644
> --- a/mm/folio.c
> +++ b/mm/folio.c
> @@ -1130,7 +1130,16 @@ static void lruvec_reparent_lru(struct lruvec *child_lruvec,
> for_each_managed_zone_pgdat(zone, NODE_DATA(nid), zid, MAX_NR_ZONES - 1) {
> unsigned long size = mem_cgroup_get_zone_lru_size(child_lruvec, lru, zid);
>
> + if (!size)
> + continue;
> +
> + /*
> + * The folios are accounted to the parent from now on, so the
> + * size has to be moved, not just copied. Leaving it behind
> + * makes the dying child describe folios it no longer owns.
> + */
> mem_cgroup_update_lru_size(parent_lruvec, lru, zid, size);
> + mem_cgroup_update_lru_size(child_lruvec, lru, zid, -(long)size);
> }
> }
>
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index fe7f0c52a18c..561eeec5628c 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -4635,7 +4635,12 @@ void lru_gen_reparent_memcg(struct mem_cgroup *memcg, struct mem_cgroup *parent,
> for_each_managed_zone_pgdat(zone, NODE_DATA(nid), zid, MAX_NR_ZONES - 1) {
> unsigned long size = mem_cgroup_get_zone_lru_size(child_lruvec, lru, zid);
>
> + if (!size)
> + continue;
> +
> + /* Move the accounting, do not duplicate it. */
> mem_cgroup_update_lru_size(parent_lruvec, lru, zid, size);
> + mem_cgroup_update_lru_size(child_lruvec, lru, zid, -(long)size);
> }
> }
> }
> --
> 2.53.0-Meta
--
Michal Hocko
SUSE Labs
prev parent reply other threads:[~2026-08-24 11:47 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-22 2:47 [PATCH] memcg: move LRU size accounting on reparenting instead of copying it Shakeel Butt
2026-08-24 11:47 ` Michal Hocko [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aowvaq5RRDAtIKIY@tiehlicka \
--to=mhocko@suse.com \
--cc=akpm@linux-foundation.org \
--cc=hannes@cmpxchg.org \
--cc=kernel-team@meta.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=muchun.song@linux.dev \
--cc=qi.zheng@linux.dev \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox