All of lore.kernel.org
 help / color / mirror / Atom feed
From: Johannes Weiner <hannes@cmpxchg.org>
To: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Andrew Morton <akpm@linux-foundation.org>,
	Michal Hocko <mhocko@suse.com>,
	Roman Gushchin <roman.gushchin@linux.dev>,
	Muchun Song <muchun.song@linux.dev>,
	Qi Zheng <qi.zheng@linux.dev>,
	Meta kernel team <kernel-team@meta.com>,
	linux-mm@kvack.org, cgroups@vger.kernel.org,
	linux-kernel@vger.kernel.org,
	Karl Erik Hofseth <karl.e.hofseth@opoint.com>,
	stable@vger.kernel.org
Subject: Re: [PATCH v2] memcg: keep folio's objcg same as its node
Date: Thu, 6 Aug 2026 14:15:26 -0400	[thread overview]
Message-ID: <anTPPr5hVV4aQLRD@cmpxchg.org> (raw)
In-Reply-To: <20260806165813.2526415-1-shakeel.butt@linux.dev>

On Thu, Aug 06, 2026 at 09:58:13AM -0700, Shakeel Butt wrote:
> memcg_reparent_objcgs() has an inherent assumption that a folio's objcg
> is the objcg of the folio's node.  Folio migration across nodes breaks
> that assumption: the new folio simply inherits the old folio's objcg
> while living on a different node.
> 
> Once the assumption is broken, the reparenting of the folio's objcg and
> the reparenting of the folio's LRU list are no longer atomic.
> memcg_reparent_objcgs() handles one node per iteration and drops all the
> locks in between, so the objcg gets reparented in the iteration for the
> objcg's node while the LRU list gets spliced in the iteration for the
> folio's node.  Any LRU operation on that folio in between resolves its
> lruvec through the objcg, and thus takes the lru_lock of the wrong
> memcg, not the lru_lock of the list the folio is actually on.
> 
> Fix this by selecting the objcg by folio_nid() at charge time, and by
> re-deriving it for the destination node in mem_cgroup_migrate() and
> mem_cgroup_replace_folio().
> 
> Reported-by: Karl Erik Hofseth <karl.e.hofseth@opoint.com>
> Closes: https://lore.kernel.org/all/anMmd1ADrDVwMO6v@work/
> Fixes: f1cf8d2f36dc ("mm: memcontrol: eliminate the problem of dying memory cgroup for LRU folios")
> Cc: stable@vger.kernel.org
> Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
> ---
> 
> Changes since v1: http://lore.kernel.org/20260806061830.3294679-1-shakeel.butt@linux.dev
> 
> - In mem_cgroup_migrate, do obj_cgroup_put at the end (Sashiko)
> - Handle scenario where destination node has been reparented to the root but the
>   source node's objcg has not yet (Sashiko)
> - Add comment explaining the race between migration and reparenting (Johannes)
> 
>  mm/memcontrol.c | 66 ++++++++++++++++++++++++++++++++++++++++---------
>  1 file changed, 54 insertions(+), 12 deletions(-)
> 
> diff --git a/mm/memcontrol.c b/mm/memcontrol.c
> index 3057396dda53..02c108cdd0f5 100644
> --- a/mm/memcontrol.c
> +++ b/mm/memcontrol.c
> @@ -2966,10 +2966,9 @@ struct mem_cgroup *mem_cgroup_from_virt(void *p)
>  	return folio_memcg_check(virt_to_folio(p));
>  }
>  
> -static struct obj_cgroup *__get_obj_cgroup_from_memcg(struct mem_cgroup *memcg)
> +static struct obj_cgroup *__get_obj_cgroup_from_memcg(struct mem_cgroup *memcg,
> +						      int nid)
>  {
> -	int nid = numa_node_id();
> -
>  	for (; memcg; memcg = parent_mem_cgroup(memcg)) {
>  		struct obj_cgroup *objcg = rcu_dereference(memcg->nodeinfo[nid]->objcg);
>  
> @@ -2980,12 +2979,13 @@ static struct obj_cgroup *__get_obj_cgroup_from_memcg(struct mem_cgroup *memcg)
>  	return NULL;
>  }
>  
> -static inline struct obj_cgroup *get_obj_cgroup_from_memcg(struct mem_cgroup *memcg)
> +static inline struct obj_cgroup *get_obj_cgroup_from_memcg(struct mem_cgroup *memcg,
> +							   int nid)
>  {
>  	struct obj_cgroup *objcg;
>  
>  	rcu_read_lock();
> -	objcg = __get_obj_cgroup_from_memcg(memcg);
> +	objcg = __get_obj_cgroup_from_memcg(memcg, nid);
>  	rcu_read_unlock();
>  
>  	return objcg;
> @@ -3029,7 +3029,7 @@ static struct obj_cgroup *current_objcg_update(void)
>  
>  		rcu_read_lock();
>  		memcg = mem_cgroup_from_task(current);
> -		objcg = __get_obj_cgroup_from_memcg(memcg);
> +		objcg = __get_obj_cgroup_from_memcg(memcg, numa_node_id());
>  		rcu_read_unlock();
>  
>  		/*
> @@ -5197,7 +5197,7 @@ static int charge_memcg(struct folio *folio, struct mem_cgroup *memcg,
>  	int ret = 0;
>  	struct obj_cgroup *objcg;
>  
> -	objcg = get_obj_cgroup_from_memcg(memcg);
> +	objcg = get_obj_cgroup_from_memcg(memcg, folio_nid(folio));
>  	/* Do not account at the root objcg level. */
>  	if (!obj_cgroup_is_root(objcg))
>  		ret = try_charge_memcg(memcg, gfp, folio_nr_pages(folio));
> @@ -5408,8 +5408,8 @@ void __mem_cgroup_uncharge_folios(struct folio_batch *folios)
>   */
>  void mem_cgroup_replace_folio(struct folio *old, struct folio *new)
>  {
> +	struct obj_cgroup *objcg, *new_objcg = NULL;
>  	struct mem_cgroup *memcg;
> -	struct obj_cgroup *objcg;
>  	long nr_pages = folio_nr_pages(new);
>  
>  	VM_BUG_ON_FOLIO(!folio_test_locked(old), old);
> @@ -5431,6 +5431,7 @@ void mem_cgroup_replace_folio(struct folio *old, struct folio *new)
>  
>  	rcu_read_lock();
>  	memcg = obj_cgroup_memcg(objcg);
> +
>  	/* Force-charge the new page. The old one will be freed soon */
>  	if (!obj_cgroup_is_root(objcg)) {
>  		page_counter_charge(&memcg->memory, nr_pages);
> @@ -5438,7 +5439,31 @@ void mem_cgroup_replace_folio(struct folio *old, struct folio *new)
>  			page_counter_charge(&memcg->memsw, nr_pages);
>  	}
>  
> -	obj_cgroup_get(objcg);
> +	/*
> +	 * memcg_reparent_objcgs() reparents a node's objcgs and its LRU lists
> +	 * together, under that node's lru_locks. If a folio's objcg is on a
> +	 * different node, the two happen in separate iterations with the locks
> +	 * dropped in between, and an LRU operation in that window takes the
> +	 * lru_lock of the wrong memcg. Keep the objcg on the folio's node.
> +	 */
> +	if (folio_nid(old) != folio_nid(new))
> +		new_objcg = __get_obj_cgroup_from_memcg(memcg, folio_nid(new));
> +
> +	/*
> +	 * LRU folios are not accounted at the root level: swapping a non-root
> +	 * objcg for the root one would skip the uncharge and leak the charge.
> +	 * Keep the old one, the root memcg never reparents. No put needed,
> +	 * tryget takes no reference on the root objcg.
> +	 */
> +	if (new_objcg && obj_cgroup_is_root(new_objcg) &&
> +	    !obj_cgroup_is_root(objcg))
> +		new_objcg = NULL;

The two versions have quite some overlap. The refcounting is
different, but IMO that part is also hard to square.

How about a helper that always returns a referenced objcg?

static struct obj_cgroup *get_migration_objcg(struct folio *old, struct folio *new)
{
	struct obj_cgroup *old_objcg, *new_objcg;
	int new_nid = folio_nid(new);
	struct mem_cgroup *memcg;

	old_objcg = get_obj_cgroup_from_folio(old);

	if (folio_nid(old) == new_nid)
		return old_objcg;

	memcg = obj_cgroup_memcg(old_objcg);
	new_objcg = __get_obj_cgroup_from_memcg(memcg, new_nid);

	if (new_objcg && obj_cgroup_is_root(new_objcg) &&
	    !obj_cgroup_is_root(old_objcg))
		return old_objcg;

	obj_cgroup_put(old_objcg);

	return new_objcg;
}

In mem_cgroup_replace_folio(), this becomes:

	rcu_read_lock();

	... page counter update ...

	objcg = get_migration_objcg(old, new);
	commit_charge(new, objcg);
	memcg1_commit_charge(new, memcg);

	rcu_read_unlock();

> +
> +	if (new_objcg)
> +		objcg = new_objcg;
> +	else
> +		obj_cgroup_get(objcg);
> +
>  	commit_charge(new, objcg);
>  	memcg1_commit_charge(new, memcg);
>  	rcu_read_unlock();
> @@ -5457,7 +5482,7 @@ void mem_cgroup_replace_folio(struct folio *old, struct folio *new)
>   */
>  void mem_cgroup_migrate(struct folio *old, struct folio *new)
>  {
> -	struct obj_cgroup *objcg;
> +	struct obj_cgroup *objcg, *new_objcg = NULL;
>  
>  	VM_BUG_ON_FOLIO(!folio_test_locked(old), old);
>  	VM_BUG_ON_FOLIO(!folio_test_locked(new), new);
> @@ -5478,12 +5503,29 @@ void mem_cgroup_migrate(struct folio *old, struct folio *new)
>  	if (!objcg)
>  		return;
>  
> -	/* Transfer the charge and the objcg ref */
> -	commit_charge(new, objcg);
> +	/* Keep the objcg on the folio's node, see mem_cgroup_replace_folio() */
> +	if (folio_nid(old) != folio_nid(new)) {
> +		rcu_read_lock();
> +		new_objcg = __get_obj_cgroup_from_memcg(obj_cgroup_memcg(objcg),
> +							folio_nid(new));
> +		rcu_read_unlock();
> +
> +		/* No root swap, see mem_cgroup_replace_folio(). */
> +		if (new_objcg && obj_cgroup_is_root(new_objcg) &&
> +		    !obj_cgroup_is_root(objcg))
> +			new_objcg = NULL;
> +	}
> +
> +	/* Transfer the charge and, unless it was swapped, the objcg ref */
> +	commit_charge(new, new_objcg ? : objcg);
>  
>  	/* Warning should never happen, so don't worry about refcount non-0 */
>  	WARN_ON_ONCE(folio_unqueue_deferred_split(old));
>  	old->memcg_data = 0;
> +
> +	/* @new took its own reference, drop @old's. */
> +	if (new_objcg)
> +		obj_cgroup_put(objcg);

And here it becomes:

	rcu_read_lock();
	objcg = get_migration_objcg(old, new);
	rcu_read_unlock();
	commit_charge(new, objcg);

	objcg = folio_objcg(old);
	old->memcg_data = 0;
	obj_cgroup_put(objcg);

  reply	other threads:[~2026-08-06 18:15 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-06 16:58 [PATCH v2] memcg: keep folio's objcg same as its node Shakeel Butt
2026-08-06 18:15 ` Johannes Weiner [this message]
2026-08-06 19:22   ` Johannes Weiner
2026-08-06 20:03     ` Shakeel Butt

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anTPPr5hVV4aQLRD@cmpxchg.org \
    --to=hannes@cmpxchg.org \
    --cc=akpm@linux-foundation.org \
    --cc=cgroups@vger.kernel.org \
    --cc=karl.e.hofseth@opoint.com \
    --cc=kernel-team@meta.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=mhocko@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=qi.zheng@linux.dev \
    --cc=roman.gushchin@linux.dev \
    --cc=shakeel.butt@linux.dev \
    --cc=stable@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.