From: "Vlastimil Babka (SUSE)" <vbabka@kernel.org>
To: "Harry Yoo (Meta)" <harry@kernel.org>,
Andrew Morton <akpm@linux-foundation.org>,
Hao Li <hao.li@linux.dev>, Christoph Lameter <cl@gentwo.org>,
David Rientjes <rientjes@google.com>,
Roman Gushchin <roman.gushchin@linux.dev>
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org,
Hyunwoo Kim <imv4bel@gmail.com>
Subject: Re: [PATCH v2] mm/slab: take n->list_lock in __slab_try_return_freelist() to avoid race
Date: Wed, 2 Sep 2026 18:57:53 +0200 [thread overview]
Message-ID: <b85f9954-4bbd-45f4-80dc-6c9d32cefcd5@kernel.org> (raw)
In-Reply-To: <20260902-slab-fix-aba-v2-1-d1ece15a8417@kernel.org>
On 9/2/26 18:41, Harry Yoo (Meta) wrote:
> Commit ba7425312607 ("mm, slab: add an optimistic
> __slab_try_return_freelist()") incorrectly assumed that nobody has freed
> an object to the slab as long as slab->freelist is NULL and cmpxchg
> succeeds.
>
> However, as reported by Hyunwoo Kim [1], other CPUs might have freed
> an object to the slab, insert the slab to the partial list, then
> allocated an object from the slab, and be in the middle of removing
> the slab from the list under n->list_lock.
>
> Since __refill_objects_node() puts the slab back on pc.slabs
> outside n->list_lock, it might insert the slab into that list while
> the slab is concurrently being removed from n->partial.
> This led to a list corruption, as reported by Hyunwoo Kim [1]:
"as reported by ..." was already said above, we could keep just the [1] or
drop completely?
> list_add corruption. next->prev should be prev
> (ffff888100000248), but was dead000000000122.
> (next=ffffea000416e410).
> kernel BUG at lib/list_debug.c:29!
> Oops: invalid opcode: 0000 [#1] SMP NOPTI
> CPU: 1 UID: 65534 PID: 144 Comm: poc Not tainted
> 7.2.0-16172-gcf72cbb39da8-dirty #1 PREEMPT(lazy)
> RIP: 0010:__list_add_valid_or_report+0x80/0xd0
> ...
> Call Trace:
> alloc_from_new_slab+0x183/0x300
> ___slab_alloc+0x31c/0x890
> __kmalloc_noprof+0x3d4/0x800
> lsm_blob_alloc+0x2d/0x50
> security_msg_msg_alloc+0x26/0x90
> load_msg+0x1aa/0x210
> do_msgsnd+0x91/0x800
> do_syscall_64+0x109/0x5d0
> entry_SYSCALL_64_after_hwframe+0x77/0x7f
> ...
> Kernel panic - not syncing: Fatal exception
>
> This is a classic ABA problem where cmpxchg succeeds but the state has
> changed since __refill_objects_node() took the freelist from the slab.
>
> As Vlastimil Babka mentioned [2], it should be rare to return more than
> one slab (due to the racy read of slab->counters in
> get_partial_node_bulk()). Therefore, instead of introducing additional
> complexity, acquire and release n->list_lock twice in the worst case.
>
> Return the slab directly to the partial list and hold n->list_lock
> across the cmpxchg and add_partial(). This is similar to the initial
> version of commit ba7425312607 [3]. This is enough to avoid the race as
> the list manipulation is serialized by n->list_lock. While at it,
> bring back unlikely() hint now that the condition is unlikely.
Ah nicely spotted.
> Reported-by: Hyunwoo Kim <imv4bel@gmail.com>
> Closes: https://lore.kernel.org/linux-mm/apPa-cGLcyt90l-E@v4bel [1]
> Link: https://lore.kernel.org/linux-mm/ae25c193-b95f-40c1-83b6-1c2546467e41@kernel.org [2]
> Link: https://lore.kernel.org/all/20260421-b4-refill-optimistic-return-v1-1-24f0bfc1acff@kernel.org [3]
> Fixes: ba7425312607 ("mm, slab: add an optimistic __slab_try_return_freelist()")
> Cc: stable@vger.kernel.org
> Signed-off-by: Harry Yoo (Meta) <harry@kernel.org>
> ---
> Changes in v2:
> - EDITME: describe what is new in this series revision.
> - EDITME: use bulletpoints and terse descriptions.
> - Link to v1: https://lore.kernel.org/r/apPa-cGLcyt90l-E@v4bel
> ---
> mm/slub.c | 17 ++++++++++++-----
> 1 file changed, 12 insertions(+), 5 deletions(-)
>
> diff --git a/mm/slub.c b/mm/slub.c
> index f9b56cb439e7..5cbbacb8ee32 100644
> --- a/mm/slub.c
> +++ b/mm/slub.c
> @@ -5684,6 +5684,8 @@ static bool __slab_try_return_freelist(struct kmem_cache *s, struct slab *slab,
> void *head, int cnt)
> {
> struct freelist_counters old, new;
> + struct kmem_cache_node *n;
> + unsigned long flags;
>
> old.freelist = slab->freelist;
> old.counters = slab->counters;
> @@ -5695,9 +5697,16 @@ static bool __slab_try_return_freelist(struct kmem_cache *s, struct slab *slab,
> new.counters = old.counters;
> new.inuse -= cnt;
>
> - if (!slab_update_freelist(s, slab, &old, &new, "__slab_try_return_freelist"))
> + n = get_node(s, slab_nid(slab));
Wonder if we could just pass the 'n' to __slab_try_return_freelist() that we
already have in the caller.
> + spin_lock_irqsave(&n->list_lock, flags);
> +
> + if (!slab_update_freelist(s, slab, &old, &new, "__slab_try_return_freelist")) {
> + spin_unlock_irqrestore(&n->list_lock, flags);
> return false;
> + }
>
> + add_partial(n, slab, ADD_TO_TAIL);
> + spin_unlock_irqrestore(&n->list_lock, flags);
> return true;
> }
>
> @@ -7296,10 +7305,8 @@ __refill_objects_node(struct kmem_cache *s, void **p, gfp_t gfp, unsigned int mi
> void *head = object;
> void *tail;
>
> - if (__slab_try_return_freelist(s, slab, head, count)) {
> - list_add(&slab->slab_list, &pc.slabs);
> + if (__slab_try_return_freelist(s, slab, head, count))
> break;
> - }
>
> do {
> tail = object;
> @@ -7312,7 +7319,7 @@ __refill_objects_node(struct kmem_cache *s, void **p, gfp_t gfp, unsigned int mi
> break;
> }
>
> - if (!list_empty(&pc.slabs)) {
> + if (unlikely(!list_empty(&pc.slabs))) {
> spin_lock_irqsave(&n->list_lock, flags);
>
> list_for_each_entry(slab, &pc.slabs, slab_list)
>
> ---
> base-commit: cee9395acd8043be0644b25c34bfa86623f2b935
> change-id: 20260902-slab-fix-aba-d60f39fc9dee
>
> Best regards,
next prev parent reply other threads:[~2026-09-02 16:58 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 16:41 [PATCH v2] mm/slab: take n->list_lock in __slab_try_return_freelist() to avoid race Harry Yoo (Meta)
2026-09-02 16:57 ` Vlastimil Babka (SUSE) [this message]
2026-09-03 10:22 ` Harry Yoo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=b85f9954-4bbd-45f4-80dc-6c9d32cefcd5@kernel.org \
--to=vbabka@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=cl@gentwo.org \
--cc=hao.li@linux.dev \
--cc=harry@kernel.org \
--cc=imv4bel@gmail.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=rientjes@google.com \
--cc=roman.gushchin@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.