From: Alexei Starovoitov <alexei.starovoitov@gmail.com>
To: bpf@vger.kernel.org, linux-mm@kvack.org
Cc: vbabka@suse.cz, harry.yoo@oracle.com, shakeel.butt@linux.dev,
mhocko@suse.com, bigeasy@linutronix.de, andrii@kernel.org,
memxor@gmail.com, akpm@linux-foundation.org,
peterz@infradead.org, rostedt@goodmis.org, hannes@cmpxchg.org
Subject: [PATCH v4 0/6] slab: Re-entrant kmalloc_nolock()
Date: Thu, 17 Jul 2025 19:16:40 -0700 [thread overview]
Message-ID: <20250718021646.73353-1-alexei.starovoitov@gmail.com> (raw)
From: Alexei Starovoitov <ast@kernel.org>
v3->v4:
- Converted local_lock_cpu_slab() to macro
- Reordered patches 5 and 6
- Emphasized that kfree_nolock() shouldn't be used on kmalloc()-ed objects
- Addressed other comments and improved commit logs
- Fixed build issues reported by bots
v3:
https://lore.kernel.org/bpf/20250716022950.69330-1-alexei.starovoitov@gmail.com/
v2->v3:
- Adopted Sebastian's local_lock_cpu_slab(), but dropped gfpflags
to avoid extra branch for performance reasons,
and added local_unlock_cpu_slab() for symmetry.
- Dropped local_lock_lockdep_start/end() pair and switched to
per kmem_cache lockdep class on PREEMPT_RT to silence false positive
when the same cpu/task acquires two local_lock-s.
- Refactorred defer_free per Sebastian's suggestion
- Fixed slab leak when it needs to be deactivated via irq_work and llist
as Vlastimil proposed. Including defer_free_barrier().
- Use kmem_cache->offset for llist_node pointer when linking objects
instead of zero offset, since whole object could be used for slabs
with ctors and other cases.
- Fixed "cnt = 1; goto redo;" issue.
- Fixed slab leak in alloc_single_from_new_slab().
- Retested with slab_debug, RT, !RT, lockdep, kasan, slab_tiny
- Added acks to patches 1-4 that should be good to go.
v2:
https://lore.kernel.org/bpf/20250709015303.8107-1-alexei.starovoitov@gmail.com/
v1->v2:
Added more comments for this non-trivial logic and addressed earlier comments.
In particular:
- Introduce alloc_frozen_pages_nolock() to avoid refcnt race
- alloc_pages_nolock() defaults to GFP_COMP
- Support SLUB_TINY
- Added more variants to stress tester to discover that kfree_nolock() can
OOM, because deferred per-slab llist won't be serviced if kfree_nolock()
gets unlucky long enough. Scraped previous approach and switched to
global per-cpu llist with immediate irq_work_queue() to process all
object sizes.
- Reentrant kmalloc cannot deactivate_slab(). In v1 the node hint was
downgraded to NUMA_NO_NODE before calling slab_alloc(). Realized it's not
good enough. There are odd cases that can trigger deactivate. Rewrote
this part.
- Struggled with SLAB_NO_CMPXCHG. Thankfully Harry had a great suggestion:
https://lore.kernel.org/bpf/aFvfr1KiNrLofavW@hyeyoo/
which was adopted. So slab_debug works now.
- In v1 I had to s/local_lock_irqsave/local_lock_irqsave_check/ in a bunch
of places in mm/slub.c to avoid lockdep false positives.
Came up with much cleaner approach to silence invalid lockdep reports
without sacrificing lockdep coverage. See local_lock_lockdep_start/end().
v1:
https://lore.kernel.org/bpf/20250501032718.65476-1-alexei.starovoitov@gmail.com/
Alexei Starovoitov (6):
locking/local_lock: Expose dep_map in local_trylock_t.
locking/local_lock: Introduce local_lock_is_locked().
mm: Allow GFP_ACCOUNT to be used in alloc_pages_nolock().
mm: Introduce alloc_frozen_pages_nolock()
slab: Make slub local_(try)lock more precise for LOCKDEP
slab: Introduce kmalloc_nolock() and kfree_nolock().
include/linux/gfp.h | 2 +-
include/linux/kasan.h | 13 +-
include/linux/local_lock.h | 2 +
include/linux/local_lock_internal.h | 16 +-
include/linux/rtmutex.h | 10 +
include/linux/slab.h | 4 +
kernel/bpf/syscall.c | 2 +-
kernel/locking/rtmutex_common.h | 9 -
mm/Kconfig | 1 +
mm/internal.h | 4 +
mm/kasan/common.c | 5 +-
mm/page_alloc.c | 54 ++--
mm/slab.h | 7 +
mm/slab_common.c | 3 +
mm/slub.c | 486 +++++++++++++++++++++++++---
15 files changed, 528 insertions(+), 90 deletions(-)
--
2.47.1
next reply other threads:[~2025-07-18 2:16 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-07-18 2:16 Alexei Starovoitov [this message]
2025-07-18 2:16 ` [PATCH v4 1/6] locking/local_lock: Expose dep_map in local_trylock_t Alexei Starovoitov
2025-07-18 2:16 ` [PATCH v4 2/6] locking/local_lock: Introduce local_lock_is_locked() Alexei Starovoitov
2025-07-18 2:16 ` [PATCH v4 3/6] mm: Allow GFP_ACCOUNT to be used in alloc_pages_nolock() Alexei Starovoitov
2025-07-18 2:16 ` [PATCH v4 4/6] mm: Introduce alloc_frozen_pages_nolock() Alexei Starovoitov
2025-07-18 2:16 ` [PATCH v4 5/6] slab: Make slub local_(try)lock more precise for LOCKDEP Alexei Starovoitov
2025-07-18 2:16 ` [PATCH v4 6/6] slab: Introduce kmalloc_nolock() and kfree_nolock() Alexei Starovoitov
2025-07-22 15:52 ` Harry Yoo
2025-08-06 2:40 ` Alexei Starovoitov
2025-08-12 15:11 ` Harry Yoo
2025-08-12 17:08 ` Harry Yoo
2025-09-09 0:08 ` Alexei Starovoitov
2025-09-09 2:05 ` Harry Yoo
2025-09-09 2:32 ` Alexei Starovoitov
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20250718021646.73353-1-alexei.starovoitov@gmail.com \
--to=alexei.starovoitov@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=andrii@kernel.org \
--cc=bigeasy@linutronix.de \
--cc=bpf@vger.kernel.org \
--cc=hannes@cmpxchg.org \
--cc=harry.yoo@oracle.com \
--cc=linux-mm@kvack.org \
--cc=memxor@gmail.com \
--cc=mhocko@suse.com \
--cc=peterz@infradead.org \
--cc=rostedt@goodmis.org \
--cc=shakeel.butt@linux.dev \
--cc=vbabka@suse.cz \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.