* [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace
@ 2026-09-30 9:59 Alexei Starovoitov
2026-09-30 9:59 ` [PATCH bpf 1/3] bpf: Factor out __do_call_rcu_ttrace() Alexei Starovoitov
` (3 more replies)
0 siblings, 4 replies; 5+ messages in thread
From: Alexei Starovoitov @ 2026-09-30 9:59 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor
From: Alexei Starovoitov <ast@kernel.org>
The objects that bpf_mem_alloc frees while RCU tasks trace GP is in flight
stay in free_by_rcu_ttrace list until the same bpf_mem_cache frees or
allocates in bulk again.
Patch 1 - refactoring. No functional change.
Patch 2 - the fix.
Patch 3 - selftest.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Alexei Starovoitov (3):
bpf: Factor out __do_call_rcu_ttrace()
bpf: Fix objects stuck in free_by_rcu_ttrace
selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace
kernel/bpf/memalloc.c | 57 ++++++++++++++----
.../selftests/bpf/prog_tests/bpf_ma_ttrace.c | 60 +++++++++++++++++++
.../selftests/bpf/progs/bpf_ma_ttrace.c | 50 ++++++++++++++++
3 files changed, 157 insertions(+), 10 deletions(-)
create mode 100644 tools/testing/selftests/bpf/prog_tests/bpf_ma_ttrace.c
create mode 100644 tools/testing/selftests/bpf/progs/bpf_ma_ttrace.c
--
2.55.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH bpf 1/3] bpf: Factor out __do_call_rcu_ttrace()
2026-09-30 9:59 [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
@ 2026-09-30 9:59 ` Alexei Starovoitov
2026-09-30 9:59 ` [PATCH bpf 2/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
` (2 subsequent siblings)
3 siblings, 0 replies; 5+ messages in thread
From: Alexei Starovoitov @ 2026-09-30 9:59 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor
From: Alexei Starovoitov <ast@kernel.org>
Move the part of do_call_rcu_ttrace() that runs after
call_rcu_ttrace_in_progress is set into __do_call_rcu_ttrace().
The next patch will call it from __free_rcu().
No functional change.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
kernel/bpf/memalloc.c | 27 +++++++++++++++++----------
1 file changed, 17 insertions(+), 10 deletions(-)
diff --git a/kernel/bpf/memalloc.c b/kernel/bpf/memalloc.c
index 8a8f088e83e6..08e4dde66cd5 100644
--- a/kernel/bpf/memalloc.c
+++ b/kernel/bpf/memalloc.c
@@ -298,19 +298,10 @@ static void enque_to_free(struct bpf_mem_cache *c, void *obj)
llist_add(llnode, &c->free_by_rcu_ttrace);
}
-static void do_call_rcu_ttrace(struct bpf_mem_cache *c)
+static void __do_call_rcu_ttrace(struct bpf_mem_cache *c)
{
struct llist_node *llnode, *t;
- if (atomic_xchg(&c->call_rcu_ttrace_in_progress, 1)) {
- if (unlikely(READ_ONCE(c->draining))) {
- scoped_guard(raw_spinlock_irqsave, &c->lock)
- llnode = llist_del_all(&c->free_by_rcu_ttrace);
- free_all(c, llnode, !!c->percpu_size);
- }
- return;
- }
-
WARN_ON_ONCE(!llist_empty(&c->waiting_for_gp_ttrace));
llist_for_each_safe(llnode, t, llist_del_all(&c->free_by_rcu_ttrace))
llist_add(llnode, &c->waiting_for_gp_ttrace);
@@ -328,6 +319,22 @@ static void do_call_rcu_ttrace(struct bpf_mem_cache *c)
call_rcu_tasks_trace(&c->rcu_ttrace, __free_rcu);
}
+static void do_call_rcu_ttrace(struct bpf_mem_cache *c)
+{
+ struct llist_node *llnode;
+
+ if (atomic_xchg(&c->call_rcu_ttrace_in_progress, 1)) {
+ if (unlikely(READ_ONCE(c->draining))) {
+ scoped_guard(raw_spinlock_irqsave, &c->lock)
+ llnode = llist_del_all(&c->free_by_rcu_ttrace);
+ free_all(c, llnode, !!c->percpu_size);
+ }
+ return;
+ }
+
+ __do_call_rcu_ttrace(c);
+}
+
static void free_bulk(struct bpf_mem_cache *c)
{
struct bpf_mem_cache *tgt = c->tgt;
--
2.55.0
^ permalink raw reply related [flat|nested] 5+ messages in thread
* [PATCH bpf 2/3] bpf: Fix objects stuck in free_by_rcu_ttrace
2026-09-30 9:59 [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
2026-09-30 9:59 ` [PATCH bpf 1/3] bpf: Factor out __do_call_rcu_ttrace() Alexei Starovoitov
@ 2026-09-30 9:59 ` Alexei Starovoitov
2026-09-30 9:59 ` [PATCH bpf 3/3] selftests/bpf: Add a test for " Alexei Starovoitov
2026-10-01 16:50 ` [PATCH bpf 0/3] bpf: Fix " patchwork-bot+netdevbpf
3 siblings, 0 replies; 5+ messages in thread
From: Alexei Starovoitov @ 2026-09-30 9:59 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor
From: Alexei Starovoitov <ast@kernel.org>
do_call_rcu_ttrace() returns early when call_rcu_ttrace_in_progress is set
and leaves the objects in free_by_rcu_ttrace. __free_rcu() frees
waiting_for_gp_ttrace only and clears the flag. Hence the objects that
free_bulk() or __free_by_rcu() added while RCU tasks trace GP was in flight
stay in free_by_rcu_ttrace until free_bulk() or alloc_bulk() is called for
the same bpf_mem_cache again, which may never happen. The number of such
objects is not bounded.
Turn call_rcu_ttrace_in_progress into three states:
0 - idle
1 - __free_rcu() is queued
2 - __free_rcu() is queued and free_by_rcu_ttrace got more objects since
do_call_rcu_ttrace() sets 2. __free_rcu() does cmpxchg(1 -> 0) and starts
the next GP when it fails. It cannot clear the flag first and check
free_by_rcu_ttrace later, since bpf_mem_alloc_destroy() frees bpf_mem_cache
without waiting for RCU callbacks when the flag is zero.
Now __free_rcu() queues itself, so the one that didn't see 'draining' may
do call_rcu_tasks_trace() after rcu_barrier_tasks_trace() in
free_mem_alloc(). Queue it under rcu_read_lock() and do synchronize_rcu()
before the barriers. Calling rcu_barrier_tasks_trace() twice works too, but
creating and destroying hash maps in a loop on many cpus slows down to one
free_mem_alloc() per GP and kworkers pile up.
Fixes: 8d5a8011b35d ("bpf: Batch call_rcu callbacks instead of SLAB_TYPESAFE_BY_RCU.")
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
kernel/bpf/memalloc.c | 34 ++++++++++++++++++++++++++++++++--
1 file changed, 32 insertions(+), 2 deletions(-)
diff --git a/kernel/bpf/memalloc.c b/kernel/bpf/memalloc.c
index 08e4dde66cd5..15684d0fc883 100644
--- a/kernel/bpf/memalloc.c
+++ b/kernel/bpf/memalloc.c
@@ -118,6 +118,11 @@ struct bpf_mem_cache {
struct llist_head free_by_rcu_ttrace;
struct llist_head waiting_for_gp_ttrace;
struct rcu_head rcu_ttrace;
+ /*
+ * 0 - idle
+ * 1 - __free_rcu() is queued
+ * 2 - __free_rcu() is queued and free_by_rcu_ttrace got more objects since
+ */
atomic_t call_rcu_ttrace_in_progress;
raw_spinlock_t lock;
};
@@ -276,6 +281,8 @@ static int free_all(struct bpf_mem_cache *c, struct llist_node *llnode, bool per
return cnt;
}
+static void __do_call_rcu_ttrace(struct bpf_mem_cache *c);
+
static void __free_rcu(struct rcu_head *head)
{
struct bpf_mem_cache *c = container_of(head, struct bpf_mem_cache, rcu_ttrace);
@@ -285,7 +292,19 @@ static void __free_rcu(struct rcu_head *head)
llnode = llist_del_all(&c->waiting_for_gp_ttrace);
free_all(c, llnode, !!c->percpu_size);
- atomic_set(&c->call_rcu_ttrace_in_progress, 0);
+
+ /*
+ * do_call_rcu_ttrace() that ran while GP was in flight left its objects
+ * in free_by_rcu_ttrace. This cache may never free or alloc in bulk
+ * again, so start the next GP from here.
+ * 'c' can be freed as soon as call_rcu_ttrace_in_progress is zero.
+ */
+ if (atomic_cmpxchg(&c->call_rcu_ttrace_in_progress, 1, 0) == 1)
+ return;
+
+ /* Pairs with synchronize_rcu() in free_mem_alloc() */
+ guard(rcu)();
+ __do_call_rcu_ttrace(c);
}
static void enque_to_free(struct bpf_mem_cache *c, void *obj)
@@ -302,6 +321,12 @@ static void __do_call_rcu_ttrace(struct bpf_mem_cache *c)
{
struct llist_node *llnode, *t;
+ /*
+ * Must be done before llist_del_all(). Objects that it misses were
+ * added by do_call_rcu_ttrace() that will set 2 after this store.
+ */
+ atomic_set(&c->call_rcu_ttrace_in_progress, 1);
+
WARN_ON_ONCE(!llist_empty(&c->waiting_for_gp_ttrace));
llist_for_each_safe(llnode, t, llist_del_all(&c->free_by_rcu_ttrace))
llist_add(llnode, &c->waiting_for_gp_ttrace);
@@ -323,7 +348,7 @@ static void do_call_rcu_ttrace(struct bpf_mem_cache *c)
{
struct llist_node *llnode;
- if (atomic_xchg(&c->call_rcu_ttrace_in_progress, 1)) {
+ if (atomic_xchg(&c->call_rcu_ttrace_in_progress, 2)) {
if (unlikely(READ_ONCE(c->draining))) {
scoped_guard(raw_spinlock_irqsave, &c->lock)
llnode = llist_del_all(&c->free_by_rcu_ttrace);
@@ -707,7 +732,12 @@ static void free_mem_alloc(struct bpf_mem_alloc *ma)
* to wait for the pending __free_by_rcu(), and __free_rcu(). RCU Tasks
* Trace grace period implies RCU grace period, so all __free_rcu don't
* need extra call_rcu() (and thus extra rcu_barrier() here).
+ *
+ * __free_rcu() queues itself again unless it sees 'draining'. After
+ * synchronize_rcu() it either did that already or will not do it, so
+ * rcu_barrier_tasks_trace() cannot miss it.
*/
+ synchronize_rcu();
rcu_barrier(); /* wait for __free_by_rcu */
rcu_barrier_tasks_trace(); /* wait for __free_rcu */
free_mem_alloc_no_barrier(ma);
--
2.55.0
^ permalink raw reply related [flat|nested] 5+ messages in thread
* [PATCH bpf 3/3] selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace
2026-09-30 9:59 [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
2026-09-30 9:59 ` [PATCH bpf 1/3] bpf: Factor out __do_call_rcu_ttrace() Alexei Starovoitov
2026-09-30 9:59 ` [PATCH bpf 2/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
@ 2026-09-30 9:59 ` Alexei Starovoitov
2026-10-01 16:50 ` [PATCH bpf 0/3] bpf: Fix " patchwork-bot+netdevbpf
3 siblings, 0 replies; 5+ messages in thread
From: Alexei Starovoitov @ 2026-09-30 9:59 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor
From: Alexei Starovoitov <ast@kernel.org>
Delete all elements of BPF_F_NO_PREALLOC hash map in one batch. The first
free_bulk() starts RCU tasks trace GP and the rest of the elements are
freed while it's in flight. Wait for call_rcu_ttrace_in_progress to clear
in bpf_mem_cache of every cpu and check that free_by_rcu_ttrace and
waiting_for_gp_ttrace lists are empty.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
.../selftests/bpf/prog_tests/bpf_ma_ttrace.c | 60 +++++++++++++++++++
.../selftests/bpf/progs/bpf_ma_ttrace.c | 50 ++++++++++++++++
2 files changed, 110 insertions(+)
create mode 100644 tools/testing/selftests/bpf/prog_tests/bpf_ma_ttrace.c
create mode 100644 tools/testing/selftests/bpf/progs/bpf_ma_ttrace.c
diff --git a/tools/testing/selftests/bpf/prog_tests/bpf_ma_ttrace.c b/tools/testing/selftests/bpf/prog_tests/bpf_ma_ttrace.c
new file mode 100644
index 000000000000..a1d41b109940
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/bpf_ma_ttrace.c
@@ -0,0 +1,60 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <test_progs.h>
+#include "bpf_ma_ttrace.skel.h"
+
+#define NR_ELEMS 4096
+
+/*
+ * The first free_bulk() starts RCU tasks trace GP. The rest of the elements are
+ * deleted while it's in flight. They should be freed without further alloc or
+ * free from this map.
+ */
+void test_bpf_ma_ttrace(void)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ struct bpf_ma_ttrace *skel;
+ __u32 cnt = NR_ELEMS;
+ long *vals = NULL;
+ int *keys = NULL;
+ int i, err, fd, nr_cpus;
+
+ skel = bpf_ma_ttrace__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "open_and_load"))
+ return;
+ nr_cpus = libbpf_num_possible_cpus();
+ if (!ASSERT_GT(nr_cpus, 0, "nr_cpus"))
+ goto out;
+ skel->bss->nr_cpus = nr_cpus;
+
+ keys = calloc(NR_ELEMS, sizeof(*keys));
+ vals = calloc(NR_ELEMS, sizeof(*vals));
+ if (!ASSERT_OK_PTR(keys, "keys") || !ASSERT_OK_PTR(vals, "vals"))
+ goto out;
+ for (i = 0; i < NR_ELEMS; i++)
+ keys[i] = i;
+
+ fd = bpf_map__fd(skel->maps.htab);
+ err = bpf_map_update_batch(fd, keys, vals, &cnt, NULL);
+ if (!ASSERT_OK(err, "update_batch") || !ASSERT_EQ(cnt, NR_ELEMS, "update_cnt"))
+ goto out;
+ err = bpf_map_delete_batch(fd, keys, &cnt, NULL);
+ if (!ASSERT_OK(err, "delete_batch") || !ASSERT_EQ(cnt, NR_ELEMS, "delete_cnt"))
+ goto out;
+
+ /* Wait for all __free_rcu() callbacks to finish */
+ for (i = 0; i < 300; i++) {
+ err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.check_ttrace), &opts);
+ if (!ASSERT_OK(err, "test_run") || !ASSERT_OK(opts.retval, "retval"))
+ goto out;
+ if (!skel->bss->in_progress)
+ break;
+ usleep(100000);
+ }
+ ASSERT_EQ(skel->bss->nr_caches, nr_cpus, "nr_caches");
+ ASSERT_EQ(skel->bss->in_progress, 0, "in_progress");
+ ASSERT_EQ(skel->bss->not_freed, 0, "not_freed");
+out:
+ free(keys);
+ free(vals);
+ bpf_ma_ttrace__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/progs/bpf_ma_ttrace.c b/tools/testing/selftests/bpf/progs/bpf_ma_ttrace.c
new file mode 100644
index 000000000000..31b2a962e339
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/bpf_ma_ttrace.c
@@ -0,0 +1,50 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_core_read.h>
+#include "bpf_kfuncs.h"
+
+struct {
+ __uint(type, BPF_MAP_TYPE_HASH);
+ __uint(map_flags, BPF_F_NO_PREALLOC);
+ __uint(max_entries, 4096);
+ __type(key, int);
+ __type(value, long);
+} htab SEC(".maps");
+
+extern const void __per_cpu_offset __ksym;
+
+int nr_cpus;
+int nr_caches;
+int in_progress;
+int not_freed;
+
+/* Look at bpf_mem_cache of every cpu that htab allocates its elements from */
+SEC("syscall")
+int check_ttrace(void *ctx)
+{
+ struct bpf_htab *h = bpf_core_cast(&htab, struct bpf_htab);
+ unsigned long cache = (unsigned long)BPF_CORE_READ(h, ma.cache);
+ const unsigned long *offsets = &__per_cpu_offset;
+ struct bpf_mem_cache *c;
+ unsigned long off;
+ int cpu;
+
+ nr_caches = 0;
+ in_progress = 0;
+ not_freed = 0;
+ bpf_for(cpu, 0, nr_cpus) {
+ if (bpf_probe_read_kernel(&off, sizeof(off), offsets + cpu))
+ return -1;
+ c = bpf_core_cast((void *)(cache + off), struct bpf_mem_cache);
+ if (c->unit_size)
+ nr_caches++;
+ if (c->call_rcu_ttrace_in_progress.counter)
+ in_progress++;
+ if (c->free_by_rcu_ttrace.first || c->waiting_for_gp_ttrace.first)
+ not_freed++;
+ }
+ return 0;
+}
+
+char _license[] SEC("license") = "GPL";
--
2.55.0
^ permalink raw reply related [flat|nested] 5+ messages in thread
* Re: [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace
2026-09-30 9:59 [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
` (2 preceding siblings ...)
2026-09-30 9:59 ` [PATCH bpf 3/3] selftests/bpf: Add a test for " Alexei Starovoitov
@ 2026-10-01 16:50 ` patchwork-bot+netdevbpf
3 siblings, 0 replies; 5+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-10-01 16:50 UTC (permalink / raw)
To: Alexei Starovoitov; +Cc: bpf, daniel, andrii, eddyz87, memxor
Hello:
This series was applied to bpf/bpf.git (master)
by Kumar Kartikeya Dwivedi <memxor@gmail.com>:
On Wed, 30 Sep 2026 09:59:17 +0000 you wrote:
> From: Alexei Starovoitov <ast@kernel.org>
>
> The objects that bpf_mem_alloc frees while RCU tasks trace GP is in flight
> stay in free_by_rcu_ttrace list until the same bpf_mem_cache frees or
> allocates in bulk again.
>
> Patch 1 - refactoring. No functional change.
> Patch 2 - the fix.
> Patch 3 - selftest.
>
> [...]
Here is the summary with links:
- [bpf,1/3] bpf: Factor out __do_call_rcu_ttrace()
https://git.kernel.org/bpf/bpf/c/47f4f695cec1
- [bpf,2/3] bpf: Fix objects stuck in free_by_rcu_ttrace
https://git.kernel.org/bpf/bpf/c/fbd97dd1d3ce
- [bpf,3/3] selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace
https://git.kernel.org/bpf/bpf/c/aeb709c767ae
You are awesome, thank you!
--
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-10-01 16:50 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30 9:59 [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
2026-09-30 9:59 ` [PATCH bpf 1/3] bpf: Factor out __do_call_rcu_ttrace() Alexei Starovoitov
2026-09-30 9:59 ` [PATCH bpf 2/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
2026-09-30 9:59 ` [PATCH bpf 3/3] selftests/bpf: Add a test for " Alexei Starovoitov
2026-10-01 16:50 ` [PATCH bpf 0/3] bpf: Fix " patchwork-bot+netdevbpf
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox