BPF List
 help / color / mirror / Atom feed
* [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace
@ 2026-09-30  9:59 Alexei Starovoitov
  2026-09-30  9:59 ` [PATCH bpf 1/3] bpf: Factor out __do_call_rcu_ttrace() Alexei Starovoitov
                   ` (3 more replies)
  0 siblings, 4 replies; 5+ messages in thread
From: Alexei Starovoitov @ 2026-09-30  9:59 UTC (permalink / raw)
  To: bpf; +Cc: daniel, andrii, eddyz87, memxor

From: Alexei Starovoitov <ast@kernel.org>

The objects that bpf_mem_alloc frees while RCU tasks trace GP is in flight
stay in free_by_rcu_ttrace list until the same bpf_mem_cache frees or
allocates in bulk again.

Patch 1 - refactoring. No functional change.
Patch 2 - the fix.
Patch 3 - selftest.

Signed-off-by: Alexei Starovoitov <ast@kernel.org>

Alexei Starovoitov (3):
  bpf: Factor out __do_call_rcu_ttrace()
  bpf: Fix objects stuck in free_by_rcu_ttrace
  selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace

 kernel/bpf/memalloc.c                         | 57 ++++++++++++++----
 .../selftests/bpf/prog_tests/bpf_ma_ttrace.c  | 60 +++++++++++++++++++
 .../selftests/bpf/progs/bpf_ma_ttrace.c       | 50 ++++++++++++++++
 3 files changed, 157 insertions(+), 10 deletions(-)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/bpf_ma_ttrace.c
 create mode 100644 tools/testing/selftests/bpf/progs/bpf_ma_ttrace.c

-- 
2.55.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH bpf 1/3] bpf: Factor out __do_call_rcu_ttrace()
  2026-09-30  9:59 [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
@ 2026-09-30  9:59 ` Alexei Starovoitov
  2026-09-30  9:59 ` [PATCH bpf 2/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
                   ` (2 subsequent siblings)
  3 siblings, 0 replies; 5+ messages in thread
From: Alexei Starovoitov @ 2026-09-30  9:59 UTC (permalink / raw)
  To: bpf; +Cc: daniel, andrii, eddyz87, memxor

From: Alexei Starovoitov <ast@kernel.org>

Move the part of do_call_rcu_ttrace() that runs after
call_rcu_ttrace_in_progress is set into __do_call_rcu_ttrace().
The next patch will call it from __free_rcu().
No functional change.

Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
 kernel/bpf/memalloc.c | 27 +++++++++++++++++----------
 1 file changed, 17 insertions(+), 10 deletions(-)

diff --git a/kernel/bpf/memalloc.c b/kernel/bpf/memalloc.c
index 8a8f088e83e6..08e4dde66cd5 100644
--- a/kernel/bpf/memalloc.c
+++ b/kernel/bpf/memalloc.c
@@ -298,19 +298,10 @@ static void enque_to_free(struct bpf_mem_cache *c, void *obj)
 	llist_add(llnode, &c->free_by_rcu_ttrace);
 }
 
-static void do_call_rcu_ttrace(struct bpf_mem_cache *c)
+static void __do_call_rcu_ttrace(struct bpf_mem_cache *c)
 {
 	struct llist_node *llnode, *t;
 
-	if (atomic_xchg(&c->call_rcu_ttrace_in_progress, 1)) {
-		if (unlikely(READ_ONCE(c->draining))) {
-			scoped_guard(raw_spinlock_irqsave, &c->lock)
-				llnode = llist_del_all(&c->free_by_rcu_ttrace);
-			free_all(c, llnode, !!c->percpu_size);
-		}
-		return;
-	}
-
 	WARN_ON_ONCE(!llist_empty(&c->waiting_for_gp_ttrace));
 	llist_for_each_safe(llnode, t, llist_del_all(&c->free_by_rcu_ttrace))
 		llist_add(llnode, &c->waiting_for_gp_ttrace);
@@ -328,6 +319,22 @@ static void do_call_rcu_ttrace(struct bpf_mem_cache *c)
 	call_rcu_tasks_trace(&c->rcu_ttrace, __free_rcu);
 }
 
+static void do_call_rcu_ttrace(struct bpf_mem_cache *c)
+{
+	struct llist_node *llnode;
+
+	if (atomic_xchg(&c->call_rcu_ttrace_in_progress, 1)) {
+		if (unlikely(READ_ONCE(c->draining))) {
+			scoped_guard(raw_spinlock_irqsave, &c->lock)
+				llnode = llist_del_all(&c->free_by_rcu_ttrace);
+			free_all(c, llnode, !!c->percpu_size);
+		}
+		return;
+	}
+
+	__do_call_rcu_ttrace(c);
+}
+
 static void free_bulk(struct bpf_mem_cache *c)
 {
 	struct bpf_mem_cache *tgt = c->tgt;
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH bpf 2/3] bpf: Fix objects stuck in free_by_rcu_ttrace
  2026-09-30  9:59 [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
  2026-09-30  9:59 ` [PATCH bpf 1/3] bpf: Factor out __do_call_rcu_ttrace() Alexei Starovoitov
@ 2026-09-30  9:59 ` Alexei Starovoitov
  2026-09-30  9:59 ` [PATCH bpf 3/3] selftests/bpf: Add a test for " Alexei Starovoitov
  2026-10-01 16:50 ` [PATCH bpf 0/3] bpf: Fix " patchwork-bot+netdevbpf
  3 siblings, 0 replies; 5+ messages in thread
From: Alexei Starovoitov @ 2026-09-30  9:59 UTC (permalink / raw)
  To: bpf; +Cc: daniel, andrii, eddyz87, memxor

From: Alexei Starovoitov <ast@kernel.org>

do_call_rcu_ttrace() returns early when call_rcu_ttrace_in_progress is set
and leaves the objects in free_by_rcu_ttrace. __free_rcu() frees
waiting_for_gp_ttrace only and clears the flag. Hence the objects that
free_bulk() or __free_by_rcu() added while RCU tasks trace GP was in flight
stay in free_by_rcu_ttrace until free_bulk() or alloc_bulk() is called for
the same bpf_mem_cache again, which may never happen. The number of such
objects is not bounded.

Turn call_rcu_ttrace_in_progress into three states:
0 - idle
1 - __free_rcu() is queued
2 - __free_rcu() is queued and free_by_rcu_ttrace got more objects since

do_call_rcu_ttrace() sets 2. __free_rcu() does cmpxchg(1 -> 0) and starts
the next GP when it fails. It cannot clear the flag first and check
free_by_rcu_ttrace later, since bpf_mem_alloc_destroy() frees bpf_mem_cache
without waiting for RCU callbacks when the flag is zero.

Now __free_rcu() queues itself, so the one that didn't see 'draining' may
do call_rcu_tasks_trace() after rcu_barrier_tasks_trace() in
free_mem_alloc(). Queue it under rcu_read_lock() and do synchronize_rcu()
before the barriers. Calling rcu_barrier_tasks_trace() twice works too, but
creating and destroying hash maps in a loop on many cpus slows down to one
free_mem_alloc() per GP and kworkers pile up.

Fixes: 8d5a8011b35d ("bpf: Batch call_rcu callbacks instead of SLAB_TYPESAFE_BY_RCU.")
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
 kernel/bpf/memalloc.c | 34 ++++++++++++++++++++++++++++++++--
 1 file changed, 32 insertions(+), 2 deletions(-)

diff --git a/kernel/bpf/memalloc.c b/kernel/bpf/memalloc.c
index 08e4dde66cd5..15684d0fc883 100644
--- a/kernel/bpf/memalloc.c
+++ b/kernel/bpf/memalloc.c
@@ -118,6 +118,11 @@ struct bpf_mem_cache {
 	struct llist_head free_by_rcu_ttrace;
 	struct llist_head waiting_for_gp_ttrace;
 	struct rcu_head rcu_ttrace;
+	/*
+	 * 0 - idle
+	 * 1 - __free_rcu() is queued
+	 * 2 - __free_rcu() is queued and free_by_rcu_ttrace got more objects since
+	 */
 	atomic_t call_rcu_ttrace_in_progress;
 	raw_spinlock_t lock;
 };
@@ -276,6 +281,8 @@ static int free_all(struct bpf_mem_cache *c, struct llist_node *llnode, bool per
 	return cnt;
 }
 
+static void __do_call_rcu_ttrace(struct bpf_mem_cache *c);
+
 static void __free_rcu(struct rcu_head *head)
 {
 	struct bpf_mem_cache *c = container_of(head, struct bpf_mem_cache, rcu_ttrace);
@@ -285,7 +292,19 @@ static void __free_rcu(struct rcu_head *head)
 		llnode = llist_del_all(&c->waiting_for_gp_ttrace);
 
 	free_all(c, llnode, !!c->percpu_size);
-	atomic_set(&c->call_rcu_ttrace_in_progress, 0);
+
+	/*
+	 * do_call_rcu_ttrace() that ran while GP was in flight left its objects
+	 * in free_by_rcu_ttrace. This cache may never free or alloc in bulk
+	 * again, so start the next GP from here.
+	 * 'c' can be freed as soon as call_rcu_ttrace_in_progress is zero.
+	 */
+	if (atomic_cmpxchg(&c->call_rcu_ttrace_in_progress, 1, 0) == 1)
+		return;
+
+	/* Pairs with synchronize_rcu() in free_mem_alloc() */
+	guard(rcu)();
+	__do_call_rcu_ttrace(c);
 }
 
 static void enque_to_free(struct bpf_mem_cache *c, void *obj)
@@ -302,6 +321,12 @@ static void __do_call_rcu_ttrace(struct bpf_mem_cache *c)
 {
 	struct llist_node *llnode, *t;
 
+	/*
+	 * Must be done before llist_del_all(). Objects that it misses were
+	 * added by do_call_rcu_ttrace() that will set 2 after this store.
+	 */
+	atomic_set(&c->call_rcu_ttrace_in_progress, 1);
+
 	WARN_ON_ONCE(!llist_empty(&c->waiting_for_gp_ttrace));
 	llist_for_each_safe(llnode, t, llist_del_all(&c->free_by_rcu_ttrace))
 		llist_add(llnode, &c->waiting_for_gp_ttrace);
@@ -323,7 +348,7 @@ static void do_call_rcu_ttrace(struct bpf_mem_cache *c)
 {
 	struct llist_node *llnode;
 
-	if (atomic_xchg(&c->call_rcu_ttrace_in_progress, 1)) {
+	if (atomic_xchg(&c->call_rcu_ttrace_in_progress, 2)) {
 		if (unlikely(READ_ONCE(c->draining))) {
 			scoped_guard(raw_spinlock_irqsave, &c->lock)
 				llnode = llist_del_all(&c->free_by_rcu_ttrace);
@@ -707,7 +732,12 @@ static void free_mem_alloc(struct bpf_mem_alloc *ma)
 	 * to wait for the pending __free_by_rcu(), and __free_rcu(). RCU Tasks
 	 * Trace grace period implies RCU grace period, so all __free_rcu don't
 	 * need extra call_rcu() (and thus extra rcu_barrier() here).
+	 *
+	 * __free_rcu() queues itself again unless it sees 'draining'. After
+	 * synchronize_rcu() it either did that already or will not do it, so
+	 * rcu_barrier_tasks_trace() cannot miss it.
 	 */
+	synchronize_rcu();
 	rcu_barrier(); /* wait for __free_by_rcu */
 	rcu_barrier_tasks_trace(); /* wait for __free_rcu */
 	free_mem_alloc_no_barrier(ma);
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH bpf 3/3] selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace
  2026-09-30  9:59 [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
  2026-09-30  9:59 ` [PATCH bpf 1/3] bpf: Factor out __do_call_rcu_ttrace() Alexei Starovoitov
  2026-09-30  9:59 ` [PATCH bpf 2/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
@ 2026-09-30  9:59 ` Alexei Starovoitov
  2026-10-01 16:50 ` [PATCH bpf 0/3] bpf: Fix " patchwork-bot+netdevbpf
  3 siblings, 0 replies; 5+ messages in thread
From: Alexei Starovoitov @ 2026-09-30  9:59 UTC (permalink / raw)
  To: bpf; +Cc: daniel, andrii, eddyz87, memxor

From: Alexei Starovoitov <ast@kernel.org>

Delete all elements of BPF_F_NO_PREALLOC hash map in one batch. The first
free_bulk() starts RCU tasks trace GP and the rest of the elements are
freed while it's in flight. Wait for call_rcu_ttrace_in_progress to clear
in bpf_mem_cache of every cpu and check that free_by_rcu_ttrace and
waiting_for_gp_ttrace lists are empty.

Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
 .../selftests/bpf/prog_tests/bpf_ma_ttrace.c  | 60 +++++++++++++++++++
 .../selftests/bpf/progs/bpf_ma_ttrace.c       | 50 ++++++++++++++++
 2 files changed, 110 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/bpf_ma_ttrace.c
 create mode 100644 tools/testing/selftests/bpf/progs/bpf_ma_ttrace.c

diff --git a/tools/testing/selftests/bpf/prog_tests/bpf_ma_ttrace.c b/tools/testing/selftests/bpf/prog_tests/bpf_ma_ttrace.c
new file mode 100644
index 000000000000..a1d41b109940
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/bpf_ma_ttrace.c
@@ -0,0 +1,60 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <test_progs.h>
+#include "bpf_ma_ttrace.skel.h"
+
+#define NR_ELEMS 4096
+
+/*
+ * The first free_bulk() starts RCU tasks trace GP. The rest of the elements are
+ * deleted while it's in flight. They should be freed without further alloc or
+ * free from this map.
+ */
+void test_bpf_ma_ttrace(void)
+{
+	LIBBPF_OPTS(bpf_test_run_opts, opts);
+	struct bpf_ma_ttrace *skel;
+	__u32 cnt = NR_ELEMS;
+	long *vals = NULL;
+	int *keys = NULL;
+	int i, err, fd, nr_cpus;
+
+	skel = bpf_ma_ttrace__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "open_and_load"))
+		return;
+	nr_cpus = libbpf_num_possible_cpus();
+	if (!ASSERT_GT(nr_cpus, 0, "nr_cpus"))
+		goto out;
+	skel->bss->nr_cpus = nr_cpus;
+
+	keys = calloc(NR_ELEMS, sizeof(*keys));
+	vals = calloc(NR_ELEMS, sizeof(*vals));
+	if (!ASSERT_OK_PTR(keys, "keys") || !ASSERT_OK_PTR(vals, "vals"))
+		goto out;
+	for (i = 0; i < NR_ELEMS; i++)
+		keys[i] = i;
+
+	fd = bpf_map__fd(skel->maps.htab);
+	err = bpf_map_update_batch(fd, keys, vals, &cnt, NULL);
+	if (!ASSERT_OK(err, "update_batch") || !ASSERT_EQ(cnt, NR_ELEMS, "update_cnt"))
+		goto out;
+	err = bpf_map_delete_batch(fd, keys, &cnt, NULL);
+	if (!ASSERT_OK(err, "delete_batch") || !ASSERT_EQ(cnt, NR_ELEMS, "delete_cnt"))
+		goto out;
+
+	/* Wait for all __free_rcu() callbacks to finish */
+	for (i = 0; i < 300; i++) {
+		err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.check_ttrace), &opts);
+		if (!ASSERT_OK(err, "test_run") || !ASSERT_OK(opts.retval, "retval"))
+			goto out;
+		if (!skel->bss->in_progress)
+			break;
+		usleep(100000);
+	}
+	ASSERT_EQ(skel->bss->nr_caches, nr_cpus, "nr_caches");
+	ASSERT_EQ(skel->bss->in_progress, 0, "in_progress");
+	ASSERT_EQ(skel->bss->not_freed, 0, "not_freed");
+out:
+	free(keys);
+	free(vals);
+	bpf_ma_ttrace__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/progs/bpf_ma_ttrace.c b/tools/testing/selftests/bpf/progs/bpf_ma_ttrace.c
new file mode 100644
index 000000000000..31b2a962e339
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/bpf_ma_ttrace.c
@@ -0,0 +1,50 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_core_read.h>
+#include "bpf_kfuncs.h"
+
+struct {
+	__uint(type, BPF_MAP_TYPE_HASH);
+	__uint(map_flags, BPF_F_NO_PREALLOC);
+	__uint(max_entries, 4096);
+	__type(key, int);
+	__type(value, long);
+} htab SEC(".maps");
+
+extern const void __per_cpu_offset __ksym;
+
+int nr_cpus;
+int nr_caches;
+int in_progress;
+int not_freed;
+
+/* Look at bpf_mem_cache of every cpu that htab allocates its elements from */
+SEC("syscall")
+int check_ttrace(void *ctx)
+{
+	struct bpf_htab *h = bpf_core_cast(&htab, struct bpf_htab);
+	unsigned long cache = (unsigned long)BPF_CORE_READ(h, ma.cache);
+	const unsigned long *offsets = &__per_cpu_offset;
+	struct bpf_mem_cache *c;
+	unsigned long off;
+	int cpu;
+
+	nr_caches = 0;
+	in_progress = 0;
+	not_freed = 0;
+	bpf_for(cpu, 0, nr_cpus) {
+		if (bpf_probe_read_kernel(&off, sizeof(off), offsets + cpu))
+			return -1;
+		c = bpf_core_cast((void *)(cache + off), struct bpf_mem_cache);
+		if (c->unit_size)
+			nr_caches++;
+		if (c->call_rcu_ttrace_in_progress.counter)
+			in_progress++;
+		if (c->free_by_rcu_ttrace.first || c->waiting_for_gp_ttrace.first)
+			not_freed++;
+	}
+	return 0;
+}
+
+char _license[] SEC("license") = "GPL";
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace
  2026-09-30  9:59 [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
                   ` (2 preceding siblings ...)
  2026-09-30  9:59 ` [PATCH bpf 3/3] selftests/bpf: Add a test for " Alexei Starovoitov
@ 2026-10-01 16:50 ` patchwork-bot+netdevbpf
  3 siblings, 0 replies; 5+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-10-01 16:50 UTC (permalink / raw)
  To: Alexei Starovoitov; +Cc: bpf, daniel, andrii, eddyz87, memxor

Hello:

This series was applied to bpf/bpf.git (master)
by Kumar Kartikeya Dwivedi <memxor@gmail.com>:

On Wed, 30 Sep 2026 09:59:17 +0000 you wrote:
> From: Alexei Starovoitov <ast@kernel.org>
> 
> The objects that bpf_mem_alloc frees while RCU tasks trace GP is in flight
> stay in free_by_rcu_ttrace list until the same bpf_mem_cache frees or
> allocates in bulk again.
> 
> Patch 1 - refactoring. No functional change.
> Patch 2 - the fix.
> Patch 3 - selftest.
> 
> [...]

Here is the summary with links:
  - [bpf,1/3] bpf: Factor out __do_call_rcu_ttrace()
    https://git.kernel.org/bpf/bpf/c/47f4f695cec1
  - [bpf,2/3] bpf: Fix objects stuck in free_by_rcu_ttrace
    https://git.kernel.org/bpf/bpf/c/fbd97dd1d3ce
  - [bpf,3/3] selftests/bpf: Add a test for objects stuck in free_by_rcu_ttrace
    https://git.kernel.org/bpf/bpf/c/aeb709c767ae

You are awesome, thank you!
-- 
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html



^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-10-01 16:50 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30  9:59 [PATCH bpf 0/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
2026-09-30  9:59 ` [PATCH bpf 1/3] bpf: Factor out __do_call_rcu_ttrace() Alexei Starovoitov
2026-09-30  9:59 ` [PATCH bpf 2/3] bpf: Fix objects stuck in free_by_rcu_ttrace Alexei Starovoitov
2026-09-30  9:59 ` [PATCH bpf 3/3] selftests/bpf: Add a test for " Alexei Starovoitov
2026-10-01 16:50 ` [PATCH bpf 0/3] bpf: Fix " patchwork-bot+netdevbpf

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox