* [PATCH bpf-next v6 1/4] bpf: Add bpf_call_rcu() kfunc
2026-09-22 20:00 [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace() Puranjay Mohan
@ 2026-09-22 20:00 ` Puranjay Mohan
2026-09-22 20:00 ` [PATCH bpf-next v6 2/4] selftests/bpf: Add tests for bpf_call_rcu() Puranjay Mohan
` (4 subsequent siblings)
5 siblings, 0 replies; 17+ messages in thread
From: Puranjay Mohan @ 2026-09-22 20:00 UTC (permalink / raw)
To: bpf, rcu
Cc: Puranjay Mohan, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman,
Kumar Kartikeya Dwivedi, Song Liu, Yonghong Song,
Harry Yoo (Oracle), Paul E. McKenney
BPF programs that manage their own objects have no way to run their own
logic once an RCU grace period has elapsed. bpf_obj_drop() defers a
free, but returning an index to an allocator or unpinning a resource
once readers are done has no equivalent. sched_ext's BPF library works
around this today by pushing freed nodes onto a list and having a
userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
BPF program to reclaim them.
Add:
int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
int (*callback)(struct bpf_map *map, void *key,
void *value));
@rh is a struct bpf_rcu_head embedded in a value of @map, so the
callback runs as callback(map, key, value) for the element it lives in
and needs no cookie. A head can only be armed once, which bounds
outstanding work by the number of elements.
struct bpf_rcu_head holds the callback state inline rather than a
pointer to it, as bpf_timer, bpf_wq and bpf_task_work do, because there
is nothing to cancel and so nothing that has to outlive the map value.
That avoids an allocation and a state machine on the arming path at the
cost of 48 bytes per element.
An RCU callback cannot be cancelled, so everything it touches has to
stay alive until it runs:
- The callback is the program's text, so arming takes a program
reference as bpf_timer, bpf_wq and bpf_task_work do, dropped once
the callback returns. bpf_prog_inc_not_zero() also fails the arm
with -EBADF once the program is dying.
- The map is held by that reference through used_maps. An inner map
is not, so bpf_rcu_head is rejected in one.
- The field is only accepted in BPF_MAP_TYPE_ARRAY, whose elements
are never freed individually. A hash element can be deleted and
recycled while a callback is queued on it.
- The head is disarmed before the callback runs so it can be armed
again from there, which takes a new program reference before the
running callback drops its own. Arming therefore fails with -EPERM
once the map is held by neither a process nor bpffs.
bpf_iter hands a program a writable pointer to the live element, which
would let it overwrite a queued head, so bpf_iter_attach_map() rejects
maps carrying one.
The callback is verified non-sleepable even when the caller is
sleepable, and RCU invokes it with BH disabled.
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
---
include/linux/bpf.h | 10 +++++
include/uapi/linux/bpf.h | 4 ++
kernel/bpf/btf.c | 7 +++
kernel/bpf/helpers.c | 75 +++++++++++++++++++++++++++++++
kernel/bpf/map_in_map.c | 4 ++
kernel/bpf/map_iter.c | 6 +++
kernel/bpf/syscall.c | 11 ++++-
kernel/bpf/verifier.c | 81 +++++++++++++++++++++++++++++++++-
tools/include/uapi/linux/bpf.h | 4 ++
9 files changed, 199 insertions(+), 3 deletions(-)
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index fd22db8bc6c50..e7c5e203edddb 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -215,6 +215,7 @@ enum btf_field_type {
BPF_UPTR = (1 << 11),
BPF_RES_SPIN_LOCK = (1 << 12),
BPF_TASK_WORK = (1 << 13),
+ BPF_RCU_HEAD = (1 << 14),
};
enum bpf_cgroup_storage_type {
@@ -269,6 +270,7 @@ struct btf_record {
int wq_off;
int refcount_off;
int task_work_off;
+ int rcu_head_off;
struct btf_field fields[];
};
@@ -374,6 +376,8 @@ static inline const char *btf_field_type_name(enum btf_field_type type)
return "bpf_refcount";
case BPF_TASK_WORK:
return "bpf_task_work";
+ case BPF_RCU_HEAD:
+ return "bpf_rcu_head";
default:
WARN_ON_ONCE(1);
return "unknown";
@@ -414,6 +418,8 @@ static inline u32 btf_field_type_size(enum btf_field_type type)
return sizeof(struct bpf_refcount);
case BPF_TASK_WORK:
return sizeof(struct bpf_task_work);
+ case BPF_RCU_HEAD:
+ return sizeof(struct bpf_rcu_head);
default:
WARN_ON_ONCE(1);
return 0;
@@ -448,6 +454,8 @@ static inline u32 btf_field_type_align(enum btf_field_type type)
return __alignof__(struct bpf_refcount);
case BPF_TASK_WORK:
return __alignof__(struct bpf_task_work);
+ case BPF_RCU_HEAD:
+ return __alignof__(struct bpf_rcu_head);
default:
WARN_ON_ONCE(1);
return 0;
@@ -480,6 +488,7 @@ static inline void bpf_obj_init_field(const struct btf_field *field, void *addr)
case BPF_KPTR_PERCPU:
case BPF_UPTR:
case BPF_TASK_WORK:
+ case BPF_RCU_HEAD:
break;
default:
WARN_ON_ONCE(1);
@@ -925,6 +934,7 @@ enum bpf_arg_type {
ARG_PTR_TO_RB_NODE, /* pointer to bpf_rb_node */
ARG_PTR_TO_WORKQUEUE, /* pointer to bpf_wq */
ARG_PTR_TO_TASK_WORK, /* pointer to bpf_task_work */
+ ARG_PTR_TO_RCU_HEAD, /* pointer to bpf_rcu_head */
ARG_PTR_TO_IRQ_FLAG, /* pointer to saved IRQ flags on the stack */
ARG_PTR_TO_RES_SPIN_LOCK, /* pointer to bpf_res_spin_lock */
ARG_PTR_TO_CTX_OUT, /* hook output argument passed through from ctx */
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 6330b7d745c57..0aaa54359aebc 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -7611,6 +7611,10 @@ struct bpf_task_work {
__u64 __opaque;
} __attribute__((aligned(8)));
+struct bpf_rcu_head {
+ __u64 __opaque[6];
+} __attribute__((aligned(8)));
+
struct bpf_wq {
__u64 __opaque[2];
} __attribute__((aligned(8)));
diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index 314ecb0e593b0..4a1fa4fbdf4e8 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -3696,6 +3696,7 @@ static int btf_get_field_type(const struct btf *btf, const struct btf_type *var_
{ BPF_TIMER, "bpf_timer", true },
{ BPF_WORKQUEUE, "bpf_wq", true },
{ BPF_TASK_WORK, "bpf_task_work", true },
+ { BPF_RCU_HEAD, "bpf_rcu_head", true },
{ BPF_LIST_HEAD, "bpf_list_head", false },
{ BPF_LIST_NODE, "bpf_list_node", false },
{ BPF_RB_ROOT, "bpf_rb_root", false },
@@ -3881,6 +3882,7 @@ static int btf_find_field_one(const struct btf *btf,
case BPF_RB_NODE:
case BPF_REFCOUNT:
case BPF_TASK_WORK:
+ case BPF_RCU_HEAD:
ret = btf_find_struct(btf, var_type, off, sz, field_type,
info_cnt ? &info[0] : &tmp);
if (ret < 0)
@@ -4176,6 +4178,7 @@ struct btf_record *btf_parse_fields(const struct btf *btf, const struct btf_type
rec->wq_off = -EINVAL;
rec->refcount_off = -EINVAL;
rec->task_work_off = -EINVAL;
+ rec->rcu_head_off = -EINVAL;
for (i = 0; i < cnt; i++) {
field_type_size = btf_field_type_size(info_arr[i].type);
if (info_arr[i].off + field_type_size > value_size) {
@@ -4219,6 +4222,10 @@ struct btf_record *btf_parse_fields(const struct btf *btf, const struct btf_type
WARN_ON_ONCE(rec->task_work_off >= 0);
rec->task_work_off = rec->fields[i].offset;
break;
+ case BPF_RCU_HEAD:
+ WARN_ON_ONCE(rec->rcu_head_off >= 0);
+ rec->rcu_head_off = rec->fields[i].offset;
+ break;
case BPF_REFCOUNT:
WARN_ON_ONCE(rec->refcount_off >= 0);
/* Cache offset for faster lookup at runtime */
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 501c7ce35cba9..8a01dd4058a03 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -4805,6 +4805,80 @@ __bpf_kfunc int bpf_task_work_schedule_resume(struct task_struct *task, struct b
return bpf_task_work_schedule(task, tw, map__const_map, callback, aux, TWA_RESUME);
}
+typedef int (*bpf_rcu_callback_t)(struct bpf_map *map, void *key, void *value);
+
+/* Actual type for struct bpf_rcu_head */
+struct bpf_rcu_head_kern {
+ struct rcu_head rcu;
+ bpf_callback_t callback_fn;
+ struct bpf_map *map;
+ struct bpf_prog *prog;
+ u32 armed;
+} __aligned(8);
+
+static void bpf_rcu_run_callback(struct rcu_head *rcu)
+{
+ struct bpf_rcu_head_kern *rh = container_of(rcu, struct bpf_rcu_head_kern, rcu);
+ bpf_callback_t callback_fn = rh->callback_fn;
+ struct bpf_prog *prog = rh->prog;
+ struct bpf_map *map = rh->map;
+ void *value, *key;
+ u32 idx;
+
+ value = (void *)rh - map->record->rcu_head_off;
+ key = map_key_from_value(map, value, &idx);
+
+ /* Pairs with the arming cmpxchg(): rh may be re-armed as soon as this store lands. */
+ smp_store_release(&rh->armed, 0);
+
+ rcu_read_lock_dont_migrate();
+ callback_fn((u64)(long)map, (u64)(long)key, (u64)(long)value, 0, 0);
+ rcu_read_unlock_migrate();
+
+ bpf_prog_put(prog);
+}
+
+/**
+ * bpf_call_rcu - Invoke a BPF callback after an RCU grace period
+ * @rh: struct bpf_rcu_head in a BPF map value
+ * @map__const_map: bpf_map that embeds struct bpf_rcu_head in the values
+ * @callback: BPF subprogram, invoked as callback(map, key, value) for the value holding @rh
+ * @aux: bpf_prog_aux of the caller, implicitly set by the verifier
+ *
+ * Return: 0, -EBUSY if @rh is already queued, -EPERM if @map is held by neither a process
+ * nor bpffs, or -EBADF if the calling program is going away.
+ */
+__bpf_kfunc int bpf_call_rcu(struct bpf_rcu_head *rh, void *map__const_map,
+ bpf_rcu_callback_t callback, struct bpf_prog_aux *aux)
+{
+ struct bpf_rcu_head_kern *rhk = (void *)rh;
+ struct bpf_map *map = map__const_map;
+ struct bpf_prog *prog;
+
+ BUILD_BUG_ON(sizeof(struct bpf_rcu_head_kern) > sizeof(struct bpf_rcu_head));
+ BUILD_BUG_ON(__alignof__(struct bpf_rcu_head_kern) != __alignof__(struct bpf_rcu_head));
+ BTF_TYPE_EMIT(struct bpf_rcu_head);
+
+ /* A queued callback cannot be cancelled, so a self-rearming one would pin prog and map. */
+ if (!atomic64_read(&map->usercnt))
+ return -EPERM;
+
+ if (cmpxchg(&rhk->armed, 0, 1))
+ return -EBUSY;
+
+ prog = bpf_prog_inc_not_zero(aux->prog);
+ if (IS_ERR(prog)) {
+ WRITE_ONCE(rhk->armed, 0);
+ return -EBADF;
+ }
+
+ rhk->callback_fn = (bpf_callback_t)(void *)callback;
+ rhk->map = map;
+ rhk->prog = prog;
+ call_rcu(&rhk->rcu, bpf_rcu_run_callback);
+ return 0;
+}
+
static int make_file_dynptr(struct file *file, u32 flags, bool may_sleep,
struct bpf_dynptr_kern *ptr)
{
@@ -5100,6 +5174,7 @@ BTF_ID_FLAGS(func, bpf_stream_vprintk, KF_IMPLICIT_ARGS | KF_SPINLOCK_SAFE)
BTF_ID_FLAGS(func, bpf_stream_print_stack, KF_IMPLICIT_ARGS | KF_SPINLOCK_SAFE)
BTF_ID_FLAGS(func, bpf_task_work_schedule_signal, KF_IMPLICIT_ARGS)
BTF_ID_FLAGS(func, bpf_task_work_schedule_resume, KF_IMPLICIT_ARGS)
+BTF_ID_FLAGS(func, bpf_call_rcu, KF_IMPLICIT_ARGS)
BTF_ID_FLAGS(func, bpf_dynptr_from_file)
BTF_ID_FLAGS(func, bpf_dynptr_file_discard, KF_RELEASE)
BTF_ID_FLAGS(func, bpf_timer_cancel_async)
diff --git a/kernel/bpf/map_in_map.c b/kernel/bpf/map_in_map.c
index d2cbab4bdf644..5ee8aae7435ba 100644
--- a/kernel/bpf/map_in_map.c
+++ b/kernel/bpf/map_in_map.c
@@ -25,6 +25,10 @@ struct bpf_map *bpf_map_meta_alloc(int inner_map_ufd)
if (!inner_map->ops->map_meta_equal)
return ERR_PTR(-ENOTSUPP);
+ /* An inner map has no used_maps reference to hold it under a queued callback. */
+ if (btf_record_has_field(inner_map->record, BPF_RCU_HEAD))
+ return ERR_PTR(-EOPNOTSUPP);
+
inner_map_meta_size = sizeof(*inner_map_meta);
/* In some cases verifier needs to access beyond just base map. */
if (inner_map->ops == &array_map_ops || inner_map->ops == &percpu_array_map_ops)
diff --git a/kernel/bpf/map_iter.c b/kernel/bpf/map_iter.c
index c19b360bad9ea..2077d46d167c7 100644
--- a/kernel/bpf/map_iter.c
+++ b/kernel/bpf/map_iter.c
@@ -117,6 +117,12 @@ static int bpf_iter_attach_map(struct bpf_prog *prog,
goto put_map;
}
+ /* The value ctx arg is writable and aliases the live element. */
+ if (btf_record_has_field(map->record, BPF_RCU_HEAD)) {
+ err = -EOPNOTSUPP;
+ goto put_map;
+ }
+
if (map->map_type == BPF_MAP_TYPE_PERCPU_HASH ||
map->map_type == BPF_MAP_TYPE_LRU_PERCPU_HASH ||
map->map_type == BPF_MAP_TYPE_PERCPU_ARRAY)
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index 113486b15d29f..74496fd716d3b 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -687,6 +687,7 @@ void btf_record_free(struct btf_record *rec)
case BPF_REFCOUNT:
case BPF_WORKQUEUE:
case BPF_TASK_WORK:
+ case BPF_RCU_HEAD:
/* Nothing to release */
break;
default:
@@ -741,6 +742,7 @@ struct btf_record *btf_record_dup(const struct btf_record *rec)
case BPF_REFCOUNT:
case BPF_WORKQUEUE:
case BPF_TASK_WORK:
+ case BPF_RCU_HEAD:
/* Nothing to acquire */
break;
default:
@@ -874,6 +876,7 @@ void bpf_obj_free_fields(const struct btf_record *rec, void *obj)
case BPF_LIST_NODE:
case BPF_RB_NODE:
case BPF_REFCOUNT:
+ case BPF_RCU_HEAD:
break;
default:
WARN_ON_ONCE(1);
@@ -1280,7 +1283,7 @@ static int map_check_btf(struct bpf_map *map, struct bpf_token *token,
map->record = btf_parse_fields(btf, value_type,
BPF_SPIN_LOCK | BPF_RES_SPIN_LOCK | BPF_TIMER | BPF_KPTR | BPF_LIST_HEAD |
BPF_RB_ROOT | BPF_REFCOUNT | BPF_WORKQUEUE | BPF_UPTR |
- BPF_TASK_WORK,
+ BPF_TASK_WORK | BPF_RCU_HEAD,
map->value_size);
if (!IS_ERR_OR_NULL(map->record)) {
int i;
@@ -1322,6 +1325,12 @@ static int map_check_btf(struct bpf_map *map, struct bpf_token *token,
goto free_map_tab;
}
break;
+ case BPF_RCU_HEAD:
+ if (map->map_type != BPF_MAP_TYPE_ARRAY) {
+ ret = -EOPNOTSUPP;
+ goto free_map_tab;
+ }
+ break;
case BPF_KPTR_UNREF:
case BPF_KPTR_REF:
case BPF_KPTR_PERCPU:
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index d62c0f74cff5e..64db47964ff9f 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -535,6 +535,7 @@ static bool is_ptr_cast_function(enum bpf_func_id func_id)
static bool is_sync_callback_calling_kfunc(u32 btf_id);
static bool is_async_callback_calling_kfunc(u32 btf_id);
+static bool is_call_rcu_kfunc(u32 btf_id);
static bool is_callback_calling_kfunc(u32 btf_id);
static bool is_bpf_wq_set_callback_kfunc(u32 btf_id);
@@ -577,6 +578,10 @@ static bool is_async_cb_sleepable(struct bpf_verifier_env *env, struct bpf_insn
if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_timer_set_callback)
return false;
+ /* bpf_call_rcu callbacks are never sleepable. */
+ if (bpf_pseudo_kfunc_call(insn) && insn->off == 0 && is_call_rcu_kfunc(insn->imm))
+ return false;
+
/* bpf_wq and bpf_task_work callbacks are always sleepable. */
if (bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
(is_bpf_wq_set_callback_kfunc(insn->imm) || is_task_work_add_kfunc(insn->imm)))
@@ -7590,6 +7595,9 @@ static int check_map_field_pointer(struct bpf_verifier_env *env, struct bpf_reg_
case BPF_TASK_WORK:
field_off = map->record->task_work_off;
break;
+ case BPF_RCU_HEAD:
+ field_off = map->record->rcu_head_off;
+ break;
case BPF_WORKQUEUE:
field_off = map->record->wq_off;
break;
@@ -8448,6 +8456,7 @@ static const struct bpf_reg_types *compatible_reg_types[__BPF_ARG_TYPE_MAX] = {
[ARG_PTR_TO_RES_SPIN_LOCK] = &map_value_or_alloc_obj_types,
[ARG_PTR_TO_WORKQUEUE] = &map_value_types,
[ARG_PTR_TO_TASK_WORK] = &map_value_types,
+ [ARG_PTR_TO_RCU_HEAD] = &map_value_types,
[ARG_PTR_TO_IRQ_FLAG] = &stack_ptr_types,
[ARG_PTR_TO_ARENA] = &arena_types,
[ARG_PTR_TO_CTX_OUT] = &ctx_out_types,
@@ -8894,6 +8903,8 @@ static int process_map_ptr_arg(struct bpf_verifier_env *env, struct bpf_reg_stat
obj_name = "timer";
else if (rec->task_work_off >= 0)
obj_name = "bpf_task_work";
+ else if (rec->rcu_head_off >= 0)
+ obj_name = "bpf_rcu_head";
verbose(env, "%s pointer in %s map_uid=%d ",
obj_name, reg_arg_name(env, obj_argno), meta->map.uid);
@@ -9442,6 +9453,11 @@ static int check_func_arg(struct bpf_verifier_env *env, u32 arg, u32 slot, u32 p
if (err < 0)
return err;
break;
+ case ARG_PTR_TO_RCU_HEAD:
+ err = check_map_field_pointer(env, reg, argno, BPF_RCU_HEAD, &meta->map);
+ if (err < 0)
+ return err;
+ break;
case ARG_PTR_TO_IRQ_FLAG:
err = process_irq_flag(env, reg, argno, meta);
if (err < 0)
@@ -10955,6 +10971,40 @@ static int set_task_work_schedule_callback_state(struct bpf_verifier_env *env,
return 0;
}
+static int set_rcu_callback_state(struct bpf_verifier_env *env,
+ struct bpf_func_state *caller,
+ struct bpf_func_state *callee,
+ int insn_idx)
+{
+ struct bpf_map *map_ptr = caller->regs[BPF_REG_2].map_ptr;
+ u32 map_uid = caller->regs[BPF_REG_2].map_uid;
+
+ /*
+ * callback_fn(struct bpf_map *map, void *key, void *value);
+ */
+ callee->regs[BPF_REG_1].type = CONST_PTR_TO_MAP;
+ __mark_reg_known_zero(&callee->regs[BPF_REG_1]);
+ callee->regs[BPF_REG_1].map_ptr = map_ptr;
+ callee->regs[BPF_REG_1].map_uid = map_uid;
+
+ callee->regs[BPF_REG_2].type = PTR_TO_MAP_KEY;
+ __mark_reg_known_zero(&callee->regs[BPF_REG_2]);
+ callee->regs[BPF_REG_2].map_ptr = map_ptr;
+ callee->regs[BPF_REG_2].map_uid = map_uid;
+
+ callee->regs[BPF_REG_3].type = PTR_TO_MAP_VALUE;
+ __mark_reg_known_zero(&callee->regs[BPF_REG_3]);
+ callee->regs[BPF_REG_3].map_ptr = map_ptr;
+ callee->regs[BPF_REG_3].map_uid = map_uid;
+
+ /* unused */
+ bpf_mark_reg_not_init(env, &callee->regs[BPF_REG_4]);
+ bpf_mark_reg_not_init(env, &callee->regs[BPF_REG_5]);
+ callee->in_async_callback_fn = true;
+ callee->callback_ret_range = retval_range(S32_MIN, S32_MAX);
+ return 0;
+}
+
static bool is_rbtree_lock_required_kfunc(u32 btf_id);
static void account_processed_insn(struct bpf_verifier_env *env)
@@ -12216,7 +12266,8 @@ enum {
KF_ARG_RES_SPIN_LOCK_ID,
KF_ARG_TASK_WORK_ID,
KF_ARG_PROG_AUX_ID,
- KF_ARG_TIMER_ID
+ KF_ARG_TIMER_ID,
+ KF_ARG_RCU_HEAD_ID
};
BTF_ID_LIST(kf_arg_btf_ids)
@@ -12230,6 +12281,7 @@ BTF_ID(struct, bpf_res_spin_lock)
BTF_ID(struct, bpf_task_work)
BTF_ID(struct, bpf_prog_aux)
BTF_ID(struct, bpf_timer)
+BTF_ID(struct, bpf_rcu_head)
static bool __is_kfunc_ptr_arg_type(const struct btf *btf,
const struct btf_param *arg, int type)
@@ -12288,6 +12340,11 @@ static bool is_kfunc_arg_task_work(const struct btf *btf, const struct btf_param
return __is_kfunc_ptr_arg_type(btf, arg, KF_ARG_TASK_WORK_ID);
}
+static bool is_kfunc_arg_rcu_head(const struct btf *btf, const struct btf_param *arg)
+{
+ return __is_kfunc_ptr_arg_type(btf, arg, KF_ARG_RCU_HEAD_ID);
+}
+
static bool is_kfunc_arg_res_spin_lock(const struct btf *btf, const struct btf_param *arg)
{
return __is_kfunc_ptr_arg_type(btf, arg, KF_ARG_RES_SPIN_LOCK_ID);
@@ -12571,6 +12628,7 @@ enum special_kfunc_type {
KF___bpf_trap,
KF_bpf_task_work_schedule_signal,
KF_bpf_task_work_schedule_resume,
+ KF_bpf_call_rcu,
KF_bpf_arena_alloc_pages,
KF_bpf_arena_free_pages,
KF_bpf_arena_reserve_pages,
@@ -12664,6 +12722,7 @@ BTF_ID(func, bpf_dynptr_file_discard)
BTF_ID(func, __bpf_trap)
BTF_ID(func, bpf_task_work_schedule_signal)
BTF_ID(func, bpf_task_work_schedule_resume)
+BTF_ID(func, bpf_call_rcu)
BTF_ID(func, bpf_arena_alloc_pages)
BTF_ID(func, bpf_arena_free_pages)
BTF_ID(func, bpf_arena_reserve_pages)
@@ -12735,6 +12794,11 @@ static bool is_bpf_rbtree_add_kfunc(u32 func_id)
func_id == special_kfunc_list[KF_bpf_rbtree_add_impl];
}
+static bool is_call_rcu_kfunc(u32 func_id)
+{
+ return func_id == special_kfunc_list[KF_bpf_call_rcu];
+}
+
static bool is_task_work_add_kfunc(u32 func_id)
{
return func_id == special_kfunc_list[KF_bpf_task_work_schedule_signal] ||
@@ -12941,6 +13005,8 @@ get_kfunc_arg_type(struct bpf_verifier_env *env, struct bpf_call_arg_meta *meta,
arg_type = ARG_PTR_TO_TIMER;
else if (is_kfunc_arg_task_work(meta->btf, &args[arg]))
arg_type = ARG_PTR_TO_TASK_WORK;
+ else if (is_kfunc_arg_rcu_head(meta->btf, &args[arg]))
+ arg_type = ARG_PTR_TO_RCU_HEAD;
else if (is_kfunc_arg_irq_flag(meta->btf, &args[arg]))
arg_type = ARG_PTR_TO_IRQ_FLAG;
else if (is_kfunc_arg_res_spin_lock(meta->btf, &args[arg]))
@@ -13442,7 +13508,8 @@ static bool is_sync_callback_calling_kfunc(u32 btf_id)
static bool is_async_callback_calling_kfunc(u32 btf_id)
{
return is_bpf_wq_set_callback_kfunc(btf_id) ||
- is_task_work_add_kfunc(btf_id);
+ is_task_work_add_kfunc(btf_id) ||
+ is_call_rcu_kfunc(btf_id);
}
bool bpf_is_throw_kfunc(struct bpf_insn *insn)
@@ -14259,6 +14326,16 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
}
}
+ if (is_call_rcu_kfunc(meta.func_id)) {
+ err = push_callback_call(env, insn, insn_idx, meta.subprogno,
+ set_rcu_callback_state);
+ if (err) {
+ verbose(env, "kfunc %s#%d failed callback verification\n",
+ func_name, meta.func_id);
+ return err;
+ }
+ }
+
rcu_lock = is_kfunc_bpf_rcu_read_lock(&meta);
rcu_unlock = is_kfunc_bpf_rcu_read_unlock(&meta);
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 6330b7d745c57..0aaa54359aebc 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -7611,6 +7611,10 @@ struct bpf_task_work {
__u64 __opaque;
} __attribute__((aligned(8)));
+struct bpf_rcu_head {
+ __u64 __opaque[6];
+} __attribute__((aligned(8)));
+
struct bpf_wq {
__u64 __opaque[2];
} __attribute__((aligned(8)));
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH bpf-next v6 2/4] selftests/bpf: Add tests for bpf_call_rcu()
2026-09-22 20:00 [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace() Puranjay Mohan
2026-09-22 20:00 ` [PATCH bpf-next v6 1/4] bpf: Add bpf_call_rcu() kfunc Puranjay Mohan
@ 2026-09-22 20:00 ` Puranjay Mohan
2026-09-22 20:00 ` [PATCH bpf-next v6 3/4] bpf: Add bpf_call_rcu_tasks_trace() kfunc Puranjay Mohan
` (3 subsequent siblings)
5 siblings, 0 replies; 17+ messages in thread
From: Puranjay Mohan @ 2026-09-22 20:00 UTC (permalink / raw)
To: bpf, rcu
Cc: Puranjay Mohan, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman,
Kumar Kartikeya Dwivedi, Song Liu, Yonghong Song,
Harry Yoo (Oracle), Paul E. McKenney
Cover the callback running after a grace period with the right map, key
and value, -EBUSY on a second arm, reuse of the head once disarmed, a
callback arming itself again, and teardown with a callback queued.
struct bpf_rcu_head is not the first member of the map value, so the
callback's recovery of the value from the head is exercised. arm()
wraps both arms in an RCU read section, otherwise a grace period may
elapse between them and the second one legitimately succeeds.
Negative tests: a hash map created with the same BTF, the map used as an
inner map, and an iterator attach, all checked for -EOPNOTSUPP; plus
verifier rejection of a mismatched map, a map with no bpf_rcu_head, a
head at the wrong offset, a head on the stack, and a sleepable callback.
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
---
.../selftests/bpf/prog_tests/call_rcu.c | 271 ++++++++++++++++++
tools/testing/selftests/bpf/progs/call_rcu.c | 93 ++++++
.../selftests/bpf/progs/call_rcu_fail.c | 114 ++++++++
3 files changed, 478 insertions(+)
create mode 100644 tools/testing/selftests/bpf/prog_tests/call_rcu.c
create mode 100644 tools/testing/selftests/bpf/progs/call_rcu.c
create mode 100644 tools/testing/selftests/bpf/progs/call_rcu_fail.c
diff --git a/tools/testing/selftests/bpf/prog_tests/call_rcu.c b/tools/testing/selftests/bpf/prog_tests/call_rcu.c
new file mode 100644
index 0000000000000..bfa6512b9052a
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/call_rcu.c
@@ -0,0 +1,271 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <test_progs.h>
+#include "call_rcu.skel.h"
+#include "call_rcu_fail.skel.h"
+
+struct elem {
+ __u64 pad;
+ struct bpf_rcu_head rh;
+ __u64 val;
+};
+
+/*
+ * Force the grace period, then poll: the callback still has to be invoked
+ * afterwards, and call_rcu() is lazy on a CONFIG_RCU_LAZY kernel.
+ */
+static bool wait_for_callbacks(struct call_rcu *skel, int expected)
+{
+ int i, got = 0;
+
+ kern_sync_rcu();
+ for (i = 0; i < 3000; i++) {
+ got = __atomic_load_n(&skel->bss->callbacks, __ATOMIC_ACQUIRE);
+ if (got >= expected)
+ return true;
+ usleep(10000);
+ }
+ return ASSERT_GE(got, expected, "callbacks");
+}
+
+static void test_call_rcu_run(void)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ struct call_rcu *skel;
+ struct elem elem;
+ __u32 key = 1;
+ int err;
+
+ skel = call_rcu__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "skel_open_and_load"))
+ return;
+
+ err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.arm), &opts);
+ if (!ASSERT_OK(err, "test_run") || !ASSERT_EQ(opts.retval, 0, "retval"))
+ goto out;
+
+ ASSERT_EQ(skel->bss->arm_err, 0, "arm_err");
+ ASSERT_EQ(skel->bss->busy_err, -EBUSY, "busy_err");
+
+ if (!wait_for_callbacks(skel, 1))
+ goto out;
+
+ ASSERT_EQ(skel->bss->cb_key, key, "cb_key");
+ ASSERT_EQ(skel->bss->cb_val, 0xdeadbeef, "cb_val");
+ ASSERT_EQ(skel->bss->cb_max_entries, bpf_map__max_entries(skel->maps.arr), "cb_map");
+
+ err = bpf_map__lookup_elem(skel->maps.arr, &key, sizeof(key), &elem, sizeof(elem), 0);
+ if (ASSERT_OK(err, "lookup"))
+ ASSERT_EQ(elem.val, 0, "value_cleared");
+
+ /* The head is disarmed before the callback runs, so it can be reused. */
+ err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.arm), &opts);
+ if (!ASSERT_OK(err, "test_run_again"))
+ goto out;
+ ASSERT_EQ(skel->bss->arm_err, 0, "rearm_err");
+ wait_for_callbacks(skel, 2);
+out:
+ call_rcu__destroy(skel);
+}
+
+static void test_call_rcu_chain(void)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ struct call_rcu *skel;
+
+ skel = call_rcu__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "skel_open_and_load"))
+ return;
+
+ skel->bss->chain = 1;
+ if (!ASSERT_OK(bpf_prog_test_run_opts(bpf_program__fd(skel->progs.arm), &opts), "test_run"))
+ goto out;
+
+ if (!wait_for_callbacks(skel, 1))
+ goto out;
+ ASSERT_EQ(skel->bss->chain_err, 0, "chain_err");
+ wait_for_callbacks(skel, 2);
+out:
+ call_rcu__destroy(skel);
+}
+
+/*
+ * A callback that keeps re-arming must stop once the map loses its last user
+ * reference, or it pins the program for good. Hold an independent fd on .bss
+ * so the refusal is still readable after the skeleton is gone.
+ */
+static void test_call_rcu_teardown(void)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ int i, err, fd = 0, bss_fd = -1;
+ struct bpf_prog_info pinfo = {};
+ struct bpf_map_info minfo = {};
+ __u32 len, prog_id, zero = 0;
+ struct call_rcu *skel;
+ char *buf = NULL;
+ size_t off, vsz;
+
+ skel = call_rcu__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "skel_open_and_load"))
+ return;
+
+ len = sizeof(pinfo);
+ if (!ASSERT_OK(bpf_prog_get_info_by_fd(bpf_program__fd(skel->progs.arm), &pinfo, &len),
+ "prog_info"))
+ goto out;
+ prog_id = pinfo.id;
+
+ len = sizeof(minfo);
+ if (!ASSERT_OK(bpf_map_get_info_by_fd(bpf_map__fd(skel->maps.bss), &minfo, &len),
+ "bss_info"))
+ goto out;
+ bss_fd = bpf_map_get_fd_by_id(minfo.id);
+ if (!ASSERT_GE(bss_fd, 0, "bss_fd"))
+ goto out;
+
+ vsz = bpf_map__value_size(skel->maps.bss);
+ off = (char *)&skel->bss->chain_err - (char *)skel->bss;
+ buf = malloc(vsz);
+ if (!ASSERT_OK_PTR(buf, "buf"))
+ goto out;
+
+ skel->bss->chain = INT_MAX;
+ err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.arm), &opts);
+ if (!ASSERT_OK(err, "test_run") || !ASSERT_EQ(opts.retval, 0, "retval"))
+ goto out;
+ if (!ASSERT_EQ(skel->bss->arm_err, 0, "arm_err"))
+ goto out;
+ /* The chain has to be running before the map reference goes away. */
+ if (!wait_for_callbacks(skel, 2))
+ goto out;
+
+ call_rcu__destroy(skel);
+ skel = NULL;
+
+ for (i = 0; i < 3000; i++) {
+ fd = bpf_prog_get_fd_by_id(prog_id);
+ if (fd < 0)
+ break;
+ close(fd);
+ usleep(10000);
+ }
+ if (!ASSERT_EQ(fd, -ENOENT, "prog_freed"))
+ goto out;
+
+ if (ASSERT_OK(bpf_map_lookup_elem(bss_fd, &zero, buf), "bss_lookup"))
+ ASSERT_EQ(*(int *)(buf + off), -EPERM, "chain_refused");
+out:
+ free(buf);
+ if (bss_fd >= 0)
+ close(bss_fd);
+ call_rcu__destroy(skel);
+}
+
+static void test_call_rcu_bad_map(void)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts);
+ struct call_rcu *skel;
+ int fd;
+
+ skel = call_rcu__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "skel_open_and_load"))
+ return;
+
+ opts.btf_fd = bpf_object__btf_fd(skel->obj);
+ opts.btf_key_type_id = bpf_map__btf_key_type_id(skel->maps.arr);
+ opts.btf_value_type_id = bpf_map__btf_value_type_id(skel->maps.arr);
+
+ fd = bpf_map_create(BPF_MAP_TYPE_HASH, "rcu_hash", sizeof(__u32),
+ bpf_map__value_size(skel->maps.arr), 1, &opts);
+ ASSERT_EQ(fd, -EOPNOTSUPP, "hash_rejected");
+ if (fd >= 0)
+ close(fd);
+
+ call_rcu__destroy(skel);
+}
+
+/* Iterating would hand the program a writable pointer to the head. */
+static void test_call_rcu_iter(void)
+{
+ LIBBPF_OPTS(bpf_iter_attach_opts, opts);
+ union bpf_iter_link_info linfo = {};
+ struct bpf_link *link;
+ struct call_rcu *skel;
+
+ skel = call_rcu__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "skel_open_and_load"))
+ return;
+
+ linfo.map.map_fd = bpf_map__fd(skel->maps.arr);
+ opts.link_info = &linfo;
+ opts.link_info_len = sizeof(linfo);
+
+ link = bpf_program__attach_iter(skel->progs.dump, &opts);
+ if (!ASSERT_ERR_PTR(link, "iter_rejected"))
+ bpf_link__destroy(link);
+ else
+ ASSERT_EQ(libbpf_get_error(link), -EOPNOTSUPP, "iter_errno");
+
+ call_rcu__destroy(skel);
+}
+
+static void test_call_rcu_inner_map(void)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts);
+ struct call_rcu *skel;
+ int fd;
+
+ skel = call_rcu__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "skel_open_and_load"))
+ return;
+
+ opts.inner_map_fd = bpf_map__fd(skel->maps.arr);
+ fd = bpf_map_create(BPF_MAP_TYPE_ARRAY_OF_MAPS, "rcu_outer",
+ sizeof(__u32), sizeof(__u32), 1, &opts);
+ ASSERT_EQ(fd, -EOPNOTSUPP, "inner_map_rejected");
+ if (fd >= 0)
+ close(fd);
+
+ call_rcu__destroy(skel);
+}
+
+/* Two heads in the same map must be independent. */
+static void test_call_rcu_two_heads(void)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ struct call_rcu *skel;
+
+ skel = call_rcu__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "skel_open_and_load"))
+ return;
+
+ if (!ASSERT_OK(bpf_prog_test_run_opts(bpf_program__fd(skel->progs.arm_both), &opts),
+ "test_run"))
+ goto out;
+ if (!ASSERT_EQ(skel->bss->arm_err, 0, "arm_err"))
+ goto out;
+ if (!wait_for_callbacks(skel, 2))
+ goto out;
+ ASSERT_EQ(skel->bss->cb_keys, 0x3, "both_keys");
+out:
+ call_rcu__destroy(skel);
+}
+
+void test_call_rcu(void)
+{
+ if (test__start_subtest("run"))
+ test_call_rcu_run();
+ if (test__start_subtest("chain"))
+ test_call_rcu_chain();
+ if (test__start_subtest("two_heads"))
+ test_call_rcu_two_heads();
+ if (test__start_subtest("teardown"))
+ test_call_rcu_teardown();
+ if (test__start_subtest("bad_map"))
+ test_call_rcu_bad_map();
+ if (test__start_subtest("iter"))
+ test_call_rcu_iter();
+ if (test__start_subtest("inner_map"))
+ test_call_rcu_inner_map();
+ RUN_TESTS(call_rcu_fail);
+}
diff --git a/tools/testing/selftests/bpf/progs/call_rcu.c b/tools/testing/selftests/bpf/progs/call_rcu.c
new file mode 100644
index 0000000000000..94d92cee88a70
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/call_rcu.c
@@ -0,0 +1,93 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+
+char _license[] SEC("license") = "GPL";
+
+/* rh is not at offset 0, so the callback's value recovery is exercised. */
+struct elem {
+ __u64 pad;
+ struct bpf_rcu_head rh;
+ __u64 val;
+};
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 2);
+ __type(key, __u32);
+ __type(value, struct elem);
+} arr SEC(".maps");
+
+__u32 cb_key;
+__u64 cb_keys;
+__u64 cb_val;
+__u32 cb_max_entries;
+int callbacks;
+int arm_err;
+int busy_err;
+int chain; /* set by userspace: number of times to re-arm from the callback */
+int chain_err;
+
+static int reclaim(struct bpf_map *map, void *key, void *value)
+{
+ struct elem *e = value;
+
+ cb_key = *(__u32 *)key;
+ __sync_fetch_and_or(&cb_keys, 1ULL << cb_key);
+ cb_val = e->val;
+ cb_max_entries = map->max_entries;
+ e->val = 0;
+
+ if (chain > 0) {
+ chain--;
+ chain_err = bpf_call_rcu(&e->rh, &arr, reclaim);
+ }
+
+ __sync_fetch_and_add(&callbacks, 1);
+ return 0;
+}
+
+SEC("syscall")
+int arm(void *ctx)
+{
+ __u32 key = 1;
+ struct elem *e;
+
+ e = bpf_map_lookup_elem(&arr, &key);
+ if (!e)
+ return 1;
+
+ e->val = 0xdeadbeef;
+ /* Keep a grace period from elapsing between the two arms. */
+ bpf_rcu_read_lock();
+ arm_err = bpf_call_rcu(&e->rh, &arr, reclaim);
+ busy_err = bpf_call_rcu(&e->rh, &arr, reclaim);
+ bpf_rcu_read_unlock();
+ return 0;
+}
+
+SEC("syscall")
+int arm_both(void *ctx)
+{
+ __u32 key0 = 0, key1 = 1;
+ struct elem *e0, *e1;
+
+ e0 = bpf_map_lookup_elem(&arr, &key0);
+ e1 = bpf_map_lookup_elem(&arr, &key1);
+ if (!e0 || !e1)
+ return 1;
+
+ e0->val = 0xdeadbeef;
+ e1->val = 0xdeadbeef;
+ arm_err = bpf_call_rcu(&e0->rh, &arr, reclaim);
+ arm_err |= bpf_call_rcu(&e1->rh, &arr, reclaim);
+ return 0;
+}
+
+SEC("iter/bpf_map_elem")
+int dump(struct bpf_iter__bpf_map_elem *ctx)
+{
+ return 0;
+}
diff --git a/tools/testing/selftests/bpf/progs/call_rcu_fail.c b/tools/testing/selftests/bpf/progs/call_rcu_fail.c
new file mode 100644
index 0000000000000..bf60a27731fb3
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/call_rcu_fail.c
@@ -0,0 +1,114 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+
+char _license[] SEC("license") = "GPL";
+
+const void *user_ptr = NULL;
+
+struct elem {
+ __u64 pad;
+ struct bpf_rcu_head rh;
+ __u64 val;
+};
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __type(key, __u32);
+ __type(value, struct elem);
+} arr SEC(".maps");
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __type(key, __u32);
+ __type(value, struct elem);
+} arr2 SEC(".maps");
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __type(key, __u32);
+ __type(value, __u64);
+} plain SEC(".maps");
+
+__u32 key = 0;
+
+static int reclaim(struct bpf_map *map, void *key, void *value)
+{
+ return 0;
+}
+
+static int sleepable_reclaim(struct bpf_map *map, void *key, void *value)
+{
+ struct elem *e = value;
+
+ bpf_copy_from_user(&e->val, sizeof(e->val), user_ptr);
+ return 0;
+}
+
+SEC("syscall")
+__failure __msg("bpf_rcu_head pointer in R1") __msg("doesn't match map pointer in R2")
+int mismatch_map(void *ctx)
+{
+ struct elem *e;
+
+ e = bpf_map_lookup_elem(&arr, &key);
+ if (!e)
+ return 0;
+ bpf_call_rcu(&e->rh, &arr2, reclaim);
+ return 0;
+}
+
+SEC("syscall")
+__failure __msg("map 'plain' has no valid bpf_rcu_head")
+int no_rcu_head(void *ctx)
+{
+ __u64 *val;
+
+ val = bpf_map_lookup_elem(&plain, &key);
+ if (!val)
+ return 0;
+ bpf_call_rcu((struct bpf_rcu_head *)val, &plain, reclaim);
+ return 0;
+}
+
+SEC("syscall")
+__failure __msg("doesn't point to 'struct bpf_rcu_head' that is at 8")
+int wrong_offset(void *ctx)
+{
+ struct elem *e;
+
+ e = bpf_map_lookup_elem(&arr, &key);
+ if (!e)
+ return 0;
+ bpf_call_rcu((struct bpf_rcu_head *)&e->pad, &arr, reclaim);
+ return 0;
+}
+
+SEC("syscall")
+__failure __msg("R1 type=fp expected=map_value")
+int rcu_head_on_stack(void *ctx)
+{
+ struct bpf_rcu_head rh;
+
+ bpf_call_rcu(&rh, &arr, reclaim);
+ return 0;
+}
+
+SEC("syscall")
+__failure __msg("sleepable helper bpf_copy_from_user") __msg("in non-sleepable prog")
+int sleepable_callback(void *ctx)
+{
+ struct elem *e;
+
+ e = bpf_map_lookup_elem(&arr, &key);
+ if (!e)
+ return 0;
+ bpf_call_rcu(&e->rh, &arr, sleepable_reclaim);
+ return 0;
+}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH bpf-next v6 3/4] bpf: Add bpf_call_rcu_tasks_trace() kfunc
2026-09-22 20:00 [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace() Puranjay Mohan
2026-09-22 20:00 ` [PATCH bpf-next v6 1/4] bpf: Add bpf_call_rcu() kfunc Puranjay Mohan
2026-09-22 20:00 ` [PATCH bpf-next v6 2/4] selftests/bpf: Add tests for bpf_call_rcu() Puranjay Mohan
@ 2026-09-22 20:00 ` Puranjay Mohan
2026-09-22 20:37 ` sashiko-bot
2026-09-22 20:00 ` [PATCH bpf-next v6 4/4] selftests/bpf: Add a test for bpf_call_rcu_tasks_trace() Puranjay Mohan
` (2 subsequent siblings)
5 siblings, 1 reply; 17+ messages in thread
From: Puranjay Mohan @ 2026-09-22 20:00 UTC (permalink / raw)
To: bpf, rcu
Cc: Puranjay Mohan, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman,
Kumar Kartikeya Dwivedi, Song Liu, Yonghong Song,
Harry Yoo (Oracle), Paul E. McKenney
Sleepable BPF programs hold rcu_read_lock_trace(), not rcu_read_lock(),
so a plain RCU grace period does not wait for them. A program whose
readers are sleepable needs this flavour to defer reclaim safely.
call_rcu_tasks_trace() is call_srcu() on rcu_tasks_trace_srcu_struct and
SRCU invokes callbacks with BH disabled, so the callback is still not
sleepable. It does run from a kworker rather than softirq or the
rcuc/rcuo kthread, so a callback must not assume anything about current.
Only the queueing call differs, so the two share struct bpf_rcu_head and
all of the verifier plumbing.
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
---
kernel/bpf/helpers.c | 57 +++++++++++++++++++++++++++++++------------
kernel/bpf/verifier.c | 7 ++++--
2 files changed, 47 insertions(+), 17 deletions(-)
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 8a01dd4058a03..301b35bd85c8a 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -4838,21 +4838,10 @@ static void bpf_rcu_run_callback(struct rcu_head *rcu)
bpf_prog_put(prog);
}
-/**
- * bpf_call_rcu - Invoke a BPF callback after an RCU grace period
- * @rh: struct bpf_rcu_head in a BPF map value
- * @map__const_map: bpf_map that embeds struct bpf_rcu_head in the values
- * @callback: BPF subprogram, invoked as callback(map, key, value) for the value holding @rh
- * @aux: bpf_prog_aux of the caller, implicitly set by the verifier
- *
- * Return: 0, -EBUSY if @rh is already queued, -EPERM if @map is held by neither a process
- * nor bpffs, or -EBADF if the calling program is going away.
- */
-__bpf_kfunc int bpf_call_rcu(struct bpf_rcu_head *rh, void *map__const_map,
- bpf_rcu_callback_t callback, struct bpf_prog_aux *aux)
+static int __bpf_call_rcu(struct bpf_rcu_head *rh, struct bpf_map *map, void *callback,
+ struct bpf_prog_aux *aux, bool trace)
{
struct bpf_rcu_head_kern *rhk = (void *)rh;
- struct bpf_map *map = map__const_map;
struct bpf_prog *prog;
BUILD_BUG_ON(sizeof(struct bpf_rcu_head_kern) > sizeof(struct bpf_rcu_head));
@@ -4872,13 +4861,50 @@ __bpf_kfunc int bpf_call_rcu(struct bpf_rcu_head *rh, void *map__const_map,
return -EBADF;
}
- rhk->callback_fn = (bpf_callback_t)(void *)callback;
+ rhk->callback_fn = (bpf_callback_t)callback;
rhk->map = map;
rhk->prog = prog;
- call_rcu(&rhk->rcu, bpf_rcu_run_callback);
+ if (trace)
+ call_rcu_tasks_trace(&rhk->rcu, bpf_rcu_run_callback);
+ else
+ call_rcu(&rhk->rcu, bpf_rcu_run_callback);
return 0;
}
+/**
+ * bpf_call_rcu - Invoke a BPF callback after an RCU grace period
+ * @rh: struct bpf_rcu_head in a BPF map value
+ * @map__const_map: bpf_map that embeds struct bpf_rcu_head in the values
+ * @callback: BPF subprogram, invoked as callback(map, key, value) for the value holding @rh
+ * @aux: bpf_prog_aux of the caller, implicitly set by the verifier
+ *
+ * Return: 0, -EBUSY if @rh is already queued, -EPERM if @map is held by neither a process
+ * nor bpffs, or -EBADF if the calling program is going away.
+ */
+__bpf_kfunc int bpf_call_rcu(struct bpf_rcu_head *rh, void *map__const_map,
+ bpf_rcu_callback_t callback, struct bpf_prog_aux *aux)
+{
+ return __bpf_call_rcu(rh, map__const_map, callback, aux, false);
+}
+
+/**
+ * bpf_call_rcu_tasks_trace - Invoke a BPF callback after an RCU tasks trace grace period
+ * @rh: struct bpf_rcu_head in a BPF map value
+ * @map__const_map: bpf_map that embeds struct bpf_rcu_head in the values
+ * @callback: BPF subprogram, invoked as callback(map, key, value) for the value holding @rh
+ * @aux: bpf_prog_aux of the caller, implicitly set by the verifier
+ *
+ * Waits for sleepable BPF programs too. The callback itself is not sleepable either way.
+ *
+ * Return: 0, -EBUSY if @rh is already queued, -EPERM if @map is held by neither a process
+ * nor bpffs, or -EBADF if the calling program is going away.
+ */
+__bpf_kfunc int bpf_call_rcu_tasks_trace(struct bpf_rcu_head *rh, void *map__const_map,
+ bpf_rcu_callback_t callback, struct bpf_prog_aux *aux)
+{
+ return __bpf_call_rcu(rh, map__const_map, callback, aux, true);
+}
+
static int make_file_dynptr(struct file *file, u32 flags, bool may_sleep,
struct bpf_dynptr_kern *ptr)
{
@@ -5175,6 +5201,7 @@ BTF_ID_FLAGS(func, bpf_stream_print_stack, KF_IMPLICIT_ARGS | KF_SPINLOCK_SAFE)
BTF_ID_FLAGS(func, bpf_task_work_schedule_signal, KF_IMPLICIT_ARGS)
BTF_ID_FLAGS(func, bpf_task_work_schedule_resume, KF_IMPLICIT_ARGS)
BTF_ID_FLAGS(func, bpf_call_rcu, KF_IMPLICIT_ARGS)
+BTF_ID_FLAGS(func, bpf_call_rcu_tasks_trace, KF_IMPLICIT_ARGS)
BTF_ID_FLAGS(func, bpf_dynptr_from_file)
BTF_ID_FLAGS(func, bpf_dynptr_file_discard, KF_RELEASE)
BTF_ID_FLAGS(func, bpf_timer_cancel_async)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 64db47964ff9f..a7c9e2d8965d5 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -578,7 +578,7 @@ static bool is_async_cb_sleepable(struct bpf_verifier_env *env, struct bpf_insn
if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_timer_set_callback)
return false;
- /* bpf_call_rcu callbacks are never sleepable. */
+ /* bpf_call_rcu and bpf_call_rcu_tasks_trace callbacks are never sleepable. */
if (bpf_pseudo_kfunc_call(insn) && insn->off == 0 && is_call_rcu_kfunc(insn->imm))
return false;
@@ -12629,6 +12629,7 @@ enum special_kfunc_type {
KF_bpf_task_work_schedule_signal,
KF_bpf_task_work_schedule_resume,
KF_bpf_call_rcu,
+ KF_bpf_call_rcu_tasks_trace,
KF_bpf_arena_alloc_pages,
KF_bpf_arena_free_pages,
KF_bpf_arena_reserve_pages,
@@ -12723,6 +12724,7 @@ BTF_ID(func, __bpf_trap)
BTF_ID(func, bpf_task_work_schedule_signal)
BTF_ID(func, bpf_task_work_schedule_resume)
BTF_ID(func, bpf_call_rcu)
+BTF_ID(func, bpf_call_rcu_tasks_trace)
BTF_ID(func, bpf_arena_alloc_pages)
BTF_ID(func, bpf_arena_free_pages)
BTF_ID(func, bpf_arena_reserve_pages)
@@ -12796,7 +12798,8 @@ static bool is_bpf_rbtree_add_kfunc(u32 func_id)
static bool is_call_rcu_kfunc(u32 func_id)
{
- return func_id == special_kfunc_list[KF_bpf_call_rcu];
+ return func_id == special_kfunc_list[KF_bpf_call_rcu] ||
+ func_id == special_kfunc_list[KF_bpf_call_rcu_tasks_trace];
}
static bool is_task_work_add_kfunc(u32 func_id)
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH bpf-next v6 3/4] bpf: Add bpf_call_rcu_tasks_trace() kfunc
2026-09-22 20:00 ` [PATCH bpf-next v6 3/4] bpf: Add bpf_call_rcu_tasks_trace() kfunc Puranjay Mohan
@ 2026-09-22 20:37 ` sashiko-bot
0 siblings, 0 replies; 17+ messages in thread
From: sashiko-bot @ 2026-09-22 20:37 UTC (permalink / raw)
To: Puranjay Mohan; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: bpf_call_rcu_tasks_trace() deadlock from NMI context
--
commit 47fa2f5e11782172e769349f78bd62857f9572eb
Author: Puranjay Mohan <puranjay@kernel.org>
bpf: Add bpf_call_rcu_tasks_trace() kfunc
This commit introduces the bpf_call_rcu_tasks_trace() kfunc to allow
sleepable BPF programs to safely defer reclaim using a tasks trace RCU
grace period. It shares verifier plumbing and struct bpf_rcu_head with
bpf_call_rcu(), differing only in the queueing call.
> diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
> --- a/kernel/bpf/helpers.c
> +++ b/kernel/bpf/helpers.c
> @@ -4872,13 +4861,50 @@ __bpf_kfunc int bpf_call_rcu(struct bpf_rcu_head *rh, void *map__const_map,
[ ... ]
> - rhk->callback_fn = (bpf_callback_t)(void *)callback;
> + rhk->callback_fn = (bpf_callback_t)callback;
> rhk->map = map;
> rhk->prog = prog;
> - call_rcu(&rhk->rcu, bpf_rcu_run_callback);
> + if (trace)
> + call_rcu_tasks_trace(&rhk->rcu, bpf_rcu_run_callback);
> + else
> + call_rcu(&rhk->rcu, bpf_rcu_run_callback);
> return 0;
> }
[Severity: High]
Can this sequence lead to a kernel deadlock if called from NMI context?
If a BPF tracing program executing in NMI context (or with IRQs disabled)
calls the bpf_call_rcu_tasks_trace() kfunc, it will unconditionally invoke
call_rcu_tasks_trace(), which is a wrapper for call_srcu().
Since call_srcu() requires acquiring a raw_spinlock_t (sdp->lock), if the
NMI interrupted a task on the same CPU that was already holding this
spinlock, the NMI handler will spin indefinitely attempting to acquire it.
Should this code defer call_rcu_tasks_trace() via irq_work_queue if
irqs_disabled() is true, similar to how it is handled in
bpf_async_refcount_put()?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260922200208.3203834-1-puranjay@kernel.org?part=3
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH bpf-next v6 4/4] selftests/bpf: Add a test for bpf_call_rcu_tasks_trace()
2026-09-22 20:00 [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace() Puranjay Mohan
` (2 preceding siblings ...)
2026-09-22 20:00 ` [PATCH bpf-next v6 3/4] bpf: Add bpf_call_rcu_tasks_trace() kfunc Puranjay Mohan
@ 2026-09-22 20:00 ` Puranjay Mohan
2026-09-22 23:41 ` [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace() Andrii Nakryiko
2026-09-22 23:50 ` patchwork-bot+netdevbpf
5 siblings, 0 replies; 17+ messages in thread
From: Puranjay Mohan @ 2026-09-22 20:00 UTC (permalink / raw)
To: bpf, rcu
Cc: Puranjay Mohan, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman,
Kumar Kartikeya Dwivedi, Song Liu, Yonghong Song,
Harry Yoo (Oracle), Paul E. McKenney
Run the same functional test against the tasks trace flavour.
Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
---
.../testing/selftests/bpf/prog_tests/call_rcu.c | 12 ++++++++----
tools/testing/selftests/bpf/progs/call_rcu.c | 17 +++++++++++++++++
2 files changed, 25 insertions(+), 4 deletions(-)
diff --git a/tools/testing/selftests/bpf/prog_tests/call_rcu.c b/tools/testing/selftests/bpf/prog_tests/call_rcu.c
index bfa6512b9052a..e4d271d43d878 100644
--- a/tools/testing/selftests/bpf/prog_tests/call_rcu.c
+++ b/tools/testing/selftests/bpf/prog_tests/call_rcu.c
@@ -28,9 +28,10 @@ static bool wait_for_callbacks(struct call_rcu *skel, int expected)
return ASSERT_GE(got, expected, "callbacks");
}
-static void test_call_rcu_run(void)
+static void test_call_rcu_run(bool trace)
{
LIBBPF_OPTS(bpf_test_run_opts, opts);
+ struct bpf_program *prog;
struct call_rcu *skel;
struct elem elem;
__u32 key = 1;
@@ -40,7 +41,8 @@ static void test_call_rcu_run(void)
if (!ASSERT_OK_PTR(skel, "skel_open_and_load"))
return;
- err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.arm), &opts);
+ prog = trace ? skel->progs.arm_trace : skel->progs.arm;
+ err = bpf_prog_test_run_opts(bpf_program__fd(prog), &opts);
if (!ASSERT_OK(err, "test_run") || !ASSERT_EQ(opts.retval, 0, "retval"))
goto out;
@@ -59,7 +61,7 @@ static void test_call_rcu_run(void)
ASSERT_EQ(elem.val, 0, "value_cleared");
/* The head is disarmed before the callback runs, so it can be reused. */
- err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.arm), &opts);
+ err = bpf_prog_test_run_opts(bpf_program__fd(prog), &opts);
if (!ASSERT_OK(err, "test_run_again"))
goto out;
ASSERT_EQ(skel->bss->arm_err, 0, "rearm_err");
@@ -254,7 +256,9 @@ static void test_call_rcu_two_heads(void)
void test_call_rcu(void)
{
if (test__start_subtest("run"))
- test_call_rcu_run();
+ test_call_rcu_run(false);
+ if (test__start_subtest("run_tasks_trace"))
+ test_call_rcu_run(true);
if (test__start_subtest("chain"))
test_call_rcu_chain();
if (test__start_subtest("two_heads"))
diff --git a/tools/testing/selftests/bpf/progs/call_rcu.c b/tools/testing/selftests/bpf/progs/call_rcu.c
index 94d92cee88a70..505008647fd6f 100644
--- a/tools/testing/selftests/bpf/progs/call_rcu.c
+++ b/tools/testing/selftests/bpf/progs/call_rcu.c
@@ -86,6 +86,23 @@ int arm_both(void *ctx)
return 0;
}
+SEC("syscall")
+int arm_trace(void *ctx)
+{
+ __u32 key = 1;
+ struct elem *e;
+
+ e = bpf_map_lookup_elem(&arr, &key);
+ if (!e)
+ return 1;
+
+ e->val = 0xdeadbeef;
+ /* No lock needed: the enclosing rcu_read_lock_trace() already blocks the grace period. */
+ arm_err = bpf_call_rcu_tasks_trace(&e->rh, &arr, reclaim);
+ busy_err = bpf_call_rcu_tasks_trace(&e->rh, &arr, reclaim);
+ return 0;
+}
+
SEC("iter/bpf_map_elem")
int dump(struct bpf_iter__bpf_map_elem *ctx)
{
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace()
2026-09-22 20:00 [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace() Puranjay Mohan
` (3 preceding siblings ...)
2026-09-22 20:00 ` [PATCH bpf-next v6 4/4] selftests/bpf: Add a test for bpf_call_rcu_tasks_trace() Puranjay Mohan
@ 2026-09-22 23:41 ` Andrii Nakryiko
2026-09-23 10:37 ` Puranjay Mohan
2026-09-22 23:50 ` patchwork-bot+netdevbpf
5 siblings, 1 reply; 17+ messages in thread
From: Andrii Nakryiko @ 2026-09-22 23:41 UTC (permalink / raw)
To: Puranjay Mohan
Cc: bpf, rcu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Song Liu, Yonghong Song, Harry Yoo (Oracle), Paul E. McKenney
On Tue, Sep 22, 2026 at 1:02 PM Puranjay Mohan <puranjay@kernel.org> wrote:
>
> Changelog:
> v5: https://lore.kernel.org/all/20260921191407.1742386-1-puranjay@kernel.org/
> Changes in v6:
> - Size struct bpf_rcu_head at 48 bytes rather than 64, matching what the
> inline state uses (Alexei)
> - Trim patch 1's changelog: drop the reasoning about possible future
> layouts and about what the other async kfuncs return
> - Rebase on bpf-next/master
>
> v4: https://lore.kernel.org/all/20260915154248.3612028-1-puranjay@kernel.org/
> Changes in v5:
> - Rename the "hash_map" subtest to "bad_map" so that it matches its
> helper test_call_rcu_bad_map() (bpf-ci)
> - Bump the callback counter after chain_err in the selftest callback.
> Userspace polls that counter and then reads chain_err, so it could
> still see the initial value before the re-arm had stored one, which
> let the chain subtest's assertion pass without checking anything
> - Say in patch 1 why embedding struct rcu_head in a uapi struct is
> acceptable here (Mykyta, Paul, Alexei)
> - Rebase on bpf-next/master
> v3: https://lore.kernel.org/all/20260915143640.36292-1-puranjay@kernel.org/
> Changes in v4:
> - Drop an unrelated hunk in bpf_async_update_prog_callback() that turned
> PTR_ERR(prog) into -EBADF. That is the shared bpf_timer/bpf_wq path
> and bpf_prog_inc_not_zero() returns -ENOENT, so it would have changed
> the errno bpf_timer_set_callback() and bpf_wq_set_callback() report to
> userspace (Sashiko). No other changes from v3
> v2: https://lore.kernel.org/all/20260915114240.3269184-1-puranjay@kernel.org/
> Changes in v3:
> - Move the bpf_call_rcu_tasks_trace() verifier bits from patch 1 to
> patch 3; patch 1 alone emitted "resolve_btfids: unresolved symbol
> bpf_call_rcu_tasks_trace" (Sashiko)
> - Poll the callback counter with an acquire load (Sashiko)
> - teardown: v2 only checked that the program was eventually freed, which
> passes even if nothing was ever armed. Also assert that the chain ran,
> and read the -EPERM back through an independent .bss fd
> - Use kern_sync_rcu() instead of open coding the grace-period wait
> - Return -EBADF rather than -ENOENT when the calling program is going
> away, matching bpf_task_work_schedule()
> - mismatch_map now pins the bpf_rcu_head label the new code emits; it
> passed with that branch removed. Add a two_heads test, and wait for
> each grace period separately in the chain test
> - Commit messages: correct the -EPERM parity claim, explain the inline
> callback state and the struct size, motivate the tasks trace flavour
> v1: https://lore.kernel.org/all/20260907134552.1772405-1-puranjay@kernel.org/
> Changes in v2:
> - Rebase on bpf-next/master
> - Use rcu_read_lock_dont_migrate() over open coding (Alexei)
> - Improve re-arming selftest to detect failure (Sashiko)
>
> BPF programs that manage their own objects have no way to run their own
> logic once an RCU grace period has elapsed. bpf_obj_drop() defers a
> free, but returning an index to an allocator or unpinning a resource
> once readers are done has no equivalent. sched_ext's BPF library works
> around this today by pushing freed nodes onto a list and having a
> userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
> BPF program to reclaim them; it is the first intended user.
>
> Add:
>
> int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
> int (*callback)(struct bpf_map *map, void *key,
> void *value));
>
> and bpf_call_rcu_tasks_trace(), same signature, which also waits for
> sleepable programs.
>
> @rh is a struct bpf_rcu_head embedded in a value of @map, so the
> callback runs as callback(map, key, value) for the element it lives in
> and needs no cookie. The field is only accepted in BPF_MAP_TYPE_ARRAY,
Isn't restricting this to BPF_MAP_TYPE_ARRAY unnecessarily restrictive
in practice? Do other APIs of similar kind have this limitation? E.g.,
bpf_task_work_schedule() or timer, I don't think they limit user to
just ARRAY maps, it's way too inflexible.
> and arming holds a reference on the calling program until the callback
> has run. Patch 1 covers the lifetime rules.
>
> This needs
>
> https://lore.kernel.org/all/20260810122758.183765-1-puranjay@kernel.org/
>
> for call_rcu() and call_srcu() to be safe from the contexts a BPF
> program can be called in.
>
> Puranjay Mohan (4):
> bpf: Add bpf_call_rcu() kfunc
> selftests/bpf: Add tests for bpf_call_rcu()
> bpf: Add bpf_call_rcu_tasks_trace() kfunc
> selftests/bpf: Add a test for bpf_call_rcu_tasks_trace()
>
> include/linux/bpf.h | 10 +
> include/uapi/linux/bpf.h | 4 +
> kernel/bpf/btf.c | 7 +
> kernel/bpf/helpers.c | 102 +++++++
> kernel/bpf/map_in_map.c | 4 +
> kernel/bpf/map_iter.c | 6 +
> kernel/bpf/syscall.c | 11 +-
> kernel/bpf/verifier.c | 84 +++++-
> tools/include/uapi/linux/bpf.h | 4 +
> .../selftests/bpf/prog_tests/call_rcu.c | 275 ++++++++++++++++++
> tools/testing/selftests/bpf/progs/call_rcu.c | 110 +++++++
> .../selftests/bpf/progs/call_rcu_fail.c | 114 ++++++++
> 12 files changed, 728 insertions(+), 3 deletions(-)
> create mode 100644 tools/testing/selftests/bpf/prog_tests/call_rcu.c
> create mode 100644 tools/testing/selftests/bpf/progs/call_rcu.c
> create mode 100644 tools/testing/selftests/bpf/progs/call_rcu_fail.c
>
>
> base-commit: 79dc258c9392051420a26f1504c647bd3d27c66a
> --
> 2.53.0-Meta
>
^ permalink raw reply [flat|nested] 17+ messages in thread* Re: [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace()
2026-09-22 23:41 ` [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace() Andrii Nakryiko
@ 2026-09-23 10:37 ` Puranjay Mohan
2026-09-23 11:29 ` Kumar Kartikeya Dwivedi
0 siblings, 1 reply; 17+ messages in thread
From: Puranjay Mohan @ 2026-09-23 10:37 UTC (permalink / raw)
To: Andrii Nakryiko
Cc: bpf, rcu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Song Liu, Yonghong Song, Harry Yoo (Oracle), Paul E. McKenney
On Wed, Sep 23, 2026 at 12:41 AM Andrii Nakryiko
<andrii.nakryiko@gmail.com> wrote:
>
> On Tue, Sep 22, 2026 at 1:02 PM Puranjay Mohan <puranjay@kernel.org> wrote:
> >
> > Changelog:
> > v5: https://lore.kernel.org/all/20260921191407.1742386-1-puranjay@kernel.org/
> > Changes in v6:
> > - Size struct bpf_rcu_head at 48 bytes rather than 64, matching what the
> > inline state uses (Alexei)
> > - Trim patch 1's changelog: drop the reasoning about possible future
> > layouts and about what the other async kfuncs return
> > - Rebase on bpf-next/master
> >
> > v4: https://lore.kernel.org/all/20260915154248.3612028-1-puranjay@kernel.org/
> > Changes in v5:
> > - Rename the "hash_map" subtest to "bad_map" so that it matches its
> > helper test_call_rcu_bad_map() (bpf-ci)
> > - Bump the callback counter after chain_err in the selftest callback.
> > Userspace polls that counter and then reads chain_err, so it could
> > still see the initial value before the re-arm had stored one, which
> > let the chain subtest's assertion pass without checking anything
> > - Say in patch 1 why embedding struct rcu_head in a uapi struct is
> > acceptable here (Mykyta, Paul, Alexei)
> > - Rebase on bpf-next/master
> > v3: https://lore.kernel.org/all/20260915143640.36292-1-puranjay@kernel.org/
> > Changes in v4:
> > - Drop an unrelated hunk in bpf_async_update_prog_callback() that turned
> > PTR_ERR(prog) into -EBADF. That is the shared bpf_timer/bpf_wq path
> > and bpf_prog_inc_not_zero() returns -ENOENT, so it would have changed
> > the errno bpf_timer_set_callback() and bpf_wq_set_callback() report to
> > userspace (Sashiko). No other changes from v3
> > v2: https://lore.kernel.org/all/20260915114240.3269184-1-puranjay@kernel.org/
> > Changes in v3:
> > - Move the bpf_call_rcu_tasks_trace() verifier bits from patch 1 to
> > patch 3; patch 1 alone emitted "resolve_btfids: unresolved symbol
> > bpf_call_rcu_tasks_trace" (Sashiko)
> > - Poll the callback counter with an acquire load (Sashiko)
> > - teardown: v2 only checked that the program was eventually freed, which
> > passes even if nothing was ever armed. Also assert that the chain ran,
> > and read the -EPERM back through an independent .bss fd
> > - Use kern_sync_rcu() instead of open coding the grace-period wait
> > - Return -EBADF rather than -ENOENT when the calling program is going
> > away, matching bpf_task_work_schedule()
> > - mismatch_map now pins the bpf_rcu_head label the new code emits; it
> > passed with that branch removed. Add a two_heads test, and wait for
> > each grace period separately in the chain test
> > - Commit messages: correct the -EPERM parity claim, explain the inline
> > callback state and the struct size, motivate the tasks trace flavour
> > v1: https://lore.kernel.org/all/20260907134552.1772405-1-puranjay@kernel.org/
> > Changes in v2:
> > - Rebase on bpf-next/master
> > - Use rcu_read_lock_dont_migrate() over open coding (Alexei)
> > - Improve re-arming selftest to detect failure (Sashiko)
> >
> > BPF programs that manage their own objects have no way to run their own
> > logic once an RCU grace period has elapsed. bpf_obj_drop() defers a
> > free, but returning an index to an allocator or unpinning a resource
> > once readers are done has no equivalent. sched_ext's BPF library works
> > around this today by pushing freed nodes onto a list and having a
> > userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
> > BPF program to reclaim them; it is the first intended user.
> >
> > Add:
> >
> > int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
> > int (*callback)(struct bpf_map *map, void *key,
> > void *value));
> >
> > and bpf_call_rcu_tasks_trace(), same signature, which also waits for
> > sleepable programs.
> >
> > @rh is a struct bpf_rcu_head embedded in a value of @map, so the
> > callback runs as callback(map, key, value) for the element it lives in
> > and needs no cookie. The field is only accepted in BPF_MAP_TYPE_ARRAY,
>
> Isn't restricting this to BPF_MAP_TYPE_ARRAY unnecessarily restrictive
> in practice? Do other APIs of similar kind have this limitation? E.g.,
> bpf_task_work_schedule() or timer, I don't think they limit user to
> just ARRAY maps, it's way too inflexible.
Right, bpf_timer and bpf_task_work take HASH/LRU_HASH/ARRAY. They can
because they can cancel: on element delete bpf_obj_free_fields() calls
bpf_task_work_cancel_and_free(), which reaches task_work_cancel(), so
a pending callback never runs against a recycled element.
A queued RCU callback cannot be cancelled. rcu_barrier() is the only
thing that waits for one, and it sleeps, so it is not callable from an
element delete. Array elements are never freed individually, which is
what makes it safe.
I can think of a complicated way to support hash maps but that would
need making the state dynamically allocated and finding a way to
cancel the callbacks. But I would do it as a follow-up if there is a
real use case.
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace()
2026-09-23 10:37 ` Puranjay Mohan
@ 2026-09-23 11:29 ` Kumar Kartikeya Dwivedi
2026-09-23 16:45 ` Andrii Nakryiko
0 siblings, 1 reply; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-23 11:29 UTC (permalink / raw)
To: Puranjay Mohan, Andrii Nakryiko
Cc: bpf, rcu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Song Liu, Yonghong Song,
Harry Yoo (Oracle), Paul E. McKenney
On Wed Sep 23, 2026 at 12:37 PM CEST, Puranjay Mohan wrote:
> On Wed, Sep 23, 2026 at 12:41 AM Andrii Nakryiko
> <andrii.nakryiko@gmail.com> wrote:
>>
>> On Tue, Sep 22, 2026 at 1:02 PM Puranjay Mohan <puranjay@kernel.org> wrote:
>> >
>> > Changelog:
>> > v5: https://lore.kernel.org/all/20260921191407.1742386-1-puranjay@kernel.org/
>> > Changes in v6:
>> > - Size struct bpf_rcu_head at 48 bytes rather than 64, matching what the
>> > inline state uses (Alexei)
>> > - Trim patch 1's changelog: drop the reasoning about possible future
>> > layouts and about what the other async kfuncs return
>> > - Rebase on bpf-next/master
>> >
>> > v4: https://lore.kernel.org/all/20260915154248.3612028-1-puranjay@kernel.org/
>> > Changes in v5:
>> > - Rename the "hash_map" subtest to "bad_map" so that it matches its
>> > helper test_call_rcu_bad_map() (bpf-ci)
>> > - Bump the callback counter after chain_err in the selftest callback.
>> > Userspace polls that counter and then reads chain_err, so it could
>> > still see the initial value before the re-arm had stored one, which
>> > let the chain subtest's assertion pass without checking anything
>> > - Say in patch 1 why embedding struct rcu_head in a uapi struct is
>> > acceptable here (Mykyta, Paul, Alexei)
>> > - Rebase on bpf-next/master
>> > v3: https://lore.kernel.org/all/20260915143640.36292-1-puranjay@kernel.org/
>> > Changes in v4:
>> > - Drop an unrelated hunk in bpf_async_update_prog_callback() that turned
>> > PTR_ERR(prog) into -EBADF. That is the shared bpf_timer/bpf_wq path
>> > and bpf_prog_inc_not_zero() returns -ENOENT, so it would have changed
>> > the errno bpf_timer_set_callback() and bpf_wq_set_callback() report to
>> > userspace (Sashiko). No other changes from v3
>> > v2: https://lore.kernel.org/all/20260915114240.3269184-1-puranjay@kernel.org/
>> > Changes in v3:
>> > - Move the bpf_call_rcu_tasks_trace() verifier bits from patch 1 to
>> > patch 3; patch 1 alone emitted "resolve_btfids: unresolved symbol
>> > bpf_call_rcu_tasks_trace" (Sashiko)
>> > - Poll the callback counter with an acquire load (Sashiko)
>> > - teardown: v2 only checked that the program was eventually freed, which
>> > passes even if nothing was ever armed. Also assert that the chain ran,
>> > and read the -EPERM back through an independent .bss fd
>> > - Use kern_sync_rcu() instead of open coding the grace-period wait
>> > - Return -EBADF rather than -ENOENT when the calling program is going
>> > away, matching bpf_task_work_schedule()
>> > - mismatch_map now pins the bpf_rcu_head label the new code emits; it
>> > passed with that branch removed. Add a two_heads test, and wait for
>> > each grace period separately in the chain test
>> > - Commit messages: correct the -EPERM parity claim, explain the inline
>> > callback state and the struct size, motivate the tasks trace flavour
>> > v1: https://lore.kernel.org/all/20260907134552.1772405-1-puranjay@kernel.org/
>> > Changes in v2:
>> > - Rebase on bpf-next/master
>> > - Use rcu_read_lock_dont_migrate() over open coding (Alexei)
>> > - Improve re-arming selftest to detect failure (Sashiko)
>> >
>> > BPF programs that manage their own objects have no way to run their own
>> > logic once an RCU grace period has elapsed. bpf_obj_drop() defers a
>> > free, but returning an index to an allocator or unpinning a resource
>> > once readers are done has no equivalent. sched_ext's BPF library works
>> > around this today by pushing freed nodes onto a list and having a
>> > userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
>> > BPF program to reclaim them; it is the first intended user.
>> >
>> > Add:
>> >
>> > int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
>> > int (*callback)(struct bpf_map *map, void *key,
>> > void *value));
>> >
>> > and bpf_call_rcu_tasks_trace(), same signature, which also waits for
>> > sleepable programs.
>> >
>> > @rh is a struct bpf_rcu_head embedded in a value of @map, so the
>> > callback runs as callback(map, key, value) for the element it lives in
>> > and needs no cookie. The field is only accepted in BPF_MAP_TYPE_ARRAY,
>>
>> Isn't restricting this to BPF_MAP_TYPE_ARRAY unnecessarily restrictive
>> in practice? Do other APIs of similar kind have this limitation? E.g.,
>> bpf_task_work_schedule() or timer, I don't think they limit user to
>> just ARRAY maps, it's way too inflexible.
>
> Right, bpf_timer and bpf_task_work take HASH/LRU_HASH/ARRAY. They can
> because they can cancel: on element delete bpf_obj_free_fields() calls
bpf_obj_cancel_fields(). bpf_obj_free_fields() is no longer called except when
the object is being freed finally.
> bpf_task_work_cancel_and_free(), which reaches task_work_cancel(), so
> a pending callback never runs against a recycled element.
> A queued RCU callback cannot be cancelled. rcu_barrier() is the only
> thing that waits for one, and it sleeps, so it is not callable from an
> element delete. Array elements are never freed individually, which is
> what makes it safe.
If cancel is a noop, can you not simply skip it on map update?
It is a bit unfortunate we can't provide cancel semantics, it will be yet
another divergence from how all other async callbacks work...sigh.
>
> I can think of a complicated way to support hash maps but that would
> need making the state dynamically allocated and finding a way to
> cancel the callbacks. But I would do it as a follow-up if there is a
> real use case.
That would mean allocation and frees, it is already quite expensive as is for a
call_rcu() primitive (with the refcount bumps and atomics).
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace()
2026-09-23 11:29 ` Kumar Kartikeya Dwivedi
@ 2026-09-23 16:45 ` Andrii Nakryiko
2026-09-23 16:53 ` Kumar Kartikeya Dwivedi
0 siblings, 1 reply; 17+ messages in thread
From: Andrii Nakryiko @ 2026-09-23 16:45 UTC (permalink / raw)
To: Kumar Kartikeya Dwivedi
Cc: Puranjay Mohan, bpf, rcu, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman, Song Liu,
Yonghong Song, Harry Yoo (Oracle), Paul E. McKenney
On Wed, Sep 23, 2026 at 4:29 AM Kumar Kartikeya Dwivedi
<memxor@gmail.com> wrote:
>
> On Wed Sep 23, 2026 at 12:37 PM CEST, Puranjay Mohan wrote:
> > On Wed, Sep 23, 2026 at 12:41 AM Andrii Nakryiko
> > <andrii.nakryiko@gmail.com> wrote:
> >>
> >> On Tue, Sep 22, 2026 at 1:02 PM Puranjay Mohan <puranjay@kernel.org> wrote:
> >> >
> >> > Changelog:
> >> > v5: https://lore.kernel.org/all/20260921191407.1742386-1-puranjay@kernel.org/
> >> > Changes in v6:
> >> > - Size struct bpf_rcu_head at 48 bytes rather than 64, matching what the
> >> > inline state uses (Alexei)
> >> > - Trim patch 1's changelog: drop the reasoning about possible future
> >> > layouts and about what the other async kfuncs return
> >> > - Rebase on bpf-next/master
> >> >
> >> > v4: https://lore.kernel.org/all/20260915154248.3612028-1-puranjay@kernel.org/
> >> > Changes in v5:
> >> > - Rename the "hash_map" subtest to "bad_map" so that it matches its
> >> > helper test_call_rcu_bad_map() (bpf-ci)
> >> > - Bump the callback counter after chain_err in the selftest callback.
> >> > Userspace polls that counter and then reads chain_err, so it could
> >> > still see the initial value before the re-arm had stored one, which
> >> > let the chain subtest's assertion pass without checking anything
> >> > - Say in patch 1 why embedding struct rcu_head in a uapi struct is
> >> > acceptable here (Mykyta, Paul, Alexei)
> >> > - Rebase on bpf-next/master
> >> > v3: https://lore.kernel.org/all/20260915143640.36292-1-puranjay@kernel.org/
> >> > Changes in v4:
> >> > - Drop an unrelated hunk in bpf_async_update_prog_callback() that turned
> >> > PTR_ERR(prog) into -EBADF. That is the shared bpf_timer/bpf_wq path
> >> > and bpf_prog_inc_not_zero() returns -ENOENT, so it would have changed
> >> > the errno bpf_timer_set_callback() and bpf_wq_set_callback() report to
> >> > userspace (Sashiko). No other changes from v3
> >> > v2: https://lore.kernel.org/all/20260915114240.3269184-1-puranjay@kernel.org/
> >> > Changes in v3:
> >> > - Move the bpf_call_rcu_tasks_trace() verifier bits from patch 1 to
> >> > patch 3; patch 1 alone emitted "resolve_btfids: unresolved symbol
> >> > bpf_call_rcu_tasks_trace" (Sashiko)
> >> > - Poll the callback counter with an acquire load (Sashiko)
> >> > - teardown: v2 only checked that the program was eventually freed, which
> >> > passes even if nothing was ever armed. Also assert that the chain ran,
> >> > and read the -EPERM back through an independent .bss fd
> >> > - Use kern_sync_rcu() instead of open coding the grace-period wait
> >> > - Return -EBADF rather than -ENOENT when the calling program is going
> >> > away, matching bpf_task_work_schedule()
> >> > - mismatch_map now pins the bpf_rcu_head label the new code emits; it
> >> > passed with that branch removed. Add a two_heads test, and wait for
> >> > each grace period separately in the chain test
> >> > - Commit messages: correct the -EPERM parity claim, explain the inline
> >> > callback state and the struct size, motivate the tasks trace flavour
> >> > v1: https://lore.kernel.org/all/20260907134552.1772405-1-puranjay@kernel.org/
> >> > Changes in v2:
> >> > - Rebase on bpf-next/master
> >> > - Use rcu_read_lock_dont_migrate() over open coding (Alexei)
> >> > - Improve re-arming selftest to detect failure (Sashiko)
> >> >
> >> > BPF programs that manage their own objects have no way to run their own
> >> > logic once an RCU grace period has elapsed. bpf_obj_drop() defers a
> >> > free, but returning an index to an allocator or unpinning a resource
> >> > once readers are done has no equivalent. sched_ext's BPF library works
> >> > around this today by pushing freed nodes onto a list and having a
> >> > userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
> >> > BPF program to reclaim them; it is the first intended user.
> >> >
> >> > Add:
> >> >
> >> > int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
> >> > int (*callback)(struct bpf_map *map, void *key,
> >> > void *value));
> >> >
> >> > and bpf_call_rcu_tasks_trace(), same signature, which also waits for
> >> > sleepable programs.
> >> >
> >> > @rh is a struct bpf_rcu_head embedded in a value of @map, so the
> >> > callback runs as callback(map, key, value) for the element it lives in
> >> > and needs no cookie. The field is only accepted in BPF_MAP_TYPE_ARRAY,
> >>
> >> Isn't restricting this to BPF_MAP_TYPE_ARRAY unnecessarily restrictive
> >> in practice? Do other APIs of similar kind have this limitation? E.g.,
> >> bpf_task_work_schedule() or timer, I don't think they limit user to
> >> just ARRAY maps, it's way too inflexible.
> >
> > Right, bpf_timer and bpf_task_work take HASH/LRU_HASH/ARRAY. They can
> > because they can cancel: on element delete bpf_obj_free_fields() calls
>
> bpf_obj_cancel_fields(). bpf_obj_free_fields() is no longer called except when
> the object is being freed finally.
>
> > bpf_task_work_cancel_and_free(), which reaches task_work_cancel(), so
> > a pending callback never runs against a recycled element.
> > A queued RCU callback cannot be cancelled. rcu_barrier() is the only
> > thing that waits for one, and it sleeps, so it is not callable from an
> > element delete. Array elements are never freed individually, which is
> > what makes it safe.
>
> If cancel is a noop, can you not simply skip it on map update?
>
> It is a bit unfortunate we can't provide cancel semantics, it will be yet
> another divergence from how all other async callbacks work...sigh.
>
> >
> > I can think of a complicated way to support hash maps but that would
> > need making the state dynamically allocated and finding a way to
> > cancel the callbacks. But I would do it as a follow-up if there is a
> > real use case.
>
> That would mean allocation and frees, it is already quite expensive as is for a
> call_rcu() primitive (with the refcount bumps and atomics).
maybe so, but as is ARRAY is a severe limitation and makes a lot of
cases either unusable or requiring extra sequential ID allocation
logic to remap some HASH map entry to index in an ARRAY map, just so
you can call rcu callback for a given kernel object stored in HASH
map.
E.g., think about keep as set of sockets, or tasks, or whatnot in HASH
or RHASHTABLE (especially the latter one with support for resizing).
How would you do bpf_call_rcu() for them with unduly complications,
limitations and a lot of waste (unlike HASH you can't have lazy memory
allocation).
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace()
2026-09-23 16:45 ` Andrii Nakryiko
@ 2026-09-23 16:53 ` Kumar Kartikeya Dwivedi
2026-09-23 16:59 ` Andrii Nakryiko
0 siblings, 1 reply; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-23 16:53 UTC (permalink / raw)
To: Andrii Nakryiko
Cc: Puranjay Mohan, bpf, rcu, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman, Song Liu,
Yonghong Song, Harry Yoo (Oracle), Paul E. McKenney
On Wed Sep 23, 2026 at 6:45 PM CEST, Andrii Nakryiko wrote:
> On Wed, Sep 23, 2026 at 4:29 AM Kumar Kartikeya Dwivedi
> <memxor@gmail.com> wrote:
>>
>> On Wed Sep 23, 2026 at 12:37 PM CEST, Puranjay Mohan wrote:
>> > On Wed, Sep 23, 2026 at 12:41 AM Andrii Nakryiko
>> > <andrii.nakryiko@gmail.com> wrote:
>> >>
>> >> On Tue, Sep 22, 2026 at 1:02 PM Puranjay Mohan <puranjay@kernel.org> wrote:
>> >> >
>> >> > Changelog:
>> >> > v5: https://lore.kernel.org/all/20260921191407.1742386-1-puranjay@kernel.org/
>> >> > Changes in v6:
>> >> > - Size struct bpf_rcu_head at 48 bytes rather than 64, matching what the
>> >> > inline state uses (Alexei)
>> >> > - Trim patch 1's changelog: drop the reasoning about possible future
>> >> > layouts and about what the other async kfuncs return
>> >> > - Rebase on bpf-next/master
>> >> >
>> >> > v4: https://lore.kernel.org/all/20260915154248.3612028-1-puranjay@kernel.org/
>> >> > Changes in v5:
>> >> > - Rename the "hash_map" subtest to "bad_map" so that it matches its
>> >> > helper test_call_rcu_bad_map() (bpf-ci)
>> >> > - Bump the callback counter after chain_err in the selftest callback.
>> >> > Userspace polls that counter and then reads chain_err, so it could
>> >> > still see the initial value before the re-arm had stored one, which
>> >> > let the chain subtest's assertion pass without checking anything
>> >> > - Say in patch 1 why embedding struct rcu_head in a uapi struct is
>> >> > acceptable here (Mykyta, Paul, Alexei)
>> >> > - Rebase on bpf-next/master
>> >> > v3: https://lore.kernel.org/all/20260915143640.36292-1-puranjay@kernel.org/
>> >> > Changes in v4:
>> >> > - Drop an unrelated hunk in bpf_async_update_prog_callback() that turned
>> >> > PTR_ERR(prog) into -EBADF. That is the shared bpf_timer/bpf_wq path
>> >> > and bpf_prog_inc_not_zero() returns -ENOENT, so it would have changed
>> >> > the errno bpf_timer_set_callback() and bpf_wq_set_callback() report to
>> >> > userspace (Sashiko). No other changes from v3
>> >> > v2: https://lore.kernel.org/all/20260915114240.3269184-1-puranjay@kernel.org/
>> >> > Changes in v3:
>> >> > - Move the bpf_call_rcu_tasks_trace() verifier bits from patch 1 to
>> >> > patch 3; patch 1 alone emitted "resolve_btfids: unresolved symbol
>> >> > bpf_call_rcu_tasks_trace" (Sashiko)
>> >> > - Poll the callback counter with an acquire load (Sashiko)
>> >> > - teardown: v2 only checked that the program was eventually freed, which
>> >> > passes even if nothing was ever armed. Also assert that the chain ran,
>> >> > and read the -EPERM back through an independent .bss fd
>> >> > - Use kern_sync_rcu() instead of open coding the grace-period wait
>> >> > - Return -EBADF rather than -ENOENT when the calling program is going
>> >> > away, matching bpf_task_work_schedule()
>> >> > - mismatch_map now pins the bpf_rcu_head label the new code emits; it
>> >> > passed with that branch removed. Add a two_heads test, and wait for
>> >> > each grace period separately in the chain test
>> >> > - Commit messages: correct the -EPERM parity claim, explain the inline
>> >> > callback state and the struct size, motivate the tasks trace flavour
>> >> > v1: https://lore.kernel.org/all/20260907134552.1772405-1-puranjay@kernel.org/
>> >> > Changes in v2:
>> >> > - Rebase on bpf-next/master
>> >> > - Use rcu_read_lock_dont_migrate() over open coding (Alexei)
>> >> > - Improve re-arming selftest to detect failure (Sashiko)
>> >> >
>> >> > BPF programs that manage their own objects have no way to run their own
>> >> > logic once an RCU grace period has elapsed. bpf_obj_drop() defers a
>> >> > free, but returning an index to an allocator or unpinning a resource
>> >> > once readers are done has no equivalent. sched_ext's BPF library works
>> >> > around this today by pushing freed nodes onto a list and having a
>> >> > userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
>> >> > BPF program to reclaim them; it is the first intended user.
>> >> >
>> >> > Add:
>> >> >
>> >> > int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
>> >> > int (*callback)(struct bpf_map *map, void *key,
>> >> > void *value));
>> >> >
>> >> > and bpf_call_rcu_tasks_trace(), same signature, which also waits for
>> >> > sleepable programs.
>> >> >
>> >> > @rh is a struct bpf_rcu_head embedded in a value of @map, so the
>> >> > callback runs as callback(map, key, value) for the element it lives in
>> >> > and needs no cookie. The field is only accepted in BPF_MAP_TYPE_ARRAY,
>> >>
>> >> Isn't restricting this to BPF_MAP_TYPE_ARRAY unnecessarily restrictive
>> >> in practice? Do other APIs of similar kind have this limitation? E.g.,
>> >> bpf_task_work_schedule() or timer, I don't think they limit user to
>> >> just ARRAY maps, it's way too inflexible.
>> >
>> > Right, bpf_timer and bpf_task_work take HASH/LRU_HASH/ARRAY. They can
>> > because they can cancel: on element delete bpf_obj_free_fields() calls
>>
>> bpf_obj_cancel_fields(). bpf_obj_free_fields() is no longer called except when
>> the object is being freed finally.
>>
>> > bpf_task_work_cancel_and_free(), which reaches task_work_cancel(), so
>> > a pending callback never runs against a recycled element.
>> > A queued RCU callback cannot be cancelled. rcu_barrier() is the only
>> > thing that waits for one, and it sleeps, so it is not callable from an
>> > element delete. Array elements are never freed individually, which is
>> > what makes it safe.
>>
>> If cancel is a noop, can you not simply skip it on map update?
>>
>> It is a bit unfortunate we can't provide cancel semantics, it will be yet
>> another divergence from how all other async callbacks work...sigh.
>>
>> >
>> > I can think of a complicated way to support hash maps but that would
>> > need making the state dynamically allocated and finding a way to
>> > cancel the callbacks. But I would do it as a follow-up if there is a
>> > real use case.
>>
>> That would mean allocation and frees, it is already quite expensive as is for a
>> call_rcu() primitive (with the refcount bumps and atomics).
>
> maybe so, but as is ARRAY is a severe limitation and makes a lot of
> cases either unusable or requiring extra sequential ID allocation
> logic to remap some HASH map entry to index in an ARRAY map, just so
> you can call rcu callback for a given kernel object stored in HASH
> map.
>
> E.g., think about keep as set of sockets, or tasks, or whatnot in HASH
> or RHASHTABLE (especially the latter one with support for resizing).
> How would you do bpf_call_rcu() for them with unduly complications,
> limitations and a lot of waste (unlike HASH you can't have lazy memory
> allocation).
Yeah, I don't disagree. We would also need to support storing them in more map
types to make it woth with arenas (which was one of the motivations for the
feature).
Perhaps it is better to make cancel a noop / unsupported and skip it when a map
element is deleted. That will be the easiest path to enabling it, but it does
create (perhaps surprising) divergence from some of the other async callback
primitives.
That said, I do think performance should be a consideration for this API; maybe
unlike other ones, I would expect that this could be called very frequently. I
had concerns about prog refcount increment as well, but didn't really bring it
up since it is something that can be addressed after the fact (maybe using pcpu
refs). But once we promise supporting cancellation, it is hard to walk that
back. I think it's less of a concern for other async cb types. In practice,
except to support map semantics I don't know if anyone will use cancellation API
either, the kernel side never grew such support.
Anyway, overall the best path to me seems to be just enabling it as is and
declaring bpf_obj_cancel_fields() on this noop. Existing synchronization should
be enough to ensure callback stops getting issued once element is finally freed
back to kernel allocator.
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace()
2026-09-23 16:53 ` Kumar Kartikeya Dwivedi
@ 2026-09-23 16:59 ` Andrii Nakryiko
2026-09-23 17:57 ` Kumar Kartikeya Dwivedi
0 siblings, 1 reply; 17+ messages in thread
From: Andrii Nakryiko @ 2026-09-23 16:59 UTC (permalink / raw)
To: Kumar Kartikeya Dwivedi
Cc: Puranjay Mohan, bpf, rcu, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman, Song Liu,
Yonghong Song, Harry Yoo (Oracle), Paul E. McKenney
On Wed, Sep 23, 2026 at 9:53 AM Kumar Kartikeya Dwivedi
<memxor@gmail.com> wrote:
>
> On Wed Sep 23, 2026 at 6:45 PM CEST, Andrii Nakryiko wrote:
> > On Wed, Sep 23, 2026 at 4:29 AM Kumar Kartikeya Dwivedi
> > <memxor@gmail.com> wrote:
> >>
> >> On Wed Sep 23, 2026 at 12:37 PM CEST, Puranjay Mohan wrote:
> >> > On Wed, Sep 23, 2026 at 12:41 AM Andrii Nakryiko
> >> > <andrii.nakryiko@gmail.com> wrote:
> >> >>
> >> >> On Tue, Sep 22, 2026 at 1:02 PM Puranjay Mohan <puranjay@kernel.org> wrote:
> >> >> >
> >> >> > Changelog:
> >> >> > v5: https://lore.kernel.org/all/20260921191407.1742386-1-puranjay@kernel.org/
> >> >> > Changes in v6:
> >> >> > - Size struct bpf_rcu_head at 48 bytes rather than 64, matching what the
> >> >> > inline state uses (Alexei)
> >> >> > - Trim patch 1's changelog: drop the reasoning about possible future
> >> >> > layouts and about what the other async kfuncs return
> >> >> > - Rebase on bpf-next/master
> >> >> >
> >> >> > v4: https://lore.kernel.org/all/20260915154248.3612028-1-puranjay@kernel.org/
> >> >> > Changes in v5:
> >> >> > - Rename the "hash_map" subtest to "bad_map" so that it matches its
> >> >> > helper test_call_rcu_bad_map() (bpf-ci)
> >> >> > - Bump the callback counter after chain_err in the selftest callback.
> >> >> > Userspace polls that counter and then reads chain_err, so it could
> >> >> > still see the initial value before the re-arm had stored one, which
> >> >> > let the chain subtest's assertion pass without checking anything
> >> >> > - Say in patch 1 why embedding struct rcu_head in a uapi struct is
> >> >> > acceptable here (Mykyta, Paul, Alexei)
> >> >> > - Rebase on bpf-next/master
> >> >> > v3: https://lore.kernel.org/all/20260915143640.36292-1-puranjay@kernel.org/
> >> >> > Changes in v4:
> >> >> > - Drop an unrelated hunk in bpf_async_update_prog_callback() that turned
> >> >> > PTR_ERR(prog) into -EBADF. That is the shared bpf_timer/bpf_wq path
> >> >> > and bpf_prog_inc_not_zero() returns -ENOENT, so it would have changed
> >> >> > the errno bpf_timer_set_callback() and bpf_wq_set_callback() report to
> >> >> > userspace (Sashiko). No other changes from v3
> >> >> > v2: https://lore.kernel.org/all/20260915114240.3269184-1-puranjay@kernel.org/
> >> >> > Changes in v3:
> >> >> > - Move the bpf_call_rcu_tasks_trace() verifier bits from patch 1 to
> >> >> > patch 3; patch 1 alone emitted "resolve_btfids: unresolved symbol
> >> >> > bpf_call_rcu_tasks_trace" (Sashiko)
> >> >> > - Poll the callback counter with an acquire load (Sashiko)
> >> >> > - teardown: v2 only checked that the program was eventually freed, which
> >> >> > passes even if nothing was ever armed. Also assert that the chain ran,
> >> >> > and read the -EPERM back through an independent .bss fd
> >> >> > - Use kern_sync_rcu() instead of open coding the grace-period wait
> >> >> > - Return -EBADF rather than -ENOENT when the calling program is going
> >> >> > away, matching bpf_task_work_schedule()
> >> >> > - mismatch_map now pins the bpf_rcu_head label the new code emits; it
> >> >> > passed with that branch removed. Add a two_heads test, and wait for
> >> >> > each grace period separately in the chain test
> >> >> > - Commit messages: correct the -EPERM parity claim, explain the inline
> >> >> > callback state and the struct size, motivate the tasks trace flavour
> >> >> > v1: https://lore.kernel.org/all/20260907134552.1772405-1-puranjay@kernel.org/
> >> >> > Changes in v2:
> >> >> > - Rebase on bpf-next/master
> >> >> > - Use rcu_read_lock_dont_migrate() over open coding (Alexei)
> >> >> > - Improve re-arming selftest to detect failure (Sashiko)
> >> >> >
> >> >> > BPF programs that manage their own objects have no way to run their own
> >> >> > logic once an RCU grace period has elapsed. bpf_obj_drop() defers a
> >> >> > free, but returning an index to an allocator or unpinning a resource
> >> >> > once readers are done has no equivalent. sched_ext's BPF library works
> >> >> > around this today by pushing freed nodes onto a list and having a
> >> >> > userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
> >> >> > BPF program to reclaim them; it is the first intended user.
> >> >> >
> >> >> > Add:
> >> >> >
> >> >> > int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
> >> >> > int (*callback)(struct bpf_map *map, void *key,
> >> >> > void *value));
> >> >> >
> >> >> > and bpf_call_rcu_tasks_trace(), same signature, which also waits for
> >> >> > sleepable programs.
> >> >> >
> >> >> > @rh is a struct bpf_rcu_head embedded in a value of @map, so the
> >> >> > callback runs as callback(map, key, value) for the element it lives in
> >> >> > and needs no cookie. The field is only accepted in BPF_MAP_TYPE_ARRAY,
> >> >>
> >> >> Isn't restricting this to BPF_MAP_TYPE_ARRAY unnecessarily restrictive
> >> >> in practice? Do other APIs of similar kind have this limitation? E.g.,
> >> >> bpf_task_work_schedule() or timer, I don't think they limit user to
> >> >> just ARRAY maps, it's way too inflexible.
> >> >
> >> > Right, bpf_timer and bpf_task_work take HASH/LRU_HASH/ARRAY. They can
> >> > because they can cancel: on element delete bpf_obj_free_fields() calls
> >>
> >> bpf_obj_cancel_fields(). bpf_obj_free_fields() is no longer called except when
> >> the object is being freed finally.
> >>
> >> > bpf_task_work_cancel_and_free(), which reaches task_work_cancel(), so
> >> > a pending callback never runs against a recycled element.
> >> > A queued RCU callback cannot be cancelled. rcu_barrier() is the only
> >> > thing that waits for one, and it sleeps, so it is not callable from an
> >> > element delete. Array elements are never freed individually, which is
> >> > what makes it safe.
> >>
> >> If cancel is a noop, can you not simply skip it on map update?
> >>
> >> It is a bit unfortunate we can't provide cancel semantics, it will be yet
> >> another divergence from how all other async callbacks work...sigh.
> >>
> >> >
> >> > I can think of a complicated way to support hash maps but that would
> >> > need making the state dynamically allocated and finding a way to
> >> > cancel the callbacks. But I would do it as a follow-up if there is a
> >> > real use case.
> >>
> >> That would mean allocation and frees, it is already quite expensive as is for a
> >> call_rcu() primitive (with the refcount bumps and atomics).
> >
> > maybe so, but as is ARRAY is a severe limitation and makes a lot of
> > cases either unusable or requiring extra sequential ID allocation
> > logic to remap some HASH map entry to index in an ARRAY map, just so
> > you can call rcu callback for a given kernel object stored in HASH
> > map.
> >
> > E.g., think about keep as set of sockets, or tasks, or whatnot in HASH
> > or RHASHTABLE (especially the latter one with support for resizing).
> > How would you do bpf_call_rcu() for them with unduly complications,
> > limitations and a lot of waste (unlike HASH you can't have lazy memory
> > allocation).
>
> Yeah, I don't disagree. We would also need to support storing them in more map
> types to make it woth with arenas (which was one of the motivations for the
> feature).
>
> Perhaps it is better to make cancel a noop / unsupported and skip it when a map
> element is deleted. That will be the easiest path to enabling it, but it does
> create (perhaps surprising) divergence from some of the other async callback
> primitives.
I don't see much need for cancellation, but a more useful/sane
behavior would be rearming on subsequent calls to bpf_call_rcu(). This
would also handle update/delete/reuse of entries. Not sure how hard it
is to support that in kernel's call_rcu() implementation, though, but
that would solve the problem, because it should always be OK to delay
call_rcu() callback, but not the other way around (which is what would
happen today because we will ignore subsequent bpf_call_rcu() calls).
>
> That said, I do think performance should be a consideration for this API; maybe
> unlike other ones, I would expect that this could be called very frequently. I
I'm not sure why this has to be super high frequency API to use, tbh.
You'd normally use this when cleaning up when some kernel object is
freed, no? Sure that can be relatively frequent, but not to the point
where we should be *that* concerned with refcount or atomics overhead
per se (multi-cpu cache bouncing of refcounting is a concern, but not
sure what you can do about that).
> had concerns about prog refcount increment as well, but didn't really bring it
> up since it is something that can be addressed after the fact (maybe using pcpu
> refs). But once we promise supporting cancellation, it is hard to walk that
> back. I think it's less of a concern for other async cb types. In practice,
> except to support map semantics I don't know if anyone will use cancellation API
> either, the kernel side never grew such support.
>
> Anyway, overall the best path to me seems to be just enabling it as is and
you mean enabling for HASH or keeping it for ARRAY only? If the
latter, my concern is that to support HASH we might need to change the
internal structure and break that 48-byte size, so we should probably
decide all this before we get this into the next Linux release.
> declaring bpf_obj_cancel_fields() on this noop. Existing synchronization should
> be enough to ensure callback stops getting issued once element is finally freed
> back to kernel allocator.
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace()
2026-09-23 16:59 ` Andrii Nakryiko
@ 2026-09-23 17:57 ` Kumar Kartikeya Dwivedi
2026-09-23 18:21 ` Andrii Nakryiko
0 siblings, 1 reply; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-23 17:57 UTC (permalink / raw)
To: Andrii Nakryiko
Cc: Puranjay Mohan, bpf, rcu, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman, Song Liu,
Yonghong Song, Harry Yoo (Oracle), Paul E. McKenney
On Wed Sep 23, 2026 at 6:59 PM CEST, Andrii Nakryiko wrote:
> On Wed, Sep 23, 2026 at 9:53 AM Kumar Kartikeya Dwivedi
> <memxor@gmail.com> wrote:
>>
>> On Wed Sep 23, 2026 at 6:45 PM CEST, Andrii Nakryiko wrote:
>> > On Wed, Sep 23, 2026 at 4:29 AM Kumar Kartikeya Dwivedi
>> > <memxor@gmail.com> wrote:
>> >>
>> >> On Wed Sep 23, 2026 at 12:37 PM CEST, Puranjay Mohan wrote:
>> >> > On Wed, Sep 23, 2026 at 12:41 AM Andrii Nakryiko
>> >> > <andrii.nakryiko@gmail.com> wrote:
>> >> >>
>> >> >> On Tue, Sep 22, 2026 at 1:02 PM Puranjay Mohan <puranjay@kernel.org> wrote:
>> >> >> >
>> >> >> > Changelog:
>> >> >> > v5: https://lore.kernel.org/all/20260921191407.1742386-1-puranjay@kernel.org/
>> >> >> > Changes in v6:
>> >> >> > - Size struct bpf_rcu_head at 48 bytes rather than 64, matching what the
>> >> >> > inline state uses (Alexei)
>> >> >> > - Trim patch 1's changelog: drop the reasoning about possible future
>> >> >> > layouts and about what the other async kfuncs return
>> >> >> > - Rebase on bpf-next/master
>> >> >> >
>> >> >> > v4: https://lore.kernel.org/all/20260915154248.3612028-1-puranjay@kernel.org/
>> >> >> > Changes in v5:
>> >> >> > - Rename the "hash_map" subtest to "bad_map" so that it matches its
>> >> >> > helper test_call_rcu_bad_map() (bpf-ci)
>> >> >> > - Bump the callback counter after chain_err in the selftest callback.
>> >> >> > Userspace polls that counter and then reads chain_err, so it could
>> >> >> > still see the initial value before the re-arm had stored one, which
>> >> >> > let the chain subtest's assertion pass without checking anything
>> >> >> > - Say in patch 1 why embedding struct rcu_head in a uapi struct is
>> >> >> > acceptable here (Mykyta, Paul, Alexei)
>> >> >> > - Rebase on bpf-next/master
>> >> >> > v3: https://lore.kernel.org/all/20260915143640.36292-1-puranjay@kernel.org/
>> >> >> > Changes in v4:
>> >> >> > - Drop an unrelated hunk in bpf_async_update_prog_callback() that turned
>> >> >> > PTR_ERR(prog) into -EBADF. That is the shared bpf_timer/bpf_wq path
>> >> >> > and bpf_prog_inc_not_zero() returns -ENOENT, so it would have changed
>> >> >> > the errno bpf_timer_set_callback() and bpf_wq_set_callback() report to
>> >> >> > userspace (Sashiko). No other changes from v3
>> >> >> > v2: https://lore.kernel.org/all/20260915114240.3269184-1-puranjay@kernel.org/
>> >> >> > Changes in v3:
>> >> >> > - Move the bpf_call_rcu_tasks_trace() verifier bits from patch 1 to
>> >> >> > patch 3; patch 1 alone emitted "resolve_btfids: unresolved symbol
>> >> >> > bpf_call_rcu_tasks_trace" (Sashiko)
>> >> >> > - Poll the callback counter with an acquire load (Sashiko)
>> >> >> > - teardown: v2 only checked that the program was eventually freed, which
>> >> >> > passes even if nothing was ever armed. Also assert that the chain ran,
>> >> >> > and read the -EPERM back through an independent .bss fd
>> >> >> > - Use kern_sync_rcu() instead of open coding the grace-period wait
>> >> >> > - Return -EBADF rather than -ENOENT when the calling program is going
>> >> >> > away, matching bpf_task_work_schedule()
>> >> >> > - mismatch_map now pins the bpf_rcu_head label the new code emits; it
>> >> >> > passed with that branch removed. Add a two_heads test, and wait for
>> >> >> > each grace period separately in the chain test
>> >> >> > - Commit messages: correct the -EPERM parity claim, explain the inline
>> >> >> > callback state and the struct size, motivate the tasks trace flavour
>> >> >> > v1: https://lore.kernel.org/all/20260907134552.1772405-1-puranjay@kernel.org/
>> >> >> > Changes in v2:
>> >> >> > - Rebase on bpf-next/master
>> >> >> > - Use rcu_read_lock_dont_migrate() over open coding (Alexei)
>> >> >> > - Improve re-arming selftest to detect failure (Sashiko)
>> >> >> >
>> >> >> > BPF programs that manage their own objects have no way to run their own
>> >> >> > logic once an RCU grace period has elapsed. bpf_obj_drop() defers a
>> >> >> > free, but returning an index to an allocator or unpinning a resource
>> >> >> > once readers are done has no equivalent. sched_ext's BPF library works
>> >> >> > around this today by pushing freed nodes onto a list and having a
>> >> >> > userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
>> >> >> > BPF program to reclaim them; it is the first intended user.
>> >> >> >
>> >> >> > Add:
>> >> >> >
>> >> >> > int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
>> >> >> > int (*callback)(struct bpf_map *map, void *key,
>> >> >> > void *value));
>> >> >> >
>> >> >> > and bpf_call_rcu_tasks_trace(), same signature, which also waits for
>> >> >> > sleepable programs.
>> >> >> >
>> >> >> > @rh is a struct bpf_rcu_head embedded in a value of @map, so the
>> >> >> > callback runs as callback(map, key, value) for the element it lives in
>> >> >> > and needs no cookie. The field is only accepted in BPF_MAP_TYPE_ARRAY,
>> >> >>
>> >> >> Isn't restricting this to BPF_MAP_TYPE_ARRAY unnecessarily restrictive
>> >> >> in practice? Do other APIs of similar kind have this limitation? E.g.,
>> >> >> bpf_task_work_schedule() or timer, I don't think they limit user to
>> >> >> just ARRAY maps, it's way too inflexible.
>> >> >
>> >> > Right, bpf_timer and bpf_task_work take HASH/LRU_HASH/ARRAY. They can
>> >> > because they can cancel: on element delete bpf_obj_free_fields() calls
>> >>
>> >> bpf_obj_cancel_fields(). bpf_obj_free_fields() is no longer called except when
>> >> the object is being freed finally.
>> >>
>> >> > bpf_task_work_cancel_and_free(), which reaches task_work_cancel(), so
>> >> > a pending callback never runs against a recycled element.
>> >> > A queued RCU callback cannot be cancelled. rcu_barrier() is the only
>> >> > thing that waits for one, and it sleeps, so it is not callable from an
>> >> > element delete. Array elements are never freed individually, which is
>> >> > what makes it safe.
>> >>
>> >> If cancel is a noop, can you not simply skip it on map update?
>> >>
>> >> It is a bit unfortunate we can't provide cancel semantics, it will be yet
>> >> another divergence from how all other async callbacks work...sigh.
>> >>
>> >> >
>> >> > I can think of a complicated way to support hash maps but that would
>> >> > need making the state dynamically allocated and finding a way to
>> >> > cancel the callbacks. But I would do it as a follow-up if there is a
>> >> > real use case.
>> >>
>> >> That would mean allocation and frees, it is already quite expensive as is for a
>> >> call_rcu() primitive (with the refcount bumps and atomics).
>> >
>> > maybe so, but as is ARRAY is a severe limitation and makes a lot of
>> > cases either unusable or requiring extra sequential ID allocation
>> > logic to remap some HASH map entry to index in an ARRAY map, just so
>> > you can call rcu callback for a given kernel object stored in HASH
>> > map.
>> >
>> > E.g., think about keep as set of sockets, or tasks, or whatnot in HASH
>> > or RHASHTABLE (especially the latter one with support for resizing).
>> > How would you do bpf_call_rcu() for them with unduly complications,
>> > limitations and a lot of waste (unlike HASH you can't have lazy memory
>> > allocation).
>>
>> Yeah, I don't disagree. We would also need to support storing them in more map
>> types to make it woth with arenas (which was one of the motivations for the
>> feature).
>>
>> Perhaps it is better to make cancel a noop / unsupported and skip it when a map
>> element is deleted. That will be the easiest path to enabling it, but it does
>> create (perhaps surprising) divergence from some of the other async callback
>> primitives.
>
> I don't see much need for cancellation, but a more useful/sane
> behavior would be rearming on subsequent calls to bpf_call_rcu(). This
> would also handle update/delete/reuse of entries. Not sure how hard it
> is to support that in kernel's call_rcu() implementation, though, but
> that would solve the problem, because it should always be OK to delay
> call_rcu() callback, but not the other way around (which is what would
> happen today because we will ignore subsequent bpf_call_rcu() calls).
>
Hm. One of the design points for this was maintaining our own lists and using
more lower level RCU primitives to poll for the grace period. I wonder if some
of this could be made easier that way. I will think a bit more about this.
We could probably also do the waiting_for_gp amortization that memalloc.c on top.
>>
>> That said, I do think performance should be a consideration for this API; maybe
>> unlike other ones, I would expect that this could be called very frequently. I
>
> I'm not sure why this has to be super high frequency API to use, tbh.
> You'd normally use this when cleaning up when some kernel object is
> freed, no? Sure that can be relatively frequent, but not to the point
> where we should be *that* concerned with refcount or atomics overhead
> per se (multi-cpu cache bouncing of refcounting is a concern, but not
> sure what you can do about that).
>
Yeah it depends, but I think you can make the other case as well. Going by
optimizations made in memalloc.c for hashtab, imagine implementing an arena hash
table using this stuff. You'd definitely want it to be as cheap as possible and
approach the kernel implementation.
If multiple CPUs dispatching it leads to constant cache line bouncing, it will
fail to scale with number of CPUs. The user then does their own batching to
amortize the cost and pace the calls, but it's just more complexity pushed down
on the callers, and it might build up memory pressure because items are now not
being freed as quickly and hit locks in the allocator (one of the main reasons
BPF maps have memory reuse semantics, to avoid exhausting caches and hitting
allocator locks).
Anyway, I know we kicked the can down the road for now on prog refcounts, and
can cache allocations etc. to amortize the cost there, but I hope that I could
illustrate why I think it might be more sensitive to performance differences
than some of the other primitives.
>> had concerns about prog refcount increment as well, but didn't really bring it
>> up since it is something that can be addressed after the fact (maybe using pcpu
>> refs). But once we promise supporting cancellation, it is hard to walk that
>> back. I think it's less of a concern for other async cb types. In practice,
>> except to support map semantics I don't know if anyone will use cancellation API
>> either, the kernel side never grew such support.
>>
>> Anyway, overall the best path to me seems to be just enabling it as is and
>
> you mean enabling for HASH or keeping it for ARRAY only? If the
> latter, my concern is that to support HASH we might need to change the
> internal structure and break that 48-byte size, so we should probably
> decide all this before we get this into the next Linux release.
I did mean enabling it in other maps, yes, sorry, I think I wrote it in a
confusing manner. I was just speculating that we can bail on cancelling and see
whether we can make it work. I probably need to spend a little more time
thinking it through.
All of that said, it seems there's a more discussion to be had about this, and
it probably landed too early. I would prefer if we could resolve these questions
without operating under some time pressure just because it might go out in the
next release.
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace()
2026-09-23 17:57 ` Kumar Kartikeya Dwivedi
@ 2026-09-23 18:21 ` Andrii Nakryiko
2026-09-23 18:30 ` Puranjay Mohan
0 siblings, 1 reply; 17+ messages in thread
From: Andrii Nakryiko @ 2026-09-23 18:21 UTC (permalink / raw)
To: Kumar Kartikeya Dwivedi
Cc: Puranjay Mohan, bpf, rcu, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman, Song Liu,
Yonghong Song, Harry Yoo (Oracle), Paul E. McKenney
On Wed, Sep 23, 2026 at 10:57 AM Kumar Kartikeya Dwivedi
<memxor@gmail.com> wrote:
>
> On Wed Sep 23, 2026 at 6:59 PM CEST, Andrii Nakryiko wrote:
> > On Wed, Sep 23, 2026 at 9:53 AM Kumar Kartikeya Dwivedi
> > <memxor@gmail.com> wrote:
> >>
> >> On Wed Sep 23, 2026 at 6:45 PM CEST, Andrii Nakryiko wrote:
> >> > On Wed, Sep 23, 2026 at 4:29 AM Kumar Kartikeya Dwivedi
> >> > <memxor@gmail.com> wrote:
> >> >>
> >> >> On Wed Sep 23, 2026 at 12:37 PM CEST, Puranjay Mohan wrote:
> >> >> > On Wed, Sep 23, 2026 at 12:41 AM Andrii Nakryiko
> >> >> > <andrii.nakryiko@gmail.com> wrote:
> >> >> >>
> >> >> >> On Tue, Sep 22, 2026 at 1:02 PM Puranjay Mohan <puranjay@kernel.org> wrote:
> >> >> >> >
> >> >> >> > Changelog:
> >> >> >> > v5: https://lore.kernel.org/all/20260921191407.1742386-1-puranjay@kernel.org/
> >> >> >> > Changes in v6:
> >> >> >> > - Size struct bpf_rcu_head at 48 bytes rather than 64, matching what the
> >> >> >> > inline state uses (Alexei)
> >> >> >> > - Trim patch 1's changelog: drop the reasoning about possible future
> >> >> >> > layouts and about what the other async kfuncs return
> >> >> >> > - Rebase on bpf-next/master
> >> >> >> >
> >> >> >> > v4: https://lore.kernel.org/all/20260915154248.3612028-1-puranjay@kernel.org/
> >> >> >> > Changes in v5:
> >> >> >> > - Rename the "hash_map" subtest to "bad_map" so that it matches its
> >> >> >> > helper test_call_rcu_bad_map() (bpf-ci)
> >> >> >> > - Bump the callback counter after chain_err in the selftest callback.
> >> >> >> > Userspace polls that counter and then reads chain_err, so it could
> >> >> >> > still see the initial value before the re-arm had stored one, which
> >> >> >> > let the chain subtest's assertion pass without checking anything
> >> >> >> > - Say in patch 1 why embedding struct rcu_head in a uapi struct is
> >> >> >> > acceptable here (Mykyta, Paul, Alexei)
> >> >> >> > - Rebase on bpf-next/master
> >> >> >> > v3: https://lore.kernel.org/all/20260915143640.36292-1-puranjay@kernel.org/
> >> >> >> > Changes in v4:
> >> >> >> > - Drop an unrelated hunk in bpf_async_update_prog_callback() that turned
> >> >> >> > PTR_ERR(prog) into -EBADF. That is the shared bpf_timer/bpf_wq path
> >> >> >> > and bpf_prog_inc_not_zero() returns -ENOENT, so it would have changed
> >> >> >> > the errno bpf_timer_set_callback() and bpf_wq_set_callback() report to
> >> >> >> > userspace (Sashiko). No other changes from v3
> >> >> >> > v2: https://lore.kernel.org/all/20260915114240.3269184-1-puranjay@kernel.org/
> >> >> >> > Changes in v3:
> >> >> >> > - Move the bpf_call_rcu_tasks_trace() verifier bits from patch 1 to
> >> >> >> > patch 3; patch 1 alone emitted "resolve_btfids: unresolved symbol
> >> >> >> > bpf_call_rcu_tasks_trace" (Sashiko)
> >> >> >> > - Poll the callback counter with an acquire load (Sashiko)
> >> >> >> > - teardown: v2 only checked that the program was eventually freed, which
> >> >> >> > passes even if nothing was ever armed. Also assert that the chain ran,
> >> >> >> > and read the -EPERM back through an independent .bss fd
> >> >> >> > - Use kern_sync_rcu() instead of open coding the grace-period wait
> >> >> >> > - Return -EBADF rather than -ENOENT when the calling program is going
> >> >> >> > away, matching bpf_task_work_schedule()
> >> >> >> > - mismatch_map now pins the bpf_rcu_head label the new code emits; it
> >> >> >> > passed with that branch removed. Add a two_heads test, and wait for
> >> >> >> > each grace period separately in the chain test
> >> >> >> > - Commit messages: correct the -EPERM parity claim, explain the inline
> >> >> >> > callback state and the struct size, motivate the tasks trace flavour
> >> >> >> > v1: https://lore.kernel.org/all/20260907134552.1772405-1-puranjay@kernel.org/
> >> >> >> > Changes in v2:
> >> >> >> > - Rebase on bpf-next/master
> >> >> >> > - Use rcu_read_lock_dont_migrate() over open coding (Alexei)
> >> >> >> > - Improve re-arming selftest to detect failure (Sashiko)
> >> >> >> >
> >> >> >> > BPF programs that manage their own objects have no way to run their own
> >> >> >> > logic once an RCU grace period has elapsed. bpf_obj_drop() defers a
> >> >> >> > free, but returning an index to an allocator or unpinning a resource
> >> >> >> > once readers are done has no equivalent. sched_ext's BPF library works
> >> >> >> > around this today by pushing freed nodes onto a list and having a
> >> >> >> > userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
> >> >> >> > BPF program to reclaim them; it is the first intended user.
> >> >> >> >
> >> >> >> > Add:
> >> >> >> >
> >> >> >> > int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
> >> >> >> > int (*callback)(struct bpf_map *map, void *key,
> >> >> >> > void *value));
> >> >> >> >
> >> >> >> > and bpf_call_rcu_tasks_trace(), same signature, which also waits for
> >> >> >> > sleepable programs.
> >> >> >> >
> >> >> >> > @rh is a struct bpf_rcu_head embedded in a value of @map, so the
> >> >> >> > callback runs as callback(map, key, value) for the element it lives in
> >> >> >> > and needs no cookie. The field is only accepted in BPF_MAP_TYPE_ARRAY,
> >> >> >>
> >> >> >> Isn't restricting this to BPF_MAP_TYPE_ARRAY unnecessarily restrictive
> >> >> >> in practice? Do other APIs of similar kind have this limitation? E.g.,
> >> >> >> bpf_task_work_schedule() or timer, I don't think they limit user to
> >> >> >> just ARRAY maps, it's way too inflexible.
> >> >> >
> >> >> > Right, bpf_timer and bpf_task_work take HASH/LRU_HASH/ARRAY. They can
> >> >> > because they can cancel: on element delete bpf_obj_free_fields() calls
> >> >>
> >> >> bpf_obj_cancel_fields(). bpf_obj_free_fields() is no longer called except when
> >> >> the object is being freed finally.
> >> >>
> >> >> > bpf_task_work_cancel_and_free(), which reaches task_work_cancel(), so
> >> >> > a pending callback never runs against a recycled element.
> >> >> > A queued RCU callback cannot be cancelled. rcu_barrier() is the only
> >> >> > thing that waits for one, and it sleeps, so it is not callable from an
> >> >> > element delete. Array elements are never freed individually, which is
> >> >> > what makes it safe.
> >> >>
> >> >> If cancel is a noop, can you not simply skip it on map update?
> >> >>
> >> >> It is a bit unfortunate we can't provide cancel semantics, it will be yet
> >> >> another divergence from how all other async callbacks work...sigh.
> >> >>
> >> >> >
> >> >> > I can think of a complicated way to support hash maps but that would
> >> >> > need making the state dynamically allocated and finding a way to
> >> >> > cancel the callbacks. But I would do it as a follow-up if there is a
> >> >> > real use case.
> >> >>
> >> >> That would mean allocation and frees, it is already quite expensive as is for a
> >> >> call_rcu() primitive (with the refcount bumps and atomics).
> >> >
> >> > maybe so, but as is ARRAY is a severe limitation and makes a lot of
> >> > cases either unusable or requiring extra sequential ID allocation
> >> > logic to remap some HASH map entry to index in an ARRAY map, just so
> >> > you can call rcu callback for a given kernel object stored in HASH
> >> > map.
> >> >
> >> > E.g., think about keep as set of sockets, or tasks, or whatnot in HASH
> >> > or RHASHTABLE (especially the latter one with support for resizing).
> >> > How would you do bpf_call_rcu() for them with unduly complications,
> >> > limitations and a lot of waste (unlike HASH you can't have lazy memory
> >> > allocation).
> >>
> >> Yeah, I don't disagree. We would also need to support storing them in more map
> >> types to make it woth with arenas (which was one of the motivations for the
> >> feature).
> >>
> >> Perhaps it is better to make cancel a noop / unsupported and skip it when a map
> >> element is deleted. That will be the easiest path to enabling it, but it does
> >> create (perhaps surprising) divergence from some of the other async callback
> >> primitives.
> >
> > I don't see much need for cancellation, but a more useful/sane
> > behavior would be rearming on subsequent calls to bpf_call_rcu(). This
> > would also handle update/delete/reuse of entries. Not sure how hard it
> > is to support that in kernel's call_rcu() implementation, though, but
> > that would solve the problem, because it should always be OK to delay
> > call_rcu() callback, but not the other way around (which is what would
> > happen today because we will ignore subsequent bpf_call_rcu() calls).
> >
>
> Hm. One of the design points for this was maintaining our own lists and using
> more lower level RCU primitives to poll for the grace period. I wonder if some
> of this could be made easier that way. I will think a bit more about this.
>
> We could probably also do the waiting_for_gp amortization that memalloc.c on top.
>
> >>
> >> That said, I do think performance should be a consideration for this API; maybe
> >> unlike other ones, I would expect that this could be called very frequently. I
> >
> > I'm not sure why this has to be super high frequency API to use, tbh.
> > You'd normally use this when cleaning up when some kernel object is
> > freed, no? Sure that can be relatively frequent, but not to the point
> > where we should be *that* concerned with refcount or atomics overhead
> > per se (multi-cpu cache bouncing of refcounting is a concern, but not
> > sure what you can do about that).
> >
>
> Yeah it depends, but I think you can make the other case as well. Going by
> optimizations made in memalloc.c for hashtab, imagine implementing an arena hash
> table using this stuff. You'd definitely want it to be as cheap as possible and
> approach the kernel implementation.
>
> If multiple CPUs dispatching it leads to constant cache line bouncing, it will
> fail to scale with number of CPUs. The user then does their own batching to
> amortize the cost and pace the calls, but it's just more complexity pushed down
> on the callers, and it might build up memory pressure because items are now not
> being freed as quickly and hit locks in the allocator (one of the main reasons
> BPF maps have memory reuse semantics, to avoid exhausting caches and hitting
> allocator locks).
>
> Anyway, I know we kicked the can down the road for now on prog refcounts, and
> can cache allocations etc. to amortize the cost there, but I hope that I could
> illustrate why I think it might be more sensitive to performance differences
> than some of the other primitives.
no, that's a good point, but I think the solution here is to make
prog's refcount more scalable by using per-cpu refcount approach,
right?
>
> >> had concerns about prog refcount increment as well, but didn't really bring it
> >> up since it is something that can be addressed after the fact (maybe using pcpu
> >> refs). But once we promise supporting cancellation, it is hard to walk that
> >> back. I think it's less of a concern for other async cb types. In practice,
> >> except to support map semantics I don't know if anyone will use cancellation API
> >> either, the kernel side never grew such support.
> >>
> >> Anyway, overall the best path to me seems to be just enabling it as is and
> >
> > you mean enabling for HASH or keeping it for ARRAY only? If the
> > latter, my concern is that to support HASH we might need to change the
> > internal structure and break that 48-byte size, so we should probably
> > decide all this before we get this into the next Linux release.
>
> I did mean enabling it in other maps, yes, sorry, I think I wrote it in a
ah ok, +1 to that
> confusing manner. I was just speculating that we can bail on cancelling and see
> whether we can make it work. I probably need to spend a little more time
> thinking it through.
>
I think there is, I'll let Puranjay provide details as we just had a
discussion offlist.
> All of that said, it seems there's a more discussion to be had about this, and
> it probably landed too early. I would prefer if we could resolve these questions
> without operating under some time pressure just because it might go out in the
> next release.
+1
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace()
2026-09-23 18:21 ` Andrii Nakryiko
@ 2026-09-23 18:30 ` Puranjay Mohan
2026-09-23 19:25 ` Andrii Nakryiko
0 siblings, 1 reply; 17+ messages in thread
From: Puranjay Mohan @ 2026-09-23 18:30 UTC (permalink / raw)
To: Andrii Nakryiko
Cc: Kumar Kartikeya Dwivedi, bpf, rcu, Alexei Starovoitov,
Daniel Borkmann, Andrii Nakryiko, Martin KaFai Lau,
Eduard Zingerman, Song Liu, Yonghong Song, Harry Yoo (Oracle),
Paul E. McKenney
On Wed, Sep 23, 2026 at 7:22 PM Andrii Nakryiko
<andrii.nakryiko@gmail.com> wrote:
>
> On Wed, Sep 23, 2026 at 10:57 AM Kumar Kartikeya Dwivedi
> <memxor@gmail.com> wrote:
> >
> > On Wed Sep 23, 2026 at 6:59 PM CEST, Andrii Nakryiko wrote:
> > > On Wed, Sep 23, 2026 at 9:53 AM Kumar Kartikeya Dwivedi
> > > <memxor@gmail.com> wrote:
> > >>
> > >> On Wed Sep 23, 2026 at 6:45 PM CEST, Andrii Nakryiko wrote:
> > >> > On Wed, Sep 23, 2026 at 4:29 AM Kumar Kartikeya Dwivedi
> > >> > <memxor@gmail.com> wrote:
> > >> >>
> > >> >> On Wed Sep 23, 2026 at 12:37 PM CEST, Puranjay Mohan wrote:
> > >> >> > On Wed, Sep 23, 2026 at 12:41 AM Andrii Nakryiko
> > >> >> > <andrii.nakryiko@gmail.com> wrote:
> > >> >> >>
> > >> >> >> On Tue, Sep 22, 2026 at 1:02 PM Puranjay Mohan <puranjay@kernel.org> wrote:
> > >> >> >> >
> > >> >> >> > Changelog:
> > >> >> >> > v5: https://lore.kernel.org/all/20260921191407.1742386-1-puranjay@kernel.org/
> > >> >> >> > Changes in v6:
> > >> >> >> > - Size struct bpf_rcu_head at 48 bytes rather than 64, matching what the
> > >> >> >> > inline state uses (Alexei)
> > >> >> >> > - Trim patch 1's changelog: drop the reasoning about possible future
> > >> >> >> > layouts and about what the other async kfuncs return
> > >> >> >> > - Rebase on bpf-next/master
> > >> >> >> >
> > >> >> >> > v4: https://lore.kernel.org/all/20260915154248.3612028-1-puranjay@kernel.org/
> > >> >> >> > Changes in v5:
> > >> >> >> > - Rename the "hash_map" subtest to "bad_map" so that it matches its
> > >> >> >> > helper test_call_rcu_bad_map() (bpf-ci)
> > >> >> >> > - Bump the callback counter after chain_err in the selftest callback.
> > >> >> >> > Userspace polls that counter and then reads chain_err, so it could
> > >> >> >> > still see the initial value before the re-arm had stored one, which
> > >> >> >> > let the chain subtest's assertion pass without checking anything
> > >> >> >> > - Say in patch 1 why embedding struct rcu_head in a uapi struct is
> > >> >> >> > acceptable here (Mykyta, Paul, Alexei)
> > >> >> >> > - Rebase on bpf-next/master
> > >> >> >> > v3: https://lore.kernel.org/all/20260915143640.36292-1-puranjay@kernel.org/
> > >> >> >> > Changes in v4:
> > >> >> >> > - Drop an unrelated hunk in bpf_async_update_prog_callback() that turned
> > >> >> >> > PTR_ERR(prog) into -EBADF. That is the shared bpf_timer/bpf_wq path
> > >> >> >> > and bpf_prog_inc_not_zero() returns -ENOENT, so it would have changed
> > >> >> >> > the errno bpf_timer_set_callback() and bpf_wq_set_callback() report to
> > >> >> >> > userspace (Sashiko). No other changes from v3
> > >> >> >> > v2: https://lore.kernel.org/all/20260915114240.3269184-1-puranjay@kernel.org/
> > >> >> >> > Changes in v3:
> > >> >> >> > - Move the bpf_call_rcu_tasks_trace() verifier bits from patch 1 to
> > >> >> >> > patch 3; patch 1 alone emitted "resolve_btfids: unresolved symbol
> > >> >> >> > bpf_call_rcu_tasks_trace" (Sashiko)
> > >> >> >> > - Poll the callback counter with an acquire load (Sashiko)
> > >> >> >> > - teardown: v2 only checked that the program was eventually freed, which
> > >> >> >> > passes even if nothing was ever armed. Also assert that the chain ran,
> > >> >> >> > and read the -EPERM back through an independent .bss fd
> > >> >> >> > - Use kern_sync_rcu() instead of open coding the grace-period wait
> > >> >> >> > - Return -EBADF rather than -ENOENT when the calling program is going
> > >> >> >> > away, matching bpf_task_work_schedule()
> > >> >> >> > - mismatch_map now pins the bpf_rcu_head label the new code emits; it
> > >> >> >> > passed with that branch removed. Add a two_heads test, and wait for
> > >> >> >> > each grace period separately in the chain test
> > >> >> >> > - Commit messages: correct the -EPERM parity claim, explain the inline
> > >> >> >> > callback state and the struct size, motivate the tasks trace flavour
> > >> >> >> > v1: https://lore.kernel.org/all/20260907134552.1772405-1-puranjay@kernel.org/
> > >> >> >> > Changes in v2:
> > >> >> >> > - Rebase on bpf-next/master
> > >> >> >> > - Use rcu_read_lock_dont_migrate() over open coding (Alexei)
> > >> >> >> > - Improve re-arming selftest to detect failure (Sashiko)
> > >> >> >> >
> > >> >> >> > BPF programs that manage their own objects have no way to run their own
> > >> >> >> > logic once an RCU grace period has elapsed. bpf_obj_drop() defers a
> > >> >> >> > free, but returning an index to an allocator or unpinning a resource
> > >> >> >> > once readers are done has no equivalent. sched_ext's BPF library works
> > >> >> >> > around this today by pushing freed nodes onto a list and having a
> > >> >> >> > userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
> > >> >> >> > BPF program to reclaim them; it is the first intended user.
> > >> >> >> >
> > >> >> >> > Add:
> > >> >> >> >
> > >> >> >> > int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
> > >> >> >> > int (*callback)(struct bpf_map *map, void *key,
> > >> >> >> > void *value));
> > >> >> >> >
> > >> >> >> > and bpf_call_rcu_tasks_trace(), same signature, which also waits for
> > >> >> >> > sleepable programs.
> > >> >> >> >
> > >> >> >> > @rh is a struct bpf_rcu_head embedded in a value of @map, so the
> > >> >> >> > callback runs as callback(map, key, value) for the element it lives in
> > >> >> >> > and needs no cookie. The field is only accepted in BPF_MAP_TYPE_ARRAY,
> > >> >> >>
> > >> >> >> Isn't restricting this to BPF_MAP_TYPE_ARRAY unnecessarily restrictive
> > >> >> >> in practice? Do other APIs of similar kind have this limitation? E.g.,
> > >> >> >> bpf_task_work_schedule() or timer, I don't think they limit user to
> > >> >> >> just ARRAY maps, it's way too inflexible.
> > >> >> >
> > >> >> > Right, bpf_timer and bpf_task_work take HASH/LRU_HASH/ARRAY. They can
> > >> >> > because they can cancel: on element delete bpf_obj_free_fields() calls
> > >> >>
> > >> >> bpf_obj_cancel_fields(). bpf_obj_free_fields() is no longer called except when
> > >> >> the object is being freed finally.
> > >> >>
> > >> >> > bpf_task_work_cancel_and_free(), which reaches task_work_cancel(), so
> > >> >> > a pending callback never runs against a recycled element.
> > >> >> > A queued RCU callback cannot be cancelled. rcu_barrier() is the only
> > >> >> > thing that waits for one, and it sleeps, so it is not callable from an
> > >> >> > element delete. Array elements are never freed individually, which is
> > >> >> > what makes it safe.
> > >> >>
> > >> >> If cancel is a noop, can you not simply skip it on map update?
> > >> >>
> > >> >> It is a bit unfortunate we can't provide cancel semantics, it will be yet
> > >> >> another divergence from how all other async callbacks work...sigh.
> > >> >>
> > >> >> >
> > >> >> > I can think of a complicated way to support hash maps but that would
> > >> >> > need making the state dynamically allocated and finding a way to
> > >> >> > cancel the callbacks. But I would do it as a follow-up if there is a
> > >> >> > real use case.
> > >> >>
> > >> >> That would mean allocation and frees, it is already quite expensive as is for a
> > >> >> call_rcu() primitive (with the refcount bumps and atomics).
> > >> >
> > >> > maybe so, but as is ARRAY is a severe limitation and makes a lot of
> > >> > cases either unusable or requiring extra sequential ID allocation
> > >> > logic to remap some HASH map entry to index in an ARRAY map, just so
> > >> > you can call rcu callback for a given kernel object stored in HASH
> > >> > map.
> > >> >
> > >> > E.g., think about keep as set of sockets, or tasks, or whatnot in HASH
> > >> > or RHASHTABLE (especially the latter one with support for resizing).
> > >> > How would you do bpf_call_rcu() for them with unduly complications,
> > >> > limitations and a lot of waste (unlike HASH you can't have lazy memory
> > >> > allocation).
> > >>
> > >> Yeah, I don't disagree. We would also need to support storing them in more map
> > >> types to make it woth with arenas (which was one of the motivations for the
> > >> feature).
> > >>
> > >> Perhaps it is better to make cancel a noop / unsupported and skip it when a map
> > >> element is deleted. That will be the easiest path to enabling it, but it does
> > >> create (perhaps surprising) divergence from some of the other async callback
> > >> primitives.
> > >
> > > I don't see much need for cancellation, but a more useful/sane
> > > behavior would be rearming on subsequent calls to bpf_call_rcu(). This
> > > would also handle update/delete/reuse of entries. Not sure how hard it
> > > is to support that in kernel's call_rcu() implementation, though, but
> > > that would solve the problem, because it should always be OK to delay
> > > call_rcu() callback, but not the other way around (which is what would
> > > happen today because we will ignore subsequent bpf_call_rcu() calls).
> > >
> >
> > Hm. One of the design points for this was maintaining our own lists and using
> > more lower level RCU primitives to poll for the grace period. I wonder if some
> > of this could be made easier that way. I will think a bit more about this.
> >
> > We could probably also do the waiting_for_gp amortization that memalloc.c on top.
> >
> > >>
> > >> That said, I do think performance should be a consideration for this API; maybe
> > >> unlike other ones, I would expect that this could be called very frequently. I
> > >
> > > I'm not sure why this has to be super high frequency API to use, tbh.
> > > You'd normally use this when cleaning up when some kernel object is
> > > freed, no? Sure that can be relatively frequent, but not to the point
> > > where we should be *that* concerned with refcount or atomics overhead
> > > per se (multi-cpu cache bouncing of refcounting is a concern, but not
> > > sure what you can do about that).
> > >
> >
> > Yeah it depends, but I think you can make the other case as well. Going by
> > optimizations made in memalloc.c for hashtab, imagine implementing an arena hash
> > table using this stuff. You'd definitely want it to be as cheap as possible and
> > approach the kernel implementation.
> >
> > If multiple CPUs dispatching it leads to constant cache line bouncing, it will
> > fail to scale with number of CPUs. The user then does their own batching to
> > amortize the cost and pace the calls, but it's just more complexity pushed down
> > on the callers, and it might build up memory pressure because items are now not
> > being freed as quickly and hit locks in the allocator (one of the main reasons
> > BPF maps have memory reuse semantics, to avoid exhausting caches and hitting
> > allocator locks).
> >
> > Anyway, I know we kicked the can down the road for now on prog refcounts, and
> > can cache allocations etc. to amortize the cost there, but I hope that I could
> > illustrate why I think it might be more sensitive to performance differences
> > than some of the other primitives.
>
> no, that's a good point, but I think the solution here is to make
> prog's refcount more scalable by using per-cpu refcount approach,
> right?
>
> >
> > >> had concerns about prog refcount increment as well, but didn't really bring it
> > >> up since it is something that can be addressed after the fact (maybe using pcpu
> > >> refs). But once we promise supporting cancellation, it is hard to walk that
> > >> back. I think it's less of a concern for other async cb types. In practice,
> > >> except to support map semantics I don't know if anyone will use cancellation API
> > >> either, the kernel side never grew such support.
> > >>
> > >> Anyway, overall the best path to me seems to be just enabling it as is and
> > >
> > > you mean enabling for HASH or keeping it for ARRAY only? If the
> > > latter, my concern is that to support HASH we might need to change the
> > > internal structure and break that 48-byte size, so we should probably
> > > decide all this before we get this into the next Linux release.
> >
> > I did mean enabling it in other maps, yes, sorry, I think I wrote it in a
>
> ah ok, +1 to that
>
> > confusing manner. I was just speculating that we can bail on cancelling and see
> > whether we can make it work. I probably need to spend a little more time
> > thinking it through.
> >
>
> I think there is, I'll let Puranjay provide details as we just had a
> discussion offlist.
I think I can come up with a way for supporting hashmaps, by keeping
the element alive until the callback has run rather than trying to
cancel it.
On delete the map asks the head first: that sets a dead bit and
reports whether a callback is still outstanding. If one is, the
element stays unlinked but is not returned to the allocator; the
callback does that at the end through a new map_release_elem(), since
only the map knows whether that is a freelist push or
bpf_mem_cache_free(). The dead bit stops it being armed again, so
there is a definite last callback and exactly one side releases the
element.
That needs one more bit of state in the head that covers the rcu_head,
which RCU dequeues before invoking so the callback can still re-arm,
and a count that covers the element, which has to live until the last
callback
returns. Both fit in the u32 already there, so struct bpf_rcu_head
does not grow and there is no allocation.
What do you think of this approach??
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace()
2026-09-23 18:30 ` Puranjay Mohan
@ 2026-09-23 19:25 ` Andrii Nakryiko
0 siblings, 0 replies; 17+ messages in thread
From: Andrii Nakryiko @ 2026-09-23 19:25 UTC (permalink / raw)
To: Puranjay Mohan
Cc: Kumar Kartikeya Dwivedi, bpf, rcu, Alexei Starovoitov,
Daniel Borkmann, Andrii Nakryiko, Martin KaFai Lau,
Eduard Zingerman, Song Liu, Yonghong Song, Harry Yoo (Oracle),
Paul E. McKenney
On Wed, Sep 23, 2026 at 11:31 AM Puranjay Mohan <puranjay12@gmail.com> wrote:
>
> On Wed, Sep 23, 2026 at 7:22 PM Andrii Nakryiko
> <andrii.nakryiko@gmail.com> wrote:
> >
> > On Wed, Sep 23, 2026 at 10:57 AM Kumar Kartikeya Dwivedi
> > <memxor@gmail.com> wrote:
> > >
> > > On Wed Sep 23, 2026 at 6:59 PM CEST, Andrii Nakryiko wrote:
> > > > On Wed, Sep 23, 2026 at 9:53 AM Kumar Kartikeya Dwivedi
> > > > <memxor@gmail.com> wrote:
> > > >>
> > > >> On Wed Sep 23, 2026 at 6:45 PM CEST, Andrii Nakryiko wrote:
> > > >> > On Wed, Sep 23, 2026 at 4:29 AM Kumar Kartikeya Dwivedi
> > > >> > <memxor@gmail.com> wrote:
> > > >> >>
> > > >> >> On Wed Sep 23, 2026 at 12:37 PM CEST, Puranjay Mohan wrote:
> > > >> >> > On Wed, Sep 23, 2026 at 12:41 AM Andrii Nakryiko
> > > >> >> > <andrii.nakryiko@gmail.com> wrote:
> > > >> >> >>
> > > >> >> >> On Tue, Sep 22, 2026 at 1:02 PM Puranjay Mohan <puranjay@kernel.org> wrote:
> > > >> >> >> >
> > > >> >> >> > Changelog:
> > > >> >> >> > v5: https://lore.kernel.org/all/20260921191407.1742386-1-puranjay@kernel.org/
> > > >> >> >> > Changes in v6:
> > > >> >> >> > - Size struct bpf_rcu_head at 48 bytes rather than 64, matching what the
> > > >> >> >> > inline state uses (Alexei)
> > > >> >> >> > - Trim patch 1's changelog: drop the reasoning about possible future
> > > >> >> >> > layouts and about what the other async kfuncs return
> > > >> >> >> > - Rebase on bpf-next/master
> > > >> >> >> >
> > > >> >> >> > v4: https://lore.kernel.org/all/20260915154248.3612028-1-puranjay@kernel.org/
> > > >> >> >> > Changes in v5:
> > > >> >> >> > - Rename the "hash_map" subtest to "bad_map" so that it matches its
> > > >> >> >> > helper test_call_rcu_bad_map() (bpf-ci)
> > > >> >> >> > - Bump the callback counter after chain_err in the selftest callback.
> > > >> >> >> > Userspace polls that counter and then reads chain_err, so it could
> > > >> >> >> > still see the initial value before the re-arm had stored one, which
> > > >> >> >> > let the chain subtest's assertion pass without checking anything
> > > >> >> >> > - Say in patch 1 why embedding struct rcu_head in a uapi struct is
> > > >> >> >> > acceptable here (Mykyta, Paul, Alexei)
> > > >> >> >> > - Rebase on bpf-next/master
> > > >> >> >> > v3: https://lore.kernel.org/all/20260915143640.36292-1-puranjay@kernel.org/
> > > >> >> >> > Changes in v4:
> > > >> >> >> > - Drop an unrelated hunk in bpf_async_update_prog_callback() that turned
> > > >> >> >> > PTR_ERR(prog) into -EBADF. That is the shared bpf_timer/bpf_wq path
> > > >> >> >> > and bpf_prog_inc_not_zero() returns -ENOENT, so it would have changed
> > > >> >> >> > the errno bpf_timer_set_callback() and bpf_wq_set_callback() report to
> > > >> >> >> > userspace (Sashiko). No other changes from v3
> > > >> >> >> > v2: https://lore.kernel.org/all/20260915114240.3269184-1-puranjay@kernel.org/
> > > >> >> >> > Changes in v3:
> > > >> >> >> > - Move the bpf_call_rcu_tasks_trace() verifier bits from patch 1 to
> > > >> >> >> > patch 3; patch 1 alone emitted "resolve_btfids: unresolved symbol
> > > >> >> >> > bpf_call_rcu_tasks_trace" (Sashiko)
> > > >> >> >> > - Poll the callback counter with an acquire load (Sashiko)
> > > >> >> >> > - teardown: v2 only checked that the program was eventually freed, which
> > > >> >> >> > passes even if nothing was ever armed. Also assert that the chain ran,
> > > >> >> >> > and read the -EPERM back through an independent .bss fd
> > > >> >> >> > - Use kern_sync_rcu() instead of open coding the grace-period wait
> > > >> >> >> > - Return -EBADF rather than -ENOENT when the calling program is going
> > > >> >> >> > away, matching bpf_task_work_schedule()
> > > >> >> >> > - mismatch_map now pins the bpf_rcu_head label the new code emits; it
> > > >> >> >> > passed with that branch removed. Add a two_heads test, and wait for
> > > >> >> >> > each grace period separately in the chain test
> > > >> >> >> > - Commit messages: correct the -EPERM parity claim, explain the inline
> > > >> >> >> > callback state and the struct size, motivate the tasks trace flavour
> > > >> >> >> > v1: https://lore.kernel.org/all/20260907134552.1772405-1-puranjay@kernel.org/
> > > >> >> >> > Changes in v2:
> > > >> >> >> > - Rebase on bpf-next/master
> > > >> >> >> > - Use rcu_read_lock_dont_migrate() over open coding (Alexei)
> > > >> >> >> > - Improve re-arming selftest to detect failure (Sashiko)
> > > >> >> >> >
> > > >> >> >> > BPF programs that manage their own objects have no way to run their own
> > > >> >> >> > logic once an RCU grace period has elapsed. bpf_obj_drop() defers a
> > > >> >> >> > free, but returning an index to an allocator or unpinning a resource
> > > >> >> >> > once readers are done has no equivalent. sched_ext's BPF library works
> > > >> >> >> > around this today by pushing freed nodes onto a list and having a
> > > >> >> >> > userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
> > > >> >> >> > BPF program to reclaim them; it is the first intended user.
> > > >> >> >> >
> > > >> >> >> > Add:
> > > >> >> >> >
> > > >> >> >> > int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
> > > >> >> >> > int (*callback)(struct bpf_map *map, void *key,
> > > >> >> >> > void *value));
> > > >> >> >> >
> > > >> >> >> > and bpf_call_rcu_tasks_trace(), same signature, which also waits for
> > > >> >> >> > sleepable programs.
> > > >> >> >> >
> > > >> >> >> > @rh is a struct bpf_rcu_head embedded in a value of @map, so the
> > > >> >> >> > callback runs as callback(map, key, value) for the element it lives in
> > > >> >> >> > and needs no cookie. The field is only accepted in BPF_MAP_TYPE_ARRAY,
> > > >> >> >>
> > > >> >> >> Isn't restricting this to BPF_MAP_TYPE_ARRAY unnecessarily restrictive
> > > >> >> >> in practice? Do other APIs of similar kind have this limitation? E.g.,
> > > >> >> >> bpf_task_work_schedule() or timer, I don't think they limit user to
> > > >> >> >> just ARRAY maps, it's way too inflexible.
> > > >> >> >
> > > >> >> > Right, bpf_timer and bpf_task_work take HASH/LRU_HASH/ARRAY. They can
> > > >> >> > because they can cancel: on element delete bpf_obj_free_fields() calls
> > > >> >>
> > > >> >> bpf_obj_cancel_fields(). bpf_obj_free_fields() is no longer called except when
> > > >> >> the object is being freed finally.
> > > >> >>
> > > >> >> > bpf_task_work_cancel_and_free(), which reaches task_work_cancel(), so
> > > >> >> > a pending callback never runs against a recycled element.
> > > >> >> > A queued RCU callback cannot be cancelled. rcu_barrier() is the only
> > > >> >> > thing that waits for one, and it sleeps, so it is not callable from an
> > > >> >> > element delete. Array elements are never freed individually, which is
> > > >> >> > what makes it safe.
> > > >> >>
> > > >> >> If cancel is a noop, can you not simply skip it on map update?
> > > >> >>
> > > >> >> It is a bit unfortunate we can't provide cancel semantics, it will be yet
> > > >> >> another divergence from how all other async callbacks work...sigh.
> > > >> >>
> > > >> >> >
> > > >> >> > I can think of a complicated way to support hash maps but that would
> > > >> >> > need making the state dynamically allocated and finding a way to
> > > >> >> > cancel the callbacks. But I would do it as a follow-up if there is a
> > > >> >> > real use case.
> > > >> >>
> > > >> >> That would mean allocation and frees, it is already quite expensive as is for a
> > > >> >> call_rcu() primitive (with the refcount bumps and atomics).
> > > >> >
> > > >> > maybe so, but as is ARRAY is a severe limitation and makes a lot of
> > > >> > cases either unusable or requiring extra sequential ID allocation
> > > >> > logic to remap some HASH map entry to index in an ARRAY map, just so
> > > >> > you can call rcu callback for a given kernel object stored in HASH
> > > >> > map.
> > > >> >
> > > >> > E.g., think about keep as set of sockets, or tasks, or whatnot in HASH
> > > >> > or RHASHTABLE (especially the latter one with support for resizing).
> > > >> > How would you do bpf_call_rcu() for them with unduly complications,
> > > >> > limitations and a lot of waste (unlike HASH you can't have lazy memory
> > > >> > allocation).
> > > >>
> > > >> Yeah, I don't disagree. We would also need to support storing them in more map
> > > >> types to make it woth with arenas (which was one of the motivations for the
> > > >> feature).
> > > >>
> > > >> Perhaps it is better to make cancel a noop / unsupported and skip it when a map
> > > >> element is deleted. That will be the easiest path to enabling it, but it does
> > > >> create (perhaps surprising) divergence from some of the other async callback
> > > >> primitives.
> > > >
> > > > I don't see much need for cancellation, but a more useful/sane
> > > > behavior would be rearming on subsequent calls to bpf_call_rcu(). This
> > > > would also handle update/delete/reuse of entries. Not sure how hard it
> > > > is to support that in kernel's call_rcu() implementation, though, but
> > > > that would solve the problem, because it should always be OK to delay
> > > > call_rcu() callback, but not the other way around (which is what would
> > > > happen today because we will ignore subsequent bpf_call_rcu() calls).
> > > >
> > >
> > > Hm. One of the design points for this was maintaining our own lists and using
> > > more lower level RCU primitives to poll for the grace period. I wonder if some
> > > of this could be made easier that way. I will think a bit more about this.
> > >
> > > We could probably also do the waiting_for_gp amortization that memalloc.c on top.
> > >
> > > >>
> > > >> That said, I do think performance should be a consideration for this API; maybe
> > > >> unlike other ones, I would expect that this could be called very frequently. I
> > > >
> > > > I'm not sure why this has to be super high frequency API to use, tbh.
> > > > You'd normally use this when cleaning up when some kernel object is
> > > > freed, no? Sure that can be relatively frequent, but not to the point
> > > > where we should be *that* concerned with refcount or atomics overhead
> > > > per se (multi-cpu cache bouncing of refcounting is a concern, but not
> > > > sure what you can do about that).
> > > >
> > >
> > > Yeah it depends, but I think you can make the other case as well. Going by
> > > optimizations made in memalloc.c for hashtab, imagine implementing an arena hash
> > > table using this stuff. You'd definitely want it to be as cheap as possible and
> > > approach the kernel implementation.
> > >
> > > If multiple CPUs dispatching it leads to constant cache line bouncing, it will
> > > fail to scale with number of CPUs. The user then does their own batching to
> > > amortize the cost and pace the calls, but it's just more complexity pushed down
> > > on the callers, and it might build up memory pressure because items are now not
> > > being freed as quickly and hit locks in the allocator (one of the main reasons
> > > BPF maps have memory reuse semantics, to avoid exhausting caches and hitting
> > > allocator locks).
> > >
> > > Anyway, I know we kicked the can down the road for now on prog refcounts, and
> > > can cache allocations etc. to amortize the cost there, but I hope that I could
> > > illustrate why I think it might be more sensitive to performance differences
> > > than some of the other primitives.
> >
> > no, that's a good point, but I think the solution here is to make
> > prog's refcount more scalable by using per-cpu refcount approach,
> > right?
> >
> > >
> > > >> had concerns about prog refcount increment as well, but didn't really bring it
> > > >> up since it is something that can be addressed after the fact (maybe using pcpu
> > > >> refs). But once we promise supporting cancellation, it is hard to walk that
> > > >> back. I think it's less of a concern for other async cb types. In practice,
> > > >> except to support map semantics I don't know if anyone will use cancellation API
> > > >> either, the kernel side never grew such support.
> > > >>
> > > >> Anyway, overall the best path to me seems to be just enabling it as is and
> > > >
> > > > you mean enabling for HASH or keeping it for ARRAY only? If the
> > > > latter, my concern is that to support HASH we might need to change the
> > > > internal structure and break that 48-byte size, so we should probably
> > > > decide all this before we get this into the next Linux release.
> > >
> > > I did mean enabling it in other maps, yes, sorry, I think I wrote it in a
> >
> > ah ok, +1 to that
> >
> > > confusing manner. I was just speculating that we can bail on cancelling and see
> > > whether we can make it work. I probably need to spend a little more time
> > > thinking it through.
> > >
> >
> > I think there is, I'll let Puranjay provide details as we just had a
> > discussion offlist.
>
> I think I can come up with a way for supporting hashmaps, by keeping
> the element alive until the callback has run rather than trying to
> cancel it.
>
> On delete the map asks the head first: that sets a dead bit and
> reports whether a callback is still outstanding. If one is, the
> element stays unlinked but is not returned to the allocator; the
> callback does that at the end through a new map_release_elem(), since
> only the map knows whether that is a freelist push or
> bpf_mem_cache_free(). The dead bit stops it being armed again, so
> there is a definite last callback and exactly one side releases the
> element.
>
> That needs one more bit of state in the head that covers the rcu_head,
> which RCU dequeues before invoking so the callback can still re-arm,
> and a count that covers the element, which has to live until the last
> callback
> returns. Both fit in the u32 already there, so struct bpf_rcu_head
> does not grow and there is no allocation.
>
> What do you think of this approach??
I haven't looked in details into the current bpf_rcu_head
implementation, but overall all this makes sense. Please give it a try
and let's see how it ends up looking and working.
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace()
2026-09-22 20:00 [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace() Puranjay Mohan
` (4 preceding siblings ...)
2026-09-22 23:41 ` [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace() Andrii Nakryiko
@ 2026-09-22 23:50 ` patchwork-bot+netdevbpf
5 siblings, 0 replies; 17+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-09-22 23:50 UTC (permalink / raw)
To: Puranjay Mohan
Cc: bpf, rcu, ast, daniel, andrii, martin.lau, eddyz87, memxor, song,
yonghong.song, harry, paulmck
Hello:
This series was applied to bpf/bpf-next.git (master)
by Alexei Starovoitov <ast@kernel.org>:
On Tue, 22 Sep 2026 13:00:49 -0700 you wrote:
> Changelog:
> v5: https://lore.kernel.org/all/20260921191407.1742386-1-puranjay@kernel.org/
> Changes in v6:
> - Size struct bpf_rcu_head at 48 bytes rather than 64, matching what the
> inline state uses (Alexei)
> - Trim patch 1's changelog: drop the reasoning about possible future
> layouts and about what the other async kfuncs return
> - Rebase on bpf-next/master
>
> [...]
Here is the summary with links:
- [bpf-next,v6,1/4] bpf: Add bpf_call_rcu() kfunc
https://git.kernel.org/bpf/bpf-next/c/1832696b2207
- [bpf-next,v6,2/4] selftests/bpf: Add tests for bpf_call_rcu()
https://git.kernel.org/bpf/bpf-next/c/908a60853b8d
- [bpf-next,v6,3/4] bpf: Add bpf_call_rcu_tasks_trace() kfunc
https://git.kernel.org/bpf/bpf-next/c/db1b20ed6b09
- [bpf-next,v6,4/4] selftests/bpf: Add a test for bpf_call_rcu_tasks_trace()
https://git.kernel.org/bpf/bpf-next/c/b22f65a009ad
You are awesome, thank you!
--
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html
^ permalink raw reply [flat|nested] 17+ messages in thread