Linux RCU subsystem development
 help / color / mirror / Atom feed
From: "Alexei Starovoitov" <alexei.starovoitov@gmail.com>
To: "Puranjay Mohan" <puranjay@kernel.org>, <bpf@vger.kernel.org>,
	<rcu@vger.kernel.org>
Cc: "Alexei Starovoitov" <ast@kernel.org>,
	"Daniel Borkmann" <daniel@iogearbox.net>,
	"Andrii Nakryiko" <andrii@kernel.org>,
	"Martin KaFai Lau" <martin.lau@linux.dev>,
	"Eduard Zingerman" <eddyz87@gmail.com>,
	"Kumar Kartikeya Dwivedi" <memxor@gmail.com>,
	"Song Liu" <song@kernel.org>,
	"Yonghong Song" <yonghong.song@linux.dev>,
	"Harry Yoo (Oracle)" <harry@kernel.org>,
	"Paul E. McKenney" <paulmck@kernel.org>
Subject: Re: [PATCH bpf-next v5 1/4] bpf: Add bpf_call_rcu() kfunc
Date: Tue, 22 Sep 2026 01:53:00 +0000	[thread overview]
Message-ID: <DLLGX9F9UG4Q.2YJWNSO4FXT1P@gmail.com> (raw)
In-Reply-To: <20260921191407.1742386-2-puranjay@kernel.org>

On Mon Sep 21, 2026 at 7:14 PM UTC, Puranjay Mohan wrote:
> BPF programs that manage their own objects have no way to run their own
> logic once an RCU grace period has elapsed.  bpf_obj_drop() defers a
> free, but returning an index to an allocator or unpinning a resource
> once readers are done has no equivalent.  sched_ext's BPF library works
> around this today by pushing freed nodes onto a list and having a
> userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a
> BPF program to reclaim them.
>
> Add:
>
> 	int bpf_call_rcu(struct bpf_rcu_head *rh, void *map,
> 			 int (*callback)(struct bpf_map *map, void *key,
> 					 void *value));
>
> @rh is a struct bpf_rcu_head embedded in a value of @map, so the
> callback runs as callback(map, key, value) for the element it lives in
> and needs no cookie.  A head can only be armed once, which bounds
> outstanding work by the number of elements.
>
> struct bpf_rcu_head holds the callback state inline rather than a
> pointer to it, as bpf_timer, bpf_wq and bpf_task_work do, because there
> is nothing to cancel and so nothing that has to outlive the map value.
> That avoids an allocation and a state machine on the arming path at the
> cost of 64 bytes per element, 48 of which are used today.  Embedding
> struct rcu_head ties part of a uapi struct to a definition outside of
> BPF, which is acceptable here only because it is two pointers, a
> callback and its argument, with no room to grow.
>
> An RCU callback cannot be cancelled, so everything it touches has to
> stay alive until it runs:
>
>   - The callback is the program's text, so arming takes a program
>     reference as bpf_timer, bpf_wq and bpf_task_work do, dropped once
>     the callback returns.  bpf_prog_inc_not_zero() also fails the arm
>     with -EBADF once the program is dying.
>
>   - The map is held by that reference through used_maps.  An inner map
>     is not, so bpf_rcu_head is rejected in one.
>
>   - The field is only accepted in BPF_MAP_TYPE_ARRAY, whose elements
>     are never freed individually.  A hash element can be deleted and
>     recycled while a callback is queued on it.
>
>   - The head is disarmed before the callback runs so it can be armed
>     again from there, which takes a new program reference before the
>     running callback drops its own.  Arming therefore fails with -EPERM
>     once the map is held by neither a process nor bpffs, which is what
>     bpf_timer and bpf_wq do at init time; bpf_task_work uses -EBUSY and
>     additionally cancels, which is not possible here.
>
> bpf_iter hands a program a writable pointer to the live element, which
> would let it overwrite a queued head, so bpf_iter_attach_map() rejects
> maps carrying one.
>
> The callback is verified non-sleepable even when the caller is
> sleepable, and RCU invokes it with BH disabled.
>
> Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
> ---
>  include/linux/bpf.h            | 10 +++++
>  include/uapi/linux/bpf.h       |  4 ++
>  kernel/bpf/btf.c               |  7 +++
>  kernel/bpf/helpers.c           | 75 +++++++++++++++++++++++++++++++
>  kernel/bpf/map_in_map.c        |  4 ++
>  kernel/bpf/map_iter.c          |  6 +++
>  kernel/bpf/syscall.c           | 11 ++++-
>  kernel/bpf/verifier.c          | 81 +++++++++++++++++++++++++++++++++-
>  tools/include/uapi/linux/bpf.h |  4 ++
>  9 files changed, 199 insertions(+), 3 deletions(-)
>
> diff --git a/include/linux/bpf.h b/include/linux/bpf.h
> index fd22db8bc6c50..e7c5e203edddb 100644
> --- a/include/linux/bpf.h
> +++ b/include/linux/bpf.h
> @@ -215,6 +215,7 @@ enum btf_field_type {
>  	BPF_UPTR       = (1 << 11),
>  	BPF_RES_SPIN_LOCK = (1 << 12),
>  	BPF_TASK_WORK  = (1 << 13),
> +	BPF_RCU_HEAD   = (1 << 14),
>  };
>  
>  enum bpf_cgroup_storage_type {
> @@ -269,6 +270,7 @@ struct btf_record {
>  	int wq_off;
>  	int refcount_off;
>  	int task_work_off;
> +	int rcu_head_off;
>  	struct btf_field fields[];
>  };
>  
> @@ -374,6 +376,8 @@ static inline const char *btf_field_type_name(enum btf_field_type type)
>  		return "bpf_refcount";
>  	case BPF_TASK_WORK:
>  		return "bpf_task_work";
> +	case BPF_RCU_HEAD:
> +		return "bpf_rcu_head";
>  	default:
>  		WARN_ON_ONCE(1);
>  		return "unknown";
> @@ -414,6 +418,8 @@ static inline u32 btf_field_type_size(enum btf_field_type type)
>  		return sizeof(struct bpf_refcount);
>  	case BPF_TASK_WORK:
>  		return sizeof(struct bpf_task_work);
> +	case BPF_RCU_HEAD:
> +		return sizeof(struct bpf_rcu_head);
>  	default:
>  		WARN_ON_ONCE(1);
>  		return 0;
> @@ -448,6 +454,8 @@ static inline u32 btf_field_type_align(enum btf_field_type type)
>  		return __alignof__(struct bpf_refcount);
>  	case BPF_TASK_WORK:
>  		return __alignof__(struct bpf_task_work);
> +	case BPF_RCU_HEAD:
> +		return __alignof__(struct bpf_rcu_head);
>  	default:
>  		WARN_ON_ONCE(1);
>  		return 0;
> @@ -480,6 +488,7 @@ static inline void bpf_obj_init_field(const struct btf_field *field, void *addr)
>  	case BPF_KPTR_PERCPU:
>  	case BPF_UPTR:
>  	case BPF_TASK_WORK:
> +	case BPF_RCU_HEAD:
>  		break;
>  	default:
>  		WARN_ON_ONCE(1);
> @@ -925,6 +934,7 @@ enum bpf_arg_type {
>  	ARG_PTR_TO_RB_NODE,	/* pointer to bpf_rb_node */
>  	ARG_PTR_TO_WORKQUEUE,	/* pointer to bpf_wq */
>  	ARG_PTR_TO_TASK_WORK,	/* pointer to bpf_task_work */
> +	ARG_PTR_TO_RCU_HEAD,	/* pointer to bpf_rcu_head */
>  	ARG_PTR_TO_IRQ_FLAG,	/* pointer to saved IRQ flags on the stack */
>  	ARG_PTR_TO_RES_SPIN_LOCK,	/* pointer to bpf_res_spin_lock */
>  	ARG_PTR_TO_CTX_OUT,	/* hook output argument passed through from ctx */
> diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
> index 6330b7d745c57..eafeba23f5da7 100644
> --- a/include/uapi/linux/bpf.h
> +++ b/include/uapi/linux/bpf.h
> @@ -7611,6 +7611,10 @@ struct bpf_task_work {
>  	__u64 __opaque;
>  } __attribute__((aligned(8)));
>  
> +struct bpf_rcu_head {
> +	__u64 __opaque[8];
> +} __attribute__((aligned(8)));
> +
>  struct bpf_wq {
>  	__u64 __opaque[2];
>  } __attribute__((aligned(8)));
> diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
> index 314ecb0e593b0..4a1fa4fbdf4e8 100644
> --- a/kernel/bpf/btf.c
> +++ b/kernel/bpf/btf.c
> @@ -3696,6 +3696,7 @@ static int btf_get_field_type(const struct btf *btf, const struct btf_type *var_
>  		{ BPF_TIMER, "bpf_timer", true },
>  		{ BPF_WORKQUEUE, "bpf_wq", true },
>  		{ BPF_TASK_WORK, "bpf_task_work", true },
> +		{ BPF_RCU_HEAD, "bpf_rcu_head", true },
>  		{ BPF_LIST_HEAD, "bpf_list_head", false },
>  		{ BPF_LIST_NODE, "bpf_list_node", false },
>  		{ BPF_RB_ROOT, "bpf_rb_root", false },
> @@ -3881,6 +3882,7 @@ static int btf_find_field_one(const struct btf *btf,
>  	case BPF_RB_NODE:
>  	case BPF_REFCOUNT:
>  	case BPF_TASK_WORK:
> +	case BPF_RCU_HEAD:
>  		ret = btf_find_struct(btf, var_type, off, sz, field_type,
>  				      info_cnt ? &info[0] : &tmp);
>  		if (ret < 0)
> @@ -4176,6 +4178,7 @@ struct btf_record *btf_parse_fields(const struct btf *btf, const struct btf_type
>  	rec->wq_off = -EINVAL;
>  	rec->refcount_off = -EINVAL;
>  	rec->task_work_off = -EINVAL;
> +	rec->rcu_head_off = -EINVAL;
>  	for (i = 0; i < cnt; i++) {
>  		field_type_size = btf_field_type_size(info_arr[i].type);
>  		if (info_arr[i].off + field_type_size > value_size) {
> @@ -4219,6 +4222,10 @@ struct btf_record *btf_parse_fields(const struct btf *btf, const struct btf_type
>  			WARN_ON_ONCE(rec->task_work_off >= 0);
>  			rec->task_work_off = rec->fields[i].offset;
>  			break;
> +		case BPF_RCU_HEAD:
> +			WARN_ON_ONCE(rec->rcu_head_off >= 0);
> +			rec->rcu_head_off = rec->fields[i].offset;
> +			break;
>  		case BPF_REFCOUNT:
>  			WARN_ON_ONCE(rec->refcount_off >= 0);
>  			/* Cache offset for faster lookup at runtime */
> diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
> index 501c7ce35cba9..8a01dd4058a03 100644
> --- a/kernel/bpf/helpers.c
> +++ b/kernel/bpf/helpers.c
> @@ -4805,6 +4805,80 @@ __bpf_kfunc int bpf_task_work_schedule_resume(struct task_struct *task, struct b
>  	return bpf_task_work_schedule(task, tw, map__const_map, callback, aux, TWA_RESUME);
>  }
>  
> +typedef int (*bpf_rcu_callback_t)(struct bpf_map *map, void *key, void *value);
> +
> +/* Actual type for struct bpf_rcu_head */
> +struct bpf_rcu_head_kern {
> +	struct rcu_head rcu;
> +	bpf_callback_t callback_fn;
> +	struct bpf_map *map;
> +	struct bpf_prog *prog;
> +	u32 armed;
> +} __aligned(8);
> +
> +static void bpf_rcu_run_callback(struct rcu_head *rcu)
> +{
> +	struct bpf_rcu_head_kern *rh = container_of(rcu, struct bpf_rcu_head_kern, rcu);
> +	bpf_callback_t callback_fn = rh->callback_fn;
> +	struct bpf_prog *prog = rh->prog;
> +	struct bpf_map *map = rh->map;
> +	void *value, *key;
> +	u32 idx;
> +
> +	value = (void *)rh - map->record->rcu_head_off;
> +	key = map_key_from_value(map, value, &idx);
> +
> +	/* Pairs with the arming cmpxchg(): rh may be re-armed as soon as this store lands. */
> +	smp_store_release(&rh->armed, 0);
> +
> +	rcu_read_lock_dont_migrate();
> +	callback_fn((u64)(long)map, (u64)(long)key, (u64)(long)value, 0, 0);
> +	rcu_read_unlock_migrate();
> +
> +	bpf_prog_put(prog);
> +}
> +
> +/**
> + * bpf_call_rcu - Invoke a BPF callback after an RCU grace period
> + * @rh: struct bpf_rcu_head in a BPF map value
> + * @map__const_map: bpf_map that embeds struct bpf_rcu_head in the values
> + * @callback: BPF subprogram, invoked as callback(map, key, value) for the value holding @rh
> + * @aux: bpf_prog_aux of the caller, implicitly set by the verifier
> + *
> + * Return: 0, -EBUSY if @rh is already queued, -EPERM if @map is held by neither a process
> + * nor bpffs, or -EBADF if the calling program is going away.
> + */
> +__bpf_kfunc int bpf_call_rcu(struct bpf_rcu_head *rh, void *map__const_map,
> +			     bpf_rcu_callback_t callback, struct bpf_prog_aux *aux)
> +{
> +	struct bpf_rcu_head_kern *rhk = (void *)rh;
> +	struct bpf_map *map = map__const_map;
> +	struct bpf_prog *prog;
> +
> +	BUILD_BUG_ON(sizeof(struct bpf_rcu_head_kern) > sizeof(struct bpf_rcu_head));
> +	BUILD_BUG_ON(__alignof__(struct bpf_rcu_head_kern) != __alignof__(struct bpf_rcu_head));
> +	BTF_TYPE_EMIT(struct bpf_rcu_head);
> +
> +	/* A queued callback cannot be cancelled, so a self-rearming one would pin prog and map. */
> +	if (!atomic64_read(&map->usercnt))
> +		return -EPERM;
> +
> +	if (cmpxchg(&rhk->armed, 0, 1))
> +		return -EBUSY;
> +
> +	prog = bpf_prog_inc_not_zero(aux->prog);
> +	if (IS_ERR(prog)) {
> +		WRITE_ONCE(rhk->armed, 0);
> +		return -EBADF;
> +	}

can we drop prog_inc and remove prog pointer from rhk ?
with something like if (prog->has_call_rcu) rcu_barrier() during prog unload?


  parent reply	other threads:[~2026-09-22  1:53 UTC|newest]

Thread overview: 13+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-21 19:14 [PATCH bpf-next v5 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace() Puranjay Mohan
2026-09-21 19:14 ` [PATCH bpf-next v5 1/4] bpf: Add bpf_call_rcu() kfunc Puranjay Mohan
2026-09-21 20:33   ` bot+bpf-ci
2026-09-22  1:53   ` Alexei Starovoitov [this message]
2026-09-22 14:17     ` Puranjay Mohan
2026-09-22 18:35       ` Alexei Starovoitov
2026-09-22 19:09         ` Puranjay Mohan
2026-09-22 23:55           ` Paul E. McKenney
2026-09-21 19:14 ` [PATCH bpf-next v5 2/4] selftests/bpf: Add tests for bpf_call_rcu() Puranjay Mohan
2026-09-21 19:14 ` [PATCH bpf-next v5 3/4] bpf: Add bpf_call_rcu_tasks_trace() kfunc Puranjay Mohan
2026-09-21 20:18   ` bot+bpf-ci
2026-09-21 20:21     ` Puranjay Mohan
2026-09-21 19:14 ` [PATCH bpf-next v5 4/4] selftests/bpf: Add a test for bpf_call_rcu_tasks_trace() Puranjay Mohan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DLLGX9F9UG4Q.2YJWNSO4FXT1P@gmail.com \
    --to=alexei.starovoitov@gmail.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=harry@kernel.org \
    --cc=martin.lau@linux.dev \
    --cc=memxor@gmail.com \
    --cc=paulmck@kernel.org \
    --cc=puranjay@kernel.org \
    --cc=rcu@vger.kernel.org \
    --cc=song@kernel.org \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox