BPF List
 help / color / mirror / Atom feed
From: "Alexei Starovoitov" <alexei.starovoitov@gmail.com>
To: "Puranjay Mohan" <puranjay@kernel.org>, <bpf@vger.kernel.org>
Cc: "Alexei Starovoitov" <ast@kernel.org>,
	"Daniel Borkmann" <daniel@iogearbox.net>,
	"Andrii Nakryiko" <andrii@kernel.org>,
	"Martin KaFai Lau" <martin.lau@linux.dev>,
	"Eduard Zingerman" <eddyz87@gmail.com>,
	"Kumar Kartikeya Dwivedi" <memxor@gmail.com>,
	"Song Liu" <song@kernel.org>,
	"Yonghong Song" <yonghong.song@linux.dev>, <tj@kernel.org>
Subject: Re: [PATCH bpf-next 1/2] bpf: Support bpf_rcu_head in hash and LRU hash maps
Date: Fri, 25 Sep 2026 01:05:05 +0000	[thread overview]
Message-ID: <DLNZS7DWLRGE.TPVKX87N0IL9@gmail.com> (raw)
In-Reply-To: <20260924162858.2435106-2-puranjay@kernel.org>

On Thu Sep 24, 2026 at 4:28 PM UTC, Puranjay Mohan wrote:
> bpf_call_rcu() is restricted to arrays because an array element is never
> freed while the map is alive. Hash elements are recycled, and an RCU
> callback cannot be cancelled, so a delete racing a queued callback would
> hand the element back to the allocator underneath it.
>
> Keep the element alive until the callback has run. Before releasing an
> element the map calls bpf_rcu_head_claim(), which marks the head dead and
> reports whether a callback is queued or running. If one is, the element is
> unlinked but not returned to the allocator; the callback does that through
> a new map_release_elem(), since only the map knows whether that is a
> freelist push or bpf_mem_cache_free(). The dead bit stops it being armed
> again, so there is a definite last callback. Timers and friends in the
> same value are still cancelled on delete.
>
> The head and the element are busy for different spans: RCU dequeues a head
> before invoking it, so it can be re-armed from inside the callback, while
> the element has to live until the callback returns. ARMED covers the head,
> RUNNING the element, and whoever clears the last of the two releases a dead
> element. RUNNING is dropped under the read locks, so only one callback runs
> on a head at a time.
>
> The dead bit lives in the element, so alloc_htab_elem() and
> prealloc_lru_pop() reset the head when they hand one out.
>
> A preallocated htab stashes the old element in a per-CPU spare on update
> rather than freeing it, which cannot be done to one with a queued callback,
> so maps carrying a head take the freelist path instead.
>
> For LRU the eviction path asks bpf_rcu_head_busy() and declines, leaving
> the element in the map.
>
> Signed-off-by: Puranjay Mohan <puranjay@kernel.org>
> ---
>  include/linux/bpf.h  |   5 ++
>  kernel/bpf/hashtab.c | 104 +++++++++++++++++++++++++++++++++-----
>  kernel/bpf/helpers.c | 117 ++++++++++++++++++++++++++++++++++++++-----
>  kernel/bpf/syscall.c |  14 +++++-
>  4 files changed, 214 insertions(+), 26 deletions(-)
>
> diff --git a/include/linux/bpf.h b/include/linux/bpf.h
> index 1e1ce2afe2ed8..91eee1d066a32 100644
> --- a/include/linux/bpf.h
> +++ b/include/linux/bpf.h
> @@ -111,6 +111,8 @@ struct bpf_map_ops {
>  	void *(*map_lookup_elem)(struct bpf_map *map, void *key);
>  	long (*map_update_elem)(struct bpf_map *map, void *key, void *value, u64 flags);
>  	long (*map_delete_elem)(struct bpf_map *map, void *key);
> +	/* Release an element a bpf_rcu_head callback was holding. */
> +	void (*map_release_elem)(struct bpf_map *map, void *value);

This is quite heavy. As you can see both bots poked plenty of holes.
I feel this will be nightmarish to support moving forward.

I'd like to hear from Tejun whether he thinks that call_rcu_only_in_array_map
is a limitation for what he wanted to do or not.

Instead of supporting in hash map we can support them in bpf_mem_alloced
objects instead. Same flexibility. A lot less pain.

pw-bot: cr

  parent reply	other threads:[~2026-09-25  1:05 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-24 16:28 [PATCH bpf-next 0/2] bpf: Support bpf_rcu_head in hash and LRU hash maps Puranjay Mohan
2026-09-24 16:28 ` [PATCH bpf-next 1/2] " Puranjay Mohan
2026-09-24 16:52   ` sashiko-bot
2026-09-24 17:24   ` bot+bpf-ci
2026-09-25  1:05   ` Alexei Starovoitov [this message]
2026-09-24 16:28 ` [PATCH bpf-next 2/2] selftests/bpf: Add bpf_call_rcu tests for " Puranjay Mohan
2026-09-24 17:23   ` bot+bpf-ci

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DLNZS7DWLRGE.TPVKX87N0IL9@gmail.com \
    --to=alexei.starovoitov@gmail.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=martin.lau@linux.dev \
    --cc=memxor@gmail.com \
    --cc=puranjay@kernel.org \
    --cc=song@kernel.org \
    --cc=tj@kernel.org \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox