From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 65FBB49A3C0; Tue, 22 Sep 2026 20:02:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790107352; cv=none; b=r5zxEZjRrk3rPLb/urMSK/K5P9OAPgPsRtXOU4R94c87KtL617bxYoINs0NXNcBM3PZ8E6FpBdR6Aq5Wjp336EbW5ktn7bZz5hOZ6ukCNOIy8U0RUF8f6Rpv/BVeQ1AUmKtowNWqYgtJaxoWJjn+4+CnlboZbTP/1j2/1BToHU4= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790107352; c=relaxed/simple; bh=w1/MEfvJY22TL1xiLJoSoCTeMJrmmlF3AaJ/bhvjcrw=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=eB8PtcoGUJRx1pNloQiBIAazLdsCoUlR8nKp8iXXm+zvBzHW6wrXEGL85i0D0vkXx3QTbsSm6TXy9LB6pRFomHHRNtqHaXDuqTrgVnk+7ILSgdythWYMeW5fNANSmAFx0VYy/ANjzjEX0+iCEvYwoBSj/YZJHEpI/hoqBtsXBow= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=KU4XW8JB; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="KU4XW8JB" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 364E21F00893; Tue, 22 Sep 2026 20:02:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790107342; bh=3nKxEyGveQKflxHQJBNJkF4hp1otR2ToGR+UzLpsrW0=; h=From:To:Cc:Subject:Date; b=KU4XW8JBHJmK2noz3qXTyStTmOFuGkxtb6cuhIWJtIUeS65HONqIsmxUZlKcldGns ubsFxXo7CsK2MgpQoB3xWa1iVhH3KrwKSvZTZh8IvVDh+sOK1NJqhRHBuCbvhKhBkI DnXfjOhoEypdHUkby1eMNDR5Ls1CEy2Ttpnakrc5rQohiAkugH4hdfgUSKWwd2nHWI 0LN5FfObAK37dgFQnP8oDr9gi2Ko9kAbANuKhsysu3nKAffilXa0SkYb2VYP7QUmqX bfxZ8E0+C0rMY6M6C9H8L7AB6ROy+/QG6TOuCUZeWtXuT/2DOm+aDgh2buzmaI7CRa IJE6PYIkbfrVg== From: Puranjay Mohan To: bpf@vger.kernel.org, rcu@vger.kernel.org Cc: Puranjay Mohan , "Alexei Starovoitov" , "Daniel Borkmann" , "Andrii Nakryiko" , "Martin KaFai Lau" , "Eduard Zingerman" , "Kumar Kartikeya Dwivedi" , "Song Liu" , "Yonghong Song" , "Harry Yoo (Oracle)" , "Paul E. McKenney" Subject: [PATCH bpf-next v6 0/4] bpf: Add bpf_call_rcu() and bpf_call_rcu_tasks_trace() Date: Tue, 22 Sep 2026 13:00:49 -0700 Message-ID: <20260922200208.3203834-1-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 Precedence: bulk X-Mailing-List: rcu@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Changelog: v5: https://lore.kernel.org/all/20260921191407.1742386-1-puranjay@kernel.org/ Changes in v6: - Size struct bpf_rcu_head at 48 bytes rather than 64, matching what the inline state uses (Alexei) - Trim patch 1's changelog: drop the reasoning about possible future layouts and about what the other async kfuncs return - Rebase on bpf-next/master v4: https://lore.kernel.org/all/20260915154248.3612028-1-puranjay@kernel.org/ Changes in v5: - Rename the "hash_map" subtest to "bad_map" so that it matches its helper test_call_rcu_bad_map() (bpf-ci) - Bump the callback counter after chain_err in the selftest callback. Userspace polls that counter and then reads chain_err, so it could still see the initial value before the re-arm had stored one, which let the chain subtest's assertion pass without checking anything - Say in patch 1 why embedding struct rcu_head in a uapi struct is acceptable here (Mykyta, Paul, Alexei) - Rebase on bpf-next/master v3: https://lore.kernel.org/all/20260915143640.36292-1-puranjay@kernel.org/ Changes in v4: - Drop an unrelated hunk in bpf_async_update_prog_callback() that turned PTR_ERR(prog) into -EBADF. That is the shared bpf_timer/bpf_wq path and bpf_prog_inc_not_zero() returns -ENOENT, so it would have changed the errno bpf_timer_set_callback() and bpf_wq_set_callback() report to userspace (Sashiko). No other changes from v3 v2: https://lore.kernel.org/all/20260915114240.3269184-1-puranjay@kernel.org/ Changes in v3: - Move the bpf_call_rcu_tasks_trace() verifier bits from patch 1 to patch 3; patch 1 alone emitted "resolve_btfids: unresolved symbol bpf_call_rcu_tasks_trace" (Sashiko) - Poll the callback counter with an acquire load (Sashiko) - teardown: v2 only checked that the program was eventually freed, which passes even if nothing was ever armed. Also assert that the chain ran, and read the -EPERM back through an independent .bss fd - Use kern_sync_rcu() instead of open coding the grace-period wait - Return -EBADF rather than -ENOENT when the calling program is going away, matching bpf_task_work_schedule() - mismatch_map now pins the bpf_rcu_head label the new code emits; it passed with that branch removed. Add a two_heads test, and wait for each grace period separately in the chain test - Commit messages: correct the -EPERM parity claim, explain the inline callback state and the struct size, motivate the tasks trace flavour v1: https://lore.kernel.org/all/20260907134552.1772405-1-puranjay@kernel.org/ Changes in v2: - Rebase on bpf-next/master - Use rcu_read_lock_dont_migrate() over open coding (Alexei) - Improve re-arming selftest to detect failure (Sashiko) BPF programs that manage their own objects have no way to run their own logic once an RCU grace period has elapsed. bpf_obj_drop() defers a free, but returning an index to an allocator or unpinning a resource once readers are done has no equivalent. sched_ext's BPF library works around this today by pushing freed nodes onto a list and having a userspace thread call membarrier(MEMBARRIER_CMD_GLOBAL) and then run a BPF program to reclaim them; it is the first intended user. Add: int bpf_call_rcu(struct bpf_rcu_head *rh, void *map, int (*callback)(struct bpf_map *map, void *key, void *value)); and bpf_call_rcu_tasks_trace(), same signature, which also waits for sleepable programs. @rh is a struct bpf_rcu_head embedded in a value of @map, so the callback runs as callback(map, key, value) for the element it lives in and needs no cookie. The field is only accepted in BPF_MAP_TYPE_ARRAY, and arming holds a reference on the calling program until the callback has run. Patch 1 covers the lifetime rules. This needs https://lore.kernel.org/all/20260810122758.183765-1-puranjay@kernel.org/ for call_rcu() and call_srcu() to be safe from the contexts a BPF program can be called in. Puranjay Mohan (4): bpf: Add bpf_call_rcu() kfunc selftests/bpf: Add tests for bpf_call_rcu() bpf: Add bpf_call_rcu_tasks_trace() kfunc selftests/bpf: Add a test for bpf_call_rcu_tasks_trace() include/linux/bpf.h | 10 + include/uapi/linux/bpf.h | 4 + kernel/bpf/btf.c | 7 + kernel/bpf/helpers.c | 102 +++++++ kernel/bpf/map_in_map.c | 4 + kernel/bpf/map_iter.c | 6 + kernel/bpf/syscall.c | 11 +- kernel/bpf/verifier.c | 84 +++++- tools/include/uapi/linux/bpf.h | 4 + .../selftests/bpf/prog_tests/call_rcu.c | 275 ++++++++++++++++++ tools/testing/selftests/bpf/progs/call_rcu.c | 110 +++++++ .../selftests/bpf/progs/call_rcu_fail.c | 114 ++++++++ 12 files changed, 728 insertions(+), 3 deletions(-) create mode 100644 tools/testing/selftests/bpf/prog_tests/call_rcu.c create mode 100644 tools/testing/selftests/bpf/progs/call_rcu.c create mode 100644 tools/testing/selftests/bpf/progs/call_rcu_fail.c base-commit: 79dc258c9392051420a26f1504c647bd3d27c66a -- 2.53.0-Meta