BPF List
 help / color / mirror / Atom feed
From: Yonghong Song <yonghong.song@linux.dev>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
	Andrii Nakryiko <andrii@kernel.org>,
	Daniel Borkmann <daniel@iogearbox.net>,
	Eduard Zingerman <eddyz87@gmail.com>,
	kernel-team@fb.com
Subject: [PATCH bpf-next v9 11/23] bpf: Dispatch cleanup pads by rewriting return addresses
Date: Thu,  8 Oct 2026 00:50:55 -0700	[thread overview]
Message-ID: <20261008075055.3000922-1-yonghong.song@linux.dev> (raw)
In-Reply-To: <20261008074959.2993751-1-yonghong.song@linux.dev>

bpf_unwind() now walks the BPF frames through arch_bpf_stack_walk_ra()
and rewrites the saved return address of each frame above the one that
called it: to the pad where a record covers the call, else to the
frame's epilogue. The walk stops after the main function. For the
verifier patch's example, main -> A -> B -> C with only A's call to B
covered:

  slot                    points into               rewritten to
  ----------------------  ------------------------  ---------------------
  bpf_unwind()'s return   C, after its unwind call  left alone
  C's return              B, after 'call C'         B's epilogue
  B's return              A, after 'call B'         P, A's pad
  A's return              main, after 'call A'      main's epilogue
  main's return           the kernel                left alone

Each frame then just returns: C through the 'r0 = 0; exit' after its
bpf_unwind(), B and main through their epilogues, A through its pad P
and P's resume.

The walk finds the frames in the calling program, not with
bpf_prog_ksym_find(). A running instance holds no reference to its
program, so user space can drop the last one, by closing the program's
fds and links, while the instance still runs: the kallsyms entries go at
once, and only freeing the program waits for the RCU or RCU Tasks Trace
grace period. An unwind in between, say in a sleepable program blocked
in bpf_copy_from_user(), would find none of its frames, so no pad would
run. So bpf_unwind() takes the program's aux as a KF_IMPLICIT_ARGS
argument, which keeps its BTF prototype void(void), and matches each
return address against aux->func[], or the program itself. To get the
implicit argument loaded, adjust_insn_aux_data() moves arg_prog, as it
moves cleanup_pad, to the slot the original instruction kept: the unwind
call stays first in its patch.

Both kfuncs become callable here. bpf_unwind() is notrace and NOKPROBE:
a tracer hooking its return would leave a trampoline in the slot of the
frame that called it, which the walk cannot rewrite.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 kernel/bpf/fixups.c  |  2 ++
 kernel/bpf/helpers.c | 67 +++++++++++++++++++++++++++++++++++++++++++-
 2 files changed, 68 insertions(+), 1 deletion(-)

diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 1cc025c45ecd..c29e14ffc475 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -247,6 +247,8 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
 			data[off + cnt - 1].non_stack_access = false;
 			data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
 			data[off + cnt - 1].cleanup_pad = 0;
+			data[i].arg_prog = data[off + cnt - 1].arg_prog;
+			data[off + cnt - 1].arg_prog = 0;
 		} else if (bpf_is_mem_insn(insn + i)) {
 			data[i].non_stack_access = true;
 		}
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 4eccd6742eba..c9efb789f36b 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -29,8 +29,10 @@
 #include <linux/task_work.h>
 #include <linux/irq_work.h>
 #include <linux/buildid.h>
+#include <linux/kprobes.h>
 
 #include "../../lib/kstrtox.h"
+#include "exception.h"
 
 /* If kernel subsystem is allowing eBPF programs to call this function,
  * inside its own verifier_ops->get_func_proto() callback it should return
@@ -3424,9 +3426,70 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
 	return false;
 }
 
-__bpf_kfunc void bpf_unwind(void)
+struct bpf_unwind_ctx {
+	const struct bpf_prog_aux *aux;
+	u32 cnt;
+};
+
+static bool bpf_unwind_ip_in(const struct bpf_prog *prog, u64 ip)
+{
+	u64 start = (u64)(long)prog->bpf_func;
+
+	return ip > start && ip <= start + prog->jited_len;
+}
+
+/*
+ * The function of the calling program that @ip returns into, looked up in the
+ * program itself: its kallsyms entries can go while an instance still runs.
+ * The main function is the outer program, which holds its table and epilogue.
+ */
+static struct bpf_prog *bpf_unwind_find_prog(const struct bpf_prog_aux *aux, u64 ip)
+{
+	u32 i;
+
+	for (i = 1; i < aux->func_cnt; i++)
+		if (bpf_unwind_ip_in(aux->func[i], ip))
+			return aux->func[i];
+	return bpf_unwind_ip_in(aux->prog, ip) ? aux->prog : NULL;
+}
+
+static bool bpf_unwind_rewrite(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra)
+{
+	const struct bpf_cleanup_range *rec;
+	struct bpf_unwind_ctx *ctx = cookie;
+	struct bpf_prog *prog;
+
+	prog = bpf_unwind_find_prog(ctx->aux, ip);
+	if (!prog)
+		return !ctx->cnt;
+	ctx->cnt++;
+
+	/*
+	 * The frame that called bpf_unwind(): bpf_exc_patch_unwind_calls()
+	 * put 'r0 = 0' and a jump to its pad, or an exit, after the call,
+	 * so leave its return address alone and let it go on there. The pad
+	 * then starts with r0 at a known zero.
+	 */
+	if (ctx->cnt == 1)
+		return bpf_is_subprog(prog);
+
+	rec = bpf_exc_pad_for_ip(prog, ip);
+	*ra = rec ? rec->pad : prog->aux->epilogue_ip;
+
+	return bpf_is_subprog(prog);
+}
+
+/*
+ * @aux is the calling program's, supplied by the verifier (KF_IMPLICIT_ARGS):
+ * programs call bpf_unwind() with no arguments.
+ */
+__bpf_kfunc notrace void bpf_unwind(struct bpf_prog_aux *aux)
 {
+	struct bpf_unwind_ctx ctx = { .aux = aux };
+
+	arch_bpf_stack_walk_ra(bpf_unwind_rewrite, &ctx);
 }
+NOKPROBE_SYMBOL(bpf_unwind);
 
 __bpf_kfunc void bpf_throw(u64 cookie)
 {
@@ -5095,6 +5158,8 @@ BTF_ID_FLAGS(func, bpf_task_get_cgroup1, KF_ACQUIRE | KF_RCU | KF_RET_NULL)
 BTF_ID_FLAGS(func, bpf_task_from_pid, KF_ACQUIRE | KF_RET_NULL)
 BTF_ID_FLAGS(func, bpf_task_from_vpid, KF_ACQUIRE | KF_RET_NULL)
 BTF_ID_FLAGS(func, bpf_throw)
+BTF_ID_FLAGS(func, bpf_unwind, KF_IMPLICIT_ARGS)
+BTF_ID_FLAGS(func, bpf_unwind_resume)
 #ifdef CONFIG_BPF_EVENTS
 BTF_ID_FLAGS(func, bpf_send_signal_task)
 #endif
-- 
2.53.0-Meta


  parent reply	other threads:[~2026-10-08  7:50 UTC|newest]

Thread overview: 38+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-08  7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-10-08  7:50 ` [PATCH bpf-next v9 01/23] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-10-08  7:50 ` [PATCH bpf-next v9 02/23] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-10-08  7:50 ` [PATCH bpf-next v9 03/23] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-10-08  7:50 ` [PATCH bpf-next v9 04/23] bpf: Keep a call site's landing pad in insn_aux_data, add lookups Yonghong Song
2026-10-08  8:01   ` sashiko-bot
2026-10-08 15:58     ` Yonghong Song
2026-10-08  7:50 ` [PATCH bpf-next v9 05/23] bpf: Mark covered call sites and check a program can take a table Yonghong Song
2026-10-08  7:50 ` [PATCH bpf-next v9 06/23] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-10-08  7:50 ` [PATCH bpf-next v9 07/23] bpf: Verify an unwind through landing pads and epilogues Yonghong Song
2026-10-08  8:57   ` bot+bpf-ci
2026-10-08 16:07     ` Yonghong Song
2026-10-08  7:50 ` [PATCH bpf-next v9 08/23] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-10-08  8:57   ` bot+bpf-ci
2026-10-08 16:11     ` Yonghong Song
2026-10-08  7:50 ` [PATCH bpf-next v9 09/23] bpf: Do not use a private stack for a program that can unwind Yonghong Song
2026-10-08  7:50 ` [PATCH bpf-next v9 10/23] bpf: Prepare JITed programs for dispatching cleanup pads Yonghong Song
2026-10-08  7:50 ` Yonghong Song [this message]
2026-10-08  7:51 ` [PATCH bpf-next v9 12/23] bpf: Refuse a trampoline that calls a subprog that can unwind Yonghong Song
2026-10-08  8:14   ` sashiko-bot
2026-10-08 16:19     ` Yonghong Song
2026-10-08  7:51 ` [PATCH bpf-next v9 13/23] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-10-08  7:51 ` [PATCH bpf-next v9 14/23] bpf, arm64: " Yonghong Song
2026-10-08  8:39   ` bot+bpf-ci
2026-10-08 16:23     ` Yonghong Song
2026-10-08  7:51 ` [PATCH bpf-next v9 15/23] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-10-08  7:51 ` [PATCH bpf-next v9 16/23] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-10-08  8:12   ` sashiko-bot
2026-10-08 16:25     ` Yonghong Song
2026-10-08  7:51 ` [PATCH bpf-next v9 17/23] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-10-08  7:51 ` [PATCH bpf-next v9 18/23] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-10-08  7:51 ` [PATCH bpf-next v9 19/23] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-10-08  8:14   ` sashiko-bot
2026-10-08 16:26     ` Yonghong Song
2026-10-08  7:51 ` [PATCH bpf-next v9 20/23] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
2026-10-08  7:51 ` [PATCH bpf-next v9 21/23] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-10-08  7:51 ` [PATCH bpf-next v9 22/23] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
2026-10-08  7:51 ` [PATCH bpf-next v9 23/23] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261008075055.3000922-1-yonghong.song@linux.dev \
    --to=yonghong.song@linux.dev \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=kernel-team@fb.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox