All of lore.kernel.org
 help / color / mirror / Atom feed
From: Yonghong Song <yonghong.song@linux.dev>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
	Andrii Nakryiko <andrii@kernel.org>,
	Daniel Borkmann <daniel@iogearbox.net>,
	Eduard Zingerman <eddyz87@gmail.com>,
	kernel-team@fb.com
Subject: [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses
Date: Thu,  1 Oct 2026 06:31:03 -0700	[thread overview]
Message-ID: <20261001133103.1340994-1-yonghong.song@linux.dev> (raw)
In-Reply-To: <20261001133006.1335369-1-yonghong.song@linux.dev>

bpf_unwind() walks the BPF frames and, for each one above the frame that
called it, rewrites the saved return address so the frame resumes where
the unwind needs it, then returns. The frame that called it goes on after
the call, where the fixups put 'r0 = 0' and then a jump to its pad, or an
exit where no record covers the call, so the pad starts with r0 at a known
zero rather than whatever bpf_unwind() left in the return register.
bpf_unwind() restores no register itself: every frame runs its own
epilogue on the way out, which is what puts its caller's r6-r9 back, so
the unwind needs no spill area and no per-frame metadata beyond the table
itself.

Where a record covers the call a frame is suspended at, it resumes at that
pad, which an earlier patch made sure ends in a resume. Where none does,
it resumes at the frame's epilogue and returns at once, its caller reached
with registers already restored. A pad's resume lowers to 'r0 = 0; exit',
so that frame returns too and the rewritten address carries the unwind on
to the next pad.

For the verifier patch's example, main -> A -> B -> C with only A's call
to B covered, bpf_unwind() in C rewrites the return addresses it finds:

  slot                    points into               rewritten to
  ----------------------  ------------------------  ---------------------
  bpf_unwind()'s return   C, after its unwind call  left alone
  C's return              B, after 'call C'         B's epilogue
  B's return              A, after 'call B'         P, A's pad
  A's return              main, after 'call A'      main's epilogue
  main's return           the kernel                left alone

Each address is looked up in the frame it points into: a record over that
call gives its pad, none gives that frame's epilogue. Then every frame
just returns:

  frame   runs
  -----   ----------------------------------------------------------
  C       'r0 = 0; exit', patched in after its bpf_unwind()
  B       its epilogue, putting back A's r6-r9
  A       P, which drops A's resources; its resume is 'r0 = 0; exit'
  main    its epilogue, returning 0 to the kernel

An epilogue therefore has to exist for every frame the walk can pass, not
only for those carrying a table: the JITs record aux->epilogue_ip for every
program in the patches that follow, and this one hands it to the outer
program with the table when jit_subprogs() compiles the main program as
func[0].

x86 emits the epilogue at a subprogram's first exit, so one the dead code
sweep leaves exitless gets none. Two shapes do that. The frame that called
bpf_unwind() loses the code after the call; the exit patched back in there
where no record covers the call is also what that frame returns through.
And a frame above one that never comes back loses its exit with
no bpf_unwind() to hang a new one on, so the last exit of every subprogram
an unwind can pass through is kept, searched for since a subprogram may
end in a jump or a gotox. One with no exit at all is refused: nothing is
left an epilogue could be emitted at. arm64 emits an epilogue either way.

Both kfuncs become callable here rather than earlier: until the walk and
the lowering exist, bpf_unwind() would return to instructions the verifier
never explored and bpf_unwind_resume() would reach its WARN_ONCE body.

arch_bpf_stack_walk_ra() hands out the return-address slot as well as the
address. It is a second entry point rather than a change to
arch_bpf_stack_walk(), so architectures that do not dispatch pads keep the
walker they have -- where it is the weak stub, the walk does nothing, so
process_bpf_unwind() now asks bpf_exc_check_prog() whether this program may
unwind at all. A cleanup table was held to that before the CFG walk; a
bpf_unwind() with no table had not been.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf.h          |  37 ++++++++++
 include/linux/bpf_verifier.h |   1 +
 include/linux/filter.h       |   2 +
 kernel/bpf/core.c            |  20 ++++-
 kernel/bpf/exception.c       |  86 ++++++++++++++++++++++
 kernel/bpf/exception.h       |   6 ++
 kernel/bpf/fixups.c          | 138 +++++++++++++++++++++++++++++++++++
 kernel/bpf/helpers.c         |  45 ++++++++++++
 kernel/bpf/verifier.c        |   7 ++
 9 files changed, 341 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 4bae3796c42f..94005cd3ad0f 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -1805,6 +1805,41 @@ enum bpf_sig_keyring {
 	BPF_SIG_KEYRING_BPF,
 };
 
+/* One cleanup region of a JITed (sub)program. */
+struct bpf_cleanup_range {
+	u64 begin;
+	u64 end;
+	u64 pad;
+};
+
+struct bpf_exception_info {
+	struct bpf_cleanup_info *info;
+	struct bpf_cleanup_range *ranges;
+	u32 nr_info;
+	u32 nr_ranges;
+};
+
+#ifdef CONFIG_BPF_SYSCALL
+int bpf_exc_attach_main_prog(struct bpf_verifier_env *env, struct bpf_prog *prog);
+void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image);
+void bpf_exc_free_info(struct bpf_prog_aux *aux);
+#else
+
+static inline int bpf_exc_attach_main_prog(struct bpf_verifier_env *env,
+					   struct bpf_prog *prog)
+{
+	return 0;
+}
+
+static inline void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+}
+
+static inline void bpf_exc_free_info(struct bpf_prog_aux *aux)
+{
+}
+#endif
+
 struct bpf_prog_aux {
 	atomic64_t refcnt;
 	u32 used_map_cnt;
@@ -1885,6 +1920,8 @@ struct bpf_prog_aux {
 	u64 (*bpf_exception_cb)(u64 cookie, u64 sp, u64 bp, u64, u64);
 	u16 stack_arg_sp_adjust;
 	u16 freplace_link_cnt; /* counts freplace links extending this prog */
+	struct bpf_exception_info *exc;
+	u64 epilogue_ip; /* native address of this (sub)program's epilogue */
 #ifdef CONFIG_SECURITY
 	void *security;
 #endif
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index ccac422737fb..7cede13f8bee 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1863,6 +1863,7 @@ int bpf_opt_subreg_zext_lo32_rnd_hi32(struct bpf_verifier_env *env, const union
 int bpf_convert_ctx_accesses(struct bpf_verifier_env *env);
 int bpf_jit_subprogs(struct bpf_verifier_env *env);
 int bpf_fixup_call_args(struct bpf_verifier_env *env);
+int bpf_exc_patch_unwind_calls(struct bpf_verifier_env *env);
 int bpf_do_misc_fixups(struct bpf_verifier_env *env);
 int bpf_insn_def32(struct bpf_prog *prog, struct bpf_insn *insn);
 
diff --git a/include/linux/filter.h b/include/linux/filter.h
index 972b3ed2a51d..0d7d949a1baa 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1290,6 +1290,8 @@ u32 bpf_jit_plan_arg_moves(const struct bpf_jit_arg_abi *abi,
 			   struct bpf_jit_arg_move *moves);
 u64 bpf_arch_uaddress_limit(void);
 void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp), void *cookie);
+void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
+			    void *cookie);
 u64 arch_bpf_timed_may_goto(void);
 u64 bpf_check_timed_may_goto(struct bpf_timed_may_goto *);
 bool bpf_helper_changes_pkt_data(enum bpf_func_id func_id);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index d813fdde29e3..60905643cb9c 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -292,6 +292,7 @@ void __bpf_prog_free(struct bpf_prog *fp)
 		mutex_destroy(&fp->aux->dst_mutex);
 		mutex_destroy(&fp->aux->st_ops_assoc_mutex);
 		kfree(fp->aux->poke_tab);
+		bpf_exc_free_info(fp->aux);
 		kfree(fp->aux);
 	}
 	free_percpu(fp->stats);
@@ -2632,9 +2633,14 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
 {
 #ifdef CONFIG_BPF_JIT
 	struct bpf_prog *orig_prog;
+	int ret;
 
-	if (!bpf_prog_need_blind(prog))
+	if (!bpf_prog_need_blind(prog)) {
+		ret = bpf_exc_attach_main_prog(env, prog);
+		if (ret)
+			return ERR_PTR(ret);
 		return bpf_int_jit_compile(env, prog);
+	}
 
 	orig_prog = prog;
 	prog = bpf_jit_blind_constants(env, prog);
@@ -2648,6 +2654,12 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
 		goto out_restore;
 	}
 
+	ret = bpf_exc_attach_main_prog(env, prog);
+	if (ret) {
+		bpf_jit_prog_release_other(orig_prog, prog);
+		return ERR_PTR(ret);
+	}
+
 	prog = bpf_int_jit_compile(env, prog);
 	if (prog->jited) {
 		bpf_jit_prog_release_other(prog, orig_prog);
@@ -3511,6 +3523,12 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
 {
 }
 
+void __weak arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp,
+						      u64 *ra),
+				   void *cookie)
+{
+}
+
 bool __weak bpf_jit_supports_cleanup_pads(void)
 {
 	return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index e824883d3981..bbe64887e069 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -402,3 +402,89 @@ int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
 
 	return pad ? (int)pad - 1 : -1;
 }
+
+/*
+ * The record covering @ip, which is a return address: the call it belongs to
+ * is the instruction before it, so a range matches on begin < ip <= end.
+ */
+const struct bpf_cleanup_range *bpf_exc_pad_for_ip(const struct bpf_prog *prog, u64 ip)
+{
+	const struct bpf_exception_info *exc = prog->aux->exc;
+	u32 l = 0, r = exc ? exc->nr_ranges : 0;
+
+	while (l < r) {
+		u32 m = l + (r - l) / 2;
+		const struct bpf_cleanup_range *rec = &exc->ranges[m];
+
+		if (ip <= rec->begin)
+			r = m;
+		else if (ip > rec->end)
+			l = m + 1;
+		else
+			return rec;
+	}
+	return NULL;
+}
+
+int bpf_exc_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt)
+{
+	struct bpf_cleanup_range *ranges;
+	struct bpf_exception_info *exc;
+
+	exc = kzalloc_obj(struct bpf_exception_info, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+	ranges = kvcalloc(cnt, sizeof(*ranges), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+	if (!exc || !ranges) {
+		kfree(exc);
+		kvfree(ranges);
+		kvfree(recs);
+		return -ENOMEM;
+	}
+
+	exc->info = recs;
+	exc->nr_info = cnt;
+	exc->ranges = ranges;
+	/* Withheld until the JIT has filled the table in. */
+	exc->nr_ranges = 0;
+	aux->exc = exc;
+	return 0;
+}
+
+void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+	struct bpf_exception_info *exc = prog->aux->exc;
+	u32 i, n;
+
+	if (!exc)
+		return;
+
+	n = exc->nr_info;
+	for (i = 0; i < n; i++) {
+		const struct bpf_cleanup_info *rec = &exc->info[i];
+
+		/*
+		 * exc_info_for_subprog() built the records from insn_aux_data
+		 * inside this subprog, so this cannot fire; if it does, no
+		 * pad is dispatched rather than one read past addrs[].
+		 */
+		if (WARN_ON_ONCE(rec->begin_off >= prog->len ||
+				 rec->end_off > prog->len ||
+				 rec->landing_pad_off >= prog->len))
+			return;
+		exc->ranges[i].begin = (u64)(long)image + addrs[rec->begin_off];
+		exc->ranges[i].end = (u64)(long)image + addrs[rec->end_off];
+		exc->ranges[i].pad = (u64)(long)image + addrs[rec->landing_pad_off];
+	}
+	exc->nr_ranges = n;
+}
+
+void bpf_exc_free_info(struct bpf_prog_aux *aux)
+{
+	struct bpf_exception_info *exc = aux->exc;
+
+	if (!exc)
+		return;
+	kvfree(exc->ranges);
+	kvfree(exc->info);
+	kfree(exc);
+	aux->exc = NULL;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index e72e68ebfe85..b97ac04785c3 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -11,6 +11,10 @@ struct bpf_verifier_env;
 struct bpf_verifier_state;
 struct bpf_func_state;
 struct bpf_insn;
+struct bpf_cleanup_info;
+struct bpf_cleanup_range;
+struct bpf_prog;
+struct bpf_prog_aux;
 
 int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
 		       bpfptr_t uattr);
@@ -25,5 +29,7 @@ bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
 bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
 int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog);
 int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn);
+int bpf_exc_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt);
+const struct bpf_cleanup_range *bpf_exc_pad_for_ip(const struct bpf_prog *prog, u64 ip);
 
 #endif /* __BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 5b7fe4ba610b..fc1d98eddf36 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -11,6 +11,7 @@
 #include <linux/sched/signal.h>
 #include <net/xdp.h>
 #include "disasm.h"
+#include "exception.h"
 
 #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
 
@@ -739,6 +740,32 @@ static void keep_funcs_with_addr_taken(struct bpf_verifier_env *env)
 	}
 }
 
+static int keep_subprog_exits(struct bpf_verifier_env *env)
+{
+	u32 i, j;
+
+	for (i = 0; i < env->subprog_cnt; i++) {
+		bool found = false;
+		u32 start;
+
+		if (!env->subprog_info[i].might_unwind)
+			continue;
+		start = env->subprog_info[i].start;
+		for (j = env->subprog_info[i + 1].start; j-- > start; ) {
+			if (env->prog->insnsi[j].code != (BPF_JMP | BPF_EXIT))
+				continue;
+			env->insn_aux_data[j].seen = env->pass_cnt;
+			found = true;
+			break;
+		}
+		if (!found) {
+			verbose(env, "subprog %u can be unwound through but has no exit\n", i);
+			return -EINVAL;
+		}
+	}
+	return 0;
+}
+
 int bpf_opt_remove_dead_code(struct bpf_verifier_env *env)
 {
 	struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
@@ -746,6 +773,9 @@ int bpf_opt_remove_dead_code(struct bpf_verifier_env *env)
 	int i, err;
 
 	keep_funcs_with_addr_taken(env);
+	err = keep_subprog_exits(env);
+	if (err)
+		return err;
 
 	for (i = 0; i < insn_cnt; i++) {
 		int j;
@@ -1286,6 +1316,53 @@ static int resolve_func_ptrs(struct bpf_verifier_env *env, struct bpf_prog *prog
 	return 0;
 }
 
+static int exc_info_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
+				u32 start, u32 end)
+{
+	struct bpf_cleanup_info *recs;
+	u32 i, cnt = 0;
+
+	if (!env->cleanup_info_cnt)
+		return 0;
+
+	for (i = start; i < end; i++) {
+		if (env->insn_aux_data[i].cleanup_pad)
+			cnt++;
+	}
+	if (!cnt)
+		return 0;
+
+	recs = kvmalloc_array(cnt, sizeof(*recs), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+	if (!recs)
+		return -ENOMEM;
+
+	for (i = start, cnt = 0; i < end; i++) {
+		u32 pad = env->insn_aux_data[i].cleanup_pad;
+
+		if (!pad)
+			continue;
+		pad--;
+		if (verifier_bug_if(pad < start || pad >= end, env,
+				    "insn %u is covered by a landing pad at %u outside its subprog [%u, %u)",
+				    i, pad, start, end)) {
+			kvfree(recs);
+			return -EFAULT;
+		}
+		recs[cnt].begin_off = i - start;
+		recs[cnt].end_off = i - start + 1;
+		recs[cnt].landing_pad_off = pad - start;
+		cnt++;
+	}
+	return bpf_exc_attach_info(sub->aux, recs, cnt);
+}
+
+int bpf_exc_attach_main_prog(struct bpf_verifier_env *env, struct bpf_prog *prog)
+{
+	if (!env || env->subprog_cnt > 1)
+		return 0;
+	return exc_info_for_subprog(env, prog, 0, prog->len);
+}
+
 static int jit_subprogs(struct bpf_verifier_env *env)
 {
 	struct bpf_prog *prog = env->prog, **func, *tmp;
@@ -1423,6 +1500,8 @@ static int jit_subprogs(struct bpf_verifier_env *env)
 		func[i]->aux->token = prog->aux->token;
 		if (!i)
 			func[i]->aux->exception_boundary = env->seen_exception;
+		if (exc_info_for_subprog(env, func[i], subprog_start, subprog_end))
+			goto out_free;
 		func[i] = bpf_int_jit_compile(env, func[i]);
 		if (!func[i]->jited) {
 			err = -ENOTSUPP;
@@ -1532,6 +1611,9 @@ static int jit_subprogs(struct bpf_verifier_env *env)
 	prog->aux->bpf_exception_cb = (void *)func[env->exception_callback_subprog]->bpf_func;
 	prog->aux->exception_boundary = func[0]->aux->exception_boundary;
 	prog->aux->stack_arg_sp_adjust = func[0]->aux->stack_arg_sp_adjust;
+	prog->aux->exc = func[0]->aux->exc;
+	func[0]->aux->exc = NULL;
+	prog->aux->epilogue_ip = func[0]->aux->epilogue_ip;
 	bpf_prog_jit_attempt_done(prog);
 	return 0;
 out_free:
@@ -1757,6 +1839,43 @@ static int may_goto_expand(struct bpf_insn *insn_buf, int off, int stack_off,
 	return cnt + tail_cnt;
 }
 
+/*
+ * Follow each bpf_unwind() call with 'r0 = 0; exit', or with
+ * 'r0 = 0; goto pad' where a record covers the call.
+ */
+int bpf_exc_patch_unwind_calls(struct bpf_verifier_env *env)
+{
+	int insn_cnt = env->prog->len;
+	struct bpf_insn insn_buf[3];
+	struct bpf_prog *new_prog;
+	int i, off, delta = 0;
+
+	for (i = 0; i < insn_cnt; i++) {
+		struct bpf_insn *insn = env->prog->insnsi + i + delta;
+		u32 pad = env->insn_aux_data[i + delta].cleanup_pad;
+
+		if (!bpf_is_unwind_kfunc(insn))
+			continue;
+
+		insn_buf[0] = *insn;
+		insn_buf[1] = BPF_MOV64_IMM(BPF_REG_0, 0);
+		insn_buf[2] = BPF_EXIT_INSN();
+		if (pad) {
+			/* Stored as index + 1; a pad after the call moves with it. */
+			pad--;
+			off = (pad > i + delta ? pad + 2 : pad) - (i + delta + 3);
+			insn_buf[2] = off == (s16)off ? BPF_JMP_A(off) : BPF_JMP32_A(off);
+		}
+
+		new_prog = bpf_patch_insn_data(env, i + delta, insn_buf, 3);
+		if (!new_prog)
+			return -ENOMEM;
+		delta += 2;
+		env->prog = new_prog;
+	}
+	return 0;
+}
+
 /* Do various post-verification rewrites in a single program pass.
  * These rewrites simplify JIT and interpreter implementations.
  */
@@ -2135,6 +2254,25 @@ int bpf_do_misc_fixups(struct bpf_verifier_env *env)
 			goto next_insn;
 		if (insn->src_reg == BPF_PSEUDO_CALL)
 			goto next_insn;
+		if (bpf_is_unwind_resume_kfunc(insn)) {
+			/*
+			 * A pad's resume is just the frame returning, to
+			 * where bpf_unwind() pointed its return address: its
+			 * caller's pad or epilogue, or the kernel from the main
+			 * program. The verifier checked this exit with r0 a
+			 * known zero, so return zero.
+			 */
+			insn_buf[0] = BPF_MOV64_IMM(BPF_REG_0, 0);
+			insn_buf[1] = BPF_EXIT_INSN();
+			cnt = 2;
+			new_prog = bpf_patch_insn_data(env, i + delta, insn_buf, cnt);
+			if (!new_prog)
+				return -ENOMEM;
+			delta += cnt - 1;
+			env->prog = prog = new_prog;
+			insn = new_prog->insnsi + i + delta;
+			goto next_insn;
+		}
 		if (insn->src_reg == BPF_PSEUDO_KFUNC_CALL) {
 			ret = bpf_fixup_kfunc_call(env, insn, insn_buf, i + delta, &cnt);
 			if (ret)
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 4eccd6742eba..c6894d6185ab 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -31,6 +31,7 @@
 #include <linux/buildid.h>
 
 #include "../../lib/kstrtox.h"
+#include "exception.h"
 
 /* If kernel subsystem is allowing eBPF programs to call this function,
  * inside its own verifier_ops->get_func_proto() callback it should return
@@ -3424,8 +3425,50 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
 	return false;
 }
 
+struct bpf_unwind_ctx {
+	u32 cnt;
+};
+
+static bool bpf_unwind_rewrite(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra)
+{
+	const struct bpf_cleanup_range *rec;
+	struct bpf_unwind_ctx *ctx = cookie;
+	struct bpf_prog *prog;
+
+	rcu_read_lock();
+	prog = bpf_prog_ksym_find(ip);
+	rcu_read_unlock();
+	if (!prog)
+		return !ctx->cnt;
+	ctx->cnt++;
+
+	/*
+	 * The frame that called bpf_unwind(): bpf_exc_patch_unwind_calls()
+	 * put 'r0 = 0' and a jump to its pad, or an exit, after the call,
+	 * so leave its return address alone and let it go on there. The pad
+	 * then starts with r0 at a known zero.
+	 */
+	if (ctx->cnt == 1)
+		return bpf_is_subprog(prog);
+
+	rec = bpf_exc_pad_for_ip(prog, ip);
+	if (rec) {
+		*ra = rec->pad;
+	} else if (prog->aux->epilogue_ip) {
+		*ra = prog->aux->epilogue_ip;
+	} else {
+		WARN_ON_ONCE(1);
+		return false;
+	}
+
+	return bpf_is_subprog(prog);
+}
+
 __bpf_kfunc void bpf_unwind(void)
 {
+	struct bpf_unwind_ctx ctx = {};
+
+	arch_bpf_stack_walk_ra(bpf_unwind_rewrite, &ctx);
 }
 
 __bpf_kfunc void bpf_throw(u64 cookie)
@@ -5095,6 +5138,8 @@ BTF_ID_FLAGS(func, bpf_task_get_cgroup1, KF_ACQUIRE | KF_RCU | KF_RET_NULL)
 BTF_ID_FLAGS(func, bpf_task_from_pid, KF_ACQUIRE | KF_RET_NULL)
 BTF_ID_FLAGS(func, bpf_task_from_vpid, KF_ACQUIRE | KF_RET_NULL)
 BTF_ID_FLAGS(func, bpf_throw)
+BTF_ID_FLAGS(func, bpf_unwind)
+BTF_ID_FLAGS(func, bpf_unwind_resume)
 #ifdef CONFIG_BPF_EVENTS
 BTF_ID_FLAGS(func, bpf_send_signal_task)
 #endif
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 488ceb9dae1b..917635adb5f6 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -19299,6 +19299,10 @@ static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
 	int pad = bpf_exc_pad_of_call(env, *insn_idx);
 	int err;
 
+	err = bpf_exc_check_prog(env);
+	if (err)
+		return err;
+
 	if (pad < 0) {
 		if (!env->cur_state->curframe) {
 			err = check_resource_leak(env, false, true,
@@ -22898,6 +22902,9 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 		/* program is valid, convert *(u32*)(ctx + off) accesses */
 		ret = bpf_convert_ctx_accesses(env);
 
+	if (ret == 0)
+		ret = bpf_exc_patch_unwind_calls(env);
+
 	if (ret == 0)
 		ret = bpf_do_misc_fixups(env);
 
-- 
2.53.0-Meta


  parent reply	other threads:[~2026-10-01 13:31 UTC|newest]

Thread overview: 50+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-10-01 13:48   ` sashiko-bot
2026-10-02 18:17     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 19:06     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier Yonghong Song
2026-10-01 13:50   ` sashiko-bot
2026-10-02 19:31     ` Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 20:49     ` Yonghong Song
2026-10-03 12:23   ` Alexei Starovoitov
2026-10-04 17:56     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 21:10     ` Yonghong Song
2026-10-03 12:25   ` Alexei Starovoitov
2026-10-04 17:59     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-10-03 12:25   ` Alexei Starovoitov
2026-10-04 18:26     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 10/22] bpf: Do not use a private stack for a program that can unwind Yonghong Song
2026-10-01 13:53   ` sashiko-bot
2026-10-02 21:38     ` Yonghong Song
2026-10-01 13:31 ` Yonghong Song [this message]
2026-10-01 14:31   ` [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses bot+bpf-ci
2026-10-02 21:48     ` Yonghong Song
2026-10-03 12:26   ` Alexei Starovoitov
2026-10-04 18:28     ` Yonghong Song
2026-10-04 18:29     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-10-01 13:49   ` sashiko-bot
2026-10-02 21:54     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 13/22] bpf, arm64: " Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-10-01 13:46   ` sashiko-bot
2026-10-02 22:09     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
2026-10-01 13:32 ` [PATCH bpf-next v8 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261001133103.1340994-1-yonghong.song@linux.dev \
    --to=yonghong.song@linux.dev \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=kernel-team@fb.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.