From: Yonghong Song <yonghong.song@linux.dev>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
Andrii Nakryiko <andrii@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Eduard Zingerman <eddyz87@gmail.com>,
kernel-team@fb.com
Subject: [PATCH bpf-next v7 07/22] bpf: Resume a covered call at its landing pad
Date: Mon, 28 Sep 2026 17:16:38 -0700 [thread overview]
Message-ID: <20260929001638.3248952-1-yonghong.song@linux.dev> (raw)
In-Reply-To: <20260929001601.3242665-1-yonghong.song@linux.dev>
A landing pad runs in the frame that owns it, entered by an ordinary
return: bpf_unwind() rewrites the frame's saved return address, so the call
the frame is suspended at comes back at the pad rather than at the next
instruction. That makes the pad a second successor of a covered call, in
the same frame, whose entry state is knowable without verifying the callee
-- the state at the call with the caller-saved registers gone, since the
callee's epilogue puts r6-r9 and the stack back on its way out.
push_cleanup_pad_branch() pushes exactly that. The call to bpf_unwind()
itself never comes back to the instruction after it: it resumes at this
frame's pad where a record covers the call, and otherwise the frame
returns at once. Where that frame is the main program's, returning at once
is the program returning, so it leaves through process_bpf_exit_full() and
the zero it returns is held to the program type.
Precision backtracking has to tell those two edges apart, and subseq_idx is
what it has to do it with. Coming back to a covered call from its own pad
stays in this frame, the callee never having been entered on that path.
Coming back to a resume is the opposite -- a resume leaves its frame the
way an exit does -- so the walk enters the callee there, and r6-r9 and the
stack stay marked in this frame's masks until it comes back out. Miss that
and the callee is walked against the caller's masks.
Answering the two here skips check_kfunc_call(), so the filter it applies
first -- whether this program may call this kfunc at all -- is split out as
check_kfunc_allowed() and applied here too. Until the patch that registers
them, that filter is what refuses them.
The CFG walk starts summarising it too: a subprogram calling bpf_unwind()
is marked might_unwind, and merge_callee_effects() carries that up to its
callers. Nothing reads it yet; later patches do.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 1 +
kernel/bpf/backtrack.c | 42 ++++++++++++++
kernel/bpf/cfg.c | 11 ++++
kernel/bpf/verifier.c | 108 ++++++++++++++++++++++++++++++++---
4 files changed, 154 insertions(+), 8 deletions(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 6d78c20e6507..0143688896b0 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -836,6 +836,7 @@ struct bpf_subprog_info {
s16 fastcall_stack_off;
bool has_tail_call: 1;
bool might_throw: 1;
+ bool might_unwind: 1;
bool tail_call_reachable: 1;
bool has_ld_abs: 1;
bool is_cb: 1;
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 0e38b9575328..90f30e152cbe 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -4,6 +4,7 @@
#include <linux/bpf_verifier.h>
#include <linux/filter.h>
#include <linux/bitmap.h>
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -434,6 +435,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
return -EFAULT;
}
+ if (bpf_exc_pad_of_call(env, idx) == subseq_idx) {
+ /*
+ * We came from this call's landing pad, which
+ * runs in the caller's frame: on that path the
+ * callee's frame was never entered, so there is
+ * no frame to leave. The call clobbered r0-r5;
+ * r6-r9 and the stack are the caller's own and
+ * keep going back from here.
+ */
+ bt_clear_reg(bt, BPF_REG_0);
+ if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
+ verifier_bug(env, "landing pad unexpected regs %x",
+ bt_reg_mask(bt));
+ return -EFAULT;
+ }
+ return 0;
+ }
+
/* callx calls static subprogs only */
if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
/* check that jump history doesn't have any
@@ -523,6 +542,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
if (bt_subprog_exit(bt))
return -EFAULT;
return 0;
+ } else if (bpf_is_unwind_resume_kfunc(insn)) {
+ /*
+ * A resume leaves its frame the way an exit does, so
+ * the walk is crossing from the caller into the callee
+ * here and has a frame to enter. The zero the resume
+ * returns is its own: nothing further back defines r0,
+ * and r1-r5 the call clobbered.
+ */
+ bt_clear_reg(bt, BPF_REG_0);
+ bt_clear_reg(bt, BPF_REG_2);
+ if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
+ verifier_bug(env, "backtracking resume unexpected regs %x",
+ bt_reg_mask(bt));
+ return -EFAULT;
+ }
+ if (bt_subprog_enter(bt))
+ return -EFAULT;
+ return 0;
} else if (opcode == BPF_CALL) {
/* kfunc with imm==0 is invalid and fixup_kfunc_call will
* catch this error later. Make backtracking conservative
@@ -956,6 +993,11 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
if (!st)
break;
+ if (verifier_bug_if(bt->frame > st->curframe, env,
+ "backtrack frame %d, state curframe %d",
+ bt->frame, st->curframe))
+ return -EFAULT;
+
for (fr = bt->frame; fr >= 0; fr--) {
func = st->frame[fr];
bitmap_from_u64(mask, bt_frame_reg_mask(bt, fr));
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index 2eb07397e874..82b5abcc736b 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -76,6 +76,14 @@ static void mark_subprog_might_throw(struct bpf_verifier_env *env, int off)
subprog->might_throw = true;
}
+static void mark_subprog_might_unwind(struct bpf_verifier_env *env, int off)
+{
+ struct bpf_subprog_info *subprog;
+
+ subprog = bpf_find_containing_subprog(env, off);
+ subprog->might_unwind = true;
+}
+
/* 't' is an index of a call-site.
* 'w' is a callee entry point.
* Eventually this function would be called when env->cfg.insn_state[w] == EXPLORED.
@@ -91,6 +99,7 @@ static void merge_callee_effects(struct bpf_verifier_env *env, int t, int w)
caller->changes_pkt_data |= callee->changes_pkt_data;
caller->might_sleep |= callee->might_sleep;
caller->might_throw |= callee->might_throw;
+ caller->might_unwind |= callee->might_unwind;
}
enum {
@@ -668,6 +677,8 @@ static int visit_insn(int t, struct bpf_verifier_env *env)
mark_subprog_changes_pkt_data(env, t);
if (ret == 0 && bpf_is_throw_kfunc(insn))
mark_subprog_might_throw(env, t);
+ if (ret == 0 && bpf_is_unwind_kfunc(insn))
+ mark_subprog_might_unwind(env, t);
}
return visit_func_call_insn(t, insns, env, insn->src_reg == BPF_PSEUDO_CALL);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index fc3df452de2e..ee074d4a936b 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -14725,6 +14725,32 @@ static int check_special_kfunc(struct bpf_verifier_env *env, struct bpf_call_arg
static int check_return_code(struct bpf_verifier_env *env, int regno, const char *reg_name);
+static int check_kfunc_allowed(struct bpf_verifier_env *env, struct bpf_insn *insn,
+ int insn_idx, struct bpf_call_arg_meta *meta)
+{
+ const char *operation;
+ int err;
+
+ err = bpf_fetch_kfunc_arg_meta(env, insn->imm, insn->off, meta);
+ if (err == -EACCES && meta->func_name) {
+ verbose(env, "calling kernel function %s is not allowed\n", meta->func_name);
+ operation = bpf_diag_fmt(env, "kfunc %s", meta->func_name);
+ bpf_diag_policy(
+ env, insn_idx, operation, "this program cannot call the kfunc",
+ "Use a kfunc allowed for this program type and attach point, or change the program context.");
+ }
+ return err;
+}
+
+/* noinline saves the caller a 200-byte struct bpf_call_arg_meta on its frame. */
+static noinline int check_kfunc_allowed_only(struct bpf_verifier_env *env,
+ struct bpf_insn *insn, int insn_idx)
+{
+ struct bpf_call_arg_meta meta;
+
+ return check_kfunc_allowed(env, insn, insn_idx, &meta);
+}
+
static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
int *insn_idx_p)
{
@@ -14746,14 +14772,7 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
if (!insn->imm)
return 0;
- err = bpf_fetch_kfunc_arg_meta(env, insn->imm, insn->off, &meta);
- if (err == -EACCES && meta.func_name) {
- verbose(env, "calling kernel function %s is not allowed\n", meta.func_name);
- operation = bpf_diag_fmt(env, "kfunc %s", meta.func_name);
- bpf_diag_policy(
- env, insn_idx, operation, "this program cannot call the kfunc",
- "Use a kfunc allowed for this program type and attach point, or change the program context.");
- }
+ err = check_kfunc_allowed(env, insn, insn_idx, &meta);
if (err)
return err;
desc_btf = meta.btf;
@@ -19167,6 +19186,57 @@ enum {
INSN_IDX_UPDATED = 2,
};
+static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
+{
+ struct bpf_verifier_state *branch;
+ struct bpf_func_state *frame;
+ int pad = bpf_exc_pad_of_call(env, insn_idx);
+
+ if (pad < 0)
+ return 0;
+ branch = push_stack(env, pad, insn_idx, false);
+ if (IS_ERR(branch))
+ return PTR_ERR(branch);
+ frame = branch->frame[branch->curframe];
+ /*
+ * The state at that call with the caller-saved registers gone: the
+ * callee's epilogue put r6-r9 and the stack back on the way out.
+ */
+ clear_caller_saved_regs(env, frame->regs);
+ mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ return 0;
+}
+
+static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
+ bool *do_print_state)
+{
+ struct bpf_func_state *frame = cur_func(env);
+ int pad = bpf_exc_pad_of_call(env, *insn_idx);
+ int err;
+
+ if (pad < 0) {
+ err = check_resource_leak(env, false, !env->cur_state->curframe,
+ "an unwind with no landing pad");
+ if (err)
+ return err;
+ if (env->cur_state->curframe)
+ return PROCESS_BPF_EXIT;
+ /*
+ * The main program's frame returns at once, which is the
+ * program returning. Mark r0 the zero the fixups leave after
+ * the call, and leave through the exit, which is what holds
+ * that zero to the program type.
+ */
+ mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
+ mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
+ return process_bpf_exit_full(env, do_print_state, false);
+ }
+ clear_caller_saved_regs(env, frame->regs);
+ mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ *insn_idx = pad;
+ return INSN_IDX_UPDATED;
+}
+
static int process_bpf_exit_full(struct bpf_verifier_env *env,
bool *do_print_state,
bool exception_exit)
@@ -19421,7 +19491,29 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
return -EINVAL;
}
}
+ if (bpf_is_unwind_kfunc(insn) || bpf_is_unwind_resume_kfunc(insn)) {
+ err = check_kfunc_allowed_only(env, insn, env->insn_idx);
+ if (err)
+ return err;
+ if (bpf_is_unwind_kfunc(insn))
+ return process_bpf_unwind(env, &env->insn_idx,
+ do_print_state);
+ /*
+ * Mark r0 a known zero -- unknown first, as
+ * the known-zero helper keeps the type it
+ * finds, which here is NOT_INIT. The fixups
+ * lower this to 'r0 = 0; exit', so the frame
+ * returns a real zero.
+ */
+ mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
+ mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
+ return process_bpf_exit_full(env, do_print_state, false);
+ }
mark_reg_scratched(env, BPF_REG_0);
+ /* An unwind out of this call resumes at the pad. */
+ err = push_cleanup_pad_branch(env, env->insn_idx);
+ if (err)
+ return err;
if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
cur_func(env)->no_stack_arg_load = true;
if (bpf_is_callx(insn))
--
2.53.0-Meta
next prev parent reply other threads:[~2026-09-29 0:16 UTC|newest]
Thread overview: 46+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-29 0:33 ` sashiko-bot
2026-09-29 21:58 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-09-29 0:31 ` sashiko-bot
2026-09-29 22:04 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-29 0:16 ` Yonghong Song [this message]
2026-09-29 0:31 ` [PATCH bpf-next v7 07/22] bpf: Resume a covered call at its landing pad sashiko-bot
2026-09-30 0:28 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
2026-09-29 0:36 ` sashiko-bot
2026-09-30 1:09 ` Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:10 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 10/22] bpf: Refuse a private stack for a program that can unwind Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
2026-09-29 1:14 ` bot+bpf-ci
2026-09-30 1:18 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-29 0:30 ` sashiko-bot
2026-09-30 1:34 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 13/22] bpf, arm64: " Yonghong Song
2026-09-29 1:14 ` bot+bpf-ci
2026-09-29 0:17 ` [PATCH bpf-next v7 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:42 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:46 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 2:19 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 3:12 ` Yonghong Song
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929001638.3248952-1-yonghong.song@linux.dev \
--to=yonghong.song@linux.dev \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=kernel-team@fb.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox