From: Yonghong Song <yonghong.song@linux.dev>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
Andrii Nakryiko <andrii@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Eduard Zingerman <eddyz87@gmail.com>,
kernel-team@fb.com
Subject: [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with
Date: Mon, 28 Sep 2026 17:16:43 -0700 [thread overview]
Message-ID: <20260929001643.3249386-1-yonghong.song@linux.dev> (raw)
In-Reply-To: <20260929001601.3242665-1-yonghong.song@linux.dev>
A landing pad's entry state is the state at the call it belongs to, taken
before the callee ran -- push_cleanup_pad_branch() snapshots it there. That
is right for the frame's own registers and stack, which the callee's
epilogue puts back, and wrong for what the program shares: nothing restores
the locks in bpf_verifier_state. So a callee that drops a lock its caller
took and then resumes leaves the caller's pad verified against a state that
still holds it. Say foo() takes an RCU read lock and calls bar(), which
unwinds:
foo:
call bpf_rcu_read_lock
1: call bar /* covered, pad at 3 */
2: r0 = 0; exit
3: call bpf_rcu_read_unlock /* pad: drops the lock foo() took */
call bpf_unwind_resume
bar:
1: call bpf_unwind /* covered, pad at 3 */
2: r0 = 0; exit
3: call bpf_rcu_read_unlock /* pad: drops it a second time */
call bpf_unwind_resume
At run time both pads run and the one lock is released twice, and no single
state shows it: bar's path ends at its resume, and foo's pad is verified
from the snapshot at its call to bar. A frame that acquires a lock and
leaves through an unwind hides the same way. Fix both by making the
snapshot true -- record what the program holds when a frame is entered and
require an unwind leaving it to have put that back, at the resume and at a
bpf_unwind() no record covers. References go by id, since ids only go up
and bpf_reference_state does not say which frame acquired one.
The frames in between need the same of them, and have nothing to run: where
no record covers the call a frame is suspended at, the JIT sends it to its
epilogue, so what it acquired since it was entered is dropped on the floor
and no path of its own arrives to say so. Ask it at the call instead, which
is where it is abandoned -- check_unwind_through_call(). Per frame rather
than of the whole stack at the unwind, since a frame that does carry a
record may hold what its pad will release. A callx counts as any subprogram
that might unwind, its target not being known there.
So every frame an unwind leaves is asked, one way of asking per way out:
- "a resume": a frame with a pad runs it and ends at bpf_unwind_resume().
The frame that raised the unwind leaves this way where a record covers
its bpf_unwind(), and so does every frame above it whose call is
covered.
- "an unwind with no landing pad": the frame that raised it with no
record over its bpf_unwind(), returning through the exit patched in
after it.
- "an unwind through this call": every frame above it whose call no
record covers, asked at the call since nothing of it runs again.
Nothing else is left to ask: a frame other than the one that raised the
unwind is suspended at a call, and the call either carries a record or does
not. The unwind reaches no further than the program it was raised in, the
walk stopping at the first frame that is not a subprogram.
One walk is then left that cannot happen. A resume ends its frame the way
an exit does, so the verifier continued the caller at the instruction
after its call -- the one place a resume does not return to, since
bpf_unwind() rewrote that address before any pad ran. Where it does return
is walked already: the caller's pad is the branch pushed at the call, and
with no pad nothing of the caller runs. Walking on anyway keeps code the
JIT never reaches out of the dead code sweep, and arrives holding what
only the pad releases, which refuses a correct program. End the path there
instead, as an unwind with no landing pad already does. The main program's
frame keeps the old way out: no caller to leave to, and the exit is what
holds its zero to the program type.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 12 ++++++
kernel/bpf/exception.c | 76 ++++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 6 +++
kernel/bpf/verifier.c | 50 +++++++++++++++++++++++-
4 files changed, 142 insertions(+), 2 deletions(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 0143688896b0..75e572a8a1ba 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -339,6 +339,18 @@ struct bpf_func_state {
bool in_async_callback_fn;
bool in_exception_callback_fn;
bool no_stack_arg_load;
+ /*
+ * What the program held when this frame was entered. An unwind leaves
+ * the frame without running anything below it, so the frame has to put
+ * these back to what it found before it goes -- otherwise a caller's
+ * landing pad, whose state was taken at the call, is wrong about them.
+ */
+ u32 entry_active_locks;
+ u32 entry_preempt_locks;
+ u32 entry_rcu_locks;
+ u32 entry_irq_id;
+ u32 entry_id_gen;
+ u32 entry_acquired_refs;
/* For callback calling functions that limit number of possible
* callback executions (e.g. bpf_loop) keeps track of current
* simulated iteration number.
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index c1779d2d02f0..c0b0b8478af5 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -12,6 +12,72 @@
BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
+void bpf_exc_record_frame_entry(const struct bpf_verifier_state *state,
+ struct bpf_func_state *frame, u32 id_gen)
+{
+ u32 i;
+
+ frame->entry_active_locks = state->active_locks;
+ frame->entry_preempt_locks = state->active_preempt_locks;
+ frame->entry_rcu_locks = state->active_rcu_locks;
+ frame->entry_irq_id = state->active_irq_id;
+
+ /* Ids only ever go up, so this one tells the frame's own apart. */
+ frame->entry_id_gen = id_gen;
+ frame->entry_acquired_refs = 0;
+ for (i = 0; i < state->acquired_refs; i++)
+ if (state->refs[i].type == REF_TYPE_PTR)
+ frame->entry_acquired_refs++;
+}
+
+int bpf_exc_check_frame_balance(struct bpf_verifier_env *env, const char *prefix)
+{
+ const struct bpf_verifier_state *state = env->cur_state;
+ const struct bpf_func_state *frame = cur_func(env);
+ u32 i, held;
+ const char *what;
+
+ if (state->active_rcu_locks != frame->entry_rcu_locks)
+ what = "bpf_rcu_read_lock";
+ else if (state->active_preempt_locks != frame->entry_preempt_locks)
+ what = "bpf_preempt_disable";
+ else if (state->active_irq_id != frame->entry_irq_id)
+ what = "bpf_local_irq_save";
+ else if (state->active_locks != frame->entry_active_locks)
+ what = "bpf_spin_lock";
+ else
+ what = NULL;
+
+ if (what) {
+ verbose(env, "%s does not leave the frame's %s state as it found it\n",
+ prefix, what);
+ return -EINVAL;
+ }
+
+ /*
+ * References the same way. ids only go up, so entry_id_gen splits
+ * refs[] in two at frame entry: nothing above that line may still be
+ * held, and the count below it has to be what it was.
+ */
+ for (i = 0, held = 0; i < state->acquired_refs; i++) {
+ if (state->refs[i].type != REF_TYPE_PTR)
+ continue;
+ if (state->refs[i].id > frame->entry_id_gen) {
+ verbose(env, "%s keeps the reference id=%d the frame acquired\n",
+ prefix, state->refs[i].id);
+ return -EINVAL;
+ }
+ held++;
+ }
+ if (held != frame->entry_acquired_refs) {
+ verbose(env, "%s does not leave the frame's references as it found it\n",
+ prefix);
+ return -EINVAL;
+ }
+
+ return 0;
+}
+
static int reject_throw(struct bpf_verifier_env *env)
{
u32 i;
@@ -46,6 +112,16 @@ static int mark_call_sites(struct bpf_verifier_env *env)
return 0;
}
+bool bpf_prog_may_unwind(const struct bpf_verifier_env *env)
+{
+ u32 i;
+
+ for (i = 0; i < env->subprog_cnt; i++)
+ if (env->subprog_info[i].might_unwind)
+ return true;
+ return false;
+}
+
int bpf_exc_check_prog(struct bpf_verifier_env *env)
{
if (bpf_prog_is_offloaded(env->prog->aux)) {
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index 1552438083d8..e93da039b5bd 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -6,9 +6,15 @@
#include <linux/types.h>
struct bpf_verifier_env;
+struct bpf_verifier_state;
+struct bpf_func_state;
int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
int bpf_exc_check_prog(struct bpf_verifier_env *env);
+bool bpf_prog_may_unwind(const struct bpf_verifier_env *env);
+void bpf_exc_record_frame_entry(const struct bpf_verifier_state *state,
+ struct bpf_func_state *frame, u32 id_gen);
+int bpf_exc_check_frame_balance(struct bpf_verifier_env *env, const char *prefix);
int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index ee074d4a936b..4bdee3f02fe9 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10774,6 +10774,7 @@ static int setup_func_entry(struct bpf_verifier_env *env, int subprog, int calls
callsite,
state->curframe + 1 /* frameno within this callchain */,
subprog /* subprog number within this prog */);
+ bpf_exc_record_frame_entry(state, callee, env->id_gen);
err = set_callee_state_cb(env, caller, callee, callsite);
if (err)
goto err_out;
@@ -19207,6 +19208,32 @@ static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
return 0;
}
+/* Can an unwind come back out of this call? */
+static bool call_may_unwind(struct bpf_verifier_env *env, const struct bpf_insn *insn,
+ int insn_idx)
+{
+ int subprog;
+
+ /* Which subprog a callx lands in is not known here, so any may be it. */
+ if (bpf_is_callx(insn))
+ return bpf_prog_may_unwind(env);
+ if (insn->src_reg != BPF_PSEUDO_CALL)
+ return false;
+ subprog = bpf_find_subprog(env, insn_idx + insn->imm + 1);
+ return subprog >= 0 && env->subprog_info[subprog].might_unwind;
+}
+
+static int check_unwind_through_call(struct bpf_verifier_env *env, int insn_idx)
+{
+ const struct bpf_insn *insn = &env->prog->insnsi[insn_idx];
+
+ if (bpf_exc_pad_of_call(env, insn_idx) >= 0)
+ return 0;
+ if (!call_may_unwind(env, insn, insn_idx))
+ return 0;
+ return bpf_exc_check_frame_balance(env, "an unwind through this call");
+}
+
static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
bool *do_print_state)
{
@@ -19215,8 +19242,13 @@ static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
int err;
if (pad < 0) {
- err = check_resource_leak(env, false, !env->cur_state->curframe,
- "an unwind with no landing pad");
+ if (!env->cur_state->curframe) {
+ err = check_resource_leak(env, false, true,
+ "an unwind with no landing pad");
+ if (err)
+ return err;
+ }
+ err = bpf_exc_check_frame_balance(env, "an unwind with no landing pad");
if (err)
return err;
if (env->cur_state->curframe)
@@ -19498,6 +19530,16 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
if (bpf_is_unwind_kfunc(insn))
return process_bpf_unwind(env, &env->insn_idx,
do_print_state);
+ err = bpf_exc_check_frame_balance(env, "a resume");
+ if (err)
+ return err;
+ /*
+ * No need to walk into the caller: its pad was
+ * pushed as a branch at its call, and with no
+ * pad nothing of it runs.
+ */
+ if (env->cur_state->curframe)
+ return PROCESS_BPF_EXIT;
/*
* Mark r0 a known zero -- unknown first, as
* the known-zero helper keeps the type it
@@ -19512,6 +19554,10 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
mark_reg_scratched(env, BPF_REG_0);
/* An unwind out of this call resumes at the pad. */
err = push_cleanup_pad_branch(env, env->insn_idx);
+ if (err)
+ return err;
+ /* Or, with no pad, leaves the frame for good. */
+ err = check_unwind_through_call(env, env->insn_idx);
if (err)
return err;
if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
--
2.53.0-Meta
next prev parent reply other threads:[~2026-09-29 0:16 UTC|newest]
Thread overview: 46+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-29 0:33 ` sashiko-bot
2026-09-29 21:58 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-09-29 0:31 ` sashiko-bot
2026-09-29 22:04 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 07/22] bpf: Resume a covered call at its landing pad Yonghong Song
2026-09-29 0:31 ` sashiko-bot
2026-09-30 0:28 ` Yonghong Song
2026-09-29 0:16 ` Yonghong Song [this message]
2026-09-29 0:36 ` [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with sashiko-bot
2026-09-30 1:09 ` Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:10 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 10/22] bpf: Refuse a private stack for a program that can unwind Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
2026-09-29 1:14 ` bot+bpf-ci
2026-09-30 1:18 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-29 0:30 ` sashiko-bot
2026-09-30 1:34 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 13/22] bpf, arm64: " Yonghong Song
2026-09-29 1:14 ` bot+bpf-ci
2026-09-29 0:17 ` [PATCH bpf-next v7 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:42 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:46 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 2:19 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 3:12 ` Yonghong Song
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929001643.3249386-1-yonghong.song@linux.dev \
--to=yonghong.song@linux.dev \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=kernel-team@fb.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.