From: Yonghong Song <yonghong.song@linux.dev>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
Andrii Nakryiko <andrii@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Eduard Zingerman <eddyz87@gmail.com>,
kernel-team@fb.com
Subject: [PATCH bpf-next v3 08/20] bpf: Walk the exception unwind in the verifier
Date: Sat, 19 Sep 2026 22:43:07 -0700 [thread overview]
Message-ID: <20260920054307.868442-1-yonghong.song@linux.dev> (raw)
In-Reply-To: <20260920054225.864535-1-yonghong.song@linux.dev>
bpf_throw() is about to stop discarding the frames it unwinds and start
running their landing pads instead. Have the verifier walk the same thing,
step for step:
at a throw, with the frame chain in env->cur_state->frame[]
if a record covers the call control is at
run that record's landing pad in this frame
on reaching its resume, pop the frame and carry on
else
pop the frame and carry on
at the boundary, deliver
Doing it this way is what keeps the resource rules unchanged. Whatever a
pad releases is released in the verifier state too, so by the time the walk
reaches the boundary the state says exactly what the program will really
hold there -- and check_resource_leak(), which used to fire the moment a
throw was seen, simply moves to the end of the walk. A frame that no record
covers contributes nothing, so anything it held is still held when the walk
ends, and that is what gets reported.
Two things about that walk have to be said out loud, because the
instruction stream does not say them.
A resume belongs to the frame whose landing pad the walker called. Each JIT
lowers bpf_unwind_resume() as the way back out of a pad -- a bare return on
x86-64, a branch to the saved address on arm64 -- and that reaches the
walker only from there. An exception being in flight is not enough to tell:
a pad may call a subprogram that has a landing pad of its own and reaches
it by ordinary control flow, and the resume in it is in a pad body by every
static measure. So the state records the frame the unwind entered a pad in,
and a resume in any other frame is refused.
The edge from a throw to a pad crosses frames with no instruction in
between to account for them, and mark_chain_precision() walks that history
backwards. Left alone it stays in the pad's frame while it reads the
callee's instructions: a request for the pad frame's r6 is cleared by the
callee's own write to r6 -- so the caller's definition never becomes
precise -- or reaches the call instruction still set and trips "static
subprog unexpected regs". The history entry for a pad therefore records how
many frames the unwind popped, and the backtrack enters that many, the way
it enters one at a time for BPF_EXIT. A throwing global subprogram gets no
frame of its own, and its exception arrives at the landing pad rather than
at the next instruction, which the global-call check has to allow for.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_cleanup_abi.h | 16 ++++
include/linux/bpf_verifier.h | 6 +-
kernel/bpf/backtrack.c | 21 ++++-
kernel/bpf/states.c | 6 ++
kernel/bpf/verifier.c | 143 ++++++++++++++++++++++++++------
5 files changed, 163 insertions(+), 29 deletions(-)
create mode 100644 include/linux/bpf_cleanup_abi.h
diff --git a/include/linux/bpf_cleanup_abi.h b/include/linux/bpf_cleanup_abi.h
new file mode 100644
index 000000000000..b6c1d589abda
--- /dev/null
+++ b/include/linux/bpf_cleanup_abi.h
@@ -0,0 +1,16 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef _LINUX_BPF_CLEANUP_ABI_H
+#define _LINUX_BPF_CLEANUP_ABI_H
+
+/*
+ * Value arch_bpf_run_cleanup_pad() leaves in r0 on the way into a landing pad.
+ * It has to be a constant the verifier knows: LLVM names r0 as both the
+ * exception pointer and the exception selector register, so every pad reads it
+ * before anything else and is free to store what it read. Kept on its own
+ * because the verifier and the arch dispatchers, which are assembly, have to
+ * agree on it.
+ */
+#define BPF_PAD_ENTRY_R0 1
+
+#endif /* _LINUX_BPF_CLEANUP_ABI_H */
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 382177b80745..0721461aaabb 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -429,7 +429,8 @@ struct bpf_jmp_history_entry {
u32 prev_idx : 20;
/* special INSN_F_xxx flags */
u32 flags : 4;
- u32 : 8;
+ u32 unwind_frames : 4; /* frames the unwind popped to get here */
+ u32 : 4;
/*
* additional registers that need precision tracking when this
* jump is backtracked, vector of five 11-bit records
@@ -509,6 +510,8 @@ struct bpf_verifier_state {
bool speculative;
bool in_sleepable;
+ bool unwinding; /* an exception is in flight */
+ u8 unwind_frameno; /* the frame whose landing pad it entered */
/* first and last insn idx of this verifier state */
u32 first_insn_idx;
@@ -991,6 +994,7 @@ struct bpf_verifier_env {
} cfg;
struct backtrack_state bt;
struct bpf_jmp_history_entry *cur_hist_ent;
+ u8 unwind_frames; /* scratch: pops for the insn about to be recorded */
/* Per-callsite copy of parent's converged at_stack_in for cross-frame fills. */
struct arg_track **callsite_at_stack;
u32 pass_cnt; /* number of times do_check() was called */
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 507a366dffa4..c19e4fe247ce 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -4,6 +4,7 @@
#include <linux/bpf_verifier.h>
#include <linux/filter.h>
#include <linux/bitmap.h>
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -47,6 +48,7 @@ int bpf_push_jmp_history(struct bpf_verifier_env *env, struct bpf_verifier_state
p->flags = insn_flags;
p->spi = spi;
p->frame = frame;
+ p->unwind_frames = 0;
p->linked_regs = linked_regs;
cur->jmp_history_cnt = cnt;
env->cur_hist_ent = p;
@@ -419,10 +421,12 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
* extra instructions from subprog; the next
* instruction after call to global subprog
* should be literally next instruction in
- * caller program
+ * caller program -- or, if the callee threw,
+ * the landing pad of this call site
*/
- verifier_bug_if(idx + 1 != subseq_idx, env,
- "extra insn from subprog");
+ verifier_bug_if(idx + 1 != subseq_idx &&
+ bpf_cleanup_pad_of_call(env, idx) != subseq_idx,
+ env, "extra insn from subprog");
/* global subprog always sets R0 */
bt_clear_reg(bt, BPF_REG_0);
/* and if it does not set R2, main pass would catch it */
@@ -888,11 +892,11 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
}
for (i = last_idx;;) {
+ hist = get_jmp_hist_entry(st, history, i);
if (skip_first) {
err = 0;
skip_first = false;
} else {
- hist = get_jmp_hist_entry(st, history, i);
err = backtrack_insn(env, i, subseq_idx, hist, bt);
}
if (err == -ENOTSUPP) {
@@ -909,6 +913,15 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
*/
return 0;
subseq_idx = i;
+ /* An exception unwind reached this insn, a landing
+ * pad, from a throw or a resume that many frames
+ * deeper. There is no insn in between to backtrack
+ * over, so enter those frames here, the way BPF_EXIT
+ * does one at a time.
+ */
+ for (fr = 0; hist && fr < hist->unwind_frames; fr++)
+ if (bt_subprog_enter(bt))
+ return -EFAULT;
i = get_prev_insn_idx(st, i, &history);
if (i == -ENOENT)
break;
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 66fb11b6c6a7..b101baa43171 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -996,6 +996,12 @@ static bool states_equal(struct bpf_verifier_env *env,
if (old->in_sleepable != cur->in_sleepable)
return false;
+ if (old->unwinding != cur->unwinding)
+ return false;
+
+ if (old->unwinding && old->unwind_frameno != cur->unwind_frameno)
+ return false;
+
if (!refsafe(old, cur, &env->idmap_scratch))
return false;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 48fa94a582e7..6ca242d70a4c 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10,6 +10,7 @@
#include <linux/slab.h>
#include <linux/bpf.h>
#include <linux/btf.h>
+#include <linux/bpf_cleanup_abi.h>
#include <linux/bpf_verifier.h>
#include <linux/filter.h>
#include <net/netlink.h>
@@ -1715,6 +1716,8 @@ int bpf_copy_verifier_state(struct bpf_verifier_state *dst_state,
return err;
dst_state->speculative = src->speculative;
dst_state->in_sleepable = src->in_sleepable;
+ dst_state->unwinding = src->unwinding;
+ dst_state->unwind_frameno = src->unwind_frameno;
dst_state->curframe = src->curframe;
dst_state->branches = src->branches;
dst_state->parent = src->parent;
@@ -10558,8 +10561,8 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
return 0;
}
-static int process_bpf_exit_full(struct bpf_verifier_env *env,
- bool *do_print_state, bool exception_exit);
+static int process_bpf_exit_full(struct bpf_verifier_env *env, bool *do_print_state);
+static int unwind_step(struct bpf_verifier_env *env, u32 callsite, int *insn_idx);
static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
int *insn_idx)
@@ -10649,7 +10652,7 @@ static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
verbose(env, "failed to push state for global subprog exception path\n");
return PTR_ERR(branch);
}
- return process_bpf_exit_full(env, NULL, true);
+ return unwind_step(env, *insn_idx, insn_idx);
}
/* continue with next insn after call */
@@ -14550,7 +14553,7 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
env->prog->call_session_cookie = true;
if (bpf_is_throw_kfunc(insn))
- return process_bpf_exit_full(env, NULL, true);
+ return unwind_step(env, insn_idx, &env->insn_idx);
return 0;
}
@@ -18475,9 +18478,103 @@ enum {
INSN_IDX_UPDATED = 2,
};
-static int process_bpf_exit_full(struct bpf_verifier_env *env,
- bool *do_print_state,
- bool exception_exit)
+static u32 unwind_pop_frame(struct bpf_verifier_env *env)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ struct bpf_func_state *callee = state->frame[state->curframe];
+ u32 callsite = callee->callsite;
+ struct bpf_func_state *caller;
+
+ caller = state->frame[state->curframe - 1];
+ account_processed_insns(env, callee, caller);
+ free_func_state(callee);
+ state->frame[state->curframe--] = NULL;
+ invalidate_outgoing_stack_args(env, caller);
+ return callsite;
+}
+
+static void unwind_enter_pad(struct bpf_verifier_env *env)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ struct bpf_func_state *frame = cur_func(env);
+
+ state->unwind_frameno = state->curframe;
+ clear_caller_saved_regs(env, frame->regs);
+ mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ __mark_reg_known(&frame->regs[BPF_REG_0], BPF_PAD_ENTRY_R0);
+}
+
+static int unwind_finish(struct bpf_verifier_env *env)
+{
+ int err = check_resource_leak(env, true, true, "bpf_throw");
+
+ if (err)
+ return err;
+ return PROCESS_BPF_EXIT;
+}
+
+static int unwind_step(struct bpf_verifier_env *env, u32 callsite, int *insn_idx)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ u8 popped = 0;
+
+ state->unwinding = true;
+ for (;;) {
+ int pad = bpf_cleanup_pad_of_call(env, callsite);
+
+ if (pad >= 0) {
+ unwind_enter_pad(env);
+ /*
+ * The edge from @callsite to the pad crosses @popped
+ * frames, and nothing in the instruction stream says
+ * so. Record it for mark_chain_precision(), which has
+ * to walk back through the same frames.
+ */
+ env->unwind_frames = popped;
+ *insn_idx = pad;
+ return INSN_IDX_UPDATED;
+ }
+ if (!state->curframe)
+ return unwind_finish(env);
+ callsite = unwind_pop_frame(env);
+ popped++;
+ }
+}
+
+static int process_cleanup_resume(struct bpf_verifier_env *env, int *insn_idx)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ int err;
+
+ /* A pad entered by ordinary control flow. */
+ if (!state->unwinding) {
+ verbose(env,
+ "bpf_unwind_resume() at insn %d reached without an exception in flight\n",
+ *insn_idx);
+ return -EINVAL;
+ }
+ /*
+ * A resume in a subprogram the pad called, which has a landing pad of
+ * its own and reached it by ordinary control flow. A JIT lowers a
+ * resume as the way back out of a pad, which reaches the walker only
+ * from the frame whose pad it called.
+ */
+ if (state->curframe != state->unwind_frameno) {
+ verbose(env,
+ "bpf_unwind_resume() at insn %d is in frame %d, not frame %d whose landing pad the exception entered\n",
+ *insn_idx, state->curframe, state->unwind_frameno);
+ return -EINVAL;
+ }
+ if (!state->curframe)
+ return unwind_finish(env);
+ err = unwind_step(env, unwind_pop_frame(env), insn_idx);
+ /* unwind_step() counted the frames it popped, not this one. */
+ if (err == INSN_IDX_UPDATED)
+ env->unwind_frames++;
+ return err;
+}
+
+static int process_bpf_exit_full(struct bpf_verifier_env *env, bool *do_print_state)
{
struct bpf_func_state *cur_frame = cur_func(env);
@@ -18487,25 +18584,11 @@ static int process_bpf_exit_full(struct bpf_verifier_env *env,
* for which reference_state must match caller reference
* state when it exits.
*/
- int err = check_resource_leak(env, exception_exit,
- exception_exit || !env->cur_state->curframe,
- exception_exit ? "bpf_throw" :
+ int err = check_resource_leak(env, false, !env->cur_state->curframe,
"BPF_EXIT instruction in main prog");
if (err)
return err;
- /* The side effect of the prepare_func_exit which is
- * being skipped is that it frees bpf_func_state.
- * Typically, process_bpf_exit will only be hit with
- * outermost exit. copy_verifier_state in pop_stack will
- * handle freeing of any extra bpf_func_state left over
- * from not processing all nested function exits. We
- * also skip return code checks as they are not needed
- * for exceptional exits.
- */
- if (exception_exit)
- return PROCESS_BPF_EXIT;
-
if (env->cur_state->curframe) {
/* exit from nested function */
err = prepare_func_exit(env, &env->insn_idx);
@@ -18679,6 +18762,8 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
env->jmps_processed++;
if (opcode == BPF_CALL) {
+ if (bpf_is_unwind_resume_kfunc(insn))
+ return process_cleanup_resume(env, &env->insn_idx);
if (env->cur_state->active_locks) {
if ((insn->src_reg == BPF_REG_0 &&
insn->imm != BPF_FUNC_spin_unlock &&
@@ -18712,7 +18797,7 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
env->insn_idx += insn->imm + 1;
return INSN_IDX_UPDATED;
} else if (opcode == BPF_EXIT) {
- return process_bpf_exit_full(env, do_print_state, false);
+ return process_bpf_exit_full(env, do_print_state);
}
return check_cond_jmp_op(env, insn, &env->insn_idx);
}
@@ -18749,10 +18834,14 @@ static int do_check(struct bpf_verifier_env *env)
for (;;) {
struct bpf_insn *insn;
struct bpf_insn_aux_data *insn_aux;
+ u8 unwind_frames;
int err;
/* reset current history entry on each new instruction */
env->cur_hist_ent = NULL;
+ /* frames the unwind popped to reach this insn, if it is a pad */
+ unwind_frames = env->unwind_frames;
+ env->unwind_frames = 0;
env->prev_insn_idx = prev_insn_idx;
if (env->insn_idx >= insn_cnt) {
@@ -18816,10 +18905,16 @@ static int do_check(struct bpf_verifier_env *env)
}
}
- if (bpf_is_jmp_point(env, env->insn_idx)) {
+ /*
+ * A landing pad is a jump point, but record the edge even if
+ * that ever stops being true: it is the only place the frames
+ * the unwind popped are written down.
+ */
+ if (bpf_is_jmp_point(env, env->insn_idx) || unwind_frames) {
err = bpf_push_jmp_history(env, state, 0, 0, 0, 0);
if (err)
return err;
+ env->cur_hist_ent->unwind_frames = unwind_frames;
}
if (signal_pending(current))
--
2.53.0-Meta
next prev parent reply other threads:[~2026-09-20 5:43 UTC|newest]
Thread overview: 45+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-20 5:42 [PATCH bpf-next v3 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
2026-09-20 5:42 ` [PATCH bpf-next v3 01/20] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-20 5:56 ` sashiko-bot
2026-09-21 15:43 ` Yonghong Song
2026-09-20 6:32 ` bot+bpf-ci
2026-09-21 15:28 ` Yonghong Song
2026-09-20 5:42 ` [PATCH bpf-next v3 02/20] bpf: Add the bpf_unwind_resume() kfunc Yonghong Song
2026-09-20 5:42 ` [PATCH bpf-next v3 03/20] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-20 6:31 ` bot+bpf-ci
2026-09-21 14:06 ` Yonghong Song
2026-09-20 5:42 ` [PATCH bpf-next v3 04/20] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-09-20 5:42 ` [PATCH bpf-next v3 05/20] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-20 5:42 ` [PATCH bpf-next v3 06/20] bpf: Explore the landing pads no call site reaches Yonghong Song
2026-09-20 5:43 ` [PATCH bpf-next v3 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch Yonghong Song
2026-09-20 6:32 ` bot+bpf-ci
2026-09-21 14:08 ` Yonghong Song
2026-09-20 5:43 ` Yonghong Song [this message]
2026-09-20 5:43 ` [PATCH bpf-next v3 09/20] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
2026-09-20 5:43 ` [PATCH bpf-next v3 10/20] bpf: Dispatch exception cleanup pads from bpf_throw() Yonghong Song
2026-09-20 6:32 ` bot+bpf-ci
2026-09-21 15:55 ` Yonghong Song
2026-09-20 5:43 ` [PATCH bpf-next v3 11/20] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-20 5:43 ` [PATCH bpf-next v3 12/20] bpf, arm64: " Yonghong Song
2026-09-20 5:43 ` [PATCH bpf-next v3 13/20] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-20 5:43 ` [PATCH bpf-next v3 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-20 6:31 ` bot+bpf-ci
2026-09-21 15:00 ` Yonghong Song
2026-09-20 5:43 ` [PATCH bpf-next v3 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-20 5:43 ` [PATCH bpf-next v3 16/20] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-20 6:01 ` sashiko-bot
2026-09-21 14:14 ` Yonghong Song
2026-09-20 5:43 ` [PATCH bpf-next v3 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-20 6:00 ` sashiko-bot
2026-09-21 14:20 ` Yonghong Song
2026-09-20 6:31 ` bot+bpf-ci
2026-09-21 14:21 ` Yonghong Song
2026-09-20 5:43 ` [PATCH bpf-next v3 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
2026-09-20 5:59 ` sashiko-bot
2026-09-20 5:44 ` [PATCH bpf-next v3 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
2026-09-20 6:01 ` sashiko-bot
2026-09-20 6:46 ` bot+bpf-ci
2026-09-21 14:27 ` Yonghong Song
2026-09-20 5:44 ` [PATCH bpf-next v3 20/20] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
2026-09-20 6:31 ` bot+bpf-ci
2026-09-21 14:30 ` Yonghong Song
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260920054307.868442-1-yonghong.song@linux.dev \
--to=yonghong.song@linux.dev \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=kernel-team@fb.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox