* [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds
@ 2026-09-29 0:16 Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
` (21 more replies)
0 siblings, 22 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:16 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
bpf_throw() walks the BPF call stack to the exception boundary and
discards every frame in between. A frame that owns something -- an RCU
read lock, a preemption-disabled section, a referenced kptr -- never gets
to give it back, so the verifier refuses to let such a frame throw at
all. That is the whole reason a Rust program cannot use bpf_throw() as
its panic path right now: Rust's Drop glue *is* that give-back, and there
is nowhere to run it.
LLVM 23 added the compiler half ([1]). A Rust function that owns a value
across a call that can unwind
fn foo() {
let _guard = RcuReadGuard::new(); /* bpf_rcu_read_lock() */
may_throw(); /* extern "C-unwind" */
} /* Drop: rcu_read_unlock */
lowers to an invoke with a cleanup landing pad holding the Drop call, and
the BPF backend writes one record per invoke region into a .bpf_cleanup
section: a flat table of 12-byte (begin, end, landing_pad) triples, each
field a byte offset into the code section. The rule is "a frame suspended
at a call in [begin, end) resumes at landing_pad when an unwind passes
through". A pad ends with a call to _Unwind_Resume(), which the kernel
provides as the bpf_unwind_resume() kfunc.
Rather than overload bpf_throw(), which keeps its own meaning -- leave for
the exception boundary with the frames in between discarded -- a new
bpf_unwind() kfunc raises the unwind this series dispatches.
This series is the kernel half: take that table at BPF_PROG_LOAD, teach
the verifier that a covered call can also go to its landing pad, and have
bpf_unwind() run the pads as it walks.
C has no unwinding, so the selftests spell out by hand what a frontend
emits -- a call site bracketed by two labels, a landing pad, and a record
tying them together. The frame above, written that way:
"call bpf_rcu_read_lock;"
"1:" "call foo3;" /* cleanup region */
"2:"
... normal path, ends in bpf_rcu_read_unlock ...
"6:" /* landing pad */
"call bpf_rcu_read_unlock;"
"call bpf_unwind_resume;"
CLEANUP_REC("1b", "2b", "6b")
Design
======
A pad runs in the frame that owns it, entered by an ordinary return.
bpf_unwind() walks the frames with arch_bpf_stack_walk_ra(), which hands
out the slot each frame's return address came from, and rewrites that
slot: to the frame's landing pad where a record covers the call it is
suspended at, and to the frame's epilogue where nothing does. Then it
returns.
Each frame therefore runs its own pad, on its own stack, and leaves
through its own epilogue -- which is what puts its caller's r6-r9 back. The
unwind needs no trampoline, no spill area and no per-frame metadata beyond
the table itself. The fixups lower a pad's bpf_unwind_resume() to
"r0 = 0; exit", so the frame returns and the address rewritten below it
carries the unwind on to the next pad.
The verifier sees the same shape with no new machinery. A covered call
gets a second successor, its landing pad, in the same frame, and its entry
state is the state at the call with the caller-saved registers gone, since
the callee's epilogue puts them back on the way out. Nothing crosses a
frame boundary that the instruction stream does not already describe, so
precision, liveness and the resource checks work there as they do on any
other branch.
1 pack bpf_insn_aux_data's flags into bit fields, so the flags this
series adds cost bits rather than bytes
2-3 uapi: cleanup_info in BPF_PROG_LOAD and struct bpf_cleanup_info;
the bpf_unwind() and bpf_unwind_resume() kfuncs
4-10 verifier: mark the covered call sites, give each an edge to its
pad, require an unwind to leave a frame as it found it, and refuse
the shapes that cannot be dispatched
11 bpf_unwind(): rewrite the return addresses as it walks
12-13 x86-64 and arm64 JITs
14-18 libbpf: collect .bpf_cleanup, pass it to the kernel, resolve
_Unwind_Resume, carry it through the light skeleton and the
static linker
19-22 selftests
Limitations
===========
- Cleanup pads only. A catch pad -- one that ends in a plain exit rather
than a resume, which is what Rust's catch_unwind would need -- is
refused: bpf_unwind() rewrites every frame's return address in one
pass, so a frame above a catch pad would resume at a pad for an unwind
that had already been caught. LLVM refuses type-specific catches and
filters on its side as well.
- A JIT that can dispatch pads is required: x86-64 (with
CONFIG_UNWINDER_ORC, which bpf_throw() already needs there) and arm64.
Anywhere else the load fails with -EOPNOTSUPP rather than silently
doing nothing.
- No offloaded programs, no private stack, and no combining a table with
an exception callback.
- In the pad's own frame: no tail call, no BPF_LD_[ABS|IND] and no
indirect jump, none of which reaches the resume. And no second unwind
anywhere above a pad, including a call to a global subprogram that can
raise one.
- The Rust toolchain does not properly support BPF exception handling
yet. The tables the selftests use are hand-written inline asm, which
the assembler turns into the same relocations the BPF AsmPrinter emits,
so libbpf and the kernel see an object indistinguishable from a
compiler-generated one.
[1] https://github.com/llvm/llvm-project/pull/192164
llvm commit 9d51c891b719 ("[BPF] Add exception handling support
with .bpf_cleanup section")
Changelog
=========
v6 -> v7:
- v6: https://lore.kernel.org/bpf/20260926050006.2213110-1-yonghong.song@linux.dev/
- New patch: require an unwind to leave a frame holding what it entered
with, including at a call it passes through with no record over it.
- A resume ends the verifier's path instead of walking back to the
instruction after the call, where an unwind never returns.
- Refuse a private stack for any program that can unwind, not only one
carrying a table.
- Refuse a table for a program whose verifier_ops plants an epilogue,
and scan the whole instruction stream for bpf_throw().
- Keep the last exit of every subprogram an unwind can pass through,
not only the one after a bpf_unwind() call, and search for it.
- Fix precision backtracking across a resume, push a landing pad by
itself so the CFG walk's DFS invariant holds, and mark callx sites.
- x86: leave a frame whose return a tracer has hooked alone; no ENDBR
at a pad head, which is only ever reached by a return.
- arm64: no BTI at a pad head; ask the build whether a return address
is signed; refuse cleanup pads where a shadow call stack is in use.
- Answer a speculative walk reaching a pad with a barrier, not a refusal.
- selftests: add the negative shapes for the above, allow repeated
__set_global()/__ret_global(), and drop duplicate accepted shapes.
v5 -> v6:
- v5: https://lore.kernel.org/bpf/20260923045846.2414643-1-yonghong.song@linux.dev/
- Run a pad in the frame that owns it: bpf_unwind() rewrites each
frame's saved return address rather than calling the pad as a
subroutine of the walker. The bpf_cleanup_pad.S trampolines, the
per-frame spill area and the pad-entry register header all go away.
- Raise the unwind with a new bpf_unwind() kfunc, so that bpf_throw()
keeps its meaning.
- Add arch_bpf_stack_walk_ra(), which also hands out the slot a return
address came from. arm64 re-signs the address it writes there.
- A pad is an ordinary second successor of a covered call, so the
verifier needs no unwind edge of its own: the cross-frame precision
and liveness work is gone, with the v5 fixes it needed, and two
patches become one.
- Refuse the undispatchable shapes per instruction in do_check(), keyed
on a per-frame mark, rather than by walking each pad.
- Allow on-stack call arguments in a pad, which now has its own frame.
- Add __set_global() and __ret_global() test tags, so RUN_TESTS() drives
the shapes and test_shapes() goes away (suggested by Eduard).
- Drop the shapes that needed a driver of their own, and the three
extension objects with them.
v4 -> v5:
- v4: https://lore.kernel.org/bpf/20260921210033.1715000-1-yonghong.song@linux.dev/
- Rebase on bpf-next.
- Clear r0 where a throw enters a landing pad: no instruction defines
it, so a precision request for it outlived the state and oopsed the
verifier.
- Defer entering the frames an unwind edge crossed until the backtrack
reads an instruction from them; entering at the landing pad could
leave bt->frame past the parent state's frames and oops the verifier.
- Stamp the popped frame count inside bpf_push_jmp_history(), so a pad's
entry carries it even when the prune path is what creates the entry.
- Replace the hand-rolled CFG traversal and its separate pass with
per-instruction checks in do_check(), keyed on the verifier's
unwinding state; kernel/bpf/exception.c halves.
- Add a first patch packing bpf_insn_aux_data's flags into one bit field
word, 144 bytes to 128, so the flags this series adds cost bits.
- Drop the per-subprogram arrays the JITs consulted for throw sites,
resume sites and pad bodies, and read insn_aux_data, which a JIT
already has.
- Drop the pad-entry r0 header: r0 at a pad is an unknown scalar, and a
dispatcher writes a defined value there only to keep a kernel one out
of BPF.
- Take the bool arguments back out of verifier_remove_insns() and the
site collector, and share pop_frame() with prepare_func_exit().
- Rename the recorded call sites to throw_call and resume_call, give the
exported functions a bpf_exc_ prefix, and drop cleanup_ from the
statics.
v3 -> v4:
- v3: https://lore.kernel.org/bpf/20260920054225.864535-1-yonghong.song@linux.dev/
- Rebase on bpf-next due to conflict.
- Reserve the throw-site spill area only in a (sub)program that calls
bpf_throw(): 40 bytes of stack per frame on x86-64, 80 on arm64.
- Bound the record count by the number of instructions in the program,
and name both in the message.
- Refuse a .bpf_cleanup section in libbpf whose record count cannot be
handed to the kernel as a count times a record size in an int.
- Fix the static linker's new bounds check, which could itself wrap, and
refuse a section too small to hold one field.
- WARN once if the body of bpf_unwind_resume() is ever reached, the way
bpf_throw() does where its exception callback should never return.
- Rename nr_pad_body to pad_body_bits, use BTF_ID_LIST_SINGLE, move
cleanup_pad out of insn_aux_data's bools, drop an arm64 include.
v2 -> v3:
- v2: https://lore.kernel.org/bpf/20260918044156.3283973-1-yonghong.song@linux.dev/
- Keep a landing pad's record when opt_remove_nops() deletes a pad that is
a nop, instead of dropping it after the verifier has already checked the
call site against it.
- Refuse a BPF_LD_[ABS|IND] in a pad body.
- Teach mark_chain_precision() about the throw-to-pad edge, which crosses
frames with no instruction to account for them.
- Refuse a bpf_unwind_resume() in any frame but the one whose landing pad
the walker entered.
- Check raw_data, alignment and bounds before the static linker writes
through a relocation in a non-executable section, which may be SHT_NOBITS.
v1 -> v2:
- v1: https://lore.kernel.org/bpf/20260917055645.3926444-1-yonghong.song@linux.dev/
- Consolidate all usages of kern_extern_name() in a single patch in libbpf.
- Avoid compiler warning and add proper cleanup_info_cnt guard in libbpf when
collecting .bpf_cleanup records.
- Add cleanup_info_cnt condition for emit_rel_store() with cleanup_info.
Yonghong Song (22):
bpf: Pack bpf_insn_aux_data flags into bit fields
bpf: Accept the compiler's exception cleanup table at program load
bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs
bpf: Add lookups for exception cleanup resumes and landing pads
bpf: Prepare for an exception cleanup table before the CFG walk
bpf: Make exception landing pads reachable in the CFG
bpf: Resume a covered call at its landing pad
bpf: Require an unwind to leave a frame holding what it entered with
bpf: Refuse a landing pad that does not resume
bpf: Refuse a private stack for a program that can unwind
bpf: Dispatch cleanup pads by rewriting return addresses
bpf, x86: Dispatch exception cleanup pads at run time
bpf, arm64: Dispatch exception cleanup pads at run time
libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
libbpf: Add cleanup_info to bpf_prog_load_opts
libbpf: Collect .bpf_cleanup records and pass them to the kernel
libbpf: Carry the exception cleanup table through the light skeleton
libbpf: Let the static linker carry .bpf_cleanup relocations
selftests/bpf: Add end-to-end and negative .bpf_cleanup exception
tests
selftests/bpf: Add __set_global() and __ret_global() test tags
selftests/bpf: Cover more accepted .bpf_cleanup exception shapes
selftests/bpf: Load an exception cleanup program from a light skeleton
arch/arm64/kernel/stacktrace.c | 103 ++
arch/arm64/net/bpf_jit_comp.c | 26 +
arch/x86/net/bpf_jit_comp.c | 42 +
include/linux/bpf.h | 37 +
include/linux/bpf_verifier.h | 70 +-
include/linux/filter.h | 3 +
include/uapi/linux/bpf.h | 9 +
kernel/bpf/Makefile | 2 +-
kernel/bpf/backtrack.c | 42 +
kernel/bpf/cfg.c | 49 +
kernel/bpf/check_btf.c | 134 +++
kernel/bpf/core.c | 25 +-
kernel/bpf/exception.c | 354 +++++++
kernel/bpf/exception.h | 30 +
kernel/bpf/fixups.c | 161 +++-
kernel/bpf/helpers.c | 60 ++
kernel/bpf/liveness.c | 21 +
kernel/bpf/syscall.c | 2 +-
kernel/bpf/verifier.c | 205 +++-
tools/include/uapi/linux/bpf.h | 9 +
tools/lib/bpf/bpf.c | 6 +-
tools/lib/bpf/bpf.h | 7 +-
tools/lib/bpf/gen_loader.c | 29 +-
tools/lib/bpf/libbpf.c | 333 ++++++-
tools/lib/bpf/libbpf_internal.h | 10 +
tools/lib/bpf/linker.c | 36 +-
tools/testing/selftests/bpf/Makefile.skel | 2 +-
.../selftests/bpf/exceptions_cleanup.h | 42 +
.../bpf/prog_tests/exceptions_cleanup.c | 115 +++
tools/testing/selftests/bpf/progs/bpf_misc.h | 7 +
.../selftests/bpf/progs/exceptions_cleanup.c | 162 ++++
.../bpf/progs/exceptions_cleanup_fail.c | 878 ++++++++++++++++++
.../bpf/progs/exceptions_cleanup_light.c | 39 +
.../bpf/progs/exceptions_cleanup_shapes.c | 607 ++++++++++++
tools/testing/selftests/bpf/test_loader.c | 314 ++++++-
35 files changed, 3924 insertions(+), 47 deletions(-)
create mode 100644 kernel/bpf/exception.c
create mode 100644 kernel/bpf/exception.h
create mode 100644 tools/testing/selftests/bpf/exceptions_cleanup.h
create mode 100644 tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
@ 2026-09-29 0:16 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
` (20 subsequent siblings)
21 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:16 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
struct bpf_insn_aux_data spreads its flags over three places: eight bools a
byte each, a u8 carrying three bit fields, and a u32 group with 25 of its
32 bits spare. Put them all in one u64 word, alu_state with them, and move
orig_idx below it. The structure goes from 128 bytes to 120, with no byte
holes and 33 of the word's 64 bits spare. Later patches in this series take
three of those bits, and add a u32 of their own that puts the size back at
128.
No functional change: a one-bit unsigned field holds 0 and 1 the way the
bool did, and nothing takes the address of any of them.
Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 43 ++++++++++++++++++------------------
1 file changed, 22 insertions(+), 21 deletions(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index c775bd757706..d85cf969bcb0 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -664,41 +664,42 @@ struct bpf_insn_aux_data {
u64 map_key_state; /* constant (32 bit) key tracking for maps */
int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
- bool nospec; /* do not execute this instruction speculatively */
- bool nospec_result; /* result is unsafe under speculation, nospec must follow */
- bool zext_dst; /* this insn zero extends dst reg */
- bool needs_zext; /* alu op needs to clear upper bits */
- bool prevent_zext; /* alu op cannot be zext (already used with 64-bit scalars) */
- bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
- bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
- bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
- u8 alu_state; /* used in combination with alu_limit */
+ u64 nospec:1; /* do not execute this instruction speculatively */
+ u64 nospec_result:1; /* result is unsafe under speculation, nospec must follow */
+ u64 zext_dst:1; /* this insn zero extends dst reg */
+ u64 needs_zext:1; /* alu op needs to clear upper bits */
+ u64 prevent_zext:1; /* alu op cannot be zext (already used with 64-bit scalars) */
+ u64 non_sleepable:1; /* helper/kfunc may be called from non-sleepable context */
+ u64 is_iter_next:1; /* bpf_iter_<type>_next() kfunc call */
+ u64 call_with_percpu_alloc_ptr:1; /* {this,per}_cpu_ptr() with prog percpu alloc */
+ u64 alu_state:8; /* used in combination with alu_limit */
/* true if STX or LDX instruction is a part of a spill/fill
* pattern for a bpf_fastcall call.
*/
- u8 fastcall_pattern:1;
+ u64 fastcall_pattern:1;
/* for CALL instructions, a number of spill/fill pairs in the
* bpf_fastcall pattern.
*/
- u8 fastcall_spills_num:3;
- u8 arg_prog:4;
+ u64 fastcall_spills_num:3;
+ u64 arg_prog:4;
- /* below fields are initialized once */
- unsigned int orig_idx; /* original instruction index */
- u32 jmp_point:1;
- u32 prune_point:1;
+ /* below flags are initialized once */
+ u64 jmp_point:1;
+ u64 prune_point:1;
/* ensure we check state equivalence and save state checkpoint and
* this instruction, regardless of any heuristics
*/
- u32 force_checkpoint:1;
+ u64 force_checkpoint:1;
/* true if instruction is a call to a helper function that
* accepts callback function as a parameter.
*/
- u32 calls_callback:1;
- u32 indirect_target:1; /* if it is an indirect jump target */
- u32 non_stack_access:1; /* instruction can access non-stack memory */
+ u64 calls_callback:1;
+ u64 indirect_target:1; /* if it is an indirect jump target */
+ u64 non_stack_access:1; /* instruction can access non-stack memory */
/* true if some jump or call instruction targets this instruction */
- u32 jump_target:1;
+ u64 jump_target:1;
+
+ unsigned int orig_idx; /* original instruction index, initialized once */
/*
* CFG strongly connected component this instruction belongs to,
* zero if it is a singleton SCC.
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 02/22] bpf: Accept the compiler's exception cleanup table at program load
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
@ 2026-09-29 0:16 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
` (19 subsequent siblings)
21 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:16 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
LLVM 23 added exception handling support for BPF with the .bpf_cleanup
section ([1]). Rust code compiled with panic=unwind runs cleanup code
(Drop glue) when an unwind passes through, and the LLVM BPF backend emits
that section from the landing pads the frontend produced. Plain C cannot
generate .bpf_cleanup unless inline asm is used. The Rust compiler does not
*properly* support BPF exception handling yet, but the kernel can support
the table today, and inline assembly is enough to test it.
Add the UAPI to carry the .bpf_cleanup table into the kernel. BPF_PROG_LOAD
grows cleanup_info, cleanup_info_cnt and cleanup_info_rec_size, and struct
bpf_cleanup_info describes one record as a triple of instruction indices:
the half-open call-site range [begin_off, end_off) and the landing_pad_off
the frame resumes at. The table arrives sorted by begin_off, with disjoint
ranges and each record's three offsets inside one subprogram;
check_cleanup_info() holds it to that at load time. Nothing reads it yet;
the patches that follow -- the CFG walk, the unwind walk and the JITs --
are its consumers.
Link: https://github.com/llvm/llvm-project/pull/192164 [1]
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 2 +
include/uapi/linux/bpf.h | 9 +++
kernel/bpf/check_btf.c | 134 +++++++++++++++++++++++++++++++++
kernel/bpf/syscall.c | 2 +-
kernel/bpf/verifier.c | 1 +
tools/include/uapi/linux/bpf.h | 9 +++
6 files changed, 156 insertions(+), 1 deletion(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index d85cf969bcb0..6ce25c96ebed 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1018,6 +1018,8 @@ struct bpf_verifier_env {
struct spill_snapshot **callsite_at_stack;
u32 pass_cnt; /* number of times do_check() was called */
u32 subprog_cnt;
+ struct bpf_cleanup_info *cleanup_info;
+ u32 cleanup_info_cnt;
/* number of instructions analyzed by the verifier */
u32 prev_insn_processed, insn_processed;
/* number of jmps, calls, exits analyzed so far */
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 4687c3310996..aca43f4f927f 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -1702,6 +1702,9 @@ union bpf_attr {
* verification.
*/
__s32 keyring_id;
+ __aligned_u64 cleanup_info; /* exception cleanup table */
+ __u32 cleanup_info_rec_size; /* userspace bpf_cleanup_info size */
+ __u32 cleanup_info_cnt; /* number of bpf_cleanup_info records */
};
struct { /* anonymous struct used by BPF_OBJ_* commands */
@@ -7638,6 +7641,12 @@ struct bpf_line_info {
__u32 line_col;
};
+struct bpf_cleanup_info {
+ __u32 begin_off;
+ __u32 end_off;
+ __u32 landing_pad_off;
+};
+
struct bpf_spin_lock {
__u32 val;
};
diff --git a/kernel/bpf/check_btf.c b/kernel/bpf/check_btf.c
index 4c1ed842f661..5c62f350d500 100644
--- a/kernel/bpf/check_btf.c
+++ b/kernel/bpf/check_btf.c
@@ -407,6 +407,136 @@ int bpf_check_core_relo(struct bpf_verifier_env *env,
return err;
}
+#define MIN_BPF_CLEANUP_INFO_SIZE 12
+#define MAX_CLEANUP_INFO_REC_SIZE MAX_FUNCINFO_REC_SIZE
+
+static int check_cleanup_info(struct bpf_verifier_env *env,
+ const union bpf_attr *attr,
+ bpfptr_t uattr)
+{
+ u32 krec_size = sizeof(struct bpf_cleanup_info);
+ u32 i, nrec, urec_size, min_size, prev_end = 0;
+ struct bpf_cleanup_info *krecord;
+ bpfptr_t urecord;
+ int ret = -EINVAL;
+
+ nrec = attr->cleanup_info_cnt;
+ if (!nrec)
+ return 0;
+ if (nrec > env->prog->len) {
+ verbose(env, "cleanup info has %u records for %u instructions\n",
+ nrec, env->prog->len);
+ return -EINVAL;
+ }
+
+ urec_size = attr->cleanup_info_rec_size;
+ if (urec_size < MIN_BPF_CLEANUP_INFO_SIZE ||
+ urec_size > MAX_CLEANUP_INFO_REC_SIZE ||
+ urec_size % sizeof(u32)) {
+ verbose(env, "invalid cleanup info rec size %u\n", urec_size);
+ return -EINVAL;
+ }
+
+ krecord = kvcalloc(nrec, krec_size, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!krecord)
+ return -ENOMEM;
+
+ min_size = min_t(u32, krec_size, urec_size);
+ urecord = make_bpfptr(attr->cleanup_info, uattr.is_kernel);
+ for (i = 0; i < nrec; i++) {
+ struct bpf_subprog_info *sb, *se, *sl;
+ struct bpf_cleanup_info *rec = &krecord[i];
+
+ ret = bpf_check_uarg_tail_zero(urecord, krec_size, urec_size);
+ if (ret) {
+ if (ret == -E2BIG) {
+ verbose(env, "nonzero tailing record in cleanup info\n");
+ if (copy_to_bpfptr_offset(uattr,
+ offsetof(union bpf_attr,
+ cleanup_info_rec_size),
+ &min_size, sizeof(min_size)))
+ ret = -EFAULT;
+ }
+ goto err_free;
+ }
+
+ if (copy_from_bpfptr(rec, urecord, min_size)) {
+ ret = -EFAULT;
+ goto err_free;
+ }
+ bpfptr_add(&urecord, urec_size);
+
+ ret = -EINVAL;
+ if (rec->begin_off >= rec->end_off) {
+ verbose(env, "cleanup_info[%u]: begin %u >= end %u\n",
+ i, rec->begin_off, rec->end_off);
+ goto err_free;
+ }
+ if (i && rec->begin_off < prev_end) {
+ verbose(env,
+ "cleanup_info[%u]: range [%u,%u) is unsorted or overlaps the previous record\n",
+ i, rec->begin_off, rec->end_off);
+ goto err_free;
+ }
+ prev_end = rec->end_off;
+
+ sb = bpf_find_containing_subprog(env, rec->begin_off);
+ se = bpf_find_containing_subprog(env, rec->end_off - 1);
+ sl = bpf_find_containing_subprog(env, rec->landing_pad_off);
+ if (!sb || !se || !sl) {
+ verbose(env, "cleanup_info[%u]: offset out of range\n", i);
+ goto err_free;
+ }
+ if (sb != se || sb != sl) {
+ verbose(env,
+ "cleanup_info[%u]: range/landing pad span multiple subprogs\n",
+ i);
+ goto err_free;
+ }
+ /*
+ * A zero opcode is the second half of a 16-byte insn, not an
+ * insn. end_off is exclusive, so it may be one past the last.
+ */
+ if (!env->prog->insnsi[rec->begin_off].code ||
+ !env->prog->insnsi[rec->landing_pad_off].code ||
+ (rec->end_off < env->prog->len &&
+ !env->prog->insnsi[rec->end_off].code)) {
+ verbose(env, "cleanup_info[%u]: points at invalid insn\n", i);
+ goto err_free;
+ }
+ }
+
+ /* Reject a landing pad inside any call-site range, its own included. */
+ ret = -EINVAL;
+ for (i = 0; i < nrec; i++) {
+ u32 pad = krecord[i].landing_pad_off;
+ u32 l = 0, r = nrec;
+
+ while (l < r) {
+ u32 m = l + (r - l) / 2;
+
+ if (pad < krecord[m].begin_off) {
+ r = m;
+ } else if (pad >= krecord[m].end_off) {
+ l = m + 1;
+ } else {
+ verbose(env,
+ "cleanup_info[%u]: landing pad %u is inside the call-site range of cleanup_info[%u]\n",
+ i, pad, m);
+ goto err_free;
+ }
+ }
+ }
+
+ env->cleanup_info = krecord;
+ env->cleanup_info_cnt = nrec;
+ return 0;
+
+err_free:
+ kvfree(krecord);
+ return ret;
+}
+
int bpf_prepare_btf_info(struct bpf_verifier_env *env,
const union bpf_attr *attr,
bpfptr_t uattr)
@@ -441,6 +571,10 @@ int bpf_check_btf_info(struct bpf_verifier_env *env,
{
int err;
+ err = check_cleanup_info(env, attr, uattr);
+ if (err)
+ return err;
+
if (!attr->func_info_cnt && !attr->line_info_cnt) {
if (check_abnormal_return(env))
return -EINVAL;
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index ac52f4ae414c..0e14afe3fdc5 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -2924,7 +2924,7 @@ int __init __used bpf_multi_func(void) { return 0; }
BTF_ID_LIST_GLOBAL_SINGLE(bpf_multi_func_btf_id, func, bpf_multi_func)
/* last field in 'union bpf_attr' used by this command */
-#define BPF_PROG_LOAD_LAST_FIELD keyring_id
+#define BPF_PROG_LOAD_LAST_FIELD cleanup_info_cnt
static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, struct bpf_log_attr *attr_log)
{
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 03dbc0e00398..6d3408f295ed 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22801,6 +22801,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
kvfree(env->callx_edges);
kvfree(env->func_ptrs);
bpf_diag_free(env);
+ kvfree(env->cleanup_info);
kvfree(env);
return ret;
}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 4687c3310996..aca43f4f927f 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -1702,6 +1702,9 @@ union bpf_attr {
* verification.
*/
__s32 keyring_id;
+ __aligned_u64 cleanup_info; /* exception cleanup table */
+ __u32 cleanup_info_rec_size; /* userspace bpf_cleanup_info size */
+ __u32 cleanup_info_cnt; /* number of bpf_cleanup_info records */
};
struct { /* anonymous struct used by BPF_OBJ_* commands */
@@ -7638,6 +7641,12 @@ struct bpf_line_info {
__u32 line_col;
};
+struct bpf_cleanup_info {
+ __u32 begin_off;
+ __u32 end_off;
+ __u32 landing_pad_off;
+};
+
struct bpf_spin_lock {
__u32 val;
};
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
@ 2026-09-29 0:16 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
` (18 subsequent siblings)
21 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:16 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
An exception cleanup needs two terminators, and the kernel provides both as
kfuncs so that a BPF program can name them.
bpf_unwind() begins an unwind: the frames between it and whatever catches
run their landing pads on the way out. It is separate from bpf_throw(),
which keeps its own meaning -- leaving for the exception boundary with the
frames in between discarded.
bpf_unwind_resume() ends a landing pad, carrying the unwind on once that
frame's cleanups have run. The compiler names it _Unwind_Resume, the base
unwind ABI's entry point for the same thing; a later libbpf patch resolves
that name to this one. Every unwind ABI hands _Unwind_Resume the exception
object, and LLVM emits that argument on BPF too. Nothing in the kernel
needs it now, but the kfunc takes it as ptr__ign so the prototype matches
the call the compiler makes.
Both are defined here and neither is registered with any program type yet.
A call to either is only valid in a particular place -- an unwind where a
record covers it, a resume inside a landing pad -- and neither does
anything until a pad can be dispatched at run time, so registration waits
for the patch that adds that.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/helpers.c | 14 ++++++++++++++
1 file changed, 14 insertions(+)
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index a284f20c97d5..08aee86a155c 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -3424,6 +3424,10 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
return false;
}
+__bpf_kfunc void bpf_unwind(void)
+{
+}
+
__bpf_kfunc void bpf_throw(u64 cookie)
{
struct bpf_throw_ctx ctx = {};
@@ -3445,6 +3449,16 @@ __bpf_kfunc void bpf_throw(u64 cookie)
WARN(1, "A call to BPF exception callback should never return\n");
}
+__bpf_kfunc void bpf_unwind_resume(void *ptr__ign)
+{
+ /*
+ * Never reached: the verifier accepts this kfunc only as a frame
+ * terminator and do_misc_fixups() lowers every one of them to
+ * 'r0 = 0; exit', so no call to this body survives to run.
+ */
+ WARN_ONCE(1, "exception cleanup resume was not lowered to a return\n");
+}
+
__bpf_kfunc int bpf_wq_init(struct bpf_wq *wq, void *p__const_map, unsigned int flags)
{
struct bpf_async_kern *async = (struct bpf_async_kern *)wq;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 04/22] bpf: Add lookups for exception cleanup resumes and landing pads
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (2 preceding siblings ...)
2026-09-29 0:16 ` [PATCH bpf-next v7 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
@ 2026-09-29 0:16 ` Yonghong Song
2026-09-29 0:33 ` sashiko-bot
2026-09-29 0:16 ` [PATCH bpf-next v7 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
` (17 subsequent siblings)
21 siblings, 1 reply; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:16 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Add two new files, exception.h and exception.c, to host the exception
handling code. Only a few helpers so far: recognising a call to
bpf_unwind_resume(), and asking which landing pad, if any, a call site
unwinds to.
The pad of a call site is kept in insn_aux_data, so the three places that
move instructions around -- bpf_patch_insn_data(), verifier_remove_insns()
and bpf_opt_remove_nops() -- learn to keep it in step.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 7 +++++++
kernel/bpf/Makefile | 2 +-
kernel/bpf/exception.c | 30 ++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 12 ++++++++++++
kernel/bpf/fixups.c | 26 +++++++++++++++++++++++++-
5 files changed, 75 insertions(+), 2 deletions(-)
create mode 100644 kernel/bpf/exception.c
create mode 100644 kernel/bpf/exception.h
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 6ce25c96ebed..6d78c20e6507 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -700,6 +700,11 @@ struct bpf_insn_aux_data {
u64 jump_target:1;
unsigned int orig_idx; /* original instruction index, initialized once */
+ /*
+ * 1 + the instruction index of the exception cleanup landing pad
+ * this call site unwinds to, or 0 for none.
+ */
+ u32 cleanup_pad;
/*
* CFG strongly connected component this instruction belongs to,
* zero if it is a singleton SCC.
@@ -1588,6 +1593,8 @@ u32 btf_func_arg_align(const struct btf *btf, const struct btf_type *t);
int bpf_find_subprog(struct bpf_verifier_env *env, int off);
bool bpf_is_throw_kfunc(struct bpf_insn *insn);
+bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
+bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
int bpf_compute_const_regs(struct bpf_verifier_env *env);
int bpf_prune_dead_branches(struct bpf_verifier_env *env);
int bpf_check_cfg(struct bpf_verifier_env *env);
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index c1f9b0d3468d..8a6947b3d13a 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -11,7 +11,7 @@ obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_
obj-$(CONFIG_BPF_SYSCALL) += hashtab.o arraymap.o percpu_freelist.o bpf_lru_list.o lpm_trie.o map_in_map.o bloom_filter.o
obj-$(CONFIG_BPF_SYSCALL) += local_storage.o queue_stack_maps.o ringbuf.o bpf_insn_array.o
obj-$(CONFIG_BPF_SYSCALL) += bpf_local_storage.o bpf_task_storage.o
-obj-$(CONFIG_BPF_SYSCALL) += fixups.o cfg.o states.o backtrack.o check_btf.o
+obj-$(CONFIG_BPF_SYSCALL) += fixups.o cfg.o states.o backtrack.o check_btf.o exception.o
obj-${CONFIG_BPF_LSM} += bpf_inode_storage.o
obj-$(CONFIG_BPF_SYSCALL) += disasm.o mprog.o
obj-$(CONFIG_BPF_JIT) += trampoline.o
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
new file mode 100644
index 000000000000..b19fcbf49b7e
--- /dev/null
+++ b/kernel/bpf/exception.c
@@ -0,0 +1,30 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bpf.h>
+#include <linux/bpf_verifier.h>
+#include <linux/btf.h>
+#include <linux/btf_ids.h>
+#include <linux/filter.h>
+#include "exception.h"
+
+BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
+BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
+
+bool bpf_is_unwind_kfunc(const struct bpf_insn *insn)
+{
+ return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
+ insn->imm == bpf_unwind_id[0];
+}
+
+bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
+{
+ return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
+ insn->imm == bpf_unwind_resume_id[0];
+}
+
+int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
+{
+ u32 pad = env->insn_aux_data[idx].cleanup_pad;
+
+ return pad ? (int)pad - 1 : -1;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
new file mode 100644
index 000000000000..634cb2c15ffc
--- /dev/null
+++ b/kernel/bpf/exception.h
@@ -0,0 +1,12 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef _LINUX_BPF_EXCEPTION_H
+#define _LINUX_BPF_EXCEPTION_H
+
+#include <linux/types.h>
+
+struct bpf_verifier_env;
+
+int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
+
+#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 37cf130ebb57..5b7fe4ba610b 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -268,11 +268,18 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
data[i].non_stack_access =
data[off + cnt - 1].non_stack_access;
data[off + cnt - 1].non_stack_access = false;
+ data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
+ data[off + cnt - 1].cleanup_pad = 0;
} else if (bpf_is_mem_insn(insn + i)) {
data[i].non_stack_access = true;
}
}
+ if (env->cleanup_info_cnt)
+ for (i = 0; i < prog_len; i++)
+ if (data[i].cleanup_pad > off + 1)
+ data[i].cleanup_pad += cnt - 1;
+
/*
* Last slot instruction could be a newly generated
* BPF_ST/BPF_LDX/BPF_STX, systematically mark it for non-stack access
@@ -619,6 +626,7 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
unsigned int orig_prog_len = env->prog->len;
int err;
+ u32 i;
if (bpf_rewrite_must_abort())
return -EINTR;
@@ -647,6 +655,17 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
sizeof(*aux_data) * (orig_prog_len - off - cnt));
env->insn_aux_data_len -= cnt;
+ if (env->cleanup_info_cnt) {
+ for (i = 0; i < env->insn_aux_data_len; i++) {
+ u32 pad = aux_data[i].cleanup_pad;
+
+ if (pad > off + cnt)
+ aux_data[i].cleanup_pad = pad - cnt;
+ else if (pad > off)
+ aux_data[i].cleanup_pad = 0;
+ }
+ }
+
return 0;
}
@@ -752,7 +771,7 @@ int bpf_opt_remove_nops(struct bpf_verifier_env *env)
struct bpf_insn *insn = env->prog->insnsi;
int insn_cnt = env->prog->len;
bool is_may_goto_0, is_ja;
- int i, err;
+ int i, j, err;
for (i = 0; i < insn_cnt; i++) {
is_may_goto_0 = !memcmp(&insn[i], &MAY_GOTO_0, sizeof(MAY_GOTO_0));
@@ -763,6 +782,11 @@ int bpf_opt_remove_nops(struct bpf_verifier_env *env)
if (aux[i].indirect_target)
continue;
+ if (env->cleanup_info_cnt)
+ for (j = 0; j < insn_cnt; j++)
+ if (env->insn_aux_data[j].cleanup_pad == i + 1)
+ env->insn_aux_data[j].cleanup_pad = i + 2;
+
err = verifier_remove_insns(env, i, 1);
if (err)
return err;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 05/22] bpf: Prepare for an exception cleanup table before the CFG walk
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (3 preceding siblings ...)
2026-09-29 0:16 ` [PATCH bpf-next v7 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
@ 2026-09-29 0:16 ` Yonghong Song
2026-09-29 0:31 ` sashiko-bot
2026-09-29 0:16 ` [PATCH bpf-next v7 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
` (16 subsequent siblings)
21 siblings, 1 reply; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:16 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Record in insn_aux_data what the later passes need from the cleanup table:
cleanup_pad, the landing pad a frame resumes at, for every call that can
unwind -- a BPF-to-BPF call, direct or indirect, or bpf_unwind() -- within
the [begin_off, end_off) range of a cleanup record. Helper and other kfunc
calls in the range cannot unwind and are skipped. Subsequent commits
consume it.
bpf_prepare_cleanup_exceptions() runs before bpf_check_cfg(), whose walk
consumes what it produces. It refuses a table on an offloaded program, and
on one whose JIT cannot dispatch landing pads or was not asked to compile
it. What survives is marked jit_required: the interpreter cannot dispatch a
pad. bpf_jit_supports_cleanup_pads() is weak here and says no; the arch
patches provide the real ones.
A table is refused for a program whose verifier_ops has a gen_epilogue, as
bpf_qdisc's do. That epilogue is planted by rewriting the exits a program
has when bpf_convert_ctx_accesses() runs, and the exits an unwind returns
through are added after it, so they would skip it.
A table is also refused alongside bpf_throw(), a second answer to what runs
on the way out: it leaves for the exception boundary without rewriting the
return addresses of the frames it passes, so no pad between the two would
run. Both the tagged exception callback and the throw itself are checked --
either can appear without the other -- over the whole instruction stream,
so a throw in a subprogram is caught as well.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/filter.h | 1 +
kernel/bpf/core.c | 5 +++
kernel/bpf/exception.c | 81 ++++++++++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 2 ++
kernel/bpf/verifier.c | 6 ++++
5 files changed, 95 insertions(+)
diff --git a/include/linux/filter.h b/include/linux/filter.h
index e42eccb0990e..972b3ed2a51d 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1248,6 +1248,7 @@ bool bpf_jit_supports_stack_args(void);
bool bpf_jit_supports_arena_args(void);
bool bpf_jit_supports_far_kfunc_call(void);
bool bpf_jit_supports_exceptions(void);
+bool bpf_jit_supports_cleanup_pads(void);
bool bpf_jit_supports_ptr_xchg(void);
bool bpf_jit_supports_arena(void);
bool bpf_jit_supports_insn(struct bpf_insn *insn, bool in_arena);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index d3b8b626ec0f..d813fdde29e3 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -3511,6 +3511,11 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
{
}
+bool __weak bpf_jit_supports_cleanup_pads(void)
+{
+ return false;
+}
+
bool __weak bpf_jit_supports_timed_may_goto(void)
{
return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index b19fcbf49b7e..c1779d2d02f0 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -7,9 +7,90 @@
#include <linux/filter.h>
#include "exception.h"
+#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
+
BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
+static int reject_throw(struct bpf_verifier_env *env)
+{
+ u32 i;
+
+ for (i = 0; i < env->prog->len; i++) {
+ if (!bpf_is_throw_kfunc(&env->prog->insnsi[i]))
+ continue;
+ verbose(env,
+ "exception cleanup table cannot be combined with bpf_throw at insn %u\n",
+ i);
+ return -EINVAL;
+ }
+ return 0;
+}
+
+static int mark_call_sites(struct bpf_verifier_env *env)
+{
+ u32 i, j;
+
+ for (i = 0; i < env->cleanup_info_cnt; i++) {
+ struct bpf_cleanup_info *rec = &env->cleanup_info[i];
+
+ for (j = rec->begin_off; j < rec->end_off; j++) {
+ struct bpf_insn *insn = &env->prog->insnsi[j];
+
+ if (!bpf_pseudo_call(insn) && !bpf_is_callx(insn) &&
+ !bpf_is_unwind_kfunc(insn))
+ continue;
+ env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
+ }
+ }
+ return 0;
+}
+
+int bpf_exc_check_prog(struct bpf_verifier_env *env)
+{
+ if (bpf_prog_is_offloaded(env->prog->aux)) {
+ verbose(env,
+ "exception cleanup is not supported for offloaded programs\n");
+ return -EINVAL;
+ }
+ if (!bpf_jit_supports_cleanup_pads() || !env->prog->jit_requested) {
+ verbose(env,
+ "exception cleanup needs a JIT that can dispatch landing pads\n");
+ return -EOPNOTSUPP;
+ }
+ if (env->ops->gen_epilogue) {
+ verbose(env,
+ "exception cleanup is not supported for a program with an epilogue\n");
+ return -EOPNOTSUPP;
+ }
+ env->prog->jit_required = 1;
+ return 0;
+}
+
+int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env)
+{
+ int err;
+
+ if (!env->cleanup_info_cnt)
+ return 0;
+
+ err = bpf_exc_check_prog(env);
+ if (err)
+ return err;
+
+ if (env->exception_callback_subprog) {
+ verbose(env,
+ "exception cleanup table cannot be combined with an exception callback\n");
+ return -EINVAL;
+ }
+
+ err = reject_throw(env);
+ if (err)
+ return err;
+
+ return mark_call_sites(env);
+}
+
bool bpf_is_unwind_kfunc(const struct bpf_insn *insn)
{
return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index 634cb2c15ffc..1552438083d8 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -7,6 +7,8 @@
struct bpf_verifier_env;
+int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
+int bpf_exc_check_prog(struct bpf_verifier_env *env);
int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 6d3408f295ed..fc3df452de2e 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -37,6 +37,7 @@
#include "diagnostics.h"
#include "disasm.h"
+#include "exception.h"
static const struct bpf_verifier_ops * const bpf_verifier_ops[] = {
#define BPF_PROG_TYPE(_id, _name, prog_ctx_type, kern_ctx_type) \
@@ -22584,6 +22585,11 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
if (ret < 0)
goto skip_full_check;
+ /* The CFG needs an edge from a call in a cleanup range to its pad. */
+ ret = bpf_prepare_cleanup_exceptions(env);
+ if (ret < 0)
+ goto skip_full_check;
+
/* Validate instructions and resolve the program's referenced resources. */
ret = check_and_resolve_insns(env);
if (ret < 0)
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 06/22] bpf: Make exception landing pads reachable in the CFG
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (4 preceding siblings ...)
2026-09-29 0:16 ` [PATCH bpf-next v7 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
@ 2026-09-29 0:16 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 07/22] bpf: Resume a covered call at its landing pad Yonghong Song
` (15 subsequent siblings)
21 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:16 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A bpf_unwind() or a bpf2bpf call inside the [begin_off, end_off) range of a
cleanup record can reach that record's landing pad. Add that edge to the
CFG walk, which explores the pad and makes both ends prune points, and to
bpf_insn_successors(), which liveness and the SCC passes walk.
The pad is pushed by itself: visit_func_call_insn() returns as soon as
visit_cleanup_pad_edge() has pushed one, and the call site is visited again
for its fall-through once the pad is explored. bpf_check_cfg() peeks the
top of its stack and re-visits until DONE_EXPLORING, so a visit pushing two
successors would leave the pad DISCOVERED while it is no longer on the path
being walked, and push_insn() reads DISCOVERED as a back-edge. A branch
from one pad into another -- how a frame with two regions chains them --
would be taken for one. The pad's own insn_state is what records that the
edge is done, so there is nothing else to remember.
Liveness needs one more thing. A frame suspended at a call normally keeps a
stack slot alive only if something reads it once the call returns. A
landing pad is not on that path: it is reached from the call itself, not
from the instruction after it. So a slot that only the pad reads looks
dead, and clean_verifier_state() poisons it while the callee runs. Ask
whether the pad reads it too. entry_pad_stack is the shape that needs
this -- it stores to its frame before the call, never reads that slot on
the way back, and reloads it in the pad.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/cfg.c | 38 ++++++++++++++++++++++++++++++++++++++
kernel/bpf/liveness.c | 21 +++++++++++++++++++++
2 files changed, 59 insertions(+)
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index b0bd9ba951df..2eb07397e874 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -6,6 +6,7 @@
#include <linux/sort.h>
#include "diagnostics.h"
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -160,6 +161,38 @@ static int push_insn(int t, int w, int e, struct bpf_verifier_env *env)
return DONE_EXPLORING;
}
+static int visit_cleanup_pad_edge(int t, struct bpf_verifier_env *env)
+{
+ int *insn_stack = env->cfg.insn_stack;
+ int *insn_state = env->cfg.insn_state;
+ int w;
+
+ if (!env->cleanup_info_cnt)
+ return DONE_EXPLORING;
+ w = bpf_exc_pad_of_call(env, t);
+ if (w < 0)
+ return DONE_EXPLORING;
+
+ /*
+ * @t is a call that may branch here, and @w is the target of that
+ * branch, so both are prune points. @w especially: every covered call
+ * site in a region unwinds to the same pad, and without a prune point
+ * at its head the verifier walks the pad again for each of them.
+ */
+ mark_prune_point(env, t);
+ mark_prune_point(env, w);
+ mark_jmp_point(env, w);
+ mark_jump_target(env, w);
+
+ if (insn_state[w])
+ return DONE_EXPLORING;
+ if (env->cfg.cur_stack >= env->prog->len)
+ return -E2BIG;
+ insn_stack[env->cfg.cur_stack++] = w;
+ insn_state[w] |= DISCOVERED;
+ return KEEP_EXPLORING;
+}
+
static int visit_func_call_insn(int t, struct bpf_insn *insns,
struct bpf_verifier_env *env,
bool visit_callee)
@@ -167,6 +200,11 @@ static int visit_func_call_insn(int t, struct bpf_insn *insns,
int ret, insn_sz;
int w;
+ /* One push per visit: @t is revisited once the pad is explored. */
+ ret = visit_cleanup_pad_edge(t, env);
+ if (ret != DONE_EXPLORING)
+ return ret;
+
insn_sz = bpf_is_ldimm64(&insns[t]) ? 2 : 1;
ret = push_insn(t, t + insn_sz, FALLTHROUGH, env);
if (ret)
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index cd9523f69298..4e0273a8ceee 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -8,6 +8,8 @@
#include <linux/slab.h>
#include <linux/sort.h>
+#include "exception.h"
+
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
/*
@@ -384,6 +386,18 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
succ->items[succ->cnt++] = exit_idx;
}
+ /*
+ * A call a cleanup record covers can leave through its landing pad.
+ * Only a call to a subprogram or to bpf_unwind() is marked, neither of
+ * which is an edge the block above adds, so succ still holds two.
+ */
+ if (unlikely(env->cleanup_info_cnt)) {
+ int pad = bpf_exc_pad_of_call(env, idx);
+
+ if (pad >= 0)
+ succ->items[succ->cnt++] = pad;
+ }
+
return succ;
}
@@ -545,6 +559,13 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
alive = callee_stack_access_at_callsite(env, callsite)
? is_live_before(instance, callsite, rel, half_spi)
: is_live_before(instance, callsite + 1, rel, half_spi);
+
+ if (!alive && unlikely(env->cleanup_info_cnt)) {
+ int pad = bpf_exc_pad_of_call(env, callsite);
+
+ if (pad >= 0)
+ alive = is_live_before(instance, pad, rel, half_spi);
+ }
if (alive)
return true;
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 07/22] bpf: Resume a covered call at its landing pad
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (5 preceding siblings ...)
2026-09-29 0:16 ` [PATCH bpf-next v7 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
@ 2026-09-29 0:16 ` Yonghong Song
2026-09-29 0:31 ` sashiko-bot
2026-09-29 0:16 ` [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
` (14 subsequent siblings)
21 siblings, 1 reply; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:16 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A landing pad runs in the frame that owns it, entered by an ordinary
return: bpf_unwind() rewrites the frame's saved return address, so the call
the frame is suspended at comes back at the pad rather than at the next
instruction. That makes the pad a second successor of a covered call, in
the same frame, whose entry state is knowable without verifying the callee
-- the state at the call with the caller-saved registers gone, since the
callee's epilogue puts r6-r9 and the stack back on its way out.
push_cleanup_pad_branch() pushes exactly that. The call to bpf_unwind()
itself never comes back to the instruction after it: it resumes at this
frame's pad where a record covers the call, and otherwise the frame
returns at once. Where that frame is the main program's, returning at once
is the program returning, so it leaves through process_bpf_exit_full() and
the zero it returns is held to the program type.
Precision backtracking has to tell those two edges apart, and subseq_idx is
what it has to do it with. Coming back to a covered call from its own pad
stays in this frame, the callee never having been entered on that path.
Coming back to a resume is the opposite -- a resume leaves its frame the
way an exit does -- so the walk enters the callee there, and r6-r9 and the
stack stay marked in this frame's masks until it comes back out. Miss that
and the callee is walked against the caller's masks.
Answering the two here skips check_kfunc_call(), so the filter it applies
first -- whether this program may call this kfunc at all -- is split out as
check_kfunc_allowed() and applied here too. Until the patch that registers
them, that filter is what refuses them.
The CFG walk starts summarising it too: a subprogram calling bpf_unwind()
is marked might_unwind, and merge_callee_effects() carries that up to its
callers. Nothing reads it yet; later patches do.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 1 +
kernel/bpf/backtrack.c | 42 ++++++++++++++
kernel/bpf/cfg.c | 11 ++++
kernel/bpf/verifier.c | 108 ++++++++++++++++++++++++++++++++---
4 files changed, 154 insertions(+), 8 deletions(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 6d78c20e6507..0143688896b0 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -836,6 +836,7 @@ struct bpf_subprog_info {
s16 fastcall_stack_off;
bool has_tail_call: 1;
bool might_throw: 1;
+ bool might_unwind: 1;
bool tail_call_reachable: 1;
bool has_ld_abs: 1;
bool is_cb: 1;
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 0e38b9575328..90f30e152cbe 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -4,6 +4,7 @@
#include <linux/bpf_verifier.h>
#include <linux/filter.h>
#include <linux/bitmap.h>
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -434,6 +435,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
return -EFAULT;
}
+ if (bpf_exc_pad_of_call(env, idx) == subseq_idx) {
+ /*
+ * We came from this call's landing pad, which
+ * runs in the caller's frame: on that path the
+ * callee's frame was never entered, so there is
+ * no frame to leave. The call clobbered r0-r5;
+ * r6-r9 and the stack are the caller's own and
+ * keep going back from here.
+ */
+ bt_clear_reg(bt, BPF_REG_0);
+ if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
+ verifier_bug(env, "landing pad unexpected regs %x",
+ bt_reg_mask(bt));
+ return -EFAULT;
+ }
+ return 0;
+ }
+
/* callx calls static subprogs only */
if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
/* check that jump history doesn't have any
@@ -523,6 +542,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
if (bt_subprog_exit(bt))
return -EFAULT;
return 0;
+ } else if (bpf_is_unwind_resume_kfunc(insn)) {
+ /*
+ * A resume leaves its frame the way an exit does, so
+ * the walk is crossing from the caller into the callee
+ * here and has a frame to enter. The zero the resume
+ * returns is its own: nothing further back defines r0,
+ * and r1-r5 the call clobbered.
+ */
+ bt_clear_reg(bt, BPF_REG_0);
+ bt_clear_reg(bt, BPF_REG_2);
+ if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
+ verifier_bug(env, "backtracking resume unexpected regs %x",
+ bt_reg_mask(bt));
+ return -EFAULT;
+ }
+ if (bt_subprog_enter(bt))
+ return -EFAULT;
+ return 0;
} else if (opcode == BPF_CALL) {
/* kfunc with imm==0 is invalid and fixup_kfunc_call will
* catch this error later. Make backtracking conservative
@@ -956,6 +993,11 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
if (!st)
break;
+ if (verifier_bug_if(bt->frame > st->curframe, env,
+ "backtrack frame %d, state curframe %d",
+ bt->frame, st->curframe))
+ return -EFAULT;
+
for (fr = bt->frame; fr >= 0; fr--) {
func = st->frame[fr];
bitmap_from_u64(mask, bt_frame_reg_mask(bt, fr));
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index 2eb07397e874..82b5abcc736b 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -76,6 +76,14 @@ static void mark_subprog_might_throw(struct bpf_verifier_env *env, int off)
subprog->might_throw = true;
}
+static void mark_subprog_might_unwind(struct bpf_verifier_env *env, int off)
+{
+ struct bpf_subprog_info *subprog;
+
+ subprog = bpf_find_containing_subprog(env, off);
+ subprog->might_unwind = true;
+}
+
/* 't' is an index of a call-site.
* 'w' is a callee entry point.
* Eventually this function would be called when env->cfg.insn_state[w] == EXPLORED.
@@ -91,6 +99,7 @@ static void merge_callee_effects(struct bpf_verifier_env *env, int t, int w)
caller->changes_pkt_data |= callee->changes_pkt_data;
caller->might_sleep |= callee->might_sleep;
caller->might_throw |= callee->might_throw;
+ caller->might_unwind |= callee->might_unwind;
}
enum {
@@ -668,6 +677,8 @@ static int visit_insn(int t, struct bpf_verifier_env *env)
mark_subprog_changes_pkt_data(env, t);
if (ret == 0 && bpf_is_throw_kfunc(insn))
mark_subprog_might_throw(env, t);
+ if (ret == 0 && bpf_is_unwind_kfunc(insn))
+ mark_subprog_might_unwind(env, t);
}
return visit_func_call_insn(t, insns, env, insn->src_reg == BPF_PSEUDO_CALL);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index fc3df452de2e..ee074d4a936b 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -14725,6 +14725,32 @@ static int check_special_kfunc(struct bpf_verifier_env *env, struct bpf_call_arg
static int check_return_code(struct bpf_verifier_env *env, int regno, const char *reg_name);
+static int check_kfunc_allowed(struct bpf_verifier_env *env, struct bpf_insn *insn,
+ int insn_idx, struct bpf_call_arg_meta *meta)
+{
+ const char *operation;
+ int err;
+
+ err = bpf_fetch_kfunc_arg_meta(env, insn->imm, insn->off, meta);
+ if (err == -EACCES && meta->func_name) {
+ verbose(env, "calling kernel function %s is not allowed\n", meta->func_name);
+ operation = bpf_diag_fmt(env, "kfunc %s", meta->func_name);
+ bpf_diag_policy(
+ env, insn_idx, operation, "this program cannot call the kfunc",
+ "Use a kfunc allowed for this program type and attach point, or change the program context.");
+ }
+ return err;
+}
+
+/* noinline saves the caller a 200-byte struct bpf_call_arg_meta on its frame. */
+static noinline int check_kfunc_allowed_only(struct bpf_verifier_env *env,
+ struct bpf_insn *insn, int insn_idx)
+{
+ struct bpf_call_arg_meta meta;
+
+ return check_kfunc_allowed(env, insn, insn_idx, &meta);
+}
+
static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
int *insn_idx_p)
{
@@ -14746,14 +14772,7 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
if (!insn->imm)
return 0;
- err = bpf_fetch_kfunc_arg_meta(env, insn->imm, insn->off, &meta);
- if (err == -EACCES && meta.func_name) {
- verbose(env, "calling kernel function %s is not allowed\n", meta.func_name);
- operation = bpf_diag_fmt(env, "kfunc %s", meta.func_name);
- bpf_diag_policy(
- env, insn_idx, operation, "this program cannot call the kfunc",
- "Use a kfunc allowed for this program type and attach point, or change the program context.");
- }
+ err = check_kfunc_allowed(env, insn, insn_idx, &meta);
if (err)
return err;
desc_btf = meta.btf;
@@ -19167,6 +19186,57 @@ enum {
INSN_IDX_UPDATED = 2,
};
+static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
+{
+ struct bpf_verifier_state *branch;
+ struct bpf_func_state *frame;
+ int pad = bpf_exc_pad_of_call(env, insn_idx);
+
+ if (pad < 0)
+ return 0;
+ branch = push_stack(env, pad, insn_idx, false);
+ if (IS_ERR(branch))
+ return PTR_ERR(branch);
+ frame = branch->frame[branch->curframe];
+ /*
+ * The state at that call with the caller-saved registers gone: the
+ * callee's epilogue put r6-r9 and the stack back on the way out.
+ */
+ clear_caller_saved_regs(env, frame->regs);
+ mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ return 0;
+}
+
+static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
+ bool *do_print_state)
+{
+ struct bpf_func_state *frame = cur_func(env);
+ int pad = bpf_exc_pad_of_call(env, *insn_idx);
+ int err;
+
+ if (pad < 0) {
+ err = check_resource_leak(env, false, !env->cur_state->curframe,
+ "an unwind with no landing pad");
+ if (err)
+ return err;
+ if (env->cur_state->curframe)
+ return PROCESS_BPF_EXIT;
+ /*
+ * The main program's frame returns at once, which is the
+ * program returning. Mark r0 the zero the fixups leave after
+ * the call, and leave through the exit, which is what holds
+ * that zero to the program type.
+ */
+ mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
+ mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
+ return process_bpf_exit_full(env, do_print_state, false);
+ }
+ clear_caller_saved_regs(env, frame->regs);
+ mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ *insn_idx = pad;
+ return INSN_IDX_UPDATED;
+}
+
static int process_bpf_exit_full(struct bpf_verifier_env *env,
bool *do_print_state,
bool exception_exit)
@@ -19421,7 +19491,29 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
return -EINVAL;
}
}
+ if (bpf_is_unwind_kfunc(insn) || bpf_is_unwind_resume_kfunc(insn)) {
+ err = check_kfunc_allowed_only(env, insn, env->insn_idx);
+ if (err)
+ return err;
+ if (bpf_is_unwind_kfunc(insn))
+ return process_bpf_unwind(env, &env->insn_idx,
+ do_print_state);
+ /*
+ * Mark r0 a known zero -- unknown first, as
+ * the known-zero helper keeps the type it
+ * finds, which here is NOT_INIT. The fixups
+ * lower this to 'r0 = 0; exit', so the frame
+ * returns a real zero.
+ */
+ mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
+ mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
+ return process_bpf_exit_full(env, do_print_state, false);
+ }
mark_reg_scratched(env, BPF_REG_0);
+ /* An unwind out of this call resumes at the pad. */
+ err = push_cleanup_pad_branch(env, env->insn_idx);
+ if (err)
+ return err;
if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
cur_func(env)->no_stack_arg_load = true;
if (bpf_is_callx(insn))
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (6 preceding siblings ...)
2026-09-29 0:16 ` [PATCH bpf-next v7 07/22] bpf: Resume a covered call at its landing pad Yonghong Song
@ 2026-09-29 0:16 ` Yonghong Song
2026-09-29 0:36 ` sashiko-bot
2026-09-29 0:52 ` bot+bpf-ci
2026-09-29 0:16 ` [PATCH bpf-next v7 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
` (13 subsequent siblings)
21 siblings, 2 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:16 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A landing pad's entry state is the state at the call it belongs to, taken
before the callee ran -- push_cleanup_pad_branch() snapshots it there. That
is right for the frame's own registers and stack, which the callee's
epilogue puts back, and wrong for what the program shares: nothing restores
the locks in bpf_verifier_state. So a callee that drops a lock its caller
took and then resumes leaves the caller's pad verified against a state that
still holds it. Say foo() takes an RCU read lock and calls bar(), which
unwinds:
foo:
call bpf_rcu_read_lock
1: call bar /* covered, pad at 3 */
2: r0 = 0; exit
3: call bpf_rcu_read_unlock /* pad: drops the lock foo() took */
call bpf_unwind_resume
bar:
1: call bpf_unwind /* covered, pad at 3 */
2: r0 = 0; exit
3: call bpf_rcu_read_unlock /* pad: drops it a second time */
call bpf_unwind_resume
At run time both pads run and the one lock is released twice, and no single
state shows it: bar's path ends at its resume, and foo's pad is verified
from the snapshot at its call to bar. A frame that acquires a lock and
leaves through an unwind hides the same way. Fix both by making the
snapshot true -- record what the program holds when a frame is entered and
require an unwind leaving it to have put that back, at the resume and at a
bpf_unwind() no record covers. References go by id, since ids only go up
and bpf_reference_state does not say which frame acquired one.
The frames in between need the same of them, and have nothing to run: where
no record covers the call a frame is suspended at, the JIT sends it to its
epilogue, so what it acquired since it was entered is dropped on the floor
and no path of its own arrives to say so. Ask it at the call instead, which
is where it is abandoned -- check_unwind_through_call(). Per frame rather
than of the whole stack at the unwind, since a frame that does carry a
record may hold what its pad will release. A callx counts as any subprogram
that might unwind, its target not being known there.
So every frame an unwind leaves is asked, one way of asking per way out:
- "a resume": a frame with a pad runs it and ends at bpf_unwind_resume().
The frame that raised the unwind leaves this way where a record covers
its bpf_unwind(), and so does every frame above it whose call is
covered.
- "an unwind with no landing pad": the frame that raised it with no
record over its bpf_unwind(), returning through the exit patched in
after it.
- "an unwind through this call": every frame above it whose call no
record covers, asked at the call since nothing of it runs again.
Nothing else is left to ask: a frame other than the one that raised the
unwind is suspended at a call, and the call either carries a record or does
not. The unwind reaches no further than the program it was raised in, the
walk stopping at the first frame that is not a subprogram.
One walk is then left that cannot happen. A resume ends its frame the way
an exit does, so the verifier continued the caller at the instruction
after its call -- the one place a resume does not return to, since
bpf_unwind() rewrote that address before any pad ran. Where it does return
is walked already: the caller's pad is the branch pushed at the call, and
with no pad nothing of the caller runs. Walking on anyway keeps code the
JIT never reaches out of the dead code sweep, and arrives holding what
only the pad releases, which refuses a correct program. End the path there
instead, as an unwind with no landing pad already does. The main program's
frame keeps the old way out: no caller to leave to, and the exit is what
holds its zero to the program type.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 12 ++++++
kernel/bpf/exception.c | 76 ++++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 6 +++
kernel/bpf/verifier.c | 50 +++++++++++++++++++++++-
4 files changed, 142 insertions(+), 2 deletions(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 0143688896b0..75e572a8a1ba 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -339,6 +339,18 @@ struct bpf_func_state {
bool in_async_callback_fn;
bool in_exception_callback_fn;
bool no_stack_arg_load;
+ /*
+ * What the program held when this frame was entered. An unwind leaves
+ * the frame without running anything below it, so the frame has to put
+ * these back to what it found before it goes -- otherwise a caller's
+ * landing pad, whose state was taken at the call, is wrong about them.
+ */
+ u32 entry_active_locks;
+ u32 entry_preempt_locks;
+ u32 entry_rcu_locks;
+ u32 entry_irq_id;
+ u32 entry_id_gen;
+ u32 entry_acquired_refs;
/* For callback calling functions that limit number of possible
* callback executions (e.g. bpf_loop) keeps track of current
* simulated iteration number.
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index c1779d2d02f0..c0b0b8478af5 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -12,6 +12,72 @@
BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
+void bpf_exc_record_frame_entry(const struct bpf_verifier_state *state,
+ struct bpf_func_state *frame, u32 id_gen)
+{
+ u32 i;
+
+ frame->entry_active_locks = state->active_locks;
+ frame->entry_preempt_locks = state->active_preempt_locks;
+ frame->entry_rcu_locks = state->active_rcu_locks;
+ frame->entry_irq_id = state->active_irq_id;
+
+ /* Ids only ever go up, so this one tells the frame's own apart. */
+ frame->entry_id_gen = id_gen;
+ frame->entry_acquired_refs = 0;
+ for (i = 0; i < state->acquired_refs; i++)
+ if (state->refs[i].type == REF_TYPE_PTR)
+ frame->entry_acquired_refs++;
+}
+
+int bpf_exc_check_frame_balance(struct bpf_verifier_env *env, const char *prefix)
+{
+ const struct bpf_verifier_state *state = env->cur_state;
+ const struct bpf_func_state *frame = cur_func(env);
+ u32 i, held;
+ const char *what;
+
+ if (state->active_rcu_locks != frame->entry_rcu_locks)
+ what = "bpf_rcu_read_lock";
+ else if (state->active_preempt_locks != frame->entry_preempt_locks)
+ what = "bpf_preempt_disable";
+ else if (state->active_irq_id != frame->entry_irq_id)
+ what = "bpf_local_irq_save";
+ else if (state->active_locks != frame->entry_active_locks)
+ what = "bpf_spin_lock";
+ else
+ what = NULL;
+
+ if (what) {
+ verbose(env, "%s does not leave the frame's %s state as it found it\n",
+ prefix, what);
+ return -EINVAL;
+ }
+
+ /*
+ * References the same way. ids only go up, so entry_id_gen splits
+ * refs[] in two at frame entry: nothing above that line may still be
+ * held, and the count below it has to be what it was.
+ */
+ for (i = 0, held = 0; i < state->acquired_refs; i++) {
+ if (state->refs[i].type != REF_TYPE_PTR)
+ continue;
+ if (state->refs[i].id > frame->entry_id_gen) {
+ verbose(env, "%s keeps the reference id=%d the frame acquired\n",
+ prefix, state->refs[i].id);
+ return -EINVAL;
+ }
+ held++;
+ }
+ if (held != frame->entry_acquired_refs) {
+ verbose(env, "%s does not leave the frame's references as it found it\n",
+ prefix);
+ return -EINVAL;
+ }
+
+ return 0;
+}
+
static int reject_throw(struct bpf_verifier_env *env)
{
u32 i;
@@ -46,6 +112,16 @@ static int mark_call_sites(struct bpf_verifier_env *env)
return 0;
}
+bool bpf_prog_may_unwind(const struct bpf_verifier_env *env)
+{
+ u32 i;
+
+ for (i = 0; i < env->subprog_cnt; i++)
+ if (env->subprog_info[i].might_unwind)
+ return true;
+ return false;
+}
+
int bpf_exc_check_prog(struct bpf_verifier_env *env)
{
if (bpf_prog_is_offloaded(env->prog->aux)) {
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index 1552438083d8..e93da039b5bd 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -6,9 +6,15 @@
#include <linux/types.h>
struct bpf_verifier_env;
+struct bpf_verifier_state;
+struct bpf_func_state;
int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
int bpf_exc_check_prog(struct bpf_verifier_env *env);
+bool bpf_prog_may_unwind(const struct bpf_verifier_env *env);
+void bpf_exc_record_frame_entry(const struct bpf_verifier_state *state,
+ struct bpf_func_state *frame, u32 id_gen);
+int bpf_exc_check_frame_balance(struct bpf_verifier_env *env, const char *prefix);
int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index ee074d4a936b..4bdee3f02fe9 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10774,6 +10774,7 @@ static int setup_func_entry(struct bpf_verifier_env *env, int subprog, int calls
callsite,
state->curframe + 1 /* frameno within this callchain */,
subprog /* subprog number within this prog */);
+ bpf_exc_record_frame_entry(state, callee, env->id_gen);
err = set_callee_state_cb(env, caller, callee, callsite);
if (err)
goto err_out;
@@ -19207,6 +19208,32 @@ static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
return 0;
}
+/* Can an unwind come back out of this call? */
+static bool call_may_unwind(struct bpf_verifier_env *env, const struct bpf_insn *insn,
+ int insn_idx)
+{
+ int subprog;
+
+ /* Which subprog a callx lands in is not known here, so any may be it. */
+ if (bpf_is_callx(insn))
+ return bpf_prog_may_unwind(env);
+ if (insn->src_reg != BPF_PSEUDO_CALL)
+ return false;
+ subprog = bpf_find_subprog(env, insn_idx + insn->imm + 1);
+ return subprog >= 0 && env->subprog_info[subprog].might_unwind;
+}
+
+static int check_unwind_through_call(struct bpf_verifier_env *env, int insn_idx)
+{
+ const struct bpf_insn *insn = &env->prog->insnsi[insn_idx];
+
+ if (bpf_exc_pad_of_call(env, insn_idx) >= 0)
+ return 0;
+ if (!call_may_unwind(env, insn, insn_idx))
+ return 0;
+ return bpf_exc_check_frame_balance(env, "an unwind through this call");
+}
+
static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
bool *do_print_state)
{
@@ -19215,8 +19242,13 @@ static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
int err;
if (pad < 0) {
- err = check_resource_leak(env, false, !env->cur_state->curframe,
- "an unwind with no landing pad");
+ if (!env->cur_state->curframe) {
+ err = check_resource_leak(env, false, true,
+ "an unwind with no landing pad");
+ if (err)
+ return err;
+ }
+ err = bpf_exc_check_frame_balance(env, "an unwind with no landing pad");
if (err)
return err;
if (env->cur_state->curframe)
@@ -19498,6 +19530,16 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
if (bpf_is_unwind_kfunc(insn))
return process_bpf_unwind(env, &env->insn_idx,
do_print_state);
+ err = bpf_exc_check_frame_balance(env, "a resume");
+ if (err)
+ return err;
+ /*
+ * No need to walk into the caller: its pad was
+ * pushed as a branch at its call, and with no
+ * pad nothing of it runs.
+ */
+ if (env->cur_state->curframe)
+ return PROCESS_BPF_EXIT;
/*
* Mark r0 a known zero -- unknown first, as
* the known-zero helper keeps the type it
@@ -19512,6 +19554,10 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
mark_reg_scratched(env, BPF_REG_0);
/* An unwind out of this call resumes at the pad. */
err = push_cleanup_pad_branch(env, env->insn_idx);
+ if (err)
+ return err;
+ /* Or, with no pad, leaves the frame for good. */
+ err = check_unwind_through_call(env, env->insn_idx);
if (err)
return err;
if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 09/22] bpf: Refuse a landing pad that does not resume
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (7 preceding siblings ...)
2026-09-29 0:16 ` [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
@ 2026-09-29 0:16 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 10/22] bpf: Refuse a private stack for a program that can unwind Yonghong Song
` (12 subsequent siblings)
21 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:16 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A cleanup pad runs drop glue and calls bpf_unwind_resume(), so the frame
returns and the unwind goes on. A catch pad runs the same drops and then
carries on in its frame, stopping the unwind. Only the first is supported:
bpf_unwind() rewrites every frame's return address in one pass, so a frame
above a catch pad would resume at a pad for an unwind already caught.
Nothing in the record says which kind a pad is, but the code does, as LLVM
emits it: a cleanup pad reaches _Unwind_Resume, a catch pad reaches a
return. bpf_unwind() and the branch pushed at a covered call are the only
ways in, and both mark the frame they enter. bpf_exc_check_insn() asks that
mark about every instruction of a program carrying a table, and refuses:
- an exit, which is how a catch pad ends
- a tail call, and a BPF_LD_[ABS|IND], which leaves through an exit on a
failed load
- an indirect jump
- a bpf_unwind(), and a call to a global subprogram that might_unwind
- an instruction reached both inside and outside a pad
do_check_insn() refuses the other half of it, a bpf_unwind_resume() the
mark does not find in a pad. That one is asked of every program rather
than only those carrying a table, since a program with no table has no pad
to be in.
The first three concern the pad's own frame, since a subprogram it calls
may do any of them and still come back; an unwind is refused anywhere above
a pad, since it never does. do_check() asks before pruning, so the mark
stays out of states_equal().
A speculative walk can reach a pad too; that is answered as do_check()
answers anything it cannot allow speculatively, by marking the instruction
for a barrier and stopping rather than refusing the program. A callback
that can unwind is refused as well, its helper frame being C with no pad.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 4 ++
kernel/bpf/exception.c | 81 ++++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 3 ++
kernel/bpf/verifier.c | 21 ++++++++++
4 files changed, 109 insertions(+)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 75e572a8a1ba..eb35aa4cfd37 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -339,6 +339,8 @@ struct bpf_func_state {
bool in_async_callback_fn;
bool in_exception_callback_fn;
bool no_stack_arg_load;
+ /* an unwind reached this frame and its landing pad is running */
+ bool in_pad;
/*
* What the program held when this frame was entered. An unwind leaves
* the frame without running anything below it, so the frame has to put
@@ -710,6 +712,8 @@ struct bpf_insn_aux_data {
u64 non_stack_access:1; /* instruction can access non-stack memory */
/* true if some jump or call instruction targets this instruction */
u64 jump_target:1;
+ u64 in_cleanup_pad:1; /* reached with a landing pad running */
+ u64 outside_cleanup_pad:1; /* reached the other way */
unsigned int orig_idx; /* original instruction index, initialized once */
/*
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index c0b0b8478af5..f2bca0242408 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -12,6 +12,15 @@
BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
+int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog)
+{
+ if (!env->subprog_info[subprog].might_unwind)
+ return 0;
+
+ verbose(env, "subprog %d may unwind and is used as a callback\n", subprog);
+ return -EINVAL;
+}
+
void bpf_exc_record_frame_entry(const struct bpf_verifier_state *state,
struct bpf_func_state *frame, u32 id_gen)
{
@@ -179,6 +188,78 @@ bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
insn->imm == bpf_unwind_resume_id[0];
}
+/* Is an unwind in flight: is this frame a landing pad, or below one? */
+static bool unwinding(const struct bpf_verifier_state *state)
+{
+ u32 i;
+
+ for (i = 0; i <= state->curframe; i++)
+ if (state->frame[i]->in_pad)
+ return true;
+ return false;
+}
+
+int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn)
+{
+ bool in_pad = cur_func(env)->in_pad;
+ struct bpf_insn_aux_data *aux;
+ u32 i = env->insn_idx;
+ const char *why = NULL;
+
+ if (unwinding(env->cur_state)) {
+ if (bpf_is_unwind_kfunc(insn)) {
+ verbose(env, "insn %u starts a second unwind while one is in flight\n", i);
+ return -EINVAL;
+ }
+ if (bpf_pseudo_call(insn)) {
+ int subprog = bpf_find_subprog(env, i + insn->imm + 1);
+
+ if (subprog >= 0 && bpf_subprog_is_global(env, subprog) &&
+ env->subprog_info[subprog].might_unwind) {
+ verbose(env,
+ "insn %u calls global subprog %d, which can unwind while an unwind is in flight\n",
+ i, subprog);
+ return -EINVAL;
+ }
+ }
+ }
+
+ aux = &env->insn_aux_data[i];
+
+ if (in_pad ? aux->outside_cleanup_pad : aux->in_cleanup_pad) {
+ verbose(env, "insn %u runs both inside and outside a landing pad\n", i);
+ return -EINVAL;
+ }
+ if (in_pad)
+ aux->in_cleanup_pad = true;
+ else
+ aux->outside_cleanup_pad = true;
+
+ if (!in_pad)
+ return 0;
+
+ if (insn->code == (BPF_JMP | BPF_EXIT)) {
+ verbose(env,
+ "exit at insn %u ends a landing pad: a catch pad is not supported yet, only cleanup pads that resume\n",
+ i);
+ return -EOPNOTSUPP;
+ }
+ if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call)
+ why = "is a tail call, which replaces the frame";
+ else if (BPF_CLASS(insn->code) == BPF_LD &&
+ (BPF_MODE(insn->code) == BPF_ABS || BPF_MODE(insn->code) == BPF_IND))
+ why = "is a BPF_LD_[ABS|IND], which can leave through the epilogue";
+ else if (insn->code == (BPF_JMP | BPF_JA | BPF_X) ||
+ insn->code == (BPF_JMP32 | BPF_JA | BPF_X))
+ why = "is an indirect jump";
+
+ if (!why)
+ return 0;
+
+ verbose(env, "insn %u %s, and is in a landing pad\n", i, why);
+ return -EINVAL;
+}
+
int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
{
u32 pad = env->insn_aux_data[idx].cleanup_pad;
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index e93da039b5bd..c5b30ff3ccc0 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -8,6 +8,7 @@
struct bpf_verifier_env;
struct bpf_verifier_state;
struct bpf_func_state;
+struct bpf_insn;
int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
int bpf_exc_check_prog(struct bpf_verifier_env *env);
@@ -16,5 +17,7 @@ void bpf_exc_record_frame_entry(const struct bpf_verifier_state *state,
struct bpf_func_state *frame, u32 id_gen);
int bpf_exc_check_frame_balance(struct bpf_verifier_env *env, const char *prefix);
int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
+int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog);
+int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn);
#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 4bdee3f02fe9..49aa76c1ea9f 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10986,6 +10986,10 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
* callbacks
*/
env->subprog_info[subprog].is_cb = true;
+ err = bpf_exc_check_callback(env, subprog);
+ if (err)
+ return err;
+
if (bpf_pseudo_kfunc_call(insn) &&
!is_callback_calling_kfunc(insn->imm)) {
verifier_bug(env, "kfunc %s#%d not marked as callback-calling",
@@ -19205,6 +19209,7 @@ static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
*/
clear_caller_saved_regs(env, frame->regs);
mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ frame->in_pad = true;
return 0;
}
@@ -19265,6 +19270,7 @@ static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
}
clear_caller_saved_regs(env, frame->regs);
mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ frame->in_pad = true;
*insn_idx = pad;
return INSN_IDX_UPDATED;
}
@@ -19530,6 +19536,11 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
if (bpf_is_unwind_kfunc(insn))
return process_bpf_unwind(env, &env->insn_idx,
do_print_state);
+ if (!cur_func(env)->in_pad) {
+ verbose(env, "resume at insn %d is not in a landing pad\n",
+ env->insn_idx);
+ return -EINVAL;
+ }
err = bpf_exc_check_frame_balance(env, "a resume");
if (err)
return err;
@@ -19664,6 +19675,16 @@ static int do_check(struct bpf_verifier_env *env)
}
}
+ if (unlikely(env->cleanup_info_cnt)) {
+ err = bpf_exc_check_insn(env, insn);
+ if (error_recoverable_with_nospec(err) && state->speculative) {
+ insn_aux->nospec = true;
+ goto process_bpf_exit;
+ }
+ if (err)
+ return err;
+ }
+
if (bpf_is_prune_point(env, env->insn_idx)) {
err = bpf_is_state_visited(env, env->insn_idx);
if (err < 0)
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 10/22] bpf: Refuse a private stack for a program that can unwind
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (8 preceding siblings ...)
2026-09-29 0:16 ` [PATCH bpf-next v7 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
@ 2026-09-29 0:16 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
` (11 subsequent siblings)
21 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:16 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
x86-64 has no register to spare for this. A private stack keeps its frame
pointer in r9 and the JIT brackets every call in push_r9/pop_r9, but an
unwind resumes a frame at its landing pad rather than at the instruction
after the call, so the pad is reached before the pop runs and would address
its frame through a stale pointer.
So no private stack for a program that can unwind, on every architecture
rather than just that one. A table is not the condition: a bpf_unwind()
with no record over it sends its callers to their epilogues just the same.
In check_max_stack_depth(), force NO_PRIV_STACK so the JIT does not use a
private stack.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/verifier.c | 11 +++++++++++
1 file changed, 11 insertions(+)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 49aa76c1ea9f..0419609beb67 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -5751,6 +5751,17 @@ static int check_max_stack_depth(struct bpf_verifier_env *env)
}
}
+ /*
+ * x86-64 has no register to spare for this: a private stack keeps its
+ * frame pointer in r9, restored by a pop after the call that an unwind
+ * skips. A frame resumed at a pad then addresses its stack through a
+ * stale pointer, and a frame sent to its epilogue instead pops its
+ * callee-saved registers one slot off. Refuse a private stack for any
+ * program that can unwind, on every arch for now.
+ */
+ if (env->cleanup_info_cnt || bpf_prog_may_unwind(env))
+ priv_stack_mode = NO_PRIV_STACK;
+
if (priv_stack_mode == PRIV_STACK_UNKNOWN)
priv_stack_mode = bpf_enable_priv_stack(env->prog);
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 11/22] bpf: Dispatch cleanup pads by rewriting return addresses
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (9 preceding siblings ...)
2026-09-29 0:16 ` [PATCH bpf-next v7 10/22] bpf: Refuse a private stack for a program that can unwind Yonghong Song
@ 2026-09-29 0:16 ` Yonghong Song
2026-09-29 1:14 ` bot+bpf-ci
2026-09-29 0:17 ` [PATCH bpf-next v7 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
` (10 subsequent siblings)
21 siblings, 1 reply; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:16 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
bpf_unwind() walks the BPF frames and, for each one, rewrites the saved
return address so the frame resumes where the unwind needs it, then
returns. Nothing restores a register: every frame runs its own epilogue on
the way out, which is what puts its caller's r6-r9 back, so the unwind
needs no spill area and no per-frame metadata beyond the table itself.
Where a record covers the call a frame is suspended at, it resumes at that
pad, which the previous patch made sure ends in a resume. Where none does,
it resumes at the frame's epilogue and returns at once, its caller reached
with registers already restored. A pad's resume lowers to 'r0 = 0; exit',
so that frame returns too and the rewritten address carries the unwind on
to the next pad.
An epilogue therefore has to exist, for every frame the walk can pass
rather than only those carrying a table: aux->epilogue_ip is recorded for
every program, since a frame need not carry a table to be on the path, and
handed to the outer program with the table when jit_subprogs() compiles the
main program as func[0].
x86 emits the epilogue at a subprogram's first exit, so one the dead code
sweep leaves exitless gets none. Two shapes do that. The frame that called
bpf_unwind() loses the exit after the call, so one is patched back in
there, which is also what that frame returns through when no record covers
the call. And a frame above one that never comes back loses its exit with
no bpf_unwind() to hang a new one on, so the last exit of every subprogram
an unwind can pass through is kept, searched for since a subprogram may
end in a jump or a gotox. arm64 emits an epilogue either way.
Both kfuncs become callable here rather than earlier: until the walk and
the lowering exist, bpf_unwind() would return to instructions the verifier
never explored and bpf_unwind_resume() would reach its WARN_ONCE body.
arch_bpf_stack_walk_ra() hands out the return-address slot as well as the
address. It is a second entry point rather than a change to
arch_bpf_stack_walk(), so architectures that do not dispatch pads keep the
walker they have -- where it is the weak stub, the walk does nothing, so
process_bpf_unwind() now asks bpf_exc_check_prog() whether this program may
unwind at all. A cleanup table was held to that before the CFG walk; a
bpf_unwind() with no table had not been.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf.h | 37 ++++++++++
include/linux/bpf_verifier.h | 1 +
include/linux/filter.h | 2 +
kernel/bpf/core.c | 20 +++++-
kernel/bpf/exception.c | 86 ++++++++++++++++++++++
kernel/bpf/exception.h | 7 ++
kernel/bpf/fixups.c | 135 +++++++++++++++++++++++++++++++++++
kernel/bpf/helpers.c | 46 ++++++++++++
kernel/bpf/verifier.c | 12 ++++
9 files changed, 345 insertions(+), 1 deletion(-)
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 4bae3796c42f..94005cd3ad0f 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -1805,6 +1805,41 @@ enum bpf_sig_keyring {
BPF_SIG_KEYRING_BPF,
};
+/* One cleanup region of a JITed (sub)program. */
+struct bpf_cleanup_range {
+ u64 begin;
+ u64 end;
+ u64 pad;
+};
+
+struct bpf_exception_info {
+ struct bpf_cleanup_info *info;
+ struct bpf_cleanup_range *ranges;
+ u32 nr_info;
+ u32 nr_ranges;
+};
+
+#ifdef CONFIG_BPF_SYSCALL
+int bpf_exc_attach_main_prog(struct bpf_verifier_env *env, struct bpf_prog *prog);
+void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image);
+void bpf_exc_free_info(struct bpf_prog_aux *aux);
+#else
+
+static inline int bpf_exc_attach_main_prog(struct bpf_verifier_env *env,
+ struct bpf_prog *prog)
+{
+ return 0;
+}
+
+static inline void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+}
+
+static inline void bpf_exc_free_info(struct bpf_prog_aux *aux)
+{
+}
+#endif
+
struct bpf_prog_aux {
atomic64_t refcnt;
u32 used_map_cnt;
@@ -1885,6 +1920,8 @@ struct bpf_prog_aux {
u64 (*bpf_exception_cb)(u64 cookie, u64 sp, u64 bp, u64, u64);
u16 stack_arg_sp_adjust;
u16 freplace_link_cnt; /* counts freplace links extending this prog */
+ struct bpf_exception_info *exc;
+ u64 epilogue_ip; /* native address of this (sub)program's epilogue */
#ifdef CONFIG_SECURITY
void *security;
#endif
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index eb35aa4cfd37..e2944b415a25 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1855,6 +1855,7 @@ int bpf_opt_subreg_zext_lo32_rnd_hi32(struct bpf_verifier_env *env, const union
int bpf_convert_ctx_accesses(struct bpf_verifier_env *env);
int bpf_jit_subprogs(struct bpf_verifier_env *env);
int bpf_fixup_call_args(struct bpf_verifier_env *env);
+int bpf_exc_keep_exits(struct bpf_verifier_env *env);
int bpf_do_misc_fixups(struct bpf_verifier_env *env);
int bpf_insn_def32(struct bpf_prog *prog, struct bpf_insn *insn);
diff --git a/include/linux/filter.h b/include/linux/filter.h
index 972b3ed2a51d..0d7d949a1baa 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1290,6 +1290,8 @@ u32 bpf_jit_plan_arg_moves(const struct bpf_jit_arg_abi *abi,
struct bpf_jit_arg_move *moves);
u64 bpf_arch_uaddress_limit(void);
void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp), void *cookie);
+void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
+ void *cookie);
u64 arch_bpf_timed_may_goto(void);
u64 bpf_check_timed_may_goto(struct bpf_timed_may_goto *);
bool bpf_helper_changes_pkt_data(enum bpf_func_id func_id);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index d813fdde29e3..60905643cb9c 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -292,6 +292,7 @@ void __bpf_prog_free(struct bpf_prog *fp)
mutex_destroy(&fp->aux->dst_mutex);
mutex_destroy(&fp->aux->st_ops_assoc_mutex);
kfree(fp->aux->poke_tab);
+ bpf_exc_free_info(fp->aux);
kfree(fp->aux);
}
free_percpu(fp->stats);
@@ -2632,9 +2633,14 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
{
#ifdef CONFIG_BPF_JIT
struct bpf_prog *orig_prog;
+ int ret;
- if (!bpf_prog_need_blind(prog))
+ if (!bpf_prog_need_blind(prog)) {
+ ret = bpf_exc_attach_main_prog(env, prog);
+ if (ret)
+ return ERR_PTR(ret);
return bpf_int_jit_compile(env, prog);
+ }
orig_prog = prog;
prog = bpf_jit_blind_constants(env, prog);
@@ -2648,6 +2654,12 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
goto out_restore;
}
+ ret = bpf_exc_attach_main_prog(env, prog);
+ if (ret) {
+ bpf_jit_prog_release_other(orig_prog, prog);
+ return ERR_PTR(ret);
+ }
+
prog = bpf_int_jit_compile(env, prog);
if (prog->jited) {
bpf_jit_prog_release_other(prog, orig_prog);
@@ -3511,6 +3523,12 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
{
}
+void __weak arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp,
+ u64 *ra),
+ void *cookie)
+{
+}
+
bool __weak bpf_jit_supports_cleanup_pads(void)
{
return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index f2bca0242408..4c1c4b9cf6c7 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -5,6 +5,7 @@
#include <linux/btf.h>
#include <linux/btf_ids.h>
#include <linux/filter.h>
+#include <linux/slab.h>
#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -266,3 +267,88 @@ int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
return pad ? (int)pad - 1 : -1;
}
+
+/*
+ * The record covering @ip, which is a return address: the call it belongs to
+ * is the instruction before it, so a range matches on begin < ip <= end.
+ */
+const struct bpf_cleanup_range *bpf_exc_pad_for_ip(const struct bpf_prog *prog, u64 ip)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+ u32 l = 0, r = exc ? exc->nr_ranges : 0;
+
+ while (l < r) {
+ u32 m = l + (r - l) / 2;
+ const struct bpf_cleanup_range *rec = &exc->ranges[m];
+
+ if (ip <= rec->begin)
+ r = m;
+ else if (ip > rec->end)
+ l = m + 1;
+ else
+ return rec;
+ }
+ return NULL;
+}
+
+int bpf_exc_alloc_info(struct bpf_prog_aux *aux)
+{
+ if (aux->exc)
+ return 0;
+ aux->exc = kzalloc_obj(struct bpf_exception_info, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ return aux->exc ? 0 : -ENOMEM;
+}
+
+int bpf_exc_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt)
+{
+ struct bpf_exception_info *exc = aux->exc;
+ struct bpf_cleanup_range *ranges;
+
+ ranges = kvcalloc(cnt, sizeof(*ranges), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!ranges) {
+ kvfree(recs);
+ return -ENOMEM;
+ }
+
+ exc->info = recs;
+ exc->nr_info = cnt;
+ exc->ranges = ranges;
+ /* Withheld until the JIT has filled the table in. */
+ exc->nr_ranges = 0;
+ return 0;
+}
+
+void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+ struct bpf_exception_info *exc = prog->aux->exc;
+ u32 i, n;
+
+ if (!exc || !exc->nr_info || !exc->ranges)
+ return;
+
+ n = exc->nr_info;
+ for (i = 0; i < n; i++) {
+ const struct bpf_cleanup_info *rec = &exc->info[i];
+
+ if (WARN_ON_ONCE(rec->begin_off >= prog->len ||
+ rec->end_off > prog->len ||
+ rec->landing_pad_off >= prog->len))
+ return;
+ exc->ranges[i].begin = (u64)(long)image + addrs[rec->begin_off];
+ exc->ranges[i].end = (u64)(long)image + addrs[rec->end_off];
+ exc->ranges[i].pad = (u64)(long)image + addrs[rec->landing_pad_off];
+ }
+ exc->nr_ranges = n;
+}
+
+void bpf_exc_free_info(struct bpf_prog_aux *aux)
+{
+ struct bpf_exception_info *exc = aux->exc;
+
+ if (!exc)
+ return;
+ kvfree(exc->ranges);
+ kvfree(exc->info);
+ kfree(exc);
+ aux->exc = NULL;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index c5b30ff3ccc0..47a6870d782f 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -9,6 +9,10 @@ struct bpf_verifier_env;
struct bpf_verifier_state;
struct bpf_func_state;
struct bpf_insn;
+struct bpf_cleanup_info;
+struct bpf_cleanup_range;
+struct bpf_prog;
+struct bpf_prog_aux;
int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
int bpf_exc_check_prog(struct bpf_verifier_env *env);
@@ -19,5 +23,8 @@ int bpf_exc_check_frame_balance(struct bpf_verifier_env *env, const char *prefix
int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog);
int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn);
+int bpf_exc_alloc_info(struct bpf_prog_aux *aux);
+int bpf_exc_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt);
+const struct bpf_cleanup_range *bpf_exc_pad_for_ip(const struct bpf_prog *prog, u64 ip);
#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 5b7fe4ba610b..bff9d2539371 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -1,5 +1,6 @@
// SPDX-License-Identifier: GPL-2.0-only
/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bitmap.h>
#include <linux/bpf.h>
#include <linux/btf.h>
#include <linux/bpf_verifier.h>
@@ -11,6 +12,7 @@
#include <linux/sched/signal.h>
#include <net/xdp.h>
#include "disasm.h"
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -739,6 +741,26 @@ static void keep_funcs_with_addr_taken(struct bpf_verifier_env *env)
}
}
+/* A JIT emits the epilogue an unwind aims at only at an exit, so keep one. */
+static void keep_subprog_exits(struct bpf_verifier_env *env)
+{
+ u32 i, j;
+
+ for (i = 0; i < env->subprog_cnt; i++) {
+ u32 start;
+
+ if (!env->subprog_info[i].might_unwind)
+ continue;
+ start = env->subprog_info[i].start;
+ for (j = env->subprog_info[i + 1].start; j-- > start; ) {
+ if (env->prog->insnsi[j].code != (BPF_JMP | BPF_EXIT))
+ continue;
+ env->insn_aux_data[j].seen = env->pass_cnt;
+ break;
+ }
+ }
+}
+
int bpf_opt_remove_dead_code(struct bpf_verifier_env *env)
{
struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
@@ -746,6 +768,7 @@ int bpf_opt_remove_dead_code(struct bpf_verifier_env *env)
int i, err;
keep_funcs_with_addr_taken(env);
+ keep_subprog_exits(env);
for (i = 0; i < insn_cnt; i++) {
int j;
@@ -1286,6 +1309,61 @@ static int resolve_func_ptrs(struct bpf_verifier_env *env, struct bpf_prog *prog
return 0;
}
+static int exc_info_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
+ u32 subprog, u32 start, u32 end)
+{
+ struct bpf_cleanup_info *recs;
+ u32 i, cnt = 0;
+ int err;
+
+ if (!env->cleanup_info_cnt)
+ return 0;
+
+ err = bpf_exc_alloc_info(sub->aux);
+ if (err)
+ return err;
+
+ for (i = start; i < end; i++) {
+ if (env->insn_aux_data[i].cleanup_pad)
+ cnt++;
+ }
+ if (!cnt)
+ return 0;
+
+ recs = kvmalloc_array(cnt, sizeof(*recs), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!recs)
+ return -ENOMEM;
+
+ for (i = start, cnt = 0; i < end; i++) {
+ u32 pad = env->insn_aux_data[i].cleanup_pad;
+
+ if (!pad)
+ continue;
+ pad--;
+ if (verifier_bug_if(pad < start || pad >= end, env,
+ "insn %u is covered by a landing pad at %u outside its subprog [%u, %u)",
+ i, pad, start, end)) {
+ kvfree(recs);
+ return -EFAULT;
+ }
+ recs[cnt].begin_off = i - start;
+ recs[cnt].end_off = i - start + 1;
+ recs[cnt].landing_pad_off = pad - start;
+ cnt++;
+ }
+ err = bpf_exc_attach_info(sub->aux, recs, cnt);
+ if (err)
+ return err;
+ return 0;
+}
+
+int bpf_exc_attach_main_prog(struct bpf_verifier_env *env, struct bpf_prog *prog)
+{
+ if (!env || env->subprog_cnt > 1)
+ return 0;
+ return exc_info_for_subprog(env, prog, 0, 0, prog->len);
+}
+
static int jit_subprogs(struct bpf_verifier_env *env)
{
struct bpf_prog *prog = env->prog, **func, *tmp;
@@ -1423,6 +1501,10 @@ static int jit_subprogs(struct bpf_verifier_env *env)
func[i]->aux->token = prog->aux->token;
if (!i)
func[i]->aux->exception_boundary = env->seen_exception;
+ err = exc_info_for_subprog(env, func[i], i, subprog_start,
+ subprog_end);
+ if (err)
+ goto out_free;
func[i] = bpf_int_jit_compile(env, func[i]);
if (!func[i]->jited) {
err = -ENOTSUPP;
@@ -1532,6 +1614,9 @@ static int jit_subprogs(struct bpf_verifier_env *env)
prog->aux->bpf_exception_cb = (void *)func[env->exception_callback_subprog]->bpf_func;
prog->aux->exception_boundary = func[0]->aux->exception_boundary;
prog->aux->stack_arg_sp_adjust = func[0]->aux->stack_arg_sp_adjust;
+ prog->aux->exc = func[0]->aux->exc;
+ func[0]->aux->exc = NULL;
+ prog->aux->epilogue_ip = func[0]->aux->epilogue_ip;
bpf_prog_jit_attempt_done(prog);
return 0;
out_free:
@@ -1757,6 +1842,37 @@ static int may_goto_expand(struct bpf_insn *insn_buf, int off, int stack_off,
return cnt + tail_cnt;
}
+/*
+ * Put an exit back after every bpf_unwind() call. Nothing reaches it, but it
+ * keeps the frame's epilogue, which is where the unwind sends a frame that
+ * has no landing pad.
+ */
+int bpf_exc_keep_exits(struct bpf_verifier_env *env)
+{
+ int insn_cnt = env->prog->len;
+ struct bpf_insn insn_buf[3];
+ struct bpf_prog *new_prog;
+ int i, delta = 0;
+
+ for (i = 0; i < insn_cnt; i++) {
+ struct bpf_insn *insn = env->prog->insnsi + i + delta;
+
+ if (!bpf_is_unwind_kfunc(insn))
+ continue;
+
+ insn_buf[0] = *insn;
+ insn_buf[1] = BPF_MOV64_IMM(BPF_REG_0, 0);
+ insn_buf[2] = BPF_EXIT_INSN();
+
+ new_prog = bpf_patch_insn_data(env, i + delta, insn_buf, 3);
+ if (!new_prog)
+ return -ENOMEM;
+ delta += 2;
+ env->prog = new_prog;
+ }
+ return 0;
+}
+
/* Do various post-verification rewrites in a single program pass.
* These rewrites simplify JIT and interpreter implementations.
*/
@@ -2135,6 +2251,25 @@ int bpf_do_misc_fixups(struct bpf_verifier_env *env)
goto next_insn;
if (insn->src_reg == BPF_PSEUDO_CALL)
goto next_insn;
+ if (bpf_is_unwind_resume_kfunc(insn)) {
+ /*
+ * A pad's resume is just the frame returning:
+ * bpf_unwind() already pointed this frame's return
+ * address at the next pad, so the ordinary epilogue
+ * carries the unwind on. The verifier checked this exit
+ * with r0 a known zero, so return zero.
+ */
+ insn_buf[0] = BPF_MOV64_IMM(BPF_REG_0, 0);
+ insn_buf[1] = BPF_EXIT_INSN();
+ cnt = 2;
+ new_prog = bpf_patch_insn_data(env, i + delta, insn_buf, cnt);
+ if (!new_prog)
+ return -ENOMEM;
+ delta += cnt - 1;
+ env->prog = prog = new_prog;
+ insn = new_prog->insnsi + i + delta;
+ goto next_insn;
+ }
if (insn->src_reg == BPF_PSEUDO_KFUNC_CALL) {
ret = bpf_fixup_kfunc_call(env, insn, insn_buf, i + delta, &cnt);
if (ret)
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 08aee86a155c..28d4b4e22ea3 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -31,6 +31,7 @@
#include <linux/buildid.h>
#include "../../lib/kstrtox.h"
+#include "exception.h"
/* If kernel subsystem is allowing eBPF programs to call this function,
* inside its own verifier_ops->get_func_proto() callback it should return
@@ -3424,8 +3425,51 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
return false;
}
+struct bpf_unwind_ctx {
+ u32 cnt;
+};
+
+static bool bpf_unwind_rewrite(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra)
+{
+ const struct bpf_cleanup_range *rec;
+ struct bpf_unwind_ctx *ctx = cookie;
+ struct bpf_exception_info *exc;
+ struct bpf_prog *prog;
+
+ rcu_read_lock();
+ prog = bpf_prog_ksym_find(ip);
+ rcu_read_unlock();
+ if (!prog)
+ return !ctx->cnt;
+ ctx->cnt++;
+
+ exc = prog->aux->exc;
+ rec = (exc && exc->nr_ranges) ? bpf_exc_pad_for_ip(prog, ip) : NULL;
+ if (rec) {
+ *ra = rec->pad;
+ } else if (ctx->cnt == 1) {
+ /*
+ * The frame that called bpf_unwind(). Its return address
+ * always names the 'r0 = 0; exit' that bpf_exc_keep_exits()
+ * put after the call, so leave it alone and let the frame
+ * return through that: running it is what sets the value
+ * the unwind returns.
+ */
+ } else if (prog->aux->epilogue_ip) {
+ *ra = prog->aux->epilogue_ip;
+ } else {
+ WARN_ON_ONCE(1);
+ return false;
+ }
+
+ return bpf_is_subprog(prog);
+}
+
__bpf_kfunc void bpf_unwind(void)
{
+ struct bpf_unwind_ctx ctx = {};
+
+ arch_bpf_stack_walk_ra(bpf_unwind_rewrite, &ctx);
}
__bpf_kfunc void bpf_throw(u64 cookie)
@@ -5095,6 +5139,8 @@ BTF_ID_FLAGS(func, bpf_task_get_cgroup1, KF_ACQUIRE | KF_RCU | KF_RET_NULL)
BTF_ID_FLAGS(func, bpf_task_from_pid, KF_ACQUIRE | KF_RET_NULL)
BTF_ID_FLAGS(func, bpf_task_from_vpid, KF_ACQUIRE | KF_RET_NULL)
BTF_ID_FLAGS(func, bpf_throw)
+BTF_ID_FLAGS(func, bpf_unwind)
+BTF_ID_FLAGS(func, bpf_unwind_resume)
#ifdef CONFIG_BPF_EVENTS
BTF_ID_FLAGS(func, bpf_send_signal_task)
#endif
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 0419609beb67..fbbfc4710a1f 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -19257,6 +19257,15 @@ static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
int pad = bpf_exc_pad_of_call(env, *insn_idx);
int err;
+ /*
+ * A table was held to this before the CFG walk; an unwind with no
+ * table reaches the same gate here, since the walk it needs exists on
+ * only some architectures.
+ */
+ err = bpf_exc_check_prog(env);
+ if (err)
+ return err;
+
if (pad < 0) {
if (!env->cur_state->curframe) {
err = check_resource_leak(env, false, true,
@@ -22868,6 +22877,9 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
/* program is valid, convert *(u32*)(ctx + off) accesses */
ret = bpf_convert_ctx_accesses(env);
+ if (ret == 0)
+ ret = bpf_exc_keep_exits(env);
+
if (ret == 0)
ret = bpf_do_misc_fixups(env);
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 12/22] bpf, x86: Dispatch exception cleanup pads at run time
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (10 preceding siblings ...)
2026-09-29 0:16 ` [PATCH bpf-next v7 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
@ 2026-09-29 0:17 ` Yonghong Song
2026-09-29 0:30 ` sashiko-bot
2026-09-29 0:17 ` [PATCH bpf-next v7 13/22] bpf, arm64: " Yonghong Song
` (9 subsequent siblings)
21 siblings, 1 reply; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:17 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
arch_bpf_stack_walk_ra() is the ORC walk with the return-address slot
alongside each frame, which unwind_get_return_address_ptr() already hands
out; writing there is what redirects a frame to its landing pad. The
feature is gated on CONFIG_UNWINDER_ORC for the same reason
arch_bpf_stack_walk() is -- there is no other unwinder here to ask.
A frame whose return a tracer has hooked is left alone, and the walk stops
there. The unwinder recovers the address such a frame will really return
to, but the slot still holds the function graph or kretprobe trampoline,
so writing it would skip the trampoline and leave its entry for the next
hooked return to pop. x86 has no flag saying a frame was hooked, so the
recovered address and the slot are compared instead.
aux->epilogue_ip comes for free: the JIT already emits one epilogue per
(sub)program and every other exit jumps to it, so the offset it keeps as
ctx->cleanup_addr, as a native address, is it. That field has always been
the epilogue and has nothing to do with the cleanup pads despite the name.
The rest is bookkeeping: build the native cleanup table from the JIT's
addrs[] once the image is final. A pad head needs no ENDBR of its own: it
is only ever reached as a return address, and IBT checks indirect jumps
and calls, not returns.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
arch/x86/net/bpf_jit_comp.c | 42 +++++++++++++++++++++++++++++++++++++
1 file changed, 42 insertions(+)
diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index 6c7a0578760e..d4feade5b5c7 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -3278,6 +3278,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
seen_exit = true;
/* Update cleanup_addr */
ctx->cleanup_addr = proglen;
+ /* Where an unwind sends a frame with no pad. */
+ bpf_prog->aux->epilogue_ip = (u64)image + proglen;
if (bpf_prog_was_classic(bpf_prog) &&
!ns_capable_noaudit(&init_user_ns, CAP_SYS_ADMIN)) {
if (emit_spectre_bhb_barrier(&prog, ip, bpf_prog))
@@ -4455,6 +4457,13 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
*/
bpf_prog_update_insn_ptrs(prog, addrs, image);
+ /*
+ * Same mapping, consumed by the bpf_unwind() walk:
+ * turn the cleanup records into native address ranges now
+ * that the image is final.
+ */
+ bpf_exc_fill_native_ranges(prog, addrs, image);
+
/*
* ctx.prog_offset is used when CFI preambles put code *before*
* the function. See emit_cfi(). For FineIBT specifically this code
@@ -4593,6 +4602,11 @@ bool bpf_jit_supports_exceptions(void)
return IS_ENABLED(CONFIG_UNWINDER_ORC);
}
+bool bpf_jit_supports_cleanup_pads(void)
+{
+ return IS_ENABLED(CONFIG_UNWINDER_ORC);
+}
+
bool bpf_jit_supports_private_stack(void)
{
return true;
@@ -4614,6 +4628,34 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
#endif
}
+void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
+ void *cookie)
+{
+#if defined(CONFIG_UNWINDER_ORC)
+ struct unwind_state state;
+ unsigned long addr, *ra;
+
+ for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
+ unwind_next_frame(&state)) {
+ addr = unwind_get_return_address(&state);
+ ra = unwind_get_return_address_ptr(&state);
+ if (!addr || !ra)
+ break;
+ /*
+ * A traced return: the unwinder recovered @addr from under a
+ * function graph or kretprobe trampoline, which is what the
+ * slot itself still holds. Writing there would skip the
+ * trampoline and leave its entry for the next hooked return
+ * to pop.
+ */
+ if (READ_ONCE_NOCHECK(*ra) != addr)
+ break;
+ if (!consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
+ break;
+ }
+#endif
+}
+
void bpf_arch_poke_desc_update(struct bpf_jit_poke_descriptor *poke,
struct bpf_prog *new, struct bpf_prog *old)
{
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 13/22] bpf, arm64: Dispatch exception cleanup pads at run time
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (11 preceding siblings ...)
2026-09-29 0:17 ` [PATCH bpf-next v7 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
@ 2026-09-29 0:17 ` Yonghong Song
2026-09-29 1:14 ` bot+bpf-ci
2026-09-29 0:17 ` [PATCH bpf-next v7 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
` (8 subsequent siblings)
21 siblings, 1 reply; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:17 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
The JIT half: build the native cleanup table from the JIT's byte offsets
once the image is final, and record the one epilogue so a frame the unwind
passes over can return through it. A pad head needs no BTI of its own: it
is only ever reached as a return address, and a return sets no BTYPE, so
no branch-target check is made.
The dispatch is arch_bpf_stack_walk_ra(), which hands the unwind the slot a
frame's return address came out of rather than just the address. arm64's
unwinder reads that address from the frame record the callee pushed, so the
slot belongs to the record the previous entry stepped through.
Writing it has to respect pointer authentication: a BPF prologue signs the
link register with PACIASP and the epilogue authenticates it, so what goes
back has to carry the same signature. Its modifier is the stack pointer the
owner was entered with, which is not known here, so recover it by
re-signing the address the unwinder stripped until that matches the slot.
Whether a slot is signed is asked of the build rather than read off the
value: CONFIG_ARM64_PTR_AUTH_KERNEL is what the prologue signs under and
what -mbranch-protection is added for. Reading it off the value instead
would take a signed address for an unsigned one whenever its PAC equalled
the bits stripping puts back. The CPU has to implement address
authentication too, since "pacia Xd, Xn" is not in the HINT space and
would be undefined without it.
Two frames are not redirected: the walk's own first frame, not returning
anywhere yet, and one the function graph tracer or a kretprobe has hooked,
whose slot holds the trampoline rather than the address the unwinder
reports -- the walk stops there.
bpf_jit_supports_cleanup_pads() can now say yes, except where a shadow
call stack is in use. JITed code restores x30 from the frame record this
walk rewrites, but bpf_unwind() is C and returns from its x18 copy
instead, so the frame that called it would carry on as though nothing had
happened while the frames above it resumed at their pads.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
arch/arm64/kernel/stacktrace.c | 103 +++++++++++++++++++++++++++++++++
arch/arm64/net/bpf_jit_comp.c | 26 +++++++++
2 files changed, 129 insertions(+)
diff --git a/arch/arm64/kernel/stacktrace.c b/arch/arm64/kernel/stacktrace.c
index 3ebcf8c53fb0..1e46a22cafbd 100644
--- a/arch/arm64/kernel/stacktrace.c
+++ b/arch/arm64/kernel/stacktrace.c
@@ -445,6 +445,109 @@ noinline noinstr void arch_bpf_stack_walk(bool (*consume_entry)(void *cookie, u6
kunwind_stack_walk(arch_bpf_unwind_consume_entry, &data, current, NULL);
}
+struct bpf_unwind_ra_consume_entry_data {
+ bool (*consume_entry)(void *cookie, u64 ip, u64 sp, u64 fp, u64 *ra);
+ void *cookie;
+ unsigned long record;
+ bool seen_first;
+};
+
+static u64 bpf_unwind_sign_ra(u64 ra, u64 modifier)
+{
+ asm volatile(ARM64_ASM_PREAMBLE
+ ".arch_extension pauth\n"
+ " pacia %0, %1"
+ : "+r" (ra) : "r" (modifier));
+ return ra;
+}
+
+/*
+ * PACIASP's modifier is the stack pointer the owner was entered with: record
+ * + 16 for a BPF prologue, but further up for bpf_unwind()'s own C frame.
+ * Recognise it by re-signing @pc, which the unwinder stripped from @stored.
+ */
+static bool bpf_unwind_ra_modifier(unsigned long record, unsigned long caller_fp,
+ u64 stored, u64 pc, u64 *modifier)
+{
+ u64 m;
+
+ for (m = record + sizeof(struct frame_record); m <= caller_fp; m += 16) {
+ if (bpf_unwind_sign_ra(pc, m) == stored) {
+ *modifier = m;
+ return true;
+ }
+ }
+ return false;
+}
+
+static bool bpf_unwind_store_ra(unsigned long record, unsigned long caller_fp,
+ u64 pc, u64 ra)
+{
+ struct frame_record *rec = (struct frame_record *)record;
+
+ /*
+ * Whether the slot holds a signed address is a property of the build,
+ * not one to be read off the value: a PAC can come out equal to the
+ * bits stripping puts back, and a signed address would then be taken
+ * for an unsigned one. What signs is CONFIG_ARM64_PTR_AUTH_KERNEL --
+ * the prologue here, and -mbranch-protection for everything the
+ * compiler emits.
+ */
+ if (IS_ENABLED(CONFIG_ARM64_PTR_AUTH_KERNEL) &&
+ system_supports_address_auth()) {
+ u64 stored = READ_ONCE(rec->lr);
+ u64 modifier;
+
+ if (WARN_ON_ONCE(!bpf_unwind_ra_modifier(record, caller_fp,
+ stored, pc, &modifier)))
+ return false;
+ ra = bpf_unwind_sign_ra(ra, modifier);
+ }
+ WRITE_ONCE(rec->lr, ra);
+ return true;
+}
+
+static bool
+arch_bpf_unwind_ra_consume_entry(const struct kunwind_state *state, void *cookie)
+{
+ struct bpf_unwind_ra_consume_entry_data *data = cookie;
+ unsigned long record = data->record;
+ bool seen_first = data->seen_first;
+ u64 ra = state->common.pc;
+ bool cont;
+
+ /* The record this frame's return address will have come out of. */
+ data->record = state->common.fp;
+ data->seen_first = true;
+
+ /* The first pc is where the walk runs, not an address it returns to. */
+ if (!seen_first)
+ return true;
+ /* A traced return: the slot holds the tracer's trampoline, not @pc. */
+ if (state->flags.fgraph || state->flags.kretprobe)
+ return false;
+
+ /* A consumer that stops still gets to redirect the frame it stopped on. */
+ cont = data->consume_entry(data->cookie, state->common.pc, 0,
+ state->common.fp, &ra);
+ if (ra != state->common.pc &&
+ !bpf_unwind_store_ra(record, state->common.fp, state->common.pc, ra))
+ return false;
+ return cont;
+}
+
+noinline noinstr void arch_bpf_stack_walk_ra(bool (*consume_entry)(void *cookie, u64 ip, u64 sp,
+ u64 fp, u64 *ra),
+ void *cookie)
+{
+ struct bpf_unwind_ra_consume_entry_data data = {
+ .consume_entry = consume_entry,
+ .cookie = cookie,
+ };
+
+ kunwind_stack_walk(arch_bpf_unwind_ra_consume_entry, &data, current, NULL);
+}
+
static const char *state_source_string(const struct kunwind_state *state)
{
switch (state->source) {
diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
index 475e70653454..e483e1e7a2d4 100644
--- a/arch/arm64/net/bpf_jit_comp.c
+++ b/arch/arm64/net/bpf_jit_comp.c
@@ -10,10 +10,12 @@
#include <linux/arm-smccc.h>
#include <linux/bitfield.h>
#include <linux/bpf.h>
+#include <linux/bpf_verifier.h>
#include <linux/cfi.h>
#include <linux/filter.h>
#include <linux/memory.h>
#include <linux/printk.h>
+#include <linux/scs.h>
#include <linux/slab.h>
#include <asm/asm-extable.h>
@@ -2423,6 +2425,17 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
* reasons, expects to point to the next instruction)
*/
bpf_prog_update_insn_ptrs(prog, ctx.offset, ctx.ro_image);
+
+ /*
+ * Same byte offsets, consumed by the bpf_unwind() walk:
+ * turn the cleanup records into native address ranges now that
+ * the image is final.
+ */
+ bpf_exc_fill_native_ranges(prog, ctx.offset, ctx.ro_image);
+
+ /* Where an unwind sends a frame with no pad. */
+ prog->aux->epilogue_ip = (u64)ctx.ro_image +
+ ctx.epilogue_offset * AARCH64_INSN_SIZE;
out_off:
if (!ro_header && priv_stack_ptr) {
free_percpu(priv_stack_ptr);
@@ -3408,6 +3421,19 @@ bool bpf_jit_supports_exceptions(void)
return true;
}
+bool bpf_jit_supports_cleanup_pads(void)
+{
+ /*
+ * An unwind redirects a frame by rewriting the frame record its
+ * callee's return address came out of. JITed code restores x30 from
+ * there, but bpf_unwind() is C: with a shadow call stack it returns
+ * from the x18 copy instead, so the frame that called it would keep
+ * going as if nothing had happened while the frames above it resumed
+ * at their pads.
+ */
+ return !scs_is_enabled();
+}
+
bool bpf_jit_supports_arena(void)
{
return true;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (12 preceding siblings ...)
2026-09-29 0:17 ` [PATCH bpf-next v7 13/22] bpf, arm64: " Yonghong Song
@ 2026-09-29 0:17 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
` (7 subsequent siblings)
21 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:17 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
LLVM terminates a cleanup landing pad with a call to _Unwind_Resume: the
base unwind ABI's entry point for carrying an unwind on once a frame's
cleanups have run. The kernel provides that terminator as a kfunc, but not
under that name: claiming _Unwind_Resume in the kernel's own symbol table,
for a function whose body never runs, would be needlessly confusing, so it
is called bpf_unwind_resume.
The compiler emits the call but not the declaration: _Unwind_Resume lands
in the object as a plain undefined symbol, with nothing in .ksyms. libbpf
takes every undefined NOTYPE symbol for an extern and refuses one it has no
BTF for ("failed to find BTF for extern '_Unwind_Resume'"), so the program
declares it itself -- extern void _Unwind_Resume(void *) __ksym; -- as the
selftests here do, and as a language runtime emitting cleanup pads has to.
What follows translates that name; it does not manufacture the declaration.
Both load paths take the detour. A direct load resolves the name against
the kernel's BTF while libbpf runs. A light skeleton instead writes the
name into the loader program's blob of bytes, for that program to resolve
when it runs, so the name recorded there has to be translated as well.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/libbpf.c | 23 ++++++++++++++++++-----
1 file changed, 18 insertions(+), 5 deletions(-)
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index 238e35c1ee3f..7eb15a013a83 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -8797,6 +8797,16 @@ static void fixup_verifier_log(struct bpf_program *prog, char *buf, size_t buf_s
}
}
+/*
+ * LLVM terminates a cleanup landing pad with a call to _Unwind_Resume, the
+ * base unwind ABI's entry point for carrying an unwind on once a frame's
+ * cleanups have run. The kernel knows it as bpf_unwind_resume.
+ */
+static const char *kern_extern_name(const char *name)
+{
+ return strcmp(name, "_Unwind_Resume") ? name : "bpf_unwind_resume";
+}
+
static int bpf_program_record_relos(struct bpf_program *prog)
{
struct bpf_object *obj = prog->obj;
@@ -8813,12 +8823,12 @@ static int bpf_program_record_relos(struct bpf_program *prog)
continue;
kind = btf_is_var(btf__type_by_id(obj->btf, ext->btf_id)) ?
BTF_KIND_VAR : BTF_KIND_FUNC;
- bpf_gen__record_extern(obj->gen_loader, ext->name,
+ bpf_gen__record_extern(obj->gen_loader, kern_extern_name(ext->name),
ext->is_weak, !ext->ksym.type_id,
true, kind, relo->insn_idx);
break;
case RELO_EXTERN_CALL:
- bpf_gen__record_extern(obj->gen_loader, ext->name,
+ bpf_gen__record_extern(obj->gen_loader, kern_extern_name(ext->name),
ext->is_weak, false, false, BTF_KIND_FUNC,
relo->insn_idx);
break;
@@ -9294,17 +9304,20 @@ static int bpf_object__resolve_ksym_func_btf_id(struct bpf_object *obj,
struct module_btf *mod_btf = NULL;
const struct btf_type *kern_func;
struct btf *kern_btf = NULL;
+ const char *local_name, *kern_name;
int ret;
local_func_proto_id = ext->ksym.type_id;
- kfunc_id = find_ksym_btf_id(obj, ext->essent_name ?: ext->name, BTF_KIND_FUNC, &kern_btf,
- &mod_btf);
+ local_name = ext->essent_name ?: ext->name;
+ kern_name = kern_extern_name(local_name);
+
+ kfunc_id = find_ksym_btf_id(obj, kern_name, BTF_KIND_FUNC, &kern_btf, &mod_btf);
if (kfunc_id < 0) {
if (kfunc_id == -ESRCH && ext->is_weak)
return 0;
pr_warn("extern (func ksym) '%s': not found in kernel or module BTFs\n",
- ext->name);
+ kern_name != local_name ? kern_name : ext->name);
return kfunc_id;
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (13 preceding siblings ...)
2026-09-29 0:17 ` [PATCH bpf-next v7 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
@ 2026-09-29 0:17 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
` (6 subsequent siblings)
21 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:17 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Let a caller hand the kernel an exception cleanup table: struct
bpf_prog_load_opts grows cleanup_info, cleanup_info_cnt and
cleanup_info_rec_size, and bpf_prog_load() passes all three on to
BPF_PROG_LOAD.
The attr size it computes grows by one field rather than three:
cleanup_info_cnt is the last of them in the BPF_PROG_LOAD attr, so
offsetofend() on it already covers the other two.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/bpf.c | 6 +++++-
tools/lib/bpf/bpf.h | 7 ++++++-
2 files changed, 11 insertions(+), 2 deletions(-)
diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
index b49822d212ae..2e761f020b14 100644
--- a/tools/lib/bpf/bpf.c
+++ b/tools/lib/bpf/bpf.c
@@ -295,7 +295,7 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
const struct bpf_insn *insns, size_t insn_cnt,
struct bpf_prog_load_opts *opts)
{
- const size_t attr_sz = offsetofend(union bpf_attr, keyring_id);
+ const size_t attr_sz = offsetofend(union bpf_attr, cleanup_info_cnt);
void *finfo = NULL, *linfo = NULL;
const char *func_info, *line_info;
__u32 log_size, log_level, attach_prog_fd, attach_btf_obj_fd;
@@ -370,6 +370,10 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
attr.fd_array = ptr_to_u64(OPTS_GET(opts, fd_array, NULL));
attr.fd_array_cnt = OPTS_GET(opts, fd_array_cnt, 0);
+ attr.cleanup_info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
+ attr.cleanup_info = ptr_to_u64(OPTS_GET(opts, cleanup_info, NULL));
+ attr.cleanup_info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
+
if (log_level) {
attr.log_buf = ptr_to_u64(log_buf);
attr.log_size = log_size;
diff --git a/tools/lib/bpf/bpf.h b/tools/lib/bpf/bpf.h
index 826d9cc9ab65..2c292d56c9a3 100644
--- a/tools/lib/bpf/bpf.h
+++ b/tools/lib/bpf/bpf.h
@@ -128,9 +128,14 @@ struct bpf_prog_load_opts {
/* if set, provides the length of fd_array */
__u32 fd_array_cnt;
+
+ /* exception cleanup table, from the .bpf_cleanup section */
+ const void *cleanup_info;
+ __u32 cleanup_info_cnt;
+ __u32 cleanup_info_rec_size;
size_t :0;
};
-#define bpf_prog_load_opts__last_field fd_array_cnt
+#define bpf_prog_load_opts__last_field cleanup_info_rec_size
LIBBPF_API int bpf_prog_load(enum bpf_prog_type prog_type,
const char *prog_name, const char *license,
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (14 preceding siblings ...)
2026-09-29 0:17 ` [PATCH bpf-next v7 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
@ 2026-09-29 0:17 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
` (5 subsequent siblings)
21 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:17 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Parse the compiler-emitted .bpf_cleanup section and hand the resulting
table to BPF_PROG_LOAD.
Each record is three 4-byte fields, and each field is a byte offset into
some code section named by a matching .rel.bpf_cleanup relocation.
bpf_object__init_cleanup_info() resolves both halves once at open time and
keeps (section, instruction index) pairs; it rejects a record whose field
has no relocation, whose relocation is not one of the two 32-bit types that
spell a data reference to a code section -- R_BPF_64_NODYLD32 from LLVM,
R_BPF_64_ABS32 from GNU as -- or whose offset is not instruction aligned.
LLVM leaves the section's sh_entsize unset, which is taken to mean the size
assumed here, so a section declaring any other is refused.
The records are sorted by begin_off once the offsets are final, since the
kernel wants the table sorted with disjoint ranges to find the record
covering a call site with a binary search, and they arrive in .bpf_cleanup
order, which says nothing about where the subprograms they describe were
appended. Overlapping ranges are reported here, where the program name and
both regions are still at hand.
bpf_object_load_prog() then passes the per-program table through the
bpf_prog_load() options added in the previous patch, with the record size
carried on the program the way func_info and line_info carry theirs.
bpf_program__clone() carries it too -- that is the load path veristat uses,
and without the table the kernel sees landing pads nothing reaches and
refuses the program with "unreachable insn".
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/libbpf.c | 310 ++++++++++++++++++++++++++++++++
tools/lib/bpf/libbpf_internal.h | 3 +
2 files changed, 313 insertions(+)
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index 7eb15a013a83..872601183112 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -516,6 +516,11 @@ struct bpf_program {
void *line_info;
__u32 line_info_rec_size;
__u32 line_info_cnt;
+
+ struct bpf_cleanup_info *cleanup_info;
+ __u32 cleanup_info_rec_size;
+ __u32 cleanup_info_cnt;
+
__u32 prog_flags;
__u8 hash[SHA256_DIGEST_LENGTH];
@@ -552,6 +557,7 @@ struct bpf_struct_ops {
#define STRUCT_OPS_SEC ".struct_ops"
#define STRUCT_OPS_LINK_SEC ".struct_ops.link"
#define ARENA_SEC ".addr_space.1"
+#define CLEANUP_SEC ".bpf_cleanup"
enum libbpf_map_type {
LIBBPF_MAP_UNSPEC,
@@ -683,6 +689,25 @@ struct elf_sec_desc {
Elf_Data *data;
};
+#define CLEANUP_REC_FIELDS (sizeof(struct bpf_cleanup_info) / sizeof(__u32))
+
+/* Index of each field of struct bpf_cleanup_info, read as an array of __u32. */
+enum {
+ CLEANUP_REC_BEGIN,
+ CLEANUP_REC_END,
+ CLEANUP_REC_PAD,
+};
+
+/*
+ * One (begin, end, landing_pad) triple from .bpf_cleanup, each field resolved
+ * from its relocation to a section and an instruction index within it. Final
+ * indices wait for subprogram placement, which differs per main program.
+ */
+struct cleanup_raw_rec {
+ int sec_idx[CLEANUP_REC_FIELDS];
+ size_t insn_idx[CLEANUP_REC_FIELDS];
+};
+
struct elf_state {
int fd;
const void *obj_buf;
@@ -702,6 +727,8 @@ struct elf_state {
bool has_st_ops;
int arena_data_shndx;
int jumptables_data_shndx;
+ Elf_Data *cleanup_data;
+ int cleanup_shndx;
};
struct usdt_manager;
@@ -779,6 +806,9 @@ struct bpf_object {
void *jumptables_data;
size_t jumptables_data_sz;
+ struct cleanup_raw_rec *cleanup_recs;
+ size_t cleanup_rec_cnt;
+
struct {
struct bpf_program *prog;
unsigned int sym_off;
@@ -848,7 +878,10 @@ static void bpf_program__exit(struct bpf_program *prog)
zfree(&prog->sec_name);
zfree(&prog->insns);
zfree(&prog->reloc_desc);
+ zfree(&prog->cleanup_info);
+ prog->cleanup_info_rec_size = 0;
+ prog->cleanup_info_cnt = 0;
prog->nr_reloc = 0;
prog->insns_cnt = 0;
prog->sec_idx = -1;
@@ -1600,6 +1633,7 @@ static struct bpf_object *bpf_object__new(const char *path,
obj->efile.obj_buf = obj_buf;
obj->efile.obj_buf_sz = obj_buf_sz;
obj->efile.btf_maps_shndx = -1;
+ obj->efile.cleanup_shndx = -1;
obj->kconfig_map_idx = -1;
obj->arena_map_idx = -1;
@@ -4100,6 +4134,9 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
sec_desc->shdr = sh;
sec_desc->data = data;
obj->efile.has_st_ops = true;
+ } else if (strcmp(name, CLEANUP_SEC) == 0) {
+ obj->efile.cleanup_data = data;
+ obj->efile.cleanup_shndx = idx;
} else if (strcmp(name, ARENA_SEC) == 0) {
obj->efile.arena_data = data;
obj->efile.arena_data_shndx = idx;
@@ -4135,6 +4172,7 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
strcmp(name, ".rel" STRUCT_OPS_LINK_SEC) &&
strcmp(name, ".rel?" STRUCT_OPS_SEC) &&
strcmp(name, ".rel?" STRUCT_OPS_LINK_SEC) &&
+ strcmp(name, ".rel" CLEANUP_SEC) &&
strcmp(name, ".rel" MAPS_ELF_SEC)) {
pr_info("elf: skipping relo section(%d) %s for section(%d) %s\n",
idx, name, targ_sec_idx,
@@ -4915,6 +4953,240 @@ static struct bpf_program *find_prog_by_sec_insn(const struct bpf_object *obj,
return NULL;
}
+static int bpf_object__init_cleanup_info(struct bpf_object *obj)
+{
+ Elf_Data *data = obj->efile.cleanup_data;
+ size_t i, nrels, nslots, nrecs;
+ struct cleanup_raw_rec *recs;
+ Elf_Data *relo = NULL;
+ const __u32 *vals;
+ Elf64_Shdr *sh;
+ int ret = 0;
+ bool native;
+
+ if (!data || obj->efile.cleanup_shndx < 0 || !data->d_size)
+ return 0;
+
+ native = is_native_endianness(obj);
+
+ for (i = 0; i < obj->efile.sec_cnt; i++) {
+ struct elf_sec_desc *sd = &obj->efile.secs[i];
+
+ if (sd->sec_type == SEC_RELO && sd->shdr &&
+ sd->shdr->sh_info == (Elf64_Word)obj->efile.cleanup_shndx) {
+ relo = sd->data;
+ break;
+ }
+ }
+ if (!relo) {
+ pr_warn("%s present without relocations\n", CLEANUP_SEC);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ /*
+ * LLVM leaves sh_entsize unset, which is read as the size assumed
+ * here, so this bites only a producer that declares a record size --
+ * which is the one able to say it means something other than three
+ * 4-byte fields.
+ */
+ sh = elf_sec_hdr(obj, elf_sec_by_idx(obj, obj->efile.cleanup_shndx));
+ if (sh && sh->sh_entsize && sh->sh_entsize != sizeof(struct bpf_cleanup_info)) {
+ pr_warn("%s record size %llu is not the expected %zu\n", CLEANUP_SEC,
+ (unsigned long long)sh->sh_entsize, sizeof(struct bpf_cleanup_info));
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ if (data->d_size % sizeof(struct bpf_cleanup_info)) {
+ pr_warn("%s size %zu is not a multiple of the record size %zu\n",
+ CLEANUP_SEC, data->d_size, sizeof(struct bpf_cleanup_info));
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ vals = data->d_buf;
+ nslots = data->d_size / sizeof(__u32);
+ nrecs = data->d_size / sizeof(struct bpf_cleanup_info);
+
+ recs = calloc(nrecs, sizeof(*recs));
+ if (!recs)
+ return -ENOMEM;
+ for (i = 0; i < nslots; i++)
+ recs[i / CLEANUP_REC_FIELDS].sec_idx[i % CLEANUP_REC_FIELDS] = -1;
+
+ /* One relocation per 4-byte field, naming the section it points into. */
+ nrels = relo->d_size / sizeof(Elf64_Rel);
+ for (i = 0; i < nrels; i++) {
+ Elf64_Rel *rel = elf_rel_by_idx(relo, i);
+ Elf64_Sym *sym = elf_sym_by_idx(obj, ELF64_R_SYM(rel->r_info));
+ size_t type = ELF64_R_TYPE(rel->r_info);
+ size_t slot = rel->r_offset / sizeof(__u32);
+ struct cleanup_raw_rec *rec;
+ size_t off;
+
+ if (type != R_BPF_64_NODYLD32 && type != R_BPF_64_ABS32) {
+ pr_warn("%s: relocation %zu has unexpected type %zu\n",
+ CLEANUP_SEC, i, type);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ if (!sym || slot >= nslots || rel->r_offset % sizeof(__u32)) {
+ pr_warn("%s: bad relocation %zu\n", CLEANUP_SEC, i);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ /*
+ * The addend lives in the section data, which libelf leaves in
+ * the object's byte order; a non-section symbol additionally
+ * contributes its own value.
+ */
+ off = (native ? vals[slot] : bswap_32(vals[slot])) + sym->st_value;
+ if (off % BPF_INSN_SZ) {
+ pr_warn("%s: field %zu offset %zu is not instruction aligned\n",
+ CLEANUP_SEC, slot, off);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ rec = &recs[slot / CLEANUP_REC_FIELDS];
+ rec->sec_idx[slot % CLEANUP_REC_FIELDS] = sym->st_shndx;
+ rec->insn_idx[slot % CLEANUP_REC_FIELDS] = off / BPF_INSN_SZ;
+ }
+
+ for (i = 0; i < nslots; i++) {
+ if (recs[i / CLEANUP_REC_FIELDS].sec_idx[i % CLEANUP_REC_FIELDS] < 0) {
+ pr_warn("%s: field %zu has no relocation\n", CLEANUP_SEC, i);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ }
+
+ obj->cleanup_recs = recs;
+ obj->cleanup_rec_cnt = nrecs;
+ return 0;
+out:
+ free(recs);
+ return ret;
+}
+
+static int cmp_cleanup_info(const void *a, const void *b)
+{
+ const struct bpf_cleanup_info *x = a, *y = b;
+
+ if (x->begin_off == y->begin_off)
+ return 0;
+ return x->begin_off < y->begin_off ? -1 : 1;
+}
+
+static int bpf_prog_collect_cleanup_info(struct bpf_object *obj,
+ struct bpf_program *prog)
+{
+ size_t i;
+ int j;
+
+ for (i = 0; i < obj->cleanup_rec_cnt; i++) {
+ struct cleanup_raw_rec *raw = &obj->cleanup_recs[i];
+ struct bpf_program *owner = NULL;
+ struct bpf_cleanup_info ci = {};
+ __u32 fields[CLEANUP_REC_FIELDS];
+ void *tmp;
+
+ if (raw->sec_idx[CLEANUP_REC_BEGIN] == raw->sec_idx[CLEANUP_REC_END] &&
+ raw->insn_idx[CLEANUP_REC_BEGIN] >= raw->insn_idx[CLEANUP_REC_END]) {
+ pr_warn("%s: record %zu is an empty range [%zu,%zu)\n",
+ CLEANUP_SEC, i, raw->insn_idx[CLEANUP_REC_BEGIN],
+ raw->insn_idx[CLEANUP_REC_END]);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ for (j = 0; j < CLEANUP_REC_FIELDS; j++) {
+ size_t idx = raw->insn_idx[j], final;
+ struct bpf_program *p;
+
+ /*
+ * An exclusive end may name the instruction past the
+ * last of a function, so ask about the last one the
+ * range covers, the way the kernel does.
+ */
+ if (j == CLEANUP_REC_END) {
+ if (!idx) {
+ pr_warn("%s: record %zu ends at instruction 0\n",
+ CLEANUP_SEC, i);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ idx--;
+ }
+
+ p = find_prog_by_sec_insn(obj, raw->sec_idx[j], idx);
+ if (!p) {
+ pr_warn("%s: record %zu field %d is not inside a function\n",
+ CLEANUP_SEC, i, j);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ if (!owner) {
+ owner = p;
+ } else if (owner != p) {
+ pr_warn("%s: record %zu spans functions '%s' and '%s'\n",
+ CLEANUP_SEC, i, owner->name, p->name);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ if (owner == prog) {
+ final = raw->insn_idx[j] - prog->sec_insn_off;
+ } else if (prog_is_subprog(obj, owner) && owner->sub_insn_off) {
+ /*
+ * sub_insn_off is where this subprogram was
+ * appended to the main program being relocated;
+ * zero means it is not part of it.
+ */
+ final = owner->sub_insn_off +
+ raw->insn_idx[j] - owner->sec_insn_off;
+ } else {
+ owner = NULL;
+ break;
+ }
+ fields[j] = final;
+ }
+ if (!owner)
+ continue;
+
+ ci.begin_off = fields[CLEANUP_REC_BEGIN];
+ ci.end_off = fields[CLEANUP_REC_END];
+ ci.landing_pad_off = fields[CLEANUP_REC_PAD];
+
+ tmp = libbpf_reallocarray(prog->cleanup_info, prog->cleanup_info_cnt + 1,
+ sizeof(*prog->cleanup_info));
+ if (!tmp)
+ return -ENOMEM;
+ prog->cleanup_info = tmp;
+ prog->cleanup_info_rec_size = sizeof(struct bpf_cleanup_info);
+ prog->cleanup_info[prog->cleanup_info_cnt++] = ci;
+
+ pr_debug("prog '%s': cleanup region [%u,%u) -> landing pad %u\n",
+ prog->name, ci.begin_off, ci.end_off, ci.landing_pad_off);
+ }
+
+ if (!prog->cleanup_info_cnt)
+ return 0;
+
+ if (prog->cleanup_info_cnt > INT32_MAX / sizeof(struct bpf_cleanup_info)) {
+ pr_warn("prog '%s': too many cleanup records: %u\n",
+ prog->name, prog->cleanup_info_cnt);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ qsort(prog->cleanup_info, prog->cleanup_info_cnt,
+ sizeof(*prog->cleanup_info), cmp_cleanup_info);
+ for (i = 1; i < prog->cleanup_info_cnt; i++) {
+ struct bpf_cleanup_info *prev = &prog->cleanup_info[i - 1];
+ struct bpf_cleanup_info *cur = &prog->cleanup_info[i];
+
+ if (cur->begin_off < prev->end_off) {
+ pr_warn("prog '%s': overlapping cleanup regions [%u,%u) and [%u,%u)\n",
+ prog->name, prev->begin_off, prev->end_off,
+ cur->begin_off, cur->end_off);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ }
+
+ return 0;
+}
+
static int
bpf_object__collect_prog_relos(struct bpf_object *obj, Elf64_Shdr *shdr, Elf_Data *data)
{
@@ -7894,6 +8166,13 @@ static int bpf_object__relocate(struct bpf_object *obj, const char *targ_btf_pat
return err;
}
}
+
+ err = bpf_prog_collect_cleanup_info(obj, prog);
+ if (err) {
+ pr_warn("prog '%s': failed to collect cleanup info: %s\n",
+ prog->name, errstr(err));
+ return err;
+ }
}
for (i = 0; i < obj->nr_programs; i++) {
prog = &obj->programs[i];
@@ -8176,6 +8455,9 @@ static int bpf_object__collect_relos(struct bpf_object *obj)
return -LIBBPF_ERRNO__INTERNAL;
}
+ if (idx == obj->efile.cleanup_shndx)
+ continue;
+
if (obj->efile.secs[idx].sec_type == SEC_RODATA)
err = bpf_object__collect_rodata_relos(obj, shdr, data);
else if (obj->efile.secs[idx].sec_type == SEC_ST_OPS)
@@ -8469,6 +8751,11 @@ static int bpf_object_load_prog(struct bpf_object *obj, struct bpf_program *prog
load_attr.line_info_rec_size = prog->line_info_rec_size;
load_attr.line_info_cnt = prog->line_info_cnt;
}
+ if (prog->cleanup_info_cnt) {
+ load_attr.cleanup_info = prog->cleanup_info;
+ load_attr.cleanup_info_cnt = prog->cleanup_info_cnt;
+ load_attr.cleanup_info_rec_size = prog->cleanup_info_rec_size;
+ }
load_attr.log_level = log_level;
load_attr.prog_flags = prog->prog_flags;
load_attr.fd_array = obj->fd_array;
@@ -9062,6 +9349,7 @@ static struct bpf_object *bpf_object_open(const char *path, const void *obj_buf,
err = err ? : bpf_object__init_maps(obj, opts);
err = err ? : bpf_object_init_progs(obj, opts);
err = err ? : bpf_object__collect_relos(obj);
+ err = err ? : bpf_object__init_cleanup_info(obj);
if (err)
goto out;
@@ -10195,6 +10483,9 @@ void bpf_object__close(struct bpf_object *obj)
zfree(&obj->jumptables_data);
obj->jumptables_data_sz = 0;
+ zfree(&obj->cleanup_recs);
+ obj->cleanup_rec_cnt = 0;
+
for (i = 0; i < obj->jumptable_map_cnt; i++)
close(obj->jumptable_maps[i].fd);
zfree(&obj->jumptable_maps);
@@ -10599,6 +10890,25 @@ int bpf_program__clone(struct bpf_program *prog, const struct bpf_prog_load_opts
attr.line_info_rec_size = info ? info_rec_size : prog->line_info_rec_size;
}
+ /* exception cleanup table */
+ info = OPTS_GET(opts, cleanup_info, NULL);
+ info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
+ info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
+ if (!!info != !!info_cnt || !!info != !!info_rec_size) {
+ pr_warn("prog '%s': cleanup_info, cleanup_info_cnt, and cleanup_info_rec_size must all be specified or all omitted\n",
+ prog->name);
+ return libbpf_err(-EINVAL);
+ }
+ if (info) {
+ attr.cleanup_info = info;
+ attr.cleanup_info_cnt = info_cnt;
+ attr.cleanup_info_rec_size = info_rec_size;
+ } else if (prog->cleanup_info_cnt) {
+ attr.cleanup_info = prog->cleanup_info;
+ attr.cleanup_info_cnt = prog->cleanup_info_cnt;
+ attr.cleanup_info_rec_size = prog->cleanup_info_rec_size;
+ }
+
/* Logging is caller-controlled; no fallback to prog/obj log settings */
attr.log_buf = OPTS_GET(opts, log_buf, NULL);
attr.log_size = OPTS_GET(opts, log_size, 0);
diff --git a/tools/lib/bpf/libbpf_internal.h b/tools/lib/bpf/libbpf_internal.h
index 546f65b95cf4..f1630f03d5f5 100644
--- a/tools/lib/bpf/libbpf_internal.h
+++ b/tools/lib/bpf/libbpf_internal.h
@@ -56,6 +56,9 @@
#ifndef R_BPF_64_ABS32
#define R_BPF_64_ABS32 3
#endif
+#ifndef R_BPF_64_NODYLD32
+#define R_BPF_64_NODYLD32 4
+#endif
#ifndef R_BPF_64_32
#define R_BPF_64_32 10
#endif
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 17/22] libbpf: Carry the exception cleanup table through the light skeleton
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (15 preceding siblings ...)
2026-09-29 0:17 ` [PATCH bpf-next v7 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
@ 2026-09-29 0:17 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
` (4 subsequent siblings)
21 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:17 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A light skeleton does not call bpf_prog_load(). bpf_gen__prog_load() builds
its own union bpf_attr field by field, and the loader program it emits is
what issues BPF_PROG_LOAD when the skeleton runs -- so a program loaded
this way reached the kernel without the table the previous patch collected
for it, and the verifier refused it with "unreachable insn", which names
neither the skeleton nor the table.
Carry it the way func_info and line_info are carried: the records go into
the loader's blob of bytes, the count and record size into the attr, and a
relocation stores the blob's address into attr.cleanup_info once that
address is known. The attr grows to its new last field, cleanup_info_cnt.
Records are 4-byte fields like the other info blobs, so a cross-endian
build has to swap them too.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/gen_loader.c | 29 +++++++++++++++++++++++++----
tools/lib/bpf/libbpf_internal.h | 7 +++++++
2 files changed, 32 insertions(+), 4 deletions(-)
diff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c
index 251392aa8b41..d10cd0672475 100644
--- a/tools/lib/bpf/gen_loader.c
+++ b/tools/lib/bpf/gen_loader.c
@@ -994,13 +994,15 @@ static void cleanup_relos(struct bpf_gen *gen, int insns)
cleanup_core_relo(gen);
}
-/* Convert func, line, and core relo info blobs to target endianness */
+/* Convert func, line, core relo and cleanup info blobs to target endianness */
static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
- int core_relos, struct bpf_prog_load_opts *load_attr)
+ int core_relos, int cleanup_info,
+ struct bpf_prog_load_opts *load_attr)
{
struct bpf_func_info *fi = gen->data_start + func_info;
struct bpf_line_info *li = gen->data_start + line_info;
struct bpf_core_relo *cr = gen->data_start + core_relos;
+ struct bpf_cleanup_info *ci = gen->data_start + cleanup_info;
int i;
for (i = 0; i < load_attr->func_info_cnt; i++)
@@ -1011,6 +1013,9 @@ static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
for (i = 0; i < gen->core_relo_cnt; i++)
bpf_core_relo_bswap(cr++);
+
+ for (i = 0; i < load_attr->cleanup_info_cnt; i++)
+ bpf_cleanup_info_bswap(ci++);
}
void bpf_gen__prog_load(struct bpf_gen *gen,
@@ -1024,8 +1029,11 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
load_attr->line_info_rec_size;
int core_relo_tot_sz = gen->core_relo_cnt *
sizeof(struct bpf_core_relo);
+ int cleanup_info_tot_sz = load_attr->cleanup_info_cnt *
+ load_attr->cleanup_info_rec_size;
int prog_load_attr, license_off, insns_off, func_info, line_info, core_relos;
- int attr_size = offsetofend(union bpf_attr, core_relo_rec_size);
+ int attr_size = offsetofend(union bpf_attr, cleanup_info_cnt);
+ int cleanup_info;
union bpf_attr attr;
memset(&attr, 0, attr_size);
@@ -1074,9 +1082,17 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
core_relos, gen->core_relo_cnt,
sizeof(struct bpf_core_relo));
+ attr.cleanup_info_rec_size = tgt_endian(load_attr->cleanup_info_rec_size);
+ attr.cleanup_info_cnt = tgt_endian(load_attr->cleanup_info_cnt);
+ cleanup_info = add_data(gen, load_attr->cleanup_info, cleanup_info_tot_sz);
+ pr_debug("gen: prog_load: cleanup_info: off %d cnt %u rec size %u\n",
+ cleanup_info, load_attr->cleanup_info_cnt,
+ load_attr->cleanup_info_rec_size);
+
/* convert all info blobs to target endianness */
if (gen->swapped_endian && !gen->error)
- info_blob_bswap(gen, func_info, line_info, core_relos, load_attr);
+ info_blob_bswap(gen, func_info, line_info, core_relos, cleanup_info,
+ load_attr);
libbpf_strlcpy(attr.prog_name, prog_name, sizeof(attr.prog_name));
prog_load_attr = add_data(gen, &attr, attr_size);
@@ -1098,6 +1114,11 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
/* populate union bpf_attr with a pointer to core_relos */
emit_rel_store(gen, attr_field(prog_load_attr, core_relos), core_relos);
+ /* with no records there is no blob of them to point the attr at */
+ if (load_attr->cleanup_info_cnt)
+ emit_rel_store(gen, attr_field(prog_load_attr, cleanup_info),
+ cleanup_info);
+
/* populate union bpf_attr fd_array with a pointer to data where map_fds are saved */
emit_rel_store(gen, attr_field(prog_load_attr, fd_array), gen->fd_array);
diff --git a/tools/lib/bpf/libbpf_internal.h b/tools/lib/bpf/libbpf_internal.h
index f1630f03d5f5..9d341839ca74 100644
--- a/tools/lib/bpf/libbpf_internal.h
+++ b/tools/lib/bpf/libbpf_internal.h
@@ -572,6 +572,13 @@ static inline void bpf_core_relo_bswap(struct bpf_core_relo *i)
i->kind = bswap_32(i->kind);
}
+static inline void bpf_cleanup_info_bswap(struct bpf_cleanup_info *i)
+{
+ i->begin_off = bswap_32(i->begin_off);
+ i->end_off = bswap_32(i->end_off);
+ i->landing_pad_off = bswap_32(i->landing_pad_off);
+}
+
enum btf_field_iter_kind {
BTF_FIELD_ITER_IDS,
BTF_FIELD_ITER_STRS,
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (16 preceding siblings ...)
2026-09-29 0:17 ` [PATCH bpf-next v7 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
@ 2026-09-29 0:17 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
` (3 subsequent siblings)
21 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:17 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
An object that carries a compiler-emitted exception cleanup table cannot be
linked today. The table's fields are byte offsets into a code section,
materialised by a 32-bit relocation against that section's symbol with the
offset itself as the implicit addend, and the linker rejects both halves of
that: the relocation type is not in the list it accepts, and a relocation
against an STT_SECTION symbol from a non-executable section is an outright
error.
Both spellings of that relocation have to be taken. LLVM emits
R_BPF_64_NODYLD32 for a .long against a section symbol; GNU as emits
R_BPF_64_ABS32, which is what bpf_reloc_type_lookup() maps BFD_RELOC_32 to.
They describe the same value, and the selftests are built with both
compilers.
Keying on the relocation type rather than the section name means any
non-executable section can reach the new arm, where an unconditional
refusal used to stand. That refusal was covering two things the arm now has
to do itself: a target may be SHT_NOBITS, which extend_sec() leaves with no
raw_data, and r_offset is alignment-checked only where the section holds
instructions. The arm rejects both, and bounds the offset against the
section size before writing through it.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/linker.c | 36 +++++++++++++++++++++++++++++++++++-
1 file changed, 35 insertions(+), 1 deletion(-)
diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
index f3f71c452f00..7c4bc232d576 100644
--- a/tools/lib/bpf/linker.c
+++ b/tools/lib/bpf/linker.c
@@ -1036,7 +1036,8 @@ static int linker_sanity_check_elf_relos(struct src_obj *obj, struct src_sec *se
size_t sym_type = ELF64_R_TYPE(relo->r_info);
if (sym_type != R_BPF_64_64 && sym_type != R_BPF_64_32 &&
- sym_type != R_BPF_64_ABS64 && sym_type != R_BPF_64_ABS32) {
+ sym_type != R_BPF_64_ABS64 && sym_type != R_BPF_64_ABS32 &&
+ sym_type != R_BPF_64_NODYLD32) {
pr_warn("ELF relo #%d in section #%zu has unexpected type %zu in %s\n",
i, sec->sec_idx, sym_type, obj->filename);
return -EINVAL;
@@ -2263,6 +2264,7 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
if (ELF64_ST_TYPE(src_sym->st_info) == STT_SECTION) {
struct src_sec *sec = &obj->secs[src_sym->st_shndx];
struct bpf_insn *insn;
+ __u32 *val;
if (src_linked_sec->shdr->sh_flags & SHF_EXECINSTR) {
/* calls to the very first static function inside
@@ -2297,6 +2299,38 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
if (linker->swapped_endian)
off = bswap_64(off);
memcpy(ptr, &off, sizeof(off));
+ } else if (sym_type == R_BPF_64_NODYLD32 ||
+ sym_type == R_BPF_64_ABS32) {
+ /*
+ * A byte offset into a code section,
+ * stored in place. LLVM spells this
+ * relocation NODYLD32 and GNU as
+ * spells it ABS32; being bytes, the
+ * section's new start goes in as it
+ * is, not scaled the way a call's
+ * instruction index is above.
+ *
+ * r_offset is checked only for an
+ * executable section, and SHT_NOBITS
+ * has no raw_data, so bound it here --
+ * subtracting, so it cannot wrap.
+ */
+ if (!dst_linked_sec->raw_data ||
+ dst_linked_sec->sec_sz < (int)sizeof(*val) ||
+ dst_rel->r_offset % sizeof(*val) ||
+ dst_rel->r_offset >
+ (size_t)dst_linked_sec->sec_sz - sizeof(*val)) {
+ pr_warn("ELF relo #%d in section #%zu points outside the data of section '%s' in %s\n",
+ j, src_sec->sec_idx,
+ dst_linked_sec->sec_name,
+ obj->filename);
+ return -EINVAL;
+ }
+ val = dst_linked_sec->raw_data + dst_rel->r_offset;
+ if (linker->swapped_endian)
+ *val = bswap_32(bswap_32(*val) + sec->dst_off);
+ else
+ *val += sec->dst_off;
} else {
pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
return -EINVAL;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (17 preceding siblings ...)
2026-09-29 0:17 ` [PATCH bpf-next v7 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
@ 2026-09-29 0:17 ` Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-29 0:17 ` [PATCH bpf-next v7 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
` (2 subsequent siblings)
21 siblings, 1 reply; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:17 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
C has no unwinding, so nothing here comes out of the frontend: the frames
that own a resource are written as __naked inline assembly, which spells
out by hand exactly what a frontend emits -- a call site bracketed by two
labels, a landing pad unreachable in the compiler's CFG, and a .bpf_cleanup
record tying them together. The assembler turns ".long <text label>" into
the same R_BPF_64_NODYLD32 relocation the BPF AsmPrinter emits, so libbpf
and the kernel see an object indistinguishable from a compiler-generated
one.
Call chain: entry -> foo1 -> foo1v -> foo2 -> foo3. foo3 holds a
non-preemptible section and unwinds inside it; foo2 holds an RCU read lock
and has two call sites sharing one pad, one of them its own unwind; foo1v
is a void frame whose pad ends in a jump to a resume block placed after an
unrelated block that ends in a plain exit; foo1 owns nothing and gets no
record; entry is the boundary.
There are also the shapes the kernel refuses:
- a catch pad, and a pad ambiguous between catch and cleanup
- a table alongside bpf_throw() or a tagged exception callback
- a subprogram that can unwind, used as a callback
- a tail call, a BPF_LD_[ABS|IND] or a gotox in a pad
- a second unwind while one is in flight, raised in the pad or below it
- a resume outside a pad, and one in a subprogram the pad called
- a kfunc stack argument in a pad
- a jump into a pad from outside it
- a pad inside another record's call-site range, and a record covering no
call that can unwind and whose pad nothing reaches
- a frame leaving through an unwind holding what it did not hold when it
was entered: a pad dropping a lock its frame never took, one forgetting
the lock it did take, a subprogram with no pad of its own leaving while
it holds one, and a pad dropping a reference the frame never reserved
- a pad-less unwind inside an RCU read-side region
- a frame holding a lock or a reference across a call an unwind passes
through with no record over it, in a subprogram's frame and in the main
program's
The test skips rather than fails where the JIT cannot dispatch a landing
pad at all.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
.../selftests/bpf/exceptions_cleanup.h | 27 +
.../bpf/prog_tests/exceptions_cleanup.c | 85 ++
.../selftests/bpf/progs/exceptions_cleanup.c | 162 ++++
.../bpf/progs/exceptions_cleanup_fail.c | 878 ++++++++++++++++++
4 files changed, 1152 insertions(+)
create mode 100644 tools/testing/selftests/bpf/exceptions_cleanup.h
create mode 100644 tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
new file mode 100644
index 000000000000..d1d40314035e
--- /dev/null
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -0,0 +1,27 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef __EXCEPTIONS_CLEANUP_H__
+#define __EXCEPTIONS_CLEANUP_H__
+
+/* progs/exceptions_cleanup.c: one bit per frame that reports it ran. */
+#define RAN_FOO3_PREEMPT 0x1
+#define RAN_FOO2_RCU 0x2
+#define RAN_FOO1V_PREEMPT 0x4
+#define RAN_FOO2_DROP 0x8
+#define RAN_BUMP 0x10
+
+#define CLEANUP_REC(begin, end, landing_pad) \
+ ".pushsection .bpf_cleanup,\"a\",@progbits;" \
+ ".long " begin ";" \
+ ".long " end ";" \
+ ".long " landing_pad ";" \
+ ".popsection;"
+
+/* Set a bit in @pads_ran. */
+#define PAD_RAN(bit) \
+ "r1 = %[pads_ran] ll;" \
+ "r2 = *(u64 *)(r1 + 0);" \
+ "r2 |= " bit ";" \
+ "*(u64 *)(r1 + 0) = r2;"
+
+#endif /* __EXCEPTIONS_CLEANUP_H__ */
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
new file mode 100644
index 000000000000..255f88d35aad
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -0,0 +1,85 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <test_progs.h>
+#include "exceptions_cleanup.h"
+#include "exceptions_cleanup.skel.h"
+#include "exceptions_cleanup_fail.skel.h"
+
+/* foo3 unwound: every frame that has a pad ran it. */
+#define PADS_FOO3_UNWOUND \
+ (RAN_FOO3_PREEMPT | RAN_FOO2_RCU | RAN_FOO1V_PREEMPT | RAN_FOO2_DROP)
+
+/* foo2 unwound after foo3 returned normally: foo3's pad must not run. */
+#define PADS_FOO2_UNWOUND \
+ (RAN_FOO2_RCU | RAN_FOO1V_PREEMPT | RAN_FOO2_DROP)
+
+static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
+ __u64 pads)
+{
+ __u64 ctx = 0;
+ int err;
+
+ LIBBPF_OPTS(bpf_test_run_opts, topts,
+ .ctx_in = &ctx,
+ .ctx_size_in = sizeof(ctx),
+ );
+
+ skel->bss->input = input;
+ skel->bss->pads_ran = 0;
+ skel->bss->result = 0;
+
+ err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.entry), &topts);
+ if (!ASSERT_OK(err, "run"))
+ return;
+ ASSERT_EQ(topts.retval, retval, "retval");
+ /* bump() is not a landing pad; it sets its bit on every run. */
+ ASSERT_EQ(skel->bss->pads_ran, pads | RAN_BUMP, "pads_ran");
+}
+
+void test_exceptions_cleanup(void)
+{
+ char log[8192] = {};
+
+ LIBBPF_OPTS(bpf_object_open_opts, opts,
+ .kernel_log_buf = log,
+ .kernel_log_size = sizeof(log));
+ struct exceptions_cleanup *skel;
+ int err;
+
+ skel = exceptions_cleanup__open_opts(&opts);
+ if (!ASSERT_OK_PTR(skel, "open"))
+ return;
+
+ err = exceptions_cleanup__load(skel);
+ if (err) {
+ if (err == -EOPNOTSUPP &&
+ strstr(log, "exception cleanup needs a JIT that can dispatch landing pads")) {
+ printf("%s:SKIP:JIT cannot dispatch exception cleanup landing pads\n",
+ __func__);
+ test__skip();
+ } else if (!ASSERT_OK(err, "load")) {
+ fprintf(stderr, "%s", log);
+ }
+ exceptions_cleanup__destroy(skel);
+ return;
+ }
+
+ /* No unwind: foo3 returns 1 ^ 1 == 0, foo2 adds one, no pad runs. */
+ if (test__start_subtest("no_unwind"))
+ run(skel, 1, 1, 0);
+
+ /* foo3 unwinds; every pad runs and entry returns zero. */
+ if (test__start_subtest("unwind_from_foo3"))
+ run(skel, 101, 0, PADS_FOO3_UNWOUND);
+
+ /*
+ * foo3 returns 2 ^ 1 == 3, so foo2 unwinds from its own second region;
+ * foo3's frame is long gone, so its pad must not run.
+ */
+ if (test__start_subtest("unwind_from_foo2"))
+ run(skel, 2, 0, PADS_FOO2_UNWOUND);
+
+ exceptions_cleanup__destroy(skel);
+
+ RUN_TESTS(exceptions_cleanup_fail);
+}
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
new file mode 100644
index 000000000000..2b22a05a0ded
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
@@ -0,0 +1,162 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_unwind();
+ bpf_rcu_read_lock();
+ bpf_rcu_read_unlock();
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume(NULL);
+}
+
+__u64 input = 0;
+__u64 pads_ran = 0;
+__u64 result = 0;
+
+static __used __noinline __u64 foo3(__u64 x)
+{
+ bpf_preempt_disable();
+ if (x > 100)
+ asm volatile (
+ "1:" "call bpf_unwind;" /* cleanup region */
+ "2:"
+ "goto 3f;"
+ "4:" /* landing pad */
+ /*
+ * r0 at pad entry is whatever the walker leaves
+ * there, carried across the call in a callee-saved
+ * register and handed to the resume, the way a
+ * compiler-emitted pad passes the exception pointer to
+ * _Unwind_Resume. The kfunc takes it and ignores it, and
+ * the two pads below do without the shuffle.
+ */
+ "r7 = r0;"
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "3:"
+ CLEANUP_REC("1b", "2b", "4b")
+ :
+ : [ran]"i"(RAN_FOO3_PREEMPT),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+ bpf_preempt_enable();
+ return x ^ 1;
+}
+
+__u64 never = 0;
+
+static __used __naked __noinline void drop_glue(void)
+{
+ asm volatile (
+ PAD_RAN("%[ran]")
+ "exit;"
+ :
+ : [ran]"i"(RAN_FOO2_DROP), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 foo2(void)
+{
+ asm volatile (
+ "r6 = r1;"
+ "call bpf_rcu_read_lock;"
+ "r1 = r6;"
+"1:" "call foo3;" /* cleanup region #1 */
+"2:"
+ "r6 = r0;"
+ "if r6 == 0 goto 5f;"
+"3:" "call bpf_unwind;" /* cleanup region #2 */
+"4:"
+ "r0 = 0;"
+ "exit;"
+"5:"
+ "call bpf_rcu_read_unlock;"
+ "r0 = r6;"
+ "r0 += 1;"
+ "exit;"
+"6:" /* landing pad, shared by both regions */
+ "call drop_glue;"
+ "call bpf_rcu_read_unlock;"
+ PAD_RAN("%[ran_rcu]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "6b")
+ CLEANUP_REC("3b", "4b", "6b")
+ :
+ : [ran_rcu]"i"(RAN_FOO2_RCU),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+static __used __naked __noinline void foo1v(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call foo2;" /* cleanup region */
+"2:"
+ "r6 = r0;"
+ "call bpf_preempt_enable;"
+ "r1 = %[result] ll;"
+ "*(u64 *)(r1 + 0) = r6;"
+ "goto 7f;"
+"8:" /* landing pad */
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "goto 9f;"
+"7:" /* the frame's own exit block */
+ "r0 = 0;"
+ "exit;"
+"9:" /* shared resume block */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "8b")
+ :
+ : [ran]"i"(RAN_FOO1V_PREEMPT), __imm_addr(input),
+ __imm_addr(result), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/*
+ * A frame with no cleanup record: an unwind leaving it runs no pad. The
+ * unwind never fires -- @never is global -- and the bit marks the return path.
+ */
+static __used __naked __noinline void bump(void)
+{
+ asm volatile (
+ PAD_RAN("%[ran]")
+ "r1 = %[never] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 == 0 goto 1f;"
+ "r1 = 0;"
+ "call bpf_unwind;"
+"1:"
+ "exit;" /* r0 deliberately left alone */
+ :
+ : [ran]"i"(RAN_BUMP), __imm_addr(never), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+__noinline __u64 foo1(void)
+{
+ bump();
+ foo1v();
+ return result;
+}
+
+SEC("syscall")
+int entry(void *ctx)
+{
+ return foo1();
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
new file mode 100644
index 000000000000..76cca8dc45a9
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
@@ -0,0 +1,878 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_experimental.h"
+#include "bpf_misc.h"
+#include "../test_kmods/bpf_testmod_kfunc.h"
+#include "exceptions_cleanup.h"
+
+__u64 input = 0;
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_unwind();
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_rcu_read_lock();
+ bpf_rcu_read_unlock();
+ bpf_unwind_resume(NULL);
+}
+
+/* An unwind raised in a callee, which is how a cleanup region gets one. */
+static __used __naked __noinline __u64 inner_unwind(void)
+{
+ asm volatile (
+ "r1 = 1;"
+ "call bpf_unwind;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static int unwinding_cb(__u32 idx, void *ctx)
+{
+ bpf_unwind();
+ return 0;
+}
+
+static __used __naked __noinline __u64 cb_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call unwinding_cb;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("may unwind and is used as a callback")
+int callback_may_unwind(void *ctx)
+{
+ bpf_loop(1, unwinding_cb, NULL, 0);
+ return cb_frame();
+}
+
+/* A pad that reaches both a resume and a plain exit. */
+static __used __naked __noinline __u64 ambiguous_pad_frame(void)
+{
+ asm volatile (
+ "r6 = r1;"
+ "call bpf_preempt_disable;"
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad: two ways out */
+ "call bpf_preempt_enable;"
+ "if r6 > 10 goto 4f;"
+ "call bpf_unwind_resume;"
+ "exit;"
+"4:"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("a catch pad is not supported yet")
+int ambiguous_landing_pad(void *ctx)
+{
+ return ambiguous_pad_frame();
+}
+
+/* A second bpf_unwind() from inside a landing pad. */
+static __used __naked __noinline __u64 unwind_in_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad that unwinds again */
+ "call bpf_preempt_enable;"
+ "r1 = 2;"
+ "call bpf_unwind;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("starts a second unwind while one is in flight")
+int unwind_from_landing_pad(void *ctx)
+{
+ return unwind_in_pad_frame();
+}
+
+__noinline int unused_exc_cb(u64 cookie)
+{
+ return 0;
+}
+
+static __used __naked __noinline __u64 cb_and_table_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 9;"
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__exception_cb(unused_exc_cb)
+__failure __msg("cannot be combined with an exception callback")
+int table_with_exception_cb(void *ctx)
+{
+ return cb_and_table_frame();
+}
+
+/*
+ * A throw and a table, with no callback tagged: the default callback is
+ * appended too late to stand in for the throw, so the program is scanned
+ * for one instead.
+ */
+static __used __naked __noinline __u64 throw_and_table_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("cannot be combined with bpf_throw")
+int table_with_throw(void *ctx)
+{
+ if (input)
+ bpf_throw(0);
+ return throw_and_table_frame();
+}
+
+__u64 never;
+
+/*
+ * A pad calling a global subprogram that can unwind. The subprogram is
+ * verified on its own, so the pad rule is what refuses it.
+ */
+__noinline void pad_callee_that_unwinds(void)
+{
+ if (never)
+ bpf_unwind();
+}
+
+static __used __naked __noinline __u64 pad_calls_unwinder_frame(void)
+{
+ asm volatile (
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call pad_callee_that_unwinds;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("which can unwind while an unwind is in flight")
+int pad_calls_unwinder(void *ctx)
+{
+ return pad_calls_unwinder_frame();
+}
+
+/*
+ * A pad calling a global subprogram that can throw. The throw is refused
+ * wherever it sits: the scan covers the subprograms too, not just the main
+ * program, so this never reaches the rules about pads.
+ */
+__noinline void pad_callee_that_throws(void)
+{
+ if (never)
+ bpf_throw(0);
+}
+
+static __used __naked __noinline __u64 pad_calls_thrower_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 11;"
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call pad_callee_that_throws;" /* ...which can throw: refused */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("cannot be combined with bpf_throw")
+int pad_calls_thrower(void *ctx)
+{
+ return pad_calls_thrower_frame();
+}
+
+static __used __naked __noinline __u64 catch_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 12;"
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* catch pad: no resume, it stops here */
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not supported yet, only cleanup pads that resume")
+int catch_landing_pad(void *ctx)
+{
+ return catch_pad_frame();
+}
+
+/* A bpf_unwind_resume() outside any landing pad. */
+static __used __naked __noinline __u64 stray_resume_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 13;"
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not in a landing pad")
+int resume_outside_pad(void *ctx)
+{
+ /* Never taken, but reachable, which is all the verifier needs. */
+ if (never)
+ bpf_unwind_resume(NULL);
+ return stray_resume_frame();
+}
+
+/* A bpf_unwind_resume() in a subprogram a landing pad calls. */
+static __used __naked __noinline void resume_in_callee(void)
+{
+ asm volatile (
+ "call bpf_unwind_resume;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 pad_calls_resumer_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 14;"
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call resume_in_callee;" /* ...which resumes: refused */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not in a landing pad")
+int resume_in_pad_callee(void *ctx)
+{
+ return pad_calls_resumer_frame();
+}
+
+static __used __naked __noinline __u64 nested_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_unwind;" /* first cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* first pad, second region's call */
+ "call bpf_preempt_enable;"
+"4:"
+ "call bpf_unwind_resume;"
+ "exit;"
+"5:" /* second pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ CLEANUP_REC("3b", "4b", "5b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is inside the call-site range of")
+int nested_landing_pad(void *ctx)
+{
+ return nested_pad_frame();
+}
+
+/* A tail call in a landing pad: the frame would never reach its resume. */
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(__u32));
+ __uint(value_size, sizeof(__u32));
+} tc_map SEC(".maps");
+
+static __used __naked __noinline __u64 tail_call_pad_frame(void)
+{
+ asm volatile (
+ "r6 = r1;"
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r1 = r6;"
+ "r2 = %[tc_map] ll;"
+ "r3 = 0;"
+ "call %[bpf_tail_call];"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm(bpf_tail_call), __imm_addr(tc_map)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("and is in a landing pad")
+int tail_call_in_pad(void *ctx)
+{
+ return tail_call_pad_frame();
+}
+
+#if defined(__BPF_FEATURE_STACK_ARGUMENT)
+
+/*
+ * A kfunc by-value argument that runs past the argument registers, in a
+ * landing pad. The pad is not what refuses it. The C call gives the extern
+ * its BTF.
+ */
+static __used __noinline void __nofit_btf_anchor(void)
+{
+ struct prog_test_pair_arg s = {};
+
+ bpf_kfunc_call_test_pair_arg_nofit(1, 2, 3, 4, s);
+}
+
+static __used __naked __noinline __u64 kfunc_arg_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_kfunc_call_test_pair_arg_nofit;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("stack arg1 is not initialized")
+int kfunc_stack_arg_in_pad(void *ctx)
+{
+ return kfunc_arg_pad_frame();
+}
+
+#endif /* __BPF_FEATURE_STACK_ARGUMENT */
+
+/* A landing pad entered by ordinary control flow, with no unwind in flight. */
+static __used __naked __noinline __u64 jump_into_pad_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "if r6 > 7 goto 4f;" /* an ordinary branch into the pad */
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+"4:" /* ... and its second instruction */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("runs both inside and outside a landing pad")
+int jump_into_pad(void *ctx)
+{
+ return jump_into_pad_frame();
+}
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)
+
+/*
+ * An indirect jump in a landing pad. A jump table entry is an offset from
+ * the program's section symbol, which has to be spelled in quotes here.
+ */
+static __used __naked __noinline void gotox_unwinder(void)
+{
+ asm volatile (
+ "r1 = 15;"
+ "call bpf_unwind;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("and is in a landing pad")
+__naked void gotox_in_pad(void)
+{
+ asm volatile (
+ ".pushsection .jumptables,\"\",@progbits;"
+"jt0_%=:"
+ ".quad l0_%= - \"?syscall\";"
+ ".quad l1_%= - \"?syscall\";"
+ ".size jt0_%=, 16;"
+ ".global jt0_%=;"
+ ".popsection;"
+
+"1:" "call gotox_unwinder;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r1 = jt0_%= ll;"
+ "r1 += 8;"
+ "r2 = *(u64 *)(r1 + 0);"
+ /*
+ * gotox r2, as raw bytes: the mnemonic only reached the LLVM
+ * assembler in llvm 22, and BPF_RAW_INSN() needs <linux/bpf.h>, which
+ * vmlinux.h rules out. dst_reg is the other nibble on a big-endian
+ * target.
+ */
+#if __BYTE_ORDER__ == __ORDER_BIG_ENDIAN__
+ ".byte 0x0d, 0x20, 0, 0, 0, 0, 0, 0;"
+#else
+ ".byte 0x0d, 0x02, 0, 0, 0, 0, 0, 0;"
+#endif
+"l0_%=:"
+ "call bpf_unwind_resume;"
+ "exit;"
+"l1_%=:"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+#endif /* x86 || arm64 */
+
+/* A BPF_LD_[ABS|IND] in a pad: a failed load leaves without resuming. */
+static __used __naked __noinline __u64 ld_abs_pad_frame(void)
+{
+ asm volatile (
+ "r6 = r1;" /* the skb BPF_LD_ABS reads */
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r0 = *(u32 *)skb[0];"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?tc")
+__failure __msg("and is in a landing pad")
+__naked void ld_abs_in_pad(void)
+{
+ asm volatile (
+ "call ld_abs_pad_frame;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/*
+ * A subprogram a landing pad calls, which unwinds on its own. The second
+ * unwind would rewrite the frames the first is walking.
+ */
+static __used __naked __noinline __u64 own_pad_callee(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* its landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 pad_calls_own_pad_frame(void)
+{
+ asm volatile (
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad, which calls the above */
+ "call own_pad_callee;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("starts a second unwind while one is in flight")
+int unwind_in_pad_callee(void *ctx)
+{
+ return pad_calls_own_pad_frame();
+}
+
+/* A record whose range holds no call that can unwind. */
+static __used __naked __noinline __u64 nounwind_rec_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call bpf_preempt_enable;" /* cleanup region: nounwind */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad, reached by nothing */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("unreachable insn")
+int nounwind_region(void *ctx)
+{
+ return nounwind_rec_frame();
+}
+
+/*
+ * An unwind with no landing pad leaves the frame with nothing run on the way
+ * out, so what the frame holds is checked as it would be at a plain exit.
+ */
+SEC("?syscall")
+__failure __msg("an unwind with no landing pad cannot be used inside bpf_rcu_read_lock-ed region")
+int unwind_no_pad_rcu(void *ctx)
+{
+ bpf_rcu_read_lock();
+ bpf_unwind();
+ bpf_rcu_read_unlock();
+ return 0;
+}
+
+/*
+ * A frame leaves through an unwind holding what it did not hold when it was
+ * entered. The frames above it resume at pads whose state was taken at their
+ * call, so they would be wrong about it.
+ */
+static __used __naked __noinline __u64 pad_drops_caller_lock_frame(void)
+{
+ asm volatile (
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad: drops a lock it never took */
+ "call bpf_rcu_read_unlock;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 caller_holds_lock_frame(void)
+{
+ asm volatile (
+ "call bpf_rcu_read_lock;"
+"1:" "call pad_drops_caller_lock_frame;"
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad */
+ "call bpf_rcu_read_unlock;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("a resume does not leave the frame's bpf_rcu_read_lock state as it found it")
+int pad_drops_caller_lock(void *ctx)
+{
+ return caller_holds_lock_frame();
+}
+
+/* The other way round: a pad that does not drop what its own frame took. */
+static __used __naked __noinline __u64 pad_keeps_own_lock_frame(void)
+{
+ asm volatile (
+ "call bpf_rcu_read_lock;"
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "call bpf_rcu_read_unlock;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad: forgets the unlock */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("a resume does not leave the frame's bpf_rcu_read_lock state as it found it")
+int pad_keeps_own_lock(void *ctx)
+{
+ return pad_keeps_own_lock_frame();
+}
+
+/* And a subprog with no pad at all, leaving through an unwind holding one. */
+static __used __naked __noinline __u64 no_pad_keeps_own_lock_frame(void)
+{
+ asm volatile (
+ "call bpf_rcu_read_lock;"
+ "call bpf_unwind;" /* no record covers it */
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure
+__msg("no landing pad does not leave the frame's bpf_rcu_read_lock state")
+int no_pad_keeps_own_lock(void *ctx)
+{
+ return no_pad_keeps_own_lock_frame();
+}
+
+struct {
+ __uint(type, BPF_MAP_TYPE_RINGBUF);
+ __uint(max_entries, 4096);
+} unwind_ringbuf SEC(".maps");
+
+/*
+ * Always unwinds, so its caller is never returned to on the modelled path --
+ * which is what keeps the release below out of the caller's post-call code.
+ * Its pad discards the record the caller reserved.
+ */
+static __used __naked __noinline __u64 pad_drops_caller_ref_frame(void)
+{
+ asm volatile (
+ "r6 = r1;" /* the caller's reserved record */
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad: drops what it never acquired */
+ "r1 = r6;"
+ "r2 = 0;"
+ "call %[bpf_ringbuf_discard];"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm(bpf_ringbuf_discard)
+ : __clobber_all);
+}
+
+/* And this frame's own pad drops it a second time. */
+static __used __naked __noinline __u64 caller_holds_ref_frame(void)
+{
+ asm volatile (
+ "r1 = %[unwind_ringbuf] ll;"
+ "r2 = 8;"
+ "r3 = 0;"
+ "call %[bpf_ringbuf_reserve];"
+ "if r0 == 0 goto 9f;"
+ "r6 = r0;"
+ "r1 = r6;"
+"1:" "call pad_drops_caller_ref_frame;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad */
+ "r1 = r6;"
+ "r2 = 0;"
+ "call %[bpf_ringbuf_discard];"
+ "call bpf_unwind_resume;"
+ "exit;"
+"9:"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm(bpf_ringbuf_reserve), __imm(bpf_ringbuf_discard),
+ __imm_addr(unwind_ringbuf)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("a resume does not leave the frame's references as it found it")
+int pad_drops_caller_ref(void *ctx)
+{
+ return caller_holds_ref_frame();
+}
+
+/*
+ * A frame an unwind returns through without a pad is abandoned where it made
+ * the call: the JIT sends it to its epilogue, so nothing of it runs again and
+ * whatever it acquired is never released. It has to hold what it entered with
+ * at every such call.
+ */
+static __used __naked __noinline __u64 pad_resumes_frame(void)
+{
+ asm volatile (
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 uncovered_holds_lock_frame(void)
+{
+ asm volatile (
+ "call bpf_rcu_read_lock;"
+ "call pad_resumes_frame;" /* no record covers this call */
+ "call bpf_rcu_read_unlock;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure
+__msg("through this call does not leave the frame's bpf_rcu_read_lock state")
+int unwind_through_call_keeps_lock(void *ctx)
+{
+ return uncovered_holds_lock_frame();
+}
+
+static __used __naked __noinline __u64 uncovered_holds_ref_frame(void)
+{
+ asm volatile (
+ "r1 = %[unwind_ringbuf] ll;"
+ "r2 = 8;"
+ "r3 = 0;"
+ "call %[bpf_ringbuf_reserve];"
+ "if r0 == 0 goto 9f;"
+ "r6 = r0;"
+ "call pad_resumes_frame;" /* no record covers this call */
+ "r1 = r6;"
+ "r2 = 0;"
+ "call %[bpf_ringbuf_discard];"
+"9:"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_ringbuf_reserve), __imm(bpf_ringbuf_discard),
+ __imm_addr(unwind_ringbuf)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("an unwind through this call keeps the reference id=")
+int unwind_through_call_keeps_ref(void *ctx)
+{
+ return uncovered_holds_ref_frame();
+}
+
+/* The main program's frame is passed by the same way. */
+SEC("?syscall")
+__failure __msg("an unwind through this call keeps the reference id=")
+int unwind_through_call_main_keeps_ref(void *ctx)
+{
+ void *rec;
+
+ rec = bpf_ringbuf_reserve(&unwind_ringbuf, 8, 0);
+ if (!rec)
+ return 0;
+ pad_resumes_frame(); /* no record covers this call */
+ bpf_ringbuf_discard(rec, 0);
+ return 0;
+}
+
+char _license[] SEC("license") = "GPL";
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (18 preceding siblings ...)
2026-09-29 0:17 ` [PATCH bpf-next v7 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
@ 2026-09-29 0:17 ` Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-29 0:17 ` [PATCH bpf-next v7 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
21 siblings, 1 reply; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:17 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A test that has to look at a program's state after it ran needs a driver of
its own: open and load a skeleton, set an input, call
bpf_prog_test_run_opts(), read a variable back, destroy the skeleton. The
loader already does all of that for __retval(), and the only thing missing
is reaching the program's globals.
__set_global(var, value) writes one before the run and __ret_global(var,
value) checks one after it, both of them implying the execution __retval()
asks for. Either may be given more than once, for a test with more than
one input or more than one thing to look at afterwards; past
MAX_GLOBAL_VARS a tag is refused rather than quietly replacing the one
before it.
The variable is found the way veristat finds one: the map whose name ends
in ".bss" or ".data", the datasec of that name in the object's BTF, then
the variable within it. Those are two lookups that nothing so far has made
agree on a size, so where the variable lands is checked against the map it
was found beside. Four and eight byte variables are supported, which is
what a counter or a bitmask needs; anything else is refused rather than
read at the wrong width.
A value is held to the range its variable can represent, signedness taken
from the BTF the way veristat's set_global_var() does, and narrowed to
what is stored. Both tags use the one rule, so a negative literal means
the same to each: without it __ret_global(err, -22) on an int could never
match, the tag parsing 64 bits while the read gave back 32. Unlike
veristat this takes no enum names and runs native-endian only.
The value may be a '|' separated list of terms, so that a bitmask reads in
the test the way it is written in the program.
Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/testing/selftests/bpf/progs/bpf_misc.h | 7 +
tools/testing/selftests/bpf/test_loader.c | 314 ++++++++++++++++++-
2 files changed, 320 insertions(+), 1 deletion(-)
diff --git a/tools/testing/selftests/bpf/progs/bpf_misc.h b/tools/testing/selftests/bpf/progs/bpf_misc.h
index f3dbc3b59bff..9afa163fac5a 100644
--- a/tools/testing/selftests/bpf/progs/bpf_misc.h
+++ b/tools/testing/selftests/bpf/progs/bpf_misc.h
@@ -93,6 +93,11 @@
* __failure Expect program load failure in privileged mode.
* __failure_unpriv Expect program load failure in unprivileged mode.
*
+ * __set_global Set a global variable of the program to a value before
+ * executing it.
+ * __ret_global Execute the program and check that a global variable
+ * holds the given value afterwards. The variable has to
+ * live in .bss or .data and be four or eight bytes wide.
* __retval Execute the program using BPF_PROG_TEST_RUN command,
* expect return value to match passed parameter:
* - a decimal number
@@ -160,6 +165,8 @@
#define __log_level(lvl) __test_tag("test_log_level=" #lvl)
#define __flag(flag) __test_tag("test_prog_flags=" #flag)
#define __retval(val) __test_tag("test_retval=" XSTR(val))
+#define __set_global(var, val) __test_tag("test_global_set=" #var ":" XSTR(val))
+#define __ret_global(var, val) __test_tag("test_global_ret=" #var ":" XSTR(val))
#define __retval_unpriv(val) __test_tag("test_retval_unpriv=" XSTR(val))
#define __auxiliary __test_tag("test_auxiliary")
#define __auxiliary_unpriv __test_tag("test_auxiliary_unpriv")
diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
index 25eeb1c1248b..2492841f7a49 100644
--- a/tools/testing/selftests/bpf/test_loader.c
+++ b/tools/testing/selftests/bpf/test_loader.c
@@ -50,6 +50,13 @@ enum stack_mode {
SMALL_STACK = 1 << 1,
};
+#define MAX_GLOBAL_VARS 8
+
+struct global_var {
+ char *name;
+ __u64 val;
+};
+
struct test_subspec {
char *name;
char *description;
@@ -62,6 +69,10 @@ struct test_subspec {
int retval;
bool execute;
__u64 caps;
+ struct global_var set_globals[MAX_GLOBAL_VARS];
+ int set_global_cnt;
+ struct global_var ret_globals[MAX_GLOBAL_VARS];
+ int ret_global_cnt;
};
struct test_spec {
@@ -115,6 +126,34 @@ void free_msgs(struct expected_msgs *msgs)
msgs->cnt = 0;
}
+static void free_global_vars(struct global_var *vars, int *cnt)
+{
+ int i;
+
+ for (i = 0; i < *cnt; i++) {
+ free(vars[i].name);
+ vars[i].name = NULL;
+ }
+ *cnt = 0;
+}
+
+static int clone_global_vars(struct global_var *dst, int *dst_cnt,
+ const struct global_var *src, int src_cnt)
+{
+ int i;
+
+ for (i = 0; i < src_cnt; i++) {
+ dst[i].name = strdup(src[i].name);
+ if (!dst[i].name) {
+ free_global_vars(dst, dst_cnt);
+ return -ENOMEM;
+ }
+ dst[i].val = src[i].val;
+ (*dst_cnt)++;
+ }
+ return 0;
+}
+
static void free_test_spec(struct test_spec *spec)
{
/* Deallocate expect_msgs arrays. */
@@ -129,6 +168,11 @@ static void free_test_spec(struct test_spec *spec)
free_msgs(&spec->unpriv.stdout);
free_msgs(&spec->priv.stdout);
+ free_global_vars(spec->priv.set_globals, &spec->priv.set_global_cnt);
+ free_global_vars(spec->priv.ret_globals, &spec->priv.ret_global_cnt);
+ free_global_vars(spec->unpriv.set_globals, &spec->unpriv.set_global_cnt);
+ free_global_vars(spec->unpriv.ret_globals, &spec->unpriv.ret_global_cnt);
+
free(spec->priv.name);
free(spec->priv.description);
free(spec->unpriv.name);
@@ -311,6 +355,216 @@ static int parse_caps(const char *str, __u64 *val, const char *name)
return 0;
}
+static int parse_global_var(const char *str, struct global_var *vars, int *cnt,
+ const char *name)
+{
+ const char *colon = strrchr(str, ':');
+ struct global_var *var;
+ char *end;
+ __u64 *val;
+
+ if (!colon || colon == str) {
+ PRINT_FAIL("expecting '<variable>:<value>' for %s, got '%s'\n", name, str);
+ return -EINVAL;
+ }
+ if (*cnt >= MAX_GLOBAL_VARS) {
+ PRINT_FAIL("too many %s tags, at most %d are supported\n",
+ name, MAX_GLOBAL_VARS);
+ return -E2BIG;
+ }
+
+ var = &vars[*cnt];
+ val = &var->val;
+ *val = 0;
+ for (const char *term = colon + 1;;) {
+ __u64 v;
+
+ errno = 0;
+ v = strtoull(term, &end, 0);
+ if (errno || end == term) {
+ PRINT_FAIL("failed to parse %s value '%s'\n", name, colon + 1);
+ return -EINVAL;
+ }
+ *val |= v;
+ while (*end == ' ')
+ end++;
+ if (!*end)
+ break;
+ if (*end != '|') {
+ PRINT_FAIL("failed to parse %s value '%s'\n", name, colon + 1);
+ return -EINVAL;
+ }
+ term = end + 1;
+ }
+
+ var->name = strndup(str, colon - str);
+ if (!var->name) {
+ PRINT_FAIL("failed to allocate %s variable name\n", name);
+ return -ENOMEM;
+ }
+ (*cnt)++;
+
+ return 0;
+}
+
+/* As veristat's is_signed_type(): anything not plainly unsigned is signed. */
+static bool global_var_is_signed(const struct btf_type *t)
+{
+ if (btf_is_int(t))
+ return btf_int_encoding(t) & BTF_INT_SIGNED;
+ if (btf_is_any_enum(t))
+ return btf_kflag(t);
+ return true;
+}
+
+static int find_global_var(struct bpf_object *obj, const char *name,
+ struct bpf_map **map, __u32 *off, __u32 *sz,
+ bool *is_signed)
+{
+ static const char * const secs[] = { ".bss", ".data" };
+ struct btf *btf = bpf_object__btf(obj);
+ int i, s;
+ __u32 vsz;
+
+ if (!btf) {
+ PRINT_FAIL("no BTF for object\n");
+ return -ENOENT;
+ }
+
+ for (s = 0; s < ARRAY_SIZE(secs); s++) {
+ const struct btf_type *sec, *vt;
+ const struct btf_var_secinfo *vsi;
+ struct bpf_map *m = NULL, *iter;
+ size_t slen = strlen(secs[s]);
+ int id;
+
+ bpf_object__for_each_map(iter, obj) {
+ const char *mname = bpf_map__name(iter);
+ size_t len = mname ? strlen(mname) : 0;
+
+ if (len >= slen && strcmp(mname + len - slen, secs[s]) == 0) {
+ m = iter;
+ break;
+ }
+ }
+ id = btf__find_by_name_kind(btf, secs[s], BTF_KIND_DATASEC);
+ if (!m || id < 0)
+ continue;
+
+ sec = btf__type_by_id(btf, id);
+ vsi = btf_var_secinfos(sec);
+ for (i = 0; i < btf_vlen(sec); i++, vsi++) {
+ const struct btf_type *var = btf__type_by_id(btf, vsi->type);
+
+ if (strcmp(btf__name_by_offset(btf, var->name_off), name))
+ continue;
+ if (vsi->size != 4 && vsi->size != 8) {
+ PRINT_FAIL("'%s' is %u bytes, only 4 and 8 are supported\n",
+ name, vsi->size);
+ return -EINVAL;
+ }
+ vt = btf__type_by_id(btf, btf__resolve_type(btf, var->type));
+ if (!vt || !(btf_is_int(vt) || btf_is_any_enum(vt))) {
+ PRINT_FAIL("'%s' is not an int or an enum\n", name);
+ return -EINVAL;
+ }
+ *is_signed = global_var_is_signed(vt);
+ /*
+ * The map is found by the suffix of its name and the
+ * section by its own, so nothing so far has made the
+ * two agree on a size.
+ */
+ vsz = bpf_map__value_size(m);
+ if (vsi->offset > vsz || vsi->size > vsz - vsi->offset) {
+ PRINT_FAIL("'%s' at %u+%u is outside '%s' of %u bytes\n",
+ name, vsi->offset, vsi->size,
+ bpf_map__name(m), vsz);
+ return -EINVAL;
+ }
+ *map = m;
+ *off = vsi->offset;
+ *sz = vsi->size;
+ return 0;
+ }
+ }
+
+ PRINT_FAIL("no global variable '%s'\n", name);
+ return -ENOENT;
+}
+
+/*
+ * A tag's value is parsed as 64 bits, but the variable may be narrower and
+ * may be signed. Hold it to the range the variable can represent, the way
+ * veristat's set_global_var() does, and narrow it to what is stored.
+ */
+static int fit_global_var(const char *name, __u32 sz, bool is_signed, __u64 *val)
+{
+ long long v = (long long)*val;
+ long long max_val;
+ __u32 bits;
+
+ if (sz >= sizeof(*val))
+ return 0;
+ bits = sz * 8 - (is_signed ? 1 : 0);
+ max_val = 1ll << bits;
+ if (v >= max_val || v < (is_signed ? -max_val : 0)) {
+ PRINT_FAIL("value %lld for '%s' is out of range [%lld; %lld]\n",
+ v, name, is_signed ? -max_val : 0, max_val - 1);
+ return -EINVAL;
+ }
+ *val = (__u32)*val;
+ return 0;
+}
+
+static int access_global_var(struct bpf_object *obj, const char *name,
+ __u64 *val, __u32 *sz_out, bool *signed_out, bool set)
+{
+ __u32 off, sz, zero = 0;
+ bool is_signed;
+ struct bpf_map *map;
+ size_t vsz;
+ void *buf;
+ int err;
+
+ err = find_global_var(obj, name, &map, &off, &sz, &is_signed);
+ if (err)
+ return err;
+ if (sz_out)
+ *sz_out = sz;
+ if (signed_out)
+ *signed_out = is_signed;
+ if (set) {
+ err = fit_global_var(name, sz, is_signed, val);
+ if (err)
+ return err;
+ }
+
+ vsz = bpf_map__value_size(map);
+ buf = calloc(1, vsz);
+ if (!buf)
+ return -ENOMEM;
+
+ err = bpf_map__lookup_elem(map, &zero, sizeof(zero), buf, vsz, 0);
+ if (err) {
+ PRINT_FAIL("failed to read '%s': %d\n", name, err);
+ goto out;
+ }
+ if (!set) {
+ *val = sz == 4 ? *(__u32 *)(buf + off) : *(__u64 *)(buf + off);
+ goto out;
+ }
+ if (sz == 4)
+ *(__u32 *)(buf + off) = *val;
+ else
+ *(__u64 *)(buf + off) = *val;
+ err = bpf_map__update_elem(map, &zero, sizeof(zero), buf, vsz, 0);
+ if (err)
+ PRINT_FAIL("failed to write '%s': %d\n", name, err);
+out:
+ free(buf);
+ return err;
+}
+
static int parse_retval(const char *str, int *val, const char *name)
{
/*
@@ -557,6 +811,22 @@ static int parse_test_spec(struct test_loader *tester,
spec->mode_mask |= UNPRIV;
spec->unpriv.execute = true;
has_unpriv_retval = true;
+ } else if ((val = str_has_pfx(s, "test_global_set="))) {
+ err = parse_global_var(val, spec->priv.set_globals,
+ &spec->priv.set_global_cnt,
+ "__set_global");
+ if (err)
+ goto cleanup;
+ spec->priv.execute = true;
+ spec->mode_mask |= PRIV;
+ } else if ((val = str_has_pfx(s, "test_global_ret="))) {
+ err = parse_global_var(val, spec->priv.ret_globals,
+ &spec->priv.ret_global_cnt,
+ "__ret_global");
+ if (err)
+ goto cleanup;
+ spec->priv.execute = true;
+ spec->mode_mask |= PRIV;
} else if ((val = str_has_pfx(s, "test_log_level="))) {
err = parse_int(val, &spec->log_level, "test log level");
if (err)
@@ -742,6 +1012,23 @@ static int parse_test_spec(struct test_loader *tester,
spec->unpriv.execute = spec->priv.execute;
}
+ if (spec->priv.set_global_cnt && !spec->unpriv.set_global_cnt) {
+ err = clone_global_vars(spec->unpriv.set_globals,
+ &spec->unpriv.set_global_cnt,
+ spec->priv.set_globals,
+ spec->priv.set_global_cnt);
+ if (err)
+ goto cleanup;
+ }
+ if (spec->priv.ret_global_cnt && !spec->unpriv.ret_global_cnt) {
+ err = clone_global_vars(spec->unpriv.ret_globals,
+ &spec->unpriv.ret_global_cnt,
+ spec->priv.ret_globals,
+ spec->priv.ret_global_cnt);
+ if (err)
+ goto cleanup;
+ }
+
if (spec->unpriv.expect_msgs.cnt == 0)
clone_msgs(&spec->priv.expect_msgs, &spec->unpriv.expect_msgs);
if (spec->unpriv.expect_xlated.cnt == 0)
@@ -1356,7 +1643,7 @@ void run_subtest(struct test_loader *tester,
struct cap_state caps = {};
struct bpf_object *tobj;
struct bpf_map *map;
- int retval, err, i;
+ int retval, err, i, gi;
int links_cnt = 0;
bool should_load;
@@ -1532,6 +1819,14 @@ void run_subtest(struct test_loader *tester,
}
}
+ for (gi = 0; gi < subspec->set_global_cnt; gi++) {
+ __u64 v = subspec->set_globals[gi].val;
+
+ if (access_global_var(tobj, subspec->set_globals[gi].name,
+ &v, NULL, NULL, true))
+ goto tobj_cleanup;
+ }
+
err = do_prog_test_run(bpf_program__fd(tprog), &retval,
bpf_program__type(tprog) == BPF_PROG_TYPE_SYSCALL ? true : false,
spec->linear_sz);
@@ -1540,6 +1835,23 @@ void run_subtest(struct test_loader *tester,
goto tobj_cleanup;
}
+ for (gi = 0; gi < subspec->ret_global_cnt; gi++) {
+ __u64 want = subspec->ret_globals[gi].val, v = 0;
+ bool is_signed = false;
+ __u32 sz = 0;
+
+ if (access_global_var(tobj, subspec->ret_globals[gi].name, &v,
+ &sz, &is_signed, false))
+ goto tobj_cleanup;
+ if (fit_global_var(subspec->ret_globals[gi].name, sz, is_signed, &want))
+ goto tobj_cleanup;
+ if (v != want) {
+ PRINT_FAIL("Unexpected %s: 0x%llx != 0x%llx\n",
+ subspec->ret_globals[gi].name, v, want);
+ goto tobj_cleanup;
+ }
+ }
+
verify_stderr(bpf_program__fd(tprog), &subspec->stderr);
if (subspec->stdout.cnt) {
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (19 preceding siblings ...)
2026-09-29 0:17 ` [PATCH bpf-next v7 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
@ 2026-09-29 0:17 ` Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-29 0:17 ` [PATCH bpf-next v7 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
21 siblings, 1 reply; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:17 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
The end-to-end test walks one call chain with a pad in most of its frames.
This adds the shapes it does not reach, in the order the new file has
them:
- a pad that reads its frame's callee-saved registers
- a region ending on a 16-byte instruction
- a pad terminated by _Unwind_Resume rather than bpf_unwind_resume
- a pad whose first instruction is a nop
- a pad that indexes its frame by a register the frame set before the
unwinding call
- two pads with an uncovered frame between them
- a precision chain crossing a resume
- a frame above an unwind that never returns, keeping no exit of its own,
ending in a jump rather than an exit so the one to keep is not the last
instruction
- a region covering an indirect call
- a pad reached only by a speculative walk, which is loaded but not run
- a pad that touches a global of each width and sign a tag can name
- a frame holding a reference across a covered call, its pad the only
place that releases it
Each shape that runs wants the same three things: an input, a return
value, and the set of landing pads that ran. That is what __set_global(),
__retval() and __ret_global() say, so they say it and RUN_TESTS() does the
rest, which also gives each shape a name of its own in the test output.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
.../selftests/bpf/exceptions_cleanup.h | 12 +
.../bpf/prog_tests/exceptions_cleanup.c | 2 +
.../bpf/progs/exceptions_cleanup_shapes.c | 607 ++++++++++++++++++
3 files changed, 621 insertions(+)
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
index d1d40314035e..616b479238c7 100644
--- a/tools/testing/selftests/bpf/exceptions_cleanup.h
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -10,6 +10,18 @@
#define RAN_FOO2_DROP 0x8
#define RAN_BUMP 0x10
+/* progs/exceptions_cleanup_shapes.c: one bit per shape. */
+#define RAN_REGS 0x1
+#define RAN_WIDE_REC 0x2
+#define RAN_RESUME_ALIAS 0x4
+#define RAN_NOP_PAD 0x8
+#define RAN_VAR_STACK 0x10
+#define RAN_GAP_INNER 0x20
+#define RAN_GAP_OUTER 0x40
+#define RAN_NO_EXIT_JA 0x80
+#define RAN_CALLX 0x100
+#define RAN_HELD_REF 0x200
+
#define CLEANUP_REC(begin, end, landing_pad) \
".pushsection .bpf_cleanup,\"a\",@progbits;" \
".long " begin ";" \
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
index 255f88d35aad..c06ec10359b9 100644
--- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -4,6 +4,7 @@
#include "exceptions_cleanup.h"
#include "exceptions_cleanup.skel.h"
#include "exceptions_cleanup_fail.skel.h"
+#include "exceptions_cleanup_shapes.skel.h"
/* foo3 unwound: every frame that has a pad ran it. */
#define PADS_FOO3_UNWOUND \
@@ -82,4 +83,5 @@ void test_exceptions_cleanup(void)
exceptions_cleanup__destroy(skel);
RUN_TESTS(exceptions_cleanup_fail);
+ RUN_TESTS(exceptions_cleanup_shapes);
}
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
new file mode 100644
index 000000000000..1978fd1102f0
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
@@ -0,0 +1,607 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_unwind();
+ bpf_rcu_read_lock();
+ bpf_rcu_read_unlock();
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume(NULL);
+}
+
+__u64 input = 0;
+__u64 magic = 0x5eed;
+__u64 pads_ran = 0;
+
+/* A pad that reads r6-r9, which the callee overwrote before it unwound. */
+#define LOAD_MAGIC_REGS \
+ "r1 = %[magic] ll;" \
+ "r6 = *(u64 *)(r1 + 0);" \
+ "r7 = r6;" \
+ "r7 += 1;" \
+ "r8 = r6;" \
+ "r8 += 2;" \
+ "r9 = r6;" \
+ "r9 += 3;"
+
+/* Set @bit only if r6-r9 still hold what LOAD_MAGIC_REGS put there. */
+#define CHECK_MAGIC_REGS(bit) \
+ "r1 = %[magic] ll;" \
+ "r2 = *(u64 *)(r1 + 0);" \
+ "if r6 != r2 goto 9f;" \
+ "r2 += 1;" \
+ "if r7 != r2 goto 9f;" \
+ "r2 += 1;" \
+ "if r8 != r2 goto 9f;" \
+ "r2 += 1;" \
+ "if r9 != r2 goto 9f;" \
+ PAD_RAN(bit) \
+ "9:"
+
+/* The callee most of the shapes below unwind out of. */
+static __used __noinline __u64 pc_unwinder(__u64 x)
+{
+ if (x > 100)
+ bpf_unwind();
+ return x + 1;
+}
+
+static __used __naked __noinline __u64 regs_unwinder(void)
+{
+ asm volatile (
+ /* Not this frame's to keep, and that is the point. */
+ "r6 = 0xdead;"
+ "r7 = 0xbeef;"
+ "r8 = 0xcafe;"
+ "r9 = 0xf00d;"
+ "call bpf_unwind;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 regs_frame(void)
+{
+ asm volatile (
+ LOAD_MAGIC_REGS
+ "call bpf_preempt_disable;"
+"1:" "call regs_unwinder;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ CHECK_MAGIC_REGS("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_REGS),
+ __imm_addr(magic), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_REGS)
+int entry_regs(void *ctx)
+{
+ return regs_frame();
+}
+
+/* A region ending on a 16-byte insn, so end - 1 names its second half. */
+static __used __naked __noinline __u64 wide_rec_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "call bpf_rcu_read_lock;"
+ "r1 = r6;"
+"1:" "call pc_unwinder;" /* cleanup region begins */
+ "r1 = %[magic] ll;" /* ... and ends on this pair */
+"2:"
+ "r6 = r0;"
+ "call bpf_rcu_read_unlock;"
+ "r0 = r6;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+ "call bpf_rcu_read_unlock;"
+ PAD_RAN("%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_WIDE_REC), __imm_addr(input), __imm_addr(magic),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_WIDE_REC)
+int entry_wide_rec(void *ctx)
+{
+ return wide_rec_frame();
+}
+
+/* The name LLVM gives the resume: _Unwind_Resume(), which libbpf maps over. */
+extern void _Unwind_Resume(void *ptr) __ksym;
+
+static __used __noinline void __resume_alias_btf_anchor(void)
+{
+ _Unwind_Resume(NULL);
+}
+
+static __used __naked __noinline __u64 resume_alias_frame(void)
+{
+ asm volatile (
+"1:" "call regs_unwinder;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ PAD_RAN("%[ran]")
+ "call _Unwind_Resume;" /* the frontend's name for it */
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_RESUME_ALIAS), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_RESUME_ALIAS)
+int entry_resume_alias(void *ctx)
+{
+ return resume_alias_frame();
+}
+
+/* A pad starting on a nop, which opt_remove_nops() drops after the walk. */
+static __used __naked __noinline __u64 nop_pad_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "if r6 < 101 goto 6f;"
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+"6:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad: a nop, then its body */
+ "goto +0;"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_NOP_PAD),
+ __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_NOP_PAD)
+int entry_nop_pad(void *ctx)
+{
+ return nop_pad_frame();
+}
+
+/* Put @magic in both of the slots a variable offset could name. */
+#define FILL_MAGIC_SLOTS \
+ "r1 = %[magic] ll;" \
+ "r1 = *(u64 *)(r1 + 0);" \
+ "*(u64 *)(r10 - 8) = r1;" \
+ "*(u64 *)(r10 - 16) = r1;"
+
+/* Set @bit if the slot @idx names, read at a variable offset, holds it. */
+#define CHECK_VAR_SLOT(idx, bit) \
+ "r1 = r10;" \
+ "r1 += " idx ";" \
+ "r2 = *(u64 *)(r1 - 16);" \
+ "r3 = %[magic] ll;" \
+ "r3 = *(u64 *)(r3 + 0);" \
+ "if r2 != r3 goto 9f;" \
+ PAD_RAN(bit) \
+ "9:"
+
+/* An unwinder that touches none of r6-r9, so the walk leaves this frame. */
+static __used __naked __noinline __u64 var_unwinder(void)
+{
+ asm volatile (
+ "if r1 < 101 goto 1f;"
+ "call bpf_unwind;"
+"1:"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 var_stack_frame(void)
+{
+ asm volatile (
+ FILL_MAGIC_SLOTS
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "r6 &= 1;" /* an unknown slot number... */
+ "r6 <<= 3;" /* ...as an aligned byte offset */
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call var_unwinder;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+ CHECK_VAR_SLOT("r6", "%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_VAR_STACK), __imm_addr(input), __imm_addr(magic),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_VAR_STACK)
+int entry_var_stack(void *ctx)
+{
+ return var_stack_frame();
+}
+
+/* Two pads with an uncovered frame between them. */
+static __used __naked __noinline __u64 gap_inner_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call pc_unwinder;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r6 = r0;"
+ PAD_RAN("%[ran]")
+ "r1 = r6;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_GAP_INNER), __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/* The frame in between, with no record of its own. */
+static __used __noinline __u64 gap_mid(void)
+{
+ return gap_inner_frame() + 1;
+}
+
+static __used __naked __noinline __u64 gap_outer_frame(void)
+{
+ asm volatile (
+"1:" "call gap_mid;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r6 = r0;"
+ PAD_RAN("%[ran]")
+ "r1 = r6;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_GAP_OUTER), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/* And one more uncovered frame between the outer pad and the boundary. */
+static __used __noinline __u64 gap_top(void)
+{
+ return gap_outer_frame() + 1;
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_GAP_INNER | RAN_GAP_OUTER)
+int entry_two_pads(void *ctx)
+{
+ return gap_top();
+}
+
+/*
+ * A precision chain crossing a resume: r6 is kept across a call whose only
+ * way back is the callee's pad, then used as a variable stack offset.
+ */
+static __used __naked __noinline __u64 prec_inner_frame(void)
+{
+ asm volatile (
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 prec_outer_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "r6 &= 0x7;"
+ "r0 = 0;"
+ "*(u64 *)(r10 - 8) = r0;"
+ "*(u64 *)(r10 - 16) = r0;"
+ "call prec_inner_frame;" /* comes back only through the pad */
+ "r2 = r10;"
+ "r2 += -16;"
+ "r2 += r6;" /* variable stack offset: r6 must be precise */
+ "*(u8 *)(r2 + 0) = 1;"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+int entry_prec_across_resume(void *ctx)
+{
+ return prec_outer_frame();
+}
+
+/* Unwinds every time, and no record covers it, so the path simply ends. */
+static __used __naked __noinline __u64 always_unwind(void)
+{
+ asm volatile (
+ "call bpf_unwind;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/*
+ * A frame with no record of its own above one that always unwinds: nothing
+ * after the call is reachable, so it keeps no exit and gets no epilogue.
+ * This one ends in a jump rather than an exit, so the exit to keep is not
+ * the last instruction.
+ */
+static __used __naked __noinline __u64 no_exit_ja_mid(void)
+{
+ asm volatile (
+ "goto 2f;"
+"1:" "r0 = 1;"
+ "exit;"
+"2:" "call always_unwind;"
+ "goto 1b;" /* the last insn, and not an exit */
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 no_exit_ja_outer_frame(void)
+{
+ asm volatile (
+"1:" "call no_exit_ja_mid;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_NO_EXIT_JA), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__ret_global(pads_ran, RAN_NO_EXIT_JA)
+int entry_no_exit_ja_mid(void *ctx)
+{
+ return no_exit_ja_outer_frame();
+}
+
+/* gcc has no indirect calls, and only these JITs emit them */
+#if defined(__clang__) && \
+ (defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64))
+
+/*
+ * A region covering an indirect call: a record names a call by its return
+ * address, which a callx leaves like any other call.
+ */
+static __used __naked __noinline __u64 callx_unwinder(void)
+{
+ asm volatile (
+ "call bpf_unwind;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 callx_region_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r2 = %[callx_unwinder] ll;"
+"1:" "callx r2;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_CALLX), __imm_addr(callx_unwinder),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__ret_global(pads_ran, RAN_CALLX)
+int entry_callx_region(void *ctx)
+{
+ return callx_region_frame();
+}
+
+#endif /* __clang__ && (x86 || arm64) */
+
+/*
+ * The jump_into_pad shape with the branch statically dead, so only a
+ * speculative walk reaches the pad: a barrier rather than a refusal.
+ */
+static __used __naked __noinline __u64 dead_jump_into_pad_frame(void)
+{
+ asm volatile (
+ "r6 = 0;"
+ "if r6 > 7 goto 4f;" /* never taken: walked speculatively */
+"1:" "call always_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+"4:" /* ... and its second instruction */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__success
+int dead_jump_into_pad(void *ctx)
+{
+ return dead_jump_into_pad_frame();
+}
+
+/* Add one to the @w-bit global at @addr. */
+#define BUMP_GLOBAL(w, addr) \
+ "r1 = " addr " ll;" \
+ "r2 = *(u" w " *)(r1 + 0);" \
+ "r2 += 1;" \
+ "*(u" w " *)(r1 + 0) = r2;"
+
+/*
+ * A pad that touches a global of each width and sign a test tag can name,
+ * so that __set_global() and __ret_global() are exercised on all four.
+ */
+int tag_i = 0;
+unsigned int tag_ui = 0;
+long tag_l = 0;
+unsigned long tag_ul = 0;
+
+static __used __naked __noinline __u64 tag_types_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call pc_unwinder;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ BUMP_GLOBAL("32", "%[tag_i]")
+ BUMP_GLOBAL("32", "%[tag_ui]")
+ BUMP_GLOBAL("64", "%[tag_l]")
+ BUMP_GLOBAL("64", "%[tag_ul]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm_addr(input), __imm_addr(tag_i), __imm_addr(tag_ui),
+ __imm_addr(tag_l), __imm_addr(tag_ul)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__set_global(input, 101)
+__set_global(tag_i, -23) __set_global(tag_ui, 0xfffffffe)
+__set_global(tag_l, -23) __set_global(tag_ul, 0xfffffffffffffffe)
+__ret_global(tag_i, -22) __ret_global(tag_ui, 0xffffffff)
+__ret_global(tag_l, -22) __ret_global(tag_ul, 0xffffffffffffffff)
+int entry_tag_types(void *ctx)
+{
+ return tag_types_frame();
+}
+
+struct {
+ __uint(type, BPF_MAP_TYPE_RINGBUF);
+ __uint(max_entries, 4096);
+} shape_ringbuf SEC(".maps");
+
+/*
+ * A frame holding a reference across a call an unwind comes out of. The record
+ * over the call is what lets it hold one: the pad releases it, where a frame
+ * with no record would be left for its epilogue still holding it.
+ */
+static __used __naked __noinline __u64 held_ref_frame(void)
+{
+ asm volatile (
+ "r1 = %[shape_ringbuf] ll;"
+ "r2 = 8;"
+ "r3 = 0;"
+ "call %[bpf_ringbuf_reserve];"
+ "if r0 == 0 goto 9f;"
+ "r6 = r0;"
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call pc_unwinder;" /* cleanup region */
+"2:"
+ "r1 = r6;"
+ "r2 = 0;"
+ "call %[bpf_ringbuf_discard];"
+ "goto 9f;"
+"3:" /* landing pad: release and resume */
+ "r1 = r6;"
+ "r2 = 0;"
+ "call %[bpf_ringbuf_discard];"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+"9:"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_HELD_REF), __imm(bpf_ringbuf_reserve),
+ __imm(bpf_ringbuf_discard), __imm_addr(shape_ringbuf),
+ __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_HELD_REF)
+int entry_held_ref(void *ctx)
+{
+ return held_ref_frame();
+}
+
+char _license[] SEC("license") = "GPL";
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* [PATCH bpf-next v7 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (20 preceding siblings ...)
2026-09-29 0:17 ` [PATCH bpf-next v7 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
@ 2026-09-29 0:17 ` Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
21 siblings, 1 reply; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 0:17 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
For a light skeleton, libbpf hands the records to bpf_gen__prog_load(),
which writes them into a blob and emits a loader program that issues
BPF_PROG_LOAD from inside the kernel. Nothing about that is shared with the
ordinary path: the attr is built field by field, and the kfunc names are
resolved by the loader program when it runs.
exceptions_cleanup_light.c is the smallest program that can tell whether a
table survives that trip: one record covering the bpf_unwind() call itself,
one pad that sets a bit and resumes. If the table arrives, the pad runs and
sets its bit; if it does not, the program does not load at all, because the
pad is then reachable from nothing.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/testing/selftests/bpf/Makefile.skel | 2 +-
.../selftests/bpf/exceptions_cleanup.h | 3 ++
.../bpf/prog_tests/exceptions_cleanup.c | 28 +++++++++++++
.../bpf/progs/exceptions_cleanup_light.c | 39 +++++++++++++++++++
4 files changed, 71 insertions(+), 1 deletion(-)
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
diff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel
index 2e22bb901bf3..41575e180efb 100644
--- a/tools/testing/selftests/bpf/Makefile.skel
+++ b/tools/testing/selftests/bpf/Makefile.skel
@@ -33,7 +33,7 @@ LINKED_SKELS := test_static_linked.skel.h linked_funcs.skel.h \
LSKELS := fexit_sleep.c trace_printk.c trace_vprintk.c map_ptr_kern.c \
core_kern.c core_kern_overflow.c test_ringbuf.c \
test_ringbuf_n.c test_ringbuf_map_key.c test_ringbuf_write.c \
- test_ringbuf_overwrite.c callx_rodata.c
+ test_ringbuf_overwrite.c callx_rodata.c exceptions_cleanup_light.c
LSKELS_SIGNED := fentry_test.c fexit_test.c atomics.c
diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
index 616b479238c7..f3dac2048f06 100644
--- a/tools/testing/selftests/bpf/exceptions_cleanup.h
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -22,6 +22,9 @@
#define RAN_CALLX 0x100
#define RAN_HELD_REF 0x200
+/* progs/exceptions_cleanup_light.c: the one pad it has. */
+#define RAN_LIGHT 0x1
+
#define CLEANUP_REC(begin, end, landing_pad) \
".pushsection .bpf_cleanup,\"a\",@progbits;" \
".long " begin ";" \
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
index c06ec10359b9..d251148279bf 100644
--- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -5,6 +5,7 @@
#include "exceptions_cleanup.skel.h"
#include "exceptions_cleanup_fail.skel.h"
#include "exceptions_cleanup_shapes.skel.h"
+#include "exceptions_cleanup_light.lskel.h"
/* foo3 unwound: every frame that has a pad ran it. */
#define PADS_FOO3_UNWOUND \
@@ -37,6 +38,30 @@ static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
ASSERT_EQ(skel->bss->pads_ran, pads | RAN_BUMP, "pads_ran");
}
+static void test_light_skeleton(void)
+{
+ struct exceptions_cleanup_light_lskel *skel;
+ __u64 ctx = 0;
+ int err;
+
+ LIBBPF_OPTS(bpf_test_run_opts, topts,
+ .ctx_in = &ctx,
+ .ctx_size_in = sizeof(ctx),
+ );
+
+ skel = exceptions_cleanup_light_lskel__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "light open_and_load"))
+ return;
+
+ err = bpf_prog_test_run_opts(skel->progs.entry_light.prog_fd, &topts);
+ if (!ASSERT_OK(err, "run"))
+ goto out;
+ ASSERT_EQ(topts.retval, 0, "retval");
+ ASSERT_EQ(skel->bss->pads_ran, RAN_LIGHT, "pads_ran");
+out:
+ exceptions_cleanup_light_lskel__destroy(skel);
+}
+
void test_exceptions_cleanup(void)
{
char log[8192] = {};
@@ -82,6 +107,9 @@ void test_exceptions_cleanup(void)
exceptions_cleanup__destroy(skel);
+ if (test__start_subtest("light_skeleton"))
+ test_light_skeleton();
+
RUN_TESTS(exceptions_cleanup_fail);
RUN_TESTS(exceptions_cleanup_shapes);
}
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
new file mode 100644
index 000000000000..19d53569bbda
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
@@ -0,0 +1,39 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_unwind();
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume(NULL);
+}
+
+__u64 pads_ran = 0;
+
+SEC("syscall")
+__naked int entry_light(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_LIGHT),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 12/22] bpf, x86: Dispatch exception cleanup pads at run time
2026-09-29 0:17 ` [PATCH bpf-next v7 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
@ 2026-09-29 0:30 ` sashiko-bot
2026-09-30 1:34 ` Yonghong Song
0 siblings, 1 reply; 46+ messages in thread
From: sashiko-bot @ 2026-09-29 0:30 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: exception unwinding aborted when bpf_unwind is traced via kretprobe
--
commit f3ba2238caee72a50c097a775a0290b3b66ab98b
Author: Yonghong Song <yonghong.song@linux.dev>
bpf, x86: Dispatch exception cleanup pads at run time
This commit updates the BPF JIT on x86 to dispatch exception cleanup pads at
run time by writing directly to the return-address slot. It introduces
arch_bpf_stack_walk_ra() to support unwinding alongside the ORC unwinder.
> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
> index 6c7a0578760e7..d4feade5b5c72 100644
> --- a/arch/x86/net/bpf_jit_comp.c
> +++ b/arch/x86/net/bpf_jit_comp.c
[ ... ]
> @@ -4614,6 +4628,34 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
> #endif
> }
>
> +void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
> + void *cookie)
> +{
> +#if defined(CONFIG_UNWINDER_ORC)
> + struct unwind_state state;
> + unsigned long addr, *ra;
> +
> + for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
> + unwind_next_frame(&state)) {
> + addr = unwind_get_return_address(&state);
> + ra = unwind_get_return_address_ptr(&state);
> + if (!addr || !ra)
> + break;
> + /*
> + * A traced return: the unwinder recovered @addr from under a
> + * function graph or kretprobe trampoline, which is what the
> + * slot itself still holds. Writing there would skip the
> + * trampoline and leave its entry for the next hooked return
> + * to pop.
> + */
> + if (READ_ONCE_NOCHECK(*ra) != addr)
> + break;
[Severity: High]
Will this unconditionally abort the unwinder if a kernel function like
bpf_unwind() is traced via a kretprobe or function graph tracer?
When a BPF program triggers an exception, bpf_unwind() invokes
arch_bpf_stack_walk_ra(). The unwinder processes kernel frames first. If
a user has legitimately attached a tracer to bpf_unwind(), the stack return
address slot (*ra) for that frame will hold the tracer's trampoline, while
the unwinder recovers the real return address (addr).
Because this evaluates to true and breaks the loop, it seems the walk stops
completely before it even reaches the BPF frames. Could this result in a
silent failure of BPF exception handling and resource leaks (e.g., unreleased
spinlocks or memory) since the cleanup pads are never executed?
> + if (!consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
> + break;
> + }
> +#endif
> +}
> +
> void bpf_arch_poke_desc_update(struct bpf_jit_poke_descriptor *poke,
> struct bpf_prog *new, struct bpf_prog *old)
> {
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260929001601.3242665-1-yonghong.song@linux.dev?part=12
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 07/22] bpf: Resume a covered call at its landing pad
2026-09-29 0:16 ` [PATCH bpf-next v7 07/22] bpf: Resume a covered call at its landing pad Yonghong Song
@ 2026-09-29 0:31 ` sashiko-bot
2026-09-30 0:28 ` Yonghong Song
0 siblings, 1 reply; 46+ messages in thread
From: sashiko-bot @ 2026-09-29 0:31 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Critical] bpf: landing pad state desync via discarded memory mutations allows OOB access
--
commit 41b3c5e14924fae7175095b7012838a14c2e5ffc
Author: Yonghong Song <yonghong.song@linux.dev>
bpf: Resume a covered call at its landing pad
This commit allows a covered call to resume at its landing pad when bpf_unwind()
is invoked by a subprogram. It pushes a cleanup pad branch with the caller's
pre-call state and handles bpf_unwind() by returning PROCESS_BPF_EXIT for
nested frames.
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index fc3df452de2e4..ee074d4a936b5 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
[ ... ]
> @@ -19167,6 +19186,57 @@ enum {
> INSN_IDX_UPDATED = 2,
> };
>
> +static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
> +{
> + struct bpf_verifier_state *branch;
> + struct bpf_func_state *frame;
> + int pad = bpf_exc_pad_of_call(env, insn_idx);
> +
> + if (pad < 0)
> + return 0;
> + branch = push_stack(env, pad, insn_idx, false);
> + if (IS_ERR(branch))
> + return PTR_ERR(branch);
> + frame = branch->frame[branch->curframe];
> + /*
> + * The state at that call with the caller-saved registers gone: the
> + * callee's epilogue put r6-r9 and the stack back on the way out.
> + */
> + clear_caller_saved_regs(env, frame->regs);
> + mark_reg_unknown(env, frame->regs, BPF_REG_0);
> + return 0;
> +}
> +
> +static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
> + bool *do_print_state)
> +{
> + struct bpf_func_state *frame = cur_func(env);
> + int pad = bpf_exc_pad_of_call(env, *insn_idx);
> + int err;
> +
> + if (pad < 0) {
> + err = check_resource_leak(env, false, !env->cur_state->curframe,
> + "an unwind with no landing pad");
> + if (err)
> + return err;
> + if (env->cur_state->curframe)
> + return PROCESS_BPF_EXIT;
[Severity: Critical]
When process_bpf_unwind() returns PROCESS_BPF_EXIT here, the verifier stops
exploring the callee and discards its state. Does this discard any memory
mutations made by the callee to pointer arguments before calling bpf_unwind()?
Since the CPU unwinds the stack without undoing memory writes at runtime,
could this desynchronize the verifier state from runtime state and allow
out-of-bounds memory accesses?
> + /*
> + * The main program's frame returns at once, which is the
> + * program returning. Mark r0 the zero the fixups leave after
> + * the call, and leave through the exit, which is what holds
> + * that zero to the program type.
> + */
> + mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
> + mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
> + return process_bpf_exit_full(env, do_print_state, false);
> + }
> + clear_caller_saved_regs(env, frame->regs);
> + mark_reg_unknown(env, frame->regs, BPF_REG_0);
> + *insn_idx = pad;
> + return INSN_IDX_UPDATED;
> +}
> +
[ ... ]
> @@ -19421,7 +19491,29 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
> return -EINVAL;
> }
> }
> + if (bpf_is_unwind_kfunc(insn) || bpf_is_unwind_resume_kfunc(insn)) {
> + err = check_kfunc_allowed_only(env, insn, env->insn_idx);
> + if (err)
> + return err;
> + if (bpf_is_unwind_kfunc(insn))
> + return process_bpf_unwind(env, &env->insn_idx,
> + do_print_state);
> + /*
> + * Mark r0 a known zero -- unknown first, as
> + * the known-zero helper keeps the type it
> + * finds, which here is NOT_INIT. The fixups
> + * lower this to 'r0 = 0; exit', so the frame
> + * returns a real zero.
> + */
> + mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
> + mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
> + return process_bpf_exit_full(env, do_print_state, false);
> + }
> mark_reg_scratched(env, BPF_REG_0);
> + /* An unwind out of this call resumes at the pad. */
> + err = push_cleanup_pad_branch(env, env->insn_idx);
> + if (err)
> + return err;
[Severity: Critical]
push_cleanup_pad_branch() pushes a branch using a snapshot of the pre-call
state. If the callee modified a stack variable (e.g., an array index) and then
unwound the stack, does the verifier evaluate the landing pad with the original
unmodified value?
Could a BPF program use this to bypass bounds checks by mutating an index in
the callee before unwinding, causing the verifier to optimize out the checks
in the landing pad while using the out-of-bounds value at runtime?
> if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
> cur_func(env)->no_stack_arg_load = true;
> if (bpf_is_callx(insn))
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260929001601.3242665-1-yonghong.song@linux.dev?part=7
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 05/22] bpf: Prepare for an exception cleanup table before the CFG walk
2026-09-29 0:16 ` [PATCH bpf-next v7 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
@ 2026-09-29 0:31 ` sashiko-bot
2026-09-29 22:04 ` Yonghong Song
0 siblings, 1 reply; 46+ messages in thread
From: sashiko-bot @ 2026-09-29 0:31 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: cleanup_pad adjustment fails for BPF_PATCH_MOVE_TARGET at off=0
--
commit fc0420a45de5f3cee47a65d96e78244ace4fbf4e
Author: Yonghong Song <yonghong.song@linux.dev>
bpf: Prepare for an exception cleanup table before the CFG walk
This patch prepares an exception cleanup table for BPF programs, adding
landing pads to the instruction auxiliary data for calls that can unwind.
It introduces the cleanup_pad field to track where execution should resume
during exception unwinding.
> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
> index b19fcbf49b7ef..c1779d2d02f05 100644
> --- a/kernel/bpf/exception.c
> +++ b/kernel/bpf/exception.c
[ ... ]
> +static int mark_call_sites(struct bpf_verifier_env *env)
> +{
> + u32 i, j;
> +
> + for (i = 0; i < env->cleanup_info_cnt; i++) {
> + struct bpf_cleanup_info *rec = &env->cleanup_info[i];
> +
> + for (j = rec->begin_off; j < rec->end_off; j++) {
> + struct bpf_insn *insn = &env->prog->insnsi[j];
> +
> + if (!bpf_pseudo_call(insn) && !bpf_is_callx(insn) &&
> + !bpf_is_unwind_kfunc(insn))
> + continue;
> + env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
[Severity: High]
How does this new cleanup_pad field handle instruction patching when a
landing pad is located at instruction 0?
If mark_call_sites() sets cleanup_pad = rec->landing_pad_off + 1, and the
landing pad is at index 0, cleanup_pad becomes 1.
If bpf_convert_ctx_accesses() in kernel/bpf/fixups.c later inserts a
prologue at instruction 0 using BPF_PATCH_MOVE_TARGET,
adjust_insn_aux_data() will try to update the aux data offsets. However,
its check:
if (data[i].cleanup_pad > off + 1)
data[i].cleanup_pad += cnt - 1;
will evaluate to 1 > 1 (false) when off=0. The required shift is skipped,
leaving the cleanup_pad pointing to the newly inserted prologue instead of
shifting to the original instruction. Unwinding to this pad could re-execute
the prologue and potentially corrupt the frame pointer or stack.
Can this sequence occur and lead to stack issues?
> + }
> + }
> + return 0;
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260929001601.3242665-1-yonghong.song@linux.dev?part=5
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 04/22] bpf: Add lookups for exception cleanup resumes and landing pads
2026-09-29 0:16 ` [PATCH bpf-next v7 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
@ 2026-09-29 0:33 ` sashiko-bot
2026-09-29 21:58 ` Yonghong Song
0 siblings, 1 reply; 46+ messages in thread
From: sashiko-bot @ 2026-09-29 0:33 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: missing cleanup_pad adjustment in BPF_PATCH_MOVE_TARGET
--
commit ddeb49ab1c3e86d5fea9b3b127432b6e34db617b
Author: Yonghong Song <yonghong.song@linux.dev>
bpf: Add lookups for exception cleanup resumes and landing pads
This commit introduces helpers to recognize exception cleanup resumes and
landing pads. It also adds logic to maintain cleanup_pad data in
insn_aux_data when instructions are patched or removed.
> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
> index 37cf130ebb57b..5b7fe4ba610b3 100644
> --- a/kernel/bpf/fixups.c
> +++ b/kernel/bpf/fixups.c
> @@ -268,11 +268,18 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
> data[i].non_stack_access =
> data[off + cnt - 1].non_stack_access;
> data[off + cnt - 1].non_stack_access = false;
> + data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
> + data[off + cnt - 1].cleanup_pad = 0;
> } else if (bpf_is_mem_insn(insn + i)) {
> data[i].non_stack_access = true;
> }
> }
>
> + if (env->cleanup_info_cnt)
> + for (i = 0; i < prog_len; i++)
> + if (data[i].cleanup_pad > off + 1)
> + data[i].cleanup_pad += cnt - 1;
[Severity: Medium]
When mode is BPF_PATCH_MOVE_TARGET and the landing pad is located exactly
at the patched instruction, does this condition fail to adjust it?
Since cleanup_pad is 1-based, a value of off + 1 means the landing pad is
exactly at off. Because this strictly checks > off + 1, it skips adjusting
data[i].cleanup_pad when it equals off + 1.
If BPF_PATCH_MOVE_TARGET is used (for example, with gen_prologue or may_goto
initialization), the original instruction is moved to the end of the patch
sequence. This means the landing pad should follow it and move to the end
as well, but it will be left pointing to the start of the inserted sequence.
Could this result in the unwinder jumping to the newly inserted prologue or
loop initialization instead of the actual exception landing pad?
> +
> /*
> * Last slot instruction could be a newly generated
> * BPF_ST/BPF_LDX/BPF_STX, systematically mark it for non-stack access
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260929001601.3242665-1-yonghong.song@linux.dev?part=4
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with
2026-09-29 0:16 ` [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
@ 2026-09-29 0:36 ` sashiko-bot
2026-09-30 1:09 ` Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
1 sibling, 1 reply; 46+ messages in thread
From: sashiko-bot @ 2026-09-29 0:36 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Critical] bpf: type confusion and memory corruption due to stale cleanup pad state
- [High] bpf: lock corruption and double-unlock via missing identity check on unwind
--
commit 585aa50d9e8811792b9a816a4d1ad6cd51d63ea7
Author: Yonghong Song <yonghong.song@linux.dev>
Subject: bpf: Require an unwind to leave a frame holding what it entered with
This commit updates the BPF verifier to record what locks and references
a program holds when a frame is entered. It enforces that an unwind must
leave the frame holding what it entered with, ensuring that landing pads
verified against pre-call snapshots remain valid.
> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 0143688896b0e..75e572a8a1bab 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -339,6 +339,18 @@ struct bpf_func_state {
> bool in_async_callback_fn;
> bool in_exception_callback_fn;
> bool no_stack_arg_load;
> + /*
> + * What the program held when this frame was entered. An unwind leaves
> + * the frame without running anything below it, so the frame has to put
> + * these back to what it found before it goes -- otherwise a caller's
> + * landing pad, whose state was taken at the call, is wrong about them.
> + */
> + u32 entry_active_locks;
> + u32 entry_preempt_locks;
> + u32 entry_rcu_locks;
> + u32 entry_irq_id;
> + u32 entry_id_gen;
> + u32 entry_acquired_refs;
[Severity: High]
Does this structure need fields to record the identity of the lock held upon
frame entry? Without this, it might be difficult to accurately enforce the
correct lock state on unwind.
Additionally, is it necessary to update func_states_equal() to compare these
new entry_* fields? Without comparing them, could the verifier unsafely prune
states with different entry locking conditions?
[ ... ]
> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
> index c1779d2d02f05..c0b0b8478af51 100644
> --- a/kernel/bpf/exception.c
> +++ b/kernel/bpf/exception.c
> @@ -12,6 +12,72 @@
> BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
> BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
>
> +void bpf_exc_record_frame_entry(const struct bpf_verifier_state *state,
> + struct bpf_func_state *frame, u32 id_gen)
> +{
[ ... ]
> +int bpf_exc_check_frame_balance(struct bpf_verifier_env *env, const char *prefix)
> +{
> + const struct bpf_verifier_state *state = env->cur_state;
> + const struct bpf_func_state *frame = cur_func(env);
> + u32 i, held;
> + const char *what;
> +
> + if (state->active_rcu_locks != frame->entry_rcu_locks)
> + what = "bpf_rcu_read_lock";
> + else if (state->active_preempt_locks != frame->entry_preempt_locks)
> + what = "bpf_preempt_disable";
> + else if (state->active_irq_id != frame->entry_irq_id)
> + what = "bpf_local_irq_save";
> + else if (state->active_locks != frame->entry_active_locks)
> + what = "bpf_spin_lock";
[Severity: High]
Does this code adequately verify lock identity? While it checks the count of
active locks, it appears to miss checking the identity of the lock
(active_lock_id and active_lock_ptr).
If a BPF program acquires lock A, passes it to a subprogram which unlocks A,
locks B, and then throws an exception (bpf_unwind()), could the unwinder
accept the frame balance since the active_locks count is still 1?
The exception would jump back to the caller's cleanup pad, which assumes
lock A is still held. The cleanup pad unlocking A could potentially cause a
double-unlock of lock A and leak lock B forever.
> + else
> + what = NULL;
[ ... ]
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index ee074d4a936b5..4bdee3f02fe9d 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
[ ... ]
> @@ -19512,6 +19554,10 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
> mark_reg_scratched(env, BPF_REG_0);
> /* An unwind out of this call resumes at the pad. */
> err = push_cleanup_pad_branch(env, env->insn_idx);
[Severity: Critical]
Does taking a pre-call snapshot of the verifier state for the cleanup pad
incorrectly ignore valid memory modifications made by the callee before
unwinding?
For example, if a caller spills a valid pointer to its stack and passes a
pointer to this stack slot to a subprogram:
- The verifier schedules the caller's cleanup pad branch by pushing a snapshot
of the verifier state (including the stack) taken BEFORE the callee runs.
- The subprogram overwrites the caller's stack slot with an attacker-controlled
scalar and calls bpf_unwind().
- At runtime, the exception jumps to the cleanup pad, which loads the scalar.
Because the cleanup pad uses the pre-call snapshot, does the verifier
incorrectly assume the loaded value is still a valid pointer, potentially
leading to type confusion and memory corruption?
> + if (err)
> + return err;
> + /* Or, with no pad, leaves the frame for good. */
> + err = check_unwind_through_call(env, env->insn_idx);
> if (err)
> return err;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260929001601.3242665-1-yonghong.song@linux.dev?part=8
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton
2026-09-29 0:17 ` [PATCH bpf-next v7 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
@ 2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 3:12 ` Yonghong Song
0 siblings, 1 reply; 46+ messages in thread
From: bot+bpf-ci @ 2026-09-29 0:52 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3318 bytes --]
> diff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel
> index 2e22bb901bf3..41575e180efb 100644
> --- a/tools/testing/selftests/bpf/Makefile.skel
> +++ b/tools/testing/selftests/bpf/Makefile.skel
> @@ -33,7 +33,7 @@ LINKED_SKELS := test_static_linked.skel.h linked_funcs.skel.h \
> LSKELS := fexit_sleep.c trace_printk.c trace_vprintk.c map_ptr_kern.c \
> core_kern.c core_kern_overflow.c test_ringbuf.c \
> test_ringbuf_n.c test_ringbuf_map_key.c test_ringbuf_write.c \
> - test_ringbuf_overwrite.c callx_rodata.c
> + test_ringbuf_overwrite.c callx_rodata.c exceptions_cleanup_light.c
[ ... ]
> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
> new file mode 100644
> index 000000000000..19d53569bbda
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
> @@ -0,0 +1,39 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
> +#include <vmlinux.h>
> +#include <bpf/bpf_helpers.h>
> +#include "bpf_misc.h"
> +#include "exceptions_cleanup.h"
> +
> +static __used __noinline void __kfunc_btf_anchor(void)
> +{
> + bpf_throw(0);
> + bpf_unwind();
> + bpf_preempt_disable();
> + bpf_preempt_enable();
> + bpf_unwind_resume(NULL);
> +}
> +
> +__u64 pads_ran = 0;
> +
> +SEC("syscall")
> +__naked int entry_light(void)
> +{
> + asm volatile (
> + "call bpf_preempt_disable;"
> +"1:" "call bpf_unwind;" /* cleanup region */
> +"2:"
> + "exit;"
> +"3:" /* landing pad */
> + "call bpf_preempt_enable;"
> + PAD_RAN("%[ran]")
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + :
> + : [ran]"i"(RAN_LIGHT),
> + __imm_addr(pads_ran)
> + : __clobber_all);
> +}
> +
> +char _license[] SEC("license") = "GPL";
A subsystem pattern flags this as potentially concerning:
This adds a new BPF program file to test loading a cleanup table through
the light skeleton path. The commit message states this is the smallest
program that can verify whether the table survives the trip through
bpf_gen__prog_load().
However, progs/exceptions_cleanup.c already covers the same behavior
more broadly - pads run on unwind and set bits in pads_ran. Makefile.skel
already has LSKELS_EXTRA for building programs as both light skeleton
and libbpf skeleton. Could exceptions_cleanup.c be built as a light
skeleton and have the existing test cases (no_unwind, unwind_from_foo3,
unwind_from_foo2) run through the gen loader path?
That approach would also exercise the light skeleton code paths this new
program never reaches. The new program has only one record in the main
function, so in bpf_prog_collect_cleanup_info() it takes the owner == prog
branch. The owner->sub_insn_off rebasing for records inside appended
subprograms, and the multi-frame bpf_unwind_rewrite() walk, would not run
under the gen loader with this minimal test.
Should this be a light-skeleton case of the existing exceptions_cleanup.c
instead of a separate file?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36502848904
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags
2026-09-29 0:17 ` [PATCH bpf-next v7 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
@ 2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:46 ` Yonghong Song
0 siblings, 1 reply; 46+ messages in thread
From: bot+bpf-ci @ 2026-09-29 0:52 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 4055 bytes --]
> diff --git a/tools/testing/selftests/bpf/progs/bpf_misc.h b/tools/testing/selftests/bpf/progs/bpf_misc.h
> index f3dbc3b59bff..9afa163fac5a 100644
> --- a/tools/testing/selftests/bpf/progs/bpf_misc.h
> +++ b/tools/testing/selftests/bpf/progs/bpf_misc.h
> @@ -93,6 +93,11 @@
> * __failure Expect program load failure in privileged mode.
> * __failure_unpriv Expect program load failure in unprivileged mode.
> *
> + * __set_global Set a global variable of the program to a value before
> + * executing it.
> + * __ret_global Execute the program and check that a global variable
> + * holds the given value afterwards. The variable has to
> + * live in .bss or .data and be four or eight bytes wide.
> * __retval Execute the program using BPF_PROG_TEST_RUN command,
> * expect return value to match passed parameter:
> * - a decimal number
[ ... ]
> diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
> index 25eeb1c1248b..2492841f7a49 100644
> --- a/tools/testing/selftests/bpf/test_loader.c
> +++ b/tools/testing/selftests/bpf/test_loader.c
[ ... ]
> +static int find_global_var(struct bpf_object *obj, const char *name,
> + struct bpf_map **map, __u32 *off, __u32 *sz,
> + bool *is_signed)
> +{
> + static const char * const secs[] = { ".bss", ".data" };
> + struct btf *btf = bpf_object__btf(obj);
> + int i, s;
> + __u32 vsz;
> +
> + if (!btf) {
> + PRINT_FAIL("no BTF for object\n");
> + return -ENOENT;
> + }
> +
> + for (s = 0; s < ARRAY_SIZE(secs); s++) {
> + const struct btf_type *sec, *vt;
> + const struct btf_var_secinfo *vsi;
> + struct bpf_map *m = NULL, *iter;
> + size_t slen = strlen(secs[s]);
> + int id;
> +
> + bpf_object__for_each_map(iter, obj) {
> + const char *mname = bpf_map__name(iter);
> + size_t len = mname ? strlen(mname) : 0;
> +
> + if (len >= slen && strcmp(mname + len - slen, secs[s]) == 0) {
> + m = iter;
> + break;
> + }
> + }
This isn't a bug, but could the map loop here be replaced with
bpf_object__find_map_by_name(obj, secs[s]), as veristat's
set_global_vars() does? That would tie the map to the DATASEC by its
section name and make the comment about the two lookups not agreeing on a
size unnecessary.
> + id = btf__find_by_name_kind(btf, secs[s], BTF_KIND_DATASEC);
> + if (!m || id < 0)
> + continue;
> +
> + sec = btf__type_by_id(btf, id);
> + vsi = btf_var_secinfos(sec);
> + for (i = 0; i < btf_vlen(sec); i++, vsi++) {
> + const struct btf_type *var = btf__type_by_id(btf, vsi->type);
> +
> + if (strcmp(btf__name_by_offset(btf, var->name_off), name))
> + continue;
> + if (vsi->size != 4 && vsi->size != 8) {
> + PRINT_FAIL("'%s' is %u bytes, only 4 and 8 are supported\n",
> + name, vsi->size);
> + return -EINVAL;
> + }
> + vt = btf__type_by_id(btf, btf__resolve_type(btf, var->type));
> + if (!vt || !(btf_is_int(vt) || btf_is_any_enum(vt))) {
> + PRINT_FAIL("'%s' is not an int or an enum\n", name);
> + return -EINVAL;
> + }
> + *is_signed = global_var_is_signed(vt);
> + /*
> + * The map is found by the suffix of its name and the
> + * section by its own, so nothing so far has made the
> + * two agree on a size.
> + */
> + vsz = bpf_map__value_size(m);
> + if (vsi->offset > vsz || vsi->size > vsz - vsi->offset) {
> + PRINT_FAIL("'%s' at %u+%u is outside '%s' of %u bytes\n",
> + name, vsi->offset, vsi->size,
> + bpf_map__name(m), vsz);
> + return -EINVAL;
> + }
> + *map = m;
> + *off = vsi->offset;
> + *sz = vsi->size;
> + return 0;
> + }
> + }
> +
> + PRINT_FAIL("no global variable '%s'\n", name);
> + return -ENOENT;
> +}
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36502848904
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes
2026-09-29 0:17 ` [PATCH bpf-next v7 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
@ 2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 2:19 ` Yonghong Song
0 siblings, 1 reply; 46+ messages in thread
From: bot+bpf-ci @ 2026-09-29 0:52 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 8900 bytes --]
> diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
> index d1d40314035e..616b479238c7 100644
> --- a/tools/testing/selftests/bpf/exceptions_cleanup.h
> +++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
[ ... ]
> diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> index 255f88d35aad..c06ec10359b9 100644
> --- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> +++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> @@ -4,6 +4,7 @@
> #include "exceptions_cleanup.h"
> #include "exceptions_cleanup.skel.h"
> #include "exceptions_cleanup_fail.skel.h"
> +#include "exceptions_cleanup_shapes.skel.h"
>
> /* foo3 unwound: every frame that has a pad ran it. */
> #define PADS_FOO3_UNWOUND \
> @@ -82,4 +83,5 @@ void test_exceptions_cleanup(void)
> exceptions_cleanup__destroy(skel);
>
> RUN_TESTS(exceptions_cleanup_fail);
> + RUN_TESTS(exceptions_cleanup_shapes);
> }
A subsystem pattern flags this as potentially concerning: The only skip
gate in test_exceptions_cleanup() is the -EOPNOTSUPP check on the main
exceptions_cleanup skeleton, which tests bpf_jit_supports_cleanup_pads().
RUN_TESTS(exceptions_cleanup_shapes) then loads entry_callx_region, which
also needs indirect-call JIT support. progs/exceptions_cleanup_shapes.c
guards that program only on the compiler and arch:
/* gcc has no indirect calls, and only these JITs emit them */
#if defined(__clang__) && \
(defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64))
On x86 the two capabilities are independent:
arch/x86/net/bpf_jit_comp.c:
bool bpf_jit_supports_callx(void)
{
return cfi_mode != CFI_FINEIBT;
}
bool bpf_jit_supports_cleanup_pads(void)
{
return IS_ENABLED(CONFIG_UNWINDER_ORC);
}
With a kCFI kernel on IBT hardware (cfi=auto resolves to CFI_FINEIBT in
arch/x86/kernel/alternative.c), the main skeleton loads, so the skip
branch is not taken. The verifier's require_callx_jit() then refuses
entry_callx_region with -EOPNOTSUPP ("JIT doesn't support callx").
test_loader's run_subtest() treats any load error on a __success program
as ASSERT_OK(err, "unexpected_load_failure"), so this subtest fails
instead of skipping, even though the host simply lacks the feature.
Could the test check for callx support separately and skip on FineIBT
hosts? The existing callx selftests (RUN_JITED(verifier_callx) in
prog_tests/verifier.c and callx_func_ptr_map.c) do not skip on FineIBT
either; they check only is_jit_enabled(). So this shape follows the
existing convention, and the failure is limited to x86 FineIBT hosts
(clang kCFI build plus IBT-capable CPU).
> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
> new file mode 100644
> index 000000000000..1978fd1102f0
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
[ ... ]
> +/*
> + * A precision chain crossing a resume: r6 is kept across a call whose only
> + * way back is the callee's pad, then used as a variable stack offset.
> + */
> +static __used __naked __noinline __u64 prec_inner_frame(void)
> +{
> + asm volatile (
> +"1:" "call bpf_unwind;" /* cleanup region */
> +"2:"
> + "r0 = 0;"
> + "exit;"
> +"3:" /* landing pad */
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + ::: __clobber_all);
> +}
> +
> +static __used __naked __noinline __u64 prec_outer_frame(void)
> +{
> + asm volatile (
> + "r1 = %[input] ll;"
> + "r6 = *(u64 *)(r1 + 0);"
> + "r6 &= 0x7;"
> + "r0 = 0;"
> + "*(u64 *)(r10 - 8) = r0;"
> + "*(u64 *)(r10 - 16) = r0;"
> + "call prec_inner_frame;" /* comes back only through the pad */
> + "r2 = r10;"
> + "r2 += -16;"
> + "r2 += r6;" /* variable stack offset: r6 must be precise */
> + "*(u8 *)(r2 + 0) = 1;"
> + "r0 = 0;"
> + "exit;"
> + :
> + : __imm_addr(input)
> + : __clobber_all);
> +}
> +
> +SEC("?syscall")
> +__success __set_global(input, 101) __retval(0)
> +int entry_prec_across_resume(void *ctx)
> +{
> + return prec_outer_frame();
> +}
A subsystem pattern flags this as potentially concerning:
entry_prec_across_resume is meant to cover "a precision chain crossing a
resume" (per the commit message), but does the verifier actually check the
code after the call?
prec_inner_frame always unwinds into its own pad. The pad ends in
bpf_unwind_resume while curframe is 1. The resume handler in
do_check_insn() ends the path there:
kernel/bpf/verifier.c:
/*
* No need to walk into the caller: its pad was
* pushed as a branch at its call, and with no
* pad nothing of it runs.
*/
if (env->cur_state->curframe)
return PROCESS_BPF_EXIT;
prec_outer_frame has no CLEANUP_REC over "call prec_inner_frame", so no
pad branch is pushed there either. The fall-through of bpf_unwind in
prec_inner_frame is never walked, because process_bpf_unwind() jumps
straight to the pad. As a result, none of these instructions are ever
verified:
"r2 = r10;"
"r2 += -16;"
"r2 += r6;"
"*(u8 *)(r2 + 0) = 1;"
They never run either: bpf_unwind_rewrite() points prec_outer_frame's
uncovered return at the epilogue.
So the variable-offset stack write never runs mark_chain_precision() on
r6. The resume-crossing branch in kernel/bpf/backtrack.c (the
bt_subprog_enter() call made on a resume) is never exercised. Would the
test still pass if "r6 &= 0x7" were deleted, or if the backtrack.c resume
handling were broken?
Both the comment "comes back only through the pad" and the header "r6 is
kept across a call whose only way back is the callee's pad" appear
inconsistent with this: nothing comes back. Could this shape be adjusted
to cover the call in prec_outer_frame with a CLEANUP_REC and do the r6
variable-offset access in prec_outer_frame's own pad, so a pad path
verifies the use and precision is propagated back through the call?
> +/*
> + * The jump_into_pad shape with the branch statically dead, so only a
> + * speculative walk reaches the pad: a barrier rather than a refusal.
> + */
> +static __used __naked __noinline __u64 dead_jump_into_pad_frame(void)
> +{
> + asm volatile (
> + "r6 = 0;"
> + "if r6 > 7 goto 4f;" /* never taken: walked speculatively */
> +"1:" "call always_unwind;" /* cleanup region */
> +"2:"
> + "r0 = 0;"
> + "exit;"
> +"3:" /* landing pad */
> + "r7 = r0;"
> +"4:" /* ... and its second instruction */
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + ::: __clobber_all);
> +}
> +
> +SEC("?syscall")
> +__success
> +int dead_jump_into_pad(void *ctx)
> +{
> + return dead_jump_into_pad_frame();
> +}
A subsystem pattern flags this as potentially concerning:
dead_jump_into_pad is meant to cover "a pad reached only by a speculative
walk" (per the commit message), but does the verifier actually do a
speculative walk here?
The program has only __success, so test_loader loads it only in privileged
mode, as root with all capabilities. The verifier sets:
env->bypass_spec_v1 = bpf_bypass_spec_v1(env->prog->aux->token);
bpf_bypass_spec_v1() is true when bpf_token_capable(token, CAP_PERFMON)
holds, which root satisfies. In check_cond_jmp_op(), the dead branch of
"if r6 > 7" (pred == 0) is pushed only when bypass_spec_v1 is false:
kernel/bpf/verifier.c:
} else if (pred == 0) {
...
if (!env->bypass_spec_v1) {
err = sanitize_speculative_path(env, insn, *insn_idx + insn->off + 1,
*insn_idx);
So label 4 is never walked. bpf_exc_check_insn() never reports "runs both
inside and outside a landing pad" on a speculative state, and the
error_recoverable_with_nospec() && state->speculative recovery in
do_check() never runs. The program loads because the verifier ignores the
dead branch entirely, not because of the barrier logic. Would the test
still pass unchanged if that recovery were removed and the error made
fatal?
The comment "never taken: walked speculatively" and the header "only a
speculative walk reaches the pad: a barrier rather than a refusal" describe
something that does not happen under the test's own settings.
Could this also run the program unprivileged with CAP_BPF kept, so it
loads but CAP_PERFMON is gone and bypass_spec_v1 is false? A syscall
program is not a perfmon program type, so CAP_BPF is enough to load it:
SEC("?syscall")
__success __caps_unpriv(CAP_BPF) __success_unpriv
int dead_jump_into_pad(void *ctx)
It could also check the verifier log for the nospec insertion.
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36502848904
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests
2026-09-29 0:17 ` [PATCH bpf-next v7 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
@ 2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:42 ` Yonghong Song
0 siblings, 1 reply; 46+ messages in thread
From: bot+bpf-ci @ 2026-09-29 0:52 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 4746 bytes --]
> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
> new file mode 100644
> index 000000000000..76cca8dc45a9
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
> @@ -0,0 +1,878 @@
[ ... ]
> +static __used __noinline void __kfunc_btf_anchor(void)
> +{
> + bpf_throw(0);
> + bpf_unwind();
> + bpf_preempt_disable();
> + bpf_preempt_enable();
> + bpf_rcu_read_lock();
> + bpf_rcu_read_unlock();
> + bpf_unwind_resume(NULL);
> +}
> +
> +/* An unwind raised in a callee, which is how a cleanup region gets one. */
> +static __used __naked __noinline __u64 inner_unwind(void)
> +{
> + asm volatile (
> + "r1 = 1;"
> + "call bpf_unwind;"
> + "r0 = 0;"
> + "exit;"
> + ::: __clobber_all);
> +}
[ ... ]
> +static __used __naked __noinline __u64 unwind_in_pad_frame(void)
> +{
> + asm volatile (
> + "call bpf_preempt_disable;"
> +"1:" "call inner_unwind;" /* cleanup region */
> +"2:"
> + "call bpf_preempt_enable;"
> + "r0 = 0;"
> + "exit;"
> +"3:" /* landing pad that unwinds again */
> + "call bpf_preempt_enable;"
> + "r1 = 2;"
> + "call bpf_unwind;"
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + ::: __clobber_all);
> +}
[ ... ]
> +static __used __naked __noinline __u64 cb_and_table_frame(void)
> +{
> + asm volatile (
> + "call bpf_preempt_disable;"
> + "r1 = 9;"
> +"1:" "call bpf_unwind;" /* cleanup region */
> +"2:"
> + "r0 = 0;"
> + "exit;"
> +"3:" /* landing pad */
> + "call bpf_preempt_enable;"
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + ::: __clobber_all);
> +}
[ ... ]
> +static __used __naked __noinline __u64 pad_calls_thrower_frame(void)
> +{
> + asm volatile (
> + "call bpf_preempt_disable;"
> + "r1 = 11;"
> +"1:" "call bpf_unwind;" /* cleanup region */
> +"2:"
> + "r0 = 0;"
> + "exit;"
> +"3:" /* landing pad */
> + "call pad_callee_that_throws;" /* ...which can throw: refused */
> + "call bpf_preempt_enable;"
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + ::: __clobber_all);
> +}
[ ... ]
> +static __used __naked __noinline __u64 catch_pad_frame(void)
> +{
> + asm volatile (
> + "call bpf_preempt_disable;"
> + "r1 = 12;"
> +"1:" "call bpf_unwind;" /* cleanup region */
> +"2:"
> + "r0 = 0;"
> + "exit;"
> +"3:" /* catch pad: no resume, it stops here */
> + "call bpf_preempt_enable;"
> + "r0 = 0;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + ::: __clobber_all);
> +}
[ ... ]
> +static __used __naked __noinline __u64 stray_resume_frame(void)
> +{
> + asm volatile (
> + "call bpf_preempt_disable;"
> + "r1 = 13;"
> +"1:" "call bpf_unwind;" /* cleanup region */
> +"2:"
> + "r0 = 0;"
> + "exit;"
> +"3:" /* landing pad */
> + "call bpf_preempt_enable;"
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + ::: __clobber_all);
> +}
[ ... ]
> +static __used __naked __noinline __u64 pad_calls_resumer_frame(void)
> +{
> + asm volatile (
> + "call bpf_preempt_disable;"
> + "r1 = 14;"
> +"1:" "call bpf_unwind;" /* cleanup region */
> +"2:"
> + "r0 = 0;"
> + "exit;"
> +"3:" /* landing pad */
> + "call bpf_preempt_enable;"
> + "call resume_in_callee;" /* ...which resumes: refused */
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + ::: __clobber_all);
> +}
[ ... ]
> +static __used __naked __noinline void gotox_unwinder(void)
> +{
> + asm volatile (
> + "r1 = 15;"
> + "call bpf_unwind;"
> + "exit;"
> + ::: __clobber_all);
> +}
This isn't a bug, but bpf_unwind() is declared as __bpf_kfunc void
bpf_unwind(void) in kernel/bpf/helpers.c and takes no arguments.
Several of these inline asm frames load r1 with a numbered value right
before calling bpf_unwind: r1=1 in inner_unwind, r1=2 in
unwind_in_pad_frame, r1=9 in cb_and_table_frame, r1=11 in
pad_calls_thrower_frame, r1=12 in catch_pad_frame, r1=13 in
stray_resume_frame, r1=14 in pad_calls_resumer_frame, and r1=15 in
gotox_unwinder.
Other bpf_unwind call sites in the same file like pad_drops_caller_lock_frame,
no_pad_keeps_own_lock_frame, and pad_drops_caller_ref_frame call bpf_unwind
without setting r1 first. The verifier checks no arguments for a zero-argument
kfunc, and neither bpf_exc_keep_exits nor the runtime walk uses r1.
Are the "r1 = N;" loads before these calls left over from an earlier
cookie-style signature? Could they be dropped so the asm matches the kfunc
prototype?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36502848904
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with
2026-09-29 0:16 ` [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
2026-09-29 0:36 ` sashiko-bot
@ 2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:10 ` Yonghong Song
1 sibling, 1 reply; 46+ messages in thread
From: bot+bpf-ci @ 2026-09-29 0:52 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 1594 bytes --]
> A landing pad's entry state is the state at the call it belongs to, taken
> before the callee ran -- push_cleanup_pad_branch() snapshots it there. That
> is right for the frame's own registers and stack, which the callee's
> epilogue puts back, and wrong for what the program shares: nothing restores
> the locks in bpf_verifier_state.
Is the claim about the callee's epilogue restoring the caller's stack
accurate?
Looking at the BPF calling convention, the callee's epilogue only restores
r6-r9 and the frame pointer. It does not restore the caller's BPF stack
slots. A callee that receives a PTR_TO_STACK into the caller's frame can
write that stack freely, and those writes are not restored on return.
For example, if the caller spills a map value pointer to fp-8 and calls
foo(fp-8) under a landing pad, foo can store an arbitrary scalar to
*(u64 *)(r1 + 0) and then call bpf_unwind(). The caller's pad would be
verified from the pre-call snapshot where fp-8 still holds the spilled
map value pointer, but at run time the pad would load the
attacker-controlled scalar.
The same issue applies to any stack slot type change that the callee makes
before unwinding, such as spilled pointers overwritten by stack writes, or
dynptr and iterator slots clobbered by the callee.
Should the commit message clarify what the epilogue actually restores?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36502848904
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 13/22] bpf, arm64: Dispatch exception cleanup pads at run time
2026-09-29 0:17 ` [PATCH bpf-next v7 13/22] bpf, arm64: " Yonghong Song
@ 2026-09-29 1:14 ` bot+bpf-ci
0 siblings, 0 replies; 46+ messages in thread
From: bot+bpf-ci @ 2026-09-29 1:14 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3177 bytes --]
> diff --git a/arch/arm64/kernel/stacktrace.c b/arch/arm64/kernel/stacktrace.c
> index 3ebcf8c53fb0..1e46a22cafbd 100644
> --- a/arch/arm64/kernel/stacktrace.c
> +++ b/arch/arm64/kernel/stacktrace.c
> @@ -445,6 +445,109 @@ noinline noinstr void arch_bpf_stack_walk(bool (*consume_entry)(void *cookie, u6
> kunwind_stack_walk(arch_bpf_unwind_consume_entry, &data, current, NULL);
> }
>
> +struct bpf_unwind_ra_consume_entry_data {
> + bool (*consume_entry)(void *cookie, u64 ip, u64 sp, u64 fp, u64 *ra);
> + void *cookie;
> + unsigned long record;
> + bool seen_first;
> +};
> +
> +static u64 bpf_unwind_sign_ra(u64 ra, u64 modifier)
> +{
> + asm volatile(ARM64_ASM_PREAMBLE
> + ".arch_extension pauth\n"
> + " pacia %0, %1"
> + : "+r" (ra) : "r" (modifier));
> + return ra;
> +}
> +
> +/*
> + * PACIASP's modifier is the stack pointer the owner was entered with: record
> + * + 16 for a BPF prologue, but further up for bpf_unwind()'s own C frame.
> + * Recognise it by re-signing @pc, which the unwinder stripped from @stored.
> + */
> +static bool bpf_unwind_ra_modifier(unsigned long record, unsigned long caller_fp,
> + u64 stored, u64 pc, u64 *modifier)
> +{
> + u64 m;
> +
> + for (m = record + sizeof(struct frame_record); m <= caller_fp; m += 16) {
> + if (bpf_unwind_sign_ra(pc, m) == stored) {
> + *modifier = m;
> + return true;
> + }
> + }
> + return false;
> +}
Can this loop find the wrong modifier for bpf_unwind()'s own frame?
The search takes the first candidate whose PAC matches. With
CONFIG_ARM64_PTR_AUTH_KERNEL, the PAC field is quite short: roughly 15 bits
with 48-bit VA, 11 bits with 52-bit VA, and only 7 bits when TBI is enabled
for KASAN tag-based modes.
For JIT frames, the true modifier is always record+16 (the prologue does
paciasp then 'stp x29, lr, [sp, #-16]!'), so they always match correctly on
the first try.
But when bpf_unwind()'s own C frame is being rewritten (which happens when
the BPF frame that called bpf_unwind() has a pad covering the call), the
compiler might place the frame record at a different offset. For example,
with the traditional GCC aarch64 layout that puts locals above the frame
record, and bpf_unwind() has a stack local (ctx) whose address is taken, the
entry SP would be record+32 rather than record+16.
The loop would then try record+16 first. If PAC(pc, record+16) happens to
equal PAC(pc, record+32) for the current task's APIA key, it returns the
wrong modifier. With a 7-bit PAC (KASAN tag-based), that's roughly a 1/128
chance per key.
When autiasp runs in bpf_unwind()'s epilogue, the authentication would fail.
With FEAT_FPAC this faults immediately; otherwise ret faults on the poisoned
LR.
The code has no ambiguity check and never derives the frame's actual entry
SP. Could the search either verify there's exactly one matching candidate,
or derive bpf_unwind()'s true entry SP directly rather than guessing?
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36502848904
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 11/22] bpf: Dispatch cleanup pads by rewriting return addresses
2026-09-29 0:16 ` [PATCH bpf-next v7 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
@ 2026-09-29 1:14 ` bot+bpf-ci
2026-09-30 1:18 ` Yonghong Song
0 siblings, 1 reply; 46+ messages in thread
From: bot+bpf-ci @ 2026-09-29 1:14 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 2093 bytes --]
> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
> index 5b7fe4ba610b..bff9d2539371 100644
> --- a/kernel/bpf/fixups.c
> +++ b/kernel/bpf/fixups.c
> @@ -1,5 +1,6 @@
> // SPDX-License-Identifier: GPL-2.0-only
> /* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
> +#include <linux/bitmap.h>
This isn't a bug, but is <linux/bitmap.h> needed? No bitmap API appears
in the file.
[ ... ]
> @@ -1286,6 +1309,61 @@ static int resolve_func_ptrs(struct bpf_verifier_env *env)
> return 0;
> }
>
> +static int exc_info_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
> + u32 subprog, u32 start, u32 end)
> +{
> + struct bpf_cleanup_info *recs;
> + u32 i, cnt = 0;
> + int err;
> +
> + if (!env->cleanup_info_cnt)
> + return 0;
> +
> + err = bpf_exc_alloc_info(sub->aux);
> + if (err)
> + return err;
> +
> + for (i = start; i < end; i++) {
> + if (env->insn_aux_data[i].cleanup_pad)
> + cnt++;
> + }
> + if (!cnt)
> + return 0;
> +
> + recs = kvmalloc_array(cnt, sizeof(*recs), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
> + if (!recs)
> + return -ENOMEM;
> +
> + for (i = start, cnt = 0; i < end; i++) {
> + u32 pad = env->insn_aux_data[i].cleanup_pad;
> +
> + if (!pad)
> + continue;
> + pad--;
> + if (verifier_bug_if(pad < start || pad >= end, env,
> + "insn %u is covered by a landing pad at %u outside its subprog [%u, %u)",
> + i, pad, start, end)) {
> + kvfree(recs);
> + return -EFAULT;
> + }
> + recs[cnt].begin_off = i - start;
> + recs[cnt].end_off = i - start + 1;
> + recs[cnt].landing_pad_off = pad - start;
> + cnt++;
> + }
> + err = bpf_exc_attach_info(sub->aux, recs, cnt);
> + if (err)
> + return err;
> + return 0;
> +}
This isn't a bug, but is the subprog parameter meant to be used? The
function body uses only start and end, yet both callers pass a value for
it.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36502848904
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 04/22] bpf: Add lookups for exception cleanup resumes and landing pads
2026-09-29 0:33 ` sashiko-bot
@ 2026-09-29 21:58 ` Yonghong Song
0 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 21:58 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/28/26 5:33 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Medium] bpf: missing cleanup_pad adjustment in BPF_PATCH_MOVE_TARGET
> --
>
> commit ddeb49ab1c3e86d5fea9b3b127432b6e34db617b
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Add lookups for exception cleanup resumes and landing pads
>
> This commit introduces helpers to recognize exception cleanup resumes and
> landing pads. It also adds logic to maintain cleanup_pad data in
> insn_aux_data when instructions are patched or removed.
>
>> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
>> index 37cf130ebb57b..5b7fe4ba610b3 100644
>> --- a/kernel/bpf/fixups.c
>> +++ b/kernel/bpf/fixups.c
>> @@ -268,11 +268,18 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
>> data[i].non_stack_access =
>> data[off + cnt - 1].non_stack_access;
>> data[off + cnt - 1].non_stack_access = false;
>> + data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
>> + data[off + cnt - 1].cleanup_pad = 0;
>> } else if (bpf_is_mem_insn(insn + i)) {
>> data[i].non_stack_access = true;
>> }
>> }
>>
>> + if (env->cleanup_info_cnt)
>> + for (i = 0; i < prog_len; i++)
>> + if (data[i].cleanup_pad > off + 1)
>> + data[i].cleanup_pad += cnt - 1;
> [Severity: Medium]
> When mode is BPF_PATCH_MOVE_TARGET and the landing pad is located exactly
> at the patched instruction, does this condition fail to adjust it?
>
> Since cleanup_pad is 1-based, a value of off + 1 means the landing pad is
> exactly at off. Because this strictly checks > off + 1, it skips adjusting
> data[i].cleanup_pad when it equals off + 1.
>
> If BPF_PATCH_MOVE_TARGET is used (for example, with gen_prologue or may_goto
> initialization), the original instruction is moved to the end of the patch
> sequence. This means the landing pad should follow it and move to the end
> as well, but it will be left pointing to the start of the inserted sequence.
>
> Could this result in the unwinder jumping to the newly inserted prologue or
> loop initialization instead of the actual exception landing pad?
Okay, commit "bpf: Retarget indirect jump targets across prologue prepends"
introduced BPF_PATCH_MOVE_TARGET and merged last Friday. It solved three
cases for ops->gen_epilogue, ops->gen_prologue || env->seen_direct_write, and
stack slots for subprogs. ops->gen_epilogue has been rejected in patch 5.
We cannot allow ops->gen_epilogue since it may silently exit.
For other cases in "bpf: Retarget indirect jump targets across prologue prepends",
The above commit should already handle this.
For the other two, cleanup_pad == off + 1 cannot happen, because MOVE only
patches an entry insn and a landing pad cannot start at insn 0: the entry is
always walked outside a pad first, so reaching it again from an unwind
fails bpf_exc_check_insn() with "insn %u runs both inside and outside a
landing pad". Pads after off are shifted by the existing `> off + 1`, and a
covered call at off moves with its insn_aux_data.
>
>> +
>> /*
>> * Last slot instruction could be a newly generated
>> * BPF_ST/BPF_LDX/BPF_STX, systematically mark it for non-stack access
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 05/22] bpf: Prepare for an exception cleanup table before the CFG walk
2026-09-29 0:31 ` sashiko-bot
@ 2026-09-29 22:04 ` Yonghong Song
0 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-29 22:04 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/28/26 5:31 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf: cleanup_pad adjustment fails for BPF_PATCH_MOVE_TARGET at off=0
> --
>
> commit fc0420a45de5f3cee47a65d96e78244ace4fbf4e
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Prepare for an exception cleanup table before the CFG walk
>
> This patch prepares an exception cleanup table for BPF programs, adding
> landing pads to the instruction auxiliary data for calls that can unwind.
> It introduces the cleanup_pad field to track where execution should resume
> during exception unwinding.
>
>> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
>> index b19fcbf49b7ef..c1779d2d02f05 100644
>> --- a/kernel/bpf/exception.c
>> +++ b/kernel/bpf/exception.c
> [ ... ]
>
>> +static int mark_call_sites(struct bpf_verifier_env *env)
>> +{
>> + u32 i, j;
>> +
>> + for (i = 0; i < env->cleanup_info_cnt; i++) {
>> + struct bpf_cleanup_info *rec = &env->cleanup_info[i];
>> +
>> + for (j = rec->begin_off; j < rec->end_off; j++) {
>> + struct bpf_insn *insn = &env->prog->insnsi[j];
>> +
>> + if (!bpf_pseudo_call(insn) && !bpf_is_callx(insn) &&
>> + !bpf_is_unwind_kfunc(insn))
>> + continue;
>> + env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
> [Severity: High]
> How does this new cleanup_pad field handle instruction patching when a
> landing pad is located at instruction 0?
>
> If mark_call_sites() sets cleanup_pad = rec->landing_pad_off + 1, and the
> landing pad is at index 0, cleanup_pad becomes 1.
>
> If bpf_convert_ctx_accesses() in kernel/bpf/fixups.c later inserts a
> prologue at instruction 0 using BPF_PATCH_MOVE_TARGET,
> adjust_insn_aux_data() will try to update the aux data offsets. However,
> its check:
>
> if (data[i].cleanup_pad > off + 1)
> data[i].cleanup_pad += cnt - 1;
>
> will evaluate to 1 > 1 (false) when off=0. The required shift is skipped,
> leaving the cleanup_pad pointing to the newly inserted prologue instead of
> shifting to the original instruction. Unwinding to this pad could re-execute
> the prologue and potentially corrupt the frame pointer or stack.
>
> Can this sequence occur and lead to stack issues?
The landing pad cannot be at insn 0. The prog entry starts with non-pad and
until it finds bpf_unwind() and goes to landing pad. In such cases, you will
see landing_pad insn with non-landing-pad and with-landing-pad. This will
be rejected.
>
>> + }
>> + }
>> + return 0;
>> +}
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 07/22] bpf: Resume a covered call at its landing pad
2026-09-29 0:31 ` sashiko-bot
@ 2026-09-30 0:28 ` Yonghong Song
0 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-30 0:28 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/28/26 5:31 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Critical] bpf: landing pad state desync via discarded memory mutations allows OOB access
> --
>
> commit 41b3c5e14924fae7175095b7012838a14c2e5ffc
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Resume a covered call at its landing pad
>
> This commit allows a covered call to resume at its landing pad when bpf_unwind()
> is invoked by a subprogram. It pushes a cleanup pad branch with the caller's
> pre-call state and handles bpf_unwind() by returning PROCESS_BPF_EXIT for
> nested frames.
>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index fc3df452de2e4..ee074d4a936b5 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
> [ ... ]
>
>> @@ -19167,6 +19186,57 @@ enum {
>> INSN_IDX_UPDATED = 2,
>> };
>>
>> +static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
>> +{
>> + struct bpf_verifier_state *branch;
>> + struct bpf_func_state *frame;
>> + int pad = bpf_exc_pad_of_call(env, insn_idx);
>> +
>> + if (pad < 0)
>> + return 0;
>> + branch = push_stack(env, pad, insn_idx, false);
>> + if (IS_ERR(branch))
>> + return PTR_ERR(branch);
>> + frame = branch->frame[branch->curframe];
>> + /*
>> + * The state at that call with the caller-saved registers gone: the
>> + * callee's epilogue put r6-r9 and the stack back on the way out.
>> + */
>> + clear_caller_saved_regs(env, frame->regs);
>> + mark_reg_unknown(env, frame->regs, BPF_REG_0);
>> + return 0;
>> +}
>> +
>> +static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
>> + bool *do_print_state)
>> +{
>> + struct bpf_func_state *frame = cur_func(env);
>> + int pad = bpf_exc_pad_of_call(env, *insn_idx);
>> + int err;
>> +
>> + if (pad < 0) {
>> + err = check_resource_leak(env, false, !env->cur_state->curframe,
>> + "an unwind with no landing pad");
>> + if (err)
>> + return err;
>> + if (env->cur_state->curframe)
>> + return PROCESS_BPF_EXIT;
> [Severity: Critical]
> When process_bpf_unwind() returns PROCESS_BPF_EXIT here, the verifier stops
> exploring the callee and discards its state. Does this discard any memory
> mutations made by the callee to pointer arguments before calling bpf_unwind()?
You are right. It is my mistake. The key problem is here (in v7):
+static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
+{
+ struct bpf_verifier_state *branch;
+ struct bpf_func_state *frame;
+ int pad = bpf_exc_pad_of_call(env, insn_idx);
+
+ if (pad < 0)
+ return 0;
+ branch = push_stack(env, pad, insn_idx, false);
+ if (IS_ERR(branch))
+ return PTR_ERR(branch);
+ frame = branch->frame[branch->curframe];
+ /*
+ * The state at that call with the caller-saved registers gone: the
+ * callee's epilogue put r6-r9 and the stack back on the way out.
+ */
+ clear_caller_saved_regs(env, frame->regs);
+ mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ return 0;
+}
esp.
branch = push_stack(env, pad, insn_idx, false);
it ignored the state change in callee, e.g., a value in stack
may get changed. This will make verification incorrect due to
such changed value.
>
> Since the CPU unwinds the stack without undoing memory writes at runtime,
> could this desynchronize the verifier state from runtime state and allow
> out-of-bounds memory accesses?
Yes, see the above.
>
>> + /*
>> + * The main program's frame returns at once, which is the
>> + * program returning. Mark r0 the zero the fixups leave after
>> + * the call, and leave through the exit, which is what holds
>> + * that zero to the program type.
>> + */
>> + mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
>> + mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
>> + return process_bpf_exit_full(env, do_print_state, false);
>> + }
>> + clear_caller_saved_regs(env, frame->regs);
>> + mark_reg_unknown(env, frame->regs, BPF_REG_0);
>> + *insn_idx = pad;
>> + return INSN_IDX_UPDATED;
>> +}
>> +
> [ ... ]
>
>> @@ -19421,7 +19491,29 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>> return -EINVAL;
>> }
>> }
>> + if (bpf_is_unwind_kfunc(insn) || bpf_is_unwind_resume_kfunc(insn)) {
>> + err = check_kfunc_allowed_only(env, insn, env->insn_idx);
>> + if (err)
>> + return err;
>> + if (bpf_is_unwind_kfunc(insn))
>> + return process_bpf_unwind(env, &env->insn_idx,
>> + do_print_state);
>> + /*
>> + * Mark r0 a known zero -- unknown first, as
>> + * the known-zero helper keeps the type it
>> + * finds, which here is NOT_INIT. The fixups
>> + * lower this to 'r0 = 0; exit', so the frame
>> + * returns a real zero.
>> + */
>> + mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
>> + mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
>> + return process_bpf_exit_full(env, do_print_state, false);
>> + }
>> mark_reg_scratched(env, BPF_REG_0);
>> + /* An unwind out of this call resumes at the pad. */
>> + err = push_cleanup_pad_branch(env, env->insn_idx);
>> + if (err)
>> + return err;
> [Severity: Critical]
> push_cleanup_pad_branch() pushes a branch using a snapshot of the pre-call
> state. If the callee modified a stack variable (e.g., an array index) and then
> unwound the stack, does the verifier evaluate the landing pad with the original
> unmodified value?
>
> Could a BPF program use this to bypass bounds checks by mutating an index in
> the callee before unwinding, causing the verifier to optimize out the checks
> in the landing pad while using the out-of-bounds value at runtime?
You are right. See the above explanation. Will fix.
>
>> if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
>> cur_func(env)->no_stack_arg_load = true;
>> if (bpf_is_callx(insn))
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with
2026-09-29 0:36 ` sashiko-bot
@ 2026-09-30 1:09 ` Yonghong Song
0 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-30 1:09 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/28/26 5:36 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
> - [Critical] bpf: type confusion and memory corruption due to stale cleanup pad state
> - [High] bpf: lock corruption and double-unlock via missing identity check on unwind
> --
>
> commit 585aa50d9e8811792b9a816a4d1ad6cd51d63ea7
> Author: Yonghong Song <yonghong.song@linux.dev>
> Subject: bpf: Require an unwind to leave a frame holding what it entered with
>
> This commit updates the BPF verifier to record what locks and references
> a program holds when a frame is entered. It enforces that an unwind must
> leave the frame holding what it entered with, ensuring that landing pads
> verified against pre-call snapshots remain valid.
>
>> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
>> index 0143688896b0e..75e572a8a1bab 100644
>> --- a/include/linux/bpf_verifier.h
>> +++ b/include/linux/bpf_verifier.h
>> @@ -339,6 +339,18 @@ struct bpf_func_state {
>> bool in_async_callback_fn;
>> bool in_exception_callback_fn;
>> bool no_stack_arg_load;
>> + /*
>> + * What the program held when this frame was entered. An unwind leaves
>> + * the frame without running anything below it, so the frame has to put
>> + * these back to what it found before it goes -- otherwise a caller's
>> + * landing pad, whose state was taken at the call, is wrong about them.
>> + */
>> + u32 entry_active_locks;
>> + u32 entry_preempt_locks;
>> + u32 entry_rcu_locks;
>> + u32 entry_irq_id;
>> + u32 entry_id_gen;
>> + u32 entry_acquired_refs;
> [Severity: High]
> Does this structure need fields to record the identity of the lock held upon
> frame entry? Without this, it might be difficult to accurately enforce the
> correct lock state on unwind.
Assuming the bug in patch 7 fixed.
Verifier flow should match lock/unlock with the same identity. Otherwise
verification will fail. I think we are okay here.
>
> Additionally, is it necessary to update func_states_equal() to compare these
> new entry_* fields? Without comparing them, could the verifier unsafely prune
> states with different entry locking conditions?
I think we are fine here. There are already some lock checking in refsafe(),
assuming patch 7 bug fixed.
>
> [ ... ]
>> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
>> index c1779d2d02f05..c0b0b8478af51 100644
>> --- a/kernel/bpf/exception.c
>> +++ b/kernel/bpf/exception.c
>> @@ -12,6 +12,72 @@
>> BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
>> BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
>>
>> +void bpf_exc_record_frame_entry(const struct bpf_verifier_state *state,
>> + struct bpf_func_state *frame, u32 id_gen)
>> +{
> [ ... ]
>> +int bpf_exc_check_frame_balance(struct bpf_verifier_env *env, const char *prefix)
>> +{
>> + const struct bpf_verifier_state *state = env->cur_state;
>> + const struct bpf_func_state *frame = cur_func(env);
>> + u32 i, held;
>> + const char *what;
>> +
>> + if (state->active_rcu_locks != frame->entry_rcu_locks)
>> + what = "bpf_rcu_read_lock";
>> + else if (state->active_preempt_locks != frame->entry_preempt_locks)
>> + what = "bpf_preempt_disable";
>> + else if (state->active_irq_id != frame->entry_irq_id)
>> + what = "bpf_local_irq_save";
>> + else if (state->active_locks != frame->entry_active_locks)
>> + what = "bpf_spin_lock";
> [Severity: High]
> Does this code adequately verify lock identity? While it checks the count of
> active locks, it appears to miss checking the identity of the lock
> (active_lock_id and active_lock_ptr).
>
> If a BPF program acquires lock A, passes it to a subprogram which unlocks A,
> locks B, and then throws an exception (bpf_unwind()), could the unwinder
> accept the frame balance since the active_locks count is still 1?
>
> The exception would jump back to the caller's cleanup pad, which assumes
> lock A is still held. The cleanup pad unlocking A could potentially cause a
> double-unlock of lock A and leak lock B forever.
This should be okay once v7, patch 7 fixed. That will be able to
avoid double unlock.
>
>> + else
>> + what = NULL;
> [ ... ]
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index ee074d4a936b5..4bdee3f02fe9d 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
> [ ... ]
>> @@ -19512,6 +19554,10 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>> mark_reg_scratched(env, BPF_REG_0);
>> /* An unwind out of this call resumes at the pad. */
>> err = push_cleanup_pad_branch(env, env->insn_idx);
> [Severity: Critical]
> Does taking a pre-call snapshot of the verifier state for the cleanup pad
> incorrectly ignore valid memory modifications made by the callee before
> unwinding?
>
> For example, if a caller spills a valid pointer to its stack and passes a
> pointer to this stack slot to a subprogram:
>
> - The verifier schedules the caller's cleanup pad branch by pushing a snapshot
> of the verifier state (including the stack) taken BEFORE the callee runs.
> - The subprogram overwrites the caller's stack slot with an attacker-controlled
> scalar and calls bpf_unwind().
> - At runtime, the exception jumps to the cleanup pad, which loads the scalar.
>
> Because the cleanup pad uses the pre-call snapshot, does the verifier
> incorrectly assume the loaded value is still a valid pointer, potentially
> leading to type confusion and memory corruption?
You are right. See comments in patch 7. Will fix.
>
>> + if (err)
>> + return err;
>> + /* Or, with no pad, leaves the frame for good. */
>> + err = check_unwind_through_call(env, env->insn_idx);
>> if (err)
>> return err;
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with
2026-09-29 0:52 ` bot+bpf-ci
@ 2026-09-30 1:10 ` Yonghong Song
0 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-30 1:10 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/28/26 5:52 PM, bot+bpf-ci@kernel.org wrote:
>> A landing pad's entry state is the state at the call it belongs to, taken
>> before the callee ran -- push_cleanup_pad_branch() snapshots it there. That
>> is right for the frame's own registers and stack, which the callee's
>> epilogue puts back, and wrong for what the program shares: nothing restores
>> the locks in bpf_verifier_state.
> Is the claim about the callee's epilogue restoring the caller's stack
> accurate?
>
> Looking at the BPF calling convention, the callee's epilogue only restores
> r6-r9 and the frame pointer. It does not restore the caller's BPF stack
> slots. A callee that receives a PTR_TO_STACK into the caller's frame can
> write that stack freely, and those writes are not restored on return.
>
> For example, if the caller spills a map value pointer to fp-8 and calls
> foo(fp-8) under a landing pad, foo can store an arbitrary scalar to
> *(u64 *)(r1 + 0) and then call bpf_unwind(). The caller's pad would be
> verified from the pre-call snapshot where fp-8 still holds the spilled
> map value pointer, but at run time the pad would load the
> attacker-controlled scalar.
>
> The same issue applies to any stack slot type change that the callee makes
> before unwinding, such as spilled pointers overwritten by stack writes, or
> dynptr and iterator slots clobbered by the callee.
>
> Should the commit message clarify what the epilogue actually restores?
Yes, this is a problem. I had some comments in patch 7. Will fix.
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36502848904
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 11/22] bpf: Dispatch cleanup pads by rewriting return addresses
2026-09-29 1:14 ` bot+bpf-ci
@ 2026-09-30 1:18 ` Yonghong Song
0 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-30 1:18 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/28/26 6:14 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
>> index 5b7fe4ba610b..bff9d2539371 100644
>> --- a/kernel/bpf/fixups.c
>> +++ b/kernel/bpf/fixups.c
>> @@ -1,5 +1,6 @@
>> // SPDX-License-Identifier: GPL-2.0-only
>> /* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
>> +#include <linux/bitmap.h>
> This isn't a bug, but is <linux/bitmap.h> needed? No bitmap API appears
> in the file.
Thanks. Will remove.
> [ ... ]
>
>> @@ -1286,6 +1309,61 @@ static int resolve_func_ptrs(struct bpf_verifier_env *env)
>> return 0;
>> }
>>
>> +static int exc_info_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
>> + u32 subprog, u32 start, u32 end)
>> +{
>> + struct bpf_cleanup_info *recs;
>> + u32 i, cnt = 0;
>> + int err;
>> +
>> + if (!env->cleanup_info_cnt)
>> + return 0;
>> +
>> + err = bpf_exc_alloc_info(sub->aux);
>> + if (err)
>> + return err;
>> +
>> + for (i = start; i < end; i++) {
>> + if (env->insn_aux_data[i].cleanup_pad)
>> + cnt++;
>> + }
>> + if (!cnt)
>> + return 0;
>> +
>> + recs = kvmalloc_array(cnt, sizeof(*recs), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
>> + if (!recs)
>> + return -ENOMEM;
>> +
>> + for (i = start, cnt = 0; i < end; i++) {
>> + u32 pad = env->insn_aux_data[i].cleanup_pad;
>> +
>> + if (!pad)
>> + continue;
>> + pad--;
>> + if (verifier_bug_if(pad < start || pad >= end, env,
>> + "insn %u is covered by a landing pad at %u outside its subprog [%u, %u)",
>> + i, pad, start, end)) {
>> + kvfree(recs);
>> + return -EFAULT;
>> + }
>> + recs[cnt].begin_off = i - start;
>> + recs[cnt].end_off = i - start + 1;
>> + recs[cnt].landing_pad_off = pad - start;
>> + cnt++;
>> + }
>> + err = bpf_exc_attach_info(sub->aux, recs, cnt);
>> + if (err)
>> + return err;
>> + return 0;
>> +}
> This isn't a bug, but is the subprog parameter meant to be used? The
> function body uses only start and end, yet both callers pass a value for
> it.
Will remove unused 'subprog' parameter.
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 12/22] bpf, x86: Dispatch exception cleanup pads at run time
2026-09-29 0:30 ` sashiko-bot
@ 2026-09-30 1:34 ` Yonghong Song
0 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-30 1:34 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/28/26 5:30 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf: exception unwinding aborted when bpf_unwind is traced via kretprobe
> --
>
> commit f3ba2238caee72a50c097a775a0290b3b66ab98b
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf, x86: Dispatch exception cleanup pads at run time
>
> This commit updates the BPF JIT on x86 to dispatch exception cleanup pads at
> run time by writing directly to the return-address slot. It introduces
> arch_bpf_stack_walk_ra() to support unwinding alongside the ORC unwinder.
>
>> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
>> index 6c7a0578760e7..d4feade5b5c72 100644
>> --- a/arch/x86/net/bpf_jit_comp.c
>> +++ b/arch/x86/net/bpf_jit_comp.c
> [ ... ]
>
>> @@ -4614,6 +4628,34 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
>> #endif
>> }
>>
>> +void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
>> + void *cookie)
>> +{
>> +#if defined(CONFIG_UNWINDER_ORC)
>> + struct unwind_state state;
>> + unsigned long addr, *ra;
>> +
>> + for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
>> + unwind_next_frame(&state)) {
>> + addr = unwind_get_return_address(&state);
>> + ra = unwind_get_return_address_ptr(&state);
>> + if (!addr || !ra)
>> + break;
>> + /*
>> + * A traced return: the unwinder recovered @addr from under a
>> + * function graph or kretprobe trampoline, which is what the
>> + * slot itself still holds. Writing there would skip the
>> + * trampoline and leave its entry for the next hooked return
>> + * to pop.
>> + */
>> + if (READ_ONCE_NOCHECK(*ra) != addr)
>> + break;
> [Severity: High]
> Will this unconditionally abort the unwinder if a kernel function like
> bpf_unwind() is traced via a kretprobe or function graph tracer?
>
> When a BPF program triggers an exception, bpf_unwind() invokes
> arch_bpf_stack_walk_ra(). The unwinder processes kernel frames first. If
> a user has legitimately attached a tracer to bpf_unwind(), the stack return
> address slot (*ra) for that frame will hold the tracer's trampoline, while
> the unwinder recovers the real return address (addr).
>
> Because this evaluates to true and breaks the loop, it seems the walk stops
> completely before it even reaches the BPF frames. Could this result in a
> silent failure of BPF exception handling and resource leaks (e.g., unreleased
> spinlocks or memory) since the cleanup pads are never executed?
Will add a WARN_ON_ONCE to expose the failure if it happens.
>
>> + if (!consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
>> + break;
>> + }
>> +#endif
>> +}
>> +
>> void bpf_arch_poke_desc_update(struct bpf_jit_poke_descriptor *poke,
>> struct bpf_prog *new, struct bpf_prog *old)
>> {
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests
2026-09-29 0:52 ` bot+bpf-ci
@ 2026-09-30 1:42 ` Yonghong Song
0 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-30 1:42 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/28/26 5:52 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
>> new file mode 100644
>> index 000000000000..76cca8dc45a9
>> --- /dev/null
>> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
>> @@ -0,0 +1,878 @@
> [ ... ]
>
>> +static __used __noinline void __kfunc_btf_anchor(void)
>> +{
>> + bpf_throw(0);
>> + bpf_unwind();
>> + bpf_preempt_disable();
>> + bpf_preempt_enable();
>> + bpf_rcu_read_lock();
>> + bpf_rcu_read_unlock();
>> + bpf_unwind_resume(NULL);
>> +}
>> +
>> +/* An unwind raised in a callee, which is how a cleanup region gets one. */
>> +static __used __naked __noinline __u64 inner_unwind(void)
>> +{
>> + asm volatile (
>> + "r1 = 1;"
>> + "call bpf_unwind;"
>> + "r0 = 0;"
>> + "exit;"
>> + ::: __clobber_all);
>> +}
> [ ... ]
>
>> +static __used __naked __noinline __u64 unwind_in_pad_frame(void)
>> +{
>> + asm volatile (
>> + "call bpf_preempt_disable;"
>> +"1:" "call inner_unwind;" /* cleanup region */
>> +"2:"
>> + "call bpf_preempt_enable;"
>> + "r0 = 0;"
>> + "exit;"
>> +"3:" /* landing pad that unwinds again */
>> + "call bpf_preempt_enable;"
>> + "r1 = 2;"
>> + "call bpf_unwind;"
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + ::: __clobber_all);
>> +}
> [ ... ]
>
>> +static __used __naked __noinline __u64 cb_and_table_frame(void)
>> +{
>> + asm volatile (
>> + "call bpf_preempt_disable;"
>> + "r1 = 9;"
>> +"1:" "call bpf_unwind;" /* cleanup region */
>> +"2:"
>> + "r0 = 0;"
>> + "exit;"
>> +"3:" /* landing pad */
>> + "call bpf_preempt_enable;"
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + ::: __clobber_all);
>> +}
> [ ... ]
>
>> +static __used __naked __noinline __u64 pad_calls_thrower_frame(void)
>> +{
>> + asm volatile (
>> + "call bpf_preempt_disable;"
>> + "r1 = 11;"
>> +"1:" "call bpf_unwind;" /* cleanup region */
>> +"2:"
>> + "r0 = 0;"
>> + "exit;"
>> +"3:" /* landing pad */
>> + "call pad_callee_that_throws;" /* ...which can throw: refused */
>> + "call bpf_preempt_enable;"
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + ::: __clobber_all);
>> +}
> [ ... ]
>
>> +static __used __naked __noinline __u64 catch_pad_frame(void)
>> +{
>> + asm volatile (
>> + "call bpf_preempt_disable;"
>> + "r1 = 12;"
>> +"1:" "call bpf_unwind;" /* cleanup region */
>> +"2:"
>> + "r0 = 0;"
>> + "exit;"
>> +"3:" /* catch pad: no resume, it stops here */
>> + "call bpf_preempt_enable;"
>> + "r0 = 0;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + ::: __clobber_all);
>> +}
> [ ... ]
>
>> +static __used __naked __noinline __u64 stray_resume_frame(void)
>> +{
>> + asm volatile (
>> + "call bpf_preempt_disable;"
>> + "r1 = 13;"
>> +"1:" "call bpf_unwind;" /* cleanup region */
>> +"2:"
>> + "r0 = 0;"
>> + "exit;"
>> +"3:" /* landing pad */
>> + "call bpf_preempt_enable;"
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + ::: __clobber_all);
>> +}
> [ ... ]
>
>> +static __used __naked __noinline __u64 pad_calls_resumer_frame(void)
>> +{
>> + asm volatile (
>> + "call bpf_preempt_disable;"
>> + "r1 = 14;"
>> +"1:" "call bpf_unwind;" /* cleanup region */
>> +"2:"
>> + "r0 = 0;"
>> + "exit;"
>> +"3:" /* landing pad */
>> + "call bpf_preempt_enable;"
>> + "call resume_in_callee;" /* ...which resumes: refused */
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + ::: __clobber_all);
>> +}
> [ ... ]
>
>> +static __used __naked __noinline void gotox_unwinder(void)
>> +{
>> + asm volatile (
>> + "r1 = 15;"
>> + "call bpf_unwind;"
>> + "exit;"
>> + ::: __clobber_all);
>> +}
> This isn't a bug, but bpf_unwind() is declared as __bpf_kfunc void
> bpf_unwind(void) in kernel/bpf/helpers.c and takes no arguments.
>
> Several of these inline asm frames load r1 with a numbered value right
> before calling bpf_unwind: r1=1 in inner_unwind, r1=2 in
> unwind_in_pad_frame, r1=9 in cb_and_table_frame, r1=11 in
> pad_calls_thrower_frame, r1=12 in catch_pad_frame, r1=13 in
> stray_resume_frame, r1=14 in pad_calls_resumer_frame, and r1=15 in
> gotox_unwinder.
>
> Other bpf_unwind call sites in the same file like pad_drops_caller_lock_frame,
> no_pad_keeps_own_lock_frame, and pad_drops_caller_ref_frame call bpf_unwind
> without setting r1 first. The verifier checks no arguments for a zero-argument
> kfunc, and neither bpf_exc_keep_exits nor the runtime walk uses r1.
>
> Are the "r1 = N;" loads before these calls left over from an earlier
> cookie-style signature? Could they be dropped so the asm matches the kfunc
> prototype?
Will fix. thanks!
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36502848904
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags
2026-09-29 0:52 ` bot+bpf-ci
@ 2026-09-30 1:46 ` Yonghong Song
0 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-30 1:46 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/28/26 5:52 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/tools/testing/selftests/bpf/progs/bpf_misc.h b/tools/testing/selftests/bpf/progs/bpf_misc.h
>> index f3dbc3b59bff..9afa163fac5a 100644
>> --- a/tools/testing/selftests/bpf/progs/bpf_misc.h
>> +++ b/tools/testing/selftests/bpf/progs/bpf_misc.h
>> @@ -93,6 +93,11 @@
>> * __failure Expect program load failure in privileged mode.
>> * __failure_unpriv Expect program load failure in unprivileged mode.
>> *
>> + * __set_global Set a global variable of the program to a value before
>> + * executing it.
>> + * __ret_global Execute the program and check that a global variable
>> + * holds the given value afterwards. The variable has to
>> + * live in .bss or .data and be four or eight bytes wide.
>> * __retval Execute the program using BPF_PROG_TEST_RUN command,
>> * expect return value to match passed parameter:
>> * - a decimal number
> [ ... ]
>
>> diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
>> index 25eeb1c1248b..2492841f7a49 100644
>> --- a/tools/testing/selftests/bpf/test_loader.c
>> +++ b/tools/testing/selftests/bpf/test_loader.c
> [ ... ]
>
>> +static int find_global_var(struct bpf_object *obj, const char *name,
>> + struct bpf_map **map, __u32 *off, __u32 *sz,
>> + bool *is_signed)
>> +{
>> + static const char * const secs[] = { ".bss", ".data" };
>> + struct btf *btf = bpf_object__btf(obj);
>> + int i, s;
>> + __u32 vsz;
>> +
>> + if (!btf) {
>> + PRINT_FAIL("no BTF for object\n");
>> + return -ENOENT;
>> + }
>> +
>> + for (s = 0; s < ARRAY_SIZE(secs); s++) {
>> + const struct btf_type *sec, *vt;
>> + const struct btf_var_secinfo *vsi;
>> + struct bpf_map *m = NULL, *iter;
>> + size_t slen = strlen(secs[s]);
>> + int id;
>> +
>> + bpf_object__for_each_map(iter, obj) {
>> + const char *mname = bpf_map__name(iter);
>> + size_t len = mname ? strlen(mname) : 0;
>> +
>> + if (len >= slen && strcmp(mname + len - slen, secs[s]) == 0) {
>> + m = iter;
>> + break;
>> + }
>> + }
> This isn't a bug, but could the map loop here be replaced with
> bpf_object__find_map_by_name(obj, secs[s]), as veristat's
> set_global_vars() does? That would tie the map to the DATASEC by its
> section name and make the comment about the two lookups not agreeing on a
> size unnecessary.
Will do. Thanks!
>
>> + id = btf__find_by_name_kind(btf, secs[s], BTF_KIND_DATASEC);
>> + if (!m || id < 0)
>> + continue;
>> +
>> + sec = btf__type_by_id(btf, id);
>> + vsi = btf_var_secinfos(sec);
>> + for (i = 0; i < btf_vlen(sec); i++, vsi++) {
>> + const struct btf_type *var = btf__type_by_id(btf, vsi->type);
>> +
>> + if (strcmp(btf__name_by_offset(btf, var->name_off), name))
>> + continue;
>> + if (vsi->size != 4 && vsi->size != 8) {
>> + PRINT_FAIL("'%s' is %u bytes, only 4 and 8 are supported\n",
>> + name, vsi->size);
>> + return -EINVAL;
>> + }
>> + vt = btf__type_by_id(btf, btf__resolve_type(btf, var->type));
>> + if (!vt || !(btf_is_int(vt) || btf_is_any_enum(vt))) {
>> + PRINT_FAIL("'%s' is not an int or an enum\n", name);
>> + return -EINVAL;
>> + }
>> + *is_signed = global_var_is_signed(vt);
>> + /*
>> + * The map is found by the suffix of its name and the
>> + * section by its own, so nothing so far has made the
>> + * two agree on a size.
>> + */
>> + vsz = bpf_map__value_size(m);
>> + if (vsi->offset > vsz || vsi->size > vsz - vsi->offset) {
>> + PRINT_FAIL("'%s' at %u+%u is outside '%s' of %u bytes\n",
>> + name, vsi->offset, vsi->size,
>> + bpf_map__name(m), vsz);
>> + return -EINVAL;
>> + }
>> + *map = m;
>> + *off = vsi->offset;
>> + *sz = vsi->size;
>> + return 0;
>> + }
>> + }
>> +
>> + PRINT_FAIL("no global variable '%s'\n", name);
>> + return -ENOENT;
>> +}
> [ ... ]
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36502848904
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes
2026-09-29 0:52 ` bot+bpf-ci
@ 2026-09-30 2:19 ` Yonghong Song
0 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-30 2:19 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/28/26 5:52 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
>> index d1d40314035e..616b479238c7 100644
>> --- a/tools/testing/selftests/bpf/exceptions_cleanup.h
>> +++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
> [ ... ]
>
>> diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
>> index 255f88d35aad..c06ec10359b9 100644
>> --- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
>> +++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
>> @@ -4,6 +4,7 @@
>> #include "exceptions_cleanup.h"
>> #include "exceptions_cleanup.skel.h"
>> #include "exceptions_cleanup_fail.skel.h"
>> +#include "exceptions_cleanup_shapes.skel.h"
>>
>> /* foo3 unwound: every frame that has a pad ran it. */
>> #define PADS_FOO3_UNWOUND \
>> @@ -82,4 +83,5 @@ void test_exceptions_cleanup(void)
>> exceptions_cleanup__destroy(skel);
>>
>> RUN_TESTS(exceptions_cleanup_fail);
>> + RUN_TESTS(exceptions_cleanup_shapes);
>> }
> A subsystem pattern flags this as potentially concerning: The only skip
> gate in test_exceptions_cleanup() is the -EOPNOTSUPP check on the main
> exceptions_cleanup skeleton, which tests bpf_jit_supports_cleanup_pads().
>
> RUN_TESTS(exceptions_cleanup_shapes) then loads entry_callx_region, which
> also needs indirect-call JIT support. progs/exceptions_cleanup_shapes.c
> guards that program only on the compiler and arch:
>
> /* gcc has no indirect calls, and only these JITs emit them */
> #if defined(__clang__) && \
> (defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64))
>
> On x86 the two capabilities are independent:
>
> arch/x86/net/bpf_jit_comp.c:
>
> bool bpf_jit_supports_callx(void)
> {
> return cfi_mode != CFI_FINEIBT;
> }
>
> bool bpf_jit_supports_cleanup_pads(void)
> {
> return IS_ENABLED(CONFIG_UNWINDER_ORC);
> }
>
> With a kCFI kernel on IBT hardware (cfi=auto resolves to CFI_FINEIBT in
> arch/x86/kernel/alternative.c), the main skeleton loads, so the skip
> branch is not taken. The verifier's require_callx_jit() then refuses
> entry_callx_region with -EOPNOTSUPP ("JIT doesn't support callx").
> test_loader's run_subtest() treats any load error on a __success program
> as ASSERT_OK(err, "unexpected_load_failure"), so this subtest fails
> instead of skipping, even though the host simply lacks the feature.
>
> Could the test check for callx support separately and skip on FineIBT
> hosts? The existing callx selftests (RUN_JITED(verifier_callx) in
> prog_tests/verifier.c and callx_func_ptr_map.c) do not skip on FineIBT
> either; they check only is_jit_enabled(). So this shape follows the
> existing convention, and the failure is limited to x86 FineIBT hosts
> (clang kCFI build plus IBT-capable CPU).
This really depends on the config. Currently CI config does not have
CONFIG_CFI so the test should be okay. This is just a CI test and we
can adjust later if config changed.
>
>> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
>> new file mode 100644
>> index 000000000000..1978fd1102f0
>> --- /dev/null
>> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
> [ ... ]
>
>> +/*
>> + * A precision chain crossing a resume: r6 is kept across a call whose only
>> + * way back is the callee's pad, then used as a variable stack offset.
>> + */
>> +static __used __naked __noinline __u64 prec_inner_frame(void)
>> +{
>> + asm volatile (
>> +"1:" "call bpf_unwind;" /* cleanup region */
>> +"2:"
>> + "r0 = 0;"
>> + "exit;"
>> +"3:" /* landing pad */
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + ::: __clobber_all);
>> +}
>> +
>> +static __used __naked __noinline __u64 prec_outer_frame(void)
>> +{
>> + asm volatile (
>> + "r1 = %[input] ll;"
>> + "r6 = *(u64 *)(r1 + 0);"
>> + "r6 &= 0x7;"
>> + "r0 = 0;"
>> + "*(u64 *)(r10 - 8) = r0;"
>> + "*(u64 *)(r10 - 16) = r0;"
>> + "call prec_inner_frame;" /* comes back only through the pad */
>> + "r2 = r10;"
>> + "r2 += -16;"
>> + "r2 += r6;" /* variable stack offset: r6 must be precise */
>> + "*(u8 *)(r2 + 0) = 1;"
>> + "r0 = 0;"
>> + "exit;"
>> + :
>> + : __imm_addr(input)
>> + : __clobber_all);
>> +}
>> +
>> +SEC("?syscall")
>> +__success __set_global(input, 101) __retval(0)
>> +int entry_prec_across_resume(void *ctx)
>> +{
>> + return prec_outer_frame();
>> +}
> A subsystem pattern flags this as potentially concerning:
> entry_prec_across_resume is meant to cover "a precision chain crossing a
> resume" (per the commit message), but does the verifier actually check the
> code after the call?
>
> prec_inner_frame always unwinds into its own pad. The pad ends in
> bpf_unwind_resume while curframe is 1. The resume handler in
> do_check_insn() ends the path there:
>
> kernel/bpf/verifier.c:
>
> /*
> * No need to walk into the caller: its pad was
> * pushed as a branch at its call, and with no
> * pad nothing of it runs.
> */
> if (env->cur_state->curframe)
> return PROCESS_BPF_EXIT;
>
> prec_outer_frame has no CLEANUP_REC over "call prec_inner_frame", so no
> pad branch is pushed there either. The fall-through of bpf_unwind in
> prec_inner_frame is never walked, because process_bpf_unwind() jumps
> straight to the pad. As a result, none of these instructions are ever
> verified:
>
> "r2 = r10;"
> "r2 += -16;"
> "r2 += r6;"
> "*(u8 *)(r2 + 0) = 1;"
>
> They never run either: bpf_unwind_rewrite() points prec_outer_frame's
> uncovered return at the epilogue.
>
> So the variable-offset stack write never runs mark_chain_precision() on
> r6. The resume-crossing branch in kernel/bpf/backtrack.c (the
> bt_subprog_enter() call made on a resume) is never exercised. Would the
> test still pass if "r6 &= 0x7" were deleted, or if the backtrack.c resume
> handling were broken?
>
> Both the comment "comes back only through the pad" and the header "r6 is
> kept across a call whose only way back is the callee's pad" appear
> inconsistent with this: nothing comes back. Could this shape be adjusted
> to cover the call in prec_outer_frame with a CLEANUP_REC and do the r6
> variable-offset access in prec_outer_frame's own pad, so a pad path
> verifies the use and precision is propagated back through the call?
Okay, will modify the test to have proper precision checking.
>
>> +/*
>> + * The jump_into_pad shape with the branch statically dead, so only a
>> + * speculative walk reaches the pad: a barrier rather than a refusal.
>> + */
>> +static __used __naked __noinline __u64 dead_jump_into_pad_frame(void)
>> +{
>> + asm volatile (
>> + "r6 = 0;"
>> + "if r6 > 7 goto 4f;" /* never taken: walked speculatively */
>> +"1:" "call always_unwind;" /* cleanup region */
>> +"2:"
>> + "r0 = 0;"
>> + "exit;"
>> +"3:" /* landing pad */
>> + "r7 = r0;"
>> +"4:" /* ... and its second instruction */
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + ::: __clobber_all);
>> +}
>> +
>> +SEC("?syscall")
>> +__success
>> +int dead_jump_into_pad(void *ctx)
>> +{
>> + return dead_jump_into_pad_frame();
>> +}
> A subsystem pattern flags this as potentially concerning:
> dead_jump_into_pad is meant to cover "a pad reached only by a speculative
> walk" (per the commit message), but does the verifier actually do a
> speculative walk here?
>
> The program has only __success, so test_loader loads it only in privileged
> mode, as root with all capabilities. The verifier sets:
>
> env->bypass_spec_v1 = bpf_bypass_spec_v1(env->prog->aux->token);
>
> bpf_bypass_spec_v1() is true when bpf_token_capable(token, CAP_PERFMON)
> holds, which root satisfies. In check_cond_jmp_op(), the dead branch of
> "if r6 > 7" (pred == 0) is pushed only when bypass_spec_v1 is false:
>
> kernel/bpf/verifier.c:
>
> } else if (pred == 0) {
> ...
> if (!env->bypass_spec_v1) {
> err = sanitize_speculative_path(env, insn, *insn_idx + insn->off + 1,
> *insn_idx);
>
> So label 4 is never walked. bpf_exc_check_insn() never reports "runs both
> inside and outside a landing pad" on a speculative state, and the
> error_recoverable_with_nospec() && state->speculative recovery in
> do_check() never runs. The program loads because the verifier ignores the
> dead branch entirely, not because of the barrier logic. Would the test
> still pass unchanged if that recovery were removed and the error made
> fatal?
>
> The comment "never taken: walked speculatively" and the header "only a
> speculative walk reaches the pad: a barrier rather than a refusal" describe
> something that does not happen under the test's own settings.
>
> Could this also run the program unprivileged with CAP_BPF kept, so it
> loads but CAP_PERFMON is gone and bypass_spec_v1 is false? A syscall
> program is not a perfmon program type, so CAP_BPF is enough to load it:
>
> SEC("?syscall")
> __success __caps_unpriv(CAP_BPF) __success_unpriv
> int dead_jump_into_pad(void *ctx)
Okay, I will try this one "__success __caps_unpriv(CAP_BPF) __success_unpriv".
>
> It could also check the verifier log for the nospec insertion.
>
> [ ... ]
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36502848904
^ permalink raw reply [flat|nested] 46+ messages in thread
* Re: [PATCH bpf-next v7 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton
2026-09-29 0:52 ` bot+bpf-ci
@ 2026-09-30 3:12 ` Yonghong Song
0 siblings, 0 replies; 46+ messages in thread
From: Yonghong Song @ 2026-09-30 3:12 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/28/26 5:52 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel
>> index 2e22bb901bf3..41575e180efb 100644
>> --- a/tools/testing/selftests/bpf/Makefile.skel
>> +++ b/tools/testing/selftests/bpf/Makefile.skel
>> @@ -33,7 +33,7 @@ LINKED_SKELS := test_static_linked.skel.h linked_funcs.skel.h \
>> LSKELS := fexit_sleep.c trace_printk.c trace_vprintk.c map_ptr_kern.c \
>> core_kern.c core_kern_overflow.c test_ringbuf.c \
>> test_ringbuf_n.c test_ringbuf_map_key.c test_ringbuf_write.c \
>> - test_ringbuf_overwrite.c callx_rodata.c
>> + test_ringbuf_overwrite.c callx_rodata.c exceptions_cleanup_light.c
> [ ... ]
>
>> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
>> new file mode 100644
>> index 000000000000..19d53569bbda
>> --- /dev/null
>> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
>> @@ -0,0 +1,39 @@
>> +// SPDX-License-Identifier: GPL-2.0
>> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
>> +#include <vmlinux.h>
>> +#include <bpf/bpf_helpers.h>
>> +#include "bpf_misc.h"
>> +#include "exceptions_cleanup.h"
>> +
>> +static __used __noinline void __kfunc_btf_anchor(void)
>> +{
>> + bpf_throw(0);
>> + bpf_unwind();
>> + bpf_preempt_disable();
>> + bpf_preempt_enable();
>> + bpf_unwind_resume(NULL);
>> +}
>> +
>> +__u64 pads_ran = 0;
>> +
>> +SEC("syscall")
>> +__naked int entry_light(void)
>> +{
>> + asm volatile (
>> + "call bpf_preempt_disable;"
>> +"1:" "call bpf_unwind;" /* cleanup region */
>> +"2:"
>> + "exit;"
>> +"3:" /* landing pad */
>> + "call bpf_preempt_enable;"
>> + PAD_RAN("%[ran]")
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + :
>> + : [ran]"i"(RAN_LIGHT),
>> + __imm_addr(pads_ran)
>> + : __clobber_all);
>> +}
>> +
>> +char _license[] SEC("license") = "GPL";
> A subsystem pattern flags this as potentially concerning:
>
> This adds a new BPF program file to test loading a cleanup table through
> the light skeleton path. The commit message states this is the smallest
> program that can verify whether the table survives the trip through
> bpf_gen__prog_load().
>
> However, progs/exceptions_cleanup.c already covers the same behavior
> more broadly - pads run on unwind and set bits in pads_ran. Makefile.skel
> already has LSKELS_EXTRA for building programs as both light skeleton
> and libbpf skeleton. Could exceptions_cleanup.c be built as a light
> skeleton and have the existing test cases (no_unwind, unwind_from_foo3,
> unwind_from_foo2) run through the gen loader path?
>
> That approach would also exercise the light skeleton code paths this new
> program never reaches. The new program has only one record in the main
> function, so in bpf_prog_collect_cleanup_info() it takes the owner == prog
> branch. The owner->sub_insn_off rebasing for records inside appended
> subprograms, and the multi-frame bpf_unwind_rewrite() walk, would not run
> under the gen loader with this minimal test.
>
> Should this be a light-skeleton case of the existing exceptions_cleanup.c
> instead of a separate file?
Okay, will follow your suggestions.
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36502848904
^ permalink raw reply [flat|nested] 46+ messages in thread
end of thread, other threads:[~2026-09-30 3:12 UTC | newest]
Thread overview: 46+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-29 0:33 ` sashiko-bot
2026-09-29 21:58 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-09-29 0:31 ` sashiko-bot
2026-09-29 22:04 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 07/22] bpf: Resume a covered call at its landing pad Yonghong Song
2026-09-29 0:31 ` sashiko-bot
2026-09-30 0:28 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
2026-09-29 0:36 ` sashiko-bot
2026-09-30 1:09 ` Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:10 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 10/22] bpf: Refuse a private stack for a program that can unwind Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
2026-09-29 1:14 ` bot+bpf-ci
2026-09-30 1:18 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-29 0:30 ` sashiko-bot
2026-09-30 1:34 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 13/22] bpf, arm64: " Yonghong Song
2026-09-29 1:14 ` bot+bpf-ci
2026-09-29 0:17 ` [PATCH bpf-next v7 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:42 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:46 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 2:19 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 3:12 ` Yonghong Song
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox