* [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
@ 2026-09-18 4:41 Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 01/20] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
` (19 more replies)
0 siblings, 20 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:41 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
bpf_throw() walks the BPF call stack to the exception boundary and
discards every frame in between. A frame that owns something -- an RCU
read lock, a preemption-disabled section, a referenced kptr -- never gets
to give it back, so the verifier refuses to let such a frame throw at
all. That is the whole reason a Rust program cannot use bpf_throw() as
its panic path right now: Rust's Drop glue *is* that give-back, and there
is nowhere to run it.
LLVM 23 added the compiler half ([1]). A Rust function that owns a value
across a call that can unwind
fn foo() {
let _guard = RcuReadGuard::new(); /* bpf_rcu_read_lock() */
may_throw(); /* extern "C-unwind" */
} /* Drop: rcu_read_unlock */
lowers to an invoke with a cleanup landing pad holding the Drop call, and
the BPF backend writes one record per invoke region into a .bpf_cleanup
section: a flat table of 12-byte (begin, end, landing_pad) triples, each
field a byte offset into the code section. The rule is "if a frame unwinds
with its return address in [begin, end), run landing_pad before discarding
it". A pad ends with a call to _Unwind_Resume(), which the kernel provides
as the bpf_unwind_resume() kfunc.
This series is the kernel half: take that table at BPF_PROG_LOAD, teach
the verifier that a covered call can also go to its landing pad, and have
bpf_throw() run the pads as it walks.
C has no unwinding, so the selftests spell out by hand what a frontend
emits -- a call site bracketed by two labels, a landing pad, and a record
tying them together. The frame above, written that way:
"call bpf_rcu_read_lock;"
"1:" "call foo3;" /* cleanup region */
"2:"
... normal path, ends in bpf_rcu_read_unlock ...
"6:" /* landing pad */
"call bpf_rcu_read_unlock;"
"call bpf_unwind_resume;"
CLEANUP_REC("1b", "2b", "6b")
Right now that program does not load: the throw inside foo3 is reported as
"bpf_throw cannot be used inside bpf_rcu_read_lock-ed region", because as
far as the verifier is concerned nothing will ever unlock. With the series
the verifier walks the unwind the same way the run time will -- into the
pad, which unlocks, then on to the next frame -- and the program both
loads and releases the lock when it throws.
Design
======
A pad is run, not lowered. bpf_throw() already walks the frames with
arch_bpf_stack_walk(); it now looks each frame's return address up in
that (sub)program's table and calls the pad as a subroutine of the walker,
with the unwinding frame's frame pointer and its callee-saved registers
restored from the spill its callee's prologue left. The pad therefore sees
its own frame but runs on the walker's stack, far below it, so nothing it
calls can disturb the frame it is cleaning up after. The JIT turns its
bpf_unwind_resume() into the way back to the walker.
The verifier walks the same thing, step for step, so the resource rules
are unchanged: whatever a pad releases is released in the verifier state
too, and check_resource_leak() simply moves from "a throw was seen" to the
end of the walk.
1-2 uapi: cleanup_info in BPF_PROG_LOAD, struct bpf_cleanup_info,
and the bpf_unwind_resume() kfunc
3-9 verifier: mark the covered call sites, give them an edge to the
pad, walk the unwind, and refuse the shapes that cannot be
dispatched
10 bpf_throw(): dispatch pads while walking
11-12 x86-64 and arm64 JITs
13-17 libbpf: collect .bpf_cleanup, pass it to the kernel, resolve
_Unwind_Resume, carry it through the light skeleton and the
static linker
18-20 selftests
Limitations
===========
- Cleanup pads only. A catch pad -- one that ends in a plain exit rather
than a resume, which is what Rust's catch_unwind would need -- is
refused: the walker calls a pad as a subroutine and cannot hand a frame
back its own execution. LLVM refuses type-specific catches and filters
on its side as well.
- A JIT that can dispatch pads is required: x86-64 (with
CONFIG_UNWINDER_ORC, which bpf_throw() already needs there) and arm64.
Anywhere else the load fails with -EOPNOTSUPP rather than silently
doing nothing.
- No offloaded programs, no private stack, and no combining a table with
an exception callback.
- In a pad body: no tail call, no indirect jump, and no on-stack call
arguments -- all of them read or write a stack the pad does not own.
- The Rust toolchain does not properly support BPF exception handling
yet. The tables the selftests use are hand-written inline asm, which
the assembler turns into the same relocations the BPF AsmPrinter emits,
so libbpf and the kernel see an object indistinguishable from a
compiler-generated one.
[1] https://github.com/llvm/llvm-project/pull/192164
llvm commit 9d51c891b719 ("[BPF] Add exception handling support
with .bpf_cleanup section")
Changelog
=========
v1 -> v2:
- v1: https://lore.kernel.org/bpf/20260917055645.3926444-1-yonghong.song@linux.dev/
- Consolidate all usages of kern_extern_name() in a single patch in libbpf.
- Avoid compiler warning and add proper cleanup_info_cnt guard in libbpf when
collecting .bpf_cleanup records.
- Add cleanup_info_cnt condition for emit_rel_store() with cleanup_info.
Yonghong Song (20):
bpf: Accept the compiler's exception cleanup table at program load
bpf: Add the bpf_unwind_resume() kfunc
bpf: Add lookups for exception cleanup resumes and landing pads
bpf: Mark the call sites an exception cleanup table covers
bpf: Make exception landing pads reachable in the CFG
bpf: Explore the landing pads no call site reaches
bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
bpf: Walk the exception unwind in the verifier
bpf: Refuse a private stack for a program with an exception cleanup
table
bpf: Dispatch exception cleanup pads from bpf_throw()
bpf, x86: Dispatch exception cleanup pads at run time
bpf, arm64: Dispatch exception cleanup pads at run time
libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
libbpf: Add cleanup_info to bpf_prog_load_opts
libbpf: Collect .bpf_cleanup records and pass them to the kernel
libbpf: Carry the exception cleanup table through the light skeleton
libbpf: Let the static linker carry .bpf_cleanup relocations
selftests/bpf: Add an end-to-end .bpf_cleanup exception test
selftests/bpf: Cover the exception cleanup shapes the chain does not
reach
selftests/bpf: Load an exception cleanup program from a light skeleton
arch/arm64/net/Makefile | 2 +-
arch/arm64/net/bpf_cleanup_pad.S | 95 ++
arch/arm64/net/bpf_jit_comp.c | 97 +-
arch/x86/net/Makefile | 2 +-
arch/x86/net/bpf_cleanup_pad.S | 74 ++
arch/x86/net/bpf_jit_comp.c | 78 +-
include/linux/bpf.h | 75 ++
include/linux/bpf_cleanup_abi.h | 16 +
include/linux/bpf_verifier.h | 11 +
include/linux/filter.h | 2 +
include/uapi/linux/bpf.h | 9 +
kernel/bpf/Makefile | 2 +-
kernel/bpf/cfg.c | 64 +-
kernel/bpf/check_btf.c | 147 +++
kernel/bpf/core.c | 35 +-
kernel/bpf/exception.c | 654 +++++++++++++
kernel/bpf/exception.h | 22 +
kernel/bpf/fixups.c | 143 +++
kernel/bpf/helpers.c | 46 +
kernel/bpf/liveness.c | 20 +
kernel/bpf/states.c | 3 +
kernel/bpf/syscall.c | 2 +-
kernel/bpf/verifier.c | 125 ++-
tools/include/uapi/linux/bpf.h | 9 +
tools/lib/bpf/bpf.c | 6 +-
tools/lib/bpf/bpf.h | 7 +-
tools/lib/bpf/gen_loader.c | 29 +-
tools/lib/bpf/libbpf.c | 312 ++++++-
tools/lib/bpf/libbpf_internal.h | 10 +
tools/lib/bpf/linker.c | 19 +-
tools/testing/selftests/bpf/Makefile | 2 +-
.../selftests/bpf/exceptions_cleanup.h | 52 ++
.../bpf/prog_tests/exceptions_cleanup.c | 421 +++++++++
.../selftests/bpf/progs/exceptions_cleanup.c | 152 +++
.../bpf/progs/exceptions_cleanup_ext_table.c | 48 +
.../bpf/progs/exceptions_cleanup_fail.c | 600 ++++++++++++
.../bpf/progs/exceptions_cleanup_freplace.c | 17 +
.../bpf/progs/exceptions_cleanup_light.c | 39 +
.../progs/exceptions_cleanup_pad_freplace.c | 17 +
.../bpf/progs/exceptions_cleanup_shapes.c | 863 ++++++++++++++++++
40 files changed, 4263 insertions(+), 64 deletions(-)
create mode 100644 arch/arm64/net/bpf_cleanup_pad.S
create mode 100644 arch/x86/net/bpf_cleanup_pad.S
create mode 100644 include/linux/bpf_cleanup_abi.h
create mode 100644 kernel/bpf/exception.c
create mode 100644 kernel/bpf/exception.h
create mode 100644 tools/testing/selftests/bpf/exceptions_cleanup.h
create mode 100644 tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_ext_table.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_freplace.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_pad_freplace.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 01/20] bpf: Accept the compiler's exception cleanup table at program load
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
@ 2026-09-18 4:42 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 02/20] bpf: Add the bpf_unwind_resume() kfunc Yonghong Song
` (18 subsequent siblings)
19 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:42 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
LLVM 23 added exception handling support for BPF with the .bpf_cleanup
section ([1]). Rust code compiled with panic=unwind runs cleanup code
(Drop glue) when bpf_throw() fires, and the LLVM BPF backend emits that
section from the landing pads the frontend produced. Plain C cannot
generate .bpf_cleanup unless inline asm is used. The Rust compiler does not
*properly* support BPF exception handling yet, but the kernel can support
the table today, and inline assembly is enough to test it.
Add the UAPI to carry the .bpf_cleanup table into the kernel. BPF_PROG_LOAD
grows cleanup_info, cleanup_info_cnt and cleanup_info_rec_size, and struct
bpf_cleanup_info describes one record as a triple of instruction indices:
the half-open call-site range [begin_off, end_off) and the landing_pad_off
the frame resumes at. check_cleanup_info() validates the table a program is
loaded with, so the rest of the kernel can rely on it.
[1] https://github.com/llvm/llvm-project/pull/192164
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 2 +
include/uapi/linux/bpf.h | 9 ++
kernel/bpf/check_btf.c | 147 +++++++++++++++++++++++++++++++++
kernel/bpf/syscall.c | 2 +-
kernel/bpf/verifier.c | 1 +
tools/include/uapi/linux/bpf.h | 9 ++
6 files changed, 169 insertions(+), 1 deletion(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index cf85141ea167..c08505b9ba82 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -987,6 +987,8 @@ struct bpf_verifier_env {
struct arg_track **callsite_at_stack;
u32 pass_cnt; /* number of times do_check() was called */
u32 subprog_cnt;
+ struct bpf_cleanup_info *cleanup_info;
+ u32 cleanup_info_cnt;
/* number of instructions analyzed by the verifier */
u32 prev_insn_processed, insn_processed;
/* number of jmps, calls, exits analyzed so far */
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 732b35cc08d1..f7dc121be094 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -1669,6 +1669,9 @@ union bpf_attr {
* verification.
*/
__s32 keyring_id;
+ __aligned_u64 cleanup_info; /* exception cleanup table */
+ __u32 cleanup_info_rec_size; /* userspace bpf_cleanup_info size */
+ __u32 cleanup_info_cnt; /* number of bpf_cleanup_info records */
};
struct { /* anonymous struct used by BPF_OBJ_* commands */
@@ -7588,6 +7591,12 @@ struct bpf_line_info {
__u32 line_col;
};
+struct bpf_cleanup_info {
+ __u32 begin_off;
+ __u32 end_off;
+ __u32 landing_pad_off;
+};
+
struct bpf_spin_lock {
__u32 val;
};
diff --git a/kernel/bpf/check_btf.c b/kernel/bpf/check_btf.c
index 0e8b3ccc7a5b..d03dc791042a 100644
--- a/kernel/bpf/check_btf.c
+++ b/kernel/bpf/check_btf.c
@@ -407,6 +407,149 @@ static int check_core_relo(struct bpf_verifier_env *env,
return err;
}
+static int cleanup_insn_subprog(struct bpf_verifier_env *env, u32 off)
+{
+ struct bpf_subprog_info *info;
+
+ if (off >= env->prog->len)
+ return -1;
+ info = bpf_find_containing_subprog(env, off);
+ return info ? info - env->subprog_info : -1;
+}
+
+#define MIN_BPF_CLEANUP_INFO_SIZE 12
+#define MAX_CLEANUP_INFO_REC_SIZE MAX_FUNCINFO_REC_SIZE
+
+static int check_cleanup_info(struct bpf_verifier_env *env,
+ const union bpf_attr *attr,
+ bpfptr_t uattr)
+{
+ u32 krec_size = sizeof(struct bpf_cleanup_info);
+ u32 i, nrec, urec_size, min_size, prev_end = 0;
+ struct bpf_cleanup_info *krecord;
+ bpfptr_t urecord;
+ int ret = -EINVAL;
+
+ nrec = attr->cleanup_info_cnt;
+ if (!nrec)
+ return 0;
+ if (nrec > INT_MAX / krec_size)
+ return -EINVAL;
+
+ urec_size = attr->cleanup_info_rec_size;
+ if (urec_size < MIN_BPF_CLEANUP_INFO_SIZE ||
+ urec_size > MAX_CLEANUP_INFO_REC_SIZE ||
+ urec_size % sizeof(u32)) {
+ verbose(env, "invalid cleanup info rec size %u\n", urec_size);
+ return -EINVAL;
+ }
+
+ krecord = kvcalloc(nrec, krec_size, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!krecord)
+ return -ENOMEM;
+
+ min_size = min_t(u32, krec_size, urec_size);
+ urecord = make_bpfptr(attr->cleanup_info, uattr.is_kernel);
+ for (i = 0; i < nrec; i++) {
+ struct bpf_cleanup_info *rec = &krecord[i];
+ int sb, se, sl;
+
+ ret = bpf_check_uarg_tail_zero(urecord, krec_size, urec_size);
+ if (ret) {
+ if (ret == -E2BIG) {
+ verbose(env, "nonzero tailing record in cleanup info\n");
+ if (copy_to_bpfptr_offset(uattr,
+ offsetof(union bpf_attr,
+ cleanup_info_rec_size),
+ &min_size, sizeof(min_size)))
+ ret = -EFAULT;
+ }
+ goto err_free;
+ }
+
+ if (copy_from_bpfptr(rec, urecord, min_size)) {
+ ret = -EFAULT;
+ goto err_free;
+ }
+ bpfptr_add(&urecord, urec_size);
+
+ ret = -EINVAL;
+ if (rec->begin_off >= rec->end_off) {
+ verbose(env, "cleanup_info[%u]: begin %u >= end %u\n",
+ i, rec->begin_off, rec->end_off);
+ goto err_free;
+ }
+ if (i && rec->begin_off < prev_end) {
+ verbose(env,
+ "cleanup_info[%u]: range [%u,%u) is unsorted or overlaps the previous record\n",
+ i, rec->begin_off, rec->end_off);
+ goto err_free;
+ }
+ prev_end = rec->end_off;
+
+ sb = cleanup_insn_subprog(env, rec->begin_off);
+ se = cleanup_insn_subprog(env, rec->end_off - 1);
+ sl = cleanup_insn_subprog(env, rec->landing_pad_off);
+ if (sb < 0 || se < 0 || sl < 0) {
+ verbose(env, "cleanup_info[%u]: offset out of range\n", i);
+ goto err_free;
+ }
+ if (sb != se || sb != sl) {
+ verbose(env,
+ "cleanup_info[%u]: range/landing pad span multiple subprogs\n",
+ i);
+ goto err_free;
+ }
+ /*
+ * The second half of a 16-byte instruction carries a zero
+ * opcode and is not an instruction of its own, so no offset
+ * may name one. end_off is exclusive, so it may also be one
+ * past the last instruction of the program.
+ */
+ if (!env->prog->insnsi[rec->begin_off].code ||
+ !env->prog->insnsi[rec->landing_pad_off].code ||
+ (rec->end_off < env->prog->len &&
+ !env->prog->insnsi[rec->end_off].code)) {
+ verbose(env, "cleanup_info[%u]: points at invalid insn\n", i);
+ goto err_free;
+ }
+ }
+
+ /*
+ * Reject a landing pad that lies inside a call-site range, its own
+ * included: it would be both a pad and a call that unwinds to one, and
+ * an exception out of it would have nowhere to go.
+ */
+ ret = -EINVAL;
+ for (i = 0; i < nrec; i++) {
+ u32 pad = krecord[i].landing_pad_off;
+ u32 l = 0, r = nrec;
+
+ while (l < r) {
+ u32 m = l + (r - l) / 2;
+
+ if (pad < krecord[m].begin_off) {
+ r = m;
+ } else if (pad >= krecord[m].end_off) {
+ l = m + 1;
+ } else {
+ verbose(env,
+ "cleanup_info[%u]: landing pad %u is inside the call-site range of cleanup_info[%u]\n",
+ i, pad, m);
+ goto err_free;
+ }
+ }
+ }
+
+ env->cleanup_info = krecord;
+ env->cleanup_info_cnt = nrec;
+ return 0;
+
+err_free:
+ kvfree(krecord);
+ return ret;
+}
+
int bpf_prepare_btf_info(struct bpf_verifier_env *env,
const union bpf_attr *attr,
bpfptr_t uattr)
@@ -441,6 +584,10 @@ int bpf_check_btf_info(struct bpf_verifier_env *env,
{
int err;
+ err = check_cleanup_info(env, attr, uattr);
+ if (err)
+ return err;
+
if (!attr->func_info_cnt && !attr->line_info_cnt) {
if (check_abnormal_return(env))
return -EINVAL;
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index def57bddb092..ac7091469766 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -2912,7 +2912,7 @@ int __init __used bpf_multi_func(void) { return 0; }
BTF_ID_LIST_GLOBAL_SINGLE(bpf_multi_func_btf_id, func, bpf_multi_func)
/* last field in 'union bpf_attr' used by this command */
-#define BPF_PROG_LOAD_LAST_FIELD keyring_id
+#define BPF_PROG_LOAD_LAST_FIELD cleanup_info_cnt
static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, struct bpf_log_attr *attr_log)
{
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 6c6b8d8520cd..c65ff2e326bf 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -21845,6 +21845,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
kvfree(env->succ);
kvfree(env->gotox_tmp_buf);
bpf_diag_free(env);
+ kvfree(env->cleanup_info);
kvfree(env);
return ret;
}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 732b35cc08d1..f7dc121be094 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -1669,6 +1669,9 @@ union bpf_attr {
* verification.
*/
__s32 keyring_id;
+ __aligned_u64 cleanup_info; /* exception cleanup table */
+ __u32 cleanup_info_rec_size; /* userspace bpf_cleanup_info size */
+ __u32 cleanup_info_cnt; /* number of bpf_cleanup_info records */
};
struct { /* anonymous struct used by BPF_OBJ_* commands */
@@ -7588,6 +7591,12 @@ struct bpf_line_info {
__u32 line_col;
};
+struct bpf_cleanup_info {
+ __u32 begin_off;
+ __u32 end_off;
+ __u32 landing_pad_off;
+};
+
struct bpf_spin_lock {
__u32 val;
};
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 02/20] bpf: Add the bpf_unwind_resume() kfunc
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 01/20] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
@ 2026-09-18 4:42 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 03/20] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
` (17 subsequent siblings)
19 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:42 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A compiler-emitted cleanup landing pad ends with a call to
_Unwind_Resume(), which carries the unwind on once that frame's cleanups
have run. The kernel provides the same terminator as a kfunc, named
bpf_unwind_resume() to keep the 'bpf_' prefix kfunc convention; a later
libbpf patch resolves the compiler's name to it.
The body never runs. A JIT emits the way back out of a landing pad in its
place.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/helpers.c | 12 ++++++++++++
1 file changed, 12 insertions(+)
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 051b6654e57c..31d84de4d318 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -3407,6 +3407,17 @@ __bpf_kfunc void bpf_throw(u64 cookie)
WARN(1, "A call to BPF exception callback should never return\n");
}
+/*
+ * Terminator of a compiler-emitted cleanup landing pad. The compiler names
+ * this _Unwind_Resume, the base unwind ABI's entry point for carrying an
+ * unwind on once a frame's cleanups have run. To match kernel kfunc
+ * convention, the kernel calls it bpf_unwind_resume and libbpf maps the
+ * compiler's name onto it.
+ */
+__bpf_kfunc void bpf_unwind_resume(void)
+{
+}
+
__bpf_kfunc int bpf_wq_init(struct bpf_wq *wq, void *p__const_map, unsigned int flags)
{
struct bpf_async_kern *async = (struct bpf_async_kern *)wq;
@@ -4853,6 +4864,7 @@ BTF_ID_FLAGS(func, bpf_task_get_cgroup1, KF_ACQUIRE | KF_RCU | KF_RET_NULL)
BTF_ID_FLAGS(func, bpf_task_from_pid, KF_ACQUIRE | KF_RET_NULL)
BTF_ID_FLAGS(func, bpf_task_from_vpid, KF_ACQUIRE | KF_RET_NULL)
BTF_ID_FLAGS(func, bpf_throw)
+BTF_ID_FLAGS(func, bpf_unwind_resume)
#ifdef CONFIG_BPF_EVENTS
BTF_ID_FLAGS(func, bpf_send_signal_task)
#endif
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 03/20] bpf: Add lookups for exception cleanup resumes and landing pads
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 01/20] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 02/20] bpf: Add the bpf_unwind_resume() kfunc Yonghong Song
@ 2026-09-18 4:42 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-18 4:42 ` [PATCH bpf-next v2 04/20] bpf: Mark the call sites an exception cleanup table covers Yonghong Song
` (16 subsequent siblings)
19 siblings, 1 reply; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:42 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Add two new files, exception.h and exception.c, to host the exception
handling code. Only a few helpers so far: recognising a call to
bpf_unwind_resume(), and asking which landing pad, if any, a call site
unwinds to.
The pad of a call site is kept in insn_aux_data, so the two places that
move instructions around -- bpf_patch_insn_data() and
verifier_remove_insns() -- learn to keep it in step. Nothing sets the mark
yet; the next patch does.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 6 ++++++
kernel/bpf/Makefile | 2 +-
kernel/bpf/exception.c | 36 ++++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 12 ++++++++++++
kernel/bpf/fixups.c | 17 +++++++++++++++++
5 files changed, 72 insertions(+), 1 deletion(-)
create mode 100644 kernel/bpf/exception.c
create mode 100644 kernel/bpf/exception.h
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index c08505b9ba82..f9bccd3e0f4d 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -681,6 +681,11 @@ struct bpf_insn_aux_data {
bool needs_zext; /* alu op needs to clear upper bits */
bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
+ /*
+ * 1 + the instruction index of the exception cleanup landing pad this
+ * call site unwinds to, or 0 for none.
+ */
+ u32 cleanup_pad;
bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
u8 alu_state; /* used in combination with alu_limit */
/* true if STX or LDX instruction is a part of a spill/fill
@@ -1518,6 +1523,7 @@ u32 btf_func_arg_align(const struct btf *btf, const struct btf_type *t);
int bpf_find_subprog(struct bpf_verifier_env *env, int off);
bool bpf_is_throw_kfunc(struct bpf_insn *insn);
+bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
int bpf_compute_const_regs(struct bpf_verifier_env *env);
int bpf_prune_dead_branches(struct bpf_verifier_env *env);
int bpf_check_cfg(struct bpf_verifier_env *env);
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index 9a92c348bbda..af9bc60428ad 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -11,7 +11,7 @@ obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_
obj-$(CONFIG_BPF_SYSCALL) += hashtab.o arraymap.o percpu_freelist.o bpf_lru_list.o lpm_trie.o map_in_map.o bloom_filter.o
obj-$(CONFIG_BPF_SYSCALL) += local_storage.o queue_stack_maps.o ringbuf.o bpf_insn_array.o
obj-$(CONFIG_BPF_SYSCALL) += bpf_local_storage.o bpf_task_storage.o
-obj-$(CONFIG_BPF_SYSCALL) += fixups.o cfg.o states.o backtrack.o check_btf.o
+obj-$(CONFIG_BPF_SYSCALL) += fixups.o cfg.o states.o backtrack.o check_btf.o exception.o
obj-${CONFIG_BPF_LSM} += bpf_inode_storage.o
obj-$(CONFIG_BPF_SYSCALL) += disasm.o mprog.o
obj-$(CONFIG_BPF_JIT) += trampoline.o
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
new file mode 100644
index 000000000000..3895b9453639
--- /dev/null
+++ b/kernel/bpf/exception.c
@@ -0,0 +1,36 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bpf.h>
+#include <linux/bpf_verifier.h>
+#include <linux/btf.h>
+#include <linux/btf_ids.h>
+#include <linux/filter.h>
+#include <linux/slab.h>
+#include "exception.h"
+
+#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
+
+enum exc_kfunc {
+ EXC_KF_bpf_unwind_resume,
+};
+
+BTF_ID_LIST(exc_kfunc_list)
+BTF_ID(func, bpf_unwind_resume)
+
+static bool insn_is_exc_kfunc(const struct bpf_insn *insn, int kf)
+{
+ return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
+ insn->imm == exc_kfunc_list[kf];
+}
+
+bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
+{
+ return insn_is_exc_kfunc(insn, EXC_KF_bpf_unwind_resume);
+}
+
+int bpf_cleanup_pad_of_call(struct bpf_verifier_env *env, u32 idx)
+{
+ u32 pad = env->insn_aux_data[idx].cleanup_pad;
+
+ return pad ? (int)pad - 1 : -1;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
new file mode 100644
index 000000000000..0f2b9624a2ce
--- /dev/null
+++ b/kernel/bpf/exception.h
@@ -0,0 +1,12 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef _LINUX_BPF_EXCEPTION_H
+#define _LINUX_BPF_EXCEPTION_H
+
+#include <linux/types.h>
+
+struct bpf_verifier_env;
+
+int bpf_cleanup_pad_of_call(struct bpf_verifier_env *env, u32 idx);
+
+#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 2add8001c3ec..82b00fac6bd6 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -261,6 +261,11 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
}
}
+ if (env->cleanup_info_cnt)
+ for (i = 0; i < prog_len; i++)
+ if (data[i].cleanup_pad > off + 1)
+ data[i].cleanup_pad += cnt - 1;
+
/*
* Last slot instruction could be a newly generated
* BPF_ST/BPF_LDX/BPF_STX, systematically mark it for non-stack access
@@ -549,6 +554,7 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
unsigned int orig_prog_len = env->prog->len;
int err;
+ u32 i;
if (bpf_prog_is_offloaded(env->prog->aux))
bpf_prog_offload_remove_insns(env, off, cnt);
@@ -573,6 +579,17 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
sizeof(*aux_data) * (orig_prog_len - off - cnt));
env->insn_aux_data_len -= cnt;
+ if (env->cleanup_info_cnt) {
+ for (i = 0; i < env->insn_aux_data_len; i++) {
+ u32 pad = aux_data[i].cleanup_pad;
+
+ if (pad > off + cnt)
+ aux_data[i].cleanup_pad = pad - cnt;
+ else if (pad > off)
+ aux_data[i].cleanup_pad = 0;
+ }
+ }
+
return 0;
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 04/20] bpf: Mark the call sites an exception cleanup table covers
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (2 preceding siblings ...)
2026-09-18 4:42 ` [PATCH bpf-next v2 03/20] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
@ 2026-09-18 4:42 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 05/20] bpf: Make exception landing pads reachable in the CFG Yonghong Song
` (15 subsequent siblings)
19 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:42 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Some plumbing work is done before bpf_check_cfg(). More specifically,
insn_aux_data records cleanup_throw_site for every bpf_throw(), and
cleanup_pad, the landing pad a frame resumes at, for every call within the
[begin_off, end_off) range of a cleanup record. Subsequent commits consume
both.
bpf_prepare_cleanup_exceptions() runs before bpf_check_cfg(), because what
it produces is what the CFG walk consumes. It refuses a table on an
offloaded program, on a program whose JIT cannot dispatch landing pads or
which the JIT was not asked to compile, and on a program that also installs
an exception callback -- two different answers to what runs on the way out.
bpf_jit_supports_cleanup_pads() is weak here and says no; the arch patches
provide the real ones.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 1 +
include/linux/filter.h | 1 +
kernel/bpf/core.c | 5 ++++
kernel/bpf/exception.c | 55 ++++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 1 +
kernel/bpf/verifier.c | 6 ++++
6 files changed, 69 insertions(+)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index f9bccd3e0f4d..7463e86d15ee 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -681,6 +681,7 @@ struct bpf_insn_aux_data {
bool needs_zext; /* alu op needs to clear upper bits */
bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
+ bool cleanup_throw_site; /* call to bpf_throw() */
/*
* 1 + the instruction index of the exception cleanup landing pad this
* call site unwinds to, or 0 for none.
diff --git a/include/linux/filter.h b/include/linux/filter.h
index b17222db2efc..9682b98ad890 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1242,6 +1242,7 @@ bool bpf_jit_supports_stack_args(void);
bool bpf_jit_supports_arena_args(void);
bool bpf_jit_supports_far_kfunc_call(void);
bool bpf_jit_supports_exceptions(void);
+bool bpf_jit_supports_cleanup_pads(void);
bool bpf_jit_supports_ptr_xchg(void);
bool bpf_jit_supports_arena(void);
bool bpf_jit_supports_insn(struct bpf_insn *insn, bool in_arena);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index 4e208cc94752..f0dd851264f4 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -3470,6 +3470,11 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
{
}
+bool __weak bpf_jit_supports_cleanup_pads(void)
+{
+ return false;
+}
+
bool __weak bpf_jit_supports_timed_may_goto(void)
{
return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index 3895b9453639..dfe9c2a9ce7d 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -23,6 +23,61 @@ static bool insn_is_exc_kfunc(const struct bpf_insn *insn, int kf)
insn->imm == exc_kfunc_list[kf];
}
+static void cleanup_mark_throw_sites(struct bpf_verifier_env *env)
+{
+ u32 i;
+
+ for (i = 0; i < env->prog->len; i++)
+ if (bpf_is_throw_kfunc(&env->prog->insnsi[i]))
+ env->insn_aux_data[i].cleanup_throw_site = true;
+}
+
+static void cleanup_mark_call_sites(struct bpf_verifier_env *env)
+{
+ u32 i, j;
+
+ for (i = 0; i < env->cleanup_info_cnt; i++) {
+ struct bpf_cleanup_info *rec = &env->cleanup_info[i];
+
+ for (j = rec->begin_off; j < rec->end_off; j++) {
+ struct bpf_insn *insn = &env->prog->insnsi[j];
+
+ if (!bpf_pseudo_call(insn) && !bpf_is_throw_kfunc(insn))
+ continue;
+ env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
+ }
+ }
+}
+
+int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env)
+{
+ if (!env->cleanup_info_cnt)
+ return 0;
+
+ if (bpf_prog_is_offloaded(env->prog->aux)) {
+ verbose(env,
+ "exception cleanup is not supported for offloaded programs\n");
+ return -EINVAL;
+ }
+
+ if (!bpf_jit_supports_cleanup_pads() || !env->prog->jit_requested) {
+ verbose(env,
+ "exception cleanup needs a JIT that can dispatch landing pads\n");
+ return -EOPNOTSUPP;
+ }
+ env->prog->jit_required = 1;
+
+ if (env->exception_callback_subprog) {
+ verbose(env,
+ "exception cleanup table cannot be combined with an exception callback\n");
+ return -EINVAL;
+ }
+
+ cleanup_mark_throw_sites(env);
+ cleanup_mark_call_sites(env);
+ return 0;
+}
+
bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
{
return insn_is_exc_kfunc(insn, EXC_KF_bpf_unwind_resume);
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index 0f2b9624a2ce..f51383fd775c 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -7,6 +7,7 @@
struct bpf_verifier_env;
+int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
int bpf_cleanup_pad_of_call(struct bpf_verifier_env *env, u32 idx);
#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index c65ff2e326bf..6496c30a10f7 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -37,6 +37,7 @@
#include "diagnostics.h"
#include "disasm.h"
+#include "exception.h"
static const struct bpf_verifier_ops * const bpf_verifier_ops[] = {
#define BPF_PROG_TYPE(_id, _name, prog_ctx_type, kern_ctx_type) \
@@ -21638,6 +21639,11 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
if (ret < 0)
goto skip_full_check;
+ /* The CFG needs an edge from a call in a cleanup range to its pad. */
+ ret = bpf_prepare_cleanup_exceptions(env);
+ if (ret < 0)
+ goto skip_full_check;
+
/* Validate instructions and resolve the program's referenced resources. */
ret = check_and_resolve_insns(env);
if (ret < 0)
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 05/20] bpf: Make exception landing pads reachable in the CFG
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (3 preceding siblings ...)
2026-09-18 4:42 ` [PATCH bpf-next v2 04/20] bpf: Mark the call sites an exception cleanup table covers Yonghong Song
@ 2026-09-18 4:42 ` Yonghong Song
2026-09-18 4:59 ` sashiko-bot
2026-09-18 5:44 ` bot+bpf-ci
2026-09-18 4:42 ` [PATCH bpf-next v2 06/20] bpf: Explore the landing pads no call site reaches Yonghong Song
` (14 subsequent siblings)
19 siblings, 2 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:42 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Both the CFG walk and liveness are involved. A bpf_throw() or a bpf2bpf
call inside the [begin_off, end_off) range of a cleanup record can reach
that record's landing pad, so both grow that edge; and a
bpf_unwind_resume() reaches nothing after it, so it has no successors.
Liveness needs one more thing. bpf_stack_slot_alive() decides whether an
outer frame's stack slot is still read after the call the frame is
suspended at by looking at the instruction after the call. A slot whose
only remaining reader is the landing pad -- which is every slot a
compiler-generated pad reloads, since only four registers survive a call --
is dead by that measure, so clean_verifier_state() poisons it while the
callee runs and the pad is then rejected for reading it. Ask about the pad
as well when the call site names one.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/cfg.c | 47 ++++++++++++++++++++++++++++++++++++++++---
kernel/bpf/liveness.c | 20 ++++++++++++++++++
2 files changed, 64 insertions(+), 3 deletions(-)
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index 842c7d1eabcc..9f8b8b54d5ea 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -6,6 +6,7 @@
#include <linux/sort.h>
#include "diagnostics.h"
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -158,17 +159,57 @@ static int push_insn(int t, int w, int e, struct bpf_verifier_env *env)
return DONE_EXPLORING;
}
+static int visit_cleanup_pad_edge(int t, struct bpf_verifier_env *env)
+{
+ int *insn_stack = env->cfg.insn_stack;
+ int *insn_state = env->cfg.insn_state;
+ int w;
+
+ if (!env->cleanup_info_cnt)
+ return DONE_EXPLORING;
+ w = bpf_cleanup_pad_of_call(env, t);
+ if (w < 0)
+ return DONE_EXPLORING;
+
+ mark_prune_point(env, t);
+ mark_jmp_point(env, w);
+ mark_jump_target(env, w);
+
+ if (insn_state[w])
+ return DONE_EXPLORING;
+ if (env->cfg.cur_stack >= env->prog->len)
+ return -E2BIG;
+ insn_stack[env->cfg.cur_stack++] = w;
+ insn_state[w] |= DISCOVERED;
+ return KEEP_EXPLORING;
+}
+
+static int merge_visit_ret(int a, int b)
+{
+ if (a < 0)
+ return a;
+ if (b < 0)
+ return b;
+ if (a == KEEP_EXPLORING || b == KEEP_EXPLORING)
+ return KEEP_EXPLORING;
+ return DONE_EXPLORING;
+}
+
static int visit_func_call_insn(int t, struct bpf_insn *insns,
struct bpf_verifier_env *env,
bool visit_callee)
{
- int ret, insn_sz;
+ int ret, insn_sz, pad_ret;
int w;
+ pad_ret = visit_cleanup_pad_edge(t, env);
+ if (pad_ret < 0)
+ return pad_ret;
+
insn_sz = bpf_is_ldimm64(&insns[t]) ? 2 : 1;
ret = push_insn(t, t + insn_sz, FALLTHROUGH, env);
if (ret)
- return ret;
+ return merge_visit_ret(pad_ret, ret);
mark_prune_point(env, t + insn_sz);
/* when we exit from subprog, we need to record non-linear history */
@@ -180,7 +221,7 @@ static int visit_func_call_insn(int t, struct bpf_insn *insns,
merge_callee_effects(env, t, w);
ret = push_insn(t, w, BRANCH, env);
}
- return ret;
+ return merge_visit_ret(pad_ret, ret);
}
struct bpf_iarray *bpf_iarray_realloc(struct bpf_iarray *old, size_t n_elem)
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index 44ecdc5b4ec2..9cfd05f970bc 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -8,6 +8,8 @@
#include <linux/slab.h>
#include <linux/sort.h>
+#include "exception.h"
+
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
struct per_frame_masks {
@@ -256,6 +258,9 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
succ = env->succ;
succ->cnt = 0;
+ if (unlikely(bpf_is_unwind_resume_kfunc(insn)))
+ return succ;
+
opcode_info = &opcode_info_tbl[BPF_CLASS(insn->code) | BPF_OP(insn->code)];
insn_sz = bpf_is_ldimm64(insn) ? 2 : 1;
if (opcode_info->can_fallthrough)
@@ -264,6 +269,13 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
if (opcode_info->can_jump)
succ->items[succ->cnt++] = idx + bpf_jmp_offset(insn) + 1;
+ if (unlikely(env->cleanup_info_cnt)) {
+ int pad = bpf_cleanup_pad_of_call(env, idx);
+
+ if (pad >= 0)
+ succ->items[succ->cnt++] = pad;
+ }
+
return succ;
}
@@ -397,6 +409,14 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
alive = bpf_calls_callback(env, callsite)
? is_live_before(instance, callsite, rel, half_spi)
: is_live_before(instance, callsite + 1, rel, half_spi);
+
+ /* Control may also go to the landing pad. */
+ if (!alive && unlikely(env->cleanup_info_cnt)) {
+ int pad = bpf_cleanup_pad_of_call(env, callsite);
+
+ if (pad >= 0)
+ alive = is_live_before(instance, pad, rel, half_spi);
+ }
if (alive)
return true;
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 06/20] bpf: Explore the landing pads no call site reaches
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (4 preceding siblings ...)
2026-09-18 4:42 ` [PATCH bpf-next v2 05/20] bpf: Make exception landing pads reachable in the CFG Yonghong Song
@ 2026-09-18 4:42 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-18 4:42 ` [PATCH bpf-next v2 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch Yonghong Song
` (13 subsequent siblings)
19 siblings, 1 reply; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:42 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A cleanup record need not cover a call an exception can unwind out of: a
frontend is free to emit a region around a helper or an ordinary kfunc,
both nounwind here. Nothing marks a call site then, and that record's
landing pad is reached by nothing at all -- leaving bpf_check_cfg() to
refuse the program over code its own frontend had no way not to emit:
0: call bpf_preempt_disable
1: call bpf_preempt_enable record = { begin = 1, end = 2, pad = 4 }
2: r0 = 0
3: exit
4: r1 = pads_ran ll landing pad
6: r2 = *(u64 *)(r1 + 0)
7: r2 |= RAN_NOUNWIND_REC
8: *(u64 *)(r1 + 0) = r2
9: call bpf_unwind_resume
10: exit
The range [1,2) holds one call, and it is a kfunc, so an exception cannot
come out of it. cleanup_mark_call_sites() marks nothing, nothing pushes an
edge to 4, and 4 through 10 are reachable from nothing: "unreachable insn
4". The pad is dead, which is correct -- no exception can ever arrive at it
-- but the program is fine and has to load.
Walk every pad the table names that the edges did not reach, the way the
walk is already re-seeded at an exception callback. From there the pad is
code like any other: do_check() never enters it, because no call site
dispatches to it, so the dead code sweep removes it along with everything
else that was not reached.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/cfg.c | 17 +++++++++++++++++
1 file changed, 17 insertions(+)
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index 9f8b8b54d5ea..5a2b48b6a9e1 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -633,6 +633,7 @@ int bpf_check_cfg(struct bpf_verifier_env *env)
int insn_cnt = env->prog->len;
int *insn_stack, *insn_state;
int ex_insn_beg, i, ret = 0;
+ u32 pad_idx = 0;
insn_state = env->cfg.insn_state = kvzalloc_objs(int, insn_cnt,
GFP_KERNEL_ACCOUNT);
@@ -688,6 +689,22 @@ int bpf_check_cfg(struct bpf_verifier_env *env)
goto walk_cfg;
}
+ /*
+ * A landing pad no call site was marked with -- a record whose range
+ * holds no call an exception can unwind out of -- is reached by
+ * nothing. Walk it from here, and let the dead code sweep remove it.
+ */
+ while (pad_idx < env->cleanup_info_cnt) {
+ u32 pad = env->cleanup_info[pad_idx++].landing_pad_off;
+
+ if (insn_state[pad] != EXPLORED) {
+ insn_state[pad] = DISCOVERED;
+ insn_stack[0] = pad;
+ env->cfg.cur_stack = 1;
+ goto walk_cfg;
+ }
+ }
+
for (i = 0; i < insn_cnt; i++) {
struct bpf_insn *insn = &env->prog->insnsi[i];
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (5 preceding siblings ...)
2026-09-18 4:42 ` [PATCH bpf-next v2 06/20] bpf: Explore the landing pads no call site reaches Yonghong Song
@ 2026-09-18 4:42 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 08/20] bpf: Walk the exception unwind in the verifier Yonghong Song
` (12 subsequent siblings)
19 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:42 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
The table is on the instructions and the landing pads are in the control
flow graph. What is left before bpf_throw() can be taught to dispatch them
is to work out what that walk will need, and refuse the shapes it could not
handle.
bpf_check_cleanup_exceptions() runs after bpf_check_cfg(). Everything it
needs is control flow, so it reads what that walk has already worked out
rather than working it out again:
subprog_info.might_throw which subprograms an exception may leave,
closed over the call graph by
merge_callee_effects() as the walk pops each
callee
cleanup_reachability() what each instruction can reach -- a resume, a
plain exit, a throw, an indirect jump
cleanup_mark_pad_bodies() which instructions only ever run with an
exception already in flight
What it refuses:
- a pad that reaches both a resume and a plain exit, or neither: nothing
says whether it is a cleanup pad or a catch pad
- a catch pad, which ends in a plain exit: the walker calls a pad as a
subroutine and cannot hand a frame back its own execution
- a throw in a pad, or a call from a pad to a subprogram that can throw:
a second unwind over frames the first is still discarding
- a bpf_unwind_resume() outside a pad body
- a subprogram that may unwind used as a helper callback: the helper's
own kernel frame would end the walk before it found a boundary
- a tail call, an indirect jump, or an outgoing on-stack call argument in
a pad body, all of which touch a stack the pad does not own
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 1 +
kernel/bpf/exception.c | 397 +++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 2 +
kernel/bpf/fixups.c | 1 +
kernel/bpf/verifier.c | 8 +
5 files changed, 409 insertions(+)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 7463e86d15ee..fe8b26351a15 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -682,6 +682,7 @@ struct bpf_insn_aux_data {
bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
bool cleanup_throw_site; /* call to bpf_throw() */
+ bool in_cleanup_pad; /* only runs with an exception in flight */
/*
* 1 + the instruction index of the exception cleanup landing pad this
* call site unwinds to, or 0 for none.
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index dfe9c2a9ce7d..b2bf831242b8 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -6,6 +6,7 @@
#include <linux/btf_ids.h>
#include <linux/filter.h>
#include <linux/slab.h>
+#include <linux/sort.h>
#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -23,6 +24,130 @@ static bool insn_is_exc_kfunc(const struct bpf_insn *insn, int kf)
insn->imm == exc_kfunc_list[kf];
}
+/* What an instruction does to intra-subprog control flow. */
+enum cleanup_insn_kind {
+ CLEANUP_INSN_PLAIN, /* the next insn runs */
+ CLEANUP_INSN_JUMP, /* unconditional jump */
+ CLEANUP_INSN_COND, /* the next insn runs, or the branch target */
+ CLEANUP_INSN_EXIT,
+ CLEANUP_INSN_THROW, /* call bpf_throw: nothing after it runs */
+ CLEANUP_INSN_RESUME, /* call bpf_unwind_resume: likewise */
+ CLEANUP_INSN_CALL, /* call to another subprog */
+ CLEANUP_INSN_GOTOX, /* indirect jump: successors not known here */
+};
+
+/* What each instruction can reach, computed once by cleanup_reachability(). */
+#define CLEANUP_REACH_RESUME BIT(0) /* a bpf_unwind_resume() call */
+#define CLEANUP_REACH_EXIT BIT(1) /* a plain BPF_EXIT */
+#define CLEANUP_REACH_UNKNOWN BIT(2) /* an indirect jump */
+#define CLEANUP_REACH_THROW BIT(3) /* a bpf_throw() call */
+
+/* Scratch shared by the analyses, sized once so no walker has to allocate. */
+struct cleanup_ctx {
+ struct bpf_verifier_env *env;
+ u8 *reach; /* per insn: CLEANUP_REACH_* mask */
+ u32 *stack; /* per insn: DFS stack */
+ void *scratch; /* the one allocation all of the above live in */
+};
+
+static bool in_pad(struct bpf_verifier_env *env, u32 i)
+{
+ return env->insn_aux_data[i].in_cleanup_pad;
+}
+
+/* One scratch array for cleanup_alloc() to hand out. */
+struct cleanup_alloc_req {
+ void **dst;
+ size_t n, sz;
+};
+
+static void *cleanup_alloc(const struct cleanup_alloc_req *tab, u32 cnt)
+{
+ size_t total = 0;
+ char *block, *p;
+ u32 i;
+
+ for (i = 0; i < cnt; i++)
+ total += round_up(tab[i].n * tab[i].sz, 8);
+
+ block = kvzalloc(total, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!block)
+ return NULL;
+
+ for (i = 0, p = block; i < cnt; i++) {
+ *tab[i].dst = p;
+ p += round_up(tab[i].n * tab[i].sz, 8);
+ }
+ return block;
+}
+
+static int cleanup_subprog_of(struct bpf_verifier_env *env, u32 off)
+{
+ struct bpf_subprog_info *info = bpf_find_containing_subprog(env, off);
+
+ return info ? info - env->subprog_info : -1;
+}
+
+/* The subprogram a linear pass is currently in. */
+struct cleanup_cursor {
+ u32 start, end; /* [start, end) of the current subprogram */
+ int sub; /* its index */
+};
+
+#define CLEANUP_CURSOR_INIT { .sub = -1 }
+
+static void cleanup_cursor_to(struct bpf_verifier_env *env, struct cleanup_cursor *c, u32 i)
+{
+ while (i >= c->end) {
+ c->sub++;
+ c->start = env->subprog_info[c->sub].start;
+ c->end = env->subprog_info[c->sub + 1].start;
+ }
+}
+
+static enum cleanup_insn_kind cleanup_classify(struct bpf_verifier_env *env, u32 i,
+ int *next, int *target)
+{
+ struct bpf_insn *insn = &env->prog->insnsi[i];
+ u8 class = BPF_CLASS(insn->code);
+
+ *next = i + 1;
+ *target = -1;
+
+ if (insn->code == (BPF_LD | BPF_IMM | BPF_DW)) {
+ *next = i + 2;
+ return CLEANUP_INSN_PLAIN;
+ }
+ if (class != BPF_JMP && class != BPF_JMP32)
+ return CLEANUP_INSN_PLAIN;
+
+ switch (BPF_OP(insn->code)) {
+ case BPF_EXIT:
+ *next = -1;
+ return CLEANUP_INSN_EXIT;
+ case BPF_JA:
+ *next = -1;
+ if (BPF_SRC(insn->code) == BPF_X)
+ return CLEANUP_INSN_GOTOX;
+ *target = class == BPF_JMP32 ? i + insn->imm + 1 : i + insn->off + 1;
+ return CLEANUP_INSN_JUMP;
+ case BPF_CALL:
+ if (bpf_is_throw_kfunc(insn)) {
+ *next = -1;
+ return CLEANUP_INSN_THROW;
+ }
+ if (insn_is_exc_kfunc(insn, EXC_KF_bpf_unwind_resume)) {
+ *next = -1;
+ return CLEANUP_INSN_RESUME;
+ }
+ return bpf_pseudo_call(insn) ? CLEANUP_INSN_CALL : CLEANUP_INSN_PLAIN;
+ default:
+ /* Conditional jump, including BPF_JCOND. */
+ *target = i + insn->off + 1;
+ return CLEANUP_INSN_COND;
+ }
+}
+
static void cleanup_mark_throw_sites(struct bpf_verifier_env *env)
{
u32 i;
@@ -32,6 +157,247 @@ static void cleanup_mark_throw_sites(struct bpf_verifier_env *env)
env->insn_aux_data[i].cleanup_throw_site = true;
}
+int bpf_cleanup_check_callback(struct bpf_verifier_env *env, int subprog)
+{
+ if (!env->cleanup_info_cnt || !env->subprog_info[subprog].might_throw)
+ return 0;
+
+ verbose(env, "subprog %d may unwind and is used as a callback\n", subprog);
+ return -EINVAL;
+}
+
+/* Intra-subprog successors of @i, or -1 each when absent. */
+static enum cleanup_insn_kind cleanup_succ(struct bpf_verifier_env *env, u32 i,
+ u32 start, u32 end, int *next, int *target)
+{
+ enum cleanup_insn_kind kind = cleanup_classify(env, i, next, target);
+
+ if (*next < (int)start || *next >= (int)end)
+ *next = -1;
+ if (*target < (int)start || *target >= (int)end)
+ *target = -1;
+ return kind;
+}
+
+static void cleanup_add_pred(u32 *head, u32 *link, u32 to, u32 e)
+{
+ link[e] = head[to];
+ head[to] = e + 1;
+}
+
+/* What every instruction can reach along intra-subprog edges, for
+ * cleanup_pad_is_catch(). One backward walk over a predecessor index, rather
+ * than a forward walk from each landing pad, which would be quadratic.
+ */
+static int cleanup_reachability(struct cleanup_ctx *ctx)
+{
+ struct bpf_verifier_env *env = ctx->env;
+ u32 len = env->prog->len;
+ struct cleanup_cursor c = CLEANUP_CURSOR_INIT;
+ u32 *head = NULL, *link = NULL;
+ bool *queued = NULL;
+ u32 i, sp = 0;
+ void *scratch;
+ const struct cleanup_alloc_req tab[] = {
+ { (void **)&head, len, sizeof(*head) },
+ { (void **)&link, 2 * (size_t)len, sizeof(*link) },
+ { (void **)&queued, len, sizeof(*queued) },
+ };
+
+ scratch = cleanup_alloc(tab, ARRAY_SIZE(tab));
+ if (!scratch)
+ return -ENOMEM;
+
+ /* Index the predecessors, and seed the walk at the terminators. */
+ for (i = 0; i < len; i++) {
+ enum cleanup_insn_kind kind;
+ int next, target;
+
+ cleanup_cursor_to(env, &c, i);
+ kind = cleanup_succ(env, i, c.start, c.end, &next, &target);
+
+ if (kind == CLEANUP_INSN_RESUME)
+ ctx->reach[i] |= CLEANUP_REACH_RESUME;
+ else if (kind == CLEANUP_INSN_EXIT)
+ ctx->reach[i] |= CLEANUP_REACH_EXIT;
+ else if (kind == CLEANUP_INSN_GOTOX)
+ ctx->reach[i] |= CLEANUP_REACH_UNKNOWN;
+ else if (kind == CLEANUP_INSN_THROW)
+ ctx->reach[i] |= CLEANUP_REACH_THROW;
+
+ if (next >= 0)
+ cleanup_add_pred(head, link, next, 2 * i);
+ if (target >= 0)
+ cleanup_add_pred(head, link, target, 2 * i + 1);
+
+ if (ctx->reach[i]) {
+ queued[i] = true;
+ ctx->stack[sp++] = i;
+ }
+ }
+
+ /* Each instruction re-enters the worklist at most once per bit it
+ * gains, so this is linear in the number of edges.
+ */
+ while (sp) {
+ u32 j = ctx->stack[--sp];
+ u8 flags = ctx->reach[j];
+ u32 e;
+
+ queued[j] = false;
+ for (e = head[j]; e; e = link[e - 1]) {
+ u32 p = (e - 1) / 2;
+
+ if ((ctx->reach[p] | flags) == ctx->reach[p])
+ continue;
+ ctx->reach[p] |= flags;
+ if (!queued[p]) {
+ queued[p] = true;
+ ctx->stack[sp++] = p;
+ }
+ }
+ }
+ kvfree(scratch);
+ return 0;
+}
+
+static int cleanup_pad_is_catch(struct cleanup_ctx *ctx, u32 pad)
+{
+ u8 reach = ctx->reach[pad];
+
+ if (reach & CLEANUP_REACH_UNKNOWN) {
+ verbose(ctx->env, "cleanup landing pad %u reaches an indirect jump\n", pad);
+ return -EINVAL;
+ }
+ if (reach & CLEANUP_REACH_THROW) {
+ verbose(ctx->env,
+ "cleanup landing pad %u can throw while an exception is in flight\n",
+ pad);
+ return -EINVAL;
+ }
+ if (!(reach & CLEANUP_REACH_RESUME) == !(reach & CLEANUP_REACH_EXIT)) {
+ verbose(ctx->env, "cleanup landing pad %u %s\n", pad,
+ (reach & CLEANUP_REACH_RESUME) ?
+ "reaches both bpf_unwind_resume() and a plain exit" :
+ "reaches neither bpf_unwind_resume() nor an exit");
+ return -EINVAL;
+ }
+ return !!(reach & CLEANUP_REACH_EXIT);
+}
+
+static int cleanup_check_pad_insn(struct bpf_verifier_env *env, u32 i)
+{
+ struct bpf_insn *insn = &env->prog->insnsi[i];
+
+ if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call) {
+ verbose(env,
+ "bpf_tail_call() at insn %u is in an exception cleanup landing pad\n",
+ i);
+ return -EINVAL;
+ }
+ /* Stack arguments are not supported. */
+ if (is_stack_arg_st(insn) || is_stack_arg_stx(insn)) {
+ verbose(env,
+ "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
+ i);
+ return -EINVAL;
+ }
+ /* Likewise, stack arguments are not supported. */
+ if (bpf_pseudo_kfunc_call(insn)) {
+ struct bpf_call_summary cs;
+
+ if (bpf_get_call_summary(env, insn, &cs) &&
+ cs.arg_slot_cnt > MAX_BPF_FUNC_REG_ARGS) {
+ verbose(env,
+ "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
+ i);
+ return -EINVAL;
+ }
+ }
+ return 0;
+}
+
+static int cleanup_mark_pad_bodies(struct cleanup_ctx *ctx)
+{
+ struct bpf_verifier_env *env = ctx->env;
+ u32 i, sp = 0;
+ int ret;
+
+ for (i = 0; i < env->cleanup_info_cnt; i++) {
+ u32 pad = env->cleanup_info[i].landing_pad_off;
+
+ if (in_pad(env, pad))
+ continue;
+
+ ret = cleanup_pad_is_catch(ctx, pad);
+ if (ret < 0)
+ return ret;
+ if (ret) {
+ verbose(env,
+ "catch landing pad %u is not supported yet, only cleanup pads that resume\n",
+ pad);
+ return -EOPNOTSUPP;
+ }
+ env->insn_aux_data[pad].in_cleanup_pad = true;
+ ctx->stack[sp++] = pad;
+ }
+
+ while (sp) {
+ u32 j = ctx->stack[--sp];
+ enum cleanup_insn_kind kind;
+ int next, target, sub;
+ u32 start, end;
+
+ ret = cleanup_check_pad_insn(env, j);
+ if (ret)
+ return ret;
+
+ sub = cleanup_subprog_of(env, j);
+ start = env->subprog_info[sub].start;
+ end = env->subprog_info[sub + 1].start;
+ kind = cleanup_succ(env, j, start, end, &next, &target);
+
+ if (kind == CLEANUP_INSN_CALL) {
+ int callee = cleanup_subprog_of(env, j + env->prog->insnsi[j].imm + 1);
+
+ if (env->subprog_info[callee].might_throw) {
+ verbose(env,
+ "cleanup landing pad calls subprog %d at insn %u, which can throw while an exception is in flight\n",
+ callee, j);
+ return -EINVAL;
+ }
+ }
+
+ if (next >= 0 && !in_pad(env, next)) {
+ env->insn_aux_data[next].in_cleanup_pad = true;
+ ctx->stack[sp++] = next;
+ }
+ if (target >= 0 && !in_pad(env, target)) {
+ env->insn_aux_data[target].in_cleanup_pad = true;
+ ctx->stack[sp++] = target;
+ }
+ }
+ return 0;
+}
+
+static int cleanup_check_resumes(struct cleanup_ctx *ctx)
+{
+ struct bpf_verifier_env *env = ctx->env;
+ u32 i;
+
+ for (i = 0; i < env->prog->len; i++) {
+ if (!insn_is_exc_kfunc(&env->prog->insnsi[i], EXC_KF_bpf_unwind_resume))
+ continue;
+ if (in_pad(env, i))
+ continue;
+ verbose(env,
+ "bpf_unwind_resume() at insn %u is not in an exception cleanup landing pad\n",
+ i);
+ return -EINVAL;
+ }
+ return 0;
+}
+
static void cleanup_mark_call_sites(struct bpf_verifier_env *env)
{
u32 i, j;
@@ -78,6 +444,37 @@ int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env)
return 0;
}
+int bpf_check_cleanup_exceptions(struct bpf_verifier_env *env)
+{
+ u32 len = env->prog->len;
+ struct cleanup_ctx ctx = { .env = env };
+ const struct cleanup_alloc_req tab[] = {
+ { (void **)&ctx.reach, len, sizeof(*ctx.reach) },
+ { (void **)&ctx.stack, len, sizeof(*ctx.stack) },
+ };
+ int ret;
+
+ if (!env->cleanup_info_cnt)
+ return 0;
+
+ ctx.scratch = cleanup_alloc(tab, ARRAY_SIZE(tab));
+ if (!ctx.scratch)
+ return -ENOMEM;
+
+ ret = cleanup_reachability(&ctx);
+ if (ret)
+ goto out;
+
+ ret = cleanup_mark_pad_bodies(&ctx);
+ if (ret)
+ goto out;
+
+ ret = cleanup_check_resumes(&ctx);
+out:
+ kvfree(ctx.scratch);
+ return ret;
+}
+
bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
{
return insn_is_exc_kfunc(insn, EXC_KF_bpf_unwind_resume);
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index f51383fd775c..7313dd2b65a1 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -8,6 +8,8 @@
struct bpf_verifier_env;
int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
+int bpf_check_cleanup_exceptions(struct bpf_verifier_env *env);
+int bpf_cleanup_check_callback(struct bpf_verifier_env *env, int subprog);
int bpf_cleanup_pad_of_call(struct bpf_verifier_env *env, u32 idx);
#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 82b00fac6bd6..f120ae66dbbc 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -252,6 +252,7 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
/* Expand insni[off]'s seen count to the patched range. */
data[i].seen = old_seen;
data[i].zext_dst = bpf_insn_def32(new_prog, insn + i) >= 0;
+ data[i].in_cleanup_pad = data[off + cnt - 1].in_cleanup_pad;
if (!memcmp(insn + i, original_insn, sizeof(struct bpf_insn))) {
data[i].non_stack_access =
data[off + cnt - 1].non_stack_access;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 6496c30a10f7..9cbdb8339701 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10485,6 +10485,10 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
* callbacks
*/
env->subprog_info[subprog].is_cb = true;
+ err = bpf_cleanup_check_callback(env, subprog);
+ if (err)
+ return err;
+
if (bpf_pseudo_kfunc_call(insn) &&
!is_callback_calling_kfunc(insn->imm)) {
verifier_bug(env, "kfunc %s#%d not marked as callback-calling",
@@ -21664,6 +21668,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
if (ret < 0)
goto skip_full_check;
+ ret = bpf_check_cleanup_exceptions(env);
+ if (ret < 0)
+ goto skip_full_check;
+
ret = bpf_compute_postorder(env);
if (ret < 0)
goto skip_full_check;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 08/20] bpf: Walk the exception unwind in the verifier
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (6 preceding siblings ...)
2026-09-18 4:42 ` [PATCH bpf-next v2 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch Yonghong Song
@ 2026-09-18 4:42 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 09/20] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
` (11 subsequent siblings)
19 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:42 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
bpf_throw() is about to stop discarding the frames it unwinds and start
running their landing pads instead. Have the verifier walk the same thing,
step for step:
at a throw, with the frame chain in env->cur_state->frame[]
if a record covers the call control is at
run that record's landing pad in this frame
on reaching its resume, pop the frame and carry on
else
pop the frame and carry on
at the boundary, deliver
Doing it this way is what keeps the resource rules unchanged. Whatever a
pad releases is released in the verifier state too, so by the time the walk
reaches the boundary the state says exactly what the program will really
hold there -- and check_resource_leak(), which used to fire the moment a
throw was seen, simply moves to the end of the walk. A frame that no record
covers contributes nothing, so anything it held is still held when the walk
ends, and that is what gets reported.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_cleanup_abi.h | 16 +++++
include/linux/bpf_verifier.h | 1 +
kernel/bpf/states.c | 3 +
kernel/bpf/verifier.c | 102 +++++++++++++++++++++++++-------
4 files changed, 99 insertions(+), 23 deletions(-)
create mode 100644 include/linux/bpf_cleanup_abi.h
diff --git a/include/linux/bpf_cleanup_abi.h b/include/linux/bpf_cleanup_abi.h
new file mode 100644
index 000000000000..b6c1d589abda
--- /dev/null
+++ b/include/linux/bpf_cleanup_abi.h
@@ -0,0 +1,16 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef _LINUX_BPF_CLEANUP_ABI_H
+#define _LINUX_BPF_CLEANUP_ABI_H
+
+/*
+ * Value arch_bpf_run_cleanup_pad() leaves in r0 on the way into a landing pad.
+ * It has to be a constant the verifier knows: LLVM names r0 as both the
+ * exception pointer and the exception selector register, so every pad reads it
+ * before anything else and is free to store what it read. Kept on its own
+ * because the verifier and the arch dispatchers, which are assembly, have to
+ * agree on it.
+ */
+#define BPF_PAD_ENTRY_R0 1
+
+#endif /* _LINUX_BPF_CLEANUP_ABI_H */
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index fe8b26351a15..09fb89fda40a 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -509,6 +509,7 @@ struct bpf_verifier_state {
bool speculative;
bool in_sleepable;
+ bool unwinding;
/* first and last insn idx of this verifier state */
u32 first_insn_idx;
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 66fb11b6c6a7..be0f529f7ecc 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -996,6 +996,9 @@ static bool states_equal(struct bpf_verifier_env *env,
if (old->in_sleepable != cur->in_sleepable)
return false;
+ if (old->unwinding != cur->unwinding)
+ return false;
+
if (!refsafe(old, cur, &env->idmap_scratch))
return false;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 9cbdb8339701..a3b34ded1392 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10,6 +10,7 @@
#include <linux/slab.h>
#include <linux/bpf.h>
#include <linux/btf.h>
+#include <linux/bpf_cleanup_abi.h>
#include <linux/bpf_verifier.h>
#include <linux/filter.h>
#include <net/netlink.h>
@@ -1715,6 +1716,7 @@ int bpf_copy_verifier_state(struct bpf_verifier_state *dst_state,
return err;
dst_state->speculative = src->speculative;
dst_state->in_sleepable = src->in_sleepable;
+ dst_state->unwinding = src->unwinding;
dst_state->curframe = src->curframe;
dst_state->branches = src->branches;
dst_state->parent = src->parent;
@@ -10540,8 +10542,8 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
return 0;
}
-static int process_bpf_exit_full(struct bpf_verifier_env *env,
- bool *do_print_state, bool exception_exit);
+static int process_bpf_exit_full(struct bpf_verifier_env *env, bool *do_print_state);
+static int unwind_step(struct bpf_verifier_env *env, u32 callsite, int *insn_idx);
static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
int *insn_idx)
@@ -10631,7 +10633,7 @@ static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
verbose(env, "failed to push state for global subprog exception path\n");
return PTR_ERR(branch);
}
- return process_bpf_exit_full(env, NULL, true);
+ return unwind_step(env, *insn_idx, insn_idx);
}
/* continue with next insn after call */
@@ -14511,7 +14513,7 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
env->prog->call_session_cookie = true;
if (bpf_is_throw_kfunc(insn))
- return process_bpf_exit_full(env, NULL, true);
+ return unwind_step(env, insn_idx, &env->insn_idx);
return 0;
}
@@ -18436,9 +18438,75 @@ enum {
INSN_IDX_UPDATED = 2,
};
-static int process_bpf_exit_full(struct bpf_verifier_env *env,
- bool *do_print_state,
- bool exception_exit)
+static u32 unwind_pop_frame(struct bpf_verifier_env *env)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ struct bpf_func_state *callee = state->frame[state->curframe];
+ u32 callsite = callee->callsite;
+ struct bpf_func_state *caller;
+
+ caller = state->frame[state->curframe - 1];
+ account_processed_insns(env, callee, caller);
+ free_func_state(callee);
+ state->frame[state->curframe--] = NULL;
+ invalidate_outgoing_stack_args(env, caller);
+ return callsite;
+}
+
+static void unwind_enter_pad(struct bpf_verifier_env *env)
+{
+ struct bpf_func_state *frame = cur_func(env);
+
+ clear_caller_saved_regs(env, frame->regs);
+ mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ __mark_reg_known(&frame->regs[BPF_REG_0], BPF_PAD_ENTRY_R0);
+}
+
+static int unwind_finish(struct bpf_verifier_env *env)
+{
+ int err = check_resource_leak(env, true, true, "bpf_throw");
+
+ if (err)
+ return err;
+ return PROCESS_BPF_EXIT;
+}
+
+static int unwind_step(struct bpf_verifier_env *env, u32 callsite, int *insn_idx)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+
+ state->unwinding = true;
+ for (;;) {
+ int pad = bpf_cleanup_pad_of_call(env, callsite);
+
+ if (pad >= 0) {
+ unwind_enter_pad(env);
+ *insn_idx = pad;
+ return INSN_IDX_UPDATED;
+ }
+ if (!state->curframe)
+ return unwind_finish(env);
+ callsite = unwind_pop_frame(env);
+ }
+}
+
+static int process_cleanup_resume(struct bpf_verifier_env *env, int *insn_idx)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+
+ /* A pad entered by ordinary control flow. */
+ if (!state->unwinding) {
+ verbose(env,
+ "bpf_unwind_resume() at insn %d reached without an exception in flight\n",
+ *insn_idx);
+ return -EINVAL;
+ }
+ if (!state->curframe)
+ return unwind_finish(env);
+ return unwind_step(env, unwind_pop_frame(env), insn_idx);
+}
+
+static int process_bpf_exit_full(struct bpf_verifier_env *env, bool *do_print_state)
{
struct bpf_func_state *cur_frame = cur_func(env);
@@ -18448,25 +18516,11 @@ static int process_bpf_exit_full(struct bpf_verifier_env *env,
* for which reference_state must match caller reference
* state when it exits.
*/
- int err = check_resource_leak(env, exception_exit,
- exception_exit || !env->cur_state->curframe,
- exception_exit ? "bpf_throw" :
+ int err = check_resource_leak(env, false, !env->cur_state->curframe,
"BPF_EXIT instruction in main prog");
if (err)
return err;
- /* The side effect of the prepare_func_exit which is
- * being skipped is that it frees bpf_func_state.
- * Typically, process_bpf_exit will only be hit with
- * outermost exit. copy_verifier_state in pop_stack will
- * handle freeing of any extra bpf_func_state left over
- * from not processing all nested function exits. We
- * also skip return code checks as they are not needed
- * for exceptional exits.
- */
- if (exception_exit)
- return PROCESS_BPF_EXIT;
-
if (env->cur_state->curframe) {
/* exit from nested function */
err = prepare_func_exit(env, &env->insn_idx);
@@ -18640,6 +18694,8 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
env->jmps_processed++;
if (opcode == BPF_CALL) {
+ if (bpf_is_unwind_resume_kfunc(insn))
+ return process_cleanup_resume(env, &env->insn_idx);
if (env->cur_state->active_locks) {
if ((insn->src_reg == BPF_REG_0 &&
insn->imm != BPF_FUNC_spin_unlock &&
@@ -18673,7 +18729,7 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
env->insn_idx += insn->imm + 1;
return INSN_IDX_UPDATED;
} else if (opcode == BPF_EXIT) {
- return process_bpf_exit_full(env, do_print_state, false);
+ return process_bpf_exit_full(env, do_print_state);
}
return check_cond_jmp_op(env, insn, &env->insn_idx);
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 09/20] bpf: Refuse a private stack for a program with an exception cleanup table
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (7 preceding siblings ...)
2026-09-18 4:42 ` [PATCH bpf-next v2 08/20] bpf: Walk the exception unwind in the verifier Yonghong Song
@ 2026-09-18 4:42 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-18 4:42 ` [PATCH bpf-next v2 10/20] bpf: Dispatch exception cleanup pads from bpf_throw() Yonghong Song
` (10 subsequent siblings)
19 siblings, 1 reply; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:42 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A landing pad is about to start running outside its own frame, on the
bpf_throw() walker's stack, with that frame's registers put back from a
spill area whose contents arch_bpf_run_cleanup_pad() knows how to read. The
x86-64 JIT addresses a private-stack program's frame through a scratch
register that it recomputes after each call rather than through rbp, and no
spill area holds that register, so a pad there would address its frame
through whatever the kernel left behind.
Refuse the combination in check_max_stack_depth(), where the choice is
made. Everywhere rather than arch-conditionally: arm64 keeps its private
stack pointer in x27, which is in the prologue spill and so survives, but a
rule that holds on one arch and not the other is not worth the second code
path when nothing is lost but an optimization.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/verifier.c | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index a3b34ded1392..43ecf79baa4a 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -5591,6 +5591,14 @@ static int check_max_stack_depth(struct bpf_verifier_env *env)
}
}
+ /*
+ * A pad rebuilds its frame from a spill area, and on x86-64 a private
+ * stack's frame pointer is in no spill area. Refused on every arch
+ * rather than just that one.
+ */
+ if (env->cleanup_info_cnt)
+ priv_stack_mode = NO_PRIV_STACK;
+
if (priv_stack_mode == PRIV_STACK_UNKNOWN)
priv_stack_mode = bpf_enable_priv_stack(env->prog);
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 10/20] bpf: Dispatch exception cleanup pads from bpf_throw()
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (8 preceding siblings ...)
2026-09-18 4:42 ` [PATCH bpf-next v2 09/20] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
@ 2026-09-18 4:42 ` Yonghong Song
2026-09-18 5:58 ` bot+bpf-ci
2026-09-18 4:42 ` [PATCH bpf-next v2 11/20] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
` (9 subsequent siblings)
19 siblings, 1 reply; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:42 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Stop discarding the frames an exception unwinds through. bpf_throw()
already walks the BPF call stack with arch_bpf_stack_walk() to find the
exception boundary; have it look each frame's return address up in that
(sub)program's cleanup table on the way and run the landing pad a matching
record names.
A pad is run as a subroutine of the walker, not jumped to. It executes with
the unwinding frame's frame pointer and BPF callee-saved registers, so
everything it reads is that frame's, but on the current stack far below it,
so nothing it calls can disturb the frame it is cleaning up after. It ends
in what the JIT emits for its bpf_unwind_resume(), which hands control back
to the walker rather than to the frame's caller. That is exactly the
procedure the LLVM commit emitting the section describes.
Restoring the frame's r6-r9 is what makes this work, and it is only
possible because the callee about to be discarded spilled them in its own
prologue. bpf_cleanup_force_spill() tells a JIT to make that spill
unconditional and of a known shape for every subprogram of a program
carrying a cleanup table, and aux->exc->spill_off records where it starts,
so the walker needs no per-frame metadata. The frame that called
bpf_throw() has no callee to have spilled anything and never runs its own
epilogue, so the JIT spills that frame's registers at the throw site
instead, in the area aux->exc->throw_spill_off names.
The main program needs the table handed to it rather than built for it.
jit_subprogs() compiles it as func[0], but the ksym covering that image is
the one bpf_prog_load() registers for the outer bpf_prog, and that is what
the walker finds -- so without the handover a landing pad in the main
program's own frame is never dispatched, silently, since the exception
still reaches the boundary and the cookie still comes back.
All of it hangs off bpf_prog_aux by a single pointer. bpf_prog_aux is on
every BPF program and this is a niche feature, so struct bpf_exception_info
exists only for a program that carries a table, and aux->exc being set is
what says the program carries one at all.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf.h | 75 +++++++++++++++++++
include/linux/filter.h | 1 +
kernel/bpf/core.c | 30 +++++++-
kernel/bpf/exception.c | 166 +++++++++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 7 ++
kernel/bpf/fixups.c | 125 +++++++++++++++++++++++++++++++
kernel/bpf/helpers.c | 34 +++++++++
7 files changed, 435 insertions(+), 3 deletions(-)
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 2a5fa346aada..83f2b0d7e596 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -1770,6 +1770,80 @@ enum bpf_sig_keyring {
BPF_SIG_KEYRING_BPF,
};
+/*
+ * One cleanup region of a JITed (sub)program: @pad is the landing pad to run
+ * for a return address in (begin, end], the native code of its call sites.
+ */
+struct bpf_cleanup_range {
+ u64 begin;
+ u64 end;
+ u64 pad;
+};
+
+struct bpf_exception_info {
+ struct bpf_cleanup_info *info;
+ struct bpf_cleanup_range *ranges;
+ /* Landing pad instruction indices, sorted and deduplicated. */
+ u32 *pad_at;
+ /* bpf_throw() call instruction indices, sorted. */
+ u32 *throw_at;
+ /* One bit per instruction that only runs while unwinding. */
+ unsigned long *pad_body;
+ u32 nr_info;
+ u32 nr_ranges;
+ u32 nr_pad_at;
+ u32 nr_throw_at;
+ u32 nr_pad_body;
+ /* Offset from a frame's FP to the caller's spilled r6-r9. */
+ s32 spill_off;
+ /* Likewise, to the registers a frame spills before calling bpf_throw(). */
+ s32 throw_spill_off;
+};
+
+#ifdef CONFIG_BPF_SYSCALL
+bool bpf_cleanup_force_spill(const struct bpf_prog *prog);
+bool bpf_cleanup_insn_is_pad(const struct bpf_prog *prog, u32 idx);
+bool bpf_cleanup_insn_in_pad(const struct bpf_prog *prog, u32 idx);
+bool bpf_cleanup_insn_is_throw(const struct bpf_prog *prog, u32 idx);
+int bpf_cleanup_attach_main_prog(struct bpf_verifier_env *env, struct bpf_prog *prog);
+void bpf_cleanup_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image);
+void bpf_cleanup_free_info(struct bpf_prog_aux *aux);
+#else
+static inline bool bpf_cleanup_force_spill(const struct bpf_prog *prog)
+{
+ return false;
+}
+
+static inline bool bpf_cleanup_insn_is_pad(const struct bpf_prog *prog, u32 idx)
+{
+ return false;
+}
+
+static inline bool bpf_cleanup_insn_in_pad(const struct bpf_prog *prog, u32 idx)
+{
+ return false;
+}
+
+static inline bool bpf_cleanup_insn_is_throw(const struct bpf_prog *prog, u32 idx)
+{
+ return false;
+}
+
+static inline int bpf_cleanup_attach_main_prog(struct bpf_verifier_env *env,
+ struct bpf_prog *prog)
+{
+ return 0;
+}
+
+static inline void bpf_cleanup_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+}
+
+static inline void bpf_cleanup_free_info(struct bpf_prog_aux *aux)
+{
+}
+#endif
+
struct bpf_prog_aux {
atomic64_t refcnt;
u32 used_map_cnt;
@@ -1850,6 +1924,7 @@ struct bpf_prog_aux {
char name[BPF_OBJ_NAME_LEN];
u64 (*bpf_exception_cb)(u64 cookie, u64 sp, u64 bp, u64, u64);
u16 stack_arg_sp_adjust;
+ struct bpf_exception_info *exc;
#ifdef CONFIG_SECURITY
void *security;
#endif
diff --git a/include/linux/filter.h b/include/linux/filter.h
index 9682b98ad890..287cd9b59aa9 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1243,6 +1243,7 @@ bool bpf_jit_supports_arena_args(void);
bool bpf_jit_supports_far_kfunc_call(void);
bool bpf_jit_supports_exceptions(void);
bool bpf_jit_supports_cleanup_pads(void);
+void arch_bpf_run_cleanup_pad(u64 pad, u64 frame_fp, u64 spill_base);
bool bpf_jit_supports_ptr_xchg(void);
bool bpf_jit_supports_arena(void);
bool bpf_jit_supports_insn(struct bpf_insn *insn, bool in_arena);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index f0dd851264f4..bd2919063cec 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -292,6 +292,7 @@ void __bpf_prog_free(struct bpf_prog *fp)
mutex_destroy(&fp->aux->dst_mutex);
mutex_destroy(&fp->aux->st_ops_assoc_mutex);
kfree(fp->aux->poke_tab);
+ bpf_cleanup_free_info(fp->aux);
kfree(fp->aux);
}
free_percpu(fp->stats);
@@ -2625,13 +2626,21 @@ static bool bpf_prog_select_interpreter(struct bpf_prog *fp)
return select_interpreter;
}
-static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struct bpf_prog *prog)
+static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struct bpf_prog *prog,
+ int *err)
{
#ifdef CONFIG_BPF_JIT
struct bpf_prog *orig_prog;
+ int ret;
- if (!bpf_prog_need_blind(prog))
+ if (!bpf_prog_need_blind(prog)) {
+ ret = bpf_cleanup_attach_main_prog(env, prog);
+ if (ret) {
+ *err = ret;
+ return prog;
+ }
return bpf_int_jit_compile(env, prog);
+ }
orig_prog = prog;
prog = bpf_jit_blind_constants(env, prog);
@@ -2642,6 +2651,13 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
if (IS_ERR(prog))
goto out_restore;
+ ret = bpf_cleanup_attach_main_prog(env, prog);
+ if (ret) {
+ *err = ret;
+ bpf_jit_prog_release_other(orig_prog, prog);
+ goto out_restore;
+ }
+
prog = bpf_int_jit_compile(env, prog);
if (prog->jited) {
bpf_jit_prog_release_other(prog, orig_prog);
@@ -2681,8 +2697,10 @@ struct bpf_prog *__bpf_prog_select_runtime(struct bpf_verifier_env *env, struct
if (*err)
return fp;
- fp = bpf_prog_jit_compile(env, fp);
+ fp = bpf_prog_jit_compile(env, fp, err);
bpf_prog_jit_attempt_done(fp);
+ if (*err)
+ return fp;
if (!fp->jited && jit_needed) {
*err = -ENOTSUPP;
return fp;
@@ -3475,6 +3493,12 @@ bool __weak bpf_jit_supports_cleanup_pads(void)
return false;
}
+/* Call @pad with the frame pointer @frame_fp and r6-r9 spilled at @spill_base. */
+void __weak arch_bpf_run_cleanup_pad(u64 pad, u64 frame_fp, u64 spill_base)
+{
+ WARN_ON_ONCE(1);
+}
+
bool __weak bpf_jit_supports_timed_may_goto(void)
{
return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index b2bf831242b8..521086d084a3 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -1,5 +1,6 @@
// SPDX-License-Identifier: GPL-2.0-only
/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bitmap.h>
#include <linux/bpf.h>
#include <linux/bpf_verifier.h>
#include <linux/btf.h>
@@ -486,3 +487,168 @@ int bpf_cleanup_pad_of_call(struct bpf_verifier_env *env, u32 idx)
return pad ? (int)pad - 1 : -1;
}
+
+/*
+ * Every subprogram of a cleanup-carrying program spills the BPF callee-saved
+ * registers, even one that never throws: a frame's spill holds its caller's
+ * registers, and that is what the walker restores before running the caller's
+ * pad. The exception callback does not, because it reuses the boundary frame
+ * rather than building one of its own.
+ */
+bool bpf_cleanup_force_spill(const struct bpf_prog *prog)
+{
+ return prog->aux->exc && !prog->aux->exception_cb;
+}
+
+const struct bpf_cleanup_range *bpf_cleanup_pad_for_ip(const struct bpf_prog *prog, u64 ip)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+ u32 l = 0, r = exc ? exc->nr_ranges : 0;
+
+ while (l < r) {
+ u32 m = l + (r - l) / 2;
+ const struct bpf_cleanup_range *rec = &exc->ranges[m];
+
+ if (ip <= rec->begin)
+ r = m;
+ else if (ip > rec->end)
+ l = m + 1;
+ else
+ return rec;
+ }
+ return NULL;
+}
+
+static int cmp_u32(const void *a, const void *b)
+{
+ u32 x = *(const u32 *)a, y = *(const u32 *)b;
+
+ return x < y ? -1 : x > y;
+}
+
+int bpf_cleanup_alloc_info(struct bpf_prog_aux *aux)
+{
+ if (aux->exc)
+ return 0;
+ aux->exc = kzalloc_obj(struct bpf_exception_info, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ return aux->exc ? 0 : -ENOMEM;
+}
+
+int bpf_cleanup_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt)
+{
+ struct bpf_exception_info *exc = aux->exc;
+ struct bpf_cleanup_range *ranges;
+ u32 i, n_at, *at;
+
+ if (!cnt) {
+ kvfree(recs);
+ return 0;
+ }
+
+ ranges = kvcalloc(cnt, sizeof(*ranges), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!ranges) {
+ kvfree(recs);
+ return -ENOMEM;
+ }
+
+ /* The pads on their own, sorted and deduplicated. */
+ at = kvmalloc_array(cnt, sizeof(*at), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!at) {
+ kvfree(ranges);
+ kvfree(recs);
+ return -ENOMEM;
+ }
+ for (i = 0; i < cnt; i++)
+ at[i] = recs[i].landing_pad_off;
+ sort(at, cnt, sizeof(*at), cmp_u32, NULL);
+ for (i = 0, n_at = 0; i < cnt; i++)
+ if (!n_at || at[n_at - 1] != at[i])
+ at[n_at++] = at[i];
+
+ exc->pad_at = at;
+ exc->nr_pad_at = n_at;
+ exc->info = recs;
+ exc->nr_info = cnt;
+ exc->ranges = ranges;
+ /* Withheld until the JIT has filled the table in. */
+ exc->nr_ranges = 0;
+ return 0;
+}
+
+void bpf_cleanup_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+ struct bpf_exception_info *exc = prog->aux->exc;
+ u32 i, n;
+
+ if (!exc || !exc->nr_info || !exc->ranges)
+ return;
+
+ n = exc->nr_info;
+ for (i = 0; i < n; i++) {
+ const struct bpf_cleanup_info *rec = &exc->info[i];
+
+ if (WARN_ON_ONCE(rec->begin_off >= prog->len ||
+ rec->end_off > prog->len ||
+ rec->landing_pad_off >= prog->len))
+ return;
+ exc->ranges[i].begin = (u64)(long)image + addrs[rec->begin_off];
+ exc->ranges[i].end = (u64)(long)image + addrs[rec->end_off];
+ exc->ranges[i].pad = (u64)(long)image + addrs[rec->landing_pad_off];
+ }
+ exc->nr_ranges = n;
+}
+
+void bpf_cleanup_free_info(struct bpf_prog_aux *aux)
+{
+ struct bpf_exception_info *exc = aux->exc;
+
+ if (!exc)
+ return;
+ kvfree(exc->ranges);
+ kvfree(exc->info);
+ kvfree(exc->pad_at);
+ kvfree(exc->throw_at);
+ bitmap_free(exc->pad_body);
+ kfree(exc);
+ aux->exc = NULL;
+}
+
+/* Is @idx in the sorted array @at of @n instruction indices? */
+static bool insn_idx_in(const u32 *at, u32 n, u32 idx)
+{
+ u32 l = 0, r = n;
+
+ while (l < r) {
+ u32 m = l + (r - l) / 2;
+
+ if (idx < at[m])
+ r = m;
+ else if (idx > at[m])
+ l = m + 1;
+ else
+ return true;
+ }
+ return false;
+}
+
+bool bpf_cleanup_insn_is_pad(const struct bpf_prog *prog, u32 idx)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+
+ return exc && insn_idx_in(exc->pad_at, exc->nr_pad_at, idx);
+}
+
+bool bpf_cleanup_insn_is_throw(const struct bpf_prog *prog, u32 idx)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+
+ return exc && insn_idx_in(exc->throw_at, exc->nr_throw_at, idx);
+}
+
+bool bpf_cleanup_insn_in_pad(const struct bpf_prog *prog, u32 idx)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+
+ return exc && exc->pad_body && idx < exc->nr_pad_body &&
+ test_bit(idx, exc->pad_body);
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index 7313dd2b65a1..c0e68ce227c8 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -5,11 +5,18 @@
#include <linux/types.h>
+struct bpf_cleanup_info;
+struct bpf_cleanup_range;
+struct bpf_prog;
+struct bpf_prog_aux;
struct bpf_verifier_env;
int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
int bpf_check_cleanup_exceptions(struct bpf_verifier_env *env);
int bpf_cleanup_check_callback(struct bpf_verifier_env *env, int subprog);
int bpf_cleanup_pad_of_call(struct bpf_verifier_env *env, u32 idx);
+int bpf_cleanup_alloc_info(struct bpf_prog_aux *aux);
+int bpf_cleanup_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt);
+const struct bpf_cleanup_range *bpf_cleanup_pad_for_ip(const struct bpf_prog *prog, u64 ip);
#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index f120ae66dbbc..134aafa6a6c9 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -1,5 +1,6 @@
// SPDX-License-Identifier: GPL-2.0-only
/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bitmap.h>
#include <linux/bpf.h>
#include <linux/btf.h>
#include <linux/bpf_verifier.h>
@@ -10,6 +11,7 @@
#include <linux/perf_event.h>
#include <net/xdp.h>
#include "disasm.h"
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -257,6 +259,11 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
data[i].non_stack_access =
data[off + cnt - 1].non_stack_access;
data[off + cnt - 1].non_stack_access = false;
+ data[i].cleanup_throw_site =
+ data[off + cnt - 1].cleanup_throw_site;
+ data[off + cnt - 1].cleanup_throw_site = false;
+ data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
+ data[off + cnt - 1].cleanup_pad = 0;
} else if (bpf_is_mem_insn(insn + i)) {
data[i].non_stack_access = true;
}
@@ -1113,6 +1120,116 @@ static void bpf_restore_subprog_starts(struct bpf_verifier_env *env, u32 *orig_s
env->subprog_info[env->subprog_cnt].start = env->prog->len;
}
+static int cleanup_throw_sites_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
+ u32 start, u32 end)
+{
+ u32 i, cnt = 0, *at;
+
+ for (i = start; i < end; i++)
+ if (env->insn_aux_data[i].cleanup_throw_site)
+ cnt++;
+ if (!cnt)
+ return 0;
+
+ at = kvmalloc_array(cnt, sizeof(*at), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!at)
+ return -ENOMEM;
+
+ for (i = start, cnt = 0; i < end; i++) {
+ if (!env->insn_aux_data[i].cleanup_throw_site)
+ continue;
+ at[cnt++] = i - start;
+ }
+
+ sub->aux->exc->throw_at = at;
+ sub->aux->exc->nr_throw_at = cnt;
+ return 0;
+}
+
+static int cleanup_pad_body_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
+ u32 start, u32 end)
+{
+ unsigned long *bits;
+ u32 i, cnt = 0;
+
+ for (i = start; i < end; i++)
+ if (env->insn_aux_data[i].in_cleanup_pad)
+ cnt++;
+ if (!cnt)
+ return 0;
+
+ bits = bitmap_zalloc(end - start, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!bits)
+ return -ENOMEM;
+
+ for (i = start; i < end; i++)
+ if (env->insn_aux_data[i].in_cleanup_pad)
+ __set_bit(i - start, bits);
+
+ sub->aux->exc->pad_body = bits;
+ sub->aux->exc->nr_pad_body = end - start;
+ return 0;
+}
+
+static int cleanup_info_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
+ u32 start, u32 end)
+{
+ struct bpf_cleanup_info *recs;
+ u32 i, cnt = 0;
+ int err;
+
+ if (!env->cleanup_info_cnt)
+ return 0;
+
+ err = bpf_cleanup_alloc_info(sub->aux);
+ if (err)
+ return err;
+
+ err = cleanup_throw_sites_for_subprog(env, sub, start, end);
+ if (err)
+ return err;
+
+ err = cleanup_pad_body_for_subprog(env, sub, start, end);
+ if (err)
+ return err;
+
+ for (i = start; i < end; i++)
+ if (env->insn_aux_data[i].cleanup_pad)
+ cnt++;
+ if (!cnt)
+ return 0;
+
+ recs = kvmalloc_array(cnt, sizeof(*recs), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!recs)
+ return -ENOMEM;
+
+ for (i = start, cnt = 0; i < end; i++) {
+ u32 pad = env->insn_aux_data[i].cleanup_pad;
+
+ if (!pad)
+ continue;
+ pad--;
+ if (verifier_bug_if(pad < start || pad >= end, env,
+ "insn %u is covered by a landing pad at %u outside its subprog [%u, %u)",
+ i, pad, start, end)) {
+ kvfree(recs);
+ return -EFAULT;
+ }
+ recs[cnt].begin_off = i - start;
+ recs[cnt].end_off = i - start + 1;
+ recs[cnt].landing_pad_off = pad - start;
+ cnt++;
+ }
+ return bpf_cleanup_attach_info(sub->aux, recs, cnt);
+}
+
+int bpf_cleanup_attach_main_prog(struct bpf_verifier_env *env, struct bpf_prog *prog)
+{
+ if (!env || env->subprog_cnt > 1)
+ return 0;
+ return cleanup_info_for_subprog(env, prog, 0, prog->len);
+}
+
static int jit_subprogs(struct bpf_verifier_env *env)
{
struct bpf_prog *prog = env->prog, **func, *tmp;
@@ -1250,6 +1367,10 @@ static int jit_subprogs(struct bpf_verifier_env *env)
func[i]->aux->token = prog->aux->token;
if (!i)
func[i]->aux->exception_boundary = env->seen_exception;
+ err = cleanup_info_for_subprog(env, func[i], subprog_start,
+ env->subprog_info[i + 1].start);
+ if (err)
+ goto out_free;
func[i] = bpf_int_jit_compile(env, func[i]);
if (!func[i]->jited) {
err = -ENOTSUPP;
@@ -1354,6 +1475,8 @@ static int jit_subprogs(struct bpf_verifier_env *env)
prog->aux->bpf_exception_cb = (void *)func[env->exception_callback_subprog]->bpf_func;
prog->aux->exception_boundary = func[0]->aux->exception_boundary;
prog->aux->stack_arg_sp_adjust = func[0]->aux->stack_arg_sp_adjust;
+ prog->aux->exc = func[0]->aux->exc;
+ func[0]->aux->exc = NULL;
bpf_prog_jit_attempt_done(prog);
return 0;
out_free:
@@ -1934,6 +2057,8 @@ int bpf_do_misc_fixups(struct bpf_verifier_env *env)
goto next_insn;
if (insn->src_reg == BPF_PSEUDO_CALL)
goto next_insn;
+ if (bpf_is_unwind_resume_kfunc(insn))
+ goto next_insn;
if (insn->src_reg == BPF_PSEUDO_KFUNC_CALL) {
ret = bpf_fixup_kfunc_call(env, insn, insn_buf, i + delta, &cnt);
if (ret)
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 31d84de4d318..ffef72804fc9 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -31,6 +31,7 @@
#include <linux/buildid.h>
#include "../../lib/kstrtox.h"
+#include "exception.h"
/* If kernel subsystem is allowing eBPF programs to call this function,
* inside its own verifier_ops->get_func_proto() callback it should return
@@ -3360,8 +3361,36 @@ struct bpf_throw_ctx {
u64 sp;
u64 bp;
int cnt;
+ const struct bpf_prog *callee;
+ u64 callee_fp;
};
+static void bpf_run_cleanup_pad(struct bpf_throw_ctx *ctx, const struct bpf_prog *prog,
+ u64 ip, u64 fp)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+ const struct bpf_cleanup_range *rec;
+ u64 spill_base;
+
+ if (!exc || !exc->nr_ranges)
+ return;
+ rec = bpf_cleanup_pad_for_ip(prog, ip);
+ if (!rec)
+ return;
+
+ /*
+ * The callee is always another subprogram of this program -- the walk
+ * ends at any frame that is not one -- so its prologue spilled these
+ * registers and its exc is there to say where.
+ */
+ if (ctx->callee)
+ spill_base = ctx->callee_fp + ctx->callee->aux->exc->spill_off;
+ else
+ spill_base = fp + exc->throw_spill_off;
+
+ arch_bpf_run_cleanup_pad(rec->pad, fp, spill_base);
+}
+
static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
{
struct bpf_throw_ctx *ctx = cookie;
@@ -3378,6 +3407,11 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
if (!prog)
return !ctx->cnt;
ctx->cnt++;
+
+ bpf_run_cleanup_pad(ctx, prog, ip, bp);
+ ctx->callee = prog;
+ ctx->callee_fp = bp;
+
if (bpf_is_subprog(prog))
return true;
ctx->aux = prog->aux;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 11/20] bpf, x86: Dispatch exception cleanup pads at run time
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (9 preceding siblings ...)
2026-09-18 4:42 ` [PATCH bpf-next v2 10/20] bpf: Dispatch exception cleanup pads from bpf_throw() Yonghong Song
@ 2026-09-18 4:42 ` Yonghong Song
2026-09-18 5:03 ` sashiko-bot
2026-09-18 5:44 ` bot+bpf-ci
2026-09-18 4:43 ` [PATCH bpf-next v2 12/20] bpf, arm64: " Yonghong Song
` (8 subsequent siblings)
19 siblings, 2 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:42 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Provide the arch half for x86-64: force the full callee-saved spill for a
program carrying a cleanup table so the walker can find a frame's r6-r9 in
its callee's prologue, record where that spill area starts, build the
native cleanup table from the JIT's addrs[], emit a bare return for a pad's
bpf_unwind_resume(), and hand control to a pad from
arch_bpf_run_cleanup_pad().
Support is gated on CONFIG_UNWINDER_ORC, the same requirement
arch_bpf_stack_walk() and therefore bpf_throw() already have here.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
arch/x86/net/Makefile | 2 +-
arch/x86/net/bpf_cleanup_pad.S | 74 ++++++++++++++++++++++++++++++++
arch/x86/net/bpf_jit_comp.c | 78 ++++++++++++++++++++++++++++------
3 files changed, 141 insertions(+), 13 deletions(-)
create mode 100644 arch/x86/net/bpf_cleanup_pad.S
diff --git a/arch/x86/net/Makefile b/arch/x86/net/Makefile
index dddbefc0f439..9d574d972df3 100644
--- a/arch/x86/net/Makefile
+++ b/arch/x86/net/Makefile
@@ -6,5 +6,5 @@
ifeq ($(CONFIG_X86_32),y)
obj-$(CONFIG_BPF_JIT) += bpf_jit_comp32.o
else
- obj-$(CONFIG_BPF_JIT) += bpf_jit_comp.o bpf_timed_may_goto.o
+ obj-$(CONFIG_BPF_JIT) += bpf_jit_comp.o bpf_timed_may_goto.o bpf_cleanup_pad.o
endif
diff --git a/arch/x86/net/bpf_cleanup_pad.S b/arch/x86/net/bpf_cleanup_pad.S
new file mode 100644
index 000000000000..da4b448ecf09
--- /dev/null
+++ b/arch/x86/net/bpf_cleanup_pad.S
@@ -0,0 +1,74 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf_cleanup_abi.h>
+#include <linux/linkage.h>
+#include <asm/nospec-branch.h>
+
+/*
+ * The x86-64 BPF JIT prologue spills, once bpf_cleanup_force_spill() makes it
+ * unconditional, r12, rbx, r13, r14 and r15 in that order -- so within the
+ * spill area the lowest address holds r15 and the highest r12. The throw-site
+ * spill the JIT emits uses the same layout, so the routine below reads both
+ * the same way:
+ *
+ * spill_base + 0 BPF r9 (r15)
+ * spill_base + 8 BPF r8 (r14)
+ * spill_base + 16 BPF r7 (r13)
+ * spill_base + 24 BPF r6 (rbx)
+ * spill_base + 32 r12 (arena base, not a BPF register)
+ */
+
+ .code64
+ .section .text, "ax"
+
+/*
+ * void arch_bpf_run_cleanup_pad(u64 pad, u64 frame_fp, u64 spill_base)
+ *
+ * rdi = native address of the landing pad
+ * rsi = frame pointer of the frame the pad belongs to
+ * rdx = spill area holding that frame's BPF callee-saved registers
+ *
+ * Give the pad the register state of its own frame and call it. The pad ends
+ * in the bare return the JIT emits for its bpf_unwind_resume(), so it comes
+ * back here rather than returning to its frame's caller. It runs on this
+ * stack, far below the frame it is cleaning up after, so nothing it calls can
+ * reach into that frame.
+ */
+SYM_FUNC_START(arch_bpf_run_cleanup_pad)
+ ANNOTATE_NOENDBR
+
+ pushq %rbp
+ pushq %rbx
+ pushq %r12
+ pushq %r13
+ pushq %r14
+ pushq %r15
+ /* Keep the pad's entry rsp congruent to a normal call's. */
+ subq $8, %rsp
+
+ movq 0(%rdx), %r15
+ movq 8(%rdx), %r14
+ movq 16(%rdx), %r13
+ movq 24(%rdx), %rbx
+ movq 32(%rdx), %r12
+ /* rbp is BPF r10, so this is the whole of the pad's frame setup. */
+ movq %rsi, %rbp
+
+ /* CALL_NOSPEC needs the target in a register; rcx is BPF r4, dead. */
+ movq %rdi, %rcx
+
+ /* BPF r0 on the way into a pad, not whatever the kernel left in rax. */
+ movl $BPF_PAD_ENTRY_R0, %eax
+
+ CALL_NOSPEC rcx
+
+ addq $8, %rsp
+ popq %r15
+ popq %r14
+ popq %r13
+ popq %r12
+ popq %rbx
+ popq %rbp
+ RET
+SYM_FUNC_END(arch_bpf_run_cleanup_pad)
diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index d4a980140b48..9d0dd54773e8 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -357,6 +357,11 @@ struct jit_context {
/* Number of bytes that will be skipped on tailcall */
#define X86_TAIL_CALL_OFFSET (12 + ENDBR_INSN_SIZE)
+/* Throw-site spill: r15, r14, r13, rbx, r12 low to high, the layout the
+ * prologue's pushes leave, so arch_bpf_run_cleanup_pad() reads both alike.
+ */
+#define X86_CLEANUP_SPILL_SZ (5 * 8)
+
static void push_r9(u8 **pprog)
{
u8 *prog = *pprog;
@@ -832,7 +837,7 @@ static void emit_bpf_tail_call_indirect(struct bpf_prog *bpf_prog,
/* Inc tail_call_cnt if the slot is populated. */
EMIT4(0x48, 0x83, 0x00, 0x01); /* add qword ptr [rax], 1 */
- if (bpf_prog->aux->exception_boundary) {
+ if (bpf_prog->aux->exception_boundary || bpf_cleanup_force_spill(bpf_prog)) {
pop_callee_regs(&prog, all_callee_regs_used);
pop_r12(&prog);
} else {
@@ -899,7 +904,7 @@ static void emit_bpf_tail_call_direct(struct bpf_prog *bpf_prog,
/* Inc tail_call_cnt if the slot is populated. */
EMIT4(0x48, 0x83, 0x00, 0x01); /* add qword ptr [rax], 1 */
- if (bpf_prog->aux->exception_boundary) {
+ if (bpf_prog->aux->exception_boundary || bpf_cleanup_force_spill(bpf_prog)) {
pop_callee_regs(&prog, all_callee_regs_used);
pop_r12(&prog);
} else {
@@ -1977,6 +1982,7 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
u8 *ip, *prog = temp;
u32 stack_depth;
int callee_saved_size;
+ u32 throw_spill, prologue_depth;
s32 outgoing_arg_base;
int err;
@@ -2015,7 +2021,10 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
detect_reg_usage(insn, insn_cnt, callee_regs_used);
- emit_prologue(&prog, image, stack_depth,
+ throw_spill = bpf_cleanup_force_spill(bpf_prog) ? X86_CLEANUP_SPILL_SZ : 0;
+ prologue_depth = stack_depth + throw_spill;
+
+ emit_prologue(&prog, image, prologue_depth,
bpf_prog_was_classic(bpf_prog), tail_call_reachable,
bpf_is_subprog(bpf_prog), bpf_prog->aux->exception_cb);
@@ -2024,7 +2033,7 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
/* Exception callback will clobber callee regs for its own use, and
* restore the original callee regs from main prog's stack frame.
*/
- if (bpf_prog->aux->exception_boundary) {
+ if (bpf_prog->aux->exception_boundary || bpf_cleanup_force_spill(bpf_prog)) {
/* We also need to save r12, which is not mapped to any BPF
* register, as we throw after entry into the kernel, which may
* overwrite r12.
@@ -2039,9 +2048,10 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
/* Compute callee-saved register area size. */
callee_saved_size = 0;
- if (bpf_prog->aux->exception_boundary || arena_vm_start)
+ if (bpf_prog->aux->exception_boundary || bpf_cleanup_force_spill(bpf_prog) ||
+ arena_vm_start)
callee_saved_size += 8; /* r12 */
- if (bpf_prog->aux->exception_boundary) {
+ if (bpf_prog->aux->exception_boundary || bpf_cleanup_force_spill(bpf_prog)) {
callee_saved_size += 4 * 8; /* rbx, r13, r14, r15 */
} else {
int j;
@@ -2063,7 +2073,19 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
* Note that tail_call_reachable is guaranteed to be false when
* stack args exist, so tcc pushes need not be accounted for.
*/
- outgoing_arg_base = -(round_up(stack_depth, 8) + callee_saved_size);
+ outgoing_arg_base = -(round_up(stack_depth, 8) + throw_spill + callee_saved_size);
+
+ /*
+ * Lowest address of each spill area, as an offset from rbp; see
+ * bpf_cleanup_pad.S for the layout. The 16 is the tail call counter
+ * pair emit_prologue_tail_call() pushes above the callee-saved one.
+ */
+ if (bpf_cleanup_force_spill(bpf_prog)) {
+ bpf_prog->aux->exc->spill_off = -(round_up(stack_depth, 8) + throw_spill +
+ (tail_call_reachable ? 16 : 0) +
+ callee_saved_size);
+ bpf_prog->aux->exc->throw_spill_off = -(round_up(stack_depth, 8) + throw_spill);
+ }
/*
* Allocate outgoing stack arg area for args 7+ only.
@@ -2110,7 +2132,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
dst_reg = X86_REG_R9;
}
- if (bpf_insn_is_indirect_target(env, bpf_prog, i - 1))
+ if (bpf_insn_is_indirect_target(env, bpf_prog, i - 1) ||
+ bpf_cleanup_insn_is_pad(bpf_prog, i - 1))
EMIT_ENDBR();
ip = image + addrs[i - 1] + (prog - temp);
@@ -2903,9 +2926,27 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
case BPF_JMP | BPF_CALL: {
const struct btf_func_model *fm = NULL;
+ if (bpf_cleanup_insn_is_throw(bpf_prog, i - 1)) {
+ /* Spill r6-r9 and r12 where the bpf_throw() walker looks. */
+ s32 off = bpf_prog->aux->exc->throw_spill_off;
+ u8 *spill = prog;
+
+ emit_stx(&prog, BPF_DW, BPF_REG_FP, BPF_REG_9, off + 0);
+ emit_stx(&prog, BPF_DW, BPF_REG_FP, BPF_REG_8, off + 8);
+ emit_stx(&prog, BPF_DW, BPF_REG_FP, BPF_REG_7, off + 16);
+ emit_stx(&prog, BPF_DW, BPF_REG_FP, BPF_REG_6, off + 24);
+ emit_stx(&prog, BPF_DW, BPF_REG_FP, X86_REG_R12, off + 32);
+ ip += prog - spill;
+ }
+
+ if (bpf_is_unwind_resume_kfunc(insn)) {
+ emit_return(&prog, image + addrs[i - 1] + (prog - temp));
+ break;
+ }
+
func = (u8 *) __bpf_call_base + imm32;
if (src_reg == BPF_PSEUDO_CALL && tail_call_reachable) {
- LOAD_TAIL_CALL_CNT_PTR(stack_depth);
+ LOAD_TAIL_CALL_CNT_PTR(prologue_depth);
ip += 7;
}
if (!imm32)
@@ -2948,13 +2989,13 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
&prog,
ip,
callee_regs_used,
- stack_depth,
+ prologue_depth,
ctx);
else
emit_bpf_tail_call_indirect(bpf_prog,
&prog,
callee_regs_used,
- stack_depth,
+ prologue_depth,
ip,
ctx);
break;
@@ -3215,7 +3256,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
}
/* Deallocate outgoing args 7+ area. */
emit_add_rsp(&prog, outgoing_rsp);
- if (bpf_prog->aux->exception_boundary) {
+ if (bpf_prog->aux->exception_boundary ||
+ bpf_cleanup_force_spill(bpf_prog)) {
pop_callee_regs(&prog, all_callee_regs_used);
pop_r12(&prog);
} else {
@@ -4385,6 +4427,13 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
*/
bpf_prog_update_insn_ptrs(prog, addrs, image);
+ /*
+ * Same mapping, consumed by the bpf_throw() frame walker:
+ * turn the cleanup records into native address ranges now
+ * that the image is final.
+ */
+ bpf_cleanup_fill_native_ranges(prog, addrs, image);
+
/*
* ctx.prog_offset is used when CFI preambles put code *before*
* the function. See emit_cfi(). For FineIBT specifically this code
@@ -4501,6 +4550,11 @@ bool bpf_jit_supports_exceptions(void)
return IS_ENABLED(CONFIG_UNWINDER_ORC);
}
+bool bpf_jit_supports_cleanup_pads(void)
+{
+ return IS_ENABLED(CONFIG_UNWINDER_ORC);
+}
+
bool bpf_jit_supports_private_stack(void)
{
return true;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 12/20] bpf, arm64: Dispatch exception cleanup pads at run time
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (10 preceding siblings ...)
2026-09-18 4:42 ` [PATCH bpf-next v2 11/20] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
@ 2026-09-18 4:43 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-18 4:43 ` [PATCH bpf-next v2 13/20] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
` (7 subsequent siblings)
19 siblings, 1 reply; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:43 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
The same arch half for arm64: force the full callee-saved spill, record
where it starts, build the native table from the JIT's byte offsets, and
hand control to a pad from arch_bpf_run_cleanup_pad().
Unlike x86-64 a pad cannot simply return: every call it makes clobbers x30,
so no return address survives to its resume. x23 carries it instead.
bpf2a64[] maps nothing to x23 or x24 -- they are pushed only to keep the
frame shape the exception callback expects -- so nothing else in generated
code touches them, and being callee-saved they survive every kfunc the pad
calls. The JIT emits "br x23" for a pad's bpf_unwind_resume().
x24 is the other one, and it is what makes a pad able to touch its own
frame at all. Generated code addresses the BPF frame through the stack
pointer -- the frame sits directly on top of it, which turns every offset
positive and each access into one instruction -- and in a pad the stack
pointer is the walker's. So the JIT has a pad recompute the equivalent of
its frame's stack pointer from BPF r10 on entry, into x24, and addresses
the frame off x24 for every instruction the previous patches marked as
running only while unwinding. One instruction per pad, and the accesses
keep the shape and the immediate range they have everywhere else.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
arch/arm64/net/Makefile | 2 +-
arch/arm64/net/bpf_cleanup_pad.S | 95 +++++++++++++++++++++++++++++++
arch/arm64/net/bpf_jit_comp.c | 97 ++++++++++++++++++++++++++++++--
3 files changed, 187 insertions(+), 7 deletions(-)
create mode 100644 arch/arm64/net/bpf_cleanup_pad.S
diff --git a/arch/arm64/net/Makefile b/arch/arm64/net/Makefile
index 3ae382bfca87..ebec2a44a52b 100644
--- a/arch/arm64/net/Makefile
+++ b/arch/arm64/net/Makefile
@@ -2,4 +2,4 @@
#
# ARM64 networking code
#
-obj-$(CONFIG_BPF_JIT) += bpf_jit_comp.o bpf_timed_may_goto.o
+obj-$(CONFIG_BPF_JIT) += bpf_jit_comp.o bpf_timed_may_goto.o bpf_cleanup_pad.o
diff --git a/arch/arm64/net/bpf_cleanup_pad.S b/arch/arm64/net/bpf_cleanup_pad.S
new file mode 100644
index 000000000000..ef441241949e
--- /dev/null
+++ b/arch/arm64/net/bpf_cleanup_pad.S
@@ -0,0 +1,95 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf_cleanup_abi.h>
+#include <linux/linkage.h>
+
+/*
+ * A frame's prologue pushes the tail call counter pair, and then -- once
+ * bpf_cleanup_force_spill() says so -- x19/x20, x21/x22, x23/x24, x25/x26 and
+ * x27/x28. Each A64_PUSH pre-decrements, so the lowest address of the spill
+ * area holds x27 and the highest x20:
+ *
+ * spill_base + 0 x27 (private stack pointer)
+ * spill_base + 8 x28 (arena base)
+ * spill_base + 16 x25 (BPF r10, the frame pointer)
+ * spill_base + 24 x26 (tail call counter pointer)
+ * spill_base + 32 x23 -- the pad's own, see below
+ * spill_base + 40 x24 -- likewise
+ * spill_base + 48 x21 (BPF r8)
+ * spill_base + 56 x22 (BPF r9)
+ * spill_base + 64 x19 (BPF r6)
+ * spill_base + 72 x20 (BPF r7)
+ *
+ * Neither x23 nor x24 is restored from that spill: the pad has its own use for
+ * both. bpf2a64[] maps nothing to either -- they are pushed only to keep the
+ * frame shape the exception callback expects -- so nothing else in generated
+ * code touches them, and being callee-saved they survive every call the pad
+ * makes.
+ *
+ * x23 is the pad's return address. Unlike x86-64 a pad cannot simply return:
+ * every call it makes clobbers x30, so nothing is left to return through by
+ * the time it reaches its resume. The JIT emits "br x23" for the pad's
+ * bpf_unwind_resume() and this routine puts .Lcleanup_pad_done there.
+ *
+ * x24 is where the pad's frame is anchored. Generated code addresses the BPF
+ * frame through the stack pointer, which here is this routine's rather than
+ * the unwinding frame's, so the JIT has the pad recompute the equivalent from
+ * BPF r10 on entry and address its frame off x24 for as long as it runs.
+ *
+ * Both of those branches are indirect, so both targets carry a BTI landing
+ * marker: the JIT emits one at each pad, and .Lcleanup_pad_done below has one
+ * of its own.
+ */
+
+ .text
+
+/*
+ * void arch_bpf_run_cleanup_pad(u64 pad, u64 frame_fp, u64 spill_base)
+ *
+ * x0 = native address of the landing pad
+ * x1 = frame pointer of the frame the pad belongs to (unused here: BPF r10 is
+ * x25, which the spill area already holds)
+ * x2 = spill area holding that frame's BPF callee-saved registers
+ *
+ * Give the pad the register state of its own frame and call it. It runs on
+ * this stack, far below the frame it is cleaning up after, so nothing it
+ * calls can reach into that frame.
+ */
+SYM_FUNC_START(arch_bpf_run_cleanup_pad)
+ /* Save the kernel's callee-saved registers; the pad owns them next. */
+ stp x29, x30, [sp, #-96]!
+ mov x29, sp
+ stp x19, x20, [sp, #16]
+ stp x21, x22, [sp, #32]
+ stp x23, x24, [sp, #48]
+ stp x25, x26, [sp, #64]
+ stp x27, x28, [sp, #80]
+
+ /* x9 is BPF_REG_AX, so the pad's address does not stay in BPF r1. */
+ mov x9, x0
+
+ ldp x27, x28, [x2, #0]
+ ldp x25, x26, [x2, #16]
+ ldp x21, x22, [x2, #48]
+ ldp x19, x20, [x2, #64]
+
+ /* BPF r0 (x8) on the way into a pad, not whatever the kernel left. */
+ mov x8, #BPF_PAD_ENTRY_R0
+
+ /* Where the pad's resume branches back to. */
+ adr x23, .Lcleanup_pad_done
+
+ br x9
+
+.Lcleanup_pad_done:
+ /* Reached by the pad's "br x23", so it is an indirect branch target. */
+ bti j
+ ldp x19, x20, [sp, #16]
+ ldp x21, x22, [sp, #32]
+ ldp x23, x24, [sp, #48]
+ ldp x25, x26, [sp, #64]
+ ldp x27, x28, [sp, #80]
+ ldp x29, x30, [sp], #96
+ ret
+SYM_FUNC_END(arch_bpf_run_cleanup_pad)
diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
index 6c04fee46876..560eba305bca 100644
--- a/arch/arm64/net/bpf_jit_comp.c
+++ b/arch/arm64/net/bpf_jit_comp.c
@@ -11,6 +11,7 @@
#include <linux/bitfield.h>
#include <linux/bpf.h>
#include <linux/cfi.h>
+#include <linux/bpf_verifier.h>
#include <linux/filter.h>
#include <linux/memory.h>
#include <linux/printk.h>
@@ -75,7 +76,21 @@ static const int bpf2a64[] = {
[ARENA_VM_START] = A64_R(28),
};
+/* Throw-site spill: the five pairs push_callee_regs() forces on, same size and
+ * slot order, so arch_bpf_run_cleanup_pad() reads both alike.
+ */
+#define A64_CLEANUP_SPILL_SZ (5 * 16)
+
+/*
+ * Where a landing pad's frame is anchored, since the stack pointer generated
+ * code normally addresses it through is the walker's inside a pad. bpf2a64[]
+ * maps nothing to x24, so nothing else in generated code touches it.
+ */
+#define A64_CLEANUP_FP A64_R(24)
+
struct jit_ctx {
+ /* Bytes reserved for the throw-site spill; see bpf_cleanup_force_spill(). */
+ u32 throw_spill;
const struct bpf_prog *prog;
int idx;
int epilogue_offset;
@@ -432,7 +447,7 @@ static void push_callee_regs(struct jit_ctx *ctx)
* Callee-saved registers as the exception callback needs to recover
* all ARM64 Callee-saved registers in its epilogue.
*/
- if (ctx->prog->aux->exception_boundary) {
+ if (ctx->prog->aux->exception_boundary || bpf_cleanup_force_spill(ctx->prog)) {
emit(A64_PUSH(A64_R(19), A64_R(20), A64_SP), ctx);
emit(A64_PUSH(A64_R(21), A64_R(22), A64_SP), ctx);
emit(A64_PUSH(A64_R(23), A64_R(24), A64_SP), ctx);
@@ -466,7 +481,8 @@ static void pop_callee_regs(struct jit_ctx *ctx)
* program's stack frame, so recover these extra registers in the above
* two cases.
*/
- if (aux->exception_boundary || aux->exception_cb) {
+ if (aux->exception_boundary || aux->exception_cb ||
+ bpf_cleanup_force_spill(ctx->prog)) {
emit(A64_POP(A64_R(27), A64_R(28), A64_SP), ctx);
emit(A64_POP(A64_R(25), A64_R(26), A64_SP), ctx);
emit(A64_POP(A64_R(23), A64_R(24), A64_SP), ctx);
@@ -602,6 +618,20 @@ static int build_prologue(struct jit_ctx *ctx, bool ebpf_from_cbpf)
emit(A64_SUB_I(1, A64_SP, A64_FP, 96), ctx);
}
+ /*
+ * Lowest address of each spill area, as an offset from A64_FP; see
+ * bpf_cleanup_pad.S for the layout. The 16 is the tail call counter
+ * pair pushed just below the frame record, and the throw-site area
+ * sits below the callee-saved one rather than in the program stack.
+ */
+ if (bpf_cleanup_force_spill(prog)) {
+ prog->aux->exc->spill_off = -(16 + A64_CLEANUP_SPILL_SZ);
+ ctx->throw_spill = A64_CLEANUP_SPILL_SZ;
+ emit(A64_SUB_I(1, A64_SP, A64_SP, ctx->throw_spill), ctx);
+ prog->aux->exc->throw_spill_off =
+ -(16 + A64_CLEANUP_SPILL_SZ) - ctx->throw_spill;
+ }
+
/* Stack must be multiples of 16B */
ctx->stack_size = round_up(prog->aux->stack_depth, 16);
@@ -691,6 +721,10 @@ static int emit_bpf_tail_call(struct jit_ctx *ctx)
if (ctx->stack_size && !ctx->priv_sp_used)
emit(A64_ADD_I(1, A64_SP, A64_SP, ctx->stack_size), ctx);
+ /* Release it for the same reason build_epilogue() does. */
+ if (ctx->throw_spill)
+ emit(A64_ADD_I(1, A64_SP, A64_SP, ctx->throw_spill), ctx);
+
pop_callee_regs(ctx);
/* goto *(prog->bpf_func + prologue_offset); */
@@ -1055,6 +1089,9 @@ static void build_epilogue(struct jit_ctx *ctx, bool was_classic)
if (ctx->stack_size && !ctx->priv_sp_used)
emit(A64_ADD_I(1, A64_SP, A64_SP, ctx->stack_size), ctx);
+ if (ctx->throw_spill)
+ emit(A64_ADD_I(1, A64_SP, A64_SP, ctx->throw_spill), ctx);
+
pop_callee_regs(ctx);
emit(A64_POP(A64_ZR, ptr, A64_SP), ctx);
@@ -1230,6 +1267,13 @@ static const u8 stack_arg_reg[] = { A64_R(5), A64_R(6), A64_R(7) };
#define NR_STACK_ARG_REGS ARRAY_SIZE(stack_arg_reg)
+/*
+ * This reads the incoming argument area off A64_FP, which in an exception
+ * cleanup landing pad would be arch_bpf_run_cleanup_pad()'s frame record
+ * rather than the unwinding frame's -- but a pad cannot contain one of these:
+ * check_stack_arg_read() requires every r11 load to come before the frame's
+ * first call, and a pad only ever runs after one.
+ */
static void emit_stack_arg_load(u8 dst, s16 bpf_off, struct jit_ctx *ctx)
{
int idx = bpf_off / sizeof(u64) - 1;
@@ -1367,6 +1411,7 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
const s16 off = insn->off;
const s32 imm = insn->imm;
const int i = insn - ctx->prog->insnsi;
+ const bool in_pad = bpf_cleanup_insn_in_pad(ctx->prog, i);
const bool is64 = BPF_CLASS(code) == BPF_ALU64 ||
BPF_CLASS(code) == BPF_JMP;
u8 jmp_cond;
@@ -1378,9 +1423,14 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
int ret;
bool sign_extend;
- if (bpf_insn_is_indirect_target(env, ctx->prog, i))
+ if (bpf_insn_is_indirect_target(env, ctx->prog, i) ||
+ bpf_cleanup_insn_is_pad(ctx->prog, i))
emit_bti(A64_BTI_J, ctx);
+ if (bpf_cleanup_insn_is_pad(ctx->prog, i))
+ emit(A64_SUB_I(1, A64_CLEANUP_FP, fp,
+ ctx->stack_size + ctx->stack_arg_size), ctx);
+
switch (code) {
/* dst = src */
case BPF_ALU | BPF_MOV | BPF_X:
@@ -1743,6 +1793,26 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
u64 func_addr;
u32 cpu_offset;
+ if (bpf_cleanup_insn_is_throw(ctx->prog, insn - ctx->prog->insnsi)) {
+ /* Spill where the bpf_throw() walker looks. */
+ const s32 off = ctx->prog->aux->exc->throw_spill_off;
+
+ emit(A64_SUB_I(1, tmp, A64_FP, -off), ctx);
+ emit(A64_STR64I(bpf2a64[PRIVATE_SP], tmp, 0), ctx);
+ emit(A64_STR64I(bpf2a64[ARENA_VM_START], tmp, 8), ctx);
+ emit(A64_STR64I(bpf2a64[BPF_REG_FP], tmp, 16), ctx);
+ emit(A64_STR64I(bpf2a64[TCCNT_PTR], tmp, 24), ctx);
+ emit(A64_STR64I(bpf2a64[BPF_REG_8], tmp, 48), ctx);
+ emit(A64_STR64I(bpf2a64[BPF_REG_9], tmp, 56), ctx);
+ emit(A64_STR64I(bpf2a64[BPF_REG_6], tmp, 64), ctx);
+ emit(A64_STR64I(bpf2a64[BPF_REG_7], tmp, 72), ctx);
+ }
+
+ if (bpf_is_unwind_resume_kfunc(insn)) {
+ emit(A64_BR(A64_R(23)), ctx);
+ break;
+ }
+
/* Implement helper call to bpf_get_smp_processor_id() inline */
if (insn->src_reg == 0 && insn->imm == BPF_FUNC_get_smp_processor_id) {
cpu_offset = offsetof(struct thread_info, cpu);
@@ -1854,7 +1924,8 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
src = tmp2;
}
if (src == fp) {
- src_adj = ctx->priv_sp_used ? priv_sp : A64_SP;
+ src_adj = ctx->priv_sp_used ? priv_sp :
+ in_pad ? A64_CLEANUP_FP : A64_SP;
off_adj = off + ctx->stack_size;
if (!ctx->priv_sp_used)
off_adj += ctx->stack_arg_size;
@@ -1952,7 +2023,8 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
dst = tmp3;
}
if (dst == fp) {
- dst_adj = ctx->priv_sp_used ? priv_sp : A64_SP;
+ dst_adj = ctx->priv_sp_used ? priv_sp :
+ in_pad ? A64_CLEANUP_FP : A64_SP;
off_adj = off + ctx->stack_size;
if (!ctx->priv_sp_used)
off_adj += ctx->stack_arg_size;
@@ -2021,7 +2093,8 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
dst = tmp2;
}
if (dst == fp) {
- dst_adj = ctx->priv_sp_used ? priv_sp : A64_SP;
+ dst_adj = ctx->priv_sp_used ? priv_sp :
+ in_pad ? A64_CLEANUP_FP : A64_SP;
off_adj = off + ctx->stack_size;
if (!ctx->priv_sp_used)
off_adj += ctx->stack_arg_size;
@@ -2410,6 +2483,13 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
* reasons, expects to point to the next instruction)
*/
bpf_prog_update_insn_ptrs(prog, ctx.offset, ctx.ro_image);
+
+ /*
+ * Same byte offsets, consumed by the bpf_throw() frame walker:
+ * turn the cleanup records into native address ranges now that
+ * the image is final.
+ */
+ bpf_cleanup_fill_native_ranges(prog, ctx.offset, ctx.ro_image);
out_off:
if (!ro_header && priv_stack_ptr) {
free_percpu(priv_stack_ptr);
@@ -3385,6 +3465,11 @@ bool bpf_jit_supports_exceptions(void)
return true;
}
+bool bpf_jit_supports_cleanup_pads(void)
+{
+ return true;
+}
+
bool bpf_jit_supports_arena(void)
{
return true;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 13/20] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (11 preceding siblings ...)
2026-09-18 4:43 ` [PATCH bpf-next v2 12/20] bpf, arm64: " Yonghong Song
@ 2026-09-18 4:43 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
` (6 subsequent siblings)
19 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:43 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
LLVM terminates a cleanup landing pad with a call to _Unwind_Resume: the
base unwind ABI's entry point for carrying an unwind on once a frame's
cleanups have run. The kernel provides that terminator as a kfunc, but
not under that name: claiming _Unwind_Resume in the kernel's own symbol
table, for a function whose body never runs, would be needlessly confusing,
so it is called bpf_unwind_resume.
Both load paths take the detour. A direct load resolves the name against
the kernel's BTF while libbpf runs. A light skeleton instead writes the
name into the loader program's blob of bytes, for that program to resolve
when it runs, so the name recorded there has to be translated as well.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/libbpf.c | 21 ++++++++++++++++-----
1 file changed, 16 insertions(+), 5 deletions(-)
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index 33ff61b151f3..27f2eec13f76 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -8351,6 +8351,15 @@ static void fixup_verifier_log(struct bpf_program *prog, char *buf, size_t buf_s
}
}
+/* LLVM terminates a cleanup landing pad with a call to _Unwind_Resume, the
+ * base unwind ABI's entry point for carrying an unwind on once a frame's
+ * cleanups have run. The kernel knows it as bpf_unwind_resume.
+ */
+static const char *kern_extern_name(const char *name)
+{
+ return strcmp(name, "_Unwind_Resume") ? name : "bpf_unwind_resume";
+}
+
static int bpf_program_record_relos(struct bpf_program *prog)
{
struct bpf_object *obj = prog->obj;
@@ -8367,12 +8376,12 @@ static int bpf_program_record_relos(struct bpf_program *prog)
continue;
kind = btf_is_var(btf__type_by_id(obj->btf, ext->btf_id)) ?
BTF_KIND_VAR : BTF_KIND_FUNC;
- bpf_gen__record_extern(obj->gen_loader, ext->name,
+ bpf_gen__record_extern(obj->gen_loader, kern_extern_name(ext->name),
ext->is_weak, !ext->ksym.type_id,
true, kind, relo->insn_idx);
break;
case RELO_EXTERN_CALL:
- bpf_gen__record_extern(obj->gen_loader, ext->name,
+ bpf_gen__record_extern(obj->gen_loader, kern_extern_name(ext->name),
ext->is_weak, false, false, BTF_KIND_FUNC,
relo->insn_idx);
break;
@@ -8812,17 +8821,19 @@ static int bpf_object__resolve_ksym_func_btf_id(struct bpf_object *obj,
struct module_btf *mod_btf = NULL;
const struct btf_type *kern_func;
struct btf *kern_btf = NULL;
+ const char *kern_name;
int ret;
local_func_proto_id = ext->ksym.type_id;
- kfunc_id = find_ksym_btf_id(obj, ext->essent_name ?: ext->name, BTF_KIND_FUNC, &kern_btf,
- &mod_btf);
+ kern_name = kern_extern_name(ext->essent_name ?: ext->name);
+
+ kfunc_id = find_ksym_btf_id(obj, kern_name, BTF_KIND_FUNC, &kern_btf, &mod_btf);
if (kfunc_id < 0) {
if (kfunc_id == -ESRCH && ext->is_weak)
return 0;
pr_warn("extern (func ksym) '%s': not found in kernel or module BTFs\n",
- ext->name);
+ strcmp(kern_name, "bpf_unwind_resume") ? ext->name : kern_name);
return kfunc_id;
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (12 preceding siblings ...)
2026-09-18 4:43 ` [PATCH bpf-next v2 13/20] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
@ 2026-09-18 4:43 ` Yonghong Song
2026-09-18 4:57 ` sashiko-bot
2026-09-18 4:43 ` [PATCH bpf-next v2 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
` (5 subsequent siblings)
19 siblings, 1 reply; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:43 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Let a caller hand the kernel an exception cleanup table. Extend struct
bpf_prog_load_opts with cleanup_info, cleanup_info_cnt and
cleanup_info_rec_size, pass them through to BPF_PROG_LOAD, and grow the
attr size bpf_prog_load() computes to cover the new fields.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/bpf.c | 6 +++++-
tools/lib/bpf/bpf.h | 7 ++++++-
2 files changed, 11 insertions(+), 2 deletions(-)
diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
index 96819c082c77..bcf490570961 100644
--- a/tools/lib/bpf/bpf.c
+++ b/tools/lib/bpf/bpf.c
@@ -295,7 +295,7 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
const struct bpf_insn *insns, size_t insn_cnt,
struct bpf_prog_load_opts *opts)
{
- const size_t attr_sz = offsetofend(union bpf_attr, keyring_id);
+ const size_t attr_sz = offsetofend(union bpf_attr, cleanup_info_cnt);
void *finfo = NULL, *linfo = NULL;
const char *func_info, *line_info;
__u32 log_size, log_level, attach_prog_fd, attach_btf_obj_fd;
@@ -370,6 +370,10 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
attr.fd_array = ptr_to_u64(OPTS_GET(opts, fd_array, NULL));
attr.fd_array_cnt = OPTS_GET(opts, fd_array_cnt, 0);
+ attr.cleanup_info = ptr_to_u64(OPTS_GET(opts, cleanup_info, NULL));
+ attr.cleanup_info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
+ attr.cleanup_info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
+
if (log_level) {
attr.log_buf = ptr_to_u64(log_buf);
attr.log_size = log_size;
diff --git a/tools/lib/bpf/bpf.h b/tools/lib/bpf/bpf.h
index 7534a593edae..6f62a99e1e3e 100644
--- a/tools/lib/bpf/bpf.h
+++ b/tools/lib/bpf/bpf.h
@@ -128,9 +128,14 @@ struct bpf_prog_load_opts {
/* if set, provides the length of fd_array */
__u32 fd_array_cnt;
+
+ /* exception cleanup table, from the .bpf_cleanup section */
+ const void *cleanup_info;
+ __u32 cleanup_info_cnt;
+ __u32 cleanup_info_rec_size;
size_t :0;
};
-#define bpf_prog_load_opts__last_field fd_array_cnt
+#define bpf_prog_load_opts__last_field cleanup_info_rec_size
LIBBPF_API int bpf_prog_load(enum bpf_prog_type prog_type,
const char *prog_name, const char *license,
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (13 preceding siblings ...)
2026-09-18 4:43 ` [PATCH bpf-next v2 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
@ 2026-09-18 4:43 ` Yonghong Song
2026-09-18 5:00 ` sashiko-bot
2026-09-18 4:43 ` [PATCH bpf-next v2 16/20] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
` (4 subsequent siblings)
19 siblings, 1 reply; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:43 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Parse the compiler-emitted .bpf_cleanup section and hand the resulting
table to BPF_PROG_LOAD.
Each record is three 4-byte fields, and each field is a byte offset into
some code section named by a matching .rel.bpf_cleanup relocation.
bpf_object__init_cleanup_info() resolves both halves once at open time and
keeps (section, instruction index) pairs; it rejects a record whose field
has no relocation, whose relocation is not one of the two 32-bit types that
spell a data reference to a code section -- R_BPF_64_NODYLD32 from LLVM,
R_BPF_64_ABS32 from GNU as -- or whose offset is not instruction aligned.
The records are sorted by begin_off once the offsets are final: the kernel
wants the table sorted with disjoint ranges so that it can find the record
covering a call site with a binary search, and records arrive in
.bpf_cleanup order, which says nothing about where the subprograms they
describe were appended. Overlapping ranges are reported here, where the
program name and both regions are still at hand.
bpf_object_load_prog() then passes the per-program table through the
bpf_prog_load() options added in the previous patch, with the record size
carried on the program the way func_info and line_info carry theirs rather
than taken from a sizeof() at the call site. bpf_program__clone() carries
it too. That is a second load path -- the one veristat uses -- and without
the table the kernel sees landing pads nothing reaches and refuses the
program with "unreachable insn".
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/libbpf.c | 291 ++++++++++++++++++++++++++++++++
tools/lib/bpf/libbpf_internal.h | 3 +
2 files changed, 294 insertions(+)
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index 27f2eec13f76..1b8b978e7b7b 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -514,6 +514,11 @@ struct bpf_program {
void *line_info;
__u32 line_info_rec_size;
__u32 line_info_cnt;
+
+ struct bpf_cleanup_info *cleanup_info;
+ __u32 cleanup_info_rec_size;
+ __u32 cleanup_info_cnt;
+
__u32 prog_flags;
__u8 hash[SHA256_DIGEST_LENGTH];
@@ -549,6 +554,7 @@ struct bpf_struct_ops {
#define STRUCT_OPS_SEC ".struct_ops"
#define STRUCT_OPS_LINK_SEC ".struct_ops.link"
#define ARENA_SEC ".addr_space.1"
+#define CLEANUP_SEC ".bpf_cleanup"
enum libbpf_map_type {
LIBBPF_MAP_UNSPEC,
@@ -677,6 +683,25 @@ struct elf_sec_desc {
Elf_Data *data;
};
+#define CLEANUP_REC_FIELDS (sizeof(struct bpf_cleanup_info) / sizeof(__u32))
+
+/* Index of each field of struct bpf_cleanup_info, read as an array of __u32. */
+enum {
+ CLEANUP_REC_BEGIN,
+ CLEANUP_REC_END,
+ CLEANUP_REC_PAD,
+};
+
+/* One (begin, end, landing_pad) triple from .bpf_cleanup, with each field
+ * resolved from its relocation to an ELF section plus a section-relative
+ * instruction index. The mapping to final program instruction indices can only
+ * happen after subprogram placement, which differs per main program.
+ */
+struct cleanup_raw_rec {
+ int sec_idx[CLEANUP_REC_FIELDS];
+ size_t insn_idx[CLEANUP_REC_FIELDS];
+};
+
struct elf_state {
int fd;
const void *obj_buf;
@@ -696,6 +721,8 @@ struct elf_state {
bool has_st_ops;
int arena_data_shndx;
int jumptables_data_shndx;
+ Elf_Data *cleanup_data;
+ int cleanup_shndx;
};
struct usdt_manager;
@@ -771,6 +798,9 @@ struct bpf_object {
void *jumptables_data;
size_t jumptables_data_sz;
+ struct cleanup_raw_rec *cleanup_recs;
+ size_t cleanup_rec_cnt;
+
struct {
struct bpf_program *prog;
unsigned int sym_off;
@@ -817,7 +847,10 @@ static void bpf_program__exit(struct bpf_program *prog)
zfree(&prog->sec_name);
zfree(&prog->insns);
zfree(&prog->reloc_desc);
+ zfree(&prog->cleanup_info);
+ prog->cleanup_info_rec_size = 0;
+ prog->cleanup_info_cnt = 0;
prog->nr_reloc = 0;
prog->insns_cnt = 0;
prog->sec_idx = -1;
@@ -1559,6 +1592,7 @@ static struct bpf_object *bpf_object__new(const char *path,
obj->efile.obj_buf = obj_buf;
obj->efile.obj_buf_sz = obj_buf_sz;
obj->efile.btf_maps_shndx = -1;
+ obj->efile.cleanup_shndx = -1;
obj->kconfig_map_idx = -1;
obj->arena_map_idx = -1;
@@ -4045,6 +4079,9 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
sec_desc->shdr = sh;
sec_desc->data = data;
obj->efile.has_st_ops = true;
+ } else if (strcmp(name, CLEANUP_SEC) == 0) {
+ obj->efile.cleanup_data = data;
+ obj->efile.cleanup_shndx = idx;
} else if (strcmp(name, ARENA_SEC) == 0) {
obj->efile.arena_data = data;
obj->efile.arena_data_shndx = idx;
@@ -4072,6 +4109,7 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
strcmp(name, ".rel" STRUCT_OPS_LINK_SEC) &&
strcmp(name, ".rel?" STRUCT_OPS_SEC) &&
strcmp(name, ".rel?" STRUCT_OPS_LINK_SEC) &&
+ strcmp(name, ".rel" CLEANUP_SEC) &&
strcmp(name, ".rel" MAPS_ELF_SEC)) {
pr_info("elf: skipping relo section(%d) %s for section(%d) %s\n",
idx, name, targ_sec_idx,
@@ -4852,6 +4890,221 @@ static struct bpf_program *find_prog_by_sec_insn(const struct bpf_object *obj,
return NULL;
}
+static int bpf_object__init_cleanup_info(struct bpf_object *obj)
+{
+ Elf_Data *data = obj->efile.cleanup_data;
+ Elf_Data *relo = NULL;
+ size_t i, nrels, nslots, nrecs;
+ struct cleanup_raw_rec *recs;
+ int *slot_sec, ret = 0;
+ size_t *slot_val;
+ const __u32 *vals;
+ bool native;
+
+ if (!data || obj->efile.cleanup_shndx < 0)
+ return 0;
+
+ native = is_native_endianness(obj);
+
+ for (i = 0; i < obj->efile.sec_cnt; i++) {
+ struct elf_sec_desc *sd = &obj->efile.secs[i];
+
+ if (sd->sec_type == SEC_RELO && sd->shdr &&
+ sd->shdr->sh_info == (Elf64_Word)obj->efile.cleanup_shndx) {
+ relo = sd->data;
+ break;
+ }
+ }
+ if (!relo) {
+ pr_warn("%s present without relocations\n", CLEANUP_SEC);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ if (data->d_size % sizeof(struct bpf_cleanup_info)) {
+ pr_warn("%s size %zu is not a multiple of the record size %zu\n",
+ CLEANUP_SEC, data->d_size, sizeof(struct bpf_cleanup_info));
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ vals = data->d_buf;
+ nslots = data->d_size / sizeof(__u32);
+ nrecs = data->d_size / sizeof(struct bpf_cleanup_info);
+
+ slot_sec = calloc(nslots, sizeof(*slot_sec));
+ slot_val = calloc(nslots, sizeof(*slot_val));
+ recs = calloc(nrecs ?: 1, sizeof(*recs));
+ if (!slot_sec || !slot_val || !recs) {
+ ret = -ENOMEM;
+ goto out;
+ }
+ for (i = 0; i < nslots; i++)
+ slot_sec[i] = -1;
+
+ /* One relocation per 4-byte field, naming the section it points into. */
+ nrels = relo->d_size / sizeof(Elf64_Rel);
+ for (i = 0; i < nrels; i++) {
+ Elf64_Rel *rel = elf_rel_by_idx(relo, i);
+ Elf64_Sym *sym = elf_sym_by_idx(obj, ELF64_R_SYM(rel->r_info));
+ size_t type = ELF64_R_TYPE(rel->r_info);
+ size_t slot = rel->r_offset / sizeof(__u32);
+
+ if (type != R_BPF_64_NODYLD32 && type != R_BPF_64_ABS32) {
+ pr_warn("%s: relocation %zu has unexpected type %zu\n",
+ CLEANUP_SEC, i, type);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ if (!sym || slot >= nslots || rel->r_offset % sizeof(__u32)) {
+ pr_warn("%s: bad relocation %zu\n", CLEANUP_SEC, i);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ slot_sec[slot] = sym->st_shndx;
+ /* The addend lives in the section data, which libelf leaves in
+ * the object's byte order; a non-section symbol additionally
+ * contributes its own value.
+ */
+ slot_val[slot] = (native ? vals[slot] : bswap_32(vals[slot])) +
+ sym->st_value;
+ }
+
+ for (i = 0; i < nslots; i++) {
+ struct cleanup_raw_rec *rec = &recs[i / CLEANUP_REC_FIELDS];
+ size_t field = i % CLEANUP_REC_FIELDS;
+
+ if (slot_sec[i] < 0) {
+ pr_warn("%s: field %zu has no relocation\n", CLEANUP_SEC, i);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ if (slot_val[i] % BPF_INSN_SZ) {
+ pr_warn("%s: field %zu offset %zu is not instruction aligned\n",
+ CLEANUP_SEC, i, slot_val[i]);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ rec->sec_idx[field] = slot_sec[i];
+ rec->insn_idx[field] = slot_val[i] / BPF_INSN_SZ;
+ }
+
+ obj->cleanup_recs = recs;
+ obj->cleanup_rec_cnt = nrecs;
+ recs = NULL;
+out:
+ free(recs);
+ free(slot_val);
+ free(slot_sec);
+ return ret;
+}
+
+static int cmp_cleanup_info(const void *a, const void *b)
+{
+ const struct bpf_cleanup_info *x = a, *y = b;
+
+ if (x->begin_off == y->begin_off)
+ return 0;
+ return x->begin_off < y->begin_off ? -1 : 1;
+}
+
+static int bpf_prog_collect_cleanup_info(struct bpf_object *obj,
+ struct bpf_program *prog)
+{
+ size_t i;
+ int j;
+
+ for (i = 0; i < obj->cleanup_rec_cnt; i++) {
+ struct cleanup_raw_rec *raw = &obj->cleanup_recs[i];
+ struct bpf_program *owner = NULL;
+ struct bpf_cleanup_info ci = {};
+ __u32 fields[CLEANUP_REC_FIELDS];
+ void *tmp;
+
+ for (j = 0; j < CLEANUP_REC_FIELDS; j++) {
+ size_t idx = raw->insn_idx[j], final;
+ struct bpf_program *p;
+
+ /* The end of a range is exclusive, so it may name the
+ * instruction just past the last one of a function,
+ * which belongs to the next function or to nothing at
+ * all. Ask about the last instruction the range covers,
+ * the way the kernel does.
+ */
+ if (j == CLEANUP_REC_END) {
+ if (!idx) {
+ pr_warn("%s: record %zu is an empty range\n",
+ CLEANUP_SEC, i);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ idx--;
+ }
+
+ p = find_prog_by_sec_insn(obj, raw->sec_idx[j], idx);
+ if (!p) {
+ pr_warn("%s: record %zu field %d is not inside a function\n",
+ CLEANUP_SEC, i, j);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ if (!owner) {
+ owner = p;
+ } else if (owner != p) {
+ pr_warn("%s: record %zu spans functions '%s' and '%s'\n",
+ CLEANUP_SEC, i, owner->name, p->name);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ if (owner == prog) {
+ final = raw->insn_idx[j] - prog->sec_insn_off;
+ } else if (prog_is_subprog(obj, owner) && owner->sub_insn_off) {
+ /* sub_insn_off is where this subprogram was
+ * appended to the main program being relocated;
+ * zero means it is not part of it.
+ */
+ final = owner->sub_insn_off +
+ raw->insn_idx[j] - owner->sec_insn_off;
+ } else {
+ owner = NULL;
+ break;
+ }
+ fields[j] = final;
+ }
+ if (!owner)
+ continue;
+
+ ci.begin_off = fields[CLEANUP_REC_BEGIN];
+ ci.end_off = fields[CLEANUP_REC_END];
+ ci.landing_pad_off = fields[CLEANUP_REC_PAD];
+
+ tmp = libbpf_reallocarray(prog->cleanup_info, prog->cleanup_info_cnt + 1,
+ sizeof(*prog->cleanup_info));
+ if (!tmp)
+ return -ENOMEM;
+ prog->cleanup_info = tmp;
+ prog->cleanup_info_rec_size = sizeof(struct bpf_cleanup_info);
+ prog->cleanup_info[prog->cleanup_info_cnt++] = ci;
+
+ pr_debug("prog '%s': cleanup region [%u,%u) -> landing pad %u\n",
+ prog->name, ci.begin_off, ci.end_off, ci.landing_pad_off);
+ }
+
+ if (!prog->cleanup_info_cnt)
+ return 0;
+
+ qsort(prog->cleanup_info, prog->cleanup_info_cnt,
+ sizeof(*prog->cleanup_info), cmp_cleanup_info);
+ for (i = 1; i < prog->cleanup_info_cnt; i++) {
+ struct bpf_cleanup_info *prev = &prog->cleanup_info[i - 1];
+ struct bpf_cleanup_info *cur = &prog->cleanup_info[i];
+
+ if (cur->begin_off < prev->end_off) {
+ pr_warn("prog '%s': overlapping cleanup regions [%u,%u) and [%u,%u)\n",
+ prog->name, prev->begin_off, prev->end_off,
+ cur->begin_off, cur->end_off);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ }
+
+ return 0;
+}
+
static int
bpf_object__collect_prog_relos(struct bpf_object *obj, Elf64_Shdr *shdr, Elf_Data *data)
{
@@ -7561,6 +7814,13 @@ static int bpf_object__relocate(struct bpf_object *obj, const char *targ_btf_pat
return err;
}
}
+
+ err = bpf_prog_collect_cleanup_info(obj, prog);
+ if (err) {
+ pr_warn("prog '%s': failed to collect cleanup info: %s\n",
+ prog->name, errstr(err));
+ return err;
+ }
}
for (i = 0; i < obj->nr_programs; i++) {
prog = &obj->programs[i];
@@ -7751,6 +8011,9 @@ static int bpf_object__collect_relos(struct bpf_object *obj)
return -LIBBPF_ERRNO__INTERNAL;
}
+ if (idx == obj->efile.cleanup_shndx)
+ continue;
+
if (obj->efile.secs[idx].sec_type == SEC_ST_OPS)
err = bpf_object__collect_st_ops_relos(obj, shdr, data);
else if (idx == obj->efile.btf_maps_shndx)
@@ -8023,6 +8286,11 @@ static int bpf_object_load_prog(struct bpf_object *obj, struct bpf_program *prog
load_attr.line_info_rec_size = prog->line_info_rec_size;
load_attr.line_info_cnt = prog->line_info_cnt;
}
+ if (prog->cleanup_info_cnt) {
+ load_attr.cleanup_info = prog->cleanup_info;
+ load_attr.cleanup_info_cnt = prog->cleanup_info_cnt;
+ load_attr.cleanup_info_rec_size = prog->cleanup_info_rec_size;
+ }
load_attr.log_level = log_level;
load_attr.prog_flags = prog->prog_flags;
load_attr.fd_array = obj->fd_array;
@@ -8579,6 +8847,7 @@ static struct bpf_object *bpf_object_open(const char *path, const void *obj_buf,
err = err ? : bpf_object__init_maps(obj, opts);
err = err ? : bpf_object_init_progs(obj, opts);
err = err ? : bpf_object__collect_relos(obj);
+ err = err ? : bpf_object__init_cleanup_info(obj);
if (err)
goto out;
@@ -9692,6 +9961,9 @@ void bpf_object__close(struct bpf_object *obj)
zfree(&obj->jumptables_data);
obj->jumptables_data_sz = 0;
+ zfree(&obj->cleanup_recs);
+ obj->cleanup_rec_cnt = 0;
+
for (i = 0; i < obj->jumptable_map_cnt; i++)
close(obj->jumptable_maps[i].fd);
zfree(&obj->jumptable_maps);
@@ -10084,6 +10356,25 @@ int bpf_program__clone(struct bpf_program *prog, const struct bpf_prog_load_opts
attr.line_info_rec_size = info ? info_rec_size : prog->line_info_rec_size;
}
+ /* exception cleanup table */
+ info = OPTS_GET(opts, cleanup_info, NULL);
+ info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
+ info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
+ if (!!info != !!info_cnt || !!info != !!info_rec_size) {
+ pr_warn("prog '%s': cleanup_info, cleanup_info_cnt, and cleanup_info_rec_size must all be specified or all omitted\n",
+ prog->name);
+ return libbpf_err(-EINVAL);
+ }
+ if (info) {
+ attr.cleanup_info = info;
+ attr.cleanup_info_cnt = info_cnt;
+ attr.cleanup_info_rec_size = info_rec_size;
+ } else if (prog->cleanup_info_cnt) {
+ attr.cleanup_info = prog->cleanup_info;
+ attr.cleanup_info_cnt = prog->cleanup_info_cnt;
+ attr.cleanup_info_rec_size = prog->cleanup_info_rec_size;
+ }
+
/* Logging is caller-controlled; no fallback to prog/obj log settings */
attr.log_buf = OPTS_GET(opts, log_buf, NULL);
attr.log_size = OPTS_GET(opts, log_size, 0);
diff --git a/tools/lib/bpf/libbpf_internal.h b/tools/lib/bpf/libbpf_internal.h
index cb4d96233844..3ba6d9090368 100644
--- a/tools/lib/bpf/libbpf_internal.h
+++ b/tools/lib/bpf/libbpf_internal.h
@@ -56,6 +56,9 @@
#ifndef R_BPF_64_ABS32
#define R_BPF_64_ABS32 3
#endif
+#ifndef R_BPF_64_NODYLD32
+#define R_BPF_64_NODYLD32 4
+#endif
#ifndef R_BPF_64_32
#define R_BPF_64_32 10
#endif
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 16/20] libbpf: Carry the exception cleanup table through the light skeleton
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (14 preceding siblings ...)
2026-09-18 4:43 ` [PATCH bpf-next v2 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
@ 2026-09-18 4:43 ` Yonghong Song
2026-09-18 5:02 ` sashiko-bot
2026-09-18 4:43 ` [PATCH bpf-next v2 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
` (3 subsequent siblings)
19 siblings, 1 reply; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:43 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A light skeleton does not call bpf_prog_load(). bpf_gen__prog_load() builds
its own union bpf_attr field by field, and the loader program it emits is
what issues BPF_PROG_LOAD when the skeleton runs -- so a program loaded
this way reached the kernel without the table the previous patch collected
for it. The load then fails, because the landing pads are code nothing
reaches and the verifier says so. It says "unreachable insn", which names
neither the skeleton nor the table.
Carry it the way func_info and line_info are carried: the records go into
the loader's blob of bytes, the count and record size into the attr, and a
relocation stores the blob's address into attr.cleanup_info once that
address is known. The attr grows to its new last field, cleanup_info_cnt.
Records are 4-byte fields like the other info blobs, so a cross-endian
build has to swap them too.
Growing the attr is what makes the relocation conditional. Every other
attr in this file stops at the last field it sets, and this one now runs
to the end of the union: a kernel that predates cleanup_info accepts an
attr longer than its own only while the tail it does not know reads as
zero. The count and the record size are already zero for a program with no
table, but the relocation stores a blob address, which never is, so a
program that has no landing pads only keeps loading on such a kernel if
the relocation is left out.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/gen_loader.c | 29 +++++++++++++++++++++++++----
tools/lib/bpf/libbpf_internal.h | 7 +++++++
2 files changed, 32 insertions(+), 4 deletions(-)
diff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c
index af3a04f161ac..2345fbdd46f5 100644
--- a/tools/lib/bpf/gen_loader.c
+++ b/tools/lib/bpf/gen_loader.c
@@ -981,13 +981,15 @@ static void cleanup_relos(struct bpf_gen *gen, int insns)
cleanup_core_relo(gen);
}
-/* Convert func, line, and core relo info blobs to target endianness */
+/* Convert func, line, core relo and cleanup info blobs to target endianness */
static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
- int core_relos, struct bpf_prog_load_opts *load_attr)
+ int core_relos, int cleanup_info,
+ struct bpf_prog_load_opts *load_attr)
{
struct bpf_func_info *fi = gen->data_start + func_info;
struct bpf_line_info *li = gen->data_start + line_info;
struct bpf_core_relo *cr = gen->data_start + core_relos;
+ struct bpf_cleanup_info *ci = gen->data_start + cleanup_info;
int i;
for (i = 0; i < load_attr->func_info_cnt; i++)
@@ -998,6 +1000,9 @@ static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
for (i = 0; i < gen->core_relo_cnt; i++)
bpf_core_relo_bswap(cr++);
+
+ for (i = 0; i < load_attr->cleanup_info_cnt; i++)
+ bpf_cleanup_info_bswap(ci++);
}
void bpf_gen__prog_load(struct bpf_gen *gen,
@@ -1011,8 +1016,11 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
load_attr->line_info_rec_size;
int core_relo_tot_sz = gen->core_relo_cnt *
sizeof(struct bpf_core_relo);
+ int cleanup_info_tot_sz = load_attr->cleanup_info_cnt *
+ load_attr->cleanup_info_rec_size;
int prog_load_attr, license_off, insns_off, func_info, line_info, core_relos;
- int attr_size = offsetofend(union bpf_attr, core_relo_rec_size);
+ int attr_size = offsetofend(union bpf_attr, cleanup_info_cnt);
+ int cleanup_info;
union bpf_attr attr;
memset(&attr, 0, attr_size);
@@ -1061,9 +1069,17 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
core_relos, gen->core_relo_cnt,
sizeof(struct bpf_core_relo));
+ attr.cleanup_info_rec_size = tgt_endian(load_attr->cleanup_info_rec_size);
+ attr.cleanup_info_cnt = tgt_endian(load_attr->cleanup_info_cnt);
+ cleanup_info = add_data(gen, load_attr->cleanup_info, cleanup_info_tot_sz);
+ pr_debug("gen: prog_load: cleanup_info: off %d cnt %u rec size %u\n",
+ cleanup_info, load_attr->cleanup_info_cnt,
+ load_attr->cleanup_info_rec_size);
+
/* convert all info blobs to target endianness */
if (gen->swapped_endian && !gen->error)
- info_blob_bswap(gen, func_info, line_info, core_relos, load_attr);
+ info_blob_bswap(gen, func_info, line_info, core_relos, cleanup_info,
+ load_attr);
libbpf_strlcpy(attr.prog_name, prog_name, sizeof(attr.prog_name));
prog_load_attr = add_data(gen, &attr, attr_size);
@@ -1085,6 +1101,11 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
/* populate union bpf_attr with a pointer to core_relos */
emit_rel_store(gen, attr_field(prog_load_attr, core_relos), core_relos);
+ /* populate union bpf_attr with a pointer to cleanup_info, if there is one */
+ if (load_attr->cleanup_info_cnt)
+ emit_rel_store(gen, attr_field(prog_load_attr, cleanup_info),
+ cleanup_info);
+
/* populate union bpf_attr fd_array with a pointer to data where map_fds are saved */
emit_rel_store(gen, attr_field(prog_load_attr, fd_array), gen->fd_array);
diff --git a/tools/lib/bpf/libbpf_internal.h b/tools/lib/bpf/libbpf_internal.h
index 3ba6d9090368..78519f24fb40 100644
--- a/tools/lib/bpf/libbpf_internal.h
+++ b/tools/lib/bpf/libbpf_internal.h
@@ -572,6 +572,13 @@ static inline void bpf_core_relo_bswap(struct bpf_core_relo *i)
i->kind = bswap_32(i->kind);
}
+static inline void bpf_cleanup_info_bswap(struct bpf_cleanup_info *i)
+{
+ i->begin_off = bswap_32(i->begin_off);
+ i->end_off = bswap_32(i->end_off);
+ i->landing_pad_off = bswap_32(i->landing_pad_off);
+}
+
enum btf_field_iter_kind {
BTF_FIELD_ITER_IDS,
BTF_FIELD_ITER_STRS,
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (15 preceding siblings ...)
2026-09-18 4:43 ` [PATCH bpf-next v2 16/20] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
@ 2026-09-18 4:43 ` Yonghong Song
2026-09-18 5:01 ` sashiko-bot
2026-09-18 5:44 ` bot+bpf-ci
2026-09-18 4:43 ` [PATCH bpf-next v2 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
` (2 subsequent siblings)
19 siblings, 2 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:43 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
An object that carries a compiler-emitted exception cleanup table cannot be
linked today. The table's fields are byte offsets into a code section,
materialised by a 32-bit relocation against that section's symbol with the
offset itself as the implicit addend, and the linker rejects both halves of
that: the relocation type is not in the list it accepts, and a relocation
against an STT_SECTION symbol from a non-executable section is an outright
error.
Both spellings of that relocation have to be taken. LLVM emits
R_BPF_64_NODYLD32 for a .long against a section symbol; GNU as emits
R_BPF_64_ABS32, which is what bpf_reloc_type_lookup() maps BFD_RELOC_32 to.
They describe the same value, and the selftests are built with both
compilers.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/linker.c | 19 ++++++++++++++++++-
1 file changed, 18 insertions(+), 1 deletion(-)
diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
index 78f92c39290a..e5c06023cb5b 100644
--- a/tools/lib/bpf/linker.c
+++ b/tools/lib/bpf/linker.c
@@ -1036,7 +1036,8 @@ static int linker_sanity_check_elf_relos(struct src_obj *obj, struct src_sec *se
size_t sym_type = ELF64_R_TYPE(relo->r_info);
if (sym_type != R_BPF_64_64 && sym_type != R_BPF_64_32 &&
- sym_type != R_BPF_64_ABS64 && sym_type != R_BPF_64_ABS32) {
+ sym_type != R_BPF_64_ABS64 && sym_type != R_BPF_64_ABS32 &&
+ sym_type != R_BPF_64_NODYLD32) {
pr_warn("ELF relo #%d in section #%zu has unexpected type %zu in %s\n",
i, sec->sec_idx, sym_type, obj->filename);
return -EINVAL;
@@ -2274,6 +2275,22 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
insn->imm += sec->dst_off / sizeof(struct bpf_insn);
else
insn->imm += sec->dst_off;
+ } else if (sym_type == R_BPF_64_NODYLD32 ||
+ sym_type == R_BPF_64_ABS32) {
+ __u32 *val;
+
+ /* Two spellings of the one thing: LLVM
+ * emits NODYLD32 for a .long against a
+ * section symbol, GNU as emits ABS32
+ * (bpf_reloc_type_lookup() maps
+ * BFD_RELOC_32 to it), and the value
+ * they describe is the same.
+ */
+ val = dst_linked_sec->raw_data + dst_rel->r_offset;
+ if (linker->swapped_endian)
+ *val = bswap_32(bswap_32(*val) + sec->dst_off);
+ else
+ *val += sec->dst_off;
} else {
pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
return -EINVAL;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (16 preceding siblings ...)
2026-09-18 4:43 ` [PATCH bpf-next v2 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
@ 2026-09-18 4:43 ` Yonghong Song
2026-09-18 4:59 ` sashiko-bot
2026-09-18 5:58 ` bot+bpf-ci
2026-09-18 4:43 ` [PATCH bpf-next v2 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 20/20] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
19 siblings, 2 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:43 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
C has no unwinding, so nothing here comes out of the frontend: the frames
that own a resource are written as __naked inline assembly, which spells
out by hand exactly what a frontend emits -- a call site bracketed by two
labels, a landing pad unreachable in the compiler's CFG, and a .bpf_cleanup
record tying them together. The assembler turns ".long <text label>" into
the same R_BPF_64_NODYLD32 relocation the BPF AsmPrinter emits, so libbpf
and the kernel see an object indistinguishable from a compiler-generated
one.
Call chain: entry -> foo1 -> foo1v -> foo2 -> foo3. foo3 holds a
non-preemptible section and throws inside it; foo2 holds an RCU read lock
and has two call sites sharing one pad, one of them its own throw; foo1v is
a void frame whose pad ends in a jump to a resume block placed after an
unrelated block that ends in a plain exit; foo1 owns nothing and gets no
record; entry is the boundary.
There are also some shapes the kernel refuses -- the ones with no correct
answer, and the ones a pad running on the bpf_throw() walker's stack cannot
express -- plus the return-value rejection that delivering an exception at
the boundary of a program type which constrains its return value produces.
The test skips rather than fails where the JIT cannot dispatch a landing
pad at all.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
.../selftests/bpf/exceptions_cleanup.h | 29 +
.../bpf/prog_tests/exceptions_cleanup.c | 79 +++
.../selftests/bpf/progs/exceptions_cleanup.c | 152 +++++
.../bpf/progs/exceptions_cleanup_fail.c | 600 ++++++++++++++++++
4 files changed, 860 insertions(+)
create mode 100644 tools/testing/selftests/bpf/exceptions_cleanup.h
create mode 100644 tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
new file mode 100644
index 000000000000..630d2e207119
--- /dev/null
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -0,0 +1,29 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef __EXCEPTIONS_CLEANUP_H__
+#define __EXCEPTIONS_CLEANUP_H__
+
+#define THROW_COOKIE 0x100
+
+/* progs/exceptions_cleanup.c: one bit per frame that reports it ran. */
+#define RAN_FOO3_PREEMPT 0x1
+#define RAN_FOO2_RCU 0x2
+#define RAN_FOO1V_PREEMPT 0x4
+#define RAN_FOO2_DROP 0x8
+#define RAN_BUMP 0x10
+
+#define CLEANUP_REC(begin, end, landing_pad) \
+ ".pushsection .bpf_cleanup,\"a\",@progbits;" \
+ ".long " begin ";" \
+ ".long " end ";" \
+ ".long " landing_pad ";" \
+ ".popsection;"
+
+/* Set a bit in @pads_ran. */
+#define PAD_RAN(bit) \
+ "r1 = %[pads_ran] ll;" \
+ "r2 = *(u64 *)(r1 + 0);" \
+ "r2 |= " bit ";" \
+ "*(u64 *)(r1 + 0) = r2;"
+
+#endif /* __EXCEPTIONS_CLEANUP_H__ */
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
new file mode 100644
index 000000000000..d1e45b765af2
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -0,0 +1,79 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <test_progs.h>
+#include "exceptions_cleanup.h"
+#include "exceptions_cleanup.skel.h"
+#include "exceptions_cleanup_fail.skel.h"
+
+/* foo3 threw: every frame that has a pad ran it. */
+#define PADS_FOO3_THREW \
+ (RAN_FOO3_PREEMPT | RAN_FOO2_RCU | RAN_FOO1V_PREEMPT | RAN_FOO2_DROP)
+
+/* foo2 threw after foo3 returned normally: foo3's pad must not run. */
+#define PADS_FOO2_THREW \
+ (RAN_FOO2_RCU | RAN_FOO1V_PREEMPT | RAN_FOO2_DROP)
+
+static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
+ __u64 pads)
+{
+ __u64 ctx = 0;
+ int err;
+
+ LIBBPF_OPTS(bpf_test_run_opts, topts,
+ .ctx_in = &ctx,
+ .ctx_size_in = sizeof(ctx),
+ );
+
+ skel->bss->input = input;
+ skel->bss->pads_ran = 0;
+ skel->bss->result = 0;
+
+ err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.entry), &topts);
+ if (!ASSERT_OK(err, "run"))
+ return;
+ ASSERT_EQ(topts.retval, retval, "retval");
+ ASSERT_EQ(skel->bss->pads_ran, pads | RAN_BUMP, "pads_ran");
+}
+
+void test_exceptions_cleanup(void)
+{
+ char log[8192] = {};
+ LIBBPF_OPTS(bpf_object_open_opts, opts,
+ .kernel_log_buf = log,
+ .kernel_log_size = sizeof(log));
+ struct exceptions_cleanup *skel;
+ int err;
+
+ skel = exceptions_cleanup__open_opts(&opts);
+ if (!ASSERT_OK_PTR(skel, "open"))
+ return;
+
+ err = exceptions_cleanup__load(skel);
+ if (err) {
+ if (err == -EOPNOTSUPP &&
+ strstr(log, "exception cleanup needs a JIT that can dispatch landing pads"))
+ test__skip();
+ else if (!ASSERT_OK(err, "load"))
+ fprintf(stderr, "%s", log);
+ exceptions_cleanup__destroy(skel);
+ return;
+ }
+
+ /* No throw: foo3 returns 1 ^ 1 == 0, foo2 adds one, no pad runs. */
+ if (test__start_subtest("no_throw"))
+ run(skel, 1, 1, 0);
+
+ /* foo3 throws; every pad runs and the cookie is delivered at entry. */
+ if (test__start_subtest("throw_from_foo3"))
+ run(skel, 101, THROW_COOKIE, PADS_FOO3_THREW);
+
+ /* foo3 returns 2 ^ 1 == 3, so foo2 throws from its own second region;
+ * foo3's frame is long gone, so its pad must not run.
+ */
+ if (test__start_subtest("throw_from_foo2"))
+ run(skel, 2, THROW_COOKIE, PADS_FOO2_THREW);
+
+ exceptions_cleanup__destroy(skel);
+
+ RUN_TESTS(exceptions_cleanup_fail);
+}
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
new file mode 100644
index 000000000000..95199a2828fa
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
@@ -0,0 +1,152 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_rcu_read_lock();
+ bpf_rcu_read_unlock();
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume();
+}
+
+__u64 input = 0;
+__u64 pads_ran = 0;
+__u64 result = 0;
+
+static __used __noinline __u64 foo3(__u64 x)
+{
+ bpf_preempt_disable();
+ if (x > 100)
+ asm volatile (
+ "r1 = %[cookie];"
+ "1:" "call bpf_throw;" /* cleanup region */
+ "2:"
+ "goto 3f;"
+ "4:" /* landing pad */
+ "r7 = r0;"
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "3:"
+ CLEANUP_REC("1b", "2b", "4b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_FOO3_PREEMPT),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+ bpf_preempt_enable();
+ return x ^ 1;
+}
+
+__u64 never = 0;
+
+static __used __naked __noinline void drop_glue(void)
+{
+ asm volatile (
+ PAD_RAN("%[ran]")
+ "exit;"
+ :
+ : [ran]"i"(RAN_FOO2_DROP), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 foo2(void)
+{
+ asm volatile (
+ "r6 = r1;"
+ "call bpf_rcu_read_lock;"
+ "r1 = r6;"
+"1:" "call foo3;" /* cleanup region #1 */
+"2:"
+ "r6 = r0;"
+ "if r6 == 0 goto 5f;"
+ "r1 = %[cookie];"
+"3:" "call bpf_throw;" /* cleanup region #2 */
+"4:"
+ "r0 = 0;"
+ "exit;"
+"5:"
+ "call bpf_rcu_read_unlock;"
+ "r0 = r6;"
+ "r0 += 1;"
+ "exit;"
+"6:" /* landing pad, shared by both regions */
+ "call drop_glue;"
+ "call bpf_rcu_read_unlock;"
+ PAD_RAN("%[ran_rcu]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "6b")
+ CLEANUP_REC("3b", "4b", "6b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran_rcu]"i"(RAN_FOO2_RCU),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+static __used __naked __noinline void foo1v(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call foo2;" /* cleanup region */
+"2:"
+ "r6 = r0;"
+ "call bpf_preempt_enable;"
+ "r1 = %[result] ll;"
+ "*(u64 *)(r1 + 0) = r6;"
+ "goto 7f;"
+"8:" /* landing pad */
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "goto 9f;"
+"7:" /* the frame's own exit block */
+ "r0 = 0;"
+ "exit;"
+"9:" /* shared resume block */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "8b")
+ :
+ : [ran]"i"(RAN_FOO1V_PREEMPT), __imm_addr(input),
+ __imm_addr(result), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+static __used __naked __noinline void bump(void)
+{
+ asm volatile (
+ PAD_RAN("%[ran]")
+ "r1 = %[never] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 == 0 goto 1f;"
+ "r1 = 0;"
+ "call bpf_throw;"
+"1:"
+ "exit;" /* r0 deliberately left alone */
+ :
+ : [ran]"i"(RAN_BUMP), __imm_addr(never), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+__noinline __u64 foo1(void)
+{
+ bump();
+ foo1v();
+ return result;
+}
+
+SEC("syscall")
+int entry(void *ctx)
+{
+ return foo1();
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
new file mode 100644
index 000000000000..ce2dac306a84
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
@@ -0,0 +1,600 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_experimental.h"
+#include "bpf_misc.h"
+#include "../test_kmods/bpf_testmod_kfunc.h"
+#include "exceptions_cleanup.h"
+
+__u64 input = 0;
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume();
+}
+
+/*
+ * 1. A subprogram that may unwind, also used as a helper callback:
+ * bpf_loop()'s own kernel frame would end the walk before it found a
+ * boundary.
+ */
+static int throwing_cb(__u32 idx, void *ctx)
+{
+ bpf_throw(0xbad);
+ return 0;
+}
+
+static __used __naked __noinline __u64 cb_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call throwing_cb;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("may unwind and is used as a callback")
+int callback_may_unwind(void *ctx)
+{
+ bpf_loop(1, throwing_cb, NULL, 0);
+ return cb_frame();
+}
+
+/*
+ * 2. A landing pad that reaches both an unwind resume and a plain exit, so
+ * nothing says whether it is a cleanup pad or a catch pad.
+ */
+static __used __naked __noinline __u64 inner_throw(void)
+{
+ asm volatile (
+ "r1 = 1;"
+ "call bpf_throw;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 ambiguous_pad_frame(void)
+{
+ asm volatile (
+ "r6 = r1;"
+ "call bpf_preempt_disable;"
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad: two ways out */
+ "call bpf_preempt_enable;"
+ "if r6 > 10 goto 4f;"
+ "call bpf_unwind_resume;"
+ "exit;"
+"4:"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("reaches both bpf_unwind_resume() and a plain exit")
+int ambiguous_landing_pad(void *ctx)
+{
+ return ambiguous_pad_frame();
+}
+
+/*
+ * 3. A throw from inside a landing pad: a second walk over the frames the
+ * first one is in the middle of discarding.
+ */
+static __used __naked __noinline __u64 throw_in_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad that throws again */
+ "call bpf_preempt_enable;"
+ "r1 = 2;"
+ "call bpf_throw;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("can throw while an exception is in flight")
+int throw_from_landing_pad(void *ctx)
+{
+ return throw_in_pad_frame();
+}
+
+/*
+ * 4. A cleanup table in a program that also installs an exception callback,
+ * two different answers to what runs on the way out.
+ */
+__noinline int unused_exc_cb(u64 cookie)
+{
+ return 0;
+}
+
+static __used __naked __noinline __u64 cb_and_table_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 9;"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__exception_cb(unused_exc_cb)
+__failure __msg("cannot be combined with an exception callback")
+int table_with_exception_cb(void *ctx)
+{
+ return cb_and_table_frame();
+}
+
+__u64 never;
+
+/*
+ * 5. A landing pad that calls a subprogram which can throw. Not case 3: the
+ * throw is in another subprogram, so what catches it is the walk of the pad's
+ * body, off subprog_info.might_throw.
+ */
+static __used __noinline void pad_callee_that_throws(void)
+{
+ if (never)
+ bpf_throw(0);
+}
+
+static __used __naked __noinline __u64 pad_calls_thrower_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 11;"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call pad_callee_that_throws;" /* ...which can throw: refused */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("which can throw while an exception is in flight")
+int pad_calls_thrower(void *ctx)
+{
+ return pad_calls_thrower_frame();
+}
+
+/*
+ * 6. A catch pad: it ends in a plain exit rather than a resume, and a walker
+ * that calls pads as subroutines cannot hand a frame back its own execution.
+ */
+static __used __naked __noinline __u64 catch_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 12;"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* catch pad: no resume, it stops here */
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not supported yet, only cleanup pads that resume")
+int catch_landing_pad(void *ctx)
+{
+ return catch_pad_frame();
+}
+
+/*
+ * 7. An exception reaching the boundary of a program type that constrains
+ * its return value: delivery makes the cookie that return value, and fentry
+ * has to return 0.
+ */
+static __used __naked __noinline __u64 boundary_throw_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 7;"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?fentry/bpf_fentry_test1")
+__failure __msg("the register R1 has smin=7 smax=7 should have been in [0, 0]")
+int boundary_delivers(void *ctx)
+{
+ return boundary_throw_frame();
+}
+
+/*
+ * 8. A bpf_unwind_resume() outside any landing pad. Both JITs lower it as
+ * the way back out of a pad, which in ordinary code leaves a live frame
+ * standing with its epilogue skipped.
+ */
+static __used __naked __noinline __u64 stray_resume_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 13;"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not in an exception cleanup landing pad")
+int resume_outside_pad(void *ctx)
+{
+ /* Never taken, but reachable, which is all the verifier needs. */
+ if (never)
+ bpf_unwind_resume();
+ return stray_resume_frame();
+}
+
+/*
+ * 9. A bpf_unwind_resume() in a subprogram a landing pad calls. The
+ * verifier's walk cannot tell it from a resume in the pad itself -- an
+ * exception is in flight either way -- so the rule is static: a resume sits in
+ * a pad body.
+ */
+static __used __naked __noinline void resume_in_callee(void)
+{
+ asm volatile (
+ "call bpf_unwind_resume;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 pad_calls_resumer_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 14;"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call resume_in_callee;" /* ...which resumes: refused */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not in an exception cleanup landing pad")
+int resume_in_pad_callee(void *ctx)
+{
+ return pad_calls_resumer_frame();
+}
+
+/*
+ * 10. A bpf_unwind_resume() in a program carrying no cleanup table, where
+ * that static rule does not run at all. do_check() refuses it on the state not
+ * unwinding, and has to: the JITs lower every one of these the same way.
+ */
+static __used __naked __noinline __u64 no_table_resume_frame(void)
+{
+ asm volatile (
+ "call bpf_unwind_resume;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("reached without an exception in flight")
+int resume_without_table(void *ctx)
+{
+ return no_table_resume_frame();
+}
+
+/*
+ * 11. A landing pad that is itself a covered call site, so an exception out
+ * of it would have nowhere to go. Hand-written only: LLVM sinks a function's
+ * pads past every range it emits.
+ */
+static __used __naked __noinline __u64 nested_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_throw;" /* first cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* first pad, second region's call */
+ "call bpf_preempt_enable;"
+"4:"
+ "call bpf_unwind_resume;"
+ "exit;"
+"5:" /* second pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ CLEANUP_REC("3b", "4b", "5b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is inside the call-site range of")
+int nested_landing_pad(void *ctx)
+{
+ return nested_pad_frame();
+}
+
+/*
+ * 12. A tail call in a landing pad: it unwinds the prologue off the stack
+ * pointer, which in a pad is the walker's.
+ */
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(__u32));
+ __uint(value_size, sizeof(__u32));
+} tc_map SEC(".maps");
+
+static __used __naked __noinline __u64 tail_call_pad_frame(void)
+{
+ asm volatile (
+ "r6 = r1;"
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r1 = r6;"
+ "r2 = %[tc_map] ll;"
+ "r3 = 0;"
+ "call %[bpf_tail_call];"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm(bpf_tail_call), __imm_addr(tc_map)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is in an exception cleanup landing pad")
+int tail_call_in_pad(void *ctx)
+{
+ return tail_call_pad_frame();
+}
+
+#if defined(__BPF_FEATURE_STACK_ARGUMENT)
+
+/*
+ * 13. A call that passes an argument on the stack, in a landing pad: the
+ * outgoing area the callee reads is not the one the caller wrote, the frame
+ * being the unwinding one and the stack pointer the walker's.
+ */
+static __used __noinline __u64 six_args(__u64 a, __u64 b, __u64 c, __u64 d,
+ __u64 e, __u64 f)
+{
+ return a + b + c + d + e + f;
+}
+
+static __used __naked __noinline __u64 stack_arg_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "r1 = 1;"
+ "r2 = 2;"
+ "r3 = 3;"
+ "r4 = 4;"
+ "r5 = 5;"
+ "*(u64 *)(r11 - 8) = 6;" /* the sixth argument */
+ "call six_args;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("on-stack call argument in an exception cleanup landing pad")
+int stack_arg_in_pad(void *ctx)
+{
+ return stack_arg_pad_frame();
+}
+
+/*
+ * 14. The same, reached the other way: a kfunc whose by-value argument runs
+ * past the five argument registers, where the JIT fills the outgoing area and
+ * the rule above has no store to catch. The C call gives the extern its BTF.
+ */
+static __used __noinline void __nofit_btf_anchor(void)
+{
+ struct prog_test_pair_arg s = {};
+
+ bpf_kfunc_call_test_pair_arg_nofit(1, 2, 3, 4, s);
+}
+
+static __used __naked __noinline __u64 kfunc_arg_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "r1 = 1;"
+ "r2 = 2;"
+ "r3 = 3;"
+ "r4 = 4;"
+ "r5 = 5;"
+ "call bpf_kfunc_call_test_pair_arg_nofit;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("on-stack call argument in an exception cleanup landing pad")
+int kfunc_stack_arg_in_pad(void *ctx)
+{
+ return kfunc_arg_pad_frame();
+}
+
+#endif /* __BPF_FEATURE_STACK_ARGUMENT */
+
+/*
+ * 15. A landing pad entered by ordinary control flow, arriving with none of
+ * what the walker sets up. Nothing static sees it -- the resume really is in a
+ * pad body -- so do_check() refuses it on the state not unwinding.
+ */
+static __used __naked __noinline __u64 jump_into_pad_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "if r6 > 7 goto 4f;" /* an ordinary branch into the pad */
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+"4:" /* ... and its second instruction */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("reached without an exception in flight")
+int jump_into_pad(void *ctx)
+{
+ return jump_into_pad_frame();
+}
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)
+
+/*
+ * 16. A landing pad that reaches an indirect jump, which cannot be told from
+ * a catch pad. SEC("socket") because a jump table entry is an offset from the
+ * program's section symbol, and "?syscall" is not a name assembly can use.
+ */
+static __used __naked __noinline void gotox_thrower(void)
+{
+ asm volatile (
+ "r1 = 15;"
+ "call bpf_throw;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("reaches an indirect jump")
+__naked void gotox_in_pad(void)
+{
+ asm volatile (
+ ".pushsection .jumptables,\"\",@progbits;"
+"jt0_%=:"
+ ".quad l0_%= - socket;"
+ ".quad l1_%= - socket;"
+ ".size jt0_%=, 16;"
+ ".global jt0_%=;"
+ ".popsection;"
+
+"1:" "call gotox_thrower;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r1 = jt0_%= ll;"
+ "r1 += 8;"
+ "r2 = *(u64 *)(r1 + 0);"
+ /* gotox r2. Spelled as a raw insn on purpose: the "gotox" mnemonic
+ * only reached the LLVM assembler in llvm 22, and BPF_RAW_INSN()
+ * needs <linux/bpf.h>, which this file cannot have -- vmlinux.h
+ * already defines the uapi enums.
+ */
+ ".8byte 0x20d;"
+"l0_%=:"
+ "call bpf_unwind_resume;"
+ "exit;"
+"l1_%=:"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+#endif /* x86 || arm64 */
+
+char _license[] SEC("license") = "GPL";
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (17 preceding siblings ...)
2026-09-18 4:43 ` [PATCH bpf-next v2 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
@ 2026-09-18 4:43 ` Yonghong Song
2026-09-18 5:01 ` sashiko-bot
2026-09-18 5:44 ` bot+bpf-ci
2026-09-18 4:43 ` [PATCH bpf-next v2 20/20] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
19 siblings, 2 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:43 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
The end-to-end test walks one call chain with a pad in most of its frames.
This adds the shapes that chain does not reach: a callee called from both a
covered and an uncovered site, a pad that reads its frame's callee-saved
registers, a tail-call-reachable callee, a tail call that is taken, an
extension standing in for a covered call, a pad in the main program's own
frame, a record covering bpf_throw() itself, a record covering only a
nounwind call, a throwing subprogram named by a BPF_PSEUDO_FUNC, a region
ending on a 16-byte instruction, a pad that reloads from and writes to its
own frame, a pad two frames up, a pad that calls a subprogram which tail
calls, an extension carrying a table of its own, and a pad terminated by
_Unwind_Resume rather than bpf_unwind_resume.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
.../selftests/bpf/exceptions_cleanup.h | 20 +
.../bpf/prog_tests/exceptions_cleanup.c | 314 +++++++
.../bpf/progs/exceptions_cleanup_ext_table.c | 48 +
.../bpf/progs/exceptions_cleanup_freplace.c | 17 +
.../progs/exceptions_cleanup_pad_freplace.c | 17 +
.../bpf/progs/exceptions_cleanup_shapes.c | 863 ++++++++++++++++++
6 files changed, 1279 insertions(+)
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_ext_table.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_freplace.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_pad_freplace.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
index 630d2e207119..5896cf83d15e 100644
--- a/tools/testing/selftests/bpf/exceptions_cleanup.h
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -4,6 +4,7 @@
#define __EXCEPTIONS_CLEANUP_H__
#define THROW_COOKIE 0x100
+#define INNER_COOKIE 0x200
/* progs/exceptions_cleanup.c: one bit per frame that reports it ran. */
#define RAN_FOO3_PREEMPT 0x1
@@ -12,6 +13,25 @@
#define RAN_FOO2_DROP 0x8
#define RAN_BUMP 0x10
+/* progs/exceptions_cleanup_shapes.c: one bit per shape, numbered its own way. */
+#define RAN_SWEEP 0x1
+#define RAN_SHARED 0x2
+#define RAN_REGS 0x4
+#define RAN_TAIL_CALL 0x8
+#define RAN_MAIN_PAD 0x10
+#define RAN_TC_TAKEN 0x20
+#define RAN_FREPLACE 0x40
+#define RAN_ADDR_TAKEN 0x80
+#define RAN_NO_SUBPROG 0x100
+#define RAN_PAD_CALLS 0x200
+#define RAN_PAD_FIRST 0x400
+#define RAN_WIDE_REC 0x800
+#define RAN_PAD_STACK 0x1000
+#define RAN_DEEP_PAD 0x2000
+#define RAN_NOUNWIND_REC 0x4000
+#define RAN_RESUME_ALIAS 0x8000
+#define RAN_PAD_TAIL_CALL 0x10000
+
#define CLEANUP_REC(begin, end, landing_pad) \
".pushsection .bpf_cleanup,\"a\",@progbits;" \
".long " begin ";" \
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
index d1e45b765af2..ffc0b9519168 100644
--- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -4,6 +4,10 @@
#include "exceptions_cleanup.h"
#include "exceptions_cleanup.skel.h"
#include "exceptions_cleanup_fail.skel.h"
+#include "exceptions_cleanup_shapes.skel.h"
+#include "exceptions_cleanup_freplace.skel.h"
+#include "exceptions_cleanup_pad_freplace.skel.h"
+#include "exceptions_cleanup_ext_table.skel.h"
/* foo3 threw: every frame that has a pad ran it. */
#define PADS_FOO3_THREW \
@@ -35,6 +39,314 @@ static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
ASSERT_EQ(skel->bss->pads_ran, pads | RAN_BUMP, "pads_ran");
}
+static void run_shape(struct exceptions_cleanup_shapes *skel, struct bpf_program *prog,
+ __u64 input, __u32 retval, __u64 pads)
+{
+ __u64 ctx = 0;
+ int err;
+
+ LIBBPF_OPTS(bpf_test_run_opts, topts,
+ .ctx_in = &ctx,
+ .ctx_size_in = sizeof(ctx),
+ );
+
+ skel->bss->input = input;
+ skel->bss->pads_ran = 0;
+
+ err = bpf_prog_test_run_opts(bpf_program__fd(prog), &topts);
+ if (!ASSERT_OK(err, "run"))
+ return;
+ ASSERT_EQ(topts.retval, retval, "retval");
+ ASSERT_EQ(skel->bss->pads_ran, pads, "pads_ran");
+}
+
+static void test_freplace(struct exceptions_cleanup_shapes *skel)
+{
+ struct exceptions_cleanup_freplace *fr;
+ struct bpf_link *link;
+ int tgt_fd;
+
+ tgt_fd = bpf_program__fd(skel->progs.entry_freplace);
+
+ fr = exceptions_cleanup_freplace__open();
+ if (!ASSERT_OK_PTR(fr, "freplace open"))
+ return;
+
+ if (!ASSERT_OK(bpf_program__set_attach_target(fr->progs.new_fr_callee,
+ tgt_fd, "fr_callee"),
+ "set_attach_target"))
+ goto out;
+ if (!ASSERT_OK(exceptions_cleanup_freplace__load(fr), "freplace load"))
+ goto out;
+
+ link = bpf_program__attach_freplace(fr->progs.new_fr_callee, tgt_fd,
+ "fr_callee");
+ if (!ASSERT_OK_PTR(link, "attach_freplace"))
+ goto out;
+
+ run_shape(skel, skel->progs.entry_freplace, 101, THROW_COOKIE, 0);
+ bpf_link__destroy(link);
+out:
+ exceptions_cleanup_freplace__destroy(fr);
+}
+
+static void test_pad_calls_freplace(struct exceptions_cleanup_shapes *skel)
+{
+ struct exceptions_cleanup_pad_freplace *fr;
+ struct bpf_link *link;
+ __u64 ctx = 0;
+ int tgt_fd, err;
+
+ LIBBPF_OPTS(bpf_test_run_opts, topts,
+ .ctx_in = &ctx,
+ .ctx_size_in = sizeof(ctx),
+ );
+
+ tgt_fd = bpf_program__fd(skel->progs.entry_pad_calls);
+
+ fr = exceptions_cleanup_pad_freplace__open();
+ if (!ASSERT_OK_PTR(fr, "pad freplace open"))
+ return;
+
+ if (!ASSERT_OK(bpf_program__set_attach_target(fr->progs.new_pad_callee,
+ tgt_fd, "pad_callee"),
+ "set_attach_target"))
+ goto out;
+ if (!ASSERT_OK(exceptions_cleanup_pad_freplace__load(fr), "pad freplace load"))
+ goto out;
+
+ link = bpf_program__attach_freplace(fr->progs.new_pad_callee, tgt_fd,
+ "pad_callee");
+ if (!ASSERT_OK_PTR(link, "attach_freplace"))
+ goto out;
+
+ skel->bss->input = 101;
+ skel->bss->pads_ran = 0;
+ skel->bss->pad_runs = 0;
+
+ err = bpf_prog_test_run_opts(tgt_fd, &topts);
+ if (!ASSERT_OK(err, "run"))
+ goto out_link;
+
+ ASSERT_EQ(skel->bss->pad_runs, 1, "pad_runs");
+ ASSERT_EQ(skel->bss->pads_ran, RAN_PAD_CALLS, "pads_ran");
+ ASSERT_EQ(topts.retval, THROW_COOKIE, "retval");
+out_link:
+ bpf_link__destroy(link);
+out:
+ exceptions_cleanup_pad_freplace__destroy(fr);
+}
+
+static void test_ext_table(struct exceptions_cleanup_shapes *skel)
+{
+ struct exceptions_cleanup_ext_table *fr;
+ struct bpf_link *link;
+ int tgt_fd;
+
+ tgt_fd = bpf_program__fd(skel->progs.entry_freplace);
+
+ fr = exceptions_cleanup_ext_table__open();
+ if (!ASSERT_OK_PTR(fr, "ext table open"))
+ return;
+
+ if (!ASSERT_OK(bpf_program__set_attach_target(fr->progs.new_fr_callee,
+ tgt_fd, "fr_callee"),
+ "set_attach_target"))
+ goto out;
+ if (!ASSERT_OK(exceptions_cleanup_ext_table__load(fr), "ext table load"))
+ goto out;
+
+ link = bpf_program__attach_freplace(fr->progs.new_fr_callee, tgt_fd,
+ "fr_callee");
+ if (!ASSERT_OK_PTR(link, "attach_freplace"))
+ goto out;
+
+ fr->bss->ext_pad_ran = 0;
+ run_shape(skel, skel->progs.entry_freplace, 101, THROW_COOKIE, 0);
+ ASSERT_EQ(fr->bss->ext_pad_ran, 1, "ext_pad_ran");
+
+ bpf_link__destroy(link);
+out:
+ exceptions_cleanup_ext_table__destroy(fr);
+}
+
+static void test_shapes(void)
+{
+ struct exceptions_cleanup_shapes *skel;
+
+ skel = exceptions_cleanup_shapes__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "shapes open_and_load"))
+ return;
+
+ /* The frame loads at all only if everything unreachable in it went. */
+ if (test__start_subtest("sweep_no_throw"))
+ run_shape(skel, skel->progs.entry_sweep, 1, 0, 0);
+ if (test__start_subtest("sweep_throw"))
+ run_shape(skel, skel->progs.entry_sweep, 101, THROW_COOKIE, RAN_SWEEP);
+
+ /* The covered call unwinds to the pad; the uncovered one never does. */
+ if (test__start_subtest("shared_callee_no_throw"))
+ run_shape(skel, skel->progs.entry_shared, 1, 2, 0);
+ if (test__start_subtest("shared_callee_throw"))
+ run_shape(skel, skel->progs.entry_shared, 101, THROW_COOKIE, RAN_SHARED);
+
+ /* The pad only sets its bit if it got the frame's own r6-r9 back. */
+ if (test__start_subtest("pad_sees_callee_saved"))
+ run_shape(skel, skel->progs.entry_regs, 101, THROW_COOKIE, RAN_REGS);
+
+ /* Same check, with a tail-call-reachable callee: its spill moves. */
+ if (test__start_subtest("tail_call_no_throw"))
+ run_shape(skel, skel->progs.entry_tail_call, 1, 0, 0);
+ if (test__start_subtest("tail_call_throw"))
+ run_shape(skel, skel->progs.entry_tail_call, 101, THROW_COOKIE,
+ RAN_TAIL_CALL);
+
+ /* A region around a nounwind call: no pad dispatched, still loads. */
+ if (test__start_subtest("nounwind_region"))
+ run_shape(skel, skel->progs.entry_nounwind_rec, 1, 0, 0);
+
+ /* A pad in the main program's own frame, not in a subprogram. */
+ if (test__start_subtest("main_program_pad"))
+ run_shape(skel, skel->progs.entry_main_pad, 101, THROW_COOKIE,
+ RAN_MAIN_PAD);
+
+ /* The same call site either way: the subprogram's throw unwinds into
+ * this frame and runs its pad, an extension's stops at its own boundary.
+ */
+ if (test__start_subtest("freplace_subprog_throws"))
+ run_shape(skel, skel->progs.entry_freplace, 7, THROW_COOKIE,
+ RAN_FREPLACE);
+ if (test__start_subtest("freplace_extension_throws"))
+ test_freplace(skel);
+
+ /* A tail call that is taken: the walk ends at the target, so the cookie
+ * comes back from there and this frame's pad does not run.
+ */
+ if (test__start_subtest("tail_call_taken")) {
+ int key = 0, prog_fd = bpf_program__fd(skel->progs.tc_target);
+
+ if (ASSERT_OK(bpf_map_update_elem(bpf_map__fd(skel->maps.taken_table),
+ &key, &prog_fd, BPF_ANY),
+ "populate taken_table"))
+ run_shape(skel, skel->progs.entry_tail_taken, 101,
+ THROW_COOKIE, 0);
+ }
+
+ /* A throwing subprog named by a BPF_PSEUDO_FUNC no helper is handed: the
+ * callback check has to look at the bpf_loop(), not at the ld_imm64.
+ */
+ if (test__start_subtest("addr_taken_no_throw"))
+ run_shape(skel, skel->progs.entry_addr_taken, 1, 2, 0);
+ if (test__start_subtest("addr_taken_throw"))
+ run_shape(skel, skel->progs.entry_addr_taken, 101, THROW_COOKIE,
+ RAN_ADDR_TAKEN);
+
+ /* A record covering bpf_throw() itself rather than a call to a frame
+ * that throws: raised, caught up with and delivered in one frame.
+ */
+ if (test__start_subtest("no_subprog_no_throw"))
+ run_shape(skel, skel->progs.entry_no_subprog, 1, 0, 0);
+ if (test__start_subtest("no_subprog_throw"))
+ run_shape(skel, skel->progs.entry_no_subprog, 101, THROW_COOKIE,
+ RAN_NO_SUBPROG);
+
+ /* A pad that calls a subprogram; with a throwing extension in its place,
+ * the nested exception has to stop there, not restart this pad.
+ */
+ if (test__start_subtest("pad_calls_subprog")) {
+ skel->bss->pad_runs = 0;
+ run_shape(skel, skel->progs.entry_pad_calls, 101, THROW_COOKIE,
+ RAN_PAD_CALLS);
+ ASSERT_EQ(skel->bss->pad_runs, 1, "pad_runs");
+ }
+ if (test__start_subtest("pad_calls_throwing_extension"))
+ test_pad_calls_freplace(skel);
+
+ /* A covered throw the sweep leaves last, where the default exception
+ * callback is patched in; the pad's bit needs r6-r9 still spilled.
+ */
+ if (test__start_subtest("pad_before_throw"))
+ run_shape(skel, skel->progs.entry_pad_first, 101, THROW_COOKIE,
+ RAN_PAD_FIRST);
+
+ /* A region whose last instruction is a 16-byte one, so that end - 1
+ * names the half of it that is not an instruction.
+ */
+ if (test__start_subtest("region_ends_on_ldimm64"))
+ run_shape(skel, skel->progs.entry_wide_rec, 101, THROW_COOKIE,
+ RAN_WIDE_REC);
+
+ /* A pad that reloads from and writes to its own frame's stack, which a
+ * JIT addressing the frame through the stack pointer gets wrong.
+ */
+ if (test__start_subtest("pad_uses_own_frame"))
+ run_shape(skel, skel->progs.entry_pad_stack, 101, THROW_COOKIE,
+ RAN_PAD_STACK);
+
+ /* The same, with an uncovered frame between the throw and the pad. */
+ if (test__start_subtest("pad_two_frames_up"))
+ run_shape(skel, skel->progs.entry_deep_pad, 101, THROW_COOKIE,
+ RAN_DEEP_PAD);
+
+ /* An extension program with a cleanup table of its own. */
+ if (test__start_subtest("extension_carries_table"))
+ test_ext_table(skel);
+
+ /* A pad terminated by _Unwind_Resume, which libbpf maps onto the kfunc;
+ * every other program here calls bpf_unwind_resume directly.
+ */
+ if (test__start_subtest("resume_alias"))
+ run_shape(skel, skel->progs.entry_resume_alias, 101,
+ THROW_COOKIE, RAN_RESUME_ALIAS);
+
+ /* A pad that calls a subprogram which tail calls, array empty and then
+ * populated: the tail call releases only the callee's own prologue.
+ */
+ if (test__start_subtest("pad_callee_tail_call")) {
+ int key = 0, prog_fd = bpf_program__fd(skel->progs.pad_tc_target);
+
+ skel->bss->pad_tc_target_ran = 0;
+ skel->bss->pad_runs = 0;
+ run_shape(skel, skel->progs.entry_pad_tail_call, 101,
+ THROW_COOKIE, RAN_PAD_TAIL_CALL);
+ ASSERT_EQ(skel->bss->pad_tc_target_ran, 0, "target not run");
+ ASSERT_EQ(skel->bss->pad_runs, 1, "pad_runs");
+
+ if (ASSERT_OK(bpf_map_update_elem(bpf_map__fd(skel->maps.pad_tc_table),
+ &key, &prog_fd, BPF_ANY),
+ "populate pad_tc_table")) {
+ skel->bss->pad_runs = 0;
+ run_shape(skel, skel->progs.entry_pad_tail_call, 101,
+ THROW_COOKIE, RAN_PAD_TAIL_CALL);
+ ASSERT_EQ(skel->bss->pad_tc_target_ran, 1, "target ran");
+ ASSERT_EQ(skel->bss->pad_runs, 1, "pad_runs");
+ }
+ }
+
+ /* The same, into a target that carries a table and throws: that target
+ * is a boundary, so the outer pad runs once, not twice.
+ */
+ if (test__start_subtest("pad_callee_tail_call_throws")) {
+ int key = 0, prog_fd = bpf_program__fd(skel->progs.pad_tc_throw_target);
+
+ if (ASSERT_OK(bpf_map_update_elem(bpf_map__fd(skel->maps.pad_tc_table),
+ &key, &prog_fd, BPF_ANY),
+ "populate pad_tc_table")) {
+ skel->bss->pad_runs = 0;
+ skel->bss->tc_target_pad_runs = 0;
+ run_shape(skel, skel->progs.entry_pad_tail_call, 101,
+ THROW_COOKIE, RAN_PAD_TAIL_CALL);
+ /* The target cleaned up after itself, once. */
+ ASSERT_EQ(skel->bss->tc_target_pad_runs, 1,
+ "tc_target_pad_runs");
+ /* And the outer pad was not started over. */
+ ASSERT_EQ(skel->bss->pad_runs, 1, "pad_runs");
+ }
+ }
+
+ exceptions_cleanup_shapes__destroy(skel);
+}
+
void test_exceptions_cleanup(void)
{
char log[8192] = {};
@@ -75,5 +387,7 @@ void test_exceptions_cleanup(void)
exceptions_cleanup__destroy(skel);
+ test_shapes();
+
RUN_TESTS(exceptions_cleanup_fail);
}
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_ext_table.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_ext_table.c
new file mode 100644
index 000000000000..d14db48d6b29
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_ext_table.c
@@ -0,0 +1,48 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+__u64 ext_pad_ran = 0;
+
+/* Without a 32-bit int in BTF, libbpf's dummy_ksym var gets type id 0. */
+int btf_int_anchor;
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume();
+}
+
+static __used __naked __noinline __u64 ext_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = %[cookie];"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "r1 = %[ext_pad_ran] ll;"
+ "r2 = 1;"
+ "*(u64 *)(r1 + 0) = r2;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [cookie]"i"(THROW_COOKIE), __imm_addr(ext_pad_ran)
+ : __clobber_all);
+}
+
+SEC("freplace/fr_callee")
+__u64 new_fr_callee(__u64 x)
+{
+ return ext_frame();
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_freplace.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_freplace.c
new file mode 100644
index 000000000000..afb358fd3d40
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_freplace.c
@@ -0,0 +1,17 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "exceptions_cleanup.h"
+
+/* Without a 32-bit int in BTF, libbpf's dummy_ksym var gets type id 0. */
+int btf_int_anchor;
+
+SEC("freplace/fr_callee")
+__u64 new_fr_callee(__u64 x)
+{
+ bpf_throw(THROW_COOKIE);
+ return 0;
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_pad_freplace.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_pad_freplace.c
new file mode 100644
index 000000000000..eabac6baabb7
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_pad_freplace.c
@@ -0,0 +1,17 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "exceptions_cleanup.h"
+
+/* Without a 32-bit int in BTF, libbpf's dummy_ksym var gets type id 0. */
+int btf_int_anchor;
+
+SEC("freplace/pad_callee")
+__u64 new_pad_callee(__u64 x)
+{
+ bpf_throw(INNER_COOKIE);
+ return 0;
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
new file mode 100644
index 000000000000..f5eb2ff15c89
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
@@ -0,0 +1,863 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+#define PAD_COUNT \
+ "r1 = %[pad_runs] ll;" \
+ "r2 = *(u64 *)(r1 + 0);" \
+ "r2 += 1;" \
+ "*(u64 *)(r1 + 0) = r2;"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_rcu_read_lock();
+ bpf_rcu_read_unlock();
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume();
+}
+
+__u64 input = 0;
+__u64 magic = 0x5eed;
+__u64 pads_ran = 0;
+__u64 pad_runs = 0;
+
+/*
+ * 1. Everything a cleanup table leaves dead: the continuation after a throw,
+ * the tail after a pad's resume, an ld_imm64 and a conditional branch inside
+ * that tail, and a block reached only by the dead continuation.
+ */
+static __used __naked __noinline __u64 sweep_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "call bpf_preempt_disable;"
+ "if r6 < 101 goto 6f;"
+ "r1 = %[cookie];"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "goto 3f;"
+"4:" /* landing pad */
+ "r7 = r0;"
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "r1 = %[pads_ran] ll;"
+ "r2 = *(u64 *)(r1 + 0);"
+ "if r2 == 0 goto 5f;"
+ "call bpf_preempt_enable;"
+ "r0 = 7;"
+ "exit;"
+"5:"
+ "r0 = 8;"
+ "exit;"
+"3:" /* dead: only the dead goto reaches it */
+ "r0 = 9;"
+ "exit;"
+"6:" /* live: the ordinary return */
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "4b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_SWEEP),
+ __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_sweep(void *ctx)
+{
+ return sweep_frame();
+}
+
+/*
+ * 2. A callee called from both a covered and an uncovered site: the pad is
+ * recorded on the call site, not on the callee. The lock sits between the two
+ * calls because the frame really would leak it if the uncovered call unwound.
+ */
+static __used __noinline __u64 shared_callee(__u64 x)
+{
+ if (x > 100)
+ bpf_throw(THROW_COOKIE);
+ return x + 1;
+}
+
+static __used __naked __noinline __u64 shared_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "r1 = 0;"
+ "call shared_callee;"
+ "call bpf_rcu_read_lock;"
+ "r1 = r6;"
+"1:" "call shared_callee;" /* cleanup region */
+"2:"
+ "r6 = r0;"
+ "call bpf_rcu_read_unlock;"
+ "r0 = r6;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+ "call bpf_rcu_read_unlock;"
+ PAD_RAN("%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_SHARED), __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_shared(void *ctx)
+{
+ return shared_frame();
+}
+
+/*
+ * 3. A landing pad that reads its frame's callee-saved registers, which only
+ * the spill in the discarded callee's prologue still holds. The callee fills
+ * r6-r9 with something else before it throws, so the pad's check passes only
+ * if the walker found that spill.
+ */
+#define LOAD_MAGIC_REGS \
+ "r1 = %[magic] ll;" \
+ "r6 = *(u64 *)(r1 + 0);" \
+ "r7 = r6;" \
+ "r7 += 1;" \
+ "r8 = r6;" \
+ "r8 += 2;" \
+ "r9 = r6;" \
+ "r9 += 3;"
+
+/* Set @bit only if r6-r9 still hold what LOAD_MAGIC_REGS put there. */
+#define CHECK_MAGIC_REGS(bit) \
+ "r1 = %[magic] ll;" \
+ "r2 = *(u64 *)(r1 + 0);" \
+ "if r6 != r2 goto 9f;" \
+ "r2 += 1;" \
+ "if r7 != r2 goto 9f;" \
+ "r2 += 1;" \
+ "if r8 != r2 goto 9f;" \
+ "r2 += 1;" \
+ "if r9 != r2 goto 9f;" \
+ PAD_RAN(bit) \
+ "9:"
+
+static __used __naked __noinline __u64 regs_thrower(void)
+{
+ asm volatile (
+ /* Not this frame's to keep, and that is the point. */
+ "r6 = 0xdead;"
+ "r7 = 0xbeef;"
+ "r8 = 0xcafe;"
+ "r9 = 0xf00d;"
+ "r1 = %[cookie];"
+ "call bpf_throw;"
+ "r0 = 0;"
+ "exit;"
+ :
+ : [cookie]"i"(THROW_COOKIE)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 regs_frame(void)
+{
+ asm volatile (
+ LOAD_MAGIC_REGS
+ "call bpf_preempt_disable;"
+"1:" "call regs_thrower;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ CHECK_MAGIC_REGS("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_REGS),
+ __imm_addr(magic), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_regs(void *ctx)
+{
+ return regs_frame();
+}
+
+/*
+ * 4. The same, with a tail-call-reachable callee: its prologue pushes the tail
+ * call counter between the program stack and the spill area, so the spill the
+ * walker reads moves. The array is left empty; being reachable is the point.
+ */
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(__u32));
+ __uint(value_size, sizeof(__u32));
+} jmp_table SEC(".maps");
+
+static __used __noinline __u64 tc_thrower(void *ctx)
+{
+ /* Never taken; its presence is what makes this frame, whose spill the
+ * walker reads, tail-call-reachable.
+ */
+ bpf_tail_call_static(ctx, &jmp_table, 0);
+ asm volatile (
+ "r6 = 0xdead;"
+ "r7 = 0xbeef;"
+ "r8 = 0xcafe;"
+ "r9 = 0xf00d;"
+ "r1 = %[cookie];"
+ "call bpf_throw;"
+ :
+ : [cookie]"i"(THROW_COOKIE)
+ : __clobber_all);
+ return 0;
+}
+
+/*
+ * The frame with the pad is the program itself, and __naked: r1 holds the
+ * context at entry, which is the only place to get one for bpf_tail_call().
+ */
+SEC("syscall")
+__naked int entry_tail_call(void)
+{
+ asm volatile (
+ "*(u64 *)(r10 - 8) = r1;" /* the context, straight from entry */
+ LOAD_MAGIC_REGS
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 8f;"
+ "r1 = *(u64 *)(r10 - 8);"
+"1:" "call tc_thrower;" /* cleanup region */
+"2:"
+"8:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ CHECK_MAGIC_REGS("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_TAIL_CALL), __imm_addr(input),
+ __imm_addr(magic), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/*
+ * 5. A landing pad in the main program's own frame. jit_subprogs() compiles it
+ * as func[0], but the ksym the walker finds is the outer bpf_prog's, so the
+ * table has to be handed over or the pad is never dispatched -- silently.
+ */
+SEC("syscall")
+__naked int entry_main_pad(void)
+{
+ asm volatile (
+ LOAD_MAGIC_REGS
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 8f;"
+"1:" "call regs_thrower;" /* cleanup region */
+"2:"
+"8:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ CHECK_MAGIC_REGS("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_MAIN_PAD), __imm_addr(input),
+ __imm_addr(magic), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/*
+ * 6. A tail call that is really taken: the target is a program in its own
+ * right, so the walk ends there and this frame's pad does not run. The callee
+ * can also throw on a path never taken, which keeps the pad out of the sweep.
+ */
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(__u32));
+ __uint(value_size, sizeof(__u32));
+} taken_table SEC(".maps");
+
+SEC("syscall")
+int tc_target(void *ctx)
+{
+ bpf_throw(THROW_COOKIE);
+ return 0;
+}
+
+static __used __noinline __u64 tc_taken_callee(void *ctx, __u64 x)
+{
+ /* Never true at run time; the verifier cannot know that, and its
+ * unwind out of here is what keeps the caller's pad alive.
+ */
+ if (x == 7)
+ bpf_throw(THROW_COOKIE);
+ bpf_tail_call_static(ctx, &taken_table, 0);
+ return 0;
+}
+
+SEC("syscall")
+__naked int entry_tail_taken(void)
+{
+ asm volatile (
+ "*(u64 *)(r10 - 8) = r1;" /* the context, straight from entry */
+ "r1 = %[input] ll;"
+ "r2 = *(u64 *)(r1 + 0);"
+ "if r2 < 101 goto 8f;"
+ "r1 = *(u64 *)(r10 - 8);"
+"1:" "call tc_taken_callee;" /* cleanup region */
+"2:"
+ "exit;" /* the cookie, delivered at tc_target */
+"8:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad: must not run */
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_TC_TAKEN), __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/*
+ * 7. An extension program over the callee of a covered call. The walk ends in
+ * the extension's frame, as it does for a tail call target, so the pad does
+ * not run; fr_callee() can also throw by itself, giving the same call site
+ * both answers.
+ */
+__noinline __u64 fr_callee(__u64 x)
+{
+ if (x == 7)
+ bpf_throw(THROW_COOKIE);
+ return x + 1;
+}
+
+SEC("syscall")
+__naked int entry_freplace(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call fr_callee;" /* cleanup region */
+"2:"
+ "exit;"
+"3:" /* landing pad */
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_FREPLACE), __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/*
+ * 8. A throwing subprogram named by a BPF_PSEUDO_FUNC on a path never taken.
+ * Handing one to a helper is what is refused, not naming it, so anything going
+ * by the ld_imm64 alone turns this program away.
+ */
+static __used __noinline int cb_thrower(__u32 idx, void *ctx)
+{
+ bpf_throw(THROW_COOKIE);
+ return 0;
+}
+
+static __used __noinline __u64 addr_taken_callee(__u64 x)
+{
+ if (x <= 100)
+ return x + 1;
+ bpf_throw(THROW_COOKIE);
+ return bpf_loop(1, cb_thrower, NULL, 0);
+}
+
+SEC("syscall")
+__naked int entry_addr_taken(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call addr_taken_callee;" /* cleanup region */
+"2:"
+ "exit;"
+"3:" /* landing pad */
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_ADDR_TAKEN), __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/*
+ * 9. A record that covers bpf_throw() itself: the frame that raises the
+ * exception is the frame the record covers and the boundary both, so the pad
+ * runs on the way to delivering the cookie out of the program it came from.
+ */
+SEC("syscall")
+__naked int entry_no_subprog(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 8f;"
+ "call bpf_preempt_disable;"
+ "r1 = %[cookie];"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "exit;"
+"8:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_NO_SUBPROG),
+ __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/*
+ * 10. A landing pad that calls a subprogram an extension can replace. The
+ * load-time rule cannot see the extension coming, so what stops a nested
+ * exception is the walk, which ends in the extension's own frame. pad_runs
+ * says the pad ran once rather than twice.
+ */
+__noinline __u64 pad_callee(__u64 x)
+{
+ return x + 1;
+}
+
+static __used __noinline __u64 pc_thrower(__u64 x)
+{
+ if (x > 100)
+ bpf_throw(THROW_COOKIE);
+ return x + 1;
+}
+
+static __used __naked __noinline __u64 pad_calls_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call pc_thrower;" /* cleanup region */
+"2:"
+ "exit;"
+"3:" /* landing pad */
+ "r6 = r0;"
+ "r1 = 1;"
+ "call pad_callee;" /* an extension can stand in here */
+ PAD_COUNT
+ PAD_RAN("%[ran]")
+ "r1 = r6;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_PAD_CALLS), __imm_addr(input), __imm_addr(pads_ran),
+ __imm_addr(pad_runs)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_pad_calls(void *ctx)
+{
+ return pad_calls_frame();
+}
+
+/*
+ * 11. A covered bpf_throw() the sweep leaves as the last instruction, where
+ * the default exception callback is then patched in -- the one patchlet that
+ * does not keep the call it replaced in the last slot, so the marks have to
+ * follow it. The r6-r9 check is what reports a lost throw site mark.
+ */
+SEC("syscall")
+__naked int entry_pad_first(void)
+{
+ asm volatile (
+ LOAD_MAGIC_REGS
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 7f;"
+ "goto 4f;"
+"3:" /* landing pad, ahead of the call */
+ CHECK_MAGIC_REGS("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+"7:"
+ "r0 = 0;"
+ "exit;"
+"4:"
+ "r1 = %[cookie];"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "exit;" /* dead: swept, leaving the call last */
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_PAD_FIRST),
+ __imm_addr(input), __imm_addr(magic), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/*
+ * 12. A cleanup region whose last instruction is a 16-byte one, so end - 1
+ * names the half that is not an instruction of its own. A well formed region
+ * that a rule against it would turn away.
+ */
+static __used __naked __noinline __u64 wide_rec_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "call bpf_rcu_read_lock;"
+ "r1 = r6;"
+"1:" "call shared_callee;" /* cleanup region begins */
+ "r1 = %[magic] ll;" /* ... and ends on this pair */
+"2:"
+ "r6 = r0;"
+ "call bpf_rcu_read_unlock;"
+ "r0 = r6;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+ "call bpf_rcu_read_unlock;"
+ PAD_RAN("%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_WIDE_REC), __imm_addr(input), __imm_addr(magic),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_wide_rec(void *ctx)
+{
+ return wide_rec_frame();
+}
+
+/*
+ * 13. A pad that works out of its own frame's stack, the shape every
+ * compiler-generated pad has. A JIT that addresses the frame through the
+ * stack pointer -- arm64 -- has to address a pad's frame some other way. Both
+ * directions are here: the reload sees the frame, and the store lands in it.
+ */
+static __used __naked __noinline __u64 pad_stack_frame(void)
+{
+ asm volatile (
+ "r1 = %[magic] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "*(u64 *)(r10 - 8) = r1;" /* what the pad will want */
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call pc_thrower;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r6 = r0;"
+ "r7 = *(u64 *)(r10 - 8);" /* reload it out of the frame */
+ "*(u64 *)(r10 - 16) = r7;" /* and write the frame while here */
+ "r1 = %[magic] ll;"
+ "r2 = *(u64 *)(r1 + 0);"
+ "if r7 != r2 goto 9f;"
+ "r3 = *(u64 *)(r10 - 16);"
+ "if r3 != r2 goto 9f;"
+ PAD_RAN("%[ran]")
+"9:"
+ "r1 = r6;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_PAD_STACK), __imm_addr(input), __imm_addr(magic),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_pad_stack(void *ctx)
+{
+ return pad_stack_frame();
+}
+
+/*
+ * 14. The same, with a frame in between that has no pad of its own, so the
+ * liveness query for an outer frame has more than one frame to walk and the
+ * pad has to be counted at every step.
+ */
+static __used __noinline __u64 deep_mid(__u64 x)
+{
+ return pc_thrower(x) + 1;
+}
+
+static __used __naked __noinline __u64 deep_frame(void)
+{
+ asm volatile (
+ "r1 = %[magic] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "*(u64 *)(r10 - 8) = r1;" /* nothing but the pad reads this */
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call deep_mid;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r6 = r0;"
+ "r7 = *(u64 *)(r10 - 8);"
+ "r1 = %[magic] ll;"
+ "r2 = *(u64 *)(r1 + 0);"
+ "if r7 != r2 goto 9f;"
+ PAD_RAN("%[ran]")
+"9:"
+ "r1 = r6;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_DEEP_PAD), __imm_addr(input), __imm_addr(magic),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_deep_pad(void *ctx)
+{
+ return deep_frame();
+}
+
+/*
+ * 15. A region around a call the kernel knows cannot unwind: no call site is
+ * marked, nothing reaches the pad, and the sweep removes it. The program is
+ * otherwise ordinary and has to load.
+ */
+static __used __naked __noinline __u64 nounwind_rec_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call bpf_preempt_enable;" /* cleanup region: nounwind */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad, never dispatched */
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_NOUNWIND_REC), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_nounwind_rec(void *ctx)
+{
+ return nounwind_rec_frame();
+}
+
+/*
+ * 16. The name a frontend gives the resume. Every pad above calls
+ * bpf_unwind_resume(); LLVM emits _Unwind_Resume() and libbpf maps one onto
+ * the other, so this program is what keeps that mapping tested.
+ */
+extern void _Unwind_Resume(void) __ksym;
+
+static __used __noinline void __resume_alias_btf_anchor(void)
+{
+ _Unwind_Resume();
+}
+
+static __used __naked __noinline __u64 resume_alias_frame(void)
+{
+ asm volatile (
+"1:" "call regs_thrower;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ PAD_RAN("%[ran]")
+ "call _Unwind_Resume;" /* the frontend's name for it */
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_RESUME_ALIAS), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_resume_alias(void *ctx)
+{
+ return resume_alias_frame();
+}
+
+/*
+ * 17. A landing pad that calls a subprogram which tail calls. What a pad may
+ * not contain is a tail call of its own, which would unwind a prologue the
+ * walker's stack never held; a callee's prologue really did run there, so its
+ * tail call releases exactly that and the target returns into the pad. The
+ * tail call counter comes out of the unwinding frame, which is one of the
+ * pad's own subprogram and so really holds one.
+ */
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(__u32));
+ __uint(value_size, sizeof(__u32));
+} pad_tc_table SEC(".maps");
+
+__u64 pad_tc_target_ran = 0;
+
+SEC("syscall")
+int pad_tc_target(void *ctx)
+{
+ pad_tc_target_ran += 1;
+ return 0;
+}
+
+static __used __noinline __u64 pad_tc_callee(void *ctx)
+{
+ /* Taken only once the test has populated the array. */
+ bpf_tail_call_static(ctx, &pad_tc_table, 0);
+ return 0;
+}
+
+/*
+ * The frame with the pad is the program itself, and __naked: r1 holds the
+ * context at entry, which is the only place to get one for bpf_tail_call().
+ * The pad reloads it from its own frame's stack.
+ */
+SEC("syscall")
+__naked int entry_pad_tail_call(void)
+{
+ asm volatile (
+ "*(u64 *)(r10 - 8) = r1;" /* the context, straight from entry */
+ LOAD_MAGIC_REGS
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 8f;"
+"1:" "call regs_thrower;" /* cleanup region */
+"2:"
+"8:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ PAD_COUNT
+ "r1 = *(u64 *)(r10 - 8);"
+ "call pad_tc_callee;"
+ /* Only if the frame survived the callee's tail call. */
+ CHECK_MAGIC_REGS("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_PAD_TAIL_CALL), __imm_addr(input),
+ __imm_addr(magic), __imm_addr(pads_ran), __imm_addr(pad_runs)
+ : __clobber_all);
+}
+
+/*
+ * 18. The other target for that same tail call: a program carrying a cleanup
+ * table of its own, which throws while the outer exception is still in flight.
+ * The tail call made it a boundary, so the inner walk runs its pad and ends in
+ * its own frame, never reaching the walker's frames above it: the outer pad is
+ * not restarted and the outer cookie is still the one delivered. The outer
+ * pad's r6-r9, which this target overwrites, come back with its frame.
+ */
+__u64 tc_target_pad_runs = 0;
+__u64 inner_magic = 0xd00d;
+
+static __used __naked __noinline __u64 inner_thrower(void)
+{
+ asm volatile (
+ /* Not this frame's to keep, the same as regs_thrower. */
+ "r6 = 0xf00d;"
+ "r7 = 0xcafe;"
+ "r8 = 0xbeef;"
+ "r9 = 0xdead;"
+ "r1 = %[cookie];"
+ "call bpf_throw;"
+ "r0 = 0;"
+ "exit;"
+ :
+ : [cookie]"i"(INNER_COOKIE)
+ : __clobber_all);
+}
+
+SEC("syscall")
+__naked int pad_tc_throw_target(void)
+{
+ asm volatile (
+ /* Distinct from the outer pad's, so neither can stand in for it. */
+ "r1 = %[inner_magic] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "r7 = r6;"
+ "r7 += 1;"
+ "r8 = r6;"
+ "r8 += 2;"
+ "r9 = r6;"
+ "r9 += 3;"
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 8f;"
+"1:" "call inner_thrower;" /* cleanup region */
+"2:"
+"8:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ /* This frame's own r6-r9, not the outer pad's. */
+ "r1 = %[inner_magic] ll;"
+ "r2 = *(u64 *)(r1 + 0);"
+ "if r6 != r2 goto 9f;"
+ "r2 += 1;"
+ "if r7 != r2 goto 9f;"
+ "r2 += 1;"
+ "if r8 != r2 goto 9f;"
+ "r2 += 1;"
+ "if r9 != r2 goto 9f;"
+ "r1 = %[tc_target_pad_runs] ll;"
+ "r2 = *(u64 *)(r1 + 0);"
+ "r2 += 1;"
+ "*(u64 *)(r1 + 0) = r2;"
+"9:"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm_addr(input), __imm_addr(inner_magic),
+ __imm_addr(tc_target_pad_runs)
+ : __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* [PATCH bpf-next v2 20/20] selftests/bpf: Load an exception cleanup program from a light skeleton
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (18 preceding siblings ...)
2026-09-18 4:43 ` [PATCH bpf-next v2 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
@ 2026-09-18 4:43 ` Yonghong Song
19 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-18 4:43 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
For a light skeleton, libbpf hands the records to bpf_gen__prog_load(),
which writes them into a blob and emits a loader program that issues
BPF_PROG_LOAD from inside the kernel. Nothing about that is shared with the
ordinary path: the attr is built field by field, and the kfunc names are
resolved by the loader program when it runs.
exceptions_cleanup_light.c is the smallest program that can tell whether a
table survives that trip: one record covering the bpf_throw() call itself,
one pad that sets a bit and resumes. If the table arrives, the pad runs and
the cookie comes back; if it does not, the program does not load at all.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/testing/selftests/bpf/Makefile | 2 +-
.../selftests/bpf/exceptions_cleanup.h | 3 ++
.../bpf/prog_tests/exceptions_cleanup.c | 28 +++++++++++++
.../bpf/progs/exceptions_cleanup_light.c | 39 +++++++++++++++++++
4 files changed, 71 insertions(+), 1 deletion(-)
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
diff --git a/tools/testing/selftests/bpf/Makefile b/tools/testing/selftests/bpf/Makefile
index 7ea5ba1df29e..981cc2643049 100644
--- a/tools/testing/selftests/bpf/Makefile
+++ b/tools/testing/selftests/bpf/Makefile
@@ -523,7 +523,7 @@ LINKED_SKELS := test_static_linked.skel.h linked_funcs.skel.h \
LSKELS := fexit_sleep.c trace_printk.c trace_vprintk.c map_ptr_kern.c \
core_kern.c core_kern_overflow.c test_ringbuf.c \
test_ringbuf_n.c test_ringbuf_map_key.c test_ringbuf_write.c \
- test_ringbuf_overwrite.c
+ test_ringbuf_overwrite.c exceptions_cleanup_light.c
LSKELS_SIGNED := fentry_test.c fexit_test.c atomics.c
diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
index 5896cf83d15e..0c088d96ca01 100644
--- a/tools/testing/selftests/bpf/exceptions_cleanup.h
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -32,6 +32,9 @@
#define RAN_RESUME_ALIAS 0x8000
#define RAN_PAD_TAIL_CALL 0x10000
+/* progs/exceptions_cleanup_light.c: the one pad it has. */
+#define RAN_LIGHT 0x1
+
#define CLEANUP_REC(begin, end, landing_pad) \
".pushsection .bpf_cleanup,\"a\",@progbits;" \
".long " begin ";" \
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
index ffc0b9519168..ad0ff949f10d 100644
--- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -8,6 +8,7 @@
#include "exceptions_cleanup_freplace.skel.h"
#include "exceptions_cleanup_pad_freplace.skel.h"
#include "exceptions_cleanup_ext_table.skel.h"
+#include "exceptions_cleanup_light.lskel.h"
/* foo3 threw: every frame that has a pad ran it. */
#define PADS_FOO3_THREW \
@@ -170,6 +171,30 @@ static void test_ext_table(struct exceptions_cleanup_shapes *skel)
exceptions_cleanup_ext_table__destroy(fr);
}
+static void test_light_skeleton(void)
+{
+ struct exceptions_cleanup_light_lskel *skel;
+ __u64 ctx = 0;
+ int err;
+
+ LIBBPF_OPTS(bpf_test_run_opts, topts,
+ .ctx_in = &ctx,
+ .ctx_size_in = sizeof(ctx),
+ );
+
+ skel = exceptions_cleanup_light_lskel__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "light open_and_load"))
+ return;
+
+ err = bpf_prog_test_run_opts(skel->progs.entry_light.prog_fd, &topts);
+ if (!ASSERT_OK(err, "run"))
+ goto out;
+ ASSERT_EQ(topts.retval, THROW_COOKIE, "retval");
+ ASSERT_EQ(skel->bss->pads_ran, RAN_LIGHT, "pads_ran");
+out:
+ exceptions_cleanup_light_lskel__destroy(skel);
+}
+
static void test_shapes(void)
{
struct exceptions_cleanup_shapes *skel;
@@ -387,6 +412,9 @@ void test_exceptions_cleanup(void)
exceptions_cleanup__destroy(skel);
+ if (test__start_subtest("light_skeleton"))
+ test_light_skeleton();
+
test_shapes();
RUN_TESTS(exceptions_cleanup_fail);
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
new file mode 100644
index 000000000000..de7ae478d67f
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
@@ -0,0 +1,39 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume();
+}
+
+__u64 pads_ran = 0;
+
+SEC("syscall")
+__naked int entry_light(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = %[cookie];"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_LIGHT),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts
2026-09-18 4:43 ` [PATCH bpf-next v2 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
@ 2026-09-18 4:57 ` sashiko-bot
2026-09-19 20:18 ` Yonghong Song
0 siblings, 1 reply; 56+ messages in thread
From: sashiko-bot @ 2026-09-18 4:57 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] libbpf: missing E2BIG fallback for cleanup_info breaks forward compatibility
--
commit d67b69538e39a329708fcc9fa943f0ad05afaa05
Author: Yonghong Song <yonghong.song@linux.dev>
libbpf: Add cleanup_info to bpf_prog_load_opts
Let a caller hand the kernel an exception cleanup table. Extend struct
bpf_prog_load_opts with cleanup_info, cleanup_info_cnt and
cleanup_info_rec_size, pass them through to BPF_PROG_LOAD, and grow the
attr size bpf_prog_load() computes to cover the new fields.
> diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
> index 96819c082c772..bcf490570961c 100644
> --- a/tools/lib/bpf/bpf.c
> +++ b/tools/lib/bpf/bpf.c
[ ... ]
> @@ -370,6 +370,10 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
> attr.fd_array = ptr_to_u64(OPTS_GET(opts, fd_array, NULL));
> attr.fd_array_cnt = OPTS_GET(opts, fd_array_cnt, 0);
>
> + attr.cleanup_info = ptr_to_u64(OPTS_GET(opts, cleanup_info, NULL));
> + attr.cleanup_info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
> + attr.cleanup_info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
[Severity: Medium]
Will this break forward compatibility for bpf_prog_load() if cleanup_info
triggers an -E2BIG error?
Later in bpf_prog_load(), there is a retry loop to handle -E2BIG when the
userspace-provided record size is larger than what the kernel supports:
while (errno == E2BIG && (!finfo || !linfo)) {
if (!finfo && attr.func_info_cnt &&
attr.func_info_rec_size < func_info_rec_size) {
/* try with corrected func info records */
...
The loop condition and body don't seem to include checks for cleanup_info.
If the kernel returns -E2BIG because of a larger cleanup_info_rec_size,
won't this hit the else { break; } block or skip the loop entirely, causing
the load to fail instead of retrying with truncated records?
> +
> if (log_level) {
> attr.log_buf = ptr_to_u64(log_buf);
> attr.log_size = log_size;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260918044156.3283973-1-yonghong.song@linux.dev?part=14
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test
2026-09-18 4:43 ` [PATCH bpf-next v2 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
@ 2026-09-18 4:59 ` sashiko-bot
2026-09-18 5:58 ` bot+bpf-ci
1 sibling, 0 replies; 56+ messages in thread
From: sashiko-bot @ 2026-09-18 4:59 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] bpf: multi-line comments do not follow modern BPF style
--
commit 8ad5e7ea23b9dcb55029206d5cb36453b231f9c8
Author: Yonghong Song <yonghong.song@linux.dev>
selftests/bpf: Add an end-to-end .bpf_cleanup exception test
This commit adds an end-to-end test for .bpf_cleanup exception handling. It
uses inline assembly to construct precise call-site brackets, landing pads,
and cleanup records that mirror compiler output. It also tests shapes that
the kernel should refuse, ensuring the JIT correctly handles or rejects
different landing pad configurations.
> diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> new file mode 100644
> index 0000000000000..d1e45b765af26
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
[ ... ]
> @@ -0,0 +1,79 @@
[ ... ]
> + /* foo3 throws; every pad runs and the cookie is delivered at entry. */
> + if (test__start_subtest("throw_from_foo3"))
> + run(skel, 101, THROW_COOKIE, PADS_FOO3_THREW);
> +
> + /* foo3 returns 2 ^ 1 == 3, so foo2 throws from its own second region;
> + * foo3's frame is long gone, so its pad must not run.
> + */
[Severity: Low]
Does this comment block follow the proper multi-line comment formatting
for the BPF subsystem?
The BPF subsystem strict guidelines require multi-line comments to have
the opening /* on its own line:
/*
* foo3 returns 2 ^ 1 == 3, so foo2 throws from its own second region;
* foo3's frame is long gone, so its pad must not run.
*/
> + if (test__start_subtest("throw_from_foo2"))
> + run(skel, 2, THROW_COOKIE, PADS_FOO2_THREW);
> +
> + exceptions_cleanup__destroy(skel);
> +
> + RUN_TESTS(exceptions_cleanup_fail);
> +}
> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
> new file mode 100644
> index 0000000000000..ce2dac306a84c
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
[ ... ]
> @@ -0,0 +1,600 @@
[ ... ]
> +SEC("socket")
> +__failure __msg("reaches an indirect jump")
> +__naked void gotox_in_pad(void)
> +{
> + asm volatile (
> + ".pushsection .jumptables,\"\",@progbits;"
> +"jt0_%=:"
> + ".quad l0_%= - socket;"
> + ".quad l1_%= - socket;"
> + ".size jt0_%=, 16;"
> + ".global jt0_%=;"
> + ".popsection;"
> +
> +"1:" "call gotox_thrower;" /* cleanup region */
> +"2:"
> + "r0 = 0;"
> + "exit;"
> +"3:" /* landing pad */
> + "r1 = jt0_%= ll;"
> + "r1 += 8;"
> + "r2 = *(u64 *)(r1 + 0);"
> + /* gotox r2. Spelled as a raw insn on purpose: the "gotox" mnemonic
> + * only reached the LLVM assembler in llvm 22, and BPF_RAW_INSN()
> + * needs <linux/bpf.h>, which this file cannot have -- vmlinux.h
> + * already defines the uapi enums.
> + */
[Severity: Low]
Does this multi-line comment format conform to the BPF style guide?
As with the previous comment in exceptions_cleanup.c, the BPF subsystem
mandates that the opening /* resides on its own line for multi-line
comments.
> + ".8byte 0x20d;"
> +"l0_%=:"
> + "call bpf_unwind_resume;"
> + "exit;"
> +"l1_%=:"
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + ::: __clobber_all);
> +}
> +
> +#endif /* x86 || arm64 */
> +
> +char _license[] SEC("license") = "GPL";
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260918044156.3283973-1-yonghong.song@linux.dev?part=18
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 05/20] bpf: Make exception landing pads reachable in the CFG
2026-09-18 4:42 ` [PATCH bpf-next v2 05/20] bpf: Make exception landing pads reachable in the CFG Yonghong Song
@ 2026-09-18 4:59 ` sashiko-bot
2026-09-19 19:17 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
1 sibling, 1 reply; 56+ messages in thread
From: sashiko-bot @ 2026-09-18 4:59 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: missing prune point at exception landing pad target
--
commit 4fbd0e36862ba7e3ee8480e535068a6821d5cad0
Author: Yonghong Song <yonghong.song@linux.dev>
bpf: Make exception landing pads reachable in the CFG
This commit updates the verifier's control flow graph and liveness analysis
to properly handle exception landing pads. It ensures that exception throwing
calls can reach cleanup records, and prevents outer frame stack slots from
being read incorrectly when control is transferred to landing pads.
> diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
> index 842c7d1eabccc..9f8b8b54d5ea7 100644
> --- a/kernel/bpf/cfg.c
> +++ b/kernel/bpf/cfg.c
[ ... ]
> @@ -158,17 +159,57 @@ static int push_insn(int t, int w, int e, struct bpf_verifier_env *env)
> return DONE_EXPLORING;
> }
>
> +static int visit_cleanup_pad_edge(int t, struct bpf_verifier_env *env)
> +{
> + int *insn_stack = env->cfg.insn_stack;
> + int *insn_state = env->cfg.insn_state;
> + int w;
> +
> + if (!env->cleanup_info_cnt)
> + return DONE_EXPLORING;
> + w = bpf_cleanup_pad_of_call(env, t);
> + if (w < 0)
> + return DONE_EXPLORING;
> +
> + mark_prune_point(env, t);
> + mark_jmp_point(env, w);
> + mark_jump_target(env, w);
[Severity: High]
Should mark_prune_point() be applied to the landing pad target w instead of
the call site t?
Because w is a branch target where control flow merges, it needs to be marked
as a prune point so the verifier knows to check for state convergence. By
marking t instead of w, the verifier will evaluate the landing pad and all
subsequent cleanup instructions independently for every exception-throwing
call that branches to w.
Can this lead to combinatorial path explosion and result in the verifier
rejecting the program with -E2BIG?
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260918044156.3283973-1-yonghong.song@linux.dev?part=5
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel
2026-09-18 4:43 ` [PATCH bpf-next v2 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
@ 2026-09-18 5:00 ` sashiko-bot
2026-09-19 20:21 ` Yonghong Song
0 siblings, 1 reply; 56+ messages in thread
From: sashiko-bot @ 2026-09-18 5:00 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] bpf: empty cleanup range check in bpf_prog_collect_cleanup_info() only catches offset 0
- [Medium] bpf: bpf_object__init_cleanup_info() rejects valid objects with 0-sized .bpf_cleanup sections
--
commit a645d3bd249aceee4909edcc78e5fe4247fd3432
Author: Yonghong Song <yonghong.song@linux.dev>
libbpf: Collect .bpf_cleanup records and pass them to the kernel
Parse the compiler-emitted .bpf_cleanup section and hand the resulting
table to BPF_PROG_LOAD.
> diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
> index 27f2eec13f76f..1b8b978e7b7b9 100644
> --- a/tools/lib/bpf/libbpf.c
> +++ b/tools/lib/bpf/libbpf.c
[ ... ]
> @@ -4852,6 +4890,221 @@ static struct bpf_program *find_prog_by_sec_insn(const struct bpf_object *obj,
> return NULL;
> }
>
> +static int bpf_object__init_cleanup_info(struct bpf_object *obj)
> +{
> + Elf_Data *data = obj->efile.cleanup_data;
> + Elf_Data *relo = NULL;
> + size_t i, nrels, nslots, nrecs;
> + struct cleanup_raw_rec *recs;
> + int *slot_sec, ret = 0;
> + size_t *slot_val;
> + const __u32 *vals;
> + bool native;
> +
> + if (!data || obj->efile.cleanup_shndx < 0)
> + return 0;
> +
> + native = is_native_endianness(obj);
> +
> + for (i = 0; i < obj->efile.sec_cnt; i++) {
> + struct elf_sec_desc *sd = &obj->efile.secs[i];
> +
> + if (sd->sec_type == SEC_RELO && sd->shdr &&
> + sd->shdr->sh_info == (Elf64_Word)obj->efile.cleanup_shndx) {
> + relo = sd->data;
> + break;
> + }
> + }
> + if (!relo) {
> + pr_warn("%s present without relocations\n", CLEANUP_SEC);
> + return -LIBBPF_ERRNO__FORMAT;
> + }
[Severity: Medium]
Could this reject valid objects if the compiler emits an empty 0-byte
.bpf_cleanup section?
If the section is empty (e.g., exceptions were optimized out), no relocations
are emitted and there is no corresponding .rel.bpf_cleanup section. However,
bpf_object__init_cleanup_info() requires the relocation section to exist if
.bpf_cleanup is present.
Would it be better to return early if data->d_size is 0 before checking for
relocations?
[ ... ]
> +static int bpf_prog_collect_cleanup_info(struct bpf_object *obj,
> + struct bpf_program *prog)
> +{
> + size_t i;
> + int j;
> +
> + for (i = 0; i < obj->cleanup_rec_cnt; i++) {
> + struct cleanup_raw_rec *raw = &obj->cleanup_recs[i];
> + struct bpf_program *owner = NULL;
> + struct bpf_cleanup_info ci = {};
> + __u32 fields[CLEANUP_REC_FIELDS];
> + void *tmp;
> +
> + for (j = 0; j < CLEANUP_REC_FIELDS; j++) {
> + size_t idx = raw->insn_idx[j], final;
> + struct bpf_program *p;
> +
> + /* The end of a range is exclusive, so it may name the
> + * instruction just past the last one of a function,
> + * which belongs to the next function or to nothing at
> + * all. Ask about the last instruction the range covers,
> + * the way the kernel does.
> + */
> + if (j == CLEANUP_REC_END) {
> + if (!idx) {
> + pr_warn("%s: record %zu is an empty range\n",
> + CLEANUP_SEC, i);
> + return -LIBBPF_ERRNO__FORMAT;
> + }
> + idx--;
> + }
[Severity: Medium]
Will this check catch empty ranges emitted by the compiler at non-zero
offsets?
If the compiler emits a .bpf_cleanup record where begin and end are both 5,
the check for !idx evaluates to false. The code will then decrement idx to 4
and look up the program for the previous instruction.
If this empty range is located at the start of a function, looking up the
previous instruction will fall into the preceding function, potentially
triggering a spurious "spans functions" error or silently accepting invalid
records.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260918044156.3283973-1-yonghong.song@linux.dev?part=15
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach
2026-09-18 4:43 ` [PATCH bpf-next v2 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
@ 2026-09-18 5:01 ` sashiko-bot
2026-09-18 5:44 ` bot+bpf-ci
1 sibling, 0 replies; 56+ messages in thread
From: sashiko-bot @ 2026-09-18 5:01 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] bpf: multi-line comment style violation
--
commit 59a07b8ba8f5c3976a5da9bea3b5d6359874813a
Author: Yonghong Song <yonghong.song@linux.dev>
selftests/bpf: Cover the exception cleanup shapes the chain does not reach
This commit adds end-to-end exception cleanup tests for various edge cases
and shapes that the primary call chain does not reach. It tests things like
tail calls, nounwind calls, subprogram pads, and landing pads that interact
with the stack.
> diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> index d1e45b765af26..ffc0b95191681 100644
> --- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> +++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
[ ... ]
> @@ -35,6 +39,314 @@ static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
> +static void test_shapes(void)
> +{
[ ... ]
> + /* The same call site either way: the subprogram's throw unwinds into
> + * this frame and runs its pad, an extension's stops at its own boundary.
> + */
[Severity: Low]
This isn't a bug, but does this comment follow the BPF subsystem coding style?
Multi-line comments should have their opening /* on its own line before the
text begins.
> + if (test__start_subtest("freplace_subprog_throws"))
> + run_shape(skel, skel->progs.entry_freplace, 7, THROW_COOKIE,
> + RAN_FREPLACE);
> + if (test__start_subtest("freplace_extension_throws"))
> + test_freplace(skel);
> +
> + /* A tail call that is taken: the walk ends at the target, so the cookie
> + * comes back from there and this frame's pad does not run.
> + */
[Severity: Low]
This isn't a bug, but is this the preferred multi-line comment format? The BPF
subsystem coding style mandates that the opening /* should be on its own line.
> + if (test__start_subtest("tail_call_taken")) {
> + int key = 0, prog_fd = bpf_program__fd(skel->progs.tc_target);
[ ... ]
> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
> new file mode 100644
> index 0000000000000..f5eb2ff15c896
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
[ ... ]
> @@ -0,0 +1,863 @@
> +static __used __noinline __u64 tc_thrower(void *ctx)
> +{
> + /* Never taken; its presence is what makes this frame, whose spill the
> + * walker reads, tail-call-reachable.
> + */
[Severity: Low]
This isn't a bug, but could this comment formatting be adjusted? Multi-line
comments in the BPF subsystem should place their opening /* on a separate line.
> + bpf_tail_call_static(ctx, &jmp_table, 0);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260918044156.3283973-1-yonghong.song@linux.dev?part=19
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations
2026-09-18 4:43 ` [PATCH bpf-next v2 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
@ 2026-09-18 5:01 ` sashiko-bot
2026-09-19 20:31 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
1 sibling, 1 reply; 56+ messages in thread
From: sashiko-bot @ 2026-09-18 5:01 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] bpf/linker: NULL pointer deref in linker_append_elf_relos()
- [Medium] bpf/linker: unaligned memory access in data relocations
- [Low] bpf/linker: multi-line comment style violation
--
commit 2eadf13e135621b2e34d8946fe67664cd72e0e3d
Author: Yonghong Song <yonghong.song@linux.dev>
libbpf: Let the static linker carry .bpf_cleanup relocations
This commit updates the BPF static linker to support .bpf_cleanup relocations.
It specifically adds handling for R_BPF_64_NODYLD32 and R_BPF_64_ABS32 to
process compiler-emitted exception cleanup tables.
> diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
> index 78f92c39290af..e5c06023cb5b1 100644
> --- a/tools/lib/bpf/linker.c
> +++ b/tools/lib/bpf/linker.c
[ ... ]
> @@ -2274,6 +2275,22 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
> insn->imm += sec->dst_off / sizeof(struct bpf_insn);
> else
> insn->imm += sec->dst_off;
> + } else if (sym_type == R_BPF_64_NODYLD32 ||
> + sym_type == R_BPF_64_ABS32) {
> + __u32 *val;
> +
> + /* Two spellings of the one thing: LLVM
[Severity: Low]
This isn't a bug, but does this violate the BPF subsystem multi-line comment
style? The preferred kernel style is to have the opening /* on its own line
rather than putting text on the same line as the marker.
> + * emits NODYLD32 for a .long against a
> + * section symbol, GNU as emits ABS32
> + * (bpf_reloc_type_lookup() maps
> + * BFD_RELOC_32 to it), and the value
> + * they describe is the same.
> + */
> + val = dst_linked_sec->raw_data + dst_rel->r_offset;
[Severity: High]
Could this lead to a NULL pointer dereference? Looking at
linker_sanity_check_elf_relos(), SHT_NOBITS sections are explicitly allowed
as relocatable sections.
For SHT_NOBITS sections (like .bss), extend_sec() typically skips allocating
raw_data, leaving it as NULL. If we process a relocation against such a
section here, we might compute val = NULL + dst_rel->r_offset and then
dereference it on the next line when handling sec->dst_off.
[Severity: Medium]
Is it also possible to hit an unaligned memory access here?
While executable sections enforce alignment checks in
linker_sanity_check_elf_relos(), there doesn't seem to be an equivalent
alignment check for data section offsets. If an ELF provides an unaligned
r_offset, the direct dereference of the unaligned val pointer might
cause a crash on architectures with strict memory alignment requirements.
> + if (linker->swapped_endian)
> + *val = bswap_32(bswap_32(*val) + sec->dst_off);
> + else
> + *val += sec->dst_off;
> } else {
> pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
> return -EINVAL;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260918044156.3283973-1-yonghong.song@linux.dev?part=17
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 16/20] libbpf: Carry the exception cleanup table through the light skeleton
2026-09-18 4:43 ` [PATCH bpf-next v2 16/20] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
@ 2026-09-18 5:02 ` sashiko-bot
2026-09-19 20:27 ` Yonghong Song
0 siblings, 1 reply; 56+ messages in thread
From: sashiko-bot @ 2026-09-18 5:02 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] libbpf: integer overflow in cleanup_info_tot_sz leads to heap OOB read/write
--
commit b9300a1f6022da2aa8b343e05882b46e9f329c64
Author: Yonghong Song <yonghong.song@linux.dev>
libbpf: Carry the exception cleanup table through the light skeleton
This patch modifies the light skeleton generation to carry the exception
cleanup table (cleanup_info), similar to func_info and line_info. It updates
bpf_gen__prog_load() to allocate space for the records, set up union bpf_attr,
and perform endianness swapping for cross-compiled cases.
> diff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c
> index af3a04f161ac1..2345fbdd46f56 100644
> --- a/tools/lib/bpf/gen_loader.c
> +++ b/tools/lib/bpf/gen_loader.c
[ ... ]
> @@ -998,6 +1000,9 @@ static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
>
> for (i = 0; i < gen->core_relo_cnt; i++)
> bpf_core_relo_bswap(cr++);
> +
> + for (i = 0; i < load_attr->cleanup_info_cnt; i++)
> + bpf_cleanup_info_bswap(ci++);
[Severity: High]
Could this loop result in an out-of-bounds heap read and write?
If an attacker crafts a BPF ELF object with an extremely large .bpf_cleanup
section, load_attr->cleanup_info_cnt can be very large. If the allocation
size calculated in bpf_gen__prog_load() overflowed, this loop will process
hundreds of millions of elements in an undersized buffer.
This could cause memory corruption when cross-compiling (gen->swapped_endian
is true).
> }
>
> void bpf_gen__prog_load(struct bpf_gen *gen,
> @@ -1011,8 +1016,11 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
> load_attr->line_info_rec_size;
> int core_relo_tot_sz = gen->core_relo_cnt *
> sizeof(struct bpf_core_relo);
> + int cleanup_info_tot_sz = load_attr->cleanup_info_cnt *
> + load_attr->cleanup_info_rec_size;
[Severity: High]
Does this unbounded 32-bit integer multiplication risk an overflow?
If load_attr->cleanup_info_cnt is excessively large, multiplying it by
cleanup_info_rec_size could wrap around modulo 2^32 to a small positive
integer. This would result in add_data() allocating an undersized heap
buffer for the cleanup info blob.
Combined with the loop in info_blob_bswap(), this could lead to crashes
or potentially arbitrary code execution when processing untrusted files
using the light skeleton generator.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260918044156.3283973-1-yonghong.song@linux.dev?part=16
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 11/20] bpf, x86: Dispatch exception cleanup pads at run time
2026-09-18 4:42 ` [PATCH bpf-next v2 11/20] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
@ 2026-09-18 5:03 ` sashiko-bot
2026-09-19 20:00 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
1 sibling, 1 reply; 56+ messages in thread
From: sashiko-bot @ 2026-09-18 5:03 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf, x86: garbage kfunc stack arguments in exception cleanup pads
--
commit 1e32297a3c2c4f6bbc4d458bc25563bed8167f6c
Author: Yonghong Song <yonghong.song@linux.dev>
bpf, x86: Dispatch exception cleanup pads at run time
This commit provides the architecture-specific support for x86-64 exception
cleanup pads. It forces a full callee-saved spill for programs with a cleanup
table and introduces arch_bpf_run_cleanup_pad to execute the pads.
> diff --git a/arch/x86/net/bpf_cleanup_pad.S b/arch/x86/net/bpf_cleanup_pad.S
> new file mode 100644
> index 0000000000000..da4b448ecf095
> --- /dev/null
> +++ b/arch/x86/net/bpf_cleanup_pad.S
[ ... ]
> +SYM_FUNC_START(arch_bpf_run_cleanup_pad)
> + ANNOTATE_NOENDBR
> +
> + pushq %rbp
> + pushq %rbx
> + pushq %r12
> + pushq %r13
> + pushq %r14
> + pushq %r15
> + /* Keep the pad's entry rsp congruent to a normal call's. */
> + subq $8, %rsp
> +
> + movq 0(%rdx), %r15
> + movq 8(%rdx), %r14
> + movq 16(%rdx), %r13
> + movq 24(%rdx), %rbx
> + movq 32(%rdx), %r12
> + /* rbp is BPF r10, so this is the whole of the pad's frame setup. */
> + movq %rsi, %rbp
> +
> + /* CALL_NOSPEC needs the target in a register; rcx is BPF r4, dead. */
> + movq %rdi, %rcx
> +
> + /* BPF r0 on the way into a pad, not whatever the kernel left in rax. */
> + movl $BPF_PAD_ENTRY_R0, %eax
> +
> + CALL_NOSPEC rcx
[Severity: High]
Does this execution environment cause kfuncs with more than 5 arguments
to read uninitialized garbage stack arguments?
In arch_bpf_run_cleanup_pad, the pad is executed on the unwinder's
stack (rsp), while rbp is set to the original BPF frame.
However, the JIT unconditionally writes outgoing kfunc stack arguments
relative to BPF_REG_FP (rbp), as seen in emit_kfunc_arg_moves in
arch/x86/net/bpf_jit_comp.c:
else if (dst_mem)
emit_stx(&prog, BPF_DW, BPF_REG_FP, reg,
stack_base + (dst - nreg) * 8);
Because rbp and rsp are desynchronized in the pad, the arguments are
written to the old frame. When the kfunc is called, the C calling
convention requires it to read stack arguments from [rsp], which points
to the unwinder's stack instead of the original frame, leading to
garbage reads.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260918044156.3283973-1-yonghong.song@linux.dev?part=11
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations
2026-09-18 4:43 ` [PATCH bpf-next v2 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-18 5:01 ` sashiko-bot
@ 2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 20:32 ` Yonghong Song
1 sibling, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-18 5:44 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 1257 bytes --]
> An object that carries a compiler-emitted exception cleanup table cannot be
> linked today. The table's fields are byte offsets into a code section,
> materialised by a 32-bit relocation against that section's symbol with the
> offset itself as the implicit addend, and the linker rejects both halves of
> that: the relocation type is not in the list it accepts, and a relocation
> against an STT_SECTION symbol from a non-executable section is an outright
> error.
>
> Both spellings of that relocation have to be taken. LLVM emits
> R_BPF_64_NODYLD32 for a .long against a section symbol; GNU as emits
> R_BPF_64_ABS32, which is what bpf_reloc_type_lookup() maps BFD_RELOC_32 to.
> They describe the same value, and the selftests are built with both
> compilers.
>
> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
This looks like a bug fix for objects that carry compiler-emitted exception
cleanup tables. Should this include:
Fixes: 9f0c317f6aa1 ("libbpf: Add support for new llvm bpf relocations")
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 03/20] bpf: Add lookups for exception cleanup resumes and landing pads
2026-09-18 4:42 ` [PATCH bpf-next v2 03/20] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
@ 2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 4:55 ` Alexei Starovoitov
0 siblings, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-18 5:44 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3799 bytes --]
> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index c08505b9ba821..f9bccd3e0f4db 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -681,6 +681,11 @@ struct bpf_insn_aux_data {
> bool needs_zext; /* alu op needs to clear upper bits */
> bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
> bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
> + /*
> + * 1 + the instruction index of the exception cleanup landing pad this
> + * call site unwinds to, or 0 for none.
> + */
> + u32 cleanup_pad;
> bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
> u8 alu_state; /* used in combination with alu_limit */
[ ... ]
> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
> index 2add8001c3ec3..82b00fac6bd6a 100644
> --- a/kernel/bpf/fixups.c
> +++ b/kernel/bpf/fixups.c
> @@ -261,6 +261,11 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
> }
> }
>
> + if (env->cleanup_info_cnt)
> + for (i = 0; i < prog_len; i++)
> + if (data[i].cleanup_pad > off + 1)
> + data[i].cleanup_pad += cnt - 1;
> +
[ ... ]
> @@ -549,6 +554,7 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
> struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
> unsigned int orig_prog_len = env->prog->len;
> int err;
> + u32 i;
>
> if (bpf_prog_is_offloaded(env->prog->aux))
> bpf_prog_offload_remove_insns(env, off, cnt);
> @@ -573,6 +579,17 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
> sizeof(*aux_data) * (orig_prog_len - off - cnt));
> env->insn_aux_data_len -= cnt;
>
> + if (env->cleanup_info_cnt) {
> + for (i = 0; i < env->insn_aux_data_len; i++) {
> + u32 pad = aux_data[i].cleanup_pad;
> +
> + if (pad > off + cnt)
> + aux_data[i].cleanup_pad = pad - cnt;
> + else if (pad > off)
> + aux_data[i].cleanup_pad = 0;
> + }
> + }
> +
> return 0;
> }
Can this logic handle NOP removal correctly? When bpf_opt_remove_nops()
calls verifier_remove_insns(env, i, 1) to remove a single fall-through
goto +0, if a cleanup record names that instruction as its landing pad,
then pad == off + 1, so pad > off is true and cleanup_pad becomes 0.
But removing the NOP leaves the pad's next instruction at the same
index, so the correct landing pad index stays unchanged rather than
being zeroed. For example:
off: some_call with cleanup_pad = off+1
off+1: goto +0 (NOP landing pad)
off+2: actual cleanup code
After removing the NOP at off+1:
off: some_call (cleanup_pad should still be off+1)
off+1: actual cleanup code (was at off+2)
The `else if (pad > off)` arm erases the record instead of preserving
it. Nothing rejects NOP landing pads during validation - a goto +0 has
code 0x05, not 0x00, so check_cleanup_info() accepts it.
The ordering makes the loss escape validation: bpf_check_cleanup_exceptions()
runs before bpf_opt_remove_nops(), so the verifier proves the cleanup
path releases acquired references, but then the NOP removal zeroes
cleanup_pad, and later cleanup_info_for_subprog() sees `if (!pad)
continue;` and omits the record from the runtime table. The program is
accepted but bpf_throw() finds no pad for that call site and skips the
cleanup, leaking whatever the pad was supposed to release.
Does the `else if (pad > off)` arm need to distinguish removals that
consume the landing pad instruction itself versus removals that happen
before it?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 05/20] bpf: Make exception landing pads reachable in the CFG
2026-09-18 4:42 ` [PATCH bpf-next v2 05/20] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-18 4:59 ` sashiko-bot
@ 2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 19:32 ` Yonghong Song
1 sibling, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-18 5:44 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3394 bytes --]
> diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
> index 44ecdc5b4ec2d..9cfd05f970bc9 100644
> --- a/kernel/bpf/liveness.c
> +++ b/kernel/bpf/liveness.c
> @@ -8,6 +8,8 @@
> #include <linux/slab.h>
> #include <linux/sort.h>
>
> +#include "exception.h"
> +
> #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
>
> struct per_frame_masks {
> @@ -256,6 +258,9 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
> succ = env->succ;
> succ->cnt = 0;
>
> + if (unlikely(bpf_is_unwind_resume_kfunc(insn)))
> + return succ;
> +
Should this early return be gated on env->cleanup_info_cnt?
bpf_unwind_resume() is an unconditionally registered generic kfunc with no
KF_ flag restricting it to programs that carry a cleanup table. Its body
is empty, so at run time the call simply returns and execution continues
at the next instruction.
visit_insn() in cfg.c still gives the call a fallthrough edge, and at this
commit nothing in do_check() rejects a bpf_unwind_resume() call that is not
the tail of a cleanup landing pad. So any CAP_BPF program with no cleanup
table can call bpf_unwind_resume() and keep executing after it.
For such a program, bpf_insn_successors() reports zero successors, so
bpf_compute_live_registers() treats everything after the call as dead. The
register-cleaning side fails closed (a poisoned slot is rejected), but the
32-bit zero-extension side fails open:
insn_aux[i].zext_dst for an insn preceding the call is only set when the
hi half of the defined register is live-after, so it stays false, and
bpf_opt_subreg_zext_lo32_rnd_hi32() (kernel/bpf/fixups.c:706) then does
'if (!aux[adj_idx].zext_dst) { if (!rnd_hi32) continue; ... }' and emits
no BPF_ZEXT_REG.
On architectures where bpf_jit_needs_zext() is true and kfunc calls are
supported (s390, riscv64, powerpc64, parisc, x86-32), the upper 32 bits of
a register the verifier proved to be 32-bit bounded are left with stale
JIT-defined contents, which defeats the range checks the verifier derived
for it.
Compare with the pad lookup ten lines below:
> opcode_info = &opcode_info_tbl[BPF_CLASS(insn->code) | BPF_OP(insn->code)];
> insn_sz = bpf_is_ldimm64(insn) ? 2 : 1;
> if (opcode_info->can_fallthrough)
> @@ -264,6 +269,13 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
> if (opcode_info->can_jump)
> succ->items[succ->cnt++] = idx + bpf_jmp_offset(insn) + 1;
>
> + if (unlikely(env->cleanup_info_cnt)) {
> + int pad = bpf_cleanup_pad_of_call(env, idx);
> +
> + if (pad >= 0)
> + succ->items[succ->cnt++] = pad;
> + }
which gates the pad lookup on cleanup_info_cnt, and with
visit_cleanup_pad_edge() in kernel/bpf/cfg.c which returns DONE_EXPLORING
when !env->cleanup_info_cnt.
Would gating the early return on env->cleanup_info_cnt fix this? That
would be sound because a program that carries a cleanup table is
additionally constrained (bpf_jit_supports_cleanup_pads() plus, from commit
fa6890df61539 onward, the requirement that a resume sit at the end of a
landing pad), whereas a program with cleanup_info_cnt == 0 must keep the
ordinary fallthrough successor of the call.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 09/20] bpf: Refuse a private stack for a program with an exception cleanup table
2026-09-18 4:42 ` [PATCH bpf-next v2 09/20] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
@ 2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 19:37 ` Yonghong Song
0 siblings, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-18 5:44 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3141 bytes --]
> A landing pad is about to start running outside its own frame, on the
> bpf_throw() walker's stack, with that frame's registers put back from a
> spill area whose contents arch_bpf_run_cleanup_pad() knows how to read. The
> x86-64 JIT addresses a private-stack program's frame through a scratch
> register that it recomputes after each call rather than through rbp, and no
> spill area holds that register, so a pad there would address its frame
> through whatever the kernel left behind.
The description says the x86-64 JIT 'recomputes after each call'. Does the
JIT actually recompute it?
Looking at the x86 JIT, emit_priv_frame_ptr() (arch/x86/net/bpf_jit_comp.c)
is emitted once in the prologue, and the value in X86_REG_R9 is carried
across calls by push_r9()/pop_r9() around each call site; BPF_REG_FP
accesses are rewritten to X86_REG_R9.
The load-bearing half of the argument is still correct -
arch_bpf_run_cleanup_pad() (added later in the series,
arch/x86/net/bpf_cleanup_pad.S) restores only r15/r14/r13/rbx/r12 and rbp
from the spill area and never restores r9, so a pad in a private-stack
program would address its frame through a stale register.
Only the described mechanism (recompute vs save/restore around calls) is
wrong, which makes the justification harder to check against the JIT.
> Everywhere rather than arch-conditionally: arm64 keeps its private
> stack pointer in x27, which is in the prologue spill and so survives, but a
> rule that holds on one arch and not the other is not worth the second code
> path when nothing is lost but an optimization.
Does 'nothing is lost but an optimization' understate the effect?
The private-stack mode also changes the verifier's stack-depth accounting,
not just JIT codegen. In check_max_stack_depth_subprog() a
PRIV_STACK_ADAPTIVE subprog is checked against MAX_BPF_STACK on its own ('if
(subprog_depth > MAX_BPF_STACK)'), while a NO_PRIV_STACK subprog accumulates
into 'depth' and the whole bpf2bpf chain must fit in MAX_BPF_STACK
('combined stack size of %d calls is %d. Too large').
So for the prog types bpf_enable_priv_stack() would have returned
PRIV_STACK_ADAPTIVE for (kprobe/tracepoint/perf_event/raw_tp, and
tracing/lsm/struct_ops with a recursion or an explicit request), a program
carrying a cleanup table whose nested subprog depths sum above 512 bytes now
fails to load rather than losing performance.
Worth stating in the changelog, since it is a load-time visible restriction
on the new feature rather than an optimization.
Related: kernel/sched/ext/ext.c sets prog->aux->priv_stack_requested for
sched_ext dispatch/sub_caps_updated, and verifier.c rejects such a request
when the JIT cannot honour it; here the request is silently ignored instead.
This looks like a bug fix for private stack support. Should this include:
Fixes: a76ab5731e32d ("bpf: Find eligible subprogs for private stack support")
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 06/20] bpf: Explore the landing pads no call site reaches
2026-09-18 4:42 ` [PATCH bpf-next v2 06/20] bpf: Explore the landing pads no call site reaches Yonghong Song
@ 2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 19:32 ` Yonghong Song
0 siblings, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-18 5:44 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 2152 bytes --]
> commit c93117f24a1a9cc39bfeaef79c42dbff02d8c455
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Explore the landing pads no call site reaches
>
> A cleanup record need not cover a call an exception can unwind out of: a
> frontend is free to emit a region around a helper or an ordinary kfunc,
> both nounwind here. Nothing marks a call site then, and that record's
> landing pad is reached by nothing at all -- leaving bpf_check_cfg() to
> refuse the program over code its own frontend had no way not to emit:
>
> 0: call bpf_preempt_disable
> 1: call bpf_preempt_enable record = { begin = 1, end = 2, pad = 4 }
> 2: r0 = 0
> 3: exit
> 4: r1 = pads_ran ll landing pad
> 6: r2 = *(u64 *)(r1 + 0)
> 7: r2 |= RAN_NOUNWIND_REC
> 8: *(u64 *)(r1 + 0) = r2
> 9: call bpf_unwind_resume
> 10: exit
>
> The range [1,2) holds one call, and it is a kfunc, so an exception cannot
> come out of it. cleanup_mark_call_sites() marks nothing, nothing pushes an
> edge to 4, and 4 through 10 are reachable from nothing: "unreachable insn
> 4". The pad is dead, which is correct -- no exception can ever arrive at it
> -- but the program is fine and has to load.
>
> Walk every pad the table names that the edges did not reach, the way the
> walk is already re-seeded at an exception callback. From there the pad is
> code like any other: do_check() never enters it, because no call site
> dispatches to it, so the dead code sweep removes it along with everything
> else that was not reached.
>
> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
This looks like a bug fix for the earlier commit that made exception landing
pads reachable in the CFG. Should this include:
Fixes: a67011bbbca1ba95 ("bpf: Make exception landing pads reachable in the CFG")
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 11/20] bpf, x86: Dispatch exception cleanup pads at run time
2026-09-18 4:42 ` [PATCH bpf-next v2 11/20] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-18 5:03 ` sashiko-bot
@ 2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 20:04 ` Yonghong Song
1 sibling, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-18 5:44 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 4703 bytes --]
> diff --git a/arch/x86/net/Makefile b/arch/x86/net/Makefile
> index dddbefc0f4398..9d574d972df37 100644
> --- a/arch/x86/net/Makefile
> +++ b/arch/x86/net/Makefile
> @@ -6,5 +6,5 @@
> ifeq ($(CONFIG_X86_32),y)
> obj-$(CONFIG_BPF_JIT) += bpf_jit_comp32.o
> else
> - obj-$(CONFIG_BPF_JIT) += bpf_jit_comp.o bpf_timed_may_goto.o
> + obj-$(CONFIG_BPF_JIT) += bpf_jit_comp.o bpf_timed_may_goto.o bpf_cleanup_pad.o
> endif
> diff --git a/arch/x86/net/bpf_cleanup_pad.S b/arch/x86/net/bpf_cleanup_pad.S
> new file mode 100644
> index 0000000000000..da4b448ecf095
> --- /dev/null
> +++ b/arch/x86/net/bpf_cleanup_pad.S
[ ... ]
> +SYM_FUNC_START(arch_bpf_run_cleanup_pad)
> + ANNOTATE_NOENDBR
> +
> + pushq %rbp
> + pushq %rbx
> + pushq %r12
> + pushq %r13
> + pushq %r14
> + pushq %r15
> + /* Keep the pad's entry rsp congruent to a normal call's. */
> + subq $8, %rsp
> +
> + movq 0(%rdx), %r15
> + movq 8(%rdx), %r14
> + movq 16(%rdx), %r13
> + movq 24(%rdx), %rbx
> + movq 32(%rdx), %r12
> + /* rbp is BPF r10, so this is the whole of the pad's frame setup. */
> + movq %rsi, %rbp
> +
> + /* CALL_NOSPEC needs the target in a register; rcx is BPF r4, dead. */
> + movq %rdi, %rcx
> +
> + /* BPF r0 on the way into a pad, not whatever the kernel left in rax. */
> + movl $BPF_PAD_ENTRY_R0, %eax
> +
> + CALL_NOSPEC rcx
> +
> + addq $8, %rsp
> + popq %r15
> + popq %r14
> + popq %r13
> + popq %r12
> + popq %rbx
> + popq %rbp
> + RET
> +SYM_FUNC_END(arch_bpf_run_cleanup_pad)
Will this build with CONFIG_STACK_VALIDATION enabled?
arch_bpf_run_cleanup_pad() is declared with SYM_FUNC_START (an STT_FUNC
symbol) and saves rbp with 'pushq %rbp', but never establishes a frame
pointer with 'mov %rsp, %rbp'. Instead, it overwrites rbp with the BPF
frame pointer (movq %rsi, %rbp) and then makes a call (CALL_NOSPEC rcx).
objtool's frame-pointer validation (CONFIG_STACK_VALIDATION with
UNWINDER_FRAME_POINTER) checks for two conditions in validate_call()
and validate_return():
tools/objtool/check.c:has_valid_stack_frame() requires cfi->cfa.base ==
CFI_BP, which is only set when the function does 'mov %rsp, %rbp' to
establish a frame. This function never does that, so the CALL_NOSPEC
would trigger:
WARN_INSN(insn, "call without frame pointer save/setup");
Additionally, the 'pushq %rbp' while cfa->base != CFI_BP sets
cfi->bp_scratch = true, which is only cleared by 'mov %rsp, %rbp'.
This flag is still set at the RET, where validate_return() warns:
WARN_INSN(insn, "BP used as a scratch register");
The Makefile adds bpf_cleanup_pad.o unconditionally under
obj-$(CONFIG_BPF_JIT), even though the feature is gated on
CONFIG_UNWINDER_ORC per bpf_jit_supports_cleanup_pads(). This means
in a CONFIG_STACK_VALIDATION=y build (frame-pointer unwinder), the
object would fail objtool validation, and with CONFIG_OBJTOOL_WERROR=y
the build would fail.
The neighbouring arch/x86/net/bpf_timed_may_goto.S does not have this
issue because it sets up a proper frame:
arch/x86/net/bpf_timed_may_goto.S:
pushq %rbp
movq %rsp, %rbp
...
leave
rbp cannot be the frame pointer here because it must carry BPF r10 into
the pad. This is the same situation arch/x86/kernel/ftrace_64.S handles
with STACK_FRAME_NON_STANDARD_FP(ftrace_caller). Could either
STACK_FRAME_NON_STANDARD_FP(arch_bpf_run_cleanup_pad) after SYM_FUNC_END
or building the object only when CONFIG_UNWINDER_ORC is set (matching the
feature gate) address this?
> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
> index d4a980140b48d..9d0dd54773e83 100644
> --- a/arch/x86/net/bpf_jit_comp.c
> +++ b/arch/x86/net/bpf_jit_comp.c
> @@ -357,6 +357,11 @@ struct jit_context {
> /* Number of bytes that will be skipped on tailcall */
> #define X86_TAIL_CALL_OFFSET (12 + ENDBR_INSN_SIZE)
>
> +/* Throw-site spill: r15, r14, r13, rbx, r12 low to high, the layout the
> + * prologue's pushes leave, so arch_bpf_run_cleanup_pad() reads both alike.
> + */
> +#define X86_CLEANUP_SPILL_SZ (5 * 8)
This isn't a bug, but could the comment format be updated to match the
BPF subsystem style? The BPF subsystem guide requires the modern kernel
multi-line comment style, where the opening /* is on its own line. The
patch already uses this format for its other new multi-line comments
(the "Lowest address of each spill area" block and the "Spill r6-r9 and
r12" block later in do_jit), so is there a reason for the inconsistency
here?
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach
2026-09-18 4:43 ` [PATCH bpf-next v2 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
2026-09-18 5:01 ` sashiko-bot
@ 2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 21:13 ` Yonghong Song
1 sibling, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-18 5:44 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 8331 bytes --]
> diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
> index 630d2e207119a..5896cf83d15ee 100644
> --- a/tools/testing/selftests/bpf/exceptions_cleanup.h
> +++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
> @@ -4,6 +4,7 @@
> #define __EXCEPTIONS_CLEANUP_H__
>
> #define THROW_COOKIE 0x100
> +#define INNER_COOKIE 0x200
>
> /* progs/exceptions_cleanup.c: one bit per frame that reports it ran. */
> #define RAN_FOO3_PREEMPT 0x1
[ ... ]
> diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> index d1e45b765af26..ffc0b95191681 100644
> --- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> +++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
[ ... ]
> +static void test_shapes(void)
> +{
> + struct exceptions_cleanup_shapes *skel;
> +
> + skel = exceptions_cleanup_shapes__open_and_load();
> + if (!ASSERT_OK_PTR(skel, "shapes open_and_load"))
> + return;
[ ... ]
> + /* A throwing subprog named by a BPF_PSEUDO_FUNC no helper is handed: the
> + * callback check has to look at the bpf_loop(), not at the ld_imm64.
> + */
> + if (test__start_subtest("addr_taken_no_throw"))
> + run_shape(skel, skel->progs.entry_addr_taken, 1, 2, 0);
> + if (test__start_subtest("addr_taken_throw"))
> + run_shape(skel, skel->progs.entry_addr_taken, 101, THROW_COOKIE,
> + RAN_ADDR_TAKEN);
Does this comment accurately describe what the addr_taken shape tests?
Looking at addr_taken_callee() in progs/exceptions_cleanup_shapes.c, the
function does hand cb_thrower to bpf_loop(). The ld_imm64 and the
bpf_loop() call are on the same path, after bpf_throw(). Since bpf_throw()
is not declared noreturn, LLVM keeps that tail, and what actually makes
the program load is that the verifier never reaches the bpf_loop() at all
(it is dead after the throw).
So push_callback_call() -> bpf_cleanup_check_callback() never fires.
The shape exercised is 'a BPF_PSEUDO_FUNC whose helper call the verifier
never reaches', not 'a BPF_PSEUDO_FUNC no helper is handed'. Could the
comment be rephrased to reflect what the verifier actually sees?
[ ... ]
> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
> new file mode 100644
> index 0000000000000..f5eb2ff15c896
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
[ ... ]
> +/*
> + * 2. A callee called from both a covered and an uncovered site: the pad is
> + * recorded on the call site, not on the callee. The lock sits between the two
> + * calls because the frame really would leak it if the uncovered call unwound.
> + */
> +static __used __noinline __u64 shared_callee(__u64 x)
> +{
> + if (x > 100)
> + bpf_throw(THROW_COOKIE);
> + return x + 1;
> +}
> +
> +static __used __naked __noinline __u64 shared_frame(void)
> +{
> + asm volatile (
> + "r1 = %[input] ll;"
> + "r6 = *(u64 *)(r1 + 0);"
> + "r1 = 0;"
> + "call shared_callee;"
> + "call bpf_rcu_read_lock;"
> + "r1 = r6;"
> +"1:" "call shared_callee;" /* cleanup region */
> +"2:"
> + "r6 = r0;"
> + "call bpf_rcu_read_unlock;"
> + "r0 = r6;"
> + "exit;"
> +"3:" /* landing pad */
> + "r7 = r0;"
> + "call bpf_rcu_read_unlock;"
> + PAD_RAN("%[ran]")
> + "r1 = r7;"
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + :
> + : [ran]"i"(RAN_SHARED), __imm_addr(input), __imm_addr(pads_ran)
> + : __clobber_all);
> +}
Does the uncovered call site verify that the pad wouldn't be wrongly
dispatched if an exception occurred there?
The uncovered site is reached with the constant 0 in r1, and shared_callee
only throws when x > 100, so the verifier proves that call cannot throw
and no unwind ever leaves the uncovered site - at verification time or at
run time.
The structural half of the property (only the second call's return address
falls inside the record's native range) is exercised, but the behavioral
half is not: a kernel that wrongly matched the uncovered site's return
address against the record and ran the pad would go unnoticed, because
that site never unwinds.
[ ... ]
> +/*
> + * 6. A tail call that is really taken: the target is a program in its own
> + * right, so the walk ends there and this frame's pad does not run. The callee
> + * can also throw on a path never taken, which keeps the pad out of the sweep.
> + */
> +struct {
> + __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
> + __uint(max_entries, 1);
> + __uint(key_size, sizeof(__u32));
> + __uint(value_size, sizeof(__u32));
> +} taken_table SEC(".maps");
> +
> +SEC("syscall")
> +int tc_target(void *ctx)
> +{
> + bpf_throw(THROW_COOKIE);
> + return 0;
> +}
> +
> +static __used __noinline __u64 tc_taken_callee(void *ctx, __u64 x)
> +{
> + /* Never true at run time; the verifier cannot know that, and its
> + * unwind out of here is what keeps the caller's pad alive.
> + */
> + if (x == 7)
> + bpf_throw(THROW_COOKIE);
> + bpf_tail_call_static(ctx, &taken_table, 0);
> + return 0;
> +}
> +
> +SEC("syscall")
> +__naked int entry_tail_taken(void)
> +{
> + asm volatile (
> + "*(u64 *)(r10 - 8) = r1;" /* the context, straight from entry */
> + "r1 = %[input] ll;"
> + "r2 = *(u64 *)(r1 + 0);"
> + "if r2 < 101 goto 8f;"
> + "r1 = *(u64 *)(r10 - 8);"
> +"1:" "call tc_taken_callee;" /* cleanup region */
> +"2:"
> + "exit;" /* the cookie, delivered at tc_target */
> +"8:"
> + "r0 = 0;"
> + "exit;"
> +"3:" /* landing pad: must not run */
> + PAD_RAN("%[ran]")
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + :
> + : [ran]"i"(RAN_TC_TAKEN), __imm_addr(input), __imm_addr(pads_ran)
> + : __clobber_all);
> +}
Can the tail_call_taken subtest fail if the kernel wrongly continues the
walk past the tail-call boundary?
entry_tail_taken narrows the argument before the covered call:
"r2 = *(u64 *)(r1 + 0);" /* r2 = input */
"if r2 < 101 goto 8f;" /* fallthrough => r2 in [101, U64_MAX] */
tc_taken_callee is a static subprog, so check_func_call() copies the
caller's r1-r5 verbatim into the callee frame. x therefore arrives with
umin_value == 101, and is_branch_taken() resolves 'if (x == 7)' to
never-taken, so the verifier never explores the bpf_throw(). Nothing then
unwinds into entry_tail_taken during verification, the landing pad at
label 3 is never marked seen, and opt_remove_dead_code() removes it.
fixups.c then clears the annotation for any call site whose cleanup_pad
falls inside the removed window:
if (env->cleanup_info_cnt) {
for (i = 0; i < env->insn_aux_data_len; i++) {
u32 pad = aux_data[i].cleanup_pad;
if (pad > off + cnt) aux_data[i].cleanup_pad = pad - cnt;
else if (pad > off) aux_data[i].cleanup_pad = 0;
}
}
so no native range is emitted for the region at all. The program still
loads (mark_subprog_might_throw() is a static analysis, so bpf_check_cfg()
still sees the pad edge and check_cleanup_info() runs before the sweep),
and pads_ran stays 0 unconditionally. A kernel that wrongly continued
bpf_stack_walker() past the tail-call boundary into entry_tail_taken's
frame would find no record to match and would still leave pads_ran == 0,
so the subtest passes on a correct and on a broken kernel alike.
The in-code comment is also wrong on both clauses: the verifier CAN know
x != 7 here, and consequently the unwind out of tc_taken_callee is NOT
what keeps the caller's pad alive - nothing does.
Contrast with shape 7 (entry_freplace, lines 862-878), the analogous 'walk
ends in the callee's frame' shape, which deliberately passes the
unnarrowed input value to fr_callee() with no 'if r1 < 101' guard, keeping
fr_callee's identical 'if (x == 7) bpf_throw()' live. Dropping the
'if r2 < 101 goto 8f' guard from entry_tail_taken (or otherwise passing a
value the verifier cannot exclude 7 from) would make the pad survive and
give the subtest something to fail on.
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 12/20] bpf, arm64: Dispatch exception cleanup pads at run time
2026-09-18 4:43 ` [PATCH bpf-next v2 12/20] bpf, arm64: " Yonghong Song
@ 2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 20:07 ` Yonghong Song
0 siblings, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-18 5:44 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 2508 bytes --]
> diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
> index 6c04fee468766..560eba305bca2 100644
> --- a/arch/arm64/net/bpf_jit_comp.c
> +++ b/arch/arm64/net/bpf_jit_comp.c
[ ... ]
> @@ -75,7 +76,21 @@ static const int bpf2a64[] = {
> [ARENA_VM_START] = A64_R(28),
> };
>
> +/* Throw-site spill: the five pairs push_callee_regs() forces on, same size and
> + * slot order, so arch_bpf_run_cleanup_pad() reads both alike.
> + */
> +#define A64_CLEANUP_SPILL_SZ (5 * 16)
> +
> +/*
> + * Where a landing pad's frame is anchored, since the stack pointer generated
> + * code normally addresses it through is the walker's inside a pad. bpf2a64[]
> + * maps nothing to x24, so nothing else in generated code touches it.
> + */
> +#define A64_CLEANUP_FP A64_R(24)
> +
> struct jit_ctx {
> + /* Bytes reserved for the throw-site spill; see bpf_cleanup_force_spill(). */
> + u32 throw_spill;
> const struct bpf_prog *prog;
> int idx;
> int epilogue_offset;
This isn't a bug, but would it read better to put throw_spill next to
stack_size and stack_arg_size, which it is always used with, rather
than ahead of prog at the head of the struct?
[ ... ]
> @@ -1367,6 +1411,7 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
> const s16 off = insn->off;
> const s32 imm = insn->imm;
> const int i = insn - ctx->prog->insnsi;
> + const bool in_pad = bpf_cleanup_insn_in_pad(ctx->prog, i);
> const bool is64 = BPF_CLASS(code) == BPF_ALU64 ||
> BPF_CLASS(code) == BPF_JMP;
> u8 jmp_cond;
> @@ -1378,9 +1423,14 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
> int ret;
> bool sign_extend;
>
> - if (bpf_insn_is_indirect_target(env, ctx->prog, i))
> + if (bpf_insn_is_indirect_target(env, ctx->prog, i) ||
> + bpf_cleanup_insn_is_pad(ctx->prog, i))
> emit_bti(A64_BTI_J, ctx);
>
> + if (bpf_cleanup_insn_is_pad(ctx->prog, i))
> + emit(A64_SUB_I(1, A64_CLEANUP_FP, fp,
> + ctx->stack_size + ctx->stack_arg_size), ctx);
> +
This isn't a bug, but could the two bpf_cleanup_insn_is_pad() calls
share a local the way in_pad does, and could the throw-site check
further down use the existing 'i' instead of recomputing
insn - ctx->prog->insnsi?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 10/20] bpf: Dispatch exception cleanup pads from bpf_throw()
2026-09-18 4:42 ` [PATCH bpf-next v2 10/20] bpf: Dispatch exception cleanup pads from bpf_throw() Yonghong Song
@ 2026-09-18 5:58 ` bot+bpf-ci
2026-09-19 19:54 ` Yonghong Song
0 siblings, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-18 5:58 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3525 bytes --]
> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
> index b2bf831242b85..521086d084a35 100644
> --- a/kernel/bpf/exception.c
> +++ b/kernel/bpf/exception.c
[ ... ]
> +static int cmp_u32(const void *a, const void *b)
> +{
> + u32 x = *(const u32 *)a, y = *(const u32 *)b;
> +
> + return x < y ? -1 : x > y;
> +}
[ ... ]
> +/* Is @idx in the sorted array @at of @n instruction indices? */
> +static bool insn_idx_in(const u32 *at, u32 n, u32 idx)
> +{
> + u32 l = 0, r = n;
> +
> + while (l < r) {
> + u32 m = l + (r - l) / 2;
> +
> + if (idx < at[m])
> + r = m;
> + else if (idx > at[m])
> + l = m + 1;
> + else
> + return true;
> + }
> + return false;
> +}
This isn't a bug, but could insn_idx_in() just use bsearch() from
<linux/bsearch.h>, reusing the cmp_u32() already defined above for the
sort()? kernel/bpf/fixups.c in this series uses bsearch() for similar
lookups.
> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
> index f120ae66dbbcb..134aafa6a6c9b 100644
> --- a/kernel/bpf/fixups.c
> +++ b/kernel/bpf/fixups.c
[ ... ]
> @@ -1113,6 +1120,116 @@ static void bpf_restore_subprog_starts(struct bpf_verifier_env *env, u32 *orig_s
> env->subprog_info[env->subprog_cnt].start = env->prog->len;
> }
>
> +static int cleanup_info_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
> + u32 start, u32 end)
> +{
> + struct bpf_cleanup_info *recs;
> + u32 i, cnt = 0;
> + int err;
> +
> + if (!env->cleanup_info_cnt)
> + return 0;
> +
> + err = bpf_cleanup_alloc_info(sub->aux);
> + if (err)
> + return err;
> +
> + err = cleanup_throw_sites_for_subprog(env, sub, start, end);
> + if (err)
> + return err;
> +
> + err = cleanup_pad_body_for_subprog(env, sub, start, end);
> + if (err)
> + return err;
> +
> + for (i = start; i < end; i++)
> + if (env->insn_aux_data[i].cleanup_pad)
> + cnt++;
> + if (!cnt)
> + return 0;
> +
> + recs = kvmalloc_array(cnt, sizeof(*recs), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
> + if (!recs)
> + return -ENOMEM;
> +
> + for (i = start, cnt = 0; i < end; i++) {
> + u32 pad = env->insn_aux_data[i].cleanup_pad;
> +
> + if (!pad)
> + continue;
> + pad--;
> + if (verifier_bug_if(pad < start || pad >= end, env,
> + "insn %u is covered by a landing pad at %u outside its subprog [%u, %u)",
> + i, pad, start, end)) {
> + kvfree(recs);
> + return -EFAULT;
> + }
> + recs[cnt].begin_off = i - start;
> + recs[cnt].end_off = i - start + 1;
> + recs[cnt].landing_pad_off = pad - start;
> + cnt++;
> + }
> + return bpf_cleanup_attach_info(sub->aux, recs, cnt);
> +}
[ ... ]
> static int jit_subprogs(struct bpf_verifier_env *env)
> {
> struct bpf_prog *prog = env->prog, **func, *tmp;
[ ... ]
> @@ -1250,6 +1367,10 @@ static int jit_subprogs(struct bpf_verifier_env *env)
> func[i]->aux->token = prog->aux->token;
> if (!i)
> func[i]->aux->exception_boundary = env->seen_exception;
> + err = cleanup_info_for_subprog(env, func[i], subprog_start,
> + env->subprog_info[i + 1].start);
This isn't a bug, but would passing subprog_end (already computed a few
lines above from env->subprog_info[i + 1].start) be clearer here? It
keeps the cleanup table's end bound explicitly tied to the same value
jit_subprogs() uses for func[i]->len.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test
2026-09-18 4:43 ` [PATCH bpf-next v2 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
2026-09-18 4:59 ` sashiko-bot
@ 2026-09-18 5:58 ` bot+bpf-ci
2026-09-19 20:34 ` Yonghong Song
1 sibling, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-18 5:58 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 2765 bytes --]
> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
> new file mode 100644
> index 0000000000000..ce2dac306a84c
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
> @@ -0,0 +1,600 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
> +#include <vmlinux.h>
> +#include <bpf/bpf_helpers.h>
> +#include "bpf_experimental.h"
> +#include "bpf_misc.h"
> +#include "../test_kmods/bpf_testmod_kfunc.h"
> +#include "exceptions_cleanup.h"
> +
> +__u64 input = 0;
> +
> +static __used __noinline void __kfunc_btf_anchor(void)
> +{
> + bpf_throw(0);
> + bpf_preempt_disable();
> + bpf_preempt_enable();
> + bpf_unwind_resume();
> +}
Is bpf_unwind_resume() declared anywhere? It is called from C here but
bpf_experimental.h declares only the other kfuncs this anchor exists for:
tools/testing/selftests/bpf/bpf_experimental.h declares bpf_throw():
extern void bpf_throw(u64 cookie) __ksym;
and the preempt pair:
void bpf_preempt_disable(void) __weak __ksym;
void bpf_preempt_enable(void) __weak __ksym;
but bpf_unwind_resume() does not appear in any header. vmlinux.h contains
type definitions only, not kfunc prototypes.
BPF objects are built with BPF_CFLAGS = -g -Wall -Werror -std=gnu11
(tools/testing/selftests/bpf/Makefile). -Wimplicit-function-declaration
is part of -Wall, and in C99+ mode modern Clang promotes it to an error,
which with -Werror becomes fatal:
progs/exceptions_cleanup_fail.c:17:2: error: call to undeclared
function 'bpf_unwind_resume'; ISO C99 and later do not support
implicit function declarations
This would break the build of the whole BPF selftests tree. The same C-level
call appears in the three other new prog files added by this series
(progs/exceptions_cleanup.c, and in follow-ups progs/exceptions_cleanup_shapes.c
and progs/exceptions_cleanup_light.c), so all of them would fail to compile.
The anchor cannot simply drop the call either: the C call is what makes Clang
emit the BTF extern FUNC for bpf_unwind_resume, which libbpf needs to resolve
the inline-asm "call bpf_unwind_resume;" relocations. The fix would be to add
the declaration:
extern void bpf_unwind_resume(void) __ksym;
in bpf_experimental.h alongside bpf_throw(). For comparison,
progs/exceptions_cleanup_shapes.c (commit 198e7e8d6d72) does declare its
sibling explicitly:
extern void _Unwind_Resume(void) __ksym;
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 03/20] bpf: Add lookups for exception cleanup resumes and landing pads
2026-09-18 5:44 ` bot+bpf-ci
@ 2026-09-19 4:55 ` Alexei Starovoitov
2026-09-19 17:42 ` Yonghong Song
0 siblings, 1 reply; 56+ messages in thread
From: Alexei Starovoitov @ 2026-09-19 4:55 UTC (permalink / raw)
To: bot+bpf-ci, yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, mason, ihor.solodrai
On Fri Sep 18, 2026 at 5:44 AM UTC, bot+bpf-ci wrote:
> > diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> > index c08505b9ba821..f9bccd3e0f4db 100644
> > --- a/include/linux/bpf_verifier.h
> > +++ b/include/linux/bpf_verifier.h
> > @@ -681,6 +681,11 @@ struct bpf_insn_aux_data {
> > bool needs_zext; /* alu op needs to clear upper bits */
> > bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
> > bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
> > + /*
> > + * 1 + the instruction index of the exception cleanup landing pad this
> > + * call site unwinds to, or 0 for none.
> > + */
> > + u32 cleanup_pad;
> > bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
> > u8 alu_state; /* used in combination with alu_limit */
>
> [ ... ]
>
> > diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
> > index 2add8001c3ec3..82b00fac6bd6a 100644
> > --- a/kernel/bpf/fixups.c
> > +++ b/kernel/bpf/fixups.c
> > @@ -261,6 +261,11 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
> > }
> > }
> >
> > + if (env->cleanup_info_cnt)
> > + for (i = 0; i < prog_len; i++)
> > + if (data[i].cleanup_pad > off + 1)
> > + data[i].cleanup_pad += cnt - 1;
> > +
>
> [ ... ]
>
> > @@ -549,6 +554,7 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
> > struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
> > unsigned int orig_prog_len = env->prog->len;
> > int err;
> > + u32 i;
> >
> > if (bpf_prog_is_offloaded(env->prog->aux))
> > bpf_prog_offload_remove_insns(env, off, cnt);
> > @@ -573,6 +579,17 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
> > sizeof(*aux_data) * (orig_prog_len - off - cnt));
> > env->insn_aux_data_len -= cnt;
> >
> > + if (env->cleanup_info_cnt) {
> > + for (i = 0; i < env->insn_aux_data_len; i++) {
> > + u32 pad = aux_data[i].cleanup_pad;
> > +
> > + if (pad > off + cnt)
> > + aux_data[i].cleanup_pad = pad - cnt;
> > + else if (pad > off)
> > + aux_data[i].cleanup_pad = 0;
> > + }
> > + }
> > +
> > return 0;
> > }
>
> Can this logic handle NOP removal correctly? When bpf_opt_remove_nops()
> calls verifier_remove_insns(env, i, 1) to remove a single fall-through
> goto +0, if a cleanup record names that instruction as its landing pad,
> then pad == off + 1, so pad > off is true and cleanup_pad becomes 0.
>
> But removing the NOP leaves the pad's next instruction at the same
> index, so the correct landing pad index stays unchanged rather than
> being zeroed. For example:
>
> off: some_call with cleanup_pad = off+1
> off+1: goto +0 (NOP landing pad)
> off+2: actual cleanup code
>
> After removing the NOP at off+1:
>
> off: some_call (cleanup_pad should still be off+1)
> off+1: actual cleanup code (was at off+2)
>
> The `else if (pad > off)` arm erases the record instead of preserving
> it. Nothing rejects NOP landing pads during validation - a goto +0 has
> code 0x05, not 0x00, so check_cleanup_info() accepts it.
>
> The ordering makes the loss escape validation: bpf_check_cleanup_exceptions()
> runs before bpf_opt_remove_nops(), so the verifier proves the cleanup
> path releases acquired references, but then the NOP removal zeroes
> cleanup_pad, and later cleanup_info_for_subprog() sees `if (!pad)
> continue;` and omits the record from the runtime table. The program is
> accepted but bpf_throw() finds no pad for that call site and skips the
> cleanup, leaking whatever the pad was supposed to release.
>
> Does the `else if (pad > off)` arm need to distinguish removals that
> consume the landing pad instruction itself versus removals that happen
> before it?
bot is correct here.
pls fix
pw-bot: cr
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 03/20] bpf: Add lookups for exception cleanup resumes and landing pads
2026-09-19 4:55 ` Alexei Starovoitov
@ 2026-09-19 17:42 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 17:42 UTC (permalink / raw)
To: Alexei Starovoitov, bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/18/26 9:55 PM, Alexei Starovoitov wrote:
> On Fri Sep 18, 2026 at 5:44 AM UTC, bot+bpf-ci wrote:
>>> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
>>> index c08505b9ba821..f9bccd3e0f4db 100644
>>> --- a/include/linux/bpf_verifier.h
>>> +++ b/include/linux/bpf_verifier.h
>>> @@ -681,6 +681,11 @@ struct bpf_insn_aux_data {
>>> bool needs_zext; /* alu op needs to clear upper bits */
>>> bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
>>> bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
>>> + /*
>>> + * 1 + the instruction index of the exception cleanup landing pad this
>>> + * call site unwinds to, or 0 for none.
>>> + */
>>> + u32 cleanup_pad;
>>> bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
>>> u8 alu_state; /* used in combination with alu_limit */
>> [ ... ]
>>
>>> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
>>> index 2add8001c3ec3..82b00fac6bd6a 100644
>>> --- a/kernel/bpf/fixups.c
>>> +++ b/kernel/bpf/fixups.c
>>> @@ -261,6 +261,11 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
>>> }
>>> }
>>>
>>> + if (env->cleanup_info_cnt)
>>> + for (i = 0; i < prog_len; i++)
>>> + if (data[i].cleanup_pad > off + 1)
>>> + data[i].cleanup_pad += cnt - 1;
>>> +
>> [ ... ]
>>
>>> @@ -549,6 +554,7 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
>>> struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
>>> unsigned int orig_prog_len = env->prog->len;
>>> int err;
>>> + u32 i;
>>>
>>> if (bpf_prog_is_offloaded(env->prog->aux))
>>> bpf_prog_offload_remove_insns(env, off, cnt);
>>> @@ -573,6 +579,17 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
>>> sizeof(*aux_data) * (orig_prog_len - off - cnt));
>>> env->insn_aux_data_len -= cnt;
>>>
>>> + if (env->cleanup_info_cnt) {
>>> + for (i = 0; i < env->insn_aux_data_len; i++) {
>>> + u32 pad = aux_data[i].cleanup_pad;
>>> +
>>> + if (pad > off + cnt)
>>> + aux_data[i].cleanup_pad = pad - cnt;
>>> + else if (pad > off)
>>> + aux_data[i].cleanup_pad = 0;
>>> + }
>>> + }
>>> +
>>> return 0;
>>> }
>> Can this logic handle NOP removal correctly? When bpf_opt_remove_nops()
>> calls verifier_remove_insns(env, i, 1) to remove a single fall-through
>> goto +0, if a cleanup record names that instruction as its landing pad,
>> then pad == off + 1, so pad > off is true and cleanup_pad becomes 0.
>>
>> But removing the NOP leaves the pad's next instruction at the same
>> index, so the correct landing pad index stays unchanged rather than
>> being zeroed. For example:
>>
>> off: some_call with cleanup_pad = off+1
>> off+1: goto +0 (NOP landing pad)
>> off+2: actual cleanup code
>>
>> After removing the NOP at off+1:
>>
>> off: some_call (cleanup_pad should still be off+1)
>> off+1: actual cleanup code (was at off+2)
>>
>> The `else if (pad > off)` arm erases the record instead of preserving
>> it. Nothing rejects NOP landing pads during validation - a goto +0 has
>> code 0x05, not 0x00, so check_cleanup_info() accepts it.
>>
>> The ordering makes the loss escape validation: bpf_check_cleanup_exceptions()
>> runs before bpf_opt_remove_nops(), so the verifier proves the cleanup
>> path releases acquired references, but then the NOP removal zeroes
>> cleanup_pad, and later cleanup_info_for_subprog() sees `if (!pad)
>> continue;` and omits the record from the runtime table. The program is
>> accepted but bpf_throw() finds no pad for that call site and skips the
>> cleanup, leaking whatever the pad was supposed to release.
>>
>> Does the `else if (pad > off)` arm need to distinguish removals that
>> consume the landing pad instruction itself versus removals that happen
>> before it?
> bot is correct here.
> pls fix
Agree with bot about for two different cases. Will fix.
>
> pw-bot: cr
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 05/20] bpf: Make exception landing pads reachable in the CFG
2026-09-18 4:59 ` sashiko-bot
@ 2026-09-19 19:17 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 19:17 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/17/26 9:59 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf: missing prune point at exception landing pad target
> --
>
> commit 4fbd0e36862ba7e3ee8480e535068a6821d5cad0
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Make exception landing pads reachable in the CFG
>
> This commit updates the verifier's control flow graph and liveness analysis
> to properly handle exception landing pads. It ensures that exception throwing
> calls can reach cleanup records, and prevents outer frame stack slots from
> being read incorrectly when control is transferred to landing pads.
>
>> diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
>> index 842c7d1eabccc..9f8b8b54d5ea7 100644
>> --- a/kernel/bpf/cfg.c
>> +++ b/kernel/bpf/cfg.c
> [ ... ]
>
>> @@ -158,17 +159,57 @@ static int push_insn(int t, int w, int e, struct bpf_verifier_env *env)
>> return DONE_EXPLORING;
>> }
>>
>> +static int visit_cleanup_pad_edge(int t, struct bpf_verifier_env *env)
>> +{
>> + int *insn_stack = env->cfg.insn_stack;
>> + int *insn_state = env->cfg.insn_state;
>> + int w;
>> +
>> + if (!env->cleanup_info_cnt)
>> + return DONE_EXPLORING;
>> + w = bpf_cleanup_pad_of_call(env, t);
>> + if (w < 0)
>> + return DONE_EXPLORING;
>> +
>> + mark_prune_point(env, t);
>> + mark_jmp_point(env, w);
>> + mark_jump_target(env, w);
> [Severity: High]
> Should mark_prune_point() be applied to the landing pad target w instead of
> the call site t?
>
> Because w is a branch target where control flow merges, it needs to be marked
> as a prune point so the verifier knows to check for state convergence. By
> marking t instead of w, the verifier will evaluate the landing pad and all
> subsequent cleanup instructions independently for every exception-throwing
> call that branches to w.
>
> Can this lead to combinatorial path explosion and result in the verifier
> rejecting the program with -E2BIG?
We can add mark_prune_point(env, w) as well, esp. there are cases there
are multiple cleanup regions but many (or all) of them may pointing into
the same landing pad.
>
> [ ... ]
>
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 05/20] bpf: Make exception landing pads reachable in the CFG
2026-09-18 5:44 ` bot+bpf-ci
@ 2026-09-19 19:32 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 19:32 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/17/26 10:44 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
>> index 44ecdc5b4ec2d..9cfd05f970bc9 100644
>> --- a/kernel/bpf/liveness.c
>> +++ b/kernel/bpf/liveness.c
>> @@ -8,6 +8,8 @@
>> #include <linux/slab.h>
>> #include <linux/sort.h>
>>
>> +#include "exception.h"
>> +
>> #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
>>
>> struct per_frame_masks {
>> @@ -256,6 +258,9 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
>> succ = env->succ;
>> succ->cnt = 0;
>>
>> + if (unlikely(bpf_is_unwind_resume_kfunc(insn)))
>> + return succ;
>> +
> Should this early return be gated on env->cleanup_info_cnt?
>
> bpf_unwind_resume() is an unconditionally registered generic kfunc with no
> KF_ flag restricting it to programs that carry a cleanup table. Its body
> is empty, so at run time the call simply returns and execution continues
> at the next instruction.
This should be okay without env->cleanup_info_cnt guard. Eventually if
bpf_unwind_resume() is not in cleanup region, a verifier error message
will be in the log.
>
> visit_insn() in cfg.c still gives the call a fallthrough edge, and at this
> commit nothing in do_check() rejects a bpf_unwind_resume() call that is not
> the tail of a cleanup landing pad. So any CAP_BPF program with no cleanup
> table can call bpf_unwind_resume() and keep executing after it.
>
> For such a program, bpf_insn_successors() reports zero successors, so
> bpf_compute_live_registers() treats everything after the call as dead. The
> register-cleaning side fails closed (a poisoned slot is rejected), but the
> 32-bit zero-extension side fails open:
>
> insn_aux[i].zext_dst for an insn preceding the call is only set when the
> hi half of the defined register is live-after, so it stays false, and
> bpf_opt_subreg_zext_lo32_rnd_hi32() (kernel/bpf/fixups.c:706) then does
> 'if (!aux[adj_idx].zext_dst) { if (!rnd_hi32) continue; ... }' and emits
> no BPF_ZEXT_REG.
>
> On architectures where bpf_jit_needs_zext() is true and kfunc calls are
> supported (s390, riscv64, powerpc64, parisc, x86-32), the upper 32 bits of
> a register the verifier proved to be 32-bit bounded are left with stale
> JIT-defined contents, which defeats the range checks the verifier derived
> for it.
>
> Compare with the pad lookup ten lines below:
>
>> opcode_info = &opcode_info_tbl[BPF_CLASS(insn->code) | BPF_OP(insn->code)];
>> insn_sz = bpf_is_ldimm64(insn) ? 2 : 1;
>> if (opcode_info->can_fallthrough)
>> @@ -264,6 +269,13 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
>> if (opcode_info->can_jump)
>> succ->items[succ->cnt++] = idx + bpf_jmp_offset(insn) + 1;
>>
>> + if (unlikely(env->cleanup_info_cnt)) {
>> + int pad = bpf_cleanup_pad_of_call(env, idx);
>> +
>> + if (pad >= 0)
>> + succ->items[succ->cnt++] = pad;
>> + }
> which gates the pad lookup on cleanup_info_cnt, and with
> visit_cleanup_pad_edge() in kernel/bpf/cfg.c which returns DONE_EXPLORING
> when !env->cleanup_info_cnt.
>
> Would gating the early return on env->cleanup_info_cnt fix this? That
> would be sound because a program that carries a cleanup table is
> additionally constrained (bpf_jit_supports_cleanup_pads() plus, from commit
> fa6890df61539 onward, the requirement that a resume sit at the end of a
> landing pad), whereas a program with cleanup_info_cnt == 0 must keep the
> ordinary fallthrough successor of the call.
As I mentioned, previous gating with "env->cleanup_info_cnt" is not
necessary.
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 06/20] bpf: Explore the landing pads no call site reaches
2026-09-18 5:44 ` bot+bpf-ci
@ 2026-09-19 19:32 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 19:32 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/17/26 10:44 PM, bot+bpf-ci@kernel.org wrote:
>> commit c93117f24a1a9cc39bfeaef79c42dbff02d8c455
>> Author: Yonghong Song <yonghong.song@linux.dev>
>>
>> bpf: Explore the landing pads no call site reaches
>>
>> A cleanup record need not cover a call an exception can unwind out of: a
>> frontend is free to emit a region around a helper or an ordinary kfunc,
>> both nounwind here. Nothing marks a call site then, and that record's
>> landing pad is reached by nothing at all -- leaving bpf_check_cfg() to
>> refuse the program over code its own frontend had no way not to emit:
>>
>> 0: call bpf_preempt_disable
>> 1: call bpf_preempt_enable record = { begin = 1, end = 2, pad = 4 }
>> 2: r0 = 0
>> 3: exit
>> 4: r1 = pads_ran ll landing pad
>> 6: r2 = *(u64 *)(r1 + 0)
>> 7: r2 |= RAN_NOUNWIND_REC
>> 8: *(u64 *)(r1 + 0) = r2
>> 9: call bpf_unwind_resume
>> 10: exit
>>
>> The range [1,2) holds one call, and it is a kfunc, so an exception cannot
>> come out of it. cleanup_mark_call_sites() marks nothing, nothing pushes an
>> edge to 4, and 4 through 10 are reachable from nothing: "unreachable insn
>> 4". The pad is dead, which is correct -- no exception can ever arrive at it
>> -- but the program is fine and has to load.
>>
>> Walk every pad the table names that the edges did not reach, the way the
>> walk is already re-seeded at an exception callback. From there the pad is
>> code like any other: do_check() never enters it, because no call site
>> dispatches to it, so the dead code sweep removes it along with everything
>> else that was not reached.
>>
>> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
> This looks like a bug fix for the earlier commit that made exception landing
> pads reachable in the CFG. Should this include:
>
> Fixes: a67011bbbca1ba95 ("bpf: Make exception landing pads reachable in the CFG")
No need for this Fixes tag.
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 09/20] bpf: Refuse a private stack for a program with an exception cleanup table
2026-09-18 5:44 ` bot+bpf-ci
@ 2026-09-19 19:37 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 19:37 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/17/26 10:44 PM, bot+bpf-ci@kernel.org wrote:
>> A landing pad is about to start running outside its own frame, on the
>> bpf_throw() walker's stack, with that frame's registers put back from a
>> spill area whose contents arch_bpf_run_cleanup_pad() knows how to read. The
>> x86-64 JIT addresses a private-stack program's frame through a scratch
>> register that it recomputes after each call rather than through rbp, and no
>> spill area holds that register, so a pad there would address its frame
>> through whatever the kernel left behind.
> The description says the x86-64 JIT 'recomputes after each call'. Does the
> JIT actually recompute it?
>
> Looking at the x86 JIT, emit_priv_frame_ptr() (arch/x86/net/bpf_jit_comp.c)
> is emitted once in the prologue, and the value in X86_REG_R9 is carried
> across calls by push_r9()/pop_r9() around each call site; BPF_REG_FP
> accesses are rewritten to X86_REG_R9.
>
> The load-bearing half of the argument is still correct -
> arch_bpf_run_cleanup_pad() (added later in the series,
> arch/x86/net/bpf_cleanup_pad.S) restores only r15/r14/r13/rbx/r12 and rbp
> from the spill area and never restores r9, so a pad in a private-stack
> program would address its frame through a stale register.
>
> Only the described mechanism (recompute vs save/restore around calls) is
> wrong, which makes the justification harder to check against the JIT.
>
>> Everywhere rather than arch-conditionally: arm64 keeps its private
>> stack pointer in x27, which is in the prologue spill and so survives, but a
>> rule that holds on one arch and not the other is not worth the second code
>> path when nothing is lost but an optimization.
> Does 'nothing is lost but an optimization' understate the effect?
>
> The private-stack mode also changes the verifier's stack-depth accounting,
> not just JIT codegen. In check_max_stack_depth_subprog() a
> PRIV_STACK_ADAPTIVE subprog is checked against MAX_BPF_STACK on its own ('if
> (subprog_depth > MAX_BPF_STACK)'), while a NO_PRIV_STACK subprog accumulates
> into 'depth' and the whole bpf2bpf chain must fit in MAX_BPF_STACK
> ('combined stack size of %d calls is %d. Too large').
>
> So for the prog types bpf_enable_priv_stack() would have returned
> PRIV_STACK_ADAPTIVE for (kprobe/tracepoint/perf_event/raw_tp, and
> tracing/lsm/struct_ops with a recursion or an explicit request), a program
> carrying a cleanup table whose nested subprog depths sum above 512 bytes now
> fails to load rather than losing performance.
>
> Worth stating in the changelog, since it is a load-time visible restriction
> on the new feature rather than an optimization.
Okay, comments and commit message may have some stale things or something
like (when nothing is lost but an optimization) is not correct. Will fix.
>
> Related: kernel/sched/ext/ext.c sets prog->aux->priv_stack_requested for
> sched_ext dispatch/sub_caps_updated, and verifier.c rejects such a request
> when the JIT cannot honour it; here the request is silently ignored instead.
>
> This looks like a bug fix for private stack support. Should this include:
>
> Fixes: a76ab5731e32d ("bpf: Find eligible subprogs for private stack support")
Fixes tag is unnecessary.
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 10/20] bpf: Dispatch exception cleanup pads from bpf_throw()
2026-09-18 5:58 ` bot+bpf-ci
@ 2026-09-19 19:54 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 19:54 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/17/26 10:58 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
>> index b2bf831242b85..521086d084a35 100644
>> --- a/kernel/bpf/exception.c
>> +++ b/kernel/bpf/exception.c
> [ ... ]
>
>> +static int cmp_u32(const void *a, const void *b)
>> +{
>> + u32 x = *(const u32 *)a, y = *(const u32 *)b;
>> +
>> + return x < y ? -1 : x > y;
>> +}
> [ ... ]
>
>> +/* Is @idx in the sorted array @at of @n instruction indices? */
>> +static bool insn_idx_in(const u32 *at, u32 n, u32 idx)
>> +{
>> + u32 l = 0, r = n;
>> +
>> + while (l < r) {
>> + u32 m = l + (r - l) / 2;
>> +
>> + if (idx < at[m])
>> + r = m;
>> + else if (idx > at[m])
>> + l = m + 1;
>> + else
>> + return true;
>> + }
>> + return false;
>> +}
> This isn't a bug, but could insn_idx_in() just use bsearch() from
> <linux/bsearch.h>, reusing the cmp_u32() already defined above for the
> sort()? kernel/bpf/fixups.c in this series uses bsearch() for similar
> lookups.
Good point. will do.
>
>> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
>> index f120ae66dbbcb..134aafa6a6c9b 100644
>> --- a/kernel/bpf/fixups.c
>> +++ b/kernel/bpf/fixups.c
> [ ... ]
>
>> @@ -1113,6 +1120,116 @@ static void bpf_restore_subprog_starts(struct bpf_verifier_env *env, u32 *orig_s
>> env->subprog_info[env->subprog_cnt].start = env->prog->len;
>> }
>>
>> +static int cleanup_info_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
>> + u32 start, u32 end)
>> +{
>> + struct bpf_cleanup_info *recs;
>> + u32 i, cnt = 0;
>> + int err;
>> +
>> + if (!env->cleanup_info_cnt)
>> + return 0;
>> +
>> + err = bpf_cleanup_alloc_info(sub->aux);
>> + if (err)
>> + return err;
>> +
>> + err = cleanup_throw_sites_for_subprog(env, sub, start, end);
>> + if (err)
>> + return err;
>> +
>> + err = cleanup_pad_body_for_subprog(env, sub, start, end);
>> + if (err)
>> + return err;
>> +
>> + for (i = start; i < end; i++)
>> + if (env->insn_aux_data[i].cleanup_pad)
>> + cnt++;
>> + if (!cnt)
>> + return 0;
>> +
>> + recs = kvmalloc_array(cnt, sizeof(*recs), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
>> + if (!recs)
>> + return -ENOMEM;
>> +
>> + for (i = start, cnt = 0; i < end; i++) {
>> + u32 pad = env->insn_aux_data[i].cleanup_pad;
>> +
>> + if (!pad)
>> + continue;
>> + pad--;
>> + if (verifier_bug_if(pad < start || pad >= end, env,
>> + "insn %u is covered by a landing pad at %u outside its subprog [%u, %u)",
>> + i, pad, start, end)) {
>> + kvfree(recs);
>> + return -EFAULT;
>> + }
>> + recs[cnt].begin_off = i - start;
>> + recs[cnt].end_off = i - start + 1;
>> + recs[cnt].landing_pad_off = pad - start;
>> + cnt++;
>> + }
>> + return bpf_cleanup_attach_info(sub->aux, recs, cnt);
>> +}
> [ ... ]
>
>> static int jit_subprogs(struct bpf_verifier_env *env)
>> {
>> struct bpf_prog *prog = env->prog, **func, *tmp;
> [ ... ]
>
>> @@ -1250,6 +1367,10 @@ static int jit_subprogs(struct bpf_verifier_env *env)
>> func[i]->aux->token = prog->aux->token;
>> if (!i)
>> func[i]->aux->exception_boundary = env->seen_exception;
>> + err = cleanup_info_for_subprog(env, func[i], subprog_start,
>> + env->subprog_info[i + 1].start);
> This isn't a bug, but would passing subprog_end (already computed a few
> lines above from env->subprog_info[i + 1].start) be clearer here? It
> keeps the cleanup table's end bound explicitly tied to the same value
> jit_subprogs() uses for func[i]->len.
Good point. subprog_end is a good choice.
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 11/20] bpf, x86: Dispatch exception cleanup pads at run time
2026-09-18 5:03 ` sashiko-bot
@ 2026-09-19 20:00 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 20:00 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/17/26 10:03 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf, x86: garbage kfunc stack arguments in exception cleanup pads
> --
>
> commit 1e32297a3c2c4f6bbc4d458bc25563bed8167f6c
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf, x86: Dispatch exception cleanup pads at run time
>
> This commit provides the architecture-specific support for x86-64 exception
> cleanup pads. It forces a full callee-saved spill for programs with a cleanup
> table and introduces arch_bpf_run_cleanup_pad to execute the pads.
>
>> diff --git a/arch/x86/net/bpf_cleanup_pad.S b/arch/x86/net/bpf_cleanup_pad.S
>> new file mode 100644
>> index 0000000000000..da4b448ecf095
>> --- /dev/null
>> +++ b/arch/x86/net/bpf_cleanup_pad.S
> [ ... ]
>
>> +SYM_FUNC_START(arch_bpf_run_cleanup_pad)
>> + ANNOTATE_NOENDBR
>> +
>> + pushq %rbp
>> + pushq %rbx
>> + pushq %r12
>> + pushq %r13
>> + pushq %r14
>> + pushq %r15
>> + /* Keep the pad's entry rsp congruent to a normal call's. */
>> + subq $8, %rsp
>> +
>> + movq 0(%rdx), %r15
>> + movq 8(%rdx), %r14
>> + movq 16(%rdx), %r13
>> + movq 24(%rdx), %rbx
>> + movq 32(%rdx), %r12
>> + /* rbp is BPF r10, so this is the whole of the pad's frame setup. */
>> + movq %rsi, %rbp
>> +
>> + /* CALL_NOSPEC needs the target in a register; rcx is BPF r4, dead. */
>> + movq %rdi, %rcx
>> +
>> + /* BPF r0 on the way into a pad, not whatever the kernel left in rax. */
>> + movl $BPF_PAD_ENTRY_R0, %eax
>> +
>> + CALL_NOSPEC rcx
> [Severity: High]
> Does this execution environment cause kfuncs with more than 5 arguments
> to read uninitialized garbage stack arguments?
We should be okay since we already have:
+ if (bpf_pseudo_kfunc_call(insn)) {
+ struct bpf_call_summary cs;
+
+ if (bpf_get_call_summary(env, insn, &cs) &&
+ cs.arg_slot_cnt > MAX_BPF_FUNC_REG_ARGS) {
+ verbose(env,
+ "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
+ i);
+ return -EINVAL;
+ }
+ }
>
> In arch_bpf_run_cleanup_pad, the pad is executed on the unwinder's
> stack (rsp), while rbp is set to the original BPF frame.
>
> However, the JIT unconditionally writes outgoing kfunc stack arguments
> relative to BPF_REG_FP (rbp), as seen in emit_kfunc_arg_moves in
> arch/x86/net/bpf_jit_comp.c:
>
> else if (dst_mem)
> emit_stx(&prog, BPF_DW, BPF_REG_FP, reg,
> stack_base + (dst - nreg) * 8);
>
> Because rbp and rsp are desynchronized in the pad, the arguments are
> written to the old frame. When the kfunc is called, the C calling
> convention requires it to read stack arguments from [rsp], which points
> to the unwinder's stack instead of the original frame, leading to
> garbage reads.
>
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 11/20] bpf, x86: Dispatch exception cleanup pads at run time
2026-09-18 5:44 ` bot+bpf-ci
@ 2026-09-19 20:04 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 20:04 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/17/26 10:44 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/arch/x86/net/Makefile b/arch/x86/net/Makefile
>> index dddbefc0f4398..9d574d972df37 100644
>> --- a/arch/x86/net/Makefile
>> +++ b/arch/x86/net/Makefile
>> @@ -6,5 +6,5 @@
>> ifeq ($(CONFIG_X86_32),y)
>> obj-$(CONFIG_BPF_JIT) += bpf_jit_comp32.o
>> else
>> - obj-$(CONFIG_BPF_JIT) += bpf_jit_comp.o bpf_timed_may_goto.o
>> + obj-$(CONFIG_BPF_JIT) += bpf_jit_comp.o bpf_timed_may_goto.o bpf_cleanup_pad.o
>> endif
>> diff --git a/arch/x86/net/bpf_cleanup_pad.S b/arch/x86/net/bpf_cleanup_pad.S
>> new file mode 100644
>> index 0000000000000..da4b448ecf095
>> --- /dev/null
>> +++ b/arch/x86/net/bpf_cleanup_pad.S
> [ ... ]
>
>> +SYM_FUNC_START(arch_bpf_run_cleanup_pad)
>> + ANNOTATE_NOENDBR
>> +
>> + pushq %rbp
>> + pushq %rbx
>> + pushq %r12
>> + pushq %r13
>> + pushq %r14
>> + pushq %r15
>> + /* Keep the pad's entry rsp congruent to a normal call's. */
>> + subq $8, %rsp
>> +
>> + movq 0(%rdx), %r15
>> + movq 8(%rdx), %r14
>> + movq 16(%rdx), %r13
>> + movq 24(%rdx), %rbx
>> + movq 32(%rdx), %r12
>> + /* rbp is BPF r10, so this is the whole of the pad's frame setup. */
>> + movq %rsi, %rbp
>> +
>> + /* CALL_NOSPEC needs the target in a register; rcx is BPF r4, dead. */
>> + movq %rdi, %rcx
>> +
>> + /* BPF r0 on the way into a pad, not whatever the kernel left in rax. */
>> + movl $BPF_PAD_ENTRY_R0, %eax
>> +
>> + CALL_NOSPEC rcx
>> +
>> + addq $8, %rsp
>> + popq %r15
>> + popq %r14
>> + popq %r13
>> + popq %r12
>> + popq %rbx
>> + popq %rbp
>> + RET
>> +SYM_FUNC_END(arch_bpf_run_cleanup_pad)
> Will this build with CONFIG_STACK_VALIDATION enabled?
>
> arch_bpf_run_cleanup_pad() is declared with SYM_FUNC_START (an STT_FUNC
> symbol) and saves rbp with 'pushq %rbp', but never establishes a frame
> pointer with 'mov %rsp, %rbp'. Instead, it overwrites rbp with the BPF
> frame pointer (movq %rsi, %rbp) and then makes a call (CALL_NOSPEC rcx).
>
> objtool's frame-pointer validation (CONFIG_STACK_VALIDATION with
> UNWINDER_FRAME_POINTER) checks for two conditions in validate_call()
> and validate_return():
>
> tools/objtool/check.c:has_valid_stack_frame() requires cfi->cfa.base ==
> CFI_BP, which is only set when the function does 'mov %rsp, %rbp' to
> establish a frame. This function never does that, so the CALL_NOSPEC
> would trigger:
>
> WARN_INSN(insn, "call without frame pointer save/setup");
>
> Additionally, the 'pushq %rbp' while cfa->base != CFI_BP sets
> cfi->bp_scratch = true, which is only cleared by 'mov %rsp, %rbp'.
> This flag is still set at the RET, where validate_return() warns:
>
> WARN_INSN(insn, "BP used as a scratch register");
>
> The Makefile adds bpf_cleanup_pad.o unconditionally under
> obj-$(CONFIG_BPF_JIT), even though the feature is gated on
> CONFIG_UNWINDER_ORC per bpf_jit_supports_cleanup_pads(). This means
> in a CONFIG_STACK_VALIDATION=y build (frame-pointer unwinder), the
> object would fail objtool validation, and with CONFIG_OBJTOOL_WERROR=y
> the build would fail.
>
> The neighbouring arch/x86/net/bpf_timed_may_goto.S does not have this
> issue because it sets up a proper frame:
>
> arch/x86/net/bpf_timed_may_goto.S:
> pushq %rbp
> movq %rsp, %rbp
> ...
> leave
>
> rbp cannot be the frame pointer here because it must carry BPF r10 into
> the pad. This is the same situation arch/x86/kernel/ftrace_64.S handles
> with STACK_FRAME_NON_STANDARD_FP(ftrace_caller). Could either
> STACK_FRAME_NON_STANDARD_FP(arch_bpf_run_cleanup_pad) after SYM_FUNC_END
> or building the object only when CONFIG_UNWINDER_ORC is set (matching the
> feature gate) address this?
Yes, jit only supports CONFIG_UNWINDER_ORC for x86, and CONFIG_STACK_VALIDATION
requires CONFIG_UNWINDER_FRAME_POINTER.
>
>> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
>> index d4a980140b48d..9d0dd54773e83 100644
>> --- a/arch/x86/net/bpf_jit_comp.c
>> +++ b/arch/x86/net/bpf_jit_comp.c
>> @@ -357,6 +357,11 @@ struct jit_context {
>> /* Number of bytes that will be skipped on tailcall */
>> #define X86_TAIL_CALL_OFFSET (12 + ENDBR_INSN_SIZE)
>>
>> +/* Throw-site spill: r15, r14, r13, rbx, r12 low to high, the layout the
>> + * prologue's pushes leave, so arch_bpf_run_cleanup_pad() reads both alike.
>> + */
>> +#define X86_CLEANUP_SPILL_SZ (5 * 8)
> This isn't a bug, but could the comment format be updated to match the
> BPF subsystem style? The BPF subsystem guide requires the modern kernel
> multi-line comment style, where the opening /* is on its own line. The
> patch already uses this format for its other new multi-line comments
> (the "Lowest address of each spill area" block and the "Spill r6-r9 and
> r12" block later in do_jit), so is there a reason for the inconsistency
> here?
Will adjust comment.
>
> [ ... ]
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 12/20] bpf, arm64: Dispatch exception cleanup pads at run time
2026-09-18 5:44 ` bot+bpf-ci
@ 2026-09-19 20:07 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 20:07 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/17/26 10:44 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
>> index 6c04fee468766..560eba305bca2 100644
>> --- a/arch/arm64/net/bpf_jit_comp.c
>> +++ b/arch/arm64/net/bpf_jit_comp.c
> [ ... ]
>
>> @@ -75,7 +76,21 @@ static const int bpf2a64[] = {
>> [ARENA_VM_START] = A64_R(28),
>> };
>>
>> +/* Throw-site spill: the five pairs push_callee_regs() forces on, same size and
>> + * slot order, so arch_bpf_run_cleanup_pad() reads both alike.
>> + */
>> +#define A64_CLEANUP_SPILL_SZ (5 * 16)
>> +
>> +/*
>> + * Where a landing pad's frame is anchored, since the stack pointer generated
>> + * code normally addresses it through is the walker's inside a pad. bpf2a64[]
>> + * maps nothing to x24, so nothing else in generated code touches it.
>> + */
>> +#define A64_CLEANUP_FP A64_R(24)
>> +
>> struct jit_ctx {
>> + /* Bytes reserved for the throw-site spill; see bpf_cleanup_force_spill(). */
>> + u32 throw_spill;
>> const struct bpf_prog *prog;
>> int idx;
>> int epilogue_offset;
> This isn't a bug, but would it read better to put throw_spill next to
> stack_size and stack_arg_size, which it is always used with, rather
> than ahead of prog at the head of the struct?
Okay, will do.
>
> [ ... ]
>
>> @@ -1367,6 +1411,7 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
>> const s16 off = insn->off;
>> const s32 imm = insn->imm;
>> const int i = insn - ctx->prog->insnsi;
>> + const bool in_pad = bpf_cleanup_insn_in_pad(ctx->prog, i);
>> const bool is64 = BPF_CLASS(code) == BPF_ALU64 ||
>> BPF_CLASS(code) == BPF_JMP;
>> u8 jmp_cond;
>> @@ -1378,9 +1423,14 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
>> int ret;
>> bool sign_extend;
>>
>> - if (bpf_insn_is_indirect_target(env, ctx->prog, i))
>> + if (bpf_insn_is_indirect_target(env, ctx->prog, i) ||
>> + bpf_cleanup_insn_is_pad(ctx->prog, i))
>> emit_bti(A64_BTI_J, ctx);
>>
>> + if (bpf_cleanup_insn_is_pad(ctx->prog, i))
>> + emit(A64_SUB_I(1, A64_CLEANUP_FP, fp,
>> + ctx->stack_size + ctx->stack_arg_size), ctx);
>> +
> This isn't a bug, but could the two bpf_cleanup_insn_is_pad() calls
> share a local the way in_pad does, and could the throw-site check
> further down use the existing 'i' instead of recomputing
> insn - ctx->prog->insnsi?
Will do.
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts
2026-09-18 4:57 ` sashiko-bot
@ 2026-09-19 20:18 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 20:18 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/17/26 9:57 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Medium] libbpf: missing E2BIG fallback for cleanup_info breaks forward compatibility
> --
>
> commit d67b69538e39a329708fcc9fa943f0ad05afaa05
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> libbpf: Add cleanup_info to bpf_prog_load_opts
>
> Let a caller hand the kernel an exception cleanup table. Extend struct
> bpf_prog_load_opts with cleanup_info, cleanup_info_cnt and
> cleanup_info_rec_size, pass them through to BPF_PROG_LOAD, and grow the
> attr size bpf_prog_load() computes to cover the new fields.
>
>> diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
>> index 96819c082c772..bcf490570961c 100644
>> --- a/tools/lib/bpf/bpf.c
>> +++ b/tools/lib/bpf/bpf.c
> [ ... ]
>> @@ -370,6 +370,10 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
>> attr.fd_array = ptr_to_u64(OPTS_GET(opts, fd_array, NULL));
>> attr.fd_array_cnt = OPTS_GET(opts, fd_array_cnt, 0);
>>
>> + attr.cleanup_info = ptr_to_u64(OPTS_GET(opts, cleanup_info, NULL));
>> + attr.cleanup_info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
>> + attr.cleanup_info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
> [Severity: Medium]
> Will this break forward compatibility for bpf_prog_load() if cleanup_info
> triggers an -E2BIG error?
>
> Later in bpf_prog_load(), there is a retry loop to handle -E2BIG when the
> userspace-provided record size is larger than what the kernel supports:
>
> while (errno == E2BIG && (!finfo || !linfo)) {
> if (!finfo && attr.func_info_cnt &&
> attr.func_info_rec_size < func_info_rec_size) {
> /* try with corrected func info records */
> ...
>
> The loop condition and body don't seem to include checks for cleanup_info.
>
> If the kernel returns -E2BIG because of a larger cleanup_info_rec_size,
> won't this hit the else { break; } block or skip the loop entirely, causing
> the load to fail instead of retrying with truncated records?
This is the first version to support .bpf_cleanup. It is too early to
talk about the future extension for .bpf_cleanup although we do have some
room to grow struct size if needed:
#define MAX_FUNCINFO_REC_SIZE 252
#define MAX_CLEANUP_INFO_REC_SIZE MAX_FUNCINFO_REC_SIZE
>
>> +
>> if (log_level) {
>> attr.log_buf = ptr_to_u64(log_buf);
>> attr.log_size = log_size;
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel
2026-09-18 5:00 ` sashiko-bot
@ 2026-09-19 20:21 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 20:21 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/17/26 10:00 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
> - [Medium] bpf: empty cleanup range check in bpf_prog_collect_cleanup_info() only catches offset 0
> - [Medium] bpf: bpf_object__init_cleanup_info() rejects valid objects with 0-sized .bpf_cleanup sections
> --
>
> commit a645d3bd249aceee4909edcc78e5fe4247fd3432
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> libbpf: Collect .bpf_cleanup records and pass them to the kernel
>
> Parse the compiler-emitted .bpf_cleanup section and hand the resulting
> table to BPF_PROG_LOAD.
>
>> diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
>> index 27f2eec13f76f..1b8b978e7b7b9 100644
>> --- a/tools/lib/bpf/libbpf.c
>> +++ b/tools/lib/bpf/libbpf.c
> [ ... ]
>
>> @@ -4852,6 +4890,221 @@ static struct bpf_program *find_prog_by_sec_insn(const struct bpf_object *obj,
>> return NULL;
>> }
>>
>> +static int bpf_object__init_cleanup_info(struct bpf_object *obj)
>> +{
>> + Elf_Data *data = obj->efile.cleanup_data;
>> + Elf_Data *relo = NULL;
>> + size_t i, nrels, nslots, nrecs;
>> + struct cleanup_raw_rec *recs;
>> + int *slot_sec, ret = 0;
>> + size_t *slot_val;
>> + const __u32 *vals;
>> + bool native;
>> +
>> + if (!data || obj->efile.cleanup_shndx < 0)
>> + return 0;
>> +
>> + native = is_native_endianness(obj);
>> +
>> + for (i = 0; i < obj->efile.sec_cnt; i++) {
>> + struct elf_sec_desc *sd = &obj->efile.secs[i];
>> +
>> + if (sd->sec_type == SEC_RELO && sd->shdr &&
>> + sd->shdr->sh_info == (Elf64_Word)obj->efile.cleanup_shndx) {
>> + relo = sd->data;
>> + break;
>> + }
>> + }
>> + if (!relo) {
>> + pr_warn("%s present without relocations\n", CLEANUP_SEC);
>> + return -LIBBPF_ERRNO__FORMAT;
>> + }
> [Severity: Medium]
> Could this reject valid objects if the compiler emits an empty 0-byte
> .bpf_cleanup section?
>
> If the section is empty (e.g., exceptions were optimized out), no relocations
> are emitted and there is no corresponding .rel.bpf_cleanup section. However,
> bpf_object__init_cleanup_info() requires the relocation section to exist if
> .bpf_cleanup is present.
>
> Would it be better to return early if data->d_size is 0 before checking for
> relocations?
I think it is okay, as you mentioned, there will be one error message later
when relocation section checks.
>
> [ ... ]
>
>> +static int bpf_prog_collect_cleanup_info(struct bpf_object *obj,
>> + struct bpf_program *prog)
>> +{
>> + size_t i;
>> + int j;
>> +
>> + for (i = 0; i < obj->cleanup_rec_cnt; i++) {
>> + struct cleanup_raw_rec *raw = &obj->cleanup_recs[i];
>> + struct bpf_program *owner = NULL;
>> + struct bpf_cleanup_info ci = {};
>> + __u32 fields[CLEANUP_REC_FIELDS];
>> + void *tmp;
>> +
>> + for (j = 0; j < CLEANUP_REC_FIELDS; j++) {
>> + size_t idx = raw->insn_idx[j], final;
>> + struct bpf_program *p;
>> +
>> + /* The end of a range is exclusive, so it may name the
>> + * instruction just past the last one of a function,
>> + * which belongs to the next function or to nothing at
>> + * all. Ask about the last instruction the range covers,
>> + * the way the kernel does.
>> + */
>> + if (j == CLEANUP_REC_END) {
>> + if (!idx) {
>> + pr_warn("%s: record %zu is an empty range\n",
>> + CLEANUP_SEC, i);
>> + return -LIBBPF_ERRNO__FORMAT;
>> + }
>> + idx--;
>> + }
> [Severity: Medium]
> Will this check catch empty ranges emitted by the compiler at non-zero
> offsets?
>
> If the compiler emits a .bpf_cleanup record where begin and end are both 5,
> the check for !idx evaluates to false. The code will then decrement idx to 4
> and look up the program for the previous instruction.
>
> If this empty range is located at the start of a function, looking up the
> previous instruction will fall into the preceding function, potentially
> triggering a spurious "spans functions" error or silently accepting invalid
> records.
Good point. Will fix this corner case.
>
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 16/20] libbpf: Carry the exception cleanup table through the light skeleton
2026-09-18 5:02 ` sashiko-bot
@ 2026-09-19 20:27 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 20:27 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/17/26 10:02 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] libbpf: integer overflow in cleanup_info_tot_sz leads to heap OOB read/write
> --
>
> commit b9300a1f6022da2aa8b343e05882b46e9f329c64
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> libbpf: Carry the exception cleanup table through the light skeleton
>
> This patch modifies the light skeleton generation to carry the exception
> cleanup table (cleanup_info), similar to func_info and line_info. It updates
> bpf_gen__prog_load() to allocate space for the records, set up union bpf_attr,
> and perform endianness swapping for cross-compiled cases.
>
>> diff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c
>> index af3a04f161ac1..2345fbdd46f56 100644
>> --- a/tools/lib/bpf/gen_loader.c
>> +++ b/tools/lib/bpf/gen_loader.c
> [ ... ]
>> @@ -998,6 +1000,9 @@ static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
>>
>> for (i = 0; i < gen->core_relo_cnt; i++)
>> bpf_core_relo_bswap(cr++);
>> +
>> + for (i = 0; i < load_attr->cleanup_info_cnt; i++)
>> + bpf_cleanup_info_bswap(ci++);
> [Severity: High]
> Could this loop result in an out-of-bounds heap read and write?
>
> If an attacker crafts a BPF ELF object with an extremely large .bpf_cleanup
> section, load_attr->cleanup_info_cnt can be very large. If the allocation
> size calculated in bpf_gen__prog_load() overflowed, this loop will process
> hundreds of millions of elements in an undersized buffer.
>
> This could cause memory corruption when cross-compiling (gen->swapped_endian
> is true).
It is user's problem if they intentionally want to have of lots and lots of
.bpf_cleanup section.
>
>> }
>>
>> void bpf_gen__prog_load(struct bpf_gen *gen,
>> @@ -1011,8 +1016,11 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
>> load_attr->line_info_rec_size;
>> int core_relo_tot_sz = gen->core_relo_cnt *
>> sizeof(struct bpf_core_relo);
>> + int cleanup_info_tot_sz = load_attr->cleanup_info_cnt *
>> + load_attr->cleanup_info_rec_size;
> [Severity: High]
> Does this unbounded 32-bit integer multiplication risk an overflow?
>
> If load_attr->cleanup_info_cnt is excessively large, multiplying it by
> cleanup_info_rec_size could wrap around modulo 2^32 to a small positive
> integer. This would result in add_data() allocating an undersized heap
> buffer for the cleanup info blob.
>
> Combined with the loop in info_blob_bswap(), this could lead to crashes
> or potentially arbitrary code execution when processing untrusted files
> using the light skeleton generator.
We should be okay. Similar to line_info, core_relo, etc., all of them
(including cleanup_info) will have a guard in add_data which ensures
the total size must be smaller than INT32_MAX.
>
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations
2026-09-18 5:01 ` sashiko-bot
@ 2026-09-19 20:31 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 20:31 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/17/26 10:01 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
> - [High] bpf/linker: NULL pointer deref in linker_append_elf_relos()
> - [Medium] bpf/linker: unaligned memory access in data relocations
> - [Low] bpf/linker: multi-line comment style violation
> --
>
> commit 2eadf13e135621b2e34d8946fe67664cd72e0e3d
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> libbpf: Let the static linker carry .bpf_cleanup relocations
>
> This commit updates the BPF static linker to support .bpf_cleanup relocations.
> It specifically adds handling for R_BPF_64_NODYLD32 and R_BPF_64_ABS32 to
> process compiler-emitted exception cleanup tables.
>
>> diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
>> index 78f92c39290af..e5c06023cb5b1 100644
>> --- a/tools/lib/bpf/linker.c
>> +++ b/tools/lib/bpf/linker.c
> [ ... ]
>
>> @@ -2274,6 +2275,22 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
>> insn->imm += sec->dst_off / sizeof(struct bpf_insn);
>> else
>> insn->imm += sec->dst_off;
>> + } else if (sym_type == R_BPF_64_NODYLD32 ||
>> + sym_type == R_BPF_64_ABS32) {
>> + __u32 *val;
>> +
>> + /* Two spellings of the one thing: LLVM
> [Severity: Low]
> This isn't a bug, but does this violate the BPF subsystem multi-line comment
> style? The preferred kernel style is to have the opening /* on its own line
> rather than putting text on the same line as the marker.
>
>> + * emits NODYLD32 for a .long against a
>> + * section symbol, GNU as emits ABS32
>> + * (bpf_reloc_type_lookup() maps
>> + * BFD_RELOC_32 to it), and the value
>> + * they describe is the same.
>> + */
>> + val = dst_linked_sec->raw_data + dst_rel->r_offset;
> [Severity: High]
> Could this lead to a NULL pointer dereference? Looking at
> linker_sanity_check_elf_relos(), SHT_NOBITS sections are explicitly allowed
> as relocatable sections.
>
> For SHT_NOBITS sections (like .bss), extend_sec() typically skips allocating
> raw_data, leaving it as NULL. If we process a relocation against such a
> section here, we might compute val = NULL + dst_rel->r_offset and then
> dereference it on the next line when handling sec->dst_off.
>
> [Severity: Medium]
> Is it also possible to hit an unaligned memory access here?
>
> While executable sections enforce alignment checks in
> linker_sanity_check_elf_relos(), there doesn't seem to be an equivalent
> alignment check for data section offsets. If an ELF provides an unaligned
> r_offset, the direct dereference of the unaligned val pointer might
> cause a crash on architectures with strict memory alignment requirements.
Okay, I will add a guard to check NULL pointer and unaligned memory access.
If this is true, will return an error.
>
>> + if (linker->swapped_endian)
>> + *val = bswap_32(bswap_32(*val) + sec->dst_off);
>> + else
>> + *val += sec->dst_off;
>> } else {
>> pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
>> return -EINVAL;
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations
2026-09-18 5:44 ` bot+bpf-ci
@ 2026-09-19 20:32 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 20:32 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/17/26 10:44 PM, bot+bpf-ci@kernel.org wrote:
>> An object that carries a compiler-emitted exception cleanup table cannot be
>> linked today. The table's fields are byte offsets into a code section,
>> materialised by a 32-bit relocation against that section's symbol with the
>> offset itself as the implicit addend, and the linker rejects both halves of
>> that: the relocation type is not in the list it accepts, and a relocation
>> against an STT_SECTION symbol from a non-executable section is an outright
>> error.
>>
>> Both spellings of that relocation have to be taken. LLVM emits
>> R_BPF_64_NODYLD32 for a .long against a section symbol; GNU as emits
>> R_BPF_64_ABS32, which is what bpf_reloc_type_lookup() maps BFD_RELOC_32 to.
>> They describe the same value, and the selftests are built with both
>> compilers.
>>
>> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
> This looks like a bug fix for objects that carry compiler-emitted exception
> cleanup tables. Should this include:
>
> Fixes: 9f0c317f6aa1 ("libbpf: Add support for new llvm bpf relocations")
No need for a Fixes tag.
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test
2026-09-18 5:58 ` bot+bpf-ci
@ 2026-09-19 20:34 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 20:34 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/17/26 10:58 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
>> new file mode 100644
>> index 0000000000000..ce2dac306a84c
>> --- /dev/null
>> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
>> @@ -0,0 +1,600 @@
>> +// SPDX-License-Identifier: GPL-2.0
>> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
>> +#include <vmlinux.h>
>> +#include <bpf/bpf_helpers.h>
>> +#include "bpf_experimental.h"
>> +#include "bpf_misc.h"
>> +#include "../test_kmods/bpf_testmod_kfunc.h"
>> +#include "exceptions_cleanup.h"
>> +
>> +__u64 input = 0;
>> +
>> +static __used __noinline void __kfunc_btf_anchor(void)
>> +{
>> + bpf_throw(0);
>> + bpf_preempt_disable();
>> + bpf_preempt_enable();
>> + bpf_unwind_resume();
>> +}
> Is bpf_unwind_resume() declared anywhere? It is called from C here but
> bpf_experimental.h declares only the other kfuncs this anchor exists for:
>
> tools/testing/selftests/bpf/bpf_experimental.h declares bpf_throw():
>
> extern void bpf_throw(u64 cookie) __ksym;
>
> and the preempt pair:
>
> void bpf_preempt_disable(void) __weak __ksym;
> void bpf_preempt_enable(void) __weak __ksym;
>
> but bpf_unwind_resume() does not appear in any header. vmlinux.h contains
> type definitions only, not kfunc prototypes.
>
> BPF objects are built with BPF_CFLAGS = -g -Wall -Werror -std=gnu11
> (tools/testing/selftests/bpf/Makefile). -Wimplicit-function-declaration
> is part of -Wall, and in C99+ mode modern Clang promotes it to an error,
> which with -Werror becomes fatal:
>
> progs/exceptions_cleanup_fail.c:17:2: error: call to undeclared
> function 'bpf_unwind_resume'; ISO C99 and later do not support
> implicit function declarations
>
> This would break the build of the whole BPF selftests tree. The same C-level
> call appears in the three other new prog files added by this series
> (progs/exceptions_cleanup.c, and in follow-ups progs/exceptions_cleanup_shapes.c
> and progs/exceptions_cleanup_light.c), so all of them would fail to compile.
>
> The anchor cannot simply drop the call either: the C call is what makes Clang
> emit the BTF extern FUNC for bpf_unwind_resume, which libbpf needs to resolve
> the inline-asm "call bpf_unwind_resume;" relocations. The fix would be to add
> the declaration:
>
> extern void bpf_unwind_resume(void) __ksym;
>
> in bpf_experimental.h alongside bpf_throw(). For comparison,
> progs/exceptions_cleanup_shapes.c (commit 198e7e8d6d72) does declare its
> sibling explicitly:
>
> extern void _Unwind_Resume(void) __ksym;
I didn't find anything wrong with this patch set and with my normal environment,
the test works fine.
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
* Re: [PATCH bpf-next v2 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach
2026-09-18 5:44 ` bot+bpf-ci
@ 2026-09-19 21:13 ` Yonghong Song
0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-19 21:13 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/17/26 10:44 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
>> index 630d2e207119a..5896cf83d15ee 100644
>> --- a/tools/testing/selftests/bpf/exceptions_cleanup.h
>> +++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
>> @@ -4,6 +4,7 @@
>> #define __EXCEPTIONS_CLEANUP_H__
>>
>> #define THROW_COOKIE 0x100
>> +#define INNER_COOKIE 0x200
>>
>> /* progs/exceptions_cleanup.c: one bit per frame that reports it ran. */
>> #define RAN_FOO3_PREEMPT 0x1
> [ ... ]
>
>> diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
>> index d1e45b765af26..ffc0b95191681 100644
>> --- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
>> +++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> [ ... ]
>
>> +static void test_shapes(void)
>> +{
>> + struct exceptions_cleanup_shapes *skel;
>> +
>> + skel = exceptions_cleanup_shapes__open_and_load();
>> + if (!ASSERT_OK_PTR(skel, "shapes open_and_load"))
>> + return;
> [ ... ]
>
>> + /* A throwing subprog named by a BPF_PSEUDO_FUNC no helper is handed: the
>> + * callback check has to look at the bpf_loop(), not at the ld_imm64.
>> + */
>> + if (test__start_subtest("addr_taken_no_throw"))
>> + run_shape(skel, skel->progs.entry_addr_taken, 1, 2, 0);
>> + if (test__start_subtest("addr_taken_throw"))
>> + run_shape(skel, skel->progs.entry_addr_taken, 101, THROW_COOKIE,
>> + RAN_ADDR_TAKEN);
> Does this comment accurately describe what the addr_taken shape tests?
>
> Looking at addr_taken_callee() in progs/exceptions_cleanup_shapes.c, the
> function does hand cb_thrower to bpf_loop(). The ld_imm64 and the
> bpf_loop() call are on the same path, after bpf_throw(). Since bpf_throw()
> is not declared noreturn, LLVM keeps that tail, and what actually makes
> the program load is that the verifier never reaches the bpf_loop() at all
> (it is dead after the throw).
>
> So push_callback_call() -> bpf_cleanup_check_callback() never fires.
>
> The shape exercised is 'a BPF_PSEUDO_FUNC whose helper call the verifier
> never reaches', not 'a BPF_PSEUDO_FUNC no helper is handed'. Could the
> comment be rephrased to reflect what the verifier actually sees?
Okay, will change comments.
>
> [ ... ]
>
>> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
>> new file mode 100644
>> index 0000000000000..f5eb2ff15c896
>> --- /dev/null
>> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
> [ ... ]
>
>> +/*
>> + * 2. A callee called from both a covered and an uncovered site: the pad is
>> + * recorded on the call site, not on the callee. The lock sits between the two
>> + * calls because the frame really would leak it if the uncovered call unwound.
>> + */
>> +static __used __noinline __u64 shared_callee(__u64 x)
>> +{
>> + if (x > 100)
>> + bpf_throw(THROW_COOKIE);
>> + return x + 1;
>> +}
>> +
>> +static __used __naked __noinline __u64 shared_frame(void)
>> +{
>> + asm volatile (
>> + "r1 = %[input] ll;"
>> + "r6 = *(u64 *)(r1 + 0);"
>> + "r1 = 0;"
>> + "call shared_callee;"
>> + "call bpf_rcu_read_lock;"
>> + "r1 = r6;"
>> +"1:" "call shared_callee;" /* cleanup region */
>> +"2:"
>> + "r6 = r0;"
>> + "call bpf_rcu_read_unlock;"
>> + "r0 = r6;"
>> + "exit;"
>> +"3:" /* landing pad */
>> + "r7 = r0;"
>> + "call bpf_rcu_read_unlock;"
>> + PAD_RAN("%[ran]")
>> + "r1 = r7;"
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + :
>> + : [ran]"i"(RAN_SHARED), __imm_addr(input), __imm_addr(pads_ran)
>> + : __clobber_all);
>> +}
> Does the uncovered call site verify that the pad wouldn't be wrongly
> dispatched if an exception occurred there?
>
> The uncovered site is reached with the constant 0 in r1, and shared_callee
> only throws when x > 100, so the verifier proves that call cannot throw
> and no unwind ever leaves the uncovered site - at verification time or at
> run time.
>
> The structural half of the property (only the second call's return address
> falls inside the record's native range) is exercised, but the behavioral
> half is not: a kernel that wrongly matched the uncovered site's return
> address against the record and ran the pad would go unnoticed, because
> that site never unwinds.
The above analysis exactly described the code so the above 'wrongly dispatched
if an exception occurred there' does not happen. I guess this tries to
capture incorrect kernel implementation.
>
> [ ... ]
>
>> +/*
>> + * 6. A tail call that is really taken: the target is a program in its own
>> + * right, so the walk ends there and this frame's pad does not run. The callee
>> + * can also throw on a path never taken, which keeps the pad out of the sweep.
>> + */
>> +struct {
>> + __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
>> + __uint(max_entries, 1);
>> + __uint(key_size, sizeof(__u32));
>> + __uint(value_size, sizeof(__u32));
>> +} taken_table SEC(".maps");
>> +
>> +SEC("syscall")
>> +int tc_target(void *ctx)
>> +{
>> + bpf_throw(THROW_COOKIE);
>> + return 0;
>> +}
>> +
>> +static __used __noinline __u64 tc_taken_callee(void *ctx, __u64 x)
>> +{
>> + /* Never true at run time; the verifier cannot know that, and its
>> + * unwind out of here is what keeps the caller's pad alive.
>> + */
>> + if (x == 7)
>> + bpf_throw(THROW_COOKIE);
>> + bpf_tail_call_static(ctx, &taken_table, 0);
>> + return 0;
>> +}
>> +
>> +SEC("syscall")
>> +__naked int entry_tail_taken(void)
>> +{
>> + asm volatile (
>> + "*(u64 *)(r10 - 8) = r1;" /* the context, straight from entry */
>> + "r1 = %[input] ll;"
>> + "r2 = *(u64 *)(r1 + 0);"
>> + "if r2 < 101 goto 8f;"
>> + "r1 = *(u64 *)(r10 - 8);"
>> +"1:" "call tc_taken_callee;" /* cleanup region */
>> +"2:"
>> + "exit;" /* the cookie, delivered at tc_target */
>> +"8:"
>> + "r0 = 0;"
>> + "exit;"
>> +"3:" /* landing pad: must not run */
>> + PAD_RAN("%[ran]")
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + :
>> + : [ran]"i"(RAN_TC_TAKEN), __imm_addr(input), __imm_addr(pads_ran)
>> + : __clobber_all);
>> +}
> Can the tail_call_taken subtest fail if the kernel wrongly continues the
> walk past the tail-call boundary?
This is a test which fixed some early regression tests. Yes, if kernel
implementation is wrong, tail_call_taken test may fail.
>
> entry_tail_taken narrows the argument before the covered call:
>
> "r2 = *(u64 *)(r1 + 0);" /* r2 = input */
> "if r2 < 101 goto 8f;" /* fallthrough => r2 in [101, U64_MAX] */
>
> tc_taken_callee is a static subprog, so check_func_call() copies the
> caller's r1-r5 verbatim into the callee frame. x therefore arrives with
> umin_value == 101, and is_branch_taken() resolves 'if (x == 7)' to
> never-taken, so the verifier never explores the bpf_throw(). Nothing then
> unwinds into entry_tail_taken during verification, the landing pad at
> label 3 is never marked seen, and opt_remove_dead_code() removes it.
>
> fixups.c then clears the annotation for any call site whose cleanup_pad
> falls inside the removed window:
>
> if (env->cleanup_info_cnt) {
> for (i = 0; i < env->insn_aux_data_len; i++) {
> u32 pad = aux_data[i].cleanup_pad;
> if (pad > off + cnt) aux_data[i].cleanup_pad = pad - cnt;
> else if (pad > off) aux_data[i].cleanup_pad = 0;
> }
> }
>
> so no native range is emitted for the region at all. The program still
> loads (mark_subprog_might_throw() is a static analysis, so bpf_check_cfg()
> still sees the pad edge and check_cleanup_info() runs before the sweep),
> and pads_ran stays 0 unconditionally. A kernel that wrongly continued
> bpf_stack_walker() past the tail-call boundary into entry_tail_taken's
> frame would find no record to match and would still leave pads_ran == 0,
> so the subtest passes on a correct and on a broken kernel alike.
>
> The in-code comment is also wrong on both clauses: the verifier CAN know
> x != 7 here, and consequently the unwind out of tc_taken_callee is NOT
> what keeps the caller's pad alive - nothing does.
The comment is:
/* Never true at run time; the verifier cannot know that, and its
* unwind out of here is what keeps the caller's pad alive.
*/
if (x == 7)
bpf_throw(THROW_COOKIE);
Yes, the 'x' will be a known value at run time. Comments need update.
>
> Contrast with shape 7 (entry_freplace, lines 862-878), the analogous 'walk
> ends in the callee's frame' shape, which deliberately passes the
> unnarrowed input value to fr_callee() with no 'if r1 < 101' guard, keeping
> fr_callee's identical 'if (x == 7) bpf_throw()' live. Dropping the
> 'if r2 < 101 goto 8f' guard from entry_tail_taken (or otherwise passing a
> value the verifier cannot exclude 7 from) would make the pad survive and
> give the subtest something to fail on.
>
> [ ... ]
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
^ permalink raw reply [flat|nested] 56+ messages in thread
end of thread, other threads:[~2026-09-19 21:13 UTC | newest]
Thread overview: 56+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 01/20] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 02/20] bpf: Add the bpf_unwind_resume() kfunc Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 03/20] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 4:55 ` Alexei Starovoitov
2026-09-19 17:42 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 04/20] bpf: Mark the call sites an exception cleanup table covers Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 05/20] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-18 4:59 ` sashiko-bot
2026-09-19 19:17 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 19:32 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 06/20] bpf: Explore the landing pads no call site reaches Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 19:32 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 08/20] bpf: Walk the exception unwind in the verifier Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 09/20] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 19:37 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 10/20] bpf: Dispatch exception cleanup pads from bpf_throw() Yonghong Song
2026-09-18 5:58 ` bot+bpf-ci
2026-09-19 19:54 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 11/20] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-18 5:03 ` sashiko-bot
2026-09-19 20:00 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 20:04 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 12/20] bpf, arm64: " Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 20:07 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 13/20] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-18 4:57 ` sashiko-bot
2026-09-19 20:18 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-18 5:00 ` sashiko-bot
2026-09-19 20:21 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 16/20] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-18 5:02 ` sashiko-bot
2026-09-19 20:27 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-18 5:01 ` sashiko-bot
2026-09-19 20:31 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 20:32 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
2026-09-18 4:59 ` sashiko-bot
2026-09-18 5:58 ` bot+bpf-ci
2026-09-19 20:34 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
2026-09-18 5:01 ` sashiko-bot
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 21:13 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 20/20] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox