* [PATCH bpf-next v9 01/23] bpf: Pack bpf_insn_aux_data flags into bit fields
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
@ 2026-10-08 7:50 ` Yonghong Song
2026-10-08 7:50 ` [PATCH bpf-next v9 02/23] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
` (21 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:50 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
struct bpf_insn_aux_data spreads its flags over nine bools, a u8 of three
bit fields and a u32 group with 25 of its 32 bits spare. Put them all in
one u64 word, alu_state with them, with orig_idx packing into the word's
upper half. The structure goes from 128 bytes to 120, and stays there
when a later patch adds a u32, which takes the tail padding.
No functional change.
Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 45 ++++++++++++++++++------------------
1 file changed, 23 insertions(+), 22 deletions(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index c51083c761cf..68636e1f048b 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -664,42 +664,43 @@ struct bpf_insn_aux_data {
u64 map_key_state; /* constant (32 bit) key tracking for maps */
int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
- bool nospec; /* do not execute this instruction speculatively */
- bool nospec_result; /* result is unsafe under speculation, nospec must follow */
- bool zext_dst; /* this insn zero extends dst reg */
- bool needs_zext; /* alu op needs to clear upper bits */
- bool prevent_zext; /* alu op cannot be zext (already used with 64-bit scalars) */
- bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
- bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
- bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
- bool arena_scalar; /* ldx/stx/st/atomic through a number, it's an address in arena */
- u8 alu_state; /* used in combination with alu_limit */
+ u64 nospec:1; /* do not execute this instruction speculatively */
+ u64 nospec_result:1; /* result is unsafe under speculation, nospec must follow */
+ u64 zext_dst:1; /* this insn zero extends dst reg */
+ u64 needs_zext:1; /* alu op needs to clear upper bits */
+ u64 prevent_zext:1; /* alu op cannot be zext (already used with 64-bit scalars) */
+ u64 non_sleepable:1; /* helper/kfunc may be called from non-sleepable context */
+ u64 is_iter_next:1; /* bpf_iter_<type>_next() kfunc call */
+ u64 call_with_percpu_alloc_ptr:1; /* {this,per}_cpu_ptr() with prog percpu alloc */
+ u64 arena_scalar:1; /* ldx/stx/st/atomic through a number, it's an address in arena */
+ u64 alu_state:8; /* used in combination with alu_limit */
/* true if STX or LDX instruction is a part of a spill/fill
* pattern for a bpf_fastcall call.
*/
- u8 fastcall_pattern:1;
+ u64 fastcall_pattern:1;
/* for CALL instructions, a number of spill/fill pairs in the
* bpf_fastcall pattern.
*/
- u8 fastcall_spills_num:3;
- u8 arg_prog:4;
+ u64 fastcall_spills_num:3;
+ u64 arg_prog:4;
- /* below fields are initialized once */
- unsigned int orig_idx; /* original instruction index */
- u32 jmp_point:1;
- u32 prune_point:1;
+ /* below flags are initialized once */
+ u64 jmp_point:1;
+ u64 prune_point:1;
/* ensure we check state equivalence and save state checkpoint and
* this instruction, regardless of any heuristics
*/
- u32 force_checkpoint:1;
+ u64 force_checkpoint:1;
/* true if instruction is a call to a helper function that
* accepts callback function as a parameter.
*/
- u32 calls_callback:1;
- u32 indirect_target:1; /* if it is an indirect jump target */
- u32 non_stack_access:1; /* instruction can access non-stack memory */
+ u64 calls_callback:1;
+ u64 indirect_target:1; /* if it is an indirect jump target */
+ u64 non_stack_access:1; /* instruction can access non-stack memory */
/* true if some jump or call instruction targets this instruction */
- u32 jump_target:1;
+ u64 jump_target:1;
+
+ unsigned int orig_idx; /* original instruction index, initialized once */
/*
* CFG strongly connected component this instruction belongs to,
* zero if it is a singleton SCC.
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 02/23] bpf: Accept the compiler's exception cleanup table at program load
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-10-08 7:50 ` [PATCH bpf-next v9 01/23] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
@ 2026-10-08 7:50 ` Yonghong Song
2026-10-08 7:50 ` [PATCH bpf-next v9 03/23] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
` (20 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:50 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
LLVM 23 added exception handling for BPF with the .bpf_cleanup section
[1]: Rust code compiled with panic=unwind runs cleanup code (Drop glue)
when an unwind passes through, and the BPF backend emits the section
from the landing pads. rustc does not fully support BPF exception
handling yet, and C can produce the section only with inline asm, which
is enough to test it.
Add the UAPI to carry the table to BPF_PROG_LOAD: cleanup_info,
cleanup_info_cnt and cleanup_info_rec_size. A record, struct
bpf_cleanup_info, says that calls in [begin_off, end_off) that unwind
resume at landing_pad_off, all instruction indices.
bpf_exc_check_info() checks the table at load time: records sorted by
begin_off, with non-empty and disjoint ranges, each inside one subprog,
no landing pad inside any range, and no offset naming the second half of
an ld_imm64. Nothing reads the table yet.
Link: https://github.com/llvm/llvm-project/pull/192164 [1]
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 2 +
include/uapi/linux/bpf.h | 14 ++++
kernel/bpf/Makefile | 2 +-
kernel/bpf/exception.c | 137 +++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 15 ++++
kernel/bpf/syscall.c | 2 +-
kernel/bpf/verifier.c | 6 ++
tools/include/uapi/linux/bpf.h | 14 ++++
8 files changed, 190 insertions(+), 2 deletions(-)
create mode 100644 kernel/bpf/exception.c
create mode 100644 kernel/bpf/exception.h
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 68636e1f048b..a5f493876993 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1021,6 +1021,8 @@ struct bpf_verifier_env {
struct spill_snapshot **callsite_at_stack;
u32 pass_cnt; /* number of times do_check() was called */
u32 subprog_cnt;
+ struct bpf_cleanup_info *cleanup_info;
+ u32 cleanup_info_cnt;
/* number of instructions analyzed by the verifier */
u32 prev_insn_processed, insn_processed;
/* number of jmps, calls, exits analyzed so far */
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index e0ed44b1bbcb..15dfba087201 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -1708,6 +1708,9 @@ union bpf_attr {
* verification.
*/
__s32 keyring_id;
+ __aligned_u64 cleanup_info; /* exception cleanup table */
+ __u32 cleanup_info_rec_size; /* userspace bpf_cleanup_info size */
+ __u32 cleanup_info_cnt; /* number of bpf_cleanup_info records */
};
struct { /* anonymous struct used by BPF_OBJ_* commands */
@@ -7644,6 +7647,17 @@ struct bpf_line_info {
__u32 line_col;
};
+/*
+ * One record of an exception cleanup table: calls in [begin_off, end_off)
+ * that unwind resume at landing_pad_off. All three are instruction offsets
+ * in the program as loaded.
+ */
+struct bpf_cleanup_info {
+ __u32 begin_off;
+ __u32 end_off;
+ __u32 landing_pad_off;
+};
+
struct bpf_spin_lock {
__u32 val;
};
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index c1f9b0d3468d..8a6947b3d13a 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -11,7 +11,7 @@ obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_
obj-$(CONFIG_BPF_SYSCALL) += hashtab.o arraymap.o percpu_freelist.o bpf_lru_list.o lpm_trie.o map_in_map.o bloom_filter.o
obj-$(CONFIG_BPF_SYSCALL) += local_storage.o queue_stack_maps.o ringbuf.o bpf_insn_array.o
obj-$(CONFIG_BPF_SYSCALL) += bpf_local_storage.o bpf_task_storage.o
-obj-$(CONFIG_BPF_SYSCALL) += fixups.o cfg.o states.o backtrack.o check_btf.o
+obj-$(CONFIG_BPF_SYSCALL) += fixups.o cfg.o states.o backtrack.o check_btf.o exception.o
obj-${CONFIG_BPF_LSM} += bpf_inode_storage.o
obj-$(CONFIG_BPF_SYSCALL) += disasm.o mprog.o
obj-$(CONFIG_BPF_JIT) += trampoline.o
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
new file mode 100644
index 000000000000..d6b8ca98e71c
--- /dev/null
+++ b/kernel/bpf/exception.c
@@ -0,0 +1,137 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bpf.h>
+#include <linux/bpf_verifier.h>
+#include <linux/slab.h>
+#include "exception.h"
+
+#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
+
+#define MIN_BPF_CLEANUP_INFO_SIZE 12
+#define MAX_CLEANUP_INFO_REC_SIZE 252 /* as MAX_FUNCINFO_REC_SIZE */
+
+int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
+ bpfptr_t uattr)
+{
+ u32 krec_size = sizeof(struct bpf_cleanup_info);
+ u32 i, nrec, urec_size, min_size, prev_end = 0;
+ struct bpf_cleanup_info *krecord;
+ bpfptr_t urecord;
+ int ret = -EINVAL;
+
+ nrec = attr->cleanup_info_cnt;
+ if (!nrec)
+ return 0;
+ if (nrec > env->prog->len) {
+ verbose(env, "cleanup info has %u records for %u instructions\n",
+ nrec, env->prog->len);
+ return -EINVAL;
+ }
+
+ urec_size = attr->cleanup_info_rec_size;
+ if (urec_size < MIN_BPF_CLEANUP_INFO_SIZE ||
+ urec_size > MAX_CLEANUP_INFO_REC_SIZE ||
+ urec_size % sizeof(u32)) {
+ verbose(env, "invalid cleanup info rec size %u\n", urec_size);
+ return -EINVAL;
+ }
+
+ krecord = kvcalloc(nrec, krec_size, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!krecord)
+ return -ENOMEM;
+
+ min_size = min_t(u32, krec_size, urec_size);
+ urecord = make_bpfptr(attr->cleanup_info, uattr.is_kernel);
+ for (i = 0; i < nrec; i++) {
+ struct bpf_subprog_info *sb, *se, *sl;
+ struct bpf_cleanup_info *rec = &krecord[i];
+
+ ret = bpf_check_uarg_tail_zero(urecord, krec_size, urec_size);
+ if (ret) {
+ if (ret == -E2BIG) {
+ verbose(env, "nonzero tailing record in cleanup info\n");
+ if (copy_to_bpfptr_offset(uattr,
+ offsetof(union bpf_attr,
+ cleanup_info_rec_size),
+ &min_size, sizeof(min_size)))
+ ret = -EFAULT;
+ }
+ goto err_free;
+ }
+
+ if (copy_from_bpfptr(rec, urecord, min_size)) {
+ ret = -EFAULT;
+ goto err_free;
+ }
+ bpfptr_add(&urecord, urec_size);
+
+ ret = -EINVAL;
+ if (rec->begin_off >= rec->end_off) {
+ verbose(env, "cleanup_info[%u]: begin %u >= end %u\n",
+ i, rec->begin_off, rec->end_off);
+ goto err_free;
+ }
+ if (i && rec->begin_off < prev_end) {
+ verbose(env,
+ "cleanup_info[%u]: range [%u,%u) is unsorted or overlaps the previous record\n",
+ i, rec->begin_off, rec->end_off);
+ goto err_free;
+ }
+ prev_end = rec->end_off;
+
+ sb = bpf_find_containing_subprog(env, rec->begin_off);
+ se = bpf_find_containing_subprog(env, rec->end_off - 1);
+ sl = bpf_find_containing_subprog(env, rec->landing_pad_off);
+ if (!sb || !se || !sl) {
+ verbose(env, "cleanup_info[%u]: offset out of range\n", i);
+ goto err_free;
+ }
+ if (sb != se || sb != sl) {
+ verbose(env,
+ "cleanup_info[%u]: range/landing pad span multiple subprogs\n",
+ i);
+ goto err_free;
+ }
+ /*
+ * A zero opcode is the second half of a 16-byte insn, not an
+ * insn. end_off is exclusive, so it may be one past the last.
+ */
+ if (!env->prog->insnsi[rec->begin_off].code ||
+ !env->prog->insnsi[rec->landing_pad_off].code ||
+ (rec->end_off < env->prog->len &&
+ !env->prog->insnsi[rec->end_off].code)) {
+ verbose(env, "cleanup_info[%u]: points at invalid insn\n", i);
+ goto err_free;
+ }
+ }
+
+ /* Reject a landing pad inside any call-site range, its own included. */
+ ret = -EINVAL;
+ for (i = 0; i < nrec; i++) {
+ u32 pad = krecord[i].landing_pad_off;
+ u32 l = 0, r = nrec;
+
+ while (l < r) {
+ u32 m = l + (r - l) / 2;
+
+ if (pad < krecord[m].begin_off) {
+ r = m;
+ } else if (pad >= krecord[m].end_off) {
+ l = m + 1;
+ } else {
+ verbose(env,
+ "cleanup_info[%u]: landing pad %u is inside the call-site range of cleanup_info[%u]\n",
+ i, pad, m);
+ goto err_free;
+ }
+ }
+ }
+
+ env->cleanup_info = krecord;
+ env->cleanup_info_cnt = nrec;
+ return 0;
+
+err_free:
+ kvfree(krecord);
+ return ret;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
new file mode 100644
index 000000000000..cf099dcc5b74
--- /dev/null
+++ b/kernel/bpf/exception.h
@@ -0,0 +1,15 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef __BPF_EXCEPTION_H
+#define __BPF_EXCEPTION_H
+
+#include <linux/bpfptr.h>
+#include <linux/types.h>
+
+union bpf_attr;
+struct bpf_verifier_env;
+
+int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
+ bpfptr_t uattr);
+
+#endif /* __BPF_EXCEPTION_H */
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index 654f896af265..10e0a693a06d 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -2948,7 +2948,7 @@ int __init __used bpf_multi_func(void) { return 0; }
BTF_ID_LIST_GLOBAL_SINGLE(bpf_multi_func_btf_id, func, bpf_multi_func)
/* last field in 'union bpf_attr' used by this command */
-#define BPF_PROG_LOAD_LAST_FIELD keyring_id
+#define BPF_PROG_LOAD_LAST_FIELD cleanup_info_cnt
static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, struct bpf_log_attr *attr_log)
{
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 353bde9ae227..60413bf0ad3f 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -37,6 +37,7 @@
#include "diagnostics.h"
#include "disasm.h"
+#include "exception.h"
static const struct bpf_verifier_ops * const bpf_verifier_ops[] = {
#define BPF_PROG_TYPE(_id, _name, prog_ctx_type, kern_ctx_type) \
@@ -22706,6 +22707,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
if (ret < 0)
goto skip_full_check;
+ ret = bpf_exc_check_info(env, attr, uattr);
+ if (ret < 0)
+ goto skip_full_check;
+
/* Validate instructions and resolve the program's referenced resources. */
ret = check_and_resolve_insns(env);
if (ret < 0)
@@ -22923,6 +22928,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
kvfree(env->callx_edges);
kvfree(env->func_ptrs);
bpf_diag_free(env);
+ kvfree(env->cleanup_info);
kvfree(env);
return ret;
}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index e0ed44b1bbcb..15dfba087201 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -1708,6 +1708,9 @@ union bpf_attr {
* verification.
*/
__s32 keyring_id;
+ __aligned_u64 cleanup_info; /* exception cleanup table */
+ __u32 cleanup_info_rec_size; /* userspace bpf_cleanup_info size */
+ __u32 cleanup_info_cnt; /* number of bpf_cleanup_info records */
};
struct { /* anonymous struct used by BPF_OBJ_* commands */
@@ -7644,6 +7647,17 @@ struct bpf_line_info {
__u32 line_col;
};
+/*
+ * One record of an exception cleanup table: calls in [begin_off, end_off)
+ * that unwind resume at landing_pad_off. All three are instruction offsets
+ * in the program as loaded.
+ */
+struct bpf_cleanup_info {
+ __u32 begin_off;
+ __u32 end_off;
+ __u32 landing_pad_off;
+};
+
struct bpf_spin_lock {
__u32 val;
};
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 03/23] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-10-08 7:50 ` [PATCH bpf-next v9 01/23] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-10-08 7:50 ` [PATCH bpf-next v9 02/23] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
@ 2026-10-08 7:50 ` Yonghong Song
2026-10-08 7:50 ` [PATCH bpf-next v9 04/23] bpf: Keep a call site's landing pad in insn_aux_data, add lookups Yonghong Song
` (19 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:50 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
An exception cleanup needs two terminators:
- bpf_unwind() begins an unwind: the frames it leaves run their landing
pads on the way out to the program's exit. It is separate from
bpf_throw(), which discards the frames in between.
- bpf_unwind_resume() ends a landing pad, carrying the unwind on. The
compiler calls it _Unwind_Resume, which a later libbpf patch resolves
to this name, and passes it the exception object, which the kfunc
takes as ptr__ign.
Neither is registered yet: that waits for the patch that dispatches pads
at run time.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/helpers.c | 14 ++++++++++++++
1 file changed, 14 insertions(+)
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index a284f20c97d5..4eccd6742eba 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -3424,6 +3424,10 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
return false;
}
+__bpf_kfunc void bpf_unwind(void)
+{
+}
+
__bpf_kfunc void bpf_throw(u64 cookie)
{
struct bpf_throw_ctx ctx = {};
@@ -3445,6 +3449,16 @@ __bpf_kfunc void bpf_throw(u64 cookie)
WARN(1, "A call to BPF exception callback should never return\n");
}
+__bpf_kfunc void bpf_unwind_resume(void *ptr__ign)
+{
+ /*
+ * Never reached: the verifier accepts this kfunc only as a frame
+ * terminator and bpf_do_misc_fixups() lowers every one of them to
+ * 'r0 = 0; exit', so no call to this body survives to run.
+ */
+ WARN_ONCE(1, "exception cleanup resume was not lowered to a return\n");
+}
+
__bpf_kfunc int bpf_wq_init(struct bpf_wq *wq, void *p__const_map, unsigned int flags)
{
struct bpf_async_kern *async = (struct bpf_async_kern *)wq;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 04/23] bpf: Keep a call site's landing pad in insn_aux_data, add lookups
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (2 preceding siblings ...)
2026-10-08 7:50 ` [PATCH bpf-next v9 03/23] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
@ 2026-10-08 7:50 ` Yonghong Song
2026-10-08 8:01 ` sashiko-bot
2026-10-08 7:50 ` [PATCH bpf-next v9 05/23] bpf: Mark covered call sites and check a program can take a table Yonghong Song
` (18 subsequent siblings)
22 siblings, 1 reply; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:50 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A call site's landing pad, if any, is kept in insn_aux_data as
cleanup_pad, so the three places that move instructions --
bpf_patch_insn_data(), verifier_remove_insns() and bpf_opt_remove_nops()
-- keep it in step. Add lookups for it and for calls to bpf_unwind() and
bpf_unwind_resume(); their users come in later patches.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 5 +++++
kernel/bpf/exception.c | 24 ++++++++++++++++++++++++
kernel/bpf/exception.h | 4 ++++
kernel/bpf/fixups.c | 26 +++++++++++++++++++++++++-
4 files changed, 58 insertions(+), 1 deletion(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index a5f493876993..e6bde421ab3b 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -701,6 +701,11 @@ struct bpf_insn_aux_data {
u64 jump_target:1;
unsigned int orig_idx; /* original instruction index, initialized once */
+ /*
+ * 1 + the instruction index of the exception cleanup landing pad
+ * this call site unwinds to, or 0 for none.
+ */
+ u32 cleanup_pad;
/*
* CFG strongly connected component this instruction belongs to,
* zero if it is a singleton SCC.
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index d6b8ca98e71c..3ea1bff5cc90 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -2,6 +2,8 @@
/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
#include <linux/bpf.h>
#include <linux/bpf_verifier.h>
+#include <linux/btf_ids.h>
+#include <linux/filter.h>
#include <linux/slab.h>
#include "exception.h"
@@ -135,3 +137,25 @@ int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
kvfree(krecord);
return ret;
}
+
+BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
+BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
+
+bool bpf_is_unwind_kfunc(const struct bpf_insn *insn)
+{
+ return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
+ insn->imm == bpf_unwind_id[0];
+}
+
+bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
+{
+ return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
+ insn->imm == bpf_unwind_resume_id[0];
+}
+
+int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
+{
+ u32 pad = env->insn_aux_data[idx].cleanup_pad;
+
+ return pad ? (int)pad - 1 : -1;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index cf099dcc5b74..d5c6ac459870 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -8,8 +8,12 @@
union bpf_attr;
struct bpf_verifier_env;
+struct bpf_insn;
int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
bpfptr_t uattr);
+int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
+bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
+bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
#endif /* __BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 206cc9a614a9..64baee1a37e6 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -244,11 +244,18 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
data[i].non_stack_access =
data[off + cnt - 1].non_stack_access;
data[off + cnt - 1].non_stack_access = false;
+ data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
+ data[off + cnt - 1].cleanup_pad = 0;
} else if (bpf_is_mem_insn(insn + i)) {
data[i].non_stack_access = true;
}
}
+ if (env->cleanup_info_cnt)
+ for (i = 0; i < prog_len; i++)
+ if (data[i].cleanup_pad > off + 1)
+ data[i].cleanup_pad += cnt - 1;
+
/*
* Last slot instruction could be a newly generated
* BPF_ST/BPF_LDX/BPF_STX, systematically mark it for non-stack access
@@ -595,6 +602,7 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
unsigned int orig_prog_len = env->prog->len;
int err;
+ u32 i;
if (bpf_rewrite_must_abort())
return -EINTR;
@@ -623,6 +631,17 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
sizeof(*aux_data) * (orig_prog_len - off - cnt));
env->insn_aux_data_len -= cnt;
+ if (env->cleanup_info_cnt) {
+ for (i = 0; i < env->insn_aux_data_len; i++) {
+ u32 pad = aux_data[i].cleanup_pad;
+
+ if (pad > off + cnt)
+ aux_data[i].cleanup_pad = pad - cnt;
+ else if (pad > off)
+ aux_data[i].cleanup_pad = 0;
+ }
+ }
+
return 0;
}
@@ -728,7 +747,7 @@ int bpf_opt_remove_nops(struct bpf_verifier_env *env)
struct bpf_insn *insn = env->prog->insnsi;
int insn_cnt = env->prog->len;
bool is_may_goto_0, is_ja;
- int i, err;
+ int i, j, err;
for (i = 0; i < insn_cnt; i++) {
is_may_goto_0 = !memcmp(&insn[i], &MAY_GOTO_0, sizeof(MAY_GOTO_0));
@@ -739,6 +758,11 @@ int bpf_opt_remove_nops(struct bpf_verifier_env *env)
if (aux[i].indirect_target)
continue;
+ if (env->cleanup_info_cnt)
+ for (j = 0; j < insn_cnt; j++)
+ if (env->insn_aux_data[j].cleanup_pad == i + 1)
+ env->insn_aux_data[j].cleanup_pad = i + 2;
+
err = verifier_remove_insns(env, i, 1);
if (err)
return err;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 04/23] bpf: Keep a call site's landing pad in insn_aux_data, add lookups
2026-10-08 7:50 ` [PATCH bpf-next v9 04/23] bpf: Keep a call site's landing pad in insn_aux_data, add lookups Yonghong Song
@ 2026-10-08 8:01 ` sashiko-bot
2026-10-08 15:58 ` Yonghong Song
0 siblings, 1 reply; 38+ messages in thread
From: sashiko-bot @ 2026-10-08 8:01 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: adjust_insn_aux_data() fails to shift cleanup_pad for BPF_PATCH_MOVE_TARGET
--
commit 1bdc64eca7dc9ba53a20b6b6c4d0e4448bdd4aa8
Author: Yonghong Song <yonghong.song@linux.dev>
bpf: Keep a call site's landing pad in insn_aux_data, add lookups
A call site's landing pad, if any, is kept in insn_aux_data as
cleanup_pad, so the three places that move instructions --
bpf_patch_insn_data(), verifier_remove_insns() and bpf_opt_remove_nops()
-- keep it in step. Add lookups for it and for calls to bpf_unwind() and
bpf_unwind_resume(); their users come in later patches.
> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
> index 206cc9a614a95..64baee1a37e60 100644
> --- a/kernel/bpf/fixups.c
> +++ b/kernel/bpf/fixups.c
> @@ -244,11 +244,18 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
> data[i].non_stack_access =
> data[off + cnt - 1].non_stack_access;
> data[off + cnt - 1].non_stack_access = false;
> + data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
> + data[off + cnt - 1].cleanup_pad = 0;
> } else if (bpf_is_mem_insn(insn + i)) {
> data[i].non_stack_access = true;
> }
> }
>
> + if (env->cleanup_info_cnt)
> + for (i = 0; i < prog_len; i++)
> + if (data[i].cleanup_pad > off + 1)
> + data[i].cleanup_pad += cnt - 1;
[Severity: Medium]
Does this logic fail to adjust the landing pad when mode is
BPF_PATCH_MOVE_TARGET?
By unconditionally using off + 1, this appears to assume
BPF_PATCH_KEEP_TARGET semantics.
If bpf_patch_insn_data() is called with BPF_PATCH_MOVE_TARGET and the landing
pad is located exactly at the patched instruction, data[i].cleanup_pad would
be equal to off + 1 (the 1-based index). The condition
data[i].cleanup_pad > off + 1 evaluates to false, meaning the landing pad
index is not shifted.
As a result, the landing pad remains pointing at the newly inserted
instructions instead of following the original instruction shifted to
off + cnt - 1.
Could this cause the unwinder to jump to the wrong instruction offset,
potentially leading to state corruption during an exception unwind?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261008074959.2993751-1-yonghong.song@linux.dev?part=4
^ permalink raw reply [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 04/23] bpf: Keep a call site's landing pad in insn_aux_data, add lookups
2026-10-08 8:01 ` sashiko-bot
@ 2026-10-08 15:58 ` Yonghong Song
0 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 15:58 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 10/8/26 10:01 AM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Medium] bpf: adjust_insn_aux_data() fails to shift cleanup_pad for BPF_PATCH_MOVE_TARGET
> --
>
> commit 1bdc64eca7dc9ba53a20b6b6c4d0e4448bdd4aa8
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Keep a call site's landing pad in insn_aux_data, add lookups
>
> A call site's landing pad, if any, is kept in insn_aux_data as
> cleanup_pad, so the three places that move instructions --
> bpf_patch_insn_data(), verifier_remove_insns() and bpf_opt_remove_nops()
> -- keep it in step. Add lookups for it and for calls to bpf_unwind() and
> bpf_unwind_resume(); their users come in later patches.
>
>> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
>> index 206cc9a614a95..64baee1a37e60 100644
>> --- a/kernel/bpf/fixups.c
>> +++ b/kernel/bpf/fixups.c
>> @@ -244,11 +244,18 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
>> data[i].non_stack_access =
>> data[off + cnt - 1].non_stack_access;
>> data[off + cnt - 1].non_stack_access = false;
>> + data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
>> + data[off + cnt - 1].cleanup_pad = 0;
>> } else if (bpf_is_mem_insn(insn + i)) {
>> data[i].non_stack_access = true;
>> }
>> }
>>
>> + if (env->cleanup_info_cnt)
>> + for (i = 0; i < prog_len; i++)
>> + if (data[i].cleanup_pad > off + 1)
>> + data[i].cleanup_pad += cnt - 1;
> [Severity: Medium]
> Does this logic fail to adjust the landing pad when mode is
> BPF_PATCH_MOVE_TARGET?
>
> By unconditionally using off + 1, this appears to assume
> BPF_PATCH_KEEP_TARGET semantics.
>
> If bpf_patch_insn_data() is called with BPF_PATCH_MOVE_TARGET and the landing
> pad is located exactly at the patched instruction, data[i].cleanup_pad would
> be equal to off + 1 (the 1-based index). The condition
> data[i].cleanup_pad > off + 1 evaluates to false, meaning the landing pad
> index is not shifted.
>
> As a result, the landing pad remains pointing at the newly inserted
> instructions instead of following the original instruction shifted to
> off + cnt - 1.
>
> Could this cause the unwinder to jump to the wrong instruction offset,
> potentially leading to state corruption during an exception unwind?
This won't happen. The below is v7 reply to sashiko. v8 does not pop
up this issue. and now v9 it comes back.
====
Okay, commit "bpf: Retarget indirect jump targets across prologue prepends"
introduced BPF_PATCH_MOVE_TARGET and merged last Friday. It solved three
cases for ops->gen_epilogue, ops->gen_prologue || env->seen_direct_write, and
stack slots for subprogs. ops->gen_epilogue has been rejected in patch 5.
We cannot allow ops->gen_epilogue since it may silently exit.
For other cases in "bpf: Retarget indirect jump targets across prologue prepends",
The above commit should already handle this.
For the other two, cleanup_pad == off + 1 cannot happen, because MOVE only
patches an entry insn and a landing pad cannot start at insn 0: the entry is
always walked outside a pad first, so reaching it again from an unwind
fails bpf_exc_check_insn() with "insn %u runs both inside and outside a
landing pad". Pads after off are shifted by the existing `> off + 1`, and a
covered call at off moves with its insn_aux_data.
====
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH bpf-next v9 05/23] bpf: Mark covered call sites and check a program can take a table
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (3 preceding siblings ...)
2026-10-08 7:50 ` [PATCH bpf-next v9 04/23] bpf: Keep a call site's landing pad in insn_aux_data, add lookups Yonghong Song
@ 2026-10-08 7:50 ` Yonghong Song
2026-10-08 7:50 ` [PATCH bpf-next v9 06/23] bpf: Make exception landing pads reachable in the CFG Yonghong Song
` (17 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:50 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
- bpf_exc_prepare(), run before the CFG walk, records the landing pad of
every call in a cleanup record's range that can unwind: a BPF-to-BPF
call, direct or indirect, or bpf_unwind().
- bpf_exc_check_prog() refuses a table where pad dispatch cannot work:
- an offloaded program, or one without a JIT that dispatches pads
- a verifier_ops gen_epilogue, as bpf_qdisc's: it is planted on the
exits bpf_convert_ctx_accesses() sees, and the exits an unwind
returns through are added after it
- an exception callback or a bpf_throw() anywhere: a throw leaves
without rewriting return addresses, so no pad would run
It runs after check_attach_btf_id(), since a struct_ops program's
gen_epilogue is only known then, and marks what passes jit_required.
bpf_jit_supports_cleanup_pads() says no until the arch patches.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/filter.h | 1 +
kernel/bpf/core.c | 5 +++
kernel/bpf/exception.c | 70 ++++++++++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 2 ++
kernel/bpf/verifier.c | 9 ++++++
5 files changed, 87 insertions(+)
diff --git a/include/linux/filter.h b/include/linux/filter.h
index 9339c6131f8f..c8ca863e352a 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1248,6 +1248,7 @@ bool bpf_jit_supports_stack_args(void);
bool bpf_jit_supports_arena_args(void);
bool bpf_jit_supports_far_kfunc_call(void);
bool bpf_jit_supports_exceptions(void);
+bool bpf_jit_supports_cleanup_pads(void);
bool bpf_jit_supports_ptr_xchg(void);
bool bpf_jit_supports_arena(void);
bool bpf_jit_supports_arena_scalar(void);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index 05c89396119a..078bcccaf242 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -3528,6 +3528,11 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
{
}
+bool __weak bpf_jit_supports_cleanup_pads(void)
+{
+ return false;
+}
+
bool __weak bpf_jit_supports_timed_may_goto(void)
{
return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index 3ea1bff5cc90..d85ca358b463 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -141,6 +141,76 @@ int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
+static int reject_throw(struct bpf_verifier_env *env)
+{
+ u32 i;
+
+ for (i = 0; i < env->prog->len; i++) {
+ if (!bpf_is_throw_kfunc(&env->prog->insnsi[i]))
+ continue;
+ verbose(env,
+ "exception cleanup cannot be combined with bpf_throw at insn %u\n",
+ i);
+ return -EINVAL;
+ }
+ return 0;
+}
+
+static void mark_call_sites(struct bpf_verifier_env *env)
+{
+ u32 i, j;
+
+ for (i = 0; i < env->cleanup_info_cnt; i++) {
+ struct bpf_cleanup_info *rec = &env->cleanup_info[i];
+
+ for (j = rec->begin_off; j < rec->end_off; j++) {
+ struct bpf_insn *insn = &env->prog->insnsi[j];
+
+ if (!bpf_pseudo_call(insn) && !bpf_is_callx(insn) &&
+ !bpf_is_unwind_kfunc(insn))
+ continue;
+ env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
+ }
+ }
+}
+
+int bpf_exc_check_prog(struct bpf_verifier_env *env)
+{
+ int err;
+
+ if (bpf_prog_is_offloaded(env->prog->aux)) {
+ verbose(env,
+ "exception cleanup is not supported for offloaded programs\n");
+ return -EINVAL;
+ }
+ if (!bpf_jit_supports_cleanup_pads() || !env->prog->jit_requested) {
+ verbose(env,
+ "exception cleanup needs a JIT that can dispatch landing pads\n");
+ return -EOPNOTSUPP;
+ }
+ if (env->ops->gen_epilogue) {
+ verbose(env,
+ "exception cleanup is not supported for a program with an epilogue\n");
+ return -EOPNOTSUPP;
+ }
+ if (env->exception_callback_subprog) {
+ verbose(env,
+ "exception cleanup cannot be combined with an exception callback\n");
+ return -EINVAL;
+ }
+ err = reject_throw(env);
+ if (err)
+ return err;
+ env->prog->jit_required = 1;
+ return 0;
+}
+
+void bpf_exc_prepare(struct bpf_verifier_env *env)
+{
+ if (env->cleanup_info_cnt)
+ mark_call_sites(env);
+}
+
bool bpf_is_unwind_kfunc(const struct bpf_insn *insn)
{
return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index d5c6ac459870..9b75e556c82e 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -12,6 +12,8 @@ struct bpf_insn;
int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
bpfptr_t uattr);
+void bpf_exc_prepare(struct bpf_verifier_env *env);
+int bpf_exc_check_prog(struct bpf_verifier_env *env);
int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 60413bf0ad3f..0cf30220b777 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22711,6 +22711,9 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
if (ret < 0)
goto skip_full_check;
+ /* The CFG needs an edge from a call in a cleanup range to its pad. */
+ bpf_exc_prepare(env);
+
/* Validate instructions and resolve the program's referenced resources. */
ret = check_and_resolve_insns(env);
if (ret < 0)
@@ -22748,6 +22751,12 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
if (ret)
goto skip_full_check;
+ if (env->cleanup_info_cnt) {
+ ret = bpf_exc_check_prog(env);
+ if (ret)
+ goto skip_full_check;
+ }
+
ret = bpf_compute_const_regs(env);
if (ret < 0)
goto skip_full_check;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 06/23] bpf: Make exception landing pads reachable in the CFG
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (4 preceding siblings ...)
2026-10-08 7:50 ` [PATCH bpf-next v9 05/23] bpf: Mark covered call sites and check a program can take a table Yonghong Song
@ 2026-10-08 7:50 ` Yonghong Song
2026-10-08 7:50 ` [PATCH bpf-next v9 07/23] bpf: Verify an unwind through landing pads and epilogues Yonghong Song
` (16 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:50 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A call that a cleanup record covers can also reach the record's landing
pad. Add that edge to the CFG walk, to bpf_insn_successors(), and to
stack liveness, where a slot is alive across a call when the call's pad
reads it.
The CFG walk pushes the pad in a visit of its own, and the fall-through
on the next visit. bpf_check_cfg() re-visits the top of its stack until
DONE_EXPLORING, so a visit pushing both would leave the pad DISCOVERED
off the current path, and push_insn() would take a branch from one pad
into another -- how a frame with two regions chains them -- for a
back-edge.
Without the liveness change, a slot only the pad reads looked dead, and
clean_verifier_state() poisoned it while the callee ran.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/cfg.c | 38 ++++++++++++++++++++++++++++++++++++++
kernel/bpf/liveness.c | 24 +++++++++++++++++++++++-
2 files changed, 61 insertions(+), 1 deletion(-)
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index d8a579680e5b..63afbc5fb296 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -6,6 +6,7 @@
#include <linux/sort.h>
#include "diagnostics.h"
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -160,6 +161,38 @@ static int push_insn(int t, int w, int e, struct bpf_verifier_env *env)
return DONE_EXPLORING;
}
+static int visit_cleanup_pad_edge(int t, struct bpf_verifier_env *env)
+{
+ int *insn_stack = env->cfg.insn_stack;
+ int *insn_state = env->cfg.insn_state;
+ int w;
+
+ if (!env->cleanup_info_cnt)
+ return DONE_EXPLORING;
+ w = bpf_exc_pad_of_call(env, t);
+ if (w < 0)
+ return DONE_EXPLORING;
+
+ /*
+ * @t is a call that may branch here, and @w is the target of that
+ * branch, so both are prune points. @w especially: every covered call
+ * site in a region unwinds to the same pad, and without a prune point
+ * at its head the verifier walks the pad again for each of them.
+ */
+ mark_prune_point(env, t);
+ mark_prune_point(env, w);
+ mark_jmp_point(env, w);
+ mark_jump_target(env, w);
+
+ if (insn_state[w])
+ return DONE_EXPLORING;
+ if (env->cfg.cur_stack >= env->prog->len)
+ return -E2BIG;
+ insn_stack[env->cfg.cur_stack++] = w;
+ insn_state[w] |= DISCOVERED;
+ return KEEP_EXPLORING;
+}
+
static int visit_func_call_insn(int t, struct bpf_insn *insns,
struct bpf_verifier_env *env,
bool visit_callee)
@@ -167,6 +200,11 @@ static int visit_func_call_insn(int t, struct bpf_insn *insns,
int ret, insn_sz;
int w;
+ /* One push per visit: @t is revisited once the pad is explored. */
+ ret = visit_cleanup_pad_edge(t, env);
+ if (ret != DONE_EXPLORING)
+ return ret;
+
insn_sz = bpf_is_ldimm64(&insns[t]) ? 2 : 1;
ret = push_insn(t, t + insn_sz, FALLTHROUGH, env);
if (ret)
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index cd9523f69298..2e7eead71c93 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -8,6 +8,8 @@
#include <linux/slab.h>
#include <linux/sort.h>
+#include "exception.h"
+
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
/*
@@ -384,6 +386,18 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
succ->items[succ->cnt++] = exit_idx;
}
+ /*
+ * A covered call can also leave through its landing pad. Only a call to
+ * a subprog or to bpf_unwind() is covered, so the two entries of
+ * env->succ are enough.
+ */
+ if (unlikely(env->cleanup_info_cnt)) {
+ int pad = bpf_exc_pad_of_call(env, idx);
+
+ if (pad >= 0)
+ succ->items[succ->cnt++] = pad;
+ }
+
return succ;
}
@@ -510,7 +524,8 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
* Slot is alive if it is read before q->insn_idx in current func instance,
* or if for some outer func instance:
* - alive before callsite if callsite calls callback or is callx, otherwise
- * - alive after callsite
+ * - alive after callsite,
+ * - or alive at the landing pad a cleanup record gives the callsite
*/
struct live_stack_query *q = &env->liveness->live_stack_query;
struct func_instance *instance, *curframe_instance;
@@ -545,6 +560,13 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
alive = callee_stack_access_at_callsite(env, callsite)
? is_live_before(instance, callsite, rel, half_spi)
: is_live_before(instance, callsite + 1, rel, half_spi);
+
+ if (!alive && unlikely(env->cleanup_info_cnt)) {
+ int pad = bpf_exc_pad_of_call(env, callsite);
+
+ if (pad >= 0)
+ alive = is_live_before(instance, pad, rel, half_spi);
+ }
if (alive)
return true;
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 07/23] bpf: Verify an unwind through landing pads and epilogues
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (5 preceding siblings ...)
2026-10-08 7:50 ` [PATCH bpf-next v9 06/23] bpf: Make exception landing pads reachable in the CFG Yonghong Song
@ 2026-10-08 7:50 ` Yonghong Song
2026-10-08 8:57 ` bot+bpf-ci
2026-10-08 7:50 ` [PATCH bpf-next v9 08/23] bpf: Refuse a landing pad that does not resume Yonghong Song
` (15 subsequent siblings)
22 siblings, 1 reply; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:50 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Once a later patch makes bpf_unwind() dispatch pads, an unwind rewrites
the return address of every frame it passes: to the pad where a record
covers the frame's call, else to the frame's epilogue. Each frame then
returns normally, a pad through its resume, which is lowered to
'r0 = 0; exit'. The verifier follows the unwind the same way: into the
pad of each frame that has one, out of each frame that does not, and out
of main as a program exit.
instruction goes on at
----------------------------------- ---------------------------------
bpf_unwind(), record over it its own frame's pad
bpf_unwind(), no record over it unwind_frames()
bpf_unwind_resume() unwind_frames()
call to a global subprog that can next insn, and the unwind from the
unwind returned state: to the call's pad,
or on through unwind_frames()
unwind_frames(), for main -> A -> B -> C where only A's call is covered:
program run time: bpf_unwind() in C
----------------------------------- ---------------------------------
main: call A main returns via its epilogue
A: 1: call B [1, 2) -> P A resumes at P
2: ...
P: <drop A's resources>
call bpf_unwind_resume
B: call C no record B returns via its epilogue
C: call bpf_unwind C returns via 'r0 = 0; exit'
frame at C's bpf_unwind() unwind_frames() at P's resume
----- ------------------- ----------------- -----------------
3 C <- curframe popped
2 B popped
1 A A <- curframe, P popped
0 main main exit with r0 = 0
P gets A's frame as B and C left it, with r0 unknown and r1-r5 cleared.
Main returning goes through process_bpf_exit_full(), the only place an
unwind's resources are checked: a frame it pops may leave something for
a caller's pad to drop.
The pad's state comes from the unwind, not from a snapshot at the call:
before unwinding, the callee may have written the caller's stack through
a pointer, overwritten a spilled pointer, reinitialised a dynptr or an
iterator, or changed packet data, none of which its epilogue restores. A
global subprog is verified on its own, so the unwind out of a call to one
starts from the state the call returns in, after check_func_call().
Precision backtracking gets three new edges:
edge frame
----------------------- ---------------------------------------------
global call <- its pad stays in the caller (subseq_idx is the pad)
unwind <- pad moves to the frame the unwind left, recorded
in its history entry under INSN_F_UNWIND
unwind <- main's exit the same, after a jump from the unwinding insn
to where the path ends
Where no frame has a pad, the path ends at E, the insn it exits from in
place of a BPF_EXIT: main's call into the popped frames, else the
unwinding insn U itself. Backtracking for the return check starts at E
without processing it, while r0 is set at U, Q being U's predecessor:
frames popped none popped
main: E: call A main: Q: r6 = 0
A: call B E = U: call bpf_unwind
B: Q: r6 = 0
U: call bpf_unwind
walk: E -> U -> Q -> ... walk: E -> U -> Q -> ...
^ skipped ^ skipped
The jump from U to E takes the walk to U rather than to the insn before
E; where E is U, it is what gets U processed at all. U's INSN_F_UNWIND
mark matters there only for an unwinding global call, which backtracking
would otherwise take for one that returned.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 9 ++-
kernel/bpf/backtrack.c | 45 ++++++++++-
kernel/bpf/cfg.c | 88 +++++++++++++++++++++
kernel/bpf/states.c | 8 +-
kernel/bpf/verifier.c | 148 ++++++++++++++++++++++++++++++++++-
5 files changed, 291 insertions(+), 7 deletions(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index e6bde421ab3b..d3740ad20235 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -383,6 +383,8 @@ enum {
INSN_F_SRC_REG_STACK = BIT(2), /* src_reg is PTR_TO_STACK */
INSN_F_STACK_ARG_ACCESS = BIT(3),
+
+ INSN_F_UNWIND = BIT(4),
};
/* Registers linked to one jump condition that a history entry can record */
@@ -393,8 +395,8 @@ struct bpf_jmp_history_entry {
u32 idx : 20;
u32 frame : 4; /* stack access frame number */
/* special INSN_F_xxx flags */
- u32 flags : 4;
- u32 : 4;
+ u32 flags : 5;
+ u32 : 3;
u32 prev_idx : 20;
u32 spi : 12; /* stack slot index */
/*
@@ -480,6 +482,8 @@ struct bpf_verifier_state {
bool speculative;
bool in_sleepable;
+ /* the path ends with an unwind returning from frame 0 */
+ bool unwind_exit;
/* first and last insn idx of this verifier state */
u32 first_insn_idx;
@@ -837,6 +841,7 @@ struct bpf_subprog_info {
s16 fastcall_stack_off;
bool has_tail_call: 1;
bool might_throw: 1;
+ bool might_unwind: 1;
bool tail_call_reachable: 1;
bool has_ld_abs: 1;
bool is_cb: 1;
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 0e38b9575328..eb6651427b12 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -4,6 +4,7 @@
#include <linux/bpf_verifier.h>
#include <linux/filter.h>
#include <linux/bitmap.h>
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -424,7 +425,31 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
if (class == BPF_STX)
bt_set_reg(bt, sreg);
} else if (class == BPF_JMP || class == BPF_JMP32) {
- if (bpf_pseudo_call(insn) || bpf_is_callx(insn)) {
+ if (hist && (hist->flags & INSN_F_UNWIND)) {
+ /*
+ * A bpf_unwind(), a resume or an unwinding global call
+ * left frame hist->frame here, for a landing pad in
+ * this one, or for the main frame's exit, which is
+ * this frame itself when it is the main one. The walk
+ * crosses back into that frame, past any frames
+ * between, which were entered and never returned
+ * from. The pad found r0 unknown and r1-r5 clobbered;
+ * r6-r9 and the stack are this frame's own and stay
+ * marked in its masks until the walk comes back out.
+ */
+ bt_clear_reg(bt, BPF_REG_0);
+ if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
+ verifier_bug(env, "backtracking unwind unexpected regs %x",
+ bt_reg_mask(bt));
+ return -EFAULT;
+ }
+ if (verifier_bug_if(hist->frame < bt->frame, env,
+ "unwind from frame %d to frame %d",
+ hist->frame, bt->frame))
+ return -EFAULT;
+ bt->frame = hist->frame;
+ return 0;
+ } else if (bpf_pseudo_call(insn) || bpf_is_callx(insn)) {
int subprog_insn_idx, subprog = -1;
if (bpf_pseudo_call(insn)) {
@@ -434,6 +459,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
return -EFAULT;
}
+ if (bpf_exc_pad_of_call(env, idx) == subseq_idx) {
+ /*
+ * We came from the landing pad of a call to a
+ * global subprog, branched to from the state
+ * the call returns in: as on its return, no
+ * frame was entered here. The call clobbered
+ * r0-r5; r6-r9 and the stack are the caller's
+ * own and keep going back from here.
+ */
+ bt_clear_reg(bt, BPF_REG_0);
+ if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
+ verifier_bug(env, "landing pad unexpected regs %x",
+ bt_reg_mask(bt));
+ return -EFAULT;
+ }
+ return 0;
+ }
+
/* callx calls static subprogs only */
if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
/* check that jump history doesn't have any
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index 63afbc5fb296..11315d190a57 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -76,6 +76,14 @@ static void mark_subprog_might_throw(struct bpf_verifier_env *env, int off)
subprog->might_throw = true;
}
+static void mark_subprog_might_unwind(struct bpf_verifier_env *env, int off)
+{
+ struct bpf_subprog_info *subprog;
+
+ subprog = bpf_find_containing_subprog(env, off);
+ subprog->might_unwind = true;
+}
+
/* 't' is an index of a call-site.
* 'w' is a callee entry point.
* Eventually this function would be called when env->cfg.insn_state[w] == EXPLORED.
@@ -91,6 +99,7 @@ static void merge_callee_effects(struct bpf_verifier_env *env, int t, int w)
caller->changes_pkt_data |= callee->changes_pkt_data;
caller->might_sleep |= callee->might_sleep;
caller->might_throw |= callee->might_throw;
+ caller->might_unwind |= callee->might_unwind;
}
enum {
@@ -668,6 +677,8 @@ static int visit_insn(int t, struct bpf_verifier_env *env)
mark_subprog_changes_pkt_data(env, t);
if (ret == 0 && bpf_is_throw_kfunc(insn))
mark_subprog_might_throw(env, t);
+ if (ret == 0 && bpf_is_unwind_kfunc(insn))
+ mark_subprog_might_unwind(env, t);
}
return visit_func_call_insn(t, insns, env, insn->src_reg == BPF_PSEUDO_CALL);
@@ -705,6 +716,82 @@ static int visit_insn(int t, struct bpf_verifier_env *env)
}
}
+static bool addr_taken_subprog_might_unwind(struct bpf_verifier_env *env)
+{
+ struct bpf_insn *insns = env->prog->insnsi;
+ struct bpf_subprog_info *callee;
+ struct bpf_func_ptr *ptrs;
+ u32 cnt, j;
+ int i;
+
+ for (i = 0; i < env->prog->len; i++) {
+ if (bpf_pseudo_func(&insns[i])) {
+ callee = bpf_find_containing_subprog(env, i + insns[i].imm + 1);
+ if (callee->might_unwind)
+ return true;
+ continue;
+ }
+ ptrs = insn_func_ptrs(env, i, &cnt);
+ for (j = 0; j < cnt; j++) {
+ callee = bpf_find_containing_subprog(env, ptrs[j].xlated_off);
+ if (callee->might_unwind)
+ return true;
+ }
+ }
+ return false;
+}
+
+/*
+ * merge_callee_effects() does not carry might_unwind to a subprog that calls
+ * through a pointer it was handed. So if any subprog a callx can reach may
+ * unwind, mark every subprog with a callx, and its callers.
+ *
+ * TODO: this is conservative: a subprog with a callx is marked even when
+ * nothing it calls unwinds. The main walk knows each callx's target and could
+ * tell those apart.
+ */
+static void mark_callx_might_unwind(struct bpf_verifier_env *env)
+{
+ struct bpf_insn *insns = env->prog->insnsi;
+ struct bpf_subprog_info *caller, *callee;
+ int i, j, len = env->prog->len;
+ struct bpf_func_ptr *ptrs;
+ bool changed;
+ u32 cnt;
+
+ if (!addr_taken_subprog_might_unwind(env))
+ return;
+
+ for (i = 0; i < len; i++)
+ if (bpf_is_callx(&insns[i]))
+ bpf_find_containing_subprog(env, i)->might_unwind = true;
+
+ do {
+ changed = false;
+ for (i = 0; i < len; i++) {
+ caller = bpf_find_containing_subprog(env, i);
+ if (caller->might_unwind)
+ continue;
+ if (bpf_pseudo_call(&insns[i]) || bpf_pseudo_func(&insns[i])) {
+ callee = bpf_find_containing_subprog(env, i + insns[i].imm + 1);
+ if (!callee->might_unwind)
+ continue;
+ caller->might_unwind = true;
+ changed = true;
+ continue;
+ }
+ ptrs = insn_func_ptrs(env, i, &cnt);
+ for (j = 0; j < cnt; j++) {
+ callee = bpf_find_containing_subprog(env, ptrs[j].xlated_off);
+ if (!callee->might_unwind)
+ continue;
+ caller->might_unwind = true;
+ changed = true;
+ }
+ }
+ } while (changed);
+}
+
/* non-recursive depth-first-search to detect loops in BPF program
* loop == back-edge in directed graph
*/
@@ -795,6 +882,7 @@ int bpf_check_cfg(struct bpf_verifier_env *env)
}
}
ret = 0; /* cfg looks good */
+ mark_callx_might_unwind(env);
env->prog->aux->changes_pkt_data = env->subprog_info[0].changes_pkt_data;
env->prog->aux->might_sleep = env->subprog_info[0].might_sleep;
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 18bf7b660c2f..4ffa5b7ebf0f 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -201,9 +201,13 @@ static int maybe_exit_scc(struct bpf_verifier_env *env, struct bpf_verifier_stat
* c. A checkpoint is reached and matched. Checkpoints are created by
* is_state_visited(), which calls maybe_enter_scc(), which allocates
* bpf_scc_visit instances for checkpoints within SCCs.
- * (c) is the only case that can reach this point.
+ * d. An unwind returns from frame 0, ending the path at a call or at
+ * the unwinding insn, which may be in an SCC as a top-level
+ * BPF_EXIT is not. Like (b), it leaves nothing to exit.
+ * (c) and (d), and a speculative path, are the only ways to reach
+ * this point.
*/
- if (!st->speculative) {
+ if (!st->speculative && !st->unwind_exit) {
verifier_bug(env, "scc exit: no visit info for call chain %s",
format_callchain(env, callchain));
return -EFAULT;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 0cf30220b777..1553c7d5bbdb 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -1744,6 +1744,7 @@ int bpf_copy_verifier_state(struct bpf_verifier_state *dst_state,
return err;
dst_state->speculative = src->speculative;
dst_state->in_sleepable = src->in_sleepable;
+ dst_state->unwind_exit = src->unwind_exit;
dst_state->curframe = src->curframe;
dst_state->branches = src->branches;
dst_state->parent = src->parent;
@@ -11058,6 +11059,9 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
static int process_bpf_exit_full(struct bpf_verifier_env *env,
bool *do_print_state, bool exception_exit);
+static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
+ bool *do_print_state);
+static int unwind_frames(struct bpf_verifier_env *env, bool *do_print_state);
/*
* Call of a static subprog. The callee is verified in the context of
@@ -15233,6 +15237,11 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
if (bpf_is_throw_kfunc(insn))
return process_bpf_exit_full(env, NULL, true);
+ if (bpf_is_unwind_kfunc(insn))
+ return process_bpf_unwind(env, insn_idx_p, NULL);
+ if (bpf_is_unwind_resume_kfunc(insn))
+ return unwind_frames(env, NULL);
+
return 0;
}
@@ -19203,6 +19212,135 @@ enum {
INSN_IDX_UPDATED = 2,
};
+static int unwind_frames(struct bpf_verifier_env *env, bool *do_print_state)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ u32 frameno = state->curframe;
+ struct bpf_func_state *callee, *caller;
+ int err, pad;
+
+ while (state->curframe) {
+ callee = cur_func(env);
+ caller = state->frame[state->curframe - 1];
+ pad = bpf_exc_pad_of_call(env, callee->callsite);
+ /* The caller is at its call now, not at this frame's insn. */
+ state->insn_idx = callee->callsite;
+ /* As on a return: the frame's dynptrs, and their slices, go with it. */
+ err = destroy_dynptrs_in_stack_slots(env, callee, 0,
+ callee->allocated_stack / BPF_REG_SIZE - 1);
+ if (err)
+ return err;
+
+ account_processed_insns(env, callee, caller);
+ free_func_state(callee);
+ state->frame[state->curframe--] = NULL;
+ invalidate_outgoing_stack_args(env, caller);
+ if (pad < 0)
+ continue;
+
+ /*
+ * Mark the unwinding insn with the frame it ran in: backtracking
+ * from the pad comes back to that insn and goes on in that frame.
+ */
+ err = bpf_push_jmp_history(env, state, INSN_F_UNWIND, 0, frameno, NULL, 0);
+ if (err)
+ return err;
+ clear_caller_saved_regs(env, caller->regs);
+ mark_reg_unknown(env, caller->regs, BPF_REG_0);
+ env->insn_idx = pad;
+ if (do_print_state)
+ *do_print_state = true;
+ return INSN_IDX_UPDATED;
+ }
+
+ /*
+ * No pad: the program ends at main's call into the popped frames, or at
+ * the unwinding insn itself if none was popped. Mark the unwinding insn
+ * as above; an unwinding global call needs that when backtracking from
+ * the end reaches it.
+ */
+ err = bpf_push_jmp_history(env, state, INSN_F_UNWIND, 0, frameno, NULL, 0);
+ if (err)
+ return err;
+ /*
+ * Then, in an entry of its own, a jump from the unwinding insn to the
+ * end, so backtracking from the end goes back to the unwinding insn.
+ */
+ env->cur_hist_ent = NULL;
+ env->prev_insn_idx = env->insn_idx;
+ env->insn_idx = state->insn_idx;
+ err = bpf_push_jmp_history(env, state, 0, 0, 0, NULL, 0);
+ if (err)
+ return err;
+ /*
+ * The end may be in a loop: unlike a plain exit, the path can stop
+ * inside an SCC with no checkpoint there.
+ */
+ state->unwind_exit = true;
+
+ /*
+ * The call clobbered r1-r5, and r0 holds the zero the fixups put
+ * there. Mark r0 unknown first: the known-zero helper keeps a NOT_INIT
+ * type.
+ */
+ clear_caller_saved_regs(env, cur_regs(env));
+ mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
+ mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
+ /*
+ * A global subprog verified on its own returns here too. One that
+ * returns in R0:R2 has R2 checked as well, though its caller never
+ * reads either: it resumes at a pad or an epilogue.
+ */
+ if (cur_func(env)->subprogno &&
+ bpf_ret_reg_pair(env, cur_func(env)->subprogno))
+ mark_reg_unknown(env, cur_regs(env), BPF_REG_2);
+
+ return process_bpf_exit_full(env, do_print_state, false);
+}
+
+static int process_global_call_unwind(struct bpf_verifier_env *env,
+ int call_idx, bool *do_print_state)
+{
+ const struct bpf_insn *insn = &env->prog->insnsi[call_idx];
+ int subprog = bpf_find_subprog(env, call_idx + insn->imm + 1);
+ struct bpf_verifier_state *branch;
+ struct bpf_func_state *frame;
+ int pad;
+
+ if (!bpf_subprog_is_global(env, subprog) ||
+ !env->subprog_info[subprog].might_unwind)
+ return 0;
+
+ /* The call returning normally is walked later. */
+ branch = push_stack(env, call_idx + 1, call_idx, false);
+ if (IS_ERR(branch))
+ return PTR_ERR(branch);
+
+ pad = bpf_exc_pad_of_call(env, call_idx);
+ if (pad < 0)
+ return unwind_frames(env, do_print_state);
+ frame = cur_func(env);
+ clear_caller_saved_regs(env, frame->regs);
+ mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ env->insn_idx = pad;
+ *do_print_state = true;
+ return INSN_IDX_UPDATED;
+}
+
+static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
+ bool *do_print_state)
+{
+ struct bpf_func_state *frame = cur_func(env);
+ int pad = bpf_exc_pad_of_call(env, *insn_idx);
+
+ if (pad < 0)
+ return unwind_frames(env, do_print_state);
+ clear_caller_saved_regs(env, frame->regs);
+ mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ *insn_idx = pad;
+ return INSN_IDX_UPDATED;
+}
+
static int process_bpf_exit_full(struct bpf_verifier_env *env,
bool *do_print_state,
bool exception_exit)
@@ -19464,8 +19602,14 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
cur_func(env)->no_stack_arg_load = true;
if (bpf_is_callx(insn))
return check_func_callx(env, insn, &env->insn_idx);
- if (insn->src_reg == BPF_PSEUDO_CALL)
- return check_func_call(env, insn, &env->insn_idx);
+ if (insn->src_reg == BPF_PSEUDO_CALL) {
+ int call_idx = env->insn_idx;
+
+ err = check_func_call(env, insn, &env->insn_idx);
+ if (err)
+ return err;
+ return process_global_call_unwind(env, call_idx, do_print_state);
+ }
if (insn->src_reg == BPF_PSEUDO_KFUNC_CALL)
return check_kfunc_call(env, insn, &env->insn_idx);
return check_helper_call(env, insn, &env->insn_idx);
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 07/23] bpf: Verify an unwind through landing pads and epilogues
2026-10-08 7:50 ` [PATCH bpf-next v9 07/23] bpf: Verify an unwind through landing pads and epilogues Yonghong Song
@ 2026-10-08 8:57 ` bot+bpf-ci
2026-10-08 16:07 ` Yonghong Song
0 siblings, 1 reply; 38+ messages in thread
From: bot+bpf-ci @ 2026-10-08 8:57 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3080 bytes --]
> bpf: Verify an unwind through landing pads and epilogues
TL;DR: In backtrack_insn(), if a landing pad directly follows a global call,
the normal return may take the new pad branch, which doesn't clear R2, so a
precise R2 could trigger verifier_bug() and -EFAULT.
> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
> --- a/kernel/bpf/backtrack.c
> +++ b/kernel/bpf/backtrack.c
[ ... ]
> @@ -434,6 +459,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
> return -EFAULT;
> }
>
> + if (bpf_exc_pad_of_call(env, idx) == subseq_idx) {
> + /*
> + * We came from the landing pad of a call to a
> + * global subprog, branched to from the state
> + * the call returns in: as on its return, no
> + * frame was entered here. The call clobbered
> + * r0-r5; r6-r9 and the stack are the caller's
> + * own and keep going back from here.
> + */
> + bt_clear_reg(bt, BPF_REG_0);
> + if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
> + verifier_bug(env, "landing pad unexpected regs %x",
> + bt_reg_mask(bt));
> + return -EFAULT;
> + }
> + return 0;
> + }
> +
> /* callx calls static subprogs only */
> if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
> /* check that jump history doesn't have any
[ ... ]
Is comparing subseq_idx with the call's pad enough to tell that the walk
came over the pad edge?
bpf_exc_check_info() accepts a record [c, c+1) -> c+1, because its pad check
treats end_off as exclusive. With that layout the pad is the insn right
after the call:
c: call global_subprog
c + 1: pad
In that case the normal return of the global call also has subseq_idx ==
idx + 1 == pad, so it now goes through the new pad branch instead of the
global call branch below it.
The pad branch clears only BPF_REG_0, while the global call branch also
clears R2. For a global subprog that returns in the R0:R2 pair
(bpf_ret_reg_pair()), the normal return path really does define R2 as a
scalar.
If precision on R2 reaches the call insn over that edge, for example from
propagate_precision() when the path is pruned against a checkpoint where R2
is precise, the pad branch finds R2 still set in the mask:
bt_clear_reg(bt, BPF_REG_0);
if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
verifier_bug(env, "landing pad unexpected regs %x", ...);
return -EFAULT;
}
That fires verifier_bug(), which is a WARN_ONCE under CONFIG_DEBUG_KERNEL
plus -EFAULT, rather than handling the edge as a normal global return as
happened before this patch.
Should the pad branch also clear R2, which is NOT_INIT on the real pad path
anyway, or should it apply only when subseq_idx != idx + 1?
I did not find a later patch in the series that touches backtrack_insn()
again to address this.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37747693645
^ permalink raw reply [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 07/23] bpf: Verify an unwind through landing pads and epilogues
2026-10-08 8:57 ` bot+bpf-ci
@ 2026-10-08 16:07 ` Yonghong Song
0 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 16:07 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 10/8/26 10:57 AM, bot+bpf-ci@kernel.org wrote:
>> bpf: Verify an unwind through landing pads and epilogues
> TL;DR: In backtrack_insn(), if a landing pad directly follows a global call,
> the normal return may take the new pad branch, which doesn't clear R2, so a
> precise R2 could trigger verifier_bug() and -EFAULT.
>
>> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
>> --- a/kernel/bpf/backtrack.c
>> +++ b/kernel/bpf/backtrack.c
> [ ... ]
>
>> @@ -434,6 +459,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
>> return -EFAULT;
>> }
>>
>> + if (bpf_exc_pad_of_call(env, idx) == subseq_idx) {
>> + /*
>> + * We came from the landing pad of a call to a
>> + * global subprog, branched to from the state
>> + * the call returns in: as on its return, no
>> + * frame was entered here. The call clobbered
>> + * r0-r5; r6-r9 and the stack are the caller's
>> + * own and keep going back from here.
>> + */
>> + bt_clear_reg(bt, BPF_REG_0);
>> + if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
>> + verifier_bug(env, "landing pad unexpected regs %x",
>> + bt_reg_mask(bt));
>> + return -EFAULT;
>> + }
>> + return 0;
>> + }
>> +
>> /* callx calls static subprogs only */
>> if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
>> /* check that jump history doesn't have any
> [ ... ]
>
> Is comparing subseq_idx with the call's pad enough to tell that the walk
> came over the pad edge?
Not sure. For global subprog, next insn could be a follow-up or could be
a landing pad, as you analysed below.
In such cases, for global subprog, it will go to non-landing pad first
and will return. This will cause problems as global subprog and landing
pad connection got lost. This is a bug and will fix it in the next revision.
Basically, we can remove "if (bpf_exc_pad_of_call(env, idx) == subseq_idx)"
and deal with it under "if (subprog >= 0 && bpf_subprog_is_global(env, subprog))".
>
> bpf_exc_check_info() accepts a record [c, c+1) -> c+1, because its pad check
> treats end_off as exclusive. With that layout the pad is the insn right
> after the call:
>
> c: call global_subprog
> c + 1: pad
>
> In that case the normal return of the global call also has subseq_idx ==
> idx + 1 == pad, so it now goes through the new pad branch instead of the
> global call branch below it.
>
> The pad branch clears only BPF_REG_0, while the global call branch also
> clears R2. For a global subprog that returns in the R0:R2 pair
> (bpf_ret_reg_pair()), the normal return path really does define R2 as a
> scalar.
>
> If precision on R2 reaches the call insn over that edge, for example from
> propagate_precision() when the path is pruned against a checkpoint where R2
> is precise, the pad branch finds R2 still set in the mask:
>
> bt_clear_reg(bt, BPF_REG_0);
> if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
> verifier_bug(env, "landing pad unexpected regs %x", ...);
> return -EFAULT;
> }
>
> That fires verifier_bug(), which is a WARN_ONCE under CONFIG_DEBUG_KERNEL
> plus -EFAULT, rather than handling the edge as a normal global return as
> happened before this patch.
>
> Should the pad branch also clear R2, which is NOT_INIT on the real pad path
> anyway, or should it apply only when subseq_idx != idx + 1?
>
> I did not find a later patch in the series that touches backtrack_insn()
> again to address this.
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37747693645
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH bpf-next v9 08/23] bpf: Refuse a landing pad that does not resume
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (6 preceding siblings ...)
2026-10-08 7:50 ` [PATCH bpf-next v9 07/23] bpf: Verify an unwind through landing pads and epilogues Yonghong Song
@ 2026-10-08 7:50 ` Yonghong Song
2026-10-08 8:57 ` bot+bpf-ci
2026-10-08 7:50 ` [PATCH bpf-next v9 09/23] bpf: Do not use a private stack for a program that can unwind Yonghong Song
` (14 subsequent siblings)
22 siblings, 1 reply; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:50 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A cleanup pad runs drop glue and calls bpf_unwind_resume(), so its frame
returns and the unwind goes on. A catch pad carries on in its frame
instead. Only the first is supported: bpf_unwind() rewrites every frame's
return address in one pass, so a catch pad's caller would still resume
at a pad.
Nothing in the record says which kind a pad is, but the code does: a
cleanup pad reaches _Unwind_Resume, a catch pad a return. So the unwind
marks the frame whose pad it enters, and refuses in the marked code:
refused where
--------------------------------------- ----------------------------
an exit: how a catch pad ends the pad's own frame
a tail call, a BPF_LD_[ABS|IND]: it the pad's own frame
leaves through an exit on a failed
load
an indirect jump: nothing a compiler the pad's own frame
frontend emits in a pad needs one
a bpf_unwind(), a call to a global the pad and its callees
subprog that might_unwind
a bpf_unwind_resume() outside a pad any program
The first three are refused only in the pad's own frame: a subprog the
pad calls may do them and still come back.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 1 +
kernel/bpf/exception.c | 68 ++++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 2 ++
kernel/bpf/states.c | 3 ++
kernel/bpf/verifier.c | 24 ++++++++++++-
5 files changed, 97 insertions(+), 1 deletion(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index d3740ad20235..6449e4babc60 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -339,6 +339,7 @@ struct bpf_func_state {
bool in_async_callback_fn;
bool in_exception_callback_fn;
bool no_stack_arg_load;
+ bool in_pad;
/* For callback calling functions that limit number of possible
* callback executions (e.g. bpf_loop) keeps track of current
* simulated iteration number.
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index d85ca358b463..6155f5d73420 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -141,6 +141,15 @@ int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
+int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog)
+{
+ if (!env->subprog_info[subprog].might_unwind)
+ return 0;
+
+ verbose(env, "subprog %d may unwind and is used as a callback\n", subprog);
+ return -EINVAL;
+}
+
static int reject_throw(struct bpf_verifier_env *env)
{
u32 i;
@@ -223,6 +232,65 @@ bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
insn->imm == bpf_unwind_resume_id[0];
}
+static bool unwinding(const struct bpf_verifier_state *state)
+{
+ u32 i;
+
+ for (i = 0; i <= state->curframe; i++)
+ if (state->frame[i]->in_pad)
+ return true;
+ return false;
+}
+
+int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn)
+{
+ bool in_pad = cur_func(env)->in_pad;
+ u32 i = env->insn_idx;
+ const char *why = NULL;
+
+ if (unwinding(env->cur_state)) {
+ if (bpf_is_unwind_kfunc(insn)) {
+ verbose(env, "insn %u starts a second unwind while one is in flight\n", i);
+ return -EINVAL;
+ }
+ if (bpf_pseudo_call(insn)) {
+ int subprog = bpf_find_subprog(env, i + insn->imm + 1);
+
+ if (subprog >= 0 && bpf_subprog_is_global(env, subprog) &&
+ env->subprog_info[subprog].might_unwind) {
+ verbose(env,
+ "insn %u calls global subprog %d, which can unwind while an unwind is in flight\n",
+ i, subprog);
+ return -EINVAL;
+ }
+ }
+ }
+
+ if (!in_pad)
+ return 0;
+
+ if (insn->code == (BPF_JMP | BPF_EXIT)) {
+ verbose(env,
+ "exit at insn %u ends a landing pad: a catch pad is not supported yet, only cleanup pads that resume\n",
+ i);
+ return -EINVAL;
+ }
+ if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call)
+ why = "is a tail call, which replaces the frame";
+ else if (BPF_CLASS(insn->code) == BPF_LD &&
+ (BPF_MODE(insn->code) == BPF_ABS || BPF_MODE(insn->code) == BPF_IND))
+ why = "is a BPF_LD_[ABS|IND], which can leave through the epilogue";
+ else if (insn->code == (BPF_JMP | BPF_JA | BPF_X) ||
+ insn->code == (BPF_JMP32 | BPF_JA | BPF_X))
+ why = "is an indirect jump, which nothing a compiler frontend emits in a pad needs";
+
+ if (!why)
+ return 0;
+
+ verbose(env, "insn %u %s, and is in a landing pad\n", i, why);
+ return -EINVAL;
+}
+
int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
{
u32 pad = env->insn_aux_data[idx].cleanup_pad;
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index 9b75e556c82e..1d7fc795619f 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -17,5 +17,7 @@ int bpf_exc_check_prog(struct bpf_verifier_env *env);
int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
+int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog);
+int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn);
#endif /* __BPF_EXCEPTION_H */
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 4ffa5b7ebf0f..eaf95d6e81bc 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -956,6 +956,9 @@ static bool func_states_equal(struct bpf_verifier_env *env, struct bpf_func_stat
if (!old->no_stack_arg_load && cur->no_stack_arg_load)
return false;
+ if (old->in_pad != cur->in_pad)
+ return false;
+
for (i = 0; i < MAX_BPF_REG; i++)
if (((1 << i) & live_regs) &&
!regsafe(env, &old->regs[i], &cur->regs[i],
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 1553c7d5bbdb..59b275bfc949 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10997,6 +10997,10 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
* callbacks
*/
env->subprog_info[subprog].is_cb = true;
+ err = bpf_exc_check_callback(env, subprog);
+ if (err)
+ return err;
+
if (bpf_pseudo_kfunc_call(insn) &&
!is_callback_calling_kfunc(insn->imm)) {
verifier_bug(env, "kfunc %s#%d not marked as callback-calling",
@@ -15239,8 +15243,13 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
if (bpf_is_unwind_kfunc(insn))
return process_bpf_unwind(env, insn_idx_p, NULL);
- if (bpf_is_unwind_resume_kfunc(insn))
+ if (bpf_is_unwind_resume_kfunc(insn)) {
+ if (!cur_func(env)->in_pad) {
+ verbose(env, "resume at insn %d is not in a landing pad\n", insn_idx);
+ return -EINVAL;
+ }
return unwind_frames(env, NULL);
+ }
return 0;
}
@@ -19247,6 +19256,7 @@ static int unwind_frames(struct bpf_verifier_env *env, bool *do_print_state)
return err;
clear_caller_saved_regs(env, caller->regs);
mark_reg_unknown(env, caller->regs, BPF_REG_0);
+ caller->in_pad = true;
env->insn_idx = pad;
if (do_print_state)
*do_print_state = true;
@@ -19322,6 +19332,7 @@ static int process_global_call_unwind(struct bpf_verifier_env *env,
frame = cur_func(env);
clear_caller_saved_regs(env, frame->regs);
mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ frame->in_pad = true;
env->insn_idx = pad;
*do_print_state = true;
return INSN_IDX_UPDATED;
@@ -19337,6 +19348,7 @@ static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
return unwind_frames(env, do_print_state);
clear_caller_saved_regs(env, frame->regs);
mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ frame->in_pad = true;
*insn_idx = pad;
return INSN_IDX_UPDATED;
}
@@ -19708,6 +19720,16 @@ static int do_check(struct bpf_verifier_env *env)
}
}
+ if (unlikely(env->cleanup_info_cnt)) {
+ err = bpf_exc_check_insn(env, insn);
+ if (error_recoverable_with_nospec(err) && state->speculative) {
+ insn_aux->nospec = true;
+ goto process_bpf_exit;
+ }
+ if (err)
+ return err;
+ }
+
if (bpf_is_prune_point(env, env->insn_idx)) {
err = bpf_is_state_visited(env, env->insn_idx);
if (err < 0)
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 08/23] bpf: Refuse a landing pad that does not resume
2026-10-08 7:50 ` [PATCH bpf-next v9 08/23] bpf: Refuse a landing pad that does not resume Yonghong Song
@ 2026-10-08 8:57 ` bot+bpf-ci
2026-10-08 16:11 ` Yonghong Song
0 siblings, 1 reply; 38+ messages in thread
From: bot+bpf-ci @ 2026-10-08 8:57 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 2773 bytes --]
> bpf: Refuse a landing pad that does not resume
TL;DR: Commit message only: the patch also makes push_callback_call() reject
callback subprogs that may unwind, even without a cleanup table, but the
message doesn't mention it. Mention it or split it out?
> bpf: Refuse a landing pad that does not resume
>
> A cleanup pad runs drop glue and calls bpf_unwind_resume(), so its frame
> returns and the unwind goes on. A catch pad carries on in its frame
> instead. Only the first is supported: bpf_unwind() rewrites every frame's
> return address in one pass, so a catch pad's caller would still resume
> at a pad.
>
> Nothing in the record says which kind a pad is, but the code does: a
> cleanup pad reaches _Unwind_Resume, a catch pad a return. So the unwind
> marks the frame whose pad it enters, and refuses in the marked code:
>
> refused where
> --------------------------------------- ----------------------------
> an exit: how a catch pad ends the pad's own frame
> a tail call, a BPF_LD_[ABS|IND]: it the pad's own frame
> leaves through an exit on a failed
> load
> an indirect jump: nothing a compiler the pad's own frame
> frontend emits in a pad needs one
> a bpf_unwind(), a call to a global the pad and its callees
> subprog that might_unwind
> a bpf_unwind_resume() outside a pad any program
>
> The first three are refused only in the pad's own frame: a subprog the
> pad calls may do them and still come back.
>
> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
The patch also adds a refusal that this table does not list.
push_callback_call() now calls bpf_exc_check_callback(), which returns
-EINVAL for any callback subprog with might_unwind set:
verbose(env, "subprog %d may unwind and is used as a callback\n", subprog);
This check is not about landing pads, and it is not gated on
env->cleanup_info_cnt. might_unwind is set for any bpf_unwind() call, so
it also applies to programs with no cleanup table. Later in the series,
bpf_prog_may_unwind() makes that case reachable.
The restriction looks needed, since bpf_unwind_rewrite() stops walking at the
helper frame between a callback and its caller, while the verifier would model
the unwind continuing into the caller's frames.
Could the commit message mention this restriction and the reason for it, or
could the callback check go into its own patch?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37747693645
^ permalink raw reply [flat|nested] 38+ messages in thread
* Re: [PATCH bpf-next v9 08/23] bpf: Refuse a landing pad that does not resume
2026-10-08 8:57 ` bot+bpf-ci
@ 2026-10-08 16:11 ` Yonghong Song
0 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 16:11 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 10/8/26 10:57 AM, bot+bpf-ci@kernel.org wrote:
>> bpf: Refuse a landing pad that does not resume
> TL;DR: Commit message only: the patch also makes push_callback_call() reject
> callback subprogs that may unwind, even without a cleanup table, but the
> message doesn't mention it. Mention it or split it out?
>
>> bpf: Refuse a landing pad that does not resume
>>
>> A cleanup pad runs drop glue and calls bpf_unwind_resume(), so its frame
>> returns and the unwind goes on. A catch pad carries on in its frame
>> instead. Only the first is supported: bpf_unwind() rewrites every frame's
>> return address in one pass, so a catch pad's caller would still resume
>> at a pad.
>>
>> Nothing in the record says which kind a pad is, but the code does: a
>> cleanup pad reaches _Unwind_Resume, a catch pad a return. So the unwind
>> marks the frame whose pad it enters, and refuses in the marked code:
>>
>> refused where
>> --------------------------------------- ----------------------------
>> an exit: how a catch pad ends the pad's own frame
>> a tail call, a BPF_LD_[ABS|IND]: it the pad's own frame
>> leaves through an exit on a failed
>> load
>> an indirect jump: nothing a compiler the pad's own frame
>> frontend emits in a pad needs one
>> a bpf_unwind(), a call to a global the pad and its callees
>> subprog that might_unwind
>> a bpf_unwind_resume() outside a pad any program
>>
>> The first three are refused only in the pad's own frame: a subprog the
>> pad calls may do them and still come back.
>>
>> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
> The patch also adds a refusal that this table does not list.
> push_callback_call() now calls bpf_exc_check_callback(), which returns
> -EINVAL for any callback subprog with might_unwind set:
>
> verbose(env, "subprog %d may unwind and is used as a callback\n", subprog);
>
> This check is not about landing pads, and it is not gated on
> env->cleanup_info_cnt. might_unwind is set for any bpf_unwind() call, so
> it also applies to programs with no cleanup table. Later in the series,
> bpf_prog_may_unwind() makes that case reachable.
>
> The restriction looks needed, since bpf_unwind_rewrite() stops walking at the
> helper frame between a callback and its caller, while the verifier would model
> the unwind continuing into the caller's frames.
>
> Could the commit message mention this restriction and the reason for it, or
> could the callback check go into its own patch?
This is not a bug as mentioned in the above. It indeed may worthwhile to
have a separate commit for callback checking since this commit is about
'refuse a landing pad that does not resume'...
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37747693645
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH bpf-next v9 09/23] bpf: Do not use a private stack for a program that can unwind
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (7 preceding siblings ...)
2026-10-08 7:50 ` [PATCH bpf-next v9 08/23] bpf: Refuse a landing pad that does not resume Yonghong Song
@ 2026-10-08 7:50 ` Yonghong Song
2026-10-08 7:50 ` [PATCH bpf-next v9 10/23] bpf: Prepare JITed programs for dispatching cleanup pads Yonghong Song
` (13 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:50 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A private stack keeps its frame pointer in %r9 on x86-64, saved and
restored around every call. An unwind skips the restore: a frame resumed
at its pad then addresses its stack through a stale pointer, and one sent
to its epilogue pops its callee-saved registers one slot off.
So force NO_PRIV_STACK for a program that can unwind, with a table or
not, on every arch for now.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/exception.c | 10 ++++++++++
kernel/bpf/exception.h | 1 +
kernel/bpf/verifier.c | 11 +++++++++++
3 files changed, 22 insertions(+)
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index 6155f5d73420..035eb4d87588 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -183,6 +183,16 @@ static void mark_call_sites(struct bpf_verifier_env *env)
}
}
+bool bpf_prog_may_unwind(const struct bpf_verifier_env *env)
+{
+ u32 i;
+
+ for (i = 0; i < env->subprog_cnt; i++)
+ if (env->subprog_info[i].might_unwind)
+ return true;
+ return false;
+}
+
int bpf_exc_check_prog(struct bpf_verifier_env *env)
{
int err;
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index 1d7fc795619f..dea45c7e4925 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -14,6 +14,7 @@ int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
bpfptr_t uattr);
void bpf_exc_prepare(struct bpf_verifier_env *env);
int bpf_exc_check_prog(struct bpf_verifier_env *env);
+bool bpf_prog_may_unwind(const struct bpf_verifier_env *env);
int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 59b275bfc949..585be741c689 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -5801,6 +5801,17 @@ static int check_max_stack_depth(struct bpf_verifier_env *env)
}
}
+ /*
+ * A private stack keeps its frame pointer in %r9 on x86-64, restored
+ * by a pop after the call that an unwind skips. A frame resumed at a
+ * pad then addresses its stack through a stale pointer, and a frame
+ * sent to its epilogue instead pops its callee-saved registers one
+ * slot off. Refuse a private stack for any program that can unwind,
+ * on every arch for now.
+ */
+ if (env->cleanup_info_cnt || bpf_prog_may_unwind(env))
+ priv_stack_mode = NO_PRIV_STACK;
+
if (priv_stack_mode == PRIV_STACK_UNKNOWN)
priv_stack_mode = bpf_enable_priv_stack(env->prog);
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 10/23] bpf: Prepare JITed programs for dispatching cleanup pads
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (8 preceding siblings ...)
2026-10-08 7:50 ` [PATCH bpf-next v9 09/23] bpf: Do not use a private stack for a program that can unwind Yonghong Song
@ 2026-10-08 7:50 ` Yonghong Song
2026-10-08 7:50 ` [PATCH bpf-next v9 11/23] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
` (12 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:50 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
The next patch makes bpf_unwind() rewrite the saved return address of
each frame above its caller: to the pad covering the call, else to the
frame's epilogue. Every frame then just returns, its epilogue restoring
its caller's r6-r9. This patch adds what that relies on:
- Each bpf_unwind() call is followed by 'r0 = 0' and a jump to its pad,
or an exit: the walk leaves the calling frame alone.
- A pad's resume lowers to 'r0 = 0; exit'.
- Each JITed function gets a table of its covered calls as native
address ranges, looked up by bpf_exc_pad_for_ip(). The main
function's table, and its epilogue_ip, go to the outer program.
- Every function the walk can pass needs an epilogue, which x86 emits
at a function's first exit. keep_subprog_exits() keeps one through
dead code removal, and refuses a function with none.
- arch_bpf_stack_walk_ra() walks the stack handing out each
return-address slot. Where an arch has only the weak stub, nothing is
dispatched, so any program that can unwind, table or not, is now held
to bpf_exc_check_prog().
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf.h | 30 ++++++++
include/linux/bpf_verifier.h | 1 +
include/linux/filter.h | 2 +
kernel/bpf/core.c | 7 ++
kernel/bpf/exception.c | 77 ++++++++++++++++++++
kernel/bpf/exception.h | 6 ++
kernel/bpf/fixups.c | 135 ++++++++++++++++++++++++++++++++++-
kernel/bpf/verifier.c | 5 +-
8 files changed, 261 insertions(+), 2 deletions(-)
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 54144372281c..7f23f4efde01 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -1841,6 +1841,34 @@ enum bpf_sig_keyring {
BPF_SIG_KEYRING_BPF,
};
+/* One cleanup region of a JITed (sub)program. */
+struct bpf_cleanup_range {
+ u64 begin;
+ u64 end;
+ u64 pad;
+};
+
+struct bpf_exception_info {
+ struct bpf_cleanup_info *info;
+ struct bpf_cleanup_range *ranges;
+ u32 nr_info;
+ u32 nr_ranges;
+};
+
+#ifdef CONFIG_BPF_SYSCALL
+void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image);
+void bpf_exc_free_info(struct bpf_prog_aux *aux);
+#else
+
+static inline void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+}
+
+static inline void bpf_exc_free_info(struct bpf_prog_aux *aux)
+{
+}
+#endif
+
struct bpf_prog_aux {
atomic64_t refcnt;
u32 used_map_cnt;
@@ -1922,6 +1950,8 @@ struct bpf_prog_aux {
u64 (*bpf_exception_cb)(u64 cookie, u64 sp, u64 bp, u64, u64);
u16 stack_arg_sp_adjust;
u16 freplace_link_cnt; /* counts freplace links extending this prog */
+ struct bpf_exception_info *exc;
+ u64 epilogue_ip; /* native address of this (sub)program's epilogue */
#ifdef CONFIG_SECURITY
void *security;
#endif
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 6449e4babc60..8f67126719fe 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1853,6 +1853,7 @@ int bpf_opt_subreg_zext_lo32_rnd_hi32(struct bpf_verifier_env *env, const union
int bpf_convert_ctx_accesses(struct bpf_verifier_env *env);
int bpf_jit_subprogs(struct bpf_verifier_env *env);
int bpf_fixup_call_args(struct bpf_verifier_env *env);
+int bpf_exc_patch_unwind_calls(struct bpf_verifier_env *env);
int bpf_do_misc_fixups(struct bpf_verifier_env *env);
int bpf_insn_def32(struct bpf_prog *prog, struct bpf_insn *insn);
diff --git a/include/linux/filter.h b/include/linux/filter.h
index c8ca863e352a..a61a78675d11 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1291,6 +1291,8 @@ u32 bpf_jit_plan_arg_moves(const struct bpf_jit_arg_abi *abi,
struct bpf_jit_arg_move *moves);
u64 bpf_arch_uaddress_limit(void);
void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp), void *cookie);
+void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
+ void *cookie);
u64 arch_bpf_timed_may_goto(void);
u64 bpf_check_timed_may_goto(struct bpf_timed_may_goto *);
bool bpf_helper_changes_pkt_data(enum bpf_func_id func_id);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index 078bcccaf242..9ed9da581ed5 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -293,6 +293,7 @@ void __bpf_prog_free(struct bpf_prog *fp)
mutex_destroy(&fp->aux->dst_mutex);
mutex_destroy(&fp->aux->st_ops_assoc_mutex);
kfree(fp->aux->poke_tab);
+ bpf_exc_free_info(fp->aux);
kfree(fp->aux);
}
free_percpu(fp->stats);
@@ -3528,6 +3529,12 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
{
}
+void __weak arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp,
+ u64 *ra),
+ void *cookie)
+{
+}
+
bool __weak bpf_jit_supports_cleanup_pads(void)
{
return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index 035eb4d87588..9f7ed78a65d0 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -307,3 +307,80 @@ int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
return pad ? (int)pad - 1 : -1;
}
+
+/*
+ * The record covering @ip, which is a return address: the call it belongs to
+ * is the instruction before it, so a range matches on begin < ip <= end.
+ */
+const struct bpf_cleanup_range *bpf_exc_pad_for_ip(const struct bpf_prog *prog, u64 ip)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+ u32 l = 0, r = exc ? exc->nr_ranges : 0;
+
+ while (l < r) {
+ u32 m = l + (r - l) / 2;
+ const struct bpf_cleanup_range *rec = &exc->ranges[m];
+
+ if (ip <= rec->begin)
+ r = m;
+ else if (ip > rec->end)
+ l = m + 1;
+ else
+ return rec;
+ }
+ return NULL;
+}
+
+int bpf_exc_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt)
+{
+ struct bpf_cleanup_range *ranges;
+ struct bpf_exception_info *exc;
+
+ exc = kzalloc_obj(struct bpf_exception_info, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ ranges = kvcalloc(cnt, sizeof(*ranges), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!exc || !ranges) {
+ kfree(exc);
+ kvfree(ranges);
+ kvfree(recs);
+ return -ENOMEM;
+ }
+
+ exc->info = recs;
+ exc->nr_info = cnt;
+ exc->ranges = ranges;
+ /* Withheld until the JIT has filled the table in. */
+ exc->nr_ranges = 0;
+ aux->exc = exc;
+ return 0;
+}
+
+void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+ struct bpf_exception_info *exc = prog->aux->exc;
+ u32 i, n;
+
+ if (!exc)
+ return;
+
+ n = exc->nr_info;
+ for (i = 0; i < n; i++) {
+ const struct bpf_cleanup_info *rec = &exc->info[i];
+
+ exc->ranges[i].begin = (u64)(long)image + addrs[rec->begin_off];
+ exc->ranges[i].end = (u64)(long)image + addrs[rec->end_off];
+ exc->ranges[i].pad = (u64)(long)image + addrs[rec->landing_pad_off];
+ }
+ exc->nr_ranges = n;
+}
+
+void bpf_exc_free_info(struct bpf_prog_aux *aux)
+{
+ struct bpf_exception_info *exc = aux->exc;
+
+ if (!exc)
+ return;
+ kvfree(exc->ranges);
+ kvfree(exc->info);
+ kfree(exc);
+ aux->exc = NULL;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index dea45c7e4925..b55ce5c2b9c6 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -9,6 +9,10 @@
union bpf_attr;
struct bpf_verifier_env;
struct bpf_insn;
+struct bpf_cleanup_info;
+struct bpf_cleanup_range;
+struct bpf_prog;
+struct bpf_prog_aux;
int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
bpfptr_t uattr);
@@ -20,5 +24,7 @@ bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog);
int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn);
+int bpf_exc_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt);
+const struct bpf_cleanup_range *bpf_exc_pad_for_ip(const struct bpf_prog *prog, u64 ip);
#endif /* __BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 64baee1a37e6..1cc025c45ecd 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -9,6 +9,7 @@
#include <linux/sched/signal.h>
#include <net/xdp.h>
#include "disasm.h"
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -715,6 +716,32 @@ static void keep_funcs_with_addr_taken(struct bpf_verifier_env *env)
}
}
+static int keep_subprog_exits(struct bpf_verifier_env *env)
+{
+ u32 i, j;
+
+ for (i = 0; i < env->subprog_cnt; i++) {
+ bool found = false;
+ u32 start;
+
+ if (!env->subprog_info[i].might_unwind)
+ continue;
+ start = env->subprog_info[i].start;
+ for (j = env->subprog_info[i + 1].start; j-- > start; ) {
+ if (env->prog->insnsi[j].code != (BPF_JMP | BPF_EXIT))
+ continue;
+ env->insn_aux_data[j].seen = env->pass_cnt;
+ found = true;
+ break;
+ }
+ if (!found) {
+ verbose(env, "subprog %u can be unwound through but has no exit\n", i);
+ return -EINVAL;
+ }
+ }
+ return 0;
+}
+
int bpf_opt_remove_dead_code(struct bpf_verifier_env *env)
{
struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
@@ -722,6 +749,9 @@ int bpf_opt_remove_dead_code(struct bpf_verifier_env *env)
int i, err;
keep_funcs_with_addr_taken(env);
+ err = keep_subprog_exits(env);
+ if (err)
+ return err;
for (i = 0; i < insn_cnt; i++) {
int j;
@@ -1262,6 +1292,40 @@ static int resolve_func_ptrs(struct bpf_verifier_env *env, struct bpf_prog *prog
return 0;
}
+static int exc_info_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
+ u32 start, u32 end)
+{
+ struct bpf_cleanup_info *recs;
+ u32 i, cnt = 0;
+
+ if (!env->cleanup_info_cnt)
+ return 0;
+
+ for (i = start; i < end; i++) {
+ if (env->insn_aux_data[i].cleanup_pad)
+ cnt++;
+ }
+ if (!cnt)
+ return 0;
+
+ recs = kvmalloc_array(cnt, sizeof(*recs), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!recs)
+ return -ENOMEM;
+
+ for (i = start, cnt = 0; i < end; i++) {
+ u32 pad = env->insn_aux_data[i].cleanup_pad;
+
+ if (!pad)
+ continue;
+ pad--;
+ recs[cnt].begin_off = i - start;
+ recs[cnt].end_off = i - start + 1;
+ recs[cnt].landing_pad_off = pad - start;
+ cnt++;
+ }
+ return bpf_exc_attach_info(sub->aux, recs, cnt);
+}
+
static int jit_subprogs(struct bpf_verifier_env *env)
{
struct bpf_prog *prog = env->prog, **func, *tmp;
@@ -1269,7 +1333,7 @@ static int jit_subprogs(struct bpf_verifier_env *env)
struct bpf_map *map_ptr;
struct bpf_insn *insn;
void *old_bpf_func;
- int err, num_exentries;
+ int err, exc_err, num_exentries;
for (i = 0, insn = prog->insnsi; i < prog->len; i++, insn++) {
if (!bpf_pseudo_func(insn) && !bpf_pseudo_call(insn))
@@ -1399,6 +1463,11 @@ static int jit_subprogs(struct bpf_verifier_env *env)
func[i]->aux->token = prog->aux->token;
if (!i)
func[i]->aux->exception_boundary = env->seen_exception;
+ exc_err = exc_info_for_subprog(env, func[i], subprog_start, subprog_end);
+ if (exc_err) {
+ err = exc_err;
+ goto out_free;
+ }
func[i] = bpf_int_jit_compile(env, func[i]);
if (!func[i]->jited) {
err = -ENOTSUPP;
@@ -1508,6 +1577,9 @@ static int jit_subprogs(struct bpf_verifier_env *env)
prog->aux->bpf_exception_cb = (void *)func[env->exception_callback_subprog]->bpf_func;
prog->aux->exception_boundary = func[0]->aux->exception_boundary;
prog->aux->stack_arg_sp_adjust = func[0]->aux->stack_arg_sp_adjust;
+ prog->aux->exc = func[0]->aux->exc;
+ func[0]->aux->exc = NULL;
+ prog->aux->epilogue_ip = func[0]->aux->epilogue_ip;
bpf_prog_jit_attempt_done(prog);
return 0;
out_free:
@@ -1733,6 +1805,48 @@ static int may_goto_expand(struct bpf_insn *insn_buf, int off, int stack_off,
return cnt + tail_cnt;
}
+/*
+ * Follow each bpf_unwind() call with 'r0 = 0; exit', or with
+ * 'r0 = 0; goto pad' where a record covers the call.
+ */
+int bpf_exc_patch_unwind_calls(struct bpf_verifier_env *env)
+{
+ int insn_cnt = env->prog->len;
+ struct bpf_insn insn_buf[3];
+ struct bpf_prog *new_prog;
+ int i, off, delta = 0;
+
+ if (!bpf_prog_may_unwind(env))
+ return 0;
+
+ for (i = 0; i < insn_cnt; i++) {
+ int call = i + delta, pad;
+ struct bpf_insn *insn = env->prog->insnsi + call;
+
+ if (!bpf_is_unwind_kfunc(insn))
+ continue;
+
+ insn_buf[0] = *insn;
+ insn_buf[1] = BPF_MOV64_IMM(BPF_REG_0, 0);
+ insn_buf[2] = BPF_EXIT_INSN();
+ pad = bpf_exc_pad_of_call(env, call);
+ if (pad >= 0) {
+ /* The goto is at call + 2; a pad after the call moves down by 2. */
+ if (pad > call)
+ pad += 2;
+ off = pad - (call + 3);
+ insn_buf[2] = off == (s16)off ? BPF_JMP_A(off) : BPF_JMP32_A(off);
+ }
+
+ new_prog = bpf_patch_insn_data(env, call, insn_buf, 3);
+ if (!new_prog)
+ return -ENOMEM;
+ delta += 2;
+ env->prog = new_prog;
+ }
+ return 0;
+}
+
/* Do various post-verification rewrites in a single program pass.
* These rewrites simplify JIT and interpreter implementations.
*/
@@ -2125,6 +2239,25 @@ int bpf_do_misc_fixups(struct bpf_verifier_env *env)
goto next_insn;
if (insn->src_reg == BPF_PSEUDO_CALL)
goto next_insn;
+ if (bpf_is_unwind_resume_kfunc(insn)) {
+ /*
+ * A pad's resume is just the frame returning, to
+ * where bpf_unwind() pointed its return address: its
+ * caller's pad or epilogue, or the kernel from the main
+ * program. The verifier checked this exit with r0 a
+ * known zero, so return zero.
+ */
+ insn_buf[0] = BPF_MOV64_IMM(BPF_REG_0, 0);
+ insn_buf[1] = BPF_EXIT_INSN();
+ cnt = 2;
+ new_prog = bpf_patch_insn_data(env, i + delta, insn_buf, cnt);
+ if (!new_prog)
+ return -ENOMEM;
+ delta += cnt - 1;
+ env->prog = prog = new_prog;
+ insn = new_prog->insnsi + i + delta;
+ goto next_insn;
+ }
if (insn->src_reg == BPF_PSEUDO_KFUNC_CALL) {
ret = bpf_fixup_kfunc_call(env, insn, insn_buf, i + delta, &cnt);
if (ret)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 585be741c689..668d811d4e4c 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22928,7 +22928,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
if (ret)
goto skip_full_check;
- if (env->cleanup_info_cnt) {
+ if (env->cleanup_info_cnt || bpf_prog_may_unwind(env)) {
ret = bpf_exc_check_prog(env);
if (ret)
goto skip_full_check;
@@ -23005,6 +23005,9 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
/* program is valid, convert *(u32*)(ctx + off) accesses */
ret = bpf_convert_ctx_accesses(env);
+ if (ret == 0)
+ ret = bpf_exc_patch_unwind_calls(env);
+
if (ret == 0)
ret = bpf_do_misc_fixups(env);
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 11/23] bpf: Dispatch cleanup pads by rewriting return addresses
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (9 preceding siblings ...)
2026-10-08 7:50 ` [PATCH bpf-next v9 10/23] bpf: Prepare JITed programs for dispatching cleanup pads Yonghong Song
@ 2026-10-08 7:50 ` Yonghong Song
2026-10-08 7:51 ` [PATCH bpf-next v9 12/23] bpf: Refuse a trampoline that calls a subprog that can unwind Yonghong Song
` (11 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:50 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
bpf_unwind() now walks the BPF frames through arch_bpf_stack_walk_ra()
and rewrites the saved return address of each frame above the one that
called it: to the pad where a record covers the call, else to the
frame's epilogue. The walk stops after the main function. For the
verifier patch's example, main -> A -> B -> C with only A's call to B
covered:
slot points into rewritten to
---------------------- ------------------------ ---------------------
bpf_unwind()'s return C, after its unwind call left alone
C's return B, after 'call C' B's epilogue
B's return A, after 'call B' P, A's pad
A's return main, after 'call A' main's epilogue
main's return the kernel left alone
Each frame then just returns: C through the 'r0 = 0; exit' after its
bpf_unwind(), B and main through their epilogues, A through its pad P
and P's resume.
The walk finds the frames in the calling program, not with
bpf_prog_ksym_find(). A running instance holds no reference to its
program, so user space can drop the last one, by closing the program's
fds and links, while the instance still runs: the kallsyms entries go at
once, and only freeing the program waits for the RCU or RCU Tasks Trace
grace period. An unwind in between, say in a sleepable program blocked
in bpf_copy_from_user(), would find none of its frames, so no pad would
run. So bpf_unwind() takes the program's aux as a KF_IMPLICIT_ARGS
argument, which keeps its BTF prototype void(void), and matches each
return address against aux->func[], or the program itself. To get the
implicit argument loaded, adjust_insn_aux_data() moves arg_prog, as it
moves cleanup_pad, to the slot the original instruction kept: the unwind
call stays first in its patch.
Both kfuncs become callable here. bpf_unwind() is notrace and NOKPROBE:
a tracer hooking its return would leave a trampoline in the slot of the
frame that called it, which the walk cannot rewrite.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/fixups.c | 2 ++
kernel/bpf/helpers.c | 67 +++++++++++++++++++++++++++++++++++++++++++-
2 files changed, 68 insertions(+), 1 deletion(-)
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 1cc025c45ecd..c29e14ffc475 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -247,6 +247,8 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
data[off + cnt - 1].non_stack_access = false;
data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
data[off + cnt - 1].cleanup_pad = 0;
+ data[i].arg_prog = data[off + cnt - 1].arg_prog;
+ data[off + cnt - 1].arg_prog = 0;
} else if (bpf_is_mem_insn(insn + i)) {
data[i].non_stack_access = true;
}
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 4eccd6742eba..c9efb789f36b 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -29,8 +29,10 @@
#include <linux/task_work.h>
#include <linux/irq_work.h>
#include <linux/buildid.h>
+#include <linux/kprobes.h>
#include "../../lib/kstrtox.h"
+#include "exception.h"
/* If kernel subsystem is allowing eBPF programs to call this function,
* inside its own verifier_ops->get_func_proto() callback it should return
@@ -3424,9 +3426,70 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
return false;
}
-__bpf_kfunc void bpf_unwind(void)
+struct bpf_unwind_ctx {
+ const struct bpf_prog_aux *aux;
+ u32 cnt;
+};
+
+static bool bpf_unwind_ip_in(const struct bpf_prog *prog, u64 ip)
+{
+ u64 start = (u64)(long)prog->bpf_func;
+
+ return ip > start && ip <= start + prog->jited_len;
+}
+
+/*
+ * The function of the calling program that @ip returns into, looked up in the
+ * program itself: its kallsyms entries can go while an instance still runs.
+ * The main function is the outer program, which holds its table and epilogue.
+ */
+static struct bpf_prog *bpf_unwind_find_prog(const struct bpf_prog_aux *aux, u64 ip)
+{
+ u32 i;
+
+ for (i = 1; i < aux->func_cnt; i++)
+ if (bpf_unwind_ip_in(aux->func[i], ip))
+ return aux->func[i];
+ return bpf_unwind_ip_in(aux->prog, ip) ? aux->prog : NULL;
+}
+
+static bool bpf_unwind_rewrite(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra)
+{
+ const struct bpf_cleanup_range *rec;
+ struct bpf_unwind_ctx *ctx = cookie;
+ struct bpf_prog *prog;
+
+ prog = bpf_unwind_find_prog(ctx->aux, ip);
+ if (!prog)
+ return !ctx->cnt;
+ ctx->cnt++;
+
+ /*
+ * The frame that called bpf_unwind(): bpf_exc_patch_unwind_calls()
+ * put 'r0 = 0' and a jump to its pad, or an exit, after the call,
+ * so leave its return address alone and let it go on there. The pad
+ * then starts with r0 at a known zero.
+ */
+ if (ctx->cnt == 1)
+ return bpf_is_subprog(prog);
+
+ rec = bpf_exc_pad_for_ip(prog, ip);
+ *ra = rec ? rec->pad : prog->aux->epilogue_ip;
+
+ return bpf_is_subprog(prog);
+}
+
+/*
+ * @aux is the calling program's, supplied by the verifier (KF_IMPLICIT_ARGS):
+ * programs call bpf_unwind() with no arguments.
+ */
+__bpf_kfunc notrace void bpf_unwind(struct bpf_prog_aux *aux)
{
+ struct bpf_unwind_ctx ctx = { .aux = aux };
+
+ arch_bpf_stack_walk_ra(bpf_unwind_rewrite, &ctx);
}
+NOKPROBE_SYMBOL(bpf_unwind);
__bpf_kfunc void bpf_throw(u64 cookie)
{
@@ -5095,6 +5158,8 @@ BTF_ID_FLAGS(func, bpf_task_get_cgroup1, KF_ACQUIRE | KF_RCU | KF_RET_NULL)
BTF_ID_FLAGS(func, bpf_task_from_pid, KF_ACQUIRE | KF_RET_NULL)
BTF_ID_FLAGS(func, bpf_task_from_vpid, KF_ACQUIRE | KF_RET_NULL)
BTF_ID_FLAGS(func, bpf_throw)
+BTF_ID_FLAGS(func, bpf_unwind, KF_IMPLICIT_ARGS)
+BTF_ID_FLAGS(func, bpf_unwind_resume)
#ifdef CONFIG_BPF_EVENTS
BTF_ID_FLAGS(func, bpf_send_signal_task)
#endif
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 12/23] bpf: Refuse a trampoline that calls a subprog that can unwind
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (10 preceding siblings ...)
2026-10-08 7:50 ` [PATCH bpf-next v9 11/23] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
@ 2026-10-08 7:51 ` Yonghong Song
2026-10-08 8:14 ` sashiko-bot
2026-10-08 7:51 ` [PATCH bpf-next v9 13/23] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
` (10 subsequent siblings)
22 siblings, 1 reply; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:51 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
bpf_unwind() walks up the stack rewriting return addresses. It expects
every frame up to the main function to be one of the program's own BPF
functions, and stops at the first frame that is not. A trampoline
attached to a subprog with fexit, fmod_ret or fsession is such a frame:
it calls the subprog, so it sits between the subprog and its caller:
main -> A -> trampoline -> B -> C C calls bpf_unwind()
The walk rewrites the return into B but stops at the trampoline, so B
returns through it to A after its call, instead of to A's pad or
epilogue: a path the verifier never walked for an unwind.
Refuse such an attachment to a subprog marked might_unwind, now copied
into each function's aux. Nothing else puts a frame between two of a
program's frames: fentry leaves none, a trampoline on main is below the
walk, freplace is a program of its own whose unwind ends as a normal
return to its caller, a callback that can unwind is refused, kprobes
and fgraph cannot hook JIT code, and a tail call replaces a frame.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf.h | 1 +
kernel/bpf/fixups.c | 1 +
kernel/bpf/verifier.c | 16 ++++++++++++++++
3 files changed, 18 insertions(+)
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 7f23f4efde01..1b3b6ee05c09 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -1910,6 +1910,7 @@ struct bpf_prog_aux {
bool priv_stack_requested;
bool changes_pkt_data;
bool might_sleep;
+ bool might_unwind;
bool kprobe_write_ctx;
struct {
s32 keyring_serial;
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index c29e14ffc475..6f719e1b083e 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -1462,6 +1462,7 @@ static int jit_subprogs(struct bpf_verifier_env *env)
func[i]->aux->exception_cb = env->subprog_info[i].is_exception_cb;
func[i]->aux->changes_pkt_data = env->subprog_info[i].changes_pkt_data;
func[i]->aux->might_sleep = env->subprog_info[i].might_sleep;
+ func[i]->aux->might_unwind = env->subprog_info[i].might_unwind;
func[i]->aux->token = prog->aux->token;
if (!i)
func[i]->aux->exception_boundary = env->seen_exception;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 668d811d4e4c..3692d9b163d4 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -21600,6 +21600,22 @@ int bpf_check_attach_target(struct bpf_verifier_log *log,
prog_extension ? "Extension" : "Tracing");
return -EINVAL;
}
+ /*
+ * A trampoline that calls its target stays as a frame between a
+ * subprog and its caller, and bpf_unwind() cannot walk past it:
+ * the frames below would return after their calls rather than at
+ * their landing pads. fentry leaves no frame, and the main
+ * program's caller is below where the walk stops.
+ */
+ if (aux->func && subprog && aux->func[subprog]->aux->might_unwind &&
+ (prog->expected_attach_type == BPF_TRACE_FEXIT ||
+ prog->expected_attach_type == BPF_MODIFY_RETURN ||
+ prog->expected_attach_type == BPF_TRACE_FSESSION)) {
+ bpf_log(log,
+ "Cannot attach fexit, fmod_ret or fsession to %s, which can unwind\n",
+ tname);
+ return -EINVAL;
+ }
conservative = aux->func_info_aux[subprog].unreliable;
if (prog_extension) {
if (conservative) {
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 12/23] bpf: Refuse a trampoline that calls a subprog that can unwind
2026-10-08 7:51 ` [PATCH bpf-next v9 12/23] bpf: Refuse a trampoline that calls a subprog that can unwind Yonghong Song
@ 2026-10-08 8:14 ` sashiko-bot
2026-10-08 16:19 ` Yonghong Song
0 siblings, 1 reply; 38+ messages in thread
From: sashiko-bot @ 2026-10-08 8:14 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Critical] bpf: x86 exception unwinder passes wrong return address pointer causing stack corruption
- [High] bpf: Missing trampoline restriction for bpf_throw() leads to NULL pointer deref
--
commit 554b5300d0510c78440f3ee71cad1d363c75cfaf
Author: Yonghong Song <yonghong.song@linux.dev>
bpf: Refuse a trampoline that calls a subprog that can unwind
This patch refuses fexit, fmod_ret, and fsession trampoline attachments
to subprograms marked with might_unwind, preventing bpf_unwind() from
tripping over unexpected trampoline frames.
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 668d811d4e4cb..3692d9b163d43 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
[ ... ]
> @@ -21600,6 +21600,22 @@ int bpf_check_attach_target(struct bpf_verifier_log *log,
> prog_extension ? "Extension" : "Tracing");
> return -EINVAL;
> }
> + /*
> + * A trampoline that calls its target stays as a frame between a
> + * subprog and its caller, and bpf_unwind() cannot walk past it:
> + * the frames below would return after their calls rather than at
> + * their landing pads. fentry leaves no frame, and the main
> + * program's caller is below where the walk stops.
> + */
> + if (aux->func && subprog && aux->func[subprog]->aux->might_unwind &&
> + (prog->expected_attach_type == BPF_TRACE_FEXIT ||
> + prog->expected_attach_type == BPF_MODIFY_RETURN ||
> + prog->expected_attach_type == BPF_TRACE_FSESSION)) {
[Severity: High]
Does this logic also need to check for bpf_throw()?
The patch correctly prevents attachments for bpf_unwind(), which is tracked
by might_unwind. However, bpf_throw() seems to be missed. It uses an
independent might_throw flag which does not appear to be propagated to
bpf_prog_aux or checked here.
If a subprogram calls bpf_throw() and has a tracing program attached
(like fexit), bpf_stack_walker() encounters the trampoline frame. At that
point, bpf_prog_ksym_find() will return NULL, and since ctx->cnt > 0, it aborts
the walk before setting ctx->aux:
kernel/bpf/helpers.c:bpf_stack_walker() {
...
if (!prog)
return !ctx->cnt;
...
}
Because the walk is aborted early, ctx.aux remains NULL, which leads to a
kernel panic when dereferenced later in bpf_throw():
kernel/bpf/helpers.c:bpf_throw() {
...
ctx.aux->bpf_exception_cb(cookie, ctx.sp + ctx.aux->stack_arg_sp_adjust,
ctx.bp, 0, 0);
...
}
> + bpf_log(log,
> + "Cannot attach fexit, fmod_ret or fsession to %s, which can unwind\n",
> + tname);
> + return -EINVAL;
> + }
> conservative = aux->func_info_aux[subprog].unreliable;
> if (prog_extension) {
[Severity: Critical]
I also have a question about the x86 exception unwinding implementation added
earlier in this series. Can the x86 unwinder corrupt the stack by passing the
wrong return address pointer?
In arch_bpf_stack_walk_ra(), the unwinder passes the current frame's return
address pointer to the callback:
arch/x86/net/bpf_jit_comp.c:arch_bpf_stack_walk_ra() {
...
ra = unwind_get_return_address_ptr(&state);
if (!ra)
break;
if (!consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
...
}
However, bpf_unwind_rewrite() seems to expect the return address pointer
of the previous frame, so it can make the previous frame return directly into
the current frame's cleanup pad.
By passing the current frame's return address instead, bpf_unwind_rewrite()
overwrites the current frame's return address with its own pad or epilogue:
kernel/bpf/helpers.c:bpf_unwind_rewrite() {
...
rec = bpf_exc_pad_for_ip(prog, ip);
*ra = rec ? rec->pad : prog->aux->epilogue_ip;
...
}
On x86, this causes the pad to execute after the current frame has already
returned via leave and ret, corrupting the caller's stack frame. For the main
program, its return address is overwritten with its own epilogue, leading to
a double-epilogue execution.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261008074959.2993751-1-yonghong.song@linux.dev?part=12
^ permalink raw reply [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 12/23] bpf: Refuse a trampoline that calls a subprog that can unwind
2026-10-08 8:14 ` sashiko-bot
@ 2026-10-08 16:19 ` Yonghong Song
0 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 16:19 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 10/8/26 10:14 AM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
> - [Critical] bpf: x86 exception unwinder passes wrong return address pointer causing stack corruption
> - [High] bpf: Missing trampoline restriction for bpf_throw() leads to NULL pointer deref
> --
>
> commit 554b5300d0510c78440f3ee71cad1d363c75cfaf
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Refuse a trampoline that calls a subprog that can unwind
>
> This patch refuses fexit, fmod_ret, and fsession trampoline attachments
> to subprograms marked with might_unwind, preventing bpf_unwind() from
> tripping over unexpected trampoline frames.
>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index 668d811d4e4cb..3692d9b163d43 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
> [ ... ]
>
>> @@ -21600,6 +21600,22 @@ int bpf_check_attach_target(struct bpf_verifier_log *log,
>> prog_extension ? "Extension" : "Tracing");
>> return -EINVAL;
>> }
>> + /*
>> + * A trampoline that calls its target stays as a frame between a
>> + * subprog and its caller, and bpf_unwind() cannot walk past it:
>> + * the frames below would return after their calls rather than at
>> + * their landing pads. fentry leaves no frame, and the main
>> + * program's caller is below where the walk stops.
>> + */
>> + if (aux->func && subprog && aux->func[subprog]->aux->might_unwind &&
>> + (prog->expected_attach_type == BPF_TRACE_FEXIT ||
>> + prog->expected_attach_type == BPF_MODIFY_RETURN ||
>> + prog->expected_attach_type == BPF_TRACE_FSESSION)) {
> [Severity: High]
> Does this logic also need to check for bpf_throw()?
>
> The patch correctly prevents attachments for bpf_unwind(), which is tracked
> by might_unwind. However, bpf_throw() seems to be missed. It uses an
> independent might_throw flag which does not appear to be propagated to
> bpf_prog_aux or checked here.
In the above, we have aux->func[subprog]->aux->might_unwind. If it is true,
that means bpf_throw() will be rejected. So I think we should be okay.
>
> If a subprogram calls bpf_throw() and has a tracing program attached
> (like fexit), bpf_stack_walker() encounters the trampoline frame. At that
> point, bpf_prog_ksym_find() will return NULL, and since ctx->cnt > 0, it aborts
> the walk before setting ctx->aux:
>
> kernel/bpf/helpers.c:bpf_stack_walker() {
> ...
> if (!prog)
> return !ctx->cnt;
> ...
> }
>
> Because the walk is aborted early, ctx.aux remains NULL, which leads to a
> kernel panic when dereferenced later in bpf_throw():
>
> kernel/bpf/helpers.c:bpf_throw() {
> ...
> ctx.aux->bpf_exception_cb(cookie, ctx.sp + ctx.aux->stack_arg_sp_adjust,
> ctx.bp, 0, 0);
> ...
> }
>
>> + bpf_log(log,
>> + "Cannot attach fexit, fmod_ret or fsession to %s, which can unwind\n",
>> + tname);
>> + return -EINVAL;
>> + }
>> conservative = aux->func_info_aux[subprog].unreliable;
>> if (prog_extension) {
> [Severity: Critical]
> I also have a question about the x86 exception unwinding implementation added
> earlier in this series. Can the x86 unwinder corrupt the stack by passing the
> wrong return address pointer?
>
> In arch_bpf_stack_walk_ra(), the unwinder passes the current frame's return
> address pointer to the callback:
>
> arch/x86/net/bpf_jit_comp.c:arch_bpf_stack_walk_ra() {
> ...
> ra = unwind_get_return_address_ptr(&state);
> if (!ra)
> break;
> if (!consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
> ...
> }
>
> However, bpf_unwind_rewrite() seems to expect the return address pointer
> of the previous frame, so it can make the previous frame return directly into
> the current frame's cleanup pad.
>
> By passing the current frame's return address instead, bpf_unwind_rewrite()
> overwrites the current frame's return address with its own pad or epilogue:
>
> kernel/bpf/helpers.c:bpf_unwind_rewrite() {
> ...
> rec = bpf_exc_pad_for_ip(prog, ip);
> *ra = rec ? rec->pad : prog->aux->epilogue_ip;
> ...
> }
>
> On x86, this causes the pad to execute after the current frame has already
> returned via leave and ret, corrupting the caller's stack frame. For the main
> program, its return address is overwritten with its own epilogue, leading to
> a double-epilogue execution.
The issue here seems more appropriate for commit
"bpf: Dispatch cleanup pads by rewriting return addresses". I didn't
find any issues. Selftest results match expectations (manually checked).
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH bpf-next v9 13/23] bpf, x86: Dispatch exception cleanup pads at run time
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (11 preceding siblings ...)
2026-10-08 7:51 ` [PATCH bpf-next v9 12/23] bpf: Refuse a trampoline that calls a subprog that can unwind Yonghong Song
@ 2026-10-08 7:51 ` Yonghong Song
2026-10-08 7:51 ` [PATCH bpf-next v9 14/23] bpf, arm64: " Yonghong Song
` (9 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:51 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
- arch_bpf_stack_walk_ra(): the ORC walk, handing out each frame's
return address, state.ip, and the slot it was read from,
unwind_get_return_address_ptr(). A BPF frame is unwound through its
frame pointer, which the JIT always sets up.
- aux->epilogue_ip: the address of the one epilogue the JIT emits per
function, at the offset it keeps as ctx->cleanup_addr.
- The native cleanup table, filled in once the image is final.
- bpf_jit_supports_cleanup_pads() says yes with CONFIG_UNWINDER_ORC, as
bpf_jit_supports_exceptions() does.
A pad needs no ENDBR: it is only reached as a return address.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
arch/x86/net/bpf_jit_comp.c | 36 ++++++++++++++++++++++++++++++++++++
1 file changed, 36 insertions(+)
diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index 083fcd6cf15b..d434871b5a11 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -15,6 +15,7 @@
#include <linux/memory.h>
#include <linux/sort.h>
#include <linux/execmem.h>
+#include <linux/kprobes.h>
#include <asm/extable.h>
#include <asm/ftrace.h>
#include <asm/set_memory.h>
@@ -3279,6 +3280,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
seen_exit = true;
/* Update cleanup_addr */
ctx->cleanup_addr = proglen;
+ /* Where an unwind sends a frame with no pad. */
+ bpf_prog->aux->epilogue_ip = (u64)image + proglen;
if (bpf_prog_was_classic(bpf_prog) &&
!ns_capable_noaudit(&init_user_ns, CAP_SYS_ADMIN)) {
if (emit_spectre_bhb_barrier(&prog, ip, bpf_prog))
@@ -4462,6 +4465,13 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
*/
bpf_prog_update_insn_ptrs(prog, addrs, image);
+ /*
+ * Same mapping, consumed by the bpf_unwind() walk:
+ * turn the cleanup records into native address ranges now
+ * that the image is final.
+ */
+ bpf_exc_fill_native_ranges(prog, addrs, image);
+
/*
* ctx.prog_offset is used when CFI preambles put code *before*
* the function. See emit_cfi(). For FineIBT specifically this code
@@ -4600,6 +4610,11 @@ bool bpf_jit_supports_exceptions(void)
return IS_ENABLED(CONFIG_UNWINDER_ORC);
}
+bool bpf_jit_supports_cleanup_pads(void)
+{
+ return IS_ENABLED(CONFIG_UNWINDER_ORC);
+}
+
bool bpf_jit_supports_private_stack(void)
{
return true;
@@ -4621,6 +4636,27 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
#endif
}
+notrace void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp,
+ u64 *ra),
+ void *cookie)
+{
+#if defined(CONFIG_UNWINDER_ORC)
+ struct unwind_state state;
+ unsigned long addr, *ra;
+
+ for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
+ unwind_next_frame(&state)) {
+ addr = state.ip;
+ ra = unwind_get_return_address_ptr(&state);
+ if (!ra)
+ break;
+ if (!consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
+ break;
+ }
+#endif
+}
+NOKPROBE_SYMBOL(arch_bpf_stack_walk_ra);
+
void bpf_arch_poke_desc_update(struct bpf_jit_poke_descriptor *poke,
struct bpf_prog *new, struct bpf_prog *old)
{
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 14/23] bpf, arm64: Dispatch exception cleanup pads at run time
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (12 preceding siblings ...)
2026-10-08 7:51 ` [PATCH bpf-next v9 13/23] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
@ 2026-10-08 7:51 ` Yonghong Song
2026-10-08 8:39 ` bot+bpf-ci
2026-10-08 7:51 ` [PATCH bpf-next v9 15/23] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
` (8 subsequent siblings)
22 siblings, 1 reply; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:51 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
- arch_bpf_stack_walk_ra(): walks the frame records with the kernel
unwinder, handing each return address to the consumer and storing a
new one back into the record it was read from. The first frame, whose
return into bpf_unwind() comes from the walk's own record, is
skipped.
- With CONFIG_ARM64_PTR_AUTH_KERNEL and a CPU that supports address
authentication, the new address is signed as the BPF prologue signs
the link register, with PACIASP and the record + 16 as modifier.
- aux->epilogue_ip and the native cleanup table.
- bpf_jit_supports_cleanup_pads() says yes, also with a shadow call
stack: only JITed frames' records are written, and JITed code keeps
no x18 copy of its return address.
A pad needs no BTI: it is only reached as a return address.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
arch/arm64/kernel/stacktrace.c | 74 ++++++++++++++++++++++++++++++++++
arch/arm64/net/bpf_jit_comp.c | 16 ++++++++
2 files changed, 90 insertions(+)
diff --git a/arch/arm64/kernel/stacktrace.c b/arch/arm64/kernel/stacktrace.c
index 3ebcf8c53fb0..56d1310eca32 100644
--- a/arch/arm64/kernel/stacktrace.c
+++ b/arch/arm64/kernel/stacktrace.c
@@ -445,6 +445,80 @@ noinline noinstr void arch_bpf_stack_walk(bool (*consume_entry)(void *cookie, u6
kunwind_stack_walk(arch_bpf_unwind_consume_entry, &data, current, NULL);
}
+struct bpf_unwind_ra_consume_entry_data {
+ bool (*consume_entry)(void *cookie, u64 ip, u64 sp, u64 fp, u64 *ra);
+ void *cookie;
+ unsigned long record;
+ bool seen_first;
+};
+
+static u64 bpf_unwind_sign_ra(u64 ra, u64 modifier)
+{
+ asm volatile(ARM64_ASM_PREAMBLE
+ ".arch_extension pauth\n"
+ " pacia %0, %1"
+ : "+r" (ra) : "r" (modifier));
+ return ra;
+}
+
+/* PACIASP's modifier is the entry sp, record + 16 for a BPF prologue. */
+static void bpf_unwind_store_ra(unsigned long record, u64 ra)
+{
+ struct frame_record *rec = (struct frame_record *)record;
+
+ /*
+ * Whether the slot holds a signed address is a property of the build,
+ * not one to be read off the value: a PAC can come out equal to the
+ * bits stripping puts back, and a signed address would then be taken
+ * for an unsigned one. What signs is CONFIG_ARM64_PTR_AUTH_KERNEL --
+ * the prologue here, and -mbranch-protection for everything the
+ * compiler emits.
+ */
+ if (IS_ENABLED(CONFIG_ARM64_PTR_AUTH_KERNEL) &&
+ system_supports_address_auth())
+ ra = bpf_unwind_sign_ra(ra, record + sizeof(struct frame_record));
+ WRITE_ONCE(rec->lr, ra);
+}
+
+static bool
+arch_bpf_unwind_ra_consume_entry(const struct kunwind_state *state, void *cookie)
+{
+ struct bpf_unwind_ra_consume_entry_data *data = cookie;
+ unsigned long record = data->record;
+ bool seen_first = data->seen_first;
+ u64 ra = state->common.pc;
+ bool cont;
+
+ /* The record this frame's return address will have come out of. */
+ data->record = state->common.fp;
+ data->seen_first = true;
+
+ /*
+ * The first pc returns into bpf_unwind(), from this walk's own frame
+ * record: not a BPF frame, and not one to redirect.
+ */
+ if (!seen_first)
+ return true;
+ /* A consumer that stops still gets to redirect the frame it stopped on. */
+ cont = data->consume_entry(data->cookie, state->common.pc, 0,
+ state->common.fp, &ra);
+ if (ra != state->common.pc)
+ bpf_unwind_store_ra(record, ra);
+ return cont;
+}
+
+noinline noinstr void arch_bpf_stack_walk_ra(bool (*consume_entry)(void *cookie, u64 ip, u64 sp,
+ u64 fp, u64 *ra),
+ void *cookie)
+{
+ struct bpf_unwind_ra_consume_entry_data data = {
+ .consume_entry = consume_entry,
+ .cookie = cookie,
+ };
+
+ kunwind_stack_walk(arch_bpf_unwind_ra_consume_entry, &data, current, NULL);
+}
+
static const char *state_source_string(const struct kunwind_state *state)
{
switch (state->source) {
diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
index 08725cde0c5a..11e94e3bf859 100644
--- a/arch/arm64/net/bpf_jit_comp.c
+++ b/arch/arm64/net/bpf_jit_comp.c
@@ -2423,6 +2423,17 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
* reasons, expects to point to the next instruction)
*/
bpf_prog_update_insn_ptrs(prog, ctx.offset, ctx.ro_image);
+
+ /*
+ * Same byte offsets, consumed by the bpf_unwind() walk:
+ * turn the cleanup records into native address ranges now that
+ * the image is final.
+ */
+ bpf_exc_fill_native_ranges(prog, ctx.offset, ctx.ro_image);
+
+ /* Where an unwind sends a frame with no pad. */
+ prog->aux->epilogue_ip = (u64)ctx.ro_image +
+ ctx.epilogue_offset * AARCH64_INSN_SIZE;
out_off:
if (!ro_header && priv_stack_ptr) {
free_percpu(priv_stack_ptr);
@@ -3415,6 +3426,11 @@ bool bpf_jit_supports_exceptions(void)
return true;
}
+bool bpf_jit_supports_cleanup_pads(void)
+{
+ return true;
+}
+
bool bpf_jit_supports_arena(void)
{
return true;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 14/23] bpf, arm64: Dispatch exception cleanup pads at run time
2026-10-08 7:51 ` [PATCH bpf-next v9 14/23] bpf, arm64: " Yonghong Song
@ 2026-10-08 8:39 ` bot+bpf-ci
2026-10-08 16:23 ` Yonghong Song
0 siblings, 1 reply; 38+ messages in thread
From: bot+bpf-ci @ 2026-10-08 8:39 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 1841 bytes --]
> bpf, arm64: Dispatch exception cleanup pads at run time
TL;DR: Commit message only, no code issues: the changelog should open with a
sentence saying what this does for arm64, and one bullet ("aux->epilogue_ip
and the native cleanup table") lacks a verb.
> bpf, arm64: Dispatch exception cleanup pads at run time
>
> - arch_bpf_stack_walk_ra(): walks the frame records with the kernel
> unwinder, handing each return address to the consumer and storing a
> new one back into the record it was read from. The first frame, whose
> return into bpf_unwind() comes from the walk's own record, is
> skipped.
> - With CONFIG_ARM64_PTR_AUTH_KERNEL and a CPU that supports address
> authentication, the new address is signed as the BPF prologue signs
> the link register, with PACIASP and the record + 16 as modifier.
> - aux->epilogue_ip and the native cleanup table.
> - bpf_jit_supports_cleanup_pads() says yes, also with a shadow call
> stack: only JITed frames' records are written, and JITed code keeps
> no x18 copy of its return address.
>
> A pad needs no BTI: it is only reached as a return address.
>
> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
this isn't a bug, but could the changelog open with a sentence on what this
does for arm64, instead of starting directly with a bullet list?
Also, the bullet "aux->epilogue_ip and the native cleanup table." has no
verb. Could it say what is done with them, for example that the JIT records
each function's epilogue address in aux->epilogue_ip and fills in the native
cleanup ranges once the image is final?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37747693645
^ permalink raw reply [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 14/23] bpf, arm64: Dispatch exception cleanup pads at run time
2026-10-08 8:39 ` bot+bpf-ci
@ 2026-10-08 16:23 ` Yonghong Song
0 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 16:23 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 10/8/26 10:39 AM, bot+bpf-ci@kernel.org wrote:
>> bpf, arm64: Dispatch exception cleanup pads at run time
> TL;DR: Commit message only, no code issues: the changelog should open with a
> sentence saying what this does for arm64, and one bullet ("aux->epilogue_ip
> and the native cleanup table") lacks a verb.
>
>> bpf, arm64: Dispatch exception cleanup pads at run time
>>
>> - arch_bpf_stack_walk_ra(): walks the frame records with the kernel
>> unwinder, handing each return address to the consumer and storing a
>> new one back into the record it was read from. The first frame, whose
>> return into bpf_unwind() comes from the walk's own record, is
>> skipped.
>> - With CONFIG_ARM64_PTR_AUTH_KERNEL and a CPU that supports address
>> authentication, the new address is signed as the BPF prologue signs
>> the link register, with PACIASP and the record + 16 as modifier.
>> - aux->epilogue_ip and the native cleanup table.
>> - bpf_jit_supports_cleanup_pads() says yes, also with a shadow call
>> stack: only JITed frames' records are written, and JITed code keeps
>> no x18 copy of its return address.
>>
>> A pad needs no BTI: it is only reached as a return address.
>>
>> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
> this isn't a bug, but could the changelog open with a sentence on what this
> does for arm64, instead of starting directly with a bullet list?
It is my fault. Should have a small paragraph to highlight the key
message before itemizing different cases.
>
> Also, the bullet "aux->epilogue_ip and the native cleanup table." has no
> verb. Could it say what is done with them, for example that the JIT records
> each function's epilogue address in aux->epilogue_ip and fills in the native
> cleanup ranges once the image is final?
I will expand to have more detailed information for this item.
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37747693645
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH bpf-next v9 15/23] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (13 preceding siblings ...)
2026-10-08 7:51 ` [PATCH bpf-next v9 14/23] bpf, arm64: " Yonghong Song
@ 2026-10-08 7:51 ` Yonghong Song
2026-10-08 7:51 ` [PATCH bpf-next v9 16/23] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
` (7 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:51 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
LLVM ends a cleanup landing pad with a call to _Unwind_Resume, the
unwind ABI's name for carrying an unwind on. The kernel's kfunc is
bpf_unwind_resume: _Unwind_Resume in the kernel's own symbol table, for a
function whose body never runs, would only confuse.
The compiler emits the call but no declaration, and libbpf refuses an
extern it has no BTF for, so the program declares it itself (extern void
_Unwind_Resume(void *) __ksym;), as a language runtime has to. libbpf
only translates the name, both where it resolves the kfunc itself and
where it records the name for a light skeleton's loader to resolve.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/libbpf.c | 45 ++++++++++++++++++++++++++++++++----------
1 file changed, 35 insertions(+), 10 deletions(-)
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index cb09ded90773..5f41d3ba2822 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -9187,6 +9187,23 @@ static void fixup_verifier_log(struct bpf_program *prog, char *buf, size_t buf_s
}
}
+/*
+ * LLVM terminates a cleanup landing pad with a call to _Unwind_Resume, the
+ * base unwind ABI's entry point for carrying an unwind on once a frame's
+ * cleanups have run. The kernel knows it as bpf_unwind_resume. Any other
+ * extern, and a variable of that name, is looked up by @name unchanged.
+ */
+static const char *kern_extern_name(const struct bpf_object *obj,
+ const struct extern_desc *ext, const char *name)
+{
+ const char *essent = ext->essent_name ?: ext->name;
+
+ if (!btf_is_func(btf__type_by_id(obj->btf, ext->btf_id)) ||
+ strcmp(essent, "_Unwind_Resume"))
+ return name;
+ return "bpf_unwind_resume";
+}
+
static int bpf_program_record_relos(struct bpf_program *prog)
{
struct bpf_object *obj = prog->obj;
@@ -9195,6 +9212,7 @@ static int bpf_program_record_relos(struct bpf_program *prog)
for (i = 0; i < prog->nr_reloc; i++) {
struct reloc_desc *relo = &prog->reloc_desc[i];
struct extern_desc *ext = &obj->externs[relo->ext_idx];
+ const char *name;
int kind;
switch (relo->type) {
@@ -9203,14 +9221,14 @@ static int bpf_program_record_relos(struct bpf_program *prog)
continue;
kind = btf_is_var(btf__type_by_id(obj->btf, ext->btf_id)) ?
BTF_KIND_VAR : BTF_KIND_FUNC;
- bpf_gen__record_extern(obj->gen_loader, ext->name,
- ext->is_weak, !ext->ksym.type_id,
- true, kind, relo->insn_idx);
+ name = kern_extern_name(obj, ext, ext->name);
+ bpf_gen__record_extern(obj->gen_loader, name, ext->is_weak,
+ !ext->ksym.type_id, true, kind, relo->insn_idx);
break;
case RELO_EXTERN_CALL:
- bpf_gen__record_extern(obj->gen_loader, ext->name,
- ext->is_weak, false, false, BTF_KIND_FUNC,
- relo->insn_idx);
+ name = kern_extern_name(obj, ext, ext->name);
+ bpf_gen__record_extern(obj->gen_loader, name, ext->is_weak, false,
+ false, BTF_KIND_FUNC, relo->insn_idx);
break;
case RELO_CORE: {
struct bpf_core_relo cr = {
@@ -9684,17 +9702,24 @@ static int bpf_object__resolve_ksym_func_btf_id(struct bpf_object *obj,
struct module_btf *mod_btf = NULL;
const struct btf_type *kern_func;
struct btf *kern_btf = NULL;
+ const char *local_name, *kern_name;
int ret;
local_func_proto_id = ext->ksym.type_id;
- kfunc_id = find_ksym_btf_id(obj, ext->essent_name ?: ext->name, BTF_KIND_FUNC, &kern_btf,
- &mod_btf);
+ local_name = ext->essent_name ?: ext->name;
+ kern_name = kern_extern_name(obj, ext, local_name);
+
+ kfunc_id = find_ksym_btf_id(obj, kern_name, BTF_KIND_FUNC, &kern_btf, &mod_btf);
if (kfunc_id < 0) {
if (kfunc_id == -ESRCH && ext->is_weak)
return 0;
- pr_warn("extern (func ksym) '%s': not found in kernel or module BTFs\n",
- ext->name);
+ if (kern_name != local_name)
+ pr_warn("extern (func ksym) '%s' ('%s' in the kernel): not found in kernel or module BTFs\n",
+ ext->name, kern_name);
+ else
+ pr_warn("extern (func ksym) '%s': not found in kernel or module BTFs\n",
+ ext->name);
return kfunc_id;
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 16/23] libbpf: Add cleanup_info to bpf_prog_load_opts
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (14 preceding siblings ...)
2026-10-08 7:51 ` [PATCH bpf-next v9 15/23] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
@ 2026-10-08 7:51 ` Yonghong Song
2026-10-08 8:12 ` sashiko-bot
2026-10-08 7:51 ` [PATCH bpf-next v9 17/23] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
` (6 subsequent siblings)
22 siblings, 1 reply; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:51 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Let a caller hand the kernel an exception cleanup table:
bpf_prog_load_opts grows cleanup_info_cnt, cleanup_info and
cleanup_info_rec_size, which bpf_prog_load() passes on to BPF_PROG_LOAD.
cleanup_info_cnt comes first to fill the hole after fd_array_cnt.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/bpf.c | 6 +++++-
tools/lib/bpf/bpf.h | 7 ++++++-
2 files changed, 11 insertions(+), 2 deletions(-)
diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
index b49822d212ae..b4031f62bdee 100644
--- a/tools/lib/bpf/bpf.c
+++ b/tools/lib/bpf/bpf.c
@@ -295,7 +295,7 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
const struct bpf_insn *insns, size_t insn_cnt,
struct bpf_prog_load_opts *opts)
{
- const size_t attr_sz = offsetofend(union bpf_attr, keyring_id);
+ const size_t attr_sz = offsetofend(union bpf_attr, cleanup_info_cnt);
void *finfo = NULL, *linfo = NULL;
const char *func_info, *line_info;
__u32 log_size, log_level, attach_prog_fd, attach_btf_obj_fd;
@@ -370,6 +370,10 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
attr.fd_array = ptr_to_u64(OPTS_GET(opts, fd_array, NULL));
attr.fd_array_cnt = OPTS_GET(opts, fd_array_cnt, 0);
+ attr.cleanup_info = ptr_to_u64(OPTS_GET(opts, cleanup_info, NULL));
+ attr.cleanup_info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
+ attr.cleanup_info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
+
if (log_level) {
attr.log_buf = ptr_to_u64(log_buf);
attr.log_size = log_size;
diff --git a/tools/lib/bpf/bpf.h b/tools/lib/bpf/bpf.h
index 826d9cc9ab65..cbe56ddc8cf7 100644
--- a/tools/lib/bpf/bpf.h
+++ b/tools/lib/bpf/bpf.h
@@ -128,9 +128,14 @@ struct bpf_prog_load_opts {
/* if set, provides the length of fd_array */
__u32 fd_array_cnt;
+
+ /* exception cleanup table, from the .bpf_cleanup section */
+ __u32 cleanup_info_cnt;
+ const void *cleanup_info;
+ __u32 cleanup_info_rec_size;
size_t :0;
};
-#define bpf_prog_load_opts__last_field fd_array_cnt
+#define bpf_prog_load_opts__last_field cleanup_info_rec_size
LIBBPF_API int bpf_prog_load(enum bpf_prog_type prog_type,
const char *prog_name, const char *license,
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 16/23] libbpf: Add cleanup_info to bpf_prog_load_opts
2026-10-08 7:51 ` [PATCH bpf-next v9 16/23] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
@ 2026-10-08 8:12 ` sashiko-bot
2026-10-08 16:25 ` Yonghong Song
0 siblings, 1 reply; 38+ messages in thread
From: sashiko-bot @ 2026-10-08 8:12 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: libbpf misaligned cleanup_info read during verifier log retry
--
commit 44e841423bd633c0b2885ef9ad198f15d3871fba
Author: Yonghong Song <yonghong.song@linux.dev>
libbpf: Add cleanup_info to bpf_prog_load_opts
This commit adds cleanup_info fields to bpf_prog_load_opts to allow
callers to pass an exception cleanup table to the kernel during program load.
The new fields are propagated to union bpf_attr for the BPF_PROG_LOAD syscall.
> diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
> index b49822d212aed..b4031f62bdee5 100644
> --- a/tools/lib/bpf/bpf.c
> +++ b/tools/lib/bpf/bpf.c
[ ... ]
> @@ -370,6 +370,10 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
> attr.fd_array = ptr_to_u64(OPTS_GET(opts, fd_array, NULL));
> attr.fd_array_cnt = OPTS_GET(opts, fd_array_cnt, 0);
>
> + attr.cleanup_info = ptr_to_u64(OPTS_GET(opts, cleanup_info, NULL));
> + attr.cleanup_info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
> + attr.cleanup_info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
> +
[Severity: Medium]
If the first sys_bpf_prog_load() fails with -E2BIG, the kernel can mutate
attr.cleanup_info_rec_size to a smaller size.
During the verifier log retry block later in bpf_prog_load():
if (log_level == 0 && log_buf) {
...
attr.log_level = 1;
fd = sys_bpf_prog_load(&attr, attr_sz, attempts);
If attr.cleanup_info_rec_size was mutated, the retry uses the smaller size
while attr.cleanup_info still points to the user's original array with the
larger stride.
Does this cause the kernel to perform misaligned reads starting from the
second record and parse garbage exception info?
> if (log_level) {
> attr.log_buf = ptr_to_u64(log_buf);
> attr.log_size = log_size;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261008074959.2993751-1-yonghong.song@linux.dev?part=16
^ permalink raw reply [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 16/23] libbpf: Add cleanup_info to bpf_prog_load_opts
2026-10-08 8:12 ` sashiko-bot
@ 2026-10-08 16:25 ` Yonghong Song
0 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 16:25 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 10/8/26 10:12 AM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Medium] bpf: libbpf misaligned cleanup_info read during verifier log retry
> --
>
> commit 44e841423bd633c0b2885ef9ad198f15d3871fba
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> libbpf: Add cleanup_info to bpf_prog_load_opts
>
> This commit adds cleanup_info fields to bpf_prog_load_opts to allow
> callers to pass an exception cleanup table to the kernel during program load.
> The new fields are propagated to union bpf_attr for the BPF_PROG_LOAD syscall.
>
>> diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
>> index b49822d212aed..b4031f62bdee5 100644
>> --- a/tools/lib/bpf/bpf.c
>> +++ b/tools/lib/bpf/bpf.c
> [ ... ]
>> @@ -370,6 +370,10 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
>> attr.fd_array = ptr_to_u64(OPTS_GET(opts, fd_array, NULL));
>> attr.fd_array_cnt = OPTS_GET(opts, fd_array_cnt, 0);
>>
>> + attr.cleanup_info = ptr_to_u64(OPTS_GET(opts, cleanup_info, NULL));
>> + attr.cleanup_info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
>> + attr.cleanup_info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
>> +
> [Severity: Medium]
> If the first sys_bpf_prog_load() fails with -E2BIG, the kernel can mutate
> attr.cleanup_info_rec_size to a smaller size.
>
> During the verifier log retry block later in bpf_prog_load():
>
> if (log_level == 0 && log_buf) {
> ...
> attr.log_level = 1;
> fd = sys_bpf_prog_load(&attr, attr_sz, attempts);
>
> If attr.cleanup_info_rec_size was mutated, the retry uses the smaller size
> while attr.cleanup_info still points to the user's original array with the
> larger stride.
>
> Does this cause the kernel to perform misaligned reads starting from the
> second record and parse garbage exception info?
This is the first time we introduced cleanup_info_rec_size. There is no
need to retry by changing attr.cleanup_info_rec_size. If users have
a bigger cleanup_info_rec_size and causing -E2BIG, it is users' problem.
>
>> if (log_level) {
>> attr.log_buf = ptr_to_u64(log_buf);
>> attr.log_size = log_size;
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH bpf-next v9 17/23] libbpf: Collect .bpf_cleanup records and pass them to the kernel
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (15 preceding siblings ...)
2026-10-08 7:51 ` [PATCH bpf-next v9 16/23] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
@ 2026-10-08 7:51 ` Yonghong Song
2026-10-08 7:51 ` [PATCH bpf-next v9 18/23] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
` (5 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:51 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Parse the compiler-emitted .bpf_cleanup section and hand the table to
BPF_PROG_LOAD:
- At open, each record field, a byte offset into a code section, is
resolved through its .rel.bpf_cleanup relocation: R_BPF_64_NODYLD32
from LLVM or R_BPF_64_ABS32 from GNU as.
- At relocation, once a main program's subprogs are appended, the
records in the main program and in those subprogs get final
instruction indices and are sorted by begin_off, as the kernel wants.
A record in code no function claims belongs to a weak function the
static linker kept after overriding it, and is skipped, as that
code's relocations are.
- At load, the table goes through the bpf_prog_load() options, from
bpf_object_load_prog() and from bpf_program__clone(), the path
veristat uses.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/libbpf.c | 285 +++++++++++++++++++++++++++++++-
tools/lib/bpf/libbpf_internal.h | 3 +
2 files changed, 286 insertions(+), 2 deletions(-)
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index 5f41d3ba2822..a0e307a4aa5a 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -516,6 +516,11 @@ struct bpf_program {
void *line_info;
__u32 line_info_rec_size;
__u32 line_info_cnt;
+
+ struct bpf_cleanup_info *cleanup_info;
+ __u32 cleanup_info_rec_size;
+ __u32 cleanup_info_cnt;
+
__u32 prog_flags;
__u8 hash[SHA256_DIGEST_LENGTH];
@@ -572,6 +577,7 @@ struct bpf_struct_ops {
* are loaded with BPF_F_ARENA_SCALAR.
*/
#define ARENA_DATA_SEC ".arena.data"
+#define CLEANUP_SEC ".bpf_cleanup"
enum libbpf_map_type {
LIBBPF_MAP_UNSPEC,
@@ -710,6 +716,25 @@ struct elf_sec_desc {
Elf_Data *data;
};
+#define CLEANUP_REC_FIELDS (sizeof(struct bpf_cleanup_info) / sizeof(__u32))
+
+/* Index of each field of struct bpf_cleanup_info, read as an array of __u32. */
+enum {
+ CLEANUP_REC_BEGIN,
+ CLEANUP_REC_END,
+ CLEANUP_REC_PAD,
+};
+
+/*
+ * One (begin, end, landing_pad) triple from .bpf_cleanup, each field resolved
+ * from its relocation to a section and an instruction index within it. Final
+ * indices wait for subprogram placement, which differs per main program.
+ */
+struct cleanup_raw_rec {
+ int sec_idx[CLEANUP_REC_FIELDS];
+ size_t insn_idx[CLEANUP_REC_FIELDS];
+};
+
struct elf_state {
int fd;
const void *obj_buf;
@@ -729,6 +754,8 @@ struct elf_state {
bool has_st_ops;
int arena_data_shndx;
int jumptables_data_shndx;
+ Elf_Data *cleanup_data;
+ int cleanup_shndx;
};
struct usdt_manager;
@@ -817,6 +844,9 @@ struct bpf_object {
void *jumptables_data;
size_t jumptables_data_sz;
+ struct cleanup_raw_rec *cleanup_recs;
+ size_t cleanup_rec_cnt;
+
struct {
struct bpf_program *prog;
unsigned int sym_off;
@@ -874,6 +904,7 @@ void bpf_program__unload(struct bpf_program *prog)
zfree(&prog->func_info);
zfree(&prog->line_info);
zfree(&prog->subprogs);
+ zfree(&prog->cleanup_info);
}
static void bpf_program__exit(struct bpf_program *prog)
@@ -1639,6 +1670,7 @@ static struct bpf_object *bpf_object__new(const char *path,
obj->efile.obj_buf_sz = obj_buf_sz;
obj->efile.btf_maps_shndx = -1;
obj->efile.arena_data_shndx = -1;
+ obj->efile.cleanup_shndx = -1;
obj->kconfig_map_idx = -1;
obj->arena_map_idx = -1;
@@ -1658,6 +1690,7 @@ static void bpf_object__elf_finish(struct bpf_object *obj)
obj->efile.ehdr = NULL;
obj->efile.symbols = NULL;
obj->efile.arena_data = NULL;
+ obj->efile.cleanup_data = NULL;
zfree(&obj->efile.secs);
obj->efile.sec_cnt = 0;
@@ -4275,6 +4308,9 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
sec_desc->shdr = sh;
sec_desc->data = data;
obj->efile.has_st_ops = true;
+ } else if (strcmp(name, CLEANUP_SEC) == 0) {
+ obj->efile.cleanup_data = data;
+ obj->efile.cleanup_shndx = idx;
} else if (strcmp(name, ARENA_SEC) == 0) {
obj->efile.arena_data = data;
obj->efile.arena_data_shndx = idx;
@@ -4300,8 +4336,9 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
/*
* Only do relo for section with exec instructions,
- * struct_ops, maps, and read-only data that might
- * have pointers to functions.
+ * struct_ops, maps, read-only data that might have
+ * pointers to functions, and the exception cleanup
+ * table.
*/
if (!section_have_execinstr(obj, targ_sec_idx) &&
!(obj->data_in_arena &&
@@ -4315,6 +4352,7 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
strcmp(name, ".rel" STRUCT_OPS_LINK_SEC) &&
strcmp(name, ".rel?" STRUCT_OPS_SEC) &&
strcmp(name, ".rel?" STRUCT_OPS_LINK_SEC) &&
+ strcmp(name, ".rel" CLEANUP_SEC) &&
strcmp(name, ".rel" MAPS_ELF_SEC)) {
pr_info("elf: skipping relo section(%d) %s for section(%d) %s\n",
idx, name, targ_sec_idx,
@@ -5099,6 +5137,217 @@ static struct bpf_program *find_prog_by_sec_insn(const struct bpf_object *obj,
return NULL;
}
+static int bpf_object_init_cleanup_info(struct bpf_object *obj)
+{
+ Elf_Data *data = obj->efile.cleanup_data;
+ size_t i, nrels, nslots, nrecs;
+ struct cleanup_raw_rec *recs;
+ Elf_Data *relo = NULL;
+ const __u32 *vals;
+ int ret = 0;
+ bool native;
+
+ if (!data || obj->efile.cleanup_shndx < 0 || !data->d_size)
+ return 0;
+
+ native = is_native_endianness(obj);
+
+ for (i = 0; i < obj->efile.sec_cnt; i++) {
+ struct elf_sec_desc *sd = &obj->efile.secs[i];
+
+ if (sd->sec_type == SEC_RELO && sd->shdr &&
+ sd->shdr->sh_info == (Elf64_Word)obj->efile.cleanup_shndx) {
+ relo = sd->data;
+ break;
+ }
+ }
+ if (!relo) {
+ pr_warn("%s present without relocations\n", CLEANUP_SEC);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ if (data->d_size % sizeof(struct bpf_cleanup_info)) {
+ pr_warn("%s size %zu is not a multiple of the record size %zu\n",
+ CLEANUP_SEC, data->d_size, sizeof(struct bpf_cleanup_info));
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ vals = data->d_buf;
+ nslots = data->d_size / sizeof(__u32);
+ nrecs = data->d_size / sizeof(struct bpf_cleanup_info);
+
+ recs = calloc(nrecs, sizeof(*recs));
+ if (!recs)
+ return -ENOMEM;
+ for (i = 0; i < nslots; i++)
+ recs[i / CLEANUP_REC_FIELDS].sec_idx[i % CLEANUP_REC_FIELDS] = -1;
+
+ /* One relocation per 4-byte field, naming the section it points into. */
+ nrels = relo->d_size / sizeof(Elf64_Rel);
+ for (i = 0; i < nrels; i++) {
+ Elf64_Rel *rel = elf_rel_by_idx(relo, i);
+ Elf64_Sym *sym = elf_sym_by_idx(obj, ELF64_R_SYM(rel->r_info));
+ size_t type = ELF64_R_TYPE(rel->r_info);
+ size_t slot = rel->r_offset / sizeof(__u32);
+ struct cleanup_raw_rec *rec;
+ size_t off;
+
+ if (type != R_BPF_64_NODYLD32 && type != R_BPF_64_ABS32) {
+ pr_warn("%s: relocation %zu has unexpected type %zu\n",
+ CLEANUP_SEC, i, type);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ if (!sym || slot >= nslots || rel->r_offset % sizeof(__u32)) {
+ pr_warn("%s: bad relocation %zu\n", CLEANUP_SEC, i);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ /*
+ * The addend lives in the section data, which libelf leaves in
+ * the object's byte order; a non-section symbol additionally
+ * contributes its own value.
+ */
+ off = (native ? vals[slot] : bswap_32(vals[slot])) + sym->st_value;
+ if (off % BPF_INSN_SZ) {
+ pr_warn("%s: field %zu offset %zu is not instruction aligned\n",
+ CLEANUP_SEC, slot, off);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ rec = &recs[slot / CLEANUP_REC_FIELDS];
+ if (rec->sec_idx[slot % CLEANUP_REC_FIELDS] >= 0) {
+ pr_warn("%s: field %zu has more than one relocation\n",
+ CLEANUP_SEC, slot);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ rec->sec_idx[slot % CLEANUP_REC_FIELDS] = sym->st_shndx;
+ rec->insn_idx[slot % CLEANUP_REC_FIELDS] = off / BPF_INSN_SZ;
+ }
+
+ for (i = 0; i < nslots; i++) {
+ if (recs[i / CLEANUP_REC_FIELDS].sec_idx[i % CLEANUP_REC_FIELDS] < 0) {
+ pr_warn("%s: field %zu has no relocation\n", CLEANUP_SEC, i);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ }
+
+ obj->cleanup_recs = recs;
+ obj->cleanup_rec_cnt = nrecs;
+ return 0;
+out:
+ free(recs);
+ return ret;
+}
+
+static int cmp_cleanup_info(const void *a, const void *b)
+{
+ const struct bpf_cleanup_info *x = a, *y = b;
+
+ if (x->begin_off == y->begin_off)
+ return 0;
+ return x->begin_off < y->begin_off ? -1 : 1;
+}
+
+static int bpf_prog_collect_cleanup_info(struct bpf_object *obj,
+ struct bpf_program *prog)
+{
+ size_t i;
+ int j;
+
+ for (i = 0; i < obj->cleanup_rec_cnt; i++) {
+ struct cleanup_raw_rec *raw = &obj->cleanup_recs[i];
+ struct bpf_program *owner = NULL;
+ struct bpf_cleanup_info ci = {};
+ __u32 fields[CLEANUP_REC_FIELDS];
+ void *tmp;
+
+ for (j = 0; j < CLEANUP_REC_FIELDS; j++) {
+ size_t idx = raw->insn_idx[j], final;
+ struct bpf_program *p;
+
+ /*
+ * An exclusive end may name the instruction past the
+ * last of a function, so ask about the last one the
+ * range covers, the way the kernel does.
+ */
+ if (j == CLEANUP_REC_END) {
+ if (!idx) {
+ pr_warn("%s: record %zu ends at instruction 0\n",
+ CLEANUP_SEC, i);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ idx--;
+ }
+
+ p = find_prog_by_sec_insn(obj, raw->sec_idx[j], idx);
+ if (!p && j == CLEANUP_REC_BEGIN) {
+ /*
+ * Code no function claims is an overridden weak
+ * function the linker kept: skip its records,
+ * as its relocations are.
+ */
+ pr_debug("%s: record %zu is in no function, probably an overridden weak function, skipping\n",
+ CLEANUP_SEC, i);
+ break;
+ }
+ if (!p) {
+ pr_warn("%s: record %zu field %d is not inside a function\n",
+ CLEANUP_SEC, i, j);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ if (!owner) {
+ owner = p;
+ } else if (owner != p) {
+ pr_warn("%s: record %zu spans functions '%s' and '%s'\n",
+ CLEANUP_SEC, i, owner->name, p->name);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ if (owner == prog) {
+ final = raw->insn_idx[j] - prog->sec_insn_off;
+ } else if (prog_is_subprog(obj, owner) && owner->sub_insn_off) {
+ /*
+ * sub_insn_off is where this subprogram was
+ * appended to the main program being relocated;
+ * zero means it is not part of it.
+ */
+ final = owner->sub_insn_off +
+ raw->insn_idx[j] - owner->sec_insn_off;
+ } else {
+ owner = NULL;
+ break;
+ }
+ fields[j] = final;
+ }
+ if (!owner)
+ continue;
+
+ ci.begin_off = fields[CLEANUP_REC_BEGIN];
+ ci.end_off = fields[CLEANUP_REC_END];
+ ci.landing_pad_off = fields[CLEANUP_REC_PAD];
+
+ tmp = libbpf_reallocarray(prog->cleanup_info, prog->cleanup_info_cnt + 1,
+ sizeof(*prog->cleanup_info));
+ if (!tmp)
+ return -ENOMEM;
+ prog->cleanup_info = tmp;
+ prog->cleanup_info_rec_size = sizeof(struct bpf_cleanup_info);
+ prog->cleanup_info[prog->cleanup_info_cnt++] = ci;
+
+ pr_debug("prog '%s': cleanup region [%u,%u) -> landing pad %u\n",
+ prog->name, ci.begin_off, ci.end_off, ci.landing_pad_off);
+ }
+
+ if (!prog->cleanup_info_cnt)
+ return 0;
+
+ qsort(prog->cleanup_info, prog->cleanup_info_cnt,
+ sizeof(*prog->cleanup_info), cmp_cleanup_info);
+ return 0;
+}
+
static int
bpf_object__collect_prog_relos(struct bpf_object *obj, Elf64_Shdr *shdr, Elf_Data *data)
{
@@ -8197,6 +8446,13 @@ static int bpf_object__relocate(struct bpf_object *obj, const char *targ_btf_pat
return err;
}
}
+
+ err = bpf_prog_collect_cleanup_info(obj, prog);
+ if (err) {
+ pr_warn("prog '%s': failed to collect cleanup info: %s\n",
+ prog->name, errstr(err));
+ return err;
+ }
}
for (i = 0; i < obj->nr_programs; i++) {
prog = &obj->programs[i];
@@ -8561,6 +8817,9 @@ static int bpf_object__collect_relos(struct bpf_object *obj)
return -LIBBPF_ERRNO__INTERNAL;
}
+ if (idx == obj->efile.cleanup_shndx)
+ continue;
+
if (obj->efile.secs[idx].sec_type == SEC_RODATA)
err = bpf_object__collect_rodata_relos(obj, shdr, data);
else if (obj->efile.secs[idx].sec_type == SEC_DATA)
@@ -8856,6 +9115,11 @@ static int bpf_object_load_prog(struct bpf_object *obj, struct bpf_program *prog
load_attr.line_info_rec_size = prog->line_info_rec_size;
load_attr.line_info_cnt = prog->line_info_cnt;
}
+ if (prog->cleanup_info_cnt) {
+ load_attr.cleanup_info = prog->cleanup_info;
+ load_attr.cleanup_info_cnt = prog->cleanup_info_cnt;
+ load_attr.cleanup_info_rec_size = prog->cleanup_info_rec_size;
+ }
load_attr.log_level = log_level;
load_attr.prog_flags = prog->prog_flags;
/* the program accesses its data in arena through plain numbers */
@@ -9460,6 +9724,7 @@ static struct bpf_object *bpf_object_open(const char *path, const void *obj_buf,
err = err ? : bpf_object__init_maps(obj, opts);
err = err ? : bpf_object_init_progs(obj, opts);
err = err ? : bpf_object__collect_relos(obj);
+ err = err ? : bpf_object_init_cleanup_info(obj);
if (err)
goto out;
@@ -10597,6 +10862,9 @@ void bpf_object__close(struct bpf_object *obj)
zfree(&obj->jumptables_data);
obj->jumptables_data_sz = 0;
+ zfree(&obj->cleanup_recs);
+ obj->cleanup_rec_cnt = 0;
+
for (i = 0; i < obj->jumptable_map_cnt; i++)
close(obj->jumptable_maps[i].fd);
zfree(&obj->jumptable_maps);
@@ -11006,6 +11274,19 @@ int bpf_program__clone(struct bpf_program *prog, const struct bpf_prog_load_opts
attr.line_info_rec_size = info ? info_rec_size : prog->line_info_rec_size;
}
+ /* exception cleanup table */
+ info = OPTS_GET(opts, cleanup_info, NULL);
+ info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
+ info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
+ if (!!info != !!info_cnt || !!info != !!info_rec_size) {
+ pr_warn("prog '%s': cleanup_info, cleanup_info_cnt, and cleanup_info_rec_size must all be specified or all omitted\n",
+ prog->name);
+ return libbpf_err(-EINVAL);
+ }
+ attr.cleanup_info = info ?: prog->cleanup_info;
+ attr.cleanup_info_cnt = info ? info_cnt : prog->cleanup_info_cnt;
+ attr.cleanup_info_rec_size = info ? info_rec_size : prog->cleanup_info_rec_size;
+
/* Logging is caller-controlled; no fallback to prog/obj log settings */
attr.log_buf = OPTS_GET(opts, log_buf, NULL);
attr.log_size = OPTS_GET(opts, log_size, 0);
diff --git a/tools/lib/bpf/libbpf_internal.h b/tools/lib/bpf/libbpf_internal.h
index 546f65b95cf4..f1630f03d5f5 100644
--- a/tools/lib/bpf/libbpf_internal.h
+++ b/tools/lib/bpf/libbpf_internal.h
@@ -56,6 +56,9 @@
#ifndef R_BPF_64_ABS32
#define R_BPF_64_ABS32 3
#endif
+#ifndef R_BPF_64_NODYLD32
+#define R_BPF_64_NODYLD32 4
+#endif
#ifndef R_BPF_64_32
#define R_BPF_64_32 10
#endif
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 18/23] libbpf: Carry the exception cleanup table through the light skeleton
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (16 preceding siblings ...)
2026-10-08 7:51 ` [PATCH bpf-next v9 17/23] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
@ 2026-10-08 7:51 ` Yonghong Song
2026-10-08 7:51 ` [PATCH bpf-next v9 19/23] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
` (4 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:51 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A light skeleton does not call bpf_prog_load(): its loader program
builds the attr itself, so a program loaded this way reached the kernel
without its table and was refused with "unreachable insn". Carry the
table the way func_info and line_info are carried: the records in the
loader's blob, the count and record size in the attr, and a relocation
for the blob's address.
Only a program with records grows the attr to cleanup_info_cnt, so light
skeletons without them are unchanged.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/gen_loader.c | 29 ++++++++++++++++++++++++++---
tools/lib/bpf/libbpf_internal.h | 7 +++++++
2 files changed, 33 insertions(+), 3 deletions(-)
diff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c
index 251392aa8b41..0aa87ea685e7 100644
--- a/tools/lib/bpf/gen_loader.c
+++ b/tools/lib/bpf/gen_loader.c
@@ -994,13 +994,15 @@ static void cleanup_relos(struct bpf_gen *gen, int insns)
cleanup_core_relo(gen);
}
-/* Convert func, line, and core relo info blobs to target endianness */
+/* Convert func, line, core relo and cleanup info blobs to target endianness */
static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
- int core_relos, struct bpf_prog_load_opts *load_attr)
+ int core_relos, int cleanup_info,
+ struct bpf_prog_load_opts *load_attr)
{
struct bpf_func_info *fi = gen->data_start + func_info;
struct bpf_line_info *li = gen->data_start + line_info;
struct bpf_core_relo *cr = gen->data_start + core_relos;
+ struct bpf_cleanup_info *ci = gen->data_start + cleanup_info;
int i;
for (i = 0; i < load_attr->func_info_cnt; i++)
@@ -1011,6 +1013,9 @@ static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
for (i = 0; i < gen->core_relo_cnt; i++)
bpf_core_relo_bswap(cr++);
+
+ for (i = 0; i < load_attr->cleanup_info_cnt; i++)
+ bpf_cleanup_info_bswap(ci++);
}
void bpf_gen__prog_load(struct bpf_gen *gen,
@@ -1024,10 +1029,15 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
load_attr->line_info_rec_size;
int core_relo_tot_sz = gen->core_relo_cnt *
sizeof(struct bpf_core_relo);
+ int cleanup_info_tot_sz = load_attr->cleanup_info_cnt *
+ load_attr->cleanup_info_rec_size;
int prog_load_attr, license_off, insns_off, func_info, line_info, core_relos;
int attr_size = offsetofend(union bpf_attr, core_relo_rec_size);
+ int cleanup_info;
union bpf_attr attr;
+ if (load_attr->cleanup_info_cnt)
+ attr_size = offsetofend(union bpf_attr, cleanup_info_cnt);
memset(&attr, 0, attr_size);
/* add license string to blob of bytes */
license_off = add_data(gen, license, strlen(license) + 1);
@@ -1074,9 +1084,17 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
core_relos, gen->core_relo_cnt,
sizeof(struct bpf_core_relo));
+ attr.cleanup_info_rec_size = tgt_endian(load_attr->cleanup_info_rec_size);
+ attr.cleanup_info_cnt = tgt_endian(load_attr->cleanup_info_cnt);
+ cleanup_info = add_data(gen, load_attr->cleanup_info, cleanup_info_tot_sz);
+ pr_debug("gen: prog_load: cleanup_info: off %d cnt %u rec size %u\n",
+ cleanup_info, load_attr->cleanup_info_cnt,
+ load_attr->cleanup_info_rec_size);
+
/* convert all info blobs to target endianness */
if (gen->swapped_endian && !gen->error)
- info_blob_bswap(gen, func_info, line_info, core_relos, load_attr);
+ info_blob_bswap(gen, func_info, line_info, core_relos, cleanup_info,
+ load_attr);
libbpf_strlcpy(attr.prog_name, prog_name, sizeof(attr.prog_name));
prog_load_attr = add_data(gen, &attr, attr_size);
@@ -1098,6 +1116,11 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
/* populate union bpf_attr with a pointer to core_relos */
emit_rel_store(gen, attr_field(prog_load_attr, core_relos), core_relos);
+ /* with no records there is no blob of them to point the attr at */
+ if (load_attr->cleanup_info_cnt)
+ emit_rel_store(gen, attr_field(prog_load_attr, cleanup_info),
+ cleanup_info);
+
/* populate union bpf_attr fd_array with a pointer to data where map_fds are saved */
emit_rel_store(gen, attr_field(prog_load_attr, fd_array), gen->fd_array);
diff --git a/tools/lib/bpf/libbpf_internal.h b/tools/lib/bpf/libbpf_internal.h
index f1630f03d5f5..9d341839ca74 100644
--- a/tools/lib/bpf/libbpf_internal.h
+++ b/tools/lib/bpf/libbpf_internal.h
@@ -572,6 +572,13 @@ static inline void bpf_core_relo_bswap(struct bpf_core_relo *i)
i->kind = bswap_32(i->kind);
}
+static inline void bpf_cleanup_info_bswap(struct bpf_cleanup_info *i)
+{
+ i->begin_off = bswap_32(i->begin_off);
+ i->end_off = bswap_32(i->end_off);
+ i->landing_pad_off = bswap_32(i->landing_pad_off);
+}
+
enum btf_field_iter_kind {
BTF_FIELD_ITER_IDS,
BTF_FIELD_ITER_STRS,
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 19/23] libbpf: Let the static linker carry .bpf_cleanup relocations
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (17 preceding siblings ...)
2026-10-08 7:51 ` [PATCH bpf-next v9 18/23] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
@ 2026-10-08 7:51 ` Yonghong Song
2026-10-08 8:14 ` sashiko-bot
2026-10-08 7:51 ` [PATCH bpf-next v9 20/23] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
` (3 subsequent siblings)
22 siblings, 1 reply; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:51 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Each field of a .bpf_cleanup record is a byte offset into a code
section: a 32-bit relocation against that section's symbol, with the
offset as the implicit addend. The linker refuses it: it does not know
LLVM's R_BPF_64_NODYLD32, and from a non-executable section it takes a
relocation against a section symbol only as a 64-bit pointer to code.
Take it in both spellings, R_BPF_64_NODYLD32 from LLVM and
R_BPF_64_ABS32 from GNU as, adding the code section's new start to the
stored byte offset. Since any non-executable section can now reach that
path, it checks itself what the old refusal covered: that the target has
data, and that r_offset is aligned and in bounds.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/linker.c | 37 ++++++++++++++++++++++++++++++++++++-
1 file changed, 36 insertions(+), 1 deletion(-)
diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
index f3f71c452f00..c607f51e3cb4 100644
--- a/tools/lib/bpf/linker.c
+++ b/tools/lib/bpf/linker.c
@@ -1036,7 +1036,8 @@ static int linker_sanity_check_elf_relos(struct src_obj *obj, struct src_sec *se
size_t sym_type = ELF64_R_TYPE(relo->r_info);
if (sym_type != R_BPF_64_64 && sym_type != R_BPF_64_32 &&
- sym_type != R_BPF_64_ABS64 && sym_type != R_BPF_64_ABS32) {
+ sym_type != R_BPF_64_ABS64 && sym_type != R_BPF_64_ABS32 &&
+ sym_type != R_BPF_64_NODYLD32) {
pr_warn("ELF relo #%d in section #%zu has unexpected type %zu in %s\n",
i, sec->sec_idx, sym_type, obj->filename);
return -EINVAL;
@@ -2263,6 +2264,7 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
if (ELF64_ST_TYPE(src_sym->st_info) == STT_SECTION) {
struct src_sec *sec = &obj->secs[src_sym->st_shndx];
struct bpf_insn *insn;
+ __u32 *val;
if (src_linked_sec->shdr->sh_flags & SHF_EXECINSTR) {
/* calls to the very first static function inside
@@ -2297,6 +2299,39 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
if (linker->swapped_endian)
off = bswap_64(off);
memcpy(ptr, &off, sizeof(off));
+ } else if ((sym_type == R_BPF_64_NODYLD32 ||
+ sym_type == R_BPF_64_ABS32) &&
+ (sec->shdr->sh_flags & SHF_EXECINSTR)) {
+ /*
+ * A byte offset into a code section,
+ * stored in place. LLVM spells this
+ * relocation NODYLD32 and GNU as
+ * spells it ABS32; being bytes, the
+ * section's new start goes in as it
+ * is, not scaled the way a call's
+ * instruction index is above.
+ *
+ * r_offset is checked only for an
+ * executable section, and SHT_NOBITS
+ * has no raw_data, so bound it here --
+ * subtracting, so it cannot wrap.
+ */
+ if (!dst_linked_sec->raw_data ||
+ dst_linked_sec->sec_sz < (int)sizeof(*val) ||
+ dst_rel->r_offset % sizeof(*val) ||
+ dst_rel->r_offset >
+ (size_t)dst_linked_sec->sec_sz - sizeof(*val)) {
+ pr_warn("ELF relo #%d in section #%zu points outside the data of section '%s' in %s\n",
+ j, src_sec->sec_idx,
+ dst_linked_sec->sec_name,
+ obj->filename);
+ return -EINVAL;
+ }
+ val = dst_linked_sec->raw_data + dst_rel->r_offset;
+ if (linker->swapped_endian)
+ *val = bswap_32(bswap_32(*val) + sec->dst_off);
+ else
+ *val += sec->dst_off;
} else {
pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
return -EINVAL;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 19/23] libbpf: Let the static linker carry .bpf_cleanup relocations
2026-10-08 7:51 ` [PATCH bpf-next v9 19/23] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
@ 2026-10-08 8:14 ` sashiko-bot
2026-10-08 16:26 ` Yonghong Song
0 siblings, 1 reply; 38+ messages in thread
From: sashiko-bot @ 2026-10-08 8:14 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] libbpf: OOB array access in linker_append_elf_relos via malformed ELF symbol
--
commit c7a7a96d1dd021ce6e4bcaa13fd9ec618fb1f8cf
Author: Yonghong Song <yonghong.song@linux.dev>
libbpf: Let the static linker carry .bpf_cleanup relocations
Each field of a .bpf_cleanup record is a byte offset into a code section: a
32-bit relocation against that section's symbol, with the offset as the
implicit addend. The patch updates the linker to handle R_BPF_64_NODYLD32
and R_BPF_64_ABS32 relocations against non-executable sections.
> diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
> index f3f71c452f007..c607f51e3cb46 100644
> --- a/tools/lib/bpf/linker.c
> +++ b/tools/lib/bpf/linker.c
[ ... ]
> @@ -2297,6 +2299,39 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
> if (linker->swapped_endian)
> off = bswap_64(off);
> memcpy(ptr, &off, sizeof(off));
> + } else if ((sym_type == R_BPF_64_NODYLD32 ||
> + sym_type == R_BPF_64_ABS32) &&
> + (sec->shdr->sh_flags & SHF_EXECINSTR)) {
[Severity: Medium]
Could this result in an out-of-bounds memory access?
If an ELF file contains a malformed STT_SECTION symbol with an st_shndx
value >= SHN_LORESERVE (such as SHN_ABS), it bypasses the bounds check against
obj->sec_cnt in linker_sanity_check_elf_symtab():
if (sym->st_shndx < SHN_LORESERVE && sym->st_shndx >= obj->sec_cnt) { ... }
When processing relocations against this symbol, linker_append_elf_relos()
indexes the obj->secs array using this out-of-bounds st_shndx:
struct src_sec *sec = &obj->secs[src_sym->st_shndx];
The newly added relocation handling branch for R_BPF_64_NODYLD32 and
R_BPF_64_ABS32 then directly dereferences the out-of-bounds sec->shdr and
sec->dst_off pointers. Can this lead to an out-of-bounds read and a
segmentation fault when linking an untrusted ELF file?
> + /*
> + * A byte offset into a code section,
> + * stored in place. LLVM spells this
> + * relocation NODYLD32 and GNU as
> + * spells it ABS32; being bytes, the
> + * section's new start goes in as it
> + * is, not scaled the way a call's
> + * instruction index is above.
> + *
> + * r_offset is checked only for an
> + * executable section, and SHT_NOBITS
> + * has no raw_data, so bound it here --
> + * subtracting, so it cannot wrap.
> + */
> + if (!dst_linked_sec->raw_data ||
> + dst_linked_sec->sec_sz < (int)sizeof(*val) ||
> + dst_rel->r_offset % sizeof(*val) ||
> + dst_rel->r_offset >
> + (size_t)dst_linked_sec->sec_sz - sizeof(*val)) {
> + pr_warn("ELF relo #%d in section #%zu points outside the data of section '%s' in %s\n",
> + j, src_sec->sec_idx,
> + dst_linked_sec->sec_name,
> + obj->filename);
> + return -EINVAL;
> + }
> + val = dst_linked_sec->raw_data + dst_rel->r_offset;
> + if (linker->swapped_endian)
> + *val = bswap_32(bswap_32(*val) + sec->dst_off);
> + else
> + *val += sec->dst_off;
> } else {
> pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
> return -EINVAL;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261008074959.2993751-1-yonghong.song@linux.dev?part=19
^ permalink raw reply [flat|nested] 38+ messages in thread* Re: [PATCH bpf-next v9 19/23] libbpf: Let the static linker carry .bpf_cleanup relocations
2026-10-08 8:14 ` sashiko-bot
@ 2026-10-08 16:26 ` Yonghong Song
0 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 16:26 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 10/8/26 10:14 AM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Medium] libbpf: OOB array access in linker_append_elf_relos via malformed ELF symbol
> --
>
> commit c7a7a96d1dd021ce6e4bcaa13fd9ec618fb1f8cf
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> libbpf: Let the static linker carry .bpf_cleanup relocations
>
> Each field of a .bpf_cleanup record is a byte offset into a code section: a
> 32-bit relocation against that section's symbol, with the offset as the
> implicit addend. The patch updates the linker to handle R_BPF_64_NODYLD32
> and R_BPF_64_ABS32 relocations against non-executable sections.
>
>> diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
>> index f3f71c452f007..c607f51e3cb46 100644
>> --- a/tools/lib/bpf/linker.c
>> +++ b/tools/lib/bpf/linker.c
> [ ... ]
>
>> @@ -2297,6 +2299,39 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
>> if (linker->swapped_endian)
>> off = bswap_64(off);
>> memcpy(ptr, &off, sizeof(off));
>> + } else if ((sym_type == R_BPF_64_NODYLD32 ||
>> + sym_type == R_BPF_64_ABS32) &&
>> + (sec->shdr->sh_flags & SHF_EXECINSTR)) {
> [Severity: Medium]
> Could this result in an out-of-bounds memory access?
>
> If an ELF file contains a malformed STT_SECTION symbol with an st_shndx
We should not care this since the elf file is *malformed*.
> value >= SHN_LORESERVE (such as SHN_ABS), it bypasses the bounds check against
> obj->sec_cnt in linker_sanity_check_elf_symtab():
>
> if (sym->st_shndx < SHN_LORESERVE && sym->st_shndx >= obj->sec_cnt) { ... }
>
> When processing relocations against this symbol, linker_append_elf_relos()
> indexes the obj->secs array using this out-of-bounds st_shndx:
>
> struct src_sec *sec = &obj->secs[src_sym->st_shndx];
>
> The newly added relocation handling branch for R_BPF_64_NODYLD32 and
> R_BPF_64_ABS32 then directly dereferences the out-of-bounds sec->shdr and
> sec->dst_off pointers. Can this lead to an out-of-bounds read and a
> segmentation fault when linking an untrusted ELF file?
>
>> + /*
>> + * A byte offset into a code section,
>> + * stored in place. LLVM spells this
>> + * relocation NODYLD32 and GNU as
>> + * spells it ABS32; being bytes, the
>> + * section's new start goes in as it
>> + * is, not scaled the way a call's
>> + * instruction index is above.
>> + *
>> + * r_offset is checked only for an
>> + * executable section, and SHT_NOBITS
>> + * has no raw_data, so bound it here --
>> + * subtracting, so it cannot wrap.
>> + */
>> + if (!dst_linked_sec->raw_data ||
>> + dst_linked_sec->sec_sz < (int)sizeof(*val) ||
>> + dst_rel->r_offset % sizeof(*val) ||
>> + dst_rel->r_offset >
>> + (size_t)dst_linked_sec->sec_sz - sizeof(*val)) {
>> + pr_warn("ELF relo #%d in section #%zu points outside the data of section '%s' in %s\n",
>> + j, src_sec->sec_idx,
>> + dst_linked_sec->sec_name,
>> + obj->filename);
>> + return -EINVAL;
>> + }
>> + val = dst_linked_sec->raw_data + dst_rel->r_offset;
>> + if (linker->swapped_endian)
>> + *val = bswap_32(bswap_32(*val) + sec->dst_off);
>> + else
>> + *val += sec->dst_off;
>> } else {
>> pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
>> return -EINVAL;
^ permalink raw reply [flat|nested] 38+ messages in thread
* [PATCH bpf-next v9 20/23] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (18 preceding siblings ...)
2026-10-08 7:51 ` [PATCH bpf-next v9 19/23] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
@ 2026-10-08 7:51 ` Yonghong Song
2026-10-08 7:51 ` [PATCH bpf-next v9 21/23] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
` (2 subsequent siblings)
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:51 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
C has no unwinding, so the frames that own a resource are written in
inline assembly, spelling out what a frontend emits: a call site between
two labels, a landing pad, and a .bpf_cleanup record tying them
together.
The end-to-end test runs a five-frame call chain whose frames hold an
RCU read lock or have preemption disabled, unwinds it from different
frames, and checks which landing pads ran. It also checks that an fentry
program can attach to a subprog an unwind passes through and an fexit
program cannot.
The negative tests cover the shapes the kernel refuses: catch pads,
unsupported instructions or a second unwind in a pad, resumes outside a
pad, callbacks that can unwind, misplaced records, stale stack slots
trusted by a pad, and resources an unwind leaves held or drops twice.
The test skips where the JIT cannot dispatch landing pads.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
.../selftests/bpf/exceptions_cleanup.h | 27 +
.../bpf/prog_tests/exceptions_cleanup.c | 116 +++
.../selftests/bpf/progs/exceptions_cleanup.c | 160 +++
.../bpf/progs/exceptions_cleanup_fail.c | 947 ++++++++++++++++++
.../bpf/progs/exceptions_cleanup_tracing.c | 23 +
5 files changed, 1273 insertions(+)
create mode 100644 tools/testing/selftests/bpf/exceptions_cleanup.h
create mode 100644 tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_tracing.c
diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
new file mode 100644
index 000000000000..96effd2c1361
--- /dev/null
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -0,0 +1,27 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef __EXCEPTIONS_CLEANUP_H__
+#define __EXCEPTIONS_CLEANUP_H__
+
+/* progs/exceptions_cleanup.c: one bit per function that reports it ran. */
+#define RAN_FOO3_PREEMPT 0x1
+#define RAN_FOO2_RCU 0x2
+#define RAN_FOO1V_PREEMPT 0x4
+#define RAN_FOO2_DROP 0x8
+#define RAN_BUMP 0x10
+
+#define CLEANUP_REC(begin, end, landing_pad) \
+ ".pushsection .bpf_cleanup,\"a\",@progbits;" \
+ ".long " begin ";" \
+ ".long " end ";" \
+ ".long " landing_pad ";" \
+ ".popsection;"
+
+/* Set a bit in @pads_ran. */
+#define PAD_RAN(bit) \
+ "r1 = %[pads_ran] ll;" \
+ "r2 = *(u64 *)(r1 + 0);" \
+ "r2 |= " bit ";" \
+ "*(u64 *)(r1 + 0) = r2;"
+
+#endif /* __EXCEPTIONS_CLEANUP_H__ */
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
new file mode 100644
index 000000000000..3e46e17ef9b4
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -0,0 +1,116 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <test_progs.h>
+#include "exceptions_cleanup.h"
+#include "exceptions_cleanup.skel.h"
+#include "exceptions_cleanup_fail.skel.h"
+#include "exceptions_cleanup_tracing.skel.h"
+
+/* foo3 unwound: every frame that has a pad ran it. */
+#define PADS_FOO3_UNWOUND \
+ (RAN_FOO3_PREEMPT | RAN_FOO2_RCU | RAN_FOO1V_PREEMPT | RAN_FOO2_DROP)
+
+/* foo2 unwound after foo3 returned normally: foo3's pad must not run. */
+#define PADS_FOO2_UNWOUND \
+ (RAN_FOO2_RCU | RAN_FOO1V_PREEMPT | RAN_FOO2_DROP)
+
+static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
+ __u64 pads)
+{
+ __u64 ctx = 0;
+ int err;
+
+ LIBBPF_OPTS(bpf_test_run_opts, topts,
+ .ctx_in = &ctx,
+ .ctx_size_in = sizeof(ctx),
+ );
+
+ skel->bss->input = input;
+ skel->bss->pads_ran = 0;
+ skel->bss->result = 0;
+
+ err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.entry), &topts);
+ if (!ASSERT_OK(err, "run"))
+ return;
+ ASSERT_EQ(topts.retval, retval, "retval");
+ /* bump() is not a landing pad; it sets its bit on every run. */
+ ASSERT_EQ(skel->bss->pads_ran, pads | RAN_BUMP, "pads_ran");
+}
+
+/*
+ * Load one tracing program against foo1(), which an unwind passes through.
+ * fexit keeps a trampoline frame between foo1() and its caller that the
+ * unwind cannot walk past, so it is refused; fentry leaves none.
+ */
+static int load_tracer(struct exceptions_cleanup *tgt, bool fexit)
+{
+ struct exceptions_cleanup_tracing *skel;
+ struct bpf_program *prog;
+ int err;
+
+ skel = exceptions_cleanup_tracing__open();
+ if (!ASSERT_OK_PTR(skel, "tracer open"))
+ return -EINVAL;
+ prog = fexit ? skel->progs.fexit_unwinding_subprog :
+ skel->progs.fentry_unwinding_subprog;
+ bpf_program__set_autoload(prog, true);
+ err = bpf_program__set_attach_target(prog, bpf_program__fd(tgt->progs.entry),
+ "foo1");
+ if (!err)
+ err = exceptions_cleanup_tracing__load(skel);
+ exceptions_cleanup_tracing__destroy(skel);
+ return err;
+}
+
+void test_exceptions_cleanup(void)
+{
+ char log[8192] = {};
+
+ LIBBPF_OPTS(bpf_object_open_opts, opts,
+ .kernel_log_buf = log,
+ .kernel_log_size = sizeof(log));
+ struct exceptions_cleanup *skel;
+ int err;
+
+ skel = exceptions_cleanup__open_opts(&opts);
+ if (!ASSERT_OK_PTR(skel, "open"))
+ return;
+
+ err = exceptions_cleanup__load(skel);
+ if (err) {
+ if (err == -EOPNOTSUPP &&
+ strstr(log, "exception cleanup needs a JIT that can dispatch landing pads")) {
+ printf("%s:SKIP:JIT cannot dispatch exception cleanup landing pads\n",
+ __func__);
+ test__skip();
+ } else if (!ASSERT_OK(err, "load")) {
+ fprintf(stderr, "%s", log);
+ }
+ exceptions_cleanup__destroy(skel);
+ return;
+ }
+
+ /* No unwind: foo3 returns 1 ^ 1 == 0, foo2 adds one, no pad runs. */
+ if (test__start_subtest("no_unwind"))
+ run(skel, 1, 1, 0);
+
+ /* foo3 unwinds; every pad runs and entry returns zero. */
+ if (test__start_subtest("unwind_from_foo3"))
+ run(skel, 101, 0, PADS_FOO3_UNWOUND);
+
+ /*
+ * foo3 returns 2 ^ 1 == 3, so foo2 unwinds from its own second region;
+ * foo3's frame is long gone, so its pad must not run.
+ */
+ if (test__start_subtest("unwind_from_foo2"))
+ run(skel, 2, 0, PADS_FOO2_UNWOUND);
+
+ if (test__start_subtest("tracing_unwinding_subprog")) {
+ ASSERT_OK(load_tracer(skel, false), "fentry");
+ ASSERT_EQ(load_tracer(skel, true), -EINVAL, "fexit");
+ }
+
+ exceptions_cleanup__destroy(skel);
+
+ RUN_TESTS(exceptions_cleanup_fail);
+}
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
new file mode 100644
index 000000000000..d065ba53c812
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
@@ -0,0 +1,160 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_unwind();
+ bpf_rcu_read_lock();
+ bpf_rcu_read_unlock();
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume(NULL);
+}
+
+__u64 input = 0;
+__u64 pads_ran = 0;
+__u64 result = 0;
+__u64 never = 0;
+
+static __used __noinline __u64 foo3(__u64 x)
+{
+ bpf_preempt_disable();
+ if (x > 100)
+ asm volatile (
+ "1:" "call bpf_unwind;" /* cleanup region */
+ "2:"
+ "goto 3f;"
+ "4:" /* landing pad */
+ /*
+ * r0 at pad entry is the zero the fixups put after the
+ * bpf_unwind() call. It is kept in a callee-saved
+ * register and handed to the resume, the way a
+ * compiler-emitted pad passes the exception pointer to
+ * _Unwind_Resume. The kfunc takes it and ignores it, and
+ * the two pads below do without the shuffle.
+ */
+ "r7 = r0;"
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "3:"
+ CLEANUP_REC("1b", "2b", "4b")
+ :
+ : [ran]"i"(RAN_FOO3_PREEMPT),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+ bpf_preempt_enable();
+ return x ^ 1;
+}
+
+static __used __naked __noinline void drop_glue(void)
+{
+ asm volatile (
+ PAD_RAN("%[ran]")
+ "exit;"
+ :
+ : [ran]"i"(RAN_FOO2_DROP), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 foo2(void)
+{
+ asm volatile (
+ "r6 = r1;"
+ "call bpf_rcu_read_lock;"
+ "r1 = r6;"
+"1:" "call foo3;" /* cleanup region #1 */
+"2:"
+ "r6 = r0;"
+ "if r6 == 0 goto 5f;"
+"3:" "call bpf_unwind;" /* cleanup region #2 */
+"4:"
+ "r0 = 0;"
+ "exit;"
+"5:"
+ "call bpf_rcu_read_unlock;"
+ "r0 = r6;"
+ "r0 += 1;"
+ "exit;"
+"6:" /* landing pad, shared by both regions */
+ "call drop_glue;"
+ "call bpf_rcu_read_unlock;"
+ PAD_RAN("%[ran_rcu]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "6b")
+ CLEANUP_REC("3b", "4b", "6b")
+ :
+ : [ran_rcu]"i"(RAN_FOO2_RCU),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+static __used __naked __noinline void foo1v(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call foo2;" /* cleanup region */
+"2:"
+ "r6 = r0;"
+ "call bpf_preempt_enable;"
+ "r1 = %[result] ll;"
+ "*(u64 *)(r1 + 0) = r6;"
+ "goto 7f;"
+"8:" /* landing pad */
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "goto 9f;"
+"7:" /* the frame's own exit block */
+ "r0 = 0;"
+ "exit;"
+"9:" /* shared resume block */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "8b")
+ :
+ : [ran]"i"(RAN_FOO1V_PREEMPT), __imm_addr(input),
+ __imm_addr(result), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/*
+ * A frame with no cleanup record: an unwind leaving it runs no pad. The
+ * unwind never fires -- @never is global -- and the bit marks the return path.
+ */
+static __used __naked __noinline void bump(void)
+{
+ asm volatile (
+ PAD_RAN("%[ran]")
+ "r1 = %[never] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 == 0 goto 1f;"
+ "call bpf_unwind;"
+"1:"
+ "exit;" /* r0 deliberately left alone */
+ :
+ : [ran]"i"(RAN_BUMP), __imm_addr(never), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+__noinline __u64 foo1(void)
+{
+ bump();
+ foo1v();
+ return result;
+}
+
+SEC("syscall")
+int entry(void *ctx)
+{
+ return foo1();
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
new file mode 100644
index 000000000000..ed724b85c692
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
@@ -0,0 +1,947 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_experimental.h"
+#include "bpf_misc.h"
+#include "../test_kmods/bpf_testmod_kfunc.h"
+#include "exceptions_cleanup.h"
+
+__u64 input = 0;
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_unwind();
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_rcu_read_lock();
+ bpf_rcu_read_unlock();
+ bpf_unwind_resume(NULL);
+}
+
+/* An unwind raised in a callee, which is how a cleanup region gets one. */
+static __used __naked __noinline __u64 inner_unwind(void)
+{
+ asm volatile (
+ "call bpf_unwind;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static int unwinding_cb(__u32 idx, void *ctx)
+{
+ bpf_unwind();
+ return 0;
+}
+
+/* Gives the program a table; the refusal is at the bpf_loop() call. */
+static __used __naked __noinline __u64 cb_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call unwinding_cb;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("may unwind and is used as a callback")
+int callback_may_unwind(void *ctx)
+{
+ bpf_loop(1, unwinding_cb, NULL, 0);
+ return cb_frame();
+}
+
+/* A second bpf_unwind() from inside a landing pad. */
+static __used __naked __noinline __u64 unwind_in_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad that unwinds again */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("starts a second unwind while one is in flight")
+int unwind_from_landing_pad(void *ctx)
+{
+ return unwind_in_pad_frame();
+}
+
+__noinline int unused_exc_cb(u64 cookie)
+{
+ return 0;
+}
+
+/* A frame with a table and a pad, for tests whose refusal lies elsewhere. */
+static __used __naked __noinline __u64 table_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__exception_cb(unused_exc_cb)
+__failure __msg("cannot be combined with an exception callback")
+int table_with_exception_cb(void *ctx)
+{
+ return table_frame();
+}
+
+/*
+ * A throw and a table, with no callback tagged: the default callback is
+ * appended too late to stand in for the throw, so the program is scanned
+ * for one instead.
+ */
+static __used __naked __noinline __u64 throw_and_table_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("cannot be combined with bpf_throw")
+int table_with_throw(void *ctx)
+{
+ if (input)
+ bpf_throw(0);
+ return throw_and_table_frame();
+}
+
+__u64 never;
+
+/*
+ * A pad calling a global subprogram that can unwind. The subprogram is
+ * verified on its own, so the pad rule is what refuses it.
+ */
+__noinline void pad_callee_that_unwinds(void)
+{
+ if (never)
+ bpf_unwind();
+}
+
+static __used __naked __noinline __u64 pad_calls_unwinder_frame(void)
+{
+ asm volatile (
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call pad_callee_that_unwinds;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("which can unwind while an unwind is in flight")
+int pad_calls_unwinder(void *ctx)
+{
+ return pad_calls_unwinder_frame();
+}
+
+/*
+ * A pad calling a global subprogram that can throw. The throw is refused
+ * wherever it sits: the scan covers the subprograms too, not just the main
+ * program, so this never reaches the rules about pads.
+ */
+__noinline void pad_callee_that_throws(void)
+{
+ if (never)
+ bpf_throw(0);
+}
+
+static __used __naked __noinline __u64 pad_calls_thrower_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call pad_callee_that_throws;" /* ...which can throw: refused */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("cannot be combined with bpf_throw")
+int pad_calls_thrower(void *ctx)
+{
+ return pad_calls_thrower_frame();
+}
+
+static __used __naked __noinline __u64 catch_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* catch pad: no resume, it stops here */
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("ends a landing pad: a catch pad is not supported yet")
+int catch_landing_pad(void *ctx)
+{
+ return catch_pad_frame();
+}
+
+/* A bpf_unwind_resume() outside any landing pad. */
+SEC("?syscall")
+__failure __msg("is not in a landing pad")
+int resume_outside_pad(void *ctx)
+{
+ /* Never taken, but reachable, which is all the verifier needs. */
+ if (never)
+ bpf_unwind_resume(NULL);
+ return table_frame();
+}
+
+/* A bpf_unwind_resume() in a subprogram a landing pad calls. */
+static __used __naked __noinline void resume_in_callee(void)
+{
+ asm volatile (
+ "call bpf_unwind_resume;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 pad_calls_resumer_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call resume_in_callee;" /* ...which resumes: refused */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not in a landing pad")
+int resume_in_pad_callee(void *ctx)
+{
+ return pad_calls_resumer_frame();
+}
+
+static __used __naked __noinline __u64 nested_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_unwind;" /* first cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* first pad, second region's call */
+ "call bpf_preempt_enable;"
+"4:"
+ "call bpf_unwind_resume;"
+ "exit;"
+"5:" /* second pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ CLEANUP_REC("3b", "4b", "5b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is inside the call-site range of")
+int nested_landing_pad(void *ctx)
+{
+ return nested_pad_frame();
+}
+
+/* A tail call in a landing pad: the frame would never reach its resume. */
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(__u32));
+ __uint(value_size, sizeof(__u32));
+} tc_map SEC(".maps");
+
+static __used __naked __noinline __u64 tail_call_pad_frame(void)
+{
+ asm volatile (
+ "r6 = r1;"
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r1 = r6;"
+ "r2 = %[tc_map] ll;"
+ "r3 = 0;"
+ "call %[bpf_tail_call];"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm(bpf_tail_call), __imm_addr(tc_map)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is a tail call, which replaces the frame, and is in a landing pad")
+int tail_call_in_pad(void *ctx)
+{
+ return tail_call_pad_frame();
+}
+
+#if defined(__BPF_FEATURE_STACK_ARGUMENT)
+
+/*
+ * A pad reading an incoming stack argument after its own bpf_unwind(): the
+ * call clobbered the register the JIT keeps it in. Every parameter is used,
+ * so the caller passes them all, and summed before the call, so the pad's is
+ * the only stack argument load after it.
+ */
+static __used __noinline __u64 stack_arg_pad_frame(__u64 a, __u64 b, __u64 c,
+ __u64 d, __u64 e, __u64 f)
+{
+ __u64 sum = a + b + c + d + e + f;
+
+ asm volatile (
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "goto 4f;"
+"3:" /* landing pad */
+ "r0 = *(u64 *)(r11 + 8);"
+ "call bpf_unwind_resume;"
+ "exit;"
+"4:"
+ CLEANUP_REC("1b", "2b", "3b")
+ : "+r"(sum) :: __clobber_common);
+ return sum;
+}
+
+SEC("?syscall")
+__failure __msg("r11 load must be before any r11 store or call insn")
+int stack_arg_load_in_pad(void *ctx)
+{
+ return stack_arg_pad_frame(1, 2, 3, 4, 5, 6);
+}
+
+#endif /* __BPF_FEATURE_STACK_ARGUMENT */
+
+/*
+ * A landing pad entered by ordinary control flow, with no unwind in flight:
+ * that path reaches the pad's resume outside a pad. Only the in_pad
+ * comparison in func_states_equal() keeps it from being pruned against the
+ * pad's own visit of that instruction.
+ */
+static __used __naked __noinline __u64 jump_into_pad_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "if r6 > 7 goto 4f;" /* an ordinary branch into the pad */
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+"4:" /* ... and its second instruction */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not in a landing pad")
+int jump_into_pad(void *ctx)
+{
+ return jump_into_pad_frame();
+}
+
+#if defined(__clang__) && \
+ (defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64))
+
+/*
+ * An indirect jump in a landing pad. A jump table entry is an offset from
+ * the program's section symbol, which has to be spelled in quotes here.
+ */
+SEC("?syscall")
+__failure __msg("is an indirect jump, which nothing a compiler frontend emits in a pad needs, and is in a landing pad")
+__naked void gotox_in_pad(void)
+{
+ asm volatile (
+ ".pushsection .jumptables,\"\",@progbits;"
+"jt0_%=:"
+ ".quad l0_%= - \"?syscall\";"
+ ".quad l1_%= - \"?syscall\";"
+ ".size jt0_%=, 16;"
+ ".global jt0_%=;"
+ ".popsection;"
+
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r1 = jt0_%= ll;"
+ "r1 += 8;"
+ "r2 = *(u64 *)(r1 + 0);"
+ /*
+ * gotox r2, as raw bytes: the mnemonic only reached the LLVM
+ * assembler in llvm 22, and BPF_RAW_INSN() needs <linux/bpf.h>, which
+ * vmlinux.h rules out. dst_reg is the other nibble on a big-endian
+ * target.
+ */
+#if __BYTE_ORDER__ == __ORDER_BIG_ENDIAN__
+ ".byte 0x0d, 0x20, 0, 0, 0, 0, 0, 0;"
+#else
+ ".byte 0x0d, 0x02, 0, 0, 0, 0, 0, 0;"
+#endif
+"l0_%=:"
+ "call bpf_unwind_resume;"
+ "exit;"
+"l1_%=:"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+#endif /* __clang__ && (x86 || arm64) */
+
+/* A BPF_LD_[ABS|IND] in a pad: a failed load leaves without resuming. */
+static __used __naked __noinline __u64 ld_abs_pad_frame(void)
+{
+ asm volatile (
+ "r6 = r1;" /* the skb BPF_LD_ABS reads */
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r0 = *(u32 *)skb[0];"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?tc")
+__failure __msg("is a BPF_LD_[ABS|IND], which can leave through the epilogue")
+__naked void ld_abs_in_pad(void)
+{
+ asm volatile (
+ "call ld_abs_pad_frame;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/*
+ * A subprogram a landing pad calls, which unwinds on its own. The second
+ * unwind would rewrite return addresses the first has already redirected.
+ */
+static __used __naked __noinline __u64 own_pad_callee(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* its landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 pad_calls_own_pad_frame(void)
+{
+ asm volatile (
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad, which calls the above */
+ "call own_pad_callee;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("starts a second unwind while one is in flight")
+int unwind_in_pad_callee(void *ctx)
+{
+ return pad_calls_own_pad_frame();
+}
+
+/* A record whose range holds no call that can unwind. */
+static __used __naked __noinline __u64 nounwind_rec_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call bpf_preempt_enable;" /* cleanup region: nounwind */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad, reached by nothing */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("unreachable insn")
+int nounwind_region(void *ctx)
+{
+ return nounwind_rec_frame();
+}
+
+/*
+ * An unwind with no landing pad in the main program returns from it, so what
+ * it holds is checked there, as at a plain exit.
+ */
+SEC("?syscall")
+__failure __msg("BPF_EXIT instruction in main prog cannot be used inside bpf_rcu_read_lock-ed region")
+int unwind_no_pad_rcu(void *ctx)
+{
+ bpf_rcu_read_lock();
+ bpf_unwind();
+ bpf_rcu_read_unlock();
+ return 0;
+}
+
+/*
+ * Lock and reference state is the program's, not a frame's: an unwind carries
+ * it through every pad it runs, and what is still held when the main program
+ * returns is refused there. A pad that drops what its caller holds is fine;
+ * the caller's own pad dropping it again is not.
+ */
+static __used __naked __noinline __u64 pad_drops_caller_lock_frame(void)
+{
+ asm volatile (
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad: drops a lock it never took */
+ "call bpf_rcu_read_unlock;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 caller_holds_lock_frame(void)
+{
+ asm volatile (
+ "call bpf_rcu_read_lock;"
+"1:" "call pad_drops_caller_lock_frame;"
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad */
+ "call bpf_rcu_read_unlock;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("unmatched rcu read unlock")
+int pad_drops_caller_lock(void *ctx)
+{
+ return caller_holds_lock_frame();
+}
+
+/* The other way round: a pad that does not drop what its own frame took. */
+static __used __naked __noinline __u64 pad_keeps_own_lock_frame(void)
+{
+ asm volatile (
+ "call bpf_rcu_read_lock;"
+"1:" "call inner_unwind;" /* cleanup region */
+"2:"
+ "call bpf_rcu_read_unlock;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad: forgets the unlock */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("BPF_EXIT instruction in main prog cannot be used inside bpf_rcu_read_lock-ed region")
+int pad_keeps_own_lock(void *ctx)
+{
+ return pad_keeps_own_lock_frame();
+}
+
+/* And a subprog with no pad at all, leaving through an unwind holding one. */
+static __used __naked __noinline __u64 no_pad_keeps_own_lock_frame(void)
+{
+ asm volatile (
+ "call bpf_rcu_read_lock;"
+ "call bpf_unwind;" /* no record covers it */
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("BPF_EXIT instruction in main prog cannot be used inside bpf_rcu_read_lock-ed region")
+int no_pad_keeps_own_lock(void *ctx)
+{
+ return no_pad_keeps_own_lock_frame();
+}
+
+struct {
+ __uint(type, BPF_MAP_TYPE_RINGBUF);
+ __uint(max_entries, 4096);
+} unwind_ringbuf SEC(".maps");
+
+/*
+ * Always unwinds, so its caller is never returned to on the modelled path,
+ * and the caller's post-call code, which never releases the record, is not
+ * walked. Its pad discards the record the caller reserved.
+ */
+static __used __naked __noinline __u64 pad_drops_caller_ref_frame(void)
+{
+ asm volatile (
+ "r6 = r1;" /* the caller's reserved record */
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad: drops what it never acquired */
+ "r1 = r6;"
+ "r2 = 0;"
+ "call %[bpf_ringbuf_discard];"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm(bpf_ringbuf_discard)
+ : __clobber_all);
+}
+
+/* And this frame's own pad drops it a second time. */
+static __used __naked __noinline __u64 caller_holds_ref_frame(void)
+{
+ asm volatile (
+ "r1 = %[unwind_ringbuf] ll;"
+ "r2 = 8;"
+ "r3 = 0;"
+ "call %[bpf_ringbuf_reserve];"
+ "if r0 == 0 goto 9f;"
+ "r6 = r0;"
+ "r1 = r6;"
+"1:" "call pad_drops_caller_ref_frame;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad */
+ "r1 = r6;"
+ "r2 = 0;"
+ "call %[bpf_ringbuf_discard];"
+ "call bpf_unwind_resume;"
+ "exit;"
+"9:"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm(bpf_ringbuf_reserve), __imm(bpf_ringbuf_discard),
+ __imm_addr(unwind_ringbuf)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("R1 type=scalar expected=ringbuf_mem")
+int pad_drops_caller_ref(void *ctx)
+{
+ return caller_holds_ref_frame();
+}
+
+/*
+ * A frame an unwind returns through without a pad is abandoned where it made
+ * the call: the JIT sends it to its epilogue, so nothing of it runs again and
+ * whatever it acquired is never released -- which the main program's exit
+ * then finds still held.
+ */
+static __used __naked __noinline __u64 pad_resumes_frame(void)
+{
+ asm volatile (
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 uncovered_holds_lock_frame(void)
+{
+ asm volatile (
+ "call bpf_rcu_read_lock;"
+ "call pad_resumes_frame;" /* no record covers this call */
+ "call bpf_rcu_read_unlock;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("BPF_EXIT instruction in main prog cannot be used inside bpf_rcu_read_lock-ed region")
+int unwind_through_call_keeps_lock(void *ctx)
+{
+ return uncovered_holds_lock_frame();
+}
+
+static __used __naked __noinline __u64 uncovered_holds_ref_frame(void)
+{
+ asm volatile (
+ "r1 = %[unwind_ringbuf] ll;"
+ "r2 = 8;"
+ "r3 = 0;"
+ "call %[bpf_ringbuf_reserve];"
+ "if r0 == 0 goto 9f;"
+ "r6 = r0;"
+ "call pad_resumes_frame;" /* no record covers this call */
+ "r1 = r6;"
+ "r2 = 0;"
+ "call %[bpf_ringbuf_discard];"
+"9:"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_ringbuf_reserve), __imm(bpf_ringbuf_discard),
+ __imm_addr(unwind_ringbuf)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("Unreleased reference id=")
+int unwind_through_call_keeps_ref(void *ctx)
+{
+ return uncovered_holds_ref_frame();
+}
+
+/* The main program's frame is passed by the same way. */
+SEC("?syscall")
+__failure __msg("Unreleased reference id=")
+int unwind_through_call_main_keeps_ref(void *ctx)
+{
+ void *rec;
+
+ rec = bpf_ringbuf_reserve(&unwind_ringbuf, 8, 0);
+ if (!rec)
+ return 0;
+ pad_resumes_frame(); /* no record covers this call */
+ bpf_ringbuf_discard(rec, 0);
+ return 0;
+}
+
+/*
+ * A callee writes its caller's stack through a pointer argument, then
+ * unwinds. The caller's pad runs after that write, so it cannot keep trusting
+ * the slot to hold the zero it held at the call.
+ */
+static __used __naked __noinline __u64 stack_writer(void)
+{
+ asm volatile (
+ "r2 = 0x10000000;"
+ "*(u64 *)(r1 + 0) = r2;" /* r1 is the caller's fp-8 */
+ "call bpf_unwind;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 stale_stack_frame(void)
+{
+ asm volatile (
+ "r6 = 0;"
+ "*(u64 *)(r10 - 8) = r6;"
+ "*(u64 *)(r10 - 64) = r6;"
+ "r1 = r10;"
+ "r1 += -8;"
+"1:" "call stack_writer;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad: fp-8 as an offset into fp-64 */
+ "r1 = *(u64 *)(r10 - 8);"
+ "r2 = r10;"
+ "r2 += -64;"
+ "r2 += r1;"
+ "r0 = *(u8 *)(r2 + 0);"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("invalid read from stack R2 off=268435392 size=1")
+int stale_stack_pad(void *ctx)
+{
+ return stale_stack_frame();
+}
+
+/* The same through a global subprog, which is not walked from its caller. */
+__noinline int global_stack_writer(__u64 *p)
+{
+ if (!p)
+ return 0;
+ *p = 0x10000000;
+ bpf_unwind();
+ return 0;
+}
+
+static __used __naked __noinline __u64 global_stale_stack_frame(void)
+{
+ asm volatile (
+ "r6 = 0;"
+ "*(u64 *)(r10 - 8) = r6;"
+ "*(u64 *)(r10 - 64) = r6;"
+ "r1 = r10;"
+ "r1 += -8;"
+"1:" "call global_stack_writer;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad: fp-8 as an offset into fp-64 */
+ "r1 = *(u64 *)(r10 - 8);"
+ "r2 = r10;"
+ "r2 += -64;"
+ "r2 += r1;"
+ "r0 = *(u8 *)(r2 + 0);"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("math between fp pointer and register with unbounded min value")
+int global_stale_stack_pad(void *ctx)
+{
+ return global_stale_stack_frame();
+}
+
+/* gcc has no indirect calls, and only these JITs emit them */
+#if defined(__clang__) && \
+ (defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64))
+
+/*
+ * A callback calling an unwinding subprog through a pointer it read from its
+ * caller's stack: it neither calls nor takes the address of inner_unwind, so
+ * only the marking of every subprog with a callx makes it might_unwind.
+ */
+static __used __naked __noinline int callx_cb(void)
+{
+ asm volatile (
+ "r1 = *(u64 *)(r2 + 0);"
+ "callx r1;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("may unwind and is used as a callback")
+__naked int callback_callx_may_unwind(void)
+{
+ asm volatile (
+ "r1 = %[inner_unwind] ll;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "r1 = 1;"
+ "r2 = %[callx_cb] ll;"
+ "r3 = r10;"
+ "r3 += -8;"
+ "r4 = 0;"
+ "call %[bpf_loop];"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm_addr(inner_unwind), __imm_addr(callx_cb), __imm(bpf_loop)
+ : __clobber_all);
+}
+
+#endif /* __clang__ && (x86 || arm64) */
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_tracing.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_tracing.c
new file mode 100644
index 000000000000..6216b09d4472
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_tracing.c
@@ -0,0 +1,23 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+
+/*
+ * Tracing programs aimed at foo1() of progs/exceptions_cleanup.c, a subprog
+ * an unwind passes through. The attach target is set at load time.
+ */
+SEC("?fexit")
+int BPF_PROG(fexit_unwinding_subprog)
+{
+ return 0;
+}
+
+SEC("?fentry")
+int BPF_PROG(fentry_unwinding_subprog)
+{
+ return 0;
+}
+
+char _license[] SEC("license") = "GPL";
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 21/23] selftests/bpf: Add __set_global() and __ret_global() test tags
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (19 preceding siblings ...)
2026-10-08 7:51 ` [PATCH bpf-next v9 20/23] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
@ 2026-10-08 7:51 ` Yonghong Song
2026-10-08 7:51 ` [PATCH bpf-next v9 22/23] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
2026-10-08 7:51 ` [PATCH bpf-next v9 23/23] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:51 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A test that looks at a program's state after it ran needs a driver of
its own, though the loader already runs programs for __retval(). What is
missing is reaching their globals:
- __set_global(var, value) writes one before the run, __ret_global(var,
value) checks one after it. Both make the loader run the program;
without __retval() it must return 0. Each may be given several times,
up to MAX_GLOBAL_VARS.
- The variable is found the way veristat finds one, in ".bss" or
".data", and must be a four or eight byte int or enum.
- A value is held to its variable's range and narrowed to its width,
signedness taken from BTF as in veristat, so a negative literal means
the same to both tags.
- A value may be a '|' separated list, so a bitmask reads as it is
written in the program.
Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/testing/selftests/bpf/progs/bpf_misc.h | 12 +
tools/testing/selftests/bpf/test_loader.c | 319 ++++++++++++++++++-
2 files changed, 330 insertions(+), 1 deletion(-)
diff --git a/tools/testing/selftests/bpf/progs/bpf_misc.h b/tools/testing/selftests/bpf/progs/bpf_misc.h
index f3dbc3b59bff..a80798a8f28f 100644
--- a/tools/testing/selftests/bpf/progs/bpf_misc.h
+++ b/tools/testing/selftests/bpf/progs/bpf_misc.h
@@ -103,6 +103,16 @@
* - POINTER_VALUE
* - TEST_DATA_LEN
* __retval_unpriv Same, but load program in unprivileged mode.
+ * __set_global Set a global variable of the program to a value before
+ * executing it. The variable has to be an int or an enum,
+ * live in .bss or .data and be four or eight bytes wide,
+ * and the value is a number or several joined by '|',
+ * without parentheses.
+ * __ret_global Execute the program and check that a global variable
+ * holds the given value afterwards, under the same rules.
+ * __set_global and __ret_global both make the loader run
+ * the program and check its return value; without
+ * __retval, the program must return 0.
*
* __description Text to be used for display and as an additional filter
* alias, while the original program name stays matchable.
@@ -161,6 +171,8 @@
#define __flag(flag) __test_tag("test_prog_flags=" #flag)
#define __retval(val) __test_tag("test_retval=" XSTR(val))
#define __retval_unpriv(val) __test_tag("test_retval_unpriv=" XSTR(val))
+#define __set_global(var, val) __test_tag("test_global_set=" #var ":" XSTR(val))
+#define __ret_global(var, val) __test_tag("test_global_ret=" #var ":" XSTR(val))
#define __auxiliary __test_tag("test_auxiliary")
#define __auxiliary_unpriv __test_tag("test_auxiliary_unpriv")
#define __btf_path(path) __test_tag("test_btf_path=" path)
diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
index a89890cd56d8..9b74b18ce538 100644
--- a/tools/testing/selftests/bpf/test_loader.c
+++ b/tools/testing/selftests/bpf/test_loader.c
@@ -50,6 +50,13 @@ enum stack_mode {
SMALL_STACK = 1 << 1,
};
+#define MAX_GLOBAL_VARS 8
+
+struct global_var {
+ char *name;
+ __u64 val;
+};
+
struct test_subspec {
char *name;
char *description;
@@ -62,6 +69,10 @@ struct test_subspec {
int retval;
bool execute;
__u64 caps;
+ struct global_var set_globals[MAX_GLOBAL_VARS];
+ int set_global_cnt;
+ struct global_var ret_globals[MAX_GLOBAL_VARS];
+ int ret_global_cnt;
};
struct test_spec {
@@ -115,6 +126,34 @@ void free_msgs(struct expected_msgs *msgs)
msgs->cnt = 0;
}
+static void free_global_vars(struct global_var *vars, int *cnt)
+{
+ int i;
+
+ for (i = 0; i < *cnt; i++) {
+ free(vars[i].name);
+ vars[i].name = NULL;
+ }
+ *cnt = 0;
+}
+
+static int clone_global_vars(struct global_var *dst, int *dst_cnt,
+ const struct global_var *src, int src_cnt)
+{
+ int i;
+
+ for (i = 0; i < src_cnt; i++) {
+ dst[i].name = strdup(src[i].name);
+ if (!dst[i].name) {
+ free_global_vars(dst, dst_cnt);
+ return -ENOMEM;
+ }
+ dst[i].val = src[i].val;
+ (*dst_cnt)++;
+ }
+ return 0;
+}
+
static void free_test_spec(struct test_spec *spec)
{
/* Deallocate expect_msgs arrays. */
@@ -129,6 +168,11 @@ static void free_test_spec(struct test_spec *spec)
free_msgs(&spec->unpriv.stdout);
free_msgs(&spec->priv.stdout);
+ free_global_vars(spec->priv.set_globals, &spec->priv.set_global_cnt);
+ free_global_vars(spec->priv.ret_globals, &spec->priv.ret_global_cnt);
+ free_global_vars(spec->unpriv.set_globals, &spec->unpriv.set_global_cnt);
+ free_global_vars(spec->unpriv.ret_globals, &spec->unpriv.ret_global_cnt);
+
free(spec->priv.name);
free(spec->priv.description);
free(spec->unpriv.name);
@@ -311,6 +355,236 @@ static int parse_caps(const char *str, __u64 *val, const char *name)
return 0;
}
+static int parse_global_var(const char *str, struct global_var *vars, int *cnt,
+ const char *name)
+{
+ const char *colon = strrchr(str, ':');
+ struct global_var *var;
+ char *end;
+ __u64 *val;
+
+ if (!colon || colon == str) {
+ PRINT_FAIL("expecting '<variable>:<value>' for %s, got '%s'\n", name, str);
+ return -EINVAL;
+ }
+ if (*cnt >= MAX_GLOBAL_VARS) {
+ PRINT_FAIL("too many %s tags, at most %d are supported\n",
+ name, MAX_GLOBAL_VARS);
+ return -E2BIG;
+ }
+
+ var = &vars[*cnt];
+ val = &var->val;
+ *val = 0;
+ for (const char *term = colon + 1;;) {
+ __u64 v;
+
+ errno = 0;
+ v = strtoull(term, &end, 0);
+ if (errno || end == term) {
+ PRINT_FAIL("failed to parse %s value '%s'\n", name, colon + 1);
+ return -EINVAL;
+ }
+ *val |= v;
+ while (*end == ' ')
+ end++;
+ if (!*end)
+ break;
+ if (*end != '|') {
+ PRINT_FAIL("failed to parse %s value '%s'\n", name, colon + 1);
+ return -EINVAL;
+ }
+ term = end + 1;
+ }
+
+ var->name = strndup(str, colon - str);
+ if (!var->name) {
+ PRINT_FAIL("failed to allocate %s variable name\n", name);
+ return -ENOMEM;
+ }
+ (*cnt)++;
+
+ return 0;
+}
+
+/* As veristat's is_signed_type(): anything not plainly unsigned is signed. */
+static bool global_var_is_signed(const struct btf_type *t)
+{
+ if (btf_is_int(t))
+ return btf_int_encoding(t) & BTF_INT_SIGNED;
+ if (btf_is_any_enum(t))
+ return btf_kflag(t);
+ return true;
+}
+
+static int find_global_var(struct bpf_object *obj, const char *name,
+ struct bpf_map **map, __u32 *off, __u32 *sz,
+ bool *is_signed)
+{
+ static const char * const secs[] = { ".bss", ".data" };
+ struct btf *btf = bpf_object__btf(obj);
+ int i, s;
+
+ if (!btf) {
+ PRINT_FAIL("no BTF for object\n");
+ return -ENOENT;
+ }
+
+ for (s = 0; s < ARRAY_SIZE(secs); s++) {
+ const struct btf_type *sec, *vt;
+ const struct btf_var_secinfo *vsi;
+ struct bpf_map *m = bpf_object__find_map_by_name(obj, secs[s]);
+ int id;
+
+ id = btf__find_by_name_kind(btf, secs[s], BTF_KIND_DATASEC);
+ if (!m || id < 0)
+ continue;
+
+ sec = btf__type_by_id(btf, id);
+ vsi = btf_var_secinfos(sec);
+ for (i = 0; i < btf_vlen(sec); i++, vsi++) {
+ const struct btf_type *var = btf__type_by_id(btf, vsi->type);
+
+ if (strcmp(btf__name_by_offset(btf, var->name_off), name))
+ continue;
+ if (vsi->size != 4 && vsi->size != 8) {
+ PRINT_FAIL("'%s' is %u bytes, only 4 and 8 are supported\n",
+ name, vsi->size);
+ return -EINVAL;
+ }
+ vt = btf__type_by_id(btf, btf__resolve_type(btf, var->type));
+ if (!vt || !(btf_is_int(vt) || btf_is_any_enum(vt))) {
+ PRINT_FAIL("'%s' is not an int or an enum\n", name);
+ return -EINVAL;
+ }
+ *is_signed = global_var_is_signed(vt);
+ *map = m;
+ *off = vsi->offset;
+ *sz = vsi->size;
+ return 0;
+ }
+ }
+
+ PRINT_FAIL("no global variable '%s'\n", name);
+ return -ENOENT;
+}
+
+/*
+ * A tag's value is parsed as 64 bits, but the variable may be narrower and
+ * may be signed. Hold it to the range the variable can represent, the way
+ * veristat's set_global_var() does, and narrow it to what is stored.
+ */
+static int fit_global_var(const char *name, __u32 sz, bool is_signed, __u64 *val)
+{
+ long long v = (long long)*val;
+ long long max_val;
+ __u32 bits;
+
+ if (sz >= sizeof(*val))
+ return 0;
+ bits = sz * 8 - (is_signed ? 1 : 0);
+ max_val = 1ll << bits;
+ if (v >= max_val || v < (is_signed ? -max_val : 0)) {
+ PRINT_FAIL("value %lld for '%s' is out of range [%lld; %lld]\n",
+ v, name, is_signed ? -max_val : 0, max_val - 1);
+ return -EINVAL;
+ }
+ *val = (__u32)*val;
+ return 0;
+}
+
+/* The value of @map's single element, which the caller frees. */
+static void *global_data(struct bpf_map *map, const char *name, size_t *vsz)
+{
+ __u32 zero = 0;
+ void *buf;
+ int err;
+
+ *vsz = bpf_map__value_size(map);
+ buf = calloc(1, *vsz);
+ if (!buf) {
+ PRINT_FAIL("failed to allocate %zu bytes for '%s'\n", *vsz, name);
+ return NULL;
+ }
+ err = bpf_map__lookup_elem(map, &zero, sizeof(zero), buf, *vsz, 0);
+ if (err) {
+ PRINT_FAIL("failed to read '%s': %d\n", name, err);
+ free(buf);
+ return NULL;
+ }
+ return buf;
+}
+
+static int read_global_var(struct bpf_map *map, const char *name, __u32 off,
+ __u32 sz, __u64 *val)
+{
+ size_t vsz;
+ void *buf;
+
+ buf = global_data(map, name, &vsz);
+ if (!buf)
+ return -EINVAL;
+ *val = sz == 4 ? *(__u32 *)(buf + off) : *(__u64 *)(buf + off);
+ free(buf);
+ return 0;
+}
+
+static int write_global_var(struct bpf_map *map, const char *name, __u32 off,
+ __u32 sz, __u64 val)
+{
+ __u32 zero = 0;
+ size_t vsz;
+ void *buf;
+ int err;
+
+ buf = global_data(map, name, &vsz);
+ if (!buf)
+ return -EINVAL;
+ if (sz == 4)
+ *(__u32 *)(buf + off) = val;
+ else
+ *(__u64 *)(buf + off) = val;
+ err = bpf_map__update_elem(map, &zero, sizeof(zero), buf, vsz, 0);
+ if (err)
+ PRINT_FAIL("failed to write '%s': %d\n", name, err);
+ free(buf);
+ return err;
+}
+
+/* Write a __set_global() value into the program's global variable. */
+static int set_global_var(struct bpf_object *obj, const struct global_var *var)
+{
+ __u64 val = var->val;
+ struct bpf_map *map;
+ __u32 off, sz;
+ bool is_signed;
+
+ if (find_global_var(obj, var->name, &map, &off, &sz, &is_signed) ||
+ fit_global_var(var->name, sz, is_signed, &val))
+ return -EINVAL;
+ return write_global_var(map, var->name, off, sz, val);
+}
+
+/* Check the program's global variable against a __ret_global() value. */
+static int check_global_var(struct bpf_object *obj, const struct global_var *var)
+{
+ __u64 want = var->val, val;
+ struct bpf_map *map;
+ __u32 off, sz;
+ bool is_signed;
+
+ if (find_global_var(obj, var->name, &map, &off, &sz, &is_signed) ||
+ fit_global_var(var->name, sz, is_signed, &want) ||
+ read_global_var(map, var->name, off, sz, &val))
+ return -EINVAL;
+ if (val != want) {
+ PRINT_FAIL("Unexpected %s: 0x%llx != 0x%llx\n", var->name,
+ (unsigned long long)val, (unsigned long long)want);
+ return -EINVAL;
+ }
+ return 0;
+}
+
static int parse_retval(const char *str, int *val, const char *name)
{
/*
@@ -557,6 +831,22 @@ static int parse_test_spec(struct test_loader *tester,
spec->mode_mask |= UNPRIV;
spec->unpriv.execute = true;
has_unpriv_retval = true;
+ } else if ((val = str_has_pfx(s, "test_global_set="))) {
+ err = parse_global_var(val, spec->priv.set_globals,
+ &spec->priv.set_global_cnt,
+ "__set_global");
+ if (err)
+ goto cleanup;
+ spec->priv.execute = true;
+ spec->mode_mask |= PRIV;
+ } else if ((val = str_has_pfx(s, "test_global_ret="))) {
+ err = parse_global_var(val, spec->priv.ret_globals,
+ &spec->priv.ret_global_cnt,
+ "__ret_global");
+ if (err)
+ goto cleanup;
+ spec->priv.execute = true;
+ spec->mode_mask |= PRIV;
} else if ((val = str_has_pfx(s, "test_log_level="))) {
err = parse_int(val, &spec->log_level, "test log level");
if (err)
@@ -744,6 +1034,23 @@ static int parse_test_spec(struct test_loader *tester,
spec->unpriv.execute = spec->priv.execute;
}
+ if (spec->priv.set_global_cnt && !spec->unpriv.set_global_cnt) {
+ err = clone_global_vars(spec->unpriv.set_globals,
+ &spec->unpriv.set_global_cnt,
+ spec->priv.set_globals,
+ spec->priv.set_global_cnt);
+ if (err)
+ goto cleanup;
+ }
+ if (spec->priv.ret_global_cnt && !spec->unpriv.ret_global_cnt) {
+ err = clone_global_vars(spec->unpriv.ret_globals,
+ &spec->unpriv.ret_global_cnt,
+ spec->priv.ret_globals,
+ spec->priv.ret_global_cnt);
+ if (err)
+ goto cleanup;
+ }
+
if (spec->unpriv.expect_msgs.cnt == 0)
clone_msgs(&spec->priv.expect_msgs, &spec->unpriv.expect_msgs);
if (spec->unpriv.expect_xlated.cnt == 0)
@@ -1358,7 +1665,7 @@ void run_subtest(struct test_loader *tester,
struct cap_state caps = {};
struct bpf_object *tobj;
struct bpf_map *map;
- int retval, err, i;
+ int retval, err, i, j;
int links_cnt = 0;
bool should_load;
@@ -1534,6 +1841,11 @@ void run_subtest(struct test_loader *tester,
}
}
+ for (j = 0; j < subspec->set_global_cnt; j++) {
+ if (set_global_var(tobj, &subspec->set_globals[j]))
+ goto tobj_cleanup;
+ }
+
err = do_prog_test_run(bpf_program__fd(tprog), &retval,
bpf_program__type(tprog) == BPF_PROG_TYPE_SYSCALL ? true : false,
spec->linear_sz);
@@ -1542,6 +1854,11 @@ void run_subtest(struct test_loader *tester,
goto tobj_cleanup;
}
+ for (j = 0; j < subspec->ret_global_cnt; j++) {
+ if (check_global_var(tobj, &subspec->ret_globals[j]))
+ goto tobj_cleanup;
+ }
+
verify_stderr(bpf_program__fd(tprog), &subspec->stderr);
if (subspec->stdout.cnt) {
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 22/23] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (20 preceding siblings ...)
2026-10-08 7:51 ` [PATCH bpf-next v9 21/23] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
@ 2026-10-08 7:51 ` Yonghong Song
2026-10-08 7:51 ` [PATCH bpf-next v9 23/23] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:51 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Add the shapes the end-to-end test does not reach:
- a pad reading its frame's callee-saved registers
- a region ending on a 16-byte instruction
- a pad ending in _Unwind_Resume rather than bpf_unwind_resume
- a pad whose first instruction is a nop
- a pad indexing its frame by a register set before the call
- two pads with an uncovered frame between them
- precision chains from a pad back across the resume that led to it,
and back into the frame the unwind left
- a frame above an unwind that never returns, whose kept exit is not
its last instruction
- a region covering an indirect call, and a subprog calling an
unwinding one through a pointer it was handed
- a frame holding a reference across a covered call, released by its
pad
- a pad reading a slot its static or global callee wrote before
unwinding
- unwinds out of a global subprog called with no record, one returning
in R0:R2 among them
- a reference moved into a callee, as rustc does for a by-value
argument
- a bpf_unwind() inside a loop, in a subprog and in main
- a frame's own pad finding r0 zero after its bpf_unwind()
- an unwind reaching main's exit in a program type that checks r0, out
of a static callee and out of a global call
- a pad touching a global of each width and sign a tag can name
- two speculative walks, into a pad and to an exit in one, loaded
without CAP_PERFMON; each checks the translated program for the
barrier, so it cannot pass without the walk
Each shape that runs is driven by __set_global(), __retval() and
__ret_global() through RUN_TESTS().
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
.../selftests/bpf/exceptions_cleanup.h | 20 +
.../bpf/prog_tests/exceptions_cleanup.c | 2 +
.../bpf/progs/exceptions_cleanup_shapes.c | 1141 +++++++++++++++++
3 files changed, 1163 insertions(+)
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
index 96effd2c1361..caa8e7b0cfb0 100644
--- a/tools/testing/selftests/bpf/exceptions_cleanup.h
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -10,6 +10,26 @@
#define RAN_FOO2_DROP 0x8
#define RAN_BUMP 0x10
+/* progs/exceptions_cleanup_shapes.c: one bit per landing pad. */
+#define RAN_REGS 0x1
+#define RAN_WIDE_REC 0x2
+#define RAN_RESUME_ALIAS 0x4
+#define RAN_NOP_PAD 0x8
+#define RAN_VAR_STACK 0x10
+#define RAN_GAP_INNER 0x20
+#define RAN_GAP_OUTER 0x40
+#define RAN_PREC_RESUME 0x80
+#define RAN_NO_EXIT_JA 0x100
+#define RAN_CALLX 0x200
+#define RAN_HELD_REF 0x400
+#define RAN_CALLEE_WRITE 0x800
+#define RAN_CALLEE_OFFSET 0x1000
+#define RAN_GLOBAL_WRITE 0x2000
+#define RAN_THROUGH_GLOBAL 0x4000
+#define RAN_OWN_PAD_R0 0x8000
+#define RAN_PAIR_GLOBAL 0x10000
+#define RAN_MOVE 0x20000
+
#define CLEANUP_REC(begin, end, landing_pad) \
".pushsection .bpf_cleanup,\"a\",@progbits;" \
".long " begin ";" \
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
index 3e46e17ef9b4..7c5dffd4b861 100644
--- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -4,6 +4,7 @@
#include "exceptions_cleanup.h"
#include "exceptions_cleanup.skel.h"
#include "exceptions_cleanup_fail.skel.h"
+#include "exceptions_cleanup_shapes.skel.h"
#include "exceptions_cleanup_tracing.skel.h"
/* foo3 unwound: every frame that has a pad ran it. */
@@ -113,4 +114,5 @@ void test_exceptions_cleanup(void)
exceptions_cleanup__destroy(skel);
RUN_TESTS(exceptions_cleanup_fail);
+ RUN_TESTS(exceptions_cleanup_shapes);
}
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
new file mode 100644
index 000000000000..183a8c8b6b13
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
@@ -0,0 +1,1141 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_unwind();
+ bpf_rcu_read_lock();
+ bpf_rcu_read_unlock();
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume(NULL);
+}
+
+__u64 input = 0;
+__u64 magic = 0x5eed;
+__u64 pads_ran = 0;
+
+/* Load r6-r9 with values derived from @magic. */
+#define LOAD_MAGIC_REGS \
+ "r1 = %[magic] ll;" \
+ "r6 = *(u64 *)(r1 + 0);" \
+ "r7 = r6;" \
+ "r7 += 1;" \
+ "r8 = r6;" \
+ "r8 += 2;" \
+ "r9 = r6;" \
+ "r9 += 3;"
+
+/* Set @bit only if r6-r9 still hold what LOAD_MAGIC_REGS put there. */
+#define CHECK_MAGIC_REGS(bit) \
+ "r1 = %[magic] ll;" \
+ "r2 = *(u64 *)(r1 + 0);" \
+ "if r6 != r2 goto 9f;" \
+ "r2 += 1;" \
+ "if r7 != r2 goto 9f;" \
+ "r2 += 1;" \
+ "if r8 != r2 goto 9f;" \
+ "r2 += 1;" \
+ "if r9 != r2 goto 9f;" \
+ PAD_RAN(bit) \
+ "9:"
+
+/* A callee that unwinds when its argument is over 100. */
+static __used __noinline __u64 pc_unwinder(__u64 x)
+{
+ if (x > 100)
+ bpf_unwind();
+ return x + 1;
+}
+
+static __used __naked __noinline __u64 regs_unwinder(void)
+{
+ asm volatile (
+ /* Not this frame's to keep, and that is the point. */
+ "r6 = 0xdead;"
+ "r7 = 0xbeef;"
+ "r8 = 0xcafe;"
+ "r9 = 0xf00d;"
+ "call bpf_unwind;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* A pad that reads r6-r9, which the callee overwrote before it unwound. */
+static __used __naked __noinline __u64 regs_frame(void)
+{
+ asm volatile (
+ LOAD_MAGIC_REGS
+ "call bpf_preempt_disable;"
+"1:" "call regs_unwinder;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ CHECK_MAGIC_REGS("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_REGS),
+ __imm_addr(magic), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__ret_global(pads_ran, RAN_REGS)
+int entry_regs(void *ctx)
+{
+ return regs_frame();
+}
+
+/* A region ending on a 16-byte insn, so end - 1 names its second half. */
+static __used __naked __noinline __u64 wide_rec_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "call bpf_rcu_read_lock;"
+ "r1 = r6;"
+"1:" "call pc_unwinder;" /* cleanup region begins */
+ "r1 = %[magic] ll;" /* ... and ends on this pair */
+"2:"
+ "r6 = r0;"
+ "call bpf_rcu_read_unlock;"
+ "r0 = r6;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+ "call bpf_rcu_read_unlock;"
+ PAD_RAN("%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_WIDE_REC), __imm_addr(input), __imm_addr(magic),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_WIDE_REC)
+int entry_wide_rec(void *ctx)
+{
+ return wide_rec_frame();
+}
+
+/* The name LLVM gives the resume: _Unwind_Resume(), which libbpf maps over. */
+extern void _Unwind_Resume(void *ptr) __ksym;
+
+static __used __noinline void __resume_alias_btf_anchor(void)
+{
+ _Unwind_Resume(NULL);
+}
+
+static __used __naked __noinline __u64 resume_alias_frame(void)
+{
+ asm volatile (
+"1:" "call regs_unwinder;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ PAD_RAN("%[ran]")
+ "call _Unwind_Resume;" /* the frontend's name for it */
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_RESUME_ALIAS), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__ret_global(pads_ran, RAN_RESUME_ALIAS)
+int entry_resume_alias(void *ctx)
+{
+ return resume_alias_frame();
+}
+
+/* A pad starting on a nop, which opt_remove_nops() drops after the walk. */
+static __used __naked __noinline __u64 nop_pad_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "if r6 < 101 goto 6f;"
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+"6:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad: a nop, then its body */
+ "goto +0;"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_NOP_PAD),
+ __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_NOP_PAD)
+int entry_nop_pad(void *ctx)
+{
+ return nop_pad_frame();
+}
+
+/* Put @magic in both of the slots a variable offset could name. */
+#define FILL_MAGIC_SLOTS \
+ "r1 = %[magic] ll;" \
+ "r1 = *(u64 *)(r1 + 0);" \
+ "*(u64 *)(r10 - 8) = r1;" \
+ "*(u64 *)(r10 - 16) = r1;"
+
+/* Set @bit if the slot @idx names, read at a variable offset, holds it. */
+#define CHECK_VAR_SLOT(idx, bit) \
+ "r1 = r10;" \
+ "r1 += " idx ";" \
+ "r2 = *(u64 *)(r1 - 16);" \
+ "r3 = %[magic] ll;" \
+ "r3 = *(u64 *)(r3 + 0);" \
+ "if r2 != r3 goto 9f;" \
+ PAD_RAN(bit) \
+ "9:"
+
+/* A callee that unwinds when r1 is at least 101, and touches none of r6-r9. */
+static __used __naked __noinline __u64 var_unwinder(void)
+{
+ asm volatile (
+ "if r1 < 101 goto 1f;"
+ "call bpf_unwind;"
+"1:"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 var_stack_frame(void)
+{
+ asm volatile (
+ FILL_MAGIC_SLOTS
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "r6 &= 1;" /* an unknown slot number... */
+ "r6 <<= 3;" /* ...as an aligned byte offset */
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call var_unwinder;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+ CHECK_VAR_SLOT("r6", "%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_VAR_STACK), __imm_addr(input), __imm_addr(magic),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_VAR_STACK)
+int entry_var_stack(void *ctx)
+{
+ return var_stack_frame();
+}
+
+/* Two pads with an uncovered frame between them. */
+static __used __naked __noinline __u64 gap_inner_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call pc_unwinder;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r6 = r0;"
+ PAD_RAN("%[ran]")
+ "r1 = r6;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_GAP_INNER), __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/* The frame in between, with no record of its own. */
+static __used __noinline __u64 gap_mid(void)
+{
+ return gap_inner_frame() + 1;
+}
+
+static __used __naked __noinline __u64 gap_outer_frame(void)
+{
+ asm volatile (
+"1:" "call gap_mid;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r6 = r0;"
+ PAD_RAN("%[ran]")
+ "r1 = r6;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_GAP_OUTER), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/* And one more uncovered frame between the outer pad and the boundary. */
+static __used __noinline __u64 gap_top(void)
+{
+ return gap_outer_frame() + 1;
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_GAP_INNER | RAN_GAP_OUTER)
+int entry_two_pads(void *ctx)
+{
+ return gap_top();
+}
+
+/*
+ * A precision chain crossing a resume: the outer frame's pad uses r6 as a
+ * variable stack offset, and the only way into that pad is the resume that
+ * ends the inner frame's pad, so backtracking goes from the pad through the
+ * inner frame and back to where r6 was bounded.
+ */
+static __used __naked __noinline __u64 prec_inner_frame(void)
+{
+ asm volatile (
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 prec_outer_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "r6 &= 0x7;"
+ "r0 = 0;"
+ "*(u64 *)(r10 - 8) = r0;"
+ "*(u64 *)(r10 - 16) = r0;"
+"1:" "call prec_inner_frame;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* pad: r6 as a variable stack offset */
+ "r2 = r10;"
+ "r2 += -16;"
+ "r2 += r6;"
+ "*(u8 *)(r2 + 0) = 1;"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_PREC_RESUME), __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_PREC_RESUME)
+__log_level(2)
+__msg("frame1: regs=r6 stack= before {{[0-9]+}}: (85) call bpf_unwind_resume")
+__msg("frame2: regs= stack= before {{[0-9]+}}: (85) call bpf_unwind#")
+__msg("frame1: regs=r6 stack= before {{[0-9]+}}: (57) r6 &= 7")
+int entry_prec_across_resume(void *ctx)
+{
+ return prec_outer_frame();
+}
+
+/* Unwinds every time, and no record covers it, so the path simply ends. */
+static __used __naked __noinline __u64 always_unwind(void)
+{
+ asm volatile (
+ "call bpf_unwind;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/*
+ * A frame with no record of its own above one that always unwinds: nothing
+ * after the call is reachable, so dead code removal would leave it no exit
+ * and no epilogue for the unwind to send it to. One is kept, and since this
+ * frame ends in a jump rather than an exit, it is not the last instruction.
+ */
+static __used __naked __noinline __u64 no_exit_ja_mid(void)
+{
+ asm volatile (
+ "goto 2f;"
+"1:" "r0 = 1;"
+ "exit;"
+"2:" "call always_unwind;"
+ "goto 1b;" /* the last insn, and not an exit */
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 no_exit_ja_outer_frame(void)
+{
+ asm volatile (
+"1:" "call no_exit_ja_mid;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_NO_EXIT_JA), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__ret_global(pads_ran, RAN_NO_EXIT_JA)
+int entry_no_exit_ja(void *ctx)
+{
+ return no_exit_ja_outer_frame();
+}
+
+/* gcc has no indirect calls, and only these JITs emit them */
+#if defined(__clang__) && \
+ (defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64))
+
+/*
+ * A region covering an indirect call: a record names a call by its return
+ * address, which a callx leaves like any other call.
+ */
+static __used __naked __noinline __u64 callx_region_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r2 = %[always_unwind] ll;"
+"1:" "callx r2;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_CALLX), __imm_addr(always_unwind),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__ret_global(pads_ran, RAN_CALLX)
+int entry_callx_region(void *ctx)
+{
+ return callx_region_frame();
+}
+
+/*
+ * A subprog calling an unwinding one through a pointer it was handed: nothing
+ * after the call runs, but the frame still needs an exit for its epilogue.
+ */
+static __used __naked __noinline __u64 callx_arg_frame(void)
+{
+ asm volatile (
+ "callx r1;" /* r1 is always_unwind */
+ "r0 = 1;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__naked int entry_callx_arg(void)
+{
+ asm volatile (
+ "r1 = %[always_unwind] ll;"
+ "call callx_arg_frame;"
+ "exit;"
+ :
+ : __imm_addr(always_unwind)
+ : __clobber_all);
+}
+
+#endif /* __clang__ && (x86 || arm64) */
+
+/*
+ * The jump_into_pad shape with the branch dead, so the jump into the pad is
+ * walked only speculatively, reaching the pad's resume outside a pad: a
+ * barrier rather than a refusal. Only a load without CAP_PERFMON walks it,
+ * hence the unprivileged run, and the branch is dead by range rather than by
+ * a constant, which const_fold would rewrite into a plain goto before any
+ * walk.
+ */
+static __used __naked __noinline __u64 dead_jump_into_pad_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "r6 &= 7;"
+ "if r6 > 7 goto 4f;" /* never taken: walked speculatively */
+"1:" "call always_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+"4:" /* ... and its second instruction */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __caps_unpriv(CAP_BPF) __success_unpriv
+__xlated_unpriv("nospec")
+int entry_dead_jump_into_pad(void *ctx)
+{
+ return dead_jump_into_pad_frame();
+}
+
+/* Add one to the @w-bit global at @addr. */
+#define BUMP_GLOBAL(w, addr) \
+ "r1 = " addr " ll;" \
+ "r2 = *(u" w " *)(r1 + 0);" \
+ "r2 += 1;" \
+ "*(u" w " *)(r1 + 0) = r2;"
+
+/*
+ * A pad that touches a global of each width and sign a test tag can name,
+ * so that __set_global() and __ret_global() are exercised on all four.
+ */
+int tag_i = 0;
+unsigned int tag_ui = 0;
+long tag_l = 0;
+unsigned long tag_ul = 0;
+
+static __used __naked __noinline __u64 tag_types_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call pc_unwinder;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ BUMP_GLOBAL("32", "%[tag_i]")
+ BUMP_GLOBAL("32", "%[tag_ui]")
+ BUMP_GLOBAL("64", "%[tag_l]")
+ BUMP_GLOBAL("64", "%[tag_ul]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm_addr(input), __imm_addr(tag_i), __imm_addr(tag_ui),
+ __imm_addr(tag_l), __imm_addr(tag_ul)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__set_global(input, 101)
+__set_global(tag_i, -23) __set_global(tag_ui, 0xfffffffe)
+__set_global(tag_l, -23) __set_global(tag_ul, 0xfffffffffffffffe)
+__ret_global(tag_i, -22) __ret_global(tag_ui, 0xffffffff)
+__ret_global(tag_l, -22) __ret_global(tag_ul, 0xffffffffffffffff)
+int entry_tag_types(void *ctx)
+{
+ return tag_types_frame();
+}
+
+struct {
+ __uint(type, BPF_MAP_TYPE_RINGBUF);
+ __uint(max_entries, 4096);
+} shape_ringbuf SEC(".maps");
+
+/*
+ * A frame holding a reference across a call an unwind comes out of. The record
+ * over the call is what lets it hold one: the pad releases it, where a frame
+ * with no record would be left for its epilogue still holding it.
+ */
+static __used __naked __noinline __u64 held_ref_frame(void)
+{
+ asm volatile (
+ "r1 = %[shape_ringbuf] ll;"
+ "r2 = 8;"
+ "r3 = 0;"
+ "call %[bpf_ringbuf_reserve];"
+ "if r0 == 0 goto 9f;"
+ "r6 = r0;"
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call pc_unwinder;" /* cleanup region */
+"2:"
+ "r1 = r6;"
+ "r2 = 0;"
+ "call %[bpf_ringbuf_discard];"
+ "goto 9f;"
+"3:" /* landing pad: release and resume */
+ "r1 = r6;"
+ "r2 = 0;"
+ "call %[bpf_ringbuf_discard];"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+"9:"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_HELD_REF), __imm(bpf_ringbuf_reserve),
+ __imm(bpf_ringbuf_discard), __imm_addr(shape_ringbuf),
+ __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_HELD_REF)
+int entry_held_ref(void *ctx)
+{
+ return held_ref_frame();
+}
+
+/*
+ * A pad reading a slot its frame's callee wrote before it unwound. The write
+ * is there when the pad runs, and the pad has to be verified that way, or
+ * the check below is taken as always failing and the bit is never set.
+ */
+static __used __naked __noinline __u64 slot_writer(void)
+{
+ asm volatile (
+ "r2 = 42;"
+ "*(u64 *)(r1 + 0) = r2;" /* r1 is the caller's fp-8 */
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 1f;"
+ "call bpf_unwind;"
+"1:"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 callee_write_frame(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "r1 = r10;"
+ "r1 += -8;"
+"1:" "call slot_writer;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r1 = *(u64 *)(r10 - 8);"
+ "if r1 != 42 goto 9f;"
+ PAD_RAN("%[ran]")
+"9:"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_CALLEE_WRITE), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_CALLEE_WRITE)
+int entry_callee_write(void *ctx)
+{
+ return callee_write_frame();
+}
+
+/*
+ * A precision chain across an unwind: the pad uses a slot the callee wrote as
+ * a variable stack offset, so backtracking follows the slot from the pad back
+ * into the frame the unwind left.
+ */
+static __used __naked __noinline __u64 offset_writer(void)
+{
+ asm volatile (
+ "r2 = %[input] ll;"
+ "r3 = *(u64 *)(r2 + 0);"
+ "r3 &= 8;"
+ "*(u64 *)(r1 + 0) = r3;" /* r1 is the caller's fp-24 */
+ "call bpf_unwind;"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 callee_offset_frame(void)
+{
+ asm volatile (
+ FILL_MAGIC_SLOTS
+ "r1 = 0;"
+ "*(u64 *)(r10 - 24) = r1;"
+ "r1 = r10;"
+ "r1 += -24;"
+"1:" "call offset_writer;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r6 = *(u64 *)(r10 - 24);"
+ CHECK_VAR_SLOT("r6", "%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_CALLEE_OFFSET), __imm_addr(magic), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_CALLEE_OFFSET)
+__log_level(2)
+__msg("frame1: regs= stack=-24 before {{[0-9]+}}: (85) call bpf_unwind#")
+__msg("frame2: regs= stack= before {{[0-9]+}}: (7b) *(u64 *)(r1 +0) = r3")
+__msg("frame2: regs=r3 stack= before {{[0-9]+}}: (57) r3 &= 8")
+int entry_callee_offset(void *ctx)
+{
+ return callee_offset_frame();
+}
+
+/* The same through a global subprog. */
+__noinline int global_slot_writer(__u64 *p)
+{
+ if (!p)
+ return 0;
+ *p = 42;
+ if (input > 100)
+ bpf_unwind();
+ return 0;
+}
+
+static __used __naked __noinline __u64 global_write_frame(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "r1 = r10;"
+ "r1 += -8;"
+"1:" "call global_slot_writer;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r1 = *(u64 *)(r10 - 8);"
+ "if r1 != 42 goto 9f;"
+ PAD_RAN("%[ran]")
+"9:"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_GLOBAL_WRITE), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_GLOBAL_WRITE)
+int entry_global_write(void *ctx)
+{
+ return global_write_frame();
+}
+
+/*
+ * An unwind raised in a global subprog, called with no record over the call
+ * from a frame whose own caller has a pad. The global subprog is verified on
+ * its own, so the unwind is taken from the state its call returns in, and it
+ * has to go on to that pad.
+ */
+__noinline int global_unwinder(int x)
+{
+ if (x > 100)
+ bpf_unwind();
+ return 0;
+}
+
+static __used __naked __noinline __u64 through_global_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "call global_unwinder;" /* no record */
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 over_global_frame(void)
+{
+ asm volatile (
+"1:" "call through_global_frame;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_THROUGH_GLOBAL), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_THROUGH_GLOBAL)
+int entry_through_global(void *ctx)
+{
+ return over_global_frame();
+}
+
+#if defined(__clang_major__) && __clang_major__ >= 23
+
+/*
+ * The same, with the global subprog returning in R0:R2: its unwind leaves
+ * R2 checked as a return value too, though its caller never reads either.
+ */
+struct u64_pair {
+ __u64 a;
+ __u64 b;
+};
+
+__noinline struct u64_pair global_pair_unwinder(int x)
+{
+ struct u64_pair p = { x, x + 1 };
+
+ if (x > 100)
+ bpf_unwind();
+ return p;
+}
+
+static __used __naked __noinline __u64 through_pair_global_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "call global_pair_unwinder;" /* no record */
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 over_pair_global_frame(void)
+{
+ asm volatile (
+"1:" "call through_pair_global_frame;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_PAIR_GLOBAL), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_PAIR_GLOBAL)
+int entry_pair_global_unwind(void *ctx)
+{
+ return over_pair_global_frame();
+}
+
+#endif /* __clang_major__ >= 23 */
+
+/*
+ * A reference moved into a callee: the caller reserves a record and hands it
+ * over with a plain call, as rustc emits for 'consume(rec)' -- it has nothing
+ * left to drop -- and the callee drops it on return and, from its pad, on an
+ * unwind.
+ */
+static __used __naked __noinline __u64 consume_frame(void)
+{
+ asm volatile (
+ "r6 = r1;" /* the record, now owned here */
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call global_unwinder;" /* cleanup region */
+"2:"
+ "r1 = r6;"
+ "r2 = 0;"
+ "call %[bpf_ringbuf_discard];"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad: drop glue */
+ PAD_RAN("%[ran]")
+ "r1 = r6;"
+ "r2 = 0;"
+ "call %[bpf_ringbuf_discard];"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_MOVE), __imm_addr(pads_ran), __imm_addr(input),
+ __imm(bpf_ringbuf_discard)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 move_frame(void)
+{
+ asm volatile (
+ "r1 = %[shape_ringbuf] ll;"
+ "r2 = 8;"
+ "r3 = 0;"
+ "call %[bpf_ringbuf_reserve];"
+ "if r0 == 0 goto 1f;"
+ "r1 = r0;"
+ "call consume_frame;" /* the record moves: no record here */
+"1:"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm_addr(shape_ringbuf), __imm(bpf_ringbuf_reserve)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_MOVE)
+int entry_move_into_callee(void *ctx)
+{
+ return move_frame();
+}
+
+/*
+ * A bpf_unwind() inside a loop, with no pad between it and the main program.
+ * The main program's frame returns from where the unwind left it, not from
+ * the loop in the subprog.
+ */
+static __used __naked __noinline __u64 loop_unwinder(void)
+{
+ asm volatile (
+ "r6 = 0;"
+"1:"
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 != r6 goto 2f;"
+ "call bpf_unwind;"
+"2:"
+ "r6 += 1;"
+ "if r6 < 4 goto 1b;"
+ "r0 = 1;"
+ "exit;"
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 2) __retval(0)
+int entry_unwind_in_loop(void *ctx)
+{
+ return loop_unwinder();
+}
+
+/*
+ * The loop in the main program itself, around a call that unwinds on one
+ * trip. The unwind returns from the main frame at that call, inside the loop,
+ * where no checkpoint need have been made yet.
+ */
+static __used __naked __noinline __u64 unwind_on_match(void)
+{
+ asm volatile (
+ "r2 = %[input] ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "if r1 != r2 goto 1f;"
+ "call bpf_unwind;"
+"1:"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 2) __retval(0)
+__naked int entry_unwind_out_of_main_loop(void)
+{
+ asm volatile (
+ "r6 = 0;"
+"1:"
+ "r1 = r6;"
+ "call unwind_on_match;"
+ "r6 += 1;"
+ "if r6 < 4 goto 1b;"
+ "r0 = 1;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* A frame's own pad, reached from its bpf_unwind(), finds r0 zero. */
+static __used __naked __noinline __u64 own_pad_r0_frame(void)
+{
+ asm volatile (
+"1:" "call bpf_unwind;" /* cleanup region */
+"2:"
+ "r0 = 1;"
+ "exit;"
+"3:" /* landing pad */
+ "if r0 != 0 goto 9f;"
+ PAD_RAN("%[ran]")
+"9:"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_OWN_PAD_R0), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__ret_global(pads_ran, RAN_OWN_PAD_R0)
+int entry_own_pad_r0(void *ctx)
+{
+ return own_pad_r0_frame();
+}
+
+/* A pad whose speculative walk reaches an exit: a barrier, not a refusal. */
+static __used __naked __noinline __u64 spec_exit_pad_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "r6 &= 7;"
+"1:" "call always_unwind;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "if r6 > 7 goto 4f;" /* never taken: walked speculatively */
+ "call bpf_unwind_resume;"
+"4:"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__success __caps_unpriv(CAP_BPF) __success_unpriv
+__xlated_unpriv("nospec")
+int entry_spec_exit_pad(void *ctx)
+{
+ return spec_exit_pad_frame();
+}
+
+/*
+ * An unwind reaching main's exit in a program type that checks r0: its
+ * precision is backtracked past an earlier call to the insn that unwound.
+ */
+static __used __naked __noinline __u64 plain_frame(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("?cgroup/skb")
+__success
+__naked int entry_unwind_to_checked_exit(void)
+{
+ asm volatile (
+ "call plain_frame;"
+ "call always_unwind;"
+ "r0 = 1;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/*
+ * The same, out of a global call, which r6 = 0 keeps from being the state's
+ * first insn, where backtracking would stop. The verifier_bug_if() this
+ * guards against only logs, so the level 2 log is checked.
+ */
+__noinline int global_always_unwind(void)
+{
+ bpf_unwind();
+ return 0;
+}
+
+SEC("?cgroup/skb")
+__success __log_level(2) __not_msg("verifier bug")
+__naked int entry_global_unwind_to_checked_exit(void)
+{
+ asm volatile (
+ "r6 = 0;"
+ "call global_always_unwind;"
+ "r0 = 1;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread* [PATCH bpf-next v9 23/23] selftests/bpf: Load an exception cleanup program from a light skeleton
2026-10-08 7:49 [PATCH bpf-next v9 00/23] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
` (21 preceding siblings ...)
2026-10-08 7:51 ` [PATCH bpf-next v9 22/23] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
@ 2026-10-08 7:51 ` Yonghong Song
22 siblings, 0 replies; 38+ messages in thread
From: Yonghong Song @ 2026-10-08 7:51 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A light skeleton loads a program by a path of its own: its loader
program builds the attr and resolves kfunc names when it runs. Build
progs/exceptions_cleanup.c as a light skeleton too, and run the case in
which every pad runs: a table that lost records or came out at the wrong
offsets would fail to load, or run the wrong pads.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/testing/selftests/bpf/Makefile.skel | 2 +-
.../bpf/prog_tests/exceptions_cleanup.c | 31 +++++++++++++++++++
2 files changed, 32 insertions(+), 1 deletion(-)
diff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel
index 06e297a17156..d9cbe8ac85b7 100644
--- a/tools/testing/selftests/bpf/Makefile.skel
+++ b/tools/testing/selftests/bpf/Makefile.skel
@@ -40,7 +40,7 @@ LSKELS_SIGNED := fentry_test.c fexit_test.c atomics.c
# Generate both light skeleton and libbpf skeleton for these
LSKELS_EXTRA := test_ksyms_module.c test_ksyms_weak.c kfunc_call_test.c \
- kfunc_call_test_subprog.c test_global_percpu_data.c
+ kfunc_call_test_subprog.c test_global_percpu_data.c exceptions_cleanup.c
SKEL_BLACKLIST += $(LSKELS) $(LSKELS_SIGNED)
test_static_linked.skel.h-deps := test_static_linked1.bpf.o test_static_linked2.bpf.o
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
index 7c5dffd4b861..1493c3efa599 100644
--- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -6,6 +6,7 @@
#include "exceptions_cleanup_fail.skel.h"
#include "exceptions_cleanup_shapes.skel.h"
#include "exceptions_cleanup_tracing.skel.h"
+#include "exceptions_cleanup.lskel.h"
/* foo3 unwound: every frame that has a pad ran it. */
#define PADS_FOO3_UNWOUND \
@@ -63,6 +64,33 @@ static int load_tracer(struct exceptions_cleanup *tgt, bool fexit)
return err;
}
+static void test_light_skeleton(void)
+{
+ struct exceptions_cleanup_lskel *skel;
+ __u64 ctx = 0;
+ int err;
+
+ LIBBPF_OPTS(bpf_test_run_opts, topts,
+ .ctx_in = &ctx,
+ .ctx_size_in = sizeof(ctx),
+ );
+
+ skel = exceptions_cleanup_lskel__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "light open_and_load"))
+ return;
+
+ /* foo3 unwinds, so every pad runs: the whole table has to arrive. */
+ skel->bss->input = 101;
+ err = bpf_prog_test_run_opts(skel->progs.entry.prog_fd, &topts);
+ if (!ASSERT_OK(err, "run"))
+ goto out;
+ ASSERT_EQ(topts.retval, 0, "retval");
+ ASSERT_EQ(skel->bss->pads_ran, PADS_FOO3_UNWOUND | RAN_BUMP,
+ "pads_ran");
+out:
+ exceptions_cleanup_lskel__destroy(skel);
+}
+
void test_exceptions_cleanup(void)
{
char log[8192] = {};
@@ -113,6 +141,9 @@ void test_exceptions_cleanup(void)
exceptions_cleanup__destroy(skel);
+ if (test__start_subtest("light_skeleton"))
+ test_light_skeleton();
+
RUN_TESTS(exceptions_cleanup_fail);
RUN_TESTS(exceptions_cleanup_shapes);
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 38+ messages in thread