* [PATCH bpf-next v4 01/20] bpf: Accept the compiler's exception cleanup table at program load
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
@ 2026-09-21 21:00 ` Yonghong Song
2026-09-21 21:56 ` bot+bpf-ci
2026-09-21 21:00 ` [PATCH bpf-next v4 02/20] bpf: Add the bpf_unwind_resume() kfunc Yonghong Song
` (19 subsequent siblings)
20 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:00 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
LLVM 23 added exception handling support for BPF with the .bpf_cleanup
section ([1]). Rust code compiled with panic=unwind runs cleanup code
(Drop glue) when bpf_throw() fires, and the LLVM BPF backend emits that
section from the landing pads the frontend produced. Plain C cannot
generate .bpf_cleanup unless inline asm is used. The Rust compiler does not
*properly* support BPF exception handling yet, but the kernel can support
the table today, and inline assembly is enough to test it.
Add the UAPI to carry the .bpf_cleanup table into the kernel. BPF_PROG_LOAD
grows cleanup_info, cleanup_info_cnt and cleanup_info_rec_size, and struct
bpf_cleanup_info describes one record as a triple of instruction indices:
the half-open call-site range [begin_off, end_off) and the landing_pad_off
the frame resumes at. check_cleanup_info() validates the table a program is
loaded with, so the rest of the kernel can rely on it.
Link: https://github.com/llvm/llvm-project/pull/192164 [1]
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 2 +
include/uapi/linux/bpf.h | 9 ++
kernel/bpf/check_btf.c | 151 +++++++++++++++++++++++++++++++++
kernel/bpf/syscall.c | 2 +-
kernel/bpf/verifier.c | 1 +
tools/include/uapi/linux/bpf.h | 9 ++
6 files changed, 173 insertions(+), 1 deletion(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 92f528c45605..de2cff5e3eca 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -987,6 +987,8 @@ struct bpf_verifier_env {
struct arg_track **callsite_at_stack;
u32 pass_cnt; /* number of times do_check() was called */
u32 subprog_cnt;
+ struct bpf_cleanup_info *cleanup_info;
+ u32 cleanup_info_cnt;
/* number of instructions analyzed by the verifier */
u32 prev_insn_processed, insn_processed;
/* number of jmps, calls, exits analyzed so far */
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 6330b7d745c5..a69cc0427575 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -1669,6 +1669,9 @@ union bpf_attr {
* verification.
*/
__s32 keyring_id;
+ __aligned_u64 cleanup_info; /* exception cleanup table */
+ __u32 cleanup_info_rec_size; /* userspace bpf_cleanup_info size */
+ __u32 cleanup_info_cnt; /* number of bpf_cleanup_info records */
};
struct { /* anonymous struct used by BPF_OBJ_* commands */
@@ -7599,6 +7602,12 @@ struct bpf_line_info {
__u32 line_col;
};
+struct bpf_cleanup_info {
+ __u32 begin_off;
+ __u32 end_off;
+ __u32 landing_pad_off;
+};
+
struct bpf_spin_lock {
__u32 val;
};
diff --git a/kernel/bpf/check_btf.c b/kernel/bpf/check_btf.c
index 0e8b3ccc7a5b..17af0a414d74 100644
--- a/kernel/bpf/check_btf.c
+++ b/kernel/bpf/check_btf.c
@@ -407,6 +407,153 @@ static int check_core_relo(struct bpf_verifier_env *env,
return err;
}
+static int cleanup_insn_subprog(struct bpf_verifier_env *env, u32 off)
+{
+ struct bpf_subprog_info *info;
+
+ if (off >= env->prog->len)
+ return -1;
+ info = bpf_find_containing_subprog(env, off);
+ return info ? info - env->subprog_info : -1;
+}
+
+#define MIN_BPF_CLEANUP_INFO_SIZE 12
+#define MAX_CLEANUP_INFO_REC_SIZE MAX_FUNCINFO_REC_SIZE
+
+static int check_cleanup_info(struct bpf_verifier_env *env,
+ const union bpf_attr *attr,
+ bpfptr_t uattr)
+{
+ u32 krec_size = sizeof(struct bpf_cleanup_info);
+ u32 i, nrec, urec_size, min_size, prev_end = 0;
+ struct bpf_cleanup_info *krecord;
+ bpfptr_t urecord;
+ int ret = -EINVAL;
+
+ nrec = attr->cleanup_info_cnt;
+ if (!nrec)
+ return 0;
+ /* Disjoint ranges, so no more records than instructions. */
+ if (nrec > env->prog->len) {
+ verbose(env, "cleanup info has %u records for %u instructions\n",
+ nrec, env->prog->len);
+ return -EINVAL;
+ }
+
+ urec_size = attr->cleanup_info_rec_size;
+ if (urec_size < MIN_BPF_CLEANUP_INFO_SIZE ||
+ urec_size > MAX_CLEANUP_INFO_REC_SIZE ||
+ urec_size % sizeof(u32)) {
+ verbose(env, "invalid cleanup info rec size %u\n", urec_size);
+ return -EINVAL;
+ }
+
+ krecord = kvcalloc(nrec, krec_size, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!krecord)
+ return -ENOMEM;
+
+ min_size = min_t(u32, krec_size, urec_size);
+ urecord = make_bpfptr(attr->cleanup_info, uattr.is_kernel);
+ for (i = 0; i < nrec; i++) {
+ struct bpf_cleanup_info *rec = &krecord[i];
+ int sb, se, sl;
+
+ ret = bpf_check_uarg_tail_zero(urecord, krec_size, urec_size);
+ if (ret) {
+ if (ret == -E2BIG) {
+ verbose(env, "nonzero tailing record in cleanup info\n");
+ if (copy_to_bpfptr_offset(uattr,
+ offsetof(union bpf_attr,
+ cleanup_info_rec_size),
+ &min_size, sizeof(min_size)))
+ ret = -EFAULT;
+ }
+ goto err_free;
+ }
+
+ if (copy_from_bpfptr(rec, urecord, min_size)) {
+ ret = -EFAULT;
+ goto err_free;
+ }
+ bpfptr_add(&urecord, urec_size);
+
+ ret = -EINVAL;
+ if (rec->begin_off >= rec->end_off) {
+ verbose(env, "cleanup_info[%u]: begin %u >= end %u\n",
+ i, rec->begin_off, rec->end_off);
+ goto err_free;
+ }
+ if (i && rec->begin_off < prev_end) {
+ verbose(env,
+ "cleanup_info[%u]: range [%u,%u) is unsorted or overlaps the previous record\n",
+ i, rec->begin_off, rec->end_off);
+ goto err_free;
+ }
+ prev_end = rec->end_off;
+
+ sb = cleanup_insn_subprog(env, rec->begin_off);
+ se = cleanup_insn_subprog(env, rec->end_off - 1);
+ sl = cleanup_insn_subprog(env, rec->landing_pad_off);
+ if (sb < 0 || se < 0 || sl < 0) {
+ verbose(env, "cleanup_info[%u]: offset out of range\n", i);
+ goto err_free;
+ }
+ if (sb != se || sb != sl) {
+ verbose(env,
+ "cleanup_info[%u]: range/landing pad span multiple subprogs\n",
+ i);
+ goto err_free;
+ }
+ /*
+ * The second half of a 16-byte instruction carries a zero
+ * opcode and is not an instruction of its own, so no offset
+ * may name one. end_off is exclusive, so it may also be one
+ * past the last instruction of the program.
+ */
+ if (!env->prog->insnsi[rec->begin_off].code ||
+ !env->prog->insnsi[rec->landing_pad_off].code ||
+ (rec->end_off < env->prog->len &&
+ !env->prog->insnsi[rec->end_off].code)) {
+ verbose(env, "cleanup_info[%u]: points at invalid insn\n", i);
+ goto err_free;
+ }
+ }
+
+ /*
+ * Reject a landing pad that lies inside a call-site range, its own
+ * included: it would be both a pad and a call that unwinds to one, and
+ * an exception out of it would have nowhere to go.
+ */
+ ret = -EINVAL;
+ for (i = 0; i < nrec; i++) {
+ u32 pad = krecord[i].landing_pad_off;
+ u32 l = 0, r = nrec;
+
+ while (l < r) {
+ u32 m = l + (r - l) / 2;
+
+ if (pad < krecord[m].begin_off) {
+ r = m;
+ } else if (pad >= krecord[m].end_off) {
+ l = m + 1;
+ } else {
+ verbose(env,
+ "cleanup_info[%u]: landing pad %u is inside the call-site range of cleanup_info[%u]\n",
+ i, pad, m);
+ goto err_free;
+ }
+ }
+ }
+
+ env->cleanup_info = krecord;
+ env->cleanup_info_cnt = nrec;
+ return 0;
+
+err_free:
+ kvfree(krecord);
+ return ret;
+}
+
int bpf_prepare_btf_info(struct bpf_verifier_env *env,
const union bpf_attr *attr,
bpfptr_t uattr)
@@ -441,6 +588,10 @@ int bpf_check_btf_info(struct bpf_verifier_env *env,
{
int err;
+ err = check_cleanup_info(env, attr, uattr);
+ if (err)
+ return err;
+
if (!attr->func_info_cnt && !attr->line_info_cnt) {
if (check_abnormal_return(env))
return -EINVAL;
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index 113486b15d29..049717dc3b7d 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -2915,7 +2915,7 @@ int __init __used bpf_multi_func(void) { return 0; }
BTF_ID_LIST_GLOBAL_SINGLE(bpf_multi_func_btf_id, func, bpf_multi_func)
/* last field in 'union bpf_attr' used by this command */
-#define BPF_PROG_LOAD_LAST_FIELD keyring_id
+#define BPF_PROG_LOAD_LAST_FIELD cleanup_info_cnt
static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, struct bpf_log_attr *attr_log)
{
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index d62c0f74cff5..f6878f903bc1 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -21919,6 +21919,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
kvfree(env->succ);
kvfree(env->gotox_tmp_buf);
bpf_diag_free(env);
+ kvfree(env->cleanup_info);
kvfree(env);
return ret;
}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 6330b7d745c5..a69cc0427575 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -1669,6 +1669,9 @@ union bpf_attr {
* verification.
*/
__s32 keyring_id;
+ __aligned_u64 cleanup_info; /* exception cleanup table */
+ __u32 cleanup_info_rec_size; /* userspace bpf_cleanup_info size */
+ __u32 cleanup_info_cnt; /* number of bpf_cleanup_info records */
};
struct { /* anonymous struct used by BPF_OBJ_* commands */
@@ -7599,6 +7602,12 @@ struct bpf_line_info {
__u32 line_col;
};
+struct bpf_cleanup_info {
+ __u32 begin_off;
+ __u32 end_off;
+ __u32 landing_pad_off;
+};
+
struct bpf_spin_lock {
__u32 val;
};
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 01/20] bpf: Accept the compiler's exception cleanup table at program load
2026-09-21 21:00 ` [PATCH bpf-next v4 01/20] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
@ 2026-09-21 21:56 ` bot+bpf-ci
2026-09-22 3:27 ` Yonghong Song
0 siblings, 1 reply; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-21 21:56 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 4560 bytes --]
> bpf: Accept the compiler's exception cleanup table at program load
>
> LLVM 23 added exception handling support for BPF with the .bpf_cleanup
> section ([1]). Rust code compiled with panic=unwind runs cleanup code
> (Drop glue) when bpf_throw() fires, and the LLVM BPF backend emits that
> section from the landing pads the frontend produced. Plain C cannot
> generate .bpf_cleanup unless inline asm is used. The Rust compiler does not
> *properly* support BPF exception handling yet, but the kernel can support
> the table today, and inline assembly is enough to test it.
>
> Add the UAPI to carry the .bpf_cleanup table into the kernel. BPF_PROG_LOAD
> grows cleanup_info, cleanup_info_cnt and cleanup_info_rec_size, and struct
> bpf_cleanup_info describes one record as a triple of instruction indices:
> the half-open call-site range [begin_off, end_off) and the landing_pad_off
> the frame resumes at. check_cleanup_info() validates the table a program is
> loaded with, so the rest of the kernel can rely on it.
This isn't a bug, but since the table is freed at the end of bpf_check()
and nothing reads env->cleanup_info yet, would it help to say here that
later patches in the series are the consumers, rather than that the rest
of the kernel can already rely on it?
>
> Link: https://github.com/llvm/llvm-project/pull/192164 [1]
> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
> diff --git a/kernel/bpf/check_btf.c b/kernel/bpf/check_btf.c
> index 0e8b3ccc7a5b..17af0a414d74 100644
> --- a/kernel/bpf/check_btf.c
> +++ b/kernel/bpf/check_btf.c
> @@ -407,6 +407,153 @@ static int check_core_relo(struct bpf_verifier_env *env,
> return err;
> }
>
> +static int cleanup_insn_subprog(struct bpf_verifier_env *env, u32 off)
> +{
> + struct bpf_subprog_info *info;
> +
> + if (off >= env->prog->len)
> + return -1;
> + info = bpf_find_containing_subprog(env, off);
> + return info ? info - env->subprog_info : -1;
> +}
This isn't a bug, but would a name like cleanup_off_subprog() or
insn_subprog_idx() read a little less like an action here, given it
returns a subprog index?
> +
> +#define MIN_BPF_CLEANUP_INFO_SIZE 12
> +#define MAX_CLEANUP_INFO_REC_SIZE MAX_FUNCINFO_REC_SIZE
> +
> +static int check_cleanup_info(struct bpf_verifier_env *env,
> + const union bpf_attr *attr,
> + bpfptr_t uattr)
> +{
> + u32 krec_size = sizeof(struct bpf_cleanup_info);
> + u32 i, nrec, urec_size, min_size, prev_end = 0;
> + struct bpf_cleanup_info *krecord;
> + bpfptr_t urecord;
> + int ret = -EINVAL;
> +
> + nrec = attr->cleanup_info_cnt;
> + if (!nrec)
> + return 0;
> + /* Disjoint ranges, so no more records than instructions. */
> + if (nrec > env->prog->len) {
> + verbose(env, "cleanup info has %u records for %u instructions\n",
> + nrec, env->prog->len);
> + return -EINVAL;
> + }
> +
> + urec_size = attr->cleanup_info_rec_size;
> + if (urec_size < MIN_BPF_CLEANUP_INFO_SIZE ||
> + urec_size > MAX_CLEANUP_INFO_REC_SIZE ||
> + urec_size % sizeof(u32)) {
> + verbose(env, "invalid cleanup info rec size %u\n", urec_size);
> + return -EINVAL;
> + }
> +
> + krecord = kvcalloc(nrec, krec_size, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
> + if (!krecord)
> + return -ENOMEM;
> +
> + min_size = min_t(u32, krec_size, urec_size);
> + urecord = make_bpfptr(attr->cleanup_info, uattr.is_kernel);
> + for (i = 0; i < nrec; i++) {
> + struct bpf_cleanup_info *rec = &krecord[i];
> + int sb, se, sl;
> +
> + ret = bpf_check_uarg_tail_zero(urecord, krec_size, urec_size);
> + if (ret) {
> + if (ret == -E2BIG) {
> + verbose(env, "nonzero tailing record in cleanup info\n");
> + if (copy_to_bpfptr_offset(uattr,
> + offsetof(union bpf_attr,
> + cleanup_info_rec_size),
> + &min_size, sizeof(min_size)))
> + ret = -EFAULT;
> + }
> + goto err_free;
> + }
> +
> + if (copy_from_bpfptr(rec, urecord, min_size)) {
> + ret = -EFAULT;
> + goto err_free;
> + }
> + bpfptr_add(&urecord, urec_size);
This isn't a bug, but with this being the fourth copy of the
rec_size-check / tail-zero / copy_from_bpfptr preamble in this file
(prepare_btf_func, check_btf_line, check_core_relo and now
check_cleanup_info), would it be worth pulling the common part into a
small helper parameterised by the record size and the bpf_attr offset?
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35656368472
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 01/20] bpf: Accept the compiler's exception cleanup table at program load
2026-09-21 21:56 ` bot+bpf-ci
@ 2026-09-22 3:27 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-22 3:27 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/21/26 2:56 PM, bot+bpf-ci@kernel.org wrote:
>> bpf: Accept the compiler's exception cleanup table at program load
>>
>> LLVM 23 added exception handling support for BPF with the .bpf_cleanup
>> section ([1]). Rust code compiled with panic=unwind runs cleanup code
>> (Drop glue) when bpf_throw() fires, and the LLVM BPF backend emits that
>> section from the landing pads the frontend produced. Plain C cannot
>> generate .bpf_cleanup unless inline asm is used. The Rust compiler does not
>> *properly* support BPF exception handling yet, but the kernel can support
>> the table today, and inline assembly is enough to test it.
>>
>> Add the UAPI to carry the .bpf_cleanup table into the kernel. BPF_PROG_LOAD
>> grows cleanup_info, cleanup_info_cnt and cleanup_info_rec_size, and struct
>> bpf_cleanup_info describes one record as a triple of instruction indices:
>> the half-open call-site range [begin_off, end_off) and the landing_pad_off
>> the frame resumes at. check_cleanup_info() validates the table a program is
>> loaded with, so the rest of the kernel can rely on it.
> This isn't a bug, but since the table is freed at the end of bpf_check()
> and nothing reads env->cleanup_info yet, would it help to say here that
> later patches in the series are the consumers, rather than that the rest
> of the kernel can already rely on it?
Okay, we can say:
the patches, which follow the CFG walk, the unwind walk and the JITs,
are its consumers.
>
>> Link: https://github.com/llvm/llvm-project/pull/192164 [1]
>> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
>> diff --git a/kernel/bpf/check_btf.c b/kernel/bpf/check_btf.c
>> index 0e8b3ccc7a5b..17af0a414d74 100644
>> --- a/kernel/bpf/check_btf.c
>> +++ b/kernel/bpf/check_btf.c
>> @@ -407,6 +407,153 @@ static int check_core_relo(struct bpf_verifier_env *env,
>> return err;
>> }
>>
>> +static int cleanup_insn_subprog(struct bpf_verifier_env *env, u32 off)
>> +{
>> + struct bpf_subprog_info *info;
>> +
>> + if (off >= env->prog->len)
>> + return -1;
>> + info = bpf_find_containing_subprog(env, off);
>> + return info ? info - env->subprog_info : -1;
>> +}
> This isn't a bug, but would a name like cleanup_off_subprog() or
> insn_subprog_idx() read a little less like an action here, given it
> returns a subprog index?
Will change cleanup_insn_subprog() to cleanup_subprog_of().
>
>> +
>> +#define MIN_BPF_CLEANUP_INFO_SIZE 12
>> +#define MAX_CLEANUP_INFO_REC_SIZE MAX_FUNCINFO_REC_SIZE
>> +
>> +static int check_cleanup_info(struct bpf_verifier_env *env,
>> + const union bpf_attr *attr,
>> + bpfptr_t uattr)
>> +{
>> + u32 krec_size = sizeof(struct bpf_cleanup_info);
>> + u32 i, nrec, urec_size, min_size, prev_end = 0;
>> + struct bpf_cleanup_info *krecord;
>> + bpfptr_t urecord;
>> + int ret = -EINVAL;
>> +
>> + nrec = attr->cleanup_info_cnt;
>> + if (!nrec)
>> + return 0;
>> + /* Disjoint ranges, so no more records than instructions. */
>> + if (nrec > env->prog->len) {
>> + verbose(env, "cleanup info has %u records for %u instructions\n",
>> + nrec, env->prog->len);
>> + return -EINVAL;
>> + }
>> +
>> + urec_size = attr->cleanup_info_rec_size;
>> + if (urec_size < MIN_BPF_CLEANUP_INFO_SIZE ||
>> + urec_size > MAX_CLEANUP_INFO_REC_SIZE ||
>> + urec_size % sizeof(u32)) {
>> + verbose(env, "invalid cleanup info rec size %u\n", urec_size);
>> + return -EINVAL;
>> + }
>> +
>> + krecord = kvcalloc(nrec, krec_size, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
>> + if (!krecord)
>> + return -ENOMEM;
>> +
>> + min_size = min_t(u32, krec_size, urec_size);
>> + urecord = make_bpfptr(attr->cleanup_info, uattr.is_kernel);
>> + for (i = 0; i < nrec; i++) {
>> + struct bpf_cleanup_info *rec = &krecord[i];
>> + int sb, se, sl;
>> +
>> + ret = bpf_check_uarg_tail_zero(urecord, krec_size, urec_size);
>> + if (ret) {
>> + if (ret == -E2BIG) {
>> + verbose(env, "nonzero tailing record in cleanup info\n");
>> + if (copy_to_bpfptr_offset(uattr,
>> + offsetof(union bpf_attr,
>> + cleanup_info_rec_size),
>> + &min_size, sizeof(min_size)))
>> + ret = -EFAULT;
>> + }
>> + goto err_free;
>> + }
>> +
>> + if (copy_from_bpfptr(rec, urecord, min_size)) {
>> + ret = -EFAULT;
>> + goto err_free;
>> + }
>> + bpfptr_add(&urecord, urec_size);
> This isn't a bug, but with this being the fourth copy of the
> rec_size-check / tail-zero / copy_from_bpfptr preamble in this file
> (prepare_btf_func, check_btf_line, check_core_relo and now
> check_cleanup_info), would it be worth pulling the common part into a
> small helper parameterised by the record size and the bpf_attr offset?
I will skip this for now.
>
> [ ... ]
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35656368472
^ permalink raw reply [flat|nested] 80+ messages in thread
* [PATCH bpf-next v4 02/20] bpf: Add the bpf_unwind_resume() kfunc
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
2026-09-21 21:00 ` [PATCH bpf-next v4 01/20] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
@ 2026-09-21 21:00 ` Yonghong Song
2026-09-21 21:56 ` bot+bpf-ci
2026-09-21 21:00 ` [PATCH bpf-next v4 03/20] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
` (18 subsequent siblings)
20 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:00 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A compiler-emitted cleanup landing pad ends with a call to
_Unwind_Resume(), which carries the unwind on once that frame's cleanups
have run. The kernel provides the same terminator as a kfunc, named
bpf_unwind_resume() to keep the 'bpf_' prefix kfunc convention; a later
libbpf patch resolves the compiler's name to it.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/helpers.c | 14 ++++++++++++++
1 file changed, 14 insertions(+)
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 501c7ce35cba..55c59f8b4c43 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -3445,6 +3445,19 @@ __bpf_kfunc void bpf_throw(u64 cookie)
WARN(1, "A call to BPF exception callback should never return\n");
}
+/*
+ * Terminator of a compiler-emitted cleanup landing pad. The compiler names
+ * this _Unwind_Resume, the base unwind ABI's entry point for carrying an
+ * unwind on once a frame's cleanups have run. To match kernel kfunc
+ * convention, the kernel calls it bpf_unwind_resume and libbpf maps the
+ * compiler's name onto it.
+ */
+__bpf_kfunc void bpf_unwind_resume(void)
+{
+ /* Never reached: a JIT emits the way out of a landing pad instead. */
+ WARN_ONCE(1, "A JIT should have replaced the exception cleanup resume\n");
+}
+
__bpf_kfunc int bpf_wq_init(struct bpf_wq *wq, void *p__const_map, unsigned int flags)
{
struct bpf_async_kern *async = (struct bpf_async_kern *)wq;
@@ -4981,6 +4994,7 @@ BTF_ID_FLAGS(func, bpf_task_get_cgroup1, KF_ACQUIRE | KF_RCU | KF_RET_NULL)
BTF_ID_FLAGS(func, bpf_task_from_pid, KF_ACQUIRE | KF_RET_NULL)
BTF_ID_FLAGS(func, bpf_task_from_vpid, KF_ACQUIRE | KF_RET_NULL)
BTF_ID_FLAGS(func, bpf_throw)
+BTF_ID_FLAGS(func, bpf_unwind_resume)
#ifdef CONFIG_BPF_EVENTS
BTF_ID_FLAGS(func, bpf_send_signal_task)
#endif
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 02/20] bpf: Add the bpf_unwind_resume() kfunc
2026-09-21 21:00 ` [PATCH bpf-next v4 02/20] bpf: Add the bpf_unwind_resume() kfunc Yonghong Song
@ 2026-09-21 21:56 ` bot+bpf-ci
2026-09-22 3:31 ` Yonghong Song
0 siblings, 1 reply; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-21 21:56 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3595 bytes --]
> diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
> index 501c7ce35cba..55c59f8b4c43 100644
> --- a/kernel/bpf/helpers.c
> +++ b/kernel/bpf/helpers.c
> @@ -3445,6 +3445,19 @@ __bpf_kfunc void bpf_throw(u64 cookie)
> WARN(1, "A call to BPF exception callback should never return\n");
> }
>
> +/*
> + * Terminator of a compiler-emitted cleanup landing pad. The compiler names
> + * this _Unwind_Resume, the base unwind ABI's entry point for carrying an
> + * unwind on once a frame's cleanups have run. To match kernel kfunc
> + * convention, the kernel calls it bpf_unwind_resume and libbpf maps the
> + * compiler's name onto it.
> + */
> +__bpf_kfunc void bpf_unwind_resume(void)
> +{
> + /* Never reached: a JIT emits the way out of a landing pad instead. */
> + WARN_ONCE(1, "A JIT should have replaced the exception cleanup resume\n");
> +}
> +
Does this WARN become reachable at this commit?
The kfunc is registered for BPF_PROG_TYPE_TRACING, BPF_PROG_TYPE_SCHED_CLS,
BPF_PROG_TYPE_XDP, BPF_PROG_TYPE_STRUCT_OPS, BPF_PROG_TYPE_SYSCALL and
BPF_PROG_TYPE_CGROUP_SKB via generic_kfunc_set in kfunc_init().
A program of any of those types can emit a plain call to bpf_unwind_resume
and load, since there appears to be no verifier handling for the kfunc at
this commit. Searching for 'unwind_resume' at 53430aa9b73a in kernel/bpf/
and include/linux/ returns only the two lines this patch adds.
bpf_is_unwind_resume_kfunc(), kernel/bpf/exception.c, and
process_cleanup_resume() do not exist yet.
No JIT lowers the call at this commit either. bpf_jit_supports_cleanup_pads()
and the x86/arm64 dispatch land later in the series (eeb1e94ae895,
3240611f5542). So the call really does reach the C function.
Running such a program, for example a BPF_PROG_TYPE_SYSCALL program via
BPF_PROG_TEST_RUN (which needs only CAP_BPF), would produce a WARN splat
and panic the box on panic_on_warn=1.
Contrast this with the neighbouring bpf_throw(), whose WARNs are
unreachable because check_kfunc_call() gates it in kernel/bpf/verifier.c:
if (meta.func_id == special_kfunc_list[KF_bpf_throw]) {
if (!bpf_jit_supports_exceptions()) {
verbose(env, "JIT does not support calling kfunc %s#%d\n", ...);
return -ENOTSUPP;
}
env->seen_exception = true;
...
}
The rest of the series does eventually add the equivalent gates.
cleanup_check_resumes() rejects "bpf_unwind_resume() at insn %u is not
in an exception cleanup landing pad" and process_cleanup_resume() rejects
"bpf_unwind_resume() at insn %d reached without an exception in flight,"
but they are not in this commit. For the 17 commits between this one and
the end of the series, the WARN is reachable.
Since the gating mechanism is known and exists later, should the kfunc
registration be deferred until the verifier check is in place? Or should
this commit carry a minimal check_kfunc_call() rejection of the call
(as bpf_throw does) rather than relying on a JIT that does not exist yet?
> @@ -4981,6 +4994,7 @@ BTF_ID_FLAGS(func, bpf_task_get_cgroup1, KF_ACQUIRE | KF_RCU | KF_RET_NULL)
> BTF_ID_FLAGS(func, bpf_task_from_pid, KF_ACQUIRE | KF_RET_NULL)
> BTF_ID_FLAGS(func, bpf_task_from_vpid, KF_ACQUIRE | KF_RET_NULL)
> BTF_ID_FLAGS(func, bpf_throw)
> +BTF_ID_FLAGS(func, bpf_unwind_resume)
> #ifdef CONFIG_BPF_EVENTS
> BTF_ID_FLAGS(func, bpf_send_signal_task)
> #endif
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35656368472
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 02/20] bpf: Add the bpf_unwind_resume() kfunc
2026-09-21 21:56 ` bot+bpf-ci
@ 2026-09-22 3:31 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-22 3:31 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/21/26 2:56 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
>> index 501c7ce35cba..55c59f8b4c43 100644
>> --- a/kernel/bpf/helpers.c
>> +++ b/kernel/bpf/helpers.c
>> @@ -3445,6 +3445,19 @@ __bpf_kfunc void bpf_throw(u64 cookie)
>> WARN(1, "A call to BPF exception callback should never return\n");
>> }
>>
>> +/*
>> + * Terminator of a compiler-emitted cleanup landing pad. The compiler names
>> + * this _Unwind_Resume, the base unwind ABI's entry point for carrying an
>> + * unwind on once a frame's cleanups have run. To match kernel kfunc
>> + * convention, the kernel calls it bpf_unwind_resume and libbpf maps the
>> + * compiler's name onto it.
>> + */
>> +__bpf_kfunc void bpf_unwind_resume(void)
>> +{
>> + /* Never reached: a JIT emits the way out of a landing pad instead. */
>> + WARN_ONCE(1, "A JIT should have replaced the exception cleanup resume\n");
>> +}
>> +
> Does this WARN become reachable at this commit?
>
> The kfunc is registered for BPF_PROG_TYPE_TRACING, BPF_PROG_TYPE_SCHED_CLS,
> BPF_PROG_TYPE_XDP, BPF_PROG_TYPE_STRUCT_OPS, BPF_PROG_TYPE_SYSCALL and
> BPF_PROG_TYPE_CGROUP_SKB via generic_kfunc_set in kfunc_init().
>
> A program of any of those types can emit a plain call to bpf_unwind_resume
> and load, since there appears to be no verifier handling for the kfunc at
> this commit. Searching for 'unwind_resume' at 53430aa9b73a in kernel/bpf/
> and include/linux/ returns only the two lines this patch adds.
> bpf_is_unwind_resume_kfunc(), kernel/bpf/exception.c, and
> process_cleanup_resume() do not exist yet.
>
> No JIT lowers the call at this commit either. bpf_jit_supports_cleanup_pads()
> and the x86/arm64 dispatch land later in the series (eeb1e94ae895,
> 3240611f5542). So the call really does reach the C function.
>
> Running such a program, for example a BPF_PROG_TYPE_SYSCALL program via
> BPF_PROG_TEST_RUN (which needs only CAP_BPF), would produce a WARN splat
> and panic the box on panic_on_warn=1.
>
> Contrast this with the neighbouring bpf_throw(), whose WARNs are
> unreachable because check_kfunc_call() gates it in kernel/bpf/verifier.c:
>
> if (meta.func_id == special_kfunc_list[KF_bpf_throw]) {
> if (!bpf_jit_supports_exceptions()) {
> verbose(env, "JIT does not support calling kfunc %s#%d\n", ...);
> return -ENOTSUPP;
> }
> env->seen_exception = true;
> ...
> }
>
> The rest of the series does eventually add the equivalent gates.
> cleanup_check_resumes() rejects "bpf_unwind_resume() at insn %u is not
> in an exception cleanup landing pad" and process_cleanup_resume() rejects
> "bpf_unwind_resume() at insn %d reached without an exception in flight,"
> but they are not in this commit. For the 17 commits between this one and
> the end of the series, the WARN is reachable.
>
> Since the gating mechanism is known and exists later, should the kfunc
> registration be deferred until the verifier check is in place? Or should
> this commit carry a minimal check_kfunc_call() rejection of the call
> (as bpf_throw does) rather than relying on a JIT that does not exist yet?
Okay, let us remove below:
BTF_ID_FLAGS(func, bpf_unwind_resume)
so bpf_unwind_resume() can still be used in the code, but actually it
will be used at runtime much later.
>
>> @@ -4981,6 +4994,7 @@ BTF_ID_FLAGS(func, bpf_task_get_cgroup1, KF_ACQUIRE | KF_RCU | KF_RET_NULL)
>> BTF_ID_FLAGS(func, bpf_task_from_pid, KF_ACQUIRE | KF_RET_NULL)
>> BTF_ID_FLAGS(func, bpf_task_from_vpid, KF_ACQUIRE | KF_RET_NULL)
>> BTF_ID_FLAGS(func, bpf_throw)
>> +BTF_ID_FLAGS(func, bpf_unwind_resume)
>> #ifdef CONFIG_BPF_EVENTS
>> BTF_ID_FLAGS(func, bpf_send_signal_task)
>> #endif
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35656368472
^ permalink raw reply [flat|nested] 80+ messages in thread
* [PATCH bpf-next v4 03/20] bpf: Add lookups for exception cleanup resumes and landing pads
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
2026-09-21 21:00 ` [PATCH bpf-next v4 01/20] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-21 21:00 ` [PATCH bpf-next v4 02/20] bpf: Add the bpf_unwind_resume() kfunc Yonghong Song
@ 2026-09-21 21:00 ` Yonghong Song
2026-09-22 4:04 ` Alexei Starovoitov
2026-09-21 21:00 ` [PATCH bpf-next v4 04/20] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
` (17 subsequent siblings)
20 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:00 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Add two new files, exception.h and exception.c, to host the exception
handling code. Only a few helpers so far: recognising a call to
bpf_unwind_resume(), and asking which landing pad, if any, a call site
unwinds to.
The pad of a call site is kept in insn_aux_data, so the two places that
move instructions around -- bpf_patch_insn_data() and
verifier_remove_insns() -- learn to keep it in step.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 6 ++++++
kernel/bpf/Makefile | 2 +-
kernel/bpf/exception.c | 31 +++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 12 ++++++++++++
kernel/bpf/fixups.c | 26 +++++++++++++++++++++++---
5 files changed, 73 insertions(+), 4 deletions(-)
create mode 100644 kernel/bpf/exception.c
create mode 100644 kernel/bpf/exception.h
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index de2cff5e3eca..325a80ffcbe2 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -675,6 +675,11 @@ struct bpf_insn_aux_data {
u64 map_key_state; /* constant (32 bit) key tracking for maps */
int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
+ /*
+ * 1 + the instruction index of the exception cleanup landing pad this
+ * call site unwinds to, or 0 for none.
+ */
+ u32 cleanup_pad;
bool nospec; /* do not execute this instruction speculatively */
bool nospec_result; /* result is unsafe under speculation, nospec must follow */
bool zext_dst; /* this insn zero extends dst reg */
@@ -1518,6 +1523,7 @@ u32 btf_func_arg_align(const struct btf *btf, const struct btf_type *t);
int bpf_find_subprog(struct bpf_verifier_env *env, int off);
bool bpf_is_throw_kfunc(struct bpf_insn *insn);
+bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
int bpf_compute_const_regs(struct bpf_verifier_env *env);
int bpf_prune_dead_branches(struct bpf_verifier_env *env);
int bpf_check_cfg(struct bpf_verifier_env *env);
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index 9a92c348bbda..af9bc60428ad 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -11,7 +11,7 @@ obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_
obj-$(CONFIG_BPF_SYSCALL) += hashtab.o arraymap.o percpu_freelist.o bpf_lru_list.o lpm_trie.o map_in_map.o bloom_filter.o
obj-$(CONFIG_BPF_SYSCALL) += local_storage.o queue_stack_maps.o ringbuf.o bpf_insn_array.o
obj-$(CONFIG_BPF_SYSCALL) += bpf_local_storage.o bpf_task_storage.o
-obj-$(CONFIG_BPF_SYSCALL) += fixups.o cfg.o states.o backtrack.o check_btf.o
+obj-$(CONFIG_BPF_SYSCALL) += fixups.o cfg.o states.o backtrack.o check_btf.o exception.o
obj-${CONFIG_BPF_LSM} += bpf_inode_storage.o
obj-$(CONFIG_BPF_SYSCALL) += disasm.o mprog.o
obj-$(CONFIG_BPF_JIT) += trampoline.o
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
new file mode 100644
index 000000000000..4b3ac93e98c1
--- /dev/null
+++ b/kernel/bpf/exception.c
@@ -0,0 +1,31 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bpf.h>
+#include <linux/bpf_verifier.h>
+#include <linux/btf.h>
+#include <linux/btf_ids.h>
+#include <linux/filter.h>
+#include <linux/slab.h>
+#include "exception.h"
+
+#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
+
+BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
+
+static bool insn_is_unwind_resume(const struct bpf_insn *insn)
+{
+ return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
+ insn->imm == bpf_unwind_resume_id[0];
+}
+
+bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
+{
+ return insn_is_unwind_resume(insn);
+}
+
+int bpf_cleanup_pad_of_call(struct bpf_verifier_env *env, u32 idx)
+{
+ u32 pad = env->insn_aux_data[idx].cleanup_pad;
+
+ return pad ? (int)pad - 1 : -1;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
new file mode 100644
index 000000000000..0f2b9624a2ce
--- /dev/null
+++ b/kernel/bpf/exception.h
@@ -0,0 +1,12 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef _LINUX_BPF_EXCEPTION_H
+#define _LINUX_BPF_EXCEPTION_H
+
+#include <linux/types.h>
+
+struct bpf_verifier_env;
+
+int bpf_cleanup_pad_of_call(struct bpf_verifier_env *env, u32 idx);
+
+#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 2add8001c3ec..5e257d5fc0ef 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -256,11 +256,18 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
data[i].non_stack_access =
data[off + cnt - 1].non_stack_access;
data[off + cnt - 1].non_stack_access = false;
+ data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
+ data[off + cnt - 1].cleanup_pad = 0;
} else if (bpf_is_mem_insn(insn + i)) {
data[i].non_stack_access = true;
}
}
+ if (env->cleanup_info_cnt)
+ for (i = 0; i < prog_len; i++)
+ if (data[i].cleanup_pad > off + 1)
+ data[i].cleanup_pad += cnt - 1;
+
/*
* Last slot instruction could be a newly generated
* BPF_ST/BPF_LDX/BPF_STX, systematically mark it for non-stack access
@@ -544,11 +551,13 @@ void bpf_clear_insn_aux_data(struct bpf_verifier_env *env, int start, int len)
}
}
-static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
+static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt,
+ bool falls_through)
{
struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
unsigned int orig_prog_len = env->prog->len;
int err;
+ u32 i;
if (bpf_prog_is_offloaded(env->prog->aux))
bpf_prog_offload_remove_insns(env, off, cnt);
@@ -573,6 +582,17 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
sizeof(*aux_data) * (orig_prog_len - off - cnt));
env->insn_aux_data_len -= cnt;
+ if (env->cleanup_info_cnt) {
+ for (i = 0; i < env->insn_aux_data_len; i++) {
+ u32 pad = aux_data[i].cleanup_pad;
+
+ if (pad > off + cnt)
+ aux_data[i].cleanup_pad = pad - cnt;
+ else if (pad > off)
+ aux_data[i].cleanup_pad = falls_through ? off + 1 : 0;
+ }
+ }
+
return 0;
}
@@ -634,7 +654,7 @@ int bpf_opt_remove_dead_code(struct bpf_verifier_env *env)
if (!j)
continue;
- err = verifier_remove_insns(env, i, j);
+ err = verifier_remove_insns(env, i, j, false);
if (err)
return err;
insn_cnt = env->prog->len;
@@ -657,7 +677,7 @@ int bpf_opt_remove_nops(struct bpf_verifier_env *env)
if (!is_may_goto_0 && !is_ja)
continue;
- err = verifier_remove_insns(env, i, 1);
+ err = verifier_remove_insns(env, i, 1, true);
if (err)
return err;
insn_cnt--;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 03/20] bpf: Add lookups for exception cleanup resumes and landing pads
2026-09-21 21:00 ` [PATCH bpf-next v4 03/20] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
@ 2026-09-22 4:04 ` Alexei Starovoitov
2026-09-22 5:28 ` Yonghong Song
0 siblings, 1 reply; 80+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 4:04 UTC (permalink / raw)
To: Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
On Mon Sep 21, 2026 at 9:00 PM UTC, Yonghong Song wrote:
> -static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
> +static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt,
> + bool falls_through)
next time pls ask AI whether Alexei (me) likes bool-s in arguments.
I hope it's training data set is new enough to know what I really don't like.
I ranted my times on the list about the reason why 'bool foo' is wrong.
pw-bot: cr
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 03/20] bpf: Add lookups for exception cleanup resumes and landing pads
2026-09-22 4:04 ` Alexei Starovoitov
@ 2026-09-22 5:28 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-22 5:28 UTC (permalink / raw)
To: Alexei Starovoitov, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
On 9/21/26 9:04 PM, Alexei Starovoitov wrote:
> On Mon Sep 21, 2026 at 9:00 PM UTC, Yonghong Song wrote:
>> -static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
>> +static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt,
>> + bool falls_through)
> next time pls ask AI whether Alexei (me) likes bool-s in arguments.
> I hope it's training data set is new enough to know what I really don't like.
> I ranted my times on the list about the reason why 'bool foo' is wrong.
Okay, will fix.
>
> pw-bot: cr
^ permalink raw reply [flat|nested] 80+ messages in thread
* [PATCH bpf-next v4 04/20] bpf: Prepare for an exception cleanup table before the CFG walk
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (2 preceding siblings ...)
2026-09-21 21:00 ` [PATCH bpf-next v4 03/20] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
@ 2026-09-21 21:00 ` Yonghong Song
2026-09-22 18:27 ` Eduard Zingerman
2026-09-21 21:00 ` [PATCH bpf-next v4 05/20] bpf: Make exception landing pads reachable in the CFG Yonghong Song
` (16 subsequent siblings)
20 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:00 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Some plumbing work is done before bpf_check_cfg(). More specifically,
insn_aux_data records cleanup_throw_site for every bpf_throw() and
cleanup_resume_site for every bpf_unwind_resume() -- the two calls a JIT
lowers its own way rather than as calls -- and cleanup_pad, the landing pad
a frame resumes at, for every call within the [begin_off, end_off) range of
a cleanup record. Subsequent commits consume all three.
Marking the two calls here, rather than recognising them in the JIT, is
what makes the recognition exact: by the time a JIT runs,
bpf_fixup_kfunc_call() has rewritten every other kfunc's imm into an offset
from __bpf_call_base, and a BTF id compared against one of those offsets
could match an unrelated call.
bpf_prepare_cleanup_exceptions() runs before bpf_check_cfg(), because what
it produces is what the CFG walk consumes. It refuses a table on an
offloaded program, on a program whose JIT cannot dispatch landing pads or
which the JIT was not asked to compile, and on a program that also installs
an exception callback -- two different answers to what runs on the way out.
bpf_jit_supports_cleanup_pads() is weak here and says no; the arch patches
provide the real ones.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 2 ++
include/linux/filter.h | 1 +
kernel/bpf/core.c | 5 +++
kernel/bpf/exception.c | 60 ++++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 1 +
kernel/bpf/fixups.c | 6 ++++
kernel/bpf/verifier.c | 6 ++++
7 files changed, 81 insertions(+)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 325a80ffcbe2..fdee9da6b45d 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -686,6 +686,8 @@ struct bpf_insn_aux_data {
bool needs_zext; /* alu op needs to clear upper bits */
bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
+ bool cleanup_throw_site; /* call to bpf_throw() */
+ bool cleanup_resume_site; /* call to bpf_unwind_resume() */
bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
u8 alu_state; /* used in combination with alu_limit */
/* true if STX or LDX instruction is a part of a spill/fill
diff --git a/include/linux/filter.h b/include/linux/filter.h
index 422284b4fa96..2582a7606e46 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1242,6 +1242,7 @@ bool bpf_jit_supports_stack_args(void);
bool bpf_jit_supports_arena_args(void);
bool bpf_jit_supports_far_kfunc_call(void);
bool bpf_jit_supports_exceptions(void);
+bool bpf_jit_supports_cleanup_pads(void);
bool bpf_jit_supports_ptr_xchg(void);
bool bpf_jit_supports_arena(void);
bool bpf_jit_supports_insn(struct bpf_insn *insn, bool in_arena);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index 227211166dcc..a379cd1ec4c6 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -3475,6 +3475,11 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
{
}
+bool __weak bpf_jit_supports_cleanup_pads(void)
+{
+ return false;
+}
+
bool __weak bpf_jit_supports_timed_may_goto(void)
{
return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index 4b3ac93e98c1..67af78baa558 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -18,6 +18,66 @@ static bool insn_is_unwind_resume(const struct bpf_insn *insn)
insn->imm == bpf_unwind_resume_id[0];
}
+static void cleanup_mark_kfunc_sites(struct bpf_verifier_env *env)
+{
+ u32 i;
+
+ for (i = 0; i < env->prog->len; i++) {
+ struct bpf_insn *insn = &env->prog->insnsi[i];
+
+ if (bpf_is_throw_kfunc(insn))
+ env->insn_aux_data[i].cleanup_throw_site = true;
+ else if (insn_is_unwind_resume(insn))
+ env->insn_aux_data[i].cleanup_resume_site = true;
+ }
+}
+
+static void cleanup_mark_call_sites(struct bpf_verifier_env *env)
+{
+ u32 i, j;
+
+ for (i = 0; i < env->cleanup_info_cnt; i++) {
+ struct bpf_cleanup_info *rec = &env->cleanup_info[i];
+
+ for (j = rec->begin_off; j < rec->end_off; j++) {
+ struct bpf_insn *insn = &env->prog->insnsi[j];
+
+ if (!bpf_pseudo_call(insn) && !bpf_is_throw_kfunc(insn))
+ continue;
+ env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
+ }
+ }
+}
+
+int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env)
+{
+ if (!env->cleanup_info_cnt)
+ return 0;
+
+ if (bpf_prog_is_offloaded(env->prog->aux)) {
+ verbose(env,
+ "exception cleanup is not supported for offloaded programs\n");
+ return -EINVAL;
+ }
+
+ if (!bpf_jit_supports_cleanup_pads() || !env->prog->jit_requested) {
+ verbose(env,
+ "exception cleanup needs a JIT that can dispatch landing pads\n");
+ return -EOPNOTSUPP;
+ }
+ env->prog->jit_required = 1;
+
+ if (env->exception_callback_subprog) {
+ verbose(env,
+ "exception cleanup table cannot be combined with an exception callback\n");
+ return -EINVAL;
+ }
+
+ cleanup_mark_kfunc_sites(env);
+ cleanup_mark_call_sites(env);
+ return 0;
+}
+
bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
{
return insn_is_unwind_resume(insn);
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index 0f2b9624a2ce..f51383fd775c 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -7,6 +7,7 @@
struct bpf_verifier_env;
+int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
int bpf_cleanup_pad_of_call(struct bpf_verifier_env *env, u32 idx);
#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 5e257d5fc0ef..18b8812f88a7 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -256,6 +256,12 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
data[i].non_stack_access =
data[off + cnt - 1].non_stack_access;
data[off + cnt - 1].non_stack_access = false;
+ data[i].cleanup_throw_site =
+ data[off + cnt - 1].cleanup_throw_site;
+ data[off + cnt - 1].cleanup_throw_site = false;
+ data[i].cleanup_resume_site =
+ data[off + cnt - 1].cleanup_resume_site;
+ data[off + cnt - 1].cleanup_resume_site = false;
data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
data[off + cnt - 1].cleanup_pad = 0;
} else if (bpf_is_mem_insn(insn + i)) {
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index f6878f903bc1..a7a3c4b4d975 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -37,6 +37,7 @@
#include "diagnostics.h"
#include "disasm.h"
+#include "exception.h"
static const struct bpf_verifier_ops * const bpf_verifier_ops[] = {
#define BPF_PROG_TYPE(_id, _name, prog_ctx_type, kern_ctx_type) \
@@ -21712,6 +21713,11 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
if (ret < 0)
goto skip_full_check;
+ /* The CFG needs an edge from a call in a cleanup range to its pad. */
+ ret = bpf_prepare_cleanup_exceptions(env);
+ if (ret < 0)
+ goto skip_full_check;
+
/* Validate instructions and resolve the program's referenced resources. */
ret = check_and_resolve_insns(env);
if (ret < 0)
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 04/20] bpf: Prepare for an exception cleanup table before the CFG walk
2026-09-21 21:00 ` [PATCH bpf-next v4 04/20] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
@ 2026-09-22 18:27 ` Eduard Zingerman
2026-09-23 3:07 ` Yonghong Song
0 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-22 18:27 UTC (permalink / raw)
To: Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Mon, 2026-09-21 at 14:00 -0700, Yonghong Song wrote:
> Some plumbing work is done before bpf_check_cfg(). More specifically,
> insn_aux_data records cleanup_throw_site for every bpf_throw() and
> cleanup_resume_site for every bpf_unwind_resume() -- the two calls a JIT
> lowers its own way rather than as calls -- and cleanup_pad, the landing pad
> a frame resumes at, for every call within the [begin_off, end_off) range of
> a cleanup record. Subsequent commits consume all three.
>
> Marking the two calls here, rather than recognising them in the JIT, is
> what makes the recognition exact: by the time a JIT runs,
> bpf_fixup_kfunc_call() has rewritten every other kfunc's imm into an offset
> from __bpf_call_base, and a BTF id compared against one of those offsets
> could match an unrelated call.
>
> bpf_prepare_cleanup_exceptions() runs before bpf_check_cfg(), because what
> it produces is what the CFG walk consumes. It refuses a table on an
> offloaded program, on a program whose JIT cannot dispatch landing pads or
> which the JIT was not asked to compile, and on a program that also installs
> an exception callback -- two different answers to what runs on the way out.
> bpf_jit_supports_cleanup_pads() is weak here and says no; the arch patches
> provide the real ones.
>
> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
> ---
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Just a few nits.
...
> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 325a80ffcbe2..fdee9da6b45d 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -686,6 +686,8 @@ struct bpf_insn_aux_data {
> bool needs_zext; /* alu op needs to clear upper bits */
> bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
> bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
> + bool cleanup_throw_site; /* call to bpf_throw() */
> + bool cleanup_resume_site; /* call to bpf_unwind_resume() */
Nit: There are 12 bools here now, followed by a 25 bits hole,
let's move `orig_idx` before or after flags and convert
all the flags to bit fields. I think it can reduce the structure
size from 144 to 136 bytes.
Also, why not simply `throw_call` and `resume_call`?
> bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
> u8 alu_state; /* used in combination with alu_limit */
> /* true if STX or LDX instruction is a part of a spill/fill
...
> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
> index 4b3ac93e98c1..67af78baa558 100644
> --- a/kernel/bpf/exception.c
> +++ b/kernel/bpf/exception.c
> @@ -18,6 +18,66 @@ static bool insn_is_unwind_resume(const struct bpf_insn *insn)
> insn->imm == bpf_unwind_resume_id[0];
> }
>
> +static void cleanup_mark_kfunc_sites(struct bpf_verifier_env *env)
Nit: the 'cleanup_' prefix in function names triggers me a bit,
as it is usually used when there are some cleanup actions are
taken by the function. Maybe drop or reword it a bit?
> +{
> + u32 i;
> +
> + for (i = 0; i < env->prog->len; i++) {
> + struct bpf_insn *insn = &env->prog->insnsi[i];
> +
> + if (bpf_is_throw_kfunc(insn))
> + env->insn_aux_data[i].cleanup_throw_site = true;
> + else if (insn_is_unwind_resume(insn))
> + env->insn_aux_data[i].cleanup_resume_site = true;
> + }
> +}
...
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 04/20] bpf: Prepare for an exception cleanup table before the CFG walk
2026-09-22 18:27 ` Eduard Zingerman
@ 2026-09-23 3:07 ` Yonghong Song
2026-09-23 3:54 ` Eduard Zingerman
0 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-23 3:07 UTC (permalink / raw)
To: Eduard Zingerman, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On 9/22/26 11:27 AM, Eduard Zingerman wrote:
> On Mon, 2026-09-21 at 14:00 -0700, Yonghong Song wrote:
>> Some plumbing work is done before bpf_check_cfg(). More specifically,
>> insn_aux_data records cleanup_throw_site for every bpf_throw() and
>> cleanup_resume_site for every bpf_unwind_resume() -- the two calls a JIT
>> lowers its own way rather than as calls -- and cleanup_pad, the landing pad
>> a frame resumes at, for every call within the [begin_off, end_off) range of
>> a cleanup record. Subsequent commits consume all three.
>>
>> Marking the two calls here, rather than recognising them in the JIT, is
>> what makes the recognition exact: by the time a JIT runs,
>> bpf_fixup_kfunc_call() has rewritten every other kfunc's imm into an offset
>> from __bpf_call_base, and a BTF id compared against one of those offsets
>> could match an unrelated call.
>>
>> bpf_prepare_cleanup_exceptions() runs before bpf_check_cfg(), because what
>> it produces is what the CFG walk consumes. It refuses a table on an
>> offloaded program, on a program whose JIT cannot dispatch landing pads or
>> which the JIT was not asked to compile, and on a program that also installs
>> an exception callback -- two different answers to what runs on the way out.
>> bpf_jit_supports_cleanup_pads() is weak here and says no; the arch patches
>> provide the real ones.
>>
>> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
>> ---
> Acked-by: Eduard Zingerman <eddyz87@gmail.com>
>
> Just a few nits.
>
> ...
>
>> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
>> index 325a80ffcbe2..fdee9da6b45d 100644
>> --- a/include/linux/bpf_verifier.h
>> +++ b/include/linux/bpf_verifier.h
>> @@ -686,6 +686,8 @@ struct bpf_insn_aux_data {
>> bool needs_zext; /* alu op needs to clear upper bits */
>> bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
>> bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
>> + bool cleanup_throw_site; /* call to bpf_throw() */
>> + bool cleanup_resume_site; /* call to bpf_unwind_resume() */
> Nit: There are 12 bools here now, followed by a 25 bits hole,
> let's move `orig_idx` before or after flags and convert
> all the flags to bit fields. I think it can reduce the structure
> size from 144 to 136 bytes.
I will make these bool as bitfield to reduce the struct size.
> Also, why not simply `throw_call` and `resume_call`?
Will do.
>
>> bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
>> u8 alu_state; /* used in combination with alu_limit */
>> /* true if STX or LDX instruction is a part of a spill/fill
> ...
>
>> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
>> index 4b3ac93e98c1..67af78baa558 100644
>> --- a/kernel/bpf/exception.c
>> +++ b/kernel/bpf/exception.c
>> @@ -18,6 +18,66 @@ static bool insn_is_unwind_resume(const struct bpf_insn *insn)
>> insn->imm == bpf_unwind_resume_id[0];
>> }
>>
>> +static void cleanup_mark_kfunc_sites(struct bpf_verifier_env *env)
> Nit: the 'cleanup_' prefix in function names triggers me a bit,
> as it is usually used when there are some cleanup actions are
> taken by the function. Maybe drop or reword it a bit?
Yes, I will try to avoid cleanup_ prefix then. For static functions,
I will not have cleanup_ prefix or may use exc_ prefix.
For global function, I will bpf_cleanup_ as prefix.
>
>> +{
>> + u32 i;
>> +
>> + for (i = 0; i < env->prog->len; i++) {
>> + struct bpf_insn *insn = &env->prog->insnsi[i];
>> +
>> + if (bpf_is_throw_kfunc(insn))
>> + env->insn_aux_data[i].cleanup_throw_site = true;
>> + else if (insn_is_unwind_resume(insn))
>> + env->insn_aux_data[i].cleanup_resume_site = true;
>> + }
>> +}
> ...
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 04/20] bpf: Prepare for an exception cleanup table before the CFG walk
2026-09-23 3:07 ` Yonghong Song
@ 2026-09-23 3:54 ` Eduard Zingerman
2026-09-23 4:05 ` Yonghong Song
0 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-23 3:54 UTC (permalink / raw)
To: Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Tue, 2026-09-22 at 20:07 -0700, Yonghong Song wrote:
...
> > > diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
> > > index 4b3ac93e98c1..67af78baa558 100644
> > > --- a/kernel/bpf/exception.c
> > > +++ b/kernel/bpf/exception.c
> > > @@ -18,6 +18,66 @@ static bool insn_is_unwind_resume(const struct bpf_insn *insn)
> > > insn->imm == bpf_unwind_resume_id[0];
> > > }
> > >
> > > +static void cleanup_mark_kfunc_sites(struct bpf_verifier_env *env)
> > Nit: the 'cleanup_' prefix in function names triggers me a bit,
> > as it is usually used when there are some cleanup actions are
> > taken by the function. Maybe drop or reword it a bit?
>
> Yes, I will try to avoid cleanup_ prefix then. For static functions,
> I will not have cleanup_ prefix or may use exc_ prefix.
> For global function, I will bpf_cleanup_ as prefix.
Yonghong, a wider naming question: should we use cleanup/cleanup_pad
or landing/landing_pad to name these things?
It appears both LLVM and GCC use landing pad.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 04/20] bpf: Prepare for an exception cleanup table before the CFG walk
2026-09-23 3:54 ` Eduard Zingerman
@ 2026-09-23 4:05 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-23 4:05 UTC (permalink / raw)
To: Eduard Zingerman, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On 9/22/26 8:54 PM, Eduard Zingerman wrote:
> On Tue, 2026-09-22 at 20:07 -0700, Yonghong Song wrote:
>
> ...
>
>>>> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
>>>> index 4b3ac93e98c1..67af78baa558 100644
>>>> --- a/kernel/bpf/exception.c
>>>> +++ b/kernel/bpf/exception.c
>>>> @@ -18,6 +18,66 @@ static bool insn_is_unwind_resume(const struct bpf_insn *insn)
>>>> insn->imm == bpf_unwind_resume_id[0];
>>>> }
>>>>
>>>> +static void cleanup_mark_kfunc_sites(struct bpf_verifier_env *env)
>>> Nit: the 'cleanup_' prefix in function names triggers me a bit,
>>> as it is usually used when there are some cleanup actions are
>>> taken by the function. Maybe drop or reword it a bit?
>> Yes, I will try to avoid cleanup_ prefix then. For static functions,
>> I will not have cleanup_ prefix or may use exc_ prefix.
>> For global function, I will bpf_cleanup_ as prefix.
> Yonghong, a wider naming question: should we use cleanup/cleanup_pad
> or landing/landing_pad to name these things?
> It appears both LLVM and GCC use landing pad.
The reason for functions with cleanup_* is due to section name '.bpf_cleanup'.
cleanup actually is a good name since it means something wrong and go to
do cleanup. Unfortunately, cleanup has some other meaning which makes it
hard to infer.
landing or landing_pad itself infers exception target (landing_pad)
but it does not infer exception source.
In my next version, I tried to use bpf_exc_* where 'exc' means exception.
^ permalink raw reply [flat|nested] 80+ messages in thread
* [PATCH bpf-next v4 05/20] bpf: Make exception landing pads reachable in the CFG
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (3 preceding siblings ...)
2026-09-21 21:00 ` [PATCH bpf-next v4 04/20] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
@ 2026-09-21 21:00 ` Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 06/20] bpf: Explore the landing pads no call site reaches Yonghong Song
` (15 subsequent siblings)
20 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:00 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Both the CFG walk and liveness are involved. A bpf_throw() or a bpf2bpf
call inside the [begin_off, end_off) range of a cleanup record can reach
that record's landing pad, so both grow that edge; and a
bpf_unwind_resume() reaches nothing after it, so it has no successors.
Liveness needs one more thing. bpf_stack_slot_alive() decides whether an
outer frame's stack slot is still read after the call the frame is
suspended at by looking at the instruction after the call. A slot whose
only remaining reader is the landing pad -- which is every slot a
compiler-generated pad reloads, since only four registers survive a call --
is dead by that measure, so clean_verifier_state() poisons it while the
callee runs and the pad is then rejected for reading it. Ask about the pad
as well when the call site names one.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/cfg.c | 54 ++++++++++++++++++++++++++++++++++++++++---
kernel/bpf/liveness.c | 20 ++++++++++++++++
2 files changed, 71 insertions(+), 3 deletions(-)
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index 842c7d1eabcc..ef02906fd5a8 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -6,6 +6,7 @@
#include <linux/sort.h>
#include "diagnostics.h"
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -158,17 +159,64 @@ static int push_insn(int t, int w, int e, struct bpf_verifier_env *env)
return DONE_EXPLORING;
}
+static int visit_cleanup_pad_edge(int t, struct bpf_verifier_env *env)
+{
+ int *insn_stack = env->cfg.insn_stack;
+ int *insn_state = env->cfg.insn_state;
+ int w;
+
+ if (!env->cleanup_info_cnt)
+ return DONE_EXPLORING;
+ w = bpf_cleanup_pad_of_call(env, t);
+ if (w < 0)
+ return DONE_EXPLORING;
+
+ /*
+ * @t is a call that may branch here, and @w is the target of that
+ * branch, so both are prune points. @w especially: every covered call
+ * site in a region unwinds to the same pad, and without a prune point
+ * at its head the verifier walks the pad again for each of them.
+ */
+ mark_prune_point(env, t);
+ mark_prune_point(env, w);
+ mark_jmp_point(env, w);
+ mark_jump_target(env, w);
+
+ if (insn_state[w])
+ return DONE_EXPLORING;
+ if (env->cfg.cur_stack >= env->prog->len)
+ return -E2BIG;
+ insn_stack[env->cfg.cur_stack++] = w;
+ insn_state[w] |= DISCOVERED;
+ return KEEP_EXPLORING;
+}
+
+static int merge_visit_ret(int a, int b)
+{
+ if (a < 0)
+ return a;
+ if (b < 0)
+ return b;
+ if (a == KEEP_EXPLORING || b == KEEP_EXPLORING)
+ return KEEP_EXPLORING;
+ return DONE_EXPLORING;
+}
+
static int visit_func_call_insn(int t, struct bpf_insn *insns,
struct bpf_verifier_env *env,
bool visit_callee)
{
- int ret, insn_sz;
+ int ret, insn_sz, pad_ret;
int w;
+ pad_ret = visit_cleanup_pad_edge(t, env);
+ if (pad_ret < 0)
+ return pad_ret;
+
insn_sz = bpf_is_ldimm64(&insns[t]) ? 2 : 1;
ret = push_insn(t, t + insn_sz, FALLTHROUGH, env);
if (ret)
- return ret;
+ return merge_visit_ret(pad_ret, ret);
mark_prune_point(env, t + insn_sz);
/* when we exit from subprog, we need to record non-linear history */
@@ -180,7 +228,7 @@ static int visit_func_call_insn(int t, struct bpf_insn *insns,
merge_callee_effects(env, t, w);
ret = push_insn(t, w, BRANCH, env);
}
- return ret;
+ return merge_visit_ret(pad_ret, ret);
}
struct bpf_iarray *bpf_iarray_realloc(struct bpf_iarray *old, size_t n_elem)
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index 44ecdc5b4ec2..2d290de8bdf1 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -8,6 +8,8 @@
#include <linux/slab.h>
#include <linux/sort.h>
+#include "exception.h"
+
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
struct per_frame_masks {
@@ -256,6 +258,9 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
succ = env->succ;
succ->cnt = 0;
+ if (unlikely(bpf_is_unwind_resume_kfunc(insn)))
+ return succ;
+
opcode_info = &opcode_info_tbl[BPF_CLASS(insn->code) | BPF_OP(insn->code)];
insn_sz = bpf_is_ldimm64(insn) ? 2 : 1;
if (opcode_info->can_fallthrough)
@@ -264,6 +269,13 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
if (opcode_info->can_jump)
succ->items[succ->cnt++] = idx + bpf_jmp_offset(insn) + 1;
+ if (unlikely(env->cleanup_info_cnt)) {
+ int pad = bpf_cleanup_pad_of_call(env, idx);
+
+ if (pad >= 0)
+ succ->items[succ->cnt++] = pad;
+ }
+
return succ;
}
@@ -397,6 +409,14 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
alive = bpf_calls_callback(env, callsite)
? is_live_before(instance, callsite, rel, half_spi)
: is_live_before(instance, callsite + 1, rel, half_spi);
+
+ /* Control may go to the landing pad. */
+ if (!alive && unlikely(env->cleanup_info_cnt)) {
+ int pad = bpf_cleanup_pad_of_call(env, callsite);
+
+ if (pad >= 0)
+ alive = is_live_before(instance, pad, rel, half_spi);
+ }
if (alive)
return true;
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* [PATCH bpf-next v4 06/20] bpf: Explore the landing pads no call site reaches
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (4 preceding siblings ...)
2026-09-21 21:00 ` [PATCH bpf-next v4 05/20] bpf: Make exception landing pads reachable in the CFG Yonghong Song
@ 2026-09-21 21:01 ` Yonghong Song
2026-09-21 23:58 ` Eduard Zingerman
2026-09-21 21:01 ` [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch Yonghong Song
` (14 subsequent siblings)
20 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:01 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A cleanup record need not cover a call an exception can unwind out of: a
frontend is free to emit a region around a helper or an ordinary kfunc,
both nounwind here. Nothing marks a call site then, and that record's
landing pad is reached by nothing at all -- leaving bpf_check_cfg() to
refuse the program over code its own frontend had no way not to emit:
0: call bpf_preempt_disable
1: call bpf_preempt_enable record = { begin = 1, end = 2, pad = 4 }
2: r0 = 0
3: exit
4: r1 = pads_ran ll landing pad
6: r2 = *(u64 *)(r1 + 0)
7: r2 |= RAN_NOUNWIND_REC
8: *(u64 *)(r1 + 0) = r2
9: call bpf_unwind_resume
10: exit
The range [1,2) holds one call, and it is a kfunc, so an exception cannot
come out of it. cleanup_mark_call_sites() marks nothing, nothing pushes an
edge to 4, and 4 through 10 are reachable from nothing: "unreachable insn
4". The pad is dead, which is correct -- no exception can ever arrive at it
-- but the program is fine and has to load.
Walk every pad the table names that the edges did not reach, the way the
walk is already re-seeded at an exception callback. From there the pad is
code like any other: do_check() never enters it, because no call site
dispatches to it, so the dead code sweep removes it along with everything
else that was not reached.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/cfg.c | 17 +++++++++++++++++
1 file changed, 17 insertions(+)
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index ef02906fd5a8..ba8f9cca2955 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -640,6 +640,7 @@ int bpf_check_cfg(struct bpf_verifier_env *env)
int insn_cnt = env->prog->len;
int *insn_stack, *insn_state;
int ex_insn_beg, i, ret = 0;
+ u32 pad_idx = 0;
insn_state = env->cfg.insn_state = kvzalloc_objs(int, insn_cnt,
GFP_KERNEL_ACCOUNT);
@@ -695,6 +696,22 @@ int bpf_check_cfg(struct bpf_verifier_env *env)
goto walk_cfg;
}
+ /*
+ * A landing pad no call site was marked with -- a record whose range
+ * holds no bpf2bpf call and no bpf_throw() -- is reached by nothing.
+ * Walk it from here, and let the dead code sweep remove it.
+ */
+ while (pad_idx < env->cleanup_info_cnt) {
+ u32 pad = env->cleanup_info[pad_idx++].landing_pad_off;
+
+ if (insn_state[pad] != EXPLORED) {
+ insn_state[pad] = DISCOVERED;
+ insn_stack[0] = pad;
+ env->cfg.cur_stack = 1;
+ goto walk_cfg;
+ }
+ }
+
for (i = 0; i < insn_cnt; i++) {
struct bpf_insn *insn = &env->prog->insnsi[i];
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 06/20] bpf: Explore the landing pads no call site reaches
2026-09-21 21:01 ` [PATCH bpf-next v4 06/20] bpf: Explore the landing pads no call site reaches Yonghong Song
@ 2026-09-21 23:58 ` Eduard Zingerman
2026-09-22 3:32 ` Yonghong Song
0 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-21 23:58 UTC (permalink / raw)
To: Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Mon, 2026-09-21 at 14:01 -0700, Yonghong Song wrote:
> A cleanup record need not cover a call an exception can unwind out of: a
> frontend is free to emit a region around a helper or an ordinary kfunc,
> both nounwind here. Nothing marks a call site then, and that record's
> landing pad is reached by nothing at all -- leaving bpf_check_cfg() to
> refuse the program over code its own frontend had no way not to emit:
>
> 0: call bpf_preempt_disable
> 1: call bpf_preempt_enable record = { begin = 1, end = 2, pad = 4 }
> 2: r0 = 0
> 3: exit
> 4: r1 = pads_ran ll landing pad
> 6: r2 = *(u64 *)(r1 + 0)
> 7: r2 |= RAN_NOUNWIND_REC
> 8: *(u64 *)(r1 + 0) = r2
> 9: call bpf_unwind_resume
> 10: exit
>
> The range [1,2) holds one call, and it is a kfunc, so an exception cannot
> come out of it. cleanup_mark_call_sites() marks nothing, nothing pushes an
> edge to 4, and 4 through 10 are reachable from nothing: "unreachable insn
> 4". The pad is dead, which is correct -- no exception can ever arrive at it
> -- but the program is fine and has to load.
>
> Walk every pad the table names that the edges did not reach, the way the
> walk is already re-seeded at an exception callback. From there the pad is
> code like any other: do_check() never enters it, because no call site
> dispatches to it, so the dead code sweep removes it along with everything
> else that was not reached.
>
> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
> ---
Yonghong, have you observed such dead code being generated by rustc?
It is a bit surprising, tbh.
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 06/20] bpf: Explore the landing pads no call site reaches
2026-09-21 23:58 ` Eduard Zingerman
@ 2026-09-22 3:32 ` Yonghong Song
2026-09-22 4:10 ` Eduard Zingerman
0 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-22 3:32 UTC (permalink / raw)
To: Eduard Zingerman, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On 9/21/26 4:58 PM, Eduard Zingerman wrote:
> On Mon, 2026-09-21 at 14:01 -0700, Yonghong Song wrote:
>> A cleanup record need not cover a call an exception can unwind out of: a
>> frontend is free to emit a region around a helper or an ordinary kfunc,
>> both nounwind here. Nothing marks a call site then, and that record's
>> landing pad is reached by nothing at all -- leaving bpf_check_cfg() to
>> refuse the program over code its own frontend had no way not to emit:
>>
>> 0: call bpf_preempt_disable
>> 1: call bpf_preempt_enable record = { begin = 1, end = 2, pad = 4 }
>> 2: r0 = 0
>> 3: exit
>> 4: r1 = pads_ran ll landing pad
>> 6: r2 = *(u64 *)(r1 + 0)
>> 7: r2 |= RAN_NOUNWIND_REC
>> 8: *(u64 *)(r1 + 0) = r2
>> 9: call bpf_unwind_resume
>> 10: exit
>>
>> The range [1,2) holds one call, and it is a kfunc, so an exception cannot
>> come out of it. cleanup_mark_call_sites() marks nothing, nothing pushes an
>> edge to 4, and 4 through 10 are reachable from nothing: "unreachable insn
>> 4". The pad is dead, which is correct -- no exception can ever arrive at it
>> -- but the program is fine and has to load.
>>
>> Walk every pad the table names that the edges did not reach, the way the
>> walk is already re-seeded at an exception callback. From there the pad is
>> code like any other: do_check() never enters it, because no call site
>> dispatches to it, so the dead code sweep removes it along with everything
>> else that was not reached.
>>
>> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
>> ---
> Yonghong, have you observed such dead code being generated by rustc?
> It is a bit surprising, tbh.
The above code is with inline asm. rustc won't be able to generate such
code.
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 06/20] bpf: Explore the landing pads no call site reaches
2026-09-22 3:32 ` Yonghong Song
@ 2026-09-22 4:10 ` Eduard Zingerman
0 siblings, 0 replies; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-22 4:10 UTC (permalink / raw)
To: Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Mon, 2026-09-21 at 20:32 -0700, Yonghong Song wrote:
>
>
> On 9/21/26 4:58 PM, Eduard Zingerman wrote:
> > On Mon, 2026-09-21 at 14:01 -0700, Yonghong Song wrote:
> > > A cleanup record need not cover a call an exception can unwind out of: a
> > > frontend is free to emit a region around a helper or an ordinary kfunc,
> > > both nounwind here. Nothing marks a call site then, and that record's
> > > landing pad is reached by nothing at all -- leaving bpf_check_cfg() to
> > > refuse the program over code its own frontend had no way not to emit:
> > >
> > > 0: call bpf_preempt_disable
> > > 1: call bpf_preempt_enable record = { begin = 1, end = 2, pad = 4 }
> > > 2: r0 = 0
> > > 3: exit
> > > 4: r1 = pads_ran ll landing pad
> > > 6: r2 = *(u64 *)(r1 + 0)
> > > 7: r2 |= RAN_NOUNWIND_REC
> > > 8: *(u64 *)(r1 + 0) = r2
> > > 9: call bpf_unwind_resume
> > > 10: exit
> > >
> > > The range [1,2) holds one call, and it is a kfunc, so an exception cannot
> > > come out of it. cleanup_mark_call_sites() marks nothing, nothing pushes an
> > > edge to 4, and 4 through 10 are reachable from nothing: "unreachable insn
> > > 4". The pad is dead, which is correct -- no exception can ever arrive at it
> > > -- but the program is fine and has to load.
> > >
> > > Walk every pad the table names that the edges did not reach, the way the
> > > walk is already re-seeded at an exception callback. From there the pad is
> > > code like any other: do_check() never enters it, because no call site
> > > dispatches to it, so the dead code sweep removes it along with everything
> > > else that was not reached.
> > >
> > > Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
> > > ---
> > Yonghong, have you observed such dead code being generated by rustc?
> > It is a bit surprising, tbh.
>
> The above code is with inline asm. rustc won't be able to generate such
> code.
I'd say we shouldn't include such mechanics only for testing.
What kinds of tests would require it?
^ permalink raw reply [flat|nested] 80+ messages in thread
* [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (5 preceding siblings ...)
2026-09-21 21:01 ` [PATCH bpf-next v4 06/20] bpf: Explore the landing pads no call site reaches Yonghong Song
@ 2026-09-21 21:01 ` Yonghong Song
2026-09-21 21:20 ` sashiko-bot
` (2 more replies)
2026-09-21 21:01 ` [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier Yonghong Song
` (13 subsequent siblings)
20 siblings, 3 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:01 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
The table is on the instructions and the landing pads are in the control
flow graph. What is left before bpf_throw() can be taught to dispatch them
is to work out what that walk will need, and refuse the shapes it could not
handle.
bpf_check_cleanup_exceptions() runs after bpf_check_cfg(). Everything it
needs is control flow, so it reads what that walk has already worked out
rather than working it out again:
subprog_info.might_throw which subprograms an exception may leave,
closed over the call graph by
merge_callee_effects() as the walk pops each
callee
cleanup_reachability() what each instruction can reach -- a resume, a
plain exit, a throw, an indirect jump
cleanup_mark_pad_bodies() which instructions only ever run with an
exception already in flight
What it refuses:
- a pad that reaches both a resume and a plain exit, or neither: nothing
says whether it is a cleanup pad or a catch pad
- a catch pad, which ends in a plain exit: the walker calls a pad as a
subroutine and cannot hand a frame back its own execution
- a throw in a pad, or a call from a pad to a subprogram that can throw:
a second unwind over frames the first is still discarding
- a bpf_unwind_resume() outside a pad body
- a subprogram that may unwind used as a helper callback: the helper's
own kernel frame would end the walk before it found a boundary
- a tail call, an indirect jump, or an outgoing on-stack call argument in
a pad body, all of which touch a stack the pad does not own
- a BPF_LD_[ABS|IND] in a pad body: a failed load leaves the subprogram
through the hidden exit gen_ld_abs() patches in, and an exit in a pad
is the epilogue, which pops off the walker's stack and returns through
the frame the walk is discarding
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_verifier.h | 1 +
kernel/bpf/exception.c | 404 +++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 2 +
kernel/bpf/fixups.c | 1 +
kernel/bpf/verifier.c | 8 +
5 files changed, 416 insertions(+)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index fdee9da6b45d..aa631bb45f76 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -688,6 +688,7 @@ struct bpf_insn_aux_data {
bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
bool cleanup_throw_site; /* call to bpf_throw() */
bool cleanup_resume_site; /* call to bpf_unwind_resume() */
+ bool in_cleanup_pad; /* reachable from an exception cleanup landing pad */
bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
u8 alu_state; /* used in combination with alu_limit */
/* true if STX or LDX instruction is a part of a spill/fill
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index 67af78baa558..da8fa6eb7e4b 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -6,6 +6,7 @@
#include <linux/btf_ids.h>
#include <linux/filter.h>
#include <linux/slab.h>
+#include <linux/sort.h>
#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -18,6 +19,130 @@ static bool insn_is_unwind_resume(const struct bpf_insn *insn)
insn->imm == bpf_unwind_resume_id[0];
}
+/* What an instruction does to intra-subprog control flow. */
+enum cleanup_insn_kind {
+ CLEANUP_INSN_PLAIN, /* the next insn runs */
+ CLEANUP_INSN_JUMP, /* unconditional jump */
+ CLEANUP_INSN_COND, /* the next insn runs, or the branch target */
+ CLEANUP_INSN_EXIT,
+ CLEANUP_INSN_THROW, /* call bpf_throw: nothing after it runs */
+ CLEANUP_INSN_RESUME, /* call bpf_unwind_resume: likewise */
+ CLEANUP_INSN_CALL, /* call to another subprog */
+ CLEANUP_INSN_GOTOX, /* indirect jump: successors not known here */
+};
+
+/* What each instruction can reach, computed once by cleanup_reachability(). */
+#define CLEANUP_REACH_RESUME BIT(0) /* a bpf_unwind_resume() call */
+#define CLEANUP_REACH_EXIT BIT(1) /* a plain BPF_EXIT */
+#define CLEANUP_REACH_UNKNOWN BIT(2) /* an indirect jump */
+#define CLEANUP_REACH_THROW BIT(3) /* a bpf_throw() call */
+
+/* Scratch shared by the analyses, sized once so no walker has to allocate. */
+struct cleanup_ctx {
+ struct bpf_verifier_env *env;
+ u8 *reach; /* per insn: CLEANUP_REACH_* mask */
+ u32 *stack; /* per insn: DFS stack */
+ void *scratch; /* the one allocation all of the above live in */
+};
+
+static bool in_pad(struct bpf_verifier_env *env, u32 i)
+{
+ return env->insn_aux_data[i].in_cleanup_pad;
+}
+
+/* One scratch array for cleanup_alloc() to hand out. */
+struct cleanup_alloc_req {
+ void **dst;
+ size_t n, sz;
+};
+
+static void *cleanup_alloc(const struct cleanup_alloc_req *tab, u32 cnt)
+{
+ size_t total = 0;
+ char *block, *p;
+ u32 i;
+
+ for (i = 0; i < cnt; i++)
+ total += round_up(tab[i].n * tab[i].sz, 8);
+
+ block = kvzalloc(total, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!block)
+ return NULL;
+
+ for (i = 0, p = block; i < cnt; i++) {
+ *tab[i].dst = p;
+ p += round_up(tab[i].n * tab[i].sz, 8);
+ }
+ return block;
+}
+
+static int cleanup_subprog_of(struct bpf_verifier_env *env, u32 off)
+{
+ struct bpf_subprog_info *info = bpf_find_containing_subprog(env, off);
+
+ return info ? info - env->subprog_info : -1;
+}
+
+/* The subprogram a linear pass is currently in. */
+struct cleanup_cursor {
+ u32 start, end; /* [start, end) of the current subprogram */
+ int sub; /* its index */
+};
+
+#define CLEANUP_CURSOR_INIT { .sub = -1 }
+
+static void cleanup_cursor_to(struct bpf_verifier_env *env, struct cleanup_cursor *c, u32 i)
+{
+ while (i >= c->end) {
+ c->sub++;
+ c->start = env->subprog_info[c->sub].start;
+ c->end = env->subprog_info[c->sub + 1].start;
+ }
+}
+
+static enum cleanup_insn_kind cleanup_classify(struct bpf_verifier_env *env, u32 i,
+ int *next, int *target)
+{
+ struct bpf_insn *insn = &env->prog->insnsi[i];
+ u8 class = BPF_CLASS(insn->code);
+
+ *next = i + 1;
+ *target = -1;
+
+ if (insn->code == (BPF_LD | BPF_IMM | BPF_DW)) {
+ *next = i + 2;
+ return CLEANUP_INSN_PLAIN;
+ }
+ if (class != BPF_JMP && class != BPF_JMP32)
+ return CLEANUP_INSN_PLAIN;
+
+ switch (BPF_OP(insn->code)) {
+ case BPF_EXIT:
+ *next = -1;
+ return CLEANUP_INSN_EXIT;
+ case BPF_JA:
+ *next = -1;
+ if (BPF_SRC(insn->code) == BPF_X)
+ return CLEANUP_INSN_GOTOX;
+ *target = class == BPF_JMP32 ? i + insn->imm + 1 : i + insn->off + 1;
+ return CLEANUP_INSN_JUMP;
+ case BPF_CALL:
+ if (bpf_is_throw_kfunc(insn)) {
+ *next = -1;
+ return CLEANUP_INSN_THROW;
+ }
+ if (insn_is_unwind_resume(insn)) {
+ *next = -1;
+ return CLEANUP_INSN_RESUME;
+ }
+ return bpf_pseudo_call(insn) ? CLEANUP_INSN_CALL : CLEANUP_INSN_PLAIN;
+ default:
+ /* Conditional jump, including BPF_JCOND. */
+ *target = i + insn->off + 1;
+ return CLEANUP_INSN_COND;
+ }
+}
+
static void cleanup_mark_kfunc_sites(struct bpf_verifier_env *env)
{
u32 i;
@@ -32,6 +157,254 @@ static void cleanup_mark_kfunc_sites(struct bpf_verifier_env *env)
}
}
+int bpf_cleanup_check_callback(struct bpf_verifier_env *env, int subprog)
+{
+ if (!env->cleanup_info_cnt || !env->subprog_info[subprog].might_throw)
+ return 0;
+
+ verbose(env, "subprog %d may unwind and is used as a callback\n", subprog);
+ return -EINVAL;
+}
+
+/* Intra-subprog successors of @i, or -1 each when absent. */
+static enum cleanup_insn_kind cleanup_succ(struct bpf_verifier_env *env, u32 i,
+ u32 start, u32 end, int *next, int *target)
+{
+ enum cleanup_insn_kind kind = cleanup_classify(env, i, next, target);
+
+ if (*next < (int)start || *next >= (int)end)
+ *next = -1;
+ if (*target < (int)start || *target >= (int)end)
+ *target = -1;
+ return kind;
+}
+
+static void cleanup_add_pred(u32 *head, u32 *link, u32 to, u32 e)
+{
+ link[e] = head[to];
+ head[to] = e + 1;
+}
+
+/* What every instruction can reach along intra-subprog edges, for
+ * cleanup_pad_is_catch(). One backward walk over a predecessor index, rather
+ * than a forward walk from each landing pad, which would be quadratic.
+ */
+static int cleanup_reachability(struct cleanup_ctx *ctx)
+{
+ struct bpf_verifier_env *env = ctx->env;
+ u32 len = env->prog->len;
+ struct cleanup_cursor c = CLEANUP_CURSOR_INIT;
+ u32 *head = NULL, *link = NULL;
+ bool *queued = NULL;
+ u32 i, sp = 0;
+ void *scratch;
+ const struct cleanup_alloc_req tab[] = {
+ { (void **)&head, len, sizeof(*head) },
+ { (void **)&link, 2 * (size_t)len, sizeof(*link) },
+ { (void **)&queued, len, sizeof(*queued) },
+ };
+
+ scratch = cleanup_alloc(tab, ARRAY_SIZE(tab));
+ if (!scratch)
+ return -ENOMEM;
+
+ /* Index the predecessors, and seed the walk at the terminators. */
+ for (i = 0; i < len; i++) {
+ enum cleanup_insn_kind kind;
+ int next, target;
+
+ cleanup_cursor_to(env, &c, i);
+ kind = cleanup_succ(env, i, c.start, c.end, &next, &target);
+
+ if (kind == CLEANUP_INSN_RESUME)
+ ctx->reach[i] |= CLEANUP_REACH_RESUME;
+ else if (kind == CLEANUP_INSN_EXIT)
+ ctx->reach[i] |= CLEANUP_REACH_EXIT;
+ else if (kind == CLEANUP_INSN_GOTOX)
+ ctx->reach[i] |= CLEANUP_REACH_UNKNOWN;
+ else if (kind == CLEANUP_INSN_THROW)
+ ctx->reach[i] |= CLEANUP_REACH_THROW;
+
+ if (next >= 0)
+ cleanup_add_pred(head, link, next, 2 * i);
+ if (target >= 0)
+ cleanup_add_pred(head, link, target, 2 * i + 1);
+
+ if (ctx->reach[i]) {
+ queued[i] = true;
+ ctx->stack[sp++] = i;
+ }
+ }
+
+ /* Each instruction re-enters the worklist at most once per bit it
+ * gains, so this is linear in the number of edges.
+ */
+ while (sp) {
+ u32 j = ctx->stack[--sp];
+ u8 flags = ctx->reach[j];
+ u32 e;
+
+ queued[j] = false;
+ for (e = head[j]; e; e = link[e - 1]) {
+ u32 p = (e - 1) / 2;
+
+ if ((ctx->reach[p] | flags) == ctx->reach[p])
+ continue;
+ ctx->reach[p] |= flags;
+ if (!queued[p]) {
+ queued[p] = true;
+ ctx->stack[sp++] = p;
+ }
+ }
+ }
+ kvfree(scratch);
+ return 0;
+}
+
+static int cleanup_pad_is_catch(struct cleanup_ctx *ctx, u32 pad)
+{
+ u8 reach = ctx->reach[pad];
+
+ if (reach & CLEANUP_REACH_UNKNOWN) {
+ verbose(ctx->env, "cleanup landing pad %u reaches an indirect jump\n", pad);
+ return -EINVAL;
+ }
+ if (reach & CLEANUP_REACH_THROW) {
+ verbose(ctx->env,
+ "cleanup landing pad %u can throw while an exception is in flight\n",
+ pad);
+ return -EINVAL;
+ }
+ if (!(reach & CLEANUP_REACH_RESUME) == !(reach & CLEANUP_REACH_EXIT)) {
+ verbose(ctx->env, "cleanup landing pad %u %s\n", pad,
+ (reach & CLEANUP_REACH_RESUME) ?
+ "reaches both bpf_unwind_resume() and a plain exit" :
+ "reaches neither bpf_unwind_resume() nor an exit");
+ return -EINVAL;
+ }
+ return !!(reach & CLEANUP_REACH_EXIT);
+}
+
+static int cleanup_check_pad_insn(struct bpf_verifier_env *env, u32 i)
+{
+ struct bpf_insn *insn = &env->prog->insnsi[i];
+
+ if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call) {
+ verbose(env,
+ "bpf_tail_call() at insn %u is in an exception cleanup landing pad\n",
+ i);
+ return -EINVAL;
+ }
+ /* A BPF_LD_[ABS|IND] can leave the frame through its epilogue. */
+ if (BPF_CLASS(insn->code) == BPF_LD &&
+ (BPF_MODE(insn->code) == BPF_ABS || BPF_MODE(insn->code) == BPF_IND)) {
+ verbose(env,
+ "BPF_LD_[ABS|IND] at insn %u is in an exception cleanup landing pad\n",
+ i);
+ return -EINVAL;
+ }
+ if (is_stack_arg_st(insn) || is_stack_arg_stx(insn)) {
+ verbose(env,
+ "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
+ i);
+ return -EINVAL;
+ }
+ if (bpf_pseudo_kfunc_call(insn)) {
+ struct bpf_call_summary cs;
+
+ if (bpf_get_call_summary(env, insn, &cs) &&
+ cs.arg_slot_cnt > MAX_BPF_FUNC_REG_ARGS) {
+ verbose(env,
+ "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
+ i);
+ return -EINVAL;
+ }
+ }
+ return 0;
+}
+
+static int cleanup_mark_pad_bodies(struct cleanup_ctx *ctx)
+{
+ struct bpf_verifier_env *env = ctx->env;
+ u32 i, sp = 0;
+ int ret;
+
+ for (i = 0; i < env->cleanup_info_cnt; i++) {
+ u32 pad = env->cleanup_info[i].landing_pad_off;
+
+ if (in_pad(env, pad))
+ continue;
+
+ ret = cleanup_pad_is_catch(ctx, pad);
+ if (ret < 0)
+ return ret;
+ if (ret) {
+ verbose(env,
+ "catch landing pad %u is not supported yet, only cleanup pads that resume\n",
+ pad);
+ return -EOPNOTSUPP;
+ }
+ env->insn_aux_data[pad].in_cleanup_pad = true;
+ ctx->stack[sp++] = pad;
+ }
+
+ while (sp) {
+ u32 j = ctx->stack[--sp];
+ enum cleanup_insn_kind kind;
+ int next, target, sub;
+ u32 start, end;
+
+ ret = cleanup_check_pad_insn(env, j);
+ if (ret)
+ return ret;
+
+ sub = cleanup_subprog_of(env, j);
+ start = env->subprog_info[sub].start;
+ end = env->subprog_info[sub + 1].start;
+ kind = cleanup_succ(env, j, start, end, &next, &target);
+
+ if (kind == CLEANUP_INSN_CALL) {
+ /* check_subprogs() registered every call target. */
+ int callee = cleanup_subprog_of(env, j + env->prog->insnsi[j].imm + 1);
+
+ if (env->subprog_info[callee].might_throw) {
+ verbose(env,
+ "cleanup landing pad calls subprog %d at insn %u, which can throw while an exception is in flight\n",
+ callee, j);
+ return -EINVAL;
+ }
+ }
+
+ if (next >= 0 && !in_pad(env, next)) {
+ env->insn_aux_data[next].in_cleanup_pad = true;
+ ctx->stack[sp++] = next;
+ }
+ if (target >= 0 && !in_pad(env, target)) {
+ env->insn_aux_data[target].in_cleanup_pad = true;
+ ctx->stack[sp++] = target;
+ }
+ }
+ return 0;
+}
+
+static int cleanup_check_resumes(struct cleanup_ctx *ctx)
+{
+ struct bpf_verifier_env *env = ctx->env;
+ u32 i;
+
+ for (i = 0; i < env->prog->len; i++) {
+ if (!insn_is_unwind_resume(&env->prog->insnsi[i]))
+ continue;
+ if (in_pad(env, i))
+ continue;
+ verbose(env,
+ "bpf_unwind_resume() at insn %u is not in an exception cleanup landing pad\n",
+ i);
+ return -EINVAL;
+ }
+ return 0;
+}
+
static void cleanup_mark_call_sites(struct bpf_verifier_env *env)
{
u32 i, j;
@@ -78,6 +451,37 @@ int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env)
return 0;
}
+int bpf_check_cleanup_exceptions(struct bpf_verifier_env *env)
+{
+ u32 len = env->prog->len;
+ struct cleanup_ctx ctx = { .env = env };
+ const struct cleanup_alloc_req tab[] = {
+ { (void **)&ctx.reach, len, sizeof(*ctx.reach) },
+ { (void **)&ctx.stack, len, sizeof(*ctx.stack) },
+ };
+ int ret;
+
+ if (!env->cleanup_info_cnt)
+ return 0;
+
+ ctx.scratch = cleanup_alloc(tab, ARRAY_SIZE(tab));
+ if (!ctx.scratch)
+ return -ENOMEM;
+
+ ret = cleanup_reachability(&ctx);
+ if (ret)
+ goto out;
+
+ ret = cleanup_mark_pad_bodies(&ctx);
+ if (ret)
+ goto out;
+
+ ret = cleanup_check_resumes(&ctx);
+out:
+ kvfree(ctx.scratch);
+ return ret;
+}
+
bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
{
return insn_is_unwind_resume(insn);
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index f51383fd775c..7313dd2b65a1 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -8,6 +8,8 @@
struct bpf_verifier_env;
int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
+int bpf_check_cleanup_exceptions(struct bpf_verifier_env *env);
+int bpf_cleanup_check_callback(struct bpf_verifier_env *env, int subprog);
int bpf_cleanup_pad_of_call(struct bpf_verifier_env *env, u32 idx);
#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 18b8812f88a7..6b5c1a1d0479 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -252,6 +252,7 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
/* Expand insni[off]'s seen count to the patched range. */
data[i].seen = old_seen;
data[i].zext_dst = bpf_insn_def32(new_prog, insn + i) >= 0;
+ data[i].in_cleanup_pad = data[off + cnt - 1].in_cleanup_pad;
if (!memcmp(insn + i, original_insn, sizeof(struct bpf_insn))) {
data[i].non_stack_access =
data[off + cnt - 1].non_stack_access;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index a7a3c4b4d975..680ae191aa3f 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10531,6 +10531,10 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
* callbacks
*/
env->subprog_info[subprog].is_cb = true;
+ err = bpf_cleanup_check_callback(env, subprog);
+ if (err)
+ return err;
+
if (bpf_pseudo_kfunc_call(insn) &&
!is_callback_calling_kfunc(insn->imm)) {
verifier_bug(env, "kfunc %s#%d not marked as callback-calling",
@@ -21738,6 +21742,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
if (ret < 0)
goto skip_full_check;
+ ret = bpf_check_cleanup_exceptions(env);
+ if (ret < 0)
+ goto skip_full_check;
+
ret = bpf_compute_postorder(env);
if (ret < 0)
goto skip_full_check;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
2026-09-21 21:01 ` [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch Yonghong Song
@ 2026-09-21 21:20 ` sashiko-bot
2026-09-22 3:39 ` Yonghong Song
2026-09-21 21:56 ` bot+bpf-ci
2026-09-22 0:30 ` Eduard Zingerman
2 siblings, 1 reply; 80+ messages in thread
From: sashiko-bot @ 2026-09-21 21:20 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: kernel stack leak via >5 arg BPF-to-BPF calls in exception pads
--
commit 961251cb485953d7f6bd99ead3712d0e8268b8eb
Author: Yonghong Song <yonghong.song@linux.dev>
bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
This commit adds verification logic to reject specific control flow shapes
that bpf_throw() and the exception cleanup handling cannot safely dispatch.
It performs post-CFG checks using bpf_check_cleanup_exceptions() to inspect
pad reachability, catch pads, and unwinding safety.
> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
> index 67af78baa558d..da8fa6eb7e4b9 100644
> --- a/kernel/bpf/exception.c
> +++ b/kernel/bpf/exception.c
[ ... ]
> +static int cleanup_check_pad_insn(struct bpf_verifier_env *env, u32 i)
> +{
> + struct bpf_insn *insn = &env->prog->insnsi[i];
> +
> + if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call) {
> + verbose(env,
> + "bpf_tail_call() at insn %u is in an exception cleanup landing pad\n",
> + i);
> + return -EINVAL;
> + }
> + /* A BPF_LD_[ABS|IND] can leave the frame through its epilogue. */
> + if (BPF_CLASS(insn->code) == BPF_LD &&
> + (BPF_MODE(insn->code) == BPF_ABS || BPF_MODE(insn->code) == BPF_IND)) {
> + verbose(env,
> + "BPF_LD_[ABS|IND] at insn %u is in an exception cleanup landing pad\n",
> + i);
> + return -EINVAL;
> + }
> + if (is_stack_arg_st(insn) || is_stack_arg_stx(insn)) {
> + verbose(env,
> + "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
> + i);
> + return -EINVAL;
> + }
> + if (bpf_pseudo_kfunc_call(insn)) {
> + struct bpf_call_summary cs;
> +
> + if (bpf_get_call_summary(env, insn, &cs) &&
> + cs.arg_slot_cnt > MAX_BPF_FUNC_REG_ARGS) {
[Severity: High]
Does this omit checking BPF-to-BPF calls (bpf_pseudo_call) for on-stack
arguments?
The JIT executes exception cleanup landing pads on the unwinder's stack (rsp),
while rbp points to the BPF frame. This check restricts kfuncs from passing
arguments on the stack, but seems to miss BPF subprograms because
bpf_get_call_summary() returns false for them.
If a pad calls a BPF subprogram requiring more than 5 arguments, could the
callee read uninitialized kernel stack memory belonging to the unwinder's
caller?
> + verbose(env,
> + "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
> + i);
> + return -EINVAL;
> + }
> + }
> + return 0;
> +}
[ ... ]
> +static int cleanup_mark_pad_bodies(struct cleanup_ctx *ctx)
> +{
> + struct bpf_verifier_env *env = ctx->env;
> + u32 i, sp = 0;
> + int ret;
[ ... ]
> + while (sp) {
> + u32 j = ctx->stack[--sp];
> + enum cleanup_insn_kind kind;
> + int next, target, sub;
> + u32 start, end;
> +
> + ret = cleanup_check_pad_insn(env, j);
> + if (ret)
> + return ret;
> +
> + sub = cleanup_subprog_of(env, j);
> + start = env->subprog_info[sub].start;
> + end = env->subprog_info[sub + 1].start;
> + kind = cleanup_succ(env, j, start, end, &next, &target);
> +
> + if (kind == CLEANUP_INSN_CALL) {
> + /* check_subprogs() registered every call target. */
> + int callee = cleanup_subprog_of(env, j + env->prog->insnsi[j].imm + 1);
> +
> + if (env->subprog_info[callee].might_throw) {
[Severity: High]
Since cleanup_mark_pad_bodies() allows BPF-to-BPF calls in landing pads as
long as they don't throw, should it also verify that the callee does not
take on-stack arguments?
For example, by checking if env->subprog_info[callee].arg_slot_cnt <=
MAX_BPF_FUNC_REG_ARGS?
Without this check, a subprogram requiring more than 5 arguments could be
called, potentially leading to the kernel stack data leak mentioned above.
> + verbose(env,
> + "cleanup landing pad calls subprog %d at insn %u, which can throw while an exception is in flight\n",
> + callee, j);
> + return -EINVAL;
> + }
> + }
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260921210033.1715000-1-yonghong.song@linux.dev?part=7
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
2026-09-21 21:20 ` sashiko-bot
@ 2026-09-22 3:39 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-22 3:39 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/21/26 2:20 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf: kernel stack leak via >5 arg BPF-to-BPF calls in exception pads
> --
>
> commit 961251cb485953d7f6bd99ead3712d0e8268b8eb
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
>
> This commit adds verification logic to reject specific control flow shapes
> that bpf_throw() and the exception cleanup handling cannot safely dispatch.
> It performs post-CFG checks using bpf_check_cleanup_exceptions() to inspect
> pad reachability, catch pads, and unwinding safety.
>
>> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
>> index 67af78baa558d..da8fa6eb7e4b9 100644
>> --- a/kernel/bpf/exception.c
>> +++ b/kernel/bpf/exception.c
> [ ... ]
>
>> +static int cleanup_check_pad_insn(struct bpf_verifier_env *env, u32 i)
>> +{
>> + struct bpf_insn *insn = &env->prog->insnsi[i];
>> +
>> + if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call) {
>> + verbose(env,
>> + "bpf_tail_call() at insn %u is in an exception cleanup landing pad\n",
>> + i);
>> + return -EINVAL;
>> + }
>> + /* A BPF_LD_[ABS|IND] can leave the frame through its epilogue. */
>> + if (BPF_CLASS(insn->code) == BPF_LD &&
>> + (BPF_MODE(insn->code) == BPF_ABS || BPF_MODE(insn->code) == BPF_IND)) {
>> + verbose(env,
>> + "BPF_LD_[ABS|IND] at insn %u is in an exception cleanup landing pad\n",
>> + i);
>> + return -EINVAL;
>> + }
>> + if (is_stack_arg_st(insn) || is_stack_arg_stx(insn)) {
>> + verbose(env,
>> + "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
>> + i);
>> + return -EINVAL;
>> + }
>> + if (bpf_pseudo_kfunc_call(insn)) {
>> + struct bpf_call_summary cs;
>> +
>> + if (bpf_get_call_summary(env, insn, &cs) &&
>> + cs.arg_slot_cnt > MAX_BPF_FUNC_REG_ARGS) {
> [Severity: High]
> Does this omit checking BPF-to-BPF calls (bpf_pseudo_call) for on-stack
> arguments?
>
> The JIT executes exception cleanup landing pads on the unwinder's stack (rsp),
> while rbp points to the BPF frame. This check restricts kfuncs from passing
> arguments on the stack, but seems to miss BPF subprograms because
> bpf_get_call_summary() returns false for them.
>
> If a pad calls a BPF subprogram requiring more than 5 arguments, could the
> callee read uninitialized kernel stack memory belonging to the unwinder's
> caller?
We should be okay here. As the code explains, if the number of reg-level argument slots
more than 5, verifier will reject.
>> + verbose(env,
>> + "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
>> + i);
>> + return -EINVAL;
>> + }
>> + }
>> + return 0;
>> +}
> [ ... ]
>
>> +static int cleanup_mark_pad_bodies(struct cleanup_ctx *ctx)
>> +{
>> + struct bpf_verifier_env *env = ctx->env;
>> + u32 i, sp = 0;
>> + int ret;
> [ ... ]
>
>> + while (sp) {
>> + u32 j = ctx->stack[--sp];
>> + enum cleanup_insn_kind kind;
>> + int next, target, sub;
>> + u32 start, end;
>> +
>> + ret = cleanup_check_pad_insn(env, j);
>> + if (ret)
>> + return ret;
>> +
>> + sub = cleanup_subprog_of(env, j);
>> + start = env->subprog_info[sub].start;
>> + end = env->subprog_info[sub + 1].start;
>> + kind = cleanup_succ(env, j, start, end, &next, &target);
>> +
>> + if (kind == CLEANUP_INSN_CALL) {
>> + /* check_subprogs() registered every call target. */
>> + int callee = cleanup_subprog_of(env, j + env->prog->insnsi[j].imm + 1);
>> +
>> + if (env->subprog_info[callee].might_throw) {
> [Severity: High]
> Since cleanup_mark_pad_bodies() allows BPF-to-BPF calls in landing pads as
> long as they don't throw, should it also verify that the callee does not
> take on-stack arguments?
>
> For example, by checking if env->subprog_info[callee].arg_slot_cnt <=
> MAX_BPF_FUNC_REG_ARGS?
>
> Without this check, a subprogram requiring more than 5 arguments could be
> called, potentially leading to the kernel stack data leak mentioned above.
We should be okay here. Commit "bpf: Walk the exception unwind in the verifier",
we have
+static u32 unwind_pop_frame(struct bpf_verifier_env *env)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ struct bpf_func_state *callee = state->frame[state->curframe];
+ u32 callsite = callee->callsite;
+ struct bpf_func_state *caller;
+
+ caller = state->frame[state->curframe - 1];
+ account_processed_insns(env, callee, caller);
+ free_func_state(callee);
+ state->frame[state->curframe--] = NULL;
+ invalidate_outgoing_stack_args(env, caller);
+ return callsite;
+}
stack arguments will be invlidated, so once pop frame and go to
landing_pad code section, verification will fail due to un-initialized
stack arguments.
>
>> + verbose(env,
>> + "cleanup landing pad calls subprog %d at insn %u, which can throw while an exception is in flight\n",
>> + callee, j);
>> + return -EINVAL;
>> + }
>> + }
> [ ... ]
>
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
2026-09-21 21:01 ` [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch Yonghong Song
2026-09-21 21:20 ` sashiko-bot
@ 2026-09-21 21:56 ` bot+bpf-ci
2026-09-22 3:44 ` Yonghong Song
2026-09-22 0:30 ` Eduard Zingerman
2 siblings, 1 reply; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-21 21:56 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 5812 bytes --]
> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
> index 67af78baa558..da8fa6eb7e4b 100644
> --- a/kernel/bpf/exception.c
> +++ b/kernel/bpf/exception.c
> @@ -3,12 +3,24 @@
> #include <linux/bpf.h>
> #include <linux/bpf_verifier.h>
> #include <linux/btf.h>
> #include <linux/btf_ids.h>
> #include <linux/filter.h>
> #include <linux/slab.h>
> +#include <linux/sort.h>
This isn't a bug, but is <linux/sort.h> used by anything in this patch,
or would it sit better in the later patch that adds the sort() call?
A scan of exception.c and exception.h at this commit finds no sort(),
sort_r(), cmp_func_t, cmp_r_func_t or swap_func_t. The first sort() call
in the file arrives three commits later, in 07c93d6858af ("bpf: Dispatch
exception cleanup pads from bpf_throw()"), which adds sort(at, cnt,
sizeof(*at), cmp_u32, NULL) without needing to touch the include block
because this patch already added the header.
[ ... ]
> @@ -29,12 +154,260 @@ static void cleanup_mark_kfunc_sites(struct bpf_verifier_env *env)
> +static int cleanup_mark_pad_bodies(struct cleanup_ctx *ctx)
> +{
> + struct bpf_verifier_env *env = ctx->env;
> + u32 i, sp = 0;
> + int ret;
> +
> + for (i = 0; i < env->cleanup_info_cnt; i++) {
> + u32 pad = env->cleanup_info[i].landing_pad_off;
> +
> + if (in_pad(env, pad))
> + continue;
> +
> + ret = cleanup_pad_is_catch(ctx, pad);
> + if (ret < 0)
> + return ret;
> + if (ret) {
> + verbose(env,
> + "catch landing pad %u is not supported yet, only cleanup pads that resume\n",
> + pad);
> + return -EOPNOTSUPP;
> + }
> + env->insn_aux_data[pad].in_cleanup_pad = true;
> + ctx->stack[sp++] = pad;
> + }
> +
> + while (sp) {
> + u32 j = ctx->stack[--sp];
> + enum cleanup_insn_kind kind;
> + int next, target, sub;
> + u32 start, end;
> +
> + ret = cleanup_check_pad_insn(env, j);
> + if (ret)
> + return ret;
> +
> + sub = cleanup_subprog_of(env, j);
> + start = env->subprog_info[sub].start;
> + end = env->subprog_info[sub + 1].start;
> + kind = cleanup_succ(env, j, start, end, &next, &target);
> +
> + if (kind == CLEANUP_INSN_CALL) {
> + /* check_subprogs() registered every call target. */
> + int callee = cleanup_subprog_of(env, j + env->prog->insnsi[j].imm + 1);
This isn't a bug, but the comment credits the wrong verifier pass.
check_subprogs() (kernel/bpf/verifier.c:3081-3146) does not register or
validate call targets at all - it explicitly skips them:
if (BPF_OP(code) == BPF_CALL)
goto next;
Its only jobs are setting has_tail_call/has_ld_abs/exit_idx and checking
that jump targets stay inside the containing subprog. The guarantee the
comment is reaching for actually comes from add_subprogs()
(verifier.c:2985-3053), which does
ret = add_subprog(env, i + insn->imm + 1);
for every bpf_pseudo_call() insn, with add_subprog() rejecting off < 0 ||
off >= insn_cnt. That is what makes cleanup_subprog_of() unable to return
-1 here, so the code is correct - only the comment is wrong.
A reader who follows the comment to check_subprogs() finds the opposite of
what it claims, which is worse than no comment, since this is the
documented reason an array index is left unchecked. Could the comment say
"add_subprogs() registered every call target" instead?
> +
> + if (env->subprog_info[callee].might_throw) {
> + verbose(env,
> + "cleanup landing pad calls subprog %d at insn %u, which can throw while an exception is in flight\n",
> + callee, j);
> + return -EINVAL;
> + }
> + }
> +
> + if (next >= 0 && !in_pad(env, next)) {
> + env->insn_aux_data[next].in_cleanup_pad = true;
> + ctx->stack[sp++] = next;
> + }
> + if (target >= 0 && !in_pad(env, target)) {
> + env->insn_aux_data[target].in_cleanup_pad = true;
> + ctx->stack[sp++] = target;
> + }
> + }
> + return 0;
> +}
[ ... ]
> @@ -154,12 +448,43 @@ int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env)
> +static int cleanup_check_pad_insn(struct bpf_verifier_env *env, u32 i)
> +{
> + struct bpf_insn *insn = &env->prog->insnsi[i];
> +
> + if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call) {
> + verbose(env,
> + "bpf_tail_call() at insn %u is in an exception cleanup landing pad\n",
> + i);
> + return -EINVAL;
> + }
> + /* A BPF_LD_[ABS|IND] can leave the frame through its epilogue. */
> + if (BPF_CLASS(insn->code) == BPF_LD &&
> + (BPF_MODE(insn->code) == BPF_ABS || BPF_MODE(insn->code) == BPF_IND)) {
> + verbose(env,
> + "BPF_LD_[ABS|IND] at insn %u is in an exception cleanup landing pad\n",
> + i);
> + return -EINVAL;
> + }
> + if (is_stack_arg_st(insn) || is_stack_arg_stx(insn)) {
> + verbose(env,
> + "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
> + i);
> + return -EINVAL;
> + }
> + if (bpf_pseudo_kfunc_call(insn)) {
> + struct bpf_call_summary cs;
> +
> + if (bpf_get_call_summary(env, insn, &cs) &&
> + cs.arg_slot_cnt > MAX_BPF_FUNC_REG_ARGS) {
> + verbose(env,
> + "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
> + i);
> + return -EINVAL;
> + }
> + }
> + return 0;
> +}
This isn't a bug, but would it be clearer to funnel both on-stack-argument
branches through a single labelled exit, or to differentiate the two
messages (e.g. mention the kfunc argument spill in the second one) so the
log says which shape was rejected?
Because the two messages are identical, the log does not distinguish which
of the two shapes was rejected, and a future edit to the wording has to be
applied in both places.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35656368472
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
2026-09-21 21:56 ` bot+bpf-ci
@ 2026-09-22 3:44 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-22 3:44 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/21/26 2:56 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
>> index 67af78baa558..da8fa6eb7e4b 100644
>> --- a/kernel/bpf/exception.c
>> +++ b/kernel/bpf/exception.c
>> @@ -3,12 +3,24 @@
>> #include <linux/bpf.h>
>> #include <linux/bpf_verifier.h>
>> #include <linux/btf.h>
>> #include <linux/btf_ids.h>
>> #include <linux/filter.h>
>> #include <linux/slab.h>
>> +#include <linux/sort.h>
> This isn't a bug, but is <linux/sort.h> used by anything in this patch,
> or would it sit better in the later patch that adds the sort() call?
>
> A scan of exception.c and exception.h at this commit finds no sort(),
> sort_r(), cmp_func_t, cmp_r_func_t or swap_func_t. The first sort() call
> in the file arrives three commits later, in 07c93d6858af ("bpf: Dispatch
> exception cleanup pads from bpf_throw()"), which adds sort(at, cnt,
> sizeof(*at), cmp_u32, NULL) without needing to touch the include block
> because this patch already added the header.
Okay, will do. But I will try Eduard's suggestion to see whether
this is necessary.
>
> [ ... ]
>
>> @@ -29,12 +154,260 @@ static void cleanup_mark_kfunc_sites(struct bpf_verifier_env *env)
>> +static int cleanup_mark_pad_bodies(struct cleanup_ctx *ctx)
>> +{
>> + struct bpf_verifier_env *env = ctx->env;
>> + u32 i, sp = 0;
>> + int ret;
>> +
>> + for (i = 0; i < env->cleanup_info_cnt; i++) {
>> + u32 pad = env->cleanup_info[i].landing_pad_off;
>> +
>> + if (in_pad(env, pad))
>> + continue;
>> +
>> + ret = cleanup_pad_is_catch(ctx, pad);
>> + if (ret < 0)
>> + return ret;
>> + if (ret) {
>> + verbose(env,
>> + "catch landing pad %u is not supported yet, only cleanup pads that resume\n",
>> + pad);
>> + return -EOPNOTSUPP;
>> + }
>> + env->insn_aux_data[pad].in_cleanup_pad = true;
>> + ctx->stack[sp++] = pad;
>> + }
>> +
>> + while (sp) {
>> + u32 j = ctx->stack[--sp];
>> + enum cleanup_insn_kind kind;
>> + int next, target, sub;
>> + u32 start, end;
>> +
>> + ret = cleanup_check_pad_insn(env, j);
>> + if (ret)
>> + return ret;
>> +
>> + sub = cleanup_subprog_of(env, j);
>> + start = env->subprog_info[sub].start;
>> + end = env->subprog_info[sub + 1].start;
>> + kind = cleanup_succ(env, j, start, end, &next, &target);
>> +
>> + if (kind == CLEANUP_INSN_CALL) {
>> + /* check_subprogs() registered every call target. */
>> + int callee = cleanup_subprog_of(env, j + env->prog->insnsi[j].imm + 1);
> This isn't a bug, but the comment credits the wrong verifier pass.
>
> check_subprogs() (kernel/bpf/verifier.c:3081-3146) does not register or
> validate call targets at all - it explicitly skips them:
>
> if (BPF_OP(code) == BPF_CALL)
> goto next;
>
> Its only jobs are setting has_tail_call/has_ld_abs/exit_idx and checking
> that jump targets stay inside the containing subprog. The guarantee the
> comment is reaching for actually comes from add_subprogs()
> (verifier.c:2985-3053), which does
>
> ret = add_subprog(env, i + insn->imm + 1);
>
> for every bpf_pseudo_call() insn, with add_subprog() rejecting off < 0 ||
> off >= insn_cnt. That is what makes cleanup_subprog_of() unable to return
> -1 here, so the code is correct - only the comment is wrong.
>
> A reader who follows the comment to check_subprogs() finds the opposite of
> what it claims, which is worse than no comment, since this is the
> documented reason an array index is left unchecked. Could the comment say
> "add_subprogs() registered every call target" instead?
Sure, we can do.
>
>> +
>> + if (env->subprog_info[callee].might_throw) {
>> + verbose(env,
>> + "cleanup landing pad calls subprog %d at insn %u, which can throw while an exception is in flight\n",
>> + callee, j);
>> + return -EINVAL;
>> + }
>> + }
>> +
>> + if (next >= 0 && !in_pad(env, next)) {
>> + env->insn_aux_data[next].in_cleanup_pad = true;
>> + ctx->stack[sp++] = next;
>> + }
>> + if (target >= 0 && !in_pad(env, target)) {
>> + env->insn_aux_data[target].in_cleanup_pad = true;
>> + ctx->stack[sp++] = target;
>> + }
>> + }
>> + return 0;
>> +}
> [ ... ]
>
>> @@ -154,12 +448,43 @@ int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env)
>> +static int cleanup_check_pad_insn(struct bpf_verifier_env *env, u32 i)
>> +{
>> + struct bpf_insn *insn = &env->prog->insnsi[i];
>> +
>> + if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call) {
>> + verbose(env,
>> + "bpf_tail_call() at insn %u is in an exception cleanup landing pad\n",
>> + i);
>> + return -EINVAL;
>> + }
>> + /* A BPF_LD_[ABS|IND] can leave the frame through its epilogue. */
>> + if (BPF_CLASS(insn->code) == BPF_LD &&
>> + (BPF_MODE(insn->code) == BPF_ABS || BPF_MODE(insn->code) == BPF_IND)) {
>> + verbose(env,
>> + "BPF_LD_[ABS|IND] at insn %u is in an exception cleanup landing pad\n",
>> + i);
>> + return -EINVAL;
>> + }
>> + if (is_stack_arg_st(insn) || is_stack_arg_stx(insn)) {
>> + verbose(env,
>> + "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
>> + i);
>> + return -EINVAL;
>> + }
>> + if (bpf_pseudo_kfunc_call(insn)) {
>> + struct bpf_call_summary cs;
>> +
>> + if (bpf_get_call_summary(env, insn, &cs) &&
>> + cs.arg_slot_cnt > MAX_BPF_FUNC_REG_ARGS) {
>> + verbose(env,
>> + "insn %u passes an on-stack call argument in an exception cleanup landing pad\n",
>> + i);
>> + return -EINVAL;
>> + }
>> + }
>> + return 0;
>> +}
> This isn't a bug, but would it be clearer to funnel both on-stack-argument
> branches through a single labelled exit, or to differentiate the two
> messages (e.g. mention the kfunc argument spill in the second one) so the
> log says which shape was rejected?
I think it is okay.
>
> Because the two messages are identical, the log does not distinguish which
> of the two shapes was rejected, and a future edit to the wording has to be
> applied in both places.
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35656368472
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
2026-09-21 21:01 ` [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch Yonghong Song
2026-09-21 21:20 ` sashiko-bot
2026-09-21 21:56 ` bot+bpf-ci
@ 2026-09-22 0:30 ` Eduard Zingerman
2026-09-22 3:45 ` Yonghong Song
2 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-22 0:30 UTC (permalink / raw)
To: Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Mon, 2026-09-21 at 14:01 -0700, Yonghong Song wrote:
> The table is on the instructions and the landing pads are in the control
> flow graph. What is left before bpf_throw() can be taught to dispatch them
> is to work out what that walk will need, and refuse the shapes it could not
> handle.
>
> bpf_check_cleanup_exceptions() runs after bpf_check_cfg(). Everything it
> needs is control flow, so it reads what that walk has already worked out
> rather than working it out again:
>
> subprog_info.might_throw which subprograms an exception may leave,
> closed over the call graph by
> merge_callee_effects() as the walk pops each
> callee
> cleanup_reachability() what each instruction can reach -- a resume, a
> plain exit, a throw, an indirect jump
> cleanup_mark_pad_bodies() which instructions only ever run with an
> exception already in flight
>
> What it refuses:
>
> - a pad that reaches both a resume and a plain exit, or neither: nothing
> says whether it is a cleanup pad or a catch pad
> - a catch pad, which ends in a plain exit: the walker calls a pad as a
> subroutine and cannot hand a frame back its own execution
> - a throw in a pad, or a call from a pad to a subprogram that can throw:
> a second unwind over frames the first is still discarding
> - a bpf_unwind_resume() outside a pad body
> - a subprogram that may unwind used as a helper callback: the helper's
> own kernel frame would end the walk before it found a boundary
> - a tail call, an indirect jump, or an outgoing on-stack call argument in
> a pad body, all of which touch a stack the pad does not own
> - a BPF_LD_[ABS|IND] in a pad body: a failed load leaves the subprogram
> through the hidden exit gen_ld_abs() patches in, and an exit in a pad
> is the epilogue, which pops off the walker's stack and returns through
> the frame the walk is discarding
>
> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
> ---
Why is it necessary to hand-roll a CFG traversal and a separate pass
for this check? Given that bpf verifier state already maintains
`unwinding' flag, the instructions properties can be checked from
do_check_insn(), e.g.:
- bpf_throw() -- reject if unwinding
- call to a throwing global -- reject if unwinding
- bpf_unwind_resume() -- require unwinding and the pad-owning frame
- BPF_EXIT -- reject in the pad-owning frame, allow in callees
- LD_{ABS,IND} -- reject if unwinding
I think that would take much less code, wdyt?
> include/linux/bpf_verifier.h | 1 +
> kernel/bpf/exception.c | 404 +++++++++++++++++++++++++++++++++++
> kernel/bpf/exception.h | 2 +
> kernel/bpf/fixups.c | 1 +
> kernel/bpf/verifier.c | 8 +
> 5 files changed, 416 insertions(+)
...
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
2026-09-22 0:30 ` Eduard Zingerman
@ 2026-09-22 3:45 ` Yonghong Song
2026-09-22 21:43 ` Eduard Zingerman
0 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-22 3:45 UTC (permalink / raw)
To: Eduard Zingerman, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On 9/21/26 5:30 PM, Eduard Zingerman wrote:
> On Mon, 2026-09-21 at 14:01 -0700, Yonghong Song wrote:
>> The table is on the instructions and the landing pads are in the control
>> flow graph. What is left before bpf_throw() can be taught to dispatch them
>> is to work out what that walk will need, and refuse the shapes it could not
>> handle.
>>
>> bpf_check_cleanup_exceptions() runs after bpf_check_cfg(). Everything it
>> needs is control flow, so it reads what that walk has already worked out
>> rather than working it out again:
>>
>> subprog_info.might_throw which subprograms an exception may leave,
>> closed over the call graph by
>> merge_callee_effects() as the walk pops each
>> callee
>> cleanup_reachability() what each instruction can reach -- a resume, a
>> plain exit, a throw, an indirect jump
>> cleanup_mark_pad_bodies() which instructions only ever run with an
>> exception already in flight
>>
>> What it refuses:
>>
>> - a pad that reaches both a resume and a plain exit, or neither: nothing
>> says whether it is a cleanup pad or a catch pad
>> - a catch pad, which ends in a plain exit: the walker calls a pad as a
>> subroutine and cannot hand a frame back its own execution
>> - a throw in a pad, or a call from a pad to a subprogram that can throw:
>> a second unwind over frames the first is still discarding
>> - a bpf_unwind_resume() outside a pad body
>> - a subprogram that may unwind used as a helper callback: the helper's
>> own kernel frame would end the walk before it found a boundary
>> - a tail call, an indirect jump, or an outgoing on-stack call argument in
>> a pad body, all of which touch a stack the pad does not own
>> - a BPF_LD_[ABS|IND] in a pad body: a failed load leaves the subprogram
>> through the hidden exit gen_ld_abs() patches in, and an exit in a pad
>> is the epilogue, which pops off the walker's stack and returns through
>> the frame the walk is discarding
>>
>> Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
>> ---
> Why is it necessary to hand-roll a CFG traversal and a separate pass
> for this check? Given that bpf verifier state already maintains
> `unwinding' flag, the instructions properties can be checked from
> do_check_insn(), e.g.:
> - bpf_throw() -- reject if unwinding
> - call to a throwing global -- reject if unwinding
> - bpf_unwind_resume() -- require unwinding and the pad-owning frame
> - BPF_EXIT -- reject in the pad-owning frame, allow in callees
> - LD_{ABS,IND} -- reject if unwinding
>
> I think that would take much less code, wdyt?
This is a good idea. Let me try. Thanks!
>
>> include/linux/bpf_verifier.h | 1 +
>> kernel/bpf/exception.c | 404 +++++++++++++++++++++++++++++++++++
>> kernel/bpf/exception.h | 2 +
>> kernel/bpf/fixups.c | 1 +
>> kernel/bpf/verifier.c | 8 +
>> 5 files changed, 416 insertions(+)
> ...
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
2026-09-22 3:45 ` Yonghong Song
@ 2026-09-22 21:43 ` Eduard Zingerman
2026-09-23 3:11 ` Yonghong Song
0 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-22 21:43 UTC (permalink / raw)
To: Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Mon, 2026-09-21 at 20:45 -0700, Yonghong Song wrote:
...
> > Why is it necessary to hand-roll a CFG traversal and a separate pass
> > for this check? Given that bpf verifier state already maintains
> > `unwinding' flag, the instructions properties can be checked from
> > do_check_insn(), e.g.:
> > - bpf_throw() -- reject if unwinding
> > - call to a throwing global -- reject if unwinding
> > - bpf_unwind_resume() -- require unwinding and the pad-owning frame
> > - BPF_EXIT -- reject in the pad-owning frame, allow in callees
> > - LD_{ABS,IND} -- reject if unwinding
> >
> > I think that would take much less code, wdyt?
>
> This is a good idea. Let me try. Thanks!
Also, regarding the tail calls. It appears that the following chain
can call bpf_throw from a landing pad:
bpf_throw() -> landing pad -> global procedure call -> tail call -> bpf_throw()
would this work as expected?
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
2026-09-22 21:43 ` Eduard Zingerman
@ 2026-09-23 3:11 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-23 3:11 UTC (permalink / raw)
To: Eduard Zingerman, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On 9/22/26 2:43 PM, Eduard Zingerman wrote:
> On Mon, 2026-09-21 at 20:45 -0700, Yonghong Song wrote:
>
> ...
>
>>> Why is it necessary to hand-roll a CFG traversal and a separate pass
>>> for this check? Given that bpf verifier state already maintains
>>> `unwinding' flag, the instructions properties can be checked from
>>> do_check_insn(), e.g.:
>>> - bpf_throw() -- reject if unwinding
>>> - call to a throwing global -- reject if unwinding
>>> - bpf_unwind_resume() -- require unwinding and the pad-owning frame
>>> - BPF_EXIT -- reject in the pad-owning frame, allow in callees
>>> - LD_{ABS,IND} -- reject if unwinding
>>>
>>> I think that would take much less code, wdyt?
This indeed reduces amount of codes by more than 2/3 for this patch.
>> This is a good idea. Let me try. Thanks!
> Also, regarding the tail calls. It appears that the following chain
> can call bpf_throw from a landing pad:
>
> bpf_throw() -> landing pad -> global procedure call -> tail call -> bpf_throw()
>
> would this work as expected?
Yes, it works. tail_call itself is the terminator as it is the main prog,
bpf_throw() won't cross main prog boundary.
^ permalink raw reply [flat|nested] 80+ messages in thread
* [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (6 preceding siblings ...)
2026-09-21 21:01 ` [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch Yonghong Song
@ 2026-09-21 21:01 ` Yonghong Song
2026-09-21 21:40 ` sashiko-bot
` (3 more replies)
2026-09-21 21:01 ` [PATCH bpf-next v4 09/20] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
` (12 subsequent siblings)
20 siblings, 4 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:01 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
bpf_throw() is about to stop discarding the frames it unwinds and start
running their landing pads instead. Have the verifier walk the same thing,
step for step:
at a throw, with the frame chain in env->cur_state->frame[]
if a record covers the call site control is at
run that record's landing pad in this frame, and on
reaching its resume, pop the frame and ask again
else
pop the frame and ask again
until frame 0 asks, which is the boundary: deliver
Doing it this way is what keeps the resource rules unchanged. Whatever a
pad releases is released in the verifier state too, so by the time the walk
reaches the boundary the state says exactly what the program will really
hold there -- and check_resource_leak(), which used to fire the moment a
throw was seen, simply moves to the end of the walk. A frame that no record
covers contributes nothing, so anything it held is still held when the walk
ends, and that is what gets reported.
A pad starts in the state the walker hands it: its own frame's stack and
callee-saved registers, which the walker puts back, nothing in the
caller-saved ones, and BPF_PAD_ENTRY_R0 in r0. LLVM names r0 as both the
exception pointer and the exception selector register, so a pad reads it
before anything else and is free to store what it read, which means the
value has to be a constant the verifier knows. It gets a header of its own
because the two sides that have to agree on it are the verifier and the
arch dispatchers, which are assembly.
Two things about that walk have to be said out loud, because the
instruction stream does not say them.
A resume belongs to the frame whose landing pad the walker called. Each JIT
lowers bpf_unwind_resume() as the way back out of a pad -- a bare return on
x86-64, a branch to the saved address on arm64 -- and that reaches the
walker only from there. An exception being in flight is not enough to tell:
a pad may call a subprogram that has a landing pad of its own and reaches
it by ordinary control flow, and the resume in it is in a pad body by every
static measure. So the state records the frame the unwind entered a pad in,
states_equal() keeps two such states apart, and a resume in any other frame
is refused.
The edge from a throw to a pad crosses frames with no instruction in
between to account for them, and mark_chain_precision() walks that history
backwards. Left alone it stays in the pad's frame while it reads the
callee's instructions: a request for the pad frame's r6 is cleared by the
callee's own write to r6 -- so the caller's definition never becomes
precise -- or reaches the call instruction still set and trips "static
subprog unexpected regs". The history entry for a pad therefore records how
many frames the unwind popped, and the backtrack enters that many, the way
it enters one at a time for BPF_EXIT. A throwing global subprogram gets no
frame of its own, and its exception arrives at the landing pad rather than
at the instruction after the call, which the check in backtrack_insn() has
to allow for.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf_cleanup_abi.h | 16 ++++
include/linux/bpf_verifier.h | 6 +-
kernel/bpf/backtrack.c | 20 ++++-
kernel/bpf/states.c | 6 ++
kernel/bpf/verifier.c | 150 +++++++++++++++++++++++++++-----
5 files changed, 169 insertions(+), 29 deletions(-)
create mode 100644 include/linux/bpf_cleanup_abi.h
diff --git a/include/linux/bpf_cleanup_abi.h b/include/linux/bpf_cleanup_abi.h
new file mode 100644
index 000000000000..40b42c78fd2b
--- /dev/null
+++ b/include/linux/bpf_cleanup_abi.h
@@ -0,0 +1,16 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef _LINUX_BPF_CLEANUP_ABI_H
+#define _LINUX_BPF_CLEANUP_ABI_H
+
+/*
+ * Value arch_bpf_run_cleanup_pad() leaves in r0 on the way into a landing pad.
+ * It has to be a constant the verifier knows: LLVM names r0 as both the
+ * exception pointer and the exception selector register, so every pad reads it
+ * before anything else and is free to store what it read. It gets a header of
+ * its own because the two sides that have to agree on it are the verifier and
+ * the arch dispatchers, which are assembly.
+ */
+#define BPF_PAD_ENTRY_R0 1
+
+#endif /* _LINUX_BPF_CLEANUP_ABI_H */
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index aa631bb45f76..076739d4d974 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -429,7 +429,8 @@ struct bpf_jmp_history_entry {
u32 prev_idx : 20;
/* special INSN_F_xxx flags */
u32 flags : 4;
- u32 : 8;
+ u32 unwind_frames : 4; /* frames the unwind popped to get here */
+ u32 : 4;
/*
* additional registers that need precision tracking when this
* jump is backtracked, vector of five 11-bit records
@@ -509,6 +510,8 @@ struct bpf_verifier_state {
bool speculative;
bool in_sleepable;
+ bool unwinding; /* an exception is in flight */
+ u8 unwind_frameno; /* the frame whose landing pad the exception entered */
/* first and last insn idx of this verifier state */
u32 first_insn_idx;
@@ -991,6 +994,7 @@ struct bpf_verifier_env {
} cfg;
struct backtrack_state bt;
struct bpf_jmp_history_entry *cur_hist_ent;
+ u8 unwind_frames; /* scratch: frames the unwind popped to reach the next insn */
/* Per-callsite copy of parent's converged at_stack_in for cross-frame fills. */
struct arg_track **callsite_at_stack;
u32 pass_cnt; /* number of times do_check() was called */
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 507a366dffa4..bf5e7e6ab78f 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -4,6 +4,7 @@
#include <linux/bpf_verifier.h>
#include <linux/filter.h>
#include <linux/bitmap.h>
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -47,6 +48,7 @@ int bpf_push_jmp_history(struct bpf_verifier_env *env, struct bpf_verifier_state
p->flags = insn_flags;
p->spi = spi;
p->frame = frame;
+ p->unwind_frames = 0;
p->linked_regs = linked_regs;
cur->jmp_history_cnt = cnt;
env->cur_hist_ent = p;
@@ -419,10 +421,12 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
* extra instructions from subprog; the next
* instruction after call to global subprog
* should be literally next instruction in
- * caller program
+ * caller program -- or, if the callee threw,
+ * the landing pad of this call site
*/
- verifier_bug_if(idx + 1 != subseq_idx, env,
- "extra insn from subprog");
+ verifier_bug_if(idx + 1 != subseq_idx &&
+ bpf_cleanup_pad_of_call(env, idx) != subseq_idx,
+ env, "extra insn from subprog");
/* global subprog always sets R0 */
bt_clear_reg(bt, BPF_REG_0);
/* and if it does not set R2, main pass would catch it */
@@ -888,11 +892,11 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
}
for (i = last_idx;;) {
+ hist = get_jmp_hist_entry(st, history, i);
if (skip_first) {
err = 0;
skip_first = false;
} else {
- hist = get_jmp_hist_entry(st, history, i);
err = backtrack_insn(env, i, subseq_idx, hist, bt);
}
if (err == -ENOTSUPP) {
@@ -909,6 +913,14 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
*/
return 0;
subseq_idx = i;
+ /* This insn is a landing pad the unwind reached from
+ * a throw or a resume hist->unwind_frames frames
+ * deeper. No insn stands between the two, so enter
+ * those frames here, the way BPF_EXIT enters one.
+ */
+ for (fr = 0; hist && fr < hist->unwind_frames; fr++)
+ if (bt_subprog_enter(bt))
+ return -EFAULT;
i = get_prev_insn_idx(st, i, &history);
if (i == -ENOENT)
break;
diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
index 66fb11b6c6a7..b101baa43171 100644
--- a/kernel/bpf/states.c
+++ b/kernel/bpf/states.c
@@ -996,6 +996,12 @@ static bool states_equal(struct bpf_verifier_env *env,
if (old->in_sleepable != cur->in_sleepable)
return false;
+ if (old->unwinding != cur->unwinding)
+ return false;
+
+ if (old->unwinding && old->unwind_frameno != cur->unwind_frameno)
+ return false;
+
if (!refsafe(old, cur, &env->idmap_scratch))
return false;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 680ae191aa3f..2bc08c18ebc8 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10,6 +10,7 @@
#include <linux/slab.h>
#include <linux/bpf.h>
#include <linux/btf.h>
+#include <linux/bpf_cleanup_abi.h>
#include <linux/bpf_verifier.h>
#include <linux/filter.h>
#include <net/netlink.h>
@@ -1721,6 +1722,8 @@ int bpf_copy_verifier_state(struct bpf_verifier_state *dst_state,
return err;
dst_state->speculative = src->speculative;
dst_state->in_sleepable = src->in_sleepable;
+ dst_state->unwinding = src->unwinding;
+ dst_state->unwind_frameno = src->unwind_frameno;
dst_state->curframe = src->curframe;
dst_state->branches = src->branches;
dst_state->parent = src->parent;
@@ -10586,8 +10589,8 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
return 0;
}
-static int process_bpf_exit_full(struct bpf_verifier_env *env,
- bool *do_print_state, bool exception_exit);
+static int process_bpf_exit_full(struct bpf_verifier_env *env, bool *do_print_state);
+static int unwind_step(struct bpf_verifier_env *env, u32 callsite, int *insn_idx);
static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
int *insn_idx)
@@ -10677,7 +10680,7 @@ static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
verbose(env, "failed to push state for global subprog exception path\n");
return PTR_ERR(branch);
}
- return process_bpf_exit_full(env, NULL, true);
+ return unwind_step(env, *insn_idx, insn_idx);
}
/* continue with next insn after call */
@@ -14585,7 +14588,7 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
env->prog->call_session_cookie = true;
if (bpf_is_throw_kfunc(insn))
- return process_bpf_exit_full(env, NULL, true);
+ return unwind_step(env, insn_idx, &env->insn_idx);
return 0;
}
@@ -18510,9 +18513,108 @@ enum {
INSN_IDX_UPDATED = 2,
};
-static int process_bpf_exit_full(struct bpf_verifier_env *env,
- bool *do_print_state,
- bool exception_exit)
+static u32 unwind_pop_frame(struct bpf_verifier_env *env)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ struct bpf_func_state *callee = state->frame[state->curframe];
+ u32 callsite = callee->callsite;
+ struct bpf_func_state *caller;
+
+ caller = state->frame[state->curframe - 1];
+ account_processed_insns(env, callee, caller);
+ free_func_state(callee);
+ state->frame[state->curframe--] = NULL;
+ invalidate_outgoing_stack_args(env, caller);
+ return callsite;
+}
+
+/*
+ * A pad runs with its frame's stack and callee-saved registers, which the
+ * walker puts back, nothing in the caller-saved ones, and r0 as the arch
+ * dispatchers leave it.
+ */
+static void unwind_enter_pad(struct bpf_verifier_env *env)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ struct bpf_func_state *frame = cur_func(env);
+
+ state->unwind_frameno = state->curframe;
+ clear_caller_saved_regs(env, frame->regs);
+ mark_reg_unknown(env, frame->regs, BPF_REG_0);
+ __mark_reg_known(&frame->regs[BPF_REG_0], BPF_PAD_ENTRY_R0);
+}
+
+static int unwind_finish(struct bpf_verifier_env *env)
+{
+ int err = check_resource_leak(env, true, true, "bpf_throw");
+
+ if (err)
+ return err;
+ return PROCESS_BPF_EXIT;
+}
+
+static int unwind_step(struct bpf_verifier_env *env, u32 callsite, int *insn_idx)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ u8 popped = 0;
+
+ state->unwinding = true;
+ for (;;) {
+ int pad = bpf_cleanup_pad_of_call(env, callsite);
+
+ if (pad >= 0) {
+ unwind_enter_pad(env);
+ /*
+ * The edge from @callsite to the pad crosses @popped
+ * frames, and nothing in the instruction stream says
+ * so. Record it for mark_chain_precision(), which has
+ * to walk back through the same frames.
+ */
+ env->unwind_frames = popped;
+ *insn_idx = pad;
+ return INSN_IDX_UPDATED;
+ }
+ if (!state->curframe)
+ return unwind_finish(env);
+ callsite = unwind_pop_frame(env);
+ popped++;
+ }
+}
+
+static int process_cleanup_resume(struct bpf_verifier_env *env, int *insn_idx)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ int err;
+
+ if (!state->unwinding) {
+ verbose(env,
+ "bpf_unwind_resume() at insn %d reached without an exception in flight\n",
+ *insn_idx);
+ return -EINVAL;
+ }
+ /*
+ * The shape this refuses: a pad calls a subprogram that has a landing
+ * pad of its own and reaches it by ordinary control flow, so the
+ * resume in it is in a pad body by every static measure. A JIT lowers
+ * a resume as the way back out of a pad, which reaches the walker only
+ * from the frame whose pad the walker called.
+ */
+ if (state->curframe != state->unwind_frameno) {
+ verbose(env,
+ "bpf_unwind_resume() at insn %d is in frame %d, not frame %d whose landing pad the exception entered\n",
+ *insn_idx, state->curframe, state->unwind_frameno);
+ return -EINVAL;
+ }
+ if (!state->curframe)
+ return unwind_finish(env);
+ err = unwind_step(env, unwind_pop_frame(env), insn_idx);
+ /* unwind_step() counted the frames it popped, not the one popped here. */
+ if (err == INSN_IDX_UPDATED)
+ env->unwind_frames++;
+ return err;
+}
+
+static int process_bpf_exit_full(struct bpf_verifier_env *env, bool *do_print_state)
{
struct bpf_func_state *cur_frame = cur_func(env);
@@ -18522,25 +18624,11 @@ static int process_bpf_exit_full(struct bpf_verifier_env *env,
* for which reference_state must match caller reference
* state when it exits.
*/
- int err = check_resource_leak(env, exception_exit,
- exception_exit || !env->cur_state->curframe,
- exception_exit ? "bpf_throw" :
+ int err = check_resource_leak(env, false, !env->cur_state->curframe,
"BPF_EXIT instruction in main prog");
if (err)
return err;
- /* The side effect of the prepare_func_exit which is
- * being skipped is that it frees bpf_func_state.
- * Typically, process_bpf_exit will only be hit with
- * outermost exit. copy_verifier_state in pop_stack will
- * handle freeing of any extra bpf_func_state left over
- * from not processing all nested function exits. We
- * also skip return code checks as they are not needed
- * for exceptional exits.
- */
- if (exception_exit)
- return PROCESS_BPF_EXIT;
-
if (env->cur_state->curframe) {
/* exit from nested function */
err = prepare_func_exit(env, &env->insn_idx);
@@ -18714,6 +18802,8 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
env->jmps_processed++;
if (opcode == BPF_CALL) {
+ if (bpf_is_unwind_resume_kfunc(insn))
+ return process_cleanup_resume(env, &env->insn_idx);
if (env->cur_state->active_locks) {
if ((insn->src_reg == BPF_REG_0 &&
insn->imm != BPF_FUNC_spin_unlock &&
@@ -18747,7 +18837,7 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
env->insn_idx += insn->imm + 1;
return INSN_IDX_UPDATED;
} else if (opcode == BPF_EXIT) {
- return process_bpf_exit_full(env, do_print_state, false);
+ return process_bpf_exit_full(env, do_print_state);
}
return check_cond_jmp_op(env, insn, &env->insn_idx);
}
@@ -18784,10 +18874,14 @@ static int do_check(struct bpf_verifier_env *env)
for (;;) {
struct bpf_insn *insn;
struct bpf_insn_aux_data *insn_aux;
+ u8 unwind_frames;
int err;
/* reset current history entry on each new instruction */
env->cur_hist_ent = NULL;
+ /* frames the unwind popped to reach this insn, if it is a pad */
+ unwind_frames = env->unwind_frames;
+ env->unwind_frames = 0;
env->prev_insn_idx = prev_insn_idx;
if (env->insn_idx >= insn_cnt) {
@@ -18851,10 +18945,18 @@ static int do_check(struct bpf_verifier_env *env)
}
}
- if (bpf_is_jmp_point(env, env->insn_idx)) {
+ /*
+ * The entry pushed here is the only record of how many frames
+ * the unwind popped to reach this insn, which the backtrack in
+ * mark_chain_precision() needs to follow the same edge. A pad
+ * is already a jump point; the second test only guards against
+ * it ever ceasing to be one.
+ */
+ if (bpf_is_jmp_point(env, env->insn_idx) || unwind_frames) {
err = bpf_push_jmp_history(env, state, 0, 0, 0, 0);
if (err)
return err;
+ env->cur_hist_ent->unwind_frames = unwind_frames;
}
if (signal_pending(current))
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier
2026-09-21 21:01 ` [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier Yonghong Song
@ 2026-09-21 21:40 ` sashiko-bot
2026-09-22 4:17 ` Yonghong Song
2026-09-21 21:56 ` bot+bpf-ci
` (2 subsequent siblings)
3 siblings, 1 reply; 80+ messages in thread
From: sashiko-bot @ 2026-09-21 21:40 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] bpf: BPF_EXIT in exception cleanup pads bypasses runtime unwinding
- [High] bpf: nested bpf_throw() in landing pads causes kernel NULL pointer dereference
- [Medium] bpf: R0 precision tracking is not cleared across exception edges
--
commit cdb6004b5b8821aea00d57d94b9ac5bcf916b6a4
Author: Yonghong Song <yonghong.song@linux.dev>
Subject: bpf: Walk the exception unwind in the verifier
This commit updates the BPF verifier to walk the exception unwinding path
step for step when bpf_throw() is called. This aligns the verifier with
the runtime unwinder to properly handle landing pads for exception cleanup.
> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
> index 507a366dffa47..bf5e7e6ab78f3 100644
> --- a/kernel/bpf/backtrack.c
> +++ b/kernel/bpf/backtrack.c
[ ... ]
> @@ -909,6 +913,14 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
> */
> return 0;
> subseq_idx = i;
> + /* This insn is a landing pad the unwind reached from
> + * a throw or a resume hist->unwind_frames frames
> + * deeper. No insn stands between the two, so enter
> + * those frames here, the way BPF_EXIT enters one.
> + */
> + for (fr = 0; hist && fr < hist->unwind_frames; fr++)
> + if (bt_subprog_enter(bt))
> + return -EFAULT;
> i = get_prev_insn_idx(st, i, &history);
> if (i == -ENOENT)
> break;
[Severity: Medium]
When traversing the exception edge here, does the backtrack correctly clear
the precision tracking for R0 in the pad's frame?
When returning from a throw to a landing pad, R0 is initialized to
BPF_PAD_ENTRY_R0 by the kernel unwinder. Since its precision might be
tracked if used in a conditional jump in the pad, this backtrack loop uses
bt_subprog_enter(bt) but does not appear to clear R0's precision mask in
the pad's frame.
This could leave R0 erroneously tracked in the pad's frame as the backtrack
proceeds, eventually reaching the program entry and triggering a
verifier_bug() that rejects valid BPF programs with -EFAULT.
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 680ae191aa3f5..2bc08c18ebc89 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
[ ... ]
> @@ -14585,7 +14588,7 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
> env->prog->call_session_cookie = true;
>
> if (bpf_is_throw_kfunc(insn))
> - return process_bpf_exit_full(env, NULL, true);
> + return unwind_step(env, insn_idx, &env->insn_idx);
>
> return 0;
> }
[Severity: High]
Is there anything preventing a nested exception from being thrown here?
There doesn't seem to be a check for env->cur_state->unwinding before
processing bpf_throw(), which allows a program to throw an exception from
within a landing pad.
At runtime, the landing pad is executed via arch_bpf_run_cleanup_pad(),
placing kernel frames on the call stack above the BPF frames. A nested call
to bpf_throw() would invoke arch_bpf_stack_walk().
Since it would encounter the non-BPF frame for arch_bpf_run_cleanup_pad(),
bpf_prog_ksym_find() returns NULL. With ctx->cnt being non-zero, the walk
halts prematurely, causing a blind dereference of ctx.aux->bpf_exception_cb
and leading to a kernel NULL pointer dereference.
[ ... ]
> @@ -18522,25 +18624,11 @@ static int process_bpf_exit_full(struct bpf_verifier_env *env,
> * for which reference_state must match caller reference
> * state when it exits.
> */
> - int err = check_resource_leak(env, exception_exit,
> - exception_exit || !env->cur_state->curframe,
> - exception_exit ? "bpf_throw" :
> + int err = check_resource_leak(env, false, !env->cur_state->curframe,
> "BPF_EXIT instruction in main prog");
> if (err)
> return err;
>
> - /* The side effect of the prepare_func_exit which is
> - * being skipped is that it frees bpf_func_state.
> - * Typically, process_bpf_exit will only be hit with
> - * outermost exit. copy_verifier_state in pop_stack will
> - * handle freeing of any extra bpf_func_state left over
> - * from not processing all nested function exits. We
> - * also skip return code checks as they are not needed
> - * for exceptional exits.
> - */
> - if (exception_exit)
> - return PROCESS_BPF_EXIT;
> -
> if (env->cur_state->curframe) {
> /* exit from nested function */
> err = prepare_func_exit(env, &env->insn_idx);
[ ... ]
> @@ -18747,7 +18837,7 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
> env->insn_idx += insn->imm + 1;
> return INSN_IDX_UPDATED;
> } else if (opcode == BPF_EXIT) {
> - return process_bpf_exit_full(env, do_print_state, false);
> + return process_bpf_exit_full(env, do_print_state);
> }
> return check_cond_jmp_op(env, insn, &env->insn_idx);
> }
[Severity: High]
Are we correctly handling BPF_EXIT when an exception is in flight?
Because process_bpf_exit_full() is called unconditionally without checking
if env->cur_state->unwinding is true, the verifier treats BPF_EXIT from a
landing pad as a normal subprogram return (when curframe > 0).
It pops the frame and continues verifying the caller's instructions, where
an attacker might place lock releases to satisfy check_resource_leak().
At runtime, BPF_EXIT emits a bare 'ret' instruction. When executed from
arch_bpf_run_cleanup_pad(), it behaves like bpf_unwind_resume() by
returning control to the kernel unwinder, which silently discards the
caller frame without executing it.
This discrepancy could allow BPF programs to bypass verifier resource
checks and leak locks or references at runtime.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260921210033.1715000-1-yonghong.song@linux.dev?part=8
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier
2026-09-21 21:40 ` sashiko-bot
@ 2026-09-22 4:17 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-22 4:17 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/21/26 2:40 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
> - [High] bpf: BPF_EXIT in exception cleanup pads bypasses runtime unwinding
> - [High] bpf: nested bpf_throw() in landing pads causes kernel NULL pointer dereference
> - [Medium] bpf: R0 precision tracking is not cleared across exception edges
> --
>
> commit cdb6004b5b8821aea00d57d94b9ac5bcf916b6a4
> Author: Yonghong Song <yonghong.song@linux.dev>
> Subject: bpf: Walk the exception unwind in the verifier
>
> This commit updates the BPF verifier to walk the exception unwinding path
> step for step when bpf_throw() is called. This aligns the verifier with
> the runtime unwinder to properly handle landing pads for exception cleanup.
>
>> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
>> index 507a366dffa47..bf5e7e6ab78f3 100644
>> --- a/kernel/bpf/backtrack.c
>> +++ b/kernel/bpf/backtrack.c
> [ ... ]
>> @@ -909,6 +913,14 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
>> */
>> return 0;
>> subseq_idx = i;
>> + /* This insn is a landing pad the unwind reached from
>> + * a throw or a resume hist->unwind_frames frames
>> + * deeper. No insn stands between the two, so enter
>> + * those frames here, the way BPF_EXIT enters one.
>> + */
>> + for (fr = 0; hist && fr < hist->unwind_frames; fr++)
>> + if (bt_subprog_enter(bt))
>> + return -EFAULT;
>> i = get_prev_insn_idx(st, i, &history);
>> if (i == -ENOENT)
>> break;
> [Severity: Medium]
> When traversing the exception edge here, does the backtrack correctly clear
> the precision tracking for R0 in the pad's frame?
>
> When returning from a throw to a landing pad, R0 is initialized to
> BPF_PAD_ENTRY_R0 by the kernel unwinder. Since its precision might be
> tracked if used in a conditional jump in the pad, this backtrack loop uses
> bt_subprog_enter(bt) but does not appear to clear R0's precision mask in
> the pad's frame.
>
> This could leave R0 erroneously tracked in the pad's frame as the backtrack
> proceeds, eventually reaching the program entry and triggering a
> verifier_bug() that rejects valid BPF programs with -EFAULT.
Good point. We should clear R0 to prevent backtracking using value 1.
>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index 680ae191aa3f5..2bc08c18ebc89 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
> [ ... ]
>> @@ -14585,7 +14588,7 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
>> env->prog->call_session_cookie = true;
>>
>> if (bpf_is_throw_kfunc(insn))
>> - return process_bpf_exit_full(env, NULL, true);
>> + return unwind_step(env, insn_idx, &env->insn_idx);
>>
>> return 0;
>> }
> [Severity: High]
> Is there anything preventing a nested exception from being thrown here?
>
> There doesn't seem to be a check for env->cur_state->unwinding before
> processing bpf_throw(), which allows a program to throw an exception from
> within a landing pad.
>
> At runtime, the landing pad is executed via arch_bpf_run_cleanup_pad(),
> placing kernel frames on the call stack above the BPF frames. A nested call
> to bpf_throw() would invoke arch_bpf_stack_walk().
>
> Since it would encounter the non-BPF frame for arch_bpf_run_cleanup_pad(),
> bpf_prog_ksym_find() returns NULL. With ctx->cnt being non-zero, the walk
> halts prematurely, causing a blind dereference of ctx.aux->bpf_exception_cb
> and leading to a kernel NULL pointer dereference.
So we do have nested exception checking:
+ if (kind == CLEANUP_INSN_CALL) {
+ /* check_subprogs() registered every call target. */
+ int callee = cleanup_subprog_of(env, j + env->prog->insnsi[j].imm + 1);
+
+ if (env->subprog_info[callee].might_throw) {
+ verbose(env,
+ "cleanup landing pad calls subprog %d at insn %u, which can throw while an ex
ception is in flight\n",
+ callee, j);
+ return -EINVAL;
+ }
+ }
and
+ if (reach & CLEANUP_REACH_THROW) {
+ verbose(ctx->env,
+ "cleanup landing pad %u can throw while an exception is in flight\n",
+ pad);
+ return -EINVAL;
+ }
We probably will be okay. The above two cases should capture nested throw's.
>
> [ ... ]
>> @@ -18522,25 +18624,11 @@ static int process_bpf_exit_full(struct bpf_verifier_env *env,
>> * for which reference_state must match caller reference
>> * state when it exits.
>> */
>> - int err = check_resource_leak(env, exception_exit,
>> - exception_exit || !env->cur_state->curframe,
>> - exception_exit ? "bpf_throw" :
>> + int err = check_resource_leak(env, false, !env->cur_state->curframe,
>> "BPF_EXIT instruction in main prog");
>> if (err)
>> return err;
>>
>> - /* The side effect of the prepare_func_exit which is
>> - * being skipped is that it frees bpf_func_state.
>> - * Typically, process_bpf_exit will only be hit with
>> - * outermost exit. copy_verifier_state in pop_stack will
>> - * handle freeing of any extra bpf_func_state left over
>> - * from not processing all nested function exits. We
>> - * also skip return code checks as they are not needed
>> - * for exceptional exits.
>> - */
>> - if (exception_exit)
>> - return PROCESS_BPF_EXIT;
>> -
>> if (env->cur_state->curframe) {
>> /* exit from nested function */
>> err = prepare_func_exit(env, &env->insn_idx);
> [ ... ]
>> @@ -18747,7 +18837,7 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>> env->insn_idx += insn->imm + 1;
>> return INSN_IDX_UPDATED;
>> } else if (opcode == BPF_EXIT) {
>> - return process_bpf_exit_full(env, do_print_state, false);
>> + return process_bpf_exit_full(env, do_print_state);
>> }
>> return check_cond_jmp_op(env, insn, &env->insn_idx);
>> }
> [Severity: High]
> Are we correctly handling BPF_EXIT when an exception is in flight?
>
> Because process_bpf_exit_full() is called unconditionally without checking
> if env->cur_state->unwinding is true, the verifier treats BPF_EXIT from a
> landing pad as a normal subprogram return (when curframe > 0).
>
> It pops the frame and continues verifying the caller's instructions, where
> an attacker might place lock releases to satisfy check_resource_leak().
>
> At runtime, BPF_EXIT emits a bare 'ret' instruction. When executed from
> arch_bpf_run_cleanup_pad(), it behaves like bpf_unwind_resume() by
> returning control to the kernel unwinder, which silently discards the
> caller frame without executing it.
>
> This discrepancy could allow BPF programs to bypass verifier resource
> checks and leak locks or references at runtime.
Similar to above.
>
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier
2026-09-21 21:01 ` [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier Yonghong Song
2026-09-21 21:40 ` sashiko-bot
@ 2026-09-21 21:56 ` bot+bpf-ci
2026-09-22 5:21 ` Yonghong Song
2026-09-22 4:08 ` Alexei Starovoitov
2026-09-22 23:43 ` Eduard Zingerman
3 siblings, 1 reply; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-21 21:56 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 12254 bytes --]
> diff --git a/include/linux/bpf_cleanup_abi.h b/include/linux/bpf_cleanup_abi.h
> new file mode 100644
> index 000000000000..40b42c78fd2b
> --- /dev/null
> +++ b/include/linux/bpf_cleanup_abi.h
[ ... ]
> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index aa631bb45f76..076739d4d974 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -429,7 +429,8 @@ struct bpf_jmp_history_entry {
> u32 prev_idx : 20;
> /* special INSN_F_xxx flags */
> u32 flags : 4;
> - u32 : 8;
> + u32 unwind_frames : 4; /* frames the unwind popped to get here */
> + u32 : 4;
[ ... ]
> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
> index 507a366dffa4..bf5e7e6ab78f 100644
> --- a/kernel/bpf/backtrack.c
> +++ b/kernel/bpf/backtrack.c
> @@ -4,6 +4,7 @@
> #include <linux/bpf_verifier.h>
> #include <linux/filter.h>
> #include <linux/bitmap.h>
> +#include "exception.h"
>
> #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
>
> @@ -47,6 +48,7 @@ int bpf_push_jmp_history(struct bpf_verifier_env *env, struct bpf_verifier_state
> p->flags = insn_flags;
> p->spi = spi;
> p->frame = frame;
> + p->unwind_frames = 0;
> p->linked_regs = linked_regs;
> cur->jmp_history_cnt = cnt;
> env->cur_hist_ent = p;
Does unwind_frames stay zero when bpf_is_state_visited() takes the pruned
path?
The landing pad is both a prune point and a jmp point from
visit_cleanup_pad_edge(), so when bpf_is_state_visited() finds an equivalent
state, its hit path in kernel/bpf/states.c pushes the history entry:
if (bpf_is_jmp_point(env, env->insn_idx))
err = bpf_push_jmp_history(env, cur, 0, 0, 0, 0);
err = err ? : propagate_precision(env, &sl->state, cur, NULL);
and do_check() then takes 'goto process_bpf_exit', skipping the block that
would write env->cur_hist_ent->unwind_frames = unwind_frames. The comment
in do_check() states this entry is the only record of how many frames the
unwind popped to reach the insn, so when the pruned path keeps the zero,
propagate_precision()->bpf_mark_chain_precision() later walks from the pad
straight to the throwing insn in the deeper frame with bt->frame still at
the pad frame?
That would let the callee's writes to r6-r9 clear the caller frame's masks,
or reach the call instruction still set and trip bt_subprog_exit() or the
static-subprog checks in backtrack_insn().
> @@ -419,10 +421,12 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
> * extra instructions from subprog; the next
> * instruction after call to global subprog
> * should be literally next instruction in
> - * caller program
> + * caller program -- or, if the callee threw,
> + * the landing pad of this call site
> */
> - verifier_bug_if(idx + 1 != subseq_idx, env,
> - "extra insn from subprog");
> + verifier_bug_if(idx + 1 != subseq_idx &&
> + bpf_cleanup_pad_of_call(env, idx) != subseq_idx,
> + env, "extra insn from subprog");
[ ... ]
> @@ -888,11 +892,11 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
> }
>
> for (i = last_idx;;) {
> + hist = get_jmp_hist_entry(st, history, i);
> if (skip_first) {
> err = 0;
> skip_first = false;
> } else {
> - hist = get_jmp_hist_entry(st, history, i);
> err = backtrack_insn(env, i, subseq_idx, hist, bt);
> }
> if (err == -ENOTSUPP) {
> @@ -909,6 +913,14 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
> */
> return 0;
> subseq_idx = i;
> + /* This insn is a landing pad the unwind reached from
> + * a throw or a resume hist->unwind_frames frames
> + * deeper. No insn stands between the two, so enter
> + * those frames here, the way BPF_EXIT enters one.
> + */
This isn't a bug, but the new multi-line comment starts its text on the
same line as the opening /*. Under kernel/bpf/ the opening /* belongs on
its own line?
> + for (fr = 0; hist && fr < hist->unwind_frames; fr++)
> + if (bt_subprog_enter(bt))
> + return -EFAULT;
> i = get_prev_insn_idx(st, i, &history);
> if (i == -ENOENT)
> break;
Can bt->frame exceed st->curframe here?
The new loop raises bt->frame by hist->unwind_frames before
get_prev_insn_idx() discovers that the pad is the first instruction of the
state. If the pad is a checkpoint's first_insn_idx, the walk leaves the
state with bt->frame at the deep frame K+P, st = st->parent is the
checkpoint that only has frames 0..K, and the parent-state loop then runs
func = st->frame[fr] with fr up to K+P:
for (fr = bt->frame; fr >= 0; fr--) {
func = st->frame[fr];
...
for (i = 0; i < func->out_stack_arg_cnt; i++)
func is NULL for every fr > K, so is this a NULL pointer dereference in
the verifier during program load?
The pad can be a state's first insn because visit_cleanup_pad_edge() does
mark_prune_point(env, w) and mark_jmp_point(env, w) on the landing pad, so
do_check() calls bpf_is_state_visited() at the pad and creates a checkpoint
whenever force_new_state is true or the 2-jumps/8-insns heuristic fires.
bpf_is_state_visited() sets cur->first_insn_idx = insn_idx and
bpf_clear_jmp_history(cur), and do_check() immediately pushes the pad's
history entry with env->cur_hist_ent->unwind_frames = unwind_frames. The
pad entry is cur->jmp_history[0] with cur->first_insn_idx == pad, so
get_prev_insn_idx() returns -ENOENT.
The parent state has no such frame because the unwind pops the callee
frames in the *previous* do_check() iteration, so cur->curframe at the pad
is the shallow pad frame, the checkpoint copied from it has the same
curframe, and bpf_copy_verifier_state() frees and NULLs
dst_state->frame[i] for i > src->curframe.
BPF_EXIT does not have this problem: the history entry belongs to the
callee (the deeper frame), and the state at that point still contains that
frame. Could this be the first case where an entry recorded in frame K
asks the backtrack to enter frames that the state at that point no longer
has?
> diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
> index 66fb11b6c6a7..b101baa43171 100644
> --- a/kernel/bpf/states.c
> +++ b/kernel/bpf/states.c
> @@ -996,6 +996,12 @@ static bool states_equal(struct bpf_verifier_env *env,
> if (old->in_sleepable != cur->in_sleepable)
> return false;
>
> + if (old->unwinding != cur->unwinding)
> + return false;
> +
> + if (old->unwinding && old->unwind_frameno != cur->unwind_frameno)
> + return false;
> +
[ ... ]
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 680ae191aa3f..2bc08c18ebc8 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
> @@ -10,6 +10,7 @@
> #include <linux/slab.h>
> #include <linux/bpf.h>
> #include <linux/btf.h>
> +#include <linux/bpf_cleanup_abi.h>
> #include <linux/bpf_verifier.h>
> #include <linux/filter.h>
> #include <net/netlink.h>
> @@ -1721,6 +1722,8 @@ int bpf_copy_verifier_state(struct bpf_verifier_state *dst_state,
> return err;
> dst_state->speculative = src->speculative;
> dst_state->in_sleepable = src->in_sleepable;
> + dst_state->unwinding = src->unwinding;
> + dst_state->unwind_frameno = src->unwind_frameno;
[ ... ]
> @@ -10677,7 +10680,7 @@ static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
> verbose(env, "failed to push state for global subprog exception path\n");
> return PTR_ERR(branch);
> }
> - return process_bpf_exit_full(env, NULL, true);
> + return unwind_step(env, *insn_idx, insn_idx);
> }
[ ... ]
> @@ -18510,9 +18513,108 @@ enum {
> INSN_IDX_UPDATED = 2,
> };
[ ... ]
> @@ -18784,10 +18874,14 @@ static int do_check(struct bpf_verifier_env *env)
> for (;;) {
> struct bpf_insn *insn;
> struct bpf_insn_aux_data *insn_aux;
> + u8 unwind_frames;
> int err;
>
> /* reset current history entry on each new instruction */
> env->cur_hist_ent = NULL;
> + /* frames the unwind popped to reach this insn, if it is a pad */
> + unwind_frames = env->unwind_frames;
> + env->unwind_frames = 0;
>
> env->prev_insn_idx = prev_insn_idx;
> if (env->insn_idx >= insn_cnt) {
> @@ -18851,10 +18945,18 @@ static int do_check(struct bpf_verifier_env *env)
> }
> }
>
> - if (bpf_is_jmp_point(env, env->insn_idx)) {
> + /*
> + * The entry pushed here is the only record of how many frames
> + * the unwind popped to reach this insn, which the backtrack in
> + * mark_chain_precision() needs to follow the same edge. A pad
> + * is already a jump point; the second test only guards against
> + * it ever ceasing to be one.
> + */
> + if (bpf_is_jmp_point(env, env->insn_idx) || unwind_frames) {
> err = bpf_push_jmp_history(env, state, 0, 0, 0, 0);
> if (err)
> return err;
> + env->cur_hist_ent->unwind_frames = unwind_frames;
> }
Can the unwind_frames count get lost before this writes it?
The pad insn can prune, and the pad is a prune point from
visit_cleanup_pad_edge()->mark_prune_point(env, w), so
bpf_is_state_visited() can return 1 at the pad. When it does, the hit
path (kernel/bpf/states.c) itself pushes the history entry and immediately
consumes it:
if (bpf_is_jmp_point(env, env->insn_idx))
err = bpf_push_jmp_history(env, cur, 0, 0, 0, 0);
err = err ? : propagate_precision(env, &sl->state, cur, NULL);
and do_check() takes 'goto process_bpf_exit' before reaching this point, so
env->cur_hist_ent->unwind_frames never gets written with the
env->unwind_frames value captured at the top of the loop? The entry pushed
on the prune path then keeps unwind_frames == 0, and
propagate_precision()->bpf_mark_chain_precision() later sees
hist->unwind_frames == 0 at the pad, does not enter the popped frames, and
walks the throwing subprogram's instructions while bt->frame is still at
the pad's frame K.
If the masks propagated from the equivalent state cover the pad frame's
callee-saved registers or stack, backtrack_insn() walks the throwing
subprogram's instructions attributing every register write to frame K; it
either reaches the static call insn that entered the throwing frame with a
non-argument mask still set and returns -EFAULT through verifier_bug(),
rejecting a valid program with an internal error, or it silently applies
the marks to the wrong frame and the pruning that produced them is unsound?
Can the pad's jmp-history entry be pushed before bpf_is_state_visited() has
a chance to create a checkpoint at the same insn?
The checkpoint is created from the state that the unwind has already popped
frames out of. cfg.c:visit_cleanup_pad_edge() does mark_prune_point(env,
w) for the pad, so do_check() calls bpf_is_state_visited(env, pad_idx)
earlier in the loop. On the add_new_state path (kernel/bpf/states.c) the
checkpoint is a bpf_copy_verifier_state() of the *current* state, whose
curframe is already the pad's frame K: bpf_copy_verifier_state() only
populates frame[0..src->curframe] of a freshly kzalloc'd state, so frames
K+1..K+P in the checkpoint are NULL. kernel/bpf/states.c then does cur->parent
= new; cur->first_insn_idx = pad_idx; bpf_clear_jmp_history(cur).
Control returns here, which pushes the pad entry as cur's *first* history
entry and stamps it with unwind_frames = P.
If any later precision request's backtrack reaches that entry, the new loop
in bpf_mark_chain_precision() (kernel/bpf/backtrack.c) raises bt->frame to
K+P. The following get_prev_insn_idx() returns -ENOENT because i ==
st->first_insn_idx and the pad entry is the only entry, so the walk breaks
out to st = st->parent, the checkpoint that has no frames above K. The
parent-state loop then runs with fr up to bt->frame = K+P:
for (fr = bt->frame; fr >= 0; fr--) {
func = st->frame[fr];
...
for (i = 0; i < func->out_stack_arg_cnt; i++) {
func is NULL for every fr > K, so is this a NULL pointer dereference in the
verifier during program load (CAP_BPF)?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35656368472
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier
2026-09-21 21:56 ` bot+bpf-ci
@ 2026-09-22 5:21 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-22 5:21 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
ihor.solodrai
On 9/21/26 2:56 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/include/linux/bpf_cleanup_abi.h b/include/linux/bpf_cleanup_abi.h
>> new file mode 100644
>> index 000000000000..40b42c78fd2b
>> --- /dev/null
>> +++ b/include/linux/bpf_cleanup_abi.h
> [ ... ]
>
>> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
>> index aa631bb45f76..076739d4d974 100644
>> --- a/include/linux/bpf_verifier.h
>> +++ b/include/linux/bpf_verifier.h
>> @@ -429,7 +429,8 @@ struct bpf_jmp_history_entry {
>> u32 prev_idx : 20;
>> /* special INSN_F_xxx flags */
>> u32 flags : 4;
>> - u32 : 8;
>> + u32 unwind_frames : 4; /* frames the unwind popped to get here */
>> + u32 : 4;
> [ ... ]
>
>> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
>> index 507a366dffa4..bf5e7e6ab78f 100644
>> --- a/kernel/bpf/backtrack.c
>> +++ b/kernel/bpf/backtrack.c
>> @@ -4,6 +4,7 @@
>> #include <linux/bpf_verifier.h>
>> #include <linux/filter.h>
>> #include <linux/bitmap.h>
>> +#include "exception.h"
>>
>> #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
>>
>> @@ -47,6 +48,7 @@ int bpf_push_jmp_history(struct bpf_verifier_env *env, struct bpf_verifier_state
>> p->flags = insn_flags;
>> p->spi = spi;
>> p->frame = frame;
>> + p->unwind_frames = 0;
>> p->linked_regs = linked_regs;
>> cur->jmp_history_cnt = cnt;
>> env->cur_hist_ent = p;
> Does unwind_frames stay zero when bpf_is_state_visited() takes the pruned
> path?
>
> The landing pad is both a prune point and a jmp point from
> visit_cleanup_pad_edge(), so when bpf_is_state_visited() finds an equivalent
> state, its hit path in kernel/bpf/states.c pushes the history entry:
>
> if (bpf_is_jmp_point(env, env->insn_idx))
> err = bpf_push_jmp_history(env, cur, 0, 0, 0, 0);
> err = err ? : propagate_precision(env, &sl->state, cur, NULL);
>
> and do_check() then takes 'goto process_bpf_exit', skipping the block that
> would write env->cur_hist_ent->unwind_frames = unwind_frames. The comment
> in do_check() states this entry is the only record of how many frames the
> unwind popped to reach the insn, so when the pruned path keeps the zero,
> propagate_precision()->bpf_mark_chain_precision() later walks from the pad
> straight to the throwing insn in the deeper frame with bt->frame still at
> the pad frame?
>
> That would let the callee's writes to r6-r9 clear the caller frame's masks,
> or reach the call instruction still set and trip bt_subprog_exit() or the
> static-subprog checks in backtrack_insn().
We do have issues here. One important thing is the return value is "R0 = 1"
at the beginning of landing pad. Then backtracking knows 'R0 = 1' and is
able to have different backtracking compared to unknown R0. If R0 becomes
unknown, more backtracking path will happen and some issues are exposed.
Will fix with additional selftests. The same as below.
>
>> @@ -419,10 +421,12 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
>> * extra instructions from subprog; the next
>> * instruction after call to global subprog
>> * should be literally next instruction in
>> - * caller program
>> + * caller program -- or, if the callee threw,
>> + * the landing pad of this call site
>> */
>> - verifier_bug_if(idx + 1 != subseq_idx, env,
>> - "extra insn from subprog");
>> + verifier_bug_if(idx + 1 != subseq_idx &&
>> + bpf_cleanup_pad_of_call(env, idx) != subseq_idx,
>> + env, "extra insn from subprog");
> [ ... ]
>
>> @@ -888,11 +892,11 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
>> }
>>
>> for (i = last_idx;;) {
>> + hist = get_jmp_hist_entry(st, history, i);
>> if (skip_first) {
>> err = 0;
>> skip_first = false;
>> } else {
>> - hist = get_jmp_hist_entry(st, history, i);
>> err = backtrack_insn(env, i, subseq_idx, hist, bt);
>> }
>> if (err == -ENOTSUPP) {
>> @@ -909,6 +913,14 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
>> */
>> return 0;
>> subseq_idx = i;
>> + /* This insn is a landing pad the unwind reached from
>> + * a throw or a resume hist->unwind_frames frames
>> + * deeper. No insn stands between the two, so enter
>> + * those frames here, the way BPF_EXIT enters one.
>> + */
> This isn't a bug, but the new multi-line comment starts its text on the
> same line as the opening /*. Under kernel/bpf/ the opening /* belongs on
> its own line?
>
>> + for (fr = 0; hist && fr < hist->unwind_frames; fr++)
>> + if (bt_subprog_enter(bt))
>> + return -EFAULT;
>> i = get_prev_insn_idx(st, i, &history);
>> if (i == -ENOENT)
>> break;
> Can bt->frame exceed st->curframe here?
>
> The new loop raises bt->frame by hist->unwind_frames before
> get_prev_insn_idx() discovers that the pad is the first instruction of the
> state. If the pad is a checkpoint's first_insn_idx, the walk leaves the
> state with bt->frame at the deep frame K+P, st = st->parent is the
> checkpoint that only has frames 0..K, and the parent-state loop then runs
> func = st->frame[fr] with fr up to K+P:
>
> for (fr = bt->frame; fr >= 0; fr--) {
> func = st->frame[fr];
> ...
> for (i = 0; i < func->out_stack_arg_cnt; i++)
>
> func is NULL for every fr > K, so is this a NULL pointer dereference in
> the verifier during program load?
>
> The pad can be a state's first insn because visit_cleanup_pad_edge() does
> mark_prune_point(env, w) and mark_jmp_point(env, w) on the landing pad, so
> do_check() calls bpf_is_state_visited() at the pad and creates a checkpoint
> whenever force_new_state is true or the 2-jumps/8-insns heuristic fires.
> bpf_is_state_visited() sets cur->first_insn_idx = insn_idx and
> bpf_clear_jmp_history(cur), and do_check() immediately pushes the pad's
> history entry with env->cur_hist_ent->unwind_frames = unwind_frames. The
> pad entry is cur->jmp_history[0] with cur->first_insn_idx == pad, so
> get_prev_insn_idx() returns -ENOENT.
>
> The parent state has no such frame because the unwind pops the callee
> frames in the *previous* do_check() iteration, so cur->curframe at the pad
> is the shallow pad frame, the checkpoint copied from it has the same
> curframe, and bpf_copy_verifier_state() frees and NULLs
> dst_state->frame[i] for i > src->curframe.
>
> BPF_EXIT does not have this problem: the history entry belongs to the
> callee (the deeper frame), and the state at that point still contains that
> frame. Could this be the first case where an entry recorded in frame K
> asks the backtrack to enter frames that the state at that point no longer
> has?
>
>> diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c
>> index 66fb11b6c6a7..b101baa43171 100644
>> --- a/kernel/bpf/states.c
>> +++ b/kernel/bpf/states.c
>> @@ -996,6 +996,12 @@ static bool states_equal(struct bpf_verifier_env *env,
>> if (old->in_sleepable != cur->in_sleepable)
>> return false;
>>
>> + if (old->unwinding != cur->unwinding)
>> + return false;
>> +
>> + if (old->unwinding && old->unwind_frameno != cur->unwind_frameno)
>> + return false;
>> +
> [ ... ]
>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index 680ae191aa3f..2bc08c18ebc8 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
>> @@ -10,6 +10,7 @@
>> #include <linux/slab.h>
>> #include <linux/bpf.h>
>> #include <linux/btf.h>
>> +#include <linux/bpf_cleanup_abi.h>
>> #include <linux/bpf_verifier.h>
>> #include <linux/filter.h>
>> #include <net/netlink.h>
>> @@ -1721,6 +1722,8 @@ int bpf_copy_verifier_state(struct bpf_verifier_state *dst_state,
>> return err;
>> dst_state->speculative = src->speculative;
>> dst_state->in_sleepable = src->in_sleepable;
>> + dst_state->unwinding = src->unwinding;
>> + dst_state->unwind_frameno = src->unwind_frameno;
> [ ... ]
>
>> @@ -10677,7 +10680,7 @@ static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
>> verbose(env, "failed to push state for global subprog exception path\n");
>> return PTR_ERR(branch);
>> }
>> - return process_bpf_exit_full(env, NULL, true);
>> + return unwind_step(env, *insn_idx, insn_idx);
>> }
> [ ... ]
>
>> @@ -18510,9 +18513,108 @@ enum {
>> INSN_IDX_UPDATED = 2,
>> };
> [ ... ]
>
>> @@ -18784,10 +18874,14 @@ static int do_check(struct bpf_verifier_env *env)
>> for (;;) {
>> struct bpf_insn *insn;
>> struct bpf_insn_aux_data *insn_aux;
>> + u8 unwind_frames;
>> int err;
>>
>> /* reset current history entry on each new instruction */
>> env->cur_hist_ent = NULL;
>> + /* frames the unwind popped to reach this insn, if it is a pad */
>> + unwind_frames = env->unwind_frames;
>> + env->unwind_frames = 0;
>>
>> env->prev_insn_idx = prev_insn_idx;
>> if (env->insn_idx >= insn_cnt) {
>> @@ -18851,10 +18945,18 @@ static int do_check(struct bpf_verifier_env *env)
>> }
>> }
>>
>> - if (bpf_is_jmp_point(env, env->insn_idx)) {
>> + /*
>> + * The entry pushed here is the only record of how many frames
>> + * the unwind popped to reach this insn, which the backtrack in
>> + * mark_chain_precision() needs to follow the same edge. A pad
>> + * is already a jump point; the second test only guards against
>> + * it ever ceasing to be one.
>> + */
>> + if (bpf_is_jmp_point(env, env->insn_idx) || unwind_frames) {
>> err = bpf_push_jmp_history(env, state, 0, 0, 0, 0);
>> if (err)
>> return err;
>> + env->cur_hist_ent->unwind_frames = unwind_frames;
>> }
> Can the unwind_frames count get lost before this writes it?
>
> The pad insn can prune, and the pad is a prune point from
> visit_cleanup_pad_edge()->mark_prune_point(env, w), so
> bpf_is_state_visited() can return 1 at the pad. When it does, the hit
> path (kernel/bpf/states.c) itself pushes the history entry and immediately
> consumes it:
>
> if (bpf_is_jmp_point(env, env->insn_idx))
> err = bpf_push_jmp_history(env, cur, 0, 0, 0, 0);
> err = err ? : propagate_precision(env, &sl->state, cur, NULL);
>
> and do_check() takes 'goto process_bpf_exit' before reaching this point, so
> env->cur_hist_ent->unwind_frames never gets written with the
> env->unwind_frames value captured at the top of the loop? The entry pushed
> on the prune path then keeps unwind_frames == 0, and
> propagate_precision()->bpf_mark_chain_precision() later sees
> hist->unwind_frames == 0 at the pad, does not enter the popped frames, and
> walks the throwing subprogram's instructions while bt->frame is still at
> the pad's frame K.
>
> If the masks propagated from the equivalent state cover the pad frame's
> callee-saved registers or stack, backtrack_insn() walks the throwing
> subprogram's instructions attributing every register write to frame K; it
> either reaches the static call insn that entered the throwing frame with a
> non-argument mask still set and returns -EFAULT through verifier_bug(),
> rejecting a valid program with an internal error, or it silently applies
> the marks to the wrong frame and the pruning that produced them is unsound?
>
> Can the pad's jmp-history entry be pushed before bpf_is_state_visited() has
> a chance to create a checkpoint at the same insn?
>
> The checkpoint is created from the state that the unwind has already popped
> frames out of. cfg.c:visit_cleanup_pad_edge() does mark_prune_point(env,
> w) for the pad, so do_check() calls bpf_is_state_visited(env, pad_idx)
> earlier in the loop. On the add_new_state path (kernel/bpf/states.c) the
> checkpoint is a bpf_copy_verifier_state() of the *current* state, whose
> curframe is already the pad's frame K: bpf_copy_verifier_state() only
> populates frame[0..src->curframe] of a freshly kzalloc'd state, so frames
> K+1..K+P in the checkpoint are NULL. kernel/bpf/states.c then does cur->parent
> = new; cur->first_insn_idx = pad_idx; bpf_clear_jmp_history(cur).
> Control returns here, which pushes the pad entry as cur's *first* history
> entry and stamps it with unwind_frames = P.
>
> If any later precision request's backtrack reaches that entry, the new loop
> in bpf_mark_chain_precision() (kernel/bpf/backtrack.c) raises bt->frame to
> K+P. The following get_prev_insn_idx() returns -ENOENT because i ==
> st->first_insn_idx and the pad entry is the only entry, so the walk breaks
> out to st = st->parent, the checkpoint that has no frames above K. The
> parent-state loop then runs with fr up to bt->frame = K+P:
>
> for (fr = bt->frame; fr >= 0; fr--) {
> func = st->frame[fr];
> ...
> for (i = 0; i < func->out_stack_arg_cnt; i++) {
>
> func is NULL for every fr > K, so is this a NULL pointer dereference in the
> verifier during program load (CAP_BPF)?
>
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35656368472
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier
2026-09-21 21:01 ` [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier Yonghong Song
2026-09-21 21:40 ` sashiko-bot
2026-09-21 21:56 ` bot+bpf-ci
@ 2026-09-22 4:08 ` Alexei Starovoitov
2026-09-22 5:25 ` Yonghong Song
2026-09-22 23:43 ` Eduard Zingerman
3 siblings, 1 reply; 80+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 4:08 UTC (permalink / raw)
To: Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
On Mon Sep 21, 2026 at 9:01 PM UTC, Yonghong Song wrote:
> --- /dev/null
> +++ b/include/linux/bpf_cleanup_abi.h
> @@ -0,0 +1,16 @@
> +/* SPDX-License-Identifier: GPL-2.0-only */
> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
> +#ifndef _LINUX_BPF_CLEANUP_ABI_H
> +#define _LINUX_BPF_CLEANUP_ABI_H
> +
> +/*
> + * Value arch_bpf_run_cleanup_pad() leaves in r0 on the way into a landing pad.
> + * It has to be a constant the verifier knows: LLVM names r0 as both the
> + * exception pointer and the exception selector register, so every pad reads it
> + * before anything else and is free to store what it read. It gets a header of
> + * its own because the two sides that have to agree on it are the verifier and
> + * the arch dispatchers, which are assembly.
> + */
> +#define BPF_PAD_ENTRY_R0 1
> +
> +#endif /* _LINUX_BPF_CLEANUP_ABI_H */
I'm not going to read the AI reasons in commit log that it came up with
to justify new .h.
I bet it doesn't need new .h. If it does, please spell it out with human voice.
And, in general, pls tell AI to be terse. and remember it forever.
I told my clanker to be like me. Terse and to the point.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier
2026-09-22 4:08 ` Alexei Starovoitov
@ 2026-09-22 5:25 ` Yonghong Song
2026-09-22 21:53 ` Eduard Zingerman
0 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-22 5:25 UTC (permalink / raw)
To: Alexei Starovoitov, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
On 9/21/26 9:08 PM, Alexei Starovoitov wrote:
> On Mon Sep 21, 2026 at 9:01 PM UTC, Yonghong Song wrote:
>> --- /dev/null
>> +++ b/include/linux/bpf_cleanup_abi.h
>> @@ -0,0 +1,16 @@
>> +/* SPDX-License-Identifier: GPL-2.0-only */
>> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
>> +#ifndef _LINUX_BPF_CLEANUP_ABI_H
>> +#define _LINUX_BPF_CLEANUP_ABI_H
>> +
>> +/*
>> + * Value arch_bpf_run_cleanup_pad() leaves in r0 on the way into a landing pad.
>> + * It has to be a constant the verifier knows: LLVM names r0 as both the
>> + * exception pointer and the exception selector register, so every pad reads it
>> + * before anything else and is free to store what it read. It gets a header of
>> + * its own because the two sides that have to agree on it are the verifier and
>> + * the arch dispatchers, which are assembly.
>> + */
>> +#define BPF_PAD_ENTRY_R0 1
>> +
>> +#endif /* _LINUX_BPF_CLEANUP_ABI_H */
> I'm not going to read the AI reasons in commit log that it came up with
> to justify new .h.
> I bet it doesn't need new .h. If it does, please spell it out with human voice.
> And, in general, pls tell AI to be terse. and remember it forever.
>
> I told my clanker to be like me. Terse and to the point.
Okay, I added this file to be shared in x86/net/bpf_cleanup_pad.S,
arm64/net/bpf_cleanup_pad.S and verifier.c. Yes, we can remove it
with single line comment in their respective files.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier
2026-09-22 5:25 ` Yonghong Song
@ 2026-09-22 21:53 ` Eduard Zingerman
2026-09-23 3:18 ` Yonghong Song
0 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-22 21:53 UTC (permalink / raw)
To: Yonghong Song, Alexei Starovoitov, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Mon, 2026-09-21 at 22:25 -0700, Yonghong Song wrote:
>
> On 9/21/26 9:08 PM, Alexei Starovoitov wrote:
> > On Mon Sep 21, 2026 at 9:01 PM UTC, Yonghong Song wrote:
> > > --- /dev/null
> > > +++ b/include/linux/bpf_cleanup_abi.h
> > > @@ -0,0 +1,16 @@
> > > +/* SPDX-License-Identifier: GPL-2.0-only */
> > > +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
> > > +#ifndef _LINUX_BPF_CLEANUP_ABI_H
> > > +#define _LINUX_BPF_CLEANUP_ABI_H
> > > +
> > > +/*
> > > + * Value arch_bpf_run_cleanup_pad() leaves in r0 on the way into a landing pad.
> > > + * It has to be a constant the verifier knows: LLVM names r0 as both the
> > > + * exception pointer and the exception selector register, so every pad reads it
> > > + * before anything else and is free to store what it read. It gets a header of
> > > + * its own because the two sides that have to agree on it are the verifier and
> > > + * the arch dispatchers, which are assembly.
> > > + */
> > > +#define BPF_PAD_ENTRY_R0 1
> > > +
> > > +#endif /* _LINUX_BPF_CLEANUP_ABI_H */
> > I'm not going to read the AI reasons in commit log that it came up with
> > to justify new .h.
> > I bet it doesn't need new .h. If it does, please spell it out with human voice.
> > And, in general, pls tell AI to be terse. and remember it forever.
> >
> > I told my clanker to be like me. Terse and to the point.
>
> Okay, I added this file to be shared in x86/net/bpf_cleanup_pad.S,
> arm64/net/bpf_cleanup_pad.S and verifier.c. Yes, we can remove it
> with single line comment in their respective files.
>
As far as I understand BPF_PAD_ENTRY_R0 is not needed at all.
As the value is unused the jits can zero out R0 upon landing
pad entry from throw or resume, verifier can initialize R0 as
an unknown scalar: mark_reg_unknown(env, frame->regs, BPF_REG_0).
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier
2026-09-22 21:53 ` Eduard Zingerman
@ 2026-09-23 3:18 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-23 3:18 UTC (permalink / raw)
To: Eduard Zingerman, Alexei Starovoitov, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On 9/22/26 2:53 PM, Eduard Zingerman wrote:
> On Mon, 2026-09-21 at 22:25 -0700, Yonghong Song wrote:
>> On 9/21/26 9:08 PM, Alexei Starovoitov wrote:
>>> On Mon Sep 21, 2026 at 9:01 PM UTC, Yonghong Song wrote:
>>>> --- /dev/null
>>>> +++ b/include/linux/bpf_cleanup_abi.h
>>>> @@ -0,0 +1,16 @@
>>>> +/* SPDX-License-Identifier: GPL-2.0-only */
>>>> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
>>>> +#ifndef _LINUX_BPF_CLEANUP_ABI_H
>>>> +#define _LINUX_BPF_CLEANUP_ABI_H
>>>> +
>>>> +/*
>>>> + * Value arch_bpf_run_cleanup_pad() leaves in r0 on the way into a landing pad.
>>>> + * It has to be a constant the verifier knows: LLVM names r0 as both the
>>>> + * exception pointer and the exception selector register, so every pad reads it
>>>> + * before anything else and is free to store what it read. It gets a header of
>>>> + * its own because the two sides that have to agree on it are the verifier and
>>>> + * the arch dispatchers, which are assembly.
>>>> + */
>>>> +#define BPF_PAD_ENTRY_R0 1
>>>> +
>>>> +#endif /* _LINUX_BPF_CLEANUP_ABI_H */
>>> I'm not going to read the AI reasons in commit log that it came up with
>>> to justify new .h.
>>> I bet it doesn't need new .h. If it does, please spell it out with human voice.
>>> And, in general, pls tell AI to be terse. and remember it forever.
>>>
>>> I told my clanker to be like me. Terse and to the point.
>> Okay, I added this file to be shared in x86/net/bpf_cleanup_pad.S,
>> arm64/net/bpf_cleanup_pad.S and verifier.c. Yes, we can remove it
>> with single line comment in their respective files.
>>
> As far as I understand BPF_PAD_ENTRY_R0 is not needed at all.
> As the value is unused the jits can zero out R0 upon landing
> pad entry from throw or resume, verifier can initialize R0 as
> an unknown scalar: mark_reg_unknown(env, frame->regs, BPF_REG_0).
Agree. R0 is not really needed so make it mark_reg_unknown is
the right approach.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier
2026-09-21 21:01 ` [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier Yonghong Song
` (2 preceding siblings ...)
2026-09-22 4:08 ` Alexei Starovoitov
@ 2026-09-22 23:43 ` Eduard Zingerman
2026-09-23 3:21 ` Yonghong Song
3 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-22 23:43 UTC (permalink / raw)
To: Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Mon, 2026-09-21 at 14:01 -0700, Yonghong Song wrote:
...
> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index aa631bb45f76..076739d4d974 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
...
> @@ -991,6 +994,7 @@ struct bpf_verifier_env {
> } cfg;
> struct backtrack_state bt;
> struct bpf_jmp_history_entry *cur_hist_ent;
> + u8 unwind_frames; /* scratch: frames the unwind popped to reach the next insn */
Instead of maintaining this variable across calls to do_check() and
unwind_step(), I think it should be possible to do bpf_push_jmp_history()
in the uwind_step() itself.
> /* Per-callsite copy of parent's converged at_stack_in for cross-frame fills. */
> struct arg_track **callsite_at_stack;
> u32 pass_cnt; /* number of times do_check() was called */
> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
...
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 680ae191aa3f..2bc08c18ebc8 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
...
> @@ -18510,9 +18513,108 @@ enum {
> INSN_IDX_UPDATED = 2,
> };
>
> -static int process_bpf_exit_full(struct bpf_verifier_env *env,
> - bool *do_print_state,
> - bool exception_exit)
> +static u32 unwind_pop_frame(struct bpf_verifier_env *env)
> +{
This function duplicates the code in prepare_func_exit(),
I'd suggest renaming it to `pop_frame` and calling it from
prepare_func_exit() as well.
> + struct bpf_verifier_state *state = env->cur_state;
> + struct bpf_func_state *callee = state->frame[state->curframe];
> + u32 callsite = callee->callsite;
> + struct bpf_func_state *caller;
> +
> + caller = state->frame[state->curframe - 1];
> + account_processed_insns(env, callee, caller);
> + free_func_state(callee);
> + state->frame[state->curframe--] = NULL;
> + invalidate_outgoing_stack_args(env, caller);
> + return callsite;
> +}
...
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier
2026-09-22 23:43 ` Eduard Zingerman
@ 2026-09-23 3:21 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-23 3:21 UTC (permalink / raw)
To: Eduard Zingerman, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On 9/22/26 4:43 PM, Eduard Zingerman wrote:
> On Mon, 2026-09-21 at 14:01 -0700, Yonghong Song wrote:
>
> ...
>
>> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
>> index aa631bb45f76..076739d4d974 100644
>> --- a/include/linux/bpf_verifier.h
>> +++ b/include/linux/bpf_verifier.h
> ...
>
>> @@ -991,6 +994,7 @@ struct bpf_verifier_env {
>> } cfg;
>> struct backtrack_state bt;
>> struct bpf_jmp_history_entry *cur_hist_ent;
>> + u8 unwind_frames; /* scratch: frames the unwind popped to reach the next insn */
> Instead of maintaining this variable across calls to do_check() and
> unwind_step(), I think it should be possible to do bpf_push_jmp_history()
> in the uwind_step() itself.
Yes, bpf_push_jmp_history() is doable. But I think my next patch with
unwind_frames and cur_unwind_frames will be simpler. We can discuss this
in detail after posting next revision.
>
>> /* Per-callsite copy of parent's converged at_stack_in for cross-frame fills. */
>> struct arg_track **callsite_at_stack;
>> u32 pass_cnt; /* number of times do_check() was called */
>> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
> ...
>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index 680ae191aa3f..2bc08c18ebc8 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
> ...
>
>> @@ -18510,9 +18513,108 @@ enum {
>> INSN_IDX_UPDATED = 2,
>> };
>>
>> -static int process_bpf_exit_full(struct bpf_verifier_env *env,
>> - bool *do_print_state,
>> - bool exception_exit)
>> +static u32 unwind_pop_frame(struct bpf_verifier_env *env)
>> +{
> This function duplicates the code in prepare_func_exit(),
> I'd suggest renaming it to `pop_frame` and calling it from
> prepare_func_exit() as well.
Good point. We want to avoid duplication.
>
>> + struct bpf_verifier_state *state = env->cur_state;
>> + struct bpf_func_state *callee = state->frame[state->curframe];
>> + u32 callsite = callee->callsite;
>> + struct bpf_func_state *caller;
>> +
>> + caller = state->frame[state->curframe - 1];
>> + account_processed_insns(env, callee, caller);
>> + free_func_state(callee);
>> + state->frame[state->curframe--] = NULL;
>> + invalidate_outgoing_stack_args(env, caller);
>> + return callsite;
>> +}
> ...
^ permalink raw reply [flat|nested] 80+ messages in thread
* [PATCH bpf-next v4 09/20] bpf: Refuse a private stack for a program with an exception cleanup table
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (7 preceding siblings ...)
2026-09-21 21:01 ` [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier Yonghong Song
@ 2026-09-21 21:01 ` Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 10/20] bpf: Dispatch exception cleanup pads from bpf_throw() Yonghong Song
` (11 subsequent siblings)
20 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:01 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A private stack is hard to support together with an exception cleanup
table on x86-64: a landing pad runs on the bpf_throw() walker's stack with
its frame's registers put back from a spill area, and carrying the private
stack's frame pointer through there as well would cost a register that
architecture cannot spare.
Force NO_PRIV_STACK in check_max_stack_depth(), where the choice is made,
and do it on every architecture for now rather than just x86-64. The
program still loads, but more than an optimization can be lost:
check_max_stack_depth_subprog() checks a PRIV_STACK_ADAPTIVE subprogram
against MAX_BPF_STACK on its own, and adds a NO_PRIV_STACK one to the
depth its callers have to fit in -- so a bpf2bpf chain deep enough to need
the private stack is now refused with "combined stack size of %d calls is
%d. Too large".
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
kernel/bpf/verifier.c | 11 +++++++++++
1 file changed, 11 insertions(+)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 2bc08c18ebc8..af234a39511c 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -5598,6 +5598,17 @@ static int check_max_stack_depth(struct bpf_verifier_env *env)
}
}
+ /*
+ * A pad rebuilds its frame from a spill area, and on x86-64 a private
+ * stack's frame pointer lives in r9, which no spill area holds.
+ * Refused on every arch rather than just that one. The subprograms
+ * below are then checked against MAX_BPF_STACK together rather than
+ * one at a time, so this can turn a program that would have loaded
+ * with a private stack into one that is too deep.
+ */
+ if (env->cleanup_info_cnt)
+ priv_stack_mode = NO_PRIV_STACK;
+
if (priv_stack_mode == PRIV_STACK_UNKNOWN)
priv_stack_mode = bpf_enable_priv_stack(env->prog);
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* [PATCH bpf-next v4 10/20] bpf: Dispatch exception cleanup pads from bpf_throw()
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (8 preceding siblings ...)
2026-09-21 21:01 ` [PATCH bpf-next v4 09/20] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
@ 2026-09-21 21:01 ` Yonghong Song
2026-09-22 21:38 ` Eduard Zingerman
2026-09-21 21:01 ` [PATCH bpf-next v4 11/20] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
` (10 subsequent siblings)
20 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:01 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Stop discarding the frames an exception unwinds through. bpf_throw()
already walks the BPF call stack with arch_bpf_stack_walk() to find the
exception boundary; have it look each frame's return address up in that
(sub)program's cleanup table on the way and run the landing pad a matching
record names.
A pad is run as a subroutine of the walker, not jumped to. It executes with
the unwinding frame's frame pointer and BPF callee-saved registers, so
everything it reads is that frame's, but on the current stack far below it,
so nothing it calls can disturb the frame it is cleaning up after. It ends
in what the JIT emits for its bpf_unwind_resume(), which hands control back
to the walker rather than to the frame's caller. That is exactly the
procedure the LLVM commit emitting the section describes.
Restoring the frame's r6-r9 is what makes this work, and it is only
possible because the callee about to be discarded spilled them in its own
prologue. bpf_cleanup_force_spill() tells a JIT to make that spill
unconditional and of a known shape for every subprogram of a program
carrying a cleanup table, and aux->exc->spill_off records where it starts,
so the walker needs no per-frame metadata. The frame that called
bpf_throw() has no callee to have spilled anything and never runs its own
epilogue, so the JIT spills that frame's registers at the throw site
instead, in the area aux->exc->throw_spill_off names.
The main program needs the table handed to it rather than built for it.
jit_subprogs() compiles it as func[0], but the ksym covering that image is
the one bpf_prog_load() registers for the outer bpf_prog, and that is what
the walker finds -- so without the handover a landing pad in the main
program's own frame is never dispatched, silently, since the exception
still reaches the boundary and the cookie still comes back.
Which instructions a JIT has to lower its own way reach it as indices into
the (sub)program, aux->exc->throw_at and ->resume_at, carried there from
the insn_aux_data marks so that patching keeps them pointing at the right
instruction.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
include/linux/bpf.h | 90 ++++++++++++++++++++++
include/linux/filter.h | 1 +
kernel/bpf/core.c | 30 +++++++-
kernel/bpf/exception.c | 168 +++++++++++++++++++++++++++++++++++++++++
kernel/bpf/exception.h | 7 ++
kernel/bpf/fixups.c | 133 ++++++++++++++++++++++++++++++++
kernel/bpf/helpers.c | 34 +++++++++
7 files changed, 460 insertions(+), 3 deletions(-)
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index fd22db8bc6c5..42c7758c7ab5 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -1777,6 +1777,95 @@ enum bpf_sig_keyring {
BPF_SIG_KEYRING_BPF,
};
+/*
+ * One cleanup region of a JITed (sub)program: @pad is the landing pad to run
+ * for a return address in (begin, end], the native code of its call sites.
+ */
+struct bpf_cleanup_range {
+ u64 begin;
+ u64 end;
+ u64 pad;
+};
+
+struct bpf_exception_info {
+ struct bpf_cleanup_info *info;
+ struct bpf_cleanup_range *ranges;
+ /* Landing pad instruction indices, sorted and deduplicated. */
+ u32 *pad_at;
+ /* bpf_throw() call instruction indices, sorted. */
+ u32 *throw_at;
+ /* Likewise for bpf_unwind_resume(), the way back out of a pad. */
+ u32 *resume_at;
+ /* One bit per instruction, set for those that only run while unwinding. */
+ unsigned long *pad_body;
+ u32 nr_info;
+ u32 nr_ranges;
+ u32 nr_pad_at;
+ u32 nr_throw_at;
+ u32 nr_resume_at;
+ u32 pad_body_bits;
+ /* Offset from a frame's FP to the caller's spilled r6-r9. */
+ s32 spill_off;
+ /* Likewise, to the registers a frame spills before calling bpf_throw(). */
+ s32 throw_spill_off;
+};
+
+#ifdef CONFIG_BPF_SYSCALL
+bool bpf_cleanup_force_spill(const struct bpf_prog *prog);
+bool bpf_cleanup_needs_throw_spill(const struct bpf_prog *prog);
+bool bpf_cleanup_insn_is_pad(const struct bpf_prog *prog, u32 idx);
+bool bpf_cleanup_insn_in_pad(const struct bpf_prog *prog, u32 idx);
+bool bpf_cleanup_insn_is_throw(const struct bpf_prog *prog, u32 idx);
+bool bpf_cleanup_insn_is_resume(const struct bpf_prog *prog, u32 idx);
+int bpf_cleanup_attach_main_prog(struct bpf_verifier_env *env, struct bpf_prog *prog);
+void bpf_cleanup_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image);
+void bpf_cleanup_free_info(struct bpf_prog_aux *aux);
+#else
+static inline bool bpf_cleanup_force_spill(const struct bpf_prog *prog)
+{
+ return false;
+}
+
+static inline bool bpf_cleanup_needs_throw_spill(const struct bpf_prog *prog)
+{
+ return false;
+}
+
+static inline bool bpf_cleanup_insn_is_pad(const struct bpf_prog *prog, u32 idx)
+{
+ return false;
+}
+
+static inline bool bpf_cleanup_insn_in_pad(const struct bpf_prog *prog, u32 idx)
+{
+ return false;
+}
+
+static inline bool bpf_cleanup_insn_is_throw(const struct bpf_prog *prog, u32 idx)
+{
+ return false;
+}
+
+static inline bool bpf_cleanup_insn_is_resume(const struct bpf_prog *prog, u32 idx)
+{
+ return false;
+}
+
+static inline int bpf_cleanup_attach_main_prog(struct bpf_verifier_env *env,
+ struct bpf_prog *prog)
+{
+ return 0;
+}
+
+static inline void bpf_cleanup_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+}
+
+static inline void bpf_cleanup_free_info(struct bpf_prog_aux *aux)
+{
+}
+#endif
+
struct bpf_prog_aux {
atomic64_t refcnt;
u32 used_map_cnt;
@@ -1857,6 +1946,7 @@ struct bpf_prog_aux {
char name[BPF_OBJ_NAME_LEN];
u64 (*bpf_exception_cb)(u64 cookie, u64 sp, u64 bp, u64, u64);
u16 stack_arg_sp_adjust;
+ struct bpf_exception_info *exc;
#ifdef CONFIG_SECURITY
void *security;
#endif
diff --git a/include/linux/filter.h b/include/linux/filter.h
index 2582a7606e46..b0c495e94f74 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1243,6 +1243,7 @@ bool bpf_jit_supports_arena_args(void);
bool bpf_jit_supports_far_kfunc_call(void);
bool bpf_jit_supports_exceptions(void);
bool bpf_jit_supports_cleanup_pads(void);
+void arch_bpf_run_cleanup_pad(u64 pad, u64 frame_fp, u64 spill_base);
bool bpf_jit_supports_ptr_xchg(void);
bool bpf_jit_supports_arena(void);
bool bpf_jit_supports_insn(struct bpf_insn *insn, bool in_arena);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index a379cd1ec4c6..f19712be12a7 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -292,6 +292,7 @@ void __bpf_prog_free(struct bpf_prog *fp)
mutex_destroy(&fp->aux->dst_mutex);
mutex_destroy(&fp->aux->st_ops_assoc_mutex);
kfree(fp->aux->poke_tab);
+ bpf_cleanup_free_info(fp->aux);
kfree(fp->aux);
}
free_percpu(fp->stats);
@@ -2625,13 +2626,21 @@ static bool bpf_prog_select_interpreter(struct bpf_prog *fp)
return select_interpreter;
}
-static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struct bpf_prog *prog)
+static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struct bpf_prog *prog,
+ int *err)
{
#ifdef CONFIG_BPF_JIT
struct bpf_prog *orig_prog;
+ int ret;
- if (!bpf_prog_need_blind(prog))
+ if (!bpf_prog_need_blind(prog)) {
+ ret = bpf_cleanup_attach_main_prog(env, prog);
+ if (ret) {
+ *err = ret;
+ return prog;
+ }
return bpf_int_jit_compile(env, prog);
+ }
orig_prog = prog;
prog = bpf_jit_blind_constants(env, prog);
@@ -2642,6 +2651,13 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
if (IS_ERR(prog))
goto out_restore;
+ ret = bpf_cleanup_attach_main_prog(env, prog);
+ if (ret) {
+ *err = ret;
+ bpf_jit_prog_release_other(orig_prog, prog);
+ goto out_restore;
+ }
+
prog = bpf_int_jit_compile(env, prog);
if (prog->jited) {
bpf_jit_prog_release_other(prog, orig_prog);
@@ -2681,8 +2697,10 @@ struct bpf_prog *__bpf_prog_select_runtime(struct bpf_verifier_env *env, struct
if (*err)
return fp;
- fp = bpf_prog_jit_compile(env, fp);
+ fp = bpf_prog_jit_compile(env, fp, err);
bpf_prog_jit_attempt_done(fp);
+ if (*err)
+ return fp;
if (!fp->jited && jit_needed) {
*err = -ENOTSUPP;
return fp;
@@ -3480,6 +3498,12 @@ bool __weak bpf_jit_supports_cleanup_pads(void)
return false;
}
+/* Call @pad with the frame pointer @frame_fp and r6-r9 spilled at @spill_base. */
+void __weak arch_bpf_run_cleanup_pad(u64 pad, u64 frame_fp, u64 spill_base)
+{
+ WARN_ON_ONCE(1);
+}
+
bool __weak bpf_jit_supports_timed_may_goto(void)
{
return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index da8fa6eb7e4b..0a9b2b943de3 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -1,7 +1,9 @@
// SPDX-License-Identifier: GPL-2.0-only
/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bitmap.h>
#include <linux/bpf.h>
#include <linux/bpf_verifier.h>
+#include <linux/bsearch.h>
#include <linux/btf.h>
#include <linux/btf_ids.h>
#include <linux/filter.h>
@@ -493,3 +495,169 @@ int bpf_cleanup_pad_of_call(struct bpf_verifier_env *env, u32 idx)
return pad ? (int)pad - 1 : -1;
}
+
+/*
+ * Every subprogram of a cleanup-carrying program spills the BPF callee-saved
+ * registers, even one that never throws: a frame's spill holds its caller's
+ * registers, and that is what the walker restores before running the caller's
+ * pad. The exception callback does not, because it reuses the boundary frame
+ * rather than building one of its own.
+ */
+bool bpf_cleanup_force_spill(const struct bpf_prog *prog)
+{
+ return prog->aux->exc && !prog->aux->exception_cb;
+}
+
+/*
+ * The throw-site spill area, on the other hand, is only ever read for the
+ * frame the walk starts in, so only a (sub)program that calls bpf_throw()
+ * needs one.
+ */
+bool bpf_cleanup_needs_throw_spill(const struct bpf_prog *prog)
+{
+ return bpf_cleanup_force_spill(prog) && prog->aux->exc->nr_throw_at;
+}
+
+const struct bpf_cleanup_range *bpf_cleanup_pad_for_ip(const struct bpf_prog *prog, u64 ip)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+ u32 l = 0, r = exc ? exc->nr_ranges : 0;
+
+ while (l < r) {
+ u32 m = l + (r - l) / 2;
+ const struct bpf_cleanup_range *rec = &exc->ranges[m];
+
+ if (ip <= rec->begin)
+ r = m;
+ else if (ip > rec->end)
+ l = m + 1;
+ else
+ return rec;
+ }
+ return NULL;
+}
+
+static int cmp_u32(const void *a, const void *b)
+{
+ u32 x = *(const u32 *)a, y = *(const u32 *)b;
+
+ return x < y ? -1 : x > y;
+}
+
+int bpf_cleanup_alloc_info(struct bpf_prog_aux *aux)
+{
+ if (aux->exc)
+ return 0;
+ aux->exc = kzalloc_obj(struct bpf_exception_info, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ return aux->exc ? 0 : -ENOMEM;
+}
+
+int bpf_cleanup_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt)
+{
+ struct bpf_exception_info *exc = aux->exc;
+ struct bpf_cleanup_range *ranges;
+ u32 i, n_at, *at;
+
+ ranges = kvcalloc(cnt, sizeof(*ranges), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!ranges) {
+ kvfree(recs);
+ return -ENOMEM;
+ }
+
+ /* The pads on their own, sorted and deduplicated. */
+ at = kvmalloc_array(cnt, sizeof(*at), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!at) {
+ kvfree(ranges);
+ kvfree(recs);
+ return -ENOMEM;
+ }
+ for (i = 0; i < cnt; i++)
+ at[i] = recs[i].landing_pad_off;
+ sort(at, cnt, sizeof(*at), cmp_u32, NULL);
+ for (i = 0, n_at = 0; i < cnt; i++)
+ if (!n_at || at[n_at - 1] != at[i])
+ at[n_at++] = at[i];
+
+ exc->pad_at = at;
+ exc->nr_pad_at = n_at;
+ exc->info = recs;
+ exc->nr_info = cnt;
+ exc->ranges = ranges;
+ /* Withheld until the JIT has filled the table in. */
+ exc->nr_ranges = 0;
+ return 0;
+}
+
+void bpf_cleanup_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+ struct bpf_exception_info *exc = prog->aux->exc;
+ u32 i, n;
+
+ if (!exc || !exc->nr_info || !exc->ranges)
+ return;
+
+ n = exc->nr_info;
+ for (i = 0; i < n; i++) {
+ const struct bpf_cleanup_info *rec = &exc->info[i];
+
+ if (WARN_ON_ONCE(rec->begin_off >= prog->len ||
+ rec->end_off > prog->len ||
+ rec->landing_pad_off >= prog->len))
+ return;
+ exc->ranges[i].begin = (u64)(long)image + addrs[rec->begin_off];
+ exc->ranges[i].end = (u64)(long)image + addrs[rec->end_off];
+ exc->ranges[i].pad = (u64)(long)image + addrs[rec->landing_pad_off];
+ }
+ exc->nr_ranges = n;
+}
+
+void bpf_cleanup_free_info(struct bpf_prog_aux *aux)
+{
+ struct bpf_exception_info *exc = aux->exc;
+
+ if (!exc)
+ return;
+ kvfree(exc->ranges);
+ kvfree(exc->info);
+ kvfree(exc->pad_at);
+ kvfree(exc->throw_at);
+ kvfree(exc->resume_at);
+ bitmap_free(exc->pad_body);
+ kfree(exc);
+ aux->exc = NULL;
+}
+
+/* Is @idx in the sorted array @at of @n instruction indices? */
+static bool insn_idx_in(const u32 *at, u32 n, u32 idx)
+{
+ return bsearch(&idx, at, n, sizeof(*at), cmp_u32);
+}
+
+bool bpf_cleanup_insn_is_pad(const struct bpf_prog *prog, u32 idx)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+
+ return exc && insn_idx_in(exc->pad_at, exc->nr_pad_at, idx);
+}
+
+bool bpf_cleanup_insn_is_throw(const struct bpf_prog *prog, u32 idx)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+
+ return exc && insn_idx_in(exc->throw_at, exc->nr_throw_at, idx);
+}
+
+bool bpf_cleanup_insn_is_resume(const struct bpf_prog *prog, u32 idx)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+
+ return exc && insn_idx_in(exc->resume_at, exc->nr_resume_at, idx);
+}
+
+bool bpf_cleanup_insn_in_pad(const struct bpf_prog *prog, u32 idx)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+
+ return exc && exc->pad_body && idx < exc->pad_body_bits &&
+ test_bit(idx, exc->pad_body);
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index 7313dd2b65a1..c0e68ce227c8 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -5,11 +5,18 @@
#include <linux/types.h>
+struct bpf_cleanup_info;
+struct bpf_cleanup_range;
+struct bpf_prog;
+struct bpf_prog_aux;
struct bpf_verifier_env;
int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
int bpf_check_cleanup_exceptions(struct bpf_verifier_env *env);
int bpf_cleanup_check_callback(struct bpf_verifier_env *env, int subprog);
int bpf_cleanup_pad_of_call(struct bpf_verifier_env *env, u32 idx);
+int bpf_cleanup_alloc_info(struct bpf_prog_aux *aux);
+int bpf_cleanup_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt);
+const struct bpf_cleanup_range *bpf_cleanup_pad_for_ip(const struct bpf_prog *prog, u64 ip);
#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 6b5c1a1d0479..abef3355bd0d 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -1,5 +1,6 @@
// SPDX-License-Identifier: GPL-2.0-only
/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bitmap.h>
#include <linux/bpf.h>
#include <linux/btf.h>
#include <linux/bpf_verifier.h>
@@ -10,6 +11,7 @@
#include <linux/perf_event.h>
#include <net/xdp.h>
#include "disasm.h"
+#include "exception.h"
#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -1122,6 +1124,129 @@ static void bpf_restore_subprog_starts(struct bpf_verifier_env *env, u32 *orig_s
env->subprog_info[env->subprog_cnt].start = env->prog->len;
}
+static bool cleanup_kfunc_site(const struct bpf_insn_aux_data *aux, bool resume)
+{
+ return resume ? aux->cleanup_resume_site : aux->cleanup_throw_site;
+}
+
+static int cleanup_kfunc_sites_for_subprog(struct bpf_verifier_env *env, u32 start, u32 end,
+ bool resume, u32 **at_p, u32 *nr_p)
+{
+ u32 i, cnt = 0, *at;
+
+ for (i = start; i < end; i++)
+ if (cleanup_kfunc_site(&env->insn_aux_data[i], resume))
+ cnt++;
+ if (!cnt)
+ return 0;
+
+ at = kvmalloc_array(cnt, sizeof(*at), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!at)
+ return -ENOMEM;
+
+ for (i = start, cnt = 0; i < end; i++) {
+ if (!cleanup_kfunc_site(&env->insn_aux_data[i], resume))
+ continue;
+ at[cnt++] = i - start;
+ }
+
+ *at_p = at;
+ *nr_p = cnt;
+ return 0;
+}
+
+static int cleanup_pad_body_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
+ u32 start, u32 end)
+{
+ unsigned long *bits;
+ u32 i, cnt = 0;
+
+ for (i = start; i < end; i++)
+ if (env->insn_aux_data[i].in_cleanup_pad)
+ cnt++;
+ if (!cnt)
+ return 0;
+
+ bits = bitmap_zalloc(end - start, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!bits)
+ return -ENOMEM;
+
+ for (i = start; i < end; i++)
+ if (env->insn_aux_data[i].in_cleanup_pad)
+ __set_bit(i - start, bits);
+
+ sub->aux->exc->pad_body = bits;
+ sub->aux->exc->pad_body_bits = end - start;
+ return 0;
+}
+
+static int cleanup_info_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
+ u32 start, u32 end)
+{
+ struct bpf_cleanup_info *recs;
+ u32 i, cnt = 0;
+ int err;
+
+ if (!env->cleanup_info_cnt)
+ return 0;
+
+ err = bpf_cleanup_alloc_info(sub->aux);
+ if (err)
+ return err;
+
+ err = cleanup_kfunc_sites_for_subprog(env, start, end, false,
+ &sub->aux->exc->throw_at,
+ &sub->aux->exc->nr_throw_at);
+ if (err)
+ return err;
+
+ err = cleanup_kfunc_sites_for_subprog(env, start, end, true,
+ &sub->aux->exc->resume_at,
+ &sub->aux->exc->nr_resume_at);
+ if (err)
+ return err;
+
+ err = cleanup_pad_body_for_subprog(env, sub, start, end);
+ if (err)
+ return err;
+
+ for (i = start; i < end; i++)
+ if (env->insn_aux_data[i].cleanup_pad)
+ cnt++;
+ if (!cnt)
+ return 0;
+
+ recs = kvmalloc_array(cnt, sizeof(*recs), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+ if (!recs)
+ return -ENOMEM;
+
+ for (i = start, cnt = 0; i < end; i++) {
+ u32 pad = env->insn_aux_data[i].cleanup_pad;
+
+ if (!pad)
+ continue;
+ pad--;
+ if (verifier_bug_if(pad < start || pad >= end, env,
+ "insn %u is covered by a landing pad at %u outside its subprog [%u, %u)",
+ i, pad, start, end)) {
+ kvfree(recs);
+ return -EFAULT;
+ }
+ recs[cnt].begin_off = i - start;
+ recs[cnt].end_off = i - start + 1;
+ recs[cnt].landing_pad_off = pad - start;
+ cnt++;
+ }
+ return bpf_cleanup_attach_info(sub->aux, recs, cnt);
+}
+
+int bpf_cleanup_attach_main_prog(struct bpf_verifier_env *env, struct bpf_prog *prog)
+{
+ if (!env || env->subprog_cnt > 1)
+ return 0;
+ return cleanup_info_for_subprog(env, prog, 0, prog->len);
+}
+
static int jit_subprogs(struct bpf_verifier_env *env)
{
struct bpf_prog *prog = env->prog, **func, *tmp;
@@ -1259,6 +1384,10 @@ static int jit_subprogs(struct bpf_verifier_env *env)
func[i]->aux->token = prog->aux->token;
if (!i)
func[i]->aux->exception_boundary = env->seen_exception;
+ err = cleanup_info_for_subprog(env, func[i], subprog_start,
+ subprog_end);
+ if (err)
+ goto out_free;
func[i] = bpf_int_jit_compile(env, func[i]);
if (!func[i]->jited) {
err = -ENOTSUPP;
@@ -1363,6 +1492,8 @@ static int jit_subprogs(struct bpf_verifier_env *env)
prog->aux->bpf_exception_cb = (void *)func[env->exception_callback_subprog]->bpf_func;
prog->aux->exception_boundary = func[0]->aux->exception_boundary;
prog->aux->stack_arg_sp_adjust = func[0]->aux->stack_arg_sp_adjust;
+ prog->aux->exc = func[0]->aux->exc;
+ func[0]->aux->exc = NULL;
bpf_prog_jit_attempt_done(prog);
return 0;
out_free:
@@ -1943,6 +2074,8 @@ int bpf_do_misc_fixups(struct bpf_verifier_env *env)
goto next_insn;
if (insn->src_reg == BPF_PSEUDO_CALL)
goto next_insn;
+ if (env->insn_aux_data[i + delta].cleanup_resume_site)
+ goto next_insn;
if (insn->src_reg == BPF_PSEUDO_KFUNC_CALL) {
ret = bpf_fixup_kfunc_call(env, insn, insn_buf, i + delta, &cnt);
if (ret)
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 55c59f8b4c43..0151d7264278 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -31,6 +31,7 @@
#include <linux/buildid.h>
#include "../../lib/kstrtox.h"
+#include "exception.h"
/* If kernel subsystem is allowing eBPF programs to call this function,
* inside its own verifier_ops->get_func_proto() callback it should return
@@ -3398,8 +3399,36 @@ struct bpf_throw_ctx {
u64 sp;
u64 bp;
int cnt;
+ const struct bpf_prog *callee;
+ u64 callee_fp;
};
+static void bpf_run_cleanup_pad(struct bpf_throw_ctx *ctx, const struct bpf_prog *prog,
+ u64 ip, u64 fp)
+{
+ const struct bpf_exception_info *exc = prog->aux->exc;
+ const struct bpf_cleanup_range *rec;
+ u64 spill_base;
+
+ if (!exc || !exc->nr_ranges)
+ return;
+ rec = bpf_cleanup_pad_for_ip(prog, ip);
+ if (!rec)
+ return;
+
+ /*
+ * The callee is always another subprogram of this program -- the walk
+ * ends at any frame that is not one -- so its prologue spilled these
+ * registers and its exc is there to say where.
+ */
+ if (ctx->callee)
+ spill_base = ctx->callee_fp + ctx->callee->aux->exc->spill_off;
+ else
+ spill_base = fp + exc->throw_spill_off;
+
+ arch_bpf_run_cleanup_pad(rec->pad, fp, spill_base);
+}
+
static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
{
struct bpf_throw_ctx *ctx = cookie;
@@ -3416,6 +3445,11 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
if (!prog)
return !ctx->cnt;
ctx->cnt++;
+
+ bpf_run_cleanup_pad(ctx, prog, ip, bp);
+ ctx->callee = prog;
+ ctx->callee_fp = bp;
+
if (bpf_is_subprog(prog))
return true;
ctx->aux = prog->aux;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 10/20] bpf: Dispatch exception cleanup pads from bpf_throw()
2026-09-21 21:01 ` [PATCH bpf-next v4 10/20] bpf: Dispatch exception cleanup pads from bpf_throw() Yonghong Song
@ 2026-09-22 21:38 ` Eduard Zingerman
2026-09-23 3:22 ` Yonghong Song
0 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-22 21:38 UTC (permalink / raw)
To: Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Mon, 2026-09-21 at 14:01 -0700, Yonghong Song wrote:
...
> +static int cleanup_pad_body_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
> + u32 start, u32 end)
> +{
pad_body{,_bits} are redundant, jits have access to insn_aux_data.
all this machinery can be dropped.
> + unsigned long *bits;
> + u32 i, cnt = 0;
> +
> + for (i = start; i < end; i++)
> + if (env->insn_aux_data[i].in_cleanup_pad)
> + cnt++;
> + if (!cnt)
> + return 0;
> +
> + bits = bitmap_zalloc(end - start, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
> + if (!bits)
> + return -ENOMEM;
> +
> + for (i = start; i < end; i++)
> + if (env->insn_aux_data[i].in_cleanup_pad)
> + __set_bit(i - start, bits);
> +
> + sub->aux->exc->pad_body = bits;
> + sub->aux->exc->pad_body_bits = end - start;
> + return 0;
> +}
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 10/20] bpf: Dispatch exception cleanup pads from bpf_throw()
2026-09-22 21:38 ` Eduard Zingerman
@ 2026-09-23 3:22 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-23 3:22 UTC (permalink / raw)
To: Eduard Zingerman, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On 9/22/26 2:38 PM, Eduard Zingerman wrote:
> On Mon, 2026-09-21 at 14:01 -0700, Yonghong Song wrote:
>
> ...
>
>> +static int cleanup_pad_body_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
>> + u32 start, u32 end)
>> +{
> pad_body{,_bits} are redundant, jits have access to insn_aux_data.
> all this machinery can be dropped.
Good point. This can simplify the code a lot!
>
>> + unsigned long *bits;
>> + u32 i, cnt = 0;
>> +
>> + for (i = start; i < end; i++)
>> + if (env->insn_aux_data[i].in_cleanup_pad)
>> + cnt++;
>> + if (!cnt)
>> + return 0;
>> +
>> + bits = bitmap_zalloc(end - start, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
>> + if (!bits)
>> + return -ENOMEM;
>> +
>> + for (i = start; i < end; i++)
>> + if (env->insn_aux_data[i].in_cleanup_pad)
>> + __set_bit(i - start, bits);
>> +
>> + sub->aux->exc->pad_body = bits;
>> + sub->aux->exc->pad_body_bits = end - start;
>> + return 0;
>> +}
^ permalink raw reply [flat|nested] 80+ messages in thread
* [PATCH bpf-next v4 11/20] bpf, x86: Dispatch exception cleanup pads at run time
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (9 preceding siblings ...)
2026-09-21 21:01 ` [PATCH bpf-next v4 10/20] bpf: Dispatch exception cleanup pads from bpf_throw() Yonghong Song
@ 2026-09-21 21:01 ` Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 12/20] bpf, arm64: " Yonghong Song
` (9 subsequent siblings)
20 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:01 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Provide the arch half for x86-64: force the full callee-saved spill for a
program carrying a cleanup table so the walker can find a frame's r6-r9 in
its callee's prologue, record where that spill area starts, build the
native cleanup table from the JIT's addrs[], emit a bare return for a pad's
bpf_unwind_resume(), and hand control to a pad from
arch_bpf_run_cleanup_pad().
Support is gated on CONFIG_UNWINDER_ORC, the same requirement
arch_bpf_stack_walk() and therefore bpf_throw() already have here, and
bpf_cleanup_pad.o is built only there. Nothing else can call it, and a
frame-pointer build would have objtool validate a routine that has to carry
BPF r10 in rbp rather than a frame pointer ("call without frame pointer
save/setup", an error under CONFIG_OBJTOOL_WERROR).
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
arch/x86/net/Makefile | 3 ++
arch/x86/net/bpf_cleanup_pad.S | 80 +++++++++++++++++++++++++++++++++
arch/x86/net/bpf_jit_comp.c | 81 +++++++++++++++++++++++++++++-----
3 files changed, 152 insertions(+), 12 deletions(-)
create mode 100644 arch/x86/net/bpf_cleanup_pad.S
diff --git a/arch/x86/net/Makefile b/arch/x86/net/Makefile
index dddbefc0f439..8bb22644ba10 100644
--- a/arch/x86/net/Makefile
+++ b/arch/x86/net/Makefile
@@ -7,4 +7,7 @@ ifeq ($(CONFIG_X86_32),y)
obj-$(CONFIG_BPF_JIT) += bpf_jit_comp32.o
else
obj-$(CONFIG_BPF_JIT) += bpf_jit_comp.o bpf_timed_may_goto.o
+ ifdef CONFIG_UNWINDER_ORC
+ obj-$(CONFIG_BPF_JIT) += bpf_cleanup_pad.o
+ endif
endif
diff --git a/arch/x86/net/bpf_cleanup_pad.S b/arch/x86/net/bpf_cleanup_pad.S
new file mode 100644
index 000000000000..025a54e718a4
--- /dev/null
+++ b/arch/x86/net/bpf_cleanup_pad.S
@@ -0,0 +1,80 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf_cleanup_abi.h>
+#include <linux/linkage.h>
+#include <asm/nospec-branch.h>
+
+/*
+ * The x86-64 BPF JIT prologue spills, once bpf_cleanup_force_spill() makes it
+ * unconditional, r12, rbx, r13, r14 and r15 in that order -- so within the
+ * spill area the lowest address holds r15 and the highest r12. The throw-site
+ * spill the JIT emits uses the same layout, so the routine below reads both
+ * the same way:
+ *
+ * spill_base + 0 BPF r9 (r15)
+ * spill_base + 8 BPF r8 (r14)
+ * spill_base + 16 BPF r7 (r13)
+ * spill_base + 24 BPF r6 (rbx)
+ * spill_base + 32 r12 (arena base, not a BPF register)
+ */
+
+ .code64
+ .section .text, "ax"
+
+/*
+ * void arch_bpf_run_cleanup_pad(u64 pad, u64 frame_fp, u64 spill_base)
+ *
+ * rdi = native address of the landing pad
+ * rsi = frame pointer of the frame the pad belongs to
+ * rdx = spill area holding that frame's BPF callee-saved registers
+ *
+ * Give the pad the register state of its own frame and call it. The pad ends
+ * in the bare return the JIT emits for its bpf_unwind_resume(), so it comes
+ * back here rather than returning to its frame's caller. It runs on this
+ * stack, far below the frame it is cleaning up after, so nothing it calls can
+ * reach into that frame.
+ *
+ * rbp addresses the pad's frame and rsp this one, so the two are not the
+ * neighbours a JITed frame expects them to be. That is why
+ * cleanup_check_pad_insn() refuses a call in a pad that passes an argument on
+ * the stack: the JIT stages those relative to rbp, and the callee reads them
+ * relative to rsp.
+ */
+SYM_FUNC_START(arch_bpf_run_cleanup_pad)
+ ANNOTATE_NOENDBR
+
+ pushq %rbp
+ pushq %rbx
+ pushq %r12
+ pushq %r13
+ pushq %r14
+ pushq %r15
+ /* Keep the pad's entry rsp congruent to a normal call's. */
+ subq $8, %rsp
+
+ movq 0(%rdx), %r15
+ movq 8(%rdx), %r14
+ movq 16(%rdx), %r13
+ movq 24(%rdx), %rbx
+ movq 32(%rdx), %r12
+ /* rbp is BPF r10, so this is the whole of the pad's frame setup. */
+ movq %rsi, %rbp
+
+ /* CALL_NOSPEC needs the target in a register; rcx is BPF r4, dead. */
+ movq %rdi, %rcx
+
+ /* BPF r0 on the way into a pad, not whatever the kernel left in rax. */
+ movl $BPF_PAD_ENTRY_R0, %eax
+
+ CALL_NOSPEC rcx
+
+ addq $8, %rsp
+ popq %r15
+ popq %r14
+ popq %r13
+ popq %r12
+ popq %rbx
+ popq %rbp
+ RET
+SYM_FUNC_END(arch_bpf_run_cleanup_pad)
diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index d4a980140b48..537bf5d20430 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -357,6 +357,12 @@ struct jit_context {
/* Number of bytes that will be skipped on tailcall */
#define X86_TAIL_CALL_OFFSET (12 + ENDBR_INSN_SIZE)
+/*
+ * Throw-site spill: r15, r14, r13, rbx, r12 low to high, the layout the
+ * prologue's pushes leave, so arch_bpf_run_cleanup_pad() reads both alike.
+ */
+#define X86_CLEANUP_SPILL_SZ (5 * 8)
+
static void push_r9(u8 **pprog)
{
u8 *prog = *pprog;
@@ -832,7 +838,7 @@ static void emit_bpf_tail_call_indirect(struct bpf_prog *bpf_prog,
/* Inc tail_call_cnt if the slot is populated. */
EMIT4(0x48, 0x83, 0x00, 0x01); /* add qword ptr [rax], 1 */
- if (bpf_prog->aux->exception_boundary) {
+ if (bpf_prog->aux->exception_boundary || bpf_cleanup_force_spill(bpf_prog)) {
pop_callee_regs(&prog, all_callee_regs_used);
pop_r12(&prog);
} else {
@@ -899,7 +905,7 @@ static void emit_bpf_tail_call_direct(struct bpf_prog *bpf_prog,
/* Inc tail_call_cnt if the slot is populated. */
EMIT4(0x48, 0x83, 0x00, 0x01); /* add qword ptr [rax], 1 */
- if (bpf_prog->aux->exception_boundary) {
+ if (bpf_prog->aux->exception_boundary || bpf_cleanup_force_spill(bpf_prog)) {
pop_callee_regs(&prog, all_callee_regs_used);
pop_r12(&prog);
} else {
@@ -1977,6 +1983,7 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
u8 *ip, *prog = temp;
u32 stack_depth;
int callee_saved_size;
+ u32 throw_spill, prologue_depth;
s32 outgoing_arg_base;
int err;
@@ -2015,7 +2022,10 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
detect_reg_usage(insn, insn_cnt, callee_regs_used);
- emit_prologue(&prog, image, stack_depth,
+ throw_spill = bpf_cleanup_needs_throw_spill(bpf_prog) ? X86_CLEANUP_SPILL_SZ : 0;
+ prologue_depth = stack_depth + throw_spill;
+
+ emit_prologue(&prog, image, prologue_depth,
bpf_prog_was_classic(bpf_prog), tail_call_reachable,
bpf_is_subprog(bpf_prog), bpf_prog->aux->exception_cb);
@@ -2024,7 +2034,7 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
/* Exception callback will clobber callee regs for its own use, and
* restore the original callee regs from main prog's stack frame.
*/
- if (bpf_prog->aux->exception_boundary) {
+ if (bpf_prog->aux->exception_boundary || bpf_cleanup_force_spill(bpf_prog)) {
/* We also need to save r12, which is not mapped to any BPF
* register, as we throw after entry into the kernel, which may
* overwrite r12.
@@ -2039,9 +2049,10 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
/* Compute callee-saved register area size. */
callee_saved_size = 0;
- if (bpf_prog->aux->exception_boundary || arena_vm_start)
+ if (bpf_prog->aux->exception_boundary || bpf_cleanup_force_spill(bpf_prog) ||
+ arena_vm_start)
callee_saved_size += 8; /* r12 */
- if (bpf_prog->aux->exception_boundary) {
+ if (bpf_prog->aux->exception_boundary || bpf_cleanup_force_spill(bpf_prog)) {
callee_saved_size += 4 * 8; /* rbx, r13, r14, r15 */
} else {
int j;
@@ -2063,7 +2074,21 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
* Note that tail_call_reachable is guaranteed to be false when
* stack args exist, so tcc pushes need not be accounted for.
*/
- outgoing_arg_base = -(round_up(stack_depth, 8) + callee_saved_size);
+ outgoing_arg_base = -(round_up(stack_depth, 8) + throw_spill + callee_saved_size);
+
+ /*
+ * Lowest address of each spill area, as an offset from rbp; see
+ * bpf_cleanup_pad.S for the layout. The 16 is the tail call counter
+ * pair emit_prologue_tail_call() pushes above the callee-saved one.
+ * Only a frame that throws has a throw-site area.
+ */
+ if (bpf_cleanup_force_spill(bpf_prog))
+ bpf_prog->aux->exc->spill_off = -(round_up(stack_depth, 8) + throw_spill +
+ (tail_call_reachable ? 16 : 0) +
+ callee_saved_size);
+ if (bpf_cleanup_needs_throw_spill(bpf_prog))
+ bpf_prog->aux->exc->throw_spill_off =
+ -(round_up(stack_depth, 8) + throw_spill);
/*
* Allocate outgoing stack arg area for args 7+ only.
@@ -2110,7 +2135,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
dst_reg = X86_REG_R9;
}
- if (bpf_insn_is_indirect_target(env, bpf_prog, i - 1))
+ if (bpf_insn_is_indirect_target(env, bpf_prog, i - 1) ||
+ bpf_cleanup_insn_is_pad(bpf_prog, i - 1))
EMIT_ENDBR();
ip = image + addrs[i - 1] + (prog - temp);
@@ -2903,9 +2929,27 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
case BPF_JMP | BPF_CALL: {
const struct btf_func_model *fm = NULL;
+ if (bpf_cleanup_insn_is_throw(bpf_prog, i - 1)) {
+ /* Spill r6-r9 and r12 where the bpf_throw() walker looks. */
+ s32 off = bpf_prog->aux->exc->throw_spill_off;
+ u8 *spill = prog;
+
+ emit_stx(&prog, BPF_DW, BPF_REG_FP, BPF_REG_9, off + 0);
+ emit_stx(&prog, BPF_DW, BPF_REG_FP, BPF_REG_8, off + 8);
+ emit_stx(&prog, BPF_DW, BPF_REG_FP, BPF_REG_7, off + 16);
+ emit_stx(&prog, BPF_DW, BPF_REG_FP, BPF_REG_6, off + 24);
+ emit_stx(&prog, BPF_DW, BPF_REG_FP, X86_REG_R12, off + 32);
+ ip += prog - spill;
+ }
+
+ if (bpf_cleanup_insn_is_resume(bpf_prog, i - 1)) {
+ emit_return(&prog, image + addrs[i - 1] + (prog - temp));
+ break;
+ }
+
func = (u8 *) __bpf_call_base + imm32;
if (src_reg == BPF_PSEUDO_CALL && tail_call_reachable) {
- LOAD_TAIL_CALL_CNT_PTR(stack_depth);
+ LOAD_TAIL_CALL_CNT_PTR(prologue_depth);
ip += 7;
}
if (!imm32)
@@ -2948,13 +2992,13 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
&prog,
ip,
callee_regs_used,
- stack_depth,
+ prologue_depth,
ctx);
else
emit_bpf_tail_call_indirect(bpf_prog,
&prog,
callee_regs_used,
- stack_depth,
+ prologue_depth,
ip,
ctx);
break;
@@ -3215,7 +3259,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
}
/* Deallocate outgoing args 7+ area. */
emit_add_rsp(&prog, outgoing_rsp);
- if (bpf_prog->aux->exception_boundary) {
+ if (bpf_prog->aux->exception_boundary ||
+ bpf_cleanup_force_spill(bpf_prog)) {
pop_callee_regs(&prog, all_callee_regs_used);
pop_r12(&prog);
} else {
@@ -4385,6 +4430,13 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
*/
bpf_prog_update_insn_ptrs(prog, addrs, image);
+ /*
+ * Same mapping, consumed by the bpf_throw() frame walker:
+ * turn the cleanup records into native address ranges now
+ * that the image is final.
+ */
+ bpf_cleanup_fill_native_ranges(prog, addrs, image);
+
/*
* ctx.prog_offset is used when CFI preambles put code *before*
* the function. See emit_cfi(). For FineIBT specifically this code
@@ -4501,6 +4553,11 @@ bool bpf_jit_supports_exceptions(void)
return IS_ENABLED(CONFIG_UNWINDER_ORC);
}
+bool bpf_jit_supports_cleanup_pads(void)
+{
+ return IS_ENABLED(CONFIG_UNWINDER_ORC);
+}
+
bool bpf_jit_supports_private_stack(void)
{
return true;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* [PATCH bpf-next v4 12/20] bpf, arm64: Dispatch exception cleanup pads at run time
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (10 preceding siblings ...)
2026-09-21 21:01 ` [PATCH bpf-next v4 11/20] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
@ 2026-09-21 21:01 ` Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 13/20] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
` (8 subsequent siblings)
20 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:01 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
The same arch half for arm64: force the full callee-saved spill, record
where it starts, build the native table from the JIT's byte offsets, and
hand control to a pad from arch_bpf_run_cleanup_pad().
Unlike x86-64 a pad cannot simply return: every call it makes clobbers x30,
so no return address survives to its resume. x23 carries it instead.
bpf2a64[] maps nothing to x23 or x24 -- they are pushed only to keep the
frame shape the exception callback expects -- so nothing else in generated
code touches them, and being callee-saved they survive every kfunc the pad
calls. The JIT emits "br x23" for a pad's bpf_unwind_resume().
x24 is the other one, and it is what makes a pad able to touch its own
frame at all. Generated code addresses the BPF frame through the stack
pointer -- the frame sits directly on top of it, which turns every offset
positive and each access into one instruction -- and in a pad the stack
pointer is the walker's. So the JIT has a pad recompute the equivalent of
its frame's stack pointer from BPF r10 on entry, into x24, and addresses
the frame off x24 for every instruction the previous patches marked as
running only while unwinding. One instruction per pad, and the accesses
keep the shape and the immediate range they have everywhere else.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
arch/arm64/net/Makefile | 2 +-
arch/arm64/net/bpf_cleanup_pad.S | 95 ++++++++++++++++++++++++++++++++
arch/arm64/net/bpf_jit_comp.c | 91 ++++++++++++++++++++++++++++--
3 files changed, 181 insertions(+), 7 deletions(-)
create mode 100644 arch/arm64/net/bpf_cleanup_pad.S
diff --git a/arch/arm64/net/Makefile b/arch/arm64/net/Makefile
index 3ae382bfca87..ebec2a44a52b 100644
--- a/arch/arm64/net/Makefile
+++ b/arch/arm64/net/Makefile
@@ -2,4 +2,4 @@
#
# ARM64 networking code
#
-obj-$(CONFIG_BPF_JIT) += bpf_jit_comp.o bpf_timed_may_goto.o
+obj-$(CONFIG_BPF_JIT) += bpf_jit_comp.o bpf_timed_may_goto.o bpf_cleanup_pad.o
diff --git a/arch/arm64/net/bpf_cleanup_pad.S b/arch/arm64/net/bpf_cleanup_pad.S
new file mode 100644
index 000000000000..ef441241949e
--- /dev/null
+++ b/arch/arm64/net/bpf_cleanup_pad.S
@@ -0,0 +1,95 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bpf_cleanup_abi.h>
+#include <linux/linkage.h>
+
+/*
+ * A frame's prologue pushes the tail call counter pair, and then -- once
+ * bpf_cleanup_force_spill() says so -- x19/x20, x21/x22, x23/x24, x25/x26 and
+ * x27/x28. Each A64_PUSH pre-decrements, so the lowest address of the spill
+ * area holds x27 and the highest x20:
+ *
+ * spill_base + 0 x27 (private stack pointer)
+ * spill_base + 8 x28 (arena base)
+ * spill_base + 16 x25 (BPF r10, the frame pointer)
+ * spill_base + 24 x26 (tail call counter pointer)
+ * spill_base + 32 x23 -- the pad's own, see below
+ * spill_base + 40 x24 -- likewise
+ * spill_base + 48 x21 (BPF r8)
+ * spill_base + 56 x22 (BPF r9)
+ * spill_base + 64 x19 (BPF r6)
+ * spill_base + 72 x20 (BPF r7)
+ *
+ * Neither x23 nor x24 is restored from that spill: the pad has its own use for
+ * both. bpf2a64[] maps nothing to either -- they are pushed only to keep the
+ * frame shape the exception callback expects -- so nothing else in generated
+ * code touches them, and being callee-saved they survive every call the pad
+ * makes.
+ *
+ * x23 is the pad's return address. Unlike x86-64 a pad cannot simply return:
+ * every call it makes clobbers x30, so nothing is left to return through by
+ * the time it reaches its resume. The JIT emits "br x23" for the pad's
+ * bpf_unwind_resume() and this routine puts .Lcleanup_pad_done there.
+ *
+ * x24 is where the pad's frame is anchored. Generated code addresses the BPF
+ * frame through the stack pointer, which here is this routine's rather than
+ * the unwinding frame's, so the JIT has the pad recompute the equivalent from
+ * BPF r10 on entry and address its frame off x24 for as long as it runs.
+ *
+ * Both of those branches are indirect, so both targets carry a BTI landing
+ * marker: the JIT emits one at each pad, and .Lcleanup_pad_done below has one
+ * of its own.
+ */
+
+ .text
+
+/*
+ * void arch_bpf_run_cleanup_pad(u64 pad, u64 frame_fp, u64 spill_base)
+ *
+ * x0 = native address of the landing pad
+ * x1 = frame pointer of the frame the pad belongs to (unused here: BPF r10 is
+ * x25, which the spill area already holds)
+ * x2 = spill area holding that frame's BPF callee-saved registers
+ *
+ * Give the pad the register state of its own frame and call it. It runs on
+ * this stack, far below the frame it is cleaning up after, so nothing it
+ * calls can reach into that frame.
+ */
+SYM_FUNC_START(arch_bpf_run_cleanup_pad)
+ /* Save the kernel's callee-saved registers; the pad owns them next. */
+ stp x29, x30, [sp, #-96]!
+ mov x29, sp
+ stp x19, x20, [sp, #16]
+ stp x21, x22, [sp, #32]
+ stp x23, x24, [sp, #48]
+ stp x25, x26, [sp, #64]
+ stp x27, x28, [sp, #80]
+
+ /* x9 is BPF_REG_AX, so the pad's address does not stay in BPF r1. */
+ mov x9, x0
+
+ ldp x27, x28, [x2, #0]
+ ldp x25, x26, [x2, #16]
+ ldp x21, x22, [x2, #48]
+ ldp x19, x20, [x2, #64]
+
+ /* BPF r0 (x8) on the way into a pad, not whatever the kernel left. */
+ mov x8, #BPF_PAD_ENTRY_R0
+
+ /* Where the pad's resume branches back to. */
+ adr x23, .Lcleanup_pad_done
+
+ br x9
+
+.Lcleanup_pad_done:
+ /* Reached by the pad's "br x23", so it is an indirect branch target. */
+ bti j
+ ldp x19, x20, [sp, #16]
+ ldp x21, x22, [sp, #32]
+ ldp x23, x24, [sp, #48]
+ ldp x25, x26, [sp, #64]
+ ldp x27, x28, [sp, #80]
+ ldp x29, x30, [sp], #96
+ ret
+SYM_FUNC_END(arch_bpf_run_cleanup_pad)
diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
index 6c04fee46876..8c280fda8a72 100644
--- a/arch/arm64/net/bpf_jit_comp.c
+++ b/arch/arm64/net/bpf_jit_comp.c
@@ -75,6 +75,19 @@ static const int bpf2a64[] = {
[ARENA_VM_START] = A64_R(28),
};
+/*
+ * Throw-site spill: the five pairs push_callee_regs() forces on, same size and
+ * slot order, so arch_bpf_run_cleanup_pad() reads both alike.
+ */
+#define A64_CLEANUP_SPILL_SZ (5 * 16)
+
+/*
+ * Where a landing pad's frame is anchored, since the stack pointer generated
+ * code normally addresses it through is the walker's inside a pad. bpf2a64[]
+ * maps nothing to x24, so nothing else in generated code touches it.
+ */
+#define A64_CLEANUP_FP A64_R(24)
+
struct jit_ctx {
const struct bpf_prog *prog;
int idx;
@@ -87,6 +100,8 @@ struct jit_ctx {
__le32 *ro_image;
u32 stack_size;
u16 stack_arg_size;
+ /* Bytes reserved for the throw-site spill; 0 unless this frame throws. */
+ u32 throw_spill;
u64 user_vm_start;
u64 arena_vm_start;
bool fp_used;
@@ -432,7 +447,7 @@ static void push_callee_regs(struct jit_ctx *ctx)
* Callee-saved registers as the exception callback needs to recover
* all ARM64 Callee-saved registers in its epilogue.
*/
- if (ctx->prog->aux->exception_boundary) {
+ if (ctx->prog->aux->exception_boundary || bpf_cleanup_force_spill(ctx->prog)) {
emit(A64_PUSH(A64_R(19), A64_R(20), A64_SP), ctx);
emit(A64_PUSH(A64_R(21), A64_R(22), A64_SP), ctx);
emit(A64_PUSH(A64_R(23), A64_R(24), A64_SP), ctx);
@@ -466,7 +481,8 @@ static void pop_callee_regs(struct jit_ctx *ctx)
* program's stack frame, so recover these extra registers in the above
* two cases.
*/
- if (aux->exception_boundary || aux->exception_cb) {
+ if (aux->exception_boundary || aux->exception_cb ||
+ bpf_cleanup_force_spill(ctx->prog)) {
emit(A64_POP(A64_R(27), A64_R(28), A64_SP), ctx);
emit(A64_POP(A64_R(25), A64_R(26), A64_SP), ctx);
emit(A64_POP(A64_R(23), A64_R(24), A64_SP), ctx);
@@ -602,6 +618,22 @@ static int build_prologue(struct jit_ctx *ctx, bool ebpf_from_cbpf)
emit(A64_SUB_I(1, A64_SP, A64_FP, 96), ctx);
}
+ /*
+ * Lowest address of each spill area, as an offset from A64_FP; see
+ * bpf_cleanup_pad.S for the layout. The 16 is the tail call counter
+ * pair pushed just below the frame record, and the throw-site area,
+ * which only a frame that throws has, sits below the callee-saved one
+ * rather than in the program stack.
+ */
+ if (bpf_cleanup_force_spill(prog))
+ prog->aux->exc->spill_off = -(16 + A64_CLEANUP_SPILL_SZ);
+ if (bpf_cleanup_needs_throw_spill(prog)) {
+ ctx->throw_spill = A64_CLEANUP_SPILL_SZ;
+ emit(A64_SUB_I(1, A64_SP, A64_SP, ctx->throw_spill), ctx);
+ prog->aux->exc->throw_spill_off =
+ -(16 + A64_CLEANUP_SPILL_SZ) - ctx->throw_spill;
+ }
+
/* Stack must be multiples of 16B */
ctx->stack_size = round_up(prog->aux->stack_depth, 16);
@@ -691,6 +723,9 @@ static int emit_bpf_tail_call(struct jit_ctx *ctx)
if (ctx->stack_size && !ctx->priv_sp_used)
emit(A64_ADD_I(1, A64_SP, A64_SP, ctx->stack_size), ctx);
+ if (ctx->throw_spill)
+ emit(A64_ADD_I(1, A64_SP, A64_SP, ctx->throw_spill), ctx);
+
pop_callee_regs(ctx);
/* goto *(prog->bpf_func + prologue_offset); */
@@ -1055,6 +1090,9 @@ static void build_epilogue(struct jit_ctx *ctx, bool was_classic)
if (ctx->stack_size && !ctx->priv_sp_used)
emit(A64_ADD_I(1, A64_SP, A64_SP, ctx->stack_size), ctx);
+ if (ctx->throw_spill)
+ emit(A64_ADD_I(1, A64_SP, A64_SP, ctx->throw_spill), ctx);
+
pop_callee_regs(ctx);
emit(A64_POP(A64_ZR, ptr, A64_SP), ctx);
@@ -1367,6 +1405,8 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
const s16 off = insn->off;
const s32 imm = insn->imm;
const int i = insn - ctx->prog->insnsi;
+ const bool in_pad = bpf_cleanup_insn_in_pad(ctx->prog, i);
+ const bool pad_head = bpf_cleanup_insn_is_pad(ctx->prog, i);
const bool is64 = BPF_CLASS(code) == BPF_ALU64 ||
BPF_CLASS(code) == BPF_JMP;
u8 jmp_cond;
@@ -1378,9 +1418,13 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
int ret;
bool sign_extend;
- if (bpf_insn_is_indirect_target(env, ctx->prog, i))
+ if (bpf_insn_is_indirect_target(env, ctx->prog, i) || pad_head)
emit_bti(A64_BTI_J, ctx);
+ if (pad_head)
+ emit(A64_SUB_I(1, A64_CLEANUP_FP, fp,
+ ctx->stack_size + ctx->stack_arg_size), ctx);
+
switch (code) {
/* dst = src */
case BPF_ALU | BPF_MOV | BPF_X:
@@ -1743,6 +1787,26 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
u64 func_addr;
u32 cpu_offset;
+ if (bpf_cleanup_insn_is_throw(ctx->prog, i)) {
+ /* Spill where the bpf_throw() walker looks. */
+ const s32 off = ctx->prog->aux->exc->throw_spill_off;
+
+ emit(A64_SUB_I(1, tmp, A64_FP, -off), ctx);
+ emit(A64_STR64I(bpf2a64[PRIVATE_SP], tmp, 0), ctx);
+ emit(A64_STR64I(bpf2a64[ARENA_VM_START], tmp, 8), ctx);
+ emit(A64_STR64I(bpf2a64[BPF_REG_FP], tmp, 16), ctx);
+ emit(A64_STR64I(bpf2a64[TCCNT_PTR], tmp, 24), ctx);
+ emit(A64_STR64I(bpf2a64[BPF_REG_8], tmp, 48), ctx);
+ emit(A64_STR64I(bpf2a64[BPF_REG_9], tmp, 56), ctx);
+ emit(A64_STR64I(bpf2a64[BPF_REG_6], tmp, 64), ctx);
+ emit(A64_STR64I(bpf2a64[BPF_REG_7], tmp, 72), ctx);
+ }
+
+ if (bpf_cleanup_insn_is_resume(ctx->prog, i)) {
+ emit(A64_BR(A64_R(23)), ctx);
+ break;
+ }
+
/* Implement helper call to bpf_get_smp_processor_id() inline */
if (insn->src_reg == 0 && insn->imm == BPF_FUNC_get_smp_processor_id) {
cpu_offset = offsetof(struct thread_info, cpu);
@@ -1854,7 +1918,8 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
src = tmp2;
}
if (src == fp) {
- src_adj = ctx->priv_sp_used ? priv_sp : A64_SP;
+ src_adj = ctx->priv_sp_used ? priv_sp :
+ in_pad ? A64_CLEANUP_FP : A64_SP;
off_adj = off + ctx->stack_size;
if (!ctx->priv_sp_used)
off_adj += ctx->stack_arg_size;
@@ -1952,7 +2017,8 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
dst = tmp3;
}
if (dst == fp) {
- dst_adj = ctx->priv_sp_used ? priv_sp : A64_SP;
+ dst_adj = ctx->priv_sp_used ? priv_sp :
+ in_pad ? A64_CLEANUP_FP : A64_SP;
off_adj = off + ctx->stack_size;
if (!ctx->priv_sp_used)
off_adj += ctx->stack_arg_size;
@@ -2021,7 +2087,8 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
dst = tmp2;
}
if (dst == fp) {
- dst_adj = ctx->priv_sp_used ? priv_sp : A64_SP;
+ dst_adj = ctx->priv_sp_used ? priv_sp :
+ in_pad ? A64_CLEANUP_FP : A64_SP;
off_adj = off + ctx->stack_size;
if (!ctx->priv_sp_used)
off_adj += ctx->stack_arg_size;
@@ -2410,6 +2477,13 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
* reasons, expects to point to the next instruction)
*/
bpf_prog_update_insn_ptrs(prog, ctx.offset, ctx.ro_image);
+
+ /*
+ * Same byte offsets, consumed by the bpf_throw() frame walker:
+ * turn the cleanup records into native address ranges now that
+ * the image is final.
+ */
+ bpf_cleanup_fill_native_ranges(prog, ctx.offset, ctx.ro_image);
out_off:
if (!ro_header && priv_stack_ptr) {
free_percpu(priv_stack_ptr);
@@ -3385,6 +3459,11 @@ bool bpf_jit_supports_exceptions(void)
return true;
}
+bool bpf_jit_supports_cleanup_pads(void)
+{
+ return true;
+}
+
bool bpf_jit_supports_arena(void)
{
return true;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* [PATCH bpf-next v4 13/20] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (11 preceding siblings ...)
2026-09-21 21:01 ` [PATCH bpf-next v4 12/20] bpf, arm64: " Yonghong Song
@ 2026-09-21 21:01 ` Yonghong Song
2026-09-21 21:13 ` sashiko-bot
2026-09-21 21:01 ` [PATCH bpf-next v4 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
` (7 subsequent siblings)
20 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:01 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
LLVM terminates a cleanup landing pad with a call to _Unwind_Resume: the
base unwind ABI's entry point for carrying an unwind on once a frame's
cleanups have run. The kernel provides that terminator as a kfunc, but
not under that name: claiming _Unwind_Resume in the kernel's own symbol
table, for a function whose body never runs, would be needlessly confusing,
so it is called bpf_unwind_resume.
Both load paths take the detour. A direct load resolves the name against
the kernel's BTF while libbpf runs. A light skeleton instead writes the
name into the loader program's blob of bytes, for that program to resolve
when it runs, so the name recorded there has to be translated as well.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/libbpf.c | 22 +++++++++++++++++-----
1 file changed, 17 insertions(+), 5 deletions(-)
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index cd1ea1bb53cb..13ae9175e968 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -8379,6 +8379,15 @@ static void fixup_verifier_log(struct bpf_program *prog, char *buf, size_t buf_s
}
}
+/* LLVM terminates a cleanup landing pad with a call to _Unwind_Resume, the
+ * base unwind ABI's entry point for carrying an unwind on once a frame's
+ * cleanups have run. The kernel knows it as bpf_unwind_resume.
+ */
+static const char *kern_extern_name(const char *name)
+{
+ return strcmp(name, "_Unwind_Resume") ? name : "bpf_unwind_resume";
+}
+
static int bpf_program_record_relos(struct bpf_program *prog)
{
struct bpf_object *obj = prog->obj;
@@ -8395,12 +8404,12 @@ static int bpf_program_record_relos(struct bpf_program *prog)
continue;
kind = btf_is_var(btf__type_by_id(obj->btf, ext->btf_id)) ?
BTF_KIND_VAR : BTF_KIND_FUNC;
- bpf_gen__record_extern(obj->gen_loader, ext->name,
+ bpf_gen__record_extern(obj->gen_loader, kern_extern_name(ext->name),
ext->is_weak, !ext->ksym.type_id,
true, kind, relo->insn_idx);
break;
case RELO_EXTERN_CALL:
- bpf_gen__record_extern(obj->gen_loader, ext->name,
+ bpf_gen__record_extern(obj->gen_loader, kern_extern_name(ext->name),
ext->is_weak, false, false, BTF_KIND_FUNC,
relo->insn_idx);
break;
@@ -8876,17 +8885,20 @@ static int bpf_object__resolve_ksym_func_btf_id(struct bpf_object *obj,
struct module_btf *mod_btf = NULL;
const struct btf_type *kern_func;
struct btf *kern_btf = NULL;
+ const char *local_name, *kern_name;
int ret;
local_func_proto_id = ext->ksym.type_id;
- kfunc_id = find_ksym_btf_id(obj, ext->essent_name ?: ext->name, BTF_KIND_FUNC, &kern_btf,
- &mod_btf);
+ local_name = ext->essent_name ?: ext->name;
+ kern_name = kern_extern_name(local_name);
+
+ kfunc_id = find_ksym_btf_id(obj, kern_name, BTF_KIND_FUNC, &kern_btf, &mod_btf);
if (kfunc_id < 0) {
if (kfunc_id == -ESRCH && ext->is_weak)
return 0;
pr_warn("extern (func ksym) '%s': not found in kernel or module BTFs\n",
- ext->name);
+ kern_name != local_name ? kern_name : ext->name);
return kfunc_id;
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 13/20] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
2026-09-21 21:01 ` [PATCH bpf-next v4 13/20] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
@ 2026-09-21 21:13 ` sashiko-bot
0 siblings, 0 replies; 80+ messages in thread
From: sashiko-bot @ 2026-09-21 21:13 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] bpf: multi-line comment style violation in kern_extern_name()
--
commit 293bdf5bcf5e7193069a28fa045283b34511449e
Author: Yonghong Song <yonghong.song@linux.dev>
libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
LLVM terminates a cleanup landing pad with a call to _Unwind_Resume.
The kernel provides that terminator as a kfunc, but under the name
bpf_unwind_resume. This patch translates the name during both direct
load and light skeleton resolutions to ensure it resolves correctly.
> diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
> index cd1ea1bb53cbf..13ae9175e9686 100644
> --- a/tools/lib/bpf/libbpf.c
> +++ b/tools/lib/bpf/libbpf.c
> @@ -8379,6 +8379,15 @@ static void fixup_verifier_log(struct bpf_program *prog, char *buf, size_t buf_s
> }
> }
>
> +/* LLVM terminates a cleanup landing pad with a call to _Unwind_Resume, the
> + * base unwind ABI's entry point for carrying an unwind on once a frame's
> + * cleanups have run. The kernel knows it as bpf_unwind_resume.
> + */
[Severity: Low]
This isn't a bug, but does this newly added comment violate the strict BPF
subsystem multi-line comment style rules? The BPF subsystem guide requires
multi-line comments to have the opening /* on its own line, with the text
starting on the subsequent line.
> +static const char *kern_extern_name(const char *name)
> +{
> + return strcmp(name, "_Unwind_Resume") ? name : "bpf_unwind_resume";
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260921210033.1715000-1-yonghong.song@linux.dev?part=13
^ permalink raw reply [flat|nested] 80+ messages in thread
* [PATCH bpf-next v4 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (12 preceding siblings ...)
2026-09-21 21:01 ` [PATCH bpf-next v4 13/20] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
@ 2026-09-21 21:01 ` Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
` (6 subsequent siblings)
20 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:01 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Let a caller hand the kernel an exception cleanup table: struct
bpf_prog_load_opts grows cleanup_info, cleanup_info_cnt and
cleanup_info_rec_size, and bpf_prog_load() passes all three on to
BPF_PROG_LOAD.
The attr size it computes grows with them, by one field rather than
three: cleanup_info_cnt is the last of the three in the BPF_PROG_LOAD
attr, so offsetofend() on it already covers cleanup_info and
cleanup_info_rec_size.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/bpf.c | 6 +++++-
tools/lib/bpf/bpf.h | 7 ++++++-
2 files changed, 11 insertions(+), 2 deletions(-)
diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
index a9de7f107cf7..03315b3a3a19 100644
--- a/tools/lib/bpf/bpf.c
+++ b/tools/lib/bpf/bpf.c
@@ -295,7 +295,7 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
const struct bpf_insn *insns, size_t insn_cnt,
struct bpf_prog_load_opts *opts)
{
- const size_t attr_sz = offsetofend(union bpf_attr, keyring_id);
+ const size_t attr_sz = offsetofend(union bpf_attr, cleanup_info_cnt);
void *finfo = NULL, *linfo = NULL;
const char *func_info, *line_info;
__u32 log_size, log_level, attach_prog_fd, attach_btf_obj_fd;
@@ -370,6 +370,10 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
attr.fd_array = ptr_to_u64(OPTS_GET(opts, fd_array, NULL));
attr.fd_array_cnt = OPTS_GET(opts, fd_array_cnt, 0);
+ attr.cleanup_info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
+ attr.cleanup_info = ptr_to_u64(OPTS_GET(opts, cleanup_info, NULL));
+ attr.cleanup_info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
+
if (log_level) {
attr.log_buf = ptr_to_u64(log_buf);
attr.log_size = log_size;
diff --git a/tools/lib/bpf/bpf.h b/tools/lib/bpf/bpf.h
index 490e8cb4ba53..ec542fd68f63 100644
--- a/tools/lib/bpf/bpf.h
+++ b/tools/lib/bpf/bpf.h
@@ -128,9 +128,14 @@ struct bpf_prog_load_opts {
/* if set, provides the length of fd_array */
__u32 fd_array_cnt;
+
+ /* exception cleanup table, from the .bpf_cleanup section */
+ const void *cleanup_info;
+ __u32 cleanup_info_cnt;
+ __u32 cleanup_info_rec_size;
size_t :0;
};
-#define bpf_prog_load_opts__last_field fd_array_cnt
+#define bpf_prog_load_opts__last_field cleanup_info_rec_size
LIBBPF_API int bpf_prog_load(enum bpf_prog_type prog_type,
const char *prog_name, const char *license,
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* [PATCH bpf-next v4 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (13 preceding siblings ...)
2026-09-21 21:01 ` [PATCH bpf-next v4 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
@ 2026-09-21 21:01 ` Yonghong Song
2026-09-21 21:20 ` sashiko-bot
2026-09-21 21:01 ` [PATCH bpf-next v4 16/20] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
` (5 subsequent siblings)
20 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:01 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
Parse the compiler-emitted .bpf_cleanup section and hand the resulting
table to BPF_PROG_LOAD.
Each record is three 4-byte fields, and each field is a byte offset into
some code section named by a matching .rel.bpf_cleanup relocation.
bpf_object__init_cleanup_info() resolves both halves once at open time and
keeps (section, instruction index) pairs; it rejects a record whose field
has no relocation, whose relocation is not one of the two 32-bit types that
spell a data reference to a code section -- R_BPF_64_NODYLD32 from LLVM,
R_BPF_64_ABS32 from GNU as -- or whose offset is not instruction aligned.
The records are sorted by begin_off once the offsets are final: the kernel
wants the table sorted with disjoint ranges so that it can find the record
covering a call site with a binary search, and records arrive in
.bpf_cleanup order, which says nothing about where the subprograms they
describe were appended. Overlapping ranges are reported here, where the
program name and both regions are still at hand.
bpf_object_load_prog() then passes the per-program table through the
bpf_prog_load() options added in the previous patch, with the record size
carried on the program the way func_info and line_info carry theirs rather
than taken from a sizeof() at the call site. bpf_program__clone() carries
it too. That is a second load path -- the one veristat uses -- and without
the table the kernel sees landing pads nothing reaches and refuses the
program with "unreachable insn".
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/libbpf.c | 305 ++++++++++++++++++++++++++++++++
tools/lib/bpf/libbpf_internal.h | 3 +
2 files changed, 308 insertions(+)
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index 13ae9175e968..e3e4db59fb5f 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -514,6 +514,11 @@ struct bpf_program {
void *line_info;
__u32 line_info_rec_size;
__u32 line_info_cnt;
+
+ struct bpf_cleanup_info *cleanup_info;
+ __u32 cleanup_info_rec_size;
+ __u32 cleanup_info_cnt;
+
__u32 prog_flags;
__u8 hash[SHA256_DIGEST_LENGTH];
@@ -549,6 +554,7 @@ struct bpf_struct_ops {
#define STRUCT_OPS_SEC ".struct_ops"
#define STRUCT_OPS_LINK_SEC ".struct_ops.link"
#define ARENA_SEC ".addr_space.1"
+#define CLEANUP_SEC ".bpf_cleanup"
enum libbpf_map_type {
LIBBPF_MAP_UNSPEC,
@@ -677,6 +683,25 @@ struct elf_sec_desc {
Elf_Data *data;
};
+#define CLEANUP_REC_FIELDS (sizeof(struct bpf_cleanup_info) / sizeof(__u32))
+
+/* Index of each field of struct bpf_cleanup_info, read as an array of __u32. */
+enum {
+ CLEANUP_REC_BEGIN,
+ CLEANUP_REC_END,
+ CLEANUP_REC_PAD,
+};
+
+/* One (begin, end, landing_pad) triple from .bpf_cleanup, with each field
+ * resolved from its relocation to an ELF section plus a section-relative
+ * instruction index. The mapping to final program instruction indices can only
+ * happen after subprogram placement, which differs per main program.
+ */
+struct cleanup_raw_rec {
+ int sec_idx[CLEANUP_REC_FIELDS];
+ size_t insn_idx[CLEANUP_REC_FIELDS];
+};
+
struct elf_state {
int fd;
const void *obj_buf;
@@ -696,6 +721,8 @@ struct elf_state {
bool has_st_ops;
int arena_data_shndx;
int jumptables_data_shndx;
+ Elf_Data *cleanup_data;
+ int cleanup_shndx;
};
struct usdt_manager;
@@ -773,6 +800,9 @@ struct bpf_object {
void *jumptables_data;
size_t jumptables_data_sz;
+ struct cleanup_raw_rec *cleanup_recs;
+ size_t cleanup_rec_cnt;
+
struct {
struct bpf_program *prog;
unsigned int sym_off;
@@ -819,7 +849,10 @@ static void bpf_program__exit(struct bpf_program *prog)
zfree(&prog->sec_name);
zfree(&prog->insns);
zfree(&prog->reloc_desc);
+ zfree(&prog->cleanup_info);
+ prog->cleanup_info_rec_size = 0;
+ prog->cleanup_info_cnt = 0;
prog->nr_reloc = 0;
prog->insns_cnt = 0;
prog->sec_idx = -1;
@@ -1561,6 +1594,7 @@ static struct bpf_object *bpf_object__new(const char *path,
obj->efile.obj_buf = obj_buf;
obj->efile.obj_buf_sz = obj_buf_sz;
obj->efile.btf_maps_shndx = -1;
+ obj->efile.cleanup_shndx = -1;
obj->kconfig_map_idx = -1;
obj->arena_map_idx = -1;
@@ -4047,6 +4081,9 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
sec_desc->shdr = sh;
sec_desc->data = data;
obj->efile.has_st_ops = true;
+ } else if (strcmp(name, CLEANUP_SEC) == 0) {
+ obj->efile.cleanup_data = data;
+ obj->efile.cleanup_shndx = idx;
} else if (strcmp(name, ARENA_SEC) == 0) {
obj->efile.arena_data = data;
obj->efile.arena_data_shndx = idx;
@@ -4074,6 +4111,7 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
strcmp(name, ".rel" STRUCT_OPS_LINK_SEC) &&
strcmp(name, ".rel?" STRUCT_OPS_SEC) &&
strcmp(name, ".rel?" STRUCT_OPS_LINK_SEC) &&
+ strcmp(name, ".rel" CLEANUP_SEC) &&
strcmp(name, ".rel" MAPS_ELF_SEC)) {
pr_info("elf: skipping relo section(%d) %s for section(%d) %s\n",
idx, name, targ_sec_idx,
@@ -4854,6 +4892,235 @@ static struct bpf_program *find_prog_by_sec_insn(const struct bpf_object *obj,
return NULL;
}
+static int bpf_object__init_cleanup_info(struct bpf_object *obj)
+{
+ Elf_Data *data = obj->efile.cleanup_data;
+ Elf_Data *relo = NULL;
+ size_t i, nrels, nslots, nrecs;
+ struct cleanup_raw_rec *recs;
+ int *slot_sec, ret = 0;
+ size_t *slot_val;
+ const __u32 *vals;
+ bool native;
+
+ if (!data || obj->efile.cleanup_shndx < 0)
+ return 0;
+
+ native = is_native_endianness(obj);
+
+ for (i = 0; i < obj->efile.sec_cnt; i++) {
+ struct elf_sec_desc *sd = &obj->efile.secs[i];
+
+ if (sd->sec_type == SEC_RELO && sd->shdr &&
+ sd->shdr->sh_info == (Elf64_Word)obj->efile.cleanup_shndx) {
+ relo = sd->data;
+ break;
+ }
+ }
+ if (!relo) {
+ pr_warn("%s present without relocations\n", CLEANUP_SEC);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ if (data->d_size % sizeof(struct bpf_cleanup_info)) {
+ pr_warn("%s size %zu is not a multiple of the record size %zu\n",
+ CLEANUP_SEC, data->d_size, sizeof(struct bpf_cleanup_info));
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ vals = data->d_buf;
+ nslots = data->d_size / sizeof(__u32);
+ nrecs = data->d_size / sizeof(struct bpf_cleanup_info);
+
+ slot_sec = calloc(nslots, sizeof(*slot_sec));
+ slot_val = calloc(nslots, sizeof(*slot_val));
+ recs = calloc(nrecs ?: 1, sizeof(*recs));
+ if (!slot_sec || !slot_val || !recs) {
+ ret = -ENOMEM;
+ goto out;
+ }
+ for (i = 0; i < nslots; i++)
+ slot_sec[i] = -1;
+
+ /* One relocation per 4-byte field, naming the section it points into. */
+ nrels = relo->d_size / sizeof(Elf64_Rel);
+ for (i = 0; i < nrels; i++) {
+ Elf64_Rel *rel = elf_rel_by_idx(relo, i);
+ Elf64_Sym *sym = elf_sym_by_idx(obj, ELF64_R_SYM(rel->r_info));
+ size_t type = ELF64_R_TYPE(rel->r_info);
+ size_t slot = rel->r_offset / sizeof(__u32);
+
+ if (type != R_BPF_64_NODYLD32 && type != R_BPF_64_ABS32) {
+ pr_warn("%s: relocation %zu has unexpected type %zu\n",
+ CLEANUP_SEC, i, type);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ if (!sym || slot >= nslots || rel->r_offset % sizeof(__u32)) {
+ pr_warn("%s: bad relocation %zu\n", CLEANUP_SEC, i);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ slot_sec[slot] = sym->st_shndx;
+ /* The addend lives in the section data, which libelf leaves in
+ * the object's byte order; a non-section symbol additionally
+ * contributes its own value.
+ */
+ slot_val[slot] = (native ? vals[slot] : bswap_32(vals[slot])) +
+ sym->st_value;
+ }
+
+ for (i = 0; i < nslots; i++) {
+ struct cleanup_raw_rec *rec = &recs[i / CLEANUP_REC_FIELDS];
+ size_t field = i % CLEANUP_REC_FIELDS;
+
+ if (slot_sec[i] < 0) {
+ pr_warn("%s: field %zu has no relocation\n", CLEANUP_SEC, i);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ if (slot_val[i] % BPF_INSN_SZ) {
+ pr_warn("%s: field %zu offset %zu is not instruction aligned\n",
+ CLEANUP_SEC, i, slot_val[i]);
+ ret = -LIBBPF_ERRNO__FORMAT;
+ goto out;
+ }
+ rec->sec_idx[field] = slot_sec[i];
+ rec->insn_idx[field] = slot_val[i] / BPF_INSN_SZ;
+ }
+
+ obj->cleanup_recs = recs;
+ obj->cleanup_rec_cnt = nrecs;
+ recs = NULL;
+out:
+ free(recs);
+ free(slot_val);
+ free(slot_sec);
+ return ret;
+}
+
+static int cmp_cleanup_info(const void *a, const void *b)
+{
+ const struct bpf_cleanup_info *x = a, *y = b;
+
+ if (x->begin_off == y->begin_off)
+ return 0;
+ return x->begin_off < y->begin_off ? -1 : 1;
+}
+
+static int bpf_prog_collect_cleanup_info(struct bpf_object *obj,
+ struct bpf_program *prog)
+{
+ size_t i;
+ int j;
+
+ for (i = 0; i < obj->cleanup_rec_cnt; i++) {
+ struct cleanup_raw_rec *raw = &obj->cleanup_recs[i];
+ struct bpf_program *owner = NULL;
+ struct bpf_cleanup_info ci = {};
+ __u32 fields[CLEANUP_REC_FIELDS];
+ void *tmp;
+
+ if (raw->sec_idx[CLEANUP_REC_BEGIN] == raw->sec_idx[CLEANUP_REC_END] &&
+ raw->insn_idx[CLEANUP_REC_BEGIN] >= raw->insn_idx[CLEANUP_REC_END]) {
+ pr_warn("%s: record %zu is an empty range [%zu,%zu)\n",
+ CLEANUP_SEC, i, raw->insn_idx[CLEANUP_REC_BEGIN],
+ raw->insn_idx[CLEANUP_REC_END]);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ for (j = 0; j < CLEANUP_REC_FIELDS; j++) {
+ size_t idx = raw->insn_idx[j], final;
+ struct bpf_program *p;
+
+ /* The end of a range is exclusive, so it may name the
+ * instruction just past the last one of a function,
+ * which belongs to the next function or to nothing at
+ * all. Ask about the last instruction the range covers,
+ * the way the kernel does.
+ */
+ if (j == CLEANUP_REC_END) {
+ if (!idx) {
+ pr_warn("%s: record %zu ends at instruction 0\n",
+ CLEANUP_SEC, i);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ idx--;
+ }
+
+ p = find_prog_by_sec_insn(obj, raw->sec_idx[j], idx);
+ if (!p) {
+ pr_warn("%s: record %zu field %d is not inside a function\n",
+ CLEANUP_SEC, i, j);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ if (!owner) {
+ owner = p;
+ } else if (owner != p) {
+ pr_warn("%s: record %zu spans functions '%s' and '%s'\n",
+ CLEANUP_SEC, i, owner->name, p->name);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ if (owner == prog) {
+ final = raw->insn_idx[j] - prog->sec_insn_off;
+ } else if (prog_is_subprog(obj, owner) && owner->sub_insn_off) {
+ /* sub_insn_off is where this subprogram was
+ * appended to the main program being relocated;
+ * zero means it is not part of it.
+ */
+ final = owner->sub_insn_off +
+ raw->insn_idx[j] - owner->sec_insn_off;
+ } else {
+ owner = NULL;
+ break;
+ }
+ fields[j] = final;
+ }
+ if (!owner)
+ continue;
+
+ ci.begin_off = fields[CLEANUP_REC_BEGIN];
+ ci.end_off = fields[CLEANUP_REC_END];
+ ci.landing_pad_off = fields[CLEANUP_REC_PAD];
+
+ tmp = libbpf_reallocarray(prog->cleanup_info, prog->cleanup_info_cnt + 1,
+ sizeof(*prog->cleanup_info));
+ if (!tmp)
+ return -ENOMEM;
+ prog->cleanup_info = tmp;
+ prog->cleanup_info_rec_size = sizeof(struct bpf_cleanup_info);
+ prog->cleanup_info[prog->cleanup_info_cnt++] = ci;
+
+ pr_debug("prog '%s': cleanup region [%u,%u) -> landing pad %u\n",
+ prog->name, ci.begin_off, ci.end_off, ci.landing_pad_off);
+ }
+
+ if (!prog->cleanup_info_cnt)
+ return 0;
+
+ if (prog->cleanup_info_cnt > INT32_MAX / sizeof(struct bpf_cleanup_info)) {
+ pr_warn("prog '%s': too many cleanup records: %u\n",
+ prog->name, prog->cleanup_info_cnt);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ qsort(prog->cleanup_info, prog->cleanup_info_cnt,
+ sizeof(*prog->cleanup_info), cmp_cleanup_info);
+ for (i = 1; i < prog->cleanup_info_cnt; i++) {
+ struct bpf_cleanup_info *prev = &prog->cleanup_info[i - 1];
+ struct bpf_cleanup_info *cur = &prog->cleanup_info[i];
+
+ if (cur->begin_off < prev->end_off) {
+ pr_warn("prog '%s': overlapping cleanup regions [%u,%u) and [%u,%u)\n",
+ prog->name, prev->begin_off, prev->end_off,
+ cur->begin_off, cur->end_off);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+ }
+
+ return 0;
+}
+
static int
bpf_object__collect_prog_relos(struct bpf_object *obj, Elf64_Shdr *shdr, Elf_Data *data)
{
@@ -7589,6 +7856,13 @@ static int bpf_object__relocate(struct bpf_object *obj, const char *targ_btf_pat
return err;
}
}
+
+ err = bpf_prog_collect_cleanup_info(obj, prog);
+ if (err) {
+ pr_warn("prog '%s': failed to collect cleanup info: %s\n",
+ prog->name, errstr(err));
+ return err;
+ }
}
for (i = 0; i < obj->nr_programs; i++) {
prog = &obj->programs[i];
@@ -7779,6 +8053,9 @@ static int bpf_object__collect_relos(struct bpf_object *obj)
return -LIBBPF_ERRNO__INTERNAL;
}
+ if (idx == obj->efile.cleanup_shndx)
+ continue;
+
if (obj->efile.secs[idx].sec_type == SEC_ST_OPS)
err = bpf_object__collect_st_ops_relos(obj, shdr, data);
else if (idx == obj->efile.btf_maps_shndx)
@@ -8051,6 +8328,11 @@ static int bpf_object_load_prog(struct bpf_object *obj, struct bpf_program *prog
load_attr.line_info_rec_size = prog->line_info_rec_size;
load_attr.line_info_cnt = prog->line_info_cnt;
}
+ if (prog->cleanup_info_cnt) {
+ load_attr.cleanup_info = prog->cleanup_info;
+ load_attr.cleanup_info_cnt = prog->cleanup_info_cnt;
+ load_attr.cleanup_info_rec_size = prog->cleanup_info_rec_size;
+ }
load_attr.log_level = log_level;
load_attr.prog_flags = prog->prog_flags;
load_attr.fd_array = obj->fd_array;
@@ -8643,6 +8925,7 @@ static struct bpf_object *bpf_object_open(const char *path, const void *obj_buf,
err = err ? : bpf_object__init_maps(obj, opts);
err = err ? : bpf_object_init_progs(obj, opts);
err = err ? : bpf_object__collect_relos(obj);
+ err = err ? : bpf_object__init_cleanup_info(obj);
if (err)
goto out;
@@ -9757,6 +10040,9 @@ void bpf_object__close(struct bpf_object *obj)
zfree(&obj->jumptables_data);
obj->jumptables_data_sz = 0;
+ zfree(&obj->cleanup_recs);
+ obj->cleanup_rec_cnt = 0;
+
for (i = 0; i < obj->jumptable_map_cnt; i++)
close(obj->jumptable_maps[i].fd);
zfree(&obj->jumptable_maps);
@@ -10155,6 +10441,25 @@ int bpf_program__clone(struct bpf_program *prog, const struct bpf_prog_load_opts
attr.line_info_rec_size = info ? info_rec_size : prog->line_info_rec_size;
}
+ /* exception cleanup table */
+ info = OPTS_GET(opts, cleanup_info, NULL);
+ info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
+ info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
+ if (!!info != !!info_cnt || !!info != !!info_rec_size) {
+ pr_warn("prog '%s': cleanup_info, cleanup_info_cnt, and cleanup_info_rec_size must all be specified or all omitted\n",
+ prog->name);
+ return libbpf_err(-EINVAL);
+ }
+ if (info) {
+ attr.cleanup_info = info;
+ attr.cleanup_info_cnt = info_cnt;
+ attr.cleanup_info_rec_size = info_rec_size;
+ } else if (prog->cleanup_info_cnt) {
+ attr.cleanup_info = prog->cleanup_info;
+ attr.cleanup_info_cnt = prog->cleanup_info_cnt;
+ attr.cleanup_info_rec_size = prog->cleanup_info_rec_size;
+ }
+
/* Logging is caller-controlled; no fallback to prog/obj log settings */
attr.log_buf = OPTS_GET(opts, log_buf, NULL);
attr.log_size = OPTS_GET(opts, log_size, 0);
diff --git a/tools/lib/bpf/libbpf_internal.h b/tools/lib/bpf/libbpf_internal.h
index cb4d96233844..3ba6d9090368 100644
--- a/tools/lib/bpf/libbpf_internal.h
+++ b/tools/lib/bpf/libbpf_internal.h
@@ -56,6 +56,9 @@
#ifndef R_BPF_64_ABS32
#define R_BPF_64_ABS32 3
#endif
+#ifndef R_BPF_64_NODYLD32
+#define R_BPF_64_NODYLD32 4
+#endif
#ifndef R_BPF_64_32
#define R_BPF_64_32 10
#endif
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel
2026-09-21 21:01 ` [PATCH bpf-next v4 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
@ 2026-09-21 21:20 ` sashiko-bot
0 siblings, 0 replies; 80+ messages in thread
From: sashiko-bot @ 2026-09-21 21:20 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] bpf: invalid multi-line comment style in libbpf.c
--
commit 640e4ca07507c0fa24643be768ba697d779b3ae5
Author: Yonghong Song <yonghong.song@linux.dev>
Subject: libbpf: Collect .bpf_cleanup records and pass them to the kernel
This commit parses the compiler-emitted .bpf_cleanup section and extracts
the table of cleanup records. It resolves section relocations for each record,
validates alignment, and passes the per-program cleanup information to the
kernel during program load.
> diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
> index 13ae9175e9686..e3e4db59fb5f4 100644
> --- a/tools/lib/bpf/libbpf.c
> +++ b/tools/lib/bpf/libbpf.c
[ ... ]
> @@ -4854,6 +4892,235 @@ static struct bpf_program *find_prog_by_sec_insn(const struct bpf_object *obj,
[ ... ]
> + slot_sec[slot] = sym->st_shndx;
> + /* The addend lives in the section data, which libelf leaves in
> + * the object's byte order; a non-section symbol additionally
> + * contributes its own value.
> + */
> + slot_val[slot] = (native ? vals[slot] : bswap_32(vals[slot])) +
> + sym->st_value;
> + }
[Severity: Low]
This isn't a bug, but does this multi-line comment follow the preferred BPF
subsystem style? I believe the opening /* is typically expected to be on its
own line.
[ ... ]
> + for (j = 0; j < CLEANUP_REC_FIELDS; j++) {
> + size_t idx = raw->insn_idx[j], final;
> + struct bpf_program *p;
> +
> + /* The end of a range is exclusive, so it may name the
> + * instruction just past the last one of a function,
> + * which belongs to the next function or to nothing at
> + * all. Ask about the last instruction the range covers,
> + * the way the kernel does.
> + */
> + if (j == CLEANUP_REC_END) {
[Severity: Low]
This isn't a bug, but should this multi-line comment be formatted with the
opening /* on a separate line to match the rest of the BPF subsystem?
[ ... ]
> + if (owner == prog) {
> + final = raw->insn_idx[j] - prog->sec_insn_off;
> + } else if (prog_is_subprog(obj, owner) && owner->sub_insn_off) {
> + /* sub_insn_off is where this subprogram was
> + * appended to the main program being relocated;
> + * zero means it is not part of it.
> + */
> + final = owner->sub_insn_off +
> + raw->insn_idx[j] - owner->sec_insn_off;
[Severity: Low]
This isn't a bug, but could we adjust this comment to have the opening /* on
its own line for consistency as well?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260921210033.1715000-1-yonghong.song@linux.dev?part=15
^ permalink raw reply [flat|nested] 80+ messages in thread
* [PATCH bpf-next v4 16/20] libbpf: Carry the exception cleanup table through the light skeleton
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (14 preceding siblings ...)
2026-09-21 21:01 ` [PATCH bpf-next v4 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
@ 2026-09-21 21:01 ` Yonghong Song
2026-09-21 21:02 ` [PATCH bpf-next v4 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
` (4 subsequent siblings)
20 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:01 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
A light skeleton does not call bpf_prog_load(). bpf_gen__prog_load() builds
its own union bpf_attr field by field, and the loader program it emits is
what issues BPF_PROG_LOAD when the skeleton runs -- so a program loaded
this way reached the kernel without the table the previous patch collected
for it. The load then fails, because the landing pads are code nothing
reaches and the verifier says so. It says "unreachable insn", which names
neither the skeleton nor the table.
Carry it the way func_info and line_info are carried: the records go into
the loader's blob of bytes, the count and record size into the attr, and a
relocation stores the blob's address into attr.cleanup_info once that
address is known. The attr grows to its new last field, cleanup_info_cnt.
Records are 4-byte fields like the other info blobs, so a cross-endian
build has to swap them too.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/gen_loader.c | 29 +++++++++++++++++++++++++----
tools/lib/bpf/libbpf_internal.h | 7 +++++++
2 files changed, 32 insertions(+), 4 deletions(-)
diff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c
index af3a04f161ac..2345fbdd46f5 100644
--- a/tools/lib/bpf/gen_loader.c
+++ b/tools/lib/bpf/gen_loader.c
@@ -981,13 +981,15 @@ static void cleanup_relos(struct bpf_gen *gen, int insns)
cleanup_core_relo(gen);
}
-/* Convert func, line, and core relo info blobs to target endianness */
+/* Convert func, line, core relo and cleanup info blobs to target endianness */
static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
- int core_relos, struct bpf_prog_load_opts *load_attr)
+ int core_relos, int cleanup_info,
+ struct bpf_prog_load_opts *load_attr)
{
struct bpf_func_info *fi = gen->data_start + func_info;
struct bpf_line_info *li = gen->data_start + line_info;
struct bpf_core_relo *cr = gen->data_start + core_relos;
+ struct bpf_cleanup_info *ci = gen->data_start + cleanup_info;
int i;
for (i = 0; i < load_attr->func_info_cnt; i++)
@@ -998,6 +1000,9 @@ static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
for (i = 0; i < gen->core_relo_cnt; i++)
bpf_core_relo_bswap(cr++);
+
+ for (i = 0; i < load_attr->cleanup_info_cnt; i++)
+ bpf_cleanup_info_bswap(ci++);
}
void bpf_gen__prog_load(struct bpf_gen *gen,
@@ -1011,8 +1016,11 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
load_attr->line_info_rec_size;
int core_relo_tot_sz = gen->core_relo_cnt *
sizeof(struct bpf_core_relo);
+ int cleanup_info_tot_sz = load_attr->cleanup_info_cnt *
+ load_attr->cleanup_info_rec_size;
int prog_load_attr, license_off, insns_off, func_info, line_info, core_relos;
- int attr_size = offsetofend(union bpf_attr, core_relo_rec_size);
+ int attr_size = offsetofend(union bpf_attr, cleanup_info_cnt);
+ int cleanup_info;
union bpf_attr attr;
memset(&attr, 0, attr_size);
@@ -1061,9 +1069,17 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
core_relos, gen->core_relo_cnt,
sizeof(struct bpf_core_relo));
+ attr.cleanup_info_rec_size = tgt_endian(load_attr->cleanup_info_rec_size);
+ attr.cleanup_info_cnt = tgt_endian(load_attr->cleanup_info_cnt);
+ cleanup_info = add_data(gen, load_attr->cleanup_info, cleanup_info_tot_sz);
+ pr_debug("gen: prog_load: cleanup_info: off %d cnt %u rec size %u\n",
+ cleanup_info, load_attr->cleanup_info_cnt,
+ load_attr->cleanup_info_rec_size);
+
/* convert all info blobs to target endianness */
if (gen->swapped_endian && !gen->error)
- info_blob_bswap(gen, func_info, line_info, core_relos, load_attr);
+ info_blob_bswap(gen, func_info, line_info, core_relos, cleanup_info,
+ load_attr);
libbpf_strlcpy(attr.prog_name, prog_name, sizeof(attr.prog_name));
prog_load_attr = add_data(gen, &attr, attr_size);
@@ -1085,6 +1101,11 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
/* populate union bpf_attr with a pointer to core_relos */
emit_rel_store(gen, attr_field(prog_load_attr, core_relos), core_relos);
+ /* populate union bpf_attr with a pointer to cleanup_info, if there is one */
+ if (load_attr->cleanup_info_cnt)
+ emit_rel_store(gen, attr_field(prog_load_attr, cleanup_info),
+ cleanup_info);
+
/* populate union bpf_attr fd_array with a pointer to data where map_fds are saved */
emit_rel_store(gen, attr_field(prog_load_attr, fd_array), gen->fd_array);
diff --git a/tools/lib/bpf/libbpf_internal.h b/tools/lib/bpf/libbpf_internal.h
index 3ba6d9090368..78519f24fb40 100644
--- a/tools/lib/bpf/libbpf_internal.h
+++ b/tools/lib/bpf/libbpf_internal.h
@@ -572,6 +572,13 @@ static inline void bpf_core_relo_bswap(struct bpf_core_relo *i)
i->kind = bswap_32(i->kind);
}
+static inline void bpf_cleanup_info_bswap(struct bpf_cleanup_info *i)
+{
+ i->begin_off = bswap_32(i->begin_off);
+ i->end_off = bswap_32(i->end_off);
+ i->landing_pad_off = bswap_32(i->landing_pad_off);
+}
+
enum btf_field_iter_kind {
BTF_FIELD_ITER_IDS,
BTF_FIELD_ITER_STRS,
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* [PATCH bpf-next v4 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (15 preceding siblings ...)
2026-09-21 21:01 ` [PATCH bpf-next v4 16/20] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
@ 2026-09-21 21:02 ` Yonghong Song
2026-09-21 21:02 ` [PATCH bpf-next v4 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
` (3 subsequent siblings)
20 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:02 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
An object that carries a compiler-emitted exception cleanup table cannot be
linked today. The table's fields are byte offsets into a code section,
materialised by a 32-bit relocation against that section's symbol with the
offset itself as the implicit addend, and the linker rejects both halves of
that: the relocation type is not in the list it accepts, and a relocation
against an STT_SECTION symbol from a non-executable section is an outright
error.
Both spellings of that relocation have to be taken. LLVM emits
R_BPF_64_NODYLD32 for a .long against a section symbol; GNU as emits
R_BPF_64_ABS32, which is what bpf_reloc_type_lookup() maps BFD_RELOC_32 to.
They describe the same value, and the selftests are built with both
compilers.
Taking them is keyed on the relocation type rather than on the section
name, so any non-executable section could reach the new arm, and what it
replaces is an unconditional refusal. Check what the refusal used to make
unnecessary: a relocated section may be SHT_NOBITS, which extend_sec()
leaves with no raw_data, and r_offset is checked for alignment only where
the section holds instructions. Refuse those rather than write through
them.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/lib/bpf/linker.c | 36 +++++++++++++++++++++++++++++++++++-
1 file changed, 35 insertions(+), 1 deletion(-)
diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
index 78f92c39290a..1f512eff5895 100644
--- a/tools/lib/bpf/linker.c
+++ b/tools/lib/bpf/linker.c
@@ -1036,7 +1036,8 @@ static int linker_sanity_check_elf_relos(struct src_obj *obj, struct src_sec *se
size_t sym_type = ELF64_R_TYPE(relo->r_info);
if (sym_type != R_BPF_64_64 && sym_type != R_BPF_64_32 &&
- sym_type != R_BPF_64_ABS64 && sym_type != R_BPF_64_ABS32) {
+ sym_type != R_BPF_64_ABS64 && sym_type != R_BPF_64_ABS32 &&
+ sym_type != R_BPF_64_NODYLD32) {
pr_warn("ELF relo #%d in section #%zu has unexpected type %zu in %s\n",
i, sec->sec_idx, sym_type, obj->filename);
return -EINVAL;
@@ -2258,6 +2259,7 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
if (ELF64_ST_TYPE(src_sym->st_info) == STT_SECTION) {
struct src_sec *sec = &obj->secs[src_sym->st_shndx];
struct bpf_insn *insn;
+ __u32 *val;
if (src_linked_sec->shdr->sh_flags & SHF_EXECINSTR) {
/* calls to the very first static function inside
@@ -2274,6 +2276,38 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
insn->imm += sec->dst_off / sizeof(struct bpf_insn);
else
insn->imm += sec->dst_off;
+ } else if (sym_type == R_BPF_64_NODYLD32 ||
+ sym_type == R_BPF_64_ABS32) {
+ /* Two spellings of the one thing: LLVM
+ * emits NODYLD32 for a .long against a
+ * section symbol, GNU as emits ABS32
+ * (bpf_reloc_type_lookup() maps
+ * BFD_RELOC_32 to it), and the value
+ * they describe is the same.
+ */
+
+ /* Only an executable section has its
+ * r_offset checked, and even there only
+ * for alignment; SHT_NOBITS has no
+ * raw_data at all. Bound it here,
+ * subtracting so it cannot wrap.
+ */
+ if (!dst_linked_sec->raw_data ||
+ dst_linked_sec->sec_sz < (int)sizeof(*val) ||
+ dst_rel->r_offset % sizeof(*val) ||
+ dst_rel->r_offset >
+ (size_t)dst_linked_sec->sec_sz - sizeof(*val)) {
+ pr_warn("ELF relo #%d in section #%zu points outside the data of section '%s' in %s\n",
+ j, src_sec->sec_idx,
+ dst_linked_sec->sec_name,
+ obj->filename);
+ return -EINVAL;
+ }
+ val = dst_linked_sec->raw_data + dst_rel->r_offset;
+ if (linker->swapped_endian)
+ *val = bswap_32(bswap_32(*val) + sec->dst_off);
+ else
+ *val += sec->dst_off;
} else {
pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
return -EINVAL;
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* [PATCH bpf-next v4 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (16 preceding siblings ...)
2026-09-21 21:02 ` [PATCH bpf-next v4 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
@ 2026-09-21 21:02 ` Yonghong Song
2026-09-21 21:22 ` sashiko-bot
2026-09-21 21:56 ` bot+bpf-ci
2026-09-21 21:02 ` [PATCH bpf-next v4 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
` (2 subsequent siblings)
20 siblings, 2 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:02 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
C has no unwinding, so nothing here comes out of the frontend: the frames
that own a resource are written as __naked inline assembly, which spells
out by hand exactly what a frontend emits -- a call site bracketed by two
labels, a landing pad unreachable in the compiler's CFG, and a .bpf_cleanup
record tying them together. The assembler turns ".long <text label>" into
the same R_BPF_64_NODYLD32 relocation the BPF AsmPrinter emits, so libbpf
and the kernel see an object indistinguishable from a compiler-generated
one.
Call chain: entry -> foo1 -> foo1v -> foo2 -> foo3. foo3 holds a
non-preemptible section and throws inside it; foo2 holds an RCU read lock
and has two call sites sharing one pad, one of them its own throw; foo1v is
a void frame whose pad ends in a jump to a resume block placed after an
unrelated block that ends in a plain exit; foo1 owns nothing and gets no
record; entry is the boundary.
There are also some shapes the kernel refuses -- the ones with no correct
answer, and the ones a pad running on the bpf_throw() walker's stack cannot
express -- plus the return-value rejection that delivering an exception at
the boundary of a program type which constrains its return value produces.
The test skips rather than fails where the JIT cannot dispatch a landing
pad at all.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
.../selftests/bpf/exceptions_cleanup.h | 29 +
.../bpf/prog_tests/exceptions_cleanup.c | 121 ++++
.../selftests/bpf/progs/exceptions_cleanup.c | 157 +++++
.../bpf/progs/exceptions_cleanup_fail.c | 662 ++++++++++++++++++
4 files changed, 969 insertions(+)
create mode 100644 tools/testing/selftests/bpf/exceptions_cleanup.h
create mode 100644 tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
new file mode 100644
index 000000000000..630d2e207119
--- /dev/null
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -0,0 +1,29 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef __EXCEPTIONS_CLEANUP_H__
+#define __EXCEPTIONS_CLEANUP_H__
+
+#define THROW_COOKIE 0x100
+
+/* progs/exceptions_cleanup.c: one bit per frame that reports it ran. */
+#define RAN_FOO3_PREEMPT 0x1
+#define RAN_FOO2_RCU 0x2
+#define RAN_FOO1V_PREEMPT 0x4
+#define RAN_FOO2_DROP 0x8
+#define RAN_BUMP 0x10
+
+#define CLEANUP_REC(begin, end, landing_pad) \
+ ".pushsection .bpf_cleanup,\"a\",@progbits;" \
+ ".long " begin ";" \
+ ".long " end ";" \
+ ".long " landing_pad ";" \
+ ".popsection;"
+
+/* Set a bit in @pads_ran. */
+#define PAD_RAN(bit) \
+ "r1 = %[pads_ran] ll;" \
+ "r2 = *(u64 *)(r1 + 0);" \
+ "r2 |= " bit ";" \
+ "*(u64 *)(r1 + 0) = r2;"
+
+#endif /* __EXCEPTIONS_CLEANUP_H__ */
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
new file mode 100644
index 000000000000..c932da7cbec1
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -0,0 +1,121 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <test_progs.h>
+#include "exceptions_cleanup.h"
+#include "exceptions_cleanup.skel.h"
+#include "exceptions_cleanup_fail.skel.h"
+
+/* foo3 threw: every frame that has a pad ran it. */
+#define PADS_FOO3_THREW \
+ (RAN_FOO3_PREEMPT | RAN_FOO2_RCU | RAN_FOO1V_PREEMPT | RAN_FOO2_DROP)
+
+/* foo2 threw after foo3 returned normally: foo3's pad must not run. */
+#define PADS_FOO2_THREW \
+ (RAN_FOO2_RCU | RAN_FOO1V_PREEMPT | RAN_FOO2_DROP)
+
+static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
+ __u64 pads)
+{
+ __u64 ctx = 0;
+ int err;
+
+ LIBBPF_OPTS(bpf_test_run_opts, topts,
+ .ctx_in = &ctx,
+ .ctx_size_in = sizeof(ctx),
+ );
+
+ skel->bss->input = input;
+ skel->bss->pads_ran = 0;
+ skel->bss->result = 0;
+
+ err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.entry), &topts);
+ if (!ASSERT_OK(err, "run"))
+ return;
+ ASSERT_EQ(topts.retval, retval, "retval");
+ /* bump() is not a landing pad; it sets its bit on every run. */
+ ASSERT_EQ(skel->bss->pads_ran, pads | RAN_BUMP, "pads_ran");
+}
+
+/* A table with more records than the program has instructions: no valid one
+ * can look like that, and it is refused before the kernel allocates for it.
+ */
+static void test_cleanup_info_cnt(void)
+{
+ struct bpf_insn insns[] = {
+ BPF_MOV64_IMM(BPF_REG_0, 0),
+ BPF_EXIT_INSN(),
+ };
+ struct bpf_cleanup_info rec = {
+ .begin_off = 0,
+ .end_off = 1,
+ .landing_pad_off = 1,
+ };
+ char log[512] = {};
+ LIBBPF_OPTS(bpf_prog_load_opts, opts,
+ .log_buf = log,
+ .log_size = sizeof(log),
+ .log_level = 1,
+ .cleanup_info = &rec,
+ .cleanup_info_cnt = 1 << 20,
+ .cleanup_info_rec_size = sizeof(rec));
+ int fd;
+
+ fd = bpf_prog_load(BPF_PROG_TYPE_SOCKET_FILTER, NULL, "GPL",
+ insns, ARRAY_SIZE(insns), &opts);
+ if (!ASSERT_LT(fd, 0, "load")) {
+ close(fd);
+ return;
+ }
+ /* Turned away on the count, rather than on whatever the records past
+ * the one below happen to hold.
+ */
+ ASSERT_HAS_SUBSTR(log, "cleanup info has 1048576 records for 2 instructions",
+ "log");
+}
+
+void test_exceptions_cleanup(void)
+{
+ char log[8192] = {};
+
+ LIBBPF_OPTS(bpf_object_open_opts, opts,
+ .kernel_log_buf = log,
+ .kernel_log_size = sizeof(log));
+ struct exceptions_cleanup *skel;
+ int err;
+
+ skel = exceptions_cleanup__open_opts(&opts);
+ if (!ASSERT_OK_PTR(skel, "open"))
+ return;
+
+ err = exceptions_cleanup__load(skel);
+ if (err) {
+ if (err == -EOPNOTSUPP &&
+ strstr(log, "exception cleanup needs a JIT that can dispatch landing pads"))
+ test__skip();
+ else if (!ASSERT_OK(err, "load"))
+ fprintf(stderr, "%s", log);
+ exceptions_cleanup__destroy(skel);
+ return;
+ }
+
+ /* No throw: foo3 returns 1 ^ 1 == 0, foo2 adds one, no pad runs. */
+ if (test__start_subtest("no_throw"))
+ run(skel, 1, 1, 0);
+
+ /* foo3 throws; every pad runs and the cookie is delivered at entry. */
+ if (test__start_subtest("throw_from_foo3"))
+ run(skel, 101, THROW_COOKIE, PADS_FOO3_THREW);
+
+ /* foo3 returns 2 ^ 1 == 3, so foo2 throws from its own second region;
+ * foo3's frame is long gone, so its pad must not run.
+ */
+ if (test__start_subtest("throw_from_foo2"))
+ run(skel, 2, THROW_COOKIE, PADS_FOO2_THREW);
+
+ exceptions_cleanup__destroy(skel);
+
+ RUN_TESTS(exceptions_cleanup_fail);
+
+ if (test__start_subtest("cleanup_info_cnt"))
+ test_cleanup_info_cnt();
+}
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
new file mode 100644
index 000000000000..a3a8c14e0db2
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
@@ -0,0 +1,157 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_rcu_read_lock();
+ bpf_rcu_read_unlock();
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume();
+}
+
+__u64 input = 0;
+__u64 pads_ran = 0;
+__u64 result = 0;
+
+static __used __noinline __u64 foo3(__u64 x)
+{
+ bpf_preempt_disable();
+ if (x > 100)
+ asm volatile (
+ "r1 = %[cookie];"
+ "1:" "call bpf_throw;" /* cleanup region */
+ "2:"
+ "goto 3f;"
+ "4:" /* landing pad */
+ "r7 = r0;"
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "3:"
+ CLEANUP_REC("1b", "2b", "4b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_FOO3_PREEMPT),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+ bpf_preempt_enable();
+ return x ^ 1;
+}
+
+__u64 never = 0;
+
+static __used __naked __noinline void drop_glue(void)
+{
+ asm volatile (
+ PAD_RAN("%[ran]")
+ "exit;"
+ :
+ : [ran]"i"(RAN_FOO2_DROP), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 foo2(void)
+{
+ asm volatile (
+ "r6 = r1;"
+ "call bpf_rcu_read_lock;"
+ "r1 = r6;"
+"1:" "call foo3;" /* cleanup region #1 */
+"2:"
+ "r6 = r0;"
+ "if r6 == 0 goto 5f;"
+ "r1 = %[cookie];"
+"3:" "call bpf_throw;" /* cleanup region #2 */
+"4:"
+ "r0 = 0;"
+ "exit;"
+"5:"
+ "call bpf_rcu_read_unlock;"
+ "r0 = r6;"
+ "r0 += 1;"
+ "exit;"
+"6:" /* landing pad, shared by both regions */
+ "call drop_glue;"
+ "call bpf_rcu_read_unlock;"
+ PAD_RAN("%[ran_rcu]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "6b")
+ CLEANUP_REC("3b", "4b", "6b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran_rcu]"i"(RAN_FOO2_RCU),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+static __used __naked __noinline void foo1v(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call foo2;" /* cleanup region */
+"2:"
+ "r6 = r0;"
+ "call bpf_preempt_enable;"
+ "r1 = %[result] ll;"
+ "*(u64 *)(r1 + 0) = r6;"
+ "goto 7f;"
+"8:" /* landing pad */
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "goto 9f;"
+"7:" /* the frame's own exit block */
+ "r0 = 0;"
+ "exit;"
+"9:" /* shared resume block */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "8b")
+ :
+ : [ran]"i"(RAN_FOO1V_PREEMPT), __imm_addr(input),
+ __imm_addr(result), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/* Called from foo1, which owns nothing and has no cleanup record: a call an
+ * exception may leave with no landing pad to run on the way. The throw is
+ * never taken -- @never is a global, so the verifier cannot prune it -- and
+ * the bit says the ordinary return path ran.
+ */
+static __used __naked __noinline void bump(void)
+{
+ asm volatile (
+ PAD_RAN("%[ran]")
+ "r1 = %[never] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 == 0 goto 1f;"
+ "r1 = 0;"
+ "call bpf_throw;"
+"1:"
+ "exit;" /* r0 deliberately left alone */
+ :
+ : [ran]"i"(RAN_BUMP), __imm_addr(never), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+__noinline __u64 foo1(void)
+{
+ bump();
+ foo1v();
+ return result;
+}
+
+SEC("syscall")
+int entry(void *ctx)
+{
+ return foo1();
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
new file mode 100644
index 000000000000..db6ac7d8bd6d
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
@@ -0,0 +1,662 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_experimental.h"
+#include "bpf_misc.h"
+#include "../test_kmods/bpf_testmod_kfunc.h"
+#include "exceptions_cleanup.h"
+
+__u64 input = 0;
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume();
+}
+
+/* 1. A subprogram that may unwind, also used as a helper callback:
+ * bpf_loop()'s own kernel frame would end the walk before it found a
+ * boundary.
+ */
+static int throwing_cb(__u32 idx, void *ctx)
+{
+ bpf_throw(0xbad);
+ return 0;
+}
+
+static __used __naked __noinline __u64 cb_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call throwing_cb;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("may unwind and is used as a callback")
+int callback_may_unwind(void *ctx)
+{
+ bpf_loop(1, throwing_cb, NULL, 0);
+ return cb_frame();
+}
+
+/* 2. A landing pad that reaches both an unwind resume and a plain exit, so
+ * nothing says whether it is a cleanup pad or a catch pad.
+ */
+static __used __naked __noinline __u64 inner_throw(void)
+{
+ asm volatile (
+ "r1 = 1;"
+ "call bpf_throw;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 ambiguous_pad_frame(void)
+{
+ asm volatile (
+ "r6 = r1;"
+ "call bpf_preempt_disable;"
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad: two ways out */
+ "call bpf_preempt_enable;"
+ "if r6 > 10 goto 4f;"
+ "call bpf_unwind_resume;"
+ "exit;"
+"4:"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("reaches both bpf_unwind_resume() and a plain exit")
+int ambiguous_landing_pad(void *ctx)
+{
+ return ambiguous_pad_frame();
+}
+
+/* 3. A throw from inside a landing pad: a second walk over the frames the
+ * first one is in the middle of discarding.
+ */
+static __used __naked __noinline __u64 throw_in_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad that throws again */
+ "call bpf_preempt_enable;"
+ "r1 = 2;"
+ "call bpf_throw;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("can throw while an exception is in flight")
+int throw_from_landing_pad(void *ctx)
+{
+ return throw_in_pad_frame();
+}
+
+/* 4. A cleanup table in a program that also installs an exception callback,
+ * two different answers to what runs on the way out.
+ */
+__noinline int unused_exc_cb(u64 cookie)
+{
+ return 0;
+}
+
+static __used __naked __noinline __u64 cb_and_table_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 9;"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__exception_cb(unused_exc_cb)
+__failure __msg("cannot be combined with an exception callback")
+int table_with_exception_cb(void *ctx)
+{
+ return cb_and_table_frame();
+}
+
+__u64 never;
+
+/* 5. A landing pad that calls a subprogram which can throw. Not case 3: the
+ * throw is in another subprogram, so what catches it is the walk of the pad's
+ * body, off subprog_info.might_throw.
+ */
+static __used __noinline void pad_callee_that_throws(void)
+{
+ if (never)
+ bpf_throw(0);
+}
+
+static __used __naked __noinline __u64 pad_calls_thrower_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 11;"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call pad_callee_that_throws;" /* ...which can throw: refused */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("which can throw while an exception is in flight")
+int pad_calls_thrower(void *ctx)
+{
+ return pad_calls_thrower_frame();
+}
+
+/* 6. A catch pad: it ends in a plain exit rather than a resume, and a walker
+ * that calls pads as subroutines cannot hand a frame back its own execution.
+ */
+static __used __naked __noinline __u64 catch_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 12;"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* catch pad: no resume, it stops here */
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not supported yet, only cleanup pads that resume")
+int catch_landing_pad(void *ctx)
+{
+ return catch_pad_frame();
+}
+
+/* 7. An exception reaching the boundary of a program type that constrains its
+ * return value: delivery makes the cookie that return value, and fentry has
+ * to return 0.
+ */
+static __used __naked __noinline __u64 boundary_throw_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 7;"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?fentry/bpf_fentry_test1")
+__failure __msg("the register R1 has smin=7 smax=7 should have been in [0, 0]")
+int boundary_delivers(void *ctx)
+{
+ return boundary_throw_frame();
+}
+
+/* 8. A bpf_unwind_resume() outside any landing pad. Both JITs lower it as the
+ * way back out of a pad, which in ordinary code leaves a live frame standing
+ * with its epilogue skipped.
+ */
+static __used __naked __noinline __u64 stray_resume_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 13;"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not in an exception cleanup landing pad")
+int resume_outside_pad(void *ctx)
+{
+ /* Never taken, but reachable, which is all the verifier needs. */
+ if (never)
+ bpf_unwind_resume();
+ return stray_resume_frame();
+}
+
+/* 9. A bpf_unwind_resume() in a subprogram a landing pad calls. The verifier's
+ * walk cannot tell it from a resume in the pad itself -- an exception is in
+ * flight either way -- so the rule is static: a resume sits in a pad body.
+ */
+static __used __naked __noinline void resume_in_callee(void)
+{
+ asm volatile (
+ "call bpf_unwind_resume;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 pad_calls_resumer_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = 14;"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "call resume_in_callee;" /* ...which resumes: refused */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not in an exception cleanup landing pad")
+int resume_in_pad_callee(void *ctx)
+{
+ return pad_calls_resumer_frame();
+}
+
+/* 10. A bpf_unwind_resume() in a program carrying no cleanup table, where
+ * that static rule does not run at all. do_check() refuses it on the state
+ * not unwinding, and has to: the JITs lower every one of these the same way.
+ */
+static __used __naked __noinline __u64 no_table_resume_frame(void)
+{
+ asm volatile (
+ "call bpf_unwind_resume;"
+ "r0 = 0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("reached without an exception in flight")
+int resume_without_table(void *ctx)
+{
+ return no_table_resume_frame();
+}
+
+/* 11. A landing pad that is itself a covered call site, so an exception out
+ * of it would have nowhere to go. Hand-written only: LLVM sinks a function's
+ * pads past every range it emits.
+ */
+static __used __naked __noinline __u64 nested_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_throw;" /* first cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* first pad, second region's call */
+ "call bpf_preempt_enable;"
+"4:"
+ "call bpf_unwind_resume;"
+ "exit;"
+"5:" /* second pad */
+ "call bpf_preempt_enable;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ CLEANUP_REC("3b", "4b", "5b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is inside the call-site range of")
+int nested_landing_pad(void *ctx)
+{
+ return nested_pad_frame();
+}
+
+/* 12. A tail call in a landing pad: it unwinds the prologue off the stack
+ * pointer, which in a pad is the walker's.
+ */
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(__u32));
+ __uint(value_size, sizeof(__u32));
+} tc_map SEC(".maps");
+
+static __used __naked __noinline __u64 tail_call_pad_frame(void)
+{
+ asm volatile (
+ "r6 = r1;"
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r1 = r6;"
+ "r2 = %[tc_map] ll;"
+ "r3 = 0;"
+ "call %[bpf_tail_call];"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm(bpf_tail_call), __imm_addr(tc_map)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is in an exception cleanup landing pad")
+int tail_call_in_pad(void *ctx)
+{
+ return tail_call_pad_frame();
+}
+
+#if defined(__BPF_FEATURE_STACK_ARGUMENT)
+
+/* 13. A call that passes an argument on the stack, in a landing pad: the
+ * outgoing area the callee reads is not the one the caller wrote, the frame
+ * being the unwinding one and the stack pointer the walker's.
+ */
+static __used __noinline __u64 six_args(__u64 a, __u64 b, __u64 c, __u64 d,
+ __u64 e, __u64 f)
+{
+ return a + b + c + d + e + f;
+}
+
+static __used __naked __noinline __u64 stack_arg_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "r1 = 1;"
+ "r2 = 2;"
+ "r3 = 3;"
+ "r4 = 4;"
+ "r5 = 5;"
+ "*(u64 *)(r11 - 8) = 6;" /* the sixth argument */
+ "call six_args;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("on-stack call argument in an exception cleanup landing pad")
+int stack_arg_in_pad(void *ctx)
+{
+ return stack_arg_pad_frame();
+}
+
+/* 14. The same, reached the other way: a kfunc whose by-value argument runs
+ * past the five argument registers, where the JIT fills the outgoing area and
+ * the rule above has no store to catch. The C call gives the extern its BTF.
+ */
+static __used __noinline void __nofit_btf_anchor(void)
+{
+ struct prog_test_pair_arg s = {};
+
+ bpf_kfunc_call_test_pair_arg_nofit(1, 2, 3, 4, s);
+}
+
+static __used __naked __noinline __u64 kfunc_arg_pad_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "r1 = 1;"
+ "r2 = 2;"
+ "r3 = 3;"
+ "r4 = 4;"
+ "r5 = 5;"
+ "call bpf_kfunc_call_test_pair_arg_nofit;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("on-stack call argument in an exception cleanup landing pad")
+int kfunc_stack_arg_in_pad(void *ctx)
+{
+ return kfunc_arg_pad_frame();
+}
+
+#endif /* __BPF_FEATURE_STACK_ARGUMENT */
+
+/* 15. A landing pad entered by ordinary control flow, arriving with none of
+ * what the walker sets up. Nothing static sees it -- the resume really is in
+ * a pad body -- so do_check() refuses it on the state not unwinding.
+ */
+static __used __naked __noinline __u64 jump_into_pad_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "if r6 > 7 goto 4f;" /* an ordinary branch into the pad */
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+"4:" /* ... and its second instruction */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("reached without an exception in flight")
+int jump_into_pad(void *ctx)
+{
+ return jump_into_pad_frame();
+}
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)
+
+/* 16. A landing pad that reaches an indirect jump, which cannot be told from
+ * a catch pad. SEC("socket") because a jump table entry is an offset from the
+ * program's section symbol, and "?syscall" is not a name assembly can use.
+ */
+static __used __naked __noinline void gotox_thrower(void)
+{
+ asm volatile (
+ "r1 = 15;"
+ "call bpf_throw;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("reaches an indirect jump")
+__naked void gotox_in_pad(void)
+{
+ asm volatile (
+ ".pushsection .jumptables,\"\",@progbits;"
+"jt0_%=:"
+ ".quad l0_%= - socket;"
+ ".quad l1_%= - socket;"
+ ".size jt0_%=, 16;"
+ ".global jt0_%=;"
+ ".popsection;"
+
+"1:" "call gotox_thrower;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r1 = jt0_%= ll;"
+ "r1 += 8;"
+ "r2 = *(u64 *)(r1 + 0);"
+ /* gotox r2, as a raw insn: the mnemonic only reached the LLVM
+ * assembler in llvm 22, and BPF_RAW_INSN() needs <linux/bpf.h>, which
+ * vmlinux.h rules out.
+ */
+ ".8byte 0x20d;"
+"l0_%=:"
+ "call bpf_unwind_resume;"
+ "exit;"
+"l1_%=:"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+#endif /* x86 || arm64 */
+
+/* 17. A BPF_LD_[ABS|IND] in a landing pad. A failed load leaves the
+ * subprogram through the hidden "r0 = 0; exit" gen_ld_abs() patches in, and
+ * that exit is the epilogue, which unwinds a stack the pad does not own.
+ */
+static __used __naked __noinline __u64 ld_abs_pad_frame(void)
+{
+ asm volatile (
+ "r6 = r1;" /* the skb BPF_LD_ABS reads */
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ /* r0 = *(u32 *)skb[0]. Spelled as a raw insn because BPF_LD_ABS()
+ * needs <linux/filter.h>, which this file cannot have -- vmlinux.h
+ * already defines the uapi enums.
+ */
+ ".8byte 0x20;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?tc")
+__failure __msg("is in an exception cleanup landing pad")
+__naked void ld_abs_in_pad(void)
+{
+ asm volatile (
+ "call ld_abs_pad_frame;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* 18. A bpf_unwind_resume() in a subprogram a landing pad called, which has a
+ * pad of its own and reached it by ordinary control flow. Not case 9: an
+ * exception is in flight, but this is not the frame whose pad the walker ran.
+ */
+static __used __naked __noinline __u64 own_pad_callee(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 == 0 goto 3f;" /* an ordinary branch into its pad */
+"1:" "call bpf_preempt_disable;" /* cleanup region: nothing that throws */
+ "call bpf_preempt_enable;"
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* its landing pad */
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm_addr(input)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 pad_calls_own_pad_frame(void)
+{
+ asm volatile (
+"1:" "call inner_throw;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad, which calls the above */
+ "call own_pad_callee;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ ::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is in frame 2, not frame 1 whose landing pad the exception entered")
+int resume_in_callee_own_pad(void *ctx)
+{
+ return pad_calls_own_pad_frame();
+}
+
+char _license[] SEC("license") = "GPL";
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test
2026-09-21 21:02 ` [PATCH bpf-next v4 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
@ 2026-09-21 21:22 ` sashiko-bot
2026-09-22 5:26 ` Yonghong Song
2026-09-21 21:56 ` bot+bpf-ci
1 sibling, 1 reply; 80+ messages in thread
From: sashiko-bot @ 2026-09-21 21:22 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] bpf: endianness bug in raw BPF instruction emission
- [Low] bpf: BPF Comment Style Violation
--
commit 322ad0f798ee8015c8f51c953009198fc70b2b9b
Author: Yonghong Song <yonghong.song@linux.dev>
selftests/bpf: Add an end-to-end .bpf_cleanup exception test
This commit adds an end-to-end test for exception cleanup using .bpf_cleanup
records. It defines inline assembly constructs mimicking compiler output to
evaluate how the kernel handles and optionally rejects various landing pad
combinations and malformed paths.
> diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> new file mode 100644
> index 0000000000000..c932da7cbec16
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> @@ -0,0 +1,121 @@
[ ... ]
> +/* A table with more records than the program has instructions: no valid one
> + * can look like that, and it is refused before the kernel allocates for it.
> + */
[Severity: Low]
Does this code adhere to the BPF comment style guidelines?
The BPF subsystem requires multi-line comments to have the opening '/*'
on its own line, with the comment text starting on the next line. In
test_cleanup_info_cnt(), the text starts on the same line as the opening.
> +static void test_cleanup_info_cnt(void)
> +{
[ ... ]
> + /* Turned away on the count, rather than on whatever the records past
> + * the one below happen to hold.
> + */
[Severity: Low]
This isn't a bug, but there's another multi-line comment in
test_cleanup_info_cnt() starting on the same line as the opening '/*'.
> + ASSERT_HAS_SUBSTR(log, "cleanup info has 1048576 records for 2 instructions",
> + "log");
> +}
> +
> +void test_exceptions_cleanup(void)
> +{
[ ... ]
> + /* foo3 returns 2 ^ 1 == 3, so foo2 throws from its own second region;
> + * foo3's frame is long gone, so its pad must not run.
> + */
[Severity: Low]
This isn't a bug, but this multi-line comment in test_exceptions_cleanup()
also starts the text on the same line as the opening '/*'.
> + if (test__start_subtest("throw_from_foo2"))
> + run(skel, 2, THROW_COOKIE, PADS_FOO2_THREW);
> +
> + exceptions_cleanup__destroy(skel);
[ ... ]
> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
> new file mode 100644
> index 0000000000000..a3a8c14e0db28
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
> @@ -0,0 +1,157 @@
[ ... ]
> +/* Called from foo1, which owns nothing and has no cleanup record: a call an
> + * exception may leave with no landing pad to run on the way. The throw is
> + * never taken -- @never is a global, so the verifier cannot prune it -- and
> + * the bit says the ordinary return path ran.
> + */
[Severity: Low]
This isn't a bug, but this multi-line comment before bump() in
exceptions_cleanup.c also violates the BPF comment style requirement.
> +static __used __naked __noinline void bump(void)
> +{
[ ... ]
> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
> new file mode 100644
> index 0000000000000..db6ac7d8bd6de
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
> @@ -0,0 +1,662 @@
[ ... ]
> +/* 1. A subprogram that may unwind, also used as a helper callback:
> + * bpf_loop()'s own kernel frame would end the walk before it found a
> + * boundary.
> + */
[Severity: Low]
This isn't a bug, but this multi-line comment before throwing_cb() in
exceptions_cleanup_fail.c starts text on the same line as the opening '/*'.
> +static int throwing_cb(__u32 idx, void *ctx)
> +{
[ ... ]
> + /* gotox r2, as a raw insn: the mnemonic only reached the LLVM
> + * assembler in llvm 22, and BPF_RAW_INSN() needs <linux/bpf.h>, which
> + * vmlinux.h rules out.
> + */
> + ".8byte 0x20d;"
[Severity: Medium]
Will this raw instruction emission work correctly on big-endian architectures?
The inline assembly in gotox_in_pad() emits a 64-bit integer literal. On a
little-endian target, 0x20d places the correct opcode bytes at the start
(offset 0). However, on a big-endian target, the 64-bit value is emitted with
leading zeros, placing 0x00 in the code byte. This corrupts the instruction
stream and will cause the BPF verifier to reject it.
> +"l0_%=:"
> + "call bpf_unwind_resume;"
> + "exit;"
[ ... ]
> + /* r0 = *(u32 *)skb[0]. Spelled as a raw insn because BPF_LD_ABS()
> + * needs <linux/filter.h>, which this file cannot have -- vmlinux.h
> + * already defines the uapi enums.
> + */
> + ".8byte 0x20;"
[Severity: Medium]
Similarly here in ld_abs_in_pad(), does emitting 0x20 as an 8-byte literal
corrupt the instruction stream on big-endian architectures?
On big-endian targets, this will evaluate with 0x00 at the code byte instead
of BPF_LD_ABS (0x20), causing malformed instructions to be passed to the
verifier.
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260921210033.1715000-1-yonghong.song@linux.dev?part=18
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test
2026-09-21 21:22 ` sashiko-bot
@ 2026-09-22 5:26 ` Yonghong Song
0 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-22 5:26 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On 9/21/26 2:22 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
> - [Medium] bpf: endianness bug in raw BPF instruction emission
> - [Low] bpf: BPF Comment Style Violation
> --
>
> commit 322ad0f798ee8015c8f51c953009198fc70b2b9b
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> selftests/bpf: Add an end-to-end .bpf_cleanup exception test
>
> This commit adds an end-to-end test for exception cleanup using .bpf_cleanup
> records. It defines inline assembly constructs mimicking compiler output to
> evaluate how the kernel handles and optionally rejects various landing pad
> combinations and malformed paths.
>
>> diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
>> new file mode 100644
>> index 0000000000000..c932da7cbec16
>> --- /dev/null
>> +++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
>> @@ -0,0 +1,121 @@
> [ ... ]
>> +/* A table with more records than the program has instructions: no valid one
>> + * can look like that, and it is refused before the kernel allocates for it.
>> + */
> [Severity: Low]
> Does this code adhere to the BPF comment style guidelines?
>
> The BPF subsystem requires multi-line comments to have the opening '/*'
> on its own line, with the comment text starting on the next line. In
> test_cleanup_info_cnt(), the text starts on the same line as the opening.
>
>
[...]
> [ ... ]
>> + /* gotox r2, as a raw insn: the mnemonic only reached the LLVM
>> + * assembler in llvm 22, and BPF_RAW_INSN() needs <linux/bpf.h>, which
>> + * vmlinux.h rules out.
>> + */
>> + ".8byte 0x20d;"
> [Severity: Medium]
> Will this raw instruction emission work correctly on big-endian architectures?
>
> The inline assembly in gotox_in_pad() emits a 64-bit integer literal. On a
> little-endian target, 0x20d places the correct opcode bytes at the start
> (offset 0). However, on a big-endian target, the 64-bit value is emitted with
> leading zeros, placing 0x00 in the code byte. This corrupts the instruction
> stream and will cause the BPF verifier to reject it.
>
>> +"l0_%=:"
>> + "call bpf_unwind_resume;"
>> + "exit;"
> [ ... ]
>> + /* r0 = *(u32 *)skb[0]. Spelled as a raw insn because BPF_LD_ABS()
>> + * needs <linux/filter.h>, which this file cannot have -- vmlinux.h
>> + * already defines the uapi enums.
>> + */
>> + ".8byte 0x20;"
> [Severity: Medium]
> Similarly here in ld_abs_in_pad(), does emitting 0x20 as an 8-byte literal
> corrupt the instruction stream on big-endian architectures?
>
> On big-endian targets, this will evaluate with 0x00 at the code byte instead
> of BPF_LD_ABS (0x20), causing malformed instructions to be passed to the
> verifier.
Will fix endian issue.
>
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test
2026-09-21 21:02 ` [PATCH bpf-next v4 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
2026-09-21 21:22 ` sashiko-bot
@ 2026-09-21 21:56 ` bot+bpf-ci
1 sibling, 0 replies; 80+ messages in thread
From: bot+bpf-ci @ 2026-09-21 21:56 UTC (permalink / raw)
To: yonghong.song, bpf
Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 5623 bytes --]
> diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> new file mode 100644
> index 000000000000..c932da7cbec1
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
[ ... ]
> @@ -0,0 +1,121 @@
[ ... ]
> + err = exceptions_cleanup__load(skel);
> + if (err) {
> + if (err == -EOPNOTSUPP &&
> + strstr(log, "exception cleanup needs a JIT that can dispatch landing pads"))
> + test__skip();
A subsystem pattern flags this as potentially concerning: the load failure
is correctly turned into a skip rather than a failure, but the skip carries
no reason. test__skip() is called with nothing printed, so the only output
is the harness' bare '#N exceptions_cleanup:SKIP' line. Is there a reason
to omit the skip message here?
A reader of a CI log cannot tell whether the test was skipped because the
JIT cannot dispatch landing pads, because the test was denylisted, or for
some other reason. Roughly a third of the existing test__skip() sites in
prog_tests/ print a reason, e.g. fprintf(stderr, "%s:SKIP: JIT cannot
dispatch landing pads\n", __func__) alongside test__skip(). (Note that
test_progs' test__skip() has no message parameter, so the reason has to be
printed separately.)
The skip decision itself is correct and precisely gated on the kernel's
message from kernel/bpf/exception.c:440, so this is a diagnosability nit
rather than a functional defect.
> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
> new file mode 100644
> index 000000000000..a3a8c14e0db2
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
[ ... ]
> +static __used __noinline __u64 foo3(__u64 x)
> +{
> + bpf_preempt_disable();
> + if (x > 100)
> + asm volatile (
> + "r1 = %[cookie];"
> + "1:" "call bpf_throw;" /* cleanup region */
> + "2:"
> + "goto 3f;"
> + "4:" /* landing pad */
> + "r7 = r0;"
> + "call bpf_preempt_enable;"
> + PAD_RAN("%[ran]")
> + "r1 = r7;"
> + "call bpf_unwind_resume;"
This isn't a bug, but is the r7/r1 shuffle in foo3()'s pad meant to
demonstrate that callee-saved state survives a pad-internal call, or can
those two instructions be dropped so all three pads call bpf_unwind_resume()
the same way?
The pad saves r0 into the callee-saved r7 across the bpf_preempt_enable()
call and restores it into r1 immediately before 'call bpf_unwind_resume'.
bpf_unwind_resume() is declared as taking no arguments, and at pad entry r0
is not the throw cookie but the fixed constant the kernel writes there
(BPF_PAD_ENTRY_R0 == 1), so both instructions appear to be dead stores.
The two other pads in the same file (foo2 label 6 and foo1v label 8) call
bpf_unwind_resume without any register setup.
> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
> new file mode 100644
> index 000000000000..db6ac7d8bd6d
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
[ ... ]
> +#if defined(__BPF_FEATURE_STACK_ARGUMENT)
[ ... ]
> +static __used __naked __noinline __u64 stack_arg_pad_frame(void)
> +{
> + asm volatile (
> + "call bpf_preempt_disable;"
> +"1:" "call inner_throw;" /* cleanup region */
> +"2:"
> + "call bpf_preempt_enable;"
> + "r0 = 0;"
> + "exit;"
> +"3:" /* landing pad */
> + "call bpf_preempt_enable;"
> + "r1 = 1;"
> + "r2 = 2;"
> + "r3 = 3;"
> + "r4 = 4;"
> + "r5 = 5;"
> + "*(u64 *)(r11 - 8) = 6;" /* the sixth argument */
> + "call six_args;"
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + ::: __clobber_all);
> +}
[ ... ]
> +static __used __naked __noinline __u64 kfunc_arg_pad_frame(void)
> +{
> + asm volatile (
> + "call bpf_preempt_disable;"
> + "r1 = 11;"
> +"1:" "call bpf_throw;" /* cleanup region */
> +"2:"
> + "r0 = 0;"
> + "exit;"
> +"3:" /* landing pad */
> + "call bpf_preempt_enable;"
> + "r1 = 1;"
> + "r2 = 2;"
> + "r3 = 3;"
> + "r4 = 4;"
> + "r5 = 5;"
> + "call bpf_kfunc_call_test_pair_arg_nofit;"
> + "call bpf_unwind_resume;"
> + "exit;"
> + CLEANUP_REC("1b", "2b", "3b")
> + ::: __clobber_all);
> +}
This isn't a bug, but since case 14 (kfunc_arg_pad_frame) is rejected on
the kfunc's BTF argument-slot count before do_check() runs, do the r1..r5
assignments there carry any weight, or were they carried over from case 13
(stack_arg_pad_frame)?
kfunc_arg_pad_frame() is instruction-for-instruction identical to
stack_arg_pad_frame() except that the '*(u64 *)(r11 - 8) = 6; call
six_args;' pair is replaced by a single 'call
bpf_kfunc_call_test_pair_arg_nofit;', and both carry the same __msg()
expectation. In case 13 the r1..r5 setup pairs with the r11 store that the
kernel rule actually matches on; in case 14 the rejection comes purely from
the kfunc's BTF argument-slot count, which the verifier computes before
do_check() ever runs, so the five register assignments carry no weight
there.
Would a shared macro for the repeated frame skeleton ('call
bpf_preempt_disable / 1: call inner_throw / 2: call bpf_preempt_enable / r0
= 0 / exit / 3: ...') make the one differing instruction in each case easier
to see? The same six-instruction frame skeleton is repeated verbatim across
roughly eleven __naked frames in the file.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35656368472
^ permalink raw reply [flat|nested] 80+ messages in thread
* [PATCH bpf-next v4 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (17 preceding siblings ...)
2026-09-21 21:02 ` [PATCH bpf-next v4 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
@ 2026-09-21 21:02 ` Yonghong Song
2026-09-21 21:19 ` sashiko-bot
2026-09-21 21:02 ` [PATCH bpf-next v4 20/20] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
2026-09-22 1:08 ` [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Eduard Zingerman
20 siblings, 1 reply; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:02 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
The end-to-end test walks one call chain with a pad in most of its frames.
This adds the shapes that chain does not reach, numbered in the new file in
the order they are listed here:
- what a cleanup table leaves dead
- a callee called from both a covered and an uncovered site
- a pad that reads its frame's callee-saved registers
- a tail-call-reachable callee
- a pad in the main program's own frame
- a tail call that is taken
- an extension standing in for a covered call
- a throwing subprogram named by a BPF_PSEUDO_FUNC
- a record covering bpf_throw() itself
- a pad that calls a subprogram an extension can replace
- a covered bpf_throw() the sweep leaves last
- a region ending on a 16-byte instruction
- a pad that reloads from and writes to its own frame
- a pad two frames up
- a record covering only a nounwind call
- a pad terminated by _Unwind_Resume rather than bpf_unwind_resume
- a pad that calls a subprogram which tail calls
- the tail-call target carrying a table of its own
- a pad whose first instruction is a nop
- a pad that indexes its frame by a register the frame set before the
throwing call
- the same over a throwing global subprogram
One more shape needs an object of its own: exceptions_cleanup_ext_table.c
is an extension carrying a cleanup table of its own, attached over the
target the extension shape above already provides.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
.../selftests/bpf/exceptions_cleanup.h | 23 +
.../bpf/prog_tests/exceptions_cleanup.c | 346 ++++++
.../bpf/progs/exceptions_cleanup_ext_table.c | 48 +
.../bpf/progs/exceptions_cleanup_freplace.c | 17 +
.../progs/exceptions_cleanup_pad_freplace.c | 17 +
.../bpf/progs/exceptions_cleanup_shapes.c | 982 ++++++++++++++++++
6 files changed, 1433 insertions(+)
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_ext_table.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_freplace.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_pad_freplace.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
index 630d2e207119..383af9edf6f5 100644
--- a/tools/testing/selftests/bpf/exceptions_cleanup.h
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -4,6 +4,7 @@
#define __EXCEPTIONS_CLEANUP_H__
#define THROW_COOKIE 0x100
+#define INNER_COOKIE 0x200
/* progs/exceptions_cleanup.c: one bit per frame that reports it ran. */
#define RAN_FOO3_PREEMPT 0x1
@@ -12,6 +13,28 @@
#define RAN_FOO2_DROP 0x8
#define RAN_BUMP 0x10
+/* progs/exceptions_cleanup_shapes.c: one bit per shape, numbered its own way. */
+#define RAN_SWEEP 0x1
+#define RAN_SHARED 0x2
+#define RAN_REGS 0x4
+#define RAN_TAIL_CALL 0x8
+#define RAN_MAIN_PAD 0x10
+#define RAN_TC_TAKEN 0x20
+#define RAN_FREPLACE 0x40
+#define RAN_ADDR_TAKEN 0x80
+#define RAN_NO_SUBPROG 0x100
+#define RAN_PAD_CALLS 0x200
+#define RAN_PAD_FIRST 0x400
+#define RAN_WIDE_REC 0x800
+#define RAN_PAD_STACK 0x1000
+#define RAN_DEEP_PAD 0x2000
+#define RAN_NOUNWIND_REC 0x4000
+#define RAN_RESUME_ALIAS 0x8000
+#define RAN_PAD_TAIL_CALL 0x10000
+#define RAN_NOP_PAD 0x20000
+#define RAN_VAR_STACK 0x40000
+#define RAN_GLOBAL_PAD 0x80000
+
#define CLEANUP_REC(begin, end, landing_pad) \
".pushsection .bpf_cleanup,\"a\",@progbits;" \
".long " begin ";" \
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
index c932da7cbec1..2825dbb53b1c 100644
--- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -4,6 +4,10 @@
#include "exceptions_cleanup.h"
#include "exceptions_cleanup.skel.h"
#include "exceptions_cleanup_fail.skel.h"
+#include "exceptions_cleanup_shapes.skel.h"
+#include "exceptions_cleanup_freplace.skel.h"
+#include "exceptions_cleanup_pad_freplace.skel.h"
+#include "exceptions_cleanup_ext_table.skel.h"
/* foo3 threw: every frame that has a pad ran it. */
#define PADS_FOO3_THREW \
@@ -36,6 +40,346 @@ static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
ASSERT_EQ(skel->bss->pads_ran, pads | RAN_BUMP, "pads_ran");
}
+static void run_shape(struct exceptions_cleanup_shapes *skel, struct bpf_program *prog,
+ __u64 input, __u32 retval, __u64 pads)
+{
+ __u64 ctx = 0;
+ int err;
+
+ LIBBPF_OPTS(bpf_test_run_opts, topts,
+ .ctx_in = &ctx,
+ .ctx_size_in = sizeof(ctx),
+ );
+
+ skel->bss->input = input;
+ skel->bss->pads_ran = 0;
+
+ err = bpf_prog_test_run_opts(bpf_program__fd(prog), &topts);
+ if (!ASSERT_OK(err, "run"))
+ return;
+ ASSERT_EQ(topts.retval, retval, "retval");
+ ASSERT_EQ(skel->bss->pads_ran, pads, "pads_ran");
+}
+
+static void test_freplace(struct exceptions_cleanup_shapes *skel)
+{
+ struct exceptions_cleanup_freplace *fr;
+ struct bpf_link *link;
+ int tgt_fd;
+
+ tgt_fd = bpf_program__fd(skel->progs.entry_freplace);
+
+ fr = exceptions_cleanup_freplace__open();
+ if (!ASSERT_OK_PTR(fr, "freplace open"))
+ return;
+
+ if (!ASSERT_OK(bpf_program__set_attach_target(fr->progs.new_fr_callee,
+ tgt_fd, "fr_callee"),
+ "set_attach_target"))
+ goto out;
+ if (!ASSERT_OK(exceptions_cleanup_freplace__load(fr), "freplace load"))
+ goto out;
+
+ link = bpf_program__attach_freplace(fr->progs.new_fr_callee, tgt_fd,
+ "fr_callee");
+ if (!ASSERT_OK_PTR(link, "attach_freplace"))
+ goto out;
+
+ run_shape(skel, skel->progs.entry_freplace, 101, THROW_COOKIE, 0);
+ bpf_link__destroy(link);
+out:
+ exceptions_cleanup_freplace__destroy(fr);
+}
+
+static void test_pad_calls_freplace(struct exceptions_cleanup_shapes *skel)
+{
+ struct exceptions_cleanup_pad_freplace *fr;
+ struct bpf_link *link;
+ __u64 ctx = 0;
+ int tgt_fd, err;
+
+ LIBBPF_OPTS(bpf_test_run_opts, topts,
+ .ctx_in = &ctx,
+ .ctx_size_in = sizeof(ctx),
+ );
+
+ tgt_fd = bpf_program__fd(skel->progs.entry_pad_calls);
+
+ fr = exceptions_cleanup_pad_freplace__open();
+ if (!ASSERT_OK_PTR(fr, "pad freplace open"))
+ return;
+
+ if (!ASSERT_OK(bpf_program__set_attach_target(fr->progs.new_pad_callee,
+ tgt_fd, "pad_callee"),
+ "set_attach_target"))
+ goto out;
+ if (!ASSERT_OK(exceptions_cleanup_pad_freplace__load(fr), "pad freplace load"))
+ goto out;
+
+ link = bpf_program__attach_freplace(fr->progs.new_pad_callee, tgt_fd,
+ "pad_callee");
+ if (!ASSERT_OK_PTR(link, "attach_freplace"))
+ goto out;
+
+ skel->bss->input = 101;
+ skel->bss->pads_ran = 0;
+ skel->bss->pad_runs = 0;
+
+ err = bpf_prog_test_run_opts(tgt_fd, &topts);
+ if (!ASSERT_OK(err, "run"))
+ goto out_link;
+
+ ASSERT_EQ(skel->bss->pad_runs, 1, "pad_runs");
+ ASSERT_EQ(skel->bss->pads_ran, RAN_PAD_CALLS, "pads_ran");
+ ASSERT_EQ(topts.retval, THROW_COOKIE, "retval");
+out_link:
+ bpf_link__destroy(link);
+out:
+ exceptions_cleanup_pad_freplace__destroy(fr);
+}
+
+static void test_ext_table(struct exceptions_cleanup_shapes *skel)
+{
+ struct exceptions_cleanup_ext_table *fr;
+ struct bpf_link *link;
+ int tgt_fd;
+
+ tgt_fd = bpf_program__fd(skel->progs.entry_freplace);
+
+ fr = exceptions_cleanup_ext_table__open();
+ if (!ASSERT_OK_PTR(fr, "ext table open"))
+ return;
+
+ if (!ASSERT_OK(bpf_program__set_attach_target(fr->progs.new_fr_callee,
+ tgt_fd, "fr_callee"),
+ "set_attach_target"))
+ goto out;
+ if (!ASSERT_OK(exceptions_cleanup_ext_table__load(fr), "ext table load"))
+ goto out;
+
+ link = bpf_program__attach_freplace(fr->progs.new_fr_callee, tgt_fd,
+ "fr_callee");
+ if (!ASSERT_OK_PTR(link, "attach_freplace"))
+ goto out;
+
+ fr->bss->ext_pad_ran = 0;
+ run_shape(skel, skel->progs.entry_freplace, 101, THROW_COOKIE, 0);
+ ASSERT_EQ(fr->bss->ext_pad_ran, 1, "ext_pad_ran");
+
+ bpf_link__destroy(link);
+out:
+ exceptions_cleanup_ext_table__destroy(fr);
+}
+
+static void test_shapes(void)
+{
+ struct exceptions_cleanup_shapes *skel;
+
+ skel = exceptions_cleanup_shapes__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "shapes open_and_load"))
+ return;
+
+ /* The frame loads at all only if everything unreachable in it went. */
+ if (test__start_subtest("sweep_no_throw"))
+ run_shape(skel, skel->progs.entry_sweep, 1, 0, 0);
+ if (test__start_subtest("sweep_throw"))
+ run_shape(skel, skel->progs.entry_sweep, 101, THROW_COOKIE, RAN_SWEEP);
+
+ /* The covered call unwinds to the pad; the uncovered one walks past. */
+ if (test__start_subtest("shared_callee_no_throw")) {
+ skel->bss->outer_input = 0;
+ run_shape(skel, skel->progs.entry_shared, 1, 2, 0);
+ }
+ if (test__start_subtest("shared_callee_throw")) {
+ skel->bss->outer_input = 0;
+ run_shape(skel, skel->progs.entry_shared, 101, THROW_COOKIE, RAN_SHARED);
+ }
+ if (test__start_subtest("shared_callee_uncovered_throw")) {
+ skel->bss->outer_input = 101;
+ run_shape(skel, skel->progs.entry_shared, 1, THROW_COOKIE, 0);
+ skel->bss->outer_input = 0;
+ }
+
+ /* The pad only sets its bit if it got the frame's own r6-r9 back. */
+ if (test__start_subtest("pad_sees_callee_saved"))
+ run_shape(skel, skel->progs.entry_regs, 101, THROW_COOKIE, RAN_REGS);
+
+ /* Same check, with a tail-call-reachable callee: its spill moves. */
+ if (test__start_subtest("tail_call_no_throw"))
+ run_shape(skel, skel->progs.entry_tail_call, 1, 0, 0);
+ if (test__start_subtest("tail_call_throw"))
+ run_shape(skel, skel->progs.entry_tail_call, 101, THROW_COOKIE,
+ RAN_TAIL_CALL);
+
+ /* A region around a nounwind call: no pad dispatched, still loads. */
+ if (test__start_subtest("nounwind_region"))
+ run_shape(skel, skel->progs.entry_nounwind_rec, 1, 0, 0);
+
+ /* A pad in the main program's own frame, not in a subprogram. */
+ if (test__start_subtest("main_program_pad"))
+ run_shape(skel, skel->progs.entry_main_pad, 101, THROW_COOKIE,
+ RAN_MAIN_PAD);
+
+ /* The same call site either way: the subprogram's throw unwinds into
+ * this frame and runs its pad, an extension's stops at its own boundary.
+ */
+ if (test__start_subtest("freplace_subprog_throws"))
+ run_shape(skel, skel->progs.entry_freplace, 7, THROW_COOKIE,
+ RAN_FREPLACE);
+ if (test__start_subtest("freplace_extension_throws"))
+ test_freplace(skel);
+
+ /* A tail call that is taken: the walk ends at the target, so the cookie
+ * comes back from there and this frame's pad does not run -- though it
+ * is reachable, so a walk past the boundary would find it.
+ */
+ if (test__start_subtest("tail_call_taken")) {
+ int key = 0, prog_fd = bpf_program__fd(skel->progs.tc_target);
+
+ if (ASSERT_OK(bpf_map_update_elem(bpf_map__fd(skel->maps.taken_table),
+ &key, &prog_fd, BPF_ANY),
+ "populate taken_table"))
+ run_shape(skel, skel->progs.entry_tail_taken, 101,
+ THROW_COOKIE, 0);
+ }
+
+ /* A throwing subprog named by a BPF_PSEUDO_FUNC and handed to a
+ * bpf_loop() the verifier never reaches: the callback check has to fire
+ * on the helper call, not on the ld_imm64.
+ */
+ if (test__start_subtest("addr_taken_no_throw"))
+ run_shape(skel, skel->progs.entry_addr_taken, 1, 2, 0);
+ if (test__start_subtest("addr_taken_throw"))
+ run_shape(skel, skel->progs.entry_addr_taken, 101, THROW_COOKIE,
+ RAN_ADDR_TAKEN);
+
+ /* A record covering bpf_throw() itself rather than a call to a frame
+ * that throws: raised, caught up with and delivered in one frame.
+ */
+ if (test__start_subtest("no_subprog_no_throw"))
+ run_shape(skel, skel->progs.entry_no_subprog, 1, 0, 0);
+ if (test__start_subtest("no_subprog_throw"))
+ run_shape(skel, skel->progs.entry_no_subprog, 101, THROW_COOKIE,
+ RAN_NO_SUBPROG);
+
+ /* A pad that calls a subprogram; with a throwing extension in its place,
+ * the nested exception has to stop there, not restart this pad.
+ */
+ if (test__start_subtest("pad_calls_subprog")) {
+ skel->bss->pad_runs = 0;
+ run_shape(skel, skel->progs.entry_pad_calls, 101, THROW_COOKIE,
+ RAN_PAD_CALLS);
+ ASSERT_EQ(skel->bss->pad_runs, 1, "pad_runs");
+ }
+ if (test__start_subtest("pad_calls_throwing_extension"))
+ test_pad_calls_freplace(skel);
+
+ /* A covered throw the sweep leaves last, where the default exception
+ * callback is patched in; the pad's bit needs r6-r9 still spilled.
+ */
+ if (test__start_subtest("pad_before_throw"))
+ run_shape(skel, skel->progs.entry_pad_first, 101, THROW_COOKIE,
+ RAN_PAD_FIRST);
+
+ /* A region whose last instruction is a 16-byte one, so that end - 1
+ * names the half of it that is not an instruction.
+ */
+ if (test__start_subtest("region_ends_on_ldimm64"))
+ run_shape(skel, skel->progs.entry_wide_rec, 101, THROW_COOKIE,
+ RAN_WIDE_REC);
+
+ /* A pad that reloads from and writes to its own frame's stack, which a
+ * JIT addressing the frame through the stack pointer gets wrong.
+ */
+ if (test__start_subtest("pad_uses_own_frame"))
+ run_shape(skel, skel->progs.entry_pad_stack, 101, THROW_COOKIE,
+ RAN_PAD_STACK);
+
+ /* The same, with an uncovered frame between the throw and the pad. */
+ if (test__start_subtest("pad_two_frames_up"))
+ run_shape(skel, skel->progs.entry_deep_pad, 101, THROW_COOKIE,
+ RAN_DEEP_PAD);
+
+ /* An extension program with a cleanup table of its own. */
+ if (test__start_subtest("extension_carries_table"))
+ test_ext_table(skel);
+
+ /* A pad terminated by _Unwind_Resume, which libbpf maps onto the kfunc;
+ * every other program here calls bpf_unwind_resume directly.
+ */
+ if (test__start_subtest("resume_alias"))
+ run_shape(skel, skel->progs.entry_resume_alias, 101,
+ THROW_COOKIE, RAN_RESUME_ALIAS);
+
+ /* A pad that calls a subprogram which tail calls, array empty and then
+ * populated: the tail call releases only the callee's own prologue.
+ */
+ if (test__start_subtest("pad_callee_tail_call")) {
+ int key = 0, prog_fd = bpf_program__fd(skel->progs.pad_tc_target);
+
+ skel->bss->pad_tc_target_ran = 0;
+ skel->bss->pad_runs = 0;
+ run_shape(skel, skel->progs.entry_pad_tail_call, 101,
+ THROW_COOKIE, RAN_PAD_TAIL_CALL);
+ ASSERT_EQ(skel->bss->pad_tc_target_ran, 0, "target not run");
+ ASSERT_EQ(skel->bss->pad_runs, 1, "pad_runs");
+
+ if (ASSERT_OK(bpf_map_update_elem(bpf_map__fd(skel->maps.pad_tc_table),
+ &key, &prog_fd, BPF_ANY),
+ "populate pad_tc_table")) {
+ skel->bss->pad_runs = 0;
+ run_shape(skel, skel->progs.entry_pad_tail_call, 101,
+ THROW_COOKIE, RAN_PAD_TAIL_CALL);
+ ASSERT_EQ(skel->bss->pad_tc_target_ran, 1, "target ran");
+ ASSERT_EQ(skel->bss->pad_runs, 1, "pad_runs");
+ }
+ }
+
+ /* The same, into a target that carries a table and throws: that target
+ * is a boundary, so the outer pad runs once, not twice.
+ */
+ if (test__start_subtest("pad_callee_tail_call_throws")) {
+ int key = 0, prog_fd = bpf_program__fd(skel->progs.pad_tc_throw_target);
+
+ if (ASSERT_OK(bpf_map_update_elem(bpf_map__fd(skel->maps.pad_tc_table),
+ &key, &prog_fd, BPF_ANY),
+ "populate pad_tc_table")) {
+ skel->bss->pad_runs = 0;
+ skel->bss->tc_target_pad_runs = 0;
+ run_shape(skel, skel->progs.entry_pad_tail_call, 101,
+ THROW_COOKIE, RAN_PAD_TAIL_CALL);
+ /* The target cleaned up after itself, once. */
+ ASSERT_EQ(skel->bss->tc_target_pad_runs, 1,
+ "tc_target_pad_runs");
+ /* And the outer pad was not started over. */
+ ASSERT_EQ(skel->bss->pad_runs, 1, "pad_runs");
+ }
+ }
+
+ /* A pad whose first instruction opt_remove_nops() deletes: the record
+ * has to follow the pad rather than be dropped with the nop.
+ */
+ if (test__start_subtest("nop_at_pad_head"))
+ run_shape(skel, skel->progs.entry_nop_pad, 101, THROW_COOKIE,
+ RAN_NOP_PAD);
+
+ /* A pad that indexes its frame's stack by a register the frame set
+ * before the throwing call, marked precise back across the unwind edge.
+ */
+ if (test__start_subtest("pad_var_stack_offset"))
+ run_shape(skel, skel->progs.entry_var_stack, 101, THROW_COOKIE,
+ RAN_VAR_STACK);
+
+ /* The same, where the throw is in a global subprogram, which the
+ * verifier walks without a frame of its own.
+ */
+ if (test__start_subtest("pad_over_global_subprog"))
+ run_shape(skel, skel->progs.entry_global_pad, 101, THROW_COOKIE,
+ RAN_GLOBAL_PAD);
+
+ exceptions_cleanup_shapes__destroy(skel);
+}
+
/* A table with more records than the program has instructions: no valid one
* can look like that, and it is refused before the kernel allocates for it.
*/
@@ -114,6 +458,8 @@ void test_exceptions_cleanup(void)
exceptions_cleanup__destroy(skel);
+ test_shapes();
+
RUN_TESTS(exceptions_cleanup_fail);
if (test__start_subtest("cleanup_info_cnt"))
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_ext_table.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_ext_table.c
new file mode 100644
index 000000000000..d14db48d6b29
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_ext_table.c
@@ -0,0 +1,48 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+__u64 ext_pad_ran = 0;
+
+/* Without a 32-bit int in BTF, libbpf's dummy_ksym var gets type id 0. */
+int btf_int_anchor;
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume();
+}
+
+static __used __naked __noinline __u64 ext_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = %[cookie];"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ "r1 = %[ext_pad_ran] ll;"
+ "r2 = 1;"
+ "*(u64 *)(r1 + 0) = r2;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [cookie]"i"(THROW_COOKIE), __imm_addr(ext_pad_ran)
+ : __clobber_all);
+}
+
+SEC("freplace/fr_callee")
+__u64 new_fr_callee(__u64 x)
+{
+ return ext_frame();
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_freplace.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_freplace.c
new file mode 100644
index 000000000000..afb358fd3d40
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_freplace.c
@@ -0,0 +1,17 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "exceptions_cleanup.h"
+
+/* Without a 32-bit int in BTF, libbpf's dummy_ksym var gets type id 0. */
+int btf_int_anchor;
+
+SEC("freplace/fr_callee")
+__u64 new_fr_callee(__u64 x)
+{
+ bpf_throw(THROW_COOKIE);
+ return 0;
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_pad_freplace.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_pad_freplace.c
new file mode 100644
index 000000000000..eabac6baabb7
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_pad_freplace.c
@@ -0,0 +1,17 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "exceptions_cleanup.h"
+
+/* Without a 32-bit int in BTF, libbpf's dummy_ksym var gets type id 0. */
+int btf_int_anchor;
+
+SEC("freplace/pad_callee")
+__u64 new_pad_callee(__u64 x)
+{
+ bpf_throw(INNER_COOKIE);
+ return 0;
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
new file mode 100644
index 000000000000..5e8ca799d9d2
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
@@ -0,0 +1,982 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+#define PAD_COUNT \
+ "r1 = %[pad_runs] ll;" \
+ "r2 = *(u64 *)(r1 + 0);" \
+ "r2 += 1;" \
+ "*(u64 *)(r1 + 0) = r2;"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_rcu_read_lock();
+ bpf_rcu_read_unlock();
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume();
+}
+
+__u64 input = 0;
+__u64 outer_input = 0;
+__u64 magic = 0x5eed;
+__u64 pads_ran = 0;
+__u64 pad_runs = 0;
+
+/* 1. Everything a cleanup table leaves dead: a throw's continuation, the tail
+ * after a resume with an ld_imm64 and a branch in it, and the block only that
+ * continuation reaches.
+ */
+static __used __naked __noinline __u64 sweep_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "call bpf_preempt_disable;"
+ "if r6 < 101 goto 6f;"
+ "r1 = %[cookie];"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "goto 3f;"
+"4:" /* landing pad */
+ "r7 = r0;"
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "r1 = %[pads_ran] ll;"
+ "r2 = *(u64 *)(r1 + 0);"
+ "if r2 == 0 goto 5f;"
+ "call bpf_preempt_enable;"
+ "r0 = 7;"
+ "exit;"
+"5:"
+ "r0 = 8;"
+ "exit;"
+"3:" /* dead: only the dead goto reaches it */
+ "r0 = 9;"
+ "exit;"
+"6:" /* live: the ordinary return */
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "4b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_SWEEP),
+ __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_sweep(void *ctx)
+{
+ return sweep_frame();
+}
+
+/* 2. A callee called from both a covered and an uncovered site: the pad
+ * belongs to the call site, not the callee. Either site can throw; the RCU
+ * lock is taken between them, so only the covered call unwinds holding it.
+ */
+static __used __noinline __u64 shared_callee(__u64 x)
+{
+ if (x > 100)
+ bpf_throw(THROW_COOKIE);
+ return x + 1;
+}
+
+static __used __naked __noinline __u64 shared_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "r1 = %[outer_input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "call shared_callee;" /* uncovered: no pad for its unwind */
+ "call bpf_rcu_read_lock;"
+ "r1 = r6;"
+"1:" "call shared_callee;" /* cleanup region */
+"2:"
+ "r6 = r0;"
+ "call bpf_rcu_read_unlock;"
+ "r0 = r6;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+ "call bpf_rcu_read_unlock;"
+ PAD_RAN("%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_SHARED), __imm_addr(input), __imm_addr(outer_input),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_shared(void *ctx)
+{
+ return shared_frame();
+}
+
+/* 3. A landing pad that reads its frame's callee-saved registers: the callee
+ * overwrites r6-r9 before it throws, so the check passes only if the walker
+ * found the spill in the discarded callee's prologue.
+ */
+#define LOAD_MAGIC_REGS \
+ "r1 = %[magic] ll;" \
+ "r6 = *(u64 *)(r1 + 0);" \
+ "r7 = r6;" \
+ "r7 += 1;" \
+ "r8 = r6;" \
+ "r8 += 2;" \
+ "r9 = r6;" \
+ "r9 += 3;"
+
+/* Set @bit only if r6-r9 still hold what LOAD_MAGIC_REGS put there. */
+#define CHECK_MAGIC_REGS(bit) \
+ "r1 = %[magic] ll;" \
+ "r2 = *(u64 *)(r1 + 0);" \
+ "if r6 != r2 goto 9f;" \
+ "r2 += 1;" \
+ "if r7 != r2 goto 9f;" \
+ "r2 += 1;" \
+ "if r8 != r2 goto 9f;" \
+ "r2 += 1;" \
+ "if r9 != r2 goto 9f;" \
+ PAD_RAN(bit) \
+ "9:"
+
+static __used __naked __noinline __u64 regs_thrower(void)
+{
+ asm volatile (
+ /* Not this frame's to keep, and that is the point. */
+ "r6 = 0xdead;"
+ "r7 = 0xbeef;"
+ "r8 = 0xcafe;"
+ "r9 = 0xf00d;"
+ "r1 = %[cookie];"
+ "call bpf_throw;"
+ "r0 = 0;"
+ "exit;"
+ :
+ : [cookie]"i"(THROW_COOKIE)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 regs_frame(void)
+{
+ asm volatile (
+ LOAD_MAGIC_REGS
+ "call bpf_preempt_disable;"
+"1:" "call regs_thrower;" /* cleanup region */
+"2:"
+ "call bpf_preempt_enable;"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ CHECK_MAGIC_REGS("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_REGS),
+ __imm_addr(magic), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_regs(void *ctx)
+{
+ return regs_frame();
+}
+
+/* 4. The same, with a tail-call-reachable callee: its prologue pushes the tail
+ * call counter, which moves the spill the walker reads. The array is left
+ * empty; being reachable is the point.
+ */
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(__u32));
+ __uint(value_size, sizeof(__u32));
+} jmp_table SEC(".maps");
+
+static __used __noinline __u64 tc_thrower(void *ctx)
+{
+ /* Never taken; its presence is what makes this frame, whose spill the
+ * walker reads, tail-call-reachable.
+ */
+ bpf_tail_call_static(ctx, &jmp_table, 0);
+ asm volatile (
+ "r6 = 0xdead;"
+ "r7 = 0xbeef;"
+ "r8 = 0xcafe;"
+ "r9 = 0xf00d;"
+ "r1 = %[cookie];"
+ "call bpf_throw;"
+ :
+ : [cookie]"i"(THROW_COOKIE)
+ : __clobber_all);
+ return 0;
+}
+
+/* The frame with the pad is the program itself, and __naked: r1 at entry is
+ * the only place to get a context for bpf_tail_call().
+ */
+SEC("syscall")
+__naked int entry_tail_call(void)
+{
+ asm volatile (
+ "*(u64 *)(r10 - 8) = r1;" /* the context, straight from entry */
+ LOAD_MAGIC_REGS
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 8f;"
+ "r1 = *(u64 *)(r10 - 8);"
+"1:" "call tc_thrower;" /* cleanup region */
+"2:"
+"8:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ CHECK_MAGIC_REGS("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_TAIL_CALL), __imm_addr(input),
+ __imm_addr(magic), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/* 5. A landing pad in the main program's own frame. jit_subprogs() compiles it
+ * as func[0], but the walker finds the outer bpf_prog's ksym, so the table has
+ * to be handed over or the pad is never dispatched.
+ */
+SEC("syscall")
+__naked int entry_main_pad(void)
+{
+ asm volatile (
+ LOAD_MAGIC_REGS
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 8f;"
+"1:" "call regs_thrower;" /* cleanup region */
+"2:"
+"8:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ CHECK_MAGIC_REGS("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_MAIN_PAD), __imm_addr(input),
+ __imm_addr(magic), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/* 6. A tail call that is really taken: the target is a program in its own
+ * right, so the walk ends there and this frame's pad does not run. The callee
+ * can also throw, which keeps the pad out of the sweep.
+ */
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(__u32));
+ __uint(value_size, sizeof(__u32));
+} taken_table SEC(".maps");
+
+SEC("syscall")
+int tc_target(void *ctx)
+{
+ bpf_throw(THROW_COOKIE);
+ return 0;
+}
+
+static __used __noinline __u64 tc_taken_callee(void *ctx, __u64 x)
+{
+ /* Never true at run time; the verifier cannot know that, and its
+ * unwind out of here is what keeps the caller's pad alive.
+ */
+ if (x == 7)
+ bpf_throw(THROW_COOKIE);
+ bpf_tail_call_static(ctx, &taken_table, 0);
+ return 0;
+}
+
+SEC("syscall")
+__naked int entry_tail_taken(void)
+{
+ asm volatile (
+ "*(u64 *)(r10 - 8) = r1;" /* the context, straight from entry */
+ "r1 = %[input] ll;"
+ /* Unnarrowed, the way entry_freplace hands fr_callee its argument: a
+ * guard here would prune the callee's throw and let the sweep take the
+ * pad, leaving no record for a walk past the boundary to match.
+ */
+ "r2 = *(u64 *)(r1 + 0);"
+ "r1 = *(u64 *)(r10 - 8);"
+"1:" "call tc_taken_callee;" /* cleanup region */
+"2:"
+ "exit;" /* the cookie, delivered at tc_target */
+"3:" /* landing pad: must not run */
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_TC_TAKEN), __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/* 7. An extension program over the callee of a covered call. The walk ends in
+ * the extension's frame, as it does for a tail call target, so the pad does
+ * not run; fr_callee() can also throw by itself.
+ */
+__noinline __u64 fr_callee(__u64 x)
+{
+ if (x == 7)
+ bpf_throw(THROW_COOKIE);
+ return x + 1;
+}
+
+SEC("syscall")
+__naked int entry_freplace(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call fr_callee;" /* cleanup region */
+"2:"
+ "exit;"
+"3:" /* landing pad */
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_FREPLACE), __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/* 8. A throwing subprogram named by a BPF_PSEUDO_FUNC on a path never taken. */
+static __used __noinline int cb_thrower(__u32 idx, void *ctx)
+{
+ bpf_throw(THROW_COOKIE);
+ return 0;
+}
+
+static __used __noinline __u64 addr_taken_callee(__u64 x)
+{
+ if (x <= 100)
+ return x + 1;
+ bpf_throw(THROW_COOKIE);
+ return bpf_loop(1, cb_thrower, NULL, 0);
+}
+
+SEC("syscall")
+__naked int entry_addr_taken(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call addr_taken_callee;" /* cleanup region */
+"2:"
+ "exit;"
+"3:" /* landing pad */
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_ADDR_TAKEN), __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/* 9. A record that covers bpf_throw() itself: the throwing frame is both the
+ * frame the record covers and the boundary, so the pad runs on the way to
+ * delivering the cookie.
+ */
+SEC("syscall")
+__naked int entry_no_subprog(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 8f;"
+ "call bpf_preempt_disable;"
+ "r1 = %[cookie];"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "exit;"
+"8:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_NO_SUBPROG),
+ __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/* 10. A landing pad that calls a subprogram an extension can replace. The
+ * load-time rule cannot see it coming, so what stops a nested exception is the
+ * walk ending in the extension's frame; pad_runs says the pad ran once.
+ */
+__noinline __u64 pad_callee(__u64 x)
+{
+ return x + 1;
+}
+
+static __used __noinline __u64 pc_thrower(__u64 x)
+{
+ if (x > 100)
+ bpf_throw(THROW_COOKIE);
+ return x + 1;
+}
+
+static __used __naked __noinline __u64 pad_calls_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call pc_thrower;" /* cleanup region */
+"2:"
+ "exit;"
+"3:" /* landing pad */
+ "r6 = r0;"
+ "r1 = 1;"
+ "call pad_callee;" /* an extension can stand in here */
+ PAD_COUNT
+ PAD_RAN("%[ran]")
+ "r1 = r6;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_PAD_CALLS), __imm_addr(input), __imm_addr(pads_ran),
+ __imm_addr(pad_runs)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_pad_calls(void *ctx)
+{
+ return pad_calls_frame();
+}
+
+/* 11. A covered bpf_throw() the sweep leaves last, where the exception
+ * callback patchlet -- the one that does not keep the call it replaced in the
+ * last slot -- has to carry the marks with it; r6-r9 reports a lost one.
+ */
+SEC("syscall")
+__naked int entry_pad_first(void)
+{
+ asm volatile (
+ LOAD_MAGIC_REGS
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 7f;"
+ "goto 4f;"
+"3:" /* landing pad, ahead of the call */
+ CHECK_MAGIC_REGS("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+"7:"
+ "r0 = 0;"
+ "exit;"
+"4:"
+ "r1 = %[cookie];"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "exit;" /* dead: swept, leaving the call last */
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_PAD_FIRST),
+ __imm_addr(input), __imm_addr(magic), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+/* 12. A cleanup region whose last instruction is a 16-byte one, so end - 1
+ * names the half that is not an instruction of its own. Well formed, and a
+ * rule against it would turn it away.
+ */
+static __used __naked __noinline __u64 wide_rec_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "call bpf_rcu_read_lock;"
+ "r1 = r6;"
+"1:" "call shared_callee;" /* cleanup region begins */
+ "r1 = %[magic] ll;" /* ... and ends on this pair */
+"2:"
+ "r6 = r0;"
+ "call bpf_rcu_read_unlock;"
+ "r0 = r6;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+ "call bpf_rcu_read_unlock;"
+ PAD_RAN("%[ran]")
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_WIDE_REC), __imm_addr(input), __imm_addr(magic),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_wide_rec(void *ctx)
+{
+ return wide_rec_frame();
+}
+
+/* 13. A pad that reloads from and stores to its own frame -- the shape every
+ * compiler-generated pad has. A JIT that addresses the frame through the stack
+ * pointer, arm64, has to find a pad's frame another way.
+ */
+static __used __naked __noinline __u64 pad_stack_frame(void)
+{
+ asm volatile (
+ "r1 = %[magic] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "*(u64 *)(r10 - 8) = r1;" /* what the pad will want */
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call pc_thrower;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r6 = r0;"
+ "r7 = *(u64 *)(r10 - 8);" /* reload it out of the frame */
+ "*(u64 *)(r10 - 16) = r7;" /* and write the frame while here */
+ "r1 = %[magic] ll;"
+ "r2 = *(u64 *)(r1 + 0);"
+ "if r7 != r2 goto 9f;"
+ "r3 = *(u64 *)(r10 - 16);"
+ "if r3 != r2 goto 9f;"
+ PAD_RAN("%[ran]")
+"9:"
+ "r1 = r6;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_PAD_STACK), __imm_addr(input), __imm_addr(magic),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_pad_stack(void *ctx)
+{
+ return pad_stack_frame();
+}
+
+/* 14. The same, with a frame in between that has no pad of its own, so the
+ * liveness query for an outer frame has more than one frame to walk.
+ */
+static __used __noinline __u64 deep_mid(__u64 x)
+{
+ return pc_thrower(x) + 1;
+}
+
+static __used __naked __noinline __u64 deep_frame(void)
+{
+ asm volatile (
+ "r1 = %[magic] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "*(u64 *)(r10 - 8) = r1;" /* nothing but the pad reads this */
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call deep_mid;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r6 = r0;"
+ "r7 = *(u64 *)(r10 - 8);"
+ "r1 = %[magic] ll;"
+ "r2 = *(u64 *)(r1 + 0);"
+ "if r7 != r2 goto 9f;"
+ PAD_RAN("%[ran]")
+"9:"
+ "r1 = r6;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_DEEP_PAD), __imm_addr(input), __imm_addr(magic),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_deep_pad(void *ctx)
+{
+ return deep_frame();
+}
+
+/* 15. A region around a call the kernel knows cannot unwind: no call site is
+ * marked, nothing reaches the pad, and the sweep removes it. The program is
+ * otherwise ordinary and has to load.
+ */
+static __used __naked __noinline __u64 nounwind_rec_frame(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+"1:" "call bpf_preempt_enable;" /* cleanup region: nounwind */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad, never dispatched */
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_NOUNWIND_REC), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_nounwind_rec(void *ctx)
+{
+ return nounwind_rec_frame();
+}
+
+/* 16. The name a frontend gives the resume: LLVM emits _Unwind_Resume() and
+ * libbpf maps it onto bpf_unwind_resume(), which every other pad here calls.
+ */
+extern void _Unwind_Resume(void) __ksym;
+
+static __used __noinline void __resume_alias_btf_anchor(void)
+{
+ _Unwind_Resume();
+}
+
+static __used __naked __noinline __u64 resume_alias_frame(void)
+{
+ asm volatile (
+"1:" "call regs_thrower;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ PAD_RAN("%[ran]")
+ "call _Unwind_Resume;" /* the frontend's name for it */
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_RESUME_ALIAS), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_resume_alias(void *ctx)
+{
+ return resume_alias_frame();
+}
+
+/* 17. A landing pad that calls a subprogram which tail calls. A tail call in
+ * the pad itself is refused, but the callee's prologue really did run, so its
+ * tail call releases exactly that and the target returns into the pad.
+ */
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(__u32));
+ __uint(value_size, sizeof(__u32));
+} pad_tc_table SEC(".maps");
+
+__u64 pad_tc_target_ran = 0;
+
+SEC("syscall")
+int pad_tc_target(void *ctx)
+{
+ pad_tc_target_ran += 1;
+ return 0;
+}
+
+static __used __noinline __u64 pad_tc_callee(void *ctx)
+{
+ /* Taken only once the test has populated the array. */
+ bpf_tail_call_static(ctx, &pad_tc_table, 0);
+ return 0;
+}
+
+/* The frame with the pad is the program itself, and __naked: r1 at entry is
+ * the only place to get a context for bpf_tail_call(); the pad reloads it
+ * from the frame's stack.
+ */
+SEC("syscall")
+__naked int entry_pad_tail_call(void)
+{
+ asm volatile (
+ "*(u64 *)(r10 - 8) = r1;" /* the context, straight from entry */
+ LOAD_MAGIC_REGS
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 8f;"
+"1:" "call regs_thrower;" /* cleanup region */
+"2:"
+"8:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ PAD_COUNT
+ "r1 = *(u64 *)(r10 - 8);"
+ "call pad_tc_callee;"
+ /* Only if the frame survived the callee's tail call. */
+ CHECK_MAGIC_REGS("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_PAD_TAIL_CALL), __imm_addr(input),
+ __imm_addr(magic), __imm_addr(pads_ran), __imm_addr(pad_runs)
+ : __clobber_all);
+}
+
+/* 18. The other target for that same tail call: a program with a table of its
+ * own, throwing inside the outer exception. The tail call made it a boundary,
+ * so the inner walk ends there and the outer pad, cookie and r6-r9 survive.
+ */
+__u64 tc_target_pad_runs = 0;
+__u64 inner_magic = 0xd00d;
+
+static __used __naked __noinline __u64 inner_thrower(void)
+{
+ asm volatile (
+ /* Not this frame's to keep, the same as regs_thrower. */
+ "r6 = 0xf00d;"
+ "r7 = 0xcafe;"
+ "r8 = 0xbeef;"
+ "r9 = 0xdead;"
+ "r1 = %[cookie];"
+ "call bpf_throw;"
+ "r0 = 0;"
+ "exit;"
+ :
+ : [cookie]"i"(INNER_COOKIE)
+ : __clobber_all);
+}
+
+SEC("syscall")
+__naked int pad_tc_throw_target(void)
+{
+ asm volatile (
+ /* Distinct from the outer pad's, so neither can stand in for it. */
+ "r1 = %[inner_magic] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "r7 = r6;"
+ "r7 += 1;"
+ "r8 = r6;"
+ "r8 += 2;"
+ "r9 = r6;"
+ "r9 += 3;"
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "if r1 < 101 goto 8f;"
+"1:" "call inner_thrower;" /* cleanup region */
+"2:"
+"8:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ /* This frame's own r6-r9, not the outer pad's. */
+ "r1 = %[inner_magic] ll;"
+ "r2 = *(u64 *)(r1 + 0);"
+ "if r6 != r2 goto 9f;"
+ "r2 += 1;"
+ "if r7 != r2 goto 9f;"
+ "r2 += 1;"
+ "if r8 != r2 goto 9f;"
+ "r2 += 1;"
+ "if r9 != r2 goto 9f;"
+ "r1 = %[tc_target_pad_runs] ll;"
+ "r2 = *(u64 *)(r1 + 0);"
+ "r2 += 1;"
+ "*(u64 *)(r1 + 0) = r2;"
+"9:"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : __imm_addr(input), __imm_addr(inner_magic),
+ __imm_addr(tc_target_pad_runs)
+ : __clobber_all);
+}
+
+/* 19. A landing pad whose first instruction is a nop. opt_remove_nops() runs
+ * long after the cleanup walk, so the record has to follow the pad to the
+ * instruction that takes its place rather than be dropped with the nop.
+ */
+static __used __naked __noinline __u64 nop_pad_frame(void)
+{
+ asm volatile (
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "if r6 < 101 goto 6f;"
+ "r1 = %[cookie];"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+"6:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad: a nop, then its body */
+ "goto +0;"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_NOP_PAD),
+ __imm_addr(input), __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_nop_pad(void *ctx)
+{
+ return nop_pad_frame();
+}
+
+/* 20. A landing pad that indexes its own frame's stack by a register the frame
+ * set before the throwing call. The offset is not a constant, so r6 has to be
+ * marked precise from inside the pad, back across the unwind edge.
+ */
+/* A thrower that touches none of r6-r9, so nothing in it can answer for the
+ * pad frame's r6 and the walk has to leave this frame to look.
+ */
+static __used __naked __noinline __u64 var_thrower(void)
+{
+ asm volatile (
+ "if r1 < 101 goto 1f;"
+ "r1 = %[cookie];"
+ "call bpf_throw;"
+"1:"
+ "r0 = 0;"
+ "exit;"
+ :
+ : [cookie]"i"(THROW_COOKIE)
+ : __clobber_all);
+}
+
+static __used __naked __noinline __u64 var_stack_frame(void)
+{
+ asm volatile (
+ "r1 = %[magic] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "*(u64 *)(r10 - 8) = r1;" /* the slot the pad will read... */
+ "*(u64 *)(r10 - 16) = r1;" /* ...whichever of the two it is */
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "r6 &= 1;" /* an unknown slot number... */
+ "r6 <<= 3;" /* ...as an aligned byte offset */
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call var_thrower;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+ "r1 = r10;"
+ "r1 += r6;" /* variable offset into the frame */
+ "r2 = *(u64 *)(r1 - 16);"
+ "r3 = %[magic] ll;"
+ "r3 = *(u64 *)(r3 + 0);"
+ "if r2 != r3 goto 9f;"
+ PAD_RAN("%[ran]")
+"9:"
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_VAR_STACK), __imm_addr(input), __imm_addr(magic),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_var_stack(void *ctx)
+{
+ return var_stack_frame();
+}
+
+/* 21. The same walk, over a call to a global subprogram: the verifier enters
+ * no frame for one, so the exception arrives at the pad from the call site
+ * rather than from a throw some frames deeper.
+ */
+__noinline __u64 global_thrower(__u64 x)
+{
+ if (x > 100)
+ bpf_throw(THROW_COOKIE);
+ return x + 1;
+}
+
+static __used __naked __noinline __u64 global_pad_frame(void)
+{
+ asm volatile (
+ "r1 = %[magic] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+ "*(u64 *)(r10 - 8) = r1;"
+ "*(u64 *)(r10 - 16) = r1;"
+ "r1 = %[input] ll;"
+ "r6 = *(u64 *)(r1 + 0);"
+ "r6 &= 1;"
+ "r6 <<= 3;"
+ "r1 = %[input] ll;"
+ "r1 = *(u64 *)(r1 + 0);"
+"1:" "call global_thrower;" /* cleanup region */
+"2:"
+ "r0 = 0;"
+ "exit;"
+"3:" /* landing pad */
+ "r7 = r0;"
+ "r1 = r10;"
+ "r1 += r6;"
+ "r2 = *(u64 *)(r1 - 16);"
+ "r3 = %[magic] ll;"
+ "r3 = *(u64 *)(r3 + 0);"
+ "if r2 != r3 goto 9f;"
+ PAD_RAN("%[ran]")
+"9:"
+ "r1 = r7;"
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [ran]"i"(RAN_GLOBAL_PAD), __imm_addr(input), __imm_addr(magic),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+SEC("syscall")
+int entry_global_pad(void *ctx)
+{
+ return global_pad_frame();
+}
+
+char _license[] SEC("license") = "GPL";
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach
2026-09-21 21:02 ` [PATCH bpf-next v4 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
@ 2026-09-21 21:19 ` sashiko-bot
0 siblings, 0 replies; 80+ messages in thread
From: sashiko-bot @ 2026-09-21 21:19 UTC (permalink / raw)
To: Yonghong Song; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] bpf: multi-line comments do not follow mandatory style
--
commit 0b5284880fa55ebb7577505990c5eb80e76e625a
Author: Yonghong Song <yonghong.song@linux.dev>
selftests/bpf: Cover the exception cleanup shapes the chain does not reach
This commit adds end-to-end tests for exception cleanup shapes that the main
call chain does not reach. It introduces various edge cases in new BPF
programs and verifies their exception handling and cleanup behaviors.
> diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> index c932da7cbec16..2825dbb53b1c6 100644
> --- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> +++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
[ ... ]
> @@ -36,6 +40,346 @@ static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
[ ... ]
> + /* The same call site either way: the subprogram's throw unwinds into
> + * this frame and runs its pad, an extension's stops at its own boundary.
> + */
[Severity: Low]
This isn't a bug, but does this multi-line comment follow the required style?
The BPF subsystem guidelines strictly require that multi-line comments must
have the opening /* on its own line, with the comment text starting on the
next line.
> + if (test__start_subtest("freplace_subprog_throws"))
> + run_shape(skel, skel->progs.entry_freplace, 7, THROW_COOKIE,
> + RAN_FREPLACE);
[ ... ]
> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
> new file mode 100644
> index 0000000000000..5e8ca799d9d2a
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
> @@ -0,0 +1,982 @@
[ ... ]
> +__u64 pads_ran = 0;
> +__u64 pad_runs = 0;
> +
> +/* 1. Everything a cleanup table leaves dead: a throw's continuation, the tail
> + * after a resume with an ld_imm64 and a branch in it, and the block only that
> + * continuation reaches.
> + */
[Severity: Low]
This isn't a bug either, but here is another instance where the multi-line
comment rule is violated ahead of sweep_frame(). Could we update these to
have the opening /* on its own line?
> +static __used __naked __noinline __u64 sweep_frame(void)
> +{
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260921210033.1715000-1-yonghong.song@linux.dev?part=19
^ permalink raw reply [flat|nested] 80+ messages in thread
* [PATCH bpf-next v4 20/20] selftests/bpf: Load an exception cleanup program from a light skeleton
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (18 preceding siblings ...)
2026-09-21 21:02 ` [PATCH bpf-next v4 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
@ 2026-09-21 21:02 ` Yonghong Song
2026-09-22 1:08 ` [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Eduard Zingerman
20 siblings, 0 replies; 80+ messages in thread
From: Yonghong Song @ 2026-09-21 21:02 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, kernel-team
For a light skeleton, libbpf hands the records to bpf_gen__prog_load(),
which writes them into a blob and emits a loader program that issues
BPF_PROG_LOAD from inside the kernel. Nothing about that is shared with the
ordinary path: the attr is built field by field, and the kfunc names are
resolved by the loader program when it runs.
exceptions_cleanup_light.c is the smallest program that can tell whether a
table survives that trip: one record covering the bpf_throw() call itself,
one pad that sets a bit and resumes. If the table arrives, the pad runs and
the cookie comes back; if it does not, the program does not load at all.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
tools/testing/selftests/bpf/Makefile.skel | 2 +-
.../selftests/bpf/exceptions_cleanup.h | 3 ++
.../bpf/prog_tests/exceptions_cleanup.c | 28 +++++++++++++
.../bpf/progs/exceptions_cleanup_light.c | 39 +++++++++++++++++++
4 files changed, 71 insertions(+), 1 deletion(-)
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
diff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel
index 580d1d82c186..6d7518fa4018 100644
--- a/tools/testing/selftests/bpf/Makefile.skel
+++ b/tools/testing/selftests/bpf/Makefile.skel
@@ -33,7 +33,7 @@ LINKED_SKELS := test_static_linked.skel.h linked_funcs.skel.h \
LSKELS := fexit_sleep.c trace_printk.c trace_vprintk.c map_ptr_kern.c \
core_kern.c core_kern_overflow.c test_ringbuf.c \
test_ringbuf_n.c test_ringbuf_map_key.c test_ringbuf_write.c \
- test_ringbuf_overwrite.c
+ test_ringbuf_overwrite.c exceptions_cleanup_light.c
LSKELS_SIGNED := fentry_test.c fexit_test.c atomics.c
diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
index 383af9edf6f5..e14b651263b8 100644
--- a/tools/testing/selftests/bpf/exceptions_cleanup.h
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -35,6 +35,9 @@
#define RAN_VAR_STACK 0x40000
#define RAN_GLOBAL_PAD 0x80000
+/* progs/exceptions_cleanup_light.c: the one pad it has. */
+#define RAN_LIGHT 0x1
+
#define CLEANUP_REC(begin, end, landing_pad) \
".pushsection .bpf_cleanup,\"a\",@progbits;" \
".long " begin ";" \
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
index 2825dbb53b1c..679cf40d1707 100644
--- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -8,6 +8,7 @@
#include "exceptions_cleanup_freplace.skel.h"
#include "exceptions_cleanup_pad_freplace.skel.h"
#include "exceptions_cleanup_ext_table.skel.h"
+#include "exceptions_cleanup_light.lskel.h"
/* foo3 threw: every frame that has a pad ran it. */
#define PADS_FOO3_THREW \
@@ -171,6 +172,30 @@ static void test_ext_table(struct exceptions_cleanup_shapes *skel)
exceptions_cleanup_ext_table__destroy(fr);
}
+static void test_light_skeleton(void)
+{
+ struct exceptions_cleanup_light_lskel *skel;
+ __u64 ctx = 0;
+ int err;
+
+ LIBBPF_OPTS(bpf_test_run_opts, topts,
+ .ctx_in = &ctx,
+ .ctx_size_in = sizeof(ctx),
+ );
+
+ skel = exceptions_cleanup_light_lskel__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "light open_and_load"))
+ return;
+
+ err = bpf_prog_test_run_opts(skel->progs.entry_light.prog_fd, &topts);
+ if (!ASSERT_OK(err, "run"))
+ goto out;
+ ASSERT_EQ(topts.retval, THROW_COOKIE, "retval");
+ ASSERT_EQ(skel->bss->pads_ran, RAN_LIGHT, "pads_ran");
+out:
+ exceptions_cleanup_light_lskel__destroy(skel);
+}
+
static void test_shapes(void)
{
struct exceptions_cleanup_shapes *skel;
@@ -458,6 +483,9 @@ void test_exceptions_cleanup(void)
exceptions_cleanup__destroy(skel);
+ if (test__start_subtest("light_skeleton"))
+ test_light_skeleton();
+
test_shapes();
RUN_TESTS(exceptions_cleanup_fail);
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
new file mode 100644
index 000000000000..de7ae478d67f
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
@@ -0,0 +1,39 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+ bpf_throw(0);
+ bpf_preempt_disable();
+ bpf_preempt_enable();
+ bpf_unwind_resume();
+}
+
+__u64 pads_ran = 0;
+
+SEC("syscall")
+__naked int entry_light(void)
+{
+ asm volatile (
+ "call bpf_preempt_disable;"
+ "r1 = %[cookie];"
+"1:" "call bpf_throw;" /* cleanup region */
+"2:"
+ "exit;"
+"3:" /* landing pad */
+ "call bpf_preempt_enable;"
+ PAD_RAN("%[ran]")
+ "call bpf_unwind_resume;"
+ "exit;"
+ CLEANUP_REC("1b", "2b", "3b")
+ :
+ : [cookie]"i"(THROW_COOKIE), [ran]"i"(RAN_LIGHT),
+ __imm_addr(pads_ran)
+ : __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-21 21:00 [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
` (19 preceding siblings ...)
2026-09-21 21:02 ` [PATCH bpf-next v4 20/20] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
@ 2026-09-22 1:08 ` Eduard Zingerman
2026-09-22 2:16 ` Alexei Starovoitov
20 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-22 1:08 UTC (permalink / raw)
To: Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Mon, 2026-09-21 at 14:00 -0700, Yonghong Song wrote:
...
> Design
> ======
>
> A pad is run, not lowered. bpf_throw() already walks the frames with
> arch_bpf_stack_walk(); it now looks each frame's return address up in
> that (sub)program's table and calls the pad as a subroutine of the walker,
> with the unwinding frame's frame pointer and its callee-saved registers
> restored from the spill its callee's prologue left. The pad therefore sees
> its own frame but runs on the walker's stack, far below it, so nothing it
> calls can disturb the frame it is cleaning up after. The JIT turns its
> bpf_unwind_resume() into the way back to the walker.
>
> The verifier walks the same thing, step for step, so the resource rules
> are unchanged: whatever a pad releases is released in the verifier state
> too, and check_resource_leak() simply moves from "a throw was seen" to the
> end of the walk.
I have two high-level questions.
1) The tables handling mechanics adds quite a lot of code to the
libbpf and initial verification phases, while at the IR level
it is basically an encoding for the invoke instruction:
invoke <target-function>
to label <where-to-go-on-return>
unwind label <where-to-go-on-unwind>
For the sake of discussion, wouldn't it be simpler for us to
just add a 16-byte invoke instruction:
word #0:
code INVOKE
dst_reg 0
src_reg BPF_PSEUDO_CALL or BPF_PSEUDO_KFUNC_CALL
off existing call meaning, including kfunc BTF fd index
imm existing call-target encoding
word #1:
code, dst_reg, src_reg, off = 0
imm signed unwind displacement, measured in 8-byte slots
With an assumption that during normal execution (not unwinding)
upon return from invoke the control flow goes to a fallthrough
instruction.
The pros are:
- much less frontend code
- if in the future we would like to manipulate BPF program
byte code, it would be significantly simpler to do in such form.
[1] https://llvm.org/docs/LangRef.html#i-invoke
2) The final goal of the BPF/Rust project is to consume whatever code
rustc generates. Ultimately, this would require supporting a way
to introduce runtime checks at arbitrary locations, whenever the
verifier can't infer that the program is safe. Such checks won't
necessarily have associated landing pads. Meaning that the unwinding
logic will have to be dynamic as in exceptions part #2 sense discussed
way back (2022?).
The argument against exceptions part #2 back then was that the code
is complex. Given the current environment, I don't think the argument
still holds.
Hence, given that dynamic abort would be necessary, and that the kernel's
rust code is already compiled with -Cpanic=abort, do we need this static
form of exceptions handling at all?
...
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-22 1:08 ` [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Eduard Zingerman
@ 2026-09-22 2:16 ` Alexei Starovoitov
2026-09-22 2:31 ` Kumar Kartikeya Dwivedi
2026-09-22 4:27 ` Eduard Zingerman
0 siblings, 2 replies; 80+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 2:16 UTC (permalink / raw)
To: Eduard Zingerman, Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Tue Sep 22, 2026 at 1:08 AM UTC, Eduard Zingerman wrote:
> On Mon, 2026-09-21 at 14:00 -0700, Yonghong Song wrote:
>
> ...
>
>> Design
>> ======
>>
>> A pad is run, not lowered. bpf_throw() already walks the frames with
>> arch_bpf_stack_walk(); it now looks each frame's return address up in
>> that (sub)program's table and calls the pad as a subroutine of the walker,
>> with the unwinding frame's frame pointer and its callee-saved registers
>> restored from the spill its callee's prologue left. The pad therefore sees
>> its own frame but runs on the walker's stack, far below it, so nothing it
>> calls can disturb the frame it is cleaning up after. The JIT turns its
>> bpf_unwind_resume() into the way back to the walker.
>>
>> The verifier walks the same thing, step for step, so the resource rules
>> are unchanged: whatever a pad releases is released in the verifier state
>> too, and check_resource_leak() simply moves from "a throw was seen" to the
>> end of the walk.
>
> I have two high-level questions.
>
> 1) The tables handling mechanics adds quite a lot of code to the
> libbpf and initial verification phases, while at the IR level
> it is basically an encoding for the invoke instruction:
>
> invoke <target-function>
> to label <where-to-go-on-return>
> unwind label <where-to-go-on-unwind>
>
> For the sake of discussion, wouldn't it be simpler for us to
> just add a 16-byte invoke instruction:
>
> word #0:
> code INVOKE
> dst_reg 0
> src_reg BPF_PSEUDO_CALL or BPF_PSEUDO_KFUNC_CALL
> off existing call meaning, including kfunc BTF fd index
> imm existing call-target encoding
>
> word #1:
> code, dst_reg, src_reg, off = 0
> imm signed unwind displacement, measured in 8-byte slots
>
> With an assumption that during normal execution (not unwinding)
> upon return from invoke the control flow goes to a fallthrough
> instruction.
>
> The pros are:
> - much less frontend code
> - if in the future we would like to manipulate BPF program
> byte code, it would be significantly simpler to do in such form.
I think the amount of code will increase a lot more with such approach.
All existing call flavors and my new callx would need to wrapped
with this new 'invoke' insn.
Also rust generates begin/end across more than single insn.
> [1] https://llvm.org/docs/LangRef.html#i-invoke
>
> 2) The final goal of the BPF/Rust project is to consume whatever code
> rustc generates. Ultimately, this would require supporting a way
> to introduce runtime checks at arbitrary locations, whenever the
> verifier can't infer that the program is safe.
so far all rustc code looks clean and definitely not arbitrary.
> Such checks won't
> necessarily have associated landing pads. Meaning that the unwinding
no landing pad? what? That's not rustc.
You're talking about some random compiler that throws garbage.
> logic will have to be dynamic as in exceptions part #2 sense discussed
> way back (2022?).
>
> The argument against exceptions part #2 back then was that the code
> is complex. Given the current environment, I don't think the argument
> still holds.
>
> Hence, given that dynamic abort would be necessary, and that the kernel's
> rust code is already compiled with -Cpanic=abort, do we need this static
> form of exceptions handling at all?
This patch is static exception handling for cases where compiler generated them.
We don't know and don't care what is in those landing pads.
rustc maybe cleaning up the objects that have no meaning for the verifier.
Maybe freeing memory (but since it's arena) we don't care,
but we will still call all the drop()s because that's what rust as a language
promised to users and we cannot break that promise.
kernel's rust with panic=abort needs to be fixed.
That's orthogonal problem and definitely not something to follow.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-22 2:16 ` Alexei Starovoitov
@ 2026-09-22 2:31 ` Kumar Kartikeya Dwivedi
2026-09-22 21:44 ` Alexei Starovoitov
2026-09-22 4:27 ` Eduard Zingerman
1 sibling, 1 reply; 80+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-22 2:31 UTC (permalink / raw)
To: Alexei Starovoitov, Eduard Zingerman, Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Tue Sep 22, 2026 at 4:16 AM CEST, Alexei Starovoitov wrote:
> On Tue Sep 22, 2026 at 1:08 AM UTC, Eduard Zingerman wrote:
>> On Mon, 2026-09-21 at 14:00 -0700, Yonghong Song wrote:
>>
>> ...
>>
>>> Design
>>> ======
>>>
>>> A pad is run, not lowered. bpf_throw() already walks the frames with
>>> arch_bpf_stack_walk(); it now looks each frame's return address up in
>>> that (sub)program's table and calls the pad as a subroutine of the walker,
>>> with the unwinding frame's frame pointer and its callee-saved registers
>>> restored from the spill its callee's prologue left. The pad therefore sees
>>> its own frame but runs on the walker's stack, far below it, so nothing it
>>> calls can disturb the frame it is cleaning up after. The JIT turns its
>>> bpf_unwind_resume() into the way back to the walker.
>>>
>>> The verifier walks the same thing, step for step, so the resource rules
>>> are unchanged: whatever a pad releases is released in the verifier state
>>> too, and check_resource_leak() simply moves from "a throw was seen" to the
>>> end of the walk.
>>
>> I have two high-level questions.
>>
>> 1) The tables handling mechanics adds quite a lot of code to the
>> libbpf and initial verification phases, while at the IR level
>> it is basically an encoding for the invoke instruction:
>>
>> invoke <target-function>
>> to label <where-to-go-on-return>
>> unwind label <where-to-go-on-unwind>
>>
>> For the sake of discussion, wouldn't it be simpler for us to
>> just add a 16-byte invoke instruction:
>>
>> word #0:
>> code INVOKE
>> dst_reg 0
>> src_reg BPF_PSEUDO_CALL or BPF_PSEUDO_KFUNC_CALL
>> off existing call meaning, including kfunc BTF fd index
>> imm existing call-target encoding
>>
>> word #1:
>> code, dst_reg, src_reg, off = 0
>> imm signed unwind displacement, measured in 8-byte slots
>>
>> With an assumption that during normal execution (not unwinding)
>> upon return from invoke the control flow goes to a fallthrough
>> instruction.
>>
>> The pros are:
>> - much less frontend code
>> - if in the future we would like to manipulate BPF program
>> byte code, it would be significantly simpler to do in such form.
>
> I think the amount of code will increase a lot more with such approach.
> All existing call flavors and my new callx would need to wrapped
> with this new 'invoke' insn.
> Also rust generates begin/end across more than single insn.
>
>> [1] https://llvm.org/docs/LangRef.html#i-invoke
>>
>> 2) The final goal of the BPF/Rust project is to consume whatever code
>> rustc generates. Ultimately, this would require supporting a way
>> to introduce runtime checks at arbitrary locations, whenever the
>> verifier can't infer that the program is safe.
>
> so far all rustc code looks clean and definitely not arbitrary.
>
>> Such checks won't
>> necessarily have associated landing pads. Meaning that the unwinding
>
> no landing pad? what? That's not rustc.
> You're talking about some random compiler that throws garbage.
>
I think you misunderstood his question.
What Eduard meant or is asking, was that (per his understanding), the verifier,
in the future, would attempt to defer to runtime checks where it cannot prove
certain patterns in the program safe, regardless of how they came to be (in C,
or Rust). Instead of rejecting the program from loading, in such a case, we
would emit some form of a runtime assertion that, for the remainder of the
program, guarantees a given value to be in the range the verifier accepts, as a
very simple example.
In such a case, the assertion and associated panic is being injected after
compilation, hence it is likely there is no associated landing pad to fall back
to and clean up program resources. Yet to safety inject such a runtime assertion
clean up would be necessary.
>> logic will have to be dynamic as in exceptions part #2 sense discussed
>> way back (2022?).
>>
>> The argument against exceptions part #2 back then was that the code
>> is complex. Given the current environment, I don't think the argument
>> still holds.
>>
>> Hence, given that dynamic abort would be necessary, and that the kernel's
>> rust code is already compiled with -Cpanic=abort, do we need this static
>> form of exceptions handling at all?
>
> This patch is static exception handling for cases where compiler generated them.
>
> We don't know and don't care what is in those landing pads.
> rustc maybe cleaning up the objects that have no meaning for the verifier.
> Maybe freeing memory (but since it's arena) we don't care,
> but we will still call all the drop()s because that's what rust as a language
> promised to users and we cannot break that promise.
I do think supporting invocation of landing pads provided by the program makes
sense, regardless of the discussion above.
> [...]
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-22 2:31 ` Kumar Kartikeya Dwivedi
@ 2026-09-22 21:44 ` Alexei Starovoitov
2026-09-23 4:36 ` Kumar Kartikeya Dwivedi
0 siblings, 1 reply; 80+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 21:44 UTC (permalink / raw)
To: Kumar Kartikeya Dwivedi, Eduard Zingerman, Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Tue Sep 22, 2026 at 2:31 AM UTC, Kumar Kartikeya Dwivedi wrote:
>
> I think you misunderstood his question.
>
> What Eduard meant or is asking, was that (per his understanding), the verifier,
> in the future, would attempt to defer to runtime checks where it cannot prove
> certain patterns in the program safe, regardless of how they came to be (in C,
> or Rust). Instead of rejecting the program from loading, in such a case, we
> would emit some form of a runtime assertion that, for the remainder of the
> program, guarantees a given value to be in the range the verifier accepts, as a
> very simple example.
>
> In such a case, the assertion and associated panic is being injected after
> compilation, hence it is likely there is no associated landing pad to fall back
> to and clean up program resources. Yet to safety inject such a runtime assertion
> clean up would be necessary.
In this case presence of landing pads from rust is especially important.
The verifier would need to rely on them to insert additional run-time checks:
if something bad between [ip_start, ip_end] goto landing_pad.
It's not possible to insert control flow edge at random points,
since the verifier cannot see original rust code and its meaning.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-22 21:44 ` Alexei Starovoitov
@ 2026-09-23 4:36 ` Kumar Kartikeya Dwivedi
2026-09-23 4:54 ` Alexei Starovoitov
0 siblings, 1 reply; 80+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-23 4:36 UTC (permalink / raw)
To: Alexei Starovoitov
Cc: Eduard Zingerman, Yonghong Song, bpf, Alexei Starovoitov,
Andrii Nakryiko, Daniel Borkmann, kernel-team
On Tue, 22 Sept 2026 at 23:44, Alexei Starovoitov
<alexei.starovoitov@gmail.com> wrote:
>
> On Tue Sep 22, 2026 at 2:31 AM UTC, Kumar Kartikeya Dwivedi wrote:
> >
> > I think you misunderstood his question.
> >
> > What Eduard meant or is asking, was that (per his understanding), the verifier,
> > in the future, would attempt to defer to runtime checks where it cannot prove
> > certain patterns in the program safe, regardless of how they came to be (in C,
> > or Rust). Instead of rejecting the program from loading, in such a case, we
> > would emit some form of a runtime assertion that, for the remainder of the
> > program, guarantees a given value to be in the range the verifier accepts, as a
> > very simple example.
> >
> > In such a case, the assertion and associated panic is being injected after
> > compilation, hence it is likely there is no associated landing pad to fall back
> > to and clean up program resources. Yet to safety inject such a runtime assertion
> > clean up would be necessary.
>
> In this case presence of landing pads from rust is especially important.
> The verifier would need to rely on them to insert additional run-time checks:
> if something bad between [ip_start, ip_end] goto landing_pad.
> It's not possible to insert control flow edge at random points,
> since the verifier cannot see original rust code and its meaning.
>
I think you're conflating several aspects again.
You can obviously insert aborts at any point in the program, provided
such a primitive works, even when the compiler doesn't see it. It is
just a way to halt program execution along a given path, and has
plenty of precedents (assert(false), std::terminate(), panic!() =
abort). That property can be used for several purposes, including
proving a condition true on the other path that does not abort, and
retaining that path condition throughout the rest of the program. That
is basically the gist of Eduard's suggestion.
Injecting random gotos and jumps to some other part of the program is
problematic, of course, and for that, your explanation above is valid.
However, we weren't discussing those.
Anyhow, let's keep focus on the current patch set and continue with
it. We need the ability to invoke user-provided landing pads in any
case.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-23 4:36 ` Kumar Kartikeya Dwivedi
@ 2026-09-23 4:54 ` Alexei Starovoitov
2026-09-23 5:20 ` Kumar Kartikeya Dwivedi
2026-09-23 6:16 ` Eduard Zingerman
0 siblings, 2 replies; 80+ messages in thread
From: Alexei Starovoitov @ 2026-09-23 4:54 UTC (permalink / raw)
To: Kumar Kartikeya Dwivedi
Cc: Eduard Zingerman, Yonghong Song, bpf, Alexei Starovoitov,
Andrii Nakryiko, Daniel Borkmann, kernel-team
On Wed Sep 23, 2026 at 4:36 AM UTC, Kumar Kartikeya Dwivedi wrote:
>
> You can obviously insert aborts at any point in the program, provided
> such a primitive works, even when the compiler doesn't see it. It is
> just a way to halt program execution along a given path, and has
> plenty of precedents (assert(false), std::terminate(), panic!() =
> abort). That property can be used for several purposes, including
> proving a condition true on the other path that does not abort, and
> retaining that path condition throughout the rest of the program. That
> is basically the gist of Eduard's suggestion.
That's only true for user space.
For bpf progs there is no such primitive. bpf_throw() is not it.
We cannot make it work from arbitrary places without introducing
massive verifier debt for automatic creation of exception tables
or via equally massive runtime penalty to remember all things to cleanup.
We're not going to support rust panic=abort for the same reasons.
If rust-bpf prog is compiled like that it will likely be rejected by the verifier.
Accept-all-rust doesn't mean accept rust that can leak resources.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-23 4:54 ` Alexei Starovoitov
@ 2026-09-23 5:20 ` Kumar Kartikeya Dwivedi
2026-09-23 6:16 ` Eduard Zingerman
1 sibling, 0 replies; 80+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-23 5:20 UTC (permalink / raw)
To: Alexei Starovoitov
Cc: Eduard Zingerman, Yonghong Song, bpf, Alexei Starovoitov,
Andrii Nakryiko, Daniel Borkmann, kernel-team
On Wed Sep 23, 2026 at 6:54 AM CEST, Alexei Starovoitov wrote:
> On Wed Sep 23, 2026 at 4:36 AM UTC, Kumar Kartikeya Dwivedi wrote:
>>
> [...]
> We're not going to support rust panic=abort for the same reasons.
> If rust-bpf prog is compiled like that it will likely be rejected by the verifier.
> Accept-all-rust doesn't mean accept rust that can leak resources.
We might find ourselves in situation where panic!() is present in code invoked
during Drop or unwinding, in case any runtime checked Rust primitive is used
(Rc, RefCell::borrow_mut(), indexing into slices, etc.). That usually translates
to complete termination.
Code is written in such a way that it won't panic!() at that point due to
various invariants, but the call would still be present.
I am not sure we can simply blame the user in such cases, since mere presence of
panic!() does not mean the panic!() occurs in practice, even though the pattern
compiles down to it wherever that primitive is used. We will simply end up
rejecting the presence, and in turn rejecting the rest of the program.
That said, we can still tackle all this later, so it's fine to continue for now,
I think.
I am not even sure what actual programs will end up looking in practice. That
said, it is something to keep in mind instead of being dismissive.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-23 4:54 ` Alexei Starovoitov
2026-09-23 5:20 ` Kumar Kartikeya Dwivedi
@ 2026-09-23 6:16 ` Eduard Zingerman
2026-09-23 6:44 ` Kumar Kartikeya Dwivedi
1 sibling, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-23 6:16 UTC (permalink / raw)
To: Alexei Starovoitov, Kumar Kartikeya Dwivedi
Cc: Yonghong Song, bpf, Alexei Starovoitov, Andrii Nakryiko,
Daniel Borkmann, kernel-team
On Wed, 2026-09-23 at 04:54 +0000, Alexei Starovoitov wrote:
> On Wed Sep 23, 2026 at 4:36 AM UTC, Kumar Kartikeya Dwivedi wrote:
> >
> > You can obviously insert aborts at any point in the program, provided
> > such a primitive works, even when the compiler doesn't see it. It is
> > just a way to halt program execution along a given path, and has
> > plenty of precedents (assert(false), std::terminate(), panic!() =
> > abort). That property can be used for several purposes, including
> > proving a condition true on the other path that does not abort, and
> > retaining that path condition throughout the rest of the program. That
> > is basically the gist of Eduard's suggestion.
>
> That's only true for user space.
> For bpf progs there is no such primitive. bpf_throw() is not it.
> We cannot make it work from arbitrary places without introducing
> massive verifier debt for automatic creation of exception tables
> or via equally massive runtime penalty to remember all things to cleanup.
I'm not sure that automatic exception tables would be all that more
complex, current implementation is not that trivial either.
Plus you mention LLMs-the-almighty yourself.
But speaking of runtime costs.
For a program like this:
main:
a()
a:
b()
b:
throw()
The series currently generates push r6-r9 at entry to each function.
This is needed to recover r6-r9 for the landing pads.
See the program and disassembly below.
Do we want to address this somehow, or is the idea that for complex
subprograms the sequence would amortize away?
On x86 the information about registers location at throw/landing-pad-entry
is stored in .eh_frame section, as far as I understand.
BPF:
jit_probe_b:
r1 = 32;
1: call bpf_throw;
2: r0 = 0;
exit;
3: call bpf_unwind_resume;
exit;
CLEANUP_REC(1b, 2b, 3b)
jit_probe_a:
r6 = 0x1234;
1: call jit_probe_b;
2: r0 = 0;
exit;
3: r1 = r6;
call bpf_unwind_resume;
exit;
CLEANUP_REC(1b, 2b, 3b)
jit_probe_main:
jit_probe_a();
return 0;
x86:
jit_probe_main:
0: endbr64
4: nopl (%rax,%rax)
9: nopl (%rax)
c: pushq %rbp
d: movq %rsp, %rbp
10: endbr64
14: pushq %r12
16: pushq %rbx
17: pushq %r13
19: pushq %r14
1b: pushq %r15
1d: callq 0x30
jit_probe_a:
0: endbr64
4: nopl (%rax,%rax)
9: nopl (%rax)
c: pushq %rbp
d: movq %rsp, %rbp
10: endbr64
14: pushq %r12
16: pushq %rbx
17: pushq %r13
19: pushq %r14
1b: pushq %r15
1d: movl $0x1234, %ebx
22: callq 0x9c
27: endbr64
2b: movq %rbx, %rdi
2e: retq
jit_probe_b:
0: endbr64
4: nopl (%rax,%rax)
9: nopl (%rax)
c: pushq %rbp
d: movq %rsp, %rbp
10: endbr64
14: subq $0x28, %rsp
1b: pushq %r12
1d: pushq %rbx
1e: pushq %r13
20: pushq %r14
22: pushq %r15
24: movl $0x20, %edi
29: movq %r15, -0x28(%rbp)
2d: movq %r14, -0x20(%rbp)
31: movq %r13, -0x18(%rbp)
35: movq %rbx, -0x10(%rbp)
39: movq %r12, -0x8(%rbp)
3d: callq 0xffffffffe18c5c5c
42: endbr64
46: retq
...
^ permalink raw reply [flat|nested] 80+ messages in thread* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-23 6:16 ` Eduard Zingerman
@ 2026-09-23 6:44 ` Kumar Kartikeya Dwivedi
0 siblings, 0 replies; 80+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-23 6:44 UTC (permalink / raw)
To: Eduard Zingerman, Alexei Starovoitov
Cc: Yonghong Song, bpf, Alexei Starovoitov, Andrii Nakryiko,
Daniel Borkmann, kernel-team
On Wed Sep 23, 2026 at 8:16 AM CEST, Eduard Zingerman wrote:
> On Wed, 2026-09-23 at 04:54 +0000, Alexei Starovoitov wrote:
>> On Wed Sep 23, 2026 at 4:36 AM UTC, Kumar Kartikeya Dwivedi wrote:
>> >
>> > You can obviously insert aborts at any point in the program, provided
>> > such a primitive works, even when the compiler doesn't see it. It is
>> > just a way to halt program execution along a given path, and has
>> > plenty of precedents (assert(false), std::terminate(), panic!() =
>> > abort). That property can be used for several purposes, including
>> > proving a condition true on the other path that does not abort, and
>> > retaining that path condition throughout the rest of the program. That
>> > is basically the gist of Eduard's suggestion.
>>
>> That's only true for user space.
>> For bpf progs there is no such primitive. bpf_throw() is not it.
>> We cannot make it work from arbitrary places without introducing
>> massive verifier debt for automatic creation of exception tables
>> or via equally massive runtime penalty to remember all things to cleanup.
>
> I'm not sure that automatic exception tables would be all that more
> complex, current implementation is not that trivial either.
> Plus you mention LLMs-the-almighty yourself.
>
> But speaking of runtime costs.
> For a program like this:
>
> main:
> a()
> a:
> b()
> b:
> throw()
>
> The series currently generates push r6-r9 at entry to each function.
> This is needed to recover r6-r9 for the landing pads.
> See the program and disassembly below.
> Do we want to address this somehow, or is the idea that for complex
> subprograms the sequence would amortize away?
> On x86 the information about registers location at throw/landing-pad-entry
> is stored in .eh_frame section, as far as I understand.
>
> BPF:
>
> jit_probe_b:
> r1 = 32;
> 1: call bpf_throw;
> 2: r0 = 0;
> exit;
> 3: call bpf_unwind_resume;
> exit;
> CLEANUP_REC(1b, 2b, 3b)
>
> jit_probe_a:
> r6 = 0x1234;
> 1: call jit_probe_b;
> 2: r0 = 0;
> exit;
> 3: r1 = r6;
> call bpf_unwind_resume;
> exit;
> CLEANUP_REC(1b, 2b, 3b)
>
> jit_probe_main:
> jit_probe_a();
> return 0;
>
> x86:
>
> jit_probe_main:
> 0: endbr64
> 4: nopl (%rax,%rax)
> 9: nopl (%rax)
> c: pushq %rbp
> d: movq %rsp, %rbp
> 10: endbr64
> 14: pushq %r12
> 16: pushq %rbx
> 17: pushq %r13
> 19: pushq %r14
> 1b: pushq %r15
> 1d: callq 0x30
>
> jit_probe_a:
> 0: endbr64
> 4: nopl (%rax,%rax)
> 9: nopl (%rax)
> c: pushq %rbp
> d: movq %rsp, %rbp
> 10: endbr64
> 14: pushq %r12
> 16: pushq %rbx
> 17: pushq %r13
> 19: pushq %r14
> 1b: pushq %r15
> 1d: movl $0x1234, %ebx
> 22: callq 0x9c
> 27: endbr64
> 2b: movq %rbx, %rdi
> 2e: retq
>
> jit_probe_b:
> 0: endbr64
> 4: nopl (%rax,%rax)
> 9: nopl (%rax)
> c: pushq %rbp
> d: movq %rsp, %rbp
> 10: endbr64
> 14: subq $0x28, %rsp
> 1b: pushq %r12
> 1d: pushq %rbx
> 1e: pushq %r13
> 20: pushq %r14
> 22: pushq %r15
> 24: movl $0x20, %edi
> 29: movq %r15, -0x28(%rbp)
> 2d: movq %r14, -0x20(%rbp)
> 31: movq %r13, -0x18(%rbp)
> 35: movq %rbx, -0x10(%rbp)
> 39: movq %r12, -0x8(%rbp)
> 3d: callq 0xffffffffe18c5c5c
> 42: endbr64
> 46: retq
>
That is really bad codegen. There might be a way to argue it's slightly ok for
x86 in terms of cost in prologue, but it looks really bad for arm64. I think we
need to find a better way of handling this...
Basically on entry to a subprog on arm64 we spill 80 bytes worth of registers,
ignoring spills around throw site which might be less problematic.
> ...
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-22 2:16 ` Alexei Starovoitov
2026-09-22 2:31 ` Kumar Kartikeya Dwivedi
@ 2026-09-22 4:27 ` Eduard Zingerman
2026-09-22 21:47 ` Alexei Starovoitov
1 sibling, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-22 4:27 UTC (permalink / raw)
To: Alexei Starovoitov, Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Tue, 2026-09-22 at 02:16 +0000, Alexei Starovoitov wrote:
> On Tue Sep 22, 2026 at 1:08 AM UTC, Eduard Zingerman wrote:
> > On Mon, 2026-09-21 at 14:00 -0700, Yonghong Song wrote:
> >
> > ...
> >
> > > Design
> > > ======
> > >
> > > A pad is run, not lowered. bpf_throw() already walks the frames with
> > > arch_bpf_stack_walk(); it now looks each frame's return address up in
> > > that (sub)program's table and calls the pad as a subroutine of the walker,
> > > with the unwinding frame's frame pointer and its callee-saved registers
> > > restored from the spill its callee's prologue left. The pad therefore sees
> > > its own frame but runs on the walker's stack, far below it, so nothing it
> > > calls can disturb the frame it is cleaning up after. The JIT turns its
> > > bpf_unwind_resume() into the way back to the walker.
> > >
> > > The verifier walks the same thing, step for step, so the resource rules
> > > are unchanged: whatever a pad releases is released in the verifier state
> > > too, and check_resource_leak() simply moves from "a throw was seen" to the
> > > end of the walk.
> >
> > I have two high-level questions.
> >
> > 1) The tables handling mechanics adds quite a lot of code to the
> > libbpf and initial verification phases, while at the IR level
> > it is basically an encoding for the invoke instruction:
> >
> > invoke <target-function>
> > to label <where-to-go-on-return>
> > unwind label <where-to-go-on-unwind>
> >
> > For the sake of discussion, wouldn't it be simpler for us to
> > just add a 16-byte invoke instruction:
> >
> > word #0:
> > code INVOKE
> > dst_reg 0
> > src_reg BPF_PSEUDO_CALL or BPF_PSEUDO_KFUNC_CALL
> > off existing call meaning, including kfunc BTF fd index
> > imm existing call-target encoding
> >
> > word #1:
> > code, dst_reg, src_reg, off = 0
> > imm signed unwind displacement, measured in 8-byte slots
> >
> > With an assumption that during normal execution (not unwinding)
> > upon return from invoke the control flow goes to a fallthrough
> > instruction.
> >
> > The pros are:
> > - much less frontend code
> > - if in the future we would like to manipulate BPF program
> > byte code, it would be significantly simpler to do in such form.
>
> I think the amount of code will increase a lot more with such approach.
> All existing call flavors and my new callx would need to wrapped
> with this new 'invoke' insn.
> Also rust generates begin/end across more than single insn.
Total size of executable sections for all Meta BPF object files used
for CI veristat testing is 10Mb. Of these there are 30K non-helper
call instructions in total.
So we are talking about increase by 8 * 30K = 240K ~ 2.4% worst case.
Note that the instruction encoding optimized for size already hearts
us in src/dst registers department: we have no room for virtual
registers, which would have simplified e.g. register allocation task
for ARM64. Point being that optimizing intermediate IR for size is not
always a right target.
> > [1] https://llvm.org/docs/LangRef.html#i-invoke
> >
> > 2) The final goal of the BPF/Rust project is to consume whatever code
> > rustc generates. Ultimately, this would require supporting a way
> > to introduce runtime checks at arbitrary locations, whenever the
> > verifier can't infer that the program is safe.
>
> so far all rustc code looks clean and definitely not arbitrary.
>
> > Such checks won't
> > necessarily have associated landing pads. Meaning that the unwinding
>
> no landing pad? what? That's not rustc.
> You're talking about some random compiler that throws garbage.
>
> > logic will have to be dynamic as in exceptions part #2 sense discussed
> > way back (2022?).
> >
> > The argument against exceptions part #2 back then was that the code
> > is complex. Given the current environment, I don't think the argument
> > still holds.
> >
> > Hence, given that dynamic abort would be necessary, and that the kernel's
> > rust code is already compiled with -Cpanic=abort, do we need this static
> > form of exceptions handling at all?
>
> This patch is static exception handling for cases where compiler generated them.
>
> We don't know and don't care what is in those landing pads.
> rustc maybe cleaning up the objects that have no meaning for the verifier.
> Maybe freeing memory (but since it's arena) we don't care,
> but we will still call all the drop()s because that's what rust as a language
> promised to users and we cannot break that promise.
> kernel's rust with panic=abort needs to be fixed.
> That's orthogonal problem and definitely not something to follow.
As Kartikeya says in a sibling email, I mean generic verifier
capability to accept any program by introducing runtime checks.
Verifier's capacity to infer safety conditions for various
instructions is orthogonal to rustc's assumptions about landing pad
boundaries. E.g. rustc/llvm might infer that certain memory access is
within bounds and optimize bound checks out, it is not a given that
verifier would come to a same conclusion and that some landing pad
would be declared for an instruction.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-22 4:27 ` Eduard Zingerman
@ 2026-09-22 21:47 ` Alexei Starovoitov
2026-09-22 23:08 ` Eduard Zingerman
0 siblings, 1 reply; 80+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 21:47 UTC (permalink / raw)
To: Eduard Zingerman, Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Tue Sep 22, 2026 at 4:27 AM UTC, Eduard Zingerman wrote:
>
> Total size of executable sections for all Meta BPF object files used
> for CI veristat testing is 10Mb. Of these there are 30K non-helper
> call instructions in total.
> So we are talking about increase by 8 * 30K = 240K ~ 2.4% worst case.
> Note that the instruction encoding optimized for size already hearts
> us in src/dst registers department: we have no room for virtual
> registers, which would have simplified e.g. register allocation task
> for ARM64. Point being that optimizing intermediate IR for size is not
> always a right target.
I wasn't talking about increase in bpf ELF size,
but the amount of the verifier work necessary to double the number
of call insns.
> As Kartikeya says in a sibling email, I mean generic verifier
> capability to accept any program by introducing runtime checks.
> Verifier's capacity to infer safety conditions for various
> instructions is orthogonal to rustc's assumptions about landing pad
> boundaries. E.g. rustc/llvm might infer that certain memory access is
> within bounds and optimize bound checks out, it is not a given that
> verifier would come to a same conclusion and that some landing pad
> would be declared for an instruction.
Replied to Kumar. I don't think the verifier will ever insert
runtime checks with control flow at random points.
If we teach rustc to insert may_goto then it will be done
at rustc/llvm level. Not by the verifier.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-22 21:47 ` Alexei Starovoitov
@ 2026-09-22 23:08 ` Eduard Zingerman
2026-09-22 23:37 ` Alexei Starovoitov
0 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-22 23:08 UTC (permalink / raw)
To: Alexei Starovoitov, Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Tue, 2026-09-22 at 21:47 +0000, Alexei Starovoitov wrote:
> On Tue Sep 22, 2026 at 4:27 AM UTC, Eduard Zingerman wrote:
> >
> > Total size of executable sections for all Meta BPF object files used
> > for CI veristat testing is 10Mb. Of these there are 30K non-helper
> > call instructions in total.
> > So we are talking about increase by 8 * 30K = 240K ~ 2.4% worst case.
> > Note that the instruction encoding optimized for size already hearts
> > us in src/dst registers department: we have no room for virtual
> > registers, which would have simplified e.g. register allocation task
> > for ARM64. Point being that optimizing intermediate IR for size is not
> > always a right target.
>
> I wasn't talking about increase in bpf ELF size,
> but the amount of the verifier work necessary to double the number
> of call insns.
I don't understand what you mean by the amount of the verifier work.
Traced program paths would remain the same, only the encoding of the
information about exceptions changes.
> > As Kartikeya says in a sibling email, I mean generic verifier
> > capability to accept any program by introducing runtime checks.
> > Verifier's capacity to infer safety conditions for various
> > instructions is orthogonal to rustc's assumptions about landing pad
> > boundaries. E.g. rustc/llvm might infer that certain memory access is
> > within bounds and optimize bound checks out, it is not a given that
> > verifier would come to a same conclusion and that some landing pad
> > would be declared for an instruction.
>
> Replied to Kumar. I don't think the verifier will ever insert
> runtime checks with control flow at random points.
> If we teach rustc to insert may_goto then it will be done
> at rustc/llvm level. Not by the verifier.
Fighting the verifier after rustc transformations would be especially fun.
Do we have a working toolchain somewhere? I can work a few examples.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-22 23:08 ` Eduard Zingerman
@ 2026-09-22 23:37 ` Alexei Starovoitov
2026-09-23 0:04 ` Eduard Zingerman
0 siblings, 1 reply; 80+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 23:37 UTC (permalink / raw)
To: Eduard Zingerman, Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Tue Sep 22, 2026 at 11:08 PM UTC, Eduard Zingerman wrote:
> On Tue, 2026-09-22 at 21:47 +0000, Alexei Starovoitov wrote:
>> On Tue Sep 22, 2026 at 4:27 AM UTC, Eduard Zingerman wrote:
>> >
>> > Total size of executable sections for all Meta BPF object files used
>> > for CI veristat testing is 10Mb. Of these there are 30K non-helper
>> > call instructions in total.
>> > So we are talking about increase by 8 * 30K = 240K ~ 2.4% worst case.
>> > Note that the instruction encoding optimized for size already hearts
>> > us in src/dst registers department: we have no room for virtual
>> > registers, which would have simplified e.g. register allocation task
>> > for ARM64. Point being that optimizing intermediate IR for size is not
>> > always a right target.
>>
>> I wasn't talking about increase in bpf ELF size,
>> but the amount of the verifier work necessary to double the number
>> of call insns.
>
> I don't understand what you mean by the amount of the verifier work.
> Traced program paths would remain the same, only the encoding of the
> information about exceptions changes.
Any new insn requires plumbing through out.
imo ld_imm64 encoding FD inline was a mistake.
We should have went with relocations through out.
Much cleaner to extend.
Same philosophy here. Landing pad is a metadata.
It's not a good idea to encode metadata into assembly instruction.
bpf ISA is not a high level IR.
> Do we have a working toolchain somewhere? I can work a few examples.
Please send SCEV upstream asap. This is way more important
then arguing over this bits.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-22 23:37 ` Alexei Starovoitov
@ 2026-09-23 0:04 ` Eduard Zingerman
2026-09-23 19:04 ` Eduard Zingerman
0 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-23 0:04 UTC (permalink / raw)
To: Alexei Starovoitov, Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Tue, 2026-09-22 at 23:37 +0000, Alexei Starovoitov wrote:
> On Tue Sep 22, 2026 at 11:08 PM UTC, Eduard Zingerman wrote:
> > On Tue, 2026-09-22 at 21:47 +0000, Alexei Starovoitov wrote:
> > > On Tue Sep 22, 2026 at 4:27 AM UTC, Eduard Zingerman wrote:
> > > >
> > > > Total size of executable sections for all Meta BPF object files used
> > > > for CI veristat testing is 10Mb. Of these there are 30K non-helper
> > > > call instructions in total.
> > > > So we are talking about increase by 8 * 30K = 240K ~ 2.4% worst case.
> > > > Note that the instruction encoding optimized for size already hearts
> > > > us in src/dst registers department: we have no room for virtual
> > > > registers, which would have simplified e.g. register allocation task
> > > > for ARM64. Point being that optimizing intermediate IR for size is not
> > > > always a right target.
> > >
> > > I wasn't talking about increase in bpf ELF size,
> > > but the amount of the verifier work necessary to double the number
> > > of call insns.
> >
> > I don't understand what you mean by the amount of the verifier work.
> > Traced program paths would remain the same, only the encoding of the
> > information about exceptions changes.
>
> Any new insn requires plumbing through out.
> imo ld_imm64 encoding FD inline was a mistake.
> We should have went with relocations through out.
> Much cleaner to extend.
> Same philosophy here. Landing pad is a metadata.
> It's not a good idea to encode metadata into assembly instruction.
> bpf ISA is not a high level IR.
Here [1] is Yonghong's series repackaged by codex with the following
encoding:
call <subprog-or-kfunc> // regular encoding
unwind <label> // a new instruction that must follow the call instruction
The amount of the verifier/libbpf code dropped from ~2K to 1K lines.
[1] https://github.com/eddyz87/bpf/tree/unwind-postfix
> > Do we have a working toolchain somewhere? I can work a few examples.
>
> Please send SCEV upstream asap. This is way more important
> then arguing over this bits.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-23 0:04 ` Eduard Zingerman
@ 2026-09-23 19:04 ` Eduard Zingerman
2026-09-23 19:24 ` Andrii Nakryiko
0 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-23 19:04 UTC (permalink / raw)
To: Alexei Starovoitov, Yonghong Song, bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann, kernel-team
On Tue, 2026-09-22 at 17:04 -0700, Eduard Zingerman wrote:
> On Tue, 2026-09-22 at 23:37 +0000, Alexei Starovoitov wrote:
> > On Tue Sep 22, 2026 at 11:08 PM UTC, Eduard Zingerman wrote:
> > > On Tue, 2026-09-22 at 21:47 +0000, Alexei Starovoitov wrote:
> > > > On Tue Sep 22, 2026 at 4:27 AM UTC, Eduard Zingerman wrote:
> > > > >
> > > > > Total size of executable sections for all Meta BPF object files used
> > > > > for CI veristat testing is 10Mb. Of these there are 30K non-helper
> > > > > call instructions in total.
> > > > > So we are talking about increase by 8 * 30K = 240K ~ 2.4% worst case.
> > > > > Note that the instruction encoding optimized for size already hearts
> > > > > us in src/dst registers department: we have no room for virtual
> > > > > registers, which would have simplified e.g. register allocation task
> > > > > for ARM64. Point being that optimizing intermediate IR for size is not
> > > > > always a right target.
> > > >
> > > > I wasn't talking about increase in bpf ELF size,
> > > > but the amount of the verifier work necessary to double the number
> > > > of call insns.
> > >
> > > I don't understand what you mean by the amount of the verifier work.
> > > Traced program paths would remain the same, only the encoding of the
> > > information about exceptions changes.
> >
> > Any new insn requires plumbing through out.
> > imo ld_imm64 encoding FD inline was a mistake.
> > We should have went with relocations through out.
> > Much cleaner to extend.
> > Same philosophy here. Landing pad is a metadata.
> > It's not a good idea to encode metadata into assembly instruction.
> > bpf ISA is not a high level IR.
>
> Here [1] is Yonghong's series repackaged by codex with the following
> encoding:
>
> call <subprog-or-kfunc> // regular encoding
> unwind <label> // a new instruction that must follow the call instruction
>
> The amount of the verifier/libbpf code dropped from ~2K to 1K lines.
>
> [1] https://github.com/eddyz87/bpf/tree/unwind-postfix
I want to highlight this once more, before it gets buried.
Whether landing pad is metadata or not is a matter of opinion,
seeing the program CFG from it's source w/o a need to consult
additional tables is definitely a plus.
Same with ld_imm64, and yes it adds a few checks in the verifier,
but those checks are trivial.
I looked a bit at what other VMs/IRs do and it's a mix:
- JVM has exception tables similar to ours (begin, end, landing-pad).
- GCC (gimple and rtl) have landing pad address attached to an
instruction / control flow graph node.
- LLVM has invoke instruction that takes three parameters:
what to call, where to jump on success, where to jump to unwind.
With that in mind, I think that the rational thing to do is to choose
based on the implementation complexity. Compared to v5 [2] of this
series the branch [1] allows to forgo:
- UAPI to copy the cleanup table from user space (patch #2)
- insn_aux_data changes for cleanup_pad address (patch #4)
- simplifies check_cfg() handling (patch #5)
- simplifies main verification pass (patch #9):
- throw raises unwinding flag
- varifier interprets `unwind` as any other instruction,
jump if unwinding, fallthrough during normal operation.
- dramatically simplifies libbpf part, dropping patches #16,17,18:
- no need to collect cleanup tables from elf
- no need to relocate these tables
- no need to pass them to the syscall
> bpf ISA is not a high level IR.
Well, BPF evolved as x86 in disguise and it was the right choice for
that moment in time. Today it is a limiting factor, just talking about
explicit stack vs SSA, it makes it harder to:
- do liveness analysis
- do range analysis
- do SCEV analysis
- use full set of registers on e.g. ARM
- the fastcall implementation is a hack
So, BPF not being a high level IR is not necessarily a good thing.
[1] https://github.com/eddyz87/bpf/tree/unwind-postfix
[2] https://lore.kernel.org/bpf/20260923045846.2414643-1-yonghong.song@linux.dev/
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-23 19:04 ` Eduard Zingerman
@ 2026-09-23 19:24 ` Andrii Nakryiko
2026-09-23 19:34 ` Kumar Kartikeya Dwivedi
0 siblings, 1 reply; 80+ messages in thread
From: Andrii Nakryiko @ 2026-09-23 19:24 UTC (permalink / raw)
To: Eduard Zingerman
Cc: Alexei Starovoitov, Yonghong Song, bpf, Alexei Starovoitov,
Andrii Nakryiko, Daniel Borkmann, kernel-team
On Wed, Sep 23, 2026 at 12:04 PM Eduard Zingerman <eddyz87@gmail.com> wrote:
>
> On Tue, 2026-09-22 at 17:04 -0700, Eduard Zingerman wrote:
> > On Tue, 2026-09-22 at 23:37 +0000, Alexei Starovoitov wrote:
> > > On Tue Sep 22, 2026 at 11:08 PM UTC, Eduard Zingerman wrote:
> > > > On Tue, 2026-09-22 at 21:47 +0000, Alexei Starovoitov wrote:
> > > > > On Tue Sep 22, 2026 at 4:27 AM UTC, Eduard Zingerman wrote:
> > > > > >
> > > > > > Total size of executable sections for all Meta BPF object files used
> > > > > > for CI veristat testing is 10Mb. Of these there are 30K non-helper
> > > > > > call instructions in total.
> > > > > > So we are talking about increase by 8 * 30K = 240K ~ 2.4% worst case.
> > > > > > Note that the instruction encoding optimized for size already hearts
> > > > > > us in src/dst registers department: we have no room for virtual
> > > > > > registers, which would have simplified e.g. register allocation task
> > > > > > for ARM64. Point being that optimizing intermediate IR for size is not
> > > > > > always a right target.
> > > > >
> > > > > I wasn't talking about increase in bpf ELF size,
> > > > > but the amount of the verifier work necessary to double the number
> > > > > of call insns.
> > > >
> > > > I don't understand what you mean by the amount of the verifier work.
> > > > Traced program paths would remain the same, only the encoding of the
> > > > information about exceptions changes.
> > >
> > > Any new insn requires plumbing through out.
> > > imo ld_imm64 encoding FD inline was a mistake.
> > > We should have went with relocations through out.
> > > Much cleaner to extend.
> > > Same philosophy here. Landing pad is a metadata.
> > > It's not a good idea to encode metadata into assembly instruction.
> > > bpf ISA is not a high level IR.
> >
> > Here [1] is Yonghong's series repackaged by codex with the following
> > encoding:
> >
> > call <subprog-or-kfunc> // regular encoding
> > unwind <label> // a new instruction that must follow the call instruction
> >
> > The amount of the verifier/libbpf code dropped from ~2K to 1K lines.
> >
> > [1] https://github.com/eddyz87/bpf/tree/unwind-postfix
>
> I want to highlight this once more, before it gets buried.
>
> Whether landing pad is metadata or not is a matter of opinion,
> seeing the program CFG from it's source w/o a need to consult
> additional tables is definitely a plus.
> Same with ld_imm64, and yes it adds a few checks in the verifier,
> but those checks are trivial.
>
> I looked a bit at what other VMs/IRs do and it's a mix:
> - JVM has exception tables similar to ours (begin, end, landing-pad).
> - GCC (gimple and rtl) have landing pad address attached to an
> instruction / control flow graph node.
> - LLVM has invoke instruction that takes three parameters:
> what to call, where to jump on success, where to jump to unwind.
>
> With that in mind, I think that the rational thing to do is to choose
> based on the implementation complexity. Compared to v5 [2] of this
> series the branch [1] allows to forgo:
> - UAPI to copy the cleanup table from user space (patch #2)
> - insn_aux_data changes for cleanup_pad address (patch #4)
> - simplifies check_cfg() handling (patch #5)
> - simplifies main verification pass (patch #9):
> - throw raises unwinding flag
> - varifier interprets `unwind` as any other instruction,
> jump if unwinding, fallthrough during normal operation.
> - dramatically simplifies libbpf part, dropping patches #16,17,18:
> - no need to collect cleanup tables from elf
> - no need to relocate these tables
> - no need to pass them to the syscall
>
All of the above sounds like a pretty compelling reason to do the
invoke instruction. 16-byte instructions is a fact of life due to
ldimm64 anyways, might as well utilize those extended instructions to
simplify a bunch of other stuff. Not having a bunch of arbitrary
ELF-side conventions is a big plus, IMO.
Unless there are some technical reasons why metadata table on the side
is objectively better, but seeing that we have precedents with LLVM
using same approach, seems like it's workable and shouldn't paint us
into the corner design-wise, no?
Given a rather lively interest and discussion, perhaps we should do
one of those long forgotten BPF office hours and go over this in a
more face-to-face-ish way?
> > bpf ISA is not a high level IR.
>
> Well, BPF evolved as x86 in disguise and it was the right choice for
> that moment in time. Today it is a limiting factor, just talking about
> explicit stack vs SSA, it makes it harder to:
> - do liveness analysis
> - do range analysis
> - do SCEV analysis
> - use full set of registers on e.g. ARM
> - the fastcall implementation is a hack
>
> So, BPF not being a high level IR is not necessarily a good thing.
>
> [1] https://github.com/eddyz87/bpf/tree/unwind-postfix
> [2] https://lore.kernel.org/bpf/20260923045846.2414643-1-yonghong.song@linux.dev/
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-23 19:24 ` Andrii Nakryiko
@ 2026-09-23 19:34 ` Kumar Kartikeya Dwivedi
2026-09-23 21:34 ` Alexei Starovoitov
0 siblings, 1 reply; 80+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-23 19:34 UTC (permalink / raw)
To: Andrii Nakryiko, Eduard Zingerman
Cc: Alexei Starovoitov, Yonghong Song, bpf, Alexei Starovoitov,
Andrii Nakryiko, Daniel Borkmann, kernel-team
On Wed Sep 23, 2026 at 9:24 PM CEST, Andrii Nakryiko wrote:
> On Wed, Sep 23, 2026 at 12:04 PM Eduard Zingerman <eddyz87@gmail.com> wrote:
>>
>> On Tue, 2026-09-22 at 17:04 -0700, Eduard Zingerman wrote:
>> > On Tue, 2026-09-22 at 23:37 +0000, Alexei Starovoitov wrote:
>> > > On Tue Sep 22, 2026 at 11:08 PM UTC, Eduard Zingerman wrote:
>> > > > On Tue, 2026-09-22 at 21:47 +0000, Alexei Starovoitov wrote:
>> > > > > On Tue Sep 22, 2026 at 4:27 AM UTC, Eduard Zingerman wrote:
>> > > > > >
>> > > > > > Total size of executable sections for all Meta BPF object files used
>> > > > > > for CI veristat testing is 10Mb. Of these there are 30K non-helper
>> > > > > > call instructions in total.
>> > > > > > So we are talking about increase by 8 * 30K = 240K ~ 2.4% worst case.
>> > > > > > Note that the instruction encoding optimized for size already hearts
>> > > > > > us in src/dst registers department: we have no room for virtual
>> > > > > > registers, which would have simplified e.g. register allocation task
>> > > > > > for ARM64. Point being that optimizing intermediate IR for size is not
>> > > > > > always a right target.
>> > > > >
>> > > > > I wasn't talking about increase in bpf ELF size,
>> > > > > but the amount of the verifier work necessary to double the number
>> > > > > of call insns.
>> > > >
>> > > > I don't understand what you mean by the amount of the verifier work.
>> > > > Traced program paths would remain the same, only the encoding of the
>> > > > information about exceptions changes.
>> > >
>> > > Any new insn requires plumbing through out.
>> > > imo ld_imm64 encoding FD inline was a mistake.
>> > > We should have went with relocations through out.
>> > > Much cleaner to extend.
>> > > Same philosophy here. Landing pad is a metadata.
>> > > It's not a good idea to encode metadata into assembly instruction.
>> > > bpf ISA is not a high level IR.
>> >
>> > Here [1] is Yonghong's series repackaged by codex with the following
>> > encoding:
>> >
>> > call <subprog-or-kfunc> // regular encoding
>> > unwind <label> // a new instruction that must follow the call instruction
>> >
>> > The amount of the verifier/libbpf code dropped from ~2K to 1K lines.
>> >
>> > [1] https://github.com/eddyz87/bpf/tree/unwind-postfix
>>
>> I want to highlight this once more, before it gets buried.
>>
>> Whether landing pad is metadata or not is a matter of opinion,
>> seeing the program CFG from it's source w/o a need to consult
>> additional tables is definitely a plus.
>> Same with ld_imm64, and yes it adds a few checks in the verifier,
>> but those checks are trivial.
>>
>> I looked a bit at what other VMs/IRs do and it's a mix:
>> - JVM has exception tables similar to ours (begin, end, landing-pad).
>> - GCC (gimple and rtl) have landing pad address attached to an
>> instruction / control flow graph node.
>> - LLVM has invoke instruction that takes three parameters:
>> what to call, where to jump on success, where to jump to unwind.
>>
>> With that in mind, I think that the rational thing to do is to choose
>> based on the implementation complexity. Compared to v5 [2] of this
>> series the branch [1] allows to forgo:
>> - UAPI to copy the cleanup table from user space (patch #2)
We will still have UAPI, in form of new instruction, so this is less salient.
>> - insn_aux_data changes for cleanup_pad address (patch #4)
>> - simplifies check_cfg() handling (patch #5)
>> - simplifies main verification pass (patch #9):
>> - throw raises unwinding flag
>> - varifier interprets `unwind` as any other instruction,
>> jump if unwinding, fallthrough during normal operation.
>> - dramatically simplifies libbpf part, dropping patches #16,17,18:
>> - no need to collect cleanup tables from elf
>> - no need to relocate these tables
>> - no need to pass them to the syscall
>>
All of these do make sense.
>
> All of the above sounds like a pretty compelling reason to do the
> invoke instruction. 16-byte instructions is a fact of life due to
> ldimm64 anyways, might as well utilize those extended instructions to
> simplify a bunch of other stuff. Not having a bunch of arbitrary
> ELF-side conventions is a big plus, IMO.
>
IIUC we just need extra unwind following the call, since it already has two
control flow edges (target and fallthrough for what it would return to), so
it might be simpler than full blown new invoke.
But yeah, we didn't shy away from adding bespoke instructions when it suited us
(may_goto, etc.) in the past.
> Unless there are some technical reasons why metadata table on the side
> is objectively better, but seeing that we have precedents with LLVM
> using same approach, seems like it's workable and shouldn't paint us
> into the corner design-wise, no?
>
> Given a rather lively interest and discussion, perhaps we should do
> one of those long forgotten BPF office hours and go over this in a
> more face-to-face-ish way?
>
>> > bpf ISA is not a high level IR.
>> [...]
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-23 19:34 ` Kumar Kartikeya Dwivedi
@ 2026-09-23 21:34 ` Alexei Starovoitov
2026-09-23 22:00 ` Eduard Zingerman
0 siblings, 1 reply; 80+ messages in thread
From: Alexei Starovoitov @ 2026-09-23 21:34 UTC (permalink / raw)
To: Kumar Kartikeya Dwivedi, Andrii Nakryiko, Eduard Zingerman
Cc: Yonghong Song, bpf, Alexei Starovoitov, Andrii Nakryiko,
Daniel Borkmann, kernel-team
On Wed Sep 23, 2026 at 7:34 PM UTC, Kumar Kartikeya Dwivedi wrote:
> On Wed Sep 23, 2026 at 9:24 PM CEST, Andrii Nakryiko wrote:
>> On Wed, Sep 23, 2026 at 12:04 PM Eduard Zingerman <eddyz87@gmail.com> wrote:
>>>
>>> On Tue, 2026-09-22 at 17:04 -0700, Eduard Zingerman wrote:
>>> > On Tue, 2026-09-22 at 23:37 +0000, Alexei Starovoitov wrote:
>>> > > On Tue Sep 22, 2026 at 11:08 PM UTC, Eduard Zingerman wrote:
>>> > > > On Tue, 2026-09-22 at 21:47 +0000, Alexei Starovoitov wrote:
>>> > > > > On Tue Sep 22, 2026 at 4:27 AM UTC, Eduard Zingerman wrote:
>>> > > > > >
>>> > > > > > Total size of executable sections for all Meta BPF object files used
>>> > > > > > for CI veristat testing is 10Mb. Of these there are 30K non-helper
>>> > > > > > call instructions in total.
>>> > > > > > So we are talking about increase by 8 * 30K = 240K ~ 2.4% worst case.
>>> > > > > > Note that the instruction encoding optimized for size already hearts
>>> > > > > > us in src/dst registers department: we have no room for virtual
>>> > > > > > registers, which would have simplified e.g. register allocation task
>>> > > > > > for ARM64. Point being that optimizing intermediate IR for size is not
>>> > > > > > always a right target.
>>> > > > >
>>> > > > > I wasn't talking about increase in bpf ELF size,
>>> > > > > but the amount of the verifier work necessary to double the number
>>> > > > > of call insns.
>>> > > >
>>> > > > I don't understand what you mean by the amount of the verifier work.
>>> > > > Traced program paths would remain the same, only the encoding of the
>>> > > > information about exceptions changes.
>>> > >
>>> > > Any new insn requires plumbing through out.
>>> > > imo ld_imm64 encoding FD inline was a mistake.
>>> > > We should have went with relocations through out.
>>> > > Much cleaner to extend.
>>> > > Same philosophy here. Landing pad is a metadata.
>>> > > It's not a good idea to encode metadata into assembly instruction.
>>> > > bpf ISA is not a high level IR.
>>> >
>>> > Here [1] is Yonghong's series repackaged by codex with the following
>>> > encoding:
>>> >
>>> > call <subprog-or-kfunc> // regular encoding
>>> > unwind <label> // a new instruction that must follow the call instruction
>>> >
>>> > The amount of the verifier/libbpf code dropped from ~2K to 1K lines.
>>> >
>>> > [1] https://github.com/eddyz87/bpf/tree/unwind-postfix
>>>
>>> I want to highlight this once more, before it gets buried.
>>>
>>> Whether landing pad is metadata or not is a matter of opinion,
>>> seeing the program CFG from it's source w/o a need to consult
>>> additional tables is definitely a plus.
>>> Same with ld_imm64, and yes it adds a few checks in the verifier,
>>> but those checks are trivial.
>>>
>>> I looked a bit at what other VMs/IRs do and it's a mix:
>>> - JVM has exception tables similar to ours (begin, end, landing-pad).
>>> - GCC (gimple and rtl) have landing pad address attached to an
>>> instruction / control flow graph node.
>>> - LLVM has invoke instruction that takes three parameters:
>>> what to call, where to jump on success, where to jump to unwind.
>>>
>>> With that in mind, I think that the rational thing to do is to choose
>>> based on the implementation complexity. Compared to v5 [2] of this
>>> series the branch [1] allows to forgo:
>>> - UAPI to copy the cleanup table from user space (patch #2)
>
> We will still have UAPI, in form of new instruction, so this is less salient.
>
>>> - insn_aux_data changes for cleanup_pad address (patch #4)
>>> - simplifies check_cfg() handling (patch #5)
>>> - simplifies main verification pass (patch #9):
>>> - throw raises unwinding flag
>>> - varifier interprets `unwind` as any other instruction,
>>> jump if unwinding, fallthrough during normal operation.
>>> - dramatically simplifies libbpf part, dropping patches #16,17,18:
>>> - no need to collect cleanup tables from elf
>>> - no need to relocate these tables
>>> - no need to pass them to the syscall
>>>
>
> All of these do make sense.
Sorry, all I hear is a knee jerk reaction to a lot of slop.
yes. a lot of it is not needed, but invoke approach is no go.
As I said couple time [ip_start, ip_end] is not a single call insn.
libbpf patches do a bunch of unnecessary copy paste,
but that's because we didn't do insn_array support cleanly.
From libbpf pov bpf_cleanup section shouldn't be any different as
"one more section with pointers to instructions".
Sorry, we're not doing new 16-byte insn. We can chat about it
during the meeting, but I feel it will be a waste of time.
New insn doesn't reduce amount of slop.
Ed's example 1k vs 2k is a counter example.
This is just one llm vs another. Both sucked and both slop.
Just different amount of it.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-23 21:34 ` Alexei Starovoitov
@ 2026-09-23 22:00 ` Eduard Zingerman
2026-09-23 23:22 ` Alexei Starovoitov
0 siblings, 1 reply; 80+ messages in thread
From: Eduard Zingerman @ 2026-09-23 22:00 UTC (permalink / raw)
To: Alexei Starovoitov, Kumar Kartikeya Dwivedi, Andrii Nakryiko
Cc: Yonghong Song, bpf, Alexei Starovoitov, Andrii Nakryiko,
Daniel Borkmann, kernel-team
On Wed, 2026-09-23 at 21:34 +0000, Alexei Starovoitov wrote:
...
> Sorry, all I hear is a knee jerk reaction to a lot of slop.
> yes. a lot of it is not needed, but invoke approach is no go.
> As I said couple time [ip_start, ip_end] is not a single call insn.
> libbpf patches do a bunch of unnecessary copy paste,
> but that's because we didn't do insn_array support cleanly.
> From libbpf pov bpf_cleanup section shouldn't be any different as
> "one more section with pointers to instructions".
>
> Sorry, we're not doing new 16-byte insn. We can chat about it
> during the meeting, but I feel it will be a waste of time.
If you simply don't like 16-byte instructions, then yes,
nothing to talk about.
> New insn doesn't reduce amount of slop.
It literally and demonstrably does, I listed the actions that are not
necessary with this encoding in the previous email. No matter which
llm you use more actions needed == more code.
> Ed's example 1k vs 2k is a counter example.
> This is just one llm vs another. Both sucked and both slop.
> Just different amount of it.
Fwiw, it was not a one-shot. The verifier part is curated.
The jit part is whatever codex decided to do with the original series.
I stand by the complexity reduction claim.
^ permalink raw reply [flat|nested] 80+ messages in thread
* Re: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
2026-09-23 22:00 ` Eduard Zingerman
@ 2026-09-23 23:22 ` Alexei Starovoitov
0 siblings, 0 replies; 80+ messages in thread
From: Alexei Starovoitov @ 2026-09-23 23:22 UTC (permalink / raw)
To: Eduard Zingerman, Kumar Kartikeya Dwivedi, Andrii Nakryiko
Cc: Yonghong Song, bpf, Alexei Starovoitov, Andrii Nakryiko,
Daniel Borkmann, kernel-team
On Wed Sep 23, 2026 at 10:00 PM UTC, Eduard Zingerman wrote:
> On Wed, 2026-09-23 at 21:34 +0000, Alexei Starovoitov wrote:
>
> ...
>
>> Sorry, all I hear is a knee jerk reaction to a lot of slop.
>> yes. a lot of it is not needed, but invoke approach is no go.
>> As I said couple time [ip_start, ip_end] is not a single call insn.
>> libbpf patches do a bunch of unnecessary copy paste,
>> but that's because we didn't do insn_array support cleanly.
>> From libbpf pov bpf_cleanup section shouldn't be any different as
>> "one more section with pointers to instructions".
>>
>> Sorry, we're not doing new 16-byte insn. We can chat about it
>> during the meeting, but I feel it will be a waste of time.
>
> If you simply don't like 16-byte instructions, then yes,
> nothing to talk about.
>
>> New insn doesn't reduce amount of slop.
>
> It literally and demonstrably does, I listed the actions that are not
> necessary with this encoding in the previous email. No matter which
> llm you use more actions needed == more code.
>
>> Ed's example 1k vs 2k is a counter example.
>> This is just one llm vs another. Both sucked and both slop.
>> Just different amount of it.
>
> Fwiw, it was not a one-shot. The verifier part is curated.
> The jit part is whatever codex decided to do with the original series.
> I stand by the complexity reduction claim.
I don't measure complexity in number of lines.
There is also llvm part that you're forgetting about
and dragging that new insn all the way.
^ permalink raw reply [flat|nested] 80+ messages in thread