BPF List
 help / color / mirror / Atom feed
* [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds
@ 2026-09-26  5:00 Yonghong Song
  2026-09-26  5:00 ` [PATCH bpf-next v6 01/21] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
                   ` (20 more replies)
  0 siblings, 21 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:00 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

bpf_throw() walks the BPF call stack to the exception boundary and
discards every frame in between. A frame that owns something -- an RCU
read lock, a preemption-disabled section, a referenced kptr -- never gets
to give it back, so the verifier refuses to let such a frame throw at
all. That is the whole reason a Rust program cannot use bpf_throw() as
its panic path right now: Rust's Drop glue *is* that give-back, and there
is nowhere to run it.

LLVM 23 added the compiler half ([1]). A Rust function that owns a value
across a call that can unwind

    fn foo() {
        let _guard = RcuReadGuard::new();  /* bpf_rcu_read_lock()   */
        may_throw();                       /* extern "C-unwind"     */
    }                                      /* Drop: rcu_read_unlock */

lowers to an invoke with a cleanup landing pad holding the Drop call, and
the BPF backend writes one record per invoke region into a .bpf_cleanup
section: a flat table of 12-byte (begin, end, landing_pad) triples, each
field a byte offset into the code section. The rule is "a frame suspended
at a call in [begin, end) resumes at landing_pad when an unwind passes
through". A pad ends with a call to _Unwind_Resume(), which the kernel
provides as the bpf_unwind_resume() kfunc.

Rather than overload bpf_throw(), which keeps its own meaning -- leave for
the exception boundary with the frames in between discarded -- a new
bpf_unwind() kfunc raises the unwind this series dispatches.

This series is the kernel half: take that table at BPF_PROG_LOAD, teach
the verifier that a covered call can also go to its landing pad, and have
bpf_unwind() run the pads as it walks.

C has no unwinding, so the selftests spell out by hand what a frontend
emits -- a call site bracketed by two labels, a landing pad, and a record
tying them together. The frame above, written that way:

        "call bpf_rcu_read_lock;"
    "1:"    "call foo3;"                /* cleanup region */
    "2:"
        ... normal path, ends in bpf_rcu_read_unlock ...
    "6:"                                /* landing pad */
        "call bpf_rcu_read_unlock;"
        "call bpf_unwind_resume;"
        CLEANUP_REC("1b", "2b", "6b")

Design
======

A pad runs in the frame that owns it, entered by an ordinary return.
bpf_unwind() walks the frames with arch_bpf_stack_walk_ra(), which hands
out the slot each frame's return address came from, and rewrites that
slot: to the frame's landing pad where a record covers the call it is
suspended at, and to the frame's epilogue where nothing does. Then it
returns.

Each frame therefore runs its own pad, on its own stack, and leaves
through its own epilogue -- which is what puts its caller's r6-r9 back. The
unwind needs no trampoline, no spill area and no per-frame metadata beyond
the table itself. The fixups lower a pad's bpf_unwind_resume() to
"r0 = 0; exit", so the frame returns and the address rewritten below it
carries the unwind on to the next pad.

The verifier sees the same shape with no new machinery. A covered call
gets a second successor, its landing pad, in the same frame, and its entry
state is the state at the call with the caller-saved registers gone, since
the callee's epilogue puts them back on the way out. Nothing crosses a
frame boundary that the instruction stream does not already describe, so
precision, liveness and the resource checks work there as they do on any
other branch.

  1      pack bpf_insn_aux_data's flags into bit fields, so the flags this
         series adds cost bits rather than bytes
  2-3    uapi: cleanup_info in BPF_PROG_LOAD and struct bpf_cleanup_info;
         the bpf_unwind() and bpf_unwind_resume() kfuncs
  4-9    verifier: mark the covered call sites, give each an edge to its
         pad, and refuse the shapes that cannot be dispatched
  10     bpf_unwind(): rewrite the return addresses as it walks
  11-12  x86-64 and arm64 JITs
  13-17  libbpf: collect .bpf_cleanup, pass it to the kernel, resolve
         _Unwind_Resume, carry it through the light skeleton and the
         static linker
  18-21  selftests

Limitations
===========

  - Cleanup pads only. A catch pad -- one that ends in a plain exit rather
    than a resume, which is what Rust's catch_unwind would need -- is
    refused: bpf_unwind() rewrites every frame's return address in one
    pass, so a frame above a catch pad would resume at a pad for an unwind
    that had already been caught. LLVM refuses type-specific catches and
    filters on its side as well.
  - A JIT that can dispatch pads is required: x86-64 (with
    CONFIG_UNWINDER_ORC, which bpf_throw() already needs there) and arm64.
    Anywhere else the load fails with -EOPNOTSUPP rather than silently
    doing nothing.
  - No offloaded programs, no private stack, and no combining a table with
    an exception callback.
  - In the pad's own frame: no tail call, no BPF_LD_[ABS|IND] and no
    indirect jump, none of which reaches the resume. And no second unwind
    anywhere above a pad, including a call to a global subprogram that can
    raise one.
  - The Rust toolchain does not properly support BPF exception handling
    yet. The tables the selftests use are hand-written inline asm, which
    the assembler turns into the same relocations the BPF AsmPrinter emits,
    so libbpf and the kernel see an object indistinguishable from a
    compiler-generated one.

  [1] https://github.com/llvm/llvm-project/pull/192164
      llvm commit 9d51c891b719 ("[BPF] Add exception handling support
      with .bpf_cleanup section")

Changelog
=========
  v5 -> v6:
    - v5: https://lore.kernel.org/bpf/20260923045846.2414643-1-yonghong.song@linux.dev/
    - Run a pad in the frame that owns it: bpf_unwind() rewrites each
      frame's saved return address rather than calling the pad as a
      subroutine of the walker. The bpf_cleanup_pad.S trampolines, the
      per-frame spill area and the pad-entry register header all go away.
    - Raise the unwind with a new bpf_unwind() kfunc, so that bpf_throw()
      keeps its meaning.
    - Add arch_bpf_stack_walk_ra(), which also hands out the slot a return
      address came from. arm64 re-signs the address it writes there.
    - A pad is an ordinary second successor of a covered call, so the
      verifier needs no unwind edge of its own: the cross-frame precision
      and liveness work is gone, with the v5 fixes it needed, and two
      patches become one.
    - Refuse the undispatchable shapes per instruction in do_check(), keyed
      on a per-frame mark, rather than by walking each pad.
    - Allow on-stack call arguments in a pad, which now has its own frame.
    - Add __set_global() and __ret_global() test tags, so RUN_TESTS() drives
      the shapes and test_shapes() goes away (suggested by Eduard).
    - Drop the shapes that needed a driver of their own, and the three
      extension objects with them.
  v4 -> v5:
    - v4: https://lore.kernel.org/bpf/20260921210033.1715000-1-yonghong.song@linux.dev/
    - Rebase on bpf-next.
    - Clear r0 where a throw enters a landing pad: no instruction defines
      it, so a precision request for it outlived the state and oopsed the
      verifier.
    - Defer entering the frames an unwind edge crossed until the backtrack
      reads an instruction from them; entering at the landing pad could
      leave bt->frame past the parent state's frames and oops the verifier.
    - Stamp the popped frame count inside bpf_push_jmp_history(), so a pad's
      entry carries it even when the prune path is what creates the entry.
    - Replace the hand-rolled CFG traversal and its separate pass with
      per-instruction checks in do_check(), keyed on the verifier's
      unwinding state; kernel/bpf/exception.c halves.
    - Add a first patch packing bpf_insn_aux_data's flags into one bit field
      word, 144 bytes to 128, so the flags this series adds cost bits.
    - Drop the per-subprogram arrays the JITs consulted for throw sites,
      resume sites and pad bodies, and read insn_aux_data, which a JIT
      already has.
    - Drop the pad-entry r0 header: r0 at a pad is an unknown scalar, and a
      dispatcher writes a defined value there only to keep a kernel one out
      of BPF.
    - Take the bool arguments back out of verifier_remove_insns() and the
      site collector, and share pop_frame() with prepare_func_exit().
    - Rename the recorded call sites to throw_call and resume_call, give the
      exported functions a bpf_exc_ prefix, and drop cleanup_ from the
      statics.
  v3 -> v4:
    - v3: https://lore.kernel.org/bpf/20260920054225.864535-1-yonghong.song@linux.dev/
    - Rebase on bpf-next due to conflict.
    - Reserve the throw-site spill area only in a (sub)program that calls
      bpf_throw(): 40 bytes of stack per frame on x86-64, 80 on arm64.
    - Bound the record count by the number of instructions in the program,
      and name both in the message.
    - Refuse a .bpf_cleanup section in libbpf whose record count cannot be
      handed to the kernel as a count times a record size in an int.
    - Fix the static linker's new bounds check, which could itself wrap, and
      refuse a section too small to hold one field.
    - WARN once if the body of bpf_unwind_resume() is ever reached, the way
      bpf_throw() does where its exception callback should never return.
    - Rename nr_pad_body to pad_body_bits, use BTF_ID_LIST_SINGLE, move
      cleanup_pad out of insn_aux_data's bools, drop an arm64 include.
  v2 -> v3:
    - v2: https://lore.kernel.org/bpf/20260918044156.3283973-1-yonghong.song@linux.dev/
    - Keep a landing pad's record when opt_remove_nops() deletes a pad that is
      a nop, instead of dropping it after the verifier has already checked the
      call site against it.
    - Refuse a BPF_LD_[ABS|IND] in a pad body.
    - Teach mark_chain_precision() about the throw-to-pad edge, which crosses
      frames with no instruction to account for them.
    - Refuse a bpf_unwind_resume() in any frame but the one whose landing pad
      the walker entered.
    - Check raw_data, alignment and bounds before the static linker writes
      through a relocation in a non-executable section, which may be SHT_NOBITS.
  v1 -> v2:
    - v1: https://lore.kernel.org/bpf/20260917055645.3926444-1-yonghong.song@linux.dev/
    - Consolidate all usages of kern_extern_name() in a single patch in libbpf.
    - Avoid compiler warning and add proper cleanup_info_cnt guard in libbpf when
      collecting .bpf_cleanup records.
    - Add cleanup_info_cnt condition for emit_rel_store() with cleanup_info.

Yonghong Song (21):
  bpf: Pack bpf_insn_aux_data flags into bit fields
  bpf: Accept the compiler's exception cleanup table at program load
  bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs
  bpf: Add lookups for exception cleanup resumes and landing pads
  bpf: Prepare for an exception cleanup table before the CFG walk
  bpf: Make exception landing pads reachable in the CFG
  bpf: Resume a covered call at its landing pad
  bpf: Refuse a landing pad that does not resume
  bpf: Refuse a private stack for a program with an exception cleanup
    table
  bpf: Dispatch cleanup pads by rewriting return addresses
  bpf, x86: Dispatch exception cleanup pads at run time
  bpf, arm64: Dispatch exception cleanup pads at run time
  libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
  libbpf: Add cleanup_info to bpf_prog_load_opts
  libbpf: Collect .bpf_cleanup records and pass them to the kernel
  libbpf: Carry the exception cleanup table through the light skeleton
  libbpf: Let the static linker carry .bpf_cleanup relocations
  selftests/bpf: Add an end-to-end .bpf_cleanup exception test
  selftests/bpf: Add __set_global() and __ret_global() test tags
  selftests/bpf: Cover the exception cleanup shapes the chain does not
    reach
  selftests/bpf: Load an exception cleanup program from a light skeleton

 arch/arm64/kernel/stacktrace.c                |  94 +++
 arch/arm64/net/bpf_jit_comp.c                 |  20 +-
 arch/x86/net/bpf_jit_comp.c                   |  35 +-
 include/linux/bpf.h                           |  45 ++
 include/linux/bpf_verifier.h                  |  59 +-
 include/linux/filter.h                        |   3 +
 include/uapi/linux/bpf.h                      |   9 +
 kernel/bpf/Makefile                           |   2 +-
 kernel/bpf/backtrack.c                        |  24 +-
 kernel/bpf/cfg.c                              |  65 +-
 kernel/bpf/check_btf.c                        | 134 ++++
 kernel/bpf/core.c                             |  25 +-
 kernel/bpf/exception.c                        | 260 +++++++
 kernel/bpf/exception.h                        |  23 +
 kernel/bpf/fixups.c                           | 141 +++-
 kernel/bpf/helpers.c                          |  61 ++
 kernel/bpf/liveness.c                         |  24 +
 kernel/bpf/syscall.c                          |   2 +-
 kernel/bpf/verifier.c                         |  93 +++
 tools/include/uapi/linux/bpf.h                |   9 +
 tools/lib/bpf/bpf.c                           |   6 +-
 tools/lib/bpf/bpf.h                           |   7 +-
 tools/lib/bpf/gen_loader.c                    |  29 +-
 tools/lib/bpf/libbpf.c                        | 342 ++++++++-
 tools/lib/bpf/libbpf_internal.h               |  10 +
 tools/lib/bpf/linker.c                        |  36 +-
 tools/testing/selftests/bpf/Makefile.skel     |   2 +-
 .../selftests/bpf/exceptions_cleanup.h        |  47 ++
 .../bpf/prog_tests/exceptions_cleanup.c       | 115 +++
 tools/testing/selftests/bpf/progs/bpf_misc.h  |   7 +
 .../selftests/bpf/progs/exceptions_cleanup.c  | 162 +++++
 .../bpf/progs/exceptions_cleanup_fail.c       | 593 ++++++++++++++++
 .../bpf/progs/exceptions_cleanup_light.c      |  39 +
 .../bpf/progs/exceptions_cleanup_shapes.c     | 665 ++++++++++++++++++
 tools/testing/selftests/bpf/test_loader.c     | 207 ++++++
 35 files changed, 3351 insertions(+), 44 deletions(-)
 create mode 100644 kernel/bpf/exception.c
 create mode 100644 kernel/bpf/exception.h
 create mode 100644 tools/testing/selftests/bpf/exceptions_cleanup.h
 create mode 100644 tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup.c
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 01/21] bpf: Pack bpf_insn_aux_data flags into bit fields
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
@ 2026-09-26  5:00 ` Yonghong Song
  2026-09-26  5:00 ` [PATCH bpf-next v6 02/21] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
                   ` (19 subsequent siblings)
  20 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:00 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

struct bpf_insn_aux_data spreads its flags over three places: eight bools a
byte each, a u8 carrying three bit fields, and a u32 group with 25 of its
32 bits spare. Put them all in one u64 word, alu_state with them, and move
orig_idx below it. The structure goes from 136 bytes to 128 with no holes.

No functional change: a one-bit unsigned field holds 0 and 1 the way the
bool did, and nothing takes the address of any of them.

Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf_verifier.h | 43 ++++++++++++++++++------------------
 1 file changed, 22 insertions(+), 21 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index c775bd757706..d85cf969bcb0 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -664,41 +664,42 @@ struct bpf_insn_aux_data {
 	u64 map_key_state; /* constant (32 bit) key tracking for maps */
 	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
 	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
-	bool nospec; /* do not execute this instruction speculatively */
-	bool nospec_result; /* result is unsafe under speculation, nospec must follow */
-	bool zext_dst; /* this insn zero extends dst reg */
-	bool needs_zext; /* alu op needs to clear upper bits */
-	bool prevent_zext; /* alu op cannot be zext (already used with 64-bit scalars) */
-	bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
-	bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
-	bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
-	u8 alu_state; /* used in combination with alu_limit */
+	u64 nospec:1; /* do not execute this instruction speculatively */
+	u64 nospec_result:1; /* result is unsafe under speculation, nospec must follow */
+	u64 zext_dst:1; /* this insn zero extends dst reg */
+	u64 needs_zext:1; /* alu op needs to clear upper bits */
+	u64 prevent_zext:1; /* alu op cannot be zext (already used with 64-bit scalars) */
+	u64 non_sleepable:1; /* helper/kfunc may be called from non-sleepable context */
+	u64 is_iter_next:1; /* bpf_iter_<type>_next() kfunc call */
+	u64 call_with_percpu_alloc_ptr:1; /* {this,per}_cpu_ptr() with prog percpu alloc */
+	u64 alu_state:8; /* used in combination with alu_limit */
 	/* true if STX or LDX instruction is a part of a spill/fill
 	 * pattern for a bpf_fastcall call.
 	 */
-	u8 fastcall_pattern:1;
+	u64 fastcall_pattern:1;
 	/* for CALL instructions, a number of spill/fill pairs in the
 	 * bpf_fastcall pattern.
 	 */
-	u8 fastcall_spills_num:3;
-	u8 arg_prog:4;
+	u64 fastcall_spills_num:3;
+	u64 arg_prog:4;
 
-	/* below fields are initialized once */
-	unsigned int orig_idx; /* original instruction index */
-	u32 jmp_point:1;
-	u32 prune_point:1;
+	/* below flags are initialized once */
+	u64 jmp_point:1;
+	u64 prune_point:1;
 	/* ensure we check state equivalence and save state checkpoint and
 	 * this instruction, regardless of any heuristics
 	 */
-	u32 force_checkpoint:1;
+	u64 force_checkpoint:1;
 	/* true if instruction is a call to a helper function that
 	 * accepts callback function as a parameter.
 	 */
-	u32 calls_callback:1;
-	u32 indirect_target:1; /* if it is an indirect jump target */
-	u32 non_stack_access:1; /* instruction can access non-stack memory */
+	u64 calls_callback:1;
+	u64 indirect_target:1; /* if it is an indirect jump target */
+	u64 non_stack_access:1; /* instruction can access non-stack memory */
 	/* true if some jump or call instruction targets this instruction */
-	u32 jump_target:1;
+	u64 jump_target:1;
+
+	unsigned int orig_idx; /* original instruction index, initialized once */
 	/*
 	 * CFG strongly connected component this instruction belongs to,
 	 * zero if it is a singleton SCC.
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 02/21] bpf: Accept the compiler's exception cleanup table at program load
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
  2026-09-26  5:00 ` [PATCH bpf-next v6 01/21] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
@ 2026-09-26  5:00 ` Yonghong Song
  2026-09-26  5:00 ` [PATCH bpf-next v6 03/21] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
                   ` (18 subsequent siblings)
  20 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:00 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

LLVM 23 added exception handling support for BPF with the .bpf_cleanup
section ([1]). Rust code compiled with panic=unwind runs cleanup code
(Drop glue) when an unwind passes through, and the LLVM BPF backend emits
that section from the landing pads the frontend produced. Plain C cannot
generate .bpf_cleanup unless inline asm is used. The Rust compiler does not
*properly* support BPF exception handling yet, but the kernel can support
the table today, and inline assembly is enough to test it.

Add the UAPI to carry the .bpf_cleanup table into the kernel. BPF_PROG_LOAD
grows cleanup_info, cleanup_info_cnt and cleanup_info_rec_size, and struct
bpf_cleanup_info describes one record as a triple of instruction indices:
the half-open call-site range [begin_off, end_off) and the landing_pad_off
the frame resumes at. The table arrives sorted by begin_off, with disjoint
ranges and each record's three offsets inside one subprogram;
check_cleanup_info() holds it to that at load time. Nothing reads it yet;
the patches that follow -- the CFG walk, the unwind walk and the JITs --
are its consumers.

Link: https://github.com/llvm/llvm-project/pull/192164 [1]
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf_verifier.h   |   2 +
 include/uapi/linux/bpf.h       |   9 +++
 kernel/bpf/check_btf.c         | 134 +++++++++++++++++++++++++++++++++
 kernel/bpf/syscall.c           |   2 +-
 kernel/bpf/verifier.c          |   1 +
 tools/include/uapi/linux/bpf.h |   9 +++
 6 files changed, 156 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index d85cf969bcb0..6ce25c96ebed 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1018,6 +1018,8 @@ struct bpf_verifier_env {
 	struct spill_snapshot **callsite_at_stack;
 	u32 pass_cnt; /* number of times do_check() was called */
 	u32 subprog_cnt;
+	struct bpf_cleanup_info *cleanup_info;
+	u32 cleanup_info_cnt;
 	/* number of instructions analyzed by the verifier */
 	u32 prev_insn_processed, insn_processed;
 	/* number of jmps, calls, exits analyzed so far */
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 4687c3310996..aca43f4f927f 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -1702,6 +1702,9 @@ union bpf_attr {
 		 * verification.
 		 */
 		__s32		keyring_id;
+		__aligned_u64	cleanup_info;	/* exception cleanup table */
+		__u32		cleanup_info_rec_size; /* userspace bpf_cleanup_info size */
+		__u32		cleanup_info_cnt; /* number of bpf_cleanup_info records */
 	};
 
 	struct { /* anonymous struct used by BPF_OBJ_* commands */
@@ -7638,6 +7641,12 @@ struct bpf_line_info {
 	__u32	line_col;
 };
 
+struct bpf_cleanup_info {
+	__u32	begin_off;
+	__u32	end_off;
+	__u32	landing_pad_off;
+};
+
 struct bpf_spin_lock {
 	__u32	val;
 };
diff --git a/kernel/bpf/check_btf.c b/kernel/bpf/check_btf.c
index 4c1ed842f661..5c62f350d500 100644
--- a/kernel/bpf/check_btf.c
+++ b/kernel/bpf/check_btf.c
@@ -407,6 +407,136 @@ int bpf_check_core_relo(struct bpf_verifier_env *env,
 	return err;
 }
 
+#define MIN_BPF_CLEANUP_INFO_SIZE	12
+#define MAX_CLEANUP_INFO_REC_SIZE	MAX_FUNCINFO_REC_SIZE
+
+static int check_cleanup_info(struct bpf_verifier_env *env,
+			      const union bpf_attr *attr,
+			      bpfptr_t uattr)
+{
+	u32 krec_size = sizeof(struct bpf_cleanup_info);
+	u32 i, nrec, urec_size, min_size, prev_end = 0;
+	struct bpf_cleanup_info *krecord;
+	bpfptr_t urecord;
+	int ret = -EINVAL;
+
+	nrec = attr->cleanup_info_cnt;
+	if (!nrec)
+		return 0;
+	if (nrec > env->prog->len) {
+		verbose(env, "cleanup info has %u records for %u instructions\n",
+			nrec, env->prog->len);
+		return -EINVAL;
+	}
+
+	urec_size = attr->cleanup_info_rec_size;
+	if (urec_size < MIN_BPF_CLEANUP_INFO_SIZE ||
+	    urec_size > MAX_CLEANUP_INFO_REC_SIZE ||
+	    urec_size % sizeof(u32)) {
+		verbose(env, "invalid cleanup info rec size %u\n", urec_size);
+		return -EINVAL;
+	}
+
+	krecord = kvcalloc(nrec, krec_size, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+	if (!krecord)
+		return -ENOMEM;
+
+	min_size = min_t(u32, krec_size, urec_size);
+	urecord = make_bpfptr(attr->cleanup_info, uattr.is_kernel);
+	for (i = 0; i < nrec; i++) {
+		struct bpf_subprog_info *sb, *se, *sl;
+		struct bpf_cleanup_info *rec = &krecord[i];
+
+		ret = bpf_check_uarg_tail_zero(urecord, krec_size, urec_size);
+		if (ret) {
+			if (ret == -E2BIG) {
+				verbose(env, "nonzero tailing record in cleanup info\n");
+				if (copy_to_bpfptr_offset(uattr,
+							  offsetof(union bpf_attr,
+								   cleanup_info_rec_size),
+							  &min_size, sizeof(min_size)))
+					ret = -EFAULT;
+			}
+			goto err_free;
+		}
+
+		if (copy_from_bpfptr(rec, urecord, min_size)) {
+			ret = -EFAULT;
+			goto err_free;
+		}
+		bpfptr_add(&urecord, urec_size);
+
+		ret = -EINVAL;
+		if (rec->begin_off >= rec->end_off) {
+			verbose(env, "cleanup_info[%u]: begin %u >= end %u\n",
+				i, rec->begin_off, rec->end_off);
+			goto err_free;
+		}
+		if (i && rec->begin_off < prev_end) {
+			verbose(env,
+				"cleanup_info[%u]: range [%u,%u) is unsorted or overlaps the previous record\n",
+				i, rec->begin_off, rec->end_off);
+			goto err_free;
+		}
+		prev_end = rec->end_off;
+
+		sb = bpf_find_containing_subprog(env, rec->begin_off);
+		se = bpf_find_containing_subprog(env, rec->end_off - 1);
+		sl = bpf_find_containing_subprog(env, rec->landing_pad_off);
+		if (!sb || !se || !sl) {
+			verbose(env, "cleanup_info[%u]: offset out of range\n", i);
+			goto err_free;
+		}
+		if (sb != se || sb != sl) {
+			verbose(env,
+				"cleanup_info[%u]: range/landing pad span multiple subprogs\n",
+				i);
+			goto err_free;
+		}
+		/*
+		 * A zero opcode is the second half of a 16-byte insn, not an
+		 * insn. end_off is exclusive, so it may be one past the last.
+		 */
+		if (!env->prog->insnsi[rec->begin_off].code ||
+		    !env->prog->insnsi[rec->landing_pad_off].code ||
+		    (rec->end_off < env->prog->len &&
+		     !env->prog->insnsi[rec->end_off].code)) {
+			verbose(env, "cleanup_info[%u]: points at invalid insn\n", i);
+			goto err_free;
+		}
+	}
+
+	/* Reject a landing pad inside any call-site range, its own included. */
+	ret = -EINVAL;
+	for (i = 0; i < nrec; i++) {
+		u32 pad = krecord[i].landing_pad_off;
+		u32 l = 0, r = nrec;
+
+		while (l < r) {
+			u32 m = l + (r - l) / 2;
+
+			if (pad < krecord[m].begin_off) {
+				r = m;
+			} else if (pad >= krecord[m].end_off) {
+				l = m + 1;
+			} else {
+				verbose(env,
+					"cleanup_info[%u]: landing pad %u is inside the call-site range of cleanup_info[%u]\n",
+					i, pad, m);
+				goto err_free;
+			}
+		}
+	}
+
+	env->cleanup_info = krecord;
+	env->cleanup_info_cnt = nrec;
+	return 0;
+
+err_free:
+	kvfree(krecord);
+	return ret;
+}
+
 int bpf_prepare_btf_info(struct bpf_verifier_env *env,
 			 const union bpf_attr *attr,
 			 bpfptr_t uattr)
@@ -441,6 +571,10 @@ int bpf_check_btf_info(struct bpf_verifier_env *env,
 {
 	int err;
 
+	err = check_cleanup_info(env, attr, uattr);
+	if (err)
+		return err;
+
 	if (!attr->func_info_cnt && !attr->line_info_cnt) {
 		if (check_abnormal_return(env))
 			return -EINVAL;
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index ac52f4ae414c..0e14afe3fdc5 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -2924,7 +2924,7 @@ int __init __used bpf_multi_func(void) { return 0; }
 BTF_ID_LIST_GLOBAL_SINGLE(bpf_multi_func_btf_id, func, bpf_multi_func)
 
 /* last field in 'union bpf_attr' used by this command */
-#define BPF_PROG_LOAD_LAST_FIELD keyring_id
+#define BPF_PROG_LOAD_LAST_FIELD cleanup_info_cnt
 
 static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, struct bpf_log_attr *attr_log)
 {
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 03dbc0e00398..6d3408f295ed 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22801,6 +22801,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	kvfree(env->callx_edges);
 	kvfree(env->func_ptrs);
 	bpf_diag_free(env);
+	kvfree(env->cleanup_info);
 	kvfree(env);
 	return ret;
 }
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 4687c3310996..aca43f4f927f 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -1702,6 +1702,9 @@ union bpf_attr {
 		 * verification.
 		 */
 		__s32		keyring_id;
+		__aligned_u64	cleanup_info;	/* exception cleanup table */
+		__u32		cleanup_info_rec_size; /* userspace bpf_cleanup_info size */
+		__u32		cleanup_info_cnt; /* number of bpf_cleanup_info records */
 	};
 
 	struct { /* anonymous struct used by BPF_OBJ_* commands */
@@ -7638,6 +7641,12 @@ struct bpf_line_info {
 	__u32	line_col;
 };
 
+struct bpf_cleanup_info {
+	__u32	begin_off;
+	__u32	end_off;
+	__u32	landing_pad_off;
+};
+
 struct bpf_spin_lock {
 	__u32	val;
 };
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 03/21] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
  2026-09-26  5:00 ` [PATCH bpf-next v6 01/21] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
  2026-09-26  5:00 ` [PATCH bpf-next v6 02/21] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
@ 2026-09-26  5:00 ` Yonghong Song
  2026-09-26  5:00 ` [PATCH bpf-next v6 04/21] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
                   ` (17 subsequent siblings)
  20 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:00 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

An exception cleanup needs two terminators, and the kernel provides both as
kfuncs so that a BPF program can name them.

bpf_unwind() begins an unwind: the frames between it and whatever catches
run their landing pads on the way out. It is separate from bpf_throw(),
which keeps its own meaning -- leaving for the exception boundary with the
frames in between discarded.

bpf_unwind_resume() ends a landing pad, carrying the unwind on once that
frame's cleanups have run. The compiler names it _Unwind_Resume, the base
unwind ABI's entry point for the same thing; a later libbpf patch resolves
that name to this one. Every unwind ABI hands _Unwind_Resume the exception
object, and LLVM emits that argument on BPF too. Nothing in the kernel
needs it now, but the kfunc takes it as ptr__ign so the prototype matches
the call the compiler makes.

Both are defined here and neither is registered with any program type yet.
A call to either is only valid in a particular place -- an unwind where a
record covers it, a resume inside a landing pad -- and the verifier cannot
say that until it knows what a landing pad is, so registration waits for
the patch that adds the rule.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 kernel/bpf/helpers.c | 14 ++++++++++++++
 1 file changed, 14 insertions(+)

diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index a284f20c97d5..08aee86a155c 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -3424,6 +3424,10 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
 	return false;
 }
 
+__bpf_kfunc void bpf_unwind(void)
+{
+}
+
 __bpf_kfunc void bpf_throw(u64 cookie)
 {
 	struct bpf_throw_ctx ctx = {};
@@ -3445,6 +3449,16 @@ __bpf_kfunc void bpf_throw(u64 cookie)
 	WARN(1, "A call to BPF exception callback should never return\n");
 }
 
+__bpf_kfunc void bpf_unwind_resume(void *ptr__ign)
+{
+	/*
+	 * Never reached: the verifier accepts this kfunc only as a frame
+	 * terminator and do_misc_fixups() lowers every one of them to
+	 * 'r0 = 0; exit', so no call to this body survives to run.
+	 */
+	WARN_ONCE(1, "exception cleanup resume was not lowered to a return\n");
+}
+
 __bpf_kfunc int bpf_wq_init(struct bpf_wq *wq, void *p__const_map, unsigned int flags)
 {
 	struct bpf_async_kern *async = (struct bpf_async_kern *)wq;
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 04/21] bpf: Add lookups for exception cleanup resumes and landing pads
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (2 preceding siblings ...)
  2026-09-26  5:00 ` [PATCH bpf-next v6 03/21] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
@ 2026-09-26  5:00 ` Yonghong Song
  2026-09-26  5:00 ` [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
                   ` (16 subsequent siblings)
  20 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:00 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

Add two new files, exception.h and exception.c, to host the exception
handling code. Only a few helpers so far: recognising a call to
bpf_unwind_resume(), and asking which landing pad, if any, a call site
unwinds to.

The pad of a call site is kept in insn_aux_data, so the two places that
move instructions around -- bpf_patch_insn_data() and
verifier_remove_insns() -- learn to keep it in step.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf_verifier.h |  7 +++++++
 kernel/bpf/Makefile          |  2 +-
 kernel/bpf/exception.c       | 30 ++++++++++++++++++++++++++++++
 kernel/bpf/exception.h       | 12 ++++++++++++
 kernel/bpf/fixups.c          | 26 +++++++++++++++++++++++++-
 5 files changed, 75 insertions(+), 2 deletions(-)
 create mode 100644 kernel/bpf/exception.c
 create mode 100644 kernel/bpf/exception.h

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 6ce25c96ebed..6d78c20e6507 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -700,6 +700,11 @@ struct bpf_insn_aux_data {
 	u64 jump_target:1;
 
 	unsigned int orig_idx; /* original instruction index, initialized once */
+	/*
+	 * 1 + the instruction index of the exception cleanup landing pad
+	 * this call site unwinds to, or 0 for none.
+	 */
+	u32 cleanup_pad;
 	/*
 	 * CFG strongly connected component this instruction belongs to,
 	 * zero if it is a singleton SCC.
@@ -1588,6 +1593,8 @@ u32 btf_func_arg_align(const struct btf *btf, const struct btf_type *t);
 
 int bpf_find_subprog(struct bpf_verifier_env *env, int off);
 bool bpf_is_throw_kfunc(struct bpf_insn *insn);
+bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
+bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
 int bpf_compute_const_regs(struct bpf_verifier_env *env);
 int bpf_prune_dead_branches(struct bpf_verifier_env *env);
 int bpf_check_cfg(struct bpf_verifier_env *env);
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index c1f9b0d3468d..8a6947b3d13a 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -11,7 +11,7 @@ obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_
 obj-$(CONFIG_BPF_SYSCALL) += hashtab.o arraymap.o percpu_freelist.o bpf_lru_list.o lpm_trie.o map_in_map.o bloom_filter.o
 obj-$(CONFIG_BPF_SYSCALL) += local_storage.o queue_stack_maps.o ringbuf.o bpf_insn_array.o
 obj-$(CONFIG_BPF_SYSCALL) += bpf_local_storage.o bpf_task_storage.o
-obj-$(CONFIG_BPF_SYSCALL) += fixups.o cfg.o states.o backtrack.o check_btf.o
+obj-$(CONFIG_BPF_SYSCALL) += fixups.o cfg.o states.o backtrack.o check_btf.o exception.o
 obj-${CONFIG_BPF_LSM}	  += bpf_inode_storage.o
 obj-$(CONFIG_BPF_SYSCALL) += disasm.o mprog.o
 obj-$(CONFIG_BPF_JIT) += trampoline.o
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
new file mode 100644
index 000000000000..b19fcbf49b7e
--- /dev/null
+++ b/kernel/bpf/exception.c
@@ -0,0 +1,30 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bpf.h>
+#include <linux/bpf_verifier.h>
+#include <linux/btf.h>
+#include <linux/btf_ids.h>
+#include <linux/filter.h>
+#include "exception.h"
+
+BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
+BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
+
+bool bpf_is_unwind_kfunc(const struct bpf_insn *insn)
+{
+	return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
+	       insn->imm == bpf_unwind_id[0];
+}
+
+bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
+{
+	return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
+	       insn->imm == bpf_unwind_resume_id[0];
+}
+
+int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
+{
+	u32 pad = env->insn_aux_data[idx].cleanup_pad;
+
+	return pad ? (int)pad - 1 : -1;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
new file mode 100644
index 000000000000..634cb2c15ffc
--- /dev/null
+++ b/kernel/bpf/exception.h
@@ -0,0 +1,12 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef _LINUX_BPF_EXCEPTION_H
+#define _LINUX_BPF_EXCEPTION_H
+
+#include <linux/types.h>
+
+struct bpf_verifier_env;
+
+int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
+
+#endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 37cf130ebb57..5b7fe4ba610b 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -268,11 +268,18 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
 			data[i].non_stack_access =
 				data[off + cnt - 1].non_stack_access;
 			data[off + cnt - 1].non_stack_access = false;
+			data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
+			data[off + cnt - 1].cleanup_pad = 0;
 		} else if (bpf_is_mem_insn(insn + i)) {
 			data[i].non_stack_access = true;
 		}
 	}
 
+	if (env->cleanup_info_cnt)
+		for (i = 0; i < prog_len; i++)
+			if (data[i].cleanup_pad > off + 1)
+				data[i].cleanup_pad += cnt - 1;
+
 	/*
 	 * Last slot instruction could be a newly generated
 	 * BPF_ST/BPF_LDX/BPF_STX, systematically mark it for non-stack access
@@ -619,6 +626,7 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
 	struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
 	unsigned int orig_prog_len = env->prog->len;
 	int err;
+	u32 i;
 
 	if (bpf_rewrite_must_abort())
 		return -EINTR;
@@ -647,6 +655,17 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
 		sizeof(*aux_data) * (orig_prog_len - off - cnt));
 	env->insn_aux_data_len -= cnt;
 
+	if (env->cleanup_info_cnt) {
+		for (i = 0; i < env->insn_aux_data_len; i++) {
+			u32 pad = aux_data[i].cleanup_pad;
+
+			if (pad > off + cnt)
+				aux_data[i].cleanup_pad = pad - cnt;
+			else if (pad > off)
+				aux_data[i].cleanup_pad = 0;
+		}
+	}
+
 	return 0;
 }
 
@@ -752,7 +771,7 @@ int bpf_opt_remove_nops(struct bpf_verifier_env *env)
 	struct bpf_insn *insn = env->prog->insnsi;
 	int insn_cnt = env->prog->len;
 	bool is_may_goto_0, is_ja;
-	int i, err;
+	int i, j, err;
 
 	for (i = 0; i < insn_cnt; i++) {
 		is_may_goto_0 = !memcmp(&insn[i], &MAY_GOTO_0, sizeof(MAY_GOTO_0));
@@ -763,6 +782,11 @@ int bpf_opt_remove_nops(struct bpf_verifier_env *env)
 		if (aux[i].indirect_target)
 			continue;
 
+		if (env->cleanup_info_cnt)
+			for (j = 0; j < insn_cnt; j++)
+				if (env->insn_aux_data[j].cleanup_pad == i + 1)
+					env->insn_aux_data[j].cleanup_pad = i + 2;
+
 		err = verifier_remove_insns(env, i, 1);
 		if (err)
 			return err;
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (3 preceding siblings ...)
  2026-09-26  5:00 ` [PATCH bpf-next v6 04/21] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
@ 2026-09-26  5:00 ` Yonghong Song
  2026-09-26  5:16   ` sashiko-bot
  2026-09-27 20:39   ` bot+bpf-ci
  2026-09-26  5:00 ` [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG Yonghong Song
                   ` (15 subsequent siblings)
  20 siblings, 2 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:00 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

Record in insn_aux_data what the later passes need from the cleanup table:
cleanup_pad, the landing pad a frame resumes at, for every call that can
unwind -- a BPF-to-BPF call or bpf_unwind() -- within the [begin_off,
end_off) range of a cleanup record. Helper and other kfunc calls in the
range cannot unwind and are skipped. Subsequent commits consume it.

bpf_prepare_cleanup_exceptions() runs before bpf_check_cfg(), whose walk
consumes what it produces. It refuses a table on an offloaded program, on
one whose JIT cannot dispatch landing pads or was not asked to compile it,
and on one that also installs an exception callback -- two different
answers to what runs on the way out. What survives is marked jit_required:
the interpreter cannot dispatch a pad. bpf_jit_supports_cleanup_pads() is
weak here and says no; the arch patches provide the real ones.

Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/filter.h |  1 +
 kernel/bpf/core.c      |  5 +++++
 kernel/bpf/exception.c | 47 ++++++++++++++++++++++++++++++++++++++++++
 kernel/bpf/exception.h |  1 +
 kernel/bpf/verifier.c  |  6 ++++++
 5 files changed, 60 insertions(+)

diff --git a/include/linux/filter.h b/include/linux/filter.h
index e42eccb0990e..972b3ed2a51d 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1248,6 +1248,7 @@ bool bpf_jit_supports_stack_args(void);
 bool bpf_jit_supports_arena_args(void);
 bool bpf_jit_supports_far_kfunc_call(void);
 bool bpf_jit_supports_exceptions(void);
+bool bpf_jit_supports_cleanup_pads(void);
 bool bpf_jit_supports_ptr_xchg(void);
 bool bpf_jit_supports_arena(void);
 bool bpf_jit_supports_insn(struct bpf_insn *insn, bool in_arena);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index d3b8b626ec0f..d813fdde29e3 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -3511,6 +3511,11 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
 {
 }
 
+bool __weak bpf_jit_supports_cleanup_pads(void)
+{
+	return false;
+}
+
 bool __weak bpf_jit_supports_timed_may_goto(void)
 {
 	return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index b19fcbf49b7e..1ed0370a171b 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -7,9 +7,56 @@
 #include <linux/filter.h>
 #include "exception.h"
 
+#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
+
 BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
 BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
 
+static void mark_call_sites(struct bpf_verifier_env *env)
+{
+	u32 i, j;
+
+	for (i = 0; i < env->cleanup_info_cnt; i++) {
+		struct bpf_cleanup_info *rec = &env->cleanup_info[i];
+
+		for (j = rec->begin_off; j < rec->end_off; j++) {
+			struct bpf_insn *insn = &env->prog->insnsi[j];
+
+			if (!bpf_pseudo_call(insn) && !bpf_is_unwind_kfunc(insn))
+				continue;
+			env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
+		}
+	}
+}
+
+int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env)
+{
+	if (!env->cleanup_info_cnt)
+		return 0;
+
+	if (bpf_prog_is_offloaded(env->prog->aux)) {
+		verbose(env,
+			"exception cleanup is not supported for offloaded programs\n");
+		return -EINVAL;
+	}
+
+	if (!bpf_jit_supports_cleanup_pads() || !env->prog->jit_requested) {
+		verbose(env,
+			"exception cleanup needs a JIT that can dispatch landing pads\n");
+		return -EOPNOTSUPP;
+	}
+	env->prog->jit_required = 1;
+
+	if (env->exception_callback_subprog) {
+		verbose(env,
+			"exception cleanup table cannot be combined with an exception callback\n");
+		return -EINVAL;
+	}
+
+	mark_call_sites(env);
+	return 0;
+}
+
 bool bpf_is_unwind_kfunc(const struct bpf_insn *insn)
 {
 	return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index 634cb2c15ffc..b96b2c429e22 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -7,6 +7,7 @@
 
 struct bpf_verifier_env;
 
+int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
 int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
 
 #endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 6d3408f295ed..fc3df452de2e 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -37,6 +37,7 @@
 
 #include "diagnostics.h"
 #include "disasm.h"
+#include "exception.h"
 
 static const struct bpf_verifier_ops * const bpf_verifier_ops[] = {
 #define BPF_PROG_TYPE(_id, _name, prog_ctx_type, kern_ctx_type) \
@@ -22584,6 +22585,11 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	if (ret < 0)
 		goto skip_full_check;
 
+	/* The CFG needs an edge from a call in a cleanup range to its pad. */
+	ret = bpf_prepare_cleanup_exceptions(env);
+	if (ret < 0)
+		goto skip_full_check;
+
 	/* Validate instructions and resolve the program's referenced resources. */
 	ret = check_and_resolve_insns(env);
 	if (ret < 0)
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (4 preceding siblings ...)
  2026-09-26  5:00 ` [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
@ 2026-09-26  5:00 ` Yonghong Song
  2026-09-26  5:21   ` sashiko-bot
  2026-09-27 20:40   ` bot+bpf-ci
  2026-09-26  5:00 ` [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad Yonghong Song
                   ` (14 subsequent siblings)
  20 siblings, 2 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:00 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

A bpf_unwind() or a bpf2bpf call inside the [begin_off, end_off) range of a
cleanup record can reach that record's landing pad. Add that edge to the
CFG walk, which explores the pad and makes both ends prune points, and to
bpf_insn_successors(), which liveness and the SCC passes walk.

Liveness needs one more thing. When bpf_stack_slot_alive()'s
is_live_before() says an outer frame's slot has no reader after the call
the frame is suspended at, the landing pad can still be one. Ask about the
pad as well when the call site names one; otherwise clean_verifier_state()
poisons the slot while the callee runs and the pad is rejected for reading
it.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 kernel/bpf/cfg.c      | 54 ++++++++++++++++++++++++++++++++++++++++---
 kernel/bpf/liveness.c | 21 +++++++++++++++++
 2 files changed, 72 insertions(+), 3 deletions(-)

diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index b0bd9ba951df..4e2b6985bc96 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -6,6 +6,7 @@
 #include <linux/sort.h>
 
 #include "diagnostics.h"
+#include "exception.h"
 
 #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
 
@@ -160,17 +161,64 @@ static int push_insn(int t, int w, int e, struct bpf_verifier_env *env)
 	return DONE_EXPLORING;
 }
 
+static int visit_cleanup_pad_edge(int t, struct bpf_verifier_env *env)
+{
+	int *insn_stack = env->cfg.insn_stack;
+	int *insn_state = env->cfg.insn_state;
+	int w;
+
+	if (!env->cleanup_info_cnt)
+		return DONE_EXPLORING;
+	w = bpf_exc_pad_of_call(env, t);
+	if (w < 0)
+		return DONE_EXPLORING;
+
+	/*
+	 * @t is a call that may branch here, and @w is the target of that
+	 * branch, so both are prune points. @w especially: every covered call
+	 * site in a region unwinds to the same pad, and without a prune point
+	 * at its head the verifier walks the pad again for each of them.
+	 */
+	mark_prune_point(env, t);
+	mark_prune_point(env, w);
+	mark_jmp_point(env, w);
+	mark_jump_target(env, w);
+
+	if (insn_state[w])
+		return DONE_EXPLORING;
+	if (env->cfg.cur_stack >= env->prog->len)
+		return -E2BIG;
+	insn_stack[env->cfg.cur_stack++] = w;
+	insn_state[w] |= DISCOVERED;
+	return KEEP_EXPLORING;
+}
+
+static int merge_visit_ret(int a, int b)
+{
+	if (a < 0)
+		return a;
+	if (b < 0)
+		return b;
+	if (a == KEEP_EXPLORING || b == KEEP_EXPLORING)
+		return KEEP_EXPLORING;
+	return DONE_EXPLORING;
+}
+
 static int visit_func_call_insn(int t, struct bpf_insn *insns,
 				struct bpf_verifier_env *env,
 				bool visit_callee)
 {
-	int ret, insn_sz;
+	int ret, insn_sz, pad_ret;
 	int w;
 
+	pad_ret = visit_cleanup_pad_edge(t, env);
+	if (pad_ret < 0)
+		return pad_ret;
+
 	insn_sz = bpf_is_ldimm64(&insns[t]) ? 2 : 1;
 	ret = push_insn(t, t + insn_sz, FALLTHROUGH, env);
 	if (ret)
-		return ret;
+		return merge_visit_ret(pad_ret, ret);
 
 	mark_prune_point(env, t + insn_sz);
 	/* when we exit from subprog, we need to record non-linear history */
@@ -182,7 +230,7 @@ static int visit_func_call_insn(int t, struct bpf_insn *insns,
 		merge_callee_effects(env, t, w);
 		ret = push_insn(t, w, BRANCH, env);
 	}
-	return ret;
+	return merge_visit_ret(pad_ret, ret);
 }
 
 struct bpf_iarray *bpf_iarray_realloc(struct bpf_iarray *old, size_t n_elem)
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index cd9523f69298..4e0273a8ceee 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -8,6 +8,8 @@
 #include <linux/slab.h>
 #include <linux/sort.h>
 
+#include "exception.h"
+
 #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
 
 /*
@@ -384,6 +386,18 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
 			succ->items[succ->cnt++] = exit_idx;
 	}
 
+	/*
+	 * A call a cleanup record covers can leave through its landing pad.
+	 * Only a call to a subprogram or to bpf_unwind() is marked, neither of
+	 * which is an edge the block above adds, so succ still holds two.
+	 */
+	if (unlikely(env->cleanup_info_cnt)) {
+		int pad = bpf_exc_pad_of_call(env, idx);
+
+		if (pad >= 0)
+			succ->items[succ->cnt++] = pad;
+	}
+
 	return succ;
 }
 
@@ -545,6 +559,13 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
 		alive = callee_stack_access_at_callsite(env, callsite)
 			? is_live_before(instance, callsite, rel, half_spi)
 			: is_live_before(instance, callsite + 1, rel, half_spi);
+
+		if (!alive && unlikely(env->cleanup_info_cnt)) {
+			int pad = bpf_exc_pad_of_call(env, callsite);
+
+			if (pad >= 0)
+				alive = is_live_before(instance, pad, rel, half_spi);
+		}
 		if (alive)
 			return true;
 	}
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (5 preceding siblings ...)
  2026-09-26  5:00 ` [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG Yonghong Song
@ 2026-09-26  5:00 ` Yonghong Song
  2026-09-26  5:15   ` sashiko-bot
  2026-09-27 20:40   ` bot+bpf-ci
  2026-09-26  5:00 ` [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume Yonghong Song
                   ` (13 subsequent siblings)
  20 siblings, 2 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:00 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

A landing pad runs in the frame that owns it, entered by an ordinary
return: bpf_unwind() rewrites the frame's saved return address, so the call
the frame is suspended at comes back at the pad instead of at the next
instruction. Nothing crosses a frame boundary that the instruction stream
does not already describe.

That makes the pad an ordinary second successor of a covered call, in the
same frame, and its entry state knowable without verifying the callee at
all: it is the state at the call with the caller-saved registers gone,
since the callee's epilogue puts r6-r9 and the stack back on its way out.
push_cleanup_pad_branch() pushes exactly that.

The call to bpf_unwind() itself never returns to the instruction after it,
having rewritten its own return address along with the rest. Control
resumes at this frame's pad when a record covers the call, and otherwise
the frame returns at once -- which its caller already explores as the other
side of its own covered call.

Resource accounting needs nothing new. Whatever the callee held it released
in its own pads before this frame's runs, so both sides of the call leave
the caller holding what it held itself, and the ordinary checks at each
exit cover the pad like any other path.

This is also where both kfuncs become callable: they are registered here,
beside the rules that say where each is allowed.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf_verifier.h |  1 +
 kernel/bpf/backtrack.c       | 24 ++++++++++++++++-
 kernel/bpf/cfg.c             | 11 ++++++++
 kernel/bpf/helpers.c         |  2 ++
 kernel/bpf/liveness.c        |  3 +++
 kernel/bpf/verifier.c        | 52 ++++++++++++++++++++++++++++++++++++
 6 files changed, 92 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 6d78c20e6507..0143688896b0 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -836,6 +836,7 @@ struct bpf_subprog_info {
 	s16 fastcall_stack_off;
 	bool has_tail_call: 1;
 	bool might_throw: 1;
+	bool might_unwind: 1;
 	bool tail_call_reachable: 1;
 	bool has_ld_abs: 1;
 	bool is_cb: 1;
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 0e38b9575328..57665c67e66b 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -4,6 +4,7 @@
 #include <linux/bpf_verifier.h>
 #include <linux/filter.h>
 #include <linux/bitmap.h>
+#include "exception.h"
 
 #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
 
@@ -434,8 +435,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
 					return -EFAULT;
 			}
 
+			if (bpf_exc_pad_of_call(env, idx) == subseq_idx) {
+				/*
+				 * We came from this call's landing pad, which
+				 * runs in the caller's frame: on that path the
+				 * callee's frame was never entered, so there is
+				 * no frame to leave. The call clobbered r0-r5;
+				 * r6-r9 and the stack are the caller's own and
+				 * keep going back from here.
+				 */
+				bt_clear_reg(bt, BPF_REG_0);
+				if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
+					verifier_bug(env, "landing pad unexpected regs %x",
+						     bt_reg_mask(bt));
+					return -EFAULT;
+				}
+				return 0;
 			/* callx calls static subprogs only */
-			if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
+			} else if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
 				/* check that jump history doesn't have any
 				 * extra instructions from subprog; the next
 				 * instruction after call to global subprog
@@ -956,6 +973,11 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
 		if (!st)
 			break;
 
+		if (verifier_bug_if(bt->frame > st->curframe, env,
+				    "backtrack frame %d, state curframe %d",
+				    bt->frame, st->curframe))
+			return -EFAULT;
+
 		for (fr = bt->frame; fr >= 0; fr--) {
 			func = st->frame[fr];
 			bitmap_from_u64(mask, bt_frame_reg_mask(bt, fr));
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index 4e2b6985bc96..cfd4fe4049ca 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -76,6 +76,14 @@ static void mark_subprog_might_throw(struct bpf_verifier_env *env, int off)
 	subprog->might_throw = true;
 }
 
+static void mark_subprog_might_unwind(struct bpf_verifier_env *env, int off)
+{
+	struct bpf_subprog_info *subprog;
+
+	subprog = bpf_find_containing_subprog(env, off);
+	subprog->might_unwind = true;
+}
+
 /* 't' is an index of a call-site.
  * 'w' is a callee entry point.
  * Eventually this function would be called when env->cfg.insn_state[w] == EXPLORED.
@@ -91,6 +99,7 @@ static void merge_callee_effects(struct bpf_verifier_env *env, int t, int w)
 	caller->changes_pkt_data |= callee->changes_pkt_data;
 	caller->might_sleep |= callee->might_sleep;
 	caller->might_throw |= callee->might_throw;
+	caller->might_unwind |= callee->might_unwind;
 }
 
 enum {
@@ -678,6 +687,8 @@ static int visit_insn(int t, struct bpf_verifier_env *env)
 				mark_subprog_changes_pkt_data(env, t);
 			if (ret == 0 && bpf_is_throw_kfunc(insn))
 				mark_subprog_might_throw(env, t);
+			if (ret == 0 && bpf_is_unwind_kfunc(insn))
+				mark_subprog_might_unwind(env, t);
 		}
 		return visit_func_call_insn(t, insns, env, insn->src_reg == BPF_PSEUDO_CALL);
 
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 08aee86a155c..f291611fe578 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -5095,6 +5095,8 @@ BTF_ID_FLAGS(func, bpf_task_get_cgroup1, KF_ACQUIRE | KF_RCU | KF_RET_NULL)
 BTF_ID_FLAGS(func, bpf_task_from_pid, KF_ACQUIRE | KF_RET_NULL)
 BTF_ID_FLAGS(func, bpf_task_from_vpid, KF_ACQUIRE | KF_RET_NULL)
 BTF_ID_FLAGS(func, bpf_throw)
+BTF_ID_FLAGS(func, bpf_unwind)
+BTF_ID_FLAGS(func, bpf_unwind_resume)
 #ifdef CONFIG_BPF_EVENTS
 BTF_ID_FLAGS(func, bpf_send_signal_task)
 #endif
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index 4e0273a8ceee..c9ee4f10f725 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -364,6 +364,9 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
 		return jt;
 	}
 
+	if (unlikely(bpf_is_unwind_resume_kfunc(insn)))
+		return succ;
+
 	opcode_info = &opcode_info_tbl[BPF_CLASS(insn->code) | BPF_OP(insn->code)];
 	insn_sz = bpf_is_ldimm64(insn) ? 2 : 1;
 	if (opcode_info->can_fallthrough)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index fc3df452de2e..77176250f866 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -19167,6 +19167,40 @@ enum {
 	INSN_IDX_UPDATED = 2,
 };
 
+static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
+{
+	struct bpf_verifier_state *branch;
+	struct bpf_func_state *frame;
+	int pad = bpf_exc_pad_of_call(env, insn_idx);
+
+	if (pad < 0)
+		return 0;
+	branch = push_stack(env, pad, insn_idx, false);
+	if (IS_ERR(branch))
+		return PTR_ERR(branch);
+	frame = branch->frame[branch->curframe];
+	/*
+	 * The state at that call with the caller-saved registers gone: the
+	 * callee's epilogue put r6-r9 and the stack back on the way out.
+	 */
+	clear_caller_saved_regs(env, frame->regs);
+	mark_reg_unknown(env, frame->regs, BPF_REG_0);
+	return 0;
+}
+
+static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx)
+{
+	struct bpf_func_state *frame = cur_func(env);
+	int pad = bpf_exc_pad_of_call(env, *insn_idx);
+
+	if (pad < 0)
+		return PROCESS_BPF_EXIT;
+	clear_caller_saved_regs(env, frame->regs);
+	mark_reg_unknown(env, frame->regs, BPF_REG_0);
+	*insn_idx = pad;
+	return INSN_IDX_UPDATED;
+}
+
 static int process_bpf_exit_full(struct bpf_verifier_env *env,
 				 bool *do_print_state,
 				 bool exception_exit)
@@ -19404,6 +19438,20 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
 
 		env->jmps_processed++;
 		if (opcode == BPF_CALL) {
+			if (bpf_is_unwind_kfunc(insn))
+				return process_bpf_unwind(env, &env->insn_idx);
+			if (bpf_is_unwind_resume_kfunc(insn)) {
+				/*
+				 * Mark r0 a known zero -- unknown first, as
+				 * the known-zero helper keeps the type it
+				 * finds, which here is NOT_INIT. The fixups
+				 * lower this to 'r0 = 0; exit', so the frame
+				 * returns a real zero.
+				 */
+				mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
+				mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
+				return process_bpf_exit_full(env, do_print_state, false);
+			}
 			if (env->cur_state->active_locks) {
 				/* similar to static subprog calls callx is allowed under a lock */
 				if (!bpf_is_callx(insn) &&
@@ -19422,6 +19470,10 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
 				}
 			}
 			mark_reg_scratched(env, BPF_REG_0);
+			/* An unwind out of this call resumes at the pad. */
+			err = push_cleanup_pad_branch(env, env->insn_idx);
+			if (err)
+				return err;
 			if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
 				cur_func(env)->no_stack_arg_load = true;
 			if (bpf_is_callx(insn))
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (6 preceding siblings ...)
  2026-09-26  5:00 ` [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad Yonghong Song
@ 2026-09-26  5:00 ` Yonghong Song
  2026-09-26  5:17   ` sashiko-bot
  2026-09-27 20:40   ` bot+bpf-ci
  2026-09-26  5:00 ` [PATCH bpf-next v6 09/21] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
                   ` (12 subsequent siblings)
  20 siblings, 2 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:00 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

A cleanup pad runs drop glue and calls bpf_unwind_resume(), so the frame
returns and the unwind goes on. A catch pad -- std::panic::catch_unwind and
its kin -- runs the same drops and then carries on in its frame, stopping
the unwind there. Only the first is supported: bpf_unwind() rewrites every
frame's return address in one pass, so a frame above a catch pad would
resume at a pad for an unwind that had already been caught.

Nothing in the record says which kind a pad is, but the code does, exactly
as LLVM emits it -- a cleanup pad reaches _Unwind_Resume and a catch pad
reaches a return. bpf_unwind() and the branch pushed at a covered call are
the only ways into a pad, and both mark the frame they enter.
bpf_exc_check_insn() asks that mark about every instruction of a program
that carries a table, and refuses:

 - an exit, which is how a catch pad ends
 - a tail call, which replaces the frame, and a BPF_LD_[ABS|IND], which on
   a failed load leaves through the "r0 = 0; exit" gen_ld_abs() patches in
 - an indirect jump, which nothing a frontend emits in a pad needs
 - a bpf_unwind(), which starts a second walk over frames the first has
   already rewritten, and a call to a global subprogram that might_unwind,
   which is verified on its own and cannot be walked into from here
 - a bpf_unwind_resume() outside a pad, where the fixups would lower it to
   a bare return rather than to a resume
 - an instruction reached both inside and outside a pad, which is a jump
   into a pad: the cleanup would run with nothing to clean up after

The first three are about the pad's own frame, since a subprogram the pad
calls may do any of them and still come back; the unwind is refused
anywhere above a pad, since it never does. do_check() asks before it may
prune the path, so an instruction is in a pad or it is not, never both, and
the mark need not join the comparison in states_equal().

A callback that can unwind is refused here too: an unwind out of one stops
at the helper's own frame, which is C and has no landing pad, so the helper
would carry on as though nothing had happened.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf_verifier.h |  4 ++
 kernel/bpf/exception.c       | 81 ++++++++++++++++++++++++++++++++++++
 kernel/bpf/exception.h       |  3 ++
 kernel/bpf/verifier.c        | 18 ++++++++
 4 files changed, 106 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 0143688896b0..4174c7d0177e 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -339,6 +339,8 @@ struct bpf_func_state {
 	bool in_async_callback_fn;
 	bool in_exception_callback_fn;
 	bool no_stack_arg_load;
+	/* an unwind reached this frame and its landing pad is running */
+	bool in_pad;
 	/* For callback calling functions that limit number of possible
 	 * callback executions (e.g. bpf_loop) keeps track of current
 	 * simulated iteration number.
@@ -698,6 +700,8 @@ struct bpf_insn_aux_data {
 	u64 non_stack_access:1; /* instruction can access non-stack memory */
 	/* true if some jump or call instruction targets this instruction */
 	u64 jump_target:1;
+	u64 in_cleanup_pad:1; /* reached with a landing pad running */
+	u64 outside_cleanup_pad:1; /* ... and the other way round */
 
 	unsigned int orig_idx; /* original instruction index, initialized once */
 	/*
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index 1ed0370a171b..86fdf847d33e 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -12,6 +12,15 @@
 BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
 BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
 
+int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog)
+{
+	if (!env->subprog_info[subprog].might_unwind)
+		return 0;
+
+	verbose(env, "subprog %d may unwind and is used as a callback\n", subprog);
+	return -EINVAL;
+}
+
 static void mark_call_sites(struct bpf_verifier_env *env)
 {
 	u32 i, j;
@@ -69,6 +78,78 @@ bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
 	       insn->imm == bpf_unwind_resume_id[0];
 }
 
+/* Is an unwind in flight: is this frame a landing pad, or below one? */
+static bool unwinding(const struct bpf_verifier_state *state)
+{
+	u32 i;
+
+	for (i = 0; i <= state->curframe; i++)
+		if (state->frame[i]->in_pad)
+			return true;
+	return false;
+}
+
+int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn)
+{
+	bool in_pad = cur_func(env)->in_pad;
+	struct bpf_insn_aux_data *aux;
+	u32 i = env->insn_idx;
+	const char *why = NULL;
+
+	if (unwinding(env->cur_state)) {
+		if (bpf_is_unwind_kfunc(insn)) {
+			verbose(env, "insn %u starts a second unwind while one is in flight\n", i);
+			return -EINVAL;
+		}
+		if (bpf_pseudo_call(insn)) {
+			int subprog = bpf_find_subprog(env, i + insn->imm + 1);
+
+			if (subprog >= 0 && bpf_subprog_is_global(env, subprog) &&
+			    env->subprog_info[subprog].might_unwind) {
+				verbose(env,
+					"insn %u calls global subprog %d, which can unwind while an unwind is in flight\n",
+					i, subprog);
+				return -EINVAL;
+			}
+		}
+	}
+
+	aux = &env->insn_aux_data[i];
+
+	if (in_pad ? aux->outside_cleanup_pad : aux->in_cleanup_pad) {
+		verbose(env, "insn %u runs both inside and outside a landing pad\n", i);
+		return -EINVAL;
+	}
+	if (in_pad)
+		aux->in_cleanup_pad = true;
+	else
+		aux->outside_cleanup_pad = true;
+
+	if (!in_pad)
+		return 0;
+
+	if (insn->code == (BPF_JMP | BPF_EXIT)) {
+		verbose(env,
+			"exit at insn %u ends a landing pad: a catch pad is not supported yet, only cleanup pads that resume\n",
+			i);
+		return -EOPNOTSUPP;
+	}
+	if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call)
+		why = "is a tail call, which replaces the frame";
+	else if (BPF_CLASS(insn->code) == BPF_LD &&
+		 (BPF_MODE(insn->code) == BPF_ABS || BPF_MODE(insn->code) == BPF_IND))
+		why = "is a BPF_LD_[ABS|IND], which can leave through the epilogue";
+	else if (insn->code == (BPF_JMP | BPF_JA | BPF_X) ||
+		 insn->code == (BPF_JMP32 | BPF_JA | BPF_X))
+		why = "is an indirect jump";
+
+	if (!why)
+		return 0;
+
+	verbose(env, "insn %u %s, and is in a landing pad\n", i, why);
+	return -EINVAL;
+}
+
 int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
 {
 	u32 pad = env->insn_aux_data[idx].cleanup_pad;
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index b96b2c429e22..f3aff0fe8ecc 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -6,8 +6,11 @@
 #include <linux/types.h>
 
 struct bpf_verifier_env;
+struct bpf_insn;
 
 int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
 int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
+int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog);
+int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn);
 
 #endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 77176250f866..cc1ed776da15 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10985,6 +10985,10 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
 	 * callbacks
 	 */
 	env->subprog_info[subprog].is_cb = true;
+	err = bpf_exc_check_callback(env, subprog);
+	if (err)
+		return err;
+
 	if (bpf_pseudo_kfunc_call(insn) &&
 	    !is_callback_calling_kfunc(insn->imm)) {
 		verifier_bug(env, "kfunc %s#%d not marked as callback-calling",
@@ -19185,6 +19189,7 @@ static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
 	 */
 	clear_caller_saved_regs(env, frame->regs);
 	mark_reg_unknown(env, frame->regs, BPF_REG_0);
+	frame->in_pad = true;
 	return 0;
 }
 
@@ -19197,6 +19202,7 @@ static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx)
 		return PROCESS_BPF_EXIT;
 	clear_caller_saved_regs(env, frame->regs);
 	mark_reg_unknown(env, frame->regs, BPF_REG_0);
+	frame->in_pad = true;
 	*insn_idx = pad;
 	return INSN_IDX_UPDATED;
 }
@@ -19441,6 +19447,11 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
 			if (bpf_is_unwind_kfunc(insn))
 				return process_bpf_unwind(env, &env->insn_idx);
 			if (bpf_is_unwind_resume_kfunc(insn)) {
+				if (!cur_func(env)->in_pad) {
+					verbose(env, "resume at insn %d is not in a landing pad\n",
+						env->insn_idx);
+					return -EINVAL;
+				}
 				/*
 				 * Mark r0 a known zero -- unknown first, as
 				 * the known-zero helper keeps the type it
@@ -19450,6 +19461,7 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
 				 */
 				mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
 				mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
+				cur_func(env)->in_pad = false;
 				return process_bpf_exit_full(env, do_print_state, false);
 			}
 			if (env->cur_state->active_locks) {
@@ -19578,6 +19590,12 @@ static int do_check(struct bpf_verifier_env *env)
 			}
 		}
 
+		if (unlikely(env->cleanup_info_cnt)) {
+			err = bpf_exc_check_insn(env, insn);
+			if (err)
+				return err;
+		}
+
 		if (bpf_is_prune_point(env, env->insn_idx)) {
 			err = bpf_is_state_visited(env, env->insn_idx);
 			if (err < 0)
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 09/21] bpf: Refuse a private stack for a program with an exception cleanup table
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (7 preceding siblings ...)
  2026-09-26  5:00 ` [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume Yonghong Song
@ 2026-09-26  5:00 ` Yonghong Song
  2026-09-26  5:00 ` [PATCH bpf-next v6 10/21] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
                   ` (11 subsequent siblings)
  20 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:00 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

An unwind resumes a frame at its landing pad rather than at the instruction
after the call. On x86-64 that skips the pop which restores r9, where a
private stack keeps its frame pointer -- the JIT brackets every call in
push_r9/pop_r9 and the pad is reached before the pop runs. The pad would
then address its frame through a stale pointer. Refused on every
architecture rather than just that one.

Force NO_PRIV_STACK in check_max_stack_depth(), where the choice is made.
The subprograms below are then checked against MAX_BPF_STACK together
rather than one at a time, so this can turn a program that would have
loaded with a private stack into one that is too deep.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 kernel/bpf/verifier.c | 13 +++++++++++++
 1 file changed, 13 insertions(+)

diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index cc1ed776da15..677bab92f624 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -5751,6 +5751,19 @@ static int check_max_stack_depth(struct bpf_verifier_env *env)
 		}
 	}
 
+	/*
+	 * An unwind resumes a frame at its landing pad rather than at the
+	 * instruction after the call, and on x86-64 that skips the pop which
+	 * restores r9 -- where a private stack keeps its frame pointer. The
+	 * pad would address its frame through a stale one. Refused on every
+	 * arch rather than just that one. The subprograms
+	 * below are then checked against MAX_BPF_STACK together rather than
+	 * one at a time, so this can turn a program that would have loaded
+	 * with a private stack into one that is too deep.
+	 */
+	if (env->cleanup_info_cnt)
+		priv_stack_mode = NO_PRIV_STACK;
+
 	if (priv_stack_mode == PRIV_STACK_UNKNOWN)
 		priv_stack_mode = bpf_enable_priv_stack(env->prog);
 
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 10/21] bpf: Dispatch cleanup pads by rewriting return addresses
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (8 preceding siblings ...)
  2026-09-26  5:00 ` [PATCH bpf-next v6 09/21] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
@ 2026-09-26  5:00 ` Yonghong Song
  2026-09-27 20:40   ` bot+bpf-ci
  2026-09-26  5:01 ` [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
                   ` (10 subsequent siblings)
  20 siblings, 1 reply; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:00 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

bpf_unwind() walks the BPF frames and, for each one, rewrites the saved
return address so that the frame resumes where the unwind needs it, then
returns. Nothing restores a register: every frame runs its own epilogue on
the way out, which is what puts its caller's r6-r9 back, so the unwind
needs no spill area and no per-frame metadata beyond the table itself.

Two things a frame can be:

  - a record covers the call it is suspended at: resume at that pad, which
    the previous patch has made sure ends in a resume
  - nothing covers it: resume at the frame's epilogue, so it returns at
    once and its caller is reached with its registers already restored

The second needs an epilogue to exist. A frame need not carry a table of
its own to be on the path, so aux->epilogue_ip is recorded for every
program, and handed to the outer program along with the table when
jit_subprogs() compiles the main program as func[0]. And a frame whose only
exit followed the bpf_unwind() call loses it to the dead code sweep, so an
exit is patched back in after every such call. That exit doubles as the
answer for the frame that called bpf_unwind() with no record over the call:
its return address already names the patched 'r0 = 0; exit', so the walk
leaves that one alone and the unwind returns zero. A pad's own resume is
lowered to the same thing, so the frame returns and the rewritten address
carries the unwind to the next pad.

arch_bpf_stack_walk_ra() is the walk that also hands out the return-address
slot. It is a second entry point rather than a change to
arch_bpf_stack_walk(), so that the architectures which do not dispatch pads
keep the walker they have.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf.h          |  45 ++++++++++++++
 include/linux/bpf_verifier.h |   2 +
 include/linux/filter.h       |   2 +
 kernel/bpf/core.c            |  20 +++++-
 kernel/bpf/exception.c       | 102 +++++++++++++++++++++++++++++++
 kernel/bpf/exception.h       |   7 +++
 kernel/bpf/fixups.c          | 115 +++++++++++++++++++++++++++++++++++
 kernel/bpf/helpers.c         |  45 ++++++++++++++
 kernel/bpf/verifier.c        |   3 +
 9 files changed, 340 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 4bae3796c42f..98de251037df 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -1805,6 +1805,49 @@ enum bpf_sig_keyring {
 	BPF_SIG_KEYRING_BPF,
 };
 
+/* One cleanup region of a JITed (sub)program. */
+struct bpf_cleanup_range {
+	u64 begin;
+	u64 end;
+	u64 pad;
+};
+
+struct bpf_exception_info {
+	struct bpf_cleanup_info *info;
+	struct bpf_cleanup_range *ranges;
+	u32 nr_info;
+	u32 nr_ranges;
+};
+
+#ifdef CONFIG_BPF_SYSCALL
+bool bpf_exc_insn_is_pad(const struct bpf_verifier_env *env,
+			 const struct bpf_prog *prog, u32 idx);
+int bpf_exc_attach_main_prog(struct bpf_verifier_env *env, struct bpf_prog *prog);
+void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image);
+void bpf_exc_free_info(struct bpf_prog_aux *aux);
+#else
+
+static inline bool bpf_exc_insn_is_pad(const struct bpf_verifier_env *env,
+				       const struct bpf_prog *prog, u32 idx)
+{
+	return false;
+}
+
+static inline int bpf_exc_attach_main_prog(struct bpf_verifier_env *env,
+					   struct bpf_prog *prog)
+{
+	return 0;
+}
+
+static inline void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+}
+
+static inline void bpf_exc_free_info(struct bpf_prog_aux *aux)
+{
+}
+#endif
+
 struct bpf_prog_aux {
 	atomic64_t refcnt;
 	u32 used_map_cnt;
@@ -1885,6 +1928,8 @@ struct bpf_prog_aux {
 	u64 (*bpf_exception_cb)(u64 cookie, u64 sp, u64 bp, u64, u64);
 	u16 stack_arg_sp_adjust;
 	u16 freplace_link_cnt; /* counts freplace links extending this prog */
+	struct bpf_exception_info *exc;
+	u64 epilogue_ip; /* native address of this (sub)program's epilogue */
 #ifdef CONFIG_SECURITY
 	void *security;
 #endif
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 4174c7d0177e..1803225feb13 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -702,6 +702,7 @@ struct bpf_insn_aux_data {
 	u64 jump_target:1;
 	u64 in_cleanup_pad:1; /* reached with a landing pad running */
 	u64 outside_cleanup_pad:1; /* ... and the other way round */
+	u64 cleanup_pad_head:1; /* first insn of a landing pad */
 
 	unsigned int orig_idx; /* original instruction index, initialized once */
 	/*
@@ -1843,6 +1844,7 @@ int bpf_opt_subreg_zext_lo32_rnd_hi32(struct bpf_verifier_env *env, const union
 int bpf_convert_ctx_accesses(struct bpf_verifier_env *env);
 int bpf_jit_subprogs(struct bpf_verifier_env *env);
 int bpf_fixup_call_args(struct bpf_verifier_env *env);
+int bpf_exc_keep_exits(struct bpf_verifier_env *env);
 int bpf_do_misc_fixups(struct bpf_verifier_env *env);
 int bpf_insn_def32(struct bpf_prog *prog, struct bpf_insn *insn);
 
diff --git a/include/linux/filter.h b/include/linux/filter.h
index 972b3ed2a51d..0d7d949a1baa 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1290,6 +1290,8 @@ u32 bpf_jit_plan_arg_moves(const struct bpf_jit_arg_abi *abi,
 			   struct bpf_jit_arg_move *moves);
 u64 bpf_arch_uaddress_limit(void);
 void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp), void *cookie);
+void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
+			    void *cookie);
 u64 arch_bpf_timed_may_goto(void);
 u64 bpf_check_timed_may_goto(struct bpf_timed_may_goto *);
 bool bpf_helper_changes_pkt_data(enum bpf_func_id func_id);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index d813fdde29e3..60905643cb9c 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -292,6 +292,7 @@ void __bpf_prog_free(struct bpf_prog *fp)
 		mutex_destroy(&fp->aux->dst_mutex);
 		mutex_destroy(&fp->aux->st_ops_assoc_mutex);
 		kfree(fp->aux->poke_tab);
+		bpf_exc_free_info(fp->aux);
 		kfree(fp->aux);
 	}
 	free_percpu(fp->stats);
@@ -2632,9 +2633,14 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
 {
 #ifdef CONFIG_BPF_JIT
 	struct bpf_prog *orig_prog;
+	int ret;
 
-	if (!bpf_prog_need_blind(prog))
+	if (!bpf_prog_need_blind(prog)) {
+		ret = bpf_exc_attach_main_prog(env, prog);
+		if (ret)
+			return ERR_PTR(ret);
 		return bpf_int_jit_compile(env, prog);
+	}
 
 	orig_prog = prog;
 	prog = bpf_jit_blind_constants(env, prog);
@@ -2648,6 +2654,12 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
 		goto out_restore;
 	}
 
+	ret = bpf_exc_attach_main_prog(env, prog);
+	if (ret) {
+		bpf_jit_prog_release_other(orig_prog, prog);
+		return ERR_PTR(ret);
+	}
+
 	prog = bpf_int_jit_compile(env, prog);
 	if (prog->jited) {
 		bpf_jit_prog_release_other(prog, orig_prog);
@@ -3511,6 +3523,12 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
 {
 }
 
+void __weak arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp,
+						      u64 *ra),
+				   void *cookie)
+{
+}
+
 bool __weak bpf_jit_supports_cleanup_pads(void)
 {
 	return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index 86fdf847d33e..c4af82c103d7 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -5,6 +5,7 @@
 #include <linux/btf.h>
 #include <linux/btf_ids.h>
 #include <linux/filter.h>
+#include <linux/slab.h>
 #include "exception.h"
 
 #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
@@ -156,3 +157,104 @@ int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
 
 	return pad ? (int)pad - 1 : -1;
 }
+
+/*
+ * The record covering @ip, which is a return address: the call it belongs to
+ * is the instruction before it, so a range matches on begin < ip <= end.
+ */
+const struct bpf_cleanup_range *bpf_exc_pad_for_ip(const struct bpf_prog *prog, u64 ip)
+{
+	const struct bpf_exception_info *exc = prog->aux->exc;
+	u32 l = 0, r = exc ? exc->nr_ranges : 0;
+
+	while (l < r) {
+		u32 m = l + (r - l) / 2;
+		const struct bpf_cleanup_range *rec = &exc->ranges[m];
+
+		if (ip <= rec->begin)
+			r = m;
+		else if (ip > rec->end)
+			l = m + 1;
+		else
+			return rec;
+	}
+	return NULL;
+}
+
+int bpf_exc_alloc_info(struct bpf_prog_aux *aux)
+{
+	if (aux->exc)
+		return 0;
+	aux->exc = kzalloc_obj(struct bpf_exception_info, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+	return aux->exc ? 0 : -ENOMEM;
+}
+
+int bpf_exc_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt)
+{
+	struct bpf_exception_info *exc = aux->exc;
+	struct bpf_cleanup_range *ranges;
+
+	ranges = kvcalloc(cnt, sizeof(*ranges), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+	if (!ranges) {
+		kvfree(recs);
+		return -ENOMEM;
+	}
+
+	exc->info = recs;
+	exc->nr_info = cnt;
+	exc->ranges = ranges;
+	/* Withheld until the JIT has filled the table in. */
+	exc->nr_ranges = 0;
+	return 0;
+}
+
+void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+	struct bpf_exception_info *exc = prog->aux->exc;
+	u32 i, n;
+
+	if (!exc || !exc->nr_info || !exc->ranges)
+		return;
+
+	n = exc->nr_info;
+	for (i = 0; i < n; i++) {
+		const struct bpf_cleanup_info *rec = &exc->info[i];
+
+		if (WARN_ON_ONCE(rec->begin_off >= prog->len ||
+				 rec->end_off > prog->len ||
+				 rec->landing_pad_off >= prog->len))
+			return;
+		exc->ranges[i].begin = (u64)(long)image + addrs[rec->begin_off];
+		exc->ranges[i].end = (u64)(long)image + addrs[rec->end_off];
+		exc->ranges[i].pad = (u64)(long)image + addrs[rec->landing_pad_off];
+	}
+	exc->nr_ranges = n;
+}
+
+void bpf_exc_free_info(struct bpf_prog_aux *aux)
+{
+	struct bpf_exception_info *exc = aux->exc;
+
+	if (!exc)
+		return;
+	kvfree(exc->ranges);
+	kvfree(exc->info);
+	kfree(exc);
+	aux->exc = NULL;
+}
+
+static const struct bpf_insn_aux_data *subprog_insn_aux(const struct bpf_verifier_env *env,
+							const struct bpf_prog *prog, u32 idx)
+{
+	if (!env || !prog->aux->exc)
+		return NULL;
+	return &env->insn_aux_data[idx + prog->aux->subprog_start];
+}
+
+bool bpf_exc_insn_is_pad(const struct bpf_verifier_env *env,
+			 const struct bpf_prog *prog, u32 idx)
+{
+	const struct bpf_insn_aux_data *aux = subprog_insn_aux(env, prog, idx);
+
+	return aux && aux->cleanup_pad_head;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index f3aff0fe8ecc..025d67dfa925 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -7,10 +7,17 @@
 
 struct bpf_verifier_env;
 struct bpf_insn;
+struct bpf_cleanup_info;
+struct bpf_cleanup_range;
+struct bpf_prog;
+struct bpf_prog_aux;
 
 int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env);
 int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
 int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog);
 int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn);
+int bpf_exc_alloc_info(struct bpf_prog_aux *aux);
+int bpf_exc_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt);
+const struct bpf_cleanup_range *bpf_exc_pad_for_ip(const struct bpf_prog *prog, u64 ip);
 
 #endif /* _LINUX_BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 5b7fe4ba610b..c93a0311d1ea 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -1,5 +1,6 @@
 // SPDX-License-Identifier: GPL-2.0-only
 /* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bitmap.h>
 #include <linux/bpf.h>
 #include <linux/btf.h>
 #include <linux/bpf_verifier.h>
@@ -11,6 +12,7 @@
 #include <linux/sched/signal.h>
 #include <net/xdp.h>
 #include "disasm.h"
+#include "exception.h"
 
 #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
 
@@ -1286,6 +1288,62 @@ static int resolve_func_ptrs(struct bpf_verifier_env *env, struct bpf_prog *prog
 	return 0;
 }
 
+static int exc_info_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
+				u32 subprog, u32 start, u32 end)
+{
+	struct bpf_cleanup_info *recs;
+	u32 i, cnt = 0;
+	int err;
+
+	if (!env->cleanup_info_cnt)
+		return 0;
+
+	err = bpf_exc_alloc_info(sub->aux);
+	if (err)
+		return err;
+
+	for (i = start; i < end; i++) {
+		if (env->insn_aux_data[i].cleanup_pad)
+			cnt++;
+	}
+	if (!cnt)
+		return 0;
+
+	recs = kvmalloc_array(cnt, sizeof(*recs), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+	if (!recs)
+		return -ENOMEM;
+
+	for (i = start, cnt = 0; i < end; i++) {
+		u32 pad = env->insn_aux_data[i].cleanup_pad;
+
+		if (!pad)
+			continue;
+		pad--;
+		if (verifier_bug_if(pad < start || pad >= end, env,
+				    "insn %u is covered by a landing pad at %u outside its subprog [%u, %u)",
+				    i, pad, start, end)) {
+			kvfree(recs);
+			return -EFAULT;
+		}
+		env->insn_aux_data[pad].cleanup_pad_head = true;
+		recs[cnt].begin_off = i - start;
+		recs[cnt].end_off = i - start + 1;
+		recs[cnt].landing_pad_off = pad - start;
+		cnt++;
+	}
+	err = bpf_exc_attach_info(sub->aux, recs, cnt);
+	if (err)
+		return err;
+	return 0;
+}
+
+int bpf_exc_attach_main_prog(struct bpf_verifier_env *env, struct bpf_prog *prog)
+{
+	if (!env || env->subprog_cnt > 1)
+		return 0;
+	return exc_info_for_subprog(env, prog, 0, 0, prog->len);
+}
+
 static int jit_subprogs(struct bpf_verifier_env *env)
 {
 	struct bpf_prog *prog = env->prog, **func, *tmp;
@@ -1423,6 +1481,10 @@ static int jit_subprogs(struct bpf_verifier_env *env)
 		func[i]->aux->token = prog->aux->token;
 		if (!i)
 			func[i]->aux->exception_boundary = env->seen_exception;
+		err = exc_info_for_subprog(env, func[i], i, subprog_start,
+					   subprog_end);
+		if (err)
+			goto out_free;
 		func[i] = bpf_int_jit_compile(env, func[i]);
 		if (!func[i]->jited) {
 			err = -ENOTSUPP;
@@ -1532,6 +1594,9 @@ static int jit_subprogs(struct bpf_verifier_env *env)
 	prog->aux->bpf_exception_cb = (void *)func[env->exception_callback_subprog]->bpf_func;
 	prog->aux->exception_boundary = func[0]->aux->exception_boundary;
 	prog->aux->stack_arg_sp_adjust = func[0]->aux->stack_arg_sp_adjust;
+	prog->aux->exc = func[0]->aux->exc;
+	func[0]->aux->exc = NULL;
+	prog->aux->epilogue_ip = func[0]->aux->epilogue_ip;
 	bpf_prog_jit_attempt_done(prog);
 	return 0;
 out_free:
@@ -1757,6 +1822,37 @@ static int may_goto_expand(struct bpf_insn *insn_buf, int off, int stack_off,
 	return cnt + tail_cnt;
 }
 
+/*
+ * Put an exit back after every bpf_unwind() call. Nothing reaches it, but it
+ * keeps the frame's epilogue, which is where the unwind sends a frame that
+ * has no landing pad.
+ */
+int bpf_exc_keep_exits(struct bpf_verifier_env *env)
+{
+	int insn_cnt = env->prog->len;
+	struct bpf_insn insn_buf[3];
+	struct bpf_prog *new_prog;
+	int i, delta = 0;
+
+	for (i = 0; i < insn_cnt; i++) {
+		struct bpf_insn *insn = env->prog->insnsi + i + delta;
+
+		if (!bpf_is_unwind_kfunc(insn))
+			continue;
+
+		insn_buf[0] = *insn;
+		insn_buf[1] = BPF_MOV64_IMM(BPF_REG_0, 0);
+		insn_buf[2] = BPF_EXIT_INSN();
+
+		new_prog = bpf_patch_insn_data(env, i + delta, insn_buf, 3);
+		if (!new_prog)
+			return -ENOMEM;
+		delta += 2;
+		env->prog = new_prog;
+	}
+	return 0;
+}
+
 /* Do various post-verification rewrites in a single program pass.
  * These rewrites simplify JIT and interpreter implementations.
  */
@@ -2135,6 +2231,25 @@ int bpf_do_misc_fixups(struct bpf_verifier_env *env)
 			goto next_insn;
 		if (insn->src_reg == BPF_PSEUDO_CALL)
 			goto next_insn;
+		if (bpf_is_unwind_resume_kfunc(insn)) {
+			/*
+			 * A pad's resume is just the frame returning:
+			 * bpf_unwind() already pointed this frame's return
+			 * address at the next pad, so the ordinary epilogue
+			 * carries the unwind on. The verifier checked this exit
+			 * with r0 a known zero, so return zero.
+			 */
+			insn_buf[0] = BPF_MOV64_IMM(BPF_REG_0, 0);
+			insn_buf[1] = BPF_EXIT_INSN();
+			cnt = 2;
+			new_prog = bpf_patch_insn_data(env, i + delta, insn_buf, cnt);
+			if (!new_prog)
+				return -ENOMEM;
+			delta += cnt - 1;
+			env->prog = prog = new_prog;
+			insn = new_prog->insnsi + i + delta;
+			goto next_insn;
+		}
 		if (insn->src_reg == BPF_PSEUDO_KFUNC_CALL) {
 			ret = bpf_fixup_kfunc_call(env, insn, insn_buf, i + delta, &cnt);
 			if (ret)
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index f291611fe578..ea5c4d81f8b7 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -31,6 +31,7 @@
 #include <linux/buildid.h>
 
 #include "../../lib/kstrtox.h"
+#include "exception.h"
 
 /* If kernel subsystem is allowing eBPF programs to call this function,
  * inside its own verifier_ops->get_func_proto() callback it should return
@@ -3416,6 +3417,7 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
 	if (!prog)
 		return !ctx->cnt;
 	ctx->cnt++;
+
 	if (bpf_is_subprog(prog))
 		return true;
 	ctx->aux = prog->aux;
@@ -3424,8 +3426,51 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
 	return false;
 }
 
+struct bpf_unwind_ctx {
+	u32 cnt;
+};
+
+static bool bpf_unwind_rewrite(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra)
+{
+	const struct bpf_cleanup_range *rec;
+	struct bpf_unwind_ctx *ctx = cookie;
+	struct bpf_exception_info *exc;
+	struct bpf_prog *prog;
+
+	rcu_read_lock();
+	prog = bpf_prog_ksym_find(ip);
+	rcu_read_unlock();
+	if (!prog)
+		return !ctx->cnt;
+	ctx->cnt++;
+
+	exc = prog->aux->exc;
+	rec = (exc && exc->nr_ranges) ? bpf_exc_pad_for_ip(prog, ip) : NULL;
+	if (rec) {
+		*ra = rec->pad;
+	} else if (ctx->cnt == 1) {
+		/*
+		 * The frame that called bpf_unwind(). Its return address
+		 * always names the 'r0 = 0; exit' that bpf_exc_keep_exits()
+		 * put after the call, so leave it alone and let the frame
+		 * return through that: running it is what sets the value
+		 * the unwind returns.
+		 */
+	} else if (prog->aux->epilogue_ip) {
+		*ra = prog->aux->epilogue_ip;
+	} else {
+		WARN_ON_ONCE(1);
+		return false;
+	}
+
+	return bpf_is_subprog(prog);
+}
+
 __bpf_kfunc void bpf_unwind(void)
 {
+	struct bpf_unwind_ctx ctx = {};
+
+	arch_bpf_stack_walk_ra(bpf_unwind_rewrite, &ctx);
 }
 
 __bpf_kfunc void bpf_throw(u64 cookie)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 677bab92f624..f2a59716fdab 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22781,6 +22781,9 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 		/* program is valid, convert *(u32*)(ctx + off) accesses */
 		ret = bpf_convert_ctx_accesses(env);
 
+	if (ret == 0)
+		ret = bpf_exc_keep_exits(env);
+
 	if (ret == 0)
 		ret = bpf_do_misc_fixups(env);
 
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (9 preceding siblings ...)
  2026-09-26  5:00 ` [PATCH bpf-next v6 10/21] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
@ 2026-09-26  5:01 ` Yonghong Song
  2026-09-26  5:15   ` sashiko-bot
  2026-09-27 20:39   ` bot+bpf-ci
  2026-09-26  5:01 ` [PATCH bpf-next v6 12/21] bpf, arm64: " Yonghong Song
                   ` (9 subsequent siblings)
  20 siblings, 2 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

arch_bpf_stack_walk_ra() is the ORC walk with the return-address slot
alongside each frame, which unwind_get_return_address_ptr() already hands
out; writing there is what redirects a frame to its landing pad. The
feature is gated on CONFIG_UNWINDER_ORC for the same reason
arch_bpf_stack_walk() is -- there is no other unwinder here to ask.

aux->epilogue_ip comes for free: the JIT already emits one epilogue per
(sub)program and every other exit jumps to it, so the offset it keeps as
ctx->cleanup_addr, as a native address, is it. That field has always been
the epilogue and has nothing to do with the cleanup pads despite the name.

The rest is bookkeeping: build the native cleanup table from the JIT's
addrs[] once the image is final, and emit an ENDBR at each pad head.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 arch/x86/net/bpf_jit_comp.c | 35 ++++++++++++++++++++++++++++++++++-
 1 file changed, 34 insertions(+), 1 deletion(-)

diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index 6c7a0578760e..fb7e8ca1aab2 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -2150,7 +2150,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
 				dst_reg = X86_REG_R9;
 		}
 
-		if (bpf_insn_is_indirect_target(env, bpf_prog, i - 1))
+		if (bpf_insn_is_indirect_target(env, bpf_prog, i - 1) ||
+		    bpf_exc_insn_is_pad(env, bpf_prog, i - 1))
 			EMIT_ENDBR();
 
 		ip = image + addrs[i - 1] + (prog - temp);
@@ -3278,6 +3279,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
 			seen_exit = true;
 			/* Update cleanup_addr */
 			ctx->cleanup_addr = proglen;
+			/* Where an unwind sends a frame with no pad. */
+			bpf_prog->aux->epilogue_ip = (u64)image + proglen;
 			if (bpf_prog_was_classic(bpf_prog) &&
 			    !ns_capable_noaudit(&init_user_ns, CAP_SYS_ADMIN)) {
 				if (emit_spectre_bhb_barrier(&prog, ip, bpf_prog))
@@ -4455,6 +4458,13 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
 		 */
 		bpf_prog_update_insn_ptrs(prog, addrs, image);
 
+		/*
+		 * Same mapping, consumed by the bpf_unwind() walk:
+		 * turn the cleanup records into native address ranges now
+		 * that the image is final.
+		 */
+		bpf_exc_fill_native_ranges(prog, addrs, image);
+
 		/*
 		 * ctx.prog_offset is used when CFI preambles put code *before*
 		 * the function. See emit_cfi(). For FineIBT specifically this code
@@ -4593,6 +4603,11 @@ bool bpf_jit_supports_exceptions(void)
 	return IS_ENABLED(CONFIG_UNWINDER_ORC);
 }
 
+bool bpf_jit_supports_cleanup_pads(void)
+{
+	return IS_ENABLED(CONFIG_UNWINDER_ORC);
+}
+
 bool bpf_jit_supports_private_stack(void)
 {
 	return true;
@@ -4614,6 +4629,24 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
 #endif
 }
 
+void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
+			    void *cookie)
+{
+#if defined(CONFIG_UNWINDER_ORC)
+	struct unwind_state state;
+	unsigned long addr, *ra;
+
+	for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
+	     unwind_next_frame(&state)) {
+		addr = unwind_get_return_address(&state);
+		ra = unwind_get_return_address_ptr(&state);
+		if (!addr || !ra ||
+		    !consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
+			break;
+	}
+#endif
+}
+
 void bpf_arch_poke_desc_update(struct bpf_jit_poke_descriptor *poke,
 			       struct bpf_prog *new, struct bpf_prog *old)
 {
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 12/21] bpf, arm64: Dispatch exception cleanup pads at run time
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (10 preceding siblings ...)
  2026-09-26  5:01 ` [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
@ 2026-09-26  5:01 ` Yonghong Song
  2026-09-26  5:14   ` sashiko-bot
  2026-09-27 20:40   ` bot+bpf-ci
  2026-09-26  5:01 ` [PATCH bpf-next v6 13/21] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
                   ` (8 subsequent siblings)
  20 siblings, 2 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

The JIT half: build the native cleanup table from the JIT's byte offsets
once the image is final, emit a BTI at each pad head, address the frame
through the private stack pointer where one is in use, and record the one
epilogue so a frame the unwind passes over can return straight through it.

The dispatch is arch_bpf_stack_walk_ra(), which hands the unwind the slot a
frame's return address came out of rather than just the address. arm64's
unwinder reads that address from the frame record the callee pushed, so the
slot belongs to the record the previous entry stepped through.

Writing it has to respect pointer authentication: a BPF prologue signs the
link register with PACIASP and the epilogue authenticates it, so what is
written back has to carry the same signature. Its modifier is the stack
pointer the owner was entered with, which is not known here -- so recover
it by re-signing the address the unwinder stripped until that matches what
the slot holds. Only where the CPU implements address authentication, and
only where the slot was signed to begin with.

Two frames are not redirected. The walk's own first frame is not returning
anywhere yet, and a frame whose return the function graph tracer or a
kretprobe has hooked holds the tracer's trampoline in its slot rather than
the address the unwinder reports, so the walk stops there.

bpf_jit_supports_cleanup_pads() can now say yes.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 arch/arm64/kernel/stacktrace.c | 94 ++++++++++++++++++++++++++++++++++
 arch/arm64/net/bpf_jit_comp.c  | 20 +++++++-
 2 files changed, 113 insertions(+), 1 deletion(-)

diff --git a/arch/arm64/kernel/stacktrace.c b/arch/arm64/kernel/stacktrace.c
index 3ebcf8c53fb0..c750c520f24d 100644
--- a/arch/arm64/kernel/stacktrace.c
+++ b/arch/arm64/kernel/stacktrace.c
@@ -445,6 +445,100 @@ noinline noinstr void arch_bpf_stack_walk(bool (*consume_entry)(void *cookie, u6
 	kunwind_stack_walk(arch_bpf_unwind_consume_entry, &data, current, NULL);
 }
 
+struct bpf_unwind_ra_consume_entry_data {
+	bool (*consume_entry)(void *cookie, u64 ip, u64 sp, u64 fp, u64 *ra);
+	void *cookie;
+	unsigned long record;
+	bool seen_first;
+};
+
+static u64 bpf_unwind_sign_ra(u64 ra, u64 modifier)
+{
+	asm volatile(ARM64_ASM_PREAMBLE
+		     ".arch_extension pauth\n"
+		     "	pacia %0, %1"
+		     : "+r" (ra) : "r" (modifier));
+	return ra;
+}
+
+/*
+ * PACIASP's modifier is the stack pointer the owner was entered with: record
+ * + 16 for a BPF prologue, but further up for bpf_unwind()'s own C frame.
+ * Recognise it by re-signing @pc, which the unwinder stripped from @stored.
+ */
+static bool bpf_unwind_ra_modifier(unsigned long record, unsigned long caller_fp,
+				   u64 stored, u64 pc, u64 *modifier)
+{
+	u64 m;
+
+	for (m = record + sizeof(struct frame_record); m <= caller_fp; m += 16) {
+		if (bpf_unwind_sign_ra(pc, m) == stored) {
+			*modifier = m;
+			return true;
+		}
+	}
+	return false;
+}
+
+static bool bpf_unwind_store_ra(unsigned long record, unsigned long caller_fp,
+				u64 pc, u64 ra)
+{
+	struct frame_record *rec = (struct frame_record *)record;
+	u64 stored = READ_ONCE(rec->lr);
+
+	if (system_supports_address_auth() && stored != pc) {
+		u64 modifier;
+
+		if (WARN_ON_ONCE(!bpf_unwind_ra_modifier(record, caller_fp,
+							 stored, pc, &modifier)))
+			return false;
+		ra = bpf_unwind_sign_ra(ra, modifier);
+	}
+	WRITE_ONCE(rec->lr, ra);
+	return true;
+}
+
+static bool
+arch_bpf_unwind_ra_consume_entry(const struct kunwind_state *state, void *cookie)
+{
+	struct bpf_unwind_ra_consume_entry_data *data = cookie;
+	unsigned long record = data->record;
+	bool seen_first = data->seen_first;
+	u64 ra = state->common.pc;
+	bool cont;
+
+	/* The record this frame's return address will have come out of. */
+	data->record = state->common.fp;
+	data->seen_first = true;
+
+	/* The first pc is where the walk runs, not an address it returns to. */
+	if (!seen_first)
+		return true;
+	/* A traced return: the slot holds the tracer's trampoline, not @pc. */
+	if (state->flags.fgraph || state->flags.kretprobe)
+		return false;
+
+	/* A consumer that stops still gets to redirect the frame it stopped on. */
+	cont = data->consume_entry(data->cookie, state->common.pc, 0,
+				   state->common.fp, &ra);
+	if (ra != state->common.pc &&
+	    !bpf_unwind_store_ra(record, state->common.fp, state->common.pc, ra))
+		return false;
+	return cont;
+}
+
+noinline noinstr void arch_bpf_stack_walk_ra(bool (*consume_entry)(void *cookie, u64 ip, u64 sp,
+								   u64 fp, u64 *ra),
+					     void *cookie)
+{
+	struct bpf_unwind_ra_consume_entry_data data = {
+		.consume_entry = consume_entry,
+		.cookie = cookie,
+	};
+
+	kunwind_stack_walk(arch_bpf_unwind_ra_consume_entry, &data, current, NULL);
+}
+
 static const char *state_source_string(const struct kunwind_state *state)
 {
 	switch (state->source) {
diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
index 475e70653454..2422a1ae1256 100644
--- a/arch/arm64/net/bpf_jit_comp.c
+++ b/arch/arm64/net/bpf_jit_comp.c
@@ -10,6 +10,7 @@
 #include <linux/arm-smccc.h>
 #include <linux/bitfield.h>
 #include <linux/bpf.h>
+#include <linux/bpf_verifier.h>
 #include <linux/cfi.h>
 #include <linux/filter.h>
 #include <linux/memory.h>
@@ -1380,7 +1381,8 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
 	int ret;
 	bool sign_extend;
 
-	if (bpf_insn_is_indirect_target(env, ctx->prog, i))
+	if (bpf_insn_is_indirect_target(env, ctx->prog, i) ||
+	    bpf_exc_insn_is_pad(env, ctx->prog, i))
 		emit_bti(A64_BTI_J, ctx);
 
 	switch (code) {
@@ -2423,6 +2425,17 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
 		 * reasons, expects to point to the next instruction)
 		 */
 		bpf_prog_update_insn_ptrs(prog, ctx.offset, ctx.ro_image);
+
+		/*
+		 * Same byte offsets, consumed by the bpf_unwind() walk:
+		 * turn the cleanup records into native address ranges now that
+		 * the image is final.
+		 */
+		bpf_exc_fill_native_ranges(prog, ctx.offset, ctx.ro_image);
+
+		/* Where an unwind sends a frame with no pad. */
+		prog->aux->epilogue_ip = (u64)ctx.ro_image +
+					 ctx.epilogue_offset * AARCH64_INSN_SIZE;
 out_off:
 		if (!ro_header && priv_stack_ptr) {
 			free_percpu(priv_stack_ptr);
@@ -3408,6 +3421,11 @@ bool bpf_jit_supports_exceptions(void)
 	return true;
 }
 
+bool bpf_jit_supports_cleanup_pads(void)
+{
+	return true;
+}
+
 bool bpf_jit_supports_arena(void)
 {
 	return true;
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 13/21] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (11 preceding siblings ...)
  2026-09-26  5:01 ` [PATCH bpf-next v6 12/21] bpf, arm64: " Yonghong Song
@ 2026-09-26  5:01 ` Yonghong Song
  2026-09-26  5:01 ` [PATCH bpf-next v6 14/21] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
                   ` (7 subsequent siblings)
  20 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

LLVM terminates a cleanup landing pad with a call to _Unwind_Resume: the
base unwind ABI's entry point for carrying an unwind on once a frame's
cleanups have run. The kernel provides that terminator as a kfunc, but not
under that name: claiming _Unwind_Resume in the kernel's own symbol table,
for a function whose body never runs, would be needlessly confusing, so it
is called bpf_unwind_resume.

The compiler emits the call but not the declaration: _Unwind_Resume lands
in the object as a plain undefined symbol, with nothing in .ksyms. libbpf
takes every undefined NOTYPE symbol for an extern and refuses one it has no
BTF for ("failed to find BTF for extern '_Unwind_Resume'"), so the program
declares it itself -- extern void _Unwind_Resume(void *) __ksym; -- as the
selftests here do, and as a language runtime emitting cleanup pads has to.
What follows translates that name; it does not manufacture the declaration.

Both load paths take the detour. A direct load resolves the name against
the kernel's BTF while libbpf runs. A light skeleton instead writes the
name into the loader program's blob of bytes, for that program to resolve
when it runs, so the name recorded there has to be translated as well.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/lib/bpf/libbpf.c | 23 ++++++++++++++++++-----
 1 file changed, 18 insertions(+), 5 deletions(-)

diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index 238e35c1ee3f..7eb15a013a83 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -8797,6 +8797,16 @@ static void fixup_verifier_log(struct bpf_program *prog, char *buf, size_t buf_s
 	}
 }
 
+/*
+ * LLVM terminates a cleanup landing pad with a call to _Unwind_Resume, the
+ * base unwind ABI's entry point for carrying an unwind on once a frame's
+ * cleanups have run. The kernel knows it as bpf_unwind_resume.
+ */
+static const char *kern_extern_name(const char *name)
+{
+	return strcmp(name, "_Unwind_Resume") ? name : "bpf_unwind_resume";
+}
+
 static int bpf_program_record_relos(struct bpf_program *prog)
 {
 	struct bpf_object *obj = prog->obj;
@@ -8813,12 +8823,12 @@ static int bpf_program_record_relos(struct bpf_program *prog)
 				continue;
 			kind = btf_is_var(btf__type_by_id(obj->btf, ext->btf_id)) ?
 				BTF_KIND_VAR : BTF_KIND_FUNC;
-			bpf_gen__record_extern(obj->gen_loader, ext->name,
+			bpf_gen__record_extern(obj->gen_loader, kern_extern_name(ext->name),
 					       ext->is_weak, !ext->ksym.type_id,
 					       true, kind, relo->insn_idx);
 			break;
 		case RELO_EXTERN_CALL:
-			bpf_gen__record_extern(obj->gen_loader, ext->name,
+			bpf_gen__record_extern(obj->gen_loader, kern_extern_name(ext->name),
 					       ext->is_weak, false, false, BTF_KIND_FUNC,
 					       relo->insn_idx);
 			break;
@@ -9294,17 +9304,20 @@ static int bpf_object__resolve_ksym_func_btf_id(struct bpf_object *obj,
 	struct module_btf *mod_btf = NULL;
 	const struct btf_type *kern_func;
 	struct btf *kern_btf = NULL;
+	const char *local_name, *kern_name;
 	int ret;
 
 	local_func_proto_id = ext->ksym.type_id;
 
-	kfunc_id = find_ksym_btf_id(obj, ext->essent_name ?: ext->name, BTF_KIND_FUNC, &kern_btf,
-				    &mod_btf);
+	local_name = ext->essent_name ?: ext->name;
+	kern_name = kern_extern_name(local_name);
+
+	kfunc_id = find_ksym_btf_id(obj, kern_name, BTF_KIND_FUNC, &kern_btf, &mod_btf);
 	if (kfunc_id < 0) {
 		if (kfunc_id == -ESRCH && ext->is_weak)
 			return 0;
 		pr_warn("extern (func ksym) '%s': not found in kernel or module BTFs\n",
-			ext->name);
+			kern_name != local_name ? kern_name : ext->name);
 		return kfunc_id;
 	}
 
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 14/21] libbpf: Add cleanup_info to bpf_prog_load_opts
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (12 preceding siblings ...)
  2026-09-26  5:01 ` [PATCH bpf-next v6 13/21] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
@ 2026-09-26  5:01 ` Yonghong Song
  2026-09-26  5:01 ` [PATCH bpf-next v6 15/21] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
                   ` (6 subsequent siblings)
  20 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

Let a caller hand the kernel an exception cleanup table: struct
bpf_prog_load_opts grows cleanup_info, cleanup_info_cnt and
cleanup_info_rec_size, and bpf_prog_load() passes all three on to
BPF_PROG_LOAD.

The attr size it computes grows by one field rather than three:
cleanup_info_cnt is the last of them in the BPF_PROG_LOAD attr, so
offsetofend() on it already covers the other two.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/lib/bpf/bpf.c | 6 +++++-
 tools/lib/bpf/bpf.h | 7 ++++++-
 2 files changed, 11 insertions(+), 2 deletions(-)

diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
index b49822d212ae..2e761f020b14 100644
--- a/tools/lib/bpf/bpf.c
+++ b/tools/lib/bpf/bpf.c
@@ -295,7 +295,7 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
 		  const struct bpf_insn *insns, size_t insn_cnt,
 		  struct bpf_prog_load_opts *opts)
 {
-	const size_t attr_sz = offsetofend(union bpf_attr, keyring_id);
+	const size_t attr_sz = offsetofend(union bpf_attr, cleanup_info_cnt);
 	void *finfo = NULL, *linfo = NULL;
 	const char *func_info, *line_info;
 	__u32 log_size, log_level, attach_prog_fd, attach_btf_obj_fd;
@@ -370,6 +370,10 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
 	attr.fd_array = ptr_to_u64(OPTS_GET(opts, fd_array, NULL));
 	attr.fd_array_cnt = OPTS_GET(opts, fd_array_cnt, 0);
 
+	attr.cleanup_info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
+	attr.cleanup_info = ptr_to_u64(OPTS_GET(opts, cleanup_info, NULL));
+	attr.cleanup_info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
+
 	if (log_level) {
 		attr.log_buf = ptr_to_u64(log_buf);
 		attr.log_size = log_size;
diff --git a/tools/lib/bpf/bpf.h b/tools/lib/bpf/bpf.h
index 826d9cc9ab65..2c292d56c9a3 100644
--- a/tools/lib/bpf/bpf.h
+++ b/tools/lib/bpf/bpf.h
@@ -128,9 +128,14 @@ struct bpf_prog_load_opts {
 
 	/* if set, provides the length of fd_array */
 	__u32 fd_array_cnt;
+
+	/* exception cleanup table, from the .bpf_cleanup section */
+	const void *cleanup_info;
+	__u32 cleanup_info_cnt;
+	__u32 cleanup_info_rec_size;
 	size_t :0;
 };
-#define bpf_prog_load_opts__last_field fd_array_cnt
+#define bpf_prog_load_opts__last_field cleanup_info_rec_size
 
 LIBBPF_API int bpf_prog_load(enum bpf_prog_type prog_type,
 			     const char *prog_name, const char *license,
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 15/21] libbpf: Collect .bpf_cleanup records and pass them to the kernel
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (13 preceding siblings ...)
  2026-09-26  5:01 ` [PATCH bpf-next v6 14/21] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
@ 2026-09-26  5:01 ` Yonghong Song
  2026-09-27 20:39   ` bot+bpf-ci
  2026-09-26  5:01 ` [PATCH bpf-next v6 16/21] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
                   ` (5 subsequent siblings)
  20 siblings, 1 reply; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

Parse the compiler-emitted .bpf_cleanup section and hand the resulting
table to BPF_PROG_LOAD.

Each record is three 4-byte fields, and each field is a byte offset into
some code section named by a matching .rel.bpf_cleanup relocation.
bpf_object__init_cleanup_info() resolves both halves once at open time and
keeps (section, instruction index) pairs; it rejects a record whose field
has no relocation, whose relocation is not one of the two 32-bit types that
spell a data reference to a code section -- R_BPF_64_NODYLD32 from LLVM,
R_BPF_64_ABS32 from GNU as -- or whose offset is not instruction aligned.
LLVM leaves the section's sh_entsize unset, so a section that declares one
of a different size is refused, that being the only way a producer can say
it means something wider.

The records are sorted by begin_off once the offsets are final, since the
kernel wants the table sorted with disjoint ranges to find the record
covering a call site with a binary search, and they arrive in .bpf_cleanup
order, which says nothing about where the subprograms they describe were
appended. Overlapping ranges are reported here, where the program name and
both regions are still at hand.

bpf_object_load_prog() then passes the per-program table through the
bpf_prog_load() options added in the previous patch, with the record size
carried on the program the way func_info and line_info carry theirs.
bpf_program__clone() carries it too -- that is the load path veristat uses,
and without the table the kernel sees landing pads nothing reaches and
refuses the program with "unreachable insn".

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/lib/bpf/libbpf.c          | 319 ++++++++++++++++++++++++++++++++
 tools/lib/bpf/libbpf_internal.h |   3 +
 2 files changed, 322 insertions(+)

diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index 7eb15a013a83..4e9584fff264 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -516,6 +516,11 @@ struct bpf_program {
 	void *line_info;
 	__u32 line_info_rec_size;
 	__u32 line_info_cnt;
+
+	struct bpf_cleanup_info *cleanup_info;
+	__u32 cleanup_info_rec_size;
+	__u32 cleanup_info_cnt;
+
 	__u32 prog_flags;
 	__u8  hash[SHA256_DIGEST_LENGTH];
 
@@ -552,6 +557,7 @@ struct bpf_struct_ops {
 #define STRUCT_OPS_SEC ".struct_ops"
 #define STRUCT_OPS_LINK_SEC ".struct_ops.link"
 #define ARENA_SEC ".addr_space.1"
+#define CLEANUP_SEC ".bpf_cleanup"
 
 enum libbpf_map_type {
 	LIBBPF_MAP_UNSPEC,
@@ -683,6 +689,25 @@ struct elf_sec_desc {
 	Elf_Data *data;
 };
 
+#define CLEANUP_REC_FIELDS	(sizeof(struct bpf_cleanup_info) / sizeof(__u32))
+
+/* Index of each field of struct bpf_cleanup_info, read as an array of __u32. */
+enum {
+	CLEANUP_REC_BEGIN,
+	CLEANUP_REC_END,
+	CLEANUP_REC_PAD,
+};
+
+/*
+ * One (begin, end, landing_pad) triple from .bpf_cleanup, each field resolved
+ * from its relocation to a section and an instruction index within it. Final
+ * indices wait for subprogram placement, which differs per main program.
+ */
+struct cleanup_raw_rec {
+	int sec_idx[CLEANUP_REC_FIELDS];
+	size_t insn_idx[CLEANUP_REC_FIELDS];
+};
+
 struct elf_state {
 	int fd;
 	const void *obj_buf;
@@ -702,6 +727,8 @@ struct elf_state {
 	bool has_st_ops;
 	int arena_data_shndx;
 	int jumptables_data_shndx;
+	Elf_Data *cleanup_data;
+	int cleanup_shndx;
 };
 
 struct usdt_manager;
@@ -779,6 +806,9 @@ struct bpf_object {
 	void *jumptables_data;
 	size_t jumptables_data_sz;
 
+	struct cleanup_raw_rec *cleanup_recs;
+	size_t cleanup_rec_cnt;
+
 	struct {
 		struct bpf_program *prog;
 		unsigned int sym_off;
@@ -848,7 +878,10 @@ static void bpf_program__exit(struct bpf_program *prog)
 	zfree(&prog->sec_name);
 	zfree(&prog->insns);
 	zfree(&prog->reloc_desc);
+	zfree(&prog->cleanup_info);
 
+	prog->cleanup_info_rec_size = 0;
+	prog->cleanup_info_cnt = 0;
 	prog->nr_reloc = 0;
 	prog->insns_cnt = 0;
 	prog->sec_idx = -1;
@@ -1600,6 +1633,7 @@ static struct bpf_object *bpf_object__new(const char *path,
 	obj->efile.obj_buf = obj_buf;
 	obj->efile.obj_buf_sz = obj_buf_sz;
 	obj->efile.btf_maps_shndx = -1;
+	obj->efile.cleanup_shndx = -1;
 	obj->kconfig_map_idx = -1;
 	obj->arena_map_idx = -1;
 
@@ -4100,6 +4134,9 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
 				sec_desc->shdr = sh;
 				sec_desc->data = data;
 				obj->efile.has_st_ops = true;
+			} else if (strcmp(name, CLEANUP_SEC) == 0) {
+				obj->efile.cleanup_data = data;
+				obj->efile.cleanup_shndx = idx;
 			} else if (strcmp(name, ARENA_SEC) == 0) {
 				obj->efile.arena_data = data;
 				obj->efile.arena_data_shndx = idx;
@@ -4135,6 +4172,7 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
 			    strcmp(name, ".rel" STRUCT_OPS_LINK_SEC) &&
 			    strcmp(name, ".rel?" STRUCT_OPS_SEC) &&
 			    strcmp(name, ".rel?" STRUCT_OPS_LINK_SEC) &&
+			    strcmp(name, ".rel" CLEANUP_SEC) &&
 			    strcmp(name, ".rel" MAPS_ELF_SEC)) {
 				pr_info("elf: skipping relo section(%d) %s for section(%d) %s\n",
 					idx, name, targ_sec_idx,
@@ -4915,6 +4953,249 @@ static struct bpf_program *find_prog_by_sec_insn(const struct bpf_object *obj,
 	return NULL;
 }
 
+static int bpf_object__init_cleanup_info(struct bpf_object *obj)
+{
+	Elf_Data *data = obj->efile.cleanup_data;
+	Elf_Data *relo = NULL;
+	size_t i, nrels, nslots, nrecs;
+	struct cleanup_raw_rec *recs;
+	int *slot_sec, ret = 0;
+	size_t *slot_val;
+	const __u32 *vals;
+	Elf64_Shdr *sh;
+	bool native;
+
+	if (!data || obj->efile.cleanup_shndx < 0 || !data->d_size)
+		return 0;
+
+	native = is_native_endianness(obj);
+
+	for (i = 0; i < obj->efile.sec_cnt; i++) {
+		struct elf_sec_desc *sd = &obj->efile.secs[i];
+
+		if (sd->sec_type == SEC_RELO && sd->shdr &&
+		    sd->shdr->sh_info == (Elf64_Word)obj->efile.cleanup_shndx) {
+			relo = sd->data;
+			break;
+		}
+	}
+	if (!relo) {
+		pr_warn("%s present without relocations\n", CLEANUP_SEC);
+		return -LIBBPF_ERRNO__FORMAT;
+	}
+	/*
+	 * LLVM leaves sh_entsize unset, so this only bites a producer that
+	 * declares a record size -- which is the one able to tell us it means
+	 * something other than three 4-byte fields.
+	 */
+	sh = elf_sec_hdr(obj, elf_sec_by_idx(obj, obj->efile.cleanup_shndx));
+	if (sh && sh->sh_entsize &&
+	    sh->sh_entsize != sizeof(struct bpf_cleanup_info)) {
+		pr_warn("%s record size %zu is not the expected %zu\n", CLEANUP_SEC,
+			(size_t)sh->sh_entsize, sizeof(struct bpf_cleanup_info));
+		return -LIBBPF_ERRNO__FORMAT;
+	}
+	if (data->d_size % sizeof(struct bpf_cleanup_info)) {
+		pr_warn("%s size %zu is not a multiple of the record size %zu\n",
+			CLEANUP_SEC, data->d_size, sizeof(struct bpf_cleanup_info));
+		return -LIBBPF_ERRNO__FORMAT;
+	}
+
+	vals = data->d_buf;
+	nslots = data->d_size / sizeof(__u32);
+	nrecs = data->d_size / sizeof(struct bpf_cleanup_info);
+
+	slot_sec = calloc(nslots, sizeof(*slot_sec));
+	slot_val = calloc(nslots, sizeof(*slot_val));
+	recs = calloc(nrecs ?: 1, sizeof(*recs));
+	if (!slot_sec || !slot_val || !recs) {
+		ret = -ENOMEM;
+		goto out;
+	}
+	for (i = 0; i < nslots; i++)
+		slot_sec[i] = -1;
+
+	/* One relocation per 4-byte field, naming the section it points into. */
+	nrels = relo->d_size / sizeof(Elf64_Rel);
+	for (i = 0; i < nrels; i++) {
+		Elf64_Rel *rel = elf_rel_by_idx(relo, i);
+		Elf64_Sym *sym = elf_sym_by_idx(obj, ELF64_R_SYM(rel->r_info));
+		size_t type = ELF64_R_TYPE(rel->r_info);
+		size_t slot = rel->r_offset / sizeof(__u32);
+
+		if (type != R_BPF_64_NODYLD32 && type != R_BPF_64_ABS32) {
+			pr_warn("%s: relocation %zu has unexpected type %zu\n",
+				CLEANUP_SEC, i, type);
+			ret = -LIBBPF_ERRNO__FORMAT;
+			goto out;
+		}
+		if (!sym || slot >= nslots || rel->r_offset % sizeof(__u32)) {
+			pr_warn("%s: bad relocation %zu\n", CLEANUP_SEC, i);
+			ret = -LIBBPF_ERRNO__FORMAT;
+			goto out;
+		}
+		slot_sec[slot] = sym->st_shndx;
+		/*
+		 * The addend lives in the section data, which libelf leaves in
+		 * the object's byte order; a non-section symbol additionally
+		 * contributes its own value.
+		 */
+		slot_val[slot] = (native ? vals[slot] : bswap_32(vals[slot])) +
+				 sym->st_value;
+	}
+
+	for (i = 0; i < nslots; i++) {
+		struct cleanup_raw_rec *rec = &recs[i / CLEANUP_REC_FIELDS];
+		size_t field = i % CLEANUP_REC_FIELDS;
+
+		if (slot_sec[i] < 0) {
+			pr_warn("%s: field %zu has no relocation\n", CLEANUP_SEC, i);
+			ret = -LIBBPF_ERRNO__FORMAT;
+			goto out;
+		}
+		if (slot_val[i] % BPF_INSN_SZ) {
+			pr_warn("%s: field %zu offset %zu is not instruction aligned\n",
+				CLEANUP_SEC, i, slot_val[i]);
+			ret = -LIBBPF_ERRNO__FORMAT;
+			goto out;
+		}
+		rec->sec_idx[field] = slot_sec[i];
+		rec->insn_idx[field] = slot_val[i] / BPF_INSN_SZ;
+	}
+
+	obj->cleanup_recs = recs;
+	obj->cleanup_rec_cnt = nrecs;
+	recs = NULL;
+out:
+	free(recs);
+	free(slot_val);
+	free(slot_sec);
+	return ret;
+}
+
+static int cmp_cleanup_info(const void *a, const void *b)
+{
+	const struct bpf_cleanup_info *x = a, *y = b;
+
+	if (x->begin_off == y->begin_off)
+		return 0;
+	return x->begin_off < y->begin_off ? -1 : 1;
+}
+
+static int bpf_prog_collect_cleanup_info(struct bpf_object *obj,
+					 struct bpf_program *prog)
+{
+	size_t i;
+	int j;
+
+	for (i = 0; i < obj->cleanup_rec_cnt; i++) {
+		struct cleanup_raw_rec *raw = &obj->cleanup_recs[i];
+		struct bpf_program *owner = NULL;
+		struct bpf_cleanup_info ci = {};
+		__u32 fields[CLEANUP_REC_FIELDS];
+		void *tmp;
+
+		if (raw->sec_idx[CLEANUP_REC_BEGIN] == raw->sec_idx[CLEANUP_REC_END] &&
+		    raw->insn_idx[CLEANUP_REC_BEGIN] >= raw->insn_idx[CLEANUP_REC_END]) {
+			pr_warn("%s: record %zu is an empty range [%zu,%zu)\n",
+				CLEANUP_SEC, i, raw->insn_idx[CLEANUP_REC_BEGIN],
+				raw->insn_idx[CLEANUP_REC_END]);
+			return -LIBBPF_ERRNO__FORMAT;
+		}
+
+		for (j = 0; j < CLEANUP_REC_FIELDS; j++) {
+			size_t idx = raw->insn_idx[j], final;
+			struct bpf_program *p;
+
+			/*
+			 * An exclusive end may name the instruction past the
+			 * last of a function, so ask about the last one the
+			 * range covers, the way the kernel does.
+			 */
+			if (j == CLEANUP_REC_END) {
+				if (!idx) {
+					pr_warn("%s: record %zu ends at instruction 0\n",
+						CLEANUP_SEC, i);
+					return -LIBBPF_ERRNO__FORMAT;
+				}
+				idx--;
+			}
+
+			p = find_prog_by_sec_insn(obj, raw->sec_idx[j], idx);
+			if (!p) {
+				pr_warn("%s: record %zu field %d is not inside a function\n",
+					CLEANUP_SEC, i, j);
+				return -LIBBPF_ERRNO__FORMAT;
+			}
+			if (!owner) {
+				owner = p;
+			} else if (owner != p) {
+				pr_warn("%s: record %zu spans functions '%s' and '%s'\n",
+					CLEANUP_SEC, i, owner->name, p->name);
+				return -LIBBPF_ERRNO__FORMAT;
+			}
+
+			if (owner == prog) {
+				final = raw->insn_idx[j] - prog->sec_insn_off;
+			} else if (prog_is_subprog(obj, owner) && owner->sub_insn_off) {
+				/*
+				 * sub_insn_off is where this subprogram was
+				 * appended to the main program being relocated;
+				 * zero means it is not part of it.
+				 */
+				final = owner->sub_insn_off +
+					raw->insn_idx[j] - owner->sec_insn_off;
+			} else {
+				owner = NULL;
+				break;
+			}
+			fields[j] = final;
+		}
+		if (!owner)
+			continue;
+
+		ci.begin_off = fields[CLEANUP_REC_BEGIN];
+		ci.end_off = fields[CLEANUP_REC_END];
+		ci.landing_pad_off = fields[CLEANUP_REC_PAD];
+
+		tmp = libbpf_reallocarray(prog->cleanup_info, prog->cleanup_info_cnt + 1,
+					  sizeof(*prog->cleanup_info));
+		if (!tmp)
+			return -ENOMEM;
+		prog->cleanup_info = tmp;
+		prog->cleanup_info_rec_size = sizeof(struct bpf_cleanup_info);
+		prog->cleanup_info[prog->cleanup_info_cnt++] = ci;
+
+		pr_debug("prog '%s': cleanup region [%u,%u) -> landing pad %u\n",
+			 prog->name, ci.begin_off, ci.end_off, ci.landing_pad_off);
+	}
+
+	if (!prog->cleanup_info_cnt)
+		return 0;
+
+	if (prog->cleanup_info_cnt > INT32_MAX / sizeof(struct bpf_cleanup_info)) {
+		pr_warn("prog '%s': too many cleanup records: %u\n",
+			prog->name, prog->cleanup_info_cnt);
+		return -LIBBPF_ERRNO__FORMAT;
+	}
+
+	qsort(prog->cleanup_info, prog->cleanup_info_cnt,
+	      sizeof(*prog->cleanup_info), cmp_cleanup_info);
+	for (i = 1; i < prog->cleanup_info_cnt; i++) {
+		struct bpf_cleanup_info *prev = &prog->cleanup_info[i - 1];
+		struct bpf_cleanup_info *cur = &prog->cleanup_info[i];
+
+		if (cur->begin_off < prev->end_off) {
+			pr_warn("prog '%s': overlapping cleanup regions [%u,%u) and [%u,%u)\n",
+				prog->name, prev->begin_off, prev->end_off,
+				cur->begin_off, cur->end_off);
+			return -LIBBPF_ERRNO__FORMAT;
+		}
+	}
+
+	return 0;
+}
+
 static int
 bpf_object__collect_prog_relos(struct bpf_object *obj, Elf64_Shdr *shdr, Elf_Data *data)
 {
@@ -7894,6 +8175,13 @@ static int bpf_object__relocate(struct bpf_object *obj, const char *targ_btf_pat
 					return err;
 			}
 		}
+
+		err = bpf_prog_collect_cleanup_info(obj, prog);
+		if (err) {
+			pr_warn("prog '%s': failed to collect cleanup info: %s\n",
+				prog->name, errstr(err));
+			return err;
+		}
 	}
 	for (i = 0; i < obj->nr_programs; i++) {
 		prog = &obj->programs[i];
@@ -8176,6 +8464,9 @@ static int bpf_object__collect_relos(struct bpf_object *obj)
 			return -LIBBPF_ERRNO__INTERNAL;
 		}
 
+		if (idx == obj->efile.cleanup_shndx)
+			continue;
+
 		if (obj->efile.secs[idx].sec_type == SEC_RODATA)
 			err = bpf_object__collect_rodata_relos(obj, shdr, data);
 		else if (obj->efile.secs[idx].sec_type == SEC_ST_OPS)
@@ -8469,6 +8760,11 @@ static int bpf_object_load_prog(struct bpf_object *obj, struct bpf_program *prog
 		load_attr.line_info_rec_size = prog->line_info_rec_size;
 		load_attr.line_info_cnt = prog->line_info_cnt;
 	}
+	if (prog->cleanup_info_cnt) {
+		load_attr.cleanup_info = prog->cleanup_info;
+		load_attr.cleanup_info_cnt = prog->cleanup_info_cnt;
+		load_attr.cleanup_info_rec_size = prog->cleanup_info_rec_size;
+	}
 	load_attr.log_level = log_level;
 	load_attr.prog_flags = prog->prog_flags;
 	load_attr.fd_array = obj->fd_array;
@@ -9062,6 +9358,7 @@ static struct bpf_object *bpf_object_open(const char *path, const void *obj_buf,
 	err = err ? : bpf_object__init_maps(obj, opts);
 	err = err ? : bpf_object_init_progs(obj, opts);
 	err = err ? : bpf_object__collect_relos(obj);
+	err = err ? : bpf_object__init_cleanup_info(obj);
 	if (err)
 		goto out;
 
@@ -10195,6 +10492,9 @@ void bpf_object__close(struct bpf_object *obj)
 	zfree(&obj->jumptables_data);
 	obj->jumptables_data_sz = 0;
 
+	zfree(&obj->cleanup_recs);
+	obj->cleanup_rec_cnt = 0;
+
 	for (i = 0; i < obj->jumptable_map_cnt; i++)
 		close(obj->jumptable_maps[i].fd);
 	zfree(&obj->jumptable_maps);
@@ -10599,6 +10899,25 @@ int bpf_program__clone(struct bpf_program *prog, const struct bpf_prog_load_opts
 		attr.line_info_rec_size = info ? info_rec_size : prog->line_info_rec_size;
 	}
 
+	/* exception cleanup table */
+	info = OPTS_GET(opts, cleanup_info, NULL);
+	info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
+	info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
+	if (!!info != !!info_cnt || !!info != !!info_rec_size) {
+		pr_warn("prog '%s': cleanup_info, cleanup_info_cnt, and cleanup_info_rec_size must all be specified or all omitted\n",
+			prog->name);
+		return libbpf_err(-EINVAL);
+	}
+	if (info) {
+		attr.cleanup_info = info;
+		attr.cleanup_info_cnt = info_cnt;
+		attr.cleanup_info_rec_size = info_rec_size;
+	} else if (prog->cleanup_info_cnt) {
+		attr.cleanup_info = prog->cleanup_info;
+		attr.cleanup_info_cnt = prog->cleanup_info_cnt;
+		attr.cleanup_info_rec_size = prog->cleanup_info_rec_size;
+	}
+
 	/* Logging is caller-controlled; no fallback to prog/obj log settings */
 	attr.log_buf = OPTS_GET(opts, log_buf, NULL);
 	attr.log_size = OPTS_GET(opts, log_size, 0);
diff --git a/tools/lib/bpf/libbpf_internal.h b/tools/lib/bpf/libbpf_internal.h
index 546f65b95cf4..f1630f03d5f5 100644
--- a/tools/lib/bpf/libbpf_internal.h
+++ b/tools/lib/bpf/libbpf_internal.h
@@ -56,6 +56,9 @@
 #ifndef R_BPF_64_ABS32
 #define R_BPF_64_ABS32 3
 #endif
+#ifndef R_BPF_64_NODYLD32
+#define R_BPF_64_NODYLD32 4
+#endif
 #ifndef R_BPF_64_32
 #define R_BPF_64_32 10
 #endif
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 16/21] libbpf: Carry the exception cleanup table through the light skeleton
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (14 preceding siblings ...)
  2026-09-26  5:01 ` [PATCH bpf-next v6 15/21] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
@ 2026-09-26  5:01 ` Yonghong Song
  2026-09-26  5:01 ` [PATCH bpf-next v6 17/21] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
                   ` (4 subsequent siblings)
  20 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

A light skeleton does not call bpf_prog_load(). bpf_gen__prog_load() builds
its own union bpf_attr field by field, and the loader program it emits is
what issues BPF_PROG_LOAD when the skeleton runs -- so a program loaded
this way reached the kernel without the table the previous patch collected
for it, and the verifier refused it with "unreachable insn", which names
neither the skeleton nor the table.

Carry it the way func_info and line_info are carried: the records go into
the loader's blob of bytes, the count and record size into the attr, and a
relocation stores the blob's address into attr.cleanup_info once that
address is known. The attr grows to its new last field, cleanup_info_cnt.
Records are 4-byte fields like the other info blobs, so a cross-endian
build has to swap them too.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/lib/bpf/gen_loader.c      | 29 +++++++++++++++++++++++++----
 tools/lib/bpf/libbpf_internal.h |  7 +++++++
 2 files changed, 32 insertions(+), 4 deletions(-)

diff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c
index 251392aa8b41..d10cd0672475 100644
--- a/tools/lib/bpf/gen_loader.c
+++ b/tools/lib/bpf/gen_loader.c
@@ -994,13 +994,15 @@ static void cleanup_relos(struct bpf_gen *gen, int insns)
 	cleanup_core_relo(gen);
 }
 
-/* Convert func, line, and core relo info blobs to target endianness */
+/* Convert func, line, core relo and cleanup info blobs to target endianness */
 static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
-			    int core_relos, struct bpf_prog_load_opts *load_attr)
+			    int core_relos, int cleanup_info,
+			    struct bpf_prog_load_opts *load_attr)
 {
 	struct bpf_func_info *fi = gen->data_start + func_info;
 	struct bpf_line_info *li = gen->data_start + line_info;
 	struct bpf_core_relo *cr = gen->data_start + core_relos;
+	struct bpf_cleanup_info *ci = gen->data_start + cleanup_info;
 	int i;
 
 	for (i = 0; i < load_attr->func_info_cnt; i++)
@@ -1011,6 +1013,9 @@ static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
 
 	for (i = 0; i < gen->core_relo_cnt; i++)
 		bpf_core_relo_bswap(cr++);
+
+	for (i = 0; i < load_attr->cleanup_info_cnt; i++)
+		bpf_cleanup_info_bswap(ci++);
 }
 
 void bpf_gen__prog_load(struct bpf_gen *gen,
@@ -1024,8 +1029,11 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
 			       load_attr->line_info_rec_size;
 	int core_relo_tot_sz = gen->core_relo_cnt *
 			       sizeof(struct bpf_core_relo);
+	int cleanup_info_tot_sz = load_attr->cleanup_info_cnt *
+				  load_attr->cleanup_info_rec_size;
 	int prog_load_attr, license_off, insns_off, func_info, line_info, core_relos;
-	int attr_size = offsetofend(union bpf_attr, core_relo_rec_size);
+	int attr_size = offsetofend(union bpf_attr, cleanup_info_cnt);
+	int cleanup_info;
 	union bpf_attr attr;
 
 	memset(&attr, 0, attr_size);
@@ -1074,9 +1082,17 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
 		 core_relos, gen->core_relo_cnt,
 		 sizeof(struct bpf_core_relo));
 
+	attr.cleanup_info_rec_size = tgt_endian(load_attr->cleanup_info_rec_size);
+	attr.cleanup_info_cnt = tgt_endian(load_attr->cleanup_info_cnt);
+	cleanup_info = add_data(gen, load_attr->cleanup_info, cleanup_info_tot_sz);
+	pr_debug("gen: prog_load: cleanup_info: off %d cnt %u rec size %u\n",
+		 cleanup_info, load_attr->cleanup_info_cnt,
+		 load_attr->cleanup_info_rec_size);
+
 	/* convert all info blobs to target endianness */
 	if (gen->swapped_endian && !gen->error)
-		info_blob_bswap(gen, func_info, line_info, core_relos, load_attr);
+		info_blob_bswap(gen, func_info, line_info, core_relos, cleanup_info,
+				load_attr);
 
 	libbpf_strlcpy(attr.prog_name, prog_name, sizeof(attr.prog_name));
 	prog_load_attr = add_data(gen, &attr, attr_size);
@@ -1098,6 +1114,11 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
 	/* populate union bpf_attr with a pointer to core_relos */
 	emit_rel_store(gen, attr_field(prog_load_attr, core_relos), core_relos);
 
+	/* with no records there is no blob of them to point the attr at */
+	if (load_attr->cleanup_info_cnt)
+		emit_rel_store(gen, attr_field(prog_load_attr, cleanup_info),
+			       cleanup_info);
+
 	/* populate union bpf_attr fd_array with a pointer to data where map_fds are saved */
 	emit_rel_store(gen, attr_field(prog_load_attr, fd_array), gen->fd_array);
 
diff --git a/tools/lib/bpf/libbpf_internal.h b/tools/lib/bpf/libbpf_internal.h
index f1630f03d5f5..9d341839ca74 100644
--- a/tools/lib/bpf/libbpf_internal.h
+++ b/tools/lib/bpf/libbpf_internal.h
@@ -572,6 +572,13 @@ static inline void bpf_core_relo_bswap(struct bpf_core_relo *i)
 	i->kind = bswap_32(i->kind);
 }
 
+static inline void bpf_cleanup_info_bswap(struct bpf_cleanup_info *i)
+{
+	i->begin_off = bswap_32(i->begin_off);
+	i->end_off = bswap_32(i->end_off);
+	i->landing_pad_off = bswap_32(i->landing_pad_off);
+}
+
 enum btf_field_iter_kind {
 	BTF_FIELD_ITER_IDS,
 	BTF_FIELD_ITER_STRS,
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 17/21] libbpf: Let the static linker carry .bpf_cleanup relocations
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (15 preceding siblings ...)
  2026-09-26  5:01 ` [PATCH bpf-next v6 16/21] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
@ 2026-09-26  5:01 ` Yonghong Song
  2026-09-26  5:01 ` [PATCH bpf-next v6 18/21] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
                   ` (3 subsequent siblings)
  20 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

An object that carries a compiler-emitted exception cleanup table cannot be
linked today. The table's fields are byte offsets into a code section,
materialised by a 32-bit relocation against that section's symbol with the
offset itself as the implicit addend, and the linker rejects both halves of
that: the relocation type is not in the list it accepts, and a relocation
against an STT_SECTION symbol from a non-executable section is an outright
error.

Both spellings of that relocation have to be taken. LLVM emits
R_BPF_64_NODYLD32 for a .long against a section symbol; GNU as emits
R_BPF_64_ABS32, which is what bpf_reloc_type_lookup() maps BFD_RELOC_32 to.
They describe the same value, and the selftests are built with both
compilers.

Keying on the relocation type rather than the section name means any
non-executable section can reach the new arm, where an unconditional
refusal used to stand. That refusal was covering two things the arm now has
to do itself: a target may be SHT_NOBITS, which extend_sec() leaves with no
raw_data, and r_offset is alignment-checked only where the section holds
instructions. The arm rejects both, and bounds the offset against the
section size before writing through it.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/lib/bpf/linker.c | 36 +++++++++++++++++++++++++++++++++++-
 1 file changed, 35 insertions(+), 1 deletion(-)

diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
index 53f64a1a1f25..26a5e53df6f0 100644
--- a/tools/lib/bpf/linker.c
+++ b/tools/lib/bpf/linker.c
@@ -1036,7 +1036,8 @@ static int linker_sanity_check_elf_relos(struct src_obj *obj, struct src_sec *se
 		size_t sym_type = ELF64_R_TYPE(relo->r_info);
 
 		if (sym_type != R_BPF_64_64 && sym_type != R_BPF_64_32 &&
-		    sym_type != R_BPF_64_ABS64 && sym_type != R_BPF_64_ABS32) {
+		    sym_type != R_BPF_64_ABS64 && sym_type != R_BPF_64_ABS32 &&
+		    sym_type != R_BPF_64_NODYLD32) {
 			pr_warn("ELF relo #%d in section #%zu has unexpected type %zu in %s\n",
 				i, sec->sec_idx, sym_type, obj->filename);
 			return -EINVAL;
@@ -2258,6 +2259,7 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
 			if (ELF64_ST_TYPE(src_sym->st_info) == STT_SECTION) {
 				struct src_sec *sec = &obj->secs[src_sym->st_shndx];
 				struct bpf_insn *insn;
+				__u32 *val;
 
 				if (src_linked_sec->shdr->sh_flags & SHF_EXECINSTR) {
 					/* calls to the very first static function inside
@@ -2292,6 +2294,38 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
 					if (linker->swapped_endian)
 						off = bswap_64(off);
 					memcpy(ptr, &off, sizeof(off));
+				} else if (sym_type == R_BPF_64_NODYLD32 ||
+					   sym_type == R_BPF_64_ABS32) {
+					/*
+					 * A byte offset into a code section,
+					 * stored in place. LLVM spells this
+					 * relocation NODYLD32 and GNU as
+					 * spells it ABS32; being bytes, the
+					 * section's new start goes in as it
+					 * is, not scaled the way a call's
+					 * instruction index is above.
+					 *
+					 * r_offset is checked only for an
+					 * executable section, and SHT_NOBITS
+					 * has no raw_data, so bound it here --
+					 * subtracting, so it cannot wrap.
+					 */
+					if (!dst_linked_sec->raw_data ||
+					    dst_linked_sec->sec_sz < (int)sizeof(*val) ||
+					    dst_rel->r_offset % sizeof(*val) ||
+					    dst_rel->r_offset >
+					    (size_t)dst_linked_sec->sec_sz - sizeof(*val)) {
+						pr_warn("ELF relo #%d in section #%zu points outside the data of section '%s' in %s\n",
+							j, src_sec->sec_idx,
+							dst_linked_sec->sec_name,
+							obj->filename);
+						return -EINVAL;
+					}
+					val = dst_linked_sec->raw_data + dst_rel->r_offset;
+					if (linker->swapped_endian)
+						*val = bswap_32(bswap_32(*val) + sec->dst_off);
+					else
+						*val += sec->dst_off;
 				} else {
 					pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
 					return -EINVAL;
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 18/21] selftests/bpf: Add an end-to-end .bpf_cleanup exception test
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (16 preceding siblings ...)
  2026-09-26  5:01 ` [PATCH bpf-next v6 17/21] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
@ 2026-09-26  5:01 ` Yonghong Song
  2026-09-26  5:01 ` [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
                   ` (2 subsequent siblings)
  20 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

C has no unwinding, so nothing here comes out of the frontend: the frames
that own a resource are written as __naked inline assembly, which spells
out by hand exactly what a frontend emits -- a call site bracketed by two
labels, a landing pad unreachable in the compiler's CFG, and a .bpf_cleanup
record tying them together. The assembler turns ".long <text label>" into
the same R_BPF_64_NODYLD32 relocation the BPF AsmPrinter emits, so libbpf
and the kernel see an object indistinguishable from a compiler-generated
one.

Call chain: entry -> foo1 -> foo1v -> foo2 -> foo3. foo3 holds a
non-preemptible section and unwinds inside it; foo2 holds an RCU read lock
and has two call sites sharing one pad, one of them its own unwind; foo1v
is a void frame whose pad ends in a jump to a resume block placed after an
unrelated block that ends in a plain exit; foo1 owns nothing and gets no
record; entry is the boundary.

There are also some shapes the kernel refuses: the ones with no correct
answer, and a table whose record covers no call that can unwind, whose pad
nothing reaches. The test skips rather than fails where the JIT cannot
dispatch a landing pad at all.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 .../selftests/bpf/exceptions_cleanup.h        |  27 +
 .../bpf/prog_tests/exceptions_cleanup.c       |  85 +++
 .../selftests/bpf/progs/exceptions_cleanup.c  | 162 +++++
 .../bpf/progs/exceptions_cleanup_fail.c       | 593 ++++++++++++++++++
 4 files changed, 867 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/exceptions_cleanup.h
 create mode 100644 tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup.c
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c

diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
new file mode 100644
index 000000000000..d1d40314035e
--- /dev/null
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -0,0 +1,27 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef __EXCEPTIONS_CLEANUP_H__
+#define __EXCEPTIONS_CLEANUP_H__
+
+/* progs/exceptions_cleanup.c: one bit per frame that reports it ran. */
+#define RAN_FOO3_PREEMPT	0x1
+#define RAN_FOO2_RCU		0x2
+#define RAN_FOO1V_PREEMPT	0x4
+#define RAN_FOO2_DROP		0x8
+#define RAN_BUMP		0x10
+
+#define CLEANUP_REC(begin, end, landing_pad)			\
+	".pushsection .bpf_cleanup,\"a\",@progbits;"		\
+	".long " begin ";"					\
+	".long " end ";"					\
+	".long " landing_pad ";"				\
+	".popsection;"
+
+/* Set a bit in @pads_ran. */
+#define PAD_RAN(bit)						\
+	"r1 = %[pads_ran] ll;"					\
+	"r2 = *(u64 *)(r1 + 0);"				\
+	"r2 |= " bit ";"					\
+	"*(u64 *)(r1 + 0) = r2;"
+
+#endif /* __EXCEPTIONS_CLEANUP_H__ */
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
new file mode 100644
index 000000000000..255f88d35aad
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -0,0 +1,85 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <test_progs.h>
+#include "exceptions_cleanup.h"
+#include "exceptions_cleanup.skel.h"
+#include "exceptions_cleanup_fail.skel.h"
+
+/* foo3 unwound: every frame that has a pad ran it. */
+#define PADS_FOO3_UNWOUND \
+	(RAN_FOO3_PREEMPT | RAN_FOO2_RCU | RAN_FOO1V_PREEMPT | RAN_FOO2_DROP)
+
+/* foo2 unwound after foo3 returned normally: foo3's pad must not run. */
+#define PADS_FOO2_UNWOUND \
+	(RAN_FOO2_RCU | RAN_FOO1V_PREEMPT | RAN_FOO2_DROP)
+
+static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
+		__u64 pads)
+{
+	__u64 ctx = 0;
+	int err;
+
+	LIBBPF_OPTS(bpf_test_run_opts, topts,
+		    .ctx_in = &ctx,
+		    .ctx_size_in = sizeof(ctx),
+	);
+
+	skel->bss->input = input;
+	skel->bss->pads_ran = 0;
+	skel->bss->result = 0;
+
+	err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.entry), &topts);
+	if (!ASSERT_OK(err, "run"))
+		return;
+	ASSERT_EQ(topts.retval, retval, "retval");
+	/* bump() is not a landing pad; it sets its bit on every run. */
+	ASSERT_EQ(skel->bss->pads_ran, pads | RAN_BUMP, "pads_ran");
+}
+
+void test_exceptions_cleanup(void)
+{
+	char log[8192] = {};
+
+	LIBBPF_OPTS(bpf_object_open_opts, opts,
+		    .kernel_log_buf = log,
+		    .kernel_log_size = sizeof(log));
+	struct exceptions_cleanup *skel;
+	int err;
+
+	skel = exceptions_cleanup__open_opts(&opts);
+	if (!ASSERT_OK_PTR(skel, "open"))
+		return;
+
+	err = exceptions_cleanup__load(skel);
+	if (err) {
+		if (err == -EOPNOTSUPP &&
+		    strstr(log, "exception cleanup needs a JIT that can dispatch landing pads")) {
+			printf("%s:SKIP:JIT cannot dispatch exception cleanup landing pads\n",
+			       __func__);
+			test__skip();
+		} else if (!ASSERT_OK(err, "load")) {
+			fprintf(stderr, "%s", log);
+		}
+		exceptions_cleanup__destroy(skel);
+		return;
+	}
+
+	/* No unwind: foo3 returns 1 ^ 1 == 0, foo2 adds one, no pad runs. */
+	if (test__start_subtest("no_unwind"))
+		run(skel, 1, 1, 0);
+
+	/* foo3 unwinds; every pad runs and entry returns zero. */
+	if (test__start_subtest("unwind_from_foo3"))
+		run(skel, 101, 0, PADS_FOO3_UNWOUND);
+
+	/*
+	 * foo3 returns 2 ^ 1 == 3, so foo2 unwinds from its own second region;
+	 * foo3's frame is long gone, so its pad must not run.
+	 */
+	if (test__start_subtest("unwind_from_foo2"))
+		run(skel, 2, 0, PADS_FOO2_UNWOUND);
+
+	exceptions_cleanup__destroy(skel);
+
+	RUN_TESTS(exceptions_cleanup_fail);
+}
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
new file mode 100644
index 000000000000..2b22a05a0ded
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
@@ -0,0 +1,162 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+	bpf_unwind();
+	bpf_rcu_read_lock();
+	bpf_rcu_read_unlock();
+	bpf_preempt_disable();
+	bpf_preempt_enable();
+	bpf_unwind_resume(NULL);
+}
+
+__u64 input = 0;
+__u64 pads_ran = 0;
+__u64 result = 0;
+
+static __used __noinline __u64 foo3(__u64 x)
+{
+	bpf_preempt_disable();
+	if (x > 100)
+		asm volatile (
+	"1:"	"call bpf_unwind;"		/* cleanup region */
+	"2:"
+		"goto 3f;"
+	"4:"					/* landing pad */
+		/*
+		 * r0 at pad entry is whatever the walker leaves
+		 * there, carried across the call in a callee-saved
+		 * register and handed to the resume, the way a
+		 * compiler-emitted pad passes the exception pointer to
+		 * _Unwind_Resume. The kfunc takes it and ignores it, and
+		 * the two pads below do without the shuffle.
+		 */
+		"r7 = r0;"
+		"call bpf_preempt_enable;"
+		PAD_RAN("%[ran]")
+		"r1 = r7;"
+		"call bpf_unwind_resume;"
+	"3:"
+		CLEANUP_REC("1b", "2b", "4b")
+		:
+		: [ran]"i"(RAN_FOO3_PREEMPT),
+		  __imm_addr(pads_ran)
+		: __clobber_all);
+	bpf_preempt_enable();
+	return x ^ 1;
+}
+
+__u64 never = 0;
+
+static __used __naked __noinline void drop_glue(void)
+{
+	asm volatile (
+	PAD_RAN("%[ran]")
+	"exit;"
+	:
+	: [ran]"i"(RAN_FOO2_DROP), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+static __used __naked __noinline __u64 foo2(void)
+{
+	asm volatile (
+	"r6 = r1;"
+	"call bpf_rcu_read_lock;"
+	"r1 = r6;"
+"1:"	"call foo3;"			/* cleanup region #1 */
+"2:"
+	"r6 = r0;"
+	"if r6 == 0 goto 5f;"
+"3:"	"call bpf_unwind;"		/* cleanup region #2 */
+"4:"
+	"r0 = 0;"
+	"exit;"
+"5:"
+	"call bpf_rcu_read_unlock;"
+	"r0 = r6;"
+	"r0 += 1;"
+	"exit;"
+"6:"					/* landing pad, shared by both regions */
+	"call drop_glue;"
+	"call bpf_rcu_read_unlock;"
+	PAD_RAN("%[ran_rcu]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "6b")
+	CLEANUP_REC("3b", "4b", "6b")
+	:
+	: [ran_rcu]"i"(RAN_FOO2_RCU),
+	  __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+static __used __naked __noinline void foo1v(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+"1:"	"call foo2;"			/* cleanup region */
+"2:"
+	"r6 = r0;"
+	"call bpf_preempt_enable;"
+	"r1 = %[result] ll;"
+	"*(u64 *)(r1 + 0) = r6;"
+	"goto 7f;"
+"8:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	PAD_RAN("%[ran]")
+	"goto 9f;"
+"7:"					/* the frame's own exit block */
+	"r0 = 0;"
+	"exit;"
+"9:"					/* shared resume block */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "8b")
+	:
+	: [ran]"i"(RAN_FOO1V_PREEMPT), __imm_addr(input),
+	  __imm_addr(result), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+/*
+ * A frame with no cleanup record: an unwind leaving it runs no pad. The
+ * unwind never fires -- @never is global -- and the bit marks the return path.
+ */
+static __used __naked __noinline void bump(void)
+{
+	asm volatile (
+	PAD_RAN("%[ran]")
+	"r1 = %[never] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"if r1 == 0 goto 1f;"
+	"r1 = 0;"
+	"call bpf_unwind;"
+"1:"
+	"exit;"				/* r0 deliberately left alone */
+	:
+	: [ran]"i"(RAN_BUMP), __imm_addr(never), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+__noinline __u64 foo1(void)
+{
+	bump();
+	foo1v();
+	return result;
+}
+
+SEC("syscall")
+int entry(void *ctx)
+{
+	return foo1();
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
new file mode 100644
index 000000000000..0d4598e9738b
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
@@ -0,0 +1,593 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_experimental.h"
+#include "bpf_misc.h"
+#include "../test_kmods/bpf_testmod_kfunc.h"
+#include "exceptions_cleanup.h"
+
+__u64 input = 0;
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+	bpf_throw(0);
+	bpf_unwind();
+	bpf_preempt_disable();
+	bpf_preempt_enable();
+	bpf_unwind_resume(NULL);
+}
+
+/* An unwind raised in a callee, which is how a cleanup region gets one. */
+static __used __naked __noinline __u64 inner_unwind(void)
+{
+	asm volatile (
+	"r1 = 1;"
+	"call bpf_unwind;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+static int unwinding_cb(__u32 idx, void *ctx)
+{
+	bpf_unwind();
+	return 0;
+}
+
+static __used __naked __noinline __u64 cb_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call unwinding_cb;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("may unwind and is used as a callback")
+int callback_may_unwind(void *ctx)
+{
+	bpf_loop(1, unwinding_cb, NULL, 0);
+	return cb_frame();
+}
+
+/* A pad that reaches both a resume and a plain exit. */
+static __used __naked __noinline __u64 ambiguous_pad_frame(void)
+{
+	asm volatile (
+	"r6 = r1;"
+	"call bpf_preempt_disable;"
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad: two ways out */
+	"call bpf_preempt_enable;"
+	"if r6 > 10 goto 4f;"
+	"call bpf_unwind_resume;"
+	"exit;"
+"4:"
+	"r0 = 0;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("a catch pad is not supported yet")
+int ambiguous_landing_pad(void *ctx)
+{
+	return ambiguous_pad_frame();
+}
+
+/* A second bpf_unwind() from inside a landing pad. */
+static __used __naked __noinline __u64 unwind_in_pad_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad that unwinds again */
+	"call bpf_preempt_enable;"
+	"r1 = 2;"
+	"call bpf_unwind;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("starts a second unwind while one is in flight")
+int unwind_from_landing_pad(void *ctx)
+{
+	return unwind_in_pad_frame();
+}
+
+__noinline int unused_exc_cb(u64 cookie)
+{
+	return 0;
+}
+
+static __used __naked __noinline __u64 cb_and_table_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+	"r1 = 9;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__exception_cb(unused_exc_cb)
+__failure __msg("cannot be combined with an exception callback")
+int table_with_exception_cb(void *ctx)
+{
+	return cb_and_table_frame();
+}
+
+__u64 never;
+
+/*
+ * A pad calling a global subprogram that can unwind. The subprogram is
+ * verified on its own, so the pad rule is what refuses it.
+ */
+__noinline void pad_callee_that_unwinds(void)
+{
+	if (never)
+		bpf_unwind();
+}
+
+static __used __naked __noinline __u64 pad_calls_unwinder_frame(void)
+{
+	asm volatile (
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call pad_callee_that_unwinds;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("which can unwind while an unwind is in flight")
+int pad_calls_unwinder(void *ctx)
+{
+	return pad_calls_unwinder_frame();
+}
+
+/*
+ * A pad calling a global subprogram that can throw. The preempt rule is
+ * what refuses it, not anything about pads.
+ */
+__noinline void pad_callee_that_throws(void)
+{
+	if (never)
+		bpf_throw(0);
+}
+
+static __used __naked __noinline __u64 pad_calls_thrower_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+	"r1 = 11;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call pad_callee_that_throws;"	/* ...which can throw: refused */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("cannot be used inside bpf_preempt_disable-ed region")
+int pad_calls_thrower(void *ctx)
+{
+	return pad_calls_thrower_frame();
+}
+
+static __used __naked __noinline __u64 catch_pad_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+	"r1 = 12;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* catch pad: no resume, it stops here */
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not supported yet, only cleanup pads that resume")
+int catch_landing_pad(void *ctx)
+{
+	return catch_pad_frame();
+}
+
+/* A bpf_unwind_resume() outside any landing pad. */
+static __used __naked __noinline __u64 stray_resume_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+	"r1 = 13;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not in a landing pad")
+int resume_outside_pad(void *ctx)
+{
+	/* Never taken, but reachable, which is all the verifier needs. */
+	if (never)
+		bpf_unwind_resume(NULL);
+	return stray_resume_frame();
+}
+
+/* A bpf_unwind_resume() in a subprogram a landing pad calls. */
+static __used __naked __noinline void resume_in_callee(void)
+{
+	asm volatile (
+	"call bpf_unwind_resume;"
+	"exit;"
+	::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 pad_calls_resumer_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+	"r1 = 14;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	"call resume_in_callee;"	/* ...which resumes: refused */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not in a landing pad")
+int resume_in_pad_callee(void *ctx)
+{
+	return pad_calls_resumer_frame();
+}
+
+static __used __naked __noinline __u64 nested_pad_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call inner_unwind;"		/* first cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* first pad, second region's call */
+	"call bpf_preempt_enable;"
+"4:"
+	"call bpf_unwind_resume;"
+	"exit;"
+"5:"					/* second pad */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	CLEANUP_REC("3b", "4b", "5b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is inside the call-site range of")
+int nested_landing_pad(void *ctx)
+{
+	return nested_pad_frame();
+}
+
+/* A tail call in a landing pad: the frame would never reach its resume. */
+struct {
+	__uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+	__uint(max_entries, 1);
+	__uint(key_size, sizeof(__u32));
+	__uint(value_size, sizeof(__u32));
+} tc_map SEC(".maps");
+
+static __used __naked __noinline __u64 tail_call_pad_frame(void)
+{
+	asm volatile (
+	"r6 = r1;"
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r1 = r6;"
+	"r2 = %[tc_map] ll;"
+	"r3 = 0;"
+	"call %[bpf_tail_call];"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: __imm(bpf_tail_call), __imm_addr(tc_map)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("and is in a landing pad")
+int tail_call_in_pad(void *ctx)
+{
+	return tail_call_pad_frame();
+}
+
+#if defined(__BPF_FEATURE_STACK_ARGUMENT)
+
+/*
+ * A kfunc by-value argument that runs past the argument registers, in a
+ * landing pad. The pad is not what refuses it. The C call gives the extern
+ * its BTF.
+ */
+static __used __noinline void __nofit_btf_anchor(void)
+{
+	struct prog_test_pair_arg s = {};
+
+	bpf_kfunc_call_test_pair_arg_nofit(1, 2, 3, 4, s);
+}
+
+static __used __naked __noinline __u64 kfunc_arg_pad_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	"call bpf_kfunc_call_test_pair_arg_nofit;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("stack arg1 is not initialized")
+int kfunc_stack_arg_in_pad(void *ctx)
+{
+	return kfunc_arg_pad_frame();
+}
+
+#endif /* __BPF_FEATURE_STACK_ARGUMENT */
+
+/* A landing pad entered by ordinary control flow, with no unwind in flight. */
+static __used __naked __noinline __u64 jump_into_pad_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"if r6 > 7 goto 4f;"		/* an ordinary branch into the pad */
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r7 = r0;"
+"4:"					/* ... and its second instruction */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: __imm_addr(input)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("runs both inside and outside a landing pad")
+int jump_into_pad(void *ctx)
+{
+	return jump_into_pad_frame();
+}
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)
+
+/*
+ * An indirect jump in a landing pad. A jump table entry is an offset from
+ * the program's section symbol, which has to be spelled in quotes here.
+ */
+static __used __naked __noinline void gotox_unwinder(void)
+{
+	asm volatile (
+	"r1 = 15;"
+	"call bpf_unwind;"
+	"exit;"
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("and is in a landing pad")
+__naked void gotox_in_pad(void)
+{
+	asm volatile (
+	".pushsection .jumptables,\"\",@progbits;"
+"jt0_%=:"
+	".quad l0_%= - \"?syscall\";"
+	".quad l1_%= - \"?syscall\";"
+	".size jt0_%=, 16;"
+	".global jt0_%=;"
+	".popsection;"
+
+"1:"	"call gotox_unwinder;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r1 = jt0_%= ll;"
+	"r1 += 8;"
+	"r2 = *(u64 *)(r1 + 0);"
+	/*
+	 * gotox r2, as raw bytes: the mnemonic only reached the LLVM
+	 * assembler in llvm 22, and BPF_RAW_INSN() needs <linux/bpf.h>, which
+	 * vmlinux.h rules out. dst_reg is the other nibble on a big-endian
+	 * target.
+	 */
+#if __BYTE_ORDER__ == __ORDER_BIG_ENDIAN__
+	".byte 0x0d, 0x20, 0, 0, 0, 0, 0, 0;"
+#else
+	".byte 0x0d, 0x02, 0, 0, 0, 0, 0, 0;"
+#endif
+"l0_%=:"
+	"call bpf_unwind_resume;"
+	"exit;"
+"l1_%=:"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+#endif /* x86 || arm64 */
+
+/* A BPF_LD_[ABS|IND] in a pad: a failed load leaves without resuming. */
+static __used __naked __noinline __u64 ld_abs_pad_frame(void)
+{
+	asm volatile (
+	"r6 = r1;"			/* the skb BPF_LD_ABS reads */
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r0 = *(u32 *)skb[0];"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?tc")
+__failure __msg("and is in a landing pad")
+__naked void ld_abs_in_pad(void)
+{
+	asm volatile (
+	"call ld_abs_pad_frame;"
+	"exit;"
+	::: __clobber_all);
+}
+
+/*
+ * A subprogram a landing pad calls, which unwinds on its own. The second
+ * unwind would rewrite the frames the first is walking.
+ */
+static __used __naked __noinline __u64 own_pad_callee(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* its landing pad */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 pad_calls_own_pad_frame(void)
+{
+	asm volatile (
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad, which calls the above */
+	"call own_pad_callee;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("starts a second unwind while one is in flight")
+int unwind_in_pad_callee(void *ctx)
+{
+	return pad_calls_own_pad_frame();
+}
+
+/* A record whose range holds no call that can unwind. */
+static __used __naked __noinline __u64 nounwind_rec_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call bpf_preempt_enable;"	/* cleanup region: nounwind */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad, reached by nothing */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("unreachable insn")
+int nounwind_region(void *ctx)
+{
+	return nounwind_rec_frame();
+}
+
+char _license[] SEC("license") = "GPL";
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (17 preceding siblings ...)
  2026-09-26  5:01 ` [PATCH bpf-next v6 18/21] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
@ 2026-09-26  5:01 ` Yonghong Song
  2026-09-26  5:18   ` sashiko-bot
  2026-09-27 20:24   ` bot+bpf-ci
  2026-09-26  5:01 ` [PATCH bpf-next v6 20/21] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
  2026-09-26  5:01 ` [PATCH bpf-next v6 21/21] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
  20 siblings, 2 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

A test that has to look at a program's state after it ran needs a driver of
its own: open and load a skeleton, set an input, call
bpf_prog_test_run_opts(), read a variable back, destroy the skeleton. The
loader already does all of that for __retval(), and the only thing missing
is reaching the program's globals.

__set_global(var, value) writes one before the run and __ret_global(var,
value) checks one after it, both of them implying the execution __retval()
asks for. The variable is found the way veristat finds one: the map whose
name ends in ".bss" or ".data", the datasec of that name in the object's
BTF, then the variable within it. Four and eight byte variables are
supported, which is what a counter or a bitmask needs; anything else is
refused rather than read at the wrong width.

The value may be a '|' separated list of terms, so that a bitmask reads in
the test the way it is written in the program.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/testing/selftests/bpf/progs/bpf_misc.h |   7 +
 tools/testing/selftests/bpf/test_loader.c    | 207 +++++++++++++++++++
 2 files changed, 214 insertions(+)

diff --git a/tools/testing/selftests/bpf/progs/bpf_misc.h b/tools/testing/selftests/bpf/progs/bpf_misc.h
index f3dbc3b59bff..9afa163fac5a 100644
--- a/tools/testing/selftests/bpf/progs/bpf_misc.h
+++ b/tools/testing/selftests/bpf/progs/bpf_misc.h
@@ -93,6 +93,11 @@
  * __failure         Expect program load failure in privileged mode.
  * __failure_unpriv  Expect program load failure in unprivileged mode.
  *
+ * __set_global      Set a global variable of the program to a value before
+ *                   executing it.
+ * __ret_global      Execute the program and check that a global variable
+ *                   holds the given value afterwards. The variable has to
+ *                   live in .bss or .data and be four or eight bytes wide.
  * __retval          Execute the program using BPF_PROG_TEST_RUN command,
  *                   expect return value to match passed parameter:
  *                   - a decimal number
@@ -160,6 +165,8 @@
 #define __log_level(lvl)	__test_tag("test_log_level=" #lvl)
 #define __flag(flag)		__test_tag("test_prog_flags=" #flag)
 #define __retval(val)		__test_tag("test_retval=" XSTR(val))
+#define __set_global(var, val)	__test_tag("test_global_set=" #var ":" XSTR(val))
+#define __ret_global(var, val)	__test_tag("test_global_ret=" #var ":" XSTR(val))
 #define __retval_unpriv(val)	__test_tag("test_retval_unpriv=" XSTR(val))
 #define __auxiliary		__test_tag("test_auxiliary")
 #define __auxiliary_unpriv	__test_tag("test_auxiliary_unpriv")
diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
index 25eeb1c1248b..65b55f7fa907 100644
--- a/tools/testing/selftests/bpf/test_loader.c
+++ b/tools/testing/selftests/bpf/test_loader.c
@@ -62,6 +62,12 @@ struct test_subspec {
 	int retval;
 	bool execute;
 	__u64 caps;
+	char *set_global_var;
+	__u64 set_global_val;
+	bool has_set_global;
+	char *ret_global_var;
+	__u64 ret_global_val;
+	bool has_ret_global;
 };
 
 struct test_spec {
@@ -129,6 +135,15 @@ static void free_test_spec(struct test_spec *spec)
 	free_msgs(&spec->unpriv.stdout);
 	free_msgs(&spec->priv.stdout);
 
+	free(spec->priv.set_global_var);
+	free(spec->priv.ret_global_var);
+	free(spec->unpriv.set_global_var);
+	free(spec->unpriv.ret_global_var);
+	spec->priv.set_global_var = NULL;
+	spec->priv.ret_global_var = NULL;
+	spec->unpriv.set_global_var = NULL;
+	spec->unpriv.ret_global_var = NULL;
+
 	free(spec->priv.name);
 	free(spec->priv.description);
 	free(spec->unpriv.name);
@@ -311,6 +326,142 @@ static int parse_caps(const char *str, __u64 *val, const char *name)
 	return 0;
 }
 
+static int parse_global_var(const char *str, char **var, __u64 *val, const char *name)
+{
+	const char *colon = strrchr(str, ':');
+	char *end;
+
+	if (!colon || colon == str) {
+		PRINT_FAIL("expecting '<variable>:<value>' for %s, got '%s'\n", name, str);
+		return -EINVAL;
+	}
+
+	*val = 0;
+	for (const char *term = colon + 1;;) {
+		__u64 v;
+
+		errno = 0;
+		v = strtoull(term, &end, 0);
+		if (errno || end == term) {
+			PRINT_FAIL("failed to parse %s value '%s'\n", name, colon + 1);
+			return -EINVAL;
+		}
+		*val |= v;
+		while (*end == ' ')
+			end++;
+		if (!*end)
+			break;
+		if (*end != '|') {
+			PRINT_FAIL("failed to parse %s value '%s'\n", name, colon + 1);
+			return -EINVAL;
+		}
+		term = end + 1;
+	}
+
+	free(*var);
+	*var = strndup(str, colon - str);
+	if (!*var) {
+		PRINT_FAIL("failed to allocate %s variable name\n", name);
+		return -ENOMEM;
+	}
+
+	return 0;
+}
+
+static int find_global_var(struct bpf_object *obj, const char *name,
+			   struct bpf_map **map, __u32 *off, __u32 *sz)
+{
+	static const char * const secs[] = { ".bss", ".data" };
+	struct btf *btf = bpf_object__btf(obj);
+	int i, s;
+
+	if (!btf) {
+		PRINT_FAIL("no BTF for object\n");
+		return -ENOENT;
+	}
+
+	for (s = 0; s < ARRAY_SIZE(secs); s++) {
+		const struct btf_var_secinfo *vsi;
+		struct bpf_map *m = NULL, *iter;
+		const struct btf_type *sec;
+		size_t slen = strlen(secs[s]);
+		int id;
+
+		bpf_object__for_each_map(iter, obj) {
+			const char *mname = bpf_map__name(iter);
+			size_t len = mname ? strlen(mname) : 0;
+
+			if (len >= slen && strcmp(mname + len - slen, secs[s]) == 0) {
+				m = iter;
+				break;
+			}
+		}
+		id = btf__find_by_name_kind(btf, secs[s], BTF_KIND_DATASEC);
+		if (!m || id < 0)
+			continue;
+
+		sec = btf__type_by_id(btf, id);
+		vsi = btf_var_secinfos(sec);
+		for (i = 0; i < btf_vlen(sec); i++, vsi++) {
+			const struct btf_type *var = btf__type_by_id(btf, vsi->type);
+
+			if (strcmp(btf__name_by_offset(btf, var->name_off), name))
+				continue;
+			if (vsi->size != 4 && vsi->size != 8) {
+				PRINT_FAIL("'%s' is %u bytes, only 4 and 8 are supported\n",
+					   name, vsi->size);
+				return -EINVAL;
+			}
+			*map = m;
+			*off = vsi->offset;
+			*sz = vsi->size;
+			return 0;
+		}
+	}
+
+	PRINT_FAIL("no global variable '%s'\n", name);
+	return -ENOENT;
+}
+
+static int access_global_var(struct bpf_object *obj, const char *name,
+			     __u64 *val, bool set)
+{
+	__u32 off, sz, zero = 0;
+	struct bpf_map *map;
+	size_t vsz;
+	void *buf;
+	int err;
+
+	err = find_global_var(obj, name, &map, &off, &sz);
+	if (err)
+		return err;
+
+	vsz = bpf_map__value_size(map);
+	buf = calloc(1, vsz);
+	if (!buf)
+		return -ENOMEM;
+
+	err = bpf_map__lookup_elem(map, &zero, sizeof(zero), buf, vsz, 0);
+	if (err) {
+		PRINT_FAIL("failed to read '%s': %d\n", name, err);
+		goto out;
+	}
+	if (!set) {
+		*val = sz == 4 ? *(__u32 *)(buf + off) : *(__u64 *)(buf + off);
+		goto out;
+	}
+	if (sz == 4)
+		*(__u32 *)(buf + off) = *val;
+	else
+		*(__u64 *)(buf + off) = *val;
+	err = bpf_map__update_elem(map, &zero, sizeof(zero), buf, vsz, 0);
+	if (err)
+		PRINT_FAIL("failed to write '%s': %d\n", name, err);
+out:
+	free(buf);
+	return err;
+}
+
 static int parse_retval(const char *str, int *val, const char *name)
 {
 	/*
@@ -557,6 +708,24 @@ static int parse_test_spec(struct test_loader *tester,
 			spec->mode_mask |= UNPRIV;
 			spec->unpriv.execute = true;
 			has_unpriv_retval = true;
+		} else if ((val = str_has_pfx(s, "test_global_set="))) {
+			err = parse_global_var(val, &spec->priv.set_global_var,
+					       &spec->priv.set_global_val,
+					       "__set_global");
+			if (err)
+				goto cleanup;
+			spec->priv.has_set_global = true;
+			spec->priv.execute = true;
+			spec->mode_mask |= PRIV;
+		} else if ((val = str_has_pfx(s, "test_global_ret="))) {
+			err = parse_global_var(val, &spec->priv.ret_global_var,
+					       &spec->priv.ret_global_val,
+					       "__ret_global");
+			if (err)
+				goto cleanup;
+			spec->priv.has_ret_global = true;
+			spec->priv.execute = true;
+			spec->mode_mask |= PRIV;
 		} else if ((val = str_has_pfx(s, "test_log_level="))) {
 			err = parse_int(val, &spec->log_level, "test log level");
 			if (err)
@@ -742,6 +911,25 @@ static int parse_test_spec(struct test_loader *tester,
 			spec->unpriv.execute = spec->priv.execute;
 		}
 
+		if (spec->priv.has_set_global && !spec->unpriv.has_set_global) {
+			spec->unpriv.set_global_var = strdup(spec->priv.set_global_var);
+			if (!spec->unpriv.set_global_var) {
+				err = -ENOMEM;
+				goto cleanup;
+			}
+			spec->unpriv.set_global_val = spec->priv.set_global_val;
+			spec->unpriv.has_set_global = true;
+		}
+		if (spec->priv.has_ret_global && !spec->unpriv.has_ret_global) {
+			spec->unpriv.ret_global_var = strdup(spec->priv.ret_global_var);
+			if (!spec->unpriv.ret_global_var) {
+				err = -ENOMEM;
+				goto cleanup;
+			}
+			spec->unpriv.ret_global_val = spec->priv.ret_global_val;
+			spec->unpriv.has_ret_global = true;
+		}
+
 		if (spec->unpriv.expect_msgs.cnt == 0)
 			clone_msgs(&spec->priv.expect_msgs, &spec->unpriv.expect_msgs);
 		if (spec->unpriv.expect_xlated.cnt == 0)
@@ -1532,6 +1720,13 @@ void run_subtest(struct test_loader *tester,
 			}
 		}
 
+		if (subspec->has_set_global) {
+			__u64 v = subspec->set_global_val;
+
+			if (access_global_var(tobj, subspec->set_global_var, &v, true))
+				goto tobj_cleanup;
+		}
+
 		err = do_prog_test_run(bpf_program__fd(tprog), &retval,
 				       bpf_program__type(tprog) == BPF_PROG_TYPE_SYSCALL ? true : false,
 				       spec->linear_sz);
@@ -1540,6 +1735,18 @@ void run_subtest(struct test_loader *tester,
 			goto tobj_cleanup;
 		}
 
+		if (subspec->has_ret_global) {
+			__u64 v = 0;
+
+			if (access_global_var(tobj, subspec->ret_global_var, &v, false))
+				goto tobj_cleanup;
+			if (v != subspec->ret_global_val) {
+				PRINT_FAIL("Unexpected %s: 0x%llx != 0x%llx\n",
+					   subspec->ret_global_var, v, subspec->ret_global_val);
+				goto tobj_cleanup;
+			}
+		}
+
 		verify_stderr(bpf_program__fd(tprog), &subspec->stderr);
 
 		if (subspec->stdout.cnt) {
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 20/21] selftests/bpf: Cover the exception cleanup shapes the chain does not reach
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (18 preceding siblings ...)
  2026-09-26  5:01 ` [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
@ 2026-09-26  5:01 ` Yonghong Song
  2026-09-27 20:40   ` bot+bpf-ci
  2026-09-26  5:01 ` [PATCH bpf-next v6 21/21] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
  20 siblings, 1 reply; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

The end-to-end test walks one call chain with a pad in most of its frames.
This adds the shapes that chain does not reach, in the order the new file
has them:

 - what a cleanup table leaves dead
 - a callee called from both a covered and an uncovered site
 - a pad that reads its frame's callee-saved registers
 - a pad in the main program's own frame
 - a covered bpf_unwind() the sweep leaves last
 - a region ending on a 16-byte instruction
 - a pad that reloads from and writes to its own frame
 - a pad terminated by _Unwind_Resume rather than bpf_unwind_resume
 - a pad whose first instruction is a nop
 - a pad that indexes its frame by a register the frame set before the
   unwinding call
 - the same over an unwinding global subprogram
 - a pad that indexes its frame by r0, which no instruction wrote
 - a region covering more than one call that can unwind
 - two pads with an uncovered frame between them

Each of them wants the same three things of a run: an input, a return
value, and the set of landing pads that ran. That is what __set_global(),
__retval() and __ret_global() say, so they say it and RUN_TESTS() does the
rest, which also gives each shape a name of its own in the test output.
The shared-callee shape takes two more programs over the same frame: one
with nothing unwinding, one unwinding from the uncovered site.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 .../selftests/bpf/exceptions_cleanup.h        |  17 +
 .../bpf/prog_tests/exceptions_cleanup.c       |   2 +
 .../bpf/progs/exceptions_cleanup_shapes.c     | 665 ++++++++++++++++++
 3 files changed, 684 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c

diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
index d1d40314035e..49b6806cde82 100644
--- a/tools/testing/selftests/bpf/exceptions_cleanup.h
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -10,6 +10,23 @@
 #define RAN_FOO2_DROP		0x8
 #define RAN_BUMP		0x10
 
+/* progs/exceptions_cleanup_shapes.c: one bit per shape. */
+#define RAN_SWEEP		0x1
+#define RAN_SHARED		0x2
+#define RAN_REGS		0x4
+#define RAN_MAIN_PAD		0x8
+#define RAN_PAD_FIRST		0x10
+#define RAN_WIDE_REC		0x20
+#define RAN_PAD_STACK		0x40
+#define RAN_RESUME_ALIAS	0x80
+#define RAN_NOP_PAD		0x100
+#define RAN_VAR_STACK		0x200
+#define RAN_GLOBAL_PAD		0x400
+#define RAN_PAD_R0		0x800
+#define RAN_MULTI_CALL		0x1000
+#define RAN_GAP_INNER		0x2000
+#define RAN_GAP_OUTER		0x4000
+
 #define CLEANUP_REC(begin, end, landing_pad)			\
 	".pushsection .bpf_cleanup,\"a\",@progbits;"		\
 	".long " begin ";"					\
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
index 255f88d35aad..c06ec10359b9 100644
--- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -4,6 +4,7 @@
 #include "exceptions_cleanup.h"
 #include "exceptions_cleanup.skel.h"
 #include "exceptions_cleanup_fail.skel.h"
+#include "exceptions_cleanup_shapes.skel.h"
 
 /* foo3 unwound: every frame that has a pad ran it. */
 #define PADS_FOO3_UNWOUND \
@@ -82,4 +83,5 @@ void test_exceptions_cleanup(void)
 	exceptions_cleanup__destroy(skel);
 
 	RUN_TESTS(exceptions_cleanup_fail);
+	RUN_TESTS(exceptions_cleanup_shapes);
 }
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
new file mode 100644
index 000000000000..37a78e035a50
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
@@ -0,0 +1,665 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+	bpf_unwind();
+	bpf_rcu_read_lock();
+	bpf_rcu_read_unlock();
+	bpf_preempt_disable();
+	bpf_preempt_enable();
+	bpf_unwind_resume(NULL);
+}
+
+__u64 input = 0;
+__u64 outer_input = 0;
+__u64 magic = 0x5eed;
+__u64 pads_ran = 0;
+
+/* The callee most of the shapes below unwind out of. */
+static __used __noinline __u64 pc_unwinder(__u64 x)
+{
+	if (x > 100)
+		bpf_unwind();
+	return x + 1;
+}
+
+/* What a cleanup table leaves dead: an unwind's continuation and its tail. */
+static __used __naked __noinline __u64 sweep_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"call bpf_preempt_disable;"
+	"if r6 < 101 goto 6f;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"goto 3f;"
+"4:"					/* landing pad */
+	"r7 = r0;"
+	"call bpf_preempt_enable;"
+	PAD_RAN("%[ran]")
+	"r1 = r7;"
+	"call bpf_unwind_resume;"
+	"r1 = %[pads_ran] ll;"
+	"r2 = *(u64 *)(r1 + 0);"
+	"if r2 == 0 goto 5f;"
+	"call bpf_preempt_enable;"
+	"r0 = 7;"
+	"exit;"
+"5:"
+	"r0 = 8;"
+	"exit;"
+"3:"					/* dead: only the dead goto reaches it */
+	"r0 = 9;"
+	"exit;"
+"6:"					/* live: the ordinary return */
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "4b")
+	:
+	: [ran]"i"(RAN_SWEEP),
+	  __imm_addr(input), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_SWEEP)
+int entry_sweep(void *ctx)
+{
+	return sweep_frame();
+}
+
+/* A callee on both a covered and an uncovered call: the pad is the site's. */
+static __used __noinline __u64 shared_callee(__u64 x)
+{
+	if (x > 100)
+		bpf_unwind();
+	return x + 1;
+}
+
+static __used __naked __noinline __u64 shared_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"r1 = %[outer_input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"call shared_callee;"		/* uncovered: no pad for its unwind */
+	"call bpf_rcu_read_lock;"
+	"r1 = r6;"
+"1:"	"call shared_callee;"		/* cleanup region */
+"2:"
+	"r6 = r0;"
+	"call bpf_rcu_read_unlock;"
+	"r0 = r6;"
+	"exit;"
+"3:"					/* landing pad */
+	"r7 = r0;"
+	"call bpf_rcu_read_unlock;"
+	PAD_RAN("%[ran]")
+	"r1 = r7;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_SHARED), __imm_addr(input), __imm_addr(outer_input),
+	  __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_SHARED)
+int entry_shared(void *ctx)
+{
+	return shared_frame();
+}
+
+/* The same shape, with no unwind. */
+SEC("?syscall")
+__success __set_global(input, 1) __retval(2)
+__ret_global(pads_ran, 0)
+int entry_shared_quiet(void *ctx)
+{
+	return shared_frame();
+}
+
+/* And the uncovered site unwinding instead: no pad runs. */
+SEC("?syscall")
+__success __set_global(input, 1) __retval(0)
+__ret_global(pads_ran, 0)
+int entry_shared_uncovered(void *ctx)
+{
+	outer_input = 101;
+	return shared_frame();
+}
+
+/* A pad that reads r6-r9, which the callee overwrote before it unwound. */
+#define LOAD_MAGIC_REGS						\
+	"r1 = %[magic] ll;"					\
+	"r6 = *(u64 *)(r1 + 0);"				\
+	"r7 = r6;"						\
+	"r7 += 1;"						\
+	"r8 = r6;"						\
+	"r8 += 2;"						\
+	"r9 = r6;"						\
+	"r9 += 3;"
+
+/* Set @bit only if r6-r9 still hold what LOAD_MAGIC_REGS put there. */
+#define CHECK_MAGIC_REGS(bit)					\
+	"r1 = %[magic] ll;"					\
+	"r2 = *(u64 *)(r1 + 0);"				\
+	"if r6 != r2 goto 9f;"					\
+	"r2 += 1;"						\
+	"if r7 != r2 goto 9f;"					\
+	"r2 += 1;"						\
+	"if r8 != r2 goto 9f;"					\
+	"r2 += 1;"						\
+	"if r9 != r2 goto 9f;"					\
+	PAD_RAN(bit)						\
+	"9:"
+
+static __used __naked __noinline __u64 regs_unwinder(void)
+{
+	asm volatile (
+	/* Not this frame's to keep, and that is the point. */
+	"r6 = 0xdead;"
+	"r7 = 0xbeef;"
+	"r8 = 0xcafe;"
+	"r9 = 0xf00d;"
+	"call bpf_unwind;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 regs_frame(void)
+{
+	asm volatile (
+	LOAD_MAGIC_REGS
+	"call bpf_preempt_disable;"
+"1:"	"call regs_unwinder;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	CHECK_MAGIC_REGS("%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_REGS),
+	  __imm_addr(magic), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_REGS)
+int entry_regs(void *ctx)
+{
+	return regs_frame();
+}
+
+/* A pad in the main program's frame, which jit_subprogs() makes func[0]. */
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_MAIN_PAD)
+__naked int entry_main_pad(void)
+{
+	asm volatile (
+	LOAD_MAGIC_REGS
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"if r1 < 101 goto 8f;"
+"1:"	"call regs_unwinder;"		/* cleanup region */
+"2:"
+"8:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	CHECK_MAGIC_REGS("%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_MAIN_PAD), __imm_addr(input),
+	  __imm_addr(magic), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+/*
+ * A covered bpf_unwind() the sweep leaves last, behind the exception callback
+ * patchlet, which has to carry the marks with it.
+ */
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_PAD_FIRST)
+__naked int entry_pad_first(void)
+{
+	asm volatile (
+	LOAD_MAGIC_REGS
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"if r1 < 101 goto 7f;"
+	"goto 4f;"
+"3:"					/* landing pad, ahead of the call */
+	CHECK_MAGIC_REGS("%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+"7:"
+	"r0 = 0;"
+	"exit;"
+"4:"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"exit;"				/* dead: swept, leaving the call last */
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_PAD_FIRST),
+	  __imm_addr(input), __imm_addr(magic), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+/* A region ending on a 16-byte insn, so end - 1 names its second half. */
+static __used __naked __noinline __u64 wide_rec_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"call bpf_rcu_read_lock;"
+	"r1 = r6;"
+"1:"	"call shared_callee;"		/* cleanup region begins */
+	"r1 = %[magic] ll;"		/* ... and ends on this pair */
+"2:"
+	"r6 = r0;"
+	"call bpf_rcu_read_unlock;"
+	"r0 = r6;"
+	"exit;"
+"3:"					/* landing pad */
+	"r7 = r0;"
+	"call bpf_rcu_read_unlock;"
+	PAD_RAN("%[ran]")
+	"r1 = r7;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_WIDE_REC), __imm_addr(input), __imm_addr(magic),
+	  __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_WIDE_REC)
+int entry_wide_rec(void *ctx)
+{
+	return wide_rec_frame();
+}
+
+/*
+ * A pad that reloads from and stores to its own frame, which arm64 addresses
+ * through the stack pointer.
+ */
+static __used __naked __noinline __u64 pad_stack_frame(void)
+{
+	asm volatile (
+	"r1 = %[magic] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"*(u64 *)(r10 - 8) = r1;"	/* what the pad will want */
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+"1:"	"call pc_unwinder;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r6 = r0;"
+	"r7 = *(u64 *)(r10 - 8);"	/* reload it out of the frame */
+	"*(u64 *)(r10 - 16) = r7;"	/* and write the frame while here */
+	"r1 = %[magic] ll;"
+	"r2 = *(u64 *)(r1 + 0);"
+	"if r7 != r2 goto 9f;"
+	"r3 = *(u64 *)(r10 - 16);"
+	"if r3 != r2 goto 9f;"
+	PAD_RAN("%[ran]")
+"9:"
+	"r1 = r6;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_PAD_STACK), __imm_addr(input), __imm_addr(magic),
+	  __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_PAD_STACK)
+int entry_pad_stack(void *ctx)
+{
+	return pad_stack_frame();
+}
+
+/* The name LLVM gives the resume: _Unwind_Resume(), which libbpf maps over. */
+extern void _Unwind_Resume(void *ptr) __ksym;
+
+static __used __noinline void __resume_alias_btf_anchor(void)
+{
+	_Unwind_Resume(NULL);
+}
+
+static __used __naked __noinline __u64 resume_alias_frame(void)
+{
+	asm volatile (
+"1:"	"call regs_unwinder;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	PAD_RAN("%[ran]")
+	"call _Unwind_Resume;"		/* the frontend's name for it */
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_RESUME_ALIAS), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_RESUME_ALIAS)
+int entry_resume_alias(void *ctx)
+{
+	return resume_alias_frame();
+}
+
+/* A pad starting on a nop, which opt_remove_nops() drops after the walk. */
+static __used __naked __noinline __u64 nop_pad_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"if r6 < 101 goto 6f;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+"6:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad: a nop, then its body */
+	"goto +0;"
+	PAD_RAN("%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_NOP_PAD),
+	  __imm_addr(input), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_NOP_PAD)
+int entry_nop_pad(void *ctx)
+{
+	return nop_pad_frame();
+}
+
+/* A pad that indexes its frame by a register set before the unwinding call. */
+
+/* An unwinder that touches none of r6-r9, so the walk leaves this frame. */
+static __used __naked __noinline __u64 var_unwinder(void)
+{
+	asm volatile (
+	"if r1 < 101 goto 1f;"
+	"call bpf_unwind;"
+"1:"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 var_stack_frame(void)
+{
+	asm volatile (
+	"r1 = %[magic] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"*(u64 *)(r10 - 8) = r1;"	/* the slot the pad will read... */
+	"*(u64 *)(r10 - 16) = r1;"	/* ...whichever of the two it is */
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"r6 &= 1;"			/* an unknown slot number... */
+	"r6 <<= 3;"			/* ...as an aligned byte offset */
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+"1:"	"call var_unwinder;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r7 = r0;"
+	"r1 = r10;"
+	"r1 += r6;"			/* variable offset into the frame */
+	"r2 = *(u64 *)(r1 - 16);"
+	"r3 = %[magic] ll;"
+	"r3 = *(u64 *)(r3 + 0);"
+	"if r2 != r3 goto 9f;"
+	PAD_RAN("%[ran]")
+"9:"
+	"r1 = r7;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_VAR_STACK), __imm_addr(input), __imm_addr(magic),
+	  __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_VAR_STACK)
+int entry_var_stack(void *ctx)
+{
+	return var_stack_frame();
+}
+
+/* The same over a global subprogram, which the verifier enters no frame for. */
+__noinline __u64 global_unwinder(__u64 x)
+{
+	if (x > 100)
+		bpf_unwind();
+	return x + 1;
+}
+
+static __used __naked __noinline __u64 global_pad_frame(void)
+{
+	asm volatile (
+	"r1 = %[magic] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"*(u64 *)(r10 - 8) = r1;"
+	"*(u64 *)(r10 - 16) = r1;"
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"r6 &= 1;"
+	"r6 <<= 3;"
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+"1:"	"call global_unwinder;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r7 = r0;"
+	"r1 = r10;"
+	"r1 += r6;"
+	"r2 = *(u64 *)(r1 - 16);"
+	"r3 = %[magic] ll;"
+	"r3 = *(u64 *)(r3 + 0);"
+	"if r2 != r3 goto 9f;"
+	PAD_RAN("%[ran]")
+"9:"
+	"r1 = r7;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_GLOBAL_PAD), __imm_addr(input), __imm_addr(magic),
+	  __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_GLOBAL_PAD)
+int entry_global_pad(void *ctx)
+{
+	return global_pad_frame();
+}
+
+/* A pad that indexes its frame by r0, which no instruction in it wrote. */
+static __used __naked __noinline __u64 pad_r0_frame(void)
+{
+	asm volatile (
+	"r1 = %[magic] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"*(u64 *)(r10 - 8) = r1;"	/* the slot the pad will read... */
+	"*(u64 *)(r10 - 16) = r1;"	/* ...whichever of the two it is */
+"1:"	"call regs_unwinder;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r0 &= 1;"			/* an unknown slot number... */
+	"r0 <<= 3;"			/* ...as an aligned byte offset */
+	"r1 = r10;"
+	"r1 += r0;"			/* variable offset into the frame */
+	"r2 = *(u64 *)(r1 - 16);"
+	"r3 = %[magic] ll;"
+	"r3 = *(u64 *)(r3 + 0);"
+	"if r2 != r3 goto 9f;"
+	PAD_RAN("%[ran]")
+"9:"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_PAD_R0), __imm_addr(magic), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_PAD_R0)
+int entry_pad_r0(void *ctx)
+{
+	return pad_r0_frame();
+}
+
+/* A region over two calls that can unwind, where the second one does. */
+static __used __naked __noinline __u64 multi_call_frame(void)
+{
+	asm volatile (
+	"r1 = 1;"
+"1:"	"call pc_unwinder;"		/* covered, and returns */
+	"r6 = r0;"
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"call pc_unwinder;"		/* covered by the same record, unwinds */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad, reached from either call */
+	"r7 = r0;"
+	PAD_RAN("%[ran]")
+	"r1 = r7;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_MULTI_CALL), __imm_addr(input), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_MULTI_CALL)
+int entry_multi_call(void *ctx)
+{
+	return multi_call_frame();
+}
+
+/* Two pads with an uncovered frame between them. */
+static __used __naked __noinline __u64 gap_inner_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+"1:"	"call pc_unwinder;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r6 = r0;"
+	PAD_RAN("%[ran]")
+	"r1 = r6;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_GAP_INNER), __imm_addr(input), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+/* The frame in between, with no record of its own. */
+static __used __noinline __u64 gap_mid(void)
+{
+	return gap_inner_frame() + 1;
+}
+
+static __used __naked __noinline __u64 gap_outer_frame(void)
+{
+	asm volatile (
+"1:"	"call gap_mid;"			/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r6 = r0;"
+	PAD_RAN("%[ran]")
+	"r1 = r6;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_GAP_OUTER), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+/* And one more uncovered frame between the outer pad and the boundary. */
+static __used __noinline __u64 gap_top(void)
+{
+	return gap_outer_frame() + 1;
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_GAP_INNER | RAN_GAP_OUTER)
+int entry_two_pads(void *ctx)
+{
+	return gap_top();
+}
+
+char _license[] SEC("license") = "GPL";
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* [PATCH bpf-next v6 21/21] selftests/bpf: Load an exception cleanup program from a light skeleton
  2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (19 preceding siblings ...)
  2026-09-26  5:01 ` [PATCH bpf-next v6 20/21] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
@ 2026-09-26  5:01 ` Yonghong Song
  20 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

For a light skeleton, libbpf hands the records to bpf_gen__prog_load(),
which writes them into a blob and emits a loader program that issues
BPF_PROG_LOAD from inside the kernel. Nothing about that is shared with the
ordinary path: the attr is built field by field, and the kfunc names are
resolved by the loader program when it runs.

exceptions_cleanup_light.c is the smallest program that can tell whether a
table survives that trip: one record covering the bpf_unwind() call itself,
one pad that sets a bit and resumes. If the table arrives, the pad runs and
sets its bit; if it does not, the program does not load at all, because the
pad is then reachable from nothing.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/testing/selftests/bpf/Makefile.skel     |  2 +-
 .../selftests/bpf/exceptions_cleanup.h        |  3 ++
 .../bpf/prog_tests/exceptions_cleanup.c       | 28 +++++++++++++
 .../bpf/progs/exceptions_cleanup_light.c      | 39 +++++++++++++++++++
 4 files changed, 71 insertions(+), 1 deletion(-)
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c

diff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel
index 3d92cdca62ed..6197ca023c2f 100644
--- a/tools/testing/selftests/bpf/Makefile.skel
+++ b/tools/testing/selftests/bpf/Makefile.skel
@@ -33,7 +33,7 @@ LINKED_SKELS := test_static_linked.skel.h linked_funcs.skel.h		\
 LSKELS := fexit_sleep.c trace_printk.c trace_vprintk.c map_ptr_kern.c 	\
 	core_kern.c core_kern_overflow.c test_ringbuf.c			\
 	test_ringbuf_n.c test_ringbuf_map_key.c test_ringbuf_write.c    \
-	test_ringbuf_overwrite.c callx_rodata.c
+	test_ringbuf_overwrite.c callx_rodata.c exceptions_cleanup_light.c
 
 LSKELS_SIGNED := fentry_test.c fexit_test.c atomics.c
 
diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
index 49b6806cde82..243ab42c7ae7 100644
--- a/tools/testing/selftests/bpf/exceptions_cleanup.h
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -27,6 +27,9 @@
 #define RAN_GAP_INNER		0x2000
 #define RAN_GAP_OUTER		0x4000
 
+/* progs/exceptions_cleanup_light.c: the one pad it has. */
+#define RAN_LIGHT		0x1
+
 #define CLEANUP_REC(begin, end, landing_pad)			\
 	".pushsection .bpf_cleanup,\"a\",@progbits;"		\
 	".long " begin ";"					\
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
index c06ec10359b9..d251148279bf 100644
--- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -5,6 +5,7 @@
 #include "exceptions_cleanup.skel.h"
 #include "exceptions_cleanup_fail.skel.h"
 #include "exceptions_cleanup_shapes.skel.h"
+#include "exceptions_cleanup_light.lskel.h"
 
 /* foo3 unwound: every frame that has a pad ran it. */
 #define PADS_FOO3_UNWOUND \
@@ -37,6 +38,30 @@ static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
 	ASSERT_EQ(skel->bss->pads_ran, pads | RAN_BUMP, "pads_ran");
 }
 
+static void test_light_skeleton(void)
+{
+	struct exceptions_cleanup_light_lskel *skel;
+	__u64 ctx = 0;
+	int err;
+
+	LIBBPF_OPTS(bpf_test_run_opts, topts,
+		    .ctx_in = &ctx,
+		    .ctx_size_in = sizeof(ctx),
+	);
+
+	skel = exceptions_cleanup_light_lskel__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "light open_and_load"))
+		return;
+
+	err = bpf_prog_test_run_opts(skel->progs.entry_light.prog_fd, &topts);
+	if (!ASSERT_OK(err, "run"))
+		goto out;
+	ASSERT_EQ(topts.retval, 0, "retval");
+	ASSERT_EQ(skel->bss->pads_ran, RAN_LIGHT, "pads_ran");
+out:
+	exceptions_cleanup_light_lskel__destroy(skel);
+}
+
 void test_exceptions_cleanup(void)
 {
 	char log[8192] = {};
@@ -82,6 +107,9 @@ void test_exceptions_cleanup(void)
 
 	exceptions_cleanup__destroy(skel);
 
+	if (test__start_subtest("light_skeleton"))
+		test_light_skeleton();
+
 	RUN_TESTS(exceptions_cleanup_fail);
 	RUN_TESTS(exceptions_cleanup_shapes);
 }
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
new file mode 100644
index 000000000000..19d53569bbda
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
@@ -0,0 +1,39 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+	bpf_throw(0);
+	bpf_unwind();
+	bpf_preempt_disable();
+	bpf_preempt_enable();
+	bpf_unwind_resume(NULL);
+}
+
+__u64 pads_ran = 0;
+
+SEC("syscall")
+__naked int entry_light(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	PAD_RAN("%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_LIGHT),
+	  __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 12/21] bpf, arm64: Dispatch exception cleanup pads at run time
  2026-09-26  5:01 ` [PATCH bpf-next v6 12/21] bpf, arm64: " Yonghong Song
@ 2026-09-26  5:14   ` sashiko-bot
  2026-09-27 20:40   ` bot+bpf-ci
  1 sibling, 0 replies; 56+ messages in thread
From: sashiko-bot @ 2026-09-26  5:14 UTC (permalink / raw)
  To: Yonghong Song; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf, arm64: flawed PAC detection heuristic in bpf_unwind_store_ra()
--

commit 8fd536108380da02aed3e75e4a2d05875d208201
Author: Yonghong Song <yonghong.song@linux.dev>

bpf, arm64: Dispatch exception cleanup pads at run time

This commit implements the runtime dispatch of exception cleanup pads
for the ARM64 BPF JIT. It integrates with the stack unwinder to replace
and correctly sign return addresses using pointer authentication (PAC)
where necessary.

> diff --git a/arch/arm64/kernel/stacktrace.c b/arch/arm64/kernel/stacktrace.c
> index 3ebcf8c53fb04..c750c520f24d9 100644
> --- a/arch/arm64/kernel/stacktrace.c
> +++ b/arch/arm64/kernel/stacktrace.c

[ ... ]

> +static bool bpf_unwind_store_ra(unsigned long record, unsigned long caller_fp,
> +				u64 pc, u64 ra)
> +{
> +	struct frame_record *rec = (struct frame_record *)record;
> +	u64 stored = READ_ONCE(rec->lr);
> +
> +	if (system_supports_address_auth() && stored != pc) {

[Severity: High]
Does this heuristic fail when the generated Pointer Authentication Code (PAC)
matches the sign-extension bits?

The hardware PAC algorithm uniformly generates pseudo-random hashes.
There is a non-zero probability (e.g., ~1/128 for a 7-bit PAC on a 48-bit VA)
that the generated PAC perfectly matches the sign-extension bits (all 1s).

When this collision occurs, the PAC-stripped pc equals the unstripped
stored value despite the pointer actually being signed. The unwinder
would then incorrectly assume the frame was unsigned and write the
replacement return address (ra) to the stack unsigned.

When the BPF epilogue subsequently executes AUTIASP, wouldn't the
authentication fail, corrupting the pointer and triggering a fatal trap
on RET?

> +		u64 modifier;
> +
> +		if (WARN_ON_ONCE(!bpf_unwind_ra_modifier(record, caller_fp,
> +							 stored, pc, &modifier)))
> +			return false;
> +		ra = bpf_unwind_sign_ra(ra, modifier);
> +	}
> +	WRITE_ONCE(rec->lr, ra);
> +	return true;
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926050006.2213110-1-yonghong.song@linux.dev?part=12

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time
  2026-09-26  5:01 ` [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
@ 2026-09-26  5:15   ` sashiko-bot
  2026-09-27  4:35     ` Yonghong Song
  2026-09-27 20:39   ` bot+bpf-ci
  1 sibling, 1 reply; 56+ messages in thread
From: sashiko-bot @ 2026-09-26  5:15 UTC (permalink / raw)
  To: Yonghong Song; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf, x86: missing JIT epilogue for exit-less subprograms crashes exception unwinder
--

commit 9bb870f715bdf91026578628a6afaf5325f80183
Author: Yonghong Song <yonghong.song@linux.dev>

bpf, x86: Dispatch exception cleanup pads at run time

This commit updates the x86 BPF JIT to dispatch exception cleanup pads at run
time, setting up native cleanup tables and populating the epilogue IP to allow
redirecting a frame to its landing pad during an ORC stack walk.

> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
> index 6c7a0578760e7..fb7e8ca1aab21 100644
> --- a/arch/x86/net/bpf_jit_comp.c
> +++ b/arch/x86/net/bpf_jit_comp.c

[ ... ]

> @@ -3278,6 +3279,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
>  			seen_exit = true;
>  			/* Update cleanup_addr */
>  			ctx->cleanup_addr = proglen;
> +			/* Where an unwind sends a frame with no pad. */
> +			bpf_prog->aux->epilogue_ip = (u64)image + proglen;

[Severity: High]
What happens if a BPF subprogram only terminates via a noreturn kfunc like
bpf_throw(), meaning it lacks a BPF_EXIT instruction entirely?

Since the x86 JIT generates the epilogue lazily when the first BPF_EXIT is
encountered, wouldn't do_jit() produce no epilogue at all for such a program,
leaving epilogue_ip as 0?

If an exception unwinds through this subprogram's frame, the unwinder relies
on epilogue_ip to rewrite the return address. Wouldn't it hit this error path
and abort the unwind?

kernel/bpf/helpers.c:bpf_unwind_rewrite() {
    ...
    } else if (prog->aux->epilogue_ip) {
        *ra = prog->aux->epilogue_ip;
    } else {
        WARN_ON_ONCE(1);
        return false;
    }
}

Could this lead to corrupted exception handling that can be triggered by
user-supplied BPF programs?

>  			if (bpf_prog_was_classic(bpf_prog) &&
>  			    !ns_capable_noaudit(&init_user_ns, CAP_SYS_ADMIN)) {
>  				if (emit_spectre_bhb_barrier(&prog, ip, bpf_prog))

[ ... ]

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926050006.2213110-1-yonghong.song@linux.dev?part=11

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad
  2026-09-26  5:00 ` [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad Yonghong Song
@ 2026-09-26  5:15   ` sashiko-bot
  2026-09-26  8:21     ` Alexei Starovoitov
  2026-09-27  0:41     ` Yonghong Song
  2026-09-27 20:40   ` bot+bpf-ci
  1 sibling, 2 replies; 56+ messages in thread
From: sashiko-bot @ 2026-09-26  5:15 UTC (permalink / raw)
  To: Yonghong Song; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
- [High] bpf: verifier bypass of resource leaks on unwinding without pad
- [High] bpf: double free via bpf_unwind_resume returning to normal path in verifier
- [Medium] bpf: precision tracking corruption when backtracking through bpf_unwind_resume
- [Medium] bpf: CFG and liveness graph mismatch for bpf_unwind_resume
--

commit c4ed3d378ed3f7a547e662e1b6885fb3e54fef13
Author: Yonghong Song <yonghong.song@linux.dev>

bpf: Resume a covered call at its landing pad

This patch introduces the ability for a BPF program to resume execution at a
landing pad after an unwinding operation. It updates the verifier's control
flow and state management to correctly account for these transitions, treating
the landing pad as an ordinary second successor in the same frame.

> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
> index 0e38b9575328c..57665c67e66b3 100644
> --- a/kernel/bpf/backtrack.c
> +++ b/kernel/bpf/backtrack.c
[ ... ]
> @@ -434,8 +435,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
>  					return -EFAULT;
>  			}
>  
> +			if (bpf_exc_pad_of_call(env, idx) == subseq_idx) {
> +				/*
> +				 * We came from this call's landing pad, which
> +				 * runs in the caller's frame: on that path the
> +				 * callee's frame was never entered, so there is
> +				 * no frame to leave. The call clobbered r0-r5;
> +				 * r6-r9 and the stack are the caller's own and
> +				 * keep going back from here.
> +				 */
> +				bt_clear_reg(bt, BPF_REG_0);
> +				if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
> +					verifier_bug(env, "landing pad unexpected regs %x",
> +						     bt_reg_mask(bt));
> +					return -EFAULT;
> +				}
> +				return 0;

[Severity: Medium]
Does the BPF_CALL backtracking logic correctly handle bpf_unwind_resume()?

When backtracking from a caller's instruction into a callee's pad and
encountering bpf_unwind_resume(), it looks like backtrack_insn() processes it
as a standard call without transitioning the backtrack frame (like
bt_subprog_enter() does for BPF_EXIT).

Could this leave the backtrack frame in the caller's context while traversing
the callee, leading to precision tracking corruption?

>  			/* callx calls static subprogs only */
> -			if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
> +			} else if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
[ ... ]
> diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
> index 4e2b6985bc964..cfd4fe4049cae 100644
> --- a/kernel/bpf/cfg.c
> +++ b/kernel/bpf/cfg.c
[ ... ]
> @@ -678,6 +687,8 @@ static int visit_insn(int t, struct bpf_verifier_env *env)
>  				mark_subprog_changes_pkt_data(env, t);
>  			if (ret == 0 && bpf_is_throw_kfunc(insn))
>  				mark_subprog_might_throw(env, t);
> +			if (ret == 0 && bpf_is_unwind_kfunc(insn))
> +				mark_subprog_might_unwind(env, t);
>  		}
>  		return visit_func_call_insn(t, insns, env, insn->src_reg == BPF_PSEUDO_CALL);

[Severity: Medium]
Is there a mismatch between the control flow graph and liveness tracking here?

In visit_insn(), bpf_unwind_resume is processed as a standard kfunc call,
which pushes a fallthrough edge into the CFG via visit_func_call_insn().

However, in bpf_insn_successors() in liveness.c, it correctly acts as an exit:

> diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
> index 4e0273a8ceee4..c9ee4f10f7255 100644
> --- a/kernel/bpf/liveness.c
> +++ b/kernel/bpf/liveness.c
[ ... ]
> @@ -364,6 +364,9 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
>  		return jt;
>  	}
>  
> +	if (unlikely(bpf_is_unwind_resume_kfunc(insn)))
> +		return succ;
> +
>  	opcode_info = &opcode_info_tbl[BPF_CLASS(insn->code) | BPF_OP(insn->code)];

Could this mismatch, where the CFG considers the path reachable but liveness
sees 0 successors, violate verifier graph invariants?

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index fc3df452de2e4..77176250f866c 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
[ ... ]
> @@ -19167,6 +19167,40 @@ enum {
>  	INSN_IDX_UPDATED = 2,
>  };
>  
> +static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
> +{
> +	struct bpf_verifier_state *branch;
> +	struct bpf_func_state *frame;
> +	int pad = bpf_exc_pad_of_call(env, insn_idx);
> +
> +	if (pad < 0)
> +		return 0;
> +	branch = push_stack(env, pad, insn_idx, false);
> +	if (IS_ERR(branch))
> +		return PTR_ERR(branch);
> +	frame = branch->frame[branch->curframe];
> +	/*
> +	 * The state at that call with the caller-saved registers gone: the
> +	 * callee's epilogue put r6-r9 and the stack back on the way out.
> +	 */
> +	clear_caller_saved_regs(env, frame->regs);
> +	mark_reg_unknown(env, frame->regs, BPF_REG_0);
> +	return 0;
> +}
> +
> +static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx)
> +{
> +	struct bpf_func_state *frame = cur_func(env);
> +	int pad = bpf_exc_pad_of_call(env, *insn_idx);
> +
> +	if (pad < 0)
> +		return PROCESS_BPF_EXIT;

[Severity: High]
Could returning PROCESS_BPF_EXIT here bypass resource leak and lock checks?

If a BPF program acquires a resource like a spinlock and calls bpf_unwind()
without a covering landing pad, pad < 0 evaluates to true. The verifier then
pops the state and terminates the branch without invoking check_resource_leak()
as process_bpf_exit_full() would normally do.

Since do_check_insn() evaluates this kfunc before the active_locks check, could
this allow the program to unwind while holding a spinlock, causing a kernel
memory leak or deadlock at runtime?

> +	clear_caller_saved_regs(env, frame->regs);
> +	mark_reg_unknown(env, frame->regs, BPF_REG_0);
> +	*insn_idx = pad;
> +	return INSN_IDX_UPDATED;
> +}
[ ... ]
> @@ -19404,6 +19438,20 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>  
>  		env->jmps_processed++;
>  		if (opcode == BPF_CALL) {
> +			if (bpf_is_unwind_kfunc(insn))
> +				return process_bpf_unwind(env, &env->insn_idx);
> +			if (bpf_is_unwind_resume_kfunc(insn)) {
> +				/*
> +				 * Mark r0 a known zero -- unknown first, as
> +				 * the known-zero helper keeps the type it
> +				 * finds, which here is NOT_INIT. The fixups
> +				 * lower this to 'r0 = 0; exit', so the frame
> +				 * returns a real zero.
> +				 */
> +				mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
> +				mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
> +				return process_bpf_exit_full(env, do_print_state, false);
> +			}

[Severity: High]
Could this simulate a normal return from the subprogram and lead to a double 
free vulnerability?

Here, bpf_unwind_resume is handled by calling process_bpf_exit_full() with
exception_exit = false, returning to the normal instruction after the call.

However, earlier in push_cleanup_pad_branch(), the caller's landing pad branch
was already pushed using the pre-call state.

At runtime, the callee's pad actually jumps to the caller's pad, not the
instruction after the call. If a callee's pad frees a resource, the verifier
will check the caller's pad using the pre-call state (where the resource was
still allocated).

Could this allow the caller's pad to free the same resource a second time,
resulting in a use-after-free or double free vulnerability?

>  			if (env->cur_state->active_locks) {
>  				/* similar to static subprog calls callx is allowed under a lock */
>  				if (!bpf_is_callx(insn) &&

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926050006.2213110-1-yonghong.song@linux.dev?part=7

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk
  2026-09-26  5:00 ` [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
@ 2026-09-26  5:16   ` sashiko-bot
  2026-09-26 23:54     ` Yonghong Song
  2026-09-27 20:39   ` bot+bpf-ci
  1 sibling, 1 reply; 56+ messages in thread
From: sashiko-bot @ 2026-09-26  5:16 UTC (permalink / raw)
  To: Yonghong Song; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: bpf_throw() bypasses exception_callback_subprog check skipping cleanup pads
--

commit bfb14bc84bb67f59ae10ca034e824ec0357d7d80
Author: Yonghong Song <yonghong.song@linux.dev>

bpf: Prepare for an exception cleanup table before the CFG walk

This commit introduces bpf_prepare_cleanup_exceptions() to record landing pads
for BPF-to-BPF calls and bpf_unwind() within a cleanup record's range. It
ensures the JIT supports dispatching landing pads and refuses to combine
tables with an exception callback.

> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
> index b19fcbf49b7ef..1ed0370a171b7 100644
> --- a/kernel/bpf/exception.c
> +++ b/kernel/bpf/exception.c
> @@ -7,9 +7,56 @@
>  #include <linux/filter.h>
>  #include "exception.h"
>  
> +#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
> +
>  BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
>  BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
>  
> +static void mark_call_sites(struct bpf_verifier_env *env)
> +{
> +	u32 i, j;
> +
> +	for (i = 0; i < env->cleanup_info_cnt; i++) {
> +		struct bpf_cleanup_info *rec = &env->cleanup_info[i];
> +
> +		for (j = rec->begin_off; j < rec->end_off; j++) {
> +			struct bpf_insn *insn = &env->prog->insnsi[j];
> +
> +			if (!bpf_pseudo_call(insn) && !bpf_is_unwind_kfunc(insn))
> +				continue;

[Severity: Medium]
Since bpf_throw() is not matched here, does it skip receiving a landing pad?

If a program is loaded with .bpf_cleanup records and calls bpf_throw() without
specifying a custom exception callback, could this cause it to silently bypass
the cleanup execution during unwinding, leaking user-defined logical resources
like map-based counters?

> +			env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
> +		}
> +	}
> +}
> +
> +int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env)
> +{

[ ... ]

> +	if (!bpf_jit_supports_cleanup_pads() || !env->prog->jit_requested) {
> +		verbose(env,
> +			"exception cleanup needs a JIT that can dispatch landing pads\n");
> +		return -EOPNOTSUPP;
> +	}
> +	env->prog->jit_required = 1;
> +
> +	if (env->exception_callback_subprog) {

[Severity: Medium]
Does this check the state of env->exception_callback_subprog too early?

If a later pass dynamically inserts a default exception callback, could it
bypass this restriction and result in a mix of exception cleanup tables and
bpf_throw() usage?

> +		verbose(env,
> +			"exception cleanup table cannot be combined with an exception callback\n");
> +		return -EINVAL;
> +	}
> +
> +	mark_call_sites(env);
> +	return 0;
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926050006.2213110-1-yonghong.song@linux.dev?part=5

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume
  2026-09-26  5:00 ` [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume Yonghong Song
@ 2026-09-26  5:17   ` sashiko-bot
  2026-09-27  3:06     ` Yonghong Song
  2026-09-27 20:40   ` bot+bpf-ci
  1 sibling, 1 reply; 56+ messages in thread
From: sashiko-bot @ 2026-09-26  5:17 UTC (permalink / raw)
  To: Yonghong Song; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [Critical] bpf: bpf_unwind bypasses resource leak checks when no landing pad is present
- [High] bpf: bpf_unwind_resume creates invalid ghost paths by returning to caller's normal path
- [Medium] bpf: Speculative branches to landing pads bypass nospec recovery
--

commit 3aafeed4814ca53533c421881eb0f30857cc42a0
Author: Yonghong Song <yonghong.song@linux.dev>
Subject: bpf: Refuse a landing pad that does not resume

This commit restricts exception landing pads in BPF to only support cleanup
pads that resume via bpf_unwind_resume(). It adds checks in the verifier
to ensure that instructions are not reached both inside and outside a pad,
and it refuses unsupported program exits within a pad.

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
> @@ -19197,6 +19202,7 @@ static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx)
>  	int pad = bpf_exc_pad_of_call(env, *insn_idx);
>  
>  	if (pad < 0)
>  		return PROCESS_BPF_EXIT;

[Severity: Critical]
Does this early return in process_bpf_unwind() bypass resource leak checks?

If a program acquires a resource like a bpf_spin_lock and throws an
exception where there is no landing pad, it appears this directs the verifier
to pop the state without calling check_resource_leak(), allowing the
program to leak the lock at runtime.

[ ... ]

> @@ -19450,6 +19461,7 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>  				 */
>  				mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
>  				mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
> +				cur_func(env)->in_pad = false;
>  				return process_bpf_exit_full(env, do_print_state, false);

[Severity: High]
Could passing false to process_bpf_exit_full() here instruct the verifier to
continue at the callsite's normal return path instead of unwinding to the
caller's landing pad?

Because push_cleanup_pad_branch() eagerly pushed the caller's landing pad
with the pre-call state, returning to the normal success path might merge
the post-cleanup state incorrectly and cause a divergence between verifier
control flow and runtime execution.

>  			}

[ ... ]

> @@ -19578,6 +19590,12 @@ static int do_check(struct bpf_verifier_env *env)
>  			}
>  		}
>  
> +		if (unlikely(env->cleanup_info_cnt)) {
> +			err = bpf_exc_check_insn(env, insn);
> +			if (err)
> +				return err;

[Severity: Medium]
When do_check() returns an error from bpf_exc_check_insn() due to a
speculative visit to a landing pad, does it bypass the subsequent
error_recoverable_with_nospec() check entirely?

If the verifier speculatively explores a dead branch targeting a landing pad,
this might cause the verifier to fatally abort instead of properly recovering
by inserting a speculation barrier.

> +		}
> +

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926050006.2213110-1-yonghong.song@linux.dev?part=8

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags
  2026-09-26  5:01 ` [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
@ 2026-09-26  5:18   ` sashiko-bot
  2026-09-27  4:58     ` Yonghong Song
  2026-09-27 20:24   ` bot+bpf-ci
  1 sibling, 1 reply; 56+ messages in thread
From: sashiko-bot @ 2026-09-26  5:18 UTC (permalink / raw)
  To: Yonghong Song; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] bpf: selftests: out-of-bounds heap access in access_global_var
- [Medium] bpf: selftests: multiple __set_global or __ret_global tags are silently overwritten
--

commit 81344b890bb87ceeed0ef105e0aab6e9d6797e2d
Author: Yonghong Song <yonghong.song@linux.dev>

selftests/bpf: Add __set_global() and __ret_global() test tags

This commit adds __set_global() and __ret_global() tags to the BPF
test loader. These tags allow setting a program's global variable before
execution and verifying its value afterwards.

> diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
> --- a/tools/testing/selftests/bpf/test_loader.c
> +++ b/tools/testing/selftests/bpf/test_loader.c

[ ... ]

> @@ -311,6 +326,142 @@ static int parse_caps(const char *str, __u64 *val, const char *name)
> +static int parse_global_var(const char *str, char **var, __u64 *val, const char *name)
> +{
> +	const char *colon = strrchr(str, ':');
> +	char *end;

[ ... ]

> +		term = end + 1;
> +	}
> +
> +	free(*var);
> +	*var = strndup(str, colon - str);

[Severity: Medium]
If a test author uses multiple __set_global() or __ret_global() tags in a
single test, does this silently overwrite the previous allocations? Only the
last tag seems to take effect, making it impossible to configure multiple
global variables in one test.

> +	if (!*var) {
> +		PRINT_FAIL("failed to allocate %s variable name\n", name);
> +		return -ENOMEM;
> +	}
> +
> +	return 0;
> +}

[ ... ]

> +static int access_global_var(struct bpf_object *obj, const char *name,
> +			     __u64 *val, bool set)
> +{
> +	__u32 off, sz, zero = 0;
> +	struct bpf_map *map;
> +	size_t vsz;
> +	void *buf;
> +	int err;
> +
> +	err = find_global_var(obj, name, &map, &off, &sz);
> +	if (err)
> +		return err;
> +
> +	vsz = bpf_map__value_size(map);
> +	buf = calloc(1, vsz);
> +	if (!buf)
> +		return -ENOMEM;
> +
> +	err = bpf_map__lookup_elem(map, &zero, sizeof(zero), buf, vsz, 0);
> +	if (err) {
> +		PRINT_FAIL("failed to read '%s': %d\n", name, err);
> +		goto out;
> +	}
> +	if (!set) {
> +		*val = sz == 4 ? *(__u32 *)(buf + off) : *(__u64 *)(buf + off);

[Severity: Medium]
Is it possible to have an out-of-bounds heap read/write here? The function
find_global_var() extracts off and sz from BTF, but off + sz is never
validated against the buffer size vsz before accessing buf + off.

> +		goto out;
> +	}
> +	if (sz == 4)
> +		*(__u32 *)(buf + off) = *val;
> +	else
> +		*(__u64 *)(buf + off) = *val;
> +	err = bpf_map__update_elem(map, &zero, sizeof(zero), buf, vsz, 0);
> +	if (err)
> +		PRINT_FAIL("failed to write '%s': %d\n", name, err);
> +out:
> +	free(buf);
> +	return err;
> +}

[ ... ]

> @@ -557,6 +708,24 @@ static int parse_test_spec(struct test_loader *tester,
>  			spec->mode_mask |= UNPRIV;
>  			spec->unpriv.execute = true;
>  			has_unpriv_retval = true;
> +		} else if ((val = str_has_pfx(s, "test_global_set="))) {
> +			err = parse_global_var(val, &spec->priv.set_global_var,
> +					       &spec->priv.set_global_val,
> +					       "__set_global");

[Severity: Medium]
Storing the variable in a singular field here also means that using multiple
__set_global() or __ret_global() tags will result in the earlier ones being
ignored. Should this support a list of variables instead?

> +			if (err)
> +				goto cleanup;
> +			spec->priv.has_set_global = true;
> +			spec->priv.execute = true;
> +			spec->mode_mask |= PRIV;

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926050006.2213110-1-yonghong.song@linux.dev?part=19

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG
  2026-09-26  5:00 ` [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG Yonghong Song
@ 2026-09-26  5:21   ` sashiko-bot
  2026-09-27  0:02     ` Yonghong Song
  2026-09-27 20:40   ` bot+bpf-ci
  1 sibling, 1 reply; 56+ messages in thread
From: sashiko-bot @ 2026-09-26  5:21 UTC (permalink / raw)
  To: Yonghong Song; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: DFS traversal invariant violation in visit_func_call_insn
--

commit 9f815c7b6f7049f1604884b98686a7516748e963
Author: Yonghong Song <yonghong.song@linux.dev>

bpf: Make exception landing pads reachable in the CFG

A bpf_unwind() or a bpf2bpf call inside the [begin_off, end_off) range of a
cleanup record can reach that record's landing pad. Add that edge to the
CFG walk, which explores the pad and makes both ends prune points, and to
bpf_insn_successors(), which liveness and the SCC passes walk.

> diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
> index b0bd9ba951dfb..4e2b6985bc964 100644
> --- a/kernel/bpf/cfg.c
> +++ b/kernel/bpf/cfg.c

[ ... ]

> @@ -160,17 +161,64 @@ static int push_insn(int t, int w, int e, struct bpf_verifier_env *env)
>  	return DONE_EXPLORING;
>  }
>  
> +static int visit_cleanup_pad_edge(int t, struct bpf_verifier_env *env)
> +{
> +	int *insn_stack = env->cfg.insn_stack;
> +	int *insn_state = env->cfg.insn_state;
> +	int w;
> +
> +	if (!env->cleanup_info_cnt)
> +		return DONE_EXPLORING;
> +	w = bpf_exc_pad_of_call(env, t);
> +	if (w < 0)
> +		return DONE_EXPLORING;
> +
> +	/*
> +	 * @t is a call that may branch here, and @w is the target of that
> +	 * branch, so both are prune points. @w especially: every covered call
> +	 * site in a region unwinds to the same pad, and without a prune point
> +	 * at its head the verifier walks the pad again for each of them.
> +	 */
> +	mark_prune_point(env, t);
> +	mark_prune_point(env, w);
> +	mark_jmp_point(env, w);
> +	mark_jump_target(env, w);
> +
> +	if (insn_state[w])
> +		return DONE_EXPLORING;
> +	if (env->cfg.cur_stack >= env->prog->len)
> +		return -E2BIG;
> +	insn_stack[env->cfg.cur_stack++] = w;
> +	insn_state[w] |= DISCOVERED;

[Severity: Medium]
Does setting the DISCOVERED flag here in visit_cleanup_pad_edge() mark the
pad prematurely?

By marking it DISCOVERED before fully exploring its path, any valid
cross-edges to the cleanup pad (e.g., from shared cleanup code) might be
incorrectly classified as back-edges and rejected by the verifier.

> +	return KEEP_EXPLORING;
> +}
> +

[ ... ]

>  static int visit_func_call_insn(int t, struct bpf_insn *insns,
>  				struct bpf_verifier_env *env,
>  				bool visit_callee)
>  {
> -	int ret, insn_sz;
> +	int ret, insn_sz, pad_ret;
>  	int w;
>  
> +	pad_ret = visit_cleanup_pad_edge(t, env);
> +	if (pad_ret < 0)
> +		return pad_ret;
> +
>  	insn_sz = bpf_is_ldimm64(&insns[t]) ? 2 : 1;
>  	ret = push_insn(t, t + insn_sz, FALLTHROUGH, env);

[Severity: Medium]
Is it safe to push the fallthrough edge even if visit_cleanup_pad_edge()
has just returned KEEP_EXPLORING?

The CFG checker typically relies on returning KEEP_EXPLORING immediately
after pushing a new node to maintain the invariant that DISCOVERED nodes are
strictly ancestors. By ignoring whether pad_ret is KEEP_EXPLORING and
eagerly pushing the fallthrough edge, this leaves the cleanup pad marked
DISCOVERED while the fallthrough path is being explored.

Could this lead to loop-free programs being incorrectly rejected if a
cross-edge on the fallthrough path is misidentified as a back-edge?

>  	if (ret)
> -		return ret;
> +		return merge_visit_ret(pad_ret, ret);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926050006.2213110-1-yonghong.song@linux.dev?part=6

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad
  2026-09-26  5:15   ` sashiko-bot
@ 2026-09-26  8:21     ` Alexei Starovoitov
  2026-09-27  0:04       ` Yonghong Song
  2026-09-27  0:41     ` Yonghong Song
  1 sibling, 1 reply; 56+ messages in thread
From: Alexei Starovoitov @ 2026-09-26  8:21 UTC (permalink / raw)
  To: sashiko-reviews, Yonghong Song; +Cc: bpf

On Sat Sep 26, 2026 at 5:15 AM UTC, sashiko-bot wrote:
> Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
> - [High] bpf: verifier bypass of resource leaks on unwinding without pad
> - [High] bpf: double free via bpf_unwind_resume returning to normal path in verifier
> - [Medium] bpf: precision tracking corruption when backtracking through bpf_unwind_resume
> - [Medium] bpf: CFG and liveness graph mismatch for bpf_unwind_resume

These and other bot findings look real.
Please tell your clanker to look for bugs harder.
sashiko shouldn't be finding such severe bugs.

pw-bot: cr

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk
  2026-09-26  5:16   ` sashiko-bot
@ 2026-09-26 23:54     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-26 23:54 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf



On 9/25/26 10:16 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Medium] bpf: bpf_throw() bypasses exception_callback_subprog check skipping cleanup pads
> --
>
> commit bfb14bc84bb67f59ae10ca034e824ec0357d7d80
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Prepare for an exception cleanup table before the CFG walk
>
> This commit introduces bpf_prepare_cleanup_exceptions() to record landing pads
> for BPF-to-BPF calls and bpf_unwind() within a cleanup record's range. It
> ensures the JIT supports dispatching landing pads and refuses to combine
> tables with an exception callback.
>
>> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
>> index b19fcbf49b7ef..1ed0370a171b7 100644
>> --- a/kernel/bpf/exception.c
>> +++ b/kernel/bpf/exception.c
>> @@ -7,9 +7,56 @@
>>   #include <linux/filter.h>
>>   #include "exception.h"
>>   
>> +#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
>> +
>>   BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
>>   BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
>>   
>> +static void mark_call_sites(struct bpf_verifier_env *env)
>> +{
>> +	u32 i, j;
>> +
>> +	for (i = 0; i < env->cleanup_info_cnt; i++) {
>> +		struct bpf_cleanup_info *rec = &env->cleanup_info[i];
>> +
>> +		for (j = rec->begin_off; j < rec->end_off; j++) {
>> +			struct bpf_insn *insn = &env->prog->insnsi[j];
>> +
>> +			if (!bpf_pseudo_call(insn) && !bpf_is_unwind_kfunc(insn))
>> +				continue;
> [Severity: Medium]
> Since bpf_throw() is not matched here, does it skip receiving a landing pad?
>
> If a program is loaded with .bpf_cleanup records and calls bpf_throw() without
> specifying a custom exception callback, could this cause it to silently bypass
> the cleanup execution during unwinding, leaking user-defined logical resources
> like map-based counters?

Yes, when bpf_throw() is triggered and landing pad has some resource to be
released, it is possible those resource is not released. Will reject
bpf_throw() if *any* landing pad exists in the program.

>
>> +			env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
>> +		}
>> +	}
>> +}
>> +
>> +int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env)
>> +{
> [ ... ]
>
>> +	if (!bpf_jit_supports_cleanup_pads() || !env->prog->jit_requested) {
>> +		verbose(env,
>> +			"exception cleanup needs a JIT that can dispatch landing pads\n");
>> +		return -EOPNOTSUPP;
>> +	}
>> +	env->prog->jit_required = 1;
>> +
>> +	if (env->exception_callback_subprog) {
> [Severity: Medium]
> Does this check the state of env->exception_callback_subprog too early?
>
> If a later pass dynamically inserts a default exception callback, could it
> bypass this restriction and result in a mix of exception cleanup tables and
> bpf_throw() usage?

In the beginning of bpf_prepare_cleanup_exceptions(), we have

+       if (!env->cleanup_info_cnt)
+               return 0;

so landing_pad will exist in the prog. With previous explanation, that
means bpf_throw() will be rejected. So reject exception_callback_subprog
too.

>
>> +		verbose(env,
>> +			"exception cleanup table cannot be combined with an exception callback\n");
>> +		return -EINVAL;
>> +	}
>> +
>> +	mark_call_sites(env);
>> +	return 0;
>> +}


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG
  2026-09-26  5:21   ` sashiko-bot
@ 2026-09-27  0:02     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-27  0:02 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf



On 9/25/26 10:21 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Medium] bpf: DFS traversal invariant violation in visit_func_call_insn
> --
>
> commit 9f815c7b6f7049f1604884b98686a7516748e963
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Make exception landing pads reachable in the CFG
>
> A bpf_unwind() or a bpf2bpf call inside the [begin_off, end_off) range of a
> cleanup record can reach that record's landing pad. Add that edge to the
> CFG walk, which explores the pad and makes both ends prune points, and to
> bpf_insn_successors(), which liveness and the SCC passes walk.
>
>> diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
>> index b0bd9ba951dfb..4e2b6985bc964 100644
>> --- a/kernel/bpf/cfg.c
>> +++ b/kernel/bpf/cfg.c
> [ ... ]
>
>> @@ -160,17 +161,64 @@ static int push_insn(int t, int w, int e, struct bpf_verifier_env *env)
>>   	return DONE_EXPLORING;
>>   }
>>   
>> +static int visit_cleanup_pad_edge(int t, struct bpf_verifier_env *env)
>> +{
>> +	int *insn_stack = env->cfg.insn_stack;
>> +	int *insn_state = env->cfg.insn_state;
>> +	int w;
>> +
>> +	if (!env->cleanup_info_cnt)
>> +		return DONE_EXPLORING;
>> +	w = bpf_exc_pad_of_call(env, t);
>> +	if (w < 0)
>> +		return DONE_EXPLORING;
>> +
>> +	/*
>> +	 * @t is a call that may branch here, and @w is the target of that
>> +	 * branch, so both are prune points. @w especially: every covered call
>> +	 * site in a region unwinds to the same pad, and without a prune point
>> +	 * at its head the verifier walks the pad again for each of them.
>> +	 */
>> +	mark_prune_point(env, t);
>> +	mark_prune_point(env, w);
>> +	mark_jmp_point(env, w);
>> +	mark_jump_target(env, w);
>> +
>> +	if (insn_state[w])
>> +		return DONE_EXPLORING;
>> +	if (env->cfg.cur_stack >= env->prog->len)
>> +		return -E2BIG;
>> +	insn_stack[env->cfg.cur_stack++] = w;
>> +	insn_state[w] |= DISCOVERED;
> [Severity: Medium]
> Does setting the DISCOVERED flag here in visit_cleanup_pad_edge() mark the
> pad prematurely?
>
> By marking it DISCOVERED before fully exploring its path, any valid
> cross-edges to the cleanup pad (e.g., from shared cleanup code) might be
> incorrectly classified as back-edges and rejected by the verifier.

The actual problem is below due to premature dfs stop's. See below.

>
>> +	return KEEP_EXPLORING;
>> +}
>> +
> [ ... ]
>
>>   static int visit_func_call_insn(int t, struct bpf_insn *insns,
>>   				struct bpf_verifier_env *env,
>>   				bool visit_callee)
>>   {
>> -	int ret, insn_sz;
>> +	int ret, insn_sz, pad_ret;
>>   	int w;
>>   
>> +	pad_ret = visit_cleanup_pad_edge(t, env);
>> +	if (pad_ret < 0)
>> +		return pad_ret;
>> +
>>   	insn_sz = bpf_is_ldimm64(&insns[t]) ? 2 : 1;
>>   	ret = push_insn(t, t + insn_sz, FALLTHROUGH, env);
> [Severity: Medium]
> Is it safe to push the fallthrough edge even if visit_cleanup_pad_edge()
> has just returned KEEP_EXPLORING?
>
> The CFG checker typically relies on returning KEEP_EXPLORING immediately
> after pushing a new node to maintain the invariant that DISCOVERED nodes are
> strictly ancestors. By ignoring whether pad_ret is KEEP_EXPLORING and
> eagerly pushing the fallthrough edge, this leaves the cleanup pad marked
> DISCOVERED while the fallthrough path is being explored.
>
> Could this lead to loop-free programs being incorrectly rejected if a
> cross-edge on the fallthrough path is misidentified as a back-edge?

The above
	if (pad_ret < 0)
		return pad_ret;
is not correct and it may cause premature dfs stop (ignoring KEEP_EXPLORING).
The correct one is
	if (pad_ret != DONE_EXPLORING)
		return pad_ret;
This way, pad_ret can be either KEEP_EXPLORING or err.
If KEEP_EXPLORING, dfs can continue with landing pad.

>
>>   	if (ret)
>> -		return ret;
>> +		return merge_visit_ret(pad_ret, ret);


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad
  2026-09-26  8:21     ` Alexei Starovoitov
@ 2026-09-27  0:04       ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-27  0:04 UTC (permalink / raw)
  To: Alexei Starovoitov, sashiko-reviews; +Cc: bpf



On 9/26/26 1:21 AM, Alexei Starovoitov wrote:
> On Sat Sep 26, 2026 at 5:15 AM UTC, sashiko-bot wrote:
>> Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
>> - [High] bpf: verifier bypass of resource leaks on unwinding without pad
>> - [High] bpf: double free via bpf_unwind_resume returning to normal path in verifier
>> - [Medium] bpf: precision tracking corruption when backtracking through bpf_unwind_resume
>> - [Medium] bpf: CFG and liveness graph mismatch for bpf_unwind_resume
> These and other bot findings look real.
> Please tell your clanker to look for bugs harder.
> sashiko shouldn't be finding such severe bugs.

Indeed, missed several critical things. Ya, need ask
more hard questions for some codes...

>
> pw-bot: cr


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad
  2026-09-26  5:15   ` sashiko-bot
  2026-09-26  8:21     ` Alexei Starovoitov
@ 2026-09-27  0:41     ` Yonghong Song
  1 sibling, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-27  0:41 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf



On 9/25/26 10:15 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
> - [High] bpf: verifier bypass of resource leaks on unwinding without pad
> - [High] bpf: double free via bpf_unwind_resume returning to normal path in verifier
> - [Medium] bpf: precision tracking corruption when backtracking through bpf_unwind_resume
> - [Medium] bpf: CFG and liveness graph mismatch for bpf_unwind_resume
> --
>
> commit c4ed3d378ed3f7a547e662e1b6885fb3e54fef13
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Resume a covered call at its landing pad
>
> This patch introduces the ability for a BPF program to resume execution at a
> landing pad after an unwinding operation. It updates the verifier's control
> flow and state management to correctly account for these transitions, treating
> the landing pad as an ordinary second successor in the same frame.
>
>> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
>> index 0e38b9575328c..57665c67e66b3 100644
>> --- a/kernel/bpf/backtrack.c
>> +++ b/kernel/bpf/backtrack.c
> [ ... ]
>> @@ -434,8 +435,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
>>   					return -EFAULT;
>>   			}
>>   
>> +			if (bpf_exc_pad_of_call(env, idx) == subseq_idx) {
>> +				/*
>> +				 * We came from this call's landing pad, which
>> +				 * runs in the caller's frame: on that path the
>> +				 * callee's frame was never entered, so there is
>> +				 * no frame to leave. The call clobbered r0-r5;
>> +				 * r6-r9 and the stack are the caller's own and
>> +				 * keep going back from here.
>> +				 */
>> +				bt_clear_reg(bt, BPF_REG_0);
>> +				if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
>> +					verifier_bug(env, "landing pad unexpected regs %x",
>> +						     bt_reg_mask(bt));
>> +					return -EFAULT;
>> +				}
>> +				return 0;
> [Severity: Medium]
> Does the BPF_CALL backtracking logic correctly handle bpf_unwind_resume()?

No. See below.

>
> When backtracking from a caller's instruction into a callee's pad and
> encountering bpf_unwind_resume(), it looks like backtrack_insn() processes it
> as a standard call without transitioning the backtrack frame (like
> bt_subprog_enter() does for BPF_EXIT).
>
> Could this leave the backtrack frame in the caller's context while traversing
> the callee, leading to precision tracking corruption?

Yes. I missed handling function bpf_is_unwind_resume_kfunc() which will
provide bt_subprog_enter().

>
>>   			/* callx calls static subprogs only */
>> -			if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
>> +			} else if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
> [ ... ]
>> diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
>> index 4e2b6985bc964..cfd4fe4049cae 100644
>> --- a/kernel/bpf/cfg.c
>> +++ b/kernel/bpf/cfg.c
> [ ... ]
>> @@ -678,6 +687,8 @@ static int visit_insn(int t, struct bpf_verifier_env *env)
>>   				mark_subprog_changes_pkt_data(env, t);
>>   			if (ret == 0 && bpf_is_throw_kfunc(insn))
>>   				mark_subprog_might_throw(env, t);
>> +			if (ret == 0 && bpf_is_unwind_kfunc(insn))
>> +				mark_subprog_might_unwind(env, t);
>>   		}
>>   		return visit_func_call_insn(t, insns, env, insn->src_reg == BPF_PSEUDO_CALL);
> [Severity: Medium]
> Is there a mismatch between the control flow graph and liveness tracking here?
>
> In visit_insn(), bpf_unwind_resume is processed as a standard kfunc call,
> which pushes a fallthrough edge into the CFG via visit_func_call_insn().
>
> However, in bpf_insn_successors() in liveness.c, it correctly acts as an exit:
>
>> diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
>> index 4e0273a8ceee4..c9ee4f10f7255 100644
>> --- a/kernel/bpf/liveness.c
>> +++ b/kernel/bpf/liveness.c
> [ ... ]
>> @@ -364,6 +364,9 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
>>   		return jt;
>>   	}
>>   
>> +	if (unlikely(bpf_is_unwind_resume_kfunc(insn)))
>> +		return succ;
>> +
>>   	opcode_info = &opcode_info_tbl[BPF_CLASS(insn->code) | BPF_OP(insn->code)];
> Could this mismatch, where the CFG considers the path reachable but liveness
> sees 0 successors, violate verifier graph invariants?

Okay, the below is an optimization:

+	if (unlikely(bpf_is_unwind_resume_kfunc(insn)))
+		return succ;

let me remove it so the number of successor's will be consistent between
cfg and liveness.

>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index fc3df452de2e4..77176250f866c 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
> [ ... ]
>> @@ -19167,6 +19167,40 @@ enum {
>>   	INSN_IDX_UPDATED = 2,
>>   };
>>   
>> +static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
>> +{
>> +	struct bpf_verifier_state *branch;
>> +	struct bpf_func_state *frame;
>> +	int pad = bpf_exc_pad_of_call(env, insn_idx);
>> +
>> +	if (pad < 0)
>> +		return 0;
>> +	branch = push_stack(env, pad, insn_idx, false);
>> +	if (IS_ERR(branch))
>> +		return PTR_ERR(branch);
>> +	frame = branch->frame[branch->curframe];
>> +	/*
>> +	 * The state at that call with the caller-saved registers gone: the
>> +	 * callee's epilogue put r6-r9 and the stack back on the way out.
>> +	 */
>> +	clear_caller_saved_regs(env, frame->regs);
>> +	mark_reg_unknown(env, frame->regs, BPF_REG_0);
>> +	return 0;
>> +}
>> +
>> +static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx)
>> +{
>> +	struct bpf_func_state *frame = cur_func(env);
>> +	int pad = bpf_exc_pad_of_call(env, *insn_idx);
>> +
>> +	if (pad < 0)
>> +		return PROCESS_BPF_EXIT;
> [Severity: High]
> Could returning PROCESS_BPF_EXIT here bypass resource leak and lock checks?
>
> If a BPF program acquires a resource like a spinlock and calls bpf_unwind()
> without a covering landing pad, pad < 0 evaluates to true. The verifier then
> pops the state and terminates the branch without invoking check_resource_leak()
> as process_bpf_exit_full() would normally do.
>
> Since do_check_insn() evaluates this kfunc before the active_locks check, could
> this allow the program to unwind while holding a spinlock, causing a kernel
> memory leak or deadlock at runtime?

Right, check_resource_leak() is a must here.

>
>> +	clear_caller_saved_regs(env, frame->regs);
>> +	mark_reg_unknown(env, frame->regs, BPF_REG_0);
>> +	*insn_idx = pad;
>> +	return INSN_IDX_UPDATED;
>> +}
> [ ... ]
>> @@ -19404,6 +19438,20 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>>   
>>   		env->jmps_processed++;
>>   		if (opcode == BPF_CALL) {
>> +			if (bpf_is_unwind_kfunc(insn))
>> +				return process_bpf_unwind(env, &env->insn_idx);
>> +			if (bpf_is_unwind_resume_kfunc(insn)) {
>> +				/*
>> +				 * Mark r0 a known zero -- unknown first, as
>> +				 * the known-zero helper keeps the type it
>> +				 * finds, which here is NOT_INIT. The fixups
>> +				 * lower this to 'r0 = 0; exit', so the frame
>> +				 * returns a real zero.
>> +				 */
>> +				mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
>> +				mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
>> +				return process_bpf_exit_full(env, do_print_state, false);
>> +			}
> [Severity: High]
> Could this simulate a normal return from the subprogram and lead to a double
> free vulnerability?
>
> Here, bpf_unwind_resume is handled by calling process_bpf_exit_full() with
> exception_exit = false, returning to the normal instruction after the call.
>
> However, earlier in push_cleanup_pad_branch(), the caller's landing pad branch
> was already pushed using the pre-call state.
>
> At runtime, the callee's pad actually jumps to the caller's pad, not the
> instruction after the call. If a callee's pad frees a resource, the verifier
> will check the caller's pad using the pre-call state (where the resource was
> still allocated).
>
> Could this allow the caller's pad to free the same resource a second time,
> resulting in a use-after-free or double free vulnerability?

Yes, it is possible.
	outer_call:				inner_call:
		bpf_rcu_read_lock
		call inner_call
						bpf_unwind
					landing_pad:
						bpf_rcu_read_unlock
						...
						bpf_unwind_resume
	landing_pad:
		bpf_rcu_read_unlock
		...

The outer_call landing_pad uses states at the beginning of 'call_inner_call'
which does have bpf_rcu_read_lock. So this caused double bpf_rcu_read_unlock.

To fix this issue, for every function call which has landing_pad,
checkpoint the entry state (bpf_rcu_read_lock, bpf_prompt_disable, etc.)
and right before bpf_unwind_resume to ensure the state is the same
as entry state. This will prevent the above double free.

>
>>   			if (env->cur_state->active_locks) {
>>   				/* similar to static subprog calls callx is allowed under a lock */
>>   				if (!bpf_is_callx(insn) &&


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume
  2026-09-26  5:17   ` sashiko-bot
@ 2026-09-27  3:06     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-27  3:06 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf



On 9/25/26 10:17 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
> - [Critical] bpf: bpf_unwind bypasses resource leak checks when no landing pad is present
> - [High] bpf: bpf_unwind_resume creates invalid ghost paths by returning to caller's normal path
> - [Medium] bpf: Speculative branches to landing pads bypass nospec recovery
> --
>
> commit 3aafeed4814ca53533c421881eb0f30857cc42a0
> Author: Yonghong Song <yonghong.song@linux.dev>
> Subject: bpf: Refuse a landing pad that does not resume
>
> This commit restricts exception landing pads in BPF to only support cleanup
> pads that resume via bpf_unwind_resume(). It adds checks in the verifier
> to ensure that instructions are not reached both inside and outside a pad,
> and it refuses unsupported program exits within a pad.
>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
>> @@ -19197,6 +19202,7 @@ static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx)
>>   	int pad = bpf_exc_pad_of_call(env, *insn_idx);
>>   
>>   	if (pad < 0)
>>   		return PROCESS_BPF_EXIT;
> [Severity: Critical]
> Does this early return in process_bpf_unwind() bypass resource leak checks?
>
> If a program acquires a resource like a bpf_spin_lock and throws an
> exception where there is no landing pad, it appears this directs the verifier
> to pop the state without calling check_resource_leak(), allowing the
> program to leak the lock at runtime.

This issue is covered in previous patch. Indeed, check_resource_leak()
is needed.

>
> [ ... ]
>
>> @@ -19450,6 +19461,7 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>>   				 */
>>   				mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
>>   				mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
>> +				cur_func(env)->in_pad = false;
>>   				return process_bpf_exit_full(env, do_print_state, false);
> [Severity: High]
> Could passing false to process_bpf_exit_full() here instruct the verifier to
> continue at the callsite's normal return path instead of unwinding to the
> caller's landing pad?

The above code is under branch:
	if (bpf_is_unwind_resume_kfunc(insn)) {
		...
                 mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
                 mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
                 return process_bpf_exit_full(env, do_print_state, false);		
	}
So control is ready to return to either bpf prog (including main) or
exit the prog itself. This one (cur_func(env)->in_pad = false)
is irrelevant.

>
> Because push_cleanup_pad_branch() eagerly pushed the caller's landing pad
> with the pre-call state, returning to the normal success path might merge
> the post-cleanup state incorrectly and cause a divergence between verifier
> control flow and runtime execution.

This is possible. This is addressed in previous commit. The idea is
the state (related to resources) at the beginning of the function
must be the same as when bpf_unwind_resume() does resource checking.

>
>>   			}
> [ ... ]
>
>> @@ -19578,6 +19590,12 @@ static int do_check(struct bpf_verifier_env *env)
>>   			}
>>   		}
>>   
>> +		if (unlikely(env->cleanup_info_cnt)) {
>> +			err = bpf_exc_check_insn(env, insn);
>> +			if (err)
>> +				return err;
> [Severity: Medium]
> When do_check() returns an error from bpf_exc_check_insn() due to a
> speculative visit to a landing pad, does it bypass the subsequent
> error_recoverable_with_nospec() check entirely?
>
> If the verifier speculatively explores a dead branch targeting a landing pad,
> this might cause the verifier to fatally abort instead of properly recovering
> by inserting a speculation barrier.
>
Yes, error_recoverable_with_nospec() should help to resolve the issue.

>> +		}
>> +


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time
  2026-09-26  5:15   ` sashiko-bot
@ 2026-09-27  4:35     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-27  4:35 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf



On 9/25/26 10:15 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf, x86: missing JIT epilogue for exit-less subprograms crashes exception unwinder
> --
>
> commit 9bb870f715bdf91026578628a6afaf5325f80183
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf, x86: Dispatch exception cleanup pads at run time
>
> This commit updates the x86 BPF JIT to dispatch exception cleanup pads at run
> time, setting up native cleanup tables and populating the epilogue IP to allow
> redirecting a frame to its landing pad during an ORC stack walk.
>
>> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
>> index 6c7a0578760e7..fb7e8ca1aab21 100644
>> --- a/arch/x86/net/bpf_jit_comp.c
>> +++ b/arch/x86/net/bpf_jit_comp.c
> [ ... ]
>
>> @@ -3278,6 +3279,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
>>   			seen_exit = true;
>>   			/* Update cleanup_addr */
>>   			ctx->cleanup_addr = proglen;
>> +			/* Where an unwind sends a frame with no pad. */
>> +			bpf_prog->aux->epilogue_ip = (u64)image + proglen;
> [Severity: High]
> What happens if a BPF subprogram only terminates via a noreturn kfunc like
> bpf_throw(), meaning it lacks a BPF_EXIT instruction entirely?

bpf_throw() will be rejected if landing_pad is available.

>
> Since the x86 JIT generates the epilogue lazily when the first BPF_EXIT is
> encountered, wouldn't do_jit() produce no epilogue at all for such a program,
> leaving epilogue_ip as 0?
>
> If an exception unwinds through this subprogram's frame, the unwinder relies
> on epilogue_ip to rewrite the return address. Wouldn't it hit this error path
> and abort the unwind?
>
> kernel/bpf/helpers.c:bpf_unwind_rewrite() {
>      ...
>      } else if (prog->aux->epilogue_ip) {
>          *ra = prog->aux->epilogue_ip;
>      } else {
>          WARN_ON_ONCE(1);
>          return false;
>      }
> }
>
> Could this lead to corrupted exception handling that can be triggered by
> user-supplied BPF programs?

Yes, it is possible that epilogue_ip is 0 if that subprog is always
in the path when unwinding happens. The fix will be keep 'exit' insn's
so they won'be removed. See bpf_jit_comp.c

                 case BPF_JMP | BPF_EXIT:
                         if (seen_exit) {
                                 jmp_offset = ctx->cleanup_addr - addrs[i];
                                 goto emit_jmp;
                         }
                         seen_exit = true;
                         /* Update cleanup_addr */
                         ctx->cleanup_addr = proglen;
                         /* Where an unwind sends a frame with no pad. */
                         bpf_prog->aux->epilogue_ip = (u64)image + proglen;
			...

>
>>   			if (bpf_prog_was_classic(bpf_prog) &&
>>   			    !ns_capable_noaudit(&init_user_ns, CAP_SYS_ADMIN)) {
>>   				if (emit_spectre_bhb_barrier(&prog, ip, bpf_prog))
> [ ... ]
>


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags
  2026-09-26  5:18   ` sashiko-bot
@ 2026-09-27  4:58     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-27  4:58 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf



On 9/25/26 10:18 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
> - [Medium] bpf: selftests: out-of-bounds heap access in access_global_var
> - [Medium] bpf: selftests: multiple __set_global or __ret_global tags are silently overwritten
> --
>
> commit 81344b890bb87ceeed0ef105e0aab6e9d6797e2d
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> selftests/bpf: Add __set_global() and __ret_global() test tags
>
> This commit adds __set_global() and __ret_global() tags to the BPF
> test loader. These tags allow setting a program's global variable before
> execution and verifying its value afterwards.
>
>> diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
>> --- a/tools/testing/selftests/bpf/test_loader.c
>> +++ b/tools/testing/selftests/bpf/test_loader.c
> [ ... ]
>
>> @@ -311,6 +326,142 @@ static int parse_caps(const char *str, __u64 *val, const char *name)
>> +static int parse_global_var(const char *str, char **var, __u64 *val, const char *name)
>> +{
>> +	const char *colon = strrchr(str, ':');
>> +	char *end;
> [ ... ]
>
>> +		term = end + 1;
>> +	}
>> +
>> +	free(*var);
>> +	*var = strndup(str, colon - str);
> [Severity: Medium]
> If a test author uses multiple __set_global() or __ret_global() tags in a
> single test, does this silently overwrite the previous allocations? Only the
> last tag seems to take effect, making it impossible to configure multiple
> global variables in one test.

Indeed, maybe an array of global's can work. I will explore.

>
>> +	if (!*var) {
>> +		PRINT_FAIL("failed to allocate %s variable name\n", name);
>> +		return -ENOMEM;
>> +	}
>> +
>> +	return 0;
>> +}
> [ ... ]
>
>> +static int access_global_var(struct bpf_object *obj, const char *name,
>> +			     __u64 *val, bool set)
>> +{
>> +	__u32 off, sz, zero = 0;
>> +	struct bpf_map *map;
>> +	size_t vsz;
>> +	void *buf;
>> +	int err;
>> +
>> +	err = find_global_var(obj, name, &map, &off, &sz);
>> +	if (err)
>> +		return err;
>> +
>> +	vsz = bpf_map__value_size(map);
>> +	buf = calloc(1, vsz);
>> +	if (!buf)
>> +		return -ENOMEM;
>> +
>> +	err = bpf_map__lookup_elem(map, &zero, sizeof(zero), buf, vsz, 0);
>> +	if (err) {
>> +		PRINT_FAIL("failed to read '%s': %d\n", name, err);
>> +		goto out;
>> +	}
>> +	if (!set) {
>> +		*val = sz == 4 ? *(__u32 *)(buf + off) : *(__u64 *)(buf + off);
> [Severity: Medium]
> Is it possible to have an out-of-bounds heap read/write here? The function
> find_global_var() extracts off and sz from BTF, but off + sz is never
> validated against the buffer size vsz before accessing buf + off.

Indeed, a bug. Will fix.

>
>> +		goto out;
>> +	}
>> +	if (sz == 4)
>> +		*(__u32 *)(buf + off) = *val;
>> +	else
>> +		*(__u64 *)(buf + off) = *val;
>> +	err = bpf_map__update_elem(map, &zero, sizeof(zero), buf, vsz, 0);
>> +	if (err)
>> +		PRINT_FAIL("failed to write '%s': %d\n", name, err);
>> +out:
>> +	free(buf);
>> +	return err;
>> +}
> [ ... ]
>
>> @@ -557,6 +708,24 @@ static int parse_test_spec(struct test_loader *tester,
>>   			spec->mode_mask |= UNPRIV;
>>   			spec->unpriv.execute = true;
>>   			has_unpriv_retval = true;
>> +		} else if ((val = str_has_pfx(s, "test_global_set="))) {
>> +			err = parse_global_var(val, &spec->priv.set_global_var,
>> +					       &spec->priv.set_global_val,
>> +					       "__set_global");
> [Severity: Medium]
> Storing the variable in a singular field here also means that using multiple
> __set_global() or __ret_global() tags will result in the earlier ones being
> ignored. Should this support a list of variables instead?

the same as previous, maybe an array of globals.

>
>> +			if (err)
>> +				goto cleanup;
>> +			spec->priv.has_set_global = true;
>> +			spec->priv.execute = true;
>> +			spec->mode_mask |= PRIV;


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags
  2026-09-26  5:01 ` [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
  2026-09-26  5:18   ` sashiko-bot
@ 2026-09-27 20:24   ` bot+bpf-ci
  2026-09-28  3:36     ` Yonghong Song
  1 sibling, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:24 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 6184 bytes --]

> diff --git a/tools/testing/selftests/bpf/progs/bpf_misc.h b/tools/testing/selftests/bpf/progs/bpf_misc.h
> index f3dbc3b59bff..9afa163fac5a 100644
> --- a/tools/testing/selftests/bpf/progs/bpf_misc.h
> +++ b/tools/testing/selftests/bpf/progs/bpf_misc.h

[ ... ]

> diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
> index 25eeb1c1248b..65b55f7fa907 100644
> --- a/tools/testing/selftests/bpf/test_loader.c
> +++ b/tools/testing/selftests/bpf/test_loader.c

[ ... ]

> @@ -311,6 +326,142 @@ static int parse_caps(const char *str, __u64 *val, const char *name)
>  	return 0;
>  }
>
> +static int access_global_var(struct bpf_object *obj, const char *name,
> +			     __u64 *val, bool set)
> +{
> +	__u32 off, sz, zero = 0;
> +	struct bpf_map *map;
> +	size_t vsz;
> +	void *buf;
> +	int err;
> +
> +	err = find_global_var(obj, name, &map, &off, &sz);
> +	if (err)
> +		return err;
> +
> +	vsz = bpf_map__value_size(map);
> +	buf = calloc(1, vsz);
> +	if (!buf)
> +		return -ENOMEM;
> +
> +	err = bpf_map__lookup_elem(map, &zero, sizeof(zero), buf, vsz, 0);
> +	if (err) {
> +		PRINT_FAIL("failed to read '%s': %d\n", name, err);
> +		goto out;
> +	}
> +	if (!set) {
> +		*val = sz == 4 ? *(__u32 *)(buf + off) : *(__u64 *)(buf + off);
> +		goto out;
> +	}
>                        ^^^^

Can __ret_global() match a negative expected value for a 4-byte variable?

parse_global_var() uses strtoull(), which takes a leading '-' and negates
it as a 64-bit unsigned value. So __ret_global(err, -22) stores
0xffffffffffffffea (64 bits).

But access_global_var() reads the 4-byte variable zero-extended:

    *val = sz == 4 ? *(__u32 *)(buf + off) : *(__u64 *)(buf + off);

and gets 0x00000000ffffffea (32 bits extended to 64). Then run_subtest()
compares the two as __u64:

    if (v != subspec->ret_global_val) {
        PRINT_FAIL("Unexpected %s: 0x%llx != 0x%llx\n", ...);

so the check always fails with "Unexpected err: 0xffffffea !=
0xffffffffffffffea", even though the program stored exactly -22.

This matters in practice. Existing selftests keep negative errnos in
4-byte globals and check them from a hand-written driver, e.g.
progs/test_bpf_nf.c has `int test_einval_reserved = 0;` checked in
prog_tests/bpf_nf.c with `ASSERT_EQ(skel->bss->test_einval_reserved,
-EINVAL, ...)`. The commit message says these tags are meant to replace
that kind of driver. The neighbouring __retval() also supports negative
values (parse_int() uses strtol() with a special case for INT_MIN), and 68
programs under progs/ use __retval(-...).

The write side does not have this problem. __set_global(err, -22)
truncates to 0xffffffea and stores -22 correctly, so the same literal
behaves differently in the two tags. 8-byte variables are fine because
both sides are 64-bit.

Should the expected value be truncated or sign-extended to the variable
width (for example, compare only the low 32 bits when sz == 4, or
sign-extend the value read back), or should out-of-range values be refused
at parse time?

A related issue: __set_global() on a 4-byte variable silently drops any
bits above 32:

> +	if (sz == 4)
> +		*(__u32 *)(buf + off) = *val;

with no error. The tree allows it, but none of the tests in this series
trigger it: they only use __u64 variables with non-negative values.

> +	else
> +		*(__u64 *)(buf + off) = *val;
> +	err = bpf_map__update_elem(map, &zero, sizeof(zero), buf, vsz, 0);
> +	if (err)
> +		PRINT_FAIL("failed to write '%s': %d\n", name, err);
> +out:
> +	free(buf);
> +	return err;
> +}

[ ... ]

> @@ -1532,6 +1720,13 @@ void run_subtest(struct test_loader *tester,
>  			}
>  		}
>
> +		if (subspec->has_set_global) {
> +			__u64 v = subspec->set_global_val;
> +
> +			if (access_global_var(tobj, subspec->set_global_var, &v, true))
> +				goto tobj_cleanup;
> +		}
> +
>  		err = do_prog_test_run(bpf_program__fd(tprog), &retval,
>  				       bpf_program__type(tprog) == BPF_PROG_TYPE_SYSCALL ? true : false,
>  				       spec->linear_sz);
> @@ -1540,6 +1735,18 @@ void run_subtest(struct test_loader *tester,
>  			goto tobj_cleanup;
>  		}
>
> +		if (subspec->has_ret_global) {
> +			__u64 v = 0;
> +
> +			if (access_global_var(tobj, subspec->ret_global_var, &v, false))
> +				goto tobj_cleanup;
> +			if (v != subspec->ret_global_val) {
> +				PRINT_FAIL("Unexpected %s: 0x%llx != 0x%llx\n",
> +					   subspec->ret_global_var, v, subspec->ret_global_val);
> +				goto tobj_cleanup;
> +			}
> +		}
>                    ^^^^

For a 4-byte variable, __ret_global() cannot match any expected value
outside [0, UINT32_MAX], and that range includes every negative literal.

parse_global_var() parses the value with strtoull(), so XSTR(-1) ("-1")
becomes 0xffffffffffffffff. access_global_var() reads a 4-byte variable as

    *val = sz == 4 ? *(__u32 *)(buf + off) : *(__u64 *)(buf + off);

which zero-extends it to 0x00000000ffffffff. The comparison here then
always fails for `int cnt; ... __ret_global(cnt, -1)`, printing
"Unexpected cnt: 0xffffffff != 0xffffffffffffffff" even when the program
set cnt to -1.

__set_global() does not have this problem: the same value truncated into a
__u32 stores -1 correctly. So the two tags do not agree on what a 4-byte
value is.

Values too wide for the variable are also not rejected:
__set_global(int_var, 0x100000000) quietly writes 0.

The commit message says 4-byte variables are supported and cites veristat
as the model, but veristat's set_global_var() checks signedness and range
("Variable %s value %lld is out of range [%lld; %lld]").

Should the expected value be truncated or sign-adjusted to the variable
width (for example, compare only the low 32 bits when sz == 4), or should
out-of-range values be refused at parse time?

No test in the series hits this today; they all use the __u64 pads_ran
with positive masks.

> +
>  		verify_stderr(bpf_program__fd(tprog), &subspec->stderr);
>
>  		if (subspec->stdout.cnt) {


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 15/21] libbpf: Collect .bpf_cleanup records and pass them to the kernel
  2026-09-26  5:01 ` [PATCH bpf-next v6 15/21] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
@ 2026-09-27 20:39   ` bot+bpf-ci
  2026-09-28  3:28     ` Yonghong Song
  0 siblings, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:39 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 4655 bytes --]

> diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
> index 7eb15a013a83..4e9584fff264 100644
> --- a/tools/lib/bpf/libbpf.c
> +++ b/tools/lib/bpf/libbpf.c

[ ... ]

> @@ -4915,6 +4953,249 @@ static struct bpf_program *find_prog_by_sec_insn(const struct bpf_object *obj,
>  	return NULL;
>  }
>
> +static int bpf_object__init_cleanup_info(struct bpf_object *obj)
> +{
> +	Elf_Data *data = obj->efile.cleanup_data;
> +	Elf_Data *relo = NULL;
> +	size_t i, nrels, nslots, nrecs;
> +	struct cleanup_raw_rec *recs;
> +	int *slot_sec, ret = 0;
> +	size_t *slot_val;
> +	const __u32 *vals;
> +	Elf64_Shdr *sh;
> +	bool native;
> +
> +	if (!data || obj->efile.cleanup_shndx < 0 || !data->d_size)
> +		return 0;
> +
> +	native = is_native_endianness(obj);
> +
> +	for (i = 0; i < obj->efile.sec_cnt; i++) {
> +		struct elf_sec_desc *sd = &obj->efile.secs[i];
> +
> +		if (sd->sec_type == SEC_RELO && sd->shdr &&
> +		    sd->shdr->sh_info == (Elf64_Word)obj->efile.cleanup_shndx) {
> +			relo = sd->data;
> +			break;
> +		}
> +	}
> +	if (!relo) {
> +		pr_warn("%s present without relocations\n", CLEANUP_SEC);
> +		return -LIBBPF_ERRNO__FORMAT;
> +	}
> +	/*
> +	 * LLVM leaves sh_entsize unset, so this only bites a producer that
> +	 * declares a record size -- which is the one able to tell us it means
> +	 * something other than three 4-byte fields.
> +	 */
> +	sh = elf_sec_hdr(obj, elf_sec_by_idx(obj, obj->efile.cleanup_shndx));
> +	if (sh && sh->sh_entsize &&
> +	    sh->sh_entsize != sizeof(struct bpf_cleanup_info)) {
> +		pr_warn("%s record size %zu is not the expected %zu\n", CLEANUP_SEC,
> +			(size_t)sh->sh_entsize, sizeof(struct bpf_cleanup_info));
> +		return -LIBBPF_ERRNO__FORMAT;
> +	}

Should sh_entsize == 0 be explicitly treated as 12, so that wider
records (if necessary in the future) end up being easier to reject?

This design question was raised by Kumar Kartikeya Dwivedi in v5:
https://lore.kernel.org/bpf/DLMGMYOHSGFW.24JML4GSMTV2K@gmail.com/

The current code accepts sh_entsize == 0 (LLVM's default) silently, but
treating it explicitly as 12 would allow future tools to cleanly reject
wider records by setting sh_entsize to a different value.

> +	if (data->d_size % sizeof(struct bpf_cleanup_info)) {
> +		pr_warn("%s size %zu is not a multiple of the record size %zu\n",
> +			CLEANUP_SEC, data->d_size, sizeof(struct bpf_cleanup_info));
> +		return -LIBBPF_ERRNO__FORMAT;
> +	}
> +
> +	vals = data->d_buf;
> +	nslots = data->d_size / sizeof(__u32);
> +	nrecs = data->d_size / sizeof(struct bpf_cleanup_info);
> +
> +	slot_sec = calloc(nslots, sizeof(*slot_sec));
> +	slot_val = calloc(nslots, sizeof(*slot_val));
> +	recs = calloc(nrecs ?: 1, sizeof(*recs));

This isn't a bug, but could the relocation loop fill recs directly
(with sec_idx initialised to -1) so the slot_sec/slot_val scratch
arrays and the copy loop go away?

The relocation loop could write straight into recs, using sec_idx = -1
as the 'no relocation' sentinel and checking alignment while the byte
offset is at hand, the way bpf_object__collect_rodata_relos() checks
(sym->st_value + addend) % BPF_INSN_SZ inline. That would leave one
allocation instead of three and drop the copy loop.

Also, is the 'nrecs ?: 1' needed? d_size has already been checked to be
a non-zero multiple of the record size.

> +	if (!slot_sec || !slot_val || !recs) {
> +		ret = -ENOMEM;
> +		goto out;
> +	}
> +	for (i = 0; i < nslots; i++)
> +		slot_sec[i] = -1;
> +
> +	/* One relocation per 4-byte field, naming the section it points into. */
> +	nrels = relo->d_size / sizeof(Elf64_Rel);
> +	for (i = 0; i < nrels; i++) {

[ ... ]

> +	for (i = 0; i < nslots; i++) {
> +		struct cleanup_raw_rec *rec = &recs[i / CLEANUP_REC_FIELDS];
> +		size_t field = i % CLEANUP_REC_FIELDS;
> +
> +		if (slot_sec[i] < 0) {
> +			pr_warn("%s: field %zu has no relocation\n", CLEANUP_SEC, i);
> +			ret = -LIBBPF_ERRNO__FORMAT;
> +			goto out;
> +		}
> +		if (slot_val[i] % BPF_INSN_SZ) {
> +			pr_warn("%s: field %zu offset %zu is not instruction aligned\n",
> +				CLEANUP_SEC, i, slot_val[i]);
> +			ret = -LIBBPF_ERRNO__FORMAT;
> +			goto out;
> +		}
> +		rec->sec_idx[field] = slot_sec[i];
> +		rec->insn_idx[field] = slot_val[i] / BPF_INSN_SZ;
> +	}
> +
> +	obj->cleanup_recs = recs;
> +	obj->cleanup_rec_cnt = nrecs;
> +	recs = NULL;
> +out:
> +	free(recs);
> +	free(slot_val);
> +	free(slot_sec);
> +	return ret;
> +}

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk
  2026-09-26  5:00 ` [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
  2026-09-26  5:16   ` sashiko-bot
@ 2026-09-27 20:39   ` bot+bpf-ci
  2026-09-28  0:01     ` Yonghong Song
  1 sibling, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:39 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 4301 bytes --]

> diff --git a/include/linux/filter.h b/include/linux/filter.h
> index e42eccb0990e..972b3ed2a51d 100644
> --- a/include/linux/filter.h
> +++ b/include/linux/filter.h
> @@ -1248,6 +1248,7 @@ bool bpf_jit_supports_stack_args(void);
>  bool bpf_jit_supports_arena_args(void);
>  bool bpf_jit_supports_far_kfunc_call(void);
>  bool bpf_jit_supports_exceptions(void);
> +bool bpf_jit_supports_cleanup_pads(void);
>  bool bpf_jit_supports_ptr_xchg(void);
>  bool bpf_jit_supports_arena(void);
>  bool bpf_jit_supports_insn(struct bpf_insn *insn, bool in_arena);

[ ... ]

> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
> index b19fcbf49b7e..1ed0370a171b 100644
> --- a/kernel/bpf/exception.c
> +++ b/kernel/bpf/exception.c
> @@ -7,9 +7,56 @@
>  #include <linux/filter.h>
>  #include "exception.h"
>
> +#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
> +
>  BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
>  BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
>
> +static void mark_call_sites(struct bpf_verifier_env *env)
> +{
> +	u32 i, j;
> +
> +	for (i = 0; i < env->cleanup_info_cnt; i++) {
> +		struct bpf_cleanup_info *rec = &env->cleanup_info[i];
> +
> +		for (j = rec->begin_off; j < rec->end_off; j++) {
> +			struct bpf_insn *insn = &env->prog->insnsi[j];
> +
> +			if (!bpf_pseudo_call(insn) && !bpf_is_unwind_kfunc(insn))
> +				continue;
> +			env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
> +		}
> +	}
> +}

Does mark_call_sites() handle indirect BPF-to-BPF calls? The function
marks only direct calls (bpf_pseudo_call) and bpf_unwind(), but skips
callx instructions (BPF_JMP|BPF_CALL|BPF_X).

check_func_callx() in the verifier takes a PTR_TO_FUNC to a static
subprog and verifies it through check_static_func_call(), the same path
as direct calls. Nothing prevents that callee from calling bpf_unwind()
since bpf_exc_check_callback() only runs for helper/kfunc callbacks, not
static subprogs.

The commit message states cleanup_pad is recorded "for every call that
can unwind -- a BPF-to-BPF call or bpf_unwind()" but a callx inside the
[begin_off, end_off) range keeps cleanup_pad == 0.

The infrastructure that consumes cleanup_pad already handles callx:

kernel/bpf/verifier.c:
visit_func_call_insn() and push_cleanup_pad_branch() run for every
BPF_CALL instruction, including callx. If cleanup_pad is not set, they
treat the call as unable to unwind.

This means:
- visit_cleanup_pad_edge() -> bpf_exc_pad_of_call() returns -1, so no
  CFG edge is created to the landing pad
- push_cleanup_pad_branch() returns 0 before check_func_callx() runs, so
  the verifier never explores the landing pad state
- bpf_unwind_rewrite() finds no cleanup range for the frame and sets
  *ra = prog->aux->epilogue_ip

Concrete scenario: a C++ or Rust program makes a virtual or trait call
(lowered to callx) inside a scope whose drop glue sits in a landing pad.
If the callee unwinds, the caller frame goes to its epilogue and the
drop glue never runs. If that drop glue releases a kernel reference or
kptr, the release is skipped. Meanwhile the verifier modeled the
callee's resume as a normal return to callx+1.

Should the test be widened to include bpf_is_callx(), or should cleanup
tables that cover a callx be rejected?

> +
> +int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env)
> +{
> +	if (!env->cleanup_info_cnt)
> +		return 0;
> +
> +	if (bpf_prog_is_offloaded(env->prog->aux)) {
> +		verbose(env,
> +			"exception cleanup is not supported for offloaded programs\n");
> +		return -EINVAL;
> +	}
> +
> +	if (!bpf_jit_supports_cleanup_pads() || !env->prog->jit_requested) {
> +		verbose(env,
> +			"exception cleanup needs a JIT that can dispatch landing pads\n");
> +		return -EOPNOTSUPP;
> +	}
> +	env->prog->jit_required = 1;
> +
> +	if (env->exception_callback_subprog) {
> +		verbose(env,
> +			"exception cleanup table cannot be combined with an exception callback\n");
> +		return -EINVAL;
> +	}
> +
> +	mark_call_sites(env);
> +	return 0;
> +}

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time
  2026-09-26  5:01 ` [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
  2026-09-26  5:15   ` sashiko-bot
@ 2026-09-27 20:39   ` bot+bpf-ci
  2026-09-28  3:10     ` Yonghong Song
  1 sibling, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:39 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 6377 bytes --]

> diff --git arch/x86/net/bpf_jit_comp.c arch/x86/net/bpf_jit_comp.c
> index 6c7a0578760e..fb7e8ca1aab2 100644
> --- arch/x86/net/bpf_jit_comp.c
> +++ arch/x86/net/bpf_jit_comp.c
> @@ -2150,7 +2150,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
>  			dst_reg = X86_REG_R9;
>  	}
>
> -	if (bpf_insn_is_indirect_target(env, bpf_prog, i - 1))
> +	if (bpf_insn_is_indirect_target(env, bpf_prog, i - 1) ||
> +	    bpf_exc_insn_is_pad(env, bpf_prog, i - 1))
>  		EMIT_ENDBR();

Landing pads are only ever entered by a ret whose stack slot
bpf_unwind_rewrite() rewrote. Nothing jumps or calls to a pad indirectly
since bpf_unwind_resume() is lowered to 'r0 = 0; exit', and
bpf_exc_check_insn() forbids gotox inside a pad. IBT only checks
indirect JMP/CALL targets, not RET targets, so the ENDBR at each pad
head should not be needed.

The same patch sends pad-less frames to aux->epilogue_ip through the
same kind of rewritten ret, and emits no ENDBR there. That shows the pad
ENDBR does no functional work. Each extra ENDBR in the middle of a JITed
function is one more valid IBT landing site in executable kernel memory,
and the kernel normally keeps those to a minimum (objtool seals unused
ENDBRs).

Is there an indirect-branch path to the pad that is not visible here?

>
>  	ip = image + addrs[i - 1] + (prog - temp);

[ ... ]

> @@ -3278,6 +3279,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
>  		seen_exit = true;
>  		/* Update cleanup_addr */
>  		ctx->cleanup_addr = proglen;
> +		/* Where an unwind sends a frame with no pad. */
> +		bpf_prog->aux->epilogue_ip = (u64)image + proglen;

Eduard Zingerman raised an optimization suggestion in v5: use register
liveness information at landing pads to avoid spilling all registers
unnecessarily. The suggestion was to use register liveness information
at call sites and collect a union of registers alive at landing pads for
all callers of this particular function.

You acknowledged and said you would try to fix it to avoid saving
unnecessary callee saved registers. However, the current v6
implementation still appears to unconditionally spill registers for
exception_boundary programs without optimizing based on which registers
are actually live at landing pads.

Note: v6 significantly changed the implementation approach from v5, so
this suggestion may need re-evaluation in the context of the new design.
Does this optimization still apply, or does the new design address this
differently?

>  		if (bpf_prog_was_classic(bpf_prog) &&
>  		    !ns_capable_noaudit(&init_user_ns, CAP_SYS_ADMIN)) {
>  			if (emit_spectre_bhb_barrier(&prog, ip, bpf_prog))

[ ... ]

> @@ -4614,6 +4629,24 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
>  #endif
>  }
>
> +void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
> +			    void *cookie)
> +{
> +#if defined(CONFIG_UNWINDER_ORC)
> +	struct unwind_state state;
> +	unsigned long addr, *ra;
> +
> +	for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
> +	     unwind_next_frame(&state)) {
> +		addr = unwind_get_return_address(&state);
> +		ra = unwind_get_return_address_ptr(&state);
> +		if (!addr || !ra ||
> +		    !consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
> +			break;
> +	}
> +#endif
> +}

arch_bpf_stack_walk_ra() pairs the unwinder's recovered return address
with the raw stack slot. On x86 ORC, unwind_next_frame() does

arch/x86/kernel/unwind_orc.c:unwind_next_frame() {
    ...
    state->ip = unwind_recover_ret_addr(state, state->ip, (unsigned long *)ip_p);
    ...
}

which turns return_to_handler (function_graph / fprobe exit on fgraph)
or the rethook trampoline (kretprobe) back into the original caller
address. But unwind_get_return_address_ptr() returns `(unsigned long
*)state->sp - 1`, i.e. ip_p itself, and that slot still holds the
tracer's trampoline address. Nothing checks that *ra == addr before
handing ra to the consumer.

bpf_unwind() is a normal __bpf_kfunc in kernel/bpf/helpers.c with no
notrace, so function_graph, kretprobe:bpf_unwind, or a kprobe.multi
return probe can hook its return. In that case the prologue replaces the
slot holding bpf_unwind's return into the BPF program. Then:

  bpf_unwind() -> arch_bpf_stack_walk_ra() -> bpf_unwind_rewrite()
    cnt == 1 frame, and the bpf_unwind call is covered by a cleanup record
    (mark_call_sites() marks bpf_is_unwind_kfunc() calls)
    -> *ra = rec->pad;   /* overwrites return_to_handler / rethook trampoline */

bpf_unwind() then returns straight into the pad and skips the tracer's
return trampoline. Its fgraph ret_stack entry (or rethook node) is left
behind. The next hooked function further up that returns through the
trampoline pops that stale entry.

For fgraph, ftrace_pop_return_trace() has no frame-pointer check on x86
(HAVE_FUNCTION_GRAPH_FP_TEST is not defined there). For rethook,
__rethook_find_ret_addr() picks the top node. Either way execution jumps
to the old BPF address after the bpf_unwind call, running on a different
function's stack, which leads to an oops or arbitrary control flow.

The arm64 version of this same walker, added later in the series
(8bd2af30a28e "bpf, arm64: Dispatch exception cleanup pads at run
time"), handles exactly this case:

arch/arm64/net/bpf_jit_comp.c:arch_bpf_stack_walk_ra() {
    ...
    if (state->flags.fgraph || state->flags.kretprobe)
        return false;
    ...
}

Its changelog says: "a frame whose return the function graph tracer or a
kretprobe has hooked holds the tracer's trampoline in its slot rather
than the address the unwinder reports, so the walk stops there." The x86
walker has no such check, and no later commit in the series touches
arch/x86/net/bpf_jit_comp.c.

Should the x86 version also check state flags or compare
READ_ONCE_NOCHECK(*ra) against the value the unwinder read, and stop the
walk when they differ?

> +
>  void bpf_arch_poke_desc_update(struct bpf_jit_poke_descriptor *poke,
>  			       struct bpf_prog *new, struct bpf_prog *old)
>  {


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 20/21] selftests/bpf: Cover the exception cleanup shapes the chain does not reach
  2026-09-26  5:01 ` [PATCH bpf-next v6 20/21] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
@ 2026-09-27 20:40   ` bot+bpf-ci
  2026-09-28  3:49     ` Yonghong Song
  0 siblings, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:40 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 5135 bytes --]

> diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
> index d1d40314035e..49b6806cde82 100644
> --- a/tools/testing/selftests/bpf/exceptions_cleanup.h
> +++ b/tools/testing/selftests/bpf/exceptions_cleanup.h

[ ... ]

> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
> new file mode 100644
> index 000000000000..37a78e035a50
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
> @@ -0,0 +1,665 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
> +#include <vmlinux.h>
> +#include <bpf/bpf_helpers.h>
> +#include "bpf_misc.h"
> +#include "exceptions_cleanup.h"

[ ... ]

> +/* The callee most of the shapes below unwind out of. */
> +static __used __noinline __u64 pc_unwinder(__u64 x)
> +{
> +	if (x > 100)
> +		bpf_unwind();
> +	return x + 1;
> +}

[ ... ]

> +/* A callee on both a covered and an uncovered call: the pad is the site's. */
> +static __used __noinline __u64 shared_callee(__u64 x)
> +{
> +	if (x > 100)
> +		bpf_unwind();
> +	return x + 1;
> +}

This isn't a bug, but shared_callee() looks the same as pc_unwinder().
Could the shared-callee shape call pc_unwinder() from both sites, or
could wide_rec_frame() use pc_unwinder() so that shared_callee() is left
to the one shape its comment describes? And should the 'most of the
shapes' comment on pc_unwinder() be narrowed to match its three callers?

[ ... ]

> +/*
> + * A covered bpf_unwind() the sweep leaves last, behind the exception callback
> + * patchlet, which has to carry the marks with it.
> + */
> +SEC("?syscall")
> +__success __set_global(input, 101) __retval(0)
> +__ret_global(pads_ran, RAN_PAD_FIRST)
> +__naked int entry_pad_first(void)
> +{

This isn't a bug, but is 'exception callback patchlet' meant to be the
exit that bpf_exc_keep_exits() puts back after bpf_unwind()? If so,
could the comment name that instead? This program never sets up an
exception callback.

[ ... ]

> +/* A pad that indexes its frame by a register set before the unwinding call. */
> +
> +/* An unwinder that touches none of r6-r9, so the walk leaves this frame. */
> +static __used __naked __noinline __u64 var_unwinder(void)
> +{
> +	asm volatile (
> +	"if r1 < 101 goto 1f;"
> +	"call bpf_unwind;"
> +"1:"
> +	"r0 = 0;"
> +	"exit;"
> +	::: __clobber_all);
> +}
> +
> +static __used __naked __noinline __u64 var_stack_frame(void)
> +{
> +	asm volatile (
> +	"r1 = %[magic] ll;"
> +	"r1 = *(u64 *)(r1 + 0);"
> +	"*(u64 *)(r10 - 8) = r1;"	/* the slot the pad will read... */
> +	"*(u64 *)(r10 - 16) = r1;"	/* ...whichever of the two it is */
> +	"r1 = %[input] ll;"
> +	"r6 = *(u64 *)(r1 + 0);"
> +	"r6 &= 1;"			/* an unknown slot number... */
> +	"r6 <<= 3;"			/* ...as an aligned byte offset */
> +	"r1 = %[input] ll;"
> +	"r1 = *(u64 *)(r1 + 0);"
> +"1:"	"call var_unwinder;"		/* cleanup region */
> +"2:"
> +	"r0 = 0;"
> +	"exit;"
> +"3:"					/* landing pad */
> +	"r7 = r0;"
> +	"r1 = r10;"
> +	"r1 += r6;"			/* variable offset into the frame */
> +	"r2 = *(u64 *)(r1 - 16);"
> +	"r3 = %[magic] ll;"
> +	"r3 = *(u64 *)(r3 + 0);"
> +	"if r2 != r3 goto 9f;"
> +	PAD_RAN("%[ran]")
> +"9:"
> +	"r1 = r7;"
> +	"call bpf_unwind_resume;"
> +	"exit;"
> +	CLEANUP_REC("1b", "2b", "3b")
> +	:
> +	: [ran]"i"(RAN_VAR_STACK), __imm_addr(input), __imm_addr(magic),
> +	  __imm_addr(pads_ran)
> +	: __clobber_all);
> +}

[ ... ]

> +/* The same over a global subprogram, which the verifier enters no frame for. */
> +__noinline __u64 global_unwinder(__u64 x)
> +{
> +	if (x > 100)
> +		bpf_unwind();
> +	return x + 1;
> +}
> +
> +static __used __naked __noinline __u64 global_pad_frame(void)
> +{
> +	asm volatile (
> +	"r1 = %[magic] ll;"
> +	"r1 = *(u64 *)(r1 + 0);"
> +	"*(u64 *)(r10 - 8) = r1;"
> +	"*(u64 *)(r10 - 16) = r1;"
> +	"r1 = %[input] ll;"
> +	"r6 = *(u64 *)(r1 + 0);"
> +	"r6 &= 1;"
> +	"r6 <<= 3;"
> +	"r1 = %[input] ll;"
> +	"r1 = *(u64 *)(r1 + 0);"
> +"1:"	"call global_unwinder;"		/* cleanup region */
> +"2:"
> +	"r0 = 0;"
> +	"exit;"
> +"3:"					/* landing pad */
> +	"r7 = r0;"
> +	"r1 = r10;"
> +	"r1 += r6;"
> +	"r2 = *(u64 *)(r1 - 16);"
> +	"r3 = %[magic] ll;"
> +	"r3 = *(u64 *)(r3 + 0);"
> +	"if r2 != r3 goto 9f;"
> +	PAD_RAN("%[ran]")
> +"9:"
> +	"r1 = r7;"
> +	"call bpf_unwind_resume;"
> +	"exit;"
> +	CLEANUP_REC("1b", "2b", "3b")
> +	:
> +	: [ran]"i"(RAN_GLOBAL_PAD), __imm_addr(input), __imm_addr(magic),
> +	  __imm_addr(pads_ran)
> +	: __clobber_all);
> +}

This isn't a bug, but global_pad_frame() looks nearly identical to
var_stack_frame(), and pad_r0_frame() repeats the same pad body. Could
these share a macro, the way LOAD_MAGIC_REGS/CHECK_MAGIC_REGS are shared
earlier in the file, so each variant only spells out the part it is
testing?

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad
  2026-09-26  5:00 ` [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad Yonghong Song
  2026-09-26  5:15   ` sashiko-bot
@ 2026-09-27 20:40   ` bot+bpf-ci
  2026-09-28  0:17     ` Yonghong Song
  1 sibling, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:40 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 6929 bytes --]

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 6d78c20e6507..0143688896b0 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -836,6 +836,7 @@ struct bpf_subprog_info {
>  	s16 fastcall_stack_off;
>  	bool has_tail_call: 1;
>  	bool might_throw: 1;
> +	bool might_unwind: 1;
>  	bool tail_call_reachable: 1;
>  	bool has_ld_abs: 1;
>  	bool is_cb: 1;

[ ... ]

> diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
> index 08aee86a155c..f291611fe578 100644
> --- a/kernel/bpf/helpers.c
> +++ b/kernel/bpf/helpers.c
> @@ -5095,6 +5095,8 @@ BTF_ID_FLAGS(func, bpf_task_from_pid, KF_ACQUIRE | KF_RET_NULL)
>  BTF_ID_FLAGS(func, bpf_task_from_vpid, KF_ACQUIRE | KF_RET_NULL)
>  BTF_ID_FLAGS(func, bpf_throw)
> +BTF_ID_FLAGS(func, bpf_unwind)
> +BTF_ID_FLAGS(func, bpf_unwind_resume)
>  #ifdef CONFIG_BPF_EVENTS
>  BTF_ID_FLAGS(func, bpf_send_signal_task)
>  #endif

Can a CAP_BPF program call bpf_unwind() at this commit? The kfuncs are
registered in generic_btf_ids, which makes them resolvable, and the body
at kernel/bpf/helpers.c is still __bpf_kfunc void bpf_unwind(void) { },
so the call returns normally to the next instruction.

However, the verifier assumes bpf_unwind() never falls through:

kernel/bpf/verifier.c:process_bpf_unwind() ends the path with
PROCESS_BPF_EXIT for an uncovered call, and jumps to the pad for a
covered call. liveness.c keeps a fall-through edge, but do_check()
never explores it. As a result, any instructions reachable only from
after the call are skipped by the verifier and removed by
bpf_opt_remove_dead_code().

At run time the call returns and execution continues into whatever the
JIT placed next, with r0-r5 clobbered, or off the end of the image.

This is fixed later in the series by 7924e9e105cf "bpf: Dispatch cleanup
pads by rewriting return addresses", which implements the real body via
arch_bpf_stack_walk_ra()/bpf_unwind_rewrite(), and bpf_exc_keep_exits()
puts 'r0 = 0; exit' after every bpf_unwind call.

[ ... ]

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index fc3df452de2e..77176250f866 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
> @@ -19167,6 +19167,40 @@ enum {
>  	INSN_IDX_UPDATED = 2,
>  };
>
> +static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
> +{
> +	struct bpf_verifier_state *branch;
> +	struct bpf_func_state *frame;
> +	int pad = bpf_exc_pad_of_call(env, insn_idx);
> +
> +	if (pad < 0)
> +		return 0;
> +	branch = push_stack(env, pad, insn_idx, false);
> +	if (IS_ERR(branch))
> +		return PTR_ERR(branch);
> +	frame = branch->frame[branch->curframe];
> +	/*
> +	 * The state at that call with the caller-saved registers gone: the
> +	 * callee's epilogue put r6-r9 and the stack back on the way out.
> +	 */
> +	clear_caller_saved_regs(env, frame->regs);
> +	mark_reg_unknown(env, frame->regs, BPF_REG_0);
> +	return 0;
> +}
> +
> +static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx)
> +{
> +	struct bpf_func_state *frame = cur_func(env);
> +	int pad = bpf_exc_pad_of_call(env, *insn_idx);
> +
> +	if (pad < 0)
> +		return PROCESS_BPF_EXIT;
> +	clear_caller_saved_regs(env, frame->regs);
> +	mark_reg_unknown(env, frame->regs, BPF_REG_0);
> +	*insn_idx = pad;
> +	return INSN_IDX_UPDATED;
> +}
> +
>  static int process_bpf_exit_full(struct bpf_verifier_env *env,
>  				 bool *do_print_state,
>  				 bool exception_exit)
> @@ -19404,6 +19438,20 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>
>  		env->jmps_processed++;
>  		if (opcode == BPF_CALL) {
> +			if (bpf_is_unwind_kfunc(insn))
> +				return process_bpf_unwind(env, &env->insn_idx);
> +			if (bpf_is_unwind_resume_kfunc(insn)) {
> +				/*
> +				 * Mark r0 a known zero -- unknown first, as
> +				 * the known-zero helper keeps the type it
> +				 * finds, which here is NOT_INIT. The fixups
> +				 * lower this to 'r0 = 0; exit', so the frame
> +				 * returns a real zero.
> +				 */
> +				mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
> +				mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
> +				return process_bpf_exit_full(env, do_print_state, false);
> +			}
>  			if (env->cur_state->active_locks) {
>  				/* similar to static subprog calls callx is allowed under a lock */
>  				if (!bpf_is_callx(insn) &&

Both new kfuncs are handled before check_kfunc_call(), so
bpf_fetch_kfunc_arg_meta() never runs for them. That function rejects a
kfunc with -EACCES "calling kernel function %s is not allowed" when
btf_kfunc_is_allowed() fails for the program type or the kfunc has no
flags. bpf_add_kfunc_call() only resolves the BTF id and does not apply
that filter.

As a result, bpf_unwind() and bpf_unwind_resume() are callable from any
program type that can make kfunc calls, including types that
generic_kfunc_set is not registered for (it is registered for TRACING,
SCHED_CLS, XDP, STRUCT_OPS, SYSCALL and CGROUP_SKB).

It also skips the usual sleepable, RCU, preempt and IRQ-context checks
in check_kfunc_call().

---

At this commit a bpf_unwind_resume() call is verified as a frame exit
with r0 = 0, but the comment says "The fixups lower this to
'r0 = 0; exit'". Does that lowering exist yet? No fixup for it exists in
kernel/bpf/fixups.c at this commit.

This call is also accepted anywhere, not only inside a landing pad. So a
CAP_BPF program can place 'call bpf_unwind_resume' in its main path. The
verifier treats the path as ended and never looks at the following
instructions, which the dead-code pass then removes or sanitizes.

At run time the real kfunc body runs WARN_ONCE(1, "exception cleanup
resume was not lowered to a return\n"). That is a user-triggerable
warning, and a panic with panic_on_warn. The call then returns normally
with r0-r5 clobbered and falls into instructions that were never
verified for this state, or off the end of the JIT image.

This is fixed later in the series by 7924e9e105cf "bpf: Dispatch cleanup
pads by rewriting return addresses", which adds the bpf_do_misc_fixups
lowering, and 2a1e4ed1e9ae "bpf: Refuse a landing pad that does not
resume", which rejects "resume at insn %d is not in a landing pad".

> @@ -19422,6 +19470,10 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>  				}
>  			}
>  			mark_reg_scratched(env, BPF_REG_0);
> +			/* An unwind out of this call resumes at the pad. */
> +			err = push_cleanup_pad_branch(env, env->insn_idx);
> +			if (err)
> +				return err;
>  			if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
>  				cur_func(env)->no_stack_arg_load = true;
>  			if (bpf_is_callx(insn))

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG
  2026-09-26  5:00 ` [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG Yonghong Song
  2026-09-26  5:21   ` sashiko-bot
@ 2026-09-27 20:40   ` bot+bpf-ci
  2026-09-28  0:12     ` Yonghong Song
  1 sibling, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:40 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 6862 bytes --]

> diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
> index b0bd9ba951df..4e2b6985bc96 100644
> --- a/kernel/bpf/cfg.c
> +++ b/kernel/bpf/cfg.c
> @@ -6,6 +6,7 @@
>  #include <linux/sort.h>
>
>  #include "diagnostics.h"
> +#include "exception.h"
>
>  #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
>

[ ... ]

> @@ -160,17 +161,64 @@ static int push_insn(int t, int w, int e, struct bpf_verifier_env *env)
>  	return DONE_EXPLORING;
>  }
>
> +static int visit_cleanup_pad_edge(int t, struct bpf_verifier_env *env)
> +{
> +	int *insn_stack = env->cfg.insn_stack;
> +	int *insn_state = env->cfg.insn_state;
> +	int w;
> +
> +	if (!env->cleanup_info_cnt)
> +		return DONE_EXPLORING;
> +	w = bpf_exc_pad_of_call(env, t);
> +	if (w < 0)
> +		return DONE_EXPLORING;
> +
> +	/*
> +	 * @t is a call that may branch here, and @w is the target of that
> +	 * branch, so both are prune points. @w especially: every covered call
> +	 * site in a region unwinds to the same pad, and without a prune point
> +	 * at its head the verifier walks the pad again for each of them.
> +	 */
> +	mark_prune_point(env, t);
> +	mark_prune_point(env, w);
> +	mark_jmp_point(env, w);
> +	mark_jump_target(env, w);
> +
> +	if (insn_state[w])
> +		return DONE_EXPLORING;
> +	if (env->cfg.cur_stack >= env->prog->len)
> +		return -E2BIG;
> +	insn_stack[env->cfg.cur_stack++] = w;
> +	insn_state[w] |= DISCOVERED;
> +	return KEEP_EXPLORING;
> +}
> +
> +static int merge_visit_ret(int a, int b)
> +{
> +	if (a < 0)
> +		return a;
> +	if (b < 0)
> +		return b;
> +	if (a == KEEP_EXPLORING || b == KEEP_EXPLORING)
> +		return KEEP_EXPLORING;
> +	return DONE_EXPLORING;
> +}
> +
>  static int visit_func_call_insn(int t, struct bpf_insn *insns,
>  				struct bpf_verifier_env *env,
>  				bool visit_callee)
>  {
> -	int ret, insn_sz;
> +	int ret, insn_sz, pad_ret;
>  	int w;
>
> +	pad_ret = visit_cleanup_pad_edge(t, env);
> +	if (pad_ret < 0)
> +		return pad_ret;

The error handling here appears to be incorrect. In the v6 discussion
thread, you identified that this should be:

	if (pad_ret != DONE_EXPLORING)
		return pad_ret;

rather than checking for negative values. The current code causes premature
DFS stop by ignoring KEEP_EXPLORING return values from
visit_cleanup_pad_edge(). When the cleanup pad edge returns KEEP_EXPLORING,
the DFS should continue, but the negative-only check misses this case and
treats it as DONE_EXPLORING.

Can this be changed to properly handle both KEEP_EXPLORING and error cases?

>
>  	insn_sz = bpf_is_ldimm64(&insns[t]) ? 2 : 1;
>  	ret = push_insn(t, t + insn_sz, FALLTHROUGH, env);
>  	if (ret)
> -		return ret;
> +		return merge_visit_ret(pad_ret, ret);
>
>  	mark_prune_point(env, t + insn_sz);
>  	/* when we exit from subprog, we need to record non-linear history */
> @@ -182,7 +230,7 @@ static int visit_func_call_insn(int t, struct bpf_insn *insns,
>  		merge_callee_effects(env, t, w);
>  		ret = push_insn(t, w, BRANCH, env);
>  	}
> -	return ret;
> +	return merge_visit_ret(pad_ret, ret);
>  }

[ ... ]

> diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
> index cd9523f69298..4e0273a8ceee 100644
> --- a/kernel/bpf/liveness.c
> +++ b/kernel/bpf/liveness.c
> @@ -8,6 +8,8 @@
>  #include <linux/slab.h>
>  #include <linux/sort.h>
>
> +#include "exception.h"
> +
>  #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)

[ ... ]

> @@ -384,6 +386,18 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
>  			succ->items[succ->cnt++] = exit_idx;
>  	}
>
> +	/*
> +	 * A call a cleanup record covers can leave through its landing pad.
> +	 * Only a call to a subprogram or to bpf_unwind() is marked, neither of
> +	 * which is an edge the block above adds, so succ still holds two.
> +	 */
> +	if (unlikely(env->cleanup_info_cnt)) {
> +		int pad = bpf_exc_pad_of_call(env, idx);
> +
> +		if (pad >= 0)
> +			succ->items[succ->cnt++] = pad;
> +	}
> +
>  	return succ;
>  }

[ ... ]

> @@ -545,6 +559,13 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
>  		alive = callee_stack_access_at_callsite(env, callsite)
>  			? is_live_before(instance, callsite, rel, half_spi)
>  			: is_live_before(instance, callsite + 1, rel, half_spi);
> +
> +		if (!alive && unlikely(env->cleanup_info_cnt)) {
> +			int pad = bpf_exc_pad_of_call(env, callsite);
> +
> +			if (pad >= 0)
> +				alive = is_live_before(instance, pad, rel, half_spi);
> +		}

Eduard Zingerman asked in v5 why this landing pad liveness check is
necessary. He stated: "If bpf_insn_succ correctly traverses call successors
including the landing pad the information should already be accounted for
by the is_live_before(instance, callsite, rel, half_spi)."

This question was not addressed in subsequent versions. Looking at the
patch series design, there's a question about which model is intended.

The commit message justifies this hunk with: "otherwise
clean_verifier_state() poisons the slot while the callee runs and the pad
is rejected for reading it."

However, at this series tip, the verifier never enters a landing pad from a
state where the pad's frame is an outer frame. There are only two entry
paths:

1. push_cleanup_pad_branch() (added in patch 5f426cf2a1ff) pushes a copy of
   the caller's state before check_func_call(), with the pad in the current
   frame.

2. process_bpf_unwind() jumps to a pad only in the current frame; otherwise
   it returns PROCESS_BPF_EXIT and bpf_unwind_resume goes through
   process_bpf_exit_full() to callsite + 1.

Since clean_verifier_state() is called only on cur in is_state_visited(),
poisoning a caller slot in cur while the callee runs cannot affect the
pushed pad copy. That copy was taken at the call site, where the caller is
the current frame, and bpf_insn_successors() already includes the pad
there.

As written, the outer-frame pad check only keeps extra caller slots alive
in callee states, which reduces pruning. It does not prevent any pad
rejection.

Either the rationale is inaccurate and the hunk is unnecessary, or the
intended model is that the pad sees the caller's frame as the callee left
it. The second reading is closer to runtime behaviour: a callee can write
the caller's stack through a PTR_TO_STACK argument before bpf_unwind(). But
that is not what push_cleanup_pad_branch() models.

Which model is intended? If the pre-call copy is the design, could this
hunk and the changelog paragraph be dropped? If the post-callee state is
the design, does the pad verification in the later verifier patch need to
change instead?

>  		if (alive)
>  			return true;
>  	}


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 12/21] bpf, arm64: Dispatch exception cleanup pads at run time
  2026-09-26  5:01 ` [PATCH bpf-next v6 12/21] bpf, arm64: " Yonghong Song
  2026-09-26  5:14   ` sashiko-bot
@ 2026-09-27 20:40   ` bot+bpf-ci
  1 sibling, 0 replies; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:40 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 7282 bytes --]

> diff --git a/arch/arm64/kernel/stacktrace.c b/arch/arm64/kernel/stacktrace.c
> index 3ebcf8c53fb0..c750c520f24d 100644
> --- a/arch/arm64/kernel/stacktrace.c
> +++ b/arch/arm64/kernel/stacktrace.c

[ ... ]

> +static bool bpf_unwind_store_ra(unsigned long record, unsigned long caller_fp,
> +				u64 pc, u64 ra)
> +{
> +	struct frame_record *rec = (struct frame_record *)record;
> +	u64 stored = READ_ONCE(rec->lr);
> +
> +	if (system_supports_address_auth() && stored != pc) {
> +		u64 modifier;
> +
> +		if (WARN_ON_ONCE(!bpf_unwind_ra_modifier(record, caller_fp,
> +							 stored, pc, &modifier)))
> +			return false;
> +		ra = bpf_unwind_sign_ra(ra, modifier);
> +	}
> +	WRITE_ONCE(rec->lr, ra);
> +	return true;
> +}

Does this redirect work under CONFIG_SHADOW_CALL_STACK?

When shadow call stack is enabled, a C function epilogue reloads x30 from
the shadow stack rather than from the frame record. For the BPF frame
that called bpf_unwind(), the frame record belongs to bpf_unwind()'s own
C frame:

arch_bpf_stack_walk_ra() is noinline, so kunwind_init_from_caller() sets
fp = __builtin_frame_address(1) inside it.

This means the WRITE_ONCE(rec->lr, ra) above stores the pad address into
a slot that bpf_unwind() never reads back when returning. The sequence
would be:

1. bpf_unwind_rewrite() finds a cleanup record (ctx->cnt == 1) and sets
   *ra = rec->pad
2. The WRITE_ONCE() stores the pad into the frame record
3. bpf_unwind() returns by reading x30 from the shadow stack via x18,
   not from the stacked lr
4. The program returns to the instruction after the call, which is the
   'r0 = 0; exit' that bpf_exc_keep_exits() added
5. The pad never runs

For a concrete example, in the exceptions_cleanup.c selftest, foo3()
calls bpf_preempt_disable(), then a covered bpf_unwind(), and its pad
calls bpf_preempt_enable(). On a kernel with CONFIG_SHADOW_CALL_STACK
enabled, the pad is skipped and the program returns with preemption still
disabled.

The same issue affects any reference acquired before a covered
bpf_unwind() and released in its pad.

Shadow call stack is enabled in configurations like Android GKI builds
(DYNAMIC_SCS on cores without PAC) and other hardened kernels. The
Kconfig help for SHADOW_CALL_STACK says the shadow stack exists "to
protect function return addresses from being overwritten", and it
requires "depends on DYNAMIC_FTRACE_WITH_ARGS ||
DYNAMIC_FTRACE_WITH_REGS || !FUNCTION_GRAPH_TRACER" which prevents the
graph tracer from rewriting stacked return addresses.

Nothing in this series tests SHADOW_CALL_STACK, and
bpf_jit_supports_cleanup_pads() returns true unconditionally.

Should the feature be disabled while SCS is active, or should both the
frame record and the shadow stack slot be updated?

> +static bool
> +arch_bpf_unwind_ra_consume_entry(const struct kunwind_state *state, void *cookie)
> +{
> +	struct bpf_unwind_ra_consume_entry_data *data = cookie;
> +	unsigned long record = data->record;
> +	bool seen_first = data->seen_first;
> +	u64 ra = state->common.pc;
> +	bool cont;
> +
> +	/* The record this frame's return address will have come out of. */
> +	data->record = state->common.fp;
> +	data->seen_first = true;
> +
> +	/* The first pc is where the walk runs, not an address it returns to. */
> +	if (!seen_first)
> +		return true;
> +	/* A traced return: the slot holds the tracer's trampoline, not @pc. */
> +	if (state->flags.fgraph || state->flags.kretprobe)
> +		return false;

Can cleanup pads work when bpf_unwind() itself is traced?

bpf_unwind() is traceable: __bpf_kfunc is "__used __retain __noclone
noinline" with no notrace, kernel/bpf/Makefile does not strip
CC_FLAGS_FTRACE from helpers.o, and bpf_unwind is not marked
NOKPROBE_SYMBOL.

When the function graph tracer is active, or a kretprobe or fprobe exit
handler is attached to bpf_unwind, the frame record holds return_to_handler
or __kretprobe_trampoline instead of the actual return address.

kunwind_recover_return_address() then sets flags.fgraph or
flags.kretprobe on the entry for the BPF frame that called bpf_unwind().
This is the first entry the callback acts on after the initial walk
frame, so it returns false immediately:

> +	if (state->flags.fgraph || state->flags.kretprobe)
> +		return false;

This happens before calling consume_entry, so the walk ends having
redirected nothing.

arch_bpf_stack_walk_ra() returns void, so bpf_unwind() cannot tell the
unwind did not happen. The consequences are the same as the SCS case:

1. bpf_unwind() returns through the tracer into the 'r0 = 0; exit' that
   the verifier never explored
2. The outer frames resume at the instruction after their calls instead
   of at their pads
3. Cleanup code in all pads is skipped

For foo3() in exceptions_cleanup.c with bpf_preempt_disable() before
the covered bpf_unwind() and bpf_preempt_enable() in the pad, the
program returns with preemption disabled whenever a privileged user
traces bpf_unwind. References released only in pads leak.

The commit message says the walk "stops there" for such frames but does
not address what happens when all cleanup is skipped.

Should bpf_unwind() be marked notrace and NOKPROBE_SYMBOL, or should
the implementation handle traced returns differently?

[ ... ]

> diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
> index 475e70653454..2422a1ae1256 100644
> --- a/arch/arm64/net/bpf_jit_comp.c
> +++ b/arch/arm64/net/bpf_jit_comp.c

[ ... ]

> @@ -2423,6 +2425,17 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
>  		 * reasons, expects to point to the next instruction)
>  		 */
>  		bpf_prog_update_insn_ptrs(prog, ctx.offset, ctx.ro_image);
> +
> +		/*
> +		 * Same byte offsets, consumed by the bpf_unwind() walk:
> +		 * turn the cleanup records into native address ranges now that
> +		 * the image is final.
> +		 */
> +		bpf_exc_fill_native_ranges(prog, ctx.offset, ctx.ro_image);
> +
> +		/* Where an unwind sends a frame with no pad. */
> +		prog->aux->epilogue_ip = (u64)ctx.ro_image +
> +					 ctx.epilogue_offset * AARCH64_INSN_SIZE;

The commit message says:

> The JIT half: build the native cleanup table from the JIT's byte offsets
> once the image is final, emit a BTI at each pad head, address the frame
> through the private stack pointer where one is in use, and record the one
> epilogue so a frame the unwind passes over can return straight through it.

Does the sentence "address the frame through the private stack pointer
where one is in use" describe code in this commit?

The diff contains no changes touching priv_sp_used, PRIVATE_SP, or
priv_stack_ptr in arch/arm64/net/bpf_jit_comp.c.

Earlier in this series, commit 88ff94dddf57 ("bpf: Refuse a private
stack for a program with an exception cleanup table") forces
NO_PRIV_STACK in check_max_stack_depth() whenever env->cleanup_info_cnt
is set, so a private stack is never in use when this code runs.

Should that sentence be removed from the changelog since the behaviour is
neither implemented nor reachable?

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 10/21] bpf: Dispatch cleanup pads by rewriting return addresses
  2026-09-26  5:00 ` [PATCH bpf-next v6 10/21] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
@ 2026-09-27 20:40   ` bot+bpf-ci
  2026-09-28  1:08     ` Yonghong Song
  0 siblings, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:40 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 11651 bytes --]

> diff --git a/include/linux/bpf.h b/include/linux/bpf.h
> index 4bae3796c42f..98de251037df 100644
> --- a/include/linux/bpf.h
> +++ b/include/linux/bpf.h

[ ... ]

> @@ -1885,6 +1928,8 @@ struct bpf_prog_aux {
>  	u64 (*bpf_exception_cb)(u64 cookie, u64 sp, u64 bp, u64, u64);
>  	u16 stack_arg_sp_adjust;
>  	u16 freplace_link_cnt; /* counts freplace links extending this prog */
> +	struct bpf_exception_info *exc;
> +	u64 epilogue_ip; /* native address of this (sub)program's epilogue */
>  #ifdef CONFIG_SECURITY
>  	void *security;
>  #endif

The commit message says epilogue_ip "is recorded for every program", but
can a caller that only calls subprogs which always unwind end up with
epilogue_ip still zero?

Looking at bpf_exc_keep_exits(), it only patches an exit after
bpf_unwind() kfunc calls:

kernel/bpf/fixups.c:bpf_exc_keep_exits() {
    for (i = 0; i < insn_cnt; i++) {
        if (!bpf_is_unwind_kfunc(insn))
            continue;
        /* patch exit after call */
    }
}

It doesn't patch exits after calls to subprogs that can only unwind.
The verifier ends that path at the subprog's bpf_unwind()
(process_bpf_unwind() returns PROCESS_BPF_EXIT), so code after the call
is never seen and gets swept. If that was the caller's only exit, it
never emits a BPF_EXIT, so the JIT never sets epilogue_ip.

At run time bpf_unwind_rewrite() would then reach:

    } else if (prog->aux->epilogue_ip) {
        *ra = prog->aux->epilogue_ip;
    } else {
        WARN_ON_ONCE(1);
        return false;
    }

with epilogue_ip still 0. The return address stays pointing just past the
call, which is the end of the JIT image.

Example: main calls a static subprog that only has `call bpf_unwind; exit`.
With no cleanup table (or no record covering either call), the main frame
would be redirected to epilogue_ip, but that field was never set.

[ ... ]

> diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
> index d813fdde29e3..60905643cb9c 100644
> --- a/kernel/bpf/core.c
> +++ b/kernel/bpf/core.c

[ ... ]

> @@ -3511,6 +3523,12 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
>  {
>  }
>
> +void __weak arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp,
> +						      u64 *ra),
> +				   void *cookie)
> +{
> +}
> +
>  bool __weak bpf_jit_supports_cleanup_pads(void)
>  {
>  	return false;

Does bpf_unwind() produce correct behavior on architectures that keep this
weak stub?

At this commit no architecture implements the walker. At the end of the
series only x86-64 with CONFIG_UNWINDER_ORC and arm64 do. x86 with the
frame-pointer unwinder, riscv, s390, powerpc and loongarch keep the stub.

Nothing stops those architectures from using bpf_unwind():

- It is in generic_btf_ids, registered for TRACING, SCHED_CLS, XDP,
  STRUCT_OPS, SYSCALL, CGROUP_SKB and others.

- bpf_prepare_cleanup_exceptions() only checks bpf_jit_supports_cleanup_pads()
  when a table was supplied:

      if (!env->cleanup_info_cnt)
          return 0;

- An uncovered call is a terminator for the verifier:

      if (pad < 0)
          return PROCESS_BPF_EXIT;

At run time the weak stub does nothing, so bpf_unwind() returns. The frame
returns through the patched 'r0 = 0; exit'. Each caller then resumes at
the instruction after its call, in a state the verifier never explored for
that path, or in code the dead code sweep already removed.

Concrete case: a static subprog that returns a checked non-NULL map value
pointer on its normal path and calls bpf_unwind() on another path. The
caller dereferences the returned pointer without a NULL check, since the
verifier only saw the pointer-returning exit. At run time the subprog
returns 0 and the caller writes through NULL in kernel context.

The commit message says the separate entry point exists "so that the
architectures which do not dispatch pads keep the walker they have."
Does bpf_unwind() need to be gated on bpf_jit_supports_cleanup_pads(), or
should the stub have a fallback that doesn't leave callers in unverified
code?

> diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
> index f291611fe578..ea5c4d81f8b7 100644
> --- a/kernel/bpf/helpers.c
> +++ b/kernel/bpf/helpers.c

[ ... ]

> @@ -3424,8 +3426,51 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
>  	return false;
>  }
>
> +struct bpf_unwind_ctx {
> +	u32 cnt;
> +};
> +
> +static bool bpf_unwind_rewrite(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra)
> +{
> +	const struct bpf_cleanup_range *rec;
> +	struct bpf_unwind_ctx *ctx = cookie;
> +	struct bpf_exception_info *exc;
> +	struct bpf_prog *prog;
> +
> +	rcu_read_lock();
> +	prog = bpf_prog_ksym_find(ip);
> +	rcu_read_unlock();
> +	if (!prog)
> +		return !ctx->cnt;
> +	ctx->cnt++;
> +
> +	exc = prog->aux->exc;
> +	rec = (exc && exc->nr_ranges) ? bpf_exc_pad_for_ip(prog, ip) : NULL;
> +	if (rec) {
> +		*ra = rec->pad;
> +	} else if (ctx->cnt == 1) {
> +		/*
> +		 * The frame that called bpf_unwind(). Its return address
> +		 * always names the 'r0 = 0; exit' that bpf_exc_keep_exits()
> +		 * put after the call, so leave it alone and let the frame
> +		 * return through that: running it is what sets the value
> +		 * the unwind returns.
> +		 */
> +	} else if (prog->aux->epilogue_ip) {
> +		*ra = prog->aux->epilogue_ip;

Does sending a frame to epilogue_ip work correctly when the verifier
treats the resume as returning to the instruction after the call?

At run time, a caller frame whose call site no cleanup record covers is
sent to aux->epilogue_ip, so it returns straight away. The verifier does
not model that: it treats a pad's bpf_unwind_resume() as an ordinary return
into the caller, at the instruction after the call. The code after the call
is therefore verified but never runs, and the return-at-once path is never
checked.

kernel/bpf/verifier.c:do_check_insn() with curframe > 0:

    cur_func(env)->in_pad = false;
    return process_bpf_exit_full(env, do_print_state, false);

process_bpf_exit_full(..., false) goes to prepare_func_exit(), which
continues the caller at callsite + 1 with r0 known to be zero. Only a
covered call site gets its pad pushed as another branch
(push_cleanup_pad_branch()). Nothing requires a call to a might_unwind
static subprog to be covered, and bpf_exc_check_insn() only restricts global
subprogs while an unwind is in flight.

Concrete case (x86-64 with ORC, or arm64, at the end of the series):

    main:  t = bpf_task_acquire(p); if (!t) return 0;
           sub();                 /* call site not covered */
           bpf_task_release(t); return 0;
    sub:   bpf_unwind();          /* covered by a record in sub */
           pad: bpf_unwind_resume(0);

The verifier accepts this because it walks main past the call and sees the
release. At run time sub's pad runs, resumes, and returns into main's
epilogue, so the task reference is leaked.

The same shape with bpf_rcu_read_lock()/unlock(), preempt disable/enable,
or bpf_spin_lock()/unlock() leaves the lock held when the program returns.
Static subprog calls are allowed under a spin lock, and the bpf_unwind
check comes before the active_locks check. bpf_spin_lock is taken with IRQs
disabled, so the next locker deadlocks.

Does the verifier need to model the epilogue_ip redirect for uncovered
callers, or should uncovered calls be refused when the callee might unwind?

---

On x86-64 with a private stack, does this redirect skip the pop r9 that
follows the call?

With a private stack, the JIT wraps every call in push_r9/pop_r9
(arch/x86/net/bpf_jit_comp.c). When bpf_unwind_rewrite() sets
`*ra = prog->aux->epilogue_ip` for a caller frame, the callee's ret lands
on the epilogue with the pushed r9 still on top of the stack. The epilogue
is `add rsp, outgoing_rsp; pop_callee_regs(); leave; ret`.
pop_callee_regs() therefore pops the saved r9 into the first callee-saved
register and shifts the rest by one slot. leave then fixes rsp, so the
function returns normally, but with the kernel caller's rbx/r13/r14/r15
corrupted.

check_max_stack_depth() refuses a private stack only when a table is
present:

    if (env->cleanup_info_cnt)
        priv_stack_mode = NO_PRIV_STACK;

Its own comment describes this exact problem for pads ("on x86-64 that
skips the pop which restores r9"). A table-less program can still reach the
redirect.

Example: an fentry or kprobe program whose main prog uses at least
BPF_PRIV_STACK_MIN_SIZE bytes of stack and holds a value in r6 across a
call to a static subprog that calls bpf_unwind(). This applies on x86-64
with ORC after the x86 enablement patch. Should the refusal cover
might_unwind programs, not only programs with a table?

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 677bab92f624..f2a59716fdab 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
> @@ -22781,6 +22781,9 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
>  		/* program is valid, convert *(u32*)(ctx + off) accesses */
>  		ret = bpf_convert_ctx_accesses(env);
>
> +	if (ret == 0)
> +		ret = bpf_exc_keep_exits(env);

Do the exits this patch adds skip a struct_ops program's BPF-level
epilogue?

bpf_convert_ctx_accesses() is the only place that epilogue is inserted:
every BPF_EXIT in the main prog is replaced with the epilogue or a jump to
it (fixups.c, "Generate epilogue for the main prog"). But
bpf_exc_keep_exits() runs after that pass and emits a plain
`r0 = 0; exit` after every bpf_unwind() call. bpf_do_misc_fixups() runs
later still and lowers each bpf_unwind_resume() to a plain `r0 = 0; exit`
too. In the main prog, both of these exits go straight to the JIT's native
epilogue. So does the `*ra = prog->aux->epilogue_ip` redirect in
bpf_unwind_rewrite().

Does this break bpf_qdisc? bpf_qdisc_gen_epilogue()
(net/sched/bpf_qdisc.c) adds a call to bpf_qdisc_reset_destroy_epilogue()
to every .reset and .destroy program. That call is the only thing that runs
qdisc_watchdog_cancel(&q->watchdog), and bpf_qdisc_validate() requires
.reset and .destroy to be BPF programs for exactly that reason.
bpf_unwind() is in generic_btf_ids, which is registered for
BPF_PROG_TYPE_STRUCT_OPS with no filter. With no cleanup table,
process_bpf_unwind() simply returns PROCESS_BPF_EXIT, so no gating applies.

Failure path:

1. A .destroy program runs `if (cond) bpf_unwind();`.

2. On any arch, bpf_unwind() either leaves the first frame's return address
   alone (bpf_unwind_rewrite(), cnt == 1) or does nothing (the weak
   arch_bpf_stack_walk_ra()). Either way the frame returns through the
   `r0 = 0; exit` that bpf_exc_keep_exits() inserted, which is not the
   gen_epilogue.

3. qdisc_watchdog_cancel() never runs, so an hrtimer that .enqueue armed
   through bpf_qdisc_watchdog_schedule() stays armed inside
   qdisc_priv(sch).

4. The qdisc is freed, the timer fires, and qdisc_watchdog()
   (net/sched/sch_api.c) dereferences wd->qdisc in freed memory. The
   hrtimer base also still links the freed timer.

The same bypass happens when a .reset/.destroy main prog resumes from a
landing pad (lowered resume), or when a covered or uncovered subprog unwinds
and the main frame is sent to epilogue_ip.

Should bpf_unwind()/cleanup tables be rejected for programs whose
verifier_ops has a gen_epilogue, or should these exits be inserted before
bpf_convert_ctx_accesses() and point epilogue_ip at the BPF-level epilogue?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume
  2026-09-26  5:00 ` [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume Yonghong Song
  2026-09-26  5:17   ` sashiko-bot
@ 2026-09-27 20:40   ` bot+bpf-ci
  2026-09-28  0:29     ` Yonghong Song
  1 sibling, 1 reply; 56+ messages in thread
From: bot+bpf-ci @ 2026-09-27 20:40 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 8590 bytes --]

> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
> index 0143688896b0..4174c7d0177e 100644
> --- a/include/linux/bpf_verifier.h
> +++ b/include/linux/bpf_verifier.h
> @@ -339,6 +339,8 @@ struct bpf_func_state {
>  	bool in_async_callback_fn;
>  	bool in_exception_callback_fn;
>  	bool no_stack_arg_load;
> +	/* an unwind reached this frame and its landing pad is running */
> +	bool in_pad;
>  	/* For callback calling functions that limit number of possible
>  	 * callback executions (e.g. bpf_loop) keeps track of current
>  	 * simulated iteration number.

[ ... ]

> @@ -698,6 +700,8 @@ struct bpf_insn_aux_data {
>  	u64 non_stack_access:1; /* instruction can access non-stack memory */
>  	/* true if some jump or call instruction targets this instruction */
>  	u64 jump_target:1;
> +	u64 in_cleanup_pad:1; /* reached with a landing pad running */
> +	u64 outside_cleanup_pad:1; /* ... and the other way round */
>
>  	unsigned int orig_idx; /* original instruction index, initialized once */

Do these two bits grow bpf_insn_aux_data by 8 bytes per instruction?

Before this patch the u64 flag word held 31 bits (nospec through
jump_target), fitting in the low 4 bytes, so orig_idx went at offset +4
inside the flag word. Adding in_cleanup_pad and outside_cleanup_pad makes
it 33 bits. Bit 32 lands in byte 4, so orig_idx moves to offset +8, and
every later member shifts by 4 bytes.

The 4-byte tail padding was already used up by cleanup_pad from 97f5b1e363fd.
The struct therefore rounds up by another 4 bytes, for +8 in total. Measured
from the start of the flag word, the tail goes from 64 bytes to 68, padded to
72.

da01eb889d50 (Suggested-by Eduard Zingerman) was written to cut this struct
from 136 bytes to 128. On those figures this patch takes it back to 136. The
array is allocated for every load as __vmalloc(array_size(sizeof(struct
bpf_insn_aux_data), len), ...), and bpf_patch_insn_data() memmoves and
reallocates it on every patch.

The two marks are only read and written by bpf_exc_check_insn(), which runs
only when env->cleanup_info_cnt is non-zero. So programs that never use
exception tables pay the cost.

The cleanup_pad_head bit added later in the series makes the word 34 bits and
does not change this. No later commit in the series restores the size.

Two possible fixes: narrow alu_state, whose flags BPF_ALU_SANITIZE_SRC
through BPF_ALU_IMMEDIATE occupy only bits 0-4, so the word stays within 32
bits; or keep the pad-membership marks in a side array allocated only when a
cleanup table is present.

> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
> index 1ed0370a171b..86fdf847d33e 100644
> --- a/kernel/bpf/exception.c
> +++ b/kernel/bpf/exception.c
> @@ -12,6 +12,15 @@
>  BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
>  BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
>
> +int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog)
> +{
> +	if (!env->subprog_info[subprog].might_unwind)
> +		return 0;
> +
> +	verbose(env, "subprog %d may unwind and is used as a callback\n", subprog);
> +	return -EINVAL;
> +}
> +
>  static void mark_call_sites(struct bpf_verifier_env *env)
>  {
>  	u32 i, j;

[ ... ]

> @@ -69,6 +78,78 @@ bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
>  	       insn->imm == bpf_unwind_resume_id[0];
>  }
>
> +/* Is an unwind in flight: is this frame a landing pad, or below one? */
> +static bool unwinding(const struct bpf_verifier_state *state)
> +{
> +	u32 i;
> +
> +	for (i = 0; i <= state->curframe; i++)
> +		if (state->frame[i]->in_pad)
> +			return true;
> +	return false;
> +}
> +
> +int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn)
> +{
> +	bool in_pad = cur_func(env)->in_pad;
> +	struct bpf_insn_aux_data *aux;
> +	u32 i = env->insn_idx;
> +	const char *why = NULL;
> +
> +	if (unwinding(env->cur_state)) {
> +		if (bpf_is_unwind_kfunc(insn)) {
> +			verbose(env, "insn %u starts a second unwind while one is in flight\n", i);
> +			return -EINVAL;
> +		}
> +		if (bpf_pseudo_call(insn)) {
> +			int subprog = bpf_find_subprog(env, i + insn->imm + 1);
> +
> +			if (subprog >= 0 && bpf_subprog_is_global(env, subprog) &&
> +			    env->subprog_info[subprog].might_unwind) {
> +				verbose(env,
> +					"insn %u calls global subprog %d, which can unwind while an unwind is in flight\n",
> +					i, subprog);
> +				return -EINVAL;
> +			}
> +		}
> +	}
> +
> +	aux = &env->insn_aux_data[i];
> +
> +	if (in_pad ? aux->outside_cleanup_pad : aux->in_cleanup_pad) {
> +		verbose(env, "insn %u runs both inside and outside a landing pad\n", i);
> +		return -EINVAL;
> +	}
> +	if (in_pad)
> +		aux->in_cleanup_pad = true;
> +	else
> +		aux->outside_cleanup_pad = true;
> +
> +	if (!in_pad)
> +		return 0;
> +
> +	if (insn->code == (BPF_JMP | BPF_EXIT)) {
> +		verbose(env,
> +			"exit at insn %u ends a landing pad: a catch pad is not supported yet, only cleanup pads that resume\n",
> +			i);
> +		return -EOPNOTSUPP;
> +	}
> +	if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call)
> +		why = "is a tail call, which replaces the frame";
> +	else if (BPF_CLASS(insn->code) == BPF_LD &&
> +		 (BPF_MODE(insn->code) == BPF_ABS || BPF_MODE(insn->code) == BPF_IND))
> +		why = "is a BPF_LD_[ABS|IND], which can leave through the epilogue";
> +	else if (insn->code == (BPF_JMP | BPF_JA | BPF_X) ||
> +		 insn->code == (BPF_JMP32 | BPF_JA | BPF_X))
> +		why = "is an indirect jump";
> +
> +	if (!why)
> +		return 0;
> +
> +	verbose(env, "insn %u %s, and is in a landing pad\n", i, why);
> +	return -EINVAL;
> +}
> +

The commit message states "bpf_unwind() and the branch pushed at a covered
call are the only ways into a pad, and both mark the frame they enter." Does
bpf_exc_check_insn() also verify callx (indirect calls)?

mark_call_sites() only sets aux->cleanup_pad for bpf_pseudo_call() and
bpf_unwind(), not for bpf_is_callx():

kernel/bpf/exception.c:mark_call_sites() {
	...
	if (!bpf_pseudo_call(insn) && !bpf_is_unwind_kfunc(insn))
		continue;
	env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
}

So bpf_exc_pad_of_call() returns -1 for a callx inside the covered range, and
push_cleanup_pad_branch() in do_check_insn() returns 0 without pushing a pad
state. The CFG and liveness successors add no pad edge for it either.

The verifier walks the callx into its static callee through
check_static_func_call(). If the callee's bpf_unwind() is not covered in the
callee, process_bpf_unwind() returns PROCESS_BPF_EXIT and the path ends
there. The caller frame is never continued, and the pad is never explored from
the state at the callx.

But at run time the unwind does reach that pad. bpf_unwind_rewrite() matches
each frame's return address with bpf_exc_pad_for_ip(), which is a pure ip
range check (begin < ip <= end) over ranges that bpf_exc_fill_native_ranges()
builds. There is no filter on call type. The callee's return address sits
inside the covered range, so it is rewritten to rec->pad. The callee's
epilogue then restores the caller's r6-r9 and stack as they were at the
callx.

Example: a covered range holds call sub_a (pseudo call) with r6 = a valid
pointer, then r6 = scalar; r2 = sub_b ll; callx r2. The pad dereferences or
stores through r6 and then calls bpf_unwind_resume(). sub_b calls
bpf_unwind() unconditionally. The pad is verified only with the state pushed
at call sub_a, where r6 is a pointer. At run time sub_b's unwind resumes the
pad with r6 holding the scalar, which gives an arbitrary kernel memory access
from a program the verifier accepted.

The same entry also gets around what this patch sets out to refuse. A pad
reached only by normal flow and by the callx is verified with in_pad false, so
it can end in exit (a catch pad), use a tail call or LD_ABS, and pass the
"both inside and outside" check, because no in-pad state ever reaches it.

No later commit in the series marks, refuses or filters callx.
mark_call_sites(), push_cleanup_pad_branch() and bpf_exc_pad_for_ip() are
unchanged in this respect, and neither the x86 nor the arm64 dispatch patch
handles callx.

Possible fixes: mark bpf_is_callx() sites in mark_call_sites(), so
push_cleanup_pad_branch() and the CFG treat them like pseudo calls.
Alternatively, refuse a callx inside a covered range when the cleanup table is
checked.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430

^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk
  2026-09-27 20:39   ` bot+bpf-ci
@ 2026-09-28  0:01     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-28  0:01 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 9/27/26 1:39 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/include/linux/filter.h b/include/linux/filter.h
>> index e42eccb0990e..972b3ed2a51d 100644
>> --- a/include/linux/filter.h
>> +++ b/include/linux/filter.h
>> @@ -1248,6 +1248,7 @@ bool bpf_jit_supports_stack_args(void);
>>   bool bpf_jit_supports_arena_args(void);
>>   bool bpf_jit_supports_far_kfunc_call(void);
>>   bool bpf_jit_supports_exceptions(void);
>> +bool bpf_jit_supports_cleanup_pads(void);
>>   bool bpf_jit_supports_ptr_xchg(void);
>>   bool bpf_jit_supports_arena(void);
>>   bool bpf_jit_supports_insn(struct bpf_insn *insn, bool in_arena);
> [ ... ]
>
>> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
>> index b19fcbf49b7e..1ed0370a171b 100644
>> --- a/kernel/bpf/exception.c
>> +++ b/kernel/bpf/exception.c
>> @@ -7,9 +7,56 @@
>>   #include <linux/filter.h>
>>   #include "exception.h"
>>
>> +#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
>> +
>>   BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
>>   BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
>>
>> +static void mark_call_sites(struct bpf_verifier_env *env)
>> +{
>> +	u32 i, j;
>> +
>> +	for (i = 0; i < env->cleanup_info_cnt; i++) {
>> +		struct bpf_cleanup_info *rec = &env->cleanup_info[i];
>> +
>> +		for (j = rec->begin_off; j < rec->end_off; j++) {
>> +			struct bpf_insn *insn = &env->prog->insnsi[j];
>> +
>> +			if (!bpf_pseudo_call(insn) && !bpf_is_unwind_kfunc(insn))
>> +				continue;
>> +			env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
>> +		}
>> +	}
>> +}
> Does mark_call_sites() handle indirect BPF-to-BPF calls? The function
> marks only direct calls (bpf_pseudo_call) and bpf_unwind(), but skips
> callx instructions (BPF_JMP|BPF_CALL|BPF_X).

I missed newly support callx in verifier. I will add it in the next revision.

>
> check_func_callx() in the verifier takes a PTR_TO_FUNC to a static
> subprog and verifies it through check_static_func_call(), the same path
> as direct calls. Nothing prevents that callee from calling bpf_unwind()
> since bpf_exc_check_callback() only runs for helper/kfunc callbacks, not
> static subprogs.
>
> The commit message states cleanup_pad is recorded "for every call that
> can unwind -- a BPF-to-BPF call or bpf_unwind()" but a callx inside the
> [begin_off, end_off) range keeps cleanup_pad == 0.
>
> The infrastructure that consumes cleanup_pad already handles callx:
>
> kernel/bpf/verifier.c:
> visit_func_call_insn() and push_cleanup_pad_branch() run for every
> BPF_CALL instruction, including callx. If cleanup_pad is not set, they
> treat the call as unable to unwind.
>
> This means:
> - visit_cleanup_pad_edge() -> bpf_exc_pad_of_call() returns -1, so no
>    CFG edge is created to the landing pad
> - push_cleanup_pad_branch() returns 0 before check_func_callx() runs, so
>    the verifier never explores the landing pad state
> - bpf_unwind_rewrite() finds no cleanup range for the frame and sets
>    *ra = prog->aux->epilogue_ip
>
> Concrete scenario: a C++ or Rust program makes a virtual or trait call
> (lowered to callx) inside a scope whose drop glue sits in a landing pad.
> If the callee unwinds, the caller frame goes to its epilogue and the
> drop glue never runs. If that drop glue releases a kernel reference or
> kptr, the release is skipped. Meanwhile the verifier modeled the
> callee's resume as a normal return to callx+1.
>
> Should the test be widened to include bpf_is_callx(), or should cleanup
> tables that cover a callx be rejected?
>
>> +
>> +int bpf_prepare_cleanup_exceptions(struct bpf_verifier_env *env)
>> +{
>> +	if (!env->cleanup_info_cnt)
>> +		return 0;
>> +
>> +	if (bpf_prog_is_offloaded(env->prog->aux)) {
>> +		verbose(env,
>> +			"exception cleanup is not supported for offloaded programs\n");
>> +		return -EINVAL;
>> +	}
>> +
>> +	if (!bpf_jit_supports_cleanup_pads() || !env->prog->jit_requested) {
>> +		verbose(env,
>> +			"exception cleanup needs a JIT that can dispatch landing pads\n");
>> +		return -EOPNOTSUPP;
>> +	}
>> +	env->prog->jit_required = 1;
>> +
>> +	if (env->exception_callback_subprog) {
>> +		verbose(env,
>> +			"exception cleanup table cannot be combined with an exception callback\n");
>> +		return -EINVAL;
>> +	}
>> +
>> +	mark_call_sites(env);
>> +	return 0;
>> +}
> [ ... ]
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG
  2026-09-27 20:40   ` bot+bpf-ci
@ 2026-09-28  0:12     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-28  0:12 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 9/27/26 1:40 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
>> index b0bd9ba951df..4e2b6985bc96 100644
>> --- a/kernel/bpf/cfg.c
>> +++ b/kernel/bpf/cfg.c
>> @@ -6,6 +6,7 @@
>>   #include <linux/sort.h>
>>
>>   #include "diagnostics.h"
>> +#include "exception.h"
>>
>>   #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
>>
> [ ... ]
>
>> @@ -160,17 +161,64 @@ static int push_insn(int t, int w, int e, struct bpf_verifier_env *env)
>>   	return DONE_EXPLORING;
>>   }
>>
>> +static int visit_cleanup_pad_edge(int t, struct bpf_verifier_env *env)
>> +{
>> +	int *insn_stack = env->cfg.insn_stack;
>> +	int *insn_state = env->cfg.insn_state;
>> +	int w;
>> +
>> +	if (!env->cleanup_info_cnt)
>> +		return DONE_EXPLORING;
>> +	w = bpf_exc_pad_of_call(env, t);
>> +	if (w < 0)
>> +		return DONE_EXPLORING;
>> +
>> +	/*
>> +	 * @t is a call that may branch here, and @w is the target of that
>> +	 * branch, so both are prune points. @w especially: every covered call
>> +	 * site in a region unwinds to the same pad, and without a prune point
>> +	 * at its head the verifier walks the pad again for each of them.
>> +	 */
>> +	mark_prune_point(env, t);
>> +	mark_prune_point(env, w);
>> +	mark_jmp_point(env, w);
>> +	mark_jump_target(env, w);
>> +
>> +	if (insn_state[w])
>> +		return DONE_EXPLORING;
>> +	if (env->cfg.cur_stack >= env->prog->len)
>> +		return -E2BIG;
>> +	insn_stack[env->cfg.cur_stack++] = w;
>> +	insn_state[w] |= DISCOVERED;
>> +	return KEEP_EXPLORING;
>> +}
>> +
>> +static int merge_visit_ret(int a, int b)
>> +{
>> +	if (a < 0)
>> +		return a;
>> +	if (b < 0)
>> +		return b;
>> +	if (a == KEEP_EXPLORING || b == KEEP_EXPLORING)
>> +		return KEEP_EXPLORING;
>> +	return DONE_EXPLORING;
>> +}
>> +
>>   static int visit_func_call_insn(int t, struct bpf_insn *insns,
>>   				struct bpf_verifier_env *env,
>>   				bool visit_callee)
>>   {
>> -	int ret, insn_sz;
>> +	int ret, insn_sz, pad_ret;
>>   	int w;
>>
>> +	pad_ret = visit_cleanup_pad_edge(t, env);
>> +	if (pad_ret < 0)
>> +		return pad_ret;
> The error handling here appears to be incorrect. In the v6 discussion
> thread, you identified that this should be:
>
> 	if (pad_ret != DONE_EXPLORING)
> 		return pad_ret;
>
> rather than checking for negative values. The current code causes premature
> DFS stop by ignoring KEEP_EXPLORING return values from
> visit_cleanup_pad_edge(). When the cleanup pad edge returns KEEP_EXPLORING,
> the DFS should continue, but the negative-only check misses this case and
> treats it as DONE_EXPLORING.
>
> Can this be changed to properly handle both KEEP_EXPLORING and error cases?

Replied earlier to Sashiko AI. Fix is the same as you suggested in the above.

>
>>   	insn_sz = bpf_is_ldimm64(&insns[t]) ? 2 : 1;
>>   	ret = push_insn(t, t + insn_sz, FALLTHROUGH, env);
>>   	if (ret)
>> -		return ret;
>> +		return merge_visit_ret(pad_ret, ret);
>>
>>   	mark_prune_point(env, t + insn_sz);
>>   	/* when we exit from subprog, we need to record non-linear history */
>> @@ -182,7 +230,7 @@ static int visit_func_call_insn(int t, struct bpf_insn *insns,
>>   		merge_callee_effects(env, t, w);
>>   		ret = push_insn(t, w, BRANCH, env);
>>   	}
>> -	return ret;
>> +	return merge_visit_ret(pad_ret, ret);
>>   }
> [ ... ]
>
>> diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
>> index cd9523f69298..4e0273a8ceee 100644
>> --- a/kernel/bpf/liveness.c
>> +++ b/kernel/bpf/liveness.c
>> @@ -8,6 +8,8 @@
>>   #include <linux/slab.h>
>>   #include <linux/sort.h>
>>
>> +#include "exception.h"
>> +
>>   #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
> [ ... ]
>
>> @@ -384,6 +386,18 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
>>   			succ->items[succ->cnt++] = exit_idx;
>>   	}
>>
>> +	/*
>> +	 * A call a cleanup record covers can leave through its landing pad.
>> +	 * Only a call to a subprogram or to bpf_unwind() is marked, neither of
>> +	 * which is an edge the block above adds, so succ still holds two.
>> +	 */
>> +	if (unlikely(env->cleanup_info_cnt)) {
>> +		int pad = bpf_exc_pad_of_call(env, idx);
>> +
>> +		if (pad >= 0)
>> +			succ->items[succ->cnt++] = pad;
>> +	}
>> +
>>   	return succ;
>>   }
> [ ... ]
>
>> @@ -545,6 +559,13 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
>>   		alive = callee_stack_access_at_callsite(env, callsite)
>>   			? is_live_before(instance, callsite, rel, half_spi)
>>   			: is_live_before(instance, callsite + 1, rel, half_spi);
>> +
>> +		if (!alive && unlikely(env->cleanup_info_cnt)) {
>> +			int pad = bpf_exc_pad_of_call(env, callsite);
>> +
>> +			if (pad >= 0)
>> +				alive = is_live_before(instance, pad, rel, half_spi);
>> +		}
> Eduard Zingerman asked in v5 why this landing pad liveness check is
> necessary. He stated: "If bpf_insn_succ correctly traverses call successors
> including the landing pad the information should already be accounted for
> by the is_live_before(instance, callsite, rel, half_spi)."
>
> This question was not addressed in subsequent versions. Looking at the
> patch series design, there's a question about which model is intended.
>
> The commit message justifies this hunk with: "otherwise
> clean_verifier_state() poisons the slot while the callee runs and the pad
> is rejected for reading it."
>
> However, at this series tip, the verifier never enters a landing pad from a
> state where the pad's frame is an outer frame. There are only two entry
> paths:
>
> 1. push_cleanup_pad_branch() (added in patch 5f426cf2a1ff) pushes a copy of
>     the caller's state before check_func_call(), with the pad in the current
>     frame.
>
> 2. process_bpf_unwind() jumps to a pad only in the current frame; otherwise
>     it returns PROCESS_BPF_EXIT and bpf_unwind_resume goes through
>     process_bpf_exit_full() to callsite + 1.
>
> Since clean_verifier_state() is called only on cur in is_state_visited(),
> poisoning a caller slot in cur while the callee runs cannot affect the
> pushed pad copy. That copy was taken at the call site, where the caller is
> the current frame, and bpf_insn_successors() already includes the pad
> there.
>
> As written, the outer-frame pad check only keeps extra caller slots alive
> in callee states, which reduces pruning. It does not prevent any pad
> rejection.

I think we are talking about case 2 in the above. The function is a
*always* unwinder, so its subsequent code is not used any more and
callee_stack_access_at_callsite() also returns false as the function
is not a callback or a callx.

static bool callee_stack_access_at_callsite(struct bpf_verifier_env *env, u32 callsite)
{
         return bpf_calls_callback(env, callsite) ||
                bpf_is_callx(&env->prog->insnsi[callsite]);
}

So the only next target is the pad.

>
> Either the rationale is inaccurate and the hunk is unnecessary, or the
> intended model is that the pad sees the caller's frame as the callee left
> it. The second reading is closer to runtime behaviour: a callee can write
> the caller's stack through a PTR_TO_STACK argument before bpf_unwind(). But
> that is not what push_cleanup_pad_branch() models.
>
> Which model is intended? If the pre-call copy is the design, could this
> hunk and the changelog paragraph be dropped? If the post-callee state is
> the design, does the pad verification in the later verifier patch need to
> change instead?
>
>>   		if (alive)
>>   			return true;
>>   	}
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad
  2026-09-27 20:40   ` bot+bpf-ci
@ 2026-09-28  0:17     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-28  0:17 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 9/27/26 1:40 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
>> index 6d78c20e6507..0143688896b0 100644
>> --- a/include/linux/bpf_verifier.h
>> +++ b/include/linux/bpf_verifier.h
>> @@ -836,6 +836,7 @@ struct bpf_subprog_info {
>>   	s16 fastcall_stack_off;
>>   	bool has_tail_call: 1;
>>   	bool might_throw: 1;
>> +	bool might_unwind: 1;
>>   	bool tail_call_reachable: 1;
>>   	bool has_ld_abs: 1;
>>   	bool is_cb: 1;
> [ ... ]
>
>> diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
>> index 08aee86a155c..f291611fe578 100644
>> --- a/kernel/bpf/helpers.c
>> +++ b/kernel/bpf/helpers.c
>> @@ -5095,6 +5095,8 @@ BTF_ID_FLAGS(func, bpf_task_from_pid, KF_ACQUIRE | KF_RET_NULL)
>>   BTF_ID_FLAGS(func, bpf_task_from_vpid, KF_ACQUIRE | KF_RET_NULL)
>>   BTF_ID_FLAGS(func, bpf_throw)
>> +BTF_ID_FLAGS(func, bpf_unwind)
>> +BTF_ID_FLAGS(func, bpf_unwind_resume)
>>   #ifdef CONFIG_BPF_EVENTS
>>   BTF_ID_FLAGS(func, bpf_send_signal_task)
>>   #endif
> Can a CAP_BPF program call bpf_unwind() at this commit? The kfuncs are
> registered in generic_btf_ids, which makes them resolvable, and the body
> at kernel/bpf/helpers.c is still __bpf_kfunc void bpf_unwind(void) { },
> so the call returns normally to the next instruction.
>
> However, the verifier assumes bpf_unwind() never falls through:
>
> kernel/bpf/verifier.c:process_bpf_unwind() ends the path with
> PROCESS_BPF_EXIT for an uncovered call, and jumps to the pad for a
> covered call. liveness.c keeps a fall-through edge, but do_check()
> never explores it. As a result, any instructions reachable only from
> after the call are skipped by the verifier and removed by
> bpf_opt_remove_dead_code().
>
> At run time the call returns and execution continues into whatever the
> JIT placed next, with r0-r5 clobbered, or off the end of the image.
>
> This is fixed later in the series by 7924e9e105cf "bpf: Dispatch cleanup
> pads by rewriting return addresses", which implements the real body via
> arch_bpf_stack_walk_ra()/bpf_unwind_rewrite(), and bpf_exc_keep_exits()
> puts 'r0 = 0; exit' after every bpf_unwind call.

As you suggested, I will move two BTF_ID_FLAGS's into commit
"bpf: Dispatch cleanup pads by rewriting return addresses".

>
> [ ... ]
>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index fc3df452de2e..77176250f866 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
>> @@ -19167,6 +19167,40 @@ enum {
>>   	INSN_IDX_UPDATED = 2,
>>   };
>>
>> +static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
>> +{
>> +	struct bpf_verifier_state *branch;
>> +	struct bpf_func_state *frame;
>> +	int pad = bpf_exc_pad_of_call(env, insn_idx);
>> +
>> +	if (pad < 0)
>> +		return 0;
>> +	branch = push_stack(env, pad, insn_idx, false);
>> +	if (IS_ERR(branch))
>> +		return PTR_ERR(branch);
>> +	frame = branch->frame[branch->curframe];
>> +	/*
>> +	 * The state at that call with the caller-saved registers gone: the
>> +	 * callee's epilogue put r6-r9 and the stack back on the way out.
>> +	 */
>> +	clear_caller_saved_regs(env, frame->regs);
>> +	mark_reg_unknown(env, frame->regs, BPF_REG_0);
>> +	return 0;
>> +}
>> +
>> +static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx)
>> +{
>> +	struct bpf_func_state *frame = cur_func(env);
>> +	int pad = bpf_exc_pad_of_call(env, *insn_idx);
>> +
>> +	if (pad < 0)
>> +		return PROCESS_BPF_EXIT;
>> +	clear_caller_saved_regs(env, frame->regs);
>> +	mark_reg_unknown(env, frame->regs, BPF_REG_0);
>> +	*insn_idx = pad;
>> +	return INSN_IDX_UPDATED;
>> +}
>> +
>>   static int process_bpf_exit_full(struct bpf_verifier_env *env,
>>   				 bool *do_print_state,
>>   				 bool exception_exit)
>> @@ -19404,6 +19438,20 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>>
>>   		env->jmps_processed++;
>>   		if (opcode == BPF_CALL) {
>> +			if (bpf_is_unwind_kfunc(insn))
>> +				return process_bpf_unwind(env, &env->insn_idx);
>> +			if (bpf_is_unwind_resume_kfunc(insn)) {
>> +				/*
>> +				 * Mark r0 a known zero -- unknown first, as
>> +				 * the known-zero helper keeps the type it
>> +				 * finds, which here is NOT_INIT. The fixups
>> +				 * lower this to 'r0 = 0; exit', so the frame
>> +				 * returns a real zero.
>> +				 */
>> +				mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
>> +				mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
>> +				return process_bpf_exit_full(env, do_print_state, false);
>> +			}
>>   			if (env->cur_state->active_locks) {
>>   				/* similar to static subprog calls callx is allowed under a lock */
>>   				if (!bpf_is_callx(insn) &&
> Both new kfuncs are handled before check_kfunc_call(), so
> bpf_fetch_kfunc_arg_meta() never runs for them. That function rejects a
> kfunc with -EACCES "calling kernel function %s is not allowed" when
> btf_kfunc_is_allowed() fails for the program type or the kfunc has no
> flags. bpf_add_kfunc_call() only resolves the BTF id and does not apply
> that filter.

You are right, we need bpf_fetch_kfunc_arg_meta check. Will fix.

>
> As a result, bpf_unwind() and bpf_unwind_resume() are callable from any
> program type that can make kfunc calls, including types that
> generic_kfunc_set is not registered for (it is registered for TRACING,
> SCHED_CLS, XDP, STRUCT_OPS, SYSCALL and CGROUP_SKB).
>
> It also skips the usual sleepable, RCU, preempt and IRQ-context checks
> in check_kfunc_call().
>
> ---
>
> At this commit a bpf_unwind_resume() call is verified as a frame exit
> with r0 = 0, but the comment says "The fixups lower this to
> 'r0 = 0; exit'". Does that lowering exist yet? No fixup for it exists in
> kernel/bpf/fixups.c at this commit.

Comment is not right. The fixup is in later patch
"bpf: Dispatch cleanup pads by rewriting return addresses".

>
> This call is also accepted anywhere, not only inside a landing pad. So a
> CAP_BPF program can place 'call bpf_unwind_resume' in its main path. The
> verifier treats the path as ended and never looks at the following
> instructions, which the dead-code pass then removes or sanitizes.
>
> At run time the real kfunc body runs WARN_ONCE(1, "exception cleanup
> resume was not lowered to a return\n"). That is a user-triggerable
> warning, and a panic with panic_on_warn. The call then returns normally
> with r0-r5 clobbered and falls into instructions that were never
> verified for this state, or off the end of the JIT image.
>
> This is fixed later in the series by 7924e9e105cf "bpf: Dispatch cleanup
> pads by rewriting return addresses", which adds the bpf_do_misc_fixups
> lowering, and 2a1e4ed1e9ae "bpf: Refuse a landing pad that does not
> resume", which rejects "resume at insn %d is not in a landing pad".
>
>> @@ -19422,6 +19470,10 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>>   				}
>>   			}
>>   			mark_reg_scratched(env, BPF_REG_0);
>> +			/* An unwind out of this call resumes at the pad. */
>> +			err = push_cleanup_pad_branch(env, env->insn_idx);
>> +			if (err)
>> +				return err;
>>   			if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
>>   				cur_func(env)->no_stack_arg_load = true;
>>   			if (bpf_is_callx(insn))
> [ ... ]
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume
  2026-09-27 20:40   ` bot+bpf-ci
@ 2026-09-28  0:29     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-28  0:29 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 9/27/26 1:40 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
>> index 0143688896b0..4174c7d0177e 100644
>> --- a/include/linux/bpf_verifier.h
>> +++ b/include/linux/bpf_verifier.h
>> @@ -339,6 +339,8 @@ struct bpf_func_state {
>>   	bool in_async_callback_fn;
>>   	bool in_exception_callback_fn;
>>   	bool no_stack_arg_load;
>> +	/* an unwind reached this frame and its landing pad is running */
>> +	bool in_pad;
>>   	/* For callback calling functions that limit number of possible
>>   	 * callback executions (e.g. bpf_loop) keeps track of current
>>   	 * simulated iteration number.
> [ ... ]
>
>> @@ -698,6 +700,8 @@ struct bpf_insn_aux_data {
>>   	u64 non_stack_access:1; /* instruction can access non-stack memory */
>>   	/* true if some jump or call instruction targets this instruction */
>>   	u64 jump_target:1;
>> +	u64 in_cleanup_pad:1; /* reached with a landing pad running */
>> +	u64 outside_cleanup_pad:1; /* ... and the other way round */
>>
>>   	unsigned int orig_idx; /* original instruction index, initialized once */
> Do these two bits grow bpf_insn_aux_data by 8 bytes per instruction?

I would like to keep the current layout for now as this is not
the important part of this patch set.

>
> Before this patch the u64 flag word held 31 bits (nospec through
> jump_target), fitting in the low 4 bytes, so orig_idx went at offset +4
> inside the flag word. Adding in_cleanup_pad and outside_cleanup_pad makes
> it 33 bits. Bit 32 lands in byte 4, so orig_idx moves to offset +8, and
> every later member shifts by 4 bytes.
>
> The 4-byte tail padding was already used up by cleanup_pad from 97f5b1e363fd.
> The struct therefore rounds up by another 4 bytes, for +8 in total. Measured
> from the start of the flag word, the tail goes from 64 bytes to 68, padded to
> 72.
>
> da01eb889d50 (Suggested-by Eduard Zingerman) was written to cut this struct
> from 136 bytes to 128. On those figures this patch takes it back to 136. The
> array is allocated for every load as __vmalloc(array_size(sizeof(struct
> bpf_insn_aux_data), len), ...), and bpf_patch_insn_data() memmoves and
> reallocates it on every patch.
>
> The two marks are only read and written by bpf_exc_check_insn(), which runs
> only when env->cleanup_info_cnt is non-zero. So programs that never use
> exception tables pay the cost.
>
> The cleanup_pad_head bit added later in the series makes the word 34 bits and
> does not change this. No later commit in the series restores the size.
>
> Two possible fixes: narrow alu_state, whose flags BPF_ALU_SANITIZE_SRC
> through BPF_ALU_IMMEDIATE occupy only bits 0-4, so the word stays within 32
> bits; or keep the pad-membership marks in a side array allocated only when a
> cleanup table is present.
>
>> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
>> index 1ed0370a171b..86fdf847d33e 100644
>> --- a/kernel/bpf/exception.c
>> +++ b/kernel/bpf/exception.c
>> @@ -12,6 +12,15 @@
>>   BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
>>   BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
>>
>> +int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog)
>> +{
>> +	if (!env->subprog_info[subprog].might_unwind)
>> +		return 0;
>> +
>> +	verbose(env, "subprog %d may unwind and is used as a callback\n", subprog);
>> +	return -EINVAL;
>> +}
>> +
>>   static void mark_call_sites(struct bpf_verifier_env *env)
>>   {
>>   	u32 i, j;
> [ ... ]
>
>> @@ -69,6 +78,78 @@ bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
>>   	       insn->imm == bpf_unwind_resume_id[0];
>>   }
>>
>> +/* Is an unwind in flight: is this frame a landing pad, or below one? */
>> +static bool unwinding(const struct bpf_verifier_state *state)
>> +{
>> +	u32 i;
>> +
>> +	for (i = 0; i <= state->curframe; i++)
>> +		if (state->frame[i]->in_pad)
>> +			return true;
>> +	return false;
>> +}
>> +
>> +int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn)
>> +{
>> +	bool in_pad = cur_func(env)->in_pad;
>> +	struct bpf_insn_aux_data *aux;
>> +	u32 i = env->insn_idx;
>> +	const char *why = NULL;
>> +
>> +	if (unwinding(env->cur_state)) {
>> +		if (bpf_is_unwind_kfunc(insn)) {
>> +			verbose(env, "insn %u starts a second unwind while one is in flight\n", i);
>> +			return -EINVAL;
>> +		}
>> +		if (bpf_pseudo_call(insn)) {
>> +			int subprog = bpf_find_subprog(env, i + insn->imm + 1);
>> +
>> +			if (subprog >= 0 && bpf_subprog_is_global(env, subprog) &&
>> +			    env->subprog_info[subprog].might_unwind) {
>> +				verbose(env,
>> +					"insn %u calls global subprog %d, which can unwind while an unwind is in flight\n",
>> +					i, subprog);
>> +				return -EINVAL;
>> +			}
>> +		}
>> +	}
>> +
>> +	aux = &env->insn_aux_data[i];
>> +
>> +	if (in_pad ? aux->outside_cleanup_pad : aux->in_cleanup_pad) {
>> +		verbose(env, "insn %u runs both inside and outside a landing pad\n", i);
>> +		return -EINVAL;
>> +	}
>> +	if (in_pad)
>> +		aux->in_cleanup_pad = true;
>> +	else
>> +		aux->outside_cleanup_pad = true;
>> +
>> +	if (!in_pad)
>> +		return 0;
>> +
>> +	if (insn->code == (BPF_JMP | BPF_EXIT)) {
>> +		verbose(env,
>> +			"exit at insn %u ends a landing pad: a catch pad is not supported yet, only cleanup pads that resume\n",
>> +			i);
>> +		return -EOPNOTSUPP;
>> +	}
>> +	if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call)
>> +		why = "is a tail call, which replaces the frame";
>> +	else if (BPF_CLASS(insn->code) == BPF_LD &&
>> +		 (BPF_MODE(insn->code) == BPF_ABS || BPF_MODE(insn->code) == BPF_IND))
>> +		why = "is a BPF_LD_[ABS|IND], which can leave through the epilogue";
>> +	else if (insn->code == (BPF_JMP | BPF_JA | BPF_X) ||
>> +		 insn->code == (BPF_JMP32 | BPF_JA | BPF_X))
>> +		why = "is an indirect jump";
>> +
>> +	if (!why)
>> +		return 0;
>> +
>> +	verbose(env, "insn %u %s, and is in a landing pad\n", i, why);
>> +	return -EINVAL;
>> +}
>> +
> The commit message states "bpf_unwind() and the branch pushed at a covered
> call are the only ways into a pad, and both mark the frame they enter." Does
> bpf_exc_check_insn() also verify callx (indirect calls)?
>
> mark_call_sites() only sets aux->cleanup_pad for bpf_pseudo_call() and
> bpf_unwind(), not for bpf_is_callx():
>
> kernel/bpf/exception.c:mark_call_sites() {
> 	...
> 	if (!bpf_pseudo_call(insn) && !bpf_is_unwind_kfunc(insn))
> 		continue;
> 	env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
> }
>
> So bpf_exc_pad_of_call() returns -1 for a callx inside the covered range, and
> push_cleanup_pad_branch() in do_check_insn() returns 0 without pushing a pad
> state. The CFG and liveness successors add no pad edge for it either.
>
> The verifier walks the callx into its static callee through
> check_static_func_call(). If the callee's bpf_unwind() is not covered in the
> callee, process_bpf_unwind() returns PROCESS_BPF_EXIT and the path ends
> there. The caller frame is never continued, and the pad is never explored from
> the state at the callx.
>
> But at run time the unwind does reach that pad. bpf_unwind_rewrite() matches
> each frame's return address with bpf_exc_pad_for_ip(), which is a pure ip
> range check (begin < ip <= end) over ranges that bpf_exc_fill_native_ranges()
> builds. There is no filter on call type. The callee's return address sits
> inside the covered range, so it is rewritten to rec->pad. The callee's
> epilogue then restores the caller's r6-r9 and stack as they were at the
> callx.
>
> Example: a covered range holds call sub_a (pseudo call) with r6 = a valid
> pointer, then r6 = scalar; r2 = sub_b ll; callx r2. The pad dereferences or
> stores through r6 and then calls bpf_unwind_resume(). sub_b calls
> bpf_unwind() unconditionally. The pad is verified only with the state pushed
> at call sub_a, where r6 is a pointer. At run time sub_b's unwind resumes the
> pad with r6 holding the scalar, which gives an arbitrary kernel memory access
> from a program the verifier accepted.
>
> The same entry also gets around what this patch sets out to refuse. A pad
> reached only by normal flow and by the callx is verified with in_pad false, so
> it can end in exit (a catch pad), use a tail call or LD_ABS, and pass the
> "both inside and outside" check, because no in-pad state ever reaches it.
>
> No later commit in the series marks, refuses or filters callx.
> mark_call_sites(), push_cleanup_pad_branch() and bpf_exc_pad_for_ip() are
> unchanged in this respect, and neither the x86 nor the arm64 dispatch patch
> handles callx.
>
> Possible fixes: mark bpf_is_callx() sites in mark_call_sites(), so
> push_cleanup_pad_branch() and the CFG treat them like pseudo calls.
> Alternatively, refuse a callx inside a covered range when the cleanup table is
> checked.

Yes, fix in the next revision by adding callx support in mark_all_sites().

>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 10/21] bpf: Dispatch cleanup pads by rewriting return addresses
  2026-09-27 20:40   ` bot+bpf-ci
@ 2026-09-28  1:08     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-28  1:08 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 9/27/26 1:40 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/include/linux/bpf.h b/include/linux/bpf.h
>> index 4bae3796c42f..98de251037df 100644
>> --- a/include/linux/bpf.h
>> +++ b/include/linux/bpf.h
> [ ... ]
>
>> @@ -1885,6 +1928,8 @@ struct bpf_prog_aux {
>>   	u64 (*bpf_exception_cb)(u64 cookie, u64 sp, u64 bp, u64, u64);
>>   	u16 stack_arg_sp_adjust;
>>   	u16 freplace_link_cnt; /* counts freplace links extending this prog */
>> +	struct bpf_exception_info *exc;
>> +	u64 epilogue_ip; /* native address of this (sub)program's epilogue */
>>   #ifdef CONFIG_SECURITY
>>   	void *security;
>>   #endif
> The commit message says epilogue_ip "is recorded for every program", but
> can a caller that only calls subprogs which always unwind end up with
> epilogue_ip still zero?
>
> Looking at bpf_exc_keep_exits(), it only patches an exit after
> bpf_unwind() kfunc calls:
>
> kernel/bpf/fixups.c:bpf_exc_keep_exits() {
>      for (i = 0; i < insn_cnt; i++) {
>          if (!bpf_is_unwind_kfunc(insn))
>              continue;
>          /* patch exit after call */
>      }
> }
>
> It doesn't patch exits after calls to subprogs that can only unwind.
> The verifier ends that path at the subprog's bpf_unwind()
> (process_bpf_unwind() returns PROCESS_BPF_EXIT), so code after the call
> is never seen and gets swept. If that was the caller's only exit, it
> never emits a BPF_EXIT, so the JIT never sets epilogue_ip.
>
> At run time bpf_unwind_rewrite() would then reach:
>
>      } else if (prog->aux->epilogue_ip) {
>          *ra = prog->aux->epilogue_ip;
>      } else {
>          WARN_ON_ONCE(1);
>          return false;
>      }
>
> with epilogue_ip still 0. The return address stays pointing just past the
> call, which is the end of the JIT image.
>
> Example: main calls a static subprog that only has `call bpf_unwind; exit`.
> With no cleanup table (or no record covering either call), the main frame
> would be redirected to epilogue_ip, but that field was never set.

Yes, this is indeed a bug. Will fix to find at least one 'exit'.

>
> [ ... ]
>
>> diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
>> index d813fdde29e3..60905643cb9c 100644
>> --- a/kernel/bpf/core.c
>> +++ b/kernel/bpf/core.c
> [ ... ]
>
>> @@ -3511,6 +3523,12 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
>>   {
>>   }
>>
>> +void __weak arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp,
>> +						      u64 *ra),
>> +				   void *cookie)
>> +{
>> +}
>> +
>>   bool __weak bpf_jit_supports_cleanup_pads(void)
>>   {
>>   	return false;
> Does bpf_unwind() produce correct behavior on architectures that keep this
> weak stub?
>
> At this commit no architecture implements the walker. At the end of the
> series only x86-64 with CONFIG_UNWINDER_ORC and arm64 do. x86 with the
> frame-pointer unwinder, riscv, s390, powerpc and loongarch keep the stub.
>
> Nothing stops those architectures from using bpf_unwind():
>
> - It is in generic_btf_ids, registered for TRACING, SCHED_CLS, XDP,
>    STRUCT_OPS, SYSCALL, CGROUP_SKB and others.
>
> - bpf_prepare_cleanup_exceptions() only checks bpf_jit_supports_cleanup_pads()
>    when a table was supplied:
>
>        if (!env->cleanup_info_cnt)
>            return 0;
>
> - An uncovered call is a terminator for the verifier:
>
>        if (pad < 0)
>            return PROCESS_BPF_EXIT;
>
> At run time the weak stub does nothing, so bpf_unwind() returns. The frame
> returns through the patched 'r0 = 0; exit'. Each caller then resumes at
> the instruction after its call, in a state the verifier never explored for
> that path, or in code the dead code sweep already removed.
>
> Concrete case: a static subprog that returns a checked non-NULL map value
> pointer on its normal path and calls bpf_unwind() on another path. The
> caller dereferences the returned pointer without a NULL check, since the
> verifier only saw the pointer-returning exit. At run time the subprog
> returns 0 and the caller writes through NULL in kernel context.
>
> The commit message says the separate entry point exists "so that the
> architectures which do not dispatch pads keep the walker they have."
> Does bpf_unwind() need to be gated on bpf_jit_supports_cleanup_pads(), or
> should the stub have a fallback that doesn't leave callers in unverified
> code?

Okay, you are correct. condition env->cleanup_info_cnt is not enough
to allow bpf_unwind() since bpf_unwind() typically won't be in cleanup
region. Will fix it by checking whether any function is might_unwind
or not.

>
>> diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
>> index f291611fe578..ea5c4d81f8b7 100644
>> --- a/kernel/bpf/helpers.c
>> +++ b/kernel/bpf/helpers.c
> [ ... ]
>
>> @@ -3424,8 +3426,51 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
>>   	return false;
>>   }
>>
>> +struct bpf_unwind_ctx {
>> +	u32 cnt;
>> +};
>> +
>> +static bool bpf_unwind_rewrite(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra)
>> +{
>> +	const struct bpf_cleanup_range *rec;
>> +	struct bpf_unwind_ctx *ctx = cookie;
>> +	struct bpf_exception_info *exc;
>> +	struct bpf_prog *prog;
>> +
>> +	rcu_read_lock();
>> +	prog = bpf_prog_ksym_find(ip);
>> +	rcu_read_unlock();
>> +	if (!prog)
>> +		return !ctx->cnt;
>> +	ctx->cnt++;
>> +
>> +	exc = prog->aux->exc;
>> +	rec = (exc && exc->nr_ranges) ? bpf_exc_pad_for_ip(prog, ip) : NULL;
>> +	if (rec) {
>> +		*ra = rec->pad;
>> +	} else if (ctx->cnt == 1) {
>> +		/*
>> +		 * The frame that called bpf_unwind(). Its return address
>> +		 * always names the 'r0 = 0; exit' that bpf_exc_keep_exits()
>> +		 * put after the call, so leave it alone and let the frame
>> +		 * return through that: running it is what sets the value
>> +		 * the unwind returns.
>> +		 */
>> +	} else if (prog->aux->epilogue_ip) {
>> +		*ra = prog->aux->epilogue_ip;
> Does sending a frame to epilogue_ip work correctly when the verifier
> treats the resume as returning to the instruction after the call?
>
> At run time, a caller frame whose call site no cleanup record covers is
> sent to aux->epilogue_ip, so it returns straight away. The verifier does
> not model that: it treats a pad's bpf_unwind_resume() as an ordinary return
> into the caller, at the instruction after the call. The code after the call
> is therefore verified but never runs, and the return-at-once path is never
> checked.
>
> kernel/bpf/verifier.c:do_check_insn() with curframe > 0:
>
>      cur_func(env)->in_pad = false;
>      return process_bpf_exit_full(env, do_print_state, false);
>
> process_bpf_exit_full(..., false) goes to prepare_func_exit(), which
> continues the caller at callsite + 1 with r0 known to be zero. Only a
> covered call site gets its pad pushed as another branch
> (push_cleanup_pad_branch()). Nothing requires a call to a might_unwind
> static subprog to be covered, and bpf_exc_check_insn() only restricts global
> subprogs while an unwind is in flight.
>
> Concrete case (x86-64 with ORC, or arm64, at the end of the series):
>
>      main:  t = bpf_task_acquire(p); if (!t) return 0;
>             sub();                 /* call site not covered */
>             bpf_task_release(t); return 0;
>      sub:   bpf_unwind();          /* covered by a record in sub */
>             pad: bpf_unwind_resume(0);
>
> The verifier accepts this because it walks main past the call and sees the
> release. At run time sub's pad runs, resumes, and returns into main's
> epilogue, so the task reference is leaked.
>
> The same shape with bpf_rcu_read_lock()/unlock(), preempt disable/enable,
> or bpf_spin_lock()/unlock() leaves the lock held when the program returns.
> Static subprog calls are allowed under a spin lock, and the bpf_unwind
> check comes before the active_locks check. bpf_spin_lock is taken with IRQs
> disabled, so the next locker deadlocks.
>
> Does the verifier need to model the epilogue_ip redirect for uncovered
> callers, or should uncovered calls be refused when the callee might unwind?

I think proper resource checking at every level is needed.

>
> ---
>
> On x86-64 with a private stack, does this redirect skip the pop r9 that
> follows the call?
>
> With a private stack, the JIT wraps every call in push_r9/pop_r9
> (arch/x86/net/bpf_jit_comp.c). When bpf_unwind_rewrite() sets
> `*ra = prog->aux->epilogue_ip` for a caller frame, the callee's ret lands
> on the epilogue with the pushed r9 still on top of the stack. The epilogue
> is `add rsp, outgoing_rsp; pop_callee_regs(); leave; ret`.
> pop_callee_regs() therefore pops the saved r9 into the first callee-saved
> register and shifts the rest by one slot. leave then fixes rsp, so the
> function returns normally, but with the kernel caller's rbx/r13/r14/r15
> corrupted.
>
> check_max_stack_depth() refuses a private stack only when a table is
> present:
>
>      if (env->cleanup_info_cnt)
>          priv_stack_mode = NO_PRIV_STACK;
>
> Its own comment describes this exact problem for pads ("on x86-64 that
> skips the pop which restores r9"). A table-less program can still reach the
> redirect.
>
> Example: an fentry or kprobe program whose main prog uses at least
> BPF_PRIV_STACK_MIN_SIZE bytes of stack and holds a value in r6 across a
> call to a static subprog that calls bpf_unwind(). This applies on x86-64
> with ORC after the x86 enablement patch. Should the refusal cover
> might_unwind programs, not only programs with a table?

Yes, we should refuse might_unwind programs with private stack.

>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index 677bab92f624..f2a59716fdab 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
>> @@ -22781,6 +22781,9 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
>>   		/* program is valid, convert *(u32*)(ctx + off) accesses */
>>   		ret = bpf_convert_ctx_accesses(env);
>>
>> +	if (ret == 0)
>> +		ret = bpf_exc_keep_exits(env);
> Do the exits this patch adds skip a struct_ops program's BPF-level
> epilogue?
>
> bpf_convert_ctx_accesses() is the only place that epilogue is inserted:
> every BPF_EXIT in the main prog is replaced with the epilogue or a jump to
> it (fixups.c, "Generate epilogue for the main prog"). But
> bpf_exc_keep_exits() runs after that pass and emits a plain
> `r0 = 0; exit` after every bpf_unwind() call. bpf_do_misc_fixups() runs
> later still and lowers each bpf_unwind_resume() to a plain `r0 = 0; exit`
> too. In the main prog, both of these exits go straight to the JIT's native
> epilogue. So does the `*ra = prog->aux->epilogue_ip` redirect in
> bpf_unwind_rewrite().
>
> Does this break bpf_qdisc? bpf_qdisc_gen_epilogue()
> (net/sched/bpf_qdisc.c) adds a call to bpf_qdisc_reset_destroy_epilogue()
> to every .reset and .destroy program. That call is the only thing that runs
> qdisc_watchdog_cancel(&q->watchdog), and bpf_qdisc_validate() requires
> .reset and .destroy to be BPF programs for exactly that reason.
> bpf_unwind() is in generic_btf_ids, which is registered for
> BPF_PROG_TYPE_STRUCT_OPS with no filter. With no cleanup table,
> process_bpf_unwind() simply returns PROCESS_BPF_EXIT, so no gating applies.
>
> Failure path:
>
> 1. A .destroy program runs `if (cond) bpf_unwind();`.
>
> 2. On any arch, bpf_unwind() either leaves the first frame's return address
>     alone (bpf_unwind_rewrite(), cnt == 1) or does nothing (the weak
>     arch_bpf_stack_walk_ra()). Either way the frame returns through the
>     `r0 = 0; exit` that bpf_exc_keep_exits() inserted, which is not the
>     gen_epilogue.
>
> 3. qdisc_watchdog_cancel() never runs, so an hrtimer that .enqueue armed
>     through bpf_qdisc_watchdog_schedule() stays armed inside
>     qdisc_priv(sch).
>
> 4. The qdisc is freed, the timer fires, and qdisc_watchdog()
>     (net/sched/sch_api.c) dereferences wd->qdisc in freed memory. The
>     hrtimer base also still links the freed timer.
>
> The same bypass happens when a .reset/.destroy main prog resumes from a
> landing pad (lowered resume), or when a covered or uncovered subprog unwinds
> and the main frame is sent to epilogue_ip.
>
> Should bpf_unwind()/cleanup tables be rejected for programs whose
> verifier_ops has a gen_epilogue, or should these exits be inserted before
> bpf_convert_ctx_accesses() and point epilogue_ip at the BPF-level epilogue?

Indeed, for such cases, we should disable cleanup exceptions to avoid
early return for qdisc epilogue.

>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time
  2026-09-27 20:39   ` bot+bpf-ci
@ 2026-09-28  3:10     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-28  3:10 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 9/27/26 1:39 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git arch/x86/net/bpf_jit_comp.c arch/x86/net/bpf_jit_comp.c
>> index 6c7a0578760e..fb7e8ca1aab2 100644
>> --- arch/x86/net/bpf_jit_comp.c
>> +++ arch/x86/net/bpf_jit_comp.c
>> @@ -2150,7 +2150,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
>>   			dst_reg = X86_REG_R9;
>>   	}
>>
>> -	if (bpf_insn_is_indirect_target(env, bpf_prog, i - 1))
>> +	if (bpf_insn_is_indirect_target(env, bpf_prog, i - 1) ||
>> +	    bpf_exc_insn_is_pad(env, bpf_prog, i - 1))
>>   		EMIT_ENDBR();
> Landing pads are only ever entered by a ret whose stack slot
> bpf_unwind_rewrite() rewrote. Nothing jumps or calls to a pad indirectly
> since bpf_unwind_resume() is lowered to 'r0 = 0; exit', and
> bpf_exc_check_insn() forbids gotox inside a pad. IBT only checks
> indirect JMP/CALL targets, not RET targets, so the ENDBR at each pad
> head should not be needed.

Yes, you are right. ENDBR is not needed here.

>
> The same patch sends pad-less frames to aux->epilogue_ip through the
> same kind of rewritten ret, and emits no ENDBR there. That shows the pad
> ENDBR does no functional work. Each extra ENDBR in the middle of a JITed
> function is one more valid IBT landing site in executable kernel memory,
> and the kernel normally keeps those to a minimum (objtool seals unused
> ENDBRs).
>
> Is there an indirect-branch path to the pad that is not visible here?
>
>>   	ip = image + addrs[i - 1] + (prog - temp);
> [ ... ]
>
>> @@ -3278,6 +3279,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
>>   		seen_exit = true;
>>   		/* Update cleanup_addr */
>>   		ctx->cleanup_addr = proglen;
>> +		/* Where an unwind sends a frame with no pad. */
>> +		bpf_prog->aux->epilogue_ip = (u64)image + proglen;
> Eduard Zingerman raised an optimization suggestion in v5: use register
> liveness information at landing pads to avoid spilling all registers
> unnecessarily. The suggestion was to use register liveness information
> at call sites and collect a union of registers alive at landing pads for
> all callers of this particular function.
>
> You acknowledged and said you would try to fix it to avoid saving
> unnecessary callee saved registers. However, the current v6
> implementation still appears to unconditionally spill registers for
> exception_boundary programs without optimizing based on which registers
> are actually live at landing pads.
>
> Note: v6 significantly changed the implementation approach from v5, so
> this suggestion may need re-evaluation in the context of the new design.
> Does this optimization still apply, or does the new design address this
> differently?

Alexei suggested a new approach in v6.

>
>>   		if (bpf_prog_was_classic(bpf_prog) &&
>>   		    !ns_capable_noaudit(&init_user_ns, CAP_SYS_ADMIN)) {
>>   			if (emit_spectre_bhb_barrier(&prog, ip, bpf_prog))
> [ ... ]
>
>> @@ -4614,6 +4629,24 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
>>   #endif
>>   }
>>
>> +void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
>> +			    void *cookie)
>> +{
>> +#if defined(CONFIG_UNWINDER_ORC)
>> +	struct unwind_state state;
>> +	unsigned long addr, *ra;
>> +
>> +	for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
>> +	     unwind_next_frame(&state)) {
>> +		addr = unwind_get_return_address(&state);
>> +		ra = unwind_get_return_address_ptr(&state);
>> +		if (!addr || !ra ||
>> +		    !consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
>> +			break;
>> +	}
>> +#endif
>> +}
> arch_bpf_stack_walk_ra() pairs the unwinder's recovered return address
> with the raw stack slot. On x86 ORC, unwind_next_frame() does
>
> arch/x86/kernel/unwind_orc.c:unwind_next_frame() {
>      ...
>      state->ip = unwind_recover_ret_addr(state, state->ip, (unsigned long *)ip_p);
>      ...
> }
>
> which turns return_to_handler (function_graph / fprobe exit on fgraph)
> or the rethook trampoline (kretprobe) back into the original caller
> address. But unwind_get_return_address_ptr() returns `(unsigned long
> *)state->sp - 1`, i.e. ip_p itself, and that slot still holds the
> tracer's trampoline address. Nothing checks that *ra == addr before
> handing ra to the consumer.
>
> bpf_unwind() is a normal __bpf_kfunc in kernel/bpf/helpers.c with no
> notrace, so function_graph, kretprobe:bpf_unwind, or a kprobe.multi
> return probe can hook its return. In that case the prologue replaces the
> slot holding bpf_unwind's return into the BPF program. Then:
>
>    bpf_unwind() -> arch_bpf_stack_walk_ra() -> bpf_unwind_rewrite()
>      cnt == 1 frame, and the bpf_unwind call is covered by a cleanup record
>      (mark_call_sites() marks bpf_is_unwind_kfunc() calls)
>      -> *ra = rec->pad;   /* overwrites return_to_handler / rethook trampoline */
>
> bpf_unwind() then returns straight into the pad and skips the tracer's
> return trampoline. Its fgraph ret_stack entry (or rethook node) is left
> behind. The next hooked function further up that returns through the
> trampoline pops that stale entry.
>
> For fgraph, ftrace_pop_return_trace() has no frame-pointer check on x86
> (HAVE_FUNCTION_GRAPH_FP_TEST is not defined there). For rethook,
> __rethook_find_ret_addr() picks the top node. Either way execution jumps
> to the old BPF address after the bpf_unwind call, running on a different
> function's stack, which leads to an oops or arbitrary control flow.
>
> The arm64 version of this same walker, added later in the series
> (8bd2af30a28e "bpf, arm64: Dispatch exception cleanup pads at run
> time"), handles exactly this case:
>
> arch/arm64/net/bpf_jit_comp.c:arch_bpf_stack_walk_ra() {
>      ...
>      if (state->flags.fgraph || state->flags.kretprobe)
>          return false;
>      ...
> }
>
> Its changelog says: "a frame whose return the function graph tracer or a
> kretprobe has hooked holds the tracer's trampoline in its slot rather
> than the address the unwinder reports, so the walk stops there." The x86
> walker has no such check, and no later commit in the series touches
> arch/x86/net/bpf_jit_comp.c.
>
> Should the x86 version also check state flags or compare
> READ_ONCE_NOCHECK(*ra) against the value the unwinder read, and stop the
> walk when they differ?

Okay, will do this since arm64 has a different guard.

>
>> +
>>   void bpf_arch_poke_desc_update(struct bpf_jit_poke_descriptor *poke,
>>   			       struct bpf_prog *new, struct bpf_prog *old)
>>   {
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 15/21] libbpf: Collect .bpf_cleanup records and pass them to the kernel
  2026-09-27 20:39   ` bot+bpf-ci
@ 2026-09-28  3:28     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-28  3:28 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 9/27/26 1:39 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
>> index 7eb15a013a83..4e9584fff264 100644
>> --- a/tools/lib/bpf/libbpf.c
>> +++ b/tools/lib/bpf/libbpf.c
> [ ... ]
>
>> @@ -4915,6 +4953,249 @@ static struct bpf_program *find_prog_by_sec_insn(const struct bpf_object *obj,
>>   	return NULL;
>>   }
>>
>> +static int bpf_object__init_cleanup_info(struct bpf_object *obj)
>> +{
>> +	Elf_Data *data = obj->efile.cleanup_data;
>> +	Elf_Data *relo = NULL;
>> +	size_t i, nrels, nslots, nrecs;
>> +	struct cleanup_raw_rec *recs;
>> +	int *slot_sec, ret = 0;
>> +	size_t *slot_val;
>> +	const __u32 *vals;
>> +	Elf64_Shdr *sh;
>> +	bool native;
>> +
>> +	if (!data || obj->efile.cleanup_shndx < 0 || !data->d_size)
>> +		return 0;
>> +
>> +	native = is_native_endianness(obj);
>> +
>> +	for (i = 0; i < obj->efile.sec_cnt; i++) {
>> +		struct elf_sec_desc *sd = &obj->efile.secs[i];
>> +
>> +		if (sd->sec_type == SEC_RELO && sd->shdr &&
>> +		    sd->shdr->sh_info == (Elf64_Word)obj->efile.cleanup_shndx) {
>> +			relo = sd->data;
>> +			break;
>> +		}
>> +	}
>> +	if (!relo) {
>> +		pr_warn("%s present without relocations\n", CLEANUP_SEC);
>> +		return -LIBBPF_ERRNO__FORMAT;
>> +	}
>> +	/*
>> +	 * LLVM leaves sh_entsize unset, so this only bites a producer that
>> +	 * declares a record size -- which is the one able to tell us it means
>> +	 * something other than three 4-byte fields.
>> +	 */
>> +	sh = elf_sec_hdr(obj, elf_sec_by_idx(obj, obj->efile.cleanup_shndx));
>> +	if (sh && sh->sh_entsize &&
>> +	    sh->sh_entsize != sizeof(struct bpf_cleanup_info)) {
>> +		pr_warn("%s record size %zu is not the expected %zu\n", CLEANUP_SEC,
>> +			(size_t)sh->sh_entsize, sizeof(struct bpf_cleanup_info));
>> +		return -LIBBPF_ERRNO__FORMAT;
>> +	}
> Should sh_entsize == 0 be explicitly treated as 12, so that wider
> records (if necessary in the future) end up being easier to reject?
>
> This design question was raised by Kumar Kartikeya Dwivedi in v5:
> https://lore.kernel.org/bpf/DLMGMYOHSGFW.24JML4GSMTV2K@gmail.com/
>
> The current code accepts sh_entsize == 0 (LLVM's default) silently, but
> treating it explicitly as 12 would allow future tools to cleanly reject
> wider records by setting sh_entsize to a different value.

Okay. This is not a bug, but rather to have better code to have
explicit 'entsize' size and then check against size of bpf_cleanup_info.

>
>> +	if (data->d_size % sizeof(struct bpf_cleanup_info)) {
>> +		pr_warn("%s size %zu is not a multiple of the record size %zu\n",
>> +			CLEANUP_SEC, data->d_size, sizeof(struct bpf_cleanup_info));
>> +		return -LIBBPF_ERRNO__FORMAT;
>> +	}
>> +
>> +	vals = data->d_buf;
>> +	nslots = data->d_size / sizeof(__u32);
>> +	nrecs = data->d_size / sizeof(struct bpf_cleanup_info);
>> +
>> +	slot_sec = calloc(nslots, sizeof(*slot_sec));
>> +	slot_val = calloc(nslots, sizeof(*slot_val));
>> +	recs = calloc(nrecs ?: 1, sizeof(*recs));
> This isn't a bug, but could the relocation loop fill recs directly
> (with sec_idx initialised to -1) so the slot_sec/slot_val scratch
> arrays and the copy loop go away?

Good point. This will make things easier.

>
> The relocation loop could write straight into recs, using sec_idx = -1
> as the 'no relocation' sentinel and checking alignment while the byte
> offset is at hand, the way bpf_object__collect_rodata_relos() checks
> (sym->st_value + addend) % BPF_INSN_SZ inline. That would leave one
> allocation instead of three and drop the copy loop.
>
> Also, is the 'nrecs ?: 1' needed? d_size has already been checked to be
> a non-zero multiple of the record size.

Will do.

>
>> +	if (!slot_sec || !slot_val || !recs) {
>> +		ret = -ENOMEM;
>> +		goto out;
>> +	}
>> +	for (i = 0; i < nslots; i++)
>> +		slot_sec[i] = -1;
>> +
>> +	/* One relocation per 4-byte field, naming the section it points into. */
>> +	nrels = relo->d_size / sizeof(Elf64_Rel);
>> +	for (i = 0; i < nrels; i++) {
> [ ... ]
>
>> +	for (i = 0; i < nslots; i++) {
>> +		struct cleanup_raw_rec *rec = &recs[i / CLEANUP_REC_FIELDS];
>> +		size_t field = i % CLEANUP_REC_FIELDS;
>> +
>> +		if (slot_sec[i] < 0) {
>> +			pr_warn("%s: field %zu has no relocation\n", CLEANUP_SEC, i);
>> +			ret = -LIBBPF_ERRNO__FORMAT;
>> +			goto out;
>> +		}
>> +		if (slot_val[i] % BPF_INSN_SZ) {
>> +			pr_warn("%s: field %zu offset %zu is not instruction aligned\n",
>> +				CLEANUP_SEC, i, slot_val[i]);
>> +			ret = -LIBBPF_ERRNO__FORMAT;
>> +			goto out;
>> +		}
>> +		rec->sec_idx[field] = slot_sec[i];
>> +		rec->insn_idx[field] = slot_val[i] / BPF_INSN_SZ;
>> +	}
>> +
>> +	obj->cleanup_recs = recs;
>> +	obj->cleanup_rec_cnt = nrecs;
>> +	recs = NULL;
>> +out:
>> +	free(recs);
>> +	free(slot_val);
>> +	free(slot_sec);
>> +	return ret;
>> +}
> [ ... ]
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags
  2026-09-27 20:24   ` bot+bpf-ci
@ 2026-09-28  3:36     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-28  3:36 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 9/27/26 1:24 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/tools/testing/selftests/bpf/progs/bpf_misc.h b/tools/testing/selftests/bpf/progs/bpf_misc.h
>> index f3dbc3b59bff..9afa163fac5a 100644
>> --- a/tools/testing/selftests/bpf/progs/bpf_misc.h
>> +++ b/tools/testing/selftests/bpf/progs/bpf_misc.h
> [ ... ]
>
>> diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
>> index 25eeb1c1248b..65b55f7fa907 100644
>> --- a/tools/testing/selftests/bpf/test_loader.c
>> +++ b/tools/testing/selftests/bpf/test_loader.c
> [ ... ]
>
>> @@ -311,6 +326,142 @@ static int parse_caps(const char *str, __u64 *val, const char *name)
>>   	return 0;
>>   }
>>
>> +static int access_global_var(struct bpf_object *obj, const char *name,
>> +			     __u64 *val, bool set)
>> +{
>> +	__u32 off, sz, zero = 0;
>> +	struct bpf_map *map;
>> +	size_t vsz;
>> +	void *buf;
>> +	int err;
>> +
>> +	err = find_global_var(obj, name, &map, &off, &sz);
>> +	if (err)
>> +		return err;
>> +
>> +	vsz = bpf_map__value_size(map);
>> +	buf = calloc(1, vsz);
>> +	if (!buf)
>> +		return -ENOMEM;
>> +
>> +	err = bpf_map__lookup_elem(map, &zero, sizeof(zero), buf, vsz, 0);
>> +	if (err) {
>> +		PRINT_FAIL("failed to read '%s': %d\n", name, err);
>> +		goto out;
>> +	}
>> +	if (!set) {
>> +		*val = sz == 4 ? *(__u32 *)(buf + off) : *(__u64 *)(buf + off);
>> +		goto out;
>> +	}
>>                         ^^^^
> Can __ret_global() match a negative expected value for a 4-byte variable?
>
> parse_global_var() uses strtoull(), which takes a leading '-' and negates
> it as a 64-bit unsigned value. So __ret_global(err, -22) stores
> 0xffffffffffffffea (64 bits).
>
> But access_global_var() reads the 4-byte variable zero-extended:
>
>      *val = sz == 4 ? *(__u32 *)(buf + off) : *(__u64 *)(buf + off);
>
> and gets 0x00000000ffffffea (32 bits extended to 64). Then run_subtest()
> compares the two as __u64:
>
>      if (v != subspec->ret_global_val) {
>          PRINT_FAIL("Unexpected %s: 0x%llx != 0x%llx\n", ...);
>
> so the check always fails with "Unexpected err: 0xffffffea !=
> 0xffffffffffffffea", even though the program stored exactly -22.
>
> This matters in practice. Existing selftests keep negative errnos in
> 4-byte globals and check them from a hand-written driver, e.g.
> progs/test_bpf_nf.c has `int test_einval_reserved = 0;` checked in
> prog_tests/bpf_nf.c with `ASSERT_EQ(skel->bss->test_einval_reserved,
> -EINVAL, ...)`. The commit message says these tags are meant to replace
> that kind of driver. The neighbouring __retval() also supports negative
> values (parse_int() uses strtol() with a special case for INT_MIN), and 68
> programs under progs/ use __retval(-...).
>
> The write side does not have this problem. __set_global(err, -22)
> truncates to 0xffffffea and stores -22 correctly, so the same literal
> behaves differently in the two tags. 8-byte variables are fine because
> both sides are 64-bit.
>
> Should the expected value be truncated or sign-extended to the variable
> width (for example, compare only the low 32 bits when sz == 4, or
> sign-extend the value read back), or should out-of-range values be refused
> at parse time?

Right. Missed 4-byte value and also missed signedness. Will fix as commit
message promised.

>
> A related issue: __set_global() on a 4-byte variable silently drops any
> bits above 32:
>
>> +	if (sz == 4)
>> +		*(__u32 *)(buf + off) = *val;
> with no error. The tree allows it, but none of the tests in this series
> trigger it: they only use __u64 variables with non-negative values.
>
>> +	else
>> +		*(__u64 *)(buf + off) = *val;
>> +	err = bpf_map__update_elem(map, &zero, sizeof(zero), buf, vsz, 0);
>> +	if (err)
>> +		PRINT_FAIL("failed to write '%s': %d\n", name, err);
>> +out:
>> +	free(buf);
>> +	return err;
>> +}
> [ ... ]
>
>> @@ -1532,6 +1720,13 @@ void run_subtest(struct test_loader *tester,
>>   			}
>>   		}
>>
>> +		if (subspec->has_set_global) {
>> +			__u64 v = subspec->set_global_val;
>> +
>> +			if (access_global_var(tobj, subspec->set_global_var, &v, true))
>> +				goto tobj_cleanup;
>> +		}
>> +
>>   		err = do_prog_test_run(bpf_program__fd(tprog), &retval,
>>   				       bpf_program__type(tprog) == BPF_PROG_TYPE_SYSCALL ? true : false,
>>   				       spec->linear_sz);
>> @@ -1540,6 +1735,18 @@ void run_subtest(struct test_loader *tester,
>>   			goto tobj_cleanup;
>>   		}
>>
>> +		if (subspec->has_ret_global) {
>> +			__u64 v = 0;
>> +
>> +			if (access_global_var(tobj, subspec->ret_global_var, &v, false))
>> +				goto tobj_cleanup;
>> +			if (v != subspec->ret_global_val) {
>> +				PRINT_FAIL("Unexpected %s: 0x%llx != 0x%llx\n",
>> +					   subspec->ret_global_var, v, subspec->ret_global_val);
>> +				goto tobj_cleanup;
>> +			}
>> +		}
>>                     ^^^^
> For a 4-byte variable, __ret_global() cannot match any expected value
> outside [0, UINT32_MAX], and that range includes every negative literal.
>
> parse_global_var() parses the value with strtoull(), so XSTR(-1) ("-1")
> becomes 0xffffffffffffffff. access_global_var() reads a 4-byte variable as
>
>      *val = sz == 4 ? *(__u32 *)(buf + off) : *(__u64 *)(buf + off);
>
> which zero-extends it to 0x00000000ffffffff. The comparison here then
> always fails for `int cnt; ... __ret_global(cnt, -1)`, printing
> "Unexpected cnt: 0xffffffff != 0xffffffffffffffff" even when the program
> set cnt to -1.
>
> __set_global() does not have this problem: the same value truncated into a
> __u32 stores -1 correctly. So the two tags do not agree on what a 4-byte
> value is.
>
> Values too wide for the variable are also not rejected:
> __set_global(int_var, 0x100000000) quietly writes 0.
>
> The commit message says 4-byte variables are supported and cites veristat
> as the model, but veristat's set_global_var() checks signedness and range
> ("Variable %s value %lld is out of range [%lld; %lld]").
>
> Should the expected value be truncated or sign-adjusted to the variable
> width (for example, compare only the low 32 bits when sz == 4), or should
> out-of-range values be refused at parse time?

Yes, should do similar thing to veristat to have sz == 4 support
and signedness support.

>
> No test in the series hits this today; they all use the __u64 pads_ran
> with positive masks.
>
>> +
>>   		verify_stderr(bpf_program__fd(tprog), &subspec->stderr);
>>
>>   		if (subspec->stdout.cnt) {
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430


^ permalink raw reply	[flat|nested] 56+ messages in thread

* Re: [PATCH bpf-next v6 20/21] selftests/bpf: Cover the exception cleanup shapes the chain does not reach
  2026-09-27 20:40   ` bot+bpf-ci
@ 2026-09-28  3:49     ` Yonghong Song
  0 siblings, 0 replies; 56+ messages in thread
From: Yonghong Song @ 2026-09-28  3:49 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 9/27/26 1:40 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
>> index d1d40314035e..49b6806cde82 100644
>> --- a/tools/testing/selftests/bpf/exceptions_cleanup.h
>> +++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
> [ ... ]
>
>> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
>> new file mode 100644
>> index 000000000000..37a78e035a50
>> --- /dev/null
>> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
>> @@ -0,0 +1,665 @@
>> +// SPDX-License-Identifier: GPL-2.0
>> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
>> +#include <vmlinux.h>
>> +#include <bpf/bpf_helpers.h>
>> +#include "bpf_misc.h"
>> +#include "exceptions_cleanup.h"
> [ ... ]
>
>> +/* The callee most of the shapes below unwind out of. */
>> +static __used __noinline __u64 pc_unwinder(__u64 x)
>> +{
>> +	if (x > 100)
>> +		bpf_unwind();
>> +	return x + 1;
>> +}
> [ ... ]
>
>> +/* A callee on both a covered and an uncovered call: the pad is the site's. */
>> +static __used __noinline __u64 shared_callee(__u64 x)
>> +{
>> +	if (x > 100)
>> +		bpf_unwind();
>> +	return x + 1;
>> +}
> This isn't a bug, but shared_callee() looks the same as pc_unwinder().
> Could the shared-callee shape call pc_unwinder() from both sites, or
> could wide_rec_frame() use pc_unwinder() so that shared_callee() is left
> to the one shape its comment describes? And should the 'most of the
> shapes' comment on pc_unwinder() be narrowed to match its three callers?

Okay, will remore shared_callee and just use pc_unwinder() then.

>
> [ ... ]
>
>> +/*
>> + * A covered bpf_unwind() the sweep leaves last, behind the exception callback
>> + * patchlet, which has to carry the marks with it.
>> + */
>> +SEC("?syscall")
>> +__success __set_global(input, 101) __retval(0)
>> +__ret_global(pads_ran, RAN_PAD_FIRST)
>> +__naked int entry_pad_first(void)
>> +{
> This isn't a bug, but is 'exception callback patchlet' meant to be the
> exit that bpf_exc_keep_exits() puts back after bpf_unwind()? If so,
> could the comment name that instead? This program never sets up an
> exception callback.

This is a mistake for comments. Will fix.

>
> [ ... ]
>
>> +/* A pad that indexes its frame by a register set before the unwinding call. */
>> +
>> +/* An unwinder that touches none of r6-r9, so the walk leaves this frame. */
>> +static __used __naked __noinline __u64 var_unwinder(void)
>> +{
>> +	asm volatile (
>> +	"if r1 < 101 goto 1f;"
>> +	"call bpf_unwind;"
>> +"1:"
>> +	"r0 = 0;"
>> +	"exit;"
>> +	::: __clobber_all);
>> +}
>> +
>> +static __used __naked __noinline __u64 var_stack_frame(void)
>> +{
>> +	asm volatile (
>> +	"r1 = %[magic] ll;"
>> +	"r1 = *(u64 *)(r1 + 0);"
>> +	"*(u64 *)(r10 - 8) = r1;"	/* the slot the pad will read... */
>> +	"*(u64 *)(r10 - 16) = r1;"	/* ...whichever of the two it is */
>> +	"r1 = %[input] ll;"
>> +	"r6 = *(u64 *)(r1 + 0);"
>> +	"r6 &= 1;"			/* an unknown slot number... */
>> +	"r6 <<= 3;"			/* ...as an aligned byte offset */
>> +	"r1 = %[input] ll;"
>> +	"r1 = *(u64 *)(r1 + 0);"
>> +"1:"	"call var_unwinder;"		/* cleanup region */
>> +"2:"
>> +	"r0 = 0;"
>> +	"exit;"
>> +"3:"					/* landing pad */
>> +	"r7 = r0;"
>> +	"r1 = r10;"
>> +	"r1 += r6;"			/* variable offset into the frame */
>> +	"r2 = *(u64 *)(r1 - 16);"
>> +	"r3 = %[magic] ll;"
>> +	"r3 = *(u64 *)(r3 + 0);"
>> +	"if r2 != r3 goto 9f;"
>> +	PAD_RAN("%[ran]")
>> +"9:"
>> +	"r1 = r7;"
>> +	"call bpf_unwind_resume;"
>> +	"exit;"
>> +	CLEANUP_REC("1b", "2b", "3b")
>> +	:
>> +	: [ran]"i"(RAN_VAR_STACK), __imm_addr(input), __imm_addr(magic),
>> +	  __imm_addr(pads_ran)
>> +	: __clobber_all);
>> +}
> [ ... ]
>
>> +/* The same over a global subprogram, which the verifier enters no frame for. */
>> +__noinline __u64 global_unwinder(__u64 x)
>> +{
>> +	if (x > 100)
>> +		bpf_unwind();
>> +	return x + 1;
>> +}
>> +
>> +static __used __naked __noinline __u64 global_pad_frame(void)
>> +{
>> +	asm volatile (
>> +	"r1 = %[magic] ll;"
>> +	"r1 = *(u64 *)(r1 + 0);"
>> +	"*(u64 *)(r10 - 8) = r1;"
>> +	"*(u64 *)(r10 - 16) = r1;"
>> +	"r1 = %[input] ll;"
>> +	"r6 = *(u64 *)(r1 + 0);"
>> +	"r6 &= 1;"
>> +	"r6 <<= 3;"
>> +	"r1 = %[input] ll;"
>> +	"r1 = *(u64 *)(r1 + 0);"
>> +"1:"	"call global_unwinder;"		/* cleanup region */
>> +"2:"
>> +	"r0 = 0;"
>> +	"exit;"
>> +"3:"					/* landing pad */
>> +	"r7 = r0;"
>> +	"r1 = r10;"
>> +	"r1 += r6;"
>> +	"r2 = *(u64 *)(r1 - 16);"
>> +	"r3 = %[magic] ll;"
>> +	"r3 = *(u64 *)(r3 + 0);"
>> +	"if r2 != r3 goto 9f;"
>> +	PAD_RAN("%[ran]")
>> +"9:"
>> +	"r1 = r7;"
>> +	"call bpf_unwind_resume;"
>> +	"exit;"
>> +	CLEANUP_REC("1b", "2b", "3b")
>> +	:
>> +	: [ran]"i"(RAN_GLOBAL_PAD), __imm_addr(input), __imm_addr(magic),
>> +	  __imm_addr(pads_ran)
>> +	: __clobber_all);
>> +}
> This isn't a bug, but global_pad_frame() looks nearly identical to
> var_stack_frame(), and pad_r0_frame() repeats the same pad body. Could
> these share a macro, the way LOAD_MAGIC_REGS/CHECK_MAGIC_REGS are shared
> earlier in the file, so each variant only spells out the part it is
> testing?

Yes, it is better indeed.

>
> [ ... ]
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430


^ permalink raw reply	[flat|nested] 56+ messages in thread

end of thread, other threads:[~2026-09-28  3:49 UTC | newest]

Thread overview: 56+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 01/21] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 02/21] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 03/21] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 04/21] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-09-26  5:16   ` sashiko-bot
2026-09-26 23:54     ` Yonghong Song
2026-09-27 20:39   ` bot+bpf-ci
2026-09-28  0:01     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-26  5:21   ` sashiko-bot
2026-09-27  0:02     ` Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  0:12     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad Yonghong Song
2026-09-26  5:15   ` sashiko-bot
2026-09-26  8:21     ` Alexei Starovoitov
2026-09-27  0:04       ` Yonghong Song
2026-09-27  0:41     ` Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  0:17     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-09-26  5:17   ` sashiko-bot
2026-09-27  3:06     ` Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  0:29     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 09/21] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 10/21] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  1:08     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-26  5:15   ` sashiko-bot
2026-09-27  4:35     ` Yonghong Song
2026-09-27 20:39   ` bot+bpf-ci
2026-09-28  3:10     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 12/21] bpf, arm64: " Yonghong Song
2026-09-26  5:14   ` sashiko-bot
2026-09-27 20:40   ` bot+bpf-ci
2026-09-26  5:01 ` [PATCH bpf-next v6 13/21] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 14/21] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 15/21] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-27 20:39   ` bot+bpf-ci
2026-09-28  3:28     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 16/21] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 17/21] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 18/21] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-09-26  5:18   ` sashiko-bot
2026-09-27  4:58     ` Yonghong Song
2026-09-27 20:24   ` bot+bpf-ci
2026-09-28  3:36     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 20/21] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  3:49     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 21/21] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox