BPF List
 help / color / mirror / Atom feed
* [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds
@ 2026-10-01 13:30 Yonghong Song
  2026-10-01 13:30 ` [PATCH bpf-next v8 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
                   ` (21 more replies)
  0 siblings, 22 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:30 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

bpf_throw() walks the BPF call stack to the exception boundary and
discards every frame in between. A frame that owns something -- an RCU
read lock, a preemption-disabled section, a referenced kptr -- never gets
to give it back, so the verifier refuses to let such a frame throw at
all. That is the whole reason a Rust program cannot use bpf_throw() as
its panic path right now: Rust's Drop glue *is* that give-back, and there
is nowhere to run it.

LLVM 23 added the compiler half ([1]). A Rust function that owns a value
across a call that can unwind

    fn foo() {
        let _guard = RcuReadGuard::new();  /* bpf_rcu_read_lock()   */
        may_throw();                       /* extern "C-unwind"     */
    }                                      /* Drop: rcu_read_unlock */

lowers to an invoke with a cleanup landing pad holding the Drop call, and
the BPF backend writes one record per invoke region into a .bpf_cleanup
section: a flat table of 12-byte (begin, end, landing_pad) triples, each
field a byte offset into the code section. The rule is "a frame suspended
at a call in [begin, end) resumes at landing_pad when an unwind passes
through". A pad ends with a call to _Unwind_Resume(), which the kernel
provides as the bpf_unwind_resume() kfunc.

Rather than overload bpf_throw(), which keeps its own meaning -- leave for
the exception boundary with the frames in between discarded -- a new
bpf_unwind() kfunc raises the unwind this series dispatches.

This series is the kernel half: take that table at BPF_PROG_LOAD, teach
the verifier to follow an unwind to the landing pad it reaches, and have
bpf_unwind() send each frame it leaves there.

C has no unwinding, so the selftests spell out by hand what a frontend
emits -- a call site bracketed by two labels, a landing pad, and a record
tying them together. The frame above, written that way:

        "call bpf_rcu_read_lock;"
    "1:"    "call foo3;"                /* cleanup region */
    "2:"
        ... normal path, ends in bpf_rcu_read_unlock ...
    "6:"                                /* landing pad */
        "call bpf_rcu_read_unlock;"
        "call bpf_unwind_resume;"
        CLEANUP_REC("1b", "2b", "6b")

Design
======

A pad runs in the frame that owns it, entered by an ordinary return.
bpf_unwind() walks the frames with arch_bpf_stack_walk_ra(), which hands
out the slot each frame's return address came from, and rewrites that
slot for every frame above the one that called it: to the landing pad
where a record covers the call the address returns into, and to that
frame's epilogue where nothing does. Then it returns. The frame that
called it goes on after the call, where the fixups put 'r0 = 0' and then
a jump to its own pad, or an exit where no record covers the call.

Each frame therefore runs its own pad, on its own stack, and leaves
through its own epilogue -- which is what puts its caller's r6-r9 back. The
unwind needs no trampoline, no spill area and no per-frame metadata beyond
the table itself. The fixups lower a pad's bpf_unwind_resume() to
'r0 = 0; exit', so the frame returns and the address rewritten below it
carries the unwind on to the next pad. For main -> A -> B -> C with only
A's call to B covered, an unwind raised in C runs:

  frame   runs
  -----   ----------------------------------------------------------
  C       'r0 = 0; exit', patched in after its bpf_unwind()
  B       its epilogue, putting back A's r6-r9
  A       its pad, which drops A's resources; the resume is an exit
  main    its epilogue, returning 0 to the kernel

The verifier follows the same path. A bpf_unwind() goes on at its own
frame's pad, or pops frames to the first pad that covers a call, or to the
main program's exit, where nothing may still be held and the zero returned
has to suit the program type. A resume goes on below its frame the same
way. The pad is entered from the state the unwind leaves, not from a
snapshot at the call: before unwinding, a callee may have written its
caller's stack through a pointer, overwritten a spilled pointer, or
changed packet data, and the callee's epilogue puts back none of that. A
global subprogram is verified on its own, so a call to one that can unwind
gets a second successor, the unwind, taken from the state the call
returns in. Precision backtracking follows the same edges back, into the
frame the unwind left.

  1      pack bpf_insn_aux_data's flags into bit fields, so the flags this
         series adds cost bits rather than bytes
  2-3    uapi: cleanup_info in BPF_PROG_LOAD and struct bpf_cleanup_info;
         the bpf_unwind() and bpf_unwind_resume() kfuncs
  4-6    mark the covered call sites, refuse what cannot carry a table, and
         make the landing pads reachable in the CFG and in liveness
  7      verifier: follow an unwind to its landing pad
  8      require an unwind to leave a frame holding what it entered with
  9-10   refuse a pad that does not resume; no private stack for a program
         that can unwind
  11     bpf_unwind(): rewrite the return addresses as it walks
  12-13  x86-64 and arm64 JITs
  14-18  libbpf: resolve _Unwind_Resume, collect .bpf_cleanup and pass it
         to the kernel, carry it through the light skeleton and the static
         linker
  19-22  selftests

Limitations
===========

  - Cleanup pads only. A catch pad, which Rust's catch_unwind would need,
    is refused: bpf_unwind() rewrites every frame in one pass.
  - Needs a JIT that dispatches pads: x86-64 with CONFIG_UNWINDER_ORC, and
    arm64. Elsewhere the load fails with -EOPNOTSUPP.
  - Not with offload, a private stack, a verifier_ops epilogue, bpf_throw()
    or an exception callback.
  - In a pad's frame: no tail call, BPF_LD_[ABS|IND] or indirect jump, and
    no second unwind in the pad or anything it calls.
  - A subprogram that can unwind, including one with a callx, cannot be a
    callback.
  - A frame an unwind leaves must hold what it held on entry, and must have
    an exit.
  - A return hooked by fgraph or a kretprobe stops the walk with a warning.
  - Until the Rust toolchain supports BPF exception handling, the selftests
    write their tables by hand in inline asm.

  [1] https://github.com/llvm/llvm-project/pull/192164
      llvm commit 9d51c891b719 ("[BPF] Add exception handling support
      with .bpf_cleanup section")

Changelog
=========
  v7 -> v8:
    - v7: https://lore.kernel.org/bpf/20260929001601.3242665-1-yonghong.song@linux.dev/
    - Follow an unwind to its pad from the state it leaves, not from a
      snapshot at the call; patch 7 is renamed to match.
    - Follow the unwind out of a call to a global subprogram as well: the
      call continues at the next insn and, from the state it returns in,
      at its pad or further out.
    - Backtrack precision across an unwind; clear r1-r5 at a main-frame
      unwind exit.
    - bpf_unwind() leaves its own caller alone; the fixups put 'r0 = 0' and
      a jump to its pad after the call.
    - Refuse bpf_throw() and an exception callback with any unwind, not
      only with a table.
    - Refuse a subprogram an unwind passes through that has no exit.
    - Patch 8's rule that a frame an unwind leaves holds what it entered
      with no longer keeps the pads safe; it stays to refuse a program at
      the frame at fault.
    - A speculative walk leaves no in-pad or outside-pad mark.
    - x86, arm64: warn on a hooked return. arm64: take the PAC modifier as
      record + 16, and allow a shadow call stack.
    - libbpf: narrow the _Unwind_Resume translation, close the opts hole,
      size the light skeleton's attr by need, and tighten the linker check.
  v6 -> v7:
    - v6: https://lore.kernel.org/bpf/20260926050006.2213110-1-yonghong.song@linux.dev/
    - New patch: require an unwind to leave a frame holding what it entered
      with, including at a call it passes through with no record over it.
    - A resume ends the verifier's path instead of walking back to the
      instruction after the call, where an unwind never returns.
    - Refuse a private stack for any program that can unwind, not only one
      carrying a table.
    - Refuse a table for a program whose verifier_ops plants an epilogue,
      and scan the whole instruction stream for bpf_throw().
    - Keep the last exit of every subprogram an unwind can pass through,
      not only the one after a bpf_unwind() call, and search for it.
    - Fix precision backtracking across a resume, push a landing pad by
      itself so the CFG walk's DFS invariant holds, and mark callx sites.
    - x86: leave a frame whose return a tracer has hooked alone; no ENDBR
      at a pad head, which is only ever reached by a return.
    - arm64: no BTI at a pad head; ask the build whether a return address
      is signed; refuse cleanup pads where a shadow call stack is in use.
    - Answer a speculative walk reaching a pad with a barrier, not a refusal.
    - selftests: add the negative shapes for the above, allow repeated
      __set_global()/__ret_global(), and drop duplicate accepted shapes.
  v5 -> v6:
    - v5: https://lore.kernel.org/bpf/20260923045846.2414643-1-yonghong.song@linux.dev/
    - Run a pad in the frame that owns it: bpf_unwind() rewrites each
      frame's saved return address rather than calling the pad as a
      subroutine of the walker. The bpf_cleanup_pad.S trampolines, the
      per-frame spill area and the pad-entry register header all go away.
    - Raise the unwind with a new bpf_unwind() kfunc, so that bpf_throw()
      keeps its meaning.
    - Add arch_bpf_stack_walk_ra(), which also hands out the slot a return
      address came from. arm64 re-signs the address it writes there.
    - A pad is an ordinary second successor of a covered call, so the
      verifier needs no unwind edge of its own: the cross-frame precision
      and liveness work is gone, with the v5 fixes it needed, and two
      patches become one.
    - Refuse the undispatchable shapes per instruction in do_check(), keyed
      on a per-frame mark, rather than by walking each pad.
    - Allow on-stack call arguments in a pad, which now has its own frame.
    - Add __set_global() and __ret_global() test tags, so RUN_TESTS() drives
      the shapes and test_shapes() goes away (suggested by Eduard).
    - Drop the shapes that needed a driver of their own, and the three
      extension objects with them.
  v4 -> v5:
    - v4: https://lore.kernel.org/bpf/20260921210033.1715000-1-yonghong.song@linux.dev/
    - Rebase on bpf-next.
    - Clear r0 where a throw enters a landing pad: no instruction defines
      it, so a precision request for it outlived the state and oopsed the
      verifier.
    - Defer entering the frames an unwind edge crossed until the backtrack
      reads an instruction from them; entering at the landing pad could
      leave bt->frame past the parent state's frames and oops the verifier.
    - Stamp the popped frame count inside bpf_push_jmp_history(), so a pad's
      entry carries it even when the prune path is what creates the entry.
    - Replace the hand-rolled CFG traversal and its separate pass with
      per-instruction checks in do_check(), keyed on the verifier's
      unwinding state; kernel/bpf/exception.c halves.
    - Add a first patch packing bpf_insn_aux_data's flags into one bit field
      word, 144 bytes to 128, so the flags this series adds cost bits.
    - Drop the per-subprogram arrays the JITs consulted for throw sites,
      resume sites and pad bodies, and read insn_aux_data, which a JIT
      already has.
    - Drop the pad-entry r0 header: r0 at a pad is an unknown scalar, and a
      dispatcher writes a defined value there only to keep a kernel one out
      of BPF.
    - Take the bool arguments back out of verifier_remove_insns() and the
      site collector, and share pop_frame() with prepare_func_exit().
    - Rename the recorded call sites to throw_call and resume_call, give the
      exported functions a bpf_exc_ prefix, and drop cleanup_ from the
      statics.
  v3 -> v4:
    - v3: https://lore.kernel.org/bpf/20260920054225.864535-1-yonghong.song@linux.dev/
    - Rebase on bpf-next due to conflict.
    - Reserve the throw-site spill area only in a (sub)program that calls
      bpf_throw(): 40 bytes of stack per frame on x86-64, 80 on arm64.
    - Bound the record count by the number of instructions in the program,
      and name both in the message.
    - Refuse a .bpf_cleanup section in libbpf whose record count cannot be
      handed to the kernel as a count times a record size in an int.
    - Fix the static linker's new bounds check, which could itself wrap, and
      refuse a section too small to hold one field.
    - WARN once if the body of bpf_unwind_resume() is ever reached, the way
      bpf_throw() does where its exception callback should never return.
    - Rename nr_pad_body to pad_body_bits, use BTF_ID_LIST_SINGLE, move
      cleanup_pad out of insn_aux_data's bools, drop an arm64 include.
  v2 -> v3:
    - v2: https://lore.kernel.org/bpf/20260918044156.3283973-1-yonghong.song@linux.dev/
    - Keep a landing pad's record when opt_remove_nops() deletes a pad that is
      a nop, instead of dropping it after the verifier has already checked the
      call site against it.
    - Refuse a BPF_LD_[ABS|IND] in a pad body.
    - Teach mark_chain_precision() about the throw-to-pad edge, which crosses
      frames with no instruction to account for them.
    - Refuse a bpf_unwind_resume() in any frame but the one whose landing pad
      the walker entered.
    - Check raw_data, alignment and bounds before the static linker writes
      through a relocation in a non-executable section, which may be SHT_NOBITS.
  v1 -> v2:
    - v1: https://lore.kernel.org/bpf/20260917055645.3926444-1-yonghong.song@linux.dev/
    - Consolidate all usages of kern_extern_name() in a single patch in libbpf.
    - Avoid compiler warning and add proper cleanup_info_cnt guard in libbpf when
      collecting .bpf_cleanup records.
    - Add cleanup_info_cnt condition for emit_rel_store() with cleanup_info.

Yonghong Song (22):
  bpf: Pack bpf_insn_aux_data flags into bit fields
  bpf: Accept the compiler's exception cleanup table at program load
  bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs
  bpf: Add lookups for exception cleanup resumes and landing pads
  bpf: Prepare for an exception cleanup table before the CFG walk
  bpf: Make exception landing pads reachable in the CFG
  bpf: Follow an unwind to its landing pad in the verifier
  bpf: Require an unwind to leave a frame holding what it entered with
  bpf: Refuse a landing pad that does not resume
  bpf: Do not use a private stack for a program that can unwind
  bpf: Dispatch cleanup pads by rewriting return addresses
  bpf, x86: Dispatch exception cleanup pads at run time
  bpf, arm64: Dispatch exception cleanup pads at run time
  libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
  libbpf: Add cleanup_info to bpf_prog_load_opts
  libbpf: Collect .bpf_cleanup records and pass them to the kernel
  libbpf: Carry the exception cleanup table through the light skeleton
  libbpf: Let the static linker carry .bpf_cleanup relocations
  selftests/bpf: Add end-to-end and negative .bpf_cleanup exception
    tests
  selftests/bpf: Add __set_global() and __ret_global() test tags
  selftests/bpf: Cover more accepted .bpf_cleanup exception shapes
  selftests/bpf: Load an exception cleanup program from a light skeleton

 arch/arm64/kernel/stacktrace.c                |  104 ++
 arch/arm64/net/bpf_jit_comp.c                 |   21 +
 arch/x86/net/bpf_jit_comp.c                   |   40 +
 include/linux/bpf.h                           |   37 +
 include/linux/bpf_verifier.h                  |   78 +-
 include/linux/filter.h                        |    3 +
 include/uapi/linux/bpf.h                      |   14 +
 kernel/bpf/Makefile                           |    2 +-
 kernel/bpf/backtrack.c                        |   49 +-
 kernel/bpf/cfg.c                              |  101 ++
 kernel/bpf/core.c                             |   25 +-
 kernel/bpf/exception.c                        |  490 ++++++++
 kernel/bpf/exception.h                        |   35 +
 kernel/bpf/fixups.c                           |  164 ++-
 kernel/bpf/helpers.c                          |   59 +
 kernel/bpf/liveness.c                         |   26 +-
 kernel/bpf/syscall.c                          |    2 +-
 kernel/bpf/verifier.c                         |  273 ++++-
 tools/include/uapi/linux/bpf.h                |   14 +
 tools/lib/bpf/bpf.c                           |    6 +-
 tools/lib/bpf/bpf.h                           |    7 +-
 tools/lib/bpf/gen_loader.c                    |   34 +-
 tools/lib/bpf/libbpf.c                        |  364 +++++-
 tools/lib/bpf/libbpf_internal.h               |   10 +
 tools/lib/bpf/linker.c                        |   37 +-
 tools/testing/selftests/bpf/Makefile.skel     |    2 +-
 .../selftests/bpf/exceptions_cleanup.h        |   45 +
 .../bpf/prog_tests/exceptions_cleanup.c       |  118 ++
 tools/testing/selftests/bpf/progs/bpf_misc.h  |   12 +
 .../selftests/bpf/progs/exceptions_cleanup.c  |  160 +++
 .../bpf/progs/exceptions_cleanup_fail.c       |  978 ++++++++++++++++
 .../bpf/progs/exceptions_cleanup_shapes.c     | 1003 +++++++++++++++++
 tools/testing/selftests/bpf/test_loader.c     |  319 +++++-
 33 files changed, 4573 insertions(+), 59 deletions(-)
 create mode 100644 kernel/bpf/exception.c
 create mode 100644 kernel/bpf/exception.h
 create mode 100644 tools/testing/selftests/bpf/exceptions_cleanup.h
 create mode 100644 tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup.c
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
@ 2026-10-01 13:30 ` Yonghong Song
  2026-10-01 13:30 ` [PATCH bpf-next v8 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
                   ` (20 subsequent siblings)
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:30 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

struct bpf_insn_aux_data spreads its flags over three places: eight bools a
byte each, a u8 carrying three bit fields, and a u32 group with 25 of its
32 bits spare. Put them all in one u64 word, alu_state with them, and move
orig_idx below it. The structure goes from 128 bytes to 120, with no byte
holes and 33 of the word's 64 bits spare. Later patches in this series take
two of those bits, and add a u32 of their own that puts the size back at
128.

No functional change: a one-bit unsigned field holds 0 and 1 the way the
bool did, and nothing takes the address of any of them.

Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf_verifier.h | 43 ++++++++++++++++++------------------
 1 file changed, 22 insertions(+), 21 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 811342e3c041..1d3130584662 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -664,41 +664,42 @@ struct bpf_insn_aux_data {
 	u64 map_key_state; /* constant (32 bit) key tracking for maps */
 	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
 	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
-	bool nospec; /* do not execute this instruction speculatively */
-	bool nospec_result; /* result is unsafe under speculation, nospec must follow */
-	bool zext_dst; /* this insn zero extends dst reg */
-	bool needs_zext; /* alu op needs to clear upper bits */
-	bool prevent_zext; /* alu op cannot be zext (already used with 64-bit scalars) */
-	bool non_sleepable; /* helper/kfunc may be called from non-sleepable context */
-	bool is_iter_next; /* bpf_iter_<type>_next() kfunc call */
-	bool call_with_percpu_alloc_ptr; /* {this,per}_cpu_ptr() with prog percpu alloc */
-	u8 alu_state; /* used in combination with alu_limit */
+	u64 nospec:1; /* do not execute this instruction speculatively */
+	u64 nospec_result:1; /* result is unsafe under speculation, nospec must follow */
+	u64 zext_dst:1; /* this insn zero extends dst reg */
+	u64 needs_zext:1; /* alu op needs to clear upper bits */
+	u64 prevent_zext:1; /* alu op cannot be zext (already used with 64-bit scalars) */
+	u64 non_sleepable:1; /* helper/kfunc may be called from non-sleepable context */
+	u64 is_iter_next:1; /* bpf_iter_<type>_next() kfunc call */
+	u64 call_with_percpu_alloc_ptr:1; /* {this,per}_cpu_ptr() with prog percpu alloc */
+	u64 alu_state:8; /* used in combination with alu_limit */
 	/* true if STX or LDX instruction is a part of a spill/fill
 	 * pattern for a bpf_fastcall call.
 	 */
-	u8 fastcall_pattern:1;
+	u64 fastcall_pattern:1;
 	/* for CALL instructions, a number of spill/fill pairs in the
 	 * bpf_fastcall pattern.
 	 */
-	u8 fastcall_spills_num:3;
-	u8 arg_prog:4;
+	u64 fastcall_spills_num:3;
+	u64 arg_prog:4;
 
-	/* below fields are initialized once */
-	unsigned int orig_idx; /* original instruction index */
-	u32 jmp_point:1;
-	u32 prune_point:1;
+	/* below flags are initialized once */
+	u64 jmp_point:1;
+	u64 prune_point:1;
 	/* ensure we check state equivalence and save state checkpoint and
 	 * this instruction, regardless of any heuristics
 	 */
-	u32 force_checkpoint:1;
+	u64 force_checkpoint:1;
 	/* true if instruction is a call to a helper function that
 	 * accepts callback function as a parameter.
 	 */
-	u32 calls_callback:1;
-	u32 indirect_target:1; /* if it is an indirect jump target */
-	u32 non_stack_access:1; /* instruction can access non-stack memory */
+	u64 calls_callback:1;
+	u64 indirect_target:1; /* if it is an indirect jump target */
+	u64 non_stack_access:1; /* instruction can access non-stack memory */
 	/* true if some jump or call instruction targets this instruction */
-	u32 jump_target:1;
+	u64 jump_target:1;
+
+	unsigned int orig_idx; /* original instruction index, initialized once */
 	/*
 	 * CFG strongly connected component this instruction belongs to,
 	 * zero if it is a singleton SCC.
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 02/22] bpf: Accept the compiler's exception cleanup table at program load
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
  2026-10-01 13:30 ` [PATCH bpf-next v8 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
@ 2026-10-01 13:30 ` Yonghong Song
  2026-10-01 13:30 ` [PATCH bpf-next v8 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
                   ` (19 subsequent siblings)
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:30 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

LLVM 23 added exception handling support for BPF with the .bpf_cleanup
section ([1]). Rust code compiled with panic=unwind runs cleanup code (Drop
glue) when an unwind passes through, and the LLVM BPF backend emits that
section from the landing pads the frontend produced. Plain C cannot
generate .bpf_cleanup unless inline asm is used. The Rust compiler does not
fully support BPF exception handling yet, but the kernel can support the
table today, and inline assembly is enough to test it.

Add the UAPI to carry the .bpf_cleanup table into the kernel. BPF_PROG_LOAD
grows cleanup_info, cleanup_info_cnt and cleanup_info_rec_size, and struct
bpf_cleanup_info describes one record as a triple of instruction indices:
the half-open call-site range [begin_off, end_off) and the landing_pad_off
the frame resumes at. The table arrives sorted by begin_off, with disjoint
ranges and each record's three offsets inside one subprogram;
bpf_exc_check_info(), in a new kernel/bpf/exception.c that the rest of the
exception code joins, holds it to that at load time. It also refuses a
landing pad inside any call-site range, its own included, an offset naming
the second half of an ld_imm64, and more records than the program has
instructions. Nothing reads the table yet; the patches that follow -- the
CFG walk, the unwind walk and the JITs -- are its consumers.

Link: https://github.com/llvm/llvm-project/pull/192164 [1]
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf_verifier.h   |   2 +
 include/uapi/linux/bpf.h       |  14 ++++
 kernel/bpf/Makefile            |   2 +-
 kernel/bpf/exception.c         | 137 +++++++++++++++++++++++++++++++++
 kernel/bpf/exception.h         |  15 ++++
 kernel/bpf/syscall.c           |   2 +-
 kernel/bpf/verifier.c          |   6 ++
 tools/include/uapi/linux/bpf.h |  14 ++++
 8 files changed, 190 insertions(+), 2 deletions(-)
 create mode 100644 kernel/bpf/exception.c
 create mode 100644 kernel/bpf/exception.h

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 1d3130584662..ad0ca8047712 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1019,6 +1019,8 @@ struct bpf_verifier_env {
 	struct spill_snapshot **callsite_at_stack;
 	u32 pass_cnt; /* number of times do_check() was called */
 	u32 subprog_cnt;
+	struct bpf_cleanup_info *cleanup_info;
+	u32 cleanup_info_cnt;
 	/* number of instructions analyzed by the verifier */
 	u32 prev_insn_processed, insn_processed;
 	/* number of jmps, calls, exits analyzed so far */
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 4687c3310996..033b5ed883db 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -1702,6 +1702,9 @@ union bpf_attr {
 		 * verification.
 		 */
 		__s32		keyring_id;
+		__aligned_u64	cleanup_info;	/* exception cleanup table */
+		__u32		cleanup_info_rec_size; /* userspace bpf_cleanup_info size */
+		__u32		cleanup_info_cnt; /* number of bpf_cleanup_info records */
 	};
 
 	struct { /* anonymous struct used by BPF_OBJ_* commands */
@@ -7638,6 +7641,17 @@ struct bpf_line_info {
 	__u32	line_col;
 };
 
+/*
+ * One record of an exception cleanup table: calls in [begin_off, end_off)
+ * that unwind resume at landing_pad_off. All three are instruction offsets
+ * in the program as loaded.
+ */
+struct bpf_cleanup_info {
+	__u32	begin_off;
+	__u32	end_off;
+	__u32	landing_pad_off;
+};
+
 struct bpf_spin_lock {
 	__u32	val;
 };
diff --git a/kernel/bpf/Makefile b/kernel/bpf/Makefile
index c1f9b0d3468d..8a6947b3d13a 100644
--- a/kernel/bpf/Makefile
+++ b/kernel/bpf/Makefile
@@ -11,7 +11,7 @@ obj-$(CONFIG_BPF_SYSCALL) += bpf_iter.o map_iter.o task_iter.o prog_iter.o link_
 obj-$(CONFIG_BPF_SYSCALL) += hashtab.o arraymap.o percpu_freelist.o bpf_lru_list.o lpm_trie.o map_in_map.o bloom_filter.o
 obj-$(CONFIG_BPF_SYSCALL) += local_storage.o queue_stack_maps.o ringbuf.o bpf_insn_array.o
 obj-$(CONFIG_BPF_SYSCALL) += bpf_local_storage.o bpf_task_storage.o
-obj-$(CONFIG_BPF_SYSCALL) += fixups.o cfg.o states.o backtrack.o check_btf.o
+obj-$(CONFIG_BPF_SYSCALL) += fixups.o cfg.o states.o backtrack.o check_btf.o exception.o
 obj-${CONFIG_BPF_LSM}	  += bpf_inode_storage.o
 obj-$(CONFIG_BPF_SYSCALL) += disasm.o mprog.o
 obj-$(CONFIG_BPF_JIT) += trampoline.o
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
new file mode 100644
index 000000000000..d6b8ca98e71c
--- /dev/null
+++ b/kernel/bpf/exception.c
@@ -0,0 +1,137 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bpf.h>
+#include <linux/bpf_verifier.h>
+#include <linux/slab.h>
+#include "exception.h"
+
+#define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
+
+#define MIN_BPF_CLEANUP_INFO_SIZE	12
+#define MAX_CLEANUP_INFO_REC_SIZE	252	/* as MAX_FUNCINFO_REC_SIZE */
+
+int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
+		       bpfptr_t uattr)
+{
+	u32 krec_size = sizeof(struct bpf_cleanup_info);
+	u32 i, nrec, urec_size, min_size, prev_end = 0;
+	struct bpf_cleanup_info *krecord;
+	bpfptr_t urecord;
+	int ret = -EINVAL;
+
+	nrec = attr->cleanup_info_cnt;
+	if (!nrec)
+		return 0;
+	if (nrec > env->prog->len) {
+		verbose(env, "cleanup info has %u records for %u instructions\n",
+			nrec, env->prog->len);
+		return -EINVAL;
+	}
+
+	urec_size = attr->cleanup_info_rec_size;
+	if (urec_size < MIN_BPF_CLEANUP_INFO_SIZE ||
+	    urec_size > MAX_CLEANUP_INFO_REC_SIZE ||
+	    urec_size % sizeof(u32)) {
+		verbose(env, "invalid cleanup info rec size %u\n", urec_size);
+		return -EINVAL;
+	}
+
+	krecord = kvcalloc(nrec, krec_size, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+	if (!krecord)
+		return -ENOMEM;
+
+	min_size = min_t(u32, krec_size, urec_size);
+	urecord = make_bpfptr(attr->cleanup_info, uattr.is_kernel);
+	for (i = 0; i < nrec; i++) {
+		struct bpf_subprog_info *sb, *se, *sl;
+		struct bpf_cleanup_info *rec = &krecord[i];
+
+		ret = bpf_check_uarg_tail_zero(urecord, krec_size, urec_size);
+		if (ret) {
+			if (ret == -E2BIG) {
+				verbose(env, "nonzero tailing record in cleanup info\n");
+				if (copy_to_bpfptr_offset(uattr,
+							  offsetof(union bpf_attr,
+								   cleanup_info_rec_size),
+							  &min_size, sizeof(min_size)))
+					ret = -EFAULT;
+			}
+			goto err_free;
+		}
+
+		if (copy_from_bpfptr(rec, urecord, min_size)) {
+			ret = -EFAULT;
+			goto err_free;
+		}
+		bpfptr_add(&urecord, urec_size);
+
+		ret = -EINVAL;
+		if (rec->begin_off >= rec->end_off) {
+			verbose(env, "cleanup_info[%u]: begin %u >= end %u\n",
+				i, rec->begin_off, rec->end_off);
+			goto err_free;
+		}
+		if (i && rec->begin_off < prev_end) {
+			verbose(env,
+				"cleanup_info[%u]: range [%u,%u) is unsorted or overlaps the previous record\n",
+				i, rec->begin_off, rec->end_off);
+			goto err_free;
+		}
+		prev_end = rec->end_off;
+
+		sb = bpf_find_containing_subprog(env, rec->begin_off);
+		se = bpf_find_containing_subprog(env, rec->end_off - 1);
+		sl = bpf_find_containing_subprog(env, rec->landing_pad_off);
+		if (!sb || !se || !sl) {
+			verbose(env, "cleanup_info[%u]: offset out of range\n", i);
+			goto err_free;
+		}
+		if (sb != se || sb != sl) {
+			verbose(env,
+				"cleanup_info[%u]: range/landing pad span multiple subprogs\n",
+				i);
+			goto err_free;
+		}
+		/*
+		 * A zero opcode is the second half of a 16-byte insn, not an
+		 * insn. end_off is exclusive, so it may be one past the last.
+		 */
+		if (!env->prog->insnsi[rec->begin_off].code ||
+		    !env->prog->insnsi[rec->landing_pad_off].code ||
+		    (rec->end_off < env->prog->len &&
+		     !env->prog->insnsi[rec->end_off].code)) {
+			verbose(env, "cleanup_info[%u]: points at invalid insn\n", i);
+			goto err_free;
+		}
+	}
+
+	/* Reject a landing pad inside any call-site range, its own included. */
+	ret = -EINVAL;
+	for (i = 0; i < nrec; i++) {
+		u32 pad = krecord[i].landing_pad_off;
+		u32 l = 0, r = nrec;
+
+		while (l < r) {
+			u32 m = l + (r - l) / 2;
+
+			if (pad < krecord[m].begin_off) {
+				r = m;
+			} else if (pad >= krecord[m].end_off) {
+				l = m + 1;
+			} else {
+				verbose(env,
+					"cleanup_info[%u]: landing pad %u is inside the call-site range of cleanup_info[%u]\n",
+					i, pad, m);
+				goto err_free;
+			}
+		}
+	}
+
+	env->cleanup_info = krecord;
+	env->cleanup_info_cnt = nrec;
+	return 0;
+
+err_free:
+	kvfree(krecord);
+	return ret;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
new file mode 100644
index 000000000000..cf099dcc5b74
--- /dev/null
+++ b/kernel/bpf/exception.h
@@ -0,0 +1,15 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef __BPF_EXCEPTION_H
+#define __BPF_EXCEPTION_H
+
+#include <linux/bpfptr.h>
+#include <linux/types.h>
+
+union bpf_attr;
+struct bpf_verifier_env;
+
+int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
+		       bpfptr_t uattr);
+
+#endif /* __BPF_EXCEPTION_H */
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index ac52f4ae414c..0e14afe3fdc5 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -2924,7 +2924,7 @@ int __init __used bpf_multi_func(void) { return 0; }
 BTF_ID_LIST_GLOBAL_SINGLE(bpf_multi_func_btf_id, func, bpf_multi_func)
 
 /* last field in 'union bpf_attr' used by this command */
-#define BPF_PROG_LOAD_LAST_FIELD keyring_id
+#define BPF_PROG_LOAD_LAST_FIELD cleanup_info_cnt
 
 static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, struct bpf_log_attr *attr_log)
 {
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index b840b3eb9b22..efc516e5ee4d 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -37,6 +37,7 @@
 
 #include "diagnostics.h"
 #include "disasm.h"
+#include "exception.h"
 
 static const struct bpf_verifier_ops * const bpf_verifier_ops[] = {
 #define BPF_PROG_TYPE(_id, _name, prog_ctx_type, kern_ctx_type) \
@@ -22545,6 +22546,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	if (ret < 0)
 		goto skip_full_check;
 
+	ret = bpf_exc_check_info(env, attr, uattr);
+	if (ret < 0)
+		goto skip_full_check;
+
 	/* Validate instructions and resolve the program's referenced resources. */
 	ret = check_and_resolve_insns(env);
 	if (ret < 0)
@@ -22762,6 +22767,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	kvfree(env->callx_edges);
 	kvfree(env->func_ptrs);
 	bpf_diag_free(env);
+	kvfree(env->cleanup_info);
 	kvfree(env);
 	return ret;
 }
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 4687c3310996..033b5ed883db 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -1702,6 +1702,9 @@ union bpf_attr {
 		 * verification.
 		 */
 		__s32		keyring_id;
+		__aligned_u64	cleanup_info;	/* exception cleanup table */
+		__u32		cleanup_info_rec_size; /* userspace bpf_cleanup_info size */
+		__u32		cleanup_info_cnt; /* number of bpf_cleanup_info records */
 	};
 
 	struct { /* anonymous struct used by BPF_OBJ_* commands */
@@ -7638,6 +7641,17 @@ struct bpf_line_info {
 	__u32	line_col;
 };
 
+/*
+ * One record of an exception cleanup table: calls in [begin_off, end_off)
+ * that unwind resume at landing_pad_off. All three are instruction offsets
+ * in the program as loaded.
+ */
+struct bpf_cleanup_info {
+	__u32	begin_off;
+	__u32	end_off;
+	__u32	landing_pad_off;
+};
+
 struct bpf_spin_lock {
 	__u32	val;
 };
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
  2026-10-01 13:30 ` [PATCH bpf-next v8 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
  2026-10-01 13:30 ` [PATCH bpf-next v8 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
@ 2026-10-01 13:30 ` Yonghong Song
  2026-10-01 13:30 ` [PATCH bpf-next v8 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
                   ` (18 subsequent siblings)
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:30 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

An exception cleanup needs two terminators, and the kernel provides both as
kfuncs so that a BPF program can name them.

bpf_unwind() begins an unwind: the frames it leaves run their landing
pads on the way out to the program's exit. It is separate from bpf_throw(),
which keeps its own meaning -- leaving for the exception boundary with the
frames in between discarded.

bpf_unwind_resume() ends a landing pad, carrying the unwind on once that
frame's cleanups have run. The compiler names it _Unwind_Resume, the base
unwind ABI's entry point for the same thing; a later libbpf patch resolves
that name to this one. Every unwind ABI hands _Unwind_Resume the exception
object, and LLVM emits that argument on BPF too. Nothing in the kernel
needs it now, but the kfunc takes it as ptr__ign so the prototype matches
the call the compiler makes.

Both are defined here and neither is registered with any program type yet.
bpf_unwind() may appear wherever a program can unwind from, and
bpf_unwind_resume() only inside a landing pad, and neither does anything
until a pad can be dispatched at run time, so registration waits for the
patch that adds that.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 kernel/bpf/helpers.c | 14 ++++++++++++++
 1 file changed, 14 insertions(+)

diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index a284f20c97d5..4eccd6742eba 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -3424,6 +3424,10 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
 	return false;
 }
 
+__bpf_kfunc void bpf_unwind(void)
+{
+}
+
 __bpf_kfunc void bpf_throw(u64 cookie)
 {
 	struct bpf_throw_ctx ctx = {};
@@ -3445,6 +3449,16 @@ __bpf_kfunc void bpf_throw(u64 cookie)
 	WARN(1, "A call to BPF exception callback should never return\n");
 }
 
+__bpf_kfunc void bpf_unwind_resume(void *ptr__ign)
+{
+	/*
+	 * Never reached: the verifier accepts this kfunc only as a frame
+	 * terminator and bpf_do_misc_fixups() lowers every one of them to
+	 * 'r0 = 0; exit', so no call to this body survives to run.
+	 */
+	WARN_ONCE(1, "exception cleanup resume was not lowered to a return\n");
+}
+
 __bpf_kfunc int bpf_wq_init(struct bpf_wq *wq, void *p__const_map, unsigned int flags)
 {
 	struct bpf_async_kern *async = (struct bpf_async_kern *)wq;
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 04/22] bpf: Add lookups for exception cleanup resumes and landing pads
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (2 preceding siblings ...)
  2026-10-01 13:30 ` [PATCH bpf-next v8 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
@ 2026-10-01 13:30 ` Yonghong Song
  2026-10-01 13:48   ` sashiko-bot
  2026-10-01 13:30 ` [PATCH bpf-next v8 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
                   ` (17 subsequent siblings)
  21 siblings, 1 reply; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:30 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

Add the first lookups to exception.c: recognising a call to bpf_unwind() or
bpf_unwind_resume(), and asking which landing pad, if any, a call site
unwinds to. Their users, and what fills in the pad of a call site, come in
later patches.

The pad of a call site is kept in insn_aux_data, so the three places that
move instructions around -- bpf_patch_insn_data(), verifier_remove_insns()
and bpf_opt_remove_nops() -- learn to keep it in step.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf_verifier.h |  5 +++++
 kernel/bpf/exception.c       | 24 ++++++++++++++++++++++++
 kernel/bpf/exception.h       |  4 ++++
 kernel/bpf/fixups.c          | 26 +++++++++++++++++++++++++-
 4 files changed, 58 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index ad0ca8047712..a22ce69e9aff 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -700,6 +700,11 @@ struct bpf_insn_aux_data {
 	u64 jump_target:1;
 
 	unsigned int orig_idx; /* original instruction index, initialized once */
+	/*
+	 * 1 + the instruction index of the exception cleanup landing pad
+	 * this call site unwinds to, or 0 for none.
+	 */
+	u32 cleanup_pad;
 	/*
 	 * CFG strongly connected component this instruction belongs to,
 	 * zero if it is a singleton SCC.
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index d6b8ca98e71c..3ea1bff5cc90 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -2,6 +2,8 @@
 /* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
 #include <linux/bpf.h>
 #include <linux/bpf_verifier.h>
+#include <linux/btf_ids.h>
+#include <linux/filter.h>
 #include <linux/slab.h>
 #include "exception.h"
 
@@ -135,3 +137,25 @@ int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
 	kvfree(krecord);
 	return ret;
 }
+
+BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
+BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
+
+bool bpf_is_unwind_kfunc(const struct bpf_insn *insn)
+{
+	return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
+	       insn->imm == bpf_unwind_id[0];
+}
+
+bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
+{
+	return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
+	       insn->imm == bpf_unwind_resume_id[0];
+}
+
+int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
+{
+	u32 pad = env->insn_aux_data[idx].cleanup_pad;
+
+	return pad ? (int)pad - 1 : -1;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index cf099dcc5b74..d5c6ac459870 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -8,8 +8,12 @@
 
 union bpf_attr;
 struct bpf_verifier_env;
+struct bpf_insn;
 
 int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
 		       bpfptr_t uattr);
+int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
+bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
+bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
 
 #endif /* __BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 37cf130ebb57..5b7fe4ba610b 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -268,11 +268,18 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
 			data[i].non_stack_access =
 				data[off + cnt - 1].non_stack_access;
 			data[off + cnt - 1].non_stack_access = false;
+			data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
+			data[off + cnt - 1].cleanup_pad = 0;
 		} else if (bpf_is_mem_insn(insn + i)) {
 			data[i].non_stack_access = true;
 		}
 	}
 
+	if (env->cleanup_info_cnt)
+		for (i = 0; i < prog_len; i++)
+			if (data[i].cleanup_pad > off + 1)
+				data[i].cleanup_pad += cnt - 1;
+
 	/*
 	 * Last slot instruction could be a newly generated
 	 * BPF_ST/BPF_LDX/BPF_STX, systematically mark it for non-stack access
@@ -619,6 +626,7 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
 	struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
 	unsigned int orig_prog_len = env->prog->len;
 	int err;
+	u32 i;
 
 	if (bpf_rewrite_must_abort())
 		return -EINTR;
@@ -647,6 +655,17 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
 		sizeof(*aux_data) * (orig_prog_len - off - cnt));
 	env->insn_aux_data_len -= cnt;
 
+	if (env->cleanup_info_cnt) {
+		for (i = 0; i < env->insn_aux_data_len; i++) {
+			u32 pad = aux_data[i].cleanup_pad;
+
+			if (pad > off + cnt)
+				aux_data[i].cleanup_pad = pad - cnt;
+			else if (pad > off)
+				aux_data[i].cleanup_pad = 0;
+		}
+	}
+
 	return 0;
 }
 
@@ -752,7 +771,7 @@ int bpf_opt_remove_nops(struct bpf_verifier_env *env)
 	struct bpf_insn *insn = env->prog->insnsi;
 	int insn_cnt = env->prog->len;
 	bool is_may_goto_0, is_ja;
-	int i, err;
+	int i, j, err;
 
 	for (i = 0; i < insn_cnt; i++) {
 		is_may_goto_0 = !memcmp(&insn[i], &MAY_GOTO_0, sizeof(MAY_GOTO_0));
@@ -763,6 +782,11 @@ int bpf_opt_remove_nops(struct bpf_verifier_env *env)
 		if (aux[i].indirect_target)
 			continue;
 
+		if (env->cleanup_info_cnt)
+			for (j = 0; j < insn_cnt; j++)
+				if (env->insn_aux_data[j].cleanup_pad == i + 1)
+					env->insn_aux_data[j].cleanup_pad = i + 2;
+
 		err = verifier_remove_insns(env, i, 1);
 		if (err)
 			return err;
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 05/22] bpf: Prepare for an exception cleanup table before the CFG walk
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (3 preceding siblings ...)
  2026-10-01 13:30 ` [PATCH bpf-next v8 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
@ 2026-10-01 13:30 ` Yonghong Song
  2026-10-01 14:31   ` bot+bpf-ci
  2026-10-01 13:30 ` [PATCH bpf-next v8 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
                   ` (16 subsequent siblings)
  21 siblings, 1 reply; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:30 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

Record in insn_aux_data what the later passes need from the cleanup table:
cleanup_pad, the landing pad a frame resumes at, for every call that can
unwind -- a BPF-to-BPF call, direct or indirect, or bpf_unwind() -- within
the [begin_off, end_off) range of a cleanup record. Helper and other kfunc
calls in the range cannot unwind and are skipped. Subsequent commits
consume it.

bpf_exc_prepare() runs before bpf_check_cfg(), whose walk consumes what it
produces. It refuses a table on an offloaded program, and on one whose JIT
cannot dispatch landing pads or was not asked to compile it. What survives
is marked jit_required: the interpreter cannot dispatch a pad.
bpf_jit_supports_cleanup_pads() is weak here and says no; the arch patches
provide the real ones.

A table is refused for a program whose verifier_ops has a gen_epilogue, as
bpf_qdisc's do. That epilogue is planted by rewriting the exits a program
has when bpf_convert_ctx_accesses() runs, and the exits an unwind returns
through are added after it, so they would skip it. A struct_ops program's
ops, and with them gen_epilogue, are only known after the CFG walk; there
it is the same check made again when bpf_unwind() is verified, from a later
patch, that refuses it.

A table is also refused alongside bpf_throw(), a second answer to what runs
on the way out: it leaves for the exception boundary without rewriting the
return addresses of the frames it passes, so no pad between the two would
run. Both the tagged exception callback and the throw itself are checked --
either can appear without the other -- over the whole instruction stream,
so a throw in a subprogram is caught as well. The checks sit in
bpf_exc_check_prog(), which a later patch also runs at every bpf_unwind(),
so they hold for a program that unwinds with no table too.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/filter.h |  1 +
 kernel/bpf/core.c      |  5 +++
 kernel/bpf/exception.c | 79 ++++++++++++++++++++++++++++++++++++++++++
 kernel/bpf/exception.h |  2 ++
 kernel/bpf/verifier.c  |  5 +++
 5 files changed, 92 insertions(+)

diff --git a/include/linux/filter.h b/include/linux/filter.h
index e42eccb0990e..972b3ed2a51d 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1248,6 +1248,7 @@ bool bpf_jit_supports_stack_args(void);
 bool bpf_jit_supports_arena_args(void);
 bool bpf_jit_supports_far_kfunc_call(void);
 bool bpf_jit_supports_exceptions(void);
+bool bpf_jit_supports_cleanup_pads(void);
 bool bpf_jit_supports_ptr_xchg(void);
 bool bpf_jit_supports_arena(void);
 bool bpf_jit_supports_insn(struct bpf_insn *insn, bool in_arena);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index d3b8b626ec0f..d813fdde29e3 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -3511,6 +3511,11 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
 {
 }
 
+bool __weak bpf_jit_supports_cleanup_pads(void)
+{
+	return false;
+}
+
 bool __weak bpf_jit_supports_timed_may_goto(void)
 {
 	return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index 3ea1bff5cc90..e12cdb12cde3 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -141,6 +141,85 @@ int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
 BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
 BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
 
+static int reject_throw(struct bpf_verifier_env *env)
+{
+	u32 i;
+
+	for (i = 0; i < env->prog->len; i++) {
+		if (!bpf_is_throw_kfunc(&env->prog->insnsi[i]))
+			continue;
+		verbose(env,
+			"exception cleanup cannot be combined with bpf_throw at insn %u\n",
+			i);
+		return -EINVAL;
+	}
+	return 0;
+}
+
+static void mark_call_sites(struct bpf_verifier_env *env)
+{
+	u32 i, j;
+
+	for (i = 0; i < env->cleanup_info_cnt; i++) {
+		struct bpf_cleanup_info *rec = &env->cleanup_info[i];
+
+		for (j = rec->begin_off; j < rec->end_off; j++) {
+			struct bpf_insn *insn = &env->prog->insnsi[j];
+
+			if (!bpf_pseudo_call(insn) && !bpf_is_callx(insn) &&
+			    !bpf_is_unwind_kfunc(insn))
+				continue;
+			env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
+		}
+	}
+}
+
+int bpf_exc_check_prog(struct bpf_verifier_env *env)
+{
+	int err;
+
+	if (bpf_prog_is_offloaded(env->prog->aux)) {
+		verbose(env,
+			"exception cleanup is not supported for offloaded programs\n");
+		return -EINVAL;
+	}
+	if (!bpf_jit_supports_cleanup_pads() || !env->prog->jit_requested) {
+		verbose(env,
+			"exception cleanup needs a JIT that can dispatch landing pads\n");
+		return -EOPNOTSUPP;
+	}
+	if (env->ops->gen_epilogue) {
+		verbose(env,
+			"exception cleanup is not supported for a program with an epilogue\n");
+		return -EOPNOTSUPP;
+	}
+	if (env->exception_callback_subprog) {
+		verbose(env,
+			"exception cleanup cannot be combined with an exception callback\n");
+		return -EINVAL;
+	}
+	err = reject_throw(env);
+	if (err)
+		return err;
+	env->prog->jit_required = 1;
+	return 0;
+}
+
+int bpf_exc_prepare(struct bpf_verifier_env *env)
+{
+	int err;
+
+	if (!env->cleanup_info_cnt)
+		return 0;
+
+	err = bpf_exc_check_prog(env);
+	if (err)
+		return err;
+
+	mark_call_sites(env);
+	return 0;
+}
+
 bool bpf_is_unwind_kfunc(const struct bpf_insn *insn)
 {
 	return bpf_pseudo_kfunc_call(insn) && insn->off == 0 &&
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index d5c6ac459870..96dac3037d75 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -12,6 +12,8 @@ struct bpf_insn;
 
 int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
 		       bpfptr_t uattr);
+int bpf_exc_prepare(struct bpf_verifier_env *env);
+int bpf_exc_check_prog(struct bpf_verifier_env *env);
 int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
 bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
 bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index efc516e5ee4d..80034429fdd0 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -22550,6 +22550,11 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 	if (ret < 0)
 		goto skip_full_check;
 
+	/* The CFG needs an edge from a call in a cleanup range to its pad. */
+	ret = bpf_exc_prepare(env);
+	if (ret < 0)
+		goto skip_full_check;
+
 	/* Validate instructions and resolve the program's referenced resources. */
 	ret = check_and_resolve_insns(env);
 	if (ret < 0)
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 06/22] bpf: Make exception landing pads reachable in the CFG
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (4 preceding siblings ...)
  2026-10-01 13:30 ` [PATCH bpf-next v8 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
@ 2026-10-01 13:30 ` Yonghong Song
  2026-10-01 13:30 ` [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier Yonghong Song
                   ` (15 subsequent siblings)
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:30 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

A bpf_unwind() or a bpf2bpf call inside the [begin_off, end_off) range of a
cleanup record can reach that record's landing pad. Add that edge to the
CFG walk, which explores the pad and makes both ends prune points, and to
bpf_insn_successors(), which liveness and the SCC passes walk.

The pad is pushed by itself: visit_func_call_insn() returns as soon as
visit_cleanup_pad_edge() has pushed one, and the call site is visited again
for its fall-through once the pad is explored. bpf_check_cfg() peeks the
top of its stack and re-visits until DONE_EXPLORING, so a visit pushing two
successors would leave the pad DISCOVERED while it is no longer on the path
being walked, and push_insn() reads DISCOVERED as a back-edge. A branch
from one pad into another -- how a frame with two regions chains them --
would be taken for one. The pad's own insn_state is what records that the
edge is done, so there is nothing else to remember.

Liveness needs one more thing. A frame suspended at a call normally keeps a
stack slot alive only if something reads it once the call returns. A
landing pad is not on that path: it is reached from the call itself, not
from the instruction after it. So a slot that only the pad reads looks
dead, and clean_verifier_state() poisons it while the callee runs. Ask
whether the pad reads it too. The entry_var_stack selftest, added later
in the series, is the shape that needs this -- it stores to its frame
before the call, never reads that slot on the way back, and reloads it in
the pad.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 kernel/bpf/cfg.c      | 38 ++++++++++++++++++++++++++++++++++++++
 kernel/bpf/liveness.c | 26 +++++++++++++++++++++++++-
 2 files changed, 63 insertions(+), 1 deletion(-)

diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index d8a579680e5b..63afbc5fb296 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -6,6 +6,7 @@
 #include <linux/sort.h>
 
 #include "diagnostics.h"
+#include "exception.h"
 
 #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
 
@@ -160,6 +161,38 @@ static int push_insn(int t, int w, int e, struct bpf_verifier_env *env)
 	return DONE_EXPLORING;
 }
 
+static int visit_cleanup_pad_edge(int t, struct bpf_verifier_env *env)
+{
+	int *insn_stack = env->cfg.insn_stack;
+	int *insn_state = env->cfg.insn_state;
+	int w;
+
+	if (!env->cleanup_info_cnt)
+		return DONE_EXPLORING;
+	w = bpf_exc_pad_of_call(env, t);
+	if (w < 0)
+		return DONE_EXPLORING;
+
+	/*
+	 * @t is a call that may branch here, and @w is the target of that
+	 * branch, so both are prune points. @w especially: every covered call
+	 * site in a region unwinds to the same pad, and without a prune point
+	 * at its head the verifier walks the pad again for each of them.
+	 */
+	mark_prune_point(env, t);
+	mark_prune_point(env, w);
+	mark_jmp_point(env, w);
+	mark_jump_target(env, w);
+
+	if (insn_state[w])
+		return DONE_EXPLORING;
+	if (env->cfg.cur_stack >= env->prog->len)
+		return -E2BIG;
+	insn_stack[env->cfg.cur_stack++] = w;
+	insn_state[w] |= DISCOVERED;
+	return KEEP_EXPLORING;
+}
+
 static int visit_func_call_insn(int t, struct bpf_insn *insns,
 				struct bpf_verifier_env *env,
 				bool visit_callee)
@@ -167,6 +200,11 @@ static int visit_func_call_insn(int t, struct bpf_insn *insns,
 	int ret, insn_sz;
 	int w;
 
+	/* One push per visit: @t is revisited once the pad is explored. */
+	ret = visit_cleanup_pad_edge(t, env);
+	if (ret != DONE_EXPLORING)
+		return ret;
+
 	insn_sz = bpf_is_ldimm64(&insns[t]) ? 2 : 1;
 	ret = push_insn(t, t + insn_sz, FALLTHROUGH, env);
 	if (ret)
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index cd9523f69298..b3091eac2cf1 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -8,6 +8,8 @@
 #include <linux/slab.h>
 #include <linux/sort.h>
 
+#include "exception.h"
+
 #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
 
 /*
@@ -384,6 +386,20 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
 			succ->items[succ->cnt++] = exit_idx;
 	}
 
+	/*
+	 * A call a cleanup record covers can leave through its landing pad.
+	 * Only a call to a subprogram, direct or through callx, or to
+	 * bpf_unwind() is marked, none of which is an edge the block above
+	 * adds, so there are at most two successors, which env->succ is sized
+	 * for.
+	 */
+	if (unlikely(env->cleanup_info_cnt)) {
+		int pad = bpf_exc_pad_of_call(env, idx);
+
+		if (pad >= 0)
+			succ->items[succ->cnt++] = pad;
+	}
+
 	return succ;
 }
 
@@ -510,7 +526,8 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
 	 * Slot is alive if it is read before q->insn_idx in current func instance,
 	 * or if for some outer func instance:
 	 * - alive before callsite if callsite calls callback or is callx, otherwise
-	 * - alive after callsite
+	 * - alive after callsite,
+	 * - or alive at the landing pad a cleanup record gives the callsite
 	 */
 	struct live_stack_query *q = &env->liveness->live_stack_query;
 	struct func_instance *instance, *curframe_instance;
@@ -545,6 +562,13 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
 		alive = callee_stack_access_at_callsite(env, callsite)
 			? is_live_before(instance, callsite, rel, half_spi)
 			: is_live_before(instance, callsite + 1, rel, half_spi);
+
+		if (!alive && unlikely(env->cleanup_info_cnt)) {
+			int pad = bpf_exc_pad_of_call(env, callsite);
+
+			if (pad >= 0)
+				alive = is_live_before(instance, pad, rel, half_spi);
+		}
 		if (alive)
 			return true;
 	}
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (5 preceding siblings ...)
  2026-10-01 13:30 ` [PATCH bpf-next v8 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
@ 2026-10-01 13:30 ` Yonghong Song
  2026-10-01 13:50   ` sashiko-bot
                     ` (2 more replies)
  2026-10-01 13:30 ` [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
                   ` (14 subsequent siblings)
  21 siblings, 3 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:30 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

Once a later patch makes bpf_unwind() dispatch pads, an unwind rewrites
the return address of every frame it passes: to the pad where a record
covers the frame's call, else to the frame's epilogue. Each frame then
returns normally. A pad ends in a resume (a later patch refuses one that
does not), lowered to 'r0 = 0; exit', so it returns too. The verifier now
follows that path.

Where the walk goes on:

  instruction                          goes on at
  -----------------------------------  ---------------------------------
  bpf_unwind(), record over it         its own frame's pad
  bpf_unwind(), no record over it      unwind_frames()
  bpf_unwind_resume()                  unwind_frames()
  call to a global subprog that can    next insn, and the unwind from the
  unwind                               returned state: to the call's pad,
                                       or on through unwind_frames()

unwind_frames(), for main -> A -> B -> C where only A's call is covered:

  program                              run time: bpf_unwind() in C
  -----------------------------------  ---------------------------------
  main:     call A                     main returns via its epilogue
  A:     1: call B     [1, 2) -> P     A resumes at P
         2: ...
         P: <drop A's resources>
            call bpf_unwind_resume
  B:        call C     no record       B returns via its epilogue
  C:        call bpf_unwind            C returns via 'r0 = 0; exit'

  frame   at C's bpf_unwind()   unwind_frames()     at P's resume
  -----   -------------------   -----------------   -----------------
    3     C  <- curframe        popped
    2     B                     popped
    1     A                     A  <- curframe, P   popped
    0     main                  main                exit with r0 = 0

  P gets A's frame as B and C left it, with r0 unknown and r1-r5
  cleared. Main returning goes through process_bpf_exit_full(): nothing
  may still be held, and r0 = 0 must suit the program type. r1-r5 are
  cleared there too, since the call main returns from clobbered them.

Why the pad's state is not taken at the call:

  the callee may, before unwinding     a snapshot at the call would trust
  -----------------------------------  ----------------------------------
  write the caller's stack via a ptr   the slot's old value
  overwrite a spilled pointer          a pointer that is now a scalar
  reinitialise a dynptr or iterator    the old dynptr or iterator
  change packet data                   stale packet pointers

  The callee's epilogue restores only r6-r9 and the frame pointer,
  none of the above. A covered call whose callee cannot unwind leaves
  its pad unreached.

  A global subprog is verified on its own, so the unwind out of it is
  taken from the state its call returns in, after check_func_call(),
  whose argument checks and packet invalidation already cover the list
  above; a dynptr passed to a global subprog is read-only to it.

Precision backtracking, three new edges:

  edge                     frame
  -----------------------  ---------------------------------------------
  global call <- its pad   stays in the caller (subseq_idx is the pad)
  unwind <- pad            moves to the frame the unwind left, recorded
                           in its history entry under INSN_F_UNWIND
  unwind <- main's exit    a history entry makes the unwinding insn,
                           which sets r0, the one before main's return

  In the example, backtracking from P:

    insn                bt->frame
    ------------------  ------------------------
    P                   1
    C: bpf_unwind()     3, from INSN_F_UNWIND
    C ... entry         2, at B's call to C
    B ... entry         1, at A's call to B
    A, before the call

Also:

 - check_kfunc_allowed() is split out of check_kfunc_call(), which the
   two kfuncs no longer reach. Until they are registered, it is what
   refuses them.
 - INSN_F_UNWIND takes a fifth flag bit from the history entry's padding.
 - The CFG walk marks a subprogram calling bpf_unwind() might_unwind,
   carried up by merge_callee_effects(). Where the program can unwind at
   all, a subprogram with a callx is marked too, since the pointer it
   calls through may have been handed to it rather than loaded there, and
   the mark is carried up to its callers again.
   unwind_out_of_global_call() reads it here, and later patches do too.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf_verifier.h |  12 ++-
 kernel/bpf/backtrack.c       |  49 +++++++++-
 kernel/bpf/cfg.c             |  63 ++++++++++++
 kernel/bpf/verifier.c        | 179 +++++++++++++++++++++++++++++++++--
 4 files changed, 290 insertions(+), 13 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index a22ce69e9aff..c90fa3f5e787 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -383,6 +383,13 @@ enum {
 	INSN_F_SRC_REG_STACK = BIT(2), /* src_reg is PTR_TO_STACK */
 
 	INSN_F_STACK_ARG_ACCESS = BIT(3),
+
+	/*
+	 * A bpf_unwind(), a resume or an unwinding global call that left frame
+	 * 'frame' for a landing pad in one of its callers; backtracking jumps
+	 * back into that frame here.
+	 */
+	INSN_F_UNWIND = BIT(4),
 };
 
 /* Registers linked to one jump condition that a history entry can record */
@@ -393,8 +400,8 @@ struct bpf_jmp_history_entry {
 	u32 idx : 20;
 	u32 frame : 4;	/* stack access frame number */
 	/* special INSN_F_xxx flags */
-	u32 flags : 4;
-	u32 : 4;
+	u32 flags : 5;
+	u32 : 3;
 	u32 prev_idx : 20;
 	u32 spi : 12;	/* stack slot index */
 	/*
@@ -836,6 +843,7 @@ struct bpf_subprog_info {
 	s16 fastcall_stack_off;
 	bool has_tail_call: 1;
 	bool might_throw: 1;
+	bool might_unwind: 1;
 	bool tail_call_reachable: 1;
 	bool has_ld_abs: 1;
 	bool is_cb: 1;
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 0e38b9575328..f5504334df90 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -4,6 +4,7 @@
 #include <linux/bpf_verifier.h>
 #include <linux/filter.h>
 #include <linux/bitmap.h>
+#include "exception.h"
 
 #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
 
@@ -424,7 +425,30 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
 		if (class == BPF_STX)
 			bt_set_reg(bt, sreg);
 	} else if (class == BPF_JMP || class == BPF_JMP32) {
-		if (bpf_pseudo_call(insn) || bpf_is_callx(insn)) {
+		if (hist && (hist->flags & INSN_F_UNWIND)) {
+			/*
+			 * A bpf_unwind(), a resume or an unwinding global call
+			 * left frame hist->frame here, for a landing pad in
+			 * this one. The walk crosses back into that frame,
+			 * past any frames between, which were entered and
+			 * never returned from. The pad found r0 unknown and
+			 * r1-r5 clobbered; r6-r9 and the stack are this
+			 * frame's own and stay marked in its masks until the
+			 * walk comes back out.
+			 */
+			bt_clear_reg(bt, BPF_REG_0);
+			if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
+				verifier_bug(env, "backtracking unwind unexpected regs %x",
+					     bt_reg_mask(bt));
+				return -EFAULT;
+			}
+			if (verifier_bug_if(hist->frame <= bt->frame, env,
+					    "unwind from frame %d to frame %d",
+					    hist->frame, bt->frame))
+				return -EFAULT;
+			bt->frame = hist->frame;
+			return 0;
+		} else if (bpf_pseudo_call(insn) || bpf_is_callx(insn)) {
 			int subprog_insn_idx, subprog = -1;
 
 			if (bpf_pseudo_call(insn)) {
@@ -434,6 +458,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
 					return -EFAULT;
 			}
 
+			if (bpf_exc_pad_of_call(env, idx) == subseq_idx) {
+				/*
+				 * We came from the landing pad of a call to a
+				 * global subprog, branched to from the state
+				 * the call returns in: as on its return, no
+				 * frame was entered here. The call clobbered
+				 * r0-r5; r6-r9 and the stack are the caller's
+				 * own and keep going back from here.
+				 */
+				bt_clear_reg(bt, BPF_REG_0);
+				if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
+					verifier_bug(env, "landing pad unexpected regs %x",
+						     bt_reg_mask(bt));
+					return -EFAULT;
+				}
+				return 0;
+			}
+
 			/* callx calls static subprogs only */
 			if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
 				/* check that jump history doesn't have any
@@ -956,6 +998,11 @@ int bpf_mark_chain_precision(struct bpf_verifier_env *env,
 		if (!st)
 			break;
 
+		if (verifier_bug_if(bt->frame > st->curframe, env,
+				    "backtrack frame %d, state curframe %d",
+				    bt->frame, st->curframe))
+			return -EFAULT;
+
 		for (fr = bt->frame; fr >= 0; fr--) {
 			func = st->frame[fr];
 			bitmap_from_u64(mask, bt_frame_reg_mask(bt, fr));
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index 63afbc5fb296..81963b3bdd5a 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -76,6 +76,14 @@ static void mark_subprog_might_throw(struct bpf_verifier_env *env, int off)
 	subprog->might_throw = true;
 }
 
+static void mark_subprog_might_unwind(struct bpf_verifier_env *env, int off)
+{
+	struct bpf_subprog_info *subprog;
+
+	subprog = bpf_find_containing_subprog(env, off);
+	subprog->might_unwind = true;
+}
+
 /* 't' is an index of a call-site.
  * 'w' is a callee entry point.
  * Eventually this function would be called when env->cfg.insn_state[w] == EXPLORED.
@@ -91,6 +99,7 @@ static void merge_callee_effects(struct bpf_verifier_env *env, int t, int w)
 	caller->changes_pkt_data |= callee->changes_pkt_data;
 	caller->might_sleep |= callee->might_sleep;
 	caller->might_throw |= callee->might_throw;
+	caller->might_unwind |= callee->might_unwind;
 }
 
 enum {
@@ -668,6 +677,8 @@ static int visit_insn(int t, struct bpf_verifier_env *env)
 				mark_subprog_changes_pkt_data(env, t);
 			if (ret == 0 && bpf_is_throw_kfunc(insn))
 				mark_subprog_might_throw(env, t);
+			if (ret == 0 && bpf_is_unwind_kfunc(insn))
+				mark_subprog_might_unwind(env, t);
 		}
 		return visit_func_call_insn(t, insns, env, insn->src_reg == BPF_PSEUDO_CALL);
 
@@ -705,6 +716,57 @@ static int visit_insn(int t, struct bpf_verifier_env *env)
 	}
 }
 
+/*
+ * merge_callee_effects() carries might_unwind to where a subprog is called or
+ * has its address taken, but not to a subprog that calls it through a pointer
+ * it was handed. So where the program can unwind at all, take every subprog
+ * with a callx to be able to, and carry that up to its callers.
+ */
+static void mark_callx_might_unwind(struct bpf_verifier_env *env)
+{
+	struct bpf_insn *insns = env->prog->insnsi;
+	struct bpf_subprog_info *caller, *callee;
+	int i, j, len = env->prog->len;
+	struct bpf_func_ptr *ptrs;
+	bool changed;
+	u32 cnt;
+
+	for (i = 0; i < env->subprog_cnt; i++)
+		if (env->subprog_info[i].might_unwind)
+			break;
+	if (i == env->subprog_cnt)
+		return;
+
+	for (i = 0; i < len; i++)
+		if (bpf_is_callx(&insns[i]))
+			bpf_find_containing_subprog(env, i)->might_unwind = true;
+
+	do {
+		changed = false;
+		for (i = 0; i < len; i++) {
+			caller = bpf_find_containing_subprog(env, i);
+			if (caller->might_unwind)
+				continue;
+			if (bpf_pseudo_call(&insns[i]) || bpf_pseudo_func(&insns[i])) {
+				callee = bpf_find_containing_subprog(env, i + insns[i].imm + 1);
+				if (!callee->might_unwind)
+					continue;
+				caller->might_unwind = true;
+				changed = true;
+				continue;
+			}
+			ptrs = insn_func_ptrs(env, i, &cnt);
+			for (j = 0; j < cnt; j++) {
+				callee = bpf_find_containing_subprog(env, ptrs[j].xlated_off);
+				if (!callee->might_unwind)
+					continue;
+				caller->might_unwind = true;
+				changed = true;
+			}
+		}
+	} while (changed);
+}
+
 /* non-recursive depth-first-search to detect loops in BPF program
  * loop == back-edge in directed graph
  */
@@ -795,6 +857,7 @@ int bpf_check_cfg(struct bpf_verifier_env *env)
 		}
 	}
 	ret = 0; /* cfg looks good */
+	mark_callx_might_unwind(env);
 	env->prog->aux->changes_pkt_data = env->subprog_info[0].changes_pkt_data;
 	env->prog->aux->might_sleep = env->subprog_info[0].might_sleep;
 
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 80034429fdd0..c7a350be538e 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -14684,6 +14684,32 @@ static int check_special_kfunc(struct bpf_verifier_env *env, struct bpf_call_arg
 
 static int check_return_code(struct bpf_verifier_env *env, int regno, const char *reg_name);
 
+static int check_kfunc_allowed(struct bpf_verifier_env *env, struct bpf_insn *insn,
+			       int insn_idx, struct bpf_call_arg_meta *meta)
+{
+	const char *operation;
+	int err;
+
+	err = bpf_fetch_kfunc_arg_meta(env, insn->imm, insn->off, meta);
+	if (err == -EACCES && meta->func_name) {
+		verbose(env, "calling kernel function %s is not allowed\n", meta->func_name);
+		operation = bpf_diag_fmt(env, "kfunc %s", meta->func_name);
+		bpf_diag_policy(
+			env, insn_idx, operation, "this program cannot call the kfunc",
+			"Use a kfunc allowed for this program type and attach point, or change the program context.");
+	}
+	return err;
+}
+
+/* noinline keeps a struct bpf_call_arg_meta off the caller's frame. */
+static noinline int check_kfunc_allowed_only(struct bpf_verifier_env *env,
+					     struct bpf_insn *insn, int insn_idx)
+{
+	struct bpf_call_arg_meta meta;
+
+	return check_kfunc_allowed(env, insn, insn_idx, &meta);
+}
+
 static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
 			    int *insn_idx_p)
 {
@@ -14705,14 +14731,7 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
 	if (!insn->imm)
 		return 0;
 
-	err = bpf_fetch_kfunc_arg_meta(env, insn->imm, insn->off, &meta);
-	if (err == -EACCES && meta.func_name) {
-		verbose(env, "calling kernel function %s is not allowed\n", meta.func_name);
-		operation = bpf_diag_fmt(env, "kfunc %s", meta.func_name);
-		bpf_diag_policy(
-			env, insn_idx, operation, "this program cannot call the kfunc",
-			"Use a kfunc allowed for this program type and attach point, or change the program context.");
-	}
+	err = check_kfunc_allowed(env, insn, insn_idx, &meta);
 	if (err)
 		return err;
 	desc_btf = meta.btf;
@@ -19126,6 +19145,127 @@ enum {
 	INSN_IDX_UPDATED = 2,
 };
 
+/*
+ * The current frame is leaving through an unwind. Its caller's saved return
+ * address now points at the pad covering the call, or, with none, at the
+ * caller's epilogue, and so on down. Follow that from the state the frame
+ * leaves in -- anything it wrote into its callers' stacks included -- to the
+ * first pad, or to the main program's frame returning.
+ */
+static int unwind_frames(struct bpf_verifier_env *env, bool *do_print_state)
+{
+	struct bpf_verifier_state *state = env->cur_state;
+	u32 frameno = state->curframe;
+	struct bpf_func_state *callee, *caller;
+	int err, pad;
+
+	while (state->curframe) {
+		callee = cur_func(env);
+		caller = state->frame[state->curframe - 1];
+		pad = bpf_exc_pad_of_call(env, callee->callsite);
+		/* The caller is at its call now, not at this frame's insn. */
+		state->insn_idx = callee->callsite;
+		account_processed_insns(env, callee, caller);
+		free_func_state(callee);
+		state->frame[state->curframe--] = NULL;
+		invalidate_outgoing_stack_args(env, caller);
+		if (pad < 0)
+			continue;
+
+		/*
+		 * The frames between were entered and never returned from,
+		 * so tell precision backtracking which one this left.
+		 */
+		err = bpf_push_jmp_history(env, state, INSN_F_UNWIND, 0, frameno, NULL, 0);
+		if (err)
+			return err;
+		clear_caller_saved_regs(env, caller->regs);
+		mark_reg_unknown(env, caller->regs, BPF_REG_0);
+		env->insn_idx = pad;
+		*do_print_state = true;
+		return INSN_IDX_UPDATED;
+	}
+
+	/*
+	 * The main frame returning ends the program, from its call if frames
+	 * were popped, else from the unwinding insn. Link that to the unwinding
+	 * insn, so that backtracking from the exit starts where r0 is set and
+	 * not in code the unwind skipped.
+	 */
+	env->cur_hist_ent = NULL;
+	env->prev_insn_idx = env->insn_idx;
+	env->insn_idx = state->insn_idx;
+	err = bpf_push_jmp_history(env, state, 0, 0, 0, NULL, 0);
+	if (err)
+		return err;
+
+	/*
+	 * The call clobbered r1-r5, and r0 holds the zero the fixups put
+	 * there. Mark r0 unknown first: the known-zero helper keeps a NOT_INIT
+	 * type.
+	 */
+	clear_caller_saved_regs(env, cur_regs(env));
+	mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
+	mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
+	return process_bpf_exit_full(env, do_print_state, false);
+}
+
+/*
+ * A global subprog is verified on its own, so an unwind out of one is not
+ * walked. Its call has a second successor instead: the unwind, taken from
+ * the state the call returns in, to this frame's pad or on out of it.
+ */
+static int unwind_out_of_global_call(struct bpf_verifier_env *env, int call_idx,
+				     bool *do_print_state)
+{
+	const struct bpf_insn *insn = &env->prog->insnsi[call_idx];
+	int subprog = bpf_find_subprog(env, call_idx + insn->imm + 1);
+	struct bpf_verifier_state *branch;
+	struct bpf_func_state *frame;
+	int pad;
+
+	if (!bpf_subprog_is_global(env, subprog) ||
+	    !env->subprog_info[subprog].might_unwind)
+		return 0;
+
+	/* The call returning normally is walked later. */
+	branch = push_stack(env, call_idx + 1, call_idx, false);
+	if (IS_ERR(branch))
+		return PTR_ERR(branch);
+
+	pad = bpf_exc_pad_of_call(env, call_idx);
+	if (pad < 0)
+		return unwind_frames(env, do_print_state);
+	frame = cur_func(env);
+	clear_caller_saved_regs(env, frame->regs);
+	mark_reg_unknown(env, frame->regs, BPF_REG_0);
+	env->insn_idx = pad;
+	*do_print_state = true;
+	return INSN_IDX_UPDATED;
+}
+
+static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
+			      bool *do_print_state)
+{
+	struct bpf_func_state *frame = cur_func(env);
+	int pad = bpf_exc_pad_of_call(env, *insn_idx);
+	int err;
+
+	if (pad < 0) {
+		if (!env->cur_state->curframe) {
+			err = check_resource_leak(env, false, true,
+						  "an unwind with no landing pad");
+			if (err)
+				return err;
+		}
+		return unwind_frames(env, do_print_state);
+	}
+	clear_caller_saved_regs(env, frame->regs);
+	mark_reg_unknown(env, frame->regs, BPF_REG_0);
+	*insn_idx = pad;
+	return INSN_IDX_UPDATED;
+}
+
 static int process_bpf_exit_full(struct bpf_verifier_env *env,
 				 bool *do_print_state,
 				 bool exception_exit)
@@ -19380,13 +19520,32 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
 					return -EINVAL;
 				}
 			}
+			if (bpf_is_unwind_kfunc(insn) || bpf_is_unwind_resume_kfunc(insn)) {
+				err = check_kfunc_allowed_only(env, insn, env->insn_idx);
+				if (err)
+					return err;
+				if (bpf_is_unwind_kfunc(insn))
+					return process_bpf_unwind(env, &env->insn_idx,
+								  do_print_state);
+				/*
+				 * The fixups lower this to 'r0 = 0; exit', and
+				 * the unwind goes on below this frame.
+				 */
+				return unwind_frames(env, do_print_state);
+			}
 			mark_reg_scratched(env, BPF_REG_0);
 			if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
 				cur_func(env)->no_stack_arg_load = true;
 			if (bpf_is_callx(insn))
 				return check_func_callx(env, insn, &env->insn_idx);
-			if (insn->src_reg == BPF_PSEUDO_CALL)
-				return check_func_call(env, insn, &env->insn_idx);
+			if (insn->src_reg == BPF_PSEUDO_CALL) {
+				int call_idx = env->insn_idx;
+
+				err = check_func_call(env, insn, &env->insn_idx);
+				if (err)
+					return err;
+				return unwind_out_of_global_call(env, call_idx, do_print_state);
+			}
 			if (insn->src_reg == BPF_PSEUDO_KFUNC_CALL)
 				return check_kfunc_call(env, insn, &env->insn_idx);
 			return check_helper_call(env, insn, &env->insn_idx);
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (6 preceding siblings ...)
  2026-10-01 13:30 ` [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier Yonghong Song
@ 2026-10-01 13:30 ` Yonghong Song
  2026-10-01 14:31   ` bot+bpf-ci
  2026-10-03 12:25   ` Alexei Starovoitov
  2026-10-01 13:30 ` [PATCH bpf-next v8 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
                   ` (13 subsequent siblings)
  21 siblings, 2 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:30 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

A landing pad is compiler output for one frame: it drops what that frame
holds and resumes. Nothing in it knows about its caller's locks and
references, or its callees'. Hold every frame an unwind leaves to that:
record what the program holds when a frame is entered, and require an
unwind leaving the frame to have put it back.

This is not what keeps a pad sound. An unwind is followed into the
caller's pad from the state the callee left, so a pad is verified against
what really holds at that point, and whatever an abandoned frame keeps
still reaches the resource checks at the program's exit. What the rule
changes is where a program gets refused, and so what the message points
at. Say foo() takes an RCU read lock and calls bar(), which unwinds:

  foo:
      call bpf_rcu_read_lock
  1:  call bar                  /* covered, pad at 3 */
  2:  r0 = 0; exit
  3:  call bpf_rcu_read_unlock  /* pad: drops the lock foo() took */
      call bpf_unwind_resume

  bar:
  1:  call bpf_unwind           /* covered, pad at 3 */
  2:  r0 = 0; exit
  3:  call bpf_rcu_read_unlock  /* pad: drops it a second time */
      call bpf_unwind_resume

Without the rule, the unwind reaches foo's pad without the lock, and that
pad's unlock is refused. The refusal is right, but it blames foo; the
mistake is in bar. With the rule, bar's resume is refused for not leaving
the RCU state as it found it. A frame that takes a lock or a reference and
leaves through an unwind is worse: without the rule it is refused only at
the program's exit, as a leak or a held lock, with nothing pointing at the
frame that dropped it. The rule does refuse a program that could be proved
safe, a callee releasing its caller's reference and a caller's pad that
relies on that, but compiled code does not do that. A pad only knows about
its own frame.

Counts are enough for the locks. An unwind is a kfunc call, and those are
refused under a spin lock, so no unwind happens with one held and a
swapped spin lock is never seen. The RCU and preemption counts are nesting
depths, and the IRQ state is compared by id. References go by id, since
ids only go up and bpf_reference_state does not say which frame acquired
one.

The frames in between need the same of them, and have nothing to run: where
no record covers the call a frame is suspended at, the unwind sends it to
its epilogue, so what it acquired since it was entered is dropped on the
floor and no path of its own arrives to say so. Ask it at the call instead,
which is where it is abandoned -- check_unwind_through_call(). Per frame
rather than of the whole stack at the unwind, since a frame that does carry
a record may hold what its pad will release. A callx counts as any
subprogram that might unwind, its target not being known there.

So every frame an unwind leaves is asked, one way of asking per way out:

  - "a resume": a frame with a pad runs it and ends at bpf_unwind_resume().
    The frame that raised the unwind leaves this way where a record covers
    its bpf_unwind(), and so does every caller whose call is covered.
  - "an unwind with no landing pad": the frame that raised it with no
    record over its bpf_unwind(), returning through the exit patched in
    after it.
  - "an unwind through this call": every caller whose call no record
    covers, asked at the call since nothing of it runs again.

Nothing else is left to ask: a frame other than the one that raised the
unwind is suspended at a call, and the call either carries a record or does
not. The unwind reaches no further than the program it was raised in:
unwind_frames() stops at the main program's frame.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf_verifier.h | 11 ++++++
 kernel/bpf/exception.c       | 76 ++++++++++++++++++++++++++++++++++++
 kernel/bpf/exception.h       |  6 +++
 kernel/bpf/verifier.c        | 37 ++++++++++++++++++
 4 files changed, 130 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index c90fa3f5e787..a625d96a5b80 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -339,6 +339,17 @@ struct bpf_func_state {
 	bool in_async_callback_fn;
 	bool in_exception_callback_fn;
 	bool no_stack_arg_load;
+	/*
+	 * What the program held when this frame was entered. A frame an unwind
+	 * leaves has to have put these back: a diagnostic, which refuses the
+	 * frame at fault rather than a later pad or the program's exit.
+	 */
+	u32 entry_active_locks;
+	u32 entry_preempt_locks;
+	u32 entry_rcu_locks;
+	u32 entry_irq_id;
+	u32 entry_id_gen;
+	u32 entry_acquired_refs;
 	/* For callback calling functions that limit number of possible
 	 * callback executions (e.g. bpf_loop) keeps track of current
 	 * simulated iteration number.
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index e12cdb12cde3..8d48cf69bef0 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -141,6 +141,72 @@ int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
 BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
 BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
 
+void bpf_exc_record_frame_entry(const struct bpf_verifier_state *state,
+				struct bpf_func_state *frame, u32 id_gen)
+{
+	u32 i;
+
+	frame->entry_active_locks = state->active_locks;
+	frame->entry_preempt_locks = state->active_preempt_locks;
+	frame->entry_rcu_locks = state->active_rcu_locks;
+	frame->entry_irq_id = state->active_irq_id;
+
+	/* Ids only ever go up, so this one tells the frame's own apart. */
+	frame->entry_id_gen = id_gen;
+	frame->entry_acquired_refs = 0;
+	for (i = 0; i < state->acquired_refs; i++)
+		if (state->refs[i].type == REF_TYPE_PTR)
+			frame->entry_acquired_refs++;
+}
+
+int bpf_exc_check_frame_balance(struct bpf_verifier_env *env, const char *prefix)
+{
+	const struct bpf_verifier_state *state = env->cur_state;
+	const struct bpf_func_state *frame = cur_func(env);
+	u32 i, held;
+	const char *what;
+
+	if (state->active_rcu_locks != frame->entry_rcu_locks)
+		what = "bpf_rcu_read_lock";
+	else if (state->active_preempt_locks != frame->entry_preempt_locks)
+		what = "bpf_preempt_disable";
+	else if (state->active_irq_id != frame->entry_irq_id)
+		what = "bpf_local_irq_save";
+	else if (state->active_locks != frame->entry_active_locks)
+		what = "bpf_spin_lock";
+	else
+		what = NULL;
+
+	if (what) {
+		verbose(env, "%s does not leave the frame's %s state as it found it\n",
+			prefix, what);
+		return -EINVAL;
+	}
+
+	/*
+	 * References the same way. ids only go up, so entry_id_gen splits
+	 * refs[] in two at frame entry: nothing above that line may still be
+	 * held, and the count below it has to be what it was.
+	 */
+	for (i = 0, held = 0; i < state->acquired_refs; i++) {
+		if (state->refs[i].type != REF_TYPE_PTR)
+			continue;
+		if (state->refs[i].id > frame->entry_id_gen) {
+			verbose(env, "%s keeps the reference id=%d the frame acquired\n",
+				prefix, state->refs[i].id);
+			return -EINVAL;
+		}
+		held++;
+	}
+	if (held != frame->entry_acquired_refs) {
+		verbose(env, "%s does not leave the frame's references as it found it\n",
+			prefix);
+		return -EINVAL;
+	}
+
+	return 0;
+}
+
 static int reject_throw(struct bpf_verifier_env *env)
 {
 	u32 i;
@@ -174,6 +240,16 @@ static void mark_call_sites(struct bpf_verifier_env *env)
 	}
 }
 
+bool bpf_prog_may_unwind(const struct bpf_verifier_env *env)
+{
+	u32 i;
+
+	for (i = 0; i < env->subprog_cnt; i++)
+		if (env->subprog_info[i].might_unwind)
+			return true;
+	return false;
+}
+
 int bpf_exc_check_prog(struct bpf_verifier_env *env)
 {
 	int err;
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index 96dac3037d75..615df30fdfdb 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -8,12 +8,18 @@
 
 union bpf_attr;
 struct bpf_verifier_env;
+struct bpf_verifier_state;
+struct bpf_func_state;
 struct bpf_insn;
 
 int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
 		       bpfptr_t uattr);
 int bpf_exc_prepare(struct bpf_verifier_env *env);
 int bpf_exc_check_prog(struct bpf_verifier_env *env);
+bool bpf_prog_may_unwind(const struct bpf_verifier_env *env);
+void bpf_exc_record_frame_entry(const struct bpf_verifier_state *state,
+				struct bpf_func_state *frame, u32 id_gen);
+int bpf_exc_check_frame_balance(struct bpf_verifier_env *env, const char *prefix);
 int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
 bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
 bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index c7a350be538e..a7b25ab04051 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10807,6 +10807,7 @@ static int setup_func_entry(struct bpf_verifier_env *env, int subprog, int calls
 			callsite,
 			state->curframe + 1 /* frameno within this callchain */,
 			subprog /* subprog number within this prog */);
+	bpf_exc_record_frame_entry(state, callee, env->id_gen);
 	err = set_callee_state_cb(env, caller, callee, callsite);
 	if (err)
 		goto err_out;
@@ -19244,6 +19245,32 @@ static int unwind_out_of_global_call(struct bpf_verifier_env *env, int call_idx,
 	return INSN_IDX_UPDATED;
 }
 
+/* Can an unwind come back out of this call? */
+static bool call_may_unwind(struct bpf_verifier_env *env, const struct bpf_insn *insn,
+			    int insn_idx)
+{
+	int subprog;
+
+	/* Which subprog a callx lands in is not known here, so any may be it. */
+	if (bpf_is_callx(insn))
+		return bpf_prog_may_unwind(env);
+	if (insn->src_reg != BPF_PSEUDO_CALL)
+		return false;
+	subprog = bpf_find_subprog(env, insn_idx + insn->imm + 1);
+	return subprog >= 0 && env->subprog_info[subprog].might_unwind;
+}
+
+static int check_unwind_through_call(struct bpf_verifier_env *env, int insn_idx)
+{
+	const struct bpf_insn *insn = &env->prog->insnsi[insn_idx];
+
+	if (bpf_exc_pad_of_call(env, insn_idx) >= 0)
+		return 0;
+	if (!call_may_unwind(env, insn, insn_idx))
+		return 0;
+	return bpf_exc_check_frame_balance(env, "an unwind through this call");
+}
+
 static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
 			      bool *do_print_state)
 {
@@ -19258,6 +19285,9 @@ static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
 			if (err)
 				return err;
 		}
+		err = bpf_exc_check_frame_balance(env, "an unwind with no landing pad");
+		if (err)
+			return err;
 		return unwind_frames(env, do_print_state);
 	}
 	clear_caller_saved_regs(env, frame->regs);
@@ -19527,6 +19557,9 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
 				if (bpf_is_unwind_kfunc(insn))
 					return process_bpf_unwind(env, &env->insn_idx,
 								  do_print_state);
+				err = bpf_exc_check_frame_balance(env, "a resume");
+				if (err)
+					return err;
 				/*
 				 * The fixups lower this to 'r0 = 0; exit', and
 				 * the unwind goes on below this frame.
@@ -19534,6 +19567,10 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
 				return unwind_frames(env, do_print_state);
 			}
 			mark_reg_scratched(env, BPF_REG_0);
+			/* An unwind with no pad leaves the frame for good. */
+			err = check_unwind_through_call(env, env->insn_idx);
+			if (err)
+				return err;
 			if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
 				cur_func(env)->no_stack_arg_load = true;
 			if (bpf_is_callx(insn))
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 09/22] bpf: Refuse a landing pad that does not resume
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (7 preceding siblings ...)
  2026-10-01 13:30 ` [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
@ 2026-10-01 13:30 ` Yonghong Song
  2026-10-03 12:25   ` Alexei Starovoitov
  2026-10-01 13:30 ` [PATCH bpf-next v8 10/22] bpf: Do not use a private stack for a program that can unwind Yonghong Song
                   ` (12 subsequent siblings)
  21 siblings, 1 reply; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:30 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

A cleanup pad runs drop glue and calls bpf_unwind_resume(), so the frame
returns and the unwind goes on. A catch pad runs the same drops and then
carries on in its frame, stopping the unwind. Only the first is supported:
bpf_unwind() rewrites every frame's return address in one pass, so a
caller of a catch pad's frame would resume at a pad for an unwind already
caught.

Nothing in the record says which kind a pad is, but the code does, as LLVM
emits it: a cleanup pad reaches _Unwind_Resume, a catch pad reaches a
return. An unwind arriving at a pad, in the frame that raised it, in a
caller, or out of a call to a global subprog, is the only way in, and it
marks the frame it enters. bpf_exc_check_insn() asks that mark about every
instruction of a program carrying a table, and refuses:

 - an exit, which is how a catch pad ends
 - a tail call, and a BPF_LD_[ABS|IND], which leaves through an exit on a
   failed load
 - an indirect jump
 - a bpf_unwind(), and a call to a global subprogram that might_unwind
 - an instruction reached both inside and outside a pad

do_check_insn() refuses the other half of it, a bpf_unwind_resume() the
mark does not find in a pad. That one is asked of every program rather
than only those carrying a table, since a program with no table has no pad
to be in.

The first three concern the pad's own frame, since a subprogram it calls
may do any of them and still come back; an unwind is refused in a pad's
callees too, since it never does. do_check() asks before pruning, so the
mark stays out of states_equal().

A speculative walk can reach a pad too; that is answered as do_check()
answers anything it cannot allow speculatively, by marking the instruction
for a barrier and stopping rather than refusing the program. Such a visit
leaves no in-pad or outside-pad mark for a real path to be refused over,
and an exit it finds in a pad gets the barrier too. A callback that can
unwind is refused as well, its helper frame being C with no pad.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf_verifier.h |  4 ++
 kernel/bpf/exception.c       | 88 ++++++++++++++++++++++++++++++++++++
 kernel/bpf/exception.h       |  2 +
 kernel/bpf/verifier.c        | 28 ++++++++++++
 4 files changed, 122 insertions(+)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index a625d96a5b80..ccac422737fb 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -339,6 +339,8 @@ struct bpf_func_state {
 	bool in_async_callback_fn;
 	bool in_exception_callback_fn;
 	bool no_stack_arg_load;
+	/* an unwind reached this frame and its landing pad is running */
+	bool in_pad;
 	/*
 	 * What the program held when this frame was entered. A frame an unwind
 	 * leaves has to have put these back: a diagnostic, which refuses the
@@ -716,6 +718,8 @@ struct bpf_insn_aux_data {
 	u64 non_stack_access:1; /* instruction can access non-stack memory */
 	/* true if some jump or call instruction targets this instruction */
 	u64 jump_target:1;
+	u64 in_cleanup_pad:1; /* reached with a landing pad running */
+	u64 outside_cleanup_pad:1; /* reached the other way */
 
 	unsigned int orig_idx; /* original instruction index, initialized once */
 	/*
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index 8d48cf69bef0..e824883d3981 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -141,6 +141,15 @@ int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
 BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
 BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
 
+int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog)
+{
+	if (!env->subprog_info[subprog].might_unwind)
+		return 0;
+
+	verbose(env, "subprog %d may unwind and is used as a callback\n", subprog);
+	return -EINVAL;
+}
+
 void bpf_exc_record_frame_entry(const struct bpf_verifier_state *state,
 				struct bpf_func_state *frame, u32 id_gen)
 {
@@ -308,6 +317,85 @@ bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
 	       insn->imm == bpf_unwind_resume_id[0];
 }
 
+/* Is an unwind in flight: is this frame running a pad, or called from one? */
+static bool unwinding(const struct bpf_verifier_state *state)
+{
+	u32 i;
+
+	for (i = 0; i <= state->curframe; i++)
+		if (state->frame[i]->in_pad)
+			return true;
+	return false;
+}
+
+int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn)
+{
+	bool in_pad = cur_func(env)->in_pad;
+	struct bpf_insn_aux_data *aux;
+	u32 i = env->insn_idx;
+	const char *why = NULL;
+
+	if (unwinding(env->cur_state)) {
+		if (bpf_is_unwind_kfunc(insn)) {
+			verbose(env, "insn %u starts a second unwind while one is in flight\n", i);
+			return -EINVAL;
+		}
+		if (bpf_pseudo_call(insn)) {
+			int subprog = bpf_find_subprog(env, i + insn->imm + 1);
+
+			if (subprog >= 0 && bpf_subprog_is_global(env, subprog) &&
+			    env->subprog_info[subprog].might_unwind) {
+				verbose(env,
+					"insn %u calls global subprog %d, which can unwind while an unwind is in flight\n",
+					i, subprog);
+				return -EINVAL;
+			}
+		}
+	}
+
+	aux = &env->insn_aux_data[i];
+
+	if (in_pad ? aux->outside_cleanup_pad : aux->in_cleanup_pad) {
+		verbose(env, "insn %u runs both inside and outside a landing pad\n", i);
+		return -EINVAL;
+	}
+	/*
+	 * Only a real path marks the insn: a speculative one that finds the
+	 * other mark gets a barrier, so it must not leave one for a real path
+	 * to be refused over.
+	 */
+	if (!env->cur_state->speculative) {
+		if (in_pad)
+			aux->in_cleanup_pad = true;
+		else
+			aux->outside_cleanup_pad = true;
+	}
+
+	if (!in_pad)
+		return 0;
+
+	if (insn->code == (BPF_JMP | BPF_EXIT)) {
+		verbose(env,
+			"exit at insn %u ends a landing pad: a catch pad is not supported yet, only cleanup pads that resume\n",
+			i);
+		return -EOPNOTSUPP;
+	}
+	if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call)
+		why = "is a tail call, which replaces the frame";
+	else if (BPF_CLASS(insn->code) == BPF_LD &&
+		 (BPF_MODE(insn->code) == BPF_ABS || BPF_MODE(insn->code) == BPF_IND))
+		why = "is a BPF_LD_[ABS|IND], which can leave through the epilogue";
+	else if (insn->code == (BPF_JMP | BPF_JA | BPF_X) ||
+		 insn->code == (BPF_JMP32 | BPF_JA | BPF_X))
+		why = "is an indirect jump";
+
+	if (!why)
+		return 0;
+
+	verbose(env, "insn %u %s, and is in a landing pad\n", i, why);
+	return -EINVAL;
+}
+
 int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
 {
 	u32 pad = env->insn_aux_data[idx].cleanup_pad;
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index 615df30fdfdb..e72e68ebfe85 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -23,5 +23,7 @@ int bpf_exc_check_frame_balance(struct bpf_verifier_env *env, const char *prefix
 int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx);
 bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
 bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
+int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog);
+int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn);
 
 #endif /* __BPF_EXCEPTION_H */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index a7b25ab04051..f3ed68960d70 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10945,6 +10945,10 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
 	 * callbacks
 	 */
 	env->subprog_info[subprog].is_cb = true;
+	err = bpf_exc_check_callback(env, subprog);
+	if (err)
+		return err;
+
 	if (bpf_pseudo_kfunc_call(insn) &&
 	    !is_callback_calling_kfunc(insn->imm)) {
 		verifier_bug(env, "kfunc %s#%d not marked as callback-calling",
@@ -19163,6 +19167,10 @@ static int unwind_frames(struct bpf_verifier_env *env, bool *do_print_state)
 	while (state->curframe) {
 		callee = cur_func(env);
 		caller = state->frame[state->curframe - 1];
+		/* A subprog that can unwind is refused as a callback. */
+		if (verifier_bug_if(callee->in_callback_fn, env,
+				    "unwind out of callback frame %d", state->curframe))
+			return -EFAULT;
 		pad = bpf_exc_pad_of_call(env, callee->callsite);
 		/* The caller is at its call now, not at this frame's insn. */
 		state->insn_idx = callee->callsite;
@@ -19182,6 +19190,7 @@ static int unwind_frames(struct bpf_verifier_env *env, bool *do_print_state)
 			return err;
 		clear_caller_saved_regs(env, caller->regs);
 		mark_reg_unknown(env, caller->regs, BPF_REG_0);
+		caller->in_pad = true;
 		env->insn_idx = pad;
 		*do_print_state = true;
 		return INSN_IDX_UPDATED;
@@ -19240,6 +19249,7 @@ static int unwind_out_of_global_call(struct bpf_verifier_env *env, int call_idx,
 	frame = cur_func(env);
 	clear_caller_saved_regs(env, frame->regs);
 	mark_reg_unknown(env, frame->regs, BPF_REG_0);
+	frame->in_pad = true;
 	env->insn_idx = pad;
 	*do_print_state = true;
 	return INSN_IDX_UPDATED;
@@ -19292,6 +19302,7 @@ static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
 	}
 	clear_caller_saved_regs(env, frame->regs);
 	mark_reg_unknown(env, frame->regs, BPF_REG_0);
+	frame->in_pad = true;
 	*insn_idx = pad;
 	return INSN_IDX_UPDATED;
 }
@@ -19557,6 +19568,11 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
 				if (bpf_is_unwind_kfunc(insn))
 					return process_bpf_unwind(env, &env->insn_idx,
 								  do_print_state);
+				if (!cur_func(env)->in_pad) {
+					verbose(env, "resume at insn %d is not in a landing pad\n",
+						env->insn_idx);
+					return -EINVAL;
+				}
 				err = bpf_exc_check_frame_balance(env, "a resume");
 				if (err)
 					return err;
@@ -19681,6 +19697,18 @@ static int do_check(struct bpf_verifier_env *env)
 			}
 		}
 
+		if (unlikely(env->cleanup_info_cnt)) {
+			err = bpf_exc_check_insn(env, insn);
+			/* An exit in a pad is refused as unsupported, not invalid. */
+			if ((error_recoverable_with_nospec(err) || err == -EOPNOTSUPP) &&
+			    state->speculative) {
+				insn_aux->nospec = true;
+				goto process_bpf_exit;
+			}
+			if (err)
+				return err;
+		}
+
 		if (bpf_is_prune_point(env, env->insn_idx)) {
 			err = bpf_is_state_visited(env, env->insn_idx);
 			if (err < 0)
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 10/22] bpf: Do not use a private stack for a program that can unwind
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (8 preceding siblings ...)
  2026-10-01 13:30 ` [PATCH bpf-next v8 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
@ 2026-10-01 13:30 ` Yonghong Song
  2026-10-01 13:53   ` sashiko-bot
  2026-10-01 13:31 ` [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
                   ` (11 subsequent siblings)
  21 siblings, 1 reply; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:30 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

A private stack keeps its frame pointer in %r9 on x86-64, and the JIT
brackets every call in push_r9/pop_r9. An unwind skips the pop: a frame
resumed at its landing pad addresses its stack through a stale pointer,
and one sent to its epilogue pops its callee-saved registers one slot off.

So no private stack for a program that can unwind, on every architecture
rather than just that one. A table is not the only condition: a
bpf_unwind() with no record over it sends its callers to their epilogues
just the same. In check_max_stack_depth(), force NO_PRIV_STACK so the JIT
does not use a private stack.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 kernel/bpf/verifier.c | 11 +++++++++++
 1 file changed, 11 insertions(+)

diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index f3ed68960d70..488ceb9dae1b 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -5769,6 +5769,17 @@ static int check_max_stack_depth(struct bpf_verifier_env *env)
 		}
 	}
 
+	/*
+	 * A private stack keeps its frame pointer in %r9 on x86-64, restored
+	 * by a pop after the call that an unwind skips. A frame resumed at a
+	 * pad then addresses its stack through a stale pointer, and a frame
+	 * sent to its epilogue instead pops its callee-saved registers one
+	 * slot off. Refuse a private stack for any program that can unwind,
+	 * on every arch for now.
+	 */
+	if (env->cleanup_info_cnt || bpf_prog_may_unwind(env))
+		priv_stack_mode = NO_PRIV_STACK;
+
 	if (priv_stack_mode == PRIV_STACK_UNKNOWN)
 		priv_stack_mode = bpf_enable_priv_stack(env->prog);
 
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (9 preceding siblings ...)
  2026-10-01 13:30 ` [PATCH bpf-next v8 10/22] bpf: Do not use a private stack for a program that can unwind Yonghong Song
@ 2026-10-01 13:31 ` Yonghong Song
  2026-10-01 14:31   ` bot+bpf-ci
  2026-10-03 12:26   ` Alexei Starovoitov
  2026-10-01 13:31 ` [PATCH bpf-next v8 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
                   ` (10 subsequent siblings)
  21 siblings, 2 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:31 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

bpf_unwind() walks the BPF frames and, for each one above the frame that
called it, rewrites the saved return address so the frame resumes where
the unwind needs it, then returns. The frame that called it goes on after
the call, where the fixups put 'r0 = 0' and then a jump to its pad, or an
exit where no record covers the call, so the pad starts with r0 at a known
zero rather than whatever bpf_unwind() left in the return register.
bpf_unwind() restores no register itself: every frame runs its own
epilogue on the way out, which is what puts its caller's r6-r9 back, so
the unwind needs no spill area and no per-frame metadata beyond the table
itself.

Where a record covers the call a frame is suspended at, it resumes at that
pad, which an earlier patch made sure ends in a resume. Where none does,
it resumes at the frame's epilogue and returns at once, its caller reached
with registers already restored. A pad's resume lowers to 'r0 = 0; exit',
so that frame returns too and the rewritten address carries the unwind on
to the next pad.

For the verifier patch's example, main -> A -> B -> C with only A's call
to B covered, bpf_unwind() in C rewrites the return addresses it finds:

  slot                    points into               rewritten to
  ----------------------  ------------------------  ---------------------
  bpf_unwind()'s return   C, after its unwind call  left alone
  C's return              B, after 'call C'         B's epilogue
  B's return              A, after 'call B'         P, A's pad
  A's return              main, after 'call A'      main's epilogue
  main's return           the kernel                left alone

Each address is looked up in the frame it points into: a record over that
call gives its pad, none gives that frame's epilogue. Then every frame
just returns:

  frame   runs
  -----   ----------------------------------------------------------
  C       'r0 = 0; exit', patched in after its bpf_unwind()
  B       its epilogue, putting back A's r6-r9
  A       P, which drops A's resources; its resume is 'r0 = 0; exit'
  main    its epilogue, returning 0 to the kernel

An epilogue therefore has to exist for every frame the walk can pass, not
only for those carrying a table: the JITs record aux->epilogue_ip for every
program in the patches that follow, and this one hands it to the outer
program with the table when jit_subprogs() compiles the main program as
func[0].

x86 emits the epilogue at a subprogram's first exit, so one the dead code
sweep leaves exitless gets none. Two shapes do that. The frame that called
bpf_unwind() loses the code after the call; the exit patched back in there
where no record covers the call is also what that frame returns through.
And a frame above one that never comes back loses its exit with
no bpf_unwind() to hang a new one on, so the last exit of every subprogram
an unwind can pass through is kept, searched for since a subprogram may
end in a jump or a gotox. One with no exit at all is refused: nothing is
left an epilogue could be emitted at. arm64 emits an epilogue either way.

Both kfuncs become callable here rather than earlier: until the walk and
the lowering exist, bpf_unwind() would return to instructions the verifier
never explored and bpf_unwind_resume() would reach its WARN_ONCE body.

arch_bpf_stack_walk_ra() hands out the return-address slot as well as the
address. It is a second entry point rather than a change to
arch_bpf_stack_walk(), so architectures that do not dispatch pads keep the
walker they have -- where it is the weak stub, the walk does nothing, so
process_bpf_unwind() now asks bpf_exc_check_prog() whether this program may
unwind at all. A cleanup table was held to that before the CFG walk; a
bpf_unwind() with no table had not been.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 include/linux/bpf.h          |  37 ++++++++++
 include/linux/bpf_verifier.h |   1 +
 include/linux/filter.h       |   2 +
 kernel/bpf/core.c            |  20 ++++-
 kernel/bpf/exception.c       |  86 ++++++++++++++++++++++
 kernel/bpf/exception.h       |   6 ++
 kernel/bpf/fixups.c          | 138 +++++++++++++++++++++++++++++++++++
 kernel/bpf/helpers.c         |  45 ++++++++++++
 kernel/bpf/verifier.c        |   7 ++
 9 files changed, 341 insertions(+), 1 deletion(-)

diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 4bae3796c42f..94005cd3ad0f 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -1805,6 +1805,41 @@ enum bpf_sig_keyring {
 	BPF_SIG_KEYRING_BPF,
 };
 
+/* One cleanup region of a JITed (sub)program. */
+struct bpf_cleanup_range {
+	u64 begin;
+	u64 end;
+	u64 pad;
+};
+
+struct bpf_exception_info {
+	struct bpf_cleanup_info *info;
+	struct bpf_cleanup_range *ranges;
+	u32 nr_info;
+	u32 nr_ranges;
+};
+
+#ifdef CONFIG_BPF_SYSCALL
+int bpf_exc_attach_main_prog(struct bpf_verifier_env *env, struct bpf_prog *prog);
+void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image);
+void bpf_exc_free_info(struct bpf_prog_aux *aux);
+#else
+
+static inline int bpf_exc_attach_main_prog(struct bpf_verifier_env *env,
+					   struct bpf_prog *prog)
+{
+	return 0;
+}
+
+static inline void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+}
+
+static inline void bpf_exc_free_info(struct bpf_prog_aux *aux)
+{
+}
+#endif
+
 struct bpf_prog_aux {
 	atomic64_t refcnt;
 	u32 used_map_cnt;
@@ -1885,6 +1920,8 @@ struct bpf_prog_aux {
 	u64 (*bpf_exception_cb)(u64 cookie, u64 sp, u64 bp, u64, u64);
 	u16 stack_arg_sp_adjust;
 	u16 freplace_link_cnt; /* counts freplace links extending this prog */
+	struct bpf_exception_info *exc;
+	u64 epilogue_ip; /* native address of this (sub)program's epilogue */
 #ifdef CONFIG_SECURITY
 	void *security;
 #endif
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index ccac422737fb..7cede13f8bee 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1863,6 +1863,7 @@ int bpf_opt_subreg_zext_lo32_rnd_hi32(struct bpf_verifier_env *env, const union
 int bpf_convert_ctx_accesses(struct bpf_verifier_env *env);
 int bpf_jit_subprogs(struct bpf_verifier_env *env);
 int bpf_fixup_call_args(struct bpf_verifier_env *env);
+int bpf_exc_patch_unwind_calls(struct bpf_verifier_env *env);
 int bpf_do_misc_fixups(struct bpf_verifier_env *env);
 int bpf_insn_def32(struct bpf_prog *prog, struct bpf_insn *insn);
 
diff --git a/include/linux/filter.h b/include/linux/filter.h
index 972b3ed2a51d..0d7d949a1baa 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1290,6 +1290,8 @@ u32 bpf_jit_plan_arg_moves(const struct bpf_jit_arg_abi *abi,
 			   struct bpf_jit_arg_move *moves);
 u64 bpf_arch_uaddress_limit(void);
 void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp), void *cookie);
+void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
+			    void *cookie);
 u64 arch_bpf_timed_may_goto(void);
 u64 bpf_check_timed_may_goto(struct bpf_timed_may_goto *);
 bool bpf_helper_changes_pkt_data(enum bpf_func_id func_id);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index d813fdde29e3..60905643cb9c 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -292,6 +292,7 @@ void __bpf_prog_free(struct bpf_prog *fp)
 		mutex_destroy(&fp->aux->dst_mutex);
 		mutex_destroy(&fp->aux->st_ops_assoc_mutex);
 		kfree(fp->aux->poke_tab);
+		bpf_exc_free_info(fp->aux);
 		kfree(fp->aux);
 	}
 	free_percpu(fp->stats);
@@ -2632,9 +2633,14 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
 {
 #ifdef CONFIG_BPF_JIT
 	struct bpf_prog *orig_prog;
+	int ret;
 
-	if (!bpf_prog_need_blind(prog))
+	if (!bpf_prog_need_blind(prog)) {
+		ret = bpf_exc_attach_main_prog(env, prog);
+		if (ret)
+			return ERR_PTR(ret);
 		return bpf_int_jit_compile(env, prog);
+	}
 
 	orig_prog = prog;
 	prog = bpf_jit_blind_constants(env, prog);
@@ -2648,6 +2654,12 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
 		goto out_restore;
 	}
 
+	ret = bpf_exc_attach_main_prog(env, prog);
+	if (ret) {
+		bpf_jit_prog_release_other(orig_prog, prog);
+		return ERR_PTR(ret);
+	}
+
 	prog = bpf_int_jit_compile(env, prog);
 	if (prog->jited) {
 		bpf_jit_prog_release_other(prog, orig_prog);
@@ -3511,6 +3523,12 @@ void __weak arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp,
 {
 }
 
+void __weak arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp,
+						      u64 *ra),
+				   void *cookie)
+{
+}
+
 bool __weak bpf_jit_supports_cleanup_pads(void)
 {
 	return false;
diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
index e824883d3981..bbe64887e069 100644
--- a/kernel/bpf/exception.c
+++ b/kernel/bpf/exception.c
@@ -402,3 +402,89 @@ int bpf_exc_pad_of_call(struct bpf_verifier_env *env, u32 idx)
 
 	return pad ? (int)pad - 1 : -1;
 }
+
+/*
+ * The record covering @ip, which is a return address: the call it belongs to
+ * is the instruction before it, so a range matches on begin < ip <= end.
+ */
+const struct bpf_cleanup_range *bpf_exc_pad_for_ip(const struct bpf_prog *prog, u64 ip)
+{
+	const struct bpf_exception_info *exc = prog->aux->exc;
+	u32 l = 0, r = exc ? exc->nr_ranges : 0;
+
+	while (l < r) {
+		u32 m = l + (r - l) / 2;
+		const struct bpf_cleanup_range *rec = &exc->ranges[m];
+
+		if (ip <= rec->begin)
+			r = m;
+		else if (ip > rec->end)
+			l = m + 1;
+		else
+			return rec;
+	}
+	return NULL;
+}
+
+int bpf_exc_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt)
+{
+	struct bpf_cleanup_range *ranges;
+	struct bpf_exception_info *exc;
+
+	exc = kzalloc_obj(struct bpf_exception_info, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+	ranges = kvcalloc(cnt, sizeof(*ranges), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+	if (!exc || !ranges) {
+		kfree(exc);
+		kvfree(ranges);
+		kvfree(recs);
+		return -ENOMEM;
+	}
+
+	exc->info = recs;
+	exc->nr_info = cnt;
+	exc->ranges = ranges;
+	/* Withheld until the JIT has filled the table in. */
+	exc->nr_ranges = 0;
+	aux->exc = exc;
+	return 0;
+}
+
+void bpf_exc_fill_native_ranges(struct bpf_prog *prog, u32 *addrs, void *image)
+{
+	struct bpf_exception_info *exc = prog->aux->exc;
+	u32 i, n;
+
+	if (!exc)
+		return;
+
+	n = exc->nr_info;
+	for (i = 0; i < n; i++) {
+		const struct bpf_cleanup_info *rec = &exc->info[i];
+
+		/*
+		 * exc_info_for_subprog() built the records from insn_aux_data
+		 * inside this subprog, so this cannot fire; if it does, no
+		 * pad is dispatched rather than one read past addrs[].
+		 */
+		if (WARN_ON_ONCE(rec->begin_off >= prog->len ||
+				 rec->end_off > prog->len ||
+				 rec->landing_pad_off >= prog->len))
+			return;
+		exc->ranges[i].begin = (u64)(long)image + addrs[rec->begin_off];
+		exc->ranges[i].end = (u64)(long)image + addrs[rec->end_off];
+		exc->ranges[i].pad = (u64)(long)image + addrs[rec->landing_pad_off];
+	}
+	exc->nr_ranges = n;
+}
+
+void bpf_exc_free_info(struct bpf_prog_aux *aux)
+{
+	struct bpf_exception_info *exc = aux->exc;
+
+	if (!exc)
+		return;
+	kvfree(exc->ranges);
+	kvfree(exc->info);
+	kfree(exc);
+	aux->exc = NULL;
+}
diff --git a/kernel/bpf/exception.h b/kernel/bpf/exception.h
index e72e68ebfe85..b97ac04785c3 100644
--- a/kernel/bpf/exception.h
+++ b/kernel/bpf/exception.h
@@ -11,6 +11,10 @@ struct bpf_verifier_env;
 struct bpf_verifier_state;
 struct bpf_func_state;
 struct bpf_insn;
+struct bpf_cleanup_info;
+struct bpf_cleanup_range;
+struct bpf_prog;
+struct bpf_prog_aux;
 
 int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
 		       bpfptr_t uattr);
@@ -25,5 +29,7 @@ bool bpf_is_unwind_kfunc(const struct bpf_insn *insn);
 bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn);
 int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog);
 int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn);
+int bpf_exc_attach_info(struct bpf_prog_aux *aux, struct bpf_cleanup_info *recs, u32 cnt);
+const struct bpf_cleanup_range *bpf_exc_pad_for_ip(const struct bpf_prog *prog, u64 ip);
 
 #endif /* __BPF_EXCEPTION_H */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 5b7fe4ba610b..fc1d98eddf36 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -11,6 +11,7 @@
 #include <linux/sched/signal.h>
 #include <net/xdp.h>
 #include "disasm.h"
+#include "exception.h"
 
 #define verbose(env, fmt, args...) bpf_verifier_log_write(env, fmt, ##args)
 
@@ -739,6 +740,32 @@ static void keep_funcs_with_addr_taken(struct bpf_verifier_env *env)
 	}
 }
 
+static int keep_subprog_exits(struct bpf_verifier_env *env)
+{
+	u32 i, j;
+
+	for (i = 0; i < env->subprog_cnt; i++) {
+		bool found = false;
+		u32 start;
+
+		if (!env->subprog_info[i].might_unwind)
+			continue;
+		start = env->subprog_info[i].start;
+		for (j = env->subprog_info[i + 1].start; j-- > start; ) {
+			if (env->prog->insnsi[j].code != (BPF_JMP | BPF_EXIT))
+				continue;
+			env->insn_aux_data[j].seen = env->pass_cnt;
+			found = true;
+			break;
+		}
+		if (!found) {
+			verbose(env, "subprog %u can be unwound through but has no exit\n", i);
+			return -EINVAL;
+		}
+	}
+	return 0;
+}
+
 int bpf_opt_remove_dead_code(struct bpf_verifier_env *env)
 {
 	struct bpf_insn_aux_data *aux_data = env->insn_aux_data;
@@ -746,6 +773,9 @@ int bpf_opt_remove_dead_code(struct bpf_verifier_env *env)
 	int i, err;
 
 	keep_funcs_with_addr_taken(env);
+	err = keep_subprog_exits(env);
+	if (err)
+		return err;
 
 	for (i = 0; i < insn_cnt; i++) {
 		int j;
@@ -1286,6 +1316,53 @@ static int resolve_func_ptrs(struct bpf_verifier_env *env, struct bpf_prog *prog
 	return 0;
 }
 
+static int exc_info_for_subprog(struct bpf_verifier_env *env, struct bpf_prog *sub,
+				u32 start, u32 end)
+{
+	struct bpf_cleanup_info *recs;
+	u32 i, cnt = 0;
+
+	if (!env->cleanup_info_cnt)
+		return 0;
+
+	for (i = start; i < end; i++) {
+		if (env->insn_aux_data[i].cleanup_pad)
+			cnt++;
+	}
+	if (!cnt)
+		return 0;
+
+	recs = kvmalloc_array(cnt, sizeof(*recs), GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
+	if (!recs)
+		return -ENOMEM;
+
+	for (i = start, cnt = 0; i < end; i++) {
+		u32 pad = env->insn_aux_data[i].cleanup_pad;
+
+		if (!pad)
+			continue;
+		pad--;
+		if (verifier_bug_if(pad < start || pad >= end, env,
+				    "insn %u is covered by a landing pad at %u outside its subprog [%u, %u)",
+				    i, pad, start, end)) {
+			kvfree(recs);
+			return -EFAULT;
+		}
+		recs[cnt].begin_off = i - start;
+		recs[cnt].end_off = i - start + 1;
+		recs[cnt].landing_pad_off = pad - start;
+		cnt++;
+	}
+	return bpf_exc_attach_info(sub->aux, recs, cnt);
+}
+
+int bpf_exc_attach_main_prog(struct bpf_verifier_env *env, struct bpf_prog *prog)
+{
+	if (!env || env->subprog_cnt > 1)
+		return 0;
+	return exc_info_for_subprog(env, prog, 0, prog->len);
+}
+
 static int jit_subprogs(struct bpf_verifier_env *env)
 {
 	struct bpf_prog *prog = env->prog, **func, *tmp;
@@ -1423,6 +1500,8 @@ static int jit_subprogs(struct bpf_verifier_env *env)
 		func[i]->aux->token = prog->aux->token;
 		if (!i)
 			func[i]->aux->exception_boundary = env->seen_exception;
+		if (exc_info_for_subprog(env, func[i], subprog_start, subprog_end))
+			goto out_free;
 		func[i] = bpf_int_jit_compile(env, func[i]);
 		if (!func[i]->jited) {
 			err = -ENOTSUPP;
@@ -1532,6 +1611,9 @@ static int jit_subprogs(struct bpf_verifier_env *env)
 	prog->aux->bpf_exception_cb = (void *)func[env->exception_callback_subprog]->bpf_func;
 	prog->aux->exception_boundary = func[0]->aux->exception_boundary;
 	prog->aux->stack_arg_sp_adjust = func[0]->aux->stack_arg_sp_adjust;
+	prog->aux->exc = func[0]->aux->exc;
+	func[0]->aux->exc = NULL;
+	prog->aux->epilogue_ip = func[0]->aux->epilogue_ip;
 	bpf_prog_jit_attempt_done(prog);
 	return 0;
 out_free:
@@ -1757,6 +1839,43 @@ static int may_goto_expand(struct bpf_insn *insn_buf, int off, int stack_off,
 	return cnt + tail_cnt;
 }
 
+/*
+ * Follow each bpf_unwind() call with 'r0 = 0; exit', or with
+ * 'r0 = 0; goto pad' where a record covers the call.
+ */
+int bpf_exc_patch_unwind_calls(struct bpf_verifier_env *env)
+{
+	int insn_cnt = env->prog->len;
+	struct bpf_insn insn_buf[3];
+	struct bpf_prog *new_prog;
+	int i, off, delta = 0;
+
+	for (i = 0; i < insn_cnt; i++) {
+		struct bpf_insn *insn = env->prog->insnsi + i + delta;
+		u32 pad = env->insn_aux_data[i + delta].cleanup_pad;
+
+		if (!bpf_is_unwind_kfunc(insn))
+			continue;
+
+		insn_buf[0] = *insn;
+		insn_buf[1] = BPF_MOV64_IMM(BPF_REG_0, 0);
+		insn_buf[2] = BPF_EXIT_INSN();
+		if (pad) {
+			/* Stored as index + 1; a pad after the call moves with it. */
+			pad--;
+			off = (pad > i + delta ? pad + 2 : pad) - (i + delta + 3);
+			insn_buf[2] = off == (s16)off ? BPF_JMP_A(off) : BPF_JMP32_A(off);
+		}
+
+		new_prog = bpf_patch_insn_data(env, i + delta, insn_buf, 3);
+		if (!new_prog)
+			return -ENOMEM;
+		delta += 2;
+		env->prog = new_prog;
+	}
+	return 0;
+}
+
 /* Do various post-verification rewrites in a single program pass.
  * These rewrites simplify JIT and interpreter implementations.
  */
@@ -2135,6 +2254,25 @@ int bpf_do_misc_fixups(struct bpf_verifier_env *env)
 			goto next_insn;
 		if (insn->src_reg == BPF_PSEUDO_CALL)
 			goto next_insn;
+		if (bpf_is_unwind_resume_kfunc(insn)) {
+			/*
+			 * A pad's resume is just the frame returning, to
+			 * where bpf_unwind() pointed its return address: its
+			 * caller's pad or epilogue, or the kernel from the main
+			 * program. The verifier checked this exit with r0 a
+			 * known zero, so return zero.
+			 */
+			insn_buf[0] = BPF_MOV64_IMM(BPF_REG_0, 0);
+			insn_buf[1] = BPF_EXIT_INSN();
+			cnt = 2;
+			new_prog = bpf_patch_insn_data(env, i + delta, insn_buf, cnt);
+			if (!new_prog)
+				return -ENOMEM;
+			delta += cnt - 1;
+			env->prog = prog = new_prog;
+			insn = new_prog->insnsi + i + delta;
+			goto next_insn;
+		}
 		if (insn->src_reg == BPF_PSEUDO_KFUNC_CALL) {
 			ret = bpf_fixup_kfunc_call(env, insn, insn_buf, i + delta, &cnt);
 			if (ret)
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 4eccd6742eba..c6894d6185ab 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -31,6 +31,7 @@
 #include <linux/buildid.h>
 
 #include "../../lib/kstrtox.h"
+#include "exception.h"
 
 /* If kernel subsystem is allowing eBPF programs to call this function,
  * inside its own verifier_ops->get_func_proto() callback it should return
@@ -3424,8 +3425,50 @@ static bool bpf_stack_walker(void *cookie, u64 ip, u64 sp, u64 bp)
 	return false;
 }
 
+struct bpf_unwind_ctx {
+	u32 cnt;
+};
+
+static bool bpf_unwind_rewrite(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra)
+{
+	const struct bpf_cleanup_range *rec;
+	struct bpf_unwind_ctx *ctx = cookie;
+	struct bpf_prog *prog;
+
+	rcu_read_lock();
+	prog = bpf_prog_ksym_find(ip);
+	rcu_read_unlock();
+	if (!prog)
+		return !ctx->cnt;
+	ctx->cnt++;
+
+	/*
+	 * The frame that called bpf_unwind(): bpf_exc_patch_unwind_calls()
+	 * put 'r0 = 0' and a jump to its pad, or an exit, after the call,
+	 * so leave its return address alone and let it go on there. The pad
+	 * then starts with r0 at a known zero.
+	 */
+	if (ctx->cnt == 1)
+		return bpf_is_subprog(prog);
+
+	rec = bpf_exc_pad_for_ip(prog, ip);
+	if (rec) {
+		*ra = rec->pad;
+	} else if (prog->aux->epilogue_ip) {
+		*ra = prog->aux->epilogue_ip;
+	} else {
+		WARN_ON_ONCE(1);
+		return false;
+	}
+
+	return bpf_is_subprog(prog);
+}
+
 __bpf_kfunc void bpf_unwind(void)
 {
+	struct bpf_unwind_ctx ctx = {};
+
+	arch_bpf_stack_walk_ra(bpf_unwind_rewrite, &ctx);
 }
 
 __bpf_kfunc void bpf_throw(u64 cookie)
@@ -5095,6 +5138,8 @@ BTF_ID_FLAGS(func, bpf_task_get_cgroup1, KF_ACQUIRE | KF_RCU | KF_RET_NULL)
 BTF_ID_FLAGS(func, bpf_task_from_pid, KF_ACQUIRE | KF_RET_NULL)
 BTF_ID_FLAGS(func, bpf_task_from_vpid, KF_ACQUIRE | KF_RET_NULL)
 BTF_ID_FLAGS(func, bpf_throw)
+BTF_ID_FLAGS(func, bpf_unwind)
+BTF_ID_FLAGS(func, bpf_unwind_resume)
 #ifdef CONFIG_BPF_EVENTS
 BTF_ID_FLAGS(func, bpf_send_signal_task)
 #endif
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 488ceb9dae1b..917635adb5f6 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -19299,6 +19299,10 @@ static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx,
 	int pad = bpf_exc_pad_of_call(env, *insn_idx);
 	int err;
 
+	err = bpf_exc_check_prog(env);
+	if (err)
+		return err;
+
 	if (pad < 0) {
 		if (!env->cur_state->curframe) {
 			err = check_resource_leak(env, false, true,
@@ -22898,6 +22902,9 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 		/* program is valid, convert *(u32*)(ctx + off) accesses */
 		ret = bpf_convert_ctx_accesses(env);
 
+	if (ret == 0)
+		ret = bpf_exc_patch_unwind_calls(env);
+
 	if (ret == 0)
 		ret = bpf_do_misc_fixups(env);
 
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 12/22] bpf, x86: Dispatch exception cleanup pads at run time
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (10 preceding siblings ...)
  2026-10-01 13:31 ` [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
@ 2026-10-01 13:31 ` Yonghong Song
  2026-10-01 13:49   ` sashiko-bot
  2026-10-01 13:31 ` [PATCH bpf-next v8 13/22] bpf, arm64: " Yonghong Song
                   ` (9 subsequent siblings)
  21 siblings, 1 reply; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:31 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

arch_bpf_stack_walk_ra() is the ORC walk with the return-address slot
alongside each frame, which unwind_get_return_address_ptr() already hands
out; writing there is what redirects a frame to its landing pad. The
feature is gated on CONFIG_UNWINDER_ORC for the same reason
arch_bpf_stack_walk() is -- there is no other unwinder here to ask -- and
bpf_jit_supports_cleanup_pads() says yes wherever it is built.

A frame whose return a tracer has hooked is left alone, and the walk stops
there. The unwinder recovers the address such a frame will really return
to, but the slot still holds the function graph or kretprobe trampoline,
so writing it would skip the trampoline and leave its entry for the next
hooked return to pop. x86 has no flag saying a frame was hooked, so the
recovered address and the slot are compared instead. Stopping leaves the
BPF frames returning to paths the verifier never walked, so it warns.

aux->epilogue_ip comes for free: the JIT already emits one epilogue per
(sub)program and every other exit jumps to it, so the offset it keeps as
ctx->cleanup_addr, as a native address, is it. That field has always been
the epilogue and has nothing to do with the cleanup pads despite the name.

The rest is bookkeeping: build the native cleanup table from the JIT's
addrs[] once the image is final. A pad head needs no ENDBR of its own: it
is only ever reached as a return address, and IBT checks indirect jumps
and calls, not returns.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 arch/x86/net/bpf_jit_comp.c | 40 +++++++++++++++++++++++++++++++++++++
 1 file changed, 40 insertions(+)

diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index 6c7a0578760e..544e8fd759ad 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -3278,6 +3278,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
 			seen_exit = true;
 			/* Update cleanup_addr */
 			ctx->cleanup_addr = proglen;
+			/* Where an unwind sends a frame with no pad. */
+			bpf_prog->aux->epilogue_ip = (u64)image + proglen;
 			if (bpf_prog_was_classic(bpf_prog) &&
 			    !ns_capable_noaudit(&init_user_ns, CAP_SYS_ADMIN)) {
 				if (emit_spectre_bhb_barrier(&prog, ip, bpf_prog))
@@ -4455,6 +4457,13 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
 		 */
 		bpf_prog_update_insn_ptrs(prog, addrs, image);
 
+		/*
+		 * Same mapping, consumed by the bpf_unwind() walk:
+		 * turn the cleanup records into native address ranges now
+		 * that the image is final.
+		 */
+		bpf_exc_fill_native_ranges(prog, addrs, image);
+
 		/*
 		 * ctx.prog_offset is used when CFI preambles put code *before*
 		 * the function. See emit_cfi(). For FineIBT specifically this code
@@ -4593,6 +4602,11 @@ bool bpf_jit_supports_exceptions(void)
 	return IS_ENABLED(CONFIG_UNWINDER_ORC);
 }
 
+bool bpf_jit_supports_cleanup_pads(void)
+{
+	return IS_ENABLED(CONFIG_UNWINDER_ORC);
+}
+
 bool bpf_jit_supports_private_stack(void)
 {
 	return true;
@@ -4614,6 +4628,32 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
 #endif
 }
 
+void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
+			    void *cookie)
+{
+#if defined(CONFIG_UNWINDER_ORC)
+	struct unwind_state state;
+	unsigned long addr, *ra;
+
+	for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
+	     unwind_next_frame(&state)) {
+		addr = unwind_get_return_address(&state);
+		ra = unwind_get_return_address_ptr(&state);
+		if (!addr || !ra)
+			break;
+		/*
+		 * A traced return: the slot holds a function graph or kretprobe
+		 * trampoline, not @addr, so it cannot be rewritten. Stopping
+		 * leaves BPF frames returning to unverified paths, so warn.
+		 */
+		if (WARN_ON_ONCE(READ_ONCE_NOCHECK(*ra) != addr))
+			break;
+		if (!consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
+			break;
+	}
+#endif
+}
+
 void bpf_arch_poke_desc_update(struct bpf_jit_poke_descriptor *poke,
 			       struct bpf_prog *new, struct bpf_prog *old)
 {
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 13/22] bpf, arm64: Dispatch exception cleanup pads at run time
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (11 preceding siblings ...)
  2026-10-01 13:31 ` [PATCH bpf-next v8 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
@ 2026-10-01 13:31 ` Yonghong Song
  2026-10-01 13:31 ` [PATCH bpf-next v8 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
                   ` (8 subsequent siblings)
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:31 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

The JIT half: build the native cleanup table from the JIT's byte offsets
once the image is final, and record the one epilogue so a frame the unwind
passes over can return through it. A pad head needs no BTI of its own: it
is only ever reached as a return address, and a return sets no BTYPE, so
no branch-target check is made.

The dispatch is arch_bpf_stack_walk_ra(). arm64's unwinder reads a frame's
return address from the frame record its callee pushed, so the walk hands
the unwind a copy and, when the unwind changes it, stores the new address
back into that record, the one the previous entry stepped through.

Writing it has to respect pointer authentication: a BPF prologue signs the
link register with PACIASP and the epilogue authenticates it, so what goes
back has to carry the same signature. Its modifier is the stack pointer the
owner was entered with, the record + 16 for a BPF prologue, and only a BPF
frame's record is written: bpf_unwind() leaves its own caller's return
address alone. Re-signing the address the unwinder stripped checks that
modifier against the slot before anything is signed with it.

Whether a slot is signed is asked of the build rather than read off the
value: CONFIG_ARM64_PTR_AUTH_KERNEL is what the prologue signs under.
Reading it off the value instead would take a signed address for an
unsigned one whenever its PAC equalled the bits stripping puts back. The
CPU has to implement address authentication too, since "pacia Xd, Xn" is
not in the HINT space and would be undefined without it.

Two frames are not redirected: the first, whose return into bpf_unwind()
comes out of the walk's own frame record, and one the function graph
tracer or a kretprobe has hooked, whose slot holds the trampoline rather
than the address the unwinder reports -- the walk stops there, with a
warning, since the BPF frames are then left returning to paths the
verifier never walked. The frame that called bpf_unwind() is left alone by
the generic code.

bpf_jit_supports_cleanup_pads() can now say yes. A shadow call stack does
not change that: only JITed frames' records are written, JITed code keeps
no x18 copy of its return address, and bpf_unwind()'s own return is never
rewritten.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 arch/arm64/kernel/stacktrace.c | 104 +++++++++++++++++++++++++++++++++
 arch/arm64/net/bpf_jit_comp.c  |  21 +++++++
 2 files changed, 125 insertions(+)

diff --git a/arch/arm64/kernel/stacktrace.c b/arch/arm64/kernel/stacktrace.c
index 3ebcf8c53fb0..66e3e2eefff4 100644
--- a/arch/arm64/kernel/stacktrace.c
+++ b/arch/arm64/kernel/stacktrace.c
@@ -445,6 +445,110 @@ noinline noinstr void arch_bpf_stack_walk(bool (*consume_entry)(void *cookie, u6
 	kunwind_stack_walk(arch_bpf_unwind_consume_entry, &data, current, NULL);
 }
 
+struct bpf_unwind_ra_consume_entry_data {
+	bool (*consume_entry)(void *cookie, u64 ip, u64 sp, u64 fp, u64 *ra);
+	void *cookie;
+	unsigned long record;
+	bool seen_first;
+};
+
+static u64 bpf_unwind_sign_ra(u64 ra, u64 modifier)
+{
+	asm volatile(ARM64_ASM_PREAMBLE
+		     ".arch_extension pauth\n"
+		     "	pacia %0, %1"
+		     : "+r" (ra) : "r" (modifier));
+	return ra;
+}
+
+/*
+ * PACIASP's modifier is the stack pointer the owner was entered with, the
+ * record + 16 for a BPF prologue. Only a BPF frame's record is rewritten --
+ * bpf_unwind() leaves its own caller's return address alone -- so that is
+ * the modifier; check it by re-signing @pc, which the unwinder stripped
+ * from @stored, before signing anything with it.
+ */
+static bool bpf_unwind_ra_modifier(unsigned long record, u64 stored, u64 pc,
+				   u64 *modifier)
+{
+	*modifier = record + sizeof(struct frame_record);
+	return bpf_unwind_sign_ra(pc, *modifier) == stored;
+}
+
+static bool bpf_unwind_store_ra(unsigned long record, u64 pc, u64 ra)
+{
+	struct frame_record *rec = (struct frame_record *)record;
+
+	/*
+	 * Whether the slot holds a signed address is a property of the build,
+	 * not one to be read off the value: a PAC can come out equal to the
+	 * bits stripping puts back, and a signed address would then be taken
+	 * for an unsigned one. What signs is CONFIG_ARM64_PTR_AUTH_KERNEL --
+	 * the prologue here, and -mbranch-protection for everything the
+	 * compiler emits.
+	 */
+	if (IS_ENABLED(CONFIG_ARM64_PTR_AUTH_KERNEL) &&
+	    system_supports_address_auth()) {
+		u64 stored = READ_ONCE(rec->lr);
+		u64 modifier;
+
+		if (WARN_ON_ONCE(!bpf_unwind_ra_modifier(record, stored, pc,
+							 &modifier)))
+			return false;
+		ra = bpf_unwind_sign_ra(ra, modifier);
+	}
+	WRITE_ONCE(rec->lr, ra);
+	return true;
+}
+
+static bool
+arch_bpf_unwind_ra_consume_entry(const struct kunwind_state *state, void *cookie)
+{
+	struct bpf_unwind_ra_consume_entry_data *data = cookie;
+	unsigned long record = data->record;
+	bool seen_first = data->seen_first;
+	u64 ra = state->common.pc;
+	bool cont;
+
+	/* The record this frame's return address will have come out of. */
+	data->record = state->common.fp;
+	data->seen_first = true;
+
+	/*
+	 * The first pc returns into bpf_unwind(), from this walk's own frame
+	 * record: not a BPF frame, and not one to redirect.
+	 */
+	if (!seen_first)
+		return true;
+	/*
+	 * A traced return: the slot holds the tracer's trampoline, not @pc.
+	 * Stopping leaves the BPF frames returning to paths the verifier
+	 * never walked, so warn.
+	 */
+	if (WARN_ON_ONCE(state->flags.fgraph || state->flags.kretprobe))
+		return false;
+
+	/* A consumer that stops still gets to redirect the frame it stopped on. */
+	cont = data->consume_entry(data->cookie, state->common.pc, 0,
+				   state->common.fp, &ra);
+	if (ra != state->common.pc &&
+	    !bpf_unwind_store_ra(record, state->common.pc, ra))
+		return false;
+	return cont;
+}
+
+noinline noinstr void arch_bpf_stack_walk_ra(bool (*consume_entry)(void *cookie, u64 ip, u64 sp,
+								   u64 fp, u64 *ra),
+					     void *cookie)
+{
+	struct bpf_unwind_ra_consume_entry_data data = {
+		.consume_entry = consume_entry,
+		.cookie = cookie,
+	};
+
+	kunwind_stack_walk(arch_bpf_unwind_ra_consume_entry, &data, current, NULL);
+}
+
 static const char *state_source_string(const struct kunwind_state *state)
 {
 	switch (state->source) {
diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
index 475e70653454..8ac98b024194 100644
--- a/arch/arm64/net/bpf_jit_comp.c
+++ b/arch/arm64/net/bpf_jit_comp.c
@@ -2423,6 +2423,17 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
 		 * reasons, expects to point to the next instruction)
 		 */
 		bpf_prog_update_insn_ptrs(prog, ctx.offset, ctx.ro_image);
+
+		/*
+		 * Same byte offsets, consumed by the bpf_unwind() walk:
+		 * turn the cleanup records into native address ranges now that
+		 * the image is final.
+		 */
+		bpf_exc_fill_native_ranges(prog, ctx.offset, ctx.ro_image);
+
+		/* Where an unwind sends a frame with no pad. */
+		prog->aux->epilogue_ip = (u64)ctx.ro_image +
+					 ctx.epilogue_offset * AARCH64_INSN_SIZE;
 out_off:
 		if (!ro_header && priv_stack_ptr) {
 			free_percpu(priv_stack_ptr);
@@ -3408,6 +3419,16 @@ bool bpf_jit_supports_exceptions(void)
 	return true;
 }
 
+bool bpf_jit_supports_cleanup_pads(void)
+{
+	/*
+	 * An unwind rewrites the return addresses in JITed frames' records,
+	 * which is what JITed code returns through, shadow call stack or not:
+	 * it keeps no x18 copy. bpf_unwind()'s own return is never rewritten.
+	 */
+	return true;
+}
+
 bool bpf_jit_supports_arena(void)
 {
 	return true;
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (12 preceding siblings ...)
  2026-10-01 13:31 ` [PATCH bpf-next v8 13/22] bpf, arm64: " Yonghong Song
@ 2026-10-01 13:31 ` Yonghong Song
  2026-10-01 13:31 ` [PATCH bpf-next v8 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
                   ` (7 subsequent siblings)
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:31 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

LLVM terminates a cleanup landing pad with a call to _Unwind_Resume: the
base unwind ABI's entry point for carrying an unwind on once a frame's
cleanups have run. The kernel provides that terminator as a kfunc, but not
under that name: claiming _Unwind_Resume in the kernel's own symbol table,
for a function whose body never runs, would be needlessly confusing, so it
is called bpf_unwind_resume.

The compiler emits the call but not the declaration: _Unwind_Resume lands
in the object as a plain undefined symbol, with nothing in .ksyms. libbpf
takes every undefined NOTYPE symbol for an extern and refuses one it has no
BTF for ("failed to find BTF for extern '_Unwind_Resume'"), so the program
declares it itself -- extern void _Unwind_Resume(void *) __ksym; -- as the
selftests here do, and as a language runtime emitting cleanup pads has to.
What follows translates that name; it does not manufacture the declaration.

Both load paths take the detour. A direct load resolves the name against
the kernel's BTF while libbpf runs. A light skeleton instead writes the
name into the loader program's blob of bytes, for that program to resolve
when it runs, so the name recorded there has to be translated as well.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/lib/bpf/libbpf.c | 45 ++++++++++++++++++++++++++++++++----------
 1 file changed, 35 insertions(+), 10 deletions(-)

diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index fdd69aac39bd..80fd0dcd0248 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -8802,6 +8802,23 @@ static void fixup_verifier_log(struct bpf_program *prog, char *buf, size_t buf_s
 	}
 }
 
+/*
+ * LLVM terminates a cleanup landing pad with a call to _Unwind_Resume, the
+ * base unwind ABI's entry point for carrying an unwind on once a frame's
+ * cleanups have run. The kernel knows it as bpf_unwind_resume. Any other
+ * extern, and a variable of that name, is looked up by @name unchanged.
+ */
+static const char *kern_extern_name(const struct bpf_object *obj,
+				    const struct extern_desc *ext, const char *name)
+{
+	const char *essent = ext->essent_name ?: ext->name;
+
+	if (!btf_is_func(btf__type_by_id(obj->btf, ext->btf_id)) ||
+	    strcmp(essent, "_Unwind_Resume"))
+		return name;
+	return "bpf_unwind_resume";
+}
+
 static int bpf_program_record_relos(struct bpf_program *prog)
 {
 	struct bpf_object *obj = prog->obj;
@@ -8810,6 +8827,7 @@ static int bpf_program_record_relos(struct bpf_program *prog)
 	for (i = 0; i < prog->nr_reloc; i++) {
 		struct reloc_desc *relo = &prog->reloc_desc[i];
 		struct extern_desc *ext = &obj->externs[relo->ext_idx];
+		const char *name;
 		int kind;
 
 		switch (relo->type) {
@@ -8818,14 +8836,14 @@ static int bpf_program_record_relos(struct bpf_program *prog)
 				continue;
 			kind = btf_is_var(btf__type_by_id(obj->btf, ext->btf_id)) ?
 				BTF_KIND_VAR : BTF_KIND_FUNC;
-			bpf_gen__record_extern(obj->gen_loader, ext->name,
-					       ext->is_weak, !ext->ksym.type_id,
-					       true, kind, relo->insn_idx);
+			name = kern_extern_name(obj, ext, ext->name);
+			bpf_gen__record_extern(obj->gen_loader, name, ext->is_weak,
+					       !ext->ksym.type_id, true, kind, relo->insn_idx);
 			break;
 		case RELO_EXTERN_CALL:
-			bpf_gen__record_extern(obj->gen_loader, ext->name,
-					       ext->is_weak, false, false, BTF_KIND_FUNC,
-					       relo->insn_idx);
+			name = kern_extern_name(obj, ext, ext->name);
+			bpf_gen__record_extern(obj->gen_loader, name, ext->is_weak, false,
+					       false, BTF_KIND_FUNC, relo->insn_idx);
 			break;
 		case RELO_CORE: {
 			struct bpf_core_relo cr = {
@@ -9299,17 +9317,24 @@ static int bpf_object__resolve_ksym_func_btf_id(struct bpf_object *obj,
 	struct module_btf *mod_btf = NULL;
 	const struct btf_type *kern_func;
 	struct btf *kern_btf = NULL;
+	const char *local_name, *kern_name;
 	int ret;
 
 	local_func_proto_id = ext->ksym.type_id;
 
-	kfunc_id = find_ksym_btf_id(obj, ext->essent_name ?: ext->name, BTF_KIND_FUNC, &kern_btf,
-				    &mod_btf);
+	local_name = ext->essent_name ?: ext->name;
+	kern_name = kern_extern_name(obj, ext, local_name);
+
+	kfunc_id = find_ksym_btf_id(obj, kern_name, BTF_KIND_FUNC, &kern_btf, &mod_btf);
 	if (kfunc_id < 0) {
 		if (kfunc_id == -ESRCH && ext->is_weak)
 			return 0;
-		pr_warn("extern (func ksym) '%s': not found in kernel or module BTFs\n",
-			ext->name);
+		if (kern_name != local_name)
+			pr_warn("extern (func ksym) '%s' ('%s' in the kernel): not found in kernel or module BTFs\n",
+				ext->name, kern_name);
+		else
+			pr_warn("extern (func ksym) '%s': not found in kernel or module BTFs\n",
+				ext->name);
 		return kfunc_id;
 	}
 
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (13 preceding siblings ...)
  2026-10-01 13:31 ` [PATCH bpf-next v8 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
@ 2026-10-01 13:31 ` Yonghong Song
  2026-10-01 13:46   ` sashiko-bot
  2026-10-01 13:31 ` [PATCH bpf-next v8 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
                   ` (6 subsequent siblings)
  21 siblings, 1 reply; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:31 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

Let a caller hand the kernel an exception cleanup table: struct
bpf_prog_load_opts grows cleanup_info_cnt, cleanup_info and
cleanup_info_rec_size, and bpf_prog_load() passes all three on to
BPF_PROG_LOAD.

The attr size it computes now ends at cleanup_info_cnt, the last of the
three new fields in the BPF_PROG_LOAD attr, so it covers all of them. In
the opts, cleanup_info_cnt comes first so that it fills the four bytes
after fd_array_cnt the pointer would otherwise leave as a hole.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/lib/bpf/bpf.c | 6 +++++-
 tools/lib/bpf/bpf.h | 7 ++++++-
 2 files changed, 11 insertions(+), 2 deletions(-)

diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
index b49822d212ae..b4031f62bdee 100644
--- a/tools/lib/bpf/bpf.c
+++ b/tools/lib/bpf/bpf.c
@@ -295,7 +295,7 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
 		  const struct bpf_insn *insns, size_t insn_cnt,
 		  struct bpf_prog_load_opts *opts)
 {
-	const size_t attr_sz = offsetofend(union bpf_attr, keyring_id);
+	const size_t attr_sz = offsetofend(union bpf_attr, cleanup_info_cnt);
 	void *finfo = NULL, *linfo = NULL;
 	const char *func_info, *line_info;
 	__u32 log_size, log_level, attach_prog_fd, attach_btf_obj_fd;
@@ -370,6 +370,10 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
 	attr.fd_array = ptr_to_u64(OPTS_GET(opts, fd_array, NULL));
 	attr.fd_array_cnt = OPTS_GET(opts, fd_array_cnt, 0);
 
+	attr.cleanup_info = ptr_to_u64(OPTS_GET(opts, cleanup_info, NULL));
+	attr.cleanup_info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
+	attr.cleanup_info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
+
 	if (log_level) {
 		attr.log_buf = ptr_to_u64(log_buf);
 		attr.log_size = log_size;
diff --git a/tools/lib/bpf/bpf.h b/tools/lib/bpf/bpf.h
index 826d9cc9ab65..cbe56ddc8cf7 100644
--- a/tools/lib/bpf/bpf.h
+++ b/tools/lib/bpf/bpf.h
@@ -128,9 +128,14 @@ struct bpf_prog_load_opts {
 
 	/* if set, provides the length of fd_array */
 	__u32 fd_array_cnt;
+
+	/* exception cleanup table, from the .bpf_cleanup section */
+	__u32 cleanup_info_cnt;
+	const void *cleanup_info;
+	__u32 cleanup_info_rec_size;
 	size_t :0;
 };
-#define bpf_prog_load_opts__last_field fd_array_cnt
+#define bpf_prog_load_opts__last_field cleanup_info_rec_size
 
 LIBBPF_API int bpf_prog_load(enum bpf_prog_type prog_type,
 			     const char *prog_name, const char *license,
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (14 preceding siblings ...)
  2026-10-01 13:31 ` [PATCH bpf-next v8 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
@ 2026-10-01 13:31 ` Yonghong Song
  2026-10-01 13:31 ` [PATCH bpf-next v8 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
                   ` (5 subsequent siblings)
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:31 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

Parse the compiler-emitted .bpf_cleanup section and hand the resulting
table to BPF_PROG_LOAD.

Each record is three 4-byte fields, and each field is a byte offset into
some code section named by a matching .rel.bpf_cleanup relocation.
bpf_object_init_cleanup_info() resolves each field through its relocation
once at open time and keeps (section, instruction index) pairs; it rejects
a record whose field has no relocation, whose relocation is not one of the
two 32-bit types that spell a data reference to a code section --
R_BPF_64_NODYLD32 from LLVM, R_BPF_64_ABS32 from GNU as -- or whose offset
is not instruction aligned. LLVM leaves the section's sh_entsize unset,
which is taken to mean the size assumed here, so a section declaring any
other is refused.

The records are sorted by begin_off once the offsets are final, since the
kernel wants the table sorted with disjoint ranges to find the record
covering a call site with a binary search, and they arrive in .bpf_cleanup
order, which says nothing about where the subprograms they describe were
appended. Overlapping ranges are reported here, where the program name and
both regions are still at hand.

bpf_object_load_prog() then passes the per-program table through the
bpf_prog_load() options added in the previous patch, with the record size
carried on the program the way func_info and line_info carry theirs.
bpf_program__clone() carries it too -- that is the load path veristat uses,
and without the table the kernel sees landing pads nothing reaches and
refuses the program with "unreachable insn".

A record field relocated twice is refused rather than silently taking the
second value, and a kernel that refuses the load with E2BIG gets a hint
that it may not know exception cleanup tables.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/lib/bpf/libbpf.c          | 319 +++++++++++++++++++++++++++++++-
 tools/lib/bpf/libbpf_internal.h |   3 +
 2 files changed, 320 insertions(+), 2 deletions(-)

diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index 80fd0dcd0248..3f3b2e84fc57 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -516,6 +516,11 @@ struct bpf_program {
 	void *line_info;
 	__u32 line_info_rec_size;
 	__u32 line_info_cnt;
+
+	struct bpf_cleanup_info *cleanup_info;
+	__u32 cleanup_info_rec_size;
+	__u32 cleanup_info_cnt;
+
 	__u32 prog_flags;
 	__u8  hash[SHA256_DIGEST_LENGTH];
 
@@ -552,6 +557,7 @@ struct bpf_struct_ops {
 #define STRUCT_OPS_SEC ".struct_ops"
 #define STRUCT_OPS_LINK_SEC ".struct_ops.link"
 #define ARENA_SEC ".addr_space.1"
+#define CLEANUP_SEC ".bpf_cleanup"
 
 enum libbpf_map_type {
 	LIBBPF_MAP_UNSPEC,
@@ -683,6 +689,25 @@ struct elf_sec_desc {
 	Elf_Data *data;
 };
 
+#define CLEANUP_REC_FIELDS	(sizeof(struct bpf_cleanup_info) / sizeof(__u32))
+
+/* Index of each field of struct bpf_cleanup_info, read as an array of __u32. */
+enum {
+	CLEANUP_REC_BEGIN,
+	CLEANUP_REC_END,
+	CLEANUP_REC_PAD,
+};
+
+/*
+ * One (begin, end, landing_pad) triple from .bpf_cleanup, each field resolved
+ * from its relocation to a section and an instruction index within it. Final
+ * indices wait for subprogram placement, which differs per main program.
+ */
+struct cleanup_raw_rec {
+	int sec_idx[CLEANUP_REC_FIELDS];
+	size_t insn_idx[CLEANUP_REC_FIELDS];
+};
+
 struct elf_state {
 	int fd;
 	const void *obj_buf;
@@ -702,6 +727,8 @@ struct elf_state {
 	bool has_st_ops;
 	int arena_data_shndx;
 	int jumptables_data_shndx;
+	Elf_Data *cleanup_data;
+	int cleanup_shndx;
 };
 
 struct usdt_manager;
@@ -779,6 +806,9 @@ struct bpf_object {
 	void *jumptables_data;
 	size_t jumptables_data_sz;
 
+	struct cleanup_raw_rec *cleanup_recs;
+	size_t cleanup_rec_cnt;
+
 	struct {
 		struct bpf_program *prog;
 		unsigned int sym_off;
@@ -836,6 +866,9 @@ void bpf_program__unload(struct bpf_program *prog)
 	zfree(&prog->func_info);
 	zfree(&prog->line_info);
 	zfree(&prog->subprogs);
+	zfree(&prog->cleanup_info);
+	prog->cleanup_info_rec_size = 0;
+	prog->cleanup_info_cnt = 0;
 }
 
 static void bpf_program__exit(struct bpf_program *prog)
@@ -1600,6 +1633,7 @@ static struct bpf_object *bpf_object__new(const char *path,
 	obj->efile.obj_buf = obj_buf;
 	obj->efile.obj_buf_sz = obj_buf_sz;
 	obj->efile.btf_maps_shndx = -1;
+	obj->efile.cleanup_shndx = -1;
 	obj->kconfig_map_idx = -1;
 	obj->arena_map_idx = -1;
 
@@ -1619,6 +1653,7 @@ static void bpf_object__elf_finish(struct bpf_object *obj)
 	obj->efile.ehdr = NULL;
 	obj->efile.symbols = NULL;
 	obj->efile.arena_data = NULL;
+	obj->efile.cleanup_data = NULL;
 
 	zfree(&obj->efile.secs);
 	obj->efile.sec_cnt = 0;
@@ -4100,6 +4135,9 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
 				sec_desc->shdr = sh;
 				sec_desc->data = data;
 				obj->efile.has_st_ops = true;
+			} else if (strcmp(name, CLEANUP_SEC) == 0) {
+				obj->efile.cleanup_data = data;
+				obj->efile.cleanup_shndx = idx;
 			} else if (strcmp(name, ARENA_SEC) == 0) {
 				obj->efile.arena_data = data;
 				obj->efile.arena_data_shndx = idx;
@@ -4123,8 +4161,9 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
 
 			/*
 			 * Only do relo for section with exec instructions,
-			 * struct_ops, maps, and read-only data that might
-			 * have pointers to functions.
+			 * struct_ops, maps, read-only data that might have
+			 * pointers to functions, and the exception cleanup
+			 * table.
 			 */
 			if (!section_have_execinstr(obj, targ_sec_idx) &&
 			    strcmp(name, ".rel" RODATA_SEC) &&
@@ -4135,6 +4174,7 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
 			    strcmp(name, ".rel" STRUCT_OPS_LINK_SEC) &&
 			    strcmp(name, ".rel?" STRUCT_OPS_SEC) &&
 			    strcmp(name, ".rel?" STRUCT_OPS_LINK_SEC) &&
+			    strcmp(name, ".rel" CLEANUP_SEC) &&
 			    strcmp(name, ".rel" MAPS_ELF_SEC)) {
 				pr_info("elf: skipping relo section(%d) %s for section(%d) %s\n",
 					idx, name, targ_sec_idx,
@@ -4915,6 +4955,246 @@ static struct bpf_program *find_prog_by_sec_insn(const struct bpf_object *obj,
 	return NULL;
 }
 
+static int bpf_object_init_cleanup_info(struct bpf_object *obj)
+{
+	Elf_Data *data = obj->efile.cleanup_data;
+	size_t i, nrels, nslots, nrecs;
+	struct cleanup_raw_rec *recs;
+	Elf_Data *relo = NULL;
+	const __u32 *vals;
+	Elf64_Shdr *sh;
+	int ret = 0;
+	bool native;
+
+	if (!data || obj->efile.cleanup_shndx < 0 || !data->d_size)
+		return 0;
+
+	native = is_native_endianness(obj);
+
+	for (i = 0; i < obj->efile.sec_cnt; i++) {
+		struct elf_sec_desc *sd = &obj->efile.secs[i];
+
+		if (sd->sec_type == SEC_RELO && sd->shdr &&
+		    sd->shdr->sh_info == (Elf64_Word)obj->efile.cleanup_shndx) {
+			relo = sd->data;
+			break;
+		}
+	}
+	if (!relo) {
+		pr_warn("%s present without relocations\n", CLEANUP_SEC);
+		return -LIBBPF_ERRNO__FORMAT;
+	}
+	/*
+	 * LLVM leaves sh_entsize unset, which is read as the size assumed
+	 * here, so this bites only a producer that declares a record size --
+	 * which is the one able to say it means something other than three
+	 * 4-byte fields.
+	 */
+	sh = elf_sec_hdr(obj, elf_sec_by_idx(obj, obj->efile.cleanup_shndx));
+	if (sh && sh->sh_entsize && sh->sh_entsize != sizeof(struct bpf_cleanup_info)) {
+		pr_warn("%s record size %llu is not the expected %zu\n", CLEANUP_SEC,
+			(unsigned long long)sh->sh_entsize, sizeof(struct bpf_cleanup_info));
+		return -LIBBPF_ERRNO__FORMAT;
+	}
+	if (data->d_size % sizeof(struct bpf_cleanup_info)) {
+		pr_warn("%s size %zu is not a multiple of the record size %zu\n",
+			CLEANUP_SEC, data->d_size, sizeof(struct bpf_cleanup_info));
+		return -LIBBPF_ERRNO__FORMAT;
+	}
+
+	vals = data->d_buf;
+	nslots = data->d_size / sizeof(__u32);
+	nrecs = data->d_size / sizeof(struct bpf_cleanup_info);
+
+	recs = calloc(nrecs, sizeof(*recs));
+	if (!recs)
+		return -ENOMEM;
+	for (i = 0; i < nslots; i++)
+		recs[i / CLEANUP_REC_FIELDS].sec_idx[i % CLEANUP_REC_FIELDS] = -1;
+
+	/* One relocation per 4-byte field, naming the section it points into. */
+	nrels = relo->d_size / sizeof(Elf64_Rel);
+	for (i = 0; i < nrels; i++) {
+		Elf64_Rel *rel = elf_rel_by_idx(relo, i);
+		Elf64_Sym *sym = elf_sym_by_idx(obj, ELF64_R_SYM(rel->r_info));
+		size_t type = ELF64_R_TYPE(rel->r_info);
+		size_t slot = rel->r_offset / sizeof(__u32);
+		struct cleanup_raw_rec *rec;
+		size_t off;
+
+		if (type != R_BPF_64_NODYLD32 && type != R_BPF_64_ABS32) {
+			pr_warn("%s: relocation %zu has unexpected type %zu\n",
+				CLEANUP_SEC, i, type);
+			ret = -LIBBPF_ERRNO__FORMAT;
+			goto out;
+		}
+		if (!sym || slot >= nslots || rel->r_offset % sizeof(__u32)) {
+			pr_warn("%s: bad relocation %zu\n", CLEANUP_SEC, i);
+			ret = -LIBBPF_ERRNO__FORMAT;
+			goto out;
+		}
+		/*
+		 * The addend lives in the section data, which libelf leaves in
+		 * the object's byte order; a non-section symbol additionally
+		 * contributes its own value.
+		 */
+		off = (native ? vals[slot] : bswap_32(vals[slot])) + sym->st_value;
+		if (off % BPF_INSN_SZ) {
+			pr_warn("%s: field %zu offset %zu is not instruction aligned\n",
+				CLEANUP_SEC, slot, off);
+			ret = -LIBBPF_ERRNO__FORMAT;
+			goto out;
+		}
+		rec = &recs[slot / CLEANUP_REC_FIELDS];
+		if (rec->sec_idx[slot % CLEANUP_REC_FIELDS] >= 0) {
+			pr_warn("%s: field %zu is relocated twice\n", CLEANUP_SEC, slot);
+			ret = -LIBBPF_ERRNO__FORMAT;
+			goto out;
+		}
+		rec->sec_idx[slot % CLEANUP_REC_FIELDS] = sym->st_shndx;
+		rec->insn_idx[slot % CLEANUP_REC_FIELDS] = off / BPF_INSN_SZ;
+	}
+
+	for (i = 0; i < nslots; i++) {
+		if (recs[i / CLEANUP_REC_FIELDS].sec_idx[i % CLEANUP_REC_FIELDS] < 0) {
+			pr_warn("%s: field %zu has no relocation\n", CLEANUP_SEC, i);
+			ret = -LIBBPF_ERRNO__FORMAT;
+			goto out;
+		}
+	}
+
+	obj->cleanup_recs = recs;
+	obj->cleanup_rec_cnt = nrecs;
+	return 0;
+out:
+	free(recs);
+	return ret;
+}
+
+static int cmp_cleanup_info(const void *a, const void *b)
+{
+	const struct bpf_cleanup_info *x = a, *y = b;
+
+	if (x->begin_off == y->begin_off)
+		return 0;
+	return x->begin_off < y->begin_off ? -1 : 1;
+}
+
+static int bpf_prog_collect_cleanup_info(struct bpf_object *obj,
+					 struct bpf_program *prog)
+{
+	size_t i;
+	int j;
+
+	for (i = 0; i < obj->cleanup_rec_cnt; i++) {
+		struct cleanup_raw_rec *raw = &obj->cleanup_recs[i];
+		struct bpf_program *owner = NULL;
+		struct bpf_cleanup_info ci = {};
+		__u32 fields[CLEANUP_REC_FIELDS];
+		void *tmp;
+
+		if (raw->sec_idx[CLEANUP_REC_BEGIN] == raw->sec_idx[CLEANUP_REC_END] &&
+		    raw->insn_idx[CLEANUP_REC_BEGIN] >= raw->insn_idx[CLEANUP_REC_END]) {
+			pr_warn("%s: record %zu is an empty range [%zu,%zu)\n",
+				CLEANUP_SEC, i, raw->insn_idx[CLEANUP_REC_BEGIN],
+				raw->insn_idx[CLEANUP_REC_END]);
+			return -LIBBPF_ERRNO__FORMAT;
+		}
+
+		for (j = 0; j < CLEANUP_REC_FIELDS; j++) {
+			size_t idx = raw->insn_idx[j], final;
+			struct bpf_program *p;
+
+			/*
+			 * An exclusive end may name the instruction past the
+			 * last of a function, so ask about the last one the
+			 * range covers, the way the kernel does.
+			 */
+			if (j == CLEANUP_REC_END) {
+				if (!idx) {
+					pr_warn("%s: record %zu ends at instruction 0\n",
+						CLEANUP_SEC, i);
+					return -LIBBPF_ERRNO__FORMAT;
+				}
+				idx--;
+			}
+
+			p = find_prog_by_sec_insn(obj, raw->sec_idx[j], idx);
+			if (!p) {
+				pr_warn("%s: record %zu field %d is not inside a function\n",
+					CLEANUP_SEC, i, j);
+				return -LIBBPF_ERRNO__FORMAT;
+			}
+			if (!owner) {
+				owner = p;
+			} else if (owner != p) {
+				pr_warn("%s: record %zu spans functions '%s' and '%s'\n",
+					CLEANUP_SEC, i, owner->name, p->name);
+				return -LIBBPF_ERRNO__FORMAT;
+			}
+
+			if (owner == prog) {
+				final = raw->insn_idx[j] - prog->sec_insn_off;
+			} else if (prog_is_subprog(obj, owner) && owner->sub_insn_off) {
+				/*
+				 * sub_insn_off is where this subprogram was
+				 * appended to the main program being relocated;
+				 * zero means it is not part of it.
+				 */
+				final = owner->sub_insn_off +
+					raw->insn_idx[j] - owner->sec_insn_off;
+			} else {
+				owner = NULL;
+				break;
+			}
+			fields[j] = final;
+		}
+		if (!owner)
+			continue;
+
+		ci.begin_off = fields[CLEANUP_REC_BEGIN];
+		ci.end_off = fields[CLEANUP_REC_END];
+		ci.landing_pad_off = fields[CLEANUP_REC_PAD];
+
+		tmp = libbpf_reallocarray(prog->cleanup_info, prog->cleanup_info_cnt + 1,
+					  sizeof(*prog->cleanup_info));
+		if (!tmp)
+			return -ENOMEM;
+		prog->cleanup_info = tmp;
+		prog->cleanup_info_rec_size = sizeof(struct bpf_cleanup_info);
+		prog->cleanup_info[prog->cleanup_info_cnt++] = ci;
+
+		pr_debug("prog '%s': cleanup region [%u,%u) -> landing pad %u\n",
+			 prog->name, ci.begin_off, ci.end_off, ci.landing_pad_off);
+	}
+
+	if (!prog->cleanup_info_cnt)
+		return 0;
+
+	/* The light skeleton's blob of the records is sized as an int. */
+	if (prog->cleanup_info_cnt > INT32_MAX / sizeof(struct bpf_cleanup_info)) {
+		pr_warn("prog '%s': too many cleanup records: %u\n",
+			prog->name, prog->cleanup_info_cnt);
+		return -LIBBPF_ERRNO__FORMAT;
+	}
+
+	qsort(prog->cleanup_info, prog->cleanup_info_cnt,
+	      sizeof(*prog->cleanup_info), cmp_cleanup_info);
+	for (i = 1; i < prog->cleanup_info_cnt; i++) {
+		struct bpf_cleanup_info *prev = &prog->cleanup_info[i - 1];
+		struct bpf_cleanup_info *cur = &prog->cleanup_info[i];
+
+		if (cur->begin_off < prev->end_off) {
+			pr_warn("prog '%s': overlapping cleanup regions [%u,%u) and [%u,%u)\n",
+				prog->name, prev->begin_off, prev->end_off,
+				cur->begin_off, cur->end_off);
+			return -LIBBPF_ERRNO__FORMAT;
+		}
+	}
+
+	return 0;
+}
+
 static int
 bpf_object__collect_prog_relos(struct bpf_object *obj, Elf64_Shdr *shdr, Elf_Data *data)
 {
@@ -7899,6 +8179,13 @@ static int bpf_object__relocate(struct bpf_object *obj, const char *targ_btf_pat
 					return err;
 			}
 		}
+
+		err = bpf_prog_collect_cleanup_info(obj, prog);
+		if (err) {
+			pr_warn("prog '%s': failed to collect cleanup info: %s\n",
+				prog->name, errstr(err));
+			return err;
+		}
 	}
 	for (i = 0; i < obj->nr_programs; i++) {
 		prog = &obj->programs[i];
@@ -8181,6 +8468,9 @@ static int bpf_object__collect_relos(struct bpf_object *obj)
 			return -LIBBPF_ERRNO__INTERNAL;
 		}
 
+		if (idx == obj->efile.cleanup_shndx)
+			continue;
+
 		if (obj->efile.secs[idx].sec_type == SEC_RODATA)
 			err = bpf_object__collect_rodata_relos(obj, shdr, data);
 		else if (obj->efile.secs[idx].sec_type == SEC_ST_OPS)
@@ -8474,6 +8764,11 @@ static int bpf_object_load_prog(struct bpf_object *obj, struct bpf_program *prog
 		load_attr.line_info_rec_size = prog->line_info_rec_size;
 		load_attr.line_info_cnt = prog->line_info_cnt;
 	}
+	if (prog->cleanup_info_cnt) {
+		load_attr.cleanup_info = prog->cleanup_info;
+		load_attr.cleanup_info_cnt = prog->cleanup_info_cnt;
+		load_attr.cleanup_info_rec_size = prog->cleanup_info_rec_size;
+	}
 	load_attr.log_level = log_level;
 	load_attr.prog_flags = prog->prog_flags;
 	load_attr.fd_array = obj->fd_array;
@@ -8584,6 +8879,9 @@ static int bpf_object_load_prog(struct bpf_object *obj, struct bpf_program *prog
 
 	pr_warn("prog '%s': BPF program load failed: %s\n", prog->name, errstr(errno));
 	pr_perm_msg(ret);
+	if (ret == -E2BIG && prog->cleanup_info_cnt)
+		pr_warn("prog '%s': the kernel may not support exception cleanup tables\n",
+			prog->name);
 
 	if (own_log_buf && log_buf && log_buf[0] != '\0') {
 		pr_warn("prog '%s': -- BEGIN PROG LOAD LOG --\n%s-- END PROG LOAD LOG --\n",
@@ -9075,6 +9373,7 @@ static struct bpf_object *bpf_object_open(const char *path, const void *obj_buf,
 	err = err ? : bpf_object__init_maps(obj, opts);
 	err = err ? : bpf_object_init_progs(obj, opts);
 	err = err ? : bpf_object__collect_relos(obj);
+	err = err ? : bpf_object_init_cleanup_info(obj);
 	if (err)
 		goto out;
 
@@ -10212,6 +10511,9 @@ void bpf_object__close(struct bpf_object *obj)
 	zfree(&obj->jumptables_data);
 	obj->jumptables_data_sz = 0;
 
+	zfree(&obj->cleanup_recs);
+	obj->cleanup_rec_cnt = 0;
+
 	for (i = 0; i < obj->jumptable_map_cnt; i++)
 		close(obj->jumptable_maps[i].fd);
 	zfree(&obj->jumptable_maps);
@@ -10616,6 +10918,19 @@ int bpf_program__clone(struct bpf_program *prog, const struct bpf_prog_load_opts
 		attr.line_info_rec_size = info ? info_rec_size : prog->line_info_rec_size;
 	}
 
+	/* exception cleanup table */
+	info = OPTS_GET(opts, cleanup_info, NULL);
+	info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
+	info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
+	if (!!info != !!info_cnt || !!info != !!info_rec_size) {
+		pr_warn("prog '%s': cleanup_info, cleanup_info_cnt, and cleanup_info_rec_size must all be specified or all omitted\n",
+			prog->name);
+		return libbpf_err(-EINVAL);
+	}
+	attr.cleanup_info = info ?: prog->cleanup_info;
+	attr.cleanup_info_cnt = info ? info_cnt : prog->cleanup_info_cnt;
+	attr.cleanup_info_rec_size = info ? info_rec_size : prog->cleanup_info_rec_size;
+
 	/* Logging is caller-controlled; no fallback to prog/obj log settings */
 	attr.log_buf = OPTS_GET(opts, log_buf, NULL);
 	attr.log_size = OPTS_GET(opts, log_size, 0);
diff --git a/tools/lib/bpf/libbpf_internal.h b/tools/lib/bpf/libbpf_internal.h
index 546f65b95cf4..f1630f03d5f5 100644
--- a/tools/lib/bpf/libbpf_internal.h
+++ b/tools/lib/bpf/libbpf_internal.h
@@ -56,6 +56,9 @@
 #ifndef R_BPF_64_ABS32
 #define R_BPF_64_ABS32 3
 #endif
+#ifndef R_BPF_64_NODYLD32
+#define R_BPF_64_NODYLD32 4
+#endif
 #ifndef R_BPF_64_32
 #define R_BPF_64_32 10
 #endif
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 17/22] libbpf: Carry the exception cleanup table through the light skeleton
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (15 preceding siblings ...)
  2026-10-01 13:31 ` [PATCH bpf-next v8 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
@ 2026-10-01 13:31 ` Yonghong Song
  2026-10-01 13:31 ` [PATCH bpf-next v8 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
                   ` (4 subsequent siblings)
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:31 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

A light skeleton does not call bpf_prog_load(). bpf_gen__prog_load() builds
its own union bpf_attr field by field, and the loader program it emits is
what issues BPF_PROG_LOAD when the skeleton runs -- so a program loaded
this way reached the kernel without the table the previous patch collected
for it, and the verifier refused it with "unreachable insn", which names
neither the skeleton nor the table.

Carry it the way func_info and line_info are carried: the records go into
the loader's blob of bytes, the count and record size into the attr, and a
relocation stores the blob's address into attr.cleanup_info once that
address is known. For a program with records the attr grows to its new
last field, cleanup_info_cnt. One without keeps the old size: the longer
attr also takes in log_true_size, which the kernel writes back, into the
loader's read-only data.

Records are 4-byte fields like the other info blobs, so a cross-endian
build has to swap them too.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/lib/bpf/gen_loader.c      | 34 ++++++++++++++++++++++++++++++---
 tools/lib/bpf/libbpf_internal.h |  7 +++++++
 2 files changed, 38 insertions(+), 3 deletions(-)

diff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c
index 251392aa8b41..4017fb366384 100644
--- a/tools/lib/bpf/gen_loader.c
+++ b/tools/lib/bpf/gen_loader.c
@@ -994,13 +994,15 @@ static void cleanup_relos(struct bpf_gen *gen, int insns)
 	cleanup_core_relo(gen);
 }
 
-/* Convert func, line, and core relo info blobs to target endianness */
+/* Convert func, line, core relo and cleanup info blobs to target endianness */
 static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
-			    int core_relos, struct bpf_prog_load_opts *load_attr)
+			    int core_relos, int cleanup_info,
+			    struct bpf_prog_load_opts *load_attr)
 {
 	struct bpf_func_info *fi = gen->data_start + func_info;
 	struct bpf_line_info *li = gen->data_start + line_info;
 	struct bpf_core_relo *cr = gen->data_start + core_relos;
+	struct bpf_cleanup_info *ci = gen->data_start + cleanup_info;
 	int i;
 
 	for (i = 0; i < load_attr->func_info_cnt; i++)
@@ -1011,6 +1013,9 @@ static void info_blob_bswap(struct bpf_gen *gen, int func_info, int line_info,
 
 	for (i = 0; i < gen->core_relo_cnt; i++)
 		bpf_core_relo_bswap(cr++);
+
+	for (i = 0; i < load_attr->cleanup_info_cnt; i++)
+		bpf_cleanup_info_bswap(ci++);
 }
 
 void bpf_gen__prog_load(struct bpf_gen *gen,
@@ -1024,10 +1029,20 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
 			       load_attr->line_info_rec_size;
 	int core_relo_tot_sz = gen->core_relo_cnt *
 			       sizeof(struct bpf_core_relo);
+	int cleanup_info_tot_sz = load_attr->cleanup_info_cnt *
+				  load_attr->cleanup_info_rec_size;
 	int prog_load_attr, license_off, insns_off, func_info, line_info, core_relos;
 	int attr_size = offsetofend(union bpf_attr, core_relo_rec_size);
+	int cleanup_info;
 	union bpf_attr attr;
 
+	/*
+	 * Reach the cleanup fields only when there are records: the attr then
+	 * also covers log_true_size, which the kernel writes back, and the
+	 * attr lives in the loader's read-only data.
+	 */
+	if (load_attr->cleanup_info_cnt)
+		attr_size = offsetofend(union bpf_attr, cleanup_info_cnt);
 	memset(&attr, 0, attr_size);
 	/* add license string to blob of bytes */
 	license_off = add_data(gen, license, strlen(license) + 1);
@@ -1074,9 +1089,17 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
 		 core_relos, gen->core_relo_cnt,
 		 sizeof(struct bpf_core_relo));
 
+	attr.cleanup_info_rec_size = tgt_endian(load_attr->cleanup_info_rec_size);
+	attr.cleanup_info_cnt = tgt_endian(load_attr->cleanup_info_cnt);
+	cleanup_info = add_data(gen, load_attr->cleanup_info, cleanup_info_tot_sz);
+	pr_debug("gen: prog_load: cleanup_info: off %d cnt %u rec size %u\n",
+		 cleanup_info, load_attr->cleanup_info_cnt,
+		 load_attr->cleanup_info_rec_size);
+
 	/* convert all info blobs to target endianness */
 	if (gen->swapped_endian && !gen->error)
-		info_blob_bswap(gen, func_info, line_info, core_relos, load_attr);
+		info_blob_bswap(gen, func_info, line_info, core_relos, cleanup_info,
+				load_attr);
 
 	libbpf_strlcpy(attr.prog_name, prog_name, sizeof(attr.prog_name));
 	prog_load_attr = add_data(gen, &attr, attr_size);
@@ -1098,6 +1121,11 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
 	/* populate union bpf_attr with a pointer to core_relos */
 	emit_rel_store(gen, attr_field(prog_load_attr, core_relos), core_relos);
 
+	/* with no records there is no blob of them to point the attr at */
+	if (load_attr->cleanup_info_cnt)
+		emit_rel_store(gen, attr_field(prog_load_attr, cleanup_info),
+			       cleanup_info);
+
 	/* populate union bpf_attr fd_array with a pointer to data where map_fds are saved */
 	emit_rel_store(gen, attr_field(prog_load_attr, fd_array), gen->fd_array);
 
diff --git a/tools/lib/bpf/libbpf_internal.h b/tools/lib/bpf/libbpf_internal.h
index f1630f03d5f5..9d341839ca74 100644
--- a/tools/lib/bpf/libbpf_internal.h
+++ b/tools/lib/bpf/libbpf_internal.h
@@ -572,6 +572,13 @@ static inline void bpf_core_relo_bswap(struct bpf_core_relo *i)
 	i->kind = bswap_32(i->kind);
 }
 
+static inline void bpf_cleanup_info_bswap(struct bpf_cleanup_info *i)
+{
+	i->begin_off = bswap_32(i->begin_off);
+	i->end_off = bswap_32(i->end_off);
+	i->landing_pad_off = bswap_32(i->landing_pad_off);
+}
+
 enum btf_field_iter_kind {
 	BTF_FIELD_ITER_IDS,
 	BTF_FIELD_ITER_STRS,
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (16 preceding siblings ...)
  2026-10-01 13:31 ` [PATCH bpf-next v8 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
@ 2026-10-01 13:31 ` Yonghong Song
  2026-10-01 13:31 ` [PATCH bpf-next v8 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
                   ` (3 subsequent siblings)
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:31 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

An object that carries a compiler-emitted exception cleanup table cannot be
linked today. The table's fields are byte offsets into a code section,
materialised by a 32-bit relocation against that section's symbol with the
offset itself as the implicit addend. The linker does not accept LLVM's
R_BPF_64_NODYLD32 at all, and from a non-executable section it takes a
relocation against an STT_SECTION symbol only as the 64-bit pointer to
code R_BPF_64_ABS64 carries: a 32-bit one is refused as not supported.

Both spellings of that relocation have to be taken. LLVM emits
R_BPF_64_NODYLD32 for a .long against a section symbol; GNU as emits
R_BPF_64_ABS32, which is what binutils' bpf_reloc_type_lookup() maps
BFD_RELOC_32 to. They describe the same value, and the selftests are built
with both compilers.

Keying on the relocation type rather than the section name means any
non-executable section can reach the new arm, where a 32-bit relocation
against a code section used to be refused; one against anything else still
is. That refusal was covering two things the arm now has to do itself: a
target may be SHT_NOBITS, which extend_sec() leaves with no raw_data, and
r_offset is alignment-checked only where the section holds instructions.
The arm rejects both, and bounds the offset against the section size before
writing through it.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/lib/bpf/linker.c | 37 ++++++++++++++++++++++++++++++++++++-
 1 file changed, 36 insertions(+), 1 deletion(-)

diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
index f3f71c452f00..c607f51e3cb4 100644
--- a/tools/lib/bpf/linker.c
+++ b/tools/lib/bpf/linker.c
@@ -1036,7 +1036,8 @@ static int linker_sanity_check_elf_relos(struct src_obj *obj, struct src_sec *se
 		size_t sym_type = ELF64_R_TYPE(relo->r_info);
 
 		if (sym_type != R_BPF_64_64 && sym_type != R_BPF_64_32 &&
-		    sym_type != R_BPF_64_ABS64 && sym_type != R_BPF_64_ABS32) {
+		    sym_type != R_BPF_64_ABS64 && sym_type != R_BPF_64_ABS32 &&
+		    sym_type != R_BPF_64_NODYLD32) {
 			pr_warn("ELF relo #%d in section #%zu has unexpected type %zu in %s\n",
 				i, sec->sec_idx, sym_type, obj->filename);
 			return -EINVAL;
@@ -2263,6 +2264,7 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
 			if (ELF64_ST_TYPE(src_sym->st_info) == STT_SECTION) {
 				struct src_sec *sec = &obj->secs[src_sym->st_shndx];
 				struct bpf_insn *insn;
+				__u32 *val;
 
 				if (src_linked_sec->shdr->sh_flags & SHF_EXECINSTR) {
 					/* calls to the very first static function inside
@@ -2297,6 +2299,39 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
 					if (linker->swapped_endian)
 						off = bswap_64(off);
 					memcpy(ptr, &off, sizeof(off));
+				} else if ((sym_type == R_BPF_64_NODYLD32 ||
+					    sym_type == R_BPF_64_ABS32) &&
+					   (sec->shdr->sh_flags & SHF_EXECINSTR)) {
+					/*
+					 * A byte offset into a code section,
+					 * stored in place. LLVM spells this
+					 * relocation NODYLD32 and GNU as
+					 * spells it ABS32; being bytes, the
+					 * section's new start goes in as it
+					 * is, not scaled the way a call's
+					 * instruction index is above.
+					 *
+					 * r_offset is checked only for an
+					 * executable section, and SHT_NOBITS
+					 * has no raw_data, so bound it here --
+					 * subtracting, so it cannot wrap.
+					 */
+					if (!dst_linked_sec->raw_data ||
+					    dst_linked_sec->sec_sz < (int)sizeof(*val) ||
+					    dst_rel->r_offset % sizeof(*val) ||
+					    dst_rel->r_offset >
+					    (size_t)dst_linked_sec->sec_sz - sizeof(*val)) {
+						pr_warn("ELF relo #%d in section #%zu points outside the data of section '%s' in %s\n",
+							j, src_sec->sec_idx,
+							dst_linked_sec->sec_name,
+							obj->filename);
+						return -EINVAL;
+					}
+					val = dst_linked_sec->raw_data + dst_rel->r_offset;
+					if (linker->swapped_endian)
+						*val = bswap_32(bswap_32(*val) + sec->dst_off);
+					else
+						*val += sec->dst_off;
 				} else {
 					pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
 					return -EINVAL;
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (17 preceding siblings ...)
  2026-10-01 13:31 ` [PATCH bpf-next v8 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
@ 2026-10-01 13:31 ` Yonghong Song
  2026-10-01 13:31 ` [PATCH bpf-next v8 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
                   ` (2 subsequent siblings)
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:31 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

C has no unwinding, so nothing here comes out of the frontend: the frames
that own a resource are written in inline assembly, which spells out by
hand exactly what a frontend emits -- a call site bracketed by two labels,
a landing pad unreachable in the compiler's CFG, and a .bpf_cleanup record
tying them together. Most are __naked; foo3, below, is C around one asm
block. Clang's assembler turns ".long <text label>" into the same
R_BPF_64_NODYLD32 relocation the BPF AsmPrinter emits, and GNU as into an
R_BPF_64_ABS32 that the static linker takes too, so libbpf and the kernel
see an object indistinguishable from a compiler-generated one.

Call chain: entry -> foo1 -> foo1v -> foo2 -> foo3. foo3 holds a
non-preemptible section and unwinds inside it; foo2 holds an RCU read lock
and has two call sites sharing one pad, one of them its own unwind; foo1v
is a void frame whose pad ends in a jump to a resume block placed after an
unrelated block that ends in a plain exit; foo1 owns nothing and gets no
record; entry is the main program, where the unwind stops. foo2's pad
calls drop_glue(), and bump() is a pad-less unwinder that is verified but
never fires at run time.

There are also the shapes the kernel refuses:

 - a catch pad, and a pad ambiguous between catch and cleanup
 - a table alongside bpf_throw() or a tagged exception callback
 - a subprogram that can unwind, used as a callback, including one that
   calls an unwinding subprogram through a pointer read from its caller
 - a tail call, a BPF_LD_[ABS|IND] or a gotox in a pad
 - a kfunc call in a pad whose by-value argument past the argument
   registers was never set, which the ordinary argument check refuses:
   pad code is verified like any other
 - a second unwind while one is in flight, raised in the pad or below it
 - a resume outside a pad, and one in a subprogram the pad called
 - a jump into a pad from outside it
 - a pad inside another record's call-site range, and a record covering no
   call that can unwind and whose pad nothing reaches
 - a frame leaving through an unwind holding what it did not hold when it
   was entered: a pad dropping a lock its frame never took, one forgetting
   the lock it did take, a subprogram with no pad of its own leaving while
   it holds one, and a pad dropping a reference the frame never reserved
 - a pad-less unwind inside an RCU read-side region
 - a frame holding a lock or a reference across a call an unwind passes
   through with no record over it: a lock or a reference in a subprogram's
   frame, and a reference in the main program's
 - a pad trusting a stack slot to hold what it held at the call, after
   the callee wrote it through a pointer and unwound, through a static
   callee and a global callee

The test skips rather than fails where the JIT cannot dispatch a landing
pad at all.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 .../selftests/bpf/exceptions_cleanup.h        |  27 +
 .../bpf/prog_tests/exceptions_cleanup.c       |  85 ++
 .../selftests/bpf/progs/exceptions_cleanup.c  | 160 +++
 .../bpf/progs/exceptions_cleanup_fail.c       | 978 ++++++++++++++++++
 4 files changed, 1250 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/exceptions_cleanup.h
 create mode 100644 tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup.c
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c

diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
new file mode 100644
index 000000000000..96effd2c1361
--- /dev/null
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -0,0 +1,27 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#ifndef __EXCEPTIONS_CLEANUP_H__
+#define __EXCEPTIONS_CLEANUP_H__
+
+/* progs/exceptions_cleanup.c: one bit per function that reports it ran. */
+#define RAN_FOO3_PREEMPT	0x1
+#define RAN_FOO2_RCU		0x2
+#define RAN_FOO1V_PREEMPT	0x4
+#define RAN_FOO2_DROP		0x8
+#define RAN_BUMP		0x10
+
+#define CLEANUP_REC(begin, end, landing_pad)			\
+	".pushsection .bpf_cleanup,\"a\",@progbits;"		\
+	".long " begin ";"					\
+	".long " end ";"					\
+	".long " landing_pad ";"				\
+	".popsection;"
+
+/* Set a bit in @pads_ran. */
+#define PAD_RAN(bit)						\
+	"r1 = %[pads_ran] ll;"					\
+	"r2 = *(u64 *)(r1 + 0);"				\
+	"r2 |= " bit ";"					\
+	"*(u64 *)(r1 + 0) = r2;"
+
+#endif /* __EXCEPTIONS_CLEANUP_H__ */
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
new file mode 100644
index 000000000000..255f88d35aad
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -0,0 +1,85 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <test_progs.h>
+#include "exceptions_cleanup.h"
+#include "exceptions_cleanup.skel.h"
+#include "exceptions_cleanup_fail.skel.h"
+
+/* foo3 unwound: every frame that has a pad ran it. */
+#define PADS_FOO3_UNWOUND \
+	(RAN_FOO3_PREEMPT | RAN_FOO2_RCU | RAN_FOO1V_PREEMPT | RAN_FOO2_DROP)
+
+/* foo2 unwound after foo3 returned normally: foo3's pad must not run. */
+#define PADS_FOO2_UNWOUND \
+	(RAN_FOO2_RCU | RAN_FOO1V_PREEMPT | RAN_FOO2_DROP)
+
+static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
+		__u64 pads)
+{
+	__u64 ctx = 0;
+	int err;
+
+	LIBBPF_OPTS(bpf_test_run_opts, topts,
+		    .ctx_in = &ctx,
+		    .ctx_size_in = sizeof(ctx),
+	);
+
+	skel->bss->input = input;
+	skel->bss->pads_ran = 0;
+	skel->bss->result = 0;
+
+	err = bpf_prog_test_run_opts(bpf_program__fd(skel->progs.entry), &topts);
+	if (!ASSERT_OK(err, "run"))
+		return;
+	ASSERT_EQ(topts.retval, retval, "retval");
+	/* bump() is not a landing pad; it sets its bit on every run. */
+	ASSERT_EQ(skel->bss->pads_ran, pads | RAN_BUMP, "pads_ran");
+}
+
+void test_exceptions_cleanup(void)
+{
+	char log[8192] = {};
+
+	LIBBPF_OPTS(bpf_object_open_opts, opts,
+		    .kernel_log_buf = log,
+		    .kernel_log_size = sizeof(log));
+	struct exceptions_cleanup *skel;
+	int err;
+
+	skel = exceptions_cleanup__open_opts(&opts);
+	if (!ASSERT_OK_PTR(skel, "open"))
+		return;
+
+	err = exceptions_cleanup__load(skel);
+	if (err) {
+		if (err == -EOPNOTSUPP &&
+		    strstr(log, "exception cleanup needs a JIT that can dispatch landing pads")) {
+			printf("%s:SKIP:JIT cannot dispatch exception cleanup landing pads\n",
+			       __func__);
+			test__skip();
+		} else if (!ASSERT_OK(err, "load")) {
+			fprintf(stderr, "%s", log);
+		}
+		exceptions_cleanup__destroy(skel);
+		return;
+	}
+
+	/* No unwind: foo3 returns 1 ^ 1 == 0, foo2 adds one, no pad runs. */
+	if (test__start_subtest("no_unwind"))
+		run(skel, 1, 1, 0);
+
+	/* foo3 unwinds; every pad runs and entry returns zero. */
+	if (test__start_subtest("unwind_from_foo3"))
+		run(skel, 101, 0, PADS_FOO3_UNWOUND);
+
+	/*
+	 * foo3 returns 2 ^ 1 == 3, so foo2 unwinds from its own second region;
+	 * foo3's frame is long gone, so its pad must not run.
+	 */
+	if (test__start_subtest("unwind_from_foo2"))
+		run(skel, 2, 0, PADS_FOO2_UNWOUND);
+
+	exceptions_cleanup__destroy(skel);
+
+	RUN_TESTS(exceptions_cleanup_fail);
+}
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
new file mode 100644
index 000000000000..d065ba53c812
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup.c
@@ -0,0 +1,160 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+	bpf_unwind();
+	bpf_rcu_read_lock();
+	bpf_rcu_read_unlock();
+	bpf_preempt_disable();
+	bpf_preempt_enable();
+	bpf_unwind_resume(NULL);
+}
+
+__u64 input = 0;
+__u64 pads_ran = 0;
+__u64 result = 0;
+__u64 never = 0;
+
+static __used __noinline __u64 foo3(__u64 x)
+{
+	bpf_preempt_disable();
+	if (x > 100)
+		asm volatile (
+	"1:"	"call bpf_unwind;"		/* cleanup region */
+	"2:"
+		"goto 3f;"
+	"4:"					/* landing pad */
+		/*
+		 * r0 at pad entry is the zero the fixups put after the
+		 * bpf_unwind() call. It is kept in a callee-saved
+		 * register and handed to the resume, the way a
+		 * compiler-emitted pad passes the exception pointer to
+		 * _Unwind_Resume. The kfunc takes it and ignores it, and
+		 * the two pads below do without the shuffle.
+		 */
+		"r7 = r0;"
+		"call bpf_preempt_enable;"
+		PAD_RAN("%[ran]")
+		"r1 = r7;"
+		"call bpf_unwind_resume;"
+	"3:"
+		CLEANUP_REC("1b", "2b", "4b")
+		:
+		: [ran]"i"(RAN_FOO3_PREEMPT),
+		  __imm_addr(pads_ran)
+		: __clobber_all);
+	bpf_preempt_enable();
+	return x ^ 1;
+}
+
+static __used __naked __noinline void drop_glue(void)
+{
+	asm volatile (
+	PAD_RAN("%[ran]")
+	"exit;"
+	:
+	: [ran]"i"(RAN_FOO2_DROP), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+static __used __naked __noinline __u64 foo2(void)
+{
+	asm volatile (
+	"r6 = r1;"
+	"call bpf_rcu_read_lock;"
+	"r1 = r6;"
+"1:"	"call foo3;"			/* cleanup region #1 */
+"2:"
+	"r6 = r0;"
+	"if r6 == 0 goto 5f;"
+"3:"	"call bpf_unwind;"		/* cleanup region #2 */
+"4:"
+	"r0 = 0;"
+	"exit;"
+"5:"
+	"call bpf_rcu_read_unlock;"
+	"r0 = r6;"
+	"r0 += 1;"
+	"exit;"
+"6:"					/* landing pad, shared by both regions */
+	"call drop_glue;"
+	"call bpf_rcu_read_unlock;"
+	PAD_RAN("%[ran_rcu]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "6b")
+	CLEANUP_REC("3b", "4b", "6b")
+	:
+	: [ran_rcu]"i"(RAN_FOO2_RCU),
+	  __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+static __used __naked __noinline void foo1v(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+"1:"	"call foo2;"			/* cleanup region */
+"2:"
+	"r6 = r0;"
+	"call bpf_preempt_enable;"
+	"r1 = %[result] ll;"
+	"*(u64 *)(r1 + 0) = r6;"
+	"goto 7f;"
+"8:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	PAD_RAN("%[ran]")
+	"goto 9f;"
+"7:"					/* the frame's own exit block */
+	"r0 = 0;"
+	"exit;"
+"9:"					/* shared resume block */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "8b")
+	:
+	: [ran]"i"(RAN_FOO1V_PREEMPT), __imm_addr(input),
+	  __imm_addr(result), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+/*
+ * A frame with no cleanup record: an unwind leaving it runs no pad. The
+ * unwind never fires -- @never is global -- and the bit marks the return path.
+ */
+static __used __naked __noinline void bump(void)
+{
+	asm volatile (
+	PAD_RAN("%[ran]")
+	"r1 = %[never] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"if r1 == 0 goto 1f;"
+	"call bpf_unwind;"
+"1:"
+	"exit;"				/* r0 deliberately left alone */
+	:
+	: [ran]"i"(RAN_BUMP), __imm_addr(never), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+__noinline __u64 foo1(void)
+{
+	bump();
+	foo1v();
+	return result;
+}
+
+SEC("syscall")
+int entry(void *ctx)
+{
+	return foo1();
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
new file mode 100644
index 000000000000..4e51e3c4c5d9
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
@@ -0,0 +1,978 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_experimental.h"
+#include "bpf_misc.h"
+#include "../test_kmods/bpf_testmod_kfunc.h"
+#include "exceptions_cleanup.h"
+
+__u64 input = 0;
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+	bpf_throw(0);
+	bpf_unwind();
+	bpf_preempt_disable();
+	bpf_preempt_enable();
+	bpf_rcu_read_lock();
+	bpf_rcu_read_unlock();
+	bpf_unwind_resume(NULL);
+}
+
+/* An unwind raised in a callee, which is how a cleanup region gets one. */
+static __used __naked __noinline __u64 inner_unwind(void)
+{
+	asm volatile (
+	"call bpf_unwind;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+static int unwinding_cb(__u32 idx, void *ctx)
+{
+	bpf_unwind();
+	return 0;
+}
+
+/* Gives the program a table; the refusal is at the bpf_loop() call. */
+static __used __naked __noinline __u64 cb_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call unwinding_cb;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("may unwind and is used as a callback")
+int callback_may_unwind(void *ctx)
+{
+	bpf_loop(1, unwinding_cb, NULL, 0);
+	return cb_frame();
+}
+
+/* A pad that reaches both a resume and a plain exit. */
+static __used __naked __noinline __u64 ambiguous_pad_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"call bpf_preempt_disable;"
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad: two ways out */
+	"call bpf_preempt_enable;"
+	"if r6 > 10 goto 4f;"
+	"call bpf_unwind_resume;"
+	"exit;"
+"4:"
+	"r0 = 0;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: __imm_addr(input)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("ends a landing pad: a catch pad is not supported yet")
+int ambiguous_landing_pad(void *ctx)
+{
+	return ambiguous_pad_frame();
+}
+
+/* A second bpf_unwind() from inside a landing pad. */
+static __used __naked __noinline __u64 unwind_in_pad_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad that unwinds again */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("starts a second unwind while one is in flight")
+int unwind_from_landing_pad(void *ctx)
+{
+	return unwind_in_pad_frame();
+}
+
+__noinline int unused_exc_cb(u64 cookie)
+{
+	return 0;
+}
+
+/* A frame with a table and a pad, for tests whose refusal lies elsewhere. */
+static __used __naked __noinline __u64 table_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__exception_cb(unused_exc_cb)
+__failure __msg("cannot be combined with an exception callback")
+int table_with_exception_cb(void *ctx)
+{
+	return table_frame();
+}
+
+/*
+ * A throw and a table, with no callback tagged: the default callback is
+ * appended too late to stand in for the throw, so the program is scanned
+ * for one instead.
+ */
+static __used __naked __noinline __u64 throw_and_table_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("cannot be combined with bpf_throw")
+int table_with_throw(void *ctx)
+{
+	if (input)
+		bpf_throw(0);
+	return throw_and_table_frame();
+}
+
+__u64 never;
+
+/*
+ * A pad calling a global subprogram that can unwind. The subprogram is
+ * verified on its own, so the pad rule is what refuses it.
+ */
+__noinline void pad_callee_that_unwinds(void)
+{
+	if (never)
+		bpf_unwind();
+}
+
+static __used __naked __noinline __u64 pad_calls_unwinder_frame(void)
+{
+	asm volatile (
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call pad_callee_that_unwinds;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("which can unwind while an unwind is in flight")
+int pad_calls_unwinder(void *ctx)
+{
+	return pad_calls_unwinder_frame();
+}
+
+/*
+ * A pad calling a global subprogram that can throw. The throw is refused
+ * wherever it sits: the scan covers the subprograms too, not just the main
+ * program, so this never reaches the rules about pads.
+ */
+__noinline void pad_callee_that_throws(void)
+{
+	if (never)
+		bpf_throw(0);
+}
+
+static __used __naked __noinline __u64 pad_calls_thrower_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call pad_callee_that_throws;"	/* ...which can throw: refused */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("cannot be combined with bpf_throw")
+int pad_calls_thrower(void *ctx)
+{
+	return pad_calls_thrower_frame();
+}
+
+static __used __naked __noinline __u64 catch_pad_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* catch pad: no resume, it stops here */
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("ends a landing pad: a catch pad is not supported yet")
+int catch_landing_pad(void *ctx)
+{
+	return catch_pad_frame();
+}
+
+/* A bpf_unwind_resume() outside any landing pad. */
+SEC("?syscall")
+__failure __msg("is not in a landing pad")
+int resume_outside_pad(void *ctx)
+{
+	/* Never taken, but reachable, which is all the verifier needs. */
+	if (never)
+		bpf_unwind_resume(NULL);
+	return table_frame();
+}
+
+/* A bpf_unwind_resume() in a subprogram a landing pad calls. */
+static __used __naked __noinline void resume_in_callee(void)
+{
+	asm volatile (
+	"call bpf_unwind_resume;"
+	"exit;"
+	::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 pad_calls_resumer_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	"call resume_in_callee;"	/* ...which resumes: refused */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is not in a landing pad")
+int resume_in_pad_callee(void *ctx)
+{
+	return pad_calls_resumer_frame();
+}
+
+static __used __naked __noinline __u64 nested_pad_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call inner_unwind;"		/* first cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* first pad, second region's call */
+	"call bpf_preempt_enable;"
+"4:"
+	"call bpf_unwind_resume;"
+	"exit;"
+"5:"					/* second pad */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	CLEANUP_REC("3b", "4b", "5b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is inside the call-site range of")
+int nested_landing_pad(void *ctx)
+{
+	return nested_pad_frame();
+}
+
+/* A tail call in a landing pad: the frame would never reach its resume. */
+struct {
+	__uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+	__uint(max_entries, 1);
+	__uint(key_size, sizeof(__u32));
+	__uint(value_size, sizeof(__u32));
+} tc_map SEC(".maps");
+
+static __used __naked __noinline __u64 tail_call_pad_frame(void)
+{
+	asm volatile (
+	"r6 = r1;"
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r1 = r6;"
+	"r2 = %[tc_map] ll;"
+	"r3 = 0;"
+	"call %[bpf_tail_call];"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: __imm(bpf_tail_call), __imm_addr(tc_map)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("is a tail call, which replaces the frame, and is in a landing pad")
+int tail_call_in_pad(void *ctx)
+{
+	return tail_call_pad_frame();
+}
+
+#if defined(__BPF_FEATURE_STACK_ARGUMENT)
+
+/*
+ * A kfunc by-value argument that runs past the argument registers, in a
+ * landing pad. The pad is not what refuses it. The C call gives the extern
+ * its BTF.
+ */
+static __used __noinline void __nofit_btf_anchor(void)
+{
+	struct prog_test_pair_arg s = {};
+
+	bpf_kfunc_call_test_pair_arg_nofit(1, 2, 3, 4, s);
+}
+
+static __used __naked __noinline __u64 kfunc_arg_pad_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	"call bpf_kfunc_call_test_pair_arg_nofit;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("stack arg1 is not initialized")
+int kfunc_stack_arg_in_pad(void *ctx)
+{
+	return kfunc_arg_pad_frame();
+}
+
+#endif /* __BPF_FEATURE_STACK_ARGUMENT */
+
+/* A landing pad entered by ordinary control flow, with no unwind in flight. */
+static __used __naked __noinline __u64 jump_into_pad_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"if r6 > 7 goto 4f;"		/* an ordinary branch into the pad */
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r7 = r0;"
+"4:"					/* ... and its second instruction */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: __imm_addr(input)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("runs both inside and outside a landing pad")
+int jump_into_pad(void *ctx)
+{
+	return jump_into_pad_frame();
+}
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)
+
+/*
+ * An indirect jump in a landing pad. A jump table entry is an offset from
+ * the program's section symbol, which has to be spelled in quotes here.
+ */
+SEC("?syscall")
+__failure __msg("is an indirect jump, and is in a landing pad")
+__naked void gotox_in_pad(void)
+{
+	asm volatile (
+	".pushsection .jumptables,\"\",@progbits;"
+"jt0_%=:"
+	".quad l0_%= - \"?syscall\";"
+	".quad l1_%= - \"?syscall\";"
+	".size jt0_%=, 16;"
+	".global jt0_%=;"
+	".popsection;"
+
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r1 = jt0_%= ll;"
+	"r1 += 8;"
+	"r2 = *(u64 *)(r1 + 0);"
+	/*
+	 * gotox r2, as raw bytes: the mnemonic only reached the LLVM
+	 * assembler in llvm 22, and BPF_RAW_INSN() needs <linux/bpf.h>, which
+	 * vmlinux.h rules out. dst_reg is the other nibble on a big-endian
+	 * target.
+	 */
+#if __BYTE_ORDER__ == __ORDER_BIG_ENDIAN__
+	".byte 0x0d, 0x20, 0, 0, 0, 0, 0, 0;"
+#else
+	".byte 0x0d, 0x02, 0, 0, 0, 0, 0, 0;"
+#endif
+"l0_%=:"
+	"call bpf_unwind_resume;"
+	"exit;"
+"l1_%=:"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+#endif /* x86 || arm64 */
+
+/* A BPF_LD_[ABS|IND] in a pad: a failed load leaves without resuming. */
+static __used __naked __noinline __u64 ld_abs_pad_frame(void)
+{
+	asm volatile (
+	"r6 = r1;"			/* the skb BPF_LD_ABS reads */
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r0 = *(u32 *)skb[0];"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?tc")
+__failure __msg("is a BPF_LD_[ABS|IND], which can leave through the epilogue")
+__naked void ld_abs_in_pad(void)
+{
+	asm volatile (
+	"call ld_abs_pad_frame;"
+	"exit;"
+	::: __clobber_all);
+}
+
+/*
+ * A subprogram a landing pad calls, which unwinds on its own. The second
+ * unwind would rewrite return addresses the first has already redirected.
+ */
+static __used __naked __noinline __u64 own_pad_callee(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* its landing pad */
+	"call bpf_preempt_enable;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 pad_calls_own_pad_frame(void)
+{
+	asm volatile (
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad, which calls the above */
+	"call own_pad_callee;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("starts a second unwind while one is in flight")
+int unwind_in_pad_callee(void *ctx)
+{
+	return pad_calls_own_pad_frame();
+}
+
+/* A record whose range holds no call that can unwind. */
+static __used __naked __noinline __u64 nounwind_rec_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+"1:"	"call bpf_preempt_enable;"	/* cleanup region: nounwind */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad, reached by nothing */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("unreachable insn")
+int nounwind_region(void *ctx)
+{
+	return nounwind_rec_frame();
+}
+
+/*
+ * An unwind with no landing pad leaves the frame with nothing run on the way
+ * out, so what the frame holds is checked as it would be at a plain exit.
+ */
+SEC("?syscall")
+__failure __msg("an unwind with no landing pad cannot be used inside bpf_rcu_read_lock-ed region")
+int unwind_no_pad_rcu(void *ctx)
+{
+	bpf_rcu_read_lock();
+	bpf_unwind();
+	bpf_rcu_read_unlock();
+	return 0;
+}
+
+/*
+ * A frame leaves through an unwind without leaving its lock or reference
+ * state as it found it.
+ */
+static __used __naked __noinline __u64 pad_drops_caller_lock_frame(void)
+{
+	asm volatile (
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* pad: drops a lock it never took */
+	"call bpf_rcu_read_unlock;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 caller_holds_lock_frame(void)
+{
+	asm volatile (
+	"call bpf_rcu_read_lock;"
+"1:"	"call pad_drops_caller_lock_frame;"
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* pad */
+	"call bpf_rcu_read_unlock;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("a resume does not leave the frame's bpf_rcu_read_lock state as it found it")
+int pad_drops_caller_lock(void *ctx)
+{
+	return caller_holds_lock_frame();
+}
+
+/* The other way round: a pad that does not drop what its own frame took. */
+static __used __naked __noinline __u64 pad_keeps_own_lock_frame(void)
+{
+	asm volatile (
+	"call bpf_rcu_read_lock;"
+"1:"	"call inner_unwind;"		/* cleanup region */
+"2:"
+	"call bpf_rcu_read_unlock;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* pad: forgets the unlock */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("a resume does not leave the frame's bpf_rcu_read_lock state as it found it")
+int pad_keeps_own_lock(void *ctx)
+{
+	return pad_keeps_own_lock_frame();
+}
+
+/* And a subprog with no pad at all, leaving through an unwind holding one. */
+static __used __naked __noinline __u64 no_pad_keeps_own_lock_frame(void)
+{
+	asm volatile (
+	"call bpf_rcu_read_lock;"
+	"call bpf_unwind;"		/* no record covers it */
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure
+__msg("no landing pad does not leave the frame's bpf_rcu_read_lock state")
+int no_pad_keeps_own_lock(void *ctx)
+{
+	return no_pad_keeps_own_lock_frame();
+}
+
+struct {
+	__uint(type, BPF_MAP_TYPE_RINGBUF);
+	__uint(max_entries, 4096);
+} unwind_ringbuf SEC(".maps");
+
+/*
+ * Always unwinds, so its caller is never returned to on the modelled path --
+ * which is what keeps the release below out of the caller's post-call code.
+ * Its pad discards the record the caller reserved.
+ */
+static __used __naked __noinline __u64 pad_drops_caller_ref_frame(void)
+{
+	asm volatile (
+	"r6 = r1;"			/* the caller's reserved record */
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* pad: drops what it never acquired */
+	"r1 = r6;"
+	"r2 = 0;"
+	"call %[bpf_ringbuf_discard];"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: __imm(bpf_ringbuf_discard)
+	: __clobber_all);
+}
+
+/* And this frame's own pad drops it a second time. */
+static __used __naked __noinline __u64 caller_holds_ref_frame(void)
+{
+	asm volatile (
+	"r1 = %[unwind_ringbuf] ll;"
+	"r2 = 8;"
+	"r3 = 0;"
+	"call %[bpf_ringbuf_reserve];"
+	"if r0 == 0 goto 9f;"
+	"r6 = r0;"
+	"r1 = r6;"
+"1:"	"call pad_drops_caller_ref_frame;"	/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* pad */
+	"r1 = r6;"
+	"r2 = 0;"
+	"call %[bpf_ringbuf_discard];"
+	"call bpf_unwind_resume;"
+	"exit;"
+"9:"
+	"r0 = 0;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: __imm(bpf_ringbuf_reserve), __imm(bpf_ringbuf_discard),
+	  __imm_addr(unwind_ringbuf)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("a resume does not leave the frame's references as it found it")
+int pad_drops_caller_ref(void *ctx)
+{
+	return caller_holds_ref_frame();
+}
+
+/*
+ * A frame an unwind returns through without a pad is abandoned where it made
+ * the call: the JIT sends it to its epilogue, so nothing of it runs again and
+ * whatever it acquired is never released. It has to hold what it entered with
+ * at every such call.
+ */
+static __used __naked __noinline __u64 pad_resumes_frame(void)
+{
+	asm volatile (
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* pad */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 uncovered_holds_lock_frame(void)
+{
+	asm volatile (
+	"call bpf_rcu_read_lock;"
+	"call pad_resumes_frame;"	/* no record covers this call */
+	"call bpf_rcu_read_unlock;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure
+__msg("through this call does not leave the frame's bpf_rcu_read_lock state")
+int unwind_through_call_keeps_lock(void *ctx)
+{
+	return uncovered_holds_lock_frame();
+}
+
+static __used __naked __noinline __u64 uncovered_holds_ref_frame(void)
+{
+	asm volatile (
+	"r1 = %[unwind_ringbuf] ll;"
+	"r2 = 8;"
+	"r3 = 0;"
+	"call %[bpf_ringbuf_reserve];"
+	"if r0 == 0 goto 9f;"
+	"r6 = r0;"
+	"call pad_resumes_frame;"	/* no record covers this call */
+	"r1 = r6;"
+	"r2 = 0;"
+	"call %[bpf_ringbuf_discard];"
+"9:"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm(bpf_ringbuf_reserve), __imm(bpf_ringbuf_discard),
+	  __imm_addr(unwind_ringbuf)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("an unwind through this call keeps the reference id=")
+int unwind_through_call_keeps_ref(void *ctx)
+{
+	return uncovered_holds_ref_frame();
+}
+
+/* The main program's frame is passed by the same way. */
+SEC("?syscall")
+__failure __msg("an unwind through this call keeps the reference id=")
+int unwind_through_call_main_keeps_ref(void *ctx)
+{
+	void *rec;
+
+	rec = bpf_ringbuf_reserve(&unwind_ringbuf, 8, 0);
+	if (!rec)
+		return 0;
+	pad_resumes_frame();		/* no record covers this call */
+	bpf_ringbuf_discard(rec, 0);
+	return 0;
+}
+
+/*
+ * A callee writes its caller's stack through a pointer argument, then
+ * unwinds. The caller's pad runs after that write, so it cannot keep trusting
+ * the slot to hold the zero it held at the call.
+ */
+static __used __naked __noinline __u64 stack_writer(void)
+{
+	asm volatile (
+	"r2 = 0x10000000;"
+	"*(u64 *)(r1 + 0) = r2;"	/* r1 is the caller's fp-8 */
+	"call bpf_unwind;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 stale_stack_frame(void)
+{
+	asm volatile (
+	"r6 = 0;"
+	"*(u64 *)(r10 - 8) = r6;"
+	"*(u64 *)(r10 - 64) = r6;"
+	"r1 = r10;"
+	"r1 += -8;"
+"1:"	"call stack_writer;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* pad: fp-8 as an offset into fp-64 */
+	"r1 = *(u64 *)(r10 - 8);"
+	"r2 = r10;"
+	"r2 += -64;"
+	"r2 += r1;"
+	"r0 = *(u8 *)(r2 + 0);"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("invalid read from stack R2 off=268435392 size=1")
+int stale_stack_pad(void *ctx)
+{
+	return stale_stack_frame();
+}
+
+/* The same through a global subprog, which is not walked from its caller. */
+__noinline int global_stack_writer(__u64 *p)
+{
+	if (!p)
+		return 0;
+	*p = 0x10000000;
+	bpf_unwind();
+	return 0;
+}
+
+static __used __naked __noinline __u64 global_stale_stack_frame(void)
+{
+	asm volatile (
+	"r6 = 0;"
+	"*(u64 *)(r10 - 8) = r6;"
+	"*(u64 *)(r10 - 64) = r6;"
+	"r1 = r10;"
+	"r1 += -8;"
+"1:"	"call global_stack_writer;"	/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* pad: fp-8 as an offset into fp-64 */
+	"r1 = *(u64 *)(r10 - 8);"
+	"r2 = r10;"
+	"r2 += -64;"
+	"r2 += r1;"
+	"r0 = *(u8 *)(r2 + 0);"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("math between fp pointer and register with unbounded min value")
+int global_stale_stack_pad(void *ctx)
+{
+	return global_stale_stack_frame();
+}
+
+/* gcc has no indirect calls, and only these JITs emit them */
+#if defined(__clang__) && \
+	(defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64))
+
+/*
+ * A callback calling an unwinding subprog through a pointer it read from its
+ * caller's stack, rather than one it loaded itself.
+ */
+static __used __naked __noinline int callx_cb(void)
+{
+	asm volatile (
+	"r1 = *(u64 *)(r2 + 0);"
+	"callx r1;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__failure __msg("may unwind and is used as a callback")
+__naked int callback_callx_may_unwind(void)
+{
+	asm volatile (
+	"r1 = %[inner_unwind] ll;"
+	"*(u64 *)(r10 - 8) = r1;"
+	"r1 = 1;"
+	"r2 = %[callx_cb] ll;"
+	"r3 = r10;"
+	"r3 += -8;"
+	"r4 = 0;"
+	"call %[bpf_loop];"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm_addr(inner_unwind), __imm_addr(callx_cb), __imm(bpf_loop)
+	: __clobber_all);
+}
+
+#endif /* __clang__ && (x86 || arm64) */
+
+char _license[] SEC("license") = "GPL";
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (18 preceding siblings ...)
  2026-10-01 13:31 ` [PATCH bpf-next v8 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
@ 2026-10-01 13:31 ` Yonghong Song
  2026-10-01 13:31 ` [PATCH bpf-next v8 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
  2026-10-01 13:32 ` [PATCH bpf-next v8 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:31 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

A test that has to look at a program's state after it ran needs a driver of
its own: open and load a skeleton, set an input, call
bpf_prog_test_run_opts(), read a variable back, destroy the skeleton. The
loader already does all of that for __retval(), and the only thing missing
is reaching the program's globals.

__set_global(var, value) writes one before the run and __ret_global(var,
value) checks one after it, and both make the loader run the program and
check its return value; without __retval(), the program must return 0. Each
tag may be given more than once, for a test with more than one input or
more than one thing to look at afterwards; past MAX_GLOBAL_VARS a tag is
refused rather than quietly replacing the one before it.

The variable is found the way veristat finds one: the map libbpf made for
the ".bss" or ".data" section, looked up by that section name, the datasec
of the same name in the object's BTF, then the variable within it. Four
and eight byte variables are supported, which is what a counter or a
bitmask needs; anything else is refused rather than read at the wrong
width.

A value is held to the range its variable can represent, signedness taken
from the BTF the way veristat's set_global_var() does, and narrowed to
what is stored. Both tags use the one rule, so a negative literal means
the same to each: without it __ret_global(err, -22) on an int could never
match, the tag parsing 64 bits while the read gave back 32. Unlike
veristat this takes no enum names and runs native-endian only.

The value may be a '|' separated list of terms, so that a bitmask reads in
the test the way it is written in the program.

Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/testing/selftests/bpf/progs/bpf_misc.h |  12 +
 tools/testing/selftests/bpf/test_loader.c    | 319 ++++++++++++++++++-
 2 files changed, 330 insertions(+), 1 deletion(-)

diff --git a/tools/testing/selftests/bpf/progs/bpf_misc.h b/tools/testing/selftests/bpf/progs/bpf_misc.h
index f3dbc3b59bff..a80798a8f28f 100644
--- a/tools/testing/selftests/bpf/progs/bpf_misc.h
+++ b/tools/testing/selftests/bpf/progs/bpf_misc.h
@@ -103,6 +103,16 @@
  *                   - POINTER_VALUE
  *                   - TEST_DATA_LEN
  * __retval_unpriv   Same, but load program in unprivileged mode.
+ * __set_global      Set a global variable of the program to a value before
+ *                   executing it. The variable has to be an int or an enum,
+ *                   live in .bss or .data and be four or eight bytes wide,
+ *                   and the value is a number or several joined by '|',
+ *                   without parentheses.
+ * __ret_global      Execute the program and check that a global variable
+ *                   holds the given value afterwards, under the same rules.
+ *                   __set_global and __ret_global both make the loader run
+ *                   the program and check its return value; without
+ *                   __retval, the program must return 0.
  *
  * __description     Text to be used for display and as an additional filter
  *                   alias, while the original program name stays matchable.
@@ -161,6 +171,8 @@
 #define __flag(flag)		__test_tag("test_prog_flags=" #flag)
 #define __retval(val)		__test_tag("test_retval=" XSTR(val))
 #define __retval_unpriv(val)	__test_tag("test_retval_unpriv=" XSTR(val))
+#define __set_global(var, val)	__test_tag("test_global_set=" #var ":" XSTR(val))
+#define __ret_global(var, val)	__test_tag("test_global_ret=" #var ":" XSTR(val))
 #define __auxiliary		__test_tag("test_auxiliary")
 #define __auxiliary_unpriv	__test_tag("test_auxiliary_unpriv")
 #define __btf_path(path)	__test_tag("test_btf_path=" path)
diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
index 25eeb1c1248b..0ff89912e6f1 100644
--- a/tools/testing/selftests/bpf/test_loader.c
+++ b/tools/testing/selftests/bpf/test_loader.c
@@ -50,6 +50,13 @@ enum stack_mode {
 	SMALL_STACK	= 1 << 1,
 };
 
+#define MAX_GLOBAL_VARS 8
+
+struct global_var {
+	char *name;
+	__u64 val;
+};
+
 struct test_subspec {
 	char *name;
 	char *description;
@@ -62,6 +69,10 @@ struct test_subspec {
 	int retval;
 	bool execute;
 	__u64 caps;
+	struct global_var set_globals[MAX_GLOBAL_VARS];
+	int set_global_cnt;
+	struct global_var ret_globals[MAX_GLOBAL_VARS];
+	int ret_global_cnt;
 };
 
 struct test_spec {
@@ -115,6 +126,34 @@ void free_msgs(struct expected_msgs *msgs)
 	msgs->cnt = 0;
 }
 
+static void free_global_vars(struct global_var *vars, int *cnt)
+{
+	int i;
+
+	for (i = 0; i < *cnt; i++) {
+		free(vars[i].name);
+		vars[i].name = NULL;
+	}
+	*cnt = 0;
+}
+
+static int clone_global_vars(struct global_var *dst, int *dst_cnt,
+			     const struct global_var *src, int src_cnt)
+{
+	int i;
+
+	for (i = 0; i < src_cnt; i++) {
+		dst[i].name = strdup(src[i].name);
+		if (!dst[i].name) {
+			free_global_vars(dst, dst_cnt);
+			return -ENOMEM;
+		}
+		dst[i].val = src[i].val;
+		(*dst_cnt)++;
+	}
+	return 0;
+}
+
 static void free_test_spec(struct test_spec *spec)
 {
 	/* Deallocate expect_msgs arrays. */
@@ -129,6 +168,11 @@ static void free_test_spec(struct test_spec *spec)
 	free_msgs(&spec->unpriv.stdout);
 	free_msgs(&spec->priv.stdout);
 
+	free_global_vars(spec->priv.set_globals, &spec->priv.set_global_cnt);
+	free_global_vars(spec->priv.ret_globals, &spec->priv.ret_global_cnt);
+	free_global_vars(spec->unpriv.set_globals, &spec->unpriv.set_global_cnt);
+	free_global_vars(spec->unpriv.ret_globals, &spec->unpriv.ret_global_cnt);
+
 	free(spec->priv.name);
 	free(spec->priv.description);
 	free(spec->unpriv.name);
@@ -311,6 +355,236 @@ static int parse_caps(const char *str, __u64 *val, const char *name)
 	return 0;
 }
 
+static int parse_global_var(const char *str, struct global_var *vars, int *cnt,
+			    const char *name)
+{
+	const char *colon = strrchr(str, ':');
+	struct global_var *var;
+	char *end;
+	__u64 *val;
+
+	if (!colon || colon == str) {
+		PRINT_FAIL("expecting '<variable>:<value>' for %s, got '%s'\n", name, str);
+		return -EINVAL;
+	}
+	if (*cnt >= MAX_GLOBAL_VARS) {
+		PRINT_FAIL("too many %s tags, at most %d are supported\n",
+			   name, MAX_GLOBAL_VARS);
+		return -E2BIG;
+	}
+
+	var = &vars[*cnt];
+	val = &var->val;
+	*val = 0;
+	for (const char *term = colon + 1;;) {
+		__u64 v;
+
+		errno = 0;
+		v = strtoull(term, &end, 0);
+		if (errno || end == term) {
+			PRINT_FAIL("failed to parse %s value '%s'\n", name, colon + 1);
+			return -EINVAL;
+		}
+		*val |= v;
+		while (*end == ' ')
+			end++;
+		if (!*end)
+			break;
+		if (*end != '|') {
+			PRINT_FAIL("failed to parse %s value '%s'\n", name, colon + 1);
+			return -EINVAL;
+		}
+		term = end + 1;
+	}
+
+	var->name = strndup(str, colon - str);
+	if (!var->name) {
+		PRINT_FAIL("failed to allocate %s variable name\n", name);
+		return -ENOMEM;
+	}
+	(*cnt)++;
+
+	return 0;
+}
+
+/* As veristat's is_signed_type(): anything not plainly unsigned is signed. */
+static bool global_var_is_signed(const struct btf_type *t)
+{
+	if (btf_is_int(t))
+		return btf_int_encoding(t) & BTF_INT_SIGNED;
+	if (btf_is_any_enum(t))
+		return btf_kflag(t);
+	return true;
+}
+
+static int find_global_var(struct bpf_object *obj, const char *name,
+			   struct bpf_map **map, __u32 *off, __u32 *sz,
+			   bool *is_signed)
+{
+	static const char * const secs[] = { ".bss", ".data" };
+	struct btf *btf = bpf_object__btf(obj);
+	int i, s;
+
+	if (!btf) {
+		PRINT_FAIL("no BTF for object\n");
+		return -ENOENT;
+	}
+
+	for (s = 0; s < ARRAY_SIZE(secs); s++) {
+		const struct btf_type *sec, *vt;
+		const struct btf_var_secinfo *vsi;
+		struct bpf_map *m = bpf_object__find_map_by_name(obj, secs[s]);
+		int id;
+
+		id = btf__find_by_name_kind(btf, secs[s], BTF_KIND_DATASEC);
+		if (!m || id < 0)
+			continue;
+
+		sec = btf__type_by_id(btf, id);
+		vsi = btf_var_secinfos(sec);
+		for (i = 0; i < btf_vlen(sec); i++, vsi++) {
+			const struct btf_type *var = btf__type_by_id(btf, vsi->type);
+
+			if (strcmp(btf__name_by_offset(btf, var->name_off), name))
+				continue;
+			if (vsi->size != 4 && vsi->size != 8) {
+				PRINT_FAIL("'%s' is %u bytes, only 4 and 8 are supported\n",
+					   name, vsi->size);
+				return -EINVAL;
+			}
+			vt = btf__type_by_id(btf, btf__resolve_type(btf, var->type));
+			if (!vt || !(btf_is_int(vt) || btf_is_any_enum(vt))) {
+				PRINT_FAIL("'%s' is not an int or an enum\n", name);
+				return -EINVAL;
+			}
+			*is_signed = global_var_is_signed(vt);
+			*map = m;
+			*off = vsi->offset;
+			*sz = vsi->size;
+			return 0;
+		}
+	}
+
+	PRINT_FAIL("no global variable '%s'\n", name);
+	return -ENOENT;
+}
+
+/*
+ * A tag's value is parsed as 64 bits, but the variable may be narrower and
+ * may be signed. Hold it to the range the variable can represent, the way
+ * veristat's set_global_var() does, and narrow it to what is stored.
+ */
+static int fit_global_var(const char *name, __u32 sz, bool is_signed, __u64 *val)
+{
+	long long v = (long long)*val;
+	long long max_val;
+	__u32 bits;
+
+	if (sz >= sizeof(*val))
+		return 0;
+	bits = sz * 8 - (is_signed ? 1 : 0);
+	max_val = 1ll << bits;
+	if (v >= max_val || v < (is_signed ? -max_val : 0)) {
+		PRINT_FAIL("value %lld for '%s' is out of range [%lld; %lld]\n",
+			   v, name, is_signed ? -max_val : 0, max_val - 1);
+		return -EINVAL;
+	}
+	*val = (__u32)*val;
+	return 0;
+}
+
+/* The value of @map's single element, which the caller frees. */
+static void *global_data(struct bpf_map *map, const char *name, size_t *vsz)
+{
+	__u32 zero = 0;
+	void *buf;
+	int err;
+
+	*vsz = bpf_map__value_size(map);
+	buf = calloc(1, *vsz);
+	if (!buf) {
+		PRINT_FAIL("failed to allocate %zu bytes for '%s'\n", *vsz, name);
+		return NULL;
+	}
+	err = bpf_map__lookup_elem(map, &zero, sizeof(zero), buf, *vsz, 0);
+	if (err) {
+		PRINT_FAIL("failed to read '%s': %d\n", name, err);
+		free(buf);
+		return NULL;
+	}
+	return buf;
+}
+
+static int read_global_var(struct bpf_map *map, const char *name, __u32 off,
+			   __u32 sz, __u64 *val)
+{
+	size_t vsz;
+	void *buf;
+
+	buf = global_data(map, name, &vsz);
+	if (!buf)
+		return -EINVAL;
+	*val = sz == 4 ? *(__u32 *)(buf + off) : *(__u64 *)(buf + off);
+	free(buf);
+	return 0;
+}
+
+static int write_global_var(struct bpf_map *map, const char *name, __u32 off,
+			    __u32 sz, __u64 val)
+{
+	__u32 zero = 0;
+	size_t vsz;
+	void *buf;
+	int err;
+
+	buf = global_data(map, name, &vsz);
+	if (!buf)
+		return -EINVAL;
+	if (sz == 4)
+		*(__u32 *)(buf + off) = val;
+	else
+		*(__u64 *)(buf + off) = val;
+	err = bpf_map__update_elem(map, &zero, sizeof(zero), buf, vsz, 0);
+	if (err)
+		PRINT_FAIL("failed to write '%s': %d\n", name, err);
+	free(buf);
+	return err;
+}
+
+/* Write a __set_global() value into the program's global variable. */
+static int set_global_var(struct bpf_object *obj, const struct global_var *var)
+{
+	__u64 val = var->val;
+	struct bpf_map *map;
+	__u32 off, sz;
+	bool is_signed;
+
+	if (find_global_var(obj, var->name, &map, &off, &sz, &is_signed) ||
+	    fit_global_var(var->name, sz, is_signed, &val))
+		return -EINVAL;
+	return write_global_var(map, var->name, off, sz, val);
+}
+
+/* Check the program's global variable against a __ret_global() value. */
+static int check_global_var(struct bpf_object *obj, const struct global_var *var)
+{
+	__u64 want = var->val, val;
+	struct bpf_map *map;
+	__u32 off, sz;
+	bool is_signed;
+
+	if (find_global_var(obj, var->name, &map, &off, &sz, &is_signed) ||
+	    fit_global_var(var->name, sz, is_signed, &want) ||
+	    read_global_var(map, var->name, off, sz, &val))
+		return -EINVAL;
+	if (val != want) {
+		PRINT_FAIL("Unexpected %s: 0x%llx != 0x%llx\n", var->name,
+			   (unsigned long long)val, (unsigned long long)want);
+		return -EINVAL;
+	}
+	return 0;
+}
+
 static int parse_retval(const char *str, int *val, const char *name)
 {
 	/*
@@ -557,6 +831,22 @@ static int parse_test_spec(struct test_loader *tester,
 			spec->mode_mask |= UNPRIV;
 			spec->unpriv.execute = true;
 			has_unpriv_retval = true;
+		} else if ((val = str_has_pfx(s, "test_global_set="))) {
+			err = parse_global_var(val, spec->priv.set_globals,
+					       &spec->priv.set_global_cnt,
+					       "__set_global");
+			if (err)
+				goto cleanup;
+			spec->priv.execute = true;
+			spec->mode_mask |= PRIV;
+		} else if ((val = str_has_pfx(s, "test_global_ret="))) {
+			err = parse_global_var(val, spec->priv.ret_globals,
+					       &spec->priv.ret_global_cnt,
+					       "__ret_global");
+			if (err)
+				goto cleanup;
+			spec->priv.execute = true;
+			spec->mode_mask |= PRIV;
 		} else if ((val = str_has_pfx(s, "test_log_level="))) {
 			err = parse_int(val, &spec->log_level, "test log level");
 			if (err)
@@ -742,6 +1032,23 @@ static int parse_test_spec(struct test_loader *tester,
 			spec->unpriv.execute = spec->priv.execute;
 		}
 
+		if (spec->priv.set_global_cnt && !spec->unpriv.set_global_cnt) {
+			err = clone_global_vars(spec->unpriv.set_globals,
+						&spec->unpriv.set_global_cnt,
+						spec->priv.set_globals,
+						spec->priv.set_global_cnt);
+			if (err)
+				goto cleanup;
+		}
+		if (spec->priv.ret_global_cnt && !spec->unpriv.ret_global_cnt) {
+			err = clone_global_vars(spec->unpriv.ret_globals,
+						&spec->unpriv.ret_global_cnt,
+						spec->priv.ret_globals,
+						spec->priv.ret_global_cnt);
+			if (err)
+				goto cleanup;
+		}
+
 		if (spec->unpriv.expect_msgs.cnt == 0)
 			clone_msgs(&spec->priv.expect_msgs, &spec->unpriv.expect_msgs);
 		if (spec->unpriv.expect_xlated.cnt == 0)
@@ -1356,7 +1663,7 @@ void run_subtest(struct test_loader *tester,
 	struct cap_state caps = {};
 	struct bpf_object *tobj;
 	struct bpf_map *map;
-	int retval, err, i;
+	int retval, err, i, j;
 	int links_cnt = 0;
 	bool should_load;
 
@@ -1532,6 +1839,11 @@ void run_subtest(struct test_loader *tester,
 			}
 		}
 
+		for (j = 0; j < subspec->set_global_cnt; j++) {
+			if (set_global_var(tobj, &subspec->set_globals[j]))
+				goto tobj_cleanup;
+		}
+
 		err = do_prog_test_run(bpf_program__fd(tprog), &retval,
 				       bpf_program__type(tprog) == BPF_PROG_TYPE_SYSCALL ? true : false,
 				       spec->linear_sz);
@@ -1540,6 +1852,11 @@ void run_subtest(struct test_loader *tester,
 			goto tobj_cleanup;
 		}
 
+		for (j = 0; j < subspec->ret_global_cnt; j++) {
+			if (check_global_var(tobj, &subspec->ret_globals[j]))
+				goto tobj_cleanup;
+		}
+
 		verify_stderr(bpf_program__fd(tprog), &subspec->stderr);
 
 		if (subspec->stdout.cnt) {
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (19 preceding siblings ...)
  2026-10-01 13:31 ` [PATCH bpf-next v8 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
@ 2026-10-01 13:31 ` Yonghong Song
  2026-10-01 13:32 ` [PATCH bpf-next v8 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:31 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

The end-to-end test walks one call chain with a pad in most of its frames.
This adds the shapes it does not reach, in the order the new file has
them:

 - a pad that reads its frame's callee-saved registers
 - a region ending on a 16-byte instruction
 - a pad terminated by _Unwind_Resume rather than bpf_unwind_resume
 - a pad whose first instruction is a nop
 - a pad that indexes its frame by a register the frame set before the
   unwinding call
 - two pads with an uncovered frame between them
 - a precision chain from a pad back across the resume that led to it
 - a frame above an unwind that never returns, which dead code removal
   would leave with no exit and no epilogue, ending in a jump rather than
   an exit so the one kept is not the last instruction
 - a region covering an indirect call
 - a subprogram calling an unwinding one through a pointer it was handed
 - a jump into a pad that only a speculative walk takes, after the unwind
   has marked the pad, loaded without CAP_PERFMON so that the walk
   happens, and not run
 - a pad that touches a global of each width and sign a tag can name
 - a frame holding a reference across a covered call, its pad releasing
   it on the unwind path
 - a pad reading a slot its frame's callee wrote before it unwound
 - a precision chain from a pad back into the frame the unwind left
 - a pad reading a slot a global callee wrote before it unwound
 - a pad reached by an unwind out of a global subprog called with no
   record, one frame above it
 - a bpf_unwind() inside a loop, with no pad between it and the main
   program
 - a frame's own pad finding r0 zero after its bpf_unwind()
 - a pad insn a speculative walk reaches from outside the pad before the
   real walk reaches it from the unwind, and a pad whose speculative walk
   reaches an exit, both loaded without CAP_PERFMON
 - an unwind reaching the main program's exit in a program type whose
   return value is checked

The three speculative shapes check the translated program for the barrier
the speculative walk leaves, so they cannot pass with no such walk.

Each shape that runs wants the same three things: an input, a return
value, and the set of landing pads that ran. That is what __set_global(),
__retval() and __ret_global() say, so they say it and RUN_TESTS() does the
rest, which also gives each shape a name of its own in the test output.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 .../selftests/bpf/exceptions_cleanup.h        |   18 +
 .../bpf/prog_tests/exceptions_cleanup.c       |    2 +
 .../bpf/progs/exceptions_cleanup_shapes.c     | 1003 +++++++++++++++++
 3 files changed, 1023 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c

diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
index 96effd2c1361..658cdf63c59b 100644
--- a/tools/testing/selftests/bpf/exceptions_cleanup.h
+++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
@@ -10,6 +10,24 @@
 #define RAN_FOO2_DROP		0x8
 #define RAN_BUMP		0x10
 
+/* progs/exceptions_cleanup_shapes.c: one bit per landing pad. */
+#define RAN_REGS		0x1
+#define RAN_WIDE_REC		0x2
+#define RAN_RESUME_ALIAS	0x4
+#define RAN_NOP_PAD		0x8
+#define RAN_VAR_STACK		0x10
+#define RAN_GAP_INNER		0x20
+#define RAN_GAP_OUTER		0x40
+#define RAN_PREC_RESUME		0x80
+#define RAN_NO_EXIT_JA		0x100
+#define RAN_CALLX		0x200
+#define RAN_HELD_REF		0x400
+#define RAN_CALLEE_WRITE	0x800
+#define RAN_CALLEE_OFFSET	0x1000
+#define RAN_GLOBAL_WRITE	0x2000
+#define RAN_THROUGH_GLOBAL	0x4000
+#define RAN_OWN_PAD_R0		0x8000
+
 #define CLEANUP_REC(begin, end, landing_pad)			\
 	".pushsection .bpf_cleanup,\"a\",@progbits;"		\
 	".long " begin ";"					\
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
index 255f88d35aad..c06ec10359b9 100644
--- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -4,6 +4,7 @@
 #include "exceptions_cleanup.h"
 #include "exceptions_cleanup.skel.h"
 #include "exceptions_cleanup_fail.skel.h"
+#include "exceptions_cleanup_shapes.skel.h"
 
 /* foo3 unwound: every frame that has a pad ran it. */
 #define PADS_FOO3_UNWOUND \
@@ -82,4 +83,5 @@ void test_exceptions_cleanup(void)
 	exceptions_cleanup__destroy(skel);
 
 	RUN_TESTS(exceptions_cleanup_fail);
+	RUN_TESTS(exceptions_cleanup_shapes);
 }
diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
new file mode 100644
index 000000000000..281691a2e1f5
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
@@ -0,0 +1,1003 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "exceptions_cleanup.h"
+
+static __used __noinline void __kfunc_btf_anchor(void)
+{
+	bpf_unwind();
+	bpf_rcu_read_lock();
+	bpf_rcu_read_unlock();
+	bpf_preempt_disable();
+	bpf_preempt_enable();
+	bpf_unwind_resume(NULL);
+}
+
+__u64 input = 0;
+__u64 magic = 0x5eed;
+__u64 pads_ran = 0;
+
+/* Load r6-r9 with values derived from @magic. */
+#define LOAD_MAGIC_REGS						\
+	"r1 = %[magic] ll;"					\
+	"r6 = *(u64 *)(r1 + 0);"				\
+	"r7 = r6;"						\
+	"r7 += 1;"						\
+	"r8 = r6;"						\
+	"r8 += 2;"						\
+	"r9 = r6;"						\
+	"r9 += 3;"
+
+/* Set @bit only if r6-r9 still hold what LOAD_MAGIC_REGS put there. */
+#define CHECK_MAGIC_REGS(bit)					\
+	"r1 = %[magic] ll;"					\
+	"r2 = *(u64 *)(r1 + 0);"				\
+	"if r6 != r2 goto 9f;"					\
+	"r2 += 1;"						\
+	"if r7 != r2 goto 9f;"					\
+	"r2 += 1;"						\
+	"if r8 != r2 goto 9f;"					\
+	"r2 += 1;"						\
+	"if r9 != r2 goto 9f;"					\
+	PAD_RAN(bit)						\
+	"9:"
+
+/* A callee that unwinds when its argument is over 100. */
+static __used __noinline __u64 pc_unwinder(__u64 x)
+{
+	if (x > 100)
+		bpf_unwind();
+	return x + 1;
+}
+
+static __used __naked __noinline __u64 regs_unwinder(void)
+{
+	asm volatile (
+	/* Not this frame's to keep, and that is the point. */
+	"r6 = 0xdead;"
+	"r7 = 0xbeef;"
+	"r8 = 0xcafe;"
+	"r9 = 0xf00d;"
+	"call bpf_unwind;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+/* A pad that reads r6-r9, which the callee overwrote before it unwound. */
+static __used __naked __noinline __u64 regs_frame(void)
+{
+	asm volatile (
+	LOAD_MAGIC_REGS
+	"call bpf_preempt_disable;"
+"1:"	"call regs_unwinder;"		/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	CHECK_MAGIC_REGS("%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_REGS),
+	  __imm_addr(magic), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__ret_global(pads_ran, RAN_REGS)
+int entry_regs(void *ctx)
+{
+	return regs_frame();
+}
+
+/* A region ending on a 16-byte insn, so end - 1 names its second half. */
+static __used __naked __noinline __u64 wide_rec_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"call bpf_rcu_read_lock;"
+	"r1 = r6;"
+"1:"	"call pc_unwinder;"		/* cleanup region begins */
+	"r1 = %[magic] ll;"		/* ... and ends on this pair */
+"2:"
+	"r6 = r0;"
+	"call bpf_rcu_read_unlock;"
+	"r0 = r6;"
+	"exit;"
+"3:"					/* landing pad */
+	"r7 = r0;"
+	"call bpf_rcu_read_unlock;"
+	PAD_RAN("%[ran]")
+	"r1 = r7;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_WIDE_REC), __imm_addr(input), __imm_addr(magic),
+	  __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_WIDE_REC)
+int entry_wide_rec(void *ctx)
+{
+	return wide_rec_frame();
+}
+
+/* The name LLVM gives the resume: _Unwind_Resume(), which libbpf maps over. */
+extern void _Unwind_Resume(void *ptr) __ksym;
+
+static __used __noinline void __resume_alias_btf_anchor(void)
+{
+	_Unwind_Resume(NULL);
+}
+
+static __used __naked __noinline __u64 resume_alias_frame(void)
+{
+	asm volatile (
+"1:"	"call regs_unwinder;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	PAD_RAN("%[ran]")
+	"call _Unwind_Resume;"		/* the frontend's name for it */
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_RESUME_ALIAS), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__ret_global(pads_ran, RAN_RESUME_ALIAS)
+int entry_resume_alias(void *ctx)
+{
+	return resume_alias_frame();
+}
+
+/* A pad starting on a nop, which opt_remove_nops() drops after the walk. */
+static __used __naked __noinline __u64 nop_pad_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"if r6 < 101 goto 6f;"
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+"6:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad: a nop, then its body */
+	"goto +0;"
+	PAD_RAN("%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_NOP_PAD),
+	  __imm_addr(input), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_NOP_PAD)
+int entry_nop_pad(void *ctx)
+{
+	return nop_pad_frame();
+}
+
+/* Put @magic in both of the slots a variable offset could name. */
+#define FILL_MAGIC_SLOTS					\
+	"r1 = %[magic] ll;"					\
+	"r1 = *(u64 *)(r1 + 0);"				\
+	"*(u64 *)(r10 - 8) = r1;"				\
+	"*(u64 *)(r10 - 16) = r1;"
+
+/* Set @bit if the slot @idx names, read at a variable offset, holds it. */
+#define CHECK_VAR_SLOT(idx, bit)				\
+	"r1 = r10;"						\
+	"r1 += " idx ";"					\
+	"r2 = *(u64 *)(r1 - 16);"				\
+	"r3 = %[magic] ll;"					\
+	"r3 = *(u64 *)(r3 + 0);"				\
+	"if r2 != r3 goto 9f;"					\
+	PAD_RAN(bit)						\
+	"9:"
+
+/* A callee that unwinds when r1 is at least 101, and touches none of r6-r9. */
+static __used __naked __noinline __u64 var_unwinder(void)
+{
+	asm volatile (
+	"if r1 < 101 goto 1f;"
+	"call bpf_unwind;"
+"1:"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 var_stack_frame(void)
+{
+	asm volatile (
+	FILL_MAGIC_SLOTS
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"r6 &= 1;"			/* an unknown slot number... */
+	"r6 <<= 3;"			/* ...as an aligned byte offset */
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+"1:"	"call var_unwinder;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r7 = r0;"
+	CHECK_VAR_SLOT("r6", "%[ran]")
+	"r1 = r7;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_VAR_STACK), __imm_addr(input), __imm_addr(magic),
+	  __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_VAR_STACK)
+int entry_var_stack(void *ctx)
+{
+	return var_stack_frame();
+}
+
+/* Two pads with an uncovered frame between them. */
+static __used __naked __noinline __u64 gap_inner_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+"1:"	"call pc_unwinder;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r6 = r0;"
+	PAD_RAN("%[ran]")
+	"r1 = r6;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_GAP_INNER), __imm_addr(input), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+/* The frame in between, with no record of its own. */
+static __used __noinline __u64 gap_mid(void)
+{
+	return gap_inner_frame() + 1;
+}
+
+static __used __naked __noinline __u64 gap_outer_frame(void)
+{
+	asm volatile (
+"1:"	"call gap_mid;"			/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r6 = r0;"
+	PAD_RAN("%[ran]")
+	"r1 = r6;"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_GAP_OUTER), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+/* And one more uncovered frame between the outer pad and the boundary. */
+static __used __noinline __u64 gap_top(void)
+{
+	return gap_outer_frame() + 1;
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_GAP_INNER | RAN_GAP_OUTER)
+int entry_two_pads(void *ctx)
+{
+	return gap_top();
+}
+
+/*
+ * A precision chain crossing a resume: the outer frame's pad uses r6 as a
+ * variable stack offset, and the only way into that pad is the resume that
+ * ends the inner frame's pad, so backtracking goes from the pad through the
+ * inner frame and back to where r6 was bounded.
+ */
+static __used __naked __noinline __u64 prec_inner_frame(void)
+{
+	asm volatile (
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 prec_outer_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"r6 &= 0x7;"
+	"r0 = 0;"
+	"*(u64 *)(r10 - 8) = r0;"
+	"*(u64 *)(r10 - 16) = r0;"
+"1:"	"call prec_inner_frame;"	/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* pad: r6 as a variable stack offset */
+	"r2 = r10;"
+	"r2 += -16;"
+	"r2 += r6;"
+	"*(u8 *)(r2 + 0) = 1;"
+	PAD_RAN("%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_PREC_RESUME), __imm_addr(input), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_PREC_RESUME)
+__log_level(2)
+__msg("frame1: regs=r6 stack= before {{[0-9]+}}: (85) call bpf_unwind_resume")
+__msg("frame2: regs= stack= before {{[0-9]+}}: (85) call bpf_unwind#")
+__msg("frame1: regs=r6 stack= before {{[0-9]+}}: (57) r6 &= 7")
+int entry_prec_across_resume(void *ctx)
+{
+	return prec_outer_frame();
+}
+
+/* Unwinds every time, and no record covers it, so the path simply ends. */
+static __used __naked __noinline __u64 always_unwind(void)
+{
+	asm volatile (
+	"call bpf_unwind;"
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+/*
+ * A frame with no record of its own above one that always unwinds: nothing
+ * after the call is reachable, so dead code removal would leave it no exit
+ * and no epilogue for the unwind to send it to. One is kept, and since this
+ * frame ends in a jump rather than an exit, it is not the last instruction.
+ */
+static __used __naked __noinline __u64 no_exit_ja_mid(void)
+{
+	asm volatile (
+	"goto 2f;"
+"1:"	"r0 = 1;"
+	"exit;"
+"2:"	"call always_unwind;"
+	"goto 1b;"			/* the last insn, and not an exit */
+	::: __clobber_all);
+}
+
+static __used __naked __noinline __u64 no_exit_ja_outer_frame(void)
+{
+	asm volatile (
+"1:"	"call no_exit_ja_mid;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	PAD_RAN("%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_NO_EXIT_JA), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__ret_global(pads_ran, RAN_NO_EXIT_JA)
+int entry_no_exit_ja(void *ctx)
+{
+	return no_exit_ja_outer_frame();
+}
+
+/* gcc has no indirect calls, and only these JITs emit them */
+#if defined(__clang__) && \
+	(defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64))
+
+/*
+ * A region covering an indirect call: a record names a call by its return
+ * address, which a callx leaves like any other call.
+ */
+static __used __naked __noinline __u64 callx_region_frame(void)
+{
+	asm volatile (
+	"call bpf_preempt_disable;"
+	"r2 = %[always_unwind] ll;"
+"1:"	"callx r2;"			/* cleanup region */
+"2:"
+	"call bpf_preempt_enable;"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"call bpf_preempt_enable;"
+	PAD_RAN("%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_CALLX), __imm_addr(always_unwind),
+	  __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__ret_global(pads_ran, RAN_CALLX)
+int entry_callx_region(void *ctx)
+{
+	return callx_region_frame();
+}
+
+/*
+ * A subprog calling an unwinding one through a pointer it was handed: nothing
+ * after the call runs, but the frame still needs an exit for its epilogue.
+ */
+static __used __naked __noinline __u64 callx_arg_frame(void)
+{
+	asm volatile (
+	"callx r1;"			/* r1 is always_unwind */
+	"r0 = 1;"
+	"exit;"
+	::: __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__naked int entry_callx_arg(void)
+{
+	asm volatile (
+	"r1 = %[always_unwind] ll;"
+	"call callx_arg_frame;"
+	"exit;"
+	:
+	: __imm_addr(always_unwind)
+	: __clobber_all);
+}
+
+#endif /* __clang__ && (x86 || arm64) */
+
+/*
+ * The jump_into_pad shape with the branch dead, so the jump into the pad is
+ * walked only speculatively, after the unwind has marked the pad: a barrier
+ * rather than a refusal. Only a load without CAP_PERFMON walks it, hence the
+ * unprivileged run, and the branch is dead by range rather than by a
+ * constant, which const_fold would rewrite into a plain goto before any walk.
+ */
+static __used __naked __noinline __u64 dead_jump_into_pad_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"r6 &= 7;"
+	"if r6 > 7 goto 4f;"		/* never taken: walked speculatively */
+"1:"	"call always_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r7 = r0;"
+"4:"					/* ... and its second instruction */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: __imm_addr(input)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __caps_unpriv(CAP_BPF) __success_unpriv
+__xlated_unpriv("nospec")
+int entry_dead_jump_into_pad(void *ctx)
+{
+	return dead_jump_into_pad_frame();
+}
+
+/* Add one to the @w-bit global at @addr. */
+#define BUMP_GLOBAL(w, addr)					\
+	"r1 = " addr " ll;"					\
+	"r2 = *(u" w " *)(r1 + 0);"				\
+	"r2 += 1;"						\
+	"*(u" w " *)(r1 + 0) = r2;"
+
+/*
+ * A pad that touches a global of each width and sign a test tag can name,
+ * so that __set_global() and __ret_global() are exercised on all four.
+ */
+int tag_i = 0;
+unsigned int tag_ui = 0;
+long tag_l = 0;
+unsigned long tag_ul = 0;
+
+static __used __naked __noinline __u64 tag_types_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+"1:"	"call pc_unwinder;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	BUMP_GLOBAL("32", "%[tag_i]")
+	BUMP_GLOBAL("32", "%[tag_ui]")
+	BUMP_GLOBAL("64", "%[tag_l]")
+	BUMP_GLOBAL("64", "%[tag_ul]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: __imm_addr(input), __imm_addr(tag_i), __imm_addr(tag_ui),
+	  __imm_addr(tag_l), __imm_addr(tag_ul)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__set_global(input, 101)
+__set_global(tag_i, -23) __set_global(tag_ui, 0xfffffffe)
+__set_global(tag_l, -23) __set_global(tag_ul, 0xfffffffffffffffe)
+__ret_global(tag_i, -22) __ret_global(tag_ui, 0xffffffff)
+__ret_global(tag_l, -22) __ret_global(tag_ul, 0xffffffffffffffff)
+int entry_tag_types(void *ctx)
+{
+	return tag_types_frame();
+}
+
+struct {
+	__uint(type, BPF_MAP_TYPE_RINGBUF);
+	__uint(max_entries, 4096);
+} shape_ringbuf SEC(".maps");
+
+/*
+ * A frame holding a reference across a call an unwind comes out of. The record
+ * over the call is what lets it hold one: the pad releases it, where a frame
+ * with no record would be left for its epilogue still holding it.
+ */
+static __used __naked __noinline __u64 held_ref_frame(void)
+{
+	asm volatile (
+	"r1 = %[shape_ringbuf] ll;"
+	"r2 = 8;"
+	"r3 = 0;"
+	"call %[bpf_ringbuf_reserve];"
+	"if r0 == 0 goto 9f;"
+	"r6 = r0;"
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+"1:"	"call pc_unwinder;"		/* cleanup region */
+"2:"
+	"r1 = r6;"
+	"r2 = 0;"
+	"call %[bpf_ringbuf_discard];"
+	"goto 9f;"
+"3:"					/* landing pad: release and resume */
+	"r1 = r6;"
+	"r2 = 0;"
+	"call %[bpf_ringbuf_discard];"
+	PAD_RAN("%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+"9:"
+	"r0 = 0;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_HELD_REF), __imm(bpf_ringbuf_reserve),
+	  __imm(bpf_ringbuf_discard), __imm_addr(shape_ringbuf),
+	  __imm_addr(input), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_HELD_REF)
+int entry_held_ref(void *ctx)
+{
+	return held_ref_frame();
+}
+
+/*
+ * A pad reading a slot its frame's callee wrote before it unwound. The write
+ * is there when the pad runs, and the pad has to be verified that way, or
+ * the check below is taken as always failing and the bit is never set.
+ */
+static __used __naked __noinline __u64 slot_writer(void)
+{
+	asm volatile (
+	"r2 = 42;"
+	"*(u64 *)(r1 + 0) = r2;"	/* r1 is the caller's fp-8 */
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"if r1 < 101 goto 1f;"
+	"call bpf_unwind;"
+"1:"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm_addr(input)
+	: __clobber_all);
+}
+
+static __used __naked __noinline __u64 callee_write_frame(void)
+{
+	asm volatile (
+	"r1 = 0;"
+	"*(u64 *)(r10 - 8) = r1;"
+	"r1 = r10;"
+	"r1 += -8;"
+"1:"	"call slot_writer;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r1 = *(u64 *)(r10 - 8);"
+	"if r1 != 42 goto 9f;"
+	PAD_RAN("%[ran]")
+"9:"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_CALLEE_WRITE), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_CALLEE_WRITE)
+int entry_callee_write(void *ctx)
+{
+	return callee_write_frame();
+}
+
+/*
+ * A precision chain across an unwind: the pad uses a slot the callee wrote as
+ * a variable stack offset, so backtracking follows the slot from the pad back
+ * into the frame the unwind left.
+ */
+static __used __naked __noinline __u64 offset_writer(void)
+{
+	asm volatile (
+	"r2 = %[input] ll;"
+	"r3 = *(u64 *)(r2 + 0);"
+	"r3 &= 8;"
+	"*(u64 *)(r1 + 0) = r3;"	/* r1 is the caller's fp-24 */
+	"call bpf_unwind;"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm_addr(input)
+	: __clobber_all);
+}
+
+static __used __naked __noinline __u64 callee_offset_frame(void)
+{
+	asm volatile (
+	FILL_MAGIC_SLOTS
+	"r1 = 0;"
+	"*(u64 *)(r10 - 24) = r1;"
+	"r1 = r10;"
+	"r1 += -24;"
+"1:"	"call offset_writer;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r6 = *(u64 *)(r10 - 24);"
+	CHECK_VAR_SLOT("r6", "%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_CALLEE_OFFSET), __imm_addr(magic), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_CALLEE_OFFSET)
+__log_level(2)
+__msg("frame1: regs= stack=-24 before {{[0-9]+}}: (85) call bpf_unwind#")
+__msg("frame2: regs= stack= before {{[0-9]+}}: (7b) *(u64 *)(r1 +0) = r3")
+__msg("frame2: regs=r3 stack= before {{[0-9]+}}: (57) r3 &= 8")
+int entry_callee_offset(void *ctx)
+{
+	return callee_offset_frame();
+}
+
+/* The same through a global subprog. */
+__noinline int global_slot_writer(__u64 *p)
+{
+	if (!p)
+		return 0;
+	*p = 42;
+	if (input > 100)
+		bpf_unwind();
+	return 0;
+}
+
+static __used __naked __noinline __u64 global_write_frame(void)
+{
+	asm volatile (
+	"r1 = 0;"
+	"*(u64 *)(r10 - 8) = r1;"
+	"r1 = r10;"
+	"r1 += -8;"
+"1:"	"call global_slot_writer;"	/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r1 = *(u64 *)(r10 - 8);"
+	"if r1 != 42 goto 9f;"
+	PAD_RAN("%[ran]")
+"9:"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_GLOBAL_WRITE), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_GLOBAL_WRITE)
+int entry_global_write(void *ctx)
+{
+	return global_write_frame();
+}
+
+/*
+ * An unwind raised in a global subprog, called with no record over the call
+ * from a frame whose own caller has a pad. The global subprog is verified on
+ * its own, so the unwind is taken from the state its call returns in, and it
+ * has to go on to that pad.
+ */
+__noinline int global_unwinder(int x)
+{
+	if (x > 100)
+		bpf_unwind();
+	return 0;
+}
+
+static __used __naked __noinline __u64 through_global_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"call global_unwinder;"		/* no record */
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm_addr(input)
+	: __clobber_all);
+}
+
+static __used __naked __noinline __u64 over_global_frame(void)
+{
+	asm volatile (
+"1:"	"call through_global_frame;"	/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	PAD_RAN("%[ran]")
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_THROUGH_GLOBAL), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 101) __retval(0)
+__ret_global(pads_ran, RAN_THROUGH_GLOBAL)
+int entry_through_global(void *ctx)
+{
+	return over_global_frame();
+}
+
+/*
+ * A bpf_unwind() inside a loop, with no pad between it and the main program.
+ * The main program's frame returns from where the unwind left it, not from
+ * the loop in the subprog.
+ */
+static __used __naked __noinline __u64 loop_unwinder(void)
+{
+	asm volatile (
+	"r6 = 0;"
+"1:"
+	"r1 = %[input] ll;"
+	"r1 = *(u64 *)(r1 + 0);"
+	"if r1 != r6 goto 2f;"
+	"call bpf_unwind;"
+"2:"
+	"r6 += 1;"
+	"if r6 < 4 goto 1b;"
+	"r0 = 1;"
+	"exit;"
+	:
+	: __imm_addr(input)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __set_global(input, 2) __retval(0)
+int entry_unwind_in_loop(void *ctx)
+{
+	return loop_unwinder();
+}
+
+/* A frame's own pad, reached from its bpf_unwind(), finds r0 zero. */
+static __used __naked __noinline __u64 own_pad_r0_frame(void)
+{
+	asm volatile (
+"1:"	"call bpf_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 1;"
+	"exit;"
+"3:"					/* landing pad */
+	"if r0 != 0 goto 9f;"
+	PAD_RAN("%[ran]")
+"9:"
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: [ran]"i"(RAN_OWN_PAD_R0), __imm_addr(pads_ran)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __retval(0)
+__ret_global(pads_ran, RAN_OWN_PAD_R0)
+int entry_own_pad_r0(void *ctx)
+{
+	return own_pad_r0_frame();
+}
+
+/*
+ * A pad's second insn reached first by a speculative walk from outside the
+ * pad, and only then by the real one from the unwind. The speculative visit
+ * gets a barrier and must leave no mark the real one is then refused over.
+ * Only a load without CAP_PERFMON walks it.
+ */
+static __used __naked __noinline __u64 spec_first_pad_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"r6 &= 7;"
+	"if r6 > 3 goto 1f;"		/* both ways: the real path is pushed */
+	"r7 = r6;"
+	"if r7 > 7 goto 4f;"		/* never taken: walked speculatively */
+	"r0 = 0;"
+	"exit;"
+"1:"	"call always_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"r7 = r0;"
+"4:"					/* ... and its second instruction */
+	"call bpf_unwind_resume;"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: __imm_addr(input)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __caps_unpriv(CAP_BPF) __success_unpriv
+__xlated_unpriv("nospec")
+int entry_spec_first_pad(void *ctx)
+{
+	return spec_first_pad_frame();
+}
+
+/* A pad whose speculative walk reaches an exit: a barrier, not a refusal. */
+static __used __naked __noinline __u64 spec_exit_pad_frame(void)
+{
+	asm volatile (
+	"r1 = %[input] ll;"
+	"r6 = *(u64 *)(r1 + 0);"
+	"r6 &= 7;"
+"1:"	"call always_unwind;"		/* cleanup region */
+"2:"
+	"r0 = 0;"
+	"exit;"
+"3:"					/* landing pad */
+	"if r6 > 7 goto 4f;"		/* never taken: walked speculatively */
+	"call bpf_unwind_resume;"
+"4:"
+	"exit;"
+	CLEANUP_REC("1b", "2b", "3b")
+	:
+	: __imm_addr(input)
+	: __clobber_all);
+}
+
+SEC("?syscall")
+__success __caps_unpriv(CAP_BPF) __success_unpriv
+__xlated_unpriv("nospec")
+int entry_spec_exit_pad(void *ctx)
+{
+	return spec_exit_pad_frame();
+}
+
+/*
+ * An unwind that reaches the main program's exit in a program type whose
+ * return value is checked: the check backtracks r0 from where main returns,
+ * past a call it made before, to the insn that unwound.
+ */
+static __used __naked __noinline __u64 plain_frame(void)
+{
+	asm volatile (
+	"r0 = 0;"
+	"exit;"
+	::: __clobber_all);
+}
+
+SEC("?cgroup/skb")
+__success
+__naked int entry_unwind_to_checked_exit(void)
+{
+	asm volatile (
+	"call plain_frame;"
+	"call always_unwind;"
+	"r0 = 1;"
+	"exit;"
+	::: __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* [PATCH bpf-next v8 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton
  2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
                   ` (20 preceding siblings ...)
  2026-10-01 13:31 ` [PATCH bpf-next v8 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
@ 2026-10-01 13:32 ` Yonghong Song
  21 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-01 13:32 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

For a light skeleton, libbpf hands the records to bpf_gen__prog_load(),
which writes them into a blob and emits a loader program that issues
BPF_PROG_LOAD from inside the kernel. Nothing about that is shared with the
ordinary path: the attr is built field by field, and the kfunc names are
resolved by the loader program when it runs.

So build progs/exceptions_cleanup.c as a light skeleton too, and run
the case where foo3 unwinds through it, which runs every pad. Its table
has several records, all inside subprograms: a table that loses some, or
lays them out at the wrong stride, does not load, and one whose offsets
come out wrong runs the wrong pads, which pads_ran shows.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 tools/testing/selftests/bpf/Makefile.skel     |  2 +-
 .../bpf/prog_tests/exceptions_cleanup.c       | 31 +++++++++++++++++++
 2 files changed, 32 insertions(+), 1 deletion(-)

diff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel
index 2e22bb901bf3..76cd4561dd7a 100644
--- a/tools/testing/selftests/bpf/Makefile.skel
+++ b/tools/testing/selftests/bpf/Makefile.skel
@@ -39,7 +39,7 @@ LSKELS_SIGNED := fentry_test.c fexit_test.c atomics.c
 
 # Generate both light skeleton and libbpf skeleton for these
 LSKELS_EXTRA := test_ksyms_module.c test_ksyms_weak.c kfunc_call_test.c \
-	kfunc_call_test_subprog.c test_global_percpu_data.c
+	kfunc_call_test_subprog.c test_global_percpu_data.c exceptions_cleanup.c
 SKEL_BLACKLIST += $(LSKELS) $(LSKELS_SIGNED)
 
 test_static_linked.skel.h-deps := test_static_linked1.bpf.o test_static_linked2.bpf.o
diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
index c06ec10359b9..bb2bb8976fd9 100644
--- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
+++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
@@ -5,6 +5,7 @@
 #include "exceptions_cleanup.skel.h"
 #include "exceptions_cleanup_fail.skel.h"
 #include "exceptions_cleanup_shapes.skel.h"
+#include "exceptions_cleanup.lskel.h"
 
 /* foo3 unwound: every frame that has a pad ran it. */
 #define PADS_FOO3_UNWOUND \
@@ -37,6 +38,33 @@ static void run(struct exceptions_cleanup *skel, __u64 input, __u32 retval,
 	ASSERT_EQ(skel->bss->pads_ran, pads | RAN_BUMP, "pads_ran");
 }
 
+static void test_light_skeleton(void)
+{
+	struct exceptions_cleanup_lskel *skel;
+	__u64 ctx = 0;
+	int err;
+
+	LIBBPF_OPTS(bpf_test_run_opts, topts,
+		    .ctx_in = &ctx,
+		    .ctx_size_in = sizeof(ctx),
+	);
+
+	skel = exceptions_cleanup_lskel__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "light open_and_load"))
+		return;
+
+	/* foo3 unwinds, so every pad runs: the whole table has to arrive. */
+	skel->bss->input = 101;
+	err = bpf_prog_test_run_opts(skel->progs.entry.prog_fd, &topts);
+	if (!ASSERT_OK(err, "run"))
+		goto out;
+	ASSERT_EQ(topts.retval, 0, "retval");
+	ASSERT_EQ(skel->bss->pads_ran, PADS_FOO3_UNWOUND | RAN_BUMP,
+		  "pads_ran");
+out:
+	exceptions_cleanup_lskel__destroy(skel);
+}
+
 void test_exceptions_cleanup(void)
 {
 	char log[8192] = {};
@@ -82,6 +110,9 @@ void test_exceptions_cleanup(void)
 
 	exceptions_cleanup__destroy(skel);
 
+	if (test__start_subtest("light_skeleton"))
+		test_light_skeleton();
+
 	RUN_TESTS(exceptions_cleanup_fail);
 	RUN_TESTS(exceptions_cleanup_shapes);
 }
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts
  2026-10-01 13:31 ` [PATCH bpf-next v8 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
@ 2026-10-01 13:46   ` sashiko-bot
  2026-10-02 22:09     ` Yonghong Song
  0 siblings, 1 reply; 50+ messages in thread
From: sashiko-bot @ 2026-10-01 13:46 UTC (permalink / raw)
  To: Yonghong Song; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: ABI breakage due to reuse of padding in bpf_prog_load_opts
--

commit 71e2541e0b207ab646c97a6cfe861ce9bf16fe7f
Author: Yonghong Song <yonghong.song@linux.dev>

libbpf: Add cleanup_info to bpf_prog_load_opts

This commit introduces an exception cleanup table to the bpf_prog_load
options. It adds cleanup_info_cnt, cleanup_info, and cleanup_info_rec_size
to struct bpf_prog_load_opts and propagates them to the kernel in
bpf_prog_load.

> diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
> index b49822d212aed..b4031f62bdee5 100644
> --- a/tools/lib/bpf/bpf.c
> +++ b/tools/lib/bpf/bpf.c
[ ... ]
> @@ -370,6 +370,10 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
>  	attr.fd_array = ptr_to_u64(OPTS_GET(opts, fd_array, NULL));
>  	attr.fd_array_cnt = OPTS_GET(opts, fd_array_cnt, 0);
>  
> +	attr.cleanup_info = ptr_to_u64(OPTS_GET(opts, cleanup_info, NULL));
> +	attr.cleanup_info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
> +	attr.cleanup_info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
> +

[Severity: High]
Does this OPTS_GET macro read uninitialized padding from older applications?
When an application compiled against an older version of libbpf passes its
opts structure here, the sz field will be set to the old struct size.

If the old struct had padding at the end that the new structure replaces
with cleanup_info_cnt, OPTS_GET might think the field is present because
the sz value includes the padding, causing it to read stack garbage.

>  	if (log_level) {
>  		attr.log_buf = ptr_to_u64(log_buf);
>  		attr.log_size = log_size;
> diff --git a/tools/lib/bpf/bpf.h b/tools/lib/bpf/bpf.h
> index 826d9cc9ab65d..cbe56ddc8cf7d 100644
> --- a/tools/lib/bpf/bpf.h
> +++ b/tools/lib/bpf/bpf.h
[ ... ]
> @@ -128,9 +128,14 @@ struct bpf_prog_load_opts {
>  
>  	/* if set, provides the length of fd_array */
>  	__u32 fd_array_cnt;
> +
> +	/* exception cleanup table, from the .bpf_cleanup section */
> +	__u32 cleanup_info_cnt;

[Severity: High]
Does placing cleanup_info_cnt immediately after fd_array_cnt break the
size-based backward compatibility check?

Because fd_array_cnt is a 4-byte field, the struct bpf_prog_load_opts in
older applications likely contained a 4-byte padding hole at the end to
align with the 8-byte size_t type.

An older application would pass an opts structure where sz includes those 4
bytes of padding. Because OPTS_HAS evaluates if sz >=
offsetofend(struct bpf_prog_load_opts, cleanup_info_cnt), it will evaluate
to true. The library will then read the uninitialized padding left by the
older application as the cleanup_info_cnt value.

If the padding contains garbage, the kernel could reject the program load
or fail unexpectedly. Could this be resolved by placing cleanup_info_cnt
after the pointer, or by introducing explicit padding?

> +	const void *cleanup_info;
> +	__u32 cleanup_info_rec_size;
>  	size_t :0;
>  };

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001133006.1335369-1-yonghong.song@linux.dev?part=15

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 04/22] bpf: Add lookups for exception cleanup resumes and landing pads
  2026-10-01 13:30 ` [PATCH bpf-next v8 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
@ 2026-10-01 13:48   ` sashiko-bot
  2026-10-02 18:17     ` Yonghong Song
  0 siblings, 1 reply; 50+ messages in thread
From: sashiko-bot @ 2026-10-01 13:48 UTC (permalink / raw)
  To: Yonghong Song; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: failure to adjust cleanup_pad for BPF_PATCH_MOVE_TARGET allows arbitrary read/crash
--

commit 1a4d7ce9c87bb54c2feec4648d42fa1046ccb8b1
Author: Yonghong Song <yonghong.song@linux.dev>

bpf: Add lookups for exception cleanup resumes and landing pads

This commit adds lookups to exception.c to recognize calls to bpf_unwind()
and bpf_unwind_resume(). It also updates bpf_patch_insn_data(),
verifier_remove_insns(), and bpf_opt_remove_nops() to keep the landing pads
of call sites properly synchronized.

> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
> index 37cf130ebb57b..5b7fe4ba610b3 100644
> --- a/kernel/bpf/fixups.c
> +++ b/kernel/bpf/fixups.c
[ ... ]
> @@ -268,11 +268,18 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
>  			data[i].non_stack_access =
>  				data[off + cnt - 1].non_stack_access;
>  			data[off + cnt - 1].non_stack_access = false;
> +			data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
> +			data[off + cnt - 1].cleanup_pad = 0;
>  		} else if (bpf_is_mem_insn(insn + i)) {
>  			data[i].non_stack_access = true;
>  		}
>  	}
>  
> +	if (env->cleanup_info_cnt)
> +		for (i = 0; i < prog_len; i++)
> +			if (data[i].cleanup_pad > off + 1)
> +				data[i].cleanup_pad += cnt - 1;

[Severity: High]
Could this fail to adjust the cleanup_pad correctly when the mode is
BPF_PATCH_MOVE_TARGET in adjust_insn_aux_data()?

The condition checks data[i].cleanup_pad > off + 1, which works well for
BPF_PATCH_KEEP_TARGET, but it appears to miss BPF_PATCH_MOVE_TARGET where
the target should follow the original instruction to the end of the patched
sequence.

When a prologue is inserted at offset 0 (for example, during
bpf_convert_ctx_accesses), exception landing pads pointing to offset 0 are
not shifted. If a user sets an exception landing pad at offset 0, unwinding
would cause execution to jump to the newly inserted prologue instead of the
original instruction.

Since the prologue for program types like TC expects R1 to contain the CTX
pointer, and R1 is a caller-saved register containing whatever was placed
there before calling bpf_unwind(), could dereferencing this
attacker-controlled R1 during an exception unwind lead to an arbitrary memory
read or a kernel panic?

Should this check be dependent on the mode, such as:
data[i].cleanup_pad > (mode == BPF_PATCH_MOVE_TARGET ? off : off + 1)
to properly account for both modes?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001133006.1335369-1-yonghong.song@linux.dev?part=4

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 12/22] bpf, x86: Dispatch exception cleanup pads at run time
  2026-10-01 13:31 ` [PATCH bpf-next v8 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
@ 2026-10-01 13:49   ` sashiko-bot
  2026-10-02 21:54     ` Yonghong Song
  0 siblings, 1 reply; 50+ messages in thread
From: sashiko-bot @ 2026-10-01 13:49 UTC (permalink / raw)
  To: Yonghong Song; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: verifier bypass via BPF exception unwind abort on traced frames
--

commit 8f24ec5b2338fd8477df1e09b72647e4de2ae3d4
Author: Yonghong Song <yonghong.song@linux.dev>

bpf, x86: Dispatch exception cleanup pads at run time

This commit introduces runtime dispatch of exception cleanup pads using the
ORC unwinder. It implements arch_bpf_stack_walk_ra() to traverse the stack
and rewrite return addresses to point to their corresponding cleanup pads.

> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
> index 6c7a0578760e7..544e8fd759ad4 100644
> --- a/arch/x86/net/bpf_jit_comp.c
> +++ b/arch/x86/net/bpf_jit_comp.c

[ ... ]

> @@ -4614,6 +4628,32 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
>  #endif
>  }
>  
> +void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
> +			    void *cookie)
> +{
> +#if defined(CONFIG_UNWINDER_ORC)
> +	struct unwind_state state;
> +	unsigned long addr, *ra;
> +
> +	for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
> +	     unwind_next_frame(&state)) {
> +		addr = unwind_get_return_address(&state);
> +		ra = unwind_get_return_address_ptr(&state);
> +		if (!addr || !ra)
> +			break;
> +		/*
> +		 * A traced return: the slot holds a function graph or kretprobe
> +		 * trampoline, not @addr, so it cannot be rewritten. Stopping
> +		 * leaves BPF frames returning to unverified paths, so warn.
> +		 */
> +		if (WARN_ON_ONCE(READ_ONCE_NOCHECK(*ra) != addr))
> +			break;

[Severity: High]
The comment above notes that this leaves BPF frames returning to unverified
paths. Since the verifier assumes bpf_throw() does not return and prunes
instructions following it, won't aborting the stack walk here allow the
execution of those unverified instructions?

Could a user with tracing permissions intentionally attach a tracer (like
kretprobe or fgraph) to trigger this break, bypassing the verifier and
forcing the execution of unverified BPF code?

> +		if (!consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
> +			break;
> +	}
> +#endif
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001133006.1335369-1-yonghong.song@linux.dev?part=12

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier
  2026-10-01 13:30 ` [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier Yonghong Song
@ 2026-10-01 13:50   ` sashiko-bot
  2026-10-02 19:31     ` Yonghong Song
  2026-10-01 14:31   ` bot+bpf-ci
  2026-10-03 12:23   ` Alexei Starovoitov
  2 siblings, 1 reply; 50+ messages in thread
From: sashiko-bot @ 2026-10-01 13:50 UTC (permalink / raw)
  To: Yonghong Song; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Critical] bpf: verifier skips unwind paths when global subprog might throw
- [Critical] bpf: missing restriction for bpf_unwind in callbacks leads to verifier divergence
--

commit d554473b8011fe90d6322a255220e23a4d7500d5
Author: Yonghong Song <yonghong.song@linux.dev>
Subject: bpf: Follow an unwind to its landing pad in the verifier

This patch updates the BPF verifier to follow unwind operations to their
landing pads. It handles how an unwind rewrites return addresses of frames
and evaluates the proper precision backtracking.

> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
[ ... ]
> @@ -19126,6 +19145,127 @@ enum {
>  	INSN_IDX_UPDATED = 2,
>  };
>  
> +/*
> + * The current frame is leaving through an unwind. Its caller's saved return
> + * address now points at the pad covering the call, or, with none, at the
> + * caller's epilogue, and so on down. Follow that from the state the frame
> + * leaves in -- anything it wrote into its callers' stacks included -- to the
> + * first pad, or to the main program's frame returning.
> + */
> +static int unwind_frames(struct bpf_verifier_env *env, bool *do_print_state)
> +{
> +	struct bpf_verifier_state *state = env->cur_state;
> +	u32 frameno = state->curframe;
> +	struct bpf_func_state *callee, *caller;
> +	int err, pad;
> +
> +	while (state->curframe) {
> +		callee = cur_func(env);
> +		caller = state->frame[state->curframe - 1];
> +		pad = bpf_exc_pad_of_call(env, callee->callsite);
> +		/* The caller is at its call now, not at this frame's insn. */
> +		state->insn_idx = callee->callsite;
> +		account_processed_insns(env, callee, caller);
> +		free_func_state(callee);
> +		state->frame[state->curframe--] = NULL;
> +		invalidate_outgoing_stack_args(env, caller);
> +		if (pad < 0)
> +			continue;

[Severity: Critical]
If a callback subprogram (such as one passed to bpf_loop()) calls
bpf_unwind(), does this loop incorrectly pop frames in the verifier state past
the C helper?

It appears bpf_unwind() and bpf_unwind_resume() lack the callback restriction
that bpf_throw() has in do_check_subprogs().

At runtime, the exception unwinder stops unwinding when it encounters the
non-BPF C helper frame and returns. The callback returns 0, the C helper
resumes, and it returns normally to the BPF program.

However, if the verifier simulates an unwind through the C helper to the
caller's landing pad, it marks the normal execution path as dead code
and replaces it with a ja -1.

Could this divergence between verifier state and runtime execution cause an
infinite loop and hang the kernel?

[ ... ]
> @@ -19380,13 +19520,32 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>  					return -EINVAL;
>  				}
>  			}
> +			if (bpf_is_unwind_kfunc(insn) || bpf_is_unwind_resume_kfunc(insn)) {
> +				err = check_kfunc_allowed_only(env, insn, env->insn_idx);
> +				if (err)
> +					return err;
> +				if (bpf_is_unwind_kfunc(insn))
> +					return process_bpf_unwind(env, &env->insn_idx,
> +								  do_print_state);
> +				/*
> +				 * The fixups lower this to 'r0 = 0; exit', and
> +				 * the unwind goes on below this frame.
> +				 */
> +				return unwind_frames(env, do_print_state);
> +			}
>  			mark_reg_scratched(env, BPF_REG_0);
>  			if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
>  				cur_func(env)->no_stack_arg_load = true;
>  			if (bpf_is_callx(insn))
>  				return check_func_callx(env, insn, &env->insn_idx);
> -			if (insn->src_reg == BPF_PSEUDO_CALL)
> -				return check_func_call(env, insn, &env->insn_idx);
> +			if (insn->src_reg == BPF_PSEUDO_CALL) {
> +				int call_idx = env->insn_idx;
> +
> +				err = check_func_call(env, insn, &env->insn_idx);
> +				if (err)
> +					return err;
> +				return unwind_out_of_global_call(env, call_idx, do_print_state);
> +			}

[Severity: Critical]
If check_func_call() processes a global subprogram that is marked with both
might_throw and might_unwind, it evaluates might_throw, calls
process_bpf_exit_full(), and returns PROCESS_BPF_EXIT.

Does this cause do_check_insn() to return early here with err = 1, completely
bypassing the unwind_out_of_global_call() logic?

If bypassed, the landing pad associated with the call is never verified and
is replaced with a ja -1 (dead code). At runtime, if the global subprogram
calls bpf_unwind(), the exception unwinder routes execution to this unverified
landing pad. 

Can this result in hitting the ja -1 infinite loop and hanging the kernel?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001133006.1335369-1-yonghong.song@linux.dev?part=7

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 10/22] bpf: Do not use a private stack for a program that can unwind
  2026-10-01 13:30 ` [PATCH bpf-next v8 10/22] bpf: Do not use a private stack for a program that can unwind Yonghong Song
@ 2026-10-01 13:53   ` sashiko-bot
  2026-10-02 21:38     ` Yonghong Song
  0 siblings, 1 reply; 50+ messages in thread
From: sashiko-bot @ 2026-10-01 13:53 UTC (permalink / raw)
  To: Yonghong Song; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: callee-saved register corruption via bpf_throw without cleanup
--

commit 9d2c0e8faeb024a4647980c48ca2caa933b90b1c
Author: Yonghong Song <yonghong.song@linux.dev>

bpf: Do not use a private stack for a program that can unwind

This commit disables private stacks for BPF programs that can unwind
because an unwind skips the pop instruction corresponding to the push r9
emitted around calls. This forces NO_PRIV_STACK in check_max_stack_depth()
for programs with cleanup records or that may unwind.

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index f3ed68960d70d..488ceb9dae1b6 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
> @@ -5769,6 +5769,17 @@ static int check_max_stack_depth(struct bpf_verifier_env *env)
>  		}
>  	}
>  
> +	/*
> +	 * A private stack keeps its frame pointer in %r9 on x86-64, restored
> +	 * by a pop after the call that an unwind skips. A frame resumed at a
> +	 * pad then addresses its stack through a stale pointer, and a frame
> +	 * sent to its epilogue instead pops its callee-saved registers one
> +	 * slot off. Refuse a private stack for any program that can unwind,
> +	 * on every arch for now.
> +	 */
> +	if (env->cleanup_info_cnt || bpf_prog_may_unwind(env))
> +		priv_stack_mode = NO_PRIV_STACK;

[Severity: High]
Does this logic miss programs that call bpf_throw() without any cleanup
records?

If a privileged program uses private stacks and calls bpf_throw() without
acquiring any resources requiring cleanup, cleanup_info_cnt will be 0.

Looking at bpf_prog_may_unwind() in kernel/bpf/exception.c, it only checks
might_unwind:

kernel/bpf/exception.c:bpf_prog_may_unwind() {
	u32 i;

	for (i = 0; i < env->subprog_cnt; i++)
		if (env->subprog_info[i].might_unwind)
			return true;
	return false;
}

Since bpf_throw() only sets might_throw, not might_unwind, it appears
bpf_prog_may_unwind() will return false.

Could this bypass the NO_PRIV_STACK enforcement and lead to the exception
callback epilogue popping the misaligned stack (shifted by the skipped
push r9) into the kernel's callee-saved registers?

> +
>  	if (priv_stack_mode == PRIV_STACK_UNKNOWN)
>  		priv_stack_mode = bpf_enable_priv_stack(env->prog);
>

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001133006.1335369-1-yonghong.song@linux.dev?part=10

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 05/22] bpf: Prepare for an exception cleanup table before the CFG walk
  2026-10-01 13:30 ` [PATCH bpf-next v8 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
@ 2026-10-01 14:31   ` bot+bpf-ci
  2026-10-02 19:06     ` Yonghong Song
  0 siblings, 1 reply; 50+ messages in thread
From: bot+bpf-ci @ 2026-10-01 14:31 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 3053 bytes --]

> bpf: Prepare for an exception cleanup table before the CFG walk

TL;DR: bpf_exc_check_prog() rescans the whole program for bpf_throw, and a
later patch calls it at every bpf_unwind() visit, so verification time may
balloon on large programs; suggests caching the result.

> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
> index 3ea1bff5cc90..e12cdb12cde3 100644
> --- a/kernel/bpf/exception.c
> +++ b/kernel/bpf/exception.c
> @@ -141,6 +141,85 @@ int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,

[ ... ]

> +static int reject_throw(struct bpf_verifier_env *env)
> +{
> +	u32 i;
> +
> +	for (i = 0; i < env->prog->len; i++) {
> +		if (!bpf_is_throw_kfunc(&env->prog->insnsi[i]))
> +			continue;
> +		verbose(env,
> +			"exception cleanup cannot be combined with bpf_throw at insn %u\n",
> +			i);
> +		return -EINVAL;
> +	}
> +	return 0;
> +}

[ ... ]

> +int bpf_exc_check_prog(struct bpf_verifier_env *env)
> +{
> +	int err;
> +

[ ... ]

> +	err = reject_throw(env);
> +	if (err)
> +		return err;
> +	env->prog->jit_required = 1;
> +	return 0;
> +}

Can this scan make verification time grow with the number of bpf_unwind()
instructions the verifier visits?

The commit message says the checks in bpf_exc_check_prog() are also run
"at every bpf_unwind()", and a later patch in the series (5de314709a62
"bpf: Dispatch cleanup pads by rewriting return addresses") does that from
process_bpf_unwind():

    static int process_bpf_unwind(struct bpf_verifier_env *env, ...)
    {
        ...
        err = bpf_exc_check_prog(env);

Each of those calls reaches reject_throw(), which walks all env->prog->len
instructions looking for bpf_throw.

None of the things bpf_exc_check_prog() tests can change once do_check()
is running: the offload flag, jit_requested, exception_callback_subprog and
whether a bpf_throw insn exists are all fixed before the main pass.  Only
env->ops can change, and only for struct_ops, in check_attach_btf_id()
before do_check().

So the scan gives the same answer each time, but every visit costs
O(prog->len) instead of O(1).  That means the BPF_COMPLEXITY_LIMIT_INSNS
budget no longer bounds verifier cost linearly; the cost becomes roughly
visits * prog_len.

For example, a program with about 500K reachable insns plus a bounded loop
of the form "if (x == i) bpf_unwind();" could reach bpf_unwind() tens of
thousands of times within the 1M processed insn budget.  Each visit would
rescan 500K insns, which is around 10^10 to 10^11 extra iterations, or
tens of seconds of extra verifier CPU for a single load.

Would it make sense to run reject_throw() once and cache the result in
env?  Alternatively, the bpf_unwind() path could re-check only the
gen_epilogue condition, which is the one that becomes known after
bpf_exc_prepare().


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36872142096

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with
  2026-10-01 13:30 ` [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
@ 2026-10-01 14:31   ` bot+bpf-ci
  2026-10-02 21:10     ` Yonghong Song
  2026-10-03 12:25   ` Alexei Starovoitov
  1 sibling, 1 reply; 50+ messages in thread
From: bot+bpf-ci @ 2026-10-01 14:31 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 3387 bytes --]

> bpf: Require an unwind to leave a frame holding what it entered with

TL;DR: call_may_unwind() treats every callx as possibly unwinding, so the
verifier may wrongly reject valid programs that hold a reference or spin lock
across a callx; the exact target is available from the register.

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index c7a350be538e..a7b25ab04051 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c

[ ... ]

> @@ -19244,6 +19245,32 @@ static int unwind_out_of_global_call(struct bpf_verifier_env *env, int call_idx,
>  	return INSN_IDX_UPDATED;
>  }
>
> +/* Can an unwind come back out of this call? */
> +static bool call_may_unwind(struct bpf_verifier_env *env, const struct bpf_insn *insn,
> +			    int insn_idx)
> +{
> +	int subprog;
> +
> +	/* Which subprog a callx lands in is not known here, so any may be it. */
> +	if (bpf_is_callx(insn))
> +		return bpf_prog_may_unwind(env);

Is it true that the target of a callx is not known here?

call_may_unwind() is reached from check_unwind_through_call() in
do_check_insn(), just before check_func_callx().  check_func_callx() takes
the exact target from the same register state a few lines later:

	reg = reg_state(env, insn->dst_reg);
	if (reg->type != PTR_TO_FUNC) ...
	/* PTR_TO_FUNC is a pointer to a static subprog */
	subprog = reg->subprogno;

and check_func_ptr_read() pushes one state per possible table target, each
with its own subprogno.  So each verified callx has exactly one target at
this point.  Only the CFG pass has to be conservative, and
mark_callx_might_unwind() already marks any target that itself contains a
callx.

With the current code, an uncovered callx is refused by
bpf_exc_check_frame_balance() whenever the calling frame has changed its
RCU, preempt, IRQ, spin lock or reference state since entry, as long as some
unrelated subprog in the program is marked might_unwind.  For example:

	foo():
	    t = bpf_task_acquire(...);  /* id > entry_id_gen */
	    r1 = vtable[i];             /* PTR_TO_FUNC -> bar, no unwind in bar */
	    callx r1                    /* no cleanup record covers it */
	    bpf_task_release(t);

	baz():
	    bpf_unwind();               /* makes bpf_prog_may_unwind() true */

Before this patch this program was accepted.  Now it is refused with "an
unwind through this call keeps the reference id=N the frame acquired", even
though no unwind can pass through that call.  callx is also explicitly
allowed under a spin lock, so a frame that holds a bpf_spin_lock across a
callx is refused as well with "an unwind through this call does not leave
the frame's bpf_spin_lock state as it found it".

Could this look at the register instead, something like:

	reg = &cur_regs(env)[insn->dst_reg];
	if (reg->type == PTR_TO_FUNC)
		return env->subprog_info[reg->subprogno].might_unwind;

and leave it to check_func_callx() to refuse any other register type, as it
does today?

> +	if (insn->src_reg != BPF_PSEUDO_CALL)
> +		return false;
> +	subprog = bpf_find_subprog(env, insn_idx + insn->imm + 1);
> +	return subprog >= 0 && env->subprog_info[subprog].might_unwind;
> +}
> +

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36872142096

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses
  2026-10-01 13:31 ` [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
@ 2026-10-01 14:31   ` bot+bpf-ci
  2026-10-02 21:48     ` Yonghong Song
  2026-10-03 12:26   ` Alexei Starovoitov
  1 sibling, 1 reply; 50+ messages in thread
From: bot+bpf-ci @ 2026-10-01 14:31 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 2674 bytes --]

> bpf: Dispatch cleanup pads by rewriting return addresses

TL;DR: Design question, not a bug: bpf_exc_attach_main_prog() in
bpf_prog_jit_compile() builds a table for single-function programs that seems
never to be read, so it and its error paths could possibly be dropped.

> diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
> index d813fdde29e3..60905643cb9c 100644
> --- a/kernel/bpf/core.c
> +++ b/kernel/bpf/core.c

[ ... ]

> @@ -2632,9 +2633,14 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
>  {
>  #ifdef CONFIG_BPF_JIT
>  	struct bpf_prog *orig_prog;
> +	int ret;
>  
> -	if (!bpf_prog_need_blind(prog))
> +	if (!bpf_prog_need_blind(prog)) {
> +		ret = bpf_exc_attach_main_prog(env, prog);
> +		if (ret)
> +			return ERR_PTR(ret);
>  		return bpf_int_jit_compile(env, prog);
> +	}
>  
>  	orig_prog = prog;
>  	prog = bpf_jit_blind_constants(env, prog);
> @@ -2648,6 +2654,12 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
>  		goto out_restore;
>  	}
>  
> +	ret = bpf_exc_attach_main_prog(env, prog);
> +	if (ret) {
> +		bpf_jit_prog_release_other(orig_prog, prog);
> +		return ERR_PTR(ret);
> +	}
> +
>  	prog = bpf_int_jit_compile(env, prog);
>  	if (prog->jited) {
>  		bpf_jit_prog_release_other(prog, orig_prog);

this isn't a bug, but is the table that bpf_exc_attach_main_prog() builds
for a single-function program read anywhere?

It only attaches a table when there are no subprograms:

kernel/bpf/fixups.c:bpf_exc_attach_main_prog() {
    if (!env || env->subprog_cnt > 1)
        return 0;
    return exc_info_for_subprog(env, prog, 0, prog->len);
}

The only reader of the table that I can find is bpf_unwind_rewrite(), and it
returns at the first BPF frame before it does the pad lookup:

kernel/bpf/helpers.c:bpf_unwind_rewrite() {
    ...
    if (ctx->cnt == 1)
        return bpf_is_subprog(prog);

    rec = bpf_exc_pad_for_ip(prog, ip);
    ...
}

With no subprograms, the main frame is the only BPF frame, so it is always
the one with ctx->cnt == 1 and bpf_exc_pad_for_ip() is not reached for it.
The bpf_unwind() calls in such a program are already sent to their pad by
the jump that bpf_exc_patch_unwind_calls() inserts.

If that is right, could bpf_exc_attach_main_prog() and its two call sites in
bpf_prog_jit_compile(), with the new 'return ERR_PTR(ret)' paths, be dropped
so that bpf_prog_jit_compile() stays as it was?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36872142096

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier
  2026-10-01 13:30 ` [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier Yonghong Song
  2026-10-01 13:50   ` sashiko-bot
@ 2026-10-01 14:31   ` bot+bpf-ci
  2026-10-02 20:49     ` Yonghong Song
  2026-10-03 12:23   ` Alexei Starovoitov
  2 siblings, 1 reply; 50+ messages in thread
From: bot+bpf-ci @ 2026-10-01 14:31 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 6056 bytes --]

> bpf: Follow an unwind to its landing pad in the verifier

TL;DR:
- Global call that unwinds without a landing pad may hit a verifier bug
  warning in backtrack_insn() on fentry/cgroup_skb programs
- bpf_unwind() skips no_stack_arg_load, so a stale R9 arg-6 load may pass
  verification

> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
> index 0e38b9575328..f5504334df90 100644
> --- a/kernel/bpf/backtrack.c
> +++ b/kernel/bpf/backtrack.c

[ ... ]

> @@ -434,6 +458,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
>  					return -EFAULT;
>  			}
>
> +			if (bpf_exc_pad_of_call(env, idx) == subseq_idx) {
> +				/*
> +				 * We came from the landing pad of a call to a
> +				 * global subprog, branched to from the state
> +				 * the call returns in: as on its return, no
> +				 * frame was entered here. The call clobbered
> +				 * r0-r5; r6-r9 and the stack are the caller's
> +				 * own and keep going back from here.
> +				 */
> +				bt_clear_reg(bt, BPF_REG_0);
> +				if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
> +					verifier_bug(env, "landing pad unexpected regs %x",
> +						     bt_reg_mask(bt));
> +					return -EFAULT;
> +				}
> +				return 0;
> +			}
> +
>  			/* callx calls static subprogs only */
>  			if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
>  				/* check that jump history doesn't have any

Can the existing idx + 1 != subseq_idx check in this global branch now fire
for a global call that unwinds with no landing pad?

    verifier_bug_if(idx + 1 != subseq_idx, env,
                    "extra insn from subprog");

In that case unwind_out_of_global_call() takes the pad < 0 path into
unwind_frames(), and when unwind_frames() reaches the main frame it pushes a
history entry with no flags:

    env->prev_insn_idx = env->insn_idx;
    env->insn_idx = state->insn_idx;
    err = bpf_push_jmp_history(env, state, 0, 0, 0, NULL, 0);
    ...
    return process_bpf_exit_full(env, do_print_state, false);

The entry's prev_idx is the global call insn, so INSN_F_UNWIND is not set.
For program types where return_retval_range() returns true (fentry/fexit,
cgroup_skb, struct_ops), process_bpf_exit_full() -> check_return_code() calls
mark_chain_precision(env, BPF_REG_0).

The backtrack then reaches backtrack_insn() with idx at the global call and
subseq_idx at either the same insn (call in main) or the call site in main
(static frames popped first).  The INSN_F_UNWIND branch does not match, the
bpf_exc_pad_of_call() == subseq_idx branch does not match because no pad
covers the call, and the walk lands in the global branch above, where
idx + 1 != subseq_idx is true.

The result is a WARN_ONCE() (a panic with panic_on_warn) and a "verifier bug"
message in the log, although the program is still accepted.  An example is an
fentry or cgroup_skb program calling a global subprog that calls
bpf_unwind() with no record covering the call, or main -> static A -> global
G with no records.

The selftests in the series all use SEC("syscall"), where
return_retval_range() returns false and r0 precision is never requested at
exit, so they would not reach this.  Nothing later in the series touches
backtrack.c.

Should the main-exit history entry get its own flag, or should this global
branch handle that entry before applying the idx + 1 check?

[ ... ]

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 80034429fdd0..c7a350be538e 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c

[ ... ]

> @@ -19380,13 +19520,32 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>  					return -EINVAL;
>  				}
>  			}
> +			if (bpf_is_unwind_kfunc(insn) || bpf_is_unwind_resume_kfunc(insn)) {
> +				err = check_kfunc_allowed_only(env, insn, env->insn_idx);
> +				if (err)
> +					return err;
> +				if (bpf_is_unwind_kfunc(insn))
> +					return process_bpf_unwind(env, &env->insn_idx,
> +								  do_print_state);
> +				/*
> +				 * The fixups lower this to 'r0 = 0; exit', and
> +				 * the unwind goes on below this frame.
> +				 */
> +				return unwind_frames(env, do_print_state);
> +			}
>  			mark_reg_scratched(env, BPF_REG_0);
>  			if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
>  				cur_func(env)->no_stack_arg_load = true;

Does this early return skip the no_stack_arg_load assignment that every
other call insn gets?

For bpf_unwind() with a pad in the same frame, process_bpf_unwind() goes on
at the pad in the frame that made the call, with no_stack_arg_load still
false:

    clear_caller_saved_regs(env, frame->regs);
    mark_reg_unknown(env, frame->regs, BPF_REG_0);
    *insn_idx = pad;
    return INSN_IDX_UPDATED;

On x86, do_jit() turns a load from BPF_REG_PARAMS at offset 8 (incoming
arg 6) into a read of R9, which is caller-saved and clobbered by any call.
check_stack_arg_read() relies on no_stack_arg_load to reject such a load
after a call.

So in a subprog with more than 5 args, would this sequence pass
verification?

    call bpf_unwind        // record covers it, pad P in this frame
  P:
    r1 = *(u64 *)(r11 + 8) // incoming arg 6
    r2 = *(u64 *)(r1 + 0)

check_stack_arg_read() gives r1 the type of caller->stack_arg_regs[0], which
can be any spilled type, including a pointer such as PTR_TO_MAP_VALUE, but
at run time r1 holds whatever bpf_unwind() left in R9.

Would it work to move the mark_reg_scratched() and no_stack_arg_load lines
above the unwind branch, or to set cur_func(env)->no_stack_arg_load in
process_bpf_unwind() before going on at the pad?

The unwind_out_of_global_call() pad path looks unaffected, since the
PSEUDO_CALL path sets the flag before check_func_call().  do_check_insn()
has the same ordering at the end of the series, where both kfuncs become
callable.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36872142096

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 04/22] bpf: Add lookups for exception cleanup resumes and landing pads
  2026-10-01 13:48   ` sashiko-bot
@ 2026-10-02 18:17     ` Yonghong Song
  0 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-02 18:17 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf



On 10/1/26 6:48 AM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf: failure to adjust cleanup_pad for BPF_PATCH_MOVE_TARGET allows arbitrary read/crash
> --
>
> commit 1a4d7ce9c87bb54c2feec4648d42fa1046ccb8b1
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Add lookups for exception cleanup resumes and landing pads
>
> This commit adds lookups to exception.c to recognize calls to bpf_unwind()
> and bpf_unwind_resume(). It also updates bpf_patch_insn_data(),
> verifier_remove_insns(), and bpf_opt_remove_nops() to keep the landing pads
> of call sites properly synchronized.
>
>> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
>> index 37cf130ebb57b..5b7fe4ba610b3 100644
>> --- a/kernel/bpf/fixups.c
>> +++ b/kernel/bpf/fixups.c
> [ ... ]
>> @@ -268,11 +268,18 @@ static void adjust_insn_aux_data(struct bpf_verifier_env *env,
>>   			data[i].non_stack_access =
>>   				data[off + cnt - 1].non_stack_access;
>>   			data[off + cnt - 1].non_stack_access = false;
>> +			data[i].cleanup_pad = data[off + cnt - 1].cleanup_pad;
>> +			data[off + cnt - 1].cleanup_pad = 0;
>>   		} else if (bpf_is_mem_insn(insn + i)) {
>>   			data[i].non_stack_access = true;
>>   		}
>>   	}
>>   
>> +	if (env->cleanup_info_cnt)
>> +		for (i = 0; i < prog_len; i++)
>> +			if (data[i].cleanup_pad > off + 1)
>> +				data[i].cleanup_pad += cnt - 1;
> [Severity: High]
> Could this fail to adjust the cleanup_pad correctly when the mode is
> BPF_PATCH_MOVE_TARGET in adjust_insn_aux_data()?
>
> The condition checks data[i].cleanup_pad > off + 1, which works well for
> BPF_PATCH_KEEP_TARGET, but it appears to miss BPF_PATCH_MOVE_TARGET where
> the target should follow the original instruction to the end of the patched
> sequence.
>
> When a prologue is inserted at offset 0 (for example, during
> bpf_convert_ctx_accesses), exception landing pads pointing to offset 0 are
> not shifted. If a user sets an exception landing pad at offset 0, unwinding
> would cause execution to jump to the newly inserted prologue instead of the
> original instruction.
>
> Since the prologue for program types like TC expects R1 to contain the CTX
> pointer, and R1 is a caller-saved register containing whatever was placed
> there before calling bpf_unwind(), could dereferencing this
> attacker-controlled R1 during an exception unwind lead to an arbitrary memory
> read or a kernel panic?
>
> Should this check be dependent on the mode, such as:
> data[i].cleanup_pad > (mode == BPF_PATCH_MOVE_TARGET ? off : off + 1)
> to properly account for both modes?
>
This patch should be okay. The following is what I replied in v8 for this patch:

Okay, commit "bpf: Retarget indirect jump targets across prologue prepends"
introduced BPF_PATCH_MOVE_TARGET and merged last Friday. It solved three
cases for ops->gen_epilogue, ops->gen_prologue || env->seen_direct_write, and
stack slots for subprogs. ops->gen_epilogue has been rejected in patch 5.
We cannot allow ops->gen_epilogue since it may silently exit.
For other cases in "bpf: Retarget indirect jump targets across prologue prepends",
The above commit should already handle this.

For the other two, cleanup_pad == off + 1 cannot happen, because MOVE only
patches an entry insn and a landing pad cannot start at insn 0: the entry is
always walked outside a pad first, so reaching it again from an unwind
fails bpf_exc_check_insn() with "insn %u runs both inside and outside a
landing pad". Pads after off are shifted by the existing `> off + 1`, and a
covered call at off moves with its insn_aux_data.

So I think this patch should be okay.


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 05/22] bpf: Prepare for an exception cleanup table before the CFG walk
  2026-10-01 14:31   ` bot+bpf-ci
@ 2026-10-02 19:06     ` Yonghong Song
  0 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-02 19:06 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 10/1/26 7:31 AM, bot+bpf-ci@kernel.org wrote:
>> bpf: Prepare for an exception cleanup table before the CFG walk
> TL;DR: bpf_exc_check_prog() rescans the whole program for bpf_throw, and a
> later patch calls it at every bpf_unwind() visit, so verification time may
> balloon on large programs; suggests caching the result.
>
>> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
>> index 3ea1bff5cc90..e12cdb12cde3 100644
>> --- a/kernel/bpf/exception.c
>> +++ b/kernel/bpf/exception.c
>> @@ -141,6 +141,85 @@ int bpf_exc_check_info(struct bpf_verifier_env *env, const union bpf_attr *attr,
> [ ... ]
>
>> +static int reject_throw(struct bpf_verifier_env *env)
>> +{
>> +	u32 i;
>> +
>> +	for (i = 0; i < env->prog->len; i++) {
>> +		if (!bpf_is_throw_kfunc(&env->prog->insnsi[i]))
>> +			continue;
>> +		verbose(env,
>> +			"exception cleanup cannot be combined with bpf_throw at insn %u\n",
>> +			i);
>> +		return -EINVAL;
>> +	}
>> +	return 0;
>> +}
> [ ... ]
>
>> +int bpf_exc_check_prog(struct bpf_verifier_env *env)
>> +{
>> +	int err;
>> +
> [ ... ]
>
>> +	err = reject_throw(env);
>> +	if (err)
>> +		return err;
>> +	env->prog->jit_required = 1;
>> +	return 0;
>> +}
> Can this scan make verification time grow with the number of bpf_unwind()
> instructions the verifier visits?
>
> The commit message says the checks in bpf_exc_check_prog() are also run
> "at every bpf_unwind()", and a later patch in the series (5de314709a62
> "bpf: Dispatch cleanup pads by rewriting return addresses") does that from
> process_bpf_unwind():
>
>      static int process_bpf_unwind(struct bpf_verifier_env *env, ...)
>      {
>          ...
>          err = bpf_exc_check_prog(env);
>
> Each of those calls reaches reject_throw(), which walks all env->prog->len
> instructions looking for bpf_throw.
>
> None of the things bpf_exc_check_prog() tests can change once do_check()
> is running: the offload flag, jit_requested, exception_callback_subprog and
> whether a bpf_throw insn exists are all fixed before the main pass.  Only
> env->ops can change, and only for struct_ops, in check_attach_btf_id()
> before do_check().
>
> So the scan gives the same answer each time, but every visit costs
> O(prog->len) instead of O(1).  That means the BPF_COMPLEXITY_LIMIT_INSNS
> budget no longer bounds verifier cost linearly; the cost becomes roughly
> visits * prog_len.
>
> For example, a program with about 500K reachable insns plus a bounded loop
> of the form "if (x == i) bpf_unwind();" could reach bpf_unwind() tens of
> thousands of times within the 1M processed insn budget.  Each visit would
> rescan 500K insns, which is around 10^10 to 10^11 extra iterations, or
> tens of seconds of extra verifier CPU for a single load.
>
> Would it make sense to run reject_throw() once and cache the result in
> env?  Alternatively, the bpf_unwind() path could re-check only the
> gen_epilogue condition, which is the one that becomes known after
> bpf_exc_prepare().

As you mentioned, yes, this is an optimization issue. We can remove
bpf_exc_check_prog() for checking per bpf_unwind() insn. We can do
a collective checking after all insns are checked. This way
bpf_exc_check_prog() will be called only once.

>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36872142096


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier
  2026-10-01 13:50   ` sashiko-bot
@ 2026-10-02 19:31     ` Yonghong Song
  0 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-02 19:31 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf



On 10/1/26 6:50 AM, sashiko-bot@kernel.org wrote:
> Subject: bpf: Follow an unwind to its landing pad in the verifier
>
> This patch updates the BPF verifier to follow unwind operations to their
> landing pads. It handles how an unwind rewrites return addresses of frames
> and evaluates the proper precision backtracking.
>
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
> [ ... ]
>> @@ -19126,6 +19145,127 @@ enum {
>>   	INSN_IDX_UPDATED = 2,
>>   };
>>   
>> +/*
>> + * The current frame is leaving through an unwind. Its caller's saved return
>> + * address now points at the pad covering the call, or, with none, at the
>> + * caller's epilogue, and so on down. Follow that from the state the frame
>> + * leaves in -- anything it wrote into its callers' stacks included -- to the
>> + * first pad, or to the main program's frame returning.
>> + */
>> +static int unwind_frames(struct bpf_verifier_env *env, bool *do_print_state)
>> +{
>> +	struct bpf_verifier_state *state = env->cur_state;
>> +	u32 frameno = state->curframe;
>> +	struct bpf_func_state *callee, *caller;
>> +	int err, pad;
>> +
>> +	while (state->curframe) {
>> +		callee = cur_func(env);
>> +		caller = state->frame[state->curframe - 1];
>> +		pad = bpf_exc_pad_of_call(env, callee->callsite);
>> +		/* The caller is at its call now, not at this frame's insn. */
>> +		state->insn_idx = callee->callsite;
>> +		account_processed_insns(env, callee, caller);
>> +		free_func_state(callee);
>> +		state->frame[state->curframe--] = NULL;
>> +		invalidate_outgoing_stack_args(env, caller);
>> +		if (pad < 0)
>> +			continue;
> [Severity: Critical]
> If a callback subprogram (such as one passed to bpf_loop()) calls
> bpf_unwind(), does this loop incorrectly pop frames in the verifier state past
> the C helper?
>
> It appears bpf_unwind() and bpf_unwind_resume() lack the callback restriction
> that bpf_throw() has in do_check_subprogs().
>
> At runtime, the exception unwinder stops unwinding when it encounters the
> non-BPF C helper frame and returns. The callback returns 0, the C helper
> resumes, and it returns normally to the BPF program.
>
> However, if the verifier simulates an unwind through the C helper to the
> caller's landing pad, it marks the normal execution path as dead code
> and replaces it with a ja -1.
>
> Could this divergence between verifier state and runtime execution cause an
> infinite loop and hang the kernel?

This won'd happen. See

int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog)
{
         if (!env->subprog_info[subprog].might_unwind)
                 return 0;
                 
         verbose(env, "subprog %d may unwind and is used as a callback\n", subprog);
         return -EINVAL;
}

If any callback function may be might_unwind which implies some callee(s)
or further down callee(s) must be a bpf_unwind(), then bpf_exc_check_callback()
will fail.

>
> [ ... ]
>> @@ -19380,13 +19520,32 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>>   					return -EINVAL;
>>   				}
>>   			}
>> +			if (bpf_is_unwind_kfunc(insn) || bpf_is_unwind_resume_kfunc(insn)) {
>> +				err = check_kfunc_allowed_only(env, insn, env->insn_idx);
>> +				if (err)
>> +					return err;
>> +				if (bpf_is_unwind_kfunc(insn))
>> +					return process_bpf_unwind(env, &env->insn_idx,
>> +								  do_print_state);
>> +				/*
>> +				 * The fixups lower this to 'r0 = 0; exit', and
>> +				 * the unwind goes on below this frame.
>> +				 */
>> +				return unwind_frames(env, do_print_state);
>> +			}
>>   			mark_reg_scratched(env, BPF_REG_0);
>>   			if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
>>   				cur_func(env)->no_stack_arg_load = true;
>>   			if (bpf_is_callx(insn))
>>   				return check_func_callx(env, insn, &env->insn_idx);
>> -			if (insn->src_reg == BPF_PSEUDO_CALL)
>> -				return check_func_call(env, insn, &env->insn_idx);
>> +			if (insn->src_reg == BPF_PSEUDO_CALL) {
>> +				int call_idx = env->insn_idx;
>> +
>> +				err = check_func_call(env, insn, &env->insn_idx);
>> +				if (err)
>> +					return err;
>> +				return unwind_out_of_global_call(env, call_idx, do_print_state);
>> +			}
> [Severity: Critical]
> If check_func_call() processes a global subprogram that is marked with both
> might_throw and might_unwind, it evaluates might_throw, calls
> process_bpf_exit_full(), and returns PROCESS_BPF_EXIT.

The global subprog is marked with both might_throw and might_unwind. The previous
suggested like

         if (env->cleanup_info_cnt || bpf_prog_may_unwind(env)) {
                 ret = bpf_exc_check_prog(env);
                 if (ret)
                         goto skip_full_check;
         }

and we have

bool bpf_prog_may_unwind(const struct bpf_verifier_env *env)
{
         u32 i;

         for (i = 0; i < env->subprog_cnt; i++)
                 if (env->subprog_info[i].might_unwind)
                         return true;
         return false;
}

so bpf_exc_check_prog() will fail if prog has both might_throw and might_unwind.


>
> Does this cause do_check_insn() to return early here with err = 1, completely
> bypassing the unwind_out_of_global_call() logic?
>
> If bypassed, the landing pad associated with the call is never verified and
> is replaced with a ja -1 (dead code). At runtime, if the global subprogram
> calls bpf_unwind(), the exception unwinder routes execution to this unverified
> landing pad.
>
> Can this result in hitting the ja -1 infinite loop and hanging the kernel?


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier
  2026-10-01 14:31   ` bot+bpf-ci
@ 2026-10-02 20:49     ` Yonghong Song
  0 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-02 20:49 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 10/1/26 7:31 AM, bot+bpf-ci@kernel.org wrote:
>> bpf: Follow an unwind to its landing pad in the verifier
> TL;DR:
> - Global call that unwinds without a landing pad may hit a verifier bug
>    warning in backtrack_insn() on fentry/cgroup_skb programs
> - bpf_unwind() skips no_stack_arg_load, so a stale R9 arg-6 load may pass
>    verification
>
>> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
>> index 0e38b9575328..f5504334df90 100644
>> --- a/kernel/bpf/backtrack.c
>> +++ b/kernel/bpf/backtrack.c
> [ ... ]
>
>> @@ -434,6 +458,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
>>   					return -EFAULT;
>>   			}
>>
>> +			if (bpf_exc_pad_of_call(env, idx) == subseq_idx) {
>> +				/*
>> +				 * We came from the landing pad of a call to a
>> +				 * global subprog, branched to from the state
>> +				 * the call returns in: as on its return, no
>> +				 * frame was entered here. The call clobbered
>> +				 * r0-r5; r6-r9 and the stack are the caller's
>> +				 * own and keep going back from here.
>> +				 */
>> +				bt_clear_reg(bt, BPF_REG_0);
>> +				if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
>> +					verifier_bug(env, "landing pad unexpected regs %x",
>> +						     bt_reg_mask(bt));
>> +					return -EFAULT;
>> +				}
>> +				return 0;
>> +			}
>> +
>>   			/* callx calls static subprogs only */
>>   			if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
>>   				/* check that jump history doesn't have any
> Can the existing idx + 1 != subseq_idx check in this global branch now fire
> for a global call that unwinds with no landing pad?
>
>      verifier_bug_if(idx + 1 != subseq_idx, env,
>                      "extra insn from subprog");
>
> In that case unwind_out_of_global_call() takes the pad < 0 path into
> unwind_frames(), and when unwind_frames() reaches the main frame it pushes a
> history entry with no flags:
>
>      env->prev_insn_idx = env->insn_idx;
>      env->insn_idx = state->insn_idx;
>      err = bpf_push_jmp_history(env, state, 0, 0, 0, NULL, 0);
>      ...
>      return process_bpf_exit_full(env, do_print_state, false);
>
> The entry's prev_idx is the global call insn, so INSN_F_UNWIND is not set.
> For program types where return_retval_range() returns true (fentry/fexit,
> cgroup_skb, struct_ops), process_bpf_exit_full() -> check_return_code() calls
> mark_chain_precision(env, BPF_REG_0).
>
> The backtrack then reaches backtrack_insn() with idx at the global call and
> subseq_idx at either the same insn (call in main) or the call site in main
> (static frames popped first).  The INSN_F_UNWIND branch does not match, the
> bpf_exc_pad_of_call() == subseq_idx branch does not match because no pad
> covers the call, and the walk lands in the global branch above, where
> idx + 1 != subseq_idx is true.
>
> The result is a WARN_ONCE() (a panic with panic_on_warn) and a "verifier bug"
> message in the log, although the program is still accepted.  An example is an
> fentry or cgroup_skb program calling a global subprog that calls
> bpf_unwind() with no record covering the call, or main -> static A -> global
> G with no records.
>
> The selftests in the series all use SEC("syscall"), where
> return_retval_range() returns false and r0 precision is never requested at
> exit, so they would not reach this.  Nothing later in the series touches
> backtrack.c.
>
> Should the main-exit history entry get its own flag, or should this global
> branch handle that entry before applying the idx + 1 check?

Yes, this global branch should handle that entry before applying the idx + 1
check (global function) and we can add INSN_F_UNWIND to this jump history
to keep later continuing doing static function.

>
> [ ... ]
>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index 80034429fdd0..c7a350be538e 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
> [ ... ]
>
>> @@ -19380,13 +19520,32 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>>   					return -EINVAL;
>>   				}
>>   			}
>> +			if (bpf_is_unwind_kfunc(insn) || bpf_is_unwind_resume_kfunc(insn)) {
>> +				err = check_kfunc_allowed_only(env, insn, env->insn_idx);
>> +				if (err)
>> +					return err;
>> +				if (bpf_is_unwind_kfunc(insn))
>> +					return process_bpf_unwind(env, &env->insn_idx,
>> +								  do_print_state);
>> +				/*
>> +				 * The fixups lower this to 'r0 = 0; exit', and
>> +				 * the unwind goes on below this frame.
>> +				 */
>> +				return unwind_frames(env, do_print_state);
>> +			}
>>   			mark_reg_scratched(env, BPF_REG_0);
>>   			if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
>>   				cur_func(env)->no_stack_arg_load = true;
> Does this early return skip the no_stack_arg_load assignment that every
> other call insn gets?
>
> For bpf_unwind() with a pad in the same frame, process_bpf_unwind() goes on
> at the pad in the frame that made the call, with no_stack_arg_load still
> false:
>
>      clear_caller_saved_regs(env, frame->regs);
>      mark_reg_unknown(env, frame->regs, BPF_REG_0);
>      *insn_idx = pad;
>      return INSN_IDX_UPDATED;
>
> On x86, do_jit() turns a load from BPF_REG_PARAMS at offset 8 (incoming
> arg 6) into a read of R9, which is caller-saved and clobbered by any call.
> check_stack_arg_read() relies on no_stack_arg_load to reject such a load
> after a call.
>
> So in a subprog with more than 5 args, would this sequence pass
> verification?
>
>      call bpf_unwind        // record covers it, pad P in this frame
>    P:
>      r1 = *(u64 *)(r11 + 8) // incoming arg 6
>      r2 = *(u64 *)(r1 + 0)
>
> check_stack_arg_read() gives r1 the type of caller->stack_arg_regs[0], which
> can be any spilled type, including a pointer such as PTR_TO_MAP_VALUE, but
> at run time r1 holds whatever bpf_unwind() left in R9.
>
> Would it work to move the mark_reg_scratched() and no_stack_arg_load lines
> above the unwind branch, or to set cur_func(env)->no_stack_arg_load in
> process_bpf_unwind() before going on at the pad?
>
> The unwind_out_of_global_call() pad path looks unaffected, since the
> PSEUDO_CALL path sets the flag before check_func_call().  do_check_insn()
> has the same ordering at the end of the series, where both kfuncs become
> callable.

Yes, you are right. The below:

  			mark_reg_scratched(env, BPF_REG_0);
  			if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
  				cur_func(env)->no_stack_arg_load = true;

should be moved earlier so later some prog checking can inherit some
choices.

>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36872142096


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with
  2026-10-01 14:31   ` bot+bpf-ci
@ 2026-10-02 21:10     ` Yonghong Song
  0 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-02 21:10 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 10/1/26 7:31 AM, bot+bpf-ci@kernel.org wrote:
>> bpf: Require an unwind to leave a frame holding what it entered with
> TL;DR: call_may_unwind() treats every callx as possibly unwinding, so the
> verifier may wrongly reject valid programs that hold a reference or spin lock
> across a callx; the exact target is available from the register.
>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index c7a350be538e..a7b25ab04051 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
> [ ... ]
>
>> @@ -19244,6 +19245,32 @@ static int unwind_out_of_global_call(struct bpf_verifier_env *env, int call_idx,
>>   	return INSN_IDX_UPDATED;
>>   }
>>
>> +/* Can an unwind come back out of this call? */
>> +static bool call_may_unwind(struct bpf_verifier_env *env, const struct bpf_insn *insn,
>> +			    int insn_idx)
>> +{
>> +	int subprog;
>> +
>> +	/* Which subprog a callx lands in is not known here, so any may be it. */
>> +	if (bpf_is_callx(insn))
>> +		return bpf_prog_may_unwind(env);
> Is it true that the target of a callx is not known here?

This is way to conservatile as if *any* subprog is marked as may_unwind, then
it will return bpf_is_callx(insn). The best thing is to follow callx itself
and find whether it is possible may_unwind or not.

>
> call_may_unwind() is reached from check_unwind_through_call() in
> do_check_insn(), just before check_func_callx().  check_func_callx() takes
> the exact target from the same register state a few lines later:
>
> 	reg = reg_state(env, insn->dst_reg);
> 	if (reg->type != PTR_TO_FUNC) ...
> 	/* PTR_TO_FUNC is a pointer to a static subprog */
> 	subprog = reg->subprogno;
>
> and check_func_ptr_read() pushes one state per possible table target, each
> with its own subprogno.  So each verified callx has exactly one target at
> this point.  Only the CFG pass has to be conservative, and
> mark_callx_might_unwind() already marks any target that itself contains a
> callx.
>
> With the current code, an uncovered callx is refused by
> bpf_exc_check_frame_balance() whenever the calling frame has changed its
> RCU, preempt, IRQ, spin lock or reference state since entry, as long as some
> unrelated subprog in the program is marked might_unwind.  For example:
>
> 	foo():
> 	    t = bpf_task_acquire(...);  /* id > entry_id_gen */
> 	    r1 = vtable[i];             /* PTR_TO_FUNC -> bar, no unwind in bar */
> 	    callx r1                    /* no cleanup record covers it */
> 	    bpf_task_release(t);
>
> 	baz():
> 	    bpf_unwind();               /* makes bpf_prog_may_unwind() true */
>
> Before this patch this program was accepted.  Now it is refused with "an
> unwind through this call keeps the reference id=N the frame acquired", even
> though no unwind can pass through that call.  callx is also explicitly
> allowed under a spin lock, so a frame that holds a bpf_spin_lock across a
> callx is refused as well with "an unwind through this call does not leave
> the frame's bpf_spin_lock state as it found it".
>
> Could this look at the register instead, something like:
>
> 	reg = &cur_regs(env)[insn->dst_reg];
> 	if (reg->type == PTR_TO_FUNC)
> 		return env->subprog_info[reg->subprogno].might_unwind;
>
> and leave it to check_func_callx() to refuse any other register type, as it
> does today?
>
>> +	if (insn->src_reg != BPF_PSEUDO_CALL)
>> +		return false;
>> +	subprog = bpf_find_subprog(env, insn_idx + insn->imm + 1);
>> +	return subprog >= 0 && env->subprog_info[subprog].might_unwind;
>> +}
>> +
> [ ... ]
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36872142096


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 10/22] bpf: Do not use a private stack for a program that can unwind
  2026-10-01 13:53   ` sashiko-bot
@ 2026-10-02 21:38     ` Yonghong Song
  0 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-02 21:38 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf



On 10/1/26 6:53 AM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf: callee-saved register corruption via bpf_throw without cleanup
> --
>
> commit 9d2c0e8faeb024a4647980c48ca2caa933b90b1c
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Do not use a private stack for a program that can unwind
>
> This commit disables private stacks for BPF programs that can unwind
> because an unwind skips the pop instruction corresponding to the push r9
> emitted around calls. This forces NO_PRIV_STACK in check_max_stack_depth()
> for programs with cleanup records or that may unwind.
>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index f3ed68960d70d..488ceb9dae1b6 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
>> @@ -5769,6 +5769,17 @@ static int check_max_stack_depth(struct bpf_verifier_env *env)
>>   		}
>>   	}
>>   
>> +	/*
>> +	 * A private stack keeps its frame pointer in %r9 on x86-64, restored
>> +	 * by a pop after the call that an unwind skips. A frame resumed at a
>> +	 * pad then addresses its stack through a stale pointer, and a frame
>> +	 * sent to its epilogue instead pops its callee-saved registers one
>> +	 * slot off. Refuse a private stack for any program that can unwind,
>> +	 * on every arch for now.
>> +	 */
>> +	if (env->cleanup_info_cnt || bpf_prog_may_unwind(env))
>> +		priv_stack_mode = NO_PRIV_STACK;
> [Severity: High]
> Does this logic miss programs that call bpf_throw() without any cleanup
> records?
>
> If a privileged program uses private stacks and calls bpf_throw() without
> acquiring any resources requiring cleanup, cleanup_info_cnt will be 0.
>
> Looking at bpf_prog_may_unwind() in kernel/bpf/exception.c, it only checks
> might_unwind:
>
> kernel/bpf/exception.c:bpf_prog_may_unwind() {
> 	u32 i;
>
> 	for (i = 0; i < env->subprog_cnt; i++)
> 		if (env->subprog_info[i].might_unwind)
> 			return true;
> 	return false;
> }
>
> Since bpf_throw() only sets might_throw, not might_unwind, it appears
> bpf_prog_may_unwind() will return false.
>
> Could this bypass the NO_PRIV_STACK enforcement and lead to the exception
> callback epilogue popping the misaligned stack (shifted by the skipped
> push r9) into the kernel's callee-saved registers?

I cannot figure out a condition to have misaligned stack (i.e. using r9)
like the above condition.

I think the current condition should be okay.

>
>> +
>>   	if (priv_stack_mode == PRIV_STACK_UNKNOWN)
>>   		priv_stack_mode = bpf_enable_priv_stack(env->prog);
>>


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses
  2026-10-01 14:31   ` bot+bpf-ci
@ 2026-10-02 21:48     ` Yonghong Song
  0 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-02 21:48 UTC (permalink / raw)
  To: bot+bpf-ci, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, martin.lau, mason,
	ihor.solodrai



On 10/1/26 7:31 AM, bot+bpf-ci@kernel.org wrote:
>> bpf: Dispatch cleanup pads by rewriting return addresses
> TL;DR: Design question, not a bug: bpf_exc_attach_main_prog() in
> bpf_prog_jit_compile() builds a table for single-function programs that seems
> never to be read, so it and its error paths could possibly be dropped.
>
>> diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
>> index d813fdde29e3..60905643cb9c 100644
>> --- a/kernel/bpf/core.c
>> +++ b/kernel/bpf/core.c
> [ ... ]
>
>> @@ -2632,9 +2633,14 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
>>   {
>>   #ifdef CONFIG_BPF_JIT
>>   	struct bpf_prog *orig_prog;
>> +	int ret;
>>   
>> -	if (!bpf_prog_need_blind(prog))
>> +	if (!bpf_prog_need_blind(prog)) {
>> +		ret = bpf_exc_attach_main_prog(env, prog);
>> +		if (ret)
>> +			return ERR_PTR(ret);
>>   		return bpf_int_jit_compile(env, prog);
>> +	}
>>   
>>   	orig_prog = prog;
>>   	prog = bpf_jit_blind_constants(env, prog);
>> @@ -2648,6 +2654,12 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc
>>   		goto out_restore;
>>   	}
>>   
>> +	ret = bpf_exc_attach_main_prog(env, prog);
>> +	if (ret) {
>> +		bpf_jit_prog_release_other(orig_prog, prog);
>> +		return ERR_PTR(ret);
>> +	}
>> +
>>   	prog = bpf_int_jit_compile(env, prog);
>>   	if (prog->jited) {
>>   		bpf_jit_prog_release_other(prog, orig_prog);
> this isn't a bug, but is the table that bpf_exc_attach_main_prog() builds
> for a single-function program read anywhere?
>
> It only attaches a table when there are no subprograms:
>
> kernel/bpf/fixups.c:bpf_exc_attach_main_prog() {
>      if (!env || env->subprog_cnt > 1)
>          return 0;
>      return exc_info_for_subprog(env, prog, 0, prog->len);
> }
>
> The only reader of the table that I can find is bpf_unwind_rewrite(), and it
> returns at the first BPF frame before it does the pad lookup:
>
> kernel/bpf/helpers.c:bpf_unwind_rewrite() {
>      ...
>      if (ctx->cnt == 1)
>          return bpf_is_subprog(prog);
>
>      rec = bpf_exc_pad_for_ip(prog, ip);
>      ...
> }
>
> With no subprograms, the main frame is the only BPF frame, so it is always
> the one with ctx->cnt == 1 and bpf_exc_pad_for_ip() is not reached for it.
> The bpf_unwind() calls in such a program are already sent to their pad by
> the jump that bpf_exc_patch_unwind_calls() inserts.
>
> If that is right, could bpf_exc_attach_main_prog() and its two call sites in
> bpf_prog_jit_compile(), with the new 'return ERR_PTR(ret)' paths, be dropped
> so that bpf_prog_jit_compile() stays as it was?

You are right. So dead code. Will remove.

>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36872142096


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 12/22] bpf, x86: Dispatch exception cleanup pads at run time
  2026-10-01 13:49   ` sashiko-bot
@ 2026-10-02 21:54     ` Yonghong Song
  0 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-02 21:54 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf



On 10/1/26 6:49 AM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf: verifier bypass via BPF exception unwind abort on traced frames
> --
>
> commit 8f24ec5b2338fd8477df1e09b72647e4de2ae3d4
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf, x86: Dispatch exception cleanup pads at run time
>
> This commit introduces runtime dispatch of exception cleanup pads using the
> ORC unwinder. It implements arch_bpf_stack_walk_ra() to traverse the stack
> and rewrite return addresses to point to their corresponding cleanup pads.
>
>> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
>> index 6c7a0578760e7..544e8fd759ad4 100644
>> --- a/arch/x86/net/bpf_jit_comp.c
>> +++ b/arch/x86/net/bpf_jit_comp.c
> [ ... ]
>
>> @@ -4614,6 +4628,32 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
>>   #endif
>>   }
>>   
>> +void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
>> +			    void *cookie)
>> +{
>> +#if defined(CONFIG_UNWINDER_ORC)
>> +	struct unwind_state state;
>> +	unsigned long addr, *ra;
>> +
>> +	for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
>> +	     unwind_next_frame(&state)) {
>> +		addr = unwind_get_return_address(&state);
>> +		ra = unwind_get_return_address_ptr(&state);
>> +		if (!addr || !ra)
>> +			break;
>> +		/*
>> +		 * A traced return: the slot holds a function graph or kretprobe
>> +		 * trampoline, not @addr, so it cannot be rewritten. Stopping
>> +		 * leaves BPF frames returning to unverified paths, so warn.
>> +		 */
>> +		if (WARN_ON_ONCE(READ_ONCE_NOCHECK(*ra) != addr))
>> +			break;
> [Severity: High]
> The comment above notes that this leaves BPF frames returning to unverified
> paths. Since the verifier assumes bpf_throw() does not return and prunes
> instructions following it, won't aborting the stack walk here allow the
> execution of those unverified instructions?
>
> Could a user with tracing permissions intentionally attach a tracer (like
> kretprobe or fgraph) to trigger this break, bypassing the verifier and
> forcing the execution of unverified BPF code?

Okay, this has been mentioned a few times. I think I should follow previous
suggestions e.g. mark bpf_unwind() as notrace etc. Will fix.

>
>> +		if (!consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
>> +			break;
>> +	}
>> +#endif
>> +}


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts
  2026-10-01 13:46   ` sashiko-bot
@ 2026-10-02 22:09     ` Yonghong Song
  0 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-02 22:09 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf



On 10/1/26 6:46 AM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf: ABI breakage due to reuse of padding in bpf_prog_load_opts
> --
>
> commit 71e2541e0b207ab646c97a6cfe861ce9bf16fe7f
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> libbpf: Add cleanup_info to bpf_prog_load_opts
>
> This commit introduces an exception cleanup table to the bpf_prog_load
> options. It adds cleanup_info_cnt, cleanup_info, and cleanup_info_rec_size
> to struct bpf_prog_load_opts and propagates them to the kernel in
> bpf_prog_load.
>
>> diff --git a/tools/lib/bpf/bpf.c b/tools/lib/bpf/bpf.c
>> index b49822d212aed..b4031f62bdee5 100644
>> --- a/tools/lib/bpf/bpf.c
>> +++ b/tools/lib/bpf/bpf.c
> [ ... ]
>> @@ -370,6 +370,10 @@ int bpf_prog_load(enum bpf_prog_type prog_type,
>>   	attr.fd_array = ptr_to_u64(OPTS_GET(opts, fd_array, NULL));
>>   	attr.fd_array_cnt = OPTS_GET(opts, fd_array_cnt, 0);
>>   
>> +	attr.cleanup_info = ptr_to_u64(OPTS_GET(opts, cleanup_info, NULL));
>> +	attr.cleanup_info_rec_size = OPTS_GET(opts, cleanup_info_rec_size, 0);
>> +	attr.cleanup_info_cnt = OPTS_GET(opts, cleanup_info_cnt, 0);
>> +
> [Severity: High]
> Does this OPTS_GET macro read uninitialized padding from older applications?
> When an application compiled against an older version of libbpf passes its
> opts structure here, the sz field will be set to the old struct size.
>
> If the old struct had padding at the end that the new structure replaces
> with cleanup_info_cnt, OPTS_GET might think the field is present because
> the sz value includes the padding, causing it to read stack garbage.
>
>>   	if (log_level) {
>>   		attr.log_buf = ptr_to_u64(log_buf);
>>   		attr.log_size = log_size;
>> diff --git a/tools/lib/bpf/bpf.h b/tools/lib/bpf/bpf.h
>> index 826d9cc9ab65d..cbe56ddc8cf7d 100644
>> --- a/tools/lib/bpf/bpf.h
>> +++ b/tools/lib/bpf/bpf.h
> [ ... ]
>> @@ -128,9 +128,14 @@ struct bpf_prog_load_opts {
>>   
>>   	/* if set, provides the length of fd_array */
>>   	__u32 fd_array_cnt;
>> +
>> +	/* exception cleanup table, from the .bpf_cleanup section */
>> +	__u32 cleanup_info_cnt;
> [Severity: High]
> Does placing cleanup_info_cnt immediately after fd_array_cnt break the
> size-based backward compatibility check?
>
> Because fd_array_cnt is a 4-byte field, the struct bpf_prog_load_opts in
> older applications likely contained a 4-byte padding hole at the end to
> align with the 8-byte size_t type.
>
> An older application would pass an opts structure where sz includes those 4
> bytes of padding. Because OPTS_HAS evaluates if sz >=
> offsetofend(struct bpf_prog_load_opts, cleanup_info_cnt), it will evaluate
> to true. The library will then read the uninitialized padding left by the
> older application as the cleanup_info_cnt value.
>
> If the padding contains garbage, the kernel could reject the program load
> or fail unexpectedly. Could this be resolved by placing cleanup_info_cnt
> after the pointer, or by introducing explicit padding?
>
>> +	const void *cleanup_info;
>> +	__u32 cleanup_info_rec_size;
>>   	size_t :0;

I think the current implementation is okay. IIUC, 'size_t :0' will ensure
to filling '0''s for tailing unnamed fields.


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier
  2026-10-01 13:30 ` [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier Yonghong Song
  2026-10-01 13:50   ` sashiko-bot
  2026-10-01 14:31   ` bot+bpf-ci
@ 2026-10-03 12:23   ` Alexei Starovoitov
  2026-10-04 17:56     ` Yonghong Song
  2 siblings, 1 reply; 50+ messages in thread
From: Alexei Starovoitov @ 2026-10-03 12:23 UTC (permalink / raw)
  To: Yonghong Song, bpf
  Cc: Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, kernel-team

On Thu, Oct 01, 2026 at 06:30 AM Yonghong Song <yonghong.song@linux.dev> wrote:

in addition to what bpf-ci found...

can we let both kfuncs go through check_kfunc_call(), like bpf_throw()
does, and follow the unwind after it?
Then there is nothing to open code here and no need for
check_kfunc_allowed() and check_kfunc_allowed_only().

pw-bot: cr

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with
  2026-10-01 13:30 ` [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
  2026-10-01 14:31   ` bot+bpf-ci
@ 2026-10-03 12:25   ` Alexei Starovoitov
  2026-10-04 17:59     ` Yonghong Song
  1 sibling, 1 reply; 50+ messages in thread
From: Alexei Starovoitov @ 2026-10-03 12:25 UTC (permalink / raw)
  To: Yonghong Song, bpf
  Cc: Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, kernel-team

On Thu, Oct 01, 2026 at 06:30 AM Yonghong Song <yonghong.song@linux.dev> wrote:
> frame that dropped it. The rule does refuse a program that could be proved
> safe, a callee releasing its caller's reference and a caller's pad that
> relies on that, but compiled code does not do that. A pad only knows about
> its own frame.

Compiled code does exactly that. It's a move.

  fn consume(rec: Record) { may_panic(); }
  fn foo() { let rec = reserve(); consume(rec); }

consume() owns rec, so the pad of consume() drops it.
foo() has nothing left to drop and rustc emits a plain call with no
record over it.
'call consume' is refused with
"an unwind through this call keeps the reference".
With a record over that call the resume in consume() is refused with
"a resume does not leave the frame's references as it found it".
pad_drops_caller_ref_frame() in patch 19 is what consume() compiles to.
RcuReadGuard passed by value is the same.

[...]

> +	u32 entry_active_locks;
> +	u32 entry_preempt_locks;
> +	u32 entry_rcu_locks;
> +	u32 entry_irq_id;
> +	u32 entry_id_gen;
> +	u32 entry_acquired_refs;

The verifier doesn't need them anymore.
"Unreleased reference id=%d alloc_insn=%d" at exit already points
at the insn that acquired it.
Drop this patch.

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 09/22] bpf: Refuse a landing pad that does not resume
  2026-10-01 13:30 ` [PATCH bpf-next v8 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
@ 2026-10-03 12:25   ` Alexei Starovoitov
  2026-10-04 18:26     ` Yonghong Song
  0 siblings, 1 reply; 50+ messages in thread
From: Alexei Starovoitov @ 2026-10-03 12:25 UTC (permalink / raw)
  To: Yonghong Song, bpf
  Cc: Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, kernel-team

On Thu, Oct 01, 2026 at 06:30 AM Yonghong Song <yonghong.song@linux.dev> wrote:
> +	u64 in_cleanup_pad:1; /* reached with a landing pad running */
> +	u64 outside_cleanup_pad:1; /* reached the other way */

can we compare in_pad in func_states_equal(), like no_stack_arg_load,
and drop these two bits, the "runs both inside and outside a landing
pad" rule and the special case for speculative paths?

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses
  2026-10-01 13:31 ` [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
  2026-10-01 14:31   ` bot+bpf-ci
@ 2026-10-03 12:26   ` Alexei Starovoitov
  2026-10-04 18:28     ` Yonghong Song
  2026-10-04 18:29     ` Yonghong Song
  1 sibling, 2 replies; 50+ messages in thread
From: Alexei Starovoitov @ 2026-10-03 12:26 UTC (permalink / raw)
  To: Yonghong Song, bpf
  Cc: Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, kernel-team

On Thu, Oct 01, 2026 at 06:31 AM Yonghong Song <yonghong.song@linux.dev> wrote:
> +	if (!bpf_prog_need_blind(prog)) {
> +		ret = bpf_exc_attach_main_prog(env, prog);
> +		if (ret)
> +			return ERR_PTR(ret);
>  		return bpf_int_jit_compile(env, prog);
> +	}

bpf-ci is right. The table of a prog without subprogs is never read.
Drop bpf_exc_attach_main_prog() and these hunks.

[...]

> +	rcu_read_lock();
> +	prog = bpf_prog_ksym_find(ip);
> +	rcu_read_unlock();
> +	if (!prog)
> +		return !ctx->cnt;

fexit attached to a subprog in the chain stops the walk here.
  main -> A -> trampoline -> B -> C
C calls bpf_unwind(). The return into B is rewritten. The next ip is in
the trampoline, bpf_prog_ksym_find() returns NULL and the returns into
A and main stay as they were.
B returns 0 through the trampoline to A at callsite + 1.
The verifier didn't see that path, or removed it as dead code.

^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier
  2026-10-03 12:23   ` Alexei Starovoitov
@ 2026-10-04 17:56     ` Yonghong Song
  0 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-04 17:56 UTC (permalink / raw)
  To: Alexei Starovoitov, bpf
  Cc: Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, kernel-team



On 10/3/26 2:23 PM, Alexei Starovoitov wrote:
> On Thu, Oct 01, 2026 at 06:30 AM Yonghong Song <yonghong.song@linux.dev> wrote:
>
> in addition to what bpf-ci found...
>
> can we let both kfuncs go through check_kfunc_call(), like bpf_throw()
> does, and follow the unwind after it?
> Then there is nothing to open code here and no need for
> check_kfunc_allowed() and check_kfunc_allowed_only().

Okay, will fix fexit issue and also add bpf_unwind() and bpf_unwind_resume()
inside check_kfunc_call().

>
> pw-bot: cr


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with
  2026-10-03 12:25   ` Alexei Starovoitov
@ 2026-10-04 17:59     ` Yonghong Song
  0 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-04 17:59 UTC (permalink / raw)
  To: Alexei Starovoitov, bpf
  Cc: Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, kernel-team



On 10/3/26 2:25 PM, Alexei Starovoitov wrote:
> On Thu, Oct 01, 2026 at 06:30 AM Yonghong Song <yonghong.song@linux.dev> wrote:
>> frame that dropped it. The rule does refuse a program that could be proved
>> safe, a callee releasing its caller's reference and a caller's pad that
>> relies on that, but compiled code does not do that. A pad only knows about
>> its own frame.
> Compiled code does exactly that. It's a move.
>
>    fn consume(rec: Record) { may_panic(); }
>    fn foo() { let rec = reserve(); consume(rec); }
>
> consume() owns rec, so the pad of consume() drops it.
> foo() has nothing left to drop and rustc emits a plain call with no
> record over it.
> 'call consume' is refused with
> "an unwind through this call keeps the reference".
> With a record over that call the resume in consume() is refused with
> "a resume does not leave the frame's references as it found it".
> pad_drops_caller_ref_frame() in patch 19 is what consume() compiles to.
> RcuReadGuard passed by value is the same.
>
> [...]
>
>> +	u32 entry_active_locks;
>> +	u32 entry_preempt_locks;
>> +	u32 entry_rcu_locks;
>> +	u32 entry_irq_id;
>> +	u32 entry_id_gen;
>> +	u32 entry_acquired_refs;
> The verifier doesn't need them anymore.
> "Unreleased reference id=%d alloc_insn=%d" at exit already points
> at the insn that acquired it.
> Drop this patch.

You are absolutely right. We only need to check resource
right before main prog return. This patch indeed not needed any more.


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 09/22] bpf: Refuse a landing pad that does not resume
  2026-10-03 12:25   ` Alexei Starovoitov
@ 2026-10-04 18:26     ` Yonghong Song
  0 siblings, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-04 18:26 UTC (permalink / raw)
  To: Alexei Starovoitov, bpf
  Cc: Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, kernel-team



On 10/3/26 2:25 PM, Alexei Starovoitov wrote:
> On Thu, Oct 01, 2026 at 06:30 AM Yonghong Song <yonghong.song@linux.dev> wrote:
>> +	u64 in_cleanup_pad:1; /* reached with a landing pad running */
>> +	u64 outside_cleanup_pad:1; /* reached the other way */
> can we compare in_pad in func_states_equal(), like no_stack_arg_load,
> and drop these two bits, the "runs both inside and outside a landing
> pad" rule and the special case for speculative paths?

Yes, will compare in_pad in func_states_equal(). special case for
speculative path can be removed. I think special case for
speculative path is not needed any more due to in_pad in func_states_equal().


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses
  2026-10-03 12:26   ` Alexei Starovoitov
@ 2026-10-04 18:28     ` Yonghong Song
  2026-10-04 18:29     ` Yonghong Song
  1 sibling, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-04 18:28 UTC (permalink / raw)
  To: Alexei Starovoitov, bpf
  Cc: Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, kernel-team



On 10/3/26 2:26 PM, Alexei Starovoitov wrote:
> On Thu, Oct 01, 2026 at 06:31 AM Yonghong Song <yonghong.song@linux.dev> wrote:
>> +	if (!bpf_prog_need_blind(prog)) {
>> +		ret = bpf_exc_attach_main_prog(env, prog);
>> +		if (ret)
>> +			return ERR_PTR(ret);
>>   		return bpf_int_jit_compile(env, prog);
>> +	}
> bpf-ci is right. The table of a prog without subprogs is never read.
> Drop bpf_exc_attach_main_prog() and these hunks.

Ok, will dropl

>
> [...]
>
>> +	rcu_read_lock();
>> +	prog = bpf_prog_ksym_find(ip);
>> +	rcu_read_unlock();
>> +	if (!prog)
>> +		return !ctx->cnt;
> fexit attached to a subprog in the chain stops the walk here.
>    main -> A -> trampoline -> B -> C
> C calls bpf_unwind(). The return into B is rewritten. The next ip is in
> the trampoline, bpf_prog_ksym_find() returns NULL and the returns into
> A and main stay as they were.
> B returns 0 through the trampoline to A at callsite + 1.
> The verifier didn't see that path, or removed it as dead code.

I will reject such a case (fexit, fmod_ret and fsession) as they
all have the same issue.


^ permalink raw reply	[flat|nested] 50+ messages in thread

* Re: [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses
  2026-10-03 12:26   ` Alexei Starovoitov
  2026-10-04 18:28     ` Yonghong Song
@ 2026-10-04 18:29     ` Yonghong Song
  1 sibling, 0 replies; 50+ messages in thread
From: Yonghong Song @ 2026-10-04 18:29 UTC (permalink / raw)
  To: Alexei Starovoitov, bpf
  Cc: Andrii Nakryiko, Daniel Borkmann, Eduard Zingerman, kernel-team



On 10/3/26 2:26 PM, Alexei Starovoitov wrote:
> On Thu, Oct 01, 2026 at 06:31 AM Yonghong Song <yonghong.song@linux.dev> wrote:
>> +	if (!bpf_prog_need_blind(prog)) {
>> +		ret = bpf_exc_attach_main_prog(env, prog);
>> +		if (ret)
>> +			return ERR_PTR(ret);
>>   		return bpf_int_jit_compile(env, prog);
>> +	}
> bpf-ci is right. The table of a prog without subprogs is never read.
> Drop bpf_exc_attach_main_prog() and these hunks.

Ok, will drop.

>
> [...]
>
>> +	rcu_read_lock();
>> +	prog = bpf_prog_ksym_find(ip);
>> +	rcu_read_unlock();
>> +	if (!prog)
>> +		return !ctx->cnt;
> fexit attached to a subprog in the chain stops the walk here.
>    main -> A -> trampoline -> B -> C
> C calls bpf_unwind(). The return into B is rewritten. The next ip is in
> the trampoline, bpf_prog_ksym_find() returns NULL and the returns into
> A and main stay as they were.
> B returns 0 through the trampoline to A at callsite + 1.
> The verifier didn't see that path, or removed it as dead code.

I will reject such a case (fexit, fmod_ret and fsession) as they
all have the same issue.


^ permalink raw reply	[flat|nested] 50+ messages in thread

end of thread, other threads:[~2026-10-04 18:29 UTC | newest]

Thread overview: 50+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-10-01 13:48   ` sashiko-bot
2026-10-02 18:17     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 19:06     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier Yonghong Song
2026-10-01 13:50   ` sashiko-bot
2026-10-02 19:31     ` Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 20:49     ` Yonghong Song
2026-10-03 12:23   ` Alexei Starovoitov
2026-10-04 17:56     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 21:10     ` Yonghong Song
2026-10-03 12:25   ` Alexei Starovoitov
2026-10-04 17:59     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-10-03 12:25   ` Alexei Starovoitov
2026-10-04 18:26     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 10/22] bpf: Do not use a private stack for a program that can unwind Yonghong Song
2026-10-01 13:53   ` sashiko-bot
2026-10-02 21:38     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 21:48     ` Yonghong Song
2026-10-03 12:26   ` Alexei Starovoitov
2026-10-04 18:28     ` Yonghong Song
2026-10-04 18:29     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-10-01 13:49   ` sashiko-bot
2026-10-02 21:54     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 13/22] bpf, arm64: " Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-10-01 13:46   ` sashiko-bot
2026-10-02 22:09     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
2026-10-01 13:32 ` [PATCH bpf-next v8 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox