From: Yonghong Song <yonghong.song@linux.dev>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
Andrii Nakryiko <andrii@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Eduard Zingerman <eddyz87@gmail.com>,
kernel-team@fb.com
Subject: [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds
Date: Mon, 21 Sep 2026 14:00:33 -0700 [thread overview]
Message-ID: <20260921210033.1715000-1-yonghong.song@linux.dev> (raw)
bpf_throw() walks the BPF call stack to the exception boundary and
discards every frame in between. A frame that owns something -- an RCU
read lock, a preemption-disabled section, a referenced kptr -- never gets
to give it back, so the verifier refuses to let such a frame throw at
all. That is the whole reason a Rust program cannot use bpf_throw() as
its panic path right now: Rust's Drop glue *is* that give-back, and there
is nowhere to run it.
LLVM 23 added the compiler half ([1]). A Rust function that owns a value
across a call that can unwind
fn foo() {
let _guard = RcuReadGuard::new(); /* bpf_rcu_read_lock() */
may_throw(); /* extern "C-unwind" */
} /* Drop: rcu_read_unlock */
lowers to an invoke with a cleanup landing pad holding the Drop call, and
the BPF backend writes one record per invoke region into a .bpf_cleanup
section: a flat table of 12-byte (begin, end, landing_pad) triples, each
field a byte offset into the code section. The rule is "if a frame unwinds
with its return address in [begin, end), run landing_pad before discarding
it". A pad ends with a call to _Unwind_Resume(), which the kernel provides
as the bpf_unwind_resume() kfunc.
This series is the kernel half: take that table at BPF_PROG_LOAD, teach
the verifier that a covered call can also go to its landing pad, and have
bpf_throw() run the pads as it walks.
C has no unwinding, so the selftests spell out by hand what a frontend
emits -- a call site bracketed by two labels, a landing pad, and a record
tying them together. The frame above, written that way:
"call bpf_rcu_read_lock;"
"1:" "call foo3;" /* cleanup region */
"2:"
... normal path, ends in bpf_rcu_read_unlock ...
"6:" /* landing pad */
"call bpf_rcu_read_unlock;"
"call bpf_unwind_resume;"
CLEANUP_REC("1b", "2b", "6b")
Right now that program does not load: the throw inside foo3 is reported as
"bpf_throw cannot be used inside bpf_rcu_read_lock-ed region", because as
far as the verifier is concerned nothing will ever unlock. With the series
the verifier walks the unwind the same way the run time will -- into the
pad, which unlocks, then on to the next frame -- and the program both
loads and releases the lock when it throws.
Design
======
A pad is run, not lowered. bpf_throw() already walks the frames with
arch_bpf_stack_walk(); it now looks each frame's return address up in
that (sub)program's table and calls the pad as a subroutine of the walker,
with the unwinding frame's frame pointer and its callee-saved registers
restored from the spill its callee's prologue left. The pad therefore sees
its own frame but runs on the walker's stack, far below it, so nothing it
calls can disturb the frame it is cleaning up after. The JIT turns its
bpf_unwind_resume() into the way back to the walker.
The verifier walks the same thing, step for step, so the resource rules
are unchanged: whatever a pad releases is released in the verifier state
too, and check_resource_leak() simply moves from "a throw was seen" to the
end of the walk.
1-2 uapi: cleanup_info in BPF_PROG_LOAD, struct bpf_cleanup_info,
and the bpf_unwind_resume() kfunc
3-9 verifier: mark the covered call sites, give them an edge to the
pad, walk the unwind, and refuse the shapes that cannot be
dispatched
10 bpf_throw(): dispatch pads while walking
11-12 x86-64 and arm64 JITs
13-17 libbpf: collect .bpf_cleanup, pass it to the kernel, resolve
_Unwind_Resume, carry it through the light skeleton and the
static linker
18-20 selftests
Limitations
===========
- Cleanup pads only. A catch pad -- one that ends in a plain exit rather
than a resume, which is what Rust's catch_unwind would need -- is
refused: the walker calls a pad as a subroutine and cannot hand a frame
back its own execution. LLVM refuses type-specific catches and filters
on its side as well.
- A JIT that can dispatch pads is required: x86-64 (with
CONFIG_UNWINDER_ORC, which bpf_throw() already needs there) and arm64.
Anywhere else the load fails with -EOPNOTSUPP rather than silently
doing nothing.
- No offloaded programs, no private stack, and no combining a table with
an exception callback.
- In a pad body: no tail call, no BPF_LD_[ABS|IND], no indirect jump, and
no on-stack call arguments -- each of them reads or writes a stack the
pad does not own.
- The Rust toolchain does not properly support BPF exception handling
yet. The tables the selftests use are hand-written inline asm, which
the assembler turns into the same relocations the BPF AsmPrinter emits,
so libbpf and the kernel see an object indistinguishable from a
compiler-generated one.
[1] https://github.com/llvm/llvm-project/pull/192164
llvm commit 9d51c891b719 ("[BPF] Add exception handling support
with .bpf_cleanup section")
Changelog
=========
v3 -> v4:
- v3: https://lore.kernel.org/bpf/20260920054225.864535-1-yonghong.song@linux.dev/
- Rebase on bpf-next due to conflict.
- Reserve the throw-site spill area only in a (sub)program that calls
bpf_throw(): 40 bytes of stack per frame on x86-64, 80 on arm64.
- Bound the record count by the number of instructions in the program,
and name both in the message.
- Refuse a .bpf_cleanup section in libbpf whose record count cannot be
handed to the kernel as a count times a record size in an int.
- Fix the static linker's new bounds check, which could itself wrap, and
refuse a section too small to hold one field.
- WARN once if the body of bpf_unwind_resume() is ever reached, the way
bpf_throw() does where its exception callback should never return.
- Rename nr_pad_body to pad_body_bits, use BTF_ID_LIST_SINGLE, move
cleanup_pad out of insn_aux_data's bools, drop an arm64 include.
v2 -> v3:
- v2: https://lore.kernel.org/bpf/20260918044156.3283973-1-yonghong.song@linux.dev/
- Keep a landing pad's record when opt_remove_nops() deletes a pad that is
a nop, instead of dropping it after the verifier has already checked the
call site against it.
- Refuse a BPF_LD_[ABS|IND] in a pad body.
- Teach mark_chain_precision() about the throw-to-pad edge, which crosses
frames with no instruction to account for them.
- Refuse a bpf_unwind_resume() in any frame but the one whose landing pad
the walker entered.
- Check raw_data, alignment and bounds before the static linker writes
through a relocation in a non-executable section, which may be SHT_NOBITS.
v1 -> v2:
- v1: https://lore.kernel.org/bpf/20260917055645.3926444-1-yonghong.song@linux.dev/
- Consolidate all usages of kern_extern_name() in a single patch in libbpf.
- Avoid compiler warning and add proper cleanup_info_cnt guard in libbpf when
collecting .bpf_cleanup records.
- Add cleanup_info_cnt condition for emit_rel_store() with cleanup_info.
Yonghong Song (20):
bpf: Accept the compiler's exception cleanup table at program load
bpf: Add the bpf_unwind_resume() kfunc
bpf: Add lookups for exception cleanup resumes and landing pads
bpf: Prepare for an exception cleanup table before the CFG walk
bpf: Make exception landing pads reachable in the CFG
bpf: Explore the landing pads no call site reaches
bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch
bpf: Walk the exception unwind in the verifier
bpf: Refuse a private stack for a program with an exception cleanup
table
bpf: Dispatch exception cleanup pads from bpf_throw()
bpf, x86: Dispatch exception cleanup pads at run time
bpf, arm64: Dispatch exception cleanup pads at run time
libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc
libbpf: Add cleanup_info to bpf_prog_load_opts
libbpf: Collect .bpf_cleanup records and pass them to the kernel
libbpf: Carry the exception cleanup table through the light skeleton
libbpf: Let the static linker carry .bpf_cleanup relocations
selftests/bpf: Add an end-to-end .bpf_cleanup exception test
selftests/bpf: Cover the exception cleanup shapes the chain does not
reach
selftests/bpf: Load an exception cleanup program from a light skeleton
arch/arm64/net/Makefile | 2 +-
arch/arm64/net/bpf_cleanup_pad.S | 95 ++
arch/arm64/net/bpf_jit_comp.c | 91 +-
arch/x86/net/Makefile | 3 +
arch/x86/net/bpf_cleanup_pad.S | 80 ++
arch/x86/net/bpf_jit_comp.c | 81 +-
include/linux/bpf.h | 90 ++
include/linux/bpf_cleanup_abi.h | 16 +
include/linux/bpf_verifier.h | 17 +-
include/linux/filter.h | 2 +
include/uapi/linux/bpf.h | 9 +
kernel/bpf/Makefile | 2 +-
kernel/bpf/backtrack.c | 20 +-
kernel/bpf/cfg.c | 71 +-
kernel/bpf/check_btf.c | 151 +++
kernel/bpf/core.c | 35 +-
kernel/bpf/exception.c | 663 ++++++++++++
kernel/bpf/exception.h | 22 +
kernel/bpf/fixups.c | 166 ++-
kernel/bpf/helpers.c | 48 +
kernel/bpf/liveness.c | 20 +
kernel/bpf/states.c | 6 +
kernel/bpf/syscall.c | 2 +-
kernel/bpf/verifier.c | 176 +++-
tools/include/uapi/linux/bpf.h | 9 +
tools/lib/bpf/bpf.c | 6 +-
tools/lib/bpf/bpf.h | 7 +-
tools/lib/bpf/gen_loader.c | 29 +-
tools/lib/bpf/libbpf.c | 327 +++++-
tools/lib/bpf/libbpf_internal.h | 10 +
tools/lib/bpf/linker.c | 36 +-
tools/testing/selftests/bpf/Makefile.skel | 2 +-
.../selftests/bpf/exceptions_cleanup.h | 55 +
.../bpf/prog_tests/exceptions_cleanup.c | 495 +++++++++
.../selftests/bpf/progs/exceptions_cleanup.c | 157 +++
.../bpf/progs/exceptions_cleanup_ext_table.c | 48 +
.../bpf/progs/exceptions_cleanup_fail.c | 662 ++++++++++++
.../bpf/progs/exceptions_cleanup_freplace.c | 17 +
.../bpf/progs/exceptions_cleanup_light.c | 39 +
.../progs/exceptions_cleanup_pad_freplace.c | 17 +
.../bpf/progs/exceptions_cleanup_shapes.c | 982 ++++++++++++++++++
41 files changed, 4694 insertions(+), 72 deletions(-)
create mode 100644 arch/arm64/net/bpf_cleanup_pad.S
create mode 100644 arch/x86/net/bpf_cleanup_pad.S
create mode 100644 include/linux/bpf_cleanup_abi.h
create mode 100644 kernel/bpf/exception.c
create mode 100644 kernel/bpf/exception.h
create mode 100644 tools/testing/selftests/bpf/exceptions_cleanup.h
create mode 100644 tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_ext_table.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_fail.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_freplace.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_light.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_pad_freplace.c
create mode 100644 tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
--
2.53.0-Meta
next reply other threads:[~2026-09-21 21:00 UTC|newest]
Thread overview: 80+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-21 21:00 Yonghong Song [this message]
2026-09-21 21:00 ` [PATCH bpf-next v4 01/20] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-21 21:56 ` bot+bpf-ci
2026-09-22 3:27 ` Yonghong Song
2026-09-21 21:00 ` [PATCH bpf-next v4 02/20] bpf: Add the bpf_unwind_resume() kfunc Yonghong Song
2026-09-21 21:56 ` bot+bpf-ci
2026-09-22 3:31 ` Yonghong Song
2026-09-21 21:00 ` [PATCH bpf-next v4 03/20] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-22 4:04 ` Alexei Starovoitov
2026-09-22 5:28 ` Yonghong Song
2026-09-21 21:00 ` [PATCH bpf-next v4 04/20] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-09-22 18:27 ` Eduard Zingerman
2026-09-23 3:07 ` Yonghong Song
2026-09-23 3:54 ` Eduard Zingerman
2026-09-23 4:05 ` Yonghong Song
2026-09-21 21:00 ` [PATCH bpf-next v4 05/20] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 06/20] bpf: Explore the landing pads no call site reaches Yonghong Song
2026-09-21 23:58 ` Eduard Zingerman
2026-09-22 3:32 ` Yonghong Song
2026-09-22 4:10 ` Eduard Zingerman
2026-09-21 21:01 ` [PATCH bpf-next v4 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch Yonghong Song
2026-09-21 21:20 ` sashiko-bot
2026-09-22 3:39 ` Yonghong Song
2026-09-21 21:56 ` bot+bpf-ci
2026-09-22 3:44 ` Yonghong Song
2026-09-22 0:30 ` Eduard Zingerman
2026-09-22 3:45 ` Yonghong Song
2026-09-22 21:43 ` Eduard Zingerman
2026-09-23 3:11 ` Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 08/20] bpf: Walk the exception unwind in the verifier Yonghong Song
2026-09-21 21:40 ` sashiko-bot
2026-09-22 4:17 ` Yonghong Song
2026-09-21 21:56 ` bot+bpf-ci
2026-09-22 5:21 ` Yonghong Song
2026-09-22 4:08 ` Alexei Starovoitov
2026-09-22 5:25 ` Yonghong Song
2026-09-22 21:53 ` Eduard Zingerman
2026-09-23 3:18 ` Yonghong Song
2026-09-22 23:43 ` Eduard Zingerman
2026-09-23 3:21 ` Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 09/20] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 10/20] bpf: Dispatch exception cleanup pads from bpf_throw() Yonghong Song
2026-09-22 21:38 ` Eduard Zingerman
2026-09-23 3:22 ` Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 11/20] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 12/20] bpf, arm64: " Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 13/20] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-21 21:13 ` sashiko-bot
2026-09-21 21:01 ` [PATCH bpf-next v4 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-21 21:01 ` [PATCH bpf-next v4 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-21 21:20 ` sashiko-bot
2026-09-21 21:01 ` [PATCH bpf-next v4 16/20] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-21 21:02 ` [PATCH bpf-next v4 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-21 21:02 ` [PATCH bpf-next v4 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
2026-09-21 21:22 ` sashiko-bot
2026-09-22 5:26 ` Yonghong Song
2026-09-21 21:56 ` bot+bpf-ci
2026-09-21 21:02 ` [PATCH bpf-next v4 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
2026-09-21 21:19 ` sashiko-bot
2026-09-21 21:02 ` [PATCH bpf-next v4 20/20] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
2026-09-22 1:08 ` [PATCH bpf-next v4 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Eduard Zingerman
2026-09-22 2:16 ` Alexei Starovoitov
2026-09-22 2:31 ` Kumar Kartikeya Dwivedi
2026-09-22 21:44 ` Alexei Starovoitov
2026-09-23 4:36 ` Kumar Kartikeya Dwivedi
2026-09-23 4:54 ` Alexei Starovoitov
2026-09-23 5:20 ` Kumar Kartikeya Dwivedi
2026-09-23 6:16 ` Eduard Zingerman
2026-09-23 6:44 ` Kumar Kartikeya Dwivedi
2026-09-22 4:27 ` Eduard Zingerman
2026-09-22 21:47 ` Alexei Starovoitov
2026-09-22 23:08 ` Eduard Zingerman
2026-09-22 23:37 ` Alexei Starovoitov
2026-09-23 0:04 ` Eduard Zingerman
2026-09-23 19:04 ` Eduard Zingerman
2026-09-23 19:24 ` Andrii Nakryiko
2026-09-23 19:34 ` Kumar Kartikeya Dwivedi
2026-09-23 21:34 ` Alexei Starovoitov
2026-09-23 22:00 ` Eduard Zingerman
2026-09-23 23:22 ` Alexei Starovoitov
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260921210033.1715000-1-yonghong.song@linux.dev \
--to=yonghong.song@linux.dev \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=kernel-team@fb.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox