From: Yonghong Song <yonghong.song@linux.dev>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
Andrii Nakryiko <andrii@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Eduard Zingerman <eddyz87@gmail.com>,
kernel-team@fb.com
Subject: [PATCH bpf-next v7 13/22] bpf, arm64: Dispatch exception cleanup pads at run time
Date: Mon, 28 Sep 2026 17:17:09 -0700 [thread overview]
Message-ID: <20260929001709.3252802-1-yonghong.song@linux.dev> (raw)
In-Reply-To: <20260929001601.3242665-1-yonghong.song@linux.dev>
The JIT half: build the native cleanup table from the JIT's byte offsets
once the image is final, and record the one epilogue so a frame the unwind
passes over can return through it. A pad head needs no BTI of its own: it
is only ever reached as a return address, and a return sets no BTYPE, so
no branch-target check is made.
The dispatch is arch_bpf_stack_walk_ra(), which hands the unwind the slot a
frame's return address came out of rather than just the address. arm64's
unwinder reads that address from the frame record the callee pushed, so the
slot belongs to the record the previous entry stepped through.
Writing it has to respect pointer authentication: a BPF prologue signs the
link register with PACIASP and the epilogue authenticates it, so what goes
back has to carry the same signature. Its modifier is the stack pointer the
owner was entered with, which is not known here, so recover it by
re-signing the address the unwinder stripped until that matches the slot.
Whether a slot is signed is asked of the build rather than read off the
value: CONFIG_ARM64_PTR_AUTH_KERNEL is what the prologue signs under and
what -mbranch-protection is added for. Reading it off the value instead
would take a signed address for an unsigned one whenever its PAC equalled
the bits stripping puts back. The CPU has to implement address
authentication too, since "pacia Xd, Xn" is not in the HINT space and
would be undefined without it.
Two frames are not redirected: the walk's own first frame, not returning
anywhere yet, and one the function graph tracer or a kretprobe has hooked,
whose slot holds the trampoline rather than the address the unwinder
reports -- the walk stops there.
bpf_jit_supports_cleanup_pads() can now say yes, except where a shadow
call stack is in use. JITed code restores x30 from the frame record this
walk rewrites, but bpf_unwind() is C and returns from its x18 copy
instead, so the frame that called it would carry on as though nothing had
happened while the frames above it resumed at their pads.
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
arch/arm64/kernel/stacktrace.c | 103 +++++++++++++++++++++++++++++++++
arch/arm64/net/bpf_jit_comp.c | 26 +++++++++
2 files changed, 129 insertions(+)
diff --git a/arch/arm64/kernel/stacktrace.c b/arch/arm64/kernel/stacktrace.c
index 3ebcf8c53fb0..1e46a22cafbd 100644
--- a/arch/arm64/kernel/stacktrace.c
+++ b/arch/arm64/kernel/stacktrace.c
@@ -445,6 +445,109 @@ noinline noinstr void arch_bpf_stack_walk(bool (*consume_entry)(void *cookie, u6
kunwind_stack_walk(arch_bpf_unwind_consume_entry, &data, current, NULL);
}
+struct bpf_unwind_ra_consume_entry_data {
+ bool (*consume_entry)(void *cookie, u64 ip, u64 sp, u64 fp, u64 *ra);
+ void *cookie;
+ unsigned long record;
+ bool seen_first;
+};
+
+static u64 bpf_unwind_sign_ra(u64 ra, u64 modifier)
+{
+ asm volatile(ARM64_ASM_PREAMBLE
+ ".arch_extension pauth\n"
+ " pacia %0, %1"
+ : "+r" (ra) : "r" (modifier));
+ return ra;
+}
+
+/*
+ * PACIASP's modifier is the stack pointer the owner was entered with: record
+ * + 16 for a BPF prologue, but further up for bpf_unwind()'s own C frame.
+ * Recognise it by re-signing @pc, which the unwinder stripped from @stored.
+ */
+static bool bpf_unwind_ra_modifier(unsigned long record, unsigned long caller_fp,
+ u64 stored, u64 pc, u64 *modifier)
+{
+ u64 m;
+
+ for (m = record + sizeof(struct frame_record); m <= caller_fp; m += 16) {
+ if (bpf_unwind_sign_ra(pc, m) == stored) {
+ *modifier = m;
+ return true;
+ }
+ }
+ return false;
+}
+
+static bool bpf_unwind_store_ra(unsigned long record, unsigned long caller_fp,
+ u64 pc, u64 ra)
+{
+ struct frame_record *rec = (struct frame_record *)record;
+
+ /*
+ * Whether the slot holds a signed address is a property of the build,
+ * not one to be read off the value: a PAC can come out equal to the
+ * bits stripping puts back, and a signed address would then be taken
+ * for an unsigned one. What signs is CONFIG_ARM64_PTR_AUTH_KERNEL --
+ * the prologue here, and -mbranch-protection for everything the
+ * compiler emits.
+ */
+ if (IS_ENABLED(CONFIG_ARM64_PTR_AUTH_KERNEL) &&
+ system_supports_address_auth()) {
+ u64 stored = READ_ONCE(rec->lr);
+ u64 modifier;
+
+ if (WARN_ON_ONCE(!bpf_unwind_ra_modifier(record, caller_fp,
+ stored, pc, &modifier)))
+ return false;
+ ra = bpf_unwind_sign_ra(ra, modifier);
+ }
+ WRITE_ONCE(rec->lr, ra);
+ return true;
+}
+
+static bool
+arch_bpf_unwind_ra_consume_entry(const struct kunwind_state *state, void *cookie)
+{
+ struct bpf_unwind_ra_consume_entry_data *data = cookie;
+ unsigned long record = data->record;
+ bool seen_first = data->seen_first;
+ u64 ra = state->common.pc;
+ bool cont;
+
+ /* The record this frame's return address will have come out of. */
+ data->record = state->common.fp;
+ data->seen_first = true;
+
+ /* The first pc is where the walk runs, not an address it returns to. */
+ if (!seen_first)
+ return true;
+ /* A traced return: the slot holds the tracer's trampoline, not @pc. */
+ if (state->flags.fgraph || state->flags.kretprobe)
+ return false;
+
+ /* A consumer that stops still gets to redirect the frame it stopped on. */
+ cont = data->consume_entry(data->cookie, state->common.pc, 0,
+ state->common.fp, &ra);
+ if (ra != state->common.pc &&
+ !bpf_unwind_store_ra(record, state->common.fp, state->common.pc, ra))
+ return false;
+ return cont;
+}
+
+noinline noinstr void arch_bpf_stack_walk_ra(bool (*consume_entry)(void *cookie, u64 ip, u64 sp,
+ u64 fp, u64 *ra),
+ void *cookie)
+{
+ struct bpf_unwind_ra_consume_entry_data data = {
+ .consume_entry = consume_entry,
+ .cookie = cookie,
+ };
+
+ kunwind_stack_walk(arch_bpf_unwind_ra_consume_entry, &data, current, NULL);
+}
+
static const char *state_source_string(const struct kunwind_state *state)
{
switch (state->source) {
diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
index 475e70653454..e483e1e7a2d4 100644
--- a/arch/arm64/net/bpf_jit_comp.c
+++ b/arch/arm64/net/bpf_jit_comp.c
@@ -10,10 +10,12 @@
#include <linux/arm-smccc.h>
#include <linux/bitfield.h>
#include <linux/bpf.h>
+#include <linux/bpf_verifier.h>
#include <linux/cfi.h>
#include <linux/filter.h>
#include <linux/memory.h>
#include <linux/printk.h>
+#include <linux/scs.h>
#include <linux/slab.h>
#include <asm/asm-extable.h>
@@ -2423,6 +2425,17 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
* reasons, expects to point to the next instruction)
*/
bpf_prog_update_insn_ptrs(prog, ctx.offset, ctx.ro_image);
+
+ /*
+ * Same byte offsets, consumed by the bpf_unwind() walk:
+ * turn the cleanup records into native address ranges now that
+ * the image is final.
+ */
+ bpf_exc_fill_native_ranges(prog, ctx.offset, ctx.ro_image);
+
+ /* Where an unwind sends a frame with no pad. */
+ prog->aux->epilogue_ip = (u64)ctx.ro_image +
+ ctx.epilogue_offset * AARCH64_INSN_SIZE;
out_off:
if (!ro_header && priv_stack_ptr) {
free_percpu(priv_stack_ptr);
@@ -3408,6 +3421,19 @@ bool bpf_jit_supports_exceptions(void)
return true;
}
+bool bpf_jit_supports_cleanup_pads(void)
+{
+ /*
+ * An unwind redirects a frame by rewriting the frame record its
+ * callee's return address came out of. JITed code restores x30 from
+ * there, but bpf_unwind() is C: with a shadow call stack it returns
+ * from the x18 copy instead, so the frame that called it would keep
+ * going as if nothing had happened while the frames above it resumed
+ * at their pads.
+ */
+ return !scs_is_enabled();
+}
+
bool bpf_jit_supports_arena(void)
{
return true;
--
2.53.0-Meta
next prev parent reply other threads:[~2026-09-29 0:17 UTC|newest]
Thread overview: 46+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-29 0:33 ` sashiko-bot
2026-09-29 21:58 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-09-29 0:31 ` sashiko-bot
2026-09-29 22:04 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 07/22] bpf: Resume a covered call at its landing pad Yonghong Song
2026-09-29 0:31 ` sashiko-bot
2026-09-30 0:28 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
2026-09-29 0:36 ` sashiko-bot
2026-09-30 1:09 ` Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:10 ` Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 10/22] bpf: Refuse a private stack for a program that can unwind Yonghong Song
2026-09-29 0:16 ` [PATCH bpf-next v7 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
2026-09-29 1:14 ` bot+bpf-ci
2026-09-30 1:18 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-29 0:30 ` sashiko-bot
2026-09-30 1:34 ` Yonghong Song
2026-09-29 0:17 ` Yonghong Song [this message]
2026-09-29 1:14 ` [PATCH bpf-next v7 13/22] bpf, arm64: " bot+bpf-ci
2026-09-29 0:17 ` [PATCH bpf-next v7 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:42 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 1:46 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 2:19 ` Yonghong Song
2026-09-29 0:17 ` [PATCH bpf-next v7 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
2026-09-29 0:52 ` bot+bpf-ci
2026-09-30 3:12 ` Yonghong Song
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929001709.3252802-1-yonghong.song@linux.dev \
--to=yonghong.song@linux.dev \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=kernel-team@fb.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.