BPF List
 help / color / mirror / Atom feed
From: Yonghong Song <yonghong.song@linux.dev>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
	Andrii Nakryiko <andrii@kernel.org>,
	Daniel Borkmann <daniel@iogearbox.net>,
	Eduard Zingerman <eddyz87@gmail.com>,
	kernel-team@fb.com
Subject: [PATCH bpf-next v8 13/22] bpf, arm64: Dispatch exception cleanup pads at run time
Date: Thu,  1 Oct 2026 06:31:13 -0700	[thread overview]
Message-ID: <20261001133113.1342612-1-yonghong.song@linux.dev> (raw)
In-Reply-To: <20261001133006.1335369-1-yonghong.song@linux.dev>

The JIT half: build the native cleanup table from the JIT's byte offsets
once the image is final, and record the one epilogue so a frame the unwind
passes over can return through it. A pad head needs no BTI of its own: it
is only ever reached as a return address, and a return sets no BTYPE, so
no branch-target check is made.

The dispatch is arch_bpf_stack_walk_ra(). arm64's unwinder reads a frame's
return address from the frame record its callee pushed, so the walk hands
the unwind a copy and, when the unwind changes it, stores the new address
back into that record, the one the previous entry stepped through.

Writing it has to respect pointer authentication: a BPF prologue signs the
link register with PACIASP and the epilogue authenticates it, so what goes
back has to carry the same signature. Its modifier is the stack pointer the
owner was entered with, the record + 16 for a BPF prologue, and only a BPF
frame's record is written: bpf_unwind() leaves its own caller's return
address alone. Re-signing the address the unwinder stripped checks that
modifier against the slot before anything is signed with it.

Whether a slot is signed is asked of the build rather than read off the
value: CONFIG_ARM64_PTR_AUTH_KERNEL is what the prologue signs under.
Reading it off the value instead would take a signed address for an
unsigned one whenever its PAC equalled the bits stripping puts back. The
CPU has to implement address authentication too, since "pacia Xd, Xn" is
not in the HINT space and would be undefined without it.

Two frames are not redirected: the first, whose return into bpf_unwind()
comes out of the walk's own frame record, and one the function graph
tracer or a kretprobe has hooked, whose slot holds the trampoline rather
than the address the unwinder reports -- the walk stops there, with a
warning, since the BPF frames are then left returning to paths the
verifier never walked. The frame that called bpf_unwind() is left alone by
the generic code.

bpf_jit_supports_cleanup_pads() can now say yes. A shadow call stack does
not change that: only JITed frames' records are written, JITed code keeps
no x18 copy of its return address, and bpf_unwind()'s own return is never
rewritten.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 arch/arm64/kernel/stacktrace.c | 104 +++++++++++++++++++++++++++++++++
 arch/arm64/net/bpf_jit_comp.c  |  21 +++++++
 2 files changed, 125 insertions(+)

diff --git a/arch/arm64/kernel/stacktrace.c b/arch/arm64/kernel/stacktrace.c
index 3ebcf8c53fb0..66e3e2eefff4 100644
--- a/arch/arm64/kernel/stacktrace.c
+++ b/arch/arm64/kernel/stacktrace.c
@@ -445,6 +445,110 @@ noinline noinstr void arch_bpf_stack_walk(bool (*consume_entry)(void *cookie, u6
 	kunwind_stack_walk(arch_bpf_unwind_consume_entry, &data, current, NULL);
 }
 
+struct bpf_unwind_ra_consume_entry_data {
+	bool (*consume_entry)(void *cookie, u64 ip, u64 sp, u64 fp, u64 *ra);
+	void *cookie;
+	unsigned long record;
+	bool seen_first;
+};
+
+static u64 bpf_unwind_sign_ra(u64 ra, u64 modifier)
+{
+	asm volatile(ARM64_ASM_PREAMBLE
+		     ".arch_extension pauth\n"
+		     "	pacia %0, %1"
+		     : "+r" (ra) : "r" (modifier));
+	return ra;
+}
+
+/*
+ * PACIASP's modifier is the stack pointer the owner was entered with, the
+ * record + 16 for a BPF prologue. Only a BPF frame's record is rewritten --
+ * bpf_unwind() leaves its own caller's return address alone -- so that is
+ * the modifier; check it by re-signing @pc, which the unwinder stripped
+ * from @stored, before signing anything with it.
+ */
+static bool bpf_unwind_ra_modifier(unsigned long record, u64 stored, u64 pc,
+				   u64 *modifier)
+{
+	*modifier = record + sizeof(struct frame_record);
+	return bpf_unwind_sign_ra(pc, *modifier) == stored;
+}
+
+static bool bpf_unwind_store_ra(unsigned long record, u64 pc, u64 ra)
+{
+	struct frame_record *rec = (struct frame_record *)record;
+
+	/*
+	 * Whether the slot holds a signed address is a property of the build,
+	 * not one to be read off the value: a PAC can come out equal to the
+	 * bits stripping puts back, and a signed address would then be taken
+	 * for an unsigned one. What signs is CONFIG_ARM64_PTR_AUTH_KERNEL --
+	 * the prologue here, and -mbranch-protection for everything the
+	 * compiler emits.
+	 */
+	if (IS_ENABLED(CONFIG_ARM64_PTR_AUTH_KERNEL) &&
+	    system_supports_address_auth()) {
+		u64 stored = READ_ONCE(rec->lr);
+		u64 modifier;
+
+		if (WARN_ON_ONCE(!bpf_unwind_ra_modifier(record, stored, pc,
+							 &modifier)))
+			return false;
+		ra = bpf_unwind_sign_ra(ra, modifier);
+	}
+	WRITE_ONCE(rec->lr, ra);
+	return true;
+}
+
+static bool
+arch_bpf_unwind_ra_consume_entry(const struct kunwind_state *state, void *cookie)
+{
+	struct bpf_unwind_ra_consume_entry_data *data = cookie;
+	unsigned long record = data->record;
+	bool seen_first = data->seen_first;
+	u64 ra = state->common.pc;
+	bool cont;
+
+	/* The record this frame's return address will have come out of. */
+	data->record = state->common.fp;
+	data->seen_first = true;
+
+	/*
+	 * The first pc returns into bpf_unwind(), from this walk's own frame
+	 * record: not a BPF frame, and not one to redirect.
+	 */
+	if (!seen_first)
+		return true;
+	/*
+	 * A traced return: the slot holds the tracer's trampoline, not @pc.
+	 * Stopping leaves the BPF frames returning to paths the verifier
+	 * never walked, so warn.
+	 */
+	if (WARN_ON_ONCE(state->flags.fgraph || state->flags.kretprobe))
+		return false;
+
+	/* A consumer that stops still gets to redirect the frame it stopped on. */
+	cont = data->consume_entry(data->cookie, state->common.pc, 0,
+				   state->common.fp, &ra);
+	if (ra != state->common.pc &&
+	    !bpf_unwind_store_ra(record, state->common.pc, ra))
+		return false;
+	return cont;
+}
+
+noinline noinstr void arch_bpf_stack_walk_ra(bool (*consume_entry)(void *cookie, u64 ip, u64 sp,
+								   u64 fp, u64 *ra),
+					     void *cookie)
+{
+	struct bpf_unwind_ra_consume_entry_data data = {
+		.consume_entry = consume_entry,
+		.cookie = cookie,
+	};
+
+	kunwind_stack_walk(arch_bpf_unwind_ra_consume_entry, &data, current, NULL);
+}
+
 static const char *state_source_string(const struct kunwind_state *state)
 {
 	switch (state->source) {
diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
index 475e70653454..8ac98b024194 100644
--- a/arch/arm64/net/bpf_jit_comp.c
+++ b/arch/arm64/net/bpf_jit_comp.c
@@ -2423,6 +2423,17 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
 		 * reasons, expects to point to the next instruction)
 		 */
 		bpf_prog_update_insn_ptrs(prog, ctx.offset, ctx.ro_image);
+
+		/*
+		 * Same byte offsets, consumed by the bpf_unwind() walk:
+		 * turn the cleanup records into native address ranges now that
+		 * the image is final.
+		 */
+		bpf_exc_fill_native_ranges(prog, ctx.offset, ctx.ro_image);
+
+		/* Where an unwind sends a frame with no pad. */
+		prog->aux->epilogue_ip = (u64)ctx.ro_image +
+					 ctx.epilogue_offset * AARCH64_INSN_SIZE;
 out_off:
 		if (!ro_header && priv_stack_ptr) {
 			free_percpu(priv_stack_ptr);
@@ -3408,6 +3419,16 @@ bool bpf_jit_supports_exceptions(void)
 	return true;
 }
 
+bool bpf_jit_supports_cleanup_pads(void)
+{
+	/*
+	 * An unwind rewrites the return addresses in JITed frames' records,
+	 * which is what JITed code returns through, shadow call stack or not:
+	 * it keeps no x18 copy. bpf_unwind()'s own return is never rewritten.
+	 */
+	return true;
+}
+
 bool bpf_jit_supports_arena(void)
 {
 	return true;
-- 
2.53.0-Meta


  parent reply	other threads:[~2026-10-01 13:31 UTC|newest]

Thread overview: 50+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-10-01 13:48   ` sashiko-bot
2026-10-02 18:17     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 19:06     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier Yonghong Song
2026-10-01 13:50   ` sashiko-bot
2026-10-02 19:31     ` Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 20:49     ` Yonghong Song
2026-10-03 12:23   ` Alexei Starovoitov
2026-10-04 17:56     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 21:10     ` Yonghong Song
2026-10-03 12:25   ` Alexei Starovoitov
2026-10-04 17:59     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-10-03 12:25   ` Alexei Starovoitov
2026-10-04 18:26     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 10/22] bpf: Do not use a private stack for a program that can unwind Yonghong Song
2026-10-01 13:53   ` sashiko-bot
2026-10-02 21:38     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 21:48     ` Yonghong Song
2026-10-03 12:26   ` Alexei Starovoitov
2026-10-04 18:28     ` Yonghong Song
2026-10-04 18:29     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 12/22] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-10-01 13:49   ` sashiko-bot
2026-10-02 21:54     ` Yonghong Song
2026-10-01 13:31 ` Yonghong Song [this message]
2026-10-01 13:31 ` [PATCH bpf-next v8 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-10-01 13:46   ` sashiko-bot
2026-10-02 22:09     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
2026-10-01 13:32 ` [PATCH bpf-next v8 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261001133113.1342612-1-yonghong.song@linux.dev \
    --to=yonghong.song@linux.dev \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=kernel-team@fb.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox