BPF List
 help / color / mirror / Atom feed
From: Yonghong Song <yonghong.song@linux.dev>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
	Andrii Nakryiko <andrii@kernel.org>,
	Daniel Borkmann <daniel@iogearbox.net>,
	Eduard Zingerman <eddyz87@gmail.com>,
	kernel-team@fb.com
Subject: [PATCH bpf-next v7 12/22] bpf, x86: Dispatch exception cleanup pads at run time
Date: Mon, 28 Sep 2026 17:17:04 -0700	[thread overview]
Message-ID: <20260929001704.3251543-1-yonghong.song@linux.dev> (raw)
In-Reply-To: <20260929001601.3242665-1-yonghong.song@linux.dev>

arch_bpf_stack_walk_ra() is the ORC walk with the return-address slot
alongside each frame, which unwind_get_return_address_ptr() already hands
out; writing there is what redirects a frame to its landing pad. The
feature is gated on CONFIG_UNWINDER_ORC for the same reason
arch_bpf_stack_walk() is -- there is no other unwinder here to ask.

A frame whose return a tracer has hooked is left alone, and the walk stops
there. The unwinder recovers the address such a frame will really return
to, but the slot still holds the function graph or kretprobe trampoline,
so writing it would skip the trampoline and leave its entry for the next
hooked return to pop. x86 has no flag saying a frame was hooked, so the
recovered address and the slot are compared instead.

aux->epilogue_ip comes for free: the JIT already emits one epilogue per
(sub)program and every other exit jumps to it, so the offset it keeps as
ctx->cleanup_addr, as a native address, is it. That field has always been
the epilogue and has nothing to do with the cleanup pads despite the name.

The rest is bookkeeping: build the native cleanup table from the JIT's
addrs[] once the image is final. A pad head needs no ENDBR of its own: it
is only ever reached as a return address, and IBT checks indirect jumps
and calls, not returns.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 arch/x86/net/bpf_jit_comp.c | 42 +++++++++++++++++++++++++++++++++++++
 1 file changed, 42 insertions(+)

diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index 6c7a0578760e..d4feade5b5c7 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -3278,6 +3278,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
 			seen_exit = true;
 			/* Update cleanup_addr */
 			ctx->cleanup_addr = proglen;
+			/* Where an unwind sends a frame with no pad. */
+			bpf_prog->aux->epilogue_ip = (u64)image + proglen;
 			if (bpf_prog_was_classic(bpf_prog) &&
 			    !ns_capable_noaudit(&init_user_ns, CAP_SYS_ADMIN)) {
 				if (emit_spectre_bhb_barrier(&prog, ip, bpf_prog))
@@ -4455,6 +4457,13 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
 		 */
 		bpf_prog_update_insn_ptrs(prog, addrs, image);
 
+		/*
+		 * Same mapping, consumed by the bpf_unwind() walk:
+		 * turn the cleanup records into native address ranges now
+		 * that the image is final.
+		 */
+		bpf_exc_fill_native_ranges(prog, addrs, image);
+
 		/*
 		 * ctx.prog_offset is used when CFI preambles put code *before*
 		 * the function. See emit_cfi(). For FineIBT specifically this code
@@ -4593,6 +4602,11 @@ bool bpf_jit_supports_exceptions(void)
 	return IS_ENABLED(CONFIG_UNWINDER_ORC);
 }
 
+bool bpf_jit_supports_cleanup_pads(void)
+{
+	return IS_ENABLED(CONFIG_UNWINDER_ORC);
+}
+
 bool bpf_jit_supports_private_stack(void)
 {
 	return true;
@@ -4614,6 +4628,34 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
 #endif
 }
 
+void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
+			    void *cookie)
+{
+#if defined(CONFIG_UNWINDER_ORC)
+	struct unwind_state state;
+	unsigned long addr, *ra;
+
+	for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
+	     unwind_next_frame(&state)) {
+		addr = unwind_get_return_address(&state);
+		ra = unwind_get_return_address_ptr(&state);
+		if (!addr || !ra)
+			break;
+		/*
+		 * A traced return: the unwinder recovered @addr from under a
+		 * function graph or kretprobe trampoline, which is what the
+		 * slot itself still holds. Writing there would skip the
+		 * trampoline and leave its entry for the next hooked return
+		 * to pop.
+		 */
+		if (READ_ONCE_NOCHECK(*ra) != addr)
+			break;
+		if (!consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
+			break;
+	}
+#endif
+}
+
 void bpf_arch_poke_desc_update(struct bpf_jit_poke_descriptor *poke,
 			       struct bpf_prog *new, struct bpf_prog *old)
 {
-- 
2.53.0-Meta


  parent reply	other threads:[~2026-09-29  0:17 UTC|newest]

Thread overview: 46+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29  0:16 [PATCH bpf-next v7 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-09-29  0:16 ` [PATCH bpf-next v7 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-09-29  0:16 ` [PATCH bpf-next v7 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-29  0:16 ` [PATCH bpf-next v7 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-09-29  0:16 ` [PATCH bpf-next v7 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-29  0:33   ` sashiko-bot
2026-09-29 21:58     ` Yonghong Song
2026-09-29  0:16 ` [PATCH bpf-next v7 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-09-29  0:31   ` sashiko-bot
2026-09-29 22:04     ` Yonghong Song
2026-09-29  0:16 ` [PATCH bpf-next v7 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-29  0:16 ` [PATCH bpf-next v7 07/22] bpf: Resume a covered call at its landing pad Yonghong Song
2026-09-29  0:31   ` sashiko-bot
2026-09-30  0:28     ` Yonghong Song
2026-09-29  0:16 ` [PATCH bpf-next v7 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
2026-09-29  0:36   ` sashiko-bot
2026-09-30  1:09     ` Yonghong Song
2026-09-29  0:52   ` bot+bpf-ci
2026-09-30  1:10     ` Yonghong Song
2026-09-29  0:16 ` [PATCH bpf-next v7 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-09-29  0:16 ` [PATCH bpf-next v7 10/22] bpf: Refuse a private stack for a program that can unwind Yonghong Song
2026-09-29  0:16 ` [PATCH bpf-next v7 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
2026-09-29  1:14   ` bot+bpf-ci
2026-09-30  1:18     ` Yonghong Song
2026-09-29  0:17 ` Yonghong Song [this message]
2026-09-29  0:30   ` [PATCH bpf-next v7 12/22] bpf, x86: Dispatch exception cleanup pads at run time sashiko-bot
2026-09-30  1:34     ` Yonghong Song
2026-09-29  0:17 ` [PATCH bpf-next v7 13/22] bpf, arm64: " Yonghong Song
2026-09-29  1:14   ` bot+bpf-ci
2026-09-29  0:17 ` [PATCH bpf-next v7 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-29  0:17 ` [PATCH bpf-next v7 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-29  0:17 ` [PATCH bpf-next v7 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-29  0:17 ` [PATCH bpf-next v7 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-29  0:17 ` [PATCH bpf-next v7 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-29  0:17 ` [PATCH bpf-next v7 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
2026-09-29  0:52   ` bot+bpf-ci
2026-09-30  1:42     ` Yonghong Song
2026-09-29  0:17 ` [PATCH bpf-next v7 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-09-29  0:52   ` bot+bpf-ci
2026-09-30  1:46     ` Yonghong Song
2026-09-29  0:17 ` [PATCH bpf-next v7 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
2026-09-29  0:52   ` bot+bpf-ci
2026-09-30  2:19     ` Yonghong Song
2026-09-29  0:17 ` [PATCH bpf-next v7 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
2026-09-29  0:52   ` bot+bpf-ci
2026-09-30  3:12     ` Yonghong Song

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260929001704.3251543-1-yonghong.song@linux.dev \
    --to=yonghong.song@linux.dev \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=kernel-team@fb.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox