All of lore.kernel.org
 help / color / mirror / Atom feed
From: Yonghong Song <yonghong.song@linux.dev>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
	Andrii Nakryiko <andrii@kernel.org>,
	Daniel Borkmann <daniel@iogearbox.net>,
	Eduard Zingerman <eddyz87@gmail.com>,
	kernel-team@fb.com
Subject: [PATCH bpf-next v8 12/22] bpf, x86: Dispatch exception cleanup pads at run time
Date: Thu,  1 Oct 2026 06:31:08 -0700	[thread overview]
Message-ID: <20261001133108.1341965-1-yonghong.song@linux.dev> (raw)
In-Reply-To: <20261001133006.1335369-1-yonghong.song@linux.dev>

arch_bpf_stack_walk_ra() is the ORC walk with the return-address slot
alongside each frame, which unwind_get_return_address_ptr() already hands
out; writing there is what redirects a frame to its landing pad. The
feature is gated on CONFIG_UNWINDER_ORC for the same reason
arch_bpf_stack_walk() is -- there is no other unwinder here to ask -- and
bpf_jit_supports_cleanup_pads() says yes wherever it is built.

A frame whose return a tracer has hooked is left alone, and the walk stops
there. The unwinder recovers the address such a frame will really return
to, but the slot still holds the function graph or kretprobe trampoline,
so writing it would skip the trampoline and leave its entry for the next
hooked return to pop. x86 has no flag saying a frame was hooked, so the
recovered address and the slot are compared instead. Stopping leaves the
BPF frames returning to paths the verifier never walked, so it warns.

aux->epilogue_ip comes for free: the JIT already emits one epilogue per
(sub)program and every other exit jumps to it, so the offset it keeps as
ctx->cleanup_addr, as a native address, is it. That field has always been
the epilogue and has nothing to do with the cleanup pads despite the name.

The rest is bookkeeping: build the native cleanup table from the JIT's
addrs[] once the image is final. A pad head needs no ENDBR of its own: it
is only ever reached as a return address, and IBT checks indirect jumps
and calls, not returns.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 arch/x86/net/bpf_jit_comp.c | 40 +++++++++++++++++++++++++++++++++++++
 1 file changed, 40 insertions(+)

diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index 6c7a0578760e..544e8fd759ad 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -3278,6 +3278,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
 			seen_exit = true;
 			/* Update cleanup_addr */
 			ctx->cleanup_addr = proglen;
+			/* Where an unwind sends a frame with no pad. */
+			bpf_prog->aux->epilogue_ip = (u64)image + proglen;
 			if (bpf_prog_was_classic(bpf_prog) &&
 			    !ns_capable_noaudit(&init_user_ns, CAP_SYS_ADMIN)) {
 				if (emit_spectre_bhb_barrier(&prog, ip, bpf_prog))
@@ -4455,6 +4457,13 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
 		 */
 		bpf_prog_update_insn_ptrs(prog, addrs, image);
 
+		/*
+		 * Same mapping, consumed by the bpf_unwind() walk:
+		 * turn the cleanup records into native address ranges now
+		 * that the image is final.
+		 */
+		bpf_exc_fill_native_ranges(prog, addrs, image);
+
 		/*
 		 * ctx.prog_offset is used when CFI preambles put code *before*
 		 * the function. See emit_cfi(). For FineIBT specifically this code
@@ -4593,6 +4602,11 @@ bool bpf_jit_supports_exceptions(void)
 	return IS_ENABLED(CONFIG_UNWINDER_ORC);
 }
 
+bool bpf_jit_supports_cleanup_pads(void)
+{
+	return IS_ENABLED(CONFIG_UNWINDER_ORC);
+}
+
 bool bpf_jit_supports_private_stack(void)
 {
 	return true;
@@ -4614,6 +4628,32 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
 #endif
 }
 
+void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
+			    void *cookie)
+{
+#if defined(CONFIG_UNWINDER_ORC)
+	struct unwind_state state;
+	unsigned long addr, *ra;
+
+	for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
+	     unwind_next_frame(&state)) {
+		addr = unwind_get_return_address(&state);
+		ra = unwind_get_return_address_ptr(&state);
+		if (!addr || !ra)
+			break;
+		/*
+		 * A traced return: the slot holds a function graph or kretprobe
+		 * trampoline, not @addr, so it cannot be rewritten. Stopping
+		 * leaves BPF frames returning to unverified paths, so warn.
+		 */
+		if (WARN_ON_ONCE(READ_ONCE_NOCHECK(*ra) != addr))
+			break;
+		if (!consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
+			break;
+	}
+#endif
+}
+
 void bpf_arch_poke_desc_update(struct bpf_jit_poke_descriptor *poke,
 			       struct bpf_prog *new, struct bpf_prog *old)
 {
-- 
2.53.0-Meta


  parent reply	other threads:[~2026-10-01 13:31 UTC|newest]

Thread overview: 50+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-01 13:30 [PATCH bpf-next v8 00/22] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 01/22] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 02/22] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 03/22] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 04/22] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-10-01 13:48   ` sashiko-bot
2026-10-02 18:17     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 05/22] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 19:06     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 06/22] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 07/22] bpf: Follow an unwind to its landing pad in the verifier Yonghong Song
2026-10-01 13:50   ` sashiko-bot
2026-10-02 19:31     ` Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 20:49     ` Yonghong Song
2026-10-03 12:23   ` Alexei Starovoitov
2026-10-04 17:56     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 08/22] bpf: Require an unwind to leave a frame holding what it entered with Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 21:10     ` Yonghong Song
2026-10-03 12:25   ` Alexei Starovoitov
2026-10-04 17:59     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 09/22] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-10-03 12:25   ` Alexei Starovoitov
2026-10-04 18:26     ` Yonghong Song
2026-10-01 13:30 ` [PATCH bpf-next v8 10/22] bpf: Do not use a private stack for a program that can unwind Yonghong Song
2026-10-01 13:53   ` sashiko-bot
2026-10-02 21:38     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 11/22] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
2026-10-01 14:31   ` bot+bpf-ci
2026-10-02 21:48     ` Yonghong Song
2026-10-03 12:26   ` Alexei Starovoitov
2026-10-04 18:28     ` Yonghong Song
2026-10-04 18:29     ` Yonghong Song
2026-10-01 13:31 ` Yonghong Song [this message]
2026-10-01 13:49   ` [PATCH bpf-next v8 12/22] bpf, x86: Dispatch exception cleanup pads at run time sashiko-bot
2026-10-02 21:54     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 13/22] bpf, arm64: " Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 14/22] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 15/22] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-10-01 13:46   ` sashiko-bot
2026-10-02 22:09     ` Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 16/22] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 17/22] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 18/22] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 19/22] selftests/bpf: Add end-to-end and negative .bpf_cleanup exception tests Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 20/22] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-10-01 13:31 ` [PATCH bpf-next v8 21/22] selftests/bpf: Cover more accepted .bpf_cleanup exception shapes Yonghong Song
2026-10-01 13:32 ` [PATCH bpf-next v8 22/22] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261001133108.1341965-1-yonghong.song@linux.dev \
    --to=yonghong.song@linux.dev \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=kernel-team@fb.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.