BPF List
 help / color / mirror / Atom feed
From: Yonghong Song <yonghong.song@linux.dev>
To: bot+bpf-ci@kernel.org, bpf@vger.kernel.org
Cc: ast@kernel.org, andrii@kernel.org, daniel@iogearbox.net,
	eddyz87@gmail.com, kernel-team@fb.com, martin.lau@kernel.org,
	mason@kernel.org, ihor.solodrai@linux.dev
Subject: Re: [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time
Date: Sun, 27 Sep 2026 20:10:36 -0700	[thread overview]
Message-ID: <f975aa30-a056-4f48-8f65-abcff0e7cc70@linux.dev> (raw)
In-Reply-To: <699b29ac22154d64cba9488294b7931090b0c50650e99e555f278a1da854647e@mail.kernel.org>



On 9/27/26 1:39 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git arch/x86/net/bpf_jit_comp.c arch/x86/net/bpf_jit_comp.c
>> index 6c7a0578760e..fb7e8ca1aab2 100644
>> --- arch/x86/net/bpf_jit_comp.c
>> +++ arch/x86/net/bpf_jit_comp.c
>> @@ -2150,7 +2150,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
>>   			dst_reg = X86_REG_R9;
>>   	}
>>
>> -	if (bpf_insn_is_indirect_target(env, bpf_prog, i - 1))
>> +	if (bpf_insn_is_indirect_target(env, bpf_prog, i - 1) ||
>> +	    bpf_exc_insn_is_pad(env, bpf_prog, i - 1))
>>   		EMIT_ENDBR();
> Landing pads are only ever entered by a ret whose stack slot
> bpf_unwind_rewrite() rewrote. Nothing jumps or calls to a pad indirectly
> since bpf_unwind_resume() is lowered to 'r0 = 0; exit', and
> bpf_exc_check_insn() forbids gotox inside a pad. IBT only checks
> indirect JMP/CALL targets, not RET targets, so the ENDBR at each pad
> head should not be needed.

Yes, you are right. ENDBR is not needed here.

>
> The same patch sends pad-less frames to aux->epilogue_ip through the
> same kind of rewritten ret, and emits no ENDBR there. That shows the pad
> ENDBR does no functional work. Each extra ENDBR in the middle of a JITed
> function is one more valid IBT landing site in executable kernel memory,
> and the kernel normally keeps those to a minimum (objtool seals unused
> ENDBRs).
>
> Is there an indirect-branch path to the pad that is not visible here?
>
>>   	ip = image + addrs[i - 1] + (prog - temp);
> [ ... ]
>
>> @@ -3278,6 +3279,8 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
>>   		seen_exit = true;
>>   		/* Update cleanup_addr */
>>   		ctx->cleanup_addr = proglen;
>> +		/* Where an unwind sends a frame with no pad. */
>> +		bpf_prog->aux->epilogue_ip = (u64)image + proglen;
> Eduard Zingerman raised an optimization suggestion in v5: use register
> liveness information at landing pads to avoid spilling all registers
> unnecessarily. The suggestion was to use register liveness information
> at call sites and collect a union of registers alive at landing pads for
> all callers of this particular function.
>
> You acknowledged and said you would try to fix it to avoid saving
> unnecessary callee saved registers. However, the current v6
> implementation still appears to unconditionally spill registers for
> exception_boundary programs without optimizing based on which registers
> are actually live at landing pads.
>
> Note: v6 significantly changed the implementation approach from v5, so
> this suggestion may need re-evaluation in the context of the new design.
> Does this optimization still apply, or does the new design address this
> differently?

Alexei suggested a new approach in v6.

>
>>   		if (bpf_prog_was_classic(bpf_prog) &&
>>   		    !ns_capable_noaudit(&init_user_ns, CAP_SYS_ADMIN)) {
>>   			if (emit_spectre_bhb_barrier(&prog, ip, bpf_prog))
> [ ... ]
>
>> @@ -4614,6 +4629,24 @@ void arch_bpf_stack_walk(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp
>>   #endif
>>   }
>>
>> +void arch_bpf_stack_walk_ra(bool (*consume_fn)(void *cookie, u64 ip, u64 sp, u64 bp, u64 *ra),
>> +			    void *cookie)
>> +{
>> +#if defined(CONFIG_UNWINDER_ORC)
>> +	struct unwind_state state;
>> +	unsigned long addr, *ra;
>> +
>> +	for (unwind_start(&state, current, NULL, NULL); !unwind_done(&state);
>> +	     unwind_next_frame(&state)) {
>> +		addr = unwind_get_return_address(&state);
>> +		ra = unwind_get_return_address_ptr(&state);
>> +		if (!addr || !ra ||
>> +		    !consume_fn(cookie, (u64)addr, (u64)state.sp, (u64)state.bp, (u64 *)ra))
>> +			break;
>> +	}
>> +#endif
>> +}
> arch_bpf_stack_walk_ra() pairs the unwinder's recovered return address
> with the raw stack slot. On x86 ORC, unwind_next_frame() does
>
> arch/x86/kernel/unwind_orc.c:unwind_next_frame() {
>      ...
>      state->ip = unwind_recover_ret_addr(state, state->ip, (unsigned long *)ip_p);
>      ...
> }
>
> which turns return_to_handler (function_graph / fprobe exit on fgraph)
> or the rethook trampoline (kretprobe) back into the original caller
> address. But unwind_get_return_address_ptr() returns `(unsigned long
> *)state->sp - 1`, i.e. ip_p itself, and that slot still holds the
> tracer's trampoline address. Nothing checks that *ra == addr before
> handing ra to the consumer.
>
> bpf_unwind() is a normal __bpf_kfunc in kernel/bpf/helpers.c with no
> notrace, so function_graph, kretprobe:bpf_unwind, or a kprobe.multi
> return probe can hook its return. In that case the prologue replaces the
> slot holding bpf_unwind's return into the BPF program. Then:
>
>    bpf_unwind() -> arch_bpf_stack_walk_ra() -> bpf_unwind_rewrite()
>      cnt == 1 frame, and the bpf_unwind call is covered by a cleanup record
>      (mark_call_sites() marks bpf_is_unwind_kfunc() calls)
>      -> *ra = rec->pad;   /* overwrites return_to_handler / rethook trampoline */
>
> bpf_unwind() then returns straight into the pad and skips the tracer's
> return trampoline. Its fgraph ret_stack entry (or rethook node) is left
> behind. The next hooked function further up that returns through the
> trampoline pops that stale entry.
>
> For fgraph, ftrace_pop_return_trace() has no frame-pointer check on x86
> (HAVE_FUNCTION_GRAPH_FP_TEST is not defined there). For rethook,
> __rethook_find_ret_addr() picks the top node. Either way execution jumps
> to the old BPF address after the bpf_unwind call, running on a different
> function's stack, which leads to an oops or arbitrary control flow.
>
> The arm64 version of this same walker, added later in the series
> (8bd2af30a28e "bpf, arm64: Dispatch exception cleanup pads at run
> time"), handles exactly this case:
>
> arch/arm64/net/bpf_jit_comp.c:arch_bpf_stack_walk_ra() {
>      ...
>      if (state->flags.fgraph || state->flags.kretprobe)
>          return false;
>      ...
> }
>
> Its changelog says: "a frame whose return the function graph tracer or a
> kretprobe has hooked holds the tracer's trampoline in its slot rather
> than the address the unwinder reports, so the walk stops there." The x86
> walker has no such check, and no later commit in the series touches
> arch/x86/net/bpf_jit_comp.c.
>
> Should the x86 version also check state flags or compare
> READ_ONCE_NOCHECK(*ra) against the value the unwinder read, and stop the
> walk when they differ?

Okay, will do this since arm64 has a different guard.

>
>> +
>>   void bpf_arch_poke_desc_update(struct bpf_jit_poke_descriptor *poke,
>>   			       struct bpf_prog *new, struct bpf_prog *old)
>>   {
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430


  reply	other threads:[~2026-09-28  3:10 UTC|newest]

Thread overview: 56+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 01/21] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 02/21] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 03/21] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 04/21] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-09-26  5:16   ` sashiko-bot
2026-09-26 23:54     ` Yonghong Song
2026-09-27 20:39   ` bot+bpf-ci
2026-09-28  0:01     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-26  5:21   ` sashiko-bot
2026-09-27  0:02     ` Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  0:12     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad Yonghong Song
2026-09-26  5:15   ` sashiko-bot
2026-09-26  8:21     ` Alexei Starovoitov
2026-09-27  0:04       ` Yonghong Song
2026-09-27  0:41     ` Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  0:17     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-09-26  5:17   ` sashiko-bot
2026-09-27  3:06     ` Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  0:29     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 09/21] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 10/21] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  1:08     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-26  5:15   ` sashiko-bot
2026-09-27  4:35     ` Yonghong Song
2026-09-27 20:39   ` bot+bpf-ci
2026-09-28  3:10     ` Yonghong Song [this message]
2026-09-26  5:01 ` [PATCH bpf-next v6 12/21] bpf, arm64: " Yonghong Song
2026-09-26  5:14   ` sashiko-bot
2026-09-27 20:40   ` bot+bpf-ci
2026-09-26  5:01 ` [PATCH bpf-next v6 13/21] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 14/21] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 15/21] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-27 20:39   ` bot+bpf-ci
2026-09-28  3:28     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 16/21] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 17/21] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 18/21] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-09-26  5:18   ` sashiko-bot
2026-09-27  4:58     ` Yonghong Song
2026-09-27 20:24   ` bot+bpf-ci
2026-09-28  3:36     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 20/21] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  3:49     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 21/21] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=f975aa30-a056-4f48-8f65-abcff0e7cc70@linux.dev \
    --to=yonghong.song@linux.dev \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bot+bpf-ci@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=ihor.solodrai@linux.dev \
    --cc=kernel-team@fb.com \
    --cc=martin.lau@kernel.org \
    --cc=mason@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox