BPF List
 help / color / mirror / Atom feed
From: Yonghong Song <yonghong.song@linux.dev>
To: sashiko-reviews@lists.linux.dev
Cc: bpf@vger.kernel.org
Subject: Re: [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad
Date: Sat, 26 Sep 2026 17:41:20 -0700	[thread overview]
Message-ID: <a20577c1-378d-4d38-b2b6-c37b950e7c08@linux.dev> (raw)
In-Reply-To: <20260926051542.D2B111F00893@smtp.kernel.org>



On 9/25/26 10:15 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
> - [High] bpf: verifier bypass of resource leaks on unwinding without pad
> - [High] bpf: double free via bpf_unwind_resume returning to normal path in verifier
> - [Medium] bpf: precision tracking corruption when backtracking through bpf_unwind_resume
> - [Medium] bpf: CFG and liveness graph mismatch for bpf_unwind_resume
> --
>
> commit c4ed3d378ed3f7a547e662e1b6885fb3e54fef13
> Author: Yonghong Song <yonghong.song@linux.dev>
>
> bpf: Resume a covered call at its landing pad
>
> This patch introduces the ability for a BPF program to resume execution at a
> landing pad after an unwinding operation. It updates the verifier's control
> flow and state management to correctly account for these transitions, treating
> the landing pad as an ordinary second successor in the same frame.
>
>> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
>> index 0e38b9575328c..57665c67e66b3 100644
>> --- a/kernel/bpf/backtrack.c
>> +++ b/kernel/bpf/backtrack.c
> [ ... ]
>> @@ -434,8 +435,24 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
>>   					return -EFAULT;
>>   			}
>>   
>> +			if (bpf_exc_pad_of_call(env, idx) == subseq_idx) {
>> +				/*
>> +				 * We came from this call's landing pad, which
>> +				 * runs in the caller's frame: on that path the
>> +				 * callee's frame was never entered, so there is
>> +				 * no frame to leave. The call clobbered r0-r5;
>> +				 * r6-r9 and the stack are the caller's own and
>> +				 * keep going back from here.
>> +				 */
>> +				bt_clear_reg(bt, BPF_REG_0);
>> +				if (bt_reg_mask(bt) & BPF_REGMASK_ARGS) {
>> +					verifier_bug(env, "landing pad unexpected regs %x",
>> +						     bt_reg_mask(bt));
>> +					return -EFAULT;
>> +				}
>> +				return 0;
> [Severity: Medium]
> Does the BPF_CALL backtracking logic correctly handle bpf_unwind_resume()?

No. See below.

>
> When backtracking from a caller's instruction into a callee's pad and
> encountering bpf_unwind_resume(), it looks like backtrack_insn() processes it
> as a standard call without transitioning the backtrack frame (like
> bt_subprog_enter() does for BPF_EXIT).
>
> Could this leave the backtrack frame in the caller's context while traversing
> the callee, leading to precision tracking corruption?

Yes. I missed handling function bpf_is_unwind_resume_kfunc() which will
provide bt_subprog_enter().

>
>>   			/* callx calls static subprogs only */
>> -			if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
>> +			} else if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
> [ ... ]
>> diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
>> index 4e2b6985bc964..cfd4fe4049cae 100644
>> --- a/kernel/bpf/cfg.c
>> +++ b/kernel/bpf/cfg.c
> [ ... ]
>> @@ -678,6 +687,8 @@ static int visit_insn(int t, struct bpf_verifier_env *env)
>>   				mark_subprog_changes_pkt_data(env, t);
>>   			if (ret == 0 && bpf_is_throw_kfunc(insn))
>>   				mark_subprog_might_throw(env, t);
>> +			if (ret == 0 && bpf_is_unwind_kfunc(insn))
>> +				mark_subprog_might_unwind(env, t);
>>   		}
>>   		return visit_func_call_insn(t, insns, env, insn->src_reg == BPF_PSEUDO_CALL);
> [Severity: Medium]
> Is there a mismatch between the control flow graph and liveness tracking here?
>
> In visit_insn(), bpf_unwind_resume is processed as a standard kfunc call,
> which pushes a fallthrough edge into the CFG via visit_func_call_insn().
>
> However, in bpf_insn_successors() in liveness.c, it correctly acts as an exit:
>
>> diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
>> index 4e0273a8ceee4..c9ee4f10f7255 100644
>> --- a/kernel/bpf/liveness.c
>> +++ b/kernel/bpf/liveness.c
> [ ... ]
>> @@ -364,6 +364,9 @@ bpf_insn_successors(struct bpf_verifier_env *env, u32 idx)
>>   		return jt;
>>   	}
>>   
>> +	if (unlikely(bpf_is_unwind_resume_kfunc(insn)))
>> +		return succ;
>> +
>>   	opcode_info = &opcode_info_tbl[BPF_CLASS(insn->code) | BPF_OP(insn->code)];
> Could this mismatch, where the CFG considers the path reachable but liveness
> sees 0 successors, violate verifier graph invariants?

Okay, the below is an optimization:

+	if (unlikely(bpf_is_unwind_resume_kfunc(insn)))
+		return succ;

let me remove it so the number of successor's will be consistent between
cfg and liveness.

>
>> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
>> index fc3df452de2e4..77176250f866c 100644
>> --- a/kernel/bpf/verifier.c
>> +++ b/kernel/bpf/verifier.c
> [ ... ]
>> @@ -19167,6 +19167,40 @@ enum {
>>   	INSN_IDX_UPDATED = 2,
>>   };
>>   
>> +static int push_cleanup_pad_branch(struct bpf_verifier_env *env, int insn_idx)
>> +{
>> +	struct bpf_verifier_state *branch;
>> +	struct bpf_func_state *frame;
>> +	int pad = bpf_exc_pad_of_call(env, insn_idx);
>> +
>> +	if (pad < 0)
>> +		return 0;
>> +	branch = push_stack(env, pad, insn_idx, false);
>> +	if (IS_ERR(branch))
>> +		return PTR_ERR(branch);
>> +	frame = branch->frame[branch->curframe];
>> +	/*
>> +	 * The state at that call with the caller-saved registers gone: the
>> +	 * callee's epilogue put r6-r9 and the stack back on the way out.
>> +	 */
>> +	clear_caller_saved_regs(env, frame->regs);
>> +	mark_reg_unknown(env, frame->regs, BPF_REG_0);
>> +	return 0;
>> +}
>> +
>> +static int process_bpf_unwind(struct bpf_verifier_env *env, int *insn_idx)
>> +{
>> +	struct bpf_func_state *frame = cur_func(env);
>> +	int pad = bpf_exc_pad_of_call(env, *insn_idx);
>> +
>> +	if (pad < 0)
>> +		return PROCESS_BPF_EXIT;
> [Severity: High]
> Could returning PROCESS_BPF_EXIT here bypass resource leak and lock checks?
>
> If a BPF program acquires a resource like a spinlock and calls bpf_unwind()
> without a covering landing pad, pad < 0 evaluates to true. The verifier then
> pops the state and terminates the branch without invoking check_resource_leak()
> as process_bpf_exit_full() would normally do.
>
> Since do_check_insn() evaluates this kfunc before the active_locks check, could
> this allow the program to unwind while holding a spinlock, causing a kernel
> memory leak or deadlock at runtime?

Right, check_resource_leak() is a must here.

>
>> +	clear_caller_saved_regs(env, frame->regs);
>> +	mark_reg_unknown(env, frame->regs, BPF_REG_0);
>> +	*insn_idx = pad;
>> +	return INSN_IDX_UPDATED;
>> +}
> [ ... ]
>> @@ -19404,6 +19438,20 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>>   
>>   		env->jmps_processed++;
>>   		if (opcode == BPF_CALL) {
>> +			if (bpf_is_unwind_kfunc(insn))
>> +				return process_bpf_unwind(env, &env->insn_idx);
>> +			if (bpf_is_unwind_resume_kfunc(insn)) {
>> +				/*
>> +				 * Mark r0 a known zero -- unknown first, as
>> +				 * the known-zero helper keeps the type it
>> +				 * finds, which here is NOT_INIT. The fixups
>> +				 * lower this to 'r0 = 0; exit', so the frame
>> +				 * returns a real zero.
>> +				 */
>> +				mark_reg_unknown(env, cur_regs(env), BPF_REG_0);
>> +				mark_reg_known_zero(env, cur_regs(env), BPF_REG_0);
>> +				return process_bpf_exit_full(env, do_print_state, false);
>> +			}
> [Severity: High]
> Could this simulate a normal return from the subprogram and lead to a double
> free vulnerability?
>
> Here, bpf_unwind_resume is handled by calling process_bpf_exit_full() with
> exception_exit = false, returning to the normal instruction after the call.
>
> However, earlier in push_cleanup_pad_branch(), the caller's landing pad branch
> was already pushed using the pre-call state.
>
> At runtime, the callee's pad actually jumps to the caller's pad, not the
> instruction after the call. If a callee's pad frees a resource, the verifier
> will check the caller's pad using the pre-call state (where the resource was
> still allocated).
>
> Could this allow the caller's pad to free the same resource a second time,
> resulting in a use-after-free or double free vulnerability?

Yes, it is possible.
	outer_call:				inner_call:
		bpf_rcu_read_lock
		call inner_call
						bpf_unwind
					landing_pad:
						bpf_rcu_read_unlock
						...
						bpf_unwind_resume
	landing_pad:
		bpf_rcu_read_unlock
		...

The outer_call landing_pad uses states at the beginning of 'call_inner_call'
which does have bpf_rcu_read_lock. So this caused double bpf_rcu_read_unlock.

To fix this issue, for every function call which has landing_pad,
checkpoint the entry state (bpf_rcu_read_lock, bpf_prompt_disable, etc.)
and right before bpf_unwind_resume to ensure the state is the same
as entry state. This will prevent the above double free.

>
>>   			if (env->cur_state->active_locks) {
>>   				/* similar to static subprog calls callx is allowed under a lock */
>>   				if (!bpf_is_callx(insn) &&


  parent reply	other threads:[~2026-09-27  0:41 UTC|newest]

Thread overview: 56+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 01/21] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 02/21] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 03/21] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 04/21] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-09-26  5:16   ` sashiko-bot
2026-09-26 23:54     ` Yonghong Song
2026-09-27 20:39   ` bot+bpf-ci
2026-09-28  0:01     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-26  5:21   ` sashiko-bot
2026-09-27  0:02     ` Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  0:12     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad Yonghong Song
2026-09-26  5:15   ` sashiko-bot
2026-09-26  8:21     ` Alexei Starovoitov
2026-09-27  0:04       ` Yonghong Song
2026-09-27  0:41     ` Yonghong Song [this message]
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  0:17     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-09-26  5:17   ` sashiko-bot
2026-09-27  3:06     ` Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  0:29     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 09/21] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 10/21] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  1:08     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-26  5:15   ` sashiko-bot
2026-09-27  4:35     ` Yonghong Song
2026-09-27 20:39   ` bot+bpf-ci
2026-09-28  3:10     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 12/21] bpf, arm64: " Yonghong Song
2026-09-26  5:14   ` sashiko-bot
2026-09-27 20:40   ` bot+bpf-ci
2026-09-26  5:01 ` [PATCH bpf-next v6 13/21] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 14/21] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 15/21] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-27 20:39   ` bot+bpf-ci
2026-09-28  3:28     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 16/21] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 17/21] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 18/21] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-09-26  5:18   ` sashiko-bot
2026-09-27  4:58     ` Yonghong Song
2026-09-27 20:24   ` bot+bpf-ci
2026-09-28  3:36     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 20/21] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  3:49     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 21/21] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=a20577c1-378d-4d38-b2b6-c37b950e7c08@linux.dev \
    --to=yonghong.song@linux.dev \
    --cc=bpf@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox