BPF List
 help / color / mirror / Atom feed
From: Yonghong Song <yonghong.song@linux.dev>
To: bot+bpf-ci@kernel.org, bpf@vger.kernel.org
Cc: ast@kernel.org, andrii@kernel.org, daniel@iogearbox.net,
	eddyz87@gmail.com, kernel-team@fb.com, martin.lau@kernel.org,
	mason@kernel.org, ihor.solodrai@linux.dev
Subject: Re: [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume
Date: Sun, 27 Sep 2026 17:29:10 -0700	[thread overview]
Message-ID: <2ca45402-d5b5-49cb-b5cb-a72d022c597f@linux.dev> (raw)
In-Reply-To: <e99037a0f0459c6d3e4c30ae417dff917a74f23163785f10598ef953425347da@mail.kernel.org>



On 9/27/26 1:40 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
>> index 0143688896b0..4174c7d0177e 100644
>> --- a/include/linux/bpf_verifier.h
>> +++ b/include/linux/bpf_verifier.h
>> @@ -339,6 +339,8 @@ struct bpf_func_state {
>>   	bool in_async_callback_fn;
>>   	bool in_exception_callback_fn;
>>   	bool no_stack_arg_load;
>> +	/* an unwind reached this frame and its landing pad is running */
>> +	bool in_pad;
>>   	/* For callback calling functions that limit number of possible
>>   	 * callback executions (e.g. bpf_loop) keeps track of current
>>   	 * simulated iteration number.
> [ ... ]
>
>> @@ -698,6 +700,8 @@ struct bpf_insn_aux_data {
>>   	u64 non_stack_access:1; /* instruction can access non-stack memory */
>>   	/* true if some jump or call instruction targets this instruction */
>>   	u64 jump_target:1;
>> +	u64 in_cleanup_pad:1; /* reached with a landing pad running */
>> +	u64 outside_cleanup_pad:1; /* ... and the other way round */
>>
>>   	unsigned int orig_idx; /* original instruction index, initialized once */
> Do these two bits grow bpf_insn_aux_data by 8 bytes per instruction?

I would like to keep the current layout for now as this is not
the important part of this patch set.

>
> Before this patch the u64 flag word held 31 bits (nospec through
> jump_target), fitting in the low 4 bytes, so orig_idx went at offset +4
> inside the flag word. Adding in_cleanup_pad and outside_cleanup_pad makes
> it 33 bits. Bit 32 lands in byte 4, so orig_idx moves to offset +8, and
> every later member shifts by 4 bytes.
>
> The 4-byte tail padding was already used up by cleanup_pad from 97f5b1e363fd.
> The struct therefore rounds up by another 4 bytes, for +8 in total. Measured
> from the start of the flag word, the tail goes from 64 bytes to 68, padded to
> 72.
>
> da01eb889d50 (Suggested-by Eduard Zingerman) was written to cut this struct
> from 136 bytes to 128. On those figures this patch takes it back to 136. The
> array is allocated for every load as __vmalloc(array_size(sizeof(struct
> bpf_insn_aux_data), len), ...), and bpf_patch_insn_data() memmoves and
> reallocates it on every patch.
>
> The two marks are only read and written by bpf_exc_check_insn(), which runs
> only when env->cleanup_info_cnt is non-zero. So programs that never use
> exception tables pay the cost.
>
> The cleanup_pad_head bit added later in the series makes the word 34 bits and
> does not change this. No later commit in the series restores the size.
>
> Two possible fixes: narrow alu_state, whose flags BPF_ALU_SANITIZE_SRC
> through BPF_ALU_IMMEDIATE occupy only bits 0-4, so the word stays within 32
> bits; or keep the pad-membership marks in a side array allocated only when a
> cleanup table is present.
>
>> diff --git a/kernel/bpf/exception.c b/kernel/bpf/exception.c
>> index 1ed0370a171b..86fdf847d33e 100644
>> --- a/kernel/bpf/exception.c
>> +++ b/kernel/bpf/exception.c
>> @@ -12,6 +12,15 @@
>>   BTF_ID_LIST_SINGLE(bpf_unwind_id, func, bpf_unwind)
>>   BTF_ID_LIST_SINGLE(bpf_unwind_resume_id, func, bpf_unwind_resume)
>>
>> +int bpf_exc_check_callback(struct bpf_verifier_env *env, int subprog)
>> +{
>> +	if (!env->subprog_info[subprog].might_unwind)
>> +		return 0;
>> +
>> +	verbose(env, "subprog %d may unwind and is used as a callback\n", subprog);
>> +	return -EINVAL;
>> +}
>> +
>>   static void mark_call_sites(struct bpf_verifier_env *env)
>>   {
>>   	u32 i, j;
> [ ... ]
>
>> @@ -69,6 +78,78 @@ bool bpf_is_unwind_resume_kfunc(const struct bpf_insn *insn)
>>   	       insn->imm == bpf_unwind_resume_id[0];
>>   }
>>
>> +/* Is an unwind in flight: is this frame a landing pad, or below one? */
>> +static bool unwinding(const struct bpf_verifier_state *state)
>> +{
>> +	u32 i;
>> +
>> +	for (i = 0; i <= state->curframe; i++)
>> +		if (state->frame[i]->in_pad)
>> +			return true;
>> +	return false;
>> +}
>> +
>> +int bpf_exc_check_insn(struct bpf_verifier_env *env, struct bpf_insn *insn)
>> +{
>> +	bool in_pad = cur_func(env)->in_pad;
>> +	struct bpf_insn_aux_data *aux;
>> +	u32 i = env->insn_idx;
>> +	const char *why = NULL;
>> +
>> +	if (unwinding(env->cur_state)) {
>> +		if (bpf_is_unwind_kfunc(insn)) {
>> +			verbose(env, "insn %u starts a second unwind while one is in flight\n", i);
>> +			return -EINVAL;
>> +		}
>> +		if (bpf_pseudo_call(insn)) {
>> +			int subprog = bpf_find_subprog(env, i + insn->imm + 1);
>> +
>> +			if (subprog >= 0 && bpf_subprog_is_global(env, subprog) &&
>> +			    env->subprog_info[subprog].might_unwind) {
>> +				verbose(env,
>> +					"insn %u calls global subprog %d, which can unwind while an unwind is in flight\n",
>> +					i, subprog);
>> +				return -EINVAL;
>> +			}
>> +		}
>> +	}
>> +
>> +	aux = &env->insn_aux_data[i];
>> +
>> +	if (in_pad ? aux->outside_cleanup_pad : aux->in_cleanup_pad) {
>> +		verbose(env, "insn %u runs both inside and outside a landing pad\n", i);
>> +		return -EINVAL;
>> +	}
>> +	if (in_pad)
>> +		aux->in_cleanup_pad = true;
>> +	else
>> +		aux->outside_cleanup_pad = true;
>> +
>> +	if (!in_pad)
>> +		return 0;
>> +
>> +	if (insn->code == (BPF_JMP | BPF_EXIT)) {
>> +		verbose(env,
>> +			"exit at insn %u ends a landing pad: a catch pad is not supported yet, only cleanup pads that resume\n",
>> +			i);
>> +		return -EOPNOTSUPP;
>> +	}
>> +	if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_tail_call)
>> +		why = "is a tail call, which replaces the frame";
>> +	else if (BPF_CLASS(insn->code) == BPF_LD &&
>> +		 (BPF_MODE(insn->code) == BPF_ABS || BPF_MODE(insn->code) == BPF_IND))
>> +		why = "is a BPF_LD_[ABS|IND], which can leave through the epilogue";
>> +	else if (insn->code == (BPF_JMP | BPF_JA | BPF_X) ||
>> +		 insn->code == (BPF_JMP32 | BPF_JA | BPF_X))
>> +		why = "is an indirect jump";
>> +
>> +	if (!why)
>> +		return 0;
>> +
>> +	verbose(env, "insn %u %s, and is in a landing pad\n", i, why);
>> +	return -EINVAL;
>> +}
>> +
> The commit message states "bpf_unwind() and the branch pushed at a covered
> call are the only ways into a pad, and both mark the frame they enter." Does
> bpf_exc_check_insn() also verify callx (indirect calls)?
>
> mark_call_sites() only sets aux->cleanup_pad for bpf_pseudo_call() and
> bpf_unwind(), not for bpf_is_callx():
>
> kernel/bpf/exception.c:mark_call_sites() {
> 	...
> 	if (!bpf_pseudo_call(insn) && !bpf_is_unwind_kfunc(insn))
> 		continue;
> 	env->insn_aux_data[j].cleanup_pad = rec->landing_pad_off + 1;
> }
>
> So bpf_exc_pad_of_call() returns -1 for a callx inside the covered range, and
> push_cleanup_pad_branch() in do_check_insn() returns 0 without pushing a pad
> state. The CFG and liveness successors add no pad edge for it either.
>
> The verifier walks the callx into its static callee through
> check_static_func_call(). If the callee's bpf_unwind() is not covered in the
> callee, process_bpf_unwind() returns PROCESS_BPF_EXIT and the path ends
> there. The caller frame is never continued, and the pad is never explored from
> the state at the callx.
>
> But at run time the unwind does reach that pad. bpf_unwind_rewrite() matches
> each frame's return address with bpf_exc_pad_for_ip(), which is a pure ip
> range check (begin < ip <= end) over ranges that bpf_exc_fill_native_ranges()
> builds. There is no filter on call type. The callee's return address sits
> inside the covered range, so it is rewritten to rec->pad. The callee's
> epilogue then restores the caller's r6-r9 and stack as they were at the
> callx.
>
> Example: a covered range holds call sub_a (pseudo call) with r6 = a valid
> pointer, then r6 = scalar; r2 = sub_b ll; callx r2. The pad dereferences or
> stores through r6 and then calls bpf_unwind_resume(). sub_b calls
> bpf_unwind() unconditionally. The pad is verified only with the state pushed
> at call sub_a, where r6 is a pointer. At run time sub_b's unwind resumes the
> pad with r6 holding the scalar, which gives an arbitrary kernel memory access
> from a program the verifier accepted.
>
> The same entry also gets around what this patch sets out to refuse. A pad
> reached only by normal flow and by the callx is verified with in_pad false, so
> it can end in exit (a catch pad), use a tail call or LD_ABS, and pass the
> "both inside and outside" check, because no in-pad state ever reaches it.
>
> No later commit in the series marks, refuses or filters callx.
> mark_call_sites(), push_cleanup_pad_branch() and bpf_exc_pad_for_ip() are
> unchanged in this respect, and neither the x86 nor the arm64 dispatch patch
> handles callx.
>
> Possible fixes: mark bpf_is_callx() sites in mark_call_sites(), so
> push_cleanup_pad_branch() and the CFG treat them like pseudo calls.
> Alternatively, refuse a callx inside a covered range when the cleanup table is
> checked.

Yes, fix in the next revision by adding callx support in mark_all_sites().

>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/36346422430


  reply	other threads:[~2026-09-28  0:29 UTC|newest]

Thread overview: 56+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-26  5:00 [PATCH bpf-next v6 00/21] bpf: Run exception cleanup landing pads when bpf_unwind() unwinds Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 01/21] bpf: Pack bpf_insn_aux_data flags into bit fields Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 02/21] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 03/21] bpf: Add the bpf_unwind() and bpf_unwind_resume() kfuncs Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 04/21] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 05/21] bpf: Prepare for an exception cleanup table before the CFG walk Yonghong Song
2026-09-26  5:16   ` sashiko-bot
2026-09-26 23:54     ` Yonghong Song
2026-09-27 20:39   ` bot+bpf-ci
2026-09-28  0:01     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 06/21] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-26  5:21   ` sashiko-bot
2026-09-27  0:02     ` Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  0:12     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 07/21] bpf: Resume a covered call at its landing pad Yonghong Song
2026-09-26  5:15   ` sashiko-bot
2026-09-26  8:21     ` Alexei Starovoitov
2026-09-27  0:04       ` Yonghong Song
2026-09-27  0:41     ` Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  0:17     ` Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 08/21] bpf: Refuse a landing pad that does not resume Yonghong Song
2026-09-26  5:17   ` sashiko-bot
2026-09-27  3:06     ` Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  0:29     ` Yonghong Song [this message]
2026-09-26  5:00 ` [PATCH bpf-next v6 09/21] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
2026-09-26  5:00 ` [PATCH bpf-next v6 10/21] bpf: Dispatch cleanup pads by rewriting return addresses Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  1:08     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 11/21] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-26  5:15   ` sashiko-bot
2026-09-27  4:35     ` Yonghong Song
2026-09-27 20:39   ` bot+bpf-ci
2026-09-28  3:10     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 12/21] bpf, arm64: " Yonghong Song
2026-09-26  5:14   ` sashiko-bot
2026-09-27 20:40   ` bot+bpf-ci
2026-09-26  5:01 ` [PATCH bpf-next v6 13/21] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 14/21] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 15/21] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-27 20:39   ` bot+bpf-ci
2026-09-28  3:28     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 16/21] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 17/21] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 18/21] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 19/21] selftests/bpf: Add __set_global() and __ret_global() test tags Yonghong Song
2026-09-26  5:18   ` sashiko-bot
2026-09-27  4:58     ` Yonghong Song
2026-09-27 20:24   ` bot+bpf-ci
2026-09-28  3:36     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 20/21] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
2026-09-27 20:40   ` bot+bpf-ci
2026-09-28  3:49     ` Yonghong Song
2026-09-26  5:01 ` [PATCH bpf-next v6 21/21] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=2ca45402-d5b5-49cb-b5cb-a72d022c597f@linux.dev \
    --to=yonghong.song@linux.dev \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bot+bpf-ci@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=ihor.solodrai@linux.dev \
    --cc=kernel-team@fb.com \
    --cc=martin.lau@kernel.org \
    --cc=mason@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox