From: Yonghong Song <yonghong.song@linux.dev>
To: bot+bpf-ci@kernel.org, bpf@vger.kernel.org
Cc: ast@kernel.org, andrii@kernel.org, daniel@iogearbox.net,
eddyz87@gmail.com, kernel-team@fb.com, martin.lau@kernel.org,
mason@kernel.org, ihor.solodrai@linux.dev
Subject: Re: [PATCH bpf-next v2 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach
Date: Sat, 19 Sep 2026 14:13:28 -0700 [thread overview]
Message-ID: <068024f3-cf0d-4df4-b3ef-6cec9fdafe96@linux.dev> (raw)
In-Reply-To: <27a9f5239c8c6a1c6039433b29d55451480f745f023c5f1ba943542502dfd780@mail.kernel.org>
On 9/17/26 10:44 PM, bot+bpf-ci@kernel.org wrote:
>> diff --git a/tools/testing/selftests/bpf/exceptions_cleanup.h b/tools/testing/selftests/bpf/exceptions_cleanup.h
>> index 630d2e207119a..5896cf83d15ee 100644
>> --- a/tools/testing/selftests/bpf/exceptions_cleanup.h
>> +++ b/tools/testing/selftests/bpf/exceptions_cleanup.h
>> @@ -4,6 +4,7 @@
>> #define __EXCEPTIONS_CLEANUP_H__
>>
>> #define THROW_COOKIE 0x100
>> +#define INNER_COOKIE 0x200
>>
>> /* progs/exceptions_cleanup.c: one bit per frame that reports it ran. */
>> #define RAN_FOO3_PREEMPT 0x1
> [ ... ]
>
>> diff --git a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
>> index d1e45b765af26..ffc0b95191681 100644
>> --- a/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
>> +++ b/tools/testing/selftests/bpf/prog_tests/exceptions_cleanup.c
> [ ... ]
>
>> +static void test_shapes(void)
>> +{
>> + struct exceptions_cleanup_shapes *skel;
>> +
>> + skel = exceptions_cleanup_shapes__open_and_load();
>> + if (!ASSERT_OK_PTR(skel, "shapes open_and_load"))
>> + return;
> [ ... ]
>
>> + /* A throwing subprog named by a BPF_PSEUDO_FUNC no helper is handed: the
>> + * callback check has to look at the bpf_loop(), not at the ld_imm64.
>> + */
>> + if (test__start_subtest("addr_taken_no_throw"))
>> + run_shape(skel, skel->progs.entry_addr_taken, 1, 2, 0);
>> + if (test__start_subtest("addr_taken_throw"))
>> + run_shape(skel, skel->progs.entry_addr_taken, 101, THROW_COOKIE,
>> + RAN_ADDR_TAKEN);
> Does this comment accurately describe what the addr_taken shape tests?
>
> Looking at addr_taken_callee() in progs/exceptions_cleanup_shapes.c, the
> function does hand cb_thrower to bpf_loop(). The ld_imm64 and the
> bpf_loop() call are on the same path, after bpf_throw(). Since bpf_throw()
> is not declared noreturn, LLVM keeps that tail, and what actually makes
> the program load is that the verifier never reaches the bpf_loop() at all
> (it is dead after the throw).
>
> So push_callback_call() -> bpf_cleanup_check_callback() never fires.
>
> The shape exercised is 'a BPF_PSEUDO_FUNC whose helper call the verifier
> never reaches', not 'a BPF_PSEUDO_FUNC no helper is handed'. Could the
> comment be rephrased to reflect what the verifier actually sees?
Okay, will change comments.
>
> [ ... ]
>
>> diff --git a/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
>> new file mode 100644
>> index 0000000000000..f5eb2ff15c896
>> --- /dev/null
>> +++ b/tools/testing/selftests/bpf/progs/exceptions_cleanup_shapes.c
> [ ... ]
>
>> +/*
>> + * 2. A callee called from both a covered and an uncovered site: the pad is
>> + * recorded on the call site, not on the callee. The lock sits between the two
>> + * calls because the frame really would leak it if the uncovered call unwound.
>> + */
>> +static __used __noinline __u64 shared_callee(__u64 x)
>> +{
>> + if (x > 100)
>> + bpf_throw(THROW_COOKIE);
>> + return x + 1;
>> +}
>> +
>> +static __used __naked __noinline __u64 shared_frame(void)
>> +{
>> + asm volatile (
>> + "r1 = %[input] ll;"
>> + "r6 = *(u64 *)(r1 + 0);"
>> + "r1 = 0;"
>> + "call shared_callee;"
>> + "call bpf_rcu_read_lock;"
>> + "r1 = r6;"
>> +"1:" "call shared_callee;" /* cleanup region */
>> +"2:"
>> + "r6 = r0;"
>> + "call bpf_rcu_read_unlock;"
>> + "r0 = r6;"
>> + "exit;"
>> +"3:" /* landing pad */
>> + "r7 = r0;"
>> + "call bpf_rcu_read_unlock;"
>> + PAD_RAN("%[ran]")
>> + "r1 = r7;"
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + :
>> + : [ran]"i"(RAN_SHARED), __imm_addr(input), __imm_addr(pads_ran)
>> + : __clobber_all);
>> +}
> Does the uncovered call site verify that the pad wouldn't be wrongly
> dispatched if an exception occurred there?
>
> The uncovered site is reached with the constant 0 in r1, and shared_callee
> only throws when x > 100, so the verifier proves that call cannot throw
> and no unwind ever leaves the uncovered site - at verification time or at
> run time.
>
> The structural half of the property (only the second call's return address
> falls inside the record's native range) is exercised, but the behavioral
> half is not: a kernel that wrongly matched the uncovered site's return
> address against the record and ran the pad would go unnoticed, because
> that site never unwinds.
The above analysis exactly described the code so the above 'wrongly dispatched
if an exception occurred there' does not happen. I guess this tries to
capture incorrect kernel implementation.
>
> [ ... ]
>
>> +/*
>> + * 6. A tail call that is really taken: the target is a program in its own
>> + * right, so the walk ends there and this frame's pad does not run. The callee
>> + * can also throw on a path never taken, which keeps the pad out of the sweep.
>> + */
>> +struct {
>> + __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
>> + __uint(max_entries, 1);
>> + __uint(key_size, sizeof(__u32));
>> + __uint(value_size, sizeof(__u32));
>> +} taken_table SEC(".maps");
>> +
>> +SEC("syscall")
>> +int tc_target(void *ctx)
>> +{
>> + bpf_throw(THROW_COOKIE);
>> + return 0;
>> +}
>> +
>> +static __used __noinline __u64 tc_taken_callee(void *ctx, __u64 x)
>> +{
>> + /* Never true at run time; the verifier cannot know that, and its
>> + * unwind out of here is what keeps the caller's pad alive.
>> + */
>> + if (x == 7)
>> + bpf_throw(THROW_COOKIE);
>> + bpf_tail_call_static(ctx, &taken_table, 0);
>> + return 0;
>> +}
>> +
>> +SEC("syscall")
>> +__naked int entry_tail_taken(void)
>> +{
>> + asm volatile (
>> + "*(u64 *)(r10 - 8) = r1;" /* the context, straight from entry */
>> + "r1 = %[input] ll;"
>> + "r2 = *(u64 *)(r1 + 0);"
>> + "if r2 < 101 goto 8f;"
>> + "r1 = *(u64 *)(r10 - 8);"
>> +"1:" "call tc_taken_callee;" /* cleanup region */
>> +"2:"
>> + "exit;" /* the cookie, delivered at tc_target */
>> +"8:"
>> + "r0 = 0;"
>> + "exit;"
>> +"3:" /* landing pad: must not run */
>> + PAD_RAN("%[ran]")
>> + "call bpf_unwind_resume;"
>> + "exit;"
>> + CLEANUP_REC("1b", "2b", "3b")
>> + :
>> + : [ran]"i"(RAN_TC_TAKEN), __imm_addr(input), __imm_addr(pads_ran)
>> + : __clobber_all);
>> +}
> Can the tail_call_taken subtest fail if the kernel wrongly continues the
> walk past the tail-call boundary?
This is a test which fixed some early regression tests. Yes, if kernel
implementation is wrong, tail_call_taken test may fail.
>
> entry_tail_taken narrows the argument before the covered call:
>
> "r2 = *(u64 *)(r1 + 0);" /* r2 = input */
> "if r2 < 101 goto 8f;" /* fallthrough => r2 in [101, U64_MAX] */
>
> tc_taken_callee is a static subprog, so check_func_call() copies the
> caller's r1-r5 verbatim into the callee frame. x therefore arrives with
> umin_value == 101, and is_branch_taken() resolves 'if (x == 7)' to
> never-taken, so the verifier never explores the bpf_throw(). Nothing then
> unwinds into entry_tail_taken during verification, the landing pad at
> label 3 is never marked seen, and opt_remove_dead_code() removes it.
>
> fixups.c then clears the annotation for any call site whose cleanup_pad
> falls inside the removed window:
>
> if (env->cleanup_info_cnt) {
> for (i = 0; i < env->insn_aux_data_len; i++) {
> u32 pad = aux_data[i].cleanup_pad;
> if (pad > off + cnt) aux_data[i].cleanup_pad = pad - cnt;
> else if (pad > off) aux_data[i].cleanup_pad = 0;
> }
> }
>
> so no native range is emitted for the region at all. The program still
> loads (mark_subprog_might_throw() is a static analysis, so bpf_check_cfg()
> still sees the pad edge and check_cleanup_info() runs before the sweep),
> and pads_ran stays 0 unconditionally. A kernel that wrongly continued
> bpf_stack_walker() past the tail-call boundary into entry_tail_taken's
> frame would find no record to match and would still leave pads_ran == 0,
> so the subtest passes on a correct and on a broken kernel alike.
>
> The in-code comment is also wrong on both clauses: the verifier CAN know
> x != 7 here, and consequently the unwind out of tc_taken_callee is NOT
> what keeps the caller's pad alive - nothing does.
The comment is:
/* Never true at run time; the verifier cannot know that, and its
* unwind out of here is what keeps the caller's pad alive.
*/
if (x == 7)
bpf_throw(THROW_COOKIE);
Yes, the 'x' will be a known value at run time. Comments need update.
>
> Contrast with shape 7 (entry_freplace, lines 862-878), the analogous 'walk
> ends in the callee's frame' shape, which deliberately passes the
> unnarrowed input value to fr_callee() with no 'if r1 < 101' guard, keeping
> fr_callee's identical 'if (x == 7) bpf_throw()' live. Dropping the
> 'if r2 < 101 goto 8f' guard from entry_tail_taken (or otherwise passing a
> value the verifier cannot exclude 7 from) would make the pad survive and
> give the subtest something to fail on.
>
> [ ... ]
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35308528711
next prev parent reply other threads:[~2026-09-19 21:13 UTC|newest]
Thread overview: 56+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-18 4:41 [PATCH bpf-next v2 00/20] bpf: Run exception cleanup landing pads when bpf_throw() unwinds Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 01/20] bpf: Accept the compiler's exception cleanup table at program load Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 02/20] bpf: Add the bpf_unwind_resume() kfunc Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 03/20] bpf: Add lookups for exception cleanup resumes and landing pads Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 4:55 ` Alexei Starovoitov
2026-09-19 17:42 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 04/20] bpf: Mark the call sites an exception cleanup table covers Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 05/20] bpf: Make exception landing pads reachable in the CFG Yonghong Song
2026-09-18 4:59 ` sashiko-bot
2026-09-19 19:17 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 19:32 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 06/20] bpf: Explore the landing pads no call site reaches Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 19:32 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 07/20] bpf: Refuse exception cleanup shapes bpf_throw() cannot dispatch Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 08/20] bpf: Walk the exception unwind in the verifier Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 09/20] bpf: Refuse a private stack for a program with an exception cleanup table Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 19:37 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 10/20] bpf: Dispatch exception cleanup pads from bpf_throw() Yonghong Song
2026-09-18 5:58 ` bot+bpf-ci
2026-09-19 19:54 ` Yonghong Song
2026-09-18 4:42 ` [PATCH bpf-next v2 11/20] bpf, x86: Dispatch exception cleanup pads at run time Yonghong Song
2026-09-18 5:03 ` sashiko-bot
2026-09-19 20:00 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 20:04 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 12/20] bpf, arm64: " Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 20:07 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 13/20] libbpf: Resolve the compiler's _Unwind_Resume to the kernel's kfunc Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 14/20] libbpf: Add cleanup_info to bpf_prog_load_opts Yonghong Song
2026-09-18 4:57 ` sashiko-bot
2026-09-19 20:18 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 15/20] libbpf: Collect .bpf_cleanup records and pass them to the kernel Yonghong Song
2026-09-18 5:00 ` sashiko-bot
2026-09-19 20:21 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 16/20] libbpf: Carry the exception cleanup table through the light skeleton Yonghong Song
2026-09-18 5:02 ` sashiko-bot
2026-09-19 20:27 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 17/20] libbpf: Let the static linker carry .bpf_cleanup relocations Yonghong Song
2026-09-18 5:01 ` sashiko-bot
2026-09-19 20:31 ` Yonghong Song
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 20:32 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 18/20] selftests/bpf: Add an end-to-end .bpf_cleanup exception test Yonghong Song
2026-09-18 4:59 ` sashiko-bot
2026-09-18 5:58 ` bot+bpf-ci
2026-09-19 20:34 ` Yonghong Song
2026-09-18 4:43 ` [PATCH bpf-next v2 19/20] selftests/bpf: Cover the exception cleanup shapes the chain does not reach Yonghong Song
2026-09-18 5:01 ` sashiko-bot
2026-09-18 5:44 ` bot+bpf-ci
2026-09-19 21:13 ` Yonghong Song [this message]
2026-09-18 4:43 ` [PATCH bpf-next v2 20/20] selftests/bpf: Load an exception cleanup program from a light skeleton Yonghong Song
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=068024f3-cf0d-4df4-b3ef-6cec9fdafe96@linux.dev \
--to=yonghong.song@linux.dev \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bot+bpf-ci@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=ihor.solodrai@linux.dev \
--cc=kernel-team@fb.com \
--cc=martin.lau@kernel.org \
--cc=mason@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox