BPF List
 help / color / mirror / Atom feed
From: "Kumar Kartikeya Dwivedi" <memxor@gmail.com>
To: <bot+bpf-ci@kernel.org>, <bpf@vger.kernel.org>
Cc: <ast@kernel.org>, <andrii@kernel.org>, <daniel@iogearbox.net>,
	<eddyz87@gmail.com>, <emil@etsalapatis.com>, <tj@kernel.org>,
	<kkd@meta.com>, <kernel-team@meta.com>, <martin.lau@kernel.org>,
	<yonghong.song@linux.dev>, <mason@kernel.org>,
	<ihor.solodrai@linux.dev>
Subject: Re: [PATCH bpf-next v2 16/18] bpf, x86: Allow programs 2 KiB of stack
Date: Thu, 24 Sep 2026 11:57:28 +0200	[thread overview]
Message-ID: <DLNGHABXA6V3.2MT1ETDTEP36D@gmail.com> (raw)
In-Reply-To: <61fbaeed4b3968122da9da2b7ac83e193a0058de88466d6f79fe0a446853d500@mail.kernel.org>

On Thu Sep 24, 2026 at 11:13 AM CEST, bot+bpf-ci wrote:
>> diff --git a/Documentation/bpf/bpf_design_QA.rst b/Documentation/bpf/bpf_design_QA.rst
>> index eb19c945f4d5..be5fc4ac00d6 100644
>> --- a/Documentation/bpf/bpf_design_QA.rst
>> +++ b/Documentation/bpf/bpf_design_QA.rst
>
> [ ... ]
>
>> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
>> index 9fbef7504e51..e2e531dd1e0b 100644
>> --- a/arch/x86/net/bpf_jit_comp.c
>> +++ b/arch/x86/net/bpf_jit_comp.c
>> @@ -4520,6 +4520,18 @@ bool bpf_jit_supports_subprog_tailcalls(void)
>>  	return true;
>>  }
>>
>> +/*
>> + * Frame sizes are 32-bit immediates in the prologue, epilogue and tail call
>> + * sequences, a tail call pops the caller's frame and lands in the target's
>> + * prologue before the target allocates its own, and private stacks are
>> + * allocated from the program's own depth, so MAX_BPF_STACK_JIT frames need
>> + * nothing special.
>> + */
>> +bool bpf_jit_supports_large_stack(void)
>> +{
>> +	return true;
>> +}
>> +
>
> Can a sched_cls or lwt_xmit program overflow the kernel stack through
> bpf_clone_redirect() recursion?
>
> A TC egress program on device X that calls bpf_clone_redirect(skb,
> X_ifindex, 0) redirects the packet back to itself. The redirect happens
> while the calling program's frame is on the stack:
>
>   bpf_clone_redirect()
>     __bpf_redirect()
>       __bpf_redirect_common()
>         __bpf_tx_skb()
>           dev_queue_xmit()
>             __dev_queue_xmit()
>               sch_handle_egress()
>                 tcx_run()
>                   bpf_prog_run() -> (same or another tc egress program)
>
> The guard in __bpf_tx_skb() (net/core/filter.c) only fails when the
> counter exceeds XMIT_RECURSION_LIMIT (8):
>
>   if (dev_xmit_recursion()) {
>       net_crit_ratelimited("bpf: recursion limit reached on datapath, buggy bpf program?\n");
>       kfree_skb(skb);
>       return -ENETDOWN;
>   }
>
> So dev_xmit_recursion() allows values 0 through 8 before it fails, which
> means the program runs at nesting levels 0 through 9. That puts ten of
> its frames on the stack at once.
>
> sched_cls programs never get a private stack:
>
>   kernel/bpf/verifier.c:bpf_enable_priv_stack() returns NO_PRIV_STACK
>   for them
>
> so their frames live on the task or IRQ stack. On x86-64 without KASAN
> that stack is 16 KiB (THREAD_SIZE_ORDER = IRQ_STACK_ORDER = 2).
>
> With this patch, a single 2048-byte frame is accepted. The BPF frames
> alone then take 10 x 2048 = 20 KiB, before any kernel frames in each
> round. That exceeds 16 KiB, so the stack overflows.
>
> Before this patch the same program was capped at 10 x 512 = 5 KiB of BPF
> frames, which fit. The recursion limit was meant to stop a buggy program
> from crashing the kernel.
>
> The commit message says nested programs "are not accounted against each
> other before or after this change; each level of nesting may now add up
> to 1.5 KiB more." It does not bound how many levels there can be. Here
> the kernel allows up to 10, which adds about 15 KiB compared to before.
>
> The same construction works through lwt_xmit programs, which also get
> bpf_clone_redirect(). Triggering the overflow needs CAP_BPF and
> CAP_NET_ADMIN to attach.
>
>

 Right, thanks; for v3 a program that calls bpf_clone_redirect() keeps the
 512-byte budget, since it can run ten frames deep through the datapath's
 recursion limit, and tests cover both bpf_clone_redirect() and bpf_redirect().

> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35976322553


  reply	other threads:[~2026-09-24  9:57 UTC|newest]

Thread overview: 34+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-24  8:25 [PATCH bpf-next v2 00/18] Raise BPF program stack size to 2KiB Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 01/18] bpf: Add accessors for verifier stack slots Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 02/18] bpf: Widen the stack slot index in the jump history Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 03/18] bpf: Store linked registers in the jump history as an array Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 04/18] bpf: Track backtracking stack slots with bitmaps Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 05/18] bpf: Track scratched stack slots with a bitmap Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 06/18] bpf: Treat unknown-size stack reads as reaching the frame top Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 07/18] bpf: Size liveness stack masks by the stack each frame uses Kumar Kartikeya Dwivedi
2026-09-24 15:12   ` Alexei Starovoitov
2026-09-24  8:25 ` [PATCH bpf-next v2 08/18] bpf: Grow the verifier id scratch on demand Kumar Kartikeya Dwivedi
2026-09-24  9:13   ` bot+bpf-ci
2026-09-24  9:55     ` Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 09/18] selftests/bpf: Cover the tail call caller stack depth limit Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 10/18] selftests/bpf: Check that narrow stack stores define no slot Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 11/18] selftests/bpf: Check liveness merge of masks with different widths Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 12/18] bpf: Size the per-frame verifier structures for a 2 KiB stack Kumar Kartikeya Dwivedi
2026-09-24  9:13   ` bot+bpf-ci
2026-09-24  9:56     ` Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 13/18] bpf: Bound program stack use by a per-program limit Kumar Kartikeya Dwivedi
2026-09-24  9:13   ` bot+bpf-ci
2026-09-24  9:56     ` Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 14/18] selftests/bpf: Add load conditions on the program stack limit Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 15/18] selftests/bpf: Give the 512-byte stack boundary tests a 2 KiB twin Kumar Kartikeya Dwivedi
2026-09-24  9:13   ` bot+bpf-ci
2026-09-24  9:56     ` Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 16/18] bpf, x86: Allow programs 2 KiB of stack Kumar Kartikeya Dwivedi
2026-09-24  9:13   ` bot+bpf-ci
2026-09-24  9:57     ` Kumar Kartikeya Dwivedi [this message]
2026-09-24  8:25 ` [PATCH bpf-next v2 17/18] bpf, arm64: " Kumar Kartikeya Dwivedi
2026-09-24  9:00   ` bot+bpf-ci
2026-09-24  9:57     ` Kumar Kartikeya Dwivedi
2026-09-24  8:25 ` [PATCH bpf-next v2 18/18] selftests/bpf: Test the 2 KiB stack budget Kumar Kartikeya Dwivedi
2026-09-24  9:13   ` bot+bpf-ci
2026-09-24  9:58     ` Kumar Kartikeya Dwivedi

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DLNGHABXA6V3.2MT1ETDTEP36D@gmail.com \
    --to=memxor@gmail.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bot+bpf-ci@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=emil@etsalapatis.com \
    --cc=ihor.solodrai@linux.dev \
    --cc=kernel-team@meta.com \
    --cc=kkd@meta.com \
    --cc=martin.lau@kernel.org \
    --cc=mason@kernel.org \
    --cc=tj@kernel.org \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox