All of lore.kernel.org
 help / color / mirror / Atom feed
From: Jakub Sitnicki <jakub@cloudflare.com>
To: "Alexei Starovoitov" <alexei.starovoitov@gmail.com>
Cc: "Jakub Kicinski" <kuba@kernel.org>,
	 "Daniel Zahka" <daniel.zahka@gmail.com>,
	 <netdev@vger.kernel.org>,
	 "Kuniyuki Iwashima" <kuniyu@google.com>,
	 "Paolo Abeni" <pabeni@redhat.com>,
	 "Stanislav Fomichev" <sdf@fomichev.me>,  <bpf@vger.kernel.org>,
	<kernel-team@cloudflare.com>,
	 "Daniel Borkmann" <daniel@iogearbox.net>,
	"John Fastabend" <john.fastabend@gmail.com>,
	 "Andrii Nakryiko" <andrii@kernel.org>,
	 "Eduard Zingerman" <eddyz87@gmail.com>,
	 "Kumar Kartikeya Dwivedi" <memxor@gmail.com>,
	 "Martin KaFai Lau" <martin.lau@linux.dev>,
	 "Song Liu" <song@kernel.org>,
	 "Yonghong Song" <yonghong.song@linux.dev>,
	 "Jiri Olsa" <jolsa@kernel.org>,
	 "Emil Tsalapatis" <emil@etsalapatis.com>,
	 "David S. Miller" <davem@davemloft.net>,
	 "Eric Dumazet" <edumazet@google.com>,
	 "Simon Horman" <horms@kernel.org>,
	 "Jesper Dangaard Brouer" <hawk@kernel.org>,
	"Willem de Bruijn" <willemdebruijn.kernel@gmail.com>,
	 "Florian Westphal" <fw@strlen.de>,
	 "Jack Wang" <163wangjack@gmail.com>
Subject: Re: [PATCH net-next v2 03/14] bpf: Make BPF skb extension survive packet scrubbing
Date: Mon, 28 Sep 2026 17:52:03 +0200	[thread overview]
Message-ID: <87jyo5bcgs.fsf@cloudflare.com> (raw)
In-Reply-To: <DLOOI89GKM5B.W4XPBNJUXRV1@gmail.com> (Alexei Starovoitov's message of "Fri, 25 Sep 2026 20:27:31 +0000")

On Fri, Sep 25, 2026 at 08:27 PM GMT, Alexei Starovoitov wrote:
> On Fri Sep 25, 2026 at 8:20 PM UTC, Jakub Kicinski wrote:
>> On Fri, 25 Sep 2026 19:51:37 +0000 Alexei Starovoitov wrote:
>>> On Fri, Sep 25, 2026 at 12:18 PM Jakub Kicinski <kuba@kernel.org> wrote:
>>> > Instead of improving the existing infra we're adding more BPF-specific
>>> > glue and another bit to the skb. Matter of perspective I suppose :/  
>>> 
>>> skb_ext version needs a bit too. SKB_EXT_BPF is the 8th id, so it takes
>>> the last bit of u8 active_extensions.
>>> skbuff_ext_cache object is sized for all ids, so it also grows by
>>> BPF_SKB_EXT_SIZE for xfrm, mptcp, psp whether bpf is used or not.
>>> 
>>> As I said earlier skb_ext is fine from bpf pov. If it can be made as
>>> cheap as bit + tracepoint I don't mind it at all.
>>> Which part of skb_ext would you improve?
>>
>> Mostly allocation speed. Either a small per-CPU cache like we have
>> for skbs or let the scalar matadata fields live inside the skb.
>
> makes sense to me. Speeding up generic infra is always a good thing.
>
>> That said, the selective tracing approach is also tempting.
>> It'd be great if we could have that for sockets.
>> Could we possibly think of a way to build that into tracepoints
>> instead of having the open coded
>>
>> obj_maybe_trace_bla()
>> {
>> 	if (obj->trace)
>> 		trace_bla();
>> }
>>
>> Having both:
>>
>> 	skb_maybe_trace_free(skb, SKB_CONSUMED);
>> 	skb_release_data(skb, SKB_CONSUMED);
>
> with a 3rd enum argument for different use cases?
> As a way to generalize different bits / different users?
>
> Also makes sense.

The benchmarks from Friday were wrong. I got bit by KVM halt polling,
which was randomly inflating the runtime cost by preventing batching.

The correct results - as in reproducible and with lower variance - are
the other way around - BPF skb ext is cheaper CPU-wise than gated skb
tracepoints. Please see:

https://patch.msgid.link/20260928-bpf-meta-gated-tracepoints-v1-0-844dbf3e1edf@cloudflare.com

Despite that, I agree that the tracepoint approach is tempting from the
UX PoV.

While an skb extension seems like a natural continuation of the XDP/TC
metadata pattern, I think that model is not a great fit for skbs
traveling through the network stack.

For XDP hook, there will be usually a single owning process of the
attached program, I think. While TC, cgroup, tracepoints can have
independent program owners, each wanting to associate their own piece of
metadata with the skb.


I can iterate on on cosmetic aspects so that we don't have blocks like:

        if (reason == SKB_CONSUMED)
                trace_consume_skb(...)
        else
                trace_kfree_skb(..., reason, ...)
        skb_maybe_trace_free(..., reason, ...)

... in a couple places where they appear in v1.

Thanks for feedback.

  reply	other threads:[~2026-09-28 15:54 UTC|newest]

Thread overview: 35+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-10 14:02 [PATCH net-next v2 00/14] skb extension for BPF metadata Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 01/14] bpf: Introduce per-packet metadata storage for BPF programs Jakub Sitnicki
2026-09-11 10:19   ` Jiayuan Chen
2026-09-10 14:02 ` [PATCH net-next v2 02/14] bpf: Allow access to bpf_sock_ops_kern->skb Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 03/14] bpf: Make BPF skb extension survive packet scrubbing Jakub Sitnicki
2026-09-23 17:46   ` Daniel Zahka
2026-09-24 16:52     ` Jakub Sitnicki
2026-09-25  1:07       ` Alexei Starovoitov
2026-09-25 19:18         ` Jakub Kicinski
2026-09-25 19:51           ` Alexei Starovoitov
2026-09-25 20:20             ` Jakub Kicinski
2026-09-25 20:27               ` Alexei Starovoitov
2026-09-28 15:52                 ` Jakub Sitnicki [this message]
2026-09-28 16:14                   ` Alexei Starovoitov
2026-09-29 20:32                     ` Jakub Sitnicki
2026-09-30  9:00                       ` Alexei Starovoitov
2026-10-02 12:37                         ` Jakub Sitnicki
2026-09-25 12:03       ` Daniel Zahka
2026-09-25 15:24         ` Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 04/14] selftests/bpf: Add tests for bpf_dynptr_from_skb_ext Jakub Sitnicki
2026-09-11 15:09   ` sashiko-bot
2026-09-10 14:02 ` [PATCH net-next v2 05/14] selftests/bpf: Test skb_ext on cloned skbs Jakub Sitnicki
2026-09-11 15:09   ` sashiko-bot
2026-09-10 14:02 ` [PATCH net-next v2 06/14] selftests/bpf: Test skb_ext survival across veth and GRE Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 07/14] selftests/bpf: Test skb_ext read from cgroup_skb and sk_filter hooks Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 08/14] selftests/bpf: Test skb_ext read from sock_ops and LSM hooks Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 09/14] selftests/bpf: Test skb_ext read from kfree_skb tracepoint Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 10/14] selftests/bpf: Test skb_ext read from netfilter hook Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 11/14] selftests/bpf: Test skb_ext from LWT in, out, and xmit hooks Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 12/14] selftests/bpf: Test skb_ext read from seg6local End.BPF hook Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 13/14] selftests/bpf: Test skb_ext read from sk_skb stream verdict hook Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 14/14] selftests/bpf: Use non-trivial test payload in xdp_context tests Jakub Sitnicki
2026-09-10 15:56 ` [PATCH net-next v2 00/14] skb extension for BPF metadata Alexei Starovoitov
2026-09-11 11:33   ` Jakub Sitnicki
2026-09-12  3:32     ` Alexei Starovoitov

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=87jyo5bcgs.fsf@cloudflare.com \
    --to=jakub@cloudflare.com \
    --cc=163wangjack@gmail.com \
    --cc=alexei.starovoitov@gmail.com \
    --cc=andrii@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel.zahka@gmail.com \
    --cc=daniel@iogearbox.net \
    --cc=davem@davemloft.net \
    --cc=eddyz87@gmail.com \
    --cc=edumazet@google.com \
    --cc=emil@etsalapatis.com \
    --cc=fw@strlen.de \
    --cc=hawk@kernel.org \
    --cc=horms@kernel.org \
    --cc=john.fastabend@gmail.com \
    --cc=jolsa@kernel.org \
    --cc=kernel-team@cloudflare.com \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=martin.lau@linux.dev \
    --cc=memxor@gmail.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=sdf@fomichev.me \
    --cc=song@kernel.org \
    --cc=willemdebruijn.kernel@gmail.com \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.