All of lore.kernel.org
 help / color / mirror / Atom feed
From: "Alexei Starovoitov" <alexei.starovoitov@gmail.com>
To: "Jakub Sitnicki" <jakub@cloudflare.com>,
	"Daniel Zahka" <daniel.zahka@gmail.com>
Cc: <netdev@vger.kernel.org>, "Alexei Starovoitov" <ast@kernel.org>,
	"Jakub Kicinski" <kuba@kernel.org>,
	"Kuniyuki Iwashima" <kuniyu@google.com>,
	"Paolo Abeni" <pabeni@redhat.com>,
	"Stanislav Fomichev" <sdf@fomichev.me>, <bpf@vger.kernel.org>,
	<kernel-team@cloudflare.com>,
	"Daniel Borkmann" <daniel@iogearbox.net>,
	"John Fastabend" <john.fastabend@gmail.com>,
	"Andrii Nakryiko" <andrii@kernel.org>,
	"Eduard Zingerman" <eddyz87@gmail.com>,
	"Kumar Kartikeya Dwivedi" <memxor@gmail.com>,
	"Martin KaFai Lau" <martin.lau@linux.dev>,
	"Song Liu" <song@kernel.org>,
	"Yonghong Song" <yonghong.song@linux.dev>,
	"Jiri Olsa" <jolsa@kernel.org>,
	"Emil Tsalapatis" <emil@etsalapatis.com>,
	"David S. Miller" <davem@davemloft.net>,
	"Eric Dumazet" <edumazet@google.com>,
	"Simon Horman" <horms@kernel.org>,
	"Jesper Dangaard Brouer" <hawk@kernel.org>,
	"Willem de Bruijn" <willemdebruijn.kernel@gmail.com>,
	"Florian Westphal" <fw@strlen.de>,
	"Jack Wang" <163wangjack@gmail.com>
Subject: Re: [PATCH net-next v2 03/14] bpf: Make BPF skb extension survive packet scrubbing
Date: Fri, 25 Sep 2026 01:07:47 +0000	[thread overview]
Message-ID: <DLNZU9RYL0GW.CHHAJ4R62R9M@gmail.com> (raw)
In-Reply-To: <87fqyypp67.fsf@cloudflare.com>

On Thu Sep 24, 2026 at 4:52 PM UTC, Jakub Sitnicki wrote:
> Hi Daniel,
>
> On Wed, Sep 23, 2026 at 01:46 PM -04, Daniel Zahka wrote:
>> On Thu Sep 10, 2026 at 10:02 AM EDT, Jakub Sitnicki wrote:
>>> skb_scrub_packet() drops all skb extensions unconditionally via
>>> skb_ext_reset(). It runs on tunnel encap/decap (ip_tunnel_rcv, vxlan_rcv,
>>> etc.) and cross-netns forwarding (dev_forward_skb).
>>>
>>> This makes it impossible for a BPF program to pass metadata via bpf_skb_ext
>>> through a tunnel or across a netns boundary. The extension is always lost
>>> at the scrub point.
>>>
>>> Introduce skb_ext_scrub(), a selective variant of skb_ext_reset(). It
>>> deletes every extension except SKB_EXT_BPF. Scrubbing is safe when the
>>> extension slab is shared with clones: deleting an extension only clears the
>>> per-skb active_extensions bit, and the shared slab payload is released
>>> lazily by __skb_ext_put() once the last reference goes away.
>>>
>>> Replace the skb_ext_reset() call in skb_scrub_packet() with skb_ext_scrub()
>>> and also switch udp_try_make_stateless() to skb_ext_scrub() as well, so the
>>> BPF metadata survives queueing onto a UDP socket receive queue and stays
>>> readable there (e.g. for a sockmap verdict program). Only mark the skb
>>> stateless when no extension survives the scrub. Otherwise skb_consume_udp()
>>> would take the __consume_stateless_skb() fast path, which skips
>>> skb_release_head_state(), and leak the extension slab.
>>>
>>> Signed-off-by: Jakub Sitnicki <jakub@cloudflare.com>
>>> ---
>>
>> Hello Jakub,
>> What are your current plans for this series? This commit solves the same
>> problem I have with wanting to preserve the PSP skb extension across
>> netns forwarding.
>
> I've implemented Alexei's idea of skb-lifecycle tracepoints that run
> only when an skb is marked/traced. Currently putting final touches on it
> before sending it out for the first round of feedback. You can take
> sneak peek at it on GH [1] to see if it meets your needs.
>
> The CPU overhead is lower compared to the skb extension, at least in my
> local runs, and the kernel changes are simpler, so it seems like a win
> overall:
>
> |                  | gated skb tps    | bpf skb ext      |
> |------------------|------------------|------------------|
> | **busy**         | **+5.44 ± 3.16** | **+8.01 ± 4.75** |
> | sys              | +2.62 ± 2.06     | +3.67 ± 2.76     |
> | soft             | +2.89 ± 1.43     | +3.83 ± 2.10     |
> | ns/pkt @146k pps | **+≈ 373**       | **+≈ 549**       |
>
> I'll be giving an update on it at LPC [1], if you're attending, and of
> course will keep you posted here on the ML.

Well, the numbers speak for themselves. I think :)


  reply	other threads:[~2026-09-25  1:07 UTC|newest]

Thread overview: 35+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-10 14:02 [PATCH net-next v2 00/14] skb extension for BPF metadata Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 01/14] bpf: Introduce per-packet metadata storage for BPF programs Jakub Sitnicki
2026-09-11 10:19   ` Jiayuan Chen
2026-09-10 14:02 ` [PATCH net-next v2 02/14] bpf: Allow access to bpf_sock_ops_kern->skb Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 03/14] bpf: Make BPF skb extension survive packet scrubbing Jakub Sitnicki
2026-09-23 17:46   ` Daniel Zahka
2026-09-24 16:52     ` Jakub Sitnicki
2026-09-25  1:07       ` Alexei Starovoitov [this message]
2026-09-25 19:18         ` Jakub Kicinski
2026-09-25 19:51           ` Alexei Starovoitov
2026-09-25 20:20             ` Jakub Kicinski
2026-09-25 20:27               ` Alexei Starovoitov
2026-09-28 15:52                 ` Jakub Sitnicki
2026-09-28 16:14                   ` Alexei Starovoitov
2026-09-29 20:32                     ` Jakub Sitnicki
2026-09-30  9:00                       ` Alexei Starovoitov
2026-10-02 12:37                         ` Jakub Sitnicki
2026-09-25 12:03       ` Daniel Zahka
2026-09-25 15:24         ` Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 04/14] selftests/bpf: Add tests for bpf_dynptr_from_skb_ext Jakub Sitnicki
2026-09-11 15:09   ` sashiko-bot
2026-09-10 14:02 ` [PATCH net-next v2 05/14] selftests/bpf: Test skb_ext on cloned skbs Jakub Sitnicki
2026-09-11 15:09   ` sashiko-bot
2026-09-10 14:02 ` [PATCH net-next v2 06/14] selftests/bpf: Test skb_ext survival across veth and GRE Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 07/14] selftests/bpf: Test skb_ext read from cgroup_skb and sk_filter hooks Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 08/14] selftests/bpf: Test skb_ext read from sock_ops and LSM hooks Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 09/14] selftests/bpf: Test skb_ext read from kfree_skb tracepoint Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 10/14] selftests/bpf: Test skb_ext read from netfilter hook Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 11/14] selftests/bpf: Test skb_ext from LWT in, out, and xmit hooks Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 12/14] selftests/bpf: Test skb_ext read from seg6local End.BPF hook Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 13/14] selftests/bpf: Test skb_ext read from sk_skb stream verdict hook Jakub Sitnicki
2026-09-10 14:02 ` [PATCH net-next v2 14/14] selftests/bpf: Use non-trivial test payload in xdp_context tests Jakub Sitnicki
2026-09-10 15:56 ` [PATCH net-next v2 00/14] skb extension for BPF metadata Alexei Starovoitov
2026-09-11 11:33   ` Jakub Sitnicki
2026-09-12  3:32     ` Alexei Starovoitov

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DLNZU9RYL0GW.CHHAJ4R62R9M@gmail.com \
    --to=alexei.starovoitov@gmail.com \
    --cc=163wangjack@gmail.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel.zahka@gmail.com \
    --cc=daniel@iogearbox.net \
    --cc=davem@davemloft.net \
    --cc=eddyz87@gmail.com \
    --cc=edumazet@google.com \
    --cc=emil@etsalapatis.com \
    --cc=fw@strlen.de \
    --cc=hawk@kernel.org \
    --cc=horms@kernel.org \
    --cc=jakub@cloudflare.com \
    --cc=john.fastabend@gmail.com \
    --cc=jolsa@kernel.org \
    --cc=kernel-team@cloudflare.com \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=martin.lau@linux.dev \
    --cc=memxor@gmail.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=sdf@fomichev.me \
    --cc=song@kernel.org \
    --cc=willemdebruijn.kernel@gmail.com \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.