From: "Emil Tsalapatis" <emil@etsalapatis.com>
To: "Kuniyuki Iwashima" <kuniyu@google.com>,
"Alexei Starovoitov" <ast@kernel.org>,
"Daniel Borkmann" <daniel@iogearbox.net>,
"Andrii Nakryiko" <andrii@kernel.org>,
"Martin KaFai Lau" <martin.lau@linux.dev>,
"Eduard Zingerman" <eddyz87@gmail.com>,
"Kumar Kartikeya Dwivedi" <memxor@gmail.com>
Cc: "Yonghong Song" <yonghong.song@linux.dev>,
"John Fastabend" <john.fastabend@gmail.com>,
"Stanislav Fomichev" <sdf@fomichev.me>,
"Eric Dumazet" <edumazet@google.com>,
"Neal Cardwell" <ncardwell@google.com>,
"Willem de Bruijn" <willemb@google.com>,
"Tenzin Ukyab" <ukyab@berkeley.edu>,
"Clément Léger" <cleger@meta.com>,
"Kuniyuki Iwashima" <kuni1840@gmail.com>,
bpf@vger.kernel.org, netdev@vger.kernel.org
Subject: Re: [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
Date: Wed, 23 Sep 2026 22:42:11 +0000 [thread overview]
Message-ID: <DLN248ZSNGXU.1ZF8YH5Q1A4XE@etsalapatis.com> (raw)
In-Reply-To: <20260923213719.224838-4-kuniyu@google.com>
On Wed Sep 23, 2026 at 9:35 PM UTC, Kuniyuki Iwashima wrote:
> Waking up a thread per packet is expensive when an application
> processes variable-length frames (e.g., RPC) that span multiple
> packets.
>
> SO_RCVLOWAT can defer wakeups, but because the frame size is
> encoded in a fixed-size descriptor at the start of each frame,
> the application has to:
>
> 1. wake up and recv() the descriptor,
> 2. raise SO_RCVLOWAT to the payload size via setsockopt(),
> 3. wake up and recv() the payload, and
> 4. reset SO_RCVLOWAT back to the descriptor size via
> setsockopt() for the next frame.
>
> This requires an extra wakeup and two setsockopt() syscalls
> for every single RPC frame.
>
> With SOCKMAP, we can parse skb and suppress wakeups in kernel,
> but SOCKMAP adds overhead and also kills zerocopy.
>
> Let's add lighter-weight opt-in callbacks to bpf_tcp_ops to
> replace that.
>
> .enqueue_rcvq(): invoked when TCP stack enqueues skb to
> sk->sk_receive_queue
>
> .dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
> is dequeued from sk->sk_receive_queue
>
> Those callbacks can be enabled on a per-socket basis by
> bpf_setsockopt():
>
> int flags = BPF_SOCK_OPS_RCVQ_CB_FLAG;
>
> bpf_setsockopt(sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
> &flags, sizeof(flags));
>
> or via the bpf_tcp_ops-specific helper added in the next patch:
>
> bpf_sock_ops_cb_flags_set(sk, BPF_SOCK_OPS_RCVQ_CB_FLAG);
>
> Later, we will add a new kfunc to adjust sk->sk_rcvlowat from
> these callbacks.
>
> This will allow the bpf_tcp_ops prog to parse each skb and
> dynamically adjust sk->sk_rcvlowat to suppress unnecessary EPOLLIN
> wakeups until sufficient data is available in the receive queue.
>
> The placement of bpf_tcp_ops_call() in tcp_ofo_queue() and
> tcp_fastopen_add_skb() is chosen to provide the same snapshot
> as tcp_queue_rcv().
>
> For example, if bpf_tcp_ops_call() were called before updating
> TCP_SKB_CB(skb)->seq in tcp_fastopen_add_skb(), BPF prog would
> need an extra branch for the unlikely TFO case to strip SYN.
>
> In addition, the TCP stack can queue overlapping skbs into recvq.
> Once rcv_nxt is updated with a new skb, BPF prog can no longer
> infer the previous rcv_nxt from skb->len.
>
> Lastly, dequeue_rcvq() is placed in tcp_cleanup_rbuf() rather
> than __tcp_cleanup_rbuf() so that it is not called for sockets
> in SOCKMAP, where calling sk->sk_data_ready() from the new
> kfunc would otherwise trigger infinite recursion.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
> ---
> include/net/tcp.h | 18 ++++++++++++++++++
> include/uapi/linux/bpf.h | 11 ++++++++++-
> net/ipv4/bpf_tcp_ops.c | 10 ++++++++++
> net/ipv4/tcp.c | 2 ++
> net/ipv4/tcp_fastopen.c | 2 ++
> net/ipv4/tcp_input.c | 4 ++++
> tools/include/uapi/linux/bpf.h | 11 ++++++++++-
> 7 files changed, 56 insertions(+), 2 deletions(-)
>
> diff --git a/include/net/tcp.h b/include/net/tcp.h
> index d61ee00052e3..f2d838bcb0a7 100644
> --- a/include/net/tcp.h
> +++ b/include/net/tcp.h
> @@ -3053,6 +3053,12 @@ struct bpf_tcp_ops {
> struct request_sock *req, struct sk_buff *syn_skb,
> enum tcp_synack_type synack_type,
> u32 opt_off);
> +
> + /* Called when an incoming skb is enqueued to sk->sk_receive_queue. */
> + void (*enqueue_rcvq)(struct sock *sk, struct sk_buff *skb);
> +
> + /* Called after data is dequeued from sk->sk_receive_queue. */
> + void (*dequeue_rcvq)(struct sock *sk);
> };
>
> #define bpf_tcp_ops_call(op, sk, ...) \
> @@ -3144,6 +3150,18 @@ static inline void tcp_bpf_rtt(struct sock *sk, long mrtt, u32 srtt)
> bpf_tcp_ops_call(rtt, sk, mrtt, srtt);
> }
>
> +static inline void bpf_tcp_ops_enqueue_rcvq(struct sock *sk, struct sk_buff *skb)
> +{
> + if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
> + bpf_tcp_ops_call(enqueue_rcvq, sk, skb);
> +}
> +
> +static inline void bpf_tcp_ops_dequeue_rcvq(struct sock *sk)
> +{
> + if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
> + bpf_tcp_ops_call(dequeue_rcvq, sk);
> +}
> +
> #if IS_ENABLED(CONFIG_SMC)
> extern struct static_key_false tcp_have_smc;
> #endif
> diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
> index 6330b7d745c5..fe122242b096 100644
> --- a/include/uapi/linux/bpf.h
> +++ b/include/uapi/linux/bpf.h
> @@ -7148,8 +7148,17 @@ enum {
> * options first before the BPF program does.
> */
> BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
> + /* Call bpf when the TCP stack enqueues/dequeues payload
> + * to/from sk->sk_receive_queue.
> + *
> + * Only bpf_tcp_ops is supported.
> + *
> + * It can be used to adjust sk->sk_rcvlowat and suppress
> + * unnecessary wakeups before sufficient data is available.
> + */
> + BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
> /* Mask of all currently supported cb flags */
> - BPF_SOCK_OPS_ALL_CB_FLAGS = 0x7F,
> + BPF_SOCK_OPS_ALL_CB_FLAGS = 0xFF,
> };
>
> enum {
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index 1ada3b781bf1..c963e2cc21b5 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
> @@ -76,6 +76,14 @@ static void write_hdr_opt_stub(struct sock *sk, struct sk_buff *skb,
> {
> }
>
> +static void enqueue_rcvq_stub(struct sock *sk, struct sk_buff *skb)
> +{
> +}
> +
> +static void dequeue_rcvq_stub(struct sock *sk)
> +{
> +}
> +
> static struct bpf_tcp_ops __bpf_tcp_ops = {
> .timeout_init = timeout_init_stub,
> .rwnd_init = rwnd_init_stub,
> @@ -90,6 +98,8 @@ static struct bpf_tcp_ops __bpf_tcp_ops = {
> .parse_hdr = parse_hdr_stub,
> .hdr_opt_len = hdr_opt_len_stub,
> .write_hdr_opt = write_hdr_opt_stub,
> + .enqueue_rcvq = enqueue_rcvq_stub,
> + .dequeue_rcvq = dequeue_rcvq_stub,
> };
>
> BPF_CALL_4(bpf_tcp_ops_store_hdr_opt, void *, ctx, const void *, from,
> diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
> index a4456b419412..a714b36a7494 100644
> --- a/net/ipv4/tcp.c
> +++ b/net/ipv4/tcp.c
> @@ -1610,6 +1610,8 @@ void tcp_cleanup_rbuf(struct sock *sk, int copied)
> "cleanup rbuf bug: copied %X seq %X rcvnxt %X\n",
> tp->copied_seq, TCP_SKB_CB(skb)->end_seq, tp->rcv_nxt);
> __tcp_cleanup_rbuf(sk, copied);
> +
> + bpf_tcp_ops_dequeue_rcvq(sk);
> }
>
> static void tcp_eat_recv_skb(struct sock *sk, struct sk_buff *skb)
> diff --git a/net/ipv4/tcp_fastopen.c b/net/ipv4/tcp_fastopen.c
> index 471c78be5513..4939bcbc81d1 100644
> --- a/net/ipv4/tcp_fastopen.c
> +++ b/net/ipv4/tcp_fastopen.c
> @@ -281,6 +281,8 @@ void tcp_fastopen_add_skb(struct sock *sk, struct sk_buff *skb)
> TCP_SKB_CB(skb)->seq++;
> TCP_SKB_CB(skb)->tcp_flags &= ~TCPHDR_SYN;
>
> + bpf_tcp_ops_enqueue_rcvq(sk, skb);
> +
> tp->rcv_nxt = TCP_SKB_CB(skb)->end_seq;
> tcp_add_receive_queue(sk, skb);
> tp->syn_data_acked = 1;
> diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
> index 6ac6f9d5b6c3..c60c61bb0a71 100644
> --- a/net/ipv4/tcp_input.c
> +++ b/net/ipv4/tcp_input.c
> @@ -5344,6 +5344,8 @@ static void tcp_ofo_queue(struct sock *sk)
> continue;
> }
>
> + bpf_tcp_ops_enqueue_rcvq(sk, skb);
> +
> tail = skb_peek_tail(&sk->sk_receive_queue);
> eaten = tail && tcp_try_coalesce(sk, tail, skb, &fragstolen);
> tcp_rcv_nxt_update(tp, TCP_SKB_CB(skb)->end_seq);
> @@ -5547,6 +5549,8 @@ static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb,
> int eaten;
> struct sk_buff *tail = skb_peek_tail(&sk->sk_receive_queue);
>
> + bpf_tcp_ops_enqueue_rcvq(sk, skb);
> +
> eaten = (tail &&
> tcp_try_coalesce(sk, tail,
> skb, fragstolen)) ? 1 : 0;
> diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
> index 6330b7d745c5..fe122242b096 100644
> --- a/tools/include/uapi/linux/bpf.h
> +++ b/tools/include/uapi/linux/bpf.h
> @@ -7148,8 +7148,17 @@ enum {
> * options first before the BPF program does.
> */
> BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
> + /* Call bpf when the TCP stack enqueues/dequeues payload
> + * to/from sk->sk_receive_queue.
> + *
> + * Only bpf_tcp_ops is supported.
> + *
> + * It can be used to adjust sk->sk_rcvlowat and suppress
> + * unnecessary wakeups before sufficient data is available.
> + */
> + BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
> /* Mask of all currently supported cb flags */
> - BPF_SOCK_OPS_ALL_CB_FLAGS = 0x7F,
> + BPF_SOCK_OPS_ALL_CB_FLAGS = 0xFF,
> };
>
> enum {
next prev parent reply other threads:[~2026-09-23 22:42 UTC|newest]
Thread overview: 26+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-23 21:35 [PATCH v2 bpf-next 0/8] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
2026-09-23 21:35 ` [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
2026-09-23 22:04 ` Emil Tsalapatis
2026-09-23 22:31 ` bot+bpf-ci
2026-09-24 15:53 ` Stanislav Fomichev
2026-09-23 21:35 ` [PATCH v2 bpf-next 2/8] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv Kuniyuki Iwashima
2026-09-23 22:11 ` Emil Tsalapatis
2026-09-23 21:35 ` [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-09-23 22:42 ` Emil Tsalapatis [this message]
2026-09-25 0:04 ` Alexei Starovoitov
2026-09-25 0:39 ` Kuniyuki Iwashima
2026-09-25 1:41 ` Alexei Starovoitov
2026-09-25 17:33 ` Amery Hung
2026-09-27 0:18 ` Kuniyuki Iwashima
2026-09-30 19:56 ` Amery Hung
2026-09-30 21:03 ` Kuniyuki Iwashima
2026-09-23 21:35 ` [PATCH v2 bpf-next 4/8] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops Kuniyuki Iwashima
2026-09-23 22:31 ` bot+bpf-ci
2026-09-23 21:35 ` [PATCH v2 bpf-next 5/8] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
2026-09-24 3:39 ` Emil Tsalapatis
2026-09-23 21:35 ` [PATCH v2 bpf-next 6/8] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG Kuniyuki Iwashima
2026-09-24 3:48 ` Emil Tsalapatis
2026-09-24 4:10 ` Kuniyuki Iwashima
2026-09-23 21:35 ` [PATCH v2 bpf-next 7/8] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
2026-09-24 0:30 ` Emil Tsalapatis
2026-09-23 21:35 ` [PATCH v2 bpf-next 8/8] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DLN248ZSNGXU.1ZF8YH5Q1A4XE@etsalapatis.com \
--to=emil@etsalapatis.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=cleger@meta.com \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=edumazet@google.com \
--cc=john.fastabend@gmail.com \
--cc=kuni1840@gmail.com \
--cc=kuniyu@google.com \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=ncardwell@google.com \
--cc=netdev@vger.kernel.org \
--cc=sdf@fomichev.me \
--cc=ukyab@berkeley.edu \
--cc=willemb@google.com \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox