Netdev List
 help / color / mirror / Atom feed
From: "Emil Tsalapatis" <emil@etsalapatis.com>
To: "Kuniyuki Iwashima" <kuniyu@google.com>,
	"Alexei Starovoitov" <ast@kernel.org>,
	"Daniel Borkmann" <daniel@iogearbox.net>,
	"Andrii Nakryiko" <andrii@kernel.org>,
	"Martin KaFai Lau" <martin.lau@linux.dev>,
	"Eduard Zingerman" <eddyz87@gmail.com>,
	"Kumar Kartikeya Dwivedi" <memxor@gmail.com>
Cc: "Yonghong Song" <yonghong.song@linux.dev>,
	"John Fastabend" <john.fastabend@gmail.com>,
	"Stanislav Fomichev" <sdf@fomichev.me>,
	"Eric Dumazet" <edumazet@google.com>,
	"Neal Cardwell" <ncardwell@google.com>,
	"Willem de Bruijn" <willemb@google.com>,
	"Tenzin Ukyab" <ukyab@berkeley.edu>,
	"Clément Léger" <cleger@meta.com>,
	"Kuniyuki Iwashima" <kuni1840@gmail.com>,
	bpf@vger.kernel.org, netdev@vger.kernel.org
Subject: Re: [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
Date: Wed, 23 Sep 2026 22:42:11 +0000	[thread overview]
Message-ID: <DLN248ZSNGXU.1ZF8YH5Q1A4XE@etsalapatis.com> (raw)
In-Reply-To: <20260923213719.224838-4-kuniyu@google.com>

On Wed Sep 23, 2026 at 9:35 PM UTC, Kuniyuki Iwashima wrote:
> Waking up a thread per packet is expensive when an application
> processes variable-length frames (e.g., RPC) that span multiple
> packets.
>
> SO_RCVLOWAT can defer wakeups, but because the frame size is
> encoded in a fixed-size descriptor at the start of each frame,
> the application has to:
>
>   1. wake up and recv() the descriptor,
>   2. raise SO_RCVLOWAT to the payload size via setsockopt(),
>   3. wake up and recv() the payload, and
>   4. reset SO_RCVLOWAT back to the descriptor size via
>      setsockopt() for the next frame.
>
> This requires an extra wakeup and two setsockopt() syscalls
> for every single RPC frame.
>
> With SOCKMAP, we can parse skb and suppress wakeups in kernel,
> but SOCKMAP adds overhead and also kills zerocopy.
>
> Let's add lighter-weight opt-in callbacks to bpf_tcp_ops to
> replace that.
>
>   .enqueue_rcvq(): invoked when TCP stack enqueues skb to
>                    sk->sk_receive_queue
>
>   .dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
>                    is dequeued from sk->sk_receive_queue
>
> Those callbacks can be enabled on a per-socket basis by
> bpf_setsockopt():
>
>   int flags = BPF_SOCK_OPS_RCVQ_CB_FLAG;
>
>   bpf_setsockopt(sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
>                  &flags, sizeof(flags));
>
> or via the bpf_tcp_ops-specific helper added in the next patch:
>
>   bpf_sock_ops_cb_flags_set(sk, BPF_SOCK_OPS_RCVQ_CB_FLAG);
>
> Later, we will add a new kfunc to adjust sk->sk_rcvlowat from
> these callbacks.
>
> This will allow the bpf_tcp_ops prog to parse each skb and
> dynamically adjust sk->sk_rcvlowat to suppress unnecessary EPOLLIN
> wakeups until sufficient data is available in the receive queue.
>
> The placement of bpf_tcp_ops_call() in tcp_ofo_queue() and
> tcp_fastopen_add_skb() is chosen to provide the same snapshot
> as tcp_queue_rcv().
>
> For example, if bpf_tcp_ops_call() were called before updating
> TCP_SKB_CB(skb)->seq in tcp_fastopen_add_skb(), BPF prog would
> need an extra branch for the unlikely TFO case to strip SYN.
>
> In addition, the TCP stack can queue overlapping skbs into recvq.
> Once rcv_nxt is updated with a new skb, BPF prog can no longer
> infer the previous rcv_nxt from skb->len.
>
> Lastly, dequeue_rcvq() is placed in tcp_cleanup_rbuf() rather
> than __tcp_cleanup_rbuf() so that it is not called for sockets
> in SOCKMAP, where calling sk->sk_data_ready() from the new
> kfunc would otherwise trigger infinite recursion.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> Acked-by: Stanislav Fomichev <sdf@fomichev.me>

Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>

> ---
>  include/net/tcp.h              | 18 ++++++++++++++++++
>  include/uapi/linux/bpf.h       | 11 ++++++++++-
>  net/ipv4/bpf_tcp_ops.c         | 10 ++++++++++
>  net/ipv4/tcp.c                 |  2 ++
>  net/ipv4/tcp_fastopen.c        |  2 ++
>  net/ipv4/tcp_input.c           |  4 ++++
>  tools/include/uapi/linux/bpf.h | 11 ++++++++++-
>  7 files changed, 56 insertions(+), 2 deletions(-)
>
> diff --git a/include/net/tcp.h b/include/net/tcp.h
> index d61ee00052e3..f2d838bcb0a7 100644
> --- a/include/net/tcp.h
> +++ b/include/net/tcp.h
> @@ -3053,6 +3053,12 @@ struct bpf_tcp_ops {
>  			      struct request_sock *req, struct sk_buff *syn_skb,
>  			      enum tcp_synack_type synack_type,
>  			      u32 opt_off);
> +
> +	/* Called when an incoming skb is enqueued to sk->sk_receive_queue. */
> +	void (*enqueue_rcvq)(struct sock *sk, struct sk_buff *skb);
> +
> +	/* Called after data is dequeued from sk->sk_receive_queue. */
> +	void (*dequeue_rcvq)(struct sock *sk);
>  };
>  
>  #define bpf_tcp_ops_call(op, sk, ...)					\
> @@ -3144,6 +3150,18 @@ static inline void tcp_bpf_rtt(struct sock *sk, long mrtt, u32 srtt)
>  	bpf_tcp_ops_call(rtt, sk, mrtt, srtt);
>  }
>  
> +static inline void bpf_tcp_ops_enqueue_rcvq(struct sock *sk, struct sk_buff *skb)
> +{
> +	if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
> +		bpf_tcp_ops_call(enqueue_rcvq, sk, skb);
> +}
> +
> +static inline void bpf_tcp_ops_dequeue_rcvq(struct sock *sk)
> +{
> +	if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
> +		bpf_tcp_ops_call(dequeue_rcvq, sk);
> +}
> +
>  #if IS_ENABLED(CONFIG_SMC)
>  extern struct static_key_false tcp_have_smc;
>  #endif
> diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
> index 6330b7d745c5..fe122242b096 100644
> --- a/include/uapi/linux/bpf.h
> +++ b/include/uapi/linux/bpf.h
> @@ -7148,8 +7148,17 @@ enum {
>  	 * options first before the BPF program does.
>  	 */
>  	BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
> +	/* Call bpf when the TCP stack enqueues/dequeues payload
> +	 * to/from sk->sk_receive_queue.
> +	 *
> +	 * Only bpf_tcp_ops is supported.
> +	 *
> +	 * It can be used to adjust sk->sk_rcvlowat and suppress
> +	 * unnecessary wakeups before sufficient data is available.
> +	 */
> +	BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
>  /* Mask of all currently supported cb flags */
> -	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0x7F,
> +	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0xFF,
>  };
>  
>  enum {
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index 1ada3b781bf1..c963e2cc21b5 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
> @@ -76,6 +76,14 @@ static void write_hdr_opt_stub(struct sock *sk, struct sk_buff *skb,
>  {
>  }
>  
> +static void enqueue_rcvq_stub(struct sock *sk, struct sk_buff *skb)
> +{
> +}
> +
> +static void dequeue_rcvq_stub(struct sock *sk)
> +{
> +}
> +
>  static struct bpf_tcp_ops __bpf_tcp_ops = {
>  	.timeout_init = timeout_init_stub,
>  	.rwnd_init = rwnd_init_stub,
> @@ -90,6 +98,8 @@ static struct bpf_tcp_ops __bpf_tcp_ops = {
>  	.parse_hdr = parse_hdr_stub,
>  	.hdr_opt_len = hdr_opt_len_stub,
>  	.write_hdr_opt = write_hdr_opt_stub,
> +	.enqueue_rcvq = enqueue_rcvq_stub,
> +	.dequeue_rcvq = dequeue_rcvq_stub,
>  };
>  
>  BPF_CALL_4(bpf_tcp_ops_store_hdr_opt, void *, ctx, const void *, from,
> diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
> index a4456b419412..a714b36a7494 100644
> --- a/net/ipv4/tcp.c
> +++ b/net/ipv4/tcp.c
> @@ -1610,6 +1610,8 @@ void tcp_cleanup_rbuf(struct sock *sk, int copied)
>  	     "cleanup rbuf bug: copied %X seq %X rcvnxt %X\n",
>  	     tp->copied_seq, TCP_SKB_CB(skb)->end_seq, tp->rcv_nxt);
>  	__tcp_cleanup_rbuf(sk, copied);
> +
> +	bpf_tcp_ops_dequeue_rcvq(sk);
>  }
>  
>  static void tcp_eat_recv_skb(struct sock *sk, struct sk_buff *skb)
> diff --git a/net/ipv4/tcp_fastopen.c b/net/ipv4/tcp_fastopen.c
> index 471c78be5513..4939bcbc81d1 100644
> --- a/net/ipv4/tcp_fastopen.c
> +++ b/net/ipv4/tcp_fastopen.c
> @@ -281,6 +281,8 @@ void tcp_fastopen_add_skb(struct sock *sk, struct sk_buff *skb)
>  	TCP_SKB_CB(skb)->seq++;
>  	TCP_SKB_CB(skb)->tcp_flags &= ~TCPHDR_SYN;
>  
> +	bpf_tcp_ops_enqueue_rcvq(sk, skb);
> +
>  	tp->rcv_nxt = TCP_SKB_CB(skb)->end_seq;
>  	tcp_add_receive_queue(sk, skb);
>  	tp->syn_data_acked = 1;
> diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
> index 6ac6f9d5b6c3..c60c61bb0a71 100644
> --- a/net/ipv4/tcp_input.c
> +++ b/net/ipv4/tcp_input.c
> @@ -5344,6 +5344,8 @@ static void tcp_ofo_queue(struct sock *sk)
>  			continue;
>  		}
>  
> +		bpf_tcp_ops_enqueue_rcvq(sk, skb);
> +
>  		tail = skb_peek_tail(&sk->sk_receive_queue);
>  		eaten = tail && tcp_try_coalesce(sk, tail, skb, &fragstolen);
>  		tcp_rcv_nxt_update(tp, TCP_SKB_CB(skb)->end_seq);
> @@ -5547,6 +5549,8 @@ static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb,
>  	int eaten;
>  	struct sk_buff *tail = skb_peek_tail(&sk->sk_receive_queue);
>  
> +	bpf_tcp_ops_enqueue_rcvq(sk, skb);
> +
>  	eaten = (tail &&
>  		 tcp_try_coalesce(sk, tail,
>  				  skb, fragstolen)) ? 1 : 0;
> diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
> index 6330b7d745c5..fe122242b096 100644
> --- a/tools/include/uapi/linux/bpf.h
> +++ b/tools/include/uapi/linux/bpf.h
> @@ -7148,8 +7148,17 @@ enum {
>  	 * options first before the BPF program does.
>  	 */
>  	BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
> +	/* Call bpf when the TCP stack enqueues/dequeues payload
> +	 * to/from sk->sk_receive_queue.
> +	 *
> +	 * Only bpf_tcp_ops is supported.
> +	 *
> +	 * It can be used to adjust sk->sk_rcvlowat and suppress
> +	 * unnecessary wakeups before sufficient data is available.
> +	 */
> +	BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
>  /* Mask of all currently supported cb flags */
> -	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0x7F,
> +	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0xFF,
>  };
>  
>  enum {


  reply	other threads:[~2026-09-23 22:42 UTC|newest]

Thread overview: 26+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-23 21:35 [PATCH v2 bpf-next 0/8] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
2026-09-23 21:35 ` [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
2026-09-23 22:04   ` Emil Tsalapatis
2026-09-23 22:31   ` bot+bpf-ci
2026-09-24 15:53   ` Stanislav Fomichev
2026-09-23 21:35 ` [PATCH v2 bpf-next 2/8] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv Kuniyuki Iwashima
2026-09-23 22:11   ` Emil Tsalapatis
2026-09-23 21:35 ` [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-09-23 22:42   ` Emil Tsalapatis [this message]
2026-09-25  0:04   ` Alexei Starovoitov
2026-09-25  0:39     ` Kuniyuki Iwashima
2026-09-25  1:41       ` Alexei Starovoitov
2026-09-25 17:33         ` Amery Hung
2026-09-27  0:18           ` Kuniyuki Iwashima
2026-09-30 19:56             ` Amery Hung
2026-09-30 21:03               ` Kuniyuki Iwashima
2026-09-23 21:35 ` [PATCH v2 bpf-next 4/8] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops Kuniyuki Iwashima
2026-09-23 22:31   ` bot+bpf-ci
2026-09-23 21:35 ` [PATCH v2 bpf-next 5/8] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
2026-09-24  3:39   ` Emil Tsalapatis
2026-09-23 21:35 ` [PATCH v2 bpf-next 6/8] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG Kuniyuki Iwashima
2026-09-24  3:48   ` Emil Tsalapatis
2026-09-24  4:10     ` Kuniyuki Iwashima
2026-09-23 21:35 ` [PATCH v2 bpf-next 7/8] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
2026-09-24  0:30   ` Emil Tsalapatis
2026-09-23 21:35 ` [PATCH v2 bpf-next 8/8] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DLN248ZSNGXU.1ZF8YH5Q1A4XE@etsalapatis.com \
    --to=emil@etsalapatis.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=cleger@meta.com \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=edumazet@google.com \
    --cc=john.fastabend@gmail.com \
    --cc=kuni1840@gmail.com \
    --cc=kuniyu@google.com \
    --cc=martin.lau@linux.dev \
    --cc=memxor@gmail.com \
    --cc=ncardwell@google.com \
    --cc=netdev@vger.kernel.org \
    --cc=sdf@fomichev.me \
    --cc=ukyab@berkeley.edu \
    --cc=willemb@google.com \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox