BPF List
 help / color / mirror / Atom feed
* [PATCH v2 bpf-next 0/8] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT.
@ 2026-09-23 21:35 Kuniyuki Iwashima
  2026-09-23 21:35 ` [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
                   ` (7 more replies)
  0 siblings, 8 replies; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-23 21:35 UTC (permalink / raw)
  To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
	bpf, netdev

This series introduces two callbacks for bpf_tcp_ops:

  .enqueue_rcvq(): invoked when TCP stack enqueues skb to
                   sk->sk_receive_queue

  .dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
                   is dequeued from sk->sk_receive_queue

Those callbacks can be enabled on a per-socket basis by
bpf_setsockopt():

  int flags = BPF_SOCK_OPS_RCVQ_CB_FLAG;

  bpf_setsockopt(sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
                 &flags, sizeof(flags));

or via the bpf_tcp_ops-specific helper added in the next patch:

  bpf_sock_ops_cb_flags_set(sk, BPF_SOCK_OPS_RCVQ_CB_FLAG);

This allows the BPF prog to dynamically adjust sk->sk_rcvlowat,
suppressing unnecessary EPOLLIN wakeups until sufficient data
is available in the receive queue.

This functionality, which we call "TCP AutoLOWAT", was originally
developed in 2020 by Tenzin Ukyab with the help of Soheil Hassas
Yeganeh, Arjun Roy, and Eric Dumazet.  It has served Google RPC
workloads for more than 5 years.

Combined with TCP RX zerocopy, this typically allows us to read an
entire RPC frame with just a single wakeup and a single system call.

While the original implementation was specialised for our
internal RPC format, this series introduces a more flexible
version by leveraging BPF.

The bpf prog in the last selftest patch closely mirrors the core
logic of the original implementation to provide a real-world
example.

Note that the new callbacks are not supported on legacy SOCK_OPS.


Changes:
  v2:
    * Add patch 1 not to allow setsockopt() from new callbacks
    * Patch 8 (selftest)
      * Avoid address comparison for a specific version of gcc.
      * Make rpc_test_cases[] static.
      * Update comment in rpc_test_case[].

  v1: https://lore.kernel.org/bpf/20260920195633.3033620-1-kuniyu@google.com/

Legacy SOCK_OPS version:
  v3: https://lore.kernel.org/bpf/20260523083001.2911931-1-kuniyu@google.com/
  v2: https://lore.kernel.org/bpf/20260522074601.1658705-1-kuniyu@google.com/
  v1: https://lore.kernel.org/bpf/20260508073355.3916746-1-kuniyu@google.com/


Kuniyuki Iwashima (8):
  bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list.
  selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv.
  bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops.
  tcp: Split out __tcp_set_rcvlowat().
  bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG.
  bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
  selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().

 include/net/tcp.h                             |  34 ++
 include/uapi/linux/bpf.h                      |  13 +-
 net/core/filter.c                             |  10 +-
 net/ipv4/bpf_tcp_ops.c                        | 135 ++++++-
 net/ipv4/tcp.c                                |  14 +-
 net/ipv4/tcp_fastopen.c                       |   2 +
 net/ipv4/tcp_input.c                          |   4 +
 tools/include/uapi/linux/bpf.h                |  13 +-
 .../selftests/bpf/prog_tests/tcp_autolowat.c  | 350 ++++++++++++++++++
 .../selftests/bpf/prog_tests/tcpbpf_user.c    |   3 +-
 .../selftests/bpf/progs/bpf_tracing_net.h     |   2 +
 .../selftests/bpf/progs/tcp_autolowat.c       | 312 ++++++++++++++++
 .../selftests/bpf/progs/test_tcpbpf_kern.c    |   3 +-
 13 files changed, 866 insertions(+), 29 deletions(-)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
 create mode 100644 tools/testing/selftests/bpf/progs/tcp_autolowat.c

-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 29+ messages in thread

* [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list.
  2026-09-23 21:35 [PATCH v2 bpf-next 0/8] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
@ 2026-09-23 21:35 ` Kuniyuki Iwashima
  2026-09-23 22:04   ` Emil Tsalapatis
                     ` (2 more replies)
  2026-09-23 21:35 ` [PATCH v2 bpf-next 2/8] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv Kuniyuki Iwashima
                   ` (6 subsequent siblings)
  7 siblings, 3 replies; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-23 21:35 UTC (permalink / raw)
  To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
	bpf, netdev

Currently, four bpf_tcp_ops callbacks are not allowed to call
bpf_setsockopt() and bpf_getsockopt().

However, the deny-list is fragile, and when a new callback is
added, we might re-open a can of worms. [1][2]

Let's convert it to allow-list.

Note that it can be written as offsetof(,X) <= moff <= offsetof(,Y)
but clang can optimise to similar code anyway.

Link: https://lore.kernel.org/r/20220929070407.965581-5-martin.lau@linux.dev/ #[0]
Link: https://lore.kernel.org/bpf/20260421155804.135786-1-kafai.wan@linux.dev/ #[1]
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 net/ipv4/bpf_tcp_ops.c | 39 ++++++++++++++++++++++++---------------
 1 file changed, 24 insertions(+), 15 deletions(-)

diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index 681fed642999..1ada3b781bf1 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -210,6 +210,24 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
 	.ret_type	= RET_INTEGER,
 };
 
+static bool is_sockopt_supported(u32 moff)
+{
+	switch (moff) {
+	case offsetof(struct bpf_tcp_ops, active_established):
+	case offsetof(struct bpf_tcp_ops, passive_established):
+	case offsetof(struct bpf_tcp_ops, rto):
+	case offsetof(struct bpf_tcp_ops, rtt):
+	case offsetof(struct bpf_tcp_ops, set_state):
+	case offsetof(struct bpf_tcp_ops, retrans):
+	case offsetof(struct bpf_tcp_ops, connect):
+	case offsetof(struct bpf_tcp_ops, listen):
+	case offsetof(struct bpf_tcp_ops, parse_hdr):
+		return true;
+	}
+
+	return false;
+}
+
 static const struct bpf_func_proto *
 get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
 {
@@ -221,22 +239,13 @@ get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
 	case BPF_FUNC_sk_storage_delete:
 		return &bpf_sk_storage_delete_proto;
 	case BPF_FUNC_setsockopt:
-		/* The sk may be an unlocked listener (synack path) or NULL
-		 * fullsock; disable for members that can run unlocked.
-		 */
-		if (moff == offsetof(struct bpf_tcp_ops, rwnd_init) ||
-		    moff == offsetof(struct bpf_tcp_ops, timeout_init) ||
-		    moff == offsetof(struct bpf_tcp_ops, hdr_opt_len) ||
-		    moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))
-			return NULL;
-		return &bpf_sk_setsockopt_proto;
+		if (is_sockopt_supported(moff))
+			return &bpf_sk_setsockopt_proto;
+		return NULL;
 	case BPF_FUNC_getsockopt:
-		if (moff == offsetof(struct bpf_tcp_ops, rwnd_init) ||
-		    moff == offsetof(struct bpf_tcp_ops, timeout_init) ||
-		    moff == offsetof(struct bpf_tcp_ops, hdr_opt_len) ||
-		    moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))
-			return NULL;
-		return &bpf_sk_getsockopt_proto;
+		if (is_sockopt_supported(moff))
+			return &bpf_sk_getsockopt_proto;
+		return NULL;
 	case BPF_FUNC_get_retval:
 		if (moff == offsetof(struct bpf_tcp_ops, timeout_init) ||
 		    moff == offsetof(struct bpf_tcp_ops, rwnd_init))
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 bpf-next 2/8] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv.
  2026-09-23 21:35 [PATCH v2 bpf-next 0/8] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
  2026-09-23 21:35 ` [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
@ 2026-09-23 21:35 ` Kuniyuki Iwashima
  2026-09-23 22:11   ` Emil Tsalapatis
  2026-09-23 21:35 ` [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
                   ` (5 subsequent siblings)
  7 siblings, 1 reply; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-23 21:35 UTC (permalink / raw)
  To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
	bpf, netdev

Once bpf_sock_ops_cb_flags_set() supports a new flag,
tcpbpf_user.c fails due to the hard-coded max value, 0x80.

Let's replace 0x80 with BPF_SOCK_OPS_ALL_CB_FLAGS + 1.

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
---
 tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c | 3 ++-
 tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c | 3 ++-
 2 files changed, 4 insertions(+), 2 deletions(-)

diff --git a/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c b/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
index 7e8fe1bad03f..e4849d2a2956 100644
--- a/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
+++ b/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
@@ -26,7 +26,8 @@ static void verify_result(struct tcpbpf_globals *result)
 	ASSERT_EQ(result->bytes_acked, 1002, "bytes_acked");
 	ASSERT_EQ(result->data_segs_in, 1, "data_segs_in");
 	ASSERT_EQ(result->data_segs_out, 1, "data_segs_out");
-	ASSERT_EQ(result->bad_cb_test_rv, 0x80, "bad_cb_test_rv");
+	ASSERT_EQ(result->bad_cb_test_rv, BPF_SOCK_OPS_ALL_CB_FLAGS + 1,
+		  "bad_cb_test_rv");
 	ASSERT_EQ(result->good_cb_test_rv, 0, "good_cb_test_rv");
 	ASSERT_EQ(result->num_listen, 1, "num_listen");
 
diff --git a/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c b/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
index 6935f32eeb8f..e30cb1fab079 100644
--- a/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
+++ b/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
@@ -92,7 +92,8 @@ int bpf_testcb(struct bpf_sock_ops *skops)
 		break;
 	case BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB:
 		/* Test failure to set largest cb flag (assumes not defined) */
-		global.bad_cb_test_rv = bpf_sock_ops_cb_flags_set(skops, 0x80);
+		global.bad_cb_test_rv = bpf_sock_ops_cb_flags_set(skops,
+								  BPF_SOCK_OPS_ALL_CB_FLAGS + 1);
 		/* Set callback */
 		global.good_cb_test_rv = bpf_sock_ops_cb_flags_set(skops,
 						 BPF_SOCK_OPS_STATE_CB_FLAG);
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  2026-09-23 21:35 [PATCH v2 bpf-next 0/8] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
  2026-09-23 21:35 ` [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
  2026-09-23 21:35 ` [PATCH v2 bpf-next 2/8] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv Kuniyuki Iwashima
@ 2026-09-23 21:35 ` Kuniyuki Iwashima
  2026-09-23 21:50   ` sashiko-bot
                     ` (2 more replies)
  2026-09-23 21:35 ` [PATCH v2 bpf-next 4/8] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops Kuniyuki Iwashima
                   ` (4 subsequent siblings)
  7 siblings, 3 replies; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-23 21:35 UTC (permalink / raw)
  To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
	bpf, netdev

Waking up a thread per packet is expensive when an application
processes variable-length frames (e.g., RPC) that span multiple
packets.

SO_RCVLOWAT can defer wakeups, but because the frame size is
encoded in a fixed-size descriptor at the start of each frame,
the application has to:

  1. wake up and recv() the descriptor,
  2. raise SO_RCVLOWAT to the payload size via setsockopt(),
  3. wake up and recv() the payload, and
  4. reset SO_RCVLOWAT back to the descriptor size via
     setsockopt() for the next frame.

This requires an extra wakeup and two setsockopt() syscalls
for every single RPC frame.

With SOCKMAP, we can parse skb and suppress wakeups in kernel,
but SOCKMAP adds overhead and also kills zerocopy.

Let's add lighter-weight opt-in callbacks to bpf_tcp_ops to
replace that.

  .enqueue_rcvq(): invoked when TCP stack enqueues skb to
                   sk->sk_receive_queue

  .dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
                   is dequeued from sk->sk_receive_queue

Those callbacks can be enabled on a per-socket basis by
bpf_setsockopt():

  int flags = BPF_SOCK_OPS_RCVQ_CB_FLAG;

  bpf_setsockopt(sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
                 &flags, sizeof(flags));

or via the bpf_tcp_ops-specific helper added in the next patch:

  bpf_sock_ops_cb_flags_set(sk, BPF_SOCK_OPS_RCVQ_CB_FLAG);

Later, we will add a new kfunc to adjust sk->sk_rcvlowat from
these callbacks.

This will allow the bpf_tcp_ops prog to parse each skb and
dynamically adjust sk->sk_rcvlowat to suppress unnecessary EPOLLIN
wakeups until sufficient data is available in the receive queue.

The placement of bpf_tcp_ops_call() in tcp_ofo_queue() and
tcp_fastopen_add_skb() is chosen to provide the same snapshot
as tcp_queue_rcv().

For example, if bpf_tcp_ops_call() were called before updating
TCP_SKB_CB(skb)->seq in tcp_fastopen_add_skb(), BPF prog would
need an extra branch for the unlikely TFO case to strip SYN.

In addition, the TCP stack can queue overlapping skbs into recvq.
Once rcv_nxt is updated with a new skb, BPF prog can no longer
infer the previous rcv_nxt from skb->len.

Lastly, dequeue_rcvq() is placed in tcp_cleanup_rbuf() rather
than __tcp_cleanup_rbuf() so that it is not called for sockets
in SOCKMAP, where calling sk->sk_data_ready() from the new
kfunc would otherwise trigger infinite recursion.

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
---
 include/net/tcp.h              | 18 ++++++++++++++++++
 include/uapi/linux/bpf.h       | 11 ++++++++++-
 net/ipv4/bpf_tcp_ops.c         | 10 ++++++++++
 net/ipv4/tcp.c                 |  2 ++
 net/ipv4/tcp_fastopen.c        |  2 ++
 net/ipv4/tcp_input.c           |  4 ++++
 tools/include/uapi/linux/bpf.h | 11 ++++++++++-
 7 files changed, 56 insertions(+), 2 deletions(-)

diff --git a/include/net/tcp.h b/include/net/tcp.h
index d61ee00052e3..f2d838bcb0a7 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -3053,6 +3053,12 @@ struct bpf_tcp_ops {
 			      struct request_sock *req, struct sk_buff *syn_skb,
 			      enum tcp_synack_type synack_type,
 			      u32 opt_off);
+
+	/* Called when an incoming skb is enqueued to sk->sk_receive_queue. */
+	void (*enqueue_rcvq)(struct sock *sk, struct sk_buff *skb);
+
+	/* Called after data is dequeued from sk->sk_receive_queue. */
+	void (*dequeue_rcvq)(struct sock *sk);
 };
 
 #define bpf_tcp_ops_call(op, sk, ...)					\
@@ -3144,6 +3150,18 @@ static inline void tcp_bpf_rtt(struct sock *sk, long mrtt, u32 srtt)
 	bpf_tcp_ops_call(rtt, sk, mrtt, srtt);
 }
 
+static inline void bpf_tcp_ops_enqueue_rcvq(struct sock *sk, struct sk_buff *skb)
+{
+	if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
+		bpf_tcp_ops_call(enqueue_rcvq, sk, skb);
+}
+
+static inline void bpf_tcp_ops_dequeue_rcvq(struct sock *sk)
+{
+	if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
+		bpf_tcp_ops_call(dequeue_rcvq, sk);
+}
+
 #if IS_ENABLED(CONFIG_SMC)
 extern struct static_key_false tcp_have_smc;
 #endif
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 6330b7d745c5..fe122242b096 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -7148,8 +7148,17 @@ enum {
 	 * options first before the BPF program does.
 	 */
 	BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
+	/* Call bpf when the TCP stack enqueues/dequeues payload
+	 * to/from sk->sk_receive_queue.
+	 *
+	 * Only bpf_tcp_ops is supported.
+	 *
+	 * It can be used to adjust sk->sk_rcvlowat and suppress
+	 * unnecessary wakeups before sufficient data is available.
+	 */
+	BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
 /* Mask of all currently supported cb flags */
-	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0x7F,
+	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0xFF,
 };
 
 enum {
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index 1ada3b781bf1..c963e2cc21b5 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -76,6 +76,14 @@ static void write_hdr_opt_stub(struct sock *sk, struct sk_buff *skb,
 {
 }
 
+static void enqueue_rcvq_stub(struct sock *sk, struct sk_buff *skb)
+{
+}
+
+static void dequeue_rcvq_stub(struct sock *sk)
+{
+}
+
 static struct bpf_tcp_ops __bpf_tcp_ops = {
 	.timeout_init = timeout_init_stub,
 	.rwnd_init = rwnd_init_stub,
@@ -90,6 +98,8 @@ static struct bpf_tcp_ops __bpf_tcp_ops = {
 	.parse_hdr = parse_hdr_stub,
 	.hdr_opt_len = hdr_opt_len_stub,
 	.write_hdr_opt = write_hdr_opt_stub,
+	.enqueue_rcvq = enqueue_rcvq_stub,
+	.dequeue_rcvq = dequeue_rcvq_stub,
 };
 
 BPF_CALL_4(bpf_tcp_ops_store_hdr_opt, void *, ctx, const void *, from,
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index a4456b419412..a714b36a7494 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -1610,6 +1610,8 @@ void tcp_cleanup_rbuf(struct sock *sk, int copied)
 	     "cleanup rbuf bug: copied %X seq %X rcvnxt %X\n",
 	     tp->copied_seq, TCP_SKB_CB(skb)->end_seq, tp->rcv_nxt);
 	__tcp_cleanup_rbuf(sk, copied);
+
+	bpf_tcp_ops_dequeue_rcvq(sk);
 }
 
 static void tcp_eat_recv_skb(struct sock *sk, struct sk_buff *skb)
diff --git a/net/ipv4/tcp_fastopen.c b/net/ipv4/tcp_fastopen.c
index 471c78be5513..4939bcbc81d1 100644
--- a/net/ipv4/tcp_fastopen.c
+++ b/net/ipv4/tcp_fastopen.c
@@ -281,6 +281,8 @@ void tcp_fastopen_add_skb(struct sock *sk, struct sk_buff *skb)
 	TCP_SKB_CB(skb)->seq++;
 	TCP_SKB_CB(skb)->tcp_flags &= ~TCPHDR_SYN;
 
+	bpf_tcp_ops_enqueue_rcvq(sk, skb);
+
 	tp->rcv_nxt = TCP_SKB_CB(skb)->end_seq;
 	tcp_add_receive_queue(sk, skb);
 	tp->syn_data_acked = 1;
diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
index 6ac6f9d5b6c3..c60c61bb0a71 100644
--- a/net/ipv4/tcp_input.c
+++ b/net/ipv4/tcp_input.c
@@ -5344,6 +5344,8 @@ static void tcp_ofo_queue(struct sock *sk)
 			continue;
 		}
 
+		bpf_tcp_ops_enqueue_rcvq(sk, skb);
+
 		tail = skb_peek_tail(&sk->sk_receive_queue);
 		eaten = tail && tcp_try_coalesce(sk, tail, skb, &fragstolen);
 		tcp_rcv_nxt_update(tp, TCP_SKB_CB(skb)->end_seq);
@@ -5547,6 +5549,8 @@ static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb,
 	int eaten;
 	struct sk_buff *tail = skb_peek_tail(&sk->sk_receive_queue);
 
+	bpf_tcp_ops_enqueue_rcvq(sk, skb);
+
 	eaten = (tail &&
 		 tcp_try_coalesce(sk, tail,
 				  skb, fragstolen)) ? 1 : 0;
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 6330b7d745c5..fe122242b096 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -7148,8 +7148,17 @@ enum {
 	 * options first before the BPF program does.
 	 */
 	BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
+	/* Call bpf when the TCP stack enqueues/dequeues payload
+	 * to/from sk->sk_receive_queue.
+	 *
+	 * Only bpf_tcp_ops is supported.
+	 *
+	 * It can be used to adjust sk->sk_rcvlowat and suppress
+	 * unnecessary wakeups before sufficient data is available.
+	 */
+	BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
 /* Mask of all currently supported cb flags */
-	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0x7F,
+	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0xFF,
 };
 
 enum {
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 bpf-next 4/8] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops.
  2026-09-23 21:35 [PATCH v2 bpf-next 0/8] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
                   ` (2 preceding siblings ...)
  2026-09-23 21:35 ` [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
@ 2026-09-23 21:35 ` Kuniyuki Iwashima
  2026-09-23 22:31   ` bot+bpf-ci
  2026-09-23 21:35 ` [PATCH v2 bpf-next 5/8] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
                   ` (3 subsequent siblings)
  7 siblings, 1 reply; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-23 21:35 UTC (permalink / raw)
  To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
	bpf, netdev

When an error occurs in bpf_tcp_ops.{enqueue,dequeue}_rcvq(),
we want to clear BPF_SOCK_OPS_RCVQ_CB_FLAG to stop invoking
the callbacks.

In addition, bpf_sock_ops_cb_flags_set() is often used during
setup or just after 3WHS completes to enable opt-in hooks.

Let's support bpf_sock_ops_cb_flags_set() in the following
callbacks: connect, listen, {active,passive}_established,
{enqueue,dequeue}_rcvq.

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
---
 include/uapi/linux/bpf.h       |  2 +-
 net/ipv4/bpf_tcp_ops.c         | 27 +++++++++++++++++++++++++++
 tools/include/uapi/linux/bpf.h |  2 +-
 3 files changed, 29 insertions(+), 2 deletions(-)

diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index fe122242b096..8cdf22667775 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -3264,7 +3264,7 @@ union bpf_attr {
  * 	Return
  * 		0
  *
- * long bpf_sock_ops_cb_flags_set(struct bpf_sock_ops *bpf_sock, int argval)
+ * long bpf_sock_ops_cb_flags_set(void *bpf_sock, int argval)
  * 	Description
  * 		Attempt to set the value of the **bpf_sock_ops_cb_flags** field
  * 		for the full TCP socket associated to *bpf_sock_ops* to
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index c963e2cc21b5..4b48711d92a2 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -220,6 +220,24 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
 	.ret_type	= RET_INTEGER,
 };
 
+BPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)
+{
+	int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
+
+	tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
+
+	return argval & ~BPF_SOCK_OPS_ALL_CB_FLAGS;
+}
+
+static const struct bpf_func_proto bpf_tcp_ops_cb_flags_set_proto = {
+	.func		= bpf_tcp_ops_cb_flags_set,
+	.gpl_only	= false,
+	.ret_type	= RET_INTEGER,
+	.arg1_type	= ARG_PTR_TO_BTF_ID,
+	.arg1_btf_id	= &btf_sock_ids[BTF_SOCK_TYPE_TCP],
+	.arg2_type	= ARG_ANYTHING,
+};
+
 static bool is_sockopt_supported(u32 moff)
 {
 	switch (moff) {
@@ -274,6 +292,15 @@ get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
 		if (moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))
 			return &bpf_tcp_ops_store_hdr_opt_proto;
 		return NULL;
+	case BPF_FUNC_sock_ops_cb_flags_set:
+		if (moff == offsetof(struct bpf_tcp_ops, connect) ||
+		    moff == offsetof(struct bpf_tcp_ops, listen) ||
+		    moff == offsetof(struct bpf_tcp_ops, active_established) ||
+		    moff == offsetof(struct bpf_tcp_ops, passive_established) ||
+		    moff == offsetof(struct bpf_tcp_ops, enqueue_rcvq) ||
+		    moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+			return &bpf_tcp_ops_cb_flags_set_proto;
+		return NULL;
 	default:
 		return bpf_base_func_proto(func_id, prog);
 	}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index fe122242b096..8cdf22667775 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -3264,7 +3264,7 @@ union bpf_attr {
  * 	Return
  * 		0
  *
- * long bpf_sock_ops_cb_flags_set(struct bpf_sock_ops *bpf_sock, int argval)
+ * long bpf_sock_ops_cb_flags_set(void *bpf_sock, int argval)
  * 	Description
  * 		Attempt to set the value of the **bpf_sock_ops_cb_flags** field
  * 		for the full TCP socket associated to *bpf_sock_ops* to
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 bpf-next 5/8] tcp: Split out __tcp_set_rcvlowat().
  2026-09-23 21:35 [PATCH v2 bpf-next 0/8] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
                   ` (3 preceding siblings ...)
  2026-09-23 21:35 ` [PATCH v2 bpf-next 4/8] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops Kuniyuki Iwashima
@ 2026-09-23 21:35 ` Kuniyuki Iwashima
  2026-09-24  3:39   ` Emil Tsalapatis
  2026-09-23 21:35 ` [PATCH v2 bpf-next 6/8] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG Kuniyuki Iwashima
                   ` (2 subsequent siblings)
  7 siblings, 1 reply; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-23 21:35 UTC (permalink / raw)
  To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
	bpf, netdev

We will add a kfunc for bpf_tcp_ops.{enqueue,dequeue}_rcvq()
to adjust sk->sk_rcvlowat.

These hooks are triggered

  * when the TCP stack enqueues an skb to sk->sk_receive_queue
  * after data is dequeued from sk->sk_receive_queue

In the enqueue path, tcp_data_ready() is always called after
the hooks in tcp_queue_rcv() and tcp_ofo_queue().

If tcp_set_rcvlowat() were used as is, tcp_data_ready() could
be called twice for the same skb, which is redundant and also
confusing.

Let's split out __tcp_set_rcvlowat() and add a flag to control
wakeup behaviour.

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
---
 include/net/tcp.h |  1 +
 net/ipv4/tcp.c    | 12 +++++++++---
 2 files changed, 10 insertions(+), 3 deletions(-)

diff --git a/include/net/tcp.h b/include/net/tcp.h
index f2d838bcb0a7..07426e8641b7 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -512,6 +512,7 @@ void tcp_set_keepalive(struct sock *sk, int val);
 void tcp_syn_ack_timeout(const struct request_sock *req);
 int tcp_recvmsg(struct sock *sk, struct msghdr *msg, size_t len,
 		int flags);
+int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup);
 int tcp_set_rcvlowat(struct sock *sk, int val);
 void tcp_set_rcvbuf(struct sock *sk, int val);
 int tcp_set_window_clamp(struct sock *sk, int val);
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index a714b36a7494..aa7593fc8334 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -1828,8 +1828,7 @@ int tcp_peek_len(struct socket *sock)
 	return tcp_inq(sock->sk);
 }
 
-/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */
-int tcp_set_rcvlowat(struct sock *sk, int val)
+int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup)
 {
 	struct tcp_sock *tp = tcp_sk(sk);
 	int space, cap;
@@ -1842,7 +1841,8 @@ int tcp_set_rcvlowat(struct sock *sk, int val)
 	WRITE_ONCE(sk->sk_rcvlowat, val ? : 1);
 
 	/* Check if we need to signal EPOLLIN right now */
-	tcp_data_ready(sk);
+	if (wakeup)
+		tcp_data_ready(sk);
 
 	if (sk->sk_userlocks & SOCK_RCVBUF_LOCK)
 		return 0;
@@ -1857,6 +1857,12 @@ int tcp_set_rcvlowat(struct sock *sk, int val)
 	return 0;
 }
 
+/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */
+int tcp_set_rcvlowat(struct sock *sk, int val)
+{
+	return __tcp_set_rcvlowat(sk, val, true);
+}
+
 void tcp_set_rcvbuf(struct sock *sk, int val)
 {
 	tcp_set_window_clamp(sk, tcp_win_from_space(sk, val));
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 bpf-next 6/8] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG.
  2026-09-23 21:35 [PATCH v2 bpf-next 0/8] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
                   ` (4 preceding siblings ...)
  2026-09-23 21:35 ` [PATCH v2 bpf-next 5/8] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
@ 2026-09-23 21:35 ` Kuniyuki Iwashima
  2026-09-24  3:48   ` Emil Tsalapatis
  2026-09-23 21:35 ` [PATCH v2 bpf-next 7/8] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
  2026-09-23 21:35 ` [PATCH v2 bpf-next 8/8] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
  7 siblings, 1 reply; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-23 21:35 UTC (permalink / raw)
  To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
	bpf, netdev

The next patch exposes a new kfunc calling __tcp_set_rcvlowat()
to bpf_tcp_ops.

MPTCP has its own sock->ops->set_rcvlowat() / mptcp_set_rcvlowat(),
so we should not allow calling __tcp_set_rcvlowat() on MPTCP
subflows.

Let's disable BPF_SOCK_OPS_RCVQ_CB_FLAG for MPTCP for now.

If needed in the future, bpf_tcp_ops_set_rcvlowat() could be
extended to properly support MPTCP.

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 include/net/tcp.h      | 15 +++++++++++++++
 net/core/filter.c      | 10 ++++++----
 net/ipv4/bpf_tcp_ops.c |  5 ++++-
 3 files changed, 25 insertions(+), 5 deletions(-)

diff --git a/include/net/tcp.h b/include/net/tcp.h
index 07426e8641b7..d3cf655da9ec 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -2932,6 +2932,16 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
 	return tcp_call_bpf(sk, op, 3, args);
 }
 
+static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
+{
+	if (sk_is_mptcp(sk) &&
+	    (val & BPF_SOCK_OPS_RCVQ_CB_FLAG))
+		return -EOPNOTSUPP;
+
+	tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
+	return 0;
+}
+
 static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
 {
 	tcp_sk(sk)->bpf_sock_ops_cb_flags = 0;
@@ -2954,6 +2964,11 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
 	return -EPERM;
 }
 
+static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
+{
+	return -EOPNOTSUPP;
+}
+
 static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
 {
 }
diff --git a/net/core/filter.c b/net/core/filter.c
index 5feb99884682..f29c061bb066 100644
--- a/net/core/filter.c
+++ b/net/core/filter.c
@@ -5588,8 +5588,7 @@ static int bpf_sol_tcp_setsockopt(struct sock *sk, int optname,
 	case TCP_BPF_SOCK_OPS_CB_FLAGS:
 		if (val & ~(BPF_SOCK_OPS_ALL_CB_FLAGS))
 			return -EINVAL;
-		tp->bpf_sock_ops_cb_flags = val;
-		break;
+		return tcp_set_sock_ops_cb_flags(sk, val);
 	default:
 		return -EINVAL;
 	}
@@ -6178,8 +6177,9 @@ static const struct bpf_func_proto bpf_sock_ops_getsockopt_proto = {
 BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
 	   int, argval)
 {
-	struct sock *sk = bpf_sock->sk;
 	int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
+	struct sock *sk = bpf_sock->sk;
+	int err;
 
 	if (!is_locked_tcp_sock_ops(bpf_sock))
 		return -EOPNOTSUPP;
@@ -6187,7 +6187,9 @@ BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
 	if (!IS_ENABLED(CONFIG_INET) || !sk_fullsock(sk))
 		return -EINVAL;
 
-	tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
+	err = tcp_set_sock_ops_cb_flags(sk, val);
+	if (err)
+		return err;
 
 	return argval & (~BPF_SOCK_OPS_ALL_CB_FLAGS);
 }
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index 4b48711d92a2..b0cade34cce6 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -223,8 +223,11 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
 BPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)
 {
 	int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
+	int err;
 
-	tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
+	err = tcp_set_sock_ops_cb_flags(sk, val);
+	if (err)
+		return err;
 
 	return argval & ~BPF_SOCK_OPS_ALL_CB_FLAGS;
 }
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 bpf-next 7/8] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
  2026-09-23 21:35 [PATCH v2 bpf-next 0/8] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
                   ` (5 preceding siblings ...)
  2026-09-23 21:35 ` [PATCH v2 bpf-next 6/8] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG Kuniyuki Iwashima
@ 2026-09-23 21:35 ` Kuniyuki Iwashima
  2026-09-23 22:02   ` sashiko-bot
  2026-09-24  0:30   ` Emil Tsalapatis
  2026-09-23 21:35 ` [PATCH v2 bpf-next 8/8] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
  7 siblings, 2 replies; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-23 21:35 UTC (permalink / raw)
  To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
	bpf, netdev

bpf_tcp_ops.{enqueue,dequeue}_rcvq() were added to parse skb and
adjust sk->sk_rcvlowat dynamically to suppress unnecessary wakeups.

Let's add a new kfunc to set sk->sk_rcvlowat.

Negative values are clamped to INT_MAX, consistent with SO_RCVLOWAT.

For enqueue_rcvq(), wakeup is set to false because:

  * tcp_data_ready() is always called after the hooks in
    tcp_queue_rcv() and tcp_ofo_queue().

  * when tcp_fastopen_add_skb() is called for TFO SYN, the socket is
    not yet accept()ed, and when called for TFO SYN+ACK, the socket
    is woken up by sk->sk_state_change() anyway.

For dequeue_rcvq(), wakeup is set to true because tcp_data_ready()
is not called in that path.

An alternative would be to support bpf_setsockopt() for these
hooks.

However, that approach involves excessive conditionals and an
unnecessary memcpy(), costs we do not want to pay for every skb
in the TCP fast path.

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Tested-by: Clément Léger <cleger@meta.com>
---
 net/ipv4/bpf_tcp_ops.c | 56 +++++++++++++++++++++++++++++++++++++++++-
 1 file changed, 55 insertions(+), 1 deletion(-)

diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index b0cade34cce6..e93c8d13684e 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -368,8 +368,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
 	.owner = THIS_MODULE,
 };
 
+__bpf_kfunc_start_defs();
+
+__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
+					 const struct bpf_prog_aux *aux)
+{
+	u32 moff = aux->attach_st_ops_member_off;
+	bool wakeup = false;
+
+	if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+		wakeup = true;
+
+	if (rcvlowat < 0)
+		rcvlowat = INT_MAX;
+
+	return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);
+}
+
+__bpf_kfunc_end_defs();
+
+BTF_KFUNCS_START(bpf_tcp_ops_rcvlowat_kfunc_set)
+BTF_ID_FLAGS(func, bpf_tcp_ops_set_rcvlowat, KF_IMPLICIT_ARGS)
+BTF_KFUNCS_END(bpf_tcp_ops_rcvlowat_kfunc_set)
+
+static int bpf_tcp_ops_rcvlowat_kfunc_filter(const struct bpf_prog *prog,
+					     u32 kfunc_id)
+{
+	u32 moff;
+
+	if (!btf_id_set8_contains(&bpf_tcp_ops_rcvlowat_kfunc_set, kfunc_id))
+		return 0;
+
+	if (prog->aux->st_ops != &bpf_tcp_ops)
+		return -EACCES;
+
+	moff = prog->aux->attach_st_ops_member_off;
+	if (moff != offsetof(struct bpf_tcp_ops, enqueue_rcvq) &&
+	    moff != offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+		return -EACCES;
+
+	return 0;
+}
+
+static const struct btf_kfunc_id_set bpf_tcp_ops_rcvlowat_kfunc_id_set = {
+	.owner = THIS_MODULE,
+	.set = &bpf_tcp_ops_rcvlowat_kfunc_set,
+	.filter = bpf_tcp_ops_rcvlowat_kfunc_filter,
+};
+
 static int __init __bpf_tcp_ops_init(void)
 {
-	return register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
+	int ret;
+
+	ret = register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS,
+					&bpf_tcp_ops_rcvlowat_kfunc_id_set);
+	ret = ret ?: register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
+
+	return ret;
 }
 late_initcall(__bpf_tcp_ops_init);
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 bpf-next 8/8] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  2026-09-23 21:35 [PATCH v2 bpf-next 0/8] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
                   ` (6 preceding siblings ...)
  2026-09-23 21:35 ` [PATCH v2 bpf-next 7/8] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
@ 2026-09-23 21:35 ` Kuniyuki Iwashima
  7 siblings, 0 replies; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-23 21:35 UTC (permalink / raw)
  To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
	bpf, netdev

The test is roughly divided into two stages, and the sequence
is as follows:

  I) Setup

    1. Attach two BPF programs to a cgroup
    2. Establish a TCP connection (@client <-> @child) within the cgroup
    3. Enable BPF_SOCK_OPS_RCVQ_CB_FLAG on @child via setsockopt()

 II) RPC frame exchange in various patterns

    4. Send a partial RPC descriptor from @client to @child
    5. Verify that epoll does NOT wake up @child
    6. Send the remaining data of the RPC frame
    7. Verify that epoll finally wakes up @child

During setup, two BPF programs are attached to simulate
a real-world scenario; one is bpf_tcp_ops and the other is
CGROUP_SOCKOPT.

While the bpf_tcp_ops prog handles the dynamic adjustment of
sk->sk_rcvlowat, the CGROUP_SOCKOPT prog is used to enable
the TCP AutoLOWAT feature via userspace setsockopt() using
pseudo options:

  #define SOL_BPF               0xdeadbeef
  #define BPF_TCP_AUTOLOWAT     0x8badf00d

  setsockopt(fd, SOL_BPF, BPF_TCP_AUTOLOWAT, &(int){1}, sizeof(int));

This reflects a common production use case where an application
decides to start parsing RPC frames only at a certain point in
the stream (e.g., after HTTP Upgrade), rather than immediately
after TCP 3WHS (BPF_SOCK_OPS_PASSIVE_ESTABLISHED_CB, etc).

When BPF_TCP_AUTOLOWAT is enabled, the BPF prog sets
BPF_SOCK_OPS_RCVQ_CB_FLAG and initializes sk_local_storage
for two sequence numbers to manage its state.

Then, for the RPC frame exchange, this test uses a simple format
defined as follows:

  0        8       16      24       32
  +--------+--------+-------+--------+ `.
  |            header size           |  |
  +--------+--------+-------+--------+   > RPC descriptor (8 bytes)
  |            payload size          |  |
  +--------+--------+-------+--------+ .'
  ~               header             ~
  +--------+--------+-------+--------+
  ~               payload            ~
  +--------+--------+-------+--------+

Every time a new skb is enqueued to sk->sk_receive_queue, the
bpf_tcp_ops prog parses it and updates these sequence numbers:

  rpc_desc_seq : the SEQ # of the start of the RPC descriptor
  rpc_end_seq  : the SEQ # of the end of the RPC frame
                 => rpc_desc_seq + 8 + header size + payload size

Assume we receive two RPC descriptors in the following pattern:

  1. When we receive skb-1, only part of the RPC descriptor is parsed.
     rpc_desc_seq is set to the first byte while rpc_end_seq is
     unknown.  Thus, sk->sk_rcvlowat is set to the size of the RPC
     descriptor (8 bytes).

   <- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
  +-----------+.................+....................+......
  |  RPC desc 1  |  header + payload  |  RPC desc 2  | ...
  +-----------+.................+....................+......
  ^              ^-.
  `- rpc_desc_seq   `- sk->sk_rcvlowat

  2. Next, we receive skb-2, which completes the first RPC descriptor.
     Now rpc_end_seq is known, so sk->sk_rcvlowat is advanced to it.

   <- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
  +-----------+-----------------+....................+......
  |  RPC desc 1  |  header + payload  |  RPC desc 2  | ...
  +-----------+-----------------+....................+......
  ^                                   ^
  '- rpc_desc_seq                     '- rpc_end_seq
                                           & sk->sk_rcvlowat

  3. Once we receive skb-3, which contains the next full RPC descriptor,
     rpc_desc_seq is advanced and rpc_end_seq is updated according
     to the size of RPC frame 2.

     Note that sk->sk_rcvlowat is NOT updated to the new rpc_end_seq
     yet.  This ensures that the application is woken up to read the
     already complete RPC frame 1.

   <- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
  +-----------+-----------------+--------------------+......
  |  RPC desc 1  |  header + payload  |  RPC desc 2  | ...   |
  +-----------+-----------------+--------------------+......
                                      ^                      ^
              rpc_desc_seq -----------'  rpc_end_seq ----...-'
                & sk->sk_rcvlowat

This sequence corresponds to the 4th test case in rpc_test_cases[],
and we can see helpful output if we "#define DEBUG":

  # cat /sys/kernel/tracing/trace_pipe | \
    awk '{ if ($0 ~ /AF_/) sub(/^.*AF_/, "AF_"); print $0 }' & \
    BGPID=$!; ./test_progs -t tcp_autolowat; kill -9 -$BGPID
  ...
  AF_INET6 rpc_test_cases[3]: Start parsing skb: seq: 0, end_seq: 1, len: 1, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_buff_len: 0
  AF_INET6 rpc_test_cases[3]: Copied 1 bytes: rpc_desc_buff_len: 1
  AF_INET6 rpc_test_cases[3]: Setting rcvlowat: tp->copied_seq: 0, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_desc_buff_len: 1
  AF_INET6 rpc_test_cases[3]: Set rcvlowat: expected: 8, actual: 8

  AF_INET6 rpc_test_cases[3]: Start parsing skb: seq: 1, end_seq: 8, len: 7, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_buff_len: 1
  AF_INET6 rpc_test_cases[3]: Copied full descriptor: rpc_desc_seq: 0, rpc_end_seq: 258, header_len: 100, payload_len: 150
  AF_INET6 rpc_test_cases[3]: No more descriptor: rpc_end_seq: 258, end_seq: 8
  AF_INET6 rpc_test_cases[3]: Setting rcvlowat: tp->copied_seq: 0, rpc_desc_seq: 0, rpc_end_seq: 258, rpc_desc_buff_len: 8
  AF_INET6 rpc_test_cases[3]: Set rcvlowat: expected: 258, actual: 258
  ...

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
---
v2:
  * Avoid address comparison for a specific version of gcc.
  * Make rpc_test_cases[] static.
  * Update comment in rpc_test_case[].
---
 .../selftests/bpf/prog_tests/tcp_autolowat.c  | 350 ++++++++++++++++++
 .../selftests/bpf/progs/bpf_tracing_net.h     |   2 +
 .../selftests/bpf/progs/tcp_autolowat.c       | 312 ++++++++++++++++
 3 files changed, 664 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
 create mode 100644 tools/testing/selftests/bpf/progs/tcp_autolowat.c

diff --git a/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
new file mode 100644
index 000000000000..337f9d34a39c
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
@@ -0,0 +1,350 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright 2026 Google LLC */
+#include <sys/epoll.h>
+
+#include "test_progs.h"
+#include "cgroup_helpers.h"
+#include "network_helpers.h"
+
+#include "tcp_autolowat.skel.h"
+
+#define SOL_BPF			0xdeadbeef
+#define BPF_TCP_AUTOLOWAT	0x8badf00d
+
+struct rpc_descriptor {
+	u32 header_len;
+	u32 payload_len;
+};
+
+enum rpc_event_type {
+	RPC_EVENT_END,
+	RPC_EVENT_AUTOLOWAT,
+	RPC_EVENT_SEND,
+	RPC_EVENT_RECV,
+	RPC_EVENT_EPOLL,
+	RPC_EVENT_RCVLOWAT,
+};
+
+struct rpc_event {
+	enum rpc_event_type type;
+	union {
+		int len;
+		int nfds;
+		int val;
+		int rcvlowat;
+	};
+};
+
+#define RPC_DESC_SIZE (sizeof(struct rpc_descriptor))
+
+static struct rpc_test_case {
+	char data[4096];
+	struct rpc_descriptor desc[32];
+	struct rpc_event event[32];
+} rpc_test_cases[] = {
+	{
+		.desc = {
+			{ .header_len = 100, .payload_len = 150 },
+		},
+		.event = {
+			{ .type = RPC_EVENT_AUTOLOWAT,	.val = 1},
+			/* Single full RPC message in skb. */
+			{ .type = RPC_EVENT_SEND,	.len = RPC_DESC_SIZE + 100 + 150},
+			{ .type = RPC_EVENT_EPOLL,	.nfds = 1},
+			{ .type = RPC_EVENT_RCVLOWAT,	.rcvlowat = RPC_DESC_SIZE + 100 + 150},
+		},
+	},
+	{
+		.desc = {
+			{.header_len = 100, .payload_len = 150},
+			{.header_len = 100, .payload_len = 150},
+			{.header_len = 100, .payload_len = 150},
+		},
+		.event = {
+			{ .type = RPC_EVENT_AUTOLOWAT,	.val = 1},
+			/* Two full RPC messages in skb. */
+			{.type = RPC_EVENT_SEND,	.len = (RPC_DESC_SIZE + 100 + 150) * 2},
+			{.type = RPC_EVENT_EPOLL,	.nfds = 1},
+			{.type = RPC_EVENT_RCVLOWAT,	.rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+			/* Single full RPC message in skb. */
+			{ .type = RPC_EVENT_SEND,	.len = RPC_DESC_SIZE + 100 + 150},
+			{ .type = RPC_EVENT_EPOLL,	.nfds = 1},
+			{ .type = RPC_EVENT_RCVLOWAT,	.rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 3},
+		},
+	},
+	{
+		.desc = {
+			{.header_len = 100, .payload_len = 150},
+			{.header_len = 100, .payload_len = 150},
+			{.header_len = 100, .payload_len = 150},
+		},
+		.event = {
+			{ .type = RPC_EVENT_AUTOLOWAT,	.val = 1},
+			/* Two full RPC messages in skb. */
+			{.type = RPC_EVENT_SEND,	.len = (RPC_DESC_SIZE + 100 + 150) * 2},
+			{.type = RPC_EVENT_EPOLL,	.nfds = 1},
+			{.type = RPC_EVENT_RCVLOWAT,	.rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+			/* Only the next descriptor in skb. */
+			{ .type = RPC_EVENT_SEND,	.len = RPC_DESC_SIZE},
+			{ .type = RPC_EVENT_EPOLL,	.nfds = 1},
+			{ .type = RPC_EVENT_RCVLOWAT,	.rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+		},
+	},
+	{
+		.desc = {
+			{.header_len = 100, .payload_len = 150},
+			{.header_len = 200, .payload_len = 500},
+		},
+		.event = {
+			{ .type = RPC_EVENT_AUTOLOWAT,	.val = 1},
+			/* The first descriptor is partial. */
+			{.type = RPC_EVENT_SEND,	.len = 1},
+			{.type = RPC_EVENT_EPOLL,	.nfds = 0},
+			{.type = RPC_EVENT_RCVLOWAT,	.rcvlowat = RPC_DESC_SIZE},
+			/* The first descriptor is available. */
+			{.type = RPC_EVENT_SEND,	.len = RPC_DESC_SIZE - 1},
+			{.type = RPC_EVENT_EPOLL,	.nfds = 0},
+			{.type = RPC_EVENT_RCVLOWAT,	.rcvlowat = RPC_DESC_SIZE + 150 + 100},
+			/* The first header is ready. */
+			{.type = RPC_EVENT_SEND,	.len = 100},
+			{.type = RPC_EVENT_EPOLL,	.nfds = 0},
+			{.type = RPC_EVENT_RCVLOWAT,	.rcvlowat = RPC_DESC_SIZE + 150 + 100},
+			/* skb has the first payload and 1 byte of the next descriptor. */
+			{.type = RPC_EVENT_SEND,	.len = 150 + 1},
+			{.type = RPC_EVENT_EPOLL,	.nfds = 1},
+			{.type = RPC_EVENT_RCVLOWAT,	.rcvlowat = RPC_DESC_SIZE + 150 + 100},
+			/* After reading the first RPC message, SO_RCVLOWAT should be RPC_DESC_SIZE. */
+			{.type = RPC_EVENT_RECV,	.len = RPC_DESC_SIZE + 150 + 100},
+			{.type = RPC_EVENT_EPOLL,	.nfds = 0},
+			{.type = RPC_EVENT_RCVLOWAT,	.rcvlowat = RPC_DESC_SIZE},
+			/* The second descriptor is available. */
+			{.type = RPC_EVENT_SEND,	.len = RPC_DESC_SIZE - 1},
+			{.type = RPC_EVENT_EPOLL,	.nfds = 0},
+			{.type = RPC_EVENT_RCVLOWAT,	.rcvlowat = RPC_DESC_SIZE + 200 + 500},
+		},
+	},
+};
+
+struct tcp_autolowat_test_cb {
+	int saved_netns;
+	union {
+		int fd[4];
+		struct {
+			int server, client, child;
+			int epoll;
+		};
+	};
+};
+
+static void tcp_autolowat_teardown_cb(struct tcp_autolowat_test_cb *cb)
+{
+	int i, err;
+
+	for (i = 0; i < ARRAY_SIZE(cb->fd); i++) {
+		if (cb->fd[i] != -1)
+			close(cb->fd[i]);
+	}
+
+	if (cb->saved_netns != -1) {
+		err = setns(cb->saved_netns, CLONE_NEWNET);
+		ASSERT_OK(err, "restore netns");
+
+		close(cb->saved_netns);
+	}
+}
+
+static int tcp_autolowat_setup_cb(struct tcp_autolowat_test_cb *cb, int family)
+{
+	struct epoll_event ev = {};
+	int err;
+	int i;
+
+	for (i = 0; i < ARRAY_SIZE(cb->fd); i++)
+		cb->fd[i] = -1;
+
+	cb->saved_netns = open("/proc/self/ns/net", O_RDONLY);
+	if (!ASSERT_OK_FD(cb->saved_netns, "save netns"))
+		goto err;
+
+	err = unshare(CLONE_NEWNET);
+	if (!ASSERT_OK(err, "unshare"))
+		goto err;
+
+	err = system("ip link set dev lo up");
+	if (!ASSERT_OK(err, "set up lo"))
+		goto err;
+
+	cb->server = start_server(family, SOCK_STREAM, NULL, 0, 0);
+	if (!ASSERT_OK_FD(cb->server, "start_server"))
+		goto err;
+
+	cb->client = connect_to_fd(cb->server, 0);
+	if (!ASSERT_OK_FD(cb->client, "connect_to_fd"))
+		goto err;
+
+	cb->child = accept(cb->server, NULL, NULL);
+	if (!ASSERT_OK_FD(cb->child, "accept"))
+		goto err;
+
+	cb->epoll = epoll_create1(0);
+	if (!ASSERT_OK_FD(cb->epoll, "epoll_create"))
+		goto err;
+
+	ev.events = EPOLLIN;
+	ev.data.fd = cb->child;
+
+	err = epoll_ctl(cb->epoll, EPOLL_CTL_ADD, cb->child, &ev);
+	if (!ASSERT_OK(err, "epoll_ctl"))
+		goto err;
+
+	return 0;
+
+err:
+	tcp_autolowat_teardown_cb(cb);
+	return -1;
+}
+
+static int tcp_autolowat_build_data(struct rpc_test_case *test_case)
+{
+	struct rpc_descriptor *desc = test_case->desc;
+	char *ptr = test_case->data;
+	int rpc_size;
+
+	memset(ptr, 0, sizeof(test_case->data));
+
+	while (desc->header_len + desc->payload_len) {
+		rpc_size = sizeof(*desc) + desc->header_len + desc->payload_len;
+
+		if (!ASSERT_LE(ptr + rpc_size - test_case->data,
+			       sizeof(test_case->data), "data overflow"))
+			return 1;
+
+		memcpy(ptr, desc, sizeof(*desc));
+		ptr += rpc_size;
+		desc++;
+	}
+
+	if (!ASSERT_GT(ptr - test_case->data, 0, "no data"))
+		return 1;
+
+	return 0;
+}
+
+static void tcp_autolowat_run_rpc_test(struct tcp_autolowat_test_cb *cb,
+				       struct rpc_test_case *test_case)
+{
+	struct rpc_event *event = test_case->event;
+	char *ptr = test_case->data;
+	struct epoll_event ev;
+	socklen_t optlen;
+	int err, optval;
+	char buf[4096];
+
+	if (tcp_autolowat_build_data(test_case))
+		return;
+
+	while (1) {
+		switch (event->type) {
+		case RPC_EVENT_END:
+			return;
+		case RPC_EVENT_AUTOLOWAT:
+			err = setsockopt(cb->child, SOL_BPF, BPF_TCP_AUTOLOWAT,
+					 &event->val, sizeof(event->val));
+			if (!ASSERT_OK(err, "setsockopt"))
+				return;
+			break;
+		case RPC_EVENT_SEND:
+			err = send(cb->client, ptr, event->len, 0);
+			if (!ASSERT_EQ(err, event->len, "send"))
+				return;
+
+			ptr += event->len;
+			break;
+		case RPC_EVENT_RECV:
+			err = recv(cb->child, buf, event->len, 0);
+			if (!ASSERT_EQ(err, event->len, "recv"))
+				return;
+			break;
+		case RPC_EVENT_EPOLL:
+			err = epoll_wait(cb->epoll, &ev, 1, 100);
+			if (!ASSERT_EQ(err, event->nfds, "epoll_wait"))
+				return;
+			break;
+		case RPC_EVENT_RCVLOWAT:
+			optval = 0;
+			optlen = sizeof(optval);
+
+			err = getsockopt(cb->child, SOL_SOCKET, SO_RCVLOWAT, &optval, &optlen);
+			if (!ASSERT_OK(err, "getsockopt") ||
+			    !ASSERT_EQ(optval, event->rcvlowat, "rcvlowat"))
+				return;
+			break;
+		}
+
+		event++;
+	}
+}
+
+static void tcp_autolowat_run_rpc_tests(struct tcp_autolowat *skel, int family)
+{
+	struct tcp_autolowat_test_cb cb;
+	int err;
+	int i;
+
+	for (i = 0; i < ARRAY_SIZE(rpc_test_cases); i++) {
+		memset(skel->bss->test_name, 0, sizeof(skel->bss->test_name));
+
+		snprintf(skel->bss->test_name, sizeof(skel->bss->test_name),
+			 "AF_INET%c rpc_test_cases[%d]",
+			 family == AF_INET ? ' ' : '6', i);
+
+		if (!test__start_subtest(skel->bss->test_name))
+			continue;
+
+		err = tcp_autolowat_setup_cb(&cb, family);
+		if (err)
+			continue;
+
+		tcp_autolowat_run_rpc_test(&cb, &rpc_test_cases[i]);
+		tcp_autolowat_teardown_cb(&cb);
+	}
+}
+
+static void tcp_autolowat_run_tests(struct tcp_autolowat *skel)
+{
+	tcp_autolowat_run_rpc_tests(skel, AF_INET);
+	tcp_autolowat_run_rpc_tests(skel, AF_INET6);
+}
+
+void test_tcp_autolowat(void)
+{
+	struct tcp_autolowat *skel;
+	struct bpf_link *link[2];
+	int cgroup;
+
+	skel = tcp_autolowat__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "open_and_load"))
+		return;
+
+	cgroup = test__join_cgroup("/tcp_autolowat");
+	if (!ASSERT_GE(cgroup, 0, "join_cgroup"))
+		goto destroy_skel;
+
+	link[0] = bpf_map__attach_cgroup_opts(skel->maps.tcp_autolowat_ops, cgroup, NULL);
+	if (!ASSERT_OK_PTR(link[0], "attach_cgroup(tcp_autolowat_ops)"))
+		goto close_cgroup;
+
+	link[1] = bpf_program__attach_cgroup(skel->progs.tcp_autolowat_setsockopt, cgroup);
+	if (!ASSERT_OK_PTR(link[1], "attach_cgroup(SETSOCKOPT)"))
+		goto destroy_sockops;
+
+	tcp_autolowat_run_tests(skel);
+
+	bpf_link__destroy(link[1]);
+destroy_sockops:
+	bpf_link__destroy(link[0]);
+close_cgroup:
+	close(cgroup);
+destroy_skel:
+	tcp_autolowat__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/progs/bpf_tracing_net.h b/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
index 593b38f90417..4c999d59cbbc 100644
--- a/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
+++ b/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
@@ -79,6 +79,8 @@
 
 #define NEXTHDR_TCP		6
 
+#define TCPHDR_FIN		0x01
+
 #define TCPOPT_NOP		1
 #define TCPOPT_EOL		0
 #define TCPOPT_MSS		2
diff --git a/tools/testing/selftests/bpf/progs/tcp_autolowat.c b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
new file mode 100644
index 000000000000..4fe7bdc98e69
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
@@ -0,0 +1,312 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright 2026 Google LLC */
+#include "vmlinux.h"
+
+#include <string.h>
+#include <limits.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+#include <bpf/bpf_core_read.h>
+
+#include "bpf_kfuncs.h"
+#include "bpf_tracing_net.h"
+
+#define SOL_BPF			0xdeadbeef
+#define BPF_TCP_AUTOLOWAT	0x8badf00d
+
+//#define DEBUG /* For verbose output. */
+
+struct rpc_descriptor {
+	u32 header_len;
+	u32 payload_len;
+};
+
+#define RPC_DESC_SIZE		(sizeof(struct rpc_descriptor))
+#define MAX_RPC_DESC_PER_SKB	100
+
+struct tcp_autolowat_cb {
+	/* Don't put this field at the end; BPF verifier complains. */
+	char rpc_desc_buf[RPC_DESC_SIZE];
+	u32 rpc_desc_seq;
+	u32 rpc_end_seq;
+#ifdef DEBUG
+	u32 isn;
+#endif
+	u8 rpc_desc_buff_len;
+};
+
+struct {
+	__uint(type, BPF_MAP_TYPE_SK_STORAGE);
+	__uint(map_flags, BPF_F_NO_PREALLOC);
+	__type(key, int);
+	__type(value, struct tcp_autolowat_cb);
+} tcp_autolowat_map SEC(".maps");
+
+char test_name[64];
+
+#ifdef DEBUG
+#define LOG(str, ...)							\
+	bpf_printk("%s: " str, test_name, ##__VA_ARGS__)
+#else
+#define LOG(...)
+#endif
+
+#define SEQ(val)				\
+	(val - cb->isn)
+#define TP_SEQ(field)				\
+	(tp->field - cb->isn)
+#define CB_SEQ(field)				\
+	(cb->field - cb->isn)
+
+static int tcp_parse_descriptor(struct tcp_autolowat_cb *cb,
+				struct bpf_dynptr *dptr,
+				u32 seq, u32 end_seq)
+{
+	struct rpc_descriptor *rpc_desc;
+	u32 rpc_copied_seq;
+	u64 copy_len; /* u32 should work, but not for no_alu32 :/ */
+	u64 rpc_len;
+	int err;
+
+	rpc_copied_seq = cb->rpc_desc_seq + cb->rpc_desc_buff_len;
+
+	if (before(cb->rpc_desc_seq + RPC_DESC_SIZE, end_seq))
+		copy_len = RPC_DESC_SIZE - cb->rpc_desc_buff_len;
+	else
+		copy_len = end_seq - rpc_copied_seq;
+
+	if (copy_len == 0)
+		goto disable; /* FIN. */
+	if (copy_len > RPC_DESC_SIZE)
+		goto disable; /* always false, only for verifier. */
+	if (cb->rpc_desc_buff_len >= RPC_DESC_SIZE)
+		goto disable; /* always false, only for verifier. */
+
+	err = bpf_dynptr_read(cb->rpc_desc_buf + cb->rpc_desc_buff_len,
+			      copy_len, dptr, rpc_copied_seq - seq, 0);
+	if (err)
+		goto disable;
+
+	cb->rpc_desc_buff_len += copy_len;
+
+	if (cb->rpc_desc_buff_len != RPC_DESC_SIZE) {
+		LOG("Copied %d bytes: rpc_desc_buff_len: %u", copy_len, cb->rpc_desc_buff_len);
+		goto partial;
+	}
+
+	rpc_desc = (struct rpc_descriptor *)cb->rpc_desc_buf;
+	rpc_len = RPC_DESC_SIZE + rpc_desc->header_len + rpc_desc->payload_len;
+
+	if (rpc_len > INT_MAX)
+		goto disable;
+
+	cb->rpc_end_seq = cb->rpc_desc_seq + rpc_len;
+
+	LOG("Copied full descriptor: rpc_desc_seq: %u, rpc_end_seq: %u, header_len: %u, payload_len: %u",
+	    CB_SEQ(rpc_desc_seq), CB_SEQ(rpc_end_seq),
+	    rpc_desc->header_len, rpc_desc->payload_len);
+
+	return 0;
+disable:
+	return -1;
+partial:
+	return 1;
+}
+
+static void tcp_set_autolowat(struct tcp_autolowat_cb *cb,
+			      struct sock *sk)
+{
+	struct tcp_sock *tp = (struct tcp_sock *)sk;
+	u32 val; /* To handle wraparound. */
+
+	LOG("Setting rcvlowat: tp->copied_seq: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_desc_buff_len: %u",
+	    TP_SEQ(copied_seq), CB_SEQ(rpc_desc_seq),
+	    CB_SEQ(rpc_end_seq), cb->rpc_desc_buff_len);
+
+	if (before(tp->copied_seq, cb->rpc_desc_seq))
+		val = cb->rpc_desc_seq - tp->copied_seq;
+	else if (cb->rpc_desc_buff_len != RPC_DESC_SIZE)
+		val = RPC_DESC_SIZE;
+	else
+		val = cb->rpc_end_seq - tp->copied_seq;
+
+	if (val != tp->inet_conn.icsk_inet.sk.sk_rcvlowat) {
+		bpf_tcp_ops_set_rcvlowat(sk, val);
+
+		LOG("Set rcvlowat: expected: %u, actual: %d\n",
+		    val, tp->inet_conn.icsk_inet.sk.sk_rcvlowat);
+	} else {
+		LOG("No need to set rcvlowat: %u\n", val);
+	}
+}
+
+static void tcp_disable_autolowat(struct sock *sk)
+{
+	struct tcp_sock *tp = (struct tcp_sock *)sk;
+	int flags;
+
+	flags = tp->bpf_sock_ops_cb_flags & ~BPF_SOCK_OPS_RCVQ_CB_FLAG;
+	bpf_sock_ops_cb_flags_set(sk, flags);
+
+	bpf_tcp_ops_set_rcvlowat(sk, 1);
+
+	LOG("Disabled autolowat");
+}
+
+static void tcp_do_autolowat(struct tcp_autolowat_cb *cb,
+			     struct sock *sk, struct sk_buff *skb)
+{
+	struct bpf_dynptr dptr;
+	struct tcp_skb_cb *tcb;
+	u32 seq, end_seq;
+	int ret = 0, i;
+
+	if (bpf_dynptr_from_skb((struct __sk_buff *)skb, 0, &dptr)) {
+		ret = -1;
+		goto update;
+	}
+
+	tcb = bpf_core_cast(skb->cb, struct tcp_skb_cb);
+	seq = tcb->seq;
+	end_seq = tcb->end_seq - !!(tcb->tcp_flags & TCPHDR_FIN);
+
+	LOG("Start parsing skb: seq: %u, end_seq: %u, len: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_buff_len: %u",
+	    SEQ(seq), SEQ(end_seq), end_seq - seq,
+	    CB_SEQ(rpc_desc_seq), CB_SEQ(rpc_end_seq), cb->rpc_desc_buff_len);
+
+	if (cb->rpc_desc_buff_len != RPC_DESC_SIZE) {
+		ret = tcp_parse_descriptor(cb, &dptr, seq, end_seq);
+		if (ret)
+			goto update;
+	}
+
+	i = 0;
+
+	while (1) {
+		if (i++ > MAX_RPC_DESC_PER_SKB) {
+			ret = -1;
+			break;
+		}
+
+		if (after(cb->rpc_end_seq, end_seq)) {
+			LOG("No more descriptor: rpc_end_seq: %u, end_seq: %u",
+			    CB_SEQ(rpc_end_seq), SEQ(end_seq));
+			break;
+		}
+
+		cb->rpc_desc_seq = cb->rpc_end_seq;
+		cb->rpc_desc_buff_len = 0;
+
+		if (cb->rpc_end_seq == end_seq)
+			break;
+
+		LOG("Found next descriptor: rpc_end_seq: %u, end_seq: %u, len: %u",
+		    CB_SEQ(rpc_end_seq), SEQ(end_seq), end_seq - cb->rpc_end_seq);
+
+		ret = tcp_parse_descriptor(cb, &dptr, seq, end_seq);
+		if (ret)
+			break;
+	}
+
+update:
+	if (ret >= 0)
+		tcp_set_autolowat(cb, sk);
+	else
+		tcp_disable_autolowat(sk);
+}
+
+SEC("struct_ops")
+void BPF_PROG(tcp_autolowat_enqueue_rcvq, struct sock *sk, struct sk_buff *skb)
+{
+	struct tcp_autolowat_cb *cb;
+
+	cb = bpf_sk_storage_get(&tcp_autolowat_map, sk, 0, 0);
+	if (!cb)
+		return;
+
+	tcp_do_autolowat(cb, sk, skb);
+}
+
+SEC("struct_ops")
+void BPF_PROG(tcp_autolowat_dequeue_rcvq, struct sock *sk)
+{
+	struct tcp_autolowat_cb *cb;
+
+	cb = bpf_sk_storage_get(&tcp_autolowat_map, sk, 0, 0);
+	if (!cb)
+		return;
+
+	tcp_set_autolowat(cb, sk);
+}
+
+SEC(".struct_ops.link")
+struct bpf_tcp_ops tcp_autolowat_ops = {
+	.enqueue_rcvq = (void *)tcp_autolowat_enqueue_rcvq,
+	.dequeue_rcvq = (void *)tcp_autolowat_dequeue_rcvq,
+};
+
+static int tcp_init_autolowat_cb(struct bpf_sockopt *sockopt,
+				 struct bpf_tcp_sock *btp)
+{
+	struct tcp_autolowat_cb *cb;
+	struct tcp_sock *tp;
+	int flags;
+
+	cb = bpf_sk_storage_get(&tcp_autolowat_map, btp, 0,
+				BPF_SK_STORAGE_GET_F_CREATE);
+	if (!cb)
+		return -1;
+
+	tp = bpf_core_cast(btp, struct tcp_sock);
+	if (!tp)
+		return -1;
+
+	cb->rpc_desc_seq = tp->copied_seq;
+	cb->rpc_end_seq = tp->copied_seq;
+#ifdef DEBUG
+	cb->isn = tp->copied_seq;
+#endif
+
+	if (bpf_getsockopt(sockopt->sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
+			   &flags, sizeof(flags)))
+		return -1;
+
+	flags |= BPF_SOCK_OPS_RCVQ_CB_FLAG;
+
+	if (bpf_setsockopt(sockopt->sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
+			   &flags, sizeof(flags)))
+		return -1;
+
+	return 0;
+}
+
+SEC("cgroup/setsockopt")
+int tcp_autolowat_setsockopt(struct bpf_sockopt *ctx)
+{
+	void *optval_end = ctx->optval_end;
+	int *optval = ctx->optval;
+	struct bpf_tcp_sock *btp;
+
+	if (ctx->level != SOL_BPF || ctx->optname != BPF_TCP_AUTOLOWAT)
+		goto out;
+
+	if (optval + 1 > optval_end)
+		return 0; /* -EPERM */
+
+	btp = bpf_tcp_sock(ctx->sk);
+	if (!btp)
+		goto out;
+
+	if (*optval && tcp_init_autolowat_cb(ctx, btp))
+		return 0; /* -EPERM */
+
+	/*
+	 * BPF has consumed this option, don't call kernel
+	 * setsockopt handler.
+	 */
+	ctx->optlen = -1;
+out:
+	return 1;
+}
+
+char _license[] SEC("license") = "GPL";
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  2026-09-23 21:35 ` [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
@ 2026-09-23 21:50   ` sashiko-bot
  2026-09-23 22:42   ` Emil Tsalapatis
  2026-09-25  0:04   ` Alexei Starovoitov
  2 siblings, 0 replies; 29+ messages in thread
From: sashiko-bot @ 2026-09-23 21:50 UTC (permalink / raw)
  To: Kuniyuki Iwashima; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] bpf: multi-line comment style violation for BPF_SOCK_OPS_RCVQ_CB_FLAG
--

commit 304bdb98cffa8e2d03b8207be4eaaadba075d6be
Author: Kuniyuki Iwashima <kuniyu@google.com>

bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().

This commit introduces opt-in callbacks to bpf_tcp_ops that are invoked when
the TCP stack enqueues or dequeues data to the receive queue. These callbacks
allow a BPF program to parse each skb and dynamically adjust sk_rcvlowat to
suppress unnecessary EPOLLIN wakeups.

[ ... ]

> diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
> index 0aaa54359aebc..3f0877b5867bd 100644
> --- a/include/uapi/linux/bpf.h
> +++ b/include/uapi/linux/bpf.h
> @@ -7148,8 +7148,17 @@ enum {
>  	 * options first before the BPF program does.
>  	 */
>  	BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
> +	/* Call bpf when the TCP stack enqueues/dequeues payload
> +	 * to/from sk->sk_receive_queue.

[Severity: Low]
This isn't a bug, but does this comment violate the BPF subsystem multi-line
comment style? The BPF guidelines request that multi-line comments have the
opening /* on its own line.

> +	 *
> +	 * Only bpf_tcp_ops is supported.
> +	 *
> +	 * It can be used to adjust sk->sk_rcvlowat and suppress
> +	 * unnecessary wakeups before sufficient data is available.
> +	 */
> +	BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),

[ ... ]

> diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
> index 0aaa54359aebc..3f0877b5867bd 100644
> --- a/tools/include/uapi/linux/bpf.h
> +++ b/tools/include/uapi/linux/bpf.h
> @@ -7148,8 +7148,17 @@ enum {
>  	 * options first before the BPF program does.
>  	 */
>  	BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
> +	/* Call bpf when the TCP stack enqueues/dequeues payload
> +	 * to/from sk->sk_receive_queue.

[Severity: Low]
This isn't a bug, but this appears to be the same multi-line comment style
violation mirrored into the tools header.

> +	 *
> +	 * Only bpf_tcp_ops is supported.
> +	 *
> +	 * It can be used to adjust sk->sk_rcvlowat and suppress
> +	 * unnecessary wakeups before sufficient data is available.
> +	 */
> +	BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923213719.224838-1-kuniyu@google.com?part=3

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 7/8] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
  2026-09-23 21:35 ` [PATCH v2 bpf-next 7/8] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
@ 2026-09-23 22:02   ` sashiko-bot
  2026-09-23 22:21     ` Kuniyuki Iwashima
  2026-09-24  0:30   ` Emil Tsalapatis
  1 sibling, 1 reply; 29+ messages in thread
From: sashiko-bot @ 2026-09-23 22:02 UTC (permalink / raw)
  To: Kuniyuki Iwashima; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: tcp: OOB write in __tcp_set_rcvlowat via unrelated socket
--

commit 68bce63bbe5802a4eee99c7c0f138096d7c16481
Author: Kuniyuki Iwashima <kuniyu@google.com>

bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.

This commit adds a new BPF kfunc bpf_tcp_ops_set_rcvlowat to allow
struct_ops programs to dynamically adjust sk->sk_rcvlowat. It clamps
negative values to INT_MAX and delegates the update to __tcp_set_rcvlowat().

> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index b0cade34cce65..e93c8d13684ef 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
[ ... ]
> @@ -368,8 +368,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
>  	.owner = THIS_MODULE,
>  };
>  
> +__bpf_kfunc_start_defs();
> +
> +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> +					 const struct bpf_prog_aux *aux)
> +{
> +	u32 moff = aux->attach_st_ops_member_off;
> +	bool wakeup = false;
> +
> +	if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
> +		wakeup = true;
> +
> +	if (rcvlowat < 0)
> +		rcvlowat = INT_MAX;
> +
> +	return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);

[Severity: High]
Does this code allow an out-of-bounds write via an unrelated socket type?

If a privileged user loads a TCP struct_ops BPF program with an enqueue_rcvq
callback, the program can extract skb->sk. After a NULL check, the verifier
treats it as a trusted struct sock.

If the packet was sent over the loopback interface by a UDP socket, skb->sk
points to a struct udp_sock. When this socket pointer is passed to
bpf_tcp_ops_set_rcvlowat(), it is handed directly to __tcp_set_rcvlowat()
which unconditionally casts it to struct tcp_sock:

net/ipv4/tcp.c:__tcp_set_rcvlowat() {
    struct tcp_sock *tp = tcp_sk(sk);
    ...
        if (tp->window_clamp && tp->window_clamp < val)
            WRITE_ONCE(tp->window_clamp, val);
    ...
}

Because struct udp_sock is much smaller than struct tcp_sock, could this
cast and subsequent WRITE_ONCE() corrupt adjacent slab objects? Should there
be a check to ensure the socket protocol is TCP before calling
__tcp_set_rcvlowat()?

> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923213719.224838-1-kuniyu@google.com?part=7

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list.
  2026-09-23 21:35 ` [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
@ 2026-09-23 22:04   ` Emil Tsalapatis
  2026-09-23 22:31   ` bot+bpf-ci
  2026-09-24 15:53   ` Stanislav Fomichev
  2 siblings, 0 replies; 29+ messages in thread
From: Emil Tsalapatis @ 2026-09-23 22:04 UTC (permalink / raw)
  To: Kuniyuki Iwashima, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman,
	Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, bpf, netdev

On Wed Sep 23, 2026 at 9:35 PM UTC, Kuniyuki Iwashima wrote:
> Currently, four bpf_tcp_ops callbacks are not allowed to call
> bpf_setsockopt() and bpf_getsockopt().
>
> However, the deny-list is fragile, and when a new callback is
> added, we might re-open a can of worms. [1][2]
>
> Let's convert it to allow-list.
>
> Note that it can be written as offsetof(,X) <= moff <= offsetof(,Y)
> but clang can optimise to similar code anyway.

Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>

Even if in the future we need to split the allowlist further it's better
to be verbose than have it break out from under us.

>
> Link: https://lore.kernel.org/r/20220929070407.965581-5-martin.lau@linux.dev/ #[0]
> Link: https://lore.kernel.org/bpf/20260421155804.135786-1-kafai.wan@linux.dev/ #[1]
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> ---
>  net/ipv4/bpf_tcp_ops.c | 39 ++++++++++++++++++++++++---------------
>  1 file changed, 24 insertions(+), 15 deletions(-)
>
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index 681fed642999..1ada3b781bf1 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
> @@ -210,6 +210,24 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
>  	.ret_type	= RET_INTEGER,
>  };
>  
> +static bool is_sockopt_supported(u32 moff)
> +{
> +	switch (moff) {
> +	case offsetof(struct bpf_tcp_ops, active_established):
> +	case offsetof(struct bpf_tcp_ops, passive_established):
> +	case offsetof(struct bpf_tcp_ops, rto):
> +	case offsetof(struct bpf_tcp_ops, rtt):
> +	case offsetof(struct bpf_tcp_ops, set_state):
> +	case offsetof(struct bpf_tcp_ops, retrans):
> +	case offsetof(struct bpf_tcp_ops, connect):
> +	case offsetof(struct bpf_tcp_ops, listen):
> +	case offsetof(struct bpf_tcp_ops, parse_hdr):
> +		return true;
> +	}
> +
> +	return false;
> +}
> +
>  static const struct bpf_func_proto *
>  get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
>  {
> @@ -221,22 +239,13 @@ get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
>  	case BPF_FUNC_sk_storage_delete:
>  		return &bpf_sk_storage_delete_proto;
>  	case BPF_FUNC_setsockopt:
> -		/* The sk may be an unlocked listener (synack path) or NULL
> -		 * fullsock; disable for members that can run unlocked.
> -		 */
> -		if (moff == offsetof(struct bpf_tcp_ops, rwnd_init) ||
> -		    moff == offsetof(struct bpf_tcp_ops, timeout_init) ||
> -		    moff == offsetof(struct bpf_tcp_ops, hdr_opt_len) ||
> -		    moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))
> -			return NULL;
> -		return &bpf_sk_setsockopt_proto;
> +		if (is_sockopt_supported(moff))
> +			return &bpf_sk_setsockopt_proto;
> +		return NULL;
>  	case BPF_FUNC_getsockopt:
> -		if (moff == offsetof(struct bpf_tcp_ops, rwnd_init) ||
> -		    moff == offsetof(struct bpf_tcp_ops, timeout_init) ||
> -		    moff == offsetof(struct bpf_tcp_ops, hdr_opt_len) ||
> -		    moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))
> -			return NULL;
> -		return &bpf_sk_getsockopt_proto;
> +		if (is_sockopt_supported(moff))
> +			return &bpf_sk_getsockopt_proto;
> +		return NULL;
>  	case BPF_FUNC_get_retval:
>  		if (moff == offsetof(struct bpf_tcp_ops, timeout_init) ||
>  		    moff == offsetof(struct bpf_tcp_ops, rwnd_init))


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 2/8] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv.
  2026-09-23 21:35 ` [PATCH v2 bpf-next 2/8] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv Kuniyuki Iwashima
@ 2026-09-23 22:11   ` Emil Tsalapatis
  0 siblings, 0 replies; 29+ messages in thread
From: Emil Tsalapatis @ 2026-09-23 22:11 UTC (permalink / raw)
  To: Kuniyuki Iwashima, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman,
	Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, bpf, netdev

On Wed Sep 23, 2026 at 9:35 PM UTC, Kuniyuki Iwashima wrote:
> Once bpf_sock_ops_cb_flags_set() supports a new flag,
> tcpbpf_user.c fails due to the hard-coded max value, 0x80.
>
> Let's replace 0x80 with BPF_SOCK_OPS_ALL_CB_FLAGS + 1.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> Acked-by: Stanislav Fomichev <sdf@fomichev.me>

Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>

> ---
>  tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c | 3 ++-
>  tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c | 3 ++-
>  2 files changed, 4 insertions(+), 2 deletions(-)
>
> diff --git a/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c b/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
> index 7e8fe1bad03f..e4849d2a2956 100644
> --- a/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
> +++ b/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
> @@ -26,7 +26,8 @@ static void verify_result(struct tcpbpf_globals *result)
>  	ASSERT_EQ(result->bytes_acked, 1002, "bytes_acked");
>  	ASSERT_EQ(result->data_segs_in, 1, "data_segs_in");
>  	ASSERT_EQ(result->data_segs_out, 1, "data_segs_out");
> -	ASSERT_EQ(result->bad_cb_test_rv, 0x80, "bad_cb_test_rv");
> +	ASSERT_EQ(result->bad_cb_test_rv, BPF_SOCK_OPS_ALL_CB_FLAGS + 1,
> +		  "bad_cb_test_rv");
>  	ASSERT_EQ(result->good_cb_test_rv, 0, "good_cb_test_rv");
>  	ASSERT_EQ(result->num_listen, 1, "num_listen");
>  
> diff --git a/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c b/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
> index 6935f32eeb8f..e30cb1fab079 100644
> --- a/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
> +++ b/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
> @@ -92,7 +92,8 @@ int bpf_testcb(struct bpf_sock_ops *skops)
>  		break;
>  	case BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB:
>  		/* Test failure to set largest cb flag (assumes not defined) */
> -		global.bad_cb_test_rv = bpf_sock_ops_cb_flags_set(skops, 0x80);
> +		global.bad_cb_test_rv = bpf_sock_ops_cb_flags_set(skops,
> +								  BPF_SOCK_OPS_ALL_CB_FLAGS + 1);
>  		/* Set callback */
>  		global.good_cb_test_rv = bpf_sock_ops_cb_flags_set(skops,
>  						 BPF_SOCK_OPS_STATE_CB_FLAG);


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 7/8] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
  2026-09-23 22:02   ` sashiko-bot
@ 2026-09-23 22:21     ` Kuniyuki Iwashima
  0 siblings, 0 replies; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-23 22:21 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: bpf

On Wed, Sep 23, 2026 at 3:02 PM <sashiko-bot@kernel.org> wrote:
>
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf: tcp: OOB write in __tcp_set_rcvlowat via unrelated socket
> --
>
> commit 68bce63bbe5802a4eee99c7c0f138096d7c16481
> Author: Kuniyuki Iwashima <kuniyu@google.com>
>
> bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
>
> This commit adds a new BPF kfunc bpf_tcp_ops_set_rcvlowat to allow
> struct_ops programs to dynamically adjust sk->sk_rcvlowat. It clamps
> negative values to INT_MAX and delegates the update to __tcp_set_rcvlowat().
>
> > diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> > index b0cade34cce65..e93c8d13684ef 100644
> > --- a/net/ipv4/bpf_tcp_ops.c
> > +++ b/net/ipv4/bpf_tcp_ops.c
> [ ... ]
> > @@ -368,8 +368,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
> >       .owner = THIS_MODULE,
> >  };
> >
> > +__bpf_kfunc_start_defs();
> > +
> > +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> > +                                      const struct bpf_prog_aux *aux)
> > +{
> > +     u32 moff = aux->attach_st_ops_member_off;
> > +     bool wakeup = false;
> > +
> > +     if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
> > +             wakeup = true;
> > +
> > +     if (rcvlowat < 0)
> > +             rcvlowat = INT_MAX;
> > +
> > +     return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);
>
> [Severity: High]
> Does this code allow an out-of-bounds write via an unrelated socket type?

No, from v1:
https://lore.kernel.org/bpf/CAAVpQUAJJ2RmL_dTH=M2s2Tj9mypL2YtYK7t6tCSaEj4bYY9Uw@mail.gmail.com/

---8<---
No, false-positive.

This kfunc is supported in encode_rcvq() and decode_rcvq()
only, and we don't support lookup helpers there.
---8<---

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list.
  2026-09-23 21:35 ` [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
  2026-09-23 22:04   ` Emil Tsalapatis
@ 2026-09-23 22:31   ` bot+bpf-ci
  2026-09-24 15:53   ` Stanislav Fomichev
  2 siblings, 0 replies; 29+ messages in thread
From: bot+bpf-ci @ 2026-09-23 22:31 UTC (permalink / raw)
  To: kuniyu, ast, daniel, andrii, martin.lau, eddyz87, memxor
  Cc: yonghong.song, john.fastabend, sdf, edumazet, ncardwell, willemb,
	ukyab, cleger, kuniyu, kuni1840, bpf, netdev, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 1174 bytes --]

> bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list.
>
> Currently, four bpf_tcp_ops callbacks are not allowed to call
> bpf_setsockopt() and bpf_getsockopt().
>
> However, the deny-list is fragile, and when a new callback is
> added, we might re-open a can of worms. [1][2]
>
> Let's convert it to allow-list.
>
> Note that it can be written as offsetof(,X) <= moff <= offsetof(,Y)
> but clang can optimise to similar code anyway.
>
> Link: https://lore.kernel.org/r/20220929070407.965581-5-martin.lau@linux.dev/ #[0]
> Link: https://lore.kernel.org/bpf/20260421155804.135786-1-kafai.wan@linux.dev/ #[1]
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>

The changelog body cites references "[1][2]", but the Link: trailers are
labeled "#[0]" and "#[1]". Nothing is labeled [2], and nothing in the
body cites [0].

Should the body say "[0][1]" instead, or should the trailers be
relabeled "#[1]" and "#[2]"?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35924977907

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 4/8] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops.
  2026-09-23 21:35 ` [PATCH v2 bpf-next 4/8] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops Kuniyuki Iwashima
@ 2026-09-23 22:31   ` bot+bpf-ci
  0 siblings, 0 replies; 29+ messages in thread
From: bot+bpf-ci @ 2026-09-23 22:31 UTC (permalink / raw)
  To: kuniyu, ast, daniel, andrii, martin.lau, eddyz87, memxor
  Cc: yonghong.song, john.fastabend, sdf, edumazet, ncardwell, willemb,
	ukyab, cleger, kuniyu, kuni1840, bpf, netdev, ast, andrii, daniel,
	martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 4119 bytes --]

> diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
> index 3f0877b5867bd..68c7e1ca418c9 100644
> --- a/include/uapi/linux/bpf.h
> +++ b/include/uapi/linux/bpf.h
> @@ -3264,7 +3264,7 @@ union bpf_attr {
>   * 	Return
>   * 		0
>   *
> - * long bpf_sock_ops_cb_flags_set(struct bpf_sock_ops *bpf_sock, int argval)
> + * long bpf_sock_ops_cb_flags_set(void *bpf_sock, int argval)
>   * 	Description
>   * 		Attempt to set the value of the **bpf_sock_ops_cb_flags** field
>   * 		for the full TCP socket associated to *bpf_sock_ops* to
>   * 		*argval*.
>   *
>   * 		The primary use of this field is to determine if there should
>   * 		be calls to eBPF programs of type
>   * 		**BPF_PROG_TYPE_SOCK_OPS** at various points in the TCP
>   * 		code. A program of the same type can change its value, per
>   * 		connection and as necessary, when the connection is
>   * 		established. This field is directly accessible for reading, but
>   * 		this helper must be used for updates in order to return an
>   * 		error if an eBPF program tries to set a callback that is not
>   * 		supported in the current kernel.
>   *
>   * 		*argval* is a flag array which can combine these flags:
>   *
>   * 		* **BPF_SOCK_OPS_RTO_CB_FLAG** (retransmission time out)
>   * 		* **BPF_SOCK_OPS_RETRANS_CB_FLAG** (retransmission)
>   * 		* **BPF_SOCK_OPS_STATE_CB_FLAG** (TCP state change)
>   * 		* **BPF_SOCK_OPS_RTT_CB_FLAG** (every RTT)

Should the documentation be updated to reflect the new dual-context use?

The prototype changed to void *bpf_sock so bpf_tcp_ops programs can pass
their struct sock * argument, but the Description still only mentions
*bpf_sock_ops* and BPF_PROG_TYPE_SOCK_OPS.

Compare bpf_setsockopt() in this file, which documents what *bpf_socket*
should be for each program type: "*bpf_socket* should be one of the
following: struct bpf_sock_ops for BPF_PROG_TYPE_SOCK_OPS ..."

Also, which bpf_tcp_ops operators can call this helper? Looking at
get_func_proto() in net/ipv4/bpf_tcp_ops.c, it returns
bpf_tcp_ops_cb_flags_set_proto only for connect, listen,
active_established, passive_established, enqueue_rcvq, and dequeue_rcvq.

Earlier patches in this series documented similar restrictions. For
example, bpf_load_hdr_opt() says "For the **bpf_tcp_ops** struct_ops,
this helper can be called from the **parse_hdr**\() and
**write_hdr_opt**\() operators."

The commit message mentions clearing BPF_SOCK_OPS_RCVQ_CB_FLAG as the
motivation for this change, but that flag doesn't appear in the *argval*
flag list above. The list only shows RTO_CB, RETRANS_CB, STATE_CB, and
RTT_CB.

Finally, the Return section says the helper returns -EINVAL if the socket
is not a full TCP socket:

>   * 	Return
>   * 		Code **-EINVAL** if the socket is not a full TCP socket;
>   * 		otherwise, a positive number containing the bits that could not
>   * 		be set is returned (which comes down to 0 if all bits were set
>   * 		as required).

But bpf_tcp_ops_cb_flags_set() in net/ipv4/bpf_tcp_ops.c just returns
'argval & ~BPF_SOCK_OPS_ALL_CB_FLAGS' without any -EINVAL path. Should
the Return documentation clarify that -EINVAL only applies to
BPF_PROG_TYPE_SOCK_OPS programs?

> [ ... ]
>
> diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
> index 3f0877b5867bd..68c7e1ca418c9 100644
> --- a/tools/include/uapi/linux/bpf.h
> +++ b/tools/include/uapi/linux/bpf.h
> @@ -3264,7 +3264,7 @@ union bpf_attr {
>   * 	Return
>   * 		0
>   *
> - * long bpf_sock_ops_cb_flags_set(struct bpf_sock_ops *bpf_sock, int argval)
> + * long bpf_sock_ops_cb_flags_set(void *bpf_sock, int argval)
>   * 	Description
>   * 		Attempt to set the value of the **bpf_sock_ops_cb_flags** field
>   * 		for the full TCP socket associated to *bpf_sock_ops* to

The same documentation text appears in tools/include/uapi/linux/bpf.h and
would need the same updates.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35924977907

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  2026-09-23 21:35 ` [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
  2026-09-23 21:50   ` sashiko-bot
@ 2026-09-23 22:42   ` Emil Tsalapatis
  2026-09-25  0:04   ` Alexei Starovoitov
  2 siblings, 0 replies; 29+ messages in thread
From: Emil Tsalapatis @ 2026-09-23 22:42 UTC (permalink / raw)
  To: Kuniyuki Iwashima, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman,
	Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, bpf, netdev

On Wed Sep 23, 2026 at 9:35 PM UTC, Kuniyuki Iwashima wrote:
> Waking up a thread per packet is expensive when an application
> processes variable-length frames (e.g., RPC) that span multiple
> packets.
>
> SO_RCVLOWAT can defer wakeups, but because the frame size is
> encoded in a fixed-size descriptor at the start of each frame,
> the application has to:
>
>   1. wake up and recv() the descriptor,
>   2. raise SO_RCVLOWAT to the payload size via setsockopt(),
>   3. wake up and recv() the payload, and
>   4. reset SO_RCVLOWAT back to the descriptor size via
>      setsockopt() for the next frame.
>
> This requires an extra wakeup and two setsockopt() syscalls
> for every single RPC frame.
>
> With SOCKMAP, we can parse skb and suppress wakeups in kernel,
> but SOCKMAP adds overhead and also kills zerocopy.
>
> Let's add lighter-weight opt-in callbacks to bpf_tcp_ops to
> replace that.
>
>   .enqueue_rcvq(): invoked when TCP stack enqueues skb to
>                    sk->sk_receive_queue
>
>   .dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
>                    is dequeued from sk->sk_receive_queue
>
> Those callbacks can be enabled on a per-socket basis by
> bpf_setsockopt():
>
>   int flags = BPF_SOCK_OPS_RCVQ_CB_FLAG;
>
>   bpf_setsockopt(sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
>                  &flags, sizeof(flags));
>
> or via the bpf_tcp_ops-specific helper added in the next patch:
>
>   bpf_sock_ops_cb_flags_set(sk, BPF_SOCK_OPS_RCVQ_CB_FLAG);
>
> Later, we will add a new kfunc to adjust sk->sk_rcvlowat from
> these callbacks.
>
> This will allow the bpf_tcp_ops prog to parse each skb and
> dynamically adjust sk->sk_rcvlowat to suppress unnecessary EPOLLIN
> wakeups until sufficient data is available in the receive queue.
>
> The placement of bpf_tcp_ops_call() in tcp_ofo_queue() and
> tcp_fastopen_add_skb() is chosen to provide the same snapshot
> as tcp_queue_rcv().
>
> For example, if bpf_tcp_ops_call() were called before updating
> TCP_SKB_CB(skb)->seq in tcp_fastopen_add_skb(), BPF prog would
> need an extra branch for the unlikely TFO case to strip SYN.
>
> In addition, the TCP stack can queue overlapping skbs into recvq.
> Once rcv_nxt is updated with a new skb, BPF prog can no longer
> infer the previous rcv_nxt from skb->len.
>
> Lastly, dequeue_rcvq() is placed in tcp_cleanup_rbuf() rather
> than __tcp_cleanup_rbuf() so that it is not called for sockets
> in SOCKMAP, where calling sk->sk_data_ready() from the new
> kfunc would otherwise trigger infinite recursion.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> Acked-by: Stanislav Fomichev <sdf@fomichev.me>

Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>

> ---
>  include/net/tcp.h              | 18 ++++++++++++++++++
>  include/uapi/linux/bpf.h       | 11 ++++++++++-
>  net/ipv4/bpf_tcp_ops.c         | 10 ++++++++++
>  net/ipv4/tcp.c                 |  2 ++
>  net/ipv4/tcp_fastopen.c        |  2 ++
>  net/ipv4/tcp_input.c           |  4 ++++
>  tools/include/uapi/linux/bpf.h | 11 ++++++++++-
>  7 files changed, 56 insertions(+), 2 deletions(-)
>
> diff --git a/include/net/tcp.h b/include/net/tcp.h
> index d61ee00052e3..f2d838bcb0a7 100644
> --- a/include/net/tcp.h
> +++ b/include/net/tcp.h
> @@ -3053,6 +3053,12 @@ struct bpf_tcp_ops {
>  			      struct request_sock *req, struct sk_buff *syn_skb,
>  			      enum tcp_synack_type synack_type,
>  			      u32 opt_off);
> +
> +	/* Called when an incoming skb is enqueued to sk->sk_receive_queue. */
> +	void (*enqueue_rcvq)(struct sock *sk, struct sk_buff *skb);
> +
> +	/* Called after data is dequeued from sk->sk_receive_queue. */
> +	void (*dequeue_rcvq)(struct sock *sk);
>  };
>  
>  #define bpf_tcp_ops_call(op, sk, ...)					\
> @@ -3144,6 +3150,18 @@ static inline void tcp_bpf_rtt(struct sock *sk, long mrtt, u32 srtt)
>  	bpf_tcp_ops_call(rtt, sk, mrtt, srtt);
>  }
>  
> +static inline void bpf_tcp_ops_enqueue_rcvq(struct sock *sk, struct sk_buff *skb)
> +{
> +	if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
> +		bpf_tcp_ops_call(enqueue_rcvq, sk, skb);
> +}
> +
> +static inline void bpf_tcp_ops_dequeue_rcvq(struct sock *sk)
> +{
> +	if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
> +		bpf_tcp_ops_call(dequeue_rcvq, sk);
> +}
> +
>  #if IS_ENABLED(CONFIG_SMC)
>  extern struct static_key_false tcp_have_smc;
>  #endif
> diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
> index 6330b7d745c5..fe122242b096 100644
> --- a/include/uapi/linux/bpf.h
> +++ b/include/uapi/linux/bpf.h
> @@ -7148,8 +7148,17 @@ enum {
>  	 * options first before the BPF program does.
>  	 */
>  	BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
> +	/* Call bpf when the TCP stack enqueues/dequeues payload
> +	 * to/from sk->sk_receive_queue.
> +	 *
> +	 * Only bpf_tcp_ops is supported.
> +	 *
> +	 * It can be used to adjust sk->sk_rcvlowat and suppress
> +	 * unnecessary wakeups before sufficient data is available.
> +	 */
> +	BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
>  /* Mask of all currently supported cb flags */
> -	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0x7F,
> +	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0xFF,
>  };
>  
>  enum {
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index 1ada3b781bf1..c963e2cc21b5 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
> @@ -76,6 +76,14 @@ static void write_hdr_opt_stub(struct sock *sk, struct sk_buff *skb,
>  {
>  }
>  
> +static void enqueue_rcvq_stub(struct sock *sk, struct sk_buff *skb)
> +{
> +}
> +
> +static void dequeue_rcvq_stub(struct sock *sk)
> +{
> +}
> +
>  static struct bpf_tcp_ops __bpf_tcp_ops = {
>  	.timeout_init = timeout_init_stub,
>  	.rwnd_init = rwnd_init_stub,
> @@ -90,6 +98,8 @@ static struct bpf_tcp_ops __bpf_tcp_ops = {
>  	.parse_hdr = parse_hdr_stub,
>  	.hdr_opt_len = hdr_opt_len_stub,
>  	.write_hdr_opt = write_hdr_opt_stub,
> +	.enqueue_rcvq = enqueue_rcvq_stub,
> +	.dequeue_rcvq = dequeue_rcvq_stub,
>  };
>  
>  BPF_CALL_4(bpf_tcp_ops_store_hdr_opt, void *, ctx, const void *, from,
> diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
> index a4456b419412..a714b36a7494 100644
> --- a/net/ipv4/tcp.c
> +++ b/net/ipv4/tcp.c
> @@ -1610,6 +1610,8 @@ void tcp_cleanup_rbuf(struct sock *sk, int copied)
>  	     "cleanup rbuf bug: copied %X seq %X rcvnxt %X\n",
>  	     tp->copied_seq, TCP_SKB_CB(skb)->end_seq, tp->rcv_nxt);
>  	__tcp_cleanup_rbuf(sk, copied);
> +
> +	bpf_tcp_ops_dequeue_rcvq(sk);
>  }
>  
>  static void tcp_eat_recv_skb(struct sock *sk, struct sk_buff *skb)
> diff --git a/net/ipv4/tcp_fastopen.c b/net/ipv4/tcp_fastopen.c
> index 471c78be5513..4939bcbc81d1 100644
> --- a/net/ipv4/tcp_fastopen.c
> +++ b/net/ipv4/tcp_fastopen.c
> @@ -281,6 +281,8 @@ void tcp_fastopen_add_skb(struct sock *sk, struct sk_buff *skb)
>  	TCP_SKB_CB(skb)->seq++;
>  	TCP_SKB_CB(skb)->tcp_flags &= ~TCPHDR_SYN;
>  
> +	bpf_tcp_ops_enqueue_rcvq(sk, skb);
> +
>  	tp->rcv_nxt = TCP_SKB_CB(skb)->end_seq;
>  	tcp_add_receive_queue(sk, skb);
>  	tp->syn_data_acked = 1;
> diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
> index 6ac6f9d5b6c3..c60c61bb0a71 100644
> --- a/net/ipv4/tcp_input.c
> +++ b/net/ipv4/tcp_input.c
> @@ -5344,6 +5344,8 @@ static void tcp_ofo_queue(struct sock *sk)
>  			continue;
>  		}
>  
> +		bpf_tcp_ops_enqueue_rcvq(sk, skb);
> +
>  		tail = skb_peek_tail(&sk->sk_receive_queue);
>  		eaten = tail && tcp_try_coalesce(sk, tail, skb, &fragstolen);
>  		tcp_rcv_nxt_update(tp, TCP_SKB_CB(skb)->end_seq);
> @@ -5547,6 +5549,8 @@ static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb,
>  	int eaten;
>  	struct sk_buff *tail = skb_peek_tail(&sk->sk_receive_queue);
>  
> +	bpf_tcp_ops_enqueue_rcvq(sk, skb);
> +
>  	eaten = (tail &&
>  		 tcp_try_coalesce(sk, tail,
>  				  skb, fragstolen)) ? 1 : 0;
> diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
> index 6330b7d745c5..fe122242b096 100644
> --- a/tools/include/uapi/linux/bpf.h
> +++ b/tools/include/uapi/linux/bpf.h
> @@ -7148,8 +7148,17 @@ enum {
>  	 * options first before the BPF program does.
>  	 */
>  	BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
> +	/* Call bpf when the TCP stack enqueues/dequeues payload
> +	 * to/from sk->sk_receive_queue.
> +	 *
> +	 * Only bpf_tcp_ops is supported.
> +	 *
> +	 * It can be used to adjust sk->sk_rcvlowat and suppress
> +	 * unnecessary wakeups before sufficient data is available.
> +	 */
> +	BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
>  /* Mask of all currently supported cb flags */
> -	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0x7F,
> +	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0xFF,
>  };
>  
>  enum {


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 7/8] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
  2026-09-23 21:35 ` [PATCH v2 bpf-next 7/8] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
  2026-09-23 22:02   ` sashiko-bot
@ 2026-09-24  0:30   ` Emil Tsalapatis
  1 sibling, 0 replies; 29+ messages in thread
From: Emil Tsalapatis @ 2026-09-24  0:30 UTC (permalink / raw)
  To: Kuniyuki Iwashima, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman,
	Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, bpf, netdev

On Wed Sep 23, 2026 at 9:35 PM UTC, Kuniyuki Iwashima wrote:
> bpf_tcp_ops.{enqueue,dequeue}_rcvq() were added to parse skb and
> adjust sk->sk_rcvlowat dynamically to suppress unnecessary wakeups.
>
> Let's add a new kfunc to set sk->sk_rcvlowat.
>
> Negative values are clamped to INT_MAX, consistent with SO_RCVLOWAT.
>
> For enqueue_rcvq(), wakeup is set to false because:
>
>   * tcp_data_ready() is always called after the hooks in
>     tcp_queue_rcv() and tcp_ofo_queue().
>
>   * when tcp_fastopen_add_skb() is called for TFO SYN, the socket is
>     not yet accept()ed, and when called for TFO SYN+ACK, the socket
>     is woken up by sk->sk_state_change() anyway.
>
> For dequeue_rcvq(), wakeup is set to true because tcp_data_ready()
> is not called in that path.
>
> An alternative would be to support bpf_setsockopt() for these
> hooks.
>
> However, that approach involves excessive conditionals and an
> unnecessary memcpy(), costs we do not want to pay for every skb
> in the TCP fast path.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> Acked-by: Stanislav Fomichev <sdf@fomichev.me>
> Tested-by: Clément Léger <cleger@meta.com>

Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>

> ---
>  net/ipv4/bpf_tcp_ops.c | 56 +++++++++++++++++++++++++++++++++++++++++-
>  1 file changed, 55 insertions(+), 1 deletion(-)
>
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index b0cade34cce6..e93c8d13684e 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
> @@ -368,8 +368,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
>  	.owner = THIS_MODULE,
>  };
>  
> +__bpf_kfunc_start_defs();
> +
> +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> +					 const struct bpf_prog_aux *aux)
> +{
> +	u32 moff = aux->attach_st_ops_member_off;
> +	bool wakeup = false;
> +
> +	if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
> +		wakeup = true;
> +
> +	if (rcvlowat < 0)
> +		rcvlowat = INT_MAX;
> +
> +	return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);
> +}
> +
> +__bpf_kfunc_end_defs();
> +
> +BTF_KFUNCS_START(bpf_tcp_ops_rcvlowat_kfunc_set)
> +BTF_ID_FLAGS(func, bpf_tcp_ops_set_rcvlowat, KF_IMPLICIT_ARGS)
> +BTF_KFUNCS_END(bpf_tcp_ops_rcvlowat_kfunc_set)
> +
> +static int bpf_tcp_ops_rcvlowat_kfunc_filter(const struct bpf_prog *prog,
> +					     u32 kfunc_id)
> +{
> +	u32 moff;
> +
> +	if (!btf_id_set8_contains(&bpf_tcp_ops_rcvlowat_kfunc_set, kfunc_id))
> +		return 0;
> +
> +	if (prog->aux->st_ops != &bpf_tcp_ops)
> +		return -EACCES;
> +
> +	moff = prog->aux->attach_st_ops_member_off;
> +	if (moff != offsetof(struct bpf_tcp_ops, enqueue_rcvq) &&
> +	    moff != offsetof(struct bpf_tcp_ops, dequeue_rcvq))
> +		return -EACCES;
> +
> +	return 0;
> +}
> +
> +static const struct btf_kfunc_id_set bpf_tcp_ops_rcvlowat_kfunc_id_set = {
> +	.owner = THIS_MODULE,
> +	.set = &bpf_tcp_ops_rcvlowat_kfunc_set,
> +	.filter = bpf_tcp_ops_rcvlowat_kfunc_filter,
> +};
> +
>  static int __init __bpf_tcp_ops_init(void)
>  {
> -	return register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
> +	int ret;
> +
> +	ret = register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS,
> +					&bpf_tcp_ops_rcvlowat_kfunc_id_set);
> +	ret = ret ?: register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
> +
> +	return ret;
>  }
>  late_initcall(__bpf_tcp_ops_init);


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 5/8] tcp: Split out __tcp_set_rcvlowat().
  2026-09-23 21:35 ` [PATCH v2 bpf-next 5/8] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
@ 2026-09-24  3:39   ` Emil Tsalapatis
  0 siblings, 0 replies; 29+ messages in thread
From: Emil Tsalapatis @ 2026-09-24  3:39 UTC (permalink / raw)
  To: Kuniyuki Iwashima, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman,
	Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, bpf, netdev

On Wed Sep 23, 2026 at 9:35 PM UTC, Kuniyuki Iwashima wrote:
> We will add a kfunc for bpf_tcp_ops.{enqueue,dequeue}_rcvq()
> to adjust sk->sk_rcvlowat.
>
> These hooks are triggered
>
>   * when the TCP stack enqueues an skb to sk->sk_receive_queue
>   * after data is dequeued from sk->sk_receive_queue
>
> In the enqueue path, tcp_data_ready() is always called after
> the hooks in tcp_queue_rcv() and tcp_ofo_queue().
>
> If tcp_set_rcvlowat() were used as is, tcp_data_ready() could
> be called twice for the same skb, which is redundant and also
> confusing.
>
> Let's split out __tcp_set_rcvlowat() and add a flag to control
> wakeup behaviour.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> Acked-by: Stanislav Fomichev <sdf@fomichev.me>

Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>

> ---
>  include/net/tcp.h |  1 +
>  net/ipv4/tcp.c    | 12 +++++++++---
>  2 files changed, 10 insertions(+), 3 deletions(-)
>
> diff --git a/include/net/tcp.h b/include/net/tcp.h
> index f2d838bcb0a7..07426e8641b7 100644
> --- a/include/net/tcp.h
> +++ b/include/net/tcp.h
> @@ -512,6 +512,7 @@ void tcp_set_keepalive(struct sock *sk, int val);
>  void tcp_syn_ack_timeout(const struct request_sock *req);
>  int tcp_recvmsg(struct sock *sk, struct msghdr *msg, size_t len,
>  		int flags);
> +int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup);
>  int tcp_set_rcvlowat(struct sock *sk, int val);
>  void tcp_set_rcvbuf(struct sock *sk, int val);
>  int tcp_set_window_clamp(struct sock *sk, int val);
> diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
> index a714b36a7494..aa7593fc8334 100644
> --- a/net/ipv4/tcp.c
> +++ b/net/ipv4/tcp.c
> @@ -1828,8 +1828,7 @@ int tcp_peek_len(struct socket *sock)
>  	return tcp_inq(sock->sk);
>  }
>  
> -/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */
> -int tcp_set_rcvlowat(struct sock *sk, int val)
> +int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup)
>  {
>  	struct tcp_sock *tp = tcp_sk(sk);
>  	int space, cap;
> @@ -1842,7 +1841,8 @@ int tcp_set_rcvlowat(struct sock *sk, int val)
>  	WRITE_ONCE(sk->sk_rcvlowat, val ? : 1);
>  
>  	/* Check if we need to signal EPOLLIN right now */
> -	tcp_data_ready(sk);
> +	if (wakeup)
> +		tcp_data_ready(sk);
>  
>  	if (sk->sk_userlocks & SOCK_RCVBUF_LOCK)
>  		return 0;
> @@ -1857,6 +1857,12 @@ int tcp_set_rcvlowat(struct sock *sk, int val)
>  	return 0;
>  }
>  
> +/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */
> +int tcp_set_rcvlowat(struct sock *sk, int val)
> +{
> +	return __tcp_set_rcvlowat(sk, val, true);
> +}
> +
>  void tcp_set_rcvbuf(struct sock *sk, int val)
>  {
>  	tcp_set_window_clamp(sk, tcp_win_from_space(sk, val));


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 6/8] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG.
  2026-09-23 21:35 ` [PATCH v2 bpf-next 6/8] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG Kuniyuki Iwashima
@ 2026-09-24  3:48   ` Emil Tsalapatis
  2026-09-24  4:10     ` Kuniyuki Iwashima
  0 siblings, 1 reply; 29+ messages in thread
From: Emil Tsalapatis @ 2026-09-24  3:48 UTC (permalink / raw)
  To: Kuniyuki Iwashima, Alexei Starovoitov, Daniel Borkmann,
	Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman,
	Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, bpf, netdev

On Wed Sep 23, 2026 at 9:35 PM UTC, Kuniyuki Iwashima wrote:
> The next patch exposes a new kfunc calling __tcp_set_rcvlowat()
> to bpf_tcp_ops.
>
> MPTCP has its own sock->ops->set_rcvlowat() / mptcp_set_rcvlowat(),
> so we should not allow calling __tcp_set_rcvlowat() on MPTCP
> subflows.
>
> Let's disable BPF_SOCK_OPS_RCVQ_CB_FLAG for MPTCP for now.
>
> If needed in the future, bpf_tcp_ops_set_rcvlowat() could be
> extended to properly support MPTCP.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> ---
>  include/net/tcp.h      | 15 +++++++++++++++
>  net/core/filter.c      | 10 ++++++----
>  net/ipv4/bpf_tcp_ops.c |  5 ++++-
>  3 files changed, 25 insertions(+), 5 deletions(-)
>
> diff --git a/include/net/tcp.h b/include/net/tcp.h
> index 07426e8641b7..d3cf655da9ec 100644
> --- a/include/net/tcp.h
> +++ b/include/net/tcp.h
> @@ -2932,6 +2932,16 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
>  	return tcp_call_bpf(sk, op, 3, args);
>  }
>  
> +static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
> +{
> +	if (sk_is_mptcp(sk) &&
> +	    (val & BPF_SOCK_OPS_RCVQ_CB_FLAG))
> +		return -EOPNOTSUPP;
> +
> +	tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
> +	return 0;
> +}
> +
>  static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
>  {
>  	tcp_sk(sk)->bpf_sock_ops_cb_flags = 0;
> @@ -2954,6 +2964,11 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
>  	return -EPERM;
>  }
>  
> +static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
> +{
> +	return -EOPNOTSUPP;
> +}
> +
>  static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
>  {
>  }
> diff --git a/net/core/filter.c b/net/core/filter.c
> index 5feb99884682..f29c061bb066 100644
> --- a/net/core/filter.c
> +++ b/net/core/filter.c
> @@ -5588,8 +5588,7 @@ static int bpf_sol_tcp_setsockopt(struct sock *sk, int optname,
>  	case TCP_BPF_SOCK_OPS_CB_FLAGS:
>  		if (val & ~(BPF_SOCK_OPS_ALL_CB_FLAGS))
>  			return -EINVAL;
> -		tp->bpf_sock_ops_cb_flags = val;
> -		break;
> +		return tcp_set_sock_ops_cb_flags(sk, val);
>  	default:
>  		return -EINVAL;
>  	}
> @@ -6178,8 +6177,9 @@ static const struct bpf_func_proto bpf_sock_ops_getsockopt_proto = {
>  BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
>  	   int, argval)
>  {
> -	struct sock *sk = bpf_sock->sk;
>  	int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
> +	struct sock *sk = bpf_sock->sk;
> +	int err;
>  
>  	if (!is_locked_tcp_sock_ops(bpf_sock))
>  		return -EOPNOTSUPP;
> @@ -6187,7 +6187,9 @@ BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
>  	if (!IS_ENABLED(CONFIG_INET) || !sk_fullsock(sk))
>  		return -EINVAL;
>  
> -	tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
> +	err = tcp_set_sock_ops_cb_flags(sk, val);
> +	if (err)
> +		return err;

Afaict the tcp_set_sock_ops_cb_flags is inconsistent with the previous return values,
in that it returns an errno instead of the flags it couldn't set (it also doesn't
set the flags that we can set for MPTCP when BPF_SOCK_OPS_RCVQ_CB_FLAG is set, which
is the convention of its callers). Would changing the error path to

	tp->bpf_sock_ops_cb_flags = val & BPF_SOCK_OPS_RCVQ_CB_FLAG;
	return (val & (~(BPF_SOCK_OPS_ALL_CB_FLAGS ^ BPF_SOCK_OPS_RCVQ_CB_FLAG);

work?


>  
>  	return argval & (~BPF_SOCK_OPS_ALL_CB_FLAGS);
>  }
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index 4b48711d92a2..b0cade34cce6 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
> @@ -223,8 +223,11 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
>  BPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)
>  {
>  	int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
> +	int err;
>  
> -	tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
> +	err = tcp_set_sock_ops_cb_flags(sk, val);
> +	if (err)
> +		return err;
>  
>  	return argval & ~BPF_SOCK_OPS_ALL_CB_FLAGS;
>  }


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 6/8] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG.
  2026-09-24  3:48   ` Emil Tsalapatis
@ 2026-09-24  4:10     ` Kuniyuki Iwashima
  0 siblings, 0 replies; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-24  4:10 UTC (permalink / raw)
  To: Emil Tsalapatis
  Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
	Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, bpf, netdev

On Wed, Sep 23, 2026 at 8:48 PM Emil Tsalapatis <emil@etsalapatis.com> wrote:
>
> On Wed Sep 23, 2026 at 9:35 PM UTC, Kuniyuki Iwashima wrote:
> > The next patch exposes a new kfunc calling __tcp_set_rcvlowat()
> > to bpf_tcp_ops.
> >
> > MPTCP has its own sock->ops->set_rcvlowat() / mptcp_set_rcvlowat(),
> > so we should not allow calling __tcp_set_rcvlowat() on MPTCP
> > subflows.
> >
> > Let's disable BPF_SOCK_OPS_RCVQ_CB_FLAG for MPTCP for now.
> >
> > If needed in the future, bpf_tcp_ops_set_rcvlowat() could be
> > extended to properly support MPTCP.
> >
> > Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> > ---
> >  include/net/tcp.h      | 15 +++++++++++++++
> >  net/core/filter.c      | 10 ++++++----
> >  net/ipv4/bpf_tcp_ops.c |  5 ++++-
> >  3 files changed, 25 insertions(+), 5 deletions(-)
> >
> > diff --git a/include/net/tcp.h b/include/net/tcp.h
> > index 07426e8641b7..d3cf655da9ec 100644
> > --- a/include/net/tcp.h
> > +++ b/include/net/tcp.h
> > @@ -2932,6 +2932,16 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
> >       return tcp_call_bpf(sk, op, 3, args);
> >  }
> >
> > +static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
> > +{
> > +     if (sk_is_mptcp(sk) &&
> > +         (val & BPF_SOCK_OPS_RCVQ_CB_FLAG))
> > +             return -EOPNOTSUPP;
> > +
> > +     tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
> > +     return 0;
> > +}
> > +
> >  static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
> >  {
> >       tcp_sk(sk)->bpf_sock_ops_cb_flags = 0;
> > @@ -2954,6 +2964,11 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
> >       return -EPERM;
> >  }
> >
> > +static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
> > +{
> > +     return -EOPNOTSUPP;
> > +}
> > +
> >  static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
> >  {
> >  }
> > diff --git a/net/core/filter.c b/net/core/filter.c
> > index 5feb99884682..f29c061bb066 100644
> > --- a/net/core/filter.c
> > +++ b/net/core/filter.c
> > @@ -5588,8 +5588,7 @@ static int bpf_sol_tcp_setsockopt(struct sock *sk, int optname,
> >       case TCP_BPF_SOCK_OPS_CB_FLAGS:
> >               if (val & ~(BPF_SOCK_OPS_ALL_CB_FLAGS))
> >                       return -EINVAL;
> > -             tp->bpf_sock_ops_cb_flags = val;
> > -             break;
> > +             return tcp_set_sock_ops_cb_flags(sk, val);
> >       default:
> >               return -EINVAL;
> >       }
> > @@ -6178,8 +6177,9 @@ static const struct bpf_func_proto bpf_sock_ops_getsockopt_proto = {
> >  BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
> >          int, argval)
> >  {
> > -     struct sock *sk = bpf_sock->sk;
> >       int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
> > +     struct sock *sk = bpf_sock->sk;
> > +     int err;
> >
> >       if (!is_locked_tcp_sock_ops(bpf_sock))
> >               return -EOPNOTSUPP;
> > @@ -6187,7 +6187,9 @@ BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
> >       if (!IS_ENABLED(CONFIG_INET) || !sk_fullsock(sk))
> >               return -EINVAL;
> >
> > -     tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
> > +     err = tcp_set_sock_ops_cb_flags(sk, val);
> > +     if (err)
> > +             return err;
>
> Afaict the tcp_set_sock_ops_cb_flags is inconsistent with the previous return values,

tcp_set_sock_ops_cb_flags() already returns error when
something goes wrong, and this is just a new validation.

The intention here is just to ensure, in the slow path, that no one
blindly uses such an unsupported combination and to avoid
unnecessary validation in kfunc called from the fast path.

Note that enabling BPF_SOCK_OPS_ALL_CB_FLAGS in the
program is clearly wrong, or it must accept that it could break
when a new (opt-in) hook is added.

So no one should stumble on the new error in practice.


> in that it returns an errno instead of the flags it couldn't set (it also doesn't
> set the flags that we can set for MPTCP when BPF_SOCK_OPS_RCVQ_CB_FLAG is set, which
> is the convention of its callers). Would changing the error path to
>
>         tp->bpf_sock_ops_cb_flags = val & BPF_SOCK_OPS_RCVQ_CB_FLAG;
>         return (val & (~(BPF_SOCK_OPS_ALL_CB_FLAGS ^ BPF_SOCK_OPS_RCVQ_CB_FLAG);
>
> work?

Regarding the consistency, this is rather inconsistent because
previously any flag in BPF_SOCK_OPS_RCVQ_CB_FLAG cannot
be returned.


>
>
> >
> >       return argval & (~BPF_SOCK_OPS_ALL_CB_FLAGS);
> >  }
> > diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> > index 4b48711d92a2..b0cade34cce6 100644
> > --- a/net/ipv4/bpf_tcp_ops.c
> > +++ b/net/ipv4/bpf_tcp_ops.c
> > @@ -223,8 +223,11 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
> >  BPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)
> >  {
> >       int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
> > +     int err;
> >
> > -     tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
> > +     err = tcp_set_sock_ops_cb_flags(sk, val);
> > +     if (err)
> > +             return err;
> >
> >       return argval & ~BPF_SOCK_OPS_ALL_CB_FLAGS;
> >  }
>

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list.
  2026-09-23 21:35 ` [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
  2026-09-23 22:04   ` Emil Tsalapatis
  2026-09-23 22:31   ` bot+bpf-ci
@ 2026-09-24 15:53   ` Stanislav Fomichev
  2 siblings, 0 replies; 29+ messages in thread
From: Stanislav Fomichev @ 2026-09-24 15:53 UTC (permalink / raw)
  To: Kuniyuki Iwashima
  Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
	Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, bpf, netdev

On 09/23, Kuniyuki Iwashima wrote:
> Currently, four bpf_tcp_ops callbacks are not allowed to call
> bpf_setsockopt() and bpf_getsockopt().
> 
> However, the deny-list is fragile, and when a new callback is
> added, we might re-open a can of worms. [1][2]
> 
> Let's convert it to allow-list.
> 
> Note that it can be written as offsetof(,X) <= moff <= offsetof(,Y)
> but clang can optimise to similar code anyway.
> 
> Link: https://lore.kernel.org/r/20220929070407.965581-5-martin.lau@linux.dev/ #[0]
> Link: https://lore.kernel.org/bpf/20260421155804.135786-1-kafai.wan@linux.dev/ #[1]
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>

Acked-by: Stanislav Fomichev <sdf@fomichev.me>

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  2026-09-23 21:35 ` [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
  2026-09-23 21:50   ` sashiko-bot
  2026-09-23 22:42   ` Emil Tsalapatis
@ 2026-09-25  0:04   ` Alexei Starovoitov
  2026-09-25  0:39     ` Kuniyuki Iwashima
  2 siblings, 1 reply; 29+ messages in thread
From: Alexei Starovoitov @ 2026-09-25  0:04 UTC (permalink / raw)
  To: Kuniyuki Iwashima, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
  Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, bpf, netdev

On Wed, Sep 23, 2026 at 09:35 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
> +	/* Call bpf when the TCP stack enqueues/dequeues payload
> +	 * to/from sk->sk_receive_queue.
> +	 *
> +	 * Only bpf_tcp_ops is supported.
> +	 *
> +	 * It can be used to adjust sk->sk_rcvlowat and suppress
> +	 * unnecessary wakeups before sufficient data is available.
> +	 */
> +	BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
>  /* Mask of all currently supported cb flags */
> -	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0x7F,
> +	BPF_SOCK_OPS_ALL_CB_FLAGS       = 0xFF,

can we drop the flag ?
hdr_opt_len() is called for every tx skb without per-socket opt-in.
The prog in patch 8 already returns when the socket has no
sk_storage.

The flag is one bit per socket shared by all bpf_tcp_ops in the
cgroup hierarchy. When one prog clears it the others stop getting
enqueue_rcvq().

Then sk_is_mptcp() check in the kfunc is enough and there is no need
to touch bpf_sock_ops_cb_flags_set() and TCP_BPF_SOCK_OPS_CB_FLAGS.

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  2026-09-25  0:04   ` Alexei Starovoitov
@ 2026-09-25  0:39     ` Kuniyuki Iwashima
  2026-09-25  1:41       ` Alexei Starovoitov
  0 siblings, 1 reply; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-25  0:39 UTC (permalink / raw)
  To: Alexei Starovoitov
  Cc: Daniel Borkmann, Andrii Nakryiko, Martin KaFai Lau,
	Eduard Zingerman, Kumar Kartikeya Dwivedi, Yonghong Song,
	John Fastabend, Stanislav Fomichev, Eric Dumazet, Neal Cardwell,
	Willem de Bruijn, Tenzin Ukyab, Clément Léger,
	Kuniyuki Iwashima, bpf, netdev

On Thu, Sep 24, 2026 at 5:04 PM Alexei Starovoitov
<alexei.starovoitov@gmail.com> wrote:
>
> On Wed, Sep 23, 2026 at 09:35 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
> > +     /* Call bpf when the TCP stack enqueues/dequeues payload
> > +      * to/from sk->sk_receive_queue.
> > +      *
> > +      * Only bpf_tcp_ops is supported.
> > +      *
> > +      * It can be used to adjust sk->sk_rcvlowat and suppress
> > +      * unnecessary wakeups before sufficient data is available.
> > +      */
> > +     BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
> >  /* Mask of all currently supported cb flags */
> > -     BPF_SOCK_OPS_ALL_CB_FLAGS       = 0x7F,
> > +     BPF_SOCK_OPS_ALL_CB_FLAGS       = 0xFF,
>
> can we drop the flag ?
> hdr_opt_len() is called for every tx skb without per-socket opt-in.

Oh, I assumed bpf_tcp_ops keeps the same behaviour.

We have a bpf prog to turn the hook on (from cgroup_skb/ingress),
only when ingress rate exceeds the allocated capacity, to signal that
via a custom TCP option and turn the hook off from the hook itself.
(and I planned to post another patch for that)

Anyway, is it because calling bpf prog is not very expensive
or per-prog switch is preferable to per-socket flag ?


> The prog in patch 8 already returns when the socket has no
> sk_storage.

Yes, but this is only for bpf verifier.

>
> The flag is one bit per socket shared by all bpf_tcp_ops in the
> cgroup hierarchy. When one prog clears it the others stop getting
> enqueue_rcvq().

Just an idea, would it make sense to add a kfunc to move bpf_tcp_ops
callback to a shadow pointer by allocating sizeof(bpf_tcp_ops) * 2 ?


>
> Then sk_is_mptcp() check in the kfunc is enough and there is no need
> to touch bpf_sock_ops_cb_flags_set() and TCP_BPF_SOCK_OPS_CB_FLAGS.

It also works, but I wanted to move the dead check to the
control path instead of checking for every rx skb / recvmsg().

and now I think we should move bpf_sock_ops_cb_flags to
a better place (maybe tcp_sock_read_txrx) in tcp_sock.

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  2026-09-25  0:39     ` Kuniyuki Iwashima
@ 2026-09-25  1:41       ` Alexei Starovoitov
  2026-09-25 17:33         ` Amery Hung
  0 siblings, 1 reply; 29+ messages in thread
From: Alexei Starovoitov @ 2026-09-25  1:41 UTC (permalink / raw)
  To: Kuniyuki Iwashima
  Cc: Daniel Borkmann, Andrii Nakryiko, Martin KaFai Lau,
	Eduard Zingerman, Kumar Kartikeya Dwivedi, Yonghong Song,
	John Fastabend, Stanislav Fomichev, Eric Dumazet, Neal Cardwell,
	Willem de Bruijn, Tenzin Ukyab, Clément Léger,
	Kuniyuki Iwashima, bpf, netdev, ameryhung

On Thu, Sep 24, 2026 at 05:39 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>> can we drop the flag ?
>> hdr_opt_len() is called for every tx skb without per-socket opt-in.
>
> Oh, I assumed bpf_tcp_ops keeps the same behaviour.

bpf_tcp_ops were moved out of BPF_SOCK_OPS_TEST_FLAG() guards on
purpose. None of the existing members has per-socket opt-in.
hdr_opt_len() runs for every tx skb, rtt() for every rtt sample.
See the log of commit 3bb54768fe3e
("bpf: tcp: Support parse/len/write header option hooks in bpf_tcp_ops"):
"A per-member/per-cgroup gate could be added later if the extra
fast-path work proves measurable."

> We have a bpf prog to turn the hook on (from cgroup_skb/ingress),
> only when ingress rate exceeds the allocated capacity, to signal that
> via a custom TCP option and turn the hook off from the hook itself.
> (and I planned to post another patch for that)

That should work today without another patch and without the flag.
cgroup_skb/ingress prog does
  bpf_sk_storage_get(&map, sk, 0, BPF_SK_STORAGE_GET_F_CREATE);
when the rate is exceeded. hdr_opt_len() and write_hdr_opt() do
  bpf_sk_storage_get(&map, sk, 0, 0);
and return when it's NULL. write_hdr_opt() calls
bpf_sk_storage_delete() after the option went out.
cg_skb_func_proto() has both helpers and get_func_proto() in
bpf_tcp_ops.c allows them in every member.
egress_read_sock_fields() in test_sock_fields.c does the cgroup_skb
part.

> Anyway, is it because calling bpf prog is not very expensive
> or per-prog switch is preferable to per-socket flag ?

The switch is still per socket. It's sk_storage instead of a bit in
tcp_sock, so every prog has its own.

While such new flag is per socket, but it's one for all progs.

And you want to take the last bit of u8 bpf_sock_ops_cb_flags...

As far as the cost. A socket that didn't opt in pays for an indirect
call into the prog and for bpf_sk_storage_get() that finds nothing,
but only in cgroups where bpf_tcp_ops with enqueue_rcvq is attached.
The members that are not set are NULL and bpf_tcp_ops_call() skips
them. That's what hdr_opt_len() costs on tx.

I don't remember what Amery measured, but worth benchmarking
for your case.

>> The prog in patch 8 already returns when the socket has no
>> sk_storage.
>
> Yes, but this is only for bpf verifier.

tcp_init_autolowat_cb() can be the only place that passes
BPF_SK_STORAGE_GET_F_CREATE.
So for a socket that didn't do setsockopt(BPF_TCP_AUTOLOWAT)
other cbs will return NULL on lookup.

tcp_disable_autolowat() will do:
  bpf_sk_storage_delete(&tcp_autolowat_map, sk);
  bpf_tcp_ops_set_rcvlowat(sk, 1);

> Just an idea, would it make sense to add a kfunc to move bpf_tcp_ops
> callback to a shadow pointer by allocating sizeof(bpf_tcp_ops) * 2 ?

Doesn't look right to me.

Amery,
please share your thoughts here.


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  2026-09-25  1:41       ` Alexei Starovoitov
@ 2026-09-25 17:33         ` Amery Hung
  2026-09-27  0:18           ` Kuniyuki Iwashima
  0 siblings, 1 reply; 29+ messages in thread
From: Amery Hung @ 2026-09-25 17:33 UTC (permalink / raw)
  To: Alexei Starovoitov
  Cc: Kuniyuki Iwashima, Daniel Borkmann, Andrii Nakryiko,
	Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
	Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
	Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
	Clément Léger, Kuniyuki Iwashima, bpf, netdev

On Thu, Sep 24, 2026 at 6:41 PM Alexei Starovoitov
<alexei.starovoitov@gmail.com> wrote:
>
> On Thu, Sep 24, 2026 at 05:39 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
> >> can we drop the flag ?
> >> hdr_opt_len() is called for every tx skb without per-socket opt-in.
> >
> > Oh, I assumed bpf_tcp_ops keeps the same behaviour.
>
> bpf_tcp_ops were moved out of BPF_SOCK_OPS_TEST_FLAG() guards on
> purpose. None of the existing members has per-socket opt-in.
> hdr_opt_len() runs for every tx skb, rtt() for every rtt sample.
> See the log of commit 3bb54768fe3e
> ("bpf: tcp: Support parse/len/write header option hooks in bpf_tcp_ops"):
> "A per-member/per-cgroup gate could be added later if the extra
> fast-path work proves measurable."
>
> > We have a bpf prog to turn the hook on (from cgroup_skb/ingress),
> > only when ingress rate exceeds the allocated capacity, to signal that
> > via a custom TCP option and turn the hook off from the hook itself.
> > (and I planned to post another patch for that)
>
> That should work today without another patch and without the flag.
> cgroup_skb/ingress prog does
>   bpf_sk_storage_get(&map, sk, 0, BPF_SK_STORAGE_GET_F_CREATE);
> when the rate is exceeded. hdr_opt_len() and write_hdr_opt() do
>   bpf_sk_storage_get(&map, sk, 0, 0);
> and return when it's NULL. write_hdr_opt() calls
> bpf_sk_storage_delete() after the option went out.
> cg_skb_func_proto() has both helpers and get_func_proto() in
> bpf_tcp_ops.c allows them in every member.
> egress_read_sock_fields() in test_sock_fields.c does the cgroup_skb
> part.
>
> > Anyway, is it because calling bpf prog is not very expensive
> > or per-prog switch is preferable to per-socket flag ?
>
> The switch is still per socket. It's sk_storage instead of a bit in
> tcp_sock, so every prog has its own.
>
> While such new flag is per socket, but it's one for all progs.
>
> And you want to take the last bit of u8 bpf_sock_ops_cb_flags...
>
> As far as the cost. A socket that didn't opt in pays for an indirect
> call into the prog and for bpf_sk_storage_get() that finds nothing,
> but only in cgroups where bpf_tcp_ops with enqueue_rcvq is attached.
> The members that are not set are NULL and bpf_tcp_ops_call() skips
> them. That's what hdr_opt_len() costs on tx.
>
> I don't remember what Amery measured, but worth benchmarking
> for your case.

I did not benchmark the initial bpf_tcp_ops patch set, so per-socket
gating was deferred. Before deciding whether to add such a gate for
AutoLOWAT, could we compare:

1. No bpf_tcp_ops attached.
2. No-op enqueue_rcvq() and dequeue_rcvq() callbacks attached.
3. The same callbacks performing an sk-local-storage lookup (or maybe
rhashtable).

>
> >> The prog in patch 8 already returns when the socket has no
> >> sk_storage.
> >
> > Yes, but this is only for bpf verifier.
>
> tcp_init_autolowat_cb() can be the only place that passes
> BPF_SK_STORAGE_GET_F_CREATE.
> So for a socket that didn't do setsockopt(BPF_TCP_AUTOLOWAT)
> other cbs will return NULL on lookup.
>
> tcp_disable_autolowat() will do:
>   bpf_sk_storage_delete(&tcp_autolowat_map, sk);
>   bpf_tcp_ops_set_rcvlowat(sk, 1);
>
> > Just an idea, would it make sense to add a kfunc to move bpf_tcp_ops
> > callback to a shadow pointer by allocating sizeof(bpf_tcp_ops) * 2 ?

A bpf_tcp_ops instance is shared by all sockets, so this would not
provide per-socket gating. Making it per-socket would require copying
the callback table for every tcp_sock, which seems more complex than
adding a gate directly to tcp_sock.

Am I understanding your idea correctly?

>
> Doesn't look right to me.
>
> Amery,
> please share your thoughts here.
>

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  2026-09-25 17:33         ` Amery Hung
@ 2026-09-27  0:18           ` Kuniyuki Iwashima
  2026-09-30 19:56             ` Amery Hung
  0 siblings, 1 reply; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-27  0:18 UTC (permalink / raw)
  To: ameryhung
  Cc: alexei.starovoitov, andrii, bpf, cleger, daniel, eddyz87,
	edumazet, john.fastabend, kuni1840, kuniyu, martin.lau, memxor,
	ncardwell, netdev, sdf, ukyab, willemb, yonghong.song

From: Amery Hung <ameryhung@gmail.com>
Date: Fri, 25 Sep 2026 10:33:50 -0700
> On Thu, Sep 24, 2026 at 6:41 PM Alexei Starovoitov
> <alexei.starovoitov@gmail.com> wrote:
> >
> > On Thu, Sep 24, 2026 at 05:39 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
> > >> can we drop the flag ?
> > >> hdr_opt_len() is called for every tx skb without per-socket opt-in.
> > >
> > > Oh, I assumed bpf_tcp_ops keeps the same behaviour.
> >
> > bpf_tcp_ops were moved out of BPF_SOCK_OPS_TEST_FLAG() guards on
> > purpose. None of the existing members has per-socket opt-in.
> > hdr_opt_len() runs for every tx skb, rtt() for every rtt sample.
> > See the log of commit 3bb54768fe3e
> > ("bpf: tcp: Support parse/len/write header option hooks in bpf_tcp_ops"):
> > "A per-member/per-cgroup gate could be added later if the extra
> > fast-path work proves measurable."
> >
> > > We have a bpf prog to turn the hook on (from cgroup_skb/ingress),
> > > only when ingress rate exceeds the allocated capacity, to signal that
> > > via a custom TCP option and turn the hook off from the hook itself.
> > > (and I planned to post another patch for that)
> >
> > That should work today without another patch and without the flag.
> > cgroup_skb/ingress prog does
> >   bpf_sk_storage_get(&map, sk, 0, BPF_SK_STORAGE_GET_F_CREATE);
> > when the rate is exceeded. hdr_opt_len() and write_hdr_opt() do
> >   bpf_sk_storage_get(&map, sk, 0, 0);
> > and return when it's NULL. write_hdr_opt() calls
> > bpf_sk_storage_delete() after the option went out.
> > cg_skb_func_proto() has both helpers and get_func_proto() in
> > bpf_tcp_ops.c allows them in every member.
> > egress_read_sock_fields() in test_sock_fields.c does the cgroup_skb
> > part.
> >
> > > Anyway, is it because calling bpf prog is not very expensive
> > > or per-prog switch is preferable to per-socket flag ?
> >
> > The switch is still per socket. It's sk_storage instead of a bit in
> > tcp_sock, so every prog has its own.
> >
> > While such new flag is per socket, but it's one for all progs.
> >
> > And you want to take the last bit of u8 bpf_sock_ops_cb_flags...
> >
> > As far as the cost. A socket that didn't opt in pays for an indirect
> > call into the prog and for bpf_sk_storage_get() that finds nothing,
> > but only in cgroups where bpf_tcp_ops with enqueue_rcvq is attached.
> > The members that are not set are NULL and bpf_tcp_ops_call() skips
> > them. That's what hdr_opt_len() costs on tx.
> >
> > I don't remember what Amery measured, but worth benchmarking
> > for your case.
> 
> I did not benchmark the initial bpf_tcp_ops patch set, so per-socket
> gating was deferred. Before deciding whether to add such a gate for
> AutoLOWAT, could we compare:
> 
> 1. No bpf_tcp_ops attached.
> 2. No-op enqueue_rcvq() and dequeue_rcvq() callbacks attached.
> 3. The same callbacks performing an sk-local-storage lookup (or maybe
> rhashtable).

I used tcp_rr (128 threads, 20000 flows) and didn't see a big
perf diff between 1. and 2., but 3. showed some costs even when
only looking up sk_storage and returning immediately:

                    |    w/o bpf    |     w/ bpf    |  diff   |  diff (abs)
 -------------------+---------------+---------------+---------+-------------
 num_transactions   | 1,209,349,692 | 1,185,180,017 |  -2.00% | -24.17M tx
 throughput (tx/s)  | 20,318,689.83 | 19,953,668.49 |  -1.80% | -365.0K tx/s
 remote_throughput  |    19,472,629 |    19,135,463 |  -1.73% | -337.2K tx/s
 latency_min        |      50.83 us |      46.08 us |  -9.34% | -4.75 us
 latency_p25        |     911.35 us |     911.35 us |   0.00% | 0.00 us
 latency_p50        |     972.79 us |     983.03 us |  +1.05% | +10.24 us
 latency_mean       |     983.64 us |   1,001.71 us |  +1.84% | +18.08 us
 latency_p90        |   1,136.63 us |   1,187.83 us |  +4.50% | +51.20 us
 latency_p95        |   1,208.31 us |   1,259.51 us |  +4.24% | +51.20 us
 latency_p99        |   1,392.63 us |   1,454.07 us |  +4.41% | +61.44 us
 latency_stddev     |     141.82 us |     158.53 us | +11.78% | +16.71 us
 cpu time (u+s, 60s)|     8,165.91s |     8,312.57s |  +1.80% | +146.66s


Given the other hooks in the fast path (hdr_opt_len(),
parse_hdr(), etc) should add costs at the same level,
I think we should guard them with flags.

But u8 is too small, so I think we could add a new field
(maybe u32?) and a new kfunc for bpf_tcp_ops only.  Also,
I'd add CONFIG_BPF_SOCK_OPS_LEGACY to save some calls.

What do you think ?


> 
> >
> > >> The prog in patch 8 already returns when the socket has no
> > >> sk_storage.
> > >
> > > Yes, but this is only for bpf verifier.
> >
> > tcp_init_autolowat_cb() can be the only place that passes
> > BPF_SK_STORAGE_GET_F_CREATE.
> > So for a socket that didn't do setsockopt(BPF_TCP_AUTOLOWAT)
> > other cbs will return NULL on lookup.
> >
> > tcp_disable_autolowat() will do:
> >   bpf_sk_storage_delete(&tcp_autolowat_map, sk);
> >   bpf_tcp_ops_set_rcvlowat(sk, 1);
> >
> > > Just an idea, would it make sense to add a kfunc to move bpf_tcp_ops
> > > callback to a shadow pointer by allocating sizeof(bpf_tcp_ops) * 2 ?
> 
> A bpf_tcp_ops instance is shared by all sockets, so this would not
> provide per-socket gating. Making it per-socket would require copying
> the callback table for every tcp_sock, which seems more complex than
> adding a gate directly to tcp_sock.
> 
> Am I understanding your idea correctly?

Yes, but this is overkill indeed.  I think most deployments
use just a single SOCK_OPS/bpf_tcp_ops, and per-scoket flag
would be enough.

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  2026-09-27  0:18           ` Kuniyuki Iwashima
@ 2026-09-30 19:56             ` Amery Hung
  2026-09-30 21:03               ` Kuniyuki Iwashima
  0 siblings, 1 reply; 29+ messages in thread
From: Amery Hung @ 2026-09-30 19:56 UTC (permalink / raw)
  To: Kuniyuki Iwashima
  Cc: alexei.starovoitov, andrii, bpf, cleger, daniel, eddyz87,
	edumazet, john.fastabend, kuni1840, martin.lau, memxor, ncardwell,
	netdev, sdf, ukyab, willemb, yonghong.song

On Sat, Sep 26, 2026 at 5:19 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>
> From: Amery Hung <ameryhung@gmail.com>
> Date: Fri, 25 Sep 2026 10:33:50 -0700
> > On Thu, Sep 24, 2026 at 6:41 PM Alexei Starovoitov
> > <alexei.starovoitov@gmail.com> wrote:
> > >
> > > On Thu, Sep 24, 2026 at 05:39 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
> > > >> can we drop the flag ?
> > > >> hdr_opt_len() is called for every tx skb without per-socket opt-in.
> > > >
> > > > Oh, I assumed bpf_tcp_ops keeps the same behaviour.
> > >
> > > bpf_tcp_ops were moved out of BPF_SOCK_OPS_TEST_FLAG() guards on
> > > purpose. None of the existing members has per-socket opt-in.
> > > hdr_opt_len() runs for every tx skb, rtt() for every rtt sample.
> > > See the log of commit 3bb54768fe3e
> > > ("bpf: tcp: Support parse/len/write header option hooks in bpf_tcp_ops"):
> > > "A per-member/per-cgroup gate could be added later if the extra
> > > fast-path work proves measurable."
> > >
> > > > We have a bpf prog to turn the hook on (from cgroup_skb/ingress),
> > > > only when ingress rate exceeds the allocated capacity, to signal that
> > > > via a custom TCP option and turn the hook off from the hook itself.
> > > > (and I planned to post another patch for that)
> > >
> > > That should work today without another patch and without the flag.
> > > cgroup_skb/ingress prog does
> > >   bpf_sk_storage_get(&map, sk, 0, BPF_SK_STORAGE_GET_F_CREATE);
> > > when the rate is exceeded. hdr_opt_len() and write_hdr_opt() do
> > >   bpf_sk_storage_get(&map, sk, 0, 0);
> > > and return when it's NULL. write_hdr_opt() calls
> > > bpf_sk_storage_delete() after the option went out.
> > > cg_skb_func_proto() has both helpers and get_func_proto() in
> > > bpf_tcp_ops.c allows them in every member.
> > > egress_read_sock_fields() in test_sock_fields.c does the cgroup_skb
> > > part.
> > >
> > > > Anyway, is it because calling bpf prog is not very expensive
> > > > or per-prog switch is preferable to per-socket flag ?
> > >
> > > The switch is still per socket. It's sk_storage instead of a bit in
> > > tcp_sock, so every prog has its own.
> > >
> > > While such new flag is per socket, but it's one for all progs.
> > >
> > > And you want to take the last bit of u8 bpf_sock_ops_cb_flags...
> > >
> > > As far as the cost. A socket that didn't opt in pays for an indirect
> > > call into the prog and for bpf_sk_storage_get() that finds nothing,
> > > but only in cgroups where bpf_tcp_ops with enqueue_rcvq is attached.
> > > The members that are not set are NULL and bpf_tcp_ops_call() skips
> > > them. That's what hdr_opt_len() costs on tx.
> > >
> > > I don't remember what Amery measured, but worth benchmarking
> > > for your case.
> >
> > I did not benchmark the initial bpf_tcp_ops patch set, so per-socket
> > gating was deferred. Before deciding whether to add such a gate for
> > AutoLOWAT, could we compare:
> >
> > 1. No bpf_tcp_ops attached.
> > 2. No-op enqueue_rcvq() and dequeue_rcvq() callbacks attached.
> > 3. The same callbacks performing an sk-local-storage lookup (or maybe
> > rhashtable).
>
> I used tcp_rr (128 threads, 20000 flows) and didn't see a big
> perf diff between 1. and 2., but 3. showed some costs even when
> only looking up sk_storage and returning immediately:

Hey, sorry for the late reply.

Your result gave me some hope that we could implement the gate in the
BPF program, since callback dispatch appeared cheap.

However, my experiment with several program-local gate implementations
produced a different result. I ran the experiment using hdr_opt_len().
Except for the baseline, every case else had bpf_tcp_ops attached, and
only the hdr_opt_len member changed. I used loopback tcp_rr with 128
threads and 20,000 flows. All lookups missed, modeling sockets that
had not opted in.

|.                     | Throughput (tx/s) | diff   | diff (abs)     |
|-----------------------|-------------------|--------|----------------|
| no tcp_ops            | 2,231,461         | --     | --             |
| attached, NULL member | 2,200,750         | -1.38% | -30,711 tx/s   |
| no-op                 | 2,172,596         | -2.64% | -58,865 tx/s   |
| sk-storage            | 2,159,986         | -3.20% | -71,475 tx/s   |
| RHASH                  | 2,156,491         | -3.36% | -74,970 tx/s   |

The NULL-member result measures the cgroup struct_ops dispatch and
array scan. The difference from NULL member to no-op measures the
indirect call and trampoline, while the remaining difference measures
the program-local lookup.

Although hdr_opt_len() has a different signature and call frequency
from the RCVQ hooks, all three layers have observable fast-path
overhead. A program-local gate cannot avoid the dispatch and
trampoline costs. Therefore, a per-socket gate before dispatch seems
worthwhile for conditional hot-path callbacks.

>
>                     |    w/o bpf    |     w/ bpf    |  diff   |  diff (abs)
>  -------------------+---------------+---------------+---------+-------------
>  num_transactions   | 1,209,349,692 | 1,185,180,017 |  -2.00% | -24.17M tx
>  throughput (tx/s)  | 20,318,689.83 | 19,953,668.49 |  -1.80% | -365.0K tx/s
>  remote_throughput  |    19,472,629 |    19,135,463 |  -1.73% | -337.2K tx/s
>  latency_min        |      50.83 us |      46.08 us |  -9.34% | -4.75 us
>  latency_p25        |     911.35 us |     911.35 us |   0.00% | 0.00 us
>  latency_p50        |     972.79 us |     983.03 us |  +1.05% | +10.24 us
>  latency_mean       |     983.64 us |   1,001.71 us |  +1.84% | +18.08 us
>  latency_p90        |   1,136.63 us |   1,187.83 us |  +4.50% | +51.20 us
>  latency_p95        |   1,208.31 us |   1,259.51 us |  +4.24% | +51.20 us
>  latency_p99        |   1,392.63 us |   1,454.07 us |  +4.41% | +61.44 us
>  latency_stddev     |     141.82 us |     158.53 us | +11.78% | +16.71 us
>  cpu time (u+s, 60s)|     8,165.91s |     8,312.57s |  +1.80% | +146.66s
>
>
> Given the other hooks in the fast path (hdr_opt_len(),
> parse_hdr(), etc) should add costs at the same level,
> I think we should guard them with flags.
>
> But u8 is too small, so I think we could add a new field
> (maybe u32?) and a new kfunc for bpf_tcp_ops only.  Also,
> I'd add CONFIG_BPF_SOCK_OPS_LEGACY to save some calls.
>
> What do you think ?

A separate bpf_tcp_ops mask in tcp_sock seems reasonable.

CONFIG_BPF_SOCK_OPS_LEGACY seems premature considering we don't have
feature parity with sockops at this moment (not saying we need to
fully duplicate though).

>
>
> >
> > >
> > > >> The prog in patch 8 already returns when the socket has no
> > > >> sk_storage.
> > > >
> > > > Yes, but this is only for bpf verifier.
> > >
> > > tcp_init_autolowat_cb() can be the only place that passes
> > > BPF_SK_STORAGE_GET_F_CREATE.
> > > So for a socket that didn't do setsockopt(BPF_TCP_AUTOLOWAT)
> > > other cbs will return NULL on lookup.
> > >
> > > tcp_disable_autolowat() will do:
> > >   bpf_sk_storage_delete(&tcp_autolowat_map, sk);
> > >   bpf_tcp_ops_set_rcvlowat(sk, 1);
> > >
> > > > Just an idea, would it make sense to add a kfunc to move bpf_tcp_ops
> > > > callback to a shadow pointer by allocating sizeof(bpf_tcp_ops) * 2 ?
> >
> > A bpf_tcp_ops instance is shared by all sockets, so this would not
> > provide per-socket gating. Making it per-socket would require copying
> > the callback table for every tcp_sock, which seems more complex than
> > adding a gate directly to tcp_sock.
> >
> > Am I understanding your idea correctly?
>
> Yes, but this is overkill indeed.  I think most deployments
> use just a single SOCK_OPS/bpf_tcp_ops, and per-scoket flag
> would be enough.

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
  2026-09-30 19:56             ` Amery Hung
@ 2026-09-30 21:03               ` Kuniyuki Iwashima
  0 siblings, 0 replies; 29+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-30 21:03 UTC (permalink / raw)
  To: Amery Hung
  Cc: alexei.starovoitov, andrii, bpf, cleger, daniel, eddyz87,
	edumazet, john.fastabend, kuni1840, martin.lau, memxor, ncardwell,
	netdev, sdf, ukyab, willemb, yonghong.song

On Wed, Sep 30, 2026 at 12:56 PM Amery Hung <ameryhung@gmail.com> wrote:
>
> On Sat, Sep 26, 2026 at 5:19 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
> >
> > From: Amery Hung <ameryhung@gmail.com>
> > Date: Fri, 25 Sep 2026 10:33:50 -0700
> > > On Thu, Sep 24, 2026 at 6:41 PM Alexei Starovoitov
> > > <alexei.starovoitov@gmail.com> wrote:
> > > >
> > > > On Thu, Sep 24, 2026 at 05:39 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
> > > > >> can we drop the flag ?
> > > > >> hdr_opt_len() is called for every tx skb without per-socket opt-in.
> > > > >
> > > > > Oh, I assumed bpf_tcp_ops keeps the same behaviour.
> > > >
> > > > bpf_tcp_ops were moved out of BPF_SOCK_OPS_TEST_FLAG() guards on
> > > > purpose. None of the existing members has per-socket opt-in.
> > > > hdr_opt_len() runs for every tx skb, rtt() for every rtt sample.
> > > > See the log of commit 3bb54768fe3e
> > > > ("bpf: tcp: Support parse/len/write header option hooks in bpf_tcp_ops"):
> > > > "A per-member/per-cgroup gate could be added later if the extra
> > > > fast-path work proves measurable."
> > > >
> > > > > We have a bpf prog to turn the hook on (from cgroup_skb/ingress),
> > > > > only when ingress rate exceeds the allocated capacity, to signal that
> > > > > via a custom TCP option and turn the hook off from the hook itself.
> > > > > (and I planned to post another patch for that)
> > > >
> > > > That should work today without another patch and without the flag.
> > > > cgroup_skb/ingress prog does
> > > >   bpf_sk_storage_get(&map, sk, 0, BPF_SK_STORAGE_GET_F_CREATE);
> > > > when the rate is exceeded. hdr_opt_len() and write_hdr_opt() do
> > > >   bpf_sk_storage_get(&map, sk, 0, 0);
> > > > and return when it's NULL. write_hdr_opt() calls
> > > > bpf_sk_storage_delete() after the option went out.
> > > > cg_skb_func_proto() has both helpers and get_func_proto() in
> > > > bpf_tcp_ops.c allows them in every member.
> > > > egress_read_sock_fields() in test_sock_fields.c does the cgroup_skb
> > > > part.
> > > >
> > > > > Anyway, is it because calling bpf prog is not very expensive
> > > > > or per-prog switch is preferable to per-socket flag ?
> > > >
> > > > The switch is still per socket. It's sk_storage instead of a bit in
> > > > tcp_sock, so every prog has its own.
> > > >
> > > > While such new flag is per socket, but it's one for all progs.
> > > >
> > > > And you want to take the last bit of u8 bpf_sock_ops_cb_flags...
> > > >
> > > > As far as the cost. A socket that didn't opt in pays for an indirect
> > > > call into the prog and for bpf_sk_storage_get() that finds nothing,
> > > > but only in cgroups where bpf_tcp_ops with enqueue_rcvq is attached.
> > > > The members that are not set are NULL and bpf_tcp_ops_call() skips
> > > > them. That's what hdr_opt_len() costs on tx.
> > > >
> > > > I don't remember what Amery measured, but worth benchmarking
> > > > for your case.
> > >
> > > I did not benchmark the initial bpf_tcp_ops patch set, so per-socket
> > > gating was deferred. Before deciding whether to add such a gate for
> > > AutoLOWAT, could we compare:
> > >
> > > 1. No bpf_tcp_ops attached.
> > > 2. No-op enqueue_rcvq() and dequeue_rcvq() callbacks attached.
> > > 3. The same callbacks performing an sk-local-storage lookup (or maybe
> > > rhashtable).
> >
> > I used tcp_rr (128 threads, 20000 flows) and didn't see a big
> > perf diff between 1. and 2., but 3. showed some costs even when
> > only looking up sk_storage and returning immediately:
>
> Hey, sorry for the late reply.
>
> Your result gave me some hope that we could implement the gate in the
> BPF program, since callback dispatch appeared cheap.
>
> However, my experiment with several program-local gate implementations
> produced a different result. I ran the experiment using hdr_opt_len().
> Except for the baseline, every case else had bpf_tcp_ops attached, and
> only the hdr_opt_len member changed. I used loopback tcp_rr with 128

Thanks for testing.  To be fair, I used two machines on the same rack,
and it seems loopback has bottleneck on CPU rather than network,
less relaxed.


> threads and 20,000 flows. All lookups missed, modeling sockets that
> had not opted in.
>
> |.                     | Throughput (tx/s) | diff   | diff (abs)     |
> |-----------------------|-------------------|--------|----------------|
> | no tcp_ops            | 2,231,461         | --     | --             |
> | attached, NULL member | 2,200,750         | -1.38% | -30,711 tx/s   |
> | no-op                 | 2,172,596         | -2.64% | -58,865 tx/s   |
> | sk-storage            | 2,159,986         | -3.20% | -71,475 tx/s   |
> | RHASH                  | 2,156,491         | -3.36% | -74,970 tx/s   |
>
> The NULL-member result measures the cgroup struct_ops dispatch and
> array scan. The difference from NULL member to no-op measures the
> indirect call and trampoline, while the remaining difference measures
> the program-local lookup.
>
> Although hdr_opt_len() has a different signature and call frequency
> from the RCVQ hooks, all three layers have observable fast-path
> overhead. A program-local gate cannot avoid the dispatch and
> trampoline costs. Therefore, a per-socket gate before dispatch seems
> worthwhile for conditional hot-path callbacks.
>
> >
> >                     |    w/o bpf    |     w/ bpf    |  diff   |  diff (abs)
> >  -------------------+---------------+---------------+---------+-------------
> >  num_transactions   | 1,209,349,692 | 1,185,180,017 |  -2.00% | -24.17M tx
> >  throughput (tx/s)  | 20,318,689.83 | 19,953,668.49 |  -1.80% | -365.0K tx/s
> >  remote_throughput  |    19,472,629 |    19,135,463 |  -1.73% | -337.2K tx/s
> >  latency_min        |      50.83 us |      46.08 us |  -9.34% | -4.75 us
> >  latency_p25        |     911.35 us |     911.35 us |   0.00% | 0.00 us
> >  latency_p50        |     972.79 us |     983.03 us |  +1.05% | +10.24 us
> >  latency_mean       |     983.64 us |   1,001.71 us |  +1.84% | +18.08 us
> >  latency_p90        |   1,136.63 us |   1,187.83 us |  +4.50% | +51.20 us
> >  latency_p95        |   1,208.31 us |   1,259.51 us |  +4.24% | +51.20 us
> >  latency_p99        |   1,392.63 us |   1,454.07 us |  +4.41% | +61.44 us
> >  latency_stddev     |     141.82 us |     158.53 us | +11.78% | +16.71 us
> >  cpu time (u+s, 60s)|     8,165.91s |     8,312.57s |  +1.80% | +146.66s
> >
> >
> > Given the other hooks in the fast path (hdr_opt_len(),
> > parse_hdr(), etc) should add costs at the same level,
> > I think we should guard them with flags.
> >
> > But u8 is too small, so I think we could add a new field
> > (maybe u32?) and a new kfunc for bpf_tcp_ops only.  Also,
> > I'd add CONFIG_BPF_SOCK_OPS_LEGACY to save some calls.
> >
> > What do you think ?
>
> A separate bpf_tcp_ops mask in tcp_sock seems reasonable.

I will add another flag and kfunc in the next version.


>
> CONFIG_BPF_SOCK_OPS_LEGACY seems premature considering we don't have
> feature parity with sockops at this moment (not saying we need to
> fully duplicate though).

This can be a separate topic, and I also think we don't need
feature parity.  I think some hooks are not used in many
deployments.


>
> >
> >
> > >
> > > >
> > > > >> The prog in patch 8 already returns when the socket has no
> > > > >> sk_storage.
> > > > >
> > > > > Yes, but this is only for bpf verifier.
> > > >
> > > > tcp_init_autolowat_cb() can be the only place that passes
> > > > BPF_SK_STORAGE_GET_F_CREATE.
> > > > So for a socket that didn't do setsockopt(BPF_TCP_AUTOLOWAT)
> > > > other cbs will return NULL on lookup.
> > > >
> > > > tcp_disable_autolowat() will do:
> > > >   bpf_sk_storage_delete(&tcp_autolowat_map, sk);
> > > >   bpf_tcp_ops_set_rcvlowat(sk, 1);
> > > >
> > > > > Just an idea, would it make sense to add a kfunc to move bpf_tcp_ops
> > > > > callback to a shadow pointer by allocating sizeof(bpf_tcp_ops) * 2 ?
> > >
> > > A bpf_tcp_ops instance is shared by all sockets, so this would not
> > > provide per-socket gating. Making it per-socket would require copying
> > > the callback table for every tcp_sock, which seems more complex than
> > > adding a gate directly to tcp_sock.
> > >
> > > Am I understanding your idea correctly?
> >
> > Yes, but this is overkill indeed.  I think most deployments
> > use just a single SOCK_OPS/bpf_tcp_ops, and per-scoket flag
> > would be enough.

^ permalink raw reply	[flat|nested] 29+ messages in thread

end of thread, other threads:[~2026-09-30 21:03 UTC | newest]

Thread overview: 29+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-23 21:35 [PATCH v2 bpf-next 0/8] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
2026-09-23 21:35 ` [PATCH v2 bpf-next 1/8] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
2026-09-23 22:04   ` Emil Tsalapatis
2026-09-23 22:31   ` bot+bpf-ci
2026-09-24 15:53   ` Stanislav Fomichev
2026-09-23 21:35 ` [PATCH v2 bpf-next 2/8] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv Kuniyuki Iwashima
2026-09-23 22:11   ` Emil Tsalapatis
2026-09-23 21:35 ` [PATCH v2 bpf-next 3/8] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-09-23 21:50   ` sashiko-bot
2026-09-23 22:42   ` Emil Tsalapatis
2026-09-25  0:04   ` Alexei Starovoitov
2026-09-25  0:39     ` Kuniyuki Iwashima
2026-09-25  1:41       ` Alexei Starovoitov
2026-09-25 17:33         ` Amery Hung
2026-09-27  0:18           ` Kuniyuki Iwashima
2026-09-30 19:56             ` Amery Hung
2026-09-30 21:03               ` Kuniyuki Iwashima
2026-09-23 21:35 ` [PATCH v2 bpf-next 4/8] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops Kuniyuki Iwashima
2026-09-23 22:31   ` bot+bpf-ci
2026-09-23 21:35 ` [PATCH v2 bpf-next 5/8] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
2026-09-24  3:39   ` Emil Tsalapatis
2026-09-23 21:35 ` [PATCH v2 bpf-next 6/8] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG Kuniyuki Iwashima
2026-09-24  3:48   ` Emil Tsalapatis
2026-09-24  4:10     ` Kuniyuki Iwashima
2026-09-23 21:35 ` [PATCH v2 bpf-next 7/8] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
2026-09-23 22:02   ` sashiko-bot
2026-09-23 22:21     ` Kuniyuki Iwashima
2026-09-24  0:30   ` Emil Tsalapatis
2026-09-23 21:35 ` [PATCH v2 bpf-next 8/8] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox