* [PATCH bpf-next 0/7] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT.
@ 2026-09-20 19:56 Kuniyuki Iwashima
2026-09-20 19:56 ` [PATCH bpf-next 1/7] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv Kuniyuki Iwashima
` (6 more replies)
0 siblings, 7 replies; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-20 19:56 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
This series introduces two callbacks for bpf_tcp_ops:
.enqueue_rcvq(): invoked when TCP stack enqueues skb to
sk->sk_receive_queue
.dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
is dequeued from sk->sk_receive_queue
Those callbacks can be enabled on a per-socket basis by
bpf_setsockopt():
int flags = BPF_SOCK_OPS_RCVQ_CB_FLAG;
bpf_setsockopt(sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
&flags, sizeof(flags));
or via the bpf_tcp_ops-specific helper added in the next patch:
bpf_sock_ops_cb_flags_set(sk, BPF_SOCK_OPS_RCVQ_CB_FLAG);
This allows the BPF prog to dynamically adjust sk->sk_rcvlowat,
suppressing unnecessary EPOLLIN wakeups until sufficient data
is available in the receive queue.
This functionality, which we call "TCP AutoLOWAT", was originally
developed in 2020 by Tenzin Ukyab with the help of Soheil Hassas
Yeganeh, Arjun Roy, and Eric Dumazet. It has served Google RPC
workloads for more than 5 years.
Combined with TCP RX zerocopy, this typically allows us to read an
entire RPC frame with just a single wakeup and a single system call.
While the original implementation was specialised for our
internal RPC format, this series introduces a more flexible
version by leveraging BPF.
The bpf prog in the last selftest patch closely mirrors the core
logic of the original implementation to provide a real-world
example.
Note that the new callbacks are not supported on legacy SOCK_OPS.
Legacy SOCK_OPS version:
v3: https://lore.kernel.org/bpf/20260523083001.2911931-1-kuniyu@google.com/
v2: https://lore.kernel.org/bpf/20260522074601.1658705-1-kuniyu@google.com/
v1: https://lore.kernel.org/bpf/20260508073355.3916746-1-kuniyu@google.com/
Kuniyuki Iwashima (7):
selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv.
bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops.
tcp: Split out __tcp_set_rcvlowat().
bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG.
bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().
include/net/tcp.h | 34 ++
include/uapi/linux/bpf.h | 13 +-
net/core/filter.c | 10 +-
net/ipv4/bpf_tcp_ops.c | 96 ++++-
net/ipv4/tcp.c | 14 +-
net/ipv4/tcp_fastopen.c | 2 +
net/ipv4/tcp_input.c | 4 +
tools/include/uapi/linux/bpf.h | 13 +-
.../selftests/bpf/prog_tests/tcp_autolowat.c | 350 ++++++++++++++++++
.../selftests/bpf/prog_tests/tcpbpf_user.c | 3 +-
.../selftests/bpf/progs/bpf_tracing_net.h | 2 +
.../selftests/bpf/progs/tcp_autolowat.c | 312 ++++++++++++++++
.../selftests/bpf/progs/test_tcpbpf_kern.c | 3 +-
13 files changed, 842 insertions(+), 14 deletions(-)
create mode 100644 tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
create mode 100644 tools/testing/selftests/bpf/progs/tcp_autolowat.c
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply [flat|nested] 28+ messages in thread
* [PATCH bpf-next 1/7] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv.
2026-09-20 19:56 [PATCH bpf-next 0/7] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
@ 2026-09-20 19:56 ` Kuniyuki Iwashima
2026-09-21 20:48 ` Stanislav Fomichev
2026-09-20 19:56 ` [PATCH bpf-next 2/7] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
` (5 subsequent siblings)
6 siblings, 1 reply; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-20 19:56 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
Once bpf_sock_ops_cb_flags_set() supports a new flag,
tcpbpf_user.c fails due to the hard-coded max value, 0x80.
Let's replace 0x80 with BPF_SOCK_OPS_ALL_CB_FLAGS + 1.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c | 3 ++-
tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c | 3 ++-
2 files changed, 4 insertions(+), 2 deletions(-)
diff --git a/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c b/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
index 7e8fe1bad03f..e4849d2a2956 100644
--- a/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
+++ b/tools/testing/selftests/bpf/prog_tests/tcpbpf_user.c
@@ -26,7 +26,8 @@ static void verify_result(struct tcpbpf_globals *result)
ASSERT_EQ(result->bytes_acked, 1002, "bytes_acked");
ASSERT_EQ(result->data_segs_in, 1, "data_segs_in");
ASSERT_EQ(result->data_segs_out, 1, "data_segs_out");
- ASSERT_EQ(result->bad_cb_test_rv, 0x80, "bad_cb_test_rv");
+ ASSERT_EQ(result->bad_cb_test_rv, BPF_SOCK_OPS_ALL_CB_FLAGS + 1,
+ "bad_cb_test_rv");
ASSERT_EQ(result->good_cb_test_rv, 0, "good_cb_test_rv");
ASSERT_EQ(result->num_listen, 1, "num_listen");
diff --git a/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c b/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
index 6935f32eeb8f..e30cb1fab079 100644
--- a/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
+++ b/tools/testing/selftests/bpf/progs/test_tcpbpf_kern.c
@@ -92,7 +92,8 @@ int bpf_testcb(struct bpf_sock_ops *skops)
break;
case BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB:
/* Test failure to set largest cb flag (assumes not defined) */
- global.bad_cb_test_rv = bpf_sock_ops_cb_flags_set(skops, 0x80);
+ global.bad_cb_test_rv = bpf_sock_ops_cb_flags_set(skops,
+ BPF_SOCK_OPS_ALL_CB_FLAGS + 1);
/* Set callback */
global.good_cb_test_rv = bpf_sock_ops_cb_flags_set(skops,
BPF_SOCK_OPS_STATE_CB_FLAG);
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH bpf-next 2/7] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-09-20 19:56 [PATCH bpf-next 0/7] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
2026-09-20 19:56 ` [PATCH bpf-next 1/7] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv Kuniyuki Iwashima
@ 2026-09-20 19:56 ` Kuniyuki Iwashima
2026-09-20 21:16 ` bot+bpf-ci
2026-09-21 20:48 ` Stanislav Fomichev
2026-09-20 19:56 ` [PATCH bpf-next 3/7] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops Kuniyuki Iwashima
` (4 subsequent siblings)
6 siblings, 2 replies; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-20 19:56 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
Waking up a thread per packet is expensive when an application
processes variable-length frames (e.g., RPC) that span multiple
packets.
SO_RCVLOWAT can defer wakeups, but because the frame size is
encoded in a fixed-size descriptor at the start of each frame,
the application has to:
1. wake up and recv() the descriptor,
2. raise SO_RCVLOWAT to the payload size via setsockopt(),
3. wake up and recv() the payload, and
4. reset SO_RCVLOWAT back to the descriptor size via
setsockopt() for the next frame.
This requires an extra wakeup and two setsockopt() syscalls
for every single RPC frame.
With SOCKMAP, we can parse skb and suppress wakeups in kernel,
but SOCKMAP adds overhead and also kills zerocopy.
Let's add lighter-weight opt-in callbacks to bpf_tcp_ops to
replace that.
.enqueue_rcvq(): invoked when TCP stack enqueues skb to
sk->sk_receive_queue
.dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
is dequeued from sk->sk_receive_queue
Those callbacks can be enabled on a per-socket basis by
bpf_setsockopt():
int flags = BPF_SOCK_OPS_RCVQ_CB_FLAG;
bpf_setsockopt(sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
&flags, sizeof(flags));
or via the bpf_tcp_ops-specific helper added in the next patch:
bpf_sock_ops_cb_flags_set(sk, BPF_SOCK_OPS_RCVQ_CB_FLAG);
Later, we will add a new kfunc to adjust sk->sk_rcvlowat from
these callbacks.
This will allow the bpf_tcp_ops prog to parse each skb and
dynamically adjust sk->sk_rcvlowat to suppress unnecessary EPOLLIN
wakeups until sufficient data is available in the receive queue.
The placement of bpf_tcp_ops_call() in tcp_ofo_queue() and
tcp_fastopen_add_skb() is chosen to provide the same snapshot
as tcp_queue_rcv().
For example, if bpf_tcp_ops_call() were called before updating
TCP_SKB_CB(skb)->seq in tcp_fastopen_add_skb(), BPF prog would
need an extra branch for the unlikely TFO case to strip SYN.
In addition, the TCP stack can queue overlapping skbs into recvq.
Once rcv_nxt is updated with a new skb, BPF prog can no longer
infer the previous rcv_nxt from skb->len.
Lastly, dequeue_rcvq() is placed in tcp_cleanup_rbuf() rather
than __tcp_cleanup_rbuf() so that it is not called for sockets
in SOCKMAP, where calling sk->sk_data_ready() from the new
kfunc would otherwise trigger infinite recursion.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
include/net/tcp.h | 18 ++++++++++++++++++
include/uapi/linux/bpf.h | 11 ++++++++++-
net/ipv4/bpf_tcp_ops.c | 10 ++++++++++
net/ipv4/tcp.c | 2 ++
net/ipv4/tcp_fastopen.c | 2 ++
net/ipv4/tcp_input.c | 4 ++++
tools/include/uapi/linux/bpf.h | 11 ++++++++++-
7 files changed, 56 insertions(+), 2 deletions(-)
diff --git a/include/net/tcp.h b/include/net/tcp.h
index d61ee00052e3..f2d838bcb0a7 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -3053,6 +3053,12 @@ struct bpf_tcp_ops {
struct request_sock *req, struct sk_buff *syn_skb,
enum tcp_synack_type synack_type,
u32 opt_off);
+
+ /* Called when an incoming skb is enqueued to sk->sk_receive_queue. */
+ void (*enqueue_rcvq)(struct sock *sk, struct sk_buff *skb);
+
+ /* Called after data is dequeued from sk->sk_receive_queue. */
+ void (*dequeue_rcvq)(struct sock *sk);
};
#define bpf_tcp_ops_call(op, sk, ...) \
@@ -3144,6 +3150,18 @@ static inline void tcp_bpf_rtt(struct sock *sk, long mrtt, u32 srtt)
bpf_tcp_ops_call(rtt, sk, mrtt, srtt);
}
+static inline void bpf_tcp_ops_enqueue_rcvq(struct sock *sk, struct sk_buff *skb)
+{
+ if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
+ bpf_tcp_ops_call(enqueue_rcvq, sk, skb);
+}
+
+static inline void bpf_tcp_ops_dequeue_rcvq(struct sock *sk)
+{
+ if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
+ bpf_tcp_ops_call(dequeue_rcvq, sk);
+}
+
#if IS_ENABLED(CONFIG_SMC)
extern struct static_key_false tcp_have_smc;
#endif
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 6330b7d745c5..fe122242b096 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -7148,8 +7148,17 @@ enum {
* options first before the BPF program does.
*/
BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
+ /* Call bpf when the TCP stack enqueues/dequeues payload
+ * to/from sk->sk_receive_queue.
+ *
+ * Only bpf_tcp_ops is supported.
+ *
+ * It can be used to adjust sk->sk_rcvlowat and suppress
+ * unnecessary wakeups before sufficient data is available.
+ */
+ BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
/* Mask of all currently supported cb flags */
- BPF_SOCK_OPS_ALL_CB_FLAGS = 0x7F,
+ BPF_SOCK_OPS_ALL_CB_FLAGS = 0xFF,
};
enum {
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index 681fed642999..c68d1fa32305 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -76,6 +76,14 @@ static void write_hdr_opt_stub(struct sock *sk, struct sk_buff *skb,
{
}
+static void enqueue_rcvq_stub(struct sock *sk, struct sk_buff *skb)
+{
+}
+
+static void dequeue_rcvq_stub(struct sock *sk)
+{
+}
+
static struct bpf_tcp_ops __bpf_tcp_ops = {
.timeout_init = timeout_init_stub,
.rwnd_init = rwnd_init_stub,
@@ -90,6 +98,8 @@ static struct bpf_tcp_ops __bpf_tcp_ops = {
.parse_hdr = parse_hdr_stub,
.hdr_opt_len = hdr_opt_len_stub,
.write_hdr_opt = write_hdr_opt_stub,
+ .enqueue_rcvq = enqueue_rcvq_stub,
+ .dequeue_rcvq = dequeue_rcvq_stub,
};
BPF_CALL_4(bpf_tcp_ops_store_hdr_opt, void *, ctx, const void *, from,
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index a4456b419412..a714b36a7494 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -1610,6 +1610,8 @@ void tcp_cleanup_rbuf(struct sock *sk, int copied)
"cleanup rbuf bug: copied %X seq %X rcvnxt %X\n",
tp->copied_seq, TCP_SKB_CB(skb)->end_seq, tp->rcv_nxt);
__tcp_cleanup_rbuf(sk, copied);
+
+ bpf_tcp_ops_dequeue_rcvq(sk);
}
static void tcp_eat_recv_skb(struct sock *sk, struct sk_buff *skb)
diff --git a/net/ipv4/tcp_fastopen.c b/net/ipv4/tcp_fastopen.c
index 471c78be5513..4939bcbc81d1 100644
--- a/net/ipv4/tcp_fastopen.c
+++ b/net/ipv4/tcp_fastopen.c
@@ -281,6 +281,8 @@ void tcp_fastopen_add_skb(struct sock *sk, struct sk_buff *skb)
TCP_SKB_CB(skb)->seq++;
TCP_SKB_CB(skb)->tcp_flags &= ~TCPHDR_SYN;
+ bpf_tcp_ops_enqueue_rcvq(sk, skb);
+
tp->rcv_nxt = TCP_SKB_CB(skb)->end_seq;
tcp_add_receive_queue(sk, skb);
tp->syn_data_acked = 1;
diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
index 6ac6f9d5b6c3..c60c61bb0a71 100644
--- a/net/ipv4/tcp_input.c
+++ b/net/ipv4/tcp_input.c
@@ -5344,6 +5344,8 @@ static void tcp_ofo_queue(struct sock *sk)
continue;
}
+ bpf_tcp_ops_enqueue_rcvq(sk, skb);
+
tail = skb_peek_tail(&sk->sk_receive_queue);
eaten = tail && tcp_try_coalesce(sk, tail, skb, &fragstolen);
tcp_rcv_nxt_update(tp, TCP_SKB_CB(skb)->end_seq);
@@ -5547,6 +5549,8 @@ static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb,
int eaten;
struct sk_buff *tail = skb_peek_tail(&sk->sk_receive_queue);
+ bpf_tcp_ops_enqueue_rcvq(sk, skb);
+
eaten = (tail &&
tcp_try_coalesce(sk, tail,
skb, fragstolen)) ? 1 : 0;
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 6330b7d745c5..fe122242b096 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -7148,8 +7148,17 @@ enum {
* options first before the BPF program does.
*/
BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG = (1<<6),
+ /* Call bpf when the TCP stack enqueues/dequeues payload
+ * to/from sk->sk_receive_queue.
+ *
+ * Only bpf_tcp_ops is supported.
+ *
+ * It can be used to adjust sk->sk_rcvlowat and suppress
+ * unnecessary wakeups before sufficient data is available.
+ */
+ BPF_SOCK_OPS_RCVQ_CB_FLAG = (1<<7),
/* Mask of all currently supported cb flags */
- BPF_SOCK_OPS_ALL_CB_FLAGS = 0x7F,
+ BPF_SOCK_OPS_ALL_CB_FLAGS = 0xFF,
};
enum {
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH bpf-next 3/7] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops.
2026-09-20 19:56 [PATCH bpf-next 0/7] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
2026-09-20 19:56 ` [PATCH bpf-next 1/7] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv Kuniyuki Iwashima
2026-09-20 19:56 ` [PATCH bpf-next 2/7] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
@ 2026-09-20 19:56 ` Kuniyuki Iwashima
2026-09-20 20:11 ` sashiko-bot
2026-09-21 20:48 ` Stanislav Fomichev
2026-09-20 19:56 ` [PATCH bpf-next 4/7] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
` (3 subsequent siblings)
6 siblings, 2 replies; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-20 19:56 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
When an error occurs in bpf_tcp_ops.{enqueue,dequeue}_rcvq(),
we want to clear BPF_SOCK_OPS_RCVQ_CB_FLAG to stop invoking
the callbacks.
In addition, bpf_sock_ops_cb_flags_set() is often used during
setup or just after 3WHS completes to enable opt-in hooks.
Let's support bpf_sock_ops_cb_flags_set() in the following
callbacks: connect, listen, {active,passive}_established,
{enqueue,dequeue}_rcvq.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
include/uapi/linux/bpf.h | 2 +-
net/ipv4/bpf_tcp_ops.c | 27 +++++++++++++++++++++++++++
tools/include/uapi/linux/bpf.h | 2 +-
3 files changed, 29 insertions(+), 2 deletions(-)
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index fe122242b096..8cdf22667775 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -3264,7 +3264,7 @@ union bpf_attr {
* Return
* 0
*
- * long bpf_sock_ops_cb_flags_set(struct bpf_sock_ops *bpf_sock, int argval)
+ * long bpf_sock_ops_cb_flags_set(void *bpf_sock, int argval)
* Description
* Attempt to set the value of the **bpf_sock_ops_cb_flags** field
* for the full TCP socket associated to *bpf_sock_ops* to
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index c68d1fa32305..6d0452441b6c 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -220,6 +220,24 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
.ret_type = RET_INTEGER,
};
+BPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)
+{
+ int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
+
+ tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
+
+ return argval & ~BPF_SOCK_OPS_ALL_CB_FLAGS;
+}
+
+static const struct bpf_func_proto bpf_tcp_ops_cb_flags_set_proto = {
+ .func = bpf_tcp_ops_cb_flags_set,
+ .gpl_only = false,
+ .ret_type = RET_INTEGER,
+ .arg1_type = ARG_PTR_TO_BTF_ID,
+ .arg1_btf_id = &btf_sock_ids[BTF_SOCK_TYPE_TCP],
+ .arg2_type = ARG_ANYTHING,
+};
+
static const struct bpf_func_proto *
get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
{
@@ -265,6 +283,15 @@ get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
if (moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))
return &bpf_tcp_ops_store_hdr_opt_proto;
return NULL;
+ case BPF_FUNC_sock_ops_cb_flags_set:
+ if (moff == offsetof(struct bpf_tcp_ops, connect) ||
+ moff == offsetof(struct bpf_tcp_ops, listen) ||
+ moff == offsetof(struct bpf_tcp_ops, active_established) ||
+ moff == offsetof(struct bpf_tcp_ops, passive_established) ||
+ moff == offsetof(struct bpf_tcp_ops, enqueue_rcvq) ||
+ moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+ return &bpf_tcp_ops_cb_flags_set_proto;
+ return NULL;
default:
return bpf_base_func_proto(func_id, prog);
}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index fe122242b096..8cdf22667775 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -3264,7 +3264,7 @@ union bpf_attr {
* Return
* 0
*
- * long bpf_sock_ops_cb_flags_set(struct bpf_sock_ops *bpf_sock, int argval)
+ * long bpf_sock_ops_cb_flags_set(void *bpf_sock, int argval)
* Description
* Attempt to set the value of the **bpf_sock_ops_cb_flags** field
* for the full TCP socket associated to *bpf_sock_ops* to
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH bpf-next 4/7] tcp: Split out __tcp_set_rcvlowat().
2026-09-20 19:56 [PATCH bpf-next 0/7] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
` (2 preceding siblings ...)
2026-09-20 19:56 ` [PATCH bpf-next 3/7] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops Kuniyuki Iwashima
@ 2026-09-20 19:56 ` Kuniyuki Iwashima
2026-09-20 21:01 ` bot+bpf-ci
2026-09-21 20:48 ` Stanislav Fomichev
2026-09-20 19:56 ` [PATCH bpf-next 5/7] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG Kuniyuki Iwashima
` (2 subsequent siblings)
6 siblings, 2 replies; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-20 19:56 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
We will add a kfunc for bpf_tcp_ops.{enqueue,dequeue}_rcvq()
to adjust sk->sk_rcvlowat.
These hooks are triggered
* when the TCP stack enqueues an skb to sk->sk_receive_queue
* after data is dequeued from sk->sk_receive_queue
In the enqueue path, tcp_data_ready() is always called after
the hooks in tcp_queue_rcv() and tcp_ofo_queue().
If tcp_set_rcvlowat() were used as is, tcp_data_ready() could
be called twice for the same skb, which is redundant and also
confusing.
Let's split out __tcp_set_rcvlowat() and add a flag to control
wakeup behaviour.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
include/net/tcp.h | 1 +
net/ipv4/tcp.c | 12 +++++++++---
2 files changed, 10 insertions(+), 3 deletions(-)
diff --git a/include/net/tcp.h b/include/net/tcp.h
index f2d838bcb0a7..07426e8641b7 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -512,6 +512,7 @@ void tcp_set_keepalive(struct sock *sk, int val);
void tcp_syn_ack_timeout(const struct request_sock *req);
int tcp_recvmsg(struct sock *sk, struct msghdr *msg, size_t len,
int flags);
+int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup);
int tcp_set_rcvlowat(struct sock *sk, int val);
void tcp_set_rcvbuf(struct sock *sk, int val);
int tcp_set_window_clamp(struct sock *sk, int val);
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index a714b36a7494..aa7593fc8334 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -1828,8 +1828,7 @@ int tcp_peek_len(struct socket *sock)
return tcp_inq(sock->sk);
}
-/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */
-int tcp_set_rcvlowat(struct sock *sk, int val)
+int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup)
{
struct tcp_sock *tp = tcp_sk(sk);
int space, cap;
@@ -1842,7 +1841,8 @@ int tcp_set_rcvlowat(struct sock *sk, int val)
WRITE_ONCE(sk->sk_rcvlowat, val ? : 1);
/* Check if we need to signal EPOLLIN right now */
- tcp_data_ready(sk);
+ if (wakeup)
+ tcp_data_ready(sk);
if (sk->sk_userlocks & SOCK_RCVBUF_LOCK)
return 0;
@@ -1857,6 +1857,12 @@ int tcp_set_rcvlowat(struct sock *sk, int val)
return 0;
}
+/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */
+int tcp_set_rcvlowat(struct sock *sk, int val)
+{
+ return __tcp_set_rcvlowat(sk, val, true);
+}
+
void tcp_set_rcvbuf(struct sock *sk, int val)
{
tcp_set_window_clamp(sk, tcp_win_from_space(sk, val));
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH bpf-next 5/7] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG.
2026-09-20 19:56 [PATCH bpf-next 0/7] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
` (3 preceding siblings ...)
2026-09-20 19:56 ` [PATCH bpf-next 4/7] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
@ 2026-09-20 19:56 ` Kuniyuki Iwashima
2026-09-20 20:05 ` sashiko-bot
2026-09-20 19:56 ` [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
2026-09-20 19:56 ` [PATCH bpf-next 7/7] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
6 siblings, 1 reply; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-20 19:56 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
The next patch exposes a new kfunc calling __tcp_set_rcvlowat()
to bpf_tcp_ops.
MPTCP has its own sock->ops->set_rcvlowat() / mptcp_set_rcvlowat(),
so we should not allow calling __tcp_set_rcvlowat() on MPTCP
subflows.
Let's disable BPF_SOCK_OPS_RCVQ_CB_FLAG for MPTCP for now.
If needed in the future, bpf_tcp_ops_set_rcvlowat() could be
extended to properly support MPTCP.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
include/net/tcp.h | 15 +++++++++++++++
net/core/filter.c | 10 ++++++----
net/ipv4/bpf_tcp_ops.c | 5 ++++-
3 files changed, 25 insertions(+), 5 deletions(-)
diff --git a/include/net/tcp.h b/include/net/tcp.h
index 07426e8641b7..d3cf655da9ec 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -2932,6 +2932,16 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
return tcp_call_bpf(sk, op, 3, args);
}
+static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
+{
+ if (sk_is_mptcp(sk) &&
+ (val & BPF_SOCK_OPS_RCVQ_CB_FLAG))
+ return -EOPNOTSUPP;
+
+ tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
+ return 0;
+}
+
static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
{
tcp_sk(sk)->bpf_sock_ops_cb_flags = 0;
@@ -2954,6 +2964,11 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
return -EPERM;
}
+static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
+{
+ return -EOPNOTSUPP;
+}
+
static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
{
}
diff --git a/net/core/filter.c b/net/core/filter.c
index 5feb99884682..f29c061bb066 100644
--- a/net/core/filter.c
+++ b/net/core/filter.c
@@ -5588,8 +5588,7 @@ static int bpf_sol_tcp_setsockopt(struct sock *sk, int optname,
case TCP_BPF_SOCK_OPS_CB_FLAGS:
if (val & ~(BPF_SOCK_OPS_ALL_CB_FLAGS))
return -EINVAL;
- tp->bpf_sock_ops_cb_flags = val;
- break;
+ return tcp_set_sock_ops_cb_flags(sk, val);
default:
return -EINVAL;
}
@@ -6178,8 +6177,9 @@ static const struct bpf_func_proto bpf_sock_ops_getsockopt_proto = {
BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
int, argval)
{
- struct sock *sk = bpf_sock->sk;
int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
+ struct sock *sk = bpf_sock->sk;
+ int err;
if (!is_locked_tcp_sock_ops(bpf_sock))
return -EOPNOTSUPP;
@@ -6187,7 +6187,9 @@ BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
if (!IS_ENABLED(CONFIG_INET) || !sk_fullsock(sk))
return -EINVAL;
- tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
+ err = tcp_set_sock_ops_cb_flags(sk, val);
+ if (err)
+ return err;
return argval & (~BPF_SOCK_OPS_ALL_CB_FLAGS);
}
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index 6d0452441b6c..b0e14b54917e 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -223,8 +223,11 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
BPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)
{
int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
+ int err;
- tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
+ err = tcp_set_sock_ops_cb_flags(sk, val);
+ if (err)
+ return err;
return argval & ~BPF_SOCK_OPS_ALL_CB_FLAGS;
}
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
2026-09-20 19:56 [PATCH bpf-next 0/7] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
` (4 preceding siblings ...)
2026-09-20 19:56 ` [PATCH bpf-next 5/7] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG Kuniyuki Iwashima
@ 2026-09-20 19:56 ` Kuniyuki Iwashima
2026-09-20 20:13 ` sashiko-bot
` (3 more replies)
2026-09-20 19:56 ` [PATCH bpf-next 7/7] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
6 siblings, 4 replies; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-20 19:56 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
bpf_tcp_ops.{enqueue,dequeue}_rcvq() were added to parse skb and
adjust sk->sk_rcvlowat dynamically to suppress unnecessary wakeups.
Let's add a new kfunc to set sk->sk_rcvlowat.
Negative values are clamped to INT_MAX, consistent with SO_RCVLOWAT.
For enqueue_rcvq(), wakeup is set to false because:
* tcp_data_ready() is always called after the hooks in
tcp_queue_rcv() and tcp_ofo_queue().
* when tcp_fastopen_add_skb() is called for TFO SYN, the socket is
not yet accept()ed, and when called for TFO SYN+ACK, the socket
is woken up by sk->sk_state_change() anyway.
For dequeue_rcvq(), wakeup is set to true because tcp_data_ready()
is not called in that path.
An alternative would be to support bpf_setsockopt() for these
hooks.
However, that approach involves excessive conditionals and an
unnecessary memcpy(), costs we do not want to pay for every skb
in the TCP fast path.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
net/ipv4/bpf_tcp_ops.c | 56 +++++++++++++++++++++++++++++++++++++++++-
1 file changed, 55 insertions(+), 1 deletion(-)
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index b0e14b54917e..3768b1440eb7 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -359,8 +359,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
.owner = THIS_MODULE,
};
+__bpf_kfunc_start_defs();
+
+__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
+ const struct bpf_prog_aux *aux)
+{
+ u32 moff = aux->attach_st_ops_member_off;
+ bool wakeup = false;
+
+ if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+ wakeup = true;
+
+ if (rcvlowat < 0)
+ rcvlowat = INT_MAX;
+
+ return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);
+}
+
+__bpf_kfunc_end_defs();
+
+BTF_KFUNCS_START(bpf_tcp_ops_rcvlowat_kfunc_set)
+BTF_ID_FLAGS(func, bpf_tcp_ops_set_rcvlowat, KF_IMPLICIT_ARGS)
+BTF_KFUNCS_END(bpf_tcp_ops_rcvlowat_kfunc_set)
+
+static int bpf_tcp_ops_rcvlowat_kfunc_filter(const struct bpf_prog *prog,
+ u32 kfunc_id)
+{
+ u32 moff;
+
+ if (!btf_id_set8_contains(&bpf_tcp_ops_rcvlowat_kfunc_set, kfunc_id))
+ return 0;
+
+ if (prog->aux->st_ops != &bpf_tcp_ops)
+ return -EACCES;
+
+ moff = prog->aux->attach_st_ops_member_off;
+ if (moff != offsetof(struct bpf_tcp_ops, enqueue_rcvq) &&
+ moff != offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+ return -EACCES;
+
+ return 0;
+}
+
+static const struct btf_kfunc_id_set bpf_tcp_ops_rcvlowat_kfunc_id_set = {
+ .owner = THIS_MODULE,
+ .set = &bpf_tcp_ops_rcvlowat_kfunc_set,
+ .filter = bpf_tcp_ops_rcvlowat_kfunc_filter,
+};
+
static int __init __bpf_tcp_ops_init(void)
{
- return register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
+ int ret;
+
+ ret = register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS,
+ &bpf_tcp_ops_rcvlowat_kfunc_id_set);
+ ret = ret ?: register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
+
+ return ret;
}
late_initcall(__bpf_tcp_ops_init);
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH bpf-next 7/7] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-09-20 19:56 [PATCH bpf-next 0/7] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
` (5 preceding siblings ...)
2026-09-20 19:56 ` [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
@ 2026-09-20 19:56 ` Kuniyuki Iwashima
2026-09-20 21:16 ` bot+bpf-ci
` (2 more replies)
6 siblings, 3 replies; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-20 19:56 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
The test is roughly divided into two stages, and the sequence
is as follows:
I) Setup
1. Attach two BPF programs to a cgroup
2. Establish a TCP connection (@client <-> @child) within the cgroup
3. Enable BPF_SOCK_OPS_RCVQ_CB_FLAG on @child via setsockopt()
II) RPC frame exchange in various patterns
4. Send a partial RPC descriptor from @client to @child
5. Verify that epoll does NOT wake up @child
6. Send the remaining data of the RPC frame
7. Verify that epoll finally wakes up @child
During setup, two BPF programs are attached to simulate
a real-world scenario; one is bpf_tcp_ops and the other is
CGROUP_SOCKOPT.
While the bpf_tcp_ops prog handles the dynamic adjustment of
sk->sk_rcvlowat, the CGROUP_SOCKOPT prog is used to enable
the TCP AutoLOWAT feature via userspace setsockopt() using
pseudo options:
#define SOL_BPF 0xdeadbeef
#define BPF_TCP_AUTOLOWAT 0x8badf00d
setsockopt(fd, SOL_BPF, BPF_TCP_AUTOLOWAT, &(int){1}, sizeof(int));
This reflects a common production use case where an application
decides to start parsing RPC frames only at a certain point in
the stream (e.g., after HTTP Upgrade), rather than immediately
after TCP 3WHS (BPF_SOCK_OPS_PASSIVE_ESTABLISHED_CB, etc).
When BPF_TCP_AUTOLOWAT is enabled, the BPF prog sets
BPF_SOCK_OPS_RCVQ_CB_FLAG and initializes sk_local_storage
for two sequence numbers to manage its state.
Then, for the RPC frame exchange, this test uses a simple format
defined as follows:
0 8 16 24 32
+--------+--------+-------+--------+ `.
| header size | |
+--------+--------+-------+--------+ > RPC descriptor (8 bytes)
| payload size | |
+--------+--------+-------+--------+ .'
~ header ~
+--------+--------+-------+--------+
~ payload ~
+--------+--------+-------+--------+
Every time a new skb is enqueued to sk->sk_receive_queue, the
bpf_tcp_ops prog parses it and updates these sequence numbers:
rpc_desc_seq : the SEQ # of the start of the RPC descriptor
rpc_end_seq : the SEQ # of the end of the RPC frame
=> rpc_desc_seq + 8 + header size + payload size
Assume we receive two RPC descriptors in the following pattern:
1. When we receive skb-1, only part of the RPC descriptor is parsed.
rpc_desc_seq is set to the first byte while rpc_end_seq is
unknown. Thus, sk->sk_rcvlowat is set to the size of the RPC
descriptor (8 bytes).
<- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
+-----------+.................+....................+......
| RPC desc 1 | header + payload | RPC desc 2 | ...
+-----------+.................+....................+......
^ ^-.
`- rpc_desc_seq `- sk->sk_rcvlowat
2. Next, we receive skb-2, which completes the first RPC descriptor.
Now rpc_end_seq is known, so sk->sk_rcvlowat is advanced to it.
<- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
+-----------+-----------------+....................+......
| RPC desc 1 | header + payload | RPC desc 2 | ...
+-----------+-----------------+....................+......
^ ^
'- rpc_desc_seq '- rpc_end_seq
& sk->sk_rcvlowat
3. Once we receive skb-3, which contains the next full RPC descriptor,
rpc_desc_seq is advanced and rpc_end_seq is updated according
to the size of RPC frame 2.
Note that sk->sk_rcvlowat is NOT updated to the new rpc_end_seq
yet. This ensures that the application is woken up to read the
already complete RPC frame 1.
<- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
+-----------+-----------------+--------------------+......
| RPC desc 1 | header + payload | RPC desc 2 | ... |
+-----------+-----------------+--------------------+......
^ ^
rpc_desc_seq -----------' rpc_end_seq ----...-'
& sk->sk_rcvlowat
This sequence corresponds to the 4th test case in rpc_test_cases[],
and we can see helpful output if we "#define DEBUG":
# cat /sys/kernel/tracing/trace_pipe | \
awk '{ if ($0 ~ /AF_/) sub(/^.*AF_/, "AF_"); print $0 }' & \
BGPID=$!; ./test_progs -t tcp_autolowat; kill -9 -$BGPID
...
AF_INET6 rpc_test_cases[3]: Start parsing skb: seq: 0, end_seq: 1, len: 1, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_buff_len: 0
AF_INET6 rpc_test_cases[3]: Copied 1 bytes: rpc_desc_buff_len: 1
AF_INET6 rpc_test_cases[3]: Setting rcvlowat: tp->copied_seq: 0, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_desc_buff_len: 1
AF_INET6 rpc_test_cases[3]: Set rcvlowat: expected: 8, actual: 8
AF_INET6 rpc_test_cases[3]: Start parsing skb: seq: 1, end_seq: 8, len: 7, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_buff_len: 1
AF_INET6 rpc_test_cases[3]: Copied full descriptor: rpc_desc_seq: 0, rpc_end_seq: 258, header_len: 100, payload_len: 150
AF_INET6 rpc_test_cases[3]: No more descriptor: rpc_end_seq: 258, end_seq: 8
AF_INET6 rpc_test_cases[3]: Setting rcvlowat: tp->copied_seq: 0, rpc_desc_seq: 0, rpc_end_seq: 258, rpc_desc_buff_len: 8
AF_INET6 rpc_test_cases[3]: Set rcvlowat: expected: 258, actual: 258
...
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
.../selftests/bpf/prog_tests/tcp_autolowat.c | 350 ++++++++++++++++++
.../selftests/bpf/progs/bpf_tracing_net.h | 2 +
.../selftests/bpf/progs/tcp_autolowat.c | 312 ++++++++++++++++
3 files changed, 664 insertions(+)
create mode 100644 tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
create mode 100644 tools/testing/selftests/bpf/progs/tcp_autolowat.c
diff --git a/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
new file mode 100644
index 000000000000..c5ead4af24f8
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
@@ -0,0 +1,350 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright 2026 Google LLC */
+#include <sys/epoll.h>
+
+#include "test_progs.h"
+#include "cgroup_helpers.h"
+#include "network_helpers.h"
+
+#include "tcp_autolowat.skel.h"
+
+#define SOL_BPF 0xdeadbeef
+#define BPF_TCP_AUTOLOWAT 0x8badf00d
+
+struct rpc_descriptor {
+ u32 header_len;
+ u32 payload_len;
+};
+
+enum rpc_event_type {
+ RPC_EVENT_END,
+ RPC_EVENT_AUTOLOWAT,
+ RPC_EVENT_SEND,
+ RPC_EVENT_RECV,
+ RPC_EVENT_EPOLL,
+ RPC_EVENT_RCVLOWAT,
+};
+
+struct rpc_event {
+ enum rpc_event_type type;
+ union {
+ int len;
+ int nfds;
+ int val;
+ int rcvlowat;
+ };
+};
+
+#define RPC_DESC_SIZE (sizeof(struct rpc_descriptor))
+
+struct rpc_test_case {
+ char data[4096];
+ struct rpc_descriptor desc[32];
+ struct rpc_event event[32];
+} rpc_test_cases[] = {
+ {
+ .desc = {
+ { .header_len = 100, .payload_len = 150 },
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* Single full RPC message in skb. */
+ { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE + 100 + 150},
+ { .type = RPC_EVENT_EPOLL, .nfds = 1},
+ { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 100 + 150},
+ },
+ },
+ {
+ .desc = {
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* Two full RPC messages in skb. */
+ {.type = RPC_EVENT_SEND, .len = (RPC_DESC_SIZE + 100 + 150) * 2},
+ {.type = RPC_EVENT_EPOLL, .nfds = 1},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+ /* Single full RPC message in skb. */
+ { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE + 100 + 150},
+ { .type = RPC_EVENT_EPOLL, .nfds = 1},
+ { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 3},
+ },
+ },
+ {
+ .desc = {
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* Two full RPC messages in skb. */
+ {.type = RPC_EVENT_SEND, .len = (RPC_DESC_SIZE + 100 + 150) * 2},
+ {.type = RPC_EVENT_EPOLL, .nfds = 1},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+ /* Single full RPC message in skb. */
+ { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE},
+ { .type = RPC_EVENT_EPOLL, .nfds = 1},
+ { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+ },
+ },
+ {
+ .desc = {
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 200, .payload_len = 500},
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* The first descriptor is partial. */
+ {.type = RPC_EVENT_SEND, .len = 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE},
+ /* The first descriptor is available. */
+ {.type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE - 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 150 + 100},
+ /* The first header is ready. */
+ {.type = RPC_EVENT_SEND, .len = 100},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 150 + 100},
+ /* skb has the first payload and 1 byte of the next descriptor. */
+ {.type = RPC_EVENT_SEND, .len = 150 + 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 1},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 150 + 100},
+ /* After reading the first RPC message, SO_RCVLOWAT should be RPC_DESC_SIZE. */
+ {.type = RPC_EVENT_RECV, .len = RPC_DESC_SIZE + 150 + 100},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE},
+ /* The second descriptor is available. */
+ {.type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE - 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 200 + 500},
+ },
+ },
+};
+
+struct tcp_autolowat_test_cb {
+ int saved_netns;
+ union {
+ int fd[4];
+ struct {
+ int server, client, child;
+ int epoll;
+ };
+ };
+};
+
+static void tcp_autolowat_teardown_cb(struct tcp_autolowat_test_cb *cb)
+{
+ int i, err;
+
+ for (i = 0; i < ARRAY_SIZE(cb->fd); i++) {
+ if (cb->fd[i] != -1)
+ close(cb->fd[i]);
+ }
+
+ if (cb->saved_netns != -1) {
+ err = setns(cb->saved_netns, CLONE_NEWNET);
+ ASSERT_OK(err, "restore netns");
+
+ close(cb->saved_netns);
+ }
+}
+
+static int tcp_autolowat_setup_cb(struct tcp_autolowat_test_cb *cb, int family)
+{
+ struct epoll_event ev = {};
+ int err;
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(cb->fd); i++)
+ cb->fd[i] = -1;
+
+ cb->saved_netns = open("/proc/self/ns/net", O_RDONLY);
+ if (!ASSERT_OK_FD(cb->saved_netns, "save netns"))
+ goto err;
+
+ err = unshare(CLONE_NEWNET);
+ if (!ASSERT_OK(err, "unshare"))
+ goto err;
+
+ err = system("ip link set dev lo up");
+ if (!ASSERT_OK(err, "set up lo"))
+ goto err;
+
+ cb->server = start_server(family, SOCK_STREAM, NULL, 0, 0);
+ if (!ASSERT_OK_FD(cb->server, "start_server"))
+ goto err;
+
+ cb->client = connect_to_fd(cb->server, 0);
+ if (!ASSERT_OK_FD(cb->client, "connect_to_fd"))
+ goto err;
+
+ cb->child = accept(cb->server, NULL, NULL);
+ if (!ASSERT_OK_FD(cb->child, "accept"))
+ goto err;
+
+ cb->epoll = epoll_create1(0);
+ if (!ASSERT_OK_FD(cb->epoll, "epoll_create"))
+ goto err;
+
+ ev.events = EPOLLIN;
+ ev.data.fd = cb->child;
+
+ err = epoll_ctl(cb->epoll, EPOLL_CTL_ADD, cb->child, &ev);
+ if (!ASSERT_OK(err, "epoll_ctl"))
+ goto err;
+
+ return 0;
+
+err:
+ tcp_autolowat_teardown_cb(cb);
+ return -1;
+}
+
+static int tcp_autolowat_build_data(struct rpc_test_case *test_case)
+{
+ struct rpc_descriptor *desc = test_case->desc;
+ char *ptr = test_case->data;
+ int rpc_size;
+
+ memset(ptr, 0, sizeof(test_case->data));
+
+ while (desc->header_len + desc->payload_len) {
+ rpc_size = sizeof(*desc) + desc->header_len + desc->payload_len;
+
+ if (!ASSERT_LE(ptr + rpc_size - test_case->data,
+ sizeof(test_case->data), "data overflow"))
+ return 1;
+
+ memcpy(ptr, desc, sizeof(*desc));
+ ptr += rpc_size;
+ desc++;
+ }
+
+ if (!ASSERT_GT(ptr - test_case->data, 0, "no data"))
+ return 1;
+
+ return 0;
+}
+
+static void tcp_autolowat_run_rpc_test(struct tcp_autolowat_test_cb *cb,
+ struct rpc_test_case *test_case)
+{
+ struct rpc_event *event = test_case->event;
+ char *ptr = test_case->data;
+ struct epoll_event ev;
+ socklen_t optlen;
+ int err, optval;
+ char buf[4096];
+
+ if (tcp_autolowat_build_data(test_case))
+ return;
+
+ while (1) {
+ switch (event->type) {
+ case RPC_EVENT_END:
+ return;
+ case RPC_EVENT_AUTOLOWAT:
+ err = setsockopt(cb->child, SOL_BPF, BPF_TCP_AUTOLOWAT,
+ &event->val, sizeof(event->val));
+ if (!ASSERT_OK(err, "setsockopt"))
+ return;
+ break;
+ case RPC_EVENT_SEND:
+ err = send(cb->client, ptr, event->len, 0);
+ if (!ASSERT_EQ(err, event->len, "send"))
+ return;
+
+ ptr += event->len;
+ break;
+ case RPC_EVENT_RECV:
+ err = recv(cb->child, buf, event->len, 0);
+ if (!ASSERT_EQ(err, event->len, "recv"))
+ return;
+ break;
+ case RPC_EVENT_EPOLL:
+ err = epoll_wait(cb->epoll, &ev, 1, 100);
+ if (!ASSERT_EQ(err, event->nfds, "epoll_wait"))
+ return;
+ break;
+ case RPC_EVENT_RCVLOWAT:
+ optval = 0;
+ optlen = sizeof(optval);
+
+ err = getsockopt(cb->child, SOL_SOCKET, SO_RCVLOWAT, &optval, &optlen);
+ if (!ASSERT_OK(err, "getsockopt") ||
+ !ASSERT_EQ(optval, event->rcvlowat, "rcvlowat"))
+ return;
+ break;
+ }
+
+ event++;
+ }
+}
+
+static void tcp_autolowat_run_rpc_tests(struct tcp_autolowat *skel, int family)
+{
+ struct tcp_autolowat_test_cb cb;
+ int err;
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(rpc_test_cases); i++) {
+ memset(skel->bss->test_name, 0, sizeof(skel->bss->test_name));
+
+ snprintf(skel->bss->test_name, sizeof(skel->bss->test_name),
+ "AF_INET%c rpc_test_cases[%d]",
+ family == AF_INET ? ' ' : '6', i);
+
+ if (!test__start_subtest(skel->bss->test_name))
+ continue;
+
+ err = tcp_autolowat_setup_cb(&cb, family);
+ if (err)
+ continue;
+
+ tcp_autolowat_run_rpc_test(&cb, &rpc_test_cases[i]);
+ tcp_autolowat_teardown_cb(&cb);
+ }
+}
+
+static void tcp_autolowat_run_tests(struct tcp_autolowat *skel)
+{
+ tcp_autolowat_run_rpc_tests(skel, AF_INET);
+ tcp_autolowat_run_rpc_tests(skel, AF_INET6);
+}
+
+void test_tcp_autolowat(void)
+{
+ struct tcp_autolowat *skel;
+ struct bpf_link *link[2];
+ int cgroup;
+
+ skel = tcp_autolowat__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "open_and_load"))
+ return;
+
+ cgroup = test__join_cgroup("/tcp_autolowat");
+ if (!ASSERT_GE(cgroup, 0, "join_cgroup"))
+ goto destroy_skel;
+
+ link[0] = bpf_map__attach_cgroup_opts(skel->maps.tcp_autolowat_ops, cgroup, NULL);
+ if (!ASSERT_OK_PTR(link[0], "attach_cgroup(tcp_autolowat_ops)"))
+ goto close_cgroup;
+
+ link[1] = bpf_program__attach_cgroup(skel->progs.tcp_autolowat_setsockopt, cgroup);
+ if (!ASSERT_OK_PTR(link[1], "attach_cgroup(SETSOCKOPT)"))
+ goto destroy_sockops;
+
+ tcp_autolowat_run_tests(skel);
+
+ bpf_link__destroy(link[1]);
+destroy_sockops:
+ bpf_link__destroy(link[0]);
+close_cgroup:
+ close(cgroup);
+destroy_skel:
+ tcp_autolowat__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/progs/bpf_tracing_net.h b/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
index 593b38f90417..4c999d59cbbc 100644
--- a/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
+++ b/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
@@ -79,6 +79,8 @@
#define NEXTHDR_TCP 6
+#define TCPHDR_FIN 0x01
+
#define TCPOPT_NOP 1
#define TCPOPT_EOL 0
#define TCPOPT_MSS 2
diff --git a/tools/testing/selftests/bpf/progs/tcp_autolowat.c b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
new file mode 100644
index 000000000000..bb96e19e7589
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
@@ -0,0 +1,312 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright 2026 Google LLC */
+#include "vmlinux.h"
+
+#include <string.h>
+#include <limits.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+#include <bpf/bpf_core_read.h>
+
+#include "bpf_kfuncs.h"
+#include "bpf_tracing_net.h"
+
+#define SOL_BPF 0xdeadbeef
+#define BPF_TCP_AUTOLOWAT 0x8badf00d
+
+//#define DEBUG /* For verbose output. */
+
+struct rpc_descriptor {
+ u32 header_len;
+ u32 payload_len;
+};
+
+#define RPC_DESC_SIZE (sizeof(struct rpc_descriptor))
+#define MAX_RPC_DESC_PER_SKB 100
+
+struct tcp_autolowat_cb {
+ /* Don't put this field at the end; BPF verifier complains. */
+ char rpc_desc_buf[RPC_DESC_SIZE];
+ u32 rpc_desc_seq;
+ u32 rpc_end_seq;
+#ifdef DEBUG
+ u32 isn;
+#endif
+ u8 rpc_desc_buff_len;
+};
+
+struct {
+ __uint(type, BPF_MAP_TYPE_SK_STORAGE);
+ __uint(map_flags, BPF_F_NO_PREALLOC);
+ __type(key, int);
+ __type(value, struct tcp_autolowat_cb);
+} tcp_autolowat_map SEC(".maps");
+
+char test_name[64];
+
+#ifdef DEBUG
+#define LOG(str, ...) \
+ bpf_printk("%s: " str, test_name, ##__VA_ARGS__)
+#else
+#define LOG(...)
+#endif
+
+#define SEQ(val) \
+ (val - cb->isn)
+#define TP_SEQ(field) \
+ (tp->field - cb->isn)
+#define CB_SEQ(field) \
+ (cb->field - cb->isn)
+
+static int tcp_parse_descriptor(struct tcp_autolowat_cb *cb,
+ struct bpf_dynptr *dptr,
+ u32 seq, u32 end_seq)
+{
+ struct rpc_descriptor *rpc_desc;
+ u32 rpc_copied_seq;
+ u64 copy_len; /* u32 should work, but not for no_alu32 :/ */
+ u64 rpc_len;
+ int err;
+
+ rpc_copied_seq = cb->rpc_desc_seq + cb->rpc_desc_buff_len;
+
+ if (before(cb->rpc_desc_seq + RPC_DESC_SIZE, end_seq))
+ copy_len = RPC_DESC_SIZE - cb->rpc_desc_buff_len;
+ else
+ copy_len = end_seq - rpc_copied_seq;
+
+ if (copy_len == 0)
+ goto disable; /* FIN. */
+ if (copy_len > RPC_DESC_SIZE)
+ goto disable; /* always false, only for verifier. */
+ if (cb->rpc_desc_buf + cb->rpc_desc_buff_len >= &cb->rpc_desc_buf[RPC_DESC_SIZE])
+ goto disable; /* always false, only for verifier. */
+
+ err = bpf_dynptr_read(cb->rpc_desc_buf + cb->rpc_desc_buff_len,
+ copy_len, dptr, rpc_copied_seq - seq, 0);
+ if (err)
+ goto disable;
+
+ cb->rpc_desc_buff_len += copy_len;
+
+ if (cb->rpc_desc_buff_len != RPC_DESC_SIZE) {
+ LOG("Copied %d bytes: rpc_desc_buff_len: %u", copy_len, cb->rpc_desc_buff_len);
+ goto partial;
+ }
+
+ rpc_desc = (struct rpc_descriptor *)cb->rpc_desc_buf;
+ rpc_len = RPC_DESC_SIZE + rpc_desc->header_len + rpc_desc->payload_len;
+
+ if (rpc_len > INT_MAX)
+ goto disable;
+
+ cb->rpc_end_seq = cb->rpc_desc_seq + rpc_len;
+
+ LOG("Copied full descriptor: rpc_desc_seq: %u, rpc_end_seq: %u, header_len: %u, payload_len: %u",
+ CB_SEQ(rpc_desc_seq), CB_SEQ(rpc_end_seq),
+ rpc_desc->header_len, rpc_desc->payload_len);
+
+ return 0;
+disable:
+ return -1;
+partial:
+ return 1;
+}
+
+static void tcp_set_autolowat(struct tcp_autolowat_cb *cb,
+ struct sock *sk)
+{
+ struct tcp_sock *tp = (struct tcp_sock *)sk;
+ u32 val; /* To handle wraparound. */
+
+ LOG("Setting rcvlowat: tp->copied_seq: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_desc_buff_len: %u",
+ TP_SEQ(copied_seq), CB_SEQ(rpc_desc_seq),
+ CB_SEQ(rpc_end_seq), cb->rpc_desc_buff_len);
+
+ if (before(tp->copied_seq, cb->rpc_desc_seq))
+ val = cb->rpc_desc_seq - tp->copied_seq;
+ else if (cb->rpc_desc_buff_len != RPC_DESC_SIZE)
+ val = RPC_DESC_SIZE;
+ else
+ val = cb->rpc_end_seq - tp->copied_seq;
+
+ if (val != tp->inet_conn.icsk_inet.sk.sk_rcvlowat) {
+ bpf_tcp_ops_set_rcvlowat(sk, val);
+
+ LOG("Set rcvlowat: expected: %u, actual: %d\n",
+ val, tp->inet_conn.icsk_inet.sk.sk_rcvlowat);
+ } else {
+ LOG("No need to set rcvlowat: %u\n", val);
+ }
+}
+
+static void tcp_disable_autolowat(struct sock *sk)
+{
+ struct tcp_sock *tp = (struct tcp_sock *)sk;
+ int flags;
+
+ flags = tp->bpf_sock_ops_cb_flags & ~BPF_SOCK_OPS_RCVQ_CB_FLAG;
+ bpf_sock_ops_cb_flags_set(sk, flags);
+
+ bpf_tcp_ops_set_rcvlowat(sk, 1);
+
+ LOG("Disabled autolowat");
+}
+
+static void tcp_do_autolowat(struct tcp_autolowat_cb *cb,
+ struct sock *sk, struct sk_buff *skb)
+{
+ struct bpf_dynptr dptr;
+ struct tcp_skb_cb *tcb;
+ u32 seq, end_seq;
+ int ret = 0, i;
+
+ if (bpf_dynptr_from_skb((struct __sk_buff *)skb, 0, &dptr)) {
+ ret = -1;
+ goto update;
+ }
+
+ tcb = bpf_core_cast(skb->cb, struct tcp_skb_cb);
+ seq = tcb->seq;
+ end_seq = tcb->end_seq - !!(tcb->tcp_flags & TCPHDR_FIN);
+
+ LOG("Start parsing skb: seq: %u, end_seq: %u, len: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_buff_len: %u",
+ SEQ(seq), SEQ(end_seq), end_seq - seq,
+ CB_SEQ(rpc_desc_seq), CB_SEQ(rpc_end_seq), cb->rpc_desc_buff_len);
+
+ if (cb->rpc_desc_buff_len != RPC_DESC_SIZE) {
+ ret = tcp_parse_descriptor(cb, &dptr, seq, end_seq);
+ if (ret)
+ goto update;
+ }
+
+ i = 0;
+
+ while (1) {
+ if (i++ > MAX_RPC_DESC_PER_SKB) {
+ ret = -1;
+ break;
+ }
+
+ if (after(cb->rpc_end_seq, end_seq)) {
+ LOG("No more descriptor: rpc_end_seq: %u, end_seq: %u",
+ CB_SEQ(rpc_end_seq), SEQ(end_seq));
+ break;
+ }
+
+ cb->rpc_desc_seq = cb->rpc_end_seq;
+ cb->rpc_desc_buff_len = 0;
+
+ if (cb->rpc_end_seq == end_seq)
+ break;
+
+ LOG("Found next descriptor: rpc_end_seq: %u, end_seq: %u, len: %u",
+ CB_SEQ(rpc_end_seq), SEQ(end_seq), end_seq - cb->rpc_end_seq);
+
+ ret = tcp_parse_descriptor(cb, &dptr, seq, end_seq);
+ if (ret)
+ break;
+ }
+
+update:
+ if (ret >= 0)
+ tcp_set_autolowat(cb, sk);
+ else
+ tcp_disable_autolowat(sk);
+}
+
+SEC("struct_ops")
+void BPF_PROG(tcp_autolowat_enqueue_rcvq, struct sock *sk, struct sk_buff *skb)
+{
+ struct tcp_autolowat_cb *cb;
+
+ cb = bpf_sk_storage_get(&tcp_autolowat_map, sk, 0, 0);
+ if (!cb)
+ return;
+
+ tcp_do_autolowat(cb, sk, skb);
+}
+
+SEC("struct_ops")
+void BPF_PROG(tcp_autolowat_dequeue_rcvq, struct sock *sk)
+{
+ struct tcp_autolowat_cb *cb;
+
+ cb = bpf_sk_storage_get(&tcp_autolowat_map, sk, 0, 0);
+ if (!cb)
+ return;
+
+ tcp_set_autolowat(cb, sk);
+}
+
+SEC(".struct_ops.link")
+struct bpf_tcp_ops tcp_autolowat_ops = {
+ .enqueue_rcvq = (void *)tcp_autolowat_enqueue_rcvq,
+ .dequeue_rcvq = (void *)tcp_autolowat_dequeue_rcvq,
+};
+
+static int tcp_init_autolowat_cb(struct bpf_sockopt *sockopt,
+ struct bpf_tcp_sock *btp)
+{
+ struct tcp_autolowat_cb *cb;
+ struct tcp_sock *tp;
+ int flags;
+
+ cb = bpf_sk_storage_get(&tcp_autolowat_map, btp, 0,
+ BPF_SK_STORAGE_GET_F_CREATE);
+ if (!cb)
+ return -1;
+
+ tp = bpf_core_cast(btp, struct tcp_sock);
+ if (!tp)
+ return -1;
+
+ cb->rpc_desc_seq = tp->copied_seq;
+ cb->rpc_end_seq = tp->copied_seq;
+#ifdef DEBUG
+ cb->isn = tp->copied_seq;
+#endif
+
+ if (bpf_getsockopt(sockopt->sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
+ &flags, sizeof(flags)))
+ return -1;
+
+ flags |= BPF_SOCK_OPS_RCVQ_CB_FLAG;
+
+ if (bpf_setsockopt(sockopt->sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
+ &flags, sizeof(flags)))
+ return -1;
+
+ return 0;
+}
+
+SEC("cgroup/setsockopt")
+int tcp_autolowat_setsockopt(struct bpf_sockopt *ctx)
+{
+ void *optval_end = ctx->optval_end;
+ int *optval = ctx->optval;
+ struct bpf_tcp_sock *btp;
+
+ if (ctx->level != SOL_BPF || ctx->optname != BPF_TCP_AUTOLOWAT)
+ goto out;
+
+ if (optval + 1 > optval_end)
+ return 0; /* -EPERM */
+
+ btp = bpf_tcp_sock(ctx->sk);
+ if (!btp)
+ goto out;
+
+ if (*optval && tcp_init_autolowat_cb(ctx, btp))
+ return 0; /* -EPERM */
+
+ /*
+ * BPF has consumed this option, don't call kernel
+ * setsockopt handler.
+ */
+ ctx->optlen = -1;
+out:
+ return 1;
+}
+
+char _license[] SEC("license") = "GPL";
--
2.55.0.1082.g2b9226bbc0-goog
^ permalink raw reply related [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 5/7] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG.
2026-09-20 19:56 ` [PATCH bpf-next 5/7] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG Kuniyuki Iwashima
@ 2026-09-20 20:05 ` sashiko-bot
2026-09-20 21:09 ` Kuniyuki Iwashima
0 siblings, 1 reply; 28+ messages in thread
From: sashiko-bot @ 2026-09-20 20:05 UTC (permalink / raw)
To: Kuniyuki Iwashima; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: tcp_set_sock_ops_cb_flags breaks ALL_CB_FLAGS for MPTCP
--
commit 578335638acc4a540df534d791edf9c63d563e32
Author: Kuniyuki Iwashima <kuniyu@google.com>
Subject: bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG.
This patch prepares for a new kfunc by disabling BPF_SOCK_OPS_RCVQ_CB_FLAG
for MPTCP subflows. It returns -EOPNOTSUPP if this flag is used on an MPTCP
socket.
> diff --git a/include/net/tcp.h b/include/net/tcp.h
> index 07426e8641b7e..d3cf655da9ecd 100644
> --- a/include/net/tcp.h
> +++ b/include/net/tcp.h
> @@ -2932,6 +2932,16 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
> return tcp_call_bpf(sk, op, 3, args);
> }
>
> +static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
> +{
> + if (sk_is_mptcp(sk) &&
> + (val & BPF_SOCK_OPS_RCVQ_CB_FLAG))
> + return -EOPNOTSUPP;
> +
[Severity: Medium]
Will this check unintentionally break existing BPF programs that use
BPF_SOCK_OPS_ALL_CB_FLAGS to enable all callbacks?
If a program passes BPF_SOCK_OPS_ALL_CB_FLAGS (which is expanded in this
patch series to include the new BPF_SOCK_OPS_RCVQ_CB_FLAG bit), this rigid
check evaluates to true for MPTCP subflows and rejects the entire multi-flag
request instead of ignoring the unsupported bit.
> + tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
> + return 0;
> +}
> +
[ ... ]
> diff --git a/net/core/filter.c b/net/core/filter.c
> index 5feb998846825..f29c061bb0664 100644
> --- a/net/core/filter.c
> +++ b/net/core/filter.c
[ ... ]
> @@ -6178,8 +6177,9 @@ static const struct bpf_func_proto bpf_sock_ops_getsockopt_proto = {
> BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
> int, argval)
> {
> - struct sock *sk = bpf_sock->sk;
> int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
> + struct sock *sk = bpf_sock->sk;
> + int err;
>
> if (!is_locked_tcp_sock_ops(bpf_sock))
> return -EOPNOTSUPP;
> @@ -6187,7 +6187,9 @@ BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
> if (!IS_ENABLED(CONFIG_INET) || !sk_fullsock(sk))
> return -EINVAL;
>
> - tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
> + err = tcp_set_sock_ops_cb_flags(sk, val);
> + if (err)
> + return err;
[Severity: Medium]
Does returning early here prevent any of the other valid callbacks requested
in val from being enabled?
Because tcp_set_sock_ops_cb_flags() returns -EOPNOTSUPP when the new flag is
present, the BPF helper propagates the error without applying any of the
previously supported flags. This effectively turns a wildcard request for
"all supported flags" into a total failure on MPTCP subflows.
>
> return argval & (~BPF_SOCK_OPS_ALL_CB_FLAGS);
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260920195633.3033620-1-kuniyu@google.com?part=5
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 3/7] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops.
2026-09-20 19:56 ` [PATCH bpf-next 3/7] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops Kuniyuki Iwashima
@ 2026-09-20 20:11 ` sashiko-bot
2026-09-20 21:05 ` Kuniyuki Iwashima
2026-09-21 20:48 ` Stanislav Fomichev
1 sibling, 1 reply; 28+ messages in thread
From: sashiko-bot @ 2026-09-20 20:11 UTC (permalink / raw)
To: Kuniyuki Iwashima; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: data race in bpf_tcp_ops_cb_flags_set() on unlocked sockets
--
commit 56d7f542a8c499237807954e17f2fae35e5e276b
Author: Kuniyuki Iwashima <kuniyu@google.com>
bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops.
This patch allows clearing or setting bpf_sock_ops_cb_flags via the
bpf_sock_ops_cb_flags_set() helper within various bpf_tcp_ops callbacks.
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index c68d1fa323058..6d0452441b6c5 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
[ ... ]
> @@ -220,6 +220,24 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
> .ret_type = RET_INTEGER,
> };
>
> +BPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)
> +{
> + int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
> +
> + tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[Severity: High]
If an unlocked socket is passed to bpf_tcp_ops_cb_flags_set(), it looks like
concurrent updates to the flags from other CPUs could be clobbered, causing
BPF TCP callbacks to fail to invoke when expected.
Does this code modify the socket flags without acquiring the socket lock
or asserting that the socket is currently owned by the caller?
> +
> + return argval & ~BPF_SOCK_OPS_ALL_CB_FLAGS;
> +}
> +
> +static const struct bpf_func_proto bpf_tcp_ops_cb_flags_set_proto = {
> + .func = bpf_tcp_ops_cb_flags_set,
> + .gpl_only = false,
> + .ret_type = RET_INTEGER,
> + .arg1_type = ARG_PTR_TO_BTF_ID,
> + .arg1_btf_id = &btf_sock_ids[BTF_SOCK_TYPE_TCP],
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[Severity: High]
Since the verifier accepts any PTR_TO_BTF_ID of type TCP here, it seems
privileged BPF programs could pass refcounted but unlocked sockets, leading
to the data race mentioned above.
Could the use of ARG_PTR_TO_BTF_ID allow a BPF program to pass an unlocked
socket obtained via helpers like bpf_sk_lookup_tcp() and bpf_skc_to_tcp_sock()?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260920195633.3033620-1-kuniyu@google.com?part=3
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
2026-09-20 19:56 ` [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
@ 2026-09-20 20:13 ` sashiko-bot
2026-09-20 21:10 ` Kuniyuki Iwashima
2026-09-20 21:16 ` bot+bpf-ci
` (2 subsequent siblings)
3 siblings, 1 reply; 28+ messages in thread
From: sashiko-bot @ 2026-09-20 20:13 UTC (permalink / raw)
To: Kuniyuki Iwashima; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: tcp: OOB access and type confusion in bpf_tcp_ops_set_rcvlowat()
--
commit c903e36ae4ed7b68c1aea9023743a57be672cd2a
Author: Kuniyuki Iwashima <kuniyu@google.com>
bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
This commit adds a new BPF kfunc bpf_tcp_ops_set_rcvlowat() to dynamically
adjust the receive low watermark for TCP sockets. It clamps negative values
to INT_MAX and configures the wakeup flag appropriately depending on the
calling hook (enqueue vs dequeue).
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index b0e14b54917eb..3768b1440eb77 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
> @@ -359,8 +359,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
> .owner = THIS_MODULE,
> };
>
> +__bpf_kfunc_start_defs();
> +
> +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> + const struct bpf_prog_aux *aux)
> +{
> + u32 moff = aux->attach_st_ops_member_off;
> + bool wakeup = false;
> +
> + if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
> + wakeup = true;
> +
> + if (rcvlowat < 0)
> + rcvlowat = INT_MAX;
> +
> + return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);
[Severity: High]
Could this lead to an out-of-bounds slab memory access if a BPF program
passes a non-TCP socket?
Because the kfunc accepts a generic struct sock *sk parameter, the BPF
verifier allows passing any socket type. For example, a trusted UDP
socket could be obtained via bpf_sk_lookup_udp() and passed directly here.
However, __tcp_set_rcvlowat() blindly casts the socket to struct tcp_sock *:
net/ipv4/tcp.c:__tcp_set_rcvlowat() {
struct tcp_sock *tp = tcp_sk(sk);
...
if (tp->window_clamp && tp->window_clamp < val)
WRITE_ONCE(tp->window_clamp, val);
}
This writes to tp->window_clamp, which resides well beyond the bounds of a
UDP socket structure.
Additionally, does this permit concurrent modification of the socket state
since the looked-up socket isn't guaranteed to be locked?
> +}
> +
> +__bpf_kfunc_end_defs();
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260920195633.3033620-1-kuniyu@google.com?part=6
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 4/7] tcp: Split out __tcp_set_rcvlowat().
2026-09-20 19:56 ` [PATCH bpf-next 4/7] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
@ 2026-09-20 21:01 ` bot+bpf-ci
2026-09-20 21:12 ` Kuniyuki Iwashima
2026-09-21 20:48 ` Stanislav Fomichev
1 sibling, 1 reply; 28+ messages in thread
From: bot+bpf-ci @ 2026-09-20 21:01 UTC (permalink / raw)
To: kuniyu, ast, daniel, andrii, martin.lau, eddyz87, memxor
Cc: yonghong.song, john.fastabend, sdf, edumazet, ncardwell, willemb,
ukyab, cleger, kuniyu, kuni1840, bpf, netdev, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 1584 bytes --]
> In the enqueue path, tcp_data_ready() is always called after
> the hooks in tcp_queue_rcv() and tcp_ofo_queue().
Does this "always" claim hold for every enqueue-side hook site?
Looking at the callers of tcp_queue_rcv() in net/ipv4/tcp_input.c:
- tcp_data_queue() -> tcp_data_ready() at line 5709,
but only if (!sock_flag(sk, SOCK_DEAD))
- tcp_rcv_established() fast path -> tcp_data_ready() at line 6694
- tcp_send_rcvq() (line 5565, TCP repair, tp->repair_queue == TCP_RECV_QUEUE)
calls tcp_queue_rcv(sk, skb, &fragstolen) at line 5601 but never calls
tcp_data_ready() or sk->sk_data_ready() at all
And there's a third enqueue hook site the message doesn't mention:
bpf_tcp_ops_enqueue_rcvq() in tcp_fastopen_add_skb()
(net/ipv4/tcp_fastopen.c:284), which also has no following tcp_data_ready().
So wakeup == false doesn't always mean "a wakeup happens later" on the repair
and TFO paths. There's no functional regression since neither tcp_send_rcvq()
nor tcp_fastopen_add_skb() signalled the reader before this series either, and
the later commit 0390d95e0979c explains the tcp_fastopen_add_skb() case.
Could the changelog be more precise about which paths it covers? Perhaps naming
tcp_data_queue()/tcp_rcv_established() specifically, and noting that
tcp_send_rcvq()/tcp_fastopen_add_skb() have no reader to wake?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35535082194
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 3/7] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops.
2026-09-20 20:11 ` sashiko-bot
@ 2026-09-20 21:05 ` Kuniyuki Iwashima
0 siblings, 0 replies; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-20 21:05 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On Sun, Sep 20, 2026 at 1:11 PM <sashiko-bot@kernel.org> wrote:
>
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf: data race in bpf_tcp_ops_cb_flags_set() on unlocked sockets
> --
>
> commit 56d7f542a8c499237807954e17f2fae35e5e276b
> Author: Kuniyuki Iwashima <kuniyu@google.com>
>
> bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops.
>
> This patch allows clearing or setting bpf_sock_ops_cb_flags via the
> bpf_sock_ops_cb_flags_set() helper within various bpf_tcp_ops callbacks.
>
> > diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> > index c68d1fa323058..6d0452441b6c5 100644
> > --- a/net/ipv4/bpf_tcp_ops.c
> > +++ b/net/ipv4/bpf_tcp_ops.c
> [ ... ]
> > @@ -220,6 +220,24 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
> > .ret_type = RET_INTEGER,
> > };
> >
> > +BPF_CALL_2(bpf_tcp_ops_cb_flags_set, struct sock *, sk, int, argval)
> > +{
> > + int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
> > +
> > + tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
> ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
>
> [Severity: High]
> If an unlocked socket is passed to bpf_tcp_ops_cb_flags_set(),
This should not happen.
The hook is only supported on hooks under sock_owned_by_me(),
and bpf_tcp_ops does not support socket lookup helpers.
> it looks like
> concurrent updates to the flags from other CPUs could be clobbered, causing
> BPF TCP callbacks to fail to invoke when expected.
>
> Does this code modify the socket flags without acquiring the socket lock
> or asserting that the socket is currently owned by the caller?
>
> > +
> > + return argval & ~BPF_SOCK_OPS_ALL_CB_FLAGS;
> > +}
> > +
> > +static const struct bpf_func_proto bpf_tcp_ops_cb_flags_set_proto = {
> > + .func = bpf_tcp_ops_cb_flags_set,
> > + .gpl_only = false,
> > + .ret_type = RET_INTEGER,
> > + .arg1_type = ARG_PTR_TO_BTF_ID,
> > + .arg1_btf_id = &btf_sock_ids[BTF_SOCK_TYPE_TCP],
> ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
>
> [Severity: High]
> Since the verifier accepts any PTR_TO_BTF_ID of type TCP here, it seems
> privileged BPF programs could pass refcounted but unlocked sockets, leading
> to the data race mentioned above.
>
> Could the use of ARG_PTR_TO_BTF_ID allow a BPF program to pass an unlocked
> socket obtained via helpers like bpf_sk_lookup_tcp() and bpf_skc_to_tcp_sock()?
Same false-positive. bpf_tcp_ops does not support such lookup helpers.
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 5/7] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG.
2026-09-20 20:05 ` sashiko-bot
@ 2026-09-20 21:09 ` Kuniyuki Iwashima
0 siblings, 0 replies; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-20 21:09 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On Sun, Sep 20, 2026 at 1:05 PM <sashiko-bot@kernel.org> wrote:
>
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Medium] bpf: tcp_set_sock_ops_cb_flags breaks ALL_CB_FLAGS for MPTCP
> --
>
> commit 578335638acc4a540df534d791edf9c63d563e32
> Author: Kuniyuki Iwashima <kuniyu@google.com>
> Subject: bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG.
>
> This patch prepares for a new kfunc by disabling BPF_SOCK_OPS_RCVQ_CB_FLAG
> for MPTCP subflows. It returns -EOPNOTSUPP if this flag is used on an MPTCP
> socket.
>
> > diff --git a/include/net/tcp.h b/include/net/tcp.h
> > index 07426e8641b7e..d3cf655da9ecd 100644
> > --- a/include/net/tcp.h
> > +++ b/include/net/tcp.h
> > @@ -2932,6 +2932,16 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
> > return tcp_call_bpf(sk, op, 3, args);
> > }
> >
> > +static inline int tcp_set_sock_ops_cb_flags(struct sock *sk, int val)
> > +{
> > + if (sk_is_mptcp(sk) &&
> > + (val & BPF_SOCK_OPS_RCVQ_CB_FLAG))
> > + return -EOPNOTSUPP;
> > +
>
> [Severity: Medium]
> Will this check unintentionally break existing BPF programs that use
> BPF_SOCK_OPS_ALL_CB_FLAGS to enable all callbacks?
This is rather a userspace problem.
BPF_SOCK_OPS_ALL_CB_FLAGS should not be used in such a way.
The flag is to be used to opt-in hooks.
>
> If a program passes BPF_SOCK_OPS_ALL_CB_FLAGS (which is expanded in this
> patch series to include the new BPF_SOCK_OPS_RCVQ_CB_FLAG bit), this rigid
> check evaluates to true for MPTCP subflows and rejects the entire multi-flag
> request instead of ignoring the unsupported bit.
>
> > + tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
> > + return 0;
> > +}
> > +
>
> [ ... ]
>
> > diff --git a/net/core/filter.c b/net/core/filter.c
> > index 5feb998846825..f29c061bb0664 100644
> > --- a/net/core/filter.c
> > +++ b/net/core/filter.c
>
> [ ... ]
>
> > @@ -6178,8 +6177,9 @@ static const struct bpf_func_proto bpf_sock_ops_getsockopt_proto = {
> > BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
> > int, argval)
> > {
> > - struct sock *sk = bpf_sock->sk;
> > int val = argval & BPF_SOCK_OPS_ALL_CB_FLAGS;
> > + struct sock *sk = bpf_sock->sk;
> > + int err;
> >
> > if (!is_locked_tcp_sock_ops(bpf_sock))
> > return -EOPNOTSUPP;
> > @@ -6187,7 +6187,9 @@ BPF_CALL_2(bpf_sock_ops_cb_flags_set, struct bpf_sock_ops_kern *, bpf_sock,
> > if (!IS_ENABLED(CONFIG_INET) || !sk_fullsock(sk))
> > return -EINVAL;
> >
> > - tcp_sk(sk)->bpf_sock_ops_cb_flags = val;
> > + err = tcp_set_sock_ops_cb_flags(sk, val);
> > + if (err)
> > + return err;
>
> [Severity: Medium]
> Does returning early here prevent any of the other valid callbacks requested
> in val from being enabled?
Same, this is intentional.
>
> Because tcp_set_sock_ops_cb_flags() returns -EOPNOTSUPP when the new flag is
> present, the BPF helper propagates the error without applying any of the
> previously supported flags. This effectively turns a wildcard request for
> "all supported flags" into a total failure on MPTCP subflows.
>
> >
> > return argval & (~BPF_SOCK_OPS_ALL_CB_FLAGS);
> > }
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260920195633.3033620-1-kuniyu@google.com?part=5
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
2026-09-20 20:13 ` sashiko-bot
@ 2026-09-20 21:10 ` Kuniyuki Iwashima
0 siblings, 0 replies; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-20 21:10 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On Sun, Sep 20, 2026 at 1:13 PM <sashiko-bot@kernel.org> wrote:
>
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] bpf: tcp: OOB access and type confusion in bpf_tcp_ops_set_rcvlowat()
> --
>
> commit c903e36ae4ed7b68c1aea9023743a57be672cd2a
> Author: Kuniyuki Iwashima <kuniyu@google.com>
>
> bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
>
> This commit adds a new BPF kfunc bpf_tcp_ops_set_rcvlowat() to dynamically
> adjust the receive low watermark for TCP sockets. It clamps negative values
> to INT_MAX and configures the wakeup flag appropriately depending on the
> calling hook (enqueue vs dequeue).
>
> > diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> > index b0e14b54917eb..3768b1440eb77 100644
> > --- a/net/ipv4/bpf_tcp_ops.c
> > +++ b/net/ipv4/bpf_tcp_ops.c
> > @@ -359,8 +359,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
> > .owner = THIS_MODULE,
> > };
> >
> > +__bpf_kfunc_start_defs();
> > +
> > +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> > + const struct bpf_prog_aux *aux)
> > +{
> > + u32 moff = aux->attach_st_ops_member_off;
> > + bool wakeup = false;
> > +
> > + if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
> > + wakeup = true;
> > +
> > + if (rcvlowat < 0)
> > + rcvlowat = INT_MAX;
> > +
> > + return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);
>
> [Severity: High]
> Could this lead to an out-of-bounds slab memory access if a BPF program
> passes a non-TCP socket?
No, false-positive.
This kfunc is supported in encode_rcvq() and decode_rcvq()
only, and we don't support lookup helpers there.
>
> Because the kfunc accepts a generic struct sock *sk parameter, the BPF
> verifier allows passing any socket type. For example, a trusted UDP
> socket could be obtained via bpf_sk_lookup_udp() and passed directly here.
>
> However, __tcp_set_rcvlowat() blindly casts the socket to struct tcp_sock *:
>
> net/ipv4/tcp.c:__tcp_set_rcvlowat() {
> struct tcp_sock *tp = tcp_sk(sk);
> ...
> if (tp->window_clamp && tp->window_clamp < val)
> WRITE_ONCE(tp->window_clamp, val);
> }
>
> This writes to tp->window_clamp, which resides well beyond the bounds of a
> UDP socket structure.
>
> Additionally, does this permit concurrent modification of the socket state
> since the looked-up socket isn't guaranteed to be locked?
>
> > +}
> > +
> > +__bpf_kfunc_end_defs();
>
> [ ... ]
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260920195633.3033620-1-kuniyu@google.com?part=6
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 4/7] tcp: Split out __tcp_set_rcvlowat().
2026-09-20 21:01 ` bot+bpf-ci
@ 2026-09-20 21:12 ` Kuniyuki Iwashima
0 siblings, 0 replies; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-20 21:12 UTC (permalink / raw)
To: bot+bpf-ci
Cc: ast, daniel, andrii, martin.lau, eddyz87, memxor, yonghong.song,
john.fastabend, sdf, edumazet, ncardwell, willemb, ukyab, cleger,
kuni1840, bpf, netdev, martin.lau, mason, ihor.solodrai
On Sun, Sep 20, 2026 at 2:01 PM <bot+bpf-ci@kernel.org> wrote:
>
> > In the enqueue path, tcp_data_ready() is always called after
> > the hooks in tcp_queue_rcv() and tcp_ofo_queue().
>
> Does this "always" claim hold for every enqueue-side hook site?
>
> Looking at the callers of tcp_queue_rcv() in net/ipv4/tcp_input.c:
>
> - tcp_data_queue() -> tcp_data_ready() at line 5709,
> but only if (!sock_flag(sk, SOCK_DEAD))
> - tcp_rcv_established() fast path -> tcp_data_ready() at line 6694
> - tcp_send_rcvq() (line 5565, TCP repair, tp->repair_queue == TCP_RECV_QUEUE)
> calls tcp_queue_rcv(sk, skb, &fragstolen) at line 5601 but never calls
> tcp_data_ready() or sk->sk_data_ready() at all
>
> And there's a third enqueue hook site the message doesn't mention:
> bpf_tcp_ops_enqueue_rcvq() in tcp_fastopen_add_skb()
> (net/ipv4/tcp_fastopen.c:284), which also has no following tcp_data_ready().
>
> So wakeup == false doesn't always mean "a wakeup happens later" on the repair
> and TFO paths.
This is mentioned in the kfunc patch. TFO path calls
sk->sk_state_change() later and wakes up the thread anyway.
> There's no functional regression since neither tcp_send_rcvq()
> nor tcp_fastopen_add_skb() signalled the reader before this series either, and
> the later commit 0390d95e0979c explains the tcp_fastopen_add_skb() case.
>
> Could the changelog be more precise about which paths it covers? Perhaps naming
> tcp_data_queue()/tcp_rcv_established() specifically, and noting that
> tcp_send_rcvq()/tcp_fastopen_add_skb() have no reader to wake?
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35535082194
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
2026-09-20 19:56 ` [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
2026-09-20 20:13 ` sashiko-bot
@ 2026-09-20 21:16 ` bot+bpf-ci
2026-09-21 20:49 ` Stanislav Fomichev
2026-09-22 7:29 ` Clément Léger
3 siblings, 0 replies; 28+ messages in thread
From: bot+bpf-ci @ 2026-09-20 21:16 UTC (permalink / raw)
To: kuniyu, ast, daniel, andrii, martin.lau, eddyz87, memxor
Cc: yonghong.song, john.fastabend, sdf, edumazet, ncardwell, willemb,
ukyab, cleger, kuniyu, kuni1840, bpf, netdev, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3144 bytes --]
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index b0e14b54917eb..3768b1440eb77 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
> @@ -359,8 +359,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
> .owner = THIS_MODULE,
> };
>
> +__bpf_kfunc_start_defs();
> +
> +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> + const struct bpf_prog_aux *aux)
> +{
> + u32 moff = aux->attach_st_ops_member_off;
> + bool wakeup = false;
> +
> + if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
> + wakeup = true;
> +
> + if (rcvlowat < 0)
> + rcvlowat = INT_MAX;
> +
> + return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);
> +}
The commit message says:
> For enqueue_rcvq(), wakeup is set to false because:
>
> * tcp_data_ready() is always called after the hooks in
> tcp_queue_rcv() and tcp_ofo_queue().
Does the "always" claim hold for every tcp_queue_rcv() call site?
Looking at tcp_send_rcvq() in net/ipv4/tcp_input.c, which handles
receive-queue injection via TCP_REPAIR:
tcp_send_rcvq() {
skb_set_owner_r(skb, sk);
tcp_queue_rcv(sk, skb, &fragstolen); /* enqueue_rcvq hook fires */
return size;
}
tcp_send_rcvq() never calls tcp_data_ready(), and its only caller at
net/ipv4/tcp.c:1205 (the TCP_REPAIR receive-queue injection path in
tcp_sendmsg_locked()) just returns after err = tcp_send_rcvq(sk, msg,
size).
So when a BPF enqueue_rcvq program lowers sk->sk_rcvlowat while data is
injected via TCP_REPAIR, __tcp_set_rcvlowat() writes the new value with
wakeup = false and nothing re-evaluates EPOLLIN. A poll()/epoll waiter
that was blocked by the higher sk_rcvlowat won't wake until the next
real packet arrives.
The non-BPF TCP_REPAIR path already issues no wakeup (tcp_send_rcvq()
itself never calls tcp_data_ready()), so this is arguably a pre-existing
gap in the repair path rather than behavior newly broken by this kfunc.
Either way, should the commit message's "always" be qualified to exclude
the repair path, or should bpf_tcp_ops_set_rcvlowat() pass wakeup = true
whenever the new value is lower than the current READ_ONCE(sk->sk_rcvlowat)
so the decision doesn't depend on assumptions about all call sites?
The other two justifications were verified:
* TFO SYN: tcp_fastopen_create_child() runs before accept(), so no
waiter exists.
* TFO SYN+ACK: tcp_rcv_synsent_state_process() in net/ipv4/tcp_input.c
calls tcp_rcv_fastopen_synack() (which reaches tcp_fastopen_add_skb())
before calling sk->sk_state_change(sk), so the wakeup does follow.
* dequeue_rcvq: none of the tcp_cleanup_rbuf() callers calls
tcp_data_ready(), so wakeup = true is required.
> +
> +__bpf_kfunc_end_defs();
> +
> +BTF_KFUNCS_START(bpf_tcp_ops_rcvlowat_kfunc_set)
> +BTF_ID_FLAGS(func, bpf_tcp_ops_set_rcvlowat, KF_IMPLICIT_ARGS)
> +BTF_KFUNCS_END(bpf_tcp_ops_rcvlowat_kfunc_set)
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35535082194
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 2/7] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-09-20 19:56 ` [PATCH bpf-next 2/7] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
@ 2026-09-20 21:16 ` bot+bpf-ci
2026-09-21 20:48 ` Stanislav Fomichev
1 sibling, 0 replies; 28+ messages in thread
From: bot+bpf-ci @ 2026-09-20 21:16 UTC (permalink / raw)
To: kuniyu, ast, daniel, andrii, martin.lau, eddyz87, memxor
Cc: yonghong.song, john.fastabend, sdf, edumazet, ncardwell, willemb,
ukyab, cleger, kuniyu, kuni1840, bpf, netdev, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 7757 bytes --]
> diff --git a/include/net/tcp.h b/include/net/tcp.h
> index d61ee00052e33..f2d838bcb0a7f 100644
> --- a/include/net/tcp.h
> +++ b/include/net/tcp.h
> @@ -3053,6 +3053,12 @@ struct bpf_tcp_ops {
> struct request_sock *req, struct sk_buff *syn_skb,
> enum tcp_synack_type synack_type,
> u32 opt_off);
> +
> + /* Called when an incoming skb is enqueued to sk->sk_receive_queue. */
> + void (*enqueue_rcvq)(struct sock *sk, struct sk_buff *skb);
> +
> + /* Called after data is dequeued from sk->sk_receive_queue. */
> + void (*dequeue_rcvq)(struct sock *sk);
> };
>
> #define bpf_tcp_ops_call(op, sk, ...) \
> @@ -3144,6 +3150,18 @@ static inline void tcp_bpf_rtt(struct sock *sk, long mrtt, u32 srtt)
> bpf_tcp_ops_call(rtt, sk, mrtt, srtt);
> }
>
> +static inline void bpf_tcp_ops_enqueue_rcvq(struct sock *sk, struct sk_buff *skb)
> +{
> + if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
> + bpf_tcp_ops_call(enqueue_rcvq, sk, skb);
> +}
> +
> +static inline void bpf_tcp_ops_dequeue_rcvq(struct sock *sk)
> +{
> + if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RCVQ_CB_FLAG))
> + bpf_tcp_ops_call(dequeue_rcvq, sk);
> +}
> +
[ ... ]
> diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
> index a4456b4194122..a714b36a74940 100644
> --- a/net/ipv4/tcp.c
> +++ b/net/ipv4/tcp.c
> @@ -1610,6 +1610,8 @@ void tcp_cleanup_rbuf(struct sock *sk, int copied)
> "cleanup rbuf bug: copied %X seq %X rcvnxt %X\n",
> tp->copied_seq, TCP_SKB_CB(skb)->end_seq, tp->rcv_nxt);
> __tcp_cleanup_rbuf(sk, copied);
> +
> + bpf_tcp_ops_dequeue_rcvq(sk);
> }
The comment at the top of the file documents dequeue_rcvq() as "Called
after data is dequeued from sk->sk_receive_queue". Is it accurate to
invoke this callback unconditionally at the tail of tcp_cleanup_rbuf()?
Looking at the callers, tcp_cleanup_rbuf(sk, copied) is invoked:
1. From __tcp_sock_set_quickack():
net/ipv4/tcp.c:__tcp_sock_set_quickack() {
...
tcp_cleanup_rbuf(sk, 1);
}
This call forces an ACK from setsockopt(TCP_QUICKACK) or TCP_NODELAY
without touching the receive queue.
2. From tcp_recvmsg_locked() with copied == 0:
net/ipv4/tcp.c:tcp_recvmsg_locked() {
...
tcp_cleanup_rbuf(sk, copied);
...
sk_wait_data(sk, &timeo, last);
}
This happens before sk_wait_data() when no data has been dequeued yet.
3. From tcp_recvmsg_locked() with MSG_PEEK:
The final tcp_cleanup_rbuf(sk, copied) runs with copied > 0, but the
skbs remain queued and copied_seq is unchanged because MSG_PEEK leaves
the data in the queue.
With the kfunc added later in the series, dequeue_rcvq() can call
__tcp_set_rcvlowat(..., wakeup=true), which triggers tcp_data_ready()
-> sk->sk_data_ready(). If a BPF program lowers sk_rcvlowat from the
dequeue_rcvq() callback, would these no-dequeue paths cause spurious
EPOLLIN wakeups that the feature is meant to suppress?
Should the hook be gated on an actual dequeue (copied > 0 and not
MSG_PEEK), or should the documented contract be adjusted?
[ ... ]
> diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
> index 6ac6f9d5b6c38..c60c61bb0a715 100644
> --- a/net/ipv4/tcp_input.c
> +++ b/net/ipv4/tcp_input.c
> @@ -5344,6 +5344,8 @@ static void tcp_ofo_queue(struct sock *sk)
> continue;
> }
>
> + bpf_tcp_ops_enqueue_rcvq(sk, skb);
> +
> tail = skb_peek_tail(&sk->sk_receive_queue);
> eaten = tail && tcp_try_coalesce(sk, tail, skb, &fragstolen);
> tcp_rcv_nxt_update(tp, TCP_SKB_CB(skb)->end_seq);
> @@ -5547,6 +5549,8 @@ static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb,
> int eaten;
> struct sk_buff *tail = skb_peek_tail(&sk->sk_receive_queue);
>
> + bpf_tcp_ops_enqueue_rcvq(sk, skb);
> +
> eaten = (tail &&
> tcp_try_coalesce(sk, tail,
> skb, fragstolen)) ? 1 : 0;
In tcp_queue_rcv(), bpf_tcp_ops_enqueue_rcvq(sk, skb) is invoked after
tail has been cached from sk->sk_receive_queue but before tail is
dereferenced by tcp_try_coalesce(). Can the BPF callback synchronously
drain sk->sk_receive_queue, leaving tail pointing at freed memory?
Concrete path (verified in this commit):
1. net/ipv4/bpf_tcp_ops.c only blocks bpf_setsockopt() for
rwnd_init/timeout_init/hdr_opt_len/write_hdr_opt:
net/ipv4/bpf_tcp_ops.c:bpf_tcp_ops_get_func_proto() {
case BPF_FUNC_setsockopt:
if (moff == offsetof(struct bpf_tcp_ops, rwnd_init) || ... write_hdr_opt))
return NULL;
return &bpf_sk_setsockopt_proto;
}
enqueue_rcvq is not in the blocked list, so a prog attached to
.enqueue_rcvq may call bpf_setsockopt(). The commit message states the
intended use is to adjust sk->sk_rcvlowat, so
bpf_setsockopt(SOL_SOCKET, SO_RCVLOWAT, ...) is permitted.
2. bpf_sk_setsockopt -> _bpf_setsockopt -> __bpf_setsockopt ->
sol_socket_sockopt() accepts SO_RCVLOWAT and calls sk_setsockopt().
3. sk_setsockopt() SO_RCVLOWAT calls tcp_set_rcvlowat() for a TCP
socket with a non-NULL sk->sk_socket.
4. tcp_set_rcvlowat() calls tcp_data_ready(sk):
net/ipv4/tcp.c:tcp_set_rcvlowat() {
...
tcp_data_ready(sk);
}
net/ipv4/tcp_input.c:tcp_data_ready() {
if (tcp_epollin_ready(sk, sk->sk_rcvlowat) || sock_flag(sk, SOCK_DONE))
READ_ONCE(sk->sk_data_ready)(sk);
}
Since tail != NULL means unread bytes are queued, rcv_nxt - copied_seq >
0, and the prog chooses the new (small) rcvlowat, tcp_epollin_ready()
returns true.
5. For a socket whose sk_data_ready drains the receive queue:
SOCKMAP verdict mode: sk_psock_verdict_data_ready() ->
tcp_read_skb() unlinks and consumes every queued skb:
net/core/skmsg.c:sk_psock_verdict_data_ready() {
...
ops->read_skb(sk, sk_psock_verdict_recv);
}
net/ipv4/tcp.c:tcp_read_skb() {
while ((skb = skb_peek(&sk->sk_receive_queue)) != NULL) {
...
__skb_unlink(skb, &sk->sk_receive_queue);
...
if (used < 0 || !skb->len) {
tcp_eat_recv_skb(sk, skb);
...
}
...
}
}
There is no pause, ownership guard, or re-entrancy protection.
SOCKMAP strparser/kTLS: sk_psock_strp_data_ready() ->
strp_data_ready() -> strp_read_sock() -> tcp_bpf_strp_read_sock() ->
tcp_eat_recv_skb() -> __skb_unlink() + __kfree_skb()/
skb_attempt_defer_free(). strp_data_ready() only defers to a workqueue
when the socket is owned by a user; in the softirq receive path it runs
inline.
6. Control returns to tcp_queue_rcv(), which then evaluates:
eaten = (tail && tcp_try_coalesce(sk, tail, skb, fragstolen)) ? 1 : 0;
tcp_try_coalesce() immediately dereferences the stale pointer:
net/ipv4/tcp_input.c:tcp_try_coalesce() {
if (TCP_SKB_CB(to)->end_seq != TCP_SKB_CB(from)->seq)
return false;
...
}
This causes a use-after-free if the skb was freed, or silent data
duplication/loss plus sk_rmem_alloc mis-accounting if the skb was
unlinked but not yet freed (handed to the psock ingress queue).
Note that the sibling hook in tcp_ofo_queue() is not affected because
tail is read *after* bpf_tcp_ops_enqueue_rcvq(). Would reading tail
after the callback in tcp_queue_rcv() as well close this window? The skb
being queued is still owned by the caller and is not on any queue, so it
cannot be consumed by the nested drain.
Also note that bpf_sk_setsockopt_nodelay_proto was introduced
specifically to block TCP_NODELAY because "TCP_NODELAY triggers
tcp_push_pending_frames() and re-enters CA_EVENT_TX_START in
bpf_tcp_cc", but bpf_tcp_ops uses the unfiltered
bpf_sk_setsockopt_proto. Is this difference intentional?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35535082194
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 7/7] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-09-20 19:56 ` [PATCH bpf-next 7/7] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
@ 2026-09-20 21:16 ` bot+bpf-ci
2026-09-21 20:50 ` Stanislav Fomichev
2026-09-22 23:14 ` Kuniyuki Iwashima
2 siblings, 0 replies; 28+ messages in thread
From: bot+bpf-ci @ 2026-09-20 21:16 UTC (permalink / raw)
To: kuniyu, ast, daniel, andrii, martin.lau, eddyz87, memxor
Cc: yonghong.song, john.fastabend, sdf, edumazet, ncardwell, willemb,
ukyab, cleger, kuniyu, kuni1840, bpf, netdev, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 8273 bytes --]
> diff --git a/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
> new file mode 100644
> index 0000000000000..c5ead4af24f85
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
> @@ -0,0 +1,350 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/* Copyright 2026 Google LLC */
> +#include <sys/epoll.h>
Does this file need to include <sched.h> for unshare() and CLONE_NEWNET?
The calls are at tcp_autolowat_setup_cb() but there's no direct include.
Every other bpf prog_test that unshares a netns includes <sched.h>
explicitly (bind_perm.c, setget_sockopt.c, sock_fields.c, tcp_hdr_options.c,
tcp_custom_syncookie.c). It compiles today only because test_progs.h pulls
in <pthread.h>, which brings <sched.h> transitively under _GNU_SOURCE.
> +
> +#include "test_progs.h"
> +#include "cgroup_helpers.h"
> +#include "network_helpers.h"
> +
> +#include "tcp_autolowat.skel.h"
[ ... ]
> +struct rpc_test_case {
> + char data[4096];
> + struct rpc_descriptor desc[32];
> + struct rpc_event event[32];
> +} rpc_test_cases[] = {
Should rpc_test_cases[] be static? It's currently defined with external
linkage, and test_progs links every prog_tests/*.o together. Every
comparable table in neighboring tests is static (test_cases[] in
ctx_rewrite.c, sk_bypass_prot_mem.c, crafted_cases[] in reg_bounds.c,
bond_test_cases[] in xdp_bonding.c). All other file-scope objects in this
file are already static.
> + {
> + .desc = {
> + { .header_len = 100, .payload_len = 150 },
> + },
> + .event = {
> + { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
> + /* Single full RPC message in skb. */
> + { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE + 100 + 150},
> + { .type = RPC_EVENT_EPOLL, .nfds = 1},
> + { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 100 + 150},
> + },
> + },
> + {
> + .desc = {
> + {.header_len = 100, .payload_len = 150},
> + {.header_len = 100, .payload_len = 150},
> + {.header_len = 100, .payload_len = 150},
> + },
> + .event = {
> + { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
> + /* Two full RPC messages in skb. */
> + {.type = RPC_EVENT_SEND, .len = (RPC_DESC_SIZE + 100 + 150) * 2},
> + {.type = RPC_EVENT_EPOLL, .nfds = 1},
> + {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
> + /* Single full RPC message in skb. */
> + { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE + 100 + 150},
> + { .type = RPC_EVENT_EPOLL, .nfds = 1},
> + { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 3},
> + },
> + },
> + {
> + .desc = {
> + {.header_len = 100, .payload_len = 150},
> + {.header_len = 100, .payload_len = 150},
> + {.header_len = 100, .payload_len = 150},
> + },
> + .event = {
> + { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
> + /* Two full RPC messages in skb. */
> + {.type = RPC_EVENT_SEND, .len = (RPC_DESC_SIZE + 100 + 150) * 2},
> + {.type = RPC_EVENT_EPOLL, .nfds = 1},
> + {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
> + /* Single full RPC message in skb. */
> + { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE},
Does this comment match the event? The comment says "Single full RPC
message in skb" but the send is only RPC_DESC_SIZE (8 bytes) without the
250-byte header+payload. This is what distinguishes this case from
rpc_test_cases[1] where the identical comment sits above
.len = RPC_DESC_SIZE + 100 + 150. The expected rcvlowat that follows
(unchanged at (RPC_DESC_SIZE + 100 + 150) * 2) confirms the intent: only
the next descriptor arrives, so SO_RCVLOWAT stays at the still-unread
complete frames. Since the comment is the only thing explaining this case's
purpose, does it need to say "Only the next descriptor in skb" instead?
> + { .type = RPC_EVENT_EPOLL, .nfds = 1},
> + { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
> + },
> + },
[ ... ]
> +static int tcp_autolowat_setup_cb(struct tcp_autolowat_test_cb *cb, int family)
> +{
> + struct epoll_event ev = {};
> + int err;
> + int i;
> +
> + for (i = 0; i < ARRAY_SIZE(cb->fd); i++)
> + cb->fd[i] = -1;
> +
> + cb->saved_netns = open("/proc/self/ns/net", O_RDONLY);
> + if (!ASSERT_OK_FD(cb->saved_netns, "save netns"))
> + goto err;
> +
> + err = unshare(CLONE_NEWNET);
> + if (!ASSERT_OK(err, "unshare"))
> + goto err;
> +
> + err = system("ip link set dev lo up");
> + if (!ASSERT_OK(err, "set up lo"))
> + goto err;
A subsystem pattern flags this as potentially concerning: network-namespace
setup is open-coded here with open("/proc/self/ns/net") + unshare() +
system("ip link ..."), and tcp_autolowat_teardown_cb() does the matching
setns()+close(). The BPF selftest tree has a SYS() wrapper in
test_progs.h that is exactly ASSERT_OK(system(cmd), cmd) plus a goto
label, so the system() call could use that. Nine other bpf prog_tests
already do the unshare(CLONE_NEWNET)+saved-fd idiom the same way this does
(bind_perm.c, bpf_iter_setsockopt.c, flow_dissector_reattach.c,
setget_sockopt.c, sk_lookup.c, sock_fields.c, sk_assign.c,
tcp_custom_syncookie.c, tcp_hdr_options.c), and network_helpers'
open_netns()/close_netns() operate on a named netns (ip netns add), not an
anonymous unshare. Is the current pattern acceptable here, or would using
SYS() for the ip command be an improvement?
[ ... ]
> diff --git a/tools/testing/selftests/bpf/progs/tcp_autolowat.c b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
> new file mode 100644
> index 0000000000000..bb96e19e7589d
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
> @@ -0,0 +1,312 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/* Copyright 2026 Google LLC */
> +#include "vmlinux.h"
> +
> +#include <string.h>
> +#include <limits.h>
> +#include <bpf/bpf_helpers.h>
> +#include <bpf/bpf_tracing.h>
> +#include <bpf/bpf_core_read.h>
> +
> +#include "bpf_kfuncs.h"
> +#include "bpf_tracing_net.h"
[ ... ]
> +static void tcp_set_autolowat(struct tcp_autolowat_cb *cb,
> + struct sock *sk)
> +{
> + struct tcp_sock *tp = (struct tcp_sock *)sk;
> + u32 val; /* To handle wraparound. */
> +
> + LOG("Setting rcvlowat: tp->copied_seq: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_desc_buff_len: %u",
> + TP_SEQ(copied_seq), CB_SEQ(rpc_desc_seq),
> + CB_SEQ(rpc_end_seq), cb->rpc_desc_buff_len);
> +
> + if (before(tp->copied_seq, cb->rpc_desc_seq))
> + val = cb->rpc_desc_seq - tp->copied_seq;
> + else if (cb->rpc_desc_buff_len != RPC_DESC_SIZE)
> + val = RPC_DESC_SIZE;
> + else
> + val = cb->rpc_end_seq - tp->copied_seq;
> +
> + if (val != tp->inet_conn.icsk_inet.sk.sk_rcvlowat) {
> + bpf_tcp_ops_set_rcvlowat(sk, val);
Does this compile? The call to bpf_tcp_ops_set_rcvlowat() has no
declaration anywhere in the translation unit. Checking:
- vmlinux.h is generated by bpftool btf dump, which never emits function
prototypes (libbpf's BTF dumper marks BTF_KIND_FUNC as ORDERED and emits
nothing for it).
- bpf_helpers.h/bpf_helper_defs.h only declare BPF helpers, not kfuncs.
- bpf_kfuncs.h declares each kfunc explicitly with __ksym, but has no
entry for bpf_tcp_ops_set_rcvlowat.
- Grep over tools/ finds the identifier only at the two call sites in this
file (here and tcp_disable_autolowat()).
All of progs/*.c are built with -Wall -Werror -std=gnu11 (Makefile:463-472),
and a call to an undeclared function in C99+/gnu11 is an error. Even if the
diagnostic were suppressed, an implicitly declared function carries no
__ksym attribute, so libbpf would not emit a kfunc relocation and the
program could not be loaded. The call itself looks correct: the kernel
kfunc is int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
const struct bpf_prog_aux *aux) registered with KF_IMPLICIT_ARGS, and
resolve_btfids drops the implicit argument from the BPF-visible signature.
Doesn't this just need the declaration, like:
extern int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat) __ksym;
in this file or in bpf_kfuncs.h?
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35535082194
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 1/7] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv.
2026-09-20 19:56 ` [PATCH bpf-next 1/7] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv Kuniyuki Iwashima
@ 2026-09-21 20:48 ` Stanislav Fomichev
0 siblings, 0 replies; 28+ messages in thread
From: Stanislav Fomichev @ 2026-09-21 20:48 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 09/20, Kuniyuki Iwashima wrote:
> Once bpf_sock_ops_cb_flags_set() supports a new flag,
> tcpbpf_user.c fails due to the hard-coded max value, 0x80.
>
> Let's replace 0x80 with BPF_SOCK_OPS_ALL_CB_FLAGS + 1.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 2/7] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-09-20 19:56 ` [PATCH bpf-next 2/7] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-09-20 21:16 ` bot+bpf-ci
@ 2026-09-21 20:48 ` Stanislav Fomichev
1 sibling, 0 replies; 28+ messages in thread
From: Stanislav Fomichev @ 2026-09-21 20:48 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 09/20, Kuniyuki Iwashima wrote:
> Waking up a thread per packet is expensive when an application
> processes variable-length frames (e.g., RPC) that span multiple
> packets.
>
> SO_RCVLOWAT can defer wakeups, but because the frame size is
> encoded in a fixed-size descriptor at the start of each frame,
> the application has to:
>
> 1. wake up and recv() the descriptor,
> 2. raise SO_RCVLOWAT to the payload size via setsockopt(),
> 3. wake up and recv() the payload, and
> 4. reset SO_RCVLOWAT back to the descriptor size via
> setsockopt() for the next frame.
>
> This requires an extra wakeup and two setsockopt() syscalls
> for every single RPC frame.
>
> With SOCKMAP, we can parse skb and suppress wakeups in kernel,
> but SOCKMAP adds overhead and also kills zerocopy.
>
> Let's add lighter-weight opt-in callbacks to bpf_tcp_ops to
> replace that.
>
> .enqueue_rcvq(): invoked when TCP stack enqueues skb to
> sk->sk_receive_queue
>
> .dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
> is dequeued from sk->sk_receive_queue
>
> Those callbacks can be enabled on a per-socket basis by
> bpf_setsockopt():
>
> int flags = BPF_SOCK_OPS_RCVQ_CB_FLAG;
>
> bpf_setsockopt(sk, SOL_TCP, TCP_BPF_SOCK_OPS_CB_FLAGS,
> &flags, sizeof(flags));
>
> or via the bpf_tcp_ops-specific helper added in the next patch:
>
> bpf_sock_ops_cb_flags_set(sk, BPF_SOCK_OPS_RCVQ_CB_FLAG);
>
> Later, we will add a new kfunc to adjust sk->sk_rcvlowat from
> these callbacks.
>
> This will allow the bpf_tcp_ops prog to parse each skb and
> dynamically adjust sk->sk_rcvlowat to suppress unnecessary EPOLLIN
> wakeups until sufficient data is available in the receive queue.
>
> The placement of bpf_tcp_ops_call() in tcp_ofo_queue() and
> tcp_fastopen_add_skb() is chosen to provide the same snapshot
> as tcp_queue_rcv().
>
> For example, if bpf_tcp_ops_call() were called before updating
> TCP_SKB_CB(skb)->seq in tcp_fastopen_add_skb(), BPF prog would
> need an extra branch for the unlikely TFO case to strip SYN.
>
> In addition, the TCP stack can queue overlapping skbs into recvq.
> Once rcv_nxt is updated with a new skb, BPF prog can no longer
> infer the previous rcv_nxt from skb->len.
>
> Lastly, dequeue_rcvq() is placed in tcp_cleanup_rbuf() rather
> than __tcp_cleanup_rbuf() so that it is not called for sockets
> in SOCKMAP, where calling sk->sk_data_ready() from the new
> kfunc would otherwise trigger infinite recursion.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 3/7] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops.
2026-09-20 19:56 ` [PATCH bpf-next 3/7] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops Kuniyuki Iwashima
2026-09-20 20:11 ` sashiko-bot
@ 2026-09-21 20:48 ` Stanislav Fomichev
1 sibling, 0 replies; 28+ messages in thread
From: Stanislav Fomichev @ 2026-09-21 20:48 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 09/20, Kuniyuki Iwashima wrote:
> When an error occurs in bpf_tcp_ops.{enqueue,dequeue}_rcvq(),
> we want to clear BPF_SOCK_OPS_RCVQ_CB_FLAG to stop invoking
> the callbacks.
>
> In addition, bpf_sock_ops_cb_flags_set() is often used during
> setup or just after 3WHS completes to enable opt-in hooks.
>
> Let's support bpf_sock_ops_cb_flags_set() in the following
> callbacks: connect, listen, {active,passive}_established,
> {enqueue,dequeue}_rcvq.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 4/7] tcp: Split out __tcp_set_rcvlowat().
2026-09-20 19:56 ` [PATCH bpf-next 4/7] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
2026-09-20 21:01 ` bot+bpf-ci
@ 2026-09-21 20:48 ` Stanislav Fomichev
1 sibling, 0 replies; 28+ messages in thread
From: Stanislav Fomichev @ 2026-09-21 20:48 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 09/20, Kuniyuki Iwashima wrote:
> We will add a kfunc for bpf_tcp_ops.{enqueue,dequeue}_rcvq()
> to adjust sk->sk_rcvlowat.
>
> These hooks are triggered
>
> * when the TCP stack enqueues an skb to sk->sk_receive_queue
> * after data is dequeued from sk->sk_receive_queue
>
> In the enqueue path, tcp_data_ready() is always called after
> the hooks in tcp_queue_rcv() and tcp_ofo_queue().
>
> If tcp_set_rcvlowat() were used as is, tcp_data_ready() could
> be called twice for the same skb, which is redundant and also
> confusing.
>
> Let's split out __tcp_set_rcvlowat() and add a flag to control
> wakeup behaviour.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
2026-09-20 19:56 ` [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
2026-09-20 20:13 ` sashiko-bot
2026-09-20 21:16 ` bot+bpf-ci
@ 2026-09-21 20:49 ` Stanislav Fomichev
2026-09-21 21:36 ` Kuniyuki Iwashima
2026-09-22 7:29 ` Clément Léger
3 siblings, 1 reply; 28+ messages in thread
From: Stanislav Fomichev @ 2026-09-21 20:49 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 09/20, Kuniyuki Iwashima wrote:
> bpf_tcp_ops.{enqueue,dequeue}_rcvq() were added to parse skb and
> adjust sk->sk_rcvlowat dynamically to suppress unnecessary wakeups.
>
> Let's add a new kfunc to set sk->sk_rcvlowat.
>
> Negative values are clamped to INT_MAX, consistent with SO_RCVLOWAT.
>
> For enqueue_rcvq(), wakeup is set to false because:
>
> * tcp_data_ready() is always called after the hooks in
> tcp_queue_rcv() and tcp_ofo_queue().
>
> * when tcp_fastopen_add_skb() is called for TFO SYN, the socket is
> not yet accept()ed, and when called for TFO SYN+ACK, the socket
> is woken up by sk->sk_state_change() anyway.
>
> For dequeue_rcvq(), wakeup is set to true because tcp_data_ready()
> is not called in that path.
>
> An alternative would be to support bpf_setsockopt() for these
> hooks.
>
> However, that approach involves excessive conditionals and an
> unnecessary memcpy(), costs we do not want to pay for every skb
> in the TCP fast path.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> ---
> net/ipv4/bpf_tcp_ops.c | 56 +++++++++++++++++++++++++++++++++++++++++-
> 1 file changed, 55 insertions(+), 1 deletion(-)
>
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index b0e14b54917e..3768b1440eb7 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
> @@ -359,8 +359,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
> .owner = THIS_MODULE,
> };
>
> +__bpf_kfunc_start_defs();
> +
> +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> + const struct bpf_prog_aux *aux)
> +{
> + u32 moff = aux->attach_st_ops_member_off;
> + bool wakeup = false;
> +
> + if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
> + wakeup = true;
[..]
> + if (rcvlowat < 0)
> + rcvlowat = INT_MAX;
nit: the same check exists in setsockopt. Any reason not to move it
to your new __tcp_set_rcvlowat ?
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 7/7] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-09-20 19:56 ` [PATCH bpf-next 7/7] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-09-20 21:16 ` bot+bpf-ci
@ 2026-09-21 20:50 ` Stanislav Fomichev
2026-09-22 23:14 ` Kuniyuki Iwashima
2 siblings, 0 replies; 28+ messages in thread
From: Stanislav Fomichev @ 2026-09-21 20:50 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 09/20, Kuniyuki Iwashima wrote:
> The test is roughly divided into two stages, and the sequence
> is as follows:
>
> I) Setup
>
> 1. Attach two BPF programs to a cgroup
> 2. Establish a TCP connection (@client <-> @child) within the cgroup
> 3. Enable BPF_SOCK_OPS_RCVQ_CB_FLAG on @child via setsockopt()
>
> II) RPC frame exchange in various patterns
>
> 4. Send a partial RPC descriptor from @client to @child
> 5. Verify that epoll does NOT wake up @child
> 6. Send the remaining data of the RPC frame
> 7. Verify that epoll finally wakes up @child
>
> During setup, two BPF programs are attached to simulate
> a real-world scenario; one is bpf_tcp_ops and the other is
> CGROUP_SOCKOPT.
>
> While the bpf_tcp_ops prog handles the dynamic adjustment of
> sk->sk_rcvlowat, the CGROUP_SOCKOPT prog is used to enable
> the TCP AutoLOWAT feature via userspace setsockopt() using
> pseudo options:
>
> #define SOL_BPF 0xdeadbeef
> #define BPF_TCP_AUTOLOWAT 0x8badf00d
>
> setsockopt(fd, SOL_BPF, BPF_TCP_AUTOLOWAT, &(int){1}, sizeof(int));
>
> This reflects a common production use case where an application
> decides to start parsing RPC frames only at a certain point in
> the stream (e.g., after HTTP Upgrade), rather than immediately
> after TCP 3WHS (BPF_SOCK_OPS_PASSIVE_ESTABLISHED_CB, etc).
>
> When BPF_TCP_AUTOLOWAT is enabled, the BPF prog sets
> BPF_SOCK_OPS_RCVQ_CB_FLAG and initializes sk_local_storage
> for two sequence numbers to manage its state.
>
> Then, for the RPC frame exchange, this test uses a simple format
> defined as follows:
>
> 0 8 16 24 32
> +--------+--------+-------+--------+ `.
> | header size | |
> +--------+--------+-------+--------+ > RPC descriptor (8 bytes)
> | payload size | |
> +--------+--------+-------+--------+ .'
> ~ header ~
> +--------+--------+-------+--------+
> ~ payload ~
> +--------+--------+-------+--------+
>
> Every time a new skb is enqueued to sk->sk_receive_queue, the
> bpf_tcp_ops prog parses it and updates these sequence numbers:
>
> rpc_desc_seq : the SEQ # of the start of the RPC descriptor
> rpc_end_seq : the SEQ # of the end of the RPC frame
> => rpc_desc_seq + 8 + header size + payload size
>
> Assume we receive two RPC descriptors in the following pattern:
>
> 1. When we receive skb-1, only part of the RPC descriptor is parsed.
> rpc_desc_seq is set to the first byte while rpc_end_seq is
> unknown. Thus, sk->sk_rcvlowat is set to the size of the RPC
> descriptor (8 bytes).
>
> <- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
> +-----------+.................+....................+......
> | RPC desc 1 | header + payload | RPC desc 2 | ...
> +-----------+.................+....................+......
> ^ ^-.
> `- rpc_desc_seq `- sk->sk_rcvlowat
>
> 2. Next, we receive skb-2, which completes the first RPC descriptor.
> Now rpc_end_seq is known, so sk->sk_rcvlowat is advanced to it.
>
> <- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
> +-----------+-----------------+....................+......
> | RPC desc 1 | header + payload | RPC desc 2 | ...
> +-----------+-----------------+....................+......
> ^ ^
> '- rpc_desc_seq '- rpc_end_seq
> & sk->sk_rcvlowat
>
> 3. Once we receive skb-3, which contains the next full RPC descriptor,
> rpc_desc_seq is advanced and rpc_end_seq is updated according
> to the size of RPC frame 2.
>
> Note that sk->sk_rcvlowat is NOT updated to the new rpc_end_seq
> yet. This ensures that the application is woken up to read the
> already complete RPC frame 1.
>
> <- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
> +-----------+-----------------+--------------------+......
> | RPC desc 1 | header + payload | RPC desc 2 | ... |
> +-----------+-----------------+--------------------+......
> ^ ^
> rpc_desc_seq -----------' rpc_end_seq ----...-'
> & sk->sk_rcvlowat
>
> This sequence corresponds to the 4th test case in rpc_test_cases[],
> and we can see helpful output if we "#define DEBUG":
>
> # cat /sys/kernel/tracing/trace_pipe | \
> awk '{ if ($0 ~ /AF_/) sub(/^.*AF_/, "AF_"); print $0 }' & \
> BGPID=$!; ./test_progs -t tcp_autolowat; kill -9 -$BGPID
> ...
> AF_INET6 rpc_test_cases[3]: Start parsing skb: seq: 0, end_seq: 1, len: 1, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_buff_len: 0
> AF_INET6 rpc_test_cases[3]: Copied 1 bytes: rpc_desc_buff_len: 1
> AF_INET6 rpc_test_cases[3]: Setting rcvlowat: tp->copied_seq: 0, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_desc_buff_len: 1
> AF_INET6 rpc_test_cases[3]: Set rcvlowat: expected: 8, actual: 8
>
> AF_INET6 rpc_test_cases[3]: Start parsing skb: seq: 1, end_seq: 8, len: 7, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_buff_len: 1
> AF_INET6 rpc_test_cases[3]: Copied full descriptor: rpc_desc_seq: 0, rpc_end_seq: 258, header_len: 100, payload_len: 150
> AF_INET6 rpc_test_cases[3]: No more descriptor: rpc_end_seq: 258, end_seq: 8
> AF_INET6 rpc_test_cases[3]: Setting rcvlowat: tp->copied_seq: 0, rpc_desc_seq: 0, rpc_end_seq: 258, rpc_desc_buff_len: 8
> AF_INET6 rpc_test_cases[3]: Set rcvlowat: expected: 258, actual: 258
> ...
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
2026-09-21 20:49 ` Stanislav Fomichev
@ 2026-09-21 21:36 ` Kuniyuki Iwashima
0 siblings, 0 replies; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-21 21:36 UTC (permalink / raw)
To: Stanislav Fomichev
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On Mon, Sep 21, 2026 at 1:50 PM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
>
> On 09/20, Kuniyuki Iwashima wrote:
> > bpf_tcp_ops.{enqueue,dequeue}_rcvq() were added to parse skb and
> > adjust sk->sk_rcvlowat dynamically to suppress unnecessary wakeups.
> >
> > Let's add a new kfunc to set sk->sk_rcvlowat.
> >
> > Negative values are clamped to INT_MAX, consistent with SO_RCVLOWAT.
> >
> > For enqueue_rcvq(), wakeup is set to false because:
> >
> > * tcp_data_ready() is always called after the hooks in
> > tcp_queue_rcv() and tcp_ofo_queue().
> >
> > * when tcp_fastopen_add_skb() is called for TFO SYN, the socket is
> > not yet accept()ed, and when called for TFO SYN+ACK, the socket
> > is woken up by sk->sk_state_change() anyway.
> >
> > For dequeue_rcvq(), wakeup is set to true because tcp_data_ready()
> > is not called in that path.
> >
> > An alternative would be to support bpf_setsockopt() for these
> > hooks.
> >
> > However, that approach involves excessive conditionals and an
> > unnecessary memcpy(), costs we do not want to pay for every skb
> > in the TCP fast path.
> >
> > Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> > ---
> > net/ipv4/bpf_tcp_ops.c | 56 +++++++++++++++++++++++++++++++++++++++++-
> > 1 file changed, 55 insertions(+), 1 deletion(-)
> >
> > diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> > index b0e14b54917e..3768b1440eb7 100644
> > --- a/net/ipv4/bpf_tcp_ops.c
> > +++ b/net/ipv4/bpf_tcp_ops.c
> > @@ -359,8 +359,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
> > .owner = THIS_MODULE,
> > };
> >
> > +__bpf_kfunc_start_defs();
> > +
> > +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> > + const struct bpf_prog_aux *aux)
> > +{
> > + u32 moff = aux->attach_st_ops_member_off;
> > + bool wakeup = false;
> > +
> > + if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
> > + wakeup = true;
>
> [..]
>
> > + if (rcvlowat < 0)
> > + rcvlowat = INT_MAX;
>
> nit: the same check exists in setsockopt. Any reason not to move it
> to your new __tcp_set_rcvlowat ?
Maybe we could duplicate it for each protocol, but I feel the logic should
reside in the core part.
$ chgrep -E "\.set_rcvlowat.*?="
net/vmw_vsock/af_vsock.c:2672: .set_rcvlowat = vsock_set_rcvlowat,
net/mptcp/protocol.c:4730: .set_rcvlowat = mptcp_set_rcvlowat,
net/ipv6/af_inet6.c:691: .set_rcvlowat = tcp_set_rcvlowat,
>
> Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Thanks !
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
2026-09-20 19:56 ` [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
` (2 preceding siblings ...)
2026-09-21 20:49 ` Stanislav Fomichev
@ 2026-09-22 7:29 ` Clément Léger
3 siblings, 0 replies; 28+ messages in thread
From: Clément Léger @ 2026-09-22 7:29 UTC (permalink / raw)
To: Kuniyuki Iwashima, Alexei Starovoitov, Daniel Borkmann,
Andrii Nakryiko, Martin KaFai Lau, Eduard Zingerman,
Kumar Kartikeya Dwivedi
Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab, Kuniyuki Iwashima,
bpf, netdev
On 9/20/26 21:56, Kuniyuki Iwashima wrote:
> >
> bpf_tcp_ops.{enqueue,dequeue}_rcvq() were added to parse skb and
> adjust sk->sk_rcvlowat dynamically to suppress unnecessary wakeups.
>
> Let's add a new kfunc to set sk->sk_rcvlowat.
>
> Negative values are clamped to INT_MAX, consistent with SO_RCVLOWAT.
>
> For enqueue_rcvq(), wakeup is set to false because:
>
> * tcp_data_ready() is always called after the hooks in
> tcp_queue_rcv() and tcp_ofo_queue().
>
> * when tcp_fastopen_add_skb() is called for TFO SYN, the socket is
> not yet accept()ed, and when called for TFO SYN+ACK, the socket
> is woken up by sk->sk_state_change() anyway.
>
> For dequeue_rcvq(), wakeup is set to true because tcp_data_ready()
> is not called in that path.
>
> An alternative would be to support bpf_setsockopt() for these
> hooks.
>
> However, that approach involves excessive conditionals and an
> unnecessary memcpy(), costs we do not want to pay for every skb
> in the TCP fast path.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> ---
> net/ipv4/bpf_tcp_ops.c | 56 +++++++++++++++++++++++++++++++++++++++++-
> 1 file changed, 55 insertions(+), 1 deletion(-)
>
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index b0e14b54917e..3768b1440eb7 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
> @@ -359,8 +359,62 @@ static struct bpf_struct_ops bpf_tcp_ops = {
> .owner = THIS_MODULE,
> };
>
> +__bpf_kfunc_start_defs();
> +
> +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> + const struct bpf_prog_aux *aux)
> +{
> + u32 moff = aux->attach_st_ops_member_off;
> + bool wakeup = false;
> +
> + if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
> + wakeup = true;
> +
> + if (rcvlowat < 0)
> + rcvlowat = INT_MAX;
> +
> + return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);
> +}
> +
> +__bpf_kfunc_end_defs();
> +
> +BTF_KFUNCS_START(bpf_tcp_ops_rcvlowat_kfunc_set)
> +BTF_ID_FLAGS(func, bpf_tcp_ops_set_rcvlowat, KF_IMPLICIT_ARGS)
> +BTF_KFUNCS_END(bpf_tcp_ops_rcvlowat_kfunc_set)
> +
> +static int bpf_tcp_ops_rcvlowat_kfunc_filter(const struct bpf_prog *prog,
> + u32 kfunc_id)
> +{
> + u32 moff;
> +
> + if (!btf_id_set8_contains(&bpf_tcp_ops_rcvlowat_kfunc_set, kfunc_id))
> + return 0;
> +
> + if (prog->aux->st_ops != &bpf_tcp_ops)
> + return -EACCES;
> +
> + moff = prog->aux->attach_st_ops_member_off;
> + if (moff != offsetof(struct bpf_tcp_ops, enqueue_rcvq) &&
> + moff != offsetof(struct bpf_tcp_ops, dequeue_rcvq))
> + return -EACCES;
> +
> + return 0;
> +}
> +
> +static const struct btf_kfunc_id_set bpf_tcp_ops_rcvlowat_kfunc_id_set = {
> + .owner = THIS_MODULE,
> + .set = &bpf_tcp_ops_rcvlowat_kfunc_set,
> + .filter = bpf_tcp_ops_rcvlowat_kfunc_filter,
> +};
> +
> static int __init __bpf_tcp_ops_init(void)
> {
> - return register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
> + int ret;
> +
> + ret = register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS,
> + &bpf_tcp_ops_rcvlowat_kfunc_id_set);
> + ret = ret ?: register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
> +
> + return ret;
> }
> late_initcall(__bpf_tcp_ops_init);
Hi Kuniyuki,
Tested-by: Clément Léger <cleger@meta.com>
Thanks,
Clément
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH bpf-next 7/7] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-09-20 19:56 ` [PATCH bpf-next 7/7] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-09-20 21:16 ` bot+bpf-ci
2026-09-21 20:50 ` Stanislav Fomichev
@ 2026-09-22 23:14 ` Kuniyuki Iwashima
2 siblings, 0 replies; 28+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-22 23:14 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On Sun, Sep 20, 2026 at 12:56 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
[...]
> diff --git a/tools/testing/selftests/bpf/progs/tcp_autolowat.c b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
> new file mode 100644
> index 000000000000..bb96e19e7589
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
[...]
> +static int tcp_parse_descriptor(struct tcp_autolowat_cb *cb,
> + struct bpf_dynptr *dptr,
> + u32 seq, u32 end_seq)
> +{
> + struct rpc_descriptor *rpc_desc;
> + u32 rpc_copied_seq;
> + u64 copy_len; /* u32 should work, but not for no_alu32 :/ */
> + u64 rpc_len;
> + int err;
> +
> + rpc_copied_seq = cb->rpc_desc_seq + cb->rpc_desc_buff_len;
> +
> + if (before(cb->rpc_desc_seq + RPC_DESC_SIZE, end_seq))
> + copy_len = RPC_DESC_SIZE - cb->rpc_desc_buff_len;
> + else
> + copy_len = end_seq - rpc_copied_seq;
> +
> + if (copy_len == 0)
> + goto disable; /* FIN. */
> + if (copy_len > RPC_DESC_SIZE)
> + goto disable; /* always false, only for verifier. */
> + if (cb->rpc_desc_buf + cb->rpc_desc_buff_len >= &cb->rpc_desc_buf[RPC_DESC_SIZE])
It seems clang on my local host omit the base pointer offset
from the comparison while gcc used in BPF CI didin't and CI failed.
https://github.com/kernel-patches/bpf/actions/runs/35774477359/job/106910558464?pr=13945
By changing the line to
if (cb->rpc_desc_buff_len >= RPC_DESC_SIZE)
, CI passed.
https://github.com/kernel-patches/bpf/pull/13991/commits
> + goto disable; /* always false, only for verifier. */
> +
> + err = bpf_dynptr_read(cb->rpc_desc_buf + cb->rpc_desc_buff_len,
> + copy_len, dptr, rpc_copied_seq - seq, 0);
> + if (err)
> + goto disable;
^ permalink raw reply [flat|nested] 28+ messages in thread
end of thread, other threads:[~2026-09-22 23:14 UTC | newest]
Thread overview: 28+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-20 19:56 [PATCH bpf-next 0/7] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
2026-09-20 19:56 ` [PATCH bpf-next 1/7] selftest: bpf: Use BPF_SOCK_OPS_ALL_CB_FLAGS + 1 for bad_cb_test_rv Kuniyuki Iwashima
2026-09-21 20:48 ` Stanislav Fomichev
2026-09-20 19:56 ` [PATCH bpf-next 2/7] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-09-20 21:16 ` bot+bpf-ci
2026-09-21 20:48 ` Stanislav Fomichev
2026-09-20 19:56 ` [PATCH bpf-next 3/7] bpf: tcp: Support bpf_sock_ops_cb_flags_set() for bpf_tcp_ops Kuniyuki Iwashima
2026-09-20 20:11 ` sashiko-bot
2026-09-20 21:05 ` Kuniyuki Iwashima
2026-09-21 20:48 ` Stanislav Fomichev
2026-09-20 19:56 ` [PATCH bpf-next 4/7] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
2026-09-20 21:01 ` bot+bpf-ci
2026-09-20 21:12 ` Kuniyuki Iwashima
2026-09-21 20:48 ` Stanislav Fomichev
2026-09-20 19:56 ` [PATCH bpf-next 5/7] bpf: mptcp: Don't support BPF_SOCK_OPS_RCVQ_CB_FLAG Kuniyuki Iwashima
2026-09-20 20:05 ` sashiko-bot
2026-09-20 21:09 ` Kuniyuki Iwashima
2026-09-20 19:56 ` [PATCH bpf-next 6/7] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
2026-09-20 20:13 ` sashiko-bot
2026-09-20 21:10 ` Kuniyuki Iwashima
2026-09-20 21:16 ` bot+bpf-ci
2026-09-21 20:49 ` Stanislav Fomichev
2026-09-21 21:36 ` Kuniyuki Iwashima
2026-09-22 7:29 ` Clément Léger
2026-09-20 19:56 ` [PATCH bpf-next 7/7] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-09-20 21:16 ` bot+bpf-ci
2026-09-21 20:50 ` Stanislav Fomichev
2026-09-22 23:14 ` Kuniyuki Iwashima
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.