* [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT.
@ 2026-10-06 19:24 Kuniyuki Iwashima
2026-10-06 19:24 ` [PATCH v4 bpf-next 01/10] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
` (10 more replies)
0 siblings, 11 replies; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-06 19:24 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
This series introduces two callbacks for bpf_tcp_ops:
.enqueue_rcvq(): invoked when TCP stack enqueues skb to
sk->sk_receive_queue
.dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
is dequeued from sk->sk_receive_queue
Those callbacks can be enabled on a per-socket basis by
a new kfunc bpf_tcp_ops_set_flags():
bpf_tcp_ops_set_flags((struct tcp_sock *)sk,
BPF_TCP_OPS_FLAG_RCVQ, 0);
This allows the BPF prog to dynamically adjust sk->sk_rcvlowat,
suppressing unnecessary EPOLLIN wakeups until sufficient data
is available in the receive queue.
This functionality, which we call "TCP AutoLOWAT", was originally
developed in 2020 by Tenzin Ukyab with the help of Soheil Hassas
Yeganeh, Arjun Roy, and Eric Dumazet. It has served Google RPC
workloads for more than 5 years.
Combined with TCP RX zerocopy, this typically allows us to read an
entire RPC frame with just a single wakeup and a single system call.
While the original implementation was specialised for our
internal RPC format, this series introduces a more flexible
version by leveraging BPF.
The bpf prog in the last selftest patch closely mirrors the core
logic of the original implementation to provide a real-world
example.
Note that the new callbacks are not supported on legacy SOCK_OPS.
Changes:
v4:
* Patch 2
* Allow-listed BPF_PROG_TYPE_CGROUP_SOCKOPT and _SKB
in bpf_tcp_ops_set_flags_kfunc_filter()
* Add doc for flag enum and bpf_tcp_ops members
* Patch 4
* Check static key first and then flags
* Patch 5 (new)
* Bubble up static key check for SOCK_OPS similar to bpf_tcp_ops
* Patch 6
* Use bpf_tcp_ops_call_flag()
v3: https://lore.kernel.org/bpf/20261005154533.4147685-1-kuniyu@google.com/
* Add patch 2 ~ 4 (guard bpf_tcp_ops w/ a new flag)
* Patch 5 & 9 : Adapt to a new flag and kfunc
v2: https://lore.kernel.org/netdev/20260923213719.224838-1-kuniyu@google.com/
* Add patch 1 not to allow setsockopt() from new callbacks
* Patch 8 (selftest)
* Avoid address comparison for a specific version of gcc.
* Make rpc_test_cases[] static.
* Update comment in rpc_test_case[].
v1: https://lore.kernel.org/bpf/20260920195633.3033620-1-kuniyu@google.com/
Legacy SOCK_OPS version:
v3: https://lore.kernel.org/bpf/20260523083001.2911931-1-kuniyu@google.com/
v2: https://lore.kernel.org/bpf/20260522074601.1658705-1-kuniyu@google.com/
v1: https://lore.kernel.org/bpf/20260508073355.3916746-1-kuniyu@google.com/
Kuniyuki Iwashima (10):
bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list.
bpf: tcp: Add a new per-socket flag and kfunc for bpf_tcp_ops.
selftest: bpf: Use bpf_tcp_ops_set_flags() in bpf_tcp_ops_hdr.c.
bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag.
bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after
cgroup_bpf_enabled(CGROUP_SOCK_OPS).
bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
tcp: Split out __tcp_set_rcvlowat().
bpf: mptcp: Don't support BPF_TCP_OPS_FLAG_RCVQ.
bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().
.../networking/net_cachelines/tcp_sock.rst | 1 +
include/linux/bpf-cgroup.h | 36 +-
include/linux/tcp.h | 7 +
include/net/tcp.h | 122 +++---
include/uapi/linux/bpf.h | 25 ++
net/ipv4/af_inet.c | 2 +-
net/ipv4/bpf_tcp_ops.c | 159 +++++++-
net/ipv4/tcp.c | 20 +-
net/ipv4/tcp_fastopen.c | 2 +
net/ipv4/tcp_input.c | 37 +-
net/ipv4/tcp_nv.c | 2 +-
net/ipv4/tcp_output.c | 61 ++-
net/ipv4/tcp_timer.c | 6 +-
tools/include/uapi/linux/bpf.h | 25 ++
.../selftests/bpf/prog_tests/tcp_autolowat.c | 350 ++++++++++++++++++
.../selftests/bpf/progs/bpf_tcp_ops_hdr.c | 20 +
.../selftests/bpf/progs/bpf_tracing_net.h | 2 +
.../selftests/bpf/progs/tcp_autolowat.c | 294 +++++++++++++++
18 files changed, 1033 insertions(+), 138 deletions(-)
create mode 100644 tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
create mode 100644 tools/testing/selftests/bpf/progs/tcp_autolowat.c
--
2.56.0.360.g66cac248cb-goog
^ permalink raw reply [flat|nested] 36+ messages in thread
* [PATCH v4 bpf-next 01/10] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list.
2026-10-06 19:24 [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
@ 2026-10-06 19:24 ` Kuniyuki Iwashima
2026-10-06 23:09 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 02/10] bpf: tcp: Add a new per-socket flag and kfunc for bpf_tcp_ops Kuniyuki Iwashima
` (9 subsequent siblings)
10 siblings, 1 reply; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-06 19:24 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev, Emil Tsalapatis
Currently, four bpf_tcp_ops callbacks are not allowed to call
bpf_setsockopt() and bpf_getsockopt().
However, the deny-list is fragile, and when a new callback is
added, we might re-open a can of worms. [0][1]
Let's convert it to allow-list.
Note that it can be written as offsetof(,X) <= moff <= offsetof(,Y)
but clang can optimise to similar code anyway.
Link: https://lore.kernel.org/r/20220929070407.965581-5-martin.lau@linux.dev/ #[0]
Link: https://lore.kernel.org/bpf/20260421155804.135786-1-kafai.wan@linux.dev/ #[1]
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
---
net/ipv4/bpf_tcp_ops.c | 39 ++++++++++++++++++++++++---------------
1 file changed, 24 insertions(+), 15 deletions(-)
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index 681fed642999..1ada3b781bf1 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -210,6 +210,24 @@ const struct bpf_func_proto bpf_tcp_ops_get_retval_proto = {
.ret_type = RET_INTEGER,
};
+static bool is_sockopt_supported(u32 moff)
+{
+ switch (moff) {
+ case offsetof(struct bpf_tcp_ops, active_established):
+ case offsetof(struct bpf_tcp_ops, passive_established):
+ case offsetof(struct bpf_tcp_ops, rto):
+ case offsetof(struct bpf_tcp_ops, rtt):
+ case offsetof(struct bpf_tcp_ops, set_state):
+ case offsetof(struct bpf_tcp_ops, retrans):
+ case offsetof(struct bpf_tcp_ops, connect):
+ case offsetof(struct bpf_tcp_ops, listen):
+ case offsetof(struct bpf_tcp_ops, parse_hdr):
+ return true;
+ }
+
+ return false;
+}
+
static const struct bpf_func_proto *
get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
{
@@ -221,22 +239,13 @@ get_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
case BPF_FUNC_sk_storage_delete:
return &bpf_sk_storage_delete_proto;
case BPF_FUNC_setsockopt:
- /* The sk may be an unlocked listener (synack path) or NULL
- * fullsock; disable for members that can run unlocked.
- */
- if (moff == offsetof(struct bpf_tcp_ops, rwnd_init) ||
- moff == offsetof(struct bpf_tcp_ops, timeout_init) ||
- moff == offsetof(struct bpf_tcp_ops, hdr_opt_len) ||
- moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))
- return NULL;
- return &bpf_sk_setsockopt_proto;
+ if (is_sockopt_supported(moff))
+ return &bpf_sk_setsockopt_proto;
+ return NULL;
case BPF_FUNC_getsockopt:
- if (moff == offsetof(struct bpf_tcp_ops, rwnd_init) ||
- moff == offsetof(struct bpf_tcp_ops, timeout_init) ||
- moff == offsetof(struct bpf_tcp_ops, hdr_opt_len) ||
- moff == offsetof(struct bpf_tcp_ops, write_hdr_opt))
- return NULL;
- return &bpf_sk_getsockopt_proto;
+ if (is_sockopt_supported(moff))
+ return &bpf_sk_getsockopt_proto;
+ return NULL;
case BPF_FUNC_get_retval:
if (moff == offsetof(struct bpf_tcp_ops, timeout_init) ||
moff == offsetof(struct bpf_tcp_ops, rwnd_init))
--
2.56.0.360.g66cac248cb-goog
^ permalink raw reply related [flat|nested] 36+ messages in thread
* [PATCH v4 bpf-next 02/10] bpf: tcp: Add a new per-socket flag and kfunc for bpf_tcp_ops.
2026-10-06 19:24 [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
2026-10-06 19:24 ` [PATCH v4 bpf-next 01/10] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
@ 2026-10-06 19:24 ` Kuniyuki Iwashima
2026-10-06 22:04 ` Stanislav Fomichev
2026-10-06 23:10 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 03/10] selftest: bpf: Use bpf_tcp_ops_set_flags() in bpf_tcp_ops_hdr.c Kuniyuki Iwashima
` (8 subsequent siblings)
10 siblings, 2 replies; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-06 19:24 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
The legacy SOCK_OPS guards some hooks with a per-socket flag,
tp->bpf_sock_ops_cb_flags.
In contrast, bpf_tcp_ops was initially designed without per-socket
flags so that users can simply define only the callbacks they need.
However, it turned out that even attaching a bpf_tcp_ops with a NULL
callback incurs measurable overhead in the fast path. [0]
We could reuse tp->bpf_sock_ops_cb_flags, but it is u8 and has only
1 bit left for a new opt-in callback, while not all of the 7 existing
opt-in hooks are in the fast path and really need a flag guard.
Moreover, tp->bpf_sock_ops_cb_flags has two bpf helper functions to
update it, both of which require sock_owned_by_me(sk):
* bpf_sock_ops_cb_flags_set()
* bpf_setsockopt(TCP_BPF_SOCK_OPS_CB_FLAGS)
Both helpers only overwrite the field, which leads to reading and
modifying the flags in the BPF prog and then writing them back
via the helper. This prevents use from the fast path (tc or
cgroup_skb) hooks where bh_lock_sock() alone cannot prevent races
with process context.
Let's add a new u32 field, tp->bpf_tcp_ops_flags, at the end of the
tcp_sock_read_txrx cacheline group, along with a new kfunc,
bpf_tcp_ops_set_flags().
bpf_tcp_ops_set_flags() takes bitmasks of flags to enable and
disable and updates tp->bpf_tcp_ops_flags atomically via
try_cmpxchg() without relying on lock_sock().
The kfunc is exposed to bpf_tcp_ops, BPF_PROG_TYPE_CGROUP_SOCKOPT,
and BPF_PROG_TYPE_CGROUP_SKB.
The first argument is struct tcp_sock * so that bpf_tcp_sock() etc
is required for cgroup hooks, but not for bpf_tcp_ops where struct
sock * is promoted to struct tcp_sock * automatically.
The fast-path callbacks will be guarded by the new flags in a later
patch after updating the existing selftest to avoid breaking bisection.
Link: https://lore.kernel.org/netdev/CAMB2axMwBuz3X4Uwn5uzZUqm91EbiMnbQX832wUFOQ44fOwDnQ@mail.gmail.com/ #[0]
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
v4:
* Allow-listed BPF_PROG_TYPE_CGROUP_SOCKOPT and _SKB
in bpf_tcp_ops_set_flags_kfunc_filter()
* Add doc for flag enum and bpf_tcp_ops members
---
.../networking/net_cachelines/tcp_sock.rst | 1 +
include/linux/tcp.h | 7 +++
include/net/tcp.h | 11 +++-
include/uapi/linux/bpf.h | 23 +++++++
net/ipv4/bpf_tcp_ops.c | 61 ++++++++++++++++++-
net/ipv4/tcp.c | 3 +
tools/include/uapi/linux/bpf.h | 23 +++++++
7 files changed, 127 insertions(+), 2 deletions(-)
diff --git a/Documentation/networking/net_cachelines/tcp_sock.rst b/Documentation/networking/net_cachelines/tcp_sock.rst
index 0f6088c4ab8b..420fc6278148 100644
--- a/Documentation/networking/net_cachelines/tcp_sock.rst
+++ b/Documentation/networking/net_cachelines/tcp_sock.rst
@@ -151,6 +151,7 @@ u32 urg_seq
unsigned_int keepalive_time
unsigned_int keepalive_intvl
int linger2
+u32 bpf_tcp_ops_flags read_mostly read_mostly bpf_tcp_ops_hdr_opt_len,bpf_skops_write_hdr_opt(tx);bpf_tcp_ops_parse_hdr,tcp_bpf_rtt(rx);
u8 bpf_sock_ops_cb_flags
u8:1 bpf_chg_cc_inprogress
u16 timeout_rehash
diff --git a/include/linux/tcp.h b/include/linux/tcp.h
index 6a8c77719322..24bb751cc65a 100644
--- a/include/linux/tcp.h
+++ b/include/linux/tcp.h
@@ -233,6 +233,13 @@ struct tcp_sock {
is_sack_reneg:1, /* in recovery from loss with SACK reneg? */
is_cwnd_limited:1,/* forward progress limited by snd_cwnd? */
recvmsg_inq : 1;/* Indicate # of bytes in queue upon recvmsg */
+#ifdef CONFIG_BPF
+ u32 bpf_tcp_ops_flags;
+#define BPF_TCP_OPS_TEST_FLAG(TP, ARG) \
+ (READ_ONCE((TP)->bpf_tcp_ops_flags) & BPF_TCP_OPS_FLAG_ ## ARG)
+#else
+#define BPF_TCP_OPS_TEST_FLAG(TP, ARG) (0)
+#endif
__cacheline_group_end(tcp_sock_read_txrx);
/* RX read-mostly hotpath cache lines */
diff --git a/include/net/tcp.h b/include/net/tcp.h
index 4ecabf4989de..55ad9db99596 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -2934,6 +2934,7 @@ static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
{
tcp_sk(sk)->bpf_sock_ops_cb_flags = 0;
+ WRITE_ONCE(tcp_sk(sk)->bpf_tcp_ops_flags, 0);
}
#else
@@ -2989,7 +2990,7 @@ struct bpf_tcp_ops {
/* Called when the retransmission timer fires. */
void (*rto)(struct sock *sk);
- /* Called on every RTT sample.
+ /* Called on every RTT sample if BPF_TCP_OPS_FLAG_RTT is enabled.
* @mrtt: the measured RTT, in microseconds.
* @srtt: the updated smoothed RTT.
*/
@@ -3016,6 +3017,10 @@ struct bpf_tcp_ops {
* Parse the TCP header options of an incoming skb received on an
* established connection. Use bpf_dynptr_from_skb()/bpf_skb_load_bytes()
* to access the options.
+ *
+ * Called if BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_ALL is enabled, or if
+ * BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_UNKNOWN is enabled and an unknown
+ * option is received.
*/
void (*parse_hdr)(struct sock *sk, struct sk_buff *skb);
@@ -3023,6 +3028,8 @@ struct bpf_tcp_ops {
* Reserve space in the outgoing TCP header for options to be written
* later by write_hdr_opt(). Call bpf_reserve_hdr_opt() to reserve bytes.
*
+ * Called if BPF_TCP_OPS_FLAG_WRITE_HDR_OPT is enabled.
+ *
* @skb: outgoing packet. NULL when called from tcp_current_mss()
* (MSS sizing).
* @req: request_sock on the synack path; NULL otherwise.
@@ -3041,6 +3048,8 @@ struct bpf_tcp_ops {
* Use bpf_store_hdr_opt() to write; it appends within the reserved window
* shared with legacy SOCKOPS.
*
+ * Called if BPF_TCP_OPS_FLAG_WRITE_HDR_OPT is enabled.
+ *
* @skb: outgoing packet.
* @req: request_sock on the synack path; NULL otherwise.
* @syn_skb: incoming SYN on the synack path; NULL otherwise.
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index e0ed44b1bbcb..6963c146311e 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -7340,6 +7340,29 @@ enum {
*/
};
+/*
+ * Most callbacks of struct bpf_tcp_ops are placed in the slow
+ * path (e.g., one-shot connection setup or unlikely events like
+ * timers) and are invoked simply by defining non-NULL callbacks.
+ *
+ * Callbacks in the fast path, however, would incur noticeable
+ * overhead even when set to NULL, so they are disabled by default
+ * and must be explicitly enabled per socket via bpf_tcp_ops_set_flags().
+ *
+ * The flags are per-socket and shared by all effective bpf_tcp_ops
+ * programs.
+ */
+enum {
+ /* .rtt() */
+ BPF_TCP_OPS_FLAG_RTT = (1 << 0),
+ /* .parse_hdr() */
+ BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_ALL = (1 << 1),
+ BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_UNKNOWN = (1 << 2),
+ /* .hdr_opt_len() and .write_hdr_opt() */
+ BPF_TCP_OPS_FLAG_WRITE_HDR_OPT = (1 << 3),
+ BPF_TCP_OPS_FLAG_ALL = (1 << 4) - 1,
+};
+
/* List of TCP states. There is a build check in net/ipv4/tcp.c to detect
* changes between the TCP and BPF versions. Ideally this should never happen.
* If it does, we need to add code to convert them before calling
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index 1ada3b781bf1..aaf96304c69f 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -328,8 +328,67 @@ static struct bpf_struct_ops bpf_tcp_ops = {
.owner = THIS_MODULE,
};
+__bpf_kfunc_start_defs();
+
+__bpf_kfunc int bpf_tcp_ops_set_flags(struct tcp_sock *tp, u32 enable, u32 disable)
+{
+ u32 old, new;
+
+ if ((enable & disable) || (enable | disable) & ~BPF_TCP_OPS_FLAG_ALL)
+ return -EINVAL;
+
+ old = READ_ONCE(tp->bpf_tcp_ops_flags);
+
+ do {
+ new = (old | enable) & ~disable;
+ if (new == old)
+ break;
+ } while (!try_cmpxchg(&tp->bpf_tcp_ops_flags, &old, new));
+
+ return 0;
+}
+
+__bpf_kfunc_end_defs();
+
+BTF_KFUNCS_START(bpf_tcp_ops_set_flags_kfunc_set)
+BTF_ID_FLAGS(func, bpf_tcp_ops_set_flags)
+BTF_KFUNCS_END(bpf_tcp_ops_set_flags_kfunc_set)
+
+static int bpf_tcp_ops_set_flags_kfunc_filter(const struct bpf_prog *prog,
+ u32 kfunc_id)
+{
+ if (!btf_id_set8_contains(&bpf_tcp_ops_set_flags_kfunc_set, kfunc_id))
+ return 0;
+
+ if (prog->type == BPF_PROG_TYPE_STRUCT_OPS &&
+ prog->aux->st_ops == &bpf_tcp_ops)
+ return 0;
+
+ if (prog->type == BPF_PROG_TYPE_CGROUP_SOCKOPT ||
+ prog->type == BPF_PROG_TYPE_CGROUP_SKB)
+ return 0;
+
+ return -EACCES;
+}
+
+static const struct btf_kfunc_id_set bpf_tcp_ops_set_flags_kfunc_id_set = {
+ .owner = THIS_MODULE,
+ .set = &bpf_tcp_ops_set_flags_kfunc_set,
+ .filter = bpf_tcp_ops_set_flags_kfunc_filter,
+};
+
static int __init __bpf_tcp_ops_init(void)
{
- return register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
+ int ret;
+
+ ret = register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS,
+ &bpf_tcp_ops_set_flags_kfunc_id_set);
+ /* BPF_PROG_TYPE_CGROUP_{SOCKOPT,SKB} share BTF_KFUNC_HOOK_CGROUP. */
+ ret = ret ?: register_btf_kfunc_id_set(BPF_PROG_TYPE_CGROUP_SOCKOPT,
+ &bpf_tcp_ops_set_flags_kfunc_id_set);
+ ret = ret ?: register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
+
+ return ret;
}
+
late_initcall(__bpf_tcp_ops_init);
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index 650a2e89950a..fa69961c47d3 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -5217,6 +5217,9 @@ static void __init tcp_struct_check(void)
CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_read_txrx, lost_out);
CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_read_txrx, sacked_out);
CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_read_txrx, scaling_ratio);
+#ifdef CONFIG_BPF
+ CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_read_txrx, bpf_tcp_ops_flags);
+#endif
/* RX read-mostly hotpath cache lines */
CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_read_rx, copied_seq);
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index e0ed44b1bbcb..6963c146311e 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -7340,6 +7340,29 @@ enum {
*/
};
+/*
+ * Most callbacks of struct bpf_tcp_ops are placed in the slow
+ * path (e.g., one-shot connection setup or unlikely events like
+ * timers) and are invoked simply by defining non-NULL callbacks.
+ *
+ * Callbacks in the fast path, however, would incur noticeable
+ * overhead even when set to NULL, so they are disabled by default
+ * and must be explicitly enabled per socket via bpf_tcp_ops_set_flags().
+ *
+ * The flags are per-socket and shared by all effective bpf_tcp_ops
+ * programs.
+ */
+enum {
+ /* .rtt() */
+ BPF_TCP_OPS_FLAG_RTT = (1 << 0),
+ /* .parse_hdr() */
+ BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_ALL = (1 << 1),
+ BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_UNKNOWN = (1 << 2),
+ /* .hdr_opt_len() and .write_hdr_opt() */
+ BPF_TCP_OPS_FLAG_WRITE_HDR_OPT = (1 << 3),
+ BPF_TCP_OPS_FLAG_ALL = (1 << 4) - 1,
+};
+
/* List of TCP states. There is a build check in net/ipv4/tcp.c to detect
* changes between the TCP and BPF versions. Ideally this should never happen.
* If it does, we need to add code to convert them before calling
--
2.56.0.360.g66cac248cb-goog
^ permalink raw reply related [flat|nested] 36+ messages in thread
* [PATCH v4 bpf-next 03/10] selftest: bpf: Use bpf_tcp_ops_set_flags() in bpf_tcp_ops_hdr.c.
2026-10-06 19:24 [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
2026-10-06 19:24 ` [PATCH v4 bpf-next 01/10] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
2026-10-06 19:24 ` [PATCH v4 bpf-next 02/10] bpf: tcp: Add a new per-socket flag and kfunc for bpf_tcp_ops Kuniyuki Iwashima
@ 2026-10-06 19:24 ` Kuniyuki Iwashima
2026-10-06 22:04 ` Stanislav Fomichev
2026-10-06 23:12 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 04/10] bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag Kuniyuki Iwashima
` (7 subsequent siblings)
10 siblings, 2 replies; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-06 19:24 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
The next patch guards bpf_tcp_ops callbacks in the fast
path with per-socket flags.
Let's use bpf_tcp_ops_set_flags() to enable flags for
parse_hdr, write_hdr_opt, and hdr_opt_len.
Note that BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_ALL and
BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_UNKNOWN are used on the
server and client, respectively, for better coverage.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
.../selftests/bpf/progs/bpf_tcp_ops_hdr.c | 20 +++++++++++++++++++
1 file changed, 20 insertions(+)
diff --git a/tools/testing/selftests/bpf/progs/bpf_tcp_ops_hdr.c b/tools/testing/selftests/bpf/progs/bpf_tcp_ops_hdr.c
index f3e3dff13784..706e94d794e4 100644
--- a/tools/testing/selftests/bpf/progs/bpf_tcp_ops_hdr.c
+++ b/tools/testing/selftests/bpf/progs/bpf_tcp_ops_hdr.c
@@ -18,6 +18,24 @@ int found_cnt;
__u8 found_d0;
__u8 found_d1;
+SEC("struct_ops")
+void BPF_PROG(test_listen, struct sock *sk)
+{
+ bpf_tcp_ops_set_flags((struct tcp_sock *)sk,
+ BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_ALL |
+ BPF_TCP_OPS_FLAG_WRITE_HDR_OPT,
+ 0);
+}
+
+SEC("struct_ops")
+void BPF_PROG(test_connect, struct sock *sk)
+{
+ bpf_tcp_ops_set_flags((struct tcp_sock *)sk,
+ BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_UNKNOWN |
+ BPF_TCP_OPS_FLAG_WRITE_HDR_OPT,
+ 0);
+}
+
SEC("struct_ops")
void BPF_PROG(test_hdr_opt_len, struct sock *sk, struct sk_buff *skb,
struct request_sock *req, struct sk_buff *syn_skb,
@@ -78,6 +96,8 @@ void BPF_PROG(test_parse_hdr, struct sock *sk, struct sk_buff *skb)
SEC(".struct_ops.link")
struct bpf_tcp_ops test_hdr_ops = {
+ .listen = (void *)test_listen,
+ .connect = (void *)test_connect,
.hdr_opt_len = (void *)test_hdr_opt_len,
.write_hdr_opt = (void *)test_write_hdr_opt,
.parse_hdr = (void *)test_parse_hdr,
--
2.56.0.360.g66cac248cb-goog
^ permalink raw reply related [flat|nested] 36+ messages in thread
* [PATCH v4 bpf-next 04/10] bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag.
2026-10-06 19:24 [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
` (2 preceding siblings ...)
2026-10-06 19:24 ` [PATCH v4 bpf-next 03/10] selftest: bpf: Use bpf_tcp_ops_set_flags() in bpf_tcp_ops_hdr.c Kuniyuki Iwashima
@ 2026-10-06 19:24 ` Kuniyuki Iwashima
2026-10-06 22:05 ` Stanislav Fomichev
2026-10-06 23:25 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 05/10] bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after cgroup_bpf_enabled(CGROUP_SOCK_OPS) Kuniyuki Iwashima
` (6 subsequent siblings)
10 siblings, 2 replies; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-06 19:24 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
bpf_tcp_ops.{parse_hdr,hdr_opt_len} are called for every
incoming / outgoing skb.
bpf_tcp_ops.rtt is called once per RTT, which is every
incoming skb in ping-pong workloads like tcp_rr.
Even attaching NULL callbacks in the fast path hurts performance.
Let's guard them (and write_hdr_opt) with the new per-socket flags.
__bpf_tcp_ops_call() and bpf_tcp_ops_call_flag() are added to check
cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS) first and avoid accessing
tp->bpf_tcp_ops_flags when no bpf_tcp_ops is attached.
Note that bpf_tcp_ops_hdr_opt_len() has 3 callers and previously
checked the static key both in tcp_established_options() and inside
bpf_tcp_ops_call(). Now the static key is checked once at the
beginning of bpf_tcp_ops_hdr_opt_len(), and the redundant check in
tcp_established_options() is removed.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
v4: Check static key first and then flags
---
include/net/tcp.h | 44 ++++++++++++++++++++++++++++---------------
net/ipv4/tcp_input.c | 12 +++++++++++-
net/ipv4/tcp_output.c | 34 ++++++++++++++++-----------------
3 files changed, 57 insertions(+), 33 deletions(-)
diff --git a/include/net/tcp.h b/include/net/tcp.h
index 55ad9db99596..b7c0f1a8797a 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -3064,22 +3064,33 @@ struct bpf_tcp_ops {
u32 opt_off);
};
-#define bpf_tcp_ops_call(op, sk, ...) \
+#define __bpf_tcp_ops_call(op, sk, ...) \
do { \
- if (cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS)) { \
- const struct bpf_prog_array_item *item; \
- const struct bpf_tcp_ops *tcp_ops; \
- struct cgroup *cgrp; \
+ const struct bpf_prog_array_item *item; \
+ const struct bpf_tcp_ops *tcp_ops; \
+ struct cgroup *cgrp; \
\
- cgrp = sock_cgroup_ptr(&sk->sk_cgrp_data); \
- rcu_read_lock_dont_migrate(); \
- bpf_cgroup_struct_ops_foreach(tcp_ops, item, cgrp, \
- CGROUP_TCP_SOCK_OPS) { \
- if (tcp_ops->op) \
- tcp_ops->op(sk, ##__VA_ARGS__); \
- } \
- rcu_read_unlock_migrate(); \
+ cgrp = sock_cgroup_ptr(&sk->sk_cgrp_data); \
+ rcu_read_lock_dont_migrate(); \
+ bpf_cgroup_struct_ops_foreach(tcp_ops, item, cgrp, \
+ CGROUP_TCP_SOCK_OPS) { \
+ if (tcp_ops->op) \
+ tcp_ops->op(sk, ##__VA_ARGS__); \
} \
+ rcu_read_unlock_migrate(); \
+} while (0)
+
+#define bpf_tcp_ops_call(op, sk, ...) \
+do { \
+ if (cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS)) \
+ __bpf_tcp_ops_call(op, sk, ##__VA_ARGS__); \
+} while (0)
+
+#define bpf_tcp_ops_call_flag(op, flag, sk, ...) \
+do { \
+ if (cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS) && \
+ BPF_TCP_OPS_TEST_FLAG(tcp_sk(sk), flag)) \
+ __bpf_tcp_ops_call(op, sk, ##__VA_ARGS__); \
} while (0)
#define bpf_tcp_ops_call_int(op, init_retval, sk, ...) \
@@ -3115,7 +3126,9 @@ do { \
})
#else
-#define bpf_tcp_ops_call(op, sk, ...) do { } while (0)
+#define __bpf_tcp_ops_call(op, sk, ...) do { } while (0)
+#define bpf_tcp_ops_call(op, sk, ...) do { } while (0)
+#define bpf_tcp_ops_call_flag(op, flag, sk, ...) do { } while (0)
#define bpf_tcp_ops_call_int(op, init_retval, sk, ...) (init_retval)
#endif
@@ -3150,7 +3163,8 @@ static inline void tcp_bpf_rtt(struct sock *sk, long mrtt, u32 srtt)
{
if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RTT_CB_FLAG))
tcp_call_bpf_2arg(sk, BPF_SOCK_OPS_RTT_CB, mrtt, srtt);
- bpf_tcp_ops_call(rtt, sk, mrtt, srtt);
+
+ bpf_tcp_ops_call_flag(rtt, RTT, sk, mrtt, srtt);
}
#if IS_ENABLED(CONFIG_SMC)
diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
index 209db8effcc4..4478d3f3d4b0 100644
--- a/net/ipv4/tcp_input.c
+++ b/net/ipv4/tcp_input.c
@@ -210,6 +210,11 @@ static void bpf_skops_established(struct sock *sk, int bpf_op,
static void bpf_tcp_ops_parse_hdr(struct sock *sk, struct sk_buff *skb)
{
+ const struct tcp_sock *tp;
+
+ if (!cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS))
+ return;
+
switch (sk->sk_state) {
case TCP_SYN_RECV:
case TCP_SYN_SENT:
@@ -217,7 +222,12 @@ static void bpf_tcp_ops_parse_hdr(struct sock *sk, struct sk_buff *skb)
return;
}
- bpf_tcp_ops_call(parse_hdr, sk, skb);
+ tp = tcp_sk(sk);
+
+ if ((tp->rx_opt.saw_unknown &&
+ BPF_TCP_OPS_TEST_FLAG(tp, PARSE_HDR_OPT_UNKNOWN)) ||
+ BPF_TCP_OPS_TEST_FLAG(tp, PARSE_HDR_OPT_ALL))
+ __bpf_tcp_ops_call(parse_hdr, sk, skb);
}
static __cold void tcp_gro_dev_warn(const struct sock *sk, const struct sk_buff *skb,
diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c
index 3ac465513bf5..8770f3084efe 100644
--- a/net/ipv4/tcp_output.c
+++ b/net/ipv4/tcp_output.c
@@ -581,8 +581,8 @@ static void bpf_skops_write_hdr_opt(struct sock *sk, struct sk_buff *skb,
* writer's bytes). The writer finds the append point by scanning from
* first_opt_off + nr_written to the first NOP.
*/
- bpf_tcp_ops_call(write_hdr_opt, sk, skb, req, syn_skb, synack_type,
- first_opt_off + nr_written);
+ bpf_tcp_ops_call_flag(write_hdr_opt, WRITE_HDR_OPT, sk, skb, req,
+ syn_skb, synack_type, first_opt_off + nr_written);
}
#else
static u32 bpf_skops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,
@@ -613,11 +613,13 @@ static u32 bpf_tcp_ops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,
{
unsigned int remaining_out = remaining, reserved;
- if (!remaining)
- return 0;
+ if (!cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS) ||
+ !BPF_TCP_OPS_TEST_FLAG(tcp_sk(sk), WRITE_HDR_OPT)||
+ !remaining)
+ return remaining;
/* bpf_tcp_ops_reserve_hdr_opt() reserves space via remaining_out */
- bpf_tcp_ops_call(hdr_opt_len, sk, skb, req, syn_skb, synack_type, &remaining_out);
+ __bpf_tcp_ops_call(hdr_opt_len, sk, skb, req, syn_skb, synack_type, &remaining_out);
reserved = remaining - remaining_out;
if (!reserved)
@@ -1193,8 +1195,9 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
struct tcp_key *key)
{
struct tcp_sock *tp = tcp_sk(sk);
- unsigned int size = 0;
unsigned int eff_sacks;
+ unsigned int remaining;
+ unsigned int size = 0;
opts->options = 0;
opts->bpf_opt_len = 0;
@@ -1224,10 +1227,10 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
* left.
*/
if (sk_is_mptcp(sk)) {
- unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
bool has_ts = opts->options & OPTION_TS;
int opt_size;
+ remaining = MAX_TCP_OPTION_SPACE - size;
opts->mptcp.drop_ts = 0;
opt_size = mptcp_established_options(sk, skb, remaining, has_ts,
@@ -1244,7 +1247,8 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
eff_sacks = tp->rx_opt.num_sacks + tp->rx_opt.dsack;
if (unlikely(eff_sacks)) {
- const unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
+ remaining = MAX_TCP_OPTION_SPACE - size;
+
if (likely(remaining >= TCPOLEN_SACK_BASE_ALIGNED +
TCPOLEN_SACK_PERBLOCK)) {
opts->num_sack_blocks =
@@ -1277,22 +1281,18 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
if (unlikely(BPF_SOCK_OPS_TEST_FLAG(tp,
BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG))) {
- unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
-
+ remaining = MAX_TCP_OPTION_SPACE - size;
remaining = bpf_skops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
remaining);
size = MAX_TCP_OPTION_SPACE - remaining;
}
- if (cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS)) {
- unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
-
- remaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
- remaining);
+ remaining = MAX_TCP_OPTION_SPACE - size;
+ remaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
+ remaining);
- size = MAX_TCP_OPTION_SPACE - remaining;
- }
+ size = MAX_TCP_OPTION_SPACE - remaining;
return size;
}
--
2.56.0.360.g66cac248cb-goog
^ permalink raw reply related [flat|nested] 36+ messages in thread
* [PATCH v4 bpf-next 05/10] bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after cgroup_bpf_enabled(CGROUP_SOCK_OPS).
2026-10-06 19:24 [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
` (3 preceding siblings ...)
2026-10-06 19:24 ` [PATCH v4 bpf-next 04/10] bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag Kuniyuki Iwashima
@ 2026-10-06 19:24 ` Kuniyuki Iwashima
2026-10-06 20:11 ` bot+bpf-ci
` (2 more replies)
2026-10-06 19:24 ` [PATCH v4 bpf-next 06/10] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
` (5 subsequent siblings)
10 siblings, 3 replies; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-06 19:24 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
BPF_CGROUP_RUN_PROG_XXX() macros guard __cgroup_bpf_run_filter_XXX()
with cgroup_bpf_enabled().
However, even when no SOCK_OPS prog is attached, callers still
initialise struct bpf_sock_ops_kern (memset(), etc.) or evaluate
BPF_SOCK_OPS_TEST_FLAG(), which loads tp->bpf_sock_ops_cb_flags
from a cold cacheline near the end of struct tcp_sock.
Similar to bpf_tcp_ops, let's check cgroup_bpf_enabled() before
BPF_SOCK_OPS_TEST_FLAG() and struct bpf_sock_ops_kern setup, and
rename BPF_CGROUP_RUN_PROG_SOCK_OPS{,_SK}() with __ prefix.
Since all direct callers of tcp_call_bpf() pass 0 and NULL for
nargs and args, they can be folded into the new tcp_call_bpf()
macro.
All callers of tcp_call_bpf_{2,3}arg() check BPF_SOCK_OPS_TEST_FLAG()
and do not need the return value. These are replaced with the new
tcp_call_bpf_flag() macro.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
include/linux/bpf-cgroup.h | 36 ++++++++++--------------
include/net/tcp.h | 56 ++++++++++++++++----------------------
net/ipv4/af_inet.c | 2 +-
net/ipv4/tcp.c | 3 +-
net/ipv4/tcp_input.c | 21 ++++++++------
net/ipv4/tcp_nv.c | 2 +-
net/ipv4/tcp_output.c | 31 +++++++++------------
net/ipv4/tcp_timer.c | 6 ++--
8 files changed, 69 insertions(+), 88 deletions(-)
diff --git a/include/linux/bpf-cgroup.h b/include/linux/bpf-cgroup.h
index 8a75a6cd7309..02ef7f899b3d 100644
--- a/include/linux/bpf-cgroup.h
+++ b/include/linux/bpf-cgroup.h
@@ -345,27 +345,19 @@ static inline bool cgroup_bpf_sock_enabled(struct sock *sk,
* calling bpf_setsockopt on listener-sk will not make sense anyway,
* so passing 'sock_ops->sk == req_sk' to the bpf prog is appropriate here.
*/
-#define BPF_CGROUP_RUN_PROG_SOCK_OPS_SK(sock_ops, sk) \
-({ \
- int __ret = 0; \
- if (cgroup_bpf_enabled(CGROUP_SOCK_OPS)) \
- __ret = __cgroup_bpf_run_filter_sock_ops(sk, \
- sock_ops, \
- CGROUP_SOCK_OPS); \
- __ret; \
-})
-
-#define BPF_CGROUP_RUN_PROG_SOCK_OPS(sock_ops) \
-({ \
- int __ret = 0; \
- if (cgroup_bpf_enabled(CGROUP_SOCK_OPS) && (sock_ops)->sk) { \
- typeof(sk) __sk = sk_to_full_sk((sock_ops)->sk); \
- if (__sk && sk_fullsock(__sk)) \
- __ret = __cgroup_bpf_run_filter_sock_ops(__sk, \
- sock_ops, \
- CGROUP_SOCK_OPS); \
- } \
- __ret; \
+#define __BPF_CGROUP_RUN_PROG_SOCK_OPS_SK(sock_ops, sk) \
+ __cgroup_bpf_run_filter_sock_ops(sk, sock_ops, \
+ CGROUP_SOCK_OPS)
+
+#define __BPF_CGROUP_RUN_PROG_SOCK_OPS(sock_ops) \
+({ \
+ int __ret = 0; \
+ typeof(sk) __sk = sk_to_full_sk((sock_ops)->sk); \
+ if (__sk && sk_fullsock(__sk)) \
+ __ret = __cgroup_bpf_run_filter_sock_ops(__sk, \
+ sock_ops, \
+ CGROUP_SOCK_OPS); \
+ __ret; \
})
#define BPF_CGROUP_RUN_PROG_DEVICE_CGROUP(atype, major, minor, access) \
@@ -529,7 +521,7 @@ static inline int cgroup_bpf_struct_ops_attach(struct bpf_map *map,
#define BPF_CGROUP_RUN_PROG_UDP4_RECVMSG_LOCK(sk, uaddr, uaddrlen) ({ 0; })
#define BPF_CGROUP_RUN_PROG_UDP6_RECVMSG_LOCK(sk, uaddr, uaddrlen) ({ 0; })
#define BPF_CGROUP_RUN_PROG_UNIX_RECVMSG_LOCK(sk, uaddr, uaddrlen) ({ 0; })
-#define BPF_CGROUP_RUN_PROG_SOCK_OPS(sock_ops) ({ 0; })
+#define __BPF_CGROUP_RUN_PROG_SOCK_OPS(sock_ops) ({ 0; })
#define BPF_CGROUP_RUN_PROG_DEVICE_CGROUP(atype, major, minor, access) ({ 0; })
#define BPF_CGROUP_RUN_PROG_SYSCTL(head,table,write,buf,count,pos) ({ 0; })
#define BPF_CGROUP_RUN_PROG_GETSOCKOPT(sock, level, optname, optval, \
diff --git a/include/net/tcp.h b/include/net/tcp.h
index b7c0f1a8797a..85b4bfe963d3 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -2891,7 +2891,7 @@ static inline void bpf_skops_init_skb(struct bpf_sock_ops_kern *skops,
* program loaded).
*/
#ifdef CONFIG_BPF
-static inline int tcp_call_bpf(struct sock *sk, int op, u32 nargs, u32 *args)
+static inline int __tcp_call_bpf(struct sock *sk, int op, u32 nargs, u32 *args)
{
struct bpf_sock_ops_kern sock_ops;
int ret;
@@ -2908,7 +2908,7 @@ static inline int tcp_call_bpf(struct sock *sk, int op, u32 nargs, u32 *args)
if (nargs > 0)
memcpy(sock_ops.args, args, nargs * sizeof(*args));
- ret = BPF_CGROUP_RUN_PROG_SOCK_OPS(&sock_ops);
+ ret = __BPF_CGROUP_RUN_PROG_SOCK_OPS(&sock_ops);
if (ret == 0)
ret = sock_ops.reply;
else
@@ -2916,20 +2916,23 @@ static inline int tcp_call_bpf(struct sock *sk, int op, u32 nargs, u32 *args)
return ret;
}
-static inline int tcp_call_bpf_2arg(struct sock *sk, int op, u32 arg1, u32 arg2)
-{
- u32 args[2] = {arg1, arg2};
-
- return tcp_call_bpf(sk, op, 2, args);
-}
-
-static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
- u32 arg3)
-{
- u32 args[3] = {arg1, arg2, arg3};
+#define tcp_call_bpf(sk, op) \
+({ \
+ int __ret = 0; \
+ if (cgroup_bpf_enabled(CGROUP_SOCK_OPS)) { \
+ __ret = __tcp_call_bpf(sk, op, 0, NULL); \
+ } \
+ __ret; \
+})
- return tcp_call_bpf(sk, op, 3, args);
-}
+#define tcp_call_bpf_flag(sk, op, ...) \
+do { \
+ if (cgroup_bpf_enabled(CGROUP_SOCK_OPS) && \
+ BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), op ## _FLAG)) { \
+ u32 __args[] = { __VA_ARGS__ }; \
+ __tcp_call_bpf(sk, op, ARRAY_SIZE(__args), __args); \
+ } \
+} while (0)
static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
{
@@ -2938,21 +2941,12 @@ static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
}
#else
-static inline int tcp_call_bpf(struct sock *sk, int op, u32 nargs, u32 *args)
-{
- return -EPERM;
-}
-
-static inline int tcp_call_bpf_2arg(struct sock *sk, int op, u32 arg1, u32 arg2)
+static inline int tcp_call_bpf(struct sock *sk, int op)
{
return -EPERM;
}
-static inline int tcp_call_bpf_3arg(struct sock *sk, int op, u32 arg1, u32 arg2,
- u32 arg3)
-{
- return -EPERM;
-}
+#define tcp_call_bpf_flag(sk, op, ...) do { } while (0)
static inline void tcp_clear_sock_ops_cb_flags(struct sock *sk)
{
@@ -3136,7 +3130,7 @@ static inline u32 tcp_timeout_init(struct sock *sk)
{
int timeout;
- timeout = tcp_call_bpf(sk, BPF_SOCK_OPS_TIMEOUT_INIT, 0, NULL);
+ timeout = tcp_call_bpf(sk, BPF_SOCK_OPS_TIMEOUT_INIT);
timeout = bpf_tcp_ops_call_int(timeout_init, timeout, sk);
if (timeout <= 0)
timeout = TCP_TIMEOUT_INIT;
@@ -3147,7 +3141,7 @@ static inline u32 tcp_rwnd_init_bpf(struct sock *sk)
{
int rwnd;
- rwnd = tcp_call_bpf(sk, BPF_SOCK_OPS_RWND_INIT, 0, NULL);
+ rwnd = tcp_call_bpf(sk, BPF_SOCK_OPS_RWND_INIT);
rwnd = bpf_tcp_ops_call_int(rwnd_init, rwnd, sk);
if (rwnd < 0)
rwnd = 0;
@@ -3156,14 +3150,12 @@ static inline u32 tcp_rwnd_init_bpf(struct sock *sk)
static inline bool tcp_bpf_ca_needs_ecn(struct sock *sk)
{
- return (tcp_call_bpf(sk, BPF_SOCK_OPS_NEEDS_ECN, 0, NULL) == 1);
+ return (tcp_call_bpf(sk, BPF_SOCK_OPS_NEEDS_ECN) == 1);
}
static inline void tcp_bpf_rtt(struct sock *sk, long mrtt, u32 srtt)
{
- if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_RTT_CB_FLAG))
- tcp_call_bpf_2arg(sk, BPF_SOCK_OPS_RTT_CB, mrtt, srtt);
-
+ tcp_call_bpf_flag(sk, BPF_SOCK_OPS_RTT_CB, mrtt, srtt);
bpf_tcp_ops_call_flag(rtt, RTT, sk, mrtt, srtt);
}
diff --git a/net/ipv4/af_inet.c b/net/ipv4/af_inet.c
index cdcfc7d3c6d2..035c83338e0f 100644
--- a/net/ipv4/af_inet.c
+++ b/net/ipv4/af_inet.c
@@ -226,7 +226,7 @@ int __inet_listen_sk(struct sock *sk, int backlog)
if (err)
return err;
- tcp_call_bpf(sk, BPF_SOCK_OPS_TCP_LISTEN_CB, 0, NULL);
+ tcp_call_bpf(sk, BPF_SOCK_OPS_TCP_LISTEN_CB);
bpf_tcp_ops_call(listen, sk);
}
return 0;
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index fa69961c47d3..5d9d3bcde8f7 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -2994,8 +2994,7 @@ void tcp_set_state(struct sock *sk, int state)
*/
BTF_TYPE_EMIT_ENUM(BPF_TCP_ESTABLISHED);
- if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_STATE_CB_FLAG))
- tcp_call_bpf_2arg(sk, BPF_SOCK_OPS_STATE_CB, oldstate, state);
+ tcp_call_bpf_flag(sk, BPF_SOCK_OPS_STATE_CB, oldstate, state);
bpf_tcp_ops_call(set_state, sk, state);
switch (state) {
diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
index 4478d3f3d4b0..f374257013b1 100644
--- a/net/ipv4/tcp_input.c
+++ b/net/ipv4/tcp_input.c
@@ -146,14 +146,16 @@ EXPORT_SYMBOL_GPL(clean_acked_data_flush);
#ifdef CONFIG_CGROUP_BPF
static void bpf_skops_parse_hdr(struct sock *sk, struct sk_buff *skb)
{
- bool unknown_opt = tcp_sk(sk)->rx_opt.saw_unknown &&
- BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk),
- BPF_SOCK_OPS_PARSE_UNKNOWN_HDR_OPT_CB_FLAG);
- bool parse_all_opt = BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk),
- BPF_SOCK_OPS_PARSE_ALL_HDR_OPT_CB_FLAG);
struct bpf_sock_ops_kern sock_ops;
+ const struct tcp_sock *tp;
+
+ if (!cgroup_bpf_enabled(CGROUP_SOCK_OPS))
+ return;
- if (likely(!unknown_opt && !parse_all_opt))
+ tp = tcp_sk(sk);
+ if (!(tp->rx_opt.saw_unknown &&
+ BPF_SOCK_OPS_TEST_FLAG(tp, BPF_SOCK_OPS_PARSE_UNKNOWN_HDR_OPT_CB_FLAG)) &&
+ !BPF_SOCK_OPS_TEST_FLAG(tp, BPF_SOCK_OPS_PARSE_ALL_HDR_OPT_CB_FLAG))
return;
/* The skb will be handled in the
@@ -176,7 +178,7 @@ static void bpf_skops_parse_hdr(struct sock *sk, struct sk_buff *skb)
sock_ops.sk = sk;
bpf_skops_init_skb(&sock_ops, skb, tcp_hdrlen(skb));
- BPF_CGROUP_RUN_PROG_SOCK_OPS(&sock_ops);
+ __BPF_CGROUP_RUN_PROG_SOCK_OPS(&sock_ops);
}
static void bpf_skops_established(struct sock *sk, int bpf_op,
@@ -184,6 +186,9 @@ static void bpf_skops_established(struct sock *sk, int bpf_op,
{
struct bpf_sock_ops_kern sock_ops;
+ if (!cgroup_bpf_enabled(CGROUP_SOCK_OPS))
+ return;
+
sock_owned_by_me(sk);
memset(&sock_ops, 0, offsetof(struct bpf_sock_ops_kern, temp));
@@ -195,7 +200,7 @@ static void bpf_skops_established(struct sock *sk, int bpf_op,
if (skb)
bpf_skops_init_skb(&sock_ops, skb, tcp_hdrlen(skb));
- BPF_CGROUP_RUN_PROG_SOCK_OPS(&sock_ops);
+ __BPF_CGROUP_RUN_PROG_SOCK_OPS(&sock_ops);
}
#else
static void bpf_skops_parse_hdr(struct sock *sk, struct sk_buff *skb)
diff --git a/net/ipv4/tcp_nv.c b/net/ipv4/tcp_nv.c
index f345897a68df..7b0dae23d9aa 100644
--- a/net/ipv4/tcp_nv.c
+++ b/net/ipv4/tcp_nv.c
@@ -146,7 +146,7 @@ static void tcpnv_init(struct sock *sk)
* within a datacenter, where we have reasonable estimates of
* RTTs
*/
- base_rtt = tcp_call_bpf(sk, BPF_SOCK_OPS_BASE_RTT, 0, NULL);
+ base_rtt = tcp_call_bpf(sk, BPF_SOCK_OPS_BASE_RTT);
if (base_rtt > 0) {
ca->nv_base_rtt = base_rtt;
ca->nv_lower_bound_rtt = (base_rtt * 205) >> 8; /* 80% */
diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c
index 8770f3084efe..db4fd1825d99 100644
--- a/net/ipv4/tcp_output.c
+++ b/net/ipv4/tcp_output.c
@@ -476,8 +476,9 @@ static u32 bpf_skops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,
struct bpf_sock_ops_kern sock_ops;
int err;
- if (likely(!BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk),
- BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG)) ||
+ if (!cgroup_bpf_enabled(CGROUP_SOCK_OPS) ||
+ !BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk),
+ BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG)||
!remaining)
return remaining;
@@ -518,7 +519,7 @@ static u32 bpf_skops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,
if (skb)
bpf_skops_init_skb(&sock_ops, skb, 0);
- err = BPF_CGROUP_RUN_PROG_SOCK_OPS_SK(&sock_ops, sk);
+ err = __BPF_CGROUP_RUN_PROG_SOCK_OPS_SK(&sock_ops, sk);
if (err || sock_ops.remaining_opt_len == remaining)
return remaining;
@@ -543,7 +544,8 @@ static void bpf_skops_write_hdr_opt(struct sock *sk, struct sk_buff *skb,
first_opt_off = tcp_hdrlen(skb) - max_opt_len;
- if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk),
+ if (cgroup_bpf_enabled(CGROUP_SOCK_OPS) &&
+ BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk),
BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG)) {
struct bpf_sock_ops_kern sock_ops;
int err;
@@ -567,7 +569,7 @@ static void bpf_skops_write_hdr_opt(struct sock *sk, struct sk_buff *skb,
sock_ops.remaining_opt_len = max_opt_len;
bpf_skops_init_skb(&sock_ops, skb, first_opt_off);
- err = BPF_CGROUP_RUN_PROG_SOCK_OPS_SK(&sock_ops, sk);
+ err = __BPF_CGROUP_RUN_PROG_SOCK_OPS_SK(&sock_ops, sk);
if (!err)
nr_written = max_opt_len - sock_ops.remaining_opt_len;
}
@@ -1279,16 +1281,10 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
}
}
- if (unlikely(BPF_SOCK_OPS_TEST_FLAG(tp,
- BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG))) {
- remaining = MAX_TCP_OPTION_SPACE - size;
- remaining = bpf_skops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
- remaining);
-
- size = MAX_TCP_OPTION_SPACE - remaining;
- }
-
remaining = MAX_TCP_OPTION_SPACE - size;
+
+ remaining = bpf_skops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
+ remaining);
remaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
remaining);
@@ -3725,9 +3721,8 @@ int __tcp_retransmit_skb(struct sock *sk, struct sk_buff *skb, int segs)
err = tcp_transmit_skb(sk, skb, 1, GFP_ATOMIC);
}
- if (BPF_SOCK_OPS_TEST_FLAG(tp, BPF_SOCK_OPS_RETRANS_CB_FLAG))
- tcp_call_bpf_3arg(sk, BPF_SOCK_OPS_RETRANS_CB,
- TCP_SKB_CB(skb)->seq, segs, err);
+ tcp_call_bpf_flag(sk, BPF_SOCK_OPS_RETRANS_CB,
+ TCP_SKB_CB(skb)->seq, segs, err);
bpf_tcp_ops_call(retrans, sk, skb, err);
if (unlikely(err) && err != -EBUSY)
@@ -4355,7 +4350,7 @@ int tcp_connect(struct sock *sk)
struct sk_buff *buff;
int err;
- tcp_call_bpf(sk, BPF_SOCK_OPS_TCP_CONNECT_CB, 0, NULL);
+ tcp_call_bpf(sk, BPF_SOCK_OPS_TCP_CONNECT_CB);
bpf_tcp_ops_call(connect, sk);
#if defined(CONFIG_TCP_MD5SIG) && defined(CONFIG_TCP_AO)
diff --git a/net/ipv4/tcp_timer.c b/net/ipv4/tcp_timer.c
index 3d49adc51766..00debcba122b 100644
--- a/net/ipv4/tcp_timer.c
+++ b/net/ipv4/tcp_timer.c
@@ -286,10 +286,8 @@ static int tcp_write_timeout(struct sock *sk)
tcp_fastopen_active_detect_blackhole(sk, expired);
mptcp_active_detect_blackhole(sk, expired);
- if (BPF_SOCK_OPS_TEST_FLAG(tp, BPF_SOCK_OPS_RTO_CB_FLAG))
- tcp_call_bpf_3arg(sk, BPF_SOCK_OPS_RTO_CB,
- icsk->icsk_retransmits,
- icsk->icsk_rto, (int)expired);
+ tcp_call_bpf_flag(sk, BPF_SOCK_OPS_RTO_CB,
+ icsk->icsk_retransmits, icsk->icsk_rto, (int)expired);
bpf_tcp_ops_call(rto, sk);
if (expired) {
--
2.56.0.360.g66cac248cb-goog
^ permalink raw reply related [flat|nested] 36+ messages in thread
* [PATCH v4 bpf-next 06/10] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-10-06 19:24 [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
` (4 preceding siblings ...)
2026-10-06 19:24 ` [PATCH v4 bpf-next 05/10] bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after cgroup_bpf_enabled(CGROUP_SOCK_OPS) Kuniyuki Iwashima
@ 2026-10-06 19:24 ` Kuniyuki Iwashima
2026-10-06 22:06 ` Stanislav Fomichev
2026-10-06 23:26 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 07/10] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
` (4 subsequent siblings)
10 siblings, 2 replies; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-06 19:24 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
Waking up a thread per packet is expensive when an application
processes variable-length frames (e.g., RPC) that span multiple
packets.
SO_RCVLOWAT can defer wakeups, but because the frame size is
encoded in a fixed-size descriptor at the start of each frame,
the application has to:
1. wake up and recv() the descriptor,
2. raise SO_RCVLOWAT to the payload size via setsockopt(),
3. wake up and recv() the payload, and
4. reset SO_RCVLOWAT back to the descriptor size via
setsockopt() for the next frame.
This requires an extra wakeup and two setsockopt() syscalls
for every single RPC frame.
With SOCKMAP, we can parse skb and suppress wakeups in kernel,
but SOCKMAP adds overhead and also kills zerocopy.
Let's add lighter-weight opt-in callbacks to bpf_tcp_ops to
replace that.
.enqueue_rcvq(): invoked when TCP stack enqueues skb to
sk->sk_receive_queue
.dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
is dequeued from sk->sk_receive_queue
Those callbacks can be enabled on a per-socket basis by
bpf_tcp_ops_set_flags():
bpf_tcp_ops_set_flags((struct tcp_sock *)sk,
BPF_TCP_OPS_FLAG_RCVQ, 0);
Later, we will add a new kfunc to adjust sk->sk_rcvlowat from
these callbacks.
This will allow the bpf_tcp_ops prog to parse each skb and
dynamically adjust sk->sk_rcvlowat to suppress unnecessary EPOLLIN
wakeups until sufficient data is available in the receive queue.
The placement of bpf_tcp_ops_call() in tcp_ofo_queue() and
tcp_fastopen_add_skb() is chosen to provide the same snapshot
as tcp_queue_rcv().
For example, if bpf_tcp_ops_call() were called before updating
TCP_SKB_CB(skb)->seq in tcp_fastopen_add_skb(), BPF prog would
need an extra branch for the unlikely TFO case to strip SYN.
In addition, the TCP stack can queue overlapping skbs into recvq.
Once rcv_nxt is updated with a new skb, BPF prog can no longer
infer the previous rcv_nxt from skb->len.
Lastly, dequeue_rcvq() is placed in tcp_cleanup_rbuf() rather
than __tcp_cleanup_rbuf() so that it is not called for sockets
in SOCKMAP, where calling sk->sk_data_ready() from the new
kfunc would otherwise trigger infinite recursion.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
v4: Use bpf_tcp_ops_call_flag()
v3: Switch to BPF_TCP_OPS_FLAG_RCVQ
---
Documentation/networking/net_cachelines/tcp_sock.rst | 2 +-
include/net/tcp.h | 12 ++++++++++++
include/uapi/linux/bpf.h | 4 +++-
net/ipv4/bpf_tcp_ops.c | 10 ++++++++++
net/ipv4/tcp.c | 2 ++
net/ipv4/tcp_fastopen.c | 2 ++
net/ipv4/tcp_input.c | 4 ++++
tools/include/uapi/linux/bpf.h | 4 +++-
8 files changed, 37 insertions(+), 3 deletions(-)
diff --git a/Documentation/networking/net_cachelines/tcp_sock.rst b/Documentation/networking/net_cachelines/tcp_sock.rst
index 420fc6278148..e4776905e882 100644
--- a/Documentation/networking/net_cachelines/tcp_sock.rst
+++ b/Documentation/networking/net_cachelines/tcp_sock.rst
@@ -151,7 +151,7 @@ u32 urg_seq
unsigned_int keepalive_time
unsigned_int keepalive_intvl
int linger2
-u32 bpf_tcp_ops_flags read_mostly read_mostly bpf_tcp_ops_hdr_opt_len,bpf_skops_write_hdr_opt(tx);bpf_tcp_ops_parse_hdr,tcp_bpf_rtt(rx);
+u32 bpf_tcp_ops_flags read_mostly read_mostly bpf_tcp_ops_hdr_opt_len,bpf_skops_write_hdr_opt(tx);bpf_tcp_ops_parse_hdr,tcp_bpf_rtt,tcp_cleanup_rbuf,tcp_queue_rcv,tcp_ofo_queue,tcp_fastopen_add_skb(rx);
u8 bpf_sock_ops_cb_flags
u8:1 bpf_chg_cc_inprogress
u16 timeout_rehash
diff --git a/include/net/tcp.h b/include/net/tcp.h
index 85b4bfe963d3..d95cbe4da96e 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -3056,6 +3056,18 @@ struct bpf_tcp_ops {
struct request_sock *req, struct sk_buff *syn_skb,
enum tcp_synack_type synack_type,
u32 opt_off);
+
+ /*
+ * Called when an incoming skb is enqueued to sk->sk_receive_queue
+ * if BPF_TCP_OPS_FLAG_RCVQ is enabled.
+ */
+ void (*enqueue_rcvq)(struct sock *sk, struct sk_buff *skb);
+
+ /*
+ * Called after data is dequeued from sk->sk_receive_queue
+ * if BPF_TCP_OPS_FLAG_RCVQ is enabled.
+ */
+ void (*dequeue_rcvq)(struct sock *sk);
};
#define __bpf_tcp_ops_call(op, sk, ...) \
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 6963c146311e..6f70db7515dd 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -7360,7 +7360,9 @@ enum {
BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_UNKNOWN = (1 << 2),
/* .hdr_opt_len() and .write_hdr_opt() */
BPF_TCP_OPS_FLAG_WRITE_HDR_OPT = (1 << 3),
- BPF_TCP_OPS_FLAG_ALL = (1 << 4) - 1,
+ /* .enqueue_rcvq() and .dequeue_rcvq() */
+ BPF_TCP_OPS_FLAG_RCVQ = (1 << 4),
+ BPF_TCP_OPS_FLAG_ALL = (1 << 5) - 1,
};
/* List of TCP states. There is a build check in net/ipv4/tcp.c to detect
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index aaf96304c69f..9a6e1c47d7c0 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -76,6 +76,14 @@ static void write_hdr_opt_stub(struct sock *sk, struct sk_buff *skb,
{
}
+static void enqueue_rcvq_stub(struct sock *sk, struct sk_buff *skb)
+{
+}
+
+static void dequeue_rcvq_stub(struct sock *sk)
+{
+}
+
static struct bpf_tcp_ops __bpf_tcp_ops = {
.timeout_init = timeout_init_stub,
.rwnd_init = rwnd_init_stub,
@@ -90,6 +98,8 @@ static struct bpf_tcp_ops __bpf_tcp_ops = {
.parse_hdr = parse_hdr_stub,
.hdr_opt_len = hdr_opt_len_stub,
.write_hdr_opt = write_hdr_opt_stub,
+ .enqueue_rcvq = enqueue_rcvq_stub,
+ .dequeue_rcvq = dequeue_rcvq_stub,
};
BPF_CALL_4(bpf_tcp_ops_store_hdr_opt, void *, ctx, const void *, from,
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index 5d9d3bcde8f7..5d907a0a0461 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -1609,6 +1609,8 @@ void tcp_cleanup_rbuf(struct sock *sk, int copied)
"cleanup rbuf bug: copied %X seq %X rcvnxt %X\n",
tp->copied_seq, TCP_SKB_CB(skb)->end_seq, tp->rcv_nxt);
__tcp_cleanup_rbuf(sk, copied);
+
+ bpf_tcp_ops_call_flag(dequeue_rcvq, RCVQ, sk);
}
static void tcp_eat_recv_skb(struct sock *sk, struct sk_buff *skb)
diff --git a/net/ipv4/tcp_fastopen.c b/net/ipv4/tcp_fastopen.c
index 471c78be5513..6dfe40322fb5 100644
--- a/net/ipv4/tcp_fastopen.c
+++ b/net/ipv4/tcp_fastopen.c
@@ -281,6 +281,8 @@ void tcp_fastopen_add_skb(struct sock *sk, struct sk_buff *skb)
TCP_SKB_CB(skb)->seq++;
TCP_SKB_CB(skb)->tcp_flags &= ~TCPHDR_SYN;
+ bpf_tcp_ops_call_flag(enqueue_rcvq, RCVQ, sk, skb);
+
tp->rcv_nxt = TCP_SKB_CB(skb)->end_seq;
tcp_add_receive_queue(sk, skb);
tp->syn_data_acked = 1;
diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
index f374257013b1..46f5c4fafa8d 100644
--- a/net/ipv4/tcp_input.c
+++ b/net/ipv4/tcp_input.c
@@ -5365,6 +5365,8 @@ static void tcp_ofo_queue(struct sock *sk)
continue;
}
+ bpf_tcp_ops_call_flag(enqueue_rcvq, RCVQ, sk, skb);
+
tail = skb_peek_tail(&sk->sk_receive_queue);
eaten = tail && tcp_try_coalesce(sk, tail, skb, &fragstolen);
tcp_rcv_nxt_update(tp, TCP_SKB_CB(skb)->end_seq);
@@ -5568,6 +5570,8 @@ static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb,
int eaten;
struct sk_buff *tail = skb_peek_tail(&sk->sk_receive_queue);
+ bpf_tcp_ops_call_flag(enqueue_rcvq, RCVQ, sk, skb);
+
eaten = (tail &&
tcp_try_coalesce(sk, tail,
skb, fragstolen)) ? 1 : 0;
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 6963c146311e..6f70db7515dd 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -7360,7 +7360,9 @@ enum {
BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_UNKNOWN = (1 << 2),
/* .hdr_opt_len() and .write_hdr_opt() */
BPF_TCP_OPS_FLAG_WRITE_HDR_OPT = (1 << 3),
- BPF_TCP_OPS_FLAG_ALL = (1 << 4) - 1,
+ /* .enqueue_rcvq() and .dequeue_rcvq() */
+ BPF_TCP_OPS_FLAG_RCVQ = (1 << 4),
+ BPF_TCP_OPS_FLAG_ALL = (1 << 5) - 1,
};
/* List of TCP states. There is a build check in net/ipv4/tcp.c to detect
--
2.56.0.360.g66cac248cb-goog
^ permalink raw reply related [flat|nested] 36+ messages in thread
* [PATCH v4 bpf-next 07/10] tcp: Split out __tcp_set_rcvlowat().
2026-10-06 19:24 [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
` (5 preceding siblings ...)
2026-10-06 19:24 ` [PATCH v4 bpf-next 06/10] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
@ 2026-10-06 19:24 ` Kuniyuki Iwashima
2026-10-07 22:43 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 08/10] bpf: mptcp: Don't support BPF_TCP_OPS_FLAG_RCVQ Kuniyuki Iwashima
` (3 subsequent siblings)
10 siblings, 1 reply; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-06 19:24 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev, Emil Tsalapatis
We will add a kfunc for bpf_tcp_ops.{enqueue,dequeue}_rcvq()
to adjust sk->sk_rcvlowat.
These hooks are triggered
* when the TCP stack enqueues an skb to sk->sk_receive_queue
* after data is dequeued from sk->sk_receive_queue
In the enqueue path, tcp_data_ready() is always called after
the hooks in tcp_queue_rcv() and tcp_ofo_queue().
If tcp_set_rcvlowat() were used as is, tcp_data_ready() could
be called twice for the same skb, which is redundant and also
confusing.
Let's split out __tcp_set_rcvlowat() and add a flag to control
wakeup behaviour.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
---
include/net/tcp.h | 1 +
net/ipv4/tcp.c | 12 +++++++++---
2 files changed, 10 insertions(+), 3 deletions(-)
diff --git a/include/net/tcp.h b/include/net/tcp.h
index d95cbe4da96e..f7624ab9b420 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -512,6 +512,7 @@ void tcp_set_keepalive(struct sock *sk, int val);
void tcp_syn_ack_timeout(const struct request_sock *req);
int tcp_recvmsg(struct sock *sk, struct msghdr *msg, size_t len,
int flags);
+int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup);
int tcp_set_rcvlowat(struct sock *sk, int val);
void tcp_set_rcvbuf(struct sock *sk, int val);
int tcp_set_window_clamp(struct sock *sk, int val);
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index 5d907a0a0461..c3c756a1ba8c 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -1827,8 +1827,7 @@ int tcp_peek_len(struct socket *sock)
return tcp_inq(sock->sk);
}
-/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */
-int tcp_set_rcvlowat(struct sock *sk, int val)
+int __tcp_set_rcvlowat(struct sock *sk, int val, bool wakeup)
{
struct tcp_sock *tp = tcp_sk(sk);
int space, cap;
@@ -1841,7 +1840,8 @@ int tcp_set_rcvlowat(struct sock *sk, int val)
WRITE_ONCE(sk->sk_rcvlowat, val ? : 1);
/* Check if we need to signal EPOLLIN right now */
- tcp_data_ready(sk);
+ if (wakeup)
+ tcp_data_ready(sk);
if (sk->sk_userlocks & SOCK_RCVBUF_LOCK)
return 0;
@@ -1856,6 +1856,12 @@ int tcp_set_rcvlowat(struct sock *sk, int val)
return 0;
}
+/* Make sure sk_rcvbuf is big enough to satisfy SO_RCVLOWAT hint */
+int tcp_set_rcvlowat(struct sock *sk, int val)
+{
+ return __tcp_set_rcvlowat(sk, val, true);
+}
+
void tcp_set_rcvbuf(struct sock *sk, int val)
{
tcp_set_window_clamp(sk, tcp_win_from_space(sk, val));
--
2.56.0.360.g66cac248cb-goog
^ permalink raw reply related [flat|nested] 36+ messages in thread
* [PATCH v4 bpf-next 08/10] bpf: mptcp: Don't support BPF_TCP_OPS_FLAG_RCVQ.
2026-10-06 19:24 [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
` (6 preceding siblings ...)
2026-10-06 19:24 ` [PATCH v4 bpf-next 07/10] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
@ 2026-10-06 19:24 ` Kuniyuki Iwashima
2026-10-06 22:06 ` Stanislav Fomichev
2026-10-07 22:43 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 09/10] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
` (2 subsequent siblings)
10 siblings, 2 replies; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-06 19:24 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
The next patch exposes a new kfunc calling __tcp_set_rcvlowat()
to bpf_tcp_ops.
MPTCP has its own sock->ops->set_rcvlowat() / mptcp_set_rcvlowat(),
so we should not allow calling __tcp_set_rcvlowat() on MPTCP
subflows.
Let's disable BPF_TCP_OPS_FLAG_RCVQ for MPTCP for now.
If needed in the future, bpf_tcp_ops_set_rcvlowat() could be
extended to properly support MPTCP.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
net/ipv4/bpf_tcp_ops.c | 3 +++
1 file changed, 3 insertions(+)
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index 9a6e1c47d7c0..8182037c4269 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -347,6 +347,9 @@ __bpf_kfunc int bpf_tcp_ops_set_flags(struct tcp_sock *tp, u32 enable, u32 disab
if ((enable & disable) || (enable | disable) & ~BPF_TCP_OPS_FLAG_ALL)
return -EINVAL;
+ if (sk_is_mptcp((struct sock *)tp) && (enable & BPF_TCP_OPS_FLAG_RCVQ))
+ return -EOPNOTSUPP;
+
old = READ_ONCE(tp->bpf_tcp_ops_flags);
do {
--
2.56.0.360.g66cac248cb-goog
^ permalink raw reply related [flat|nested] 36+ messages in thread
* [PATCH v4 bpf-next 09/10] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
2026-10-06 19:24 [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
` (7 preceding siblings ...)
2026-10-06 19:24 ` [PATCH v4 bpf-next 08/10] bpf: mptcp: Don't support BPF_TCP_OPS_FLAG_RCVQ Kuniyuki Iwashima
@ 2026-10-06 19:24 ` Kuniyuki Iwashima
2026-10-06 19:42 ` sashiko-bot
2026-10-07 23:06 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 10/10] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-10-07 13:11 ` [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Jakub Sitnicki
10 siblings, 2 replies; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-06 19:24 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev, Emil Tsalapatis
bpf_tcp_ops.{enqueue,dequeue}_rcvq() were added to parse skb and
adjust sk->sk_rcvlowat dynamically to suppress unnecessary wakeups.
Let's add a new kfunc to set sk->sk_rcvlowat.
Negative values are clamped to INT_MAX, consistent with SO_RCVLOWAT.
For enqueue_rcvq(), wakeup is set to false because:
* tcp_data_ready() is always called after the hooks in
tcp_queue_rcv() and tcp_ofo_queue().
* when tcp_fastopen_add_skb() is called for TFO SYN, the socket is
not yet accept()ed, and when called for TFO SYN+ACK, the socket
is woken up by sk->sk_state_change() anyway.
For dequeue_rcvq(), wakeup is set to true because tcp_data_ready()
is not called in that path.
An alternative would be to support bpf_setsockopt() for these
hooks.
However, that approach involves excessive conditionals and an
unnecessary memcpy(), costs we do not want to pay for every skb
in the TCP fast path.
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Tested-by: Clément Léger <cleger@meta.com>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
---
net/ipv4/bpf_tcp_ops.c | 46 ++++++++++++++++++++++++++++++++++++++++++
1 file changed, 46 insertions(+)
diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
index 8182037c4269..2ba73dd6c52c 100644
--- a/net/ipv4/bpf_tcp_ops.c
+++ b/net/ipv4/bpf_tcp_ops.c
@@ -361,12 +361,31 @@ __bpf_kfunc int bpf_tcp_ops_set_flags(struct tcp_sock *tp, u32 enable, u32 disab
return 0;
}
+__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
+ const struct bpf_prog_aux *aux)
+{
+ u32 moff = aux->attach_st_ops_member_off;
+ bool wakeup = false;
+
+ if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+ wakeup = true;
+
+ if (rcvlowat < 0)
+ rcvlowat = INT_MAX;
+
+ return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);
+}
+
__bpf_kfunc_end_defs();
BTF_KFUNCS_START(bpf_tcp_ops_set_flags_kfunc_set)
BTF_ID_FLAGS(func, bpf_tcp_ops_set_flags)
BTF_KFUNCS_END(bpf_tcp_ops_set_flags_kfunc_set)
+BTF_KFUNCS_START(bpf_tcp_ops_set_rcvlowat_kfunc_set)
+BTF_ID_FLAGS(func, bpf_tcp_ops_set_rcvlowat, KF_IMPLICIT_ARGS)
+BTF_KFUNCS_END(bpf_tcp_ops_set_rcvlowat_kfunc_set)
+
static int bpf_tcp_ops_set_flags_kfunc_filter(const struct bpf_prog *prog,
u32 kfunc_id)
{
@@ -390,6 +409,31 @@ static const struct btf_kfunc_id_set bpf_tcp_ops_set_flags_kfunc_id_set = {
.filter = bpf_tcp_ops_set_flags_kfunc_filter,
};
+static int bpf_tcp_ops_set_rcvlowat_kfunc_filter(const struct bpf_prog *prog,
+ u32 kfunc_id)
+{
+ u32 moff;
+
+ if (!btf_id_set8_contains(&bpf_tcp_ops_set_rcvlowat_kfunc_set, kfunc_id))
+ return 0;
+
+ if (prog->aux->st_ops != &bpf_tcp_ops)
+ return -EACCES;
+
+ moff = prog->aux->attach_st_ops_member_off;
+ if (moff != offsetof(struct bpf_tcp_ops, enqueue_rcvq) &&
+ moff != offsetof(struct bpf_tcp_ops, dequeue_rcvq))
+ return -EACCES;
+
+ return 0;
+}
+
+static const struct btf_kfunc_id_set bpf_tcp_ops_set_rcvlowat_kfunc_id_set = {
+ .owner = THIS_MODULE,
+ .set = &bpf_tcp_ops_set_rcvlowat_kfunc_set,
+ .filter = bpf_tcp_ops_set_rcvlowat_kfunc_filter,
+};
+
static int __init __bpf_tcp_ops_init(void)
{
int ret;
@@ -399,6 +443,8 @@ static int __init __bpf_tcp_ops_init(void)
/* BPF_PROG_TYPE_CGROUP_{SOCKOPT,SKB} share BTF_KFUNC_HOOK_CGROUP. */
ret = ret ?: register_btf_kfunc_id_set(BPF_PROG_TYPE_CGROUP_SOCKOPT,
&bpf_tcp_ops_set_flags_kfunc_id_set);
+ ret = ret ?: register_btf_kfunc_id_set(BPF_PROG_TYPE_STRUCT_OPS,
+ &bpf_tcp_ops_set_rcvlowat_kfunc_id_set);
ret = ret ?: register_bpf_struct_ops(&bpf_tcp_ops, bpf_tcp_ops);
return ret;
--
2.56.0.360.g66cac248cb-goog
^ permalink raw reply related [flat|nested] 36+ messages in thread
* [PATCH v4 bpf-next 10/10] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-10-06 19:24 [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
` (8 preceding siblings ...)
2026-10-06 19:24 ` [PATCH v4 bpf-next 09/10] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
@ 2026-10-06 19:24 ` Kuniyuki Iwashima
2026-10-06 19:43 ` sashiko-bot
2026-10-06 22:06 ` Stanislav Fomichev
2026-10-07 13:11 ` [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Jakub Sitnicki
10 siblings, 2 replies; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-06 19:24 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, Kuniyuki Iwashima,
bpf, netdev
The test is roughly divided into two stages, and the sequence
is as follows:
I) Setup
1. Attach two BPF programs to a cgroup
2. Establish a TCP connection (@client <-> @child) within the cgroup
3. Enable BPF_TCP_OPS_FLAG_RCVQ on @child via setsockopt()
II) RPC frame exchange in various patterns
4. Send a partial RPC descriptor from @client to @child
5. Verify that epoll does NOT wake up @child
6. Send the remaining data of the RPC frame
7. Verify that epoll finally wakes up @child
During setup, two BPF programs are attached to simulate
a real-world scenario; one is bpf_tcp_ops and the other is
CGROUP_SOCKOPT.
While the bpf_tcp_ops prog handles the dynamic adjustment of
sk->sk_rcvlowat, the CGROUP_SOCKOPT prog is used to enable
the TCP AutoLOWAT feature via userspace setsockopt() using
pseudo options:
#define SOL_BPF 0xdeadbeef
#define BPF_TCP_AUTOLOWAT 0x8badf00d
setsockopt(fd, SOL_BPF, BPF_TCP_AUTOLOWAT, &(int){1}, sizeof(int));
This reflects a common production use case where an application
decides to start parsing RPC frames only at a certain point in
the stream (e.g., after HTTP Upgrade), rather than immediately
after TCP 3WHS (->passive_established(), etc).
When BPF_TCP_AUTOLOWAT is enabled, the BPF prog sets
BPF_TCP_OPS_FLAG_RCVQ and initializes sk_local_storage
for two sequence numbers to manage its state.
Then, for the RPC frame exchange, this test uses a simple format
defined as follows:
0 8 16 24 32
+--------+--------+-------+--------+ `.
| header size | |
+--------+--------+-------+--------+ > RPC descriptor (8 bytes)
| payload size | |
+--------+--------+-------+--------+ .'
~ header ~
+--------+--------+-------+--------+
~ payload ~
+--------+--------+-------+--------+
Every time a new skb is enqueued to sk->sk_receive_queue, the
bpf_tcp_ops prog parses it and updates these sequence numbers:
rpc_desc_seq : the SEQ # of the start of the RPC descriptor
rpc_end_seq : the SEQ # of the end of the RPC frame
=> rpc_desc_seq + 8 + header size + payload size
Assume we receive two RPC descriptors in the following pattern:
1. When we receive skb-1, only part of the RPC descriptor is parsed.
rpc_desc_seq is set to the first byte while rpc_end_seq is
unknown. Thus, sk->sk_rcvlowat is set to the size of the RPC
descriptor (8 bytes).
<- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
+-----------+.................+....................+......
| RPC desc 1 | header + payload | RPC desc 2 | ...
+-----------+.................+....................+......
^ ^-.
`- rpc_desc_seq `- sk->sk_rcvlowat
2. Next, we receive skb-2, which completes the first RPC descriptor.
Now rpc_end_seq is known, so sk->sk_rcvlowat is advanced to it.
<- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
+-----------+-----------------+....................+......
| RPC desc 1 | header + payload | RPC desc 2 | ...
+-----------+-----------------+....................+......
^ ^
'- rpc_desc_seq '- rpc_end_seq
& sk->sk_rcvlowat
3. Once we receive skb-3, which contains the next full RPC descriptor,
rpc_desc_seq is advanced and rpc_end_seq is updated according
to the size of RPC frame 2.
Note that sk->sk_rcvlowat is NOT updated to the new rpc_end_seq
yet. This ensures that the application is woken up to read the
already complete RPC frame 1.
<- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
+-----------+-----------------+--------------------+......
| RPC desc 1 | header + payload | RPC desc 2 | ... |
+-----------+-----------------+--------------------+......
^ ^
rpc_desc_seq -----------' rpc_end_seq ----...-'
& sk->sk_rcvlowat
This sequence corresponds to the 4th test case in rpc_test_cases[],
and we can see helpful output if we "#define DEBUG":
# cat /sys/kernel/tracing/trace_pipe | \
awk '{ if ($0 ~ /AF_/) sub(/^.*AF_/, "AF_"); print $0 }' & \
BGPID=$!; ./test_progs -t tcp_autolowat; kill -9 -$BGPID
...
AF_INET6 rpc_test_cases[3]: Start parsing skb: seq: 0, end_seq: 1, len: 1, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_buff_len: 0
AF_INET6 rpc_test_cases[3]: Copied 1 bytes: rpc_desc_buff_len: 1
AF_INET6 rpc_test_cases[3]: Setting rcvlowat: tp->copied_seq: 0, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_desc_buff_len: 1
AF_INET6 rpc_test_cases[3]: Set rcvlowat: expected: 8, actual: 8
AF_INET6 rpc_test_cases[3]: Start parsing skb: seq: 1, end_seq: 8, len: 7, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_buff_len: 1
AF_INET6 rpc_test_cases[3]: Copied full descriptor: rpc_desc_seq: 0, rpc_end_seq: 258, header_len: 100, payload_len: 150
AF_INET6 rpc_test_cases[3]: No more descriptor: rpc_end_seq: 258, end_seq: 8
AF_INET6 rpc_test_cases[3]: Setting rcvlowat: tp->copied_seq: 0, rpc_desc_seq: 0, rpc_end_seq: 258, rpc_desc_buff_len: 8
AF_INET6 rpc_test_cases[3]: Set rcvlowat: expected: 258, actual: 258
...
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
v3:
* Use bpf_tcp_ops_set_flags()
v2:
* Avoid address comparison for a specific version of gcc.
* Make rpc_test_cases[] static.
* Update comment in rpc_test_case[].
---
.../selftests/bpf/prog_tests/tcp_autolowat.c | 350 ++++++++++++++++++
.../selftests/bpf/progs/bpf_tracing_net.h | 2 +
.../selftests/bpf/progs/tcp_autolowat.c | 294 +++++++++++++++
3 files changed, 646 insertions(+)
create mode 100644 tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
create mode 100644 tools/testing/selftests/bpf/progs/tcp_autolowat.c
diff --git a/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
new file mode 100644
index 000000000000..337f9d34a39c
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/tcp_autolowat.c
@@ -0,0 +1,350 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright 2026 Google LLC */
+#include <sys/epoll.h>
+
+#include "test_progs.h"
+#include "cgroup_helpers.h"
+#include "network_helpers.h"
+
+#include "tcp_autolowat.skel.h"
+
+#define SOL_BPF 0xdeadbeef
+#define BPF_TCP_AUTOLOWAT 0x8badf00d
+
+struct rpc_descriptor {
+ u32 header_len;
+ u32 payload_len;
+};
+
+enum rpc_event_type {
+ RPC_EVENT_END,
+ RPC_EVENT_AUTOLOWAT,
+ RPC_EVENT_SEND,
+ RPC_EVENT_RECV,
+ RPC_EVENT_EPOLL,
+ RPC_EVENT_RCVLOWAT,
+};
+
+struct rpc_event {
+ enum rpc_event_type type;
+ union {
+ int len;
+ int nfds;
+ int val;
+ int rcvlowat;
+ };
+};
+
+#define RPC_DESC_SIZE (sizeof(struct rpc_descriptor))
+
+static struct rpc_test_case {
+ char data[4096];
+ struct rpc_descriptor desc[32];
+ struct rpc_event event[32];
+} rpc_test_cases[] = {
+ {
+ .desc = {
+ { .header_len = 100, .payload_len = 150 },
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* Single full RPC message in skb. */
+ { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE + 100 + 150},
+ { .type = RPC_EVENT_EPOLL, .nfds = 1},
+ { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 100 + 150},
+ },
+ },
+ {
+ .desc = {
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* Two full RPC messages in skb. */
+ {.type = RPC_EVENT_SEND, .len = (RPC_DESC_SIZE + 100 + 150) * 2},
+ {.type = RPC_EVENT_EPOLL, .nfds = 1},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+ /* Single full RPC message in skb. */
+ { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE + 100 + 150},
+ { .type = RPC_EVENT_EPOLL, .nfds = 1},
+ { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 3},
+ },
+ },
+ {
+ .desc = {
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 100, .payload_len = 150},
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* Two full RPC messages in skb. */
+ {.type = RPC_EVENT_SEND, .len = (RPC_DESC_SIZE + 100 + 150) * 2},
+ {.type = RPC_EVENT_EPOLL, .nfds = 1},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+ /* Only the next descriptor in skb. */
+ { .type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE},
+ { .type = RPC_EVENT_EPOLL, .nfds = 1},
+ { .type = RPC_EVENT_RCVLOWAT, .rcvlowat = (RPC_DESC_SIZE + 100 + 150) * 2},
+ },
+ },
+ {
+ .desc = {
+ {.header_len = 100, .payload_len = 150},
+ {.header_len = 200, .payload_len = 500},
+ },
+ .event = {
+ { .type = RPC_EVENT_AUTOLOWAT, .val = 1},
+ /* The first descriptor is partial. */
+ {.type = RPC_EVENT_SEND, .len = 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE},
+ /* The first descriptor is available. */
+ {.type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE - 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 150 + 100},
+ /* The first header is ready. */
+ {.type = RPC_EVENT_SEND, .len = 100},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 150 + 100},
+ /* skb has the first payload and 1 byte of the next descriptor. */
+ {.type = RPC_EVENT_SEND, .len = 150 + 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 1},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 150 + 100},
+ /* After reading the first RPC message, SO_RCVLOWAT should be RPC_DESC_SIZE. */
+ {.type = RPC_EVENT_RECV, .len = RPC_DESC_SIZE + 150 + 100},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE},
+ /* The second descriptor is available. */
+ {.type = RPC_EVENT_SEND, .len = RPC_DESC_SIZE - 1},
+ {.type = RPC_EVENT_EPOLL, .nfds = 0},
+ {.type = RPC_EVENT_RCVLOWAT, .rcvlowat = RPC_DESC_SIZE + 200 + 500},
+ },
+ },
+};
+
+struct tcp_autolowat_test_cb {
+ int saved_netns;
+ union {
+ int fd[4];
+ struct {
+ int server, client, child;
+ int epoll;
+ };
+ };
+};
+
+static void tcp_autolowat_teardown_cb(struct tcp_autolowat_test_cb *cb)
+{
+ int i, err;
+
+ for (i = 0; i < ARRAY_SIZE(cb->fd); i++) {
+ if (cb->fd[i] != -1)
+ close(cb->fd[i]);
+ }
+
+ if (cb->saved_netns != -1) {
+ err = setns(cb->saved_netns, CLONE_NEWNET);
+ ASSERT_OK(err, "restore netns");
+
+ close(cb->saved_netns);
+ }
+}
+
+static int tcp_autolowat_setup_cb(struct tcp_autolowat_test_cb *cb, int family)
+{
+ struct epoll_event ev = {};
+ int err;
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(cb->fd); i++)
+ cb->fd[i] = -1;
+
+ cb->saved_netns = open("/proc/self/ns/net", O_RDONLY);
+ if (!ASSERT_OK_FD(cb->saved_netns, "save netns"))
+ goto err;
+
+ err = unshare(CLONE_NEWNET);
+ if (!ASSERT_OK(err, "unshare"))
+ goto err;
+
+ err = system("ip link set dev lo up");
+ if (!ASSERT_OK(err, "set up lo"))
+ goto err;
+
+ cb->server = start_server(family, SOCK_STREAM, NULL, 0, 0);
+ if (!ASSERT_OK_FD(cb->server, "start_server"))
+ goto err;
+
+ cb->client = connect_to_fd(cb->server, 0);
+ if (!ASSERT_OK_FD(cb->client, "connect_to_fd"))
+ goto err;
+
+ cb->child = accept(cb->server, NULL, NULL);
+ if (!ASSERT_OK_FD(cb->child, "accept"))
+ goto err;
+
+ cb->epoll = epoll_create1(0);
+ if (!ASSERT_OK_FD(cb->epoll, "epoll_create"))
+ goto err;
+
+ ev.events = EPOLLIN;
+ ev.data.fd = cb->child;
+
+ err = epoll_ctl(cb->epoll, EPOLL_CTL_ADD, cb->child, &ev);
+ if (!ASSERT_OK(err, "epoll_ctl"))
+ goto err;
+
+ return 0;
+
+err:
+ tcp_autolowat_teardown_cb(cb);
+ return -1;
+}
+
+static int tcp_autolowat_build_data(struct rpc_test_case *test_case)
+{
+ struct rpc_descriptor *desc = test_case->desc;
+ char *ptr = test_case->data;
+ int rpc_size;
+
+ memset(ptr, 0, sizeof(test_case->data));
+
+ while (desc->header_len + desc->payload_len) {
+ rpc_size = sizeof(*desc) + desc->header_len + desc->payload_len;
+
+ if (!ASSERT_LE(ptr + rpc_size - test_case->data,
+ sizeof(test_case->data), "data overflow"))
+ return 1;
+
+ memcpy(ptr, desc, sizeof(*desc));
+ ptr += rpc_size;
+ desc++;
+ }
+
+ if (!ASSERT_GT(ptr - test_case->data, 0, "no data"))
+ return 1;
+
+ return 0;
+}
+
+static void tcp_autolowat_run_rpc_test(struct tcp_autolowat_test_cb *cb,
+ struct rpc_test_case *test_case)
+{
+ struct rpc_event *event = test_case->event;
+ char *ptr = test_case->data;
+ struct epoll_event ev;
+ socklen_t optlen;
+ int err, optval;
+ char buf[4096];
+
+ if (tcp_autolowat_build_data(test_case))
+ return;
+
+ while (1) {
+ switch (event->type) {
+ case RPC_EVENT_END:
+ return;
+ case RPC_EVENT_AUTOLOWAT:
+ err = setsockopt(cb->child, SOL_BPF, BPF_TCP_AUTOLOWAT,
+ &event->val, sizeof(event->val));
+ if (!ASSERT_OK(err, "setsockopt"))
+ return;
+ break;
+ case RPC_EVENT_SEND:
+ err = send(cb->client, ptr, event->len, 0);
+ if (!ASSERT_EQ(err, event->len, "send"))
+ return;
+
+ ptr += event->len;
+ break;
+ case RPC_EVENT_RECV:
+ err = recv(cb->child, buf, event->len, 0);
+ if (!ASSERT_EQ(err, event->len, "recv"))
+ return;
+ break;
+ case RPC_EVENT_EPOLL:
+ err = epoll_wait(cb->epoll, &ev, 1, 100);
+ if (!ASSERT_EQ(err, event->nfds, "epoll_wait"))
+ return;
+ break;
+ case RPC_EVENT_RCVLOWAT:
+ optval = 0;
+ optlen = sizeof(optval);
+
+ err = getsockopt(cb->child, SOL_SOCKET, SO_RCVLOWAT, &optval, &optlen);
+ if (!ASSERT_OK(err, "getsockopt") ||
+ !ASSERT_EQ(optval, event->rcvlowat, "rcvlowat"))
+ return;
+ break;
+ }
+
+ event++;
+ }
+}
+
+static void tcp_autolowat_run_rpc_tests(struct tcp_autolowat *skel, int family)
+{
+ struct tcp_autolowat_test_cb cb;
+ int err;
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(rpc_test_cases); i++) {
+ memset(skel->bss->test_name, 0, sizeof(skel->bss->test_name));
+
+ snprintf(skel->bss->test_name, sizeof(skel->bss->test_name),
+ "AF_INET%c rpc_test_cases[%d]",
+ family == AF_INET ? ' ' : '6', i);
+
+ if (!test__start_subtest(skel->bss->test_name))
+ continue;
+
+ err = tcp_autolowat_setup_cb(&cb, family);
+ if (err)
+ continue;
+
+ tcp_autolowat_run_rpc_test(&cb, &rpc_test_cases[i]);
+ tcp_autolowat_teardown_cb(&cb);
+ }
+}
+
+static void tcp_autolowat_run_tests(struct tcp_autolowat *skel)
+{
+ tcp_autolowat_run_rpc_tests(skel, AF_INET);
+ tcp_autolowat_run_rpc_tests(skel, AF_INET6);
+}
+
+void test_tcp_autolowat(void)
+{
+ struct tcp_autolowat *skel;
+ struct bpf_link *link[2];
+ int cgroup;
+
+ skel = tcp_autolowat__open_and_load();
+ if (!ASSERT_OK_PTR(skel, "open_and_load"))
+ return;
+
+ cgroup = test__join_cgroup("/tcp_autolowat");
+ if (!ASSERT_GE(cgroup, 0, "join_cgroup"))
+ goto destroy_skel;
+
+ link[0] = bpf_map__attach_cgroup_opts(skel->maps.tcp_autolowat_ops, cgroup, NULL);
+ if (!ASSERT_OK_PTR(link[0], "attach_cgroup(tcp_autolowat_ops)"))
+ goto close_cgroup;
+
+ link[1] = bpf_program__attach_cgroup(skel->progs.tcp_autolowat_setsockopt, cgroup);
+ if (!ASSERT_OK_PTR(link[1], "attach_cgroup(SETSOCKOPT)"))
+ goto destroy_sockops;
+
+ tcp_autolowat_run_tests(skel);
+
+ bpf_link__destroy(link[1]);
+destroy_sockops:
+ bpf_link__destroy(link[0]);
+close_cgroup:
+ close(cgroup);
+destroy_skel:
+ tcp_autolowat__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/progs/bpf_tracing_net.h b/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
index 593b38f90417..4c999d59cbbc 100644
--- a/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
+++ b/tools/testing/selftests/bpf/progs/bpf_tracing_net.h
@@ -79,6 +79,8 @@
#define NEXTHDR_TCP 6
+#define TCPHDR_FIN 0x01
+
#define TCPOPT_NOP 1
#define TCPOPT_EOL 0
#define TCPOPT_MSS 2
diff --git a/tools/testing/selftests/bpf/progs/tcp_autolowat.c b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
new file mode 100644
index 000000000000..8a55f3cee260
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
@@ -0,0 +1,294 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright 2026 Google LLC */
+#include "vmlinux.h"
+
+#include <string.h>
+#include <limits.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+#include <bpf/bpf_core_read.h>
+
+#include "bpf_tracing_net.h"
+
+#define SOL_BPF 0xdeadbeef
+#define BPF_TCP_AUTOLOWAT 0x8badf00d
+
+//#define DEBUG /* For verbose output. */
+
+struct rpc_descriptor {
+ u32 header_len;
+ u32 payload_len;
+};
+
+#define RPC_DESC_SIZE (sizeof(struct rpc_descriptor))
+#define MAX_RPC_DESC_PER_SKB 100
+
+struct tcp_autolowat_cb {
+ /* Don't put this field at the end; BPF verifier complains. */
+ char rpc_desc_buf[RPC_DESC_SIZE];
+ u32 rpc_desc_seq;
+ u32 rpc_end_seq;
+#ifdef DEBUG
+ u32 isn;
+#endif
+ u8 rpc_desc_buff_len;
+};
+
+struct {
+ __uint(type, BPF_MAP_TYPE_SK_STORAGE);
+ __uint(map_flags, BPF_F_NO_PREALLOC);
+ __type(key, int);
+ __type(value, struct tcp_autolowat_cb);
+} tcp_autolowat_map SEC(".maps");
+
+char test_name[64];
+
+#ifdef DEBUG
+#define LOG(str, ...) \
+ bpf_printk("%s: " str, test_name, ##__VA_ARGS__)
+#else
+#define LOG(...)
+#endif
+
+#define SEQ(val) \
+ (val - cb->isn)
+#define TP_SEQ(field) \
+ (tp->field - cb->isn)
+#define CB_SEQ(field) \
+ (cb->field - cb->isn)
+
+static int tcp_parse_descriptor(struct tcp_autolowat_cb *cb,
+ struct bpf_dynptr *dptr,
+ u32 seq, u32 end_seq)
+{
+ struct rpc_descriptor *rpc_desc;
+ u32 rpc_copied_seq;
+ u64 copy_len; /* u32 should work, but not for no_alu32 :/ */
+ u64 rpc_len;
+ int err;
+
+ rpc_copied_seq = cb->rpc_desc_seq + cb->rpc_desc_buff_len;
+
+ if (before(cb->rpc_desc_seq + RPC_DESC_SIZE, end_seq))
+ copy_len = RPC_DESC_SIZE - cb->rpc_desc_buff_len;
+ else
+ copy_len = end_seq - rpc_copied_seq;
+
+ if (copy_len == 0)
+ goto disable; /* FIN. */
+ if (copy_len > RPC_DESC_SIZE)
+ goto disable; /* always false, only for verifier. */
+ if (cb->rpc_desc_buff_len >= RPC_DESC_SIZE)
+ goto disable; /* always false, only for verifier. */
+
+ err = bpf_dynptr_read(cb->rpc_desc_buf + cb->rpc_desc_buff_len,
+ copy_len, dptr, rpc_copied_seq - seq, 0);
+ if (err)
+ goto disable;
+
+ cb->rpc_desc_buff_len += copy_len;
+
+ if (cb->rpc_desc_buff_len != RPC_DESC_SIZE) {
+ LOG("Copied %d bytes: rpc_desc_buff_len: %u", copy_len, cb->rpc_desc_buff_len);
+ goto partial;
+ }
+
+ rpc_desc = (struct rpc_descriptor *)cb->rpc_desc_buf;
+ rpc_len = RPC_DESC_SIZE + rpc_desc->header_len + rpc_desc->payload_len;
+
+ if (rpc_len > INT_MAX)
+ goto disable;
+
+ cb->rpc_end_seq = cb->rpc_desc_seq + rpc_len;
+
+ LOG("Copied full descriptor: rpc_desc_seq: %u, rpc_end_seq: %u, header_len: %u, payload_len: %u",
+ CB_SEQ(rpc_desc_seq), CB_SEQ(rpc_end_seq),
+ rpc_desc->header_len, rpc_desc->payload_len);
+
+ return 0;
+disable:
+ return -1;
+partial:
+ return 1;
+}
+
+static void tcp_set_autolowat(struct tcp_autolowat_cb *cb,
+ struct sock *sk)
+{
+ struct tcp_sock *tp = (struct tcp_sock *)sk;
+ u32 val; /* To handle wraparound. */
+
+ LOG("Setting rcvlowat: tp->copied_seq: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_desc_buff_len: %u",
+ TP_SEQ(copied_seq), CB_SEQ(rpc_desc_seq),
+ CB_SEQ(rpc_end_seq), cb->rpc_desc_buff_len);
+
+ if (before(tp->copied_seq, cb->rpc_desc_seq))
+ val = cb->rpc_desc_seq - tp->copied_seq;
+ else if (cb->rpc_desc_buff_len != RPC_DESC_SIZE)
+ val = RPC_DESC_SIZE;
+ else
+ val = cb->rpc_end_seq - tp->copied_seq;
+
+ if (val != tp->inet_conn.icsk_inet.sk.sk_rcvlowat) {
+ bpf_tcp_ops_set_rcvlowat(sk, val);
+
+ LOG("Set rcvlowat: expected: %u, actual: %d\n",
+ val, tp->inet_conn.icsk_inet.sk.sk_rcvlowat);
+ } else {
+ LOG("No need to set rcvlowat: %u\n", val);
+ }
+}
+
+static void tcp_disable_autolowat(struct sock *sk)
+{
+ bpf_tcp_ops_set_flags((struct tcp_sock *)sk, 0, BPF_TCP_OPS_FLAG_RCVQ);
+
+ bpf_tcp_ops_set_rcvlowat(sk, 1);
+
+ LOG("Disabled autolowat");
+}
+
+static void tcp_do_autolowat(struct tcp_autolowat_cb *cb,
+ struct sock *sk, struct sk_buff *skb)
+{
+ struct bpf_dynptr dptr;
+ struct tcp_skb_cb *tcb;
+ u32 seq, end_seq;
+ int ret = 0, i;
+
+ if (bpf_dynptr_from_skb((struct __sk_buff *)skb, 0, &dptr)) {
+ ret = -1;
+ goto update;
+ }
+
+ tcb = bpf_core_cast(skb->cb, struct tcp_skb_cb);
+ seq = tcb->seq;
+ end_seq = tcb->end_seq - !!(tcb->tcp_flags & TCPHDR_FIN);
+
+ LOG("Start parsing skb: seq: %u, end_seq: %u, len: %u, rpc_desc_seq: %u, rpc_end_seq: %u, rpc_buff_len: %u",
+ SEQ(seq), SEQ(end_seq), end_seq - seq,
+ CB_SEQ(rpc_desc_seq), CB_SEQ(rpc_end_seq), cb->rpc_desc_buff_len);
+
+ if (cb->rpc_desc_buff_len != RPC_DESC_SIZE) {
+ ret = tcp_parse_descriptor(cb, &dptr, seq, end_seq);
+ if (ret)
+ goto update;
+ }
+
+ i = 0;
+
+ while (1) {
+ if (i++ > MAX_RPC_DESC_PER_SKB) {
+ ret = -1;
+ break;
+ }
+
+ if (after(cb->rpc_end_seq, end_seq)) {
+ LOG("No more descriptor: rpc_end_seq: %u, end_seq: %u",
+ CB_SEQ(rpc_end_seq), SEQ(end_seq));
+ break;
+ }
+
+ cb->rpc_desc_seq = cb->rpc_end_seq;
+ cb->rpc_desc_buff_len = 0;
+
+ if (cb->rpc_end_seq == end_seq)
+ break;
+
+ LOG("Found next descriptor: rpc_end_seq: %u, end_seq: %u, len: %u",
+ CB_SEQ(rpc_end_seq), SEQ(end_seq), end_seq - cb->rpc_end_seq);
+
+ ret = tcp_parse_descriptor(cb, &dptr, seq, end_seq);
+ if (ret)
+ break;
+ }
+
+update:
+ if (ret >= 0)
+ tcp_set_autolowat(cb, sk);
+ else
+ tcp_disable_autolowat(sk);
+}
+
+SEC("struct_ops")
+void BPF_PROG(tcp_autolowat_enqueue_rcvq, struct sock *sk, struct sk_buff *skb)
+{
+ struct tcp_autolowat_cb *cb;
+
+ cb = bpf_sk_storage_get(&tcp_autolowat_map, sk, 0, 0);
+ if (!cb)
+ return;
+
+ tcp_do_autolowat(cb, sk, skb);
+}
+
+SEC("struct_ops")
+void BPF_PROG(tcp_autolowat_dequeue_rcvq, struct sock *sk)
+{
+ struct tcp_autolowat_cb *cb;
+
+ cb = bpf_sk_storage_get(&tcp_autolowat_map, sk, 0, 0);
+ if (!cb)
+ return;
+
+ tcp_set_autolowat(cb, sk);
+}
+
+SEC(".struct_ops.link")
+struct bpf_tcp_ops tcp_autolowat_ops = {
+ .enqueue_rcvq = (void *)tcp_autolowat_enqueue_rcvq,
+ .dequeue_rcvq = (void *)tcp_autolowat_dequeue_rcvq,
+};
+
+static int tcp_init_autolowat_cb(struct bpf_tcp_sock *btp)
+{
+ struct tcp_autolowat_cb *cb;
+ struct tcp_sock *tp;
+
+ cb = bpf_sk_storage_get(&tcp_autolowat_map, btp, 0,
+ BPF_SK_STORAGE_GET_F_CREATE);
+ if (!cb)
+ return -1;
+
+ tp = bpf_core_cast(btp, struct tcp_sock);
+
+ cb->rpc_desc_seq = tp->copied_seq;
+ cb->rpc_end_seq = tp->copied_seq;
+#ifdef DEBUG
+ cb->isn = tp->copied_seq;
+#endif
+
+ return bpf_tcp_ops_set_flags((struct tcp_sock *)btp,
+ BPF_TCP_OPS_FLAG_RCVQ, 0);
+}
+
+SEC("cgroup/setsockopt")
+int tcp_autolowat_setsockopt(struct bpf_sockopt *ctx)
+{
+ void *optval_end = ctx->optval_end;
+ int *optval = ctx->optval;
+ struct bpf_tcp_sock *btp;
+
+ if (ctx->level != SOL_BPF || ctx->optname != BPF_TCP_AUTOLOWAT)
+ goto out;
+
+ if (optval + 1 > optval_end)
+ return 0; /* -EPERM */
+
+ btp = bpf_tcp_sock(ctx->sk);
+ if (!btp)
+ goto out;
+
+ if (*optval && tcp_init_autolowat_cb(btp))
+ return 0; /* -EPERM */
+
+ /*
+ * BPF has consumed this option, don't call kernel
+ * setsockopt handler.
+ */
+ ctx->optlen = -1;
+out:
+ return 1;
+}
+
+char _license[] SEC("license") = "GPL";
--
2.56.0.360.g66cac248cb-goog
^ permalink raw reply related [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 09/10] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
2026-10-06 19:24 ` [PATCH v4 bpf-next 09/10] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
@ 2026-10-06 19:42 ` sashiko-bot
2026-10-07 23:06 ` Amery Hung
1 sibling, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-06 19:42 UTC (permalink / raw)
To: Kuniyuki Iwashima; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: bpf_tcp_ops_set_rcvlowat() allows out-of-bounds writes and data races via arbitrary sockets
--
commit f1017b98cfed57e0845856c2710e8fad69bfcbe9
Author: Kuniyuki Iwashima <kuniyu@google.com>
bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
This commit adds a new BPF kfunc bpf_tcp_ops_set_rcvlowat() to dynamically
adjust a socket's sk_rcvlowat to suppress unnecessary wakeups. The function
determines whether to trigger a wakeup based on the BPF struct_ops hook
context.
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index 8182037c42690..2ba73dd6c52cc 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
[ ... ]
> @@ -361,12 +361,31 @@ __bpf_kfunc int bpf_tcp_ops_set_flags(struct tcp_sock *tp, u32 enable, u32 disab
> return 0;
> }
>
> +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> + const struct bpf_prog_aux *aux)
> +{
> + u32 moff = aux->attach_st_ops_member_off;
> + bool wakeup = false;
> +
> + if (moff == offsetof(struct bpf_tcp_ops, dequeue_rcvq))
> + wakeup = true;
> +
> + if (rcvlowat < 0)
> + rcvlowat = INT_MAX;
> +
> + return __tcp_set_rcvlowat(sk, rcvlowat, wakeup);
[Severity: High]
Does this code allow out-of-bounds writes if a non-TCP socket is passed?
The BPF verifier allows passing any trusted socket to kfuncs expecting a
struct sock *, such as those obtained via bpf_sk_lookup_udp() or
bpf_sk_lookup_tcp(). This means the sk argument is not restricted to the
locked context socket passed to the enqueue_rcvq or dequeue_rcvq hooks.
If a UDP socket is passed in, __tcp_set_rcvlowat() unsafely casts it to
struct tcp_sock and modifies tp->window_clamp:
net/ipv4/tcp.c:__tcp_set_rcvlowat() {
struct tcp_sock *tp = tcp_sk(sk);
...
if (space > sk->sk_rcvbuf) {
...
if (tp->window_clamp && tp->window_clamp < val)
WRITE_ONCE(tp->window_clamp, val);
}
}
Since struct udp_sock is smaller and lacks the window_clamp field, will this
write out of bounds? Should the signature take a struct tcp_sock *tp instead,
similar to bpf_tcp_ops_set_flags()?
Additionally, if a valid TCP socket is passed via a lookup, its socket lock
is not held. __tcp_set_rcvlowat() modifies sk_rcvlowat, sk_rcvbuf, and
window_clamp using logic that relies on the socket lock being held. Is there
a risk of data races and state corruption if it is called on an unlocked
socket?
Finally, passing an MPTCP subflow could bypass the intended restriction
since MPTCP subflows typically need to use mptcp_set_rcvlowat().
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261006192601.1875100-1-kuniyu@google.com?part=9
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 10/10] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-10-06 19:24 ` [PATCH v4 bpf-next 10/10] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
@ 2026-10-06 19:43 ` sashiko-bot
2026-10-06 22:06 ` Stanislav Fomichev
1 sibling, 0 replies; 36+ messages in thread
From: sashiko-bot @ 2026-10-06 19:43 UTC (permalink / raw)
To: Kuniyuki Iwashima; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] selftest: bpf: tcp_autolowat missing rpc_desc_buff_len reset
--
commit 379ea90d2efb24d101b9be64b55e86c49d5beee6
Author: Kuniyuki Iwashima <kuniyu@google.com>
selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().
This commit introduces a BPF selftest to verify the behavior of the
bpf_tcp_ops enqueue and dequeue callbacks dynamically adjusting sk_rcvlowat
for RPC frames. It sets up a client-server connection within a cgroup and
uses setsockopt with a custom BPF option to enable TCP AutoLOWAT.
> diff --git a/tools/testing/selftests/bpf/progs/tcp_autolowat.c b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
> new file mode 100644
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/tcp_autolowat.c
[ ... ]
> +static int tcp_init_autolowat_cb(struct bpf_tcp_sock *btp)
> +{
> + struct tcp_autolowat_cb *cb;
> + struct tcp_sock *tp;
> +
> + cb = bpf_sk_storage_get(&tcp_autolowat_map, btp, 0,
> + BPF_SK_STORAGE_GET_F_CREATE);
> + if (!cb)
> + return -1;
> +
> + tp = bpf_core_cast(btp, struct tcp_sock);
> +
> + cb->rpc_desc_seq = tp->copied_seq;
> + cb->rpc_end_seq = tp->copied_seq;
> +#ifdef DEBUG
> + cb->isn = tp->copied_seq;
> +#endif
> +
[Severity: Medium]
Should cb->rpc_desc_buff_len be reset to 0 here as well?
If an application re-initializes parsing by calling setsockopt again on
the same socket, tcp_init_autolowat_cb() resets the sequence numbers but
leaves the buffer length intact. If the storage entry already existed
from a previous operation that had read a partial RPC descriptor, the stale
parser state will persist and could corrupt the next parse operation.
> + return bpf_tcp_ops_set_flags((struct tcp_sock *)btp,
> + BPF_TCP_OPS_FLAG_RCVQ, 0);
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261006192601.1875100-1-kuniyu@google.com?part=10
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 05/10] bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after cgroup_bpf_enabled(CGROUP_SOCK_OPS).
2026-10-06 19:24 ` [PATCH v4 bpf-next 05/10] bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after cgroup_bpf_enabled(CGROUP_SOCK_OPS) Kuniyuki Iwashima
@ 2026-10-06 20:11 ` bot+bpf-ci
2026-10-06 22:05 ` Stanislav Fomichev
2026-10-07 22:38 ` Amery Hung
2 siblings, 0 replies; 36+ messages in thread
From: bot+bpf-ci @ 2026-10-06 20:11 UTC (permalink / raw)
To: kuniyu, ast, daniel, andrii, martin.lau, eddyz87, memxor
Cc: ameryhung, yonghong.song, john.fastabend, sdf, edumazet,
ncardwell, willemb, ukyab, cleger, kuniyu, kuni1840, bpf, netdev,
ast, andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 1605 bytes --]
> bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after cgroup_bpf_enabled(CGROUP_SOCK_OPS).
TL;DR: Nit only: a comment in tcp_set_state() still mentions
tcp_call_bpf_2arg(), which this patch removes; it may need to name
tcp_call_bpf_flag() instead.
> diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
> index fa69961c47d3..5d9d3bcde8f7 100644
> --- a/net/ipv4/tcp.c
> +++ b/net/ipv4/tcp.c
> @@ -2994,8 +2994,7 @@ void tcp_set_state(struct sock *sk, int state)
> */
> BTF_TYPE_EMIT_ENUM(BPF_TCP_ESTABLISHED);
>
> - if (BPF_SOCK_OPS_TEST_FLAG(tcp_sk(sk), BPF_SOCK_OPS_STATE_CB_FLAG))
> - tcp_call_bpf_2arg(sk, BPF_SOCK_OPS_STATE_CB, oldstate, state);
> + tcp_call_bpf_flag(sk, BPF_SOCK_OPS_STATE_CB, oldstate, state);
> bpf_tcp_ops_call(set_state, sk, state);
>
> switch (state) {
This isn't a bug, but the comment earlier in tcp_set_state() still refers to
a helper that this patch removes:
/* We defined a new enum for TCP states that are exported in BPF
* so as not force the internal TCP states to be frozen. The
* following checks will detect if an internal state value ever
* differs from the BPF value. If this ever happens, then we will
* need to remap the internal value to the BPF value before calling
* tcp_call_bpf_2arg.
*/
tcp_call_bpf_2arg() no longer exists after this change. Should the comment
now say tcp_call_bpf_flag() instead?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37521083031
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 02/10] bpf: tcp: Add a new per-socket flag and kfunc for bpf_tcp_ops.
2026-10-06 19:24 ` [PATCH v4 bpf-next 02/10] bpf: tcp: Add a new per-socket flag and kfunc for bpf_tcp_ops Kuniyuki Iwashima
@ 2026-10-06 22:04 ` Stanislav Fomichev
2026-10-06 23:10 ` Amery Hung
1 sibling, 0 replies; 36+ messages in thread
From: Stanislav Fomichev @ 2026-10-06 22:04 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 10/06, Kuniyuki Iwashima wrote:
> The legacy SOCK_OPS guards some hooks with a per-socket flag,
> tp->bpf_sock_ops_cb_flags.
>
> In contrast, bpf_tcp_ops was initially designed without per-socket
> flags so that users can simply define only the callbacks they need.
>
> However, it turned out that even attaching a bpf_tcp_ops with a NULL
> callback incurs measurable overhead in the fast path. [0]
>
> We could reuse tp->bpf_sock_ops_cb_flags, but it is u8 and has only
> 1 bit left for a new opt-in callback, while not all of the 7 existing
> opt-in hooks are in the fast path and really need a flag guard.
>
> Moreover, tp->bpf_sock_ops_cb_flags has two bpf helper functions to
> update it, both of which require sock_owned_by_me(sk):
>
> * bpf_sock_ops_cb_flags_set()
> * bpf_setsockopt(TCP_BPF_SOCK_OPS_CB_FLAGS)
>
> Both helpers only overwrite the field, which leads to reading and
> modifying the flags in the BPF prog and then writing them back
> via the helper. This prevents use from the fast path (tc or
> cgroup_skb) hooks where bh_lock_sock() alone cannot prevent races
> with process context.
>
> Let's add a new u32 field, tp->bpf_tcp_ops_flags, at the end of the
> tcp_sock_read_txrx cacheline group, along with a new kfunc,
> bpf_tcp_ops_set_flags().
>
> bpf_tcp_ops_set_flags() takes bitmasks of flags to enable and
> disable and updates tp->bpf_tcp_ops_flags atomically via
> try_cmpxchg() without relying on lock_sock().
>
> The kfunc is exposed to bpf_tcp_ops, BPF_PROG_TYPE_CGROUP_SOCKOPT,
> and BPF_PROG_TYPE_CGROUP_SKB.
>
> The first argument is struct tcp_sock * so that bpf_tcp_sock() etc
> is required for cgroup hooks, but not for bpf_tcp_ops where struct
> sock * is promoted to struct tcp_sock * automatically.
>
> The fast-path callbacks will be guarded by the new flags in a later
> patch after updating the existing selftest to avoid breaking bisection.
>
> Link: https://lore.kernel.org/netdev/CAMB2axMwBuz3X4Uwn5uzZUqm91EbiMnbQX832wUFOQ44fOwDnQ@mail.gmail.com/ #[0]
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 03/10] selftest: bpf: Use bpf_tcp_ops_set_flags() in bpf_tcp_ops_hdr.c.
2026-10-06 19:24 ` [PATCH v4 bpf-next 03/10] selftest: bpf: Use bpf_tcp_ops_set_flags() in bpf_tcp_ops_hdr.c Kuniyuki Iwashima
@ 2026-10-06 22:04 ` Stanislav Fomichev
2026-10-06 23:12 ` Amery Hung
1 sibling, 0 replies; 36+ messages in thread
From: Stanislav Fomichev @ 2026-10-06 22:04 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 10/06, Kuniyuki Iwashima wrote:
> The next patch guards bpf_tcp_ops callbacks in the fast
> path with per-socket flags.
>
> Let's use bpf_tcp_ops_set_flags() to enable flags for
> parse_hdr, write_hdr_opt, and hdr_opt_len.
>
> Note that BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_ALL and
> BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_UNKNOWN are used on the
> server and client, respectively, for better coverage.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 04/10] bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag.
2026-10-06 19:24 ` [PATCH v4 bpf-next 04/10] bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag Kuniyuki Iwashima
@ 2026-10-06 22:05 ` Stanislav Fomichev
2026-10-06 23:25 ` Amery Hung
1 sibling, 0 replies; 36+ messages in thread
From: Stanislav Fomichev @ 2026-10-06 22:05 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 10/06, Kuniyuki Iwashima wrote:
> bpf_tcp_ops.{parse_hdr,hdr_opt_len} are called for every
> incoming / outgoing skb.
>
> bpf_tcp_ops.rtt is called once per RTT, which is every
> incoming skb in ping-pong workloads like tcp_rr.
>
> Even attaching NULL callbacks in the fast path hurts performance.
>
> Let's guard them (and write_hdr_opt) with the new per-socket flags.
>
> __bpf_tcp_ops_call() and bpf_tcp_ops_call_flag() are added to check
> cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS) first and avoid accessing
> tp->bpf_tcp_ops_flags when no bpf_tcp_ops is attached.
>
> Note that bpf_tcp_ops_hdr_opt_len() has 3 callers and previously
> checked the static key both in tcp_established_options() and inside
> bpf_tcp_ops_call(). Now the static key is checked once at the
> beginning of bpf_tcp_ops_hdr_opt_len(), and the redundant check in
> tcp_established_options() is removed.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 05/10] bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after cgroup_bpf_enabled(CGROUP_SOCK_OPS).
2026-10-06 19:24 ` [PATCH v4 bpf-next 05/10] bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after cgroup_bpf_enabled(CGROUP_SOCK_OPS) Kuniyuki Iwashima
2026-10-06 20:11 ` bot+bpf-ci
@ 2026-10-06 22:05 ` Stanislav Fomichev
2026-10-07 22:38 ` Amery Hung
2 siblings, 0 replies; 36+ messages in thread
From: Stanislav Fomichev @ 2026-10-06 22:05 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 10/06, Kuniyuki Iwashima wrote:
> BPF_CGROUP_RUN_PROG_XXX() macros guard __cgroup_bpf_run_filter_XXX()
> with cgroup_bpf_enabled().
>
> However, even when no SOCK_OPS prog is attached, callers still
> initialise struct bpf_sock_ops_kern (memset(), etc.) or evaluate
> BPF_SOCK_OPS_TEST_FLAG(), which loads tp->bpf_sock_ops_cb_flags
> from a cold cacheline near the end of struct tcp_sock.
>
> Similar to bpf_tcp_ops, let's check cgroup_bpf_enabled() before
> BPF_SOCK_OPS_TEST_FLAG() and struct bpf_sock_ops_kern setup, and
> rename BPF_CGROUP_RUN_PROG_SOCK_OPS{,_SK}() with __ prefix.
>
> Since all direct callers of tcp_call_bpf() pass 0 and NULL for
> nargs and args, they can be folded into the new tcp_call_bpf()
> macro.
>
> All callers of tcp_call_bpf_{2,3}arg() check BPF_SOCK_OPS_TEST_FLAG()
> and do not need the return value. These are replaced with the new
> tcp_call_bpf_flag() macro.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 06/10] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-10-06 19:24 ` [PATCH v4 bpf-next 06/10] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
@ 2026-10-06 22:06 ` Stanislav Fomichev
2026-10-06 23:26 ` Amery Hung
1 sibling, 0 replies; 36+ messages in thread
From: Stanislav Fomichev @ 2026-10-06 22:06 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 10/06, Kuniyuki Iwashima wrote:
> Waking up a thread per packet is expensive when an application
> processes variable-length frames (e.g., RPC) that span multiple
> packets.
>
> SO_RCVLOWAT can defer wakeups, but because the frame size is
> encoded in a fixed-size descriptor at the start of each frame,
> the application has to:
>
> 1. wake up and recv() the descriptor,
> 2. raise SO_RCVLOWAT to the payload size via setsockopt(),
> 3. wake up and recv() the payload, and
> 4. reset SO_RCVLOWAT back to the descriptor size via
> setsockopt() for the next frame.
>
> This requires an extra wakeup and two setsockopt() syscalls
> for every single RPC frame.
>
> With SOCKMAP, we can parse skb and suppress wakeups in kernel,
> but SOCKMAP adds overhead and also kills zerocopy.
>
> Let's add lighter-weight opt-in callbacks to bpf_tcp_ops to
> replace that.
>
> .enqueue_rcvq(): invoked when TCP stack enqueues skb to
> sk->sk_receive_queue
>
> .dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
> is dequeued from sk->sk_receive_queue
>
> Those callbacks can be enabled on a per-socket basis by
> bpf_tcp_ops_set_flags():
>
> bpf_tcp_ops_set_flags((struct tcp_sock *)sk,
> BPF_TCP_OPS_FLAG_RCVQ, 0);
>
> Later, we will add a new kfunc to adjust sk->sk_rcvlowat from
> these callbacks.
>
> This will allow the bpf_tcp_ops prog to parse each skb and
> dynamically adjust sk->sk_rcvlowat to suppress unnecessary EPOLLIN
> wakeups until sufficient data is available in the receive queue.
>
> The placement of bpf_tcp_ops_call() in tcp_ofo_queue() and
> tcp_fastopen_add_skb() is chosen to provide the same snapshot
> as tcp_queue_rcv().
>
> For example, if bpf_tcp_ops_call() were called before updating
> TCP_SKB_CB(skb)->seq in tcp_fastopen_add_skb(), BPF prog would
> need an extra branch for the unlikely TFO case to strip SYN.
>
> In addition, the TCP stack can queue overlapping skbs into recvq.
> Once rcv_nxt is updated with a new skb, BPF prog can no longer
> infer the previous rcv_nxt from skb->len.
>
> Lastly, dequeue_rcvq() is placed in tcp_cleanup_rbuf() rather
> than __tcp_cleanup_rbuf() so that it is not called for sockets
> in SOCKMAP, where calling sk->sk_data_ready() from the new
> kfunc would otherwise trigger infinite recursion.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 10/10] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-10-06 19:24 ` [PATCH v4 bpf-next 10/10] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-10-06 19:43 ` sashiko-bot
@ 2026-10-06 22:06 ` Stanislav Fomichev
1 sibling, 0 replies; 36+ messages in thread
From: Stanislav Fomichev @ 2026-10-06 22:06 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 10/06, Kuniyuki Iwashima wrote:
> The test is roughly divided into two stages, and the sequence
> is as follows:
>
> I) Setup
>
> 1. Attach two BPF programs to a cgroup
> 2. Establish a TCP connection (@client <-> @child) within the cgroup
> 3. Enable BPF_TCP_OPS_FLAG_RCVQ on @child via setsockopt()
>
> II) RPC frame exchange in various patterns
>
> 4. Send a partial RPC descriptor from @client to @child
> 5. Verify that epoll does NOT wake up @child
> 6. Send the remaining data of the RPC frame
> 7. Verify that epoll finally wakes up @child
>
> During setup, two BPF programs are attached to simulate
> a real-world scenario; one is bpf_tcp_ops and the other is
> CGROUP_SOCKOPT.
>
> While the bpf_tcp_ops prog handles the dynamic adjustment of
> sk->sk_rcvlowat, the CGROUP_SOCKOPT prog is used to enable
> the TCP AutoLOWAT feature via userspace setsockopt() using
> pseudo options:
>
> #define SOL_BPF 0xdeadbeef
> #define BPF_TCP_AUTOLOWAT 0x8badf00d
>
> setsockopt(fd, SOL_BPF, BPF_TCP_AUTOLOWAT, &(int){1}, sizeof(int));
>
> This reflects a common production use case where an application
> decides to start parsing RPC frames only at a certain point in
> the stream (e.g., after HTTP Upgrade), rather than immediately
> after TCP 3WHS (->passive_established(), etc).
>
> When BPF_TCP_AUTOLOWAT is enabled, the BPF prog sets
> BPF_TCP_OPS_FLAG_RCVQ and initializes sk_local_storage
> for two sequence numbers to manage its state.
>
> Then, for the RPC frame exchange, this test uses a simple format
> defined as follows:
>
> 0 8 16 24 32
> +--------+--------+-------+--------+ `.
> | header size | |
> +--------+--------+-------+--------+ > RPC descriptor (8 bytes)
> | payload size | |
> +--------+--------+-------+--------+ .'
> ~ header ~
> +--------+--------+-------+--------+
> ~ payload ~
> +--------+--------+-------+--------+
>
> Every time a new skb is enqueued to sk->sk_receive_queue, the
> bpf_tcp_ops prog parses it and updates these sequence numbers:
>
> rpc_desc_seq : the SEQ # of the start of the RPC descriptor
> rpc_end_seq : the SEQ # of the end of the RPC frame
> => rpc_desc_seq + 8 + header size + payload size
>
> Assume we receive two RPC descriptors in the following pattern:
>
> 1. When we receive skb-1, only part of the RPC descriptor is parsed.
> rpc_desc_seq is set to the first byte while rpc_end_seq is
> unknown. Thus, sk->sk_rcvlowat is set to the size of the RPC
> descriptor (8 bytes).
>
> <- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
> +-----------+.................+....................+......
> | RPC desc 1 | header + payload | RPC desc 2 | ...
> +-----------+.................+....................+......
> ^ ^-.
> `- rpc_desc_seq `- sk->sk_rcvlowat
>
> 2. Next, we receive skb-2, which completes the first RPC descriptor.
> Now rpc_end_seq is known, so sk->sk_rcvlowat is advanced to it.
>
> <- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
> +-----------+-----------------+....................+......
> | RPC desc 1 | header + payload | RPC desc 2 | ...
> +-----------+-----------------+....................+......
> ^ ^
> '- rpc_desc_seq '- rpc_end_seq
> & sk->sk_rcvlowat
>
> 3. Once we receive skb-3, which contains the next full RPC descriptor,
> rpc_desc_seq is advanced and rpc_end_seq is updated according
> to the size of RPC frame 2.
>
> Note that sk->sk_rcvlowat is NOT updated to the new rpc_end_seq
> yet. This ensures that the application is woken up to read the
> already complete RPC frame 1.
>
> <- skb-1 -> <---- skb-2 ----> <------ skb-3 ----->
> +-----------+-----------------+--------------------+......
> | RPC desc 1 | header + payload | RPC desc 2 | ... |
> +-----------+-----------------+--------------------+......
> ^ ^
> rpc_desc_seq -----------' rpc_end_seq ----...-'
> & sk->sk_rcvlowat
>
> This sequence corresponds to the 4th test case in rpc_test_cases[],
> and we can see helpful output if we "#define DEBUG":
>
> # cat /sys/kernel/tracing/trace_pipe | \
> awk '{ if ($0 ~ /AF_/) sub(/^.*AF_/, "AF_"); print $0 }' & \
> BGPID=$!; ./test_progs -t tcp_autolowat; kill -9 -$BGPID
> ...
> AF_INET6 rpc_test_cases[3]: Start parsing skb: seq: 0, end_seq: 1, len: 1, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_buff_len: 0
> AF_INET6 rpc_test_cases[3]: Copied 1 bytes: rpc_desc_buff_len: 1
> AF_INET6 rpc_test_cases[3]: Setting rcvlowat: tp->copied_seq: 0, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_desc_buff_len: 1
> AF_INET6 rpc_test_cases[3]: Set rcvlowat: expected: 8, actual: 8
>
> AF_INET6 rpc_test_cases[3]: Start parsing skb: seq: 1, end_seq: 8, len: 7, rpc_desc_seq: 0, rpc_end_seq: 0, rpc_buff_len: 1
> AF_INET6 rpc_test_cases[3]: Copied full descriptor: rpc_desc_seq: 0, rpc_end_seq: 258, header_len: 100, payload_len: 150
> AF_INET6 rpc_test_cases[3]: No more descriptor: rpc_end_seq: 258, end_seq: 8
> AF_INET6 rpc_test_cases[3]: Setting rcvlowat: tp->copied_seq: 0, rpc_desc_seq: 0, rpc_end_seq: 258, rpc_desc_buff_len: 8
> AF_INET6 rpc_test_cases[3]: Set rcvlowat: expected: 258, actual: 258
> ...
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 08/10] bpf: mptcp: Don't support BPF_TCP_OPS_FLAG_RCVQ.
2026-10-06 19:24 ` [PATCH v4 bpf-next 08/10] bpf: mptcp: Don't support BPF_TCP_OPS_FLAG_RCVQ Kuniyuki Iwashima
@ 2026-10-06 22:06 ` Stanislav Fomichev
2026-10-07 22:43 ` Amery Hung
1 sibling, 0 replies; 36+ messages in thread
From: Stanislav Fomichev @ 2026-10-06 22:06 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On 10/06, Kuniyuki Iwashima wrote:
> The next patch exposes a new kfunc calling __tcp_set_rcvlowat()
> to bpf_tcp_ops.
>
> MPTCP has its own sock->ops->set_rcvlowat() / mptcp_set_rcvlowat(),
> so we should not allow calling __tcp_set_rcvlowat() on MPTCP
> subflows.
>
> Let's disable BPF_TCP_OPS_FLAG_RCVQ for MPTCP for now.
>
> If needed in the future, bpf_tcp_ops_set_rcvlowat() could be
> extended to properly support MPTCP.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Acked-by: Stanislav Fomichev <sdf@fomichev.me>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 01/10] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list.
2026-10-06 19:24 ` [PATCH v4 bpf-next 01/10] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
@ 2026-10-06 23:09 ` Amery Hung
0 siblings, 0 replies; 36+ messages in thread
From: Amery Hung @ 2026-10-06 23:09 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev,
Emil Tsalapatis
On Tue, Oct 6, 2026 at 12:26 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>
> Currently, four bpf_tcp_ops callbacks are not allowed to call
> bpf_setsockopt() and bpf_getsockopt().
>
> However, the deny-list is fragile, and when a new callback is
> added, we might re-open a can of worms. [0][1]
>
> Let's convert it to allow-list.
>
> Note that it can be written as offsetof(,X) <= moff <= offsetof(,Y)
> but clang can optimise to similar code anyway.
>
> Link: https://lore.kernel.org/r/20220929070407.965581-5-martin.lau@linux.dev/ #[0]
> Link: https://lore.kernel.org/bpf/20260421155804.135786-1-kafai.wan@linux.dev/ #[1]
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
> Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Reviewed-by: Amery Hung <ameryhung@gmail.com>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 02/10] bpf: tcp: Add a new per-socket flag and kfunc for bpf_tcp_ops.
2026-10-06 19:24 ` [PATCH v4 bpf-next 02/10] bpf: tcp: Add a new per-socket flag and kfunc for bpf_tcp_ops Kuniyuki Iwashima
2026-10-06 22:04 ` Stanislav Fomichev
@ 2026-10-06 23:10 ` Amery Hung
1 sibling, 0 replies; 36+ messages in thread
From: Amery Hung @ 2026-10-06 23:10 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On Tue, Oct 6, 2026 at 12:26 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>
> The legacy SOCK_OPS guards some hooks with a per-socket flag,
> tp->bpf_sock_ops_cb_flags.
>
> In contrast, bpf_tcp_ops was initially designed without per-socket
> flags so that users can simply define only the callbacks they need.
>
> However, it turned out that even attaching a bpf_tcp_ops with a NULL
> callback incurs measurable overhead in the fast path. [0]
>
> We could reuse tp->bpf_sock_ops_cb_flags, but it is u8 and has only
> 1 bit left for a new opt-in callback, while not all of the 7 existing
> opt-in hooks are in the fast path and really need a flag guard.
>
> Moreover, tp->bpf_sock_ops_cb_flags has two bpf helper functions to
> update it, both of which require sock_owned_by_me(sk):
>
> * bpf_sock_ops_cb_flags_set()
> * bpf_setsockopt(TCP_BPF_SOCK_OPS_CB_FLAGS)
>
> Both helpers only overwrite the field, which leads to reading and
> modifying the flags in the BPF prog and then writing them back
> via the helper. This prevents use from the fast path (tc or
> cgroup_skb) hooks where bh_lock_sock() alone cannot prevent races
> with process context.
>
> Let's add a new u32 field, tp->bpf_tcp_ops_flags, at the end of the
> tcp_sock_read_txrx cacheline group, along with a new kfunc,
> bpf_tcp_ops_set_flags().
>
> bpf_tcp_ops_set_flags() takes bitmasks of flags to enable and
> disable and updates tp->bpf_tcp_ops_flags atomically via
> try_cmpxchg() without relying on lock_sock().
>
> The kfunc is exposed to bpf_tcp_ops, BPF_PROG_TYPE_CGROUP_SOCKOPT,
> and BPF_PROG_TYPE_CGROUP_SKB.
>
> The first argument is struct tcp_sock * so that bpf_tcp_sock() etc
> is required for cgroup hooks, but not for bpf_tcp_ops where struct
> sock * is promoted to struct tcp_sock * automatically.
>
> The fast-path callbacks will be guarded by the new flags in a later
> patch after updating the existing selftest to avoid breaking bisection.
>
> Link: https://lore.kernel.org/netdev/CAMB2axMwBuz3X4Uwn5uzZUqm91EbiMnbQX832wUFOQ44fOwDnQ@mail.gmail.com/ #[0]
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Amery Hung <ameryhung@gmail.com>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 03/10] selftest: bpf: Use bpf_tcp_ops_set_flags() in bpf_tcp_ops_hdr.c.
2026-10-06 19:24 ` [PATCH v4 bpf-next 03/10] selftest: bpf: Use bpf_tcp_ops_set_flags() in bpf_tcp_ops_hdr.c Kuniyuki Iwashima
2026-10-06 22:04 ` Stanislav Fomichev
@ 2026-10-06 23:12 ` Amery Hung
1 sibling, 0 replies; 36+ messages in thread
From: Amery Hung @ 2026-10-06 23:12 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On Tue, Oct 6, 2026 at 12:26 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>
> The next patch guards bpf_tcp_ops callbacks in the fast
> path with per-socket flags.
>
> Let's use bpf_tcp_ops_set_flags() to enable flags for
> parse_hdr, write_hdr_opt, and hdr_opt_len.
>
> Note that BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_ALL and
> BPF_TCP_OPS_FLAG_PARSE_HDR_OPT_UNKNOWN are used on the
> server and client, respectively, for better coverage.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Amery Hung <ameryhung@gmail.com>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 04/10] bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag.
2026-10-06 19:24 ` [PATCH v4 bpf-next 04/10] bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag Kuniyuki Iwashima
2026-10-06 22:05 ` Stanislav Fomichev
@ 2026-10-06 23:25 ` Amery Hung
2026-10-07 2:48 ` Kuniyuki Iwashima
1 sibling, 1 reply; 36+ messages in thread
From: Amery Hung @ 2026-10-06 23:25 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
> +++ b/net/ipv4/tcp_input.c
> @@ -210,6 +210,11 @@ static void bpf_skops_established(struct sock *sk, int bpf_op,
>
> static void bpf_tcp_ops_parse_hdr(struct sock *sk, struct sk_buff *skb)
> {
> + const struct tcp_sock *tp;
> +
> + if (!cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS))
> + return;
> +
> switch (sk->sk_state) {
> case TCP_SYN_RECV:
> case TCP_SYN_SENT:
> @@ -217,7 +222,12 @@ static void bpf_tcp_ops_parse_hdr(struct sock *sk, struct sk_buff *skb)
> return;
> }
>
> - bpf_tcp_ops_call(parse_hdr, sk, skb);
> + tp = tcp_sk(sk);
> +
> + if ((tp->rx_opt.saw_unknown &&
> + BPF_TCP_OPS_TEST_FLAG(tp, PARSE_HDR_OPT_UNKNOWN)) ||
> + BPF_TCP_OPS_TEST_FLAG(tp, PARSE_HDR_OPT_ALL))
nit: ^ missing a space
> + __bpf_tcp_ops_call(parse_hdr, sk, skb);
> }
>
> static __cold void tcp_gro_dev_warn(const struct sock *sk, const struct sk_buff *skb,
> diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c
> index 3ac465513bf5..8770f3084efe 100644
> --- a/net/ipv4/tcp_output.c
> +++ b/net/ipv4/tcp_output.c
> @@ -581,8 +581,8 @@ static void bpf_skops_write_hdr_opt(struct sock *sk, struct sk_buff *skb,
> * writer's bytes). The writer finds the append point by scanning from
> * first_opt_off + nr_written to the first NOP.
> */
> - bpf_tcp_ops_call(write_hdr_opt, sk, skb, req, syn_skb, synack_type,
> - first_opt_off + nr_written);
> + bpf_tcp_ops_call_flag(write_hdr_opt, WRITE_HDR_OPT, sk, skb, req,
> + syn_skb, synack_type, first_opt_off + nr_written);
> }
> #else
> static u32 bpf_skops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,
> @@ -613,11 +613,13 @@ static u32 bpf_tcp_ops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,
> {
> unsigned int remaining_out = remaining, reserved;
>
> - if (!remaining)
> - return 0;
> + if (!cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS) ||
> + !BPF_TCP_OPS_TEST_FLAG(tcp_sk(sk), WRITE_HDR_OPT)||
nit: missing a space before ||
> + !remaining)
> + return remaining;
>
> /* bpf_tcp_ops_reserve_hdr_opt() reserves space via remaining_out */
> - bpf_tcp_ops_call(hdr_opt_len, sk, skb, req, syn_skb, synack_type, &remaining_out);
> + __bpf_tcp_ops_call(hdr_opt_len, sk, skb, req, syn_skb, synack_type, &remaining_out);
>
> reserved = remaining - remaining_out;
> if (!reserved)
> @@ -1193,8 +1195,9 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> struct tcp_key *key)
> {
> struct tcp_sock *tp = tcp_sk(sk);
> - unsigned int size = 0;
> unsigned int eff_sacks;
> + unsigned int remaining;
> + unsigned int size = 0;
>
> opts->options = 0;
> opts->bpf_opt_len = 0;
> @@ -1224,10 +1227,10 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> * left.
> */
> if (sk_is_mptcp(sk)) {
> - unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> bool has_ts = opts->options & OPTION_TS;
> int opt_size;
>
> + remaining = MAX_TCP_OPTION_SPACE - size;
> opts->mptcp.drop_ts = 0;
>
> opt_size = mptcp_established_options(sk, skb, remaining, has_ts,
> @@ -1244,7 +1247,8 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
>
> eff_sacks = tp->rx_opt.num_sacks + tp->rx_opt.dsack;
> if (unlikely(eff_sacks)) {
> - const unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> + remaining = MAX_TCP_OPTION_SPACE - size;
> +
> if (likely(remaining >= TCPOLEN_SACK_BASE_ALIGNED +
> TCPOLEN_SACK_PERBLOCK)) {
> opts->num_sack_blocks =
> @@ -1277,22 +1281,18 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
>
> if (unlikely(BPF_SOCK_OPS_TEST_FLAG(tp,
> BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG))) {
> - unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> -
> + remaining = MAX_TCP_OPTION_SPACE - size;
> remaining = bpf_skops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
> remaining);
>
> size = MAX_TCP_OPTION_SPACE - remaining;
> }
>
> - if (cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS)) {
> - unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> -
> - remaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
> - remaining);
> + remaining = MAX_TCP_OPTION_SPACE - size;
> + remaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
> + remaining);
>
> - size = MAX_TCP_OPTION_SPACE - remaining;
> - }
> + size = MAX_TCP_OPTION_SPACE - remaining;
This static-key check is not redundant unless
bpf_tcp_ops_hdr_opt_len() is inlined. In my build,
tcp_established_options() unconditionally calls the out-of-line
helper, so sockets pay the call/prologue cost even when no bpf_tcp_ops
is attached.
>
> return size;
> }
> --
> 2.56.0.360.g66cac248cb-goog
>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 06/10] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq().
2026-10-06 19:24 ` [PATCH v4 bpf-next 06/10] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-10-06 22:06 ` Stanislav Fomichev
@ 2026-10-06 23:26 ` Amery Hung
1 sibling, 0 replies; 36+ messages in thread
From: Amery Hung @ 2026-10-06 23:26 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On Tue, Oct 6, 2026 at 12:26 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>
> Waking up a thread per packet is expensive when an application
> processes variable-length frames (e.g., RPC) that span multiple
> packets.
>
> SO_RCVLOWAT can defer wakeups, but because the frame size is
> encoded in a fixed-size descriptor at the start of each frame,
> the application has to:
>
> 1. wake up and recv() the descriptor,
> 2. raise SO_RCVLOWAT to the payload size via setsockopt(),
> 3. wake up and recv() the payload, and
> 4. reset SO_RCVLOWAT back to the descriptor size via
> setsockopt() for the next frame.
>
> This requires an extra wakeup and two setsockopt() syscalls
> for every single RPC frame.
>
> With SOCKMAP, we can parse skb and suppress wakeups in kernel,
> but SOCKMAP adds overhead and also kills zerocopy.
>
> Let's add lighter-weight opt-in callbacks to bpf_tcp_ops to
> replace that.
>
> .enqueue_rcvq(): invoked when TCP stack enqueues skb to
> sk->sk_receive_queue
>
> .dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
> is dequeued from sk->sk_receive_queue
>
> Those callbacks can be enabled on a per-socket basis by
> bpf_tcp_ops_set_flags():
>
> bpf_tcp_ops_set_flags((struct tcp_sock *)sk,
> BPF_TCP_OPS_FLAG_RCVQ, 0);
>
> Later, we will add a new kfunc to adjust sk->sk_rcvlowat from
> these callbacks.
>
> This will allow the bpf_tcp_ops prog to parse each skb and
> dynamically adjust sk->sk_rcvlowat to suppress unnecessary EPOLLIN
> wakeups until sufficient data is available in the receive queue.
>
> The placement of bpf_tcp_ops_call() in tcp_ofo_queue() and
> tcp_fastopen_add_skb() is chosen to provide the same snapshot
> as tcp_queue_rcv().
>
> For example, if bpf_tcp_ops_call() were called before updating
> TCP_SKB_CB(skb)->seq in tcp_fastopen_add_skb(), BPF prog would
> need an extra branch for the unlikely TFO case to strip SYN.
>
> In addition, the TCP stack can queue overlapping skbs into recvq.
> Once rcv_nxt is updated with a new skb, BPF prog can no longer
> infer the previous rcv_nxt from skb->len.
>
> Lastly, dequeue_rcvq() is placed in tcp_cleanup_rbuf() rather
> than __tcp_cleanup_rbuf() so that it is not called for sockets
> in SOCKMAP, where calling sk->sk_data_ready() from the new
> kfunc would otherwise trigger infinite recursion.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Amery Hung <ameryhung@gmail.com>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 04/10] bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag.
2026-10-06 23:25 ` Amery Hung
@ 2026-10-07 2:48 ` Kuniyuki Iwashima
2026-10-07 22:36 ` Amery Hung
0 siblings, 1 reply; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-07 2:48 UTC (permalink / raw)
To: Amery Hung
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On Tue, Oct 6, 2026 at 4:25 PM Amery Hung <ameryhung@gmail.com> wrote:
>
> > +++ b/net/ipv4/tcp_input.c
> > @@ -210,6 +210,11 @@ static void bpf_skops_established(struct sock *sk, int bpf_op,
> >
> > static void bpf_tcp_ops_parse_hdr(struct sock *sk, struct sk_buff *skb)
> > {
> > + const struct tcp_sock *tp;
> > +
> > + if (!cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS))
> > + return;
> > +
> > switch (sk->sk_state) {
> > case TCP_SYN_RECV:
> > case TCP_SYN_SENT:
> > @@ -217,7 +222,12 @@ static void bpf_tcp_ops_parse_hdr(struct sock *sk, struct sk_buff *skb)
> > return;
> > }
> >
> > - bpf_tcp_ops_call(parse_hdr, sk, skb);
> > + tp = tcp_sk(sk);
> > +
> > + if ((tp->rx_opt.saw_unknown &&
> > + BPF_TCP_OPS_TEST_FLAG(tp, PARSE_HDR_OPT_UNKNOWN)) ||
> > + BPF_TCP_OPS_TEST_FLAG(tp, PARSE_HDR_OPT_ALL))
>
> nit: ^ missing a space
this indentation is correct.
>
> > + __bpf_tcp_ops_call(parse_hdr, sk, skb);
> > }
> >
> > static __cold void tcp_gro_dev_warn(const struct sock *sk, const struct sk_buff *skb,
> > diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c
> > index 3ac465513bf5..8770f3084efe 100644
> > --- a/net/ipv4/tcp_output.c
> > +++ b/net/ipv4/tcp_output.c
> > @@ -581,8 +581,8 @@ static void bpf_skops_write_hdr_opt(struct sock *sk, struct sk_buff *skb,
> > * writer's bytes). The writer finds the append point by scanning from
> > * first_opt_off + nr_written to the first NOP.
> > */
> > - bpf_tcp_ops_call(write_hdr_opt, sk, skb, req, syn_skb, synack_type,
> > - first_opt_off + nr_written);
> > + bpf_tcp_ops_call_flag(write_hdr_opt, WRITE_HDR_OPT, sk, skb, req,
> > + syn_skb, synack_type, first_opt_off + nr_written);
> > }
> > #else
> > static u32 bpf_skops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,
> > @@ -613,11 +613,13 @@ static u32 bpf_tcp_ops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,
> > {
> > unsigned int remaining_out = remaining, reserved;
> >
> > - if (!remaining)
> > - return 0;
> > + if (!cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS) ||
> > + !BPF_TCP_OPS_TEST_FLAG(tcp_sk(sk), WRITE_HDR_OPT)||
>
> nit: missing a space before ||
oops, due to last minute edit :/
>
> > + !remaining)
> > + return remaining;
> >
> > /* bpf_tcp_ops_reserve_hdr_opt() reserves space via remaining_out */
> > - bpf_tcp_ops_call(hdr_opt_len, sk, skb, req, syn_skb, synack_type, &remaining_out);
> > + __bpf_tcp_ops_call(hdr_opt_len, sk, skb, req, syn_skb, synack_type, &remaining_out);
> >
> > reserved = remaining - remaining_out;
> > if (!reserved)
> > @@ -1193,8 +1195,9 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> > struct tcp_key *key)
> > {
> > struct tcp_sock *tp = tcp_sk(sk);
> > - unsigned int size = 0;
> > unsigned int eff_sacks;
> > + unsigned int remaining;
> > + unsigned int size = 0;
> >
> > opts->options = 0;
> > opts->bpf_opt_len = 0;
> > @@ -1224,10 +1227,10 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> > * left.
> > */
> > if (sk_is_mptcp(sk)) {
> > - unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> > bool has_ts = opts->options & OPTION_TS;
> > int opt_size;
> >
> > + remaining = MAX_TCP_OPTION_SPACE - size;
> > opts->mptcp.drop_ts = 0;
> >
> > opt_size = mptcp_established_options(sk, skb, remaining, has_ts,
> > @@ -1244,7 +1247,8 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> >
> > eff_sacks = tp->rx_opt.num_sacks + tp->rx_opt.dsack;
> > if (unlikely(eff_sacks)) {
> > - const unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> > + remaining = MAX_TCP_OPTION_SPACE - size;
> > +
> > if (likely(remaining >= TCPOLEN_SACK_BASE_ALIGNED +
> > TCPOLEN_SACK_PERBLOCK)) {
> > opts->num_sack_blocks =
> > @@ -1277,22 +1281,18 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> >
> > if (unlikely(BPF_SOCK_OPS_TEST_FLAG(tp,
> > BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG))) {
> > - unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> > -
> > + remaining = MAX_TCP_OPTION_SPACE - size;
> > remaining = bpf_skops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
> > remaining);
> >
> > size = MAX_TCP_OPTION_SPACE - remaining;
> > }
> >
> > - if (cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS)) {
> > - unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> > -
> > - remaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
> > - remaining);
> > + remaining = MAX_TCP_OPTION_SPACE - size;
> > + remaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
> > + remaining);
> >
> > - size = MAX_TCP_OPTION_SPACE - remaining;
> > - }
> > + size = MAX_TCP_OPTION_SPACE - remaining;
>
> This static-key check is not redundant unless
> bpf_tcp_ops_hdr_opt_len() is inlined. In my build,
> tcp_established_options() unconditionally calls the out-of-line
> helper, so sockets pay the call/prologue cost even when no bpf_tcp_ops
> is attached.
Interesting, I guess you don't use FDO ?
On my FDO build, even tcp_established_options() is inlined
to __tcp_transmit_skb() and both bpf_tcp_ops_hdr_opt_len()
and bpf_tcp_ops_hdr_opt_len() are NOPs as expected.
ffffffff8224a511: 41 bc 28 00 00 00 movl $0x28, %r12d
# remaining = MAX_TCP_OPTION_SPACE (40)
ffffffff8224a517: 45 29 ec subl %r13d, %r12d
# remaining -= size
ffffffff8224a51a: 0f 1f 44 00 00 nopl (%rax,%rax)
# bpf_skops_hdr_opt_len()
ffffffff8224a51f: 44 89 64 24 78 movl %r12d, 0x78(%rsp)
# remaining_out = remaining
ffffffff8224a524: 0f 1f 44 00 00 nopl (%rax,%rax)
# bpf_tcp_ops_hdr_opt_len()
I will add a static inline function and move both
cgroup_bpf_enabled() there and reuse it in 3 places.
(same for bpf_tcp_ops_parse_hdr())
Then, the unnecessary movl will also disappear.
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT.
2026-10-06 19:24 [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
` (9 preceding siblings ...)
2026-10-06 19:24 ` [PATCH v4 bpf-next 10/10] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
@ 2026-10-07 13:11 ` Jakub Sitnicki
10 siblings, 0 replies; 36+ messages in thread
From: Jakub Sitnicki @ 2026-10-07 13:11 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Amery Hung, Yonghong Song, John Fastabend, Stanislav Fomichev,
Eric Dumazet, Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On Tue, Oct 06, 2026 at 07:24 PM GMT, Kuniyuki Iwashima wrote:
> This series introduces two callbacks for bpf_tcp_ops:
>
> .enqueue_rcvq(): invoked when TCP stack enqueues skb to
> sk->sk_receive_queue
>
> .dequeue_rcvq(): invoked in tcp_cleanup_rbuf() after data
> is dequeued from sk->sk_receive_queue
>
> Those callbacks can be enabled on a per-socket basis by
> a new kfunc bpf_tcp_ops_set_flags():
>
> bpf_tcp_ops_set_flags((struct tcp_sock *)sk,
> BPF_TCP_OPS_FLAG_RCVQ, 0);
>
> This allows the BPF prog to dynamically adjust sk->sk_rcvlowat,
> suppressing unnecessary EPOLLIN wakeups until sufficient data
> is available in the receive queue.
>
> This functionality, which we call "TCP AutoLOWAT", was originally
> developed in 2020 by Tenzin Ukyab with the help of Soheil Hassas
> Yeganeh, Arjun Roy, and Eric Dumazet. It has served Google RPC
> workloads for more than 5 years.
>
> Combined with TCP RX zerocopy, this typically allows us to read an
> entire RPC frame with just a single wakeup and a single system call.
Nice! Just yesterday we saw at LPC that even with SK_SKB/SOCKMAP
overhead, you can observe an RPS gain from waking up once there is a
full L7 request to read out [1].
[1] slide 72, https://lpc.events/event/20/contributions/2428/attachments/2064/4781/Beeper.pdf
[...]
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 04/10] bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag.
2026-10-07 2:48 ` Kuniyuki Iwashima
@ 2026-10-07 22:36 ` Amery Hung
2026-10-08 2:10 ` Kuniyuki Iwashima
0 siblings, 1 reply; 36+ messages in thread
From: Amery Hung @ 2026-10-07 22:36 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On Tue, Oct 6, 2026 at 7:48 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>
> On Tue, Oct 6, 2026 at 4:25 PM Amery Hung <ameryhung@gmail.com> wrote:
> >
> > > +++ b/net/ipv4/tcp_input.c
> > > @@ -210,6 +210,11 @@ static void bpf_skops_established(struct sock *sk, int bpf_op,
> > >
> > > static void bpf_tcp_ops_parse_hdr(struct sock *sk, struct sk_buff *skb)
> > > {
> > > + const struct tcp_sock *tp;
> > > +
> > > + if (!cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS))
> > > + return;
> > > +
> > > switch (sk->sk_state) {
> > > case TCP_SYN_RECV:
> > > case TCP_SYN_SENT:
> > > @@ -217,7 +222,12 @@ static void bpf_tcp_ops_parse_hdr(struct sock *sk, struct sk_buff *skb)
> > > return;
> > > }
> > >
> > > - bpf_tcp_ops_call(parse_hdr, sk, skb);
> > > + tp = tcp_sk(sk);
> > > +
> > > + if ((tp->rx_opt.saw_unknown &&
> > > + BPF_TCP_OPS_TEST_FLAG(tp, PARSE_HDR_OPT_UNKNOWN)) ||
> > > + BPF_TCP_OPS_TEST_FLAG(tp, PARSE_HDR_OPT_ALL))
> >
> > nit: ^ missing a space
>
> this indentation is correct.
Sorry. False alarm.
>
> >
> > > + __bpf_tcp_ops_call(parse_hdr, sk, skb);
> > > }
> > >
> > > static __cold void tcp_gro_dev_warn(const struct sock *sk, const struct sk_buff *skb,
> > > diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c
> > > index 3ac465513bf5..8770f3084efe 100644
> > > --- a/net/ipv4/tcp_output.c
> > > +++ b/net/ipv4/tcp_output.c
> > > @@ -581,8 +581,8 @@ static void bpf_skops_write_hdr_opt(struct sock *sk, struct sk_buff *skb,
> > > * writer's bytes). The writer finds the append point by scanning from
> > > * first_opt_off + nr_written to the first NOP.
> > > */
> > > - bpf_tcp_ops_call(write_hdr_opt, sk, skb, req, syn_skb, synack_type,
> > > - first_opt_off + nr_written);
> > > + bpf_tcp_ops_call_flag(write_hdr_opt, WRITE_HDR_OPT, sk, skb, req,
> > > + syn_skb, synack_type, first_opt_off + nr_written);
> > > }
> > > #else
> > > static u32 bpf_skops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,
> > > @@ -613,11 +613,13 @@ static u32 bpf_tcp_ops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,
> > > {
> > > unsigned int remaining_out = remaining, reserved;
> > >
> > > - if (!remaining)
> > > - return 0;
> > > + if (!cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS) ||
> > > + !BPF_TCP_OPS_TEST_FLAG(tcp_sk(sk), WRITE_HDR_OPT)||
> >
> > nit: missing a space before ||
>
> oops, due to last minute edit :/
>
> >
> > > + !remaining)
> > > + return remaining;
> > >
> > > /* bpf_tcp_ops_reserve_hdr_opt() reserves space via remaining_out */
> > > - bpf_tcp_ops_call(hdr_opt_len, sk, skb, req, syn_skb, synack_type, &remaining_out);
> > > + __bpf_tcp_ops_call(hdr_opt_len, sk, skb, req, syn_skb, synack_type, &remaining_out);
> > >
> > > reserved = remaining - remaining_out;
> > > if (!reserved)
> > > @@ -1193,8 +1195,9 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> > > struct tcp_key *key)
> > > {
> > > struct tcp_sock *tp = tcp_sk(sk);
> > > - unsigned int size = 0;
> > > unsigned int eff_sacks;
> > > + unsigned int remaining;
> > > + unsigned int size = 0;
> > >
> > > opts->options = 0;
> > > opts->bpf_opt_len = 0;
> > > @@ -1224,10 +1227,10 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> > > * left.
> > > */
> > > if (sk_is_mptcp(sk)) {
> > > - unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> > > bool has_ts = opts->options & OPTION_TS;
> > > int opt_size;
> > >
> > > + remaining = MAX_TCP_OPTION_SPACE - size;
> > > opts->mptcp.drop_ts = 0;
> > >
> > > opt_size = mptcp_established_options(sk, skb, remaining, has_ts,
> > > @@ -1244,7 +1247,8 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> > >
> > > eff_sacks = tp->rx_opt.num_sacks + tp->rx_opt.dsack;
> > > if (unlikely(eff_sacks)) {
> > > - const unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> > > + remaining = MAX_TCP_OPTION_SPACE - size;
> > > +
> > > if (likely(remaining >= TCPOLEN_SACK_BASE_ALIGNED +
> > > TCPOLEN_SACK_PERBLOCK)) {
> > > opts->num_sack_blocks =
> > > @@ -1277,22 +1281,18 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> > >
> > > if (unlikely(BPF_SOCK_OPS_TEST_FLAG(tp,
> > > BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG))) {
> > > - unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> > > -
> > > + remaining = MAX_TCP_OPTION_SPACE - size;
> > > remaining = bpf_skops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
> > > remaining);
> > >
> > > size = MAX_TCP_OPTION_SPACE - remaining;
> > > }
> > >
> > > - if (cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS)) {
> > > - unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> > > -
> > > - remaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
> > > - remaining);
> > > + remaining = MAX_TCP_OPTION_SPACE - size;
> > > + remaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
> > > + remaining);
> > >
> > > - size = MAX_TCP_OPTION_SPACE - remaining;
> > > - }
> > > + size = MAX_TCP_OPTION_SPACE - remaining;
> >
> > This static-key check is not redundant unless
> > bpf_tcp_ops_hdr_opt_len() is inlined. In my build,
> > tcp_established_options() unconditionally calls the out-of-line
> > helper, so sockets pay the call/prologue cost even when no bpf_tcp_ops
> > is attached.
>
> Interesting, I guess you don't use FDO ?
Thanks for explaining. Yes. I did not use FDO for my local build.
>
> On my FDO build, even tcp_established_options() is inlined
> to __tcp_transmit_skb() and both bpf_tcp_ops_hdr_opt_len()
> and bpf_tcp_ops_hdr_opt_len() are NOPs as expected.
>
> ffffffff8224a511: 41 bc 28 00 00 00 movl $0x28, %r12d
> # remaining = MAX_TCP_OPTION_SPACE (40)
> ffffffff8224a517: 45 29 ec subl %r13d, %r12d
> # remaining -= size
> ffffffff8224a51a: 0f 1f 44 00 00 nopl (%rax,%rax)
> # bpf_skops_hdr_opt_len()
> ffffffff8224a51f: 44 89 64 24 78 movl %r12d, 0x78(%rsp)
> # remaining_out = remaining
> ffffffff8224a524: 0f 1f 44 00 00 nopl (%rax,%rax)
> # bpf_tcp_ops_hdr_opt_len()
>
> I will add a static inline function and move both
> cgroup_bpf_enabled() there and reuse it in 3 places.
> (same for bpf_tcp_ops_parse_hdr())
>
> Then, the unnecessary movl will also disappear.
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 05/10] bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after cgroup_bpf_enabled(CGROUP_SOCK_OPS).
2026-10-06 19:24 ` [PATCH v4 bpf-next 05/10] bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after cgroup_bpf_enabled(CGROUP_SOCK_OPS) Kuniyuki Iwashima
2026-10-06 20:11 ` bot+bpf-ci
2026-10-06 22:05 ` Stanislav Fomichev
@ 2026-10-07 22:38 ` Amery Hung
2 siblings, 0 replies; 36+ messages in thread
From: Amery Hung @ 2026-10-07 22:38 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On Tue, Oct 6, 2026 at 12:26 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>
> BPF_CGROUP_RUN_PROG_XXX() macros guard __cgroup_bpf_run_filter_XXX()
> with cgroup_bpf_enabled().
>
> However, even when no SOCK_OPS prog is attached, callers still
> initialise struct bpf_sock_ops_kern (memset(), etc.) or evaluate
> BPF_SOCK_OPS_TEST_FLAG(), which loads tp->bpf_sock_ops_cb_flags
> from a cold cacheline near the end of struct tcp_sock.
>
> Similar to bpf_tcp_ops, let's check cgroup_bpf_enabled() before
> BPF_SOCK_OPS_TEST_FLAG() and struct bpf_sock_ops_kern setup, and
> rename BPF_CGROUP_RUN_PROG_SOCK_OPS{,_SK}() with __ prefix.
>
> Since all direct callers of tcp_call_bpf() pass 0 and NULL for
> nargs and args, they can be folded into the new tcp_call_bpf()
> macro.
>
> All callers of tcp_call_bpf_{2,3}arg() check BPF_SOCK_OPS_TEST_FLAG()
> and do not need the return value. These are replaced with the new
> tcp_call_bpf_flag() macro.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Amery Hung <ameryhung@gmail.com>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 08/10] bpf: mptcp: Don't support BPF_TCP_OPS_FLAG_RCVQ.
2026-10-06 19:24 ` [PATCH v4 bpf-next 08/10] bpf: mptcp: Don't support BPF_TCP_OPS_FLAG_RCVQ Kuniyuki Iwashima
2026-10-06 22:06 ` Stanislav Fomichev
@ 2026-10-07 22:43 ` Amery Hung
1 sibling, 0 replies; 36+ messages in thread
From: Amery Hung @ 2026-10-07 22:43 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On Tue, Oct 6, 2026 at 12:26 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>
> The next patch exposes a new kfunc calling __tcp_set_rcvlowat()
> to bpf_tcp_ops.
>
> MPTCP has its own sock->ops->set_rcvlowat() / mptcp_set_rcvlowat(),
> so we should not allow calling __tcp_set_rcvlowat() on MPTCP
> subflows.
>
> Let's disable BPF_TCP_OPS_FLAG_RCVQ for MPTCP for now.
>
> If needed in the future, bpf_tcp_ops_set_rcvlowat() could be
> extended to properly support MPTCP.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
Reviewed-by: Amery Hung <ameryhung@gmail.com>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 07/10] tcp: Split out __tcp_set_rcvlowat().
2026-10-06 19:24 ` [PATCH v4 bpf-next 07/10] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
@ 2026-10-07 22:43 ` Amery Hung
0 siblings, 0 replies; 36+ messages in thread
From: Amery Hung @ 2026-10-07 22:43 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev,
Emil Tsalapatis
On Tue, Oct 6, 2026 at 12:26 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>
> We will add a kfunc for bpf_tcp_ops.{enqueue,dequeue}_rcvq()
> to adjust sk->sk_rcvlowat.
>
> These hooks are triggered
>
> * when the TCP stack enqueues an skb to sk->sk_receive_queue
> * after data is dequeued from sk->sk_receive_queue
>
> In the enqueue path, tcp_data_ready() is always called after
> the hooks in tcp_queue_rcv() and tcp_ofo_queue().
>
> If tcp_set_rcvlowat() were used as is, tcp_data_ready() could
> be called twice for the same skb, which is redundant and also
> confusing.
>
> Let's split out __tcp_set_rcvlowat() and add a flag to control
> wakeup behaviour.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> Acked-by: Stanislav Fomichev <sdf@fomichev.me>
> Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Reviewed-by: Amery Hung <ameryhung@gmail.com>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 09/10] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
2026-10-06 19:24 ` [PATCH v4 bpf-next 09/10] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
2026-10-06 19:42 ` sashiko-bot
@ 2026-10-07 23:06 ` Amery Hung
2026-10-07 23:21 ` Kuniyuki Iwashima
1 sibling, 1 reply; 36+ messages in thread
From: Amery Hung @ 2026-10-07 23:06 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev,
Emil Tsalapatis
On Tue, Oct 6, 2026 at 12:26 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>
> bpf_tcp_ops.{enqueue,dequeue}_rcvq() were added to parse skb and
> adjust sk->sk_rcvlowat dynamically to suppress unnecessary wakeups.
>
> Let's add a new kfunc to set sk->sk_rcvlowat.
>
> Negative values are clamped to INT_MAX, consistent with SO_RCVLOWAT.
>
> For enqueue_rcvq(), wakeup is set to false because:
>
> * tcp_data_ready() is always called after the hooks in
> tcp_queue_rcv() and tcp_ofo_queue().
>
> * when tcp_fastopen_add_skb() is called for TFO SYN, the socket is
> not yet accept()ed, and when called for TFO SYN+ACK, the socket
> is woken up by sk->sk_state_change() anyway.
>
> For dequeue_rcvq(), wakeup is set to true because tcp_data_ready()
> is not called in that path.
>
> An alternative would be to support bpf_setsockopt() for these
> hooks.
>
> However, that approach involves excessive conditionals and an
> unnecessary memcpy(), costs we do not want to pay for every skb
> in the TCP fast path.
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> Acked-by: Stanislav Fomichev <sdf@fomichev.me>
> Tested-by: Clément Léger <cleger@meta.com>
> Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
> ---
> net/ipv4/bpf_tcp_ops.c | 46 ++++++++++++++++++++++++++++++++++++++++++
> 1 file changed, 46 insertions(+)
>
> diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> index 8182037c4269..2ba73dd6c52c 100644
> --- a/net/ipv4/bpf_tcp_ops.c
> +++ b/net/ipv4/bpf_tcp_ops.c
> @@ -361,12 +361,31 @@ __bpf_kfunc int bpf_tcp_ops_set_flags(struct tcp_sock *tp, u32 enable, u32 disab
> return 0;
> }
>
> +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> + const struct bpf_prog_aux *aux)
> +{
Is there a reason this takes struct sock * instead of struct tcp_sock
*, like bpf_tcp_ops_set_flags() in patch 2?
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 09/10] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
2026-10-07 23:06 ` Amery Hung
@ 2026-10-07 23:21 ` Kuniyuki Iwashima
2026-10-07 23:42 ` Amery Hung
0 siblings, 1 reply; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-07 23:21 UTC (permalink / raw)
To: Amery Hung
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev,
Emil Tsalapatis
On Wed, Oct 7, 2026 at 4:06 PM Amery Hung <ameryhung@gmail.com> wrote:
>
> On Tue, Oct 6, 2026 at 12:26 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
> >
> > bpf_tcp_ops.{enqueue,dequeue}_rcvq() were added to parse skb and
> > adjust sk->sk_rcvlowat dynamically to suppress unnecessary wakeups.
> >
> > Let's add a new kfunc to set sk->sk_rcvlowat.
> >
> > Negative values are clamped to INT_MAX, consistent with SO_RCVLOWAT.
> >
> > For enqueue_rcvq(), wakeup is set to false because:
> >
> > * tcp_data_ready() is always called after the hooks in
> > tcp_queue_rcv() and tcp_ofo_queue().
> >
> > * when tcp_fastopen_add_skb() is called for TFO SYN, the socket is
> > not yet accept()ed, and when called for TFO SYN+ACK, the socket
> > is woken up by sk->sk_state_change() anyway.
> >
> > For dequeue_rcvq(), wakeup is set to true because tcp_data_ready()
> > is not called in that path.
> >
> > An alternative would be to support bpf_setsockopt() for these
> > hooks.
> >
> > However, that approach involves excessive conditionals and an
> > unnecessary memcpy(), costs we do not want to pay for every skb
> > in the TCP fast path.
> >
> > Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> > Acked-by: Stanislav Fomichev <sdf@fomichev.me>
> > Tested-by: Clément Léger <cleger@meta.com>
> > Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
> > ---
> > net/ipv4/bpf_tcp_ops.c | 46 ++++++++++++++++++++++++++++++++++++++++++
> > 1 file changed, 46 insertions(+)
> >
> > diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> > index 8182037c4269..2ba73dd6c52c 100644
> > --- a/net/ipv4/bpf_tcp_ops.c
> > +++ b/net/ipv4/bpf_tcp_ops.c
> > @@ -361,12 +361,31 @@ __bpf_kfunc int bpf_tcp_ops_set_flags(struct tcp_sock *tp, u32 enable, u32 disab
> > return 0;
> > }
> >
> > +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> > + const struct bpf_prog_aux *aux)
> > +{
>
> Is there a reason this takes struct sock * instead of struct tcp_sock
> *, like bpf_tcp_ops_set_flags() in patch 2?
It also works, but bpf prog needs to cast sk to tcp_sock.
This patch is straightforward and rather patch 2 has a
reason to enforce bpf tcp_sk helper use in cgroup hooks.
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 09/10] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat.
2026-10-07 23:21 ` Kuniyuki Iwashima
@ 2026-10-07 23:42 ` Amery Hung
0 siblings, 0 replies; 36+ messages in thread
From: Amery Hung @ 2026-10-07 23:42 UTC (permalink / raw)
To: Kuniyuki Iwashima
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev,
Emil Tsalapatis
On Wed, Oct 7, 2026 at 4:22 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>
> On Wed, Oct 7, 2026 at 4:06 PM Amery Hung <ameryhung@gmail.com> wrote:
> >
> > On Tue, Oct 6, 2026 at 12:26 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
> > >
> > > bpf_tcp_ops.{enqueue,dequeue}_rcvq() were added to parse skb and
> > > adjust sk->sk_rcvlowat dynamically to suppress unnecessary wakeups.
> > >
> > > Let's add a new kfunc to set sk->sk_rcvlowat.
> > >
> > > Negative values are clamped to INT_MAX, consistent with SO_RCVLOWAT.
> > >
> > > For enqueue_rcvq(), wakeup is set to false because:
> > >
> > > * tcp_data_ready() is always called after the hooks in
> > > tcp_queue_rcv() and tcp_ofo_queue().
> > >
> > > * when tcp_fastopen_add_skb() is called for TFO SYN, the socket is
> > > not yet accept()ed, and when called for TFO SYN+ACK, the socket
> > > is woken up by sk->sk_state_change() anyway.
> > >
> > > For dequeue_rcvq(), wakeup is set to true because tcp_data_ready()
> > > is not called in that path.
> > >
> > > An alternative would be to support bpf_setsockopt() for these
> > > hooks.
> > >
> > > However, that approach involves excessive conditionals and an
> > > unnecessary memcpy(), costs we do not want to pay for every skb
> > > in the TCP fast path.
> > >
> > > Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> > > Acked-by: Stanislav Fomichev <sdf@fomichev.me>
> > > Tested-by: Clément Léger <cleger@meta.com>
> > > Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
> > > ---
> > > net/ipv4/bpf_tcp_ops.c | 46 ++++++++++++++++++++++++++++++++++++++++++
> > > 1 file changed, 46 insertions(+)
> > >
> > > diff --git a/net/ipv4/bpf_tcp_ops.c b/net/ipv4/bpf_tcp_ops.c
> > > index 8182037c4269..2ba73dd6c52c 100644
> > > --- a/net/ipv4/bpf_tcp_ops.c
> > > +++ b/net/ipv4/bpf_tcp_ops.c
> > > @@ -361,12 +361,31 @@ __bpf_kfunc int bpf_tcp_ops_set_flags(struct tcp_sock *tp, u32 enable, u32 disab
> > > return 0;
> > > }
> > >
> > > +__bpf_kfunc int bpf_tcp_ops_set_rcvlowat(struct sock *sk, int rcvlowat,
> > > + const struct bpf_prog_aux *aux)
> > > +{
> >
> > Is there a reason this takes struct sock * instead of struct tcp_sock
> > *, like bpf_tcp_ops_set_flags() in patch 2?
>
> It also works, but bpf prog needs to cast sk to tcp_sock.
>
> This patch is straightforward and rather patch 2 has a
> reason to enforce bpf tcp_sk helper use in cgroup hooks.
Make sense. Thanks.
Reviewed-by: Amery Hung <ameryhung@gmail.com>
^ permalink raw reply [flat|nested] 36+ messages in thread
* Re: [PATCH v4 bpf-next 04/10] bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag.
2026-10-07 22:36 ` Amery Hung
@ 2026-10-08 2:10 ` Kuniyuki Iwashima
0 siblings, 0 replies; 36+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-08 2:10 UTC (permalink / raw)
To: Amery Hung
Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Martin KaFai Lau, Eduard Zingerman, Kumar Kartikeya Dwivedi,
Yonghong Song, John Fastabend, Stanislav Fomichev, Eric Dumazet,
Neal Cardwell, Willem de Bruijn, Tenzin Ukyab,
Clément Léger, Kuniyuki Iwashima, bpf, netdev
On Wed, Oct 7, 2026 at 3:36 PM Amery Hung <ameryhung@gmail.com> wrote:
>
> On Tue, Oct 6, 2026 at 7:48 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
> >
> > On Tue, Oct 6, 2026 at 4:25 PM Amery Hung <ameryhung@gmail.com> wrote:
> > >
> > > > +++ b/net/ipv4/tcp_input.c
> > > > @@ -210,6 +210,11 @@ static void bpf_skops_established(struct sock *sk, int bpf_op,
> > > >
> > > > static void bpf_tcp_ops_parse_hdr(struct sock *sk, struct sk_buff *skb)
> > > > {
> > > > + const struct tcp_sock *tp;
> > > > +
> > > > + if (!cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS))
> > > > + return;
> > > > +
> > > > switch (sk->sk_state) {
> > > > case TCP_SYN_RECV:
> > > > case TCP_SYN_SENT:
> > > > @@ -217,7 +222,12 @@ static void bpf_tcp_ops_parse_hdr(struct sock *sk, struct sk_buff *skb)
> > > > return;
> > > > }
> > > >
> > > > - bpf_tcp_ops_call(parse_hdr, sk, skb);
> > > > + tp = tcp_sk(sk);
> > > > +
> > > > + if ((tp->rx_opt.saw_unknown &&
> > > > + BPF_TCP_OPS_TEST_FLAG(tp, PARSE_HDR_OPT_UNKNOWN)) ||
> > > > + BPF_TCP_OPS_TEST_FLAG(tp, PARSE_HDR_OPT_ALL))
> > >
> > > nit: ^ missing a space
> >
> > this indentation is correct.
>
> Sorry. False alarm.
>
> >
> > >
> > > > + __bpf_tcp_ops_call(parse_hdr, sk, skb);
> > > > }
> > > >
> > > > static __cold void tcp_gro_dev_warn(const struct sock *sk, const struct sk_buff *skb,
> > > > diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c
> > > > index 3ac465513bf5..8770f3084efe 100644
> > > > --- a/net/ipv4/tcp_output.c
> > > > +++ b/net/ipv4/tcp_output.c
> > > > @@ -581,8 +581,8 @@ static void bpf_skops_write_hdr_opt(struct sock *sk, struct sk_buff *skb,
> > > > * writer's bytes). The writer finds the append point by scanning from
> > > > * first_opt_off + nr_written to the first NOP.
> > > > */
> > > > - bpf_tcp_ops_call(write_hdr_opt, sk, skb, req, syn_skb, synack_type,
> > > > - first_opt_off + nr_written);
> > > > + bpf_tcp_ops_call_flag(write_hdr_opt, WRITE_HDR_OPT, sk, skb, req,
> > > > + syn_skb, synack_type, first_opt_off + nr_written);
> > > > }
> > > > #else
> > > > static u32 bpf_skops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,
> > > > @@ -613,11 +613,13 @@ static u32 bpf_tcp_ops_hdr_opt_len(struct sock *sk, struct sk_buff *skb,
> > > > {
> > > > unsigned int remaining_out = remaining, reserved;
> > > >
> > > > - if (!remaining)
> > > > - return 0;
> > > > + if (!cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS) ||
> > > > + !BPF_TCP_OPS_TEST_FLAG(tcp_sk(sk), WRITE_HDR_OPT)||
> > >
> > > nit: missing a space before ||
> >
> > oops, due to last minute edit :/
> >
> > >
> > > > + !remaining)
> > > > + return remaining;
> > > >
> > > > /* bpf_tcp_ops_reserve_hdr_opt() reserves space via remaining_out */
> > > > - bpf_tcp_ops_call(hdr_opt_len, sk, skb, req, syn_skb, synack_type, &remaining_out);
> > > > + __bpf_tcp_ops_call(hdr_opt_len, sk, skb, req, syn_skb, synack_type, &remaining_out);
> > > >
> > > > reserved = remaining - remaining_out;
> > > > if (!reserved)
> > > > @@ -1193,8 +1195,9 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> > > > struct tcp_key *key)
> > > > {
> > > > struct tcp_sock *tp = tcp_sk(sk);
> > > > - unsigned int size = 0;
> > > > unsigned int eff_sacks;
> > > > + unsigned int remaining;
> > > > + unsigned int size = 0;
> > > >
> > > > opts->options = 0;
> > > > opts->bpf_opt_len = 0;
> > > > @@ -1224,10 +1227,10 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> > > > * left.
> > > > */
> > > > if (sk_is_mptcp(sk)) {
> > > > - unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> > > > bool has_ts = opts->options & OPTION_TS;
> > > > int opt_size;
> > > >
> > > > + remaining = MAX_TCP_OPTION_SPACE - size;
> > > > opts->mptcp.drop_ts = 0;
> > > >
> > > > opt_size = mptcp_established_options(sk, skb, remaining, has_ts,
> > > > @@ -1244,7 +1247,8 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> > > >
> > > > eff_sacks = tp->rx_opt.num_sacks + tp->rx_opt.dsack;
> > > > if (unlikely(eff_sacks)) {
> > > > - const unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> > > > + remaining = MAX_TCP_OPTION_SPACE - size;
> > > > +
> > > > if (likely(remaining >= TCPOLEN_SACK_BASE_ALIGNED +
> > > > TCPOLEN_SACK_PERBLOCK)) {
> > > > opts->num_sack_blocks =
> > > > @@ -1277,22 +1281,18 @@ static unsigned int tcp_established_options(struct sock *sk, struct sk_buff *skb
> > > >
> > > > if (unlikely(BPF_SOCK_OPS_TEST_FLAG(tp,
> > > > BPF_SOCK_OPS_WRITE_HDR_OPT_CB_FLAG))) {
> > > > - unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> > > > -
> > > > + remaining = MAX_TCP_OPTION_SPACE - size;
> > > > remaining = bpf_skops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
> > > > remaining);
> > > >
> > > > size = MAX_TCP_OPTION_SPACE - remaining;
> > > > }
> > > >
> > > > - if (cgroup_bpf_enabled(CGROUP_TCP_SOCK_OPS)) {
> > > > - unsigned int remaining = MAX_TCP_OPTION_SPACE - size;
> > > > -
> > > > - remaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
> > > > - remaining);
> > > > + remaining = MAX_TCP_OPTION_SPACE - size;
> > > > + remaining = bpf_tcp_ops_hdr_opt_len(sk, skb, NULL, NULL, 0, opts,
> > > > + remaining);
> > > >
> > > > - size = MAX_TCP_OPTION_SPACE - remaining;
> > > > - }
> > > > + size = MAX_TCP_OPTION_SPACE - remaining;
> > >
> > > This static-key check is not redundant unless
> > > bpf_tcp_ops_hdr_opt_len() is inlined. In my build,
> > > tcp_established_options() unconditionally calls the out-of-line
> > > helper, so sockets pay the call/prologue cost even when no bpf_tcp_ops
> > > is attached.
> >
> > Interesting, I guess you don't use FDO ?
>
> Thanks for explaining. Yes. I did not use FDO for my local build.
I updated patch 4&5 and now I don't see callq for
bpf_tcp_ops_hdr_opt_len() even without FDO.
I still see callq for bpf_skops_write_hdr_opt() but it can be
follow-up if needed.
I'll post v5 shortly.
Thanks !
>
> >
> > On my FDO build, even tcp_established_options() is inlined
> > to __tcp_transmit_skb() and both bpf_tcp_ops_hdr_opt_len()
> > and bpf_tcp_ops_hdr_opt_len() are NOPs as expected.
> >
> > ffffffff8224a511: 41 bc 28 00 00 00 movl $0x28, %r12d
> > # remaining = MAX_TCP_OPTION_SPACE (40)
> > ffffffff8224a517: 45 29 ec subl %r13d, %r12d
> > # remaining -= size
> > ffffffff8224a51a: 0f 1f 44 00 00 nopl (%rax,%rax)
> > # bpf_skops_hdr_opt_len()
> > ffffffff8224a51f: 44 89 64 24 78 movl %r12d, 0x78(%rsp)
> > # remaining_out = remaining
> > ffffffff8224a524: 0f 1f 44 00 00 nopl (%rax,%rax)
> > # bpf_tcp_ops_hdr_opt_len()
> >
> > I will add a static inline function and move both
> > cgroup_bpf_enabled() there and reuse it in 3 places.
> > (same for bpf_tcp_ops_parse_hdr())
> >
> > Then, the unnecessary movl will also disappear.
^ permalink raw reply [flat|nested] 36+ messages in thread
end of thread, other threads:[~2026-10-08 2:10 UTC | newest]
Thread overview: 36+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-06 19:24 [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Kuniyuki Iwashima
2026-10-06 19:24 ` [PATCH v4 bpf-next 01/10] bpf: tcp: Convert deny-list for bpf_{get,set}sockopt() to allow-list Kuniyuki Iwashima
2026-10-06 23:09 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 02/10] bpf: tcp: Add a new per-socket flag and kfunc for bpf_tcp_ops Kuniyuki Iwashima
2026-10-06 22:04 ` Stanislav Fomichev
2026-10-06 23:10 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 03/10] selftest: bpf: Use bpf_tcp_ops_set_flags() in bpf_tcp_ops_hdr.c Kuniyuki Iwashima
2026-10-06 22:04 ` Stanislav Fomichev
2026-10-06 23:12 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 04/10] bpf: tcp: Guard fast-path bpf_tcp_ops_call() under per-socket flag Kuniyuki Iwashima
2026-10-06 22:05 ` Stanislav Fomichev
2026-10-06 23:25 ` Amery Hung
2026-10-07 2:48 ` Kuniyuki Iwashima
2026-10-07 22:36 ` Amery Hung
2026-10-08 2:10 ` Kuniyuki Iwashima
2026-10-06 19:24 ` [PATCH v4 bpf-next 05/10] bpf: tcp: Check BPF_SOCK_OPS_TEST_FLAG() after cgroup_bpf_enabled(CGROUP_SOCK_OPS) Kuniyuki Iwashima
2026-10-06 20:11 ` bot+bpf-ci
2026-10-06 22:05 ` Stanislav Fomichev
2026-10-07 22:38 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 06/10] bpf: tcp: Introduce bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-10-06 22:06 ` Stanislav Fomichev
2026-10-06 23:26 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 07/10] tcp: Split out __tcp_set_rcvlowat() Kuniyuki Iwashima
2026-10-07 22:43 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 08/10] bpf: mptcp: Don't support BPF_TCP_OPS_FLAG_RCVQ Kuniyuki Iwashima
2026-10-06 22:06 ` Stanislav Fomichev
2026-10-07 22:43 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 09/10] bpf: tcp: Add kfunc to adjust sk->sk_rcvlowat Kuniyuki Iwashima
2026-10-06 19:42 ` sashiko-bot
2026-10-07 23:06 ` Amery Hung
2026-10-07 23:21 ` Kuniyuki Iwashima
2026-10-07 23:42 ` Amery Hung
2026-10-06 19:24 ` [PATCH v4 bpf-next 10/10] selftest: bpf: Add test for bpf_tcp_ops.{enqueue,dequeue}_rcvq() Kuniyuki Iwashima
2026-10-06 19:43 ` sashiko-bot
2026-10-06 22:06 ` Stanislav Fomichev
2026-10-07 13:11 ` [PATCH v4 bpf-next 00/10] bpf: Add bpf_tcp_ops hooks for TCP AutoLOWAT Jakub Sitnicki
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox