* [PATCH v6 net-next 1/1] tcp: Replace min_tso_segs() with tso_segs() CC callback
@ 2026-07-30 11:49 chia-yu.chang
2026-07-31 11:49 ` sashiko-bot
2026-08-04 14:08 ` Paolo Abeni
0 siblings, 2 replies; 4+ messages in thread
From: chia-yu.chang @ 2026-07-30 11:49 UTC (permalink / raw)
To: emil, jolsa, yonghong.song, song, linux-kselftest, memxor, shuah,
martin.lau, ast, daniel, andrii, eddyz87, horms, dsahern, bpf,
netdev, pabeni, jhs, kuba, stephen, davem, edumazet,
andrew+netdev, donald.hunter, kuniyu, ij, ncardwell,
koen.de_schepper, g.white, ingemar.s.johansson, mirja.kuehlewind,
cheshire, rs.ietf, Jason_Livingood, vidhi_goel
Cc: Chia-Yu Chang
From: Chia-Yu Chang <chia-yu.chang@nokia-bell-labs.com>
This patch replaces existing min_tso_segs() with tso_segs() CC callback
for CC algorithm to provide explicit tso segment number of each data
burst and overrides tcp_tso_autosize().
This change provides below impacts on BPF struct_ops users:
- The callback is renamed from min_tso_segs() to tso_segs()
- The signature gains an extra u32 mss_now argument
- The return value semantics is changed from "floor value passed into
tcp_tso_autosize()" to "final tso_segs value", bypassing autosizing
As a result, BPF programs shall be updated, because returning a small
constant will now directly limit the final tso_segs value instead of
specifying the minimum value passed to tcp_tso_autosize().
Signed-off-by: Ilpo Järvinen <ij@kernel.org>
Signed-off-by: Chia-Yu Chang <chia-yu.chang@nokia-bell-labs.com>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
---
v6:
- Add clamp_t to avoid returning 0 by ca_ops->tso_segs()
- Update commit message
v5:
- Revert back to v3 and add Reviewed-by tag
v4:
- Use union for both min_tso_segs() and tso_segs() and a
v3:
- Update bpf_tcp_ca_tso_segs() to use tcp_tso_autosize()
- Add divide by 0 protection in case of mss_now=0
v2:
- Export tcp_tso_autosize()
---
include/net/tcp.h | 12 ++++++++++--
net/ipv4/bpf_tcp_ca.c | 9 ++++++---
net/ipv4/tcp_bbr.c | 13 ++++++++++---
net/ipv4/tcp_output.c | 15 ++++++++-------
tools/testing/selftests/bpf/progs/tcp_ca_kfunc.c | 8 ++++----
5 files changed, 38 insertions(+), 19 deletions(-)
diff --git a/include/net/tcp.h b/include/net/tcp.h
index 2c5b889530b5..ef9975112f15 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -824,6 +824,9 @@ unsigned int tcp_sync_mss(struct sock *sk, u32 pmtu);
unsigned int tcp_current_mss(struct sock *sk);
u32 tcp_clamp_probe0_to_user_timeout(const struct sock *sk, u32 when);
+u32 tcp_tso_autosize(const struct sock *sk, unsigned int mss_now,
+ int min_tso_segs);
+
/* Bound MSS / TSO packet size with the half of the window */
static inline int tcp_bound_to_half_wnd(struct tcp_sock *tp, int pktsize)
{
@@ -1361,8 +1364,13 @@ struct tcp_congestion_ops {
/* hook for packet ack accounting (optional) */
void (*pkts_acked)(struct sock *sk, const struct ack_sample *sample);
- /* override sysctl_tcp_min_tso_segs (optional) */
- u32 (*min_tso_segs)(struct sock *sk);
+ /* Override tcp_tso_autosize (optional)
+ *
+ * If provided, this callback returns the final TSO segment number
+ * and will bypass tcp_tso_autosize() entirely. The implementation
+ * must derive an appropriate value and ensure the result is valid.
+ */
+ u32 (*tso_segs)(struct sock *sk, u32 mss_now);
/* new value of cwnd after loss (required) */
u32 (*undo_cwnd)(struct sock *sk);
diff --git a/net/ipv4/bpf_tcp_ca.c b/net/ipv4/bpf_tcp_ca.c
index 791e15063237..e5faa38bdd25 100644
--- a/net/ipv4/bpf_tcp_ca.c
+++ b/net/ipv4/bpf_tcp_ca.c
@@ -284,9 +284,12 @@ static void bpf_tcp_ca_pkts_acked(struct sock *sk, const struct ack_sample *samp
{
}
-static u32 bpf_tcp_ca_min_tso_segs(struct sock *sk)
+static u32 bpf_tcp_ca_tso_segs(struct sock *sk, u32 mss_now)
{
- return 0;
+ if (unlikely(!mss_now))
+ /* No TSO sizing decision. Let caller apply its limits. */
+ return U32_MAX;
+ return tcp_tso_autosize(sk, mss_now, 0);
}
static void bpf_tcp_ca_cong_control(struct sock *sk, u32 ack, int flag,
@@ -320,7 +323,7 @@ static struct tcp_congestion_ops __bpf_ops_tcp_congestion_ops = {
.cwnd_event_tx_start = bpf_tcp_ca_cwnd_event_tx_start,
.in_ack_event = bpf_tcp_ca_in_ack_event,
.pkts_acked = bpf_tcp_ca_pkts_acked,
- .min_tso_segs = bpf_tcp_ca_min_tso_segs,
+ .tso_segs = bpf_tcp_ca_tso_segs,
.cong_control = bpf_tcp_ca_cong_control,
.undo_cwnd = bpf_tcp_ca_undo_cwnd,
.sndbuf_expand = bpf_tcp_ca_sndbuf_expand,
diff --git a/net/ipv4/tcp_bbr.c b/net/ipv4/tcp_bbr.c
index 82378a2bfd1e..ecf11be46f38 100644
--- a/net/ipv4/tcp_bbr.c
+++ b/net/ipv4/tcp_bbr.c
@@ -297,11 +297,18 @@ static void bbr_set_pacing_rate(struct sock *sk, u32 bw, int gain)
}
/* override sysctl_tcp_min_tso_segs */
-__bpf_kfunc static u32 bbr_min_tso_segs(struct sock *sk)
+static u32 bbr_min_tso_segs(struct sock *sk)
{
return READ_ONCE(sk->sk_pacing_rate) < (bbr_min_tso_rate >> 3) ? 1 : 2;
}
+__bpf_kfunc static u32 bbr_tso_segs(struct sock *sk, u32 mss_now)
+{
+ if (unlikely(!mss_now))
+ return bbr_min_tso_segs(sk);
+ return tcp_tso_autosize(sk, mss_now, bbr_min_tso_segs(sk));
+}
+
static u32 bbr_tso_segs_goal(struct sock *sk)
{
struct tcp_sock *tp = tcp_sk(sk);
@@ -1151,7 +1158,7 @@ static struct tcp_congestion_ops tcp_bbr_cong_ops __read_mostly = {
.undo_cwnd = bbr_undo_cwnd,
.cwnd_event_tx_start = bbr_cwnd_event_tx_start,
.ssthresh = bbr_ssthresh,
- .min_tso_segs = bbr_min_tso_segs,
+ .tso_segs = bbr_tso_segs,
.get_info = bbr_get_info,
.set_state = bbr_set_state,
};
@@ -1163,7 +1170,7 @@ BTF_ID_FLAGS(func, bbr_sndbuf_expand)
BTF_ID_FLAGS(func, bbr_undo_cwnd)
BTF_ID_FLAGS(func, bbr_cwnd_event_tx_start)
BTF_ID_FLAGS(func, bbr_ssthresh)
-BTF_ID_FLAGS(func, bbr_min_tso_segs)
+BTF_ID_FLAGS(func, bbr_tso_segs)
BTF_ID_FLAGS(func, bbr_set_state)
BTF_KFUNCS_END(tcp_bbr_check_kfunc_ids)
diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c
index d7c1444b5e30..d9bda3a84f26 100644
--- a/net/ipv4/tcp_output.c
+++ b/net/ipv4/tcp_output.c
@@ -2253,8 +2253,8 @@ static bool tcp_nagle_check(bool partial, const struct tcp_sock *tp,
* for every 2^9 usec (aka 512 us) of RTT, so that the RTT-based allowance
* is below 1500 bytes after 6 * ~500 usec = 3ms.
*/
-static u32 tcp_tso_autosize(const struct sock *sk, unsigned int mss_now,
- int min_tso_segs)
+u32 tcp_tso_autosize(const struct sock *sk, unsigned int mss_now,
+ int min_tso_segs)
{
unsigned long bytes;
u32 r;
@@ -2269,6 +2269,7 @@ static u32 tcp_tso_autosize(const struct sock *sk, unsigned int mss_now,
return max_t(u32, bytes / mss_now, min_tso_segs);
}
+EXPORT_SYMBOL(tcp_tso_autosize);
/* Return the number of segments we want in the skb we are transmitting.
* See if congestion control module wants to decide; otherwise, autosize.
@@ -2278,12 +2279,12 @@ static u32 tcp_tso_segs(struct sock *sk, unsigned int mss_now)
const struct tcp_congestion_ops *ca_ops = inet_csk(sk)->icsk_ca_ops;
u32 min_tso, tso_segs;
- min_tso = ca_ops->min_tso_segs ?
- ca_ops->min_tso_segs(sk) :
- READ_ONCE(sock_net(sk)->ipv4.sysctl_tcp_min_tso_segs);
+ min_tso = READ_ONCE(sock_net(sk)->ipv4.sysctl_tcp_min_tso_segs);
- tso_segs = tcp_tso_autosize(sk, mss_now, min_tso);
- return min_t(u32, tso_segs, sk->sk_gso_max_segs);
+ tso_segs = ca_ops->tso_segs ?
+ ca_ops->tso_segs(sk, mss_now) :
+ tcp_tso_autosize(sk, mss_now, min_tso);
+ return clamp_t(u32, tso_segs, 1, sk->sk_gso_max_segs);
}
/* Returns the portion of skb which can be sent right away */
diff --git a/tools/testing/selftests/bpf/progs/tcp_ca_kfunc.c b/tools/testing/selftests/bpf/progs/tcp_ca_kfunc.c
index 0a3e9d35bf6f..58262e490336 100644
--- a/tools/testing/selftests/bpf/progs/tcp_ca_kfunc.c
+++ b/tools/testing/selftests/bpf/progs/tcp_ca_kfunc.c
@@ -10,7 +10,7 @@ extern u32 bbr_sndbuf_expand(struct sock *sk) __ksym;
extern u32 bbr_undo_cwnd(struct sock *sk) __ksym;
extern void bbr_cwnd_event_tx_start(struct sock *sk) __ksym;
extern u32 bbr_ssthresh(struct sock *sk) __ksym;
-extern u32 bbr_min_tso_segs(struct sock *sk) __ksym;
+extern u32 bbr_tso_segs(struct sock *sk, u32 mss_now) __ksym;
extern void bbr_set_state(struct sock *sk, u8 new_state) __ksym;
extern void dctcp_init(struct sock *sk) __ksym;
@@ -90,9 +90,9 @@ u32 BPF_PROG(ssthresh, struct sock *sk)
}
SEC("struct_ops")
-u32 BPF_PROG(min_tso_segs, struct sock *sk)
+u32 BPF_PROG(tso_segs, struct sock *sk, u32 mss_now)
{
- return bbr_min_tso_segs(sk);
+ return bbr_tso_segs(sk, mss_now);
}
SEC("struct_ops")
@@ -120,7 +120,7 @@ struct tcp_congestion_ops tcp_ca_kfunc = {
.cwnd_event = (void *)cwnd_event,
.cwnd_event_tx_start = (void *)cwnd_event_tx_start,
.ssthresh = (void *)ssthresh,
- .min_tso_segs = (void *)min_tso_segs,
+ .tso_segs = (void *)tso_segs,
.set_state = (void *)set_state,
.pkts_acked = (void *)pkts_acked,
.name = "tcp_ca_kfunc",
--
2.34.1
^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [PATCH v6 net-next 1/1] tcp: Replace min_tso_segs() with tso_segs() CC callback
2026-07-30 11:49 [PATCH v6 net-next 1/1] tcp: Replace min_tso_segs() with tso_segs() CC callback chia-yu.chang
@ 2026-07-31 11:49 ` sashiko-bot
2026-08-04 14:08 ` Paolo Abeni
1 sibling, 0 replies; 4+ messages in thread
From: sashiko-bot @ 2026-07-31 11:49 UTC (permalink / raw)
To: chia-yu.chang; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] tcp_tso_autosize is not exported as a BPF kfunc, making it impossible for custom BPF CC programs to utilize standard TSO autosizing.
--
--- Patch [1]: [PATCH v6 net-next 1/1] tcp: Replace min_tso_segs() with tso_segs() CC callback ---
commit e9e6269217c20aef163ee8ebfb7cab71f918807f
Author: Chia-Yu Chang <chia-yu.chang@nokia-bell-labs.com>
tcp: Replace min_tso_segs() with tso_segs() CC callback
This patch replaces the existing min_tso_segs() CC callback with a new
tso_segs() callback. The new callback provides the explicit TSO segment
number for each data burst, completely bypassing tcp_tso_autosize().
> diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c
> index d7c1444b5e304..d9bda3a84f261 100644
> --- a/net/ipv4/tcp_output.c
> +++ b/net/ipv4/tcp_output.c
[ ... ]
> @@ -2269,6 +2269,7 @@ static u32 tcp_tso_autosize(const struct sock *sk, unsigned int mss_now,
>
> return max_t(u32, bytes / mss_now, min_tso_segs);
> }
> +EXPORT_SYMBOL(tcp_tso_autosize);
[Severity: High]
Since BPF congestion control programs now need to compute the final TSO
segments value, do they need access to tcp_tso_autosize() to retain standard
autosizing?
The commit exports tcp_tso_autosize for native kernel modules, but it appears
missing from the BTF kfuncs set (bpf_tcp_ca_check_kfunc_ids) in
net/ipv4/bpf_tcp_ca.c:
BTF_KFUNCS_START(bpf_tcp_ca_check_kfunc_ids)
BTF_ID_FLAGS(func, tcp_reno_ssthresh)
BTF_ID_FLAGS(func, tcp_reno_cong_avoid)
...
BTF_KFUNCS_END(bpf_tcp_ca_check_kfunc_ids)
Without being registered here, custom BPF CC programs cannot call
tcp_tso_autosize(). This might prevent them from properly migrating to the
new tso_segs callback if they want to retain standard TSO autosizing.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260730114919.658373-1-chia-yu.chang@nokia-bell-labs.com?part=1
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH v6 net-next 1/1] tcp: Replace min_tso_segs() with tso_segs() CC callback
2026-07-30 11:49 [PATCH v6 net-next 1/1] tcp: Replace min_tso_segs() with tso_segs() CC callback chia-yu.chang
2026-07-31 11:49 ` sashiko-bot
@ 2026-08-04 14:08 ` Paolo Abeni
2026-08-13 12:15 ` Chia-Yu Chang (Nokia)
1 sibling, 1 reply; 4+ messages in thread
From: Paolo Abeni @ 2026-08-04 14:08 UTC (permalink / raw)
To: chia-yu.chang, emil, jolsa, yonghong.song, song, linux-kselftest,
memxor, shuah, martin.lau, ast, daniel, andrii, eddyz87, horms,
dsahern, bpf, netdev, jhs, kuba, stephen, davem, edumazet,
andrew+netdev, donald.hunter, kuniyu, ij, ncardwell,
koen.de_schepper, g.white, ingemar.s.johansson, mirja.kuehlewind,
cheshire, rs.ietf, Jason_Livingood, vidhi_goel
On 7/30/26 1:49 PM, chia-yu.chang@nokia-bell-labs.com wrote:
> From: Chia-Yu Chang <chia-yu.chang@nokia-bell-labs.com>
>
> This patch replaces existing min_tso_segs() with tso_segs() CC callback
> for CC algorithm to provide explicit tso segment number of each data
> burst and overrides tcp_tso_autosize().
>
> This change provides below impacts on BPF struct_ops users:
> - The callback is renamed from min_tso_segs() to tso_segs()
> - The signature gains an extra u32 mss_now argument
> - The return value semantics is changed from "floor value passed into
> tcp_tso_autosize()" to "final tso_segs value", bypassing autosizing
>
> As a result, BPF programs shall be updated, because returning a small
> constant will now directly limit the final tso_segs value instead of
> specifying the minimum value passed to tcp_tso_autosize().
>
> Signed-off-by: Ilpo Järvinen <ij@kernel.org>
> Signed-off-by: Chia-Yu Chang <chia-yu.chang@nokia-bell-labs.com>
> Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
Sashiko nipa has more feedback, that looks relevant to me even if it's
tagged not hi prio:
https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260730114919.658373-1-chia-yu.chang%40nokia-bell-labs.com
/P
^ permalink raw reply [flat|nested] 4+ messages in thread
* RE: [PATCH v6 net-next 1/1] tcp: Replace min_tso_segs() with tso_segs() CC callback
2026-08-04 14:08 ` Paolo Abeni
@ 2026-08-13 12:15 ` Chia-Yu Chang (Nokia)
0 siblings, 0 replies; 4+ messages in thread
From: Chia-Yu Chang (Nokia) @ 2026-08-13 12:15 UTC (permalink / raw)
To: Paolo Abeni, emil@etsalapatis.com, jolsa@kernel.org,
yonghong.song@linux.dev, song@kernel.org,
linux-kselftest@vger.kernel.org, memxor@gmail.com,
shuah@kernel.org, martin.lau@linux.dev, ast@kernel.org,
daniel@iogearbox.net, andrii@kernel.org, eddyz87@gmail.com,
horms@kernel.org, dsahern@kernel.org, bpf@vger.kernel.org,
netdev@vger.kernel.org, jhs@mojatatu.com, kuba@kernel.org,
stephen@networkplumber.org, davem@davemloft.net,
edumazet@google.com, andrew+netdev@lunn.ch,
donald.hunter@gmail.com, kuniyu@google.com, ij@kernel.org,
ncardwell@google.com, Koen De Schepper (Nokia),
g.white@cablelabs.com, ingemar.s.johansson@ericsson.com,
mirja.kuehlewind@ericsson.com, cheshire@apple.com, rs.ietf@gmx.at,
Jason_Livingood@comcast.com, vidhi_goel@apple.com
> -----Original Message-----
> From: Paolo Abeni <pabeni@redhat.com>
> Sent: Tuesday, August 4, 2026 4:09 PM
> To: Chia-Yu Chang (Nokia) <chia-yu.chang@nokia-bell-labs.com>; emil@etsalapatis.com; jolsa@kernel.org; yonghong.song@linux.dev; song@kernel.org; linux-kselftest@vger.kernel.org; memxor@gmail.com; shuah@kernel.org; martin.lau@linux.dev; ast@kernel.org; daniel@iogearbox.net; andrii@kernel.org; eddyz87@gmail.com; horms@kernel.org; dsahern@kernel.org; bpf@vger.kernel.org; netdev@vger.kernel.org; jhs@mojatatu.com; kuba@kernel.org; stephen@networkplumber.org; davem@davemloft.net; edumazet@google.com; andrew+netdev@lunn.ch; donald.hunter@gmail.com; kuniyu@google.com; ij@kernel.org; ncardwell@google.com; Koen De Schepper (Nokia) <koen.de_schepper@nokia-bell-labs.com>; g.white@cablelabs.com; ingemar.s.johansson@ericsson.com; mirja.kuehlewind@ericsson.com; cheshire@apple.com; rs.ietf@gmx.at; Jason_Livingood@comcast.com; vidhi_goel@apple.com
> Subject: Re: [PATCH v6 net-next 1/1] tcp: Replace min_tso_segs() with tso_segs() CC callback
>
>
> CAUTION: This is an external email. Please be very careful when clicking links or opening attachments. See the URL nok.it/ext for additional information.
>
>
>
> On 7/30/26 1:49 PM, chia-yu.chang@nokia-bell-labs.com wrote:
> > From: Chia-Yu Chang <chia-yu.chang@nokia-bell-labs.com>
> >
> > This patch replaces existing min_tso_segs() with tso_segs() CC
> > callback for CC algorithm to provide explicit tso segment number of
> > each data burst and overrides tcp_tso_autosize().
> >
> > This change provides below impacts on BPF struct_ops users:
> > - The callback is renamed from min_tso_segs() to tso_segs()
> > - The signature gains an extra u32 mss_now argument
> > - The return value semantics is changed from "floor value passed into
> > tcp_tso_autosize()" to "final tso_segs value", bypassing autosizing
> >
> > As a result, BPF programs shall be updated, because returning a small
> > constant will now directly limit the final tso_segs value instead of
> > specifying the minimum value passed to tcp_tso_autosize().
> >
> > Signed-off-by: Ilpo Järvinen <ij@kernel.org>
> > Signed-off-by: Chia-Yu Chang <chia-yu.chang@nokia-bell-labs.com>
> > Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
> Sashiko nipa has more feedback, that looks relevant to me even if it's tagged not hi prio:
>
> https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260730114919.658373-1-chia-yu.chang%40nokia-bell-labs.com
>
> /P
Hi Paolo,
Thanks for the pointer and I have a look on all comments.
About the first medium comment, this is indeed an API change because we were asked not to introduce an additional callback.
Emil reviewed this from the BPF perspective and agreed with the approach.
I do agree that the callback comment can be documented more clearly.
So I'll clarify in v7 that implementations must handle mss_now==0, and that returning 0 results in the caller clamping the value to a single segment.
Regarding the second medium comment, this version does not expose tcp_tso_autosize() as a BPF kfunc.
The initial goal was to support in-kernel congestion control only.
If BPF maintainers think it would be useful, I will expose it as a BPF kfunc in a separated patch of the same series.
For the low-priority comments, I agree and will:
- Change bpf_tcp_ca_tso_segs() to return 0.
- Change EXPORT_SYMBOL() to EXPORT_SYMBOL_GPL().
- Move the READ_ONCE() into the fallback branch.
Thanks!
Chia-Yu
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-08-13 12:15 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-30 11:49 [PATCH v6 net-next 1/1] tcp: Replace min_tso_segs() with tso_segs() CC callback chia-yu.chang
2026-07-31 11:49 ` sashiko-bot
2026-08-04 14:08 ` Paolo Abeni
2026-08-13 12:15 ` Chia-Yu Chang (Nokia)
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.