From: Wang Zhan <wang.zhan@smartx.com>
To: netdev@vger.kernel.org
Cc: davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
pabeni@redhat.com, horms@kernel.org, keyong.sun@smartx.com,
Willem de Bruijn <willemdebruijn.kernel@gmail.com>,
Jason Wang <jasowangio@gmail.com>,
Andrew Lunn <andrew+netdev@lunn.ch>,
Aaron Conole <aconole@redhat.com>,
Eelco Chaudron <echaudro@redhat.com>,
Ilya Maximets <i.maximets@ovn.org>,
dev@openvswitch.org, Daniel Borkmann <daniel@iogearbox.net>,
Neal Cardwell <ncardwell@google.com>,
Kuniyuki Iwashima <kuniyu@google.com>,
Alice Mikityanska <alice@isovalent.com>,
David Laight <david.laight.linux@gmail.com>,
Wang Zhan <wang.zhan@smartx.com>,
Willem de Bruijn <willemb@google.com>
Subject: [PATCH net-next v5 5/6] net: core: re-segment oversized TCP GSO skbs
Date: Thu, 8 Oct 2026 18:06:50 +0800 [thread overview]
Message-ID: <20261008100651.2534957-6-wang.zhan@smartx.com> (raw)
In-Reply-To: <20261008100651.2534957-1-wang.zhan@smartx.com>
An skb which exceeds an egress device limit loses its GSO feature mask and
is segmented into individual packets, although the device can still offload
smaller TCP GSO skbs. A BIG TCP hop which feeds a non-BIG TCP one pays
that on every flow.
An unencapsulated TCP GSO skb which exceeds gso_max_size or gso_max_segs is
re-segmented with a max_segs which keeps each output inside that limit. An
encapsulated or frag-list skb, a GSO type the device cannot offload, a
max_segs which leaves room for a single MSS and an skb whose transport header
is unset or stale all keep today's segmentation.
This path emits a plain GSO skb, not a BIG TCP one: inet_gso_segment() and
ipv6_gso_segment() write each output's whole length into the 16-bit L3
length field, so the size limit is also capped at what that field can
express.
The helper runs on the skb handed to the driver and only from the
netif_needs_gso() branch: an skb the device takes as it is pays nothing, and
an over-limit one pays the TCP header read and the features recomputation.
Measured on a veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP
enabled on the veth endpoints and left off in the guest, so the skbs which
the veth hop accepts have to be segmented before the TAP device. A single
iperf3 TCP flow, six alternating runs per state (`-t 15 -O 5`, fixed CPU
affinity and port tuple). The middle column is the same tree with the
re-segmentation disabled:
protocol no BIG TCP mixed, no reseg mixed, resegmented
TCP/IPv4 51.550 Gbps 15.850 Gbps 52.617 Gbps
TCP/IPv6 52.050 Gbps 15.783 Gbps 51.933 Gbps
Coefficient of variation for the two mixed columns was 0.48% and 0.82% for
IPv4 and 0.44% and 0.44% for IPv6. A BIG TCP hop which feeds a 64 KiB hop
loses 69% of the throughput of a path which never enables BIG TCP at all;
re-segmentation recovers it, 3.3x over the existing segmentation path and
within noise of the no BIG TCP baseline.
Reviewed-by: Willem de Bruijn <willemb@google.com>
Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
---
v4: https://lore.kernel.org/20260930111526.2183107-5-wang.zhan@smartx.com/
v3: https://lore.kernel.org/20260928044102.1004310-5-wang.zhan@smartx.com/
v2: https://lore.kernel.org/20260918084651.3022878-4-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-4-wang.zhan@smartx.com/
---
net/core/dev.c | 53 +++++++++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 52 insertions(+), 1 deletion(-)
diff --git a/net/core/dev.c b/net/core/dev.c
index d6e39df55cd8cf..ef64834520949f 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3996,6 +3996,53 @@ __netif_skb_features(struct sk_buff *skb, bool check_gso_limits)
}
EXPORT_SYMBOL(__netif_skb_features);
+static unsigned int
+skb_gso_output_max_segs(struct sk_buff *skb, struct net_device *dev)
+{
+ unsigned int mss = skb_shinfo(skb)->gso_size;
+ unsigned int gso_max_size, hdr_len, max_segs;
+ netdev_features_t features;
+ struct tcphdr _tcph, *th;
+
+ if (!skb_is_gso(skb) || !skb_is_gso_tcp(skb) ||
+ skb->encapsulation || skb_has_frag_list(skb))
+ return 0;
+
+ /*
+ * The transport offset can be unset or stale, so hdr_len is only taken
+ * from a header at the checksum start.
+ */
+ if (!skb_transport_header_was_set(skb) ||
+ skb_checksum_start(skb) != skb_transport_header(skb))
+ return 0;
+
+ th = skb_header_pointer(skb, skb_transport_offset(skb), sizeof(_tcph),
+ &_tcph);
+ if (!th || __tcp_hdrlen(th) < sizeof(*th))
+ return 0;
+
+ hdr_len = skb_transport_header(skb) - skb_mac_header(skb) +
+ __tcp_hdrlen(th);
+
+ /* The output is still a GSO skb: recheck without the limit checks */
+ features = __netif_skb_features(skb, false) | NETIF_F_GSO_ROBUST;
+ if (!net_gso_ok(features, skb_shinfo(skb)->gso_type))
+ return 0;
+
+ gso_max_size = netif_get_gso_max_size(dev, skb);
+ gso_max_size = min(gso_max_size, GSO_LEGACY_MAX_SIZE);
+
+ /*
+ * gso_within_dev_limits() accepts gso_segs == gso_max_segs but
+ * rejects skb->len >= gso_max_size, so only the size limit needs the
+ * - 1; the inner min() keeps that subtraction from wrapping.
+ */
+ max_segs = (gso_max_size - min(gso_max_size, hdr_len + 1)) / mss;
+ max_segs = min(max_segs, READ_ONCE(dev->gso_max_segs));
+
+ return max_segs;
+}
+
static int xmit_one(struct sk_buff *skb, struct net_device *dev,
struct netdev_queue *txq, bool more)
{
@@ -4159,9 +4206,13 @@ static struct sk_buff *validate_xmit_skb(struct sk_buff *skb, struct net_device
goto out_null;
if (netif_needs_gso(skb, features)) {
+ unsigned int max_segs = 0;
struct sk_buff *segs;
- segs = skb_gso_segment(skb, features);
+ if (unlikely(!gso_within_dev_limits(skb, dev)))
+ max_segs = skb_gso_output_max_segs(skb, dev);
+
+ segs = __skb_gso_segment(skb, features, true, max_segs);
if (IS_ERR(segs)) {
goto out_kfree_skb;
} else if (segs) {
--
2.47.3
next prev parent reply other threads:[~2026-10-08 10:07 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-08 10:06 [PATCH net-next v5 0/6] net: re-segment oversized TCP GSO skbs Wang Zhan
2026-10-08 10:06 ` [PATCH net-next v5 1/6] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
2026-10-08 10:06 ` [PATCH net-next v5 2/6] net: core: factor out the GSO device limit check Wang Zhan
2026-10-08 10:06 ` [PATCH net-next v5 3/6] net: gso: support re-segmentation of TCP GSO skbs Wang Zhan
2026-10-08 10:06 ` [PATCH net-next v5 4/6] net: gso: clear the GSO state of the last output segment Wang Zhan
2026-10-08 18:37 ` Willem de Bruijn
2026-10-08 10:06 ` Wang Zhan [this message]
2026-10-08 10:06 ` [PATCH net-next v5 6/6] net: net_test: add tests for TCP re-segmentation Wang Zhan
2026-10-08 18:38 ` Willem de Bruijn
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261008100651.2534957-6-wang.zhan@smartx.com \
--to=wang.zhan@smartx.com \
--cc=aconole@redhat.com \
--cc=alice@isovalent.com \
--cc=andrew+netdev@lunn.ch \
--cc=daniel@iogearbox.net \
--cc=davem@davemloft.net \
--cc=david.laight.linux@gmail.com \
--cc=dev@openvswitch.org \
--cc=echaudro@redhat.com \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=i.maximets@ovn.org \
--cc=jasowangio@gmail.com \
--cc=keyong.sun@smartx.com \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=ncardwell@google.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=willemb@google.com \
--cc=willemdebruijn.kernel@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox