From: Wang Zhan <wang.zhan@smartx.com>
To: netdev@vger.kernel.org
Cc: davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
pabeni@redhat.com, horms@kernel.org, keyong.sun@smartx.com,
Willem de Bruijn <willemdebruijn.kernel@gmail.com>,
Jason Wang <jasowangio@gmail.com>,
Andrew Lunn <andrew+netdev@lunn.ch>,
Aaron Conole <aconole@redhat.com>,
Eelco Chaudron <echaudro@redhat.com>,
Ilya Maximets <i.maximets@ovn.org>,
dev@openvswitch.org, Daniel Borkmann <daniel@iogearbox.net>,
Neal Cardwell <ncardwell@google.com>,
Kuniyuki Iwashima <kuniyu@google.com>,
Alice Mikityanska <alice@isovalent.com>,
David Laight <david.laight.linux@gmail.com>,
Wang Zhan <wang.zhan@smartx.com>
Subject: [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs
Date: Mon, 28 Sep 2026 12:40:57 +0800 [thread overview]
Message-ID: <20260928044102.1004310-1-wang.zhan@smartx.com> (raw)
BIG TCP is negotiated per netdevice, so BIG TCP and non-BIG TCP ports can
coexist in one path. When an skb exceeds the GSO limits of the port it is
sent to, it loses its GSO feature mask and is segmented into individual MSS
sized packets, which that port then sends without TSO.
This series cuts such a TCP GSO skb into GSO skbs which fit the device
limits instead, so the rest of the path keeps using TSO. On a
veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP enabled on the
veth endpoints and left off in the guest, a single iperf3 TCP flow, six
alternating runs per state (-t 15 -O 5, fixed CPU affinity and port tuple):
protocol no BIG TCP mixed, no reseg mixed, resegmented
TCP/IPv4 51.550 Gbps 15.850 Gbps 52.617 Gbps
TCP/IPv6 52.050 Gbps 15.783 Gbps 51.933 Gbps
The middle column is the same tree with the bounded path disabled. A BIG
TCP hop which feeds a 64 KiB hop loses 69% of the throughput of a path
which never enables BIG TCP; bounded resegmentation recovers it.
1/5 stands on its own as a fix: the size limit is picked from
skb->protocol, which is the VLAN ethertype for a frame which already
carries its tag in the packet, so an IPv6 frame was measured against the
IPv4 limit. 2/5 is the preparation which lets the limit tests be skipped
for one caller, and carries no functional change.
The new path is taken only when the skb is an unencapsulated TCP GSO skb
without a frag_list, it exceeds gso_max_size or gso_max_segs, and the
device offloads that GSO type. Everything else keeps today's
segmentation.
The output obeys the GSO feature and limit contract the device already
advertises, so this needs no new UAPI, no device state and no driver
change, and it applies automatically. The bound only says how the output
is grouped; an over-limit skb pays one extra ndo_features_check() in
exchange for staying a GSO skb.
Alternatives considered:
- The caller could set skb_shinfo(skb)->gso_size to ~64K and adjust the
gso bits in the shared info afterwards, which would need
skb_unclone(), a repeat of the grouping logic skb_segment() already
has, and a recomputed IPv4 ID for the DF=0 case.
Patch layout:
[1/5] the GSO size limit follows the packet's L3 protocol
[2/5] factor the device limit check out of gso_features_check()
[3/5] let the GSO engine bound the MSS segments per output skb
[4/5] apply that bound to oversized TCP GSO skbs in the TX path
[5/5] KUnit coverage for the bound, the device limits and the TCP path
---
v3:
- patch 1 is new: the size limit follows the packet's L3 protocol
- patch 2: move the check_gso_limits flag and the wrapper split in from patch 4
- patch 3: drop the tcp_gso_segment() exception, the caller keeps the features
- patch 3: cap the bound with GSO_MAX_SEGS and drop the output reset
- patch 3: document the bound as TCP only and note the frag_list gate
- patch 4: enter from the limit predicate gso_features_check() uses
- patch 4: keep the features, so the bound shapes the output only
- patch 4: drop the SG and checksum tests, fold the MSS minimum
- patch 4: cap each output at GSO_LEGACY_MAX_SIZE, not the BIG TCP size
- patch 5: cover the IPv4 and the IPv6 limit, with and without the tag
- patch 5: skip the TCP cases without CONFIG_INET, reserve headroom
- patch 5: free on the failure paths, check the ungrouped single MSS
v2: https://lore.kernel.org/20260918084651.3022878-1-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-1-wang.zhan@smartx.com/
Wang Zhan (5):
net: core: use the packet's L3 protocol for the GSO size limit
net: core: factor out the GSO device limit check
net: gso: support bounded TCP segmentation
net: core: resegment oversized TCP GSO skbs
net: net_test: add tests for bounded GSO segmentation
drivers/net/tap.c | 3 +-
include/linux/netdevice.h | 4 +-
include/net/gso.h | 6 +-
include/net/udp.h | 2 +-
net/core/dev.c | 91 ++++++++--
net/core/gso.c | 6 +-
net/core/net_test.c | 347 +++++++++++++++++++++++++++++++++++++
net/core/skbuff.c | 14 +-
net/openvswitch/datapath.c | 2 +-
9 files changed, 455 insertions(+), 20 deletions(-)
base-commit: 014d795c73837ea2339a4ea8e8f82c6e959b845d
--
2.47.3
next reply other threads:[~2026-09-28 4:41 UTC|newest]
Thread overview: 27+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 4:40 Wang Zhan [this message]
2026-09-28 4:40 ` [PATCH net-next v3 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
2026-09-28 23:36 ` Willem de Bruijn
2026-09-29 3:44 ` Wang Zhan
2026-09-29 4:01 ` Wang Zhan
2026-09-29 14:59 ` Willem de Bruijn
2026-09-28 4:40 ` [PATCH net-next v3 2/5] net: core: factor out the GSO device limit check Wang Zhan
2026-09-28 23:37 ` Willem de Bruijn
2026-09-29 7:47 ` Paolo Abeni
2026-09-28 4:41 ` [PATCH net-next v3 3/5] net: gso: support bounded TCP segmentation Wang Zhan
2026-09-28 23:39 ` Willem de Bruijn
2026-09-30 4:41 ` netdev-bot+sashiko
2026-09-28 4:41 ` [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs Wang Zhan
2026-09-28 23:47 ` Willem de Bruijn
2026-09-29 10:25 ` Wang Zhan
2026-09-29 15:00 ` Willem de Bruijn
2026-09-30 4:41 ` netdev-bot+sashiko
2026-09-28 4:41 ` [PATCH net-next v3 5/5] net: net_test: add tests for bounded GSO segmentation Wang Zhan
2026-09-28 23:59 ` Willem de Bruijn
2026-09-29 10:30 ` Wang Zhan
2026-09-29 15:01 ` Willem de Bruijn
2026-09-30 4:41 ` netdev-bot+sashiko
2026-09-28 4:45 ` [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs netdev-bot+sinfo
2026-09-28 5:49 ` Wang Zhan
2026-09-28 23:34 ` Willem de Bruijn
2026-09-29 7:27 ` Paolo Abeni
2026-09-29 11:50 ` Wang Zhan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928044102.1004310-1-wang.zhan@smartx.com \
--to=wang.zhan@smartx.com \
--cc=aconole@redhat.com \
--cc=alice@isovalent.com \
--cc=andrew+netdev@lunn.ch \
--cc=daniel@iogearbox.net \
--cc=davem@davemloft.net \
--cc=david.laight.linux@gmail.com \
--cc=dev@openvswitch.org \
--cc=echaudro@redhat.com \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=i.maximets@ovn.org \
--cc=jasowangio@gmail.com \
--cc=keyong.sun@smartx.com \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=ncardwell@google.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=willemdebruijn.kernel@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox