Netdev List
 help / color / mirror / Atom feed
* [PATCH net-next v5 0/6] net: re-segment oversized TCP GSO skbs
@ 2026-10-08 10:06 Wang Zhan
  2026-10-08 10:06 ` [PATCH net-next v5 1/6] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
                   ` (5 more replies)
  0 siblings, 6 replies; 9+ messages in thread
From: Wang Zhan @ 2026-10-08 10:06 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

BIG TCP is negotiated per netdevice, so BIG TCP and non-BIG TCP ports can
coexist in one path.  When an skb exceeds the GSO limits of the port it is
sent to, it loses its GSO feature mask and is segmented into individual MSS
sized packets, which that port then sends without TSO.

This series cuts such a TCP GSO skb into GSO skbs which fit the device
limits instead, so the rest of the path keeps using TSO.  On a
veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP enabled on the
veth endpoints and left off in the guest, a single iperf3 TCP flow, six
alternating runs per state (-t 15 -O 5, fixed CPU affinity and port tuple):

  protocol  no BIG TCP   mixed, no reseg  mixed, resegmented
  TCP/IPv4  51.550 Gbps  15.850 Gbps      52.617 Gbps
  TCP/IPv6  52.050 Gbps  15.783 Gbps      51.933 Gbps

The middle column is the same tree with the re-segmentation disabled.  A
BIG TCP hop which feeds a 64 KiB hop loses 69% of the throughput of a path
which never enables BIG TCP; re-segmentation recovers it.

1/6 moves the GSO size limit lookup into dev.c and makes it follow the
packet's L3 protocol, so where the tag sits does not decide which limit
applies.  2/6 is the preparation which lets the limit tests be skipped for
one caller, and carries no functional change.  3/6 lets the GSO engine
group several MSS into one output skb, and 4/6 clears the whole GSO state
of the last output of such a group.

The new path is taken only when the skb is an unencapsulated TCP GSO skb
without a frag_list, it exceeds gso_max_size or gso_max_segs, and the
device offloads that GSO type.  Everything else keeps today's segmentation.

The output obeys the GSO feature and limit contract the device already
advertises, so this needs no new UAPI, no device state and no driver
change, and it applies automatically.  max_segs only says how the output is
grouped; an over-limit skb pays one extra ndo_features_check() in exchange
for staying a GSO skb.

Alternatives considered:

  - The caller could set skb_shinfo(skb)->gso_size to ~64K and adjust the
    gso bits in the shared info afterwards, which would need
    skb_unclone(), a repeat of the grouping logic skb_segment() already
    has, and a recomputed IPv4 ID for the DF=0 case.

Patch layout:

  [1/6] the GSO size limit follows the packet's L3 protocol
  [2/6] factor the device limit check out of gso_features_check()
  [3/6] let the GSO engine group several MSS into one output skb
  [4/6] clear the whole GSO state of the last output segment
  [5/6] re-segment oversized TCP GSO skbs in the TX path
  [6/6] KUnit coverage for re-segmentation

---

v5:
- patch 3: make max_segs a u8, skb_gso_cb has one byte left
- patch 3: only warn about the input segment count without max_segs
- patch 4: new, clear the whole GSO state of the last output segment
- patch 6: add a plain-tail case, drop the TCP path tests
v4: https://lore.kernel.org/20260930111526.2183107-1-wang.zhan@smartx.com/
v3: https://lore.kernel.org/20260928044102.1004310-1-wang.zhan@smartx.com/
v2: https://lore.kernel.org/20260918084651.3022878-1-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-1-wang.zhan@smartx.com/

Wang Zhan (6):
  net: core: use the packet's L3 protocol for the GSO size limit
  net: core: factor out the GSO device limit check
  net: gso: support re-segmentation of TCP GSO skbs
  net: gso: clear the GSO state of the last output segment
  net: core: re-segment oversized TCP GSO skbs
  net: net_test: add tests for TCP re-segmentation

 drivers/net/tap.c          |  3 +-
 include/linux/netdevice.h  | 18 ++++----
 include/net/gso.h          |  7 ++-
 include/net/udp.h          |  2 +-
 net/core/dev.c             | 92 +++++++++++++++++++++++++++++++++-----
 net/core/gso.c             |  6 ++-
 net/core/net_test.c        | 87 +++++++++++++++++++++++++++++++++++
 net/core/skbuff.c          | 12 +++--
 net/openvswitch/datapath.c |  2 +-
 9 files changed, 198 insertions(+), 31 deletions(-)


base-commit: 8df0638138d3e0344fd1fb36cf2d1ca1cf5028f0
-- 
2.47.3

^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH net-next v5 1/6] net: core: use the packet's L3 protocol for the GSO size limit
  2026-10-08 10:06 [PATCH net-next v5 0/6] net: re-segment oversized TCP GSO skbs Wang Zhan
@ 2026-10-08 10:06 ` Wang Zhan
  2026-10-08 10:06 ` [PATCH net-next v5 2/6] net: core: factor out the GSO device limit check Wang Zhan
                   ` (4 subsequent siblings)
  5 siblings, 0 replies; 9+ messages in thread
From: Wang Zhan @ 2026-10-08 10:06 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan, Willem de Bruijn

gso_features_check() compares the frame length against
netif_get_gso_max_size(), which picks the IPv4 or the IPv6 limit from
skb->protocol.  A tag which is already inside the frame replaces that field
with the VLAN ethertype, as skb_vlan_push() does, and an IPv6 packet is
then measured against the IPv4 limit and segmented although the device
could send it as one TSO frame.

The lookup moves from netdevice.h into dev.c: vlan_get_protocol() reads the
L3 protocol from behind an in-frame tag, while if_vlan.h includes
netdevice.h, so an inline in that header cannot call it.  dev.c is the only
file which calls the lookup, so it becomes a static helper there.

Reviewed-by: Willem de Bruijn <willemb@google.com>
Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>

---
v4: https://lore.kernel.org/20260930111526.2183107-2-wang.zhan@smartx.com/
v3: https://lore.kernel.org/20260928044102.1004310-2-wang.zhan@smartx.com/
---
 include/linux/netdevice.h | 9 ---------
 net/core/dev.c            | 9 +++++++++
 2 files changed, 9 insertions(+), 9 deletions(-)

diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index 8cae9b00211eee..4819acbc06ead9 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -5564,15 +5564,6 @@ netif_get_gro_max_size(const struct net_device *dev, const struct sk_buff *skb)
 	       READ_ONCE(dev->gro_ipv4_max_size);
 }
 
-static inline unsigned int
-netif_get_gso_max_size(const struct net_device *dev, const struct sk_buff *skb)
-{
-	/* pairs with WRITE_ONCE() in netif_set_gso(_ipv4)_max_size() */
-	return skb->protocol == htons(ETH_P_IPV6) ?
-	       READ_ONCE(dev->gso_max_size) :
-	       READ_ONCE(dev->gso_ipv4_max_size);
-}
-
 static inline bool netif_is_macsec(const struct net_device *dev)
 {
 	return dev->priv_flags & IFF_MACSEC;
diff --git a/net/core/dev.c b/net/core/dev.c
index f587645e930a15..0aedea7c9ffb79 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3888,6 +3888,15 @@ static bool skb_has_ipv6_extension_hdr(const struct sk_buff *skb)
 	return false;
 }
 
+static unsigned int
+netif_get_gso_max_size(const struct net_device *dev, const struct sk_buff *skb)
+{
+	/* pairs with WRITE_ONCE() in netif_set_gso(_ipv4)_max_size() */
+	return vlan_get_protocol(skb) == htons(ETH_P_IPV6) ?
+	       READ_ONCE(dev->gso_max_size) :
+	       READ_ONCE(dev->gso_ipv4_max_size);
+}
+
 static netdev_features_t gso_features_check(const struct sk_buff *skb,
 					    struct net_device *dev,
 					    netdev_features_t features)
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH net-next v5 2/6] net: core: factor out the GSO device limit check
  2026-10-08 10:06 [PATCH net-next v5 0/6] net: re-segment oversized TCP GSO skbs Wang Zhan
  2026-10-08 10:06 ` [PATCH net-next v5 1/6] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
@ 2026-10-08 10:06 ` Wang Zhan
  2026-10-08 10:06 ` [PATCH net-next v5 3/6] net: gso: support re-segmentation of TCP GSO skbs Wang Zhan
                   ` (3 subsequent siblings)
  5 siblings, 0 replies; 9+ messages in thread
From: Wang Zhan @ 2026-10-08 10:06 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan, Willem de Bruijn

gso_features_check() decides whether an egress device can offload a GSO
skb as a single TSO frame by comparing the segment count and the frame
length against the device limits.  The resegmentation path added by a
later patch caps the MSS segments per output skb and needs the same
features computed without those limits, so put the two tests in a helper
and gate them with one flag.

__netif_skb_features() takes that flag and takes the place of
netif_skb_features(), which becomes an inline wrapper in netdevice.h and
passes the flag set.  Existing callers keep their call, and the
resegmentation path asks for the features without the limit checks.

No functional changes.

Reviewed-by: Willem de Bruijn <willemb@google.com>
Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>

---
v4: https://lore.kernel.org/20260930111526.2183107-3-wang.zhan@smartx.com/
v3: https://lore.kernel.org/20260928044102.1004310-3-wang.zhan@smartx.com/
v2: https://lore.kernel.org/20260918084651.3022878-2-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-2-wang.zhan@smartx.com/
---
 include/linux/netdevice.h |  9 ++++++++-
 net/core/dev.c            | 30 ++++++++++++++++++++----------
 2 files changed, 28 insertions(+), 11 deletions(-)

diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index 4819acbc06ead9..feb5fa99c2c629 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -5498,7 +5498,14 @@ void netif_stacked_transfer_operstate(const struct net_device *rootdev,
 netdev_features_t passthru_features_check(struct sk_buff *skb,
 					  struct net_device *dev,
 					  netdev_features_t features);
-netdev_features_t netif_skb_features(struct sk_buff *skb);
+netdev_features_t __netif_skb_features(struct sk_buff *skb,
+				       bool check_gso_limits);
+
+static inline netdev_features_t netif_skb_features(struct sk_buff *skb)
+{
+	return __netif_skb_features(skb, true);
+}
+
 void skb_warn_bad_offload(const struct sk_buff *skb);
 
 static inline bool net_gso_ok(netdev_features_t features, int gso_type)
diff --git a/net/core/dev.c b/net/core/dev.c
index 0aedea7c9ffb79..d6e39df55cd8cf 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3897,16 +3897,24 @@ netif_get_gso_max_size(const struct net_device *dev, const struct sk_buff *skb)
 	       READ_ONCE(dev->gso_ipv4_max_size);
 }
 
+static bool
+gso_within_dev_limits(const struct sk_buff *skb, const struct net_device *dev)
+{
+	if (skb_shinfo(skb)->gso_segs > READ_ONCE(dev->gso_max_segs))
+		return false;
+
+	if (unlikely(skb->len >= netif_get_gso_max_size(dev, skb)))
+		return false;
+
+	return true;
+}
+
 static netdev_features_t gso_features_check(const struct sk_buff *skb,
 					    struct net_device *dev,
-					    netdev_features_t features)
+					    netdev_features_t features,
+					    bool check_limits)
 {
-	u16 gso_segs = skb_shinfo(skb)->gso_segs;
-
-	if (gso_segs > READ_ONCE(dev->gso_max_segs))
-		return features & ~NETIF_F_GSO_MASK;
-
-	if (unlikely(skb->len >= netif_get_gso_max_size(dev, skb)))
+	if (check_limits && !gso_within_dev_limits(skb, dev))
 		return features & ~NETIF_F_GSO_MASK;
 
 	if (!skb_shinfo(skb)->gso_type) {
@@ -3955,13 +3963,15 @@ static netdev_features_t gso_features_check(const struct sk_buff *skb,
 	return features;
 }
 
-netdev_features_t netif_skb_features(struct sk_buff *skb)
+netdev_features_t
+__netif_skb_features(struct sk_buff *skb, bool check_gso_limits)
 {
 	struct net_device *dev = skb->dev;
 	netdev_features_t features = dev->features;
 
 	if (skb_is_gso(skb))
-		features = gso_features_check(skb, dev, features);
+		features = gso_features_check(skb, dev, features,
+					      check_gso_limits);
 
 	/* If encapsulation offload request, verify we are testing
 	 * hardware encapsulation features instead of standard
@@ -3984,7 +3994,7 @@ netdev_features_t netif_skb_features(struct sk_buff *skb)
 
 	return harmonize_features(skb, features);
 }
-EXPORT_SYMBOL(netif_skb_features);
+EXPORT_SYMBOL(__netif_skb_features);
 
 static int xmit_one(struct sk_buff *skb, struct net_device *dev,
 		    struct netdev_queue *txq, bool more)
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH net-next v5 3/6] net: gso: support re-segmentation of TCP GSO skbs
  2026-10-08 10:06 [PATCH net-next v5 0/6] net: re-segment oversized TCP GSO skbs Wang Zhan
  2026-10-08 10:06 ` [PATCH net-next v5 1/6] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
  2026-10-08 10:06 ` [PATCH net-next v5 2/6] net: core: factor out the GSO device limit check Wang Zhan
@ 2026-10-08 10:06 ` Wang Zhan
  2026-10-08 10:06 ` [PATCH net-next v5 4/6] net: gso: clear the GSO state of the last output segment Wang Zhan
                   ` (2 subsequent siblings)
  5 siblings, 0 replies; 9+ messages in thread
From: Wang Zhan @ 2026-10-08 10:06 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan, Willem de Bruijn

A caller which re-segments an oversized TCP GSO skb needs the engine to
group several MSS into one output skb, so let it pass how many MSS segments
each output skb may carry through __skb_gso_segment().  max_segs is a
per-call u8 in the skb_gso_cb scratch context, and only an unencapsulated
TCP skb without a frag_list sets it; every other caller passes zero.

Without max_segs, skb_segment() groups several MSS into one output skb only
for a device which advertises NETIF_F_GSO_PARTIAL, or for a frag_list which
splits into uniform pieces.  max_segs skips that test.

The output stays a GSO skb: gso_size keeps the original MSS and gso_segs the
number of MSS it holds, so a downstream device can still perform TSO.  The
caller keeps the output within the 64 KiB its L3 length field can express.

Reviewed-by: Willem de Bruijn <willemb@google.com>
Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>

---
v5:
- make max_segs a u8, skb_gso_cb has one byte left
- only warn about the input segment count without max_segs
v4: https://lore.kernel.org/20260930111526.2183107-4-wang.zhan@smartx.com/
v3: https://lore.kernel.org/20260928044102.1004310-4-wang.zhan@smartx.com/
v2: https://lore.kernel.org/20260918084651.3022878-3-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-3-wang.zhan@smartx.com/
---
 drivers/net/tap.c          |  3 ++-
 include/net/gso.h          |  7 +++++--
 include/net/udp.h          |  2 +-
 net/core/gso.c             |  6 +++++-
 net/core/skbuff.c          | 10 +++++++---
 net/openvswitch/datapath.c |  2 +-
 6 files changed, 21 insertions(+), 9 deletions(-)

diff --git a/drivers/net/tap.c b/drivers/net/tap.c
index 832439b8a8988d..2f19fb07bb2d42 100644
--- a/drivers/net/tap.c
+++ b/drivers/net/tap.c
@@ -278,9 +278,10 @@ rx_handler_result_t tap_handle_frame(struct sk_buff **pskb)
 	if (q->flags & IFF_VNET_HDR)
 		features |= tap->tap_features;
 	if (netif_needs_gso(skb, features)) {
-		struct sk_buff *segs = __skb_gso_segment(skb, features, false);
+		struct sk_buff *segs;
 		struct sk_buff *next;
 
+		segs = __skb_gso_segment(skb, features, false, 0);
 		if (IS_ERR(segs)) {
 			drop_reason = SKB_DROP_REASON_SKB_GSO_SEG;
 			goto drop;
diff --git a/include/net/gso.h b/include/net/gso.h
index 0749230d414ecb..fe3260b381d4ed 100644
--- a/include/net/gso.h
+++ b/include/net/gso.h
@@ -21,6 +21,8 @@ struct skb_gso_cb {
 	__u16	csum_start;
 	/* Number of IPv4/IPv6 GSO handler entries for this packet. */
 	u8	recursion_counter;
+	/* Max MSS segs per output skb, 0 = no limit */
+	u8	max_segs;
 };
 #define SKB_GSO_CB_OFFSET	32
 #define SKB_GSO_CB(skb) ((struct skb_gso_cb *)((skb)->cb + SKB_GSO_CB_OFFSET))
@@ -83,12 +85,13 @@ static inline __sum16 gso_make_checksum(struct sk_buff *skb, __wsum res)
 }
 
 struct sk_buff *__skb_gso_segment(struct sk_buff *skb,
-				  netdev_features_t features, bool tx_path);
+				  netdev_features_t features, bool tx_path,
+				  unsigned int max_segs);
 
 static inline struct sk_buff *skb_gso_segment(struct sk_buff *skb,
 					      netdev_features_t features)
 {
-	return __skb_gso_segment(skb, features, true);
+	return __skb_gso_segment(skb, features, true, 0);
 }
 
 struct sk_buff *skb_eth_gso_segment(struct sk_buff *skb,
diff --git a/include/net/udp.h b/include/net/udp.h
index 1fee17274745f0..5bc25dcf25fbaa 100644
--- a/include/net/udp.h
+++ b/include/net/udp.h
@@ -613,7 +613,7 @@ static inline struct sk_buff *udp_rcv_segment(struct sock *sk,
 	/* the GSO CB lays after the UDP one, no need to save and restore any
 	 * CB fragment
 	 */
-	segs = __skb_gso_segment(skb, features, false);
+	segs = __skb_gso_segment(skb, features, false, 0);
 	if (IS_ERR_OR_NULL(segs)) {
 		drop_count = skb_shinfo(skb)->gso_segs;
 		goto drop;
diff --git a/net/core/gso.c b/net/core/gso.c
index e96ef635006487..b5bd269e56e4ec 100644
--- a/net/core/gso.c
+++ b/net/core/gso.c
@@ -77,6 +77,8 @@ static bool skb_needs_check(const struct sk_buff *skb, bool tx_path)
  *	@skb: buffer to segment
  *	@features: features for the output path (see dev->features)
  *	@tx_path: whether it is called in TX path
+ *	@max_segs: maximum MSS segments per output GSO skb, 0 means no limit;
+ *		   set only for an unencapsulated TCP skb without a frag_list
  *
  *	This function segments the given skb and returns a list of segments.
  *
@@ -86,7 +88,8 @@ static bool skb_needs_check(const struct sk_buff *skb, bool tx_path)
  *	Segmentation preserves SKB_GSO_CB_OFFSET bytes of previous skb cb.
  */
 struct sk_buff *__skb_gso_segment(struct sk_buff *skb,
-				  netdev_features_t features, bool tx_path)
+				  netdev_features_t features, bool tx_path,
+				  unsigned int max_segs)
 {
 	struct sk_buff *segs;
 
@@ -118,6 +121,7 @@ struct sk_buff *__skb_gso_segment(struct sk_buff *skb,
 	SKB_GSO_CB(skb)->mac_offset = skb_headroom(skb);
 	SKB_GSO_CB(skb)->encap_level = 0;
 	SKB_GSO_CB(skb)->recursion_counter = 0;
+	SKB_GSO_CB(skb)->max_segs = min(max_segs, U8_MAX);
 
 	skb_reset_mac_header(skb);
 	skb_reset_mac_len(skb);
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 43ebe61c7fc481..1feb038d9ff680 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -4793,6 +4793,7 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
 	struct sk_buff *segs = NULL;
 	struct sk_buff *tail = NULL;
 	struct sk_buff *list_skb = skb_shinfo(head_skb)->frag_list;
+	unsigned int max_segs = SKB_GSO_CB(head_skb)->max_segs;
 	unsigned int mss = skb_shinfo(head_skb)->gso_size;
 	bool gso_by_frags = mss == GSO_BY_FRAGS;
 	unsigned int doffset = head_skb->data - skb_mac_header(head_skb);
@@ -4839,7 +4840,7 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
 	csum = !!can_checksum_protocol(features, proto);
 
 	if (sg && csum && !gso_by_frags)  {
-		if (!(features & NETIF_F_GSO_PARTIAL)) {
+		if (!max_segs && !(features & NETIF_F_GSO_PARTIAL)) {
 			struct sk_buff *iter;
 			unsigned int frag_len;
 
@@ -4873,8 +4874,11 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
 		 * doesn't fit into an MSS sized block, so take care of that
 		 * now.
 		 */
-		DEBUG_NET_WARN_ON_ONCE(len / mss > GSO_MAX_SEGS);
-		partial_segs = min(len / mss, GSO_MAX_SEGS);
+		DEBUG_NET_WARN_ON_ONCE(!max_segs && len / mss > GSO_MAX_SEGS);
+		if (max_segs)
+			partial_segs = min(len / mss, max_segs);
+		else
+			partial_segs = min(len / mss, GSO_MAX_SEGS);
 		if (partial_segs > 1)
 			mss *= partial_segs;
 		else
diff --git a/net/openvswitch/datapath.c b/net/openvswitch/datapath.c
index 21870341432552..e793aead68372d 100644
--- a/net/openvswitch/datapath.c
+++ b/net/openvswitch/datapath.c
@@ -375,7 +375,7 @@ static int queue_gso_packets(struct datapath *dp, struct sk_buff *skb,
 	int err;
 
 	BUILD_BUG_ON(sizeof(*OVS_CB(skb)) > SKB_GSO_CB_OFFSET);
-	segs = __skb_gso_segment(skb, NETIF_F_SG, false);
+	segs = __skb_gso_segment(skb, NETIF_F_SG, false, 0);
 	if (IS_ERR(segs))
 		return PTR_ERR(segs);
 	if (segs == NULL)
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH net-next v5 4/6] net: gso: clear the GSO state of the last output segment
  2026-10-08 10:06 [PATCH net-next v5 0/6] net: re-segment oversized TCP GSO skbs Wang Zhan
                   ` (2 preceding siblings ...)
  2026-10-08 10:06 ` [PATCH net-next v5 3/6] net: gso: support re-segmentation of TCP GSO skbs Wang Zhan
@ 2026-10-08 10:06 ` Wang Zhan
  2026-10-08 18:37   ` Willem de Bruijn
  2026-10-08 10:06 ` [PATCH net-next v5 5/6] net: core: re-segment oversized TCP GSO skbs Wang Zhan
  2026-10-08 10:06 ` [PATCH net-next v5 6/6] net: net_test: add tests for TCP re-segmentation Wang Zhan
  5 siblings, 1 reply; 9+ messages in thread
From: Wang Zhan @ 2026-10-08 10:06 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan, Willem de Bruijn

skb_segment() clears gso_size on the last output of a group which holds one
MSS or less, but leaves gso_segs and gso_type as the grouping set them, so
the tail is still counted as the whole group.  Readers which take gso_segs
without testing gso_size first, such as the ECT statistics in ip_rcv_core()
and tp->rcv_ooopack in tcp_rcv_established(), see that stale count, and a
re-segmented output reaches them when a veth or bridge port delivers it back
into the receive path.

Clear gso_size, gso_segs and gso_type together with skb_gso_reset().

Suggested-by: Willem de Bruijn <willemb@google.com>
Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
---
 net/core/skbuff.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 1feb038d9ff680..43d63cd29f694f 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -5125,7 +5125,7 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
 		}
 
 		if (tail->len - doffset <= gso_size)
-			skb_shinfo(tail)->gso_size = 0;
+			skb_gso_reset(tail);
 		else if (tail != segs)
 			skb_shinfo(tail)->gso_segs = DIV_ROUND_UP(tail->len - doffset, gso_size);
 	}
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH net-next v5 5/6] net: core: re-segment oversized TCP GSO skbs
  2026-10-08 10:06 [PATCH net-next v5 0/6] net: re-segment oversized TCP GSO skbs Wang Zhan
                   ` (3 preceding siblings ...)
  2026-10-08 10:06 ` [PATCH net-next v5 4/6] net: gso: clear the GSO state of the last output segment Wang Zhan
@ 2026-10-08 10:06 ` Wang Zhan
  2026-10-08 10:06 ` [PATCH net-next v5 6/6] net: net_test: add tests for TCP re-segmentation Wang Zhan
  5 siblings, 0 replies; 9+ messages in thread
From: Wang Zhan @ 2026-10-08 10:06 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan, Willem de Bruijn

An skb which exceeds an egress device limit loses its GSO feature mask and
is segmented into individual packets, although the device can still offload
smaller TCP GSO skbs.  A BIG TCP hop which feeds a non-BIG TCP one pays
that on every flow.

An unencapsulated TCP GSO skb which exceeds gso_max_size or gso_max_segs is
re-segmented with a max_segs which keeps each output inside that limit.  An
encapsulated or frag-list skb, a GSO type the device cannot offload, a
max_segs which leaves room for a single MSS and an skb whose transport header
is unset or stale all keep today's segmentation.

This path emits a plain GSO skb, not a BIG TCP one: inet_gso_segment() and
ipv6_gso_segment() write each output's whole length into the 16-bit L3
length field, so the size limit is also capped at what that field can
express.

The helper runs on the skb handed to the driver and only from the
netif_needs_gso() branch: an skb the device takes as it is pays nothing, and
an over-limit one pays the TCP header read and the features recomputation.

Measured on a veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP
enabled on the veth endpoints and left off in the guest, so the skbs which
the veth hop accepts have to be segmented before the TAP device.  A single
iperf3 TCP flow, six alternating runs per state (`-t 15 -O 5`, fixed CPU
affinity and port tuple).  The middle column is the same tree with the
re-segmentation disabled:

  protocol  no BIG TCP   mixed, no reseg  mixed, resegmented
  TCP/IPv4  51.550 Gbps  15.850 Gbps      52.617 Gbps
  TCP/IPv6  52.050 Gbps  15.783 Gbps      51.933 Gbps

Coefficient of variation for the two mixed columns was 0.48% and 0.82% for
IPv4 and 0.44% and 0.44% for IPv6.  A BIG TCP hop which feeds a 64 KiB hop
loses 69% of the throughput of a path which never enables BIG TCP at all;
re-segmentation recovers it, 3.3x over the existing segmentation path and
within noise of the no BIG TCP baseline.

Reviewed-by: Willem de Bruijn <willemb@google.com>
Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>

---
v4: https://lore.kernel.org/20260930111526.2183107-5-wang.zhan@smartx.com/
v3: https://lore.kernel.org/20260928044102.1004310-5-wang.zhan@smartx.com/
v2: https://lore.kernel.org/20260918084651.3022878-4-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-4-wang.zhan@smartx.com/
---
 net/core/dev.c | 53 +++++++++++++++++++++++++++++++++++++++++++++++++-
 1 file changed, 52 insertions(+), 1 deletion(-)

diff --git a/net/core/dev.c b/net/core/dev.c
index d6e39df55cd8cf..ef64834520949f 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3996,6 +3996,53 @@ __netif_skb_features(struct sk_buff *skb, bool check_gso_limits)
 }
 EXPORT_SYMBOL(__netif_skb_features);
 
+static unsigned int
+skb_gso_output_max_segs(struct sk_buff *skb, struct net_device *dev)
+{
+	unsigned int mss = skb_shinfo(skb)->gso_size;
+	unsigned int gso_max_size, hdr_len, max_segs;
+	netdev_features_t features;
+	struct tcphdr _tcph, *th;
+
+	if (!skb_is_gso(skb) || !skb_is_gso_tcp(skb) ||
+	    skb->encapsulation || skb_has_frag_list(skb))
+		return 0;
+
+	/*
+	 * The transport offset can be unset or stale, so hdr_len is only taken
+	 * from a header at the checksum start.
+	 */
+	if (!skb_transport_header_was_set(skb) ||
+	    skb_checksum_start(skb) != skb_transport_header(skb))
+		return 0;
+
+	th = skb_header_pointer(skb, skb_transport_offset(skb), sizeof(_tcph),
+				&_tcph);
+	if (!th || __tcp_hdrlen(th) < sizeof(*th))
+		return 0;
+
+	hdr_len = skb_transport_header(skb) - skb_mac_header(skb) +
+		  __tcp_hdrlen(th);
+
+	/* The output is still a GSO skb: recheck without the limit checks */
+	features = __netif_skb_features(skb, false) | NETIF_F_GSO_ROBUST;
+	if (!net_gso_ok(features, skb_shinfo(skb)->gso_type))
+		return 0;
+
+	gso_max_size = netif_get_gso_max_size(dev, skb);
+	gso_max_size = min(gso_max_size, GSO_LEGACY_MAX_SIZE);
+
+	/*
+	 * gso_within_dev_limits() accepts gso_segs == gso_max_segs but
+	 * rejects skb->len >= gso_max_size, so only the size limit needs the
+	 * - 1; the inner min() keeps that subtraction from wrapping.
+	 */
+	max_segs = (gso_max_size - min(gso_max_size, hdr_len + 1)) / mss;
+	max_segs = min(max_segs, READ_ONCE(dev->gso_max_segs));
+
+	return max_segs;
+}
+
 static int xmit_one(struct sk_buff *skb, struct net_device *dev,
 		    struct netdev_queue *txq, bool more)
 {
@@ -4159,9 +4206,13 @@ static struct sk_buff *validate_xmit_skb(struct sk_buff *skb, struct net_device
 		goto out_null;
 
 	if (netif_needs_gso(skb, features)) {
+		unsigned int max_segs = 0;
 		struct sk_buff *segs;
 
-		segs = skb_gso_segment(skb, features);
+		if (unlikely(!gso_within_dev_limits(skb, dev)))
+			max_segs = skb_gso_output_max_segs(skb, dev);
+
+		segs = __skb_gso_segment(skb, features, true, max_segs);
 		if (IS_ERR(segs)) {
 			goto out_kfree_skb;
 		} else if (segs) {
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH net-next v5 6/6] net: net_test: add tests for TCP re-segmentation
  2026-10-08 10:06 [PATCH net-next v5 0/6] net: re-segment oversized TCP GSO skbs Wang Zhan
                   ` (4 preceding siblings ...)
  2026-10-08 10:06 ` [PATCH net-next v5 5/6] net: core: re-segment oversized TCP GSO skbs Wang Zhan
@ 2026-10-08 10:06 ` Wang Zhan
  2026-10-08 18:38   ` Willem de Bruijn
  5 siblings, 1 reply; 9+ messages in thread
From: Wang Zhan @ 2026-10-08 10:06 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

The GSO engine can now be asked to group several MSS segments into one
output skb, which is what re-segmentation needs.  Add KUnit coverage for
the max_segs parameter.

The parameterized GSO test gains a max_segs input and four cases: a limit
of two MSS applied to a 3003 byte and to a 5003 byte payload, a tail which
holds one MSS and is therefore not a GSO skb, and a max_segs of one MSS
which must leave the output unchanged.  It drives skb_segment() directly,
because the synthetic protocol it uses has no gso_segment callback.

Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>

---
v5:
- new case: last output not a GSO skb
- free extra outputs instead of leaking them
- drop the TCP path tests, coverage in a later series
v4: https://lore.kernel.org/20260930111526.2183107-6-wang.zhan@smartx.com/
v3: https://lore.kernel.org/20260928044102.1004310-6-wang.zhan@smartx.com/
v2: https://lore.kernel.org/20260918084651.3022878-5-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-5-wang.zhan@smartx.com/
---
 net/core/net_test.c | 87 +++++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 87 insertions(+)

diff --git a/net/core/net_test.c b/net/core/net_test.c
index 9c3a590865d269..bc2ef7612e7dfe 100644
--- a/net/core/net_test.c
+++ b/net/core/net_test.c
@@ -5,6 +5,7 @@
 /* GSO */
 
 #include <linux/skbuff.h>
+#include <net/gso.h>
 
 static const char hdr[] = "abcdefgh";
 #define GSO_TEST_SIZE 1000
@@ -34,6 +35,10 @@ enum gso_test_nr {
 	GSO_TEST_FRAG_LIST_PURE,
 	GSO_TEST_FRAG_LIST_NON_UNIFORM,
 	GSO_TEST_GSO_BY_FRAGS,
+	GSO_TEST_RESEGMENT,
+	GSO_TEST_RESEGMENT_MULTI,
+	GSO_TEST_RESEGMENT_TAIL,
+	GSO_TEST_RESEGMENT_ONE_MSS,
 };
 
 struct gso_test_case {
@@ -46,10 +51,12 @@ struct gso_test_case {
 	const unsigned int *frags;
 	unsigned int nr_frag_skbs;
 	const unsigned int *frag_skbs;
+	unsigned int max_segs;
 
 	/* output as expected */
 	unsigned int nr_segs;
 	const unsigned int *segs;
+	const unsigned int *gso_segs;
 };
 
 static struct gso_test_case cases[] = {
@@ -135,6 +142,68 @@ static struct gso_test_case cases[] = {
 		.nr_segs = 4,
 		.segs = (const unsigned int[]) { 100, 200, 300, 400 },
 	},
+	{
+		.id = GSO_TEST_RESEGMENT,
+		.name = "resegment",
+		.linear_len = GSO_TEST_SIZE,
+		.nr_frags = 3,
+		.frags = (const unsigned int[]) {
+			GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
+		},
+		.max_segs = 2,
+		.nr_segs = 2,
+		.segs = (const unsigned int[]) {
+			2 * GSO_TEST_SIZE, GSO_TEST_SIZE + 3,
+		},
+		.gso_segs = (const unsigned int[]) { 2, 2 },
+	},
+	{
+		.id = GSO_TEST_RESEGMENT_MULTI,
+		.name = "resegment_multi",
+		.linear_len = 2 * GSO_TEST_SIZE,
+		.nr_frags = 4,
+		.frags = (const unsigned int[]) {
+			GSO_TEST_SIZE, GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
+		},
+		.max_segs = 2,
+		.nr_segs = 3,
+		.segs = (const unsigned int[]) {
+			2 * GSO_TEST_SIZE, 2 * GSO_TEST_SIZE, GSO_TEST_SIZE + 3,
+		},
+		.gso_segs = (const unsigned int[]) { 2, 2, 2 },
+	},
+	{
+		/* The last output is not a GSO skb. */
+		.id = GSO_TEST_RESEGMENT_TAIL,
+		.name = "resegment_tail",
+		.linear_len = GSO_TEST_SIZE,
+		.nr_frags = 2,
+		.frags = (const unsigned int[]) {
+			GSO_TEST_SIZE, 400,
+		},
+		.max_segs = 2,
+		.nr_segs = 2,
+		.segs = (const unsigned int[]) {
+			2 * GSO_TEST_SIZE, 400,
+		},
+		.gso_segs = (const unsigned int[]) { 2, 0 },
+	},
+	{
+		/* max_segs of 1 is the same as no limit. */
+		.id = GSO_TEST_RESEGMENT_ONE_MSS,
+		.name = "resegment_one_mss",
+		.linear_len = GSO_TEST_SIZE,
+		.nr_frags = 3,
+		.frags = (const unsigned int[]) {
+			GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
+		},
+		.max_segs = 1,
+		.nr_segs = 4,
+		.segs = (const unsigned int[]) {
+			GSO_TEST_SIZE, GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
+		},
+		.gso_segs = (const unsigned int[]) { 0, 0, 0, 0 },
+	},
 };
 
 static void gso_test_case_to_desc(struct gso_test_case *t, char *desc)
@@ -226,6 +295,7 @@ static void gso_test_func(struct kunit *test)
 	if (tcase->id == GSO_TEST_FRAG_LIST_NON_UNIFORM)
 		features &= ~NETIF_F_SG;
 
+	SKB_GSO_CB(skb)->max_segs = tcase->max_segs;
 	segs = skb_segment(skb, features);
 	if (IS_ERR(segs)) {
 		KUNIT_FAIL(test, "segs error %pe", segs);
@@ -239,6 +309,9 @@ static void gso_test_func(struct kunit *test)
 	for (cur = segs, i = 0; cur; cur = next, i++) {
 		next = cur->next;
 
+		if (i >= tcase->nr_segs)
+			goto consume;
+
 		KUNIT_ASSERT_EQ(test, cur->len, sizeof(hdr) + tcase->segs[i]);
 
 		/* segs have skb->data pointing to the mac header */
@@ -247,11 +320,25 @@ static void gso_test_func(struct kunit *test)
 
 		/* header was copied to all segs */
 		KUNIT_ASSERT_EQ(test, memcmp(skb_mac_header(cur), hdr, sizeof(hdr)), 0);
+		if (tcase->gso_segs) {
+			unsigned int gso_segs = tcase->gso_segs[i];
+
+			KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_segs,
+					gso_segs);
+			if (!gso_segs) {
+				KUNIT_EXPECT_FALSE(test, skb_is_gso(cur));
+			} else {
+				KUNIT_EXPECT_TRUE(test, skb_is_gso(cur));
+				KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_size,
+						GSO_TEST_SIZE);
+			}
+		}
 
 		/* last seg can be found through segs->prev pointer */
 		if (!next)
 			KUNIT_ASSERT_PTR_EQ(test, cur, last);
 
+consume:
 		consume_skb(cur);
 	}
 
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* Re: [PATCH net-next v5 4/6] net: gso: clear the GSO state of the last output segment
  2026-10-08 10:06 ` [PATCH net-next v5 4/6] net: gso: clear the GSO state of the last output segment Wang Zhan
@ 2026-10-08 18:37   ` Willem de Bruijn
  0 siblings, 0 replies; 9+ messages in thread
From: Willem de Bruijn @ 2026-10-08 18:37 UTC (permalink / raw)
  To: Wang Zhan, netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan, Willem de Bruijn

Wang Zhan wrote:
> skb_segment() clears gso_size on the last output of a group which holds one
> MSS or less, but leaves gso_segs and gso_type as the grouping set them, so
> the tail is still counted as the whole group.  Readers which take gso_segs
> without testing gso_size first, such as the ECT statistics in ip_rcv_core()
> and tp->rcv_ooopack in tcp_rcv_established(), see that stale count, and a
> re-segmented output reaches them when a veth or bridge port delivers it back
> into the receive path.
> 
> Clear gso_size, gso_segs and gso_type together with skb_gso_reset().
> 
> Suggested-by: Willem de Bruijn <willemb@google.com>
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>

Reviewed-by: Willem de Bruijn <willemb@google.com>

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH net-next v5 6/6] net: net_test: add tests for TCP re-segmentation
  2026-10-08 10:06 ` [PATCH net-next v5 6/6] net: net_test: add tests for TCP re-segmentation Wang Zhan
@ 2026-10-08 18:38   ` Willem de Bruijn
  0 siblings, 0 replies; 9+ messages in thread
From: Willem de Bruijn @ 2026-10-08 18:38 UTC (permalink / raw)
  To: Wang Zhan, netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

Wang Zhan wrote:
> The GSO engine can now be asked to group several MSS segments into one
> output skb, which is what re-segmentation needs.  Add KUnit coverage for
> the max_segs parameter.
> 
> The parameterized GSO test gains a max_segs input and four cases: a limit
> of two MSS applied to a 3003 byte and to a 5003 byte payload, a tail which
> holds one MSS and is therefore not a GSO skb, and a max_segs of one MSS
> which must leave the output unchanged.  It drives skb_segment() directly,
> because the synthetic protocol it uses has no gso_segment callback.
> 
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>

Reviewed-by: Willem de Bruijn <willemb@google.com>

^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-10-08 18:38 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-08 10:06 [PATCH net-next v5 0/6] net: re-segment oversized TCP GSO skbs Wang Zhan
2026-10-08 10:06 ` [PATCH net-next v5 1/6] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
2026-10-08 10:06 ` [PATCH net-next v5 2/6] net: core: factor out the GSO device limit check Wang Zhan
2026-10-08 10:06 ` [PATCH net-next v5 3/6] net: gso: support re-segmentation of TCP GSO skbs Wang Zhan
2026-10-08 10:06 ` [PATCH net-next v5 4/6] net: gso: clear the GSO state of the last output segment Wang Zhan
2026-10-08 18:37   ` Willem de Bruijn
2026-10-08 10:06 ` [PATCH net-next v5 5/6] net: core: re-segment oversized TCP GSO skbs Wang Zhan
2026-10-08 10:06 ` [PATCH net-next v5 6/6] net: net_test: add tests for TCP re-segmentation Wang Zhan
2026-10-08 18:38   ` Willem de Bruijn

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox