Netdev List
 help / color / mirror / Atom feed
* [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs
@ 2026-09-28  4:40 Wang Zhan
  2026-09-28  4:40 ` [PATCH net-next v3 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
                   ` (5 more replies)
  0 siblings, 6 replies; 27+ messages in thread
From: Wang Zhan @ 2026-09-28  4:40 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

BIG TCP is negotiated per netdevice, so BIG TCP and non-BIG TCP ports can
coexist in one path.  When an skb exceeds the GSO limits of the port it is
sent to, it loses its GSO feature mask and is segmented into individual MSS
sized packets, which that port then sends without TSO.

This series cuts such a TCP GSO skb into GSO skbs which fit the device
limits instead, so the rest of the path keeps using TSO.  On a
veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP enabled on the
veth endpoints and left off in the guest, a single iperf3 TCP flow, six
alternating runs per state (-t 15 -O 5, fixed CPU affinity and port tuple):

  protocol  no BIG TCP   mixed, no reseg  mixed, resegmented
  TCP/IPv4  51.550 Gbps  15.850 Gbps      52.617 Gbps
  TCP/IPv6  52.050 Gbps  15.783 Gbps      51.933 Gbps

The middle column is the same tree with the bounded path disabled.  A BIG
TCP hop which feeds a 64 KiB hop loses 69% of the throughput of a path
which never enables BIG TCP; bounded resegmentation recovers it.

1/5 stands on its own as a fix: the size limit is picked from
skb->protocol, which is the VLAN ethertype for a frame which already
carries its tag in the packet, so an IPv6 frame was measured against the
IPv4 limit.  2/5 is the preparation which lets the limit tests be skipped
for one caller, and carries no functional change.

The new path is taken only when the skb is an unencapsulated TCP GSO skb
without a frag_list, it exceeds gso_max_size or gso_max_segs, and the
device offloads that GSO type.  Everything else keeps today's
segmentation.

The output obeys the GSO feature and limit contract the device already
advertises, so this needs no new UAPI, no device state and no driver
change, and it applies automatically.  The bound only says how the output
is grouped; an over-limit skb pays one extra ndo_features_check() in
exchange for staying a GSO skb.

Alternatives considered:

  - The caller could set skb_shinfo(skb)->gso_size to ~64K and adjust the
    gso bits in the shared info afterwards, which would need
    skb_unclone(), a repeat of the grouping logic skb_segment() already
    has, and a recomputed IPv4 ID for the DF=0 case.

Patch layout:

  [1/5] the GSO size limit follows the packet's L3 protocol
  [2/5] factor the device limit check out of gso_features_check()
  [3/5] let the GSO engine bound the MSS segments per output skb
  [4/5] apply that bound to oversized TCP GSO skbs in the TX path
  [5/5] KUnit coverage for the bound, the device limits and the TCP path

---
v3:
- patch 1 is new: the size limit follows the packet's L3 protocol
- patch 2: move the check_gso_limits flag and the wrapper split in from patch 4
- patch 3: drop the tcp_gso_segment() exception, the caller keeps the features
- patch 3: cap the bound with GSO_MAX_SEGS and drop the output reset
- patch 3: document the bound as TCP only and note the frag_list gate
- patch 4: enter from the limit predicate gso_features_check() uses
- patch 4: keep the features, so the bound shapes the output only
- patch 4: drop the SG and checksum tests, fold the MSS minimum
- patch 4: cap each output at GSO_LEGACY_MAX_SIZE, not the BIG TCP size
- patch 5: cover the IPv4 and the IPv6 limit, with and without the tag
- patch 5: skip the TCP cases without CONFIG_INET, reserve headroom
- patch 5: free on the failure paths, check the ungrouped single MSS
v2: https://lore.kernel.org/20260918084651.3022878-1-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-1-wang.zhan@smartx.com/

Wang Zhan (5):
  net: core: use the packet's L3 protocol for the GSO size limit
  net: core: factor out the GSO device limit check
  net: gso: support bounded TCP segmentation
  net: core: resegment oversized TCP GSO skbs
  net: net_test: add tests for bounded GSO segmentation

 drivers/net/tap.c          |   3 +-
 include/linux/netdevice.h  |   4 +-
 include/net/gso.h          |   6 +-
 include/net/udp.h          |   2 +-
 net/core/dev.c             |  91 ++++++++--
 net/core/gso.c             |   6 +-
 net/core/net_test.c        | 347 +++++++++++++++++++++++++++++++++++++
 net/core/skbuff.c          |  14 +-
 net/openvswitch/datapath.c |   2 +-
 9 files changed, 455 insertions(+), 20 deletions(-)


base-commit: 014d795c73837ea2339a4ea8e8f82c6e959b845d
-- 
2.47.3

^ permalink raw reply	[flat|nested] 27+ messages in thread

* [PATCH net-next v3 1/5] net: core: use the packet's L3 protocol for the GSO size limit
  2026-09-28  4:40 [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs Wang Zhan
@ 2026-09-28  4:40 ` Wang Zhan
  2026-09-28 23:36   ` Willem de Bruijn
  2026-09-28  4:40 ` [PATCH net-next v3 2/5] net: core: factor out the GSO device limit check Wang Zhan
                   ` (4 subsequent siblings)
  5 siblings, 1 reply; 27+ messages in thread
From: Wang Zhan @ 2026-09-28  4:40 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

gso_features_check() compares the frame length against
netif_get_gso_max_size(), which picks the IPv4 or the IPv6 limit from
skb->protocol.  A tag which is already inside the frame replaces that field
with the VLAN ethertype, as skb_vlan_push() does, and an IPv6 packet is
then measured against the IPv4 limit and segmented although the device
could send it as one TSO frame.

Let the limit lookup take the protocol as an argument, and pass the L3
protocol, so where the tag sits does not decide which limit applies.  The
helper cannot look behind the tag itself: it lives in netdevice.h, which
cannot include if_vlan.h because that header includes netdevice.h.

Fixes: e609c959a9396 ("net: Fix gso_features_check to check for both dev->gso_{ipv4_,}max_size")
Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>

---
v3:
- new patch: the L3 protocol fix split out of the resegmentation patch
v2: https://lore.kernel.org/20260918084651.3022878-4-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-4-wang.zhan@smartx.com/
---
 include/linux/netdevice.h | 4 ++--
 net/core/dev.c            | 3 ++-
 2 files changed, 4 insertions(+), 3 deletions(-)

diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index d037faff7c44b..1c28861fcbe2c 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -5562,10 +5562,10 @@ netif_get_gro_max_size(const struct net_device *dev, const struct sk_buff *skb)
 }
 
 static inline unsigned int
-netif_get_gso_max_size(const struct net_device *dev, const struct sk_buff *skb)
+netif_get_gso_max_size(const struct net_device *dev, __be16 protocol)
 {
 	/* pairs with WRITE_ONCE() in netif_set_gso(_ipv4)_max_size() */
-	return skb->protocol == htons(ETH_P_IPV6) ?
+	return protocol == htons(ETH_P_IPV6) ?
 	       READ_ONCE(dev->gso_max_size) :
 	       READ_ONCE(dev->gso_ipv4_max_size);
 }
diff --git a/net/core/dev.c b/net/core/dev.c
index f660fccfc0dbc..ffa9b0c27788c 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3843,7 +3843,8 @@ static netdev_features_t gso_features_check(const struct sk_buff *skb,
 	if (gso_segs > READ_ONCE(dev->gso_max_segs))
 		return features & ~NETIF_F_GSO_MASK;
 
-	if (unlikely(skb->len >= netif_get_gso_max_size(dev, skb)))
+	if (unlikely(skb->len >=
+		     netif_get_gso_max_size(dev, vlan_get_protocol(skb))))
 		return features & ~NETIF_F_GSO_MASK;
 
 	if (!skb_shinfo(skb)->gso_type) {
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH net-next v3 2/5] net: core: factor out the GSO device limit check
  2026-09-28  4:40 [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs Wang Zhan
  2026-09-28  4:40 ` [PATCH net-next v3 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
@ 2026-09-28  4:40 ` Wang Zhan
  2026-09-28 23:37   ` Willem de Bruijn
  2026-09-29  7:47   ` Paolo Abeni
  2026-09-28  4:41 ` [PATCH net-next v3 3/5] net: gso: support bounded TCP segmentation Wang Zhan
                   ` (3 subsequent siblings)
  5 siblings, 2 replies; 27+ messages in thread
From: Wang Zhan @ 2026-09-28  4:40 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

gso_features_check() decides whether an egress device can offload a GSO
skb as a single TSO frame by comparing the segment count and the frame
length against the device limits.  The resegmentation path added by a
later patch caps the MSS segments per output skb and needs the same
features computed without those limits, so put the two tests in a helper
and gate them with one flag.

No functional changes.

Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>

---
v3:
- rename the helper to gso_within_dev_limits
- move the check_gso_limits flag and the wrapper split here from patch 4
v2: https://lore.kernel.org/20260918084651.3022878-2-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-2-wang.zhan@smartx.com/
---
 net/core/dev.c | 29 +++++++++++++++++++----------
 1 file changed, 19 insertions(+), 10 deletions(-)

diff --git a/net/core/dev.c b/net/core/dev.c
index ffa9b0c27788c..d66b667071837 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3834,17 +3834,19 @@ static bool skb_gso_has_extension_hdr(const struct sk_buff *skb)
 			 skb_inner_network_header_len(skb) != sizeof(struct ipv6hdr)));
 }
 
+static bool
+gso_within_dev_limits(const struct sk_buff *skb, const struct net_device *dev)
+{
+	return skb_shinfo(skb)->gso_segs <= READ_ONCE(dev->gso_max_segs) &&
+	       skb->len < netif_get_gso_max_size(dev, vlan_get_protocol(skb));
+}
+
 static netdev_features_t gso_features_check(const struct sk_buff *skb,
 					    struct net_device *dev,
-					    netdev_features_t features)
+					    netdev_features_t features,
+					    bool check_limits)
 {
-	u16 gso_segs = skb_shinfo(skb)->gso_segs;
-
-	if (gso_segs > READ_ONCE(dev->gso_max_segs))
-		return features & ~NETIF_F_GSO_MASK;
-
-	if (unlikely(skb->len >=
-		     netif_get_gso_max_size(dev, vlan_get_protocol(skb))))
+	if (check_limits && !gso_within_dev_limits(skb, dev))
 		return features & ~NETIF_F_GSO_MASK;
 
 	if (!skb_shinfo(skb)->gso_type) {
@@ -3893,13 +3895,15 @@ static netdev_features_t gso_features_check(const struct sk_buff *skb,
 	return features;
 }
 
-netdev_features_t netif_skb_features(struct sk_buff *skb)
+static netdev_features_t
+__netif_skb_features(struct sk_buff *skb, bool check_gso_limits)
 {
 	struct net_device *dev = skb->dev;
 	netdev_features_t features = dev->features;
 
 	if (skb_is_gso(skb))
-		features = gso_features_check(skb, dev, features);
+		features = gso_features_check(skb, dev, features,
+					      check_gso_limits);
 
 	/* If encapsulation offload request, verify we are testing
 	 * hardware encapsulation features instead of standard
@@ -3922,6 +3926,11 @@ netdev_features_t netif_skb_features(struct sk_buff *skb)
 
 	return harmonize_features(skb, features);
 }
+
+netdev_features_t netif_skb_features(struct sk_buff *skb)
+{
+	return __netif_skb_features(skb, true);
+}
 EXPORT_SYMBOL(netif_skb_features);
 
 static int xmit_one(struct sk_buff *skb, struct net_device *dev,
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH net-next v3 3/5] net: gso: support bounded TCP segmentation
  2026-09-28  4:40 [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs Wang Zhan
  2026-09-28  4:40 ` [PATCH net-next v3 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
  2026-09-28  4:40 ` [PATCH net-next v3 2/5] net: core: factor out the GSO device limit check Wang Zhan
@ 2026-09-28  4:41 ` Wang Zhan
  2026-09-28 23:39   ` Willem de Bruijn
  2026-09-30  4:41   ` netdev-bot+sashiko
  2026-09-28  4:41 ` [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs Wang Zhan
                   ` (2 subsequent siblings)
  5 siblings, 2 replies; 27+ messages in thread
From: Wang Zhan @ 2026-09-28  4:41 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

The bounded resegmentation added by the next patch splits an oversized TCP
GSO skb into several GSO skbs which fit the device limits.  That needs the
GSO engine to group several MSS segments into one output skb, so let
callers bound the number of MSS segments each output skb carries and pass
the bound through the existing __skb_gso_segment() entry point.  Ordinary
callers use zero for no limit.

Without a bound, skb_segment() groups several MSS into one output skb only
for a device which advertises NETIF_F_GSO_PARTIAL or for a skb with a
frag_list which can be split into uniform pieces, and falls back to one
segment per skb otherwise.  A caller which passes a bound asks for the
grouping regardless, so the block which makes that decision is skipped
when max_segs is set.  Callers which pass no bound keep it, and the bounded
path only runs for skbs which carry no frag_list.

The output stays a GSO skb: gso_size is the original MSS and gso_segs is
the number of MSS it holds, so a downstream device can still perform
ordinary TSO.  Store the bound in the existing skb_gso_cb scratch context,
alongside the call-local data_offset and mac_offset fields, so that the
segmentation methods keep their signature.  A zero max_segs value means
that no bound is active; it is not a persistent skb flag.

Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>

---
v3:
- drop the tcp_gso_segment() exception: the caller keeps the features
- cap the bound with GSO_MAX_SEGS instead of U16_MAX
- drop the reset of the field in the output skbs: every caller initializes it
- note in the kernel-doc that a bound is for TCP GSO skbs only
- note at the frag_list gate that a bound must not go with a frag_list skb
v2: https://lore.kernel.org/20260918084651.3022878-3-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-3-wang.zhan@smartx.com/
---
 drivers/net/tap.c          |  3 ++-
 include/net/gso.h          |  6 ++++--
 include/net/udp.h          |  2 +-
 net/core/gso.c             |  6 +++++-
 net/core/skbuff.c          | 14 ++++++++++++--
 net/openvswitch/datapath.c |  2 +-
 6 files changed, 25 insertions(+), 8 deletions(-)

diff --git a/drivers/net/tap.c b/drivers/net/tap.c
index ff67d99deb39e..bc111495ebbce 100644
--- a/drivers/net/tap.c
+++ b/drivers/net/tap.c
@@ -278,9 +278,10 @@ rx_handler_result_t tap_handle_frame(struct sk_buff **pskb)
 	if (q->flags & IFF_VNET_HDR)
 		features |= tap->tap_features;
 	if (netif_needs_gso(skb, features)) {
-		struct sk_buff *segs = __skb_gso_segment(skb, features, false);
+		struct sk_buff *segs;
 		struct sk_buff *next;
 
+		segs = __skb_gso_segment(skb, features, false, 0);
 		if (IS_ERR(segs)) {
 			drop_reason = SKB_DROP_REASON_SKB_GSO_SEG;
 			goto drop;
diff --git a/include/net/gso.h b/include/net/gso.h
index 29975440cad51..fccb37889965f 100644
--- a/include/net/gso.h
+++ b/include/net/gso.h
@@ -19,6 +19,7 @@ struct skb_gso_cb {
 	int	encap_level;
 	__wsum	csum;
 	__u16	csum_start;
+	__u16	max_segs;	/* Max MSS segs per output skb, 0 = no limit */
 };
 #define SKB_GSO_CB_OFFSET	32
 #define SKB_GSO_CB(skb) ((struct skb_gso_cb *)((skb)->cb + SKB_GSO_CB_OFFSET))
@@ -75,12 +76,13 @@ static inline __sum16 gso_make_checksum(struct sk_buff *skb, __wsum res)
 }
 
 struct sk_buff *__skb_gso_segment(struct sk_buff *skb,
-				  netdev_features_t features, bool tx_path);
+				  netdev_features_t features, bool tx_path,
+				  unsigned int max_segs);
 
 static inline struct sk_buff *skb_gso_segment(struct sk_buff *skb,
 					      netdev_features_t features)
 {
-	return __skb_gso_segment(skb, features, true);
+	return __skb_gso_segment(skb, features, true, 0);
 }
 
 struct sk_buff *skb_eth_gso_segment(struct sk_buff *skb,
diff --git a/include/net/udp.h b/include/net/udp.h
index 1fee17274745f..5bc25dcf25fba 100644
--- a/include/net/udp.h
+++ b/include/net/udp.h
@@ -613,7 +613,7 @@ static inline struct sk_buff *udp_rcv_segment(struct sock *sk,
 	/* the GSO CB lays after the UDP one, no need to save and restore any
 	 * CB fragment
 	 */
-	segs = __skb_gso_segment(skb, features, false);
+	segs = __skb_gso_segment(skb, features, false, 0);
 	if (IS_ERR_OR_NULL(segs)) {
 		drop_count = skb_shinfo(skb)->gso_segs;
 		goto drop;
diff --git a/net/core/gso.c b/net/core/gso.c
index bcd156372f4df..42fbce17b508c 100644
--- a/net/core/gso.c
+++ b/net/core/gso.c
@@ -77,6 +77,8 @@ static bool skb_needs_check(const struct sk_buff *skb, bool tx_path)
  *	@skb: buffer to segment
  *	@features: features for the output path (see dev->features)
  *	@tx_path: whether it is called in TX path
+ *	@max_segs: maximum MSS segments per output GSO skb, 0 means no limit;
+ *		   must only be set for TCP GSO skbs
  *
  *	This function segments the given skb and returns a list of segments.
  *
@@ -86,7 +88,8 @@ static bool skb_needs_check(const struct sk_buff *skb, bool tx_path)
  *	Segmentation preserves SKB_GSO_CB_OFFSET bytes of previous skb cb.
  */
 struct sk_buff *__skb_gso_segment(struct sk_buff *skb,
-				  netdev_features_t features, bool tx_path)
+				  netdev_features_t features, bool tx_path,
+				  unsigned int max_segs)
 {
 	struct sk_buff *segs;
 
@@ -117,6 +120,7 @@ struct sk_buff *__skb_gso_segment(struct sk_buff *skb,
 
 	SKB_GSO_CB(skb)->mac_offset = skb_headroom(skb);
 	SKB_GSO_CB(skb)->encap_level = 0;
+	SKB_GSO_CB(skb)->max_segs = min(max_segs, GSO_MAX_SEGS);
 
 	skb_reset_mac_header(skb);
 	skb_reset_mac_len(skb);
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 8912a66cd9097..b16a6843f3192 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -4793,6 +4793,7 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
 	struct sk_buff *segs = NULL;
 	struct sk_buff *tail = NULL;
 	struct sk_buff *list_skb = skb_shinfo(head_skb)->frag_list;
+	unsigned int max_segs = SKB_GSO_CB(head_skb)->max_segs;
 	unsigned int mss = skb_shinfo(head_skb)->gso_size;
 	bool gso_by_frags = mss == GSO_BY_FRAGS;
 	unsigned int doffset = head_skb->data - skb_mac_header(head_skb);
@@ -4839,7 +4840,13 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
 	csum = !!can_checksum_protocol(features, proto);
 
 	if (sg && csum && !gso_by_frags)  {
-		if (!(features & NETIF_F_GSO_PARTIAL)) {
+		/*
+		 * A call with max_segs set groups the MSS segments below,
+		 * and the frag_list split in this block does not carry the
+		 * bound, so a bound must not be passed for a skb with a
+		 * frag_list.
+		 */
+		if (!max_segs && !(features & NETIF_F_GSO_PARTIAL)) {
 			struct sk_buff *iter;
 			unsigned int frag_len;
 
@@ -4874,7 +4881,10 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
 		 * now.
 		 */
 		DEBUG_NET_WARN_ON_ONCE(len / mss > GSO_MAX_SEGS);
-		partial_segs = min(len / mss, GSO_MAX_SEGS);
+		if (max_segs)
+			partial_segs = min(len / mss, max_segs);
+		else
+			partial_segs = min(len / mss, GSO_MAX_SEGS);
 		if (partial_segs > 1)
 			mss *= partial_segs;
 		else
diff --git a/net/openvswitch/datapath.c b/net/openvswitch/datapath.c
index 2187034143255..e793aead68372 100644
--- a/net/openvswitch/datapath.c
+++ b/net/openvswitch/datapath.c
@@ -375,7 +375,7 @@ static int queue_gso_packets(struct datapath *dp, struct sk_buff *skb,
 	int err;
 
 	BUILD_BUG_ON(sizeof(*OVS_CB(skb)) > SKB_GSO_CB_OFFSET);
-	segs = __skb_gso_segment(skb, NETIF_F_SG, false);
+	segs = __skb_gso_segment(skb, NETIF_F_SG, false, 0);
 	if (IS_ERR(segs))
 		return PTR_ERR(segs);
 	if (segs == NULL)
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs
  2026-09-28  4:40 [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs Wang Zhan
                   ` (2 preceding siblings ...)
  2026-09-28  4:41 ` [PATCH net-next v3 3/5] net: gso: support bounded TCP segmentation Wang Zhan
@ 2026-09-28  4:41 ` Wang Zhan
  2026-09-28 23:47   ` Willem de Bruijn
  2026-09-30  4:41   ` netdev-bot+sashiko
  2026-09-28  4:41 ` [PATCH net-next v3 5/5] net: net_test: add tests for bounded GSO segmentation Wang Zhan
  2026-09-28  4:45 ` [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs netdev-bot+sinfo
  5 siblings, 2 replies; 27+ messages in thread
From: Wang Zhan @ 2026-09-28  4:41 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

A GSO skb which exceeds an egress device limit loses its GSO feature mask
and is segmented into individual packets.  This is unnecessarily expensive
when the device can still offload smaller TCP GSO skbs, which is easy to
hit once one hop of a BIG TCP path raises gso_max_size and the next one
does not.

For an unencapsulated TCP GSO skb which exceeds gso_max_size or
gso_max_segs, work out how many MSS segments each output skb may carry and
resegment the skb with that max_segs bound instead.  Encapsulated and
frag-list skbs, GSO types the device cannot offload and bounds which leave
room for a single MSS keep today's segmentation.  The output obeys the GSO
feature and limit contract the device already advertises, so it applies
automatically, without extra device state or a userspace control.

The output is a plain GSO skb, so its whole length lands in the 16-bit L3
length field which inet_gso_segment() and ipv6_gso_segment() write.  An
egress limit above 64 KiB would give outputs whose length truncates, so
the size limit is also capped at what that field can express.

The helper runs on the skb which is handed to the driver, after
validate_xmit_vlan() and sk_validate_xmit_skb(), and only from the
netif_needs_gso() branch, so an skb which is not segmented pays nothing.
An over-limit skb pays one device limit test and one ndo_features_check()
for the bound, in exchange for keeping the output a GSO skb.

Measured on a veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP
enabled on the veth endpoints and left off in the guest, so the skbs which
the veth hop accepts have to be segmented before the TAP device.  A single
iperf3 TCP flow, six alternating runs per state (`-t 15 -O 5`, fixed CPU
affinity and port tuple).  The middle column is the same tree with the
resegmentation disabled:

  protocol  no BIG TCP   mixed, no reseg  mixed, resegmented
  TCP/IPv4  51.550 Gbps  15.850 Gbps      52.617 Gbps
  TCP/IPv6  52.050 Gbps  15.783 Gbps      51.933 Gbps

Coefficient of variation for the two mixed columns was 0.48% and 0.82% for
IPv4 and 0.44% and 0.44% for IPv6.  A BIG TCP hop which feeds a 64 KiB hop
loses 69% of the throughput of a path which never enables BIG TCP at all;
bounded resegmentation recovers it, 3.3x over the existing segmentation
path and within noise of the no BIG TCP baseline.

Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>

---
v3:
- the L3 protocol change moved to patch 1
- fold skb_can_gso_resegment() into the helper, now skb_gso_output_max_segs()
- pass the caller's features on unchanged, the bound only shapes the output
- drop the scatter-gather and checksum tests skb_segment() applies itself
- saturate the size limit with min(), drop the sub-MSS guard
- cap the output size at GSO_LEGACY_MAX_SIZE, these outputs are not BIG TCP
v2: https://lore.kernel.org/20260918084651.3022878-4-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-4-wang.zhan@smartx.com/
---
 net/core/dev.c | 63 +++++++++++++++++++++++++++++++++++++++++++++++++-
 1 file changed, 62 insertions(+), 1 deletion(-)

diff --git a/net/core/dev.c b/net/core/dev.c
index d66b667071837..728260772f349 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3933,6 +3933,63 @@ netdev_features_t netif_skb_features(struct sk_buff *skb)
 }
 EXPORT_SYMBOL(netif_skb_features);
 
+static unsigned int
+skb_gso_output_max_segs(struct sk_buff *skb, struct net_device *dev)
+{
+	unsigned int mss = skb_shinfo(skb)->gso_size;
+	unsigned int gso_max_size, hdr_len, max_segs;
+	netdev_features_t features;
+	struct tcphdr _tcph, *th;
+
+	/*
+	 * The TCP frag-list path segments through skb_segment_list(), which
+	 * does not carry max_segs, so bounded calls skip those skbs.
+	 */
+	if (!skb_is_gso(skb) || !skb_is_gso_tcp(skb) ||
+	    skb->encapsulation || skb_has_frag_list(skb) ||
+	    !skb_mac_header_was_set(skb) ||
+	    !skb_transport_header_was_set(skb))
+		return 0;
+
+	th = skb_header_pointer(skb, skb_transport_offset(skb), sizeof(_tcph),
+				&_tcph);
+	if (!th || th->doff < sizeof(*th) / 4)
+		return 0;
+
+	hdr_len = skb_transport_header(skb) - skb_mac_header(skb) +
+		  th->doff * 4;
+
+	/*
+	 * The output stays a GSO skb, so the device has to offload the GSO
+	 * type.  The caller's features cannot tell that: this skb is over
+	 * the device limits, so gso_features_check() has cleared their GSO
+	 * bits.  Compute the features again without the limit checks.
+	 */
+	features = __netif_skb_features(skb, false);
+	if (!net_gso_ok(features | NETIF_F_GSO_ROBUST,
+			skb_shinfo(skb)->gso_type))
+		return 0;
+
+	/*
+	 * The output is a plain GSO skb, so its whole length lands in the
+	 * 16-bit L3 length field: only a BIG TCP skb may exceed the legacy
+	 * GSO size, and this path does not emit one.
+	 */
+	gso_max_size = netif_get_gso_max_size(dev, vlan_get_protocol(skb));
+	gso_max_size = min(gso_max_size, GSO_LEGACY_MAX_SIZE);
+
+	/*
+	 * gso_within_dev_limits() accepts gso_segs == gso_max_segs but
+	 * rejects skb->len >= gso_max_size, so only the size bound needs - 1.
+	 * The min() keeps that subtraction from wrapping when the device
+	 * limit is smaller than the headers.
+	 */
+	max_segs = (gso_max_size - min(gso_max_size, hdr_len + 1)) / mss;
+	max_segs = min(max_segs, READ_ONCE(dev->gso_max_segs));
+
+	return max_segs;
+}
+
 static int xmit_one(struct sk_buff *skb, struct net_device *dev,
 		    struct netdev_queue *txq, bool more)
 {
@@ -4097,9 +4154,13 @@ static struct sk_buff *validate_xmit_skb(struct sk_buff *skb, struct net_device
 		goto out_null;
 
 	if (netif_needs_gso(skb, features)) {
+		unsigned int max_segs = 0;
 		struct sk_buff *segs;
 
-		segs = skb_gso_segment(skb, features);
+		if (unlikely(!gso_within_dev_limits(skb, dev)))
+			max_segs = skb_gso_output_max_segs(skb, dev);
+
+		segs = __skb_gso_segment(skb, features, true, max_segs);
 		if (IS_ERR(segs)) {
 			goto out_kfree_skb;
 		} else if (segs) {
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH net-next v3 5/5] net: net_test: add tests for bounded GSO segmentation
  2026-09-28  4:40 [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs Wang Zhan
                   ` (3 preceding siblings ...)
  2026-09-28  4:41 ` [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs Wang Zhan
@ 2026-09-28  4:41 ` Wang Zhan
  2026-09-28 23:59   ` Willem de Bruijn
  2026-09-30  4:41   ` netdev-bot+sashiko
  2026-09-28  4:45 ` [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs netdev-bot+sinfo
  5 siblings, 2 replies; 27+ messages in thread
From: Wang Zhan @ 2026-09-28  4:41 UTC (permalink / raw)
  To: netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

The GSO engine can now be asked to bound the number of MSS segments which
go into each output skb.  Add KUnit coverage for it.

The parameterized GSO test gains a max_segs input and three cases: two
bounds which cut the input into two and three output skbs, and a bound of
one MSS, which must leave the ungrouped output of the unbounded path
alone.  It drives skb_segment() directly, because the synthetic protocol
it uses has no gso_segment callback, and stores the bound in the GSO
control block itself.

The TCP test drives __skb_gso_segment() with a bound of two MSS and checks
that every output skb stays GSO, keeps its gso_size, and stays within the
bound.  The length test runs a 200 KiB TCP skb through
validate_xmit_skb_list(), the caller which sets the bound, and checks that
the length declared by every output matches the L3 length of that output.
The limit test checks that the GSO size limit which netif_skb_features()
applies follows the packet's L3 protocol, for an IPv4 and an IPv6 skb,
also with the tag inside the frame.

Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>

---
v3:
- add cases for the length a bounded output declares and for a bound of one
  MSS, which must leave the output ungrouped
- cover the IPv4 half of the limit test, checking the protocol's own bit
- skip the TCP cases without CONFIG_INET and bound the expected index in
  the parameterized loop
- free the skb and the device on the failure paths
- reserve headroom so the VLAN step needs no atomic allocation
- drop NETIF_F_TSO: the transmit path passes the features without it
v2: https://lore.kernel.org/20260918084651.3022878-5-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-5-wang.zhan@smartx.com/
---
 net/core/net_test.c | 347 ++++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 347 insertions(+)

diff --git a/net/core/net_test.c b/net/core/net_test.c
index 9c3a590865d26..a4f61398a90ab 100644
--- a/net/core/net_test.c
+++ b/net/core/net_test.c
@@ -4,7 +4,14 @@
 
 /* GSO */
 
+#include <linux/if_ether.h>
+#include <linux/if_vlan.h>
+#include <linux/ip.h>
+#include <linux/ipv6.h>
+#include <linux/netdevice.h>
 #include <linux/skbuff.h>
+#include <linux/tcp.h>
+#include <net/gso.h>
 
 static const char hdr[] = "abcdefgh";
 #define GSO_TEST_SIZE 1000
@@ -34,6 +41,9 @@ enum gso_test_nr {
 	GSO_TEST_FRAG_LIST_PURE,
 	GSO_TEST_FRAG_LIST_NON_UNIFORM,
 	GSO_TEST_GSO_BY_FRAGS,
+	GSO_TEST_BOUNDED,
+	GSO_TEST_BOUNDED_MULTI,
+	GSO_TEST_BOUNDED_ONE_MSS,
 };
 
 struct gso_test_case {
@@ -46,10 +56,12 @@ struct gso_test_case {
 	const unsigned int *frags;
 	unsigned int nr_frag_skbs;
 	const unsigned int *frag_skbs;
+	unsigned int max_segs;
 
 	/* output as expected */
 	unsigned int nr_segs;
 	const unsigned int *segs;
+	bool segs_are_gso;
 };
 
 static struct gso_test_case cases[] = {
@@ -135,6 +147,54 @@ static struct gso_test_case cases[] = {
 		.nr_segs = 4,
 		.segs = (const unsigned int[]) { 100, 200, 300, 400 },
 	},
+	{
+		.id = GSO_TEST_BOUNDED,
+		.name = "bounded",
+		.linear_len = GSO_TEST_SIZE,
+		.nr_frags = 3,
+		.frags = (const unsigned int[]) {
+			GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
+		},
+		.max_segs = 2,
+		.nr_segs = 2,
+		.segs = (const unsigned int[]) {
+			2 * GSO_TEST_SIZE, GSO_TEST_SIZE + 3,
+		},
+		.segs_are_gso = true,
+	},
+	{
+		.id = GSO_TEST_BOUNDED_MULTI,
+		.name = "bounded_multi",
+		.linear_len = 2 * GSO_TEST_SIZE,
+		.nr_frags = 4,
+		.frags = (const unsigned int[]) {
+			GSO_TEST_SIZE, GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
+		},
+		.max_segs = 2,
+		.nr_segs = 3,
+		.segs = (const unsigned int[]) {
+			2 * GSO_TEST_SIZE, 2 * GSO_TEST_SIZE, GSO_TEST_SIZE + 3,
+		},
+		.segs_are_gso = true,
+	},
+	{
+		/*
+		 * One MSS per skb is what the unbounded path produces, so a
+		 * bound of a single segment must not change the output.
+		 */
+		.id = GSO_TEST_BOUNDED_ONE_MSS,
+		.name = "bounded_one_mss",
+		.linear_len = GSO_TEST_SIZE,
+		.nr_frags = 3,
+		.frags = (const unsigned int[]) {
+			GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
+		},
+		.max_segs = 1,
+		.nr_segs = 4,
+		.segs = (const unsigned int[]) {
+			GSO_TEST_SIZE, GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
+		},
+	},
 };
 
 static void gso_test_case_to_desc(struct gso_test_case *t, char *desc)
@@ -226,6 +286,7 @@ static void gso_test_func(struct kunit *test)
 	if (tcase->id == GSO_TEST_FRAG_LIST_NON_UNIFORM)
 		features &= ~NETIF_F_SG;
 
+	SKB_GSO_CB(skb)->max_segs = tcase->max_segs;
 	segs = skb_segment(skb, features);
 	if (IS_ERR(segs)) {
 		KUNIT_FAIL(test, "segs error %pe", segs);
@@ -239,6 +300,7 @@ static void gso_test_func(struct kunit *test)
 	for (cur = segs, i = 0; cur; cur = next, i++) {
 		next = cur->next;
 
+		KUNIT_ASSERT_LT(test, i, tcase->nr_segs);
 		KUNIT_ASSERT_EQ(test, cur->len, sizeof(hdr) + tcase->segs[i]);
 
 		/* segs have skb->data pointing to the mac header */
@@ -247,6 +309,17 @@ static void gso_test_func(struct kunit *test)
 
 		/* header was copied to all segs */
 		KUNIT_ASSERT_EQ(test, memcmp(skb_mac_header(cur), hdr, sizeof(hdr)), 0);
+		if (tcase->segs_are_gso) {
+			KUNIT_EXPECT_TRUE(test, skb_is_gso(cur));
+			KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_size,
+					GSO_TEST_SIZE);
+			KUNIT_EXPECT_LE(test, skb_shinfo(cur)->gso_segs,
+					tcase->max_segs);
+			KUNIT_EXPECT_FALSE(test, skb_shinfo(cur)->gso_type &
+					 SKB_GSO_PARTIAL);
+		} else if (tcase->max_segs) {
+			KUNIT_EXPECT_FALSE(test, skb_is_gso(cur));
+		}
 
 		/* last seg can be found through segs->prev pointer */
 		if (!next)
@@ -261,6 +334,277 @@ static void gso_test_func(struct kunit *test)
 	consume_skb(skb);
 }
 
+#define GSO_TCP_HDR_LEN \
+	(ETH_HLEN + sizeof(struct iphdr) + sizeof(struct tcphdr))
+
+static struct sk_buff *gso_tcp_skb_new(unsigned int payload_len)
+{
+	struct sk_buff *skb;
+	struct ethhdr *eth;
+	struct tcphdr *th;
+	struct iphdr *iph;
+
+	skb = alloc_skb(GSO_TCP_HDR_LEN + payload_len, GFP_KERNEL);
+	if (!skb)
+		return NULL;
+	skb_put_zero(skb, GSO_TCP_HDR_LEN + payload_len);
+
+	skb_reset_mac_header(skb);
+	eth = eth_hdr(skb);
+	eth->h_proto = htons(ETH_P_IP);
+	skb->protocol = eth->h_proto;
+
+	skb_set_network_header(skb, ETH_HLEN);
+	iph = ip_hdr(skb);
+	iph->version = 4;
+	iph->ihl = sizeof(*iph) / 4;
+	iph->protocol = IPPROTO_TCP;
+	iph->tot_len = htons(sizeof(*iph) + sizeof(*th) + payload_len);
+
+	skb_set_transport_header(skb, ETH_HLEN + sizeof(*iph));
+	th = tcp_hdr(skb);
+	th->doff = sizeof(*th) / 4;
+
+	skb->ip_summed = CHECKSUM_PARTIAL;
+	skb->csum_start = skb_transport_header(skb) - skb->head;
+	skb->csum_offset = offsetof(struct tcphdr, check);
+	skb_shinfo(skb)->gso_type = SKB_GSO_TCPV4;
+	skb_shinfo(skb)->gso_size = GSO_TEST_SIZE;
+	skb_shinfo(skb)->gso_segs = DIV_ROUND_UP(payload_len, GSO_TEST_SIZE);
+
+	return skb;
+}
+
+/*
+ * The transmit path passes the features of the device which rejects this
+ * skb, with the GSO bits cleared, so the engine segments it.
+ */
+static void gso_test_tcp_bounded_segment(struct kunit *test)
+{
+	netdev_features_t features = NETIF_F_SG | NETIF_F_HW_CSUM;
+	const unsigned int payload_len = 3 * GSO_TEST_SIZE + 3;
+	struct sk_buff *skb, *segs, *cur, *next;
+	const unsigned int expected[] = {
+		2 * GSO_TEST_SIZE, GSO_TEST_SIZE + 3,
+	};
+	const unsigned int max_segs = 2;
+	int i = 0;
+
+	if (!IS_ENABLED(CONFIG_INET))
+		kunit_skip(test, "requires CONFIG_INET");
+
+	skb = gso_tcp_skb_new(payload_len);
+	if (!skb) {
+		KUNIT_FAIL(test, "no skb");
+		return;
+	}
+
+	segs = __skb_gso_segment(skb, features, true, max_segs);
+	if (IS_ERR_OR_NULL(segs)) {
+		KUNIT_FAIL(test, "segs error %pe", segs);
+		consume_skb(skb);
+		return;
+	}
+
+	for (cur = segs; cur; cur = next, i++) {
+		next = cur->next;
+
+		KUNIT_ASSERT_LT(test, i, ARRAY_SIZE(expected));
+		KUNIT_EXPECT_EQ(test, cur->len,
+				GSO_TCP_HDR_LEN + expected[i]);
+		KUNIT_EXPECT_TRUE(test, skb_is_gso(cur));
+		KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_size,
+				GSO_TEST_SIZE);
+		KUNIT_EXPECT_LE(test, skb_shinfo(cur)->gso_segs, max_segs);
+
+		consume_skb(cur);
+	}
+
+	KUNIT_EXPECT_EQ(test, i, ARRAY_SIZE(expected));
+	consume_skb(skb);
+}
+
+#define GSO_TCP6_HDR_LEN \
+	(ETH_HLEN + sizeof(struct ipv6hdr) + sizeof(struct tcphdr))
+
+static struct sk_buff *gso_tcp6_skb_new(unsigned int payload_len)
+{
+	struct ipv6hdr *ip6h;
+	struct sk_buff *skb;
+	struct ethhdr *eth;
+	struct tcphdr *th;
+
+	skb = alloc_skb(GSO_TCP6_HDR_LEN + payload_len, GFP_KERNEL);
+	if (!skb)
+		return NULL;
+	skb_reserve(skb, NET_SKB_PAD);
+	skb_put_zero(skb, GSO_TCP6_HDR_LEN + payload_len);
+
+	skb_reset_mac_header(skb);
+	eth = eth_hdr(skb);
+	eth->h_proto = htons(ETH_P_IPV6);
+	skb->protocol = eth->h_proto;
+
+	skb_set_network_header(skb, ETH_HLEN);
+	ip6h = ipv6_hdr(skb);
+	ip6h->version = 6;
+	ip6h->nexthdr = IPPROTO_TCP;
+	ip6h->payload_len = htons(sizeof(*th) + payload_len);
+
+	skb_set_transport_header(skb, ETH_HLEN + sizeof(*ip6h));
+	th = tcp_hdr(skb);
+	th->doff = sizeof(*th) / 4;
+
+	skb->ip_summed = CHECKSUM_PARTIAL;
+	skb->csum_start = skb_transport_header(skb) - skb->head;
+	skb->csum_offset = offsetof(struct tcphdr, check);
+	skb_shinfo(skb)->gso_type = SKB_GSO_TCPV6;
+	skb_shinfo(skb)->gso_size = GSO_TEST_SIZE;
+	skb_shinfo(skb)->gso_segs = DIV_ROUND_UP(payload_len, GSO_TEST_SIZE);
+
+	return skb;
+}
+
+/*
+ * The device GSO size limit is per L3 protocol, so only a lowered limit of
+ * the packet's own protocol may cost it the TSO bit, and a tag inside the
+ * frame, which replaces skb->protocol with the ethertype, must not change
+ * which limit applies.  Takes ownership of @skb.
+ */
+static void gso_test_tcp_l3_limit(struct kunit *test, struct net_device *dev,
+				  struct sk_buff *skb, netdev_features_t tso,
+				  bool ipv6)
+{
+	unsigned int other = GSO_LEGACY_MAX_SIZE;
+	unsigned int own = GSO_MAX_SIZE;
+
+	/* The skb fits the limit of its own L3 protocol, not the other's. */
+	dev->gso_max_size = ipv6 ? own : other;
+	dev->gso_ipv4_max_size = ipv6 ? other : own;
+	KUNIT_EXPECT_TRUE(test, netif_skb_features(skb) & tso);
+
+	/* ...and the other way around. */
+	swap(own, other);
+	dev->gso_max_size = ipv6 ? own : other;
+	dev->gso_ipv4_max_size = ipv6 ? other : own;
+	KUNIT_EXPECT_FALSE(test, netif_skb_features(skb) & tso);
+
+	skb = vlan_insert_tag_set_proto(skb, htons(ETH_P_8021Q), 0);
+	if (!skb) {
+		KUNIT_FAIL(test, "no tagged skb");
+		return;
+	}
+	KUNIT_EXPECT_TRUE(test, skb->protocol == htons(ETH_P_8021Q));
+
+	swap(own, other);
+	dev->gso_max_size = ipv6 ? own : other;
+	dev->gso_ipv4_max_size = ipv6 ? other : own;
+	KUNIT_EXPECT_TRUE(test, netif_skb_features(skb) & tso);
+
+	swap(own, other);
+	dev->gso_max_size = ipv6 ? own : other;
+	dev->gso_ipv4_max_size = ipv6 ? other : own;
+	KUNIT_EXPECT_FALSE(test, netif_skb_features(skb) & tso);
+
+	consume_skb(skb);
+}
+
+static void gso_test_tcp_limit_l3_proto(struct kunit *test)
+{
+	static const struct net_device_ops dummy_netdev_ops = { };
+	const unsigned int payload_len = 100 * 1024;
+	struct net_device *dev;
+	struct sk_buff *skb;
+
+	dev = alloc_etherdev(0);
+	if (!dev) {
+		KUNIT_FAIL(test, "no net_device");
+		return;
+	}
+	dev->netdev_ops = &dummy_netdev_ops;
+	dev->hw_features = NETIF_F_SG | NETIF_F_HW_CSUM | NETIF_F_TSO |
+			   NETIF_F_TSO6;
+	dev->features = dev->hw_features;
+	dev->vlan_features = dev->hw_features;
+
+	skb = gso_tcp6_skb_new(payload_len);
+	if (!skb) {
+		KUNIT_FAIL(test, "no IPv6 skb");
+		goto free_dev;
+	}
+	skb->dev = dev;
+	gso_test_tcp_l3_limit(test, dev, skb, NETIF_F_TSO6, true);
+
+	skb = gso_tcp_skb_new(payload_len);
+	if (!skb) {
+		KUNIT_FAIL(test, "no IPv4 skb");
+		goto free_dev;
+	}
+	skb->dev = dev;
+	gso_test_tcp_l3_limit(test, dev, skb, NETIF_F_TSO, false);
+
+free_dev:
+	free_netdev(dev);
+}
+
+/*
+ * A bounded output is a plain GSO skb, so the whole length lands in the 16-bit
+ * L3 length field.  A device which accepts GSO skbs above 64 KiB would
+ * otherwise be handed outputs whose length field truncates.
+ */
+static void gso_test_tcp_resegment_l3_len(struct kunit *test)
+{
+	static const struct net_device_ops dummy_netdev_ops = { };
+	const unsigned int payload_len = 200 * 1024;
+	struct sk_buff *skb, *segs, *cur, *next;
+	unsigned int nr_segs = 0;
+	struct net_device *dev;
+	bool again = false;
+
+	if (!IS_ENABLED(CONFIG_INET))
+		kunit_skip(test, "requires CONFIG_INET");
+
+	dev = alloc_etherdev(0);
+	if (!dev) {
+		KUNIT_FAIL(test, "no net_device");
+		return;
+	}
+	dev->netdev_ops = &dummy_netdev_ops;
+	dev->hw_features = NETIF_F_SG | NETIF_F_HW_CSUM | NETIF_F_TSO;
+	dev->features = dev->hw_features;
+	dev->vlan_features = dev->hw_features;
+	dev->gso_ipv4_max_size = 120 * 1024;
+
+	skb = gso_tcp_skb_new(payload_len);
+	if (!skb) {
+		KUNIT_FAIL(test, "no skb");
+		goto free_dev;
+	}
+	skb->dev = dev;
+
+	segs = validate_xmit_skb_list(skb, dev, &again);
+	if (IS_ERR_OR_NULL(segs)) {
+		KUNIT_FAIL(test, "segs error %pe", segs);
+		goto free_dev;
+	}
+
+	for (cur = segs; cur; cur = next, nr_segs++) {
+		next = cur->next;
+		cur->next = NULL;
+
+		KUNIT_EXPECT_TRUE(test, skb_is_gso(cur));
+		KUNIT_EXPECT_EQ(test, ntohs(ip_hdr(cur)->tot_len),
+				cur->len - (skb_network_header(cur) -
+					    skb_mac_header(cur)));
+		consume_skb(cur);
+	}
+
+	KUNIT_EXPECT_EQ(test, nr_segs, 4);
+
+free_dev:
+	free_netdev(dev);
+}
+
 /* IP tunnel flags */
 
 #include <net/ip_tunnels.h>
@@ -372,6 +716,9 @@ static void ip_tunnel_flags_test_run(struct kunit *test)
 
 static struct kunit_case net_test_cases[] = {
 	KUNIT_CASE_PARAM(gso_test_func, gso_test_gen_params),
+	KUNIT_CASE(gso_test_tcp_bounded_segment),
+	KUNIT_CASE(gso_test_tcp_limit_l3_proto),
+	KUNIT_CASE(gso_test_tcp_resegment_l3_len),
 	KUNIT_CASE_PARAM(ip_tunnel_flags_test_run,
 			 ip_tunnel_flags_test_gen_params),
 	{ },
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs
  2026-09-28  4:40 [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs Wang Zhan
                   ` (4 preceding siblings ...)
  2026-09-28  4:41 ` [PATCH net-next v3 5/5] net: net_test: add tests for bounded GSO segmentation Wang Zhan
@ 2026-09-28  4:45 ` netdev-bot+sinfo
  2026-09-28  5:49   ` Wang Zhan
  5 siblings, 1 reply; 27+ messages in thread
From: netdev-bot+sinfo @ 2026-09-28  4:45 UTC (permalink / raw)
  To: Wang Zhan
  Cc: netdev, davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight

Hi!

This is an automated message. This series looks like a fix, but its
commit messages seem to be missing some information:

 - How the issue was discovered, e.g. hit in production, hit during
   development, syzbot report, manual code inspection, LLM or static
   analysis tool scan.

 - Whether the issue was actually triggered, or is only theoretical
   (e.g. found by code inspection). If it was triggered please include
   the symptoms, like the stack trace or error messages.

Please do not repost the series just to address the above. Instead,
reply to this email with the missing information, so that reviewers
can take it into account. If the series needs another revision for
other reasons, please include the information in the commit messages
then.

The evaluation is done by an LLM so it may be wrong, if you think
that is the case please reply and explain.

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs
  2026-09-28  4:45 ` [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs netdev-bot+sinfo
@ 2026-09-28  5:49   ` Wang Zhan
  2026-09-28 23:34     ` Willem de Bruijn
  0 siblings, 1 reply; 27+ messages in thread
From: Wang Zhan @ 2026-09-28  5:49 UTC (permalink / raw)
  To: netdev-bot+sinfo, netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

On Mon, 28 Sep 2026 04:45:18 +0000 netdev-bot+sinfo@kernel.org wrote:
> This is an automated message. This series looks like a fix, but its
> commit messages seem to be missing some information:
>
>  - How the issue was discovered, e.g. hit in production, hit during
>    development, syzbot report, manual code inspection, LLM or static
>    analysis tool scan.
>
>  - Whether the issue was actually triggered, or is only theoretical
>    (e.g. found by code inspection). If it was triggered please include
>    the symptoms, like the stack trace or error messages.

Only 1/5 carries a Fixes tag; the rest is new feature work.  1/5 was found
by code inspection while writing the resegmentation path, and it is covered
by the KUnit case gso_test_tcp_limit_l3_proto() in 5/5, which fails on the
unpatched tree: a tagged IPv6 frame is measured against the IPv4 limit and
segmented although the device could send it as one TSO frame, no crash.

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs
  2026-09-28  5:49   ` Wang Zhan
@ 2026-09-28 23:34     ` Willem de Bruijn
  2026-09-29  7:27       ` Paolo Abeni
  2026-09-29 11:50       ` Wang Zhan
  0 siblings, 2 replies; 27+ messages in thread
From: Willem de Bruijn @ 2026-09-28 23:34 UTC (permalink / raw)
  To: Wang Zhan, netdev-bot+sinfo, netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

Wang Zhan wrote:
> On Mon, 28 Sep 2026 04:45:18 +0000 netdev-bot+sinfo@kernel.org wrote:
> > This is an automated message. This series looks like a fix, but its
> > commit messages seem to be missing some information:
> >
> >  - How the issue was discovered, e.g. hit in production, hit during
> >    development, syzbot report, manual code inspection, LLM or static
> >    analysis tool scan.
> >
> >  - Whether the issue was actually triggered, or is only theoretical
> >    (e.g. found by code inspection). If it was triggered please include
> >    the symptoms, like the stack trace or error messages.
> 
> Only 1/5 carries a Fixes tag; the rest is new feature work.  1/5 was found
> by code inspection while writing the resegmentation path, and it is covered
> by the KUnit case gso_test_tcp_limit_l3_proto() in 5/5, which fails on the
> unpatched tree: a tagged IPv6 frame is measured against the IPv4 limit and
> segmented although the device could send it as one TSO frame, no crash.

If 1/5 fixes a bug that is reachable today (does it?) then it needs to
go to net on its own.

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 1/5] net: core: use the packet's L3 protocol for the GSO size limit
  2026-09-28  4:40 ` [PATCH net-next v3 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
@ 2026-09-28 23:36   ` Willem de Bruijn
  2026-09-29  3:44     ` Wang Zhan
  2026-09-29  4:01     ` Wang Zhan
  0 siblings, 2 replies; 27+ messages in thread
From: Willem de Bruijn @ 2026-09-28 23:36 UTC (permalink / raw)
  To: Wang Zhan, netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

Wang Zhan wrote:
> gso_features_check() compares the frame length against
> netif_get_gso_max_size(), which picks the IPv4 or the IPv6 limit from
> skb->protocol.  A tag which is already inside the frame replaces that field
> with the VLAN ethertype, as skb_vlan_push() does, and an IPv6 packet is
> then measured against the IPv4 limit and segmented although the device
> could send it as one TSO frame.
> 
> Let the limit lookup take the protocol as an argument, and pass the L3
> protocol, so where the tag sits does not decide which limit applies.  The
> helper cannot look behind the tag itself: it lives in netdevice.h, which
> cannot include if_vlan.h because that header includes netdevice.h.
> 
> Fixes: e609c959a9396 ("net: Fix gso_features_check to check for both dev->gso_{ipv4_,}max_size")
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
> 
> ---
> v3:
> - new patch: the L3 protocol fix split out of the resegmentation patch
> v2: https://lore.kernel.org/20260918084651.3022878-4-wang.zhan@smartx.com/
> v1: https://lore.kernel.org/20260917063854.2011613-4-wang.zhan@smartx.com/
> ---
>  include/linux/netdevice.h | 4 ++--
>  net/core/dev.c            | 3 ++-
>  2 files changed, 4 insertions(+), 3 deletions(-)
> 
> diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
> index d037faff7c44b..1c28861fcbe2c 100644
> --- a/include/linux/netdevice.h
> +++ b/include/linux/netdevice.h
> @@ -5562,10 +5562,10 @@ netif_get_gro_max_size(const struct net_device *dev, const struct sk_buff *skb)
>  }
>  
>  static inline unsigned int
> -netif_get_gso_max_size(const struct net_device *dev, const struct sk_buff *skb)
> +netif_get_gso_max_size(const struct net_device *dev, __be16 protocol)
>  {
>  	/* pairs with WRITE_ONCE() in netif_set_gso(_ipv4)_max_size() */
> -	return skb->protocol == htons(ETH_P_IPV6) ?
> +	return protocol == htons(ETH_P_IPV6) ?
>  	       READ_ONCE(dev->gso_max_size) :
>  	       READ_ONCE(dev->gso_ipv4_max_size);
>  }
> diff --git a/net/core/dev.c b/net/core/dev.c
> index f660fccfc0dbc..ffa9b0c27788c 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -3843,7 +3843,8 @@ static netdev_features_t gso_features_check(const struct sk_buff *skb,
>  	if (gso_segs > READ_ONCE(dev->gso_max_segs))
>  		return features & ~NETIF_F_GSO_MASK;
>  
> -	if (unlikely(skb->len >= netif_get_gso_max_size(dev, skb)))
> +	if (unlikely(skb->len >=
> +		     netif_get_gso_max_size(dev, vlan_get_protocol(skb))))

If netif_get_gso_max_size needs to inspect the result from 
vlan_get_protocol, call it inside that function directly? More robust
and a one line change.

>  		return features & ~NETIF_F_GSO_MASK;
>  
>  	if (!skb_shinfo(skb)->gso_type) {
> -- 
> 2.47.3
> 



^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 2/5] net: core: factor out the GSO device limit check
  2026-09-28  4:40 ` [PATCH net-next v3 2/5] net: core: factor out the GSO device limit check Wang Zhan
@ 2026-09-28 23:37   ` Willem de Bruijn
  2026-09-29  7:47   ` Paolo Abeni
  1 sibling, 0 replies; 27+ messages in thread
From: Willem de Bruijn @ 2026-09-28 23:37 UTC (permalink / raw)
  To: Wang Zhan, netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

Wang Zhan wrote:
> gso_features_check() decides whether an egress device can offload a GSO
> skb as a single TSO frame by comparing the segment count and the frame
> length against the device limits.  The resegmentation path added by a
> later patch caps the MSS segments per output skb and needs the same
> features computed without those limits, so put the two tests in a helper
> and gate them with one flag.
> 
> No functional changes.
> 
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
> 
> ---
> v3:
> - rename the helper to gso_within_dev_limits
> - move the check_gso_limits flag and the wrapper split here from patch 4
> v2: https://lore.kernel.org/20260918084651.3022878-2-wang.zhan@smartx.com/
> v1: https://lore.kernel.org/20260917063854.2011613-2-wang.zhan@smartx.com/
> ---
>  net/core/dev.c | 29 +++++++++++++++++++----------
>  1 file changed, 19 insertions(+), 10 deletions(-)
> 
> diff --git a/net/core/dev.c b/net/core/dev.c
> index ffa9b0c27788c..d66b667071837 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -3834,17 +3834,19 @@ static bool skb_gso_has_extension_hdr(const struct sk_buff *skb)
>  			 skb_inner_network_header_len(skb) != sizeof(struct ipv6hdr)));
>  }
>  
> +static bool
> +gso_within_dev_limits(const struct sk_buff *skb, const struct net_device *dev)
> +{
> +	return skb_shinfo(skb)->gso_segs <= READ_ONCE(dev->gso_max_segs) &&
> +	       skb->len < netif_get_gso_max_size(dev, vlan_get_protocol(skb));

nit: keep the separate return statements as before. It's easier to read.
> +}
> +
>  static netdev_features_t gso_features_check(const struct sk_buff *skb,
>  					    struct net_device *dev,
> -					    netdev_features_t features)
> +					    netdev_features_t features,
> +					    bool check_limits)
>  {
> -	u16 gso_segs = skb_shinfo(skb)->gso_segs;
> -
> -	if (gso_segs > READ_ONCE(dev->gso_max_segs))
> -		return features & ~NETIF_F_GSO_MASK;
> -
> -	if (unlikely(skb->len >=
> -		     netif_get_gso_max_size(dev, vlan_get_protocol(skb))))
> +	if (check_limits && !gso_within_dev_limits(skb, dev))
>  		return features & ~NETIF_F_GSO_MASK;
>  
>  	if (!skb_shinfo(skb)->gso_type) {
> @@ -3893,13 +3895,15 @@ static netdev_features_t gso_features_check(const struct sk_buff *skb,
>  	return features;
>  }
>  
> -netdev_features_t netif_skb_features(struct sk_buff *skb)
> +static netdev_features_t
> +__netif_skb_features(struct sk_buff *skb, bool check_gso_limits)
>  {
>  	struct net_device *dev = skb->dev;
>  	netdev_features_t features = dev->features;
>  
>  	if (skb_is_gso(skb))
> -		features = gso_features_check(skb, dev, features);
> +		features = gso_features_check(skb, dev, features,
> +					      check_gso_limits);
>  
>  	/* If encapsulation offload request, verify we are testing
>  	 * hardware encapsulation features instead of standard
> @@ -3922,6 +3926,11 @@ netdev_features_t netif_skb_features(struct sk_buff *skb)
>  
>  	return harmonize_features(skb, features);
>  }
> +
> +netdev_features_t netif_skb_features(struct sk_buff *skb)
> +{
> +	return __netif_skb_features(skb, true);
> +}
>  EXPORT_SYMBOL(netif_skb_features);
>  
>  static int xmit_one(struct sk_buff *skb, struct net_device *dev,
> -- 
> 2.47.3
> 



^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 3/5] net: gso: support bounded TCP segmentation
  2026-09-28  4:41 ` [PATCH net-next v3 3/5] net: gso: support bounded TCP segmentation Wang Zhan
@ 2026-09-28 23:39   ` Willem de Bruijn
  2026-09-30  4:41   ` netdev-bot+sashiko
  1 sibling, 0 replies; 27+ messages in thread
From: Willem de Bruijn @ 2026-09-28 23:39 UTC (permalink / raw)
  To: Wang Zhan, netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

Wang Zhan wrote:
> The bounded resegmentation added by the next patch splits an oversized TCP
> GSO skb into several GSO skbs which fit the device limits.  That needs the
> GSO engine to group several MSS segments into one output skb, so let
> callers bound the number of MSS segments each output skb carries and pass
> the bound through the existing __skb_gso_segment() entry point.  Ordinary
> callers use zero for no limit.
> 
> Without a bound, skb_segment() groups several MSS into one output skb only
> for a device which advertises NETIF_F_GSO_PARTIAL or for a skb with a
> frag_list which can be split into uniform pieces, and falls back to one
> segment per skb otherwise.  A caller which passes a bound asks for the
> grouping regardless, so the block which makes that decision is skipped
> when max_segs is set.  Callers which pass no bound keep it, and the bounded
> path only runs for skbs which carry no frag_list.
> 
> The output stays a GSO skb: gso_size is the original MSS and gso_segs is
> the number of MSS it holds, so a downstream device can still perform
> ordinary TSO.  Store the bound in the existing skb_gso_cb scratch context,
> alongside the call-local data_offset and mac_offset fields, so that the
> segmentation methods keep their signature.  A zero max_segs value means
> that no bound is active; it is not a persistent skb flag.
> 
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
> 
> ---
> v3:
> - drop the tcp_gso_segment() exception: the caller keeps the features
> - cap the bound with GSO_MAX_SEGS instead of U16_MAX
> - drop the reset of the field in the output skbs: every caller initializes it
> - note in the kernel-doc that a bound is for TCP GSO skbs only
> - note at the frag_list gate that a bound must not go with a frag_list skb
> v2: https://lore.kernel.org/20260918084651.3022878-3-wang.zhan@smartx.com/
> v1: https://lore.kernel.org/20260917063854.2011613-3-wang.zhan@smartx.com/

> @@ -4839,7 +4840,13 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
>  	csum = !!can_checksum_protocol(features, proto);
>  
>  	if (sg && csum && !gso_by_frags)  {
> -		if (!(features & NETIF_F_GSO_PARTIAL)) {
> +		/*
> +		 * A call with max_segs set groups the MSS segments below,
> +		 * and the frag_list split in this block does not carry the
> +		 * bound, so a bound must not be passed for a skb with a
> +		 * frag_list.
> +		 */

To me this comment confuses rather than helps.

For one, it refers to "the bound" without any context.

Probably just drop.

> +		if (!max_segs && !(features & NETIF_F_GSO_PARTIAL)) {
>  			struct sk_buff *iter;
>  			unsigned int frag_len;
>  


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs
  2026-09-28  4:41 ` [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs Wang Zhan
@ 2026-09-28 23:47   ` Willem de Bruijn
  2026-09-29 10:25     ` Wang Zhan
  2026-09-30  4:41   ` netdev-bot+sashiko
  1 sibling, 1 reply; 27+ messages in thread
From: Willem de Bruijn @ 2026-09-28 23:47 UTC (permalink / raw)
  To: Wang Zhan, netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

Wang Zhan wrote:
> A GSO skb which exceeds an egress device limit loses its GSO feature mask
> and is segmented into individual packets.  This is unnecessarily expensive
> when the device can still offload smaller TCP GSO skbs, which is easy to
> hit once one hop of a BIG TCP path raises gso_max_size and the next one
> does not.
> 
> For an unencapsulated TCP GSO skb which exceeds gso_max_size or
> gso_max_segs, work out how many MSS segments each output skb may carry and
> resegment the skb with that max_segs bound instead.  Encapsulated and
> frag-list skbs, GSO types the device cannot offload and bounds which leave
> room for a single MSS keep today's segmentation.  The output obeys the GSO
> feature and limit contract the device already advertises, so it applies
> automatically, without extra device state or a userspace control.
> 
> The output is a plain GSO skb, so its whole length lands in the 16-bit L3
> length field which inet_gso_segment() and ipv6_gso_segment() write.  An
> egress limit above 64 KiB would give outputs whose length truncates, so
> the size limit is also capped at what that field can express.
> 
> The helper runs on the skb which is handed to the driver, after
> validate_xmit_vlan() and sk_validate_xmit_skb(), and only from the
> netif_needs_gso() branch, so an skb which is not segmented pays nothing.
> An over-limit skb pays one device limit test and one ndo_features_check()
> for the bound, in exchange for keeping the output a GSO skb.
> 
> Measured on a veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP
> enabled on the veth endpoints and left off in the guest, so the skbs which
> the veth hop accepts have to be segmented before the TAP device.  A single
> iperf3 TCP flow, six alternating runs per state (`-t 15 -O 5`, fixed CPU
> affinity and port tuple).  The middle column is the same tree with the
> resegmentation disabled:
> 
>   protocol  no BIG TCP   mixed, no reseg  mixed, resegmented
>   TCP/IPv4  51.550 Gbps  15.850 Gbps      52.617 Gbps
>   TCP/IPv6  52.050 Gbps  15.783 Gbps      51.933 Gbps
> 
> Coefficient of variation for the two mixed columns was 0.48% and 0.82% for
> IPv4 and 0.44% and 0.44% for IPv6.  A BIG TCP hop which feeds a 64 KiB hop
> loses 69% of the throughput of a path which never enables BIG TCP at all;
> bounded resegmentation recovers it, 3.3x over the existing segmentation
> path and within noise of the no BIG TCP baseline.
> 
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
> 
> ---
> v3:
> - the L3 protocol change moved to patch 1
> - fold skb_can_gso_resegment() into the helper, now skb_gso_output_max_segs()
> - pass the caller's features on unchanged, the bound only shapes the output
> - drop the scatter-gather and checksum tests skb_segment() applies itself
> - saturate the size limit with min(), drop the sub-MSS guard
> - cap the output size at GSO_LEGACY_MAX_SIZE, these outputs are not BIG TCP
> v2: https://lore.kernel.org/20260918084651.3022878-4-wang.zhan@smartx.com/
> v1: https://lore.kernel.org/20260917063854.2011613-4-wang.zhan@smartx.com/
> ---
>  net/core/dev.c | 63 +++++++++++++++++++++++++++++++++++++++++++++++++-
>  1 file changed, 62 insertions(+), 1 deletion(-)
> 
> diff --git a/net/core/dev.c b/net/core/dev.c
> index d66b667071837..728260772f349 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -3933,6 +3933,63 @@ netdev_features_t netif_skb_features(struct sk_buff *skb)
>  }
>  EXPORT_SYMBOL(netif_skb_features);
>  
> +static unsigned int
> +skb_gso_output_max_segs(struct sk_buff *skb, struct net_device *dev)
> +{
> +	unsigned int mss = skb_shinfo(skb)->gso_size;
> +	unsigned int gso_max_size, hdr_len, max_segs;
> +	netdev_features_t features;
> +	struct tcphdr _tcph, *th;
> +
> +	/*
> +	 * The TCP frag-list path segments through skb_segment_list(), which
> +	 * does not carry max_segs, so bounded calls skip those skbs.
> +	 */

This comment answers only one of six conditions. And one that is
pretty straightforward. I'd drop.

In general, drop all too-obvious comments. AI has a habit of adding
a lot more, and more low information, comments than is customary in
kernel code (where we also have commit messages). Generally, repeating
what the code does is of little value.

> +	if (!skb_is_gso(skb) || !skb_is_gso_tcp(skb) ||
> +	    skb->encapsulation || skb_has_frag_list(skb) ||
> +	    !skb_mac_header_was_set(skb) ||
> +	    !skb_transport_header_was_set(skb))

Conversely, they last two conditions are less obvious. Are they not
always true for a TSO packet?

> +		return 0;
> +
> +	th = skb_header_pointer(skb, skb_transport_offset(skb), sizeof(_tcph),
> +				&_tcph);
> +	if (!th || th->doff < sizeof(*th) / 4)
> +		return 0;
> +
> +	hdr_len = skb_transport_header(skb) - skb_mac_header(skb) +
> +		  th->doff * 4;
> +
> +	/*
> +	 * The output stays a GSO skb, so the device has to offload the GSO
> +	 * type.  The caller's features cannot tell that: this skb is over
> +	 * the device limits, so gso_features_check() has cleared their GSO
> +	 * bits.  Compute the features again without the limit checks.
> +	 */
> +	features = __netif_skb_features(skb, false);
> +	if (!net_gso_ok(features | NETIF_F_GSO_ROBUST,
> +			skb_shinfo(skb)->gso_type))
> +		return 0;
> +
> +	/*
> +	 * The output is a plain GSO skb, so its whole length lands in the
> +	 * 16-bit L3 length field: only a BIG TCP skb may exceed the legacy
> +	 * GSO size, and this path does not emit one.
> +	 */
> +	gso_max_size = netif_get_gso_max_size(dev, vlan_get_protocol(skb));

Third time this is now called in validate_xmit_skb. Not sure if that can
easily be avoided.

> +	gso_max_size = min(gso_max_size, GSO_LEGACY_MAX_SIZE);
> +
> +	/*
> +	 * gso_within_dev_limits() accepts gso_segs == gso_max_segs but
> +	 * rejects skb->len >= gso_max_size, so only the size bound needs - 1.
> +	 * The min() keeps that subtraction from wrapping when the device
> +	 * limit is smaller than the headers.
> +	 */
> +	max_segs = (gso_max_size - min(gso_max_size, hdr_len + 1)) / mss;
> +	max_segs = min(max_segs, READ_ONCE(dev->gso_max_segs));
> +
> +	return max_segs;
> +}
> +
>  static int xmit_one(struct sk_buff *skb, struct net_device *dev,
>  		    struct netdev_queue *txq, bool more)
>  {
> @@ -4097,9 +4154,13 @@ static struct sk_buff *validate_xmit_skb(struct sk_buff *skb, struct net_device
>  		goto out_null;
>  
>  	if (netif_needs_gso(skb, features)) {
> +		unsigned int max_segs = 0;
>  		struct sk_buff *segs;
>  
> -		segs = skb_gso_segment(skb, features);
> +		if (unlikely(!gso_within_dev_limits(skb, dev)))
> +			max_segs = skb_gso_output_max_segs(skb, dev);
> +
> +		segs = __skb_gso_segment(skb, features, true, max_segs);
>  		if (IS_ERR(segs)) {
>  			goto out_kfree_skb;
>  		} else if (segs) {
> -- 
> 2.47.3
> 



^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 5/5] net: net_test: add tests for bounded GSO segmentation
  2026-09-28  4:41 ` [PATCH net-next v3 5/5] net: net_test: add tests for bounded GSO segmentation Wang Zhan
@ 2026-09-28 23:59   ` Willem de Bruijn
  2026-09-29 10:30     ` Wang Zhan
  2026-09-30  4:41   ` netdev-bot+sashiko
  1 sibling, 1 reply; 27+ messages in thread
From: Willem de Bruijn @ 2026-09-28 23:59 UTC (permalink / raw)
  To: Wang Zhan, netdev
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
	Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
	Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
	Wang Zhan

Wang Zhan wrote:
> The GSO engine can now be asked to bound the number of MSS segments which
> go into each output skb.  Add KUnit coverage for it.
> 
> The parameterized GSO test gains a max_segs input and three cases: two
> bounds which cut the input into two and three output skbs, and a bound of
> one MSS, which must leave the ungrouped output of the unbounded path
> alone.  It drives skb_segment() directly, because the synthetic protocol
> it uses has no gso_segment callback, and stores the bound in the GSO
> control block itself.
> 
> The TCP test drives __skb_gso_segment() with a bound of two MSS and checks
> that every output skb stays GSO, keeps its gso_size, and stays within the
> bound.  The length test runs a 200 KiB TCP skb through
> validate_xmit_skb_list(), the caller which sets the bound, and checks that
> the length declared by every output matches the L3 length of that output.
> The limit test checks that the GSO size limit which netif_skb_features()
> applies follows the packet's L3 protocol, for an IPv4 and an IPv6 skb,
> also with the tag inside the frame.
> 
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
> 
> ---
> v3:
> - add cases for the length a bounded output declares and for a bound of one
>   MSS, which must leave the output ungrouped
> - cover the IPv4 half of the limit test, checking the protocol's own bit
> - skip the TCP cases without CONFIG_INET and bound the expected index in
>   the parameterized loop
> - free the skb and the device on the failure paths
> - reserve headroom so the VLAN step needs no atomic allocation
> - drop NETIF_F_TSO: the transmit path passes the features without it
> v2: https://lore.kernel.org/20260918084651.3022878-5-wang.zhan@smartx.com/
> v1: https://lore.kernel.org/20260917063854.2011613-5-wang.zhan@smartx.com/
> ---
>  net/core/net_test.c | 347 ++++++++++++++++++++++++++++++++++++++++++++
>  1 file changed, 347 insertions(+)
> 
> diff --git a/net/core/net_test.c b/net/core/net_test.c
> index 9c3a590865d26..a4f61398a90ab 100644
> --- a/net/core/net_test.c
> +++ b/net/core/net_test.c
> @@ -4,7 +4,14 @@
>  
>  /* GSO */
>  
> +#include <linux/if_ether.h>
> +#include <linux/if_vlan.h>
> +#include <linux/ip.h>
> +#include <linux/ipv6.h>
> +#include <linux/netdevice.h>
>  #include <linux/skbuff.h>
> +#include <linux/tcp.h>
> +#include <net/gso.h>
>  
>  static const char hdr[] = "abcdefgh";
>  #define GSO_TEST_SIZE 1000
> @@ -34,6 +41,9 @@ enum gso_test_nr {
>  	GSO_TEST_FRAG_LIST_PURE,
>  	GSO_TEST_FRAG_LIST_NON_UNIFORM,
>  	GSO_TEST_GSO_BY_FRAGS,
> +	GSO_TEST_BOUNDED,
> +	GSO_TEST_BOUNDED_MULTI,
> +	GSO_TEST_BOUNDED_ONE_MSS,

nit: bound is not a helpful name for this feature.

The default segments to a stream of skbs of MSS 1, so segmenting to
a stream of skbs larger MSS to me is not bounding. Quite the opposite.

Not a comment only about this test patch.

Perhaps partial or re-segmentation better captures it.

> +	{
> +		/*
> +		 * One MSS per skb is what the unbounded path produces, so a
> +		 * bound of a single segment must not change the output.
> +		 */

so setting max_segs = 0 is equivalent to setting max_segs = 1.

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 1/5] net: core: use the packet's L3 protocol for the GSO size limit
  2026-09-28 23:36   ` Willem de Bruijn
@ 2026-09-29  3:44     ` Wang Zhan
  2026-09-29  4:01     ` Wang Zhan
  1 sibling, 0 replies; 27+ messages in thread
From: Wang Zhan @ 2026-09-29  3:44 UTC (permalink / raw)
  To: netdev, Willem de Bruijn
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun, Wang Zhan,
	Jason Wang, Andrew Lunn, Aaron Conole, Eelco Chaudron,
	Ilya Maximets, dev, Daniel Borkmann, Neal Cardwell,
	Kuniyuki Iwashima, Alice Mikityanska, David Laight

On Mon, 28 Sep 2026 19:36:01 -0400 Willem de Bruijn wrote:
> If netif_get_gso_max_size needs to inspect the result from
> vlan_get_protocol, call it inside that function directly? More robust
> and a one line change.

That's a cycle: vlan_get_protocol() is in if_vlan.h and
netif_get_gso_max_size() in netdevice.h, and if_vlan.h includes
netdevice.h.  The commit message says that.

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 1/5] net: core: use the packet's L3 protocol for the GSO size limit
  2026-09-28 23:36   ` Willem de Bruijn
  2026-09-29  3:44     ` Wang Zhan
@ 2026-09-29  4:01     ` Wang Zhan
  2026-09-29 14:59       ` Willem de Bruijn
  1 sibling, 1 reply; 27+ messages in thread
From: Wang Zhan @ 2026-09-29  4:01 UTC (permalink / raw)
  To: netdev, Willem de Bruijn
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun, Wang Zhan,
	Jason Wang, Andrew Lunn, Aaron Conole, Eelco Chaudron,
	Ilya Maximets, dev, Daniel Borkmann, Neal Cardwell,
	Kuniyuki Iwashima, Alice Mikityanska, David Laight

On Mon, 28 Sep 2026 19:36:01 -0400 Willem de Bruijn wrote:
> If netif_get_gso_max_size needs to inspect the result from
> vlan_get_protocol, call it inside that function directly? More robust
> and a one line change.

Or we can move it to dev.c, only one caller in-tree (before this series).

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs
  2026-09-28 23:34     ` Willem de Bruijn
@ 2026-09-29  7:27       ` Paolo Abeni
  2026-09-29 11:50       ` Wang Zhan
  1 sibling, 0 replies; 27+ messages in thread
From: Paolo Abeni @ 2026-09-29  7:27 UTC (permalink / raw)
  To: Willem de Bruijn, Wang Zhan, netdev-bot+sinfo, netdev
  Cc: davem, edumazet, kuba, horms, keyong.sun, Jason Wang, Andrew Lunn,
	Aaron Conole, Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
	Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight

On 9/29/26 01:34, Willem de Bruijn wrote:
> Wang Zhan wrote:
>> On Mon, 28 Sep 2026 04:45:18 +0000 netdev-bot+sinfo@kernel.org wrote:
>>> This is an automated message. This series looks like a fix, but its
>>> commit messages seem to be missing some information:
>>>
>>>   - How the issue was discovered, e.g. hit in production, hit during
>>>     development, syzbot report, manual code inspection, LLM or static
>>>     analysis tool scan.
>>>
>>>   - Whether the issue was actually triggered, or is only theoretical
>>>     (e.g. found by code inspection). If it was triggered please include
>>>     the symptoms, like the stack trace or error messages.
>>
>> Only 1/5 carries a Fixes tag; the rest is new feature work.  1/5 was found
>> by code inspection while writing the resegmentation path, and it is covered
>> by the KUnit case gso_test_tcp_limit_l3_proto() in 5/5, which fails on the
>> unpatched tree: a tagged IPv6 frame is measured against the IPv4 limit and
>> segmented although the device could send it as one TSO frame, no crash.
> 
> If 1/5 fixes a bug that is reachable today (does it?) then it needs to
> go to net on its own.

FTR I don't think patch 1/5 is stable material: I would keep it in this
series stripping the fixes tag.

/P


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 2/5] net: core: factor out the GSO device limit check
  2026-09-28  4:40 ` [PATCH net-next v3 2/5] net: core: factor out the GSO device limit check Wang Zhan
  2026-09-28 23:37   ` Willem de Bruijn
@ 2026-09-29  7:47   ` Paolo Abeni
  1 sibling, 0 replies; 27+ messages in thread
From: Paolo Abeni @ 2026-09-29  7:47 UTC (permalink / raw)
  To: Wang Zhan, netdev
  Cc: davem, edumazet, kuba, horms, keyong.sun, Willem de Bruijn,
	Jason Wang, Andrew Lunn, Aaron Conole, Eelco Chaudron,
	Ilya Maximets, dev, Daniel Borkmann, Neal Cardwell,
	Kuniyuki Iwashima, Alice Mikityanska, David Laight

On 9/28/26 06:40, Wang Zhan wrote:
> @@ -3922,6 +3926,11 @@ netdev_features_t netif_skb_features(struct sk_buff *skb)
>   
>   	return harmonize_features(skb, features);
>   }
> +
> +netdev_features_t netif_skb_features(struct sk_buff *skb)
> +{
> +	return __netif_skb_features(skb, true);
> +}
>   EXPORT_SYMBOL(netif_skb_features);

Nit: I *think* that moving the above to an header file and exporting
__netif_skb_features() will generate better code, as the latter should
not be inlined (large functions with 2 callers in the same compilation
unit).

/P


^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs
  2026-09-28 23:47   ` Willem de Bruijn
@ 2026-09-29 10:25     ` Wang Zhan
  2026-09-29 15:00       ` Willem de Bruijn
  0 siblings, 1 reply; 27+ messages in thread
From: Wang Zhan @ 2026-09-29 10:25 UTC (permalink / raw)
  To: netdev, Willem de Bruijn
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun, Wang Zhan,
	Jason Wang, Andrew Lunn, Aaron Conole, Eelco Chaudron,
	Ilya Maximets, dev, Daniel Borkmann, Neal Cardwell,
	Kuniyuki Iwashima, Alice Mikityanska, David Laight

On Mon, 28 Sep 2026 19:47:43 -0400 Willem de Bruijn wrote:
> > +	/*
> > +	 * The TCP frag-list path segments through skb_segment_list(), which
> > +	 * does not carry max_segs, so bounded calls skip those skbs.
> > +	 */
>
> This comment answers only one of six conditions. And one that is
> pretty straightforward. I'd drop.
>
> In general, drop all too-obvious comments. AI has a habit of adding
> a lot more, and more low information, comments than is customary in
> kernel code (where we also have commit messages). Generally, repeating
> what the code does is of little value.

Dropped in v4.  I will check all the comments in the series.

> > +	if (!skb_is_gso(skb) || !skb_is_gso_tcp(skb) ||
> > +	    skb->encapsulation || skb_has_frag_list(skb) ||
> > +	    !skb_mac_header_was_set(skb) ||
> > +	    !skb_transport_header_was_set(skb))
>
> Conversely, they last two conditions are less obvious. Are they not
> always true for a TSO packet?

The transport header can be missing.  qdisc_pkt_len_segs_init() does the
same check on this path (net/core/dev.c:4245, a0dce8752193e).

The mac header is always set.  It can be dropped in v4.

> > +	gso_max_size = netif_get_gso_max_size(dev, vlan_get_protocol(skb));
>
> Third time this is now called in validate_xmit_skb. Not sure if that can
> easily be avoided.

Maybe we can pass the oversize and gso_max_size flags out of
__netif_skb_features, but it would be a bit ugly.  I think the current
cost is acceptable.

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 5/5] net: net_test: add tests for bounded GSO segmentation
  2026-09-28 23:59   ` Willem de Bruijn
@ 2026-09-29 10:30     ` Wang Zhan
  2026-09-29 15:01       ` Willem de Bruijn
  0 siblings, 1 reply; 27+ messages in thread
From: Wang Zhan @ 2026-09-29 10:30 UTC (permalink / raw)
  To: netdev, Willem de Bruijn
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun, Wang Zhan,
	Jason Wang, Andrew Lunn, Aaron Conole, Eelco Chaudron,
	Ilya Maximets, dev, Daniel Borkmann, Neal Cardwell,
	Kuniyuki Iwashima, Alice Mikityanska, David Laight

On Mon, 28 Sep 2026 19:59:54 -0400 Willem de Bruijn wrote:
> > +	GSO_TEST_BOUNDED,
> > +	GSO_TEST_BOUNDED_MULTI,
> > +	GSO_TEST_BOUNDED_ONE_MSS,
>
> nit: bound is not a helpful name for this feature.
>
> The default segments to a stream of skbs of MSS 1, so segmenting to
> a stream of skbs larger MSS to me is not bounding. Quite the opposite.
>
> Not a comment only about this test patch.
>
> Perhaps partial or re-segmentation better captures it.

Agreed.  I think re-segmentation is descriptive enough.  Partial is
already for NETIF_F_GSO_PARTIAL.

> > +	{
> > +		/*
> > +		 * One MSS per skb is what the unbounded path produces, so a
> > +		 * bound of a single segment must not change the output.
> > +		 */
>
> so setting max_segs = 0 is equivalent to setting max_segs = 1.

Right.  The default is one MSS per skb, so 0 and 1 behave the same.

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs
  2026-09-28 23:34     ` Willem de Bruijn
  2026-09-29  7:27       ` Paolo Abeni
@ 2026-09-29 11:50       ` Wang Zhan
  1 sibling, 0 replies; 27+ messages in thread
From: Wang Zhan @ 2026-09-29 11:50 UTC (permalink / raw)
  To: netdev, Willem de Bruijn
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun, Wang Zhan,
	Jason Wang, Andrew Lunn, Aaron Conole, Eelco Chaudron,
	Ilya Maximets, dev, Daniel Borkmann, Neal Cardwell,
	Kuniyuki Iwashima, Alice Mikityanska, David Laight

On Mon, 28 Sep 2026 19:34:21 -0400 Willem de Bruijn wrote:
> If 1/5 fixes a bug that is reachable today (does it?) then it needs to
> go to net on its own.

Before this series, it is reachable when the tag is already in the frame,
for example a stacked VLAN device or an OVS QinQ path.  I reproduced it
with the following script:

  ip netns add g1; ip netns add g2
  ip link add v0 type veth peer name v1
  ip link set v0 netns g1; ip link set v1 netns g2
  for n in g1 g2; do
      r=v0; [ $n = g2 ] && r=v1
      ip -n $n link add link $r name vlan200 type vlan id 200
      ip -n $n link add link vlan200 name vlan100 type vlan id 100
      for d in $r vlan200 vlan100; do
          ip -n $n link set $d gso_max_size 524280
      done
      ip -n $n link set $r up; ip -n $n link set vlan200 up
      ip -n $n link set vlan100 up
  done
  ip -n g1 addr add 2001:db8::1/64 dev vlan100
  ip -n g2 addr add 2001:db8::2/64 dev vlan100
  ip netns exec g2 iperf3 -s -D
  ip -n g1 link set v0 gso_ipv4_max_size 65536
  ip netns exec g1 iperf3 -c 2001:db8::2 -t 6
  ip -n g1 link set v0 gso_ipv4_max_size 524280
  ip netns exec g1 iperf3 -c 2001:db8::2 -t 6

With gso_ipv4_max_size = 65536, iperf3 reports 442 Mbps; with 524280, it
reports 39.4 Gbps.  So gso_ipv4_max_size clearly affects this IPv6 flow.

I will keep 1/5 in the series and drop the Fixes tag, as Paolo suggests.

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 1/5] net: core: use the packet's L3 protocol for the GSO size limit
  2026-09-29  4:01     ` Wang Zhan
@ 2026-09-29 14:59       ` Willem de Bruijn
  0 siblings, 0 replies; 27+ messages in thread
From: Willem de Bruijn @ 2026-09-29 14:59 UTC (permalink / raw)
  To: Wang Zhan, netdev, Willem de Bruijn
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun, Wang Zhan,
	Jason Wang, Andrew Lunn, Aaron Conole, Eelco Chaudron,
	Ilya Maximets, dev, Daniel Borkmann, Neal Cardwell,
	Kuniyuki Iwashima, Alice Mikityanska, David Laight

Wang Zhan wrote:
> On Mon, 28 Sep 2026 19:36:01 -0400 Willem de Bruijn wrote:
> > If netif_get_gso_max_size needs to inspect the result from
> > vlan_get_protocol, call it inside that function directly? More robust
> > and a one line change.
> 
> Or we can move it to dev.c, only one caller in-tree (before this series).

Nice. Yeah that works.

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs
  2026-09-29 10:25     ` Wang Zhan
@ 2026-09-29 15:00       ` Willem de Bruijn
  0 siblings, 0 replies; 27+ messages in thread
From: Willem de Bruijn @ 2026-09-29 15:00 UTC (permalink / raw)
  To: Wang Zhan, netdev, Willem de Bruijn
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun, Wang Zhan,
	Jason Wang, Andrew Lunn, Aaron Conole, Eelco Chaudron,
	Ilya Maximets, dev, Daniel Borkmann, Neal Cardwell,
	Kuniyuki Iwashima, Alice Mikityanska, David Laight

Wang Zhan wrote:
> On Mon, 28 Sep 2026 19:47:43 -0400 Willem de Bruijn wrote:
> > > +	/*
> > > +	 * The TCP frag-list path segments through skb_segment_list(), which
> > > +	 * does not carry max_segs, so bounded calls skip those skbs.
> > > +	 */
> >
> > This comment answers only one of six conditions. And one that is
> > pretty straightforward. I'd drop.
> >
> > In general, drop all too-obvious comments. AI has a habit of adding
> > a lot more, and more low information, comments than is customary in
> > kernel code (where we also have commit messages). Generally, repeating
> > what the code does is of little value.
> 
> Dropped in v4.  I will check all the comments in the series.
> 
> > > +	if (!skb_is_gso(skb) || !skb_is_gso_tcp(skb) ||
> > > +	    skb->encapsulation || skb_has_frag_list(skb) ||
> > > +	    !skb_mac_header_was_set(skb) ||
> > > +	    !skb_transport_header_was_set(skb))
> >
> > Conversely, they last two conditions are less obvious. Are they not
> > always true for a TSO packet?
> 
> The transport header can be missing.  qdisc_pkt_len_segs_init() does the
> same check on this path (net/core/dev.c:4245, a0dce8752193e).
> 
> The mac header is always set.  It can be dropped in v4.
> 
> > > +	gso_max_size = netif_get_gso_max_size(dev, vlan_get_protocol(skb));
> >
> > Third time this is now called in validate_xmit_skb. Not sure if that can
> > easily be avoided.
> 
> Maybe we can pass the oversize and gso_max_size flags out of
> __netif_skb_features, but it would be a bit ugly.  I think the current
> cost is acceptable.

Agreed. I was hoping otherwise, but don't see an easy fix (and didn't
explore more deeply to be fair).

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 5/5] net: net_test: add tests for bounded GSO segmentation
  2026-09-29 10:30     ` Wang Zhan
@ 2026-09-29 15:01       ` Willem de Bruijn
  0 siblings, 0 replies; 27+ messages in thread
From: Willem de Bruijn @ 2026-09-29 15:01 UTC (permalink / raw)
  To: Wang Zhan, netdev, Willem de Bruijn
  Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun, Wang Zhan,
	Jason Wang, Andrew Lunn, Aaron Conole, Eelco Chaudron,
	Ilya Maximets, dev, Daniel Borkmann, Neal Cardwell,
	Kuniyuki Iwashima, Alice Mikityanska, David Laight

Wang Zhan wrote:
> On Mon, 28 Sep 2026 19:59:54 -0400 Willem de Bruijn wrote:
> > > +	GSO_TEST_BOUNDED,
> > > +	GSO_TEST_BOUNDED_MULTI,
> > > +	GSO_TEST_BOUNDED_ONE_MSS,
> >
> > nit: bound is not a helpful name for this feature.
> >
> > The default segments to a stream of skbs of MSS 1, so segmenting to
> > a stream of skbs larger MSS to me is not bounding. Quite the opposite.
> >
> > Not a comment only about this test patch.
> >
> > Perhaps partial or re-segmentation better captures it.
> 
> Agreed.  I think re-segmentation is descriptive enough.  Partial is
> already for NETIF_F_GSO_PARTIAL.

Come to think of it, that may have always been the intention of partial?
Segment, but not yet fully.

Technically this would more be pre-segment than resegment.

Anyway, I don't mean to bikeshed this and don't have a better name.
Either will do.
 
> > > +	{
> > > +		/*
> > > +		 * One MSS per skb is what the unbounded path produces, so a
> > > +		 * bound of a single segment must not change the output.
> > > +		 */
> >
> > so setting max_segs = 0 is equivalent to setting max_segs = 1.
> 
> Right.  The default is one MSS per skb, so 0 and 1 behave the same.

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 3/5] net: gso: support bounded TCP segmentation
  2026-09-28  4:41 ` [PATCH net-next v3 3/5] net: gso: support bounded TCP segmentation Wang Zhan
  2026-09-28 23:39   ` Willem de Bruijn
@ 2026-09-30  4:41   ` netdev-bot+sashiko
  1 sibling, 0 replies; 27+ messages in thread
From: netdev-bot+sashiko @ 2026-09-30  4:41 UTC (permalink / raw)
  To: wang.zhan
  Cc: netdev, davem, edumazet, kuba, pabeni, horms, keyong.sun,
	willemdebruijn.kernel, jasowangio, andrew+netdev, aconole,
	echaudro, i.maximets, dev, daniel, ncardwell, kuniyu, alice,
	david.laight.linux

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] The public @max_segs kernel-doc for __skb_gso_segment()
  (net/core/gso.c) says only "must only be set for TCP GSO skbs".
- [Low] Rejected as a false positive.
- [Low] The comment in tcp_gso_segment() (net/ipv4/tcp_offload.c) that
  explains `mss *= skb_shinfo(segs)->gso_segs` still names only two ways…

--

net: gso: support bounded TCP segmentation

This patch adds a max_segs argument to __skb_gso_segment() and stores it
in struct skb_gso_cb, so that skb_segment() can put a bounded number of
MSS segments into each output GSO skb. When max_segs is set, skb_segment()
skips the GSO_PARTIAL/frag_list decision block and caps partial_segs with
the bound. Existing callers pass 0 for no limit.

> diff --git a/net/core/gso.c b/net/core/gso.c
> index bcd156372f4df..42fbce17b508c 100644
> --- a/net/core/gso.c
> +++ b/net/core/gso.c
> @@ -77,6 +77,8 @@ static bool skb_needs_check(const struct sk_buff *skb, bool tx_path)
>   *	@skb: buffer to segment
>   *	@features: features for the output path (see dev->features)
>   *	@tx_path: whether it is called in TX path
> + *	@max_segs: maximum MSS segments per output GSO skb, 0 means no limit;
> + *		   must only be set for TCP GSO skbs
>   *
>   *	This function segments the given skb and returns a list of segments.
>   *

[ ... ]

> @@ -117,6 +120,7 @@ struct sk_buff *__skb_gso_segment(struct sk_buff *skb,
>  
>  	SKB_GSO_CB(skb)->mac_offset = skb_headroom(skb);
>  	SKB_GSO_CB(skb)->encap_level = 0;
> +	SKB_GSO_CB(skb)->max_segs = min(max_segs, GSO_MAX_SEGS);

[Severity: Low]
Should the kernel-doc for @max_segs also say that a bound must not be
passed for an skb with a frag_list?

Right now that rule is stated only in the new comment inside
skb_segment():

    /*
     * A call with max_segs set groups the MSS segments below,
     * and the frag_list split in this block does not carry the
     * bound, so a bound must not be passed for a skb with a
     * frag_list.
     */
    if (!max_segs && !(features & NETIF_F_GSO_PARTIAL)) {

Neither the TCP rule nor the frag_list rule is enforced, here or in
skb_segment(). The bound is stored without a skb_has_frag_list() or
gso_type check. A non-zero max_segs also skips the frag_list uniformity
and head_frag checks.

A future caller could follow the kernel-doc and bound a TCP GSO skb that
carries a frag_list, which is common after GRO. Could that caller then hit
the BUG_ON()s in the frag_list walk of skb_segment()?

This can't happen in the series as posted. The only caller that passes a
non-zero bound is skb_gso_output_max_segs(), added in "net: core:
resegment oversized TCP GSO skbs". It returns 0 for skb_has_frag_list(),
non-TCP and encapsulated skbs. Every other caller passes 0.

Would it be worth adding the frag_list rule to the kernel-doc? Another
option is a check like DEBUG_NET_WARN_ON_ONCE(max_segs && list_skb) in
skb_segment() that falls back to the normal path.

> diff --git a/net/core/skbuff.c b/net/core/skbuff.c
> index 8912a66cd9097..b16a6843f3192 100644
> --- a/net/core/skbuff.c
> +++ b/net/core/skbuff.c

[ ... ]

> @@ -4874,7 +4881,10 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
>  		 * now.
>  		 */
>  		DEBUG_NET_WARN_ON_ONCE(len / mss > GSO_MAX_SEGS);
> -		partial_segs = min(len / mss, GSO_MAX_SEGS);
> +		if (max_segs)
> +			partial_segs = min(len / mss, max_segs);
> +		else
> +			partial_segs = min(len / mss, GSO_MAX_SEGS);

[Severity: Low]
This isn't a bug, but the comment in tcp_gso_segment() in
net/ipv4/tcp_offload.c still lists only two ways that skb_segment()
returns GSO skbs:

    /* GSO partial and frag_list segmentation only requires splitting
     * the frame into an MSS multiple and possibly a remainder, both
     * cases return a GSO skb. So update the mss now.
     */
    if (skb_is_gso(segs))
        mss *= skb_shinfo(segs)->gso_segs;

The bounded grouping added here is a third case, with neither
NETIF_F_GSO_PARTIAL nor a frag_list. The code still works because it
tests skb_is_gso(segs). However, the comment isn't updated in this patch
or later in the series.

Could the comment be updated to include the max_segs case?

>  		if (partial_segs > 1)
>  			mss *= partial_segs;
>  		else

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928044102.1004310-1-wang.zhan%40smartx.com

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs
  2026-09-28  4:41 ` [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs Wang Zhan
  2026-09-28 23:47   ` Willem de Bruijn
@ 2026-09-30  4:41   ` netdev-bot+sashiko
  1 sibling, 0 replies; 27+ messages in thread
From: netdev-bot+sashiko @ 2026-09-30  4:41 UTC (permalink / raw)
  To: wang.zhan
  Cc: netdev, davem, edumazet, kuba, pabeni, horms, keyong.sun,
	willemdebruijn.kernel, jasowangio, andrew+netdev, aconole,
	echaudro, i.maximets, dev, daniel, ncardwell, kuniyu, alice,
	david.laight.linux

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 1 · Medium: 0 · Low: 1

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] skb_gso_output_max_segs() (net/core/dev.c:3954-3960) works out
  hdr_len from a TCP header read at skb->transport_header.
- [Low] Parts of the commit message's cost claims do not match the code.

--

net: core: resegment oversized TCP GSO skbs

When an unencapsulated TCP GSO skb exceeds the egress device's
gso_max_size or gso_max_segs, validate_xmit_skb() now resegments it into
smaller GSO skbs instead of segmenting it into individual packets. A new
helper, skb_gso_output_max_segs(), works out the per-output segment bound
from the header length, the MSS and the device limits. The size limit is
capped at GSO_LEGACY_MAX_SIZE.

> The helper runs on the skb which is handed to the driver, after
> validate_xmit_vlan() and sk_validate_xmit_skb(), and only from the
> netif_needs_gso() branch, so an skb which is not segmented pays nothing.
> An over-limit skb pays one device limit test and one ndo_features_check()
> for the bound, in exchange for keeping the output a GSO skb.

[Severity: Low]
Is this description of the cost accurate?

Take SKB_GSO_DODGY TCP skbs, for example from tap/virtio guests or packet
sockets, sent through a device without NETIF_F_GSO_ROBUST. They fail
skb_gso_ok() and take the netif_needs_gso() branch, so they now pay the
gso_within_dev_limits() test in validate_xmit_skb().

They are then not segmented, because tcp_gso_segment() returns early:

    if (skb_gso_ok(skb, features | NETIF_F_GSO_ROBUST)) {
        /* Packet is from an untrusted source, reset gso_segs. */
        skb_shinfo(skb)->gso_segs = DIV_ROUND_UP(skb->len, mss);
        segs = NULL;
        goto out;
    }

CHECKSUM_NONE GSO skbs on a TSO device go through the same path.

In the over-limit case, the helper calls __netif_skb_features(skb, false).
That reruns gso_features_check(), the VLAN feature intersection,
ndo_features_check() or dflt_features_check(), and harmonize_features().
The helper also does an extra skb_header_pointer() read for the TCP header.

Could the commit message describe this cost more precisely?

> diff --git a/net/core/dev.c b/net/core/dev.c
> index d66b667071837..728260772f349 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -3933,6 +3933,63 @@ netdev_features_t netif_skb_features(struct sk_buff *skb)
>  }
>  EXPORT_SYMBOL(netif_skb_features);
>  
> +static unsigned int
> +skb_gso_output_max_segs(struct sk_buff *skb, struct net_device *dev)
> +{
> +	unsigned int mss = skb_shinfo(skb)->gso_size;
> +	unsigned int gso_max_size, hdr_len, max_segs;
> +	netdev_features_t features;
> +	struct tcphdr _tcph, *th;
> +
> +	/*
> +	 * The TCP frag-list path segments through skb_segment_list(), which
> +	 * does not carry max_segs, so bounded calls skip those skbs.
> +	 */
> +	if (!skb_is_gso(skb) || !skb_is_gso_tcp(skb) ||
> +	    skb->encapsulation || skb_has_frag_list(skb) ||
> +	    !skb_mac_header_was_set(skb) ||
> +	    !skb_transport_header_was_set(skb))
> +		return 0;
> +
> +	th = skb_header_pointer(skb, skb_transport_offset(skb), sizeof(_tcph),
> +				&_tcph);
> +	if (!th || th->doff < sizeof(*th) / 4)
> +		return 0;
> +
> +	hdr_len = skb_transport_header(skb) - skb_mac_header(skb) +
> +		  th->doff * 4;

[Severity: High]
Can hdr_len end up computed from a header that isn't the TCP header?

skb_transport_header_was_set() only says that the offset was set. It does
not say that the offset points at the TCP header.

The segmentation code does not use this offset. inet_gso_segment() and
ipv6_gso_segment() re-parse L3 and reset the transport header themselves,
so skb_segment() builds the outputs from the real header length.

There seem to be two ways a forwarded TCP GSO skb can arrive here with a
stale transport offset.

The first is IPv6 forwarding with a Destination Options or Routing header.
ipv6_gro_receive() sets the transport header to the real TCP header.
ip6_rcv_core() then overwrites it:

    skb->transport_header = skb->network_header + sizeof(*hdr);

Only a Hop-by-Hop header moves it past that point. So on the
ip6_forward()->validate_xmit_skb() path, this helper reads doff from
extension header bytes that the remote sender controls.

The second is VXLAN decap followed by bridge forwarding, with GRO off on
the vxlan device or an XDP prog attached. gro_cells_receive() does:

    skb_unset_transport_header(skb);

    if (!gcells->cells || skb_cloned(skb) || netif_elide_gro(dev)) {
        res = netif_rx(skb);

Without CONFIG_DEBUG_NET, __netif_receive_skb_core() then resets the
transport header to skb->data, which is the inner IP header, and the
bridge leaves it there.

doff then comes from the high nibble of the inner IPv4 saddr. An address
in 80.x to 95.x gives doff = 5, so hdr_len is 34 instead of 54 or more.

In both cases hdr_len comes out too small, so max_segs comes out too large.
Take gso_ipv4_max_size = 32160, MSS 1460 and no TCP options. This helper
returns 22 instead of 21, and skb_segment() emits a 32174 byte GSO skb.
That is above the limit the device advertises.

In the BIG TCP to 64 KiB case, the GSO_LEGACY_MAX_SIZE cap does not bound
the real output either. A Destination Options header of about 48 bytes is
enough in the IPv6 case for the output to go past what the 16-bit length
field can hold. ipv6_gso_segment() would then truncate here:

    payload_len = skb->len - nhoff - sizeof(*ipv6h);
    ipv6h->payload_len = htons(payload_len);

The same applies to iph->tot_len in inet_gso_segment(). The commit message
says the output "obeys the GSO feature and limit contract the device
already advertises". Before this patch, these over-limit skbs were fully
segmented.

Would it be safer to find the TCP header the same way the GSO code does?
One option is to require CHECKSUM_PARTIAL and use
skb_checksum_start_offset(), since tcp_gso_segment() already requires
csum_start to match the TCP header. Another is to return 0 unless
skb_checksum_start(skb) == skb_transport_header(skb).

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928044102.1004310-1-wang.zhan%40smartx.com

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH net-next v3 5/5] net: net_test: add tests for bounded GSO segmentation
  2026-09-28  4:41 ` [PATCH net-next v3 5/5] net: net_test: add tests for bounded GSO segmentation Wang Zhan
  2026-09-28 23:59   ` Willem de Bruijn
@ 2026-09-30  4:41   ` netdev-bot+sashiko
  1 sibling, 0 replies; 27+ messages in thread
From: netdev-bot+sashiko @ 2026-09-30  4:41 UTC (permalink / raw)
  To: wang.zhan
  Cc: netdev, davem, edumazet, kuba, pabeni, horms, keyong.sun,
	willemdebruijn.kernel, jasowangio, andrew+netdev, aconole,
	echaudro, i.maximets, dev, daniel, ncardwell, kuniyu, alice,
	david.laight.linux

Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 3

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] The v3 changelog says headroom is now reserved "so the VLAN
  step needs no atomic allocation".
- [Low] The new bounded tests only check `gso_segs <= max_segs`
  (KUNIT_EXPECT_LE in gso_test_func() and…
- [Low] gso_tcp6_skb_new() allocates exactly GSO_TCP6_HDR_LEN +
  payload_len bytes, calls skb_reserve(skb, NET_SKB_PAD), and then calls…
- [Low] gso_test_tcp_bounded_segment() calls KUNIT_ASSERT_LT(test, i,
  ARRAY_SIZE(expected)) inside the loop that walks the segment list it…

--

net: net_test: add tests for bounded GSO segmentation

This commit adds KUnit coverage for GSO segmentation with a bound on the
number of MSS segments per output skb. It adds three parameterized
skb_segment() cases with max_segs and a TCP __skb_gso_segment() test with a
bound of two MSS. It also adds a test that pushes a 200 KiB TCP skb through
validate_xmit_skb_list(), and a test for the per-L3 GSO size limit that
netif_skb_features() applies.

> diff --git a/net/core/net_test.c b/net/core/net_test.c
> index 9c3a590865d26..a4f61398a90ab 100644
> --- a/net/core/net_test.c
> +++ b/net/core/net_test.c

[ ... ]

> @@ -247,6 +309,17 @@ static void gso_test_func(struct kunit *test)
>  
>  		/* header was copied to all segs */
>  		KUNIT_ASSERT_EQ(test, memcmp(skb_mac_header(cur), hdr, sizeof(hdr)), 0);
> +		if (tcase->segs_are_gso) {
> +			KUNIT_EXPECT_TRUE(test, skb_is_gso(cur));
> +			KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_size,
> +					GSO_TEST_SIZE);
> +			KUNIT_EXPECT_LE(test, skb_shinfo(cur)->gso_segs,
> +					tcase->max_segs);

[Severity: Low]
This isn't a bug, but would it be worth checking the exact gso_segs of each
output here instead of only an upper bound?

skb_segment() sets gso_segs on the grouped outputs:

	skb_shinfo(iter)->gso_segs = partial_segs;

and tcp_gso_segment() uses that value as the per-output sequence stride:

	if (skb_is_gso(segs))
		mss *= skb_shinfo(segs)->gso_segs;

If that accounting regressed and left gso_segs at 0 or 1, these checks would
still pass even though the outputs carried the wrong sequence numbers.
gso_test_tcp_bounded_segment() has the same KUNIT_EXPECT_LE() check, and it
doesn't check th->seq either.

With the current skb_segment() the grouped outputs get 2, and the tail gets
DIV_ROUND_UP(1003, 1000) = 2, so the tests pass today for the right reason.

> +			KUNIT_EXPECT_FALSE(test, skb_shinfo(cur)->gso_type &
> +					 SKB_GSO_PARTIAL);
> +		} else if (tcase->max_segs) {
> +			KUNIT_EXPECT_FALSE(test, skb_is_gso(cur));
> +		}

[ ... ]

> +static struct sk_buff *gso_tcp_skb_new(unsigned int payload_len)
> +{
> +	struct sk_buff *skb;
> +	struct ethhdr *eth;
> +	struct tcphdr *th;
> +	struct iphdr *iph;
> +
> +	skb = alloc_skb(GSO_TCP_HDR_LEN + payload_len, GFP_KERNEL);
> +	if (!skb)
> +		return NULL;
> +	skb_put_zero(skb, GSO_TCP_HDR_LEN + payload_len);

[Severity: Medium]
The v3 changelog says:

  reserve headroom so the VLAN step needs no atomic allocation

Only gso_tcp6_skb_new() calls skb_reserve(skb, NET_SKB_PAD). This IPv4
builder has no reserve, so skb_headroom() is 0.

gso_test_tcp_limit_l3_proto() still passes the 100 KiB IPv4 skb into
gso_test_tcp_l3_limit(), which inserts the tag:

gso_test_tcp_l3_limit()
  vlan_insert_tag_set_proto()
    vlan_insert_tag()
      vlan_insert_inner_tag()
        __vlan_insert_inner_tag()
          skb_cow_head(skb, meta_len + VLAN_HLEN)
            __skb_cow()
              pskb_expand_head(skb, ALIGN(delta, NET_SKB_PAD), 0, GFP_ATOMIC)

The head from alloc_skb(54 + 100 KiB) is already rounded up to 128 KiB.
Adding the new headroom and the shinfo pushes the reallocation past that, so
it becomes a 256 KiB (order-6) GFP_ATOMIC allocation.

If that allocation fails, vlan_insert_inner_tag() frees the skb and returns
NULL, and the test reports "no tagged skb".

Can the IPv4 half of the limit test fail this way for reasons unrelated to
GSO? Should gso_tcp_skb_new() reserve headroom the same way the IPv6 builder
does?

> +
> +	skb_reset_mac_header(skb);

[ ... ]

> +	segs = __skb_gso_segment(skb, features, true, max_segs);
> +	if (IS_ERR_OR_NULL(segs)) {
> +		KUNIT_FAIL(test, "segs error %pe", segs);
> +		consume_skb(skb);
> +		return;
> +	}
> +
> +	for (cur = segs; cur; cur = next, i++) {
> +		next = cur->next;
> +
> +		KUNIT_ASSERT_LT(test, i, ARRAY_SIZE(expected));

[Severity: Low]
Does this assertion leak cur, the rest of the segment list and the original
skb when it fires?

If __skb_gso_segment() returns more outputs than expected, which is the
regression this test exists to catch, KUNIT_ASSERT_LT() aborts the case
through kunit_try_catch_throw(). That skips consume_skb() for cur and every
later segment, and it also skips the final consume_skb(skb). None of these
are registered as KUnit resources.

The v3 changelog says it now frees "the skb and the device on the failure
paths". Could this use KUNIT_EXPECT_LT() and free the rest of the list
before bailing out?

> +		KUNIT_EXPECT_EQ(test, cur->len,
> +				GSO_TCP_HDR_LEN + expected[i]);

[ ... ]

> +static struct sk_buff *gso_tcp6_skb_new(unsigned int payload_len)
> +{
> +	struct ipv6hdr *ip6h;
> +	struct sk_buff *skb;
> +	struct ethhdr *eth;
> +	struct tcphdr *th;
> +
> +	skb = alloc_skb(GSO_TCP6_HDR_LEN + payload_len, GFP_KERNEL);
> +	if (!skb)
> +		return NULL;
> +	skb_reserve(skb, NET_SKB_PAD);
> +	skb_put_zero(skb, GSO_TCP6_HDR_LEN + payload_len);

[Severity: Low]
Is the NET_SKB_PAD reserved here missing from the alloc_skb() size?

skb_reserve() takes its bytes out of the tailroom. The full-length
skb_put_zero() only fits because kmalloc_reserve() rounds the head up with
kmalloc_size_roundup().

With today's only caller (payload_len = 100 KiB) the head rounds up to
128 KiB, which leaves plenty of slack. Some other payload_len could put the
aligned head size within NET_SKB_PAD of a kmalloc bucket or page-order
boundary. skb_put_zero() would then hit skb_over_panic().

Should this be alloc_skb(NET_SKB_PAD + GSO_TCP6_HDR_LEN + payload_len,
GFP_KERNEL)?

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928044102.1004310-1-wang.zhan%40smartx.com

^ permalink raw reply	[flat|nested] 27+ messages in thread

end of thread, other threads:[~2026-09-30  4:41 UTC | newest]

Thread overview: 27+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28  4:40 [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs Wang Zhan
2026-09-28  4:40 ` [PATCH net-next v3 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
2026-09-28 23:36   ` Willem de Bruijn
2026-09-29  3:44     ` Wang Zhan
2026-09-29  4:01     ` Wang Zhan
2026-09-29 14:59       ` Willem de Bruijn
2026-09-28  4:40 ` [PATCH net-next v3 2/5] net: core: factor out the GSO device limit check Wang Zhan
2026-09-28 23:37   ` Willem de Bruijn
2026-09-29  7:47   ` Paolo Abeni
2026-09-28  4:41 ` [PATCH net-next v3 3/5] net: gso: support bounded TCP segmentation Wang Zhan
2026-09-28 23:39   ` Willem de Bruijn
2026-09-30  4:41   ` netdev-bot+sashiko
2026-09-28  4:41 ` [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs Wang Zhan
2026-09-28 23:47   ` Willem de Bruijn
2026-09-29 10:25     ` Wang Zhan
2026-09-29 15:00       ` Willem de Bruijn
2026-09-30  4:41   ` netdev-bot+sashiko
2026-09-28  4:41 ` [PATCH net-next v3 5/5] net: net_test: add tests for bounded GSO segmentation Wang Zhan
2026-09-28 23:59   ` Willem de Bruijn
2026-09-29 10:30     ` Wang Zhan
2026-09-29 15:01       ` Willem de Bruijn
2026-09-30  4:41   ` netdev-bot+sashiko
2026-09-28  4:45 ` [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs netdev-bot+sinfo
2026-09-28  5:49   ` Wang Zhan
2026-09-28 23:34     ` Willem de Bruijn
2026-09-29  7:27       ` Paolo Abeni
2026-09-29 11:50       ` Wang Zhan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox