* [PATCH net-next v4 0/5] net: re-segment oversized TCP GSO skbs
@ 2026-09-30 11:15 Wang Zhan
2026-09-30 11:15 ` [PATCH net-next v4 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
` (4 more replies)
0 siblings, 5 replies; 17+ messages in thread
From: Wang Zhan @ 2026-09-30 11:15 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
Wang Zhan
BIG TCP is negotiated per netdevice, so BIG TCP and non-BIG TCP ports can
coexist in one path. When an skb exceeds the GSO limits of the port it is
sent to, it loses its GSO feature mask and is segmented into individual MSS
sized packets, which that port then sends without TSO.
This series cuts such a TCP GSO skb into GSO skbs which fit the device
limits instead, so the rest of the path keeps using TSO. On a
veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP enabled on the
veth endpoints and left off in the guest, a single iperf3 TCP flow, six
alternating runs per state (-t 15 -O 5, fixed CPU affinity and port tuple):
protocol no BIG TCP mixed, no reseg mixed, resegmented
TCP/IPv4 51.550 Gbps 15.850 Gbps 52.617 Gbps
TCP/IPv6 52.050 Gbps 15.783 Gbps 51.933 Gbps
The middle column is the same tree with the re-segmentation disabled. A
BIG TCP hop which feeds a 64 KiB hop loses 69% of the throughput of a path
which never enables BIG TCP; re-segmentation recovers it.
1/5 moves the GSO size limit lookup into dev.c and makes it follow the
packet's L3 protocol, so where the tag sits does not decide which limit
applies. 2/5 is the preparation which lets the limit tests be skipped for
one caller, and carries no functional change.
The new path is taken only when the skb is an unencapsulated TCP GSO skb
without a frag_list, it exceeds gso_max_size or gso_max_segs, and the
device offloads that GSO type. Everything else keeps today's segmentation.
The output obeys the GSO feature and limit contract the device already
advertises, so this needs no new UAPI, no device state and no driver
change, and it applies automatically. max_segs only says how the output is
grouped; an over-limit skb pays one extra ndo_features_check() in exchange
for staying a GSO skb.
Alternatives considered:
- The caller could set skb_shinfo(skb)->gso_size to ~64K and adjust the
gso bits in the shared info afterwards, which would need
skb_unclone(), a repeat of the grouping logic skb_segment() already
has, and a recomputed IPv4 ID for the DF=0 case.
Patch layout:
[1/5] the GSO size limit follows the packet's L3 protocol
[2/5] factor the device limit check out of gso_features_check()
[3/5] let the GSO engine group several MSS into one output skb
[4/5] re-segment oversized TCP GSO skbs in the TX path
[5/5] KUnit coverage for re-segmentation and the TCP path
---
v4:
- patch 1: drop the Fixes tag
- patch 1: move the limit lookup into dev.c, so it can use vlan_get_protocol()
- patch 2: keep the two limit tests as separate returns
- patch 2: move the wrapper to netdevice.h, export the callee
- patch 3: call the feature re-segmentation rather than bound
- patch 3: drop the comment at the frag_list gate
- patch 3: say in the @max_segs kernel-doc which skbs may set it
- patch 4: drop the comment above the frag_list check
- patch 4: drop the mac header test, it is always set on this path
- patch 4: take the TCP header from the checksum start the GSO engine uses
- patch 4: keep the transport header test, its accessor warns when unset
- patch 4: shorten the comment above the features lookup
- patch 5: name the cases and the descriptions after re-segmentation
- patch 5: reserve the builder headroom the VLAN step needs
- patch 5: check the exact gso_segs of every bounded output
- patch 5: free the segments when a bounded case does not match
v3: https://lore.kernel.org/20260928044102.1004310-1-wang.zhan@smartx.com/
v2: https://lore.kernel.org/20260918084651.3022878-1-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-1-wang.zhan@smartx.com/
Wang Zhan (5):
net: core: use the packet's L3 protocol for the GSO size limit
net: core: factor out the GSO device limit check
net: gso: support re-segmentation of TCP GSO skbs
net: core: re-segment oversized TCP GSO skbs
net: net_test: add tests for TCP re-segmentation
drivers/net/tap.c | 3 +-
include/linux/netdevice.h | 18 +-
include/net/gso.h | 6 +-
include/net/udp.h | 2 +-
net/core/dev.c | 92 ++++++++--
net/core/gso.c | 6 +-
net/core/net_test.c | 351 +++++++++++++++++++++++++++++++++++++
net/core/skbuff.c | 8 +-
net/openvswitch/datapath.c | 2 +-
9 files changed, 459 insertions(+), 29 deletions(-)
base-commit: 47a1446725732cd3996edf607e8739334bbf4d78
--
2.47.3
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH net-next v4 1/5] net: core: use the packet's L3 protocol for the GSO size limit
2026-09-30 11:15 [PATCH net-next v4 0/5] net: re-segment oversized TCP GSO skbs Wang Zhan
@ 2026-09-30 11:15 ` Wang Zhan
2026-09-30 19:40 ` Willem de Bruijn
2026-09-30 11:15 ` [PATCH net-next v4 2/5] net: core: factor out the GSO device limit check Wang Zhan
` (3 subsequent siblings)
4 siblings, 1 reply; 17+ messages in thread
From: Wang Zhan @ 2026-09-30 11:15 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
Wang Zhan
gso_features_check() compares the frame length against
netif_get_gso_max_size(), which picks the IPv4 or the IPv6 limit from
skb->protocol. A tag which is already inside the frame replaces that field
with the VLAN ethertype, as skb_vlan_push() does, and an IPv6 packet is
then measured against the IPv4 limit and segmented although the device
could send it as one TSO frame.
The lookup moves from netdevice.h into dev.c: vlan_get_protocol() reads the
L3 protocol from behind an in-frame tag, while if_vlan.h includes
netdevice.h, so an inline in that header cannot call it. dev.c is the only
file which calls the lookup, so it becomes a static helper there.
Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
---
v4:
- drop the Fixes tag: not stable material
- move the limit lookup into dev.c, so it can use vlan_get_protocol()
v3: https://lore.kernel.org/20260928044102.1004310-2-wang.zhan@smartx.com/
v2: https://lore.kernel.org/20260918084651.3022878-4-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-4-wang.zhan@smartx.com/
---
include/linux/netdevice.h | 9 ---------
net/core/dev.c | 9 +++++++++
2 files changed, 9 insertions(+), 9 deletions(-)
diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index d037faff7c44b6..421c0f5952463e 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -5561,15 +5561,6 @@ netif_get_gro_max_size(const struct net_device *dev, const struct sk_buff *skb)
READ_ONCE(dev->gro_ipv4_max_size);
}
-static inline unsigned int
-netif_get_gso_max_size(const struct net_device *dev, const struct sk_buff *skb)
-{
- /* pairs with WRITE_ONCE() in netif_set_gso(_ipv4)_max_size() */
- return skb->protocol == htons(ETH_P_IPV6) ?
- READ_ONCE(dev->gso_max_size) :
- READ_ONCE(dev->gso_ipv4_max_size);
-}
-
static inline bool netif_is_macsec(const struct net_device *dev)
{
return dev->priv_flags & IFF_MACSEC;
diff --git a/net/core/dev.c b/net/core/dev.c
index a8eb382f40caf3..8475d5da64fdf0 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3869,6 +3869,15 @@ static bool skb_gso_has_extension_hdr(const struct sk_buff *skb)
skb_inner_network_header_len(skb) != sizeof(struct ipv6hdr)));
}
+static unsigned int
+netif_get_gso_max_size(const struct net_device *dev, const struct sk_buff *skb)
+{
+ /* pairs with WRITE_ONCE() in netif_set_gso(_ipv4)_max_size() */
+ return vlan_get_protocol(skb) == htons(ETH_P_IPV6) ?
+ READ_ONCE(dev->gso_max_size) :
+ READ_ONCE(dev->gso_ipv4_max_size);
+}
+
static netdev_features_t gso_features_check(const struct sk_buff *skb,
struct net_device *dev,
netdev_features_t features)
--
2.47.3
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH net-next v4 2/5] net: core: factor out the GSO device limit check
2026-09-30 11:15 [PATCH net-next v4 0/5] net: re-segment oversized TCP GSO skbs Wang Zhan
2026-09-30 11:15 ` [PATCH net-next v4 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
@ 2026-09-30 11:15 ` Wang Zhan
2026-09-30 19:40 ` Willem de Bruijn
2026-09-30 11:15 ` [PATCH net-next v4 3/5] net: gso: support re-segmentation of TCP GSO skbs Wang Zhan
` (2 subsequent siblings)
4 siblings, 1 reply; 17+ messages in thread
From: Wang Zhan @ 2026-09-30 11:15 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
Wang Zhan
gso_features_check() decides whether an egress device can offload a GSO
skb as a single TSO frame by comparing the segment count and the frame
length against the device limits. The resegmentation path added by a
later patch caps the MSS segments per output skb and needs the same
features computed without those limits, so put the two tests in a helper
and gate them with one flag.
__netif_skb_features() takes that flag and takes the place of
netif_skb_features(), which becomes an inline wrapper in netdevice.h and
passes the flag set. Existing callers keep their call, and the
resegmentation path asks for the features without the limit checks.
No functional changes.
Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
---
v4:
- keep the two limit tests as separate returns
- move the wrapper into netdevice.h, export __netif_skb_features()
v3: https://lore.kernel.org/20260928044102.1004310-3-wang.zhan@smartx.com/
v2: https://lore.kernel.org/20260918084651.3022878-2-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-2-wang.zhan@smartx.com/
---
include/linux/netdevice.h | 9 ++++++++-
net/core/dev.c | 30 ++++++++++++++++++++----------
2 files changed, 28 insertions(+), 11 deletions(-)
diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index 421c0f5952463e..5b7b0c87e4f637 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -5495,7 +5495,14 @@ void netif_stacked_transfer_operstate(const struct net_device *rootdev,
netdev_features_t passthru_features_check(struct sk_buff *skb,
struct net_device *dev,
netdev_features_t features);
-netdev_features_t netif_skb_features(struct sk_buff *skb);
+netdev_features_t __netif_skb_features(struct sk_buff *skb,
+ bool check_gso_limits);
+
+static inline netdev_features_t netif_skb_features(struct sk_buff *skb)
+{
+ return __netif_skb_features(skb, true);
+}
+
void skb_warn_bad_offload(const struct sk_buff *skb);
static inline bool net_gso_ok(netdev_features_t features, int gso_type)
diff --git a/net/core/dev.c b/net/core/dev.c
index 8475d5da64fdf0..a6213c9ed5e721 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3878,16 +3878,24 @@ netif_get_gso_max_size(const struct net_device *dev, const struct sk_buff *skb)
READ_ONCE(dev->gso_ipv4_max_size);
}
+static bool
+gso_within_dev_limits(const struct sk_buff *skb, const struct net_device *dev)
+{
+ if (skb_shinfo(skb)->gso_segs > READ_ONCE(dev->gso_max_segs))
+ return false;
+
+ if (unlikely(skb->len >= netif_get_gso_max_size(dev, skb)))
+ return false;
+
+ return true;
+}
+
static netdev_features_t gso_features_check(const struct sk_buff *skb,
struct net_device *dev,
- netdev_features_t features)
+ netdev_features_t features,
+ bool check_limits)
{
- u16 gso_segs = skb_shinfo(skb)->gso_segs;
-
- if (gso_segs > READ_ONCE(dev->gso_max_segs))
- return features & ~NETIF_F_GSO_MASK;
-
- if (unlikely(skb->len >= netif_get_gso_max_size(dev, skb)))
+ if (check_limits && !gso_within_dev_limits(skb, dev))
return features & ~NETIF_F_GSO_MASK;
if (!skb_shinfo(skb)->gso_type) {
@@ -3936,13 +3944,15 @@ static netdev_features_t gso_features_check(const struct sk_buff *skb,
return features;
}
-netdev_features_t netif_skb_features(struct sk_buff *skb)
+netdev_features_t
+__netif_skb_features(struct sk_buff *skb, bool check_gso_limits)
{
struct net_device *dev = skb->dev;
netdev_features_t features = dev->features;
if (skb_is_gso(skb))
- features = gso_features_check(skb, dev, features);
+ features = gso_features_check(skb, dev, features,
+ check_gso_limits);
/* If encapsulation offload request, verify we are testing
* hardware encapsulation features instead of standard
@@ -3965,7 +3975,7 @@ netdev_features_t netif_skb_features(struct sk_buff *skb)
return harmonize_features(skb, features);
}
-EXPORT_SYMBOL(netif_skb_features);
+EXPORT_SYMBOL(__netif_skb_features);
static int xmit_one(struct sk_buff *skb, struct net_device *dev,
struct netdev_queue *txq, bool more)
--
2.47.3
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH net-next v4 3/5] net: gso: support re-segmentation of TCP GSO skbs
2026-09-30 11:15 [PATCH net-next v4 0/5] net: re-segment oversized TCP GSO skbs Wang Zhan
2026-09-30 11:15 ` [PATCH net-next v4 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
2026-09-30 11:15 ` [PATCH net-next v4 2/5] net: core: factor out the GSO device limit check Wang Zhan
@ 2026-09-30 11:15 ` Wang Zhan
2026-09-30 19:52 ` Willem de Bruijn
2026-10-02 11:16 ` netdev-bot+sashiko
2026-09-30 11:15 ` [PATCH net-next v4 4/5] net: core: re-segment oversized " Wang Zhan
2026-09-30 11:15 ` [PATCH net-next v4 5/5] net: net_test: add tests for TCP re-segmentation Wang Zhan
4 siblings, 2 replies; 17+ messages in thread
From: Wang Zhan @ 2026-09-30 11:15 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
Wang Zhan
The re-segmentation added by the next patch splits an oversized TCP GSO skb
into several GSO skbs which fit the device limits. That needs the GSO
engine to group several MSS segments into one output skb, so let callers
set the number of MSS segments each output skb may carry and pass it
through the existing __skb_gso_segment() entry point. Ordinary callers
pass zero for no limit.
Without max_segs, skb_segment() groups several MSS into one output skb only
for a device which advertises NETIF_F_GSO_PARTIAL or for a skb with a
frag_list which can be split into uniform pieces, and falls back to one
segment per skb otherwise. A caller which sets max_segs asks for the
grouping regardless, so the block which makes that decision is skipped.
Callers which pass zero keep it, and the re-segmentation path only runs for
an unencapsulated TCP skb without a frag_list.
The output stays a GSO skb: gso_size is the original MSS and gso_segs is
the number of MSS it holds, so a downstream device can still perform
ordinary TSO. Store max_segs in the existing skb_gso_cb scratch context,
alongside the call-local data_offset and mac_offset fields, so that the
segmentation methods keep their signature. A zero max_segs value means
that no limit is active; it is not a persistent skb flag.
Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
---
v4:
- drop the comment at the frag_list gate
- call the feature re-segmentation rather than bound
- say in the kernel-doc which skbs may set max_segs
v3: https://lore.kernel.org/20260928044102.1004310-4-wang.zhan@smartx.com/
v2: https://lore.kernel.org/20260918084651.3022878-3-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-3-wang.zhan@smartx.com/
---
drivers/net/tap.c | 3 ++-
include/net/gso.h | 6 ++++--
include/net/udp.h | 2 +-
net/core/gso.c | 6 +++++-
net/core/skbuff.c | 8 ++++++--
net/openvswitch/datapath.c | 2 +-
6 files changed, 19 insertions(+), 8 deletions(-)
diff --git a/drivers/net/tap.c b/drivers/net/tap.c
index ff67d99deb39ec..bc111495ebbce5 100644
--- a/drivers/net/tap.c
+++ b/drivers/net/tap.c
@@ -278,9 +278,10 @@ rx_handler_result_t tap_handle_frame(struct sk_buff **pskb)
if (q->flags & IFF_VNET_HDR)
features |= tap->tap_features;
if (netif_needs_gso(skb, features)) {
- struct sk_buff *segs = __skb_gso_segment(skb, features, false);
+ struct sk_buff *segs;
struct sk_buff *next;
+ segs = __skb_gso_segment(skb, features, false, 0);
if (IS_ERR(segs)) {
drop_reason = SKB_DROP_REASON_SKB_GSO_SEG;
goto drop;
diff --git a/include/net/gso.h b/include/net/gso.h
index 29975440cad51e..fccb37889965f5 100644
--- a/include/net/gso.h
+++ b/include/net/gso.h
@@ -19,6 +19,7 @@ struct skb_gso_cb {
int encap_level;
__wsum csum;
__u16 csum_start;
+ __u16 max_segs; /* Max MSS segs per output skb, 0 = no limit */
};
#define SKB_GSO_CB_OFFSET 32
#define SKB_GSO_CB(skb) ((struct skb_gso_cb *)((skb)->cb + SKB_GSO_CB_OFFSET))
@@ -75,12 +76,13 @@ static inline __sum16 gso_make_checksum(struct sk_buff *skb, __wsum res)
}
struct sk_buff *__skb_gso_segment(struct sk_buff *skb,
- netdev_features_t features, bool tx_path);
+ netdev_features_t features, bool tx_path,
+ unsigned int max_segs);
static inline struct sk_buff *skb_gso_segment(struct sk_buff *skb,
netdev_features_t features)
{
- return __skb_gso_segment(skb, features, true);
+ return __skb_gso_segment(skb, features, true, 0);
}
struct sk_buff *skb_eth_gso_segment(struct sk_buff *skb,
diff --git a/include/net/udp.h b/include/net/udp.h
index 1fee17274745f0..5bc25dcf25fbaa 100644
--- a/include/net/udp.h
+++ b/include/net/udp.h
@@ -613,7 +613,7 @@ static inline struct sk_buff *udp_rcv_segment(struct sock *sk,
/* the GSO CB lays after the UDP one, no need to save and restore any
* CB fragment
*/
- segs = __skb_gso_segment(skb, features, false);
+ segs = __skb_gso_segment(skb, features, false, 0);
if (IS_ERR_OR_NULL(segs)) {
drop_count = skb_shinfo(skb)->gso_segs;
goto drop;
diff --git a/net/core/gso.c b/net/core/gso.c
index bcd156372f4df0..21a259ca392f43 100644
--- a/net/core/gso.c
+++ b/net/core/gso.c
@@ -77,6 +77,8 @@ static bool skb_needs_check(const struct sk_buff *skb, bool tx_path)
* @skb: buffer to segment
* @features: features for the output path (see dev->features)
* @tx_path: whether it is called in TX path
+ * @max_segs: maximum MSS segments per output GSO skb, 0 means no limit;
+ * set only for an unencapsulated TCP skb without a frag_list
*
* This function segments the given skb and returns a list of segments.
*
@@ -86,7 +88,8 @@ static bool skb_needs_check(const struct sk_buff *skb, bool tx_path)
* Segmentation preserves SKB_GSO_CB_OFFSET bytes of previous skb cb.
*/
struct sk_buff *__skb_gso_segment(struct sk_buff *skb,
- netdev_features_t features, bool tx_path)
+ netdev_features_t features, bool tx_path,
+ unsigned int max_segs)
{
struct sk_buff *segs;
@@ -117,6 +120,7 @@ struct sk_buff *__skb_gso_segment(struct sk_buff *skb,
SKB_GSO_CB(skb)->mac_offset = skb_headroom(skb);
SKB_GSO_CB(skb)->encap_level = 0;
+ SKB_GSO_CB(skb)->max_segs = min(max_segs, GSO_MAX_SEGS);
skb_reset_mac_header(skb);
skb_reset_mac_len(skb);
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 8912a66cd90972..fd10a7cdc1883a 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -4793,6 +4793,7 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
struct sk_buff *segs = NULL;
struct sk_buff *tail = NULL;
struct sk_buff *list_skb = skb_shinfo(head_skb)->frag_list;
+ unsigned int max_segs = SKB_GSO_CB(head_skb)->max_segs;
unsigned int mss = skb_shinfo(head_skb)->gso_size;
bool gso_by_frags = mss == GSO_BY_FRAGS;
unsigned int doffset = head_skb->data - skb_mac_header(head_skb);
@@ -4839,7 +4840,7 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
csum = !!can_checksum_protocol(features, proto);
if (sg && csum && !gso_by_frags) {
- if (!(features & NETIF_F_GSO_PARTIAL)) {
+ if (!max_segs && !(features & NETIF_F_GSO_PARTIAL)) {
struct sk_buff *iter;
unsigned int frag_len;
@@ -4874,7 +4875,10 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
* now.
*/
DEBUG_NET_WARN_ON_ONCE(len / mss > GSO_MAX_SEGS);
- partial_segs = min(len / mss, GSO_MAX_SEGS);
+ if (max_segs)
+ partial_segs = min(len / mss, max_segs);
+ else
+ partial_segs = min(len / mss, GSO_MAX_SEGS);
if (partial_segs > 1)
mss *= partial_segs;
else
diff --git a/net/openvswitch/datapath.c b/net/openvswitch/datapath.c
index 21870341432552..e793aead68372d 100644
--- a/net/openvswitch/datapath.c
+++ b/net/openvswitch/datapath.c
@@ -375,7 +375,7 @@ static int queue_gso_packets(struct datapath *dp, struct sk_buff *skb,
int err;
BUILD_BUG_ON(sizeof(*OVS_CB(skb)) > SKB_GSO_CB_OFFSET);
- segs = __skb_gso_segment(skb, NETIF_F_SG, false);
+ segs = __skb_gso_segment(skb, NETIF_F_SG, false, 0);
if (IS_ERR(segs))
return PTR_ERR(segs);
if (segs == NULL)
--
2.47.3
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH net-next v4 4/5] net: core: re-segment oversized TCP GSO skbs
2026-09-30 11:15 [PATCH net-next v4 0/5] net: re-segment oversized TCP GSO skbs Wang Zhan
` (2 preceding siblings ...)
2026-09-30 11:15 ` [PATCH net-next v4 3/5] net: gso: support re-segmentation of TCP GSO skbs Wang Zhan
@ 2026-09-30 11:15 ` Wang Zhan
2026-09-30 19:52 ` Willem de Bruijn
2026-10-02 11:16 ` netdev-bot+sashiko
2026-09-30 11:15 ` [PATCH net-next v4 5/5] net: net_test: add tests for TCP re-segmentation Wang Zhan
4 siblings, 2 replies; 17+ messages in thread
From: Wang Zhan @ 2026-09-30 11:15 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
Wang Zhan
A GSO skb which exceeds an egress device limit loses its GSO feature mask
and is segmented into individual packets. This is unnecessarily expensive
when the device can still offload smaller TCP GSO skbs, which is easy to
hit once one hop of a BIG TCP path raises gso_max_size and the next one
does not.
For an unencapsulated TCP GSO skb which exceeds gso_max_size or
gso_max_segs, work out how many MSS segments each output skb may carry and
re-segment the skb with that max_segs instead. The bound is measured from
the transport header, so an skb whose transport header is unset or stale
keeps today's segmentation.
An encapsulated or frag-list skb, a GSO type the device cannot offload and
a max_segs which leaves room for a single MSS all keep today's segmentation
as well. The GSO type test runs on the features without the limit checks,
because gso_features_check() has already cleared the GSO bits of an skb
which exceeds them.
This path emits a plain GSO skb, not a BIG TCP one: inet_gso_segment() and
ipv6_gso_segment() write the whole length of each output into the 16-bit L3
length field. An egress limit above 64 KiB would give outputs whose length
truncates, so the size limit is also capped at what that field can express.
The helper runs on the skb which is handed to the driver, after
validate_xmit_vlan() and sk_validate_xmit_skb(), and only from the
netif_needs_gso() branch: an skb which the device takes as it is pays
nothing. A skb which reaches that branch pays one device limit test, and
an over-limit one pays the TCP header read and the features recomputation
for max_segs, in exchange for keeping the output a GSO skb.
Measured on a veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP
enabled on the veth endpoints and left off in the guest, so the skbs which
the veth hop accepts have to be segmented before the TAP device. A single
iperf3 TCP flow, six alternating runs per state (`-t 15 -O 5`, fixed CPU
affinity and port tuple). The middle column is the same tree with the
re-segmentation disabled:
protocol no BIG TCP mixed, no reseg mixed, resegmented
TCP/IPv4 51.550 Gbps 15.850 Gbps 52.617 Gbps
TCP/IPv6 52.050 Gbps 15.783 Gbps 51.933 Gbps
Coefficient of variation for the two mixed columns was 0.48% and 0.82% for
IPv4 and 0.44% and 0.44% for IPv6. A BIG TCP hop which feeds a 64 KiB hop
loses 69% of the throughput of a path which never enables BIG TCP at all;
re-segmentation recovers it, 3.3x over the existing segmentation path and
within noise of the no BIG TCP baseline.
Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
---
v4:
- drop the comment above the frag_list check
- drop the mac header test, it is always set on this path
- take the TCP header from the checksum start the GSO engine uses
- test the transport header first, skb_transport_header() warns when unset
- shorten the comment above the features lookup
v3: https://lore.kernel.org/20260928044102.1004310-5-wang.zhan@smartx.com/
v2: https://lore.kernel.org/20260918084651.3022878-4-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-4-wang.zhan@smartx.com/
---
net/core/dev.c | 53 +++++++++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 52 insertions(+), 1 deletion(-)
diff --git a/net/core/dev.c b/net/core/dev.c
index a6213c9ed5e721..ff9bce2506f670 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3977,6 +3977,53 @@ __netif_skb_features(struct sk_buff *skb, bool check_gso_limits)
}
EXPORT_SYMBOL(__netif_skb_features);
+static unsigned int
+skb_gso_output_max_segs(struct sk_buff *skb, struct net_device *dev)
+{
+ unsigned int mss = skb_shinfo(skb)->gso_size;
+ unsigned int gso_max_size, hdr_len, max_segs;
+ netdev_features_t features;
+ struct tcphdr _tcph, *th;
+
+ if (!skb_is_gso(skb) || !skb_is_gso_tcp(skb) ||
+ skb->encapsulation || skb_has_frag_list(skb))
+ return 0;
+
+ /*
+ * The transport offset can be unset or stale, so hdr_len is only taken
+ * from a header at the checksum start.
+ */
+ if (!skb_transport_header_was_set(skb) ||
+ skb_checksum_start(skb) != skb_transport_header(skb))
+ return 0;
+
+ th = skb_header_pointer(skb, skb_transport_offset(skb), sizeof(_tcph),
+ &_tcph);
+ if (!th || __tcp_hdrlen(th) < sizeof(*th))
+ return 0;
+
+ hdr_len = skb_transport_header(skb) - skb_mac_header(skb) +
+ __tcp_hdrlen(th);
+
+ /* The output is still a GSO skb: recheck without the limit checks */
+ features = __netif_skb_features(skb, false) | NETIF_F_GSO_ROBUST;
+ if (!net_gso_ok(features, skb_shinfo(skb)->gso_type))
+ return 0;
+
+ gso_max_size = netif_get_gso_max_size(dev, skb);
+ gso_max_size = min(gso_max_size, GSO_LEGACY_MAX_SIZE);
+
+ /*
+ * gso_within_dev_limits() accepts gso_segs == gso_max_segs but
+ * rejects skb->len >= gso_max_size, so only the size limit needs the
+ * - 1; the inner min() keeps that subtraction from wrapping.
+ */
+ max_segs = (gso_max_size - min(gso_max_size, hdr_len + 1)) / mss;
+ max_segs = min(max_segs, READ_ONCE(dev->gso_max_segs));
+
+ return max_segs;
+}
+
static int xmit_one(struct sk_buff *skb, struct net_device *dev,
struct netdev_queue *txq, bool more)
{
@@ -4141,9 +4188,13 @@ static struct sk_buff *validate_xmit_skb(struct sk_buff *skb, struct net_device
goto out_null;
if (netif_needs_gso(skb, features)) {
+ unsigned int max_segs = 0;
struct sk_buff *segs;
- segs = skb_gso_segment(skb, features);
+ if (unlikely(!gso_within_dev_limits(skb, dev)))
+ max_segs = skb_gso_output_max_segs(skb, dev);
+
+ segs = __skb_gso_segment(skb, features, true, max_segs);
if (IS_ERR(segs)) {
goto out_kfree_skb;
} else if (segs) {
--
2.47.3
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH net-next v4 5/5] net: net_test: add tests for TCP re-segmentation
2026-09-30 11:15 [PATCH net-next v4 0/5] net: re-segment oversized TCP GSO skbs Wang Zhan
` (3 preceding siblings ...)
2026-09-30 11:15 ` [PATCH net-next v4 4/5] net: core: re-segment oversized " Wang Zhan
@ 2026-09-30 11:15 ` Wang Zhan
2026-09-30 19:53 ` Willem de Bruijn
2026-10-02 11:16 ` netdev-bot+sashiko
4 siblings, 2 replies; 17+ messages in thread
From: Wang Zhan @ 2026-09-30 11:15 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
Wang Zhan
The GSO engine can now be asked to group several MSS segments into one
output skb, which is what re-segmentation needs. Add KUnit coverage for
the max_segs parameter.
The parameterized GSO test gains a max_segs input and three cases: two
limits which cut the input into two and three output skbs, and a max_segs
of one MSS, which must leave the output unchanged. It drives
skb_segment() directly, because the synthetic protocol it uses has no
gso_segment callback, and stores max_segs in the GSO control block itself.
The TCP test drives __skb_gso_segment() with a max_segs of two MSS and
checks that every output skb stays GSO, keeps its gso_size, and stays
within the limit. The length test runs a 200 KiB TCP skb through
validate_xmit_skb_list(), the caller which sets max_segs, and checks that
the length declared by every output matches the L3 length of that output.
The limit test checks that the GSO size limit which netif_skb_features()
applies follows the packet's L3 protocol, for an IPv4 and an IPv6 skb,
also with the tag inside the frame.
Assisted-by: LLM
Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
---
v4:
- name the cases and the descriptions after re-segmentation
- reserve the headroom the VLAN step needs in both builders
- check the exact gso_segs of each output
- free the segments when a bounded case does not match
v3: https://lore.kernel.org/20260928044102.1004310-6-wang.zhan@smartx.com/
v2: https://lore.kernel.org/20260918084651.3022878-5-wang.zhan@smartx.com/
v1: https://lore.kernel.org/20260917063854.2011613-5-wang.zhan@smartx.com/
---
net/core/net_test.c | 351 ++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 351 insertions(+)
diff --git a/net/core/net_test.c b/net/core/net_test.c
index 9c3a590865d269..6a91db83e53017 100644
--- a/net/core/net_test.c
+++ b/net/core/net_test.c
@@ -4,7 +4,14 @@
/* GSO */
+#include <linux/if_ether.h>
+#include <linux/if_vlan.h>
+#include <linux/ip.h>
+#include <linux/ipv6.h>
+#include <linux/netdevice.h>
#include <linux/skbuff.h>
+#include <linux/tcp.h>
+#include <net/gso.h>
static const char hdr[] = "abcdefgh";
#define GSO_TEST_SIZE 1000
@@ -34,6 +41,9 @@ enum gso_test_nr {
GSO_TEST_FRAG_LIST_PURE,
GSO_TEST_FRAG_LIST_NON_UNIFORM,
GSO_TEST_GSO_BY_FRAGS,
+ GSO_TEST_RESEGMENT,
+ GSO_TEST_RESEGMENT_MULTI,
+ GSO_TEST_RESEGMENT_ONE_MSS,
};
struct gso_test_case {
@@ -46,10 +56,12 @@ struct gso_test_case {
const unsigned int *frags;
unsigned int nr_frag_skbs;
const unsigned int *frag_skbs;
+ unsigned int max_segs;
/* output as expected */
unsigned int nr_segs;
const unsigned int *segs;
+ bool segs_are_gso;
};
static struct gso_test_case cases[] = {
@@ -135,6 +147,51 @@ static struct gso_test_case cases[] = {
.nr_segs = 4,
.segs = (const unsigned int[]) { 100, 200, 300, 400 },
},
+ {
+ .id = GSO_TEST_RESEGMENT,
+ .name = "resegment",
+ .linear_len = GSO_TEST_SIZE,
+ .nr_frags = 3,
+ .frags = (const unsigned int[]) {
+ GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
+ },
+ .max_segs = 2,
+ .nr_segs = 2,
+ .segs = (const unsigned int[]) {
+ 2 * GSO_TEST_SIZE, GSO_TEST_SIZE + 3,
+ },
+ .segs_are_gso = true,
+ },
+ {
+ .id = GSO_TEST_RESEGMENT_MULTI,
+ .name = "resegment_multi",
+ .linear_len = 2 * GSO_TEST_SIZE,
+ .nr_frags = 4,
+ .frags = (const unsigned int[]) {
+ GSO_TEST_SIZE, GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
+ },
+ .max_segs = 2,
+ .nr_segs = 3,
+ .segs = (const unsigned int[]) {
+ 2 * GSO_TEST_SIZE, 2 * GSO_TEST_SIZE, GSO_TEST_SIZE + 3,
+ },
+ .segs_are_gso = true,
+ },
+ {
+ /* A max_segs of one leaves one MSS per skb, like no limit. */
+ .id = GSO_TEST_RESEGMENT_ONE_MSS,
+ .name = "resegment_one_mss",
+ .linear_len = GSO_TEST_SIZE,
+ .nr_frags = 3,
+ .frags = (const unsigned int[]) {
+ GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
+ },
+ .max_segs = 1,
+ .nr_segs = 4,
+ .segs = (const unsigned int[]) {
+ GSO_TEST_SIZE, GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
+ },
+ },
};
static void gso_test_case_to_desc(struct gso_test_case *t, char *desc)
@@ -226,6 +283,7 @@ static void gso_test_func(struct kunit *test)
if (tcase->id == GSO_TEST_FRAG_LIST_NON_UNIFORM)
features &= ~NETIF_F_SG;
+ SKB_GSO_CB(skb)->max_segs = tcase->max_segs;
segs = skb_segment(skb, features);
if (IS_ERR(segs)) {
KUNIT_FAIL(test, "segs error %pe", segs);
@@ -239,6 +297,7 @@ static void gso_test_func(struct kunit *test)
for (cur = segs, i = 0; cur; cur = next, i++) {
next = cur->next;
+ KUNIT_ASSERT_LT(test, i, tcase->nr_segs);
KUNIT_ASSERT_EQ(test, cur->len, sizeof(hdr) + tcase->segs[i]);
/* segs have skb->data pointing to the mac header */
@@ -247,6 +306,18 @@ static void gso_test_func(struct kunit *test)
/* header was copied to all segs */
KUNIT_ASSERT_EQ(test, memcmp(skb_mac_header(cur), hdr, sizeof(hdr)), 0);
+ if (tcase->segs_are_gso) {
+ KUNIT_EXPECT_TRUE(test, skb_is_gso(cur));
+ KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_size,
+ GSO_TEST_SIZE);
+ KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_segs,
+ DIV_ROUND_UP(tcase->segs[i],
+ GSO_TEST_SIZE));
+ KUNIT_EXPECT_FALSE(test, skb_shinfo(cur)->gso_type &
+ SKB_GSO_PARTIAL);
+ } else if (tcase->max_segs) {
+ KUNIT_EXPECT_FALSE(test, skb_is_gso(cur));
+ }
/* last seg can be found through segs->prev pointer */
if (!next)
@@ -261,6 +332,283 @@ static void gso_test_func(struct kunit *test)
consume_skb(skb);
}
+#define GSO_TCP_HDR_LEN \
+ (ETH_HLEN + sizeof(struct iphdr) + sizeof(struct tcphdr))
+
+static struct sk_buff *gso_tcp_skb_new(unsigned int payload_len)
+{
+ struct sk_buff *skb;
+ struct ethhdr *eth;
+ struct tcphdr *th;
+ struct iphdr *iph;
+
+ skb = alloc_skb(NET_SKB_PAD + GSO_TCP_HDR_LEN + payload_len,
+ GFP_KERNEL);
+ if (!skb)
+ return NULL;
+ skb_reserve(skb, NET_SKB_PAD);
+ skb_put_zero(skb, GSO_TCP_HDR_LEN + payload_len);
+
+ skb_reset_mac_header(skb);
+ eth = eth_hdr(skb);
+ eth->h_proto = htons(ETH_P_IP);
+ skb->protocol = eth->h_proto;
+
+ skb_set_network_header(skb, ETH_HLEN);
+ iph = ip_hdr(skb);
+ iph->version = 4;
+ iph->ihl = sizeof(*iph) / 4;
+ iph->protocol = IPPROTO_TCP;
+ iph->tot_len = htons(sizeof(*iph) + sizeof(*th) + payload_len);
+
+ skb_set_transport_header(skb, ETH_HLEN + sizeof(*iph));
+ th = tcp_hdr(skb);
+ th->doff = sizeof(*th) / 4;
+
+ skb->ip_summed = CHECKSUM_PARTIAL;
+ skb->csum_start = skb_transport_header(skb) - skb->head;
+ skb->csum_offset = offsetof(struct tcphdr, check);
+ skb_shinfo(skb)->gso_type = SKB_GSO_TCPV4;
+ skb_shinfo(skb)->gso_size = GSO_TEST_SIZE;
+ skb_shinfo(skb)->gso_segs = DIV_ROUND_UP(payload_len, GSO_TEST_SIZE);
+
+ return skb;
+}
+
+/*
+ * The transmit path passes the features of the device which rejects this
+ * skb, with the GSO bits cleared, so the engine segments it.
+ */
+static void gso_test_tcp_resegment(struct kunit *test)
+{
+ netdev_features_t features = NETIF_F_SG | NETIF_F_HW_CSUM;
+ const unsigned int payload_len = 3 * GSO_TEST_SIZE + 3;
+ struct sk_buff *skb, *segs, *cur, *next;
+ const unsigned int expected[] = {
+ 2 * GSO_TEST_SIZE, GSO_TEST_SIZE + 3,
+ };
+ const unsigned int max_segs = 2;
+ int i = 0;
+
+ if (!IS_ENABLED(CONFIG_INET))
+ kunit_skip(test, "requires CONFIG_INET");
+
+ skb = gso_tcp_skb_new(payload_len);
+ if (!skb) {
+ KUNIT_FAIL(test, "no skb");
+ return;
+ }
+
+ segs = __skb_gso_segment(skb, features, true, max_segs);
+ if (IS_ERR_OR_NULL(segs)) {
+ KUNIT_FAIL(test, "segs error %pe", segs);
+ consume_skb(skb);
+ return;
+ }
+
+ for (cur = segs; cur; cur = next, i++) {
+ next = cur->next;
+
+ if (i < ARRAY_SIZE(expected)) {
+ KUNIT_EXPECT_EQ(test, cur->len,
+ GSO_TCP_HDR_LEN + expected[i]);
+ KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_segs,
+ DIV_ROUND_UP(expected[i],
+ GSO_TEST_SIZE));
+ }
+ KUNIT_EXPECT_TRUE(test, skb_is_gso(cur));
+ KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_size,
+ GSO_TEST_SIZE);
+
+ consume_skb(cur);
+ }
+
+ KUNIT_EXPECT_EQ(test, i, ARRAY_SIZE(expected));
+ consume_skb(skb);
+}
+
+#define GSO_TCP6_HDR_LEN \
+ (ETH_HLEN + sizeof(struct ipv6hdr) + sizeof(struct tcphdr))
+
+static struct sk_buff *gso_tcp6_skb_new(unsigned int payload_len)
+{
+ struct ipv6hdr *ip6h;
+ struct sk_buff *skb;
+ struct ethhdr *eth;
+ struct tcphdr *th;
+
+ skb = alloc_skb(NET_SKB_PAD + GSO_TCP6_HDR_LEN + payload_len,
+ GFP_KERNEL);
+ if (!skb)
+ return NULL;
+ skb_reserve(skb, NET_SKB_PAD);
+ skb_put_zero(skb, GSO_TCP6_HDR_LEN + payload_len);
+
+ skb_reset_mac_header(skb);
+ eth = eth_hdr(skb);
+ eth->h_proto = htons(ETH_P_IPV6);
+ skb->protocol = eth->h_proto;
+
+ skb_set_network_header(skb, ETH_HLEN);
+ ip6h = ipv6_hdr(skb);
+ ip6h->version = 6;
+ ip6h->nexthdr = IPPROTO_TCP;
+ ip6h->payload_len = htons(sizeof(*th) + payload_len);
+
+ skb_set_transport_header(skb, ETH_HLEN + sizeof(*ip6h));
+ th = tcp_hdr(skb);
+ th->doff = sizeof(*th) / 4;
+
+ skb->ip_summed = CHECKSUM_PARTIAL;
+ skb->csum_start = skb_transport_header(skb) - skb->head;
+ skb->csum_offset = offsetof(struct tcphdr, check);
+ skb_shinfo(skb)->gso_type = SKB_GSO_TCPV6;
+ skb_shinfo(skb)->gso_size = GSO_TEST_SIZE;
+ skb_shinfo(skb)->gso_segs = DIV_ROUND_UP(payload_len, GSO_TEST_SIZE);
+
+ return skb;
+}
+
+/*
+ * The device GSO size limit is per L3 protocol, so only a lowered limit of
+ * the packet's own protocol may cost it the TSO bit, and a tag inside the
+ * frame, which replaces skb->protocol with the ethertype, must not change
+ * which limit applies. Takes ownership of @skb.
+ */
+static void gso_test_tcp_l3_limit(struct kunit *test, struct net_device *dev,
+ struct sk_buff *skb, netdev_features_t tso,
+ bool ipv6)
+{
+ unsigned int other = GSO_LEGACY_MAX_SIZE;
+ unsigned int own = GSO_MAX_SIZE;
+
+ /* The skb fits the limit of its own L3 protocol, not the other's. */
+ dev->gso_max_size = ipv6 ? own : other;
+ dev->gso_ipv4_max_size = ipv6 ? other : own;
+ KUNIT_EXPECT_TRUE(test, netif_skb_features(skb) & tso);
+
+ /* ...and the other way around. */
+ swap(own, other);
+ dev->gso_max_size = ipv6 ? own : other;
+ dev->gso_ipv4_max_size = ipv6 ? other : own;
+ KUNIT_EXPECT_FALSE(test, netif_skb_features(skb) & tso);
+
+ skb = vlan_insert_tag_set_proto(skb, htons(ETH_P_8021Q), 0);
+ if (!skb) {
+ KUNIT_FAIL(test, "no tagged skb");
+ return;
+ }
+ KUNIT_EXPECT_TRUE(test, skb->protocol == htons(ETH_P_8021Q));
+
+ swap(own, other);
+ dev->gso_max_size = ipv6 ? own : other;
+ dev->gso_ipv4_max_size = ipv6 ? other : own;
+ KUNIT_EXPECT_TRUE(test, netif_skb_features(skb) & tso);
+
+ swap(own, other);
+ dev->gso_max_size = ipv6 ? own : other;
+ dev->gso_ipv4_max_size = ipv6 ? other : own;
+ KUNIT_EXPECT_FALSE(test, netif_skb_features(skb) & tso);
+
+ consume_skb(skb);
+}
+
+static void gso_test_tcp_limit_l3_proto(struct kunit *test)
+{
+ static const struct net_device_ops dummy_netdev_ops = { };
+ const unsigned int payload_len = 100 * 1024;
+ struct net_device *dev;
+ struct sk_buff *skb;
+
+ dev = alloc_etherdev(0);
+ if (!dev) {
+ KUNIT_FAIL(test, "no net_device");
+ return;
+ }
+ dev->netdev_ops = &dummy_netdev_ops;
+ dev->hw_features = NETIF_F_SG | NETIF_F_HW_CSUM | NETIF_F_TSO |
+ NETIF_F_TSO6;
+ dev->features = dev->hw_features;
+ dev->vlan_features = dev->hw_features;
+
+ skb = gso_tcp6_skb_new(payload_len);
+ if (!skb) {
+ KUNIT_FAIL(test, "no IPv6 skb");
+ goto free_dev;
+ }
+ skb->dev = dev;
+ gso_test_tcp_l3_limit(test, dev, skb, NETIF_F_TSO6, true);
+
+ skb = gso_tcp_skb_new(payload_len);
+ if (!skb) {
+ KUNIT_FAIL(test, "no IPv4 skb");
+ goto free_dev;
+ }
+ skb->dev = dev;
+ gso_test_tcp_l3_limit(test, dev, skb, NETIF_F_TSO, false);
+
+free_dev:
+ free_netdev(dev);
+}
+
+/*
+ * A re-segmented output is a plain GSO skb, so the whole length lands in the
+ * 16-bit L3 length field. A device which accepts GSO skbs above 64 KiB would
+ * otherwise be handed outputs whose length field truncates.
+ */
+static void gso_test_tcp_resegment_l3_len(struct kunit *test)
+{
+ static const struct net_device_ops dummy_netdev_ops = { };
+ const unsigned int payload_len = 200 * 1024;
+ struct sk_buff *skb, *segs, *cur, *next;
+ unsigned int nr_segs = 0;
+ struct net_device *dev;
+ bool again = false;
+
+ if (!IS_ENABLED(CONFIG_INET))
+ kunit_skip(test, "requires CONFIG_INET");
+
+ dev = alloc_etherdev(0);
+ if (!dev) {
+ KUNIT_FAIL(test, "no net_device");
+ return;
+ }
+ dev->netdev_ops = &dummy_netdev_ops;
+ dev->hw_features = NETIF_F_SG | NETIF_F_HW_CSUM | NETIF_F_TSO;
+ dev->features = dev->hw_features;
+ dev->vlan_features = dev->hw_features;
+ dev->gso_ipv4_max_size = 120 * 1024;
+
+ skb = gso_tcp_skb_new(payload_len);
+ if (!skb) {
+ KUNIT_FAIL(test, "no skb");
+ goto free_dev;
+ }
+ skb->dev = dev;
+
+ segs = validate_xmit_skb_list(skb, dev, &again);
+ if (IS_ERR_OR_NULL(segs)) {
+ KUNIT_FAIL(test, "segs error %pe", segs);
+ goto free_dev;
+ }
+
+ for (cur = segs; cur; cur = next, nr_segs++) {
+ next = cur->next;
+ cur->next = NULL;
+
+ KUNIT_EXPECT_TRUE(test, skb_is_gso(cur));
+ KUNIT_EXPECT_EQ(test, ntohs(ip_hdr(cur)->tot_len),
+ cur->len - (skb_network_header(cur) -
+ skb_mac_header(cur)));
+ consume_skb(cur);
+ }
+
+ KUNIT_EXPECT_EQ(test, nr_segs, 4);
+
+free_dev:
+ free_netdev(dev);
+}
+
/* IP tunnel flags */
#include <net/ip_tunnels.h>
@@ -372,6 +720,9 @@ static void ip_tunnel_flags_test_run(struct kunit *test)
static struct kunit_case net_test_cases[] = {
KUNIT_CASE_PARAM(gso_test_func, gso_test_gen_params),
+ KUNIT_CASE(gso_test_tcp_resegment),
+ KUNIT_CASE(gso_test_tcp_limit_l3_proto),
+ KUNIT_CASE(gso_test_tcp_resegment_l3_len),
KUNIT_CASE_PARAM(ip_tunnel_flags_test_run,
ip_tunnel_flags_test_gen_params),
{ },
--
2.47.3
^ permalink raw reply related [flat|nested] 17+ messages in thread
* Re: [PATCH net-next v4 1/5] net: core: use the packet's L3 protocol for the GSO size limit
2026-09-30 11:15 ` [PATCH net-next v4 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
@ 2026-09-30 19:40 ` Willem de Bruijn
0 siblings, 0 replies; 17+ messages in thread
From: Willem de Bruijn @ 2026-09-30 19:40 UTC (permalink / raw)
To: Wang Zhan, netdev
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
Wang Zhan
Wang Zhan wrote:
> gso_features_check() compares the frame length against
> netif_get_gso_max_size(), which picks the IPv4 or the IPv6 limit from
> skb->protocol. A tag which is already inside the frame replaces that field
> with the VLAN ethertype, as skb_vlan_push() does, and an IPv6 packet is
> then measured against the IPv4 limit and segmented although the device
> could send it as one TSO frame.
>
> The lookup moves from netdevice.h into dev.c: vlan_get_protocol() reads the
> L3 protocol from behind an in-frame tag, while if_vlan.h includes
> netdevice.h, so an inline in that header cannot call it. dev.c is the only
> file which calls the lookup, so it becomes a static helper there.
>
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
>
> ---
> v4:
> - drop the Fixes tag: not stable material
> - move the limit lookup into dev.c, so it can use vlan_get_protocol()
> v3: https://lore.kernel.org/20260928044102.1004310-2-wang.zhan@smartx.com/
> v2: https://lore.kernel.org/20260918084651.3022878-4-wang.zhan@smartx.com/
> v1: https://lore.kernel.org/20260917063854.2011613-4-wang.zhan@smartx.com/
Reviewed-by: Willem de Bruijn <willemb@google.com>
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH net-next v4 2/5] net: core: factor out the GSO device limit check
2026-09-30 11:15 ` [PATCH net-next v4 2/5] net: core: factor out the GSO device limit check Wang Zhan
@ 2026-09-30 19:40 ` Willem de Bruijn
0 siblings, 0 replies; 17+ messages in thread
From: Willem de Bruijn @ 2026-09-30 19:40 UTC (permalink / raw)
To: Wang Zhan, netdev
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
Wang Zhan
Wang Zhan wrote:
> gso_features_check() decides whether an egress device can offload a GSO
> skb as a single TSO frame by comparing the segment count and the frame
> length against the device limits. The resegmentation path added by a
> later patch caps the MSS segments per output skb and needs the same
> features computed without those limits, so put the two tests in a helper
> and gate them with one flag.
>
> __netif_skb_features() takes that flag and takes the place of
> netif_skb_features(), which becomes an inline wrapper in netdevice.h and
> passes the flag set. Existing callers keep their call, and the
> resegmentation path asks for the features without the limit checks.
>
> No functional changes.
>
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
>
> ---
> v4:
> - keep the two limit tests as separate returns
> - move the wrapper into netdevice.h, export __netif_skb_features()
> v3: https://lore.kernel.org/20260928044102.1004310-3-wang.zhan@smartx.com/
> v2: https://lore.kernel.org/20260918084651.3022878-2-wang.zhan@smartx.com/
> v1: https://lore.kernel.org/20260917063854.2011613-2-wang.zhan@smartx.com/
Reviewed-by: Willem de Bruijn <willemb@google.com>
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH net-next v4 3/5] net: gso: support re-segmentation of TCP GSO skbs
2026-09-30 11:15 ` [PATCH net-next v4 3/5] net: gso: support re-segmentation of TCP GSO skbs Wang Zhan
@ 2026-09-30 19:52 ` Willem de Bruijn
2026-10-02 16:51 ` Wang Zhan
2026-10-02 11:16 ` netdev-bot+sashiko
1 sibling, 1 reply; 17+ messages in thread
From: Willem de Bruijn @ 2026-09-30 19:52 UTC (permalink / raw)
To: Wang Zhan, netdev
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
Wang Zhan
Wang Zhan wrote:
> The re-segmentation added by the next patch splits an oversized TCP GSO skb
> into several GSO skbs which fit the device limits. That needs the GSO
> engine to group several MSS segments into one output skb, so let callers
> set the number of MSS segments each output skb may carry and pass it
> through the existing __skb_gso_segment() entry point. Ordinary callers
> pass zero for no limit.
>
> Without max_segs, skb_segment() groups several MSS into one output skb only
> for a device which advertises NETIF_F_GSO_PARTIAL or for a skb with a
> frag_list which can be split into uniform pieces, and falls back to one
> segment per skb otherwise. A caller which sets max_segs asks for the
> grouping regardless, so the block which makes that decision is skipped.
> Callers which pass zero keep it, and the re-segmentation path only runs for
> an unencapsulated TCP skb without a frag_list.
>
> The output stays a GSO skb: gso_size is the original MSS and gso_segs is
> the number of MSS it holds, so a downstream device can still perform
> ordinary TSO. Store max_segs in the existing skb_gso_cb scratch context,
> alongside the call-local data_offset and mac_offset fields, so that the
> segmentation methods keep their signature. A zero max_segs value means
> that no limit is active; it is not a persistent skb flag.
>
> Assisted-by: LLM
Reviewed-by: Willem de Bruijn <willemb@google.com>
If respinning, consider asking the LLM to rewrite commit messages to be
more concise.
> @@ -4839,7 +4840,7 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
> csum = !!can_checksum_protocol(features, proto);
>
> if (sg && csum && !gso_by_frags) {
> - if (!(features & NETIF_F_GSO_PARTIAL)) {
> + if (!max_segs && !(features & NETIF_F_GSO_PARTIAL)) {
> struct sk_buff *iter;
> unsigned int frag_len;
>
> @@ -4874,7 +4875,10 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
> * now.
> */
> DEBUG_NET_WARN_ON_ONCE(len / mss > GSO_MAX_SEGS);
> - partial_segs = min(len / mss, GSO_MAX_SEGS);
> + if (max_segs)
> + partial_segs = min(len / mss, max_segs);
> + else
> + partial_segs = min(len / mss, GSO_MAX_SEGS);
> if (partial_segs > 1)
> mss *= partial_segs;
> else
Not new with this series, nor high priority, so only if respinning:
Consider adding
@@ -5121,7 +5121,7 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
}
if (tail->len - doffset <= gso_size)
- skb_shinfo(tail)->gso_size = 0;
+ skb_gso_reset(tail);
else if (tail != segs)
Because packets looped into the Rx path will sometimes have their
gso_segs used without first checking gso_size / skb_is_gso. For instance
in ip_rcv_core __IP_ADD_STATS.
Additionally for patch 5/5, consider one testcase that resegments, but for
which the last skb is not a GSO skb.
> diff --git a/net/openvswitch/datapath.c b/net/openvswitch/datapath.c
> index 21870341432552..e793aead68372d 100644
> --- a/net/openvswitch/datapath.c
> +++ b/net/openvswitch/datapath.c
> @@ -375,7 +375,7 @@ static int queue_gso_packets(struct datapath *dp, struct sk_buff *skb,
> int err;
>
> BUILD_BUG_ON(sizeof(*OVS_CB(skb)) > SKB_GSO_CB_OFFSET);
> - segs = __skb_gso_segment(skb, NETIF_F_SG, false);
> + segs = __skb_gso_segment(skb, NETIF_F_SG, false, 0);
> if (IS_ERR(segs))
> return PTR_ERR(segs);
> if (segs == NULL)
> --
> 2.47.3
>
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH net-next v4 4/5] net: core: re-segment oversized TCP GSO skbs
2026-09-30 11:15 ` [PATCH net-next v4 4/5] net: core: re-segment oversized " Wang Zhan
@ 2026-09-30 19:52 ` Willem de Bruijn
2026-10-02 11:16 ` netdev-bot+sashiko
1 sibling, 0 replies; 17+ messages in thread
From: Willem de Bruijn @ 2026-09-30 19:52 UTC (permalink / raw)
To: Wang Zhan, netdev
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
Wang Zhan
Wang Zhan wrote:
> A GSO skb which exceeds an egress device limit loses its GSO feature mask
> and is segmented into individual packets. This is unnecessarily expensive
> when the device can still offload smaller TCP GSO skbs, which is easy to
> hit once one hop of a BIG TCP path raises gso_max_size and the next one
> does not.
>
> For an unencapsulated TCP GSO skb which exceeds gso_max_size or
> gso_max_segs, work out how many MSS segments each output skb may carry and
> re-segment the skb with that max_segs instead. The bound is measured from
> the transport header, so an skb whose transport header is unset or stale
> keeps today's segmentation.
>
> An encapsulated or frag-list skb, a GSO type the device cannot offload and
> a max_segs which leaves room for a single MSS all keep today's segmentation
> as well. The GSO type test runs on the features without the limit checks,
> because gso_features_check() has already cleared the GSO bits of an skb
> which exceeds them.
>
> This path emits a plain GSO skb, not a BIG TCP one: inet_gso_segment() and
> ipv6_gso_segment() write the whole length of each output into the 16-bit L3
> length field. An egress limit above 64 KiB would give outputs whose length
> truncates, so the size limit is also capped at what that field can express.
>
> The helper runs on the skb which is handed to the driver, after
> validate_xmit_vlan() and sk_validate_xmit_skb(), and only from the
> netif_needs_gso() branch: an skb which the device takes as it is pays
> nothing. A skb which reaches that branch pays one device limit test, and
> an over-limit one pays the TCP header read and the features recomputation
> for max_segs, in exchange for keeping the output a GSO skb.
>
> Measured on a veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP
> enabled on the veth endpoints and left off in the guest, so the skbs which
> the veth hop accepts have to be segmented before the TAP device. A single
> iperf3 TCP flow, six alternating runs per state (`-t 15 -O 5`, fixed CPU
> affinity and port tuple). The middle column is the same tree with the
> re-segmentation disabled:
>
> protocol no BIG TCP mixed, no reseg mixed, resegmented
> TCP/IPv4 51.550 Gbps 15.850 Gbps 52.617 Gbps
> TCP/IPv6 52.050 Gbps 15.783 Gbps 51.933 Gbps
>
> Coefficient of variation for the two mixed columns was 0.48% and 0.82% for
> IPv4 and 0.44% and 0.44% for IPv6. A BIG TCP hop which feeds a 64 KiB hop
> loses 69% of the throughput of a path which never enables BIG TCP at all;
> re-segmentation recovers it, 3.3x over the existing segmentation path and
> within noise of the no BIG TCP baseline.
>
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
>
> ---
> v4:
> - drop the comment above the frag_list check
> - drop the mac header test, it is always set on this path
> - take the TCP header from the checksum start the GSO engine uses
> - test the transport header first, skb_transport_header() warns when unset
> - shorten the comment above the features lookup
> v3: https://lore.kernel.org/20260928044102.1004310-5-wang.zhan@smartx.com/
> v2: https://lore.kernel.org/20260918084651.3022878-4-wang.zhan@smartx.com/
> v1: https://lore.kernel.org/20260917063854.2011613-4-wang.zhan@smartx.com/
Reviewed-by: Willem de Bruijn <willemb@google.com>
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH net-next v4 5/5] net: net_test: add tests for TCP re-segmentation
2026-09-30 11:15 ` [PATCH net-next v4 5/5] net: net_test: add tests for TCP re-segmentation Wang Zhan
@ 2026-09-30 19:53 ` Willem de Bruijn
2026-10-02 17:59 ` Wang Zhan
2026-10-02 11:16 ` netdev-bot+sashiko
1 sibling, 1 reply; 17+ messages in thread
From: Willem de Bruijn @ 2026-09-30 19:53 UTC (permalink / raw)
To: Wang Zhan, netdev
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun,
Willem de Bruijn, Jason Wang, Andrew Lunn, Aaron Conole,
Eelco Chaudron, Ilya Maximets, dev, Daniel Borkmann,
Neal Cardwell, Kuniyuki Iwashima, Alice Mikityanska, David Laight,
Wang Zhan
Wang Zhan wrote:
> The GSO engine can now be asked to group several MSS segments into one
> output skb, which is what re-segmentation needs. Add KUnit coverage for
> the max_segs parameter.
>
> The parameterized GSO test gains a max_segs input and three cases: two
> limits which cut the input into two and three output skbs, and a max_segs
> of one MSS, which must leave the output unchanged. It drives
> skb_segment() directly, because the synthetic protocol it uses has no
> gso_segment callback, and stores max_segs in the GSO control block itself.
>
> The TCP test drives __skb_gso_segment() with a max_segs of two MSS and
> checks that every output skb stays GSO, keeps its gso_size, and stays
> within the limit. The length test runs a 200 KiB TCP skb through
> validate_xmit_skb_list(), the caller which sets max_segs, and checks that
> the length declared by every output matches the L3 length of that output.
> The limit test checks that the GSO size limit which netif_skb_features()
> applies follows the packet's L3 protocol, for an IPv4 and an IPv6 skb,
> also with the tag inside the frame.
>
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
>
> ---
> v4:
> - name the cases and the descriptions after re-segmentation
> - reserve the headroom the VLAN step needs in both builders
> - check the exact gso_segs of each output
> - free the segments when a bounded case does not match
> v3: https://lore.kernel.org/20260928044102.1004310-6-wang.zhan@smartx.com/
> v2: https://lore.kernel.org/20260918084651.3022878-5-wang.zhan@smartx.com/
> v1: https://lore.kernel.org/20260917063854.2011613-5-wang.zhan@smartx.com/
> ---
> net/core/net_test.c | 351 ++++++++++++++++++++++++++++++++++++++++++++
> 1 file changed, 351 insertions(+)
>
> diff --git a/net/core/net_test.c b/net/core/net_test.c
> index 9c3a590865d269..6a91db83e53017 100644
> --- a/net/core/net_test.c
> +++ b/net/core/net_test.c
> @@ -4,7 +4,14 @@
>
> /* GSO */
>
> +#include <linux/if_ether.h>
> +#include <linux/if_vlan.h>
> +#include <linux/ip.h>
> +#include <linux/ipv6.h>
> +#include <linux/netdevice.h>
> #include <linux/skbuff.h>
> +#include <linux/tcp.h>
> +#include <net/gso.h>
>
> static const char hdr[] = "abcdefgh";
> #define GSO_TEST_SIZE 1000
> @@ -34,6 +41,9 @@ enum gso_test_nr {
> GSO_TEST_FRAG_LIST_PURE,
> GSO_TEST_FRAG_LIST_NON_UNIFORM,
> GSO_TEST_GSO_BY_FRAGS,
> + GSO_TEST_RESEGMENT,
> + GSO_TEST_RESEGMENT_MULTI,
> + GSO_TEST_RESEGMENT_ONE_MSS,
> };
>
> struct gso_test_case {
> @@ -46,10 +56,12 @@ struct gso_test_case {
> const unsigned int *frags;
> unsigned int nr_frag_skbs;
> const unsigned int *frag_skbs;
> + unsigned int max_segs;
>
> /* output as expected */
> unsigned int nr_segs;
> const unsigned int *segs;
> + bool segs_are_gso;
> };
>
> static struct gso_test_case cases[] = {
> @@ -135,6 +147,51 @@ static struct gso_test_case cases[] = {
> .nr_segs = 4,
> .segs = (const unsigned int[]) { 100, 200, 300, 400 },
> },
> + {
> + .id = GSO_TEST_RESEGMENT,
> + .name = "resegment",
> + .linear_len = GSO_TEST_SIZE,
> + .nr_frags = 3,
> + .frags = (const unsigned int[]) {
> + GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
> + },
> + .max_segs = 2,
> + .nr_segs = 2,
> + .segs = (const unsigned int[]) {
> + 2 * GSO_TEST_SIZE, GSO_TEST_SIZE + 3,
> + },
> + .segs_are_gso = true,
> + },
> + {
> + .id = GSO_TEST_RESEGMENT_MULTI,
> + .name = "resegment_multi",
> + .linear_len = 2 * GSO_TEST_SIZE,
> + .nr_frags = 4,
> + .frags = (const unsigned int[]) {
> + GSO_TEST_SIZE, GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
> + },
> + .max_segs = 2,
> + .nr_segs = 3,
> + .segs = (const unsigned int[]) {
> + 2 * GSO_TEST_SIZE, 2 * GSO_TEST_SIZE, GSO_TEST_SIZE + 3,
> + },
> + .segs_are_gso = true,
> + },
> + {
> + /* A max_segs of one leaves one MSS per skb, like no limit. */
> + .id = GSO_TEST_RESEGMENT_ONE_MSS,
> + .name = "resegment_one_mss",
> + .linear_len = GSO_TEST_SIZE,
> + .nr_frags = 3,
> + .frags = (const unsigned int[]) {
> + GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
> + },
> + .max_segs = 1,
> + .nr_segs = 4,
> + .segs = (const unsigned int[]) {
> + GSO_TEST_SIZE, GSO_TEST_SIZE, GSO_TEST_SIZE, 3,
> + },
> + },
> };
Given that we have the above (concise) testcases:
Do we need to add the below verbose code? What does it add, just the
TCP specific path?
> static void gso_test_case_to_desc(struct gso_test_case *t, char *desc)
> @@ -226,6 +283,7 @@ static void gso_test_func(struct kunit *test)
> if (tcase->id == GSO_TEST_FRAG_LIST_NON_UNIFORM)
> features &= ~NETIF_F_SG;
>
> + SKB_GSO_CB(skb)->max_segs = tcase->max_segs;
> segs = skb_segment(skb, features);
> if (IS_ERR(segs)) {
> KUNIT_FAIL(test, "segs error %pe", segs);
> @@ -239,6 +297,7 @@ static void gso_test_func(struct kunit *test)
> for (cur = segs, i = 0; cur; cur = next, i++) {
> next = cur->next;
>
> + KUNIT_ASSERT_LT(test, i, tcase->nr_segs);
> KUNIT_ASSERT_EQ(test, cur->len, sizeof(hdr) + tcase->segs[i]);
>
> /* segs have skb->data pointing to the mac header */
> @@ -247,6 +306,18 @@ static void gso_test_func(struct kunit *test)
>
> /* header was copied to all segs */
> KUNIT_ASSERT_EQ(test, memcmp(skb_mac_header(cur), hdr, sizeof(hdr)), 0);
> + if (tcase->segs_are_gso) {
> + KUNIT_EXPECT_TRUE(test, skb_is_gso(cur));
> + KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_size,
> + GSO_TEST_SIZE);
> + KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_segs,
> + DIV_ROUND_UP(tcase->segs[i],
> + GSO_TEST_SIZE));
> + KUNIT_EXPECT_FALSE(test, skb_shinfo(cur)->gso_type &
> + SKB_GSO_PARTIAL);
> + } else if (tcase->max_segs) {
> + KUNIT_EXPECT_FALSE(test, skb_is_gso(cur));
> + }
>
> /* last seg can be found through segs->prev pointer */
> if (!next)
> @@ -261,6 +332,283 @@ static void gso_test_func(struct kunit *test)
> consume_skb(skb);
> }
>
> +#define GSO_TCP_HDR_LEN \
> + (ETH_HLEN + sizeof(struct iphdr) + sizeof(struct tcphdr))
> +
> +static struct sk_buff *gso_tcp_skb_new(unsigned int payload_len)
> +{
> + struct sk_buff *skb;
> + struct ethhdr *eth;
> + struct tcphdr *th;
> + struct iphdr *iph;
> +
> + skb = alloc_skb(NET_SKB_PAD + GSO_TCP_HDR_LEN + payload_len,
> + GFP_KERNEL);
> + if (!skb)
> + return NULL;
> + skb_reserve(skb, NET_SKB_PAD);
> + skb_put_zero(skb, GSO_TCP_HDR_LEN + payload_len);
> +
> + skb_reset_mac_header(skb);
> + eth = eth_hdr(skb);
> + eth->h_proto = htons(ETH_P_IP);
> + skb->protocol = eth->h_proto;
> +
> + skb_set_network_header(skb, ETH_HLEN);
> + iph = ip_hdr(skb);
> + iph->version = 4;
> + iph->ihl = sizeof(*iph) / 4;
> + iph->protocol = IPPROTO_TCP;
> + iph->tot_len = htons(sizeof(*iph) + sizeof(*th) + payload_len);
> +
> + skb_set_transport_header(skb, ETH_HLEN + sizeof(*iph));
> + th = tcp_hdr(skb);
> + th->doff = sizeof(*th) / 4;
> +
> + skb->ip_summed = CHECKSUM_PARTIAL;
> + skb->csum_start = skb_transport_header(skb) - skb->head;
> + skb->csum_offset = offsetof(struct tcphdr, check);
> + skb_shinfo(skb)->gso_type = SKB_GSO_TCPV4;
> + skb_shinfo(skb)->gso_size = GSO_TEST_SIZE;
> + skb_shinfo(skb)->gso_segs = DIV_ROUND_UP(payload_len, GSO_TEST_SIZE);
> +
> + return skb;
> +}
> +
> +/*
> + * The transmit path passes the features of the device which rejects this
> + * skb, with the GSO bits cleared, so the engine segments it.
> + */
> +static void gso_test_tcp_resegment(struct kunit *test)
> +{
> + netdev_features_t features = NETIF_F_SG | NETIF_F_HW_CSUM;
> + const unsigned int payload_len = 3 * GSO_TEST_SIZE + 3;
> + struct sk_buff *skb, *segs, *cur, *next;
> + const unsigned int expected[] = {
> + 2 * GSO_TEST_SIZE, GSO_TEST_SIZE + 3,
> + };
> + const unsigned int max_segs = 2;
> + int i = 0;
> +
> + if (!IS_ENABLED(CONFIG_INET))
> + kunit_skip(test, "requires CONFIG_INET");
> +
> + skb = gso_tcp_skb_new(payload_len);
> + if (!skb) {
> + KUNIT_FAIL(test, "no skb");
> + return;
> + }
> +
> + segs = __skb_gso_segment(skb, features, true, max_segs);
> + if (IS_ERR_OR_NULL(segs)) {
> + KUNIT_FAIL(test, "segs error %pe", segs);
> + consume_skb(skb);
> + return;
> + }
> +
> + for (cur = segs; cur; cur = next, i++) {
> + next = cur->next;
> +
> + if (i < ARRAY_SIZE(expected)) {
> + KUNIT_EXPECT_EQ(test, cur->len,
> + GSO_TCP_HDR_LEN + expected[i]);
> + KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_segs,
> + DIV_ROUND_UP(expected[i],
> + GSO_TEST_SIZE));
> + }
> + KUNIT_EXPECT_TRUE(test, skb_is_gso(cur));
> + KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_size,
> + GSO_TEST_SIZE);
> +
> + consume_skb(cur);
> + }
> +
> + KUNIT_EXPECT_EQ(test, i, ARRAY_SIZE(expected));
> + consume_skb(skb);
> +}
> +
> +#define GSO_TCP6_HDR_LEN \
> + (ETH_HLEN + sizeof(struct ipv6hdr) + sizeof(struct tcphdr))
> +
> +static struct sk_buff *gso_tcp6_skb_new(unsigned int payload_len)
> +{
> + struct ipv6hdr *ip6h;
> + struct sk_buff *skb;
> + struct ethhdr *eth;
> + struct tcphdr *th;
> +
> + skb = alloc_skb(NET_SKB_PAD + GSO_TCP6_HDR_LEN + payload_len,
> + GFP_KERNEL);
> + if (!skb)
> + return NULL;
> + skb_reserve(skb, NET_SKB_PAD);
> + skb_put_zero(skb, GSO_TCP6_HDR_LEN + payload_len);
> +
> + skb_reset_mac_header(skb);
> + eth = eth_hdr(skb);
> + eth->h_proto = htons(ETH_P_IPV6);
> + skb->protocol = eth->h_proto;
> +
> + skb_set_network_header(skb, ETH_HLEN);
> + ip6h = ipv6_hdr(skb);
> + ip6h->version = 6;
> + ip6h->nexthdr = IPPROTO_TCP;
> + ip6h->payload_len = htons(sizeof(*th) + payload_len);
> +
> + skb_set_transport_header(skb, ETH_HLEN + sizeof(*ip6h));
> + th = tcp_hdr(skb);
> + th->doff = sizeof(*th) / 4;
> +
> + skb->ip_summed = CHECKSUM_PARTIAL;
> + skb->csum_start = skb_transport_header(skb) - skb->head;
> + skb->csum_offset = offsetof(struct tcphdr, check);
> + skb_shinfo(skb)->gso_type = SKB_GSO_TCPV6;
> + skb_shinfo(skb)->gso_size = GSO_TEST_SIZE;
> + skb_shinfo(skb)->gso_segs = DIV_ROUND_UP(payload_len, GSO_TEST_SIZE);
> +
> + return skb;
> +}
> +
> +/*
> + * The device GSO size limit is per L3 protocol, so only a lowered limit of
> + * the packet's own protocol may cost it the TSO bit, and a tag inside the
> + * frame, which replaces skb->protocol with the ethertype, must not change
> + * which limit applies. Takes ownership of @skb.
> + */
> +static void gso_test_tcp_l3_limit(struct kunit *test, struct net_device *dev,
> + struct sk_buff *skb, netdev_features_t tso,
> + bool ipv6)
> +{
> + unsigned int other = GSO_LEGACY_MAX_SIZE;
> + unsigned int own = GSO_MAX_SIZE;
> +
> + /* The skb fits the limit of its own L3 protocol, not the other's. */
> + dev->gso_max_size = ipv6 ? own : other;
> + dev->gso_ipv4_max_size = ipv6 ? other : own;
> + KUNIT_EXPECT_TRUE(test, netif_skb_features(skb) & tso);
> +
> + /* ...and the other way around. */
> + swap(own, other);
> + dev->gso_max_size = ipv6 ? own : other;
> + dev->gso_ipv4_max_size = ipv6 ? other : own;
> + KUNIT_EXPECT_FALSE(test, netif_skb_features(skb) & tso);
> +
> + skb = vlan_insert_tag_set_proto(skb, htons(ETH_P_8021Q), 0);
> + if (!skb) {
> + KUNIT_FAIL(test, "no tagged skb");
> + return;
> + }
> + KUNIT_EXPECT_TRUE(test, skb->protocol == htons(ETH_P_8021Q));
> +
> + swap(own, other);
> + dev->gso_max_size = ipv6 ? own : other;
> + dev->gso_ipv4_max_size = ipv6 ? other : own;
> + KUNIT_EXPECT_TRUE(test, netif_skb_features(skb) & tso);
> +
> + swap(own, other);
> + dev->gso_max_size = ipv6 ? own : other;
> + dev->gso_ipv4_max_size = ipv6 ? other : own;
> + KUNIT_EXPECT_FALSE(test, netif_skb_features(skb) & tso);
> +
> + consume_skb(skb);
> +}
> +
> +static void gso_test_tcp_limit_l3_proto(struct kunit *test)
> +{
> + static const struct net_device_ops dummy_netdev_ops = { };
> + const unsigned int payload_len = 100 * 1024;
> + struct net_device *dev;
> + struct sk_buff *skb;
> +
> + dev = alloc_etherdev(0);
> + if (!dev) {
> + KUNIT_FAIL(test, "no net_device");
> + return;
> + }
> + dev->netdev_ops = &dummy_netdev_ops;
> + dev->hw_features = NETIF_F_SG | NETIF_F_HW_CSUM | NETIF_F_TSO |
> + NETIF_F_TSO6;
> + dev->features = dev->hw_features;
> + dev->vlan_features = dev->hw_features;
> +
> + skb = gso_tcp6_skb_new(payload_len);
> + if (!skb) {
> + KUNIT_FAIL(test, "no IPv6 skb");
> + goto free_dev;
> + }
> + skb->dev = dev;
> + gso_test_tcp_l3_limit(test, dev, skb, NETIF_F_TSO6, true);
> +
> + skb = gso_tcp_skb_new(payload_len);
> + if (!skb) {
> + KUNIT_FAIL(test, "no IPv4 skb");
> + goto free_dev;
> + }
> + skb->dev = dev;
> + gso_test_tcp_l3_limit(test, dev, skb, NETIF_F_TSO, false);
> +
> +free_dev:
> + free_netdev(dev);
> +}
> +
> +/*
> + * A re-segmented output is a plain GSO skb, so the whole length lands in the
> + * 16-bit L3 length field. A device which accepts GSO skbs above 64 KiB would
> + * otherwise be handed outputs whose length field truncates.
> + */
> +static void gso_test_tcp_resegment_l3_len(struct kunit *test)
> +{
> + static const struct net_device_ops dummy_netdev_ops = { };
> + const unsigned int payload_len = 200 * 1024;
> + struct sk_buff *skb, *segs, *cur, *next;
> + unsigned int nr_segs = 0;
> + struct net_device *dev;
> + bool again = false;
> +
> + if (!IS_ENABLED(CONFIG_INET))
> + kunit_skip(test, "requires CONFIG_INET");
> +
> + dev = alloc_etherdev(0);
> + if (!dev) {
> + KUNIT_FAIL(test, "no net_device");
> + return;
> + }
> + dev->netdev_ops = &dummy_netdev_ops;
> + dev->hw_features = NETIF_F_SG | NETIF_F_HW_CSUM | NETIF_F_TSO;
> + dev->features = dev->hw_features;
> + dev->vlan_features = dev->hw_features;
> + dev->gso_ipv4_max_size = 120 * 1024;
> +
> + skb = gso_tcp_skb_new(payload_len);
> + if (!skb) {
> + KUNIT_FAIL(test, "no skb");
> + goto free_dev;
> + }
> + skb->dev = dev;
> +
> + segs = validate_xmit_skb_list(skb, dev, &again);
> + if (IS_ERR_OR_NULL(segs)) {
> + KUNIT_FAIL(test, "segs error %pe", segs);
> + goto free_dev;
> + }
> +
> + for (cur = segs; cur; cur = next, nr_segs++) {
> + next = cur->next;
> + cur->next = NULL;
> +
> + KUNIT_EXPECT_TRUE(test, skb_is_gso(cur));
> + KUNIT_EXPECT_EQ(test, ntohs(ip_hdr(cur)->tot_len),
> + cur->len - (skb_network_header(cur) -
> + skb_mac_header(cur)));
> + consume_skb(cur);
> + }
> +
> + KUNIT_EXPECT_EQ(test, nr_segs, 4);
> +
> +free_dev:
> + free_netdev(dev);
> +}
> +
> /* IP tunnel flags */
>
> #include <net/ip_tunnels.h>
> @@ -372,6 +720,9 @@ static void ip_tunnel_flags_test_run(struct kunit *test)
>
> static struct kunit_case net_test_cases[] = {
> KUNIT_CASE_PARAM(gso_test_func, gso_test_gen_params),
> + KUNIT_CASE(gso_test_tcp_resegment),
> + KUNIT_CASE(gso_test_tcp_limit_l3_proto),
> + KUNIT_CASE(gso_test_tcp_resegment_l3_len),
> KUNIT_CASE_PARAM(ip_tunnel_flags_test_run,
> ip_tunnel_flags_test_gen_params),
> { },
> --
> 2.47.3
>
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH net-next v4 3/5] net: gso: support re-segmentation of TCP GSO skbs
2026-09-30 11:15 ` [PATCH net-next v4 3/5] net: gso: support re-segmentation of TCP GSO skbs Wang Zhan
2026-09-30 19:52 ` Willem de Bruijn
@ 2026-10-02 11:16 ` netdev-bot+sashiko
1 sibling, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02 11:16 UTC (permalink / raw)
To: wang.zhan
Cc: netdev, davem, edumazet, kuba, pabeni, horms, keyong.sun,
willemdebruijn.kernel, jasowangio, andrew+netdev, aconole,
echaudro, i.maximets, dev, daniel, ncardwell, kuniyu, alice,
david.laight.linux
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] Possible truncation of the 16-bit IP length for grouped outputs
larger than 64 KiB (fixed by the next patch in the series).
--
net: gso: support re-segmentation of TCP GSO skbs
This commit adds a max_segs argument to __skb_gso_segment() and stores it
in the skb_gso_cb scratch area. When max_segs is non-zero, skb_segment()
skips the partial/frag_list decision block and groups up to max_segs MSS
segments into each output GSO skb. Existing callers pass zero to keep
their current behaviour.
> diff --git a/net/core/gso.c b/net/core/gso.c
> index bcd156372f4df..21a259ca392f4 100644
> --- a/net/core/gso.c
> +++ b/net/core/gso.c
[ ... ]
> @@ -117,6 +120,7 @@ struct sk_buff *__skb_gso_segment(struct sk_buff *skb,
>
> SKB_GSO_CB(skb)->mac_offset = skb_headroom(skb);
> SKB_GSO_CB(skb)->encap_level = 0;
> + SKB_GSO_CB(skb)->max_segs = min(max_segs, GSO_MAX_SEGS);
>
> skb_reset_mac_header(skb);
> skb_reset_mac_len(skb);
> diff --git a/net/core/skbuff.c b/net/core/skbuff.c
> index 8912a66cd9097..fd10a7cdc1883 100644
> --- a/net/core/skbuff.c
> +++ b/net/core/skbuff.c
[ ... ]
> @@ -4874,7 +4875,10 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
> * now.
> */
> DEBUG_NET_WARN_ON_ONCE(len / mss > GSO_MAX_SEGS);
> - partial_segs = min(len / mss, GSO_MAX_SEGS);
> + if (max_segs)
> + partial_segs = min(len / mss, max_segs);
> + else
> + partial_segs = min(len / mss, GSO_MAX_SEGS);
> if (partial_segs > 1)
> mss *= partial_segs;
> else
[Severity: Low]
When max_segs is non-zero, the grouped outputs are not SKB_GSO_PARTIAL
skbs. Their size is bounded only by max_segs * gso_size. The only clamp in
__skb_gso_segment() is on the segment count (GSO_MAX_SEGS), and nothing
limits the number of bytes.
For non-partial GSO outputs, the L3 length is taken from the whole output
skb and stored in a 16-bit field:
net/ipv4/af_inet.c:inet_gso_segment() {
...
else
tot_len = skb->len - nhoff;
...
iph->tot_len = htons(tot_len);
ip_send_check(iph);
...
}
net/ipv6/ip6_offload.c:ipv6_gso_segment() {
...
else
payload_len = skb->len - nhoff - sizeof(*ipv6h);
ipv6h->payload_len = htons(payload_len);
...
}
If a caller passed a max_segs where max_segs * gso_size plus headers is
more than 65535, would tot_len and payload_len be truncated, and would the
IPv4 header checksum then be computed over the wrong length?
No caller in this patch passes a non-zero value. The only non-zero caller
in the series is skb_gso_output_max_segs(), added by the next patch ("net:
core: re-segment oversized TCP GSO skbs"). It caps gso_max_size at
GSO_LEGACY_MAX_SIZE before it computes max_segs, so this cannot be hit
anywhere in the series. What is left is the API contract of
__skb_gso_segment().
Should the kernel-doc for max_segs say that the caller must keep
max_segs * gso_size plus headers within 64KB? Alternatively, could
__skb_gso_segment() or skb_segment() enforce a byte limit, so a future
caller cannot produce outputs with a truncated IP length?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260930111526.2183107-1-wang.zhan%40smartx.com
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH net-next v4 4/5] net: core: re-segment oversized TCP GSO skbs
2026-09-30 11:15 ` [PATCH net-next v4 4/5] net: core: re-segment oversized " Wang Zhan
2026-09-30 19:52 ` Willem de Bruijn
@ 2026-10-02 11:16 ` netdev-bot+sashiko
1 sibling, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02 11:16 UTC (permalink / raw)
To: wang.zhan
Cc: netdev, davem, edumazet, kuba, pabeni, horms, keyong.sun,
willemdebruijn.kernel, jasowangio, andrew+netdev, aconole,
echaudro, i.maximets, dev, daniel, ncardwell, kuniyu, alice,
david.laight.linux
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] The re-segmentation path in validate_xmit_skb() /
skb_gso_output_max_segs() lets untrusted SKB_GSO_DODGY TCP skbs reach…
--
net: core: re-segment oversized TCP GSO skbs
With this patch, validate_xmit_skb() no longer segments an unencapsulated
TCP GSO skb into individual packets when it exceeds the egress device's
gso_max_size or gso_max_segs. It re-segments the skb into smaller GSO skbs
instead. A new helper, skb_gso_output_max_segs(), works out how many MSS
segments each output skb may carry and passes that to __skb_gso_segment()
as max_segs.
> diff --git a/net/core/dev.c b/net/core/dev.c
> index a6213c9ed5e72..ff9bce2506f67 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -3977,6 +3977,53 @@ __netif_skb_features(struct sk_buff *skb, bool check_gso_limits)
> }
> EXPORT_SYMBOL(__netif_skb_features);
>
> +static unsigned int
> +skb_gso_output_max_segs(struct sk_buff *skb, struct net_device *dev)
> +{
> + unsigned int mss = skb_shinfo(skb)->gso_size;
> + unsigned int gso_max_size, hdr_len, max_segs;
> + netdev_features_t features;
> + struct tcphdr _tcph, *th;
> +
> + if (!skb_is_gso(skb) || !skb_is_gso_tcp(skb) ||
> + skb->encapsulation || skb_has_frag_list(skb))
> + return 0;
[ ... ]
> + gso_max_size = netif_get_gso_max_size(dev, skb);
> + gso_max_size = min(gso_max_size, GSO_LEGACY_MAX_SIZE);
> +
> + /*
> + * gso_within_dev_limits() accepts gso_segs == gso_max_segs but
> + * rejects skb->len >= gso_max_size, so only the size limit needs the
> + * - 1; the inner min() keeps that subtraction from wrapping.
> + */
> + max_segs = (gso_max_size - min(gso_max_size, hdr_len + 1)) / mss;
> + max_segs = min(max_segs, READ_ONCE(dev->gso_max_segs));
> +
> + return max_segs;
> +}
[Severity: Medium]
Can untrusted SKB_GSO_DODGY skbs now reach the DEBUG_NET_WARN_ON_ONCE() in
skb_segment()?
When max_segs is non-zero, skb_segment() skips the frag_list block. It then
runs the assertion against the full payload length:
net/core/skbuff.c:skb_segment() {
...
if (sg && csum && !gso_by_frags) {
if (!max_segs && !(features & NETIF_F_GSO_PARTIAL)) {
...
if (!list_skb ||
!net_gso_ok(features, skb_shinfo(head_skb)->gso_type))
goto normal;
...
}
...
DEBUG_NET_WARN_ON_ONCE(len / mss > GSO_MAX_SEGS);
...
}
Two things keep that assertion true today. First,
__virtio_net_hdr_to_skb() clamps TCP gso_size to at least TCP_MIN_GSO_SIZE,
guarded by:
BUILD_BUG_ON(TCP_MIN_GSO_SIZE * GSO_MAX_SEGS < GSO_MAX_SIZE);
Second, the grouping path was only used for skbs that passed
gso_within_dev_limits(), so len was below GSO_MAX_SIZE. The new path runs
only for skbs that fail those limits, so their length is no longer bounded
by GSO_MAX_SIZE.
With PACKET_VNET_HDR, packet_snd() does not check the length against the
MTU when a gso_type is set. packet_alloc_skb() accepts up to
MAX_SKB_FRAGS * (PAGE_SIZE << PAGE_ALLOC_COSTLY_ORDER) bytes of frags, plus
a larger linear part.
So a CAP_NET_RAW user, including user namespace root, can build a
TCPv4/v6 GSO skb with NEEDS_CSUM, gso_size 8 and a payload of 524288 bytes
or more. tun/tap goes through the same virtio_net_hdr_to_skb() helper.
For such an skb the path looks like this:
dev_queue_xmit()
validate_xmit_skb()
gso_within_dev_limits() fails
skb_gso_output_max_segs() returns about (65536 - hdr_len - 1) / 8
__skb_gso_segment(skb, features, true, max_segs)
tcp_gso_segment()
skb_segment()
DEBUG_NET_WARN_ON_ONCE(len / mss > GSO_MAX_SEGS)
Every check in the helper passes here. csum_start equals the transport
header set by skb_partial_csum_set(), and net_gso_ok() succeeds on a TSO
device.
Before this patch the same skb had max_segs == 0 and no GSO_PARTIAL. It
took the goto normal path and never reached the assertion.
The output segments should still be correct, since
partial_segs = min(len / mss, max_segs). But with CONFIG_DEBUG_NET and
panic_on_warn, the new warning would panic the kernel.
Could skb_gso_output_max_segs() return 0 when (skb->len - hdr_len) / mss
exceeds GSO_MAX_SEGS?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260930111526.2183107-1-wang.zhan%40smartx.com
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH net-next v4 5/5] net: net_test: add tests for TCP re-segmentation
2026-09-30 11:15 ` [PATCH net-next v4 5/5] net: net_test: add tests for TCP re-segmentation Wang Zhan
2026-09-30 19:53 ` Willem de Bruijn
@ 2026-10-02 11:16 ` netdev-bot+sashiko
1 sibling, 0 replies; 17+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02 11:16 UTC (permalink / raw)
To: wang.zhan
Cc: netdev, davem, edumazet, kuba, pabeni, horms, keyong.sun,
willemdebruijn.kernel, jasowangio, andrew+netdev, aconole,
echaudro, i.maximets, dev, daniel, ncardwell, kuniyu, alice,
david.laight.linux
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] The commit message says the parameterized GSO test gains "two
limits which cut the input into two and three output skbs".
- [Low] The new exact gso_segs checks cannot catch a broken gso_segs
value on the last re-segmented output.
- [Low] In gso_test_func(), the new `KUNIT_ASSERT_LT(test, i,
tcase->nr_segs);` inside the segment-freeing loop…
--
net: net_test: add tests for TCP re-segmentation
This commit adds KUnit coverage for the max_segs parameter of the GSO
engine. It adds three re-segmentation cases to the parameterized
skb_segment() test. It also adds TCP tests that drive __skb_gso_segment()
and validate_xmit_skb_list(), and a test that checks that
netif_skb_features() applies the GSO size limit of the packet's L3
protocol.
> The parameterized GSO test gains a max_segs input and three cases: two
> limits which cut the input into two and three output skbs, and a max_segs
> of one MSS, which must leave the output unchanged.
[Severity: Low]
Is "two limits" accurate here? GSO_TEST_RESEGMENT and
GSO_TEST_RESEGMENT_MULTI both use the same limit:
.max_segs = 2,
That limit is applied to a 3003 byte payload and to a 5003 byte payload.
The only other value tested is max_segs = 1 in GSO_TEST_RESEGMENT_ONE_MSS.
Could the message say that one limit is applied to two inputs?
> diff --git a/net/core/net_test.c b/net/core/net_test.c
> index 9c3a590865d26..6a91db83e5301 100644
> --- a/net/core/net_test.c
> +++ b/net/core/net_test.c
[ ... ]
> @@ -239,6 +297,7 @@ static void gso_test_func(struct kunit *test)
> for (cur = segs, i = 0; cur; cur = next, i++) {
> next = cur->next;
>
> + KUNIT_ASSERT_LT(test, i, tcase->nr_segs);
> KUNIT_ASSERT_EQ(test, cur->len, sizeof(hdr) + tcase->segs[i]);
[Severity: Low]
If skb_segment() returns more outputs than tcase->nr_segs, this assertion
aborts the test through kunit_try_catch_throw(). Does that leak cur and
every later skb on the ->next chain, along with their page frag
references? It would also skip freeing the input skb at free_gso_skb.
The v4 changelog says "free the segments when a bounded case does not
match". gso_test_tcp_resegment() does this: it uses EXPECT checks and
consumes every segment, but gso_test_func() still aborts here.
The other ASSERTs in this loop already abort the same way. This one does
stop the read past the end of tcase->segs[] on the next line.
Would an EXPECT here, with the per-index checks skipped once i reaches
tcase->nr_segs, keep both the bounds protection and the cleanup?
>
> /* segs have skb->data pointing to the mac header */
> @@ -247,6 +306,18 @@ static void gso_test_func(struct kunit *test)
>
> /* header was copied to all segs */
> KUNIT_ASSERT_EQ(test, memcmp(skb_mac_header(cur), hdr, sizeof(hdr)), 0);
> + if (tcase->segs_are_gso) {
> + KUNIT_EXPECT_TRUE(test, skb_is_gso(cur));
> + KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_size,
> + GSO_TEST_SIZE);
> + KUNIT_EXPECT_EQ(test, skb_shinfo(cur)->gso_segs,
> + DIV_ROUND_UP(tcase->segs[i],
> + GSO_TEST_SIZE));
[Severity: Low]
Can this check detect a wrong gso_segs on the last output? skb_segment()
first gives every output partial_segs and then fixes up the tail:
net/core/skbuff.c:skb_segment() {
...
for (iter = segs; iter; iter = iter->next) {
skb_shinfo(iter)->gso_size = gso_size;
skb_shinfo(iter)->gso_segs = partial_segs;
...
if (tail->len - doffset <= gso_size)
skb_shinfo(tail)->gso_size = 0;
else if (tail != segs)
skb_shinfo(tail)->gso_segs = DIV_ROUND_UP(tail->len - doffset, gso_size);
...
}
GSO_TEST_RESEGMENT, GSO_TEST_RESEGMENT_MULTI and gso_test_tcp_resegment()
all use max_segs = 2 and end with a 1003 byte tail. DIV_ROUND_UP(1003,
1000) is 2, the same as partial_segs, so these checks would still pass if
the tail fix-up were dropped.
The one test whose tail count differs is gso_test_tcp_resegment_l3_len().
Its outputs are 65000, 65000, 65000 and 9800 bytes, so the tail should have
gso_segs of 10 rather than 65. That test does not check gso_segs.
The v4 changelog says "check the exact gso_segs of each output". Could one
of the vectors use a tail that gives a different count? For example,
max_segs = 3 with a 4003 byte payload should give gso_segs of 3 and 2.
> + KUNIT_EXPECT_FALSE(test, skb_shinfo(cur)->gso_type &
> + SKB_GSO_PARTIAL);
> + } else if (tcase->max_segs) {
> + KUNIT_EXPECT_FALSE(test, skb_is_gso(cur));
> + }
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260930111526.2183107-1-wang.zhan%40smartx.com
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH net-next v4 3/5] net: gso: support re-segmentation of TCP GSO skbs
2026-09-30 19:52 ` Willem de Bruijn
@ 2026-10-02 16:51 ` Wang Zhan
0 siblings, 0 replies; 17+ messages in thread
From: Wang Zhan @ 2026-10-02 16:51 UTC (permalink / raw)
To: netdev, Willem de Bruijn
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun, Wang Zhan,
Jason Wang, Andrew Lunn, Aaron Conole, Eelco Chaudron,
Ilya Maximets, dev, Daniel Borkmann, Neal Cardwell,
Kuniyuki Iwashima, Alice Mikityanska, David Laight
On Wed, 30 Sep 2026 15:52:01 -0400 Willem de Bruijn wrote:
> If respinning, consider asking the LLM to rewrite commit messages to be
> more concise.
Sure, will do in v5.
> Not new with this series, nor high priority, so only if respinning:
>
> Consider adding
>
> @@ -5121,7 +5121,7 @@ struct sk_buff *skb_segment(struct sk_buff *head_skb,
> }
>
> if (tail->len - doffset <= gso_size)
> - skb_shinfo(tail)->gso_size = 0;
> + skb_gso_reset(tail);
> else if (tail != segs)
>
> Because packets looped into the Rx path will sometimes have their
> gso_segs used without first checking gso_size / skb_is_gso. For instance
> in ip_rcv_core __IP_ADD_STATS.
Agreed. Several places count with gso_segs without checking gso_size
(ip_rcv_core, tp->rcv_ooopack), and with this series the re-segmented output
reaches them in-tree. I will add it as a new patch in the series.
> Additionally for patch 5/5, consider one testcase that resegments, but for
> which the last skb is not a GSO skb.
Sure, will add in v5.
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH net-next v4 5/5] net: net_test: add tests for TCP re-segmentation
2026-09-30 19:53 ` Willem de Bruijn
@ 2026-10-02 17:59 ` Wang Zhan
2026-10-03 19:10 ` Willem de Bruijn
0 siblings, 1 reply; 17+ messages in thread
From: Wang Zhan @ 2026-10-02 17:59 UTC (permalink / raw)
To: netdev, Willem de Bruijn
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun, Wang Zhan,
Jason Wang, Andrew Lunn, Aaron Conole, Eelco Chaudron,
Ilya Maximets, dev, Daniel Borkmann, Neal Cardwell,
Kuniyuki Iwashima, Alice Mikityanska, David Laight
On Wed, 30 Sep 2026 15:53:09 -0400 Willem de Bruijn wrote:
> Given that we have the above (concise) testcases:
>
> Do we need to add the below verbose code? What does it add, just the
> TCP specific path?
gso_test_tcp_resegment() is a v1 leftover, when tcp_gso_segment() needed a
max_segs exception in its "untrusted source" check. max_segs is passed
through now, so it is unnecessary and can be dropped in v5.
gso_test_tcp_limit_l3_proto() was added in v3 for the L3 protocol limit
lookup fix (1/5).
gso_test_tcp_resegment_l3_len() was added in v3 for the 64 KiB cap on the
grouped output (4/5), from this sashiko report on v2:
https://lore.kernel.org/179002380484.2160803.13955762820082413555@kernel.org/
You are right that they are verbose and one-off. I do not have a concrete plan
yet: I will try to trim them, or just drop them and add a selftest later. Any
suggestions?
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH net-next v4 5/5] net: net_test: add tests for TCP re-segmentation
2026-10-02 17:59 ` Wang Zhan
@ 2026-10-03 19:10 ` Willem de Bruijn
0 siblings, 0 replies; 17+ messages in thread
From: Willem de Bruijn @ 2026-10-03 19:10 UTC (permalink / raw)
To: Wang Zhan, netdev, Willem de Bruijn
Cc: davem, edumazet, kuba, pabeni, horms, keyong.sun, Wang Zhan,
Jason Wang, Andrew Lunn, Aaron Conole, Eelco Chaudron,
Ilya Maximets, dev, Daniel Borkmann, Neal Cardwell,
Kuniyuki Iwashima, Alice Mikityanska, David Laight
Wang Zhan wrote:
> On Wed, 30 Sep 2026 15:53:09 -0400 Willem de Bruijn wrote:
> > Given that we have the above (concise) testcases:
> >
> > Do we need to add the below verbose code? What does it add, just the
> > TCP specific path?
>
> gso_test_tcp_resegment() is a v1 leftover, when tcp_gso_segment() needed a
> max_segs exception in its "untrusted source" check. max_segs is passed
> through now, so it is unnecessary and can be dropped in v5.
>
> gso_test_tcp_limit_l3_proto() was added in v3 for the L3 protocol limit
> lookup fix (1/5).
>
> gso_test_tcp_resegment_l3_len() was added in v3 for the 64 KiB cap on the
> grouped output (4/5), from this sashiko report on v2:
> https://lore.kernel.org/179002380484.2160803.13955762820082413555@kernel.org/
>
> You are right that they are verbose and one-off. I do not have a concrete plan
> yet: I will try to trim them, or just drop them and add a selftest later. Any
> suggestions?
The trade-off in added coverage vs added test code (that also has a
maintenance cost) is subjective.
I would maybe ask the LLM whether it can generate a significantly more
concise test that gives the same coverage. Initially they don't see to
optimize for size. It has worked for me before.
Otherwise, I'd drop 64KB limit test gso_test_tcp_resegment_l3_len, and
possibly both.
But as said subjective, your call.
^ permalink raw reply [flat|nested] 17+ messages in thread
end of thread, other threads:[~2026-10-03 19:10 UTC | newest]
Thread overview: 17+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30 11:15 [PATCH net-next v4 0/5] net: re-segment oversized TCP GSO skbs Wang Zhan
2026-09-30 11:15 ` [PATCH net-next v4 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
2026-09-30 19:40 ` Willem de Bruijn
2026-09-30 11:15 ` [PATCH net-next v4 2/5] net: core: factor out the GSO device limit check Wang Zhan
2026-09-30 19:40 ` Willem de Bruijn
2026-09-30 11:15 ` [PATCH net-next v4 3/5] net: gso: support re-segmentation of TCP GSO skbs Wang Zhan
2026-09-30 19:52 ` Willem de Bruijn
2026-10-02 16:51 ` Wang Zhan
2026-10-02 11:16 ` netdev-bot+sashiko
2026-09-30 11:15 ` [PATCH net-next v4 4/5] net: core: re-segment oversized " Wang Zhan
2026-09-30 19:52 ` Willem de Bruijn
2026-10-02 11:16 ` netdev-bot+sashiko
2026-09-30 11:15 ` [PATCH net-next v4 5/5] net: net_test: add tests for TCP re-segmentation Wang Zhan
2026-09-30 19:53 ` Willem de Bruijn
2026-10-02 17:59 ` Wang Zhan
2026-10-03 19:10 ` Willem de Bruijn
2026-10-02 11:16 ` netdev-bot+sashiko
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox