From: Willem de Bruijn <willemdebruijn.kernel@gmail.com>
To: Wang Zhan <wang.zhan@smartx.com>, netdev@vger.kernel.org
Cc: davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
pabeni@redhat.com, horms@kernel.org, keyong.sun@smartx.com,
Willem de Bruijn <willemdebruijn.kernel@gmail.com>,
Jason Wang <jasowangio@gmail.com>,
Andrew Lunn <andrew+netdev@lunn.ch>,
Aaron Conole <aconole@redhat.com>,
Eelco Chaudron <echaudro@redhat.com>,
Ilya Maximets <i.maximets@ovn.org>,
dev@openvswitch.org, Daniel Borkmann <daniel@iogearbox.net>,
Neal Cardwell <ncardwell@google.com>,
Kuniyuki Iwashima <kuniyu@google.com>,
Alice Mikityanska <alice@isovalent.com>,
David Laight <david.laight.linux@gmail.com>,
Wang Zhan <wang.zhan@smartx.com>
Subject: Re: [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs
Date: Mon, 28 Sep 2026 19:47:43 -0400 [thread overview]
Message-ID: <willemdebruijn.kernel.3b556463ed93f@gmail.com> (raw)
In-Reply-To: <20260928044102.1004310-5-wang.zhan@smartx.com>
Wang Zhan wrote:
> A GSO skb which exceeds an egress device limit loses its GSO feature mask
> and is segmented into individual packets. This is unnecessarily expensive
> when the device can still offload smaller TCP GSO skbs, which is easy to
> hit once one hop of a BIG TCP path raises gso_max_size and the next one
> does not.
>
> For an unencapsulated TCP GSO skb which exceeds gso_max_size or
> gso_max_segs, work out how many MSS segments each output skb may carry and
> resegment the skb with that max_segs bound instead. Encapsulated and
> frag-list skbs, GSO types the device cannot offload and bounds which leave
> room for a single MSS keep today's segmentation. The output obeys the GSO
> feature and limit contract the device already advertises, so it applies
> automatically, without extra device state or a userspace control.
>
> The output is a plain GSO skb, so its whole length lands in the 16-bit L3
> length field which inet_gso_segment() and ipv6_gso_segment() write. An
> egress limit above 64 KiB would give outputs whose length truncates, so
> the size limit is also capped at what that field can express.
>
> The helper runs on the skb which is handed to the driver, after
> validate_xmit_vlan() and sk_validate_xmit_skb(), and only from the
> netif_needs_gso() branch, so an skb which is not segmented pays nothing.
> An over-limit skb pays one device limit test and one ndo_features_check()
> for the bound, in exchange for keeping the output a GSO skb.
>
> Measured on a veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP
> enabled on the veth endpoints and left off in the guest, so the skbs which
> the veth hop accepts have to be segmented before the TAP device. A single
> iperf3 TCP flow, six alternating runs per state (`-t 15 -O 5`, fixed CPU
> affinity and port tuple). The middle column is the same tree with the
> resegmentation disabled:
>
> protocol no BIG TCP mixed, no reseg mixed, resegmented
> TCP/IPv4 51.550 Gbps 15.850 Gbps 52.617 Gbps
> TCP/IPv6 52.050 Gbps 15.783 Gbps 51.933 Gbps
>
> Coefficient of variation for the two mixed columns was 0.48% and 0.82% for
> IPv4 and 0.44% and 0.44% for IPv6. A BIG TCP hop which feeds a 64 KiB hop
> loses 69% of the throughput of a path which never enables BIG TCP at all;
> bounded resegmentation recovers it, 3.3x over the existing segmentation
> path and within noise of the no BIG TCP baseline.
>
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
>
> ---
> v3:
> - the L3 protocol change moved to patch 1
> - fold skb_can_gso_resegment() into the helper, now skb_gso_output_max_segs()
> - pass the caller's features on unchanged, the bound only shapes the output
> - drop the scatter-gather and checksum tests skb_segment() applies itself
> - saturate the size limit with min(), drop the sub-MSS guard
> - cap the output size at GSO_LEGACY_MAX_SIZE, these outputs are not BIG TCP
> v2: https://lore.kernel.org/20260918084651.3022878-4-wang.zhan@smartx.com/
> v1: https://lore.kernel.org/20260917063854.2011613-4-wang.zhan@smartx.com/
> ---
> net/core/dev.c | 63 +++++++++++++++++++++++++++++++++++++++++++++++++-
> 1 file changed, 62 insertions(+), 1 deletion(-)
>
> diff --git a/net/core/dev.c b/net/core/dev.c
> index d66b667071837..728260772f349 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -3933,6 +3933,63 @@ netdev_features_t netif_skb_features(struct sk_buff *skb)
> }
> EXPORT_SYMBOL(netif_skb_features);
>
> +static unsigned int
> +skb_gso_output_max_segs(struct sk_buff *skb, struct net_device *dev)
> +{
> + unsigned int mss = skb_shinfo(skb)->gso_size;
> + unsigned int gso_max_size, hdr_len, max_segs;
> + netdev_features_t features;
> + struct tcphdr _tcph, *th;
> +
> + /*
> + * The TCP frag-list path segments through skb_segment_list(), which
> + * does not carry max_segs, so bounded calls skip those skbs.
> + */
This comment answers only one of six conditions. And one that is
pretty straightforward. I'd drop.
In general, drop all too-obvious comments. AI has a habit of adding
a lot more, and more low information, comments than is customary in
kernel code (where we also have commit messages). Generally, repeating
what the code does is of little value.
> + if (!skb_is_gso(skb) || !skb_is_gso_tcp(skb) ||
> + skb->encapsulation || skb_has_frag_list(skb) ||
> + !skb_mac_header_was_set(skb) ||
> + !skb_transport_header_was_set(skb))
Conversely, they last two conditions are less obvious. Are they not
always true for a TSO packet?
> + return 0;
> +
> + th = skb_header_pointer(skb, skb_transport_offset(skb), sizeof(_tcph),
> + &_tcph);
> + if (!th || th->doff < sizeof(*th) / 4)
> + return 0;
> +
> + hdr_len = skb_transport_header(skb) - skb_mac_header(skb) +
> + th->doff * 4;
> +
> + /*
> + * The output stays a GSO skb, so the device has to offload the GSO
> + * type. The caller's features cannot tell that: this skb is over
> + * the device limits, so gso_features_check() has cleared their GSO
> + * bits. Compute the features again without the limit checks.
> + */
> + features = __netif_skb_features(skb, false);
> + if (!net_gso_ok(features | NETIF_F_GSO_ROBUST,
> + skb_shinfo(skb)->gso_type))
> + return 0;
> +
> + /*
> + * The output is a plain GSO skb, so its whole length lands in the
> + * 16-bit L3 length field: only a BIG TCP skb may exceed the legacy
> + * GSO size, and this path does not emit one.
> + */
> + gso_max_size = netif_get_gso_max_size(dev, vlan_get_protocol(skb));
Third time this is now called in validate_xmit_skb. Not sure if that can
easily be avoided.
> + gso_max_size = min(gso_max_size, GSO_LEGACY_MAX_SIZE);
> +
> + /*
> + * gso_within_dev_limits() accepts gso_segs == gso_max_segs but
> + * rejects skb->len >= gso_max_size, so only the size bound needs - 1.
> + * The min() keeps that subtraction from wrapping when the device
> + * limit is smaller than the headers.
> + */
> + max_segs = (gso_max_size - min(gso_max_size, hdr_len + 1)) / mss;
> + max_segs = min(max_segs, READ_ONCE(dev->gso_max_segs));
> +
> + return max_segs;
> +}
> +
> static int xmit_one(struct sk_buff *skb, struct net_device *dev,
> struct netdev_queue *txq, bool more)
> {
> @@ -4097,9 +4154,13 @@ static struct sk_buff *validate_xmit_skb(struct sk_buff *skb, struct net_device
> goto out_null;
>
> if (netif_needs_gso(skb, features)) {
> + unsigned int max_segs = 0;
> struct sk_buff *segs;
>
> - segs = skb_gso_segment(skb, features);
> + if (unlikely(!gso_within_dev_limits(skb, dev)))
> + max_segs = skb_gso_output_max_segs(skb, dev);
> +
> + segs = __skb_gso_segment(skb, features, true, max_segs);
> if (IS_ERR(segs)) {
> goto out_kfree_skb;
> } else if (segs) {
> --
> 2.47.3
>
next prev parent reply other threads:[~2026-09-28 23:47 UTC|newest]
Thread overview: 27+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 4:40 [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs Wang Zhan
2026-09-28 4:40 ` [PATCH net-next v3 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
2026-09-28 23:36 ` Willem de Bruijn
2026-09-29 3:44 ` Wang Zhan
2026-09-29 4:01 ` Wang Zhan
2026-09-29 14:59 ` Willem de Bruijn
2026-09-28 4:40 ` [PATCH net-next v3 2/5] net: core: factor out the GSO device limit check Wang Zhan
2026-09-28 23:37 ` Willem de Bruijn
2026-09-29 7:47 ` Paolo Abeni
2026-09-28 4:41 ` [PATCH net-next v3 3/5] net: gso: support bounded TCP segmentation Wang Zhan
2026-09-28 23:39 ` Willem de Bruijn
2026-09-30 4:41 ` netdev-bot+sashiko
2026-09-28 4:41 ` [PATCH net-next v3 4/5] net: core: resegment oversized TCP GSO skbs Wang Zhan
2026-09-28 23:47 ` Willem de Bruijn [this message]
2026-09-29 10:25 ` Wang Zhan
2026-09-29 15:00 ` Willem de Bruijn
2026-09-30 4:41 ` netdev-bot+sashiko
2026-09-28 4:41 ` [PATCH net-next v3 5/5] net: net_test: add tests for bounded GSO segmentation Wang Zhan
2026-09-28 23:59 ` Willem de Bruijn
2026-09-29 10:30 ` Wang Zhan
2026-09-29 15:01 ` Willem de Bruijn
2026-09-30 4:41 ` netdev-bot+sashiko
2026-09-28 4:45 ` [PATCH net-next v3 0/5] net: resegment oversized TCP GSO skbs netdev-bot+sinfo
2026-09-28 5:49 ` Wang Zhan
2026-09-28 23:34 ` Willem de Bruijn
2026-09-29 7:27 ` Paolo Abeni
2026-09-29 11:50 ` Wang Zhan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=willemdebruijn.kernel.3b556463ed93f@gmail.com \
--to=willemdebruijn.kernel@gmail.com \
--cc=aconole@redhat.com \
--cc=alice@isovalent.com \
--cc=andrew+netdev@lunn.ch \
--cc=daniel@iogearbox.net \
--cc=davem@davemloft.net \
--cc=david.laight.linux@gmail.com \
--cc=dev@openvswitch.org \
--cc=echaudro@redhat.com \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=i.maximets@ovn.org \
--cc=jasowangio@gmail.com \
--cc=keyong.sun@smartx.com \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=ncardwell@google.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=wang.zhan@smartx.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox