From: Willem de Bruijn <willemdebruijn.kernel@gmail.com>
To: Wang Zhan <wang.zhan@smartx.com>, netdev@vger.kernel.org
Cc: davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
pabeni@redhat.com, horms@kernel.org, keyong.sun@smartx.com,
Willem de Bruijn <willemdebruijn.kernel@gmail.com>,
Jason Wang <jasowangio@gmail.com>,
Andrew Lunn <andrew+netdev@lunn.ch>,
Aaron Conole <aconole@redhat.com>,
Eelco Chaudron <echaudro@redhat.com>,
Ilya Maximets <i.maximets@ovn.org>,
dev@openvswitch.org, Daniel Borkmann <daniel@iogearbox.net>,
Neal Cardwell <ncardwell@google.com>,
Kuniyuki Iwashima <kuniyu@google.com>,
Alice Mikityanska <alice@isovalent.com>,
David Laight <david.laight.linux@gmail.com>,
Wang Zhan <wang.zhan@smartx.com>
Subject: Re: [PATCH net-next v4 4/5] net: core: re-segment oversized TCP GSO skbs
Date: Wed, 30 Sep 2026 15:52:09 -0400 [thread overview]
Message-ID: <willemdebruijn.kernel.66e7c7b6f68b@gmail.com> (raw)
In-Reply-To: <20260930111526.2183107-5-wang.zhan@smartx.com>
Wang Zhan wrote:
> A GSO skb which exceeds an egress device limit loses its GSO feature mask
> and is segmented into individual packets. This is unnecessarily expensive
> when the device can still offload smaller TCP GSO skbs, which is easy to
> hit once one hop of a BIG TCP path raises gso_max_size and the next one
> does not.
>
> For an unencapsulated TCP GSO skb which exceeds gso_max_size or
> gso_max_segs, work out how many MSS segments each output skb may carry and
> re-segment the skb with that max_segs instead. The bound is measured from
> the transport header, so an skb whose transport header is unset or stale
> keeps today's segmentation.
>
> An encapsulated or frag-list skb, a GSO type the device cannot offload and
> a max_segs which leaves room for a single MSS all keep today's segmentation
> as well. The GSO type test runs on the features without the limit checks,
> because gso_features_check() has already cleared the GSO bits of an skb
> which exceeds them.
>
> This path emits a plain GSO skb, not a BIG TCP one: inet_gso_segment() and
> ipv6_gso_segment() write the whole length of each output into the 16-bit L3
> length field. An egress limit above 64 KiB would give outputs whose length
> truncates, so the size limit is also capped at what that field can express.
>
> The helper runs on the skb which is handed to the driver, after
> validate_xmit_vlan() and sk_validate_xmit_skb(), and only from the
> netif_needs_gso() branch: an skb which the device takes as it is pays
> nothing. A skb which reaches that branch pays one device limit test, and
> an over-limit one pays the TCP header read and the features recomputation
> for max_segs, in exchange for keeping the output a GSO skb.
>
> Measured on a veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP
> enabled on the veth endpoints and left off in the guest, so the skbs which
> the veth hop accepts have to be segmented before the TAP device. A single
> iperf3 TCP flow, six alternating runs per state (`-t 15 -O 5`, fixed CPU
> affinity and port tuple). The middle column is the same tree with the
> re-segmentation disabled:
>
> protocol no BIG TCP mixed, no reseg mixed, resegmented
> TCP/IPv4 51.550 Gbps 15.850 Gbps 52.617 Gbps
> TCP/IPv6 52.050 Gbps 15.783 Gbps 51.933 Gbps
>
> Coefficient of variation for the two mixed columns was 0.48% and 0.82% for
> IPv4 and 0.44% and 0.44% for IPv6. A BIG TCP hop which feeds a 64 KiB hop
> loses 69% of the throughput of a path which never enables BIG TCP at all;
> re-segmentation recovers it, 3.3x over the existing segmentation path and
> within noise of the no BIG TCP baseline.
>
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
>
> ---
> v4:
> - drop the comment above the frag_list check
> - drop the mac header test, it is always set on this path
> - take the TCP header from the checksum start the GSO engine uses
> - test the transport header first, skb_transport_header() warns when unset
> - shorten the comment above the features lookup
> v3: https://lore.kernel.org/20260928044102.1004310-5-wang.zhan@smartx.com/
> v2: https://lore.kernel.org/20260918084651.3022878-4-wang.zhan@smartx.com/
> v1: https://lore.kernel.org/20260917063854.2011613-4-wang.zhan@smartx.com/
Reviewed-by: Willem de Bruijn <willemb@google.com>
next prev parent reply other threads:[~2026-09-30 19:52 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 11:15 [PATCH net-next v4 0/5] net: re-segment oversized TCP GSO skbs Wang Zhan
2026-09-30 11:15 ` [PATCH net-next v4 1/5] net: core: use the packet's L3 protocol for the GSO size limit Wang Zhan
2026-09-30 19:40 ` Willem de Bruijn
2026-09-30 11:15 ` [PATCH net-next v4 2/5] net: core: factor out the GSO device limit check Wang Zhan
2026-09-30 19:40 ` Willem de Bruijn
2026-09-30 11:15 ` [PATCH net-next v4 3/5] net: gso: support re-segmentation of TCP GSO skbs Wang Zhan
2026-09-30 19:52 ` Willem de Bruijn
2026-10-02 16:51 ` Wang Zhan
2026-10-02 11:16 ` netdev-bot+sashiko
2026-09-30 11:15 ` [PATCH net-next v4 4/5] net: core: re-segment oversized " Wang Zhan
2026-09-30 19:52 ` Willem de Bruijn [this message]
2026-10-02 11:16 ` netdev-bot+sashiko
2026-09-30 11:15 ` [PATCH net-next v4 5/5] net: net_test: add tests for TCP re-segmentation Wang Zhan
2026-09-30 19:53 ` Willem de Bruijn
2026-10-02 17:59 ` Wang Zhan
2026-10-03 19:10 ` Willem de Bruijn
2026-10-02 11:16 ` netdev-bot+sashiko
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=willemdebruijn.kernel.66e7c7b6f68b@gmail.com \
--to=willemdebruijn.kernel@gmail.com \
--cc=aconole@redhat.com \
--cc=alice@isovalent.com \
--cc=andrew+netdev@lunn.ch \
--cc=daniel@iogearbox.net \
--cc=davem@davemloft.net \
--cc=david.laight.linux@gmail.com \
--cc=dev@openvswitch.org \
--cc=echaudro@redhat.com \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=i.maximets@ovn.org \
--cc=jasowangio@gmail.com \
--cc=keyong.sun@smartx.com \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=ncardwell@google.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=wang.zhan@smartx.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox