Netdev List
 help / color / mirror / Atom feed
From: Willem de Bruijn <willemdebruijn.kernel@gmail.com>
To: Wang Zhan <wang.zhan@smartx.com>,  netdev@vger.kernel.org
Cc: davem@davemloft.net,  edumazet@google.com,  kuba@kernel.org,
	 pabeni@redhat.com,  horms@kernel.org,  keyong.sun@smartx.com,
	 Ilya Maximets <i.maximets@ovn.org>,
	 Aaron Conole <aconole@redhat.com>,
	 Eelco Chaudron <echaudro@redhat.com>,
	 dev@openvswitch.org,  Andrew Lunn <andrew+netdev@lunn.ch>,
	 Jason Wang <jasowangio@gmail.com>,
	 Willem de Bruijn <willemdebruijn.kernel@gmail.com>,
	 Neal Cardwell <ncardwell@google.com>,
	 Kuniyuki Iwashima <kuniyu@google.com>,
	 Alice Mikityanska <alice@isovalent.com>,
	 Wang Zhan <wang.zhan@smartx.com>
Subject: Re: [PATCH net-next v2 3/4] net: core: resegment oversized TCP GSO skbs
Date: Sat, 19 Sep 2026 11:37:45 -0400	[thread overview]
Message-ID: <willemdebruijn.kernel.32be703ea17ab@gmail.com> (raw)
In-Reply-To: <20260918084651.3022878-4-wang.zhan@smartx.com>

Wang Zhan wrote:
> A GSO skb which exceeds an egress device limit loses its GSO feature mask
> and is segmented into individual packets. This is unnecessarily expensive
> when the device can still offload smaller TCP GSO skbs, which is easy to
> hit once one hop of a BIG TCP path raises gso_max_size and the next one
> does not.
> 
> For an unencapsulated TCP GSO skb which exceeds gso_max_size or
> gso_max_segs, work out how many MSS segments each output skb may carry and
> resegment the skb with that bound instead. Keep the features computed
> without the GSO limit checks, which say whether the device offloads the
> GSO type at all. Encapsulated and frag-list skbs, GSO types the device
> cannot offload, and bounds below two segments keep the existing full
> segmentation path. The result obeys the GSO feature and limit contract the
> device already advertises, so apply it automatically, without extra device
> state or a userspace control.
> 
> The check runs on the skb which is handed to the driver, after
> validate_xmit_vlan() and sk_validate_xmit_skb(), and costs one extra
> ndo_features_check() on the oversized path, against segmenting the skb
> into individual packets. That position is also why the limit follows the
> L3 protocol rather than skb->protocol: validate_xmit_vlan() replaces the
> latter with the VLAN ethertype when it pushes the tag inside the skb.
> 
> Measured on a veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP
> enabled on the veth endpoints and left off in the guest, so the skbs which
> the veth hop accepts have to be segmented before the TAP device. A single
> iperf3 TCP flow, six alternating runs per state (`-t 15 -O 5`, fixed CPU
> affinity and port tuple). The middle column is the same tree with the
> resegmentation disabled:
> 
>   protocol  no BIG TCP   mixed, no reseg  mixed, resegmented
>   TCP/IPv4  51.550 Gbps  15.850 Gbps      52.617 Gbps
>   TCP/IPv6  52.050 Gbps  15.783 Gbps      51.933 Gbps
> 
> Coefficient of variation for the two mixed columns was 0.48% and 0.82%
> for IPv4 and 0.44% and 0.44% for IPv6. A BIG TCP hop which feeds a 64 KiB
> hop loses 69% of the throughput of a path which never enables BIG TCP at
> all; bounded resegmentation recovers it, 3.3x over the existing
> segmentation path and within noise of the no BIG TCP baseline.
> 
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <wang.zhan@smartx.com>
> ---
>  net/core/dev.c | 113 ++++++++++++++++++++++++++++++++++++++++++++++---
>  1 file changed, 106 insertions(+), 7 deletions(-)
> 
> diff --git a/net/core/dev.c b/net/core/dev.c
> index 16685888b2812..548db4d4e874c 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -3834,18 +3834,24 @@ static bool skb_gso_has_extension_hdr(const struct sk_buff *skb)
>  			 skb_inner_network_header_len(skb) != sizeof(struct ipv6hdr)));
>  }
>  
> +/*
> + * Does @skb fit the GSO limits of @dev?  The size limit depends on the L3
> + * protocol, which validate_xmit_vlan() replaces with the VLAN ethertype when
> + * it pushes the tag inside the skb, so look behind the tag.
> + */
>  static bool gso_within_device_limits(const struct sk_buff *skb,
>  				     const struct net_device *dev)
>  {
>  	return skb_shinfo(skb)->gso_segs <= READ_ONCE(dev->gso_max_segs) &&
> -	       skb->len < netif_get_gso_max_size(dev, skb->protocol);
> +	       skb->len < netif_get_gso_max_size(dev, vlan_get_protocol(skb));

If this change is needed, it is not new for this feature and should be
a separate commit.

> +static bool skb_can_gso_resegment(struct sk_buff *skb,
> +				  netdev_features_t features)
> +{
> +	__be16 protocol;
> +
> +	if (!net_gso_ok(features | NETIF_F_GSO_ROBUST,
> +			skb_shinfo(skb)->gso_type))
> +		return false;
> +
> +	if (!(features & NETIF_F_SG))
> +		return false;
> +
> +	protocol = skb_network_protocol(skb, NULL);
> +	if (!protocol || !can_checksum_protocol(features, protocol))
> +		return false;
> +
> +	/*
> +	 * The TCP frag-list path does not carry the bounded segment
> +	 * limit through skb_segment_list(). Keep bounded resegmentation
> +	 * on the regular skb path until that support is added.
> +	 */
> +	if (skb_has_frag_list(skb))
> +		return false;
> +
> +	return true;
> +}
> +
> +static unsigned int
> +skb_gso_resegment_max_segs(struct sk_buff *skb, struct net_device *dev,
> +			   netdev_features_t features)
> +{
> +	unsigned int mss = skb_shinfo(skb)->gso_size;
> +	unsigned int hdr_len, max_segs;
> +	unsigned int gso_max_size;
> +	struct tcphdr _tcph, *th;
> +
> +	gso_max_size = netif_get_gso_max_size(dev, vlan_get_protocol(skb));
> +
> +	if (!skb_is_gso(skb) || !skb_is_gso_tcp(skb) ||
> +	    skb->encapsulation || mss == GSO_BY_FRAGS ||
> +	    !skb_mac_header_was_set(skb) ||
> +	    !skb_transport_header_was_set(skb) ||
> +	    !skb_can_gso_resegment(skb, features))
> +		return 0;

Is this duplicating/extending skb_can_gso_resegment

> +
> +	th = skb_header_pointer(skb, skb_transport_offset(skb), sizeof(_tcph),
> +				&_tcph);
> +	if (!th || th->doff < sizeof(*th) / 4)
> +		return 0;
> +
> +	hdr_len = skb_transport_header(skb) - skb_mac_header(skb) +
> +		  th->doff * 4;
> +	if (gso_max_size <= hdr_len + mss)
> +		return 0;
> +
> +	/*
> +	 * gso_within_device_limits() accepts gso_segs == gso_max_segs but
> +	 * rejects skb->len >= gso_max_size, so only the size bound needs - 1.
> +	 */
> +	max_segs = (gso_max_size - hdr_len - 1) / mss;
> +	max_segs = min_t(unsigned int, max_segs,
> +			 READ_ONCE(dev->gso_max_segs));
> +
> +	return max_segs > 1 ? max_segs : 0;
> +}
> +
>  static int xmit_one(struct sk_buff *skb, struct net_device *dev,
>  		    struct netdev_queue *txq, bool more)
>  {
> @@ -4073,6 +4152,7 @@ static struct sk_buff *validate_xmit_unreadable_skb(struct sk_buff *skb,
>   */
>  static struct sk_buff *validate_xmit_skb(struct sk_buff *skb, struct net_device *dev, bool *again)
>  {
> +	unsigned int resegment_max_segs = 0;
>  	netdev_features_t features;
>  
>  	skb = validate_xmit_unreadable_skb(skb, dev);
> @@ -4088,10 +4168,29 @@ static struct sk_buff *validate_xmit_skb(struct sk_buff *skb, struct net_device
>  	if (unlikely(!skb))
>  		goto out_null;
>  
> -	if (netif_needs_gso(skb, features)) {
> +	/*
> +	 * An oversized skb loses its GSO feature bits and is segmented
> +	 * down to MSS sized skbs below.  A TCP skb can instead be split
> +	 * into GSO skbs which do fit the device, so keep the bits and
> +	 * bound the resegmentation.  The features computed without the
> +	 * limit checks say whether the device offloads the GSO type at
> +	 * all.
> +	 */
> +	if (skb_is_gso(skb) && skb_is_gso_tcp(skb) && !skb->encapsulation &&
> +	    !gso_within_device_limits(skb, dev)) {
> +		netdev_features_t offload = __netif_skb_features(skb, false);
> +
> +		resegment_max_segs =
> +			skb_gso_resegment_max_segs(skb, dev, offload);
> +		if (resegment_max_segs)
> +			features = offload;
> +	}
> +

This is a lot to put in the hot path for a rare use case. Consider how
to make this less expensive.

> +	if (resegment_max_segs || netif_needs_gso(skb, features)) {
>  		struct sk_buff *segs;
>  
> -		segs = skb_gso_segment(skb, features);
> +		segs = __skb_gso_segment(skb, features, true,
> +					 resegment_max_segs);
>  		if (IS_ERR(segs)) {
>  			goto out_kfree_skb;
>  		} else if (segs) {
> -- 
> 2.47.3
> 



  reply	other threads:[~2026-09-19 15:37 UTC|newest]

Thread overview: 27+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-18  8:46 [PATCH net-next v2 0/4] net: resegment oversized TCP GSO skbs Wang Zhan
2026-09-18  8:46 ` [PATCH net-next v2 1/4] net: core: factor out the GSO device limit check Wang Zhan
2026-09-18  8:46 ` [PATCH net-next v2 2/4] net: gso: support bounded TCP segmentation Wang Zhan
2026-09-19 15:35   ` Willem de Bruijn
2026-09-20 13:12     ` Wang Zhan
2026-09-21 20:36       ` Willem de Bruijn
2026-09-23  9:45         ` Wang Zhan
2026-09-23 16:42           ` Willem de Bruijn
2026-09-24  9:03             ` Wang Zhan
2026-09-21 21:07       ` Willem de Bruijn
2026-09-23 10:38         ` Wang Zhan
2026-09-23 16:44           ` Willem de Bruijn
2026-09-24  9:27             ` Wang Zhan
2026-09-21 20:50   ` netdev-bot+sashiko
2026-09-23 16:50   ` Willem de Bruijn
2026-09-24  9:09     ` Wang Zhan
2026-09-24 10:53   ` David Laight
2026-09-24 12:15     ` Wang Zhan
2026-09-24 14:09   ` Paolo Abeni
2026-09-25  7:45     ` Wang Zhan
2026-09-18  8:46 ` [PATCH net-next v2 3/4] net: core: resegment oversized TCP GSO skbs Wang Zhan
2026-09-19 15:37   ` Willem de Bruijn [this message]
2026-09-20 13:31     ` Wang Zhan
2026-09-24 14:02     ` Paolo Abeni
2026-09-21 20:50   ` netdev-bot+sashiko
2026-09-18  8:46 ` [PATCH net-next v2 4/4] net: net_test: add tests for bounded GSO segmentation Wang Zhan
2026-09-21 20:50   ` netdev-bot+sashiko

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=willemdebruijn.kernel.32be703ea17ab@gmail.com \
    --to=willemdebruijn.kernel@gmail.com \
    --cc=aconole@redhat.com \
    --cc=alice@isovalent.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=davem@davemloft.net \
    --cc=dev@openvswitch.org \
    --cc=echaudro@redhat.com \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=i.maximets@ovn.org \
    --cc=jasowangio@gmail.com \
    --cc=keyong.sun@smartx.com \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=ncardwell@google.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=wang.zhan@smartx.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox