Netdev List
 help / color / mirror / Atom feed
From: Wang Zhan <wang.zhan@smartx.com>
To: netdev@vger.kernel.org
Cc: davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
	pabeni@redhat.com, horms@kernel.org, keyong.sun@smartx.com,
	Ilya Maximets <i.maximets@ovn.org>,
	Aaron Conole <aconole@redhat.com>,
	Eelco Chaudron <echaudro@redhat.com>,
	dev@openvswitch.org, Andrew Lunn <andrew+netdev@lunn.ch>,
	Jason Wang <jasowangio@gmail.com>,
	Willem de Bruijn <willemdebruijn.kernel@gmail.com>,
	Neal Cardwell <ncardwell@google.com>,
	Kuniyuki Iwashima <kuniyu@google.com>,
	Alice Mikityanska <alice@isovalent.com>,
	Wang Zhan <wang.zhan@smartx.com>
Subject: [PATCH net-next v2 0/4] net: resegment oversized TCP GSO skbs
Date: Fri, 18 Sep 2026 16:46:47 +0800	[thread overview]
Message-ID: <20260918084651.3022878-1-wang.zhan@smartx.com> (raw)

BIG TCP is negotiated per netdevice, so BIG TCP and non-BIG TCP ports can
coexist in one path. When a BIG TCP packet goes to a port without BIG TCP
support, it loses its GSO state and the stack segments it into individual
MSS sized packets, so the port TSO is left unused.

This series cuts an oversized unencapsulated TCP GSO skb into GSO skbs
which fit the device limits instead of single segments, so the per-segment
work stays on the device TSO. On a veth -> bridge -> TAP -> guest
virtio-net path, with BIG TCP enabled on the veth endpoints and left off
in the guest, a single iperf3 TCP flow, six alternating runs per state
(-t 15 -O 5, fixed CPU affinity and port tuple):

  protocol  no BIG TCP   mixed, no reseg  mixed, resegmented
  TCP/IPv4  51.550 Gbps  15.850 Gbps      52.617 Gbps
  TCP/IPv6  52.050 Gbps  15.783 Gbps      51.933 Gbps

The middle column comes from the same kernel with the bounded path
disabled for the comparison. A BIG TCP hop which feeds a 64 KiB hop loses
69% of the throughput of a path which never enables BIG TCP, and bounded
resegmentation recovers it.

The new path is taken only when the skb is a plain TCP GSO skb which
exceeds gso_max_size or gso_max_segs, the device offloads that GSO type
and has scatter-gather and checksum offload for the protocol, and the
bound leaves room for at least two MSS segments per output skb.
Encapsulated skbs, frag-list skbs, GSO types the device cannot offload and
bounds below two segments keep the existing full segmentation path
unchanged.

The output obeys the GSO feature and limit contract the device already
advertises - gso_size stays at the MSS, gso_segs stays within the bound,
and the frame length stays below gso_max_size - so this needs no new UAPI,
no device state and no driver change, and it is applied automatically.
I considered a per-device switch and decided against it: it would spend
netlink ABI and net_device state on a decision the stack can make from
capabilities the device already advertises, and it would have to be
configured on every device created later.

Patch layout:

  [1/4] factor the device GSO limit check out of gso_features_check()
  [2/4] let the GSO engine bound the MSS segments per output skb
  [3/4] apply that bound to oversized TCP GSO skbs in the TX path
  [4/4] KUnit coverage for the bound, the device limits and the TCP path

1/4 is a preparation patch with no functional change.

v2:
- patch 2 and patch 4: fix the lines over 80 columns reported by checkpatch
- patch 4: use KUNIT_ASSERT_TRUE() for the __be16 check, EQ warns in sparse
- Cc the full get_maintainer list (patch 2 was missing dev@openvswitch.org)
v1: https://lore.kernel.org/20260917063854.2011613-1-wang.zhan@smartx.com/

Wang Zhan (4):
  net: core: factor out the GSO device limit check
  net: gso: support bounded TCP segmentation
  net: core: resegment oversized TCP GSO skbs
  net: net_test: add tests for bounded GSO segmentation

 drivers/net/tap.c          |   3 +-
 include/linux/netdevice.h  |   4 +-
 include/net/gso.h          |   6 +-
 include/net/udp.h          |   2 +-
 net/core/dev.c             | 123 +++++++++++++++++--
 net/core/gso.c             |   5 +-
 net/core/net_test.c        | 243 +++++++++++++++++++++++++++++++++++++
 net/core/skbuff.c          |  14 ++-
 net/ipv4/tcp_offload.c     |   3 +-
 net/openvswitch/datapath.c |   2 +-
 10 files changed, 383 insertions(+), 22 deletions(-)


base-commit: 4bb9710c6a68d35207f123aef55dcd50e7195ec5
-- 
2.47.3

             reply	other threads:[~2026-09-18  8:47 UTC|newest]

Thread overview: 27+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-18  8:46 Wang Zhan [this message]
2026-09-18  8:46 ` [PATCH net-next v2 1/4] net: core: factor out the GSO device limit check Wang Zhan
2026-09-18  8:46 ` [PATCH net-next v2 2/4] net: gso: support bounded TCP segmentation Wang Zhan
2026-09-19 15:35   ` Willem de Bruijn
2026-09-20 13:12     ` Wang Zhan
2026-09-21 20:36       ` Willem de Bruijn
2026-09-23  9:45         ` Wang Zhan
2026-09-23 16:42           ` Willem de Bruijn
2026-09-24  9:03             ` Wang Zhan
2026-09-21 21:07       ` Willem de Bruijn
2026-09-23 10:38         ` Wang Zhan
2026-09-23 16:44           ` Willem de Bruijn
2026-09-24  9:27             ` Wang Zhan
2026-09-21 20:50   ` netdev-bot+sashiko
2026-09-23 16:50   ` Willem de Bruijn
2026-09-24  9:09     ` Wang Zhan
2026-09-24 10:53   ` David Laight
2026-09-24 12:15     ` Wang Zhan
2026-09-24 14:09   ` Paolo Abeni
2026-09-25  7:45     ` Wang Zhan
2026-09-18  8:46 ` [PATCH net-next v2 3/4] net: core: resegment oversized TCP GSO skbs Wang Zhan
2026-09-19 15:37   ` Willem de Bruijn
2026-09-20 13:31     ` Wang Zhan
2026-09-24 14:02     ` Paolo Abeni
2026-09-21 20:50   ` netdev-bot+sashiko
2026-09-18  8:46 ` [PATCH net-next v2 4/4] net: net_test: add tests for bounded GSO segmentation Wang Zhan
2026-09-21 20:50   ` netdev-bot+sashiko

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260918084651.3022878-1-wang.zhan@smartx.com \
    --to=wang.zhan@smartx.com \
    --cc=aconole@redhat.com \
    --cc=alice@isovalent.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=davem@davemloft.net \
    --cc=dev@openvswitch.org \
    --cc=echaudro@redhat.com \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=i.maximets@ovn.org \
    --cc=jasowangio@gmail.com \
    --cc=keyong.sun@smartx.com \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=ncardwell@google.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=willemdebruijn.kernel@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox