Netdev List
 help / color / mirror / Atom feed
From: Mariusz Klimek <maklimek97@gmail.com>
To: Paolo Abeni <pabeni@redhat.com>
Cc: Eric Dumazet <edumazet@google.com>,
	netdev@vger.kernel.org, andrew+netdev@lunn.ch,
	davem@davemloft.net, dsahern@kernel.org, idosch@nvidia.com,
	ncardwell@google.com, shuah@kernel.org, kuniyu@google.com,
	alice@isovalent.com, Jakub Kicinski <kuba@kernel.org>
Subject: Re: [PATCH net-next 00/10] tcp: support non-GSO jumbograms
Date: Mon, 5 Oct 2026 15:15:24 +0200	[thread overview]
Message-ID: <0b7aadda-c768-45bb-a322-24df31f8f823@gmail.com> (raw)
In-Reply-To: <f926259d-4612-44d2-9711-4a408e41e7f6@redhat.com>

On 6/10/26 22:10, Paolo Abeni wrote:

> On 6/10/26 4:16 PM, Mariusz Klimek wrote:
>> On 6/10/26 00:27, Jakub Kicinski wrote:
>>> On Tue, 9 Jun 2026 19:01:02 +0200 Mariusz Klimek wrote:
>>>> I understand your concerns and they're reasonable. I would still like to
>>>> proceed with this series, though. What can I do for you to consider this
>>>> series? Would it help to submit a few less-significant patches first before
>>>> resubmitting this patch series? Should I resubmit it as an RFC?
>>> AI tools changed how much we value code, but they did not change much
>>> about how one joins a community.
>> OK, I'll start with smaller patches to establish myself more
>> in the community. I am still interested in eventually getting this
>> series through as it can speed up workloads relevant to our company.
>> Do you think this patch could potentially be considered down the line?
> If you are strongly motivated to do progresses in this area, one thing
> you could look at is the root cause of the performance difference
> between the veth jumbogram and the same size veth big tcp test: I'm
> quite surprised some delta is visible there.
>
> Possibly perf can outline some bottle-neck in the big tcp path that
> could be ironed-out.
>
> Also it could be interesting if some devices could do H/W GRO based big TCP.
>
> /P
>
Hi Paolo, it's been a long time.

After some research I found out why jumbograms outperform BIG TCP in the 
iperf3
over veth test: it's because Nagle splits BIG TCP packets but not 
jumbograms.

BIG TCP assumes packets will eventually be sent over the network, and so 
applies
the Nagle test over the real MSS of 1500 bytes. If a send occurs before the
last ACK is received, Nagle splits the BIG TCP packet into two packets: one
whose size is a multiple of the MSS, and one that contains the remainder
(tcp_mss_split_point and tso_fragment). The remainder packet is queued 
and is
only sent once the ACK is received.

In the jumbogram case, the jumbogram's size is already smaller than or equal
to the MSS size, so Nagle doesn't split the packet. The ~5% performance gain
comes from avoiding the packet being split in two and having one packet
traverse the network stack instead of two.

A simple way to prove the above is disabling Nagle with the 
`TCP_NODELAY` flag
and seeing that the performance gap resolves. I also ran iperf3 on a kernel
with the `limit >= max_segs * tp->mss_cache`, `limit >= chunk`, and
`limit > tcp_max_tso_deferred_mss(tp) tp->mss_cache` checks removed from
`tcp_tso_should_defer`, and it also resolved the performance gap. Perhaps
tcp_tso_should_defer should be fixed to so that these BIG TCP packets aren't
split?

While debugging this issue I tried restricting the iperf3 client and server
processes to one CPU. That reduced the number of packets that were split to
around 10% (from 60%), but surprisingly the throughput also increased
significantly and was much larger than the other tests. It appears that
restricting the iperf3 client and server to the same CPU fixes something 
else
and provides a 12% improvement over the base case and a ~7% improvement over
jumbograms. I can't explain this improvement yet, but the reduction in 
packet
splitting is likely due to less racy execution that makes the Nagle case 
happen
less frequently.

Measurements (output of a python wrapper around iperf3):

[ROUNDS=84 MTU=1500 GSO_SIZE=524280]
IPERF3 BITRATE: 241.35 [95.0% CI 239.81 - 242.91] ±7.17 Gbits/sec

SERVER RX PACKETS SZ=862:    406030.32 [95.0% CI 374158.02 - 437902.63] 
±144141.51
SERVER RX PACKETS SZ=511310: 401084.13 [95.0% CI 368103.20 - 434065.07] 
±150101.57
SERVER RX PACKETS SZ=512086: 197816.40 [95.0% CI 162980.59 - 232652.22] 
±160523.86

CLIENT RX PACKETS SZ=86:     647034.43 [95.0% CI 636767.16 - 657301.70] 
±47311.70

[ROUNDS=284 MTU=524280 GSO_OFF]
IPERF3 BITRATE: 252.56 [95.0% CI 251.67 - 253.44] ±7.55 Gbits/sec

SERVER RX PACKETS SZ=512094: 616549.29 [95.0% CI 614399.99 - 618698.58] 
±18401.15

CLIENT RX PACKETS SZ=86:     616754.45 [95.0% CI 614599.64 - 618909.26] 
±18448.43

[ROUNDS=84 MTU=1500 GSO_SIZE=524280 NODELAY]
IPERF3 BITRATE: 252.96 [95.0% CI 251.36 - 254.55] ±7.36 Gbits/sec

SERVER RX PACKETS SZ=512086: 617684.17 [95.0% CI 613782.72 - 621585.61] 
±17977.91

CLIENT RX PACKETS SZ=86:     617721.74 [95.0% CI 613817.65 - 621625.83] 
±17990.11

[ROUNDS=84 MTU=524280 GSO_OFF NODELAY]
IPERF3 BITRATE: 252.21 [95.0% CI 250.54 - 253.88] ±7.69 Gbits/sec

SERVER RX PACKETS SZ=512094: 615726.77 [95.0% CI 611620.41 - 619833.14] 
±18922.17

CLIENT RX PACKETS SZ=86:     615912.33 [95.0% CI 611836.51 - 619988.15] 
±18781.42

[ROUNDS=125 MTU=1500 GSO_SIZE=524280 NCPUS=1]
IPERF3 BITRATE: 272.80 [95.0% CI 272.37 - 273.23] ±2.42 Gbits/sec

SERVER RX PACKETS SZ=511310: 6074.89 [95.0% CI 5988.63 - 6161.14] ±487.24
SERVER RX PACKETS SZ=512086: 652866.16 [95.0% CI 651750.82 - 653981.50] 
±6300.23
SERVER RX PACKETS SZ=512738: 6454.82 [95.0% CI 6355.94 - 6553.71] ±558.58

CLIENT RX PACKETS SZ=86:     655050.50 [95.0% CI 653940.67 - 656160.32] 
±6269.06

[ROUNDS=64 MTU=1500 GSO_SIZE=524280 DELAY_PACKETS_UNTIL_ACK]
IPERF3 BITRATE: 251.64 [95.0% CI 250.45 - 252.83] ±4.78 Gbits/sec

SERVER RX PACKETS SZ=512086: 614318.41 [95.0% CI 611402.89 - 617233.93] 
±11671.76

CLIENT RX PACKETS SZ=86:     614633.78 [95.0% CI 611697.43 - 617570.13] 
±11755.14

-- 
Mariusz K.



      reply	other threads:[~2026-10-05 13:15 UTC|newest]

Thread overview: 18+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-06-08 16:33 [PATCH net-next 00/10] tcp: support non-GSO jumbograms Mariusz Klimek
2026-06-08 13:07 ` [PATCH net-next 01/10] ipv6: do not fragment packets into jumbograms Mariusz Klimek
2026-06-08 13:07 ` [PATCH net-next 02/10] ipv6: allow route exceptions with MTUs above 65535 Mariusz Klimek
2026-06-08 13:07 ` [PATCH net-next 03/10] ipv6: add jumbo payload option to non-gso jumbograms Mariusz Klimek
2026-06-08 13:07 ` [PATCH net-next 04/10] tcp: decouple TSO segment length from MSS Mariusz Klimek
2026-06-08 13:07 ` [PATCH net-next 05/10] tcp: split jumbograms with urgent pointer correctly Mariusz Klimek
2026-06-08 13:07 ` [PATCH net-next 06/10] tcp: set MSS correctly for PMTU above 65535 Mariusz Klimek
2026-06-08 13:07 ` [PATCH net-next 07/10] veth: raise the max MTU " Mariusz Klimek
2026-06-08 13:07 ` [PATCH net-next 08/10] selftests/net: test sending TCP jumbograms over veth Mariusz Klimek
2026-06-08 13:07 ` [PATCH net-next 09/10] selftests/net: add test cases with MTU above 65535 to big_tcp.sh Mariusz Klimek
2026-06-08 13:07 ` [PATCH net-next 10/10] selftests/net: add jumbogram test case to msg_zerocopy.sh Mariusz Klimek
2026-06-09  1:15 ` [PATCH net-next 00/10] tcp: support non-GSO jumbograms Eric Dumazet
2026-06-09 17:01   ` Mariusz Klimek
2026-06-09 22:27     ` Jakub Kicinski
2026-06-10 14:16       ` Mariusz Klimek
2026-06-10 17:11         ` Eric Dumazet
2026-06-10 20:10         ` Paolo Abeni
2026-10-05 13:15           ` Mariusz Klimek [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=0b7aadda-c768-45bb-a322-24df31f8f823@gmail.com \
    --to=maklimek97@gmail.com \
    --cc=alice@isovalent.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=davem@davemloft.net \
    --cc=dsahern@kernel.org \
    --cc=edumazet@google.com \
    --cc=idosch@nvidia.com \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=ncardwell@google.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=shuah@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox