From: "Vlastimil Babka (SUSE)" <vbabka@kernel.org>
To: Eric Dumazet <edumazet@google.com>
Cc: "David S . Miller" <davem@davemloft.net>,
Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
Simon Horman <horms@kernel.org>,
Kuniyuki Iwashima <kuniyu@google.com>,
netdev@vger.kernel.org, eric.dumazet@gmail.com,
David Rientjes <rientjes@google.com>,
Roman Gushchin <roman.gushchin@linux.dev>,
Harry Yoo <harry.yoo@oracle.com>, Hao Li <hao.li@linux.dev>
Subject: Re: [PATCH net-next] net: adopt SLUB sheaves for skbuff_small_head
Date: Thu, 1 Oct 2026 12:37:37 +0200 [thread overview]
Message-ID: <aff0a6c6-ab34-4578-850b-86e7d94deb7a@kernel.org> (raw)
In-Reply-To: <CANn89iKH92jjRDD_jDNeioSy9U5mxKAYeyBJ3hHAFR_z4+wccA@mail.gmail.com>
On 3/1/26 17:30, Eric Dumazet wrote:
> On Sun, Mar 1, 2026 at 12:24 PM Vlastimil Babka <vbabka@suse.com> wrote:
>>
>> On 2/28/26 15:12, Eric Dumazet wrote:
>> > skbuff_small_head is used both on receive and send paths,
>> > serving potentially 80 million allocations and frees per second.
>> >
>> > Tuning it on large servers has been problematic, especially
>> > on AMD Turins platforms, where "lock cmpxch16b" latency can
>> > be over 30,000 cycles.
>>
>> Huh, really? That sounds insane. Any pointers about that?
>
> Yes, obviously on semi-contended cache lines.
>
>>
>> > Switching to SLUB sheaves fixes the issue nicely.
>> >
>> > tcp_rr benchmark with 10,000 flows goes from 25 Mpps to 40 Mpps
>> > on AMD Turin.
>> >
>> > Other platforms show benefits with tcp_rr with more than 30,000
>> > flows.
>>
>> That's nice, thanks!
>>
>> However I must point out some caveates. I assume you did this on 6.19, where
>> sheaves are still opt-in. But also, when you opt-in, the pre-existing
>> per-cpu caching layer of percpu slab and percpu partial slabs is also still
>> there, so effectively the amount of percpu cached slab objects increase,
>> which can be the main performance difference for some workloads, and not the
>> difference between sheaves and percpu (partial) slabs implementation.
>>
>
> Tests are on 6.18 LTS kernel, on which our latest production kernel is based.
>
>> Note: but hopefully for your workload it's really the implementation.
>> "(lock) cmpxch16b" should be avoided, until you start freeing NUMA-remote
>> (to the freeing cpu) objects in significant volumes.
>
> Right, __slab_free() is absolutely not 'slow path' when we have
> ~80,000 in-flight objects
> on a 512 cpu host.
>
>>
>> In 7.0-rc1 sheaves are enabled for every cache automatically, and cpu
>> (partial) caches are gone completely. Their size is calculated to roughly
>> match the average amount of percpu caching the old scheme achieved (but that
>> effectively depended on the workload too, so can't be exactly translated)
>> and the result is visible in /sys/kernel/slab/$cache/sheaf_capacity
>> the args.sheaf_capacity can override that automatic sizing, if the specified
>> one is larger.
>>
>
> Nice, I did not know that (I am not following lkml traffic)
>
>> So what I would suggest is checking the performance betwen 6.19 and 7.0-rc1
>> without this patch (hope there won't be any other factors in the upgrade
>> influencing this much), noting the auto-calculated capacity. If it still
>> looks good, you don't need to do anything, otherwise you can try making the
>> capacity larger and see what happens.
>
> I can not test this yet using 7.0-rc1.
>
> I guess we will carry this patch privately, and will come back in a
> few months when
> I can get our infra ready.
Hi, just wondering if since then you were able to test a 7.0+ kernel and
thus see if the refactoring of all caches to use sheaves had the same
benefit as your patch on top of 6.18?
Thanks,
Vlastimil
> Thanks.
prev parent reply other threads:[~2026-10-01 10:37 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-02-28 14:12 [PATCH net-next] net: adopt SLUB sheaves for skbuff_small_head Eric Dumazet
2026-02-28 19:51 ` Kuniyuki Iwashima
2026-03-01 8:32 ` Jason Xing
2026-03-01 11:24 ` Vlastimil Babka
2026-03-01 16:30 ` Eric Dumazet
2026-10-01 10:37 ` Vlastimil Babka (SUSE) [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aff0a6c6-ab34-4578-850b-86e7d94deb7a@kernel.org \
--to=vbabka@kernel.org \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=eric.dumazet@gmail.com \
--cc=hao.li@linux.dev \
--cc=harry.yoo@oracle.com \
--cc=horms@kernel.org \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=rientjes@google.com \
--cc=roman.gushchin@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox