* [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket
@ 2026-10-06 9:20 Kees Cook
2026-10-06 9:20 ` [PATCH net-next v6 8/8] " Kees Cook
2026-10-08 21:25 ` [PATCH net-next v6 0/8] " Harry Yoo
0 siblings, 2 replies; 5+ messages in thread
From: Kees Cook @ 2026-10-06 9:20 UTC (permalink / raw)
To: Vlastimil Babka
Cc: Kees Cook, Harry Yoo, David S. Miller, Andrew Morton, Hao Li,
Christoph Lameter, David Rientjes, Roman Gushchin, Pedro Falcato,
Kuniyuki Iwashima, Christian Brauner, Jan Kara, Johannes Weiner,
Michal Hocko, Shakeel Butt, Muchun Song, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Willem de Bruijn,
Jason Xing, cgroups, netdev, linux-mm, linux-kernel,
linux-hardening
Hi!
This gets the buckets able to handle memcg (GFP_KERNEL_ACCOUNT) with
isolation (since it's common due to AF_UNIX), and GFP_DMA with fall back
(since it's rare). It gave me an excuse to build out bucket kunit tests
too, and that (and LLM review) found a couple other issues that needed
fixing too, including msg_msg allocations going uncharged to their memcg
when CONFIG_SLAB_BUCKETS=n (fixed in 2/8).
Harry, on your v4 question[1] about bucket users giving their own
alignment: I tried that in v5, but a set's allocations don't always
come from its own caches. With CONFIG_SLAB_BUCKETS=n, after a failed
kmem_buckets_create(), and for the DMA and reclaimable fallbacks, they
come from the general kmalloc caches, which can only give kmalloc()'s
alignment. So v6 goes back to mirroring the kmalloc cache's alignment,
and drops the ctor and flags arguments for the same reason. And the whole
exploration made me realize I had a completely wrong understanding of
how memcg worked. :P
The bulk of this is mm/slab, but the final patch is netdev, which Paolo
acked in v4, so I'm hoping this whole series can go via slab?
Thanks!
-Kees
v6:
- drop v5's 6/7 ("Let a bucket set handle __GFP_ACCOUNT") and its
kmem_buckets_create_types(): memcg charges each object in whatever
cache serves it, so accounted allocations can stay in a set's single
row of caches, and the fallback now covers only DMA, reclaimable, and
no-obj-ext allocations (Sashiko)
- 2/8: new: account msg_msg with GFP_KERNEL_ACCOUNT again; with
CONFIG_SLAB_BUCKETS=n it went uncharged, since its accounting lived in
SLAB_ACCOUNT on bucket caches that are not created (Sashiko)
- 3/8: new: drop the ctor and flags arguments from kmem_buckets_create();
neither reaches allocations that fall back to the general kmalloc
caches, and no caller needs them any more (Sashiko)
- 4/8: go back to v4's form: no alignment argument, and each bucket cache
takes the alignment of the kmalloc cache it mirrors, since the fallbacks
to kmalloc can give no other (Sashiko, Harry)
- 5/8: say in the teardown comment that cache sharing comes from kmalloc
rounding sizes up to a larger class (Sashiko)
- 6/8: drop the explicit alignment tests; check that each size lands in
the cache of the size kmalloc() rounds it up to, not just a big enough
one; and in the destroy test, assert on the allocation, skip when KFENCE
serves it, and tear the set down through its KUnit cleanup action
(Sashiko)
- 7/8: keep __GFP_ACCOUNT allocations in the set, document what an
allocation that falls back loses, and test the reclaimable fallback
(Sashiko)
- 8/8: create the skb_data set with kmem_buckets_create(), and make the
comment above kmalloc_reserve() name no allocator (Sashiko)
- v5..v6 diff: https://git.kernel.org/pub/scm/linux/kernel/git/kees/linux.git/diff/?id=dev/v7.3-rc2/skb-buckets/v6&id2=dev/v7.3-rc2/skb-buckets/v5
v5: https://lore.kernel.org/all/20261002231120.late.500-kees@kernel.org/
v4: https://lore.kernel.org/all/20260921075811.too.775-kees@kernel.org/
v3: https://lore.kernel.org/all/20260702170728.168755-1-pfalcato@suse.de/
[1] https://lore.kernel.org/all/arViR2Miz61-3fV4@thinkstation/
Kees Cook (7):
mm/slab: Mark the kmem_buckets_create() context as a Context: section
ipc, msg: Account msg_msg allocations with GFP_KERNEL_ACCOUNT again
mm/slab: Drop the ctor and flags arguments from kmem_buckets_create()
mm/slab: Give bucket caches the alignment of the caches they mirror
mm/slab: Add kmem_buckets_destroy()
mm/slab: Add tests for the existing kmem_buckets behaviour
mm/slab: Provide kmalloc type fallback for bucket allocations
Pedro Falcato (1):
net: skb: isolate skb data area allocations into a separate bucket
include/linux/slab.h | 6 +-
mm/slab.h | 19 ++-
ipc/msgutil.c | 8 +-
lib/tests/slub_kunit.c | 291 +++++++++++++++++++++++++++++++++++++++++
mm/slab_common.c | 74 ++++++++---
mm/util.c | 2 +-
net/core/skbuff.c | 10 +-
7 files changed, 378 insertions(+), 32 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH net-next v6 8/8] net: skb: isolate skb data area allocations into a separate bucket
2026-10-06 9:20 [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
@ 2026-10-06 9:20 ` Kees Cook
2026-10-08 21:25 ` [PATCH net-next v6 0/8] " Harry Yoo
1 sibling, 0 replies; 5+ messages in thread
From: Kees Cook @ 2026-10-06 9:20 UTC (permalink / raw)
To: Vlastimil Babka
Cc: Kees Cook, Pedro Falcato, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Willem de Bruijn,
Jason Xing, netdev, Kuniyuki Iwashima, linux-hardening, linux-mm,
linux-kernel
From: Pedro Falcato <pfalcato@suse.de>
SKB data area allocations (as done from alloc_skb()) use kmalloc().
These allocations can be variably sized and their contents can be more
or less controlled from userspace, which makes them useful for attackers
that want to overwrite a use-after-free'd object from the same kmalloc slab
(which often just requires the sizes to roughly match into the same kmalloc
bucket). [0] is an easy example of an exploit that uses netlink skb
allocation to target another similarly-sized accidentally freed object.
While other mitigations like CONFIG_RANDOM_KMALLOC_CACHES exist, these are
probabilistic. Use the existing kmem buckets API to further isolate these
allocations in a guaranteed fashion, when CONFIG_SLAB_BUCKETS=y.
AF_UNIX sets sk_allocation to GFP_KERNEL_ACCOUNT, and those skb data
areas, the ones most worth isolating, stay in the set, where memcg
charges them as it would in the general caches. GFP_DMA falls back to
the general caches, being passed to an skb allocator only by rare
devices.
Link: https://github.com/google/security-research/blob/master/pocs/linux/kernelctf/CVE-2023-4207_lts_cos_mitigation_2/docs/exploit.md [0]
Reviewed-by: Kees Cook <kees@kernel.org>
Signed-off-by: Pedro Falcato <pfalcato@suse.de>
Acked-by: Paolo Abeni <pabeni@redhat.com>
Signed-off-by: Kees Cook <kees@kernel.org>
---
net/core/skbuff.c | 10 +++++++---
1 file changed, 7 insertions(+), 3 deletions(-)
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 966af3beed94..6f6b5f4cb39f 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -586,6 +586,8 @@ struct sk_buff *napi_build_skb(void *data, unsigned int frag_size)
}
EXPORT_SYMBOL(napi_build_skb);
+static kmem_buckets *skb_data_buckets __ro_after_init;
+
static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
{
if (!gfp_pfmemalloc_allowed(flags))
@@ -593,11 +595,12 @@ static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
if (!obj_size)
return kmem_cache_alloc_node(net_hotdata.skb_small_head_cache,
flags, node);
- return kmalloc_node_track_caller(obj_size, flags, node);
+ return kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
+ flags, node);
}
/*
- * kmalloc_reserve is a wrapper around kmalloc_node_track_caller that tells
+ * kmalloc_reserve is a wrapper around a caller-tracked kmalloc that tells
* the caller if emergency pfmemalloc reserves are being used. If it is and
* the socket is later found to be SOCK_MEMALLOC then PFMEMALLOC reserves
* may be used. Otherwise, the packet data may be discarded until enough
@@ -634,7 +637,7 @@ static void *kmalloc_reserve(unsigned int *size, gfp_t flags, int node,
* Try a regular allocation, when that fails and we're not entitled
* to the reserves, fail.
*/
- obj = kmalloc_node_track_caller(obj_size,
+ obj = kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
flags | __GFP_NOMEMALLOC | __GFP_NOWARN,
node);
if (likely(obj))
@@ -5235,6 +5238,7 @@ void __init skb_init(void)
0,
SKB_SMALL_HEAD_HEADROOM,
NULL);
+ skb_data_buckets = kmem_buckets_create("skb_data", 0, INT_MAX);
skb_extensions_init();
}
--
2.55.0
^ permalink raw reply related [flat|nested] 5+ messages in thread
* Re: [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket
2026-10-06 9:20 [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
2026-10-06 9:20 ` [PATCH net-next v6 8/8] " Kees Cook
@ 2026-10-08 21:25 ` Harry Yoo
2026-10-09 6:36 ` Kees Cook
2026-10-09 7:02 ` Vlastimil Babka
1 sibling, 2 replies; 5+ messages in thread
From: Harry Yoo @ 2026-10-08 21:25 UTC (permalink / raw)
To: Kees Cook
Cc: Vlastimil Babka, David S. Miller, Andrew Morton, Hao Li,
Christoph Lameter, David Rientjes, Roman Gushchin, Pedro Falcato,
Kuniyuki Iwashima, Christian Brauner, Jan Kara, Johannes Weiner,
Michal Hocko, Shakeel Butt, Muchun Song, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Willem de Bruijn,
Jason Xing, cgroups, netdev, linux-mm, linux-kernel,
linux-hardening
On Tue, Oct 06, 2026 at 02:20:26AM -0700, Kees Cook wrote:
> Hi!
Hi Kees!
Was hoping to say hi to you at LPC but I missed the chance ;)
Maybe next time. Safe travels!
Uh, my mailbox stopped working for a few days as I forgot to renew
subscription. (Kiryl told me it's bouncing, thanks!) Hopefully I didn't
miss too much...
> This gets the buckets able to handle memcg (GFP_KERNEL_ACCOUNT) with
> isolation (since it's common due to AF_UNIX),
Cool!
> and GFP_DMA with fall back
> (since it's rare).
> It gave me an excuse to build out bucket kunit tests
> too, and that (and LLM review) found a couple other issues that needed
> fixing too, including msg_msg allocations going uncharged to their memcg
> when CONFIG_SLAB_BUCKETS=n (fixed in 2/8).
Oh.
> Harry, on your v4 question[1] about bucket users giving their own
> alignment: I tried that in v5, but a set's allocations don't always
> come from its own caches. With CONFIG_SLAB_BUCKETS=n, after a failed
> kmem_buckets_create(), and for the DMA and reclaimable fallbacks, they
> come from the general kmalloc caches, which can only give kmalloc()'s
> alignment.
>
> So v6 goes back to mirroring the kmalloc cache's alignment,
I might be missing something, but why is that a problem?
For kmem_buckets users, the reason* to specify alignment is because
they might need less strict alignment than kmalloc.
(*Perhaps it's nice to document that in the comment)
However, because kmem_buckets can fall back to kmalloc on e.g. kernels
w/o CONFIG_SLAB_BUCKETS, it should be fine to fall back. No?
Creating kmem_buckets with more strict alignment than
kmalloc doesn't make sense.
> and drops the ctor and flags arguments for the same reason.
Uh, for ctor and flags, yes. We can't have them in kmem_buckets.
> And the whole
> exploration made me realize I had a completely wrong understanding of
> how memcg worked. :P
>
> The bulk of this is mm/slab, but the final patch is netdev, which Paolo
> acked in v4, so I'm hoping this whole series can go via slab?
Going thorough slab/for-next sounds reasonable to me once it gets
some reviews.
Vlastimil?
--
Cheers,
Harry / Hyeonggon
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket
2026-10-08 21:25 ` [PATCH net-next v6 0/8] " Harry Yoo
@ 2026-10-09 6:36 ` Kees Cook
2026-10-09 7:02 ` Vlastimil Babka
1 sibling, 0 replies; 5+ messages in thread
From: Kees Cook @ 2026-10-09 6:36 UTC (permalink / raw)
To: Harry Yoo
Cc: Vlastimil Babka, David S. Miller, Andrew Morton, Hao Li,
Christoph Lameter, David Rientjes, Roman Gushchin, Pedro Falcato,
Kuniyuki Iwashima, Christian Brauner, Jan Kara, Johannes Weiner,
Michal Hocko, Shakeel Butt, Muchun Song, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Willem de Bruijn,
Jason Xing, cgroups, netdev, linux-mm, linux-kernel,
linux-hardening
On Thu, Oct 08, 2026 at 11:25:13PM +0200, Harry Yoo wrote:
> On Tue, Oct 06, 2026 at 02:20:26AM -0700, Kees Cook wrote:
> > Hi!
>
> Hi Kees!
>
> Was hoping to say hi to you at LPC but I missed the chance ;)
> Maybe next time. Safe travels!
Hi! Yes, I kept trying to find you and Vlastimil but it never worked out.
LPC is a non-stop hallway track usually. :) I will try again next year!
> > So v6 goes back to mirroring the kmalloc cache's alignment,
>
> I might be missing something, but why is that a problem?
>
> For kmem_buckets users, the reason* to specify alignment is because
> they might need less strict alignment than kmalloc.
>
> (*Perhaps it's nice to document that in the comment)
>
> However, because kmem_buckets can fall back to kmalloc on e.g. kernels
> w/o CONFIG_SLAB_BUCKETS, it should be fine to fall back. No?
>
> Creating kmem_buckets with more strict alignment than
> kmalloc doesn't make sense.
>
> > and drops the ctor and flags arguments for the same reason.
>
> Uh, for ctor and flags, yes. We can't have them in kmem_buckets.
Yeah, and given that these two, I'd just prefer to keep it a direct
mirror for alignment too and not allow for any configurability here:
they are supposed to be direct stand-ins for the general cache.
> Going thorough slab/for-next sounds reasonable to me once it gets
> some reviews.
Thanks!
-Kees
--
Kees Cook
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket
2026-10-08 21:25 ` [PATCH net-next v6 0/8] " Harry Yoo
2026-10-09 6:36 ` Kees Cook
@ 2026-10-09 7:02 ` Vlastimil Babka
1 sibling, 0 replies; 5+ messages in thread
From: Vlastimil Babka @ 2026-10-09 7:02 UTC (permalink / raw)
To: Harry Yoo, Kees Cook
Cc: Vlastimil Babka, David S. Miller, Andrew Morton, Hao Li,
Christoph Lameter, David Rientjes, Roman Gushchin, Pedro Falcato,
Kuniyuki Iwashima, Christian Brauner, Jan Kara, Johannes Weiner,
Michal Hocko, Shakeel Butt, Muchun Song, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Willem de Bruijn,
Jason Xing, cgroups, netdev, linux-mm, linux-kernel,
linux-hardening
On October 8, 2026 11:25:13 PM GMT+02:00, Harry Yoo <harry@kernel.org> wrote:
>On Tue, Oct 06, 2026 at 02:20:26AM -0700, Kees Cook wrote:
>> Hi!
>
>Hi Kees!
>
>Was hoping to say hi to you at LPC but I missed the chance ;)
>Maybe next time. Safe travels!
>
>Uh, my mailbox stopped working for a few days as I forgot to renew
>subscription. (Kiryl told me it's bouncing, thanks!) Hopefully I didn't
>miss too much...
>
>> This gets the buckets able to handle memcg (GFP_KERNEL_ACCOUNT) with
>> isolation (since it's common due to AF_UNIX),
>
>Cool!
>
>> and GFP_DMA with fall back
>> (since it's rare).
>
>> It gave me an excuse to build out bucket kunit tests
>> too, and that (and LLM review) found a couple other issues that needed
>> fixing too, including msg_msg allocations going uncharged to their memcg
>> when CONFIG_SLAB_BUCKETS=n (fixed in 2/8).
>
>Oh.
>
>> Harry, on your v4 question[1] about bucket users giving their own
>> alignment: I tried that in v5, but a set's allocations don't always
>> come from its own caches. With CONFIG_SLAB_BUCKETS=n, after a failed
>> kmem_buckets_create(), and for the DMA and reclaimable fallbacks, they
>> come from the general kmalloc caches, which can only give kmalloc()'s
>> alignment.
>>
>> So v6 goes back to mirroring the kmalloc cache's alignment,
>
>I might be missing something, but why is that a problem?
>
>For kmem_buckets users, the reason* to specify alignment is because
>they might need less strict alignment than kmalloc.
>
>(*Perhaps it's nice to document that in the comment)
>
>However, because kmem_buckets can fall back to kmalloc on e.g. kernels
>w/o CONFIG_SLAB_BUCKETS, it should be fine to fall back. No?
>
>Creating kmem_buckets with more strict alignment than
>kmalloc doesn't make sense.
>
>> and drops the ctor and flags arguments for the same reason.
>
>Uh, for ctor and flags, yes. We can't have them in kmem_buckets.
>
>> And the whole
>> exploration made me realize I had a completely wrong understanding of
>> how memcg worked. :P
>>
>> The bulk of this is mm/slab, but the final patch is netdev, which Paolo
>> acked in v4, so I'm hoping this whole series can go via slab?
>
>Going thorough slab/for-next sounds reasonable to me once it gets
>some reviews.
>
>Vlastimil?
Sure! If you review and feel it's ready to be added there, please do so. I couldn't yet due to conferencing, should be able next week.
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-10-09 7:02 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-06 9:20 [PATCH net-next v6 0/8] net: skb: isolate skb data area allocations into a separate bucket Kees Cook
2026-10-06 9:20 ` [PATCH net-next v6 8/8] " Kees Cook
2026-10-08 21:25 ` [PATCH net-next v6 0/8] " Harry Yoo
2026-10-09 6:36 ` Kees Cook
2026-10-09 7:02 ` Vlastimil Babka
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox