Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* Re: [PATCH net-next v3] net: skb: isolate skb data area allocations into a separate bucket
       [not found] ` <04debe19-bbe8-4b5f-9668-753d1f97832d@redhat.com>
@ 2026-07-08 11:16   ` Pedro Falcato
  2026-07-08 13:27     ` Harry Yoo
  0 siblings, 1 reply; 4+ messages in thread
From: Pedro Falcato @ 2026-07-08 11:16 UTC (permalink / raw)
  To: Paolo Abeni
  Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Simon Horman,
	Jason Xing, Kuniyuki Iwashima, netdev, linux-kernel,
	linux-hardening, Kees Cook, linux-mm, Vlastimil Babka, Harry Yoo

On Wed, Jul 08, 2026 at 10:30:50AM +0200, Paolo Abeni wrote:
> On 7/2/26 7:07 PM, Pedro Falcato wrote:> @@ -586,6 +586,8 @@ struct
> sk_buff *napi_build_skb(void *data, unsigned int frag_size)
> >  }
> >  EXPORT_SYMBOL(napi_build_skb);
> >  
> > +static kmem_buckets *skb_data_buckets __ro_after_init;
> > +
> >  static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
> >  {
> >  	if (!gfp_pfmemalloc_allowed(flags))
> > @@ -593,7 +595,8 @@ static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
> >  	if (!obj_size)
> >  		return kmem_cache_alloc_node(net_hotdata.skb_small_head_cache,
> >  					     flags, node);
> > -	return kmalloc_node_track_caller(obj_size, flags, node);
> > +	return kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
> > +						    flags, node);
> 
> Sashiko noted that some drivers may require GFP_DMA buckets, and the
> above may break them:
> 
> https://sashiko.dev/#/patchset/20260702170728.168755-1-pfalcato%40suse.de

Oh, this is really awkward. Adding linux-mm and slab maintainers for input here.

Considering the current slab bucketing does not seem to duplicate DMA or
CGROUP caches, could it make sense to duplicate those as well? Otherwise we
could add a branch like:

if (gfp_flags & __GFP_DMA)
	/* use the global dma kmalloc caches */

> 
> >  }
> >  
> >  /*
> > @@ -634,7 +637,7 @@ static void *kmalloc_reserve(unsigned int *size, gfp_t flags, int node,
> >  	 * Try a regular allocation, when that fails and we're not entitled
> >  	 * to the reserves, fail.
> >  	 */
> > -	obj = kmalloc_node_track_caller(obj_size,
> > +	obj = kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
> >  					flags | __GFP_NOMEMALLOC | __GFP_NOWARN,
> >  					node);
> 
> Minor nit: checkpatch laments WRT brackets alignment.

Will fix, thanks.

-- 
Pedro


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH net-next v3] net: skb: isolate skb data area allocations into a separate bucket
  2026-07-08 11:16   ` [PATCH net-next v3] net: skb: isolate skb data area allocations into a separate bucket Pedro Falcato
@ 2026-07-08 13:27     ` Harry Yoo
  2026-07-15 11:07       ` Pedro Falcato
  0 siblings, 1 reply; 4+ messages in thread
From: Harry Yoo @ 2026-07-08 13:27 UTC (permalink / raw)
  To: Pedro Falcato, Paolo Abeni
  Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Simon Horman,
	Jason Xing, Kuniyuki Iwashima, netdev, linux-kernel,
	linux-hardening, Kees Cook, linux-mm, Vlastimil Babka



On 7/8/26 8:16 PM, Pedro Falcato wrote:
> On Wed, Jul 08, 2026 at 10:30:50AM +0200, Paolo Abeni wrote:
>> On 7/2/26 7:07 PM, Pedro Falcato wrote:> @@ -586,6 +586,8 @@ struct
>> sk_buff *napi_build_skb(void *data, unsigned int frag_size)
>>>  }
>>>  EXPORT_SYMBOL(napi_build_skb);
>>>  
>>> +static kmem_buckets *skb_data_buckets __ro_after_init;
>>> +
>>>  static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
>>>  {
>>>  	if (!gfp_pfmemalloc_allowed(flags))
>>> @@ -593,7 +595,8 @@ static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
>>>  	if (!obj_size)
>>>  		return kmem_cache_alloc_node(net_hotdata.skb_small_head_cache,
>>>  					     flags, node);
>>> -	return kmalloc_node_track_caller(obj_size, flags, node);
>>> +	return kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
>>> +						    flags, node);
>>
>> Sashiko noted that some drivers may require GFP_DMA buckets, and the
>> above may break them:
>>
>> https://sashiko.dev/#/patchset/20260702170728.168755-1-pfalcato%40suse.de
> 
> Oh, this is really awkward. Adding linux-mm and slab maintainers for input here.
> 
> Considering the current slab bucketing does not seem to duplicate DMA or
> CGROUP caches, could it make sense to duplicate those as well?

Could we specify what kmalloc types the user needs when creating
kmem_buckets and duplicate caches for the requested kmalloc types only?

> Otherwise we could add a branch like:
>
> if (gfp_flags & __GFP_DMA)
> 	/* use the global dma kmalloc caches */

-- 
Cheers,
Harry / Hyeonggon


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH net-next v3] net: skb: isolate skb data area allocations into a separate bucket
  2026-07-08 13:27     ` Harry Yoo
@ 2026-07-15 11:07       ` Pedro Falcato
  2026-07-16  2:10         ` Harry Yoo
  0 siblings, 1 reply; 4+ messages in thread
From: Pedro Falcato @ 2026-07-15 11:07 UTC (permalink / raw)
  To: Harry Yoo
  Cc: Paolo Abeni, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Simon Horman, Jason Xing, Kuniyuki Iwashima, netdev, linux-kernel,
	linux-hardening, Kees Cook, linux-mm, Vlastimil Babka

On Wed, Jul 08, 2026 at 10:27:54PM +0900, Harry Yoo wrote:
> 
> 
> On 7/8/26 8:16 PM, Pedro Falcato wrote:
> > On Wed, Jul 08, 2026 at 10:30:50AM +0200, Paolo Abeni wrote:
> >> On 7/2/26 7:07 PM, Pedro Falcato wrote:> @@ -586,6 +586,8 @@ struct
> >> sk_buff *napi_build_skb(void *data, unsigned int frag_size)
> >>>  }
> >>>  EXPORT_SYMBOL(napi_build_skb);
> >>>  
> >>> +static kmem_buckets *skb_data_buckets __ro_after_init;
> >>> +
> >>>  static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
> >>>  {
> >>>  	if (!gfp_pfmemalloc_allowed(flags))
> >>> @@ -593,7 +595,8 @@ static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
> >>>  	if (!obj_size)
> >>>  		return kmem_cache_alloc_node(net_hotdata.skb_small_head_cache,
> >>>  					     flags, node);
> >>> -	return kmalloc_node_track_caller(obj_size, flags, node);
> >>> +	return kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
> >>> +						    flags, node);
> >>
> >> Sashiko noted that some drivers may require GFP_DMA buckets, and the
> >> above may break them:
> >>
> >> https://sashiko.dev/#/patchset/20260702170728.168755-1-pfalcato%40suse.de
> > 
> > Oh, this is really awkward. Adding linux-mm and slab maintainers for input here.
> > 
> > Considering the current slab bucketing does not seem to duplicate DMA or
> > CGROUP caches, could it make sense to duplicate those as well?
> 
> Could we specify what kmalloc types the user needs when creating
> kmem_buckets and duplicate caches for the requested kmalloc types only?

Perhaps. But do the users themselves know? alloc_skb() allows users to specify
random __GFP flags. We're bound to see some random caller do
alloc_skb(__GFP_ACCOUNT) ;)

In all honesty, I'm not quite sure what the best way forward here is. The most
transparent way is to bucket those other kmalloc types as well, but that might
very trivially result in a lot more caches (and possibly memory usage) for no
great reason. So perhaps specifying caches might do.

-- 
Pedro


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH net-next v3] net: skb: isolate skb data area allocations into a separate bucket
  2026-07-15 11:07       ` Pedro Falcato
@ 2026-07-16  2:10         ` Harry Yoo
  0 siblings, 0 replies; 4+ messages in thread
From: Harry Yoo @ 2026-07-16  2:10 UTC (permalink / raw)
  To: Pedro Falcato
  Cc: Paolo Abeni, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Simon Horman, Jason Xing, Kuniyuki Iwashima, netdev, linux-kernel,
	linux-hardening, Kees Cook, linux-mm, Vlastimil Babka


[-- Attachment #1.1: Type: text/plain, Size: 2815 bytes --]



On 7/15/26 8:07 PM, Pedro Falcato wrote:
> On Wed, Jul 08, 2026 at 10:27:54PM +0900, Harry Yoo wrote:
>> On 7/8/26 8:16 PM, Pedro Falcato wrote:
>>> On Wed, Jul 08, 2026 at 10:30:50AM +0200, Paolo Abeni wrote:
>>>> On 7/2/26 7:07 PM, Pedro Falcato wrote:> @@ -586,6 +586,8 @@ struct
>>>> sk_buff *napi_build_skb(void *data, unsigned int frag_size)
>>>>>  }
>>>>>  EXPORT_SYMBOL(napi_build_skb);
>>>>>  
>>>>> +static kmem_buckets *skb_data_buckets __ro_after_init;
>>>>> +
>>>>>  static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
>>>>>  {
>>>>>  	if (!gfp_pfmemalloc_allowed(flags))
>>>>> @@ -593,7 +595,8 @@ static void *kmalloc_pfmemalloc(size_t obj_size, gfp_t flags, int node)
>>>>>  	if (!obj_size)
>>>>>  		return kmem_cache_alloc_node(net_hotdata.skb_small_head_cache,
>>>>>  					     flags, node);
>>>>> -	return kmalloc_node_track_caller(obj_size, flags, node);
>>>>> +	return kmem_buckets_alloc_node_track_caller(skb_data_buckets, obj_size,
>>>>> +						    flags, node);
>>>>
>>>> Sashiko noted that some drivers may require GFP_DMA buckets, and the
>>>> above may break them:
>>>>
>>>> https://sashiko.dev/#/patchset/20260702170728.168755-1-pfalcato%40suse.de
>>>
>>> Oh, this is really awkward. Adding linux-mm and slab maintainers for input here.
>>>
>>> Considering the current slab bucketing does not seem to duplicate DMA or
>>> CGROUP caches, could it make sense to duplicate those as well?
>>
>> Could we specify what kmalloc types the user needs when creating
>> kmem_buckets and duplicate caches for the requested kmalloc types only?
> 
> Perhaps. But do the users themselves know? alloc_skb() allows users to specify
> random __GFP flags. We're bound to see some random caller do
> alloc_skb(__GFP_ACCOUNT) ;)

Other users don't expose the buckets to drivers, so I thought only
alloc_skb() would create the buckets for each kmalloc type.

> In all honesty, I'm not quite sure what the best way forward here is. The most
> transparent way is to bucket those other kmalloc types as well, but that might
> very trivially result in a lot more caches (and possibly memory usage) for no
> great reason. So perhaps specifying caches might do.

Another direction could be merging those buckets.

If we want to protect kmalloc objects from user-controllable
allocations, can we create buckets for each kmalloc type during the boot
process and let the kmem_buckets users share them?

That doesn't sound like creating too many kmalloc caches, while
providing a decent separation. We already have two buckets users,
one w/ SLAB_ACCOUNT and the other w/o SLAB_ACCOUNT.

If you really want each bucket to have a separate set of caches, you
have to sacrifice some memory for security :)

-- 
Cheers,
Harry / Hyeonggon

[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 228 bytes --]

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-07-16  2:10 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
     [not found] <20260702170728.168755-1-pfalcato@suse.de>
     [not found] ` <04debe19-bbe8-4b5f-9668-753d1f97832d@redhat.com>
2026-07-08 11:16   ` [PATCH net-next v3] net: skb: isolate skb data area allocations into a separate bucket Pedro Falcato
2026-07-08 13:27     ` Harry Yoo
2026-07-15 11:07       ` Pedro Falcato
2026-07-16  2:10         ` Harry Yoo

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox