BPF List
 help / color / mirror / Atom feed
* [PATCH bpf-next] bpf: Speed up htab hashing for u32/u64 keys
@ 2026-09-21 15:28 Mykyta Yatsenko
  2026-09-21 16:45 ` Anton Protopopov
       [not found] ` <179000542526.3934731.994864591824564797@gmail.com>
  0 siblings, 2 replies; 4+ messages in thread
From: Mykyta Yatsenko @ 2026-09-21 15:28 UTC (permalink / raw)
  To: bpf, ast, andrii, daniel, kernel-team, eddyz87, memxor; +Cc: Mykyta Yatsenko

From: Mykyta Yatsenko <yatsenko@meta.com>

Four- and eight-byte keys are common enough to warrant avoiding the
generic jhash2() path. Fold JHASH_INITVAL and the key length into
hashrnd when allocating maps with these key sizes, then use
__jhash_nwords() during lookup.

This preserves the jhash2() result while avoiding its length handling
and repeated state setup.

Measure this with six runs of:

  ./bench -w3 -d10 -a bpf-hashmap-lookup \
      --key_size KEY_SIZE --max_entries MAX_ENTRIES \
      --nr_entries NR_ENTRIES --nr_loops NR_LOOPS \
      --map_flags 0x40

inside a vng guest with two vCPUs and 4 GiB of RAM. Keep each map 50%
full with these workloads:

             max entries   entries       loops
  small              512       256   8,388,608
  medium          10,000     5,000   8,000,000
  large          100,000    50,000   8,000,000

Mean throughput in million lookups per second is:

  Baseline:
                         key size (bytes)
                      1        4        8       10
    small         87.55    85.65   102.20    85.74
    medium       120.88    79.02    89.12    76.52
    large        126.61    52.28    57.78    47.54

  Optimized:
                         key size (bytes)
                      1        4        8       10
    small         85.97   106.06   129.24    84.47
    medium       115.97    91.30   101.41    74.17
    large        120.27    58.16    62.03    45.99

  Change:
                         key size (bytes)
                      1        4        8       10
    small         -1.80%  +23.83%  +26.46%   -1.48%
    medium        -4.06%  +15.54%  +13.79%   -3.06%
    large         -5.00%  +11.24%   +7.36%   -3.26%

The targeted four-byte lookups improve by 11-24%, and eight-byte
lookups improve by 7-26%. Untargeted one- and ten-byte lookups are
1-5% slower, consistent with the extra key-size checks on those paths.

For bigger key sizes the overhead of optimized key size checks
should be less visible.

Signed-off-by: Mykyta Yatsenko <yatsenko@meta.com>
---
 kernel/bpf/hashtab.c | 21 +++++++++++++++++++--
 1 file changed, 19 insertions(+), 2 deletions(-)

diff --git a/kernel/bpf/hashtab.c b/kernel/bpf/hashtab.c
index 6f331c80130d..4f1dbd6ebc07 100644
--- a/kernel/bpf/hashtab.c
+++ b/kernel/bpf/hashtab.c
@@ -612,6 +612,13 @@ static struct bpf_map *htab_map_alloc(union bpf_attr *attr)
 		htab->hashrnd = 0;
 	else
 		htab->hashrnd = get_random_u32();
+	/*
+	 * Fold the jhash constant and key length into hashrnd once for the
+	 * fixed-size fast paths instead of doing it on every lookup.
+	 */
+	if (htab->map.key_size == sizeof(u32) ||
+	    htab->map.key_size == sizeof(u64))
+		htab->hashrnd += JHASH_INITVAL + htab->map.key_size;
 
 	htab_init_buckets(htab);
 
@@ -679,9 +686,19 @@ static struct bpf_map *htab_map_alloc(union bpf_attr *attr)
 
 static inline u32 htab_map_hash(const void *key, u32 key_len, u32 hashrnd)
 {
-	if (likely(key_len % 4 == 0))
+	const u32 *k = key;
+	u32 b;
+
+	if (key_len == sizeof(u32))
+		b = 0;
+	else if (key_len == sizeof(u64))
+		b = k[1];
+	else if (likely(key_len % 4 == 0))
 		return jhash2(key, key_len / 4, hashrnd);
-	return jhash(key, key_len, hashrnd);
+	else
+		return jhash(key, key_len, hashrnd);
+
+	return __jhash_nwords(k[0], b, 0, hashrnd);
 }
 
 static inline struct bucket *__select_bucket(struct bpf_htab *htab, u32 hash)

---
base-commit: 6849a44e5946ebf646e82c6a0296840d2b3bda5f
change-id: 20260916-hashtab_fast_hashfn-9e52b3c2276e

Best regards,
--  
Mykyta Yatsenko <yatsenko@meta.com>


^ permalink raw reply related	[flat|nested] 4+ messages in thread

* Re: [PATCH bpf-next] bpf: Speed up htab hashing for u32/u64 keys
  2026-09-21 15:28 [PATCH bpf-next] bpf: Speed up htab hashing for u32/u64 keys Mykyta Yatsenko
@ 2026-09-21 16:45 ` Anton Protopopov
       [not found] ` <179000542526.3934731.994864591824564797@gmail.com>
  1 sibling, 0 replies; 4+ messages in thread
From: Anton Protopopov @ 2026-09-21 16:45 UTC (permalink / raw)
  To: Mykyta Yatsenko
  Cc: bpf, ast, andrii, daniel, kernel-team, eddyz87, memxor,
	Mykyta Yatsenko

On 26/09/21 08:28AM, Mykyta Yatsenko wrote:
> From: Mykyta Yatsenko <yatsenko@meta.com>
> 
> Four- and eight-byte keys are common enough to warrant avoiding the
> generic jhash2() path. Fold JHASH_INITVAL and the key length into
> hashrnd when allocating maps with these key sizes, then use
> __jhash_nwords() during lookup.
> 
> This preserves the jhash2() result while avoiding its length handling
> and repeated state setup.
> 
> Measure this with six runs of:
> 
>   ./bench -w3 -d10 -a bpf-hashmap-lookup \
>       --key_size KEY_SIZE --max_entries MAX_ENTRIES \
>       --nr_entries NR_ENTRIES --nr_loops NR_LOOPS \
>       --map_flags 0x40
> 
> inside a vng guest with two vCPUs and 4 GiB of RAM. Keep each map 50%
> full with these workloads:
> 
>              max entries   entries       loops
>   small              512       256   8,388,608
>   medium          10,000     5,000   8,000,000
>   large          100,000    50,000   8,000,000
> 
> Mean throughput in million lookups per second is:
> 
>   Baseline:
>                          key size (bytes)
>                       1        4        8       10
>     small         87.55    85.65   102.20    85.74
>     medium       120.88    79.02    89.12    76.52
>     large        126.61    52.28    57.78    47.54
> 
>   Optimized:
>                          key size (bytes)
>                       1        4        8       10
>     small         85.97   106.06   129.24    84.47
>     medium       115.97    91.30   101.41    74.17
>     large        120.27    58.16    62.03    45.99
> 
>   Change:
>                          key size (bytes)
>                       1        4        8       10
>     small         -1.80%  +23.83%  +26.46%   -1.48%
>     medium        -4.06%  +15.54%  +13.79%   -3.06%
>     large         -5.00%  +11.24%   +7.36%   -3.26%
> 
> The targeted four-byte lookups improve by 11-24%, and eight-byte
> lookups improve by 7-26%. Untargeted one- and ten-byte lookups are
> 1-5% slower, consistent with the extra key-size checks on those paths.
> 
> For bigger key sizes the overhead of optimized key size checks
> should be less visible.
> 
> Signed-off-by: Mykyta Yatsenko <yatsenko@meta.com>
> ---
>  kernel/bpf/hashtab.c | 21 +++++++++++++++++++--
>  1 file changed, 19 insertions(+), 2 deletions(-)
> 
> diff --git a/kernel/bpf/hashtab.c b/kernel/bpf/hashtab.c
> index 6f331c80130d..4f1dbd6ebc07 100644
> --- a/kernel/bpf/hashtab.c
> +++ b/kernel/bpf/hashtab.c
> @@ -612,6 +612,13 @@ static struct bpf_map *htab_map_alloc(union bpf_attr *attr)
>  		htab->hashrnd = 0;
>  	else
>  		htab->hashrnd = get_random_u32();
> +	/*
> +	 * Fold the jhash constant and key length into hashrnd once for the
> +	 * fixed-size fast paths instead of doing it on every lookup.
> +	 */
> +	if (htab->map.key_size == sizeof(u32) ||
> +	    htab->map.key_size == sizeof(u64))
> +		htab->hashrnd += JHASH_INITVAL + htab->map.key_size;
>  
>  	htab_init_buckets(htab);
>  
> @@ -679,9 +686,19 @@ static struct bpf_map *htab_map_alloc(union bpf_attr *attr)
>  
>  static inline u32 htab_map_hash(const void *key, u32 key_len, u32 hashrnd)
>  {
> -	if (likely(key_len % 4 == 0))
> +	const u32 *k = key;
> +	u32 b;
> +
> +	if (key_len == sizeof(u32))
> +		b = 0;
> +	else if (key_len == sizeof(u64))
> +		b = k[1];
> +	else if (likely(key_len % 4 == 0))
>  		return jhash2(key, key_len / 4, hashrnd);
> -	return jhash(key, key_len, hashrnd);
> +	else
> +		return jhash(key, key_len, hashrnd);
> +
> +	return __jhash_nwords(k[0], b, 0, hashrnd);
>  }
>  
>  static inline struct bucket *__select_bucket(struct bpf_htab *htab, u32 hash)
> 
> ---
> base-commit: 6849a44e5946ebf646e82c6a0296840d2b3bda5f
> change-id: 20260916-hashtab_fast_hashfn-9e52b3c2276e
> 
> Best regards,
> --  
> Mykyta Yatsenko <yatsenko@meta.com>
> 

Nice!

Acked-by: Anton Protopopov <a.s.protopopov@gmail.com>

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH bpf-next] bpf: Speed up htab hashing for u32/u64 keys
       [not found] ` <179000542526.3934731.994864591824564797@gmail.com>
@ 2026-09-21 17:00   ` Alexei Starovoitov
  2026-09-21 17:10     ` Mykyta Yatsenko
  0 siblings, 1 reply; 4+ messages in thread
From: Alexei Starovoitov @ 2026-09-21 17:00 UTC (permalink / raw)
  To: Mykyta Yatsenko, bpf, andrii, daniel, kernel-team, eddyz87,
	memxor
  Cc: Mykyta Yatsenko

On Mon Sep 21, 2026 at 3:43 PM UTC, Alexei Starovoitov wrote:
> >  static inline u32 htab_map_hash(const void *key, u32 key_len, u32 hashrnd)
> >  {
> > -	if (likely(key_len % 4 == 0))
> > +	const u32 *k = key;
> > +	u32 b;
> > +
> > +	if (key_len == sizeof(u32))
> > +		b = 0;
> > +	else if (key_len == sizeof(u64))
> > +		b = k[1];
> > +	else if (likely(key_len % 4 == 0))
> >  		return jhash2(key, key_len / 4, hashrnd);

Nice improvement, but key_size is known at verification time.
Let's do a step further and specialize htab_map_gen_lookup
for small key sizes?


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH bpf-next] bpf: Speed up htab hashing for u32/u64 keys
  2026-09-21 17:00   ` Alexei Starovoitov
@ 2026-09-21 17:10     ` Mykyta Yatsenko
  0 siblings, 0 replies; 4+ messages in thread
From: Mykyta Yatsenko @ 2026-09-21 17:10 UTC (permalink / raw)
  To: Alexei Starovoitov, bpf, andrii, daniel, kernel-team, eddyz87,
	memxor
  Cc: Mykyta Yatsenko



On 9/21/26 6:00 PM, Alexei Starovoitov wrote:
> On Mon Sep 21, 2026 at 3:43 PM UTC, Alexei Starovoitov wrote:
>>>  static inline u32 htab_map_hash(const void *key, u32 key_len, u32 hashrnd)
>>>  {
>>> -	if (likely(key_len % 4 == 0))
>>> +	const u32 *k = key;
>>> +	u32 b;
>>> +
>>> +	if (key_len == sizeof(u32))
>>> +		b = 0;
>>> +	else if (key_len == sizeof(u64))
>>> +		b = k[1];
>>> +	else if (likely(key_len % 4 == 0))
>>>  		return jhash2(key, key_len / 4, hashrnd);
> 
> Nice improvement, but key_size is known at verification time.
> Let's do a step further and specialize htab_map_gen_lookup
> for small key sizes?
> 
Sure, let me try that.

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-09-21 17:10 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-21 15:28 [PATCH bpf-next] bpf: Speed up htab hashing for u32/u64 keys Mykyta Yatsenko
2026-09-21 16:45 ` Anton Protopopov
     [not found] ` <179000542526.3934731.994864591824564797@gmail.com>
2026-09-21 17:00   ` Alexei Starovoitov
2026-09-21 17:10     ` Mykyta Yatsenko

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox