From: Jiri Olsa <olsajiri@gmail.com>
To: "Mike Rapoport (Microsoft)" <rppt@kernel.org>
Cc: Alexei Starovoitov <ast@kernel.org>,
Andrii Nakryiko <andrii@kernel.org>,
Andy Lutomirski <luto@kernel.org>, Borislav Petkov <bp@alien8.de>,
Daniel Borkmann <daniel@iogearbox.net>,
Dave Hansen <dave.hansen@linux.intel.com>,
Eduard Zingerman <eddyz87@gmail.com>,
Ingo Molnar <mingo@redhat.com>,
Kumar Kartikeya Dwivedi <memxor@gmail.com>,
Peter Zijlstra <peterz@infradead.org>, Song Liu <song@kernel.org>,
Thomas Gleixner <tglx@kernel.org>,
Emil Tsalapatis <emil@etsalapatis.com>,
John Fastabend <john.fastabend@gmail.com>,
Martin KaFai Lau <martin.lau@linux.dev>,
"H. Peter Anvin" <hpa@zytor.com>,
Yonghong Song <yonghong.song@linux.dev>,
bpf@vger.kernel.org, linux-kernel@vger.kernel.org,
x86@kernel.org, Leon Hwang <leon.hwang@linux.dev>
Subject: Re: [PATCH bpf-next v3 4/5] bpf, x86: make sure allocation in arch_bpf_trampoline_size() is writable
Date: Fri, 14 Aug 2026 16:53:32 +0200 [thread overview]
Message-ID: <an8r7EODLIL-bZM3@krava> (raw)
In-Reply-To: <20260716-execmem-x86-rox-bpf-v0-v3-4-4e76158c01c5@kernel.org>
On Thu, Jul 16, 2026 at 10:51:38AM +0300, Mike Rapoport (Microsoft) wrote:
> arch_bpf_trampoline_size() allocates a buffer to get actual size required
> for a trampoline.
>
> This buffer must be in the module address space because
> __arch_prepare_bpf_trampoline() calculates rel32 offsets relatively to
> that buffer.
>
> In preparation for enabling ROX mode for EXECMEM_BPF make sure that the
> allocated memory is writable.
>
> Add bpf_jit_alloc_exec_rw() wrapper for execmem_alloc_rw() and use it for > buffer allocation in arch_bpf_trampoline_size().
>
> Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
> ---
> arch/x86/net/bpf_jit_comp.c | 5 ++---
> include/linux/filter.h | 1 +
> kernel/bpf/core.c | 5 +++++
> 3 files changed, 8 insertions(+), 3 deletions(-)
>
> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
> index de7515ea1bea..b2feec81e231 100644
> --- a/arch/x86/net/bpf_jit_comp.c
> +++ b/arch/x86/net/bpf_jit_comp.c
> @@ -3703,13 +3703,12 @@ int arch_bpf_trampoline_size(const struct btf_func_model *m, u32 flags,
> int ret;
>
> /* Allocate a temporary buffer for __arch_prepare_bpf_trampoline().
> - * This will NOT cause fragmentation in direct map, as we do not
> - * call set_memory_*() on this buffer.
> *
> * We cannot use kvmalloc here, because we need image to be in
> * module memory range.
> + * Since it must be writable use bpf_jit_alloc_exec_rw().
> */
> - image = bpf_jit_alloc_exec(PAGE_SIZE);
> + image = bpf_jit_alloc_exec_rw(PAGE_SIZE);
hi,
this change (this particular patch plus possibly others in this set) is
causing tracing_multi attachment bench slowdown
the benchmark allocates huge number of trampolines and I'm seeing extra
arch_bpf_trampoline_size code paths in the attached perf profile
I'm not that familiar with the allocator, but following hack makes the
benchmark ok again:
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index 6a94370a2448..bbff3c9c6681 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -1130,7 +1130,7 @@ void *bpf_jit_alloc_exec(unsigned long size)
void *bpf_jit_alloc_exec_rw(unsigned long size)
{
- return execmem_alloc_rw(EXECMEM_BPF, size);
+ return execmem_alloc(EXECMEM_MODULE_DATA, size);
}
void bpf_jit_free_exec(void *addr)
I still need to do more checks, but I'm wondering if we could actually fix
this by not allocating image data in arch_bpf_trampoline_size at all..
and just teach __arch_prepare_bpf_trampoline to survive NULL image data
and just return the size in such case
thanks,
jirka
---
11.11% test_progs [kernel.kallsyms] [k] decay_va_pool_node
|
---decay_va_pool_node
__purge_vmap_area_lazy
_vm_unmap_aliases
change_page_attr_set_clr
|
|--4.58%--set_memory_rw
| |
| |--2.33%--execmem_alloc_rw
| | arch_bpf_trampoline_size
| | bpf_trampoline_update
| | __bpf_trampoline_link_prog
| | bpf_trampoline_multi_attach
| | bpf_tracing_multi_attach
| | __sys_bpf
| | __x64_sys_bpf
| | do_syscall_64
| | entry_SYSCALL_64_after_hwframe
| | syscall
| | sys_bpf_fd
| | bpf_link_create
| | bpf_program__attach_tracing_multi
| | serial_test_tracing_multi_bench_attach
| | run_one_test
| | main
| | __libc_start_call_main
| | __libc_start_main@@GLIBC_2.34
| | _start
| |
| --2.25%--__execmem_cache_free
| execmem_free
| arch_bpf_trampoline_size
| bpf_trampoline_update
| __bpf_trampoline_link_prog
| bpf_trampoline_multi_attach
| bpf_tracing_multi_attach
| __sys_bpf
| __x64_sys_bpf
| do_syscall_64
| entry_SYSCALL_64_after_hwframe
| syscall
| sys_bpf_fd
| bpf_link_create
| bpf_program__attach_tracing_multi
| serial_test_tracing_multi_bench_attach
| run_one_test
| main
| __libc_start_call_main
| __libc_start_main@@GLIBC_2.34
| _start
|
|--4.52%--set_memory_nx
| |
| |--2.32%--execmem_alloc_rw
| | arch_bpf_trampoline_size
| | bpf_trampoline_update
| | __bpf_trampoline_link_prog
| | bpf_trampoline_multi_attach
| | bpf_tracing_multi_attach
| | __sys_bpf
| | __x64_sys_bpf
| | do_syscall_64
| | entry_SYSCALL_64_after_hwframe
| | syscall
| | sys_bpf_fd
| | bpf_link_create
| | bpf_program__attach_tracing_multi
| | serial_test_tracing_multi_bench_attach
| | run_one_test
| | main
| | __libc_start_call_main
| | __libc_start_main@@GLIBC_2.34
| | _start
| |
| --2.20%--__execmem_cache_free
| execmem_free
| arch_bpf_trampoline_size
| bpf_trampoline_update
| __bpf_trampoline_link_prog
| bpf_trampoline_multi_attach
| bpf_tracing_multi_attach
| __sys_bpf
| __x64_sys_bpf
| do_syscall_64
| entry_SYSCALL_64_after_hwframe
| syscall
| sys_bpf_fd
| bpf_link_create
| bpf_program__attach_tracing_multi
| serial_test_tracing_multi_bench_attach
| run_one_test
| main
| __libc_start_call_main
| __libc_start_main@@GLIBC_2.34
| _start
|
--2.01%--set_memory_rox
|
--2.01%--__execmem_cache_free
execmem_free
arch_bpf_trampoline_size
bpf_trampoline_update
__bpf_trampoline_link_prog
bpf_trampoline_multi_attach
bpf_tracing_multi_attach
__sys_bpf
__x64_sys_bpf
do_syscall_64
entry_SYSCALL_64_after_hwframe
syscall
sys_bpf_fd
bpf_link_create
bpf_program__attach_tracing_multi
serial_test_tracing_multi_bench_attach
run_one_test
main
__libc_start_call_main
__libc_start_main@@GLIBC_2.34
_start
6.15% test_progs [kernel.kallsyms] [k] smp_call_function_many_cond
|
---smp_call_function_many_cond
on_each_cpu_cond_mask
|
|--4.80%--cpa_flush
| change_page_attr_set_clr
| |
| |--2.37%--set_memory_rox
| | __execmem_cache_free
| | execmem_free
| | arch_bpf_trampoline_size
| | bpf_trampoline_update
| | __bpf_trampoline_link_prog
| | bpf_trampoline_multi_attach
| | bpf_tracing_multi_attach
| | __sys_bpf
| | __x64_sys_bpf
| | do_syscall_64
| | entry_SYSCALL_64_after_hwframe
| | syscall
| | sys_bpf_fd
| | bpf_link_create
| | bpf_program__attach_tracing_multi
| | serial_test_tracing_multi_bench_attach
| | run_one_test
| | main
| | __libc_start_call_main
| | __libc_start_main@@GLIBC_2.34
| | _start
| |
| |--1.21%--set_memory_nx
| | execmem_alloc_rw
| | arch_bpf_trampoline_size
| | bpf_trampoline_update
| | __bpf_trampoline_link_prog
| | bpf_trampoline_multi_attach
| | bpf_tracing_multi_attach
| | __sys_bpf
| | __x64_sys_bpf
| | do_syscall_64
| | entry_SYSCALL_64_after_hwframe
| | syscall
| | sys_bpf_fd
| | bpf_link_create
| | bpf_program__attach_tracing_multi
| | serial_test_tracing_multi_bench_attach
| | run_one_test
| | main
| | __libc_start_call_main
| | __libc_start_main@@GLIBC_2.34
| | _start
| |
| --1.21%--set_memory_rw
| execmem_alloc_rw
| |
| --1.21%--arch_bpf_trampoline_size
| bpf_trampoline_update
| __bpf_trampoline_link_prog
| bpf_trampoline_multi_attach
| bpf_tracing_multi_attach
| __sys_bpf
| __x64_sys_bpf
| do_syscall_64
| entry_SYSCALL_64_after_hwframe
| syscall
| sys_bpf_fd
| bpf_link_create
| bpf_program__attach_tracing_multi
| serial_test_tracing_multi_bench_attach
| run_one_test
| main
| __libc_start_call_main
| __libc_start_main@@GLIBC_2.34
| _start
|
--1.21%--__change_page_attr_set_clr
change_page_attr_set_clr
set_memory_nx
execmem_alloc_rw
|
--1.20%--arch_bpf_trampoline_size
bpf_trampoline_update
__bpf_trampoline_link_prog
bpf_trampoline_multi_attach
bpf_tracing_multi_attach
__sys_bpf
__x64_sys_bpf
do_syscall_64
entry_SYSCALL_64_after_hwframe
syscall
sys_bpf_fd
bpf_link_create
bpf_program__attach_tracing_multi
serial_test_tracing_multi_bench_attach
run_one_test
main
__libc_start_call_main
__libc_start_main@@GLIBC_2.34
_start
5.01% test_progs [kernel.kallsyms] [k] __raw_callee_save___pv_queued_spin_unlock
|
---__raw_callee_save___pv_queued_spin_unlock
do_raw_spin_unlock
|
|--4.20%--_raw_spin_unlock
| |
| --4.10%--decay_va_pool_node
| __purge_vmap_area_lazy
| _vm_unmap_aliases
| change_page_attr_set_clr
| |
| |--1.73%--set_memory_nx
| | |
| | |--0.95%--__execmem_cache_free
| | | execmem_free
| | | arch_bpf_trampoline_size
| | | bpf_trampoline_update
| | | __bpf_trampoline_link_prog
| | | bpf_trampoline_multi_attach
| | | bpf_tracing_multi_attach
| | | __sys_bpf
| | | __x64_sys_bpf
| | | do_syscall_64
| | | entry_SYSCALL_64_after_hwframe
| | | syscall
| | | sys_bpf_fd
| | | bpf_link_create
| | | bpf_program__attach_tracing_multi
| | | serial_test_tracing_multi_bench_attach
| | | run_one_test
| | | main
| | | __libc_start_call_main
| | | __libc_start_main@@GLIBC_2.34
| | | _start
| | |
| | --0.78%--execmem_alloc_rw
| | arch_bpf_trampoline_size
| | bpf_trampoline_update
| | __bpf_trampoline_link_prog
| | bpf_trampoline_multi_attach
| | bpf_tracing_multi_attach
| | __sys_bpf
| | __x64_sys_bpf
| | do_syscall_64
| | entry_SYSCALL_64_after_hwframe
| | syscall
| | sys_bpf_fd
| | bpf_link_create
| | bpf_program__attach_tracing_multi
| | serial_test_tracing_multi_bench_attach
| | run_one_test
| | main
| | __libc_start_call_main
| | __libc_start_main@@GLIBC_2.34
| | _start
| |
| |--1.51%--set_memory_rw
| | |
| | |--0.76%--__execmem_cache_free
| | | execmem_free
| | | arch_bpf_trampoline_size
| | | bpf_trampoline_update
| | | __bpf_trampoline_link_prog
| | | bpf_trampoline_multi_attach
| | | bpf_tracing_multi_attach
| | | __sys_bpf
| | | __x64_sys_bpf
| | | do_syscall_64
| | | entry_SYSCALL_64_after_hwframe
| | | syscall
| | | sys_bpf_fd
| | | bpf_link_create
| | | bpf_program__attach_tracing_multi
| | | serial_test_tracing_multi_bench_attach
| | | run_one_test
| | | main
| | | __libc_start_call_main
| | | __libc_start_main@@GLIBC_2.34
| | | _start
| | |
| | --0.75%--execmem_alloc_rw
| | arch_bpf_trampoline_size
| | bpf_trampoline_update
| | __bpf_trampoline_link_prog
| | bpf_trampoline_multi_attach
| | bpf_tracing_multi_attach
| | __sys_bpf
| | __x64_sys_bpf
| | do_syscall_64
| | entry_SYSCALL_64_after_hwframe
| | syscall
| | sys_bpf_fd
| | bpf_link_create
| | bpf_program__attach_tracing_multi
| | serial_test_tracing_multi_bench_attach
| | run_one_test
| | main
| | __libc_start_call_main
| | __libc_start_main@@GLIBC_2.34
| | _start
| |
| --0.86%--set_memory_rox
| __execmem_cache_free
| execmem_free
| arch_bpf_trampoline_size
| bpf_trampoline_update
| __bpf_trampoline_link_prog
| bpf_trampoline_multi_attach
| bpf_tracing_multi_attach
| __sys_bpf
| __x64_sys_bpf
| do_syscall_64
| entry_SYSCALL_64_after_hwframe
| syscall
| sys_bpf_fd
| bpf_link_create
| bpf_program__attach_tracing_multi
| serial_test_tracing_multi_bench_attach
| run_one_test
| main
| __libc_start_call_main
| __libc_start_main@@GLIBC_2.34
| _start
|
--0.80%--_raw_spin_unlock_irqrestore
4.88% test_progs [kernel.kallsyms] [k] get_symbol_offset
next prev parent reply other threads:[~2026-08-14 14:53 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-16 7:51 [PATCH bpf-next v3 0/5] bpf, x86: enable EXECMEM_ROX_CACHE for BPF allocations Mike Rapoport (Microsoft)
2026-07-16 7:51 ` [PATCH bpf-next v3 1/5] bpf: dispatcher: allocate bpf_dispatcher->rw_image with vzalloc() Mike Rapoport (Microsoft)
2026-07-16 8:52 ` bot+bpf-ci
2026-07-16 9:27 ` Mike Rapoport
2026-07-16 9:49 ` Kumar Kartikeya Dwivedi
2026-07-16 7:51 ` [PATCH bpf-next v3 2/5] bpf: drop __weak from bpf_jit_alloc_exec() and bpf_jit_free_exec() Mike Rapoport (Microsoft)
2026-07-16 7:51 ` [PATCH bpf-next v3 3/5] bpf: alloc_prog_pack(): skip ROX management for already ROX memory Mike Rapoport (Microsoft)
2026-07-16 7:51 ` [PATCH bpf-next v3 4/5] bpf, x86: make sure allocation in arch_bpf_trampoline_size() is writable Mike Rapoport (Microsoft)
2026-08-14 14:53 ` Jiri Olsa [this message]
2026-07-16 7:51 ` [PATCH bpf-next v3 5/5] x86/bpf: enable EXECMEM_ROX_CACHE for BPF allocations Mike Rapoport (Microsoft)
2026-07-17 0:00 ` [PATCH bpf-next v3 0/5] bpf, x86: " Song Liu
2026-07-17 6:41 ` Mike Rapoport
2026-07-17 7:27 ` Song Liu
2026-07-17 9:29 ` Mike Rapoport
2026-07-17 17:50 ` Song Liu
2026-07-19 9:23 ` Mike Rapoport
2026-07-21 17:59 ` Song Liu
2026-07-22 8:05 ` Mike Rapoport
2026-07-22 15:04 ` Song Liu
2026-07-27 8:35 ` Mike Rapoport
2026-07-22 15:40 ` patchwork-bot+netdevbpf
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=an8r7EODLIL-bZM3@krava \
--to=olsajiri@gmail.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bp@alien8.de \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=dave.hansen@linux.intel.com \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=hpa@zytor.com \
--cc=john.fastabend@gmail.com \
--cc=leon.hwang@linux.dev \
--cc=linux-kernel@vger.kernel.org \
--cc=luto@kernel.org \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rppt@kernel.org \
--cc=song@kernel.org \
--cc=tglx@kernel.org \
--cc=x86@kernel.org \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.