From: Leon Hwang <leon.hwang@linux.dev>
To: Ihor Solodrai <ihor.solodrai@linux.dev>,
Catalin Marinas <catalin.marinas@arm.com>,
Will Deacon <will@kernel.org>,
Steven Rostedt <rostedt@goodmis.org>,
Masami Hiramatsu <mhiramat@kernel.org>,
Mark Rutland <mark.rutland@arm.com>,
Alexei Starovoitov <ast@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Andrii Nakryiko <andrii@kernel.org>,
Eduard Zingerman <eddyz87@gmail.com>,
Kumar Kartikeya Dwivedi <memxor@gmail.com>,
Martin KaFai Lau <martin.lau@linux.dev>,
Song Liu <song@kernel.org>,
Yonghong Song <yonghong.song@linux.dev>,
Jiri Olsa <jolsa@kernel.org>,
Emil Tsalapatis <emil@etsalapatis.com>,
Shuah Khan <shuah@kernel.org>,
"Jose Fernandez (Anthropic)" <jose.fernandez@linux.dev>
Cc: bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org,
linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org,
linux-kselftest@vger.kernel.org
Subject: Re: [PATCH 1/2] arm64: ftrace: enable single ftrace_ops for direct calls
Date: Fri, 31 Jul 2026 10:21:37 +0800 [thread overview]
Message-ID: <93b6496e-8f52-4983-95e1-01db68609bc0@linux.dev> (raw)
In-Reply-To: <77edc25d-a521-4f65-84ed-dbb016593902@linux.dev>
On 31/7/26 07:03, Ihor Solodrai wrote:
> On 7/27/26 7:28 AM, Leon Hwang wrote:
>> The BPF tracing multi link updates several direct-call sites through one
>> ftrace_ops. Its implementation is therefore gated by
>> HAVE_SINGLE_FTRACE_DIRECT_OPS in addition to
>> DYNAMIC_FTRACE_WITH_DIRECT_CALLS.
>>
>> Select HAVE_SINGLE_FTRACE_DIRECT_OPS whenever arm64 enables dynamic ftrace
>> direct calls. This enables BPF tracing multi links on arm64. Also
>> generalize the unreachable-trampoline comment because the single-ops path
>> does not use ops->direct_call.
>
> Hi Leon,
>
> I don't think this change can land as is yet. The series doesn't even
> apply cleanly to bpf-next, but that's minor.
>
> More importantly, it depends on Jose's series [1], which is not in the
> mainline yet. And there Mark has raised performance concerns [2] and
> the discussion still seems to be open.
>
> [1] https://lore.kernel.org/all/20260609-arm64-ftrace-direct-calls-v1-0-4a46f266697f@linux.dev/
> [2] https://lore.kernel.org/all/amjnf5gz0xP5PTSB@J2N7QTR9R3/
>
>>
>> Assisted-by: Codex:gpt-5.6-sol
>> Signed-off-by: Leon Hwang <leon.hwang@linux.dev>
>> ---
>> arch/arm64/Kconfig | 2 ++
>> arch/arm64/kernel/ftrace.c | 3 +--
>> 2 files changed, 3 insertions(+), 2 deletions(-)
>>
>> diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig
>> index 0de419ed780f..c98dca76859b 100644
>> --- a/arch/arm64/Kconfig
>> +++ b/arch/arm64/Kconfig
>> @@ -188,6 +188,8 @@ config ARM64
>> CLANG_SUPPORTS_DYNAMIC_FTRACE_WITH_ARGS)
>> select HAVE_DYNAMIC_FTRACE_WITH_DIRECT_CALLS \
>> if DYNAMIC_FTRACE_WITH_ARGS
>> + select HAVE_SINGLE_FTRACE_DIRECT_OPS \
>> + if DYNAMIC_FTRACE_WITH_DIRECT_CALLS\
>
> The select is only conditional on DYNAMIC_FTRACE_WITH_DIRECT_CALLS, so
> it can be set along with HAVE_DYNAMIC_FTRACE_WITH_CALL_OPS. And AFAIU
> this would make the fast path effectively dead: every BPF direct call
> routed through ftrace_caller now goes through the slow path.
>
> As Mark noted in the other thread, on arm64 trampolines come from
> EXECMEM_BPF, so they always land out of BL range (chance of landing in
> range is 256M/terabytes).
>
> I vibe-slop-coded a benchmark and ran it on a Neoverse V2 machine, and
> toggling your config change seems to be causing a 1.3x regression in
> the tracing overhead:
>
> do-nothing fentry (r0=0; exit) on __arm64_sys_getpid, 20 M calls, min-of-N, several boots. Results:
>
> untraced (base) traced overhead
> Kernel A ~128 ns/call ~145.5 ns ~17.6 ns (fast path: br x17)
> Kernel B ~127 ns/call ~150.4 ns ~23.4 ns (slow path: save regs + call_direct_funcs + hash)
Thanks for your testing.
>
> This confirms Jiri's suspicion.
True.
>
> However my understanding is the regression should mostly disappear in
> case some version of in-range trampoline allocation on arm64 lands.
>
> So, I think the landing sequence should be something like follows:
> * in-range BPF-trampoline allocation that Jose proposed [3]
> * then HAVE_SINGLE_FTRACE_DIRECT_OPS selection
>
> After all of that reaches mainline, then a selftest patch can go
> through the bpf-next.
Sounds reasonable.
This series is based on Jose's series and is intended for the arm64
tree, rather than bpf-next. Like Jose's series, this series aims to
enable BPF tracing_multi link on arm64. And yes, with in-range
BPF-trampoline allocation, the regression should mostly disappear.
I'll follow the sequence and repost the patches afterward.
Thanks,
Leon
>
> [3] https://lore.kernel.org/all/amn7cBQ-cqs2XlKV@linux.dev/
>
>> select HAVE_DYNAMIC_FTRACE_WITH_CALL_OPS \
>> if (DYNAMIC_FTRACE_WITH_ARGS && !CFI && \
>> (CC_IS_CLANG || !CC_OPTIMIZE_FOR_SIZE))
>> diff --git a/arch/arm64/kernel/ftrace.c b/arch/arm64/kernel/ftrace.c
>> index e1a3c0b3a051..56ba72a87dfa 100644
>> --- a/arch/arm64/kernel/ftrace.c
>> +++ b/arch/arm64/kernel/ftrace.c
>> @@ -301,8 +301,7 @@ static bool ftrace_find_callable_addr(struct dyn_ftrace *rec,
>>
>> /*
>> * If a custom trampoline is unreachable, rely on the ftrace_caller
>> - * trampoline which knows how to indirectly reach that trampoline
>> - * through ops->direct_call.
>> + * trampoline which knows how to indirectly reach that trampoline.
>> */
>> if (*addr != FTRACE_ADDR && !reachable_by_bl(*addr, pc))
>> *addr = FTRACE_ADDR;
>
next prev parent reply other threads:[~2026-07-31 2:22 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-27 14:28 [PATCH 0/2] arm64: ftrace: enable single ftrace_ops for direct calls Leon Hwang
2026-07-27 14:28 ` [PATCH 1/2] " Leon Hwang
2026-07-27 14:52 ` sashiko-bot
2026-07-28 2:19 ` Leon Hwang
2026-07-28 15:46 ` Steven Rostedt
2026-07-29 6:27 ` Leon Hwang
2026-07-29 11:56 ` Jiri Olsa
2026-07-30 3:05 ` Leon Hwang
2026-07-30 10:58 ` Jiri Olsa
2026-07-31 2:30 ` Leon Hwang
2026-07-30 23:03 ` Ihor Solodrai
2026-07-31 2:21 ` Leon Hwang [this message]
2026-07-27 14:28 ` [PATCH 2/2] selftests/bpf: Enable tracing_multi tests on arm64 Leon Hwang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=93b6496e-8f52-4983-95e1-01db68609bc0@linux.dev \
--to=leon.hwang@linux.dev \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=catalin.marinas@arm.com \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=ihor.solodrai@linux.dev \
--cc=jolsa@kernel.org \
--cc=jose.fernandez@linux.dev \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=mhiramat@kernel.org \
--cc=rostedt@goodmis.org \
--cc=shuah@kernel.org \
--cc=song@kernel.org \
--cc=will@kernel.org \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.