From: Yusheng Zheng <yunwei356@gmail.com>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Andrii Nakryiko <andrii@kernel.org>,
Eduard Zingerman <eddyz87@gmail.com>,
Kumar Kartikeya Dwivedi <memxor@gmail.com>,
Martin KaFai Lau <martin.lau@linux.dev>,
Song Liu <song@kernel.org>,
Yonghong Song <yonghong.song@linux.dev>,
Jiri Olsa <jolsa@kernel.org>,
John Fastabend <john.fastabend@gmail.com>,
Emil Tsalapatis <emil@etsalapatis.com>,
Ihor Solodrai <ihor.solodrai@linux.dev>,
x86@kernel.org, Thomas Gleixner <tglx@kernel.org>,
Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
"H . Peter Anvin" <hpa@zytor.com>,
Leon Hwang <leon.hwang@linux.dev>,
Puranjay Mohan <puranjay@kernel.org>,
Hao Sun <sunhao.th@gmail.com>,
Yusheng Zheng <yunwei356@gmail.com>
Subject: [RFC PATCH bpf-next 0/7] bpf: Inline kfuncs that have a BPF body
Date: Mon, 5 Oct 2026 07:22:12 -0700 [thread overview]
Message-ID: <20261005142219.33451-1-yunwei356@gmail.com> (raw)
The BPF JIT translates one BPF instruction at a time. Operations that
the CPU does in one instruction, such as a rotate, a conditional select
or a big-endian load, take several BPF instructions, and they stay
several after the JIT. On 27 small compute benchmarks, each compiled
from the same C source to BPF and to native code, BPF ran 1.71x
(x86-64) and 2.05x (arm64) slower than native code, while compiling the
same BPF bytecode with LLVM came within 7% of native code [1]. The loss
is in the JIT, not in the bytecode. Our talk in the eBPF track at
LPC 2026 [5] covers this gap and the approach of this series.
A kfunc could do such an operation, but the call costs about as much as
the operation, and the verifier knows nothing about the result. After
off = bpf_select64(c, 4, 8), it does not know that off is 4 or 8, so a
packet load at data + off needs another bounds check.
This series lets a kfunc have a body: a short sequence of BPF
instructions that computes the same result, and optionally a callback
that emits native code for a call. For example:
/* x << n | x >> (-n & 63) */
static const struct bpf_insn rol64_body[] = { ... };
static const struct bpf_kfunc_body bodies[] = {
{ &body_ids[0], rol64_body, ARRAY_SIZE(rol64_body), rol64_emit },
};
Before its analysis, the verifier replaces each call of such a kfunc
with its body, so it analyzes the result like any other BPF code. After
the analysis, the call is restored, and the JIT puts native code in its
place. The native code uses the registers that already hold the
operands, so the argument and result moves around the call go away. If
the JIT has no native code for the call, as on other architectures or
with constant blinding, the verified body stays in the program instead
of a call.
Programs call these kfuncs like any other kfunc; arguments named __k
must be known constants. There is no new instruction encoding and no
UAPI change. Because the verifier analyzes the body at every call, a
program that uses these kfuncs costs as much to verify as the same
operations written in BPF, which is more than an opaque call. JITs
already inline some helper calls, and the verifier already inlines the
numeric iterator kfuncs as BPF. This series combines the two, and in
every case the body is what tells the verifier the result. Instructions
that BPF lacks are the first use, but any small kfunc can have a body,
including kfuncs in modules.
The native code comes with the kfunc, not from the JIT. By default the
x86-64 JIT copies the compiled kfunc and renames its registers, as was
suggested for the bitops kfuncs [2]. A copy keeps constant arguments in
registers, cannot write the result into the register of an argument,
and has only the baseline x86-64 instructions that the kernel is built
for, so no movbe or bextr. A kfunc can therefore have an emit callback
instead, which gets the registers that hold the operands and the
values of the constant arguments. Five of the kfuncs in this series
have one. The kfuncs and their x86-64 code are in
kernel/bpf/insn_kfuncs/, which can be built as a module.
In a loop of dependent operations on x86-64 (Core Ultra 9 285K, KVM;
numbers vary by about 5% between boots), plain is the body as the JIT
translates it today, emit the code from the callback and copy the
copied kfunc:
x86-64 code ns per operation
operation emitted copied plain emit copy
rotate by 13 rol $0xd,%rbx 5 insns 0.38 0.19 0.20
bit extract mov $0x110d,%eax; bextr 10 insns 0.57 0.57 0.57
big-endian load movbe (%rdi),%rax 4 insns 0.57 0.57 1.53
lea lea 0x10(%rbx,%rsi,8),%rbx 8 insns 0.41 0.37 0.41
select test; mov; cmovne same 3.29 1.22 1.21
The emitted code is much shorter, but a copy is about as fast, except
for the big-endian load, which is 2.7x slower as a copy, and lea, which
is about 10% slower. The select row uses an unpredictable condition;
with a predictable one, the branch in plain BPF is 3x faster than cmov.
The paper's prototype also had the verifier check a BPF version of each
operation and ran native code in its place [1]. It made the 27
benchmarks 1.24x (x86-64) and 1.22x (arm64) faster, and raised the
throughput of the Cilium datapath on x86-64 by 7.4%. Where the native
code pays off depends on the CPU and the workload: for Katran on arm64,
21 chosen call sites gave 7.3% more throughput, but all 62 gave 0.5%
less.
- Patch 1 lets a kfunc set give its kfuncs a body and native code.
- Patch 2 is the verifier side, in kernel/bpf/kfunc_inline.c.
- Patch 3 makes the x86-64 JIT inline the native code of a kfunc, or
else a copy of the compiled kfunc.
- Patch 4 adds kfuncs with bodies for rotate, select, bit field
extract, big-endian load, prefetch, 16-byte copy and lea, under
CONFIG_BPF_INSN_KFUNCS.
- Patch 5 adds x86-64 code for five of them.
- Patches 6 and 7 add selftests and documentation.
Testing on x86-64: verifier_kfunc_inline (24 subtests) and
kfunc_body_registration pass with CONFIG_BPF_INSN_KFUNCS=y. With
bpf_jit_harden=2 the bodies stay in the program, so the subtests that
check the JIT output fail as expected, and all results are correct.
Questions:
1. Is it acceptable that the verifier replaces calls with their
bodies before the analysis and restores them afterwards?
2. Should kfuncs bring native code through an emit callback, or
should the JIT only copy compiled kfuncs? The table shows what
copying alone costs.
3. Native code is trusted like the rest of the JIT, and the selftests
compare it with the bodies on random inputs. Is that enough?
4. The bodies are written by hand. Should they rather be compiled from
the kfunc's C source, as in the inlinable kfuncs RFC [3]?
Not done yet: arm64 has no native code and runs the bodies. Copying
compiled kfuncs would be the first step there.
Earlier work:
- Leon Hwang's bitops kfuncs [2] emitted native code for kfuncs in the
JIT. The review asked to keep the C version as a fallback and
suggested copying the compiled kfunc. Here the fallback is the body,
and the native code is kept with the kfuncs instead of in the JIT.
- Eduard Zingerman's inlinable kfuncs [3] and Puranjay Mohan's
iterator inlining [4] inline kfuncs as BPF, so that the verifier
sees what they compute. This series does the same and then lets the
JIT put native code in place of the call.
[1] https://arxiv.org/abs/2606.24213
[2] https://lore.kernel.org/bpf/20260219142933.13904-1-leon.hwang@linux.dev/
[3] https://lore.kernel.org/bpf/20241107175040.1659341-1-eddyz87@gmail.com/
[4] https://lore.kernel.org/bpf/20260804134601.2305303-1-puranjay@kernel.org/
[5] https://lpc.events/event/20/contributions/2445/
Yusheng Zheng (7):
bpf: Let kfunc sets give kfuncs a BPF body
bpf: Verify calls of kfuncs with a body through the body
bpf, x86: Inline native code for kfuncs that have a body
bpf: Add kfuncs with bodies for common operations
bpf, x86: Add native code for some inline kfuncs
selftests/bpf: Test inline kfuncs
Documentation/bpf: Describe inline kfuncs
Documentation/bpf/kfuncs.rst | 53 ++
arch/x86/net/bpf_jit_comp.c | 181 ++++++
include/linux/bpf_verifier.h | 34 ++
include/linux/btf.h | 27 +
include/linux/filter.h | 2 +
kernel/bpf/Kconfig | 1 +
kernel/bpf/Makefile | 3 +-
kernel/bpf/btf.c | 126 ++++
kernel/bpf/const_fold.c | 4 +
kernel/bpf/core.c | 9 +
kernel/bpf/insn_kfuncs/Kconfig | 16 +
kernel/bpf/insn_kfuncs/Makefile | 5 +
kernel/bpf/insn_kfuncs/bpf_insn_kfuncs.c | 184 ++++++
kernel/bpf/insn_kfuncs/x86/insn_kfuncs.h | 155 +++++
kernel/bpf/kfunc_inline.c | 273 +++++++++
kernel/bpf/verifier.c | 53 +-
tools/testing/selftests/bpf/Makefile | 2 +-
tools/testing/selftests/bpf/config | 1 +
.../bpf/prog_tests/kfunc_body_registration.c | 18 +
.../selftests/bpf/prog_tests/verifier.c | 2 +
.../bpf/progs/verifier_kfunc_inline.c | 543 ++++++++++++++++++
.../testing/selftests/bpf/test_kmods/Makefile | 2 +-
.../bpf/test_kmods/bpf_test_kfunc_body.c | 121 ++++
.../selftests/bpf/test_kmods/bpf_testmod.c | 77 +++
.../bpf/test_kmods/bpf_testmod_kfunc.h | 4 +
25 files changed, 1877 insertions(+), 19 deletions(-)
create mode 100644 kernel/bpf/insn_kfuncs/Kconfig
create mode 100644 kernel/bpf/insn_kfuncs/Makefile
create mode 100644 kernel/bpf/insn_kfuncs/bpf_insn_kfuncs.c
create mode 100644 kernel/bpf/insn_kfuncs/x86/insn_kfuncs.h
create mode 100644 kernel/bpf/kfunc_inline.c
create mode 100644 tools/testing/selftests/bpf/prog_tests/kfunc_body_registration.c
create mode 100644 tools/testing/selftests/bpf/progs/verifier_kfunc_inline.c
create mode 100644 tools/testing/selftests/bpf/test_kmods/bpf_test_kfunc_body.c
base-commit: 99dc1ba542420db6b8df209744f55cc52466ad91
--
2.51.1
next reply other threads:[~2026-10-05 14:22 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-05 14:22 Yusheng Zheng [this message]
2026-10-05 14:22 ` [RFC PATCH bpf-next 1/7] bpf: Let kfunc sets give kfuncs a BPF body Yusheng Zheng
2026-10-05 14:38 ` sashiko-bot
2026-10-05 15:16 ` bot+bpf-ci
2026-10-05 14:22 ` [RFC PATCH bpf-next 2/7] bpf: Verify calls of kfuncs with a body through the body Yusheng Zheng
2026-10-05 14:41 ` sashiko-bot
2026-10-05 14:22 ` [RFC PATCH bpf-next 3/7] bpf, x86: Inline native code for kfuncs that have a body Yusheng Zheng
2026-10-05 14:22 ` [RFC PATCH bpf-next 4/7] bpf: Add kfuncs with bodies for common operations Yusheng Zheng
2026-10-05 14:39 ` sashiko-bot
2026-10-05 14:22 ` [RFC PATCH bpf-next 5/7] bpf, x86: Add native code for some inline kfuncs Yusheng Zheng
2026-10-05 14:22 ` [RFC PATCH bpf-next 6/7] selftests/bpf: Test " Yusheng Zheng
2026-10-05 14:22 ` [RFC PATCH bpf-next 7/7] Documentation/bpf: Describe " Yusheng Zheng
2026-10-05 15:16 ` bot+bpf-ci
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261005142219.33451-1-yunwei356@gmail.com \
--to=yunwei356@gmail.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bp@alien8.de \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=dave.hansen@linux.intel.com \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=hpa@zytor.com \
--cc=ihor.solodrai@linux.dev \
--cc=john.fastabend@gmail.com \
--cc=jolsa@kernel.org \
--cc=leon.hwang@linux.dev \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=mingo@redhat.com \
--cc=puranjay@kernel.org \
--cc=song@kernel.org \
--cc=sunhao.th@gmail.com \
--cc=tglx@kernel.org \
--cc=x86@kernel.org \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox