BPF List
 help / color / mirror / Atom feed
* [RFC PATCH bpf-next 0/7] bpf: Inline kfuncs that have a BPF body
@ 2026-10-05 14:22 Yusheng Zheng
  2026-10-05 14:22 ` [RFC PATCH bpf-next 1/7] bpf: Let kfunc sets give kfuncs " Yusheng Zheng
                   ` (6 more replies)
  0 siblings, 7 replies; 13+ messages in thread
From: Yusheng Zheng @ 2026-10-05 14:22 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
	Song Liu, Yonghong Song, Jiri Olsa, John Fastabend,
	Emil Tsalapatis, Ihor Solodrai, x86, Thomas Gleixner, Ingo Molnar,
	Borislav Petkov, Dave Hansen, H . Peter Anvin, Leon Hwang,
	Puranjay Mohan, Hao Sun, Yusheng Zheng

The BPF JIT translates one BPF instruction at a time. Operations that
the CPU does in one instruction, such as a rotate, a conditional select
or a big-endian load, take several BPF instructions, and they stay
several after the JIT. On 27 small compute benchmarks, each compiled
from the same C source to BPF and to native code, BPF ran 1.71x
(x86-64) and 2.05x (arm64) slower than native code, while compiling the
same BPF bytecode with LLVM came within 7% of native code [1]. The loss
is in the JIT, not in the bytecode. Our talk in the eBPF track at
LPC 2026 [5] covers this gap and the approach of this series.

A kfunc could do such an operation, but the call costs about as much as
the operation, and the verifier knows nothing about the result. After
off = bpf_select64(c, 4, 8), it does not know that off is 4 or 8, so a
packet load at data + off needs another bounds check.

This series lets a kfunc have a body: a short sequence of BPF
instructions that computes the same result, and optionally a callback
that emits native code for a call. For example:

  /* x << n | x >> (-n & 63) */
  static const struct bpf_insn rol64_body[] = { ... };

  static const struct bpf_kfunc_body bodies[] = {
    { &body_ids[0], rol64_body, ARRAY_SIZE(rol64_body), rol64_emit },
  };

Before its analysis, the verifier replaces each call of such a kfunc
with its body, so it analyzes the result like any other BPF code. After
the analysis, the call is restored, and the JIT puts native code in its
place. The native code uses the registers that already hold the
operands, so the argument and result moves around the call go away. If
the JIT has no native code for the call, as on other architectures or
with constant blinding, the verified body stays in the program instead
of a call.

Programs call these kfuncs like any other kfunc; arguments named __k
must be known constants. There is no new instruction encoding and no
UAPI change. Because the verifier analyzes the body at every call, a
program that uses these kfuncs costs as much to verify as the same
operations written in BPF, which is more than an opaque call. JITs
already inline some helper calls, and the verifier already inlines the
numeric iterator kfuncs as BPF. This series combines the two, and in
every case the body is what tells the verifier the result. Instructions
that BPF lacks are the first use, but any small kfunc can have a body,
including kfuncs in modules.

The native code comes with the kfunc, not from the JIT. By default the
x86-64 JIT copies the compiled kfunc and renames its registers, as was
suggested for the bitops kfuncs [2]. A copy keeps constant arguments in
registers, cannot write the result into the register of an argument,
and has only the baseline x86-64 instructions that the kernel is built
for, so no movbe or bextr. A kfunc can therefore have an emit callback
instead, which gets the registers that hold the operands and the
values of the constant arguments. Five of the kfuncs in this series
have one. The kfuncs and their x86-64 code are in
kernel/bpf/insn_kfuncs/, which can be built as a module.

In a loop of dependent operations on x86-64 (Core Ultra 9 285K, KVM;
numbers vary by about 5% between boots), plain is the body as the JIT
translates it today, emit the code from the callback and copy the
copied kfunc:

                   x86-64 code                           ns per operation
  operation        emitted                      copied   plain emit  copy
  rotate by 13     rol $0xd,%rbx                5 insns  0.38  0.19  0.20
  bit extract      mov $0x110d,%eax; bextr     10 insns  0.57  0.57  0.57
  big-endian load  movbe (%rdi),%rax            4 insns  0.57  0.57  1.53
  lea              lea 0x10(%rbx,%rsi,8),%rbx   8 insns  0.41  0.37  0.41
  select           test; mov; cmovne            same     3.29  1.22  1.21

The emitted code is much shorter, but a copy is about as fast, except
for the big-endian load, which is 2.7x slower as a copy, and lea, which
is about 10% slower. The select row uses an unpredictable condition;
with a predictable one, the branch in plain BPF is 3x faster than cmov.

The paper's prototype also had the verifier check a BPF version of each
operation and ran native code in its place [1]. It made the 27
benchmarks 1.24x (x86-64) and 1.22x (arm64) faster, and raised the
throughput of the Cilium datapath on x86-64 by 7.4%. Where the native
code pays off depends on the CPU and the workload: for Katran on arm64,
21 chosen call sites gave 7.3% more throughput, but all 62 gave 0.5%
less.

  - Patch 1 lets a kfunc set give its kfuncs a body and native code.
  - Patch 2 is the verifier side, in kernel/bpf/kfunc_inline.c.
  - Patch 3 makes the x86-64 JIT inline the native code of a kfunc, or
    else a copy of the compiled kfunc.
  - Patch 4 adds kfuncs with bodies for rotate, select, bit field
    extract, big-endian load, prefetch, 16-byte copy and lea, under
    CONFIG_BPF_INSN_KFUNCS.
  - Patch 5 adds x86-64 code for five of them.
  - Patches 6 and 7 add selftests and documentation.

Testing on x86-64: verifier_kfunc_inline (24 subtests) and
kfunc_body_registration pass with CONFIG_BPF_INSN_KFUNCS=y. With
bpf_jit_harden=2 the bodies stay in the program, so the subtests that
check the JIT output fail as expected, and all results are correct.

Questions:
  1. Is it acceptable that the verifier replaces calls with their
     bodies before the analysis and restores them afterwards?
  2. Should kfuncs bring native code through an emit callback, or
     should the JIT only copy compiled kfuncs? The table shows what
     copying alone costs.
  3. Native code is trusted like the rest of the JIT, and the selftests
     compare it with the bodies on random inputs. Is that enough?
  4. The bodies are written by hand. Should they rather be compiled from
     the kfunc's C source, as in the inlinable kfuncs RFC [3]?

Not done yet: arm64 has no native code and runs the bodies. Copying
compiled kfuncs would be the first step there.

Earlier work:
  - Leon Hwang's bitops kfuncs [2] emitted native code for kfuncs in the
    JIT. The review asked to keep the C version as a fallback and
    suggested copying the compiled kfunc. Here the fallback is the body,
    and the native code is kept with the kfuncs instead of in the JIT.
  - Eduard Zingerman's inlinable kfuncs [3] and Puranjay Mohan's
    iterator inlining [4] inline kfuncs as BPF, so that the verifier
    sees what they compute. This series does the same and then lets the
    JIT put native code in place of the call.

[1] https://arxiv.org/abs/2606.24213
[2] https://lore.kernel.org/bpf/20260219142933.13904-1-leon.hwang@linux.dev/
[3] https://lore.kernel.org/bpf/20241107175040.1659341-1-eddyz87@gmail.com/
[4] https://lore.kernel.org/bpf/20260804134601.2305303-1-puranjay@kernel.org/
[5] https://lpc.events/event/20/contributions/2445/

Yusheng Zheng (7):
  bpf: Let kfunc sets give kfuncs a BPF body
  bpf: Verify calls of kfuncs with a body through the body
  bpf, x86: Inline native code for kfuncs that have a body
  bpf: Add kfuncs with bodies for common operations
  bpf, x86: Add native code for some inline kfuncs
  selftests/bpf: Test inline kfuncs
  Documentation/bpf: Describe inline kfuncs

 Documentation/bpf/kfuncs.rst                  |  53 ++
 arch/x86/net/bpf_jit_comp.c                   | 181 ++++++
 include/linux/bpf_verifier.h                  |  34 ++
 include/linux/btf.h                           |  27 +
 include/linux/filter.h                        |   2 +
 kernel/bpf/Kconfig                            |   1 +
 kernel/bpf/Makefile                           |   3 +-
 kernel/bpf/btf.c                              | 126 ++++
 kernel/bpf/const_fold.c                       |   4 +
 kernel/bpf/core.c                             |   9 +
 kernel/bpf/insn_kfuncs/Kconfig                |  16 +
 kernel/bpf/insn_kfuncs/Makefile               |   5 +
 kernel/bpf/insn_kfuncs/bpf_insn_kfuncs.c      | 184 ++++++
 kernel/bpf/insn_kfuncs/x86/insn_kfuncs.h      | 155 +++++
 kernel/bpf/kfunc_inline.c                     | 273 +++++++++
 kernel/bpf/verifier.c                         |  53 +-
 tools/testing/selftests/bpf/Makefile          |   2 +-
 tools/testing/selftests/bpf/config            |   1 +
 .../bpf/prog_tests/kfunc_body_registration.c  |  18 +
 .../selftests/bpf/prog_tests/verifier.c       |   2 +
 .../bpf/progs/verifier_kfunc_inline.c         | 543 ++++++++++++++++++
 .../testing/selftests/bpf/test_kmods/Makefile |   2 +-
 .../bpf/test_kmods/bpf_test_kfunc_body.c      | 121 ++++
 .../selftests/bpf/test_kmods/bpf_testmod.c    |  77 +++
 .../bpf/test_kmods/bpf_testmod_kfunc.h        |   4 +
 25 files changed, 1877 insertions(+), 19 deletions(-)
 create mode 100644 kernel/bpf/insn_kfuncs/Kconfig
 create mode 100644 kernel/bpf/insn_kfuncs/Makefile
 create mode 100644 kernel/bpf/insn_kfuncs/bpf_insn_kfuncs.c
 create mode 100644 kernel/bpf/insn_kfuncs/x86/insn_kfuncs.h
 create mode 100644 kernel/bpf/kfunc_inline.c
 create mode 100644 tools/testing/selftests/bpf/prog_tests/kfunc_body_registration.c
 create mode 100644 tools/testing/selftests/bpf/progs/verifier_kfunc_inline.c
 create mode 100644 tools/testing/selftests/bpf/test_kmods/bpf_test_kfunc_body.c


base-commit: 99dc1ba542420db6b8df209744f55cc52466ad91
-- 
2.51.1


^ permalink raw reply	[flat|nested] 13+ messages in thread

end of thread, other threads:[~2026-10-05 15:16 UTC | newest]

Thread overview: 13+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-05 14:22 [RFC PATCH bpf-next 0/7] bpf: Inline kfuncs that have a BPF body Yusheng Zheng
2026-10-05 14:22 ` [RFC PATCH bpf-next 1/7] bpf: Let kfunc sets give kfuncs " Yusheng Zheng
2026-10-05 14:38   ` sashiko-bot
2026-10-05 15:16   ` bot+bpf-ci
2026-10-05 14:22 ` [RFC PATCH bpf-next 2/7] bpf: Verify calls of kfuncs with a body through the body Yusheng Zheng
2026-10-05 14:41   ` sashiko-bot
2026-10-05 14:22 ` [RFC PATCH bpf-next 3/7] bpf, x86: Inline native code for kfuncs that have a body Yusheng Zheng
2026-10-05 14:22 ` [RFC PATCH bpf-next 4/7] bpf: Add kfuncs with bodies for common operations Yusheng Zheng
2026-10-05 14:39   ` sashiko-bot
2026-10-05 14:22 ` [RFC PATCH bpf-next 5/7] bpf, x86: Add native code for some inline kfuncs Yusheng Zheng
2026-10-05 14:22 ` [RFC PATCH bpf-next 6/7] selftests/bpf: Test " Yusheng Zheng
2026-10-05 14:22 ` [RFC PATCH bpf-next 7/7] Documentation/bpf: Describe " Yusheng Zheng
2026-10-05 15:16   ` bot+bpf-ci

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox