From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f45.google.com (mail-wm1-f45.google.com [209.85.128.45]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9038E49B1ED for ; Mon, 5 Oct 2026 14:22:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.45 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791210150; cv=none; b=QH/Ef9XctnXHeGkx7c4KvWStptm+CPmnLQ3KonBZaJMjKdiAd3uigEtAlDanYd23NKX2h/LjVrRRD69zUqd34vPpVCNo/JUhnET6ZI7ak06qVkPZwfRJtTpkwhttt4Of0YfGCFMrcyeZwEgTGFad6EIIaOWT0z9I7KuLrvIT3r4= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791210150; c=relaxed/simple; bh=/ON5VvuE2202uNP/Jo7Gk3gRb5YeJ5IXyS46n2UeLb4=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=cbkGDmIzDXH2kNuJyd6SSZ3RU3DtkwmJHWpnk/LvMZEt3FKKGJnKSsOJUAJWEJlyu6c89Q47DN9nUEvLl5+2ImHQKjDQ7IdnUbfp9adFRjEElG2sIlzEGbmq3CDrFQZlhdEwOeywhBlMU7MZS6yH/nid0mUucX0XVOMH2YHddtM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=GLOvYwfI; arc=none smtp.client-ip=209.85.128.45 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="GLOvYwfI" Received: by mail-wm1-f45.google.com with SMTP id 5b1f17b1804b1-49d0da752ffso15728515e9.3 for ; Mon, 05 Oct 2026 07:22:23 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1791210141; x=1791814941; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=kHDFlSag/ot+2MjJ8umaIa1QqDpwIFqGW2Q5oil11EQ=; b=GLOvYwfIfvUPFcb7+M0mSB76Stt5N4OT4w4aG8PMr9bR4kNai2KPr61eEOitu9UdGj claiorxPY5bqBANKZ0t+ogqcKRPSof/RgjpFSRhYz8BOZiIlw2SUr5VZ7Yy866mfo9s3 vdBb5Ain5MbKNEyZHB68ioSF1ud53m6qyXxEIAXbDHfqNivaHkQ+LNUu3JnkC3Qe+nz0 C8v0aa61iKOeKPKe3E3mBGwm4DPhrBgqfSxGOA3MJ9YEOxWwqBcQtmwoR+vqd0v8ejlL oJOq95a8ufpbQoRBRGY5wMLhMNZkjmxDsTY5A4wch/WEtUg71YZWScc5onkS6bENUt26 VIKg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1791210141; x=1791814941; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=kHDFlSag/ot+2MjJ8umaIa1QqDpwIFqGW2Q5oil11EQ=; b=VzMyRZf61k4x/K6YpLMEEBVkzwQJf91w843YKNhoV39z+5X+DNagfSymlA4Q0wLWIl C88vh733UIJlyZ8J+P3JbV+8v/f7rQ3uhaCHR4vYONzfyuv9NCR2WQcTTiVqFhUeEZnY xy48FrpUtFle2j3AVgdYMCNCZfdlUgzS038zLNXbsqiSSLiMfqCrfjD/uagFaoGgBSnK HsoxK0Aev7cZ2CbatI4EGMT24+lk+FRsdezt2nXRn0TiIjqgENB/EQ/V0oOhERE93F/N 2lnBZMg13b7p2cjnbCpklVK5HLjcEydQMlFaYBwd1zUxcFPBzCgTwERp74AKRbj8mTVZ jGBQ== X-Gm-Message-State: AFuF++mLLIKbU5MtymBITGoKZbRmmtgK1Bp9TsN6jSUE2FXgTGJsm8hp GHHjS5XWdKiwyvIerLmu3xppKufDP5w/GQgx77IOABLZ+Zg9ieJENJHhSdpNZXd+fEQXvFfI X-Gm-Gg: AYBFou0M0gtXRNv1keRj2+hBNizXrHsCQGx28I1VRjI+n2Hx6X3+Km/5wrYKnYvc7SW CiIbryIM2fdn/2VFn5Cvq3WKg7xIl22SaTFfWrcIRFjDr+a79qfYFFbxfDE5uYtPJCVAyAbYXSm 5ImJ+Lb9sW26DA0580hyZeNdHv+oBCL13SvduluuIiIHRvRXKhbWZXCo+ZjX/L7iZHo2BjmSSj1 jo+fXMPDiBkLMUw9gqnIXL8IfPGR3/LwJpgsUKRoFzrdMERJu/q2GZC5rulj37vo0hr7QIn9DXv FbJGnJhwtZ36YYvbt7sIM90WhmNYdldJs0EfyPp9qS0sTVokU/psr0iBdqglk+h5+sOVOZ+IoXH qaxu8xMncWlpiOg8j4eLchxMcEck5h6+mlhsCPoQfeldfbm9kQcs1NK7Quii7fuXFV2dbDD5VmJ hM6yBY/P3XvmcdXFMKcAL9KvTMy5FrwvJCX1KnadgZmIQ0UenDWKbih83rCk1e02nhM1GINjDcw 2Bt0dsf6nggnBIw9p0gtzvA0rG5gqlkko7oax1Ci8qj25WNUV8wZdk+gkxtNnU8rZ/S8YNvDfYx KEcaHyriy4fsZ/dqZ9mnhR+hMxwxwG0ouxKyVw== X-Received: by 2002:a05:600c:1d0a:b0:4a1:6ff4:886a with SMTP id 5b1f17b1804b1-4a16ff48af5mr72658765e9.16.1791210140886; Mon, 05 Oct 2026 07:22:20 -0700 (PDT) Received: from macbook (90-182-211-1.rcp.o2.cz. [90.182.211.1]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-48c622ab5a1sm3881623f8f.36.2026.10.05.07.22.19 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 05 Oct 2026 07:22:20 -0700 (PDT) From: Yusheng Zheng To: bpf@vger.kernel.org Cc: Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , John Fastabend , Emil Tsalapatis , Ihor Solodrai , x86@kernel.org, Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , Leon Hwang , Puranjay Mohan , Hao Sun , Yusheng Zheng Subject: [RFC PATCH bpf-next 0/7] bpf: Inline kfuncs that have a BPF body Date: Mon, 5 Oct 2026 07:22:12 -0700 Message-ID: <20261005142219.33451-1-yunwei356@gmail.com> X-Mailer: git-send-email 2.54.0.windows.1 Precedence: bulk X-Mailing-List: bpf@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit The BPF JIT translates one BPF instruction at a time. Operations that the CPU does in one instruction, such as a rotate, a conditional select or a big-endian load, take several BPF instructions, and they stay several after the JIT. On 27 small compute benchmarks, each compiled from the same C source to BPF and to native code, BPF ran 1.71x (x86-64) and 2.05x (arm64) slower than native code, while compiling the same BPF bytecode with LLVM came within 7% of native code [1]. The loss is in the JIT, not in the bytecode. Our talk in the eBPF track at LPC 2026 [5] covers this gap and the approach of this series. A kfunc could do such an operation, but the call costs about as much as the operation, and the verifier knows nothing about the result. After off = bpf_select64(c, 4, 8), it does not know that off is 4 or 8, so a packet load at data + off needs another bounds check. This series lets a kfunc have a body: a short sequence of BPF instructions that computes the same result, and optionally a callback that emits native code for a call. For example: /* x << n | x >> (-n & 63) */ static const struct bpf_insn rol64_body[] = { ... }; static const struct bpf_kfunc_body bodies[] = { { &body_ids[0], rol64_body, ARRAY_SIZE(rol64_body), rol64_emit }, }; Before its analysis, the verifier replaces each call of such a kfunc with its body, so it analyzes the result like any other BPF code. After the analysis, the call is restored, and the JIT puts native code in its place. The native code uses the registers that already hold the operands, so the argument and result moves around the call go away. If the JIT has no native code for the call, as on other architectures or with constant blinding, the verified body stays in the program instead of a call. Programs call these kfuncs like any other kfunc; arguments named __k must be known constants. There is no new instruction encoding and no UAPI change. Because the verifier analyzes the body at every call, a program that uses these kfuncs costs as much to verify as the same operations written in BPF, which is more than an opaque call. JITs already inline some helper calls, and the verifier already inlines the numeric iterator kfuncs as BPF. This series combines the two, and in every case the body is what tells the verifier the result. Instructions that BPF lacks are the first use, but any small kfunc can have a body, including kfuncs in modules. The native code comes with the kfunc, not from the JIT. By default the x86-64 JIT copies the compiled kfunc and renames its registers, as was suggested for the bitops kfuncs [2]. A copy keeps constant arguments in registers, cannot write the result into the register of an argument, and has only the baseline x86-64 instructions that the kernel is built for, so no movbe or bextr. A kfunc can therefore have an emit callback instead, which gets the registers that hold the operands and the values of the constant arguments. Five of the kfuncs in this series have one. The kfuncs and their x86-64 code are in kernel/bpf/insn_kfuncs/, which can be built as a module. In a loop of dependent operations on x86-64 (Core Ultra 9 285K, KVM; numbers vary by about 5% between boots), plain is the body as the JIT translates it today, emit the code from the callback and copy the copied kfunc: x86-64 code ns per operation operation emitted copied plain emit copy rotate by 13 rol $0xd,%rbx 5 insns 0.38 0.19 0.20 bit extract mov $0x110d,%eax; bextr 10 insns 0.57 0.57 0.57 big-endian load movbe (%rdi),%rax 4 insns 0.57 0.57 1.53 lea lea 0x10(%rbx,%rsi,8),%rbx 8 insns 0.41 0.37 0.41 select test; mov; cmovne same 3.29 1.22 1.21 The emitted code is much shorter, but a copy is about as fast, except for the big-endian load, which is 2.7x slower as a copy, and lea, which is about 10% slower. The select row uses an unpredictable condition; with a predictable one, the branch in plain BPF is 3x faster than cmov. The paper's prototype also had the verifier check a BPF version of each operation and ran native code in its place [1]. It made the 27 benchmarks 1.24x (x86-64) and 1.22x (arm64) faster, and raised the throughput of the Cilium datapath on x86-64 by 7.4%. Where the native code pays off depends on the CPU and the workload: for Katran on arm64, 21 chosen call sites gave 7.3% more throughput, but all 62 gave 0.5% less. - Patch 1 lets a kfunc set give its kfuncs a body and native code. - Patch 2 is the verifier side, in kernel/bpf/kfunc_inline.c. - Patch 3 makes the x86-64 JIT inline the native code of a kfunc, or else a copy of the compiled kfunc. - Patch 4 adds kfuncs with bodies for rotate, select, bit field extract, big-endian load, prefetch, 16-byte copy and lea, under CONFIG_BPF_INSN_KFUNCS. - Patch 5 adds x86-64 code for five of them. - Patches 6 and 7 add selftests and documentation. Testing on x86-64: verifier_kfunc_inline (24 subtests) and kfunc_body_registration pass with CONFIG_BPF_INSN_KFUNCS=y. With bpf_jit_harden=2 the bodies stay in the program, so the subtests that check the JIT output fail as expected, and all results are correct. Questions: 1. Is it acceptable that the verifier replaces calls with their bodies before the analysis and restores them afterwards? 2. Should kfuncs bring native code through an emit callback, or should the JIT only copy compiled kfuncs? The table shows what copying alone costs. 3. Native code is trusted like the rest of the JIT, and the selftests compare it with the bodies on random inputs. Is that enough? 4. The bodies are written by hand. Should they rather be compiled from the kfunc's C source, as in the inlinable kfuncs RFC [3]? Not done yet: arm64 has no native code and runs the bodies. Copying compiled kfuncs would be the first step there. Earlier work: - Leon Hwang's bitops kfuncs [2] emitted native code for kfuncs in the JIT. The review asked to keep the C version as a fallback and suggested copying the compiled kfunc. Here the fallback is the body, and the native code is kept with the kfuncs instead of in the JIT. - Eduard Zingerman's inlinable kfuncs [3] and Puranjay Mohan's iterator inlining [4] inline kfuncs as BPF, so that the verifier sees what they compute. This series does the same and then lets the JIT put native code in place of the call. [1] https://arxiv.org/abs/2606.24213 [2] https://lore.kernel.org/bpf/20260219142933.13904-1-leon.hwang@linux.dev/ [3] https://lore.kernel.org/bpf/20241107175040.1659341-1-eddyz87@gmail.com/ [4] https://lore.kernel.org/bpf/20260804134601.2305303-1-puranjay@kernel.org/ [5] https://lpc.events/event/20/contributions/2445/ Yusheng Zheng (7): bpf: Let kfunc sets give kfuncs a BPF body bpf: Verify calls of kfuncs with a body through the body bpf, x86: Inline native code for kfuncs that have a body bpf: Add kfuncs with bodies for common operations bpf, x86: Add native code for some inline kfuncs selftests/bpf: Test inline kfuncs Documentation/bpf: Describe inline kfuncs Documentation/bpf/kfuncs.rst | 53 ++ arch/x86/net/bpf_jit_comp.c | 181 ++++++ include/linux/bpf_verifier.h | 34 ++ include/linux/btf.h | 27 + include/linux/filter.h | 2 + kernel/bpf/Kconfig | 1 + kernel/bpf/Makefile | 3 +- kernel/bpf/btf.c | 126 ++++ kernel/bpf/const_fold.c | 4 + kernel/bpf/core.c | 9 + kernel/bpf/insn_kfuncs/Kconfig | 16 + kernel/bpf/insn_kfuncs/Makefile | 5 + kernel/bpf/insn_kfuncs/bpf_insn_kfuncs.c | 184 ++++++ kernel/bpf/insn_kfuncs/x86/insn_kfuncs.h | 155 +++++ kernel/bpf/kfunc_inline.c | 273 +++++++++ kernel/bpf/verifier.c | 53 +- tools/testing/selftests/bpf/Makefile | 2 +- tools/testing/selftests/bpf/config | 1 + .../bpf/prog_tests/kfunc_body_registration.c | 18 + .../selftests/bpf/prog_tests/verifier.c | 2 + .../bpf/progs/verifier_kfunc_inline.c | 543 ++++++++++++++++++ .../testing/selftests/bpf/test_kmods/Makefile | 2 +- .../bpf/test_kmods/bpf_test_kfunc_body.c | 121 ++++ .../selftests/bpf/test_kmods/bpf_testmod.c | 77 +++ .../bpf/test_kmods/bpf_testmod_kfunc.h | 4 + 25 files changed, 1877 insertions(+), 19 deletions(-) create mode 100644 kernel/bpf/insn_kfuncs/Kconfig create mode 100644 kernel/bpf/insn_kfuncs/Makefile create mode 100644 kernel/bpf/insn_kfuncs/bpf_insn_kfuncs.c create mode 100644 kernel/bpf/insn_kfuncs/x86/insn_kfuncs.h create mode 100644 kernel/bpf/kfunc_inline.c create mode 100644 tools/testing/selftests/bpf/prog_tests/kfunc_body_registration.c create mode 100644 tools/testing/selftests/bpf/progs/verifier_kfunc_inline.c create mode 100644 tools/testing/selftests/bpf/test_kmods/bpf_test_kfunc_body.c base-commit: 99dc1ba542420db6b8df209744f55cc52466ad91 -- 2.51.1