BPF List
 help / color / mirror / Atom feed
From: Yusheng Zheng <yunwei356@gmail.com>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
	Daniel Borkmann <daniel@iogearbox.net>,
	Andrii Nakryiko <andrii@kernel.org>,
	Eduard Zingerman <eddyz87@gmail.com>,
	Kumar Kartikeya Dwivedi <memxor@gmail.com>,
	Martin KaFai Lau <martin.lau@linux.dev>,
	Song Liu <song@kernel.org>,
	Yonghong Song <yonghong.song@linux.dev>,
	Jiri Olsa <jolsa@kernel.org>,
	John Fastabend <john.fastabend@gmail.com>,
	Emil Tsalapatis <emil@etsalapatis.com>,
	Ihor Solodrai <ihor.solodrai@linux.dev>,
	x86@kernel.org, Thomas Gleixner <tglx@kernel.org>,
	Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
	Dave Hansen <dave.hansen@linux.intel.com>,
	"H . Peter Anvin" <hpa@zytor.com>,
	Leon Hwang <leon.hwang@linux.dev>,
	Puranjay Mohan <puranjay@kernel.org>,
	Hao Sun <sunhao.th@gmail.com>,
	Yusheng Zheng <yunwei356@gmail.com>
Subject: [RFC PATCH bpf-next 3/7] bpf, x86: Inline native code for kfuncs that have a body
Date: Mon,  5 Oct 2026 07:22:15 -0700	[thread overview]
Message-ID: <20261005142219.33451-4-yunwei356@gmail.com> (raw)
In-Reply-To: <20261005142219.33451-1-yunwei356@gmail.com>

Implement bpf_jit_inline_kfunc() for x86-64. The native code for a call
comes from the emit callback of the kfunc's body, which gets the x86
registers that the operands are bound to and the values of the constant
arguments. Without a callback, when it returns an error, or when the
constant arguments differ between paths, the JIT copies the code that
the compiler produced for the kfunc, as was suggested for the bitops
kfuncs [1], with the registers of R0-R5 renamed. Only straight-line
moves, ALU instructions, cmovcc, bswap, prefetch and lea up to the
return are copied: no control flow, rip-relative addressing or
division, no registers but those of R0-R5, r10 and r11, and no implicit
operand, such as %rax in "and $imm, %eax", in a renamed register.
Otherwise the body stays. Copying needs the instruction decoder.

The JIT has no code for any particular kfunc.

[1] https://lore.kernel.org/bpf/CAADnVQLNmQGKf5S5ZNwHYzScYBhnWFmnzLg=5Xxy4SgYKE3EfQ@mail.gmail.com/

Assisted-by: LLM
Signed-off-by: Yusheng Zheng <yunwei356@gmail.com>
---
 arch/x86/net/bpf_jit_comp.c | 181 ++++++++++++++++++++++++++++++++++++
 1 file changed, 181 insertions(+)

diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index 083fcd6cf15b7..6a57109caa97a 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -15,8 +15,12 @@
 #include <linux/memory.h>
 #include <linux/sort.h>
 #include <linux/execmem.h>
+#include <linux/kallsyms.h>
+#include <linux/uaccess.h>
 #include <asm/extable.h>
 #include <asm/ftrace.h>
+#include <asm/insn.h>
+#include <asm/insn-eval.h>
 #include <asm/set_memory.h>
 #include <asm/nospec-branch.h>
 #include <asm/text-patching.h>
@@ -2000,6 +2004,174 @@ static int emit_kfunc_arena_args(struct bpf_prog *bpf_prog,
 	return prog - start;
 }
 
+/*
+ * Registers that copied kfunc code may use: those of R0-R5, which hold the
+ * arguments and the result as for a call, and the scratch registers r10 and
+ * r11. r9 can hold the private frame pointer.
+ */
+#define KFUNC_COPY_REGS	(BIT(0) | BIT(1) | BIT(2) | BIT(6) | BIT(7) | BIT(8) | \
+			 BIT(10) | BIT(11))
+
+/*
+ * Write an instruction of a compiled kfunc with the registers of R0-R5
+ * renamed by @map to those that the verifier bound them to. It must be a
+ * move, ALU or address computation without control flow, prefixes other than
+ * operand size, rip-relative addressing or registers other than
+ * KFUNC_COPY_REGS, and no division, which can trap.
+ */
+static u8 *kfunc_copy_insn(u8 *p, const struct insn *insn, const u8 *c, const u8 *map)
+{
+	u8 op = insn->opcode.bytes[insn->opcode.nbytes - 1], rex = insn->rex_prefix.bytes[0];
+	u8 modrm = insn->modrm.value, sib = insn->sib.value, ext = X86_MODRM_REG(modrm);
+	u8 mod = X86_MODRM_MOD(modrm), reg = ext, rm = 0, idx = 4;
+	bool two = insn->opcode.nbytes == 2, group, opreg, disp8, implicit;
+	u16 regs = 0;
+	int head;
+
+	if (insn->vex_prefix.nbytes || insn->opcode.nbytes > 2 || insn->prefixes.nbytes > 1 ||
+	    (insn->prefixes.nbytes && insn->prefixes.bytes[0] != 0x66))
+		return NULL;
+	if (two) {
+		/* cmovcc, imul, movzx/movsx of words, bswap, prefetch */
+		group = op == 0x18;
+		if ((op & 0xf0) != 0x40 && op != 0xaf && op != 0xb7 && op != 0xbf &&
+		    (op & 0xf8) != 0xc8 && !(group && ext < 4))
+			return NULL;
+	} else {
+		/* add, or, and, sub, xor, cmp, mov, lea, test, imul, shifts, not, neg, mul */
+		group = op == 0x81 || op == 0x83 || op == 0xc1 || op == 0xc7 ||
+			op == 0xd1 || op == 0xd3 || op == 0xf7;
+		if (!(op < 0x40 && (op & 0xf0) != 0x10 && (op & 7) % 2 && (op & 7) < 6) &&
+		    op != 0x63 && op != 0x69 && op != 0x6b && op != 0x85 && op != 0x89 &&
+		    op != 0x8b && op != 0x8d && op != 0x98 && op != 0x99 && op != 0xa9 &&
+		    (op & 0xf8) != 0xb8 &&
+		    !(group && op != 0xc7 && op != 0xf7 && ext != 2 && ext != 3) &&
+		    !(op == 0xc7 && !ext) && !(op == 0xf7 && ext != 1 && ext < 6))
+			return NULL;
+	}
+
+	/* uses %rax, %rcx or %rdx without naming it, as in "and $imm, %eax" */
+	implicit = !two && ((op < 0x40 && (op & 7) == 5) || op == 0x98 || op == 0x99 ||
+			    op == 0xa9 || op == 0xd3 || (op == 0xf7 && (ext == 4 || ext == 5)));
+
+	opreg = (op & 0xf8) == (two ? 0xc8 : 0xb8);
+	if (opreg)
+		rm = (op & 7) + (X86_REX_B(rex) ? 8 : 0);
+	if (insn->modrm.nbytes) {
+		reg = group ? ext : ext + (X86_REX_R(rex) ? 8 : 0);
+		rm = (insn->sib.nbytes ? X86_SIB_BASE(sib) : X86_MODRM_RM(modrm)) +
+		     (X86_REX_B(rex) ? 8 : 0);
+		if (insn->sib.nbytes)
+			idx = X86_SIB_INDEX(sib) + (X86_REX_X(rex) ? 8 : 0);
+		/* rip-relative or absolute */
+		if (!mod && (rm & 7) == 5)
+			return NULL;
+		regs = (group ? 0 : BIT(reg)) | (idx != 4 ? BIT(idx) : 0);
+	}
+	if ((regs | (opreg || insn->modrm.nbytes ? BIT(rm) : 0)) & ~KFUNC_COPY_REGS)
+		return NULL;
+	/* which cannot be renamed, so they must hold their operands as for a call */
+	if (implicit && (map[0] != 0 || map[1] != 1 || map[2] != 2))
+		return NULL;
+
+	/* the instruction again, with the registers renamed */
+	if (!group)
+		reg = map[reg];
+	rm = map[rm];
+	idx = idx != 4 ? map[idx] : 4;
+	if (insn->prefixes.nbytes)
+		*p++ = 0x66;
+	rex = 0x40 | X86_REX_W(rex) | (reg & 8 ? 4 : 0) | (idx & 8 ? 2 : 0) | (rm & 8 ? 1 : 0);
+	if (rex != 0x40)
+		*p++ = rex;
+	if (two)
+		*p++ = 0x0f;
+	*p++ = opreg ? (op & 0xf8) | (rm & 7) : op;
+	if (insn->modrm.nbytes) {
+		/* a base of r13 needs a displacement */
+		disp8 = !mod && (rm & 7) == 5;
+		*p++ = (disp8 ? 1 : mod) << 6 | (reg & 7) << 3 | (insn->sib.nbytes ? 4 : rm & 7);
+		if (insn->sib.nbytes)
+			*p++ = (sib & 0xc0) | (idx & 7) << 3 | (rm & 7);
+		if (disp8)
+			*p++ = 0;
+	}
+	head = insn->prefixes.nbytes + insn->rex_prefix.nbytes + insn->opcode.nbytes +
+	       insn->modrm.nbytes + insn->sib.nbytes;
+	memcpy(p, c + head, insn->length - head);
+	return p + insn->length - head;
+}
+
+/*
+ * Copy the compiled kfunc of an inlined call up to its return, without the
+ * ENDBR and NOPs at its entry, with the operands renamed by @map.
+ */
+static int kfunc_copy(const struct bpf_kfunc_inline *in, const u8 *map, u8 *buf)
+{
+	u8 code[BPF_KFUNC_INLINE_MAX + MAX_INSN_SIZE], *c, *p = buf;
+	unsigned long size, off, ret;
+	struct insn insn;
+	int pos;
+
+	if (!kallsyms_lookup_size_offset(in->addr, &size, &off) || off)
+		return -EINVAL;
+	size = min(size, sizeof(code));
+	if (copy_from_kernel_nofault(code, (void *)in->addr, size))
+		return -EFAULT;
+	for (pos = 0; pos < size; pos += insn.length) {
+		if (insn_decode(&insn, code + pos, size - pos, INSN_MODE_64))
+			return -EINVAL;
+		c = code + pos;
+		/* ret, or a jump to the return thunk */
+		ret = in->addr + pos + insn.length + insn.immediate.value;
+		if ((insn.length == 1 && c[0] == 0xc3) ||
+		    (c[0] == 0xe9 && (ret == (unsigned long)x86_return_thunk ||
+				      ret == (unsigned long)__x86_return_thunk)))
+			return p > buf ? p - buf : -EINVAL;
+		if ((insn.length == 4 && is_endbr((u32 *)c)) || insn_is_nop(&insn))
+			continue;
+		if (insn_is_rex2(&insn))
+			return -EINVAL;
+		/* renaming adds at most a REX prefix and a displacement */
+		if (p - buf + insn.length + 2 > BPF_KFUNC_INLINE_MAX)
+			return -E2BIG;
+		p = kfunc_copy_insn(p, &insn, c, map);
+		if (!p)
+			return -EINVAL;
+	}
+	return -EINVAL;
+}
+
+static u8 x86_reg(u32 reg)
+{
+	return reg2hex[reg] + (is_ereg(reg) ? 8 : 0);
+}
+
+/*
+ * Get native code for an inlined kfunc call, with the operands in the x86
+ * registers that the verifier bound them to: the code from the emit callback
+ * of the kfunc's body, or else a copy of the compiled kfunc.
+ */
+int bpf_jit_inline_kfunc(const struct bpf_kfunc_inline *in, u8 *buf)
+{
+	u8 reg[MAX_BPF_FUNC_REG_ARGS + 1], map[16];
+	int i, len;
+
+	for (i = 0; i < ARRAY_SIZE(map); i++)
+		map[i] = i;
+	for (i = BPF_REG_0; i <= BPF_REG_5; i++) {
+		reg[i] = x86_reg(in->reg[i]);
+		map[x86_reg(i)] = reg[i];
+	}
+	/* copying needs the instruction decoder */
+	if (in->copy && IS_ENABLED(CONFIG_INSTRUCTION_DECODER))
+		return kfunc_copy(in, map, buf);
+	if (in->copy || !in->body->emit)
+		return -EOPNOTSUPP;
+	len = in->body->emit(reg, in->imm, buf);
+	return len > 0 && len <= BPF_KFUNC_INLINE_MAX ? len : -EINVAL;
+}
+
 static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *addrs, u8 *image,
 		  u8 *rw_image, int oldproglen, struct jit_context *ctx, bool jmp_padding)
 {
@@ -2952,6 +3124,15 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
 			if (!imm32)
 				return -EINVAL;
 			if (src_reg == BPF_PSEUDO_KFUNC_CALL) {
+				const struct bpf_kfunc_inline *in;
+
+				/* an inlined kfunc call gets its native code */
+				in = bpf_kfunc_native(env, insn_idx);
+				if (in) {
+					memcpy(prog, in->image, in->image_len);
+					prog += in->image_len;
+					break;
+				}
 				fm = bpf_jit_find_kfunc_model(bpf_prog, insn);
 				if (!fm)
 					return -EINVAL;
-- 
2.51.1


  parent reply	other threads:[~2026-10-05 14:22 UTC|newest]

Thread overview: 13+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-05 14:22 [RFC PATCH bpf-next 0/7] bpf: Inline kfuncs that have a BPF body Yusheng Zheng
2026-10-05 14:22 ` [RFC PATCH bpf-next 1/7] bpf: Let kfunc sets give kfuncs " Yusheng Zheng
2026-10-05 14:38   ` sashiko-bot
2026-10-05 15:16   ` bot+bpf-ci
2026-10-05 14:22 ` [RFC PATCH bpf-next 2/7] bpf: Verify calls of kfuncs with a body through the body Yusheng Zheng
2026-10-05 14:41   ` sashiko-bot
2026-10-05 14:22 ` Yusheng Zheng [this message]
2026-10-05 14:22 ` [RFC PATCH bpf-next 4/7] bpf: Add kfuncs with bodies for common operations Yusheng Zheng
2026-10-05 14:39   ` sashiko-bot
2026-10-05 14:22 ` [RFC PATCH bpf-next 5/7] bpf, x86: Add native code for some inline kfuncs Yusheng Zheng
2026-10-05 14:22 ` [RFC PATCH bpf-next 6/7] selftests/bpf: Test " Yusheng Zheng
2026-10-05 14:22 ` [RFC PATCH bpf-next 7/7] Documentation/bpf: Describe " Yusheng Zheng
2026-10-05 15:16   ` bot+bpf-ci

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261005142219.33451-4-yunwei356@gmail.com \
    --to=yunwei356@gmail.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bp@alien8.de \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=dave.hansen@linux.intel.com \
    --cc=eddyz87@gmail.com \
    --cc=emil@etsalapatis.com \
    --cc=hpa@zytor.com \
    --cc=ihor.solodrai@linux.dev \
    --cc=john.fastabend@gmail.com \
    --cc=jolsa@kernel.org \
    --cc=leon.hwang@linux.dev \
    --cc=martin.lau@linux.dev \
    --cc=memxor@gmail.com \
    --cc=mingo@redhat.com \
    --cc=puranjay@kernel.org \
    --cc=song@kernel.org \
    --cc=sunhao.th@gmail.com \
    --cc=tglx@kernel.org \
    --cc=x86@kernel.org \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox