BPF List
 help / color / mirror / Atom feed
From: Yonghong Song <yonghong.song@linux.dev>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
	Andrii Nakryiko <andrii@kernel.org>,
	Daniel Borkmann <daniel@iogearbox.net>,
	Eduard Zingerman <eddyz87@gmail.com>,
	kernel-team@fb.com
Subject: [PATCH bpf-next v2 11/13] selftests/bpf: Add inline-asm and subprog tests for R0:R2 returns
Date: Tue,  4 Aug 2026 13:36:18 -0700	[thread overview]
Message-ID: <20260804203618.1878749-1-yonghong.song@linux.dev> (raw)
In-Reply-To: <20260804203522.1869244-1-yonghong.song@linux.dev>

Add inline-asm tests, which do not depend on the compiler lowering a
register-pair return and so run regardless of the LLVM version, covering
what the C tests cannot reach. aggregate_ret_func.c exercises BPF-to-BPF
returns: a global subprogram whose R2 the caller may read, ones that leave
R2 uninitialised or holding a pointer, a static subprogram whose R2 stays
precise under backtracking, R2 liveness across a call, and a >8 byte return
at program exit. Six kfuncs returning aggregates by value are added to
bpf_testmod, and aggregate_ret_run.c calls them from inline asm to check
what comes back in R0:R2.

A negative arena test is added as well: a global subprogram with a
register-pair return that leaves an arena pointer in R2 is rejected, since
an arena pointer is only a valid return value when it is returned in R0
alone.

Three cases cover the boundaries of the new convention:

 - A return value larger than 16 bytes does not fit in R0:R2 and is
   rejected by btf_distill_func_proto(), ahead of the KF_FASTCALL and
   JIT-capability checks; one of the new kfuncs returns a 24-byte struct
   for this. The equivalent for a BPF subprogram cannot be written in C:
   from LLVM 23 on, a by-value return larger than 16 bytes is lowered to an
   sret pointer argument and the BTF the verifier reads says the function
   returns void, so the size bound in btf_validate_return_type() only
   guards hand-crafted BTF.

 - A static subprogram returning a struct that contains a pointer is
   accepted, and the caller can use the returned pointer. Unlike a global
   subprogram, whose caller models the return as an opaque scalar pair, a
   static one is verified inline, so prepare_func_exit() hands the caller
   real register state and the pointer stays tracked.

 - An extension cannot replace a function returning more than 8 bytes.
   btf_check_type_match() does not catch this, since it compares return
   types by btf_type->info alone and both an __int128 and a __u64 are
   BTF_KIND_INT with no vlen, so the rejection has to come from
   bpf_check_attach_target(). The test reuses the freplace failure harness
   in fexit_bpf2bpf.c, with aggregate_ret_target.c providing a target whose
   global subprogram returns in R0:R2.

The kfunc tests need the JIT to place the second half of a return value
into R2, which bpf_add_kfunc_call() only allows where
bpf_jit_supports_kfunc_ret_reg_pair() is true. In aggregate_ret_kfunc.c the
two tests that depend on getting past that check are tagged
__arch_x86_64/__arch_arm64/__arch_riscv64; the others are rejected earlier
(on KF_FASTCALL, on a >16 byte return, and on reading R2 after an 8-byte
struct return) and run everywhere. In aggregate_ret_run.c the kfunc-calling
programs are dropped from the object when the load reports -EOPNOTSUPP and
their subtests are skipped, and the __int128 inline-asm test is split into
a BPF-to-BPF half, which needs no JIT capability and runs everywhere, and a
kfunc half.

Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 .../selftests/bpf/prog_tests/aggregate_ret.c  |  51 +++
 .../selftests/bpf/prog_tests/fexit_bpf2bpf.c  |  15 +
 .../selftests/bpf/progs/aggregate_ret_func.c  | 420 ++++++++++++++++++
 .../selftests/bpf/progs/aggregate_ret_kfunc.c | 127 ++++++
 .../selftests/bpf/progs/aggregate_ret_run.c   | 168 +++++++
 .../bpf/progs/aggregate_ret_target.c          |  29 ++
 .../selftests/bpf/progs/freplace_ret_pair.c   |  20 +
 .../selftests/bpf/progs/verifier_arena.c      |  38 ++
 .../selftests/bpf/test_kmods/bpf_testmod.c    |  50 +++
 .../bpf/test_kmods/bpf_testmod_kfunc.h        |  35 ++
 10 files changed, 953 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/progs/aggregate_ret_func.c
 create mode 100644 tools/testing/selftests/bpf/progs/aggregate_ret_kfunc.c
 create mode 100644 tools/testing/selftests/bpf/progs/aggregate_ret_run.c
 create mode 100644 tools/testing/selftests/bpf/progs/aggregate_ret_target.c
 create mode 100644 tools/testing/selftests/bpf/progs/freplace_ret_pair.c

diff --git a/tools/testing/selftests/bpf/prog_tests/aggregate_ret.c b/tools/testing/selftests/bpf/prog_tests/aggregate_ret.c
index 42017f89c4da..2a0c6996f0b9 100644
--- a/tools/testing/selftests/bpf/prog_tests/aggregate_ret.c
+++ b/tools/testing/selftests/bpf/prog_tests/aggregate_ret.c
@@ -5,6 +5,9 @@
 #include "aggregate_ret_struct_c.skel.h"
 #include "aggregate_ret_union_c.skel.h"
 #include "aggregate_ret_kfunc_c.skel.h"
+#include "aggregate_ret_run.skel.h"
+#include "aggregate_ret_func.skel.h"
+#include "aggregate_ret_kfunc.skel.h"
 
 static void run_prog(struct bpf_program *prog, bool supported)
 {
@@ -116,10 +119,58 @@ static void test_kfunc_c(void)
 	aggregate_ret_kfunc_c__destroy(skel);
 }
 
+static void test_run(void)
+{
+	struct aggregate_ret_run *skel;
+	bool kfunc_ok = true;
+	int err;
+
+	skel = aggregate_ret_run__open();
+	if (!ASSERT_OK_PTR(skel, "skel_run_open"))
+		return;
+
+	err = aggregate_ret_run__load(skel);
+	if (err == -EOPNOTSUPP) {
+		kfunc_ok = false;
+		aggregate_ret_run__destroy(skel);
+
+		skel = aggregate_ret_run__open();
+		if (!ASSERT_OK_PTR(skel, "skel_run_reopen"))
+			return;
+
+		bpf_program__set_autoload(skel->progs.aggregate_ret_asm_kfunc_test, false);
+		bpf_program__set_autoload(skel->progs.aggregate_ret_struct_test, false);
+		bpf_program__set_autoload(skel->progs.aggregate_ret_union_test, false);
+
+		err = aggregate_ret_run__load(skel);
+	}
+	if (!ASSERT_OK(err, "skel_run_load"))
+		goto out;
+
+	if (test__start_subtest("asm"))
+		run_prog(skel->progs.aggregate_ret_asm_test, true);
+
+	if (test__start_subtest("asm_kfunc"))
+		run_prog(skel->progs.aggregate_ret_asm_kfunc_test, kfunc_ok);
+
+	if (test__start_subtest("struct"))
+		run_prog(skel->progs.aggregate_ret_struct_test, kfunc_ok);
+
+	if (test__start_subtest("union"))
+		run_prog(skel->progs.aggregate_ret_union_test, kfunc_ok);
+
+out:
+	aggregate_ret_run__destroy(skel);
+}
+
 void test_aggregate_ret(void)
 {
 	test_int128_c();
 	test_struct_c();
 	test_union_c();
 	test_kfunc_c();
+	test_run();
+
+	RUN_TESTS(aggregate_ret_func);
+	RUN_TESTS(aggregate_ret_kfunc);
 }
diff --git a/tools/testing/selftests/bpf/prog_tests/fexit_bpf2bpf.c b/tools/testing/selftests/bpf/prog_tests/fexit_bpf2bpf.c
index 4a87d7163c8c..a40d63a87d7f 100644
--- a/tools/testing/selftests/bpf/prog_tests/fexit_bpf2bpf.c
+++ b/tools/testing/selftests/bpf/prog_tests/fexit_bpf2bpf.c
@@ -440,6 +440,19 @@ static void test_func_replace_int_with_void(void)
 				     " doesn't match type INT of global_func2()");
 }
 
+static void test_func_replace_ret_pair(void)
+{
+	const char *msg = "Cannot replace function agg_ret_target_func with a >8 byte return";
+
+	/*
+	 * An extension cannot replace a function whose return value comes back
+	 * in the R0:R2 pair: the extension's own return is capped at 8 bytes,
+	 * so it would leave R2 stale for the target's callers.
+	 */
+	test_obj_load_failure_common("freplace_ret_pair.bpf.o",
+				     "./aggregate_ret_target.bpf.o", msg);
+}
+
 static int find_prog_btf_id(const char *name, __u32 attach_prog_fd)
 {
 	struct bpf_prog_info info = {};
@@ -605,6 +618,8 @@ void serial_test_fexit_bpf2bpf(void)
 		test_func_replace_progmap();
 	if (test__start_subtest("freplace_int_with_void"))
 		test_func_replace_int_with_void();
+	if (test__start_subtest("freplace_ret_pair"))
+		test_func_replace_ret_pair();
 	if (test__start_subtest("freplace_void"))
 		test_func_replace_void();
 }
diff --git a/tools/testing/selftests/bpf/progs/aggregate_ret_func.c b/tools/testing/selftests/bpf/progs/aggregate_ret_func.c
new file mode 100644
index 000000000000..0bc18450a46e
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/aggregate_ret_func.c
@@ -0,0 +1,420 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+
+typedef unsigned __int128 u128;
+
+__naked u128 global_agg_good(void)
+{
+	asm volatile (
+	"r0 = 0x1234;"	/* low 64 bits */
+	"r2 = 0x5678;"	/* high 64 bits */
+	"exit;"
+	);
+}
+
+__naked u128 global_agg_bad(void)
+{
+	asm volatile (
+	"r0 = 0;"
+	"exit;"
+	);
+}
+
+__naked u128 global_agg_bad_ptr(void)
+{
+	asm volatile (
+	"r0 = 0;"
+	"r2 = r10;"
+	"exit;"
+	);
+}
+
+SEC("tc")
+__success __retval(0)
+int aggregate_ret_global(void *ctx)
+{
+	__u64 lo, hi;
+
+	asm volatile (
+	"call %[global_agg_good];"
+	"%[lo] = r0;"
+	"%[hi] = r2;"
+	: [lo]"=r"(lo), [hi]"=r"(hi)
+	: __imm(global_agg_good)
+	: "r0", "r1", "r2", "r3", "r4", "r5");
+	if (lo != 0x1234)
+		return 1;
+	if (hi != 0x5678)
+		return 2;
+	return 0;
+}
+
+SEC("tc")
+__failure __msg("R2 !read_ok")
+__naked int aggregate_ret_global_fail(void)
+{
+	asm volatile (
+	"call %[global_agg_bad];"
+	"r0 = r2;"
+	"exit;"
+	:
+	: __imm(global_agg_bad)
+	: __clobber_all);
+}
+
+SEC("tc")
+__failure __msg("At subprogram exit the register R2 is not a scalar value")
+__naked int aggregate_ret_global_ptr_fail(void)
+{
+	asm volatile (
+	"call %[global_agg_bad_ptr];"
+	"r0 = r2;"
+	"exit;"
+	:
+	: __imm(global_agg_bad_ptr)
+	: __clobber_all);
+}
+
+static __naked __noinline u128 static_agg_bad_ptr(void)
+{
+	asm volatile (
+	"r0 = 0;"
+	"r2 = r10;"	/* stack pointer placed in the second return register */
+	"exit;"
+	);
+}
+
+/*
+ * R2 is caller-saved and only copied from the callee at exit; a PTR_TO_STACK
+ * left in it is turned into an uninitialized R2 in the caller. A caller that
+ * never reads R2 is therefore unaffected and loads fine.
+ */
+SEC("tc")
+__success __retval(0)
+__naked int aggregate_ret_static_ptr_unused(void)
+{
+	asm volatile (
+	"call %[static_agg_bad_ptr];"
+	"r0 = 0;"		/* R2 holds a stack pointer but is never read */
+	"exit;"
+	:
+	: __imm(static_agg_bad_ptr)
+	: __clobber_all);
+}
+
+/* But a caller that does read the returned stack pointer is rejected. */
+SEC("tc")
+__failure __msg("R2 !read_ok")
+__naked int aggregate_ret_static_ptr_read_fail(void)
+{
+	asm volatile (
+	"call %[static_agg_bad_ptr];"
+	"r0 = r2;"		/* using the returned stack pointer is rejected */
+	"exit;"
+	:
+	: __imm(static_agg_bad_ptr)
+	: __clobber_all);
+}
+
+static __naked __noinline u128 static_agg_no_r2(void)
+{
+	asm volatile (
+	"r0 = 0;"
+	"exit;"
+	);
+}
+
+SEC("tc")
+__failure __msg("R2 !read_ok")
+__naked int aggregate_ret_static_uninit_fail(void)
+{
+	asm volatile (
+	"call %[static_agg_no_r2];"
+	"r0 = r2;"
+	"exit;"
+	:
+	: __imm(static_agg_no_r2)
+	: __clobber_all);
+}
+
+static __naked __noinline u128 static_agg_precise(void)
+{
+	asm volatile (
+	"r0 = 0;"
+	"r2 = 4;"	/* second half; its value is made precise below */
+	"exit;"
+	);
+}
+
+SEC("tc")
+__success __retval(0)
+__log_level(2)
+__msg("mark_precise: frame0: last_idx 5 first_idx 0 subseq_idx -1")
+__msg("mark_precise: frame0: regs=r6 stack= before 4: (07) r1 += -8")
+__msg("mark_precise: frame0: regs=r6 stack= before 3: (bf) r1 = r10")
+__msg("mark_precise: frame0: regs=r6 stack= before 2: (57) r6 &= 7")
+__msg("mark_precise: frame0: regs=r6 stack= before 1: (bf) r6 = r2")
+__msg("mark_precise: frame0: regs=r2 stack= before 12: (95) exit")
+__msg("mark_precise: frame1: regs=r2 stack= before 11: (b7) r2 = 4")
+__naked int aggregate_ret_static_precise(void)
+{
+	asm volatile (
+	"call %[static_agg_precise];"
+	"r6 = r2;"		/* derived from the aggregate's second half */
+	"r6 &= 7;"		/* keep it in [0, 7] to index the stack */
+	"r1 = r10;"
+	"r1 += -8;"
+	"r1 += r6;"		/* ptr += scalar marks r6 (hence R2) precise */
+	"r0 = 0;"
+	"*(u8 *)(r1 + 0) = r0;"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm(static_agg_precise)
+	: __clobber_all);
+}
+
+SEC("tc")
+__success __retval(0)
+__log_level(2)
+__msg("mark_precise: frame0: last_idx 5 first_idx 0 subseq_idx -1")
+__msg("mark_precise: frame0: regs=r6 stack= before 4: (07) r1 += -8")
+__msg("mark_precise: frame0: regs=r6 stack= before 3: (bf) r1 = r10")
+__msg("mark_precise: frame0: regs=r6 stack= before 2: (57) r6 &= 7")
+__msg("mark_precise: frame0: regs=r6 stack= before 1: (bf) r6 = r2")
+__msg("mark_precise: frame0: regs=r2 stack= before 0: (85) call pc+9")
+__naked int aggregate_ret_global_precise(void)
+{
+	asm volatile (
+	"call %[global_agg_good];"
+	"r6 = r2;"		/* derived from the aggregate's second half */
+	"r6 &= 7;"		/* keep it in [0, 7] to index the stack */
+	"r1 = r10;"
+	"r1 += -8;"
+	"r1 += r6;"		/* ptr += scalar marks r6 (hence R2) precise */
+	"r0 = 0;"
+	"*(u8 *)(r1 + 0) = r0;"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm(global_agg_good)
+	: __clobber_all);
+}
+
+SEC("tc")
+__failure __msg("return value larger than 8 bytes is not supported at program exit")
+__naked u128 aggregate_ret_entry_fail(void)
+{
+	asm volatile (
+	"r0 = 0;"
+	"r2 = 0;"
+	"exit;"
+	);
+}
+
+#if __clang_major__ >= 23
+
+struct pair {
+	__u64 hi;
+	__u64 lo;
+};
+
+union upair {
+	__u64 halves[2];
+	struct {
+		__u64 lo;
+		__u64 hi;
+	} parts;
+};
+
+/* A by-value struct that smuggles a pointer, which must be rejected. */
+struct with_ptr {
+	void *p;
+	__u64 x;
+};
+
+/* A by-value union that smuggles a pointer, which must be rejected too. */
+union upair_with_ptr {
+	void *p;
+	__u64 halves[2];
+};
+
+/* Global subprogram returning a scalar-only 16-byte struct in R0:R2. */
+__naked struct pair global_ret_struct(void)
+{
+	asm volatile (
+	"r0 = 0x1234;"	/* struct's first half */
+	"r2 = 0x5678;"	/* struct's second half */
+	"exit;"
+	);
+}
+
+/* Global subprogram returning a scalar-only 16-byte union in R0:R2. */
+__naked union upair global_ret_union(void)
+{
+	asm volatile (
+	"r0 = 0x1234;"
+	"r2 = 0x5678;"
+	"exit;"
+	);
+}
+
+SEC("tc")
+__success __retval(0)
+int aggregate_ret_global_struct(void *ctx)
+{
+	__u64 lo, hi;
+
+	asm volatile (
+	"call %[global_ret_struct];"
+	"%[lo] = r0;"
+	"%[hi] = r2;"
+	: [lo]"=r"(lo), [hi]"=r"(hi)
+	: __imm(global_ret_struct)
+	: "r0", "r1", "r2", "r3", "r4", "r5");
+	if (lo != 0x1234)
+		return 1;
+	if (hi != 0x5678)
+		return 2;
+	return 0;
+}
+
+SEC("tc")
+__success __retval(0)
+int aggregate_ret_global_union(void *ctx)
+{
+	__u64 lo, hi;
+
+	asm volatile (
+	"call %[global_ret_union];"
+	"%[lo] = r0;"
+	"%[hi] = r2;"
+	: [lo]"=r"(lo), [hi]"=r"(hi)
+	: __imm(global_ret_union)
+	: "r0", "r1", "r2", "r3", "r4", "r5");
+	if (lo != 0x1234)
+		return 1;
+	if (hi != 0x5678)
+		return 2;
+	return 0;
+}
+
+__naked struct with_ptr global_ret_struct_ptr(void)
+{
+	asm volatile (
+	"r0 = 0;"
+	"r2 = 0;"
+	"exit;"
+	);
+}
+
+SEC("tc")
+__failure __msg("Global function global_ret_struct_ptr() has unsupported return type")
+__naked int aggregate_ret_global_struct_ptr_fail(void)
+{
+	asm volatile (
+	"call %[global_ret_struct_ptr];"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm(global_ret_struct_ptr)
+	: __clobber_all);
+}
+
+__naked union upair_with_ptr global_ret_union_ptr(void)
+{
+	asm volatile (
+	"r0 = 0;"
+	"r2 = 0;"
+	"exit;"
+	);
+}
+
+SEC("tc")
+__failure __msg("Global function global_ret_union_ptr() has unsupported return type")
+__naked int aggregate_ret_global_union_ptr_fail(void)
+{
+	asm volatile (
+	"call %[global_ret_union_ptr];"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm(global_ret_union_ptr)
+	: __clobber_all);
+}
+
+#endif /* __clang_major__ >= 23 */
+
+static __naked u128 agg_callee(void)
+{
+	asm volatile (
+	"r0 = 1;"
+	"r2 = 2;"
+	"exit;"
+	);
+}
+
+SEC("tc")
+__log_level(2)
+__msg("Live regs before insn:")
+/*
+ * R2 is read at the exit of agg_callee() (insn 5), which returns a pair, but
+ * not at the exit of this program (insn 2), which returns an int.
+ */
+__msg("0: .12345.... (85) call pc+2")
+__msg("1: ..2....... (bf) r0 = r2")
+__msg("2: 0......... (95) exit")
+__msg("3: .......... (b7) r0 = 1")
+__msg("4: 0......... (b7) r2 = 2")
+__msg("5: 0.2....... (95) exit")
+__naked int aggregate_ret_live(void)
+{
+	asm volatile (
+	"call %[agg_callee];"
+	"r0 = r2;"
+	"exit;"
+	:
+	: [agg_callee]"i"(agg_callee)
+	: __clobber_all);
+}
+
+/*
+ * A static subprogram is verified inline, so prepare_func_exit() hands the
+ * caller the callee's actual R0:R2 register state rather than an opaque scalar
+ * pair. A pointer in the returned struct therefore stays tracked and is usable
+ * by the caller, which is why btf_validate_return_type() does not apply the
+ * scalar-only restriction to a local function. Return the context pointer as
+ * the upper half and dereference it in the caller.
+ */
+struct ptr_pair {
+	void *p;
+	__u64 x;
+};
+
+static __naked __noinline struct ptr_pair static_ret_ptr_pair(void)
+{
+	asm volatile (
+	"r0 = 0;"
+	"r2 = r1;"
+	"exit;"
+	);
+}
+
+SEC("tc")
+__success __retval(0)
+__naked int aggregate_ret_static_ptr_pair(void)
+{
+	asm volatile (
+	"call %[static_ret_ptr_pair];"
+	"r1 = *(u32 *)(r2 + 0);"	/* deref the returned ctx pointer */
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm(static_ret_ptr_pair)
+	: __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/aggregate_ret_kfunc.c b/tools/testing/selftests/bpf/progs/aggregate_ret_kfunc.c
new file mode 100644
index 000000000000..ef10d765f738
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/aggregate_ret_kfunc.c
@@ -0,0 +1,127 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "../test_kmods/bpf_testmod_kfunc.h"
+
+/*
+ * Reference kfunc addresses to force those BTF to be emitted. Taking the address
+ * (rather than calling) avoids any dependence on the compiler lowering an
+ * __int128 or struct return value, which the BPF backend only supports from
+ * LLVM 23 on.
+ */
+void __kfunc_btf_root(void)
+{
+	asm volatile (""
+	:
+	: "r"(&bpf_kfunc_call_test_i128),
+	  "r"(&bpf_kfunc_call_test_ret_fastcall),
+	  "r"(&bpf_kfunc_call_test_ret_ptr),
+	  "r"(&bpf_kfunc_call_test_ret_ii),
+	  "r"(&bpf_kfunc_call_test_ret_big));
+}
+
+/*
+ * bpf_add_kfunc_call() rejects a kfunc returning more than 8 bytes unless the
+ * JIT advertises bpf_jit_supports_kfunc_ret_reg_pair(), so a test that has to
+ * get past it is tagged with the architectures implementing it. The two tests
+ * below that are rejected earlier, on KF_FASTCALL or on reading R2 after an
+ * 8-byte struct return, behave the same everywhere and are not tagged.
+ */
+
+SEC("tc")
+__arch_x86_64 __arch_arm64 __arch_riscv64
+__success __retval(0)
+__log_level(2)
+__msg("mark_precise: frame0: last_idx 7 first_idx 0 subseq_idx -1")
+__msg("mark_precise: frame0: regs=r6 stack= before 6: (07) r1 += -8")
+__msg("mark_precise: frame0: regs=r6 stack= before 5: (bf) r1 = r10")
+__msg("mark_precise: frame0: regs=r6 stack= before 4: (57) r6 &= 7")
+__msg("mark_precise: frame0: regs=r6 stack= before 3: (bf) r6 = r2")
+__msg("mark_precise: frame0: regs=r2 stack= before 2: (85) call bpf_kfunc_call_test_i128")
+__naked int aggregate_ret_kfunc_precise(void)
+{
+	asm volatile (
+	"r1 = 1;"
+	"r2 = 2;"
+	"call %[bpf_kfunc_call_test_i128];"
+	"r6 = r2;"		/* second return half */
+	"r6 &= 7;"		/* keep it in [0, 7] to index the stack */
+	"r1 = r10;"
+	"r1 += -8;"
+	"r1 += r6;"		/* ptr += scalar marks r6 (hence R2) precise */
+	"r0 = 0;"
+	"*(u8 *)(r1 + 0) = r0;"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm(bpf_kfunc_call_test_i128)
+	: __clobber_all);
+}
+
+SEC("tc")
+__failure __msg("kfunc bpf_kfunc_call_test_ret_fastcall with >8-byte return is not supported with KF_FASTCALL")
+__naked int aggregate_ret_kfunc_fastcall_fail(void)
+{
+	asm volatile (
+	"r1 = 1;"
+	"r2 = 2;"
+	"call %[bpf_kfunc_call_test_ret_fastcall];"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm(bpf_kfunc_call_test_ret_fastcall)
+	: __clobber_all);
+}
+
+SEC("tc")
+__arch_x86_64 __arch_arm64 __arch_riscv64
+__failure __msg("is not composed of scalars")
+__naked int aggregate_ret_kfunc_ptr_fail(void)
+{
+	asm volatile (
+	"r1 = 0;"
+	"call %[bpf_kfunc_call_test_ret_ptr];"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm(bpf_kfunc_call_test_ret_ptr)
+	: __clobber_all);
+}
+
+SEC("tc")
+__failure __msg("R2 !read_ok")
+__naked int aggregate_ret_kfunc_small_no_r2(void)
+{
+	asm volatile (
+	"r1 = 0;"
+	"r2 = 0;"
+	"call %[bpf_kfunc_call_test_ret_ii];"
+	"r0 = r2;"	/* R2 is not a return register for a <=8 byte struct */
+	"exit;"
+	:
+	: __imm(bpf_kfunc_call_test_ret_ii)
+	: __clobber_all);
+}
+
+/*
+ * A return value larger than 16 bytes does not fit in R0:R2 and is rejected by
+ * btf_distill_func_proto(), before the KF_FASTCALL and JIT-capability checks,
+ * so this behaves the same on every architecture.
+ */
+SEC("tc")
+__failure __msg("The function bpf_kfunc_call_test_ret_big return type STRUCT is unsupported")
+__naked int aggregate_ret_kfunc_too_big_fail(void)
+{
+	asm volatile (
+	"r1 = 0;"
+	"call %[bpf_kfunc_call_test_ret_big];"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm(bpf_kfunc_call_test_ret_big)
+	: __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/aggregate_ret_run.c b/tools/testing/selftests/bpf/progs/aggregate_ret_run.c
new file mode 100644
index 000000000000..fac813531b4f
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/aggregate_ret_run.c
@@ -0,0 +1,168 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "../test_kmods/bpf_testmod_kfunc.h"
+
+typedef unsigned __int128 u128;
+
+/*
+ * Reference kfunc addresses to force those BTF to be emitted. Taking the address
+ * (rather than calling) avoids any dependence on the compiler lowering an __int128
+ * or struct return value, which the BPF backend only supports from LLVM 23 on.
+ */
+void __kfunc_btf_root(void)
+{
+	asm volatile (""
+	:
+	: "r"(&bpf_kfunc_call_test_i128),
+	  "r"(&bpf_kfunc_call_test_ret_pair),
+	  "r"(&bpf_kfunc_call_test_ret_li),
+	  "r"(&bpf_kfunc_call_test_ret_ii),
+	  "r"(&bpf_kfunc_call_test_ret_uu));
+}
+
+#define I128_ASM_LO 0xABCDabcd12345678ULL
+#define I128_ASM_HI 0x1234567890abcdefULL
+
+static __naked __noinline u128 make_i128_asm(void)
+{
+	asm volatile (
+	"r0 = %[lo] ll;"	/* low 64 bits */
+	"r2 = %[hi] ll;"	/* high 64 bits */
+	"exit;"
+	:
+	: __imm_const(lo, I128_ASM_LO), __imm_const(hi, I128_ASM_HI)
+	);
+}
+
+SEC("tc")
+int aggregate_ret_asm_test(struct __sk_buff *skb)
+{
+	__u64 lo, hi;
+
+	asm volatile (
+	"call %[callee];"
+	"%[lo] = r0;"
+	"%[hi] = r2;"
+	: [lo]"=r"(lo), [hi]"=r"(hi)
+	: [callee]"i"(make_i128_asm)
+	: "r0", "r1", "r2", "r3", "r4", "r5"
+	);
+	if (lo != I128_ASM_LO)
+		return 1;
+	if (hi != I128_ASM_HI)
+		return 2;
+
+	return 0;
+}
+
+SEC("tc")
+int aggregate_ret_asm_kfunc_test(struct __sk_buff *skb)
+{
+	__u64 a = skb->len;
+	__u64 b = skb->len ^ 0xdeadbeefULL;
+	__u64 lo, hi;
+
+	asm volatile (
+	"r1 = %[a];"
+	"r2 = %[b];"
+	"call %[kfunc];"
+	"%[lo] = r0;"
+	"%[hi] = r2;"
+	: [lo]"=r"(lo), [hi]"=r"(hi)
+	: [a]"r"(a), [b]"r"(b), [kfunc]"i"(bpf_kfunc_call_test_i128)
+	: "r0", "r1", "r2", "r3", "r4", "r5"
+	);
+	if (hi != a + b)
+		return 1;
+	if (lo != a - b)
+		return 2;
+
+	return 0;
+}
+
+SEC("tc")
+int aggregate_ret_struct_test(struct __sk_buff *skb)
+{
+	__u64 a = skb->len;
+	__u64 b = skb->len ^ 0xdeadbeefULL;
+	__u64 lo, hi;
+
+	/* struct { u64 hi; u64 lo; }: R0 = hi, R2 = lo. */
+	asm volatile (
+	"r1 = %[a];"
+	"r2 = %[b];"
+	"call %[kfunc];"
+	"%[lo] = r0;"
+	"%[hi] = r2;"
+	: [lo]"=r"(lo), [hi]"=r"(hi)
+	: [a]"r"(a), [b]"r"(b), [kfunc]"i"(bpf_kfunc_call_test_ret_pair)
+	: "r0", "r1", "r2", "r3", "r4", "r5"
+	);
+	if (lo != a + b)
+		return 1;
+	if (hi != a - b)
+		return 2;
+
+	/* struct { u64 a; int b; }: R0 = a, low 32 bits of R2 = b. */
+	asm volatile (
+	"r1 = %[a];"
+	"r2 = %[b];"
+	"call %[kfunc];"
+	"%[lo] = r0;"
+	"%[hi] = r2;"
+	: [lo]"=r"(lo), [hi]"=r"(hi)
+	: [a]"r"(a), [b]"r"(b), [kfunc]"i"(bpf_kfunc_call_test_ret_li)
+	: "r0", "r1", "r2", "r3", "r4", "r5"
+	);
+	if (lo != a)
+		return 3;
+	if ((int)hi != ~(int)b)
+		return 4;
+
+	/* struct { int a; int b; }: 8 bytes, packed into R0; R2 is not used. */
+	asm volatile (
+	"r1 = %[a];"
+	"r2 = %[b];"
+	"call %[kfunc];"
+	"%[lo] = r0;"
+	: [lo]"=r"(lo)
+	: [a]"r"(a), [b]"r"(b), [kfunc]"i"(bpf_kfunc_call_test_ret_ii)
+	: "r0", "r1", "r2", "r3", "r4", "r5"
+	);
+	if ((int)lo != (int)a)
+		return 5;
+	if ((int)(lo >> 32) != (int)b)
+		return 6;
+
+	return 0;
+}
+
+SEC("tc")
+int aggregate_ret_union_test(struct __sk_buff *skb)
+{
+	__u64 a = skb->len;
+	__u64 b = skb->len ^ 0xdeadbeefULL;
+	__u64 lo, hi;
+
+	asm volatile (
+	"r1 = %[a];"
+	"r2 = %[b];"
+	"call %[kfunc];"
+	"%[lo] = r0;"
+	"%[hi] = r2;"
+	: [lo]"=r"(lo), [hi]"=r"(hi)
+	: [a]"r"(a), [b]"r"(b), [kfunc]"i"(bpf_kfunc_call_test_ret_uu)
+	: "r0", "r1", "r2", "r3", "r4", "r5"
+	);
+	if (lo != a + b)
+		return 1;
+	if (hi != a - b)
+		return 2;
+
+	return 0;
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/aggregate_ret_target.c b/tools/testing/selftests/bpf/progs/aggregate_ret_target.c
new file mode 100644
index 000000000000..cffd8d7d3241
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/aggregate_ret_target.c
@@ -0,0 +1,29 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+
+/* freplace target: a global subprogram returning 16 bytes in R0:R2. */
+__naked unsigned __int128 agg_ret_target_func(void)
+{
+	asm volatile (
+	"r0 = 0x1234;"
+	"r2 = 0x5678;"
+	"exit;"
+	);
+}
+
+SEC("tc")
+__naked int agg_ret_target(void)
+{
+	asm volatile (
+	"call %[agg_ret_target_func];"
+	"r0 = 0;"
+	"exit;"
+	:
+	: __imm(agg_ret_target_func)
+	: __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/freplace_ret_pair.c b/tools/testing/selftests/bpf/progs/freplace_ret_pair.c
new file mode 100644
index 000000000000..84b701402ca6
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/freplace_ret_pair.c
@@ -0,0 +1,20 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+
+/*
+ * An extension replaces its target outright, so it has to match the target's
+ * return convention. Its own return value is capped at 8 bytes, so it can
+ * never fill the R0:R2 pair that the target's callers read, and the attach is
+ * rejected. btf_check_type_match() cannot catch this: it compares return types
+ * by btf_type->info only, and an int carries no vlen, so the __u64 here and
+ * the target's __int128 compare equal.
+ */
+SEC("freplace/agg_ret_target_func")
+__u64 new_agg_ret_target_func(void)
+{
+	return 0;
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/verifier_arena.c b/tools/testing/selftests/bpf/progs/verifier_arena.c
index b241bbcf54a8..455b55296f35 100644
--- a/tools/testing/selftests/bpf/progs/verifier_arena.c
+++ b/tools/testing/selftests/bpf/progs/verifier_arena.c
@@ -704,4 +704,42 @@ int check_arena_arg_ret(void *ctx)
 	return 0;
 }
 
+struct arena_ret_pair {
+	__u64 lo;
+	__u64 hi;
+};
+
+/*
+ * A 16-byte value is returned in the R0:R2 register pair. A global subprogram
+ * may return an arena pointer in R0, but R2 holds the upper half of a scalar
+ * pair, so an arena pointer there is not a valid return value. The ld_imm64 of
+ * the arena map is what links the arena to the program, without which the
+ * addr_space_cast insn is not allowed.
+ */
+__naked struct arena_ret_pair global_ret_arena_ptr_in_r2(void)
+{
+	asm volatile (
+		"r1 = %[arena] ll;"
+		"r2 = 8192;"
+		"r2 = addr_space_cast(r2, 0x0, 0x1);"
+		"r0 = 0;"
+		"exit;"
+		:
+		: __imm_addr(arena)
+		: __clobber_all);
+}
+
+SEC("syscall")
+__failure __msg("At subprogram exit the register R2 is not a scalar value (arena)")
+__naked int check_global_ret_arena_ptr_in_r2(void)
+{
+	asm volatile (
+		"call %[global_ret_arena_ptr_in_r2];"
+		"r0 = 0;"
+		"exit;"
+		:
+		: __imm(global_ret_arena_ptr_in_r2)
+		: __clobber_all);
+}
+
 char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c b/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
index 2ceb34df472e..5d13482b0d2e 100644
--- a/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
+++ b/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
@@ -869,6 +869,50 @@ __bpf_kfunc struct prog_test_ret_pair bpf_kfunc_call_test_ret_pair(u64 a, u64 b)
 	return r;
 }
 
+__bpf_kfunc struct prog_test_ret_pair bpf_kfunc_call_test_ret_fastcall(u64 a, u64 b)
+{
+	struct prog_test_ret_pair r = { .hi = a + b, .lo = a - b };
+
+	return r;
+}
+
+__bpf_kfunc struct prog_test_ret_li bpf_kfunc_call_test_ret_li(u64 a, int b)
+{
+	struct prog_test_ret_li r = { .a = a, .b = ~b };
+
+	return r;
+}
+
+__bpf_kfunc struct prog_test_ret_ii bpf_kfunc_call_test_ret_ii(int a, int b)
+{
+	struct prog_test_ret_ii r = { .a = a, .b = b };
+
+	return r;
+}
+
+__bpf_kfunc union prog_test_ret_uu bpf_kfunc_call_test_ret_uu(u64 a, u64 b)
+{
+	union prog_test_ret_uu r;
+
+	r.halves[0] = a + b;
+	r.halves[1] = a - b;
+	return r;
+}
+
+__bpf_kfunc struct prog_test_ret_ptr bpf_kfunc_call_test_ret_ptr(u64 tag)
+{
+	struct prog_test_ret_ptr r = { .p = NULL, .tag = tag };
+
+	return r;
+}
+
+__bpf_kfunc struct prog_test_ret_big bpf_kfunc_call_test_ret_big(u64 a)
+{
+	struct prog_test_ret_big r = { .a = a, .b = a, .c = a };
+
+	return r;
+}
+
 __bpf_kfunc u64 bpf_kfunc_call_stack_arg(u64 a, u64 b, u64 c, u64 d,
 					 u64 e, u64 f, u64 g, u64 h,
 					 u64 i, u64 j)
@@ -1404,6 +1448,12 @@ BTF_ID_FLAGS(func, bpf_kfunc_call_test4)
 BTF_ID_FLAGS(func, bpf_kfunc_call_test5)
 BTF_ID_FLAGS(func, bpf_kfunc_call_test_i128)
 BTF_ID_FLAGS(func, bpf_kfunc_call_test_ret_pair)
+BTF_ID_FLAGS(func, bpf_kfunc_call_test_ret_fastcall, KF_FASTCALL)
+BTF_ID_FLAGS(func, bpf_kfunc_call_test_ret_li)
+BTF_ID_FLAGS(func, bpf_kfunc_call_test_ret_ii)
+BTF_ID_FLAGS(func, bpf_kfunc_call_test_ret_uu)
+BTF_ID_FLAGS(func, bpf_kfunc_call_test_ret_ptr)
+BTF_ID_FLAGS(func, bpf_kfunc_call_test_ret_big)
 BTF_ID_FLAGS(func, bpf_kfunc_call_stack_arg)
 BTF_ID_FLAGS(func, bpf_kfunc_call_stack_arg_ptr)
 BTF_ID_FLAGS(func, bpf_kfunc_call_stack_arg_mix)
diff --git a/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h b/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h
index e6e59fdce33c..917e943bd963 100644
--- a/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h
+++ b/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h
@@ -64,6 +64,35 @@ struct prog_test_ret_pair {
 	__u64 lo;
 };
 
+struct prog_test_ret_li {	/* 16 bytes: R0:R2 */
+	__u64 a;
+	int b;
+};
+
+struct prog_test_ret_ii {	/* 8 bytes: R0 only */
+	int a;
+	int b;
+};
+
+union prog_test_ret_uu {	/* 16 bytes: R0:R2 */
+	__u64 halves[2];
+	struct {
+		__u64 lo;
+		__u64 hi;
+	} parts;
+};
+
+struct prog_test_ret_ptr {	/* 16 bytes: contains a pointer */
+	void *p;
+	__u64 tag;
+};
+
+struct prog_test_ret_big {	/* 24 bytes: too large for R0:R2 */
+	__u64 a;
+	__u64 b;
+	__u64 c;
+};
+
 struct prog_test_fail1 {
 	void *p;
 	int x;
@@ -129,6 +158,12 @@ long bpf_kfunc_call_test4(signed char a, short b, int c, long d) __ksym;
 int bpf_kfunc_call_test5(__u8 a, __u16 b, __u32 c) __ksym;
 __int128 bpf_kfunc_call_test_i128(__u64 a, __u64 b) __ksym;
 struct prog_test_ret_pair bpf_kfunc_call_test_ret_pair(__u64 a, __u64 b) __ksym;
+struct prog_test_ret_pair bpf_kfunc_call_test_ret_fastcall(__u64 a, __u64 b) __ksym;
+struct prog_test_ret_li bpf_kfunc_call_test_ret_li(__u64 a, int b) __ksym;
+struct prog_test_ret_ii bpf_kfunc_call_test_ret_ii(int a, int b) __ksym;
+union prog_test_ret_uu bpf_kfunc_call_test_ret_uu(__u64 a, __u64 b) __ksym;
+struct prog_test_ret_ptr bpf_kfunc_call_test_ret_ptr(__u64 tag) __ksym;
+struct prog_test_ret_big bpf_kfunc_call_test_ret_big(__u64 a) __ksym;
 __u64 bpf_kfunc_call_stack_arg(__u64 a, __u64 b, __u64 c, __u64 d,
 			       __u64 e, __u64 f, __u64 g, __u64 h,
 			       __u64 i, __u64 j) __ksym;
-- 
2.53.0-Meta


  parent reply	other threads:[~2026-08-04 20:36 UTC|newest]

Thread overview: 19+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-04 20:35 [PATCH v2 00/13] bpf: Support aggregate return values up to 16 bytes Yonghong Song
2026-08-04 20:35 ` [PATCH bpf-next v2 01/13] bpf: Factor check_global_ret_scalar_reg() out of the global return check Yonghong Song
2026-08-04 20:35 ` [PATCH bpf-next v2 02/13] bpf: Add helpers to describe the R0:R2 return register pair Yonghong Song
2026-08-04 20:35 ` [PATCH bpf-next v2 03/13] bpf: Wire up JIT support for 16-byte kfunc returns Yonghong Song
2026-08-04 20:35 ` [PATCH bpf-next v2 04/13] bpf: Track R2 of register-pair returns in precision backtracking Yonghong Song
2026-08-04 20:35 ` [PATCH bpf-next v2 05/13] bpf: Account R2 of register-pair returns in live register analysis Yonghong Song
2026-08-04 21:14   ` sashiko-bot
2026-08-04 20:35 ` [PATCH bpf-next v2 06/13] bpf: Reject callbacks returning more than 8 bytes Yonghong Song
2026-08-04 21:54   ` bot+bpf-ci
2026-08-04 20:35 ` [PATCH bpf-next v2 07/13] bpf: Add verifier support for 16-byte returns in R0:R2 Yonghong Song
2026-08-04 20:52   ` sashiko-bot
2026-08-04 20:36 ` [PATCH bpf-next v2 08/13] bpf: Reject register-pair returns when the subprog BTF is unreliable Yonghong Song
2026-08-04 20:36 ` [PATCH bpf-next v2 09/13] bpf: Enable aggregate return types up to 16 bytes Yonghong Song
2026-08-04 20:36 ` [PATCH bpf-next v2 10/13] selftests/bpf: Add C tests for 16-byte returns in R0:R2 Yonghong Song
2026-08-04 20:47   ` sashiko-bot
2026-08-04 20:36 ` Yonghong Song [this message]
2026-08-04 20:52   ` [PATCH bpf-next v2 11/13] selftests/bpf: Add inline-asm and subprog tests for R0:R2 returns sashiko-bot
2026-08-04 20:36 ` [PATCH bpf-next v2 12/13] selftests/bpf: Add tests for callbacks returning more than 8 bytes Yonghong Song
2026-08-04 20:36 ` [PATCH bpf-next v2 13/13] Documentation/bpf: Document up to 16-byte kfunc return values in R0:R2 Yonghong Song

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260804203618.1878749-1-yonghong.song@linux.dev \
    --to=yonghong.song@linux.dev \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=kernel-team@fb.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox