All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH bpf-next v5 0/3] bpf: Fix trampoline handling of 128-bit values
@ 2026-07-29  5:01 Yonghong Song
  2026-07-29  5:01 ` [PATCH bpf-next v5 1/3] bpf: Reject >8 byte return values on return-reading trampoline paths Yonghong Song
                   ` (2 more replies)
  0 siblings, 3 replies; 5+ messages in thread
From: Yonghong Song @ 2026-07-29  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team

The BPF trampoline preserves only 8 bytes of a target function's return
value (R0), and its register save area under-allocates space for 128-bit
arguments for x86_64. These two problems lead to memory corruption or
incorrect values observed by BPF programs and the real caller.

This series fixes both issues and adds two selftests, otherwise, each of
them will fail if without the corresponding fix.

Changelogs:
  v4 -> v5:
    - v4: https://lore.kernel.org/bpf/1c4223ae-a5ba-48a4-95d3-57c8ff241055@linux.dev/
    - For function test_fexit_int128_ret(), guard with __x86_64__ and __aarch64__
      to avoid s390x failure
  v3 -> v4:
    - v3: https://lore.kernel.org/bpf/20260710225206.4013062-1-yonghong.song@linux.dev/
    - Add Ack from Leon Hwang
  v2 -> v3:
    - v2: https://lore.kernel.org/bpf/20260710182204.1085329-1-yonghong.song@linux.dev/
    - Align __int128 argument at even position enforced by arm64.
  v1 -> v2:
    - v1: https://lore.kernel.org/bpf/20260710144404.2579671-1-yonghong.song@linux.dev/
    - Also handle __int128 arguments for x86_64.

Yonghong Song (3):
  bpf: Reject >8 byte return values on return-reading trampoline paths
  bpf, x86: Fix trampoline stack size for 128-bit arguments
  selftests/bpf: Add tests for >8 byte return value and 128-bit
    arguments

 arch/x86/net/bpf_jit_comp.c                   |  7 ++--
 kernel/bpf/bpf_struct_ops.c                   | 12 +++++++
 kernel/bpf/verifier.c                         | 25 +++++++++++++
 .../bpf/prog_tests/tracing_failure.c          | 20 +++++++++++
 .../selftests/bpf/prog_tests/tracing_struct.c | 36 +++++++++++++++++++
 .../selftests/bpf/progs/tracing_failure.c     |  6 ++++
 .../bpf/progs/tracing_struct_int128.c         | 18 ++++++++++
 .../selftests/bpf/test_kmods/bpf_testmod.c    | 32 +++++++++++++++++
 8 files changed, 151 insertions(+), 5 deletions(-)
 create mode 100644 tools/testing/selftests/bpf/progs/tracing_struct_int128.c

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH bpf-next v5 1/3] bpf: Reject >8 byte return values on return-reading trampoline paths
  2026-07-29  5:01 [PATCH bpf-next v5 0/3] bpf: Fix trampoline handling of 128-bit values Yonghong Song
@ 2026-07-29  5:01 ` Yonghong Song
  2026-07-29  5:02 ` [PATCH bpf-next v5 2/3] bpf, x86: Fix trampoline stack size for 128-bit arguments Yonghong Song
  2026-07-29  5:02 ` [PATCH bpf-next v5 3/3] selftests/bpf: Add tests for >8 byte return value and " Yonghong Song
  2 siblings, 0 replies; 5+ messages in thread
From: Yonghong Song @ 2026-07-29  5:01 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team, Leon Hwang

btf_distill_func_proto() builds the function model used for the
fentry/fexit/fmod_ret/fsession trampolines and struct_ops. It has
accepted a 16-byte __int128 return value since the trampoline was
introduced: __get_type_size() returns the integer's type size, and the
return-type check only rejected ret < 0.

But the BPF trampoline preserves only 8 bytes of the return value (RAX on
x86, i.e. R0). For an attach type that reads the target's return value the
second half (RDX / R3) is neither saved nor restored, so a program
attached to a function returning a 16-byte value corrupts the value seen
by the real caller and itself observes only half of it. struct_ops
trampolines have the same limitation.

This affects the attach types that read the target's return value: fexit,
fmod_ret and fsession (plus the _multi variants of fexit and fsession),
and struct_ops. fentry/fentry_multi run before the target returns and are
unaffected.

Reject a >8 byte return value for these attach types in
bpf_check_attach_target() and bpf_check_attach_btf_id_multi(), and for
struct_ops in bpf_struct_ops_desc_init().

Fixes: fec56f5890d9 ("bpf: Introduce BPF trampoline")
Acked-by: Leon Hwang <leon.hwang@linux.dev>
Reviewed-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 kernel/bpf/bpf_struct_ops.c | 12 ++++++++++++
 kernel/bpf/verifier.c       | 25 +++++++++++++++++++++++++
 2 files changed, 37 insertions(+)

diff --git a/kernel/bpf/bpf_struct_ops.c b/kernel/bpf/bpf_struct_ops.c
index 51b16e5f5534..4e7a48c02be5 100644
--- a/kernel/bpf/bpf_struct_ops.c
+++ b/kernel/bpf/bpf_struct_ops.c
@@ -445,6 +445,18 @@ int bpf_struct_ops_desc_init(struct bpf_struct_ops_desc *st_ops_desc,
 			goto errout;
 		}
 
+		/*
+		 * A >8 byte return value is passed back in a register pair,
+		 * which the struct_ops trampoline does not preserve (only
+		 * 8 bytes of the return value are saved and restored).
+		 */
+		if (st_ops->func_models[i].ret_size > 8) {
+			pr_warn("func ptr %s in struct %s has a >8 byte return value, which is not supported\n",
+				mname, st_ops->name);
+			err = -EOPNOTSUPP;
+			goto errout;
+		}
+
 		stub_func_addr = *(void **)(st_ops->cfi_stubs + moff);
 		err = prepare_arg_info(btf, st_ops->name, mname,
 				       func_proto, stub_func_addr,
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index e6f35f4e715b..8d0635ee48c7 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -19027,6 +19027,20 @@ btf_attach_func_proto(struct bpf_verifier_log *log, struct btf *btf, u32 func_id
 	return btf_type_by_id(btf, func->type);
 }
 
+static bool attach_uses_trampoline_retval(enum bpf_attach_type type)
+{
+	switch (type) {
+	case BPF_MODIFY_RETURN:
+	case BPF_TRACE_FEXIT:
+	case BPF_TRACE_FEXIT_MULTI:
+	case BPF_TRACE_FSESSION:
+	case BPF_TRACE_FSESSION_MULTI:
+		return true;
+	default:
+		return false;
+	}
+}
+
 int bpf_check_attach_target(struct bpf_verifier_log *log,
 			    const struct bpf_prog *prog,
 			    const struct bpf_prog *tgt_prog,
@@ -19291,6 +19305,14 @@ int bpf_check_attach_target(struct bpf_verifier_log *log,
 		if (ret < 0)
 			return ret;
 
+		if (tgt_info->fmodel.ret_size > 8 &&
+		    attach_uses_trampoline_retval(prog->expected_attach_type)) {
+			bpf_log(log,
+				"Attach to function %s with a >8 byte return value is not supported for this attach type\n",
+				tname);
+			return -EOPNOTSUPP;
+		}
+
 		/*
 		 * *.multi programs don't need an address during program
 		 * verification, we just take the module ref if needed.
@@ -19565,6 +19587,9 @@ int bpf_check_attach_btf_id_multi(struct btf *btf, struct bpf_prog *prog, u32 bt
 	err = btf_distill_func_proto(NULL, btf, t, tname, &tgt_info->fmodel);
 	if (err < 0)
 		return err;
+	if (tgt_info->fmodel.ret_size > 8 &&
+	    attach_uses_trampoline_retval(prog->expected_attach_type))
+		return -EOPNOTSUPP;
 	if (btf_is_module(btf)) {
 		/* The bpf program already holds reference to module. */
 		if (WARN_ON_ONCE(!prog->aux->mod))
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH bpf-next v5 2/3] bpf, x86: Fix trampoline stack size for 128-bit arguments
  2026-07-29  5:01 [PATCH bpf-next v5 0/3] bpf: Fix trampoline handling of 128-bit values Yonghong Song
  2026-07-29  5:01 ` [PATCH bpf-next v5 1/3] bpf: Reject >8 byte return values on return-reading trampoline paths Yonghong Song
@ 2026-07-29  5:02 ` Yonghong Song
  2026-07-29  6:20   ` bot+bpf-ci
  2026-07-29  5:02 ` [PATCH bpf-next v5 3/3] selftests/bpf: Add tests for >8 byte return value and " Yonghong Song
  2 siblings, 1 reply; 5+ messages in thread
From: Yonghong Song @ 2026-07-29  5:02 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team, Leon Hwang

btf_distill_func_proto() accepts a function argument up to 16 bytes, so a
128-bit scalar such as __int128 reaches the x86 trampoline with
arg_size == 16. But the current implementation assumes an __int128
argument only needs one register, so the register save area is
under-allocated and save_args() overwrites adjacent stack slots.

Compute the register count from arg_size for all arguments to fix it.

Fixes: a9c5ad31fbdc ("bpf: x86: Support in-register struct arguments in trampoline programs")
Acked-by: Leon Hwang <leon.hwang@linux.dev>
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 arch/x86/net/bpf_jit_comp.c | 7 ++-----
 1 file changed, 2 insertions(+), 5 deletions(-)

diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index b2feec81e231..01e7ce569c1e 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -3369,11 +3369,8 @@ static int __arch_prepare_bpf_trampoline(struct bpf_tramp_image *im, void *rw_im
 	WARN_ON_ONCE((flags & BPF_TRAMP_F_INDIRECT) &&
 		     (flags & ~(BPF_TRAMP_F_INDIRECT | BPF_TRAMP_F_RET_FENTRY_RET)));
 
-	/* extra registers for struct arguments */
-	for (i = 0; i < m->nr_args; i++) {
-		if (m->arg_flags[i] & BTF_FMODEL_STRUCT_ARG)
-			nr_regs += (m->arg_size[i] + 7) / 8 - 1;
-	}
+	for (i = 0; i < m->nr_args; i++)
+		nr_regs += (m->arg_size[i] + 7) / 8 - 1;
 
 	/* x86-64 supports up to MAX_BPF_FUNC_ARGS arguments. 1-6
 	 * are passed through regs, the remains are through stack.
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH bpf-next v5 3/3] selftests/bpf: Add tests for >8 byte return value and 128-bit arguments
  2026-07-29  5:01 [PATCH bpf-next v5 0/3] bpf: Fix trampoline handling of 128-bit values Yonghong Song
  2026-07-29  5:01 ` [PATCH bpf-next v5 1/3] bpf: Reject >8 byte return values on return-reading trampoline paths Yonghong Song
  2026-07-29  5:02 ` [PATCH bpf-next v5 2/3] bpf, x86: Fix trampoline stack size for 128-bit arguments Yonghong Song
@ 2026-07-29  5:02 ` Yonghong Song
  2 siblings, 0 replies; 5+ messages in thread
From: Yonghong Song @ 2026-07-29  5:02 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, kernel-team, Leon Hwang

The BPF trampoline preserves only 8 bytes of the target's return value
(R0), so attaching an fexit/fmod_ret/fsession program to a function that
returns a >8 byte value is now rejected by the verifier. Add a bpf_testmod
function returning __int128 and an fexit program that targets it. The
program is expected to fail to load with the "with a >8 byte return value
is not supported for this attach type" message.

A 128-bit __int128 argument is passed in a register pair and occupies two
trampoline context slots. Add a bpf_testmod function taking a leading
__int128 argument followed by an int and a long, and an fexit program that
reads those two trailing arguments and the return value, verifying that the
trampoline reserves enough stack for the 128-bit argument and places the
following arguments and the return value at the right context slots.

__int128 is only available on 64-bit targets (where the compiler defines
__SIZEOF_INT128__). The argument test additionally depends on the calling
convention: x86_64 and arm64 pass an __int128 in a register pair as the
trampoline expects, while other architectures pass it differently (e.g.
s390x passes larger arguments by reference), so that subtest runs only on
x86_64 and arm64 and is skipped elsewhere.

Acked-by: Leon Hwang <leon.hwang@linux.dev>
Signed-off-by: Yonghong Song <yonghong.song@linux.dev>
---
 .../bpf/prog_tests/tracing_failure.c          | 20 +++++++++++
 .../selftests/bpf/prog_tests/tracing_struct.c | 36 +++++++++++++++++++
 .../selftests/bpf/progs/tracing_failure.c     |  6 ++++
 .../bpf/progs/tracing_struct_int128.c         | 18 ++++++++++
 .../selftests/bpf/test_kmods/bpf_testmod.c    | 32 +++++++++++++++++
 5 files changed, 112 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/progs/tracing_struct_int128.c

diff --git a/tools/testing/selftests/bpf/prog_tests/tracing_failure.c b/tools/testing/selftests/bpf/prog_tests/tracing_failure.c
index f9f9e1cb87bf..eb585918f0d4 100644
--- a/tools/testing/selftests/bpf/prog_tests/tracing_failure.c
+++ b/tools/testing/selftests/bpf/prog_tests/tracing_failure.c
@@ -76,6 +76,24 @@ static void test_fexit_noreturns(void)
 			       "Attaching fexit/fsession/fmod_ret to __noreturn function 'do_exit' is rejected.");
 }
 
+static void test_fexit_int128_ret(void)
+{
+	/*
+	 * __int128 is returned in a register pair on x86_64 and arm64, so
+	 * bpf_testmod_test_int128_ret() is BTF-encoded and attachable and the
+	 * verifier can reject its >8 byte return value. Other architectures
+	 * return a __int128 differently (e.g. s390x returns larger values by
+	 * reference, which makes pahole skip BTF encoding of the function), so
+	 * only exercise this on x86_64 and arm64.
+	 */
+#if defined(__x86_64__) || defined(__aarch64__)
+	test_tracing_fail_prog("fexit_int128_ret",
+			       "with a >8 byte return value is not supported for this attach type");
+#else
+	test__skip();
+#endif
+}
+
 void test_tracing_failure(void)
 {
 	if (test__start_subtest("bpf_spin_lock"))
@@ -86,4 +104,6 @@ void test_tracing_failure(void)
 		test_tracing_deny();
 	if (test__start_subtest("fexit_noreturns"))
 		test_fexit_noreturns();
+	if (test__start_subtest("fexit_int128_ret"))
+		test_fexit_int128_ret();
 }
diff --git a/tools/testing/selftests/bpf/prog_tests/tracing_struct.c b/tools/testing/selftests/bpf/prog_tests/tracing_struct.c
index 6f8c0bfb0415..15b95d0235b5 100644
--- a/tools/testing/selftests/bpf/prog_tests/tracing_struct.c
+++ b/tools/testing/selftests/bpf/prog_tests/tracing_struct.c
@@ -4,6 +4,7 @@
 #include <test_progs.h>
 #include "tracing_struct.skel.h"
 #include "tracing_struct_many_args.skel.h"
+#include "tracing_struct_int128.skel.h"
 
 static void test_struct_args(void)
 {
@@ -112,6 +113,39 @@ static void test_struct_many_args(void)
 	tracing_struct_many_args__destroy(skel);
 }
 
+static void test_int128_args(void)
+{
+	/*
+	 * __int128 arguments are passed in a register pair on x86_64 and
+	 * arm64, which the trampoline packs into two context slots. Other
+	 * architectures pass a __int128 differently (e.g. s390x passes larger
+	 * arguments by reference), so only exercise this on x86_64 and arm64.
+	 */
+#if defined(__x86_64__) || defined(__aarch64__)
+	struct tracing_struct_int128 *skel;
+	int err;
+
+	skel = tracing_struct_int128__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "tracing_struct_int128__open_and_load"))
+		return;
+
+	err = tracing_struct_int128__attach(skel);
+	if (!ASSERT_OK(err, "tracing_struct_int128__attach"))
+		goto destroy_skel;
+
+	ASSERT_OK(trigger_module_test_read(256), "trigger_read");
+
+	ASSERT_EQ(skel->bss->t_b, 2, "t:b");
+	ASSERT_EQ(skel->bss->t_c, 3, "t:c");
+	ASSERT_EQ(skel->bss->t_ret, 6, "t ret");
+
+destroy_skel:
+	tracing_struct_int128__destroy(skel);
+#else
+	test__skip();
+#endif
+}
+
 static void test_union_args(void)
 {
 	struct tracing_struct *skel;
@@ -145,6 +179,8 @@ void test_tracing_struct(void)
 		test_struct_args();
 	if (test__start_subtest("struct_many_args"))
 		test_struct_many_args();
+	if (test__start_subtest("int128_args"))
+		test_int128_args();
 	if (test__start_subtest("union_args"))
 		test_union_args();
 }
diff --git a/tools/testing/selftests/bpf/progs/tracing_failure.c b/tools/testing/selftests/bpf/progs/tracing_failure.c
index 65e485c4468c..f7a095767679 100644
--- a/tools/testing/selftests/bpf/progs/tracing_failure.c
+++ b/tools/testing/selftests/bpf/progs/tracing_failure.c
@@ -30,3 +30,9 @@ int BPF_PROG(fexit_noreturns)
 {
 	return 0;
 }
+
+SEC("?fexit/bpf_testmod_test_int128_ret")
+int BPF_PROG(fexit_int128_ret)
+{
+	return 0;
+}
diff --git a/tools/testing/selftests/bpf/progs/tracing_struct_int128.c b/tools/testing/selftests/bpf/progs/tracing_struct_int128.c
new file mode 100644
index 000000000000..4638dfec1f38
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/tracing_struct_int128.c
@@ -0,0 +1,18 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_tracing.h>
+#include <bpf/bpf_helpers.h>
+
+long t_b, t_c, t_ret;
+
+SEC("fexit/bpf_testmod_test_int128_arg")
+int test_int128_arg_fexit(unsigned long long *ctx)
+{
+	t_b = (int)ctx[2];
+	t_c = (long)ctx[3];
+	t_ret = (long)ctx[4];
+	return 0;
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c b/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
index 30f1cd23093c..eb0f9b5e18d8 100644
--- a/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
+++ b/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
@@ -161,6 +161,33 @@ bpf_testmod_test_arg_ptr_to_struct(struct bpf_testmod_struct_arg_1 *a) {
 	return bpf_testmod_test_struct_arg_result;
 }
 
+#ifdef __SIZEOF_INT128__
+noinline __int128
+bpf_testmod_test_int128_ret(int a)
+{
+	bpf_testmod_test_struct_arg_result = a;
+	return (__int128)a;
+}
+
+/*
+ * The __int128 'a' is the first argument on purpose. On arm64 a 16-byte
+ * argument must start in an even-numbered register pair, so placing it
+ * after a single-register scalar would leave a padding register (x1)
+ * unused. pahole maps parameters to registers positionally and would then
+ * see the following argument in an "unexpected" register and skip BTF
+ * encoding of the whole function, making it unattachable. Keeping the
+ * __int128 first (x0:x1) avoids the padding while still exercising the
+ * trampoline packing of a 128-bit argument together with the trailing
+ * int and long arguments.
+ */
+noinline long
+bpf_testmod_test_int128_arg(__int128 a, int b, long c)
+{
+	bpf_testmod_test_struct_arg_result = (long)a + b + c;
+	return bpf_testmod_test_struct_arg_result;
+}
+#endif
+
 __weak noinline void bpf_testmod_looooooooooooooooooooooooooooooong_name(void)
 {
 }
@@ -514,6 +541,11 @@ bpf_testmod_test_read(struct file *file, struct kobject *kobj,
 
 	(void)bpf_testmod_test_arg_ptr_to_struct(&struct_arg1_2);
 
+#ifdef __SIZEOF_INT128__
+	(void)bpf_testmod_test_int128_ret(i);
+	(void)bpf_testmod_test_int128_arg((__int128)1, 2, 3);
+#endif
+
 	(void)trace_bpf_testmod_test_raw_tp_null_tp(NULL);
 
 	bpf_testmod_test_struct_ops3();
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH bpf-next v5 2/3] bpf, x86: Fix trampoline stack size for 128-bit arguments
  2026-07-29  5:02 ` [PATCH bpf-next v5 2/3] bpf, x86: Fix trampoline stack size for 128-bit arguments Yonghong Song
@ 2026-07-29  6:20   ` bot+bpf-ci
  0 siblings, 0 replies; 5+ messages in thread
From: bot+bpf-ci @ 2026-07-29  6:20 UTC (permalink / raw)
  To: yonghong.song, bpf
  Cc: ast, andrii, daniel, eddyz87, kernel-team, leon.hwang, ast,
	andrii, daniel, martin.lau, eddyz87, yonghong.song, clm,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 3749 bytes --]

> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
> index b2feec81e231..01e7ce569c1e 100644
> --- a/arch/x86/net/bpf_jit_comp.c
> +++ b/arch/x86/net/bpf_jit_comp.c
> @@ -3369,11 +3369,8 @@ static int __arch_prepare_bpf_trampoline(struct bpf_tramp_image *im, void *rw_im
>  	WARN_ON_ONCE((flags & BPF_TRAMP_F_INDIRECT) &&
>  		     (flags & ~(BPF_TRAMP_F_INDIRECT | BPF_TRAMP_F_RET_FENTRY_RET)));
>
> -	/* extra registers for struct arguments */
> -	for (i = 0; i < m->nr_args; i++) {
> -		if (m->arg_flags[i] & BTF_FMODEL_STRUCT_ARG)
> -			nr_regs += (m->arg_size[i] + 7) / 8 - 1;
> -	}
> +	for (i = 0; i < m->nr_args; i++)
> +		nr_regs += (m->arg_size[i] + 7) / 8 - 1;
>
>  	/* x86-64 supports up to MAX_BPF_FUNC_ARGS arguments. 1-6
>  	 * are passed through regs, the remains are through stack.

Does the Fixes: tag target the commit that introduced the bug described in
the commit message?

The changelog says save_args() overwrites adjacent stack slots because the
register save area is under-allocated. That mismatch requires the writer of
the save area to count 2 slots for a 16-byte argument while the accounting
loop counts 1.

In commit a9c5ad31fbdc (v6.1-era), both sides were symmetric and save_args()
did not exist yet. The save_regs() function in a9c5ad31fbdc used:

arch/x86/net/bpf_jit_comp.c, save_regs():
  if (m->arg_flags[i] & BTF_FMODEL_STRUCT_ARG) {
          nr_regs = (m->arg_size[i] + 7) / 8;
          arg_size = 8;
  } else {
          nr_regs = 1;
          arg_size = m->arg_size[i];
  }

arch_prepare_bpf_trampoline():
  if (m->arg_flags[i] & BTF_FMODEL_STRUCT_ARG)
          extra_nregs += (m->arg_size[i] + 7) / 8 - 1;

Both used the same BTF_FMODEL_STRUCT_ARG condition, so a non-struct 16-byte
__int128 was accounted for 1 slot and written to 1 slot. No adjacent slot
was overwritten.

The asymmetry the changelog describes was introduced later, when save_regs()
was replaced by save_args()/restore_regs() that compute the slot count
unconditionally:

arch/x86/net/bpf_jit_comp.c (current):
  save_args():
      arg_regs = (m->arg_size[i] + 7) / 8;
  restore_regs():
      arg_regs = (m->arg_size[i] + 7) / 8;
  get_nr_used_regs():
      arg_regs = (m->arg_size[i] + 7) / 8;

versus the STRUCT_ARG-conditional accounting loop this patch deletes. The
commit that made save_args() unconditional is the one that created the
mismatch ('bpf, x86: allow function arguments up to 12 for TRACING', Menglong
Dong, v6.6-era).

The Fixes: tag points at a v6.1-era commit while the real culprit is
v6.6-era. On a v6.1-v6.5 tree this patch changes nr_regs for an __int128
from 1 to 2 while save_regs() still writes only 1 slot, which over-allocates
the save area rather than fixing anything.

---

Kumar Kartikeya Dwivedi raised a concern in v3 about a pahole BTF encoding
issue with the function bpf_testmod_test_int128_arg(__int128 a, int b, long c)
in tools/testing/selftests/bpf/test_kmods/bpf_testmod.c.

On x86-64, parameter 'a' consumes rdi:rsi, and 'b' ends up in rdx (not the
positionally-expected rsi). pahole maps the second parameter positionally to
rsi, sees 'b' in rdx, and skips BTF encoding for the function.

The comment in the code claiming that putting __int128 first prevents the
pahole issue is incorrect. Kumar suggested using bpf_testmod_test_int128_arg
(int b, long c, __int128 a) instead.

Is the pahole patch set 'Encode true signatures in kernel BTF' that you
mentioned in your response now merged? Without it, will this test fail in CI?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/30424886481

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-07-29  6:20 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-29  5:01 [PATCH bpf-next v5 0/3] bpf: Fix trampoline handling of 128-bit values Yonghong Song
2026-07-29  5:01 ` [PATCH bpf-next v5 1/3] bpf: Reject >8 byte return values on return-reading trampoline paths Yonghong Song
2026-07-29  5:02 ` [PATCH bpf-next v5 2/3] bpf, x86: Fix trampoline stack size for 128-bit arguments Yonghong Song
2026-07-29  6:20   ` bot+bpf-ci
2026-07-29  5:02 ` [PATCH bpf-next v5 3/3] selftests/bpf: Add tests for >8 byte return value and " Yonghong Song

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.