* [PATCH bpf-next 01/17] bpf: Fix infinite loop in check_max_stack_depth()
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-22 2:01 ` bot+bpf-ci
2026-09-23 22:35 ` Eduard Zingerman
2026-09-22 1:13 ` [PATCH bpf-next 02/17] selftests/bpf: Test recursion through a global function and a callback Alexei Starovoitov
` (15 subsequent siblings)
16 siblings, 2 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
sort_subprogs_topo() rejects recursion via direct calls, but allows cycles
that go through ld_imm64 BPF_PSEUDO_FUNC, since the address could be taken
for an async callback. The main verifier pass limits the number of frames,
but it doesn't follow calls into global functions.
So the following program passes both:
static int cb(u32 i, void *ctx) { return foo(i); }
static int foo(int x) { return bar(x); }
int bar(int x) { bpf_loop(1, cb, NULL, 0); return 0; }
SEC("socket") int prog(void *ctx) { return cb(0, NULL); }
cb -> foo -> bar -> bpf_loop -> cb is unbounded recursion.
check_max_stack_depth_subprog() walks this cycle too. None of
the functions use stack and the frame counter is reset by every global
function, so neither stack size nor frame limit is hit and bpf_prog_load()
spins forever.
Reject a call into a subprog that is already in the chain of callers.
Async callbacks are skipped earlier and are not affected.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
kernel/bpf/verifier.c | 16 ++++++++++++++++
1 file changed, 16 insertions(+)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index d62c0f74cff5..72ea662baea5 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -5505,6 +5505,22 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
}
}
+ /*
+ * sort_subprogs_topo() tolerates cycles in the call graph that
+ * go through the address of a function being taken, since it
+ * doesn't know what it is taken for. Such cycle is a recursion
+ * unless it's an async callback, which are skipped above.
+ * The main verification pass limits the depth of the recursion,
+ * but it doesn't follow calls of global functions.
+ */
+ for (tmp = idx; tmp >= 0; tmp = dinfo[tmp].caller) {
+ if (tmp != sidx)
+ continue;
+ verbose(env, "recursive call from %s() to %s()\n",
+ bpf_subprog_name(env, idx), bpf_subprog_name(env, sidx));
+ return -EINVAL;
+ }
+
/* store caller info for after we return from callee */
dinfo[idx].frame = frame;
dinfo[idx].ret_insn = i + 1;
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 01/17] bpf: Fix infinite loop in check_max_stack_depth()
2026-09-22 1:13 ` [PATCH bpf-next 01/17] bpf: Fix infinite loop in check_max_stack_depth() Alexei Starovoitov
@ 2026-09-22 2:01 ` bot+bpf-ci
2026-09-22 2:56 ` Alexei Starovoitov
2026-09-23 22:35 ` Eduard Zingerman
1 sibling, 1 reply; 44+ messages in thread
From: bot+bpf-ci @ 2026-09-22 2:01 UTC (permalink / raw)
To: alexei.starovoitov, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, ast, andrii,
daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3389 bytes --]
> commit c85fc3663258fcad6e896f09408e2d78dc70eb8b
> Author: Alexei Starovoitov <ast@kernel.org>
> Date: Tue Sep 22 01:13:07 2026 +0000
>
> bpf: Fix infinite loop in check_max_stack_depth()
>
> sort_subprogs_topo() rejects recursion via direct calls, but allows cycles
> that go through ld_imm64 BPF_PSEUDO_FUNC, since the address could be taken
> for an async callback. The main verifier pass limits the number of frames,
> but it doesn't follow calls into global functions.
> So the following program passes both:
>
> static int cb(u32 i, void *ctx) { return foo(i); }
> static int foo(int x) { return bar(x); }
> int bar(int x) { bpf_loop(1, cb, NULL, 0); return 0; }
>
> SEC("socket") int prog(void *ctx) { return cb(0, NULL); }
>
> cb -> foo -> bar -> bpf_loop -> cb is unbounded recursion.
> check_max_stack_depth_subprog() walks this cycle too. None of
> the functions use stack and the frame counter is reset by every global
> function, so neither stack size nor frame limit is hit and bpf_prog_load()
> spins forever.
>
> Reject a call into a subprog that is already in the chain of callers.
> Async callbacks are skipped earlier and are not affected.
>
> Signed-off-by: Alexei Starovoitov <ast@kernel.org>
This looks like a bug fix for async callback support. Should this include:
Fixes: 7ddc80a476c2 ("bpf: Teach stack depth check about async callbacks.")
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index d62c0f74cff5..72ea662baea5 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
> @@ -5505,6 +5505,22 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
> }
> }
>
> + /*
> + * sort_subprogs_topo() tolerates cycles in the call graph that
> + * go through the address of a function being taken, since it
> + * doesn't know what it is taken for. Such cycle is a recursion
> + * unless it's an async callback, which are skipped above.
> + * The main verification pass limits the depth of the recursion,
> + * but it doesn't follow calls of global functions.
> + */
> + for (tmp = idx; tmp >= 0; tmp = dinfo[tmp].caller) {
> + if (tmp != sidx)
> + continue;
> + verbose(env, "recursive call from %s() to %s()\n",
> + bpf_subprog_name(env, idx), bpf_subprog_name(env, sidx));
> + return -EINVAL;
> + }
> +
This isn't a bug, but would it be worth attaching a bpf_diag_program_structure()
hint here so both recursion rejections look the same to users?
sort_subprogs_topo() emits the same "recursive call from %s() to %s()" message
paired with a bpf_diag_program_structure() call that gives users a title, a
hint like "Rewrite the recursion as an explicit bounded loop...", and an
explanation. This new site emits only the bare verbose() line. bpf_diag is
still usable at this point in bpf_check() - check_max_stack_depth() runs
before the "instruction rewrites happen after this point" boundary.
> /* store caller info for after we return from callee */
> dinfo[idx].frame = frame;
> dinfo[idx].ret_insn = i + 1;
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35675730306
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 01/17] bpf: Fix infinite loop in check_max_stack_depth()
2026-09-22 2:01 ` bot+bpf-ci
@ 2026-09-22 2:56 ` Alexei Starovoitov
0 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 2:56 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, martin.lau,
yonghong.song, mason, ihor.solodrai
On Tue, Sep 22, 2026 at 02:01 AM bot+bpf-ci@kernel.org <bot+bpf-ci@kernel.org> wrote:
> This looks like a bug fix for async callback support. Should this include:
>
> Fixes: 7ddc80a476c2 ("bpf: Teach stack depth check about async callbacks.")
Nope.
No fixes tags was deliberate.
^ permalink raw reply [flat|nested] 44+ messages in thread
* Re: [PATCH bpf-next 01/17] bpf: Fix infinite loop in check_max_stack_depth()
2026-09-22 1:13 ` [PATCH bpf-next 01/17] bpf: Fix infinite loop in check_max_stack_depth() Alexei Starovoitov
2026-09-22 2:01 ` bot+bpf-ci
@ 2026-09-23 22:35 ` Eduard Zingerman
1 sibling, 0 replies; 44+ messages in thread
From: Eduard Zingerman @ 2026-09-23 22:35 UTC (permalink / raw)
To: Alexei Starovoitov, bpf; +Cc: daniel, andrii, memxor, a.s.protopopov
On Tue, 2026-09-22 at 01:13 +0000, Alexei Starovoitov wrote:
> From: Alexei Starovoitov <ast@kernel.org>
>
> sort_subprogs_topo() rejects recursion via direct calls, but allows cycles
> that go through ld_imm64 BPF_PSEUDO_FUNC, since the address could be taken
> for an async callback. The main verifier pass limits the number of frames,
> but it doesn't follow calls into global functions.
> So the following program passes both:
>
> static int cb(u32 i, void *ctx) { return foo(i); }
> static int foo(int x) { return bar(x); }
> int bar(int x) { bpf_loop(1, cb, NULL, 0); return 0; }
>
> SEC("socket") int prog(void *ctx) { return cb(0, NULL); }
>
> cb -> foo -> bar -> bpf_loop -> cb is unbounded recursion.
> check_max_stack_depth_subprog() walks this cycle too. None of
> the functions use stack and the frame counter is reset by every global
> function, so neither stack size nor frame limit is hit and bpf_prog_load()
> spins forever.
>
> Reject a call into a subprog that is already in the chain of callers.
> Async callbacks are skipped earlier and are not affected.
>
> Signed-off-by: Alexei Starovoitov <ast@kernel.org>
> ---
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
...
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH bpf-next 02/17] selftests/bpf: Test recursion through a global function and a callback
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
2026-09-22 1:13 ` [PATCH bpf-next 01/17] bpf: Fix infinite loop in check_max_stack_depth() Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-22 1:13 ` [PATCH bpf-next 03/17] bpf: Don't fold loads from insn_array maps into constants Alexei Starovoitov
` (14 subsequent siblings)
16 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
Add a test for recursion
callback -> static func -> global func -> bpf_loop() -> callback
that used to hang check_max_stack_depth().
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
.../bpf/progs/verifier_global_subprogs.c | 31 +++++++++++++++++++
1 file changed, 31 insertions(+)
diff --git a/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c b/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c
index 966f49348787..27fbe54e8795 100644
--- a/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c
+++ b/tools/testing/selftests/bpf/progs/verifier_global_subprogs.c
@@ -536,4 +536,35 @@ int return_from_void_global(struct __sk_buff *skb)
return 0;
}
+int global_calls_loop(int x);
+
+static __noinline int static_calls_global(int x)
+{
+ return global_calls_loop(x);
+}
+
+static __noinline int loop_cb_calls_static(u32 i, void *ctx)
+{
+ return static_calls_global(i);
+}
+
+__noinline int global_calls_loop(int x)
+{
+ bpf_loop(1, loop_cb_calls_static, NULL, 0);
+ return 0;
+}
+
+/*
+ * loop_cb_calls_static() -> static_calls_global() -> global_calls_loop() ->
+ * bpf_loop() -> loop_cb_calls_static() is an unbounded recursion that the
+ * main verification pass can't see, because it doesn't follow calls of global
+ * functions. None of the functions use stack.
+ */
+SEC("?raw_tp")
+__failure __msg("recursive call from global_calls_loop() to loop_cb_calls_static()")
+int recursion_via_global_func_and_callback(const void *ctx)
+{
+ return loop_cb_calls_static(0, NULL);
+}
+
char _license[] SEC("license") = "GPL";
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* [PATCH bpf-next 03/17] bpf: Don't fold loads from insn_array maps into constants
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
2026-09-22 1:13 ` [PATCH bpf-next 01/17] bpf: Fix infinite loop in check_max_stack_depth() Alexei Starovoitov
2026-09-22 1:13 ` [PATCH bpf-next 02/17] selftests/bpf: Test recursion through a global function and a callback Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-23 22:39 ` Eduard Zingerman
2026-09-24 0:19 ` bot+bpf-ci
2026-09-22 1:13 ` [PATCH bpf-next 04/17] bpf: Prepare static analysis passes for callx instruction Alexei Starovoitov
` (13 subsequent siblings)
16 siblings, 2 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
bpf_compute_const_regs() constant-folds loads from frozen read-only maps
via bpf_map_direct_read(). insn_array maps are frozen and read-only too,
but their values are addresses of jitted insns that are not known during
verification. bpf_map_direct_read() returns zero, the register loaded
from the first two elements of a jump table becomes const 0 and
bpf_prune_dead_branches() may rewrite a conditional jump on it.
check_map_mem_read() in the main pass marks such register as PTR_TO_INSN.
Treat it as unknown in const_fold as well.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
kernel/bpf/const_fold.c | 5 +++++
1 file changed, 5 insertions(+)
diff --git a/kernel/bpf/const_fold.c b/kernel/bpf/const_fold.c
index 7f1b30059cc8..982125026eaf 100644
--- a/kernel/bpf/const_fold.c
+++ b/kernel/bpf/const_fold.c
@@ -181,7 +181,12 @@ static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *
int off = src->val + insn->off;
u64 val = 0;
+ /*
+ * Values of insn_array map are addresses of jitted instructions,
+ * which are not known until the program is jitted.
+ */
if (!bpf_map_is_rdonly(map) || !map->ops->map_direct_value_addr ||
+ map->map_type == BPF_MAP_TYPE_INSN_ARRAY ||
off < 0 || off + size > map->value_size ||
bpf_map_direct_read(map, off, size, &val, is_ldsx)) {
*dst = unknown;
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 03/17] bpf: Don't fold loads from insn_array maps into constants
2026-09-22 1:13 ` [PATCH bpf-next 03/17] bpf: Don't fold loads from insn_array maps into constants Alexei Starovoitov
@ 2026-09-23 22:39 ` Eduard Zingerman
2026-09-23 23:12 ` Alexei Starovoitov
2026-09-24 0:19 ` bot+bpf-ci
1 sibling, 1 reply; 44+ messages in thread
From: Eduard Zingerman @ 2026-09-23 22:39 UTC (permalink / raw)
To: Alexei Starovoitov, bpf; +Cc: daniel, andrii, memxor, a.s.protopopov
On Tue, 2026-09-22 at 01:13 +0000, Alexei Starovoitov wrote:
> From: Alexei Starovoitov <ast@kernel.org>
>
> bpf_compute_const_regs() constant-folds loads from frozen read-only maps
> via bpf_map_direct_read(). insn_array maps are frozen and read-only too,
> but their values are addresses of jitted insns that are not known during
> verification. bpf_map_direct_read() returns zero, the register loaded
> from the first two elements of a jump table becomes const 0 and
> bpf_prune_dead_branches() may rewrite a conditional jump on it.
>
> check_map_mem_read() in the main pass marks such register as PTR_TO_INSN.
> Treat it as unknown in const_fold as well.
>
> Signed-off-by: Alexei Starovoitov <ast@kernel.org>
> ---
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
This one does deserve fixes tag, though.
...
^ permalink raw reply [flat|nested] 44+ messages in thread
* Re: [PATCH bpf-next 03/17] bpf: Don't fold loads from insn_array maps into constants
2026-09-23 22:39 ` Eduard Zingerman
@ 2026-09-23 23:12 ` Alexei Starovoitov
0 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-23 23:12 UTC (permalink / raw)
To: Eduard Zingerman, bpf; +Cc: daniel, andrii, memxor, a.s.protopopov
On Wed, Sep 23, 2026 at 03:39 PM Eduard Zingerman <eddyz87@gmail.com> wrote:
>> verification. bpf_map_direct_read() returns zero, the register loaded
>> from the first two elements of a jump table becomes const 0 and
>> bpf_prune_dead_branches() may rewrite a conditional jump on it.
[...]
> Acked-by: Eduard Zingerman <eddyz87@gmail.com>
>
> This one does deserve fixes tag, though.
Sorry that was confusion.
This check is already part of bpf_map_direct_read().
Will drop this patch.
^ permalink raw reply [flat|nested] 44+ messages in thread
* Re: [PATCH bpf-next 03/17] bpf: Don't fold loads from insn_array maps into constants
2026-09-22 1:13 ` [PATCH bpf-next 03/17] bpf: Don't fold loads from insn_array maps into constants Alexei Starovoitov
2026-09-23 22:39 ` Eduard Zingerman
@ 2026-09-24 0:19 ` bot+bpf-ci
1 sibling, 0 replies; 44+ messages in thread
From: bot+bpf-ci @ 2026-09-24 0:19 UTC (permalink / raw)
To: alexei.starovoitov, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, ast, andrii,
daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3372 bytes --]
> bpf: Don't fold loads from insn_array maps into constants
>
> bpf_compute_const_regs() constant-folds loads from frozen read-only maps
> via bpf_map_direct_read(). insn_array maps are frozen and read-only too,
> but their values are addresses of jitted insns that are not known during
> verification. bpf_map_direct_read() returns zero, the register loaded
> from the first two elements of a jump table becomes const 0 and
> bpf_prune_dead_branches() may rewrite a conditional jump on it.
> check_map_mem_read() in the main pass marks such register as PTR_TO_INSN.
> Treat it as unknown in const_fold as well.
> Signed-off-by: Alexei Starovoitov <ast@kernel.org>
> Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Does the commit message accurately describe the bug being fixed?
The message says bpf_map_direct_read() returns zero for insn_array maps,
causing const_fold to fold the load into const 0. However,
bpf_map_direct_read() in kernel/bpf/verifier.c already rejects this map
type before calling map_direct_value_addr:
int bpf_map_direct_read(struct bpf_map *map, int off, int size, u64 *val,
bool is_ldsx)
{
...
if (map->map_type == BPF_MAP_TYPE_INSN_ARRAY ||
map->map_type == BPF_MAP_TYPE_PERCPU_ARRAY)
return -EINVAL;
...
}
In const_reg_xfer() the BPF_LDX case already ends with:
if (!bpf_map_is_rdonly(map) || !map->ops->map_direct_value_addr ||
off < 0 || off + size > map->value_size ||
bpf_map_direct_read(map, off, size, &val, is_ldsx)) {
*dst = unknown;
break;
}
Before this patch, a load from an insn_array map already produced
*dst = unknown through the non-zero return of bpf_map_direct_read().
The new check only short-circuits earlier into the same outcome.
The register was never const 0, and bpf_prune_dead_branches() could not
rewrite a jump based on it. Is this patch intended as a bug fix, or as a
clarity/defence-in-depth change?
Additionally, Eduard Zingerman noted in the v1 review that this patch
"does deserve fixes tag, though." You responded that "This check is already
part of bpf_map_direct_read(). Will drop this patch." However, the patch
was not dropped and no Fixes tag was added. Should this patch be dropped as
originally planned, or if it's being kept for other reasons, should the
commit message be reworded to reflect that this is a clarity change rather
than a bug fix?
> diff --git a/kernel/bpf/const_fold.c b/kernel/bpf/const_fold.c
> index 7f1b30059cc8..982125026eaf 100644
> --- a/kernel/bpf/const_fold.c
> +++ b/kernel/bpf/const_fold.c
> @@ -181,7 +181,12 @@ static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *
> int off = src->val + insn->off;
> u64 val = 0;
>
> + /*
> + * Values of insn_array map are addresses of jitted instructions,
> + * which are not known until the program is jitted.
> + */
> if (!bpf_map_is_rdonly(map) || !map->ops->map_direct_value_addr ||
> + map->map_type == BPF_MAP_TYPE_INSN_ARRAY ||
> off < 0 || off + size > map->value_size ||
> bpf_map_direct_read(map, off, size, &val, is_ldsx)) {
> *dst = unknown;
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35935998342
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH bpf-next 04/17] bpf: Prepare static analysis passes for callx instruction
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (2 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 03/17] bpf: Don't fold loads from insn_array maps into constants Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-23 23:10 ` Eduard Zingerman
2026-09-22 1:13 ` [PATCH bpf-next 05/17] bpf: Add callx instruction to call bpf subprogs indirectly Alexei Starovoitov
` (12 subsequent siblings)
16 siblings, 1 reply; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
BPF_JMP | BPF_CALL | BPF_X (callx) will be an indirect call of static
subprog with address in dst_reg. The passes that run before the main
verifier pass don't know which subprog callx calls.
Teach them to treat callx as a call with unknown callee:
- const_fold: callx clobbers R0-R5 like any other call. Otherwise
bpf_prune_dead_branches() could rewrite a live conditional jump.
- live regs: callx uses R1-R5 and dst_reg, defines R0-R5.
- stack liveness: func instances are keyed by (callsite, depth) and
cannot describe a callsite with multiple callees. Don't create
instances for callees of callx. If any callx argument is derived
from fp mark stack of all frames as read at callx insn and keep slots
of outer frames alive 'before' the callsite while the callee is
verified. Same as for callbacks that are not known statically.
The callee is analyzed as standalone instance.
- backtracking: treat callx as a call of static subprog, same as
BPF_PSEUDO_CALL.
callx is still rejected as unknown opcode. No functional change.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
include/linux/bpf_verifier.h | 6 ++++++
kernel/bpf/backtrack.c | 20 ++++++++++++--------
kernel/bpf/const_fold.c | 3 ++-
kernel/bpf/liveness.c | 35 ++++++++++++++++++++++++++++++-----
4 files changed, 50 insertions(+), 14 deletions(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 92f528c45605..6380c851ed24 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1089,6 +1089,12 @@ static inline bool bpf_pseudo_kfunc_call(const struct bpf_insn *insn)
insn->src_reg == BPF_PSEUDO_KFUNC_CALL;
}
+/* callx: indirect call of a bpf subprog whose address is in insn->dst_reg */
+static inline bool bpf_is_callx(const struct bpf_insn *insn)
+{
+ return insn->code == (BPF_JMP | BPF_CALL | BPF_X);
+}
+
__printf(2, 0) void bpf_verifier_vlog(struct bpf_verifier_log *log,
const char *fmt, va_list args);
__printf(2, 3) void bpf_verifier_log_write(struct bpf_verifier_env *env,
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 507a366dffa4..4da99dec0818 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -406,15 +406,18 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
if (class == BPF_STX)
bt_set_reg(bt, sreg);
} else if (class == BPF_JMP || class == BPF_JMP32) {
- if (bpf_pseudo_call(insn)) {
- int subprog_insn_idx, subprog;
+ if (bpf_pseudo_call(insn) || bpf_is_callx(insn)) {
+ int subprog_insn_idx, subprog = -1;
- subprog_insn_idx = idx + insn->imm + 1;
- subprog = bpf_find_subprog(env, subprog_insn_idx);
- if (subprog < 0)
- return -EFAULT;
+ if (bpf_pseudo_call(insn)) {
+ subprog_insn_idx = idx + insn->imm + 1;
+ subprog = bpf_find_subprog(env, subprog_insn_idx);
+ if (subprog < 0)
+ return -EFAULT;
+ }
- if (bpf_subprog_is_global(env, subprog)) {
+ /* callx calls static subprogs only */
+ if (subprog >= 0 && bpf_subprog_is_global(env, subprog)) {
/* check that jump history doesn't have any
* extra instructions from subprog; the next
* instruction after call to global subprog
@@ -536,7 +539,8 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
* never do that.
*/
from_subprog_call = subseq_idx - 1 >= 0 &&
- bpf_pseudo_call(&env->prog->insnsi[subseq_idx - 1]);
+ (bpf_pseudo_call(&env->prog->insnsi[subseq_idx - 1]) ||
+ bpf_is_callx(&env->prog->insnsi[subseq_idx - 1]));
r0_precise = from_subprog_call && bt_is_reg_set(bt, BPF_REG_0);
r2_precise = from_subprog_call && bt_is_reg_set(bt, BPF_REG_2);
diff --git a/kernel/bpf/const_fold.c b/kernel/bpf/const_fold.c
index 982125026eaf..f44ae8487ec6 100644
--- a/kernel/bpf/const_fold.c
+++ b/kernel/bpf/const_fold.c
@@ -196,7 +196,8 @@ static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *
dst->val = val;
break;
case BPF_JMP:
- if (opcode != BPF_CALL)
+ /* both 'call imm' and 'callx reg' clobber caller saved registers */
+ if (BPF_OP(insn->code) != BPF_CALL)
break;
process_call:
for (r = BPF_REG_0; r <= BPF_REG_5; r++)
diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
index 44ecdc5b4ec2..5aa2f68d92b3 100644
--- a/kernel/bpf/liveness.c
+++ b/kernel/bpf/liveness.c
@@ -356,12 +356,25 @@ int bpf_live_stack_query_init(struct bpf_verifier_env *env, struct bpf_verifier_
return 0;
}
+/*
+ * Stack accesses of callbacks and of callx callees are not tracked by
+ * func instances keyed by the @callsite. Callbacks might be called several
+ * times and the callee of callx is not known when stack liveness is computed.
+ * In both cases stack slots of the outer frames that might be read by the
+ * callee are accounted as read by the @callsite instruction itself.
+ */
+static bool callee_stack_access_at_callsite(struct bpf_verifier_env *env, u32 callsite)
+{
+ return bpf_calls_callback(env, callsite) ||
+ bpf_is_callx(&env->prog->insnsi[callsite]);
+}
+
bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_spi)
{
/*
* Slot is alive if it is read before q->insn_idx in current func instance,
* or if for some outer func instance:
- * - alive before callsite if callsite calls callback, otherwise
+ * - alive before callsite if callsite calls callback or is callx, otherwise
* - alive after callsite
*/
struct live_stack_query *q = &env->liveness->live_stack_query;
@@ -394,7 +407,7 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
/* Get callsite from verifier state, not from instance callchain */
callsite = q->callsites[i];
- alive = bpf_calls_callback(env, callsite)
+ alive = callee_stack_access_at_callsite(env, callsite)
? is_live_before(instance, callsite, rel, half_spi)
: is_live_before(instance, callsite + 1, rel, half_spi);
if (alive)
@@ -1439,7 +1452,15 @@ static int record_call_access(struct bpf_verifier_env *env,
if (bpf_pseudo_call(insn))
return 0;
- if (bpf_get_call_summary(env, insn, &cs))
+ if (bpf_is_callx(insn))
+ /*
+ * The callee is not known statically. Assume that all arg
+ * slots are passed and let record_arg_access() conservatively
+ * mark the stack of all frames as read if any of them is
+ * derived from a frame pointer.
+ */
+ arg_slot_cnt = MAX_BPF_FUNC_REG_ARGS + MAX_STACK_ARG_SLOTS;
+ else if (bpf_get_call_summary(env, insn, &cs))
arg_slot_cnt = cs.arg_slot_cnt;
for (r = BPF_REG_1; r < BPF_REG_1 + min(arg_slot_cnt, MAX_BPF_FUNC_REG_ARGS); r++) {
@@ -1533,7 +1554,8 @@ static void print_subprog_arg_access(struct bpf_verifier_env *env,
bool has_extra = false;
u8 cls = BPF_CLASS(insns[idx].code);
bool is_ldx_stx_call = cls == BPF_LDX || cls == BPF_STX ||
- insns[idx].code == (BPF_JMP | BPF_CALL);
+ insns[idx].code == (BPF_JMP | BPF_CALL) ||
+ bpf_is_callx(&insns[idx]);
verbose(env, "%3d: ", idx);
bpf_verbose_insn(env, &insns[idx]);
@@ -1722,7 +1744,7 @@ static int compute_subprog_args(struct bpf_verifier_env *env,
if (err)
goto err_free;
- if (insn->code == (BPF_JMP | BPF_CALL)) {
+ if (insn->code == (BPF_JMP | BPF_CALL) || bpf_is_callx(insn)) {
err = record_call_access(env, instance, at_in[i], idx);
if (err)
goto err_free;
@@ -2202,6 +2224,9 @@ static void compute_insn_live_regs(struct bpf_verifier_env *env,
use = GENMASK(min_t(u8, cs.arg_slot_cnt, MAX_BPF_FUNC_REG_ARGS), 1);
def = mask_widen(def);
use = mask_widen(use);
+ /* callx reads the address of the callee from dst_reg */
+ if (bpf_is_callx(insn))
+ use |= dst;
break;
default:
def = 0;
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 04/17] bpf: Prepare static analysis passes for callx instruction
2026-09-22 1:13 ` [PATCH bpf-next 04/17] bpf: Prepare static analysis passes for callx instruction Alexei Starovoitov
@ 2026-09-23 23:10 ` Eduard Zingerman
2026-09-23 23:51 ` Alexei Starovoitov
0 siblings, 1 reply; 44+ messages in thread
From: Eduard Zingerman @ 2026-09-23 23:10 UTC (permalink / raw)
To: Alexei Starovoitov, bpf; +Cc: daniel, andrii, memxor, a.s.protopopov
On Tue, 2026-09-22 at 01:13 +0000, Alexei Starovoitov wrote:
...
> diff --git a/kernel/bpf/liveness.c b/kernel/bpf/liveness.c
> index 44ecdc5b4ec2..5aa2f68d92b3 100644
> --- a/kernel/bpf/liveness.c
> +++ b/kernel/bpf/liveness.c
...
> @@ -394,7 +407,7 @@ bool bpf_stack_slot_alive(struct bpf_verifier_env *env, u32 frameno, u32 half_sp
> /* Get callsite from verifier state, not from instance callchain */
> callsite = q->callsites[i];
>
> - alive = bpf_calls_callback(env, callsite)
> + alive = callee_stack_access_at_callsite(env, callsite)
> ? is_live_before(instance, callsite, rel, half_spi)
> : is_live_before(instance, callsite + 1, rel, half_spi);
> if (alive)
I don't understand this change.
The original reasoning is that callback calling function can call the
callback many times and the callback might access stack slots from
outer frames. Hence is_live_before(... callsite ...).
callx only calls target once.
The rest lgtm.
...
> @@ -1533,7 +1554,8 @@ static void print_subprog_arg_access(struct bpf_verifier_env *env,
> bool has_extra = false;
> u8 cls = BPF_CLASS(insns[idx].code);
> bool is_ldx_stx_call = cls == BPF_LDX || cls == BPF_STX ||
> - insns[idx].code == (BPF_JMP | BPF_CALL);
> + insns[idx].code == (BPF_JMP | BPF_CALL) ||
> + bpf_is_callx(&insns[idx]);
Nit: it was BPF_OP(.code) == BPF_CALL in the other hunk
>
> verbose(env, "%3d: ", idx);
> bpf_verbose_insn(env, &insns[idx]);
...
^ permalink raw reply [flat|nested] 44+ messages in thread
* Re: [PATCH bpf-next 04/17] bpf: Prepare static analysis passes for callx instruction
2026-09-23 23:10 ` Eduard Zingerman
@ 2026-09-23 23:51 ` Alexei Starovoitov
0 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-23 23:51 UTC (permalink / raw)
To: Eduard Zingerman, bpf; +Cc: daniel, andrii, memxor, a.s.protopopov
On Wed, Sep 23, 2026 at 04:10 PM Eduard Zingerman <eddyz87@gmail.com> wrote:
>> - alive = bpf_calls_callback(env, callsite)
>> + alive = callee_stack_access_at_callsite(env, callsite)
>> ? is_live_before(instance, callsite, rel, half_spi)
>> : is_live_before(instance, callsite + 1, rel, half_spi);
>> if (alive)
>
> I don't understand this change.
> The original reasoning is that callback calling function can call the
> callback many times and the callback might access stack slots from
> outer frames. Hence is_live_before(... callsite ...).
> callx only calls target once.
It's not about the number of calls.
analyze_subprog() doesn't know the callee of callx and doesn't create
an instance for it at this callsite. While the verifier is in the
callee lookup_instance() finds its standalone depth 0 instance that
knows nothing about outer frames. What the callee may read in the
caller's stack is recorded as may_read of callx insn itself by
record_call_access(). That's in live_before(callsite), but not in
live_before(callsite + 1).
callx_callee_reads_caller_stack_ok in patch 15 is a test case for this.
Without this hunk fp-8 is dead.
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH bpf-next 05/17] bpf: Add callx instruction to call bpf subprogs indirectly
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (3 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 04/17] bpf: Prepare static analysis passes for callx instruction Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-24 0:09 ` Eduard Zingerman
2026-09-22 1:13 ` [PATCH bpf-next 06/17] bpf: Add callx calls to the call graph Alexei Starovoitov
` (11 subsequent siblings)
16 siblings, 1 reply; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
Introduce BPF_JMP | BPF_CALL | BPF_X (opcode 0x8d) 'callx dst_reg'
instruction: indirect call of bpf subprog with address in dst_reg.
That's the encoding LLVM emits for calls via function pointer.
src_reg, off, imm are reserved and must be zero.
dst_reg must be PTR_TO_FUNC produced by ld_imm64 BPF_PSEUDO_FUNC.
check_ld_imm() allows it for static subprogs only, so callx cannot call
global subprogs or the main prog. Since every callee has its address
taken by ld_imm64, add_subprogs() and check_cfg() see all of them before
the main pass, and might_sleep, changes_pkt_data, might_throw of
the callee are already merged into the subprog that takes the address.
reg->subprogno is the callee. Verify callx as a direct call of that
static subprog: split check_func_call() into check_static_func_call()
that is shared with new check_func_callx(). Different paths through
the same callx may call different subprogs.
Arithmetic on PTR_TO_FUNC is allowed, so check that the pointer wasn't
modified. Allow callx while holding a lock like direct calls of static
subprogs.
The interpreter doesn't support callx. Set jit_required and add
bpf_jit_supports_callx() for JITs to opt in. No JIT does yet, so callx
is still rejected.
Print it as "callx rN" in the verifier log and xlated dump.
Adjust "invalid call insn1" test_verifier test that used opcode 0x8d as
unknown opcode. It fails with "R0 !read_ok" now.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
include/linux/filter.h | 1 +
kernel/bpf/core.c | 7 +
kernel/bpf/disasm.c | 5 +-
kernel/bpf/verifier.c | 144 ++++++++++++++----
.../selftests/bpf/verifier/basic_call.c | 2 +-
5 files changed, 131 insertions(+), 28 deletions(-)
diff --git a/include/linux/filter.h b/include/linux/filter.h
index 422284b4fa96..4f0662e42897 100644
--- a/include/linux/filter.h
+++ b/include/linux/filter.h
@@ -1237,6 +1237,7 @@ bool bpf_jit_inlines_helper_call(s32 imm);
bool bpf_jit_supports_subprog_tailcalls(void);
bool bpf_jit_supports_percpu_insn(void);
bool bpf_jit_supports_kfunc_call(void);
+bool bpf_jit_supports_callx(void);
bool bpf_jit_supports_kfunc_ret_reg_pair(void);
bool bpf_jit_supports_stack_args(void);
bool bpf_jit_supports_arena_args(void);
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index 227211166dcc..868350977f0c 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -1831,6 +1831,7 @@ bool bpf_opcode_in_insntable(u8 code)
[BPF_LD | BPF_IND | BPF_H] = true,
[BPF_LD | BPF_IND | BPF_W] = true,
[BPF_JMP | BPF_JA | BPF_X] = true,
+ [BPF_JMP | BPF_CALL | BPF_X] = true,
[BPF_JMP | BPF_JCOND] = true,
};
#undef BPF_INSN_3_TBL
@@ -3287,6 +3288,12 @@ bool __weak bpf_jit_supports_kfunc_call(void)
return false;
}
+/* Return TRUE if the JIT backend supports callx (indirect call) instruction. */
+bool __weak bpf_jit_supports_callx(void)
+{
+ return false;
+}
+
bool __weak bpf_jit_supports_kfunc_ret_reg_pair(void)
{
return false;
diff --git a/kernel/bpf/disasm.c b/kernel/bpf/disasm.c
index 3ce8d74b0e40..36d3228d7745 100644
--- a/kernel/bpf/disasm.c
+++ b/kernel/bpf/disasm.c
@@ -350,7 +350,10 @@ void print_bpf_insn(const struct bpf_insn_cbs *cbs,
if (opcode == BPF_CALL) {
char tmp[64];
- if (insn->src_reg == BPF_PSEUDO_CALL) {
+ if (BPF_SRC(insn->code) == BPF_X) {
+ verbose(cbs->private_data, "(%02x) callx r%d",
+ insn->code, insn->dst_reg);
+ } else if (insn->src_reg == BPF_PSEUDO_CALL) {
verbose(cbs->private_data, "(%02x) call pc%s",
insn->code,
__func_get_name(cbs, insn,
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 72ea662baea5..0d32d3921210 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -10600,13 +10600,61 @@ static int push_callback_call(struct bpf_verifier_env *env, struct bpf_insn *ins
static int process_bpf_exit_full(struct bpf_verifier_env *env,
bool *do_print_state, bool exception_exit);
-static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
- int *insn_idx)
+/*
+ * Call of a static subprog. The callee is verified in the context of
+ * the caller, hence set up a new frame and continue from the first
+ * instruction of the callee.
+ */
+static int check_static_func_call(struct bpf_verifier_env *env, int subprog,
+ int *insn_idx)
{
struct bpf_verifier_state *state = env->cur_state;
struct bpf_subprog_info *caller_info;
u16 callee_incoming, stack_arg_cnt;
struct bpf_func_state *caller;
+ int err;
+
+ caller = state->frame[state->curframe];
+
+ /*
+ * Track caller's total stack arg count (incoming + max outgoing).
+ * This is needed so the JIT knows how much stack arg space to allocate.
+ */
+ caller_info = &env->subprog_info[caller->subprogno];
+ callee_incoming = bpf_in_stack_arg_cnt(&env->subprog_info[subprog]);
+ stack_arg_cnt = bpf_in_stack_arg_cnt(caller_info) + callee_incoming;
+ if (stack_arg_cnt > caller_info->stack_arg_cnt)
+ caller_info->stack_arg_cnt = stack_arg_cnt;
+
+ /*
+ * For regular function entry setup new frame and continue
+ * from that frame.
+ */
+ err = setup_func_entry(env, subprog, *insn_idx, set_callee_state, state);
+ if (err)
+ return err;
+
+ bpf_diag_record_scrub(env, &caller->regs[BPF_REG_0], BPF_DIAG_MOD_CALLER_SAVED);
+ clear_caller_saved_regs(env, caller->regs);
+
+ /* and go analyze first insn of the callee */
+ *insn_idx = env->subprog_info[subprog].start - 1;
+
+ if (env->log.level & BPF_LOG_LEVEL) {
+ verbose(env, "caller:\n");
+ print_verifier_state(env, state, caller->frameno, true);
+ verbose(env, "callee:\n");
+ print_verifier_state(env, state, state->curframe, true);
+ }
+
+ return 0;
+}
+
+static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
+ int *insn_idx)
+{
+ struct bpf_verifier_state *state = env->cur_state;
+ struct bpf_func_state *caller;
int err, subprog, target_insn;
u32 i, nregs;
@@ -10695,37 +10743,70 @@ static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
return 0;
}
- /*
- * Track caller's total stack arg count (incoming + max outgoing).
- * This is needed so the JIT knows how much stack arg space to allocate.
- */
- caller_info = &env->subprog_info[caller->subprogno];
- callee_incoming = bpf_in_stack_arg_cnt(&env->subprog_info[subprog]);
- stack_arg_cnt = bpf_in_stack_arg_cnt(caller_info) + callee_incoming;
- if (stack_arg_cnt > caller_info->stack_arg_cnt)
- caller_info->stack_arg_cnt = stack_arg_cnt;
+ return check_static_func_call(env, subprog, insn_idx);
+}
- /* for regular function entry setup new frame and continue
- * from that frame.
- */
- err = setup_func_entry(env, subprog, *insn_idx, set_callee_state, state);
+/*
+ * callx dst_reg: call a bpf subprog whose address is in dst_reg.
+ *
+ * The address of a subprog is loaded into a register by ld_imm64 with
+ * src_reg == BPF_PSEUDO_FUNC, which is allowed for static subprogs only.
+ * Hence all possible callees of callx are discovered by add_subprogs() and
+ * are reachable in the control flow graph before the main verification pass
+ * begins. PTR_TO_FUNC register identifies the callee, so from here on callx
+ * is verified as a direct call of that static subprog.
+ */
+static int check_func_callx(struct bpf_verifier_env *env, struct bpf_insn *insn,
+ int *insn_idx)
+{
+ struct bpf_func_state *caller = cur_func(env);
+ struct bpf_reg_state *reg;
+ const char *reason;
+ int err, subprog;
+
+ err = check_reg_arg(env, insn->dst_reg, SRC_OP);
if (err)
return err;
- bpf_diag_record_scrub(env, &caller->regs[BPF_REG_0], BPF_DIAG_MOD_CALLER_SAVED);
- clear_caller_saved_regs(env, caller->regs);
+ reg = reg_state(env, insn->dst_reg);
+ if (reg->type != PTR_TO_FUNC) {
+ verbose(env, "R%d has type %s, expected func\n", insn->dst_reg,
+ reg_type_str(env, reg->type));
+ reason = bpf_diag_fmt(
+ env, "R%d holds %s, but callx can only call through the address of a static BPF function.",
+ insn->dst_reg, bpf_diag_reg_type_plain(env, reg->type));
+ bpf_diag_register_type(
+ env, *insn_idx, insn->dst_reg, "indirect call through a non-function pointer", reason,
+ "Load the address of a static BPF function into the register before callx.");
+ return -EACCES;
+ }
- /* and go analyze first insn of the callee */
- *insn_idx = env->subprog_info[subprog].start - 1;
+ /*
+ * Arithmetic on PTR_TO_FUNC is allowed, but only unmodified address
+ * of a subprog can be called.
+ */
+ err = check_ptr_off_reg(env, reg, insn->dst_reg);
+ if (err)
+ return err;
- if (env->log.level & BPF_LOG_LEVEL) {
- verbose(env, "caller:\n");
- print_verifier_state(env, state, caller->frameno, true);
- verbose(env, "callee:\n");
- print_verifier_state(env, state, state->curframe, true);
+ /* there is no support for callx in the interpreter */
+ if (!env->prog->jit_requested) {
+ verbose(env, "JIT is required to use callx\n");
+ return -EOPNOTSUPP;
}
+ if (!bpf_jit_supports_callx()) {
+ verbose(env, "JIT doesn't support callx\n");
+ return -EOPNOTSUPP;
+ }
+ env->prog->jit_required = true;
- return 0;
+ /* check_ld_imm() allows to take the address of static subprogs only */
+ subprog = reg->subprogno;
+ err = btf_check_subprog_call(env, subprog, caller->regs);
+ if (err == -EFAULT)
+ return err;
+
+ return check_static_func_call(env, subprog, insn_idx);
}
int map_set_for_each_callback_args(struct bpf_verifier_env *env,
@@ -18725,7 +18806,8 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
env->jmps_processed++;
if (opcode == BPF_CALL) {
- if (env->cur_state->active_locks) {
+ /* similar to static subprog calls callx is allowed under a lock */
+ if (env->cur_state->active_locks && !bpf_is_callx(insn)) {
if ((insn->src_reg == BPF_REG_0 &&
insn->imm != BPF_FUNC_spin_unlock &&
insn->imm != BPF_FUNC_kptr_xchg) ||
@@ -18743,6 +18825,8 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
mark_reg_scratched(env, BPF_REG_0);
if (bpf_in_stack_arg_cnt(&env->subprog_info[cur_func(env)->subprogno]))
cur_func(env)->no_stack_arg_load = true;
+ if (bpf_is_callx(insn))
+ return check_func_callx(env, insn, &env->insn_idx);
if (insn->src_reg == BPF_PSEUDO_CALL)
return check_func_call(env, insn, &env->insn_idx);
if (insn->src_reg == BPF_PSEUDO_KFUNC_CALL)
@@ -19509,6 +19593,14 @@ static int check_jmp_fields(struct bpf_verifier_env *env, struct bpf_insn *insn)
switch (opcode) {
case BPF_CALL:
+ if (bpf_is_callx(insn)) {
+ /* callx dst_reg */
+ if (insn->src_reg != BPF_REG_0 || insn->imm != 0 || insn->off != 0) {
+ verbose(env, "BPF_CALL|BPF_X uses reserved fields\n");
+ return -EINVAL;
+ }
+ return 0;
+ }
if (BPF_SRC(insn->code) != BPF_K ||
(insn->src_reg != BPF_PSEUDO_KFUNC_CALL && insn->off != 0) ||
(insn->src_reg != BPF_REG_0 && insn->src_reg != BPF_PSEUDO_CALL &&
diff --git a/tools/testing/selftests/bpf/verifier/basic_call.c b/tools/testing/selftests/bpf/verifier/basic_call.c
index a8c6ab4c1622..0f93c4551f23 100644
--- a/tools/testing/selftests/bpf/verifier/basic_call.c
+++ b/tools/testing/selftests/bpf/verifier/basic_call.c
@@ -4,7 +4,7 @@
BPF_RAW_INSN(BPF_JMP | BPF_CALL | BPF_X, 0, 0, 0, 0),
BPF_EXIT_INSN(),
},
- .errstr = "unknown opcode 8d",
+ .errstr = "R0 !read_ok",
.result = REJECT,
},
{
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 05/17] bpf: Add callx instruction to call bpf subprogs indirectly
2026-09-22 1:13 ` [PATCH bpf-next 05/17] bpf: Add callx instruction to call bpf subprogs indirectly Alexei Starovoitov
@ 2026-09-24 0:09 ` Eduard Zingerman
0 siblings, 0 replies; 44+ messages in thread
From: Eduard Zingerman @ 2026-09-24 0:09 UTC (permalink / raw)
To: Alexei Starovoitov, bpf; +Cc: daniel, andrii, memxor, a.s.protopopov
On Tue, 2026-09-22 at 01:13 +0000, Alexei Starovoitov wrote:
> From: Alexei Starovoitov <ast@kernel.org>
>
> Introduce BPF_JMP | BPF_CALL | BPF_X (opcode 0x8d) 'callx dst_reg'
> instruction: indirect call of bpf subprog with address in dst_reg.
> That's the encoding LLVM emits for calls via function pointer.
> src_reg, off, imm are reserved and must be zero.
>
> dst_reg must be PTR_TO_FUNC produced by ld_imm64 BPF_PSEUDO_FUNC.
> check_ld_imm() allows it for static subprogs only, so callx cannot call
> global subprogs or the main prog. Since every callee has its address
> taken by ld_imm64, add_subprogs() and check_cfg() see all of them before
> the main pass, and might_sleep, changes_pkt_data, might_throw of
> the callee are already merged into the subprog that takes the address.
>
> reg->subprogno is the callee. Verify callx as a direct call of that
> static subprog: split check_func_call() into check_static_func_call()
> that is shared with new check_func_callx(). Different paths through
> the same callx may call different subprogs.
>
> Arithmetic on PTR_TO_FUNC is allowed, so check that the pointer wasn't
> modified. Allow callx while holding a lock like direct calls of static
> subprogs.
>
> The interpreter doesn't support callx. Set jit_required and add
> bpf_jit_supports_callx() for JITs to opt in. No JIT does yet, so callx
> is still rejected.
>
> Print it as "callx rN" in the verifier log and xlated dump.
>
> Adjust "invalid call insn1" test_verifier test that used opcode 0x8d as
> unknown opcode. It fails with "R0 !read_ok" now.
>
> Signed-off-by: Alexei Starovoitov <ast@kernel.org>
> ---
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
...
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 72ea662baea5..0d32d3921210 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
...
> @@ -18725,7 +18806,8 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
>
> env->jmps_processed++;
> if (opcode == BPF_CALL) {
> - if (env->cur_state->active_locks) {
> + /* similar to static subprog calls callx is allowed under a lock */
> + if (env->cur_state->active_locks && !bpf_is_callx(insn)) {
Nit: moving !bpf_is_callx(insn) inside the nested 'if' would have been less surprising.
> if ((insn->src_reg == BPF_REG_0 &&
> insn->imm != BPF_FUNC_spin_unlock &&
> insn->imm != BPF_FUNC_kptr_xchg) ||
...
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH bpf-next 06/17] bpf: Add callx calls to the call graph
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (4 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 05/17] bpf: Add callx instruction to call bpf subprogs indirectly Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-22 1:27 ` sashiko-bot
2026-09-22 1:13 ` [PATCH bpf-next 07/17] bpf, x86: Add JIT support for callx Alexei Starovoitov
` (10 subsequent siblings)
16 siblings, 1 reply; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
The callee of callx is known to the main verifier pass only, but
sort_subprogs_topo() (recursion check) and check_max_stack_depth()
(stack size, number of frames, tail_call_reachable, private stack) need
the call graph. Both treat ld_imm64 BPF_PSEUDO_FUNC as a call from
the subprog that takes the address, which is not where callx calls it:
main: r1 = foo ll; call bar
bar: callx r1
The call chain is main -> bar -> foo, but only main -> bar and
main -> foo are seen.
Record caller -> callee edge in a bitmap when the main pass processes
callx. It explores all feasible paths, so all feasible edges are there.
Then:
- rerun sort_subprogs_topo() after the main pass with callx edges.
Unbounded recursion is already rejected by the frame limit. This one
catches bounded recursion via callx, same as direct recursion.
The call graph is not context sensitive: foo() that calls bar() via
callx and bar() that calls foo() is a recursion even if it never
happens in the same call chain.
- check_max_stack_depth_subprog() walks all recorded callees at each
callx.
JITs pass tail call counter to the callee in a register. callx doesn't.
Reject tail_call_reachable callees of callx.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
include/linux/bpf_verifier.h | 7 ++
kernel/bpf/verifier.c | 129 +++++++++++++++++++++++++++++------
2 files changed, 115 insertions(+), 21 deletions(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 6380c851ed24..931f305fa440 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -965,6 +965,13 @@ struct bpf_verifier_env {
struct bpf_subprog_info subprog_info[BPF_MAX_SUBPROGS + 2]; /* max + 2 for the fake and exception subprogs */
/* subprog indices sorted in topological order: leaves first, callers last */
int subprog_topo_order[BPF_MAX_SUBPROGS + 2];
+ /*
+ * Call graph edges created by callx instructions. A bitmap of
+ * subprog_cnt * subprog_cnt bits, where bit (caller * subprog_cnt + callee)
+ * is set when the main verification pass sees 'caller' calling 'callee'
+ * via callx. Allocated when the first such edge is recorded.
+ */
+ unsigned long *callx_edges;
union {
struct bpf_idmap idmap_scratch;
struct bpf_idset idset_scratch;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 0d32d3921210..12898d31e244 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -3144,12 +3144,50 @@ static int check_subprogs(struct bpf_verifier_env *env)
return 0;
}
+/*
+ * The callee of callx is known to the main verification pass only, which
+ * records the 'caller' -> 'callee' edge of the call graph for the checks
+ * that follow it: absence of recursion and the maximum stack depth.
+ */
+static int record_callx_edge(struct bpf_verifier_env *env, int caller, int callee)
+{
+ u32 cnt = env->subprog_cnt;
+
+ if (!env->callx_edges) {
+ env->callx_edges = kvcalloc(BITS_TO_LONGS(cnt * cnt), sizeof(long),
+ GFP_KERNEL_ACCOUNT);
+ if (!env->callx_edges)
+ return -ENOMEM;
+ }
+ __set_bit(caller * cnt + callee, env->callx_edges);
+ return 0;
+}
+
+/*
+ * Return the first subprog with the number >= 'from' that 'caller' calls
+ * via callx, or -1 when there is none.
+ */
+static int next_callx_callee(struct bpf_verifier_env *env, int caller, int from)
+{
+ u32 cnt = env->subprog_cnt;
+ unsigned long bit, end = (caller + 1) * cnt;
+
+ if (!env->callx_edges || from >= cnt)
+ return -1;
+ bit = find_next_bit(env->callx_edges, end, caller * cnt + from);
+ return bit < end ? bit - caller * cnt : -1;
+}
+
/*
* Sort subprogs in topological order so that leaf subprogs come first and
* their callers come later. This is a DFS post-order traversal of the call
* graph. Scan only reachable instructions (those in the computed postorder) of
* the current subprog to discover callees (direct subprogs and sync
* callbacks).
+ *
+ * The callees of callx are not known before the main verification pass.
+ * When callx is used the sort is repeated after it with the recorded callx
+ * edges added to the call graph to reject recursion through indirect calls.
*/
static int sort_subprogs_topo(struct bpf_verifier_env *env)
{
@@ -3190,12 +3228,22 @@ static int sort_subprogs_topo(struct bpf_verifier_env *env)
int idx = insn_postorder[j];
int callee;
- if (!bpf_pseudo_call(&insn[idx]) && !bpf_pseudo_func(&insn[idx]))
+ if (bpf_is_callx(&insn[idx])) {
+ /* find a callee that is not explored yet */
+ callee = -1;
+ do {
+ callee = next_callx_callee(env, cur, callee + 1);
+ } while (callee >= 0 && color[callee] == 2);
+ if (callee < 0)
+ continue;
+ } else if (bpf_pseudo_call(&insn[idx]) || bpf_pseudo_func(&insn[idx])) {
+ callee = bpf_find_subprog(env, idx + insn[idx].imm + 1);
+ if (callee < 0) {
+ ret = -EFAULT;
+ goto out;
+ }
+ } else {
continue;
- callee = bpf_find_subprog(env, idx + insn[idx].imm + 1);
- if (callee < 0) {
- ret = -EFAULT;
- goto out;
}
if (color[callee] == 2)
continue;
@@ -5367,6 +5415,9 @@ struct bpf_subprog_call_depth_info {
int ret_insn; /* caller instruction where we return to. */
int caller; /* caller subprogram idx */
int frame; /* # of consecutive static call stack frames on top of stack */
+ int callx_insn; /* callx instruction whose callees are being walked */
+ int callx_next; /* next callee of callx_insn to walk */
+ bool via_callx; /* the subprogram is entered via callx */
};
/* starting from main bpf function walk all instructions of the function
@@ -5386,11 +5437,13 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
/* no caller idx */
dinfo[idx].caller = -1;
+ dinfo[idx].via_callx = false;
i = subprog[idx].start;
if (!priv_stack_supported)
subprog[idx].priv_stack_mode = NO_PRIV_STACK;
process_func:
+ dinfo[idx].callx_insn = -1;
if (subprog[idx].has_ld_abs) {
for (tmp = idx; tmp >= 0; tmp = dinfo[tmp].caller) {
if (subprog[tmp].is_cb) {
@@ -5486,23 +5539,40 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
return -EINVAL;
}
- if (!bpf_pseudo_call(insn + i) && !bpf_pseudo_func(insn + i))
- continue;
- /* remember insn and function to return to */
-
- /* find the callee */
- next_insn = i + insn[i].imm + 1;
- sidx = bpf_find_subprog(env, next_insn);
- if (verifier_bug_if(sidx < 0, env, "callee not found at insn %d", next_insn))
- return -EFAULT;
- if (subprog[sidx].is_async_cb) {
- /* async callbacks don't increase bpf prog stack size unless called directly */
- if (!bpf_pseudo_call(insn + i))
+ if (bpf_is_callx(insn + i)) {
+ /*
+ * Walk the callees recorded by the main verification
+ * pass one by one, returning to this insn after each.
+ */
+ if (dinfo[idx].callx_insn != i) {
+ dinfo[idx].callx_insn = i;
+ dinfo[idx].callx_next = 0;
+ }
+ sidx = next_callx_callee(env, idx, dinfo[idx].callx_next);
+ if (sidx < 0)
continue;
- if (subprog[sidx].is_exception_cb) {
- verbose(env, "insn %d cannot call exception cb directly", i);
- return -EINVAL;
+ dinfo[idx].callx_next = sidx + 1;
+ dinfo[idx].ret_insn = i;
+ next_insn = subprog[sidx].start;
+ } else if (bpf_pseudo_call(insn + i) || bpf_pseudo_func(insn + i)) {
+ /* find the callee */
+ next_insn = i + insn[i].imm + 1;
+ sidx = bpf_find_subprog(env, next_insn);
+ if (verifier_bug_if(sidx < 0, env, "callee not found at insn %d", next_insn))
+ return -EFAULT;
+ if (subprog[sidx].is_async_cb) {
+ /* async callbacks don't increase bpf prog stack size unless called directly */
+ if (!bpf_pseudo_call(insn + i))
+ continue;
+ if (subprog[sidx].is_exception_cb) {
+ verbose(env, "insn %d cannot call exception cb directly", i);
+ return -EINVAL;
+ }
}
+ /* remember insn to return to */
+ dinfo[idx].ret_insn = i + 1;
+ } else {
+ continue;
}
/*
@@ -5523,10 +5593,10 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
/* store caller info for after we return from callee */
dinfo[idx].frame = frame;
- dinfo[idx].ret_insn = i + 1;
/* push caller idx into callee's dinfo */
dinfo[sidx].caller = idx;
+ dinfo[sidx].via_callx = bpf_is_callx(insn + i);
i = next_insn;
@@ -5560,6 +5630,14 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
verbose(env, "tail_calls are not allowed in programs with stack args\n");
return -EINVAL;
}
+ /*
+ * JITs pass tail call counter in a register that is
+ * not available when the callee is called via callx.
+ */
+ if (dinfo[tmp].via_callx) {
+ verbose(env, "tail_calls are not allowed in functions called via callx\n");
+ return -EINVAL;
+ }
subprog[tmp].tail_call_reachable = true;
}
} else if (!idx && subprog[0].has_tail_call && subprog[0].stack_arg_cnt) {
@@ -10806,6 +10884,10 @@ static int check_func_callx(struct bpf_verifier_env *env, struct bpf_insn *insn,
if (err == -EFAULT)
return err;
+ err = record_callx_edge(env, caller->subprogno, subprog);
+ if (err)
+ return err;
+
return check_static_func_call(env, subprog, insn_idx);
}
@@ -21884,6 +21966,10 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
ret = do_check_main(env);
ret = ret ?: do_check_subprogs(env);
+ /* reject recursion through the callx edges found by the main pass */
+ if (ret == 0 && env->callx_edges)
+ ret = sort_subprogs_topo(env);
+
if (ret == 0 && bpf_prog_is_offloaded(env->prog->aux))
ret = bpf_prog_offload_finalize(env);
@@ -22026,6 +22112,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
kvfree(env->scc_info);
kvfree(env->succ);
kvfree(env->gotox_tmp_buf);
+ kvfree(env->callx_edges);
bpf_diag_free(env);
kvfree(env);
return ret;
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 06/17] bpf: Add callx calls to the call graph
2026-09-22 1:13 ` [PATCH bpf-next 06/17] bpf: Add callx calls to the call graph Alexei Starovoitov
@ 2026-09-22 1:27 ` sashiko-bot
2026-09-22 2:54 ` Alexei Starovoitov
0 siblings, 1 reply; 44+ messages in thread
From: sashiko-bot @ 2026-09-22 1:27 UTC (permalink / raw)
To: Alexei Starovoitov; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: quadratic path explosion in check_max_stack_depth_subprog for callx
--
commit 17c2586ab967c2795c90663c91164f6b6733f760
Author: Alexei Starovoitov <ast@kernel.org>
bpf: Add callx calls to the call graph
This commit introduces call graph tracking for indirect callx calls.
It records caller-to-callee edges in a bitmap during the main verification
pass, which is then used by the recursion check and maximum stack depth checks.
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 0d32d3921210a..12898d31e244d 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
[ ... ]
> @@ -5486,23 +5539,40 @@ static int check_max_stack_depth_subprog(struct bpf_verifier_env *env, int idx,
> return -EINVAL;
> }
>
> - if (!bpf_pseudo_call(insn + i) && !bpf_pseudo_func(insn + i))
> - continue;
> - /* remember insn and function to return to */
> -
> - /* find the callee */
> - next_insn = i + insn[i].imm + 1;
> - sidx = bpf_find_subprog(env, next_insn);
> - if (verifier_bug_if(sidx < 0, env, "callee not found at insn %d", next_insn))
> - return -EFAULT;
> - if (subprog[sidx].is_async_cb) {
> - /* async callbacks don't increase bpf prog stack size unless called directly */
> - if (!bpf_pseudo_call(insn + i))
> + if (bpf_is_callx(insn + i)) {
> + /*
> + * Walk the callees recorded by the main verification
> + * pass one by one, returning to this insn after each.
> + */
> + if (dinfo[idx].callx_insn != i) {
> + dinfo[idx].callx_insn = i;
> + dinfo[idx].callx_next = 0;
> + }
> + sidx = next_callx_callee(env, idx, dinfo[idx].callx_next);
[Severity: High]
Could this lead to a quadratic path explosion and unbounded CPU time in the
verifier?
The callx_edges bitmap appears to aggregate all distinct callees for an
entire caller subprogram. If I am reading this correctly,
check_max_stack_depth_subprog iterates over this entire aggregate bitmap
for every single callx instruction in that subprogram.
If a subprogram has K callx instructions and M distinct callees, would the
verifier evaluate K * M edges here instead of just K edges? Across nested
subprogram calls, it seems like this combinatorial amplification could
evaluate paths exponentially. Since this pass lacks state pruning or
complexity bounds, could this bypass the verifier complexity limits and hang
the kernel thread?
> + if (sidx < 0)
> continue;
> - if (subprog[sidx].is_exception_cb) {
> - verbose(env, "insn %d cannot call exception cb directly", i);
> - return -EINVAL;
> + dinfo[idx].callx_next = sidx + 1;
> + dinfo[idx].ret_insn = i;
> + next_insn = subprog[sidx].start;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260922011323.1298619-1-alexei.starovoitov@gmail.com?part=6
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 06/17] bpf: Add callx calls to the call graph
2026-09-22 1:27 ` sashiko-bot
@ 2026-09-22 2:54 ` Alexei Starovoitov
0 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 2:54 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On Tue, Sep 22, 2026 at 01:27 AM sashiko-bot@kernel.org <sashiko-bot@kernel.org> wrote:
>> + if (bpf_is_callx(insn + i)) {
>> + /*
>> + * Walk the callees recorded by the main verification
>> + * pass one by one, returning to this insn after each.
>> + */
>> + if (dinfo[idx].callx_insn != i) {
>> + dinfo[idx].callx_insn = i;
>> + dinfo[idx].callx_next = 0;
>> + }
>> + sidx = next_callx_callee(env, idx, dinfo[idx].callx_next);
>
> [Severity: High]
> Could this lead to a quadratic path explosion and unbounded CPU time in the
> verifier?
>
> The callx_edges bitmap appears to aggregate all distinct callees for an
> entire caller subprogram. If I am reading this correctly,
> check_max_stack_depth_subprog iterates over this entire aggregate bitmap
> for every single callx instruction in that subprogram.
>
> If a subprogram has K callx instructions and M distinct callees, would the
> verifier evaluate K * M edges here instead of just K edges? Across nested
> subprogram calls, it seems like this combinatorial amplification could
> evaluate paths exponentially. Since this pass lacks state pruning or
> complexity bounds, could this bypass the verifier complexity limits and hang
> the kernel thread?
Not a bug. K * M is intentional. The edges are recorded per caller,
not per insn.
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH bpf-next 07/17] bpf, x86: Add JIT support for callx
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (5 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 06/17] bpf: Add callx calls to the call graph Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-22 1:27 ` sashiko-bot
2026-09-22 1:13 ` [PATCH bpf-next 08/17] bpf, arm64: " Alexei Starovoitov
` (9 subsequent siblings)
16 siblings, 1 reply; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
Emit indirect call via the register that holds the address of bpf subprog.
ld_imm64 BPF_PSEUDO_FUNC is resolved to func[i]->bpf_func by
jit_subprogs() already. It's the same address that is passed to helpers
as a callback.
Add emit_indirect_call() similar to emit_indirect_jump() used by tail
calls and gotox: use ITS or retpoline thunks when enabled.
Save/restore r9 around the call when private stack is used, like direct
calls do.
The verifier guarantees that callees of callx are not tail call
reachable, so don't pass tail call counter in rax, which may hold
the callee address.
FineIBT expects indirect callers to go through CFI preamble. Not done
yet. Return false from bpf_jit_supports_callx() in this mode, so such
progs are rejected at load time.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
arch/x86/net/bpf_jit_comp.c | 68 +++++++++++++++++++++++++++++++++++++
1 file changed, 68 insertions(+)
diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
index d4a980140b48..9fbef7504e51 100644
--- a/arch/x86/net/bpf_jit_comp.c
+++ b/arch/x86/net/bpf_jit_comp.c
@@ -749,6 +749,46 @@ static void emit_indirect_jump(u8 **pprog, int bpf_reg, u8 *ip)
*pprog = prog;
}
+static void __emit_indirect_call(u8 **pprog, int reg, bool ereg)
+{
+ u8 *prog = *pprog;
+
+ if (ereg)
+ EMIT1(0x41);
+
+ EMIT2(0xFF, 0xD0 + reg);
+
+ *pprog = prog;
+}
+
+/* call *bpf_reg */
+static int emit_indirect_call(u8 **pprog, int bpf_reg, u8 *ip)
+{
+ u8 *prog = *pprog;
+ int reg = reg2hex[bpf_reg];
+ bool ereg = is_ereg(bpf_reg);
+ int err = 0;
+
+ if (cpu_feature_enabled(X86_FEATURE_INDIRECT_THUNK_ITS)) {
+ OPTIMIZER_HIDE_VAR(reg);
+ err = emit_call(&prog, its_static_thunk(reg + 8*ereg), ip);
+ } else if (cpu_feature_enabled(X86_FEATURE_RETPOLINE_LFENCE)) {
+ EMIT_LFENCE();
+ __emit_indirect_call(&prog, reg, ereg);
+ } else if (cpu_feature_enabled(X86_FEATURE_RETPOLINE)) {
+ OPTIMIZER_HIDE_VAR(reg);
+ if (cpu_feature_enabled(X86_FEATURE_CALL_DEPTH))
+ err = emit_call(&prog, &__x86_indirect_call_thunk_array[reg + 8*ereg], ip);
+ else
+ err = emit_call(&prog, &__x86_indirect_thunk_array[reg + 8*ereg], ip);
+ } else {
+ __emit_indirect_call(&prog, reg, ereg);
+ }
+
+ *pprog = prog;
+ return err;
+}
+
static void emit_return(u8 **pprog, u8 *ip)
{
u8 *prog = *pprog;
@@ -2941,6 +2981,24 @@ static int do_jit(struct bpf_verifier_env *env, struct bpf_prog *bpf_prog, int *
break;
}
+ /* callx: call of a bpf subprog whose address is in dst_reg */
+ case BPF_JMP | BPF_CALL | BPF_X:
+ /*
+ * The verifier makes sure that callees of callx are
+ * not tail call reachable, hence unlike a direct call
+ * of a subprog there is no need to pass
+ * tail_call_cnt_ptr in rax.
+ */
+ if (priv_frame_ptr) {
+ push_r9(&prog);
+ ip += 2;
+ }
+ if (emit_indirect_call(&prog, insn->dst_reg, ip))
+ return -EINVAL;
+ if (priv_frame_ptr)
+ pop_r9(&prog);
+ break;
+
case BPF_JMP | BPF_TAIL_CALL:
if (imm32)
emit_bpf_tail_call_direct(bpf_prog,
@@ -4414,6 +4472,16 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
return prog;
}
+bool bpf_jit_supports_callx(void)
+{
+ /*
+ * FineIBT poisons ENDBR at the entry of a JITed function and expects
+ * indirect callers to go through the CFI preamble instead.
+ * callx doesn't do that yet.
+ */
+ return cfi_mode != CFI_FINEIBT;
+}
+
bool bpf_jit_supports_kfunc_call(void)
{
return true;
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 07/17] bpf, x86: Add JIT support for callx
2026-09-22 1:13 ` [PATCH bpf-next 07/17] bpf, x86: Add JIT support for callx Alexei Starovoitov
@ 2026-09-22 1:27 ` sashiko-bot
2026-09-22 2:53 ` Alexei Starovoitov
0 siblings, 1 reply; 44+ messages in thread
From: sashiko-bot @ 2026-09-22 1:27 UTC (permalink / raw)
To: Alexei Starovoitov; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: callx bypasses KCFI due to missing caller-side hash check
--
commit d4dfd1b46a3e02ea97924c833ac04239e01627f1
Author: Alexei Starovoitov <ast@kernel.org>
Subject: bpf, x86: Add JIT support for callx
This commit adds JIT support for BPF indirect calls via the register that
holds the address of the BPF subprog. It introduces the emit_indirect_call
helper to handle the translation using ITS or retpoline thunks.
> diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c
> --- a/arch/x86/net/bpf_jit_comp.c
> +++ b/arch/x86/net/bpf_jit_comp.c
[ ... ]
> @@ -749,6 +749,46 @@ static void emit_indirect_jump(u8 **pprog, int bpf_reg, u8 *ip)
> *pprog = prog;
> }
>
> +static void __emit_indirect_call(u8 **pprog, int reg, bool ereg)
> +{
> + u8 *prog = *pprog;
> +
> + if (ereg)
> + EMIT1(0x41);
> +
> + EMIT2(0xFF, 0xD0 + reg);
> +
> + *pprog = prog;
> +}
> +
> +/* call *bpf_reg */
> +static int emit_indirect_call(u8 **pprog, int bpf_reg, u8 *ip)
> +{
> + u8 *prog = *pprog;
> + int reg = reg2hex[bpf_reg];
> + bool ereg = is_ereg(bpf_reg);
> + int err = 0;
> +
> + if (cpu_feature_enabled(X86_FEATURE_INDIRECT_THUNK_ITS)) {
> + OPTIMIZER_HIDE_VAR(reg);
> + err = emit_call(&prog, its_static_thunk(reg + 8*ereg), ip);
> + } else if (cpu_feature_enabled(X86_FEATURE_RETPOLINE_LFENCE)) {
> + EMIT_LFENCE();
> + __emit_indirect_call(&prog, reg, ereg);
> + } else if (cpu_feature_enabled(X86_FEATURE_RETPOLINE)) {
> + OPTIMIZER_HIDE_VAR(reg);
> + if (cpu_feature_enabled(X86_FEATURE_CALL_DEPTH))
> + err = emit_call(&prog, &__x86_indirect_call_thunk_array[reg + 8*ereg], ip);
> + else
> + err = emit_call(&prog, &__x86_indirect_thunk_array[reg + 8*ereg], ip);
> + } else {
> + __emit_indirect_call(&prog, reg, ereg);
[Severity: High]
Does emit_indirect_call() emit raw indirect branches without KCFI hash
validation?
If KCFI is active, it seems __emit_indirect_call() emits a bare indirect
branch instruction without the preceding caller-side hash check.
If an attacker were able to overwrite a spilled subprogram pointer, could they
trigger this execution path to jump to an arbitrary kernel address and bypass
KCFI protections?
> + }
> +
> + *pprog = prog;
> + return err;
> +}
> +
[ ... ]
> @@ -4414,6 +4472,16 @@ struct bpf_prog *bpf_int_jit_compile(struct bpf_verifier_env *env, struct bpf_pr
> return prog;
> }
>
> +bool bpf_jit_supports_callx(void)
> +{
> + /*
> + * FineIBT poisons ENDBR at the entry of a JITed function and expects
> + * indirect callers to go through the CFI preamble instead.
> + * callx doesn't do that yet.
> + */
> + return cfi_mode != CFI_FINEIBT;
[Severity: High]
Can bpf_jit_supports_callx() incorrectly allow indirect calls to be compiled
when KCFI is enabled?
The current logic blocks FineIBT but permits compilation when the CFI mode is
CFI_KCFI.
Since emit_indirect_call() does not currently emit caller-side hash checks,
allowing compilation under KCFI seems to introduce uninstrumented indirect
branches into kernel space.
> +}
> +
> bool bpf_jit_supports_kfunc_call(void)
> {
> return true;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260922011323.1298619-1-alexei.starovoitov@gmail.com?part=7
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 07/17] bpf, x86: Add JIT support for callx
2026-09-22 1:27 ` sashiko-bot
@ 2026-09-22 2:53 ` Alexei Starovoitov
0 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 2:53 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On Tue, Sep 22, 2026 at 01:27 AM sashiko-bot@kernel.org <sashiko-bot@kernel.org> wrote:
>> + } else {
>> + __emit_indirect_call(&prog, reg, ereg);
>
> [Severity: High]
> Does emit_indirect_call() emit raw indirect branches without KCFI hash
> validation?
>
> If KCFI is active, it seems __emit_indirect_call() emits a bare indirect
> branch instruction without the preceding caller-side hash check.
>
> If an attacker were able to overwrite a spilled subprogram pointer, could they
> trigger this execution path to jump to an arbitrary kernel address and bypass
> KCFI protections?
Not a bug. The kCFI hash check is what the compiler emits in front of
indirect calls in C code. JIT emits the callee side only, see
emit_kcfi(), so that C callers of bpf_func and of callbacks pass their
check. None of the indirect branches emitted by JIT check the hash:
tail call, gotox, trampoline. callx is no different. The target is
PTR_TO_FUNC with zero offset. The verifier guarantees that it's the
entry of a static subprog of this prog.
Whoever can overwrite a spilled pointer on bpf stack can overwrite
the return address next to it.
[...]
>> + return cfi_mode != CFI_FINEIBT;
>
> [Severity: High]
> Can bpf_jit_supports_callx() incorrectly allow indirect calls to be compiled
> when KCFI is enabled?
No. Same as above. With CFI_KCFI bpf_func points to ENDBR after the
hash, so 'call *reg' into it works. With FineIBT that ENDBR is
poisoned and the call would fault. That's the only reason FineIBT is
excluded.
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH bpf-next 08/17] bpf, arm64: Add JIT support for callx
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (6 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 07/17] bpf, x86: Add JIT support for callx Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-22 15:05 ` Puranjay Mohan
2026-09-22 1:13 ` [PATCH bpf-next 09/17] bpf: Discover subprogs described by func_info Alexei Starovoitov
` (8 subsequent siblings)
16 siblings, 1 reply; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
Emit callx as BLR to the register with the address of the subprog and
move x0 into BPF_REG_0 like after a direct call.
That's what a direct call of a subprog out of BL range is JITed into
already. bpf functions start with BTI JC. The arguments are in x0-x7 and
on the stack. Tail call counter ptr and private stack ptr are in callee
saved x26 and x27.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
arch/arm64/net/bpf_jit_comp.c | 16 ++++++++++++++++
1 file changed, 16 insertions(+)
diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
index 6c04fee46876..9544b2f483e5 100644
--- a/arch/arm64/net/bpf_jit_comp.c
+++ b/arch/arm64/net/bpf_jit_comp.c
@@ -1789,6 +1789,17 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
emit(A64_MOV(1, r0, A64_R(0)), ctx);
break;
}
+ /* indirect call of a bpf subprog, dst holds its address */
+ case BPF_JMP | BPF_CALL | BPF_X:
+ /*
+ * It's the same as a direct call of a subprog that is out of
+ * range of BL: the subprog starts with BTI JC, the arguments
+ * are in place, and the registers that hold the tail call
+ * counter and the private stack are callee saved.
+ */
+ emit(A64_BLR(dst), ctx);
+ emit(A64_MOV(1, bpf2a64[BPF_REG_0], A64_R(0)), ctx);
+ break;
/* tail call */
case BPF_JMP | BPF_TAIL_CALL:
if (emit_bpf_tail_call(ctx))
@@ -2485,6 +2496,11 @@ bool bpf_jit_supports_subprog_tailcalls(void)
return true;
}
+bool bpf_jit_supports_callx(void)
+{
+ return true;
+}
+
static void invoke_bpf_prog(struct jit_ctx *ctx, struct bpf_tramp_node *node,
int bargs_off, int retval_off, int run_ctx_off,
bool save_ret)
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 08/17] bpf, arm64: Add JIT support for callx
2026-09-22 1:13 ` [PATCH bpf-next 08/17] bpf, arm64: " Alexei Starovoitov
@ 2026-09-22 15:05 ` Puranjay Mohan
0 siblings, 0 replies; 44+ messages in thread
From: Puranjay Mohan @ 2026-09-22 15:05 UTC (permalink / raw)
To: Alexei Starovoitov, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, Puranjay Mohan
Alexei Starovoitov <alexei.starovoitov@gmail.com> writes:
> From: Alexei Starovoitov <ast@kernel.org>
>
> Emit callx as BLR to the register with the address of the subprog and
> move x0 into BPF_REG_0 like after a direct call.
>
> That's what a direct call of a subprog out of BL range is JITed into
> already. bpf functions start with BTI JC. The arguments are in x0-x7 and
> on the stack. Tail call counter ptr and private stack ptr are in callee
> saved x26 and x27.
>
> Signed-off-by: Alexei Starovoitov <ast@kernel.org>
> ---
> arch/arm64/net/bpf_jit_comp.c | 16 ++++++++++++++++
> 1 file changed, 16 insertions(+)
>
> diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c
> index 6c04fee46876..9544b2f483e5 100644
> --- a/arch/arm64/net/bpf_jit_comp.c
> +++ b/arch/arm64/net/bpf_jit_comp.c
> @@ -1789,6 +1789,17 @@ static int build_insn(const struct bpf_verifier_env *env, const struct bpf_insn
> emit(A64_MOV(1, r0, A64_R(0)), ctx);
> break;
> }
> + /* indirect call of a bpf subprog, dst holds its address */
> + case BPF_JMP | BPF_CALL | BPF_X:
> + /*
> + * It's the same as a direct call of a subprog that is out of
> + * range of BL: the subprog starts with BTI JC, the arguments
> + * are in place, and the registers that hold the tail call
> + * counter and the private stack are callee saved.
> + */
> + emit(A64_BLR(dst), ctx);
> + emit(A64_MOV(1, bpf2a64[BPF_REG_0], A64_R(0)), ctx);
> + break;
> /* tail call */
> case BPF_JMP | BPF_TAIL_CALL:
> if (emit_bpf_tail_call(ctx))
> @@ -2485,6 +2496,11 @@ bool bpf_jit_supports_subprog_tailcalls(void)
> return true;
> }
>
> +bool bpf_jit_supports_callx(void)
> +{
> + return true;
> +}
> +
Reviewed-by: Puranjay Mohan <puranjay@kernel.org>
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH bpf-next 09/17] bpf: Discover subprogs described by func_info
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (7 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 08/17] bpf, arm64: " Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-22 1:13 ` [PATCH bpf-next 10/17] bpf: Recognize pointers to functions in read-only maps Alexei Starovoitov
` (7 subsequent siblings)
16 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
add_subprogs() finds subprogs by scanning insns for calls and ld_imm64
BPF_PSEUDO_FUNC. A function that is called via callx through a pointer
in data only, e.g. a vtable, is not referenced by any insn, but it must
be known before check_cfg() and the main pass.
func_info describes all functions of the program and libbpf provides it
for every function it puts into the prog. bpf_prepare_btf_info() reads
it before add_subprogs(). Register every function from func_info as
a subprog.
func_info had to match the subprogs found in insns exactly, so existing
programs are not affected. A function that is listed in func_info and
not referenced by anything is still rejected by check_cfg() as
unreachable.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
kernel/bpf/verifier.c | 16 ++++++++++++++++
1 file changed, 16 insertions(+)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 12898d31e244..c37a1d3eb8c8 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -3018,6 +3018,22 @@ static int add_subprogs(struct bpf_verifier_env *env)
return ret;
}
+ /*
+ * func_info describes all functions of the program. Those that are
+ * referenced only from data, e.g. from a table of functions to be
+ * called via callx, are not seen by the loop above. They are possible
+ * callees that have to be known upfront as well.
+ */
+ if (env->bpf_capable) {
+ struct bpf_prog_aux *aux = env->prog->aux;
+
+ for (i = 1; i < aux->func_info_cnt; i++) {
+ ret = add_subprog(env, aux->func_info[i].insn_off);
+ if (ret < 0)
+ return ret;
+ }
+ }
+
ret = bpf_find_exception_callback_insn_off(env);
if (ret < 0)
return ret;
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* [PATCH bpf-next 10/17] bpf: Recognize pointers to functions in read-only maps
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (8 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 09/17] bpf: Discover subprogs described by func_info Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-22 1:31 ` sashiko-bot
2026-09-24 0:46 ` bot+bpf-ci
2026-09-22 1:13 ` [PATCH bpf-next 11/17] libbpf: Support pointers to static functions in data when linking Alexei Starovoitov
` (6 subsequent siblings)
16 siblings, 2 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
Compilers put pointers to functions into read-only data: tables of
functions, struct ops, vtables, where they're mixed with sizes,
alignments and other data. Rust vtables are like that.
Let callx call through them:
r6 = vtable ll
r1 = *(u64 *)(r6 + 0) // data
r2 = *(u64 *)(r6 + 8) // pointer to a function
callx r2
libbpf keeps such data in a frozen read-only array map like .rodata and
stores the byte offset of static function within the program into
the pointer. The kernel guesses: aligned 64-bit value that is equal to
the offset of a static subprog is a pointer to it. If the guess is wrong
the program either fails to load, since it uses a pointer as a number,
or sees an address instead of the number. Safe either way.
Add resolve_func_ptrs() that scans such maps of a program with callx
after the maps are resolved and before check_cfg(), so all callees of
callx, including the ones referenced by data only, are known before
the main pass. The prog must be the only user of the map and it must
be allowed to leak pointers (CAP_PERFMON), since it reads function
addresses as data.
- check_cfg() treats ld_imm64 of such map like ld_imm64 BPF_PSEUDO_FUNC
for every function in the map: they're reachable, first insns are
prune and jump points, effects are merged into the subprog that
refers to the map.
- 64-bit load of the pointer is PTR_TO_FUNC. If the offset is variable
(tbl[i]) all possible offsets, given bounds and var_off, must be
pointers. The verifier forks a state for each.
- The value of the pointer is not known until JIT, so reject everything
that would const-fold its bytes: narrow or misaligned loads, const
strings. bpf_compute_const_regs() skips them too.
After JIT jit_subprogs() replaces the offsets in the map with addresses
of functions. Other progs that use the map were verified with its old
content const-folded, so there must be none. Add 'user' to bpf_map:
__add_used_map() records the first prog that uses the map and marks
the map as shared when another prog comes. Only the sole user can store
addresses into the map. After that no other prog can use it. The prog
that failed to load is not a user. libbpf loads it again to get the log.
The offsets are adjusted when insns are patched or removed, same as
subprog starts. Dead function is removed and the pointer becomes NULL.
Reject the prog if it reads such pointer.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
include/linux/bpf.h | 10 +
include/linux/bpf_verifier.h | 30 +++
kernel/bpf/cfg.c | 75 ++++++++
kernel/bpf/const_fold.c | 3 +
kernel/bpf/core.c | 5 +
kernel/bpf/fixups.c | 55 ++++++
kernel/bpf/verifier.c | 343 ++++++++++++++++++++++++++++++++++-
7 files changed, 513 insertions(+), 8 deletions(-)
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index fd22db8bc6c5..7747c5fc290f 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -342,8 +342,18 @@ struct bpf_map {
s64 __percpu *elem_count;
u64 cookie; /* write-once */
char *excl_prog_sha;
+ /*
+ * Which programs use the map, see bpf_map_claim(): 0 - none so far,
+ * aux of the program - only that one, the same with BPF_MAP_USER_PATCHED
+ * set - only that one and it stored the addresses of its functions into
+ * the map, BPF_MAP_USER_MANY - more than one.
+ */
+ unsigned long user;
};
+#define BPF_MAP_USER_MANY 1UL
+#define BPF_MAP_USER_PATCHED 1UL
+
static inline const char *btf_field_type_name(enum btf_field_type type)
{
switch (type) {
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 931f305fa440..38a4ba50669a 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -786,6 +786,22 @@ int bpf_log_attr_finalize(struct bpf_log_attr *attr, struct bpf_verifier_log *lo
#define BPF_MAX_SUBPROGS 256
+/*
+ * A pointer to a static subprog in the value of a frozen read-only array map:
+ * a 64-bit value that is the offset in bytes of the first instruction of
+ * the subprog in the program.
+ */
+struct bpf_func_ptr {
+ struct bpf_map *map;
+ u32 map_off; /* offset of the pointer in the value of the map */
+ u32 orig_off; /* what the map has: the first instruction of the subprog */
+ u32 xlated_off; /* the same after instructions were patched and removed */
+ bool used; /* the program reads the pointer */
+};
+
+/* the subprog that a bpf_func_ptr pointed to was removed as dead code */
+#define BPF_FUNC_PTR_DELETED ((u32)-1)
+
struct bpf_subprog_arg_info {
enum bpf_arg_type arg_type;
union {
@@ -965,6 +981,13 @@ struct bpf_verifier_env {
struct bpf_subprog_info subprog_info[BPF_MAX_SUBPROGS + 2]; /* max + 2 for the fake and exception subprogs */
/* subprog indices sorted in topological order: leaves first, callers last */
int subprog_topo_order[BPF_MAX_SUBPROGS + 2];
+ /*
+ * Pointers to static subprogs found in frozen read-only maps of the
+ * program, see resolve_func_ptrs(). Sorted by map and map_off.
+ */
+ struct bpf_func_ptr *func_ptrs;
+ u32 func_ptr_cnt;
+ bool has_callx;
/*
* Call graph edges created by callx instructions. A bitmap of
* subprog_cnt * subprog_cnt bits, where bit (caller * subprog_cnt + callee)
@@ -1324,6 +1347,13 @@ static inline bool bt_is_frame_slot_set(struct backtrack_state *bt, u32 frame, u
}
bool bpf_map_is_rdonly(const struct bpf_map *map);
+struct bpf_func_ptr *bpf_map_func_ptrs(struct bpf_verifier_env *env,
+ const struct bpf_map *map, u32 *cnt);
+struct bpf_func_ptr *bpf_map_range_func_ptrs(struct bpf_verifier_env *env,
+ const struct bpf_map *map,
+ u64 off, u64 size, u32 *cnt);
+void bpf_adjust_func_ptrs(struct bpf_verifier_env *env, u32 off, u32 len);
+void bpf_adjust_func_ptrs_after_remove(struct bpf_verifier_env *env, u32 off, u32 len);
int bpf_map_direct_read(struct bpf_map *map, int off, int size, u64 *val,
bool is_ldsx);
diff --git a/kernel/bpf/cfg.c b/kernel/bpf/cfg.c
index 842c7d1eabcc..a068c191409a 100644
--- a/kernel/bpf/cfg.c
+++ b/kernel/bpf/cfg.c
@@ -421,6 +421,75 @@ static int visit_gotox_insn(int t, struct bpf_verifier_env *env)
return keep_exploring ? KEEP_EXPLORING : DONE_EXPLORING;
}
+/*
+ * Return pointers to functions in the read-only map that ld_imm64 instruction
+ * 't' loads the address of, or of its value, if there are any.
+ */
+static struct bpf_func_ptr *insn_func_ptrs(struct bpf_verifier_env *env, int t, u32 *cnt)
+{
+ struct bpf_insn *insn = &env->prog->insnsi[t];
+
+ *cnt = 0;
+ if (!env->func_ptr_cnt || !bpf_is_ldimm64(insn))
+ return NULL;
+ if (insn->src_reg != BPF_PSEUDO_MAP_VALUE &&
+ insn->src_reg != BPF_PSEUDO_MAP_IDX_VALUE &&
+ insn->src_reg != BPF_PSEUDO_MAP_FD &&
+ insn->src_reg != BPF_PSEUDO_MAP_IDX)
+ return NULL;
+
+ return bpf_map_func_ptrs(env, env->used_maps[env->insn_aux_data[t].map_index], cnt);
+}
+
+/*
+ * ld_imm64 that loads the address of a map that has pointers to functions
+ * is similar to ld_imm64 with BPF_PSEUDO_FUNC that loads the address of one
+ * function: any of them may be read from the map and called via callx later.
+ * Treat it as a call of all of them.
+ */
+static int visit_func_ptrs_insn(int t, struct bpf_verifier_env *env,
+ struct bpf_func_ptr *ptrs, u32 cnt)
+{
+ int *insn_stack = env->cfg.insn_stack;
+ int *insn_state = env->cfg.insn_state;
+ bool keep_exploring = false;
+ int ret, w;
+ u32 i;
+
+ ret = push_insn(t, t + 2, FALLTHROUGH, env);
+ if (ret)
+ return ret;
+
+ mark_prune_point(env, t);
+ for (i = 0; i < cnt; i++) {
+ w = ptrs[i].xlated_off;
+
+ /*
+ * This function is called until all functions are explored,
+ * so the effects are complete in the end.
+ */
+ merge_callee_effects(env, t, w);
+
+ /* the same marks as push_insn() leaves on a branch target */
+ mark_prune_point(env, w);
+ mark_jmp_point(env, w);
+ mark_jump_target(env, w);
+
+ /* EXPLORED || DISCOVERED */
+ if (insn_state[w])
+ continue;
+
+ if (env->cfg.cur_stack >= env->prog->len)
+ return -E2BIG;
+
+ insn_stack[env->cfg.cur_stack++] = w;
+ insn_state[w] |= DISCOVERED;
+ keep_exploring = true;
+ }
+
+ return keep_exploring ? KEEP_EXPLORING : DONE_EXPLORING;
+}
+
/*
* Instructions that can abnormally return from a subprog (tail_call
* upon success, ld_{abs,ind} upon load failure) have a hidden exit
@@ -453,11 +522,17 @@ static int visit_abnormal_return_insn(struct bpf_verifier_env *env, int t)
static int visit_insn(int t, struct bpf_verifier_env *env)
{
struct bpf_insn *insns = env->prog->insnsi, *insn = &insns[t];
+ struct bpf_func_ptr *ptrs;
int ret, off, insn_sz;
+ u32 cnt;
if (bpf_pseudo_func(insn))
return visit_func_call_insn(t, insns, env, true);
+ ptrs = insn_func_ptrs(env, t, &cnt);
+ if (ptrs)
+ return visit_func_ptrs_insn(t, env, ptrs, cnt);
+
/* All non-branch instructions have a single fall-through edge. */
if (BPF_CLASS(insn->code) != BPF_JMP &&
BPF_CLASS(insn->code) != BPF_JMP32) {
diff --git a/kernel/bpf/const_fold.c b/kernel/bpf/const_fold.c
index f44ae8487ec6..fea639f62b3f 100644
--- a/kernel/bpf/const_fold.c
+++ b/kernel/bpf/const_fold.c
@@ -180,6 +180,7 @@ static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *
bool is_ldsx = mode == BPF_MEMSX;
int off = src->val + insn->off;
u64 val = 0;
+ u32 cnt;
/*
* Values of insn_array map are addresses of jitted instructions,
@@ -188,6 +189,8 @@ static void const_reg_xfer(struct bpf_verifier_env *env, struct const_arg_info *
if (!bpf_map_is_rdonly(map) || !map->ops->map_direct_value_addr ||
map->map_type == BPF_MAP_TYPE_INSN_ARRAY ||
off < 0 || off + size > map->value_size ||
+ /* so are the addresses of functions that the map points to */
+ bpf_map_range_func_ptrs(env, map, off, size, &cnt) ||
bpf_map_direct_read(map, off, size, &val, is_ldsx)) {
*dst = unknown;
break;
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index 868350977f0c..273f74068068 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -3029,6 +3029,11 @@ void __bpf_free_used_maps(struct bpf_prog_aux *aux,
map->ops->map_poke_untrack(map, aux);
if (sleepable)
atomic64_dec(&map->sleepable_refcnt);
+ /*
+ * The program that didn't load is not a user of the map. libbpf
+ * loads the program again to get the log of the verifier.
+ */
+ cmpxchg(&map->user, (unsigned long)aux, 0);
bpf_map_put(map);
}
}
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 2add8001c3ec..e568b9b790b5 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -361,6 +361,7 @@ struct bpf_prog *bpf_patch_insn_data(struct bpf_verifier_env *env, u32 off,
adjust_insn_aux_data(env, new_prog, off, len, &original_insn);
adjust_subprog_starts(env, off, len);
adjust_insn_arrays(env, off, len);
+ bpf_adjust_func_ptrs(env, off, len);
adjust_poke_descs(new_prog, off, len);
return new_prog;
}
@@ -559,6 +560,9 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt)
if (err)
return err;
+ /* before subprogs are adjusted, since it looks at them */
+ bpf_adjust_func_ptrs_after_remove(env, off, cnt);
+
err = adjust_subprog_starts_after_remove(env, off, cnt);
if (err)
return err;
@@ -1285,6 +1289,57 @@ static int jit_subprogs(struct bpf_verifier_env *env)
cond_resched();
}
+ /*
+ * The addresses of all functions are final. Replace the offsets of
+ * functions with them in the maps of the program, see
+ * resolve_func_ptrs(). The program must be the only user of such map.
+ * From now on no other program can use it, see bpf_map_claim().
+ */
+ for (i = 0; i < env->func_ptr_cnt; i++) {
+ struct bpf_func_ptr *ptr = &env->func_ptrs[i];
+ unsigned long me = (unsigned long)prog->aux;
+ u64 addr, old, new = 0;
+
+ /* pointers are sorted by map */
+ if ((!i || ptr->map != ptr[-1].map) &&
+ cmpxchg(&ptr->map->user, me, me | BPF_MAP_USER_PATCHED) != me) {
+ verbose(env, "map '%s' is used by another program\n", ptr->map->name);
+ err = -EBUSY;
+ goto out_free;
+ }
+
+ /* it's the address of the value of the map whatever the offset is */
+ err = ptr->map->ops->map_direct_value_addr(ptr->map, &addr, 0);
+ if (verifier_bug_if(err, env, "no value of map '%s'", ptr->map->name)) {
+ err = -EFAULT;
+ goto out_free;
+ }
+ addr += ptr->map_off;
+
+ if (ptr->xlated_off != BPF_FUNC_PTR_DELETED) {
+ subprog = bpf_find_subprog(env, ptr->xlated_off);
+ if (verifier_bug_if(subprog <= 0, env, "no function at insn %u",
+ ptr->xlated_off)) {
+ err = -EFAULT;
+ goto out_free;
+ }
+ new = (unsigned long)func[subprog]->bpf_func;
+ } else if (verifier_bug_if(ptr->used, env, "function of map '%s' offset %u is removed",
+ ptr->map->name, ptr->map_off)) {
+ /* the program that reads the pointer might call the function */
+ err = -EFAULT;
+ goto out_free;
+ }
+ /* else the function is dead code, nothing calls it, the pointer is NULL */
+
+ old = (u64)ptr->orig_off * sizeof(struct bpf_insn);
+ if (verifier_bug_if(cmpxchg64((u64 *)(unsigned long)addr, old, new) != old, env,
+ "map '%s' offset %u changed", ptr->map->name, ptr->map_off)) {
+ err = -EFAULT;
+ goto out_free;
+ }
+ }
+
/*
* Cleanup func[i]->aux fields which aren't required
* or can become invalid in future
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index c37a1d3eb8c8..8507114cf690 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -3115,6 +3115,8 @@ static int check_subprogs(struct bpf_verifier_env *env)
if (BPF_CLASS(code) == BPF_LD &&
(BPF_MODE(code) == BPF_ABS || BPF_MODE(code) == BPF_IND))
subprog[cur_subprog].has_ld_abs = true;
+ if (bpf_is_callx(&insn[i]))
+ env->has_callx = true;
if (BPF_CLASS(code) != BPF_JMP && BPF_CLASS(code) != BPF_JMP32)
goto next;
if (BPF_OP(code) == BPF_CALL)
@@ -6018,6 +6020,110 @@ int bpf_map_direct_read(struct bpf_map *map, int off, int size, u64 *val,
return 0;
}
+static int cmp_func_ptrs(const void *_a, const void *_b)
+{
+ const struct bpf_func_ptr *a = _a, *b = _b;
+
+ if (a->map != b->map)
+ return a->map < b->map ? -1 : 1;
+ if (a->map_off != b->map_off)
+ return a->map_off < b->map_off ? -1 : 1;
+ return 0;
+}
+
+/* Find the first pointer to a function at or after 'off' in the value of 'map' */
+static u32 func_ptr_lower_bound(struct bpf_verifier_env *env, const struct bpf_map *map, u64 off)
+{
+ u32 l = 0, r = env->func_ptr_cnt, m;
+ struct bpf_func_ptr *p;
+
+ while (l < r) {
+ m = l + (r - l) / 2;
+ p = &env->func_ptrs[m];
+ if (p->map < map || (p->map == map && p->map_off < off))
+ l = m + 1;
+ else
+ r = m;
+ }
+ return l;
+}
+
+/*
+ * Return pointers to functions that overlap with 'size' bytes at offset 'off'
+ * of the value of 'map' and their number in 'cnt'.
+ */
+struct bpf_func_ptr *bpf_map_range_func_ptrs(struct bpf_verifier_env *env,
+ const struct bpf_map *map,
+ u64 off, u64 size, u32 *cnt)
+{
+ u32 first, last;
+
+ *cnt = 0;
+ if (!env->func_ptr_cnt || !size)
+ return NULL;
+
+ /* a pointer that starts up to 7 bytes before 'off' overlaps too */
+ first = func_ptr_lower_bound(env, map, off >= sizeof(u64) ? off - sizeof(u64) + 1 : 0);
+ last = func_ptr_lower_bound(env, map, off + size);
+ if (first >= last)
+ return NULL;
+
+ *cnt = last - first;
+ return &env->func_ptrs[first];
+}
+
+/* Return all pointers to functions in the value of 'map' */
+struct bpf_func_ptr *bpf_map_func_ptrs(struct bpf_verifier_env *env,
+ const struct bpf_map *map, u32 *cnt)
+{
+ return bpf_map_range_func_ptrs(env, map, 0, (u64)map->value_size, cnt);
+}
+
+/* instructions [off, off + len) replaced the instruction at 'off' */
+void bpf_adjust_func_ptrs(struct bpf_verifier_env *env, u32 off, u32 len)
+{
+ struct bpf_func_ptr *p;
+ u32 i;
+
+ if (len <= 1)
+ return;
+
+ for (i = 0; i < env->func_ptr_cnt; i++) {
+ p = &env->func_ptrs[i];
+ if (p->xlated_off <= off || p->xlated_off == BPF_FUNC_PTR_DELETED)
+ continue;
+ p->xlated_off += len - 1;
+ }
+}
+
+/*
+ * Instructions [off, off + len) are about to be removed. It's called before
+ * the starts of subprogs are adjusted. A subprog is gone when all of its
+ * instructions are. Otherwise, e.g. when its first instruction is a nop,
+ * it starts where the removed instructions did.
+ */
+void bpf_adjust_func_ptrs_after_remove(struct bpf_verifier_env *env, u32 off, u32 len)
+{
+ struct bpf_func_ptr *p;
+ int subprog;
+ u32 i;
+
+ for (i = 0; i < env->func_ptr_cnt; i++) {
+ p = &env->func_ptrs[i];
+ if (p->xlated_off < off || p->xlated_off == BPF_FUNC_PTR_DELETED)
+ continue;
+ if (p->xlated_off >= off + len) {
+ p->xlated_off -= len;
+ continue;
+ }
+ subprog = bpf_find_subprog(env, p->xlated_off);
+ if (subprog > 0 && env->subprog_info[subprog + 1].start > off + len)
+ p->xlated_off = off;
+ else
+ p->xlated_off = BPF_FUNC_PTR_DELETED;
+ }
+}
+
#define BTF_TYPE_SAFE_RCU(__type) __PASTE(__type, __safe_rcu)
#define BTF_TYPE_SAFE_RCU_OR_NULL(__type) __PASTE(__type, __safe_rcu_or_null)
#define BTF_TYPE_SAFE_TRUSTED(__type) __PASTE(__type, __safe_trusted)
@@ -6527,12 +6633,100 @@ static void add_scalar_to_reg(struct bpf_reg_state *dst_reg, s64 val)
reg_bounds_sync(dst_reg);
}
+static void mark_reg_func_ptr(struct bpf_verifier_env *env, struct bpf_reg_state *regs,
+ int regno, int subprog)
+{
+ mark_reg_known_zero(env, regs, regno);
+ regs[regno].type = PTR_TO_FUNC;
+ regs[regno].subprogno = subprog;
+}
+
+/* a read from a table of functions branches into that many states at most */
+#define BPF_MAX_FUNC_PTR_TARGETS 64
+/* and the table, which might have other data in it, is that many pointers long at most */
+#define BPF_MAX_FUNC_PTR_RANGE 4096
+
+/*
+ * A read from a frozen read-only map that has pointers to functions, see
+ * resolve_func_ptrs(). A read of exactly one pointer yields PTR_TO_FUNC.
+ * When the offset is variable and only pointers can be read, which is how
+ * an element of a table of functions is loaded, the verification continues
+ * with each of them. Other reads that overlap with a pointer are rejected,
+ * because their result is not known until the program is jitted.
+ *
+ * Return -ENOENT if there are no pointers to functions in the bytes that are read.
+ */
+static int check_func_ptr_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,
+ int size, int value_regno)
+{
+ u64 min_off = reg_umin(reg) + off, max_off = reg_umax(reg) + off;
+ struct tnum offs = tnum_add(reg->var_off, tnum_const(off));
+ struct bpf_reg_state *regs = cur_regs(env);
+ struct bpf_map *map = reg->map_ptr;
+ struct bpf_verifier_state *branch;
+ int subprog, targets[BPF_MAX_FUNC_PTR_TARGETS];
+ struct bpf_func_ptr *ptrs;
+ u32 i, cnt, n = 0;
+ u64 o;
+
+ ptrs = bpf_map_range_func_ptrs(env, map, min_off, max_off - min_off + size, &cnt);
+ if (!ptrs)
+ return -ENOENT;
+
+ if (size != sizeof(u64) || value_regno < 0 || !tnum_is_aligned(offs, sizeof(u64)) ||
+ (max_off - min_off) / sizeof(u64) > BPF_MAX_FUNC_PTR_RANGE)
+ goto overlap;
+
+ /*
+ * Every offset that the read is possible at has to be the offset of
+ * a pointer. var_off tells the stride of the elements of an array.
+ */
+ for (o = round_up(min_off, sizeof(u64)), i = 0; o <= max_off; o += sizeof(u64)) {
+ if ((o ^ offs.value) & ~offs.mask)
+ continue;
+ while (i < cnt && ptrs[i].map_off < o)
+ i++;
+ if (i == cnt || ptrs[i].map_off != o)
+ goto overlap;
+ if (n == BPF_MAX_FUNC_PTR_TARGETS) {
+ verbose(env, "read from map '%s' may yield more than %d pointers to functions\n",
+ map->name, BPF_MAX_FUNC_PTR_TARGETS);
+ return -E2BIG;
+ }
+ subprog = bpf_find_subprog(env, ptrs[i].xlated_off);
+ if (verifier_bug_if(subprog <= 0, env, "no function at insn %u for map '%s' offset %u",
+ ptrs[i].xlated_off, map->name, ptrs[i].map_off))
+ return -EFAULT;
+ ptrs[i].used = true;
+ targets[n++] = subprog;
+ }
+ if (verifier_bug_if(!n, env, "no offsets to read map '%s' at", map->name))
+ return -EFAULT;
+
+ for (i = 0; i < n - 1; i++) {
+ branch = push_stack(env, env->insn_idx + 1, env->insn_idx,
+ env->cur_state->speculative);
+ if (IS_ERR(branch))
+ return PTR_ERR(branch);
+ mark_reg_func_ptr(env, branch->frame[branch->curframe]->regs, value_regno,
+ targets[i]);
+ }
+ mark_reg_func_ptr(env, regs, value_regno, targets[n - 1]);
+ return 0;
+
+overlap:
+ verbose(env, "read of %d bytes at offset [%llu,%llu] of map '%s' overlaps with a pointer to a function\n",
+ size, min_off, max_off, map->name);
+ return -EACCES;
+}
+
static int check_map_mem_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,
int bpf_size, int value_regno, bool is_ldsx)
{
struct bpf_reg_state *regs = cur_regs(env);
int size = bpf_size_to_bytes(bpf_size);
struct bpf_map *map = reg->map_ptr;
+ int err;
switch (map->map_type) {
case BPF_MAP_TYPE_INSN_ARRAY:
@@ -6550,13 +6744,18 @@ static int check_map_mem_read(struct bpf_verifier_env *env, struct bpf_reg_state
break;
}
+ if (env->func_ptr_cnt) {
+ err = check_func_ptr_read(env, reg, off, size, value_regno);
+ if (err != -ENOENT)
+ return err;
+ }
+
/* If map is read-only, track its contents as scalars. */
if (tnum_is_const(reg->var_off) &&
bpf_map_is_rdonly(map) &&
map->ops->map_direct_value_addr) {
int map_off = off + reg->var_off.value;
u64 val = 0;
- int err;
err = bpf_map_direct_read(map, map_off, size, &val, is_ldsx);
if (err)
@@ -8864,6 +9063,7 @@ static int check_arg_const_str(struct bpf_verifier_env *env,
int map_off;
u64 map_addr;
char *str_ptr;
+ u32 cnt;
if (reg->type != PTR_TO_MAP_VALUE)
return -EINVAL;
@@ -8913,6 +9113,11 @@ static int check_arg_const_str(struct bpf_verifier_env *env,
verbose(env, "string is not zero-terminated\n");
return -EINVAL;
}
+ /* the bytes of a pointer to a function are not known until the program is jitted */
+ if (bpf_map_range_func_ptrs(env, map, map_off, strlen(str_ptr + map_off) + 1, &cnt)) {
+ verbose(env, "string overlaps with a pointer to a function\n");
+ return -EACCES;
+ }
return 0;
}
@@ -10843,12 +11048,13 @@ static int check_func_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
/*
* callx dst_reg: call a bpf subprog whose address is in dst_reg.
*
- * The address of a subprog is loaded into a register by ld_imm64 with
- * src_reg == BPF_PSEUDO_FUNC, which is allowed for static subprogs only.
- * Hence all possible callees of callx are discovered by add_subprogs() and
- * are reachable in the control flow graph before the main verification pass
- * begins. PTR_TO_FUNC register identifies the callee, so from here on callx
- * is verified as a direct call of that static subprog.
+ * The address of a subprog is either loaded into a register by ld_imm64 with
+ * src_reg == BPF_PSEUDO_FUNC, or it is read from a frozen read-only map, see
+ * resolve_func_ptrs(). Both are possible for static subprogs only. Hence all
+ * possible callees of callx are discovered by add_subprogs() and are reachable
+ * in the control flow graph before the main verification pass begins.
+ * PTR_TO_FUNC register identifies the callee, so from here on callx is verified
+ * as a direct call of that static subprog.
*/
static int check_func_callx(struct bpf_verifier_env *env, struct bpf_insn *insn,
int *insn_idx)
@@ -10894,7 +11100,7 @@ static int check_func_callx(struct bpf_verifier_env *env, struct bpf_insn *insn,
}
env->prog->jit_required = true;
- /* check_ld_imm() allows to take the address of static subprogs only */
+ /* PTR_TO_FUNC is a pointer to a static subprog */
subprog = reg->subprogno;
err = btf_check_subprog_call(env, subprog, caller->regs);
if (err == -EFAULT)
@@ -19500,6 +19706,30 @@ static int check_map_prog_compatibility(struct bpf_verifier_env *env,
return 0;
}
+/*
+ * Keep track of whether the map is used by one program only. Such program may
+ * store the addresses of its functions into the map when it's frozen, see
+ * resolve_func_ptrs(), since nothing else relies on what the map has. After
+ * that the map is not available to other programs.
+ */
+static int bpf_map_claim(struct bpf_verifier_env *env, struct bpf_map *map)
+{
+ unsigned long me = (unsigned long)env->prog->aux, old;
+
+ for (;;) {
+ old = READ_ONCE(map->user);
+ if (old == me || old == BPF_MAP_USER_MANY)
+ return 0;
+ if (old & BPF_MAP_USER_PATCHED) {
+ verbose(env, "map '%s' has addresses of functions of another program\n",
+ map->name);
+ return -EBUSY;
+ }
+ if (cmpxchg(&map->user, old, old ? BPF_MAP_USER_MANY : me) == old)
+ return 0;
+ }
+}
+
static int __add_used_map(struct bpf_verifier_env *env, struct bpf_map *map)
{
int i, err;
@@ -19537,6 +19767,10 @@ static int __add_used_map(struct bpf_verifier_env *env, struct bpf_map *map)
env->used_maps[env->used_map_cnt++] = map;
+ err = bpf_map_claim(env, map);
+ if (err)
+ return err;
+
if (map->map_type == BPF_MAP_TYPE_INSN_ARRAY) {
err = bpf_insn_array_init(map, env->prog);
if (err) {
@@ -19953,6 +20187,93 @@ static int check_and_resolve_insns(struct bpf_verifier_env *env)
return 0;
}
+static int add_func_ptr(struct bpf_verifier_env *env, struct bpf_map *map, u32 map_off,
+ u32 xlated_off)
+{
+ struct bpf_func_ptr *ptrs;
+
+ /* grow by doubling, the array is sorted and searched later */
+ if (!(env->func_ptr_cnt & (env->func_ptr_cnt - 1))) {
+ ptrs = kvrealloc(env->func_ptrs,
+ array_size(max(2 * env->func_ptr_cnt, 16U), sizeof(*ptrs)),
+ GFP_KERNEL_ACCOUNT);
+ if (!ptrs)
+ return -ENOMEM;
+ env->func_ptrs = ptrs;
+ }
+ env->func_ptrs[env->func_ptr_cnt++] = (struct bpf_func_ptr){
+ .map = map,
+ .map_off = map_off,
+ .orig_off = xlated_off,
+ .xlated_off = xlated_off,
+ };
+ return 0;
+}
+
+/*
+ * Compilers put pointers to functions into read-only data: tables of functions,
+ * structures of operations, vtables, where they are mixed with other data.
+ * The loader stores such data in a frozen read-only array map and resolves
+ * a pointer to a static function to the offset in bytes of its first
+ * instruction in the program: the address of the function in the program.
+ *
+ * Find 64-bit values that look like that in the maps of a program that uses
+ * callx. It's a guess. When the value is not a pointer, the program either
+ * fails to load, because it does with a pointer what can be done with
+ * a number only, or it sees the address of a function instead of the number.
+ * It's known before the control flow graph of the program is built and the main
+ * verification pass begins which functions may be called via callx.
+ *
+ * When the program is jitted the offsets are replaced with the addresses of
+ * the functions in the map itself, see jit_subprogs(). Hence the program has to
+ * be the only user of the map, see bpf_map_claim(): nothing else may rely on
+ * what the map had.
+ *
+ * The program reads the addresses of its functions from there like any other
+ * data, so it has to be allowed to leak pointers.
+ */
+static int resolve_func_ptrs(struct bpf_verifier_env *env)
+{
+ int insn_cnt = env->prog->len;
+ int i, err, subprog;
+ struct bpf_map *map;
+ u64 addr, val;
+ u32 off;
+
+ if (!env->has_callx || !env->allow_ptr_leaks)
+ return 0;
+
+ for (i = 0; i < env->used_map_cnt; i++) {
+ map = env->used_maps[i];
+ /* coincidences in maps that are shared with other programs don't matter */
+ if (READ_ONCE(map->user) != (unsigned long)env->prog->aux)
+ continue;
+ if (map->map_type != BPF_MAP_TYPE_ARRAY || map->max_entries != 1 ||
+ !bpf_map_is_rdonly(map) || !map->ops->map_direct_value_addr ||
+ !IS_ERR_OR_NULL(map->record))
+ continue;
+ if (map->ops->map_direct_value_addr(map, &addr, 0))
+ continue;
+
+ for (off = 0; off + sizeof(u64) <= map->value_size; off += sizeof(u64)) {
+ val = *(u64 *)(unsigned long)(addr + off);
+ if (!val || val % sizeof(struct bpf_insn) ||
+ val / sizeof(struct bpf_insn) >= insn_cnt)
+ continue;
+ subprog = bpf_find_subprog(env, val / sizeof(struct bpf_insn));
+ if (subprog <= 0 || bpf_subprog_is_global(env, subprog))
+ continue;
+ err = add_func_ptr(env, map, off, val / sizeof(struct bpf_insn));
+ if (err)
+ return err;
+ }
+ }
+ if (env->func_ptr_cnt)
+ sort(env->func_ptrs, env->func_ptr_cnt, sizeof(*env->func_ptrs),
+ cmp_func_ptrs, NULL);
+ return 0;
+}
+
/* drop refcnt of maps used by the rejected program */
static void release_maps(struct bpf_verifier_env *env)
{
@@ -21923,6 +22244,11 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
if (ret < 0)
goto skip_full_check;
+ /* Find pointers to functions in the read-only maps of the program. */
+ ret = resolve_func_ptrs(env);
+ if (ret < 0)
+ goto skip_full_check;
+
/* Build kfunc prototypes after resolving program resources. */
ret = add_kfuncs(env);
if (ret < 0)
@@ -22129,6 +22455,7 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
kvfree(env->succ);
kvfree(env->gotox_tmp_buf);
kvfree(env->callx_edges);
+ kvfree(env->func_ptrs);
bpf_diag_free(env);
kvfree(env);
return ret;
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 10/17] bpf: Recognize pointers to functions in read-only maps
2026-09-22 1:13 ` [PATCH bpf-next 10/17] bpf: Recognize pointers to functions in read-only maps Alexei Starovoitov
@ 2026-09-22 1:31 ` sashiko-bot
2026-09-22 3:01 ` Alexei Starovoitov
2026-09-24 0:46 ` bot+bpf-ci
1 sibling, 1 reply; 44+ messages in thread
From: sashiko-bot @ 2026-09-22 1:31 UTC (permalink / raw)
To: Alexei Starovoitov; +Cc: bpf
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] bpf: verifier incorrectly rejects valid programs and triggers WARN_ONCE when reading uncalled function pointers
- [Medium] bpf: map user state is not restored on JIT failure, breaking libbpf log reloads
--
commit ac5645576aa732578e623a1ca8af9bcff7aa9194
Author: Alexei Starovoitov <ast@kernel.org>
bpf: Recognize pointers to functions in read-only maps
Compilers put pointers to functions into read-only data: tables of
functions, struct ops, vtables, where they're mixed with sizes,
alignments and other data. Rust vtables are like that.
Let callx call through them.
> diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
> index 868350977f0cb..273f74068068c 100644
> --- a/kernel/bpf/core.c
> +++ b/kernel/bpf/core.c
[ ... ]
> @@ -3029,6 +3029,11 @@ void __bpf_free_used_maps(struct bpf_prog_aux *aux,
> map->ops->map_poke_untrack(map, aux);
> if (sleepable)
> atomic64_dec(&map->sleepable_refcnt);
> + /*
> + * The program that didn't load is not a user of the map. libbpf
> + * loads the program again to get the log of the verifier.
> + */
> + cmpxchg(&map->user, (unsigned long)aux, 0);
[Severity: Medium]
Will this properly clear the map's user state if the BPF_MAP_USER_PATCHED
flag was set before a failure?
If jit_subprogs() sets the BPF_MAP_USER_PATCHED flag, map->user becomes
me | BPF_MAP_USER_PATCHED. If a failure happens afterward, this cmpxchg()
strictly expects map->user to equal aux, which will fail to match.
This can leave the map permanently locked in the patched state, causing
libbpf's subsequent reload attempt (to obtain the verifier log) to fail
with -EBUSY.
> bpf_map_put(map);
> }
> }
> diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
> index 2add8001c3ec3..e568b9b790b5c 100644
> --- a/kernel/bpf/fixups.c
> +++ b/kernel/bpf/fixups.c
[ ... ]
> @@ -1285,6 +1289,57 @@ static int jit_subprogs(struct bpf_verifier_env *env)
> cond_resched();
> }
>
> + /*
> + * The addresses of all functions are final. Replace the offsets of
> + * functions with them in the maps of the program, see
> + * resolve_func_ptrs(). The program must be the only user of such map.
> + * From now on no other program can use it, see bpf_map_claim().
> + */
> + for (i = 0; i < env->func_ptr_cnt; i++) {
> + struct bpf_func_ptr *ptr = &env->func_ptrs[i];
> + unsigned long me = (unsigned long)prog->aux;
> + u64 addr, old, new = 0;
> +
> + /* pointers are sorted by map */
> + if ((!i || ptr->map != ptr[-1].map) &&
> + cmpxchg(&ptr->map->user, me, me | BPF_MAP_USER_PATCHED) != me) {
> + verbose(env, "map '%s' is used by another program\n", ptr->map->name);
> + err = -EBUSY;
> + goto out_free;
> + }
> +
> + /* it's the address of the value of the map whatever the offset is */
> + err = ptr->map->ops->map_direct_value_addr(ptr->map, &addr, 0);
> + if (verifier_bug_if(err, env, "no value of map '%s'", ptr->map->name)) {
> + err = -EFAULT;
> + goto out_free;
> + }
> + addr += ptr->map_off;
> +
> + if (ptr->xlated_off != BPF_FUNC_PTR_DELETED) {
> + subprog = bpf_find_subprog(env, ptr->xlated_off);
> + if (verifier_bug_if(subprog <= 0, env, "no function at insn %u",
> + ptr->xlated_off)) {
> + err = -EFAULT;
> + goto out_free;
> + }
> + new = (unsigned long)func[subprog]->bpf_func;
> + } else if (verifier_bug_if(ptr->used, env, "function of map '%s' offset %u is removed",
> + ptr->map->name, ptr->map_off)) {
[Severity: High]
Does eagerly setting ptr->used = true in check_func_ptr_read() lead to a
verifier bug here if the function is never actually called?
If a program simply reads the function pointer from the map but does not
invoke callx, check_func_ptr_read() still flags it as used. Since the
subprogram is never called, dead-code elimination can remove it.
When jit_subprogs() later evaluates the deleted but "used" pointer, it
incorrectly triggers a WARN_ONCE and returns -EFAULT, rejecting a valid
program.
> + /* the program that reads the pointer might call the function */
> + err = -EFAULT;
> + goto out_free;
> + }
> + /* else the function is dead code, nothing calls it, the pointer is NULL */
> +
> + old = (u64)ptr->orig_off * sizeof(struct bpf_insn);
> + if (verifier_bug_if(cmpxchg64((u64 *)(unsigned long)addr, old, new) != old, env,
> + "map '%s' offset %u changed", ptr->map->name, ptr->map_off)) {
> + err = -EFAULT;
> + goto out_free;
> + }
> + }
> +
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260922011323.1298619-1-alexei.starovoitov@gmail.com?part=10
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 10/17] bpf: Recognize pointers to functions in read-only maps
2026-09-22 1:31 ` sashiko-bot
@ 2026-09-22 3:01 ` Alexei Starovoitov
0 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 3:01 UTC (permalink / raw)
To: sashiko-reviews; +Cc: bpf
On Tue, Sep 22, 2026 at 01:31 AM sashiko-bot@kernel.org <sashiko-bot@kernel.org> wrote:
>> + /*
>> + * The program that didn't load is not a user of the map. libbpf
>> + * loads the program again to get the log of the verifier.
>> + */
>> + cmpxchg(&map->user, (unsigned long)aux, 0);
>
> [Severity: Medium]
> Will this properly clear the map's user state if the BPF_MAP_USER_PATCHED
> flag was set before a failure?
>
> If jit_subprogs() sets the BPF_MAP_USER_PATCHED flag, map->user becomes
> me | BPF_MAP_USER_PATCHED. If a failure happens afterward, this cmpxchg()
> strictly expects map->user to equal aux, which will fail to match.
>
> This can leave the map permanently locked in the patched state, causing
> libbpf's subsequent reload attempt (to obtain the verifier log) to fail
> with -EBUSY.
No. Not a bug.
>> + } else if (verifier_bug_if(ptr->used, env, "function of map '%s' offset %u is removed",
>> + ptr->map->name, ptr->map_off)) {
>
> [Severity: High]
> Does eagerly setting ptr->used = true in check_func_ptr_read() lead to a
> verifier bug here if the function is never actually called?
>
> If a program simply reads the function pointer from the map but does not
> invoke callx, check_func_ptr_read() still flags it as used. Since the
> subprogram is never called, dead-code elimination can remove it.
>
> When jit_subprogs() later evaluates the deleted but "used" pointer, it
> incorrectly triggers a WARN_ONCE and returns -EFAULT, rejecting a valid
> program.
hmm. A prog can read the pointer and never call the
function, e.g. when the load is hoisted above a branch. Then the
function is dead, 'used' is set and jit_subprogs() hits
verifier_bug_if().
Will follow up with one line fix.
^ permalink raw reply [flat|nested] 44+ messages in thread
* Re: [PATCH bpf-next 10/17] bpf: Recognize pointers to functions in read-only maps
2026-09-22 1:13 ` [PATCH bpf-next 10/17] bpf: Recognize pointers to functions in read-only maps Alexei Starovoitov
2026-09-22 1:31 ` sashiko-bot
@ 2026-09-24 0:46 ` bot+bpf-ci
2026-09-24 2:12 ` Alexei Starovoitov
1 sibling, 1 reply; 44+ messages in thread
From: bot+bpf-ci @ 2026-09-24 0:46 UTC (permalink / raw)
To: alexei.starovoitov, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, ast, andrii,
daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 7295 bytes --]
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 2ebd08a3c736..958a8897cbf6 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
[ ... ]
> @@ -6532,12 +6638,100 @@ static void add_scalar_to_reg(struct bpf_reg_state *dst_reg, s64 val)
> reg_bounds_sync(dst_reg);
> }
>
> +static void mark_reg_func_ptr(struct bpf_verifier_env *env, struct bpf_reg_state *regs,
> + int regno, int subprog)
> +{
> + mark_reg_known_zero(env, regs, regno);
> + regs[regno].type = PTR_TO_FUNC;
> + regs[regno].subprogno = subprog;
> +}
> +
> +/* a read from a table of functions branches into that many states at most */
> +#define BPF_MAX_FUNC_PTR_TARGETS 64
> +/* and the table, which might have other data in it, is that many pointers long at most */
> +#define BPF_MAX_FUNC_PTR_RANGE 4096
> +
> +/*
> + * A read from a frozen read-only map that has pointers to functions, see
> + * resolve_func_ptrs(). A read of exactly one pointer yields PTR_TO_FUNC.
> + * When the offset is variable and only pointers can be read, which is how
> + * an element of a table of functions is loaded, the verification continues
> + * with each of them. Other reads that overlap with a pointer are rejected,
> + * because their result is not known until the program is jitted.
> + *
> + * Return -ENOENT if there are no pointers to functions in the bytes that are read.
> + */
> +static int check_func_ptr_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,
> + int size, int value_regno)
> +{
> + u64 min_off = reg_umin(reg) + off, max_off = reg_umax(reg) + off;
> + struct tnum offs = tnum_add(reg->var_off, tnum_const(off));
> + struct bpf_reg_state *regs = cur_regs(env);
> + struct bpf_map *map = reg->map_ptr;
> + struct bpf_verifier_state *branch;
> + int subprog, targets[BPF_MAX_FUNC_PTR_TARGETS];
> + struct bpf_func_ptr *ptrs;
> + u32 i, cnt, n = 0;
> + u64 o;
> +
> + ptrs = bpf_map_range_func_ptrs(env, map, min_off, max_off - min_off + size, &cnt);
> + if (!ptrs)
> + return -ENOENT;
> +
> + if (size != sizeof(u64) || value_regno < 0 || !tnum_is_aligned(offs, sizeof(u64)) ||
> + (max_off - min_off) / sizeof(u64) > BPF_MAX_FUNC_PTR_RANGE)
> + goto overlap;
> +
> + /*
> + * Every offset that the read is possible at has to be the offset of
> + * a pointer. var_off tells the stride of the elements of an array.
> + */
> + for (o = round_up(min_off, sizeof(u64)), i = 0; o <= max_off; o += sizeof(u64)) {
> + if ((o ^ offs.value) & ~offs.mask)
> + continue;
> + while (i < cnt && ptrs[i].map_off < o)
> + i++;
> + if (i == cnt || ptrs[i].map_off != o)
> + goto overlap;
Can reads that don't actually overlap pointers still be rejected here?
Looking at the bounds-based overlap check, when a vtable-like struct has
both data and function pointers (e.g., struct { u64 data; void (*fn)(void); }
with stride 16), reading tbl[i].data gives var_off with offsets 0, 16, 32...
where every tnum-possible offset is non-pointer data.
However, bpf_map_range_func_ptrs() returns pointers in the bounds interval
[min_off, max_off + size), which includes the fn slots at 8, 24, etc. The
check then rejects the read because ptrs[i].map_off != o for the data offsets,
even though none of the tnum-possible offsets actually touch a pointer.
The commit message mentions supporting tables where pointers are mixed with
sizes and other data, so these programs seem like valid use cases. Should
the overlap check walk only the tnum-possible offsets and return -ENOENT when
none of them overlap a pointer?
> + if (n == BPF_MAX_FUNC_PTR_TARGETS) {
> + verbose(env, "read from map '%s' may yield more than %d pointers to functions\n",
> + map->name, BPF_MAX_FUNC_PTR_TARGETS);
> + return -E2BIG;
> + }
> + subprog = bpf_find_subprog(env, ptrs[i].xlated_off);
> + if (verifier_bug_if(subprog <= 0, env, "no function at insn %u for map '%s' offset %u",
> + ptrs[i].xlated_off, map->name, ptrs[i].map_off))
> + return -EFAULT;
> + ptrs[i].used = true;
> + targets[n++] = subprog;
> + }
> + if (verifier_bug_if(!n, env, "no offsets to read map '%s' at", map->name))
> + return -EFAULT;
> +
> + for (i = 0; i < n - 1; i++) {
> + branch = push_stack(env, env->insn_idx + 1, env->insn_idx,
> + env->cur_state->speculative);
> + if (IS_ERR(branch))
> + return PTR_ERR(branch);
> + mark_reg_func_ptr(env, branch->frame[branch->curframe]->regs, value_regno,
> + targets[i]);
> + }
> + mark_reg_func_ptr(env, regs, value_regno, targets[n - 1]);
> + return 0;
> +
> +overlap:
> + verbose(env, "read of %d bytes at offset [%llu,%llu] of map '%s' overlaps with a pointer to a function\n",
> + size, min_off, max_off, map->name);
> + return -EACCES;
> +}
> +
> static int check_map_mem_read(struct bpf_verifier_env *env, struct bpf_reg_state *reg, int off,
> int bpf_size, int value_regno, bool is_ldsx)
> {
> struct bpf_reg_state *regs = cur_regs(env);
> int size = bpf_size_to_bytes(bpf_size);
> struct bpf_map *map = reg->map_ptr;
> + int err;
>
> switch (map->map_type) {
> case BPF_MAP_TYPE_INSN_ARRAY:
> @@ -6555,13 +6749,18 @@ static int check_map_mem_read(struct bpf_verifier_env *env, struct bpf_reg_state
> break;
> }
>
> + if (env->func_ptr_cnt) {
> + err = check_func_ptr_read(env, reg, off, size, value_regno);
> + if (err != -ENOENT)
> + return err;
> + }
Does a PTR_TO_FUNC loaded from a map force the program to be JITed?
check_func_ptr_read() marks the register as PTR_TO_FUNC but doesn't set
env->prog->jit_required. The map still holds the byte offset (e.g. 0x40),
not a real address until jit_subprogs() patches it.
Helpers like bpf_for_each_map_elem() and bpf_timer_set_callback() accept
any PTR_TO_FUNC via check_func_arg() -> ARG_PTR_TO_FUNC, which just sets
meta->subprogno = reg->subprogno. They call the register value as a kernel
function pointer at runtime.
bpf_fixup_call_args() falls back to the interpreter when jit_requested is
false (bpf_jit_enable=0 without CONFIG_BPF_JIT_ALWAYS_ON=y) or when
bpf_jit_subprogs() fails with anything other than -EFAULT. That includes the
new -EBUSY claim failure, -ENOMEM, or a subprog JIT failure.
The interpreter's callback check only looks for ld_imm64 BPF_PSEUDO_FUNC:
kernel/bpf/fixups.c:bpf_fixup_call_args() {
if (bpf_pseudo_func(insn)) {
verbose(env, "callbacks are not allowed in non-JITed programs\n");
return -EINVAL;
A PTR_TO_FUNC loaded from a map has no such insn, so the program loads and
runs in the interpreter with the map unpatched. The helper then calls the
raw offset (0x40) as a kernel function pointer.
A program needs only a callx on a dead branch (e.g. gated by a .rodata
constant folded to false) to enable resolve_func_ptrs(), since
check_subprogs() sets env->has_callx from a static scan. check_func_callx()
and its JIT requirement check never run.
Should check_func_ptr_read() or resolve_func_ptrs() check
jit_requested && bpf_jit_supports_callx() and set jit_required, the same way
check_func_callx() does?
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35935998342
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 10/17] bpf: Recognize pointers to functions in read-only maps
2026-09-24 0:46 ` bot+bpf-ci
@ 2026-09-24 2:12 ` Alexei Starovoitov
0 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-24 2:12 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, martin.lau,
yonghong.song, mason, ihor.solodrai
On Thu, Sep 24, 2026 at 12:46 AM bot+bpf-ci@kernel.org <bot+bpf-ci@kernel.org> wrote:
>> + while (i < cnt && ptrs[i].map_off < o)
>> + i++;
>> + if (i == cnt || ptrs[i].map_off != o)
>> + goto overlap;
>
> Can reads that don't actually overlap pointers still be rejected here?
> Looking at the bounds-based overlap check, when a vtable-like struct has
> both data and function pointers (e.g., struct { u64 data; void (*fn)(void); }
> with stride 16), reading tbl[i].data gives var_off with offsets 0, 16, 32...
> where every tnum-possible offset is non-pointer data.
That's a limitation of v1,
but sure I'll try to make it work in v2.
>> + if (env->func_ptr_cnt) {
>> + err = check_func_ptr_read(env, reg, off, size, value_regno);
>
> Does a PTR_TO_FUNC loaded from a map force the program to be JITed?
> check_func_ptr_read() marks the register as PTR_TO_FUNC but doesn't set
> env->prog->jit_required. The map still holds the byte offset (e.g. 0x40),
> not a real address until jit_subprogs() patches it.
Will fix in v2: the read of such pointer will set jit_required.
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH bpf-next 11/17] libbpf: Support pointers to static functions in data when linking
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (9 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 10/17] bpf: Recognize pointers to functions in read-only maps Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-22 1:13 ` [PATCH bpf-next 12/17] libbpf: Resolve pointers to functions in read-only data Alexei Starovoitov
` (5 subsequent siblings)
16 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
A pointer to a static function in data, e.g.
static int (* const handlers[])(void *ctx) = { foo, bar };
is R_BPF_64_ABS64 relocation against .text section symbol with the offset
of the function stored in place. The linker rejects it:
relocation against STT_SECTION in non-exec section is not supported!
so an object with a table of static functions cannot go through
'bpftool gen object', which all selftests do.
Functions move by the offset of input .text in the output section.
Add it to the pointer. Swap bytes if the object is of foreign endianness,
since data sections are kept in the byte order of the object.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
tools/lib/bpf/linker.c | 18 ++++++++++++++++++
1 file changed, 18 insertions(+)
diff --git a/tools/lib/bpf/linker.c b/tools/lib/bpf/linker.c
index 78f92c39290a..53f64a1a1f25 100644
--- a/tools/lib/bpf/linker.c
+++ b/tools/lib/bpf/linker.c
@@ -2274,6 +2274,24 @@ static int linker_append_elf_relos(struct bpf_linker *linker, struct src_obj *ob
insn->imm += sec->dst_off / sizeof(struct bpf_insn);
else
insn->imm += sec->dst_off;
+ } else if (sym_type == R_BPF_64_ABS64 &&
+ (sec->shdr->sh_flags & SHF_EXECINSTR)) {
+ /*
+ * A pointer to a static function in a data section,
+ * which is stored in place as an offset of the
+ * function in its section. Data sections are kept
+ * in the byte order of the object.
+ */
+ void *ptr = dst_linked_sec->raw_data + dst_rel->r_offset;
+ __u64 off;
+
+ memcpy(&off, ptr, sizeof(off));
+ if (linker->swapped_endian)
+ off = bswap_64(off);
+ off += sec->dst_off;
+ if (linker->swapped_endian)
+ off = bswap_64(off);
+ memcpy(ptr, &off, sizeof(off));
} else {
pr_warn("relocation against STT_SECTION in non-exec section is not supported!\n");
return -EINVAL;
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* [PATCH bpf-next 12/17] libbpf: Resolve pointers to functions in read-only data
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (10 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 11/17] libbpf: Support pointers to static functions in data when linking Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-24 0:33 ` bot+bpf-ci
2026-09-22 1:13 ` [PATCH bpf-next 13/17] libbpf: Treat .data.rel.ro as " Alexei Starovoitov
` (4 subsequent siblings)
16 siblings, 1 reply; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
The kernel recognizes a pointer to a function in frozen read-only array
map by value: byte offset of a static function in the program.
Make it work for tables of functions, struct ops, vtables in .rodata:
static const struct shape_ops square_ops = { 1, area, 10, perimeter };
...
return ops->area(x) * ops->scale;
is compiled into
r3 = *(u64 *)(r1 + 8)
callx r3
where 'area' in .rodata is R_BPF_64_ABS64 relocation against .text.
- Collect relocations in .rodata* that are aligned pointers to
functions. Relocations in data were ignored. The rest still are.
- Append all static functions that the section points to to the prog
that refers to the section, since any of them can be called. callx
cannot call global functions. Don't pull them in. Such pointer is NULL.
- Functions have different offsets in different progs, so create a copy
of the map per prog and store the offsets there. The kernel replaces
them with addresses at load and the prog must be the only user of
the map. The map of the section itself is left as-is.
- The kernel treats any aligned u64 that matches an offset of a static
function as a pointer. Warn if data has one that isn't.
Light skeleton is not supported yet.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
tools/lib/bpf/libbpf.c | 372 ++++++++++++++++++++++++++++++++++++++++-
1 file changed, 369 insertions(+), 3 deletions(-)
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index cd1ea1bb53cb..9538a21f0db4 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -601,6 +601,9 @@ struct bpf_map {
bool autoattach;
__u64 map_extra;
struct bpf_program *excl_prog;
+ /* pointers to functions in the data of an internal map, see obj->func_ptrs */
+ struct func_ptr *func_ptrs;
+ size_t func_ptr_cnt;
};
enum extern_type {
@@ -780,6 +783,29 @@ struct bpf_object {
} *jumptable_maps;
size_t jumptable_map_cnt;
+ /*
+ * Pointers to functions found in read-only data sections: tables of
+ * functions, structures of operations, vtables. Sorted by section
+ * and offset.
+ */
+ struct func_ptr {
+ int sec_idx; /* ELF section that contains the pointer */
+ size_t sec_off; /* offset of the pointer in the section */
+ size_t text_off; /* offset of the function in .text section */
+ } *func_ptrs;
+ size_t func_ptr_cnt;
+
+ /*
+ * Read-only data with pointers to functions is different for every
+ * program that uses it, because so are the offsets of the functions.
+ */
+ struct {
+ struct bpf_program *prog;
+ int map_idx;
+ int fd;
+ } *func_ptr_maps;
+ size_t func_ptr_map_cnt;
+
struct kern_feature_cache *feat_cache;
char *token_path;
int token_fd;
@@ -845,6 +871,17 @@ static bool insn_is_pseudo_func(struct bpf_insn *insn)
return is_ldimm64_insn(insn) && insn->src_reg == BPF_PSEUDO_FUNC;
}
+/*
+ * ld_imm64 that loads the address of read-only data with pointers to functions
+ * is marked by bpf_object__relocate() before the code is relocated. Compilers
+ * leave src_reg of other ld_imm64 zero and it's set by
+ * bpf_object__relocate_data() later.
+ */
+static bool insn_is_func_ptrs_addr(struct bpf_insn *insn)
+{
+ return is_ldimm64_insn(insn) && insn->src_reg == BPF_PSEUDO_MAP_VALUE;
+}
+
static int
bpf_object__init_prog(struct bpf_object *obj, struct bpf_program *prog,
const char *name, size_t sec_idx, const char *sec_name,
@@ -4068,8 +4105,14 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
targ_sec_idx >= obj->efile.sec_cnt)
return -LIBBPF_ERRNO__FORMAT;
- /* Only do relo for section with exec instructions */
+ /*
+ * Only do relo for section with exec instructions,
+ * struct_ops, maps, and read-only data that might
+ * have pointers to functions.
+ */
if (!section_have_execinstr(obj, targ_sec_idx) &&
+ strcmp(name, ".rel" RODATA_SEC) &&
+ !str_has_pfx(name, ".rel" RODATA_SEC ".") &&
strcmp(name, ".rel" STRUCT_OPS_SEC) &&
strcmp(name, ".rel" STRUCT_OPS_LINK_SEC) &&
strcmp(name, ".rel?" STRUCT_OPS_SEC) &&
@@ -6471,6 +6514,137 @@ static int create_jt_map(struct bpf_object *obj, struct bpf_program *prog, struc
return err;
}
+/*
+ * The kernel recognizes a pointer to a function in a frozen read-only map by
+ * its value: the offset in bytes of the function in the program. It makes
+ * callx work for tables of functions, structures of operations and vtables,
+ * where pointers are mixed with other data. Functions have different offsets
+ * in different programs, so create a copy of the map for the program.
+ * The kernel replaces the offsets with the addresses of the functions when it
+ * loads the program, which has to be the only user of the map.
+ */
+static int create_func_ptr_map(struct bpf_object *obj, struct bpf_program *prog, int map_idx)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts, .map_flags = BPF_F_RDONLY_PROG);
+ struct bpf_map *map = &obj->maps[map_idx];
+ __u32 value_size = map->def.value_size;
+ size_t i, j, cnt, sec_insn_off;
+ struct func_ptr *ptrs;
+ int map_fd, err, zero = 0;
+ __u64 val;
+ void *data, *tmp;
+
+ for (i = 0; i < obj->func_ptr_map_cnt; i++)
+ if (obj->func_ptr_maps[i].prog == prog &&
+ obj->func_ptr_maps[i].map_idx == map_idx)
+ return obj->func_ptr_maps[i].fd;
+
+ if (obj->gen_loader) {
+ pr_warn("prog '%s': map '%s': pointers to functions in data are not supported by light skeleton\n",
+ prog->name, map->name);
+ return -ENOTSUP;
+ }
+
+ data = malloc(value_size);
+ if (!data)
+ return -ENOMEM;
+
+ /* the content of the map is final, it's frozen already */
+ if (map->mmaped) {
+ memcpy(data, map->mmaped, value_size);
+ } else if (bpf_map_lookup_elem(map->fd, &zero, data)) {
+ err = -errno;
+ pr_warn("prog '%s': map '%s': failed to read the content: %s\n",
+ prog->name, map->name, errstr(err));
+ goto err_free;
+ }
+
+ ptrs = map->func_ptrs;
+ cnt = map->func_ptr_cnt;
+ for (i = 0; i < cnt; i++) {
+ if (ptrs[i].sec_off + sizeof(val) > value_size) {
+ err = -LIBBPF_ERRNO__FORMAT;
+ goto err_free;
+ }
+ /*
+ * Static functions were appended by bpf_object__append_func_ptrs_code().
+ * A global function is in the program only if the code refers to it.
+ */
+ sec_insn_off = ptrs[i].text_off / BPF_INSN_SZ;
+ for (j = 0; j < prog->subprog_cnt; j++)
+ if (prog->subprogs[j].sec_insn_off == sec_insn_off)
+ break;
+ if (j == prog->subprog_cnt) {
+ pr_debug("prog '%s': map '%s': no function for the pointer at offset %zu, it's NULL\n",
+ prog->name, map->name, ptrs[i].sec_off);
+ val = 0;
+ } else {
+ val = (__u64)prog->subprogs[j].sub_insn_off * BPF_INSN_SZ;
+ }
+ memcpy(data + ptrs[i].sec_off, &val, sizeof(val));
+ }
+
+ /*
+ * The kernel takes any aligned 64-bit value that is equal to the offset
+ * of a function for a pointer. Tell when it's going to get it wrong.
+ */
+ for (i = 0, j = 0; i + sizeof(val) <= value_size; i += sizeof(val)) {
+ __u32 k;
+
+ while (j < cnt && ptrs[j].sec_off < i)
+ j++;
+ if (j < cnt && ptrs[j].sec_off == i)
+ continue;
+ memcpy(&val, data + i, sizeof(val));
+ if (!val || val % BPF_INSN_SZ)
+ continue;
+ for (k = 0; k < prog->subprog_cnt; k++) {
+ if ((__u64)prog->subprogs[k].sub_insn_off * BPF_INSN_SZ != val)
+ continue;
+ pr_warn("prog '%s': map '%s': value %llu at offset %zu is the offset of a function, the kernel will treat it as a pointer to it\n",
+ prog->name, map->name, (unsigned long long)val, i);
+ break;
+ }
+ }
+
+ map_fd = bpf_map_create(BPF_MAP_TYPE_ARRAY, map->name, sizeof(int), value_size, 1, &opts);
+ if (map_fd < 0) {
+ err = map_fd;
+ goto err_free;
+ }
+
+ err = bpf_map_update_elem(map_fd, &zero, data, 0);
+ if (!err)
+ err = bpf_map_freeze(map_fd);
+ if (err) {
+ err = -errno;
+ goto err_close;
+ }
+
+ tmp = libbpf_reallocarray(obj->func_ptr_maps, obj->func_ptr_map_cnt + 1,
+ sizeof(*obj->func_ptr_maps));
+ if (!tmp) {
+ err = -ENOMEM;
+ goto err_close;
+ }
+ obj->func_ptr_maps = tmp;
+ obj->func_ptr_maps[obj->func_ptr_map_cnt].prog = prog;
+ obj->func_ptr_maps[obj->func_ptr_map_cnt].map_idx = map_idx;
+ obj->func_ptr_maps[obj->func_ptr_map_cnt].fd = map_fd;
+ obj->func_ptr_map_cnt++;
+
+ pr_debug("prog '%s': created a copy of map '%s' with %zu pointers to functions\n",
+ prog->name, map->name, cnt);
+ free(data);
+ return map_fd;
+
+err_close:
+ close(map_fd);
+err_free:
+ free(data);
+ return err;
+}
+
/* Relocate data references within program code:
* - map references;
* - global variable references;
@@ -6508,7 +6682,19 @@ bpf_object__relocate_data(struct bpf_object *obj, struct bpf_program *prog)
if (relo->map_idx == obj->arena_map_idx)
insn[1].imm += obj->arena_data_off;
- if (obj->gen_loader) {
+ if (map->autocreate && map->func_ptr_cnt) {
+ int map_fd;
+
+ /* the program gets its own map with pointers to its functions */
+ map_fd = create_func_ptr_map(obj, prog, relo->map_idx);
+ if (map_fd < 0) {
+ pr_warn("prog '%s': relo #%d: can't create a copy of map '%s' with pointers to functions\n",
+ prog->name, i, map->name);
+ return map_fd;
+ }
+ insn[0].src_reg = BPF_PSEUDO_MAP_VALUE;
+ insn[0].imm = map_fd;
+ } else if (obj->gen_loader) {
insn[0].src_reg = BPF_PSEUDO_MAP_IDX_VALUE;
insn[0].imm = relo->map_idx;
} else if (map->autocreate) {
@@ -6835,6 +7021,51 @@ bpf_object__append_subprog_code(struct bpf_object *obj, struct bpf_program *main
return 0;
}
+static int
+bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,
+ struct bpf_program *prog);
+
+/* Append to the main program all functions that the data of the map points to */
+static int
+bpf_object__append_func_ptrs_code(struct bpf_object *obj, struct bpf_program *main_prog,
+ const struct bpf_map *map)
+{
+ struct bpf_program *subprog;
+ size_t i, cnt, sec_insn_off;
+ struct func_ptr *ptrs;
+ int err;
+
+ ptrs = map->func_ptrs;
+ cnt = map->func_ptr_cnt;
+ for (i = 0; i < cnt; i++) {
+ sec_insn_off = ptrs[i].text_off / BPF_INSN_SZ;
+ subprog = find_prog_by_sec_insn(obj, obj->efile.text_shndx, sec_insn_off);
+ if (!subprog || subprog->sec_insn_off != sec_insn_off) {
+ pr_warn("prog '%s': map '%s': no function at .text+%zu for the pointer at offset %zu\n",
+ main_prog->name, map->name, ptrs[i].text_off, ptrs[i].sec_off);
+ return -LIBBPF_ERRNO__RELOC;
+ }
+
+ /*
+ * callx can't call global functions. Don't add one to the
+ * program only because the data points to it.
+ */
+ if (subprog->sym_global)
+ continue;
+
+ /* see the comment in bpf_object__reloc_code() */
+ if (subprog->sub_insn_off == 0) {
+ err = bpf_object__append_subprog_code(obj, main_prog, subprog);
+ if (err)
+ return err;
+ err = bpf_object__reloc_code(obj, main_prog, subprog);
+ if (err)
+ return err;
+ }
+ }
+ return 0;
+}
+
static int
bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,
struct bpf_program *prog)
@@ -6851,6 +7082,20 @@ bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,
for (insn_idx = 0; insn_idx < prog->sec_insn_cnt; insn_idx++) {
insn = &main_prog->insns[prog->sub_insn_off + insn_idx];
+ if (insn_is_func_ptrs_addr(insn)) {
+ /*
+ * The code that loads the address of the data might
+ * call any function that the data points to.
+ */
+ relo = find_prog_insn_relo(prog, insn_idx);
+ if (relo && relo->type == RELO_DATA) {
+ err = bpf_object__append_func_ptrs_code(obj, main_prog,
+ &obj->maps[relo->map_idx]);
+ if (err)
+ return err;
+ }
+ continue;
+ }
if (!insn_is_subprog_call(insn) && !insn_is_pseudo_func(insn))
continue;
@@ -7541,6 +7786,9 @@ static int bpf_object__relocate(struct bpf_object *obj, const char *targ_btf_pat
/* mark the insn, so it's recognized by insn_is_pseudo_func() */
if (relo->type == RELO_SUBPROG_ADDR)
insn[0].src_reg = BPF_PSEUDO_FUNC;
+ /* and by insn_is_func_ptrs_addr() */
+ if (relo->type == RELO_DATA && obj->maps[relo->map_idx].func_ptr_cnt)
+ insn[0].src_reg = BPF_PSEUDO_MAP_VALUE;
}
}
@@ -7757,6 +8005,98 @@ static int bpf_object__collect_map_relos(struct bpf_object *obj,
return 0;
}
+/*
+ * Collect pointers to functions in a read-only data section. They are
+ * R_BPF_64_ABS64 relocations against .text section, where the offset of
+ * a static function in the section is stored in place. Relocations in data
+ * sections were ignored before pointers to functions were supported. Those
+ * that are something else, e.g. pointers to data, still are.
+ */
+static int bpf_object__collect_rodata_relos(struct bpf_object *obj,
+ Elf64_Shdr *shdr, Elf_Data *data)
+{
+ size_t sec_idx = shdr->sh_info, sym_idx;
+ int i, nrels = shdr->sh_size / shdr->sh_entsize;
+ const char *relo_sec_name;
+ struct func_ptr *ptrs;
+ Elf_Data *scn_data;
+ Elf64_Sym *sym;
+ Elf64_Rel *rel;
+ __u64 addend;
+
+ relo_sec_name = elf_sec_str(obj, shdr->sh_name) ?: "<?>";
+ scn_data = obj->efile.secs[sec_idx].data;
+ if (!scn_data)
+ return -LIBBPF_ERRNO__FORMAT;
+
+ for (i = 0; i < nrels; i++) {
+ rel = elf_rel_by_idx(data, i);
+ if (!rel) {
+ pr_warn("sec '%s': failed to get relo #%d\n", relo_sec_name, i);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ sym_idx = ELF64_R_SYM(rel->r_info);
+ sym = elf_sym_by_idx(obj, sym_idx);
+ if (!sym) {
+ pr_warn("sec '%s': symbol #%zu not found for relo #%d\n",
+ relo_sec_name, sym_idx, i);
+ return -LIBBPF_ERRNO__FORMAT;
+ }
+
+ if (ELF64_R_TYPE(rel->r_info) != R_BPF_64_ABS64 ||
+ !sym_is_subprog(sym, obj->efile.text_shndx)) {
+ pr_debug("sec '%s': relo #%d: not a pointer to a function, skipping...\n",
+ relo_sec_name, i);
+ continue;
+ }
+
+ /* the kernel finds aligned pointers only */
+ if (rel->r_offset % sizeof(__u64) || rel->r_offset >= scn_data->d_size ||
+ scn_data->d_size - rel->r_offset < sizeof(__u64)) {
+ pr_debug("sec '%s': relo #%d: unsupported offset 0x%zx, skipping...\n",
+ relo_sec_name, i, (size_t)rel->r_offset);
+ continue;
+ }
+
+ memcpy(&addend, scn_data->d_buf + rel->r_offset, sizeof(addend));
+ if (!is_native_endianness(obj))
+ addend = bswap_64(addend);
+ if ((sym->st_value + addend) % BPF_INSN_SZ) {
+ pr_debug("sec '%s': relo #%d: bad pointer to a function at offset %zu+%llu, skipping...\n",
+ relo_sec_name, i, (size_t)sym->st_value,
+ (unsigned long long)addend);
+ continue;
+ }
+
+ ptrs = libbpf_reallocarray(obj->func_ptrs, obj->func_ptr_cnt + 1, sizeof(*ptrs));
+ if (!ptrs)
+ return -ENOMEM;
+ obj->func_ptrs = ptrs;
+
+ ptrs[obj->func_ptr_cnt].sec_idx = sec_idx;
+ ptrs[obj->func_ptr_cnt].sec_off = rel->r_offset;
+ ptrs[obj->func_ptr_cnt].text_off = sym->st_value + addend;
+ obj->func_ptr_cnt++;
+
+ pr_debug("sec '%s': relo #%d: pointer at offset %zu to a function at .text+%zu\n",
+ relo_sec_name, i, (size_t)rel->r_offset, (size_t)(sym->st_value + addend));
+ }
+ return 0;
+}
+
+static int cmp_func_ptrs(const void *_a, const void *_b)
+{
+ const struct func_ptr *a = _a;
+ const struct func_ptr *b = _b;
+
+ if (a->sec_idx != b->sec_idx)
+ return a->sec_idx < b->sec_idx ? -1 : 1;
+ if (a->sec_off != b->sec_off)
+ return a->sec_off < b->sec_off ? -1 : 1;
+ return 0;
+}
+
static int bpf_object__collect_relos(struct bpf_object *obj)
{
int i, err;
@@ -7779,7 +8119,9 @@ static int bpf_object__collect_relos(struct bpf_object *obj)
return -LIBBPF_ERRNO__INTERNAL;
}
- if (obj->efile.secs[idx].sec_type == SEC_ST_OPS)
+ if (obj->efile.secs[idx].sec_type == SEC_RODATA)
+ err = bpf_object__collect_rodata_relos(obj, shdr, data);
+ else if (obj->efile.secs[idx].sec_type == SEC_ST_OPS)
err = bpf_object__collect_st_ops_relos(obj, shdr, data);
else if (idx == obj->efile.btf_maps_shndx)
err = bpf_object__collect_map_relos(obj, shdr, data);
@@ -7789,6 +8131,25 @@ static int bpf_object__collect_relos(struct bpf_object *obj)
return err;
}
+ /* sort by section, so that pointers in the data of a map are next to each other */
+ if (obj->func_ptr_cnt)
+ qsort(obj->func_ptrs, obj->func_ptr_cnt, sizeof(*obj->func_ptrs), cmp_func_ptrs);
+
+ for (i = 0; i < obj->nr_maps; i++) {
+ struct bpf_map *map = &obj->maps[i];
+ size_t j;
+
+ if (map->libbpf_type != LIBBPF_MAP_RODATA)
+ continue;
+ for (j = 0; j < obj->func_ptr_cnt; j++) {
+ if (obj->func_ptrs[j].sec_idx != map->sec_idx)
+ continue;
+ if (!map->func_ptr_cnt)
+ map->func_ptrs = &obj->func_ptrs[j];
+ map->func_ptr_cnt++;
+ }
+ }
+
bpf_object__sort_relos(obj);
return 0;
}
@@ -9749,6 +10110,11 @@ void bpf_object__close(struct bpf_object *obj)
close(obj->jumptable_maps[i].fd);
zfree(&obj->jumptable_maps);
+ for (i = 0; i < obj->func_ptr_map_cnt; i++)
+ close(obj->func_ptr_maps[i].fd);
+ zfree(&obj->func_ptr_maps);
+ zfree(&obj->func_ptrs);
+
if (obj->btf_module_allowlist) {
for (i = 0; i < obj->btf_module_allowlist_cnt; i++)
zfree(&obj->btf_module_allowlist[i]);
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 12/17] libbpf: Resolve pointers to functions in read-only data
2026-09-22 1:13 ` [PATCH bpf-next 12/17] libbpf: Resolve pointers to functions in read-only data Alexei Starovoitov
@ 2026-09-24 0:33 ` bot+bpf-ci
2026-09-24 2:13 ` Alexei Starovoitov
0 siblings, 1 reply; 44+ messages in thread
From: bot+bpf-ci @ 2026-09-24 0:33 UTC (permalink / raw)
To: alexei.starovoitov, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, ast, andrii,
daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 8537 bytes --]
> diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
> index cd1ea1bb53cb..9538a21f0db4 100644
> --- a/tools/lib/bpf/libbpf.c
> +++ b/tools/lib/bpf/libbpf.c
[ ... ]
> @@ -6471,6 +6514,137 @@ static int create_jt_map(struct bpf_object *obj, struct bpf_program *prog, struc
> return err;
> }
>
> +/*
> + * The kernel recognizes a pointer to a function in a frozen read-only map by
> + * its value: the offset in bytes of the function in the program. It makes
> + * callx work for tables of functions, structures of operations and vtables,
> + * where pointers are mixed with other data. Functions have different offsets
> + * in different programs, so create a copy of the map for the program.
> + * The kernel replaces the offsets with the addresses of the functions when it
> + * loads the program, which has to be the only user of the map.
> + */
> +static int create_func_ptr_map(struct bpf_object *obj, struct bpf_program *prog, int map_idx)
> +{
> + LIBBPF_OPTS(bpf_map_create_opts, opts, .map_flags = BPF_F_RDONLY_PROG);
> + struct bpf_map *map = &obj->maps[map_idx];
> + __u32 value_size = map->def.value_size;
> + size_t i, j, cnt, sec_insn_off;
> + struct func_ptr *ptrs;
> + int map_fd, err, zero = 0;
> + __u64 val;
> + void *data, *tmp;
> +
> + for (i = 0; i < obj->func_ptr_map_cnt; i++)
> + if (obj->func_ptr_maps[i].prog == prog &&
> + obj->func_ptr_maps[i].map_idx == map_idx)
> + return obj->func_ptr_maps[i].fd;
> +
> + if (obj->gen_loader) {
> + pr_warn("prog '%s': map '%s': pointers to functions in data are not supported by light skeleton\n",
> + prog->name, map->name);
> + return -ENOTSUP;
> + }
> +
> + data = malloc(value_size);
> + if (!data)
> + return -ENOMEM;
> +
> + /* the content of the map is final, it's frozen already */
> + if (map->mmaped) {
> + memcpy(data, map->mmaped, value_size);
> + } else if (bpf_map_lookup_elem(map->fd, &zero, data)) {
> + err = -errno;
> + pr_warn("prog '%s': map '%s': failed to read the content: %s\n",
> + prog->name, map->name, errstr(err));
> + goto err_free;
> + }
Is this comment accurate about the map being frozen?
Looking at bpf_object_prepare(), create_func_ptr_map() runs during
bpf_object__relocate():
err = err ? : bpf_object__relocate(obj, ...);
err = err ? : bpf_object__sanitize_and_load_btf(obj);
err = err ? : bpf_object__create_maps(obj);
At this point the .rodata map has not been created, populated or frozen,
and map->fd is still the placeholder fd from create_placeholder_fd().
The data is correct because it comes from map->mmaped (the anonymous mmap
allocated at open time in bpf_object__init_internal_map()), but the
comment's claim and the bpf_map_lookup_elem(map->fd, ...) fallback don't
match the actual call order.
> +
> + ptrs = map->func_ptrs;
> + cnt = map->func_ptr_cnt;
[ ... ]
> +
> + map_fd = bpf_map_create(BPF_MAP_TYPE_ARRAY, map->name, sizeof(int), value_size, 1, &opts);
> + if (map_fd < 0) {
> + err = map_fd;
> + goto err_free;
> + }
Does this need to pass the object's BPF token?
Every other map creation in libbpf sets the token when present, for
example bpf_object__create_map():
create_attr.token_fd = obj->token_fd;
if (obj->token_fd)
create_attr.map_flags |= BPF_F_TOKEN_FD;
Program load does the same (load_attr.token_fd = obj->token_fd). Without
the token, map_create() in kernel/bpf/syscall.c:1493 returns -EPERM when
sysctl_unprivileged_bpf_disabled is set and the caller relies on a
delegated BPF token (user namespace with obj->token_fd set by
bpf_object_prepare_token()).
In that case, bpf_object__relocate_data() prints "can't create a copy of
map ... with pointers to functions" and the whole object fails to load,
even though the kernel side could grant allow_ptr_leaks through the token
(CAP_PERFMON). This affects any program in the object that references
.rodata with function pointers. The existing create_jt_map() has the same
gap.
[ ... ]
> @@ -6835,6 +7021,51 @@ bpf_object__append_subprog_code(struct bpf_object *obj, struct bpf_program *main
> return 0;
> }
>
> +static int
> +bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,
> + struct bpf_program *prog);
> +
> +/* Append to the main program all functions that the data of the map points to */
> +static int
> +bpf_object__append_func_ptrs_code(struct bpf_object *obj, struct bpf_program *main_prog,
> + const struct bpf_map *map)
> +{
> + struct bpf_program *subprog;
> + size_t i, cnt, sec_insn_off;
> + struct func_ptr *ptrs;
> + int err;
> +
> + ptrs = map->func_ptrs;
> + cnt = map->func_ptr_cnt;
> + for (i = 0; i < cnt; i++) {
> + sec_insn_off = ptrs[i].text_off / BPF_INSN_SZ;
> + subprog = find_prog_by_sec_insn(obj, obj->efile.text_shndx, sec_insn_off);
> + if (!subprog || subprog->sec_insn_off != sec_insn_off) {
> + pr_warn("prog '%s': map '%s': no function at .text+%zu for the pointer at offset %zu\n",
> + main_prog->name, map->name, ptrs[i].text_off, ptrs[i].sec_off);
> + return -LIBBPF_ERRNO__RELOC;
> + }
> +
> + /*
> + * callx can't call global functions. Don't add one to the
> + * program only because the data points to it.
> + */
> + if (subprog->sym_global)
> + continue;
> +
> + /* see the comment in bpf_object__reloc_code() */
> + if (subprog->sub_insn_off == 0) {
> + err = bpf_object__append_subprog_code(obj, main_prog, subprog);
> + if (err)
> + return err;
> + err = bpf_object__reloc_code(obj, main_prog, subprog);
> + if (err)
> + return err;
> + }
> + }
> + return 0;
> +}
[ ... ]
> @@ -6851,6 +7082,20 @@ bpf_object__reloc_code(struct bpf_object *obj, struct bpf_program *main_prog,
>
> for (insn_idx = 0; insn_idx < prog->sec_insn_cnt; insn_idx++) {
> insn = &main_prog->insns[prog->sub_insn_off + insn_idx];
> + if (insn_is_func_ptrs_addr(insn)) {
> + /*
> + * The code that loads the address of the data might
> + * call any function that the data points to.
> + */
> + relo = find_prog_insn_relo(prog, insn_idx);
> + if (relo && relo->type == RELO_DATA) {
> + err = bpf_object__append_func_ptrs_code(obj, main_prog,
> + &obj->maps[relo->map_idx]);
> + if (err)
> + return err;
> + }
> + continue;
> + }
> if (!insn_is_subprog_call(insn) && !insn_is_pseudo_func(insn))
> continue;
Does this append functions to programs without callx?
bpf_object__reloc_code() calls bpf_object__append_func_ptrs_code() for
every ld_imm64 of a map with func_ptr_cnt. bpf_object__relocate() marks
any RELO_DATA targeting a map with function pointers:
if (relo->type == RELO_DATA && obj->maps[relo->map_idx].func_ptr_cnt)
insn[0].src_reg = BPF_PSEUDO_MAP_VALUE;
Nothing checks whether the program contains a callx instruction. .rodata
is a single shared map for all const and const volatile globals in the
object, so any ld_imm64 that reads .rodata for a config variable pulls in
all static functions that the section points to.
On the kernel side, resolve_func_ptrs() makes those functions reachable
only when the program has callx:
if (!env->has_callx || !env->allow_ptr_leaks)
return 0;
When the program has no callx, the appended functions are not reached by
any edge, and check_cfg() in kernel/bpf/cfg.c:729 rejects it with
"unreachable insn %d". e7d77667afdf states this: "A function that is
listed in func_info and not referenced by anything is still rejected by
check_cfg() as unreachable."
For example, with tools/testing/selftests/bpf/progs/callx_rodata.c:
const volatile int bias = 1;
static op_fn const ops[] = { add_bias, mul3 };
SEC("socket") int use_table(void *ctx) { ... }
SEC("socket") int read_config(void *ctx) { return bias; }
The ld_imm64 for bias in read_config() makes libbpf append add_bias and
mul3 to it. The kernel then rejects read_config as containing unreachable
instructions, and bpf_object__load() fails for the whole object, including
use_table. Before this patch, .rel.rodata was skipped and read_config
loaded fine. The same failure hits a callx program loaded without
CAP_PERFMON (allow_ptr_leaks false).
Should the functions be appended only when the program (or one of its
subprogs) contains BPF_CALL|BPF_X, or should libbpf probe for kernel
callx support before processing func_ptrs?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35935998342
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 12/17] libbpf: Resolve pointers to functions in read-only data
2026-09-24 0:33 ` bot+bpf-ci
@ 2026-09-24 2:13 ` Alexei Starovoitov
0 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-24 2:13 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, martin.lau,
yonghong.song, mason, ihor.solodrai
On Thu, Sep 24, 2026 at 12:33 AM bot+bpf-ci@kernel.org <bot+bpf-ci@kernel.org> wrote:
> Should the functions be appended only when the program (or one of its
> subprogs) contains BPF_CALL|BPF_X, or should libbpf probe for kernel
> callx support before processing func_ptrs?
yeah. small annoying bug.
Will fix in v2: append the functions and create the copy of the map
only for progs that have callx. The rest keep using .rodata as-is.
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH bpf-next 13/17] libbpf: Treat .data.rel.ro as read-only data
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (11 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 12/17] libbpf: Resolve pointers to functions in read-only data Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-22 1:13 ` [PATCH bpf-next 14/17] libbpf: Support pointers to functions in read-only data in light skeleton Alexei Starovoitov
` (3 subsequent siblings)
16 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
PIC code keeps constants with pointers, e.g. vtables, in .data.rel.ro
instead of .rodata. Nothing writes there after relocation.
libbpf treats every .data* section as writable, while the kernel looks
for pointers to functions in frozen read-only maps only.
Make .data.rel.ro and .data.rel.ro.* maps read-only and frozen like
.rodata and collect pointers to functions there.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
tools/lib/bpf/libbpf.c | 14 ++++++++++++++
1 file changed, 14 insertions(+)
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index 9538a21f0db4..b5c580ec02bc 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -544,6 +544,7 @@ struct bpf_struct_ops {
#define PERCPU_SEC ".percpu"
#define BSS_SEC ".bss"
#define RODATA_SEC ".rodata"
+#define DATA_REL_RO_SEC ".data.rel.ro"
#define KCONFIG_SEC ".kconfig"
#define KSYMS_SEC ".ksyms"
#define STRUCT_OPS_SEC ".struct_ops"
@@ -4061,6 +4062,17 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
err = bpf_object__add_programs(obj, data, name, idx);
if (err)
return err;
+ } else if (strcmp(name, DATA_REL_RO_SEC) == 0 ||
+ str_has_pfx(name, DATA_REL_RO_SEC ".")) {
+ /*
+ * Constants with pointers in them, e.g. vtables,
+ * that position independent code keeps here to
+ * have them relocated. There is nothing that
+ * writes to it after that.
+ */
+ sec_desc->sec_type = SEC_RODATA;
+ sec_desc->shdr = sh;
+ sec_desc->data = data;
} else if (strcmp(name, DATA_SEC) == 0 ||
str_has_pfx(name, DATA_SEC ".")) {
sec_desc->sec_type = SEC_DATA;
@@ -4113,6 +4125,8 @@ static int bpf_object__elf_collect(struct bpf_object *obj)
if (!section_have_execinstr(obj, targ_sec_idx) &&
strcmp(name, ".rel" RODATA_SEC) &&
!str_has_pfx(name, ".rel" RODATA_SEC ".") &&
+ strcmp(name, ".rel" DATA_REL_RO_SEC) &&
+ !str_has_pfx(name, ".rel" DATA_REL_RO_SEC ".") &&
strcmp(name, ".rel" STRUCT_OPS_SEC) &&
strcmp(name, ".rel" STRUCT_OPS_LINK_SEC) &&
strcmp(name, ".rel?" STRUCT_OPS_SEC) &&
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* [PATCH bpf-next 14/17] libbpf: Support pointers to functions in read-only data in light skeleton
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (12 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 13/17] libbpf: Treat .data.rel.ro as " Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-22 2:01 ` bot+bpf-ci
2026-09-22 1:13 ` [PATCH bpf-next 15/17] selftests/bpf: Add tests for callx Alexei Starovoitov
` (2 subsequent siblings)
16 siblings, 1 reply; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
Light skeleton doesn't create maps. The loader prog does at run time.
Teach gen_loader to create a copy of .rodata map with offsets of
functions for a prog, see create_func_ptr_map().
- Such maps are not maps of the object and are not in the loader ctx.
Reserve slots in fd_array after the maps of the object for them and
close their fds when the loader is done or fails. The progs hold them.
- Add bpf_gen__func_ptr_map_create() that emits map create, update,
freeze and returns the index in fd_array for ld_imm64
BPF_PSEUDO_MAP_IDX_VALUE.
- User space can set .rodata before loading the skeleton, see
initial_value in bpf_map_desc. Copy it into the copy of the map as
well and store the offsets of functions on top, since user space
doesn't have them.
- Light skeleton can be generated for a target of another endianness.
Store the offsets in the byte order of the object.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
tools/lib/bpf/bpf_gen_internal.h | 13 ++-
tools/lib/bpf/gen_loader.c | 172 ++++++++++++++++++++++++-------
tools/lib/bpf/libbpf.c | 64 ++++++++++--
3 files changed, 201 insertions(+), 48 deletions(-)
diff --git a/tools/lib/bpf/bpf_gen_internal.h b/tools/lib/bpf/bpf_gen_internal.h
index 6c5ad6c55e8a..206adf28793d 100644
--- a/tools/lib/bpf/bpf_gen_internal.h
+++ b/tools/lib/bpf/bpf_gen_internal.h
@@ -51,9 +51,17 @@ struct bpf_gen {
__u32 nr_ksyms;
int fd_array;
int nr_fd_array;
+ /*
+ * Maps with pointers to functions, that programs get their own copies
+ * of, take slots in fd_array after nr_obj_maps maps of the object.
+ */
+ __u32 nr_obj_maps;
+ __u32 nr_func_ptr_maps;
+ __u32 max_func_ptr_maps;
};
-void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps);
+void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps,
+ int max_func_ptr_maps);
int bpf_gen__finish(struct bpf_gen *gen, int nr_progs, int nr_maps);
void bpf_gen__free(struct bpf_gen *gen);
void bpf_gen__load_btf(struct bpf_gen *gen, const void *raw_data, __u32 raw_size);
@@ -68,6 +76,9 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *value, __u32 value_size,
__u64 flags);
void bpf_gen__map_freeze(struct bpf_gen *gen, int map_idx);
+int bpf_gen__func_ptr_map_create(struct bpf_gen *gen, const char *map_name, int obj_map_idx,
+ void *value, __u32 value_size, const __u32 *ptr_offs,
+ const __u64 *ptr_vals, int ptr_cnt);
void bpf_gen__record_attach_target(struct bpf_gen *gen, const char *name, enum bpf_attach_type type);
void bpf_gen__record_extern(struct bpf_gen *gen, const char *name, bool is_weak,
bool is_typeless, bool is_ld64, int kind, int insn_idx);
diff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c
index af3a04f161ac..251392aa8b41 100644
--- a/tools/lib/bpf/gen_loader.c
+++ b/tools/lib/bpf/gen_loader.c
@@ -112,13 +112,23 @@ static void emit2(struct bpf_gen *gen, struct bpf_insn insn1, struct bpf_insn in
static int add_data(struct bpf_gen *gen, const void *data, __u32 size);
static void emit_sys_close_blob(struct bpf_gen *gen, int blob_off);
-void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps)
+void bpf_gen__init(struct bpf_gen *gen, int log_level, int nr_progs, int nr_maps,
+ int max_func_ptr_maps)
{
size_t stack_sz = sizeof(struct loader_stack), nr_progs_sz;
int i;
gen->fd_array = add_data(gen, NULL, MAX_FD_ARRAY_SZ * sizeof(int));
gen->log_level = log_level;
+ gen->nr_obj_maps = nr_maps;
+ gen->max_func_ptr_maps = max_func_ptr_maps;
+ if (nr_maps + max_func_ptr_maps > MAX_USED_MAPS) {
+ pr_warn("Total maps exceeds %d\n", MAX_USED_MAPS);
+ gen->error = -E2BIG;
+ return;
+ }
+ /* their fds are closed like the fds of the maps of the object when loading fails */
+ nr_maps += max_func_ptr_maps;
/* save ctx pointer into R6 */
emit(gen, BPF_MOV64_REG(BPF_REG_6, BPF_REG_1));
@@ -385,6 +395,9 @@ int bpf_gen__finish(struct bpf_gen *gen, int nr_progs, int nr_maps)
return gen->error;
}
emit_sys_close_stack(gen, stack_off(btf_fd));
+ /* programs hold their maps with pointers to functions, nothing else needs them */
+ for (i = 0; i < gen->nr_func_ptr_maps; i++)
+ emit_sys_close_blob(gen, blob_fd_array_off(gen, gen->nr_obj_maps + i));
for (i = 0; i < gen->nr_progs; i++)
move_stack2ctx(gen,
sizeof(struct bpf_loader_ctx) +
@@ -1127,12 +1140,59 @@ void bpf_gen__prog_load(struct bpf_gen *gen,
gen->nr_progs++;
}
+/*
+ * if (map_desc[map_idx].initial_value) {
+ * if (ctx->flags & BPF_SKEL_KERNEL)
+ * bpf_probe_read_kernel(value, value_size, initial_value);
+ * else
+ * bpf_copy_from_user(value, value_size, initial_value);
+ * nr_more_insns that the caller emits
+ * }
+ */
+static void emit_copy_initial_value(struct bpf_gen *gen, int map_idx, int value,
+ __u32 value_size, int nr_more_insns)
+{
+ emit(gen, BPF_LDX_MEM(BPF_DW, BPF_REG_3, BPF_REG_6,
+ sizeof(struct bpf_loader_ctx) +
+ sizeof(struct bpf_map_desc) * map_idx +
+ offsetof(struct bpf_map_desc, initial_value)));
+ emit(gen, BPF_JMP_IMM(BPF_JEQ, BPF_REG_3, 0, 8 + nr_more_insns));
+ emit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,
+ 0, 0, 0, value));
+ emit(gen, BPF_MOV64_IMM(BPF_REG_2, value_size));
+ emit(gen, BPF_LDX_MEM(BPF_W, BPF_REG_0, BPF_REG_6,
+ offsetof(struct bpf_loader_ctx, flags)));
+ emit(gen, BPF_JMP_IMM(BPF_JSET, BPF_REG_0, BPF_SKEL_KERNEL, 2));
+ emit(gen, BPF_EMIT_CALL(BPF_FUNC_copy_from_user));
+ emit(gen, BPF_JMP_IMM(BPF_JA, 0, 0, 1));
+ emit(gen, BPF_EMIT_CALL(BPF_FUNC_probe_read_kernel));
+}
+
+/* Update the element of the map whose fd is in the slot map_idx of fd_array */
+static void emit_map_update_elem(struct bpf_gen *gen, int map_idx, union bpf_attr *attr,
+ int attr_size, int key, int value, __u32 value_size)
+{
+ int map_update_attr;
+
+ map_update_attr = add_data(gen, attr, attr_size);
+ pr_debug("gen: map_update_elem: idx %d, value: off %d size %u, attr: off %d size %d\n",
+ map_idx, value, value_size, map_update_attr, attr_size);
+ move_blob2blob(gen, attr_field(map_update_attr, map_fd), 4,
+ blob_fd_array_off(gen, map_idx));
+ emit_rel_store(gen, attr_field(map_update_attr, key), key);
+ emit_rel_store(gen, attr_field(map_update_attr, value), value);
+ /* emit MAP_UPDATE_ELEM command */
+ emit_sys_bpf(gen, BPF_MAP_UPDATE_ELEM, map_update_attr, attr_size);
+ debug_ret(gen, "update_elem idx %d value_size %d", map_idx, value_size);
+ emit_check_err(gen);
+}
+
void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *pvalue,
__u32 value_size, __u64 flags)
{
int attr_size = offsetofend(union bpf_attr, flags);
- int map_update_attr, value, key;
union bpf_attr attr;
+ int value, key;
int zero = 0;
memset(&attr, 0, attr_size);
@@ -1142,47 +1202,16 @@ void bpf_gen__map_update_elem(struct bpf_gen *gen, int map_idx, void *pvalue,
key = add_data(gen, &zero, sizeof(zero));
/*
- * if (map_desc[map_idx].initial_value) {
- * if (ctx->flags & BPF_SKEL_KERNEL)
- * bpf_probe_read_kernel(value, value_size, initial_value);
- * else
- * bpf_copy_from_user(value, value_size, initial_value);
- * }
- *
* The runtime initial_value comes from the host-supplied loader
* ctx and would overwrite the blob value that the program signature
* covers and the kernel verifies at load time. For a signed loader
* (gen_hash) the attested blob value must be authoritative, so skip
* the override and leave the signed value in place.
*/
- if (!OPTS_GET(gen->opts, gen_hash, false)) {
- emit(gen, BPF_LDX_MEM(BPF_DW, BPF_REG_3, BPF_REG_6,
- sizeof(struct bpf_loader_ctx) +
- sizeof(struct bpf_map_desc) * map_idx +
- offsetof(struct bpf_map_desc, initial_value)));
- emit(gen, BPF_JMP_IMM(BPF_JEQ, BPF_REG_3, 0, 8));
- emit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,
- 0, 0, 0, value));
- emit(gen, BPF_MOV64_IMM(BPF_REG_2, value_size));
- emit(gen, BPF_LDX_MEM(BPF_W, BPF_REG_0, BPF_REG_6,
- offsetof(struct bpf_loader_ctx, flags)));
- emit(gen, BPF_JMP_IMM(BPF_JSET, BPF_REG_0, BPF_SKEL_KERNEL, 2));
- emit(gen, BPF_EMIT_CALL(BPF_FUNC_copy_from_user));
- emit(gen, BPF_JMP_IMM(BPF_JA, 0, 0, 1));
- emit(gen, BPF_EMIT_CALL(BPF_FUNC_probe_read_kernel));
- }
+ if (!OPTS_GET(gen->opts, gen_hash, false))
+ emit_copy_initial_value(gen, map_idx, value, value_size, 0);
- map_update_attr = add_data(gen, &attr, attr_size);
- pr_debug("gen: map_update_elem: idx %d, value: off %d size %u, attr: off %d size %d\n",
- map_idx, value, value_size, map_update_attr, attr_size);
- move_blob2blob(gen, attr_field(map_update_attr, map_fd), 4,
- blob_fd_array_off(gen, map_idx));
- emit_rel_store(gen, attr_field(map_update_attr, key), key);
- emit_rel_store(gen, attr_field(map_update_attr, value), value);
- /* emit MAP_UPDATE_ELEM command */
- emit_sys_bpf(gen, BPF_MAP_UPDATE_ELEM, map_update_attr, attr_size);
- debug_ret(gen, "update_elem idx %d value_size %d", map_idx, value_size);
- emit_check_err(gen);
+ emit_map_update_elem(gen, map_idx, &attr, attr_size, key, value, value_size);
}
void bpf_gen__populate_outer_map(struct bpf_gen *gen, int outer_map_idx, int slot,
@@ -1214,6 +1243,77 @@ void bpf_gen__populate_outer_map(struct bpf_gen *gen, int outer_map_idx, int slo
emit_check_err(gen);
}
+/*
+ * A copy of the read-only data map obj_map_idx for a program, with the offsets
+ * of its functions in it, see create_func_ptr_map() in libbpf.c. It's not
+ * a map of the object: it's not in the loader ctx. Its content is the content
+ * of obj_map_idx, that the host may supply when the skeleton is loaded, with
+ * ptr_cnt 64-bit ptr_vals at ptr_offs.
+ * Return the index of the map in fd_array for instructions to refer to.
+ */
+int bpf_gen__func_ptr_map_create(struct bpf_gen *gen, const char *map_name, int obj_map_idx,
+ void *pvalue, __u32 value_size, const __u32 *ptr_offs,
+ const __u64 *ptr_vals, int ptr_cnt)
+{
+ int attr_size = offsetofend(union bpf_attr, map_extra);
+ int map_create_attr, map_idx, key, value, zero = 0, i;
+ union bpf_attr attr;
+
+ if (gen->nr_func_ptr_maps == gen->max_func_ptr_maps) {
+ gen->error = -EDOM; /* internal bug */
+ return 0;
+ }
+ map_idx = gen->nr_obj_maps + gen->nr_func_ptr_maps++;
+
+ memset(&attr, 0, attr_size);
+ attr.map_type = tgt_endian(BPF_MAP_TYPE_ARRAY);
+ attr.key_size = tgt_endian((__u32)sizeof(int));
+ attr.value_size = tgt_endian(value_size);
+ attr.max_entries = tgt_endian((__u32)1);
+ attr.map_flags = tgt_endian((__u32)BPF_F_RDONLY_PROG);
+ if (map_name)
+ libbpf_strlcpy(attr.map_name, map_name, sizeof(attr.map_name));
+
+ map_create_attr = add_data(gen, &attr, attr_size);
+ pr_debug("gen: func_ptr_map_create: %s idx %d value_size %u, attr: off %d size %d\n",
+ map_name, map_idx, value_size, map_create_attr, attr_size);
+ emit_sys_bpf(gen, BPF_MAP_CREATE, map_create_attr, attr_size);
+ debug_ret(gen, "func_ptr_map_create %s idx %d value_size %d", map_name, map_idx,
+ value_size);
+ emit_check_err(gen);
+ /* remember map_fd in fd_array */
+ emit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,
+ 0, 0, 0, blob_fd_array_off(gen, map_idx)));
+ emit(gen, BPF_STX_MEM(BPF_W, BPF_REG_1, BPF_REG_7, 0));
+
+ /* pvalue has the pointers already */
+ value = add_data(gen, pvalue, value_size);
+ key = add_data(gen, &zero, sizeof(zero));
+
+ /* see bpf_gen__map_update_elem() */
+ if (!OPTS_GET(gen->opts, gen_hash, false)) {
+ /* the jump over these instructions has 16-bit offset */
+ if (ptr_cnt > 10000) {
+ gen->error = -E2BIG;
+ return 0;
+ }
+ emit_copy_initial_value(gen, obj_map_idx, value, value_size, 3 * ptr_cnt);
+ /* the content that the host supplied doesn't have them */
+ for (i = 0; i < ptr_cnt; i++) {
+ emit2(gen, BPF_LD_IMM64_RAW_FULL(BPF_REG_1, BPF_PSEUDO_MAP_IDX_VALUE,
+ 0, 0, 0, value + ptr_offs[i]));
+ emit(gen, BPF_ST_MEM(BPF_DW, BPF_REG_1, 0, ptr_vals[i]));
+ }
+ }
+
+ attr_size = offsetofend(union bpf_attr, flags);
+ memset(&attr, 0, attr_size);
+ emit_map_update_elem(gen, map_idx, &attr, attr_size, key, value, value_size);
+
+ bpf_gen__map_freeze(gen, map_idx);
+ return map_idx;
+}
+
void bpf_gen__map_freeze(struct bpf_gen *gen, int map_idx)
{
int attr_size = offsetofend(union bpf_attr, map_fd);
diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
index b5c580ec02bc..2e11808f7508 100644
--- a/tools/lib/bpf/libbpf.c
+++ b/tools/lib/bpf/libbpf.c
@@ -6553,19 +6553,19 @@ static int create_func_ptr_map(struct bpf_object *obj, struct bpf_program *prog,
obj->func_ptr_maps[i].map_idx == map_idx)
return obj->func_ptr_maps[i].fd;
- if (obj->gen_loader) {
- pr_warn("prog '%s': map '%s': pointers to functions in data are not supported by light skeleton\n",
- prog->name, map->name);
- return -ENOTSUP;
- }
-
data = malloc(value_size);
if (!data)
return -ENOMEM;
- /* the content of the map is final, it's frozen already */
+ /*
+ * The content of the map is final, it's frozen already. There is no map
+ * when light skeleton is generated, but there is what it's created with.
+ */
if (map->mmaped) {
memcpy(data, map->mmaped, value_size);
+ } else if (obj->gen_loader) {
+ err = -EINVAL;
+ goto err_free;
} else if (bpf_map_lookup_elem(map->fd, &zero, data)) {
err = -errno;
pr_warn("prog '%s': map '%s': failed to read the content: %s\n",
@@ -6595,6 +6595,9 @@ static int create_func_ptr_map(struct bpf_object *obj, struct bpf_program *prog,
} else {
val = (__u64)prog->subprogs[j].sub_insn_off * BPF_INSN_SZ;
}
+ /* light skeleton can be generated for a target of another endianness */
+ if (!is_native_endianness(obj))
+ val = bswap_64(val);
memcpy(data + ptrs[i].sec_off, &val, sizeof(val));
}
@@ -6610,6 +6613,8 @@ static int create_func_ptr_map(struct bpf_object *obj, struct bpf_program *prog,
if (j < cnt && ptrs[j].sec_off == i)
continue;
memcpy(&val, data + i, sizeof(val));
+ if (!is_native_endianness(obj))
+ val = bswap_64(val);
if (!val || val % BPF_INSN_SZ)
continue;
for (k = 0; k < prog->subprog_cnt; k++) {
@@ -6621,6 +6626,30 @@ static int create_func_ptr_map(struct bpf_object *obj, struct bpf_program *prog,
}
}
+ if (obj->gen_loader) {
+ __u32 *ptr_offs = calloc(cnt, sizeof(*ptr_offs));
+ __u64 *ptr_vals = calloc(cnt, sizeof(*ptr_vals));
+
+ if (!ptr_offs || !ptr_vals) {
+ free(ptr_offs);
+ free(ptr_vals);
+ err = -ENOMEM;
+ goto err_free;
+ }
+ for (i = 0; i < cnt; i++) {
+ ptr_offs[i] = ptrs[i].sec_off;
+ memcpy(&ptr_vals[i], data + ptrs[i].sec_off, sizeof(val));
+ if (!is_native_endianness(obj))
+ ptr_vals[i] = bswap_64(ptr_vals[i]);
+ }
+ /* it's an index in fd_array of the loader, not an fd */
+ map_fd = bpf_gen__func_ptr_map_create(obj->gen_loader, map->name, map_idx, data,
+ value_size, ptr_offs, ptr_vals, cnt);
+ free(ptr_offs);
+ free(ptr_vals);
+ goto done;
+ }
+
map_fd = bpf_map_create(BPF_MAP_TYPE_ARRAY, map->name, sizeof(int), value_size, 1, &opts);
if (map_fd < 0) {
err = map_fd;
@@ -6634,6 +6663,7 @@ static int create_func_ptr_map(struct bpf_object *obj, struct bpf_program *prog,
err = -errno;
goto err_close;
}
+done:
tmp = libbpf_reallocarray(obj->func_ptr_maps, obj->func_ptr_map_cnt + 1,
sizeof(*obj->func_ptr_maps));
@@ -6653,7 +6683,8 @@ static int create_func_ptr_map(struct bpf_object *obj, struct bpf_program *prog,
return map_fd;
err_close:
- close(map_fd);
+ if (!obj->gen_loader)
+ close(map_fd);
err_free:
free(data);
return err;
@@ -6706,7 +6737,8 @@ bpf_object__relocate_data(struct bpf_object *obj, struct bpf_program *prog)
prog->name, i, map->name);
return map_fd;
}
- insn[0].src_reg = BPF_PSEUDO_MAP_VALUE;
+ insn[0].src_reg = obj->gen_loader ? BPF_PSEUDO_MAP_IDX_VALUE :
+ BPF_PSEUDO_MAP_VALUE;
insn[0].imm = map_fd;
} else if (obj->gen_loader) {
insn[0].src_reg = BPF_PSEUDO_MAP_IDX_VALUE;
@@ -9583,8 +9615,17 @@ static int bpf_object_load(struct bpf_object *obj, int extra_log_level, const ch
* permit cross-endian creation of "light skeleton".
*/
if (obj->gen_loader) {
+ int nr_func_ptr_maps = 0, nr_progs = 0, i;
+
+ /* every program may get a copy of every map with pointers to functions */
+ for (i = 0; i < obj->nr_maps; i++)
+ if (obj->maps[i].autocreate && obj->maps[i].func_ptr_cnt)
+ nr_func_ptr_maps++;
+ for (i = 0; i < obj->nr_programs; i++)
+ if (obj->programs[i].autoload && !prog_is_subprog(obj, &obj->programs[i]))
+ nr_progs++;
bpf_gen__init(obj->gen_loader, obj->log_level | extra_log_level,
- obj->nr_programs, obj->nr_maps);
+ obj->nr_programs, obj->nr_maps, nr_func_ptr_maps * nr_progs);
} else if (!is_native_endianness(obj)) {
pr_warn("object '%s': loading non-native endianness is unsupported\n", obj->name);
return libbpf_err(-LIBBPF_ERRNO__ENDIAN);
@@ -10125,7 +10166,8 @@ void bpf_object__close(struct bpf_object *obj)
zfree(&obj->jumptable_maps);
for (i = 0; i < obj->func_ptr_map_cnt; i++)
- close(obj->func_ptr_maps[i].fd);
+ if (!obj->gen_loader)
+ close(obj->func_ptr_maps[i].fd);
zfree(&obj->func_ptr_maps);
zfree(&obj->func_ptrs);
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 14/17] libbpf: Support pointers to functions in read-only data in light skeleton
2026-09-22 1:13 ` [PATCH bpf-next 14/17] libbpf: Support pointers to functions in read-only data in light skeleton Alexei Starovoitov
@ 2026-09-22 2:01 ` bot+bpf-ci
2026-09-22 2:55 ` Alexei Starovoitov
0 siblings, 1 reply; 44+ messages in thread
From: bot+bpf-ci @ 2026-09-22 2:01 UTC (permalink / raw)
To: alexei.starovoitov, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, ast, andrii,
daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 4025 bytes --]
> diff --git a/tools/lib/bpf/bpf_gen_internal.h b/tools/lib/bpf/bpf_gen_internal.h
> index 6c5ad6c55e8a..206adf28793d 100644
> --- a/tools/lib/bpf/bpf_gen_internal.h
> +++ b/tools/lib/bpf/bpf_gen_internal.h
[ ... ]
> diff --git a/tools/lib/bpf/gen_loader.c b/tools/lib/bpf/gen_loader.c
> index af3a04f161ac..251392aa8b41 100644
> --- a/tools/lib/bpf/gen_loader.c
> +++ b/tools/lib/bpf/gen_loader.c
[ ... ]
> +int bpf_gen__func_ptr_map_create(struct bpf_gen *gen, const char *map_name, int obj_map_idx,
> + void *pvalue, __u32 value_size, const __u32 *ptr_offs,
> + const __u64 *ptr_vals, int ptr_cnt)
> +{
> + int attr_size = offsetofend(union bpf_attr, map_extra);
> + int map_create_attr, map_idx, key, value, zero = 0, i;
> + union bpf_attr attr;
> +
> + if (gen->nr_func_ptr_maps == gen->max_func_ptr_maps) {
> + gen->error = -EDOM; /* internal bug */
> + return 0;
> + }
> + map_idx = gen->nr_obj_maps + gen->nr_func_ptr_maps++;
Can max_func_ptr_maps end up being too small? bpf_object_load() counts
nr_progs using prog->autoload before calling bpf_gen__init(), but
bpf_object_prepare() can still turn autoload ON after that point.
The sequence is:
bpf_object_load() (libbpf.c)
counts nr_progs from programs[i].autoload
calls bpf_gen__init() with nr_func_ptr_maps * nr_progs
bpf_object_prepare()
bpf_object_adjust_struct_ops_autoload() can set prog->autoload = true
bpf_object__relocate()
bpf_object__relocate_data()
create_func_ptr_map()
bpf_gen__func_ptr_map_create() <-- called for newly enabled program
bpf_object_adjust_struct_ops_autoload() enables prog->autoload if there
are autocreated struct_ops maps that reference the program (libbpf.c around
line 1212). This happens for SEC("?struct_ops/...") programs that start
with autoload = false but get re-enabled during prepare.
If such a program references a read-only map with func_ptr_cnt, it needs
a copy that was never counted in the reservation.
Since the boundary check treats max_func_ptr_maps as a hard invariant and
returns -EDOM on overflow, could a light skeleton for an object with both
a func-ptr map and a SEC("?struct_ops/...") program fail with -EDOM?
[ ... ]
> diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c
> index b5c580ec02bc..2e11808f7508 100644
> --- a/tools/lib/bpf/libbpf.c
> +++ b/tools/lib/bpf/libbpf.c
[ ... ]
> @@ -9583,8 +9615,17 @@ static int bpf_object_load(struct bpf_object *obj, int extra_log_level, const ch
> * permit cross-endian creation of "light skeleton".
> */
> if (obj->gen_loader) {
> + int nr_func_ptr_maps = 0, nr_progs = 0, i;
> +
> + /* every program may get a copy of every map with pointers to functions */
> + for (i = 0; i < obj->nr_maps; i++)
> + if (obj->maps[i].autocreate && obj->maps[i].func_ptr_cnt)
> + nr_func_ptr_maps++;
> + for (i = 0; i < obj->nr_programs; i++)
> + if (obj->programs[i].autoload && !prog_is_subprog(obj, &obj->programs[i]))
> + nr_progs++;
> bpf_gen__init(obj->gen_loader, obj->log_level | extra_log_level,
> - obj->nr_programs, obj->nr_maps);
> + obj->nr_programs, obj->nr_maps, nr_func_ptr_maps * nr_progs);
> } else if (!is_native_endianness(obj)) {
> pr_warn("object '%s': loading non-native endianness is unsupported\n", obj->name);
> return libbpf_err(-LIBBPF_ERRNO__ENDIAN);
This is the same ordering issue from the other side. The nr_progs loop
uses programs[i].autoload at this point in bpf_object_load(), but the
reservation is passed to bpf_gen__init() before bpf_object_prepare()
runs.
Since bpf_object_prepare() is called after this count (at libbpf.c
around line 9632), and it can raise autoload from false to true, could
the reservation be short by one copy per struct_ops program that gets
re-enabled?
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35675730306
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 14/17] libbpf: Support pointers to functions in read-only data in light skeleton
2026-09-22 2:01 ` bot+bpf-ci
@ 2026-09-22 2:55 ` Alexei Starovoitov
0 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 2:55 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, martin.lau,
yonghong.song, mason, ihor.solodrai
On Tue, Sep 22, 2026 at 02:01 AM bot+bpf-ci@kernel.org <bot+bpf-ci@kernel.org> wrote:
>> + if (gen->nr_func_ptr_maps == gen->max_func_ptr_maps) {
>> + gen->error = -EDOM; /* internal bug */
>> + return 0;
>> + }
>> + map_idx = gen->nr_obj_maps + gen->nr_func_ptr_maps++;
>
> Can max_func_ptr_maps end up being too small? bpf_object_load() counts
> nr_progs using prog->autoload before calling bpf_gen__init(), but
> bpf_object_prepare() can still turn autoload ON after that point.
>
> [ ... ]
>
> bpf_object_adjust_struct_ops_autoload() enables prog->autoload if there
> are autocreated struct_ops maps that reference the program (libbpf.c around
> line 1212). This happens for SEC("?struct_ops/...") programs that start
> with autoload = false but get re-enabled during prepare.
Not a bug.
Light skeleton doesn't support struct_ops maps.
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH bpf-next 15/17] selftests/bpf: Add tests for callx
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (13 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 14/17] libbpf: Support pointers to functions in read-only data in light skeleton Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-24 0:33 ` bot+bpf-ci
2026-09-22 1:13 ` [PATCH bpf-next 16/17] selftests/bpf: Add tests for callx through pointers in read-only data Alexei Starovoitov
2026-09-22 1:13 ` [PATCH bpf-next 17/17] bpf, docs: Document callx instruction Alexei Starovoitov
16 siblings, 1 reply; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
Add verifier_callx tests:
- calls via r0, callee saved regs, pointer passed as argument, returned
from subprog, spilled/filled, different callees on different paths.
These are executed to test JIT.
- invalid operands: scalar, other pointers, modified and variable
PTR_TO_FUNC, global functions, reserved fields, JMP32.
- callx under lock.
- bounded and unbounded recursion via callx.
- stack depth of call chains through callx, including address taken
several frames above callx.
- tail calls in the callee and the caller of callx.
- stack liveness: caller slots read by callee stay alive.
- no const propagation across callx.
- precision backtracking through callx.
- function pointers in C.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
.../selftests/bpf/prog_tests/verifier.c | 2 +
.../selftests/bpf/progs/verifier_callx.c | 984 ++++++++++++++++++
2 files changed, 986 insertions(+)
create mode 100644 tools/testing/selftests/bpf/progs/verifier_callx.c
diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index 4f1e1c1cd5ab..dc4ed6d9e5e4 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -27,6 +27,7 @@
#include "verifier_btf_ctx_access.skel.h"
#include "verifier_btf_unreliable_prog.skel.h"
#include "verifier_call_large_imm.skel.h"
+#include "verifier_callx.skel.h"
#include "verifier_cfg.skel.h"
#include "verifier_cgroup_inv_retcode.skel.h"
#include "verifier_cgroup_skb.skel.h"
@@ -196,6 +197,7 @@ void test_verifier_bswap(void) { RUN(verifier_bswap); }
void test_verifier_btf_ctx_access(void) { RUN(verifier_btf_ctx_access); }
void test_verifier_btf_unreliable_prog(void) { RUN(verifier_btf_unreliable_prog); }
void test_verifier_call_large_imm(void) { RUN(verifier_call_large_imm); }
+void test_verifier_callx(void) { RUN(verifier_callx); }
void test_verifier_cfg(void) { RUN(verifier_cfg); }
void test_verifier_cgroup_inv_retcode(void) { RUN(verifier_cgroup_inv_retcode); }
void test_verifier_cgroup_skb(void) { RUN(verifier_cgroup_skb); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_callx.c b/tools/testing/selftests/bpf/progs/verifier_callx.c
new file mode 100644
index 000000000000..ec63d2107e6c
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_callx.c
@@ -0,0 +1,984 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Tests for callx: indirect calls of bpf subprogs */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "../../../include/linux/filter.h"
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)
+
+#define CALLX_INSN(DST, SRC, OFF, IMM) \
+ BPF_RAW_INSN(BPF_JMP | BPF_CALL | BPF_X, DST, SRC, OFF, IMM)
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __type(key, int);
+ __type(value, long long);
+} map_array SEC(".maps");
+
+struct val_with_lock {
+ struct bpf_spin_lock lock;
+ int cnt;
+};
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __type(key, int);
+ __type(value, struct val_with_lock);
+} map_lock SEC(".maps");
+
+struct {
+ __uint(type, BPF_MAP_TYPE_PROG_ARRAY);
+ __uint(max_entries, 1);
+ __uint(key_size, sizeof(int));
+ __uint(value_size, sizeof(int));
+} map_prog SEC(".maps");
+
+__naked __noinline __used
+static unsigned long add1(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "r0 += 1;"
+ "exit;"
+ );
+}
+
+__naked __noinline __used
+static unsigned long add2(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "r0 += 2;"
+ "exit;"
+ );
+}
+
+/* apply(fn, x) { return fn(x); } */
+__naked __noinline __used
+static unsigned long apply(void)
+{
+ asm volatile (
+ "r3 = r1;"
+ "r1 = r2;"
+ "callx r3;"
+ "exit;"
+ );
+}
+
+SEC("socket")
+__success __retval(6)
+__naked void callx_basic(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* callx is printed by the verifier log and xlated dump */
+SEC("socket")
+__success __log_level(2)
+__msg("(8d) callx r2")
+__xlated("callx r2")
+__naked void callx_disasm(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* different callees are called by the same callx on different paths */
+SEC("socket")
+__success __retval(11)
+__naked void callx_two_callees(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r6 = r0;"
+ "r6 &= 1;"
+ "r2 = %[add1] ll;"
+ "if r6 == 0 goto +2;"
+ "r2 = %[add2] ll;"
+ "r1 = 10;"
+ "callx r2;"
+ /* add1(10) - 0 or add2(10) - 1 */
+ "r0 -= r6;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(add1),
+ __imm_addr(add2)
+ : __clobber_all);
+}
+
+/* pointer to a function is passed as an argument */
+SEC("socket")
+__success __retval(45)
+__naked void callx_fn_as_arg(void)
+{
+ asm volatile (
+ "r1 = %[add2] ll;"
+ "r2 = 40;"
+ "call apply;"
+ "r6 = r0;"
+ "r1 = %[add1] ll;"
+ "r2 = 2;"
+ "call apply;"
+ "r0 += r6;"
+ "exit;"
+ :
+ : __imm_addr(add1),
+ __imm_addr(add2)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long get_add1(void)
+{
+ asm volatile (
+ "r0 = %[add1] ll;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* pointer to a function is returned from a subprog and called via r0 */
+SEC("socket")
+__success __retval(2)
+__naked void callx_r0(void)
+{
+ asm volatile (
+ "call get_add1;"
+ "r1 = 1;"
+ "callx r0;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* pointer to a function survives spill/fill */
+SEC("socket")
+__success __retval(9)
+__naked void callx_spill_fill(void)
+{
+ asm volatile (
+ "r2 = %[add2] ll;"
+ "*(u64 *)(r10 - 8) = r2;"
+ "call %[bpf_get_prandom_u32];"
+ "r1 = 7;"
+ "r9 = *(u64 *)(r10 - 8);"
+ "callx r9;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(add2)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long clobber_callee_saved(void)
+{
+ asm volatile (
+ "r6 = 100;"
+ "r7 = 100;"
+ "r8 = 100;"
+ "r9 = 100;"
+ "r0 = r1;"
+ "exit;"
+ );
+}
+
+/* r6-r9 are preserved across callx */
+SEC("socket")
+__success __retval(11)
+__naked void callx_callee_saved_regs(void)
+{
+ asm volatile (
+ "r6 = 1;"
+ "r7 = 2;"
+ "r8 = 3;"
+ "r9 = 4;"
+ "r1 = 1;"
+ "r2 = %[clobber_callee_saved] ll;"
+ "callx r2;"
+ "r0 += r6;"
+ "r0 += r7;"
+ "r0 += r8;"
+ "r0 += r9;"
+ "exit;"
+ :
+ : __imm_addr(clobber_callee_saved)
+ : __clobber_all);
+}
+
+/* r1-r5 are scratched by callx */
+SEC("socket")
+__failure __msg("R1 !read_ok")
+__naked void callx_scratches_args(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "r0 = r1;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* return value of the callee is tracked */
+SEC("socket")
+__success __log_level(2)
+__msg("R0=6")
+__naked void callx_retval_is_tracked(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long write42(void)
+{
+ asm volatile (
+ "r2 = 42;"
+ "*(u64 *)(r1 + 0) = r2;"
+ "r0 = 0;"
+ "exit;"
+ );
+}
+
+/* callee writes into the stack of the caller */
+SEC("socket")
+__success __retval(42)
+__naked void callx_callee_writes_caller_stack(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = %[write42] ll;"
+ "callx r2;"
+ "r0 = *(u64 *)(r10 - 8);"
+ "exit;"
+ :
+ : __imm_addr(write42)
+ : __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("R1 has type scalar, expected func")
+__naked void callx_scalar(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "callx r1;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("R2 !read_ok")
+__naked void callx_uninit_reg(void)
+{
+ asm volatile (
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("R10 has type fp, expected func")
+__naked void callx_fp(void)
+{
+ asm volatile (
+ "callx r10;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("R1 has type map_value, expected func")
+__naked void callx_map_value(void)
+{
+ asm volatile (
+ "r1 = %[map_array] ll;"
+ "r2 = r10;"
+ "r2 += -4;"
+ "r3 = 0;"
+ "*(u32 *)(r2 + 0) = r3;"
+ "call %[bpf_map_lookup_elem];"
+ "if r0 == 0 goto 1f;"
+ "r1 = r0;"
+ "callx r1;"
+ "1:"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_map_lookup_elem),
+ __imm_addr(map_array)
+ : __clobber_all);
+}
+
+/* the address of a subprog can't be modified before the call */
+SEC("socket")
+__failure __msg("dereference of modified func ptr R2 off=8 disallowed")
+__naked void callx_modified_ptr(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ "r2 += 8;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("variable func access var_off=")
+__naked void callx_variable_ptr(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r0 &= 8;"
+ "r2 = %[add1] ll;"
+ "r2 += r0;"
+ "r1 = 5;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(add1)
+ : __clobber_all);
+}
+
+__noinline __used
+int global_add3(int x)
+{
+ return x + 3;
+}
+
+/* only static subprogs can be called via callx */
+SEC("socket")
+__failure __msg("callback function not static")
+__naked void callx_global_func(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[global_add3] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(global_add3)
+ : __clobber_all);
+}
+
+#define DEFINE_CALLX_RESERVED_FIELDS_PROG(NAME, SRC_REG, OFF, IMM) \
+ SEC("socket") \
+ __failure __msg("BPF_CALL|BPF_X uses reserved fields") \
+ __naked void callx_reserved_field_ ## NAME(void) \
+ { \
+ asm volatile ( \
+ "r1 = 5;" \
+ "r2 = %[add1] ll;" \
+ ".8byte %[callx_r2];" \
+ "exit;" \
+ : \
+ : __imm_addr(add1), \
+ __imm_insn(callx_r2, CALLX_INSN(BPF_REG_2, (SRC_REG), (OFF), (IMM))) \
+ : __clobber_all); \
+ }
+
+DEFINE_CALLX_RESERVED_FIELDS_PROG(src_reg, BPF_REG_1, 0, 0)
+DEFINE_CALLX_RESERVED_FIELDS_PROG(off, BPF_REG_0, 1, 0)
+DEFINE_CALLX_RESERVED_FIELDS_PROG(imm, BPF_REG_0, 0, 1)
+
+SEC("socket")
+__failure __msg("unknown opcode 8e")
+__naked void callx_jmp32(void)
+{
+ asm volatile (
+ "r1 = 5;"
+ "r2 = %[add1] ll;"
+ ".8byte %[callx32_r2];"
+ "exit;"
+ :
+ : __imm_addr(add1),
+ __imm_insn(callx32_r2,
+ BPF_RAW_INSN(BPF_JMP32 | BPF_CALL | BPF_X, BPF_REG_2, 0, 0, 0))
+ : __clobber_all);
+}
+
+/* similar to calls of static subprogs callx is allowed under a lock */
+SEC("tc")
+__success __retval(3)
+__naked void callx_under_lock(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "*(u32 *)(r10 - 4) = r1;"
+ "r2 = r10;"
+ "r2 += -4;"
+ "r1 = %[map_lock] ll;"
+ "call %[bpf_map_lookup_elem];"
+ "if r0 != 0 goto 1f;"
+ "exit;"
+ "1:"
+ "r6 = r0;"
+ "r1 = r6;"
+ "call %[bpf_spin_lock];"
+ "r1 = 1;"
+ "r2 = %[add2] ll;"
+ "callx r2;"
+ "r7 = r0;"
+ "r1 = r6;"
+ "call %[bpf_spin_unlock];"
+ "r0 = r7;"
+ "exit;"
+ :
+ : __imm(bpf_map_lookup_elem),
+ __imm(bpf_spin_lock),
+ __imm(bpf_spin_unlock),
+ __imm_addr(map_lock),
+ __imm_addr(add2)
+ : __clobber_all);
+}
+
+/* helpers are still not allowed under a lock in the callee */
+__naked __noinline __used
+static unsigned long call_helper(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+SEC("tc")
+__failure __msg("function calls are not allowed while holding a lock")
+__naked void callx_helper_under_lock(void)
+{
+ asm volatile (
+ "r1 = 0;"
+ "*(u32 *)(r10 - 4) = r1;"
+ "r2 = r10;"
+ "r2 += -4;"
+ "r1 = %[map_lock] ll;"
+ "call %[bpf_map_lookup_elem];"
+ "if r0 != 0 goto 1f;"
+ "exit;"
+ "1:"
+ "r6 = r0;"
+ "r1 = r6;"
+ "call %[bpf_spin_lock];"
+ "r2 = %[call_helper] ll;"
+ "callx r2;"
+ "r1 = r6;"
+ "call %[bpf_spin_unlock];"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_map_lookup_elem),
+ __imm(bpf_spin_lock),
+ __imm(bpf_spin_unlock),
+ __imm_addr(map_lock),
+ __imm_addr(call_helper)
+ : __clobber_all);
+}
+
+/* self(fn) { return fn(fn); } */
+__naked __noinline __used
+static unsigned long self(void)
+{
+ asm volatile (
+ "callx r1;"
+ "exit;"
+ );
+}
+
+/* unbounded recursion is caught by the main verification pass */
+SEC("socket")
+__failure __msg("frames is too deep")
+__naked void callx_unbounded_recursion(void)
+{
+ asm volatile (
+ "r1 = %[self] ll;"
+ "call self;"
+ "exit;"
+ :
+ : __imm_addr(self)
+ : __clobber_all);
+}
+
+/* countdown(fn, n) { return n ? fn(fn, n - 1) : 0; } */
+__naked __noinline __used
+static unsigned long countdown(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "if r2 == 0 goto 1f;"
+ "r2 += -1;"
+ "callx r1;"
+ "1:"
+ "exit;"
+ );
+}
+
+/*
+ * The depth of the recursion is known to the main verification pass,
+ * but recursive calls are not allowed regardless.
+ */
+SEC("socket")
+__failure __msg("recursive call from countdown() to countdown()")
+__naked void callx_bounded_recursion(void)
+{
+ asm volatile (
+ "r1 = %[countdown] ll;"
+ "r2 = 2;"
+ "call countdown;"
+ "exit;"
+ :
+ : __imm_addr(countdown)
+ : __clobber_all);
+}
+
+/* ping(fn1, fn2, n) { return n ? fn2(fn2, fn1, n - 1) : 0; } */
+__naked __noinline __used
+static unsigned long ping(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "if r3 == 0 goto 1f;"
+ "r3 += -1;"
+ "r4 = r1;"
+ "r1 = r2;"
+ "r2 = r4;"
+ "callx r1;"
+ "1:"
+ "exit;"
+ );
+}
+
+__naked __noinline __used
+static unsigned long pong(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "if r3 == 0 goto 1f;"
+ "r3 += -1;"
+ "r4 = r1;"
+ "r1 = r2;"
+ "r2 = r4;"
+ "callx r1;"
+ "1:"
+ "exit;"
+ );
+}
+
+SEC("socket")
+__failure __msg("recursive call from")
+__naked void callx_mutual_recursion(void)
+{
+ asm volatile (
+ "r1 = %[ping] ll;"
+ "r2 = %[pong] ll;"
+ "r3 = 3;"
+ "call ping;"
+ "exit;"
+ :
+ : __imm_addr(ping),
+ __imm_addr(pong)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long use_stack_304(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "exit;"
+ );
+}
+
+/* stack of the callee of callx is accounted */
+SEC("socket")
+__failure __msg("combined stack size of 2 calls is")
+__naked void callx_stack_depth(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "r2 = %[use_stack_304] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(use_stack_304)
+ : __clobber_all);
+}
+
+/* apply_stack_304(fn) { char buf[304]; return fn(); } */
+__naked __noinline __used
+static unsigned long apply_stack_304(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "callx r1;"
+ "exit;"
+ );
+}
+
+/*
+ * The address of use_stack_304() is taken by the main prog that doesn't
+ * use stack, but it is called from apply_stack_304().
+ */
+SEC("socket")
+__failure __msg("combined stack size of 3 calls is")
+__naked void callx_stack_depth_nested(void)
+{
+ asm volatile (
+ "r1 = %[use_stack_304] ll;"
+ "call apply_stack_304;"
+ "exit;"
+ :
+ : __imm_addr(use_stack_304)
+ : __clobber_all);
+}
+
+SEC("socket")
+__success __retval(0)
+__naked void callx_stack_depth_ok(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 200) = r0;"
+ "r2 = %[use_stack_304] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(use_stack_304)
+ : __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long do_tail_call(void)
+{
+ asm volatile (
+ "r2 = %[map_prog] ll;"
+ "r3 = 0;"
+ "call %[bpf_tail_call];"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_tail_call),
+ __imm_addr(map_prog)
+ : __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("tail_calls are not allowed in functions called via callx")
+__naked void callx_tail_call_in_callee(void)
+{
+ asm volatile (
+ "r2 = %[do_tail_call] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm_addr(do_tail_call)
+ : __clobber_all);
+}
+
+/* tail call in the caller of callx is fine */
+SEC("socket")
+__success __retval(3)
+__naked void callx_tail_call_in_caller(void)
+{
+ asm volatile (
+ "r6 = r1;"
+ "r1 = 2;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ "r7 = r0;"
+ "r1 = r6;"
+ "r2 = %[map_prog] ll;"
+ "r3 = 0;"
+ "call %[bpf_tail_call];"
+ "r0 = r7;"
+ "exit;"
+ :
+ : __imm(bpf_tail_call),
+ __imm_addr(map_prog),
+ __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* read_idx(p) { return ((char *)map_value)[*p]; } */
+__naked __noinline __used
+static unsigned long read_idx(void)
+{
+ asm volatile (
+ "r6 = *(u64 *)(r1 + 0);"
+ "r1 = 0;"
+ "*(u32 *)(r10 - 4) = r1;"
+ "r2 = r10;"
+ "r2 += -4;"
+ "r1 = %[map_array] ll;"
+ "call %[bpf_map_lookup_elem];"
+ "if r0 == 0 goto 1f;"
+ "r0 += r6;"
+ "r0 = *(u8 *)(r0 + 0);"
+ "1:"
+ "exit;"
+ :
+ : __imm(bpf_map_lookup_elem),
+ __imm_addr(map_array)
+ : __clobber_all);
+}
+
+/*
+ * Stack slots of the caller that might be read by the callee of callx
+ * have to be considered alive at the checkpoints before callx and
+ * inside of the callee. Otherwise the state with fp-8 == 1000 is pruned
+ * and out of bounds access in read_idx() goes unnoticed.
+ */
+SEC("socket")
+__failure __msg("invalid access to map value, value_size=8 off=1000 size=1")
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void callx_callee_reads_caller_stack(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r1 = 1000;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "if r0 == 0 goto 1f;"
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "1:"
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = %[read_idx] ll;"
+ "callx r2;"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(read_idx)
+ : __clobber_all);
+}
+
+/* in bounds access in read_idx() is fine */
+SEC("socket")
+__success __retval(0)
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void callx_callee_reads_caller_stack_ok(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r1 = 7;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "if r0 == 0 goto 1f;"
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "1:"
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = %[read_idx] ll;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(read_idx)
+ : __clobber_all);
+}
+
+/* same as above, but the pointer to the stack is passed through one more frame */
+SEC("socket")
+__failure __msg("invalid access to map value, value_size=8 off=1000 size=1")
+__flag(BPF_F_TEST_STATE_FREQ)
+__naked void callx_callee_reads_caller_stack_nested(void)
+{
+ asm volatile (
+ "call %[bpf_get_prandom_u32];"
+ "r1 = 1000;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "if r0 == 0 goto 1f;"
+ "r1 = 0;"
+ "*(u64 *)(r10 - 8) = r1;"
+ "1:"
+ "r1 = %[read_idx] ll;"
+ "r2 = r10;"
+ "r2 += -8;"
+ "call apply;"
+ "r0 = 0;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32),
+ __imm_addr(read_idx)
+ : __clobber_all);
+}
+
+/* registers that are constant before callx are not constant after it */
+SEC("socket")
+__success __retval(1)
+__naked void callx_clobbers_const_regs(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "r1 = 0;"
+ "r2 = %[add1] ll;"
+ "callx r2;"
+ /* dead branch pruning must not assume that r0 is still 0 */
+ "if r0 == 0 goto 1f;"
+ "r0 = 1;"
+ "exit;"
+ "1:"
+ "r0 = 2;"
+ "exit;"
+ :
+ : __imm_addr(add1)
+ : __clobber_all);
+}
+
+/* scalar argument passed through callx is tracked precisely */
+__naked __noinline __used
+static unsigned long identity(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "exit;"
+ );
+}
+
+long long vals[] SEC(".data.vals") = {1, 2, 3, 4};
+
+SEC("socket")
+__success __log_level(2)
+__msg("mark_precise: frame0: regs=r0 stack= before 12: (95) exit")
+__msg("mark_precise: frame1: regs=r0 stack= before 11: (bf) r0 = r1")
+__msg("mark_precise: frame1: regs=r1 stack= before 4: (8d) callx r2")
+__msg("mark_precise: frame0: regs=r1 stack= before 3: (bf) r1 = r6")
+__msg("mark_precise: frame0: regs=r6 stack= before 2: (b7) r6 = 3")
+__retval(4)
+__naked void callx_precision(void)
+{
+ asm volatile (
+ "r2 = %[identity] ll;"
+ "r6 = 3;"
+ "r1 = r6;"
+ "callx r2;"
+ "r0 *= 8;"
+ "r1 = %[vals] ll;"
+ "r1 += r0;"
+ "r0 = *(u64 *)(r1 + 0);"
+ "exit;"
+ :
+ : __imm_addr(identity),
+ __imm_addr(vals)
+ : __clobber_all);
+}
+
+/* function pointers in C */
+
+typedef int (*op_fn)(int);
+
+static __noinline int mul3(int x)
+{
+ return x * 3;
+}
+
+static __noinline int sub7(int x)
+{
+ return x - 7;
+}
+
+static __noinline int apply_op(op_fn op, int x)
+{
+ return op(x);
+}
+
+SEC("socket")
+__success __retval(36)
+int callx_c_fn_as_arg(void *ctx)
+{
+ /* (5 * 3) + (28 - 7) */
+ return apply_op(mul3, 5) + apply_op(sub7, 28);
+}
+
+SEC("socket")
+__success __retval(30)
+int callx_c_select(void *ctx)
+{
+ __u32 rnd = bpf_get_prandom_u32() & 1;
+ op_fn op = rnd ? mul3 : sub7;
+ int x = rnd ? 10 : 37;
+
+ /* 10 * 3 or 37 - 7 */
+ return op(x);
+}
+
+struct ops {
+ op_fn first;
+ op_fn second;
+ int bias;
+};
+
+static __noinline int run_ops(const struct ops *ops, int x)
+{
+ return ops->second(ops->first(x)) + ops->bias;
+}
+
+SEC("socket")
+__success __retval(100)
+int callx_c_ops_on_stack(void *ctx)
+{
+ struct ops a, b;
+
+ /* avoid an initializer with function pointers in .rodata */
+ a.first = mul3;
+ a.second = sub7;
+ a.bias = 10;
+ b.first = sub7;
+ b.second = mul3;
+ b.bias = 1;
+
+ /* ((7 * 3) - 7 + 10) + ((32 - 7) * 3 + 1) */
+ return run_ops(&a, 7) + run_ops(&b, 32);
+}
+
+#else
+
+SEC("socket")
+__success
+int dummy(void *ctx)
+{
+ return 0;
+}
+
+#endif
+
+char _license[] SEC("license") = "GPL";
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 15/17] selftests/bpf: Add tests for callx
2026-09-22 1:13 ` [PATCH bpf-next 15/17] selftests/bpf: Add tests for callx Alexei Starovoitov
@ 2026-09-24 0:33 ` bot+bpf-ci
2026-09-24 2:13 ` Alexei Starovoitov
0 siblings, 1 reply; 44+ messages in thread
From: bot+bpf-ci @ 2026-09-24 0:33 UTC (permalink / raw)
To: alexei.starovoitov, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, ast, andrii,
daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 2091 bytes --]
> diff --git a/tools/testing/selftests/bpf/progs/verifier_callx.c b/tools/testing/selftests/bpf/progs/verifier_callx.c
> new file mode 100644
> index 000000000000..ec63d2107e6c
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/verifier_callx.c
[ ... ]
> +/* return value of the callee is tracked */
> +SEC("socket")
> +__success __log_level(2)
> +__msg("R0=6")
> +__naked void callx_retval_is_tracked(void)
> +{
> + asm volatile (
> + "r1 = 5;"
> + "r2 = %[add1] ll;"
> + "callx r2;"
> + "exit;"
> + :
> + : __imm_addr(add1)
> + : __clobber_all);
> +}
The comment says this test checks that the callee's return value is
tracked in the caller, but does __msg("R0=6") actually verify that?
The level 2 log first prints the callee's frame state. Since add1 is
appended at insn 5, the log contains a line from inside the callee:
6: (07) r0 += 1 ; frame1: R0=6
This line comes from print_insn_state() and matches the __msg check, even
though it's in the callee's frame, not the caller's.
The line that should be checked appears later, after prepare_func_exit()
copies R0 back to the caller:
to caller at 4:
R0=6 R10=fp0
from 7 to 4: R0=6 R10=fp0
validate_msgs() takes the first match, so the test passes even if the
caller never sees R0=6. If check_func_callx() or the return path left the
caller's R0 as an unknown scalar, the program would still be accepted (and
__msg("R0=6") would still pass), but the test would fail to catch that
regression. callx_basic already covers the runtime result with __retval(6).
Could the message be anchored to the caller context? For example,
__msg("to caller at 4:") followed by __msg("R0=6 R10=fp0"), or
__msg("from 7 to 4: R0=6") would specifically check that the caller sees
the tracked return value, similar to how verifier_precision.c anchors with
instruction indices.
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35935998342
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 15/17] selftests/bpf: Add tests for callx
2026-09-24 0:33 ` bot+bpf-ci
@ 2026-09-24 2:13 ` Alexei Starovoitov
0 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-24 2:13 UTC (permalink / raw)
To: bot+bpf-ci, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, martin.lau,
yonghong.song, mason, ihor.solodrai
On Thu, Sep 24, 2026 at 12:33 AM bot+bpf-ci@kernel.org <bot+bpf-ci@kernel.org> wrote:
>> +/* return value of the callee is tracked */
>> +SEC("socket")
>> +__success __log_level(2)
>> +__msg("R0=6")
>> +__naked void callx_retval_is_tracked(void)
>
> The level 2 log first prints the callee's frame state. Since add1 is
> appended at insn 5, the log contains a line from inside the callee:
>
> 6: (07) r0 += 1 ; frame1: R0=6
>
> This line comes from print_insn_state() and matches the __msg check, even
> though it's in the callee's frame, not the caller's.
Right. The first R0=6 in the log is from the callee.
Will add __msg("to caller at 4:") in front of it in v2.
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH bpf-next 16/17] selftests/bpf: Add tests for callx through pointers in read-only data
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (14 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 15/17] selftests/bpf: Add tests for callx Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
2026-09-22 2:01 ` bot+bpf-ci
2026-09-24 0:33 ` bot+bpf-ci
2026-09-22 1:13 ` [PATCH bpf-next 17/17] bpf, docs: Document callx instruction Alexei Starovoitov
16 siblings, 2 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
Add verifier_callx_rodata tests where pointers to functions are in
.rodata, libbpf stores offsets there and the kernel recognizes them:
- vtable-like data in .rodata and .data.rel.ro, one of two vtables picked
at run time, const and variable index, array of structs where
the index skips the data, address of an element, table without
a symbol. These are executed.
- data next to a pointer is still a constant.
- all functions the index can select are verified. The rest are not.
- index that may select non-pointer, partial and misaligned reads,
arithmetic on the pointer.
- pointer to a global function is not recognized.
- recursion and stack depth via a table.
- CAP_BPF without CAP_PERFMON.
- C: array of functions, NULL check of an element that
bpf_prune_dead_branches() must not fold, struct ops with data, switch
that may become a table, packed struct with misaligned pointers.
Add callx_func_ptr_map test that doesn't use libbpf logic: map that
another prog uses is not scanned, the prog that failed to load is not
a user, the offset in the map is replaced at load, the data is not, no
other prog including second instance of the same one can use the map
after that.
Add callx_rodata_lskel test: two progs with a table of functions in
.rodata loaded by light skeleton, with .rodata set by user space.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
tools/testing/selftests/bpf/Makefile.skel | 2 +-
.../bpf/prog_tests/callx_func_ptr_map.c | 192 ++++++
.../bpf/prog_tests/callx_rodata_lskel.c | 55 ++
.../selftests/bpf/prog_tests/verifier.c | 2 +
.../selftests/bpf/progs/callx_rodata.c | 43 ++
.../bpf/progs/verifier_callx_rodata.c | 651 ++++++++++++++++++
6 files changed, 944 insertions(+), 1 deletion(-)
create mode 100644 tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c
create mode 100644 tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c
create mode 100644 tools/testing/selftests/bpf/progs/callx_rodata.c
create mode 100644 tools/testing/selftests/bpf/progs/verifier_callx_rodata.c
diff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel
index 580d1d82c186..3d92cdca62ed 100644
--- a/tools/testing/selftests/bpf/Makefile.skel
+++ b/tools/testing/selftests/bpf/Makefile.skel
@@ -33,7 +33,7 @@ LINKED_SKELS := test_static_linked.skel.h linked_funcs.skel.h \
LSKELS := fexit_sleep.c trace_printk.c trace_vprintk.c map_ptr_kern.c \
core_kern.c core_kern_overflow.c test_ringbuf.c \
test_ringbuf_n.c test_ringbuf_map_key.c test_ringbuf_write.c \
- test_ringbuf_overwrite.c
+ test_ringbuf_overwrite.c callx_rodata.c
LSKELS_SIGNED := fentry_test.c fexit_test.c atomics.c
diff --git a/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c b/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c
new file mode 100644
index 000000000000..19f96792df5d
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c
@@ -0,0 +1,192 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * The kernel replaces the offsets of functions in a frozen read-only map with
+ * their addresses when the program is loaded. The program has to be the only
+ * user of the map, and no other program can use the map after that.
+ */
+#include <test_progs.h>
+#include <linux/filter.h>
+#include <bpf/btf.h>
+
+#if defined(__x86_64__) || defined(__aarch64__)
+
+#define CALLEE_INSN 6
+#define DATA 0x1234
+
+/*
+ * main: r2 = &value; r2 = *(u64 *)(r2 + 8); r1 = 10; callx r2; exit
+ * add1: r0 = r1; r0 += 1; exit
+ *
+ * where value is { DATA, offset of add1 in the program }.
+ */
+static const struct bpf_insn callx_insns[] = {
+ BPF_LD_MAP_VALUE(BPF_REG_2, 0, 0),
+ BPF_LDX_MEM(BPF_DW, BPF_REG_2, BPF_REG_2, 8),
+ BPF_MOV64_IMM(BPF_REG_1, 10),
+ BPF_RAW_INSN(BPF_JMP | BPF_CALL | BPF_X, BPF_REG_2, 0, 0, 0),
+ BPF_EXIT_INSN(),
+ BPF_MOV64_REG(BPF_REG_0, BPF_REG_1),
+ BPF_ALU64_IMM(BPF_ADD, BPF_REG_0, 1),
+ BPF_EXIT_INSN(),
+};
+
+/* reads the data of the map */
+static const struct bpf_insn reader_insns[] = {
+ BPF_LD_MAP_VALUE(BPF_REG_2, 0, 0),
+ BPF_LDX_MEM(BPF_DW, BPF_REG_0, BPF_REG_2, 0),
+ BPF_EXIT_INSN(),
+};
+
+/* refers to the map and is rejected: r0 is not set */
+static const struct bpf_insn bad_insns[] = {
+ BPF_LD_MAP_VALUE(BPF_REG_2, 0, 0),
+ BPF_EXIT_INSN(),
+};
+
+static char log_buf[16 * 1024];
+static struct bpf_func_info func_info[2];
+static int btf_fd;
+
+static int create_map(void)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts, .map_flags = BPF_F_RDONLY_PROG);
+ __u64 value[2] = { DATA, CALLEE_INSN * sizeof(struct bpf_insn) };
+ int fd, zero = 0;
+
+ fd = bpf_map_create(BPF_MAP_TYPE_ARRAY, "callx_rodata", sizeof(int), sizeof(value), 1,
+ &opts);
+ if (!ASSERT_OK_FD(fd, "map_create"))
+ return -1;
+ if (!ASSERT_OK(bpf_map_update_elem(fd, &zero, value, 0), "map_update") ||
+ !ASSERT_OK(bpf_map_freeze(fd), "map_freeze")) {
+ close(fd);
+ return -1;
+ }
+ return fd;
+}
+
+static int load(const struct bpf_insn *prog_insns, int cnt, int map_fd)
+{
+ LIBBPF_OPTS(bpf_prog_load_opts, opts,
+ .log_buf = log_buf,
+ .log_size = sizeof(log_buf),
+ .log_level = 1,
+ );
+ struct bpf_insn insns[ARRAY_SIZE(callx_insns)];
+
+ if (prog_insns == callx_insns) {
+ /* add1() is referred to by the data only and is found through func_info */
+ opts.prog_btf_fd = btf_fd;
+ opts.func_info = func_info;
+ opts.func_info_cnt = 2;
+ opts.func_info_rec_size = sizeof(func_info[0]);
+ }
+ memcpy(insns, prog_insns, cnt * sizeof(insns[0]));
+ insns[0].imm = map_fd;
+ log_buf[0] = 0;
+ return bpf_prog_load(BPF_PROG_TYPE_SOCKET_FILTER, "callx_map", "GPL", insns, cnt, &opts);
+}
+
+#define LOAD(insns, map_fd) load(insns, ARRAY_SIZE(insns), map_fd)
+
+static void run(int prog_fd, int expected, const char *name)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, topts);
+ char pkt[64] = {};
+
+ topts.data_in = pkt;
+ topts.data_size_in = sizeof(pkt);
+ if (ASSERT_OK(bpf_prog_test_run_opts(prog_fd, &topts), name))
+ ASSERT_EQ(topts.retval, expected, name);
+}
+
+void test_callx_func_ptr_map(void)
+{
+ int int_id, proto_id, map_fd = -1, prog_fd = -1, reader_fd = -1, fd, zero = 0;
+ struct btf *btf;
+ __u64 value[2];
+
+ btf = btf__new_empty();
+ if (!ASSERT_OK_PTR(btf, "btf_new"))
+ return;
+ int_id = btf__add_int(btf, "int", 4, BTF_INT_SIGNED);
+ proto_id = btf__add_func_proto(btf, int_id);
+ func_info[0].insn_off = 0;
+ func_info[0].type_id = btf__add_func(btf, "main_prog", BTF_FUNC_GLOBAL, proto_id);
+ func_info[1].insn_off = CALLEE_INSN;
+ func_info[1].type_id = btf__add_func(btf, "add1", BTF_FUNC_STATIC, proto_id);
+ if (!ASSERT_GT(func_info[1].type_id, 0, "btf_add_func") ||
+ !ASSERT_OK(btf__load_into_kernel(btf), "btf_load"))
+ goto out;
+ btf_fd = btf__fd(btf);
+
+ /*
+ * Another program relies on what the map has: it's verified with
+ * the data folded into constants. Pointers to functions are not looked
+ * for in such map.
+ */
+ map_fd = create_map();
+ if (map_fd < 0)
+ goto out;
+ reader_fd = LOAD(reader_insns, map_fd);
+ if (!ASSERT_OK_FD(reader_fd, "load_reader"))
+ goto out;
+ fd = LOAD(callx_insns, map_fd);
+ if (!ASSERT_LT(fd, 0, "load_shared"))
+ close(fd);
+ ASSERT_HAS_SUBSTR(log_buf, "unreachable insn 6", "log_shared");
+ run(reader_fd, DATA, "run_reader");
+ close(reader_fd);
+ reader_fd = -1;
+ close(map_fd);
+
+ map_fd = create_map();
+ if (map_fd < 0)
+ goto out;
+
+ /* a program that is rejected is not a user, libbpf loads it again to get the log */
+ fd = LOAD(bad_insns, map_fd);
+ if (!ASSERT_LT(fd, 0, "load_bad"))
+ close(fd);
+
+ prog_fd = LOAD(callx_insns, map_fd);
+ if (!ASSERT_OK_FD(prog_fd, "load_callx")) {
+ printf("%s\n", log_buf);
+ goto out;
+ }
+ run(prog_fd, 11, "run_callx");
+
+ /* the offset of the function is gone from the map, the data is intact */
+ if (ASSERT_OK(bpf_map_lookup_elem(map_fd, &zero, value), "map_lookup")) {
+ ASSERT_EQ(value[0], DATA, "data");
+ ASSERT_NEQ(value[1], CALLEE_INSN * sizeof(struct bpf_insn), "pointer");
+ }
+
+ /* no other program can use the map now, another instance of the same one too */
+ fd = LOAD(reader_insns, map_fd);
+ if (!ASSERT_EQ(fd, -EBUSY, "load_reader_after"))
+ close(fd);
+ ASSERT_HAS_SUBSTR(log_buf, "has addresses of functions of another program", "log_reader");
+ fd = LOAD(callx_insns, map_fd);
+ if (!ASSERT_EQ(fd, -EBUSY, "load_second_instance"))
+ close(fd);
+
+ run(prog_fd, 11, "run_callx_again");
+out:
+ if (reader_fd >= 0)
+ close(reader_fd);
+ if (prog_fd >= 0)
+ close(prog_fd);
+ if (map_fd >= 0)
+ close(map_fd);
+ btf__free(btf);
+}
+
+#else
+
+void test_callx_func_ptr_map(void)
+{
+ test__skip();
+}
+
+#endif
diff --git a/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c b/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c
new file mode 100644
index 000000000000..5c60f352ff44
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c
@@ -0,0 +1,55 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <test_progs.h>
+
+#include "callx_rodata.lskel.h"
+
+#if defined(__x86_64__) || defined(__aarch64__)
+
+static void run(int prog_fd, int expected, const char *name)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, topts);
+ char pkt[64] = {};
+
+ topts.data_in = pkt;
+ topts.data_size_in = sizeof(pkt);
+ if (ASSERT_OK(bpf_prog_test_run_opts(prog_fd, &topts), name))
+ ASSERT_EQ(topts.retval, expected, name);
+}
+
+/*
+ * Every program gets its own copy of .rodata with the offsets of its functions.
+ * The copies have what user space puts into .rodata before the load.
+ */
+void test_callx_rodata_lskel(void)
+{
+ struct callx_rodata_lskel *skel;
+
+ skel = callx_rodata_lskel__open();
+ if (!ASSERT_OK_PTR(skel, "open"))
+ return;
+
+ skel->rodata->bias = 7;
+
+ if (!ASSERT_OK(callx_rodata_lskel__load(skel), "load"))
+ goto out;
+
+ skel->bss->op_idx = 0;
+ run(skel->progs.select_op.prog_fd, 17, "add_bias");
+ skel->bss->op_idx = 1;
+ run(skel->progs.select_op.prog_fd, 30, "mul3");
+ skel->bss->op_idx = 2;
+ run(skel->progs.select_op.prog_fd, -1, "out_of_range");
+ /* mul3(add_bias(4)) */
+ run(skel->progs.both_ops.prog_fd, 33, "both_ops");
+out:
+ callx_rodata_lskel__destroy(skel);
+}
+
+#else
+
+void test_callx_rodata_lskel(void)
+{
+ test__skip();
+}
+
+#endif
diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index dc4ed6d9e5e4..3e07b957ee80 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -28,6 +28,7 @@
#include "verifier_btf_unreliable_prog.skel.h"
#include "verifier_call_large_imm.skel.h"
#include "verifier_callx.skel.h"
+#include "verifier_callx_rodata.skel.h"
#include "verifier_cfg.skel.h"
#include "verifier_cgroup_inv_retcode.skel.h"
#include "verifier_cgroup_skb.skel.h"
@@ -198,6 +199,7 @@ void test_verifier_btf_ctx_access(void) { RUN(verifier_btf_ctx_access); }
void test_verifier_btf_unreliable_prog(void) { RUN(verifier_btf_unreliable_prog); }
void test_verifier_call_large_imm(void) { RUN(verifier_call_large_imm); }
void test_verifier_callx(void) { RUN(verifier_callx); }
+void test_verifier_callx_rodata(void) { RUN(verifier_callx_rodata); }
void test_verifier_cfg(void) { RUN(verifier_cfg); }
void test_verifier_cgroup_inv_retcode(void) { RUN(verifier_cgroup_inv_retcode); }
void test_verifier_cgroup_skb(void) { RUN(verifier_cgroup_skb); }
diff --git a/tools/testing/selftests/bpf/progs/callx_rodata.c b/tools/testing/selftests/bpf/progs/callx_rodata.c
new file mode 100644
index 000000000000..7475c87cdaed
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/callx_rodata.c
@@ -0,0 +1,43 @@
+// SPDX-License-Identifier: GPL-2.0
+/* callx through pointers to functions in .rodata, loaded by light skeleton */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+
+typedef int (*op_fn)(int);
+
+/* set by user space before the programs are loaded, it's in .rodata too */
+const volatile int bias = 1;
+
+int op_idx;
+
+static __noinline int add_bias(int x)
+{
+ return x + bias;
+}
+
+static __noinline int mul3(int x)
+{
+ return x * 3;
+}
+
+static op_fn const ops[] = { add_bias, mul3 };
+
+SEC("socket")
+int select_op(void *ctx)
+{
+ unsigned int i = op_idx;
+
+ if (i >= sizeof(ops) / sizeof(ops[0]))
+ return -1;
+ return ops[i](10);
+}
+
+/* functions have other offsets in this program */
+SEC("socket")
+int both_ops(void *ctx)
+{
+ return ops[1](ops[0](4));
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c b/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c
new file mode 100644
index 000000000000..419a90b1f894
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c
@@ -0,0 +1,651 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Tests for callx through pointers to functions in read-only data */
+
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "../../../include/linux/filter.h"
+
+#if defined(__TARGET_ARCH_x86) || defined(__TARGET_ARCH_arm64)
+
+/*
+ * Read-only data with pointers to functions, where the compiler puts tables
+ * of functions, structures of operations and vtables. libbpf resolves
+ * a pointer to the offset of the function in the program and the kernel
+ * recognizes it by that value.
+ */
+#define DATA(SECTION, NAME, ...) \
+ ".pushsection " SECTION ",@progbits;" \
+ ".balign 8;" \
+ #NAME "_%=:" \
+ __VA_ARGS__ \
+ ".type " #NAME "_%=, @object;" \
+ ".size " #NAME "_%=, .-" #NAME "_%=;" \
+ ".popsection;"
+
+#define RODATA(NAME, ...) DATA(".rodata,\"a\"", NAME, __VA_ARGS__)
+
+#define FUNC_TABLE2(NAME, F0, F1) RODATA(NAME, ".quad " #F0 "; .quad " #F1 ";")
+
+__naked __noinline __used
+static unsigned long add1(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "r0 += 1;"
+ "exit;"
+ );
+}
+
+__naked __noinline __used
+static unsigned long add2(void)
+{
+ asm volatile (
+ "r0 = r1;"
+ "r0 += 2;"
+ "exit;"
+ );
+}
+
+/* the second element of the table is called */
+SEC("socket")
+__success __retval(12)
+__naked void callx_rodata_const_index(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 8);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/*
+ * The program reads the address of a function from the data, so pointers to
+ * functions in data are recognized only for programs that may leak pointers.
+ * Otherwise nothing refers to the functions.
+ */
+SEC("socket")
+__success __retval(12)
+__failure_unpriv __msg_unpriv("unreachable insn")
+__caps_unpriv(CAP_BPF)
+__naked void callx_rodata_needs_perfmon(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 8);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* the address of an element is used instead of the address of the table */
+SEC("socket")
+__success __retval(12)
+__naked void callx_rodata_elem_addr(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= + 8 ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* pointers to functions are mixed with other data like in a vtable */
+SEC("socket")
+__success __retval(23)
+__log_level(2)
+__msg("r1 = *(u64 *)(r6 +0) ; R1=7")
+__msg("r2 = *(u64 *)(r6 +8) ; R2=func()")
+__naked void callx_rodata_mixed_with_data(void)
+{
+ asm volatile (
+ RODATA(vt, ".quad 7; .quad add1; .quad 13; .quad add2;")
+ "r6 = vt_%= ll;"
+ "r1 = *(u64 *)(r6 + 0);"
+ "r2 = *(u64 *)(r6 + 8);"
+ /* add1(7) */
+ "callx r2;"
+ "r7 = r0;"
+ "r1 = *(u64 *)(r6 + 16);"
+ "r2 = *(u64 *)(r6 + 24);"
+ /* add2(13) */
+ "callx r2;"
+ "r0 += r7;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/*
+ * Position independent code keeps constants with pointers in .data.rel.ro.
+ * libbpf treats it as read-only data.
+ */
+SEC("socket")
+__success __retval(23)
+__naked void callx_data_rel_ro(void)
+{
+ asm volatile (
+ DATA(".data.rel.ro,\"aw\"", vt, ".quad 7; .quad add1; .quad 13; .quad add2;")
+ "r6 = vt_%= ll;"
+ "r1 = *(u64 *)(r6 + 0);"
+ "r2 = *(u64 *)(r6 + 8);"
+ "callx r2;"
+ "r7 = r0;"
+ "r1 = *(u64 *)(r6 + 16);"
+ "r2 = *(u64 *)(r6 + 24);"
+ "callx r2;"
+ "r0 += r7;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* an array of structures: the index selects one of the functions, not the data */
+SEC("socket")
+__success __retval(11)
+__naked void callx_rodata_array_of_structs(void)
+{
+ asm volatile (
+ RODATA(arr, ".quad add1; .quad 0x1111; .quad add2; .quad 0x2222;")
+ "call %[bpf_get_prandom_u32];"
+ "r6 = r0;"
+ "r6 &= 1;"
+ "r3 = r6;"
+ "r3 <<= 4;"
+ "r2 = arr_%= ll;"
+ "r2 += r3;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ /* add1(10) - 0 or add2(10) - 1 */
+ "r0 -= r6;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* dynamic dispatch: which vtable is used is not known until run time */
+SEC("socket")
+__success __retval(2)
+__naked void callx_rodata_two_vtables(void)
+{
+ asm volatile (
+ RODATA(vt_a, ".quad 1; .quad add1;")
+ RODATA(vt_b, ".quad 0; .quad add2;")
+ "call %[bpf_get_prandom_u32];"
+ "r6 = vt_a_%= ll;"
+ "r0 &= 1;"
+ "if r0 == 0 goto +2;"
+ "r6 = vt_b_%= ll;"
+ "r1 = *(u64 *)(r6 + 0);"
+ "r2 = *(u64 *)(r6 + 8);"
+ /* add1(1) or add2(0) */
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* both functions are verified, either of them is called */
+SEC("socket")
+__success __retval(11)
+__naked void callx_rodata_var_index(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "call %[bpf_get_prandom_u32];"
+ "r6 = r0;"
+ "r6 &= 1;"
+ "r3 = r6;"
+ "r3 <<= 3;"
+ "r2 = tbl_%= ll;"
+ "r2 += r3;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ /* add1(10) - 0 or add2(10) - 1 */
+ "r0 -= r6;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* a table that has no symbol, the compiler generates such for a switch statement */
+SEC("socket")
+__success __retval(11)
+__naked void callx_rodata_no_symbol(void)
+{
+ asm volatile (
+ ".pushsection .rodata,\"a\",@progbits;"
+ ".balign 8;"
+ ".Lanon_%=:"
+ ".quad add1;"
+ ".quad add2;"
+ ".popsection;"
+ "call %[bpf_get_prandom_u32];"
+ "r6 = r0;"
+ "r6 &= 1;"
+ "r3 = r6;"
+ "r3 <<= 3;"
+ "r2 = .Lanon_%= ll;"
+ "r2 += r3;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ "r0 -= r6;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* the first instruction of the function is removed by the verifier */
+__naked __noinline __used
+static unsigned long nop_add1(void)
+{
+ asm volatile (
+ "goto +0;"
+ "r0 = r1;"
+ "r0 += 1;"
+ "exit;"
+ );
+}
+
+/* the pointer follows the function when instructions are removed */
+SEC("socket")
+__success __retval(11)
+__naked void callx_rodata_func_starts_with_nop(void)
+{
+ asm volatile (
+ RODATA(tbl, ".quad nop_add1; .quad 0;")
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* can't be called with a scalar in r1 */
+__naked __noinline __used
+static unsigned long deref_r1(void)
+{
+ asm volatile (
+ "r0 = *(u64 *)(r1 + 0);"
+ "exit;"
+ );
+}
+
+__naked __noinline __used
+static unsigned long ret0(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "exit;"
+ );
+}
+
+/* every function that might be called is verified */
+SEC("socket")
+__failure __msg("R1 invalid mem access 'scalar'")
+__naked void callx_rodata_all_callees_verified(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, ret0, deref_r1)
+ "call %[bpf_get_prandom_u32];"
+ "r0 &= 1;"
+ "r0 <<= 3;"
+ "r2 = tbl_%= ll;"
+ "r2 += r0;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 0;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* a function that is never called is not verified, it's dead code */
+SEC("socket")
+__success __retval(0)
+__naked void callx_rodata_unused_callee(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, ret0, deref_r1)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 0;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* the index may select what is not a pointer to a function */
+SEC("socket")
+__failure __msg("overlaps with a pointer to a function")
+__naked void callx_rodata_index_beyond_table(void)
+{
+ asm volatile (
+ RODATA(tbl, ".quad add1; .quad add2; .quad 0x1234; .quad 0x5678;")
+ "call %[bpf_get_prandom_u32];"
+ "r0 &= 3;"
+ "r0 <<= 3;"
+ "r2 = tbl_%= ll;"
+ "r2 += r0;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* the address of a function is not known until the program is jitted */
+SEC("socket")
+__failure __msg("read of 4 bytes at offset")
+__msg("overlaps with a pointer to a function")
+__naked void callx_rodata_partial_read(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= ll;"
+ "r0 = *(u32 *)(r2 + 0);"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("read of 8 bytes at offset")
+__msg("overlaps with a pointer to a function")
+__naked void callx_rodata_misaligned_read(void)
+{
+ asm volatile (
+ RODATA(tbl, ".quad add1; .quad add2; .quad 0;")
+ "r2 = tbl_%= ll;"
+ "r0 = *(u64 *)(r2 + 4);"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* the data next to a pointer is still known to the verifier */
+SEC("socket")
+__success __retval(0x1234)
+__log_level(2)
+__msg("R0=4660")
+__naked void callx_rodata_data_is_const(void)
+{
+ asm volatile (
+ RODATA(tbl, ".quad add1; .quad 0x1234;")
+ "r2 = tbl_%= ll;"
+ "r0 = *(u64 *)(r2 + 8);"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* the pointer read from the data can't be modified */
+SEC("socket")
+__failure __msg("dereference of modified func ptr R2 off=8 disallowed")
+__naked void callx_rodata_modified_ptr(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, add2)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "r2 += 8;"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+__noinline __used
+int global_add3(int x)
+{
+ return x + 3;
+}
+
+/* a pointer to a global function is not recognized, it's a number */
+SEC("socket")
+__failure __msg("R2 has type scalar, expected func")
+__naked void callx_rodata_global_func(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, add1, global_add3)
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 8);"
+ "r1 = 10;"
+ "callx r2;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+/* selfcall(n) { return n ? tbl[0](n - 1) : 0; }, where tbl[0] == selfcall */
+__naked __noinline __used
+static unsigned long selfcall(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, selfcall, ret0)
+ "r0 = 0;"
+ "if r1 == 0 goto 1f;"
+ "r1 += -1;"
+ "r2 = tbl_%= ll;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "callx r2;"
+ "1:"
+ "exit;"
+ ::: __clobber_all);
+}
+
+SEC("socket")
+__failure __msg("recursive call from selfcall() to selfcall()")
+__naked void callx_rodata_recursion(void)
+{
+ asm volatile (
+ "r1 = 2;"
+ "call selfcall;"
+ "exit;"
+ ::: __clobber_all);
+}
+
+__naked __noinline __used
+static unsigned long use_stack_304(void)
+{
+ asm volatile (
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "exit;"
+ );
+}
+
+/* stack of all possible callees is accounted */
+SEC("socket")
+__failure __msg("combined stack size of 2 calls is")
+__naked void callx_rodata_stack_depth(void)
+{
+ asm volatile (
+ FUNC_TABLE2(tbl, ret0, use_stack_304)
+ "r0 = 0;"
+ "*(u64 *)(r10 - 304) = r0;"
+ "call %[bpf_get_prandom_u32];"
+ "r0 &= 1;"
+ "r0 <<= 3;"
+ "r2 = tbl_%= ll;"
+ "r2 += r0;"
+ "r2 = *(u64 *)(r2 + 0);"
+ "callx r2;"
+ "exit;"
+ :
+ : __imm(bpf_get_prandom_u32)
+ : __clobber_all);
+}
+
+/* pointers to functions in C */
+
+typedef int (*op_fn)(int);
+
+#define DEFINE_OP(N) static __noinline int op##N(int x) { return x * (N + 2) + N; }
+
+DEFINE_OP(0) DEFINE_OP(1) DEFINE_OP(2) DEFINE_OP(3)
+DEFINE_OP(4) DEFINE_OP(5) DEFINE_OP(6) DEFINE_OP(7)
+DEFINE_OP(8) DEFINE_OP(9) DEFINE_OP(10) DEFINE_OP(11)
+DEFINE_OP(12) DEFINE_OP(13) DEFINE_OP(14) DEFINE_OP(15)
+
+static op_fn const ops[] = {
+ op0, op1, op2, op3, op4, op5, op6, op7,
+ op8, op9, op10, op11, op12, op13, op14, op15,
+};
+
+int op_idx = 11;
+
+SEC("socket")
+__success __retval(50)
+int callx_c_table(void *ctx)
+{
+ unsigned int i = op_idx;
+
+ if (i >= sizeof(ops) / sizeof(ops[0]))
+ return -1;
+ /* op11(3) = 3 * 13 + 11 */
+ return ops[i](3);
+}
+
+SEC("socket")
+__success __retval(50)
+int callx_c_table_null_check(void *ctx)
+{
+ unsigned int i = op_idx;
+ op_fn op;
+
+ if (i >= sizeof(ops) / sizeof(ops[0]))
+ return -1;
+ /* the address of a function is not known until the program is jitted */
+ op = ops[i];
+ if (!op)
+ return -2;
+ return op(3);
+}
+
+/* a structure of operations, where pointers to functions are mixed with data */
+struct shape_ops {
+ int id;
+ op_fn area;
+ long scale;
+ op_fn perimeter;
+};
+
+static const struct shape_ops square_ops = { 1, op1, 10, op2 };
+static const struct shape_ops circle_ops = { 2, op3, 20, op1 };
+
+static __noinline int use_shape(const struct shape_ops *ops, int x)
+{
+ return ops->area(x) * ops->scale + ops->perimeter(ops->id);
+}
+
+SEC("socket")
+__success __retval(433)
+int callx_c_ops_mixed_with_data(void *ctx)
+{
+ /*
+ * square: op1(2) * 10 + op2(1) = 7 * 10 + 6 = 76
+ * circle: op3(3) * 20 + op1(2) = 18 * 20 + 7 = 367
+ * minus 10 when op_idx is not what it is set to
+ */
+ return use_shape(&square_ops, 2) + use_shape(&circle_ops, 3) - (op_idx == 11 ? 10 : 0);
+}
+
+/* the ops are selected at run time */
+SEC("socket")
+__success __retval(367)
+int callx_c_ops_selected(void *ctx)
+{
+ const struct shape_ops *ops = op_idx == 11 ? &circle_ops : &square_ops;
+
+ return use_shape(ops, 3);
+}
+
+/* the compiler might turn the switch into a table that has no symbol */
+static __noinline int call_by_switch(unsigned int idx, int x)
+{
+ op_fn op;
+
+ switch (idx) {
+ case 0:
+ op = op8;
+ break;
+ case 1:
+ op = op9;
+ break;
+ case 2:
+ op = op10;
+ break;
+ case 3:
+ op = op11;
+ break;
+ case 4:
+ op = op12;
+ break;
+ case 5:
+ op = op13;
+ break;
+ case 6:
+ op = op14;
+ break;
+ case 7:
+ op = op15;
+ break;
+ default:
+ return -1;
+ }
+ return op(x);
+}
+
+SEC("socket")
+__success __retval(50)
+int callx_c_switch_table(void *ctx)
+{
+ /* op11(3) = 3 * 13 + 11 */
+ return call_by_switch(op_idx - 8, 3);
+}
+
+/*
+ * Misaligned pointers to functions are ignored, the rest of the data is
+ * accessible as before.
+ */
+static const struct {
+ char tag;
+ op_fn op;
+} __attribute__((packed)) packed_ops[] = {
+ { 5, op0 },
+ { 7, op1 },
+};
+
+SEC("socket")
+__success __retval(7)
+int callx_c_packed_struct(void *ctx)
+{
+ return packed_ops[op_idx & 1].tag;
+}
+
+#else
+
+SEC("socket")
+__success
+int dummy(void *ctx)
+{
+ return 0;
+}
+
+#endif
+
+char _license[] SEC("license") = "GPL";
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 16/17] selftests/bpf: Add tests for callx through pointers in read-only data
2026-09-22 1:13 ` [PATCH bpf-next 16/17] selftests/bpf: Add tests for callx through pointers in read-only data Alexei Starovoitov
@ 2026-09-22 2:01 ` bot+bpf-ci
2026-09-24 0:33 ` bot+bpf-ci
1 sibling, 0 replies; 44+ messages in thread
From: bot+bpf-ci @ 2026-09-22 2:01 UTC (permalink / raw)
To: alexei.starovoitov, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, ast, andrii,
daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 8508 bytes --]
> diff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel
> index 580d1d82c186..3d92cdca62ed 100644
> --- a/tools/testing/selftests/bpf/Makefile.skel
> +++ b/tools/testing/selftests/bpf/Makefile.skel
> @@ -33,7 +33,7 @@ LINKED_SKELS := test_static_linked.skel.h linked_funcs.skel.h \
> LSKELS := fexit_sleep.c trace_printk.c trace_vprintk.c map_ptr_kern.c \
> core_kern.c core_kern_overflow.c test_ringbuf.c \
> test_ringbuf_n.c test_ringbuf_map_key.c test_ringbuf_write.c \
> - test_ringbuf_overwrite.c
> + test_ringbuf_overwrite.c callx_rodata.c
>
> LSKELS_SIGNED := fentry_test.c fexit_test.c atomics.c
>
> diff --git a/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c b/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c
> new file mode 100644
> index 000000000000..19f96792df5d
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/prog_tests/callx_func_ptr_map.c
> @@ -0,0 +1,192 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * The kernel replaces the offsets of functions in a frozen read-only map with
> + * their addresses when the program is loaded. The program has to be the only
> + * user of the map, and no other program can use the map after that.
> + */
> +#include <test_progs.h>
> +#include <linux/filter.h>
> +#include <bpf/btf.h>
> +
> +#if defined(__x86_64__) || defined(__aarch64__)
[ ... ]
> + int_id = btf__add_int(btf, "int", 4, BTF_INT_SIGNED);
> + proto_id = btf__add_func_proto(btf, int_id);
> + func_info[0].insn_off = 0;
> + func_info[0].type_id = btf__add_func(btf, "main_prog", BTF_FUNC_GLOBAL, proto_id);
> + func_info[1].insn_off = CALLEE_INSN;
> + func_info[1].type_id = btf__add_func(btf, "add1", BTF_FUNC_STATIC, proto_id);
> + if (!ASSERT_GT(func_info[1].type_id, 0, "btf_add_func") ||
Can this assertion fail? func_info[1].type_id is __u32 (from struct
bpf_func_info in include/uapi/linux/bpf.h), so when btf__add_func() returns
-EINVAL and that gets stored in the u32 field it becomes 4294967274, which
compares greater than 0. Every error value passes this check; the test
would fail later at btf__load_into_kernel() instead.
> + !ASSERT_OK(btf__load_into_kernel(btf), "btf_load"))
> + goto out;
> + btf_fd = btf__fd(btf);
> +
> + /*
> + * Another program relies on what the map has: it's verified with
> + * the data folded into constants. Pointers to functions are not looked
> + * for in such map.
> + */
> + map_fd = create_map();
> + if (map_fd < 0)
> + goto out;
> + reader_fd = LOAD(reader_insns, map_fd);
> + if (!ASSERT_OK_FD(reader_fd, "load_reader"))
> + goto out;
> + fd = LOAD(callx_insns, map_fd);
> + if (!ASSERT_LT(fd, 0, "load_shared"))
> + close(fd);
> + ASSERT_HAS_SUBSTR(log_buf, "unreachable insn 6", "log_shared");
> + run(reader_fd, DATA, "run_reader");
> + close(reader_fd);
> + reader_fd = -1;
> + close(map_fd);
> +
> + map_fd = create_map();
> + if (map_fd < 0)
> + goto out;
> +
> + /* a program that is rejected is not a user, libbpf loads it again to get the log */
> + fd = LOAD(bad_insns, map_fd);
> + if (!ASSERT_LT(fd, 0, "load_bad"))
> + close(fd);
> +
> + prog_fd = LOAD(callx_insns, map_fd);
> + if (!ASSERT_OK_FD(prog_fd, "load_callx")) {
Does this test properly handle the absence of callx support? The test is
gated only by architecture at line 11, not by run-time capability. When
the kernel returns -EOPNOTSUPP (either because net.core.bpf_jit_enable=0,
which happens on any CONFIG_BPF_JIT=y kernel without
CONFIG_BPF_JIT_DEFAULT_ON/ALWAYS_ON, or because bpf_jit_supports_callx()
is false, which on x86_64 happens on any IBT-capable CFI kernel where
cfi_mode resolves to CFI_FINEIBT), this ASSERT_OK_FD fails and the test
reports a false regression instead of skipping.
check_func_callx() in kernel/bpf/verifier.c explicitly returns -EOPNOTSUPP
in both of these documented capability-absent scenarios. The documentation
added in the same series states "callx requires the BPF JIT".
The infrastructure already exists: is_jit_enabled() combined with
test__skip(), as used by prog_tests/kasan.c. The #else arm at line 187
already calls test__skip() for other architectures.
> + printf("%s\n", log_buf);
> + goto out;
> + }
[ ... ]
> diff --git a/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c b/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c
> new file mode 100644
> index 000000000000..5c60f352ff44
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/prog_tests/callx_rodata_lskel.c
> @@ -0,0 +1,55 @@
> +// SPDX-License-Identifier: GPL-2.0
> +#include <test_progs.h>
> +
> +#include "callx_rodata.lskel.h"
> +
> +#if defined(__x86_64__) || defined(__aarch64__)
[ ... ]
> +void test_callx_rodata_lskel(void)
> +{
> + struct callx_rodata_lskel *skel;
> +
> + skel = callx_rodata_lskel__open();
> + if (!ASSERT_OK_PTR(skel, "open"))
> + return;
> +
> + skel->rodata->bias = 7;
> +
> + if (!ASSERT_OK(callx_rodata_lskel__load(skel), "load"))
Does this test properly handle the absence of callx support? The only gate
is the architecture check at line 6, not a run-time capability check.
progs/callx_rodata.c calls through static op_fn const ops[], so every
program in it needs callx, and check_func_callx() returns -EOPNOTSUPP
either when net.core.bpf_jit_enable=0 or when bpf_jit_supports_callx() is
false (which on x86_64 happens when cfi_mode is CFI_FINEIBT on IBT-capable
CFI kernels).
callx_rodata_lskel__load(skel) then returns that error and ASSERT_OK(...)
reports a FAILURE, where the correct behavior for an absent capability is
test__skip().
The convention is available in the same directory: is_jit_enabled() plus
test__skip(), as used by prog_tests/kasan.c. The #else arm at line 50
already calls test__skip() for other architectures. The documentation added
later in the same series records "callx requires the BPF JIT" as a
limitation.
> + goto out;
[ ... ]
> diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
> index dc4ed6d9e5e4..3e07b957ee80 100644
> --- a/tools/testing/selftests/bpf/prog_tests/verifier.c
> +++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
> @@ -28,6 +28,7 @@
> #include "verifier_btf_unreliable_prog.skel.h"
> #include "verifier_call_large_imm.skel.h"
> #include "verifier_callx.skel.h"
> +#include "verifier_callx_rodata.skel.h"
> #include "verifier_cfg.skel.h"
> #include "verifier_cgroup_inv_retcode.skel.h"
> #include "verifier_cgroup_skb.skel.h"
> @@ -198,6 +199,7 @@ void test_verifier_btf_ctx_access(void) { RUN(verifier_btf_ctx_access); }
> void test_verifier_btf_unreliable_prog(void) { RUN(verifier_btf_unreliable_prog); }
> void test_verifier_call_large_imm(void) { RUN(verifier_call_large_imm); }
> void test_verifier_callx(void) { RUN(verifier_callx); }
> +void test_verifier_callx_rodata(void) { RUN(verifier_callx_rodata); }
> void test_verifier_cfg(void) { RUN(verifier_cfg); }
> void test_verifier_cgroup_inv_retcode(void) { RUN(verifier_cgroup_inv_retcode); }
> void test_verifier_cgroup_skb(void) { RUN(verifier_cgroup_skb); }
[ ... ]
> diff --git a/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c b/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c
> new file mode 100644
> index 000000000000..419a90b1f894
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c
[ ... ]
> +SEC("socket")
> +__success __retval(433)
> +int callx_c_ops_mixed_with_data(void *ctx)
> +{
> + /*
> + * square: op1(2) * 10 + op2(1) = 7 * 10 + 6 = 76
> + * circle: op3(3) * 20 + op1(2) = 18 * 20 + 7 = 367
> + * minus 10 when op_idx is not what it is set to
The comment states the condition inverted with respect to the code. op_idx
is defined as int op_idx = 11 at line 509 and nothing changes it, so
(op_idx == 11 ? 10 : 0) evaluates to 10, and 10 is subtracted exactly when
op_idx IS what it is set to. The expected value confirms it: 76 + 367 =
443, and __retval(433) = 443 - 10. If op_idx were not 11 the result would
be 443, not 433. Should this read "minus 10, since op_idx is 11" or
similar?
> + */
> + return use_shape(&square_ops, 2) + use_shape(&circle_ops, 3) - (op_idx == 11 ? 10 : 0);
> +}
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35675730306
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH bpf-next 16/17] selftests/bpf: Add tests for callx through pointers in read-only data
2026-09-22 1:13 ` [PATCH bpf-next 16/17] selftests/bpf: Add tests for callx through pointers in read-only data Alexei Starovoitov
2026-09-22 2:01 ` bot+bpf-ci
@ 2026-09-24 0:33 ` bot+bpf-ci
1 sibling, 0 replies; 44+ messages in thread
From: bot+bpf-ci @ 2026-09-24 0:33 UTC (permalink / raw)
To: alexei.starovoitov, bpf
Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov, ast, andrii,
daniel, martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 6511 bytes --]
> diff --git a/tools/testing/selftests/bpf/Makefile.skel b/tools/testing/selftests/bpf/Makefile.skel
> index 580d1d82c186..3d92cdca62ed 100644
> --- a/tools/testing/selftests/bpf/Makefile.skel
> +++ b/tools/testing/selftests/bpf/Makefile.skel
[ ... ]
> diff --git a/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c b/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c
> new file mode 100644
> index 000000000000..419a90b1f894
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/verifier_callx_rodata.c
> @@ -0,0 +1,651 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/* Tests for callx through pointers to functions in read-only data */
[ ... ]
> +__noinline __used
> +int global_add3(int x)
> +{
> + return x + 3;
> +}
> +
> +/* a pointer to a global function is not recognized, it's a number */
> +SEC("socket")
> +__failure __msg("R2 has type scalar, expected func")
> +__naked void callx_rodata_global_func(void)
> +{
> + asm volatile (
> + FUNC_TABLE2(tbl, add1, global_add3)
> + "r2 = tbl_%= ll;"
> + "r2 = *(u64 *)(r2 + 8);"
> + "r1 = 10;"
> + "callx r2;"
> + "exit;"
> + ::: __clobber_all);
> +}
Does this test actually verify that the kernel skips global functions?
Looking at bpf_object__append_func_ptrs_code() in libbpf, it explicitly
skips global functions:
if (subprog->sym_global)
continue;
Since global_add3() is not referenced directly anywhere in the program,
libbpf never appends it. create_func_ptr_map() then finds no subprog for
that slot and writes 0 there:
if (j == prog->subprog_cnt) {
pr_debug("... no function for the pointer at offset %zu, it's NULL\n", ...);
val = 0;
}
So the kernel receives { offset(add1), 0 }. In resolve_func_ptrs(), the 0
is dropped by `if (!val || ...) continue;` before the
`bpf_subprog_is_global(env, subprog)` skip is ever reached. A kernel
without that skip would give the same "R2 has type scalar, expected func"
message.
To actually test global function handling, the program would need to also
reference global_add3() directly (e.g., with a `call global_add3`) so that
libbpf includes it and stores its offset in the table.
> +/* the compiler might turn the switch into a table that has no symbol */
> +static __noinline int call_by_switch(unsigned int idx, int x)
> +{
> + op_fn op;
> +
> + switch (idx) {
> + case 0:
> + op = op8;
[ ... ]
> + return op(x);
> +}
> +
> +SEC("socket")
> +__success __retval(50)
> +int callx_c_table_null_check(void *ctx)
> +{
> + unsigned int i = op_idx;
> + op_fn op;
> +
> + if (i >= sizeof(ops) / sizeof(ops[0]))
> + return -1;
> + /* the address of a function is not known until the program is jitted */
> + op = ops[i];
> + if (!op)
> + return -2;
> + return op(3);
> +}
Does the NULL check reach the const_fold logic it claims to test?
The commit message says this test covers a "NULL check of an element that
bpf_prune_dead_branches() must not fold". The comment points at the
const_fold.c check in const_reg_xfer():
/* so are the addresses of functions that the map points to */
bpf_map_range_func_ptrs(env, map, off, size, &cnt) ||
However, the index comes from `int op_idx = 11;`, which is in writable
.data, so const_fold treats `i` as unknown. The element address is built
as `r2 = ops ll; r2 += rX`. That BPF_ADD|BPF_X goes to the
`default: *dst = unknown;` case of const_reg_xfer(). The following BPF_LDX
then leaves at `if (!ci_is_map_value(src)) { *dst = unknown; break; }`
before bpf_map_range_func_ptrs() is ever called.
So `op` is never treated as a constant, and bpf_prune_dead_branches()
never considers the `if (!op)` conditional, whether or not the
function-pointer exclusion exists.
To cover the claim, the test would need a constant index into the table,
so that const_fold reaches the func-ptr read, followed by a conditional
jump on the loaded value.
> +/* a structure of operations, where pointers to functions are mixed with data */
> +struct shape_ops {
> + int id;
> + op_fn area;
> + long scale;
> + op_fn perimeter;
> +};
> +
> +static const struct shape_ops square_ops = { 1, op1, 10, op2 };
> +static const struct shape_ops circle_ops = { 2, op3, 20, op1 };
> +
> +static __noinline int use_shape(const struct shape_ops *ops, int x)
> +{
> + return ops->area(x) * ops->scale + ops->perimeter(ops->id);
> +}
> +
> +SEC("socket")
> +__success __retval(433)
> +int callx_c_ops_mixed_with_data(void *ctx)
> +{
> + /*
> + * square: op1(2) * 10 + op2(1) = 7 * 10 + 6 = 76
> + * circle: op3(3) * 20 + op1(2) = 18 * 20 + 7 = 367
> + * minus 10 when op_idx is not what it is set to
> + */
> + return use_shape(&square_ops, 2) + use_shape(&circle_ops, 3) - (op_idx == 11 ? 10 : 0);
> +}
[ ... ]
> +SEC("socket")
> +__success __retval(50)
> +int callx_c_switch_table(void *ctx)
> +{
> + /* op11(3) = 3 * 13 + 11 */
> + return call_by_switch(op_idx - 8, 3);
> +}
A subsystem pattern flags this as potentially concerning: callx_c_switch_table()
compiles to callx, but is annotated only `__success __retval(50)`, with no
`__load_if_JITed()`.
check_callx() in kernel/bpf/verifier.c rejects callx with -EOPNOTSUPP when
the program is not going to be JITed:
/* there is no support for callx in the interpreter */
if (!env->prog->jit_requested) {
verbose(env, "JIT is required to use callx\n");
return -EOPNOTSUPP;
}
With the JIT off (a kernel without CONFIG_BPF_JIT_ALWAYS_ON and
net.core.bpf_jit_enable=0), run_subtest() in test_loader.c computes
current_runtime = NO_JITED. The default load_mask is JITED | NO_JITED, so
the load-mode skip does not apply. The load goes ahead, bpf_object__load()
returns -EOPNOTSUPP, and ASSERT_OK(err, "unexpected_load_failure") reports
a failure.
The other callx-based tests in the file (callx_c_table,
callx_c_table_null_check, callx_c_ops_mixed_with_data, callx_c_ops_selected
and the asm callx_rodata_* success tests) have the same pattern. The file
already has the mechanism: `#define __load_if_JITed()
__test_tag("load_mode=jited")` in bpf_misc.h, which about 30 other verifier
progs use.
This is unlikely to show up in BPF CI (tools/testing/selftests/bpf/config.x86_64
and config.aarch64 both set CONFIG_BPF_JIT_ALWAYS_ON=y), but the test
should skip when JIT support is unavailable.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35935998342
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH bpf-next 17/17] bpf, docs: Document callx instruction
2026-09-22 1:13 [PATCH bpf-next 00/17] bpf: Indirect calls of bpf subprogs (callx) Alexei Starovoitov
` (15 preceding siblings ...)
2026-09-22 1:13 ` [PATCH bpf-next 16/17] selftests/bpf: Add tests for callx through pointers in read-only data Alexei Starovoitov
@ 2026-09-22 1:13 ` Alexei Starovoitov
16 siblings, 0 replies; 44+ messages in thread
From: Alexei Starovoitov @ 2026-09-22 1:13 UTC (permalink / raw)
To: bpf; +Cc: daniel, andrii, eddyz87, memxor, a.s.protopopov
From: Alexei Starovoitov <ast@kernel.org>
Document callx: what it does, where the callee address comes from and
what is not supported.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
Documentation/bpf/clang-notes.rst | 7 +++++--
Documentation/bpf/linux-notes.rst | 31 +++++++++++++++++++++++++++----
2 files changed, 32 insertions(+), 6 deletions(-)
diff --git a/Documentation/bpf/clang-notes.rst b/Documentation/bpf/clang-notes.rst
index 2c872a1ee08e..3ccc7b09d19e 100644
--- a/Documentation/bpf/clang-notes.rst
+++ b/Documentation/bpf/clang-notes.rst
@@ -23,8 +23,11 @@ For CPU versions prior to 3, Clang v7.0 and later can enable ``BPF_ALU`` support
Jump instructions
=================
-If ``-O0`` is used, Clang will generate the ``BPF_CALL | BPF_X | BPF_JMP`` (0x8d)
-instruction, which is not supported by the Linux kernel verifier.
+Clang generates the ``BPF_CALL | BPF_X | BPF_JMP`` (0x8d) instruction for calls
+through a function pointer. The Linux kernel verifier accepts it only when it
+can prove that the register holds the address of a static BPF function, see
+Documentation/bpf/linux-notes.rst. If ``-O0`` is used, Clang will generate this
+instruction for helper calls as well, which is not supported.
Atomic operations
=================
diff --git a/Documentation/bpf/linux-notes.rst b/Documentation/bpf/linux-notes.rst
index 00d2693de025..6c036b54a29f 100644
--- a/Documentation/bpf/linux-notes.rst
+++ b/Documentation/bpf/linux-notes.rst
@@ -15,10 +15,33 @@ Byte swap instructions
Jump instructions
=================
-``BPF_CALL | BPF_X | BPF_JMP`` (0x8d), where the helper function
-integer would be read from a specified register, is not currently supported
-by the verifier. Any programs with this instruction will fail to load
-until such support is added.
+``BPF_CALL | BPF_X | BPF_JMP`` (0x8d), ``callx dst``, performs an indirect
+call of a BPF function whose address is held in the ``dst`` register. The
+``src``, ``offset`` and ``imm`` fields are reserved and must be zero.
+
+The address of a BPF function gets into a register in one of two ways:
+
+* it is loaded by a 64-bit immediate instruction with ``src`` =
+ ``BPF_PSEUDO_FUNC``;
+* it is read, with a 64-bit load, from a frozen read-only array map, that no
+ other program uses, that holds its read-only data: tables of functions,
+ structures of operations, vtables, where pointers to functions may be mixed
+ with other data. In the map a pointer to a function is the offset in bytes
+ of its first instruction in the program, and that is how the verifier
+ recognizes it. It is replaced with the address of the function when
+ the program is loaded. The program reads it from there, which requires
+ ``CAP_PERFMON``.
+
+In both cases only static functions can be referenced. Therefore all functions
+that can be called indirectly are known to the verifier before it starts to
+analyze the program, and ``callx`` is verified as a direct call of every
+function that ``dst`` may point to at that instruction. The same rules apply:
+the calls can not be recursive, and the depth of the call chain and its
+combined stack size are limited.
+
+Calling helper or kernel functions through a register, indirect calls of global
+functions, and tail calls in functions that are called via ``callx`` are not
+supported. ``callx`` requires the BPF JIT.
Maps
====
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread