* [PATCH bpf-next v5 01/11] selftests/bpf: Allow privileged preparation for capability tests
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
@ 2026-09-21 2:38 ` Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 02/11] bpf: Record raw memory arguments during argument checking Kumar Kartikeya Dwivedi
` (10 subsequent siblings)
11 siblings, 0 replies; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-21 2:38 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, Emil Tsalapatis, Tejun Heo, Amery Hung, kkd,
kernel-team
The annotation-driven loader drops capabilities before libbpf prepares an
object. Resolving bpf_testmod kfuncs requires CAP_SYS_ADMIN to enumerate and
open module BTF, so tests without that capability fail before reaching the
verifier.
Add an opt-in __prepare_priv annotation. Call bpf_object__prepare() with the
fixture's initial capabilities, then apply __caps_unpriv before loading the
programs. This uses libbpf's explicit prepare/load boundary. In particular,
CAP_SYS_ADMIN must be dropped along with CAP_PERFMON to test uninitialized
stack checks, since CAP_SYS_ADMIN satisfies the verifier's CAP_PERFMON check.
Preparation also creates maps and loads BTF. Keep it opt-in so existing tests
continue checking those operations with reduced capabilities. The existing
pre-execution callback runs after program loading and is too late for this.
Allow tests retaining CAP_BPF to run when the unprivileged-BPF sysctl is set.
Check CPU mitigations separately: disabled or undetectable mitigations must
still skip these tests, because CAP_BPF does not restore speculative
execution checks.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
tools/testing/selftests/bpf/progs/bpf_misc.h | 9 +++-
tools/testing/selftests/bpf/test_loader.c | 48 ++++++++++++++------
tools/testing/selftests/bpf/unpriv_helpers.c | 16 +++++--
tools/testing/selftests/bpf/unpriv_helpers.h | 2 +
4 files changed, 55 insertions(+), 20 deletions(-)
diff --git a/tools/testing/selftests/bpf/progs/bpf_misc.h b/tools/testing/selftests/bpf/progs/bpf_misc.h
index eb88d9ce6c34..2ced1d751ace 100644
--- a/tools/testing/selftests/bpf/progs/bpf_misc.h
+++ b/tools/testing/selftests/bpf/progs/bpf_misc.h
@@ -9,7 +9,8 @@
#define QUOTE(str) #str
#define EXPAND_QUOTE(str) QUOTE(str)
-/* This set of attributes controls behavior of the
+/*
+ * This set of attributes controls behavior of the
* test_loader.c:test_loader__run_subtests().
*
* The test_loader sequentially loads each program in a skeleton.
@@ -131,6 +132,11 @@
* Several __arch_* annotations could be specified at once.
* When test case is not run on current arch it is marked as skipped.
* __caps_unpriv Specify the capabilities that should be set when running the test.
+ * __prepare_priv In unprivileged mode, prepare the object with the fixture's
+ * initial capabilities before dropping them for program loading.
+ * Preparation includes map creation and BTF/kfunc resolution;
+ * these operations are not tested at the reduced capabilities.
+ * Program loading uses the normal __caps_unpriv selection.
*
* __linear_size Specify the size of the linear area of non-linear skbs, or
* 0 for linear skbs.
@@ -166,6 +172,7 @@
#define __arch_s390x __arch("s390x")
#define __arch_loongarch __arch("LOONGARCH")
#define __caps_unpriv(caps) __test_tag("test_caps_unpriv=" EXPAND_QUOTE(caps))
+#define __prepare_priv __test_tag("test_prepare_priv")
#define __load_if_JITed() __test_tag("load_mode=jited")
#define __load_if_no_JITed() __test_tag("load_mode=no_jited")
#define __stderr(msg) __test_tag("test_expect_stderr=" msg)
diff --git a/tools/testing/selftests/bpf/test_loader.c b/tools/testing/selftests/bpf/test_loader.c
index 28724de06322..a6e3fcc1079c 100644
--- a/tools/testing/selftests/bpf/test_loader.c
+++ b/tools/testing/selftests/bpf/test_loader.c
@@ -33,6 +33,7 @@ static inline const char *str_has_pfx(const char *str, const char *pfx)
#endif
static int sysctl_unpriv_disabled = -1;
+static int unpriv_mitigations_disabled = -1;
enum mode {
PRIV = 1,
@@ -71,6 +72,7 @@ struct test_spec {
int load_mask;
int linear_sz;
const char *skip_reason;
+ bool prepare_priv;
bool auxiliary;
bool valid;
};
@@ -606,6 +608,8 @@ static int parse_test_spec(struct test_loader *tester,
if (err)
goto cleanup;
spec->mode_mask |= UNPRIV;
+ } else if (strcmp(s, "test_prepare_priv") == 0) {
+ spec->prepare_priv = true;
} else if ((val = str_has_pfx(s, "load_mode="))) {
if (strcmp(val, "jited") == 0) {
load_mask = JITED;
@@ -1015,10 +1019,10 @@ struct cap_state {
bool initialized;
};
-static int drop_capabilities(struct cap_state *caps)
+static int drop_capabilities(struct cap_state *caps, __u64 keep_caps)
{
const __u64 caps_to_drop = (1ULL << CAP_SYS_ADMIN | 1ULL << CAP_NET_ADMIN |
- 1ULL << CAP_PERFMON | 1ULL << CAP_BPF);
+ 1ULL << CAP_PERFMON | 1ULL << CAP_BPF) & ~keep_caps;
int err;
err = cap_disable_effective(caps_to_drop, &caps->old_caps);
@@ -1028,6 +1032,13 @@ static int drop_capabilities(struct cap_state *caps)
}
caps->initialized = true;
+ if (keep_caps) {
+ err = cap_enable_effective(keep_caps, NULL);
+ if (err) {
+ PRINT_FAIL("failed to set capabilities: %i, %s\n", err, strerror(-err));
+ return err;
+ }
+ }
return 0;
}
@@ -1048,8 +1059,12 @@ static int restore_capabilities(struct cap_state *caps)
static bool can_execute_unpriv(struct test_loader *tester, struct test_spec *spec)
{
if (sysctl_unpriv_disabled < 0)
- sysctl_unpriv_disabled = get_unpriv_disabled() ? 1 : 0;
- if (sysctl_unpriv_disabled)
+ sysctl_unpriv_disabled = get_unpriv_sysctl_disabled();
+ if (sysctl_unpriv_disabled && !(spec->unpriv.caps & (1ULL << CAP_BPF)))
+ return false;
+ if (unpriv_mitigations_disabled < 0)
+ unpriv_mitigations_disabled = get_unpriv_mitigations_disabled();
+ if (unpriv_mitigations_disabled)
return false;
if ((spec->prog_flags & BPF_F_ANY_ALIGNMENT) && !EFFICIENT_UNALIGNED_ACCESS)
return false;
@@ -1351,17 +1366,8 @@ void run_subtest(struct test_loader *tester,
test__end_subtest();
return;
}
- if (drop_capabilities(&caps)) {
- test__end_subtest();
- return;
- }
- if (subspec->caps) {
- err = cap_enable_effective(subspec->caps, NULL);
- if (err) {
- PRINT_FAIL("failed to set capabilities: %i, %s\n", err, strerror(-err));
- goto subtest_cleanup;
- }
- }
+ if (!spec->prepare_priv && drop_capabilities(&caps, subspec->caps))
+ goto subtest_cleanup;
}
/* Implicitly reset to NULL if next test case doesn't specify.
@@ -1414,6 +1420,18 @@ void run_subtest(struct test_loader *tester,
bpf_object__for_each_map(map, tobj)
bpf_map__set_autocreate(map, !unpriv || is_unpriv_capable_map(map));
+ if (unpriv && spec->prepare_priv) {
+ /*
+ * Module BTF lookup needs CAP_SYS_ADMIN. Allow tests to prepare
+ * their objects first, then verify programs with the requested caps.
+ */
+ err = bpf_object__prepare(tobj);
+ if (!ASSERT_OK(err, "obj_prepare"))
+ goto tobj_cleanup;
+ if (drop_capabilities(&caps, subspec->caps))
+ goto tobj_cleanup;
+ }
+
err = bpf_object__load(tobj);
if (subspec->expect_failure) {
if (!ASSERT_ERR(err, "unexpected_load_success")) {
diff --git a/tools/testing/selftests/bpf/unpriv_helpers.c b/tools/testing/selftests/bpf/unpriv_helpers.c
index 2c8c5edb8751..c8dd5d848584 100644
--- a/tools/testing/selftests/bpf/unpriv_helpers.c
+++ b/tools/testing/selftests/bpf/unpriv_helpers.c
@@ -111,9 +111,8 @@ static int get_mitigations_off(void)
return !enabled_in_config;
}
-bool get_unpriv_disabled(void)
+bool get_unpriv_sysctl_disabled(void)
{
- int mitigations_off;
bool disabled;
char buf[2];
FILE *fd;
@@ -127,8 +126,12 @@ bool get_unpriv_disabled(void)
disabled = true;
}
- if (disabled)
- return true;
+ return disabled;
+}
+
+bool get_unpriv_mitigations_disabled(void)
+{
+ int mitigations_off;
/*
* Some unpriv tests rely on spectre mitigations being on.
@@ -144,6 +147,11 @@ bool get_unpriv_disabled(void)
return mitigations_off;
}
+bool get_unpriv_disabled(void)
+{
+ return get_unpriv_sysctl_disabled() || get_unpriv_mitigations_disabled();
+}
+
bool get_kasan_jit_enabled(void)
{
return config_contains("CONFIG_BPF_JIT_KASAN=y") == 1;
diff --git a/tools/testing/selftests/bpf/unpriv_helpers.h b/tools/testing/selftests/bpf/unpriv_helpers.h
index a7ceb51577cd..c24d53e14f3a 100644
--- a/tools/testing/selftests/bpf/unpriv_helpers.h
+++ b/tools/testing/selftests/bpf/unpriv_helpers.h
@@ -5,5 +5,7 @@
#define UNPRIV_SYSCTL "kernel/unprivileged_bpf_disabled"
bool get_unpriv_disabled(void);
+bool get_unpriv_sysctl_disabled(void);
+bool get_unpriv_mitigations_disabled(void);
bool get_kasan_jit_enabled(void);
bool get_kasan_multi_shot_enabled(void);
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH bpf-next v5 02/11] bpf: Record raw memory arguments during argument checking
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 01/11] selftests/bpf: Allow privileged preparation for capability tests Kumar Kartikeya Dwivedi
@ 2026-09-21 2:38 ` Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 03/11] bpf: Check __uninit kfunc output buffers as write-only Kumar Kartikeya Dwivedi
` (9 subsequent siblings)
11 siblings, 0 replies; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-21 2:38 UTC (permalink / raw)
To: bpf
Cc: Eduard Zingerman, Alexei Starovoitov, Andrii Nakryiko,
Daniel Borkmann, Emil Tsalapatis, Tejun Heo, Amery Hung, kkd,
kernel-team
From: Eduard Zingerman <eddyz87@gmail.com>
Record raw outputs in check_func_arg() after resolving their memory type,
instead of pre-recording them in check_raw_mode_ok(). Keep prototype
validation separate from per-call output tracking so kfuncs can share the
latter in a following patch.
For bpf_strtol(), bpf_strtoul() and bpf_kallsyms_lookup_name(), a variable-size
input is checked before the arg4 output. Recording outputs as their arguments
are checked prevents that input from clearing the output's raw mode early.
Only clear raw mode when a variable size belongs to the recorded output:
an unrelated later input must also preserve an earlier output descriptor.
For bpf_map_peek_elem() on a bloom filter, type resolution removes MEM_UNINIT
and MEM_WRITE because the value is an input. Recording the resolved type makes
the later bloom-filter-specific raw-mode reset unnecessary. Reuse the same
raw-memory predicate for output recording and prototype validation.
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
kernel/bpf/verifier.c | 31 ++++++++++++++-----------------
1 file changed, 14 insertions(+), 17 deletions(-)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 18aad4886f9c..67c1abfde922 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -7217,12 +7217,13 @@ static int check_mem_size_reg(struct bpf_verifier_env *env,
*/
meta->msize_max_value = reg_umax(size_reg);
- /* The register is SCALAR_VALUE; the access check happens using
- * its boundaries. For unprivileged variable accesses, disable
- * raw mode so that the program is required to initialize all
- * the memory that the helper could just partially fill up.
+ /*
+ * A variable size does not guarantee that the call initializes the whole
+ * checked range. Disable raw mode for this output and apply the ordinary
+ * stack initialization checks, including their privilege exceptions.
*/
- if (!tnum_is_const(size_reg->var_off))
+ if (!tnum_is_const(size_reg->var_off) &&
+ meta->arg_raw_mem.regno == reg_from_argno(mem_argno))
meta->arg_raw_mem.regno = 0;
if (reg_smin(size_reg) < 0) {
@@ -8940,6 +8941,9 @@ static int check_func_arg(struct bpf_verifier_env *env, u32 arg, u32 slot, u32 p
if (err)
return err;
+ if (!meta->btf && arg_type_is_raw_mem(arg_type))
+ meta->arg_raw_mem.regno = slot + 1;
+
if (bpf_register_is_null(reg) && type_may_be_null(arg_type)) {
err = mark_arg_precision(env, argno);
if (err)
@@ -9029,14 +9033,6 @@ static int check_func_arg(struct bpf_verifier_env *env, u32 arg, u32 slot, u32 p
return -EFAULT;
}
- /*
- * Disable raw mode for bpf_map_peek_elem() on a bloom filter. The helper reads
- * the value buffer as an input rather than filling it.
- */
- if (is_helper_call(meta, BPF_FUNC_map_peek_elem) &&
- meta->map.ptr->map_type == BPF_MAP_TYPE_BLOOM_FILTER)
- meta->arg_raw_mem.regno = 0;
-
err = check_helper_mem_access(env, reg, argno, meta->map.ptr->value_size,
arg_type & MEM_WRITE ? BPF_WRITE : BPF_READ,
false, meta, NULL);
@@ -9860,8 +9856,9 @@ static int check_map_func_compatibility(struct bpf_verifier_env *env,
return -EINVAL;
}
-static bool check_raw_mode_ok(const struct bpf_func_proto *fn, struct bpf_call_arg_meta *meta)
+static bool check_raw_mode_ok(const struct bpf_func_proto *fn)
{
+ bool seen = false;
int i;
for (i = 0; i < ARRAY_SIZE(fn->arg_type); i++) {
@@ -9869,9 +9866,9 @@ static bool check_raw_mode_ok(const struct bpf_func_proto *fn, struct bpf_call_a
break;
if (!arg_type_is_raw_mem(fn->arg_type[i]))
continue;
- if (meta->arg_raw_mem.regno)
+ if (seen)
return false;
- meta->arg_raw_mem.regno = i + 1;
+ seen = true;
}
return true;
@@ -10003,7 +10000,7 @@ static int check_func_proto(struct bpf_verifier_env *env, const struct bpf_func_
struct bpf_call_arg_meta *meta)
{
return check_arg_prog_aux(env, fn) &&
- check_raw_mode_ok(fn, meta) &&
+ check_raw_mode_ok(fn) &&
check_arg_pair_ok(fn) &&
check_mem_arg_rw_flag_ok(fn) &&
check_proto_release_reg(fn, meta) &&
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH bpf-next v5 03/11] bpf: Check __uninit kfunc output buffers as write-only
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 01/11] selftests/bpf: Allow privileged preparation for capability tests Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 02/11] bpf: Record raw memory arguments during argument checking Kumar Kartikeya Dwivedi
@ 2026-09-21 2:38 ` Kumar Kartikeya Dwivedi
2026-09-21 3:54 ` bot+bpf-ci
2026-09-21 2:38 ` [PATCH bpf-next v5 04/11] bpf: Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
` (8 subsequent siblings)
11 siblings, 1 reply; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-21 2:38 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, Emil Tsalapatis, Tejun Heo, Amery Hung, kkd,
kernel-team
Generic kfunc memory arguments are checked for both read and write access,
including __uninit outputs. An output pointing into a write-only map value
is therefore rejected, and a stack output must already hold readable
contents even though the kfunc only writes it.
Mark generic buffers in generated kfunc prototypes with MEM_WRITE, and keep
MEM_UNINIT when a struct pointer resolves to generic memory. Check __uninit
buffers as write-only while ordinary kfunc buffers retain read/write
access. Helper access selection is unchanged. Definite initialization of
stack outputs after the call is addressed separately.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
kernel/bpf/verifier.c | 12 +++++++-----
1 file changed, 7 insertions(+), 5 deletions(-)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 67c1abfde922..de40785d6a7b 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -9190,7 +9190,8 @@ static int check_func_arg(struct bpf_verifier_env *env, u32 arg, u32 slot, u32 p
break;
access_type = arg_type & MEM_WRITE ? BPF_WRITE : BPF_READ;
- if (meta->btf)
+ /* Ordinary kfunc buffers are input/output; __uninit buffers are outputs. */
+ if (meta->btf && !(arg_type & MEM_UNINIT))
access_type = BPF_READ | BPF_WRITE;
err = check_mem_reg(env, reg, argno, arg_size, access_type, meta, &known_memory);
@@ -9238,7 +9239,7 @@ static int check_func_arg(struct bpf_verifier_env *env, u32 arg, u32 slot, u32 p
break;
access_type = fn->arg_type[arg - 1] & MEM_WRITE ? BPF_WRITE : BPF_READ;
- if (meta->btf)
+ if (meta->btf && !(fn->arg_type[arg - 1] & MEM_UNINIT))
access_type = BPF_READ | BPF_WRITE;
zero_size_allowed = meta->btf || base_type(arg_type) == ARG_MEM_SIZE_OR_ZERO;
@@ -12457,7 +12458,8 @@ static int resolve_func_arg_type(struct bpf_verifier_env *env,
PTR_ERR(resolve_ret));
return -EINVAL;
}
- *arg_type = ARG_PTR_TO_MEM | MEM_FIXED_SIZE | (*arg_type & PTR_MAYBE_NULL);
+ *arg_type = ARG_PTR_TO_MEM | MEM_FIXED_SIZE | MEM_WRITE |
+ (*arg_type & (PTR_MAYBE_NULL | MEM_UNINIT));
return 0;
}
@@ -12970,7 +12972,7 @@ get_kfunc_arg_type(struct bpf_verifier_env *env, struct bpf_call_arg_meta *meta,
reg_arg_name(env, argno), btf_type_str(ref_t), ref_tname);
return -EINVAL;
}
- arg_type = ARG_PTR_TO_MEM;
+ arg_type = ARG_PTR_TO_MEM | MEM_WRITE;
} else if (btf_type_is_struct(ref_t))
/* A pointer to a struct without a size argument is classified as ARG_PTR_TO_BTF_ID */
arg_type = ARG_PTR_TO_BTF_ID;
@@ -12995,7 +12997,7 @@ get_kfunc_arg_type(struct bpf_verifier_env *env, struct bpf_call_arg_meta *meta,
return -EINVAL;
}
proto->arg_size[arg] = type_size;
- arg_type = ARG_PTR_TO_MEM | MEM_FIXED_SIZE;
+ arg_type = ARG_PTR_TO_MEM | MEM_FIXED_SIZE | MEM_WRITE;
}
if (is_kfunc_arg_uninit(meta->btf, &args[arg]))
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH bpf-next v5 03/11] bpf: Check __uninit kfunc output buffers as write-only
2026-09-21 2:38 ` [PATCH bpf-next v5 03/11] bpf: Check __uninit kfunc output buffers as write-only Kumar Kartikeya Dwivedi
@ 2026-09-21 3:54 ` bot+bpf-ci
0 siblings, 0 replies; 17+ messages in thread
From: bot+bpf-ci @ 2026-09-21 3:54 UTC (permalink / raw)
To: memxor, bpf
Cc: ast, andrii, daniel, eddyz87, emil, tj, ameryhung, kkd,
kernel-team, ast, andrii, daniel, martin.lau, eddyz87,
yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3331 bytes --]
> bpf: Check __uninit kfunc output buffers as write-only
>
> Generic kfunc memory arguments are checked for both read and write access,
> including __uninit outputs. An output pointing into a write-only map value
> is therefore rejected, and a stack output must already hold readable
> contents even though the kfunc only writes it.
> Mark generic buffers in generated kfunc prototypes with MEM_WRITE, and keep
> MEM_UNINIT when a struct pointer resolves to generic memory. Check __uninit
> buffers as write-only while ordinary kfunc buffers retain read/write
> access. Helper access selection is unchanged. Definite initialization of
> stack outputs after the call is addressed separately.
> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
The changelog states two problems: (1) '__uninit output pointing into a
write-only map value is rejected' and (2) 'a stack output must already
hold readable contents even though the kfunc only writes it'. Does this
patch actually fix (2)?
With access_type == BPF_WRITE, check_helper_mem_access() calls
check_stack_range_initialized() with clobber = true, but the acceptance
test for uninitialized bytes is independent of clobber:
if ((*stype == STACK_ZERO) ||
(*stype == STACK_INVALID && env->allow_uninit_stack)) {
if (clobber)
*stype = STACK_MISC;
goto mark;
}
So a genuinely uninitialized stack buffer passed to a __uninit kfunc
argument is still rejected with 'invalid read from stack' unless
env->allow_uninit_stack (requires CAP_PERFMON).
The verifier's raw-mode escape hatch cannot help either: at this commit
meta->arg_raw_mem is only recorded for helpers (in check_func_arg, 'if
(!meta->btf && arg_type_is_raw_mem(arg_type))'), so raw_mode is always
false for a kfunc argument.
Additionally, switching to BPF_WRITE makes clobber true, which also makes
allow_poison true. STACK_POISON bytes therefore pass the argument check,
but the poison branch does not convert them to STACK_MISC, and nothing
initializes them after the kfunc call. Since bpf_kfunc_stack_access_bytes()
reports a __uninit argument as a write, stack liveness kills the program's
own initialization of the buffer, so those bytes end up poisoned.
The net effect is that a generic __uninit kfunc output still cannot be read
back after the call; the failure merely moves from the call site to the
subsequent load.
Both gaps appear to be closed later in the series (de26ebd0567d3 adds
mark_raw_stack() to check_kfunc_call() and drops the '!meta->btf' guard,
and 06eb5f8a1b1b4 changes the STACK_INVALID test to include uninit).
Could the commit message be more precise about what is actually fixed at
this step versus what remains work-in-progress? The existing caveat
('Definite initialization of stack outputs after the call is addressed
separately') covers post-call marking, not the pre-call requirement that
the buffer already be readable.
This looks like a bug fix for kfunc argument checking. Should this include:
Fixes: ab7fff217fd1 ("bpf: Check helper and kfunc arguments in one path")
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35555832550
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH bpf-next v5 04/11] bpf: Fix generic __uninit kfunc output buffers
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
` (2 preceding siblings ...)
2026-09-21 2:38 ` [PATCH bpf-next v5 03/11] bpf: Check __uninit kfunc output buffers as write-only Kumar Kartikeya Dwivedi
@ 2026-09-21 2:38 ` Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 05/11] selftests/bpf: Cover generic __uninit output initialization Kumar Kartikeya Dwivedi
` (7 subsequent siblings)
11 siblings, 0 replies; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-21 2:38 UTC (permalink / raw)
To: bpf
Cc: Tejun Heo, Eduard Zingerman, Alexei Starovoitov, Andrii Nakryiko,
Daniel Borkmann, Emil Tsalapatis, Amery Hung, kkd, kernel-team
Stack liveness treats __uninit kfunc arguments as writes. Even with
write-only access checks, ordinary clobber handling leaves poisoned stack
bytes poisoned after the call, so the output cannot be read.
Reuse the helper output descriptor for generic kfunc buffers. Record the
output after resolving its memory type, and initialize its stack bytes
only after all arguments have been checked. This prevents an output from
making an aliased, uninitialized input appear valid.
Track the output by its ABI argument slot, since a kfunc pointer may be
passed on the stack or follow a multi-slot by-value argument. Keep the
single-output restriction in prototype validation; multiple-output
tracking is a separate extension.
Document that output buffers must be fully initialized on every return
path, including error returns and padding.
Fixes: 2cb27158adb3 ("bpf: poison dead stack slots")
Reported-by: Tejun Heo <tj@kernel.org>
Link: https://lore.kernel.org/bpf/86d966ec88bbf27d21b2bb4e18c8aa00@kernel.org
Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
Documentation/bpf/kfuncs.rst | 26 ++++++++----
include/linux/bpf_verifier.h | 7 ++--
kernel/bpf/verifier.c | 80 ++++++++++++++++++++++++++----------
3 files changed, 81 insertions(+), 32 deletions(-)
diff --git a/Documentation/bpf/kfuncs.rst b/Documentation/bpf/kfuncs.rst
index 3f300118a623..f393c3c3d3b4 100644
--- a/Documentation/bpf/kfuncs.rst
+++ b/Documentation/bpf/kfuncs.rst
@@ -164,19 +164,31 @@ suffix should be used.
2.3.3 __uninit Annotation
-------------------------
-This annotation is used to indicate that the argument will be treated as
-uninitialized.
+Use ``__uninit`` on a pointer parameter for an output that the kfunc
+initializes without reading its incoming contents.
-An example is given below::
+For generic memory buffers, the kfunc must initialize every byte in the
+declared range on every return path, including error returns and struct
+padding. The range is determined by the pointed-to type or the associated
+``__sz`` or ``__szk`` size argument.
+
+The annotation does not change the accepted pointer types. A stack-backed
+struct passed as a generic memory buffer must still be scalar-only.
+
+A stack buffer with a verifier-known constant offset and size may be
+uninitialized before the call and is considered initialized afterwards.
+For variable offsets or sizes, the usual stack-initialization and
+variable-offset restrictions still apply.
+
+For dynptr parameters, ``__uninit`` indicates that the kfunc constructs a
+dynptr in the supplied storage. For example::
- __bpf_kfunc int bpf_dynptr_from_skb(..., struct bpf_dynptr_kern *ptr__uninit)
+ __bpf_kfunc int bpf_dynptr_from_skb(..., struct bpf_dynptr *ptr__uninit)
{
...
}
-Here, the dynptr will be treated as an uninitialized dynptr. Without this
-annotation, the verifier will reject the program if the dynptr passed in is
-not initialized.
+Without this annotation, a dynptr argument must already be initialized.
2.3.4 __nullable Annotation
---------------------------
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index cf85141ea167..6ff1c227d298 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1548,10 +1548,11 @@ struct ref_obj_desc {
/*
* A memory argument a call fills in. The verifier allows the stack to be uninitialized if
- * the range is a known constant. Stack slots are marked as STACK_MISC by check_mem_access().
+ * the range is a known constant. Stack slots are marked as STACK_MISC by check_mem_access()
+ * after all arguments have been checked.
*/
struct arg_raw_mem_desc {
- u8 regno;
+ u8 argno; /* One-based ABI argument slot; zero means no output. */
int size;
};
@@ -1579,6 +1580,7 @@ struct bpf_call_arg_meta {
struct bpf_dynptr_desc dynptr;
struct ref_obj_desc ref_obj;
struct ret_mem_desc ret_mem;
+ struct arg_raw_mem_desc arg_raw_mem;
/* Only set by kfunc */
bool r0_rdonly;
@@ -1617,7 +1619,6 @@ struct bpf_call_arg_meta {
s64 const_map_key;
struct btf *ret_btf;
struct btf_field *kptr_field;
- struct arg_raw_mem_desc arg_raw_mem;
};
int bpf_get_helper_proto(struct bpf_verifier_env *env, int func_id,
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index de40785d6a7b..f4b88e402ff5 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -302,6 +302,12 @@ static int arg_idx_from_argno(argno_t a)
return arg_from_argno(a) - 1;
}
+/* Normalize helper register numbers and kfunc argument numbers to ABI slots. */
+static u32 arg_slot_from_argno(argno_t a)
+{
+ return abs(a.argno) - 1;
+}
+
static const char *btf_type_name(const struct btf *btf, u32 id)
{
return btf_name_by_offset(btf, btf_type_by_id(btf, id)->name_off);
@@ -6988,7 +6994,7 @@ static int check_stack_range_initialized(
*/
bool allow_poison = access_size < 0 || clobber;
/* The call will initialize the memory; uninitialized stack allowed */
- bool raw_mode = meta && meta->arg_raw_mem.regno == reg_from_argno(argno);
+ bool raw_mode = meta && meta->arg_raw_mem.argno == arg_slot_from_argno(argno) + 1;
access_size = abs(access_size);
@@ -7017,11 +7023,9 @@ static int check_stack_range_initialized(
reg_arg_name(env, argno), tn_buf);
return -EACCES;
}
- /* Only initialized buffer on stack is allowed to be accessed
- * with variable offset. With uninitialized buffer it's hard to
- * guarantee that whole memory is marked as initialized on
- * helper return since specific bounds are unknown what may
- * cause uninitialized stack leaking.
+ /*
+ * The call may touch any byte in the possible range, but does not
+ * definitely initialize all of it. Fall back to ordinary stack checks.
*/
raw_mode = false;
@@ -7223,8 +7227,8 @@ static int check_mem_size_reg(struct bpf_verifier_env *env,
* stack initialization checks, including their privilege exceptions.
*/
if (!tnum_is_const(size_reg->var_off) &&
- meta->arg_raw_mem.regno == reg_from_argno(mem_argno))
- meta->arg_raw_mem.regno = 0;
+ meta->arg_raw_mem.argno == arg_slot_from_argno(mem_argno) + 1)
+ meta->arg_raw_mem.argno = 0;
if (reg_smin(size_reg) < 0) {
verbose(env, "%s min value is negative, either use unsigned or 'var &= const'\n",
@@ -8941,8 +8945,8 @@ static int check_func_arg(struct bpf_verifier_env *env, u32 arg, u32 slot, u32 p
if (err)
return err;
- if (!meta->btf && arg_type_is_raw_mem(arg_type))
- meta->arg_raw_mem.regno = slot + 1;
+ if (arg_type_is_raw_mem(arg_type))
+ meta->arg_raw_mem.argno = slot + 1;
if (bpf_register_is_null(reg) && type_may_be_null(arg_type)) {
err = mark_arg_precision(env, argno);
@@ -9565,6 +9569,33 @@ static int check_func_args(struct bpf_verifier_env *env, struct bpf_call_arg_met
return 0;
}
+static int mark_raw_stack(struct bpf_verifier_env *env, struct bpf_call_arg_meta *meta,
+ int insn_idx)
+{
+ struct bpf_func_state *caller = cur_func(env);
+ struct bpf_reg_state *reg;
+ u32 slot = meta->arg_raw_mem.argno - 1;
+ int i, err;
+
+ if (!meta->arg_raw_mem.size)
+ return 0;
+ reg = get_func_arg_reg(caller, cur_regs(env), slot);
+
+ /*
+ * Validate every argument before initializing outputs: an input argument
+ * may alias an output buffer. Use the normal stack-write checks to discard
+ * stale spills and preserve the rules for special stack objects.
+ */
+ for (i = 0; i < meta->arg_raw_mem.size; i++) {
+ err = check_mem_access(env, insn_idx, reg, argno_from_arg(slot + 1), i, BPF_B,
+ BPF_WRITE, -1, false, false);
+ if (err)
+ return err;
+ }
+
+ return 0;
+}
+
static bool may_update_sockmap(struct bpf_verifier_env *env, int func_id)
{
enum bpf_attach_type eatype = env->prog->expected_attach_type;
@@ -9863,9 +9894,13 @@ static bool check_raw_mode_ok(const struct bpf_func_proto *fn)
int i;
for (i = 0; i < ARRAY_SIZE(fn->arg_type); i++) {
- if (fn->arg_type[i] == ARG_UNUSED)
+ enum bpf_arg_type type = fn->arg_type[i];
+
+ if (type == ARG_UNUSED)
break;
- if (!arg_type_is_raw_mem(fn->arg_type[i]))
+ /* Struct pointers may resolve to generic memory during argument checking. */
+ if (!arg_type_is_raw_mem(type) &&
+ !(base_type(type) == ARG_PTR_TO_BTF_ID && (type & MEM_UNINIT)))
continue;
if (seen)
return false;
@@ -11588,16 +11623,9 @@ static int check_helper_call(struct bpf_verifier_env *env, struct bpf_insn *insn
regs = cur_regs(env);
- /* Mark slots with STACK_MISC in case of raw mode, stack offset
- * is inferred from register state.
- */
- for (i = 0; i < meta.arg_raw_mem.size; i++) {
- err = check_mem_access(env, insn_idx, regs + meta.arg_raw_mem.regno,
- argno_from_reg(meta.arg_raw_mem.regno), i, BPF_B,
- BPF_WRITE, -1, false, false);
- if (err)
- return err;
- }
+ err = mark_raw_stack(env, &meta, insn_idx);
+ if (err)
+ return err;
if (meta.release_regno) {
struct bpf_reg_state *reg = ®s[meta.release_regno];
@@ -13101,6 +13129,10 @@ static int gen_kfunc_arg_proto(struct bpf_verifier_env *env, struct bpf_call_arg
return -ENOTSUPP;
}
+ if (!check_raw_mode_ok(proto)) {
+ verbose(env, "multiple __uninit buffers are not supported\n");
+ return -EINVAL;
+ }
return check_arg_prog_aux(env, proto) ? 0 : -EINVAL;
}
@@ -14178,6 +14210,10 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
if (err < 0)
return err;
+ err = mark_raw_stack(env, &meta, insn_idx);
+ if (err)
+ return err;
+
if ((is_bpf_obj_drop_kfunc(meta.func_id) ||
is_bpf_percpu_obj_drop_kfunc(meta.func_id)) && (is_tracing_prog_type(prog_type) ||
/* is_tracing_prog_type() for now doesn't cover non-iterator tracing progs. */
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH bpf-next v5 05/11] selftests/bpf: Cover generic __uninit output initialization
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
` (3 preceding siblings ...)
2026-09-21 2:38 ` [PATCH bpf-next v5 04/11] bpf: Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
@ 2026-09-21 2:38 ` Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 06/11] bpf: Support multiple __uninit kfunc output arguments Kumar Kartikeya Dwivedi
` (6 subsequent siblings)
11 siblings, 0 replies; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-21 2:38 UTC (permalink / raw)
To: bpf
Cc: Amery Hung, Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, Emil Tsalapatis, Tejun Heo, kkd, kernel-team
Exercise the struct and sized-buffer cases where stack liveness poisons an
output before a kfunc call. Check that the verifier accepts these outputs
and that the kfunc initializes the memory read after the call.
Verify that an uninitialized input aliasing an output is still rejected
without CAP_PERFMON or CAP_SYS_ADMIN. Include an initialized alias as a
positive control, using an int-width store so its value is independent of
endianness.
Use __prepare_priv to resolve the test module's BTF before dropping to
CAP_BPF and CAP_NET_ADMIN for program loading. Keep multiple-output and
argument-slot coverage separate from these immediate regression tests.
Reviewed-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
.../selftests/bpf/prog_tests/verifier.c | 2 +
.../bpf/progs/verifier_kfunc_uninit.c | 100 ++++++++++++++++++
.../selftests/bpf/test_kmods/bpf_testmod.c | 25 +++++
.../bpf/test_kmods/bpf_testmod_kfunc.h | 3 +
4 files changed, 130 insertions(+)
create mode 100644 tools/testing/selftests/bpf/progs/verifier_kfunc_uninit.c
diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index 973bbeda9318..ffeba2d1464a 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -56,6 +56,7 @@
#include "verifier_jeq_infer_not_null.skel.h"
#include "verifier_jit_convergence.skel.h"
#include "verifier_kfunc_packet_access.skel.h"
+#include "verifier_kfunc_uninit.skel.h"
#include "verifier_ld_ind.skel.h"
#include "verifier_ldsx.skel.h"
#include "verifier_leak_ptr.skel.h"
@@ -223,6 +224,7 @@ void test_verifier_iterating_callbacks(void) { RUN(verifier_iterating_callbacks
void test_verifier_jeq_infer_not_null(void) { RUN(verifier_jeq_infer_not_null); }
void test_verifier_jit_convergence(void) { RUN(verifier_jit_convergence); }
void test_verifier_kfunc_packet_access(void) { RUN_TESTS(verifier_kfunc_packet_access); }
+void test_verifier_kfunc_uninit(void) { RUN_TESTS(verifier_kfunc_uninit); }
void test_verifier_load_acquire(void) { RUN(verifier_load_acquire); }
void test_verifier_ld_ind(void) { RUN(verifier_ld_ind); }
void test_verifier_ldsx(void) { RUN(verifier_ldsx); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit.c b/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit.c
new file mode 100644
index 000000000000..f7818303e703
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit.c
@@ -0,0 +1,100 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "../test_kmods/bpf_testmod_kfunc.h"
+
+/* Keep the kfunc BTF records used by the inline assembly. */
+void __kfunc_btf_root(void)
+{
+ asm volatile ("" :
+ : "r"(&bpf_kfunc_test_uninit_struct),
+ "r"(&bpf_kfunc_test_uninit_mem),
+ "r"(&bpf_kfunc_test_uninit_alias));
+}
+
+SEC("tc")
+__success __retval(10)
+__flag(BPF_F_TEST_STATE_FREQ)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__success_unpriv
+__naked void struct_poisoned_at_checkpoint(void)
+{
+ asm volatile (
+ "*(u64 *)(r10 - 16) = 0;"
+ "*(u64 *)(r10 - 8) = 0;"
+ "goto +0;"
+ "r1 = r10;"
+ "r1 += -16;"
+ "call %[bpf_kfunc_test_uninit_struct];"
+ "r0 = *(u32 *)(r10 - 16);"
+ "r1 = *(u32 *)(r10 - 12);"
+ "r0 += r1;"
+ "r1 = *(u32 *)(r10 - 8);"
+ "r0 += r1;"
+ "r1 = *(u32 *)(r10 - 4);"
+ "r0 += r1;"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_struct) : __clobber_all);
+}
+
+SEC("tc")
+__success __retval(0x2a2a2a2a)
+__flag(BPF_F_TEST_STATE_FREQ)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__success_unpriv
+__naked void sized_buffer_poisoned_at_checkpoint(void)
+{
+ asm volatile (
+ "*(u64 *)(r10 - 8) = 0;"
+ "goto +0;"
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = 8;"
+ "call %[bpf_kfunc_test_uninit_mem];"
+ "r0 = *(u32 *)(r10 - 8);"
+ "r1 = *(u32 *)(r10 - 4);"
+ "if r0 == r1 goto +1;"
+ "r0 = 0;"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_mem) : __clobber_all);
+}
+
+SEC("tc")
+__success __retval(7)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__success_unpriv
+__naked void initialized_input_alias(void)
+{
+ asm volatile (
+ "*(u32 *)(r10 - 8) = 7;"
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = r1;"
+ "call %[bpf_kfunc_test_uninit_alias];"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_alias) : __clobber_all);
+}
+
+SEC("tc")
+__success
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__failure_unpriv __msg_unpriv("invalid read from stack")
+__naked void uninitialized_input_alias(void)
+{
+ asm volatile (
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = r1;"
+ "call %[bpf_kfunc_test_uninit_alias];"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_alias) : __clobber_all);
+}
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c b/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
index fd2c0cdc91b1..542edeb28b27 100644
--- a/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
+++ b/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
@@ -17,6 +17,7 @@
#include <linux/in.h>
#include <linux/in6.h>
#include <linux/un.h>
+#include <linux/unaligned.h>
#include <linux/filter.h>
#include <linux/rcupdate_trace.h>
#include <net/sock.h>
@@ -1316,6 +1317,27 @@ __bpf_kfunc void bpf_kfunc_call_test_pass2(struct prog_test_pass2 *p)
{
}
+__bpf_kfunc void bpf_kfunc_test_uninit_struct(struct prog_test_pass1 *out__uninit)
+{
+ out__uninit->x0 = 1;
+ out__uninit->x1 = 2;
+ out__uninit->x2 = 3;
+ out__uninit->x3 = 4;
+}
+
+__bpf_kfunc void bpf_kfunc_test_uninit_mem(void *out__uninit, u32 out__sz)
+{
+ memset(out__uninit, 0x2a, out__sz);
+}
+
+__bpf_kfunc int bpf_kfunc_test_uninit_alias(int *out__uninit, const int *in)
+{
+ int value = get_unaligned(in);
+
+ put_unaligned(42, out__uninit);
+ return value;
+}
+
__bpf_kfunc void bpf_kfunc_call_test_fail1(struct prog_test_fail1 *p)
{
}
@@ -1747,6 +1769,9 @@ BTF_ID_FLAGS(func, bpf_kfunc_call_int_mem_release, KF_RELEASE)
BTF_ID_FLAGS(func, bpf_kfunc_call_test_pass_ctx)
BTF_ID_FLAGS(func, bpf_kfunc_call_test_pass1)
BTF_ID_FLAGS(func, bpf_kfunc_call_test_pass2)
+BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_struct)
+BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_mem)
+BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_alias)
BTF_ID_FLAGS(func, bpf_kfunc_call_test_fail1)
BTF_ID_FLAGS(func, bpf_kfunc_call_test_fail2)
BTF_ID_FLAGS(func, bpf_kfunc_call_test_fail3)
diff --git a/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h b/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h
index d3696d5254c9..91b64e783123 100644
--- a/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h
+++ b/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h
@@ -289,6 +289,9 @@ __u64 bpf_kfunc_call_stack_arg_big(__u64 a, __u64 b, __u64 c, __u64 d, __u64 e,
void bpf_kfunc_call_test_pass_ctx(struct __sk_buff *skb) __ksym;
void bpf_kfunc_call_test_pass1(struct prog_test_pass1 *p) __ksym;
void bpf_kfunc_call_test_pass2(struct prog_test_pass2 *p) __ksym;
+void bpf_kfunc_test_uninit_struct(struct prog_test_pass1 *out__uninit) __ksym;
+void bpf_kfunc_test_uninit_mem(void *out__uninit, __u32 out__sz) __ksym;
+int bpf_kfunc_test_uninit_alias(int *out__uninit, const int *in) __ksym;
void bpf_kfunc_call_test_mem_len_fail2(__u64 *mem, int len) __ksym;
void bpf_kfunc_call_test_destructive(void) __ksym;
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH bpf-next v5 06/11] bpf: Support multiple __uninit kfunc output arguments
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
` (4 preceding siblings ...)
2026-09-21 2:38 ` [PATCH bpf-next v5 05/11] selftests/bpf: Cover generic __uninit output initialization Kumar Kartikeya Dwivedi
@ 2026-09-21 2:38 ` Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 07/11] selftests/bpf: Cover __uninit kfunc output argument slots Kumar Kartikeya Dwivedi
` (5 subsequent siblings)
11 siblings, 0 replies; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-21 2:38 UTC (permalink / raw)
To: bpf
Cc: Amery Hung, Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, Emil Tsalapatis, Tejun Heo, kkd, kernel-team
A single output descriptor cannot retain the initialization ranges of two
__uninit memory arguments. Track raw-mode eligibility and the definite
output size separately for each ABI slot, so a variable-size output leaves
independent constant-size outputs intact.
Use the shared argument-checking path to record outputs for both helpers
and kfuncs. With per-slot tracking, neither needs the single-output
prototype restriction. Initialize every recorded output after checking all
arguments, skipping empty slots before looking up their register state.
Use the resolved BTF parameter when looking up __uninit for stack liveness.
A preceding by-value parameter can consume multiple ABI slots, so its slot
number cannot index the BTF parameter array.
Suggested-by: Amery Hung <ameryhung@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
include/linux/bpf_verifier.h | 10 +++---
kernel/bpf/verifier.c | 68 ++++++++++++------------------------
2 files changed, 28 insertions(+), 50 deletions(-)
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 6ff1c227d298..9ddbb20ec1e9 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1547,13 +1547,13 @@ struct ref_obj_desc {
};
/*
- * A memory argument a call fills in. The verifier allows the stack to be uninitialized if
- * the range is a known constant. Stack slots are marked as STACK_MISC by check_mem_access()
- * after all arguments have been checked.
+ * Memory arguments a call fills in, indexed by argument slot. The verifier allows the
+ * stack to be uninitialized if the range is a known constant. Stack slots are marked as
+ * STACK_MISC by check_mem_access() after all arguments have been checked.
*/
struct arg_raw_mem_desc {
- u8 argno; /* One-based ABI argument slot; zero means no output. */
- int size;
+ u16 mask;
+ int size[MAX_BPF_FUNC_ARGS];
};
/* Size of PTR_TO_MEM returned, taken from a constant allocation-size argument */
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index f4b88e402ff5..0c94f1214cc3 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -6994,7 +6994,9 @@ static int check_stack_range_initialized(
*/
bool allow_poison = access_size < 0 || clobber;
/* The call will initialize the memory; uninitialized stack allowed */
- bool raw_mode = meta && meta->arg_raw_mem.argno == arg_slot_from_argno(argno) + 1;
+ u32 arg_slot = arg_slot_from_argno(argno);
+ bool raw_mode = meta && arg_slot < MAX_BPF_FUNC_ARGS &&
+ (meta->arg_raw_mem.mask & BIT(arg_slot));
access_size = abs(access_size);
@@ -7034,7 +7036,7 @@ static int check_stack_range_initialized(
}
if (raw_mode) {
- meta->arg_raw_mem.size = access_size;
+ meta->arg_raw_mem.size[arg_slot] = access_size;
return 0;
}
@@ -7226,9 +7228,8 @@ static int check_mem_size_reg(struct bpf_verifier_env *env,
* checked range. Disable raw mode for this output and apply the ordinary
* stack initialization checks, including their privilege exceptions.
*/
- if (!tnum_is_const(size_reg->var_off) &&
- meta->arg_raw_mem.argno == arg_slot_from_argno(mem_argno) + 1)
- meta->arg_raw_mem.argno = 0;
+ if (!tnum_is_const(size_reg->var_off))
+ meta->arg_raw_mem.mask &= ~BIT(arg_slot_from_argno(mem_argno));
if (reg_smin(size_reg) < 0) {
verbose(env, "%s min value is negative, either use unsigned or 'var &= const'\n",
@@ -8946,7 +8947,7 @@ static int check_func_arg(struct bpf_verifier_env *env, u32 arg, u32 slot, u32 p
return err;
if (arg_type_is_raw_mem(arg_type))
- meta->arg_raw_mem.argno = slot + 1;
+ meta->arg_raw_mem.mask |= BIT(slot);
if (bpf_register_is_null(reg) && type_may_be_null(arg_type)) {
err = mark_arg_precision(env, argno);
@@ -9573,24 +9574,28 @@ static int mark_raw_stack(struct bpf_verifier_env *env, struct bpf_call_arg_meta
int insn_idx)
{
struct bpf_func_state *caller = cur_func(env);
- struct bpf_reg_state *reg;
- u32 slot = meta->arg_raw_mem.argno - 1;
+ u32 slot;
int i, err;
- if (!meta->arg_raw_mem.size)
- return 0;
- reg = get_func_arg_reg(caller, cur_regs(env), slot);
-
/*
* Validate every argument before initializing outputs: an input argument
* may alias an output buffer. Use the normal stack-write checks to discard
* stale spills and preserve the rules for special stack objects.
*/
- for (i = 0; i < meta->arg_raw_mem.size; i++) {
- err = check_mem_access(env, insn_idx, reg, argno_from_arg(slot + 1), i, BPF_B,
- BPF_WRITE, -1, false, false);
- if (err)
- return err;
+ for (slot = 0; slot < MAX_BPF_FUNC_ARGS; slot++) {
+ struct bpf_reg_state *reg;
+ argno_t argno = argno_from_arg(slot + 1);
+
+ if (!meta->arg_raw_mem.size[slot])
+ continue;
+ reg = get_func_arg_reg(caller, cur_regs(env), slot);
+
+ for (i = 0; i < meta->arg_raw_mem.size[slot]; i++) {
+ err = check_mem_access(env, insn_idx, reg, argno, i, BPF_B,
+ BPF_WRITE, -1, false, false);
+ if (err)
+ return err;
+ }
}
return 0;
@@ -9888,28 +9893,6 @@ static int check_map_func_compatibility(struct bpf_verifier_env *env,
return -EINVAL;
}
-static bool check_raw_mode_ok(const struct bpf_func_proto *fn)
-{
- bool seen = false;
- int i;
-
- for (i = 0; i < ARRAY_SIZE(fn->arg_type); i++) {
- enum bpf_arg_type type = fn->arg_type[i];
-
- if (type == ARG_UNUSED)
- break;
- /* Struct pointers may resolve to generic memory during argument checking. */
- if (!arg_type_is_raw_mem(type) &&
- !(base_type(type) == ARG_PTR_TO_BTF_ID && (type & MEM_UNINIT)))
- continue;
- if (seen)
- return false;
- seen = true;
- }
-
- return true;
-}
-
static bool check_args_pair_invalid(const struct bpf_func_proto *fn, int arg)
{
bool is_fixed = fn->arg_type[arg] & MEM_FIXED_SIZE;
@@ -10036,7 +10019,6 @@ static int check_func_proto(struct bpf_verifier_env *env, const struct bpf_func_
struct bpf_call_arg_meta *meta)
{
return check_arg_prog_aux(env, fn) &&
- check_raw_mode_ok(fn) &&
check_arg_pair_ok(fn) &&
check_mem_arg_rw_flag_ok(fn) &&
check_proto_release_reg(fn, meta) &&
@@ -13129,10 +13111,6 @@ static int gen_kfunc_arg_proto(struct bpf_verifier_env *env, struct bpf_call_arg
return -ENOTSUPP;
}
- if (!check_raw_mode_ok(proto)) {
- verbose(env, "multiple __uninit buffers are not supported\n");
- return -EINVAL;
- }
return check_arg_prog_aux(env, proto) ? 0 : -EINVAL;
}
@@ -13929,7 +13907,7 @@ s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn *
/* KF_ITER_NEW kfuncs initialize the iterator state at arg 0 */
if (arg == 0 && meta.kfunc_flags & KF_ITER_NEW)
return -size;
- if (is_kfunc_arg_uninit(btf, &args[arg]))
+ if (is_kfunc_arg_uninit(btf, &args[i]))
return -size;
return size;
}
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH bpf-next v5 07/11] selftests/bpf: Cover __uninit kfunc output argument slots
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
` (5 preceding siblings ...)
2026-09-21 2:38 ` [PATCH bpf-next v5 06/11] bpf: Support multiple __uninit kfunc output arguments Kumar Kartikeya Dwivedi
@ 2026-09-21 2:38 ` Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 08/11] bpf: Preserve stack initialization for generic output buffers Kumar Kartikeya Dwivedi
` (4 subsequent siblings)
11 siblings, 0 replies; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-21 2:38 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, Emil Tsalapatis, Tejun Heo, Amery Hung, kkd,
kernel-team
Keep coverage for per-slot output tracking separate from the immediate
single-output regression tests. Check that both constant-size outputs are
initialized, and that a variable-size output does not disable initialization
of an independent constant-size output.
Also exercise an output following a by-value parameter that occupies two
argument slots, and an output pointer passed on the stack. Run each case
with normal capabilities and with CAP_BPF and CAP_NET_ADMIN only. Leave the
stack-passed output uninitialized so its reduced-capability case fails if
the verifier treats it as an ordinary input buffer.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
.../selftests/bpf/prog_tests/verifier.c | 2 +
.../bpf/progs/verifier_kfunc_uninit_multi.c | 113 ++++++++++++++++++
.../selftests/bpf/test_kmods/bpf_testmod.c | 20 ++++
.../bpf/test_kmods/bpf_testmod_kfunc.h | 4 +
4 files changed, 139 insertions(+)
create mode 100644 tools/testing/selftests/bpf/progs/verifier_kfunc_uninit_multi.c
diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index ffeba2d1464a..4f1e1c1cd5ab 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -57,6 +57,7 @@
#include "verifier_jit_convergence.skel.h"
#include "verifier_kfunc_packet_access.skel.h"
#include "verifier_kfunc_uninit.skel.h"
+#include "verifier_kfunc_uninit_multi.skel.h"
#include "verifier_ld_ind.skel.h"
#include "verifier_ldsx.skel.h"
#include "verifier_leak_ptr.skel.h"
@@ -225,6 +226,7 @@ void test_verifier_jeq_infer_not_null(void) { RUN(verifier_jeq_infer_not_null)
void test_verifier_jit_convergence(void) { RUN(verifier_jit_convergence); }
void test_verifier_kfunc_packet_access(void) { RUN_TESTS(verifier_kfunc_packet_access); }
void test_verifier_kfunc_uninit(void) { RUN_TESTS(verifier_kfunc_uninit); }
+void test_verifier_kfunc_uninit_multi(void) { RUN_TESTS(verifier_kfunc_uninit_multi); }
void test_verifier_load_acquire(void) { RUN(verifier_load_acquire); }
void test_verifier_ld_ind(void) { RUN(verifier_ld_ind); }
void test_verifier_ldsx(void) { RUN(verifier_ldsx); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit_multi.c b/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit_multi.c
new file mode 100644
index 000000000000..18ced0d0a3b2
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit_multi.c
@@ -0,0 +1,113 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_misc.h"
+#include "../test_kmods/bpf_testmod_kfunc.h"
+
+/* Keep the kfunc BTF records used by the inline assembly. */
+void __kfunc_btf_root(void)
+{
+ asm volatile ("" :
+ : "r"(&bpf_kfunc_test_uninit_multi),
+ "r"(&bpf_kfunc_test_uninit_pair),
+ "r"(&bpf_kfunc_test_uninit_stack));
+}
+
+SEC("tc")
+__success __retval(42)
+__flag(BPF_F_TEST_STATE_FREQ)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__success_unpriv
+__naked void multiple_outputs(void)
+{
+ asm volatile (
+ "*(u64 *)(r10 - 16) = 0;"
+ "*(u64 *)(r10 - 8) = 0;"
+ "goto +0;"
+ "r1 = r10;"
+ "r1 += -16;"
+ "r2 = r10;"
+ "r2 += -8;"
+ "r3 = 8;"
+ "call %[bpf_kfunc_test_uninit_multi];"
+ "r0 = *(u32 *)(r10 - 16);"
+ "r1 = *(u32 *)(r10 - 8);"
+ "if r1 != 0x2a2a2a2a goto 1f;"
+ "r1 = *(u32 *)(r10 - 4);"
+ "if r1 == 0x2a2a2a2a goto 2f;"
+ "1:;"
+ "r0 = 0;"
+ "2:;"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_multi) : __clobber_all);
+}
+
+SEC("tc")
+__success __retval(42)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__success_unpriv
+__naked void variable_size_preserves_other_output(void)
+{
+ asm volatile (
+ "r3 = *(u32 *)(r1 + 0);"
+ "r3 &= 7;"
+ "*(u64 *)(r10 - 8) = 0;"
+ "r1 = r10;"
+ "r1 += -16;"
+ "r2 = r10;"
+ "r2 += -8;"
+ "call %[bpf_kfunc_test_uninit_multi];"
+ "r0 = *(u32 *)(r10 - 16);"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_multi) : __clobber_all);
+}
+
+SEC("tc")
+__arch_x86_64 __arch_arm64
+__success __retval(42)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__success_unpriv
+__naked void output_after_by_value_argument(void)
+{
+ asm volatile (
+ "r1 = 20;"
+ "r2 = 22;"
+ "r3 = r10;"
+ "r3 += -8;"
+ "call %[bpf_kfunc_test_uninit_pair];"
+ "r0 = *(u32 *)(r10 - 8);"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_pair) : __clobber_all);
+}
+
+#if defined(__BPF_FEATURE_STACK_ARGUMENT)
+SEC("tc")
+__arch_x86_64 __arch_arm64 __arch_riscv64
+__success __retval(15)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__success_unpriv
+__naked void output_passed_on_stack(void)
+{
+ asm volatile (
+ "r1 = 1;"
+ "r2 = 2;"
+ "r3 = 3;"
+ "r4 = 4;"
+ "r5 = 5;"
+ "r6 = r10;"
+ "r6 += -8;"
+ "*(u64 *)(r11 - 8) = r6;"
+ "call %[bpf_kfunc_test_uninit_stack];"
+ "r0 = *(u32 *)(r10 - 8);"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_stack) : __clobber_all);
+}
+#endif
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c b/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
index 542edeb28b27..211886a8ee87 100644
--- a/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
+++ b/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
@@ -1338,6 +1338,23 @@ __bpf_kfunc int bpf_kfunc_test_uninit_alias(int *out__uninit, const int *in)
return value;
}
+__bpf_kfunc void bpf_kfunc_test_uninit_multi(int *a__uninit, void *b__uninit, u32 b__sz)
+{
+ put_unaligned(42, a__uninit);
+ memset(b__uninit, 0x2a, b__sz);
+}
+
+__bpf_kfunc void bpf_kfunc_test_uninit_pair(struct prog_test_pair_arg p, int *out__uninit)
+{
+ put_unaligned((int)(p.lo + p.hi), out__uninit);
+}
+
+__bpf_kfunc void bpf_kfunc_test_uninit_stack(u64 a, u64 b, u64 c, u64 d, u64 e,
+ int *out__uninit)
+{
+ put_unaligned((int)(a + b + c + d + e), out__uninit);
+}
+
__bpf_kfunc void bpf_kfunc_call_test_fail1(struct prog_test_fail1 *p)
{
}
@@ -1772,6 +1789,9 @@ BTF_ID_FLAGS(func, bpf_kfunc_call_test_pass2)
BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_struct)
BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_mem)
BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_alias)
+BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_multi)
+BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_pair)
+BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_stack)
BTF_ID_FLAGS(func, bpf_kfunc_call_test_fail1)
BTF_ID_FLAGS(func, bpf_kfunc_call_test_fail2)
BTF_ID_FLAGS(func, bpf_kfunc_call_test_fail3)
diff --git a/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h b/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h
index 91b64e783123..524f2cb9bdf4 100644
--- a/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h
+++ b/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h
@@ -292,6 +292,10 @@ void bpf_kfunc_call_test_pass2(struct prog_test_pass2 *p) __ksym;
void bpf_kfunc_test_uninit_struct(struct prog_test_pass1 *out__uninit) __ksym;
void bpf_kfunc_test_uninit_mem(void *out__uninit, __u32 out__sz) __ksym;
int bpf_kfunc_test_uninit_alias(int *out__uninit, const int *in) __ksym;
+void bpf_kfunc_test_uninit_multi(int *a__uninit, void *b__uninit, __u32 b__sz) __ksym;
+void bpf_kfunc_test_uninit_pair(struct prog_test_pair_arg p, int *out__uninit) __ksym;
+void bpf_kfunc_test_uninit_stack(__u64 a, __u64 b, __u64 c, __u64 d, __u64 e,
+ int *out__uninit) __ksym;
void bpf_kfunc_call_test_mem_len_fail2(__u64 *mem, int len) __ksym;
void bpf_kfunc_call_test_destructive(void) __ksym;
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH bpf-next v5 08/11] bpf: Preserve stack initialization for generic output buffers
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
` (6 preceding siblings ...)
2026-09-21 2:38 ` [PATCH bpf-next v5 07/11] selftests/bpf: Cover __uninit kfunc output argument slots Kumar Kartikeya Dwivedi
@ 2026-09-21 2:38 ` Kumar Kartikeya Dwivedi
2026-09-21 3:54 ` bot+bpf-ci
2026-09-21 2:38 ` [PATCH bpf-next v5 09/11] selftests/bpf: Cover generic output stack initialization Kumar Kartikeya Dwivedi
` (3 subsequent siblings)
11 siblings, 1 reply; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-21 2:38 UTC (permalink / raw)
To: bpf
Cc: Eduard Zingerman, Alexei Starovoitov, Andrii Nakryiko,
Daniel Borkmann, Emil Tsalapatis, Tejun Heo, Amery Hung, kkd,
kernel-team
Partial-output helpers such as bpf_snprintf() do not read incoming buffer
contents, but may leave some bytes untouched. Let them accept uninitialized
storage without promising full initialization to callers that cannot read
uninitialized stack memory.
Make this the default for generic MEM_UNINIT buffers, including __uninit
kfunc arguments. Allow invalid stack bytes through the output check but
leave them invalid when uninitialized stack reads are not allowed. Scrub
initialized bytes and scalar spills as possible writes, retaining the
existing restrictions on spilled pointers and special stack objects.
Retain the output annotation for variable-sized arguments and track their
raw-mode eligibility separately. Callers allowed uninitialized stack reads
can continue treating the potentially written range as initialized. For
constant ranges, defer that initialization until all inputs are checked.
Keep prior stack contents live for generic outputs when the caller cannot
read uninitialized stack memory. Such calls do not define the entire range,
so liveness must preserve initialization facts that remain relevant after
the call. Dynptr and iterator constructors still define their storage.
Annotate the snprintf, sysctl name, d_path, snprintf_btf and branch-record
destinations with MEM_UNINIT, and document that generic __uninit kfuncs may
leave bytes untouched. This also changes readback from existing full-writing
helpers: without permission to read uninitialized stack memory, programs
must initialize those bytes themselves before reading them after a call.
Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
Documentation/bpf/kfuncs.rst | 21 +++++++++--------
include/linux/bpf.h | 5 +++-
include/linux/bpf_verifier.h | 8 ++++---
kernel/bpf/cgroup.c | 2 +-
kernel/bpf/helpers.c | 2 +-
kernel/bpf/verifier.c | 44 ++++++++++++++++++++++++------------
kernel/trace/bpf_trace.c | 6 ++---
7 files changed, 54 insertions(+), 34 deletions(-)
diff --git a/Documentation/bpf/kfuncs.rst b/Documentation/bpf/kfuncs.rst
index f393c3c3d3b4..c27663cc6cc5 100644
--- a/Documentation/bpf/kfuncs.rst
+++ b/Documentation/bpf/kfuncs.rst
@@ -164,21 +164,22 @@ suffix should be used.
2.3.3 __uninit Annotation
-------------------------
-Use ``__uninit`` on a pointer parameter for an output that the kfunc
-initializes without reading its incoming contents.
+Use ``__uninit`` on a pointer parameter for an output buffer whose incoming
+contents the kfunc does not read.
-For generic memory buffers, the kfunc must initialize every byte in the
-declared range on every return path, including error returns and struct
-padding. The range is determined by the pointed-to type or the associated
-``__sz`` or ``__szk`` size argument.
+For generic memory buffers, the kfunc may leave bytes untouched, including
+on error returns. The writable range is determined by the pointed-to type
+or the associated ``__sz`` or ``__szk`` size argument.
The annotation does not change the accepted pointer types. A stack-backed
struct passed as a generic memory buffer must still be scalar-only.
-A stack buffer with a verifier-known constant offset and size may be
-uninitialized before the call and is considered initialized afterwards.
-For variable offsets or sizes, the usual stack-initialization and
-variable-offset restrictions still apply.
+Generic stack buffers may be uninitialized before the call. The call does
+not make previously uninitialized bytes readable unless the program is
+allowed to read uninitialized stack memory (normally requiring
+``CAP_PERFMON``). Other callers must initialize those bytes themselves
+before reading them. Stack bounds and variable-offset restrictions still
+apply.
For dynptr parameters, ``__uninit`` indicates that the kfunc constructs a
dynptr in the supplied storage. For example::
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 5033b934ffd9..fd22db8bc6c5 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -781,7 +781,10 @@ enum bpf_type_flag {
*/
PTR_UNTRUSTED = BIT(6 + BPF_BASE_TYPE_BITS),
- /* MEM can be uninitialized. */
+ /*
+ * MEM can be uninitialized. Generic memory outputs need not be fully
+ * initialized by the callee.
+ */
MEM_UNINIT = BIT(7 + BPF_BASE_TYPE_BITS),
/* DYNPTR points to memory local to the bpf program. */
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 9ddbb20ec1e9..92f528c45605 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1547,12 +1547,14 @@ struct ref_obj_desc {
};
/*
- * Memory arguments a call fills in, indexed by argument slot. The verifier allows the
- * stack to be uninitialized if the range is a known constant. Stack slots are marked as
- * STACK_MISC by check_mem_access() after all arguments have been checked.
+ * Generic MEM_UNINIT arguments, indexed by ABI slot. var_size_mask excludes
+ * variable-sized buffers from raw mode without losing the output annotation.
+ * size records constant ranges to mark initialized after checking all arguments,
+ * only when the caller is allowed to read uninitialized stack memory.
*/
struct arg_raw_mem_desc {
u16 mask;
+ u16 var_size_mask;
int size[MAX_BPF_FUNC_ARGS];
};
diff --git a/kernel/bpf/cgroup.c b/kernel/bpf/cgroup.c
index 696b27383974..1cb5e6a6ffc1 100644
--- a/kernel/bpf/cgroup.c
+++ b/kernel/bpf/cgroup.c
@@ -2491,7 +2491,7 @@ static const struct bpf_func_proto bpf_sysctl_get_name_proto = {
.gpl_only = false,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
- .arg2_type = ARG_PTR_TO_MEM | MEM_WRITE,
+ .arg2_type = ARG_PTR_TO_MEM | MEM_UNINIT | MEM_WRITE,
.arg3_type = ARG_MEM_SIZE,
.arg4_type = ARG_ANYTHING,
};
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 82402d97ce67..501c7ce35cba 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -1130,7 +1130,7 @@ const struct bpf_func_proto bpf_snprintf_proto = {
.func = bpf_snprintf,
.gpl_only = true,
.ret_type = RET_INTEGER,
- .arg1_type = ARG_PTR_TO_MEM_OR_NULL | MEM_WRITE,
+ .arg1_type = ARG_PTR_TO_MEM_OR_NULL | MEM_UNINIT | MEM_WRITE,
.arg2_type = ARG_MEM_SIZE_OR_ZERO,
.arg3_type = ARG_PTR_TO_CONST_STR,
.arg4_type = ARG_PTR_TO_MEM | PTR_MAYBE_NULL | MEM_RDONLY,
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 0c94f1214cc3..7f1cc115456f 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -6993,10 +6993,11 @@ static int check_stack_range_initialized(
* but BTF based global subprog validation isn't accurate enough.
*/
bool allow_poison = access_size < 0 || clobber;
- /* The call will initialize the memory; uninitialized stack allowed */
u32 arg_slot = arg_slot_from_argno(argno);
- bool raw_mode = meta && arg_slot < MAX_BPF_FUNC_ARGS &&
- (meta->arg_raw_mem.mask & BIT(arg_slot));
+ bool uninit = clobber && meta && arg_slot < MAX_BPF_FUNC_ARGS &&
+ (meta->arg_raw_mem.mask & BIT(arg_slot));
+ bool raw_mode = uninit && env->allow_uninit_stack &&
+ !(meta->arg_raw_mem.var_size_mask & BIT(arg_slot));
access_size = abs(access_size);
@@ -7035,6 +7036,7 @@ static int check_stack_range_initialized(
max_off = reg_smax(reg) + off;
}
+ /* Unprivileged outputs retain each byte's initialization state. */
if (raw_mode) {
meta->arg_raw_mem.size[arg_slot] = access_size;
return 0;
@@ -7054,8 +7056,8 @@ static int check_stack_range_initialized(
if (*stype == STACK_MISC)
goto mark;
if ((*stype == STACK_ZERO) ||
- (*stype == STACK_INVALID && env->allow_uninit_stack)) {
- if (clobber) {
+ (*stype == STACK_INVALID && (uninit || env->allow_uninit_stack))) {
+ if (clobber && (*stype != STACK_INVALID || env->allow_uninit_stack)) {
/* helper can write anything into the stack */
*stype = STACK_MISC;
}
@@ -7074,8 +7076,11 @@ static int check_stack_range_initialized(
}
if (*stype == STACK_POISON) {
- if (allow_poison)
+ if (allow_poison) {
+ if (uninit && env->allow_uninit_stack)
+ *stype = STACK_MISC;
goto mark;
+ }
verbose(env, "reading from stack %s off %d+%d size %d, slot poisoned by dead code elimination\n",
reg_arg_name(env, argno), min_off, i - min_off, access_size);
} else if (tnum_is_const(reg->var_off)) {
@@ -7224,12 +7229,12 @@ static int check_mem_size_reg(struct bpf_verifier_env *env,
meta->msize_max_value = reg_umax(size_reg);
/*
- * A variable size does not guarantee that the call initializes the whole
- * checked range. Disable raw mode for this output and apply the ordinary
- * stack initialization checks, including their privilege exceptions.
+ * Check variable ranges byte by byte instead of using raw mode. Keep the
+ * MEM_UNINIT annotation so invalid bytes are accepted without marking them
+ * initialized when the caller cannot read uninitialized stack memory.
*/
if (!tnum_is_const(size_reg->var_off))
- meta->arg_raw_mem.mask &= ~BIT(arg_slot_from_argno(mem_argno));
+ meta->arg_raw_mem.var_size_mask |= BIT(arg_slot_from_argno(mem_argno));
if (reg_smin(size_reg) < 0) {
verbose(env, "%s min value is negative, either use unsigned or 'var &= const'\n",
@@ -13730,12 +13735,20 @@ s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn
struct bpf_insn_aux_data *aux = &env->insn_aux_data[insn_idx];
const struct bpf_func_proto *fn;
enum bpf_arg_type at;
+ bool full_write;
s64 size;
if (bpf_get_helper_proto(env, insn->imm, &fn) < 0)
return S64_MIN;
at = fn->arg_type[arg];
+ /*
+ * Generic outputs may leave bytes untouched. Keep prior initialization
+ * live when the caller cannot read uninitialized bytes. Constructors of
+ * special objects, such as dynptrs, still define their storage.
+ */
+ full_write = (at & MEM_UNINIT) &&
+ (!arg_type_is_raw_mem(at) || env->allow_uninit_stack);
switch (base_type(at)) {
case ARG_PTR_TO_MAP_KEY:
@@ -13804,7 +13817,7 @@ s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn
* Size arg is const on each path but differs across merged
* paths. MAX_BPF_STACK is a safe upper bound for reads.
*/
- if (at & MEM_UNINIT)
+ if (full_write)
return 0;
return MAX_BPF_STACK;
}
@@ -13824,10 +13837,10 @@ s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn
}
out:
/*
- * MEM_UNINIT args are write-only: the helper initializes the
- * buffer without reading it.
+ * Other accesses keep the previous state live, including untouched bytes
+ * of an unprivileged generic output.
*/
- if (at & MEM_UNINIT)
+ if (full_write)
return -size;
return size;
}
@@ -13907,7 +13920,8 @@ s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn *
/* KF_ITER_NEW kfuncs initialize the iterator state at arg 0 */
if (arg == 0 && meta.kfunc_flags & KF_ITER_NEW)
return -size;
- if (is_kfunc_arg_uninit(btf, &args[i]))
+ if (is_kfunc_arg_uninit(btf, &args[i]) &&
+ (is_kfunc_arg_dynptr(btf, &args[i]) || env->allow_uninit_stack))
return -size;
return size;
}
diff --git a/kernel/trace/bpf_trace.c b/kernel/trace/bpf_trace.c
index 29260951aa87..195f78db9bda 100644
--- a/kernel/trace/bpf_trace.c
+++ b/kernel/trace/bpf_trace.c
@@ -995,7 +995,7 @@ static const struct bpf_func_proto bpf_d_path_proto = {
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_BTF_ID,
.arg1_btf_id = &bpf_d_path_btf_ids[0],
- .arg2_type = ARG_PTR_TO_MEM | MEM_WRITE,
+ .arg2_type = ARG_PTR_TO_MEM | MEM_UNINIT | MEM_WRITE,
.arg3_type = ARG_MEM_SIZE_OR_ZERO,
.allowed = bpf_d_path_allowed,
};
@@ -1052,7 +1052,7 @@ const struct bpf_func_proto bpf_snprintf_btf_proto = {
.func = bpf_snprintf_btf,
.gpl_only = false,
.ret_type = RET_INTEGER,
- .arg1_type = ARG_PTR_TO_MEM | MEM_WRITE,
+ .arg1_type = ARG_PTR_TO_MEM | MEM_UNINIT | MEM_WRITE,
.arg2_type = ARG_MEM_SIZE,
.arg3_type = ARG_PTR_TO_MEM | MEM_RDONLY,
.arg4_type = ARG_MEM_SIZE,
@@ -1565,7 +1565,7 @@ static const struct bpf_func_proto bpf_read_branch_records_proto = {
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
- .arg2_type = ARG_PTR_TO_MEM_OR_NULL | MEM_WRITE,
+ .arg2_type = ARG_PTR_TO_MEM_OR_NULL | MEM_UNINIT | MEM_WRITE,
.arg3_type = ARG_MEM_SIZE_OR_ZERO,
.arg4_type = ARG_ANYTHING,
};
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH bpf-next v5 08/11] bpf: Preserve stack initialization for generic output buffers
2026-09-21 2:38 ` [PATCH bpf-next v5 08/11] bpf: Preserve stack initialization for generic output buffers Kumar Kartikeya Dwivedi
@ 2026-09-21 3:54 ` bot+bpf-ci
0 siblings, 0 replies; 17+ messages in thread
From: bot+bpf-ci @ 2026-09-21 3:54 UTC (permalink / raw)
To: memxor, bpf
Cc: eddyz87, ast, andrii, daniel, emil, tj, ameryhung, kkd,
kernel-team, ast, andrii, daniel, martin.lau, eddyz87,
yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 6582 bytes --]
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 0c94f1214cc33..7f1cc115456f4 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
> @@ -6993,10 +6993,11 @@ static int check_stack_range_initialized(
> * but BTF based global subprog validation isn't accurate enough.
> */
> bool allow_poison = access_size < 0 || clobber;
> - /* The call will initialize the memory; uninitialized stack allowed */
> u32 arg_slot = arg_slot_from_argno(argno);
> - bool raw_mode = meta && arg_slot < MAX_BPF_FUNC_ARGS &&
> - (meta->arg_raw_mem.mask & BIT(arg_slot));
> + bool uninit = clobber && meta && arg_slot < MAX_BPF_FUNC_ARGS &&
> + (meta->arg_raw_mem.mask & BIT(arg_slot));
> + bool raw_mode = uninit && env->allow_uninit_stack &&
> + !(meta->arg_raw_mem.var_size_mask & BIT(arg_slot));
Can gating raw_mode on env->allow_uninit_stack cause unprivileged
programs to miss the fastcall contract check?
Before this commit, a generic MEM_UNINIT argument with a constant size
always took raw mode and was later written byte by byte by
mark_raw_stack():
kernel/bpf/verifier.c:mark_raw_stack() {
for (i = 0; i < MAX_BPF_FUNC_ARGS; i++) {
if (!meta->arg_raw_mem.size[slot])
continue;
...
err = check_mem_access(env, insn_idx, reg, argno, i, BPF_B,
BPF_WRITE, -1, false, false);
}
}
This check_mem_access() call reaches check_stack_write_fixed_off(),
which calls check_fastcall_stack_contract() to disable the fastcall
rewrite for stack ranges touched by the output buffer. That call to
check_fastcall_stack_contract() is the only mechanism that disables
the fastcall rewrite for output buffers, since
check_stack_range_initialized() itself never calls it.
After this commit, raw_mode is false whenever env->allow_uninit_stack
is false (which is true only for CAP_PERFMON). The byte loop below then
accepts STACK_INVALID slots but leaves them STACK_INVALID, and never
sets arg_raw_mem.size[slot], so mark_raw_stack() skips the slot
entirely and no stack access is ever recorded for the output range.
Meanwhile, mark_fastcall_patterns() still recognizes a fastcall pattern
from BPF_STX stores to R10, and subprog->fastcall_stack_off is set from
them. If a helper output buffer is the only thing touching a stack
region, that region contributes nothing to lowest_off, so the fastcall
pattern is still recognized.
Then check_stack_access_within_bounds() -> grow_stack_state() raises
env->subprog_info[].stack_depth to cover the buffer, but
bpf_remove_fastcall_spills_fills() later overwrites it:
kernel/bpf/fixups.c:bpf_remove_fastcall_spills_fills() {
if (modified && !subprog->keep_fastcall_stack)
subprog->stack_depth = -subprog->fastcall_stack_off;
}
and do_misc_fixups() propagates the shrunken value with
env->prog->aux->stack_depth = subprogs[0].stack_depth.
This means the JIT emits a frame of the reduced size while the helper
writes outside that frame. For example, a CAP_BPF | CAP_NET_ADMIN
program (without CAP_PERFMON) that does:
*(u64 *)(r10 - 8) = r1; /* fastcall pattern recognized */
call bpf_get_smp_processor_id;
r1 = *(u64 *)(r10 - 8);
r2 = r10;
r2 += -512; /* output buffer */
r3 = 64;
call bpf_skb_load_bytes; /* arg3 = ARG_PTR_TO_UNINIT_MEM */
would have stack_depth reach 512 during do_check(), but the contract
check never fires, so the rewrite is applied and stack_depth is reset
to -fastcall_stack_off. The JIT then emits a frame of that reduced
size, while bpf_skb_load_bytes() writes 64 bytes at r10-512, outside
the frame.
Does this produce kernel stack corruption from a non-CAP_PERFMON
program?
>
> access_size = abs(access_size);
>
> @@ -7035,6 +7036,7 @@ static int check_stack_range_initialized(
> max_off = reg_smax(reg) + off;
> }
>
> + /* Unprivileged outputs retain each byte's initialization state. */
> if (raw_mode) {
> meta->arg_raw_mem.size[arg_slot] = access_size;
> return 0;
This isn't a bug, but the comment describes the opposite of the code
path it precedes. raw_mode is only true when env->allow_uninit_stack
is set, which requires CAP_PERFMON, so this is the privileged path.
The block records the range so mark_raw_stack() can mark the whole
buffer STACK_MISC, which is the opposite of retaining each byte's
initialization state. The behaviour the comment describes belongs to
the unprivileged fall-through byte loop below.
> @@ -7054,8 +7056,8 @@ static int check_stack_range_initialized(
> if (*stype == STACK_MISC)
> goto mark;
> if ((*stype == STACK_ZERO) ||
> - (*stype == STACK_INVALID && env->allow_uninit_stack)) {
> - if (clobber) {
> + (*stype == STACK_INVALID && (uninit || env->allow_uninit_stack))) {
> + if (clobber && (*stype != STACK_INVALID || env->allow_uninit_stack)) {
> /* helper can write anything into the stack */
> *stype = STACK_MISC;
> }
Can unprivileged programs still read back buffers filled by helpers
that fully initialize them?
Because raw_mode now requires env->allow_uninit_stack, STACK_INVALID
slots of an output argument are accepted but left STACK_INVALID (the
clobber condition (*stype != STACK_INVALID || env->allow_uninit_stack)
is false), so any later load from the buffer fails with "invalid read
from stack ..." in check_stack_read_fixed_off().
This change affects helpers that write every byte they are given, not
just partial writers. For example, bpf_skb_load_bytes() with a constant
size writes every byte, yet its output is now treated as
possibly-uninitialized for unprivileged callers.
The commit message acknowledges this: "This also changes readback from
existing full-writing helpers: without permission to read uninitialized
stack memory, programs must initialize those bytes themselves before
reading them after a call."
However, the stated motivation is partial-output helpers such as
bpf_snprintf() that may leave bytes untouched. Those helpers are newly
annotated with MEM_UNINIT by this same commit and previously required
a pre-initialized buffer anyway, so for them the change is a pure
relaxation.
Would distinguishing partial writers from full writers (e.g. a separate
flag) avoid breaking existing unprivileged users of
bpf_skb_load_bytes() and similar helpers?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35555832550
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH bpf-next v5 09/11] selftests/bpf: Cover generic output stack initialization
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
` (7 preceding siblings ...)
2026-09-21 2:38 ` [PATCH bpf-next v5 08/11] bpf: Preserve stack initialization for generic output buffers Kumar Kartikeya Dwivedi
@ 2026-09-21 2:38 ` Kumar Kartikeya Dwivedi
2026-09-21 3:54 ` bot+bpf-ci
2026-09-21 2:38 ` [PATCH bpf-next v5 10/11] bpf: Check read access for helper input/output buffers Kumar Kartikeya Dwivedi
` (2 subsequent siblings)
11 siblings, 1 reply; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-21 2:38 UTC (permalink / raw)
To: bpf
Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
Eduard Zingerman, Emil Tsalapatis, Tejun Heo, Amery Hung, kkd,
kernel-team
Exercise generic output buffers with and without CAP_PERFMON, using both
helpers and __uninit kfuncs. Check that initialized bytes stay readable
and lose stale value information, while invalid bytes remain unreadable
without permission to read uninitialized stack memory. Cover constant and
variable sizes, scalar spills, pointer spills, special stack objects, and
privileged variable offsets.
Add a kfunc that writes only the first byte of its output. Its runtime tests
read preinitialized bytes, checking both the written byte and an untouched
tail byte across a liveness checkpoint. Verify that the sysctl name helper
and uninitialized fixed and variable-sized kfunc outputs are accepted when
their contents are not read back.
Update existing __uninit readback expectations and exercise skb_load_bytes
with reduced capabilities. Even fully-writing generic outputs now preserve
invalid bytes in the verifier, so reading those bytes requires prior
initialization by the BPF program.
Use map_update_elem inputs for helper_arg_fallback_keeps_scanning. Its
original snprintf argument no longer reads the buffer, and a privileged
output with an unknown size does not trigger the whole-stack read fallback.
A variable-offset key and parent-frame value retain the intended assertion.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
.../progs/verifier_helper_access_var_len.c | 189 ++++++++++++++++++
.../bpf/progs/verifier_kfunc_uninit.c | 136 +++++++++++++
.../bpf/progs/verifier_kfunc_uninit_multi.c | 6 +-
.../selftests/bpf/progs/verifier_live_stack.c | 33 ++-
.../selftests/bpf/progs/verifier_raw_stack.c | 4 +
.../selftests/bpf/test_kmods/bpf_testmod.c | 7 +
.../bpf/test_kmods/bpf_testmod_kfunc.h | 1 +
7 files changed, 356 insertions(+), 20 deletions(-)
diff --git a/tools/testing/selftests/bpf/progs/verifier_helper_access_var_len.c b/tools/testing/selftests/bpf/progs/verifier_helper_access_var_len.c
index d1452ef6f2f9..e87eb6f221e4 100644
--- a/tools/testing/selftests/bpf/progs/verifier_helper_access_var_len.c
+++ b/tools/testing/selftests/bpf/progs/verifier_helper_access_var_len.c
@@ -822,4 +822,193 @@ __naked void bytes_no_leak_init_memory(void)
: __clobber_all);
}
+SEC("cgroup/sysctl")
+__success
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__success_unpriv
+__naked void sysctl_initialized_stack(void)
+{
+ asm volatile (
+ "*(u64 *)(r10 - 16) = 0;"
+ "*(u64 *)(r10 - 8) = 0;"
+ "r2 = r10;"
+ "r2 += -16;"
+ "r3 = 16;"
+ "r4 = 0;"
+ "call %[bpf_sysctl_get_name];"
+ "r0 = *(u8 *)(r10 - 1);"
+ "r0 &= 1;"
+ "exit;"
+ : : __imm(bpf_sysctl_get_name) : __clobber_all);
+}
+
+SEC("cgroup/sysctl")
+__success
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__success_unpriv
+__naked void sysctl_uninitialized_stack(void)
+{
+ asm volatile (
+ "r2 = r10;"
+ "r2 += -16;"
+ "r3 = 16;"
+ "r4 = 0;"
+ "call %[bpf_sysctl_get_name];"
+ "r0 = 0;"
+ "exit;"
+ : : __imm(bpf_sysctl_get_name) : __clobber_all);
+}
+
+SEC("cgroup/sysctl")
+__success
+__flag(BPF_F_TEST_STATE_FREQ)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__success_unpriv
+__naked void sysctl_partial_initialized_bytes(void)
+{
+ asm volatile (
+ "*(u32 *)(r10 - 16) = 0;"
+ "*(u8 *)(r10 - 1) = 1;"
+ "goto +0;"
+ "r2 = r10;"
+ "r2 += -16;"
+ "r3 = 16;"
+ "r4 = 0;"
+ "call %[bpf_sysctl_get_name];"
+ "r0 = *(u32 *)(r10 - 16);"
+ "r1 = *(u8 *)(r10 - 1);"
+ "r0 += r1;"
+ "r0 &= 1;"
+ "exit;"
+ : : __imm(bpf_sysctl_get_name) : __clobber_all);
+}
+
+SEC("cgroup/sysctl")
+__success
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__failure_unpriv __msg_unpriv("invalid read from stack off -1+0 size 1")
+__naked void sysctl_partial_invalid_bytes(void)
+{
+ asm volatile (
+ "*(u32 *)(r10 - 16) = 0;"
+ "r2 = r10;"
+ "r2 += -16;"
+ "r3 = 16;"
+ "r4 = 0;"
+ "call %[bpf_sysctl_get_name];"
+ "r0 = *(u8 *)(r10 - 1);"
+ "r0 &= 1;"
+ "exit;"
+ : : __imm(bpf_sysctl_get_name) : __clobber_all);
+}
+
+SEC("cgroup/sysctl")
+__success
+__flag(BPF_F_TEST_STATE_FREQ)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__success_unpriv
+__naked void sysctl_partial_variable_size(void)
+{
+ asm volatile (
+ "r3 = *(u32 *)(r1 + 0);"
+ "r3 &= 15;"
+ "r3 += 1;"
+ "*(u8 *)(r10 - 1) = 1;"
+ "goto +0;"
+ "r2 = r10;"
+ "r2 += -16;"
+ "r4 = 0;"
+ "call %[bpf_sysctl_get_name];"
+ "r0 = *(u8 *)(r10 - 1);"
+ "r0 &= 1;"
+ "exit;"
+ : : __imm(bpf_sysctl_get_name) : __clobber_all);
+}
+
+SEC("cgroup/sysctl")
+__success
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__failure_unpriv __msg_unpriv("invalid read from stack off -1+0 size 1")
+__naked void sysctl_partial_variable_invalid(void)
+{
+ asm volatile (
+ "r3 = *(u32 *)(r1 + 0);"
+ "r3 &= 15;"
+ "r3 += 1;"
+ "r2 = r10;"
+ "r2 += -16;"
+ "r4 = 0;"
+ "call %[bpf_sysctl_get_name];"
+ "r0 = *(u8 *)(r10 - 1);"
+ "r0 &= 1;"
+ "exit;"
+ : : __imm(bpf_sysctl_get_name) : __clobber_all);
+}
+
+SEC("cgroup/sysctl")
+__success
+__flag(BPF_F_TEST_STATE_FREQ)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__failure_unpriv __msg_unpriv("invalid read from stack R2")
+__naked void sysctl_partial_spilled_pointer(void)
+{
+ asm volatile (
+ "*(u64 *)(r10 - 8) = r1;"
+ "goto +0;"
+ "r2 = r10;"
+ "r2 += -8;"
+ "r3 = 8;"
+ "r4 = 0;"
+ "call %[bpf_sysctl_get_name];"
+ "r0 = 0;"
+ "exit;"
+ : : __imm(bpf_sysctl_get_name) : __clobber_all);
+}
+
+SEC("cgroup/sysctl")
+__success
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__failure_unpriv __msg_unpriv("invalid read from stack R2")
+int sysctl_partial_dynptr(struct bpf_sysctl *ctx)
+{
+ struct bpf_dynptr ptr;
+ long long key = 0, *data;
+
+ data = bpf_map_lookup_elem(&map_hash_8b, &key);
+ if (!data)
+ return 0;
+ bpf_dynptr_from_mem(data, sizeof(*data), 0, &ptr);
+ bpf_sysctl_get_name(ctx, (char *)&ptr, sizeof(ptr), 0);
+ return 0;
+}
+
+SEC("cgroup/sysctl")
+__success
+__naked void sysctl_partial_variable_offset(void)
+{
+ asm volatile (
+ "r2 = *(u32 *)(r1 + 0);"
+ "r2 &= 8;"
+ "r2 += r10;"
+ "r2 += -24;"
+ "r3 = 16;"
+ "r4 = 0;"
+ "call %[bpf_sysctl_get_name];"
+ "r0 = *(u8 *)(r10 - 16);"
+ "r0 &= 1;"
+ "exit;"
+ : : __imm(bpf_sysctl_get_name) : __clobber_all);
+}
+
+SEC("tc")
+__success __retval(42)
+int snprintf_partial_output_runtime(struct __sk_buff *ctx)
+{
+ char buf[16];
+
+ buf[15] = 42;
+ bpf_snprintf(buf, sizeof(buf), "ok", NULL, 0);
+ return buf[15];
+}
+
char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit.c b/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit.c
index f7818303e703..ff2fb36a0260 100644
--- a/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit.c
+++ b/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit.c
@@ -12,6 +12,7 @@ void __kfunc_btf_root(void)
asm volatile ("" :
: "r"(&bpf_kfunc_test_uninit_struct),
"r"(&bpf_kfunc_test_uninit_mem),
+ "r"(&bpf_kfunc_test_uninit_partial),
"r"(&bpf_kfunc_test_uninit_alias));
}
@@ -97,4 +98,139 @@ __naked void uninitialized_input_alias(void)
: : __imm(bpf_kfunc_test_uninit_alias) : __clobber_all);
}
+SEC("tc")
+__success __retval(0)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__success_unpriv
+__naked void uninitialized_struct_output_ignored(void)
+{
+ asm volatile (
+ "r1 = r10;"
+ "r1 += -16;"
+ "call %[bpf_kfunc_test_uninit_struct];"
+ "r0 = 0;"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_struct) : __clobber_all);
+}
+
+SEC("tc")
+__success __retval(0)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__success_unpriv
+__naked void variable_size_uninitialized_output_ignored(void)
+{
+ asm volatile (
+ "r2 = *(u32 *)(r1 + 0);"
+ "r2 &= 7;"
+ "r2 += 1;"
+ "r1 = r10;"
+ "r1 += -8;"
+ "call %[bpf_kfunc_test_uninit_mem];"
+ "r0 = 0;"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_mem) : __clobber_all);
+}
+
+SEC("tc")
+__success __retval(49)
+__flag(BPF_F_TEST_STATE_FREQ)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__success_unpriv
+__naked void partial_output_preserves_initialized_bytes(void)
+{
+ asm volatile (
+ "*(u8 *)(r10 - 8) = 0;"
+ "*(u8 *)(r10 - 1) = 7;"
+ "goto +0;"
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = 8;"
+ "call %[bpf_kfunc_test_uninit_partial];"
+ "r0 = *(u8 *)(r10 - 8);"
+ "r1 = *(u8 *)(r10 - 1);"
+ "r0 += r1;"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_partial) : __clobber_all);
+}
+
+SEC("tc")
+__success __retval(42)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__failure_unpriv __msg_unpriv("invalid read from stack off -8+0 size 1")
+__naked void uninitialized_sized_output_read(void)
+{
+ asm volatile (
+ "r1 = r10;"
+ "r1 += -8;"
+ "r2 = 8;"
+ "call %[bpf_kfunc_test_uninit_mem];"
+ "r0 = *(u8 *)(r10 - 8);"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_mem) : __clobber_all);
+}
+
+SEC("tc")
+__success __retval(42)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__failure_unpriv __msg_unpriv("invalid read from stack off -8+0 size 1")
+__naked void variable_size_uninitialized_output_read(void)
+{
+ asm volatile (
+ "r2 = *(u32 *)(r1 + 0);"
+ "r2 &= 7;"
+ "r2 += 1;"
+ "r1 = r10;"
+ "r1 += -8;"
+ "call %[bpf_kfunc_test_uninit_mem];"
+ "r0 = *(u8 *)(r10 - 8);"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_mem) : __clobber_all);
+}
+
+SEC("tc")
+__success __retval(1)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__failure_unpriv __msg_unpriv("invalid read from stack off -16+0 size 4")
+__naked void fixed_struct_uninitialized_output_read(void)
+{
+ asm volatile (
+ "r1 = r10;"
+ "r1 += -16;"
+ "call %[bpf_kfunc_test_uninit_struct];"
+ "r0 = *(u32 *)(r10 - 16);"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_struct) : __clobber_all);
+}
+
+SEC("tc")
+__success __retval(49)
+__flag(BPF_F_TEST_STATE_FREQ)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__prepare_priv
+__success_unpriv
+__naked void partial_output_variable_size(void)
+{
+ asm volatile (
+ "r2 = *(u32 *)(r1 + 0);"
+ "r2 &= 7;"
+ "r2 += 1;"
+ "*(u8 *)(r10 - 8) = 0;"
+ "*(u8 *)(r10 - 1) = 7;"
+ "goto +0;"
+ "r1 = r10;"
+ "r1 += -8;"
+ "call %[bpf_kfunc_test_uninit_partial];"
+ "r0 = *(u8 *)(r10 - 8);"
+ "r1 = *(u8 *)(r10 - 1);"
+ "r0 += r1;"
+ "exit;"
+ : : __imm(bpf_kfunc_test_uninit_partial) : __clobber_all);
+}
+
char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit_multi.c b/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit_multi.c
index 18ced0d0a3b2..ab3bb81a11fd 100644
--- a/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit_multi.c
+++ b/tools/testing/selftests/bpf/progs/verifier_kfunc_uninit_multi.c
@@ -49,7 +49,7 @@ SEC("tc")
__success __retval(42)
__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
__prepare_priv
-__success_unpriv
+__failure_unpriv __msg_unpriv("invalid read from stack off -16+0 size 4")
__naked void variable_size_preserves_other_output(void)
{
asm volatile (
@@ -71,7 +71,7 @@ __arch_x86_64 __arch_arm64
__success __retval(42)
__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
__prepare_priv
-__success_unpriv
+__failure_unpriv __msg_unpriv("invalid read from stack off -8+0 size 4")
__naked void output_after_by_value_argument(void)
{
asm volatile (
@@ -91,7 +91,7 @@ __arch_x86_64 __arch_arm64 __arch_riscv64
__success __retval(15)
__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
__prepare_priv
-__success_unpriv
+__failure_unpriv __msg_unpriv("invalid read from stack off -8+0 size 4")
__naked void output_passed_on_stack(void)
{
asm volatile (
diff --git a/tools/testing/selftests/bpf/progs/verifier_live_stack.c b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
index bc3dfdc1a536..7a1a0670f851 100644
--- a/tools/testing/selftests/bpf/progs/verifier_live_stack.c
+++ b/tools/testing/selftests/bpf/progs/verifier_live_stack.c
@@ -21,8 +21,6 @@ struct {
__type(value, __u64);
} array_map_8b SEC(".maps");
-const char snprintf_u64_fmt[] = "%llu";
-
SEC("socket")
__log_level(2)
__msg("0: (79) r1 = *(u64 *)(r10 -8) ; use: fp0-8")
@@ -1947,15 +1945,15 @@ static __used __naked void fwd_parent_key_to_helper(void)
/*
* Regression for keeping later helper args after a whole-stack fallback
- * on an earlier local arg. The first bpf_snprintf() arg is a local
+ * on an earlier local arg. The bpf_map_update_elem() key is a local
* frame-derived pointer with offset-imprecise tracking (`fp1 ?`), which
- * conservatively marks the whole local stack live. The fourth arg still
+ * conservatively marks the whole local stack live. The value arg still
* forwards &parent_fp-8 and must contribute nonlocal_use[0]=0:3.
*/
SEC("socket")
__log_level(2)
__success
-__msg("call bpf_snprintf{{.*}} ; use: fp1-8..-512 fp0-8")
+__msg("call bpf_map_update_elem{{.*}}; use: fp1-8..-512 fp0-8")
__naked void helper_arg_fallback_keeps_scanning(void)
{
asm volatile (
@@ -1963,32 +1961,33 @@ __naked void helper_arg_fallback_keeps_scanning(void)
"*(u64 *)(r10 - 8) = r1;"
"r1 = r10;"
"r1 += -8;"
- "call helper_snprintf_parent_after_local_fallback;"
+ "call helper_update_parent_after_local_fallback;"
"r0 = 0;"
"exit;"
::: __clobber_all);
}
-static __used __naked void helper_snprintf_parent_after_local_fallback(void)
+static __used __naked void helper_update_parent_after_local_fallback(void)
{
asm volatile (
"r6 = r1;" /* save &parent_fp-8 */
"call %[bpf_get_prandom_u32];"
"r0 &= 8;"
- "r1 = r10;"
- "r1 += -16;"
- "r1 += r0;" /* local fp, offset-imprecise */
- "r2 = 8;"
- "r3 = %[snprintf_u64_fmt] ll;"
- "r4 = r6;" /* later arg: parent fp-8 */
- "r5 = 8;"
- "call %[bpf_snprintf];"
+ "*(u64 *)(r10 - 16) = 0;"
+ "*(u64 *)(r10 - 8) = 0;"
+ "r2 = r10;"
+ "r2 += -16;"
+ "r2 += r0;" /* local fp, offset-imprecise */
+ "r1 = %[array_map_8b] ll;"
+ "r3 = r6;" /* later arg: parent fp-8 */
+ "r4 = 0;"
+ "call %[bpf_map_update_elem];"
"r0 = 0;"
"exit;"
:
: __imm(bpf_get_prandom_u32),
- __imm(bpf_snprintf),
- __imm_addr(snprintf_u64_fmt)
+ __imm(bpf_map_update_elem),
+ __imm_addr(array_map_8b)
: __clobber_all);
}
diff --git a/tools/testing/selftests/bpf/progs/verifier_raw_stack.c b/tools/testing/selftests/bpf/progs/verifier_raw_stack.c
index c689665e07b9..9f0f48ecb421 100644
--- a/tools/testing/selftests/bpf/progs/verifier_raw_stack.c
+++ b/tools/testing/selftests/bpf/progs/verifier_raw_stack.c
@@ -84,6 +84,8 @@ __naked void skb_load_bytes_zero_len(void)
SEC("tc")
__description("raw_stack: skb_load_bytes, no init")
__success __retval(0)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__failure_unpriv __msg_unpriv("invalid read from stack off -8+0 size 8")
__naked void skb_load_bytes_no_init(void)
{
asm volatile (" \
@@ -103,6 +105,8 @@ __naked void skb_load_bytes_no_init(void)
SEC("tc")
__description("raw_stack: skb_load_bytes, init")
__success __retval(0)
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__success_unpriv
__naked void stack_skb_load_bytes_init(void)
{
asm volatile (" \
diff --git a/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c b/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
index 211886a8ee87..93847ca6293b 100644
--- a/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
+++ b/tools/testing/selftests/bpf/test_kmods/bpf_testmod.c
@@ -1330,6 +1330,12 @@ __bpf_kfunc void bpf_kfunc_test_uninit_mem(void *out__uninit, u32 out__sz)
memset(out__uninit, 0x2a, out__sz);
}
+__bpf_kfunc void bpf_kfunc_test_uninit_partial(void *out__uninit, u32 out__sz)
+{
+ if (out__sz)
+ *(u8 *)out__uninit = 42;
+}
+
__bpf_kfunc int bpf_kfunc_test_uninit_alias(int *out__uninit, const int *in)
{
int value = get_unaligned(in);
@@ -1788,6 +1794,7 @@ BTF_ID_FLAGS(func, bpf_kfunc_call_test_pass1)
BTF_ID_FLAGS(func, bpf_kfunc_call_test_pass2)
BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_struct)
BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_mem)
+BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_partial)
BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_alias)
BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_multi)
BTF_ID_FLAGS(func, bpf_kfunc_test_uninit_pair)
diff --git a/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h b/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h
index 524f2cb9bdf4..67c02a421d13 100644
--- a/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h
+++ b/tools/testing/selftests/bpf/test_kmods/bpf_testmod_kfunc.h
@@ -291,6 +291,7 @@ void bpf_kfunc_call_test_pass1(struct prog_test_pass1 *p) __ksym;
void bpf_kfunc_call_test_pass2(struct prog_test_pass2 *p) __ksym;
void bpf_kfunc_test_uninit_struct(struct prog_test_pass1 *out__uninit) __ksym;
void bpf_kfunc_test_uninit_mem(void *out__uninit, __u32 out__sz) __ksym;
+void bpf_kfunc_test_uninit_partial(void *out__uninit, __u32 out__sz) __ksym;
int bpf_kfunc_test_uninit_alias(int *out__uninit, const int *in) __ksym;
void bpf_kfunc_test_uninit_multi(int *a__uninit, void *b__uninit, __u32 b__sz) __ksym;
void bpf_kfunc_test_uninit_pair(struct prog_test_pair_arg p, int *out__uninit) __ksym;
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH bpf-next v5 09/11] selftests/bpf: Cover generic output stack initialization
2026-09-21 2:38 ` [PATCH bpf-next v5 09/11] selftests/bpf: Cover generic output stack initialization Kumar Kartikeya Dwivedi
@ 2026-09-21 3:54 ` bot+bpf-ci
0 siblings, 0 replies; 17+ messages in thread
From: bot+bpf-ci @ 2026-09-21 3:54 UTC (permalink / raw)
To: memxor, bpf
Cc: ast, andrii, daniel, eddyz87, emil, tj, ameryhung, kkd,
kernel-team, ast, andrii, daniel, martin.lau, eddyz87,
yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 2499 bytes --]
> selftests/bpf: Cover generic output stack initialization
>
> Exercise generic output buffers with and without CAP_PERFMON, using both
> helpers and __uninit kfuncs. Check that initialized bytes stay readable
> and lose stale value information, while invalid bytes remain unreadable
> without permission to read uninitialized stack memory. Cover constant and
> variable sizes, scalar spills, pointer spills, special stack objects, and
> privileged variable offsets.
The changelog says the tests cover "scalar spills", but does the commit
actually add a scalar spill test?
Looking through all seven files, the only register spill added is the
pointer spill in sysctl_partial_spilled_pointer():
*(u64 *)(r10 - 8) = r1; /* r1 == PTR_TO_CTX */
Every other added stack store is an immediate write (*(u32 *)(r10 - 16) = 0;,
*(u8 *)(r10 - 1) = 1;, etc.), which the verifier records as STACK_ZERO or
STACK_MISC, not STACK_SPILL.
The other four items in that sentence each have a matching test: constant
sizes (sysctl_partial_initialized_bytes, sysctl_partial_invalid_bytes),
variable sizes (sysctl_partial_variable_size, sysctl_partial_variable_size
_invalid), special stack objects (sysctl_partial_dynptr), and privileged
variable offsets (sysctl_partial_variable_offset). The missing "scalar
spills" case stands out.
The scalar-spill arm in check_stack_range_initialized() is a distinct
branch:
if (bpf_is_spilled_reg(&state->stack[spi]) &&
(state->stack[spi].spilled_ptr.type == SCALAR_VALUE ||
env->allow_ptr_leaks)) {
if (clobber) {
__mark_reg_unknown(env, &state->stack[spi].spilled_ptr);
for (j = 0; j < BPF_REG_SIZE; j++)
scrub_spilled_slot(&state->stack[spi].slot_type[j]);
}
goto mark;
}
A SCALAR_VALUE spill is accepted even without CAP_PERFMON and is scrubbed
to an unknown scalar, so the readback-after-scalar-spill behaviour is not
pinned by any test in this series.
Could you add a test that spills a scalar register into the output buffer
(for example, r1 = 1; *(u64 *)(r10 - 8) = r1; before the call, then read
the slot back in both the priv and unpriv arms), or drop "scalar spills"
from the changelog?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35555832550
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH bpf-next v5 10/11] bpf: Check read access for helper input/output buffers
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
` (8 preceding siblings ...)
2026-09-21 2:38 ` [PATCH bpf-next v5 09/11] selftests/bpf: Cover generic output stack initialization Kumar Kartikeya Dwivedi
@ 2026-09-21 2:38 ` Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 11/11] selftests/bpf: Cover helper memory access permissions Kumar Kartikeya Dwivedi
2026-09-21 17:20 ` [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers patchwork-bot+netdevbpf
11 siblings, 0 replies; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-21 2:38 UTC (permalink / raw)
To: bpf
Cc: Eduard Zingerman, Alexei Starovoitov, Andrii Nakryiko,
Daniel Borkmann, Emil Tsalapatis, Tejun Heo, Amery Hung, kkd,
kernel-team
From: Eduard Zingerman <eddyz87@gmail.com>
MEM_WRITE without MEM_UNINIT denotes memory that is read as well as
written. Helper argument checking requests only BPF_WRITE for it, which
checks stack initialization but omits read permission checks on map values.
For example, bpf_check_mtu() reads its mtu_len argument before overwriting
it, but the verifier permits that argument to point into a BPF_F_WRONLY_PROG
map. bpf_fib_lookup() and bpf_load_hdr_opt() are affected the same way.
Derive generic memory access from argument flags: read-only for inputs,
write-only for MEM_WRITE | MEM_UNINIT, and read/write for MEM_WRITE alone.
Partial-output helpers such as bpf_snprintf() now carry MEM_UNINIT, so
their write-only map destinations remain valid. Generated kfunc prototypes
already encode the same distinction, which retires the kfunc-specific
access override.
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
kernel/bpf/verifier.c | 22 +++++++++++++++-------
1 file changed, 15 insertions(+), 7 deletions(-)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 7f1cc115456f..d62c0f74cff5 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -8907,6 +8907,19 @@ static int process_map_ptr_arg(struct bpf_verifier_env *env, struct bpf_reg_stat
return 0;
}
+/*
+ * MEM_WRITE alone denotes an input/output buffer. MEM_UNINIT marks an output
+ * whose incoming contents are not read.
+ */
+static enum bpf_access_type func_arg_access_type(enum bpf_arg_type arg_type)
+{
+ if (!(arg_type & MEM_WRITE))
+ return BPF_READ;
+ if (arg_type & MEM_UNINIT)
+ return BPF_WRITE;
+ return BPF_READ | BPF_WRITE;
+}
+
static int check_func_arg(struct bpf_verifier_env *env, u32 arg, u32 slot, u32 prev_slot,
struct bpf_call_arg_meta *meta,
int insn_idx)
@@ -9199,10 +9212,7 @@ static int check_func_arg(struct bpf_verifier_env *env, u32 arg, u32 slot, u32 p
if (!(arg_type & MEM_FIXED_SIZE))
break;
- access_type = arg_type & MEM_WRITE ? BPF_WRITE : BPF_READ;
- /* Ordinary kfunc buffers are input/output; __uninit buffers are outputs. */
- if (meta->btf && !(arg_type & MEM_UNINIT))
- access_type = BPF_READ | BPF_WRITE;
+ access_type = func_arg_access_type(arg_type);
err = check_mem_reg(env, reg, argno, arg_size, access_type, meta, &known_memory);
if (err < 0) {
@@ -9248,9 +9258,7 @@ static int check_func_arg(struct bpf_verifier_env *env, u32 arg, u32 slot, u32 p
if (meta->btf && bpf_register_is_null(buff_reg))
break;
- access_type = fn->arg_type[arg - 1] & MEM_WRITE ? BPF_WRITE : BPF_READ;
- if (meta->btf && !(fn->arg_type[arg - 1] & MEM_UNINIT))
- access_type = BPF_READ | BPF_WRITE;
+ access_type = func_arg_access_type(fn->arg_type[arg - 1]);
zero_size_allowed = meta->btf || base_type(arg_type) == ARG_MEM_SIZE_OR_ZERO;
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* [PATCH bpf-next v5 11/11] selftests/bpf: Cover helper memory access permissions
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
` (9 preceding siblings ...)
2026-09-21 2:38 ` [PATCH bpf-next v5 10/11] bpf: Check read access for helper input/output buffers Kumar Kartikeya Dwivedi
@ 2026-09-21 2:38 ` Kumar Kartikeya Dwivedi
2026-09-21 3:54 ` bot+bpf-ci
2026-09-21 17:20 ` [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers patchwork-bot+netdevbpf
11 siblings, 1 reply; 17+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-21 2:38 UTC (permalink / raw)
To: bpf
Cc: Eduard Zingerman, Alexei Starovoitov, Andrii Nakryiko,
Daniel Borkmann, Emil Tsalapatis, Tejun Heo, Amery Hung, kkd,
kernel-team
From: Eduard Zingerman <eddyz87@gmail.com>
Check that fixed-size and sized helper input/output buffers require both
read and write permission. Cover the MTU and FIB helpers, including read-only
and write-only rejection and read/write positive controls. Exercise the XDP
prototypes and the sock_ops header-option input/output argument as well.
Keep write-only maps usable as destinations for partial-output helpers
bpf_snprintf() and bpf_sysctl_get_name(), while rejecting a read-only
snprintf destination. Also retain a write-only-map destination for
bpf_get_current_comm(), whose output is annotated MEM_UNINIT and fully
initialized by the helper.
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
.../progs/verifier_helper_access_var_len.c | 143 ++++++++++++++++++
.../selftests/bpf/progs/verifier_mtu.c | 88 +++++++++++
2 files changed, 231 insertions(+)
diff --git a/tools/testing/selftests/bpf/progs/verifier_helper_access_var_len.c b/tools/testing/selftests/bpf/progs/verifier_helper_access_var_len.c
index e87eb6f221e4..2782faf0a528 100644
--- a/tools/testing/selftests/bpf/progs/verifier_helper_access_var_len.c
+++ b/tools/testing/selftests/bpf/progs/verifier_helper_access_var_len.c
@@ -1,4 +1,5 @@
// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
/* Converted from tools/testing/selftests/bpf/verifier/helper_access_var_len.c */
#include <linux/bpf.h>
@@ -822,6 +823,148 @@ __naked void bytes_no_leak_init_memory(void)
: __clobber_all);
}
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __uint(map_flags, BPF_F_WRONLY_PROG);
+ __type(key, __u32);
+ __type(value, struct bpf_fib_lookup);
+} map_fib_wo SEC(".maps");
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __type(key, __u32);
+ __type(value, struct bpf_fib_lookup);
+} map_fib_rw SEC(".maps");
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __uint(map_flags, BPF_F_RDONLY_PROG);
+ __type(key, __u32);
+ __type(value, struct bpf_fib_lookup);
+} map_fib_ro SEC(".maps");
+
+SEC("tc")
+__failure __msg("read from map forbidden")
+int writeonly_sized_input(struct __sk_buff *ctx)
+{
+ struct bpf_fib_lookup *params;
+ __u32 key = 0;
+
+ params = bpf_map_lookup_elem(&map_fib_wo, &key);
+ if (params)
+ bpf_fib_lookup(ctx, params, sizeof(*params), 0);
+ return 0;
+}
+
+SEC("tc")
+__success
+int readwrite_sized_input(struct __sk_buff *ctx)
+{
+ struct bpf_fib_lookup *params;
+ __u32 key = 0;
+
+ params = bpf_map_lookup_elem(&map_fib_rw, &key);
+ if (params)
+ bpf_fib_lookup(ctx, params, sizeof(*params), 0);
+ return 0;
+}
+
+SEC("tc")
+__failure __msg("write into map forbidden")
+int readonly_sized_output(struct __sk_buff *ctx)
+{
+ struct bpf_fib_lookup *params;
+ __u32 key = 0;
+
+ params = bpf_map_lookup_elem(&map_fib_ro, &key);
+ if (params)
+ bpf_fib_lookup(ctx, params, sizeof(*params), 0);
+ return 0;
+}
+
+SEC("xdp")
+__failure __msg("read from map forbidden")
+int xdp_writeonly_sized_input(struct xdp_md *ctx)
+{
+ struct bpf_fib_lookup *params;
+ __u32 key = 0;
+
+ params = bpf_map_lookup_elem(&map_fib_wo, &key);
+ if (params)
+ bpf_fib_lookup(ctx, params, sizeof(*params), 0);
+ return XDP_PASS;
+}
+
+SEC("sockops")
+__failure __msg("read from map forbidden")
+int writeonly_header_option(struct bpf_sock_ops *ctx)
+{
+ struct bpf_fib_lookup *buf;
+ __u32 key = 0;
+
+ buf = bpf_map_lookup_elem(&map_fib_wo, &key);
+ if (buf)
+ bpf_load_hdr_opt(ctx, buf, sizeof(*buf), 0);
+ return 0;
+}
+
+SEC("sockops")
+__success
+int readwrite_header_option(struct bpf_sock_ops *ctx)
+{
+ struct bpf_fib_lookup *buf;
+ __u32 key = 0;
+
+ buf = bpf_map_lookup_elem(&map_fib_rw, &key);
+ if (buf)
+ bpf_load_hdr_opt(ctx, buf, sizeof(*buf), 0);
+ return 0;
+}
+
+SEC("tc")
+__success
+int snprintf_writeonly_output(struct __sk_buff *ctx)
+{
+ void *buf;
+ __u32 key = 0;
+
+ buf = bpf_map_lookup_elem(&map_fib_wo, &key);
+ if (buf)
+ bpf_snprintf(buf, 16, "ok", NULL, 0);
+ return 0;
+}
+
+SEC("tc")
+__failure __msg("write into map forbidden")
+int snprintf_readonly_output(struct __sk_buff *ctx)
+{
+ void *buf;
+ __u32 key = 0;
+
+ buf = bpf_map_lookup_elem(&map_fib_ro, &key);
+ if (buf)
+ bpf_snprintf(buf, 16, "ok", NULL, 0);
+ return 0;
+}
+
+SEC("cgroup/sysctl")
+__success
+__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
+__success_unpriv
+int sysctl_writeonly_output(struct bpf_sysctl *ctx)
+{
+ void *buf;
+ __u32 key = 0;
+
+ buf = bpf_map_lookup_elem(&map_fib_wo, &key);
+ if (buf)
+ bpf_sysctl_get_name(ctx, buf, 16, 0);
+ return 0;
+}
+
SEC("cgroup/sysctl")
__success
__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
diff --git a/tools/testing/selftests/bpf/progs/verifier_mtu.c b/tools/testing/selftests/bpf/progs/verifier_mtu.c
index 256956ea1ac5..2f71f70302e4 100644
--- a/tools/testing/selftests/bpf/progs/verifier_mtu.c
+++ b/tools/testing/selftests/bpf/progs/verifier_mtu.c
@@ -17,4 +17,92 @@ int tc_uninit_mtu(struct __sk_buff *ctx)
return TCX_PASS;
}
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __uint(map_flags, BPF_F_WRONLY_PROG);
+ __type(key, __u32);
+ __type(value, __u32);
+} map_mtu_wo SEC(".maps");
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __type(key, __u32);
+ __type(value, __u32);
+} map_mtu_rw SEC(".maps");
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARRAY);
+ __uint(max_entries, 1);
+ __uint(map_flags, BPF_F_RDONLY_PROG);
+ __type(key, __u32);
+ __type(value, __u32);
+} map_mtu_ro SEC(".maps");
+
+SEC("tc/ingress")
+__failure __msg("read from map forbidden")
+int tc_writeonly_mtu(struct __sk_buff *ctx)
+{
+ __u32 key = 0;
+ __u32 *mtu;
+
+ mtu = bpf_map_lookup_elem(&map_mtu_wo, &key);
+ if (mtu)
+ bpf_check_mtu(ctx, 0, mtu, 0, 0);
+ return TCX_PASS;
+}
+
+SEC("tc/ingress")
+__success
+int tc_readwrite_mtu(struct __sk_buff *ctx)
+{
+ __u32 key = 0;
+ __u32 *mtu;
+
+ mtu = bpf_map_lookup_elem(&map_mtu_rw, &key);
+ if (mtu)
+ bpf_check_mtu(ctx, 0, mtu, 0, 0);
+ return TCX_PASS;
+}
+
+SEC("xdp")
+__failure __msg("read from map forbidden")
+int xdp_writeonly_mtu(struct xdp_md *ctx)
+{
+ __u32 key = 0;
+ __u32 *mtu;
+
+ mtu = bpf_map_lookup_elem(&map_mtu_wo, &key);
+ if (mtu)
+ bpf_check_mtu(ctx, 0, mtu, 0, 0);
+ return XDP_PASS;
+}
+
+SEC("tc/ingress")
+__failure __msg("write into map forbidden")
+int tc_readonly_mtu(struct __sk_buff *ctx)
+{
+ __u32 key = 0;
+ __u32 *mtu;
+
+ mtu = bpf_map_lookup_elem(&map_mtu_ro, &key);
+ if (mtu)
+ bpf_check_mtu(ctx, 0, mtu, 0, 0);
+ return TCX_PASS;
+}
+
+SEC("tc/ingress")
+__success
+int tc_writeonly_output(struct __sk_buff *ctx)
+{
+ __u32 key = 0;
+ __u32 *out;
+
+ out = bpf_map_lookup_elem(&map_mtu_wo, &key);
+ if (out)
+ bpf_get_current_comm(out, sizeof(*out));
+ return TCX_PASS;
+}
+
char LICENSE[] SEC("license") = "GPL";
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread* Re: [PATCH bpf-next v5 11/11] selftests/bpf: Cover helper memory access permissions
2026-09-21 2:38 ` [PATCH bpf-next v5 11/11] selftests/bpf: Cover helper memory access permissions Kumar Kartikeya Dwivedi
@ 2026-09-21 3:54 ` bot+bpf-ci
0 siblings, 0 replies; 17+ messages in thread
From: bot+bpf-ci @ 2026-09-21 3:54 UTC (permalink / raw)
To: memxor, bpf
Cc: eddyz87, ast, andrii, daniel, emil, tj, ameryhung, kkd,
kernel-team, ast, andrii, daniel, martin.lau, eddyz87,
yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 5365 bytes --]
> Check that fixed-size and sized helper input/output buffers require both
> read and write permission. Cover the MTU and FIB helpers, including read-only
> and write-only rejection and read/write positive controls. Exercise the XDP
> prototypes and the sock_ops header-option input/output argument as well.
>
> Keep write-only maps usable as destinations for partial-output helpers
> bpf_snprintf() and bpf_sysctl_get_name(), while rejecting a read-only
> snprintf destination. Also retain a write-only-map destination for
> bpf_get_current_comm(), whose output is annotated MEM_UNINIT and fully
> initialized by the helper.
This isn't a bug, but could this sentence drop the "fully initialized by the
helper" part, since after the earlier patch in the series the write-only
destination follows from the MEM_UNINIT annotation rather than from full
initialization?
>
> Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
> Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
> diff --git a/tools/testing/selftests/bpf/progs/verifier_helper_access_var_len.c b/tools/testing/selftests/bpf/progs/verifier_helper_access_var_len.c
> index e87eb6f221e48..2782faf0a528b 100644
> --- a/tools/testing/selftests/bpf/progs/verifier_helper_access_var_len.c
> +++ b/tools/testing/selftests/bpf/progs/verifier_helper_access_var_len.c
> @@ -1,4 +1,5 @@
> // SPDX-License-Identifier: GPL-2.0
> +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
> /* Converted from tools/testing/selftests/bpf/verifier/helper_access_var_len.c */
>
> #include <linux/bpf.h>
> @@ -822,6 +823,148 @@ __naked void bytes_no_leak_init_memory(void)
> : __clobber_all);
> }
>
> +struct {
> + __uint(type, BPF_MAP_TYPE_ARRAY);
> + __uint(max_entries, 1);
> + __uint(map_flags, BPF_F_WRONLY_PROG);
> + __type(key, __u32);
> + __type(value, struct bpf_fib_lookup);
> +} map_fib_wo SEC(".maps");
> +
> +struct {
> + __uint(type, BPF_MAP_TYPE_ARRAY);
> + __uint(max_entries, 1);
> + __type(key, __u32);
> + __type(value, struct bpf_fib_lookup);
> +} map_fib_rw SEC(".maps");
> +
> +struct {
> + __uint(type, BPF_MAP_TYPE_ARRAY);
> + __uint(max_entries, 1);
> + __uint(map_flags, BPF_F_RDONLY_PROG);
> + __type(key, __u32);
> + __type(value, struct bpf_fib_lookup);
> +} map_fib_ro SEC(".maps");
This isn't a bug, but would a neutral name and value type (for example a
plain char[64] scratch buffer named map_wo/map_rw/map_ro) read better, given
these maps are also used for bpf_load_hdr_opt(), bpf_snprintf() and
bpf_sysctl_get_name()?
> +
> +SEC("tc")
> +__failure __msg("read from map forbidden")
> +int writeonly_sized_input(struct __sk_buff *ctx)
> +{
> + struct bpf_fib_lookup *params;
> + __u32 key = 0;
> +
> + params = bpf_map_lookup_elem(&map_fib_wo, &key);
> + if (params)
> + bpf_fib_lookup(ctx, params, sizeof(*params), 0);
> + return 0;
> +}
[ ... ]
> +SEC("sockops")
> +__failure __msg("read from map forbidden")
> +int writeonly_header_option(struct bpf_sock_ops *ctx)
> +{
> + struct bpf_fib_lookup *buf;
> + __u32 key = 0;
> +
> + buf = bpf_map_lookup_elem(&map_fib_wo, &key);
> + if (buf)
> + bpf_load_hdr_opt(ctx, buf, sizeof(*buf), 0);
> + return 0;
> +}
> +
> +SEC("sockops")
> +__success
> +int readwrite_header_option(struct bpf_sock_ops *ctx)
> +{
> + struct bpf_fib_lookup *buf;
> + __u32 key = 0;
> +
> + buf = bpf_map_lookup_elem(&map_fib_rw, &key);
> + if (buf)
> + bpf_load_hdr_opt(ctx, buf, sizeof(*buf), 0);
> + return 0;
> +}
> +
> +SEC("tc")
> +__success
> +int snprintf_writeonly_output(struct __sk_buff *ctx)
> +{
> + void *buf;
> + __u32 key = 0;
> +
> + buf = bpf_map_lookup_elem(&map_fib_wo, &key);
> + if (buf)
> + bpf_snprintf(buf, 16, "ok", NULL, 0);
> + return 0;
> +}
> +
> +SEC("tc")
> +__failure __msg("write into map forbidden")
> +int snprintf_readonly_output(struct __sk_buff *ctx)
> +{
> + void *buf;
> + __u32 key = 0;
> +
> + buf = bpf_map_lookup_elem(&map_fib_ro, &key);
> + if (buf)
> + bpf_snprintf(buf, 16, "ok", NULL, 0);
> + return 0;
> +}
> +
> +SEC("cgroup/sysctl")
> +__success
> +__caps_unpriv(CAP_BPF | CAP_NET_ADMIN)
> +__success_unpriv
> +int sysctl_writeonly_output(struct bpf_sysctl *ctx)
> +{
> + void *buf;
> + __u32 key = 0;
> +
> + buf = bpf_map_lookup_elem(&map_fib_wo, &key);
> + if (buf)
> + bpf_sysctl_get_name(ctx, buf, 16, 0);
> + return 0;
> +}
[ ... ]
> diff --git a/tools/testing/selftests/bpf/progs/verifier_mtu.c b/tools/testing/selftests/bpf/progs/verifier_mtu.c
> index 256956ea1ac51..2f71f70302e48 100644
> --- a/tools/testing/selftests/bpf/progs/verifier_mtu.c
> +++ b/tools/testing/selftests/bpf/progs/verifier_mtu.c
[ ... ]
> +SEC("tc/ingress")
> +__success
> +int tc_writeonly_output(struct __sk_buff *ctx)
> +{
> + __u32 key = 0;
> + __u32 *out;
> +
> + out = bpf_map_lookup_elem(&map_mtu_wo, &key);
> + if (out)
> + bpf_get_current_comm(out, sizeof(*out));
> + return TCX_PASS;
> +}
This isn't a bug, but would this bpf_get_current_comm() case be easier to
find next to the snprintf/sysctl output tests in
verifier_helper_access_var_len.c, or at least under a name that mentions the
helper?
> +
> char LICENSE[] SEC("license") = "GPL";
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35555832550
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
` (10 preceding siblings ...)
2026-09-21 2:38 ` [PATCH bpf-next v5 11/11] selftests/bpf: Cover helper memory access permissions Kumar Kartikeya Dwivedi
@ 2026-09-21 17:20 ` patchwork-bot+netdevbpf
11 siblings, 0 replies; 17+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-09-21 17:20 UTC (permalink / raw)
To: Kumar Kartikeya Dwivedi
Cc: bpf, ast, andrii, daniel, eddyz87, emil, tj, ameryhung, kkd,
kernel-team
Hello:
This series was applied to bpf/bpf-next.git (master)
by Alexei Starovoitov <ast@kernel.org>:
On Mon, 21 Sep 2026 04:38:24 +0200 you wrote:
> Generic __uninit kfunc arguments are output buffers. Stack liveness treats
> them as writes, but argument checking still requires readable contents and
> does not record definite initialization after the call. Check these
> arguments as write-only and record privileged output initialization after
> validating all inputs, including inputs that alias an output.
>
> Following Eduard's rework, helpers and kfuncs record generic outputs in the
> same argument-checking path after type resolution. Generated kfunc
> prototypes now mark generic buffers with MEM_WRITE, so __uninit buffers are
> checked as write-only ahead of the fix while ordinary buffers stay
> read/write.
>
> [...]
Here is the summary with links:
- [bpf-next,v5,01/11] selftests/bpf: Allow privileged preparation for capability tests
https://git.kernel.org/bpf/bpf-next/c/06046a8a50e8
- [bpf-next,v5,02/11] bpf: Record raw memory arguments during argument checking
https://git.kernel.org/bpf/bpf-next/c/e2b4aaa75040
- [bpf-next,v5,03/11] bpf: Check __uninit kfunc output buffers as write-only
https://git.kernel.org/bpf/bpf-next/c/fc8dc2be101b
- [bpf-next,v5,04/11] bpf: Fix generic __uninit kfunc output buffers
https://git.kernel.org/bpf/bpf-next/c/fc670d4b6c31
- [bpf-next,v5,05/11] selftests/bpf: Cover generic __uninit output initialization
https://git.kernel.org/bpf/bpf-next/c/a8bbd9ee1bbc
- [bpf-next,v5,06/11] bpf: Support multiple __uninit kfunc output arguments
https://git.kernel.org/bpf/bpf-next/c/bca730ba6df6
- [bpf-next,v5,07/11] selftests/bpf: Cover __uninit kfunc output argument slots
https://git.kernel.org/bpf/bpf-next/c/89835699ef4e
- [bpf-next,v5,08/11] bpf: Preserve stack initialization for generic output buffers
https://git.kernel.org/bpf/bpf-next/c/5da4a9f26fca
- [bpf-next,v5,09/11] selftests/bpf: Cover generic output stack initialization
https://git.kernel.org/bpf/bpf-next/c/a25a61385b04
- [bpf-next,v5,10/11] bpf: Check read access for helper input/output buffers
https://git.kernel.org/bpf/bpf-next/c/7a6cef39b760
- [bpf-next,v5,11/11] selftests/bpf: Cover helper memory access permissions
https://git.kernel.org/bpf/bpf-next/c/ad36589c3f39
You are awesome, thank you!
--
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html
^ permalink raw reply [flat|nested] 17+ messages in thread