BPF List
 help / color / mirror / Atom feed
From: Kumar Kartikeya Dwivedi <memxor@gmail.com>
To: bpf@vger.kernel.org
Cc: Eduard Zingerman <eddyz87@gmail.com>,
	Alexei Starovoitov <ast@kernel.org>,
	Andrii Nakryiko <andrii@kernel.org>,
	Daniel Borkmann <daniel@iogearbox.net>,
	Emil Tsalapatis <emil@etsalapatis.com>, Tejun Heo <tj@kernel.org>,
	Amery Hung <ameryhung@gmail.com>,
	kkd@meta.com, kernel-team@meta.com
Subject: [PATCH bpf-next v5 08/11] bpf: Preserve stack initialization for generic output buffers
Date: Mon, 21 Sep 2026 04:38:32 +0200	[thread overview]
Message-ID: <20260921023843.411943-9-memxor@gmail.com> (raw)
In-Reply-To: <20260921023843.411943-1-memxor@gmail.com>

Partial-output helpers such as bpf_snprintf() do not read incoming buffer
contents, but may leave some bytes untouched. Let them accept uninitialized
storage without promising full initialization to callers that cannot read
uninitialized stack memory.

Make this the default for generic MEM_UNINIT buffers, including __uninit
kfunc arguments. Allow invalid stack bytes through the output check but
leave them invalid when uninitialized stack reads are not allowed. Scrub
initialized bytes and scalar spills as possible writes, retaining the
existing restrictions on spilled pointers and special stack objects.

Retain the output annotation for variable-sized arguments and track their
raw-mode eligibility separately. Callers allowed uninitialized stack reads
can continue treating the potentially written range as initialized. For
constant ranges, defer that initialization until all inputs are checked.

Keep prior stack contents live for generic outputs when the caller cannot
read uninitialized stack memory. Such calls do not define the entire range,
so liveness must preserve initialization facts that remain relevant after
the call. Dynptr and iterator constructors still define their storage.

Annotate the snprintf, sysctl name, d_path, snprintf_btf and branch-record
destinations with MEM_UNINIT, and document that generic __uninit kfuncs may
leave bytes untouched. This also changes readback from existing full-writing
helpers: without permission to read uninitialized stack memory, programs
must initialize those bytes themselves before reading them after a call.

Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 Documentation/bpf/kfuncs.rst | 21 +++++++++--------
 include/linux/bpf.h          |  5 +++-
 include/linux/bpf_verifier.h |  8 ++++---
 kernel/bpf/cgroup.c          |  2 +-
 kernel/bpf/helpers.c         |  2 +-
 kernel/bpf/verifier.c        | 44 ++++++++++++++++++++++++------------
 kernel/trace/bpf_trace.c     |  6 ++---
 7 files changed, 54 insertions(+), 34 deletions(-)

diff --git a/Documentation/bpf/kfuncs.rst b/Documentation/bpf/kfuncs.rst
index f393c3c3d3b4..c27663cc6cc5 100644
--- a/Documentation/bpf/kfuncs.rst
+++ b/Documentation/bpf/kfuncs.rst
@@ -164,21 +164,22 @@ suffix should be used.
 2.3.3 __uninit Annotation
 -------------------------
 
-Use ``__uninit`` on a pointer parameter for an output that the kfunc
-initializes without reading its incoming contents.
+Use ``__uninit`` on a pointer parameter for an output buffer whose incoming
+contents the kfunc does not read.
 
-For generic memory buffers, the kfunc must initialize every byte in the
-declared range on every return path, including error returns and struct
-padding. The range is determined by the pointed-to type or the associated
-``__sz`` or ``__szk`` size argument.
+For generic memory buffers, the kfunc may leave bytes untouched, including
+on error returns. The writable range is determined by the pointed-to type
+or the associated ``__sz`` or ``__szk`` size argument.
 
 The annotation does not change the accepted pointer types. A stack-backed
 struct passed as a generic memory buffer must still be scalar-only.
 
-A stack buffer with a verifier-known constant offset and size may be
-uninitialized before the call and is considered initialized afterwards.
-For variable offsets or sizes, the usual stack-initialization and
-variable-offset restrictions still apply.
+Generic stack buffers may be uninitialized before the call. The call does
+not make previously uninitialized bytes readable unless the program is
+allowed to read uninitialized stack memory (normally requiring
+``CAP_PERFMON``). Other callers must initialize those bytes themselves
+before reading them. Stack bounds and variable-offset restrictions still
+apply.
 
 For dynptr parameters, ``__uninit`` indicates that the kfunc constructs a
 dynptr in the supplied storage. For example::
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 5033b934ffd9..fd22db8bc6c5 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -781,7 +781,10 @@ enum bpf_type_flag {
 	 */
 	PTR_UNTRUSTED		= BIT(6 + BPF_BASE_TYPE_BITS),
 
-	/* MEM can be uninitialized. */
+	/*
+	 * MEM can be uninitialized. Generic memory outputs need not be fully
+	 * initialized by the callee.
+	 */
 	MEM_UNINIT		= BIT(7 + BPF_BASE_TYPE_BITS),
 
 	/* DYNPTR points to memory local to the bpf program. */
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 9ddbb20ec1e9..92f528c45605 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1547,12 +1547,14 @@ struct ref_obj_desc {
 };
 
 /*
- * Memory arguments a call fills in, indexed by argument slot. The verifier allows the
- * stack to be uninitialized if the range is a known constant. Stack slots are marked as
- * STACK_MISC by check_mem_access() after all arguments have been checked.
+ * Generic MEM_UNINIT arguments, indexed by ABI slot. var_size_mask excludes
+ * variable-sized buffers from raw mode without losing the output annotation.
+ * size records constant ranges to mark initialized after checking all arguments,
+ * only when the caller is allowed to read uninitialized stack memory.
  */
 struct arg_raw_mem_desc {
 	u16 mask;
+	u16 var_size_mask;
 	int size[MAX_BPF_FUNC_ARGS];
 };
 
diff --git a/kernel/bpf/cgroup.c b/kernel/bpf/cgroup.c
index 696b27383974..1cb5e6a6ffc1 100644
--- a/kernel/bpf/cgroup.c
+++ b/kernel/bpf/cgroup.c
@@ -2491,7 +2491,7 @@ static const struct bpf_func_proto bpf_sysctl_get_name_proto = {
 	.gpl_only	= false,
 	.ret_type	= RET_INTEGER,
 	.arg1_type	= ARG_PTR_TO_CTX,
-	.arg2_type	= ARG_PTR_TO_MEM | MEM_WRITE,
+	.arg2_type	= ARG_PTR_TO_MEM | MEM_UNINIT | MEM_WRITE,
 	.arg3_type	= ARG_MEM_SIZE,
 	.arg4_type	= ARG_ANYTHING,
 };
diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index 82402d97ce67..501c7ce35cba 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -1130,7 +1130,7 @@ const struct bpf_func_proto bpf_snprintf_proto = {
 	.func		= bpf_snprintf,
 	.gpl_only	= true,
 	.ret_type	= RET_INTEGER,
-	.arg1_type	= ARG_PTR_TO_MEM_OR_NULL | MEM_WRITE,
+	.arg1_type	= ARG_PTR_TO_MEM_OR_NULL | MEM_UNINIT | MEM_WRITE,
 	.arg2_type	= ARG_MEM_SIZE_OR_ZERO,
 	.arg3_type	= ARG_PTR_TO_CONST_STR,
 	.arg4_type	= ARG_PTR_TO_MEM | PTR_MAYBE_NULL | MEM_RDONLY,
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 0c94f1214cc3..7f1cc115456f 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -6993,10 +6993,11 @@ static int check_stack_range_initialized(
 	 * but BTF based global subprog validation isn't accurate enough.
 	 */
 	bool allow_poison = access_size < 0 || clobber;
-	/* The call will initialize the memory; uninitialized stack allowed */
 	u32 arg_slot = arg_slot_from_argno(argno);
-	bool raw_mode = meta && arg_slot < MAX_BPF_FUNC_ARGS &&
-		       (meta->arg_raw_mem.mask & BIT(arg_slot));
+	bool uninit = clobber && meta && arg_slot < MAX_BPF_FUNC_ARGS &&
+		      (meta->arg_raw_mem.mask & BIT(arg_slot));
+	bool raw_mode = uninit && env->allow_uninit_stack &&
+			!(meta->arg_raw_mem.var_size_mask & BIT(arg_slot));
 
 	access_size = abs(access_size);
 
@@ -7035,6 +7036,7 @@ static int check_stack_range_initialized(
 		max_off = reg_smax(reg) + off;
 	}
 
+	/* Unprivileged outputs retain each byte's initialization state. */
 	if (raw_mode) {
 		meta->arg_raw_mem.size[arg_slot] = access_size;
 		return 0;
@@ -7054,8 +7056,8 @@ static int check_stack_range_initialized(
 		if (*stype == STACK_MISC)
 			goto mark;
 		if ((*stype == STACK_ZERO) ||
-		    (*stype == STACK_INVALID && env->allow_uninit_stack)) {
-			if (clobber) {
+		    (*stype == STACK_INVALID && (uninit || env->allow_uninit_stack))) {
+			if (clobber && (*stype != STACK_INVALID || env->allow_uninit_stack)) {
 				/* helper can write anything into the stack */
 				*stype = STACK_MISC;
 			}
@@ -7074,8 +7076,11 @@ static int check_stack_range_initialized(
 		}
 
 		if (*stype == STACK_POISON) {
-			if (allow_poison)
+			if (allow_poison) {
+				if (uninit && env->allow_uninit_stack)
+					*stype = STACK_MISC;
 				goto mark;
+			}
 			verbose(env, "reading from stack %s off %d+%d size %d, slot poisoned by dead code elimination\n",
 				reg_arg_name(env, argno), min_off, i - min_off, access_size);
 		} else if (tnum_is_const(reg->var_off)) {
@@ -7224,12 +7229,12 @@ static int check_mem_size_reg(struct bpf_verifier_env *env,
 	meta->msize_max_value = reg_umax(size_reg);
 
 	/*
-	 * A variable size does not guarantee that the call initializes the whole
-	 * checked range. Disable raw mode for this output and apply the ordinary
-	 * stack initialization checks, including their privilege exceptions.
+	 * Check variable ranges byte by byte instead of using raw mode. Keep the
+	 * MEM_UNINIT annotation so invalid bytes are accepted without marking them
+	 * initialized when the caller cannot read uninitialized stack memory.
 	 */
 	if (!tnum_is_const(size_reg->var_off))
-		meta->arg_raw_mem.mask &= ~BIT(arg_slot_from_argno(mem_argno));
+		meta->arg_raw_mem.var_size_mask |= BIT(arg_slot_from_argno(mem_argno));
 
 	if (reg_smin(size_reg) < 0) {
 		verbose(env, "%s min value is negative, either use unsigned or 'var &= const'\n",
@@ -13730,12 +13735,20 @@ s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn
 	struct bpf_insn_aux_data *aux = &env->insn_aux_data[insn_idx];
 	const struct bpf_func_proto *fn;
 	enum bpf_arg_type at;
+	bool full_write;
 	s64 size;
 
 	if (bpf_get_helper_proto(env, insn->imm, &fn) < 0)
 		return S64_MIN;
 
 	at = fn->arg_type[arg];
+	/*
+	 * Generic outputs may leave bytes untouched. Keep prior initialization
+	 * live when the caller cannot read uninitialized bytes. Constructors of
+	 * special objects, such as dynptrs, still define their storage.
+	 */
+	full_write = (at & MEM_UNINIT) &&
+		     (!arg_type_is_raw_mem(at) || env->allow_uninit_stack);
 
 	switch (base_type(at)) {
 	case ARG_PTR_TO_MAP_KEY:
@@ -13804,7 +13817,7 @@ s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn
 			 * Size arg is const on each path but differs across merged
 			 * paths. MAX_BPF_STACK is a safe upper bound for reads.
 			 */
-			if (at & MEM_UNINIT)
+			if (full_write)
 				return 0;
 			return MAX_BPF_STACK;
 		}
@@ -13824,10 +13837,10 @@ s64 bpf_helper_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn
 	}
 out:
 	/*
-	 * MEM_UNINIT args are write-only: the helper initializes the
-	 * buffer without reading it.
+	 * Other accesses keep the previous state live, including untouched bytes
+	 * of an unprivileged generic output.
 	 */
-	if (at & MEM_UNINIT)
+	if (full_write)
 		return -size;
 	return size;
 }
@@ -13907,7 +13920,8 @@ s64 bpf_kfunc_stack_access_bytes(struct bpf_verifier_env *env, struct bpf_insn *
 	/* KF_ITER_NEW kfuncs initialize the iterator state at arg 0 */
 	if (arg == 0 && meta.kfunc_flags & KF_ITER_NEW)
 		return -size;
-	if (is_kfunc_arg_uninit(btf, &args[i]))
+	if (is_kfunc_arg_uninit(btf, &args[i]) &&
+	    (is_kfunc_arg_dynptr(btf, &args[i]) || env->allow_uninit_stack))
 		return -size;
 	return size;
 }
diff --git a/kernel/trace/bpf_trace.c b/kernel/trace/bpf_trace.c
index 29260951aa87..195f78db9bda 100644
--- a/kernel/trace/bpf_trace.c
+++ b/kernel/trace/bpf_trace.c
@@ -995,7 +995,7 @@ static const struct bpf_func_proto bpf_d_path_proto = {
 	.ret_type	= RET_INTEGER,
 	.arg1_type	= ARG_PTR_TO_BTF_ID,
 	.arg1_btf_id	= &bpf_d_path_btf_ids[0],
-	.arg2_type	= ARG_PTR_TO_MEM | MEM_WRITE,
+	.arg2_type	= ARG_PTR_TO_MEM | MEM_UNINIT | MEM_WRITE,
 	.arg3_type	= ARG_MEM_SIZE_OR_ZERO,
 	.allowed	= bpf_d_path_allowed,
 };
@@ -1052,7 +1052,7 @@ const struct bpf_func_proto bpf_snprintf_btf_proto = {
 	.func		= bpf_snprintf_btf,
 	.gpl_only	= false,
 	.ret_type	= RET_INTEGER,
-	.arg1_type	= ARG_PTR_TO_MEM | MEM_WRITE,
+	.arg1_type	= ARG_PTR_TO_MEM | MEM_UNINIT | MEM_WRITE,
 	.arg2_type	= ARG_MEM_SIZE,
 	.arg3_type	= ARG_PTR_TO_MEM | MEM_RDONLY,
 	.arg4_type	= ARG_MEM_SIZE,
@@ -1565,7 +1565,7 @@ static const struct bpf_func_proto bpf_read_branch_records_proto = {
 	.gpl_only       = true,
 	.ret_type       = RET_INTEGER,
 	.arg1_type      = ARG_PTR_TO_CTX,
-	.arg2_type      = ARG_PTR_TO_MEM_OR_NULL | MEM_WRITE,
+	.arg2_type      = ARG_PTR_TO_MEM_OR_NULL | MEM_UNINIT | MEM_WRITE,
 	.arg3_type      = ARG_MEM_SIZE_OR_ZERO,
 	.arg4_type      = ARG_ANYTHING,
 };
-- 
2.53.0


  parent reply	other threads:[~2026-09-21  2:39 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-21  2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
2026-09-21  2:38 ` [PATCH bpf-next v5 01/11] selftests/bpf: Allow privileged preparation for capability tests Kumar Kartikeya Dwivedi
2026-09-21  2:38 ` [PATCH bpf-next v5 02/11] bpf: Record raw memory arguments during argument checking Kumar Kartikeya Dwivedi
2026-09-21  2:38 ` [PATCH bpf-next v5 03/11] bpf: Check __uninit kfunc output buffers as write-only Kumar Kartikeya Dwivedi
2026-09-21  3:54   ` bot+bpf-ci
2026-09-21  2:38 ` [PATCH bpf-next v5 04/11] bpf: Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
2026-09-21  2:38 ` [PATCH bpf-next v5 05/11] selftests/bpf: Cover generic __uninit output initialization Kumar Kartikeya Dwivedi
2026-09-21  2:38 ` [PATCH bpf-next v5 06/11] bpf: Support multiple __uninit kfunc output arguments Kumar Kartikeya Dwivedi
2026-09-21  2:38 ` [PATCH bpf-next v5 07/11] selftests/bpf: Cover __uninit kfunc output argument slots Kumar Kartikeya Dwivedi
2026-09-21  2:38 ` Kumar Kartikeya Dwivedi [this message]
2026-09-21  3:54   ` [PATCH bpf-next v5 08/11] bpf: Preserve stack initialization for generic output buffers bot+bpf-ci
2026-09-21  2:38 ` [PATCH bpf-next v5 09/11] selftests/bpf: Cover generic output stack initialization Kumar Kartikeya Dwivedi
2026-09-21  3:54   ` bot+bpf-ci
2026-09-21  2:38 ` [PATCH bpf-next v5 10/11] bpf: Check read access for helper input/output buffers Kumar Kartikeya Dwivedi
2026-09-21  2:38 ` [PATCH bpf-next v5 11/11] selftests/bpf: Cover helper memory access permissions Kumar Kartikeya Dwivedi
2026-09-21  3:54   ` bot+bpf-ci
2026-09-21 17:20 ` [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers patchwork-bot+netdevbpf

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260921023843.411943-9-memxor@gmail.com \
    --to=memxor@gmail.com \
    --cc=ameryhung@gmail.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=emil@etsalapatis.com \
    --cc=kernel-team@meta.com \
    --cc=kkd@meta.com \
    --cc=tj@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox