From: Kumar Kartikeya Dwivedi <memxor@gmail.com>
To: bpf@vger.kernel.org
Cc: Tejun Heo <tj@kernel.org>, Eduard Zingerman <eddyz87@gmail.com>,
Alexei Starovoitov <ast@kernel.org>,
Andrii Nakryiko <andrii@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Emil Tsalapatis <emil@etsalapatis.com>,
Amery Hung <ameryhung@gmail.com>,
kkd@meta.com, kernel-team@meta.com
Subject: [PATCH bpf-next v5 04/11] bpf: Fix generic __uninit kfunc output buffers
Date: Mon, 21 Sep 2026 04:38:28 +0200 [thread overview]
Message-ID: <20260921023843.411943-5-memxor@gmail.com> (raw)
In-Reply-To: <20260921023843.411943-1-memxor@gmail.com>
Stack liveness treats __uninit kfunc arguments as writes. Even with
write-only access checks, ordinary clobber handling leaves poisoned stack
bytes poisoned after the call, so the output cannot be read.
Reuse the helper output descriptor for generic kfunc buffers. Record the
output after resolving its memory type, and initialize its stack bytes
only after all arguments have been checked. This prevents an output from
making an aliased, uninitialized input appear valid.
Track the output by its ABI argument slot, since a kfunc pointer may be
passed on the stack or follow a multi-slot by-value argument. Keep the
single-output restriction in prototype validation; multiple-output
tracking is a separate extension.
Document that output buffers must be fully initialized on every return
path, including error returns and padding.
Fixes: 2cb27158adb3 ("bpf: poison dead stack slots")
Reported-by: Tejun Heo <tj@kernel.org>
Link: https://lore.kernel.org/bpf/86d966ec88bbf27d21b2bb4e18c8aa00@kernel.org
Suggested-by: Eduard Zingerman <eddyz87@gmail.com>
Acked-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
Documentation/bpf/kfuncs.rst | 26 ++++++++----
include/linux/bpf_verifier.h | 7 ++--
kernel/bpf/verifier.c | 80 ++++++++++++++++++++++++++----------
3 files changed, 81 insertions(+), 32 deletions(-)
diff --git a/Documentation/bpf/kfuncs.rst b/Documentation/bpf/kfuncs.rst
index 3f300118a623..f393c3c3d3b4 100644
--- a/Documentation/bpf/kfuncs.rst
+++ b/Documentation/bpf/kfuncs.rst
@@ -164,19 +164,31 @@ suffix should be used.
2.3.3 __uninit Annotation
-------------------------
-This annotation is used to indicate that the argument will be treated as
-uninitialized.
+Use ``__uninit`` on a pointer parameter for an output that the kfunc
+initializes without reading its incoming contents.
-An example is given below::
+For generic memory buffers, the kfunc must initialize every byte in the
+declared range on every return path, including error returns and struct
+padding. The range is determined by the pointed-to type or the associated
+``__sz`` or ``__szk`` size argument.
+
+The annotation does not change the accepted pointer types. A stack-backed
+struct passed as a generic memory buffer must still be scalar-only.
+
+A stack buffer with a verifier-known constant offset and size may be
+uninitialized before the call and is considered initialized afterwards.
+For variable offsets or sizes, the usual stack-initialization and
+variable-offset restrictions still apply.
+
+For dynptr parameters, ``__uninit`` indicates that the kfunc constructs a
+dynptr in the supplied storage. For example::
- __bpf_kfunc int bpf_dynptr_from_skb(..., struct bpf_dynptr_kern *ptr__uninit)
+ __bpf_kfunc int bpf_dynptr_from_skb(..., struct bpf_dynptr *ptr__uninit)
{
...
}
-Here, the dynptr will be treated as an uninitialized dynptr. Without this
-annotation, the verifier will reject the program if the dynptr passed in is
-not initialized.
+Without this annotation, a dynptr argument must already be initialized.
2.3.4 __nullable Annotation
---------------------------
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index cf85141ea167..6ff1c227d298 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1548,10 +1548,11 @@ struct ref_obj_desc {
/*
* A memory argument a call fills in. The verifier allows the stack to be uninitialized if
- * the range is a known constant. Stack slots are marked as STACK_MISC by check_mem_access().
+ * the range is a known constant. Stack slots are marked as STACK_MISC by check_mem_access()
+ * after all arguments have been checked.
*/
struct arg_raw_mem_desc {
- u8 regno;
+ u8 argno; /* One-based ABI argument slot; zero means no output. */
int size;
};
@@ -1579,6 +1580,7 @@ struct bpf_call_arg_meta {
struct bpf_dynptr_desc dynptr;
struct ref_obj_desc ref_obj;
struct ret_mem_desc ret_mem;
+ struct arg_raw_mem_desc arg_raw_mem;
/* Only set by kfunc */
bool r0_rdonly;
@@ -1617,7 +1619,6 @@ struct bpf_call_arg_meta {
s64 const_map_key;
struct btf *ret_btf;
struct btf_field *kptr_field;
- struct arg_raw_mem_desc arg_raw_mem;
};
int bpf_get_helper_proto(struct bpf_verifier_env *env, int func_id,
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index de40785d6a7b..f4b88e402ff5 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -302,6 +302,12 @@ static int arg_idx_from_argno(argno_t a)
return arg_from_argno(a) - 1;
}
+/* Normalize helper register numbers and kfunc argument numbers to ABI slots. */
+static u32 arg_slot_from_argno(argno_t a)
+{
+ return abs(a.argno) - 1;
+}
+
static const char *btf_type_name(const struct btf *btf, u32 id)
{
return btf_name_by_offset(btf, btf_type_by_id(btf, id)->name_off);
@@ -6988,7 +6994,7 @@ static int check_stack_range_initialized(
*/
bool allow_poison = access_size < 0 || clobber;
/* The call will initialize the memory; uninitialized stack allowed */
- bool raw_mode = meta && meta->arg_raw_mem.regno == reg_from_argno(argno);
+ bool raw_mode = meta && meta->arg_raw_mem.argno == arg_slot_from_argno(argno) + 1;
access_size = abs(access_size);
@@ -7017,11 +7023,9 @@ static int check_stack_range_initialized(
reg_arg_name(env, argno), tn_buf);
return -EACCES;
}
- /* Only initialized buffer on stack is allowed to be accessed
- * with variable offset. With uninitialized buffer it's hard to
- * guarantee that whole memory is marked as initialized on
- * helper return since specific bounds are unknown what may
- * cause uninitialized stack leaking.
+ /*
+ * The call may touch any byte in the possible range, but does not
+ * definitely initialize all of it. Fall back to ordinary stack checks.
*/
raw_mode = false;
@@ -7223,8 +7227,8 @@ static int check_mem_size_reg(struct bpf_verifier_env *env,
* stack initialization checks, including their privilege exceptions.
*/
if (!tnum_is_const(size_reg->var_off) &&
- meta->arg_raw_mem.regno == reg_from_argno(mem_argno))
- meta->arg_raw_mem.regno = 0;
+ meta->arg_raw_mem.argno == arg_slot_from_argno(mem_argno) + 1)
+ meta->arg_raw_mem.argno = 0;
if (reg_smin(size_reg) < 0) {
verbose(env, "%s min value is negative, either use unsigned or 'var &= const'\n",
@@ -8941,8 +8945,8 @@ static int check_func_arg(struct bpf_verifier_env *env, u32 arg, u32 slot, u32 p
if (err)
return err;
- if (!meta->btf && arg_type_is_raw_mem(arg_type))
- meta->arg_raw_mem.regno = slot + 1;
+ if (arg_type_is_raw_mem(arg_type))
+ meta->arg_raw_mem.argno = slot + 1;
if (bpf_register_is_null(reg) && type_may_be_null(arg_type)) {
err = mark_arg_precision(env, argno);
@@ -9565,6 +9569,33 @@ static int check_func_args(struct bpf_verifier_env *env, struct bpf_call_arg_met
return 0;
}
+static int mark_raw_stack(struct bpf_verifier_env *env, struct bpf_call_arg_meta *meta,
+ int insn_idx)
+{
+ struct bpf_func_state *caller = cur_func(env);
+ struct bpf_reg_state *reg;
+ u32 slot = meta->arg_raw_mem.argno - 1;
+ int i, err;
+
+ if (!meta->arg_raw_mem.size)
+ return 0;
+ reg = get_func_arg_reg(caller, cur_regs(env), slot);
+
+ /*
+ * Validate every argument before initializing outputs: an input argument
+ * may alias an output buffer. Use the normal stack-write checks to discard
+ * stale spills and preserve the rules for special stack objects.
+ */
+ for (i = 0; i < meta->arg_raw_mem.size; i++) {
+ err = check_mem_access(env, insn_idx, reg, argno_from_arg(slot + 1), i, BPF_B,
+ BPF_WRITE, -1, false, false);
+ if (err)
+ return err;
+ }
+
+ return 0;
+}
+
static bool may_update_sockmap(struct bpf_verifier_env *env, int func_id)
{
enum bpf_attach_type eatype = env->prog->expected_attach_type;
@@ -9863,9 +9894,13 @@ static bool check_raw_mode_ok(const struct bpf_func_proto *fn)
int i;
for (i = 0; i < ARRAY_SIZE(fn->arg_type); i++) {
- if (fn->arg_type[i] == ARG_UNUSED)
+ enum bpf_arg_type type = fn->arg_type[i];
+
+ if (type == ARG_UNUSED)
break;
- if (!arg_type_is_raw_mem(fn->arg_type[i]))
+ /* Struct pointers may resolve to generic memory during argument checking. */
+ if (!arg_type_is_raw_mem(type) &&
+ !(base_type(type) == ARG_PTR_TO_BTF_ID && (type & MEM_UNINIT)))
continue;
if (seen)
return false;
@@ -11588,16 +11623,9 @@ static int check_helper_call(struct bpf_verifier_env *env, struct bpf_insn *insn
regs = cur_regs(env);
- /* Mark slots with STACK_MISC in case of raw mode, stack offset
- * is inferred from register state.
- */
- for (i = 0; i < meta.arg_raw_mem.size; i++) {
- err = check_mem_access(env, insn_idx, regs + meta.arg_raw_mem.regno,
- argno_from_reg(meta.arg_raw_mem.regno), i, BPF_B,
- BPF_WRITE, -1, false, false);
- if (err)
- return err;
- }
+ err = mark_raw_stack(env, &meta, insn_idx);
+ if (err)
+ return err;
if (meta.release_regno) {
struct bpf_reg_state *reg = ®s[meta.release_regno];
@@ -13101,6 +13129,10 @@ static int gen_kfunc_arg_proto(struct bpf_verifier_env *env, struct bpf_call_arg
return -ENOTSUPP;
}
+ if (!check_raw_mode_ok(proto)) {
+ verbose(env, "multiple __uninit buffers are not supported\n");
+ return -EINVAL;
+ }
return check_arg_prog_aux(env, proto) ? 0 : -EINVAL;
}
@@ -14178,6 +14210,10 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
if (err < 0)
return err;
+ err = mark_raw_stack(env, &meta, insn_idx);
+ if (err)
+ return err;
+
if ((is_bpf_obj_drop_kfunc(meta.func_id) ||
is_bpf_percpu_obj_drop_kfunc(meta.func_id)) && (is_tracing_prog_type(prog_type) ||
/* is_tracing_prog_type() for now doesn't cover non-iterator tracing progs. */
--
2.53.0
next prev parent reply other threads:[~2026-09-21 2:38 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-21 2:38 [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 01/11] selftests/bpf: Allow privileged preparation for capability tests Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 02/11] bpf: Record raw memory arguments during argument checking Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 03/11] bpf: Check __uninit kfunc output buffers as write-only Kumar Kartikeya Dwivedi
2026-09-21 3:54 ` bot+bpf-ci
2026-09-21 2:38 ` Kumar Kartikeya Dwivedi [this message]
2026-09-21 2:38 ` [PATCH bpf-next v5 05/11] selftests/bpf: Cover generic __uninit output initialization Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 06/11] bpf: Support multiple __uninit kfunc output arguments Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 07/11] selftests/bpf: Cover __uninit kfunc output argument slots Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 08/11] bpf: Preserve stack initialization for generic output buffers Kumar Kartikeya Dwivedi
2026-09-21 3:54 ` bot+bpf-ci
2026-09-21 2:38 ` [PATCH bpf-next v5 09/11] selftests/bpf: Cover generic output stack initialization Kumar Kartikeya Dwivedi
2026-09-21 3:54 ` bot+bpf-ci
2026-09-21 2:38 ` [PATCH bpf-next v5 10/11] bpf: Check read access for helper input/output buffers Kumar Kartikeya Dwivedi
2026-09-21 2:38 ` [PATCH bpf-next v5 11/11] selftests/bpf: Cover helper memory access permissions Kumar Kartikeya Dwivedi
2026-09-21 3:54 ` bot+bpf-ci
2026-09-21 17:20 ` [PATCH bpf-next v5 00/11] Fix generic __uninit kfunc output buffers patchwork-bot+netdevbpf
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260921023843.411943-5-memxor@gmail.com \
--to=memxor@gmail.com \
--cc=ameryhung@gmail.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=kernel-team@meta.com \
--cc=kkd@meta.com \
--cc=tj@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox