BPF List
 help / color / mirror / Atom feed
* [RFC PATCH bpf-next v1 00/16] BPF typed arenas
@ 2026-09-26 23:34 Kumar Kartikeya Dwivedi
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 01/16] mm/vmalloc: Add get_vm_area_align() Kumar Kartikeya Dwivedi
                   ` (15 more replies)
  0 siblings, 16 replies; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

NOTE: Patch marked as RFC, since this not ready for proper review. For
now, reviewers are advised to ignore concrete details, but instead focus
on the general approach, design decisions, and tradeoffs made in the
sandboxing approach to support objects that can hold trusted kernel objects.

This set adds typed arenas: kernel-only memory next to an arena map that
holds objects of a program's own struct types with special fields, kptrs
for now, at fixed slots. Programs reach the objects through native
pointers, the verifier trusts the special fields in them, and it does not
track the lifetime of an object, because the memory behind a typed arena
pointer is always valid.

Today, a struct with a kptr cannot be placed in arena memory.
The arena is mapped into user space, and access arbitrarily by the BPF
program with no constraints, which can write anything into it, so the
verifier cannot trust a pointer it loads from arena memory, and a
reference kept in such a field could be overwritten, duplicated or
dropped.

A program that builds its data structures in the arena and wants to
attach a task, a socket or an allocated object to a node keeps the
reference in a separate map and looks it up by key every time.
sched-ext has introduced id based indirection for task structs to work
around similar limitations.

The other way to get typed objects with kptrs, bpf_obj_new(), spills the
lifetime of every object into the static analysis: the pointer is owned
by whoever holds it, it must be dropped or moved into a map, a list or
an rbtree before the program exits, sharing it needs bpf_refcount, and
reading a shared object needs RCU or a lock. Every program that touches
such an object is verified against those rules, and data structures
built from allocated objects are written around them. This has proven
pretty inflexible in practice, and requires kernel changes for each new
data structure.

A typed arena is arena memory that is only mapped in the kernel, so that
only the program writes it and under the verifier's rules, and that is
always valid: every address in it resolves to an object of its struct,
allocated or not. A pointer to a typed arena object therefore does not
dangle, does not have to be released and does not have to be checked for
being alive, so there is no ownership, no reference and no RCU state for
the verifier to track.

The verifier still has to ensure is that a value used as a pointer to
struct T points at the start of an object of T. It does not do that by
tracking where the value came from, but by construction: a new
instruction, typed_arena_cast, masks any 64-bit value into T's slice of
the typed arena and rounds it to an object boundary. The instruction
lowers to a mask with an immediate and an add of the slice's base, four
instructions, with no branch and no bounds check, and the result cannot
point outside the slice or into the middle of an object. Any value goes
through it, a pointer loaded from the raw arena that user space
corrupted, an integer, a pointer of another struct, and comes out as a
trusted pointer to a whole object. Pointers are stored as they are, in
the raw arena, in maps, on the stack or in typed objects; only their use
as an address is sanitized.

The typed arena region is sparsely mapped, and populated on demand at
runtime using new allocation functions mirroring arena allocation
kfuncs. In case of faults, scratch pages are used to supply underlying
memory to allow such objects to be accessed freely from both the BPF
program and the kernel without issues.

Layout
------

Each arena map reserves a 4 GiB kernel-only region ahead of its
user-mappable window, which user space cannot map. A new mm helper
aligns the vm area to the region, so that every power-of-two slice in it
is aligned to its own size and the cast needs only a mask. A struct with
special fields gets one slice per program BTF and type in the program's
arena map, registered at program load when the verifier first meets the
type in a cast, in a load from a typed pointer field or in an
allocation.  The slice's size comes from a "typed_arena_size:<size>"
declaration tag on the struct, 128 MiB by default, a power of two of at
least a page and at most 2 GiB. Objects sit at slots of the struct's
size rounded up to a power of two. The unit of backing is a chunk, the
larger of a slot and a page: an object larger than a page is one
contiguous block, and an object never spans two allocations. The page
tables of a slice are populated at registration and charged to the map,
together with two bitmaps, because a fault on an unbacked typed page is
recovered in atomic context.

A chunk nobody allocated reads as zeroed dummy objects: a fault on it
maps in the slice's scratch chunk, one per slice, at the same position,
so that a scalar field of one dummy object does not alias another. Two
kfuncs back a range with real objects and release it, in page counts
like the raw arena kfuncs, rounded up to whole chunks, with the granted
count written back:

  void *bpf_typed_arena_alloc_pages(void *map, __u64 type_id, void *addr,
                                    __u32 *page_cnt, int node_id);
  void bpf_typed_arena_free_pages(void *map, __u64 type_id, void *ptr,
                                  __u32 page_cnt);

A release is deferred behind an RCU and an RCU-tasks-trace grace period,
so that every invocation that started before it keeps valid objects, and
a range is released once, since a second release queued behind the first
would take the range away from its next owner. The kptrs of released
objects are dropped once the chunk is unmapped. A pointer to a released
object that a program kept still names a whole object of the type, of
whoever claims the chunk next; that is program logic, like a stale index
into an array, not a safety problem.

Access
------

A pointer member of a typed struct whose pointee is itself a typed struct
is a typed pointer field, found from BTF alone, without annotation. The
verifier lets nothing but a trusted pointer to the pointee at offset
zero, or the constant 0, be stored into it, the memory is kernel-only,
and a reused chunk comes back zeroed, so a load from the field is
trusted to be an object or 0 without a cast. The loaded pointer is
canonicalized in place only where it is used as an address; a
comparison, a copy or a store of it as a value sees the raw value, so a
list is walked with plain loads and ends on the usual test against NULL.
A program that uses the value without checking for 0 lands on object 0
of the slice, as a cast of 0 would. Whether to check is left to the
program, since nothing it can do with the value reaches outside the
slice; a maybe-NULL type that must be checked before every use would
cost a check on every hop that the sanitization makes unnecessary.

Scalar fields are loaded and stored natively, atomics included, with no
probe mode and no exception table entry. kptr fields go through
bpf_kptr_xchg(), as in map values.

Compiler Changes
----------------

Programs do not write the cast. Behind the typed-arena target feature,
Clang inserts it where a pointer to a struct with special fields is used
as an address, and where a value of another type is converted to such a
pointer, which is where a value loaded from the raw arena or returned by
the allocator becomes a typed pointer. The same struct also lives in
allocated objects, map values and on the stack, and the compiler cannot
tell which, so the verifier decides per path what a cast costs: a value
it does not trust is sanitized, a pointer it already trusts is copied,
and a program that only uses bpf_obj_new() objects of the struct loads
as before. The builtin, __builtin_bpf_typed_arena_cast(), remains for
turning a scalar or a raw arena pointer into a typed pointer on purpose.
The feature defines __BPF_FEATURE_TYPED_ARENA_CAST; libbpf relocates the
instruction and provides bpf_typed_arena_cast(); the selftests probe for
the feature and are skipped without it.

The compiler changes are on the typed-arena branch of
https://github.com/kkdwvd/llvm-project, two commits on top of main:

  4f5f6b0db95c BPF: Add the typed_arena_cast builtin and instruction
  cb29a13b73bd BPF: Insert typed_arena_cast at uses of typed record pointers

Notes
-----

 * BPF kptrs are the only special fields a typed arena object may have. Timers,
   workqueues, task work and RCU heads keep kernel state that a release or
   a fault could pull from under a kfunc, so they are refused inline for now,
   this will be made to work in subsequent revisions.

   Pages can be unmapped while a kernel is operating on an object
   embedded inline a typed arena object sitting in a page that was
   requested to be freed. Some protocol to signal when operating on such
   objects from the kernel side will be necessary to ensure that new
   operations cannot begin on such fields, and whether existing
   operations are finished.

   Spin locks, lists, rbtrees, etc. are unnecessary to support as
   special fields, since they have their own analogues in arenas.

 * A slot is a power of two, so an object is padded by up to half its slot.
   This is the trade the slab allocator makes for its size classes, and it
   is what keeps the cast at a mask and an add. Denser packing can be added
   behind the same interface, since programs never see the slot, but it
   is likely not worth it.
 * Slices are fixed at registration, neither grown nor moved, and a
   map's slices share its 4 GiB region. Objects larger than 4 MiB are
   refused.
 * The scratch chunk is shared by every unallocated chunk of a slice, so
   a store through a stray pointer is visible through every unallocated
   chunk, and a kptr stored there is dropped with the map.
 * The sanitizing sequence runs on every value that enters from
   untrusted memory and, for now, before every use of a loaded typed
   pointer as an address.
 * The page tables of a slice cost 8 bytes per page, 256 KiB for the
   default size, charged to the map.
 * Nothing in the set is UAPI: the kfuncs, the declaration tag and the
   instruction encoding may change.
 * Only x86-64 has been tested. No JIT changes are needed: the cast is
   lowered by the verifier, accesses are native loads and stores, and
   the fault path lives in the arena map.

Follow-ups
----------

 * Elide the sanitization on loads through a typed pointer that nothing
   writes to, with a probed load, so that a chain of reads costs nothing
   beyond the loads.

Patch 1 adds an mm helper for an aligned sparse vm area. Patches 2 and 3
add the region, the registry and the scratch backing. Patches 4 to 8 add
the cast instruction, object access, special fields, typed pointer fields
and their canonicalization at use. Patch 9 adds the kfuncs and patch 10
the free cast on trusted pointers. Patch 11 adds the libbpf support, patch
12 the selftests' compiler probe, and patches 13 to 16 the selftests.

Kumar Kartikeya Dwivedi (16):
  mm/vmalloc: Add get_vm_area_align()
  bpf: Introduce BPF typed arenas
  bpf: Back typed arena chunks with scratch on demand
  bpf: Add the typed_arena_cast instruction
  bpf: Allow scalar and atomic access to typed arena objects
  bpf: Support special fields in typed arena objects
  bpf: Trust typed pointer fields of typed arena objects
  bpf: Canonicalize loaded typed arena pointers where they are used
  bpf: Add typed arena page allocation and release kfuncs
  bpf: Let typed_arena_cast copy pointers the verifier already trusts
  libbpf: Support the typed_arena_cast instruction
  selftests/bpf: Build BPF objects with compiler-inserted typed arena
    casts
  selftests/bpf: Test typed arena casts and registration
  selftests/bpf: Test typed arena object access, kptrs and typed pointer
    fields
  selftests/bpf: Test typed arena page allocation and release
  selftests/bpf: Exercise typed arenas at run time

 include/linux/bpf.h                           |   91 ++
 include/linux/bpf_verifier.h                  |   35 +
 include/linux/vmalloc.h                       |    2 +
 include/uapi/linux/bpf.h                      |    6 +
 kernel/bpf/arena.c                            |  772 +++++++++-
 kernel/bpf/backtrack.c                        |    8 +-
 kernel/bpf/btf.c                              |  140 +-
 kernel/bpf/core.c                             |    2 +
 kernel/bpf/disasm.c                           |    9 +
 kernel/bpf/fixups.c                           |   64 +
 kernel/bpf/log.c                              |    6 +-
 kernel/bpf/syscall.c                          |    3 +
 kernel/bpf/verifier.c                         |  498 ++++++-
 mm/vmalloc.c                                  |   19 +
 tools/include/uapi/linux/bpf.h                |    6 +
 tools/lib/bpf/bpf_helpers.h                   |   13 +
 tools/lib/bpf/relo_core.c                     |   19 +-
 tools/testing/selftests/bpf/Makefile          |   13 +-
 .../testing/selftests/bpf/bpf_experimental.h  |   18 +
 .../selftests/bpf/prog_tests/typed_arena.c    |  489 +++++++
 .../selftests/bpf/prog_tests/verifier.c       |   22 +
 .../testing/selftests/bpf/progs/typed_arena.c |  354 +++++
 .../bpf/progs/verifier_typed_arena.c          | 1304 +++++++++++++++++
 23 files changed, 3852 insertions(+), 41 deletions(-)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/typed_arena.c
 create mode 100644 tools/testing/selftests/bpf/progs/typed_arena.c
 create mode 100644 tools/testing/selftests/bpf/progs/verifier_typed_arena.c


base-commit: ea9358e1270ab2c3ba6f36bd9bdda68617665516
-- 
2.53.0


^ permalink raw reply	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 01/16] mm/vmalloc: Add get_vm_area_align()
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:42   ` sashiko-bot
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 02/16] bpf: Introduce BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (14 subsequent siblings)
  15 siblings, 1 reply; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

Add a variant of get_vm_area() that takes the alignment of the area's start.
The internal allocator has supported an alignment all along, through
__get_vm_area_node(), but every caller of the public API passes one page, so
a user that needs a naturally aligned region has to over-reserve and skip the
slack in front of the aligned address, wasting up to the alignment in virtual
space per area.

BPF arena maps are about to reserve a 4 GiB kernel-only region next to the
existing arena window, holding one power-of-two slice per program struct
type. Each slice is placed at a multiple of its own size inside the region,
which makes the region's alignment the largest slice size, so that a BPF
program can turn any 64-bit value into an object address in a slice with a
mask and an add. Reserving the region aligned avoids carrying 4 GiB of slack
in every arena map.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 include/linux/vmalloc.h |  2 ++
 mm/vmalloc.c            | 19 +++++++++++++++++++
 2 files changed, 21 insertions(+)

diff --git a/include/linux/vmalloc.h b/include/linux/vmalloc.h
index aed121d729b0..e8aa955fe062 100644
--- a/include/linux/vmalloc.h
+++ b/include/linux/vmalloc.h
@@ -246,6 +246,8 @@ static inline size_t get_vm_area_size(const struct vm_struct *area)
 extern struct vm_struct *get_vm_area(unsigned long size, unsigned long flags);
 extern struct vm_struct *get_vm_area_caller(unsigned long size,
 					unsigned long flags, const void *caller);
+extern struct vm_struct *get_vm_area_align(unsigned long size,
+					unsigned long align, unsigned long flags);
 extern struct vm_struct *__get_vm_area_caller(unsigned long size,
 					unsigned long flags,
 					unsigned long start, unsigned long end,
diff --git a/mm/vmalloc.c b/mm/vmalloc.c
index bea9f76ed7e7..b70e4e27ec36 100644
--- a/mm/vmalloc.c
+++ b/mm/vmalloc.c
@@ -3301,6 +3301,25 @@ struct vm_struct *get_vm_area_caller(unsigned long size, unsigned long flags,
 				  NUMA_NO_NODE, GFP_KERNEL, caller);
 }
 
+/**
+ * get_vm_area_align - reserve a contiguous kernel virtual area at an aligned start
+ * @size:	 size of the area
+ * @align:	 alignment of the start of the area, a power of two
+ * @flags:	 %VM_IOREMAP for I/O mappings or VM_ALLOC
+ *
+ * Like get_vm_area(), with the start of the area aligned to @align.
+ *
+ * Return: the area descriptor on success or %NULL on failure.
+ */
+struct vm_struct *get_vm_area_align(unsigned long size, unsigned long align,
+				    unsigned long flags)
+{
+	return __get_vm_area_node(size, align, PAGE_SHIFT, flags,
+				  VMALLOC_START, VMALLOC_END,
+				  NUMA_NO_NODE, GFP_KERNEL,
+				  __builtin_return_address(0));
+}
+
 /**
  * find_vm_area - find a continuous kernel virtual area
  * @addr:	  base address
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 02/16] bpf: Introduce BPF typed arenas
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 01/16] mm/vmalloc: Add get_vm_area_align() Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 03/16] bpf: Back typed arena chunks with scratch on demand Kumar Kartikeya Dwivedi
                   ` (13 subsequent siblings)
  15 siblings, 0 replies; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

Objects with special fields cannot live in arena memory: user space maps the
arena and programs store arbitrary bytes into it, so the verifier trusts
nothing loaded from there, and a kptr field is a kernel pointer the kernel
must be able to trust. Give such objects a home next to the arena instead.

Reserve a typed arena region ahead of the arena's raw window, a kernel-only
4 GiB of virtual space user space never maps, and hand out one typed arena
per program-BTF struct with special fields: a power-of-two slice of the
region holding one object per power-of-two slot. Objects are reached through
native kernel pointers that the verifier constructs, so the region's memory
carries no user-visible representation and objects need no translation on
load or store; a later patch adds the instruction that turns any 64-bit value
into such a pointer, and another the fields the verifier trusts inside them.

The slice is aligned to its own size, which is what makes that cast a mask
and an add: masking a value with the slice's size less its slot lands on a
slot of the slice once the base is added, with no bounds check, no branch and
no subtraction. Alignment inside the region comes from placing a slice at a
multiple of its size, first-fit over the registry kept in address order, and
alignment of the region itself comes from reserving the arena's vm_area
aligned to the region with get_vm_area_align(), instead of over-reserving and
skipping slack in front. A slice is at most half the region, so that the mask
is a positive 32-bit immediate; the raw window and its guards keep their
layout and offset math, shifted behind the region.

The slot is the object's size rounded up to a power of two. That pads an
object by up to half its slot, the same trade the slab allocator makes for
its size classes, and it buys the four-instruction cast above: any denser
packing needs a reciprocal multiply to find an object boundary and a scratch
layout that repeats with the least common multiple of slot and page. Both
can be added later behind the same interface, since nothing about the slot
is visible to programs. The typed arena's size comes from the struct's
"typed_arena_size:" declaration tag, parsed by the verifier, with a 128 MiB
default; it decides how many objects the type can hold.

A typed arena is keyed by the BTF object and the type ID, so two programs
loaded from different objects get distinct slices even for structurally equal
structs, and it is registered at program load, where the verifier first meets
the type. A program whose load fails drops its references; a program that
loaded keeps the typed arena alive for the map's lifetime, so that objects
survive program reloads. The page tables of the whole slice are populated at
registration, because a fault on an unbacked typed page is recovered in
atomic context, which cannot allocate them then; they cost eight bytes per
page of the slice, 256 KiB for the default size, and are charged to the map
together with the two chunk bitmaps as memory fixed for the map's lifetime.

The chunk, the larger of the slot and a page, is the unit at which the slice
is backed: a chunk holds whole objects, so an object is never split between a
real page and a page backed some other way, and a chunk larger than a page is
one allocation of its order, so that a multi-page object is contiguous in the
direct map. Map teardown walks every slice, drops the special fields of the
objects of each real chunk through the arena mapping, still in place, and
frees the chunk's allocation.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 include/linux/bpf.h |  62 ++++++++++
 kernel/bpf/arena.c  | 274 +++++++++++++++++++++++++++++++++++++++++++-
 2 files changed, 331 insertions(+), 5 deletions(-)

diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 4bae3796c42f..f06d138b1f57 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -45,6 +45,7 @@ struct bpf_prog;
 struct bpf_prog_aux;
 struct bpf_map;
 struct bpf_arena;
+struct bpf_typed_arena;
 struct sock;
 struct seq_file;
 struct btf;
@@ -275,6 +276,67 @@ struct btf_record {
 	struct btf_field fields[];
 };
 
+/*
+ * Typed arenas: kernel-only object storage next to an arena map. A program-BTF
+ * struct with special fields that a program casts to, or allocates, gets a
+ * typed arena: a power-of-two slice of the map's typed region, naturally
+ * aligned, holding one object per power-of-two slot. Objects are reached
+ * through native kernel pointers that the verifier constructs; a cast masks
+ * any 64-bit value into a slot of the slice, so every value names an object
+ * of the type. The chunk, max(slot, page), is the unit of backing: a chunk no
+ * program allocated reads as the zeroed scratch chunk, one dummy object per
+ * slot, once an access faults it in. The chunks bitmap marks every chunk that
+ * holds a mapping, real or scratch; pending marks the chunks whose release is
+ * queued, which stay mapped and marked until it has run. record is the
+ * struct's own special-field record, from the BTF that is retained here.
+ */
+#define BPF_TYPED_ARENA_SIZE_TAG "typed_arena_size:"
+#define BPF_TYPED_ARENA_DEFAULT_SIZE SZ_128M
+
+struct bpf_typed_arena {
+	struct list_head node;
+	struct bpf_map *map;
+	refcount_t refcnt;
+	struct btf *btf;
+	u32 btf_id;
+	u8 slot_shift;
+	u8 chunk_shift;
+	u8 size_shift;
+	void *base;
+	void *scratch;
+	struct page **scratch_pages;
+	unsigned long *chunks;
+	unsigned long *pending;
+	const struct btf_record *record;
+};
+
+static inline u64 bpf_typed_arena_size(const struct bpf_typed_arena *ta)
+{
+	return 1ull << ta->size_shift;
+}
+
+static inline u32 bpf_typed_arena_slot(const struct bpf_typed_arena *ta)
+{
+	return 1u << ta->slot_shift;
+}
+
+static inline u32 bpf_typed_arena_chunk(const struct bpf_typed_arena *ta)
+{
+	return 1u << ta->chunk_shift;
+}
+
+/* The cast mask keeps the slot index and drops the offset within the slot and the bits above the slice. */
+static inline u64 bpf_typed_arena_mask(const struct bpf_typed_arena *ta)
+{
+	return bpf_typed_arena_size(ta) - bpf_typed_arena_slot(ta);
+}
+
+struct bpf_typed_arena *bpf_typed_arena_get(struct bpf_map *map, struct btf *btf, u32 btf_id,
+					    const struct btf_record *record, u64 size);
+void bpf_typed_arena_put(struct bpf_map *map, struct bpf_typed_arena *ta);
+int bpf_typed_arena_cast_insns(const struct bpf_typed_arena *ta, u8 dst, u8 src,
+			       struct bpf_insn *buf);
+
 /* Non-opaque version of bpf_rb_node in uapi/linux/bpf.h */
 struct bpf_rb_node_kern {
 	struct rb_node rb_node;
diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index c6369ea5e208..9cd1b1ce434c 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -40,11 +40,33 @@
  * bpf program can allocate a page via bpf_arena_alloc_pages() kfunc
  * which will insert it into kernel vm_area.
  * The later fault-in from user space will populate that page into user vma.
+ *
+ * Ahead of the lower guard, the same vm_area holds the typed arena region, a
+ * kernel-only space user space never maps:
+ *
+ *   [ typed region 4 GiB ][ GUARD_SZ/2 ][ raw window 4 GiB ][ GUARD_SZ/2 ]
+ *   ^ kern_vm->addr                     ^ kern_vm_start
+ *
+ * A program-BTF struct with special fields that a program casts to gets a
+ * typed arena there: a power-of-two slice holding one object per power-of-two
+ * slot, reached through native kernel pointers. The region is aligned to its
+ * size and a slice sits at a multiple of its own size, so the slice is aligned
+ * to itself and a cast is a mask and an add: any 64-bit value masked with the
+ * slice's size less its slot, plus the base, is an object of the slice. Raw
+ * arena pointers and typed pointers never mix: a typed access is bounded by
+ * its slice, so the guard between the region and the window keeps it off the
+ * window, and the window's 32-bit offsets cannot reach the region.
  */
 
 /* number of bytes addressable by LDX/STX insn with 16-bit 'off' field */
 #define GUARD_SZ round_up(1ull << sizeof_field(struct bpf_insn, off) * 8, PAGE_SIZE << 1)
-#define KERN_VM_SZ (SZ_4G + GUARD_SZ)
+/*
+ * A slice is at most half the region: the cast masks with the slice's size
+ * less its slot, which must be a positive 32-bit immediate.
+ */
+#define TYPED_ARENA_REGION_SZ SZ_4G
+#define TYPED_ARENA_MAX_SZ SZ_2G
+#define KERN_VM_SZ (TYPED_ARENA_REGION_SZ + SZ_4G + GUARD_SZ)
 
 static void arena_free_pages(struct bpf_arena *arena, long uaddr, long page_cnt, bool sleepable);
 
@@ -60,7 +82,11 @@ struct bpf_arena {
 	/* number of pages currently populated in the arena */
 	u64 nr_pages;
 	struct list_head vma_list;
-	/* protects vma_list */
+	/* typed arenas in address order; mutated under lock, walked under RCU */
+	struct list_head typed_arenas;
+	/* memory of the typed arenas fixed at registration, for accounting */
+	u64 typed_arena_mem;
+	/* protects vma_list and typed_arenas */
 	struct mutex lock;
 	u64 zap_gen;
 	struct mutex zap_mutex;
@@ -78,9 +104,14 @@ struct arena_free_span {
 	u32 page_cnt;
 };
 
+static unsigned long typed_arena_region(struct bpf_arena *arena)
+{
+	return (unsigned long)arena->kern_vm->addr;
+}
+
 u64 bpf_arena_get_kern_vm_start(struct bpf_arena *arena)
 {
-	return arena ? (u64) (long) arena->kern_vm->addr + GUARD_SZ / 2 : 0;
+	return arena ? typed_arena_region(arena) + TYPED_ARENA_REGION_SZ + GUARD_SZ / 2 : 0;
 }
 
 u64 bpf_arena_get_user_vm_start(struct bpf_arena *arena)
@@ -263,6 +294,236 @@ static int populate_pgtable_except_pte(struct bpf_arena *arena)
 				   SZ_4G + GUARD_SZ / 2, apply_range_set_cb, NULL);
 }
 
+static unsigned long typed_arena_nr_chunks(const struct bpf_typed_arena *ta)
+{
+	return bpf_typed_arena_size(ta) >> ta->chunk_shift;
+}
+
+/* A chunk is one allocation of this order: a multi-page object is contiguous in the direct map. */
+static unsigned int typed_arena_chunk_order(const struct bpf_typed_arena *ta)
+{
+	return ta->chunk_shift - PAGE_SHIFT;
+}
+
+/*
+ * The memory a typed arena takes for the map's lifetime: a page table entry
+ * per page of the slice and the two chunk bitmaps.
+ */
+static u64 typed_arena_static_mem(const struct bpf_typed_arena *ta)
+{
+	return (bpf_typed_arena_size(ta) >> PAGE_SHIFT) * sizeof(pte_t) +
+	       2 * BITS_TO_LONGS(typed_arena_nr_chunks(ta)) * sizeof(long);
+}
+
+/*
+ * First-fit search for a slice of the region aligned to its own size. The
+ * registry is sorted by address, so its gaps are visited in order.
+ */
+static s64 typed_arena_find_slice(struct bpf_arena *arena, u64 size)
+{
+	unsigned long region = typed_arena_region(arena);
+	struct bpf_typed_arena *ta;
+	u64 off = 0, start;
+
+	list_for_each_entry(ta, &arena->typed_arenas, node) {
+		start = ALIGN(off, size);
+		if (start + size <= (unsigned long)ta->base - region)
+			return start;
+		off = (unsigned long)ta->base - region + bpf_typed_arena_size(ta);
+	}
+	start = ALIGN(off, size);
+	return start + size <= TYPED_ARENA_REGION_SZ ? start : -ENOSPC;
+}
+
+static void typed_arena_insert(struct bpf_arena *arena, struct bpf_typed_arena *ta)
+{
+	struct bpf_typed_arena *pos;
+
+	list_for_each_entry(pos, &arena->typed_arenas, node) {
+		if (pos->base > ta->base) {
+			list_add_tail_rcu(&ta->node, &pos->node);
+			return;
+		}
+	}
+	list_add_tail_rcu(&ta->node, &arena->typed_arenas);
+}
+
+/* Drop the special fields of every object of the chunk mapped at @chunk. */
+static void typed_arena_free_objects(const struct bpf_typed_arena *ta, void *chunk)
+{
+	u32 off;
+
+	for (off = 0; off < bpf_typed_arena_chunk(ta); off += bpf_typed_arena_slot(ta))
+		bpf_obj_free_fields(ta->record, chunk + off);
+}
+
+/*
+ * Free the real chunks of a typed arena at map teardown. A chunk is backed as
+ * a whole by one allocation of its order, so its first entry decides: the
+ * fields of its objects are dropped through the arena mapping, still in
+ * place, and the allocation is returned. The entries are not cleared;
+ * free_vm_area() does that for the whole area.
+ */
+static int typed_arena_teardown_cb(pte_t *ptep, unsigned long addr, void *data)
+{
+	struct bpf_typed_arena *ta = data;
+	pte_t pte = ptep_get(ptep);
+
+	if ((addr - (unsigned long)ta->base) & (bpf_typed_arena_chunk(ta) - 1))
+		return 0;
+	if (!pte_present(pte))
+		return 0;
+	typed_arena_free_objects(ta, (void *)addr);
+	__free_pages(pte_page(pte), typed_arena_chunk_order(ta));
+	return 0;
+}
+
+static void typed_arena_free(struct bpf_arena *arena, struct bpf_typed_arena *ta)
+{
+	WRITE_ONCE(arena->typed_arena_mem, arena->typed_arena_mem - typed_arena_static_mem(ta));
+	apply_to_existing_page_range(&init_mm, (unsigned long)ta->base, bpf_typed_arena_size(ta),
+				     typed_arena_teardown_cb, ta);
+	bitmap_free(ta->chunks);
+	bitmap_free(ta->pending);
+	btf_put(ta->btf);
+	kfree(ta);
+}
+
+/*
+ * Find or create the typed arena of a program-BTF struct. A typed arena is
+ * keyed by the BTF object and the type ID: two programs with different BTF
+ * objects get distinct slices even for structurally equal structs, and one
+ * BTF declares one size for a struct. The caller validated the record and
+ * parsed the size, a power of two between a page and TYPED_ARENA_MAX_SZ. The
+ * slot is the power of two covering the object, the chunk the larger of the
+ * slot and a page, so that a chunk holds whole objects and an object is never
+ * split between a real page and a scratch one. The page tables of the slice
+ * are populated here, in a sleepable context, because a fault on an unbacked
+ * typed page is recovered in atomic context, which cannot allocate them; they
+ * are charged to the map's memory, together with the bitmaps, for the map's
+ * lifetime. The reference returned is dropped with bpf_typed_arena_put().
+ */
+struct bpf_typed_arena *bpf_typed_arena_get(struct bpf_map *map, struct btf *btf, u32 btf_id,
+					    const struct btf_record *record, u64 size)
+{
+	struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
+	const struct btf_type *t = btf_type_by_id(btf, btf_id);
+	struct bpf_typed_arena *ta;
+	u64 slot, chunk;
+	s64 off;
+	int err;
+
+	if (!t || !t->size || !record)
+		return ERR_PTR(-EINVAL);
+	if (!is_power_of_2(size) || size < PAGE_SIZE || size > TYPED_ARENA_MAX_SZ)
+		return ERR_PTR(-EINVAL);
+	slot = roundup_pow_of_two(t->size);
+	chunk = max_t(u64, slot, PAGE_SIZE);
+	if (slot > size || chunk > PAGE_SIZE << MAX_PAGE_ORDER)
+		return ERR_PTR(-E2BIG);
+
+	guard(mutex)(&arena->lock);
+
+	list_for_each_entry(ta, &arena->typed_arenas, node) {
+		if (ta->btf != btf || ta->btf_id != btf_id)
+			continue;
+		refcount_inc(&ta->refcnt);
+		return ta;
+	}
+
+	off = typed_arena_find_slice(arena, size);
+	if (off < 0)
+		return ERR_PTR(off);
+
+	ta = kzalloc(sizeof(*ta), GFP_KERNEL_ACCOUNT);
+	if (!ta)
+		return ERR_PTR(-ENOMEM);
+	ta->base = (void *)(typed_arena_region(arena) + off);
+	ta->size_shift = ilog2(size);
+	ta->slot_shift = ilog2(slot);
+	ta->chunk_shift = ilog2(chunk);
+	ta->chunks = bitmap_zalloc(typed_arena_nr_chunks(ta), GFP_KERNEL_ACCOUNT);
+	ta->pending = bitmap_zalloc(typed_arena_nr_chunks(ta), GFP_KERNEL_ACCOUNT);
+	if (!ta->chunks || !ta->pending) {
+		err = -ENOMEM;
+		goto free;
+	}
+	err = apply_to_page_range(&init_mm, (unsigned long)ta->base, size, apply_range_set_cb, NULL);
+	if (err)
+		goto free;
+	refcount_set(&ta->refcnt, 1);
+	ta->map = map;
+	btf_get(btf);
+	ta->btf = btf;
+	ta->btf_id = btf_id;
+	ta->record = record;
+	typed_arena_insert(arena, ta);
+	WRITE_ONCE(arena->typed_arena_mem, arena->typed_arena_mem + typed_arena_static_mem(ta));
+	return ta;
+
+free:
+	bitmap_free(ta->chunks);
+	bitmap_free(ta->pending);
+	kfree(ta);
+	return ERR_PTR(err);
+}
+
+/*
+ * Drop a reference taken by bpf_typed_arena_get(). Only a program whose load
+ * was rejected drops its references; a program that loaded keeps the typed
+ * arena alive for the map's lifetime, so that its objects survive program
+ * reloads. The last reference retracts the slice. Its page tables stay
+ * populated until the map is freed and serve the next slice placed there.
+ */
+void bpf_typed_arena_put(struct bpf_map *map, struct bpf_typed_arena *ta)
+{
+	struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
+
+	guard(mutex)(&arena->lock);
+	if (!refcount_dec_and_test(&ta->refcnt))
+		return;
+	/*
+	 * No loaded program holds a pointer into a typed arena it did not
+	 * register, so nothing can fault in this slice; walkers of the registry
+	 * can still be stepping past this node.
+	 */
+	list_del_rcu(&ta->node);
+	synchronize_rcu();
+	typed_arena_free(arena, ta);
+}
+
+/*
+ * Lower a typed_arena_cast to plain BPF: dst = (src & mask) + base. The mask
+ * is the slice's size less its slot: it rounds any value down to a slot and,
+ * as a positive 64-bit AND, clears everything above the slice, so every input
+ * lands on an object of the type. The base is aligned to the slice's size, so
+ * nothing needs subtracting first. Return the number of instructions written.
+ */
+int bpf_typed_arena_cast_insns(const struct bpf_typed_arena *ta, u8 dst, u8 src,
+			       struct bpf_insn *buf)
+{
+	struct bpf_insn base[2] = { BPF_LD_IMM64(BPF_REG_AX, (unsigned long)ta->base) };
+	int cnt = 0;
+
+	if (dst != src)
+		buf[cnt++] = BPF_MOV64_REG(dst, src);
+	buf[cnt++] = BPF_ALU64_IMM(BPF_AND, dst, bpf_typed_arena_mask(ta));
+	buf[cnt++] = base[0];
+	buf[cnt++] = base[1];
+	buf[cnt++] = BPF_ALU64_REG(BPF_ADD, dst, BPF_REG_AX);
+	return cnt;
+}
+
+static void typed_arenas_free(struct bpf_arena *arena)
+{
+	struct bpf_typed_arena *ta, *tmp;
+
+	list_for_each_entry_safe(ta, tmp, &arena->typed_arenas, node) {
+		list_del(&ta->node);
+		typed_arena_free(arena, ta);
+	}
+}
+
 static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
 {
 	struct vm_struct *kern_vm;
@@ -293,7 +554,7 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
 		/* user vma must not cross 32-bit boundary */
 		return ERR_PTR(-ERANGE);
 
-	kern_vm = get_vm_area(KERN_VM_SZ, VM_SPARSE | VM_USERMAP);
+	kern_vm = get_vm_area_align(KERN_VM_SZ, TYPED_ARENA_REGION_SZ, VM_SPARSE | VM_USERMAP);
 	if (!kern_vm)
 		return ERR_PTR(-ENOMEM);
 
@@ -307,6 +568,7 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
 		arena->user_vm_end = arena->user_vm_start + vm_range;
 
 	INIT_LIST_HEAD(&arena->vma_list);
+	INIT_LIST_HEAD(&arena->typed_arenas);
 	init_llist_head(&arena->free_spans);
 	init_irq_work(&arena->free_irq, arena_free_irq);
 	INIT_WORK(&arena->free_work, arena_free_worker);
@@ -392,6 +654,7 @@ static void arena_map_free(struct bpf_map *map)
 	 */
 	apply_to_existing_page_range(&init_mm, bpf_arena_get_kern_vm_start(arena),
 				     SZ_4G + GUARD_SZ / 2, existing_page_cb, arena);
+	typed_arenas_free(arena);
 	free_vm_area(arena->kern_vm);
 	range_tree_destroy(&arena->rt);
 	__free_page(arena->scratch_page);
@@ -419,7 +682,8 @@ static u64 arena_map_mem_usage(const struct bpf_map *map)
 {
 	struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
 
-	return (u64)READ_ONCE(arena->nr_pages) << PAGE_SHIFT;
+	return ((u64)READ_ONCE(arena->nr_pages) << PAGE_SHIFT) +
+	       READ_ONCE(arena->typed_arena_mem);
 }
 
 struct vma_list {
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 03/16] bpf: Back typed arena chunks with scratch on demand
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 01/16] mm/vmalloc: Add get_vm_area_align() Kumar Kartikeya Dwivedi
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 02/16] bpf: Introduce BPF typed arenas Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:56   ` sashiko-bot
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 04/16] bpf: Add the typed_arena_cast instruction Kumar Kartikeya Dwivedi
                   ` (12 subsequent siblings)
  15 siblings, 1 reply; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

The cast that produces a typed arena pointer never fails: it masks any value
into a slot of the slice, so a program can reach every object the slice
could hold, whether or not anything allocated it. Every such pointer must
still denote a valid object of the type, or the native accesses through it
would fault, and a kptr exchange would write a kernel pointer into unbacked
memory. So back the slice on demand: an access to an unbacked page of the
slice faults, and the kernel fault recovery of the arena, reached through
bpf_arena_handle_page_fault(), installs scratch memory there and resumes
the access, reporting it on the program's stderr stream as a use of an
unallocated object.

The scratch memory is one zeroed chunk per typed arena, and the page of the
scratch chunk installed at a faulting page is the one at the same position in
the chunk. A chunk holds whole objects laid out from its start, so the dummy
objects a program sees are laid out exactly like real ones: a kptr field of
one dummy object never aliases a scalar field of another, which matters
because a kptr stored through one pointer must not be readable as a scalar
through another, and the special fields of the dummy objects can be dropped
at map teardown by walking the scratch chunk like a real one. This is why
the chunk, rather than the page, is the unit of backing, and why the slot is
a power of two: with objects straddling pages at arbitrary offsets the
layout would differ from page to page, and the scratch memory would have to
grow to the period at which it repeats. The scratch chunk is shared by every
unallocated chunk of the slice, so the dummy objects behave as one sink: a
program that writes through a stray pointer sees its writes through every
other stray pointer to the same position, and a kptr it leaves there lives
until the map is freed.

The fault path marks the chunk taken in the bitmap before installing the
page, with an atomic bit set because a program can fault in any context it
runs in, including NMI, and the allocator, added next, only backs chunks
that are not marked. A real page never replaces a scratch page either: a
program may be in the middle of using the dummy object, nothing can wait for
that use to end because the next stray access faults the scratch page right
back in, and the effect of the mapping changing under a kernel operation on
the object has not been reasoned through. The raw arena can replace its
scratch page because its contents are raw data. A chunk the fault path took
is given back with the release kfunc like any other, after which it can be
allocated or faulted again.

Refuse to register typed arenas on architectures that do not provide the
atomic PTE installer, since those do not wire up the arena fault path and a
typed access there would oops. The guard between the region and the raw
window stays outside recovery: a typed access is bounded by its slice and
cannot reach it, so a fault there is still a bug to oops on.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 kernel/bpf/arena.c | 115 ++++++++++++++++++++++++++++++++++++++++++---
 1 file changed, 109 insertions(+), 6 deletions(-)

diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index 9cd1b1ce434c..c757c1932360 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -67,6 +67,16 @@
 #define TYPED_ARENA_REGION_SZ SZ_4G
 #define TYPED_ARENA_MAX_SZ SZ_2G
 #define KERN_VM_SZ (TYPED_ARENA_REGION_SZ + SZ_4G + GUARD_SZ)
+/*
+ * Typed accesses are native loads and stores, so a fault on an unbacked typed
+ * page can only be recovered where the arena's kernel fault path is wired up,
+ * which is where the architecture provides an atomic PTE installer.
+ */
+#ifdef ptep_try_set
+#define TYPED_ARENA_SUPPORTED true
+#else
+#define TYPED_ARENA_SUPPORTED false
+#endif
 
 static void arena_free_pages(struct bpf_arena *arena, long uaddr, long page_cnt, bool sleepable);
 
@@ -305,13 +315,28 @@ static unsigned int typed_arena_chunk_order(const struct bpf_typed_arena *ta)
 	return ta->chunk_shift - PAGE_SHIFT;
 }
 
+static u32 typed_arena_chunk_pages(const struct bpf_typed_arena *ta)
+{
+	return bpf_typed_arena_chunk(ta) >> PAGE_SHIFT;
+}
+
+/* The page of the scratch chunk that backs the page at @addr while its chunk is unallocated. */
+static struct page *typed_arena_scratch_page(const struct bpf_typed_arena *ta, unsigned long addr)
+{
+	unsigned long off = addr - (unsigned long)ta->base;
+
+	return ta->scratch_pages[(off & (bpf_typed_arena_chunk(ta) - 1)) >> PAGE_SHIFT];
+}
+
 /*
  * The memory a typed arena takes for the map's lifetime: a page table entry
- * per page of the slice and the two chunk bitmaps.
+ * per page of the slice, the scratch chunk with its page array, and the two
+ * chunk bitmaps.
  */
 static u64 typed_arena_static_mem(const struct bpf_typed_arena *ta)
 {
 	return (bpf_typed_arena_size(ta) >> PAGE_SHIFT) * sizeof(pte_t) +
+	       bpf_typed_arena_chunk(ta) + typed_arena_chunk_pages(ta) * sizeof(struct page *) +
 	       2 * BITS_TO_LONGS(typed_arena_nr_chunks(ta)) * sizeof(long);
 }
 
@@ -371,7 +396,7 @@ static int typed_arena_teardown_cb(pte_t *ptep, unsigned long addr, void *data)
 
 	if ((addr - (unsigned long)ta->base) & (bpf_typed_arena_chunk(ta) - 1))
 		return 0;
-	if (!pte_present(pte))
+	if (!pte_present(pte) || pte_page(pte) == typed_arena_scratch_page(ta, addr))
 		return 0;
 	typed_arena_free_objects(ta, (void *)addr);
 	__free_pages(pte_page(pte), typed_arena_chunk_order(ta));
@@ -383,6 +408,10 @@ static void typed_arena_free(struct bpf_arena *arena, struct bpf_typed_arena *ta
 	WRITE_ONCE(arena->typed_arena_mem, arena->typed_arena_mem - typed_arena_static_mem(ta));
 	apply_to_existing_page_range(&init_mm, (unsigned long)ta->base, bpf_typed_arena_size(ta),
 				     typed_arena_teardown_cb, ta);
+	/* Programs may have stored kptrs into the dummy objects. */
+	typed_arena_free_objects(ta, ta->scratch);
+	vfree(ta->scratch);
+	kfree(ta->scratch_pages);
 	bitmap_free(ta->chunks);
 	bitmap_free(ta->pending);
 	btf_put(ta->btf);
@@ -408,11 +437,14 @@ struct bpf_typed_arena *bpf_typed_arena_get(struct bpf_map *map, struct btf *btf
 {
 	struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
 	const struct btf_type *t = btf_type_by_id(btf, btf_id);
+	struct mem_cgroup *new_memcg, *old_memcg;
 	struct bpf_typed_arena *ta;
 	u64 slot, chunk;
 	s64 off;
-	int err;
+	int err, i;
 
+	if (!TYPED_ARENA_SUPPORTED)
+		return ERR_PTR(-EOPNOTSUPP);
 	if (!t || !t->size || !record)
 		return ERR_PTR(-EINVAL);
 	if (!is_power_of_2(size) || size < PAGE_SIZE || size > TYPED_ARENA_MAX_SZ)
@@ -444,10 +476,17 @@ struct bpf_typed_arena *bpf_typed_arena_get(struct bpf_map *map, struct btf *btf
 	ta->chunk_shift = ilog2(chunk);
 	ta->chunks = bitmap_zalloc(typed_arena_nr_chunks(ta), GFP_KERNEL_ACCOUNT);
 	ta->pending = bitmap_zalloc(typed_arena_nr_chunks(ta), GFP_KERNEL_ACCOUNT);
-	if (!ta->chunks || !ta->pending) {
+	ta->scratch_pages = kcalloc(typed_arena_chunk_pages(ta), sizeof(*ta->scratch_pages),
+				    GFP_KERNEL_ACCOUNT);
+	bpf_map_memcg_enter(map, &old_memcg, &new_memcg);
+	ta->scratch = __vmalloc(chunk, GFP_KERNEL_ACCOUNT | __GFP_ZERO);
+	bpf_map_memcg_exit(old_memcg, new_memcg);
+	if (!ta->chunks || !ta->pending || !ta->scratch_pages || !ta->scratch) {
 		err = -ENOMEM;
 		goto free;
 	}
+	for (i = 0; i < typed_arena_chunk_pages(ta); i++)
+		ta->scratch_pages[i] = vmalloc_to_page(ta->scratch + i * PAGE_SIZE);
 	err = apply_to_page_range(&init_mm, (unsigned long)ta->base, size, apply_range_set_cb, NULL);
 	if (err)
 		goto free;
@@ -462,6 +501,8 @@ struct bpf_typed_arena *bpf_typed_arena_get(struct bpf_map *map, struct btf *btf
 	return ta;
 
 free:
+	vfree(ta->scratch);
+	kfree(ta->scratch_pages);
 	bitmap_free(ta->chunks);
 	bitmap_free(ta->pending);
 	kfree(ta);
@@ -1460,6 +1501,64 @@ static void __bpf_prog_report_arena_violation(struct bpf_prog *prog, bool write,
 	}));
 }
 
+static struct bpf_typed_arena *typed_arena_lookup(struct bpf_arena *arena, unsigned long addr)
+{
+	struct bpf_typed_arena *ta;
+
+	list_for_each_entry_rcu(ta, &arena->typed_arenas, node)
+		if (addr - (unsigned long)ta->base < bpf_typed_arena_size(ta))
+			return ta;
+	return NULL;
+}
+
+static void __bpf_prog_report_typed_arena_violation(struct bpf_prog *prog,
+						    const struct bpf_typed_arena *ta,
+						    bool write, unsigned long addr)
+{
+	const struct btf_type *t = btf_type_by_id(ta->btf, ta->btf_id);
+	struct bpf_stream_stage ss;
+
+	/* Use main prog for stream access */
+	prog = prog->aux->main_prog_aux->prog;
+
+	bpf_stream_stage(ss, prog, BPF_STDERR, ({
+		bpf_stream_printk(ss, "ERROR: Typed arena %s access to unallocated struct %s at 0x%lx\n",
+				  write ? "WRITE" : "READ", btf_name_by_offset(ta->btf, t->name_off),
+				  addr & ~((unsigned long)bpf_typed_arena_slot(ta) - 1));
+		bpf_stream_dump_stack(ss);
+	}));
+}
+
+/*
+ * A typed access reached an unbacked page, so the program used a pointer to an
+ * object nobody allocated, which the cast permits by design. The pointer must
+ * still denote an object of the type: back the page with the page of the
+ * scratch chunk at the same position, so that the dummy objects the program
+ * sees are laid out like real ones and a kptr field of one never aliases a
+ * scalar field of another. Mark the chunk taken first, so that the allocator
+ * never puts real pages where a program may be in the middle of using the
+ * dummy object; nothing could wait for that use to end, since the next stray
+ * access would fault the scratch page right back in. The mark is an atomic
+ * bit because the fault can happen in any context a program runs in.
+ */
+static bool typed_arena_handle_page_fault(struct bpf_arena *arena, struct bpf_prog *prog,
+					  unsigned long addr, bool is_write)
+{
+	unsigned long page_addr = addr & PAGE_MASK;
+	struct bpf_typed_arena *ta;
+
+	guard(rcu)();
+	ta = typed_arena_lookup(arena, page_addr);
+	if (!ta)
+		return false;
+	set_bit((page_addr - (unsigned long)ta->base) >> ta->chunk_shift, ta->chunks);
+	apply_to_page_range(&init_mm, page_addr, PAGE_SIZE, apply_range_set_scratch_cb,
+			    typed_arena_scratch_page(ta, page_addr));
+	flush_vmap_cache(page_addr, PAGE_SIZE);
+	__bpf_prog_report_typed_arena_violation(prog, ta, is_write, addr);
+	return true;
+}
+
 bool bpf_arena_handle_page_fault(unsigned long addr, bool is_write, unsigned long fault_ip)
 {
 	struct bpf_arena *arena;
@@ -1476,12 +1575,16 @@ bool bpf_arena_handle_page_fault(unsigned long addr, bool is_write, unsigned lon
 	if (!arena)
 		return false;
 
+	if (page_addr - typed_arena_region(arena) < TYPED_ARENA_REGION_SZ)
+		return typed_arena_handle_page_fault(arena, prog, addr, is_write);
+
 	kbase = bpf_arena_get_kern_vm_start(arena);
 
 	/*
 	 * Recovery covers the 4 GiB mappable band plus the upper half-guard.
-	 * Lower guard is unreachable from kfuncs; an address there indicates
-	 * a different bug class - leave it to the regular kernel oops path.
+	 * Lower guard is unreachable from typed accesses and from kfuncs; an
+	 * address there indicates a different bug class - leave it to the
+	 * regular kernel oops path.
 	 */
 	if (page_addr < kbase || page_addr >= kbase + SZ_4G + GUARD_SZ / 2)
 		return false;
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 04/16] bpf: Add the typed_arena_cast instruction
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (2 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 03/16] bpf: Back typed arena chunks with scratch on demand Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:55   ` sashiko-bot
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 05/16] bpf: Allow scalar and atomic access to typed arena objects Kumar Kartikeya Dwivedi
                   ` (11 subsequent siblings)
  15 siblings, 1 reply; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

Add the instruction that turns an untrusted 64-bit value into a
verifier-trusted pointer to a typed arena object:

  dst = typed_arena_cast(src, imm)

encoded as a 64-bit BPF_MOV with a reserved off value, the value in src, the
program-local BTF type ID of the struct in imm, and the pointer in dst,
which may be src. The compiler emits the instruction with a CO-RE type ID
relocation landing in imm, so the type is part of the instruction and a
cast names one typed arena on every path. The verifier registers the typed
arena for the struct on first sight, records it in the instruction's aux
data, and lowers the instruction after verification to the sanitizing
sequence of that typed arena, provided by the registry: the value is masked
with the typed arena's size less its slot size and the base is added. That
keeps the slot index bits and drops everything else, so any value lands on
the start of an object of the type. The result is trusted and never NULL.

Any value casts. A typed arena holds objects of one struct at every slot,
and a chunk nobody allocated reads as the zeroed scratch chunk, so the
verifier need not know where a value came from: a pointer loaded from the
raw arena that user space corrupted, a pointer of another type, an integer,
or a pointer to this type that arithmetic moved inside an object, which the
cast rounds back to the object. This is the whole trust model for values
that enter from untrusted memory, in four instructions on every entry, and
the reason no NULL check follows a cast.

Pointers to typed objects are stored as-is, in the raw arena, in maps, on
the stack and in typed objects. There is no handle form and no translation
on the way to memory: a load gives back the 64-bit value, and a cast makes a
pointer of it wherever the value is not trusted. A 32-bit view of a typed
pointer is what it is for any kernel pointer, a truncated scalar under the
leak rules, and storing the pointer in memory user space can read is a
pointer leak the arena's privilege already permits. The alternative, a
32-bit slot offset as the stored form, needs a conversion on every store and
a second instruction to obtain it, for no gain in what the cast can trust.

The registration validates the struct once per program: it must be a struct
with special fields, since a struct without them belongs in the raw arena,
and every special field must be one a typed arena supports, which for now is
a kptr; locks, timers, workqueues, lists, rbtrees, refcounts and uptrs are
refused as not supported. The typed arena's size comes from the struct's
"typed_arena_size:<bytes>" decl tag with the usual suffixes, 128 MiB by
default, and must be a power of two of at least a page, because the cast
bounds the value with one AND and a page is the smallest unit of backing.
The registry's answer, a slice of the typed arena region, is logged with its
slot, chunk, object count and size, so the padding a power-of-two slot costs
is visible. A program keeps a reference on each typed arena it registers
until its load fails, at which point the references are dropped before the
arena map's; a loaded program leaves the typed arena to the map for the
map's lifetime, so objects survive program reloads.

Object access, kptr fields and pointer fields come in the following patches;
until then dereferencing a typed pointer is refused.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 include/linux/bpf.h            |  17 +++
 include/linux/bpf_verifier.h   |  19 +++
 include/uapi/linux/bpf.h       |   6 +
 kernel/bpf/backtrack.c         |   8 +-
 kernel/bpf/core.c              |   2 +
 kernel/bpf/disasm.c            |   9 ++
 kernel/bpf/fixups.c            |  33 ++++++
 kernel/bpf/log.c               |   5 +-
 kernel/bpf/verifier.c          | 207 ++++++++++++++++++++++++++++++++-
 tools/include/uapi/linux/bpf.h |   6 +
 10 files changed, 306 insertions(+), 6 deletions(-)

diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index f06d138b1f57..307e0e7c9445 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -945,6 +945,9 @@ enum bpf_type_flag {
 	/* DYNPTR points to file */
 	DYNPTR_TYPE_FILE	= BIT(20 + BPF_BASE_TYPE_BITS),
 
+	/* MEM is an object in a typed arena, reached through a native pointer. */
+	MEM_ARENA		= BIT(21 + BPF_BASE_TYPE_BITS),
+
 	__BPF_TYPE_FLAG_MAX,
 	__BPF_TYPE_LAST_FLAG	= __BPF_TYPE_FLAG_MAX - 1,
 };
@@ -1916,6 +1919,9 @@ struct bpf_prog_aux {
 	u64 prog_array_member_cnt; /* counts how many times as member of prog_array */
 	struct mutex ext_mutex; /* mutex for freplace_link_cnt and prog_array_member_cnt */
 	struct bpf_arena *arena;
+	/* typed arenas this program casts to or allocates from; referenced until the load fails */
+	struct bpf_typed_arena **typed_arenas;
+	u32 typed_arena_cnt;
 	void (*recursion_detected)(struct bpf_prog *prog); /* callback if recursion is detected */
 	/* BTF_KIND_FUNC_PROTO for valid attach_btf_id */
 	const struct btf_type *attach_func_proto;
@@ -1991,6 +1997,17 @@ struct bpf_prog_aux {
 
 #define BPF_NR_CONTEXTS        4       /* normal, softirq, hardirq, NMI */
 
+/* The typed arena of a program-BTF struct this program registered, or NULL. */
+static inline struct bpf_typed_arena *bpf_prog_typed_arena(const struct bpf_prog_aux *aux, u32 btf_id)
+{
+	u32 i;
+
+	for (i = 0; i < aux->typed_arena_cnt; i++)
+		if (aux->typed_arenas[i]->btf_id == btf_id)
+			return aux->typed_arenas[i];
+	return NULL;
+}
+
 struct bpf_prog {
 	u16			pages;		/* Number of allocated pages */
 	u32			jited:1,	/* Is our filter JIT'ed? */
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index c775bd757706..135049628313 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -661,6 +661,7 @@ struct bpf_insn_aux_data {
 		u64 insert_off;
 	};
 	struct btf_struct_meta *kptr_struct_meta;
+	struct bpf_typed_arena *typed_arena; /* named by a cast or a typed arena kfunc call */
 	u64 map_key_state; /* constant (32 bit) key tracking for maps */
 	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
 	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
@@ -1462,6 +1463,23 @@ static inline bool type_is_ptr_alloc_obj(u32 type)
 	       !(type_flag(type) & PTR_UNTRUSTED);
 }
 
+/* A pointer to an object in a typed arena: trusted, never NULL once checked, offset within the object. */
+static inline bool type_is_typed_arena_obj(u32 type)
+{
+	return base_type(type) == PTR_TO_BTF_ID && type_flag(type) & MEM_ARENA;
+}
+
+/* An object of a program-BTF struct: allocated by the program, or in a typed arena. */
+static inline bool type_is_local_obj(u32 type)
+{
+	return type & (MEM_ALLOC | MEM_ARENA);
+}
+
+static inline bool insn_is_typed_arena_cast(const struct bpf_insn *insn)
+{
+	return insn->code == (BPF_ALU64 | BPF_MOV | BPF_X) && insn->off == BPF_TYPED_ARENA_CAST;
+}
+
 static inline bool type_is_non_owning_ref(u32 type)
 {
 	return type_is_ptr_alloc_obj(type) && type_flag(type) & NON_OWN_REF;
@@ -1824,6 +1842,7 @@ int bpf_optimize_bpf_loop(struct bpf_verifier_env *env);
 void bpf_opt_hard_wire_dead_code_branches(struct bpf_verifier_env *env);
 int bpf_opt_remove_dead_code(struct bpf_verifier_env *env);
 int bpf_opt_remove_nops(struct bpf_verifier_env *env);
+int bpf_lower_typed_arena_insns(struct bpf_verifier_env *env);
 int bpf_opt_subreg_zext_lo32_rnd_hi32(struct bpf_verifier_env *env, const union bpf_attr *attr);
 int bpf_convert_ctx_accesses(struct bpf_verifier_env *env);
 int bpf_jit_subprogs(struct bpf_verifier_env *env);
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 4687c3310996..4bfcd0143400 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -1423,6 +1423,12 @@ enum {
 
 enum bpf_addr_space_cast {
 	BPF_ADDR_SPACE_CAST = 1,
+	/*
+	 * dst = typed_arena_cast(src, imm): src holds any 64-bit value, dst
+	 * becomes a pointer to the object it names in the typed arena of the
+	 * struct whose program-BTF type ID is imm. dst may be src.
+	 */
+	BPF_TYPED_ARENA_CAST = 2,
 };
 
 /* flags for BPF_MAP_UPDATE_ELEM command */
diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
index 0e38b9575328..38984e52ee37 100644
--- a/kernel/bpf/backtrack.c
+++ b/kernel/bpf/backtrack.c
@@ -319,7 +319,7 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
 			 */
 			return 0;
 		} else if (opcode == BPF_MOV) {
-			if (BPF_SRC(insn->code) == BPF_X) {
+			if (BPF_SRC(insn->code) == BPF_X && insn->off != BPF_TYPED_ARENA_CAST) {
 				/* dreg = sreg or dreg = (s8, s16, s32)sreg
 				 * dreg needs precision after this insn
 				 * sreg needs precision before this insn
@@ -328,11 +328,13 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
 				if (sreg != BPF_REG_FP)
 					bt_set_reg(bt, sreg);
 			} else {
-				/* dreg = K
+				/* dreg = K, or dreg = typed_arena_cast(sreg, imm)
 				 * dreg needs precision after this insn.
 				 * Corresponding register is already marked
 				 * as precise=true in this verifier state.
-				 * No further markings in parent are necessary
+				 * No further markings in parent are necessary;
+				 * a cast yields a pointer that is safe for any
+				 * value of sreg, which needs no precision.
 				 */
 				bt_clear_reg(bt, dreg);
 			}
diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
index d3b8b626ec0f..619f7c6a778d 100644
--- a/kernel/bpf/core.c
+++ b/kernel/bpf/core.c
@@ -3085,6 +3085,8 @@ static void bpf_prog_free_deferred(struct work_struct *work)
 	aux = container_of(work, struct bpf_prog_aux, work);
 #ifdef CONFIG_BPF_SYSCALL
 	bpf_free_kfunc_btf_tab(aux->kfunc_btf_tab);
+	/* The typed arenas outlive the program; only the load's failure retracts them. */
+	kfree(aux->typed_arenas);
 #endif
 #ifdef CONFIG_CGROUP_BPF
 	if (aux->cgroup_atype != CGROUP_BPF_ATTACH_TYPE_INVALID)
diff --git a/kernel/bpf/disasm.c b/kernel/bpf/disasm.c
index 36d3228d7745..b2ac521a4fb2 100644
--- a/kernel/bpf/disasm.c
+++ b/kernel/bpf/disasm.c
@@ -175,6 +175,12 @@ static bool is_addr_space_cast(const struct bpf_insn *insn)
 		insn->off == BPF_ADDR_SPACE_CAST;
 }
 
+static bool is_typed_arena_cast(const struct bpf_insn *insn)
+{
+	return insn->code == (BPF_ALU64 | BPF_MOV | BPF_X) &&
+		insn->off == BPF_TYPED_ARENA_CAST;
+}
+
 /* Special (internal-only) form of mov, used to resolve per-CPU addrs:
  * dst_reg = src_reg + <percpu_base_off>
  * BPF_ADDR_PERCPU is used as a special insn->off value.
@@ -208,6 +214,9 @@ void print_bpf_insn(const struct bpf_insn_cbs *cbs,
 			verbose(cbs->private_data, "(%02x) r%d = addr_space_cast(r%d, %u, %u)",
 				insn->code, insn->dst_reg,
 				insn->src_reg, ((u32)insn->imm) >> 16, (u16)insn->imm);
+		} else if (is_typed_arena_cast(insn)) {
+			verbose(cbs->private_data, "(%02x) r%d = typed_arena_cast(r%d, %d)",
+				insn->code, insn->dst_reg, insn->src_reg, insn->imm);
 		} else if (is_mov_percpu_addr(insn)) {
 			verbose(cbs->private_data, "(%02x) r%d = &(void __percpu *)(r%d)",
 				insn->code, insn->dst_reg, insn->src_reg);
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 37cf130ebb57..0b4636498ddc 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -876,6 +876,39 @@ int bpf_opt_subreg_zext_lo32_rnd_hi32(struct bpf_verifier_env *env,
  *     struct __sk_buff    -> struct sk_buff
  *     struct bpf_sock_ops -> struct sock
  */
+/*
+ * Replace every typed_arena_cast with the sanitizing sequence of the typed
+ * arena the verifier registered for it. The type is part of the instruction,
+ * so a cast has exactly one typed arena by the time it gets here.
+ */
+int bpf_lower_typed_arena_insns(struct bpf_verifier_env *env)
+{
+	struct bpf_insn *insn = env->prog->insnsi;
+	int i, cnt, delta = 0, insn_cnt = env->prog->len;
+	const struct bpf_typed_arena *ta;
+	struct bpf_insn insn_buf[8];
+	struct bpf_prog *new_prog;
+
+	for (i = 0; i < insn_cnt; i++, insn++) {
+		if (!insn_is_typed_arena_cast(insn))
+			continue;
+		ta = env->insn_aux_data[i + delta].typed_arena;
+		if (!ta) {
+			verifier_bug(env, "typed_arena_cast at insn %d has no typed arena", i);
+			return -EFAULT;
+		}
+		cnt = bpf_typed_arena_cast_insns(ta, insn->dst_reg, insn->src_reg, insn_buf);
+		new_prog = bpf_patch_insn_data(env, i + delta, insn_buf, cnt);
+		if (!new_prog)
+			return -ENOMEM;
+		delta += cnt - 1;
+		env->prog = new_prog;
+		insn = new_prog->insnsi + i + delta;
+	}
+
+	return 0;
+}
+
 int bpf_convert_ctx_accesses(struct bpf_verifier_env *env)
 {
 	struct bpf_subprog_info *subprogs = env->subprog_info;
diff --git a/kernel/bpf/log.c b/kernel/bpf/log.c
index d850a7863d2e..1ae29a08a607 100644
--- a/kernel/bpf/log.c
+++ b/kernel/bpf/log.c
@@ -432,14 +432,15 @@ const char *reg_type_str(struct bpf_verifier_env *env, enum bpf_reg_type type)
 			strscpy(postfix, "_or_null");
 	}
 
-	snprintf(prefix, sizeof(prefix), "%s%s%s%s%s%s%s",
+	snprintf(prefix, sizeof(prefix), "%s%s%s%s%s%s%s%s",
 		 type & MEM_RDONLY ? "rdonly_" : "",
 		 type & MEM_RINGBUF ? "ringbuf_" : "",
 		 type & MEM_USER ? "user_" : "",
 		 type & MEM_PERCPU ? "percpu_" : "",
 		 type & MEM_RCU ? "rcu_" : "",
 		 type & PTR_UNTRUSTED ? "untrusted_" : "",
-		 type & PTR_TRUSTED ? "trusted_" : ""
+		 type & PTR_TRUSTED ? "trusted_" : "",
+		 type & MEM_ARENA ? "typed_arena_" : ""
 	);
 
 	snprintf(env->tmp_str_buf, TMP_STR_BUF_LEN, "%s%s%s",
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 03dbc0e00398..a0069983f103 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -6341,6 +6341,12 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
 	u32 btf_id = 0;
 	int ret;
 
+	/* The access rules for typed arena objects come with a later patch. */
+	if (type_is_typed_arena_obj(reg->type)) {
+		verbose(env, "typed arena access is not supported yet\n");
+		return -EACCES;
+	}
+
 	if (!env->allow_ptr_leaks) {
 		verbose(env,
 			"'struct %s' access is allowed only to CAP_PERFMON and CAP_SYS_ADMIN\n",
@@ -16934,6 +16940,180 @@ static int adjust_reg_min_max_vals(struct bpf_verifier_env *env,
 	return 0;
 }
 
+/*
+ * Every field kind program BTF can describe, and the ones a typed arena object
+ * may hold: those whose operations are single-word atomics, which is what
+ * keeps them sound while the chunk under the object comes and goes.
+ */
+#define BPF_TYPED_ARENA_FIELDS \
+	(BPF_SPIN_LOCK | BPF_RES_SPIN_LOCK | BPF_TIMER | BPF_KPTR | BPF_LIST_HEAD | \
+	 BPF_LIST_NODE | BPF_RB_ROOT | BPF_RB_NODE | BPF_REFCOUNT | BPF_WORKQUEUE | \
+	 BPF_UPTR | BPF_TASK_WORK | BPF_RCU_HEAD)
+#define BPF_TYPED_ARENA_SUPPORTED_FIELDS BPF_KPTR
+
+/*
+ * The size of a struct's typed arena is a declared resource, read from the
+ * struct's "typed_arena_size:<bytes>" decl tag with the usual K/M/G suffixes,
+ * with a default. It must be a power of two, so that the cast bounds a value
+ * with one AND, and at least a page, the smallest unit of backing. The upper
+ * bound is the typed arena region, which the registry enforces.
+ */
+static int typed_arena_size(struct bpf_verifier_env *env, const struct btf *btf,
+			    const struct btf_type *t, const char *tname, u64 *size,
+			    const char **value)
+{
+	char *end;
+	u64 sz;
+
+	*value = btf_find_decl_tag_value(btf, t, -1, BPF_TYPED_ARENA_SIZE_TAG);
+	if (IS_ERR(*value)) {
+		if (PTR_ERR(*value) == -ENOENT) {
+			*value = "default";
+			*size = BPF_TYPED_ARENA_DEFAULT_SIZE;
+			return 0;
+		}
+		verbose(env, "struct %s has conflicting typed arena size declarations\n", tname);
+		return PTR_ERR(*value);
+	}
+	sz = memparse(*value, &end);
+	if (end == *value || *end || !is_power_of_2(sz) || sz < PAGE_SIZE) {
+		verbose(env, "struct %s has invalid typed arena size '%s'\n", tname, *value);
+		return -EINVAL;
+	}
+	*size = sz;
+	return 0;
+}
+
+/*
+ * Register the typed arena for a program-BTF struct the first time this
+ * program names it, and hold a reference on it until the load has either
+ * succeeded, after which the typed arena lives as long as the map, or failed.
+ */
+static struct bpf_typed_arena *typed_arena_register(struct bpf_verifier_env *env, u32 btf_id)
+{
+	struct bpf_prog_aux *aux = env->prog->aux;
+	struct bpf_typed_arena *ta, **tas;
+	struct btf_struct_meta *meta;
+	struct btf *btf = aux->btf;
+	const struct btf_type *t;
+	struct btf_record *record;
+	const char *tname, *value;
+	u64 size;
+	u32 i;
+	int err;
+
+	ta = bpf_prog_typed_arena(aux, btf_id);
+	if (ta)
+		return ta;
+
+	t = btf_type_by_id(btf, btf_id);
+	if (!t || !__btf_type_is_struct(t)) {
+		verbose(env, "typed_arena_cast type ID %u is not a struct\n", btf_id);
+		return ERR_PTR(-EINVAL);
+	}
+	tname = btf_name_by_offset(btf, t->name_off);
+
+	record = btf_parse_fields(btf, t, BPF_TYPED_ARENA_FIELDS, t->size);
+	if (IS_ERR(record)) {
+		verbose(env, "struct %s has invalid special fields\n", tname);
+		return ERR_CAST(record);
+	}
+	if (!record) {
+		verbose(env, "struct %s has no special fields and needs no typed arena\n", tname);
+		return ERR_PTR(-EINVAL);
+	}
+	for (i = 0; i < record->cnt; i++) {
+		if (record->fields[i].type & BPF_TYPED_ARENA_SUPPORTED_FIELDS)
+			continue;
+		verbose(env, "struct %s field %s is not supported in a typed arena\n", tname,
+			btf_field_type_name(record->fields[i].type));
+		btf_record_free(record);
+		return ERR_PTR(-EOPNOTSUPP);
+	}
+	btf_record_free(record);
+	/* BTF keeps a record for every struct with these fields. */
+	meta = btf_find_struct_meta(btf, btf_id);
+	if (!meta) {
+		verifier_bug(env, "struct %s has special fields but no metadata", tname);
+		return ERR_PTR(-EFAULT);
+	}
+
+	err = typed_arena_size(env, btf, t, tname, &size, &value);
+	if (err)
+		return ERR_PTR(err);
+
+	tas = krealloc_array(aux->typed_arenas, aux->typed_arena_cnt + 1, sizeof(*tas),
+			     GFP_KERNEL_ACCOUNT);
+	if (!tas)
+		return ERR_PTR(-ENOMEM);
+	aux->typed_arenas = tas;
+
+	ta = bpf_typed_arena_get(bpf_prog_arena(env->prog), btf, btf_id, meta->record, size);
+	if (IS_ERR(ta)) {
+		switch (PTR_ERR(ta)) {
+		case -E2BIG:
+			verbose(env, "struct %s does not fit its typed arena: slot %lu bytes, size %llu bytes\n",
+				tname, roundup_pow_of_two(t->size), size);
+			break;
+		case -EINVAL:
+			verbose(env, "struct %s has invalid typed arena size '%s'\n", tname, value);
+			break;
+		case -ENOSPC:
+			verbose(env, "no room in the typed arena region for struct %s\n", tname);
+			break;
+		case -EOPNOTSUPP:
+			verbose(env, "typed arenas are not supported on this architecture\n");
+			break;
+		default:
+			verbose(env, "cannot register a typed arena for struct %s: %ld\n",
+				tname, PTR_ERR(ta));
+		}
+		return ta;
+	}
+	aux->typed_arenas[aux->typed_arena_cnt++] = ta;
+	verbose(env, "typed arena for struct %s: slot %u bytes, chunk %u bytes, %llu objects, size %llu bytes\n",
+		tname, bpf_typed_arena_slot(ta), bpf_typed_arena_chunk(ta),
+		bpf_typed_arena_size(ta) >> ta->slot_shift, bpf_typed_arena_size(ta));
+	return ta;
+}
+
+/*
+ * dst = typed_arena_cast(src, imm): sanitize the value in src into a pointer
+ * to an object of the struct whose program-BTF type ID is imm. Any value
+ * casts, since the lowering masks it to a slot inside the typed arena, so the
+ * result is trusted and never NULL; a pointer that already names an object of
+ * the type comes back rounded to the object it points into. The type is part
+ * of the instruction, so an instruction casts to one type on every path, and
+ * the lowering can be chosen once.
+ */
+static int check_typed_arena_cast(struct bpf_verifier_env *env, struct bpf_insn *insn)
+{
+	struct bpf_insn_aux_data *aux = &env->insn_aux_data[env->insn_idx];
+	struct bpf_reg_state *regs = cur_regs(env);
+	struct bpf_reg_state *dst = &regs[insn->dst_reg];
+	struct bpf_typed_arena *ta;
+
+	/* The arena itself needs CAP_PERFMON, so the leak rules already permit a kernel pointer. */
+	if (!env->prog->aux->arena) {
+		verbose(env, "typed_arena_cast insn can only be used in a program that has an associated arena\n");
+		return -EINVAL;
+	}
+	if (!env->prog->aux->btf) {
+		verbose(env, "typed_arena_cast insn requires program BTF\n");
+		return -EINVAL;
+	}
+	ta = typed_arena_register(env, insn->imm);
+	if (IS_ERR(ta))
+		return PTR_ERR(ta);
+	aux->typed_arena = ta;
+
+	mark_reg_known_zero(env, regs, insn->dst_reg);
+	dst->type = PTR_TO_BTF_ID | MEM_ARENA;
+	dst->btf = env->prog->aux->btf;
+	dst->btf_id = ta->btf_id;
+	return 0;
+}
+
 /* check validity of 32-bit and 64-bit arithmetic operations */
 static int check_alu_op(struct bpf_verifier_env *env, struct bpf_insn *insn)
 {
@@ -16993,7 +17173,9 @@ static int check_alu_op(struct bpf_verifier_env *env, struct bpf_insn *insn)
 			struct bpf_reg_state *dst_reg = regs + insn->dst_reg;
 
 			if (BPF_CLASS(insn->code) == BPF_ALU64) {
-				if (insn->imm) {
+				if (insn->off == BPF_TYPED_ARENA_CAST) {
+					return check_typed_arena_cast(env, insn);
+				} else if (insn->imm) {
 					/* off == BPF_ADDR_SPACE_CAST */
 					mark_reg_unknown(env, regs, insn->dst_reg);
 					if (insn->imm == 1) /* cast from as(1) to as(0) */
@@ -20172,6 +20354,8 @@ static int check_alu_fields(struct bpf_verifier_env *env, struct bpf_insn *insn)
 					verbose(env, "addr_space_cast insn can only convert between address space 1 and 0\n");
 					return -EINVAL;
 				}
+			} else if (insn->off == BPF_TYPED_ARENA_CAST) {
+				/* imm holds the type ID; src is the value, and dst may be src */
 			} else if ((insn->off != 0 && insn->off != 8 &&
 				    insn->off != 16 && insn->off != 32) || insn->imm) {
 				verbose(env, "BPF_MOV uses reserved fields\n");
@@ -20580,8 +20764,26 @@ static int resolve_func_ptrs(struct bpf_verifier_env *env)
 }
 
 /* drop refcnt of maps used by the rejected program */
+/*
+ * Drop the rejected program's typed arena references before its arena map
+ * reference. A typed arena nobody else registered is retracted; one that a
+ * loaded program registered lives on for the map's lifetime.
+ */
+static void release_typed_arenas(struct bpf_verifier_env *env)
+{
+	struct bpf_prog_aux *aux = env->prog->aux;
+	u32 i;
+
+	for (i = 0; i < aux->typed_arena_cnt; i++)
+		bpf_typed_arena_put(bpf_prog_arena(env->prog), aux->typed_arenas[i]);
+	kfree(aux->typed_arenas);
+	aux->typed_arenas = NULL;
+	aux->typed_arena_cnt = 0;
+}
+
 static void release_maps(struct bpf_verifier_env *env)
 {
+	release_typed_arenas(env);
 	__bpf_free_used_maps(env->prog->aux, env->used_maps,
 			     env->used_map_cnt);
 }
@@ -22688,6 +22890,9 @@ int bpf_check(struct bpf_prog **prog, union bpf_attr *attr, bpfptr_t uattr,
 			sanitize_dead_code(env);
 	}
 
+	if (ret == 0)
+		ret = bpf_lower_typed_arena_insns(env);
+
 	if (ret == 0)
 		/* program is valid, convert *(u32*)(ctx + off) accesses */
 		ret = bpf_convert_ctx_accesses(env);
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 4687c3310996..4bfcd0143400 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -1423,6 +1423,12 @@ enum {
 
 enum bpf_addr_space_cast {
 	BPF_ADDR_SPACE_CAST = 1,
+	/*
+	 * dst = typed_arena_cast(src, imm): src holds any 64-bit value, dst
+	 * becomes a pointer to the object it names in the typed arena of the
+	 * struct whose program-BTF type ID is imm. dst may be src.
+	 */
+	BPF_TYPED_ARENA_CAST = 2,
 };
 
 /* flags for BPF_MAP_UPDATE_ELEM command */
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 05/16] bpf: Allow scalar and atomic access to typed arena objects
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (3 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 04/16] bpf: Add the typed_arena_cast instruction Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 06/16] bpf: Support special fields in " Kumar Kartikeya Dwivedi
                   ` (10 subsequent siblings)
  15 siblings, 0 replies; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

Let a program read and write the ordinary fields of a typed arena object,
and run atomics on them, through the native pointer the cast produced.
Every byte the pointer can reach is mapped, either with real objects or
with the type's scratch chunk, so the accesses are plain loads and stores:
no probe mode, no exception table entry, and no address check at the
access. The verifier's bounds are the whole check. It already keeps a BTF
pointer's offset constant and inside the object, and the object is inside
its slot by construction, so a typed pointer never leaves its typed arena.
This is what the sanitizing cast buys: the cost of trust is paid once, where
a value enters, and every access after it is as cheap as a load from the
stack.

Reuse the rules for program-allocated objects. A typed arena object is a
program-BTF struct with a special-field record, as an allocated object is,
so give both the same treatment where the code asked whether a pointer is
allocated: reads and writes of scalar fields are permitted, an access that
overlaps a special field is rejected, a pointer field loads as a scalar
rather than as a pointer, and flexible arrays are not walked. Writes are
permitted because the memory under the object is never unmapped while a
program that saw it can still run, which is what allocated objects need a
reference for. The fault-prone dereference predicate does not match the
class, so the context conversion pass leaves the accesses alone and
load-acquire is allowed on it. Helpers and kfuncs keep rejecting the class
as plain memory, since none of the argument type tables lists it; a kfunc
that wants a typed arena pointer asks for it by type, as the page kfuncs and
the kptr exchange do.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 kernel/bpf/btf.c      |  9 +++++----
 kernel/bpf/verifier.c | 18 ++++++++----------
 2 files changed, 13 insertions(+), 14 deletions(-)

diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index 9bcfefdfb734..6a29a9d87702 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -7782,7 +7782,7 @@ int btf_struct_access(struct bpf_verifier_log *log,
 	u32 id = reg->btf_id;
 	int err;
 
-	while (type_is_alloc(reg->type)) {
+	while (type_is_local_obj(reg->type)) {
 		struct btf_struct_meta *meta;
 		struct btf_record *rec;
 		int i;
@@ -7807,14 +7807,15 @@ int btf_struct_access(struct bpf_verifier_log *log,
 	t = btf_type_by_id(btf, id);
 	do {
 		err = btf_struct_walk(log, btf, t, off, size, &id, &tmp_flag,
-				      field_name, !type_is_alloc(reg->type));
+				      field_name, !type_is_local_obj(reg->type));
 
 		switch (err) {
 		case WALK_PTR:
-			/* For local types, the destination register cannot
+			/*
+			 * For local types, the destination register cannot
 			 * become a pointer again.
 			 */
-			if (type_is_alloc(reg->type))
+			if (type_is_local_obj(reg->type))
 				return SCALAR_VALUE;
 			/* If we found the pointer or scalar on t+off,
 			 * we're done.
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index a0069983f103..75697e52a2df 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -6341,12 +6341,6 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
 	u32 btf_id = 0;
 	int ret;
 
-	/* The access rules for typed arena objects come with a later patch. */
-	if (type_is_typed_arena_obj(reg->type)) {
-		verbose(env, "typed arena access is not supported yet\n");
-		return -EACCES;
-	}
-
 	if (!env->allow_ptr_leaks) {
 		verbose(env,
 			"'struct %s' access is allowed only to CAP_PERFMON and CAP_SYS_ADMIN\n",
@@ -6398,7 +6392,7 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
 		return -EACCES;
 	}
 
-	if (env->ops->btf_struct_access && !type_is_alloc(reg->type) && atype == BPF_WRITE) {
+	if (env->ops->btf_struct_access && !type_is_local_obj(reg->type) && atype == BPF_WRITE) {
 		if (!btf_is_kernel(reg->btf)) {
 			verifier_bug(env, "reg->btf must be kernel btf");
 			return -EFAULT;
@@ -6409,10 +6403,14 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
 				"%s cannot write into ptr_%s at off=%d size=%d\n",
 				reg_arg_name(env, argno), tname, off, size);
 	} else {
-		/* Writes are permitted with default btf_struct_access for
-		 * program allocated objects (which always have id > 0).
+		/*
+		 * Writes are permitted with default btf_struct_access for
+		 * program allocated objects (which always have id > 0) and
+		 * for typed arena objects, whose memory is never unmapped
+		 * under a program.
 		 */
-		if (atype != BPF_READ && !type_is_ptr_alloc_obj(reg->type)) {
+		if (atype != BPF_READ && !type_is_ptr_alloc_obj(reg->type) &&
+		    !type_is_typed_arena_obj(reg->type)) {
 			verbose(env, "only read is supported\n");
 			return -EACCES;
 		}
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 06/16] bpf: Support special fields in typed arena objects
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (4 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 05/16] bpf: Allow scalar and atomic access to typed arena objects Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:59   ` sashiko-bot
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 07/16] bpf: Trust typed pointer fields of " Kumar Kartikeya Dwivedi
                   ` (9 subsequent siblings)
  15 siblings, 1 reply; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

Let bpf_kptr_xchg() take a pointer to a referenced or percpu kptr field of a
typed arena object, so that objects in arena memory can own references to
kernel objects and to program-allocated objects. The exchange is a single
atomic word swap on a native address, so it needs nothing from the arena:
the field's record comes from the struct's BTF as it does for an allocated
object, the destination may carry the field's offset, and the value is
matched against the field's type as before. Direct loads and stores of the
field stay rejected, as they are for allocated objects; the exchange is the
only access. A reference left in an object is dropped when its chunk is
released or the map is destroyed, including one stored into a dummy object
of the scratch chunk by a program that used a pointer nobody allocated.

Kptrs are the only special fields a typed arena object may hold, and the
reason is how its memory behaves. A chunk is released after a grace period
that covers the invocations that started before the release was requested,
not the ones that start after it and cast into the range, and a fault in
the middle of a kfunc swaps the page under the object for the scratch
chunk, so the rest of the kfunc runs on the dummy object that every
unallocated slot shares. A field survives that only if each operation on it
is one atomic instruction, so that it lands on the real page or on the
dummy but is never split, and if the kernel keeps no pointer into the
object, so that nothing dangles once the memory is gone or reused. A kptr
is exactly that: the exchange is the only operation, and the referenced
object is held through the pointer value alone.

Timers, workqueues, task work and RCU heads are not. Their kfuncs take a
lock in the field and touch it several times, the async callback object
records the object's address for the callback, and an RCU head is linked
into the RCU callback list in place. A release or a fault in the middle
leaves a lock taken on a real page and dropped on the dummy, a callback
running on memory that now holds another object, or the dummy's RCU head
queued twice. Locks, lists, rbtrees and refcounts stay refused as well; the
memory is native, and a program builds those from typed pointer fields and
atomics without kernel help.

The way to give a typed object a timer, a workqueue, task work or an RCU
head is to allocate an object holding them with bpf_obj_new() and keep it
in a kptr field of the typed object. Allocated memory is owned by the
allocator and freed only by bpf_obj_drop() after its fields are cancelled,
so nothing swaps it or reuses it under a kfunc or a callback. That needs
those fields to be allowed in allocated objects, which a later patch adds.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 kernel/bpf/verifier.c | 6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 75697e52a2df..f854d8419fff 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -386,7 +386,7 @@ static struct btf_record *reg_btf_record(const struct bpf_reg_state *reg)
 
 	if (reg->type == PTR_TO_MAP_VALUE) {
 		rec = reg->map_ptr->record;
-	} else if (type_is_ptr_alloc_obj(reg->type)) {
+	} else if (type_is_ptr_alloc_obj(reg->type) || type_is_typed_arena_obj(reg->type)) {
 		meta = btf_find_struct_meta(reg->btf, reg->btf_id);
 		if (meta)
 			rec = meta->record;
@@ -8063,7 +8063,7 @@ static int process_kptr_func(struct bpf_verifier_env *env, int regno,
 	struct btf_record *rec;
 	u32 kptr_off;
 
-	if (type_is_ptr_alloc_obj(reg->type)) {
+	if (type_is_ptr_alloc_obj(reg->type) || type_is_typed_arena_obj(reg->type)) {
 		rec = reg_btf_record(reg);
 	} else { /* PTR_TO_MAP_VALUE */
 		map_ptr = reg->map_ptr;
@@ -8834,6 +8834,7 @@ static const struct bpf_reg_types kptr_xchg_dest_types = {
 		PTR_TO_BTF_ID | MEM_ALLOC,
 		PTR_TO_BTF_ID | MEM_ALLOC | NON_OWN_REF,
 		PTR_TO_BTF_ID | MEM_ALLOC | NON_OWN_REF | MEM_RCU,
+		PTR_TO_BTF_ID | MEM_ARENA,
 	}
 };
 static const struct bpf_reg_types dynptr_types = {
@@ -9154,6 +9155,7 @@ static int check_func_arg_reg_off(struct bpf_verifier_env *env,
 	case PTR_TO_BTF_ID | MEM_RCU:
 	case PTR_TO_BTF_ID | MEM_ALLOC | NON_OWN_REF:
 	case PTR_TO_BTF_ID | MEM_ALLOC | NON_OWN_REF | MEM_RCU:
+	case PTR_TO_BTF_ID | MEM_ARENA:
 		/* When referenced PTR_TO_BTF_ID is passed to release function,
 		 * its fixed offset must be 0. bpf_refcount_acquire() returns the
 		 * pointer it was given while incrementing the refcount at the
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 07/16] bpf: Trust typed pointer fields of typed arena objects
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (5 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 06/16] bpf: Support special fields in " Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-27  0:03   ` sashiko-bot
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 08/16] bpf: Canonicalize loaded typed arena pointers where they are used Kumar Kartikeya Dwivedi
                   ` (8 subsequent siblings)
  15 siblings, 1 reply; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

A pointer member of a typed arena object whose pointee is itself a struct
with a typed arena is a typed pointer field: it holds a pointer to an object
of that struct, or 0, and nothing else, because the verifier lets nothing
else be written there. Typed objects live in memory only programs can
write, every store into such a field is a 64-bit store of a typed arena
pointer to the pointee struct at offset zero or of the constant 0, and a
chunk comes back zeroed when it is reused. A stale pointer, to an object
whose chunk was released and claimed again, still names a whole object of
the type. So a 64-bit load from the field gives back a value that is an
object of the pointee type or 0 without a check, and a linked structure in
a typed arena is walked with plain loads: no cast per hop, no annotation.

Detect the fields from BTF alone. A struct is typed when it has a special
field, a kptr today, or a pointer member to a typed struct, taken to the
closure: the closure runs when the program's BTF is parsed, before the
struct records, and is kept in the BTF, so that parsing a record can ask
whether a pointee is typed. A struct whose only pointers are to itself and
that has no special field of its own never joins, and pointer members
tagged as arena pointers are user addresses and stay scalars. The field
becomes a new record kind, BPF_TYPED_PTR, eight bytes aligned to eight like
a kptr, allowed in arrays and nested structs, with nothing to acquire,
release or free. It counts as a special field, so a struct holding only a
pointer to a typed struct, a list head for instance, gets a typed arena of
its own. The same member of a struct allocated with bpf_obj_new() keeps its
old meaning, a plain pointer that loads as a scalar, and so does the same
shape in a map value or in the raw arena, whose memory is not disciplined.

The load yields a new form of typed arena pointer, PTR_UNSANITIZED: the
value is trusted to be an object or 0 but is not canonical, since nothing
masked it into its slot, and it may be 0, which no NULL check is required
to exclude. It survives 64-bit copies, spills and fills, and 64-bit
compares on the raw value, so a loop ends with the usual test against 0. It
may be stored back into a typed pointer field or anywhere a pointer store is
allowed. What it cannot do yet is serve as an address: a load or store
through it, arithmetic on it, or passing it to a call is refused until the
next patch teaches the lowering to sanitize it in place at that point.

Narrower loads and stores of a typed pointer field, stores of anything but
a typed pointer to the pointee or 0, and atomics on the field are rejected,
as the field's invariant depends on them. A struct whose special fields are
only inherited from a nested struct has no record of its own and is refused
at registration rather than treated as a verifier bug.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 include/linux/bpf.h          |  12 ++++
 include/linux/bpf_verifier.h |   6 ++
 kernel/bpf/btf.c             | 131 +++++++++++++++++++++++++++++++----
 kernel/bpf/log.c             |   3 +-
 kernel/bpf/syscall.c         |   3 +
 kernel/bpf/verifier.c        | 124 +++++++++++++++++++++++++++++++--
 6 files changed, 260 insertions(+), 19 deletions(-)

diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 307e0e7c9445..f4c65fabedbf 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -218,6 +218,7 @@ enum btf_field_type {
 	BPF_RES_SPIN_LOCK = (1 << 12),
 	BPF_TASK_WORK  = (1 << 13),
 	BPF_RCU_HEAD   = (1 << 14),
+	BPF_TYPED_PTR  = (1 << 15),
 };
 
 enum bpf_cgroup_storage_type {
@@ -451,6 +452,8 @@ static inline const char *btf_field_type_name(enum btf_field_type type)
 		return "bpf_task_work";
 	case BPF_RCU_HEAD:
 		return "bpf_rcu_head";
+	case BPF_TYPED_PTR:
+		return "typed_ptr";
 	default:
 		WARN_ON_ONCE(1);
 		return "unknown";
@@ -478,6 +481,7 @@ static inline u32 btf_field_type_size(enum btf_field_type type)
 	case BPF_KPTR_REF:
 	case BPF_KPTR_PERCPU:
 	case BPF_UPTR:
+	case BPF_TYPED_PTR:
 		return sizeof(u64);
 	case BPF_LIST_HEAD:
 		return sizeof(struct bpf_list_head);
@@ -514,6 +518,7 @@ static inline u32 btf_field_type_align(enum btf_field_type type)
 	case BPF_KPTR_REF:
 	case BPF_KPTR_PERCPU:
 	case BPF_UPTR:
+	case BPF_TYPED_PTR:
 		return __alignof__(u64);
 	case BPF_LIST_HEAD:
 		return __alignof__(struct bpf_list_head);
@@ -562,6 +567,7 @@ static inline void bpf_obj_init_field(const struct btf_field *field, void *addr)
 	case BPF_UPTR:
 	case BPF_TASK_WORK:
 	case BPF_RCU_HEAD:
+	case BPF_TYPED_PTR:
 		break;
 	default:
 		WARN_ON_ONCE(1);
@@ -948,6 +954,12 @@ enum bpf_type_flag {
 	/* MEM is an object in a typed arena, reached through a native pointer. */
 	MEM_ARENA		= BIT(21 + BPF_BASE_TYPE_BITS),
 
+	/*
+	 * MEM_ARENA pointer as loaded from a typed pointer field: an object of
+	 * the type or 0, not yet masked into its slot.
+	 */
+	PTR_UNSANITIZED		= BIT(22 + BPF_BASE_TYPE_BITS),
+
 	__BPF_TYPE_FLAG_MAX,
 	__BPF_TYPE_LAST_FLAG	= __BPF_TYPE_FLAG_MAX - 1,
 };
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 135049628313..77e7c4b45a1f 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1475,6 +1475,12 @@ static inline bool type_is_local_obj(u32 type)
 	return type & (MEM_ALLOC | MEM_ARENA);
 }
 
+/* A typed arena pointer as a typed pointer field holds it, not yet rounded into its slot. */
+static inline bool type_is_unsanitized_arena_obj(u32 type)
+{
+	return type_is_typed_arena_obj(type) && type_flag(type) & PTR_UNSANITIZED;
+}
+
 static inline bool insn_is_typed_arena_cast(const struct bpf_insn *insn)
 {
 	return insn->code == (BPF_ALU64 | BPF_MOV | BPF_X) && insn->off == BPF_TYPED_ARENA_CAST;
diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index 6a29a9d87702..20338a2657cd 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -272,6 +272,8 @@ struct btf {
 	struct btf_struct_metas *struct_meta_tab;
 	struct btf_struct_ops_tab *struct_ops_tab;
 	struct btf_layout *layout;
+	/* structs with special fields, closed over plain pointers to them; see btf_find_kptr() */
+	unsigned long *typed_structs;
 
 	/* split BTF support */
 	struct btf *base_btf;
@@ -1890,6 +1892,7 @@ static void btf_free_struct_ops_tab(struct btf *btf)
 static void btf_free(struct btf *btf)
 {
 	btf_free_struct_meta_tab(btf);
+	bitmap_free(btf->typed_structs);
 	btf_free_dtor_kfunc_tab(btf);
 	btf_free_kfunc_set_tab(btf);
 	btf_free_struct_ops_tab(btf);
@@ -3578,6 +3581,11 @@ bool btf_type_is_arena_ptr(const struct btf *btf, const struct btf_type *t)
 	return false;
 }
 
+static bool btf_struct_is_typed(const struct btf *btf, u32 id)
+{
+	return btf->typed_structs && id < btf_nr_types(btf) && test_bit(id, btf->typed_structs);
+}
+
 static int btf_find_kptr(const struct btf *btf, const struct btf_type *t,
 			 u32 off, int sz, struct btf_field_info *info, u32 field_mask)
 {
@@ -3589,6 +3597,7 @@ static int btf_find_kptr(const struct btf *btf, const struct btf_type *t,
 	};
 	struct btf_type_tag_walk_ctx ctx;
 	enum btf_field_type type = 0;
+	const struct btf_type *ptr;
 	int err;
 	u32 res_id;
 
@@ -3598,6 +3607,7 @@ static int btf_find_kptr(const struct btf *btf, const struct btf_type *t,
 	/* For PTR, sz is always == 8 */
 	if (!btf_type_is_ptr(t))
 		return BTF_FIELD_IGNORE;
+	ptr = t;
 
 	ctx.t = t;
 	err = btf_type_tag_walk(btf, &ctx, kptr_type_tags,
@@ -3609,6 +3619,15 @@ static int btf_find_kptr(const struct btf *btf, const struct btf_type *t,
 	res_id = ctx.id;
 	type = ctx.res;
 
+	/*
+	 * A plain pointer to a struct that has a typed arena is a typed pointer
+	 * field: an object of that struct, or NULL, whatever the program wrote.
+	 * Arena pointers are user addresses and stay plain scalars.
+	 */
+	if (!type && field_mask & BPF_TYPED_PTR && __btf_type_is_struct(t) &&
+	    !btf_type_is_arena_ptr(btf, ptr) && btf_struct_is_typed(btf, res_id))
+		type = BPF_TYPED_PTR;
+
 	if (!(type & field_mask))
 		return BTF_FIELD_IGNORE;
 
@@ -3748,7 +3767,7 @@ static int btf_get_field_type(const struct btf *btf, const struct btf_type *var_
 	}
 
 	/* Only return BPF_KPTR when all other types with matchable names fail */
-	if (field_mask & (BPF_KPTR | BPF_UPTR) && !__btf_type_is_struct(var_type)) {
+	if (field_mask & (BPF_KPTR | BPF_UPTR | BPF_TYPED_PTR) && !__btf_type_is_struct(var_type)) {
 		type = BPF_KPTR_REF;
 		goto end;
 	}
@@ -3780,6 +3799,7 @@ static int btf_repeat_fields(struct btf_field_info *info, int info_cnt,
 		case BPF_KPTR_REF:
 		case BPF_KPTR_PERCPU:
 		case BPF_UPTR:
+		case BPF_TYPED_PTR:
 		case BPF_LIST_HEAD:
 		case BPF_RB_ROOT:
 			break;
@@ -3915,6 +3935,7 @@ static int btf_find_field_one(const struct btf *btf,
 	case BPF_KPTR_REF:
 	case BPF_KPTR_PERCPU:
 	case BPF_UPTR:
+	case BPF_TYPED_PTR:
 		ret = btf_find_kptr(btf, var_type, off, sz,
 				    info_cnt ? &info[0] : &tmp, field_mask);
 		if (ret < 0)
@@ -4262,6 +4283,11 @@ struct btf_record *btf_parse_fields(const struct btf *btf, const struct btf_type
 			if (ret < 0)
 				goto end;
 			break;
+		case BPF_TYPED_PTR:
+			/* The pointee is in this BTF, which outlives the record. */
+			rec->fields[i].kptr.btf = (struct btf *)btf;
+			rec->fields[i].kptr.btf_id = info_arr[i].kptr.type_id;
+			break;
 		case BPF_LIST_HEAD:
 			ret = btf_parse_list_head(btf, &rec->fields[i], &info_arr[i]);
 			if (ret < 0)
@@ -6162,12 +6188,29 @@ static const char * const alloc_obj_fields[] = {
 	"bpf_refcount",
 };
 
+/*
+ * Whether a struct holds a typed pointer field, given the typed structs known
+ * so far. Every struct of the BTF is scanned, kernel structs pulled in from
+ * vmlinux.h included, and the scan can fail on members a typed arena object
+ * could not have, such as pointers with kernel type tags: a failure does not
+ * make a struct typed. A struct that a program registers is parsed again then,
+ * and that parse reports what is wrong with it.
+ */
+static bool btf_struct_has_typed_ptr(const struct btf *btf, const struct btf_type *t)
+{
+	struct btf_field_info info[BTF_FIELDS_MAX];
+
+	return btf_find_field(btf, t, BPF_TYPED_PTR, info, ARRAY_SIZE(info)) > 0;
+}
+
 static struct btf_struct_metas *
 btf_parse_struct_metas(struct bpf_verifier_log *log, struct btf *btf)
 {
 	struct btf_struct_metas *tab = NULL;
+	unsigned long *typed = NULL;
 	struct btf_id_set *aof;
 	int i, n, id, ret;
+	bool changed;
 
 	BUILD_BUG_ON(offsetof(struct btf_id_set, cnt) != 0);
 	BUILD_BUG_ON(sizeof(struct btf_id_set) != sizeof(u32));
@@ -6231,26 +6274,65 @@ btf_parse_struct_metas(struct bpf_verifier_log *log, struct btf *btf)
 	}
 	sort(&aof->ids, aof->cnt, sizeof(aof->ids[0]), btf_id_cmp_func, NULL);
 
+	/*
+	 * The structs with special fields, and their closure over plain
+	 * pointers: a pointer member to a struct with special fields is a typed
+	 * pointer field, itself a special field, so the struct holding it joins
+	 * the set, and so on until nothing changes. A struct whose only pointers
+	 * are to itself, with no special field of its own, never joins. The set
+	 * is published in the btf before the records are parsed, since parsing
+	 * consults it for every pointer member.
+	 */
+	typed = bitmap_zalloc(n, GFP_KERNEL | __GFP_NOWARN);
+	if (!typed) {
+		ret = -ENOMEM;
+		goto free_aof;
+	}
 	for (i = 1; i < n; i++) {
-		struct btf_struct_metas *new_tab;
 		const struct btf_member *member;
-		struct btf_struct_meta *type;
-		struct btf_record *record;
 		const struct btf_type *t;
-		int j, tab_cnt;
+		int j;
 
 		t = btf_type_by_id(btf, i);
 		if (!__btf_type_is_struct(t))
 			continue;
+		for_each_member(j, t, member) {
+			if (btf_id_set_contains(aof, member->type)) {
+				set_bit(i, typed);
+				break;
+			}
+		}
+	}
+	btf->typed_structs = typed;
+	do {
+		changed = false;
+		for (i = 1; i < n; i++) {
+			const struct btf_type *t;
+
+			t = btf_type_by_id(btf, i);
+			if (!__btf_type_is_struct(t) || test_bit(i, typed))
+				continue;
+			cond_resched();
+			if (btf_struct_has_typed_ptr(btf, t)) {
+				set_bit(i, typed);
+				changed = true;
+			}
+		}
+	} while (changed);
+
+	for (i = 1; i < n; i++) {
+		struct btf_struct_metas *new_tab;
+		struct btf_struct_meta *type;
+		struct btf_record *record;
+		const struct btf_type *t;
+		int tab_cnt;
+
+		if (!test_bit(i, typed))
+			continue;
+		t = btf_type_by_id(btf, i);
 
 		cond_resched();
 
-		for_each_member(j, t, member) {
-			if (btf_id_set_contains(aof, member->type))
-				goto parse;
-		}
-		continue;
-	parse:
 		tab_cnt = tab ? tab->cnt : 0;
 		new_tab = krealloc(tab, struct_size(new_tab, types, tab_cnt + 1),
 				   GFP_KERNEL | __GFP_NOWARN);
@@ -6266,10 +6348,12 @@ btf_parse_struct_metas(struct bpf_verifier_log *log, struct btf *btf)
 		type->btf_id = i;
 		record = btf_parse_fields(btf, t, BPF_SPIN_LOCK | BPF_RES_SPIN_LOCK | BPF_LIST_HEAD | BPF_LIST_NODE |
 						  BPF_RB_ROOT | BPF_RB_NODE | BPF_REFCOUNT |
-						  BPF_KPTR, t->size);
+						  BPF_KPTR | BPF_TYPED_PTR, t->size);
 		/* The record cannot be unset, treat it as an error if so */
 		if (IS_ERR_OR_NULL(record)) {
 			ret = PTR_ERR_OR_ZERO(record) ?: -EFAULT;
+			bpf_log(log, "struct %s has an invalid layout of special fields: %d\n",
+				__btf_name_by_offset(btf, t->name_off), ret);
 			goto free;
 		}
 		type->record = record;
@@ -6279,6 +6363,8 @@ btf_parse_struct_metas(struct bpf_verifier_log *log, struct btf *btf)
 	return tab;
 free:
 	btf_struct_metas_free(tab);
+	bitmap_free(typed);
+	btf->typed_structs = NULL;
 free_aof:
 	kfree(aof);
 	return ERR_PTR(ret);
@@ -7782,6 +7868,7 @@ int btf_struct_access(struct bpf_verifier_log *log,
 	u32 id = reg->btf_id;
 	int err;
 
+	t = btf_type_by_id(btf, id);
 	while (type_is_local_obj(reg->type)) {
 		struct btf_struct_meta *meta;
 		struct btf_record *rec;
@@ -7795,6 +7882,25 @@ int btf_struct_access(struct bpf_verifier_log *log,
 			struct btf_field *field = &rec->fields[i];
 			u32 offset = field->offset;
 			if (off < offset + field->size && offset < off + size) {
+				if (field->type == BPF_TYPED_PTR) {
+					/*
+					 * A typed pointer field is only special in a
+					 * typed arena object, where it is read and
+					 * written whole. In an allocated object it is
+					 * the plain pointer member it always was.
+					 */
+					if (!type_is_typed_arena_obj(reg->type))
+						continue;
+					if (off != offset || size != field->size) {
+						bpf_log(log,
+							"typed pointer field of struct %s must be accessed with a 64-bit load or store\n",
+							__btf_name_by_offset(btf, t->name_off));
+						return -EACCES;
+					}
+					*next_btf_id = field->kptr.btf_id;
+					*flag = MEM_ARENA | PTR_UNSANITIZED;
+					return PTR_TO_BTF_ID;
+				}
 				bpf_log(log,
 					"direct access to %s is disallowed\n",
 					btf_field_type_name(field->type));
@@ -7804,7 +7910,6 @@ int btf_struct_access(struct bpf_verifier_log *log,
 		break;
 	}
 
-	t = btf_type_by_id(btf, id);
 	do {
 		err = btf_struct_walk(log, btf, t, off, size, &id, &tmp_flag,
 				      field_name, !type_is_local_obj(reg->type));
diff --git a/kernel/bpf/log.c b/kernel/bpf/log.c
index 1ae29a08a607..f2319e9f0fa6 100644
--- a/kernel/bpf/log.c
+++ b/kernel/bpf/log.c
@@ -432,7 +432,7 @@ const char *reg_type_str(struct bpf_verifier_env *env, enum bpf_reg_type type)
 			strscpy(postfix, "_or_null");
 	}
 
-	snprintf(prefix, sizeof(prefix), "%s%s%s%s%s%s%s%s",
+	snprintf(prefix, sizeof(prefix), "%s%s%s%s%s%s%s%s%s",
 		 type & MEM_RDONLY ? "rdonly_" : "",
 		 type & MEM_RINGBUF ? "ringbuf_" : "",
 		 type & MEM_USER ? "user_" : "",
@@ -440,6 +440,7 @@ const char *reg_type_str(struct bpf_verifier_env *env, enum bpf_reg_type type)
 		 type & MEM_RCU ? "rcu_" : "",
 		 type & PTR_UNTRUSTED ? "untrusted_" : "",
 		 type & PTR_TRUSTED ? "trusted_" : "",
+		 type & PTR_UNSANITIZED ? "unsanitized_" : "",
 		 type & MEM_ARENA ? "typed_arena_" : ""
 	);
 
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index ac52f4ae414c..c989a29e1ba7 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -688,6 +688,7 @@ void btf_record_free(struct btf_record *rec)
 		case BPF_WORKQUEUE:
 		case BPF_TASK_WORK:
 		case BPF_RCU_HEAD:
+		case BPF_TYPED_PTR:
 			/* Nothing to release */
 			break;
 		default:
@@ -743,6 +744,7 @@ struct btf_record *btf_record_dup(const struct btf_record *rec)
 		case BPF_WORKQUEUE:
 		case BPF_TASK_WORK:
 		case BPF_RCU_HEAD:
+		case BPF_TYPED_PTR:
 			/* Nothing to acquire */
 			break;
 		default:
@@ -877,6 +879,7 @@ void bpf_obj_free_fields(const struct btf_record *rec, void *obj)
 		case BPF_RB_NODE:
 		case BPF_REFCOUNT:
 		case BPF_RCU_HEAD:
+		case BPF_TYPED_PTR:
 			break;
 		default:
 			WARN_ON_ONCE(1);
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index f854d8419fff..5ca5fc4696d9 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -358,6 +358,9 @@ static bool reg_not_null(struct bpf_verifier_env *env, const struct bpf_reg_stat
 	type = reg->type;
 	if (type_may_be_null(type))
 		return false;
+	/* A typed pointer field holds an object or 0, and its load is neither checked nor masked. */
+	if (type_flag(type) & PTR_UNSANITIZED)
+		return false;
 
 	/*
 	 * The types below guarantee a non-NULL base, an unbounded offset can
@@ -379,6 +382,23 @@ static bool reg_not_null(struct bpf_verifier_env *env, const struct bpf_reg_stat
 		type == CONST_PTR_TO_MAP;
 }
 
+/*
+ * @regno is about to serve as an address, or go to a call. A typed arena
+ * pointer loaded from a typed pointer field is not canonical yet: its slice
+ * and slot are what the field's discipline promises, not what the lowering has
+ * masked, and until the lowering learns to sanitize it in place such a use is
+ * refused.
+ */
+static int typed_arena_use(struct bpf_verifier_env *env, int regno)
+{
+	struct bpf_reg_state *reg = &cur_regs(env)[regno];
+
+	if (!type_is_unsanitized_arena_obj(reg->type))
+		return 0;
+	verbose(env, "R%d unsanitized typed arena pointer cannot be dereferenced\n", regno);
+	return -EACCES;
+}
+
 static struct btf_record *reg_btf_record(const struct bpf_reg_state *reg)
 {
 	struct btf_record *rec = NULL;
@@ -6328,6 +6348,61 @@ static bool type_is_trusted_or_null(struct bpf_verifier_env *env,
 					  "__safe_trusted_or_null");
 }
 
+/*
+ * A typed pointer field of a typed arena object holds a pointer to an object
+ * of the pointee struct, or 0, and nothing else. Only a 64-bit store of a
+ * typed arena pointer to that struct at offset 0, sanitized or not, or of the
+ * constant 0, may write it, and only a plain 64-bit load may read it. The load
+ * yields the value as stored: an unsanitized typed arena pointer, which the
+ * write discipline makes trustworthy without a check.
+ */
+static int check_typed_ptr_field_access(struct bpf_verifier_env *env, struct bpf_reg_state *regs,
+					struct bpf_reg_state *reg, const char *tname, u32 btf_id,
+					enum bpf_access_type atype, int value_regno)
+{
+	struct bpf_insn *insn = &env->prog->insnsi[env->insn_idx];
+	u8 class = BPF_CLASS(insn->code);
+	struct bpf_reg_state *val;
+	struct bpf_typed_arena *ta;
+
+	if (BPF_MODE(insn->code) != BPF_MEM || BPF_SIZE(insn->code) != BPF_DW ||
+	    (class != BPF_LDX && class != BPF_STX && class != BPF_ST)) {
+		verbose(env, "typed pointer field of struct %s must be accessed with a 64-bit load or store\n",
+			tname);
+		return -EACCES;
+	}
+	if (atype == BPF_READ) {
+		/*
+		 * The pointer leads into the pointee's typed arena, which the
+		 * program may never cast to or allocate from: register it here,
+		 * as a cast would, so that the sanitization has a slice to mask
+		 * into and the slice exists for the map's lifetime.
+		 */
+		ta = typed_arena_register(env, btf_id);
+		if (IS_ERR(ta))
+			return PTR_ERR(ta);
+		return mark_btf_ld_reg(env, regs, value_regno, PTR_TO_BTF_ID, reg->btf, btf_id,
+				       MEM_ARENA | PTR_UNSANITIZED);
+	}
+
+	if (class == BPF_ST) {
+		if (!insn->imm)
+			return 0;
+	} else {
+		val = &regs[value_regno];
+		if (val->type == SCALAR_VALUE && tnum_is_const(val->var_off) && !val->var_off.value)
+			return 0;
+		if (type_is_typed_arena_obj(val->type) &&
+		    !(type_flag(val->type) & ~(MEM_ARENA | PTR_UNSANITIZED)) &&
+		    val->btf == reg->btf && val->btf_id == btf_id &&
+		    tnum_is_const(val->var_off) && !val->var_off.value)
+			return 0;
+	}
+	verbose(env, "store into typed pointer field of struct %s expects a typed arena pointer to struct %s or NULL\n",
+		tname, btf_name_by_offset(reg->btf, btf_type_by_id(reg->btf, btf_id)->name_off));
+	return -EACCES;
+}
+
 static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
 				   struct bpf_reg_state *regs, struct bpf_reg_state *reg,
 				   argno_t argno, int off, int size,
@@ -6433,6 +6508,9 @@ static int check_ptr_to_btf_access(struct bpf_verifier_env *env,
 	if (ret < 0)
 		return ret;
 
+	if (ret == PTR_TO_BTF_ID && flag & MEM_ARENA)
+		return check_typed_ptr_field_access(env, regs, reg, tname, btf_id, atype, value_regno);
+
 	if (ret != PTR_TO_BTF_ID) {
 		/* just mark; */
 
@@ -7170,6 +7248,10 @@ static int check_load_mem(struct bpf_verifier_env *env, struct bpf_insn *insn,
 	if (err)
 		return err;
 
+	err = typed_arena_use(env, insn->src_reg);
+	if (err)
+		return err;
+
 	/* check dst operand */
 	err = check_reg_arg(env, insn->dst_reg, DST_OP_NO_MARK);
 	if (err)
@@ -7221,6 +7303,10 @@ static int check_store_reg(struct bpf_verifier_env *env, struct bpf_insn *insn,
 	if (err)
 		return err;
 
+	err = typed_arena_use(env, insn->dst_reg);
+	if (err)
+		return err;
+
 	dst_reg_type = regs[insn->dst_reg].type;
 
 	/* Check if (dst_reg + off) is writeable. */
@@ -7273,6 +7359,10 @@ static int check_atomic_rmw(struct bpf_verifier_env *env,
 		return -EACCES;
 	}
 
+	err = typed_arena_use(env, insn->dst_reg);
+	if (err)
+		return err;
+
 	if (!atomic_ptr_type_ok(env, insn->dst_reg, insn)) {
 		verbose(env, "BPF_ATOMIC stores into R%d %s is not allowed\n",
 			insn->dst_reg,
@@ -9386,6 +9476,13 @@ static int check_func_arg(struct bpf_verifier_env *env, u32 arg, u32 slot, u32 p
 		return 0;
 	}
 
+	/* Every register argument reaches the callee, an ignored one included. */
+	if (regno >= 0) {
+		err = typed_arena_use(env, regno);
+		if (err)
+			return err;
+	}
+
 	if (arg_type == ARG_IGNORE)
 		return 0;
 
@@ -16811,6 +16908,15 @@ static int adjust_reg_min_max_vals(struct bpf_verifier_env *env,
 		aux->prevent_zext = true;
 	}
 
+	/* Arithmetic moves a typed arena pointer within its object, which needs the object. */
+	if (BPF_CLASS(insn->code) == BPF_ALU64) {
+		err = typed_arena_use(env, insn->dst_reg);
+		if (!err && src_reg)
+			err = typed_arena_use(env, insn->src_reg);
+		if (err)
+			return err;
+	}
+
 	if (dst_reg->type != SCALAR_VALUE)
 		ptr_reg = dst_reg;
 
@@ -16948,8 +17054,8 @@ static int adjust_reg_min_max_vals(struct bpf_verifier_env *env,
 #define BPF_TYPED_ARENA_FIELDS \
 	(BPF_SPIN_LOCK | BPF_RES_SPIN_LOCK | BPF_TIMER | BPF_KPTR | BPF_LIST_HEAD | \
 	 BPF_LIST_NODE | BPF_RB_ROOT | BPF_RB_NODE | BPF_REFCOUNT | BPF_WORKQUEUE | \
-	 BPF_UPTR | BPF_TASK_WORK | BPF_RCU_HEAD)
-#define BPF_TYPED_ARENA_SUPPORTED_FIELDS BPF_KPTR
+	 BPF_UPTR | BPF_TASK_WORK | BPF_RCU_HEAD | BPF_TYPED_PTR)
+#define BPF_TYPED_ARENA_SUPPORTED_FIELDS (BPF_KPTR | BPF_TYPED_PTR)
 
 /*
  * The size of a struct's typed arena is a declared resource, read from the
@@ -17031,11 +17137,15 @@ static struct bpf_typed_arena *typed_arena_register(struct bpf_verifier_env *env
 		return ERR_PTR(-EOPNOTSUPP);
 	}
 	btf_record_free(record);
-	/* BTF keeps a record for every struct with these fields. */
+	/*
+	 * BTF keeps a record for every struct with these fields as direct
+	 * members; one that only inherits them from a nested struct has none.
+	 */
 	meta = btf_find_struct_meta(btf, btf_id);
 	if (!meta) {
-		verifier_bug(env, "struct %s has special fields but no metadata", tname);
-		return ERR_PTR(-EFAULT);
+		verbose(env, "struct %s has special fields only in a nested struct, which a typed arena does not support\n",
+			tname);
+		return ERR_PTR(-EOPNOTSUPP);
 	}
 
 	err = typed_arena_size(env, btf, t, tname, &size, &value);
@@ -19569,6 +19679,10 @@ static int do_check_insn(struct bpf_verifier_env *env, bool *do_print_state)
 		if (err)
 			return err;
 
+		err = typed_arena_use(env, insn->dst_reg);
+		if (err)
+			return err;
+
 		dst_reg_type = cur_regs(env)[insn->dst_reg].type;
 
 		err = check_mem_access(env, env->insn_idx, cur_regs(env) + insn->dst_reg, argno_from_reg(insn->dst_reg),
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 08/16] bpf: Canonicalize loaded typed arena pointers where they are used
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (6 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 07/16] bpf: Trust typed pointer fields of " Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 09/16] bpf: Add typed arena page allocation and release kfuncs Kumar Kartikeya Dwivedi
                   ` (7 subsequent siblings)
  15 siblings, 0 replies; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

A typed arena pointer loaded from a typed pointer field is an object of its
type or 0, but it is not canonical: nothing masked it into its slot, and it
may be 0. The previous patch refused to use it as an address. Sanitize it
instead where it is used, in place: when an instruction loads or stores
through such a pointer, or an atomic runs on it, when arithmetic moves it,
or when it goes to a helper or kfunc, the lowering prepends the typed
arena's cast sequence to that instruction, on that register, and on that
path the register is canonical from then on. Compares, copies, spills and
stores of the value as a value need nothing, so a list is walked with plain
loads and ends with the usual test against 0, and a pointer that is only
passed along is never touched.

This is the whole of the NULL question. The field's discipline says the
value is an object or 0. A program that tests for 0 sees the raw value and
takes its branch. A program that does not test, or tests wrongly, uses the
value as an address, the sequence maps 0 to object 0 of the slice, and the
access lands on a whole object of the type, the same outcome as feeding 0
to a cast. Nothing the program does with the value can reach outside the
slice, so the verifier asks for no annotation and no check: the SFI scheme
decides the result, not the type system, and a NULL check is a matter of
program logic alone. The alternative, a maybe-NULL pointer type that must
be checked before every use, would put a check on every hop of a
traversal that the sanitization makes unnecessary.

The sequence is lowered once per instruction and runs on every path through
it, so it must be harmless on every path. It is idempotent on a canonical
pointer of the same typed arena at offset zero, which is what a path that
got its pointer from a cast or from an earlier sanitization brings. It is
wrong on a pointer with an offset into its object, since the mask rounds to
the object base, on a pointer of another typed arena, whose base and mask
differ, and on any other value. So the verifier records the register and
its typed arena in the instruction's aux data when a path needs the
sequence, and rejects the program when another path brings that register
in a form the sequence would corrupt, or when a second register needs it at
the same instruction, which no compiler-generated code does. Arithmetic
sanitizes before it moves the pointer for the same reason: the offset must
be added to the object base, not masked away later.

The later optimization phase builds on the unsanitized form: a load through
such a pointer that nothing writes to can be a probed load instead of a
sanitized one, and a chain of reads then costs nothing beyond the loads. A
store of a typed pointer into a typed object always sanitizes its
destination, since the field's discipline depends on the store landing in a
real object.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 include/linux/bpf_verifier.h | 10 +++++
 kernel/bpf/fixups.c          | 38 +++++++++++++------
 kernel/bpf/verifier.c        | 71 ++++++++++++++++++++++++++++++++----
 3 files changed, 99 insertions(+), 20 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 77e7c4b45a1f..bf0f23929671 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -662,6 +662,16 @@ struct bpf_insn_aux_data {
 	};
 	struct btf_struct_meta *kptr_struct_meta;
 	struct bpf_typed_arena *typed_arena; /* named by a cast or a typed arena kfunc call */
+	/*
+	 * Lazy sanitization of a typed arena pointer used as an address here:
+	 * the register and its typed arena, whether a path brought the pointer
+	 * unsanitized, and the registers some path brought in a form the
+	 * sanitizing sequence would corrupt.
+	 */
+	struct bpf_typed_arena *sanitize_arena;
+	u16 sanitize_plain;
+	u8 sanitize_reg;
+	bool sanitize_needed;
 	u64 map_key_state; /* constant (32 bit) key tracking for maps */
 	int ctx_field_size; /* the ctx field size for load insn, maybe 0 */
 	u32 seen; /* this insn was processed by the verifier at env->pass_cnt */
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 0b4636498ddc..fd0ba8376c1c 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -877,27 +877,41 @@ int bpf_opt_subreg_zext_lo32_rnd_hi32(struct bpf_verifier_env *env,
  *     struct bpf_sock_ops -> struct sock
  */
 /*
- * Replace every typed_arena_cast with the sanitizing sequence of the typed
- * arena the verifier registered for it. The type is part of the instruction,
- * so a cast has exactly one typed arena by the time it gets here.
+ * Lower the instructions the verifier tied to a typed arena. A typed_arena_cast
+ * becomes the sanitizing sequence of the typed arena the verifier registered
+ * for it; the type is part of the instruction, so a cast has exactly one. An
+ * instruction that uses a typed pointer loaded from a typed pointer field as
+ * an address gets the same sequence prepended, in place on that register: the
+ * verifier checked that every path brings that register here as an
+ * unsanitized pointer of that typed arena, or as a canonical one at offset
+ * zero, on which the sequence is a no-op.
  */
 int bpf_lower_typed_arena_insns(struct bpf_verifier_env *env)
 {
 	struct bpf_insn *insn = env->prog->insnsi;
 	int i, cnt, delta = 0, insn_cnt = env->prog->len;
-	const struct bpf_typed_arena *ta;
-	struct bpf_insn insn_buf[8];
+	struct bpf_insn insn_buf[16];
+	struct bpf_insn_aux_data *aux;
 	struct bpf_prog *new_prog;
 
 	for (i = 0; i < insn_cnt; i++, insn++) {
-		if (!insn_is_typed_arena_cast(insn))
-			continue;
-		ta = env->insn_aux_data[i + delta].typed_arena;
-		if (!ta) {
-			verifier_bug(env, "typed_arena_cast at insn %d has no typed arena", i);
-			return -EFAULT;
+		aux = &env->insn_aux_data[i + delta];
+		cnt = 0;
+		if (aux->sanitize_needed)
+			cnt = bpf_typed_arena_cast_insns(aux->sanitize_arena, aux->sanitize_reg,
+							 aux->sanitize_reg, insn_buf);
+		if (insn_is_typed_arena_cast(insn)) {
+			if (!aux->typed_arena) {
+				verifier_bug(env, "typed_arena_cast at insn %d has no typed arena", i);
+				return -EFAULT;
+			}
+			cnt += bpf_typed_arena_cast_insns(aux->typed_arena, insn->dst_reg,
+							  insn->src_reg, insn_buf + cnt);
+		} else if (cnt) {
+			insn_buf[cnt++] = *insn;
 		}
-		cnt = bpf_typed_arena_cast_insns(ta, insn->dst_reg, insn->src_reg, insn_buf);
+		if (!cnt)
+			continue;
 		new_prog = bpf_patch_insn_data(env, i + delta, insn_buf, cnt);
 		if (!new_prog)
 			return -ENOMEM;
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 5ca5fc4696d9..2d406034556e 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -383,20 +383,75 @@ static bool reg_not_null(struct bpf_verifier_env *env, const struct bpf_reg_stat
 }
 
 /*
- * @regno is about to serve as an address, or go to a call. A typed arena
- * pointer loaded from a typed pointer field is not canonical yet: its slice
- * and slot are what the field's discipline promises, not what the lowering has
- * masked, and until the lowering learns to sanitize it in place such a use is
- * refused.
+ * @regno is about to serve as an address, be moved by arithmetic, or go to a
+ * call. A typed arena pointer loaded from a typed pointer field is not
+ * canonical: it is an object of its type or 0, as the field's discipline
+ * promises, but nothing masked it into its slot. Sanitize it here, in place:
+ * the lowering prepends the typed arena's cast sequence to this instruction,
+ * and on this path the register is canonical from here on. The sequence maps
+ * 0 to object 0 of the slice, as the cast maps any value, so a NULL the
+ * program did not test costs no more than a stray cast would.
+ *
+ * The sequence is lowered once per instruction and runs on every path through
+ * it, so it must be harmless on every path. It is a no-op on a canonical
+ * pointer of the same typed arena at offset zero, and wrong on anything else:
+ * a path that brings another typed arena, a pointer with an offset into its
+ * object, or a value of another kind to this register here is rejected, and
+ * so is a second register in need of the sequence at the same instruction.
+ * Where nothing brings an unsanitized pointer, nothing is lowered.
  */
 static int typed_arena_use(struct bpf_verifier_env *env, int regno)
 {
+	struct bpf_insn_aux_data *aux = &env->insn_aux_data[env->insn_idx];
 	struct bpf_reg_state *reg = &cur_regs(env)[regno];
+	struct bpf_typed_arena *ta;
+	bool canonical;
+
+	canonical = !(type_flag(reg->type) & PTR_UNSANITIZED);
+	if (!type_is_typed_arena_obj(reg->type) ||
+	    (canonical && (!tnum_is_const(reg->var_off) || reg->var_off.value))) {
+		if (aux->sanitize_needed && aux->sanitize_reg == regno) {
+			verbose(env, "insn %d sanitizes a typed arena pointer only on some paths\n",
+				env->insn_idx);
+			return -EINVAL;
+		}
+		aux->sanitize_plain |= BIT(regno);
+		return 0;
+	}
 
-	if (!type_is_unsanitized_arena_obj(reg->type))
+	ta = bpf_prog_typed_arena(env->prog->aux, reg->btf_id);
+	if (verifier_bug_if(!ta, env, "R%d typed arena pointer has no typed arena", regno))
+		return -EFAULT;
+	if (aux->sanitize_arena && aux->sanitize_reg == regno && aux->sanitize_arena != ta) {
+		verbose(env, "insn %d sanitizes typed arena pointers of different types on different paths\n",
+			env->insn_idx);
+		return -EINVAL;
+	}
+	if (aux->sanitize_arena && aux->sanitize_reg != regno) {
+		if (canonical)
+			return 0;
+		if (aux->sanitize_needed) {
+			verbose(env, "insn %d needs more than one typed arena pointer sanitized\n",
+				env->insn_idx);
+			return -EINVAL;
+		}
+		/* A canonical pointer only booked the register; the one in need takes it. */
+		aux->sanitize_arena = NULL;
+	}
+	if (!aux->sanitize_arena) {
+		aux->sanitize_arena = ta;
+		aux->sanitize_reg = regno;
+	}
+	if (canonical)
 		return 0;
-	verbose(env, "R%d unsanitized typed arena pointer cannot be dereferenced\n", regno);
-	return -EACCES;
+	if (aux->sanitize_plain & BIT(regno)) {
+		verbose(env, "insn %d sanitizes a typed arena pointer only on some paths\n",
+			env->insn_idx);
+		return -EINVAL;
+	}
+	aux->sanitize_needed = true;
+	reg->type &= ~PTR_UNSANITIZED;
+	return 0;
 }
 
 static struct btf_record *reg_btf_record(const struct bpf_reg_state *reg)
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 09/16] bpf: Add typed arena page allocation and release kfuncs
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (7 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 08/16] bpf: Canonicalize loaded typed arena pointers where they are used Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:55   ` sashiko-bot
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 10/16] bpf: Let typed_arena_cast copy pointers the verifier already trusts Kumar Kartikeya Dwivedi
                   ` (6 subsequent siblings)
  15 siblings, 1 reply; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

Give programs control over the memory behind a typed arena. So far every
chunk of one reads as the dummy objects of the scratch chunk once touched;
now a program can back a range with real objects and give it back:

  void *bpf_typed_arena_alloc_pages(void *map, u64 local_type_id__k,
                                    void *addr, u32 *page_cnt, int node_id);
  void bpf_typed_arena_free_pages(void *map, u64 local_type_id__k,
                                  void *ptr, u32 page_cnt);

The prototypes follow the raw page kfuncs, with the struct's local BTF type
ID naming the typed arena inside the map as a verifier-known constant, as
bpf_obj_new() takes its type. The verifier requires the program to have an
associated arena map and BTF, registers the typed arena for the struct as a
cast does, records it in the call's aux data, and refuses a call site that
two paths reach with two type IDs, since the fixup replaces the type ID
register with the typed arena once per call site: the kfunc receives the
registered typed arena and only checks that it belongs to the map it was
given. The address is a typed pointer naming the chunk of the first object,
NULL for any range, ignored by the verifier and validated by the kfunc; a
loaded, unsanitized pointer passed there is sanitized before the call like
any call argument. The request is in pages, like the raw allocator's, but
the unit of backing is the chunk, so the request is rounded up to whole
chunks and the count granted is written back through the pointer: an object
larger than a page is never half allocated, and a program learns how many
objects it got. The generic kfunc argument rules classify a pointer to a
scalar as writable fixed-size memory, so a spilled constant in that slot
does not survive the call as known. The allocator returns a typed pointer to
the first object of the range, or NULL, checked once and then used natively,
and the objects follow at slot strides. Both kfuncs are safe under a spin
lock: the allocator takes the arena's resilient spinlock and uses the
lock-free allocators, and the release only queues work. The flags argument
of the raw allocator is left out to stay within five arguments; a variant
can add it, kfuncs not being ABI.

Allocation claims the chunk range in the type's bitmap under the spinlock
and installs each chunk as one allocation of its order into empty entries
only: a chunk the fault path faulted to scratch is skipped by a search or
fails a fixed request, because a scratch page is never replaced by a real
one. The fault path marks its chunk without the lock, so it can still slip
in between the claim and the install; the chunks installed so far are then
backed out under the lock, the chunk it took keeps its mark, and a search
retries a few times while a fixed request fails. The bits are set one at a
time so that a word-wide update cannot undo the fault path's atomic mark.
One allocation per chunk also removes the array of pages the raw allocator
sizes against the lock-free kmalloc limit, and makes a multi-page object
contiguous in the direct map, which is what lets its special fields be
dropped after its mapping is gone.

Release is always deferred. The chunks stay mapped and marked until a worker
has waited for an RCU and an RCU-tasks-trace grace period, so every
invocation that started before the release keeps valid objects, sleepable
ones included. The worker then clears the entries and the marks under the
spinlock, flushes the TLB, drops the special fields of every object of each
real chunk from its block, and frees it. A range is released once: the
release marks its chunks pending under the spinlock, and a further release of
any of them is refused until the worker has cleared the marks, since a second
release queued behind the first would run after the first has let the range
be claimed again and take it away from its new owner. An invocation that
casts into the range after the release was requested is racing with it: it
stays memory safe, and it sees the real objects before the clear and the
dummy object after. Releasing a chunk that the fault path took only clears
its entries, after which it can be claimed or faulted again. A release that
cannot queue its work in a non-sleepable context leaves the chunks allocated
until the map is freed. Map teardown flushes the worker.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 kernel/bpf/arena.c    | 391 ++++++++++++++++++++++++++++++++++++++++++
 kernel/bpf/verifier.c |  60 +++++++
 2 files changed, 451 insertions(+)

diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index c757c1932360..0b8edfda3c1a 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -11,6 +11,7 @@
 #include <linux/vmalloc.h>
 #include <linux/pagemap.h>
 #include <asm/tlbflush.h>
+#include <linux/rcupdate_wait.h>
 #include "range_tree.h"
 
 /*
@@ -103,6 +104,9 @@ struct bpf_arena {
 	struct irq_work     free_irq;
 	struct work_struct  free_work;
 	struct llist_head   free_spans;
+	struct irq_work     typed_free_irq;
+	struct work_struct  typed_free_work;
+	struct llist_head   typed_free_spans;
 };
 
 static void arena_free_worker(struct work_struct *work);
@@ -114,6 +118,17 @@ struct arena_free_span {
 	u32 page_cnt;
 };
 
+static void typed_arena_free_worker(struct work_struct *work);
+static void typed_arena_free_irq(struct irq_work *iw);
+
+/* A chunk range of a typed arena whose release is queued. */
+struct typed_arena_free_span {
+	struct llist_node node;
+	struct bpf_typed_arena *ta;
+	u32 coff;
+	u32 chunk_cnt;
+};
+
 static unsigned long typed_arena_region(struct bpf_arena *arena)
 {
 	return (unsigned long)arena->kern_vm->addr;
@@ -565,6 +580,347 @@ static void typed_arenas_free(struct bpf_arena *arena)
 	}
 }
 
+static bool typed_arena_chunks_free(const struct bpf_typed_arena *ta, unsigned long coff,
+				    unsigned long cnt)
+{
+	return find_next_bit(ta->chunks, coff + cnt, coff) >= coff + cnt;
+}
+
+static bool typed_arena_chunks_taken(const struct bpf_typed_arena *ta, unsigned long coff,
+				     unsigned long cnt)
+{
+	return find_next_zero_bit(ta->chunks, coff + cnt, coff) >= coff + cnt;
+}
+
+static bool typed_arena_chunks_pending(const struct bpf_typed_arena *ta, unsigned long coff,
+				       unsigned long cnt)
+{
+	return find_next_bit(ta->pending, coff + cnt, coff) < coff + cnt;
+}
+
+/*
+ * The bits are set one at a time: the fault path marks a chunk it faults to
+ * scratch with an atomic set from any context, and a word-wide update here
+ * could undo that mark.
+ */
+static void typed_arena_chunks_mark(struct bpf_typed_arena *ta, unsigned long coff,
+				    unsigned long cnt, bool taken)
+{
+	unsigned long i;
+
+	for (i = coff; i < coff + cnt; i++) {
+		if (taken)
+			set_bit(i, ta->chunks);
+		else
+			clear_bit(i, ta->chunks);
+	}
+}
+
+struct typed_install_data {
+	struct bpf_arena *arena;
+	struct page *head;
+	unsigned long start;
+	int i;
+};
+
+/*
+ * The typed arena installer maps the pages of one chunk's allocation and only
+ * fills empty entries: a scratch page is never replaced, see
+ * typed_arena_handle_page_fault(). Pairs with the atomic scratch installer.
+ */
+static int apply_range_set_typed_cb(pte_t *pte, unsigned long addr, void *data)
+{
+	struct typed_install_data *d = data;
+	struct page *page = d->head + ((addr - d->start) >> PAGE_SHIFT);
+
+	if (!ptep_try_set(pte, mk_pte(page, PAGE_KERNEL)))
+		return -EBUSY;
+	d->i++;
+	WRITE_ONCE(d->arena->nr_pages, d->arena->nr_pages + 1);
+	return 0;
+}
+
+struct typed_clear_data {
+	struct bpf_arena *arena;
+	const struct bpf_typed_arena *ta;
+	struct page *head;
+};
+
+/*
+ * Clear the entries of one chunk. A real chunk is one allocation, so the page
+ * at the chunk's first entry is its head, which is reported for freeing once
+ * the mapping is gone; scratch pages are left to the scratch chunk.
+ */
+static int apply_range_clear_typed_cb(pte_t *pte, unsigned long addr, void *data)
+{
+	struct typed_clear_data *d = data;
+	pte_t old_pte;
+	struct page *page;
+
+	old_pte = ptep_get_and_clear(&init_mm, addr, pte);
+	if (pte_none(old_pte) || !pte_present(old_pte))
+		return 0;
+	page = pte_page(old_pte);
+	if (page == typed_arena_scratch_page(d->ta, addr))
+		return 0;
+	if (!((addr - (unsigned long)d->ta->base) & (bpf_typed_arena_chunk(d->ta) - 1)))
+		d->head = page;
+	WRITE_ONCE(d->arena->nr_pages, d->arena->nr_pages - 1);
+	return 0;
+}
+
+/*
+ * Back a chunk range of a typed arena with zeroed objects and return the
+ * address of its first object, or 0. @addr names the chunk of the first
+ * object, or is 0 for any range; *@page_cnt is the request in pages, rounded
+ * up to whole chunks, and receives the count granted. The range is claimed in
+ * the chunk bitmap under the arena's spinlock before anything is installed, so
+ * a chunk the fault path faulted to scratch is skipped by a search, or fails a
+ * fixed request. The fault path marks its chunk without the lock, so it can
+ * still win the race between the claim and the install; the chunks installed
+ * so far are then backed out, the chunk it took stays taken, and a search is
+ * retried a few times while a fixed request fails. Each chunk is one
+ * allocation of its order from the lock-free allocator, which the kfunc needs
+ * because it may run under a spin lock; there is no array of pages to size.
+ */
+static unsigned long typed_arena_alloc_pages(struct bpf_typed_arena *ta, unsigned long addr,
+					     u32 *page_cnt, int node_id)
+{
+	unsigned long base = (unsigned long)ta->base, nr_chunks = typed_arena_nr_chunks(ta);
+	struct bpf_arena *arena = container_of(ta->map, struct bpf_arena, map);
+	unsigned int order = typed_arena_chunk_order(ta);
+	unsigned long chunk = bpf_typed_arena_chunk(ta);
+	struct mem_cgroup *new_memcg, *old_memcg;
+	unsigned long chunk_cnt, coff = 0, start, done, i;
+	struct typed_install_data data;
+	struct typed_clear_data cdata;
+	struct llist_node *pos, *t;
+	struct llist_head freed;
+	unsigned long flags;
+	struct page *head;
+	int ret, attempt = 0;
+
+	if (node_id != NUMA_NO_NODE &&
+	    ((unsigned int)node_id >= nr_node_ids || !node_online(node_id)))
+		return 0;
+	if (!*page_cnt)
+		return 0;
+	chunk_cnt = DIV_ROUND_UP((unsigned long)*page_cnt, typed_arena_chunk_pages(ta));
+	if (chunk_cnt > nr_chunks)
+		return 0;
+	if (addr) {
+		if (addr - base >= bpf_typed_arena_size(ta))
+			return 0;
+		coff = (addr - base) >> ta->chunk_shift;
+		if (chunk_cnt > nr_chunks - coff)
+			return 0;
+	}
+
+	bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg);
+retry:
+	if (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
+		goto out;
+
+	if (addr) {
+		if (!typed_arena_chunks_free(ta, coff, chunk_cnt))
+			goto out_unlock;
+	} else {
+		coff = bitmap_find_next_zero_area(ta->chunks, nr_chunks, 0, chunk_cnt, 0);
+		if (coff >= nr_chunks)
+			goto out_unlock;
+	}
+	typed_arena_chunks_mark(ta, coff, chunk_cnt, true);
+	start = base + (coff << ta->chunk_shift);
+
+	data.arena = arena;
+	for (done = 0; done < chunk_cnt; done++) {
+		head = alloc_pages_nolock(__GFP_ACCOUNT, node_id, order);
+		if (!head) {
+			ret = -ENOMEM;
+			goto back_out;
+		}
+		data.head = head;
+		data.start = start + (done << ta->chunk_shift);
+		data.i = 0;
+		ret = apply_to_page_range(&init_mm, data.start, chunk, apply_range_set_typed_cb,
+					  &data);
+		if (ret) {
+			/* The fault path took this chunk: give back the allocation whole. */
+			cdata.arena = arena;
+			cdata.ta = ta;
+			if (data.i)
+				apply_to_existing_page_range(&init_mm, data.start,
+							     (unsigned long)data.i << PAGE_SHIFT,
+							     apply_range_clear_typed_cb, &cdata);
+			free_pages_nolock(head, order);
+			goto back_out;
+		}
+	}
+	flush_vmap_cache(start, chunk_cnt << ta->chunk_shift);
+	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
+	bpf_map_memcg_exit(old_memcg, new_memcg);
+	*page_cnt = chunk_cnt * typed_arena_chunk_pages(ta);
+	return start;
+
+back_out:
+	/*
+	 * Back out the chunks installed before the one that failed. On -EBUSY
+	 * that chunk is the one the fault path took, and it keeps its mark.
+	 */
+	init_llist_head(&freed);
+	cdata.arena = arena;
+	cdata.ta = ta;
+	for (i = 0; i < chunk_cnt; i++)
+		if (ret != -EBUSY || i != done)
+			clear_bit(coff + i, ta->chunks);
+	for (i = 0; i < done; i++) {
+		cdata.head = NULL;
+		apply_to_existing_page_range(&init_mm, start + (i << ta->chunk_shift), chunk,
+					     apply_range_clear_typed_cb, &cdata);
+		if (cdata.head)
+			__llist_add(&cdata.head->pcp_llist, &freed);
+	}
+	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
+	flush_tlb_kernel_range(start, start + (chunk_cnt << ta->chunk_shift));
+	llist_for_each_safe(pos, t, __llist_del_all(&freed))
+		free_pages_nolock(llist_entry(pos, struct page, pcp_llist), order);
+	if (ret == -EBUSY && !addr && ++attempt < 3)
+		goto retry;
+	goto out;
+
+out_unlock:
+	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
+out:
+	bpf_map_memcg_exit(old_memcg, new_memcg);
+	return 0;
+}
+
+/*
+ * Queue the release of the chunks covering a page range of a typed arena. The
+ * chunks stay mapped and marked until the worker has waited for the grace
+ * periods, so that every invocation that started before this call keeps its
+ * objects. Every chunk of the range must be taken, by real pages or by the
+ * fault path, and none of them may have a release queued already: a second
+ * release of a chunk would run after the first has let it be claimed again,
+ * and take it away from its new owner. Releasing a chunk the fault path took
+ * clears its entries, after which it can be claimed or faulted again.
+ */
+static void typed_arena_free_pages(struct bpf_typed_arena *ta, unsigned long addr, u32 page_cnt)
+{
+	unsigned long base = (unsigned long)ta->base, size = bpf_typed_arena_size(ta);
+	struct bpf_arena *arena = container_of(ta->map, struct bpf_arena, map);
+	unsigned long off, first, last, flags;
+	struct typed_arena_free_span *s;
+
+	if (!page_cnt || addr - base >= size)
+		return;
+	off = addr - base;
+	if (page_cnt > (size - off) >> PAGE_SHIFT)
+		return;
+	first = off >> ta->chunk_shift;
+	last = (off + ((unsigned long)page_cnt << PAGE_SHIFT) - 1) >> ta->chunk_shift;
+
+	s = kmalloc_nolock(sizeof(*s), __GFP_ACCOUNT, NUMA_NO_NODE);
+	if (!s)
+		/*
+		 * The chunks stay allocated until the map is freed; nothing can
+		 * be retried from here.
+		 */
+		return;
+	if (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
+		goto free_span;
+	if (!typed_arena_chunks_taken(ta, first, last - first + 1) ||
+	    typed_arena_chunks_pending(ta, first, last - first + 1)) {
+		raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
+		goto free_span;
+	}
+	/* Only this path and the worker touch the pending bits, both under the lock. */
+	bitmap_set(ta->pending, first, last - first + 1);
+	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
+
+	s->ta = ta;
+	s->coff = first;
+	s->chunk_cnt = last - first + 1;
+	llist_add(&s->node, &arena->typed_free_spans);
+	irq_work_queue(&arena->typed_free_irq);
+	return;
+
+free_span:
+	kfree_nolock(s);
+}
+
+/*
+ * Release the queued chunk ranges once every invocation that started before
+ * a release was requested is done with its objects: after an RCU and an RCU
+ * tasks trace grace period, since sleepable programs run under the latter.
+ * The entries are cleared and the marks dropped under the spinlock, chunk by
+ * chunk, then the TLB is flushed and each real chunk's allocation, one
+ * contiguous block, has the special fields of its objects dropped and is
+ * freed. An invocation that casts into the range after the release was
+ * requested races with it: it stays memory safe, and sees the real objects
+ * before the clear and the dummy object after.
+ */
+static void typed_arena_free_worker(struct work_struct *work)
+{
+	struct bpf_arena *arena = container_of(work, struct bpf_arena, typed_free_work);
+	struct mem_cgroup *new_memcg, *old_memcg;
+	struct llist_node *list, *pos, *t, *p, *pn;
+	struct typed_arena_free_span *s;
+	struct typed_clear_data cdata;
+	struct bpf_typed_arena *ta;
+	struct llist_head heads;
+	unsigned long flags, start, i;
+	struct page *head;
+
+	list = llist_del_all(&arena->typed_free_spans);
+	if (!list)
+		return;
+
+	synchronize_rcu_mult(call_rcu, call_rcu_tasks_trace);
+
+	bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg);
+	llist_for_each_safe(pos, t, list) {
+		s = llist_entry(pos, struct typed_arena_free_span, node);
+		ta = s->ta;
+		start = (unsigned long)ta->base + ((unsigned long)s->coff << ta->chunk_shift);
+
+		init_llist_head(&heads);
+		cdata.arena = arena;
+		cdata.ta = ta;
+
+		while (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
+			cpu_relax();
+		for (i = 0; i < s->chunk_cnt; i++) {
+			cdata.head = NULL;
+			apply_to_existing_page_range(&init_mm, start + (i << ta->chunk_shift),
+						     bpf_typed_arena_chunk(ta),
+						     apply_range_clear_typed_cb, &cdata);
+			if (cdata.head)
+				__llist_add(&cdata.head->pcp_llist, &heads);
+		}
+		typed_arena_chunks_mark(ta, s->coff, s->chunk_cnt, false);
+		bitmap_clear(ta->pending, s->coff, s->chunk_cnt);
+		raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
+
+		flush_tlb_kernel_range(start,
+				       start + ((unsigned long)s->chunk_cnt << ta->chunk_shift));
+		llist_for_each_safe(p, pn, __llist_del_all(&heads)) {
+			head = llist_entry(p, struct page, pcp_llist);
+			typed_arena_free_objects(ta, page_address(head));
+			__free_pages(head, typed_arena_chunk_order(ta));
+		}
+		kfree_nolock(s);
+	}
+	bpf_map_memcg_exit(old_memcg, new_memcg);
+}
+
+static void typed_arena_free_irq(struct irq_work *iw)
+{
+	struct bpf_arena *arena = container_of(iw, struct bpf_arena, typed_free_irq);
+
+	schedule_work(&arena->typed_free_work);
+}
+
 static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
 {
 	struct vm_struct *kern_vm;
@@ -613,6 +969,9 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
 	init_llist_head(&arena->free_spans);
 	init_irq_work(&arena->free_irq, arena_free_irq);
 	INIT_WORK(&arena->free_work, arena_free_worker);
+	init_llist_head(&arena->typed_free_spans);
+	init_irq_work(&arena->typed_free_irq, typed_arena_free_irq);
+	INIT_WORK(&arena->typed_free_work, typed_arena_free_worker);
 	bpf_map_init_from_attr(&arena->map, attr);
 
 	err = bpf_map_alloc_pages(&arena->map, NUMA_NO_NODE, 1, &arena->scratch_page);
@@ -686,6 +1045,8 @@ static void arena_map_free(struct bpf_map *map)
 	/* Ensure no pending deferred frees */
 	irq_work_sync(&arena->free_irq);
 	flush_work(&arena->free_work);
+	irq_work_sync(&arena->typed_free_irq);
+	flush_work(&arena->typed_free_work);
 
 	/*
 	 * free_vm_area() calls remove_vm_area() that calls free_unmap_vmap_area().
@@ -1463,12 +1824,42 @@ __bpf_kfunc int bpf_arena_reserve_pages(void *p__map, void *ptr__ign, u32 page_c
 
 	return arena_reserve_pages(arena, (long)ptr__ign, page_cnt);
 }
+
+/*
+ * The verifier registers the typed arena of local_type_id__k at load, checks
+ * that the map is the program's arena, and replaces the type ID register with
+ * the registered typed arena before the call. addr__ign is a typed pointer
+ * into that arena naming the chunk of the first object, or NULL for any
+ * range; *page_cnt is the request in pages and receives the count granted,
+ * rounded up to whole objects.
+ */
+__bpf_kfunc void *bpf_typed_arena_alloc_pages(void *p__map, u64 local_type_id__k, void *addr__ign,
+					      u32 *page_cnt, int node_id)
+{
+	struct bpf_typed_arena *ta = (struct bpf_typed_arena *)local_type_id__k;
+
+	if (ta->map != p__map)
+		return NULL;
+	return (void *)typed_arena_alloc_pages(ta, (unsigned long)addr__ign, page_cnt, node_id);
+}
+
+__bpf_kfunc void bpf_typed_arena_free_pages(void *p__map, u64 local_type_id__k, void *ptr__ign,
+					    u32 page_cnt)
+{
+	struct bpf_typed_arena *ta = (struct bpf_typed_arena *)local_type_id__k;
+
+	if (ta->map != p__map)
+		return;
+	typed_arena_free_pages(ta, (unsigned long)ptr__ign, page_cnt);
+}
 __bpf_kfunc_end_defs();
 
 BTF_KFUNCS_START(arena_kfuncs)
 BTF_ID_FLAGS(func, bpf_arena_alloc_pages, KF_ARENA_RET | KF_ARENA_ARG2 | KF_SPINLOCK_SAFE)
 BTF_ID_FLAGS(func, bpf_arena_free_pages, KF_ARENA_ARG2 | KF_SPINLOCK_SAFE)
 BTF_ID_FLAGS(func, bpf_arena_reserve_pages, KF_ARENA_ARG2 | KF_SPINLOCK_SAFE)
+BTF_ID_FLAGS(func, bpf_typed_arena_alloc_pages, KF_RET_NULL | KF_SPINLOCK_SAFE)
+BTF_ID_FLAGS(func, bpf_typed_arena_free_pages, KF_SPINLOCK_SAFE)
 BTF_KFUNCS_END(arena_kfuncs)
 
 static const struct btf_kfunc_id_set common_kfunc_set = {
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 2d406034556e..b85d99f81ef5 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -216,6 +216,7 @@ static bool is_tracing_prog_type(enum bpf_prog_type type);
 static int ref_set_non_owning(struct bpf_verifier_env *env,
 			      struct bpf_reg_state *reg);
 static bool is_trusted_reg(struct bpf_verifier_env *env, const struct bpf_reg_state *reg);
+static struct bpf_typed_arena *typed_arena_register(struct bpf_verifier_env *env, u32 btf_id);
 static inline bool in_sleepable_context(struct bpf_verifier_env *env);
 static const char *non_sleepable_context_description(struct bpf_verifier_env *env);
 static void scalar32_min_max_add(struct bpf_reg_state *dst_reg, struct bpf_reg_state *src_reg);
@@ -13334,6 +13335,8 @@ enum special_kfunc_type {
 	KF_bpf_arena_alloc_pages,
 	KF_bpf_arena_free_pages,
 	KF_bpf_arena_reserve_pages,
+	KF_bpf_typed_arena_alloc_pages,
+	KF_bpf_typed_arena_free_pages,
 	KF_bpf_session_is_return,
 	KF_bpf_stream_vprintk,
 	KF_bpf_stream_print_stack,
@@ -13429,6 +13432,8 @@ BTF_ID(func, bpf_call_rcu_tasks_trace)
 BTF_ID(func, bpf_arena_alloc_pages)
 BTF_ID(func, bpf_arena_free_pages)
 BTF_ID(func, bpf_arena_reserve_pages)
+BTF_ID(func, bpf_typed_arena_alloc_pages)
+BTF_ID(func, bpf_typed_arena_free_pages)
 #ifdef CONFIG_BPF_EVENTS
 BTF_ID(func, bpf_session_is_return)
 #else
@@ -13458,6 +13463,12 @@ static bool is_bpf_obj_new_kfunc(u32 func_id)
 	       func_id == special_kfunc_list[KF_bpf_obj_new_impl];
 }
 
+static bool is_typed_arena_pages_kfunc(u32 func_id)
+{
+	return func_id == special_kfunc_list[KF_bpf_typed_arena_alloc_pages] ||
+	       func_id == special_kfunc_list[KF_bpf_typed_arena_free_pages];
+}
+
 static bool is_bpf_percpu_obj_new_kfunc(u32 func_id)
 {
 	return func_id == special_kfunc_list[KF_bpf_percpu_obj_new] ||
@@ -14800,6 +14811,11 @@ static int check_special_kfunc(struct bpf_verifier_env *env, struct bpf_call_arg
 
 		insn_aux->obj_new_size = ret_t->size;
 		insn_aux->kptr_struct_meta = struct_meta;
+	} else if (is_kfunc_call(meta, special_kfunc_list[KF_bpf_typed_arena_alloc_pages])) {
+		mark_reg_known_zero(env, regs, BPF_REG_0);
+		regs[BPF_REG_0].type = PTR_TO_BTF_ID | MEM_ARENA;
+		regs[BPF_REG_0].btf = env->prog->aux->btf;
+		regs[BPF_REG_0].btf_id = insn_aux->typed_arena->btf_id;
 	} else if (is_bpf_refcount_acquire_kfunc(meta->func_id)) {
 		mark_reg_known_zero(env, regs, BPF_REG_0);
 		regs[BPF_REG_0].type = PTR_TO_BTF_ID | MEM_ALLOC;
@@ -15130,6 +15146,38 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
 		ref_convert_owning_non_owning(env, id);
 	}
 
+	if (meta.btf == btf_vmlinux && is_typed_arena_pages_kfunc(meta.func_id)) {
+		struct bpf_typed_arena *ta;
+
+		if (!env->prog->aux->arena) {
+			verbose(env, "kfunc %s can only be used in a program that has an associated arena\n",
+				func_name);
+			return -EINVAL;
+		}
+		if (!env->prog->aux->btf) {
+			verbose(env, "kfunc %s requires program BTF\n", func_name);
+			return -EINVAL;
+		}
+		if (((u64)(u32)meta.arg_constant.value) != meta.arg_constant.value) {
+			verbose(env, "local type ID argument must be in range [0, U32_MAX]\n");
+			return -EINVAL;
+		}
+		ta = typed_arena_register(env, meta.arg_constant.value);
+		if (IS_ERR(ta))
+			return PTR_ERR(ta);
+		/*
+		 * The type ID is a register value, so it may differ between
+		 * paths; the fixup replaces it with the typed arena once per
+		 * call site, and the allocator's return value is typed by it.
+		 */
+		if (insn_aux->typed_arena && insn_aux->typed_arena != ta) {
+			verbose(env, "kfunc %s at insn %d is called with different types on different paths\n",
+				func_name, insn_idx);
+			return -EINVAL;
+		}
+		insn_aux->typed_arena = ta;
+	}
+
 	if (meta.func_id == special_kfunc_list[KF_bpf_throw]) {
 		if (!bpf_jit_supports_exceptions()) {
 			verbose(env, "JIT does not support calling kfunc %s#%d\n",
@@ -22493,6 +22541,18 @@ int bpf_fixup_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
 		insn_buf[2] = addr[1];
 		insn_buf[3] = *insn;
 		*cnt = 4;
+	} else if (is_typed_arena_pages_kfunc(desc->func_id)) {
+		struct bpf_typed_arena *ta = env->insn_aux_data[insn_idx].typed_arena;
+		struct bpf_insn addr[2] = { BPF_LD_IMM64(BPF_REG_2, (long)ta) };
+
+		if (!ta) {
+			verifier_bug(env, "typed arena kfunc at insn %d has no typed arena", insn_idx);
+			return -EFAULT;
+		}
+		insn_buf[0] = addr[0];
+		insn_buf[1] = addr[1];
+		insn_buf[2] = *insn;
+		*cnt = 3;
 	} else if (is_bpf_obj_drop_kfunc(desc->func_id) ||
 		   is_bpf_percpu_obj_drop_kfunc(desc->func_id) ||
 		   is_bpf_refcount_acquire_kfunc(desc->func_id)) {
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 10/16] bpf: Let typed_arena_cast copy pointers the verifier already trusts
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (8 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 09/16] bpf: Add typed arena page allocation and release kfuncs Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:49   ` sashiko-bot
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 11/16] libbpf: Support the typed_arena_cast instruction Kumar Kartikeya Dwivedi
                   ` (5 subsequent siblings)
  15 siblings, 1 reply; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

The compiler inserts the cast at every use of a pointer to a struct with
special fields, so that programs never write one, and it cannot tell what
the pointer is: the same struct lives in typed arenas, in allocated objects,
in map values and on the stack. So the verifier decides per path what the
cast takes. A scalar, a raw arena pointer, a typed arena pointer loaded from
a field and not yet sanitized, one of another struct, or one that
arithmetic moved inside its object is sanitized as before. A sanitized
pointer of the struct's own typed arena at offset zero is already what the
cast produces, so it is copied, and since the sequence is a no-op on it,
such a path may share the instruction with sanitizing ones. Any other
verified pointer, to an allocated object, a map value, the stack or
anything else, is copied with its state, offset and references: the cast
has nothing to add to what the verifier knows, so it needs neither an arena
map nor a registration, and a program that only allocates objects of the
struct loads as before. The lowering is per instruction and runs on every
path through it, so a path that copies such a pointer cannot share the
instruction with one that sanitizes, since the sequence would mask a
pointer that is not in the slice; that program is rejected. An instruction
that only copied lowers to a move, or to nothing when dst is src, which
takes removing the instruction, as the nop removal pass has already run.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 kernel/bpf/fixups.c   | 35 ++++++++++++++++++++++--------
 kernel/bpf/verifier.c | 50 +++++++++++++++++++++++++++++++++++++++++++
 2 files changed, 76 insertions(+), 9 deletions(-)

diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index fd0ba8376c1c..a0856bd18ac2 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -897,17 +897,34 @@ int bpf_lower_typed_arena_insns(struct bpf_verifier_env *env)
 	for (i = 0; i < insn_cnt; i++, insn++) {
 		aux = &env->insn_aux_data[i + delta];
 		cnt = 0;
-		if (aux->sanitize_needed)
-			cnt = bpf_typed_arena_cast_insns(aux->sanitize_arena, aux->sanitize_reg,
-							 aux->sanitize_reg, insn_buf);
 		if (insn_is_typed_arena_cast(insn)) {
-			if (!aux->typed_arena) {
-				verifier_bug(env, "typed_arena_cast at insn %d has no typed arena", i);
-				return -EFAULT;
+			/*
+			 * A cast that sanitized on some path lowers to the
+			 * sequence, a no-op on the paths that brought a sanitized
+			 * pointer of the same typed arena. One that only copied
+			 * verified pointers is a move, or nothing at all.
+			 */
+			if (aux->sanitize_needed) {
+				if (!aux->typed_arena) {
+					verifier_bug(env, "typed_arena_cast at insn %d has no typed arena", i);
+					return -EFAULT;
+				}
+				cnt = bpf_typed_arena_cast_insns(aux->typed_arena, insn->dst_reg,
+								 insn->src_reg, insn_buf);
+			} else if (insn->dst_reg != insn->src_reg) {
+				insn_buf[cnt++] = BPF_MOV64_REG(insn->dst_reg, insn->src_reg);
+			} else {
+				int err = verifier_remove_insns(env, i + delta, 1);
+
+				if (err)
+					return err;
+				delta--;
+				insn = env->prog->insnsi + i + delta;
+				continue;
 			}
-			cnt += bpf_typed_arena_cast_insns(aux->typed_arena, insn->dst_reg,
-							  insn->src_reg, insn_buf + cnt);
-		} else if (cnt) {
+		} else if (aux->sanitize_needed) {
+			cnt = bpf_typed_arena_cast_insns(aux->sanitize_arena, aux->sanitize_reg,
+							 aux->sanitize_reg, insn_buf);
 			insn_buf[cnt++] = *insn;
 		}
 		if (!cnt)
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index b85d99f81ef5..aceb17e62948 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -17299,13 +17299,51 @@ static struct bpf_typed_arena *typed_arena_register(struct bpf_verifier_env *env
  * of the instruction, so an instruction casts to one type on every path, and
  * the lowering can be chosen once.
  */
+/*
+ * dst = typed_arena_cast(src, imm) makes a trusted pointer to an object of the
+ * struct imm names out of the value in src. What that takes depends on the
+ * value, and the compiler that inserts the cast at every use of a pointer to
+ * the struct cannot tell, since the same struct lives in typed arenas, in
+ * allocated objects, in map values and on the stack. So the verifier decides
+ * per path. A scalar, a raw arena pointer, or a typed arena pointer that is
+ * unsanitized, of another struct, or moved inside its object is sanitized:
+ * dst becomes a pointer to the start of an object of the struct's typed
+ * arena, and the lowering masks and rebases the value. A sanitized pointer of
+ * that typed arena at offset zero is already what the cast makes, so dst is a
+ * copy of src; the sequence is a no-op on it, which lets such a path share
+ * the instruction with sanitizing ones. Any other verified pointer, to an
+ * allocated object, a map value, the stack, or anything else, is copied as it
+ * is, with its state, offset and references: the cast has nothing to add to
+ * what the verifier already knows, and needs no arena and no registration. A
+ * typed arena pointer that may be NULL, an allocation's result, is one of
+ * these: sanitizing it would turn a failed allocation into object 0, so it is
+ * copied and keeps its check.
+ *
+ * The lowering is per instruction and runs on every path through it, so a
+ * path that copies a verified pointer of another kind cannot share the
+ * instruction with a path that sanitizes: the sequence would mask a pointer
+ * that is not in the slice. Such a program is rejected.
+ */
 static int check_typed_arena_cast(struct bpf_verifier_env *env, struct bpf_insn *insn)
 {
 	struct bpf_insn_aux_data *aux = &env->insn_aux_data[env->insn_idx];
 	struct bpf_reg_state *regs = cur_regs(env);
+	struct bpf_reg_state *src = &regs[insn->src_reg];
 	struct bpf_reg_state *dst = &regs[insn->dst_reg];
 	struct bpf_typed_arena *ta;
 
+	if (src->type != SCALAR_VALUE && base_type(src->type) != PTR_TO_ARENA &&
+	    (!type_is_typed_arena_obj(src->type) || type_flag(src->type) & PTR_MAYBE_NULL)) {
+		if (aux->sanitize_needed) {
+			verbose(env, "insn %d casts values that need different treatment on different paths\n",
+				env->insn_idx);
+			return -EINVAL;
+		}
+		aux->sanitize_plain |= BIT(insn->src_reg);
+		*dst = *src;
+		return 0;
+	}
+
 	/* The arena itself needs CAP_PERFMON, so the leak rules already permit a kernel pointer. */
 	if (!env->prog->aux->arena) {
 		verbose(env, "typed_arena_cast insn can only be used in a program that has an associated arena\n");
@@ -17320,6 +17358,18 @@ static int check_typed_arena_cast(struct bpf_verifier_env *env, struct bpf_insn
 		return PTR_ERR(ta);
 	aux->typed_arena = ta;
 
+	if (src->type == (PTR_TO_BTF_ID | MEM_ARENA) && src->btf_id == ta->btf_id &&
+	    tnum_is_const(src->var_off) && !src->var_off.value) {
+		*dst = *src;
+		return 0;
+	}
+	if (aux->sanitize_plain) {
+		verbose(env, "insn %d casts values that need different treatment on different paths\n",
+			env->insn_idx);
+		return -EINVAL;
+	}
+	aux->sanitize_needed = true;
+
 	mark_reg_known_zero(env, regs, insn->dst_reg);
 	dst->type = PTR_TO_BTF_ID | MEM_ARENA;
 	dst->btf = env->prog->aux->btf;
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 11/16] libbpf: Support the typed_arena_cast instruction
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (9 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 10/16] bpf: Let typed_arena_cast copy pointers the verifier already trusts Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 12/16] selftests/bpf: Build BPF objects with compiler-inserted typed arena casts Kumar Kartikeya Dwivedi
                   ` (4 subsequent siblings)
  15 siblings, 0 replies; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

The typed_arena_cast instruction turns any 64-bit value into a pointer to an
object of a program-BTF struct in the struct's typed arena, the kernel-only
region of an arena map. It is encoded as a 64-bit register move with an off
of BPF_TYPED_ARENA_CAST and the struct's local BTF type ID in imm, and the
kernel takes the ID from there, so a compiler that emits the instruction
records a BPF_CORE_TYPE_ID_LOCAL relocation against it, the way it does for
the ld_imm64 behind __builtin_btf_type_id().

Let the relocation patcher write a local type ID into that instruction. It
is the one ALU form with a register source that carries a relocation, and it
takes no other relocation kind; nothing else about the patching of ALU
immediates changes.

Add bpf_typed_arena_cast() to bpf_helpers.h, wrapping the compiler builtin
under its feature macro, so that programs cast through one name whatever the
toolchain.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 tools/lib/bpf/bpf_helpers.h | 13 +++++++++++++
 tools/lib/bpf/relo_core.c   | 19 +++++++++++++++++--
 2 files changed, 30 insertions(+), 2 deletions(-)

diff --git a/tools/lib/bpf/bpf_helpers.h b/tools/lib/bpf/bpf_helpers.h
index 9d160b5b9c0e..1a440e32ebeb 100644
--- a/tools/lib/bpf/bpf_helpers.h
+++ b/tools/lib/bpf/bpf_helpers.h
@@ -340,6 +340,19 @@ enum libbpf_tristate {
 /* Helper macro to print out debug messages */
 #define bpf_printk(fmt, args...) ___bpf_pick_printk(args)(fmt, ##args)
 
+/*
+ * Turn any 64-bit value into a pointer to an object of the given struct type
+ * in its typed arena, with the typed_arena_cast instruction: the kernel masks
+ * the value into the struct's slice of the arena map's typed region, so the
+ * result names an object of that type whatever the value was, and the
+ * verifier trusts it. The struct's local BTF type ID is relocated into the
+ * instruction. Requires compiler support.
+ */
+#ifdef __BPF_FEATURE_TYPED_ARENA_CAST
+#define bpf_typed_arena_cast(v, type)					\
+	((typeof(type) *)__builtin_bpf_typed_arena_cast((v), *(typeof(type) *)0))
+#endif
+
 struct bpf_iter_num;
 
 extern int bpf_iter_num_new(struct bpf_iter_num *it, int start, int end) __weak __ksym;
diff --git a/tools/lib/bpf/relo_core.c b/tools/lib/bpf/relo_core.c
index 2672623a4198..e872e432729b 100644
--- a/tools/lib/bpf/relo_core.c
+++ b/tools/lib/bpf/relo_core.c
@@ -1028,6 +1028,12 @@ static int insn_bytes_to_bpf_size(__u32 sz)
 	}
 }
 
+/* rX = typed_arena_cast(rY, <local type ID in imm>) */
+static bool is_typed_arena_cast_insn(struct bpf_insn *insn)
+{
+	return insn->code == (BPF_ALU64 | BPF_MOV | BPF_X) && insn->off == BPF_TYPED_ARENA_CAST;
+}
+
 /*
  * Patch relocatable BPF instruction.
  *
@@ -1043,7 +1049,8 @@ static int insn_bytes_to_bpf_size(__u32 sz)
  * 3. rX = <imm64> (load with 64-bit immediate value);
  * 4. rX = *(T *)(rY + <off>), where T is one of {u8, u16, u32, u64};
  * 5. *(T *)(rX + <off>) = rY, where T is one of {u8, u16, u32, u64};
- * 6. *(T *)(rX + <off>) = <imm>, where T is one of {u8, u16, u32, u64}.
+ * 6. *(T *)(rX + <off>) = <imm>, where T is one of {u8, u16, u32, u64};
+ * 7. rX = typed_arena_cast(rY, <imm>), for a local type ID relocation only.
  */
 int bpf_core_patch_insn(const char *prog_name, struct bpf_insn *insn,
 			int insn_idx, const struct bpf_core_relo *relo,
@@ -1060,8 +1067,16 @@ int bpf_core_patch_insn(const char *prog_name, struct bpf_insn *insn,
 	switch (class) {
 	case BPF_ALU:
 	case BPF_ALU64:
-		if (BPF_SRC(insn->code) != BPF_K)
+		/*
+		 * The typed arena cast is a register move whose imm names the
+		 * struct by its local type ID; the kernel takes the ID from there.
+		 */
+		if (is_typed_arena_cast_insn(insn)) {
+			if (relo->kind != BPF_CORE_TYPE_ID_LOCAL)
+				goto bad_insn;
+		} else if (BPF_SRC(insn->code) != BPF_K) {
 			goto bad_insn;
+		}
 		if (res->poison)
 			return bpf_core_poison_insn(prog_name, relo_idx, insn, insn_idx);
 		if (res->validate && insn->imm != orig_val) {
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 12/16] selftests/bpf: Build BPF objects with compiler-inserted typed arena casts
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (10 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 11/16] libbpf: Support the typed_arena_cast instruction Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:46   ` sashiko-bot
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 13/16] selftests/bpf: Test typed arena casts and registration Kumar Kartikeya Dwivedi
                   ` (3 subsequent siblings)
  15 siblings, 1 reply; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

Clang inserts typed_arena_cast where a pointer to a typed record is used,
but only with the +typed-arena target feature, and it defines
__BPF_FEATURE_TYPED_ARENA_CAST only then: a struct with kptr fields lives in
map values of programs that run on kernels without the instruction, so the
insertion has to be opted into. Probe the feature the way cpu v4 is probed,
by asking the compiler for the macro, and pass it to every flavor when the
compiler has it. The typed arena tests are gated on the macro and skip with
a compiler that lacks the feature.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 tools/testing/selftests/bpf/Makefile | 13 ++++++++++---
 1 file changed, 10 insertions(+), 3 deletions(-)

diff --git a/tools/testing/selftests/bpf/Makefile b/tools/testing/selftests/bpf/Makefile
index afa589a27b15..bd1f84f55edc 100644
--- a/tools/testing/selftests/bpf/Makefile
+++ b/tools/testing/selftests/bpf/Makefile
@@ -15,6 +15,13 @@ ifneq ($(shell $(CLANG) --target=bpf -mcpu=help 2>&1 | grep 'v4'),)
 CLANG_CPUV4 := 1
 endif
 
+# Check whether clang inserts typed arena casts; the feature defines the macro
+# the typed arena tests are gated on, so they skip with a clang without it
+ifneq ($(shell echo | $(CLANG) --target=bpf -Xclang -target-feature -Xclang +typed-arena \
+		-dM -E - 2>/dev/null | grep __BPF_FEATURE_TYPED_ARENA_CAST),)
+CLANG_TYPED_ARENA := -Xclang -target-feature -Xclang +typed-arena
+endif
+
 # Order correspond to 'make run_tests' order
 TEST_GEN_PROGS = test_verifier test_tag test_maps test_lru_map test_progs \
 	test_sockmap \
@@ -510,7 +517,7 @@ BINARY := test_progs
 BPF_CC := $(CLANG)
 BPF_CC_MSG := CLNG-BPF
 BPF_SYS_INCLUDES := $(CLANG_SYS_INCLUDES)
-BPF_CC_FLAGS := -O2 $(BPF_TARGET_ENDIAN) -mcpu=v3
+BPF_CC_FLAGS := -O2 $(BPF_TARGET_ENDIAN) -mcpu=v3 $(CLANG_TYPED_ARENA)
 BPF_DEFINES := -DENABLE_ATOMICS_TESTS
 include Makefile.skel
 
@@ -526,12 +533,12 @@ $(OUTPUT)/test_progs: $(RUNNER_PREREQS) $(BPF_OBJS) $(ALL_SKELS) FORCE
 
 $(OUTPUT)/test_progs-no_alu32: $(RUNNER_PREREQS) FORCE
 	+$(Q)$(RUNNER_MAKE) RUNNER=test_progs FLAVOR=no_alu32 TESTS_DIR=prog_tests \
-		$(CLANG_RUNNER_ARGS) BPF_CC_FLAGS='-O2 $(BPF_TARGET_ENDIAN) -mcpu=v2'
+		$(CLANG_RUNNER_ARGS) BPF_CC_FLAGS='-O2 $(BPF_TARGET_ENDIAN) -mcpu=v2 $(CLANG_TYPED_ARENA)'
 
 ifneq ($(CLANG_CPUV4),)
 $(OUTPUT)/test_progs-cpuv4: $(RUNNER_PREREQS) FORCE
 	+$(Q)$(RUNNER_MAKE) RUNNER=test_progs FLAVOR=cpuv4 TESTS_DIR=prog_tests \
-		$(CLANG_RUNNER_ARGS) BPF_CC_FLAGS='-O2 $(BPF_TARGET_ENDIAN) -mcpu=v4' \
+		$(CLANG_RUNNER_ARGS) BPF_CC_FLAGS='-O2 $(BPF_TARGET_ENDIAN) -mcpu=v4 $(CLANG_TYPED_ARENA)' \
 		BPF_DEFINES=-DENABLE_ATOMICS_TESTS
 endif
 
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 13/16] selftests/bpf: Test typed arena casts and registration
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (11 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 12/16] selftests/bpf: Build BPF objects with compiler-inserted typed arena casts Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 14/16] selftests/bpf: Test typed arena object access, kptrs and typed pointer fields Kumar Kartikeya Dwivedi
                   ` (2 subsequent siblings)
  15 siblings, 0 replies; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

The compiler inserts a typed_arena_cast where a value is converted to a
pointer to a typed struct and before every use of such a pointer as an
address, so the programs write no cast: an object is whatever value user
space hands in, a raw arena pointer, or a pointer of another type, and the
casts the verifier sees are all the compiler's.

Convert a value user space hands in and see it lowered to the mask, the
base load and the add, with the pointer typed as the struct's object;
convert a raw arena pointer, and a typed pointer passed through an opaque
value and to another type, which land on objects like any other value. See
the registration logged with the slot size. A cast of a pointer the
verifier already trusts is the identity: an allocated object, a typed
struct inside a map value and one on the stack keep their type and state
through the cast the compiler inserts, and such programs need no arena.

Reject a cast in a program without an arena map, and structs with the
special fields a typed arena does not carry: spin locks, resilient spin
locks, timers, refcounts, list heads and rbtree roots. A struct without
special fields is not typed to the compiler and no cast names it; the page
kfuncs, which take a type ID, reject one in their own tests. Reject typed
arena sizes that are not a power of two, below a page, above the region, or
too small for one object, and two size declarations on one struct; accept
sizes with a suffix and see the object count follow, take an object larger
than a page in a slot of whole pages, and see the region refuse a third
type once two 2 GiB ones have filled it.

A typed pointer is a kernel pointer: a 32-bit copy is the scalar a privileged
program may take of any pointer, and a narrow store to the stack is an
invalid spill.

The values user space would hand in stay unset: the verifier does not care
where a value lands, and at run time an unset one names object 0 of its
slice. The tests exist only when the compiler emits the cast; the runner
checks a flag the object exports and skips otherwise. Add the typed arena
size tag macro to bpf_experimental.h.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 .../testing/selftests/bpf/bpf_experimental.h  |   6 +
 .../selftests/bpf/prog_tests/verifier.c       |  22 +
 .../bpf/progs/verifier_typed_arena.c          | 528 ++++++++++++++++++
 3 files changed, 556 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/progs/verifier_typed_arena.c

diff --git a/tools/testing/selftests/bpf/bpf_experimental.h b/tools/testing/selftests/bpf/bpf_experimental.h
index 2893bf06ff25..128654997328 100644
--- a/tools/testing/selftests/bpf/bpf_experimental.h
+++ b/tools/testing/selftests/bpf/bpf_experimental.h
@@ -9,6 +9,12 @@
 
 #define __contains(name, node) __attribute__((btf_decl_tag("contains:" #name ":" #node)))
 
+/*
+ * Size of a struct's typed arena in bytes: a power of two with an optional K,
+ * M or G suffix, 128M unless declared.
+ */
+#define __typed_arena_size(sz) __attribute__((btf_decl_tag("typed_arena_size:" #sz)))
+
 /* Convenience macro to wrap over bpf_obj_new */
 #define bpf_obj_new(type) ((type *)bpf_obj_new(bpf_core_type_id_local(type)))
 
diff --git a/tools/testing/selftests/bpf/prog_tests/verifier.c b/tools/testing/selftests/bpf/prog_tests/verifier.c
index 8a6d341b754a..3a4270fca559 100644
--- a/tools/testing/selftests/bpf/prog_tests/verifier.c
+++ b/tools/testing/selftests/bpf/prog_tests/verifier.c
@@ -120,6 +120,7 @@
 #include "verifier_subreg.skel.h"
 #include "verifier_tailcall.skel.h"
 #include "verifier_tailcall_jit.skel.h"
+#include "verifier_typed_arena.skel.h"
 #include "verifier_typedef.skel.h"
 #include "verifier_uninit.skel.h"
 #include "verifier_unpriv.skel.h"
@@ -303,6 +304,27 @@ void test_verifier_subprog_topo(void)        { RUN(verifier_subprog_topo); }
 void test_verifier_subreg(void)               { RUN(verifier_subreg); }
 void test_verifier_tailcall(void)             { RUN(verifier_tailcall); }
 void test_verifier_tailcall_jit(void)         { RUN(verifier_tailcall_jit); }
+
+/*
+ * The typed arena tests need a compiler that emits the cast; without one the
+ * object holds no programs.
+ */
+void test_verifier_typed_arena(void)
+{
+	struct verifier_typed_arena *skel;
+	bool supported;
+
+	skel = verifier_typed_arena__open();
+	if (!ASSERT_OK_PTR(skel, "open"))
+		return;
+	supported = skel->rodata->typed_arena_supported;
+	verifier_typed_arena__destroy(skel);
+	if (!supported) {
+		test__skip();
+		return;
+	}
+	RUN(verifier_typed_arena);
+}
 void test_verifier_typedef(void)              { RUN(verifier_typedef); }
 void test_verifier_uninit(void)               { RUN(verifier_uninit); }
 void test_verifier_unpriv(void)               { RUN(verifier_unpriv); }
diff --git a/tools/testing/selftests/bpf/progs/verifier_typed_arena.c b/tools/testing/selftests/bpf/progs/verifier_typed_arena.c
new file mode 100644
index 000000000000..22a8c3bfd493
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/verifier_typed_arena.c
@@ -0,0 +1,528 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+
+#include <vmlinux.h>
+#include <errno.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+#include "bpf_misc.h"
+#include "bpf_experimental.h"
+#include "bpf_arena_common.h"
+
+#ifdef __TARGET_ARCH_arm64
+#define ARENA_VM_START ((1ull << 32) | (~0u - __PAGE_SIZE * 2 + 1))
+#else
+#define ARENA_VM_START ((1ull << 44) | (~0u - __PAGE_SIZE * 2 + 1))
+#endif
+
+struct {
+	__uint(type, BPF_MAP_TYPE_ARENA);
+	__uint(map_flags, BPF_F_MMAPABLE);
+	__uint(max_entries, 2);
+	__ulong(map_extra, ARENA_VM_START);
+} arena SEC(".maps");
+
+struct {
+	__uint(type, BPF_MAP_TYPE_ARRAY);
+	__uint(max_entries, 1);
+	__type(key, __u32);
+	__type(value, __u64);
+} not_an_arena SEC(".maps");
+
+/* Tells the runner whether the compiler emits the cast, and so whether the tests below exist. */
+const volatile bool typed_arena_supported =
+#ifdef __BPF_FEATURE_TYPED_ARENA_CAST
+	true;
+#else
+	false;
+#endif
+
+#ifdef __BPF_FEATURE_TYPED_ARENA_CAST
+
+/* 16-byte slot, 8 Mi objects in the default 128 MiB typed arena: the cast mask is 134217712 */
+struct typed_obj {
+	struct task_struct __kptr *task;
+	__u64 value;
+};
+
+/* 32-byte slot */
+struct other_obj {
+	struct task_struct __kptr *task;
+	__u64 a;
+	__u64 b;
+};
+
+struct plain_obj {
+	__u64 value;
+};
+
+/*
+ * A program is associated with an arena by referencing the map. Programs that
+ * only convert values they already hold reference it explicitly.
+ */
+#define arena_bind() asm volatile("r0 = %[m] ll" :: [m] "i"(&arena) : "r0")
+
+/*
+ * Objects as the opaque 64-bit values user space hands in. Converting one to a
+ * pointer to a typed struct is where the compiler inserts the typed_arena_cast,
+ * as it does before every use of such a pointer as an address. A cast of a
+ * pointer the verifier already trusts lowers to nothing, so a lowered program
+ * shows one sequence per value that enters. The values stay unset here: the
+ * verifier does not care where a value lands, and at run time an unset one
+ * names object 0 of its slice.
+ */
+void *ptr;
+void *ptr2;
+
+SEC("syscall")
+__description("a cast masks the value into a slot and adds the typed arena base")
+__success __retval(0) __log_level(2)
+__msg("typed arena for struct typed_obj: slot 16 bytes")
+__msg("R{{[0-9]}}=typed_arena_ptr_typed_obj(")
+__xlated("r{{[0-9]}} &= 134217712")
+__xlated("r12 = 0x{{[0-9a-f]+}}")
+__xlated("r{{[0-9]}} += r12")
+int cast_lowers_to_mask_and_base(void *ctx)
+{
+	struct typed_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj->value;
+}
+
+SEC("syscall")
+__description("any 64-bit value casts: a raw arena pointer lands on an object")
+__success __retval(0)
+int cast_raw_arena_pointer(void *ctx)
+{
+	struct typed_obj *obj;
+	void __arena *page;
+
+	page = bpf_arena_alloc_pages(&arena, NULL, 1, NUMA_NO_NODE, 0);
+	if (!page)
+		return 1;
+	obj = (void *)page;
+	obj->value = 3;
+	return obj->value - 3;
+}
+
+SEC("syscall")
+__description("a typed pointer casts to itself, and to another type without knowing the source")
+__success __retval(0)
+int cast_typed_pointer_again(void *ctx)
+{
+	struct typed_obj *obj, *again;
+	struct other_obj *other;
+	void *opaque;
+
+	arena_bind();
+	obj = ptr;
+	opaque = obj;
+	again = opaque;
+	if (again != obj)
+		return 1;
+	other = (struct other_obj *)obj;
+	other->a = 1;
+	return other->a - 1;
+}
+
+/*
+ * The same struct lives in allocated objects, in map values and on the stack,
+ * and the compiler casts every use of a pointer to it there too. Such a cast
+ * is the identity: the pointer keeps its type and its state, nothing is
+ * lowered, and the access follows that pointer's own rules. These programs
+ * reference no arena, since nothing in them is sanitized.
+ */
+SEC("syscall")
+__description("a cast of an allocated object of a typed struct is the identity")
+__success __retval(0) __log_level(2)
+__msg("typed_arena_cast(r{{[0-9]}}, {{[0-9]+}})")
+__msg("R{{[0-9]}}=ptr_typed_obj(")
+int cast_allocated_object_is_identity(void *ctx)
+{
+	struct typed_obj *obj;
+
+	obj = bpf_obj_new(struct typed_obj);
+	if (!obj)
+		return 1;
+	obj->value = 1;
+	bpf_obj_drop(obj);
+	return 0;
+}
+
+struct value_with_typed {
+	__u64 pad;
+	struct typed_obj obj;
+};
+
+struct {
+	__uint(type, BPF_MAP_TYPE_ARRAY);
+	__uint(max_entries, 1);
+	__type(key, __u32);
+	__type(value, struct value_with_typed);
+} typed_in_map SEC(".maps");
+
+SEC("syscall")
+__description("a cast of a typed struct inside a map value is the identity")
+__success __retval(0) __log_level(2)
+__msg("typed_arena_cast(r{{[0-9]}}, {{[0-9]+}})")
+__msg("R{{[0-9]}}=map_value(")
+int cast_map_value_is_identity(void *ctx)
+{
+	struct value_with_typed *v;
+	struct typed_obj *obj;
+	__u32 key = 0;
+
+	v = bpf_map_lookup_elem(&typed_in_map, &key);
+	if (!v)
+		return 1;
+	obj = &v->obj;
+	obj->value = 1;
+	return 0;
+}
+
+SEC("syscall")
+__description("a cast of a typed struct on the stack is the identity")
+__success __retval(0) __log_level(2)
+__msg("typed_arena_cast(r{{[0-9]}}, {{[0-9]+}})")
+__msg("R{{[0-9]}}=fp-")
+int cast_stack_object_is_identity(void *ctx)
+{
+	struct typed_obj local = {};
+	struct typed_obj *obj = &local;
+
+	barrier_var(obj);
+	obj->value = 1;
+	return local.value - 1;
+}
+
+SEC("syscall")
+__description("the cast needs the program's arena")
+__failure __msg("typed_arena_cast insn can only be used in a program that has an associated arena")
+int cast_needs_arena(void *ctx)
+{
+	struct typed_obj *obj;
+
+	obj = ptr;
+	return obj->value;
+}
+
+struct locked_obj {
+	struct bpf_spin_lock lock;
+	__u64 value;
+};
+
+SEC("syscall")
+__description("spin locks are not supported in typed arena objects yet")
+__failure __msg("struct locked_obj field bpf_spin_lock is not supported in a typed arena")
+int cast_rejects_spin_lock(void *ctx)
+{
+	struct locked_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj->value;
+}
+
+struct res_locked_obj {
+	struct bpf_res_spin_lock lock;
+	__u64 value;
+};
+
+SEC("syscall")
+__description("resilient spin locks are not supported in typed arena objects yet")
+__failure __msg("struct res_locked_obj field bpf_res_spin_lock is not supported in a typed arena")
+int cast_rejects_res_spin_lock(void *ctx)
+{
+	struct res_locked_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj->value;
+}
+
+struct timer_obj {
+	struct bpf_timer timer;
+	__u64 value;
+};
+
+SEC("syscall")
+__description("timers are not supported in typed arena objects")
+__failure __msg("struct timer_obj field bpf_timer is not supported in a typed arena")
+int cast_rejects_timer(void *ctx)
+{
+	struct timer_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj->value;
+}
+
+struct refcount_obj {
+	struct bpf_refcount ref;
+	struct task_struct __kptr *task;
+	__u64 value;
+};
+
+SEC("syscall")
+__description("refcounts are not supported in typed arena objects")
+__failure __msg("struct refcount_obj field bpf_refcount is not supported in a typed arena")
+int cast_rejects_refcount(void *ctx)
+{
+	struct refcount_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj->value;
+}
+
+struct list_node_obj {
+	struct bpf_list_node node;
+	__u64 value;
+};
+
+struct list_obj {
+	struct bpf_list_head head __contains(list_node_obj, node);
+	struct bpf_spin_lock lock;
+};
+
+/* The node structs must reach the BTF for the contains tags to resolve; nothing else uses them. */
+struct list_node_obj *list_node_in_btf;
+
+SEC("syscall")
+__description("list heads are not supported in typed arena objects")
+__failure __msg("struct list_obj field bpf_list_head is not supported in a typed arena")
+int cast_rejects_list_head(void *ctx)
+{
+	struct list_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj == NULL;
+}
+
+struct rb_node_obj {
+	struct bpf_rb_node node;
+	__u64 value;
+};
+
+struct rb_obj {
+	struct bpf_rb_root root __contains(rb_node_obj, node);
+	struct bpf_spin_lock lock;
+};
+
+struct rb_node_obj *rb_node_in_btf;
+
+SEC("syscall")
+__description("rbtree roots are not supported in typed arena objects")
+__failure __msg("struct rb_obj field bpf_rb_root is not supported in a typed arena")
+int cast_rejects_rb_root(void *ctx)
+{
+	struct rb_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj == NULL;
+}
+
+struct size_not_pow2_obj {
+	struct task_struct __kptr *task;
+	__u64 value;
+} __typed_arena_size(3M);
+
+SEC("syscall")
+__description("a typed arena size must be a power of two")
+__failure __msg("struct size_not_pow2_obj has invalid typed arena size '3M'")
+int size_rejects_not_pow2(void *ctx)
+{
+	struct size_not_pow2_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj->value;
+}
+
+struct size_below_page_obj {
+	struct task_struct __kptr *task;
+	__u64 value;
+} __typed_arena_size(2048);
+
+SEC("syscall")
+__description("a typed arena is at least a page")
+__failure __msg("struct size_below_page_obj has invalid typed arena size '2048'")
+int size_rejects_below_page(void *ctx)
+{
+	struct size_below_page_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj->value;
+}
+
+struct size_above_region_obj {
+	struct task_struct __kptr *task;
+	__u64 value;
+} __typed_arena_size(8G);
+
+SEC("syscall")
+__description("a typed arena is at most the region")
+__failure __msg("struct size_above_region_obj has invalid typed arena size '8G'")
+int size_rejects_above_region(void *ctx)
+{
+	struct size_above_region_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj->value;
+}
+
+/* 128 KiB slot in a 64 KiB typed arena */
+struct size_below_slot_obj {
+	struct task_struct __kptr *task;
+	char pad[65536 + 8];
+} __typed_arena_size(64K);
+
+SEC("syscall")
+__description("a typed arena holds at least one object")
+__failure __msg("struct size_below_slot_obj does not fit its typed arena: slot 131072 bytes, size 65536 bytes")
+int size_rejects_below_slot(void *ctx)
+{
+	struct size_below_slot_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj->pad[0];
+}
+
+struct size_conflict_obj {
+	struct task_struct __kptr *task;
+	__u64 value;
+} __typed_arena_size(64K) __typed_arena_size(128K);
+
+SEC("syscall")
+__description("a struct declares one typed arena size")
+__failure __msg("struct size_conflict_obj has conflicting typed arena size declarations")
+int size_rejects_conflict(void *ctx)
+{
+	struct size_conflict_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj->value;
+}
+
+struct size_64k_obj {
+	struct task_struct __kptr *task;
+	__u64 value;
+} __typed_arena_size(64K);
+
+struct size_2m_obj {
+	struct task_struct __kptr *task;
+	__u64 value;
+} __typed_arena_size(2M);
+
+SEC("syscall")
+__description("a typed arena size takes a suffix and sets the object count")
+__success __retval(0) __log_level(2)
+__msg("typed arena for struct size_64k_obj: slot 16 bytes")
+__msg("size 65536 bytes")
+__msg("typed arena for struct size_2m_obj: slot 16 bytes")
+__msg("size 2097152 bytes")
+int size_accepts_suffix(void *ctx)
+{
+	struct size_64k_obj *small;
+	struct size_2m_obj *large;
+
+	arena_bind();
+	small = ptr;
+	large = ptr;
+	small->value = 1;
+	large->value = 2;
+	return small->value + large->value - 3;
+}
+
+/* 16 KiB slot: an object of more than a page, four of them in a 64 KiB typed arena */
+struct big_obj {
+	struct task_struct __kptr *task;
+	char pad[8192];
+} __typed_arena_size(64K);
+
+SEC("syscall")
+__description("an object larger than a page takes a slot of whole pages")
+__success __retval(0) __log_level(2)
+__msg("typed arena for struct big_obj: slot 16384 bytes")
+int size_object_beyond_a_page(void *ctx)
+{
+	struct big_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	obj->pad[0] = 1;
+	obj->pad[8191] = 2;
+	return obj->pad[0] + obj->pad[8191] - 3;
+}
+
+/* A slice is capped at 2 GiB, half the region: two of them fill it. */
+struct huge_obj {
+	struct task_struct __kptr *task;
+	__u64 value;
+} __typed_arena_size(2G);
+
+struct huge_obj2 {
+	struct task_struct __kptr *task;
+	__u64 value;
+} __typed_arena_size(2G);
+
+SEC("syscall")
+__description("the region has room for 4 GiB of typed arenas")
+__failure __msg("no room in the typed arena region for struct typed_obj")
+int size_region_runs_out(void *ctx)
+{
+	struct huge_obj *huge;
+	struct huge_obj2 *huge2;
+	struct typed_obj *obj;
+
+	arena_bind();
+	huge = ptr;
+	huge2 = ptr;
+	obj = ptr;
+	return huge->value + huge2->value + obj->value;
+}
+
+SEC("syscall")
+__description("a 32-bit copy of a typed pointer is a scalar, as for any pointer a privileged program holds")
+__success __retval(0) __log_level(2)
+__msg("R2=scalar(smin=0,smax=umax=0xffffffff,var_off=(0x0; 0xffffffff))")
+int narrow_copy_yields_scalar(void *ctx)
+{
+	struct typed_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	asm volatile("r1 = %[p];"
+		     "w2 = w1;"
+		     :: [p] "r"(obj)
+		     : "r1", "r2");
+	return 0;
+}
+
+SEC("syscall")
+__description("a narrow store of a typed pointer to the stack is an invalid spill")
+__failure __msg("invalid size of register spill")
+int narrow_store_is_invalid_spill(void *ctx)
+{
+	struct typed_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	asm volatile("r1 = %[p];"
+		     "*(u32 *)(r10 - 8) = r1;"
+		     :: [p] "r"(obj)
+		     : "r1", "memory");
+	return 0;
+}
+
+#endif /* __BPF_FEATURE_TYPED_ARENA_CAST */
+
+char _license[] SEC("license") = "GPL";
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 14/16] selftests/bpf: Test typed arena object access, kptrs and typed pointer fields
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (12 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 13/16] selftests/bpf: Test typed arena casts and registration Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 15/16] selftests/bpf: Test typed arena page allocation and release Kumar Kartikeya Dwivedi
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 16/16] selftests/bpf: Exercise typed arenas at run time Kumar Kartikeya Dwivedi
  15 siblings, 0 replies; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

Write and read a scalar field natively on a chunk faulted to scratch, run an
atomic on it, and move the pointer within the object; reject an access
beyond the object, a negative offset and a variable one. Walk a nested
struct and see a pointer to a kernel struct load as a scalar. See helpers
refuse the pointer as memory.

Exchange a kernel kptr into an object and out again, leave a local kptr in a
dummy object for the map to drop, and reject direct loads and stores of a
kptr field, an exchange on a scalar field and one with the wrong type.

Load a typed pointer field and see the pointer stay unsanitized until it is
used as an address: the dereference carries the mask, base load and add in
place, a compare of the value does not, a store of the raw value written by
hand is accepted, a second dereference through the same register does not
sanitize again, and pointer arithmetic sanitizes before it adds. A NULL
dereferences object 0 of the slice, whether held directly or loaded from
the field. The field takes a typed pointer, a loaded pointer or NULL, and
the compiler casts what a program assigns to it, so a scalar and a pointer
to another typed struct are accepted from C and rejected only when stored
by hand; reject a narrow load or store and an atomic. The same member in
raw arena memory is data: a stored pointer loads as a scalar and casts back
to its object where it is used. A struct whose only special field is a
pointer to a typed struct gets a typed arena of its own. One instruction
sanitizes one typed arena: reject it when two paths bring pointers of
different types, or a typed pointer on one path only, and reject the cast
the compiler inserts when one path brings a loaded arena pointer and
another an allocated object.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 .../bpf/progs/verifier_typed_arena.c          | 634 ++++++++++++++++++
 1 file changed, 634 insertions(+)

diff --git a/tools/testing/selftests/bpf/progs/verifier_typed_arena.c b/tools/testing/selftests/bpf/progs/verifier_typed_arena.c
index 22a8c3bfd493..808a65611d4e 100644
--- a/tools/testing/selftests/bpf/progs/verifier_typed_arena.c
+++ b/tools/testing/selftests/bpf/progs/verifier_typed_arena.c
@@ -523,6 +523,640 @@ int narrow_store_is_invalid_spill(void *ctx)
 	return 0;
 }
 
+SEC("syscall")
+__description("a scalar field is written and read natively, on a chunk faulted to scratch")
+__success __retval(7)
+__stderr("ERROR: Typed arena WRITE access to unallocated struct typed_obj at 0x{{[0-9a-f]+}}")
+int access_scalar_field(void *ctx)
+{
+	struct typed_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	obj->value = 7;
+	return obj->value;
+}
+
+SEC("syscall")
+__description("atomics run natively on a scalar field")
+__success __retval(3)
+int access_atomic(void *ctx)
+{
+	struct typed_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	obj->value = 1;
+	__sync_fetch_and_add(&obj->value, 2);
+	return obj->value;
+}
+
+SEC("syscall")
+__description("pointer arithmetic stays inside the object")
+__success __retval(9)
+int access_after_arithmetic(void *ctx)
+{
+	struct typed_obj *obj;
+	void *p;
+
+	arena_bind();
+	obj = ptr;
+	p = obj;
+	asm volatile("%[p] += 8" : [p] "+r"(p));
+	*(__u64 *)p = 9;
+	return obj->value;
+}
+
+SEC("syscall")
+__description("an access beyond the object is rejected")
+__failure __msg("access beyond struct typed_obj")
+int access_beyond_object(void *ctx)
+{
+	struct typed_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	((__u64 *)obj)[2] = 1;
+	return 0;
+}
+
+SEC("syscall")
+__description("a negative offset is rejected")
+__failure __msg("invalid negative access")
+int access_negative_offset(void *ctx)
+{
+	struct typed_obj *obj;
+	void *p;
+
+	arena_bind();
+	obj = ptr;
+	p = (void *)obj - 8;
+	return *(__u64 *)p;
+}
+
+SEC("syscall")
+__description("a variable offset is rejected")
+__failure __msg("{{variable (offset|typed_arena_ptr_ access)}}")
+int access_variable_offset(void *ctx)
+{
+	struct typed_obj *obj;
+	void *p;
+
+	arena_bind();
+	obj = ptr;
+	p = (void *)obj + (bpf_get_prandom_u32() & 8);
+	return *(__u64 *)p;
+}
+
+struct mixed_obj {
+	struct task_struct __kptr *task;
+	struct {
+		__u32 a;
+		__u32 b;
+	} inner;
+	struct task_struct *ptr;
+};
+
+SEC("syscall")
+__description("nested structs are walked and pointers to kernel structs load as scalars")
+__success __retval(3)
+int access_nested_and_kernel_pointer_fields(void *ctx)
+{
+	struct mixed_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	obj->inner.b = 3;
+	if (obj->ptr)
+		return 1;
+	return obj->inner.b;
+}
+
+SEC("syscall")
+__description("helpers do not take typed arena pointers as memory")
+__failure __msg("R1 type=typed_arena_ptr_ expected=")
+int helper_rejects_typed_pointer(void *ctx)
+{
+	struct typed_obj *obj;
+	__u64 src = 0;
+
+	arena_bind();
+	obj = ptr;
+	bpf_probe_read_kernel(&obj->value, sizeof(obj->value), &src);
+	return 0;
+}
+
+struct arena_node {
+	__u64 v;
+};
+
+struct kptr_obj {
+	struct task_struct __kptr *task;
+	struct arena_node __kptr *node;
+	__u64 value;
+};
+
+SEC("syscall")
+__description("a kernel kptr is exchanged into and out of an object")
+__success __retval(0)
+int kptr_xchg_task(void *ctx)
+{
+	struct task_struct *task, *old;
+	struct kptr_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	task = bpf_task_acquire(bpf_get_current_task_btf());
+	if (!task)
+		return 2;
+	old = bpf_kptr_xchg(&obj->task, task);
+	if (old)
+		bpf_task_release(old);
+	old = bpf_kptr_xchg(&obj->task, NULL);
+	if (!old)
+		return 3;
+	bpf_task_release(old);
+	return 0;
+}
+
+SEC("syscall")
+__description("a local kptr left in a dummy object is dropped with the map")
+__success __retval(0)
+int kptr_xchg_local(void *ctx)
+{
+	struct arena_node *n, *old;
+	struct kptr_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	n = bpf_obj_new(struct arena_node);
+	if (!n)
+		return 2;
+	n->v = 42;
+	old = bpf_kptr_xchg(&obj->node, n);
+	if (old)
+		bpf_obj_drop(old);
+	return 0;
+}
+
+SEC("syscall")
+__description("a kptr field is not read directly")
+__failure __msg("direct access to kptr is disallowed")
+int kptr_read_directly(void *ctx)
+{
+	struct kptr_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	return obj->task != NULL;
+}
+
+SEC("syscall")
+__description("a kptr field is not written directly")
+__failure __msg("direct access to kptr is disallowed")
+int kptr_write_directly(void *ctx)
+{
+	struct kptr_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	obj->task = NULL;
+	return 0;
+}
+
+SEC("syscall")
+__description("an exchange needs a kptr field")
+__failure __msg("off=16 doesn't point to kptr")
+int kptr_xchg_scalar_field(void *ctx)
+{
+	struct kptr_obj *obj;
+
+	arena_bind();
+	obj = ptr;
+	bpf_kptr_xchg(&obj->value, NULL);
+	return 0;
+}
+
+SEC("syscall")
+__description("an exchange checks the value against the field's type")
+__failure __msg("invalid kptr access, R2 type=ptr_arena_node expected=ptr_task_struct")
+int kptr_xchg_wrong_type(void *ctx)
+{
+	struct kptr_obj *obj;
+	struct arena_node *n;
+
+	arena_bind();
+	obj = ptr;
+	n = bpf_obj_new(struct arena_node);
+	if (!n)
+		return 2;
+	n = bpf_kptr_xchg(&obj->task, n);
+	if (n)
+		bpf_obj_drop(n);
+	return 0;
+}
+
+/* 32-byte slot, linked through a typed pointer field: the sanitize mask is 134217696 */
+struct node_obj {
+	struct task_struct __kptr *task;
+	struct node_obj *next;
+	__u64 value;
+};
+
+/* The same layout as node_obj, a different typed arena */
+struct pair_obj {
+	struct task_struct __kptr *task;
+	struct pair_obj *next;
+	__u64 value;
+};
+
+/* 8-byte slot, typed only by its pointer to a typed struct: the cast mask is 134217720 */
+struct head_obj {
+	struct node_obj *first;
+};
+
+SEC("syscall")
+__description("a typed pointer field loads unsanitized, and a dereference sanitizes it in place")
+__success __retval(0) __log_level(2)
+__msg("R{{[0-9]}}=unsanitized_typed_arena_ptr_node_obj(")
+__xlated("r{{[0-9]}} &= 134217720")
+__xlated("r12 = 0x{{[0-9a-f]+}}")
+__xlated("r{{[0-9]}} += r12")
+__xlated("...")
+__xlated("r{{[0-9]}} = *(u64 *)(r{{[0-9]}} +0)")
+__xlated("r{{[0-9]}} &= 134217696")
+__xlated("r12 = 0x{{[0-9a-f]+}}")
+__xlated("r{{[0-9]}} += r12")
+__xlated("r{{[0-9]}} = *(u64 *)(r{{[0-9]}} +16)")
+int ptr_field_deref_sanitizes(void *ctx)
+{
+	struct head_obj *h;
+	struct node_obj *n;
+
+	arena_bind();
+	h = ptr;
+	n = h->first;
+	return n->value;
+}
+
+SEC("syscall")
+__description("a compare and a store of a loaded typed pointer use the raw value")
+__success __retval(0) __log_level(2)
+__msg("R{{[0-9]}}=unsanitized_typed_arena_ptr_node_obj(")
+__msg("if r{{[0-9]}} == 0x0 goto")
+__msg("R{{[0-9]}}=unsanitized_typed_arena_ptr_node_obj(")
+__msg("*(u64 *)(r1 +8) = r2")
+int ptr_field_compare_and_store_stay_raw(void *ctx)
+{
+	struct node_obj *n, *m;
+
+	arena_bind();
+	n = ptr;
+	m = n->next;
+	/* Opaque to the compiler, so that the compare is emitted. */
+	barrier_var(m);
+	if (!m)
+		return 0;
+	/* The store is written by hand: the compiler would cast the value first. */
+	asm volatile("r1 = %[n];"
+		     "r2 = %[m];"
+		     "*(u64 *)(r1 + 8) = r2;"
+		     :: [n] "r"(n), [m] "r"(m)
+		     : "r1", "r2", "memory");
+	return 0;
+}
+
+SEC("syscall")
+__description("a register sanitized by one dereference is not sanitized again by the next")
+__success __retval(5)
+__xlated("r1 = *(u64 *)(r1 +8)")
+__xlated("r2 = 5")
+__xlated("r1 &= 134217696")
+__xlated("r12 = 0x{{[0-9a-f]+}}")
+__xlated("r1 += r12")
+__xlated("*(u64 *)(r1 +16) = r2")
+__xlated("r0 = *(u64 *)(r1 +16)")
+int ptr_field_second_deref_not_sanitized_again(void *ctx)
+{
+	struct node_obj *n;
+	__u64 ret;
+
+	arena_bind();
+	n = ptr;
+	/* Written by hand: the compiler would cast copies rather than reuse the register. */
+	asm volatile("r1 = %[n];"
+		     "r1 = *(u64 *)(r1 + 8);"
+		     "r2 = 5;"
+		     "*(u64 *)(r1 + 16) = r2;"
+		     "r0 = *(u64 *)(r1 + 16);"
+		     "%[ret] = r0;"
+		     : [ret] "=r"(ret) : [n] "r"(n)
+		     : "r0", "r1", "r2", "memory");
+	return ret;
+}
+
+SEC("syscall")
+__description("pointer arithmetic on a loaded typed pointer sanitizes it first")
+__success __retval(0)
+__xlated("r{{[0-9]}} = *(u64 *)(r{{[0-9]}} +8)")
+__xlated("...")
+__xlated("r1 &= 134217696")
+__xlated("r12 = 0x{{[0-9a-f]+}}")
+__xlated("r1 += r12")
+__xlated("r1 += 16")
+__xlated("*(u64 *)(r1 +0) = r2")
+int ptr_field_arithmetic_sanitizes(void *ctx)
+{
+	struct node_obj *n, *m;
+
+	arena_bind();
+	n = ptr;
+	m = n->next;
+	asm volatile("r1 = %[m];"
+		     "r2 = 0;"
+		     "r1 += 16;"
+		     "*(u64 *)(r1 + 0) = r2;"
+		     :: [m] "r"(m)
+		     : "r1", "r2", "memory");
+	return 0;
+}
+
+SEC("syscall")
+__description("a NULL typed pointer dereferences object 0 of the slice, held directly or loaded from a field")
+__success __retval(42)
+int ptr_field_null_lands_on_object_zero(void *ctx)
+{
+	struct node_obj *zero = NULL, *n, *m;
+
+	arena_bind();
+	/*
+	 * The compiler treats a NULL dereference as undefined and would drop
+	 * the store, fold the stored NULL into the load and the load into
+	 * nothing; the barriers keep each value opaque so that the accesses
+	 * are emitted.
+	 */
+	barrier_var(zero);
+	zero->value = 42;
+	n = ptr;
+	n->next = NULL;
+	barrier_var(n);
+	m = n->next;
+	barrier_var(m);
+	return m->value;
+}
+
+SEC("syscall")
+__description("a typed pointer field takes a typed pointer, a loaded one, or NULL")
+__success __retval(0)
+int ptr_field_store_accepted(void *ctx)
+{
+	struct node_obj *n, *m, *p;
+
+	arena_bind();
+	n = ptr;
+	m = ptr2;
+	n->next = m;
+	p = m->next;
+	n->next = p;
+	n->next = NULL;
+	return 0;
+}
+
+SEC("syscall")
+__description("a typed pointer field does not take a scalar")
+__failure __msg("store into typed pointer field of struct node_obj expects a typed arena pointer to struct node_obj or NULL")
+int ptr_field_store_scalar_rejected(void *ctx)
+{
+	struct node_obj *n;
+
+	arena_bind();
+	n = ptr;
+	/* Written by hand: the compiler would cast the value before the store. */
+	asm volatile("r6 = %[n];"
+		     "call %[bpf_get_prandom_u32];"
+		     "r1 = r6;"
+		     "*(u64 *)(r1 + 8) = r0;"
+		     :: [n] "r"(n), __imm(bpf_get_prandom_u32)
+		     : "r0", "r1", "r2", "r3", "r4", "r5", "r6", "memory");
+	return 0;
+}
+
+SEC("syscall")
+__description("a scalar assigned to a typed pointer field is cast by the compiler first")
+__success __retval(0)
+__xlated("call unknown")
+__xlated("...")
+__xlated("r{{[0-9]}} &= 134217696")
+__xlated("r12 = 0x{{[0-9a-f]+}}")
+__xlated("r{{[0-9]}} += r12")
+__xlated("*(u64 *)(r{{[0-9]}} +8) = r{{[0-9]}}")
+int ptr_field_store_scalar_cast_by_compiler(void *ctx)
+{
+	struct node_obj *n;
+
+	arena_bind();
+	n = ptr;
+	n->next = (void *)(long)bpf_get_prandom_u32();
+	return 0;
+}
+
+SEC("syscall")
+__description("a typed pointer field does not take a pointer to another typed struct")
+__failure __msg("store into typed pointer field of struct node_obj expects a typed arena pointer to struct node_obj or NULL")
+int ptr_field_store_other_type_rejected(void *ctx)
+{
+	struct typed_obj *other;
+	struct node_obj *n;
+
+	arena_bind();
+	n = ptr;
+	other = ptr;
+	/* Written by hand: the compiler would re-cast the value to node_obj first. */
+	asm volatile("r1 = %[n];"
+		     "r2 = %[o];"
+		     "*(u64 *)(r1 + 8) = r2;"
+		     :: [n] "r"(n), [o] "r"(other)
+		     : "r1", "r2", "memory");
+	return 0;
+}
+
+SEC("syscall")
+__description("a pointer to another typed struct assigned to a typed pointer field is re-cast by the compiler")
+__success __retval(0)
+int ptr_field_store_other_type_recast(void *ctx)
+{
+	struct typed_obj *other;
+	struct node_obj *n;
+
+	arena_bind();
+	n = ptr;
+	other = ptr;
+	n->next = (struct node_obj *)other;
+	return n->next == NULL;
+}
+
+SEC("syscall")
+__description("a typed pointer field is loaded whole")
+__failure __msg("typed pointer field of struct node_obj must be accessed with a 64-bit load or store")
+int ptr_field_narrow_load_rejected(void *ctx)
+{
+	struct node_obj *n;
+
+	arena_bind();
+	n = ptr;
+	asm volatile("r1 = %[n];"
+		     "w2 = *(u32 *)(r1 + 8);"
+		     :: [n] "r"(n)
+		     : "r1", "r2");
+	return 0;
+}
+
+SEC("syscall")
+__description("a typed pointer field is stored whole")
+__failure __msg("typed pointer field of struct node_obj must be accessed with a 64-bit load or store")
+int ptr_field_narrow_store_rejected(void *ctx)
+{
+	struct node_obj *n;
+
+	arena_bind();
+	n = ptr;
+	asm volatile("r1 = %[n];"
+		     "w2 = 0;"
+		     "*(u32 *)(r1 + 8) = w2;"
+		     :: [n] "r"(n)
+		     : "r1", "r2", "memory");
+	return 0;
+}
+
+SEC("syscall")
+__description("a typed pointer field takes no atomic operation")
+__failure __msg("typed pointer field of struct node_obj")
+int ptr_field_atomic_rejected(void *ctx)
+{
+	struct node_obj *n, *m;
+
+	arena_bind();
+	n = ptr;
+	m = ptr2;
+	__sync_val_compare_and_swap((__u64 *)&n->next, 0, (__u64)m);
+	return 0;
+}
+
+/* The same member outside a typed object is data */
+struct raw_holder {
+	struct node_obj *n;
+	__u64 v;
+};
+
+SEC("syscall")
+__description("a pointer stored in raw arena memory loads as a scalar and casts back to its object")
+__success __retval(0)
+int ptr_field_in_raw_memory_is_scalar(void *ctx)
+{
+	struct raw_holder __arena *h;
+	struct node_obj *n, *again;
+
+	h = bpf_arena_alloc_pages(&arena, NULL, 1, NUMA_NO_NODE, 0);
+	if (!h)
+		return 1;
+	n = ptr;
+	n->value = 7;
+	h->n = n;
+	again = h->n;
+	return again != n || again->value != 7;
+}
+
+SEC("syscall")
+__description("a struct typed only by a pointer to a typed struct gets a typed arena of its own")
+__success __retval(0) __log_level(2)
+__msg("typed arena for struct head_obj: slot 8 bytes")
+int ptr_field_makes_struct_typed(void *ctx)
+{
+	struct head_obj *h;
+	struct node_obj *n;
+
+	arena_bind();
+	h = ptr;
+	n = ptr2;
+	h->first = n;
+	return h->first != n;
+}
+
+/*
+ * The cast the compiler inserts is one instruction on every path through it:
+ * a loaded typed pointer needs the sanitizing sequence there, an allocated
+ * object needs the identity, and the two cannot share a lowering.
+ */
+SEC("syscall")
+__description("one cast cannot both sanitize an arena pointer and pass an allocated object")
+__failure __msg("casts values that need different treatment on different paths")
+int cast_conflicts_between_arena_and_allocated(void *ctx)
+{
+	struct node_obj *n, *a;
+	struct head_obj *h;
+
+	arena_bind();
+	a = bpf_obj_new(struct node_obj);
+	if (!a)
+		return 1;
+	h = ptr;
+	if (bpf_get_prandom_u32() & 1)
+		n = h->first;
+	else
+		n = a;
+	n->value = 1;
+	bpf_obj_drop(a);
+	return 0;
+}
+
+SEC("syscall")
+__description("one instruction sanitizes one typed arena: two types on two paths are rejected")
+__failure __msg("sanitizes typed arena pointers of different types on different paths")
+int ptr_field_sanitize_two_types_at_one_insn(void *ctx)
+{
+	struct node_obj *n;
+	struct pair_obj *p;
+
+	arena_bind();
+	n = ptr;
+	p = ptr;
+	asm volatile("call %[bpf_get_prandom_u32];"
+		     "if r0 == 0 goto 1f;"
+		     "r1 = %[n];"
+		     "r1 = *(u64 *)(r1 + 8);"
+		     "goto 2f;"
+		     "1: r1 = %[p];"
+		     "r1 = *(u64 *)(r1 + 8);"
+		     "2: r2 = *(u64 *)(r1 + 16);"
+		     :: [n] "r"(n), [p] "r"(p), __imm(bpf_get_prandom_u32)
+		     : "r0", "r1", "r2", "r3", "r4", "r5", "memory");
+	return 0;
+}
+
+SEC("syscall")
+__description("one instruction sanitizes one typed arena: a typed pointer on one path only is rejected")
+__failure __msg("sanitizes a typed arena pointer only on some paths")
+int ptr_field_sanitize_on_one_path(void *ctx)
+{
+	struct node_obj *n;
+
+	arena_bind();
+	n = ptr;
+	asm volatile("r2 = 0;"
+		     "*(u64 *)(r10 - 16) = r2;"
+		     "call %[bpf_get_prandom_u32];"
+		     "if r0 == 0 goto 1f;"
+		     "r1 = %[n];"
+		     "r1 = *(u64 *)(r1 + 8);"
+		     "goto 2f;"
+		     "1: r1 = r10;"
+		     "r1 += -32;"
+		     "2: r2 = *(u64 *)(r1 + 16);"
+		     :: [n] "r"(n), __imm(bpf_get_prandom_u32)
+		     : "r0", "r1", "r2", "r3", "r4", "r5", "memory");
+	return 0;
+}
+
 #endif /* __BPF_FEATURE_TYPED_ARENA_CAST */
 
 char _license[] SEC("license") = "GPL";
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 15/16] selftests/bpf: Test typed arena page allocation and release
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (13 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 14/16] selftests/bpf: Test typed arena object access, kptrs and typed pointer fields Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:46   ` sashiko-bot
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 16/16] selftests/bpf: Exercise typed arenas at run time Kumar Kartikeya Dwivedi
  15 siblings, 1 reply; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

Allocate a page of typed arena objects, write one, reach it again through
an opaque copy of its address and its neighbour by arithmetic on that copy,
and check that the kfunc receives the registered typed arena in place of
the type ID. Take the chunk a value names and see the same chunk refused
twice, once for the chunk and once for a page inside it. Ask for one page
of an object that spans several and see the request rounded up to the
object's chunk, with the granted count written back. Release a chunk and
see it keep its objects and stay taken within the same invocation, since
the release is deferred behind a grace period.

Reject a use of the returned pointer without a NULL check, a struct without
special fields, and a map that is not the program's arena. Declare the two
kfuncs in bpf_experimental.h.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 .../testing/selftests/bpf/bpf_experimental.h  |  12 ++
 .../bpf/progs/verifier_typed_arena.c          | 142 ++++++++++++++++++
 2 files changed, 154 insertions(+)

diff --git a/tools/testing/selftests/bpf/bpf_experimental.h b/tools/testing/selftests/bpf/bpf_experimental.h
index 128654997328..cc1537a5e07a 100644
--- a/tools/testing/selftests/bpf/bpf_experimental.h
+++ b/tools/testing/selftests/bpf/bpf_experimental.h
@@ -15,6 +15,18 @@
  */
 #define __typed_arena_size(sz) __attribute__((btf_decl_tag("typed_arena_size:" #sz)))
 
+/*
+ * Back the chunks covering *page_cnt pages of the typed arena of the struct
+ * whose local BTF type ID is type_id with zeroed objects, at addr, a typed
+ * pointer naming the first chunk, or anywhere for NULL; round the count up
+ * to whole chunks and write it back; return a pointer to the first object,
+ * or NULL. Release such a range after a grace period, after which its
+ * objects read as the dummy object. Both are safe under a spin lock.
+ */
+extern void *bpf_typed_arena_alloc_pages(void *map, __u64 type_id, void *addr, __u32 *page_cnt,
+					 int node_id) __ksym;
+extern void bpf_typed_arena_free_pages(void *map, __u64 type_id, void *ptr, __u32 page_cnt) __ksym;
+
 /* Convenience macro to wrap over bpf_obj_new */
 #define bpf_obj_new(type) ((type *)bpf_obj_new(bpf_core_type_id_local(type)))
 
diff --git a/tools/testing/selftests/bpf/progs/verifier_typed_arena.c b/tools/testing/selftests/bpf/progs/verifier_typed_arena.c
index 808a65611d4e..8ec75fd1f97b 100644
--- a/tools/testing/selftests/bpf/progs/verifier_typed_arena.c
+++ b/tools/testing/selftests/bpf/progs/verifier_typed_arena.c
@@ -1157,6 +1157,148 @@ int ptr_field_sanitize_on_one_path(void *ctx)
 	return 0;
 }
 
+#define TYPE_ID(T) bpf_core_type_id_local(T)
+
+SEC("syscall")
+__description("allocated pages hold real objects, reachable through any value that lands in them")
+__success __retval(0)
+__xlated("r2 = 0x{{[0-9a-f]+[0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f]}}")
+__xlated("call kernel-function")
+int pages_alloc(void *ctx)
+{
+	struct typed_obj *obj, *again, *next;
+	void *opaque;
+	__u32 cnt = 1;
+
+	obj = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct typed_obj), NULL, &cnt,
+					  NUMA_NO_NODE);
+	if (!obj)
+		return 1;
+	if (cnt != 1)
+		return 2;
+	obj->value = 5;
+	/* The object through an opaque value, and its neighbor by arithmetic on that value */
+	opaque = obj;
+	again = opaque;
+	if (again != obj || again->value != 5)
+		return 3;
+	next = opaque + sizeof(*obj);
+	next->value = 6;
+	if (next == obj || next->value != 6 || obj->value != 5)
+		return 4;
+	return 0;
+}
+
+SEC("syscall")
+__description("a fixed request takes its chunk once")
+__success __retval(0)
+int pages_alloc_fixed(void *ctx)
+{
+	struct typed_obj *obj, *hint;
+	__u32 cnt;
+
+	/* The chunk user space names, the first one here */
+	hint = ptr;
+	cnt = 2;
+	obj = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct typed_obj), hint, &cnt,
+					  NUMA_NO_NODE);
+	if (!obj)
+		return 1;
+	if (obj != hint || cnt != 2)
+		return 2;
+	cnt = 1;
+	if (bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct typed_obj), hint, &cnt,
+					NUMA_NO_NODE))
+		return 3;
+	/* The page after it, part of the first request */
+	hint = (void *)hint + __PAGE_SIZE;
+	cnt = 1;
+	if (bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct typed_obj), hint, &cnt,
+					NUMA_NO_NODE))
+		return 4;
+	return 0;
+}
+
+SEC("syscall")
+__description("a request is rounded up to whole chunks and the granted count written back")
+__success __retval(0)
+int pages_alloc_granted_count(void *ctx)
+{
+	__u32 cnt = 1, chunk_pages = 16384 > __PAGE_SIZE ? 16384 / __PAGE_SIZE : 1;
+	struct big_obj *obj;
+
+	obj = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct big_obj), NULL, &cnt,
+					  NUMA_NO_NODE);
+	if (!obj)
+		return 1;
+	if (cnt != chunk_pages)
+		return 2;
+	obj->pad[8191] = 1;
+	return obj->pad[8191] - 1;
+}
+
+SEC("syscall")
+__description("a released chunk keeps its objects and stays taken until the grace period has passed")
+__success __retval(0)
+int pages_free(void *ctx)
+{
+	struct typed_obj *obj;
+	__u32 cnt = 1;
+
+	obj = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct typed_obj), NULL, &cnt,
+					  NUMA_NO_NODE);
+	if (!obj)
+		return 1;
+	obj->value = 7;
+	bpf_typed_arena_free_pages(&arena, TYPE_ID(struct typed_obj), obj, 1);
+	if (obj->value != 7)
+		return 2;
+	cnt = 1;
+	if (bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct typed_obj), obj, &cnt, NUMA_NO_NODE))
+		return 3;
+	return 0;
+}
+
+SEC("syscall")
+__description("an allocation must be checked for NULL")
+__failure __msg("invalid mem access 'typed_arena_ptr_or_null_'")
+int pages_alloc_null_check(void *ctx)
+{
+	struct typed_obj *obj;
+	__u32 cnt = 1;
+
+	obj = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct typed_obj), NULL, &cnt,
+					  NUMA_NO_NODE);
+	obj->value = 1;
+	return 0;
+}
+
+SEC("syscall")
+__description("the page kfuncs register the type like a cast: a struct without special fields belongs in the raw arena")
+__failure __msg("struct plain_obj has no special fields and needs no typed arena")
+int pages_alloc_plain_struct(void *ctx)
+{
+	struct plain_obj *obj;
+	__u32 cnt = 1;
+
+	obj = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct plain_obj), NULL, &cnt,
+					  NUMA_NO_NODE);
+	return obj == NULL;
+}
+
+SEC("syscall")
+__description("the page kfuncs need the program's arena")
+__failure __msg("can only be used in a program that has an associated arena")
+int pages_alloc_needs_arena(void *ctx)
+{
+	struct typed_obj *obj;
+	__u32 cnt = 1;
+
+	obj = bpf_typed_arena_alloc_pages(&not_an_arena, TYPE_ID(struct typed_obj), NULL, &cnt,
+					  NUMA_NO_NODE);
+	return obj == NULL;
+}
+
 #endif /* __BPF_FEATURE_TYPED_ARENA_CAST */
 
 char _license[] SEC("license") = "GPL";
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [RFC PATCH bpf-next v1 16/16] selftests/bpf: Exercise typed arenas at run time
  2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
                   ` (14 preceding siblings ...)
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 15/16] selftests/bpf: Test typed arena page allocation and release Kumar Kartikeya Dwivedi
@ 2026-09-26 23:34 ` Kumar Kartikeya Dwivedi
  2026-09-26 23:50   ` sashiko-bot
  15 siblings, 1 reply; 27+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-26 23:34 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, kkd, kernel-team

Drive typed arenas from user space through a skeleton. Check that
registration at load is accounted in the map's memory usage, that an
allocated chunk adds its pages to it and a released one takes them away,
that a value written through one invocation is read by another, and that a
chunk is refused to a fixed request while it is taken. Leave a reference to
the test module's object in a kptr field, release the chunk, and poll the
object's reference count until the deferred release has dropped it, which
proves the grace period and the field teardown ran; then claim the same
chunk again and see it come back zeroed. Read an object nobody allocated and
see the dummy object, see the chunk it faulted in refused to the allocator,
release it, and poll until it can be claimed. Load the same object twice on
one arena map and check that the two program BTFs get distinct typed arenas
for the same struct.

Release the pages of one allocation one at a time and see them all come
back, since the worker walks every span of a batch. Fill a typed arena, ask
to release one page more than it holds past the address, and see the request
refused while a release of the last chunk alone goes through. Release a
chunk twice, the second time while the worker is waiting out the grace
period of the first, claim the chunk again as soon as it is free, and see an
object written to it survive the repeated request. Allocate more pages in
one request than the allocator's batch holds and see the whole range served
and returned. Allocate an object of two pages with a request for one and see
the whole object granted, both pages usable, and the chunk returned.

Carve objects out of an allocated page, link them into a list through a
typed pointer field and walk it with no cast: the sum of the values only
reads, a second pass writes through every pointer, and the NULL that ends
the list, dereferenced on purpose, reads object 0 of the slice.

User space hands objects around as opaque 64-bit values; converting one to
a typed pointer casts it, and any value casts to an object, so a small
offset names one as well as a pointer does. Programs that only convert
values they hold reference the arena map explicitly, since that is what
associates the arena with a program. The tests need a compiler that emits
the cast and are skipped otherwise.

Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
 .../selftests/bpf/prog_tests/typed_arena.c    | 489 ++++++++++++++++++
 .../testing/selftests/bpf/progs/typed_arena.c | 354 +++++++++++++
 2 files changed, 843 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/typed_arena.c
 create mode 100644 tools/testing/selftests/bpf/progs/typed_arena.c

diff --git a/tools/testing/selftests/bpf/prog_tests/typed_arena.c b/tools/testing/selftests/bpf/prog_tests/typed_arena.c
new file mode 100644
index 000000000000..8fcb6bb6a51f
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/typed_arena.c
@@ -0,0 +1,489 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <test_progs.h>
+#include <sys/user.h>
+#ifndef PAGE_SIZE /* on some archs it comes in sys/user.h */
+#include <unistd.h>
+#define PAGE_SIZE getpagesize()
+#endif
+
+#include "typed_arena.skel.h"
+
+/* The sizes progs/typed_arena.c declares */
+#define OBJ_ARENA_SIZE (256 * 1024)
+#define WIDE_ARENA_SIZE (8 * 1024 * 1024)
+#define BIG_OBJ_SIZE 8192
+#define LIST_LEN 64
+
+/* The map's memory usage as the "memlock:" line of its fdinfo. */
+static long map_memlock(int map_fd)
+{
+	char path[64], line[128];
+	long memlock = -1;
+	FILE *f;
+
+	snprintf(path, sizeof(path), "/proc/self/fdinfo/%d", map_fd);
+	f = fopen(path, "r");
+	if (!ASSERT_OK_PTR(f, "open_fdinfo"))
+		return -1;
+	while (fgets(line, sizeof(line), f)) {
+		if (sscanf(line, "memlock:\t%ld", &memlock) == 1)
+			break;
+	}
+	fclose(f);
+	ASSERT_NEQ(memlock, -1, "parse_memlock");
+	return memlock;
+}
+
+static int run_ret(struct bpf_program *prog, const char *name)
+{
+	LIBBPF_OPTS(bpf_test_run_opts, opts);
+	int err = bpf_prog_test_run_opts(bpf_program__fd(prog), &opts);
+
+	if (!ASSERT_OK(err, name))
+		return -1;
+	return opts.retval;
+}
+
+static int run(struct bpf_program *prog, const char *name)
+{
+	int ret = run_ret(prog, name);
+
+	if (!ASSERT_OK(ret, name))
+		return -1;
+	return 0;
+}
+
+/* Poll the module object's reference count until a deferred release has run. */
+static long wait_ref_cnt(struct typed_arena *skel, long want)
+{
+	int i;
+
+	for (i = 0; i < 500; i++) {
+		if (run(skel->progs.read_ref_cnt, "read_ref_cnt"))
+			return -1;
+		if (skel->bss->ref_cnt == want)
+			break;
+		usleep(10000);
+	}
+	return skel->bss->ref_cnt;
+}
+
+/* Poll the map's memory usage until it reaches the expected value. */
+static long wait_memlock(int map_fd, long want)
+{
+	long memlock = -1;
+	int i;
+
+	for (i = 0; i < 500; i++) {
+		memlock = map_memlock(map_fd);
+		if (memlock == want)
+			break;
+		usleep(10000);
+	}
+	return memlock;
+}
+
+/* Poll a fixed allocation until the release of the chunk it names has run. */
+static int wait_alloc_at(struct typed_arena *skel)
+{
+	int i, ret = -1;
+
+	for (i = 0; i < 500; i++) {
+		ret = run_ret(skel->progs.alloc_at, "alloc_at");
+		if (ret <= 0)
+			break;
+		usleep(10000);
+	}
+	return ret;
+}
+
+/* Objects persist across invocations, and chunks come and go around them. */
+static void test_pages(void)
+{
+	long ps = PAGE_SIZE, base, base_cnt;
+	struct typed_arena *skel;
+	void *p;
+	int fd;
+
+	skel = typed_arena__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "open_load"))
+		return;
+	fd = bpf_map__fd(skel->maps.arena);
+
+	/* Registration at load accounts the page tables and the scratch chunks. */
+	base = map_memlock(fd);
+	ASSERT_GE(base, ps, "registered");
+	if (run(skel->progs.read_ref_cnt, "base_ref_cnt"))
+		goto out;
+	base_cnt = skel->bss->ref_cnt;
+
+	if (run(skel->progs.alloc, "alloc"))
+		goto out;
+	ASSERT_EQ(skel->data->page_cnt, 1, "granted");
+	p = skel->bss->ptr;
+	ASSERT_EQ(map_memlock(fd), base + ps, "after_alloc");
+
+	skel->bss->value = 42;
+	if (run(skel->progs.write_value, "write_value"))
+		goto out;
+	skel->bss->value = 0;
+	if (run(skel->progs.read_value, "read_value"))
+		goto out;
+	ASSERT_EQ(skel->bss->value, 42, "persisted");
+
+	/* The chunk is taken until released. */
+	ASSERT_EQ(run_ret(skel->progs.alloc_at, "alloc_at"), 1, "taken");
+
+	if (run(skel->progs.stash_ref, "stash_ref"))
+		goto out;
+	ASSERT_EQ(wait_ref_cnt(skel, base_cnt + 1), base_cnt + 1, "ref_stashed");
+
+	/* Release drops the reference after the grace period and returns the memory. */
+	if (run(skel->progs.free_pages, "free_pages"))
+		goto out;
+	ASSERT_EQ(wait_ref_cnt(skel, base_cnt), base_cnt, "ref_dropped");
+	ASSERT_EQ(map_memlock(fd), base, "after_free");
+
+	/* The same chunk can be claimed again, and comes back zeroed. */
+	ASSERT_EQ(wait_alloc_at(skel), 0, "reclaimed");
+	ASSERT_EQ(skel->bss->ptr, p, "same_object");
+	if (run(skel->progs.read_value, "read_fresh"))
+		goto out;
+	ASSERT_EQ(skel->bss->value, 0, "fresh");
+	if (run(skel->progs.free_pages, "free_again"))
+		goto out;
+
+	/*
+	 * An object nobody allocated reads as the dummy object, and the chunk
+	 * it faults in is taken until released in turn. Any value casts to an
+	 * object, so a small offset names one as well as a pointer does.
+	 */
+	skel->bss->ptr = (void *)(7 * ps);
+	skel->bss->value = 1;
+	if (run(skel->progs.read_value, "read_unallocated"))
+		goto out;
+	ASSERT_EQ(skel->bss->value, 0, "dummy");
+	ASSERT_EQ(run_ret(skel->progs.alloc_at, "alloc_scratch_chunk"), 1, "scratch_taken");
+	if (run(skel->progs.free_pages, "free_scratch_chunk"))
+		goto out;
+	ASSERT_EQ(wait_alloc_at(skel), 0, "scratch_released");
+out:
+	typed_arena__destroy(skel);
+}
+
+/*
+ * A typed arena belongs to the program BTF that registered it. Two loads of
+ * the same object sharing one arena map carry two BTF objects, so a pointer
+ * of one names an object in a different typed arena when the other casts it.
+ */
+static void test_identity(void)
+{
+	struct typed_arena *skel1, *skel2 = NULL;
+
+	skel1 = typed_arena__open_and_load();
+	if (!ASSERT_OK_PTR(skel1, "open_load1"))
+		return;
+	skel2 = typed_arena__open();
+	if (!ASSERT_OK_PTR(skel2, "open2"))
+		goto out;
+	if (!ASSERT_OK(bpf_map__reuse_fd(skel2->maps.arena, bpf_map__fd(skel1->maps.arena)),
+		       "reuse_fd"))
+		goto out;
+	if (!ASSERT_OK(typed_arena__load(skel2), "load2"))
+		goto out;
+
+	if (run(skel1->progs.alloc, "alloc1"))
+		goto out;
+	skel1->bss->value = 7;
+	if (run(skel1->progs.write_value, "write1"))
+		goto out;
+	skel2->bss->ptr = skel1->bss->ptr;
+	if (run(skel2->progs.read_value, "read2"))
+		goto out;
+	ASSERT_EQ(skel2->bss->value, 0, "distinct_typed_arena");
+	if (run(skel1->progs.read_value, "read1"))
+		goto out;
+	ASSERT_EQ(skel1->bss->value, 7, "own_typed_arena");
+out:
+	typed_arena__destroy(skel2);
+	typed_arena__destroy(skel1);
+}
+
+/* Every span of a batch of releases is processed, not only the first one with real pages. */
+static void test_release_spans(void)
+{
+	struct typed_arena *skel;
+	long ps = PAGE_SIZE, base;
+	void *p;
+	int fd, i;
+
+	skel = typed_arena__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "open_load"))
+		return;
+	fd = bpf_map__fd(skel->maps.arena);
+	base = map_memlock(fd);
+
+	skel->data->page_cnt = 8;
+	if (run(skel->progs.alloc, "alloc"))
+		goto out;
+	ASSERT_EQ(skel->data->page_cnt, 8, "granted");
+	p = skel->bss->ptr;
+	ASSERT_EQ(map_memlock(fd), base + 8 * ps, "after_alloc");
+
+	/* Release the pages one at a time, so that the worker sees a span per page. */
+	skel->data->page_cnt = 1;
+	for (i = 0; i < 8; i++) {
+		skel->bss->ptr = p + i * ps;
+		if (run(skel->progs.free_pages, "free_pages"))
+			goto out;
+	}
+	ASSERT_EQ(wait_memlock(fd, base), base, "all_released");
+out:
+	typed_arena__destroy(skel);
+}
+
+/* A release of more pages than the typed arena holds past the address is refused. */
+static void test_release_bounds(void)
+{
+	long ps = PAGE_SIZE, base, nr = OBJ_ARENA_SIZE / PAGE_SIZE;
+	struct typed_arena *skel;
+	int fd;
+
+	skel = typed_arena__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "open_load"))
+		return;
+	fd = bpf_map__fd(skel->maps.arena);
+	base = map_memlock(fd);
+
+	/* Fill the typed arena, so that a scan of its bitmap for a free chunk runs to the end. */
+	skel->bss->ptr = NULL;
+	skel->data->page_cnt = nr;
+	if (run(skel->progs.alloc_at, "fill"))
+		goto out;
+	ASSERT_EQ(map_memlock(fd), base + nr * ps, "full");
+
+	skel->data->page_cnt = nr + 1;
+	if (run(skel->progs.free_pages, "free_oversized"))
+		goto out;
+
+	/*
+	 * A release of the last chunk alone goes through. Once it has run,
+	 * anything queued before it has run too.
+	 */
+	skel->bss->ptr = (void *)((nr - 1) * ps);
+	skel->data->page_cnt = 1;
+	if (run(skel->progs.free_pages, "free_last"))
+		goto out;
+	ASSERT_EQ(wait_alloc_at(skel), 0, "last_released");
+	ASSERT_EQ(map_memlock(fd), base + nr * ps, "still_full");
+	skel->bss->ptr = NULL;
+	ASSERT_EQ(run_ret(skel->progs.alloc_at, "alloc_at"), 1, "first_still_taken");
+
+	skel->data->page_cnt = nr;
+	if (run(skel->progs.free_pages, "free_all"))
+		goto out;
+	ASSERT_EQ(wait_memlock(fd, base), base, "released");
+out:
+	typed_arena__destroy(skel);
+}
+
+/*
+ * A chunk is released once: a second release of a chunk whose release is
+ * queued is refused, so that it cannot take away whatever claims the chunk
+ * after the first release has run. The repeated request is made while the
+ * worker is waiting out the grace period of the first, which puts it in a
+ * later batch; a request that arrives before the worker starts joins the
+ * first batch instead and is harmless either way.
+ */
+static void test_release_once(void)
+{
+	struct typed_arena *skel;
+	void *p1, *p2;
+	long base;
+	int fd;
+
+	skel = typed_arena__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "open_load"))
+		return;
+	fd = bpf_map__fd(skel->maps.arena);
+	base = map_memlock(fd);
+
+	if (run(skel->progs.alloc, "alloc1"))
+		goto out;
+	p1 = skel->bss->ptr;
+	if (run(skel->progs.alloc, "alloc2"))
+		goto out;
+	p2 = skel->bss->ptr;
+
+	skel->bss->ptr = p1;
+	if (run(skel->progs.free_pages, "free1"))
+		goto out;
+	usleep(1000);
+	if (run(skel->progs.free_pages, "free1_again"))
+		goto out;
+	skel->bss->ptr = p2;
+	if (run(skel->progs.free_pages, "free2"))
+		goto out;
+
+	/* Claim the first chunk again as soon as its release has run, and write to it. */
+	skel->bss->ptr = p1;
+	ASSERT_EQ(wait_alloc_at(skel), 0, "reclaimed");
+	skel->bss->value = 42;
+	if (run(skel->progs.write_value, "write_value"))
+		goto out;
+
+	/*
+	 * The second chunk's release was queued after the repeated one. Once it
+	 * has run, so has the repeated one, if it was accepted.
+	 */
+	skel->bss->ptr = p2;
+	ASSERT_EQ(wait_alloc_at(skel), 0, "marker_released");
+
+	skel->bss->ptr = p1;
+	skel->bss->value = 0;
+	if (run(skel->progs.read_value, "read_value"))
+		goto out;
+	ASSERT_EQ(skel->bss->value, 42, "kept");
+
+	if (run(skel->progs.free_pages, "free1_final"))
+		goto out;
+	skel->bss->ptr = p2;
+	if (run(skel->progs.free_pages, "free2_final"))
+		goto out;
+	ASSERT_EQ(wait_memlock(fd, base), base, "released");
+out:
+	typed_arena__destroy(skel);
+}
+
+/* A request larger than one batch of the allocator is served whole. */
+static void test_batch_alloc(void)
+{
+	long ps = PAGE_SIZE, base, nr = WIDE_ARENA_SIZE / PAGE_SIZE;
+	struct typed_arena *skel;
+	void *p;
+	int fd;
+
+	skel = typed_arena__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "open_load"))
+		return;
+	fd = bpf_map__fd(skel->maps.arena);
+	base = map_memlock(fd);
+
+	skel->data->page_cnt = nr;
+	if (run(skel->progs.wide_alloc, "wide_alloc"))
+		goto out;
+	ASSERT_EQ(skel->data->page_cnt, nr, "granted");
+	p = skel->bss->ptr;
+	ASSERT_EQ(map_memlock(fd), base + nr * ps, "after_alloc");
+
+	/* The last page of the range is usable. */
+	skel->bss->ptr = p + (nr - 1) * ps;
+	skel->bss->value = 42;
+	if (run(skel->progs.wide_touch, "wide_touch"))
+		goto out;
+
+	skel->bss->ptr = p;
+	if (run(skel->progs.wide_free, "wide_free"))
+		goto out;
+	ASSERT_EQ(wait_memlock(fd, base), base, "released");
+out:
+	typed_arena__destroy(skel);
+}
+
+/* An object of more than a page is backed and released as one chunk. */
+static void test_multipage(void)
+{
+	long ps = PAGE_SIZE, base, granted;
+	struct typed_arena *skel;
+	int fd;
+
+	granted = BIG_OBJ_SIZE > ps ? BIG_OBJ_SIZE / ps : 1;
+
+	skel = typed_arena__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "open_load"))
+		return;
+	fd = bpf_map__fd(skel->maps.arena);
+	base = map_memlock(fd);
+
+	/* One page requested, the whole object granted. */
+	skel->data->page_cnt = 1;
+	if (run(skel->progs.big_alloc, "big_alloc"))
+		goto out;
+	ASSERT_EQ(skel->data->page_cnt, granted, "granted");
+	ASSERT_EQ(map_memlock(fd), base + granted * ps, "after_alloc");
+
+	skel->bss->value = 42;
+	if (run(skel->progs.big_touch, "big_touch"))
+		goto out;
+
+	if (run(skel->progs.big_free, "big_free"))
+		goto out;
+	ASSERT_EQ(wait_memlock(fd, base), base, "released");
+out:
+	typed_arena__destroy(skel);
+}
+
+/*
+ * A list linked through typed pointer fields is walked with no cast, read
+ * and written, and its NULL end dereferences object 0 of the slice.
+ */
+static void test_fields(void)
+{
+	struct typed_arena *skel;
+
+	skel = typed_arena__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "open_load"))
+		return;
+	if (run(skel->progs.list_build, "list_build"))
+		goto out;
+	if (run(skel->progs.list_sum_values, "list_sum"))
+		goto out;
+	ASSERT_EQ(skel->bss->list_sum, LIST_LEN * (LIST_LEN + 1) / 2, "sum");
+	if (run(skel->progs.list_bump_values, "list_bump"))
+		goto out;
+	if (run(skel->progs.list_sum_values, "list_sum_bumped"))
+		goto out;
+	ASSERT_EQ(skel->bss->list_sum, LIST_LEN * (LIST_LEN + 1) / 2 + LIST_LEN, "sum_bumped");
+
+	skel->bss->value = 4242;
+	if (run(skel->progs.list_deref_end, "list_deref_end"))
+		goto out;
+	ASSERT_EQ(skel->bss->list_sum, 4242, "null_is_object_zero");
+out:
+	typed_arena__destroy(skel);
+}
+
+void test_typed_arena(void)
+{
+	struct typed_arena *skel;
+	bool supported;
+
+	/* The programs need a compiler that emits the cast. */
+	skel = typed_arena__open();
+	if (!ASSERT_OK_PTR(skel, "open"))
+		return;
+	supported = skel->rodata->typed_arena_supported;
+	typed_arena__destroy(skel);
+	if (!supported) {
+		test__skip();
+		return;
+	}
+
+	if (test__start_subtest("pages"))
+		test_pages();
+	if (test__start_subtest("identity"))
+		test_identity();
+	if (test__start_subtest("release_spans"))
+		test_release_spans();
+	if (test__start_subtest("release_bounds"))
+		test_release_bounds();
+	if (test__start_subtest("release_once"))
+		test_release_once();
+	if (test__start_subtest("batch_alloc"))
+		test_batch_alloc();
+	if (test__start_subtest("multipage"))
+		test_multipage();
+	if (test__start_subtest("fields"))
+		test_fields();
+}
diff --git a/tools/testing/selftests/bpf/progs/typed_arena.c b/tools/testing/selftests/bpf/progs/typed_arena.c
new file mode 100644
index 000000000000..ed57ed4ee24b
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/typed_arena.c
@@ -0,0 +1,354 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include "bpf_experimental.h"
+#include "bpf_arena_common.h"
+#include "../test_kmods/bpf_testmod_kfunc.h"
+
+struct {
+	__uint(type, BPF_MAP_TYPE_ARENA);
+	__uint(map_flags, BPF_F_MMAPABLE);
+	__uint(max_entries, 8);
+#ifdef __TARGET_ARCH_arm64
+	__ulong(map_extra, 0x1ull << 32);
+#else
+	__ulong(map_extra, 0x1ull << 44);
+#endif
+} arena SEC(".maps");
+
+/* Tells the runner whether the compiler emits the cast, and so whether the programs below exist. */
+const volatile bool typed_arena_supported =
+#ifdef __BPF_FEATURE_TYPED_ARENA_CAST
+	true;
+#else
+	false;
+#endif
+
+/* 16-byte slots: a 256 KiB typed arena, 64 pages of 4 KiB, one word of the chunk bitmap */
+struct obj {
+	struct prog_test_ref_kfunc __kptr *ref;
+	__u64 value;
+} __typed_arena_size(256K);
+
+/* 16-byte slots: an 8 MiB typed arena, 2048 pages of 4 KiB, more than one batch of the allocator */
+struct wide_obj {
+	struct prog_test_ref_kfunc __kptr *ref;
+	__u64 value;
+} __typed_arena_size(8M);
+
+/* An object of 8 KiB, two pages of 4 KiB: its chunk is the object */
+struct big_obj {
+	struct prog_test_ref_kfunc __kptr *ref;
+	__u64 value;
+	char pad[8192 - 16];
+} __typed_arena_size(64K);
+
+/* 32-byte slots, linked through a typed pointer field */
+struct arena_list_node {
+	struct prog_test_ref_kfunc __kptr *ref;
+	struct arena_list_node *next;
+	__u64 value;
+} __typed_arena_size(64K);
+
+#define LIST_LEN 64
+/* Objects sit at slot strides, the power of two covering the struct, not at sizeof. */
+#define NODE_SLOT 32
+
+/* An object as an opaque 64-bit value user space hands back; any value casts to an object. */
+void *ptr;
+/* Pages requested; the allocator writes back what it granted. */
+__u32 page_cnt = 1;
+__u64 value;
+__u64 ref_cnt;
+void *list_head;
+__u64 list_sum;
+
+#ifdef __BPF_FEATURE_TYPED_ARENA_CAST
+
+/*
+ * A program is associated with an arena by referencing the map. Programs that
+ * only cast values they hold reference it explicitly.
+ */
+#define arena_bind() asm volatile("r0 = %[m] ll" :: [m] "i"(&arena) : "r0")
+
+#define TYPE_ID(T) bpf_core_type_id_local(T)
+
+SEC("syscall")
+int alloc(void *ctx)
+{
+	struct obj *o;
+
+	o = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct obj), NULL, &page_cnt, NUMA_NO_NODE);
+	if (!o)
+		return 1;
+	ptr = o;
+	return 0;
+}
+
+/* Take the chunks at ptr: 0 when granted, 1 when refused. */
+SEC("syscall")
+int alloc_at(void *ctx)
+{
+	struct obj *hint, *o;
+
+	hint = ptr;
+	o = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct obj), hint, &page_cnt, NUMA_NO_NODE);
+	if (!o)
+		return 1;
+	return o != hint;
+}
+
+SEC("syscall")
+int free_pages(void *ctx)
+{
+	struct obj *o = ptr;
+
+	bpf_typed_arena_free_pages(&arena, TYPE_ID(struct obj), o, page_cnt);
+	return 0;
+}
+
+SEC("syscall")
+int write_value(void *ctx)
+{
+	struct obj *o;
+
+	arena_bind();
+	o = ptr;
+	o->value = value;
+	return 0;
+}
+
+SEC("syscall")
+int read_value(void *ctx)
+{
+	struct obj *o;
+
+	arena_bind();
+	o = ptr;
+	value = o->value;
+	return 0;
+}
+
+/* Leave a reference to the test module's object in the kptr field. */
+SEC("syscall")
+int stash_ref(void *ctx)
+{
+	struct prog_test_ref_kfunc *p, *old;
+	unsigned long sp = 0;
+	struct obj *o;
+
+	arena_bind();
+	p = bpf_kfunc_call_test_acquire(&sp);
+	if (!p)
+		return 1;
+	o = ptr;
+	old = bpf_kptr_xchg(&o->ref, p);
+	if (old)
+		bpf_kfunc_call_test_release(old);
+	return 0;
+}
+
+SEC("syscall")
+int read_ref_cnt(void *ctx)
+{
+	struct prog_test_ref_kfunc *p;
+	unsigned long sp = 0;
+
+	p = bpf_kfunc_call_test_acquire(&sp);
+	if (!p)
+		return 1;
+	/* the acquire above holds one */
+	ref_cnt = p->cnt.refs.counter - 1;
+	bpf_kfunc_call_test_release(p);
+	return 0;
+}
+
+SEC("syscall")
+int wide_alloc(void *ctx)
+{
+	struct wide_obj *o;
+
+	o = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct wide_obj), NULL, &page_cnt,
+					NUMA_NO_NODE);
+	if (!o)
+		return 1;
+	ptr = o;
+	return 0;
+}
+
+/* Write the value into the object at ptr and read it back. */
+SEC("syscall")
+int wide_touch(void *ctx)
+{
+	struct wide_obj *o;
+
+	arena_bind();
+	o = ptr;
+	o->value = value;
+	return o->value != value;
+}
+
+SEC("syscall")
+int wide_free(void *ctx)
+{
+	struct wide_obj *o = ptr;
+
+	bpf_typed_arena_free_pages(&arena, TYPE_ID(struct wide_obj), o, page_cnt);
+	return 0;
+}
+
+SEC("syscall")
+int big_alloc(void *ctx)
+{
+	struct big_obj *o;
+
+	o = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct big_obj), NULL, &page_cnt,
+					NUMA_NO_NODE);
+	if (!o)
+		return 1;
+	ptr = o;
+	return 0;
+}
+
+/* Write both pages of the object at ptr and read them back. */
+SEC("syscall")
+int big_touch(void *ctx)
+{
+	struct big_obj *o;
+
+	arena_bind();
+	o = ptr;
+	o->value = value;
+	o->pad[sizeof(o->pad) - 1] = 1;
+	return o->value != value || o->pad[sizeof(o->pad) - 1] != 1;
+}
+
+SEC("syscall")
+int big_free(void *ctx)
+{
+	struct big_obj *o = ptr;
+
+	bpf_typed_arena_free_pages(&arena, TYPE_ID(struct big_obj), o, page_cnt);
+	return 0;
+}
+
+/* A page of nodes, the first LIST_LEN linked in order with values 1..LIST_LEN, ending in NULL. */
+SEC("syscall")
+int list_build(void *ctx)
+{
+	struct arena_list_node *n;
+	__u32 cnt = 1;
+	void *page;
+	int i;
+
+	page = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct arena_list_node), NULL, &cnt,
+					   NUMA_NO_NODE);
+	if (!page)
+		return 1;
+	for (i = 0; i < LIST_LEN; i++) {
+		n = page + i * NODE_SLOT;
+		n->value = i + 1;
+		if (i + 1 < LIST_LEN)
+			n->next = page + (i + 1) * NODE_SLOT;
+		else
+			n->next = NULL;
+	}
+	list_head = page;
+	return 0;
+}
+
+/*
+ * Walk the list through its typed pointer fields with no cast; the NULL that
+ * ends it is compared raw.
+ */
+SEC("syscall")
+int list_sum_values(void *ctx)
+{
+	struct arena_list_node *n;
+	__u64 sum = 0;
+	int i;
+
+	arena_bind();
+	n = list_head;
+	for (i = 0; n && i < LIST_LEN + 1; i++) {
+		sum += n->value;
+		n = n->next;
+	}
+	list_sum = sum;
+	return 0;
+}
+
+SEC("syscall")
+int list_bump_values(void *ctx)
+{
+	struct arena_list_node *n;
+	int i;
+
+	arena_bind();
+	n = list_head;
+	for (i = 0; n && i < LIST_LEN + 1; i++) {
+		n->value += 1;
+		n = n->next;
+	}
+	return 0;
+}
+
+/*
+ * Dereference the NULL that ends the list: it lands on object 0 of the node
+ * slice, marked with the value first through a NULL pointer held directly.
+ */
+SEC("syscall")
+int list_deref_end(void *ctx)
+{
+	struct arena_list_node *zero = NULL, *n, *m;
+	int i;
+
+	arena_bind();
+	/*
+	 * The compiler treats a NULL dereference as undefined and would drop
+	 * the store, the checks and the last hop; the barriers keep each
+	 * pointer opaque so that the store, the walk and the final dereference
+	 * are emitted.
+	 */
+	barrier_var(zero);
+	zero->value = value;
+	n = list_head;
+	for (i = 0; i < LIST_LEN; i++) {
+		m = n->next;
+		barrier_var(m);
+		if (!m)
+			break;
+		n = m;
+	}
+	m = n->next;
+	barrier_var(m);
+	list_sum = m->value;
+	return 0;
+}
+
+#else /* !__BPF_FEATURE_TYPED_ARENA_CAST */
+
+/* Keep the skeleton's shape without the compiler support; the runner skips the tests. */
+#define STUB(name) SEC("syscall") int name(void *ctx) { return 0; }
+STUB(alloc)
+STUB(alloc_at)
+STUB(free_pages)
+STUB(write_value)
+STUB(read_value)
+STUB(stash_ref)
+STUB(read_ref_cnt)
+STUB(wide_alloc)
+STUB(wide_touch)
+STUB(wide_free)
+STUB(big_alloc)
+STUB(big_touch)
+STUB(big_free)
+STUB(list_build)
+STUB(list_sum_values)
+STUB(list_bump_values)
+STUB(list_deref_end)
+
+#endif /* __BPF_FEATURE_TYPED_ARENA_CAST */
+
+char _license[] SEC("license") = "GPL";
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* Re: [RFC PATCH bpf-next v1 01/16] mm/vmalloc: Add get_vm_area_align()
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 01/16] mm/vmalloc: Add get_vm_area_align() Kumar Kartikeya Dwivedi
@ 2026-09-26 23:42   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-26 23:42 UTC (permalink / raw)
  To: Kumar Kartikeya Dwivedi; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] mm: get_vm_area_align() silently ignores align if VM_IOREMAP is passed
--

commit 60a153047292d8f40c9a75614c11c4a932d74c1f
Author: Kumar Kartikeya Dwivedi <memxor@gmail.com>

mm/vmalloc: Add get_vm_area_align()

This commit introduces get_vm_area_align(), a new variant of get_vm_area()
that allows callers to specify a custom alignment for the start of the
virtual memory area. This avoids wasting virtual address space when naturally
aligned regions are required, such as for BPF arena maps.

> diff --git a/mm/vmalloc.c b/mm/vmalloc.c
> index bea9f76ed7e74..b70e4e27ec362 100644
> --- a/mm/vmalloc.c
> +++ b/mm/vmalloc.c

[ ... ]

> @@ -3301,6 +3301,25 @@ struct vm_struct *get_vm_area_caller(unsigned long size, unsigned long flags,
>  				  NUMA_NO_NODE, GFP_KERNEL, caller);
>  }
>  
> +/**
> + * get_vm_area_align - reserve a contiguous kernel virtual area at an aligned start
> + * @size:	 size of the area
> + * @align:	 alignment of the start of the area, a power of two
> + * @flags:	 %VM_IOREMAP for I/O mappings or VM_ALLOC

[Severity: Medium]
Does this API work as intended if a caller passes VM_IOREMAP?

The documentation explicitly invites the caller to pass VM_IOREMAP, but
__get_vm_area_node() unconditionally overwrites the requested alignment
when this flag is set:

mm/vmalloc.c:__get_vm_area_node() {
    ...
    if (flags & VM_IOREMAP)
        align = 1ul << clamp_t(int, get_count_order_long(size),
                               PAGE_SHIFT, IOREMAP_MAX_ORDER);
    ...
}

Will this establish a broken contract where the caller's explicitly
requested alignment is silently ignored and overridden based on the
size of the area?

> + *
> + * Like get_vm_area(), with the start of the area aligned to @align.
> + *
> + * Return: the area descriptor on success or %NULL on failure.
> + */
> +struct vm_struct *get_vm_area_align(unsigned long size, unsigned long align,
> +				    unsigned long flags)
> +{
> +	return __get_vm_area_node(size, align, PAGE_SHIFT, flags,
> +				  VMALLOC_START, VMALLOC_END,
> +				  NUMA_NO_NODE, GFP_KERNEL,
> +				  __builtin_return_address(0));
> +}
> +

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926233503.3114147-1-memxor@gmail.com?part=1

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [RFC PATCH bpf-next v1 12/16] selftests/bpf: Build BPF objects with compiler-inserted typed arena casts
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 12/16] selftests/bpf: Build BPF objects with compiler-inserted typed arena casts Kumar Kartikeya Dwivedi
@ 2026-09-26 23:46   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-26 23:46 UTC (permalink / raw)
  To: Kumar Kartikeya Dwivedi; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] bpf: missing typed-arena feature flag propagation to libarena sub-make
--

commit b75aa4c5a8315082b286ae6d14b710be82785867
Author: Kumar Kartikeya Dwivedi <memxor@gmail.com>

selftests/bpf: Build BPF objects with compiler-inserted typed arena casts

This commit probes clang for typed arena cast support and defines
CLANG_TYPED_ARENA if available. It then appends this new flag to BPF_CC_FLAGS
when building test_progs and its various flavors.

> diff --git a/tools/testing/selftests/bpf/Makefile b/tools/testing/selftests/bpf/Makefile
> index afa589a27b151..bd1f84f55edc5 100644
> --- a/tools/testing/selftests/bpf/Makefile
> +++ b/tools/testing/selftests/bpf/Makefile

[ ... ]

> @@ -510,7 +517,7 @@ BINARY := test_progs
>  BPF_CC := $(CLANG)
>  BPF_CC_MSG := CLNG-BPF
>  BPF_SYS_INCLUDES := $(CLANG_SYS_INCLUDES)
> -BPF_CC_FLAGS := -O2 $(BPF_TARGET_ENDIAN) -mcpu=v3
> +BPF_CC_FLAGS := -O2 $(BPF_TARGET_ENDIAN) -mcpu=v3 $(CLANG_TYPED_ARENA)

[Severity: Low]
Should the newly introduced CLANG_TYPED_ARENA flag also be propagated to the
libarena sub-make?

When building libarena, BPF compilation flags are explicitly passed down
without the new compiler feature flag:

LIBARENA_MAKE_ARGS = \
...
		CLANG="$(CLANG)" \
		BPF_CFLAGS="$(BPF_CFLAGS) $(CLANG_CFLAGS)" \
...

While there are currently no typed arena tests within libarena, omitting this
flag means any future typed arena programs added to libarena might fail to
compile or get unexpectedly skipped. Does it make sense to include it here
for consistency?

>  BPF_DEFINES := -DENABLE_ATOMICS_TESTS
>  include Makefile.skel

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926233503.3114147-1-memxor@gmail.com?part=12

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [RFC PATCH bpf-next v1 15/16] selftests/bpf: Test typed arena page allocation and release
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 15/16] selftests/bpf: Test typed arena page allocation and release Kumar Kartikeya Dwivedi
@ 2026-09-26 23:46   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-26 23:46 UTC (permalink / raw)
  To: Kumar Kartikeya Dwivedi; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Critical] bpf: physical pages freed before TLB flush in typed_arena_alloc_pages()
- [High] bpf: bpf_typed_arena_free_pages() frees adjacent chunks due to unaligned addr
--

commit 45f79d175e13ad1f2540fd16c264cce33991c839
Author: Kumar Kartikeya Dwivedi <memxor@gmail.com>

selftests/bpf: Test typed arena page allocation and release

Allocate a page of typed arena objects, write one, reach it again through
an opaque copy of its address and its neighbour by arithmetic on that copy,
and check that the kfunc receives the registered typed arena in place of
the type ID. Take the chunk a value names and see the same chunk refused
twice, once for the chunk and once for a page inside it. Ask for one page
of an object that spans several and see the request rounded up to the
object's chunk, with the granted count written back. Release a chunk and
see it keep its objects and stay taken within the same invocation, since
the release is deferred behind a grace period.

Reject a use of the returned pointer without a NULL check, a struct without
special fields, and a map that is not the program's arena. Declare the two
kfuncs in bpf_experimental.h.

> diff --git a/tools/testing/selftests/bpf/progs/verifier_typed_arena.c b/tools/testing/selftests/bpf/progs/verifier_typed_arena.c
> index 808a65611d4ef..8ec75fd1f97bf 100644
> --- a/tools/testing/selftests/bpf/progs/verifier_typed_arena.c
> +++ b/tools/testing/selftests/bpf/progs/verifier_typed_arena.c
> @@ -1157,6 +1157,148 @@ int ptr_field_sanitize_on_one_path(void *ctx)
>  	return 0;
>  }
>  
> +#define TYPE_ID(T) bpf_core_type_id_local(T)
> +
> +SEC("syscall")
> +__description("allocated pages hold real objects, reachable through any value that lands in them")
> +__success __retval(0)
> +__xlated("r2 = 0x{{[0-9a-f]+[0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f][0-9a-f]}}")
> +__xlated("call kernel-function")
> +int pages_alloc(void *ctx)
> +{
> +	struct typed_obj *obj, *again, *next;
> +	void *opaque;
> +	__u32 cnt = 1;
> +
> +	obj = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct typed_obj), NULL, &cnt,
> +					  NUMA_NO_NODE);

[Severity: Critical]
Is it possible for physical pages to be freed before the TLB flush occurs?

Looking at typed_arena_alloc_pages() in kernel/bpf/arena.c, when an
allocation fails midway:

kernel/bpf/arena.c:typed_arena_alloc_pages() {
    ...
			if (data.i)
				apply_to_existing_page_range(&init_mm, data.start,
							     (unsigned long)data.i << PAGE_SHIFT,
							     apply_range_clear_typed_cb, &cdata);
			free_pages_nolock(head, order);
			goto back_out;
    ...
}

If a concurrent page fault occurs, can this result in the physical pages being
returned to the allocator before the TLB flush in the back_out path?

> +	if (!obj)
> +		return 1;

[ ... ]

> +SEC("syscall")
> +__description("a released chunk keeps its objects and stays taken until the grace period has passed")
> +__success __retval(0)
> +int pages_free(void *ctx)
> +{
> +	struct typed_obj *obj;
> +	__u32 cnt = 1;
> +
> +	obj = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct typed_obj), NULL, &cnt,
> +					  NUMA_NO_NODE);
> +	if (!obj)
> +		return 1;
> +	obj->value = 7;
> +	bpf_typed_arena_free_pages(&arena, TYPE_ID(struct typed_obj), obj, 1);

[Severity: High]
Does this code handle unaligned offset values correctly? 

Looking at typed_arena_free_pages() in kernel/bpf/arena.c, when an unaligned
addr is passed, the calculation for the last chunk might spill over into the
next chunk:

kernel/bpf/arena.c:typed_arena_free_pages() {
	off = addr - base;
	if (page_cnt > (size - off) >> PAGE_SHIFT)
		return;
	first = off >> ta->chunk_shift;
	last = (off + ((unsigned long)page_cnt << PAGE_SHIFT) - 1) >> ta->chunk_shift;
    ...
}

Can this lead to unintentional freeing of adjacent chunks?

> +	if (obj->value != 7)
> +		return 2;
> +	cnt = 1;
> +	if (bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct typed_obj), obj, &cnt, NUMA_NO_NODE))
> +		return 3;
> +	return 0;
> +}

[ ... ]

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926233503.3114147-1-memxor@gmail.com?part=15

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [RFC PATCH bpf-next v1 10/16] bpf: Let typed_arena_cast copy pointers the verifier already trusts
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 10/16] bpf: Let typed_arena_cast copy pointers the verifier already trusts Kumar Kartikeya Dwivedi
@ 2026-09-26 23:49   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-26 23:49 UTC (permalink / raw)
  To: Kumar Kartikeya Dwivedi; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: check_typed_arena_cast() misses src->off check causing state mismatch
--

commit e8c7f600f2244019888ad241fa81b63bc3bccf3d
Author: Kumar Kartikeya Dwivedi <memxor@gmail.com>

bpf: Let typed_arena_cast copy pointers the verifier already trusts

The compiler inserts the cast at every use of a pointer to a struct with
special fields, so that programs never write one, and it cannot tell what
the pointer is: the same struct lives in typed arenas, in allocated objects,
in map values and on the stack. So the verifier decides per path what the
cast takes.

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index b85d99f81ef54..aceb17e629481 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
[ ... ]
> @@ -17320,6 +17358,18 @@ static int check_typed_arena_cast(struct bpf_verifier_env *env, struct bpf_insn
>  		return PTR_ERR(ta);
>  	aux->typed_arena = ta;
>  
> +	if (src->type == (PTR_TO_BTF_ID | MEM_ARENA) && src->btf_id == ta->btf_id &&
> +	    tnum_is_const(src->var_off) && !src->var_off.value) {

[Severity: High]
Does check_typed_arena_cast() need to ensure src->off == 0 here?

If a BPF program uses typed_arena_cast on a control flow path where one branch
provides a SCALAR_VALUE (setting aux->sanitize_needed = true) and the other
branch provides a PTR_TO_BTF_ID | MEM_ARENA with a non-zero fixed offset
(src->off > 0), the runtime sanitization sequence drops the offset by masking
it to 0, but the verifier retains the non-zero offset.

Could this state mismatch allow a program to subtract the offset and access
memory before the arena object, bypassing verifier bounds checks?

> +		*dst = *src;
> +		return 0;
> +	}
> +	if (aux->sanitize_plain) {
> +		verbose(env, "insn %d casts values that need different treatment on different paths\n",
> +			env->insn_idx);
> +		return -EINVAL;
> +	}
> +	aux->sanitize_needed = true;
> +

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926233503.3114147-1-memxor@gmail.com?part=10

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [RFC PATCH bpf-next v1 16/16] selftests/bpf: Exercise typed arenas at run time
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 16/16] selftests/bpf: Exercise typed arenas at run time Kumar Kartikeya Dwivedi
@ 2026-09-26 23:50   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-26 23:50 UTC (permalink / raw)
  To: Kumar Kartikeya Dwivedi; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] selftests/bpf: test_release_bounds unconditionally fails due to logic error in alloc_at expectation
--

commit a5c65698dd674c11891a999b72347849ae68ce84
Author: Kumar Kartikeya Dwivedi <memxor@gmail.com>

selftests/bpf: Exercise typed arenas at run time

Drive typed arenas from user space through a skeleton, testing map memory
usage, chunk allocation, release, multipage allocations, bounds checking,
and pointer fields.

> diff --git a/tools/testing/selftests/bpf/prog_tests/typed_arena.c b/tools/testing/selftests/bpf/prog_tests/typed_arena.c
> new file mode 100644
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/prog_tests/typed_arena.c

[ ... ]

> +static void test_release_bounds(void)
> +{
> +	long ps = PAGE_SIZE, base, nr = OBJ_ARENA_SIZE / PAGE_SIZE;
> +	struct typed_arena *skel;
> +	int fd;
> +
> +	skel = typed_arena__open_and_load();
> +	if (!ASSERT_OK_PTR(skel, "open_load"))
> +		return;
> +	fd = bpf_map__fd(skel->maps.arena);
> +	base = map_memlock(fd);
> +
> +	/* Fill the typed arena, so that a scan of its bitmap for a free chunk runs to the end. */
> +	skel->bss->ptr = NULL;
> +	skel->data->page_cnt = nr;
> +	if (run(skel->progs.alloc_at, "fill"))
> +		goto out;

[Severity: Medium]
Will this unconditionally abort the test?

When skel->bss->ptr is set to NULL in test_release_bounds(), the BPF program
alloc_at() executes with its hint variable as NULL. If the allocation is
successful, o will be non-NULL, which makes the condition o != hint evaluate
to 1 (true):

tools/testing/selftests/bpf/progs/typed_arena.c:alloc_at() {
    ...
	hint = ptr;
	o = bpf_typed_arena_alloc_pages(&arena, TYPE_ID(struct obj), hint, &page_cnt, NUMA_NO_NODE);
	if (!o)
		return 1;
	return o != hint;
}

The run() helper wraps ASSERT_OK(ret), which expects a return value of 0.
Since it receives 1, the assertion fails and the test executes goto out,
leaving the rest of test_release_bounds() unexecuted.

Could the return logic in alloc_at() be adjusted when hint is NULL, or
should this test use a different program or helper to fill the arena?

> +	ASSERT_EQ(map_memlock(fd), base + nr * ps, "full");
> +
> +	skel->data->page_cnt = nr + 1;
> +	if (run(skel->progs.free_pages, "free_oversized"))
> +		goto out;

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926233503.3114147-1-memxor@gmail.com?part=16

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [RFC PATCH bpf-next v1 09/16] bpf: Add typed arena page allocation and release kfuncs
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 09/16] bpf: Add typed arena page allocation and release kfuncs Kumar Kartikeya Dwivedi
@ 2026-09-26 23:55   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-26 23:55 UTC (permalink / raw)
  To: Kumar Kartikeya Dwivedi; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 7 potential issue(s) to consider:
- [Critical] bpf: stale TLB use-after-free window in allocation error path
- [High] bpf: uninitialized kernel memory leak via arena page allocation
- [High] bpf: TLB shootdown race due to premature state publication
- [High] bpf: TOCTOU vulnerability via double fetch of user-provided page_cnt
- [High] bpf: state machine corruption race between worker and page fault handler
- [Medium] bpf: memory leak when resilient spinlock acquisition fails
- [Medium] bpf: incorrect memcg charging in tracking struct allocation
--

commit fcd33d7dbae3719eae8ea173b7e1a7f8cbe911c9
Author: Kumar Kartikeya Dwivedi <memxor@gmail.com>

bpf: Add typed arena page allocation and release kfuncs

This patch introduces typed arenas for BPF programs, allowing them to
allocate and release contiguous memory regions mapped as real objects
rather than scratch pages. It exposes two new kfuncs, 
bpf_typed_arena_alloc_pages() and bpf_typed_arena_free_pages(), which 
interface with a lock-free page allocator and defer the release via RCU
workers.

diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -686,6 +1045,8 @@ static unsigned long typed_arena_alloc_pages(struct bpf_typed_arena *ta, unsigne
 					     u32 *page_cnt, int node_id)
 {
[ ... ]
> +	if (node_id != NUMA_NO_NODE &&
> +	    ((unsigned int)node_id >= nr_node_ids || !node_online(node_id)))
> +		return 0;
> +	if (!*page_cnt)
> +		return 0;
> +	chunk_cnt = DIV_ROUND_UP((unsigned long)*page_cnt, typed_arena_chunk_pages(ta));

[Severity: High]
Does this code introduce a TOCTOU vulnerability? Because page_cnt is a pointer
to memory that can be shared with BPF programs or userspace, it could be
modified concurrently. If the value is changed to zero between the
if (!*page_cnt) check and the DIV_ROUND_UP() calculation, chunk_cnt will
evaluate to 0. This would cause bitmap_find_next_zero_area() to return 0 and
grant the BPF program a valid pointer to the arena's base address without
actually allocating any pages, potentially leading to memory corruption.

[ ... ]
> +	data.arena = arena;
> +	for (done = 0; done < chunk_cnt; done++) {
> +		head = alloc_pages_nolock(__GFP_ACCOUNT, node_id, order);
> +		if (!head) {

[Severity: High]
Does this allocation leak uninitialized buddy allocator memory?
The call to alloc_pages_nolock() omits the __GFP_ZERO flag. Since arena
memory is exposed to BPF programs (and potentially mapped to userspace),
untrusted code could read stale kernel data. Furthermore, during teardown,
the kernel might invoke bpf_obj_free_fields() on this uninitialized garbage,
which could lead to panics if random memory is treated as valid kptrs.

> +			ret = -ENOMEM;
> +			goto back_out;
> +		}
> +		data.head = head;
> +		data.start = start + (done << ta->chunk_shift);
> +		data.i = 0;
> +		ret = apply_to_page_range(&init_mm, data.start, chunk, apply_range_set_typed_cb,
> +					  &data);
> +		if (ret) {
> +			/* The fault path took this chunk: give back the allocation whole. */
> +			cdata.arena = arena;
> +			cdata.ta = ta;
> +			if (data.i)
> +				apply_to_existing_page_range(&init_mm, data.start,
> +							     (unsigned long)data.i << PAGE_SHIFT,
> +							     apply_range_clear_typed_cb, &cdata);
> +			free_pages_nolock(head, order);
> +			goto back_out;
> +		}

[Severity: Critical]
Does this code clear PTEs and free physical pages before flushing the TLB?
When apply_to_page_range() returns -EBUSY, the partially installed PTEs are
cleared and the pages are immediately freed to the buddy allocator via
free_pages_nolock(). The actual TLB flush happens later at the back_out
label. This creates a window where freed pages can be reallocated
elsewhere in the kernel, while stale, writable VMA translations remain cached
in the TLB, allowing arbitrary memory overwrite.

[ ... ]
> +back_out:
> +	/*
> +	 * Back out the chunks installed before the one that failed. On -EBUSY
> +	 * that chunk is the one the fault path took, and it keeps its mark.
> +	 */
> +	init_llist_head(&freed);
> +	cdata.arena = arena;
> +	cdata.ta = ta;
> +	for (i = 0; i < chunk_cnt; i++)
> +		if (ret != -EBUSY || i != done)
> +			clear_bit(coff + i, ta->chunks);
> +	for (i = 0; i < done; i++) {
> +		cdata.head = NULL;
> +		apply_to_existing_page_range(&init_mm, start + (i << ta->chunk_shift), chunk,
> +					     apply_range_clear_typed_cb, &cdata);
> +		if (cdata.head)
> +			__llist_add(&cdata.head->pcp_llist, &freed);
> +	}
> +	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> +	flush_tlb_kernel_range(start, start + (chunk_cnt << ta->chunk_shift));

[Severity: High]
Does this sequence create a TLB shootdown race? By clearing the allocation
state bits (ta->chunks) and dropping the spinlock before calling
flush_tlb_kernel_range(), it creates a window where concurrent allocations can
immediately claim the chunk and expose new mappings. Other CPUs might retain
stale TLB translations to the old pages that are about to be freed, leading to
silent data corruption or use-after-free conditions.

[ ... ]
> +static void typed_arena_free_pages(struct bpf_typed_arena *ta, unsigned long addr, u32 page_cnt)
> +{
[ ... ]
> +	first = off >> ta->chunk_shift;
> +	last = (off + ((unsigned long)page_cnt << PAGE_SHIFT) - 1) >> ta->chunk_shift;
> +
> +	s = kmalloc_nolock(sizeof(*s), __GFP_ACCOUNT, NUMA_NO_NODE);

[Severity: Medium]
Does this structure allocation incorrectly charge the current context's memory
cgroup? Because this is a generic kfunc that can be invoked from arbitrary
execution contexts (like tracepoints or network packets), __GFP_ACCOUNT will
charge whichever random user task happens to be interrupted. Shouldn't this be
wrapped with bpf_map_memcg_enter() and bpf_map_memcg_exit() to charge the
map's memcg?

> +	if (!s)
> +		/*
> +		 * The chunks stay allocated until the map is freed; nothing can
> +		 * be retried from here.
> +		 */
> +		return;
> +	if (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
> +		goto free_span;
[ ... ]
> +	irq_work_queue(&arena->typed_free_irq);
> +	return;
> +
> +free_span:
> +	kfree_nolock(s);
> +}

[Severity: Medium]
Does failing to acquire the resilient spinlock cause a silent memory leak?
If raw_res_spin_lock_irqsave() fails (e.g., due to lock contention), the
tracking struct is freed and the function returns without error. The requested
chunks remain marked as taken indefinitely, but the BPF program assumes they
were successfully queued for release, leaking arena chunks until map teardown.

[ ... ]
> +static void typed_arena_free_worker(struct work_struct *work)
> +{
[ ... ]
> +		while (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
> +			cpu_relax();
> +		for (i = 0; i < s->chunk_cnt; i++) {
> +			cdata.head = NULL;
> +			apply_to_existing_page_range(&init_mm, start + (i << ta->chunk_shift),
> +						     bpf_typed_arena_chunk(ta),
> +						     apply_range_clear_typed_cb, &cdata);
> +			if (cdata.head)
> +				__llist_add(&cdata.head->pcp_llist, &heads);
> +		}
> +		typed_arena_chunks_mark(ta, s->coff, s->chunk_cnt, false);

[Severity: High]
Does this create a destructive state machine race with the page fault handler?
The worker clears the PTE mappings under the spinlock, but the page fault
handler (typed_arena_handle_page_fault()) runs locklessly. If a page fault
triggers on a cleared chunk just before typed_arena_chunks_mark(..., false) is
called, the fault handler will install a scratch page and set the chunk bit.
The worker will then blindly force the bit to 0. The chunk becomes marked as
'free' but contains a scratch page, causing all future allocations of this
chunk to permanently fail with -EBUSY.

> +		bitmap_clear(ta->pending, s->coff, s->chunk_cnt);
> +		raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> +
> +		flush_tlb_kernel_range(start,
> +				       start + ((unsigned long)s->chunk_cnt << ta->chunk_shift));

[Severity: High]
Can this sequence lead to a TLB shootdown race? Similar to the allocation
error path, clearing the allocation state bits and dropping the spinlock
before flushing the TLB allows concurrent allocations to claim the chunks
while stale translations remain cached, introducing memory corruption risks.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926233503.3114147-1-memxor@gmail.com?part=9

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [RFC PATCH bpf-next v1 04/16] bpf: Add the typed_arena_cast instruction
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 04/16] bpf: Add the typed_arena_cast instruction Kumar Kartikeya Dwivedi
@ 2026-09-26 23:55   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-26 23:55 UTC (permalink / raw)
  To: Kumar Kartikeya Dwivedi; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] bpf: typed arenas leak on late verification failure
- [Medium] bpf: diagnostic log state leak in check_alu_op() on typed arena cast
- [Low] bpf: comment style violation in backtrack_insn()
--

commit 1039fff673a440b6fdcfc452a1630f899213bc0c
Author: Kumar Kartikeya Dwivedi <memxor@gmail.com>

bpf: Add the typed_arena_cast instruction

This commit adds the typed_arena_cast instruction, which turns an untrusted
64-bit value into a verifier-trusted pointer to a typed arena object. The
verifier registers the typed arena on first sight, logs it, and lowers the
instruction after verification.

> diff --git a/kernel/bpf/backtrack.c b/kernel/bpf/backtrack.c
> index 0e38b9575328c..38984e52ee375 100644
> --- a/kernel/bpf/backtrack.c
> +++ b/kernel/bpf/backtrack.c
> @@ -328,11 +328,13 @@ static int backtrack_insn(struct bpf_verifier_env *env, int idx, int subseq_idx,
>  				if (sreg != BPF_REG_FP)
>  					bt_set_reg(bt, sreg);
>  			} else {
> -				/* dreg = K
> +				/* dreg = K, or dreg = typed_arena_cast(sreg, imm)
>  				 * dreg needs precision after this insn.
>  				 * Corresponding register is already marked
>  				 * as precise=true in this verifier state.
> -				 * No further markings in parent are necessary
> +				 * No further markings in parent are necessary;
> +				 * a cast yields a pointer that is safe for any
> +				 * value of sreg, which needs no precision.
>  				 */
                                 ^^^

[Severity: Low]
Does this multi-line comment violate the BPF subsystem comment formatting rules?
The opening /* should typically be on its own line for multi-line comments in
BPF subsystem code.

[ ... ]

> diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c
> index d3b8b626ec0f7..619f7c6a778d9 100644
> --- a/kernel/bpf/core.c
> +++ b/kernel/bpf/core.c
> @@ -3085,6 +3085,8 @@ static void bpf_prog_free_deferred(struct work_struct *work)
>  	aux = container_of(work, struct bpf_prog_aux, work);
>  #ifdef CONFIG_BPF_SYSCALL
>  	bpf_free_kfunc_btf_tab(aux->kfunc_btf_tab);
> +	/* The typed arenas outlive the program; only the load's failure retracts them. */
> +	kfree(aux->typed_arenas);
>  #endif
>  #ifdef CONFIG_CGROUP_BPF
>  	if (aux->cgroup_atype != CGROUP_BPF_ATTACH_TYPE_INVALID)

[ ... ]

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 03dbc0e003986..a0069983f1036 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
> @@ -16993,7 +17173,9 @@ static int check_alu_op(struct bpf_verifier_env *env, struct bpf_insn *insn)
>  			struct bpf_reg_state *dst_reg = regs + insn->dst_reg;
>  
>  			if (BPF_CLASS(insn->code) == BPF_ALU64) {
> -				if (insn->imm) {
> +				if (insn->off == BPF_TYPED_ARENA_CAST) {
> +					return check_typed_arena_cast(env, insn);
                                        ^^^^

[Severity: Medium]
Does this early return bypass the paired bpf_diag_mod_end(env) cleanup at the
end of check_alu_op()? Leaving the diagnostic state active here fails to record
the register modification in the diagnostic event log.

> +				} else if (insn->imm) {
>  					/* off == BPF_ADDR_SPACE_CAST */
>  					mark_reg_unknown(env, regs, insn->dst_reg);
>  					if (insn->imm == 1) /* cast from as(1) to as(0) */

[ ... ]

> @@ -20580,8 +20764,26 @@ static int resolve_func_ptrs(struct bpf_verifier_env *env)
>  }
>  
>  /* drop refcnt of maps used by the rejected program */
> +/*
> + * Drop the rejected program's typed arena references before its arena map
> + * reference. A typed arena nobody else registered is retracted; one that a
> + * loaded program registered lives on for the map's lifetime.
> + */
> +static void release_typed_arenas(struct bpf_verifier_env *env)
> +{
> +	struct bpf_prog_aux *aux = env->prog->aux;
> +	u32 i;
> +
> +	for (i = 0; i < aux->typed_arena_cnt; i++)
> +		bpf_typed_arena_put(bpf_prog_arena(env->prog), aux->typed_arenas[i]);
> +	kfree(aux->typed_arenas);
> +	aux->typed_arenas = NULL;
> +	aux->typed_arena_cnt = 0;
> +}
> +
>  static void release_maps(struct bpf_verifier_env *env)
>  {
> +	release_typed_arenas(env);
        ^^^^

[Severity: High]
Can this leak the typed arena references?

In bpf_check(), if verification fails after env->prog->aux->used_maps is
populated, the err_release_maps path bypasses calling release_maps(env):

err_release_maps:
	if (ret)
		release_insn_arrays(env);
	if (!env->prog->aux->used_maps)
		release_maps(env);

If release_maps(env) is skipped, release_typed_arenas(env) will not be called.
Then, bpf_prog_free_deferred() only does kfree(aux->typed_arenas) without
dropping the actual bpf_typed_arena_put() refcounts, permanently leaking
the typed arenas and their backing resources.

>  	__bpf_free_used_maps(env->prog->aux, env->used_maps,
>  			     env->used_map_cnt);
>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926233503.3114147-1-memxor@gmail.com?part=4

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [RFC PATCH bpf-next v1 03/16] bpf: Back typed arena chunks with scratch on demand
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 03/16] bpf: Back typed arena chunks with scratch on demand Kumar Kartikeya Dwivedi
@ 2026-09-26 23:56   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-26 23:56 UTC (permalink / raw)
  To: Kumar Kartikeya Dwivedi; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Critical] bpf: Use-After-Free due to missing PTE cleanup in typed_arena_teardown_cb()
--

commit 5ad8213f9e65607de0b6ea3647cc93fae67f349a
Author: Kumar Kartikeya Dwivedi <memxor@gmail.com>

bpf: Back typed arena chunks with scratch on demand

This patch introduces demand-paging for typed BPF arenas, where unbacked
pages fault and are populated with scratch memory. It handles the fault
recovery, marks chunks in a bitmap, and adds allocation and freeing
mechanisms for these chunks during the map's lifecycle.

> diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
> index 9cd1b1ce434c4..c757c19323602 100644
> --- a/kernel/bpf/arena.c
> +++ b/kernel/bpf/arena.c

[ ... ]

> @@ -371,7 +396,7 @@ static int typed_arena_teardown_cb(pte_t *ptep, unsigned long addr, void *data)
>  
>  	if ((addr - (unsigned long)ta->base) & (bpf_typed_arena_chunk(ta) - 1))
>  		return 0;
> -	if (!pte_present(pte))
> +	if (!pte_present(pte) || pte_page(pte) == typed_arena_scratch_page(ta, addr))
>  		return 0;
>  	typed_arena_free_objects(ta, (void *)addr);
>  	__free_pages(pte_page(pte), typed_arena_chunk_order(ta));

[Severity: Critical]
Does this leave the kernel leaf PTEs pointing to freed memory?

This callback frees the physical pages for allocated chunks using
__free_pages(), but it doesn't appear to clear the PTEs using pte_clear()
or ptep_get_and_clear().

If a typed arena is destroyed (for example, if a BPF program load fails
after allocation), bpf_typed_arena_put() calls typed_arena_free(), which
then uses this callback. The underlying memory is freed, but the page
tables remain populated with valid mappings.

Because typed_arena_find_slice() searches for and reuses freed address
gaps, a subsequently loaded BPF program could allocate an arena of the
same size, inherit the same virtual address slice, and access these stale
PTEs to manipulate the freed memory.

> @@ -383,6 +408,10 @@ static void typed_arena_free(struct bpf_arena *arena, struct bpf_typed_arena *ta
>  	WRITE_ONCE(arena->typed_arena_mem, arena->typed_arena_mem - typed_arena_static_mem(ta));
>  	apply_to_existing_page_range(&init_mm, (unsigned long)ta->base, bpf_typed_arena_size(ta),
>  				     typed_arena_teardown_cb, ta);
> +	/* Programs may have stored kptrs into the dummy objects. */
> +	typed_arena_free_objects(ta, ta->scratch);
> +	vfree(ta->scratch);

[Severity: Critical]
Could this similarly leave stale mappings to freed scratch space?

The scratch pages are freed via vfree(), but the aliased PTEs within the
arena's kernel VM area that point to these scratch pages are never cleared
during the teardown process.

> +	kfree(ta->scratch_pages);
>  	bitmap_free(ta->chunks);
>  	bitmap_free(ta->pending);
>  	btf_put(ta->btf);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926233503.3114147-1-memxor@gmail.com?part=3

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [RFC PATCH bpf-next v1 06/16] bpf: Support special fields in typed arena objects
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 06/16] bpf: Support special fields in " Kumar Kartikeya Dwivedi
@ 2026-09-26 23:59   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-26 23:59 UTC (permalink / raw)
  To: Kumar Kartikeya Dwivedi; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Critical] bpf: bpf_kptr_xchg allows user-space to forge kernel pointers via typed arena memory
- [High] bpf: bpf_kptr_xchg on arena pointers causes kernel panic on archs lacking inline support
--

commit 2d0dbe5c6fd18e1e3096285d4ab48687f876b4ce
Author: Kumar Kartikeya Dwivedi <memxor@gmail.com>

bpf: Support special fields in typed arena objects

This commit allows bpf_kptr_xchg() to take a pointer to a kptr field of a
typed arena object. It modifies the verifier to permit objects in arena
memory to own references to kernel objects and program-allocated objects.

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 75697e52a2df6..f854d8419fff0 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
[ ... ]
> @@ -8063,7 +8063,7 @@ static int process_kptr_func(struct bpf_verifier_env *env, int regno,
>  	struct btf_record *rec;
>  	u32 kptr_off;
>  
> -	if (type_is_ptr_alloc_obj(reg->type)) {
> +	if (type_is_ptr_alloc_obj(reg->type) || type_is_typed_arena_obj(reg->type)) {
>  		rec = reg_btf_record(reg);

[Severity: Critical]
Does this allow user-space to forge kernel pointers?

Because arena memory (BPF_MAP_TYPE_ARENA) is inherently shared between the
kernel and user-space, it can be mmapped with PROT_READ and PROT_WRITE.
Furthermore, arena maps lack a map-level btf_record (it is NULL), which
bypasses the standard check in bpf_map_mmap() that forbids mmapping maps
containing special fields.

Could a malicious user-space process mmap the arena, write an arbitrary
64-bit value to the offset of the kptr field, and then have the BPF program
call bpf_kptr_xchg() to read this user-supplied value? The verifier would
unconditionally track the returned value as a trusted kernel pointer.

Could this forged pointer then be dereferenced or passed to bpf_obj_drop(),
leading to memory corruption or KASLR bypass?

>  	} else { /* PTR_TO_MAP_VALUE */
>  		map_ptr = reg->map_ptr;
[ ... ]
> @@ -8834,6 +8834,7 @@ static const struct bpf_reg_types kptr_xchg_dest_types = {
>  		PTR_TO_BTF_ID | MEM_ALLOC,
>  		PTR_TO_BTF_ID | MEM_ALLOC | NON_OWN_REF,
>  		PTR_TO_BTF_ID | MEM_ALLOC | NON_OWN_REF | MEM_RCU,
> +		PTR_TO_BTF_ID | MEM_ARENA,

[Severity: High]
Could this cause a kernel panic on architectures that lack inline support
for ptr_xchg?

On architectures that support arenas but do not override
bpf_jit_supports_ptr_xchg() to return true (e.g., PowerPC and s390x), the JIT
will fall back to calling the bpf_kptr_xchg() C kernel helper.

The bpf_kptr_xchg() helper uses the raw xchg() macro on the provided pointer,
which resides in demand-paged arena memory:

kernel/bpf/helpers.c:bpf_kptr_xchg() {
    ...
    return xchg(kptr, (unsigned long)ptr);
}

If the accessed arena page is unmapped, a kernel page fault occurs. Because
bpf_arena_handle_page_fault() is only wired up in the architecture-specific
fault handlers for arm64 and x86, it is completely ignored on PowerPC and
s390x.

Since the xchg() macro lacks an exception table entry inside the helper,
wouldn't an unmapped arena page result in an unhandled fault and a direct
kernel panic?

>  	}
>  };

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926233503.3114147-1-memxor@gmail.com?part=6

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [RFC PATCH bpf-next v1 07/16] bpf: Trust typed pointer fields of typed arena objects
  2026-09-26 23:34 ` [RFC PATCH bpf-next v1 07/16] bpf: Trust typed pointer fields of " Kumar Kartikeya Dwivedi
@ 2026-09-27  0:03   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-27  0:03 UTC (permalink / raw)
  To: Kumar Kartikeya Dwivedi; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] bpf: refcount leak in bpf_typed_arena causes saturation and WARN
--

commit 824870737e5e6082cced029879731300754bd906
Author: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Subject: bpf: Trust typed pointer fields of typed arena objects

This commit introduces logic to track and trust typed pointer fields
within typed arena objects. A load from these fields returns an
unsanitized pointer which is considered trustworthy without additional
checks, and writes ensure only valid typed arena pointers or null are
stored.

> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
[ ... ]
> @@ -6328,6 +6348,61 @@ static bool type_is_trusted_or_null(struct bpf_verifier_env *env,
[ ... ]
> +static int check_typed_ptr_field_access(struct bpf_verifier_env *env, struct bpf_reg_state *regs,
> +					struct bpf_reg_state *reg, const char *tname, u32 btf_id,
> +					enum bpf_access_type atype, int value_regno)
> +{
> +	struct bpf_insn *insn = &env->prog->insnsi[env->insn_idx];
> +	u8 class = BPF_CLASS(insn->code);
> +	struct bpf_reg_state *val;
> +	struct bpf_typed_arena *ta;
> +
> +	if (BPF_MODE(insn->code) != BPF_MEM || BPF_SIZE(insn->code) != BPF_DW ||
> +	    (class != BPF_LDX && class != BPF_STX && class != BPF_ST)) {
> +		verbose(env, "typed pointer field of struct %s must be accessed with a 64-bit load or store\n",
> +			tname);
> +		return -EACCES;
> +	}
> +	if (atype == BPF_READ) {
> +		/*
> +		 * The pointer leads into the pointee's typed arena, which the
> +		 * program may never cast to or allocate from: register it here,
> +		 * as a cast would, so that the sanitization has a slice to mask
> +		 * into and the slice exists for the map's lifetime.
> +		 */
> +		ta = typed_arena_register(env, btf_id);

[Severity: Medium]
Could this lead to a refcount leak for the typed arena when the program
is unloaded?

When typed_arena_register() is called here, it acquires a reference to
the typed arena via bpf_typed_arena_get() which increments ta->refcnt.
However, the program unload path in bpf_prog_free_deferred() skips
releasing the reference acquired during load:

kernel/bpf/core.c:bpf_prog_free_deferred() {
    ...
    /* The typed arenas outlive the program; only the load's failure retracts them. */
    kfree(aux->typed_arenas);
    ...
}

Because the reference is never decremented after a successful load, the
refcount will monotonically increase by one per load and unload cycle.

Since this path is reachable by any user capable of loading BPF programs,
is it possible for a user to repeatedly load and unload a program to
saturate the refcount and trigger a kernel WARN?

> +		if (IS_ERR(ta))
> +			return PTR_ERR(ta);
> +		return mark_btf_ld_reg(env, regs, value_regno, PTR_TO_BTF_ID, reg->btf, btf_id,
> +				       MEM_ARENA | PTR_UNSANITIZED);
> +	}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260926233503.3114147-1-memxor@gmail.com?part=7

^ permalink raw reply	[flat|nested] 27+ messages in thread

end of thread, other threads:[~2026-09-27  0:03 UTC | newest]

Thread overview: 27+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 01/16] mm/vmalloc: Add get_vm_area_align() Kumar Kartikeya Dwivedi
2026-09-26 23:42   ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 02/16] bpf: Introduce BPF typed arenas Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 03/16] bpf: Back typed arena chunks with scratch on demand Kumar Kartikeya Dwivedi
2026-09-26 23:56   ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 04/16] bpf: Add the typed_arena_cast instruction Kumar Kartikeya Dwivedi
2026-09-26 23:55   ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 05/16] bpf: Allow scalar and atomic access to typed arena objects Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 06/16] bpf: Support special fields in " Kumar Kartikeya Dwivedi
2026-09-26 23:59   ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 07/16] bpf: Trust typed pointer fields of " Kumar Kartikeya Dwivedi
2026-09-27  0:03   ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 08/16] bpf: Canonicalize loaded typed arena pointers where they are used Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 09/16] bpf: Add typed arena page allocation and release kfuncs Kumar Kartikeya Dwivedi
2026-09-26 23:55   ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 10/16] bpf: Let typed_arena_cast copy pointers the verifier already trusts Kumar Kartikeya Dwivedi
2026-09-26 23:49   ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 11/16] libbpf: Support the typed_arena_cast instruction Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 12/16] selftests/bpf: Build BPF objects with compiler-inserted typed arena casts Kumar Kartikeya Dwivedi
2026-09-26 23:46   ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 13/16] selftests/bpf: Test typed arena casts and registration Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 14/16] selftests/bpf: Test typed arena object access, kptrs and typed pointer fields Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 15/16] selftests/bpf: Test typed arena page allocation and release Kumar Kartikeya Dwivedi
2026-09-26 23:46   ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 16/16] selftests/bpf: Exercise typed arenas at run time Kumar Kartikeya Dwivedi
2026-09-26 23:50   ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox