From: Kumar Kartikeya Dwivedi <memxor@gmail.com>
To: bpf@vger.kernel.org
Cc: Alexei Starovoitov <ast@kernel.org>,
Andrii Nakryiko <andrii@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Eduard Zingerman <eddyz87@gmail.com>,
Emil Tsalapatis <emil@etsalapatis.com>,
kkd@meta.com, kernel-team@meta.com
Subject: [RFC PATCH bpf-next v1 03/16] bpf: Back typed arena chunks with scratch on demand
Date: Sun, 27 Sep 2026 01:34:41 +0200 [thread overview]
Message-ID: <20260926233503.3114147-4-memxor@gmail.com> (raw)
In-Reply-To: <20260926233503.3114147-1-memxor@gmail.com>
The cast that produces a typed arena pointer never fails: it masks any value
into a slot of the slice, so a program can reach every object the slice
could hold, whether or not anything allocated it. Every such pointer must
still denote a valid object of the type, or the native accesses through it
would fault, and a kptr exchange would write a kernel pointer into unbacked
memory. So back the slice on demand: an access to an unbacked page of the
slice faults, and the kernel fault recovery of the arena, reached through
bpf_arena_handle_page_fault(), installs scratch memory there and resumes
the access, reporting it on the program's stderr stream as a use of an
unallocated object.
The scratch memory is one zeroed chunk per typed arena, and the page of the
scratch chunk installed at a faulting page is the one at the same position in
the chunk. A chunk holds whole objects laid out from its start, so the dummy
objects a program sees are laid out exactly like real ones: a kptr field of
one dummy object never aliases a scalar field of another, which matters
because a kptr stored through one pointer must not be readable as a scalar
through another, and the special fields of the dummy objects can be dropped
at map teardown by walking the scratch chunk like a real one. This is why
the chunk, rather than the page, is the unit of backing, and why the slot is
a power of two: with objects straddling pages at arbitrary offsets the
layout would differ from page to page, and the scratch memory would have to
grow to the period at which it repeats. The scratch chunk is shared by every
unallocated chunk of the slice, so the dummy objects behave as one sink: a
program that writes through a stray pointer sees its writes through every
other stray pointer to the same position, and a kptr it leaves there lives
until the map is freed.
The fault path marks the chunk taken in the bitmap before installing the
page, with an atomic bit set because a program can fault in any context it
runs in, including NMI, and the allocator, added next, only backs chunks
that are not marked. A real page never replaces a scratch page either: a
program may be in the middle of using the dummy object, nothing can wait for
that use to end because the next stray access faults the scratch page right
back in, and the effect of the mapping changing under a kernel operation on
the object has not been reasoned through. The raw arena can replace its
scratch page because its contents are raw data. A chunk the fault path took
is given back with the release kfunc like any other, after which it can be
allocated or faulted again.
Refuse to register typed arenas on architectures that do not provide the
atomic PTE installer, since those do not wire up the arena fault path and a
typed access there would oops. The guard between the region and the raw
window stays outside recovery: a typed access is bounded by its slice and
cannot reach it, so a fault there is still a bug to oops on.
Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
---
kernel/bpf/arena.c | 115 ++++++++++++++++++++++++++++++++++++++++++---
1 file changed, 109 insertions(+), 6 deletions(-)
diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index 9cd1b1ce434c..c757c1932360 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -67,6 +67,16 @@
#define TYPED_ARENA_REGION_SZ SZ_4G
#define TYPED_ARENA_MAX_SZ SZ_2G
#define KERN_VM_SZ (TYPED_ARENA_REGION_SZ + SZ_4G + GUARD_SZ)
+/*
+ * Typed accesses are native loads and stores, so a fault on an unbacked typed
+ * page can only be recovered where the arena's kernel fault path is wired up,
+ * which is where the architecture provides an atomic PTE installer.
+ */
+#ifdef ptep_try_set
+#define TYPED_ARENA_SUPPORTED true
+#else
+#define TYPED_ARENA_SUPPORTED false
+#endif
static void arena_free_pages(struct bpf_arena *arena, long uaddr, long page_cnt, bool sleepable);
@@ -305,13 +315,28 @@ static unsigned int typed_arena_chunk_order(const struct bpf_typed_arena *ta)
return ta->chunk_shift - PAGE_SHIFT;
}
+static u32 typed_arena_chunk_pages(const struct bpf_typed_arena *ta)
+{
+ return bpf_typed_arena_chunk(ta) >> PAGE_SHIFT;
+}
+
+/* The page of the scratch chunk that backs the page at @addr while its chunk is unallocated. */
+static struct page *typed_arena_scratch_page(const struct bpf_typed_arena *ta, unsigned long addr)
+{
+ unsigned long off = addr - (unsigned long)ta->base;
+
+ return ta->scratch_pages[(off & (bpf_typed_arena_chunk(ta) - 1)) >> PAGE_SHIFT];
+}
+
/*
* The memory a typed arena takes for the map's lifetime: a page table entry
- * per page of the slice and the two chunk bitmaps.
+ * per page of the slice, the scratch chunk with its page array, and the two
+ * chunk bitmaps.
*/
static u64 typed_arena_static_mem(const struct bpf_typed_arena *ta)
{
return (bpf_typed_arena_size(ta) >> PAGE_SHIFT) * sizeof(pte_t) +
+ bpf_typed_arena_chunk(ta) + typed_arena_chunk_pages(ta) * sizeof(struct page *) +
2 * BITS_TO_LONGS(typed_arena_nr_chunks(ta)) * sizeof(long);
}
@@ -371,7 +396,7 @@ static int typed_arena_teardown_cb(pte_t *ptep, unsigned long addr, void *data)
if ((addr - (unsigned long)ta->base) & (bpf_typed_arena_chunk(ta) - 1))
return 0;
- if (!pte_present(pte))
+ if (!pte_present(pte) || pte_page(pte) == typed_arena_scratch_page(ta, addr))
return 0;
typed_arena_free_objects(ta, (void *)addr);
__free_pages(pte_page(pte), typed_arena_chunk_order(ta));
@@ -383,6 +408,10 @@ static void typed_arena_free(struct bpf_arena *arena, struct bpf_typed_arena *ta
WRITE_ONCE(arena->typed_arena_mem, arena->typed_arena_mem - typed_arena_static_mem(ta));
apply_to_existing_page_range(&init_mm, (unsigned long)ta->base, bpf_typed_arena_size(ta),
typed_arena_teardown_cb, ta);
+ /* Programs may have stored kptrs into the dummy objects. */
+ typed_arena_free_objects(ta, ta->scratch);
+ vfree(ta->scratch);
+ kfree(ta->scratch_pages);
bitmap_free(ta->chunks);
bitmap_free(ta->pending);
btf_put(ta->btf);
@@ -408,11 +437,14 @@ struct bpf_typed_arena *bpf_typed_arena_get(struct bpf_map *map, struct btf *btf
{
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
const struct btf_type *t = btf_type_by_id(btf, btf_id);
+ struct mem_cgroup *new_memcg, *old_memcg;
struct bpf_typed_arena *ta;
u64 slot, chunk;
s64 off;
- int err;
+ int err, i;
+ if (!TYPED_ARENA_SUPPORTED)
+ return ERR_PTR(-EOPNOTSUPP);
if (!t || !t->size || !record)
return ERR_PTR(-EINVAL);
if (!is_power_of_2(size) || size < PAGE_SIZE || size > TYPED_ARENA_MAX_SZ)
@@ -444,10 +476,17 @@ struct bpf_typed_arena *bpf_typed_arena_get(struct bpf_map *map, struct btf *btf
ta->chunk_shift = ilog2(chunk);
ta->chunks = bitmap_zalloc(typed_arena_nr_chunks(ta), GFP_KERNEL_ACCOUNT);
ta->pending = bitmap_zalloc(typed_arena_nr_chunks(ta), GFP_KERNEL_ACCOUNT);
- if (!ta->chunks || !ta->pending) {
+ ta->scratch_pages = kcalloc(typed_arena_chunk_pages(ta), sizeof(*ta->scratch_pages),
+ GFP_KERNEL_ACCOUNT);
+ bpf_map_memcg_enter(map, &old_memcg, &new_memcg);
+ ta->scratch = __vmalloc(chunk, GFP_KERNEL_ACCOUNT | __GFP_ZERO);
+ bpf_map_memcg_exit(old_memcg, new_memcg);
+ if (!ta->chunks || !ta->pending || !ta->scratch_pages || !ta->scratch) {
err = -ENOMEM;
goto free;
}
+ for (i = 0; i < typed_arena_chunk_pages(ta); i++)
+ ta->scratch_pages[i] = vmalloc_to_page(ta->scratch + i * PAGE_SIZE);
err = apply_to_page_range(&init_mm, (unsigned long)ta->base, size, apply_range_set_cb, NULL);
if (err)
goto free;
@@ -462,6 +501,8 @@ struct bpf_typed_arena *bpf_typed_arena_get(struct bpf_map *map, struct btf *btf
return ta;
free:
+ vfree(ta->scratch);
+ kfree(ta->scratch_pages);
bitmap_free(ta->chunks);
bitmap_free(ta->pending);
kfree(ta);
@@ -1460,6 +1501,64 @@ static void __bpf_prog_report_arena_violation(struct bpf_prog *prog, bool write,
}));
}
+static struct bpf_typed_arena *typed_arena_lookup(struct bpf_arena *arena, unsigned long addr)
+{
+ struct bpf_typed_arena *ta;
+
+ list_for_each_entry_rcu(ta, &arena->typed_arenas, node)
+ if (addr - (unsigned long)ta->base < bpf_typed_arena_size(ta))
+ return ta;
+ return NULL;
+}
+
+static void __bpf_prog_report_typed_arena_violation(struct bpf_prog *prog,
+ const struct bpf_typed_arena *ta,
+ bool write, unsigned long addr)
+{
+ const struct btf_type *t = btf_type_by_id(ta->btf, ta->btf_id);
+ struct bpf_stream_stage ss;
+
+ /* Use main prog for stream access */
+ prog = prog->aux->main_prog_aux->prog;
+
+ bpf_stream_stage(ss, prog, BPF_STDERR, ({
+ bpf_stream_printk(ss, "ERROR: Typed arena %s access to unallocated struct %s at 0x%lx\n",
+ write ? "WRITE" : "READ", btf_name_by_offset(ta->btf, t->name_off),
+ addr & ~((unsigned long)bpf_typed_arena_slot(ta) - 1));
+ bpf_stream_dump_stack(ss);
+ }));
+}
+
+/*
+ * A typed access reached an unbacked page, so the program used a pointer to an
+ * object nobody allocated, which the cast permits by design. The pointer must
+ * still denote an object of the type: back the page with the page of the
+ * scratch chunk at the same position, so that the dummy objects the program
+ * sees are laid out like real ones and a kptr field of one never aliases a
+ * scalar field of another. Mark the chunk taken first, so that the allocator
+ * never puts real pages where a program may be in the middle of using the
+ * dummy object; nothing could wait for that use to end, since the next stray
+ * access would fault the scratch page right back in. The mark is an atomic
+ * bit because the fault can happen in any context a program runs in.
+ */
+static bool typed_arena_handle_page_fault(struct bpf_arena *arena, struct bpf_prog *prog,
+ unsigned long addr, bool is_write)
+{
+ unsigned long page_addr = addr & PAGE_MASK;
+ struct bpf_typed_arena *ta;
+
+ guard(rcu)();
+ ta = typed_arena_lookup(arena, page_addr);
+ if (!ta)
+ return false;
+ set_bit((page_addr - (unsigned long)ta->base) >> ta->chunk_shift, ta->chunks);
+ apply_to_page_range(&init_mm, page_addr, PAGE_SIZE, apply_range_set_scratch_cb,
+ typed_arena_scratch_page(ta, page_addr));
+ flush_vmap_cache(page_addr, PAGE_SIZE);
+ __bpf_prog_report_typed_arena_violation(prog, ta, is_write, addr);
+ return true;
+}
+
bool bpf_arena_handle_page_fault(unsigned long addr, bool is_write, unsigned long fault_ip)
{
struct bpf_arena *arena;
@@ -1476,12 +1575,16 @@ bool bpf_arena_handle_page_fault(unsigned long addr, bool is_write, unsigned lon
if (!arena)
return false;
+ if (page_addr - typed_arena_region(arena) < TYPED_ARENA_REGION_SZ)
+ return typed_arena_handle_page_fault(arena, prog, addr, is_write);
+
kbase = bpf_arena_get_kern_vm_start(arena);
/*
* Recovery covers the 4 GiB mappable band plus the upper half-guard.
- * Lower guard is unreachable from kfuncs; an address there indicates
- * a different bug class - leave it to the regular kernel oops path.
+ * Lower guard is unreachable from typed accesses and from kfuncs; an
+ * address there indicates a different bug class - leave it to the
+ * regular kernel oops path.
*/
if (page_addr < kbase || page_addr >= kbase + SZ_4G + GUARD_SZ / 2)
return false;
--
2.53.0
next prev parent reply other threads:[~2026-09-26 23:35 UTC|newest]
Thread overview: 27+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 01/16] mm/vmalloc: Add get_vm_area_align() Kumar Kartikeya Dwivedi
2026-09-26 23:42 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 02/16] bpf: Introduce BPF typed arenas Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` Kumar Kartikeya Dwivedi [this message]
2026-09-26 23:56 ` [RFC PATCH bpf-next v1 03/16] bpf: Back typed arena chunks with scratch on demand sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 04/16] bpf: Add the typed_arena_cast instruction Kumar Kartikeya Dwivedi
2026-09-26 23:55 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 05/16] bpf: Allow scalar and atomic access to typed arena objects Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 06/16] bpf: Support special fields in " Kumar Kartikeya Dwivedi
2026-09-26 23:59 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 07/16] bpf: Trust typed pointer fields of " Kumar Kartikeya Dwivedi
2026-09-27 0:03 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 08/16] bpf: Canonicalize loaded typed arena pointers where they are used Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 09/16] bpf: Add typed arena page allocation and release kfuncs Kumar Kartikeya Dwivedi
2026-09-26 23:55 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 10/16] bpf: Let typed_arena_cast copy pointers the verifier already trusts Kumar Kartikeya Dwivedi
2026-09-26 23:49 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 11/16] libbpf: Support the typed_arena_cast instruction Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 12/16] selftests/bpf: Build BPF objects with compiler-inserted typed arena casts Kumar Kartikeya Dwivedi
2026-09-26 23:46 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 13/16] selftests/bpf: Test typed arena casts and registration Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 14/16] selftests/bpf: Test typed arena object access, kptrs and typed pointer fields Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 15/16] selftests/bpf: Test typed arena page allocation and release Kumar Kartikeya Dwivedi
2026-09-26 23:46 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 16/16] selftests/bpf: Exercise typed arenas at run time Kumar Kartikeya Dwivedi
2026-09-26 23:50 ` sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260926233503.3114147-4-memxor@gmail.com \
--to=memxor@gmail.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=kernel-team@meta.com \
--cc=kkd@meta.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox