From: sashiko-bot@kernel.org
To: "Kumar Kartikeya Dwivedi" <memxor@gmail.com>
Cc: bpf@vger.kernel.org
Subject: Re: [RFC PATCH bpf-next v1 09/16] bpf: Add typed arena page allocation and release kfuncs
Date: Sat, 26 Sep 2026 23:55:15 +0000 [thread overview]
Message-ID: <20260926235516.28EAA1F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20260926233503.3114147-10-memxor@gmail.com>
Thank you for your contribution! Sashiko AI review found 7 potential issue(s) to consider:
- [Critical] bpf: stale TLB use-after-free window in allocation error path
- [High] bpf: uninitialized kernel memory leak via arena page allocation
- [High] bpf: TLB shootdown race due to premature state publication
- [High] bpf: TOCTOU vulnerability via double fetch of user-provided page_cnt
- [High] bpf: state machine corruption race between worker and page fault handler
- [Medium] bpf: memory leak when resilient spinlock acquisition fails
- [Medium] bpf: incorrect memcg charging in tracking struct allocation
--
commit fcd33d7dbae3719eae8ea173b7e1a7f8cbe911c9
Author: Kumar Kartikeya Dwivedi <memxor@gmail.com>
bpf: Add typed arena page allocation and release kfuncs
This patch introduces typed arenas for BPF programs, allowing them to
allocate and release contiguous memory regions mapped as real objects
rather than scratch pages. It exposes two new kfuncs,
bpf_typed_arena_alloc_pages() and bpf_typed_arena_free_pages(), which
interface with a lock-free page allocator and defer the release via RCU
workers.
diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -686,6 +1045,8 @@ static unsigned long typed_arena_alloc_pages(struct bpf_typed_arena *ta, unsigne
u32 *page_cnt, int node_id)
{
[ ... ]
> + if (node_id != NUMA_NO_NODE &&
> + ((unsigned int)node_id >= nr_node_ids || !node_online(node_id)))
> + return 0;
> + if (!*page_cnt)
> + return 0;
> + chunk_cnt = DIV_ROUND_UP((unsigned long)*page_cnt, typed_arena_chunk_pages(ta));
[Severity: High]
Does this code introduce a TOCTOU vulnerability? Because page_cnt is a pointer
to memory that can be shared with BPF programs or userspace, it could be
modified concurrently. If the value is changed to zero between the
if (!*page_cnt) check and the DIV_ROUND_UP() calculation, chunk_cnt will
evaluate to 0. This would cause bitmap_find_next_zero_area() to return 0 and
grant the BPF program a valid pointer to the arena's base address without
actually allocating any pages, potentially leading to memory corruption.
[ ... ]
> + data.arena = arena;
> + for (done = 0; done < chunk_cnt; done++) {
> + head = alloc_pages_nolock(__GFP_ACCOUNT, node_id, order);
> + if (!head) {
[Severity: High]
Does this allocation leak uninitialized buddy allocator memory?
The call to alloc_pages_nolock() omits the __GFP_ZERO flag. Since arena
memory is exposed to BPF programs (and potentially mapped to userspace),
untrusted code could read stale kernel data. Furthermore, during teardown,
the kernel might invoke bpf_obj_free_fields() on this uninitialized garbage,
which could lead to panics if random memory is treated as valid kptrs.
> + ret = -ENOMEM;
> + goto back_out;
> + }
> + data.head = head;
> + data.start = start + (done << ta->chunk_shift);
> + data.i = 0;
> + ret = apply_to_page_range(&init_mm, data.start, chunk, apply_range_set_typed_cb,
> + &data);
> + if (ret) {
> + /* The fault path took this chunk: give back the allocation whole. */
> + cdata.arena = arena;
> + cdata.ta = ta;
> + if (data.i)
> + apply_to_existing_page_range(&init_mm, data.start,
> + (unsigned long)data.i << PAGE_SHIFT,
> + apply_range_clear_typed_cb, &cdata);
> + free_pages_nolock(head, order);
> + goto back_out;
> + }
[Severity: Critical]
Does this code clear PTEs and free physical pages before flushing the TLB?
When apply_to_page_range() returns -EBUSY, the partially installed PTEs are
cleared and the pages are immediately freed to the buddy allocator via
free_pages_nolock(). The actual TLB flush happens later at the back_out
label. This creates a window where freed pages can be reallocated
elsewhere in the kernel, while stale, writable VMA translations remain cached
in the TLB, allowing arbitrary memory overwrite.
[ ... ]
> +back_out:
> + /*
> + * Back out the chunks installed before the one that failed. On -EBUSY
> + * that chunk is the one the fault path took, and it keeps its mark.
> + */
> + init_llist_head(&freed);
> + cdata.arena = arena;
> + cdata.ta = ta;
> + for (i = 0; i < chunk_cnt; i++)
> + if (ret != -EBUSY || i != done)
> + clear_bit(coff + i, ta->chunks);
> + for (i = 0; i < done; i++) {
> + cdata.head = NULL;
> + apply_to_existing_page_range(&init_mm, start + (i << ta->chunk_shift), chunk,
> + apply_range_clear_typed_cb, &cdata);
> + if (cdata.head)
> + __llist_add(&cdata.head->pcp_llist, &freed);
> + }
> + raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> + flush_tlb_kernel_range(start, start + (chunk_cnt << ta->chunk_shift));
[Severity: High]
Does this sequence create a TLB shootdown race? By clearing the allocation
state bits (ta->chunks) and dropping the spinlock before calling
flush_tlb_kernel_range(), it creates a window where concurrent allocations can
immediately claim the chunk and expose new mappings. Other CPUs might retain
stale TLB translations to the old pages that are about to be freed, leading to
silent data corruption or use-after-free conditions.
[ ... ]
> +static void typed_arena_free_pages(struct bpf_typed_arena *ta, unsigned long addr, u32 page_cnt)
> +{
[ ... ]
> + first = off >> ta->chunk_shift;
> + last = (off + ((unsigned long)page_cnt << PAGE_SHIFT) - 1) >> ta->chunk_shift;
> +
> + s = kmalloc_nolock(sizeof(*s), __GFP_ACCOUNT, NUMA_NO_NODE);
[Severity: Medium]
Does this structure allocation incorrectly charge the current context's memory
cgroup? Because this is a generic kfunc that can be invoked from arbitrary
execution contexts (like tracepoints or network packets), __GFP_ACCOUNT will
charge whichever random user task happens to be interrupted. Shouldn't this be
wrapped with bpf_map_memcg_enter() and bpf_map_memcg_exit() to charge the
map's memcg?
> + if (!s)
> + /*
> + * The chunks stay allocated until the map is freed; nothing can
> + * be retried from here.
> + */
> + return;
> + if (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
> + goto free_span;
[ ... ]
> + irq_work_queue(&arena->typed_free_irq);
> + return;
> +
> +free_span:
> + kfree_nolock(s);
> +}
[Severity: Medium]
Does failing to acquire the resilient spinlock cause a silent memory leak?
If raw_res_spin_lock_irqsave() fails (e.g., due to lock contention), the
tracking struct is freed and the function returns without error. The requested
chunks remain marked as taken indefinitely, but the BPF program assumes they
were successfully queued for release, leaking arena chunks until map teardown.
[ ... ]
> +static void typed_arena_free_worker(struct work_struct *work)
> +{
[ ... ]
> + while (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
> + cpu_relax();
> + for (i = 0; i < s->chunk_cnt; i++) {
> + cdata.head = NULL;
> + apply_to_existing_page_range(&init_mm, start + (i << ta->chunk_shift),
> + bpf_typed_arena_chunk(ta),
> + apply_range_clear_typed_cb, &cdata);
> + if (cdata.head)
> + __llist_add(&cdata.head->pcp_llist, &heads);
> + }
> + typed_arena_chunks_mark(ta, s->coff, s->chunk_cnt, false);
[Severity: High]
Does this create a destructive state machine race with the page fault handler?
The worker clears the PTE mappings under the spinlock, but the page fault
handler (typed_arena_handle_page_fault()) runs locklessly. If a page fault
triggers on a cleared chunk just before typed_arena_chunks_mark(..., false) is
called, the fault handler will install a scratch page and set the chunk bit.
The worker will then blindly force the bit to 0. The chunk becomes marked as
'free' but contains a scratch page, causing all future allocations of this
chunk to permanently fail with -EBUSY.
> + bitmap_clear(ta->pending, s->coff, s->chunk_cnt);
> + raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> +
> + flush_tlb_kernel_range(start,
> + start + ((unsigned long)s->chunk_cnt << ta->chunk_shift));
[Severity: High]
Can this sequence lead to a TLB shootdown race? Similar to the allocation
error path, clearing the allocation state bits and dropping the spinlock
before flushing the TLB allows concurrent allocations to claim the chunks
while stale translations remain cached, introducing memory corruption risks.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260926233503.3114147-1-memxor@gmail.com?part=9
next prev parent reply other threads:[~2026-09-26 23:55 UTC|newest]
Thread overview: 27+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-26 23:34 [RFC PATCH bpf-next v1 00/16] BPF typed arenas Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 01/16] mm/vmalloc: Add get_vm_area_align() Kumar Kartikeya Dwivedi
2026-09-26 23:42 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 02/16] bpf: Introduce BPF typed arenas Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 03/16] bpf: Back typed arena chunks with scratch on demand Kumar Kartikeya Dwivedi
2026-09-26 23:56 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 04/16] bpf: Add the typed_arena_cast instruction Kumar Kartikeya Dwivedi
2026-09-26 23:55 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 05/16] bpf: Allow scalar and atomic access to typed arena objects Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 06/16] bpf: Support special fields in " Kumar Kartikeya Dwivedi
2026-09-26 23:59 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 07/16] bpf: Trust typed pointer fields of " Kumar Kartikeya Dwivedi
2026-09-27 0:03 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 08/16] bpf: Canonicalize loaded typed arena pointers where they are used Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 09/16] bpf: Add typed arena page allocation and release kfuncs Kumar Kartikeya Dwivedi
2026-09-26 23:55 ` sashiko-bot [this message]
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 10/16] bpf: Let typed_arena_cast copy pointers the verifier already trusts Kumar Kartikeya Dwivedi
2026-09-26 23:49 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 11/16] libbpf: Support the typed_arena_cast instruction Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 12/16] selftests/bpf: Build BPF objects with compiler-inserted typed arena casts Kumar Kartikeya Dwivedi
2026-09-26 23:46 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 13/16] selftests/bpf: Test typed arena casts and registration Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 14/16] selftests/bpf: Test typed arena object access, kptrs and typed pointer fields Kumar Kartikeya Dwivedi
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 15/16] selftests/bpf: Test typed arena page allocation and release Kumar Kartikeya Dwivedi
2026-09-26 23:46 ` sashiko-bot
2026-09-26 23:34 ` [RFC PATCH bpf-next v1 16/16] selftests/bpf: Exercise typed arenas at run time Kumar Kartikeya Dwivedi
2026-09-26 23:50 ` sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260926235516.28EAA1F000FF@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=memxor@gmail.com \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox