From: Puranjay Mohan <puranjay@kernel.org>
To: Emil Tsalapatis <emil@etsalapatis.com>, bpf@vger.kernel.org
Cc: ast@kernel.org, andrii@kernel.org, memxor@gmail.com,
daniel@iogearbox.net, eddyz87@gmail.com,
nickolay.lysenko@gmail.com,
Emil Tsalapatis <emil@etsalapatis.com>,
Puranjay Mohan <puranjay12@gmail.com>
Subject: Re: [PATCH bpf-next 3/5] bpf: Fix arena race between page free and alloc leading to incoherency
Date: Mon, 07 Sep 2026 12:41:57 +0100 [thread overview]
Message-ID: <m2ik4hqoe2.fsf@kernel.org> (raw)
In-Reply-To: <20260902070239.16968-4-emil@etsalapatis.com>
Emil Tsalapatis <emil@etsalapatis.com> writes:
> Existing arena kfunc code has an underlying race condition
> that can lead to writes being lost from the BPF program's
> point of view:
>
> a) A memory range gets gets freed by operation (1), and its
> range is added back to the arena range tree.
>
> b) A concurrent allocation (2) reallocates the range, and does
> writes to it. Writes from that CPU may follow the stale TLB
> entries into the pages that are about to be freed.
>
> c) (1) invalidates the TLB. The old pages, and any writes done
> to them, are now inaccessible. zap_pages() simlarly removes the
> mappings for userspace threads.
>
> This can be triggered by particularly demanding BPF arena data
> structures that constantly allocate and deallocate memory, like
> hash table allocations.
>
> Solve this ABA problem by preventing range reallocation until
> TLB invalidation/unmapping is complete. First, mark the range
> freed but unavailable. Afterwards, drop the spinlock lock and
> flush the kernel TLB and zap user page tables. Then pick up
> the lock again and mark the ranges as available once again,
> completing the free operation.
>
> Reported-by: Mykola Lysenko <nickolay.lysenko@gmail.com>
> Fixes: 317460317a02 ("bpf: Introduce bpf_arena.")
This bug was introduced by yours truly :D in
b8467290edab ("bpf: arena: make arena kfuncs any context safe")
before this commit everything was serialized.
> Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
> ---
> kernel/bpf/arena.c | 94 +++++++++++++++++++++++++++++++++++++++++-----
> 1 file changed, 85 insertions(+), 9 deletions(-)
>
> diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
> index f49b52fa8586..d22b71a791db 100644
> --- a/kernel/bpf/arena.c
> +++ b/kernel/bpf/arena.c
> @@ -76,6 +76,7 @@ struct arena_free_span {
> struct llist_node node;
> unsigned long uaddr;
> u32 page_cnt;
> + bool release_only;
> };
>
> u64 bpf_arena_get_kern_vm_start(struct bpf_arena *arena)
> @@ -855,6 +856,7 @@ static void arena_free_pages(struct bpf_arena *arena, long uaddr, long page_cnt,
> struct arena_free_span *s;
> struct clear_range_data cdata;
> unsigned long flags;
> + bool release_only = false;
> int ret = 0;
>
> /* only aligned lower 32-bit are relevant */
> @@ -881,7 +883,15 @@ static void arena_free_pages(struct bpf_arena *arena, long uaddr, long page_cnt,
> if (ret)
> goto defer;
>
> - range_tree_set_avail(&arena->rt, pgoff, page_cnt);
> + ret = range_tree_set_unavail(&arena->rt, pgoff, page_cnt);
> + if (ret) {
> + raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> + if (ret == -ENOMEM)
> + goto defer;
> + WARN_ON_ONCE(ret);
> + bpf_map_memcg_exit(old_memcg, new_memcg);
> + return;
> + }
>
> init_llist_head(&free_pages);
> cdata.arena = arena;
> @@ -911,6 +921,16 @@ static void arena_free_pages(struct bpf_arena *arena, long uaddr, long page_cnt,
> zap_pages(arena, full_uaddr, 1);
> __free_page(page);
> }
> +
> + ret = raw_res_spin_lock_irqsave(&arena->spinlock, flags);
> + if (ret) {
> + release_only = true;
> + goto defer;
> + }
> +
> + ret = range_tree_make_avail(&arena->rt, pgoff, page_cnt);
> + raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> + WARN_ON_ONCE(ret);
> bpf_map_memcg_exit(old_memcg, new_memcg);
>
> return;
> @@ -928,6 +948,7 @@ static void arena_free_pages(struct bpf_arena *arena, long uaddr, long page_cnt,
>
> s->page_cnt = page_cnt;
> s->uaddr = uaddr;
> + s->release_only = release_only;
> llist_add(&s->node, &arena->free_spans);
> irq_work_queue(&arena->free_irq);
> }
> @@ -977,12 +998,13 @@ static void arena_free_worker(struct work_struct *work)
> struct llist_node *list, *pos, *t;
> struct arena_free_span *s;
> u64 arena_vm_start, user_vm_start;
> - struct llist_head free_pages;
> + struct llist_head free_pages, teardown_spans, release_spans;
> struct clear_range_data cdata;
> struct page *page;
> unsigned long full_uaddr;
> long kaddr, page_cnt, pgoff;
> unsigned long flags;
> + int ret;
>
> if (raw_res_spin_lock_irqsave(&arena->spinlock, flags)) {
> schedule_work(work);
> @@ -992,28 +1014,51 @@ static void arena_free_worker(struct work_struct *work)
> bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg);
>
> init_llist_head(&free_pages);
> + init_llist_head(&teardown_spans);
> + init_llist_head(&release_spans);
> cdata.arena = arena;
> cdata.free_pages = &free_pages;
> arena_vm_start = bpf_arena_get_kern_vm_start(arena);
> user_vm_start = bpf_arena_get_user_vm_start(arena);
>
> list = llist_del_all(&arena->free_spans);
> - llist_for_each(pos, list) {
> + llist_for_each_safe(pos, t, list) {
> s = llist_entry(pos, struct arena_free_span, node);
> page_cnt = s->page_cnt;
> - kaddr = arena_vm_start + s->uaddr;
> pgoff = compute_pgoff(arena, s->uaddr);
>
> + if (s->release_only) {
> + ret = range_tree_make_avail(&arena->rt, pgoff, page_cnt);
> + WARN_ON_ONCE(ret);
> + kfree_nolock(s);
> + continue;
> + }
> +
> + kaddr = arena_vm_start + s->uaddr;
> +
> + ret = range_tree_set_unavail(&arena->rt, pgoff, page_cnt);
> + if (ret) {
> + /*
> + * An -ENOMEM failure is the same failure mode as in
> + * the defer: path of arena_free_pages(). Do not treat
> + * the leak as a bug.
> + */
> + if (ret != -ENOMEM)
> + WARN_ON_ONCE(ret);
> +
> + kfree_nolock(s);
> + continue;
> + }
> +
> /* clear ptes and collect pages in free_pages llist */
> apply_to_existing_page_range(&init_mm, kaddr, page_cnt << PAGE_SHIFT,
> apply_range_clear_cb, &cdata);
> -
> - range_tree_set_avail(&arena->rt, pgoff, page_cnt);
> + __llist_add(pos, &teardown_spans);
> }
> raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
>
> - /* Iterate the list again without holding spinlock to do the tlb flush and zap_pages */
> - llist_for_each_safe(pos, t, list) {
> + /* Keep ranges unavailable until their stale translations are gone. */
> + llist_for_each_safe(pos, t, __llist_del_all(&teardown_spans)) {
> s = llist_entry(pos, struct arena_free_span, node);
> page_cnt = s->page_cnt;
> full_uaddr = clear_lo32(user_vm_start) + s->uaddr;
> @@ -1025,7 +1070,7 @@ static void arena_free_worker(struct work_struct *work)
> /* remove pages from user vmas */
> zap_pages(arena, full_uaddr, page_cnt);
>
> - kfree_nolock(s);
> + __llist_add(pos, &release_spans);
> }
This isn't a bug, but does release_spans carry any information?
This loop has no condition in it, so every span that goes into it comes
back out into release_spans. The two lists always hold the same set, and
the split only exists because llist_for_each_safe() consumes the list for
a pass that never frees anything.
Would a plain llist_for_each() over teardown_spans work here, leaving the
list intact for the final make_avail pass and dropping release_spans
entirely?
If release_only became a small stage enum on struct arena_free_span rather
than a bool, the s->release_only = true fixup in the lock-failure path
below would go away as well: the stage can be advanced once, in the first
loop, right after range_tree_set_unavail() succeeds. The worker then reads
as three passes over one list - reserve under the lock, tear down unlocked,
release under the lock.
There may be more to gain from factoring those three steps into helpers,
since arena_free_pages() open-codes the same sequence for a single span
that arena_free_worker() open-codes for a batch, including both of the
requeue paths between them.
> /* free all pages collected by apply_to_existing_page_range() in the first loop */
> @@ -1034,6 +1079,37 @@ static void arena_free_worker(struct work_struct *work)
> __free_page(page);
> }
>
> + if (!llist_empty(&release_spans)) {
> + if (raw_res_spin_lock_irqsave(&arena->spinlock, flags)) {
> + llist_for_each_safe(pos, t, __llist_del_all(&release_spans)) {
> + s = llist_entry(pos, struct arena_free_span, node);
> + s->release_only = true;
> + llist_add(pos, &arena->free_spans);
> + }
> +
> + schedule_work(work);
> + bpf_map_memcg_exit(old_memcg, new_memcg);
> + return;
> + }
> +
> + llist_for_each_safe(pos, t, __llist_del_all(&release_spans)) {
> + s = llist_entry(pos, struct arena_free_span, node);
> + page_cnt = s->page_cnt;
> + pgoff = compute_pgoff(arena, s->uaddr);
> + /*
> + * This range tree operation does not allocate memory,
> + * and so should never fail regardless of contention
> + * or memory pressure. This is in contrast to regular
> + * inserts that _can_ fail under memory pressure and
> + * force us to defer the free.
> + */
> + ret = range_tree_make_avail(&arena->rt, pgoff, page_cnt);
> + WARN_ON_ONCE(ret);
> + kfree_nolock(s);
> + }
> + raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> + }
> +
> bpf_map_memcg_exit(old_memcg, new_memcg);
> }
>
> --
> 2.55.0
next prev parent reply other threads:[~2026-09-07 11:42 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 7:02 [PATCH bpf-next 0/5] bpf: Fix arena memory incoherence Emil Tsalapatis
2026-09-02 7:02 ` [PATCH bpf-next 1/5] bpf: Update is_range_tree_set to work for consecutive ranges Emil Tsalapatis
2026-09-02 8:01 ` bot+bpf-ci
2026-09-02 7:02 ` [PATCH bpf-next 2/5] bpf: Track availability information for ranges in range tree Emil Tsalapatis
2026-09-02 8:20 ` bot+bpf-ci
2026-09-02 7:02 ` [PATCH bpf-next 3/5] bpf: Fix arena race between page free and alloc leading to incoherency Emil Tsalapatis
2026-09-02 8:20 ` bot+bpf-ci
2026-09-07 11:41 ` Puranjay Mohan [this message]
2026-09-02 7:02 ` [PATCH bpf-next 4/5] bpf: Atomically update PTE and range tree in arena VM fault handler Emil Tsalapatis
2026-09-02 7:19 ` sashiko-bot
2026-09-07 11:45 ` Puranjay Mohan
2026-09-02 7:02 ` [PATCH bpf-next 5/5] selftests/bpf: Add arena allocation race tests Emil Tsalapatis
2026-09-02 7:14 ` sashiko-bot
2026-09-02 8:20 ` bot+bpf-ci
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=m2ik4hqoe2.fsf@kernel.org \
--to=puranjay@kernel.org \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=memxor@gmail.com \
--cc=nickolay.lysenko@gmail.com \
--cc=puranjay12@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox