BPF List
 help / color / mirror / Atom feed
From: Emil Tsalapatis <emil@etsalapatis.com>
To: bpf@vger.kernel.org
Cc: ast@kernel.org, andrii@kernel.org, eddyz87@gmail.com,
	memxor@gmail.com, daniel@iogearbox.net,
	Emil Tsalapatis <emil@etsalapatis.com>
Subject: [PATCH bpf-next v2 6/7] bpf: Avoid unavailable range leakage during arena_free_pages()
Date: Wed, 23 Sep 2026 06:14:24 +0000	[thread overview]
Message-ID: <20260923061425.7045-7-emil@etsalapatis.com> (raw)
In-Reply-To: <20260923061425.7045-1-emil@etsalapatis.com>

The arena_free_pages() call is designed to leak the attempted
freed page when it is unable to either directly complete or
defer the operation. This is valid for ranges for which freeing
has not started. For those already marked as unavailable in the
range tree, leaving them permanently unavailable breaks the
assumption that unavailable ranges denote an operation currently
in progress.

Atomically turn the range available instead. The only point a
range can be leaked as unavailable is when failing to regrab the
arena spinlock after zapping the range. At that point, the only
operation left to complete the free is to mark the range available.
The spinlock is required to merge the range with adjacent free ones
to avoid range space fragmentation, but this is not necessary for
correctness or safety. Just flip the range node back to available
without the spinlock then.

This is safe because unavailable nodes are not affected by operations
done to the other nodes, and vice versa for changing the node's available
state. We can thus treat pointers to newly created unavailable nodes
as transient references we drop when we mark the node available.
We do not make the layout of the range_node part of the tree's public
API.

The new operation causes fragmentation in the range tree. This is
preferable to leaking the range, even setting aside the correctness
issue.

Fixes: 226d4d4d1272 ("bpf: Fix arena race between page free and alloc leading to incoherency")
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
---
 kernel/bpf/arena.c      | 42 +++++++++++++++++++++++++++++++----------
 kernel/bpf/range_tree.c | 36 ++++++++++++++++++++++++++++++-----
 kernel/bpf/range_tree.h |  5 ++++-
 3 files changed, 67 insertions(+), 16 deletions(-)

diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index 8594b76dae61..299525f93bbd 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -489,6 +489,7 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf)
 	struct bpf_map *map = vmf->vma->vm_file->private_data;
 	struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
 	struct mem_cgroup *new_memcg, *old_memcg;
+	struct range_node *unavail_node;
 	struct page *page;
 	long kbase, kaddr;
 	unsigned long flags;
@@ -551,10 +552,11 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf)
 out:
 	/* Reserve the page while installing its user PTE without the arena lock. */
 	bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg);
-	ret = range_tree_set_unavail(&arena->rt, vmf->pgoff, 1);
+	unavail_node = range_tree_set_unavail(&arena->rt, vmf->pgoff, 1);
 	bpf_map_memcg_exit(old_memcg, new_memcg);
 	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
-	if (ret) {
+	if (IS_ERR(unavail_node)) {
+		ret = PTR_ERR(unavail_node);
 		if (ret == -EAGAIN)
 			goto retry;
 		return VM_FAULT_OOM;
@@ -895,6 +897,7 @@ static void zap_pages(struct bpf_arena *arena, long uaddr, long page_cnt)
 static void arena_free_pages(struct bpf_arena *arena, long uaddr, long page_cnt, bool sleepable)
 {
 	struct mem_cgroup *new_memcg, *old_memcg;
+	struct range_node *unavail_node = NULL;
 	u64 full_uaddr, uaddr_end;
 	long kaddr, pgoff;
 	struct page *page;
@@ -930,8 +933,9 @@ static void arena_free_pages(struct bpf_arena *arena, long uaddr, long page_cnt,
 	if (ret)
 		goto defer;
 
-	ret = range_tree_set_unavail(&arena->rt, pgoff, page_cnt);
-	if (ret) {
+	unavail_node = range_tree_set_unavail(&arena->rt, pgoff, page_cnt);
+	if (IS_ERR(unavail_node)) {
+		ret = PTR_ERR(unavail_node);
 		raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
 		/*
 		 * For -EAGAIN: An overlapping fault reserves
@@ -989,13 +993,29 @@ static void arena_free_pages(struct bpf_arena *arena, long uaddr, long page_cnt,
 defer:
 	s = kmalloc_nolock(sizeof(struct arena_free_span), __GFP_ACCOUNT, -1);
 	bpf_map_memcg_exit(old_memcg, new_memcg);
-	if (!s)
+	if (!s) {
+		/*
+		 * Directly mark the region available. Unavailable regions must
+		 * always be transient, and a permanent one breaks the assumptions
+		 * made in the allocation/faulting code. The operation is safe because
+		 * unavailable nodes are not modified by the range tree code without a reference
+		 * to the node. The modification from unavailable to available also does
+		 * not require touching anything in the tree apart from the node itself,
+		 * so it can be done locklessly. The downside is fragmentation in the tree,
+		 * since we cannot merge with neighbors, but this is a) safe and b) better
+		 * than leaking the range.
+		 */
+		if (release_only)
+			range_node_mark_available(unavail_node);
+
 		/*
-		 * If allocation fails in non-sleepable context, pages are intentionally left
-		 * inaccessible (leaked) until the arena is destroyed. Cleanup or retries are not
-		 * possible here, so we intentionally omit them for safety.
+		 * If we haven't even zapped the pages, intentionally leave them
+		 * inaccessible (leaked) until the arena is destroyed. Cleanup or
+		 * retries are not possible here, so we intentionally omit them for safety.
 		 */
+
 		return;
+	}
 
 	s->page_cnt = page_cnt;
 	s->uaddr = uaddr;
@@ -1048,6 +1068,7 @@ static void arena_free_worker(struct work_struct *work)
 	struct mem_cgroup *new_memcg, *old_memcg;
 	struct llist_node *list, *pos, *t;
 	struct arena_free_span *s;
+	struct range_node *unavail_node;
 	u64 arena_vm_start, user_vm_start;
 	struct llist_head free_pages;
 	struct clear_range_data cdata;
@@ -1081,8 +1102,9 @@ static void arena_free_worker(struct work_struct *work)
 		pgoff = compute_pgoff(arena, s->uaddr);
 		kaddr = arena_vm_start + s->uaddr;
 
-		ret = range_tree_set_unavail(&arena->rt, pgoff, page_cnt);
-		if (ret) {
+		unavail_node = range_tree_set_unavail(&arena->rt, pgoff, page_cnt);
+		if (IS_ERR(unavail_node)) {
+			ret = PTR_ERR(unavail_node);
 			/* Kick off another attempt at the end of this call. */
 			if (ret == -EAGAIN)
 				continue;
diff --git a/kernel/bpf/range_tree.c b/kernel/bpf/range_tree.c
index 46a32623efb2..62ebf4df51cb 100644
--- a/kernel/bpf/range_tree.c
+++ b/kernel/bpf/range_tree.c
@@ -3,6 +3,7 @@
 #include <linux/interval_tree_generic.h>
 #include <linux/slab.h>
 #include <linux/bpf.h>
+#include <linux/err.h>
 #include "range_tree.h"
 
 /*
@@ -45,7 +46,21 @@ struct range_node {
 /* Is the range available for merging? */
 static inline bool range_available(struct range_node *rn)
 {
-	return rn && rn->available;
+	/* Pairs with smp_store_release() in range_node_mark_available(). */
+	return rn && smp_load_acquire(&rn->available);
+}
+
+/*
+ * Mark a node as available. Designed to be used locklessly
+ * as a last-ditch effort to avoid leaking unavailable range
+ * tree nodes when all attempts to directly or indirectly
+ * free an arena region has failed. See arena_free_pages()
+ * for more info.
+ */
+void range_node_mark_available(struct range_node *rn)
+{
+	/* Pairs with smp_load_acquire() in range_available(). */
+	smp_store_release(&rn->available, true);
 }
 
 static struct range_node *rb_to_range_node(struct rb_node *rb)
@@ -321,7 +336,8 @@ int range_tree_make_avail(struct range_tree *rt, u32 start, u32 len)
 }
 
 /* Set the range in this range tree */
-static int range_tree_set(struct range_tree *rt, u32 start, u32 len, bool available)
+static int range_tree_set(struct range_tree *rt, u32 start, u32 len, bool available,
+			  struct range_node **new_rn)
 {
 	u32 last = start + len - 1;
 	struct range_node *right;
@@ -364,18 +380,28 @@ static int range_tree_set(struct range_tree *rt, u32 start, u32 len, bool availa
 	left->rn_start = start;
 	left->rn_last = last;
 	range_it_insert(left, rt);
+	if (new_rn)
+		*new_rn = left;
 
 	return 0;
 }
 
 int range_tree_set_avail(struct range_tree *rt, u32 start, u32 len)
 {
-	return range_tree_set(rt, start, len, true);
+	return range_tree_set(rt, start, len, true, NULL);
 }
 
-int range_tree_set_unavail(struct range_tree *rt, u32 start, u32 len)
+struct range_node *range_tree_set_unavail(struct range_tree *rt, u32 start, u32 len)
 {
-	return range_tree_set(rt, start, len, false);
+	struct range_node *rn = NULL;
+	int err;
+
+	err = range_tree_set(rt, start, len, false, &rn);
+	if (err)
+		return ERR_PTR(err);
+	if (WARN_ON_ONCE(!rn))
+		return ERR_PTR(-EINVAL);
+	return rn;
 }
 
 int range_tree_remove_unavail(struct range_tree *rt, u32 start, u32 len)
diff --git a/kernel/bpf/range_tree.h b/kernel/bpf/range_tree.h
index 4b12ef51cc0b..79295abe3684 100644
--- a/kernel/bpf/range_tree.h
+++ b/kernel/bpf/range_tree.h
@@ -10,12 +10,15 @@ struct range_tree {
 	struct rb_root_cached range_size_root;
 };
 
+struct range_node;
+
 void range_tree_init(struct range_tree *rt);
 void range_tree_destroy(struct range_tree *rt);
 
 int range_tree_clear(struct range_tree *rt, u32 start, u32 len);
 int range_tree_set_avail(struct range_tree *rt, u32 start, u32 len);
-int range_tree_set_unavail(struct range_tree *rt, u32 start, u32 len);
+struct range_node *range_tree_set_unavail(struct range_tree *rt, u32 start, u32 len);
+void range_node_mark_available(struct range_node *rn);
 int range_tree_remove_unavail(struct range_tree *rt, u32 start, u32 len);
 int range_tree_make_avail(struct range_tree *rt, u32 start, u32 len);
 int is_range_tree_set(struct range_tree *rt, u32 start, u32 len);
-- 
2.54.0


  parent reply	other threads:[~2026-09-23  6:14 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-23  6:14 [PATCH bpf-next v2 0/7] bpf: Fix arena memory incoherence Emil Tsalapatis
2026-09-23  6:14 ` [PATCH bpf-next v2 1/7] bpf: Update is_range_tree_set to work for consecutive ranges Emil Tsalapatis
2026-09-23  6:14 ` [PATCH bpf-next v2 2/7] bpf: Track availability information for ranges in range tree Emil Tsalapatis
2026-09-23  6:14 ` [PATCH bpf-next v2 3/7] bpf: Fix arena race between page free and alloc leading to incoherency Emil Tsalapatis
2026-09-23  6:14 ` [PATCH bpf-next v2 4/7] bpf: Add explicit state machine for arena free spans Emil Tsalapatis
2026-09-23  6:14 ` [PATCH bpf-next v2 5/7] bpf: Atomically update PTE and range tree in arena VM fault handler Emil Tsalapatis
2026-09-23  6:14 ` Emil Tsalapatis [this message]
2026-09-23  6:14 ` [PATCH bpf-next v2 7/7] selftests/bpf: Add arena allocation race tests Emil Tsalapatis

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260923061425.7045-7-emil@etsalapatis.com \
    --to=emil@etsalapatis.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=memxor@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox