bpf.vger.kernel.org archive mirror
 help / color / mirror / Atom feed
* [PATCH bpf v3 0/7] Make sleepable arena paths use sleepable alloc_pages
@ 2026-09-25 20:39 Emil Tsalapatis
  2026-09-25 20:39 ` [PATCH bpf v3 1/7] bpf: Use an llist for page allocations Emil Tsalapatis
                   ` (6 more replies)
  0 siblings, 7 replies; 11+ messages in thread
From: Emil Tsalapatis @ 2026-09-25 20:39 UTC (permalink / raw)
  To: bpf; +Cc: ast, andrii, eddyz87, memxor, daniel, Emil Tsalapatis

The arena_alloc_pages() call takes a sleepable argument based on whether
its caller is a sleepable BPF function. This flag, along with the context
the kfunc is called in, decides whether the call will try to fulfill the
allocation using the regular or the _nolock variant of the alloc_pages
API, by means of bpf_map_alloc_pages().

However, the arena_alloc_pages() call currently only makes allocations
inside an IRQ-disabled critical section. This forces all allocations to
use the _nolock() API, which may eagerly fail where the regular variant
would eventually succeed. There have been reports of this happening for
sched-ext schedulers.

Restructure arena_alloc_pages() to use the _nolock() page allocation API
only when necessary. This requires moving allocations outside of the 
spinlock critical section for sleepable calls, which in turn requires
slightly different logic in the allocation path. Make a separate
sleepable code path that implements this logic within
arena_alloc_pages().

Also fix kfunc specialization to not unnecessarily force the nonsleepable version
of bpf_arena_alloc_pages() for call sites that do not need it. This requires
making specialization per-call site instead of overwriting the descriptor
during fixups. 

Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>

v1 -> v2 (https://lore.kernel.org/bpf/20260824082530.47553-1-emil@etsalapatis.com/)

- Keep the sleepable and non-sleepable allocation paths within
  arena_alloc_pages (Alexei)
- Incorporate bot feedback on selftests (bot-ci)

v2 -> v3 (https://lore.kernel.org/bpf/20260923191125.5311-1-emil@etsalapatis.com/)

- Remove the intermediate page array allocation in bpf_arena_alloc_pages() and
  use the pcp_llist pointer instead (Alexei)
- Fix function specialization to only specialize to the nonsleepable version
  when necessary


Emil Tsalapatis (7):
  bpf: Use an llist for page allocations
  bpf: Add sleepable argument to bpf_alloc_pages()
  bpf: Add sleepable arena page allocation path
  selftests/bpf: Test large allocations for both sleepable/nonsleepable
    arena users
  bpf: Support call-site kfunc specialization for near calls
  bpf: Support call-site kfunc specialization for far calls
  selftests/bpf: Test per-call site function specialization

 include/linux/bpf.h                           |  10 +-
 include/linux/bpf_verifier.h                  |  15 +-
 kernel/bpf/arena.c                            | 186 +++++++++++-------
 kernel/bpf/fixups.c                           |  18 +-
 kernel/bpf/syscall.c                          |  43 ++--
 kernel/bpf/verifier.c                         |  82 +++++++-
 .../selftests/bpf/prog_tests/file_reader.c    |  15 ++
 .../testing/selftests/bpf/progs/file_reader.c | 129 ++++++++++++
 .../bpf/progs/verifier_arena_large.c          |  64 ++++--
 9 files changed, 434 insertions(+), 128 deletions(-)

-- 
2.52.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH bpf v3 1/7] bpf: Use an llist for page allocations
  2026-09-25 20:39 [PATCH bpf v3 0/7] Make sleepable arena paths use sleepable alloc_pages Emil Tsalapatis
@ 2026-09-25 20:39 ` Emil Tsalapatis
  2026-09-25 20:39 ` [PATCH bpf v3 2/7] bpf: Add sleepable argument to bpf_alloc_pages() Emil Tsalapatis
                   ` (5 subsequent siblings)
  6 siblings, 0 replies; 11+ messages in thread
From: Emil Tsalapatis @ 2026-09-25 20:39 UTC (permalink / raw)
  To: bpf; +Cc: ast, andrii, eddyz87, memxor, daniel, Emil Tsalapatis

bpf_map_alloc_pages() does not use its map argument. Storing allocated
pages in an array also forces arena callers to allocate a separate
pointer array.

Expose the single-page allocator as bpf_alloc_page(), rename the bulk
helper to bpf_alloc_pages(), and return bulk allocations through an
llist using page->pcp_llist. Add bpf_free_pages() to safely release all
pages remaining on such a list.

Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
---
 include/linux/bpf.h  |  6 ++-
 kernel/bpf/arena.c   | 95 +++++++++++++++++++-------------------------
 kernel/bpf/syscall.c | 39 ++++++++++--------
 3 files changed, 67 insertions(+), 73 deletions(-)

diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index e7c5e203eddd..52242d88cb51 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -2884,8 +2884,10 @@ struct bpf_map *bpf_map_get_curr_or_next(u32 *id);
 struct bpf_prog *bpf_prog_get_curr_or_next(u32 *id);
 
 
-int bpf_map_alloc_pages(const struct bpf_map *map, int nid,
-			unsigned long nr_pages, struct page **page_array);
+struct page *bpf_alloc_page(int nid);
+int bpf_alloc_pages(int nid, unsigned long nr_pages,
+		    struct llist_head *pages);
+void bpf_free_pages(struct llist_head *pages);
 #ifdef CONFIG_MEMCG
 void bpf_map_memcg_enter(const struct bpf_map *map, struct mem_cgroup **old_memcg,
 			 struct mem_cgroup **new_memcg);
diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index c6369ea5e208..de4f7c7f68f5 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -146,7 +146,7 @@ static long compute_pgoff(struct bpf_arena *arena, long uaddr)
 
 struct apply_range_data {
 	struct bpf_arena *arena;
-	struct page **pages;
+	struct llist_head *pages;
 	int i;
 };
 
@@ -158,13 +158,17 @@ struct clear_range_data {
 static int apply_range_set_cb(pte_t *pte, unsigned long addr, void *data)
 {
 	struct apply_range_data *d = data;
+	struct llist_node *node;
 	struct page *page;
 	pte_t pteval;
 
 	if (!data)
 		return 0;
 
-	page = d->pages[d->i];
+	node = READ_ONCE(d->pages->first);
+	if (WARN_ON_ONCE(!node))
+		return -EINVAL;
+	page = llist_entry(node, struct page, pcp_llist);
 	/* paranoia, similar to vmap_pages_pte_range() */
 	if (WARN_ON_ONCE(!pfn_valid(page_to_pfn(page))))
 		return -EINVAL;
@@ -197,6 +201,7 @@ static int apply_range_set_cb(pte_t *pte, unsigned long addr, void *data)
 		return -EBUSY;
 	set_pte_at(&init_mm, addr, pte, pteval);
 #endif
+	WARN_ON_ONCE(llist_del_first(d->pages) != node);
 	d->i++;
 	WRITE_ONCE(d->arena->nr_pages, d->arena->nr_pages + 1);
 	return 0;
@@ -312,8 +317,8 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
 	INIT_WORK(&arena->free_work, arena_free_worker);
 	bpf_map_init_from_attr(&arena->map, attr);
 
-	err = bpf_map_alloc_pages(&arena->map, NUMA_NO_NODE, 1, &arena->scratch_page);
-	if (err)
+	arena->scratch_page = bpf_alloc_page(NUMA_NO_NODE);
+	if (!arena->scratch_page)
 		goto err_free_arena;
 
 	range_tree_init(&arena->rt);
@@ -481,7 +486,9 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf)
 	struct bpf_map *map = vmf->vma->vm_file->private_data;
 	struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
 	struct mem_cgroup *new_memcg, *old_memcg;
+	LLIST_HEAD(pages);
 	struct page *page, *new_page = NULL;
+	struct apply_range_data data;
 	vm_fault_t fault_ret;
 	long kbase, kaddr;
 	unsigned long flags;
@@ -543,8 +550,8 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf)
 		 * The probed page was freed meanwhile or preallocation failed;
 		 * try the non-blocking allocator, we cannot sleep here.
 		 */
-		ret = bpf_map_alloc_pages(map, map->numa_node, 1, &new_page);
-		if (ret) {
+		new_page = bpf_alloc_page(map->numa_node);
+		if (!new_page) {
 			fault_ret = VM_FAULT_SIGBUS;
 			goto out_err_locked_memcg;
 		}
@@ -555,12 +562,14 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf)
 		fault_ret = VM_FAULT_SIGBUS;
 		goto out_err_locked_memcg;
 	}
-	struct apply_range_data data = {
-		.arena = arena, .pages = &new_page, .i = 0
-	};
+	llist_add(&new_page->pcp_llist, &pages);
+	data.arena = arena;
+	data.pages = &pages;
+	data.i = 0;
 
 	ret = apply_to_page_range(&init_mm, kaddr, PAGE_SIZE, apply_range_set_cb, &data);
 	if (ret) {
+		llist_del_first(&pages);
 		range_tree_set(&arena->rt, vmf->pgoff, 1);
 		fault_ret = VM_FAULT_SIGBUS;
 		goto out_err_locked_memcg;
@@ -716,13 +725,12 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
 	u64 kern_vm_start = bpf_arena_get_kern_vm_start(arena);
 	struct mem_cgroup *new_memcg, *old_memcg;
 	struct apply_range_data data;
-	struct page **pages = NULL;
-	long remaining, mapped = 0;
-	long alloc_pages;
+	LLIST_HEAD(pages);
+	long mapped = 0;
 	unsigned long flags;
 	long pgoff = 0;
 	u32 uaddr32;
-	int ret, i;
+	int ret;
 
 	if (node_id != NUMA_NO_NODE &&
 	    ((unsigned int)node_id >= nr_node_ids || !node_online(node_id)))
@@ -741,15 +749,9 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
 	}
 
 	bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg);
-	/* Cap allocation size to KMALLOC_MAX_CACHE_SIZE so kmalloc_nolock() can succeed. */
-	alloc_pages = min(page_cnt, KMALLOC_MAX_CACHE_SIZE / sizeof(struct page *));
-	pages = kmalloc_nolock(alloc_pages * sizeof(struct page *), __GFP_ACCOUNT, NUMA_NO_NODE);
-	if (!pages) {
-		bpf_map_memcg_exit(old_memcg, new_memcg);
-		return 0;
-	}
 	data.arena = arena;
-	data.pages = pages;
+	data.pages = &pages;
+	data.i = 0;
 
 	if (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
 		goto out_free_pages;
@@ -767,45 +769,28 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
 	if (ret)
 		goto out_unlock_free_pages;
 
-	remaining = page_cnt;
 	uaddr32 = (u32)(arena->user_vm_start + pgoff * PAGE_SIZE);
 
-	while (remaining) {
-		long this_batch = min(remaining, alloc_pages);
-
-		/* zeroing is needed, since alloc_pages_bulk() only fills in non-zero entries */
-		memset(pages, 0, this_batch * sizeof(struct page *));
-
-		ret = bpf_map_alloc_pages(&arena->map, node_id, this_batch, pages);
-		if (ret)
-			goto out;
+	ret = bpf_alloc_pages(node_id, page_cnt, &pages);
+	if (ret)
+		goto out;
 
-		/*
-		 * Earlier checks made sure that uaddr32 + page_cnt * PAGE_SIZE - 1
-		 * will not overflow 32-bit. Lower 32-bit need to represent
-		 * contiguous user address range.
-		 * Map these pages at kern_vm_start base.
-		 * kern_vm_start + uaddr32 + page_cnt * PAGE_SIZE - 1 can overflow
-		 * lower 32-bit and it's ok.
-		 */
-		data.i = 0;
-		ret = apply_to_page_range(&init_mm,
-					  kern_vm_start + uaddr32 + (mapped << PAGE_SHIFT),
-					  this_batch << PAGE_SHIFT, apply_range_set_cb, &data);
-		if (ret) {
-			/* data.i pages were mapped, account them and free the remaining */
-			mapped += data.i;
-			for (i = data.i; i < this_batch; i++)
-				free_pages_nolock(pages[i], 0);
-			goto out;
-		}
+	/*
+	 * Earlier checks made sure that uaddr32 + page_cnt * PAGE_SIZE - 1
+	 * will not overflow 32-bit. Lower 32-bit need to represent
+	 * contiguous user address range.
+	 * Map these pages at kern_vm_start base.
+	 * kern_vm_start + uaddr32 + page_cnt * PAGE_SIZE - 1 can overflow
+	 * lower 32-bit and it's ok.
+	 */
+	ret = apply_to_page_range(&init_mm, kern_vm_start + uaddr32,
+				  page_cnt << PAGE_SHIFT, apply_range_set_cb, &data);
+	mapped = data.i;
+	if (ret)
+		goto out;
 
-		mapped += this_batch;
-		remaining -= this_batch;
-	}
 	flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
 	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
-	kfree_nolock(pages);
 	bpf_map_memcg_exit(old_memcg, new_memcg);
 	return clear_lo32(arena->user_vm_start) + uaddr32;
 out:
@@ -819,7 +804,7 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
 out_unlock_free_pages:
 	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
 out_free_pages:
-	kfree_nolock(pages);
+	bpf_free_pages(&pages);
 	bpf_map_memcg_exit(old_memcg, new_memcg);
 	return 0;
 }
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index 74496fd716d3..80cae0979c23 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -602,7 +602,7 @@ static bool can_alloc_pages(void)
 		!IS_ENABLED(CONFIG_PREEMPT_RT);
 }
 
-static struct page *__bpf_alloc_page(int nid)
+struct page *bpf_alloc_page(int nid)
 {
 	if (!can_alloc_pages())
 		return alloc_pages_nolock(__GFP_ACCOUNT, nid, 0);
@@ -613,27 +613,34 @@ static struct page *__bpf_alloc_page(int nid)
 				0);
 }
 
-int bpf_map_alloc_pages(const struct bpf_map *map, int nid,
-			unsigned long nr_pages, struct page **pages)
+void bpf_free_pages(struct llist_head *pages)
 {
-	unsigned long i, j;
+	struct llist_node *node;
+	struct page *page, *tmp;
+
+	node = llist_del_all(pages);
+	llist_for_each_entry_safe(page, tmp, node, pcp_llist)
+		free_pages_nolock(page, 0);
+}
+
+int bpf_alloc_pages(int nid, unsigned long nr_pages,
+		    struct llist_head *pages)
+{
+	unsigned long i;
 	struct page *pg;
-	int ret = 0;
 
 	for (i = 0; i < nr_pages; i++) {
-		pg = __bpf_alloc_page(nid);
-
-		if (pg) {
-			pages[i] = pg;
-			continue;
-		}
-		for (j = 0; j < i; j++)
-			free_pages_nolock(pages[j], 0);
-		ret = -ENOMEM;
-		break;
+		pg = bpf_alloc_page(nid);
+		if (!pg)
+			goto free_pages;
+		llist_add(&pg->pcp_llist, pages);
 	}
 
-	return ret;
+	return 0;
+
+free_pages:
+	bpf_free_pages(pages);
+	return -ENOMEM;
 }
 
 static int btf_field_cmp(const void *a, const void *b)
-- 
2.52.0


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH bpf v3 2/7] bpf: Add sleepable argument to bpf_alloc_pages()
  2026-09-25 20:39 [PATCH bpf v3 0/7] Make sleepable arena paths use sleepable alloc_pages Emil Tsalapatis
  2026-09-25 20:39 ` [PATCH bpf v3 1/7] bpf: Use an llist for page allocations Emil Tsalapatis
@ 2026-09-25 20:39 ` Emil Tsalapatis
  2026-09-25 20:39 ` [PATCH bpf v3 3/7] bpf: Add sleepable arena page allocation path Emil Tsalapatis
                   ` (4 subsequent siblings)
  6 siblings, 0 replies; 11+ messages in thread
From: Emil Tsalapatis @ 2026-09-25 20:39 UTC (permalink / raw)
  To: bpf; +Cc: ast, andrii, eddyz87, memxor, daniel, Emil Tsalapatis

bpf_alloc_pages() currently decides whether it may use the blocking page allocator from the current execution context alone. Let callers further restrict that choice by passing whether their context is sleepable.

Use the blocking allocator only when both the caller and runtime context allow sleeping. Add __GFP_RETRY_MAYFAIL so this path can reclaim without invoking the OOM killer when the allocation is charged to another memcg. Existing non-sleepable callers retain the no-lock allocation behavior.

Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
---
 include/linux/bpf.h  |  4 ++--
 kernel/bpf/arena.c   |  6 +++---
 kernel/bpf/syscall.c | 10 +++++-----
 3 files changed, 10 insertions(+), 10 deletions(-)

diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 52242d88cb51..7df74e47ecb0 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -2884,9 +2884,9 @@ struct bpf_map *bpf_map_get_curr_or_next(u32 *id);
 struct bpf_prog *bpf_prog_get_curr_or_next(u32 *id);
 
 
-struct page *bpf_alloc_page(int nid);
+struct page *bpf_alloc_page(int nid, bool sleepable);
 int bpf_alloc_pages(int nid, unsigned long nr_pages,
-		    struct llist_head *pages);
+		    struct llist_head *pages, bool sleepable);
 void bpf_free_pages(struct llist_head *pages);
 #ifdef CONFIG_MEMCG
 void bpf_map_memcg_enter(const struct bpf_map *map, struct mem_cgroup **old_memcg,
diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index de4f7c7f68f5..c556df7730c4 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -317,7 +317,7 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
 	INIT_WORK(&arena->free_work, arena_free_worker);
 	bpf_map_init_from_attr(&arena->map, attr);
 
-	arena->scratch_page = bpf_alloc_page(NUMA_NO_NODE);
+	arena->scratch_page = bpf_alloc_page(NUMA_NO_NODE, true);
 	if (!arena->scratch_page)
 		goto err_free_arena;
 
@@ -550,7 +550,7 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf)
 		 * The probed page was freed meanwhile or preallocation failed;
 		 * try the non-blocking allocator, we cannot sleep here.
 		 */
-		new_page = bpf_alloc_page(map->numa_node);
+		new_page = bpf_alloc_page(map->numa_node, false);
 		if (!new_page) {
 			fault_ret = VM_FAULT_SIGBUS;
 			goto out_err_locked_memcg;
@@ -771,7 +771,7 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
 
 	uaddr32 = (u32)(arena->user_vm_start + pgoff * PAGE_SIZE);
 
-	ret = bpf_alloc_pages(node_id, page_cnt, &pages);
+	ret = bpf_alloc_pages(node_id, page_cnt, &pages, false);
 	if (ret)
 		goto out;
 
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index 80cae0979c23..b3b0cd349eef 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -602,14 +602,14 @@ static bool can_alloc_pages(void)
 		!IS_ENABLED(CONFIG_PREEMPT_RT);
 }
 
-struct page *bpf_alloc_page(int nid)
+struct page *bpf_alloc_page(int nid, bool sleepable)
 {
-	if (!can_alloc_pages())
+	if (!sleepable || !can_alloc_pages())
 		return alloc_pages_nolock(__GFP_ACCOUNT, nid, 0);
 
 	return alloc_pages_node(nid,
 				GFP_KERNEL | __GFP_ZERO | __GFP_ACCOUNT
-				| __GFP_NOWARN,
+				| __GFP_NOWARN | __GFP_RETRY_MAYFAIL,
 				0);
 }
 
@@ -624,13 +624,13 @@ void bpf_free_pages(struct llist_head *pages)
 }
 
 int bpf_alloc_pages(int nid, unsigned long nr_pages,
-		    struct llist_head *pages)
+		    struct llist_head *pages, bool sleepable)
 {
 	unsigned long i;
 	struct page *pg;
 
 	for (i = 0; i < nr_pages; i++) {
-		pg = bpf_alloc_page(nid);
+		pg = bpf_alloc_page(nid, sleepable);
 		if (!pg)
 			goto free_pages;
 		llist_add(&pg->pcp_llist, pages);
-- 
2.52.0


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH bpf v3 3/7] bpf: Add sleepable arena page allocation path
  2026-09-25 20:39 [PATCH bpf v3 0/7] Make sleepable arena paths use sleepable alloc_pages Emil Tsalapatis
  2026-09-25 20:39 ` [PATCH bpf v3 1/7] bpf: Use an llist for page allocations Emil Tsalapatis
  2026-09-25 20:39 ` [PATCH bpf v3 2/7] bpf: Add sleepable argument to bpf_alloc_pages() Emil Tsalapatis
@ 2026-09-25 20:39 ` Emil Tsalapatis
  2026-09-25 21:25   ` Alexei Starovoitov
  2026-09-25 20:39 ` [PATCH bpf v3 4/7] selftests/bpf: Test large allocations for both sleepable/nonsleepable arena users Emil Tsalapatis
                   ` (3 subsequent siblings)
  6 siblings, 1 reply; 11+ messages in thread
From: Emil Tsalapatis @ 2026-09-25 20:39 UTC (permalink / raw)
  To: bpf; +Cc: ast, andrii, eddyz87, memxor, daniel, Emil Tsalapatis

The bpf_arena_alloc_pages() function currently only allocates pages
inside a spinlock critical section with IRQs off. This forces the use of
alloc_pages_nolock() in the BPF allocator, even when the caller is a
sleepable BPF function. This in turn causes allocation failures even in
cases where falling into the allocator slow path and possibly sleeping
would eventually succeed. This can be triggered consistently by heavy
BPF arena users like scx.

Allocate the arena pages before taking the critical section and pass
whether the caller can sleep to bpf_alloc_pages(). This lets sleepable
callers use the blocking allocator while non-sleepable callers retain
the no-lock allocation behavior.

Since allocation now happens before range-tree arbitration, concurrent
callers can allocate pages for competing requests. Bound the aggregate
number of pages in this pre-arbitration window to the configured maximum
size of the arena.

Fixes: b8467290edab ("bpf: arena: make arena kfuncs any context safe")
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
---
 kernel/bpf/arena.c | 99 ++++++++++++++++++++++++++++++++++------------
 1 file changed, 74 insertions(+), 25 deletions(-)

diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index c556df7730c4..3301445deebb 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -59,6 +59,8 @@ struct bpf_arena {
 	rqspinlock_t spinlock;
 	/* number of pages currently populated in the arena */
 	u64 nr_pages;
+	/* number of pages being allocated before range-tree arbitration */
+	atomic_long_t nr_pages_inflight;
 	struct list_head vma_list;
 	/* protects vma_list */
 	struct mutex lock;
@@ -316,6 +318,7 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
 	init_irq_work(&arena->free_irq, arena_free_irq);
 	INIT_WORK(&arena->free_work, arena_free_worker);
 	bpf_map_init_from_attr(&arena->map, attr);
+	atomic_long_set(&arena->nr_pages_inflight, 0);
 
 	arena->scratch_page = bpf_alloc_page(NUMA_NO_NODE, true);
 	if (!arena->scratch_page)
@@ -713,6 +716,48 @@ static u64 clear_lo32(u64 val)
 	return val & ~(u64)~0U;
 }
 
+static bool arena_reserve_inflight_pages(struct bpf_arena *arena, long page_cnt)
+{
+	long limit = arena->map.max_entries;
+	long old;
+
+	old = atomic_long_read(&arena->nr_pages_inflight);
+	do {
+		if (old > limit || page_cnt > limit - old)
+			return false;
+	} while (!atomic_long_try_cmpxchg(&arena->nr_pages_inflight, &old,
+					 old + page_cnt));
+
+	return true;
+}
+
+static void arena_release_inflight_pages(struct bpf_arena *arena, long page_cnt)
+{
+	WARN_ON_ONCE(atomic_long_sub_return(page_cnt,
+					    &arena->nr_pages_inflight) < 0);
+}
+
+static int arena_adjust_tree(struct bpf_arena *arena, long uaddr, long page_cnt, long *pgoff)
+{
+	int ret;
+
+	/* Special case where user is requesting specific range. */
+	if (uaddr) {
+		ret = is_range_tree_set(&arena->rt, *pgoff, page_cnt);
+		if (ret)
+			return ret;
+		return range_tree_clear(&arena->rt, *pgoff, page_cnt);
+	}
+
+	ret = range_tree_find(&arena->rt, page_cnt);
+	if (ret < 0)
+		return ret;
+
+	*pgoff = ret;
+
+	return range_tree_clear(&arena->rt, *pgoff, page_cnt);
+}
+
 /*
  * Allocate pages and vmap them into kernel vmalloc area.
  * Later the pages will be mmaped into user space vma.
@@ -730,6 +775,7 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
 	unsigned long flags;
 	long pgoff = 0;
 	u32 uaddr32;
+	long addr = 0;
 	int ret;
 
 	if (node_id != NUMA_NO_NODE &&
@@ -747,8 +793,15 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
 			/* requested address will be outside of user VMA */
 			return 0;
 	}
+	if (!arena_reserve_inflight_pages(arena, page_cnt))
+		return 0;
 
 	bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg);
+
+	ret = bpf_alloc_pages(node_id, page_cnt, &pages, sleepable);
+	if (ret)
+		goto out_memcg;
+
 	data.arena = arena;
 	data.pages = &pages;
 	data.i = 0;
@@ -756,25 +809,14 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
 	if (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
 		goto out_free_pages;
 
-	if (uaddr) {
-		ret = is_range_tree_set(&arena->rt, pgoff, page_cnt);
-		if (ret)
-			goto out_unlock_free_pages;
-		ret = range_tree_clear(&arena->rt, pgoff, page_cnt);
-	} else {
-		ret = pgoff = range_tree_find(&arena->rt, page_cnt);
-		if (pgoff >= 0)
-			ret = range_tree_clear(&arena->rt, pgoff, page_cnt);
+	ret = arena_adjust_tree(arena, uaddr, page_cnt, &pgoff);
+	if (ret) {
+		raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
+		goto out_free_pages;
 	}
-	if (ret)
-		goto out_unlock_free_pages;
 
 	uaddr32 = (u32)(arena->user_vm_start + pgoff * PAGE_SIZE);
 
-	ret = bpf_alloc_pages(node_id, page_cnt, &pages, false);
-	if (ret)
-		goto out;
-
 	/*
 	 * Earlier checks made sure that uaddr32 + page_cnt * PAGE_SIZE - 1
 	 * will not overflow 32-bit. Lower 32-bit need to represent
@@ -787,26 +829,33 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
 				  page_cnt << PAGE_SHIFT, apply_range_set_cb, &data);
 	mapped = data.i;
 	if (ret)
-		goto out;
+		goto out_unmap;
 
 	flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
 	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
-	bpf_map_memcg_exit(old_memcg, new_memcg);
-	return clear_lo32(arena->user_vm_start) + uaddr32;
-out:
+
+	addr = clear_lo32(arena->user_vm_start) + uaddr32;
+	goto out_memcg;
+
+out_unmap:
+	if (sleepable)
+		flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
 	range_tree_set(&arena->rt, pgoff + mapped, page_cnt - mapped);
 	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
-	if (mapped) {
-		flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
+	if (mapped || sleepable) {
+		if (!sleepable)
+			flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
 		arena_free_pages(arena, uaddr32, mapped, sleepable);
 	}
-	goto out_free_pages;
-out_unlock_free_pages:
-	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
+
 out_free_pages:
 	bpf_free_pages(&pages);
+
+out_memcg:
 	bpf_map_memcg_exit(old_memcg, new_memcg);
-	return 0;
+	arena_release_inflight_pages(arena, page_cnt);
+
+	return addr;
 }
 
 /*
-- 
2.52.0


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH bpf v3 4/7] selftests/bpf: Test large allocations for both sleepable/nonsleepable arena users
  2026-09-25 20:39 [PATCH bpf v3 0/7] Make sleepable arena paths use sleepable alloc_pages Emil Tsalapatis
                   ` (2 preceding siblings ...)
  2026-09-25 20:39 ` [PATCH bpf v3 3/7] bpf: Add sleepable arena page allocation path Emil Tsalapatis
@ 2026-09-25 20:39 ` Emil Tsalapatis
  2026-09-25 20:39 ` [PATCH bpf v3 5/7] bpf: Support call-site kfunc specialization for near calls Emil Tsalapatis
                   ` (2 subsequent siblings)
  6 siblings, 0 replies; 11+ messages in thread
From: Emil Tsalapatis @ 2026-09-25 20:39 UTC (permalink / raw)
  To: bpf; +Cc: ast, andrii, eddyz87, memxor, daniel, Emil Tsalapatis

We have relaxed the limitation of 1024-page batches for both sleepable
and nonsleepable arena allocations by removing the intermediate page
array. Add a test to confirm that large allocations succeed for both.

Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
---
 .../bpf/progs/verifier_arena_large.c          | 64 +++++++++++++++----
 1 file changed, 50 insertions(+), 14 deletions(-)

diff --git a/tools/testing/selftests/bpf/progs/verifier_arena_large.c b/tools/testing/selftests/bpf/progs/verifier_arena_large.c
index 6ab8730d4878..dbdea14ca76f 100644
--- a/tools/testing/selftests/bpf/progs/verifier_arena_large.c
+++ b/tools/testing/selftests/bpf/progs/verifier_arena_large.c
@@ -10,6 +10,9 @@
 #include <bpf_arena_common.h>
 
 #define ARENA_SIZE (1ull << 32)
+#define LARGE_PAGE_CNT 1025
+
+volatile int zero = 0;
 
 struct {
 	__uint(type, BPF_MAP_TYPE_ARENA);
@@ -284,6 +287,7 @@ int big_alloc2(void *ctx)
 	return 0;
 }
 
+/* Nonsleepable because it binds to a socket program. */
 SEC("socket")
 __success __retval(0)
 int big_alloc3(void *ctx)
@@ -291,24 +295,56 @@ int big_alloc3(void *ctx)
 #if defined(__BPF_FEATURE_ADDR_SPACE_CAST)
 	char __arena *pages;
 	u64 i;
+	int err = 0;
 
-	/*
-	 * Allocate 2051 pages in one go to check how kmalloc_nolock() handles large requests.
-	 * Since kmalloc_nolock() can allocate up to 1024 struct page * at a time, this call should
-	 * result in three batches: two batches of 1024 pages each, followed by a final batch of 3
-	 * pages.
-	 */
-	pages = bpf_arena_alloc_pages(&arena, NULL, 2051, NUMA_NO_NODE, 0);
+	/* Verify that a nonsleepable allocation larger than 1024 pages succeeds. */
+	pages = bpf_arena_alloc_pages(&arena, NULL, LARGE_PAGE_CNT, NUMA_NO_NODE, 0);
 	if (!pages)
-		return 0;
+		return 1;
+
+	for (i = zero; i < LARGE_PAGE_CNT && can_loop; i++)
+		pages[i * PAGE_SIZE] = 123;
+
+	for (i = zero; i < LARGE_PAGE_CNT && can_loop; i++) {
+		if (pages[i * PAGE_SIZE] == 123)
+			continue;
+		err = 2;
+		break;
+	}
+
+	bpf_arena_free_pages(&arena, pages, LARGE_PAGE_CNT);
+	return err;
+#endif
+	return 0;
+}
 
-	bpf_for(i, 0, 2051)
-			pages[i * PAGE_SIZE] = 123;
-	bpf_for(i, 0, 2051)
-			if (pages[i * PAGE_SIZE] != 123)
-				return i;
+/* SYSCALL programs are always sleepable. */
+SEC("syscall")
+__success __retval(0)
+int big_alloc4(void *ctx)
+{
+#if defined(__BPF_FEATURE_ADDR_SPACE_CAST)
+	char __arena *pages;
+	u64 i;
+	int err = 0;
+
+	/* Verify that a sleepable allocation larger than 1024 pages succeeds. */
+	pages = bpf_arena_alloc_pages(&arena, NULL, LARGE_PAGE_CNT, NUMA_NO_NODE, 0);
+	if (!pages)
+		return 1;
+
+	for (i = zero; i < LARGE_PAGE_CNT && can_loop; i++)
+		pages[i * PAGE_SIZE] = 123;
+
+	for (i = zero; i < LARGE_PAGE_CNT && can_loop; i++) {
+		if (pages[i * PAGE_SIZE] == 123)
+			continue;
+		err = 2;
+		break;
+	}
 
-	bpf_arena_free_pages(&arena, pages, 2051);
+	bpf_arena_free_pages(&arena, pages, LARGE_PAGE_CNT);
+	return err;
 #endif
 	return 0;
 }
-- 
2.52.0


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH bpf v3 5/7] bpf: Support call-site kfunc specialization for near calls
  2026-09-25 20:39 [PATCH bpf v3 0/7] Make sleepable arena paths use sleepable alloc_pages Emil Tsalapatis
                   ` (3 preceding siblings ...)
  2026-09-25 20:39 ` [PATCH bpf v3 4/7] selftests/bpf: Test large allocations for both sleepable/nonsleepable arena users Emil Tsalapatis
@ 2026-09-25 20:39 ` Emil Tsalapatis
  2026-09-25 20:39 ` [PATCH bpf v3 6/7] bpf: Support call-site kfunc specialization for far calls Emil Tsalapatis
  2026-09-25 20:39 ` [PATCH bpf v3 7/7] selftests/bpf: Test per-call site function specialization Emil Tsalapatis
  6 siblings, 0 replies; 11+ messages in thread
From: Emil Tsalapatis @ 2026-09-25 20:39 UTC (permalink / raw)
  To: bpf; +Cc: ast, andrii, eddyz87, memxor, daniel, Emil Tsalapatis

specialize_kfunc() currently updates the canonical kfunc descriptor
in place. It is not currently possible to swtich different
specializations of a kfunc per call site in the same program.
In fact, specializations are order-dependent: Once a function is
specialized, all subsequent call sites are specialized even if they
wouldn't trigger specialization themselves. This is especially an
issue for bpf_arena_alloc_pages() that is specialized into its
non-sleepable for all call sites after a single non-sleepable one.

Allow per-call site kfunc specialization for JITs that use near calls.
Implement this by keeping two versions of the kfunc table, one with just
the initial kfuncs and one with all valid specializations for the
program. We currently assume 2 concurrent specializations for each
kfunc. This is a conservative estimate, since most of them do not
specialize at all.

Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
---
 include/linux/bpf_verifier.h | 13 ++++---
 kernel/bpf/verifier.c        | 67 +++++++++++++++++++++++++++++++++---
 2 files changed, 71 insertions(+), 9 deletions(-)

diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index 92f528c45605..e36936936418 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1714,6 +1714,8 @@ enum bpf_reg_arg_type {
 };
 
 #define MAX_KFUNC_DESCS 256
+/* Each kfunc can have its canonical and one specialized call target. */
+#define MAX_KFUNC_CALL_DESCS (MAX_KFUNC_DESCS * 2)
 
 struct bpf_kfunc_desc {
 	struct btf_func_model func_model;
@@ -1726,12 +1728,15 @@ struct bpf_kfunc_desc {
 
 struct bpf_kfunc_desc_tab {
 	u32 nr_descs;
+	u32 nr_base_descs;
 	/* Sorted by func_id (BTF ID) and offset (fd_array offset) during
-	 * verification. JITs do lookups by bpf_insn, where func_id may not be
-	 * available, therefore at the end of verification do_misc_fixups()
-	 * sorts this by imm and offset.
+	 * verification. The first nr_base_descs entries are the canonical
+	 * descriptors used for verifier lookups. Call specialization may append
+	 * immutable descriptors for additional targets. Near-call JITs look up
+	 * descriptors by imm and offset after do_misc_fixups() sorts the table.
 	 *
-	 * Grown one entry at a time by bpf_add_kfunc_call().
+	 * Grown one entry at a time by bpf_add_kfunc_call() and during
+	 * call specialization.
 	 */
 	struct bpf_kfunc_desc descs[];
 };
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index a7c9e2d8965d..5e7c589991e9 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -2570,7 +2570,7 @@ find_kfunc_desc(const struct bpf_prog *prog, u32 func_id, u16 offset)
 	struct bpf_kfunc_desc_tab *tab;
 
 	tab = prog->aux->kfunc_tab;
-	return bsearch(&desc, tab->descs, tab->nr_descs,
+	return bsearch(&desc, tab->descs, tab->nr_base_descs,
 		       sizeof(tab->descs[0]), kfunc_desc_cmp_by_id_off);
 }
 
@@ -2920,10 +2920,12 @@ int bpf_add_kfunc_call(struct bpf_verifier_env *env, u32 func_id, u16 offset)
 	if (find_kfunc_desc(env->prog, func_id, offset))
 		return 0;
 
-	if (tab->nr_descs == MAX_KFUNC_DESCS) {
+	if (tab->nr_base_descs == MAX_KFUNC_DESCS) {
 		verbose(env, "too many different kernel function calls\n");
 		return -E2BIG;
 	}
+	if (WARN_ON_ONCE(tab->nr_descs != tab->nr_base_descs))
+		return -EFAULT;
 
 	err = fetch_kfunc_meta(env, func_id, offset, &kfunc);
 	if (err)
@@ -2981,7 +2983,8 @@ int bpf_add_kfunc_call(struct bpf_verifier_env *env, u32 func_id, u16 offset)
 	desc->addr = addr;
 	desc->func_model = func_model;
 	tab->nr_descs++;
-	sort(tab->descs, tab->nr_descs, sizeof(tab->descs[0]),
+	tab->nr_base_descs++;
+	sort(tab->descs, tab->nr_base_descs, sizeof(tab->descs[0]),
 	     kfunc_desc_cmp_by_id_off, NULL);
 	return 0;
 }
@@ -21299,6 +21302,40 @@ static int specialize_kfunc(struct bpf_verifier_env *env, struct bpf_kfunc_desc
 	return 0;
 }
 
+static int add_kfunc_desc_target(struct bpf_verifier_env *env,
+				 const struct bpf_kfunc_desc *target_desc)
+{
+	struct bpf_kfunc_desc desc = *target_desc;
+	struct bpf_kfunc_desc_tab *new_tab;
+	struct bpf_kfunc_desc_tab *tab;
+	struct bpf_prog_aux *prog_aux;
+	u32 i;
+
+	prog_aux = env->prog->aux;
+	tab = prog_aux->kfunc_tab;
+	for (i = 0; i < tab->nr_descs; i++) {
+		if (tab->descs[i].func_id == desc.func_id &&
+		    tab->descs[i].offset == desc.offset &&
+		    tab->descs[i].addr == desc.addr)
+			return 0;
+	}
+
+	if (tab->nr_descs == MAX_KFUNC_CALL_DESCS) {
+		verbose(env, "too many different kernel function call targets\n");
+		return -E2BIG;
+	}
+
+	new_tab = krealloc(tab, struct_size(tab, descs, tab->nr_descs + 1),
+			   GFP_KERNEL_ACCOUNT);
+	if (!new_tab)
+		return -ENOMEM;
+	tab = new_tab;
+	prog_aux->kfunc_tab = tab;
+
+	tab->descs[tab->nr_descs++] = desc;
+	return 0;
+}
+
 static void __fixup_collection_insert_kfunc(struct bpf_insn_aux_data *insn_aux,
 					    u16 struct_meta_reg,
 					    u16 node_offset_reg,
@@ -21319,7 +21356,10 @@ static void __fixup_collection_insert_kfunc(struct bpf_insn_aux_data *insn_aux,
 int bpf_fixup_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
 		     struct bpf_insn *insn_buf, int insn_idx, int *cnt)
 {
+	struct bpf_kfunc_desc desc_copy;
 	struct bpf_kfunc_desc *desc;
+	unsigned long call_imm;
+	bool near_call;
 	int err;
 
 	if (!insn->imm) {
@@ -21340,12 +21380,29 @@ int bpf_fixup_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
 		return -EFAULT;
 	}
 
+	near_call = !bpf_jit_supports_far_kfunc_call();
+	if (near_call) {
+		desc_copy = *desc;
+		desc = &desc_copy;
+	}
+
 	err = specialize_kfunc(env, desc, insn_idx);
 	if (err)
 		return err;
 
-	if (!bpf_jit_supports_far_kfunc_call())
-		insn->imm = BPF_CALL_IMM(desc->addr);
+	if (near_call) {
+		call_imm = BPF_CALL_IMM(desc->addr);
+		if ((unsigned long)(s32)call_imm != call_imm) {
+			verbose(env, "address of kernel func_id %u is out of range\n",
+				desc->func_id);
+			return -EINVAL;
+		}
+		insn->imm = call_imm;
+
+		err = add_kfunc_desc_target(env, desc);
+		if (err)
+			return err;
+	}
 
 	if (is_bpf_obj_new_kfunc(desc->func_id) || is_bpf_percpu_obj_new_kfunc(desc->func_id)) {
 		struct btf_struct_meta *kptr_struct_meta = env->insn_aux_data[insn_idx].kptr_struct_meta;
-- 
2.52.0


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH bpf v3 6/7] bpf: Support call-site kfunc specialization for far calls
  2026-09-25 20:39 [PATCH bpf v3 0/7] Make sleepable arena paths use sleepable alloc_pages Emil Tsalapatis
                   ` (4 preceding siblings ...)
  2026-09-25 20:39 ` [PATCH bpf v3 5/7] bpf: Support call-site kfunc specialization for near calls Emil Tsalapatis
@ 2026-09-25 20:39 ` Emil Tsalapatis
  2026-09-25 20:39 ` [PATCH bpf v3 7/7] selftests/bpf: Test per-call site function specialization Emil Tsalapatis
  6 siblings, 0 replies; 11+ messages in thread
From: Emil Tsalapatis @ 2026-09-25 20:39 UTC (permalink / raw)
  To: bpf; +Cc: ast, andrii, eddyz87, memxor, daniel, Emil Tsalapatis

Far-call JITs keep the kfunc BTF ID in the call immediate,
so it cannot distinguish multiple specialized targets for
the same kfunc and BTF object.

Store the finalized target descriptor index in the call offset.
Use that index for address and function-model lookup, and keep
the descriptor table in verification order so the indices remain
stable.

Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
---
 include/linux/bpf.h          |  4 ++--
 include/linux/bpf_verifier.h |  2 ++
 kernel/bpf/fixups.c          | 18 ++++++++++++++----
 kernel/bpf/verifier.c        | 35 ++++++++++++++++++++++-------------
 4 files changed, 40 insertions(+), 19 deletions(-)

diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index 7df74e47ecb0..5a9580b1769d 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -3283,7 +3283,7 @@ const struct btf_func_model *
 bpf_jit_find_kfunc_model(const struct bpf_prog *prog,
 			 const struct bpf_insn *insn);
 int bpf_get_kfunc_addr(const struct bpf_prog *prog, u32 func_id,
-		       u16 btf_fd_idx, u8 **func_addr);
+		       u16 desc_idx, u8 **func_addr);
 
 struct bpf_core_ctx {
 	struct bpf_verifier_log *log;
@@ -3626,7 +3626,7 @@ bpf_jit_find_kfunc_model(const struct bpf_prog *prog,
 
 static inline int
 bpf_get_kfunc_addr(const struct bpf_prog *prog, u32 func_id,
-		   u16 btf_fd_idx, u8 **func_addr)
+		   u16 desc_idx, u8 **func_addr)
 {
 	return -ENOTSUPP;
 }
diff --git a/include/linux/bpf_verifier.h b/include/linux/bpf_verifier.h
index e36936936418..da219eb0a9cb 100644
--- a/include/linux/bpf_verifier.h
+++ b/include/linux/bpf_verifier.h
@@ -1734,6 +1734,8 @@ struct bpf_kfunc_desc_tab {
 	 * descriptors used for verifier lookups. Call specialization may append
 	 * immutable descriptors for additional targets. Near-call JITs look up
 	 * descriptors by imm and offset after do_misc_fixups() sorts the table.
+	 * Far-call JITs use the descriptor index stored in the finalized call's
+	 * off field, so their table remains in verification order.
 	 *
 	 * Grown one entry at a time by bpf_add_kfunc_call() and during
 	 * call specialization.
diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c
index 2add8001c3ec..b9f76eee5016 100644
--- a/kernel/bpf/fixups.c
+++ b/kernel/bpf/fixups.c
@@ -140,6 +140,13 @@ bpf_jit_find_kfunc_model(const struct bpf_prog *prog,
 	struct bpf_kfunc_desc_tab *tab;
 
 	tab = prog->aux->kfunc_tab;
+	if (bpf_jit_supports_far_kfunc_call()) {
+		if (insn->off < 0 || insn->off >= tab->nr_descs)
+			return NULL;
+		res = &tab->descs[insn->off];
+		return res->func_id == insn->imm ? &res->func_model : NULL;
+	}
+
 	res = bsearch(&desc, tab->descs, tab->nr_descs,
 		      sizeof(tab->descs[0]), kfunc_desc_cmp_by_imm_off);
 
@@ -2517,11 +2524,14 @@ int bpf_do_misc_fixups(struct bpf_verifier_env *env)
 		}
 	}
 
-	ret = sort_kfunc_descs_by_imm_off(env);
-	if (ret)
-		return ret;
+	/*
+	 * Do not change kfunc desc position into the table for far JIT.
+	 * because we use the indices in the instructions.
+	 */
+	if (bpf_jit_supports_far_kfunc_call())
+		return 0;
 
-	return 0;
+	return sort_kfunc_descs_by_imm_off(env);
 }
 
 static struct bpf_prog *inline_bpf_loop(struct bpf_verifier_env *env,
diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
index 5e7c589991e9..12e33a568a3e 100644
--- a/kernel/bpf/verifier.c
+++ b/kernel/bpf/verifier.c
@@ -2575,12 +2575,16 @@ find_kfunc_desc(const struct bpf_prog *prog, u32 func_id, u16 offset)
 }
 
 int bpf_get_kfunc_addr(const struct bpf_prog *prog, u32 func_id,
-		       u16 btf_fd_idx, u8 **func_addr)
+		       u16 desc_idx, u8 **func_addr)
 {
+	struct bpf_kfunc_desc_tab *tab;
 	const struct bpf_kfunc_desc *desc;
 
-	desc = find_kfunc_desc(prog, func_id, btf_fd_idx);
-	if (!desc)
+	tab = prog->aux->kfunc_tab;
+	if (desc_idx >= tab->nr_descs)
+		return -EFAULT;
+	desc = &tab->descs[desc_idx];
+	if (desc->func_id != func_id)
 		return -EFAULT;
 
 	*func_addr = (u8 *)desc->addr;
@@ -21303,7 +21307,8 @@ static int specialize_kfunc(struct bpf_verifier_env *env, struct bpf_kfunc_desc
 }
 
 static int add_kfunc_desc_target(struct bpf_verifier_env *env,
-				 const struct bpf_kfunc_desc *target_desc)
+				 const struct bpf_kfunc_desc *target_desc,
+				 u16 *desc_idx)
 {
 	struct bpf_kfunc_desc desc = *target_desc;
 	struct bpf_kfunc_desc_tab *new_tab;
@@ -21316,8 +21321,10 @@ static int add_kfunc_desc_target(struct bpf_verifier_env *env,
 	for (i = 0; i < tab->nr_descs; i++) {
 		if (tab->descs[i].func_id == desc.func_id &&
 		    tab->descs[i].offset == desc.offset &&
-		    tab->descs[i].addr == desc.addr)
+		    tab->descs[i].addr == desc.addr) {
+			*desc_idx = i;
 			return 0;
+		}
 	}
 
 	if (tab->nr_descs == MAX_KFUNC_CALL_DESCS) {
@@ -21332,6 +21339,7 @@ static int add_kfunc_desc_target(struct bpf_verifier_env *env,
 	tab = new_tab;
 	prog_aux->kfunc_tab = tab;
 
+	*desc_idx = tab->nr_descs;
 	tab->descs[tab->nr_descs++] = desc;
 	return 0;
 }
@@ -21359,6 +21367,7 @@ int bpf_fixup_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
 	struct bpf_kfunc_desc desc_copy;
 	struct bpf_kfunc_desc *desc;
 	unsigned long call_imm;
+	u16 desc_idx;
 	bool near_call;
 	int err;
 
@@ -21381,10 +21390,8 @@ int bpf_fixup_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
 	}
 
 	near_call = !bpf_jit_supports_far_kfunc_call();
-	if (near_call) {
-		desc_copy = *desc;
-		desc = &desc_copy;
-	}
+	desc_copy = *desc;
+	desc = &desc_copy;
 
 	err = specialize_kfunc(env, desc, insn_idx);
 	if (err)
@@ -21398,12 +21405,14 @@ int bpf_fixup_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn,
 			return -EINVAL;
 		}
 		insn->imm = call_imm;
-
-		err = add_kfunc_desc_target(env, desc);
-		if (err)
-			return err;
 	}
 
+	err = add_kfunc_desc_target(env, desc, &desc_idx);
+	if (err)
+		return err;
+	if (!near_call)
+		insn->off = desc_idx;
+
 	if (is_bpf_obj_new_kfunc(desc->func_id) || is_bpf_percpu_obj_new_kfunc(desc->func_id)) {
 		struct btf_struct_meta *kptr_struct_meta = env->insn_aux_data[insn_idx].kptr_struct_meta;
 		struct bpf_insn addr[2] = { BPF_LD_IMM64(BPF_REG_2, (long)kptr_struct_meta) };
-- 
2.52.0


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH bpf v3 7/7] selftests/bpf: Test per-call site function specialization
  2026-09-25 20:39 [PATCH bpf v3 0/7] Make sleepable arena paths use sleepable alloc_pages Emil Tsalapatis
                   ` (5 preceding siblings ...)
  2026-09-25 20:39 ` [PATCH bpf v3 6/7] bpf: Support call-site kfunc specialization for far calls Emil Tsalapatis
@ 2026-09-25 20:39 ` Emil Tsalapatis
  6 siblings, 0 replies; 11+ messages in thread
From: Emil Tsalapatis @ 2026-09-25 20:39 UTC (permalink / raw)
  To: bpf; +Cc: ast, andrii, eddyz87, memxor, daniel, Emil Tsalapatis

Add a test to ensure function call specialization is
done per-call site. Use bpf_dynptr_from_file that has
observably different behavior between its sleepable and
nonsleepable versions. The sleepable path fails in the
sleepable __kernel_read() call with -EIO, while the
nonsleepable fails in the page-cache lookup path with
-EFAULT. Test that whatever the order the nonsleepable
and sleepable calls are made in the program, both
call sites use the correct specialized kfunc.

Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
---
 .../selftests/bpf/prog_tests/file_reader.c    |  15 ++
 .../testing/selftests/bpf/progs/file_reader.c | 129 ++++++++++++++++++
 2 files changed, 144 insertions(+)

diff --git a/tools/testing/selftests/bpf/prog_tests/file_reader.c b/tools/testing/selftests/bpf/prog_tests/file_reader.c
index 48aae7ea0e4b..e59c6c87e9d5 100644
--- a/tools/testing/selftests/bpf/prog_tests/file_reader.c
+++ b/tools/testing/selftests/bpf/prog_tests/file_reader.c
@@ -7,10 +7,12 @@
 #include "file_reader_fail.skel.h"
 #include <dlfcn.h>
 #include <sys/mman.h>
+#include <sys/stat.h>
 
 const char *user_ptr = "hello world";
 char file_contents[256000];
 void *addr;
+__u64 beyond_eof_offset;
 
 void *get_executable_base_addr(void)
 {
@@ -26,12 +28,18 @@ void *get_executable_base_addr(void)
 
 static int initialize_file_contents(void)
 {
+	struct stat st;
 	int fd, page_sz = sysconf(_SC_PAGESIZE);
 	ssize_t n = 0, cur;
 
 	fd = open("/proc/self/exe", O_RDONLY);
 	if (!ASSERT_OK_FD(fd, "Open /proc/self/exe\n"))
 		return 1;
+	if (!ASSERT_OK(fstat(fd, &st), "fstat /proc/self/exe")) {
+		close(fd);
+		return 1;
+	}
+	beyond_eof_offset = st.st_size + (1ULL << 30);
 
 	do {
 		cur = read(fd, file_contents + n, sizeof(file_contents) - n);
@@ -75,6 +83,7 @@ static void run_test(const char *prog_name)
 
 	memcpy(skel->bss->user_buf, file_contents, sizeof(file_contents));
 	skel->bss->pid = getpid();
+	skel->bss->beyond_eof_offset = beyond_eof_offset;
 
 	err = file_reader__load(skel);
 	if (!ASSERT_OK(err, "file_reader__load"))
@@ -110,6 +119,12 @@ void test_file_reader(void)
 	if (test__start_subtest("on_open_validate_file_read"))
 		run_test("on_open_validate_file_read");
 
+	if (test__start_subtest("on_open_non_sleepable_first"))
+		run_test("on_open_non_sleepable_first");
+
+	if (test__start_subtest("on_open_sleepable_first"))
+		run_test("on_open_sleepable_first");
+
 	if (test__start_subtest("negative"))
 		RUN_TESTS(file_reader_fail);
 }
diff --git a/tools/testing/selftests/bpf/progs/file_reader.c b/tools/testing/selftests/bpf/progs/file_reader.c
index aa2c05cce2b3..8b972fd26d73 100644
--- a/tools/testing/selftests/bpf/progs/file_reader.c
+++ b/tools/testing/selftests/bpf/progs/file_reader.c
@@ -27,9 +27,14 @@ char tmp_buf[256000];
 
 int pid = 0;
 int err, run_success = 0;
+__u64 beyond_eof_offset;
 
 static int validate_file_read(struct file *file);
 static int task_work_callback(struct bpf_map *map, void *key, void *value);
+static int sleepable_second_callback(struct bpf_map *map, void *key, void *value);
+
+void bpf_rcu_read_lock(void) __ksym;
+void bpf_rcu_read_unlock(void) __ksym;
 
 SEC("lsm/file_open")
 int on_open_expect_fault(void *c)
@@ -81,6 +86,101 @@ int on_open_validate_file_read(void *c)
 	return 0;
 }
 
+/*
+ * Exercise bpf_dynptr_from_file() first from a non-sleepable LSM program and
+ * then from its sleepable task-work callback. Reading beyond EOF makes the two
+ * backing implementations return different errors.
+ */
+SEC("lsm/file_open")
+int on_open_non_sleepable_first(void *c)
+{
+	struct task_struct *task = bpf_get_current_task_btf();
+	struct bpf_dynptr dynptr;
+	struct elem *work;
+	struct file *file;
+	int key = 0;
+	int ret;
+
+	if (bpf_get_current_pid_tgid() >> 32 != pid)
+		return 0;
+
+	file = bpf_get_task_exe_file(task);
+	if (!file) {
+		err = 1;
+		return 0;
+	}
+
+	/* The non-sleepable reader cannot fault in an uncached folio. */
+	ret = bpf_dynptr_from_file(file, 0, &dynptr);
+	if (!ret)
+		ret = bpf_dynptr_read(tmp_buf, 1, &dynptr, beyond_eof_offset, 0);
+	bpf_dynptr_file_discard(&dynptr);
+	bpf_put_file(file);
+	if (ret != -EFAULT) {
+		err = 2;
+		return 0;
+	}
+
+	work = bpf_map_lookup_elem(&arrmap, &key);
+	if (!work) {
+		err = 3;
+		return 0;
+	}
+
+	ret = bpf_task_work_schedule_signal(task, &work->tw, &arrmap,
+					    sleepable_second_callback);
+	if (ret)
+		err = 4;
+	return 0;
+}
+
+/*
+ * Exercise the opposite fixup order: the first call is made from a sleepable
+ * LSM program, while the RCU read-side section makes the second non-sleepable.
+ */
+SEC("lsm.s/file_open")
+int on_open_sleepable_first(void *c)
+{
+	struct task_struct *task = bpf_get_current_task_btf();
+	struct bpf_dynptr dynptr;
+	struct file *file;
+	int ret;
+
+	if (bpf_get_current_pid_tgid() >> 32 != pid)
+		return 0;
+
+	file = bpf_get_task_exe_file(task);
+	if (!file) {
+		err = 7;
+		return 0;
+	}
+
+	ret = bpf_dynptr_from_file(file, 0, &dynptr);
+	if (!ret)
+		ret = bpf_dynptr_read(tmp_buf, 1, &dynptr, beyond_eof_offset, 0);
+	bpf_dynptr_file_discard(&dynptr);
+	if (ret != -EIO) {
+		err = 8;
+		goto out;
+	}
+
+	bpf_rcu_read_lock();
+	ret = bpf_dynptr_from_file(file, 0, &dynptr);
+	bpf_rcu_read_unlock();
+	if (!ret)
+		ret = bpf_dynptr_read(tmp_buf, 1, &dynptr, beyond_eof_offset, 0);
+	bpf_dynptr_file_discard(&dynptr);
+	if (ret != -EFAULT) {
+		err = 9;
+		goto out;
+	}
+
+	run_success = 1;
+out:
+	bpf_put_file(file);
+	return 0;
+}
+
 /* Called in a sleepable context, read 256K bytes, cross check with user space read data */
 static int task_work_callback(struct bpf_map *map, void *key, void *value)
 {
@@ -97,6 +197,35 @@ static int task_work_callback(struct bpf_map *map, void *key, void *value)
 	return 0;
 }
 
+/* Task-work callbacks are verified as sleepable. */
+static int sleepable_second_callback(struct bpf_map *map, void *key, void *value)
+{
+	struct task_struct *task = bpf_get_current_task_btf();
+	struct bpf_dynptr dynptr;
+	struct file *file;
+	int ret;
+
+	file = bpf_get_task_exe_file(task);
+	if (!file) {
+		err = 5;
+		return 0;
+	}
+
+	/* freader_fetch() converts __kernel_read()'s short read at EOF to -EIO. */
+	ret = bpf_dynptr_from_file(file, 0, &dynptr);
+	if (!ret)
+		ret = bpf_dynptr_read(tmp_buf, 1, &dynptr, beyond_eof_offset, 0);
+	bpf_dynptr_file_discard(&dynptr);
+	bpf_put_file(file);
+	if (ret != -EIO) {
+		err = 6;
+		return 0;
+	}
+
+	run_success = 1;
+	return 0;
+}
+
 static int verify_dynptr_read(struct bpf_dynptr *ptr, u32 off, char *user_buf, u32 len)
 {
 	int i;
-- 
2.52.0


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* Re: [PATCH bpf v3 3/7] bpf: Add sleepable arena page allocation path
  2026-09-25 20:39 ` [PATCH bpf v3 3/7] bpf: Add sleepable arena page allocation path Emil Tsalapatis
@ 2026-09-25 21:25   ` Alexei Starovoitov
  2026-09-25 21:57     ` Emil Tsalapatis
  0 siblings, 1 reply; 11+ messages in thread
From: Alexei Starovoitov @ 2026-09-25 21:25 UTC (permalink / raw)
  To: Emil Tsalapatis, bpf; +Cc: andrii, eddyz87, memxor, daniel

On Fri, Sep 25, 2026 at 08:39 PM Emil Tsalapatis <emil@etsalapatis.com> wrote:
> Since allocation now happens before range-tree arbitration, concurrent
> callers can allocate pages for competing requests. Bound the aggregate
> number of pages in this pre-arbitration window to the configured maximum
> size of the arena.

The pages are allocated with __GFP_ACCOUNT. memcg bounds them already.
Let's drop nr_pages_inflight. Seems like unnecessary complication.

> +out_unmap:
> +	if (sleepable)
> +		flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
>  	range_tree_set(&arena->rt, pgoff + mapped, page_cnt - mapped);
>  	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> -	if (mapped) {
> -		flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
> +	if (mapped || sleepable) {
> +		if (!sleepable)
> +			flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
>  		arena_free_pages(arena, uaddr32, mapped, sleepable);
>  	}

v2 leftovers? It doesn't depend on sleepable ?

pw-bot: cr

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH bpf v3 3/7] bpf: Add sleepable arena page allocation path
  2026-09-25 21:25   ` Alexei Starovoitov
@ 2026-09-25 21:57     ` Emil Tsalapatis
  2026-09-26  7:47       ` Alexei Starovoitov
  0 siblings, 1 reply; 11+ messages in thread
From: Emil Tsalapatis @ 2026-09-25 21:57 UTC (permalink / raw)
  To: Alexei Starovoitov, Emil Tsalapatis, bpf; +Cc: andrii, eddyz87, memxor, daniel

On Fri Sep 25, 2026 at 9:25 PM UTC, Alexei Starovoitov wrote:
> On Fri, Sep 25, 2026 at 08:39 PM Emil Tsalapatis <emil@etsalapatis.com> wrote:
>> Since allocation now happens before range-tree arbitration, concurrent
>> callers can allocate pages for competing requests. Bound the aggregate
>> number of pages in this pre-arbitration window to the configured maximum
>> size of the arena.
>
> The pages are allocated with __GFP_ACCOUNT. memcg bounds them already.
> Let's drop nr_pages_inflight. Seems like unnecessary complication.
>

The nr_pages_in_flight is there for the edge case where multiple concurrent
allocations in the arena inflate memcg usage close to the limit, then a non-arena
related allocation fails/triggers OOM. If we're fine with that happy to remove
it since it's a bit of a tacked-on hack.

I'll replace it with __GFP_RETRY_MAYFAIL on alloc_pages() so that the allocation
doesn't accidentally trigger the OOM killer. AFAICT it can do so otherwise.

>> +out_unmap:
>> +	if (sleepable)
>> +		flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
>>  	range_tree_set(&arena->rt, pgoff + mapped, page_cnt - mapped);
>>  	raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
>> -	if (mapped) {
>> -		flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
>> +	if (mapped || sleepable) {
>> +		if (!sleepable)
>> +			flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
>>  		arena_free_pages(arena, uaddr32, mapped, sleepable);
>>  	}
>
> v2 leftovers? It doesn't depend on sleepable ?

Yes, I will remove that part.

>
> pw-bot: cr


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH bpf v3 3/7] bpf: Add sleepable arena page allocation path
  2026-09-25 21:57     ` Emil Tsalapatis
@ 2026-09-26  7:47       ` Alexei Starovoitov
  0 siblings, 0 replies; 11+ messages in thread
From: Alexei Starovoitov @ 2026-09-26  7:47 UTC (permalink / raw)
  To: Emil Tsalapatis, bpf; +Cc: andrii, eddyz87, memxor, daniel

On Fri, Sep 25, 2026 at 09:57 PM Emil Tsalapatis <emil@etsalapatis.com> wrote:
> The nr_pages_in_flight is there for the edge case where multiple concurrent
> allocations in the arena inflate memcg usage close to the limit, then a non-arena
> related allocation fails/triggers OOM. If we're fine with that happy to remove
> it since it's a bit of a tacked-on hack.

Yes. Let's remove it.
An arena that is populated up to the memcg limit does the same
to other allocations in that memcg.

> I'll replace it with __GFP_RETRY_MAYFAIL on alloc_pages() so that the allocation
> doesn't accidentally trigger the OOM killer. AFAICT it can do so otherwise.

confused...
Patch 2 already adds __GFP_RETRY_MAYFAIL to bpf_alloc_page().

^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-09-26  7:47 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-25 20:39 [PATCH bpf v3 0/7] Make sleepable arena paths use sleepable alloc_pages Emil Tsalapatis
2026-09-25 20:39 ` [PATCH bpf v3 1/7] bpf: Use an llist for page allocations Emil Tsalapatis
2026-09-25 20:39 ` [PATCH bpf v3 2/7] bpf: Add sleepable argument to bpf_alloc_pages() Emil Tsalapatis
2026-09-25 20:39 ` [PATCH bpf v3 3/7] bpf: Add sleepable arena page allocation path Emil Tsalapatis
2026-09-25 21:25   ` Alexei Starovoitov
2026-09-25 21:57     ` Emil Tsalapatis
2026-09-26  7:47       ` Alexei Starovoitov
2026-09-25 20:39 ` [PATCH bpf v3 4/7] selftests/bpf: Test large allocations for both sleepable/nonsleepable arena users Emil Tsalapatis
2026-09-25 20:39 ` [PATCH bpf v3 5/7] bpf: Support call-site kfunc specialization for near calls Emil Tsalapatis
2026-09-25 20:39 ` [PATCH bpf v3 6/7] bpf: Support call-site kfunc specialization for far calls Emil Tsalapatis
2026-09-25 20:39 ` [PATCH bpf v3 7/7] selftests/bpf: Test per-call site function specialization Emil Tsalapatis

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).