* [PATCH bpf v2 0/2] Make sleepable arena paths use sleepable alloc_pages
@ 2026-09-24 5:36 Emil Tsalapatis
2026-09-24 5:36 ` [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path Emil Tsalapatis
2026-09-24 5:36 ` [PATCH bpf v2 2/2] selftests/bpf: Test large allocations for both sleepable/nonsleepable arena users Emil Tsalapatis
0 siblings, 2 replies; 9+ messages in thread
From: Emil Tsalapatis @ 2026-09-24 5:36 UTC (permalink / raw)
To: bpf; +Cc: ast, andrii, eddyz87, memxor, daniel, Emil Tsalapatis
The arena_alloc_pages() call takes a sleepable argument based on whether
its caller is a sleepable BPF function. This flag, along with the context
the kfunc is called in, decides whether the call will try to fulfill the
allocation using the regular or the _nolock variant of the alloc_pages
API, by means of bpf_map_alloc_pages().
However, the arena_alloc_pages() call currently only makes allocations
inside an IRQ-disabled critical section. This forces all allocations to
use the _nolock() API, which may eagerly fail where the regular variant
would eventually succeed. There have been reports of this happening for
sched-ext schedulers.
Restructure arena_alloc_pages() to use the _nolock() page allocation API
only when necessary. This requires moving allocations outside of the
spinlock critical section for sleepable calls, which in turn requires
slightly different logic in the allocation path. Make a separate
sleepable code path that implements this logic within
arena_alloc_pages().
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
v1 -> v2 (https://lore.kernel.org/bpf/20260824082530.47553-1-emil@etsalapatis.com/)
- Keep the sleepable and non-sleepable allocation paths within
arena_alloc_pages (Alexei)
- Incorporate bot feedback on selftests (bot-ci)
Emil Tsalapatis (2):
bpf: Add sleepable arena page allocation path
selftests/bpf: Test large allocations for both sleepable/nonsleepable
arena users
include/linux/bpf.h | 6 +
kernel/bpf/arena.c | 123 ++++++++++++------
kernel/bpf/syscall.c | 8 +-
.../bpf/progs/verifier_arena_large.c | 59 +++++++--
4 files changed, 141 insertions(+), 55 deletions(-)
--
2.52.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path
2026-09-24 5:36 [PATCH bpf v2 0/2] Make sleepable arena paths use sleepable alloc_pages Emil Tsalapatis
@ 2026-09-24 5:36 ` Emil Tsalapatis
2026-09-24 6:29 ` bot+bpf-ci
` (2 more replies)
2026-09-24 5:36 ` [PATCH bpf v2 2/2] selftests/bpf: Test large allocations for both sleepable/nonsleepable arena users Emil Tsalapatis
1 sibling, 3 replies; 9+ messages in thread
From: Emil Tsalapatis @ 2026-09-24 5:36 UTC (permalink / raw)
To: bpf; +Cc: ast, andrii, eddyz87, memxor, daniel, Emil Tsalapatis
The bpf_arena_alloc_pages() function currently only allocates pages
inside a spinlock critical section with IRQs off. This forces the use
of alloc_pages_nolock() in the BPF allocator, even when the caller is
a sleepable BPF function. This in turn causes allocation failures even
in cases where falling into the allocator slow path and possibly
sleeping would eventually succeed. This can be triggered consistently
by heavy BPF arena users like scx.
Add a separate arena page allocation path just for sleepable callers.
The path preallocates the arena memory to be added to the tree before
taking the critical section.
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
---
include/linux/bpf.h | 6 +++
kernel/bpf/arena.c | 123 ++++++++++++++++++++++++++++++-------------
kernel/bpf/syscall.c | 8 +--
3 files changed, 92 insertions(+), 45 deletions(-)
diff --git a/include/linux/bpf.h b/include/linux/bpf.h
index e7c5e203eddd..b37e34cf1086 100644
--- a/include/linux/bpf.h
+++ b/include/linux/bpf.h
@@ -720,6 +720,12 @@ void bpf_map_free_internal_structs(struct bpf_map *map, void *obj);
int bpf_dynptr_from_file_sleepable(struct file *file, u32 flags,
struct bpf_dynptr *ptr__uninit);
+static inline bool is_bpf_alloc_nonsleepable(void)
+{
+ return preempt_count() > 0 || irqs_disabled() ||
+ IS_ENABLED(CONFIG_PREEMPT_RT);
+}
+
#if defined(CONFIG_MMU) && defined(CONFIG_64BIT)
void *bpf_arena_alloc_pages_non_sleepable(void *p__map, void *addr__ign, u32 page_cnt, int node_id,
u64 flags);
diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index c6369ea5e208..eaeae16bbe7a 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -704,6 +704,27 @@ static u64 clear_lo32(u64 val)
return val & ~(u64)~0U;
}
+static int arena_adjust_tree(struct bpf_arena *arena, long uaddr, long page_cnt, long *pgoff)
+{
+ int ret;
+
+ /* Special case where user is requesting specific range. */
+ if (uaddr) {
+ ret = is_range_tree_set(&arena->rt, *pgoff, page_cnt);
+ if (ret)
+ return ret;
+ return range_tree_clear(&arena->rt, *pgoff, page_cnt);
+ }
+
+ ret = range_tree_find(&arena->rt, page_cnt);
+ if (ret < 0)
+ return ret;
+
+ *pgoff = ret;
+
+ return range_tree_clear(&arena->rt, *pgoff, page_cnt);
+}
+
/*
* Allocate pages and vmap them into kernel vmalloc area.
* Later the pages will be mmaped into user space vma.
@@ -721,7 +742,9 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
long alloc_pages;
unsigned long flags;
long pgoff = 0;
+ bool can_sleep;
u32 uaddr32;
+ long addr = 0;
int ret, i;
if (node_id != NUMA_NO_NODE &&
@@ -741,31 +764,38 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
}
bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg);
- /* Cap allocation size to KMALLOC_MAX_CACHE_SIZE so kmalloc_nolock() can succeed. */
- alloc_pages = min(page_cnt, KMALLOC_MAX_CACHE_SIZE / sizeof(struct page *));
- pages = kmalloc_nolock(alloc_pages * sizeof(struct page *), __GFP_ACCOUNT, NUMA_NO_NODE);
- if (!pages) {
- bpf_map_memcg_exit(old_memcg, new_memcg);
- return 0;
+
+ can_sleep = sleepable && !is_bpf_alloc_nonsleepable();
+ if (can_sleep) {
+ alloc_pages = page_cnt;
+ pages = kvcalloc(page_cnt, sizeof(struct page *), GFP_KERNEL_ACCOUNT);
+ if (!pages)
+ goto out_memcg;
+
+ ret = bpf_map_alloc_pages(&arena->map, node_id, page_cnt, pages);
+ if (ret)
+ goto out_free_array;
+ data.i = 0;
+ } else {
+ /* Cap allocation size so kmalloc_nolock() can succeed. */
+ alloc_pages = min(page_cnt, KMALLOC_MAX_CACHE_SIZE / sizeof(struct page *));
+ pages = kmalloc_nolock(alloc_pages * sizeof(struct page *), __GFP_ACCOUNT,
+ NUMA_NO_NODE);
+ if (!pages)
+ goto out_memcg;
}
+
data.arena = arena;
data.pages = pages;
if (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
goto out_free_pages;
- if (uaddr) {
- ret = is_range_tree_set(&arena->rt, pgoff, page_cnt);
- if (ret)
- goto out_unlock_free_pages;
- ret = range_tree_clear(&arena->rt, pgoff, page_cnt);
- } else {
- ret = pgoff = range_tree_find(&arena->rt, page_cnt);
- if (pgoff >= 0)
- ret = range_tree_clear(&arena->rt, pgoff, page_cnt);
+ ret = arena_adjust_tree(arena, uaddr, page_cnt, &pgoff);
+ if (ret) {
+ raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
+ goto out_free_pages;
}
- if (ret)
- goto out_unlock_free_pages;
remaining = page_cnt;
uaddr32 = (u32)(arena->user_vm_start + pgoff * PAGE_SIZE);
@@ -773,12 +803,14 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
while (remaining) {
long this_batch = min(remaining, alloc_pages);
- /* zeroing is needed, since alloc_pages_bulk() only fills in non-zero entries */
- memset(pages, 0, this_batch * sizeof(struct page *));
+ if (!can_sleep) {
+ /* alloc_pages_bulk() only fills in non-zero entries. */
+ memset(pages, 0, this_batch * sizeof(struct page *));
- ret = bpf_map_alloc_pages(&arena->map, node_id, this_batch, pages);
- if (ret)
- goto out;
+ ret = bpf_map_alloc_pages(&arena->map, node_id, this_batch, pages);
+ if (ret)
+ goto out_unmap;
+ }
/*
* Earlier checks made sure that uaddr32 + page_cnt * PAGE_SIZE - 1
@@ -793,35 +825,50 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
kern_vm_start + uaddr32 + (mapped << PAGE_SHIFT),
this_batch << PAGE_SHIFT, apply_range_set_cb, &data);
if (ret) {
- /* data.i pages were mapped, account them and free the remaining */
+ /* data.i pages were mapped, account them and free the remaining. */
mapped += data.i;
- for (i = data.i; i < this_batch; i++)
- free_pages_nolock(pages[i], 0);
- goto out;
+ if (!can_sleep)
+ for (i = data.i; i < this_batch; i++)
+ free_pages_nolock(pages[i], 0);
+ goto out_unmap;
}
mapped += this_batch;
remaining -= this_batch;
}
+
flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
- kfree_nolock(pages);
- bpf_map_memcg_exit(old_memcg, new_memcg);
- return clear_lo32(arena->user_vm_start) + uaddr32;
-out:
+
+ addr = clear_lo32(arena->user_vm_start) + uaddr32;
+ goto out_free_array;
+
+out_unmap:
+ if (can_sleep)
+ flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
range_tree_set(&arena->rt, pgoff + mapped, page_cnt - mapped);
raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
- if (mapped) {
- flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
- arena_free_pages(arena, uaddr32, mapped, sleepable);
+ if (mapped || can_sleep) {
+ if (!can_sleep)
+ flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
+ arena_free_pages(arena, uaddr32, mapped, can_sleep);
}
- goto out_free_pages;
-out_unlock_free_pages:
- raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
+
out_free_pages:
- kfree_nolock(pages);
+ if (can_sleep)
+ for (i = data.i; i < page_cnt; i++)
+ __free_page(pages[i]);
+
+out_free_array:
+ if (can_sleep)
+ kvfree(pages);
+ else
+ kfree_nolock(pages);
+
+out_memcg:
bpf_map_memcg_exit(old_memcg, new_memcg);
- return 0;
+
+ return addr;
}
/*
diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
index 74496fd716d3..d193d84cc33e 100644
--- a/kernel/bpf/syscall.c
+++ b/kernel/bpf/syscall.c
@@ -596,15 +596,9 @@ static void bpf_map_release_memcg(struct bpf_map *map)
}
#endif
-static bool can_alloc_pages(void)
-{
- return preempt_count() == 0 && !irqs_disabled() &&
- !IS_ENABLED(CONFIG_PREEMPT_RT);
-}
-
static struct page *__bpf_alloc_page(int nid)
{
- if (!can_alloc_pages())
+ if (is_bpf_alloc_nonsleepable())
return alloc_pages_nolock(__GFP_ACCOUNT, nid, 0);
return alloc_pages_node(nid,
--
2.52.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* [PATCH bpf v2 2/2] selftests/bpf: Test large allocations for both sleepable/nonsleepable arena users
2026-09-24 5:36 [PATCH bpf v2 0/2] Make sleepable arena paths use sleepable alloc_pages Emil Tsalapatis
2026-09-24 5:36 ` [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path Emil Tsalapatis
@ 2026-09-24 5:36 ` Emil Tsalapatis
1 sibling, 0 replies; 9+ messages in thread
From: Emil Tsalapatis @ 2026-09-24 5:36 UTC (permalink / raw)
To: bpf; +Cc: ast, andrii, eddyz87, memxor, daniel, Emil Tsalapatis
We now use different code paths in the internal allocator when
allocating arena memory, depending on whether the caller is sleepable
or not. These paths mostly differ functionally for large allocations,
so add extra testing for that case.
Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
---
.../bpf/progs/verifier_arena_large.c | 59 +++++++++++++++----
1 file changed, 49 insertions(+), 10 deletions(-)
diff --git a/tools/testing/selftests/bpf/progs/verifier_arena_large.c b/tools/testing/selftests/bpf/progs/verifier_arena_large.c
index 6ab8730d4878..e002815b6929 100644
--- a/tools/testing/selftests/bpf/progs/verifier_arena_large.c
+++ b/tools/testing/selftests/bpf/progs/verifier_arena_large.c
@@ -11,6 +11,8 @@
#define ARENA_SIZE (1ull << 32)
+volatile int zero = 0;
+
struct {
__uint(type, BPF_MAP_TYPE_ARENA);
__uint(map_flags, BPF_F_MMAPABLE);
@@ -284,6 +286,7 @@ int big_alloc2(void *ctx)
return 0;
}
+/* Nonsleepable because it binds to a socket program. */
SEC("socket")
__success __retval(0)
int big_alloc3(void *ctx)
@@ -291,24 +294,60 @@ int big_alloc3(void *ctx)
#if defined(__BPF_FEATURE_ADDR_SPACE_CAST)
char __arena *pages;
u64 i;
+ int err = 0;
/*
- * Allocate 2051 pages in one go to check how kmalloc_nolock() handles large requests.
- * Since kmalloc_nolock() can allocate up to 1024 struct page * at a time, this call should
- * result in three batches: two batches of 1024 pages each, followed by a final batch of 3
- * pages.
+ * Allocate 1025 pages in one go to check how kmalloc_nolock() handles large requests.
+ * Since kmalloc_nolock() can allocate up to 1024 struct page * at a time, this is the
+ * smallest request that exercises multiple batches, limiting the time spent with IRQs
+ * disabled.
*/
+ pages = bpf_arena_alloc_pages(&arena, NULL, 1025, NUMA_NO_NODE, 0);
+ if (!pages)
+ return 1;
+
+ for (i = zero; i < 1025 && can_loop; i++)
+ pages[i * PAGE_SIZE] = 123;
+
+ for (i = zero; i < 1025 && can_loop; i++) {
+ if (pages[i * PAGE_SIZE] == 123)
+ continue;
+ err = 2;
+ break;
+ }
+
+ bpf_arena_free_pages(&arena, pages, 1025);
+ return err;
+#endif
+ return 0;
+}
+
+/* SYSCALL programs are always sleepable. */
+SEC("syscall")
+__success __retval(0)
+int big_alloc4(void *ctx)
+{
+#if defined(__BPF_FEATURE_ADDR_SPACE_CAST)
+ char __arena *pages;
+ u64 i;
+ int err = 0;
+
pages = bpf_arena_alloc_pages(&arena, NULL, 2051, NUMA_NO_NODE, 0);
if (!pages)
- return 0;
+ return 1;
- bpf_for(i, 0, 2051)
- pages[i * PAGE_SIZE] = 123;
- bpf_for(i, 0, 2051)
- if (pages[i * PAGE_SIZE] != 123)
- return i;
+ for (i = zero; i < 2051 && can_loop; i++)
+ pages[i * PAGE_SIZE] = 123;
+
+ for (i = zero; i < 2051 && can_loop; i++) {
+ if (pages[i * PAGE_SIZE] == 123)
+ continue;
+ err = 2;
+ break;
+ }
bpf_arena_free_pages(&arena, pages, 2051);
+ return err;
#endif
return 0;
}
--
2.52.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* Re: [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path
2026-09-24 5:36 ` [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path Emil Tsalapatis
@ 2026-09-24 6:29 ` bot+bpf-ci
2026-09-24 18:29 ` Emil Tsalapatis
2026-09-24 10:11 ` Jiayuan Chen
2026-09-24 21:33 ` Alexei Starovoitov
2 siblings, 1 reply; 9+ messages in thread
From: bot+bpf-ci @ 2026-09-24 6:29 UTC (permalink / raw)
To: emil, bpf
Cc: ast, andrii, eddyz87, memxor, daniel, emil, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 3969 bytes --]
> diff --git a/include/linux/bpf.h b/include/linux/bpf.h
> index 1d2676782d70..71688190a7f0 100644
> --- a/include/linux/bpf.h
> +++ b/include/linux/bpf.h
[ ... ]
> diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
> index 7b6847200b43..013faacc4551 100644
> --- a/kernel/bpf/arena.c
> +++ b/kernel/bpf/arena.c
[ ... ]
> @@ -698,31 +721,38 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
> }
>
> bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg);
> - /* Cap allocation size to KMALLOC_MAX_CACHE_SIZE so kmalloc_nolock() can succeed. */
> - alloc_pages = min(page_cnt, KMALLOC_MAX_CACHE_SIZE / sizeof(struct page *));
> - pages = kmalloc_nolock(alloc_pages * sizeof(struct page *), __GFP_ACCOUNT, NUMA_NO_NODE);
> - if (!pages) {
> - bpf_map_memcg_exit(old_memcg, new_memcg);
> - return 0;
> +
> + can_sleep = sleepable && !is_bpf_alloc_nonsleepable();
> + if (can_sleep) {
> + alloc_pages = page_cnt;
> + pages = kvcalloc(page_cnt, sizeof(struct page *), GFP_KERNEL_ACCOUNT);
> + if (!pages)
> + goto out_memcg;
> +
> + ret = bpf_map_alloc_pages(&arena->map, node_id, page_cnt, pages);
> + if (ret)
> + goto out_free_array;
> + data.i = 0;
> + } else {
> + /* Cap allocation size so kmalloc_nolock() can succeed. */
> + alloc_pages = min(page_cnt, KMALLOC_MAX_CACHE_SIZE / sizeof(struct page *));
> + pages = kmalloc_nolock(alloc_pages * sizeof(struct page *), __GFP_ACCOUNT,
> + NUMA_NO_NODE);
> + if (!pages)
> + goto out_memcg;
> }
> +
> data.arena = arena;
> data.pages = pages;
>
> if (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
> goto out_free_pages;
>
> - if (uaddr) {
> - ret = is_range_tree_set(&arena->rt, pgoff, page_cnt);
> - if (ret)
> - goto out_unlock_free_pages;
> - ret = range_tree_clear(&arena->rt, pgoff, page_cnt);
> - } else {
> - ret = pgoff = range_tree_find(&arena->rt, page_cnt);
> - if (pgoff >= 0)
> - ret = range_tree_clear(&arena->rt, pgoff, page_cnt);
> + ret = arena_adjust_tree(arena, uaddr, page_cnt, &pgoff);
> + if (ret) {
> + raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> + goto out_free_pages;
> }
> - if (ret)
> - goto out_unlock_free_pages;
In the sleepable path, the code now preallocates the entire pointer
array via kvcalloc(page_cnt, ..., GFP_KERNEL_ACCOUNT) and all page_cnt
pages via bpf_map_alloc_pages() before checking whether the requested
arena range is actually available.
The range check happens later in arena_adjust_tree() under the spinlock.
That check rejects requests where a uaddr hint overlaps already
allocated pages (is_range_tree_set() returns -ESRCH) or where no
contiguous free range exists (range_tree_find() returns -ENOENT).
The old code ran these checks under the lock before allocating anything,
so rejected requests returned NULL at minimal cost. With the new
sleepable path, each rejected request allocates and zeroes all the pages
first via GFP_KERNEL allocations, potentially pushing the system into
direct reclaim or OOM, then frees them all through the out_free_pages
loop.
Can a sleepable program that retries a fixed-address allocation or
probes an almost-full arena repeatedly cause large reclaim and OOM
pressure for requests that were never going to succeed?
One approach would be to reserve the range under the lock first via
arena_adjust_tree() and allocate afterwards, restoring the range with
range_tree_set() if allocation fails. Another option would be an
unlocked or locked precheck against the range tree before
preallocating.
[ ... ]
> diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
> index 244a939b9d2d..9eeb65f2d929 100644
> --- a/kernel/bpf/syscall.c
> +++ b/kernel/bpf/syscall.c
[ ... ]
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35961888778
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path
2026-09-24 5:36 ` [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path Emil Tsalapatis
2026-09-24 6:29 ` bot+bpf-ci
@ 2026-09-24 10:11 ` Jiayuan Chen
2026-09-24 17:56 ` Emil Tsalapatis
2026-09-24 21:33 ` Alexei Starovoitov
2 siblings, 1 reply; 9+ messages in thread
From: Jiayuan Chen @ 2026-09-24 10:11 UTC (permalink / raw)
To: Emil Tsalapatis, bpf; +Cc: ast, andrii, eddyz87, memxor, daniel
On 9/24/26 1:36 PM, Emil Tsalapatis wrote:
> The bpf_arena_alloc_pages() function currently only allocates pages
> inside a spinlock critical section with IRQs off. This forces the use
> of alloc_pages_nolock() in the BPF allocator, even when the caller is
> a sleepable BPF function. This in turn causes allocation failures even
> in cases where falling into the allocator slow path and possibly
> sleeping would eventually succeed. This can be triggered consistently
> by heavy BPF arena users like scx.
>
> Add a separate arena page allocation path just for sleepable callers.
> The path preallocates the arena memory to be added to the tree before
> taking the critical section.
>
> Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
> ---
> include/linux/bpf.h | 6 +++
> kernel/bpf/arena.c | 123 ++++++++++++++++++++++++++++++-------------
> kernel/bpf/syscall.c | 8 +--
> 3 files changed, 92 insertions(+), 45 deletions(-)
>
> diff --git a/include/linux/bpf.h b/include/linux/bpf.h
> index e7c5e203eddd..b37e34cf1086 100644
> --- a/include/linux/bpf.h
> +++ b/include/linux/bpf.h
> @@ -720,6 +720,12 @@ void bpf_map_free_internal_structs(struct bpf_map *map, void *obj);
> int bpf_dynptr_from_file_sleepable(struct file *file, u32 flags,
> struct bpf_dynptr *ptr__uninit);
>
> +static inline bool is_bpf_alloc_nonsleepable(void)
> +{
> + return preempt_count() > 0 || irqs_disabled() ||
> + IS_ENABLED(CONFIG_PREEMPT_RT);
> +}
> +
A runtime check cannot tell whether sleeping is allowed: without
CONFIG_PREEMPT_COUNT preempt_count() does not see spinlocks or
rcu_read_lock(), which is why preemptible() is 0 there.
How about passing 'sleepable' down to bpf_map_alloc_pages() instead?
The verifier already proves it, and the code gets simpler.
> #if defined(CONFIG_MMU) && defined(CONFIG_64BIT)
> void *bpf_arena_alloc_pages_non_sleepable(void *p__map, void *addr__ign, u32 page_cnt, int node_id,
> u64 flags);
> diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
> index c6369ea5e208..eaeae16bbe7a 100644
> --- a/kernel/bpf/arena.c
> +++ b/kernel/bpf/arena.c
> @@ -704,6 +704,27 @@ static u64 clear_lo32(u64 val)
> return val & ~(u64)~0U;
> }
>
> +static int arena_adjust_tree(struct bpf_arena *arena, long uaddr, long page_cnt, long *pgoff)
> +{
> + int ret;
> +
> + /* Special case where user is requesting specific range. */
> + if (uaddr) {
> + ret = is_range_tree_set(&arena->rt, *pgoff, page_cnt);
> + if (ret)
> + return ret;
> + return range_tree_clear(&arena->rt, *pgoff, page_cnt);
> + }
> +
> + ret = range_tree_find(&arena->rt, page_cnt);
> + if (ret < 0)
> + return ret;
> +
> + *pgoff = ret;
> +
> + return range_tree_clear(&arena->rt, *pgoff, page_cnt);
> +}
> +
> /*
> * Allocate pages and vmap them into kernel vmalloc area.
> * Later the pages will be mmaped into user space vma.
> @@ -721,7 +742,9 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
> long alloc_pages;
> unsigned long flags;
> long pgoff = 0;
> + bool can_sleep;
> u32 uaddr32;
> + long addr = 0;
> int ret, i;
>
> if (node_id != NUMA_NO_NODE &&
> @@ -741,31 +764,38 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
> }
>
> bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg);
> - /* Cap allocation size to KMALLOC_MAX_CACHE_SIZE so kmalloc_nolock() can succeed. */
> - alloc_pages = min(page_cnt, KMALLOC_MAX_CACHE_SIZE / sizeof(struct page *));
> - pages = kmalloc_nolock(alloc_pages * sizeof(struct page *), __GFP_ACCOUNT, NUMA_NO_NODE);
> - if (!pages) {
> - bpf_map_memcg_exit(old_memcg, new_memcg);
> - return 0;
> +
> + can_sleep = sleepable && !is_bpf_alloc_nonsleepable();
> + if (can_sleep) {
> + alloc_pages = page_cnt;
> + pages = kvcalloc(page_cnt, sizeof(struct page *), GFP_KERNEL_ACCOUNT);
> + if (!pages)
> + goto out_memcg;
> +
> + ret = bpf_map_alloc_pages(&arena->map, node_id, page_cnt, pages);
> + if (ret)
> + goto out_free_array;
> + data.i = 0;
> + } else {
> + /* Cap allocation size so kmalloc_nolock() can succeed. */
> + alloc_pages = min(page_cnt, KMALLOC_MAX_CACHE_SIZE / sizeof(struct page *));
> + pages = kmalloc_nolock(alloc_pages * sizeof(struct page *), __GFP_ACCOUNT,
> + NUMA_NO_NODE);
> + if (!pages)
> + goto out_memcg;
> }
> +
> data.arena = arena;
> data.pages = pages;
>
> if (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
> goto out_free_pages;
>
> - if (uaddr) {
> - ret = is_range_tree_set(&arena->rt, pgoff, page_cnt);
> - if (ret)
> - goto out_unlock_free_pages;
> - ret = range_tree_clear(&arena->rt, pgoff, page_cnt);
> - } else {
> - ret = pgoff = range_tree_find(&arena->rt, page_cnt);
> - if (pgoff >= 0)
> - ret = range_tree_clear(&arena->rt, pgoff, page_cnt);
> + ret = arena_adjust_tree(arena, uaddr, page_cnt, &pgoff);
> + if (ret) {
> + raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> + goto out_free_pages;
> }
> - if (ret)
> - goto out_unlock_free_pages;
>
> remaining = page_cnt;
> uaddr32 = (u32)(arena->user_vm_start + pgoff * PAGE_SIZE);
> @@ -773,12 +803,14 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
> while (remaining) {
> long this_batch = min(remaining, alloc_pages);
>
> - /* zeroing is needed, since alloc_pages_bulk() only fills in non-zero entries */
> - memset(pages, 0, this_batch * sizeof(struct page *));
> + if (!can_sleep) {
> + /* alloc_pages_bulk() only fills in non-zero entries. */
> + memset(pages, 0, this_batch * sizeof(struct page *));
>
> - ret = bpf_map_alloc_pages(&arena->map, node_id, this_batch, pages);
> - if (ret)
> - goto out;
> + ret = bpf_map_alloc_pages(&arena->map, node_id, this_batch, pages);
> + if (ret)
> + goto out_unmap;
> + }
>
> /*
> * Earlier checks made sure that uaddr32 + page_cnt * PAGE_SIZE - 1
> @@ -793,35 +825,50 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
> kern_vm_start + uaddr32 + (mapped << PAGE_SHIFT),
> this_batch << PAGE_SHIFT, apply_range_set_cb, &data);
> if (ret) {
> - /* data.i pages were mapped, account them and free the remaining */
> + /* data.i pages were mapped, account them and free the remaining. */
> mapped += data.i;
> - for (i = data.i; i < this_batch; i++)
> - free_pages_nolock(pages[i], 0);
> - goto out;
> + if (!can_sleep)
> + for (i = data.i; i < this_batch; i++)
> + free_pages_nolock(pages[i], 0);
> + goto out_unmap;
> }
>
> mapped += this_batch;
> remaining -= this_batch;
> }
> +
> flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
> raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> - kfree_nolock(pages);
> - bpf_map_memcg_exit(old_memcg, new_memcg);
> - return clear_lo32(arena->user_vm_start) + uaddr32;
> -out:
> +
> + addr = clear_lo32(arena->user_vm_start) + uaddr32;
> + goto out_free_array;
> +
> +out_unmap:
> + if (can_sleep)
> + flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
> range_tree_set(&arena->rt, pgoff + mapped, page_cnt - mapped);
> raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> - if (mapped) {
> - flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
> - arena_free_pages(arena, uaddr32, mapped, sleepable);
> + if (mapped || can_sleep) {
> + if (!can_sleep)
> + flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
> + arena_free_pages(arena, uaddr32, mapped, can_sleep);
> }
> - goto out_free_pages;
> -out_unlock_free_pages:
> - raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
> +
> out_free_pages:
> - kfree_nolock(pages);
> + if (can_sleep)
> + for (i = data.i; i < page_cnt; i++)
> + __free_page(pages[i]);
> +
> +out_free_array:
> + if (can_sleep)
> + kvfree(pages);
> + else
> + kfree_nolock(pages);
> +
> +out_memcg:
> bpf_map_memcg_exit(old_memcg, new_memcg);
> - return 0;
> +
> + return addr;
> }
>
> /*
> diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
> index 74496fd716d3..d193d84cc33e 100644
> --- a/kernel/bpf/syscall.c
> +++ b/kernel/bpf/syscall.c
> @@ -596,15 +596,9 @@ static void bpf_map_release_memcg(struct bpf_map *map)
> }
> #endif
>
> -static bool can_alloc_pages(void)
> -{
> - return preempt_count() == 0 && !irqs_disabled() &&
> - !IS_ENABLED(CONFIG_PREEMPT_RT);
> -}
> -
> static struct page *__bpf_alloc_page(int nid)
> {
> - if (!can_alloc_pages())
> + if (is_bpf_alloc_nonsleepable())
> return alloc_pages_nolock(__GFP_ACCOUNT, nid, 0);
>
> return alloc_pages_node(nid,
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path
2026-09-24 10:11 ` Jiayuan Chen
@ 2026-09-24 17:56 ` Emil Tsalapatis
0 siblings, 0 replies; 9+ messages in thread
From: Emil Tsalapatis @ 2026-09-24 17:56 UTC (permalink / raw)
To: Jiayuan Chen, Emil Tsalapatis, bpf; +Cc: ast, andrii, eddyz87, memxor, daniel
On Thu Sep 24, 2026 at 10:11 AM UTC, Jiayuan Chen wrote:
>
> On 9/24/26 1:36 PM, Emil Tsalapatis wrote:
>> The bpf_arena_alloc_pages() function currently only allocates pages
>> inside a spinlock critical section with IRQs off. This forces the use
>> of alloc_pages_nolock() in the BPF allocator, even when the caller is
>> a sleepable BPF function. This in turn causes allocation failures even
>> in cases where falling into the allocator slow path and possibly
>> sleeping would eventually succeed. This can be triggered consistently
>> by heavy BPF arena users like scx.
>>
>> Add a separate arena page allocation path just for sleepable callers.
>> The path preallocates the arena memory to be added to the tree before
>> taking the critical section.
>>
>> Signed-off-by: Emil Tsalapatis <emil@etsalapatis.com>
>> ---
>> include/linux/bpf.h | 6 +++
>> kernel/bpf/arena.c | 123 ++++++++++++++++++++++++++++++-------------
>> kernel/bpf/syscall.c | 8 +--
>> 3 files changed, 92 insertions(+), 45 deletions(-)
>>
>> diff --git a/include/linux/bpf.h b/include/linux/bpf.h
>> index e7c5e203eddd..b37e34cf1086 100644
>> --- a/include/linux/bpf.h
>> +++ b/include/linux/bpf.h
>> @@ -720,6 +720,12 @@ void bpf_map_free_internal_structs(struct bpf_map *map, void *obj);
>> int bpf_dynptr_from_file_sleepable(struct file *file, u32 flags,
>> struct bpf_dynptr *ptr__uninit);
>>
>> +static inline bool is_bpf_alloc_nonsleepable(void)
>> +{
>> + return preempt_count() > 0 || irqs_disabled() ||
>> + IS_ENABLED(CONFIG_PREEMPT_RT);
>> +}
>> +
>
>
> A runtime check cannot tell whether sleeping is allowed: without
> CONFIG_PREEMPT_COUNT preempt_count() does not see spinlocks or
> rcu_read_lock(), which is why preemptible() is 0 there.
>
> How about passing 'sleepable' down to bpf_map_alloc_pages() instead?
> The verifier already proves it, and the code gets simpler.
Sure, we can keep it there. Afaict what you're proposing is to keep
the can_alloc_pages() call where it is and pass the sleepable flag
down to bpf_map_alloc_pages() instead, where now we're hoisting
can_alloc_pages() up.
>
>
>> #if defined(CONFIG_MMU) && defined(CONFIG_64BIT)
>> void *bpf_arena_alloc_pages_non_sleepable(void *p__map, void *addr__ign, u32 page_cnt, int node_id,
>> u64 flags);
>> diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
>> index c6369ea5e208..eaeae16bbe7a 100644
>> --- a/kernel/bpf/arena.c
>> +++ b/kernel/bpf/arena.c
>> @@ -704,6 +704,27 @@ static u64 clear_lo32(u64 val)
>> return val & ~(u64)~0U;
>> }
>>
>> +static int arena_adjust_tree(struct bpf_arena *arena, long uaddr, long page_cnt, long *pgoff)
>> +{
>> + int ret;
>> +
>> + /* Special case where user is requesting specific range. */
>> + if (uaddr) {
>> + ret = is_range_tree_set(&arena->rt, *pgoff, page_cnt);
>> + if (ret)
>> + return ret;
>> + return range_tree_clear(&arena->rt, *pgoff, page_cnt);
>> + }
>> +
>> + ret = range_tree_find(&arena->rt, page_cnt);
>> + if (ret < 0)
>> + return ret;
>> +
>> + *pgoff = ret;
>> +
>> + return range_tree_clear(&arena->rt, *pgoff, page_cnt);
>> +}
>> +
>> /*
>> * Allocate pages and vmap them into kernel vmalloc area.
>> * Later the pages will be mmaped into user space vma.
>> @@ -721,7 +742,9 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
>> long alloc_pages;
>> unsigned long flags;
>> long pgoff = 0;
>> + bool can_sleep;
>> u32 uaddr32;
>> + long addr = 0;
>> int ret, i;
>>
>> if (node_id != NUMA_NO_NODE &&
>> @@ -741,31 +764,38 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
>> }
>>
>> bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg);
>> - /* Cap allocation size to KMALLOC_MAX_CACHE_SIZE so kmalloc_nolock() can succeed. */
>> - alloc_pages = min(page_cnt, KMALLOC_MAX_CACHE_SIZE / sizeof(struct page *));
>> - pages = kmalloc_nolock(alloc_pages * sizeof(struct page *), __GFP_ACCOUNT, NUMA_NO_NODE);
>> - if (!pages) {
>> - bpf_map_memcg_exit(old_memcg, new_memcg);
>> - return 0;
>> +
>> + can_sleep = sleepable && !is_bpf_alloc_nonsleepable();
>> + if (can_sleep) {
>> + alloc_pages = page_cnt;
>> + pages = kvcalloc(page_cnt, sizeof(struct page *), GFP_KERNEL_ACCOUNT);
>> + if (!pages)
>> + goto out_memcg;
>> +
>> + ret = bpf_map_alloc_pages(&arena->map, node_id, page_cnt, pages);
>> + if (ret)
>> + goto out_free_array;
>> + data.i = 0;
>> + } else {
>> + /* Cap allocation size so kmalloc_nolock() can succeed. */
>> + alloc_pages = min(page_cnt, KMALLOC_MAX_CACHE_SIZE / sizeof(struct page *));
>> + pages = kmalloc_nolock(alloc_pages * sizeof(struct page *), __GFP_ACCOUNT,
>> + NUMA_NO_NODE);
>> + if (!pages)
>> + goto out_memcg;
>> }
>> +
>> data.arena = arena;
>> data.pages = pages;
>>
>> if (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
>> goto out_free_pages;
>>
>> - if (uaddr) {
>> - ret = is_range_tree_set(&arena->rt, pgoff, page_cnt);
>> - if (ret)
>> - goto out_unlock_free_pages;
>> - ret = range_tree_clear(&arena->rt, pgoff, page_cnt);
>> - } else {
>> - ret = pgoff = range_tree_find(&arena->rt, page_cnt);
>> - if (pgoff >= 0)
>> - ret = range_tree_clear(&arena->rt, pgoff, page_cnt);
>> + ret = arena_adjust_tree(arena, uaddr, page_cnt, &pgoff);
>> + if (ret) {
>> + raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
>> + goto out_free_pages;
>> }
>> - if (ret)
>> - goto out_unlock_free_pages;
>>
>> remaining = page_cnt;
>> uaddr32 = (u32)(arena->user_vm_start + pgoff * PAGE_SIZE);
>> @@ -773,12 +803,14 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
>> while (remaining) {
>> long this_batch = min(remaining, alloc_pages);
>>
>> - /* zeroing is needed, since alloc_pages_bulk() only fills in non-zero entries */
>> - memset(pages, 0, this_batch * sizeof(struct page *));
>> + if (!can_sleep) {
>> + /* alloc_pages_bulk() only fills in non-zero entries. */
>> + memset(pages, 0, this_batch * sizeof(struct page *));
>>
>> - ret = bpf_map_alloc_pages(&arena->map, node_id, this_batch, pages);
>> - if (ret)
>> - goto out;
>> + ret = bpf_map_alloc_pages(&arena->map, node_id, this_batch, pages);
>> + if (ret)
>> + goto out_unmap;
>> + }
>>
>> /*
>> * Earlier checks made sure that uaddr32 + page_cnt * PAGE_SIZE - 1
>> @@ -793,35 +825,50 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
>> kern_vm_start + uaddr32 + (mapped << PAGE_SHIFT),
>> this_batch << PAGE_SHIFT, apply_range_set_cb, &data);
>> if (ret) {
>> - /* data.i pages were mapped, account them and free the remaining */
>> + /* data.i pages were mapped, account them and free the remaining. */
>> mapped += data.i;
>> - for (i = data.i; i < this_batch; i++)
>> - free_pages_nolock(pages[i], 0);
>> - goto out;
>> + if (!can_sleep)
>> + for (i = data.i; i < this_batch; i++)
>> + free_pages_nolock(pages[i], 0);
>> + goto out_unmap;
>> }
>>
>> mapped += this_batch;
>> remaining -= this_batch;
>> }
>> +
>> flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
>> raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
>> - kfree_nolock(pages);
>> - bpf_map_memcg_exit(old_memcg, new_memcg);
>> - return clear_lo32(arena->user_vm_start) + uaddr32;
>> -out:
>> +
>> + addr = clear_lo32(arena->user_vm_start) + uaddr32;
>> + goto out_free_array;
>> +
>> +out_unmap:
>> + if (can_sleep)
>> + flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
>> range_tree_set(&arena->rt, pgoff + mapped, page_cnt - mapped);
>> raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
>> - if (mapped) {
>> - flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
>> - arena_free_pages(arena, uaddr32, mapped, sleepable);
>> + if (mapped || can_sleep) {
>> + if (!can_sleep)
>> + flush_vmap_cache(kern_vm_start + uaddr32, mapped << PAGE_SHIFT);
>> + arena_free_pages(arena, uaddr32, mapped, can_sleep);
>> }
>> - goto out_free_pages;
>> -out_unlock_free_pages:
>> - raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
>> +
>> out_free_pages:
>> - kfree_nolock(pages);
>> + if (can_sleep)
>> + for (i = data.i; i < page_cnt; i++)
>> + __free_page(pages[i]);
>> +
>> +out_free_array:
>> + if (can_sleep)
>> + kvfree(pages);
>> + else
>> + kfree_nolock(pages);
>> +
>> +out_memcg:
>> bpf_map_memcg_exit(old_memcg, new_memcg);
>> - return 0;
>> +
>> + return addr;
>> }
>>
>> /*
>> diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
>> index 74496fd716d3..d193d84cc33e 100644
>> --- a/kernel/bpf/syscall.c
>> +++ b/kernel/bpf/syscall.c
>> @@ -596,15 +596,9 @@ static void bpf_map_release_memcg(struct bpf_map *map)
>> }
>> #endif
>>
>> -static bool can_alloc_pages(void)
>> -{
>> - return preempt_count() == 0 && !irqs_disabled() &&
>> - !IS_ENABLED(CONFIG_PREEMPT_RT);
>> -}
>> -
>> static struct page *__bpf_alloc_page(int nid)
>> {
>> - if (!can_alloc_pages())
>> + if (is_bpf_alloc_nonsleepable())
>> return alloc_pages_nolock(__GFP_ACCOUNT, nid, 0);
>>
>> return alloc_pages_node(nid,
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path
2026-09-24 6:29 ` bot+bpf-ci
@ 2026-09-24 18:29 ` Emil Tsalapatis
0 siblings, 0 replies; 9+ messages in thread
From: Emil Tsalapatis @ 2026-09-24 18:29 UTC (permalink / raw)
To: bot+bpf-ci, emil, bpf
Cc: ast, andrii, eddyz87, memxor, daniel, martin.lau, yonghong.song,
mason, ihor.solodrai
On Thu Sep 24, 2026 at 6:29 AM UTC, bot+bpf-ci wrote:
>> diff --git a/include/linux/bpf.h b/include/linux/bpf.h
>> index 1d2676782d70..71688190a7f0 100644
>> --- a/include/linux/bpf.h
>> +++ b/include/linux/bpf.h
>
> [ ... ]
>
>> diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
>> index 7b6847200b43..013faacc4551 100644
>> --- a/kernel/bpf/arena.c
>> +++ b/kernel/bpf/arena.c
>
> [ ... ]
>
>> @@ -698,31 +721,38 @@ static long arena_alloc_pages(struct bpf_arena *arena, long uaddr, long page_cnt
>> }
>>
>> bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg);
>> - /* Cap allocation size to KMALLOC_MAX_CACHE_SIZE so kmalloc_nolock() can succeed. */
>> - alloc_pages = min(page_cnt, KMALLOC_MAX_CACHE_SIZE / sizeof(struct page *));
>> - pages = kmalloc_nolock(alloc_pages * sizeof(struct page *), __GFP_ACCOUNT, NUMA_NO_NODE);
>> - if (!pages) {
>> - bpf_map_memcg_exit(old_memcg, new_memcg);
>> - return 0;
>> +
>> + can_sleep = sleepable && !is_bpf_alloc_nonsleepable();
>> + if (can_sleep) {
>> + alloc_pages = page_cnt;
>> + pages = kvcalloc(page_cnt, sizeof(struct page *), GFP_KERNEL_ACCOUNT);
>> + if (!pages)
>> + goto out_memcg;
>> +
>> + ret = bpf_map_alloc_pages(&arena->map, node_id, page_cnt, pages);
>> + if (ret)
>> + goto out_free_array;
>> + data.i = 0;
>> + } else {
>> + /* Cap allocation size so kmalloc_nolock() can succeed. */
>> + alloc_pages = min(page_cnt, KMALLOC_MAX_CACHE_SIZE / sizeof(struct page *));
>> + pages = kmalloc_nolock(alloc_pages * sizeof(struct page *), __GFP_ACCOUNT,
>> + NUMA_NO_NODE);
>> + if (!pages)
>> + goto out_memcg;
>> }
>> +
>> data.arena = arena;
>> data.pages = pages;
>>
>> if (raw_res_spin_lock_irqsave(&arena->spinlock, flags))
>> goto out_free_pages;
>>
>> - if (uaddr) {
>> - ret = is_range_tree_set(&arena->rt, pgoff, page_cnt);
>> - if (ret)
>> - goto out_unlock_free_pages;
>> - ret = range_tree_clear(&arena->rt, pgoff, page_cnt);
>> - } else {
>> - ret = pgoff = range_tree_find(&arena->rt, page_cnt);
>> - if (pgoff >= 0)
>> - ret = range_tree_clear(&arena->rt, pgoff, page_cnt);
>> + ret = arena_adjust_tree(arena, uaddr, page_cnt, &pgoff);
>> + if (ret) {
>> + raw_res_spin_unlock_irqrestore(&arena->spinlock, flags);
>> + goto out_free_pages;
>> }
>> - if (ret)
>> - goto out_unlock_free_pages;
>
> In the sleepable path, the code now preallocates the entire pointer
> array via kvcalloc(page_cnt, ..., GFP_KERNEL_ACCOUNT) and all page_cnt
> pages via bpf_map_alloc_pages() before checking whether the requested
> arena range is actually available.
>
> The range check happens later in arena_adjust_tree() under the spinlock.
> That check rejects requests where a uaddr hint overlaps already
> allocated pages (is_range_tree_set() returns -ESRCH) or where no
> contiguous free range exists (range_tree_find() returns -ENOENT).
>
> The old code ran these checks under the lock before allocating anything,
> so rejected requests returned NULL at minimal cost. With the new
> sleepable path, each rejected request allocates and zeroes all the pages
> first via GFP_KERNEL allocations, potentially pushing the system into
> direct reclaim or OOM, then frees them all through the out_free_pages
> loop.
>
> Can a sleepable program that retries a fixed-address allocation or
> probes an almost-full arena repeatedly cause large reclaim and OOM
> pressure for requests that were never going to succeed?
>
> One approach would be to reserve the range under the lock first via
> arena_adjust_tree() and allocate afterwards, restoring the range with
> range_tree_set() if allocation fails. Another option would be an
> unlocked or locked precheck against the range tree before
> preallocating.
Not possible. Reserved ranges can be pages in with minor page faults from
userspace, causing the allocation to fail halfway. Any pages we install
after dropping the lock will be visible because apply_to_page_range will
have completed for them, preventing us from rolling back the allocation
as they may have been modified by userspace.
>
> [ ... ]
>
>> diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
>> index 244a939b9d2d..9eeb65f2d929 100644
>> --- a/kernel/bpf/syscall.c
>> +++ b/kernel/bpf/syscall.c
>
> [ ... ]
>
>
> ---
> AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
> See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
>
> CI run summary: https://github.com/kernel-patches/bpf/actions/runs/35961888778
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path
2026-09-24 5:36 ` [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path Emil Tsalapatis
2026-09-24 6:29 ` bot+bpf-ci
2026-09-24 10:11 ` Jiayuan Chen
@ 2026-09-24 21:33 ` Alexei Starovoitov
2026-09-24 22:03 ` Emil Tsalapatis
2 siblings, 1 reply; 9+ messages in thread
From: Alexei Starovoitov @ 2026-09-24 21:33 UTC (permalink / raw)
To: Emil Tsalapatis, bpf; +Cc: andrii, eddyz87, memxor, daniel
On Thu, Sep 24, 2026 at 05:36 AM Emil Tsalapatis <emil@etsalapatis.com> wrote:
> Add a separate arena page allocation path just for sleepable callers.
> The path preallocates the arena memory to be added to the tree before
> taking the critical section.
Another idea...
can we drop pages[] and link the pages via page->pcp_llist
the way apply_range_clear_cb() does on the free side ?
Then both sleepable and non-sleepable allocate all pages before
taking the lock and there is no need for kvcalloc vs kmalloc_nolock,
batching, etc.
The whole function will probably be much simpler.
pw-bot: cr
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path
2026-09-24 21:33 ` Alexei Starovoitov
@ 2026-09-24 22:03 ` Emil Tsalapatis
0 siblings, 0 replies; 9+ messages in thread
From: Emil Tsalapatis @ 2026-09-24 22:03 UTC (permalink / raw)
To: Alexei Starovoitov, Emil Tsalapatis, bpf; +Cc: andrii, eddyz87, memxor, daniel
On Thu Sep 24, 2026 at 9:33 PM UTC, Alexei Starovoitov wrote:
> On Thu, Sep 24, 2026 at 05:36 AM Emil Tsalapatis <emil@etsalapatis.com> wrote:
>> Add a separate arena page allocation path just for sleepable callers.
>> The path preallocates the arena memory to be added to the tree before
>> taking the critical section.
>
> Another idea...
> can we drop pages[] and link the pages via page->pcp_llist
> the way apply_range_clear_cb() does on the free side ?
>
> Then both sleepable and non-sleepable allocate all pages before
> taking the lock and there is no need for kvcalloc vs kmalloc_nolock,
> batching, etc.
> The whole function will probably be much simpler.
That would be ideal. Most of the complexity of this has been shaping the
function around the fact we need these intermediate arrays that force us
to drop the lock to refill/possibly require significant memory allocations.
pcp_llist should be available to use after allocating the page. I'll adjust
and resend.
>
> pw-bot: cr
^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2026-09-24 22:04 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24 5:36 [PATCH bpf v2 0/2] Make sleepable arena paths use sleepable alloc_pages Emil Tsalapatis
2026-09-24 5:36 ` [PATCH bpf v2 1/2] bpf: Add sleepable arena page allocation path Emil Tsalapatis
2026-09-24 6:29 ` bot+bpf-ci
2026-09-24 18:29 ` Emil Tsalapatis
2026-09-24 10:11 ` Jiayuan Chen
2026-09-24 17:56 ` Emil Tsalapatis
2026-09-24 21:33 ` Alexei Starovoitov
2026-09-24 22:03 ` Emil Tsalapatis
2026-09-24 5:36 ` [PATCH bpf v2 2/2] selftests/bpf: Test large allocations for both sleepable/nonsleepable arena users Emil Tsalapatis
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox