From: Khawar Ahemad <ahemadkhawar123@gmail.com>
To: bpf@vger.kernel.org
Cc: linux-kernel@vger.kernel.org, ast@kernel.org,
daniel@iogearbox.net, andrii@kernel.org, eddyz87@gmail.com,
jiayuan.chen@linux.dev, emil@etsalapatis.com,
ahemadkhawar123@gmail.com
Subject: [PATCH bpf-next v6 0/4] bpf: arena: handle memory.max on fault-in with reclaim/OOM
Date: Tue, 25 Aug 2026 15:19:51 +0530 [thread overview]
Message-ID: <20260825094955.83240-1-ahemadkhawar123@gmail.com> (raw)
This series fixes an issue where accessing a valid BPF arena page under
memory pressure (hitting a cgroup's memory.max limit) incorrectly results
in a SIGSEGV crash instead of invoking memory reclaim or the memcg OOM
killer.
Problem:
========
arena_vm_fault() performed page allocation while holding arena->spinlock
(with interrupts disabled), restricting the allocation to the non-blocking
alloc_pages_nolock() path, which cannot sleep, reclaim memory, or invoke
the memcg OOM killer. When the cgroup's memory.max limit was reached, the
allocation returned -ENOMEM, which arena_vm_fault() converted to
VM_FAULT_SIGSEGV, killing the process on a valid virtual address. Hitting
memory.max is routine (e.g., page cache growth from reading a large file),
so this killed innocent processes.
Solution:
=========
1. Add bpf_map_alloc_page_sleepable() for callers that know they are in a
sleepable context. can_alloc_pages() is a conservative guess for BPF
program context -- notably it is always false on PREEMPT_RT -- so going
through bpf_map_alloc_pages() there needlessly picks the non-blocking
allocator. The new helper uses BPF_PAGE_GFP (GFP_KERNEL | __GFP_ZERO |
__GFP_ACCOUNT | __GFP_NOWARN) directly.
2. Rework arena_vm_fault() to preallocate the page outside arena->spinlock
with bpf_map_alloc_page_sleepable(), like do_anonymous_page() does, so
it can sleep, reclaim, and invoke the memcg OOM killer instead of
turning a routine memory.max event into a spurious SIGSEGV.
3. A lockless probe (vmalloc_to_page()) before allocation skips
preallocation when a page is already mapped (e.g., allocated by the bpf
program), so the common case wastes no allocation. The rare race where
such a page is freed before the lock is taken falls back to the
non-blocking allocator under the lock.
4. Return VM_FAULT_SIGBUS for non-recoverable errors (lock failure,
range-tree and page-table failures) instead of VM_FAULT_SIGSEGV;
only BPF_F_SEGV_ON_FAULT, and a scratch-page hole under that flag,
is a real user addressing error and keeps VM_FAULT_SIGSEGV.
5. Add a BPF selftest (arena_memcg) that joins a child process into a
memcg capped 64 MiB above its post-load usage, faults arena pages
until the budget is exhausted, and verifies the child is OOM-killed
(memcg oom_kill event) rather than receiving SIGSEGV.
v6 vs v5:
- Fixed redundant #include <unistd.h> inside #ifndef PAGE_SIZE guard
in the selftest (the header was already unconditionally included).
- All patches carry both original author and reviewer Signed-off-by.
Signed-off-by: Khawar Ahemad <ahemadkhawar123@gmail.com>
Jiayuan Chen (4):
bpf: Add a sleepable page allocator for map memory
bpf: arena: allocate the fault-in page outside the lock
selftests/bpf: Add read_cgroup_file() to cgroup_helpers
selftests/bpf: Add a test for arena fault-in under memory.max
include/linux/bpf.h | 1 +
kernel/bpf/arena.c | 90 +++++++---
kernel/bpf/syscall.c | 21 ++-
tools/testing/selftests/bpf/cgroup_helpers.c | 67 ++++++++
tools/testing/selftests/bpf/cgroup_helpers.h | 4 +
.../selftests/bpf/prog_tests/arena_memcg.c | 157 ++++++++++++++++++
.../testing/selftests/bpf/progs/arena_memcg.c | 24 +++
7 files changed, 339 insertions(+), 25 deletions(-)
create mode 100644 tools/testing/selftests/bpf/prog_tests/arena_memcg.c
create mode 100644 tools/testing/selftests/bpf/progs/arena_memcg.c
--
2.54.0 (Apple Git-157)
next reply other threads:[~2026-08-25 9:50 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-25 9:49 Khawar Ahemad [this message]
2026-08-25 9:49 ` [PATCH bpf-next v6 1/4] bpf: Add a sleepable page allocator for map memory Khawar Ahemad
2026-08-25 10:11 ` sashiko-bot
2026-08-25 10:22 ` Khawar Ahemad
2026-08-25 10:32 ` bot+bpf-ci
2026-08-25 9:49 ` [PATCH bpf-next v6 2/4] bpf: arena: allocate the fault-in page outside the lock Khawar Ahemad
2026-08-25 10:32 ` bot+bpf-ci
2026-08-25 9:49 ` [PATCH bpf-next v6 3/4] selftests/bpf: Add read_cgroup_file() to cgroup_helpers Khawar Ahemad
2026-08-25 9:49 ` [PATCH bpf-next v6 4/4] selftests/bpf: Add a test for arena fault-in under memory.max Khawar Ahemad
2026-08-25 10:32 ` bot+bpf-ci
2026-08-25 10:43 ` Khawar Ahemad
2026-08-25 11:19 ` Jiayuan Chen
2026-08-25 12:23 ` Kumar Kartikeya Dwivedi
2026-08-25 11:28 ` Khawar Ahemad
2026-08-25 9:55 ` [PATCH bpf-next v6 0/4] bpf: arena: handle memory.max on fault-in with reclaim/OOM Jiayuan Chen
-- strict thread matches above, loose matches on Subject: below --
2026-08-25 9:19 Khawar Ahemad
2026-08-25 9:16 Khawar Ahemad
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260825094955.83240-1-ahemadkhawar123@gmail.com \
--to=ahemadkhawar123@gmail.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=jiayuan.chen@linux.dev \
--cc=linux-kernel@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.