From: Khawar Ahemad <ahemadkhawar123@gmail.com>
To: bpf@vger.kernel.org
Cc: linux-kernel@vger.kernel.org, ast@kernel.org,
daniel@iogearbox.net, andrii@kernel.org, eddyz87@gmail.com,
jiayuan.chen@linux.dev, emil@etsalapatis.com,
ahemadkhawar123@gmail.com
Subject: [PATCH bpf-next v6 0/4] bpf: arena: handle memory.max on fault-in with reclaim/OOM
Date: Tue, 25 Aug 2026 15:19:51 +0530 [thread overview]
Message-ID: <20260825094955.83240-1-ahemadkhawar123@gmail.com> (raw)
This series fixes an issue where accessing a valid BPF arena page under
memory pressure (hitting a cgroup's memory.max limit) incorrectly results
in a SIGSEGV crash instead of invoking memory reclaim or the memcg OOM
killer.
Problem:
========
arena_vm_fault() performed page allocation while holding arena->spinlock
(with interrupts disabled), restricting the allocation to the non-blocking
alloc_pages_nolock() path, which cannot sleep, reclaim memory, or invoke
the memcg OOM killer. When the cgroup's memory.max limit was reached, the
allocation returned -ENOMEM, which arena_vm_fault() converted to
VM_FAULT_SIGSEGV, killing the process on a valid virtual address. Hitting
memory.max is routine (e.g., page cache growth from reading a large file),
so this killed innocent processes.
Solution:
=========
1. Add bpf_map_alloc_page_sleepable() for callers that know they are in a
sleepable context. can_alloc_pages() is a conservative guess for BPF
program context -- notably it is always false on PREEMPT_RT -- so going
through bpf_map_alloc_pages() there needlessly picks the non-blocking
allocator. The new helper uses BPF_PAGE_GFP (GFP_KERNEL | __GFP_ZERO |
__GFP_ACCOUNT | __GFP_NOWARN) directly.
2. Rework arena_vm_fault() to preallocate the page outside arena->spinlock
with bpf_map_alloc_page_sleepable(), like do_anonymous_page() does, so
it can sleep, reclaim, and invoke the memcg OOM killer instead of
turning a routine memory.max event into a spurious SIGSEGV.
3. A lockless probe (vmalloc_to_page()) before allocation skips
preallocation when a page is already mapped (e.g., allocated by the bpf
program), so the common case wastes no allocation. The rare race where
such a page is freed before the lock is taken falls back to the
non-blocking allocator under the lock.
4. Return VM_FAULT_SIGBUS for non-recoverable errors (lock failure,
range-tree and page-table failures) instead of VM_FAULT_SIGSEGV;
only BPF_F_SEGV_ON_FAULT, and a scratch-page hole under that flag,
is a real user addressing error and keeps VM_FAULT_SIGSEGV.
5. Add a BPF selftest (arena_memcg) that joins a child process into a
memcg capped 64 MiB above its post-load usage, faults arena pages
until the budget is exhausted, and verifies the child is OOM-killed
(memcg oom_kill event) rather than receiving SIGSEGV.
v6 vs v5:
- Fixed redundant #include <unistd.h> inside #ifndef PAGE_SIZE guard
in the selftest (the header was already unconditionally included).
- All patches carry both original author and reviewer Signed-off-by.
Signed-off-by: Khawar Ahemad <ahemadkhawar123@gmail.com>
Jiayuan Chen (4):
bpf: Add a sleepable page allocator for map memory
bpf: arena: allocate the fault-in page outside the lock
selftests/bpf: Add read_cgroup_file() to cgroup_helpers
selftests/bpf: Add a test for arena fault-in under memory.max
include/linux/bpf.h | 1 +
kernel/bpf/arena.c | 90 +++++++---
kernel/bpf/syscall.c | 21 ++-
tools/testing/selftests/bpf/cgroup_helpers.c | 67 ++++++++
tools/testing/selftests/bpf/cgroup_helpers.h | 4 +
.../selftests/bpf/prog_tests/arena_memcg.c | 157 ++++++++++++++++++
.../testing/selftests/bpf/progs/arena_memcg.c | 24 +++
7 files changed, 339 insertions(+), 25 deletions(-)
create mode 100644 tools/testing/selftests/bpf/prog_tests/arena_memcg.c
create mode 100644 tools/testing/selftests/bpf/progs/arena_memcg.c
--
2.54.0 (Apple Git-157)
next reply other threads:[~2026-08-25 9:50 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-25 9:49 Khawar Ahemad [this message]
2026-08-25 9:49 ` [PATCH bpf-next v6 1/4] bpf: Add a sleepable page allocator for map memory Khawar Ahemad
2026-08-25 10:11 ` sashiko-bot
2026-08-25 10:22 ` Khawar Ahemad
2026-08-25 10:32 ` bot+bpf-ci
2026-08-25 9:49 ` [PATCH bpf-next v6 2/4] bpf: arena: allocate the fault-in page outside the lock Khawar Ahemad
2026-08-25 10:32 ` bot+bpf-ci
2026-08-25 9:49 ` [PATCH bpf-next v6 3/4] selftests/bpf: Add read_cgroup_file() to cgroup_helpers Khawar Ahemad
2026-08-25 9:49 ` [PATCH bpf-next v6 4/4] selftests/bpf: Add a test for arena fault-in under memory.max Khawar Ahemad
2026-08-25 10:32 ` bot+bpf-ci
2026-08-25 10:43 ` Khawar Ahemad
2026-08-25 11:19 ` Jiayuan Chen
2026-08-25 12:23 ` Kumar Kartikeya Dwivedi
2026-08-25 11:28 ` Khawar Ahemad
2026-08-25 9:55 ` [PATCH bpf-next v6 0/4] bpf: arena: handle memory.max on fault-in with reclaim/OOM Jiayuan Chen
-- strict thread matches above, loose matches on Subject: below --
2026-08-25 9:19 Khawar Ahemad
2026-08-25 9:16 Khawar Ahemad
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260825094955.83240-1-ahemadkhawar123@gmail.com \
--to=ahemadkhawar123@gmail.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=jiayuan.chen@linux.dev \
--cc=linux-kernel@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox