From: Jiayuan Chen <jiayuan.chen@linux.dev>
To: bpf@vger.kernel.org
Cc: Jiayuan Chen <jiayuan.chen@linux.dev>,
Alexei Starovoitov <ast@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
John Fastabend <john.fastabend@gmail.com>,
Andrii Nakryiko <andrii@kernel.org>,
Eduard Zingerman <eddyz87@gmail.com>,
Kumar Kartikeya Dwivedi <memxor@gmail.com>,
Martin KaFai Lau <martin.lau@linux.dev>,
Song Liu <song@kernel.org>,
Yonghong Song <yonghong.song@linux.dev>,
Jiri Olsa <jolsa@kernel.org>,
Emil Tsalapatis <emil@etsalapatis.com>,
Ihor Solodrai <ihor.solodrai@linux.dev>,
Shuah Khan <shuah@kernel.org>,
Sebastian Andrzej Siewior <bigeasy@linutronix.de>,
Clark Williams <clrkwllms@kernel.org>,
Steven Rostedt <rostedt@goodmis.org>,
linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org,
linux-rt-devel@lists.linux.dev
Subject: [PATCH bpf-next v5 0/4] bpf: arena: handle memory.max on fault-in with reclaim/OOM
Date: Mon, 24 Aug 2026 17:29:51 +0800 [thread overview]
Message-ID: <20260824093122.362135-1-jiayuan.chen@linux.dev> (raw)
Since commit e66fe1bc6d25 ("bpf: arena: Reintroduce memcg accounting"),
arena pages are charged to the memcg of the process that created the arena.
That accounting exposes two problems in the arena user page fault path.
1. The fault-in allocation runs under arena->spinlock, so it can only use the
non-blocking allocator, which never reclaims. Once memory.current is at
memory.max the allocation simply fails. Reaching memory.max is completely
normal for a healthy application - e.g. reading a large file fills
memory.current with page cache - so the process ends up killed for no real
reason.
2. That failure is turned into VM_FAULT_SIGSEGV, which is misleading: the
faulting address is a perfectly valid arena address. The memcg OOM killer
should run instead, and a genuine out-of-memory should surface as a
non-recoverable fault, not a bogus segfault.
Preallocate the page outside the lock (patch 2), the way do_anonymous_page()
does, so the allocation can sleep, reclaim and run the memcg OOM killer. On a
genuine failure the fault is non-recoverable and returns VM_FAULT_SIGBUS; a
task faulting its own arena is unaffected, the OOM killer picks it inside the
allocation and it dies by SIGKILL. This needs a sleepable allocator (patch 1),
because can_alloc_pages() is a conservative guess for BPF program context and
always forces the non-blocking allocator under PREEMPT_RT. patch 3&4 adds a
selftest that faults an arena in under a memory.max limit: without the fix the
child gets SIGSEGV on a valid address, with it the child is killed by the
memcg OOM killer.
v4 -> v5:
- arena: return VM_FAULT_SIGBUS instead of VM_FAULT_OOM when the fault-in
allocation fails. The allocation already ran reclaim and the OOM killer,
so the failure is non-recoverable; VM_FAULT_OOM would be retried by the
fault path and can livelock a task faulting a shared arena whose memcg
OOM killer cannot reach it. (reported by the Sashiko AI review)
- selftest: cap the arena at 50000 pages so it stays under the 4G arena
limit on 64K-page kernels.
- selftest: ASSERT_OK setup_cgroup_environment(), print a reason on the
no-memory-controller skip, and flush only stderr in the child.
v3 -> v4:
- rebase bpf-next and fix conflict
- add Reviewed-by tag from Emil Tsalapatis
v2 -> v3:
- selftest: check the memcg OOM via memory.events "oom_kill" instead of
the exit signal; it only aims to pass on the fixed kernel, since the
unfixed SIGSEGV is racy.
v1 -> v2:
- Rebase on the separate deadlock fix (found by the Sashiko AI review),
now applied to bpf-next.
- Honor the map's NUMA node on fault-in.
- Return VM_FAULT_SIGBUS for the non-recoverable faults (lock, range-tree
and page-table failures); a scratch-page hole stays VM_FAULT_SIGSEGV
only under BPF_F_SEGV_ON_FAULT. (Kumar Kartikeya Dwivedi)
- Add read_cgroup_file() to cgroup_helpers instead of open-coding the
/mnt/... path in the test. (Emil Tsalapatis)
- Dump the cgroup memory stats on test failure to ease debugging.
v4:
https://lore.kernel.org/bpf/20260821050250.35112-1-jiayuan.chen@linux.dev/T/#t
v2:
https://lore.kernel.org/bpf/20260805091720.139924-1-jiayuan.chen@linux.dev/
v1:
https://lore.kernel.org/bpf/20260727062521.376231-1-jiayuan.chen@linux.dev/
Jiayuan Chen (4):
bpf: Add a sleepable page allocator for map memory
bpf: arena: allocate the fault-in page outside the lock
selftests/bpf: Add read_cgroup_file() to cgroup_helpers
selftests/bpf: Add a test for arena fault-in under memory.max
include/linux/bpf.h | 1 +
kernel/bpf/arena.c | 90 +++++++---
kernel/bpf/syscall.c | 21 ++-
tools/testing/selftests/bpf/cgroup_helpers.c | 67 ++++++++
tools/testing/selftests/bpf/cgroup_helpers.h | 4 +
.../selftests/bpf/prog_tests/arena_memcg.c | 158 ++++++++++++++++++
.../testing/selftests/bpf/progs/arena_memcg.c | 24 +++
7 files changed, 340 insertions(+), 25 deletions(-)
create mode 100644 tools/testing/selftests/bpf/prog_tests/arena_memcg.c
create mode 100644 tools/testing/selftests/bpf/progs/arena_memcg.c
--
2.43.0
next reply other threads:[~2026-08-24 9:32 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-24 9:29 Jiayuan Chen [this message]
2026-08-24 9:29 ` [PATCH bpf-next v5 1/4] bpf: Add a sleepable page allocator for map memory Jiayuan Chen
2026-08-24 9:29 ` [PATCH bpf-next v5 2/4] bpf: arena: allocate the fault-in page outside the lock Jiayuan Chen
2026-08-24 10:33 ` bot+bpf-ci
2026-08-24 9:29 ` [PATCH bpf-next v5 3/4] selftests/bpf: Add read_cgroup_file() to cgroup_helpers Jiayuan Chen
2026-08-24 9:29 ` [PATCH bpf-next v5 4/4] selftests/bpf: Add a test for arena fault-in under memory.max Jiayuan Chen
2026-08-24 10:18 ` bot+bpf-ci
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260824093122.362135-1-jiayuan.chen@linux.dev \
--to=jiayuan.chen@linux.dev \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bigeasy@linutronix.de \
--cc=bpf@vger.kernel.org \
--cc=clrkwllms@kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=ihor.solodrai@linux.dev \
--cc=john.fastabend@gmail.com \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-rt-devel@lists.linux.dev \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=rostedt@goodmis.org \
--cc=shuah@kernel.org \
--cc=song@kernel.org \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox