From: "Emil Tsalapatis" <emil@etsalapatis.com>
To: "Jiayuan Chen" <jiayuan.chen@linux.dev>, <bpf@vger.kernel.org>
Cc: "Alexei Starovoitov" <ast@kernel.org>,
"Daniel Borkmann" <daniel@iogearbox.net>,
"John Fastabend" <john.fastabend@gmail.com>,
"Andrii Nakryiko" <andrii@kernel.org>,
"Eduard Zingerman" <eddyz87@gmail.com>,
"Kumar Kartikeya Dwivedi" <memxor@gmail.com>,
"Martin KaFai Lau" <martin.lau@linux.dev>,
"Song Liu" <song@kernel.org>,
"Yonghong Song" <yonghong.song@linux.dev>,
"Jiri Olsa" <jolsa@kernel.org>,
"Emil Tsalapatis" <emil@etsalapatis.com>,
"Shuah Khan" <shuah@kernel.org>,
"Sebastian Andrzej Siewior" <bigeasy@linutronix.de>,
"Clark Williams" <clrkwllms@kernel.org>,
"Steven Rostedt" <rostedt@goodmis.org>,
<linux-kernel@vger.kernel.org>, <linux-kselftest@vger.kernel.org>,
<linux-rt-devel@lists.linux.dev>
Subject: Re: [PATCH bpf-next 1/3] bpf: Add a sleepable page allocator for map memory
Date: Mon, 27 Jul 2026 20:48:51 -0400 [thread overview]
Message-ID: <DK9SHN2RV21T.485J38W373WQ@etsalapatis.com> (raw)
In-Reply-To: <20260727062521.376231-2-jiayuan.chen@linux.dev>
On Mon Jul 27, 2026 at 2:24 AM EDT, Jiayuan Chen wrote:
> bpf_map_alloc_pages() picks the allocator via can_alloc_pages(), a
> conservative guess for BPF program context that is always false under
> PREEMPT_RT. So even a caller that really is sleepable gets the
> non-blocking allocator, which never reclaims and never engages the OOM
> machinery.
>
> Add bpf_map_alloc_page_sleepable() for callers that know they are
> sleepable. The next patch uses it from the arena page fault handler.
>
> Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Reviewed-by: Emil Tsalapatis <emil@etsalapatis.com>
> ---
> include/linux/bpf.h | 1 +
> kernel/bpf/syscall.c | 21 +++++++++++++++++----
> 2 files changed, 18 insertions(+), 4 deletions(-)
>
> diff --git a/include/linux/bpf.h b/include/linux/bpf.h
> index d9542127dfdf..9ba03622709b 100644
> --- a/include/linux/bpf.h
> +++ b/include/linux/bpf.h
> @@ -2779,6 +2779,7 @@ struct bpf_prog *bpf_prog_get_curr_or_next(u32 *id);
>
> int bpf_map_alloc_pages(const struct bpf_map *map, int nid,
> unsigned long nr_pages, struct page **page_array);
> +struct page *bpf_map_alloc_page_sleepable(const struct bpf_map *map, int nid);
> #ifdef CONFIG_MEMCG
> void bpf_map_memcg_enter(const struct bpf_map *map, struct mem_cgroup **old_memcg,
> struct mem_cgroup **new_memcg);
> diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c
> index 0ff9e3aa293d..cec88450e4ce 100644
> --- a/kernel/bpf/syscall.c
> +++ b/kernel/bpf/syscall.c
> @@ -602,15 +602,14 @@ static bool can_alloc_pages(void)
> !IS_ENABLED(CONFIG_PREEMPT_RT);
> }
>
> +#define BPF_PAGE_GFP (GFP_KERNEL | __GFP_ZERO | __GFP_ACCOUNT | __GFP_NOWARN)
> +
> static struct page *__bpf_alloc_page(int nid)
> {
> if (!can_alloc_pages())
> return alloc_pages_nolock(__GFP_ACCOUNT, nid, 0);
>
> - return alloc_pages_node(nid,
> - GFP_KERNEL | __GFP_ZERO | __GFP_ACCOUNT
> - | __GFP_NOWARN,
> - 0);
> + return alloc_pages_node(nid, BPF_PAGE_GFP, 0);
> }
>
> int bpf_map_alloc_pages(const struct bpf_map *map, int nid,
> @@ -636,6 +635,20 @@ int bpf_map_alloc_pages(const struct bpf_map *map, int nid,
> return ret;
> }
>
> +/*
> + * For callers that know they run in a sleepable context, e.g. a user page
> + * fault handler. can_alloc_pages() is a conservative guess made for BPF
> + * program context - notably it is always false on PREEMPT_RT - so going
> + * through bpf_map_alloc_pages() there would needlessly pick the
> + * non-blocking allocator, which never reclaims and never engages the OOM
> + * machinery.
> + */
> +struct page *bpf_map_alloc_page_sleepable(const struct bpf_map *map, int nid)
> +{
> + might_sleep();
> + return alloc_pages_node(nid, BPF_PAGE_GFP, 0);
> +}
> +
>
> static int btf_field_cmp(const void *a, const void *b)
> {
next prev parent reply other threads:[~2026-07-28 0:48 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-27 6:24 [PATCH bpf-next 0/3] bpf: arena: handle memory.max on fault-in with reclaim/OOM Jiayuan Chen
2026-07-27 6:24 ` [PATCH bpf-next 1/3] bpf: Add a sleepable page allocator for map memory Jiayuan Chen
2026-07-27 6:37 ` sashiko-bot
2026-07-27 7:30 ` Jiayuan Chen
2026-07-28 0:48 ` Emil Tsalapatis [this message]
2026-07-27 6:24 ` [PATCH bpf-next 2/3] bpf: arena: allocate the fault-in page outside the lock Jiayuan Chen
2026-07-27 6:42 ` sashiko-bot
2026-07-27 8:00 ` Jiayuan Chen
2026-07-28 1:00 ` Emil Tsalapatis
2026-07-27 7:10 ` bot+bpf-ci
2026-07-28 1:22 ` Emil Tsalapatis
2026-08-03 6:43 ` Jiayuan Chen
2026-07-27 6:24 ` [PATCH bpf-next 3/3] selftests/bpf: Add a test for arena fault-in under memory.max Jiayuan Chen
2026-07-28 0:54 ` Emil Tsalapatis
2026-08-03 11:42 ` Jiayuan Chen
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DK9SHN2RV21T.485J38W373WQ@etsalapatis.com \
--to=emil@etsalapatis.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=bigeasy@linutronix.de \
--cc=bpf@vger.kernel.org \
--cc=clrkwllms@kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=jiayuan.chen@linux.dev \
--cc=john.fastabend@gmail.com \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-rt-devel@lists.linux.dev \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=rostedt@goodmis.org \
--cc=shuah@kernel.org \
--cc=song@kernel.org \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.