From: Shakeel Butt <shakeel.butt@linux.dev>
To: Hui Zhu <hui.zhu@linux.dev>
Cc: Roman Gushchin <roman.gushchin@linux.dev>,
JP Kobryn <inwardvessel@gmail.com>,
Andrew Morton <akpm@linux-foundation.org>,
Andrii Nakryiko <andrii@kernel.org>,
Eduard Zingerman <eddyz87@gmail.com>,
Ihor Solodrai <ihor.solodrai@linux.dev>,
Alexei Starovoitov <ast@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Kumar Kartikeya Dwivedi <memxor@gmail.com>,
Martin KaFai Lau <martin.lau@linux.dev>,
Song Liu <song@kernel.org>,
Yonghong Song <yonghong.song@linux.dev>,
Jiri Olsa <jolsa@kernel.org>,
Emil Tsalapatis <emil@etsalapatis.com>,
Shuah Khan <shuah@kernel.org>, Barry Song <baohua@kernel.org>,
Geliang Tang <geliang@kernel.org>,
linux-kernel@vger.kernel.org, bpf@vger.kernel.org,
linux-mm@kvack.org, linux-kselftest@vger.kernel.org,
Hui Zhu <zhuhui@kylinos.cn>
Subject: Re: [PATCH bpf-next v5 1/2] mm/bpf: Add bpf_proactive_reclaim kfunc
Date: Fri, 28 Aug 2026 12:53:19 -0700 [thread overview]
Message-ID: <apHk86Y8araBlXcq@linux.dev> (raw)
In-Reply-To: <5dfdc7800469eac4e9a240f2422ba65d4ef4c4ba.1787826402.git.zhuhui@kylinos.cn>
On Thu, Aug 27, 2026 at 06:36:29PM +0800, Hui Zhu wrote:
> From: Hui Zhu <zhuhui@kylinos.cn>
>
> Add bpf_proactive_reclaim(), a sleepable kfunc which performs one
> proactive reclaim pass on a given memory cgroup, similar to a write
> to memory.reclaim but without retrying until the target is reached.
>
> The kfunc refuses to reclaim if the calling task is already in a
> reclaim context, as a nested reclaim would corrupt the outer reclaim
> state.
>
> Signed-off-by: Hui Zhu <zhuhui@kylinos.cn>
> ---
> mm/bpf_memcontrol.c | 46 +++++++++++++++++++++++++++++++++++++++++++++
> 1 file changed, 46 insertions(+)
>
> diff --git a/mm/bpf_memcontrol.c b/mm/bpf_memcontrol.c
> index 716df49d7647..297ff7f05042 100644
> --- a/mm/bpf_memcontrol.c
> +++ b/mm/bpf_memcontrol.c
> @@ -6,6 +6,7 @@
> */
>
> #include <linux/memcontrol.h>
> +#include <linux/swap.h>
> #include <linux/bpf.h>
>
> __bpf_kfunc_start_defs();
> @@ -159,6 +160,49 @@ __bpf_kfunc void bpf_mem_cgroup_flush_stats(struct mem_cgroup *memcg)
> mem_cgroup_flush_stats(memcg);
> }
>
> +/*
> + * Reclaim must not recurse: try_to_free_mem_cgroup_pages() overwrites
> + * current->reclaim_state, so a nested call would corrupt the outer
> + * reclaim state. Reclaim windows are marked with PF_MEMALLOC;
> + * reclaim_state is also checked because it is installed slightly
> + * before PF_MEMALLOC.
> + */
> +static bool bpf_in_reclaim_context(void)
> +{
> + return (current->flags & PF_MEMALLOC) || current->reclaim_state;
> +}
> +
> +/**
> + * bpf_proactive_reclaim - proactively reclaim memory from a memory
> + * cgroup
> + * @memcg: the target memory cgroup to reclaim from
> + * @size: the amount of memory to reclaim, in bytes
> + *
> + * Trigger one proactive reclaim pass on @memcg, similar to a write to
> + * memory.reclaim, but without retrying until @size is reached.
> + * Must not be called with a filesystem lock held: the reclaim path
> + * may deadlock on it via filesystem shrinkers.
> + *
> + * Return: The amount of memory reclaimed, in bytes, or 0 if @size is
> + * smaller than a page or the task is already in a reclaim context.
> + */
> +__bpf_kfunc unsigned long bpf_proactive_reclaim(struct mem_cgroup *memcg,
> + unsigned long size)
> +{
> + unsigned long nr_reclaimed;
> +
> + if (size < PAGE_SIZE || unlikely(bpf_in_reclaim_context()))
> + return 0;
I have been thinking about this more and more and after looking at the reasoning
behind your check current->reclaim_state and also Sashiko's comment on NOIO/NOFS
contexts, I am more convinced that this kfunc can not be a simple sleepable
function. We need more than that. We need clean process context as well similar
to the userspace poking memory.reclaim. Something like kthread or workqueue.
Kumar & Andrii, is there a way to restrict a kfunc to only be called from
special BPF threads/workqueues? Is there some similar concept in BPF world?
> +
> + nr_reclaimed = try_to_free_mem_cgroup_pages(memcg, size / PAGE_SIZE,
> + GFP_KERNEL,
> + MEMCG_RECLAIM_MAY_SWAP |
> + MEMCG_RECLAIM_PROACTIVE,
> + NULL);
> +
> + return nr_reclaimed * PAGE_SIZE;
> +}
> +
> __bpf_kfunc_end_defs();
>
> BTF_KFUNCS_START(bpf_memcontrol_kfuncs)
> @@ -172,6 +216,8 @@ BTF_ID_FLAGS(func, bpf_mem_cgroup_usage)
> BTF_ID_FLAGS(func, bpf_mem_cgroup_page_state)
> BTF_ID_FLAGS(func, bpf_mem_cgroup_flush_stats, KF_SLEEPABLE)
>
> +BTF_ID_FLAGS(func, bpf_proactive_reclaim, KF_SLEEPABLE)
> +
> BTF_KFUNCS_END(bpf_memcontrol_kfuncs)
>
> static const struct btf_kfunc_id_set bpf_memcontrol_kfunc_set = {
> --
> 2.53.0
>
next prev parent reply other threads:[~2026-08-28 19:53 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-27 10:36 [PATCH bpf-next v5 0/2] bpf: BPF-driven proactive memcg reclaim Hui Zhu
2026-08-27 10:36 ` [PATCH bpf-next v5 1/2] mm/bpf: Add bpf_proactive_reclaim kfunc Hui Zhu
2026-08-27 10:46 ` sashiko-bot
2026-08-27 11:37 ` bot+bpf-ci
2026-08-28 19:53 ` Shakeel Butt [this message]
2026-08-28 21:07 ` Kumar Kartikeya Dwivedi
2026-09-01 2:59 ` Hui Zhu
2026-08-27 10:36 ` [PATCH bpf-next v5 2/2] selftests/bpf: Add memcg async reclaim test Hui Zhu
2026-08-27 10:46 ` sashiko-bot
2026-08-27 11:37 ` bot+bpf-ci
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=apHk86Y8araBlXcq@linux.dev \
--to=shakeel.butt@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=baohua@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=geliang@kernel.org \
--cc=hui.zhu@linux.dev \
--cc=ihor.solodrai@linux.dev \
--cc=inwardvessel@gmail.com \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=roman.gushchin@linux.dev \
--cc=shuah@kernel.org \
--cc=song@kernel.org \
--cc=yonghong.song@linux.dev \
--cc=zhuhui@kylinos.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.