From: "Hui Zhu" <hui.zhu@linux.dev>
To: Roman Gushchin <roman.gushchin@linux.dev>,
JP Kobryn <inwardvessel@gmail.com>,
Shakeel Butt <shakeel.butt@linux.dev>,
Andrew Morton <akpm@linux-foundation.org>,
Andrii Nakryiko <andrii@kernel.org>,
Eduard Zingerman <eddyz87@gmail.com>,
Ihor Solodrai <ihor.solodrai@linux.dev>,
Alexei Starovoitov <ast@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Kumar Kartikeya Dwivedi <memxor@gmail.com>,
Martin KaFai Lau <martin.lau@linux.dev>,
Song Liu <song@kernel.org>,
Yonghong Song <yonghong.song@linux.dev>,
Jiri Olsa <jolsa@kernel.org>,
Emil Tsalapatis <emil@etsalapatis.com>,
Shuah Khan <shuah@kernel.org>, Barry Song <baohua@kernel.org>,
Geliang Tang <geliang@kernel.org>,
linux-kernel@vger.kernel.org, bpf@vger.kernel.org,
linux-mm@kvack.org, linux-kselftest@vger.kernel.org
Cc: Hui Zhu <zhuhui@kylinos.cn>
Subject: [PATCH bpf-next v4 0/2] bpf: BPF-driven proactive memcg reclaim
Date: Thu, 20 Aug 2026 14:12:25 +0800 [thread overview]
Message-ID: <cover.1787205002.git.zhuhui@kylinos.cn> (raw)
From: Hui Zhu <zhuhui@kylinos.cn>
This series lets a BPF program decide when to trigger memcg reclaim
and how aggressively to do it, based on whatever runtime signal it
chooses to observe -- rather than reclaim only being triggered once a
cgroup's usage crosses a fixed threshold. The core idea is a pair of
new kfuncs, bpf_proactive_reclaim() and
bpf_proactive_reclaim_swappiness(), which give BPF direct access to
the proactive reclaim path so this decision can be made in BPF policy
rather than hard-coded threshold logic.
This was originally part of a larger series posted here [1].
That series also adds a memcg BPF struct_ops (memcg_charged,
memcg_uncharged, below_low, below_min) for synchronous, in-line memory
protection decisions. That mechanism and this one solve different
problems -- struct_ops hooks run inline on the charge/reclaim path,
while the kfuncs here are for asynchronous, out-of-band reclaim
decided independently by a BPF program -- so they are reviewed as
separate series. This series carries only the async reclaim piece.
Compared to v1, the kfunc interface has been reworked based on review
feedback: instead of a thin wrapper around
try_to_free_mem_cgroup_pages() exposing raw gfp/reclaim-option knobs,
the series now provides use-case-driven kfuncs that perform one
proactive reclaim pass with the same parameters memory.reclaim uses.
The bpf_thread_wq patches from v1 (old patches 2-3) are dropped from
this series: following the discussion in [2], the cgroup-aware
workqueue is being superseded by a disaggregated set of async
primitives (bpf_kthread/bpf_waitq) that will be developed separately
(discussion in [3]), and the selftest now queues its reclaim work
through bpf_wq.
Patch 1 adds bpf_proactive_reclaim() and
bpf_proactive_reclaim_swappiness(), sleepable kfuncs that perform one
reclaim pass on a target memcg, like a write to memory.reclaim: swap
is allowed, and the anon/file balance follows the cgroup's swappiness
or an explicit override in [MIN_SWAPPINESS, MAX_SWAPPINESS] plus
SWAPPINESS_ANON_ONLY. Both go through a shared helper,
bpf_proactive_reclaim_pages(), which guards against reclaim recursion
and calls try_to_free_mem_cgroup_pages() with GFP_KERNEL and
MEMCG_RECLAIM_MAY_SWAP | MEMCG_RECLAIM_PROACTIVE, the same parameters
user_proactive_reclaim() uses, and unlike memory.reclaim they do not
retry until the requested size is reached. Both refuse to run when
the caller already holds PF_MEMALLOC or has a non-NULL
current->reclaim_state, since a nested try_to_free_mem_cgroup_pages()
would clobber the outer reclaim's current->reclaim_state (e.g. MGLRU
dereferences current->reclaim_state->mm_walk); the reclaim_state
check also closes the window where try_to_free_mem_cgroup_pages() has
installed it but not yet set PF_MEMALLOC, reachable by a sleepable
program attaching fentry to the generated trace iterator function.
The size argument and the return value are both in bytes, matching
the byte-based unit of bpf_mem_cgroup_usage() and
bpf_mem_cgroup_page_state() so callers can mix them without manual
page/byte conversions. An out-of-range swappiness is reported with
(unsigned long)-1 rather than 0, since 0 cannot be told apart from a
pass that reclaimed nothing.
Patch 2 (selftests/bpf: add memcg async reclaim test) ties the kfuncs
into a worked example: it watches the WORKINGSET_REFAULT_FILE counter
of a high-priority cgroup as a proxy for memory-pressure impact, and
once it starts climbing, proactively reclaims memory from a
low-priority cgroup via bpf_proactive_reclaim(), with the reclaim
work queued asynchronously through bpf_wq. The test asserts that the
monitored cgroup's workload finishes faster once async reclaim kicks
in, and -- as timing alone cannot distinguish a working reclaim from
a no-op one -- that the BPF program actually made reclaim calls and
reclaimed bytes, via counters it exports through its .bss. This
demonstrates the end-to-end use case: BPF observes pressure on the
cgroup it wants to protect, and reclaims from the cgroup it wants to
reclaim from, in one self-contained mechanism. Note that, without
bpf_thread_wq, the CPU cost of the reclaim work is not yet attributed
to a chosen cgroup; that part waits for the async primitives work
mentioned above.
Changelog:
v4:
According to the comments of bot+bpf-ci and sashiko, also check
current->reclaim_state to close the fentry-on-trace-iter recursion
window in bpf_in_reclaim_context.
Return bytes instead of pages ( nr * PAGE_SIZE ) in
bpf_proactive_reclaim_pages and bpf_proactive_reclaim_swappiness.
Return (unsigned long)-1 on out-of-range swappiness (was 0).
Kdoc of both kfuncs: updated Return descriptions; added FS-lock deadlock
warning to bpf_proactive_reclaim.
Fix potential child process leak in selftests.
Use _exit() instead of exit() in forked children in selftests.
Rename reclaimed_pages to reclaimed_bytes in selftests.
Fix comments issues in selftests.
v3:
According to the comments of bot+bpf-ci, add a shared helper
bpf_proactive_reclaim_pages() that is called by bpf_proactive_reclaim
and bpf_proactive_reclaim_swappiness.
According to the comments of sashiko and bot+bpf-ci, fix the issues of
selftests.
v2:
According to the comments of Shakeel Butt, replace
bpf_try_to_free_mem_cgroup_pages() with
bpf_proactive_reclaim(memcg, size) and
bpf_proactive_reclaim_swappiness(memcg, size, swappiness).
According to the comments of Kumar Kartikeya Dwivedi, drop patch 2
and patch 3.
Remove bpf_thread_wq code in patch 4.
According to the comments of sashiko-bot, fix the issues of selftests.
[1] https://sashiko.dev/#/message/cover.1779760876.git.zhuhui%40kylinos.cn
[2] https://sashiko.dev/#/message/1b58d56976202f26818d31dbd0da2ecb2e2460f5%40linux.dev
[3] https://sashiko.dev/#/message/DKNHV09PBQZP.IRQL20BY574I%40gmail.com
Hui Zhu (2):
mm/bpf: Add bpf_proactive_reclaim kfuncs
selftests/bpf: add memcg async reclaim test
mm/bpf_memcontrol.c | 118 +++++
.../bpf/prog_tests/memcg_async_reclaim.c | 480 ++++++++++++++++++
.../selftests/bpf/progs/memcg_async_reclaim.c | 181 +++++++
3 files changed, 779 insertions(+)
create mode 100644 tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c
create mode 100644 tools/testing/selftests/bpf/progs/memcg_async_reclaim.c
--
2.53.0
next reply other threads:[~2026-08-20 6:12 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-20 6:12 Hui Zhu [this message]
2026-08-20 6:12 ` [PATCH bpf-next v4 1/2] mm/bpf: Add bpf_proactive_reclaim kfuncs Hui Zhu
2026-08-20 6:24 ` sashiko-bot
2026-08-20 7:05 ` bot+bpf-ci
2026-08-20 6:12 ` [PATCH bpf-next v4 2/2] selftests/bpf: add memcg async reclaim test Hui Zhu
2026-08-20 6:22 ` sashiko-bot
2026-08-20 7:05 ` bot+bpf-ci
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=cover.1787205002.git.zhuhui@kylinos.cn \
--to=hui.zhu@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=baohua@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=geliang@kernel.org \
--cc=ihor.solodrai@linux.dev \
--cc=inwardvessel@gmail.com \
--cc=jolsa@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=shuah@kernel.org \
--cc=song@kernel.org \
--cc=yonghong.song@linux.dev \
--cc=zhuhui@kylinos.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.