Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH bpf-next v2 0/2] bpf: BPF-driven proactive memcg reclaim
@ 2026-08-18  8:36 Hui Zhu
  2026-08-18  8:36 ` [PATCH bpf-next v2 1/2] mm/bpf: Add bpf_proactive_reclaim kfuncs Hui Zhu
  2026-08-18  8:36 ` [PATCH bpf-next v2 2/2] selftests/bpf: add memcg async reclaim test Hui Zhu
  0 siblings, 2 replies; 5+ messages in thread
From: Hui Zhu @ 2026-08-18  8:36 UTC (permalink / raw)
  To: Roman Gushchin, JP Kobryn, Shakeel Butt, Andrew Morton,
	Andrii Nakryiko, Eduard Zingerman, Ihor Solodrai,
	Alexei Starovoitov, Daniel Borkmann, Kumar Kartikeya Dwivedi,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Shuah Khan, Barry Song, Geliang Tang,
	linux-kernel, bpf, linux-mm, linux-kselftest
  Cc: Hui Zhu

From: Hui Zhu <zhuhui@kylinos.cn>

This series lets a BPF program decide when to trigger memcg reclaim
and how aggressively to do it, based on whatever runtime signal it
chooses to observe -- rather than reclaim only being triggered once a
cgroup's usage crosses a fixed threshold. The core idea is a pair of
new kfuncs, bpf_proactive_reclaim() and
bpf_proactive_reclaim_swappiness(), which give BPF direct access to
the proactive reclaim path so this decision can be made in BPF policy
rather than hard-coded threshold logic.

This was originally part of a larger series posted here [1].
That series also adds a memcg BPF struct_ops (memcg_charged,
memcg_uncharged, below_low, below_min) for synchronous, in-line memory
protection decisions. That mechanism and this one solve different
problems -- struct_ops hooks run inline on the charge/reclaim path,
while the kfuncs here are for asynchronous, out-of-band reclaim
decided independently by a BPF program -- so they are reviewed as
separate series. This series carries only the async reclaim piece.

Compared to v1, the kfunc interface has been reworked based on review
feedback: instead of a thin wrapper around
try_to_free_mem_cgroup_pages() exposing raw gfp/reclaim-option knobs,
the series now provides use-case-driven kfuncs that perform one
proactive reclaim pass with the same parameters memory.reclaim uses.
The bpf_thread_wq patches from v1 (old patches 2-3) are dropped from
this series: following the discussion in [2], the cgroup-aware
workqueue is being superseded by a disaggregated set of async
primitives (bpf_kthread/bpf_waitq) that will be developed separately
(discussion in [3]), and the selftest now queues its reclaim work
through bpf_wq.

Patch 1 adds bpf_proactive_reclaim() and
bpf_proactive_reclaim_swappiness(), sleepable kfuncs that perform one
reclaim pass on a target memcg, like a write to memory.reclaim: swap
is allowed, and the anon/file balance follows the cgroup's swappiness
or an explicit override in [MIN_SWAPPINESS, MAX_SWAPPINESS] plus
SWAPPINESS_ANON_ONLY. Both delegate to try_to_free_mem_cgroup_pages()
with GFP_KERNEL and MEMCG_RECLAIM_MAY_SWAP | MEMCG_RECLAIM_PROACTIVE,
the same parameters user_proactive_reclaim() uses, and unlike
memory.reclaim they do not retry until the requested size is reached.
Both refuse to run when the caller already holds PF_MEMALLOC, since a
nested try_to_free_mem_cgroup_pages() would clobber the outer
reclaim's current->reclaim_state (e.g. MGLRU dereferences
current->reclaim_state->mm_walk).

Patch 2 (selftests/bpf: add memcg async reclaim test) ties the kfuncs
into a worked example: it watches the WORKINGSET_REFAULT_FILE counter
of a high-priority cgroup as a proxy for memory-pressure impact, and
once it starts climbing, proactively reclaims pages from a
low-priority cgroup via bpf_proactive_reclaim(), with the reclaim
work queued asynchronously through bpf_wq. The test asserts that the
monitored cgroup's workload finishes faster once async reclaim kicks
in. This demonstrates the end-to-end use case: BPF observes pressure
on the cgroup it wants to protect, and reclaims from the cgroup it
wants to reclaim from, in one self-contained mechanism. Note that,
without bpf_thread_wq, the CPU cost of the reclaim work is not yet
attributed to a chosen cgroup; that part waits for the async
primitives work mentioned above.

Changelog:
v2:
According to the comments of Shakeel Butt, replace
bpf_try_to_free_mem_cgroup_pages() with
bpf_proactive_reclaim(memcg, size) and
bpf_proactive_reclaim_swappiness(memcg, size, swappiness).
According to the comments of Kumar Kartikeya Dwivedi, drop patch 2
and patch 3.
Remove bpf_thread_wq code in patch 4.
According to the comments of sashiko-bot, fix the issues of selftests.

[1] https://sashiko.dev/#/message/cover.1779760876.git.zhuhui%40kylinos.cn
[2] https://sashiko.dev/#/message/1b58d56976202f26818d31dbd0da2ecb2e2460f5%40linux.dev
[3] https://sashiko.dev/#/message/DKNHV09PBQZP.IRQL20BY574I%40gmail.com

Hui Zhu (2):
  mm/bpf: Add bpf_proactive_reclaim kfuncs
  selftests/bpf: add memcg async reclaim test

 mm/bpf_memcontrol.c                           |  95 +++++
 .../bpf/prog_tests/memcg_async_reclaim.c      | 382 ++++++++++++++++++
 .../selftests/bpf/progs/memcg_async_reclaim.c | 167 ++++++++
 3 files changed, 644 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c
 create mode 100644 tools/testing/selftests/bpf/progs/memcg_async_reclaim.c

-- 
2.53.0



^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-08-18  9:26 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-18  8:36 [PATCH bpf-next v2 0/2] bpf: BPF-driven proactive memcg reclaim Hui Zhu
2026-08-18  8:36 ` [PATCH bpf-next v2 1/2] mm/bpf: Add bpf_proactive_reclaim kfuncs Hui Zhu
2026-08-18  9:26   ` bot+bpf-ci
2026-08-18  8:36 ` [PATCH bpf-next v2 2/2] selftests/bpf: add memcg async reclaim test Hui Zhu
2026-08-18  9:26   ` bot+bpf-ci

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox