BPF List
 help / color / mirror / Atom feed
* [PATCH bpf-next v7 0/2] bpf: BPF-driven proactive memcg reclaim
@ 2026-09-04 10:20 Hui Zhu
  2026-09-04 10:20 ` [PATCH bpf-next v7 1/2] mm/bpf: Add bpf_proactive_reclaim kfunc Hui Zhu
  2026-09-04 10:20 ` [PATCH bpf-next v7 2/2] selftests/bpf: Add memcg async reclaim test Hui Zhu
  0 siblings, 2 replies; 7+ messages in thread
From: Hui Zhu @ 2026-09-04 10:20 UTC (permalink / raw)
  To: Roman Gushchin, JP Kobryn, Shakeel Butt, Andrew Morton,
	Andrii Nakryiko, Eduard Zingerman, Ihor Solodrai,
	Alexei Starovoitov, Daniel Borkmann, Kumar Kartikeya Dwivedi,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Shuah Khan, Barry Song, Geliang Tang,
	linux-kernel, bpf, linux-mm, linux-kselftest
  Cc: Hui Zhu

From: Hui Zhu <zhuhui@kylinos.cn>

BPF programs can observe memory pressure on a cgroup (e.g. refault
stats via bpf_mem_cgroup_page_state()), but cannot act on it:
triggering reclaim on a chosen cgroup requires writing to
memory.reclaim, which BPF cannot do. This series adds
bpf_proactive_reclaim(), a sleepable kfunc performing one proactive
reclaim pass on a target memcg, so when and how hard to reclaim is
BPF policy rather than hard-coded thresholds.

The kfunc is restricted to BPF_PROG_TYPE_SYSCALL so that reclaim
always runs in a clean process context: generic sleepable programs
may execute with filesystem locks held or in NOFS/NOIO contexts,
where the reclaim path could deadlock in filesystem shrinkers. The
bpf_wq and task_work callbacks of a SYSCALL program keep its program
type and run in process context, so reclaim work can still be queued
asynchronously through them, as the selftest does with bpf_wq.

The use case we are looking at is protecting high-priority workloads:
a BPF program monitors the state of a high-priority cgroup and, when
it degrades (e.g. PSI rises or refaults increase, as in the
selftest), asynchronously reclaims memory from low-priority cgroups
via bpf_wq and bpf_proactive_reclaim(), giving the pressured cgroup
more free pages.

Another use case: several vendor-maintained kernels carry private
implementations that trigger asynchronous reclaim when a memcg enters
a certain state. These exist for historical and partly psychological
reasons, but the underlying demand is real. We expect BPF-driven
proactive reclaim, combined with the BPF hooks for the memory
controller currently under discussion and development, to serve these
needs in mainline, reducing kernel fragmentation and improving kernel
maintainability.

Changelog:
v7:
According to the comments of JP, clamp the reclaim target of one
bpf_proactive_reclaim() call to MEMCG_CHARGE_BATCH so each call is
a bounded unit of work, and document the batching policy in the kfunc.
selftest: check the target cgroup for dying state before reclaiming
from it, and add the memcg_async_reclaim_dying test covering target
removal while reclaim is running.
v6:
According to the comments of Kumar and Shakeel, Restrict
bpf_proactive_reclaim() to BPF_PROG_TYPE_SYSCALL by moving it
to a dedicated kfunc set registered for that program type only, and
document the clean-process-context requirement in its kerneldoc.
v5:
According to the comments of Andrii, Kumar and Shakeel, remove
bpf_proactive_reclaim_swappiness.
v4:
According to the comments of bot+bpf-ci and sashiko, also check
current->reclaim_state to close the fentry-on-trace-iter recursion
window in bpf_in_reclaim_context.
Return bytes instead of pages ( nr * PAGE_SIZE ) in
bpf_proactive_reclaim_pages and bpf_proactive_reclaim_swappiness.
Return (unsigned long)-1 on out-of-range swappiness (was 0).
Kdoc of both kfuncs: updated Return descriptions; added FS-lock deadlock
warning to bpf_proactive_reclaim.
Fix potential child process leak in selftests.
Use _exit() instead of exit() in forked children in selftests.
Rename reclaimed_pages to reclaimed_bytes in selftests.
Fix comments issues in selftests.
v3:
According to the comments of bot+bpf-ci, add a shared helper
bpf_proactive_reclaim_pages() that is called by bpf_proactive_reclaim
and bpf_proactive_reclaim_swappiness.
According to the comments of sashiko and bot+bpf-ci, fix the issues of
selftests.
v2:
According to the comments of Shakeel Butt, replace
bpf_try_to_free_mem_cgroup_pages() with
bpf_proactive_reclaim(memcg, size) and
bpf_proactive_reclaim_swappiness(memcg, size, swappiness).
According to the comments of Kumar Kartikeya Dwivedi, drop patch 2
and patch 3.
Remove bpf_thread_wq code in patch 4.
According to the comments of sashiko-bot, fix the issues of selftests.

Hui Zhu (2):
  mm/bpf: Add bpf_proactive_reclaim kfunc
  selftests/bpf: Add memcg async reclaim test

 mm/bpf_memcontrol.c                           |  95 ++-
 .../bpf/prog_tests/memcg_async_reclaim.c      | 686 ++++++++++++++++++
 .../selftests/bpf/progs/memcg_async_reclaim.c | 259 +++++++
 3 files changed, 1038 insertions(+), 2 deletions(-)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c
 create mode 100644 tools/testing/selftests/bpf/progs/memcg_async_reclaim.c

-- 
2.53.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH bpf-next v7 1/2] mm/bpf: Add bpf_proactive_reclaim kfunc
  2026-09-04 10:20 [PATCH bpf-next v7 0/2] bpf: BPF-driven proactive memcg reclaim Hui Zhu
@ 2026-09-04 10:20 ` Hui Zhu
  2026-09-04 10:36   ` sashiko-bot
                     ` (2 more replies)
  2026-09-04 10:20 ` [PATCH bpf-next v7 2/2] selftests/bpf: Add memcg async reclaim test Hui Zhu
  1 sibling, 3 replies; 7+ messages in thread
From: Hui Zhu @ 2026-09-04 10:20 UTC (permalink / raw)
  To: Roman Gushchin, JP Kobryn, Shakeel Butt, Andrew Morton,
	Andrii Nakryiko, Eduard Zingerman, Ihor Solodrai,
	Alexei Starovoitov, Daniel Borkmann, Kumar Kartikeya Dwivedi,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Shuah Khan, Barry Song, Geliang Tang,
	linux-kernel, bpf, linux-mm, linux-kselftest
  Cc: Hui Zhu

From: Hui Zhu <zhuhui@kylinos.cn>

Add bpf_proactive_reclaim(), a sleepable kfunc which performs one
proactive reclaim pass on a given memory cgroup, similar to a write
to memory.reclaim but without retrying until the target is reached.

The kfunc is restricted to BPF_PROG_TYPE_SYSCALL so that reclaim
always runs in a clean process context. Generic sleepable programs
may execute with filesystem locks held or in NOFS/NOIO contexts,
where the reclaim path could deadlock in filesystem shrinkers. A
SYSCALL program can still drive reclaim asynchronously through
bpf_wq or task_work callbacks, which run in process context and
keep the SYSCALL program type, so they can call the kfunc too.

The kfunc refuses to reclaim if the calling task is already in a
reclaim context, as a nested reclaim would corrupt the outer reclaim
state.

The reclaim target of a single call is clamped to MEMCG_CHARGE_BATCH.
try_to_free_mem_cgroup_pages() takes nr_to_reclaim as a lower bound,
so nothing else limits how long one call scans. An unbounded call
stalls other work on a shared workqueue, and since lru_lock is held
with interrupts disabled the contention delays IPI handling and can
cause CSD lock stalls. MEMCG_CHARGE_BATCH is the bound already used
by high_work_func(), so each call becomes a bounded unit of work.

Reclaiming more than one batch is left to the BPF program. This is
documented in the kfunc rather than enforced: call the kfunc once
per bpf_wq callback and requeue the same work item for the next
batch instead of looping inside a callback, and give each target
memcg its own bpf_wq item. Keeping the policy in BPF lets a program
decide when to reclaim and when to stop, e.g. once the target cgroup
is dying.

Signed-off-by: Hui Zhu <zhuhui@kylinos.cn>
---
 mm/bpf_memcontrol.c | 95 ++++++++++++++++++++++++++++++++++++++++++++-
 1 file changed, 93 insertions(+), 2 deletions(-)

diff --git a/mm/bpf_memcontrol.c b/mm/bpf_memcontrol.c
index 716df49d7647..92f35ba66309 100644
--- a/mm/bpf_memcontrol.c
+++ b/mm/bpf_memcontrol.c
@@ -6,6 +6,7 @@
  */
 
 #include <linux/memcontrol.h>
+#include <linux/swap.h>
 #include <linux/bpf.h>
 
 __bpf_kfunc_start_defs();
@@ -159,6 +160,74 @@ __bpf_kfunc void bpf_mem_cgroup_flush_stats(struct mem_cgroup *memcg)
 	mem_cgroup_flush_stats(memcg);
 }
 
+/*
+ * Reclaim must not recurse: try_to_free_mem_cgroup_pages() overwrites
+ * current->reclaim_state, so a nested call would corrupt the outer
+ * reclaim state. Reclaim windows are marked with PF_MEMALLOC;
+ * reclaim_state is also checked because it is installed slightly
+ * before PF_MEMALLOC.
+ */
+static bool bpf_in_reclaim_context(void)
+{
+	return (current->flags & PF_MEMALLOC) || current->reclaim_state;
+}
+
+/**
+ * bpf_proactive_reclaim - proactively reclaim memory from a memory
+ *                         cgroup
+ * @memcg: the target memory cgroup to reclaim from
+ * @size:  the amount of memory to reclaim, in bytes, clamped to
+ *         MEMCG_CHARGE_BATCH (64 pages)
+ *
+ * Trigger one proactive reclaim pass on @memcg, similar to a write to
+ * memory.reclaim, but without retrying until @size is reached.
+ *
+ * @size is clamped so that one call is a bounded unit of work, matching
+ * the memory.high workqueue fallback in high_work_func(). To reclaim
+ * more, call this kfunc repeatedly instead of passing a larger @size.
+ *
+ * This kfunc is restricted to BPF_PROG_TYPE_SYSCALL to ensure it runs
+ * in a clean process context. The SYSCALL program can schedule the
+ * actual reclaim work via bpf_wq or timers, which also execute in
+ * safe process context (workqueue, task_work).
+ *
+ * When reclaim is driven from a bpf_wq, call this kfunc once per
+ * callback and requeue the same work item for the next batch rather
+ * than looping inside the callback: a long-running callback stalls
+ * other work on the shared workqueue, and because lru_lock is held with
+ * interrupts disabled the resulting contention also delays IPI
+ * handling. Give each target memcg its own bpf_wq item, so that
+ * reclaiming one memcg neither serializes behind nor piles up on top of
+ * another. Deciding whether to submit the next batch is up to the BPF
+ * program, which can stop at any point, e.g. once the target cgroup is
+ * dying.
+ *
+ * Must not be called with a filesystem lock held: the reclaim path
+ * may deadlock on it via filesystem shrinkers.
+ *
+ * Return: The amount of memory reclaimed, in bytes, or 0 if @size is
+ * smaller than a page or the task is already in a reclaim context.
+ */
+__bpf_kfunc unsigned long bpf_proactive_reclaim(struct mem_cgroup *memcg,
+						unsigned long size)
+{
+	unsigned long nr_reclaimed;
+	unsigned long nr_pages;
+
+	if (size < PAGE_SIZE || unlikely(bpf_in_reclaim_context()))
+		return 0;
+
+	nr_pages = min(size / PAGE_SIZE, (unsigned long)MEMCG_CHARGE_BATCH);
+
+	nr_reclaimed = try_to_free_mem_cgroup_pages(memcg, nr_pages,
+						    GFP_KERNEL,
+						    MEMCG_RECLAIM_MAY_SWAP |
+						    MEMCG_RECLAIM_PROACTIVE,
+						    NULL);
+
+	return nr_reclaimed * PAGE_SIZE;
+}
+
 __bpf_kfunc_end_defs();
 
 BTF_KFUNCS_START(bpf_memcontrol_kfuncs)
@@ -171,22 +240,44 @@ BTF_ID_FLAGS(func, bpf_mem_cgroup_memory_events)
 BTF_ID_FLAGS(func, bpf_mem_cgroup_usage)
 BTF_ID_FLAGS(func, bpf_mem_cgroup_page_state)
 BTF_ID_FLAGS(func, bpf_mem_cgroup_flush_stats, KF_SLEEPABLE)
-
 BTF_KFUNCS_END(bpf_memcontrol_kfuncs)
 
+/*
+ * Proactive reclaim needs a clean process context, so it is restricted
+ * to BPF_PROG_TYPE_SYSCALL. The bpf_wq and task_work callbacks that a
+ * SYSCALL program schedules run as the same program type, so they can
+ * still invoke it; generic sleepable programs (e.g. fentry on reclaim
+ * paths, inode_rmdir) cannot.
+ */
+BTF_KFUNCS_START(bpf_memcontrol_reclaim_kfuncs)
+BTF_ID_FLAGS(func, bpf_proactive_reclaim, KF_SLEEPABLE)
+BTF_KFUNCS_END(bpf_memcontrol_reclaim_kfuncs)
+
 static const struct btf_kfunc_id_set bpf_memcontrol_kfunc_set = {
 	.owner          = THIS_MODULE,
 	.set            = &bpf_memcontrol_kfuncs,
 };
 
+static const struct btf_kfunc_id_set bpf_memcontrol_reclaim_kfunc_set = {
+	.owner          = THIS_MODULE,
+	.set            = &bpf_memcontrol_reclaim_kfuncs,
+};
+
 static int __init bpf_memcontrol_init(void)
 {
 	int err;
 
 	err = register_btf_kfunc_id_set(BPF_PROG_TYPE_UNSPEC,
 					&bpf_memcontrol_kfunc_set);
-	if (err)
+	if (err) {
 		pr_warn("error while registering bpf memcontrol kfuncs: %d", err);
+		return err;
+	}
+
+	err = register_btf_kfunc_id_set(BPF_PROG_TYPE_SYSCALL,
+					&bpf_memcontrol_reclaim_kfunc_set);
+	if (err)
+		pr_warn("error registering bpf reclaim kfuncs: %d", err);
 
 	return err;
 }
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 7+ messages in thread

* [PATCH bpf-next v7 2/2] selftests/bpf: Add memcg async reclaim test
  2026-09-04 10:20 [PATCH bpf-next v7 0/2] bpf: BPF-driven proactive memcg reclaim Hui Zhu
  2026-09-04 10:20 ` [PATCH bpf-next v7 1/2] mm/bpf: Add bpf_proactive_reclaim kfunc Hui Zhu
@ 2026-09-04 10:20 ` Hui Zhu
  2026-09-04 11:25   ` bot+bpf-ci
  1 sibling, 1 reply; 7+ messages in thread
From: Hui Zhu @ 2026-09-04 10:20 UTC (permalink / raw)
  To: Roman Gushchin, JP Kobryn, Shakeel Butt, Andrew Morton,
	Andrii Nakryiko, Eduard Zingerman, Ihor Solodrai,
	Alexei Starovoitov, Daniel Borkmann, Kumar Kartikeya Dwivedi,
	Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
	Emil Tsalapatis, Shuah Khan, Barry Song, Geliang Tang,
	linux-kernel, bpf, linux-mm, linux-kselftest
  Cc: Hui Zhu

From: Hui Zhu <zhuhui@kylinos.cn>

Add the memcg_async_reclaim selftest, which verifies that BPF-driven
async proactive reclaim mitigates refault-induced slowdown under
memory pressure: a BPF program monitors the refault stats of a
memory-pressured cgroup and, once they grow, asynchronously reclaims
another cgroup via bpf_wq and bpf_proactive_reclaim(), letting the
pressured workload finish faster.

The BPF program also handles a dying reclaim target. Looking the
target up by id is not enough: bpf_cgroup_from_id() keeps handing
back a cgroup until its last reference is dropped, so the program
checks the css flags and skips reclaim once the target is offlined
or dying. A second test, memcg_async_reclaim_dying, keeps reclaim
rounds running against the target, removes the target cgroup while
reclaim is in flight, and verifies that reclaim stops on the removed
target instead of reclaiming from it.

Signed-off-by: Hui Zhu <zhuhui@kylinos.cn>
---
 .../bpf/prog_tests/memcg_async_reclaim.c      | 686 ++++++++++++++++++
 .../selftests/bpf/progs/memcg_async_reclaim.c | 259 +++++++
 2 files changed, 945 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c
 create mode 100644 tools/testing/selftests/bpf/progs/memcg_async_reclaim.c

diff --git a/tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c b/tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c
new file mode 100644
index 000000000000..65f500684463
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c
@@ -0,0 +1,686 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Memory controller eBPF async reclaim test
+ */
+
+#include <test_progs.h>
+#include <sys/mman.h>
+#include <sys/stat.h>
+#include <sys/vfs.h>
+#include <sys/wait.h>
+#include <fcntl.h>
+#include <signal.h>
+#include <time.h>
+#include <unistd.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <limits.h>
+#include <linux/magic.h>
+
+#include "cgroup_helpers.h"
+
+struct bpf_args {
+	u64 high_cgroup_id;
+	u64 low_cgroup_id;
+	u64 event_delta_threshold;
+	u64 check_ns;
+};
+
+#include "memcg_async_reclaim.skel.h"
+
+#define FILE_SIZE (32 * 1024 * 1024ul)
+#define BUFFER_SIZE (4096)
+#define CG_LIMIT (32 * 1024 * 1024ul)
+#define READ_TIMES 50
+
+#define CG_DIR "/memcg_async_reclaim"
+#define CG_HIGH_DIR CG_DIR "/high"
+#define CG_LOW_DIR CG_DIR "/low"
+
+#define CG_DYING_DIR "/memcg_async_reclaim_dying"
+#define CG_DYING_TRIGGER_DIR CG_DYING_DIR "/trigger"
+#define CG_DYING_TARGET_DIR CG_DYING_DIR "/target"
+
+#define CHECK_PERIOD_NS (2 * 1000 * 1000ull)
+#define EVENT_DELTA_THRESHOLD 1
+
+/*
+ * Timing for the dying test: after the target cgroup is removed, give
+ * in-flight reclaim passes time to drain, then wait for a reclaim round
+ * to hit the removed target. The keepalive reader keeps the trigger
+ * cgroup refaulting, and the timer fires every CHECK_PERIOD_NS, so
+ * such a round must show up within a few timer periods.
+ */
+#define DYING_SETTLE_US (200 * 1000)
+#define DYING_POLL_ITERS 500
+#define DYING_POLL_INTERVAL_US (10 * 1000)
+
+/*
+ * The workload files must sit on a regular filesystem: with swap
+ * disabled for the cgroup, tmpfs/ramfs pages are unevictable and would
+ * OOM the cgroup instead of exercising reclaim; they are also charged
+ * as anonymous memory, so they never raise the WORKINGSET_REFAULT_FILE
+ * events the BPF program monitors. Fall back to the current directory
+ * when /tmp is backed by such a filesystem.
+ */
+static const char *workload_files_dir(void)
+{
+	struct statfs st;
+
+	if (!statfs("/tmp", &st) &&
+	    (st.f_type == TMPFS_MAGIC || st.f_type == RAMFS_MAGIC))
+		return ".";
+	return "/tmp";
+}
+
+/*
+ * The workload children run after test_progs hijacked stdio, so
+ * anything they print is lost with their private copy of the hijacked
+ * buffer. The exit status is the only diagnostics channel that reaches
+ * the parent, so each failing step gets its own code.
+ */
+enum child_exit_code {
+	CHILD_EXIT_OK = 0,
+	CHILD_EXIT_JOIN_CGROUP,
+	CHILD_EXIT_WRITE_FILE,
+	CHILD_EXIT_READ_FILE,
+	CHILD_EXIT_TIME_FILE,
+};
+
+static const char *child_exit_str(int code)
+{
+	switch (code) {
+	case CHILD_EXIT_OK:
+		return "success";
+	case CHILD_EXIT_JOIN_CGROUP:
+		return "join cgroup";
+	case CHILD_EXIT_WRITE_FILE:
+		return "write data file";
+	case CHILD_EXIT_READ_FILE:
+		return "read data file";
+	case CHILD_EXIT_TIME_FILE:
+		return "write time file";
+	default:
+		return "unknown";
+	}
+}
+
+static int setup_high_low_cgroups(u64 *high_cgroup_id, u64 *low_cgroup_id)
+{
+	int ret;
+	char limit_buf[20];
+
+	ret = setup_cgroup_environment();
+	if (!ASSERT_OK(ret, "setup_cgroup_environment"))
+		goto cleanup;
+
+	ret = create_and_get_cgroup(CG_DIR);
+	if (!ASSERT_GE(ret, 0, "create_and_get_cgroup " CG_DIR))
+		goto cleanup;
+	close(ret);
+
+	ret = enable_controllers(CG_DIR, "memory");
+	if (!ASSERT_OK(ret, "enable_controllers"))
+		goto cleanup;
+
+	snprintf(limit_buf, sizeof(limit_buf), "%lu", CG_LIMIT);
+	ret = write_cgroup_file(CG_DIR, "memory.max", limit_buf);
+	if (!ASSERT_OK(ret, "write_cgroup_file memory.max"))
+		goto cleanup;
+
+	/*
+	 * Keep the workloads from swapping out. With CONFIG_SWAP=n the
+	 * memory.swap.max file does not exist, and no swap can happen
+	 * anyway, so skip the write.
+	 */
+	if (!access("/proc/swaps", F_OK)) {
+		ret = write_cgroup_file(CG_DIR, "memory.swap.max", "0");
+		if (!ASSERT_OK(ret, "write_cgroup_file memory.swap.max"))
+			goto cleanup;
+	}
+
+	ret = create_and_get_cgroup(CG_HIGH_DIR);
+	if (!ASSERT_GE(ret, 0, "create_and_get_cgroup " CG_HIGH_DIR))
+		goto cleanup;
+	close(ret);
+
+	*high_cgroup_id = get_cgroup_id(CG_HIGH_DIR);
+	if (!ASSERT_GT(*high_cgroup_id, 0, "get_cgroup_id"))
+		goto cleanup;
+
+	ret = create_and_get_cgroup(CG_LOW_DIR);
+	if (!ASSERT_GE(ret, 0, "create_and_get_cgroup " CG_LOW_DIR))
+		goto cleanup;
+	close(ret);
+
+	*low_cgroup_id = get_cgroup_id(CG_LOW_DIR);
+	if (!ASSERT_GT(*low_cgroup_id, 0, "get_cgroup_id"))
+		goto cleanup;
+
+	return 0;
+
+cleanup:
+	cleanup_cgroup_environment();
+	return -1;
+}
+
+/*
+ * The dying test needs an empty reclaim target plus a cgroup that keeps
+ * refaulting while the target is removed, so reclaim rounds keep
+ * starting and run into the removed target. The two have to be separate
+ * cgroups: the target must hold no processes to be removed, and v2's
+ * no-internal-process constraint keeps the refaulting workload out of
+ * any parent that has domain children.
+ */
+static int setup_dying_cgroups(u64 *trigger_cgroup_id, u64 *target_cgroup_id)
+{
+	int ret;
+	char limit_buf[20];
+
+	ret = setup_cgroup_environment();
+	if (!ASSERT_OK(ret, "setup_cgroup_environment"))
+		goto cleanup;
+
+	ret = create_and_get_cgroup(CG_DYING_DIR);
+	if (!ASSERT_GE(ret, 0, "create_and_get_cgroup " CG_DYING_DIR))
+		goto cleanup;
+	close(ret);
+
+	ret = enable_controllers(CG_DYING_DIR, "memory");
+	if (!ASSERT_OK(ret, "enable_controllers"))
+		goto cleanup;
+
+	snprintf(limit_buf, sizeof(limit_buf), "%lu", CG_LIMIT);
+	ret = write_cgroup_file(CG_DYING_DIR, "memory.max", limit_buf);
+	if (!ASSERT_OK(ret, "write_cgroup_file memory.max"))
+		goto cleanup;
+
+	/* See the matching write in setup_high_low_cgroups(). */
+	if (!access("/proc/swaps", F_OK)) {
+		ret = write_cgroup_file(CG_DYING_DIR, "memory.swap.max", "0");
+		if (!ASSERT_OK(ret, "write_cgroup_file memory.swap.max"))
+			goto cleanup;
+	}
+
+	ret = create_and_get_cgroup(CG_DYING_TRIGGER_DIR);
+	if (!ASSERT_GE(ret, 0, "create_and_get_cgroup " CG_DYING_TRIGGER_DIR))
+		goto cleanup;
+	close(ret);
+
+	*trigger_cgroup_id = get_cgroup_id(CG_DYING_TRIGGER_DIR);
+	if (!ASSERT_GT(*trigger_cgroup_id, 0, "get_cgroup_id"))
+		goto cleanup;
+
+	ret = create_and_get_cgroup(CG_DYING_TARGET_DIR);
+	if (!ASSERT_GE(ret, 0, "create_and_get_cgroup " CG_DYING_TARGET_DIR))
+		goto cleanup;
+	close(ret);
+
+	*target_cgroup_id = get_cgroup_id(CG_DYING_TARGET_DIR);
+	if (!ASSERT_GT(*target_cgroup_id, 0, "get_cgroup_id"))
+		goto cleanup;
+
+	return 0;
+
+cleanup:
+	cleanup_cgroup_environment();
+	return -1;
+}
+
+static int write_file(const char *filename)
+{
+	int ret = -1;
+	size_t written = 0;
+	char *buffer;
+	FILE *fp;
+
+	fp = fopen(filename, "wb");
+	if (!fp)
+		goto out;
+
+	buffer = malloc(BUFFER_SIZE);
+	if (!buffer)
+		goto cleanup_fp;
+
+	memset(buffer, 'A', BUFFER_SIZE);
+
+	while (written < FILE_SIZE) {
+		size_t to_write = FILE_SIZE - written < BUFFER_SIZE ?
+				  FILE_SIZE - written : BUFFER_SIZE;
+
+		if (fwrite(buffer, 1, to_write, fp) != to_write)
+			goto cleanup;
+		written += to_write;
+	}
+
+	ret = 0;
+cleanup:
+	free(buffer);
+cleanup_fp:
+	fclose(fp);
+out:
+	return ret;
+}
+
+static int read_file(const char *filename, int iterations)
+{
+	int ret = -1;
+	long page_size = sysconf(_SC_PAGESIZE);
+	char *map;
+	size_t i;
+	int fd;
+	struct stat sb;
+
+	fd = open(filename, O_RDONLY);
+	if (fd == -1)
+		goto out;
+
+	if (fstat(fd, &sb) == -1)
+		goto cleanup_fd;
+
+	if (sb.st_size != FILE_SIZE) {
+		fprintf(stderr, "File size mismatch: expected %lu, got %lu\n",
+			(unsigned long)FILE_SIZE, (unsigned long)sb.st_size);
+		goto cleanup_fd;
+	}
+
+	map = mmap(NULL, FILE_SIZE, PROT_READ, MAP_PRIVATE, fd, 0);
+	if (map == MAP_FAILED)
+		goto cleanup_fd;
+
+	for (int iter = 0; iter < iterations; iter++) {
+		for (i = 0; i < FILE_SIZE; i += page_size) {
+			/* access a byte to trigger page fault */
+			volatile char v = map[i];
+			(void)v;
+		}
+	}
+
+	if (munmap(map, FILE_SIZE) == -1)
+		goto cleanup_fd;
+
+	ret = 0;
+
+cleanup_fd:
+	close(fd);
+out:
+	return ret;
+}
+
+static int real_test_child_work(const char *cgroup_path, char *data_filename,
+				char *time_filename, int read_times)
+{
+	struct timespec start, end;
+	double elapsed;
+	FILE *fp;
+
+	if (join_parent_cgroup(cgroup_path))
+		return CHILD_EXIT_JOIN_CGROUP;
+
+	clock_gettime(CLOCK_MONOTONIC, &start);
+
+	if (write_file(data_filename))
+		return CHILD_EXIT_WRITE_FILE;
+
+	if (read_file(data_filename, read_times))
+		return CHILD_EXIT_READ_FILE;
+
+	clock_gettime(CLOCK_MONOTONIC, &end);
+
+	if (!time_filename)
+		return CHILD_EXIT_OK;
+
+	elapsed = (end.tv_sec - start.tv_sec) +
+		  (end.tv_nsec - start.tv_nsec) / 1000000000.0;
+	printf("%.6f\n", elapsed);
+
+	fp = fopen(time_filename, "w");
+	if (!fp)
+		return CHILD_EXIT_TIME_FILE;
+	fprintf(fp, "%.6f", elapsed);
+	fclose(fp);
+
+	return CHILD_EXIT_OK;
+}
+
+static int get_time(char *time_filename, double *time)
+{
+	int ret = -1;
+	FILE *fp;
+	char buf[64];
+
+	fp = fopen(time_filename, "r");
+	if (!ASSERT_OK_PTR(fp, "fopen"))
+		goto out;
+
+	if (!ASSERT_OK_PTR(fgets(buf, sizeof(buf), fp), "fgets"))
+		goto cleanup;
+
+	if (sscanf(buf, "%lf", time) != 1) {
+		PRINT_FAIL("sscanf %s", buf);
+		goto cleanup;
+	}
+
+	ret = 0;
+cleanup:
+	fclose(fp);
+out:
+	return ret;
+}
+
+static int
+run_high_low_workload(double *high_elapsed, double *low_elapsed, int read_times)
+{
+	char high_data_file[PATH_MAX];
+	char low_data_file[PATH_MAX];
+	char high_time_file[PATH_MAX];
+	char low_time_file[PATH_MAX];
+	const char *dir = workload_files_dir();
+	pid_t high_pid = -1, low_pid = -1;
+	pid_t wait_ret;
+	int fd, status;
+	int ret = -1;
+
+	snprintf(high_data_file, sizeof(high_data_file),
+		 "%s/memcg_async_high_data_XXXXXX", dir);
+	snprintf(low_data_file, sizeof(low_data_file),
+		 "%s/memcg_async_low_data_XXXXXX", dir);
+	snprintf(high_time_file, sizeof(high_time_file),
+		 "%s/memcg_async_high_time_XXXXXX", dir);
+	snprintf(low_time_file, sizeof(low_time_file),
+		 "%s/memcg_async_low_time_XXXXXX", dir);
+
+	fd = mkstemp(high_data_file);
+	if (!ASSERT_GE(fd, 0, "mkstemp"))
+		goto cleanup;
+	close(fd);
+
+	fd = mkstemp(low_data_file);
+	if (!ASSERT_GE(fd, 0, "mkstemp"))
+		goto cleanup;
+	close(fd);
+
+	fd = mkstemp(high_time_file);
+	if (!ASSERT_GE(fd, 0, "mkstemp"))
+		goto cleanup;
+	close(fd);
+
+	fd = mkstemp(low_time_file);
+	if (!ASSERT_GE(fd, 0, "mkstemp"))
+		goto cleanup;
+	close(fd);
+
+	low_pid = fork();
+	if (!ASSERT_GE(low_pid, 0, "fork low"))
+		goto cleanup;
+	if (low_pid == 0)
+		_exit(real_test_child_work(CG_LOW_DIR, low_data_file,
+					  low_time_file, read_times));
+
+	high_pid = fork();
+	if (!ASSERT_GE(high_pid, 0, "fork high"))
+		goto cleanup;
+	if (high_pid == 0)
+		_exit(real_test_child_work(CG_HIGH_DIR, high_data_file,
+					  high_time_file, read_times));
+
+	wait_ret = waitpid(low_pid, &status, 0);
+	if (!ASSERT_GT(wait_ret, 0, "low waitpid"))
+		goto cleanup;
+	/*
+	 * The child has been reaped and its PID can already be reused,
+	 * so mark it to keep cleanup from signaling an unrelated process.
+	 */
+	low_pid = -1;
+	if (!ASSERT_TRUE(WIFEXITED(status), "low exited"))
+		goto cleanup;
+	if (WEXITSTATUS(status) != CHILD_EXIT_OK) {
+		PRINT_FAIL("low child failed at: %s (exit status %d)",
+			   child_exit_str(WEXITSTATUS(status)),
+			   WEXITSTATUS(status));
+		goto cleanup;
+	}
+
+	wait_ret = waitpid(high_pid, &status, 0);
+	if (!ASSERT_GT(wait_ret, 0, "high waitpid"))
+		goto cleanup;
+	/* Same as above: the reaped PID must not be signaled again. */
+	high_pid = -1;
+	if (!ASSERT_TRUE(WIFEXITED(status), "high exited"))
+		goto cleanup;
+	if (WEXITSTATUS(status) != CHILD_EXIT_OK) {
+		PRINT_FAIL("high child failed at: %s (exit status %d)",
+			   child_exit_str(WEXITSTATUS(status)),
+			   WEXITSTATUS(status));
+		goto cleanup;
+	}
+
+	if (get_time(high_time_file, high_elapsed))
+		goto cleanup;
+	if (get_time(low_time_file, low_elapsed))
+		goto cleanup;
+
+	ret = 0;
+
+cleanup:
+	/* On failure, make sure no child process is left behind */
+	if (ret) {
+		if (high_pid > 0) {
+			kill(high_pid, SIGKILL);
+			(void)waitpid(high_pid, NULL, 0);
+		}
+		if (low_pid > 0) {
+			kill(low_pid, SIGKILL);
+			(void)waitpid(low_pid, NULL, 0);
+		}
+	}
+	unlink(low_time_file);
+	unlink(high_time_file);
+	unlink(low_data_file);
+	unlink(high_data_file);
+	return ret;
+}
+
+static int
+setup_bpf(u64 high_cgroup_id, u64 low_cgroup_id,
+	  struct memcg_async_reclaim **skel_ptr)
+{
+	struct memcg_async_reclaim *skel;
+	struct bpf_args args = {
+		.high_cgroup_id = high_cgroup_id,
+		.low_cgroup_id = low_cgroup_id,
+		.event_delta_threshold = EVENT_DELTA_THRESHOLD,
+		.check_ns = CHECK_PERIOD_NS,
+	};
+	LIBBPF_OPTS(bpf_test_run_opts, run_opts,
+		.ctx_in = &args,
+		.ctx_size_in = sizeof(args));
+	int prog_init_fd, err;
+
+	skel = memcg_async_reclaim__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "memcg_async_reclaim__open_and_load"))
+		return -1;
+
+	prog_init_fd = bpf_program__fd(skel->progs.wq_prog_init);
+
+	err = bpf_prog_test_run_opts(prog_init_fd, &run_opts);
+	if (!ASSERT_OK(err, "bpf_prog_test_run_opts"))
+		goto error_out;
+	if (!ASSERT_EQ(run_opts.retval, 0, "prog_init retval"))
+		goto error_out;
+
+	*skel_ptr = skel;
+	return 0;
+
+error_out:
+	memcg_async_reclaim__destroy(skel);
+	return -1;
+}
+
+void test_memcg_async_reclaim(void)
+{
+	u64 high_cgroup_id, low_cgroup_id;
+	int err;
+	double high_time = 0.0, low_time = 0.0;
+	struct memcg_async_reclaim *skel = NULL;
+
+	err = setup_high_low_cgroups(&high_cgroup_id, &low_cgroup_id);
+	if (!ASSERT_OK(err, "setup_high_low_cgroups reclaim"))
+		return;
+
+	err = setup_bpf(high_cgroup_id, low_cgroup_id, &skel);
+	if (!ASSERT_OK(err, "setup_bpf"))
+		goto out;
+
+	err = run_high_low_workload(&high_time, &low_time, READ_TIMES);
+	if (!ASSERT_OK(err, "run_high_low_workload reclaim"))
+		goto out;
+
+	/*
+	 * The timing comparison below alone cannot distinguish a working
+	 * reclaim from a no-op one, so require that the BPF program
+	 * actually reclaimed memory from the low cgroup.
+	 */
+	if (!ASSERT_GT(skel->bss->reclaim_calls, 0, "reclaim_calls"))
+		goto out;
+	if (!ASSERT_GT(skel->bss->reclaimed_bytes, 0, "reclaimed_bytes"))
+		goto out;
+
+	if (high_time >= low_time)
+		PRINT_FAIL("high cgroup not improved: high=%f low=%f",
+			   high_time, low_time);
+
+out:
+	if (skel)
+		memcg_async_reclaim__destroy(skel);
+	cleanup_cgroup_environment();
+}
+
+/*
+ * Keep refaults flowing through the trigger cgroup so reclaim rounds
+ * keep being triggered while the target cgroup is being removed. The
+ * child joins the trigger cgroup and writes the data file there, so
+ * that the file pages are charged to the trigger cgroup and actually
+ * come under its memory limit; then it re-reads the file in a loop
+ * until it is killed.
+ */
+static pid_t spawn_keepalive_reader(const char *data_file)
+{
+	pid_t pid = fork();
+
+	if (pid != 0)
+		return pid;
+
+	if (join_parent_cgroup(CG_DYING_TRIGGER_DIR))
+		_exit(CHILD_EXIT_JOIN_CGROUP);
+	if (write_file(data_file))
+		_exit(CHILD_EXIT_WRITE_FILE);
+	for (;;) {
+		if (read_file(data_file, READ_TIMES))
+			_exit(CHILD_EXIT_READ_FILE);
+	}
+}
+
+/*
+ * Remove the reclaim target while the BPF program keeps running and
+ * verify that reclaim stops on the dying/removed cgroup instead of
+ * reclaiming from it.
+ *
+ * The target stays empty; the workload lives in the trigger cgroup and
+ * only keeps refaults flowing so that reclaim rounds keep starting,
+ * both before and after the target is removed. reclaim_calls growing
+ * while the target is alive proves that rounds really run (the kfunc
+ * returns 0 on the empty target, but the call is still counted), and
+ * after the removal the skip counters must grow while reclaim_calls
+ * and reclaimed_bytes stay frozen.
+ */
+void test_memcg_async_reclaim_dying(void)
+{
+	u64 trigger_cgroup_id, target_cgroup_id;
+	u64 calls_before, bytes_before;
+	char data_file[PATH_MAX] = "";
+	struct memcg_async_reclaim *skel = NULL;
+	pid_t reader_pid = -1;
+	int err, fd, i;
+
+	err = setup_dying_cgroups(&trigger_cgroup_id, &target_cgroup_id);
+	if (!ASSERT_OK(err, "setup_dying_cgroups"))
+		return;
+
+	err = setup_bpf(trigger_cgroup_id, target_cgroup_id, &skel);
+	if (!ASSERT_OK(err, "setup_bpf"))
+		goto out;
+
+	snprintf(data_file, sizeof(data_file),
+		 "%s/memcg_async_dying_XXXXXX", workload_files_dir());
+	fd = mkstemp(data_file);
+	if (!ASSERT_GE(fd, 0, "mkstemp"))
+		goto out;
+	close(fd);
+
+	reader_pid = spawn_keepalive_reader(data_file);
+	if (!ASSERT_GT(reader_pid, 0, "fork keepalive reader"))
+		goto out;
+
+	/* Wait for reclaim rounds to reach the live target cgroup. */
+	for (i = 0; i < DYING_POLL_ITERS; i++) {
+		if (skel->bss->reclaim_calls > 0)
+			break;
+		usleep(DYING_POLL_INTERVAL_US);
+	}
+	if (!ASSERT_GT(skel->bss->reclaim_calls, 0, "reclaim_calls"))
+		goto out;
+
+	remove_cgroup(CG_DYING_TARGET_DIR);
+
+	/* Let reclaim passes that were already in flight drain. */
+	usleep(DYING_SETTLE_US);
+
+	calls_before = skel->bss->reclaim_calls;
+	bytes_before = skel->bss->reclaimed_bytes;
+
+	/* Wait for reclaim rounds to hit the removed cgroup. */
+	for (i = 0; i < DYING_POLL_ITERS; i++) {
+		if (skel->bss->reclaim_target_gone ||
+		    skel->bss->reclaim_skipped_dying)
+			break;
+		usleep(DYING_POLL_INTERVAL_US);
+	}
+
+	if (!skel->bss->reclaim_target_gone &&
+	    !skel->bss->reclaim_skipped_dying) {
+		PRINT_FAIL("no reclaim round hit the removed cgroup (gone=%llu, dying=%llu)",
+			   (unsigned long long)skel->bss->reclaim_target_gone,
+			   (unsigned long long)skel->bss->reclaim_skipped_dying);
+		goto out;
+	}
+
+	/*
+	 * reclaim_skipped_dying shows that the CSS_DYING/CSS_ONLINE check
+	 * caught the cgroup mid-teardown. Whether it is hit is timing
+	 * dependent, because the cgroup may already be fully released, so
+	 * only the combined skip count above is asserted.
+	 */
+	printf("memcg_async_reclaim_dying: skips on removed cgroup: gone=%llu, dying=%llu\n",
+	       (unsigned long long)skel->bss->reclaim_target_gone,
+	       (unsigned long long)skel->bss->reclaim_skipped_dying);
+
+	/* Nothing may have been reclaimed from the removed target. */
+	if (!ASSERT_EQ(skel->bss->reclaim_calls, calls_before, "reclaim_calls"))
+		goto out;
+	if (!ASSERT_EQ(skel->bss->reclaimed_bytes, bytes_before,
+		       "reclaimed_bytes"))
+		goto out;
+
+out:
+	if (reader_pid > 0) {
+		kill(reader_pid, SIGKILL);
+		(void)waitpid(reader_pid, NULL, 0);
+	}
+	if (data_file[0])
+		unlink(data_file);
+	if (skel)
+		memcg_async_reclaim__destroy(skel);
+	cleanup_cgroup_environment();
+}
diff --git a/tools/testing/selftests/bpf/progs/memcg_async_reclaim.c b/tools/testing/selftests/bpf/progs/memcg_async_reclaim.c
new file mode 100644
index 000000000000..e6839ade472b
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/memcg_async_reclaim.c
@@ -0,0 +1,259 @@
+// SPDX-License-Identifier: GPL-2.0
+
+#include "vmlinux.h"
+#include "bpf_experimental.h"
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+#include <bpf/bpf_core_read.h>
+
+#define CLOCK_MONOTONIC_ID	1
+#define PAGE_SIZE		4096UL
+/*
+ * One reclaim round targets RECLAIM_MAX_ITER batches of RECLAIM_SIZE
+ * each. Each bpf_wq callback reclaims a single batch and requeues the
+ * same work item for the next one, so no callback runs longer than one
+ * bounded reclaim pass.
+ */
+#define RECLAIM_SIZE		(32 * PAGE_SIZE)
+#define RECLAIM_MAX_ITER	32
+
+struct bpf_args {
+	u64 high_cgroup_id;
+	u64 low_cgroup_id;
+	u64 event_delta_threshold;
+	u64 check_ns;
+};
+
+struct cgroup_memcg {
+	struct cgroup *cgrp;
+	struct mem_cgroup *memcg;
+};
+
+static u64 wq_high_cgroup_id;
+static u64 wq_low_cgroup_id;
+
+/*
+ * Statistics exposed to userspace through .bss, so the test can verify
+ * that reclaim actually happened instead of relying on timing alone.
+ */
+u64 reclaim_calls;
+u64 reclaimed_bytes;
+/*
+ * Reclaim attempts skipped because the target cgroup is dying or has
+ * been removed. reclaim_skipped_dying counts lookups that still found
+ * the cgroup while it is being torn down, reclaim_target_gone counts
+ * lookups that found nothing. The test removes the target cgroup while
+ * reclaim is running and checks that reclaim stops via these counters.
+ */
+u64 reclaim_skipped_dying;
+u64 reclaim_target_gone;
+
+static int get_cgroup_memcg_from_id(u64 cgroup_id, struct cgroup_memcg *cm)
+{
+	cm->cgrp = bpf_cgroup_from_id(cgroup_id);
+	if (!cm->cgrp)
+		return -1;
+
+	cm->memcg = bpf_get_mem_cgroup(&cm->cgrp->self);
+	if (!cm->memcg) {
+		bpf_cgroup_release(cm->cgrp);
+		return -1;
+	}
+
+	return 0;
+}
+
+static void put_cgroup_memcg(struct cgroup_memcg *cm)
+{
+	bpf_put_mem_cgroup(cm->memcg);
+	bpf_cgroup_release(cm->cgrp);
+}
+
+static int get_cgroup_event(u64 cgroup_id, u64 *val)
+{
+	struct cgroup_memcg cm;
+
+	if (get_cgroup_memcg_from_id(cgroup_id, &cm))
+		return -1;
+	bpf_mem_cgroup_flush_stats(cm.memcg);
+	*val = bpf_mem_cgroup_page_state(cm.memcg,
+		bpf_core_enum_value(enum node_stat_item,
+				    WORKINGSET_REFAULT_FILE));
+	put_cgroup_memcg(&cm);
+
+	return 0;
+}
+
+static bool
+should_reclaim_cgroup(u64 cgroup_id, u64 *prev_event, u64 event_delta_threshold)
+{
+	u64 cur, delta;
+
+	if (get_cgroup_event(cgroup_id, &cur))
+		return false;
+
+	delta = cur - *prev_event;
+	*prev_event = cur;
+
+	return delta >= event_delta_threshold;
+}
+
+/*
+ * A cgroup is dying once it has been offlined (CSS_ONLINE cleared) or
+ * CSS_DYING has been raised, mirroring cgroup_is_dead()/css_is_dying()
+ * in include/linux/cgroup.h. bpf_cgroup_from_id() can still hand back
+ * such a cgroup, because it only fails once the last reference has been
+ * dropped, so reclaim has to check these flags instead of relying on
+ * the lookup failing.
+ *
+ * CSS_ONLINE and CSS_DYING come from vmlinux.h: the kernel defines them
+ * in an anonymous enum, so bpf_core_enum_value() has no enum type to
+ * bind to, and redeclaring them locally would clash with the vmlinux.h
+ * enumerators. vmlinux.h is generated from the running kernel's BTF, so
+ * the values already match the target kernel.
+ */
+static bool cgroup_is_dying(struct cgroup *cgrp)
+{
+	unsigned int flags = cgrp->self.flags;
+
+	return (flags & CSS_DYING) || !(flags & CSS_ONLINE);
+}
+
+/*
+ * Reclaim one batch from the target cgroup. Returns the number of
+ * bytes reclaimed, or 0 if the cgroup is dying or gone or nothing was
+ * reclaimed.
+ */
+static u64 reclaim_cgroup(u64 cgroup_id, u64 size)
+{
+	struct cgroup_memcg cm;
+	u64 nr = 0;
+
+	if (get_cgroup_memcg_from_id(cgroup_id, &cm)) {
+		reclaim_target_gone++;
+		return 0;
+	}
+
+	if (cgroup_is_dying(cm.cgrp)) {
+		reclaim_skipped_dying++;
+		put_cgroup_memcg(&cm);
+		return 0;
+	}
+
+	reclaim_calls++;
+	nr = bpf_proactive_reclaim(cm.memcg, size);
+	reclaimed_bytes += nr;
+
+	put_cgroup_memcg(&cm);
+
+	return nr;
+}
+
+struct wq_elem {
+	struct bpf_timer timer;
+	struct bpf_wq work;
+	u64 prev_event;
+	u64 event_delta_threshold;
+	u64 check_ns;
+	/*
+	 * Bytes still to reclaim in the current round, carried across
+	 * requeues. 0 means no round is in progress; the timer path
+	 * starts a new round by resetting it, requeued work only looks
+	 * at it.
+	 */
+	u64 remaining;
+};
+
+struct {
+	__uint(type, BPF_MAP_TYPE_ARRAY);
+	__uint(max_entries, 1);
+	__type(key, __u32);
+	__type(value, struct wq_elem);
+} wq_map SEC(".maps");
+
+static int reclaim_work_fn(void *map, int *key, void *value)
+{
+	struct wq_elem *elem = value;
+	u64 nr, size;
+
+	if (!elem->remaining) {
+		/*
+		 * Timer-triggered entry: start a new round only when the
+		 * high cgroup refaults enough. Requeued entries skip this
+		 * check and only look at remaining, so the refault delta
+		 * is consumed once per round.
+		 */
+		if (!should_reclaim_cgroup(wq_high_cgroup_id, &elem->prev_event,
+			elem->event_delta_threshold))
+			return 0;
+		elem->remaining = RECLAIM_MAX_ITER * RECLAIM_SIZE;
+	}
+
+	/* One bounded reclaim pass per callback */
+	size = elem->remaining < RECLAIM_SIZE ? elem->remaining : RECLAIM_SIZE;
+	nr = reclaim_cgroup(wq_low_cgroup_id, size);
+	if (!nr) {
+		elem->remaining = 0;
+		return 0;
+	}
+
+	/* try_to_free_mem_cgroup_pages() may reclaim more than requested */
+	if (nr >= elem->remaining)
+		elem->remaining = 0;
+	else
+		elem->remaining -= nr;
+
+	/* Requeue the same work item for the next batch */
+	if (elem->remaining)
+		bpf_wq_start(&elem->work, 0);
+
+	return 0;
+}
+
+static int wq_timer_cb(void *map, int *key, struct wq_elem *elem)
+{
+	bpf_wq_start(&elem->work, 0);
+	bpf_timer_start(&elem->timer, elem->check_ns, 0);
+
+	return 0;
+}
+
+SEC("syscall")
+int wq_prog_init(struct bpf_args *ctx)
+{
+	struct wq_elem *elem;
+	__u32 key = 0;
+	int ret;
+
+	elem = bpf_map_lookup_elem(&wq_map, &key);
+	if (!elem)
+		return -1;
+
+	ret = bpf_wq_init(&elem->work, &wq_map, 0);
+	if (ret)
+		return ret;
+
+	ret = bpf_wq_set_callback(&elem->work, reclaim_work_fn, 0);
+	if (ret)
+		return ret;
+
+	ret = bpf_timer_init(&elem->timer, &wq_map, CLOCK_MONOTONIC_ID);
+	if (ret)
+		return ret;
+
+	ret = bpf_timer_set_callback(&elem->timer, wq_timer_cb);
+	if (ret)
+		return ret;
+
+	elem->prev_event = 0;
+	elem->remaining = 0;
+	elem->event_delta_threshold = ctx->event_delta_threshold;
+	elem->check_ns = ctx->check_ns;
+
+	wq_high_cgroup_id = ctx->high_cgroup_id;
+	wq_low_cgroup_id = ctx->low_cgroup_id;
+
+	return bpf_timer_start(&elem->timer, elem->check_ns, 0);
+}
+
+char _license[] SEC("license") = "GPL";
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next v7 1/2] mm/bpf: Add bpf_proactive_reclaim kfunc
  2026-09-04 10:20 ` [PATCH bpf-next v7 1/2] mm/bpf: Add bpf_proactive_reclaim kfunc Hui Zhu
@ 2026-09-04 10:36   ` sashiko-bot
  2026-09-04 11:25   ` bot+bpf-ci
  2026-09-04 18:03   ` Andrew Morton
  2 siblings, 0 replies; 7+ messages in thread
From: sashiko-bot @ 2026-09-04 10:36 UTC (permalink / raw)
  To: Hui Zhu; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:

Pre-existing issues:
- [High] The kfuncs `bpf_mem_cgroup_memory_events` and `bpf_mem_cgroup_vm_events` perform out-of-bounds array reads when negative enum values are provided.
--

commit d1e15e132a96633de13a710a4e0aff1aa016fbaa
Author: Hui Zhu <zhuhui@kylinos.cn>

mm/bpf: Add bpf_proactive_reclaim kfunc

Add bpf_proactive_reclaim(), a sleepable kfunc which performs one
proactive reclaim pass on a given memory cgroup, similar to a write
to memory.reclaim but without retrying until the target is reached.
The kfunc is restricted to BPF_PROG_TYPE_SYSCALL to ensure it runs
in a clean process context.

> diff --git a/mm/bpf_memcontrol.c b/mm/bpf_memcontrol.c
> index 716df49d76477..92f35ba66309e 100644
> --- a/mm/bpf_memcontrol.c
> +++ b/mm/bpf_memcontrol.c
[ ... ]
> @@ -171,22 +240,44 @@ BTF_ID_FLAGS(func, bpf_mem_cgroup_memory_events)

[Severity: High]
This is a pre-existing issue, but do the implementations of
bpf_mem_cgroup_memory_events() and bpf_mem_cgroup_vm_events() handle negative
enum values correctly?

The bounds check in bpf_mem_cgroup_memory_events() does not cast to unsigned,
so a negative index bypasses the check:

mm/bpf_memcontrol.c:bpf_mem_cgroup_memory_events() {
	if (unlikely(event >= MEMCG_NR_MEMORY_EVENTS))
		return (unsigned long)-1;

	return atomic_long_read(&memcg->memory_events[event]);
}

Because the BPF verifier does not validate that scalar arguments fall within
valid enum ranges, a negative value passed from a BPF program can bypass
this check and cause an out-of-bounds read on the memory_events array.
This can overlap with the vmstats pointer in struct mem_cgroup, potentially
leaking a kernel heap pointer to the BPF program.

Similarly, bpf_mem_cgroup_vm_events() also appears to be affected, where
the upper-bounds check fails to filter out negative values before calling
memcg_events_index():

mm/memcontrol.c:memcg_vm_event_item_valid() {
	if (idx >= NR_VM_EVENT_ITEMS)
		return false;

	return !BAD_STAT_IDX(memcg_events_index(idx));
}

Which leads to an out-of-bounds array access using the negative index:

mm/memcontrol.c:memcg_events_index() {
	return mem_cgroup_events_index[idx];
}

Could we add proper lower bound checks or cast the index to unsigned before
comparing against the array size in these functions?

>  BTF_ID_FLAGS(func, bpf_mem_cgroup_usage)
>  BTF_ID_FLAGS(func, bpf_mem_cgroup_page_state)
>  BTF_ID_FLAGS(func, bpf_mem_cgroup_flush_stats, KF_SLEEPABLE)

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1788516769.git.zhuhui@kylinos.cn?part=1

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next v7 1/2] mm/bpf: Add bpf_proactive_reclaim kfunc
  2026-09-04 10:20 ` [PATCH bpf-next v7 1/2] mm/bpf: Add bpf_proactive_reclaim kfunc Hui Zhu
  2026-09-04 10:36   ` sashiko-bot
@ 2026-09-04 11:25   ` bot+bpf-ci
  2026-09-04 18:03   ` Andrew Morton
  2 siblings, 0 replies; 7+ messages in thread
From: bot+bpf-ci @ 2026-09-04 11:25 UTC (permalink / raw)
  To: hui.zhu, roman.gushchin, inwardvessel, shakeel.butt, akpm, andrii,
	eddyz87, ihor.solodrai, ast, daniel, memxor, martin.lau, song,
	yonghong.song, jolsa, emil, shuah, baohua, geliang, linux-kernel,
	bpf, linux-mm, linux-kselftest
  Cc: zhuhui, ast, andrii, daniel, martin.lau, eddyz87, yonghong.song,
	mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 4023 bytes --]

> diff --git a/mm/bpf_memcontrol.c b/mm/bpf_memcontrol.c
> index 716df49d76477..92f35ba66309e 100644
> --- a/mm/bpf_memcontrol.c
> +++ b/mm/bpf_memcontrol.c
> @@ -6,6 +6,7 @@
>   */
>
>  #include <linux/memcontrol.h>
> +#include <linux/swap.h>
>  #include <linux/bpf.h>
>
>  __bpf_kfunc_start_defs();

[ ... ]

> +/**
> + * bpf_proactive_reclaim - proactively reclaim memory from a memory
> + *                         cgroup
> + * @memcg: the target memory cgroup to reclaim from
> + * @size:  the amount of memory to reclaim, in bytes, clamped to
> + *         MEMCG_CHARGE_BATCH (64 pages)
> + *
> + * Trigger one proactive reclaim pass on @memcg, similar to a write to
> + * memory.reclaim, but without retrying until @size is reached.
> + *
> + * @size is clamped so that one call is a bounded unit of work, matching
> + * the memory.high workqueue fallback in high_work_func(). To reclaim
> + * more, call this kfunc repeatedly instead of passing a larger @size.
> + *
> + * This kfunc is restricted to BPF_PROG_TYPE_SYSCALL to ensure it runs
> + * in a clean process context. The SYSCALL program can schedule the
> + * actual reclaim work via bpf_wq or timers, which also execute in
> + * safe process context (workqueue, task_work).

Does the kernel-doc accurately describe when this kfunc can be called?

The doc says "bpf_wq or timers" can call this kfunc and attributes "safe
process context (workqueue, task_work)" to both. bpf_timer callbacks run
in softirq context (HRTIMER_MODE_REL_SOFT in kernel/bpf/helpers.c), not
process context, and the verifier rejects the documented pattern:

kernel/bpf/verifier.c:is_async_cb_sleepable() {
    /* bpf_timer callbacks are never sleepable. */
    if (bpf_helper_call(insn) && insn->imm == BPF_FUNC_timer_set_callback)
            return false;

which then causes check_kfunc_call() to reject the KF_SLEEPABLE kfunc:

    if (sleepable && !in_sleepable(env)) {
            verbose(env, "program must be sleepable to call sleepable kfunc %s\n", func_name);
            return -EACCES;

So a bpf_timer callback calling bpf_proactive_reclaim() fails verification
regardless of the program being marked sleepable.

The commit message says "A SYSCALL program can still drive reclaim
asynchronously through bpf_wq or task_work callbacks", and the block
comment above BTF_KFUNCS_START(bpf_memcontrol_reclaim_kfuncs) says "The
bpf_wq and task_work callbacks that a SYSCALL program schedules run as the
same program type". Only the kernel-doc substitutes "timers" for
task_work, so the three descriptions disagree.

The follow-up selftest (5ed7b6799588 "selftests/bpf: Add memcg async
reclaim test") demonstrates the actual constraint: its bpf_timer callback
wq_timer_cb() only calls bpf_wq_start()/bpf_timer_start(), and
bpf_proactive_reclaim() is called solely from the bpf_wq callback
reclaim_work_fn().

Could "or timers" be dropped (or reworded to "a bpf_timer callback can
only kick a bpf_wq; the kfunc itself must be called from the bpf_wq or
task_work callback") so the kernel-doc matches both the verifier and the
other two descriptions in the patch?

> + *
> + * When reclaim is driven from a bpf_wq, call this kfunc once per
> + * callback and requeue the same work item for the next batch rather
> + * than looping inside the callback: a long-running callback stalls
> + * other work on the shared workqueue, and because lru_lock is held with
> + * interrupts disabled the resulting contention also delays IPI
> + * handling. Give each target memcg its own bpf_wq item, so that
> + * reclaiming one memcg neither serializes behind nor piles up on top of
> + * another. Deciding whether to submit the next batch is up to the BPF
> + * program, which can stop at any point, e.g. once the target cgroup is
> + * dying.

[ ... ]


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/33864052510

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next v7 2/2] selftests/bpf: Add memcg async reclaim test
  2026-09-04 10:20 ` [PATCH bpf-next v7 2/2] selftests/bpf: Add memcg async reclaim test Hui Zhu
@ 2026-09-04 11:25   ` bot+bpf-ci
  0 siblings, 0 replies; 7+ messages in thread
From: bot+bpf-ci @ 2026-09-04 11:25 UTC (permalink / raw)
  To: hui.zhu, roman.gushchin, inwardvessel, shakeel.butt, akpm, andrii,
	eddyz87, ihor.solodrai, ast, daniel, memxor, martin.lau, song,
	yonghong.song, jolsa, emil, shuah, baohua, geliang, linux-kernel,
	bpf, linux-mm, linux-kselftest
  Cc: zhuhui, ast, andrii, daniel, martin.lau, eddyz87, yonghong.song,
	mason, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 12460 bytes --]

> diff --git a/tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c b/tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c
> new file mode 100644
> index 0000000000000..65f5006844636
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c

[ ... ]

> +static int write_file(const char *filename)
> +{
> +	int ret = -1;
> +	size_t written = 0;
> +	char *buffer;
> +	FILE *fp;
> +
> +	fp = fopen(filename, "wb");
> +	if (!fp)
> +		goto out;
> +
> +	buffer = malloc(BUFFER_SIZE);
> +	if (!buffer)
> +		goto cleanup_fp;
> +
> +	memset(buffer, 'A', BUFFER_SIZE);
> +
> +	while (written < FILE_SIZE) {
> +		size_t to_write = FILE_SIZE - written < BUFFER_SIZE ?
> +				  FILE_SIZE - written : BUFFER_SIZE;
> +
> +		if (fwrite(buffer, 1, to_write, fp) != to_write)
> +			goto cleanup;
> +		written += to_write;
> +	}
> +
> +	ret = 0;
> +cleanup:
> +	free(buffer);
> +cleanup_fp:
> +	fclose(fp);
> +out:
> +	return ret;
> +}

Does write_file() report buffered flush failures? It sets ret = 0
before fclose() and then discards fclose()'s return value, so if the
final flush fails (for example, ENOSPC when writing the 32 MiB file
into /tmp), the function still returns 0.

The truncated file is caught later by read_file()'s st_size check, but
the child exits with CHILD_EXIT_READ_FILE instead of
CHILD_EXIT_WRITE_FILE, so child_exit_str() reports "read data file"
for what was really a write failure. The per-step exit codes added by
this patch exist precisely to pin down which step failed.

[ ... ]

> +/*
> + * Remove the reclaim target while the BPF program keeps running and
> + * verify that reclaim stops on the dying/removed cgroup instead of
> + * reclaiming from it.
> + *
> + * The target stays empty; the workload lives in the trigger cgroup and
> + * only keeps refaults flowing so that reclaim rounds keep starting,
> + * both before and after the target is removed. reclaim_calls growing
> + * while the target is alive proves that rounds really run (the kfunc
> + * returns 0 on the empty target, but the call is still counted), and
> + * after the removal the skip counters must grow while reclaim_calls
> + * and reclaimed_bytes stay frozen.
> + */
> +void test_memcg_async_reclaim_dying(void)
> +{
> +	u64 trigger_cgroup_id, target_cgroup_id;
> +	u64 calls_before, bytes_before;
> +	char data_file[PATH_MAX] = "";
> +	struct memcg_async_reclaim *skel = NULL;
> +	pid_t reader_pid = -1;
> +	int err, fd, i;
> +
> +	err = setup_dying_cgroups(&trigger_cgroup_id, &target_cgroup_id);
> +	if (!ASSERT_OK(err, "setup_dying_cgroups"))
> +		return;
> +
> +	err = setup_bpf(trigger_cgroup_id, target_cgroup_id, &skel);
> +	if (!ASSERT_OK(err, "setup_bpf"))
> +		goto out;

[ ... ]

> +	/* Wait for reclaim rounds to reach the live target cgroup. */
> +	for (i = 0; i < DYING_POLL_ITERS; i++) {
> +		if (skel->bss->reclaim_calls > 0)
> +			break;
> +		usleep(DYING_POLL_INTERVAL_US);
> +	}
> +	if (!ASSERT_GT(skel->bss->reclaim_calls, 0, "reclaim_calls"))
> +		goto out;
> +
> +	remove_cgroup(CG_DYING_TARGET_DIR);
> +
> +	/* Let reclaim passes that were already in flight drain. */
> +	usleep(DYING_SETTLE_US);
> +
> +	calls_before = skel->bss->reclaim_calls;
> +	bytes_before = skel->bss->reclaimed_bytes;
> +
> +	/* Wait for reclaim rounds to hit the removed cgroup. */
> +	for (i = 0; i < DYING_POLL_ITERS; i++) {
> +		if (skel->bss->reclaim_target_gone ||
> +		    skel->bss->reclaim_skipped_dying)
> +			break;
> +		usleep(DYING_POLL_INTERVAL_US);
> +	}
> +
> +	if (!skel->bss->reclaim_target_gone &&
> +	    !skel->bss->reclaim_skipped_dying) {
> +		PRINT_FAIL("no reclaim round hit the removed cgroup (gone=%llu, dying=%llu)",
> +			   (unsigned long long)skel->bss->reclaim_target_gone,
> +			   (unsigned long long)skel->bss->reclaim_skipped_dying);
> +		goto out;
> +	}
> +
> +	/*
> +	 * reclaim_skipped_dying shows that the CSS_DYING/CSS_ONLINE check
> +	 * caught the cgroup mid-teardown. Whether it is hit is timing
> +	 * dependent, because the cgroup may already be fully released, so
> +	 * only the combined skip count above is asserted.
> +	 */
> +	printf("memcg_async_reclaim_dying: skips on removed cgroup: gone=%llu, dying=%llu\n",
> +	       (unsigned long long)skel->bss->reclaim_target_gone,
> +	       (unsigned long long)skel->bss->reclaim_skipped_dying);
> +
> +	/* Nothing may have been reclaimed from the removed target. */
> +	if (!ASSERT_EQ(skel->bss->reclaim_calls, calls_before, "reclaim_calls"))
> +		goto out;
> +	if (!ASSERT_EQ(skel->bss->reclaimed_bytes, bytes_before,
> +		       "reclaimed_bytes"))
> +		goto out;

Can this test distinguish a BPF program that implements the
CSS_DYING/CSS_ONLINE check from one that does not?

The commit message says "Looking the target up by id is not enough:
bpf_cgroup_from_id() keeps handing back a cgroup until its last
reference is dropped, so the program checks the css flags and skips
reclaim once the target is offlined or dying", and the comment above
this test says it "verifies that reclaim stops on the removed target
instead of reclaiming from it".

But bpf_cgroup_from_id() does not keep handing back the cgroup after
rmdir. The lookup sequence is:

  bpf_cgroup_from_id()
    -> cgroup_get_from_id()
      -> __cgroup_get_from_id() in kernel/cgroup/cgroup.c
        -> kernfs_find_and_get_node_by_id()

kernfs_find_and_get_node_by_id() in fs/kernfs/dir.c fails on an
inactive node:

    if (unlikely(!__kernfs_active(kn) || !atomic_inc_not_zero(&kn->count)))
        goto err_unlock;

cgroup_destroy_locked() runs synchronously inside the rmdir syscall and
calls kernfs_remove(cgrp->kn) before it returns. By the time
remove_cgroup(CG_DYING_TARGET_DIR) at line 635 returns, every subsequent
reclaim round takes the get_cgroup_memcg_from_id() failure path in
progs/memcg_async_reclaim.c:132, bumping reclaim_target_gone and leaving
reclaim_calls / reclaimed_bytes frozen.

That happens whether or not cgroup_is_dying() exists in the program.
reclaim_skipped_dying is only reachable through a sub-microsecond race
in reclaim_cgroup() where the reference is acquired before rmdir but the
flags are read after, and the deliberate usleep(DYING_SETTLE_US) = 200
ms at line 638, placed before the calls_before/bytes_before snapshot,
makes even that window essentially unobservable.

So the break condition at lines 645-647 and all three assertions (lines
651-657, 670, 672) are satisfied by the lookup-failure path alone.
Deleting cgroup_is_dying() from the BPF program would not make this test
fail.

To actually exercise the dying window the test would have to keep the
target cgroup pinned so that bpf_cgroup_from_id() still resolves it
while it is offline (for example, hold a cgroup fd or a task in it, or
drive the removal concurrently with reclaim without the 200 ms settle)
and then assert on reclaim_skipped_dying specifically rather than on the
OR of the two counters.

> diff --git a/tools/testing/selftests/bpf/progs/memcg_async_reclaim.c b/tools/testing/selftests/bpf/progs/memcg_async_reclaim.c
> new file mode 100644
> index 0000000000000..e6839ade472bb
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/memcg_async_reclaim.c

[ ... ]

> +/*
> + * One reclaim round targets RECLAIM_MAX_ITER batches of RECLAIM_SIZE
> + * each. Each bpf_wq callback reclaims a single batch and requeues the
> + * same work item for the next one, so no callback runs longer than one
> + * bounded reclaim pass.
> + */
> +#define RECLAIM_SIZE		(32 * PAGE_SIZE)
> +#define RECLAIM_MAX_ITER	32

[ ... ]

> +/*
> + * A cgroup is dying once it has been offlined (CSS_ONLINE cleared) or
> + * CSS_DYING has been raised, mirroring cgroup_is_dead()/css_is_dying()
> + * in include/linux/cgroup.h. bpf_cgroup_from_id() can still hand back
> + * such a cgroup, because it only fails once the last reference has been
> + * dropped, so reclaim has to check these flags instead of relying on
> + * the lookup failing.
> + *
> + * CSS_ONLINE and CSS_DYING come from vmlinux.h: the kernel defines them
> + * in an anonymous enum, so bpf_core_enum_value() has no enum type to
> + * bind to, and redeclaring them locally would clash with the vmlinux.h
> + * enumerators. vmlinux.h is generated from the running kernel's BTF, so
> + * the values already match the target kernel.
> + */
> +static bool cgroup_is_dying(struct cgroup *cgrp)
> +{
> +	unsigned int flags = cgrp->self.flags;
> +
> +	return (flags & CSS_DYING) || !(flags & CSS_ONLINE);
> +}

Does the (flags & CSS_DYING) term ever fire?

CSS_DYING is never set on cgrp->self. The only assignment in the tree is
css->flags |= CSS_DYING in kill_css_sync() at kernel/cgroup/cgroup.c:6147.
kill_css_sync() has exactly two callers, and both pass a subsystem css,
never the cgroup's own self css:

  - cgroup_destroy_locked() calls it under for_each_css(css, ssid, cgrp),
    and that macro only ever yields cgrp->subsys[ssid].
  - cgroup_apply_control_disable() calls it with css = cgroup_css(dsct, ss)
    where ss comes from for_each_subsys(), so again dsct->subsys[ss->id].

kill_css_sync() also unconditionally dereferences css->ss->css_killed,
which would NULL-deref for &cgrp->self whose ss is NULL. So cgrp->self.flags
& CSS_DYING is always 0 and cgroup_is_dying() reduces to cgroup_is_dead().

Also, the comment above the helper claims bpf_cgroup_from_id() keeps
handing back the cgroup until the last reference is dropped. That is not
what the lookup does. bpf_cgroup_from_id() -> cgroup_get_from_id() ->
__cgroup_get_from_id() is kernfs_find_and_get_node_by_id() followed by
cgroup_tryget(). It starts failing as soon as cgroup_destroy_locked()
reaches kernfs_remove(cgrp->kn) or percpu_ref_kill(&cgrp->self.refcnt),
both long before the last reference is dropped. The window in which the
surviving !CSS_ONLINE check can still observe a cgroup is between
cgrp->self.flags &= ~CSS_ONLINE and kernfs_remove() a few statements
later, all under cgroup_mutex.

The comment at line 112 also states "vmlinux.h is generated from the
running kernel's BTF". tools/testing/selftests/bpf/vmlinux.h is generated
at build time from $(VMLINUX_BTF), which for a cross-built or packaged
selftest is not necessarily the kernel the test runs on.

[ ... ]

> +static int reclaim_work_fn(void *map, int *key, void *value)
> +{
> +	struct wq_elem *elem = value;
> +	u64 nr, size;
> +
> +	if (!elem->remaining) {
> +		/*
> +		 * Timer-triggered entry: start a new round only when the
> +		 * high cgroup refaults enough. Requeued entries skip this
> +		 * check and only look at remaining, so the refault delta
> +		 * is consumed once per round.
> +		 */
> +		if (!should_reclaim_cgroup(wq_high_cgroup_id, &elem->prev_event,
> +			elem->event_delta_threshold))
> +			return 0;
> +		elem->remaining = RECLAIM_MAX_ITER * RECLAIM_SIZE;
> +	}
> +
> +	/* One bounded reclaim pass per callback */
> +	size = elem->remaining < RECLAIM_SIZE ? elem->remaining : RECLAIM_SIZE;
> +	nr = reclaim_cgroup(wq_low_cgroup_id, size);
> +	if (!nr) {
> +		elem->remaining = 0;
> +		return 0;
> +	}
> +
> +	/* try_to_free_mem_cgroup_pages() may reclaim more than requested */
> +	if (nr >= elem->remaining)
> +		elem->remaining = 0;
> +	else
> +		elem->remaining -= nr;
> +
> +	/* Requeue the same work item for the next batch */
> +	if (elem->remaining)
> +		bpf_wq_start(&elem->work, 0);
> +
> +	return 0;
> +}

Does RECLAIM_MAX_ITER actually bound an iteration count?

The macro is only ever used as a byte multiplier:

  elem->remaining = RECLAIM_MAX_ITER * RECLAIM_SIZE = 32 * 128 KiB = 4 MiB

and the requeue loop is bounded by bytes actually reclaimed, not by
callback count. bpf_proactive_reclaim() returns the bytes that
try_to_free_mem_cgroup_pages() actually freed, which is frequently less
than the requested 128 KiB. A single page (4 KiB) is enough to keep the
round alive, so a round in which each pass frees one page needs 4 MiB /
4 KiB = 1024 callbacks, i.e. 32x more than the name and the file header
comment suggest.

Total work per round is still bounded at 4 MiB and each individual
callback is still one bounded pass, so there is no runaway. The defect
is that the documented iteration bound does not exist, which makes the
requeue budget hard to reason about.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/33864052510

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next v7 1/2] mm/bpf: Add bpf_proactive_reclaim kfunc
  2026-09-04 10:20 ` [PATCH bpf-next v7 1/2] mm/bpf: Add bpf_proactive_reclaim kfunc Hui Zhu
  2026-09-04 10:36   ` sashiko-bot
  2026-09-04 11:25   ` bot+bpf-ci
@ 2026-09-04 18:03   ` Andrew Morton
  2 siblings, 0 replies; 7+ messages in thread
From: Andrew Morton @ 2026-09-04 18:03 UTC (permalink / raw)
  To: Hui Zhu
  Cc: Roman Gushchin, JP Kobryn, Shakeel Butt, Andrii Nakryiko,
	Eduard Zingerman, Ihor Solodrai, Alexei Starovoitov,
	Daniel Borkmann, Kumar Kartikeya Dwivedi, Martin KaFai Lau,
	Song Liu, Yonghong Song, Jiri Olsa, Emil Tsalapatis, Shuah Khan,
	Barry Song, Geliang Tang, linux-kernel, bpf, linux-mm,
	linux-kselftest, Hui Zhu

On Fri,  4 Sep 2026 18:20:19 +0800 "Hui Zhu" <hui.zhu@linux.dev> wrote:

> @@ -159,6 +160,74 @@ __bpf_kfunc void bpf_mem_cgroup_flush_stats(struct mem_cgroup *memcg)
>  	mem_cgroup_flush_stats(memcg);
>  }
>  
> +/*
> + * Reclaim must not recurse: try_to_free_mem_cgroup_pages() overwrites
> + * current->reclaim_state, so a nested call would corrupt the outer
> + * reclaim state. Reclaim windows are marked with PF_MEMALLOC;
> + * reclaim_state is also checked because it is installed slightly
> + * before PF_MEMALLOC.
> + */
> +static bool bpf_in_reclaim_context(void)
> +{
> +	return (current->flags & PF_MEMALLOC) || current->reclaim_state;
> +}
> +

Would life improve if try_to_free_mem_cgroup_pages() didn't do that? 
If try_to_free_mem_cgroup_pages() (or some variant of it) were to
permit nesting?

	rs = new_thing(current, &sc.reclaim_state);
	...
	set_task_reclaim_state(current, rs);

?

^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-09-04 18:03 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-04 10:20 [PATCH bpf-next v7 0/2] bpf: BPF-driven proactive memcg reclaim Hui Zhu
2026-09-04 10:20 ` [PATCH bpf-next v7 1/2] mm/bpf: Add bpf_proactive_reclaim kfunc Hui Zhu
2026-09-04 10:36   ` sashiko-bot
2026-09-04 11:25   ` bot+bpf-ci
2026-09-04 18:03   ` Andrew Morton
2026-09-04 10:20 ` [PATCH bpf-next v7 2/2] selftests/bpf: Add memcg async reclaim test Hui Zhu
2026-09-04 11:25   ` bot+bpf-ci

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox