From: Johannes Weiner <hannes@cmpxchg.org>
To: Hao Jia <jiahao.kernel@gmail.com>
Cc: akpm@linux-foundation.org, chengming.zhou@linux.dev,
jiahao1@lixiang.com, linux-doc@vger.kernel.org,
linux-kernel@vger.kernel.org, linux-mm@kvack.org,
mhocko@kernel.org, mkoutny@suse.com, muchun.song@linux.dev,
nphamcs@gmail.com, roman.gushchin@linux.dev,
shakeel.butt@linux.dev, stable@vger.kernel.org, tj@kernel.org,
yosry@kernel.org
Subject: Re: [PATCH v3 2/2] mm/zswap: Support batch writeback in shrink_memcg()
Date: Fri, 31 Jul 2026 11:17:20 -0400 [thread overview]
Message-ID: <amy8gAN33W4QsM5R@cmpxchg.org> (raw)
In-Reply-To: <20260731071900.38942-1-jiahao.kernel@gmail.com>
On Fri, Jul 31, 2026 at 03:19:00PM +0800, Hao Jia wrote:
> From: Hao Jia <jiahao1@lixiang.com>
>
> Currently, shrink_memcg() writes back at most one entry per-node during
> its traversal. This makes shrink_worker() inefficient, as it must
> repeatedly re-enter shrink_memcg() to make any substantial progress.
> Under high memory pressure, this can cause the writeback speed to be
> too slow to keep up with refaults, leading to zswap store failures and
> forcing pages to skip zswap and go directly to disk, which results in
> an LRU inversion.
>
> To address this, extend shrink_memcg() and rewrite its LRU iteration logic,
> enabling batch writeback for both the shrink_worker() and zswap_store() paths.
> To prevent shrink unfairness across NUMA nodes caused by a shared global scan
> quota, limit scanning to up to SWAP_CLUSTER_MAX pages per node and write back
> any reclaimable entries found.
>
> Test Setup:
> - Total memory: 32 GB, 1 NUMA node.
> - zswap settings: accept_threshold_percent=50, shrinker_enabled=N.
>
> Test Case 1:
> Set max_pool_percent=1, allocate 512MB of anonymous pages, and fill them
> with random data (to avoid compression). Then, use cgroup memory.reclaim
> to force a large amount of anonymous pages into zswap. At an interval of
> 2ms, allocate a 4K anonymous page where the first 4 bytes are random numbers
> and the rest are zeros, and then trigger reclamation of this 4K page through
> cgroup memory.reclaim. When the pool threshold is reached, shrink_memcg()
> will be triggered.
> The test data after running for 120s is as follows:
> Baseline Patched
> shrink_worker wakeups 5,363 169
> shrink_memcg calls 11,373,201 350,703
> written_back pages 40,212 40,241
> zswap_store calls 161,190 163,753
> store succeeded (ret=1) 102,743 117,183
> store rejected (ret=0) 58,447 46,570
> store reject rate ~36% ~28%
> pool_limit_hit delta 55,826 33,760
> pswpout 98,659 86,811
> pswpin 2 0
>
> Test Case 2:
> We evaluated the following two sub-configurations using stress-ng inside
> a cgroup capped at memory.max=1G for 120 seconds:
> Test Case 2a (max_pool_percent=1): Continuously triggers the global
> zswap pool limit, thereby waking up shrink_worker() to perform asynchronous
> shrinking.
> Test Case 2b (zswap.max=320M, max_pool_percent=50): Continuously triggers
> the cgroup's zswap.max limit, thereby invoking synchronous shrinking.
> Command executed for both setups:
> bash -c 'echo $$ > /sys/fs/cgroup/zswaptest/cgroup.procs ; \
> exec stress-ng --vm 4 --vm-bytes 4G --vm-keep --vm-method rand-set -t \
> 120s -q'
>
> Test Case 2a (max_pool_percent=1):
> Baseline Patched
> shrink_worker wakeups 5,640 1,308
> shrink_memcg calls 8,481,500 3,140,972
> written_back pages 260 468,216
> zswap_store calls 2,742,756 2,011,269
> store succeeded (ret=1) 934,640 947,988
> store rejected (ret=0) 1,808,116 1,063,281
> store reject rate ~66% ~52%
> pool_limit_hit delta 1,181,310 196,882
> pswpout 1,808,376 1,531,497
> pswpin 4,288,497 3,635,365
> Test Case 2b (zswap.max=320M, max_pool_percent=50):
> Baseline Patched
> shrink_worker wakeups 0 0
> shrink_memcg calls 687,608 54,002
> written_back pages 639,176 846,663
> zswap_store calls 1,224,222 1,228,548
> store succeeded (ret=1) 992,816 1,208,123
> store rejected (ret=0) 231,431 20,425
> store reject rate ~19% ~2%
> pool_limit_hit delta 0 0
> pswpout 870,745 867,360
> pswpin 1,707,823 1,216,814
>
> Under identical workloads and runtimes, batched zswap shrinking
> exhibits a significant reduction in both shrink_worker() wakeups
> and shrink_memcg() calls. Furthermore, the sharp drop in both pswpin
> and zswap_store() rejections demonstrates that batching zswap shrink
> operations effectively mitigates zswap_store() failures caused by
> hitting the pool limit. This significantly prevents pages from bypassing
> zswap and falling back directly to disk, thereby reducing LRU inversion.
>
> Suggested-by: Yosry Ahmed <yosry@kernel.org>
> Acked-by: Yosry Ahmed <yosry@kernel.org>
> Acked-by: Nhat Pham <nphamcs@gmail.com>
> Signed-off-by: Hao Jia <jiahao1@lixiang.com>
> ---
> mm/zswap.c | 30 ++++++++++++++++++++++++++++--
> 1 file changed, 28 insertions(+), 2 deletions(-)
>
> diff --git a/mm/zswap.c b/mm/zswap.c
> index 48fc7b575e24..d406c14925d8 100644
> --- a/mm/zswap.c
> +++ b/mm/zswap.c
> @@ -1275,6 +1275,21 @@ static struct shrinker *zswap_alloc_shrinker(void)
> return shrinker;
> }
>
> +/*
> + * Scan up to SWAP_CLUSTER_MAX pages on each per-node zswap LRU of @memcg
> + * and write back the reclaimable ones.
> + *
> + * Since the second-chance algorithm rotates referenced entries to the
> + * LRU tail, the per-node scan is capped at the current LRU length so
> + * each entry is scanned at most once per call. It is up to the caller
> + * to handle retries, deciding whether to scan another memcg to complete
> + * the full iteration, or to rescan the current memcg to drain its zswap
> + * entries.
> + *
> + * Return: 0 if at least one entry was written back, -EAGAIN if entries
> + * were scanned but none could be written back, or -ENOENT if @memcg has
> + * writeback disabled, is a zombie cgroup, or has empty zswap LRUs.
> + */
> static int shrink_memcg(struct mem_cgroup *memcg)
> {
> int nid, shrunk = 0, scanned = 0;
> @@ -1290,13 +1305,24 @@ static int shrink_memcg(struct mem_cgroup *memcg)
> return -ENOENT;
>
> for_each_node_state(nid, N_NORMAL_MEMORY) {
> - unsigned long nr_to_walk = 1;
> + unsigned long nr_to_walk, node_budget;
> +
> + /*
> + * Cap the scan at the per-node LRU length so each entry is
> + * scanned at most once per call.
> + */
> + node_budget = min(SWAP_CLUSTER_MAX,
> + list_lru_count_one(&zswap_list_lru, nid, memcg));
AFAICS you can just do unsigned long nr_to_walk = SWAP_CLUSTER_MAX.
__list_lru_walk_one() does a list_for_each_safe() that will exit the
same way whether you hit !nr_to_walk or run out of items.
> + if (!node_budget)
> + continue;
>
> + nr_to_walk = node_budget;
> shrunk += list_lru_walk_one(&zswap_list_lru, nid, memcg,
> &shrink_memcg_cb, NULL, &nr_to_walk);
> - scanned += 1 - nr_to_walk;
> + scanned += node_budget - nr_to_walk;
scanned += SWAP_CLUSTER_MAX - nr_to_walk;
next prev parent reply other threads:[~2026-07-31 15:17 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-29 8:42 [PATCH v3 0/2] mm/zswap: Fixes and improves the zswap shrink Hao Jia
2026-07-29 8:42 ` [PATCH v3 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
2026-07-29 22:58 ` Andrew Morton
2026-07-30 0:30 ` Yosry Ahmed
2026-07-30 1:17 ` Andrew Morton
2026-07-30 16:52 ` Yosry Ahmed
2026-07-30 17:59 ` Andrew Morton
2026-07-30 18:02 ` Yosry Ahmed
2026-07-30 6:31 ` Hao Jia
2026-07-30 17:48 ` Yosry Ahmed
2026-07-30 19:01 ` Johannes Weiner
2026-07-31 7:19 ` [PATCH v3 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
2026-07-31 15:17 ` Johannes Weiner [this message]
2026-07-31 7:25 ` [PATCH v3 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
2026-07-29 8:42 ` [PATCH v3 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=amy8gAN33W4QsM5R@cmpxchg.org \
--to=hannes@cmpxchg.org \
--cc=akpm@linux-foundation.org \
--cc=chengming.zhou@linux.dev \
--cc=jiahao.kernel@gmail.com \
--cc=jiahao1@lixiang.com \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mhocko@kernel.org \
--cc=mkoutny@suse.com \
--cc=muchun.song@linux.dev \
--cc=nphamcs@gmail.com \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=stable@vger.kernel.org \
--cc=tj@kernel.org \
--cc=yosry@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.