From: Hao Jia <jiahao.kernel@gmail.com>
To: Johannes Weiner <hannes@cmpxchg.org>
Cc: akpm@linux-foundation.org, chengming.zhou@linux.dev,
jiahao1@lixiang.com, linux-doc@vger.kernel.org,
linux-kernel@vger.kernel.org, linux-mm@kvack.org,
mhocko@kernel.org, mkoutny@suse.com, muchun.song@linux.dev,
nphamcs@gmail.com, roman.gushchin@linux.dev,
shakeel.butt@linux.dev, stable@vger.kernel.org, tj@kernel.org,
yosry@kernel.org
Subject: Re: [PATCH v3 2/2] mm/zswap: Support batch writeback in shrink_memcg()
Date: Sat, 1 Aug 2026 08:31:09 +0800 [thread overview]
Message-ID: <6e16a760-d84d-98e1-433e-3233a9dd2236@gmail.com> (raw)
In-Reply-To: <amy8gAN33W4QsM5R@cmpxchg.org>
On 2026/7/31 23:17, Johannes Weiner wrote:
> On Fri, Jul 31, 2026 at 03:19:00PM +0800, Hao Jia wrote:
>> From: Hao Jia <jiahao1@lixiang.com>
>>
>> Currently, shrink_memcg() writes back at most one entry per-node during
>> its traversal. This makes shrink_worker() inefficient, as it must
>> repeatedly re-enter shrink_memcg() to make any substantial progress.
>> Under high memory pressure, this can cause the writeback speed to be
>> too slow to keep up with refaults, leading to zswap store failures and
>> forcing pages to skip zswap and go directly to disk, which results in
>> an LRU inversion.
>>
>> To address this, extend shrink_memcg() and rewrite its LRU iteration logic,
>> enabling batch writeback for both the shrink_worker() and zswap_store() paths.
>> To prevent shrink unfairness across NUMA nodes caused by a shared global scan
>> quota, limit scanning to up to SWAP_CLUSTER_MAX pages per node and write back
>> any reclaimable entries found.
>>
>> Test Setup:
>> - Total memory: 32 GB, 1 NUMA node.
>> - zswap settings: accept_threshold_percent=50, shrinker_enabled=N.
>>
>> Test Case 1:
>> Set max_pool_percent=1, allocate 512MB of anonymous pages, and fill them
>> with random data (to avoid compression). Then, use cgroup memory.reclaim
>> to force a large amount of anonymous pages into zswap. At an interval of
>> 2ms, allocate a 4K anonymous page where the first 4 bytes are random numbers
>> and the rest are zeros, and then trigger reclamation of this 4K page through
>> cgroup memory.reclaim. When the pool threshold is reached, shrink_memcg()
>> will be triggered.
>> The test data after running for 120s is as follows:
>> Baseline Patched
>> shrink_worker wakeups 5,363 169
>> shrink_memcg calls 11,373,201 350,703
>> written_back pages 40,212 40,241
>> zswap_store calls 161,190 163,753
>> store succeeded (ret=1) 102,743 117,183
>> store rejected (ret=0) 58,447 46,570
>> store reject rate ~36% ~28%
>> pool_limit_hit delta 55,826 33,760
>> pswpout 98,659 86,811
>> pswpin 2 0
>>
>> Test Case 2:
>> We evaluated the following two sub-configurations using stress-ng inside
>> a cgroup capped at memory.max=1G for 120 seconds:
>> Test Case 2a (max_pool_percent=1): Continuously triggers the global
>> zswap pool limit, thereby waking up shrink_worker() to perform asynchronous
>> shrinking.
>> Test Case 2b (zswap.max=320M, max_pool_percent=50): Continuously triggers
>> the cgroup's zswap.max limit, thereby invoking synchronous shrinking.
>> Command executed for both setups:
>> bash -c 'echo $$ > /sys/fs/cgroup/zswaptest/cgroup.procs ; \
>> exec stress-ng --vm 4 --vm-bytes 4G --vm-keep --vm-method rand-set -t \
>> 120s -q'
>>
>> Test Case 2a (max_pool_percent=1):
>> Baseline Patched
>> shrink_worker wakeups 5,640 1,308
>> shrink_memcg calls 8,481,500 3,140,972
>> written_back pages 260 468,216
>> zswap_store calls 2,742,756 2,011,269
>> store succeeded (ret=1) 934,640 947,988
>> store rejected (ret=0) 1,808,116 1,063,281
>> store reject rate ~66% ~52%
>> pool_limit_hit delta 1,181,310 196,882
>> pswpout 1,808,376 1,531,497
>> pswpin 4,288,497 3,635,365
>> Test Case 2b (zswap.max=320M, max_pool_percent=50):
>> Baseline Patched
>> shrink_worker wakeups 0 0
>> shrink_memcg calls 687,608 54,002
>> written_back pages 639,176 846,663
>> zswap_store calls 1,224,222 1,228,548
>> store succeeded (ret=1) 992,816 1,208,123
>> store rejected (ret=0) 231,431 20,425
>> store reject rate ~19% ~2%
>> pool_limit_hit delta 0 0
>> pswpout 870,745 867,360
>> pswpin 1,707,823 1,216,814
>>
>> Under identical workloads and runtimes, batched zswap shrinking
>> exhibits a significant reduction in both shrink_worker() wakeups
>> and shrink_memcg() calls. Furthermore, the sharp drop in both pswpin
>> and zswap_store() rejections demonstrates that batching zswap shrink
>> operations effectively mitigates zswap_store() failures caused by
>> hitting the pool limit. This significantly prevents pages from bypassing
>> zswap and falling back directly to disk, thereby reducing LRU inversion.
>>
>> Suggested-by: Yosry Ahmed <yosry@kernel.org>
>> Acked-by: Yosry Ahmed <yosry@kernel.org>
>> Acked-by: Nhat Pham <nphamcs@gmail.com>
>> Signed-off-by: Hao Jia <jiahao1@lixiang.com>
>> ---
>> mm/zswap.c | 30 ++++++++++++++++++++++++++++--
>> 1 file changed, 28 insertions(+), 2 deletions(-)
>>
>> diff --git a/mm/zswap.c b/mm/zswap.c
>> index 48fc7b575e24..d406c14925d8 100644
>> --- a/mm/zswap.c
>> +++ b/mm/zswap.c
>> @@ -1275,6 +1275,21 @@ static struct shrinker *zswap_alloc_shrinker(void)
>> return shrinker;
>> }
>>
>> +/*
>> + * Scan up to SWAP_CLUSTER_MAX pages on each per-node zswap LRU of @memcg
>> + * and write back the reclaimable ones.
>> + *
>> + * Since the second-chance algorithm rotates referenced entries to the
>> + * LRU tail, the per-node scan is capped at the current LRU length so
>> + * each entry is scanned at most once per call. It is up to the caller
>> + * to handle retries, deciding whether to scan another memcg to complete
>> + * the full iteration, or to rescan the current memcg to drain its zswap
>> + * entries.
>> + *
>> + * Return: 0 if at least one entry was written back, -EAGAIN if entries
>> + * were scanned but none could be written back, or -ENOENT if @memcg has
>> + * writeback disabled, is a zombie cgroup, or has empty zswap LRUs.
>> + */
>> static int shrink_memcg(struct mem_cgroup *memcg)
>> {
>> int nid, shrunk = 0, scanned = 0;
>> @@ -1290,13 +1305,24 @@ static int shrink_memcg(struct mem_cgroup *memcg)
>> return -ENOENT;
>>
>> for_each_node_state(nid, N_NORMAL_MEMORY) {
>> - unsigned long nr_to_walk = 1;
>> + unsigned long nr_to_walk, node_budget;
>> +
>> + /*
>> + * Cap the scan at the per-node LRU length so each entry is
>> + * scanned at most once per call.
>> + */
>> + node_budget = min(SWAP_CLUSTER_MAX,
>> + list_lru_count_one(&zswap_list_lru, nid, memcg));
>
> AFAICS you can just do unsigned long nr_to_walk = SWAP_CLUSTER_MAX.
>
> __list_lru_walk_one() does a list_for_each_safe() that will exit the
> same way whether you hit !nr_to_walk or run out of items.
Wouldn't it be better to ensure that each entry is scanned at most once
per call, particularly when list_lru_count_one(nid) < SWAP_CLUSTER_MAX?
On one hand, this avoids scanning the same entry multiple times within a
single pass. Since the second-chance algorithm rotates referenced
entries to the tail of the LRU, entries on nodes with a large number of
zswap entries require at least two shrink_memcg() calls to be written
back, whereas entries on nodes with fewer entries might get written back
in a single shrink_memcg() call instead. On the other hand, we also
avoid spinning repeatedly on entries that fail writeback.
Thanks,
Hao
>
>> + if (!node_budget)
>> + continue;
>>
>> + nr_to_walk = node_budget;
>> shrunk += list_lru_walk_one(&zswap_list_lru, nid, memcg,
>> &shrink_memcg_cb, NULL, &nr_to_walk);
>> - scanned += 1 - nr_to_walk;
>> + scanned += node_budget - nr_to_walk;
>
> scanned += SWAP_CLUSTER_MAX - nr_to_walk;
next prev parent reply other threads:[~2026-08-01 0:31 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-29 8:42 [PATCH v3 0/2] mm/zswap: Fixes and improves the zswap shrink Hao Jia
2026-07-29 8:42 ` [PATCH v3 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
2026-07-29 22:58 ` Andrew Morton
2026-07-30 0:30 ` Yosry Ahmed
2026-07-30 1:17 ` Andrew Morton
2026-07-30 16:52 ` Yosry Ahmed
2026-07-30 17:59 ` Andrew Morton
2026-07-30 18:02 ` Yosry Ahmed
2026-07-30 6:31 ` Hao Jia
2026-07-30 17:48 ` Yosry Ahmed
2026-07-30 19:01 ` Johannes Weiner
2026-07-31 7:19 ` [PATCH v3 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
2026-07-31 15:17 ` Johannes Weiner
2026-08-01 0:31 ` Hao Jia [this message]
2026-07-31 7:25 ` [PATCH v3 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
2026-07-29 8:42 ` [PATCH v3 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=6e16a760-d84d-98e1-433e-3233a9dd2236@gmail.com \
--to=jiahao.kernel@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=chengming.zhou@linux.dev \
--cc=hannes@cmpxchg.org \
--cc=jiahao1@lixiang.com \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mhocko@kernel.org \
--cc=mkoutny@suse.com \
--cc=muchun.song@linux.dev \
--cc=nphamcs@gmail.com \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=stable@vger.kernel.org \
--cc=tj@kernel.org \
--cc=yosry@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox