The Linux Kernel Mailing List
 help / color / mirror / Atom feed
* [PATCH v4 0/2] mm/zswap: Fixes and improves the zswap shrink
@ 2026-08-06  7:09 Hao Jia
  2026-08-06  7:09 ` [PATCH v4 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
                   ` (2 more replies)
  0 siblings, 3 replies; 5+ messages in thread
From: Hao Jia @ 2026-08-06  7:09 UTC (permalink / raw)
  To: akpm, tj, hannes, shakeel.butt, mhocko, yosry, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin
  Cc: linux-mm, linux-kernel, linux-doc, Hao Jia

From: Hao Jia <jiahao1@lixiang.com>

This series fixes and improves the zswap global shrinker (shrink_worker()):
Patch 1: Fix missing global shrinker when memory cgroup is disabled.
Patch 2: Extend shrink_memcg() to support batch writeback and thereby improving
         the writeback efficiency in the shrink_worker() and zswap_store() paths.

v3->v4:
    - Remove the nr_to_scan parameter from shrink_memcg() and cap the per-node
      scan count at SWAP_CLUSTER_MAX to avoid writeback unfairness across NUMA
      nodes caused by a shared global scan quota.
v2->v3:
    - Added user impact to the commit 1 message.
    - Updated writeback batch size to SWAP_CLUSTER_MAX to avoid introducing new
      magic macros. And enabled batched shrinking in the zswap_store() path as well.
    - Re-benchmarked performance data with the updated patch.
v1->v2:
    - Add a reschedule check to the -ENOENT return path in shrink_memcg() to
      handle the theoretical issue of prolonged heavy concurrent zswap stores.
    - Remove the shrink_memcg() return value changes part, and include a more
       detailed test report in the commit message.

[v3] https://lore.kernel.org/all/20260729084206.77793-1-jiahao.kernel@gmail.com
[v2] https://lore.kernel.org/all/20260717085151.22822-1-jiahao.kernel@gmail.com
[v1] https://lore.kernel.org/all/20260714081510.16895-1-jiahao.kernel@gmail.com

Hao Jia (2):
  mm/zswap: Fix global shrinker when memory cgroup is disabled
  mm/zswap: Support batch writeback in shrink_memcg()

 mm/zswap.c | 26 ++++++++++++++++++--------
 1 file changed, 18 insertions(+), 8 deletions(-)

-- 
2.34.1


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v4 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled
  2026-08-06  7:09 [PATCH v4 0/2] mm/zswap: Fixes and improves the zswap shrink Hao Jia
@ 2026-08-06  7:09 ` Hao Jia
  2026-08-06  7:09 ` [PATCH v4 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
  2026-08-06 22:41 ` [PATCH v4 0/2] mm/zswap: Fixes and improves the zswap shrink Andrew Morton
  2 siblings, 0 replies; 5+ messages in thread
From: Hao Jia @ 2026-08-06  7:09 UTC (permalink / raw)
  To: akpm, tj, hannes, shakeel.butt, mhocko, yosry, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin
  Cc: linux-mm, linux-kernel, linux-doc, Hao Jia, stable

From: Hao Jia <jiahao1@lixiang.com>

Zswap writeback when the global pool limit is hit fails when memory
cgroup is disabled. The pool remains full until it is organically
drained by swapins or memory freeing, leading to zswap store failures
and pages bypassing getting written directly to the backing swap device,
causing LRU inversion (hotter pages with higher fault latency).

This happens because mem_cgroup_iter() always returns NULL when
memory cgroups are disabled. As a result, the global shrinker
shrink_worker() repeatedly takes empty walks. After MAX_RECLAIM_RETRIES
failed attempts, the worker gives up without writing back any pages.

Therefore, when memory cgroup is disabled, fall through with the !memcg
branch and shrink the root memcg directly.

With memcg disabled, shrink_memcg() only returns -ENOENT when the root
LRU is empty, which means the total pages are already below thr. In the
absence of heavy concurrent zswap stores, the loop then safely bails out
via the zswap_total_pages() <= thr check; otherwise, it will resume
shrinking the memcg after processing the reschedule check. For any other
return value from shrink_memcg(), the loop is guaranteed to terminate,
either after MAX_RECLAIM_RETRIES failures or once the threshold is met.

Fixes: a65b0e7607cc ("zswap: make shrinking memcg-aware")
Cc: stable@vger.kernel.org
Suggested-by: Nhat Pham <nphamcs@gmail.com>
Acked-by: Nhat Pham <nphamcs@gmail.com>
Acked-by: Yosry Ahmed <yosry@kernel.org>
Reported-by: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Hao Jia <jiahao1@lixiang.com>
---
 mm/zswap.c | 13 +++++++------
 1 file changed, 7 insertions(+), 6 deletions(-)

diff --git a/mm/zswap.c b/mm/zswap.c
index b5a17ea20237..48fc7b575e24 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1356,11 +1356,12 @@ static void shrink_worker(struct work_struct *w)
 		} while (memcg && !mem_cgroup_tryget_online(memcg));
 		spin_unlock(&zswap_shrink_lock);
 
-		if (!memcg) {
-			/*
-			 * Continue shrinking without incrementing failures if
-			 * we found candidate memcgs in the last tree walk.
-			 */
+		/*
+		 * A NULL memcg ends a full hierarchy pass (except when memcg is
+		 * disabled, where it is always NULL: fall through to the root LRU).
+		 * Count a failure only if the last pass found no candidates.
+		 */
+		if (!memcg && !mem_cgroup_disabled()) {
 			if (!attempts && ++failures == MAX_RECLAIM_RETRIES)
 				break;
 
@@ -1379,7 +1380,7 @@ static void shrink_worker(struct work_struct *w)
 		 * and failures.
 		 */
 		if (ret == -ENOENT)
-			continue;
+			goto resched;
 		++attempts;
 
 		if (ret && ++failures == MAX_RECLAIM_RETRIES)
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v4 2/2] mm/zswap: Support batch writeback in shrink_memcg()
  2026-08-06  7:09 [PATCH v4 0/2] mm/zswap: Fixes and improves the zswap shrink Hao Jia
  2026-08-06  7:09 ` [PATCH v4 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
@ 2026-08-06  7:09 ` Hao Jia
  2026-08-06 22:41 ` [PATCH v4 0/2] mm/zswap: Fixes and improves the zswap shrink Andrew Morton
  2 siblings, 0 replies; 5+ messages in thread
From: Hao Jia @ 2026-08-06  7:09 UTC (permalink / raw)
  To: akpm, tj, hannes, shakeel.butt, mhocko, yosry, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin
  Cc: linux-mm, linux-kernel, linux-doc, Hao Jia

From: Hao Jia <jiahao1@lixiang.com>

Currently, shrink_memcg() writes back at most one entry per-node during
its traversal. This makes shrink_worker() inefficient, as it must
repeatedly re-enter shrink_memcg() to make any substantial progress.
Under high memory pressure, this can cause the writeback speed to be
too slow to keep up with refaults, leading to zswap store failures and
forcing pages to skip zswap and go directly to disk, which results in
an LRU inversion.

To address this, extend the per-node scan budget in shrink_memcg() from a
single entry to up to SWAP_CLUSTER_MAX pages, enabling batch writeback for
both the shrink_worker() and zswap_store() paths.

Test Setup:
- Total memory: 32 GB, 1 NUMA node.
- zswap settings: accept_threshold_percent=50, shrinker_enabled=N.

Test Case 1:
Set max_pool_percent=1, allocate 512MB of anonymous pages, and fill them
with random data (to avoid compression). Then, use cgroup memory.reclaim
to force a large amount of anonymous pages into zswap. At an interval of
2ms, allocate a 4K anonymous page where the first 4 bytes are random numbers
and the rest are zeros, and then trigger reclamation of this 4K page through
cgroup memory.reclaim. When the pool threshold is reached, shrink_memcg()
will be triggered.
The test data after running for 120s is as follows:
                                Baseline      Patched
shrink_worker wakeups              5,363          169
shrink_memcg calls            11,373,201      350,703
written_back pages                40,212       40,241
zswap_store calls                161,190      163,753
   store succeeded (ret=1)       102,743      117,183
   store rejected (ret=0)         58,447       46,570
   store reject rate                ~36%        ~28%
pool_limit_hit delta              55,826       33,760
pswpout                           98,659       86,811
pswpin                                 2            0

Test Case 2:
We evaluated the following two sub-configurations using stress-ng inside
a cgroup capped at memory.max=1G for 120 seconds:
  Test Case 2a (max_pool_percent=1): Continuously triggers the global
  zswap pool limit, thereby waking up shrink_worker() to perform asynchronous
  shrinking.
  Test Case 2b (zswap.max=320M, max_pool_percent=50): Continuously triggers
  the cgroup's zswap.max limit, thereby invoking synchronous shrinking.
Command executed for both setups:
   bash -c 'echo $$ > /sys/fs/cgroup/zswaptest/cgroup.procs ; \
   exec stress-ng --vm 4 --vm-bytes 4G --vm-keep --vm-method rand-set -t \
120s -q'

Test Case 2a (max_pool_percent=1):
                                Baseline       Patched
shrink_worker wakeups              5,640         1,308
shrink_memcg calls             8,481,500     3,140,972
written_back pages                   260       468,216
zswap_store calls              2,742,756     2,011,269
   store succeeded (ret=1)       934,640       947,988
   store rejected (ret=0)      1,808,116     1,063,281
   store reject rate                ~66%          ~52%
pool_limit_hit delta           1,181,310       196,882
pswpout                        1,808,376     1,531,497
pswpin                         4,288,497     3,635,365
Test Case 2b (zswap.max=320M, max_pool_percent=50):
                                Baseline       Patched
shrink_worker wakeups                 0              0
shrink_memcg calls              687,608         54,002
written_back pages              639,176        846,663
zswap_store calls             1,224,222      1,228,548
   store succeeded (ret=1)      992,816      1,208,123
   store rejected (ret=0)       231,431         20,425
   store reject rate               ~19%            ~2%
pool_limit_hit delta                  0              0
pswpout                         870,745        867,360
pswpin                        1,707,823      1,216,814

Under identical workloads and runtimes, batched zswap shrinking
exhibits a significant reduction in both shrink_worker() wakeups
and shrink_memcg() calls. Furthermore, the sharp drop in both pswpin
and zswap_store() rejections demonstrates that batching zswap shrink
operations effectively mitigates zswap_store() failures caused by
hitting the pool limit. This significantly prevents pages from bypassing
zswap and falling back directly to disk, thereby reducing LRU inversion.

Suggested-by: Yosry Ahmed <yosry@kernel.org>
Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Acked-by: Yosry Ahmed <yosry@kernel.org>
Acked-by: Nhat Pham <nphamcs@gmail.com>
Signed-off-by: Hao Jia <jiahao1@lixiang.com>
---
 mm/zswap.c | 13 +++++++++++--
 1 file changed, 11 insertions(+), 2 deletions(-)

diff --git a/mm/zswap.c b/mm/zswap.c
index 48fc7b575e24..ebbfe85ba7e8 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1275,6 +1275,14 @@ static struct shrinker *zswap_alloc_shrinker(void)
 	return shrinker;
 }
 
+/*
+ * Scan up to SWAP_CLUSTER_MAX pages on each per-node zswap LRU of @memcg
+ * and write back the reclaimable ones.
+ *
+ * Return: 0 if at least one entry was written back, -EAGAIN if entries
+ * were scanned but none could be written back, or -ENOENT if @memcg has
+ * writeback disabled, is a zombie cgroup, or has empty zswap LRUs.
+ */
 static int shrink_memcg(struct mem_cgroup *memcg)
 {
 	int nid, shrunk = 0, scanned = 0;
@@ -1290,13 +1298,14 @@ static int shrink_memcg(struct mem_cgroup *memcg)
 		return -ENOENT;
 
 	for_each_node_state(nid, N_NORMAL_MEMORY) {
-		unsigned long nr_to_walk = 1;
+		unsigned long nr_to_walk = SWAP_CLUSTER_MAX;
 
 		shrunk += list_lru_walk_one(&zswap_list_lru, nid, memcg,
 					    &shrink_memcg_cb, NULL, &nr_to_walk);
-		scanned += 1 - nr_to_walk;
+		scanned += SWAP_CLUSTER_MAX - nr_to_walk;
 	}
 
+	/* Nothing was scanned: every LRU under @memcg was empty. */
 	if (!scanned)
 		return -ENOENT;
 
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH v4 0/2] mm/zswap: Fixes and improves the zswap shrink
  2026-08-06  7:09 [PATCH v4 0/2] mm/zswap: Fixes and improves the zswap shrink Hao Jia
  2026-08-06  7:09 ` [PATCH v4 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
  2026-08-06  7:09 ` [PATCH v4 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
@ 2026-08-06 22:41 ` Andrew Morton
  2026-08-06 22:47   ` Yosry Ahmed
  2 siblings, 1 reply; 5+ messages in thread
From: Andrew Morton @ 2026-08-06 22:41 UTC (permalink / raw)
  To: Hao Jia
  Cc: tj, hannes, shakeel.butt, mhocko, yosry, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Thu,  6 Aug 2026 15:09:41 +0800 Hao Jia <jiahao.kernel@gmail.com> wrote:

> From: Hao Jia <jiahao1@lixiang.com>
> 
> This series fixes and improves the zswap global shrinker (shrink_worker()):
> Patch 1: Fix missing global shrinker when memory cgroup is disabled.
> Patch 2: Extend shrink_memcg() to support batch writeback and thereby improving
>          the writeback efficiency in the shrink_worker() and zswap_store() paths.

Thanks.

Why is a -stable backport proposed for [1/2]?  Its changelog should
describe the effect of the bug upon our users so that others can
understand why this was requested.

AI review suggests there may be an issue in [1/2].  Please check?
	https://sashiko.dev/#/patchset/20260806070943.95542-1-jiahao.kernel@gmail.com



^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v4 0/2] mm/zswap: Fixes and improves the zswap shrink
  2026-08-06 22:41 ` [PATCH v4 0/2] mm/zswap: Fixes and improves the zswap shrink Andrew Morton
@ 2026-08-06 22:47   ` Yosry Ahmed
  0 siblings, 0 replies; 5+ messages in thread
From: Yosry Ahmed @ 2026-08-06 22:47 UTC (permalink / raw)
  To: Andrew Morton
  Cc: Hao Jia, tj, hannes, shakeel.butt, mhocko, mkoutny, nphamcs,
	chengming.zhou, muchun.song, roman.gushchin, linux-mm,
	linux-kernel, linux-doc, Hao Jia

On Thu, Aug 6, 2026 at 3:41 PM Andrew Morton <akpm@linux-foundation.org> wrote:
>
> On Thu,  6 Aug 2026 15:09:41 +0800 Hao Jia <jiahao.kernel@gmail.com> wrote:
>
> > From: Hao Jia <jiahao1@lixiang.com>
> >
> > This series fixes and improves the zswap global shrinker (shrink_worker()):
> > Patch 1: Fix missing global shrinker when memory cgroup is disabled.
> > Patch 2: Extend shrink_memcg() to support batch writeback and thereby improving
> >          the writeback efficiency in the shrink_worker() and zswap_store() paths.
>
> Thanks.
>
> Why is a -stable backport proposed for [1/2]?  Its changelog should
> describe the effect of the bug upon our users so that others can
> understand why this was requested.

It's a potential performance regression for people using zswap without
memcg that was introduced by the commit in "Fixes".

>
> AI review suggests there may be an issue in [1/2].  Please check?
>         https://sashiko.dev/#/patchset/20260806070943.95542-1-jiahao.kernel@gmail.com

Same thing from previous versions, shouldn't be a problem in practice.

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-08-06 22:47 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-06  7:09 [PATCH v4 0/2] mm/zswap: Fixes and improves the zswap shrink Hao Jia
2026-08-06  7:09 ` [PATCH v4 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
2026-08-06  7:09 ` [PATCH v4 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia
2026-08-06 22:41 ` [PATCH v4 0/2] mm/zswap: Fixes and improves the zswap shrink Andrew Morton
2026-08-06 22:47   ` Yosry Ahmed

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox