Linux Documentation
 help / color / mirror / Atom feed
From: Hao Jia <jiahao.kernel@gmail.com>
To: Yosry Ahmed <yosry@kernel.org>,
	Andrew Morton <akpm@linux-foundation.org>
Cc: tj@kernel.org, hannes@cmpxchg.org, shakeel.butt@linux.dev,
	mhocko@kernel.org, mkoutny@suse.com, nphamcs@gmail.com,
	chengming.zhou@linux.dev, muchun.song@linux.dev,
	roman.gushchin@linux.dev, linux-mm@kvack.org,
	linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org,
	Hao Jia <jiahao1@lixiang.com>,
	stable@vger.kernel.org
Subject: Re: [PATCH v3 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled
Date: Thu, 30 Jul 2026 14:31:26 +0800	[thread overview]
Message-ID: <c4c95102-72e2-a44f-f16e-4134ba6bb060@gmail.com> (raw)
In-Reply-To: <CAO9r8zNhBCghP8T_98z2p==Cfs_=kamcRN_3ROxaj3R8PN2Qvw@mail.gmail.com>



On 2026/7/30 08:30, Yosry Ahmed wrote:
> On Wed, Jul 29, 2026 at 3:58 PM Andrew Morton <akpm@linux-foundation.org> wrote:
>>
>> On Wed, 29 Jul 2026 16:42:05 +0800 Hao Jia <jiahao.kernel@gmail.com> wrote:
>>
>>> Zswap writeback when the global pool limit is hit fails when memory
>>> cgroup is disabled. The pool remains full until it is organically
>>> drained by swapins or memory freeing, leading to zswap store failures
>>> and pages bypassing getting written directly to the backing swap device,
>>> causing LRU inversion (hotter pages with higher fault latency).
>>>
>>> This happens because mem_cgroup_iter() always returns NULL when
>>> memory cgroups are disabled. As a result, the global shrinker
>>> shrink_worker() repeatedly takes empty walks. After MAX_RECLAIM_RETRIES
>>> failed attempts, the worker gives up without writing back any pages.
>>>
>>> Therefore, when memory cgroup is disabled, fall through with the !memcg
>>> branch and shrink the root memcg directly.
>>>
>>> With memcg disabled, shrink_memcg() only returns -ENOENT when the root
>>> LRU is empty, which means the total pages are already below thr. In the
>>> absence of heavy concurrent zswap stores, the loop then safely bails out
>>> via the zswap_total_pages() <= thr check; otherwise, it will resume
>>> shrinking the memcg after processing the reschedule check. For any other
>>> return value from shrink_memcg(), the loop is guaranteed to terminate,
>>> either after MAX_RECLAIM_RETRIES failures or once the threshold is met.
>>>
>>> Fixes: a65b0e7607cc ("zswap: make shrinking memcg-aware")
>>> Cc: stable@vger.kernel.org
>>
>> How does this affect users?  What behavior do they observe when it
>> occurs?
> 
> I think the first paragraph sums it up pretty well, especially the
> last sentence "hotter pages with higher fault latency".
> 
>>
>>> Closes: https://lore.kernel.org/all/CAO9r8zPVzMKFbCixxD-qgtRrkFxWVrHiZZeLc=eyTPKPVQgX4g@mail.gmail.com
>>
>> hm, that isn't really a bug report and doesn't answer the above
>> question.
> 
> Yeah, it isn't. Probably we should drop "Closes". I assume Hao added
> it because checkpatch annoyingly complains if you add "Reported-by"
> without "Closes", so Hao just linked to the thread where I pointed out
> the bug.

Yeah, checkpatch will complain if it's missing.

> 
>>
>> AI review asked a couple of questions:
>>          https://sashiko.dev/#/patchset/20260729084206.77793-1-jiahao.kernel@gmail.com
> 
> The review on patch #1 is something theoretical, we discussed it at
> length in previous versions.
> 
> For patch #2:
> 
>> Does this batching logic break NUMA fairness?
>>
>> Because for_each_node_state() always starts from the lowest node
>> ID and breaks when the scan budget is exhausted, subsequent
>> calls to shrink_memcg() will restart at the lowest node ID again.
>>
>> If the lowest node (typically Node 0) consistently has enough
>> items to exhaust the scan budget, wouldn't we exclusively evict
>> pages from it while ignoring older pages on other nodes? Could
>> this cause LRU inversion across nodes, keeping older pages in
>> memory on Node 1 while hot pages on Node 0 are evicted?
> 
> Yes, unfairness is possible.
> 
> For global shrinking, it's probably not an issue. We reclaim until we
> hit the acceptance threshold and it's very unlikely this will happen
> before iterating all nodes (given that the batch size is 32 pages).
> However, with the shrink_memcg() path, we only reclaim one batch, so
> there's a chance we'll always reclaim it from node 0.
> 
> Maybe we should just drop the early bailout and accept potentially
> doing more writeback than needed. Hao, WDYT?
> 
If we scan and attempt to write back SWAP_CLUSTER_MAX zswap entries per 
node, it might lead to excessive writeback on machines with many NUMA 
nodes. Furthermore, I'm concerned about introducing higher latency in 
synchronous shrink paths like zswap_store()—especially on systems with a 
large number of NUMA nodes, where it could end up writing back hundreds 
of pages in a single call.

Maybe we could do something like this instead? That way, in the 
worst-case scenario, it falls back to the baseline behavior without 
introducing any extra latency risks.

  static int shrink_memcg(struct mem_cgroup *memcg)
  {
-       int nid, shrunk = 0, scanned = 0;
+       unsigned long node_batch, scanned = 0;
+       int nid, shrunk = 0;

         if (!mem_cgroup_zswap_writeback_enabled(memcg))
                 return -ENOENT;
@@ -1289,14 +1313,26 @@ static int shrink_memcg(struct mem_cgroup *memcg)
         if (memcg && !mem_cgroup_online(memcg))
                 return -ENOENT;

+       node_batch = max(1UL, SWAP_CLUSTER_MAX / 
num_node_state(N_NORMAL_MEMORY));
         for_each_node_state(nid, N_NORMAL_MEMORY) {
-               unsigned long nr_to_walk = 1;
+               unsigned long nr_to_walk, budget;
+
+               /*
+                * Cap the scan at the per-node LRU length so each entry is
+                * scanned at most once per call.
+                */
+               budget = min(node_batch,
+                            list_lru_count_one(&zswap_list_lru, nid, 
memcg));
+               if (!budget)
+                       continue;

+               nr_to_walk = budget;
                 shrunk += list_lru_walk_one(&zswap_list_lru, nid, memcg,
                                             &shrink_memcg_cb, NULL, 
&nr_to_walk);
-               scanned += 1 - nr_to_walk;
+               scanned += budget - nr_to_walk;
         }

+       /* Nothing was scanned: every LRU under @memcg was empty. */
         if (!scanned)
                 return -ENOENT;

Thanks,
Hao
> If you respin, please also drop the batch size argument to
> shrink_memcg() as it's now always SWAP_CLUSTER_MAX.

  parent reply	other threads:[~2026-07-30  6:31 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-29  8:42 [PATCH v3 0/2] mm/zswap: Fixes and improves the zswap shrink Hao Jia
2026-07-29  8:42 ` [PATCH v3 1/2] mm/zswap: Fix global shrinker when memory cgroup is disabled Hao Jia
2026-07-29 22:58   ` Andrew Morton
2026-07-30  0:30     ` Yosry Ahmed
2026-07-30  1:17       ` Andrew Morton
2026-07-30  6:31       ` Hao Jia [this message]
2026-07-29  8:42 ` [PATCH v3 2/2] mm/zswap: Support batch writeback in shrink_memcg() Hao Jia

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=c4c95102-72e2-a44f-f16e-4134ba6bb060@gmail.com \
    --to=jiahao.kernel@gmail.com \
    --cc=akpm@linux-foundation.org \
    --cc=chengming.zhou@linux.dev \
    --cc=hannes@cmpxchg.org \
    --cc=jiahao1@lixiang.com \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=mhocko@kernel.org \
    --cc=mkoutny@suse.com \
    --cc=muchun.song@linux.dev \
    --cc=nphamcs@gmail.com \
    --cc=roman.gushchin@linux.dev \
    --cc=shakeel.butt@linux.dev \
    --cc=stable@vger.kernel.org \
    --cc=tj@kernel.org \
    --cc=yosry@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox