From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f175.google.com (mail-pf1-f175.google.com [209.85.210.175]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D800F3E0C4A for ; Thu, 6 Aug 2026 07:10:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.175 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786000222; cv=none; b=E94eTOcomNfdKHp1DBfLeW/lZEMHN8y/X1b08cUhN9N6K31YijOL8obKvRbbyRooR5wjRU8QKw8rTAzvoADrG0hFdp37N41fbtz0sQa2HNFsxpMDT0uNdWy5MMf9sF/mA997rryP30I8yhT9N9PnzH1TA7unsaWMbnJs4+Tr900= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786000222; c=relaxed/simple; bh=iuubh3GxaCB8kP2Iv2dCi5tAwFrSn8rh02T1IyQsamw=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=p6QRio2fP8wS6RfYGCjeGct9KO2yuknUPvyEk1xfDhMxGQTz8TCyenDj3aFaYRYyZyJ17KnwX2clzIBLOjQWwwiYVA7+KBGKdJTu81AJOzcQIWx4UDG30398wT+LRzDPoG7Kg9XynuOKIVA2HzvpZEHCGaXmca6UxFGd2BOmIhM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=ppNQQdKH; arc=none smtp.client-ip=209.85.210.175 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="ppNQQdKH" Received: by mail-pf1-f175.google.com with SMTP id d2e1a72fcca58-84862b0d5f8so1752848b3a.3 for ; Thu, 06 Aug 2026 00:10:20 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786000220; x=1786605020; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=NholKuWtQvEG00t+vpLZ7cw0VYR49OMbVtszbQdNqzA=; b=ppNQQdKHGXm0ae8PpYwlLpbJm4wKHgDX6/SQcuZAvOwd2Y5mfNz45CjqiHiHxlWShB fJlaF7VLkV0UoFAi/ai0KTGTk5UjoU/g6ZdfcsHFF193SL5tR6rBOVW7G1UvGQK/6o99 Kt/jhs0kOwkpx/Tvve7OsNeetF4zwJA6hOJ8Iii3g3j0w8Qya8j3jauu3Vx25IiUkQOs 93GGkhCxcLZU1cVDWKfgoo0/fjHavl+wB+pm8PekxVj1p6FpDcwPMFjIwef7pYlJv3c2 G4u1fFivjPZGvKUcfNN8F9ZlYa6YdQay8o2AtE3QUE+W5/FBzxxViqnJjHNwTJyPF194 P+lQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786000220; x=1786605020; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=NholKuWtQvEG00t+vpLZ7cw0VYR49OMbVtszbQdNqzA=; b=Zse/toXBPM7T1V5e/x7M1G6X2EnGfELdB0AfAIjkkRL/oPs+9b0u434Skhzc4Q4349 u+XgDesBp0GY1K518Z3mW7wBsQ0096ANkZCoq6vJsWsAW52F+Fk3pId5wBnH6lKwS+q9 RnFSWDrZzNDv/6JC5k4T6XiCT2LWReje1jNzr4r6NNfvOJNpNeMIm4mxGb3i2VjJ0qhf pgqQMPWGuy3w5nonJBZ7fwx4LzQ1Py0vtCkgw6nbjE8xK1Zl61hx8AS/+epotAm8aOPu 6tcE2WO+qkKxDiRmHn0EdGUkhuf72TJopQD/jYIVOPHDxPcYKvQCKcWbBFmmIxeC/WyH yx6Q== X-Forwarded-Encrypted: i=1; AHgh+Rpd97UQaZ6QlK1EaBF11zNko3qdWaPxHvfqS6OmvHFjoBM1oxwRD0oeTKyz+ggXFQUxudptS3YwpVzSX4I=@vger.kernel.org X-Gm-Message-State: AOJu0YyfY0cy7gMohZsm/9R+Rqr8JMBFna3ZMuaDdglYDMpilCuheT1s jMbgnfeJmvGnOl9jEZ5UJQrS2+H2sbti7WX9sdmn0IbaNY9sSo5XvCND X-Gm-Gg: AR+sD11Y96VmrbuitUHfmj7fg0n456D7YD0S2mHuRIWUar8VNKgjkn8+iylz4ygP0ph 6hEDcMtkjmKRktAdlNJXDplAp3IHqiw1neoyi9xM+8xLeT2rOVnZkwhaCxsqEdPp/tUcyQSNQq0 b+0a3jmXDhbtXMfDpGByvNOkLDm71IEDfRNZ4kTnhiI11uBYy0U3R2Uh9f6KWKanV0OHUxvA40/ XWxAmO6GJuOf+AYwEU/FsF9ilWAyaZ92LAlyIAol0Vjc8QuPp0QQvxw6BF4288oQCe+sR+SQji6 Oe1RLOQoodHVPLuOQ97JXvH7OWMazMbXKcrFqKrrE5WIiOeloqulOmC180neDDEhlb/or/yv8If ARUQKMmQ6bmbFw+bWGQV37HNwuhR+UpS3si9kGeSmhdfyOIPJu+gMR6goX5k91wQu41qVrfJke0 sS6mlP9CHWT7BE2QGNhKubWtTuVEvpK9mwSIa79+754G34+100y8f75+iSRS4abQ2ndXXd0TXOz MmIN2LjMT9gSSFnFcV5G6M9M0I= X-Received: by 2002:a05:6a00:2e1a:b0:848:42a7:1854 with SMTP id d2e1a72fcca58-84f2e0c9c90mr13382273b3a.39.1786000219033; Thu, 06 Aug 2026 00:10:19 -0700 (PDT) Received: from localhost.localdomain ([210.184.73.204]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-84f45bc9cffsm760154b3a.59.2026.08.06.00.10.11 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Thu, 06 Aug 2026 00:10:18 -0700 (PDT) From: Hao Jia To: akpm@linux-foundation.org, tj@kernel.org, hannes@cmpxchg.org, shakeel.butt@linux.dev, mhocko@kernel.org, yosry@kernel.org, mkoutny@suse.com, nphamcs@gmail.com, chengming.zhou@linux.dev, muchun.song@linux.dev, roman.gushchin@linux.dev Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, Hao Jia Subject: [PATCH v4 2/2] mm/zswap: Support batch writeback in shrink_memcg() Date: Thu, 6 Aug 2026 15:09:43 +0800 Message-Id: <20260806070943.95542-3-jiahao.kernel@gmail.com> X-Mailer: git-send-email 2.39.2 (Apple Git-143) In-Reply-To: <20260806070943.95542-1-jiahao.kernel@gmail.com> References: <20260806070943.95542-1-jiahao.kernel@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit From: Hao Jia Currently, shrink_memcg() writes back at most one entry per-node during its traversal. This makes shrink_worker() inefficient, as it must repeatedly re-enter shrink_memcg() to make any substantial progress. Under high memory pressure, this can cause the writeback speed to be too slow to keep up with refaults, leading to zswap store failures and forcing pages to skip zswap and go directly to disk, which results in an LRU inversion. To address this, extend the per-node scan budget in shrink_memcg() from a single entry to up to SWAP_CLUSTER_MAX pages, enabling batch writeback for both the shrink_worker() and zswap_store() paths. Test Setup: - Total memory: 32 GB, 1 NUMA node. - zswap settings: accept_threshold_percent=50, shrinker_enabled=N. Test Case 1: Set max_pool_percent=1, allocate 512MB of anonymous pages, and fill them with random data (to avoid compression). Then, use cgroup memory.reclaim to force a large amount of anonymous pages into zswap. At an interval of 2ms, allocate a 4K anonymous page where the first 4 bytes are random numbers and the rest are zeros, and then trigger reclamation of this 4K page through cgroup memory.reclaim. When the pool threshold is reached, shrink_memcg() will be triggered. The test data after running for 120s is as follows: Baseline Patched shrink_worker wakeups 5,363 169 shrink_memcg calls 11,373,201 350,703 written_back pages 40,212 40,241 zswap_store calls 161,190 163,753 store succeeded (ret=1) 102,743 117,183 store rejected (ret=0) 58,447 46,570 store reject rate ~36% ~28% pool_limit_hit delta 55,826 33,760 pswpout 98,659 86,811 pswpin 2 0 Test Case 2: We evaluated the following two sub-configurations using stress-ng inside a cgroup capped at memory.max=1G for 120 seconds: Test Case 2a (max_pool_percent=1): Continuously triggers the global zswap pool limit, thereby waking up shrink_worker() to perform asynchronous shrinking. Test Case 2b (zswap.max=320M, max_pool_percent=50): Continuously triggers the cgroup's zswap.max limit, thereby invoking synchronous shrinking. Command executed for both setups: bash -c 'echo $$ > /sys/fs/cgroup/zswaptest/cgroup.procs ; \ exec stress-ng --vm 4 --vm-bytes 4G --vm-keep --vm-method rand-set -t \ 120s -q' Test Case 2a (max_pool_percent=1): Baseline Patched shrink_worker wakeups 5,640 1,308 shrink_memcg calls 8,481,500 3,140,972 written_back pages 260 468,216 zswap_store calls 2,742,756 2,011,269 store succeeded (ret=1) 934,640 947,988 store rejected (ret=0) 1,808,116 1,063,281 store reject rate ~66% ~52% pool_limit_hit delta 1,181,310 196,882 pswpout 1,808,376 1,531,497 pswpin 4,288,497 3,635,365 Test Case 2b (zswap.max=320M, max_pool_percent=50): Baseline Patched shrink_worker wakeups 0 0 shrink_memcg calls 687,608 54,002 written_back pages 639,176 846,663 zswap_store calls 1,224,222 1,228,548 store succeeded (ret=1) 992,816 1,208,123 store rejected (ret=0) 231,431 20,425 store reject rate ~19% ~2% pool_limit_hit delta 0 0 pswpout 870,745 867,360 pswpin 1,707,823 1,216,814 Under identical workloads and runtimes, batched zswap shrinking exhibits a significant reduction in both shrink_worker() wakeups and shrink_memcg() calls. Furthermore, the sharp drop in both pswpin and zswap_store() rejections demonstrates that batching zswap shrink operations effectively mitigates zswap_store() failures caused by hitting the pool limit. This significantly prevents pages from bypassing zswap and falling back directly to disk, thereby reducing LRU inversion. Suggested-by: Yosry Ahmed Suggested-by: Johannes Weiner Acked-by: Yosry Ahmed Acked-by: Nhat Pham Signed-off-by: Hao Jia --- mm/zswap.c | 13 +++++++++++-- 1 file changed, 11 insertions(+), 2 deletions(-) diff --git a/mm/zswap.c b/mm/zswap.c index 48fc7b575e24..ebbfe85ba7e8 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1275,6 +1275,14 @@ static struct shrinker *zswap_alloc_shrinker(void) return shrinker; } +/* + * Scan up to SWAP_CLUSTER_MAX pages on each per-node zswap LRU of @memcg + * and write back the reclaimable ones. + * + * Return: 0 if at least one entry was written back, -EAGAIN if entries + * were scanned but none could be written back, or -ENOENT if @memcg has + * writeback disabled, is a zombie cgroup, or has empty zswap LRUs. + */ static int shrink_memcg(struct mem_cgroup *memcg) { int nid, shrunk = 0, scanned = 0; @@ -1290,13 +1298,14 @@ static int shrink_memcg(struct mem_cgroup *memcg) return -ENOENT; for_each_node_state(nid, N_NORMAL_MEMORY) { - unsigned long nr_to_walk = 1; + unsigned long nr_to_walk = SWAP_CLUSTER_MAX; shrunk += list_lru_walk_one(&zswap_list_lru, nid, memcg, &shrink_memcg_cb, NULL, &nr_to_walk); - scanned += 1 - nr_to_walk; + scanned += SWAP_CLUSTER_MAX - nr_to_walk; } + /* Nothing was scanned: every LRU under @memcg was empty. */ if (!scanned) return -ENOENT; -- 2.34.1