* [RFC PATCH v4 0/3] mm/zswap: shrink zswap_entry via a fixed pool index
@ 2026-08-30 11:47 Jianyue Wu
2026-08-30 11:47 ` [RFC PATCH v4 1/3] mm/zswap: release retired pools via queue_rcu_work() instead of synchronize_rcu() Jianyue Wu
` (2 more replies)
0 siblings, 3 replies; 12+ messages in thread
From: Jianyue Wu @ 2026-08-30 11:47 UTC (permalink / raw)
To: Johannes Weiner, Yosry Ahmed, Nhat Pham, Chengming Zhou,
Andrew Morton
Cc: Jianyue Wu, Chris Li, linux-mm, linux-kernel
Every stored page has a struct zswap_entry, so its size is pure per-page
overhead. On x86_64 it is currently 56 bytes, of which 8 bytes are a
pointer to the owning zswap_pool.
Only a handful of pools are ever live: a new pool is created only when
the compressor is (re)set, and pools are reused across compressor
switches. That makes a per-entry pool pointer more expensive than it
needs to be, and the RCU list that currently tracks pools is more
machinery than this needs once each pool already has a stable slot.
This series:
1. Releases retired pools with queue_rcu_work() instead of
synchronize_rcu() so the last put no longer blocks on an RCU grace
period.
2. Replaces the zswap_pools list with a fixed ZSWAP_MAX_POOLS (16)
array and a separate RCU-protected current-pool pointer, giving
each pool a stable slot index. Slot 0 is left unused, and pool
creation publishes the fully constructed pool directly into a free
slot.
3. Stores that u8 slot index in each zswap_entry instead of the pool
pointer. The u8 fits in padding after the bool referenced field,
so the entry shrinks from 56 to 48 bytes on x86_64 (~2MiB of
metadata saved per 1GiB of data held in zswap).
Runtime compressor switching is preserved, but the fixed array now
bounds the number of simultaneously live distinct compressor pools to 15
(slot 0 is reserved). If all usable slots are full, creating a pool for
another compressor fails and the compressor switch is rejected.
Benchmark (x86_64, compressor=lzo, MADV_PAGEOUT store + fault-in load):
- zswap_entry object_size: 56 -> 48 bytes
- e2e store+load median latency: no measurable regression vs baseline
at matched stored_delta
The extra cost per store/free/decompress is one array-index load
instead of a pointer dereference. With a single (or few) live pool(s)
that does not show up against (de)compression.
Testing
=======
- sizeof_check: 56 -> 48 bytes on x86_64
- Boot with DEBUG_ATOMIC_SLEEP + lockdep/PROVE_RCU + KASAN:
zswap store/load and compressor switch (retire + reuse) pass
This series is based on akpm/mm-unstable as of 2026-08-30
(42d64d4fef83).
Signed-off-by: Jianyue Wu <wujianyue000@gmail.com>
Changes since RFC v3:
- Retire pools with queue_rcu_work() instead of call_rcu().
- Queue the deferred RCU release work on system_percpu_wq.
- Use rcu_dereference_check() with entry->pool_idx != 0 in
zswap_entry_pool() instead of rcu_dereference_protected().
- Simplify pool-slot publishing after confirming compressor parameter
updates are serialized by the module parameter lock.
Link: https://lore.kernel.org/all/20260815-shrink_zswap_entry_0815_v2-v3-3-0171bd86a667@gmail.com/
Link: https://lore.kernel.org/all/20260731-shrink_zswap_entry_v2-0-0-v2-0-e72083aa8734@gmail.com/
Link: https://lore.kernel.org/all/20260726-shrink_zswap_entry_v1-0-0-v1-1-30957e4d0cb6@gmail.com/
Jianyue Wu (3):
mm/zswap: release retired pools via queue_rcu_work() instead of
synchronize_rcu()
mm/zswap: replace the zswap_pools list with a fixed pools array
mm/zswap: reference the pool by index to shrink struct zswap_entry
mm/zswap.c | 140 ++++++++++++++++++++++++++++++++++++++---------------
1 file changed, 101 insertions(+), 39 deletions(-)
base-commit: 42d64d4fef83a241c919c8693fdf0a21b2cb6061
--
2.43.0
^ permalink raw reply [flat|nested] 12+ messages in thread
* [RFC PATCH v4 1/3] mm/zswap: release retired pools via queue_rcu_work() instead of synchronize_rcu()
2026-08-30 11:47 [RFC PATCH v4 0/3] mm/zswap: shrink zswap_entry via a fixed pool index Jianyue Wu
@ 2026-08-30 11:47 ` Jianyue Wu
2026-08-31 15:20 ` Yosry Ahmed
2026-09-01 15:38 ` Johannes Weiner
2026-08-30 11:47 ` [RFC PATCH v4 2/3] mm/zswap: replace the zswap_pools list with a fixed pools array Jianyue Wu
2026-08-30 11:47 ` [RFC PATCH v4 3/3] mm/zswap: reference the pool by index to shrink struct zswap_entry Jianyue Wu
2 siblings, 2 replies; 12+ messages in thread
From: Jianyue Wu @ 2026-08-30 11:47 UTC (permalink / raw)
To: Johannes Weiner, Yosry Ahmed, Nhat Pham, Chengming Zhou,
Andrew Morton
Cc: Jianyue Wu, Chris Li, linux-mm, linux-kernel
When a pool's last reference is dropped, __zswap_pool_empty() removes it
from the pool list and schedules __zswap_pool_release(), which calls
synchronize_rcu() to wait for readers before tearing the pool down.
synchronize_rcu() is a synchronous, potentially long wait. Replace it
with queue_rcu_work(): __zswap_pool_empty() hands the pool to
queue_rcu_work(), which waits for a grace period asynchronously and then
runs __zswap_pool_release() from a worker for the sleepable teardown
(__zswap_pool_empty() can run in atomic context and must not block).
The grace-period guarantee is unchanged; the retirement path just no
longer blocks on it.
Suggested-by: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Jianyue Wu <wujianyue000@gmail.com>
---
mm/zswap.c | 12 +++++-------
1 file changed, 5 insertions(+), 7 deletions(-)
diff --git a/mm/zswap.c b/mm/zswap.c
index 37f34e406c8e..0bb30e58950a 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -155,7 +155,7 @@ struct zswap_pool {
struct crypto_acomp_ctx __percpu *acomp_ctx;
struct percpu_ref ref;
struct list_head list;
- struct work_struct release_work;
+ struct rcu_work release_work;
struct hlist_node node;
char tfm_name[CRYPTO_MAX_ALG_NAME];
};
@@ -379,10 +379,8 @@ static void zswap_pool_destroy(struct zswap_pool *pool)
static void __zswap_pool_release(struct work_struct *work)
{
- struct zswap_pool *pool = container_of(work, typeof(*pool),
- release_work);
-
- synchronize_rcu();
+ struct zswap_pool *pool = container_of(to_rcu_work(work),
+ typeof(*pool), release_work);
/* nobody should have been able to get a ref... */
WARN_ON(!percpu_ref_is_zero(&pool->ref));
@@ -406,8 +404,8 @@ static void __zswap_pool_empty(struct percpu_ref *ref)
list_del_rcu(&pool->list);
- INIT_WORK(&pool->release_work, __zswap_pool_release);
- schedule_work(&pool->release_work);
+ INIT_RCU_WORK(&pool->release_work, __zswap_pool_release);
+ queue_rcu_work(system_percpu_wq, &pool->release_work);
spin_unlock_bh(&zswap_pools_lock);
}
--
2.43.0
^ permalink raw reply related [flat|nested] 12+ messages in thread
* [RFC PATCH v4 2/3] mm/zswap: replace the zswap_pools list with a fixed pools array
2026-08-30 11:47 [RFC PATCH v4 0/3] mm/zswap: shrink zswap_entry via a fixed pool index Jianyue Wu
2026-08-30 11:47 ` [RFC PATCH v4 1/3] mm/zswap: release retired pools via queue_rcu_work() instead of synchronize_rcu() Jianyue Wu
@ 2026-08-30 11:47 ` Jianyue Wu
2026-08-31 15:28 ` Yosry Ahmed
2026-09-01 16:13 ` Johannes Weiner
2026-08-30 11:47 ` [RFC PATCH v4 3/3] mm/zswap: reference the pool by index to shrink struct zswap_entry Jianyue Wu
2 siblings, 2 replies; 12+ messages in thread
From: Jianyue Wu @ 2026-08-30 11:47 UTC (permalink / raw)
To: Johannes Weiner, Yosry Ahmed, Nhat Pham, Chengming Zhou,
Andrew Morton
Cc: Jianyue Wu, Chris Li, linux-mm, linux-kernel
Originally zswap holds its pools on an RCU list whose head also serves
as the "current pool". Only a handful of pools are ever live at once,
since a new pool is only created when the compressor is (re)set and
pools are reused across compressor switches.
Hold the pools in a fixed ZSWAP_MAX_POOLS-element array so each pool
has a stable slot number, and track the current pool with a separate
rcu-protected pointer.
Slot 0 is intentionally left unused (always NULL): a zeroed or
incorrectly initialized pool index then resolves to NULL and trips a
WARN rather than silently aliasing a live pool in another slot.
The array keeps the same RCU publish/retire discipline the list had,
so lookup and teardown stay equivalent. A fully-constructed pool is
stored into its slot as the last step of zswap_pool_create(), so array
walkers only ever observe a NULL slot or a ready pool. Pool creation
is serialized by the module-wide kernel param mutex (all built-in
params share one lock) and otherwise only happens during
single-threaded init, so no two creators race for a slot.
zswap_pools_lock still serializes the store against a retiring pool
clearing its slot in __zswap_pool_empty().
Behavior change: the fixed array bounds the number of simultaneously
live pools at ZSWAP_MAX_POOLS - 1 (15, since slot 0 is reserved),
whereas the old list was unbounded. A pool is only live while it is
the current pool or still has stored pages referencing it, and pools
are reused across compressor switches, so 15 is far more than any real
configuration needs. Once all slots are occupied, creating a pool for
a 16th distinct compressor fails: zswap_pool_create() errors and
returns NULL, and the compressor switch is rejected with -EINVAL
rather than silently succeeding. The cap can be raised by increasing
ZSWAP_MAX_POOLS (bounded by the u8 slot index, so up to 256).
Suggested-by: Nhat Pham <nphamcs@gmail.com>
Suggested-by: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Jianyue Wu <wujianyue000@gmail.com>
---
mm/zswap.c | 97 ++++++++++++++++++++++++++++++++++++++++--------------
1 file changed, 72 insertions(+), 25 deletions(-)
diff --git a/mm/zswap.c b/mm/zswap.c
index 0bb30e58950a..b3b5e2887c00 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -13,6 +13,7 @@
#define pr_fmt(fmt) KBUILD_MODNAME ": " fmt
+#include <linux/cleanup.h>
#include <linux/module.h>
#include <linux/cpu.h>
#include <linux/highmem.h>
@@ -154,12 +155,27 @@ struct zswap_pool {
struct zs_pool *zs_pool;
struct crypto_acomp_ctx __percpu *acomp_ctx;
struct percpu_ref ref;
- struct list_head list;
struct rcu_work release_work;
struct hlist_node node;
+ u8 idx;
char tfm_name[CRYPTO_MAX_ALG_NAME];
};
+#define ZSWAP_MAX_POOLS 16
+/*
+ * Slot 0 is intentionally never used: it stays NULL so that a zeroed or
+ * incorrectly initialized pool->idx resolves to NULL (and trips a WARN)
+ * instead of silently aliasing a live pool in another slot.
+ */
+#define ZSWAP_FIRST_POOL_SLOT 1
+static struct zswap_pool __rcu *zswap_pools[ZSWAP_MAX_POOLS];
+static_assert(ZSWAP_MAX_POOLS - 1 <= U8_MAX);
+/*
+ * The current pool (NULL if none): an alias of one zswap_pools[] slot.
+ * It always holds a ref, so a pool is never retired while it is current.
+ */
+static struct zswap_pool __rcu *zswap_current_pool;
+
/* Global LRU lists shared by all zswap pools. */
static struct list_lru zswap_list_lru;
@@ -200,9 +216,6 @@ struct zswap_entry {
static struct xarray *zswap_trees[MAX_SWAPFILES];
static unsigned int nr_zswap_trees[MAX_SWAPFILES];
-/* RCU-protected iteration */
-static LIST_HEAD(zswap_pools);
-/* protects zswap_pools list modification */
static DEFINE_SPINLOCK(zswap_pools_lock);
/* pool counter to provide unique names to zsmalloc */
static atomic_t zswap_pools_count = ATOMIC_INIT(0);
@@ -270,6 +283,31 @@ static void acomp_ctx_free(struct crypto_acomp_ctx *acomp_ctx)
acomp_ctx->buffer = NULL;
}
+/*
+ * Publish a fully-constructed pool into a free array slot. Pool creation is
+ * serialized by the module-wide kernel param mutex (all built-in params share
+ * one lock) and only otherwise happens during single-threaded init, so no two
+ * creators race for a slot. The pool is complete before it is stored, and
+ * zswap_pools_lock still serializes this store against a concurrent retiring
+ * pool clearing its slot in __zswap_pool_empty(), so array walkers only ever
+ * observe a NULL slot or a ready pool.
+ */
+static int zswap_pool_assign_slot(struct zswap_pool *pool)
+{
+ int i;
+
+ guard(spinlock_bh)(&zswap_pools_lock);
+ for (i = ZSWAP_FIRST_POOL_SLOT; i < ZSWAP_MAX_POOLS; i++) {
+ if (!rcu_access_pointer(zswap_pools[i])) {
+ pool->idx = i;
+ rcu_assign_pointer(zswap_pools[i], pool);
+ return i;
+ }
+ }
+
+ return -ENOSPC;
+}
+
static struct zswap_pool *zswap_pool_create(char *compressor)
{
struct zswap_pool *pool;
@@ -313,19 +351,29 @@ static struct zswap_pool *zswap_pool_create(char *compressor)
if (ret)
goto cpuhp_add_fail;
- /* being the current pool takes 1 ref; this func expects the
- * caller to always add the new pool as the current pool
+ /*
+ * The initial ref keeps the pool alive while it is current. Stored
+ * entries take additional refs so a retired pool remains alive while
+ * any entries still reference it.
*/
ret = percpu_ref_init(&pool->ref, __zswap_pool_empty,
PERCPU_REF_ALLOW_REINIT, GFP_KERNEL);
if (ret)
goto ref_fail;
- INIT_LIST_HEAD(&pool->list);
+
+ ret = zswap_pool_assign_slot(pool);
+ if (ret < 0) {
+ pr_err("cannot create more than %d pools\n",
+ ZSWAP_MAX_POOLS - ZSWAP_FIRST_POOL_SLOT);
+ goto slot_fail;
+ }
zswap_pool_debug("created", pool);
return pool;
+slot_fail:
+ percpu_ref_exit(&pool->ref);
ref_fail:
cpuhp_state_remove_instance(CPUHP_MM_ZSWP_POOL_PREPARE, &pool->node);
@@ -386,7 +434,6 @@ static void __zswap_pool_release(struct work_struct *work)
WARN_ON(!percpu_ref_is_zero(&pool->ref));
percpu_ref_exit(&pool->ref);
- /* pool is now off zswap_pools list and has no references. */
zswap_pool_destroy(pool);
}
@@ -402,7 +449,7 @@ static void __zswap_pool_empty(struct percpu_ref *ref)
WARN_ON(pool == zswap_pool_current());
- list_del_rcu(&pool->list);
+ rcu_assign_pointer(zswap_pools[pool->idx], NULL);
INIT_RCU_WORK(&pool->release_work, __zswap_pool_release);
queue_rcu_work(system_percpu_wq, &pool->release_work);
@@ -433,7 +480,8 @@ static struct zswap_pool *__zswap_pool_current(void)
{
struct zswap_pool *pool;
- pool = list_first_or_null_rcu(&zswap_pools, typeof(*pool), list);
+ pool = rcu_dereference_check(zswap_current_pool,
+ lockdep_is_held(&zswap_pools_lock));
WARN_ONCE(!pool && zswap_has_pool,
"%s: no page storage pool!\n", __func__);
@@ -466,11 +514,12 @@ static struct zswap_pool *zswap_pool_current_get(void)
static struct zswap_pool *zswap_pool_find_get(char *compressor)
{
struct zswap_pool *pool;
+ int i;
- assert_spin_locked(&zswap_pools_lock);
-
- list_for_each_entry_rcu(pool, &zswap_pools, list) {
- if (strcmp(pool->tfm_name, compressor))
+ for (i = ZSWAP_FIRST_POOL_SLOT; i < ZSWAP_MAX_POOLS; i++) {
+ pool = rcu_dereference_protected(zswap_pools[i],
+ lockdep_is_held(&zswap_pools_lock));
+ if (!pool || strcmp(pool->tfm_name, compressor))
continue;
/* if we can't get it, it's about to be destroyed */
if (!zswap_pool_tryget(pool))
@@ -495,10 +544,14 @@ unsigned long zswap_total_pages(void)
{
struct zswap_pool *pool;
unsigned long total = 0;
+ int i;
rcu_read_lock();
- list_for_each_entry_rcu(pool, &zswap_pools, list)
- total += zs_get_total_pages(pool->zs_pool);
+ for (i = ZSWAP_FIRST_POOL_SLOT; i < ZSWAP_MAX_POOLS; i++) {
+ pool = rcu_dereference(zswap_pools[i]);
+ if (pool)
+ total += zs_get_total_pages(pool->zs_pool);
+ }
rcu_read_unlock();
return total;
@@ -560,7 +613,6 @@ static int zswap_compressor_param_set(const char *val, const struct kernel_param
if (pool) {
zswap_pool_debug("using existing", pool);
WARN_ON(pool == zswap_pool_current());
- list_del_rcu(&pool->list);
}
spin_unlock_bh(&zswap_pools_lock);
@@ -588,15 +640,9 @@ static int zswap_compressor_param_set(const char *val, const struct kernel_param
if (!ret) {
put_pool = zswap_pool_current();
- list_add_rcu(&pool->list, &zswap_pools);
+ rcu_assign_pointer(zswap_current_pool, pool);
zswap_has_pool = true;
} else if (pool) {
- /*
- * Add the possibly pre-existing pool to the end of the pools
- * list; if it's new (and empty) then it'll be removed and
- * destroyed by the put after we drop the lock
- */
- list_add_tail_rcu(&pool->list, &zswap_pools);
put_pool = pool;
}
@@ -1801,7 +1847,8 @@ static int zswap_setup(void)
pool = __zswap_pool_create_fallback();
if (pool) {
pr_info("loaded using pool %s\n", pool->tfm_name);
- list_add(&pool->list, &zswap_pools);
+ /* zswap_pool_create() already stored the pool in its array slot. */
+ rcu_assign_pointer(zswap_current_pool, pool);
zswap_has_pool = true;
static_branch_enable(&zswap_ever_enabled);
} else {
--
2.43.0
^ permalink raw reply related [flat|nested] 12+ messages in thread
* [RFC PATCH v4 3/3] mm/zswap: reference the pool by index to shrink struct zswap_entry
2026-08-30 11:47 [RFC PATCH v4 0/3] mm/zswap: shrink zswap_entry via a fixed pool index Jianyue Wu
2026-08-30 11:47 ` [RFC PATCH v4 1/3] mm/zswap: release retired pools via queue_rcu_work() instead of synchronize_rcu() Jianyue Wu
2026-08-30 11:47 ` [RFC PATCH v4 2/3] mm/zswap: replace the zswap_pools list with a fixed pools array Jianyue Wu
@ 2026-08-30 11:47 ` Jianyue Wu
2026-08-31 15:30 ` Yosry Ahmed
2 siblings, 1 reply; 12+ messages in thread
From: Jianyue Wu @ 2026-08-30 11:47 UTC (permalink / raw)
To: Johannes Weiner, Yosry Ahmed, Nhat Pham, Chengming Zhou,
Andrew Morton
Cc: Jianyue Wu, Chris Li, linux-mm, linux-kernel
struct zswap_entry is one allocation per stored page, so its size is
pure overhead. It currently embeds an 8-byte pool pointer, even though
the live pools now sit in a small fixed array indexed by a u8 slot
number.
Replace the per-entry pool pointer with that u8 slot index and resolve
it through the fixed pool array. A live entry holds a reference to its
pool, so the slot cannot be reused under it. The lookup therefore needs
no RCU read-side section or zswap_pools_lock.
The u8 fits in the padding after the bool referenced field, shrinking
the entry from 56 to 48 bytes on x86_64. This raises objs_per_slab from
73 to 85 and saves about 2MiB of metadata per 1GiB of data held in
zswap.
Suggested-by: Chris Li <chrisl@kernel.org>
Signed-off-by: Jianyue Wu <wujianyue000@gmail.com>
---
mm/zswap.c | 31 ++++++++++++++++++++++++-------
1 file changed, 24 insertions(+), 7 deletions(-)
diff --git a/mm/zswap.c b/mm/zswap.c
index b3b5e2887c00..521e0187bcd1 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -198,7 +198,7 @@ static struct shrinker *zswap_shrinker;
* writeback logic. The entry is only reclaimed by the writeback
* logic if referenced is unset. See comments in the shrinker
* section for context.
- * pool - the zswap_pool the entry's data is in
+ * pool_idx - slot of the zswap_pool that the entry's data is in.
* handle - zsmalloc allocation handle that stores the compressed page data
* objcg - the obj_cgroup that the compressed memory is charged to
* lru - handle to the pool's lru used to evict pages.
@@ -207,12 +207,22 @@ struct zswap_entry {
swp_entry_t swpentry;
unsigned int length;
bool referenced;
- struct zswap_pool *pool;
+ u8 pool_idx;
unsigned long handle;
struct obj_cgroup *objcg;
struct list_head lru;
};
+static struct zswap_pool *zswap_entry_pool(struct zswap_entry *entry)
+{
+ /*
+ * A live entry holds a pool reference, so the slot stays valid with no
+ * RCU read-side section. The != 0 check marks access protected by
+ * the reference. A live entry never uses the reserved slot 0.
+ */
+ return rcu_dereference_check(zswap_pools[entry->pool_idx], entry->pool_idx != 0);
+}
+
static struct xarray *zswap_trees[MAX_SWAPFILES];
static unsigned int nr_zswap_trees[MAX_SWAPFILES];
@@ -808,9 +818,13 @@ static void zswap_entry_cache_free(struct zswap_entry *entry)
*/
static void zswap_entry_free(struct zswap_entry *entry)
{
+ struct zswap_pool *pool = zswap_entry_pool(entry);
+
zswap_lru_del(entry);
- zs_free(entry->pool->zs_pool, entry->handle);
- zswap_pool_put(entry->pool);
+ if (!WARN_ON_ONCE(!pool)) {
+ zs_free(pool->zs_pool, entry->handle);
+ zswap_pool_put(pool);
+ }
if (entry->objcg) {
obj_cgroup_uncharge_zswap(entry->objcg, entry->length);
obj_cgroup_put(entry->objcg);
@@ -967,12 +981,15 @@ static bool zswap_compress(struct page *page, struct zswap_entry *entry,
static bool zswap_decompress(struct zswap_entry *entry, struct folio *folio)
{
- struct zswap_pool *pool = entry->pool;
+ struct zswap_pool *pool = zswap_entry_pool(entry);
struct scatterlist input[2]; /* zsmalloc returns an SG list 1-2 entries */
struct scatterlist output;
struct crypto_acomp_ctx *acomp_ctx;
int ret = 0, dlen;
+ if (WARN_ON_ONCE(!pool))
+ return false;
+
acomp_ctx = raw_cpu_ptr(pool->acomp_ctx);
mutex_lock(&acomp_ctx->mutex);
zs_obj_read_sg_begin(pool->zs_pool, entry->handle, input, entry->length);
@@ -1008,7 +1025,7 @@ static bool zswap_decompress(struct zswap_entry *entry, struct folio *folio)
pr_alert_ratelimited("Decompression error from zswap (%d:%lu %s %u->%d)\n",
swp_type(entry->swpentry),
swp_offset(entry->swpentry),
- entry->pool->tfm_name,
+ pool->tfm_name,
entry->length, dlen);
return false;
}
@@ -1511,7 +1528,7 @@ static bool zswap_store_page(struct page *page,
* The publishing order matters to prevent writeback from seeing
* an incoherent entry.
*/
- entry->pool = pool;
+ entry->pool_idx = pool->idx;
entry->swpentry = page_swpentry;
entry->objcg = objcg;
entry->referenced = true;
--
2.43.0
^ permalink raw reply related [flat|nested] 12+ messages in thread
* Re: [RFC PATCH v4 1/3] mm/zswap: release retired pools via queue_rcu_work() instead of synchronize_rcu()
2026-08-30 11:47 ` [RFC PATCH v4 1/3] mm/zswap: release retired pools via queue_rcu_work() instead of synchronize_rcu() Jianyue Wu
@ 2026-08-31 15:20 ` Yosry Ahmed
2026-09-01 14:33 ` Jianyue Wu
2026-09-01 15:38 ` Johannes Weiner
1 sibling, 1 reply; 12+ messages in thread
From: Yosry Ahmed @ 2026-08-31 15:20 UTC (permalink / raw)
To: Jianyue Wu
Cc: Johannes Weiner, Nhat Pham, Chengming Zhou, Andrew Morton,
Chris Li, linux-mm, linux-kernel
On Sun, Aug 30, 2026 at 4:47 AM Jianyue Wu <wujianyue000@gmail.com> wrote:
>
> When a pool's last reference is dropped, __zswap_pool_empty() removes it
> from the pool list and schedules __zswap_pool_release(), which calls
> synchronize_rcu() to wait for readers before tearing the pool down.
>
> synchronize_rcu() is a synchronous, potentially long wait. Replace it
> with queue_rcu_work(): __zswap_pool_empty() hands the pool to
> queue_rcu_work(), which waits for a grace period asynchronously and then
> runs __zswap_pool_release() from a worker for the sleepable teardown
> (__zswap_pool_empty() can run in atomic context and must not block).
> The grace-period guarantee is unchanged; the retirement path just no
> longer blocks on it.
>
> Suggested-by: Yosry Ahmed <yosry@kernel.org>
> Signed-off-by: Jianyue Wu <wujianyue000@gmail.com>
Why are the patches still tagged RFC?
Anyway, with one nit below:
Acked-by: Yosry Ahmed <yosry@kernel.org>
> ---
> mm/zswap.c | 12 +++++-------
> 1 file changed, 5 insertions(+), 7 deletions(-)
>
> diff --git a/mm/zswap.c b/mm/zswap.c
> index 37f34e406c8e..0bb30e58950a 100644
> --- a/mm/zswap.c
> +++ b/mm/zswap.c
> @@ -155,7 +155,7 @@ struct zswap_pool {
> struct crypto_acomp_ctx __percpu *acomp_ctx;
> struct percpu_ref ref;
> struct list_head list;
> - struct work_struct release_work;
> + struct rcu_work release_work;
Seems like the convention is to use rwork in the name instead of work.
> struct hlist_node node;
> char tfm_name[CRYPTO_MAX_ALG_NAME];
> };
> @@ -379,10 +379,8 @@ static void zswap_pool_destroy(struct zswap_pool *pool)
>
> static void __zswap_pool_release(struct work_struct *work)
> {
> - struct zswap_pool *pool = container_of(work, typeof(*pool),
> - release_work);
> -
> - synchronize_rcu();
> + struct zswap_pool *pool = container_of(to_rcu_work(work),
> + typeof(*pool), release_work);
>
> /* nobody should have been able to get a ref... */
> WARN_ON(!percpu_ref_is_zero(&pool->ref));
> @@ -406,8 +404,8 @@ static void __zswap_pool_empty(struct percpu_ref *ref)
>
> list_del_rcu(&pool->list);
>
> - INIT_WORK(&pool->release_work, __zswap_pool_release);
> - schedule_work(&pool->release_work);
> + INIT_RCU_WORK(&pool->release_work, __zswap_pool_release);
> + queue_rcu_work(system_percpu_wq, &pool->release_work);
>
> spin_unlock_bh(&zswap_pools_lock);
> }
> --
> 2.43.0
>
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH v4 2/3] mm/zswap: replace the zswap_pools list with a fixed pools array
2026-08-30 11:47 ` [RFC PATCH v4 2/3] mm/zswap: replace the zswap_pools list with a fixed pools array Jianyue Wu
@ 2026-08-31 15:28 ` Yosry Ahmed
2026-09-01 16:13 ` Johannes Weiner
1 sibling, 0 replies; 12+ messages in thread
From: Yosry Ahmed @ 2026-08-31 15:28 UTC (permalink / raw)
To: Jianyue Wu
Cc: Johannes Weiner, Nhat Pham, Chengming Zhou, Andrew Morton,
Chris Li, linux-mm, linux-kernel
On Sun, Aug 30, 2026 at 4:47 AM Jianyue Wu <wujianyue000@gmail.com> wrote:
>
> Originally zswap holds its pools on an RCU list whose head also serves
> as the "current pool". Only a handful of pools are ever live at once,
> since a new pool is only created when the compressor is (re)set and
> pools are reused across compressor switches.
>
> Hold the pools in a fixed ZSWAP_MAX_POOLS-element array so each pool
> has a stable slot number, and track the current pool with a separate
> rcu-protected pointer.
>
> Slot 0 is intentionally left unused (always NULL): a zeroed or
> incorrectly initialized pool index then resolves to NULL and trips a
> WARN rather than silently aliasing a live pool in another slot.
>
> The array keeps the same RCU publish/retire discipline the list had,
> so lookup and teardown stay equivalent. A fully-constructed pool is
> stored into its slot as the last step of zswap_pool_create(), so array
> walkers only ever observe a NULL slot or a ready pool. Pool creation
> is serialized by the module-wide kernel param mutex (all built-in
> params share one lock) and otherwise only happens during
> single-threaded init, so no two creators race for a slot.
> zswap_pools_lock still serializes the store against a retiring pool
> clearing its slot in __zswap_pool_empty().
>
> Behavior change: the fixed array bounds the number of simultaneously
> live pools at ZSWAP_MAX_POOLS - 1 (15, since slot 0 is reserved),
> whereas the old list was unbounded. A pool is only live while it is
> the current pool or still has stored pages referencing it, and pools
> are reused across compressor switches, so 15 is far more than any real
> configuration needs. Once all slots are occupied, creating a pool for
> a 16th distinct compressor fails: zswap_pool_create() errors and
> returns NULL, and the compressor switch is rejected with -EINVAL
> rather than silently succeeding. The cap can be raised by increasing
> ZSWAP_MAX_POOLS (bounded by the u8 slot index, so up to 256).
>
> Suggested-by: Nhat Pham <nphamcs@gmail.com>
> Suggested-by: Yosry Ahmed <yosry@kernel.org>
> Signed-off-by: Jianyue Wu <wujianyue000@gmail.com>
Acked-by: Yosry Ahmed <yosry@kernel.org>
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH v4 3/3] mm/zswap: reference the pool by index to shrink struct zswap_entry
2026-08-30 11:47 ` [RFC PATCH v4 3/3] mm/zswap: reference the pool by index to shrink struct zswap_entry Jianyue Wu
@ 2026-08-31 15:30 ` Yosry Ahmed
0 siblings, 0 replies; 12+ messages in thread
From: Yosry Ahmed @ 2026-08-31 15:30 UTC (permalink / raw)
To: Jianyue Wu
Cc: Johannes Weiner, Nhat Pham, Chengming Zhou, Andrew Morton,
Chris Li, linux-mm, linux-kernel
On Sun, Aug 30, 2026 at 4:48 AM Jianyue Wu <wujianyue000@gmail.com> wrote:
>
> struct zswap_entry is one allocation per stored page, so its size is
> pure overhead. It currently embeds an 8-byte pool pointer, even though
> the live pools now sit in a small fixed array indexed by a u8 slot
> number.
>
> Replace the per-entry pool pointer with that u8 slot index and resolve
> it through the fixed pool array. A live entry holds a reference to its
> pool, so the slot cannot be reused under it. The lookup therefore needs
> no RCU read-side section or zswap_pools_lock.
>
> The u8 fits in the padding after the bool referenced field, shrinking
> the entry from 56 to 48 bytes on x86_64. This raises objs_per_slab from
> 73 to 85 and saves about 2MiB of metadata per 1GiB of data held in
> zswap.
>
> Suggested-by: Chris Li <chrisl@kernel.org>
> Signed-off-by: Jianyue Wu <wujianyue000@gmail.com>
Acked-by: Yosry Ahmed <yosry@kernel.org>
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH v4 1/3] mm/zswap: release retired pools via queue_rcu_work() instead of synchronize_rcu()
2026-08-31 15:20 ` Yosry Ahmed
@ 2026-09-01 14:33 ` Jianyue Wu
0 siblings, 0 replies; 12+ messages in thread
From: Jianyue Wu @ 2026-09-01 14:33 UTC (permalink / raw)
To: Yosry Ahmed
Cc: Johannes Weiner, Nhat Pham, Chengming Zhou, Andrew Morton,
Chris Li, linux-mm, linux-kernel
On Mon, Aug 31, 2026 at 11:20 PM Yosry Ahmed <yosry@kernel.org> wrote:
>
> On Sun, Aug 30, 2026 at 4:47 AM Jianyue Wu <wujianyue000@gmail.com> wrote:
> >
> > When a pool's last reference is dropped, __zswap_pool_empty() removes it
> > from the pool list and schedules __zswap_pool_release(), which calls
> > synchronize_rcu() to wait for readers before tearing the pool down.
> >
> > synchronize_rcu() is a synchronous, potentially long wait. Replace it
> > with queue_rcu_work(): __zswap_pool_empty() hands the pool to
> > queue_rcu_work(), which waits for a grace period asynchronously and then
> > runs __zswap_pool_release() from a worker for the sleepable teardown
> > (__zswap_pool_empty() can run in atomic context and must not block).
> > The grace-period guarantee is unchanged; the retirement path just no
> > longer blocks on it.
> >
> > Suggested-by: Yosry Ahmed <yosry@kernel.org>
> > Signed-off-by: Jianyue Wu <wujianyue000@gmail.com>
>
> Why are the patches still tagged RFC?
Exactly, should be removed, I dropped the RFC tag in the new version.
> Anyway, with one nit below:
>
> Acked-by: Yosry Ahmed <yosry@kernel.org>
>
> > ---
> > mm/zswap.c | 12 +++++-------
> > 1 file changed, 5 insertions(+), 7 deletions(-)
> >
> > diff --git a/mm/zswap.c b/mm/zswap.c
> > index 37f34e406c8e..0bb30e58950a 100644
> > --- a/mm/zswap.c
> > +++ b/mm/zswap.c
> > @@ -155,7 +155,7 @@ struct zswap_pool {
> > struct crypto_acomp_ctx __percpu *acomp_ctx;
> > struct percpu_ref ref;
> > struct list_head list;
> > - struct work_struct release_work;
> > + struct rcu_work release_work;
>
> Seems like the convention is to use rwork in the name instead of work.
Good point, I renamed it to release_rwork in the new version.
Best regards,
Jianyue
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH v4 1/3] mm/zswap: release retired pools via queue_rcu_work() instead of synchronize_rcu()
2026-08-30 11:47 ` [RFC PATCH v4 1/3] mm/zswap: release retired pools via queue_rcu_work() instead of synchronize_rcu() Jianyue Wu
2026-08-31 15:20 ` Yosry Ahmed
@ 2026-09-01 15:38 ` Johannes Weiner
2026-09-02 0:53 ` Jianyue Wu
1 sibling, 1 reply; 12+ messages in thread
From: Johannes Weiner @ 2026-09-01 15:38 UTC (permalink / raw)
To: Jianyue Wu
Cc: Yosry Ahmed, Nhat Pham, Chengming Zhou, Andrew Morton, Chris Li,
linux-mm, linux-kernel
On Sun, Aug 30, 2026 at 07:47:29PM +0800, Jianyue Wu wrote:
> When a pool's last reference is dropped, __zswap_pool_empty() removes it
> from the pool list and schedules __zswap_pool_release(), which calls
> synchronize_rcu() to wait for readers before tearing the pool down.
>
> synchronize_rcu() is a synchronous, potentially long wait. Replace it
> with queue_rcu_work(): __zswap_pool_empty() hands the pool to
> queue_rcu_work(), which waits for a grace period asynchronously and then
> runs __zswap_pool_release() from a worker for the sleepable teardown
> (__zswap_pool_empty() can run in atomic context and must not block).
> The grace-period guarantee is unchanged; the retirement path just no
> longer blocks on it.
>
> Suggested-by: Yosry Ahmed <yosry@kernel.org>
> Signed-off-by: Jianyue Wu <wujianyue000@gmail.com>
With "release_rwork",
Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH v4 2/3] mm/zswap: replace the zswap_pools list with a fixed pools array
2026-08-30 11:47 ` [RFC PATCH v4 2/3] mm/zswap: replace the zswap_pools list with a fixed pools array Jianyue Wu
2026-08-31 15:28 ` Yosry Ahmed
@ 2026-09-01 16:13 ` Johannes Weiner
2026-09-02 0:50 ` Jianyue Wu
1 sibling, 1 reply; 12+ messages in thread
From: Johannes Weiner @ 2026-09-01 16:13 UTC (permalink / raw)
To: Jianyue Wu
Cc: Yosry Ahmed, Nhat Pham, Chengming Zhou, Andrew Morton, Chris Li,
linux-mm, linux-kernel
On Sun, Aug 30, 2026 at 07:47:30PM +0800, Jianyue Wu wrote:
> Originally zswap holds its pools on an RCU list whose head also serves
> as the "current pool". Only a handful of pools are ever live at once,
> since a new pool is only created when the compressor is (re)set and
> pools are reused across compressor switches.
>
> Hold the pools in a fixed ZSWAP_MAX_POOLS-element array so each pool
> has a stable slot number, and track the current pool with a separate
> rcu-protected pointer.
>
> Slot 0 is intentionally left unused (always NULL): a zeroed or
> incorrectly initialized pool index then resolves to NULL and trips a
> WARN rather than silently aliasing a live pool in another slot.
>
> The array keeps the same RCU publish/retire discipline the list had,
> so lookup and teardown stay equivalent. A fully-constructed pool is
> stored into its slot as the last step of zswap_pool_create(), so array
> walkers only ever observe a NULL slot or a ready pool. Pool creation
> is serialized by the module-wide kernel param mutex (all built-in
> params share one lock) and otherwise only happens during
> single-threaded init, so no two creators race for a slot.
> zswap_pools_lock still serializes the store against a retiring pool
> clearing its slot in __zswap_pool_empty().
>
> Behavior change: the fixed array bounds the number of simultaneously
> live pools at ZSWAP_MAX_POOLS - 1 (15, since slot 0 is reserved),
> whereas the old list was unbounded. A pool is only live while it is
> the current pool or still has stored pages referencing it, and pools
> are reused across compressor switches, so 15 is far more than any real
> configuration needs. Once all slots are occupied, creating a pool for
> a 16th distinct compressor fails: zswap_pool_create() errors and
> returns NULL, and the compressor switch is rejected with -EINVAL
> rather than silently succeeding. The cap can be raised by increasing
> ZSWAP_MAX_POOLS (bounded by the u8 slot index, so up to 256).
>
> Suggested-by: Nhat Pham <nphamcs@gmail.com>
> Suggested-by: Yosry Ahmed <yosry@kernel.org>
> Signed-off-by: Jianyue Wu <wujianyue000@gmail.com>
> ---
> mm/zswap.c | 97 ++++++++++++++++++++++++++++++++++++++++--------------
> 1 file changed, 72 insertions(+), 25 deletions(-)
>
> diff --git a/mm/zswap.c b/mm/zswap.c
> index 0bb30e58950a..b3b5e2887c00 100644
> --- a/mm/zswap.c
> +++ b/mm/zswap.c
> @@ -13,6 +13,7 @@
>
> #define pr_fmt(fmt) KBUILD_MODNAME ": " fmt
>
> +#include <linux/cleanup.h>
> #include <linux/module.h>
> #include <linux/cpu.h>
> #include <linux/highmem.h>
> @@ -154,12 +155,27 @@ struct zswap_pool {
> struct zs_pool *zs_pool;
> struct crypto_acomp_ctx __percpu *acomp_ctx;
> struct percpu_ref ref;
> - struct list_head list;
> struct rcu_work release_work;
> struct hlist_node node;
> + u8 idx;
> char tfm_name[CRYPTO_MAX_ALG_NAME];
> };
>
> +#define ZSWAP_MAX_POOLS 16
It's unlikely to happen, but this is a super annoying failure
mode. User would have to kill something, delete shmem/tmpfs, or
swapoff. And it's not obvious which entries are in which pool.
Wouldn't an idr make more sense?
> @@ -270,6 +283,31 @@ static void acomp_ctx_free(struct crypto_acomp_ctx *acomp_ctx)
> acomp_ctx->buffer = NULL;
> }
>
> +/*
> + * Publish a fully-constructed pool into a free array slot. Pool creation is
> + * serialized by the module-wide kernel param mutex (all built-in params share
> + * one lock) and only otherwise happens during single-threaded init, so no two
> + * creators race for a slot. The pool is complete before it is stored, and
> + * zswap_pools_lock still serializes this store against a concurrent retiring
> + * pool clearing its slot in __zswap_pool_empty(), so array walkers only ever
> + * observe a NULL slot or a ready pool.
> + */
> +static int zswap_pool_assign_slot(struct zswap_pool *pool)
> +{
> + int i;
> +
> + guard(spinlock_bh)(&zswap_pools_lock);
> + for (i = ZSWAP_FIRST_POOL_SLOT; i < ZSWAP_MAX_POOLS; i++) {
> + if (!rcu_access_pointer(zswap_pools[i])) {
> + pool->idx = i;
> + rcu_assign_pointer(zswap_pools[i], pool);
> + return i;
> + }
> + }
> +
> + return -ENOSPC;
> +}
It was kind of overdue, but with this now requiring a pool walk as
well, it would be better to factor out a find_or_create function?
Something like:
static struct zswap_pool *zswap_pool_find_or_create(char *compressor)
{
struct zswap_pool *pool, *new_pool = NULL;
u8 id, new_id = 0;
insert_new:
spin_lock_bh(&zswap_pools_lock);
idr_for_each_entry(&zswap_pools, pool, id) {
if (pool && !strcmp(pool->tfm_name, compressor) && zswap_pool_tryget(pool)) {
if (new_pool) {
pool_put(new_pool);
idr_free(&zswap_pools, new_id);
}
spin_unlock_bh(&zswap_pools_lock);
return pool;
}
}
if (new_pool) {
idr_replace(&zswap_pools, new_pool, new_id);
spin_unlock_bh(&zswap_pools_lock);
return new_pool;
}
spin_unlock_bh(&zswap_pools_lock);
new_pool = pool_alloc();
if (!new_pool)
...
new_id = idr_alloc(&zswap_pools, NULL, 1, 256, GFP_KERNEL);
if (new_id < 0)
...
goto insert_new;
}
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH v4 2/3] mm/zswap: replace the zswap_pools list with a fixed pools array
2026-09-01 16:13 ` Johannes Weiner
@ 2026-09-02 0:50 ` Jianyue Wu
0 siblings, 0 replies; 12+ messages in thread
From: Jianyue Wu @ 2026-09-02 0:50 UTC (permalink / raw)
To: Johannes Weiner
Cc: Yosry Ahmed, Nhat Pham, Chengming Zhou, Andrew Morton, Chris Li,
linux-mm, linux-kernel
On Wed, Sep 2, 2026 at 12:13 AM Johannes Weiner <hannes@cmpxchg.org> wrote:
>
> On Sun, Aug 30, 2026 at 07:47:30PM +0800, Jianyue Wu wrote:
> > Originally zswap holds its pools on an RCU list whose head also serves
> > as the "current pool". Only a handful of pools are ever live at once,
> > since a new pool is only created when the compressor is (re)set and
> > pools are reused across compressor switches.
> >
> > Hold the pools in a fixed ZSWAP_MAX_POOLS-element array so each pool
> > has a stable slot number, and track the current pool with a separate
> > rcu-protected pointer.
> >
> > Slot 0 is intentionally left unused (always NULL): a zeroed or
> > incorrectly initialized pool index then resolves to NULL and trips a
> > WARN rather than silently aliasing a live pool in another slot.
> >
> > The array keeps the same RCU publish/retire discipline the list had,
> > so lookup and teardown stay equivalent. A fully-constructed pool is
> > stored into its slot as the last step of zswap_pool_create(), so array
> > walkers only ever observe a NULL slot or a ready pool. Pool creation
> > is serialized by the module-wide kernel param mutex (all built-in
> > params share one lock) and otherwise only happens during
> > single-threaded init, so no two creators race for a slot.
> > zswap_pools_lock still serializes the store against a retiring pool
> > clearing its slot in __zswap_pool_empty().
> >
> > Behavior change: the fixed array bounds the number of simultaneously
> > live pools at ZSWAP_MAX_POOLS - 1 (15, since slot 0 is reserved),
> > whereas the old list was unbounded. A pool is only live while it is
> > the current pool or still has stored pages referencing it, and pools
> > are reused across compressor switches, so 15 is far more than any real
> > configuration needs. Once all slots are occupied, creating a pool for
> > a 16th distinct compressor fails: zswap_pool_create() errors and
> > returns NULL, and the compressor switch is rejected with -EINVAL
> > rather than silently succeeding. The cap can be raised by increasing
> > ZSWAP_MAX_POOLS (bounded by the u8 slot index, so up to 256).
> >
> > Suggested-by: Nhat Pham <nphamcs@gmail.com>
> > Suggested-by: Yosry Ahmed <yosry@kernel.org>
> > Signed-off-by: Jianyue Wu <wujianyue000@gmail.com>
> > ---
> > mm/zswap.c | 97 ++++++++++++++++++++++++++++++++++++++++--------------
> > 1 file changed, 72 insertions(+), 25 deletions(-)
> >
> > diff --git a/mm/zswap.c b/mm/zswap.c
> > index 0bb30e58950a..b3b5e2887c00 100644
> > --- a/mm/zswap.c
> > +++ b/mm/zswap.c
> > @@ -13,6 +13,7 @@
> >
> > #define pr_fmt(fmt) KBUILD_MODNAME ": " fmt
> >
> > +#include <linux/cleanup.h>
> > #include <linux/module.h>
> > #include <linux/cpu.h>
> > #include <linux/highmem.h>
> > @@ -154,12 +155,27 @@ struct zswap_pool {
> > struct zs_pool *zs_pool;
> > struct crypto_acomp_ctx __percpu *acomp_ctx;
> > struct percpu_ref ref;
> > - struct list_head list;
> > struct rcu_work release_work;
> > struct hlist_node node;
> > + u8 idx;
> > char tfm_name[CRYPTO_MAX_ALG_NAME];
> > };
> >
> > +#define ZSWAP_MAX_POOLS 16
>
> It's unlikely to happen, but this is a super annoying failure
> mode. User would have to kill something, delete shmem/tmpfs, or
> swapoff. And it's not obvious which entries are in which pool.
>
> Wouldn't an idr make more sense?
>
> > @@ -270,6 +283,31 @@ static void acomp_ctx_free(struct crypto_acomp_ctx *acomp_ctx)
> > acomp_ctx->buffer = NULL;
> > }
> >
> > +/*
> > + * Publish a fully-constructed pool into a free array slot. Pool creation is
> > + * serialized by the module-wide kernel param mutex (all built-in params share
> > + * one lock) and only otherwise happens during single-threaded init, so no two
> > + * creators race for a slot. The pool is complete before it is stored, and
> > + * zswap_pools_lock still serializes this store against a concurrent retiring
> > + * pool clearing its slot in __zswap_pool_empty(), so array walkers only ever
> > + * observe a NULL slot or a ready pool.
> > + */
> > +static int zswap_pool_assign_slot(struct zswap_pool *pool)
> > +{
> > + int i;
> > +
> > + guard(spinlock_bh)(&zswap_pools_lock);
> > + for (i = ZSWAP_FIRST_POOL_SLOT; i < ZSWAP_MAX_POOLS; i++) {
> > + if (!rcu_access_pointer(zswap_pools[i])) {
> > + pool->idx = i;
> > + rcu_assign_pointer(zswap_pools[i], pool);
> > + return i;
> > + }
> > + }
> > +
> > + return -ENOSPC;
> > +}
>
> It was kind of overdue, but with this now requiring a pool walk as
> well, it would be better to factor out a find_or_create function?
>
> Something like:
>
> static struct zswap_pool *zswap_pool_find_or_create(char *compressor)
> {
> struct zswap_pool *pool, *new_pool = NULL;
> u8 id, new_id = 0;
>
> insert_new:
> spin_lock_bh(&zswap_pools_lock);
> idr_for_each_entry(&zswap_pools, pool, id) {
> if (pool && !strcmp(pool->tfm_name, compressor) && zswap_pool_tryget(pool)) {
> if (new_pool) {
> pool_put(new_pool);
> idr_free(&zswap_pools, new_id);
> }
> spin_unlock_bh(&zswap_pools_lock);
> return pool;
> }
> }
> if (new_pool) {
> idr_replace(&zswap_pools, new_pool, new_id);
> spin_unlock_bh(&zswap_pools_lock);
> return new_pool;
> }
> spin_unlock_bh(&zswap_pools_lock);
>
> new_pool = pool_alloc();
> if (!new_pool)
> ...
> new_id = idr_alloc(&zswap_pools, NULL, 1, 256, GFP_KERNEL);
> if (new_id < 0)
> ...
> goto insert_new;
> }
Agreed. Hitting the 15-pool cap would be an annoying failure mode, even
if it should be rare.
I will switch this to an IDR with IDs in the 1..255 range, keeping 0 as
the invalid entry value, and factor the lookup/allocation path into a
zswap_pool_find_or_create() helper.
Best regards,
Jianyue
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [RFC PATCH v4 1/3] mm/zswap: release retired pools via queue_rcu_work() instead of synchronize_rcu()
2026-09-01 15:38 ` Johannes Weiner
@ 2026-09-02 0:53 ` Jianyue Wu
0 siblings, 0 replies; 12+ messages in thread
From: Jianyue Wu @ 2026-09-02 0:53 UTC (permalink / raw)
To: Johannes Weiner
Cc: Yosry Ahmed, Nhat Pham, Chengming Zhou, Andrew Morton, Chris Li,
linux-mm, linux-kernel
On Tue, Sep 1, 2026 at 11:38 PM Johannes Weiner <hannes@cmpxchg.org> wrote:
>
> On Sun, Aug 30, 2026 at 07:47:29PM +0800, Jianyue Wu wrote:
> > When a pool's last reference is dropped, __zswap_pool_empty() removes it
> > from the pool list and schedules __zswap_pool_release(), which calls
> > synchronize_rcu() to wait for readers before tearing the pool down.
> >
> > synchronize_rcu() is a synchronous, potentially long wait. Replace it
> > with queue_rcu_work(): __zswap_pool_empty() hands the pool to
> > queue_rcu_work(), which waits for a grace period asynchronously and then
> > runs __zswap_pool_release() from a worker for the sleepable teardown
> > (__zswap_pool_empty() can run in atomic context and must not block).
> > The grace-period guarantee is unchanged; the retirement path just no
> > longer blocks on it.
> >
> > Suggested-by: Yosry Ahmed <yosry@kernel.org>
> > Signed-off-by: Jianyue Wu <wujianyue000@gmail.com>
>
> With "release_rwork",
>
> Reviewed-by: Johannes Weiner <hannes@cmpxchg.org>
Thanks, I will use release_rwork instead.
Best regards,
Jianyue
^ permalink raw reply [flat|nested] 12+ messages in thread
end of thread, other threads:[~2026-09-02 0:54 UTC | newest]
Thread overview: 12+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-30 11:47 [RFC PATCH v4 0/3] mm/zswap: shrink zswap_entry via a fixed pool index Jianyue Wu
2026-08-30 11:47 ` [RFC PATCH v4 1/3] mm/zswap: release retired pools via queue_rcu_work() instead of synchronize_rcu() Jianyue Wu
2026-08-31 15:20 ` Yosry Ahmed
2026-09-01 14:33 ` Jianyue Wu
2026-09-01 15:38 ` Johannes Weiner
2026-09-02 0:53 ` Jianyue Wu
2026-08-30 11:47 ` [RFC PATCH v4 2/3] mm/zswap: replace the zswap_pools list with a fixed pools array Jianyue Wu
2026-08-31 15:28 ` Yosry Ahmed
2026-09-01 16:13 ` Johannes Weiner
2026-09-02 0:50 ` Jianyue Wu
2026-08-30 11:47 ` [RFC PATCH v4 3/3] mm/zswap: reference the pool by index to shrink struct zswap_entry Jianyue Wu
2026-08-31 15:30 ` Yosry Ahmed
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox