* [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I)
@ 2026-10-03 0:31 Baoquan He
2026-10-03 0:31 ` [PATCH v4 01/14] mm: xswap support for zswap Baoquan He
` (14 more replies)
0 siblings, 15 replies; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
A normal swap device has a fixed size, set when it is enabled. This
does not fit zswap well. With zswap, the swapped data stays in RAM in
compressed form. So how much swap a machine can really use depends on
how well the workload compresses. It does not depend on a number we
pick before we know the workload. If we make the device large enough
for the worst case, most of it is never used. If we make it just large
enough for the average case, the machine runs out of swap while zswap
still has free room.
xswap is a swap device with no backing file. Its swap entries hold pages
in zswap, and its cluster_info array lives in a VM_SPARSE area that is
mapped a chunk at a time, so the device covers a large address space
while costing only the chunks it has actually used. The device grows as
the workload needs it and gives the tail back when it does not.
This series is the device itself: the sparse cluster array, growth and
shrink, and the interfaces that create and destroy one. It is phase I
of three; the physical backend and the swap accounting builds on it.
Design notes
------------
The device is only an address space. si->max is set when the device is
created and does not move for its lifetime; what moves is how much of it
is mapped. It defaults to twice RAM, and xswap.max= sets another size at
boot, in bytes or as a percentage of RAM:
xswap.max=8G 8 GiB
xswap.max=300% three times RAM
That is a boot parameter but not a runtime knob because the address space
cannot move once the device exists; it is fixed at creation and every
mapped index stays valid for the device's life. swap_info_struct carries
nr_clusters_max (the address space) and nr_clusters_mapped (the mapped
prefix), and the two walkers that touch cluster_info are bounded by the
latter.
Growth starts when 85% of the mapped range is in use. The call comes from
cluster_alloc_swap_entry(). Growing maps pages into the VM_SPARSE area.
Mapping can sleep, so it cannot be done inside the allocation that would
otherwise fail. If it were, the swap-out path would have nothing to fall
back on.
Shrink gives the free tail back when the mapped range is at most 50% in
use. It is triggered when a cluster becomes completely free. Shrinking on
a smaller drop is not useful. Growth happens on demand, so the same
clusters would just be mapped again, and each unmap also costs an RCU
grace period.
A file-less device has no swapon, so /sys/kernel/mm/xswap/create makes
one and destroy takes its swap type. The teardown is shared with
sys_swapoff() rather than duplicated.
The only argument create takes is a priority. The size comes from
xswap.max= at boot, so there is one place to ask for a size and one to
ask for a priority, and a priority can be given per device while a size
cannot:
# echo > /sys/kernel/mm/xswap/create default priority
# echo 100 > /sys/kernel/mm/xswap/create priority 100
The default is the highest priority, SWAP_FLAG_PRIO_MASK. An xswap
device is then the first one tried, ahead of any ordinary swap device.
Setting a lower priority is allowed.
The device does not move once it exists, so the address space cannot be
changed from this interface either. xswap.max= at boot and this together
are the whole configuration.
The device needs zswap. Without it swapout takes a swap entry and frees
no memory, so creating one fails.
Testing
-------
qemu KVM guest, 8G RAM, booted with xswap.max=10G. memhog is a small
local helper: it faults <total_gb> of anon, fills it with a fixed
pattern, and holds it.
Set up zswap and create a device:
# echo 1 > /sys/module/zswap/parameters/enabled
# echo > /sys/kernel/mm/xswap/create
# awk 'NR == 1 || /xswap/' /proc/swaps
Filename Type Size Used Priority
xswap0 xswap 10485756 0 32767
# cat /sys/kernel/debug/xswap/type0
clusters_max 5120
clusters_mapped 73
usage_pages 0
tail_free 72
grows 0
shrinks 0
Swap 9G of anon through a 2G cgroup:
# mkdir -p /sys/fs/cgroup/xswap_limit
# echo 2G > /sys/fs/cgroup/xswap_limit/memory.max
# echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
# ( echo $BASHPID > /sys/fs/cgroup/xswap_limit/cgroup.procs
# exec env MEMHOG_FILL=pattern numactl --cpunodebind=0 \
# ./memhog 9 600 ) &
clusters_mapped grows to cover the pages that go to zswap, while
SwapTotal stays at 10485756. Throughout, SwapTotal never moves and
SwapFree stays in [0, SwapTotal]. Pushing past the size leaves the
cgroup out of room and the OOM killer takes the workload, which is the
pass signal there.
Growth and shrink, measured with two cgroups so that one keeps its pages
while the other goes away:
idle mapped=73 occ=0.000 grows=0 shrinks=0
2G in a 1G cgroup mapped=803 occ=0.737 grows=8 shrinks=0
3G more in a second 1G cgroup mapped=2190 occ=0.777 grows=13 shrinks=0
kill the second cgroup mapped=876 occ=0.675 grows=13 shrinks=1
The second cgroup is started later, so its clusters sit at the tail, and
killing it leaves a free tail to take back. The first stays alive, so
usage falls but not to zero: that is the state the shrink has to land in.
occ is usage over the mapped range, which is what the grow and shrink
thresholds are expressed in.
# echo 0 > /sys/kernel/mm/xswap/destroy
Also tested things as below, and nothing crashed or warned:
- a shrink racing a swapoff of the same device, 20 rounds
- a cgroup that cannot zswap: it must get no xswap slot, and the
allocator walk must not spin
- a create/destroy cycle: it must not leak per-cluster state
Open items
----------
- xswap_create() sets si->pages to the whole address space (twice RAM by
default) and _enable_swap_info() adds it to nr_swap_pages and
total_swap_pages, which is what __vm_enough_memory() uses for
overcommit. That raises the commit limit by 2xRAM for a device whose
real capacity is bounded by the zswap pool and by how well the workload
compresses. The question is which semantics overcommit should follow:
worst-case backing, or the address space.
- destroy takes a bare swap type. The type is an internal index that
alloc_swap_info() recycles as soon as SWP_USED clears, so a value read
earlier can name a different device by the time it is written. The
swapoff ABI uses a path for this reason. It wants a generation or
another stable handle.
- The new interfaces (/sys/kernel/mm/xswap/create and destroy, xswap.max=,
the new type string in /proc/swaps) are not documented yet.
Changelog
=========
v3 -> v4:
- Rebased onto the latest mm-new.
- Add boot parameter xswap.max= .
- The device has one size now: set at creation, fixed, with xswap.max= to
ask for another one at boot. v3's runtime ceiling, the per-device limit,
the clamp it wrote and the shrink to that ceiling are all gone as
Johannes suggested (v3's patches 12, 13 and 14).
- v3's patches 2 and 4 are the new patch 3: neither builds a device
alone. It also bounds find_next_to_unuse() and wait_for_allocation(),
which walked up to si->max past the mapped range. Chris suggested
this.
- New patches 11 to 13: a cgroup that cannot zswap is not given an xswap
slot, swap_info_struct max and pages are unsigned long, and debugfs
counters for the mapped range and the grow and shrink counts.
- v3's patches 1 and 3, and 5 to 11, are unchanged; the merge shifts them
to 1 and 2, and 4 to 10.
- Minor comment and cleanup changes.
v2 -> v3:
- Rebased onto the latest mm-new.
- The grow path now honors the user-set ceiling (si->nr_clusters) instead
of growing up to nr_clusters_max, and a ceiling below the mapped range
is unmapped exactly instead of rounded to a chunk (patches 12 and 14).
- The limit write clamps the ceiling up to the clusters covering the pages
in use, replacing the earlier WARN_ONCE; si->pages becomes mutable at
runtime (patch 13).
- Minor comment and cleanup changes.
v1->v2:
- Patch 1 (mm: zswap: return -ENOENT when the swap device is gone) is not
part of this series; it was posted separately.
- There is only one size knob now. The runtime ceiling and the debugfs
per-device limit are gone. All that is left is the optional per-device
cap, /sys/kernel/mm/xswap/type<N>/limit. Grow and shrink work without
it.
- The shrink no longer keeps its own count of the free tail. It scans the
tail instead, and dropping the counter also removes a call from the
cluster allocation path.
- The priority is no longer a patch of its own. The create attribute
takes it:
echo 100 > /sys/kernel/mm/xswap/create
RFC v3 -> RFC v2
- Add patch 16 to support setting xswap device priority at creation.
The create sysfs interface (/sys/kernel/mm/xswap/create) previously
hardcoded every new device's priority to DEF_SWAP_PRIO, it now
accepts an optional priority:
echo "<percent> [<prio>]" > /sys/kernel/mm/xswap/create
- Bug fix: xswap_lock init ordering. mutex_init(&si->xswap_lock) was called
after xswap_map_clusters() (which locks it), i.e. locking an uninitialized
mutex. Init now before the first xswap_map_clusters() call. Thanks to Klara.
- Bug fix: Fixes a compile error in !CONFIG_XSWAP builds. xswap_debugfs_root
is declared inside CONFIG_XSWAP ifdeffery scope, so the ungarded use
caused error when CONFIG_XSWAP is off.
RFC v2-> RFC v3:
- Replace the "header-only swap file + swapon" creation hack with a
proper file-less device created and destroyed via sysfs
(/sys/kernel/mm/xswap/{create,destroy}). This required the
__swapoff() refactor and the free_swap_cluster_info() signature
change (patches 4, 6, 14).
- Require zswap: refuse to create an xswap device when zswap is
unavailable (patch 15).
- Split the unrelated zswap -ENOENT fix out of the series into a
standalone patch (patch 1).
- Fix nr_free_tail over-counting on concurrent grow, shrink leaking
detached clusters on early bail-out, a re-init race on cluster
spinlocks in xswap_map_clusters(), the nr_clusters_mapped update
ordering, and swapoff accessing the shrinker-unmapped cluster tail.
- Minor cleanups (checkpatch, /proc/swaps alignment, commit messages).
RFC v1-> RFC v2:
- Added __GFP_HIGH | __GFP_NOMEMALLOC to alloc_page() and kmalloc_array()
in the grow path, plus memalloc_noreclaim_save/restore() wrapping,
to prevent the grow path from consuming emergency memory reserves
or recursing into swap under PF_MEMALLOC. This is folded into patch 3.
This was pointed out by Nhat.
- Folded the mutex serialization fix into the cluster grow patch (patch
3). This is suggested by Nhat.
- Fixed coding style issues: corrected indentation of declarations in
xswap_unmap_clusters(), removed unnecessary block scope around the
err variable in xswap_map_clusters().
- Rebased onto mm-unstable
Baoquan He (13):
mm, swap: refactor free_swap_cluster_info to take swap_info_struct
mm, swap: back the cluster_info array with a sparse VM_SPARSE area
mm, swap: add sysfs create interface for xswap
mm, swap: add xswap grow trigger on cluster allocation
mm, swap: add xswap_try_shrink and shrink trigger on cluster free
mm, swap: free backing pages in xswap_unmap_clusters
mm, swap: defer xswap shrink to workqueue to avoid lock recursion
mm, swap: refactor swapoff and add xswap_destroy
mm, swap: require zswap for xswap devices
mm, swap: do not give an xswap slot to a cgroup that cannot zswap
mm, swap: widen swap_info_struct max/pages to unsigned long
mm, swap: add debugfs counters for xswap
mm, swap: let the command line set the xswap device size
Chris Li (1):
mm: xswap support for zswap
include/linux/swap.h | 16 +-
mm/Kconfig | 9 +
mm/page_io.c | 19 +
mm/swap_state.c | 9 +
mm/swapfile.c | 1221 +++++++++++++++++++++++++++++++++++++++---
mm/zswap.c | 11 +-
6 files changed, 1205 insertions(+), 80 deletions(-)
base-commit: 9a0542b19a541eda3f82934ee747c61ec3f18847
--
2.54.0
^ permalink raw reply [flat|nested] 21+ messages in thread
* [PATCH v4 01/14] mm: xswap support for zswap
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-07 4:22 ` KunWu Chan
2026-10-03 0:31 ` [PATCH v4 02/14] mm, swap: refactor free_swap_cluster_info to take swap_info_struct Baoquan He
` (13 subsequent siblings)
14 siblings, 1 reply; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
From: Chris Li <chrisl@kernel.org>
Introduce extendable swap device support - xswap.
An xswap device has no backing storage and no swap data section, so
it wastes no disk space. Creation is via a sysfs interface added in a
later patch.
Zswap writeback is gated on whether a real (non-xswap) swap device is
active. nr_real_swapfiles counts such devices and is maintained at
swapon/swapoff only, so the gate reflects "a device exists to write
back to" rather than "a device currently has free slots". This keeps
writeback working even when the real swap device is full, and avoids
a double decrement when a full device is swapped off.
An xswap entry is not added to the zswap writeback LRU: with no backing
store there is nothing to write it back to. zswap_lru_del() tolerates
that, since an entry that was never added is on no list.
Co-developed-by: Baoquan He <hebaoquan@kylinos.cn>
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
Signed-off-by: Chris Li <chrisl@kernel.org>
---
include/linux/swap.h | 2 ++
mm/page_io.c | 19 +++++++++++++++++++
mm/swap_state.c | 9 +++++++++
mm/swapfile.c | 28 ++++++++++++++++++++++++----
mm/zswap.c | 11 +++++++++--
5 files changed, 63 insertions(+), 6 deletions(-)
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 4f686709d71f..9073d29377d8 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -201,6 +201,7 @@ enum {
SWP_STABLE_WRITES = (1 << 11), /* no overwrite PG_writeback pages */
SWP_SYNCHRONOUS_IO = (1 << 12), /* synchronous IO is efficient */
SWP_HIBERNATION = (1 << 13), /* pinned for hibernation */
+ SWP_XSWAP = (1 << 14), /* extendable swap device */
/* add others here before... */
};
@@ -378,6 +379,7 @@ void free_folio_and_swap_cache(struct folio *folio);
void free_pages_and_swap_cache(struct encoded_page **, int);
/* linux/mm/swapfile.c */
extern atomic_long_t nr_swap_pages;
+extern atomic_t nr_real_swapfiles;
extern long total_swap_pages;
extern atomic_t nr_rotate_swap;
diff --git a/mm/page_io.c b/mm/page_io.c
index c6824fcd483e..25fa9b82ed46 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -248,6 +248,15 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
}
rcu_read_unlock();
+ /*
+ * xswap has no backing store: keep the folio. ctx->sis is not set
+ * yet, so look the device up from the entry.
+ */
+ if (unlikely(__swap_entry_to_info(folio->swap)->flags & SWP_XSWAP)) {
+ folio_mark_dirty(folio);
+ return AOP_WRITEPAGE_ACTIVATE;
+ }
+
__swap_writeout(ctx, folio);
return 0;
out_unlock:
@@ -486,6 +495,16 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct folio *folio)
if (zswap_load(folio) != -ENOENT)
goto finish;
+ if (unlikely(sis->flags & SWP_XSWAP)) {
+ /*
+ * An xswap entry only ever lives in zswap, so zswap_load()
+ * must have found it. Unlock and let the caller retry.
+ */
+ WARN_ON_ONCE(1);
+ folio_unlock(folio);
+ goto finish;
+ }
+
/* We have to read from slower devices. Increase zswap protection. */
zswap_folio_swapin(folio);
swap_add_folio(ctx, folio, READ);
diff --git a/mm/swap_state.c b/mm/swap_state.c
index ebef568cd62e..3a4e9e447b0b 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -884,6 +884,10 @@ struct folio *swap_cluster_readahead(swp_entry_t entry, gfp_t gfp_mask,
struct blk_plug plug;
swp_entry_t ra_entry;
+ /* xswap entries live only in zswap; readahead does not help. */
+ if (si->flags & SWP_XSWAP)
+ goto skip;
+
mask = swapin_nr_pages(offset) - 1;
if (!mask)
goto skip;
@@ -969,6 +973,7 @@ static int swap_vma_ra_win(struct vm_fault *vmf, unsigned long *start,
static struct folio *swap_vma_readahead(swp_entry_t targ_entry, gfp_t gfp_mask,
struct mempolicy *mpol, pgoff_t targ_ilx, struct vm_fault *vmf)
{
+ struct swap_info_struct *si = __swap_entry_to_info(targ_entry);
struct swap_io_ctx ctx = {};
struct blk_plug plug;
struct folio *folio;
@@ -977,6 +982,10 @@ static struct folio *swap_vma_readahead(swp_entry_t targ_entry, gfp_t gfp_mask,
unsigned long start, end, addr;
pgoff_t ilx = targ_ilx;
+ /* xswap entries live only in zswap; readahead does not help. */
+ if (si->flags & SWP_XSWAP)
+ goto skip;
+
win = swap_vma_ra_win(vmf, &start, &end);
if (win == 1)
goto skip;
diff --git a/mm/swapfile.c b/mm/swapfile.c
index c3288910b3e3..32c3133e1211 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -66,6 +66,7 @@ static void move_cluster(struct swap_info_struct *si,
static DEFINE_SPINLOCK(swap_lock);
static unsigned int nr_swapfiles;
atomic_long_t nr_swap_pages;
+atomic_t nr_real_swapfiles;
/*
* Some modules use swappable objects and may try to swap them out under
* memory pressure (via the shrinker). Before doing so, they may wish to
@@ -733,7 +734,8 @@ static void free_cluster(struct swap_info_struct *si, struct swap_cluster_info *
/*
* If the swap is discardable, prepare discard the cluster
* instead of free it immediately. The cluster will be freed
- * after discard.
+ * after discard. xswap has no bdev and never sets
+ * SWP_PAGE_DISCARD, so it always takes the free path below.
*/
if ((si->flags & (SWP_WRITEOK | SWP_PAGE_DISCARD)) ==
(SWP_WRITEOK | SWP_PAGE_DISCARD)) {
@@ -1216,6 +1218,9 @@ static void del_from_avail_list(struct swap_info_struct *si, bool swapoff)
*/
lockdep_assert_held(&si->lock);
si->flags &= ~SWP_WRITEOK;
+ /* Count active devices, not merely those on the avail list. */
+ if (!(si->flags & SWP_XSWAP))
+ atomic_sub(1, &nr_real_swapfiles);
atomic_long_or(SWAP_USAGE_OFFLIST_BIT, &si->inuse_pages);
} else {
/*
@@ -1273,6 +1278,8 @@ static void add_to_avail_list(struct swap_info_struct *si, bool swapon)
}
plist_add(&si->avail_list, &swap_avail_head);
+ if (swapon && !(si->flags & SWP_XSWAP))
+ atomic_add(1, &nr_real_swapfiles);
skip:
spin_unlock(&swap_avail_lock);
@@ -3268,7 +3275,8 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
destroy_swap_extents(p, p->swap_file);
- if (!(p->flags & SWP_SOLIDSTATE))
+ if (!(p->flags & SWP_XSWAP) &&
+ !(p->flags & SWP_SOLIDSTATE))
atomic_dec(&nr_rotate_swap);
mutex_lock(&swapon_mutex);
@@ -3378,6 +3386,19 @@ static void swap_stop(struct seq_file *swap, void *v)
mutex_unlock(&swapon_mutex);
}
+static const char *swap_type_str(struct swap_info_struct *si)
+{
+ struct file *file = si->swap_file;
+
+ if (si->flags & SWP_XSWAP)
+ return "xswap\t";
+
+ if (S_ISBLK(file_inode(file)->i_mode))
+ return "partition";
+
+ return "file\t";
+}
+
static int swap_show(struct seq_file *swap, void *v)
{
struct swap_info_struct *si = v;
@@ -3397,8 +3418,7 @@ static int swap_show(struct seq_file *swap, void *v)
len = seq_file_path(swap, file, " \t\n\\");
seq_printf(swap, "%*s%s\t%lu\t%s%lu\t%s%d\n",
len < 40 ? 40 - len : 1, " ",
- S_ISBLK(file_inode(file)->i_mode) ?
- "partition" : "file\t",
+ swap_type_str(si),
bytes, bytes < 10000000 ? "\t" : "",
inuse, inuse < 10000000 ? "\t" : "",
si->prio);
diff --git a/mm/zswap.c b/mm/zswap.c
index ae19e301fced..7fe25f0b157c 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1018,6 +1018,11 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
if (IS_ERR_OR_NULL(si))
return -ENOENT;
+ if (si->flags & SWP_XSWAP) {
+ put_swap_device(si);
+ return -EINVAL;
+ }
+
mpol = get_task_policy(current);
folio = __swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol,
NO_INTERLEAVE_INDEX);
@@ -1511,7 +1516,9 @@ static bool zswap_store_page(struct folio *folio, long index,
entry->referenced = true;
if (entry->length) {
INIT_LIST_HEAD(&entry->lru);
- zswap_lru_add(entry);
+ /* No backing store: nothing to write these back to. */
+ if (!(__swap_entry_to_info(page_swpentry)->flags & SWP_XSWAP))
+ zswap_lru_add(entry);
}
return true;
@@ -1581,7 +1588,7 @@ bool zswap_store(struct folio *folio)
zswap_pool_put(pool);
put_objcg:
obj_cgroup_put(objcg);
- if (!ret && zswap_pool_reached_full)
+ if (!ret && zswap_pool_reached_full && atomic_read(&nr_real_swapfiles))
queue_work(shrink_wq, &zswap_shrink_work);
check_old:
/*
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 02/14] mm, swap: refactor free_swap_cluster_info to take swap_info_struct
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
2026-10-03 0:31 ` [PATCH v4 01/14] mm: xswap support for zswap Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-07 4:32 ` KunWu Chan
2026-10-03 0:31 ` [PATCH v4 03/14] mm, swap: back the cluster_info array with a sparse VM_SPARSE area Baoquan He
` (12 subsequent siblings)
14 siblings, 1 reply; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
Change free_swap_cluster_info() to take struct swap_info_struct* instead
of (cluster_info, maxpages). It now extracts the fields from si and
clears si->cluster_info after freeing to avoid a double free on the
swapon() error path.
The new parameter also lets the xswap path access si->flags in the
function.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
Reviewed-by: Chris Li <chrisl@kernel.org>
---
mm/swapfile.c | 26 ++++++++++++++------------
1 file changed, 14 insertions(+), 12 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 32c3133e1211..13ec5c0f3c2f 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -3144,14 +3144,17 @@ static void wait_for_allocation(struct swap_info_struct *si)
}
}
-static void free_swap_cluster_info(struct swap_cluster_info *cluster_info,
- unsigned long maxpages)
+static void free_swap_cluster_info(struct swap_info_struct *si)
{
+ struct swap_cluster_info *cluster_info = si->cluster_info;
+ unsigned long maxpages = si->max;
struct swap_cluster_info *ci;
- int i, nr_clusters = DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER);
+ int i, nr_clusters;
if (!cluster_info)
return;
+
+ nr_clusters = DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER);
for (i = 0; i < nr_clusters; i++) {
ci = cluster_info + i;
/* Cluster with bad marks count will have a remaining table */
@@ -3163,6 +3166,7 @@ static void free_swap_cluster_info(struct swap_cluster_info *cluster_info,
spin_unlock(&ci->lock);
}
kvfree(cluster_info);
+ si->cluster_info = NULL;
}
/*
@@ -3190,11 +3194,9 @@ static void flush_percpu_swap_cluster(struct swap_info_struct *si)
SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
{
struct swap_info_struct *p = NULL;
- struct swap_cluster_info *cluster_info;
struct file *swap_file, *victim;
struct address_space *mapping;
struct inode *inode;
- unsigned int maxpages;
int err, found = 0;
if (!capable(CAP_SYS_ADMIN))
@@ -3286,10 +3288,6 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
swap_file = p->swap_file;
p->swap_file = NULL;
- maxpages = p->max;
- cluster_info = p->cluster_info;
- p->max = 0;
- p->cluster_info = NULL;
spin_unlock(&p->lock);
spin_unlock(&swap_lock);
arch_swap_invalidate_area(p->type);
@@ -3297,7 +3295,9 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
mutex_unlock(&swapon_mutex);
kfree(p->global_cluster);
p->global_cluster = NULL;
- free_swap_cluster_info(cluster_info, maxpages);
+ free_swap_cluster_info(p);
+ p->max = 0;
+ p->cluster_info = NULL;
inode = mapping->host;
@@ -3664,6 +3664,8 @@ static int setup_swap_clusters_info(struct swap_info_struct *si,
if (!cluster_info)
goto err;
+ si->cluster_info = cluster_info;
+
for (i = 0; i < nr_clusters; i++)
spin_lock_init(&cluster_info[i].lock);
@@ -3727,7 +3729,7 @@ static int setup_swap_clusters_info(struct swap_info_struct *si,
si->cluster_info = cluster_info;
return 0;
err:
- free_swap_cluster_info(cluster_info, maxpages);
+ free_swap_cluster_info(si);
return err;
}
@@ -3949,7 +3951,7 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
si->global_cluster = NULL;
inode = NULL;
destroy_swap_extents(si, swap_file);
- free_swap_cluster_info(si->cluster_info, si->max);
+ free_swap_cluster_info(si);
si->cluster_info = NULL;
/*
* Clear the SWP_USED flag after all resources are freed so
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 03/14] mm, swap: back the cluster_info array with a sparse VM_SPARSE area
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
2026-10-03 0:31 ` [PATCH v4 01/14] mm: xswap support for zswap Baoquan He
2026-10-03 0:31 ` [PATCH v4 02/14] mm, swap: refactor free_swap_cluster_info to take swap_info_struct Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-03 0:31 ` [PATCH v4 04/14] mm, swap: add sysfs create interface for xswap Baoquan He
` (11 subsequent siblings)
14 siblings, 0 replies; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
An xswap device has no backing storage, so its size is not fixed up
front: the cluster_info array does not have to cover the whole address
space at once. Back it with a VM_SPARSE area that is populated lazily in
chunks, so only an initial chunk is mapped at device setup and the rest
is mapped on demand.
Add CONFIG_XSWAP, which depends on SWAP && 64BIT && ZSWAP (xswap devices
are backed by zswap) and on SYSFS (currently the only way to create one),
and the fields this needs:
- cluster_vm: the VM_SPARSE vm_struct backing the array
- nr_clusters_max: total number of clusters in the address space
- nr_clusters_mapped: number of clusters currently mapped
- xswap_lock: serializes map/unmap
The two cluster_info walkers, find_next_to_unuse() and
wait_for_allocation(), both walk up to si->max today. Bound them to
nr_clusters_mapped instead, so they stay inside the mapped range.
The grow path rejects stale ranges and cleans up a partial mapping on
failure.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
include/linux/swap.h | 6 +
mm/Kconfig | 9 ++
mm/swapfile.c | 296 +++++++++++++++++++++++++++++++++++++++++--
3 files changed, 302 insertions(+), 9 deletions(-)
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 9073d29377d8..382a578140b5 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -242,6 +242,12 @@ struct swap_info_struct {
signed char type; /* strange name for an index */
unsigned int max; /* size of this swap device */
struct swap_cluster_info *cluster_info; /* array, one entry per cluster */
+#ifdef CONFIG_XSWAP
+ struct vm_struct *cluster_vm; /* VM_SPARSE area for cluster_info */
+ unsigned long nr_clusters_max;/* total clusters in the xswap address space */
+ unsigned long nr_clusters_mapped; /* currently mapped cluster count */
+ struct mutex xswap_lock; /* serialize map/unmap operations */
+#endif
struct list_head free_clusters; /* free clusters list */
struct list_head full_clusters; /* full clusters list */
struct list_head nonfull_clusters[SWAP_NR_ORDERS];
diff --git a/mm/Kconfig b/mm/Kconfig
index acefc994d9a8..39927b5c75ac 100644
--- a/mm/Kconfig
+++ b/mm/Kconfig
@@ -122,6 +122,15 @@ config ZSWAP_COMPRESSOR_DEFAULT
default "zstd" if ZSWAP_COMPRESSOR_DEFAULT_ZSTD
default ""
+config XSWAP
+ bool "Extendable (virtual) swap device"
+ depends on SWAP && 64BIT && ZSWAP && SYSFS
+ help
+ Adds support for extendable swap devices (xswap) that decouple
+ PTE swap entries from physical backing storage. The cluster_info
+ array is backed by a sparse vmalloc area that grows and shrinks
+ on demand, avoiding static pre-allocation overhead.
+
config ZSMALLOC
tristate
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 13ec5c0f3c2f..767d3f46877f 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -49,6 +49,25 @@
#include "internal.h"
#include "swap.h"
+#ifdef CONFIG_XSWAP
+/*
+ * xswap: dynamically grow the cluster_info array via a VM_SPARSE area.
+ *
+ * XSWAP_GROW_CLUSTERS is the number of clusters to map in one grow
+ * operation. It is set to the number of cluster_info structs that
+ * fit in a single page (at least 16), so that the vmalloc page table
+ * overhead is proportional to the number of clusters mapped.
+ */
+#define XSWAP_GROW_CLUSTERS \
+ max_t(unsigned long, PAGE_SIZE / sizeof(struct swap_cluster_info), 16)
+
+static int xswap_map_clusters(struct swap_info_struct *si,
+ unsigned long start_idx, unsigned long nr);
+static void xswap_unmap_clusters(struct swap_info_struct *si,
+ unsigned long start_idx, unsigned long nr);
+static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data);
+#endif
+
static void swap_range_alloc(struct swap_info_struct *si,
unsigned int nr_entries);
static bool folio_swapcache_freeable(struct folio *folio);
@@ -2799,10 +2818,27 @@ static unsigned int find_next_to_unuse(struct swap_info_struct *si,
unsigned int prev)
{
struct swap_cluster_info *ci;
- unsigned long i, end;
+ unsigned long i, cluster_end, end;
unsigned int ci_off;
unsigned long swp_tb;
+ end = si->max;
+#ifdef CONFIG_XSWAP
+ /*
+ * An xswap device maps its cluster_info in chunks as it grows, so
+ * the tail past nr_clusters_mapped has nothing behind it.
+ */
+ if (si->flags & SWP_XSWAP) {
+ unsigned long mapped_end;
+
+ /* Pairs with the smp_store_release() in xswap_map_clusters(). */
+ mapped_end = smp_load_acquire(&si->nr_clusters_mapped) *
+ SWAPFILE_CLUSTER;
+ if (mapped_end < end)
+ end = mapped_end;
+ }
+#endif
+
/*
* No need for swap_lock here: we're just looking
* for whether an entry is in use, not modifying it; false
@@ -2810,11 +2846,11 @@ static unsigned int find_next_to_unuse(struct swap_info_struct *si,
* allocations from this area (while holding swap_lock).
*/
i = prev + 1;
- while (i < si->max) {
+ while (i < end) {
ci = __swap_offset_to_cluster(si, i);
- end = min_t(unsigned long,
- ALIGN_DOWN(i, SWAPFILE_CLUSTER) + SWAPFILE_CLUSTER,
- si->max);
+ cluster_end = min_t(unsigned long,
+ ALIGN_DOWN(i, SWAPFILE_CLUSTER) + SWAPFILE_CLUSTER,
+ end);
/*
* An empty cluster has no slot in use, so skip it whole.
@@ -2824,13 +2860,13 @@ static unsigned int find_next_to_unuse(struct swap_info_struct *si,
* enough, unlike in every other cluster_is_empty() caller.
*/
if (!READ_ONCE(ci->count)) {
- i = end;
+ i = cluster_end;
cond_resched();
continue;
}
ci_off = i % SWAPFILE_CLUSTER;
- for (; i < end; ci_off++, i++) {
+ for (; i < cluster_end; ci_off++, i++) {
swp_tb = swap_table_get(ci, ci_off);
if (!swp_tb_is_null(swp_tb) && !swp_tb_is_bad(swp_tb))
return i;
@@ -2840,7 +2876,6 @@ static unsigned int find_next_to_unuse(struct swap_info_struct *si,
return 0;
}
-
static int try_to_unuse(unsigned int type)
{
struct mm_struct *prev_mm;
@@ -3138,6 +3173,14 @@ static void wait_for_allocation(struct swap_info_struct *si)
BUG_ON(si->flags & SWP_WRITEOK);
+#ifdef CONFIG_XSWAP
+ if (si->flags & SWP_XSWAP) {
+ /* Pairs with the smp_store_release() in xswap_map_clusters(). */
+ end = min(end, smp_load_acquire(&si->nr_clusters_mapped) *
+ SWAPFILE_CLUSTER);
+ }
+#endif
+
for (offset = 0; offset < end; offset += SWAPFILE_CLUSTER) {
ci = swap_cluster_lock(si, offset);
swap_cluster_unlock(ci);
@@ -3149,11 +3192,41 @@ static void free_swap_cluster_info(struct swap_info_struct *si)
struct swap_cluster_info *cluster_info = si->cluster_info;
unsigned long maxpages = si->max;
struct swap_cluster_info *ci;
- int i, nr_clusters;
+ unsigned long i, nr_clusters;
if (!cluster_info)
return;
+#ifdef CONFIG_XSWAP
+ if (si->flags & SWP_XSWAP) {
+ unsigned long nr_mapped;
+
+ /*
+ * Cluster 0 keeps the bad header slot, so it never empties
+ * and __free_cluster() never frees its table.
+ */
+ /* Pairs with the smp_store_release() in xswap_map_clusters(). */
+ nr_mapped = smp_load_acquire(&si->nr_clusters_mapped);
+ for (i = 0; i < nr_mapped; i++) {
+ ci = &cluster_info[i];
+ spin_lock(&ci->lock);
+ if (cluster_table_is_alloced(ci)) {
+ swap_cluster_assert_empty(ci, 0, SWAPFILE_CLUSTER, true);
+ swap_cluster_free_table(ci);
+ }
+ spin_unlock(&ci->lock);
+ }
+ /* Unmap all mapped clusters and free the VM_SPARSE area */
+ if (si->nr_clusters_mapped > 0)
+ xswap_unmap_clusters(si, 0, si->nr_clusters_mapped);
+ free_vm_area(si->cluster_vm);
+ si->cluster_vm = NULL;
+ si->cluster_info = NULL;
+ si->nr_clusters_mapped = 0;
+ return;
+ }
+#endif
+
nr_clusters = DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER);
for (i = 0; i < nr_clusters; i++) {
ci = cluster_info + i;
@@ -3651,6 +3724,148 @@ static unsigned long read_swap_header(struct swap_info_struct *si,
return maxpages;
}
+#ifdef CONFIG_XSWAP
+static int xswap_map_clusters(struct swap_info_struct *si,
+ unsigned long start_idx, unsigned long nr)
+{
+ unsigned long start_addr = (unsigned long)si->cluster_info +
+ (size_t)start_idx * sizeof(struct swap_cluster_info);
+ unsigned long end_addr = start_addr + (size_t)nr * sizeof(struct swap_cluster_info);
+ unsigned long vm_start = PAGE_ALIGN(start_addr);
+ unsigned long vm_end = PAGE_ALIGN(end_addr);
+ unsigned int noreclaim_flags;
+ unsigned long mapped_end;
+ unsigned long npages;
+ struct page **pages;
+ unsigned long i;
+ int err;
+
+ mutex_lock(&si->xswap_lock);
+
+ /* Refuse a stale range: the boundary moved since the caller read it. */
+ if (start_idx != READ_ONCE(si->nr_clusters_mapped)) {
+ mutex_unlock(&si->xswap_lock);
+ return -EAGAIN;
+ }
+ if (start_idx + nr > si->nr_clusters_max) {
+ mutex_unlock(&si->xswap_lock);
+ return -EAGAIN;
+ }
+
+ /*
+ * Page-granular mapping can cover clusters past the previous chunk.
+ * Find the already-mapped prefix and map only the rest.
+ */
+ mapped_end = vm_start;
+ if (vm_start < vm_end)
+ apply_to_existing_page_range(&init_mm, vm_start,
+ vm_end - vm_start,
+ xswap_mapped_end, &mapped_end);
+ if (vm_start >= vm_end || mapped_end == vm_end)
+ goto mapped;
+ vm_start = mapped_end;
+
+ npages = (vm_end - vm_start) >> PAGE_SHIFT;
+
+ /* Prevent recursive reclaim during vmap page table allocation. */
+ noreclaim_flags = memalloc_noreclaim_save();
+
+ pages = kmalloc_array(npages, sizeof(*pages),
+ __GFP_HIGH | __GFP_NOMEMALLOC | GFP_KERNEL);
+ if (!pages) {
+ memalloc_noreclaim_restore(noreclaim_flags);
+ mutex_unlock(&si->xswap_lock);
+ return -ENOMEM;
+ }
+
+ for (i = 0; i < npages; i++) {
+ /* __GFP_ZERO: cluster_info pointer fields must start NULL. */
+ pages[i] = alloc_page(__GFP_HIGH | __GFP_NOMEMALLOC |
+ GFP_KERNEL | __GFP_ZERO);
+ if (!pages[i])
+ goto fail;
+ }
+
+ err = vm_area_map_pages(si->cluster_vm, vm_start, vm_end, pages);
+ if (err) {
+ i = npages;
+ goto fail;
+ }
+
+ for (i = start_idx; i < start_idx + nr; i++)
+ spin_lock_init(&si->cluster_info[i].lock);
+
+ kfree(pages);
+ memalloc_noreclaim_restore(noreclaim_flags);
+
+ /*
+ * Publish the new mappings and cluster lock initialization before
+ * the count; walkers without xswap_lock use smp_load_acquire().
+ */
+ smp_store_release(&si->nr_clusters_mapped, start_idx + nr);
+ mutex_unlock(&si->xswap_lock);
+ return 0;
+
+mapped:
+ for (i = start_idx; i < start_idx + nr; i++)
+ spin_lock_init(&si->cluster_info[i].lock);
+
+ /* Publish the advanced count. */
+ smp_store_release(&si->nr_clusters_mapped, start_idx + nr);
+ mutex_unlock(&si->xswap_lock);
+ return 0;
+
+fail:
+ /* Clear PTEs vm_area_map_pages() may have left before freeing pages. */
+ vm_area_unmap_pages(si->cluster_vm, vm_start, vm_end);
+ while (i > 0) {
+ i--;
+ if (pages[i])
+ __free_page(pages[i]);
+ }
+ memalloc_noreclaim_restore(noreclaim_flags);
+ kfree(pages);
+ mutex_unlock(&si->xswap_lock);
+ return -ENOMEM;
+}
+
+static void xswap_unmap_clusters(struct swap_info_struct *si,
+ unsigned long start_idx, unsigned long nr)
+{
+ unsigned long start_addr = (unsigned long)si->cluster_info +
+ (size_t)start_idx * sizeof(struct swap_cluster_info);
+ unsigned long end_addr = start_addr + (size_t)nr * sizeof(struct swap_cluster_info);
+ unsigned long vm_start = PAGE_ALIGN(start_addr);
+ unsigned long vm_end = PAGE_ALIGN(end_addr);
+
+ mutex_lock(&si->xswap_lock);
+
+ if (vm_start >= vm_end) {
+ WRITE_ONCE(si->nr_clusters_mapped, start_idx);
+ mutex_unlock(&si->xswap_lock);
+ return;
+ }
+
+ vm_area_unmap_pages(si->cluster_vm, vm_start, vm_end);
+ /* vm_area_unmap_pages() clears PTEs but does not free pages. */
+ /* TODO: free backing pages via page table walk or tracking bitmap */
+
+ WRITE_ONCE(si->nr_clusters_mapped, start_idx);
+ mutex_unlock(&si->xswap_lock);
+}
+
+/* Track the end of the run of pages that is already mapped. */
+static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data)
+{
+ unsigned long *mapped_end = data;
+
+ if (!pte_present(ptep_get(pte)))
+ return 0;
+ *mapped_end = addr + PAGE_SIZE;
+ return 0;
+}
+#endif /* CONFIG_XSWAP */
+
static int setup_swap_clusters_info(struct swap_info_struct *si,
union swap_header *swap_header,
unsigned long maxpages)
@@ -3660,6 +3875,69 @@ static int setup_swap_clusters_info(struct swap_info_struct *si,
int err = -ENOMEM;
unsigned long i;
+#ifdef CONFIG_XSWAP
+ if (si->flags & SWP_XSWAP) {
+ unsigned long size = PAGE_ALIGN(nr_clusters * sizeof(*cluster_info));
+ struct vm_struct *vm;
+
+ vm = get_vm_area(size, VM_SPARSE);
+ if (!vm)
+ goto err;
+
+ cluster_info = vm->addr;
+ si->cluster_vm = vm;
+ si->nr_clusters_max = nr_clusters;
+ si->cluster_info = cluster_info;
+
+ /* Must be initialized before xswap_map_clusters() locks it. */
+ mutex_init(&si->xswap_lock);
+
+ if (xswap_map_clusters(si, 0, min_t(unsigned long,
+ XSWAP_GROW_CLUSTERS, nr_clusters)))
+ goto err_free_vm;
+
+ /* xswap: only cluster 0 slot 0 is bad */
+ err = swap_cluster_setup_bad_slot(si, cluster_info, 0, false);
+ if (err)
+ goto err_unmap;
+
+ INIT_LIST_HEAD(&si->free_clusters);
+ INIT_LIST_HEAD(&si->full_clusters);
+ INIT_LIST_HEAD(&si->discard_clusters);
+ for (i = 0; i < SWAP_NR_ORDERS; i++) {
+ INIT_LIST_HEAD(&si->nonfull_clusters[i]);
+ INIT_LIST_HEAD(&si->frag_clusters[i]);
+ }
+
+ /*
+ * Cluster 0 holds the header slot, which is marked bad, so it
+ * is not entirely free. si->max is cluster aligned, so no
+ * cluster holds a slot past the end of the device.
+ */
+ for (i = 0; i < si->nr_clusters_mapped; i++) {
+ struct swap_cluster_info *ci = &cluster_info[i];
+
+ if (ci->count) {
+ ci->flags = CLUSTER_FLAG_NONFULL;
+ list_add_tail(&ci->list, &si->nonfull_clusters[0]);
+ } else {
+ ci->flags = CLUSTER_FLAG_FREE;
+ list_add_tail(&ci->list, &si->free_clusters);
+ }
+ }
+
+ return 0;
+
+err_unmap:
+ xswap_unmap_clusters(si, 0, si->nr_clusters_mapped);
+err_free_vm:
+ free_vm_area(si->cluster_vm);
+ si->cluster_vm = NULL;
+ si->cluster_info = NULL;
+ return err;
+ }
+#endif /* CONFIG_XSWAP */
+
cluster_info = kvzalloc_objs(*cluster_info, nr_clusters);
if (!cluster_info)
goto err;
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 04/14] mm, swap: add sysfs create interface for xswap
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
` (2 preceding siblings ...)
2026-10-03 0:31 ` [PATCH v4 03/14] mm, swap: back the cluster_info array with a sparse VM_SPARSE area Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-07 6:32 ` KunWu Chan
2026-10-03 0:31 ` [PATCH v4 05/14] mm, swap: add xswap grow trigger on cluster allocation Baoquan He
` (10 subsequent siblings)
14 siblings, 1 reply; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
xswap devices have no backing storage, so there is no file to swapon.
Add a sysfs interface to create them directly:
/sys/kernel/mm/xswap/create write an optional priority to create
a device, empty for the default
A new device is created with si->max and nr_clusters_max both equal to
twice RAM: the cluster_info array is a sparse VM_SPARSE area mapped
lazily, so an idle device costs nothing. The device shows up in
/proc/swaps as "xswap<N>".
It takes the highest priority by default. Swapout picks the
highest-priority device that has room, so a real device ahead of an xswap
one would take the page directly: zswap would still cache it, but the
slot on the real device is taken either way, and not taking it is the
point of xswap. A lower priority can be asked for, not a higher one.
Hibernation device discovery skips xswap devices: they have no bdev
to carry a resume image.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 169 ++++++++++++++++++++++++++++++++++++++++++++++++--
1 file changed, 164 insertions(+), 5 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 767d3f46877f..9c204e455ef9 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -16,6 +16,8 @@
#include <linux/kernel_stat.h>
#include <linux/swap.h>
#include <linux/vmalloc.h>
+#include <linux/kobject.h>
+#include <linux/sysfs.h>
#include <linux/pagemap.h>
#include <linux/namei.h>
#include <linux/shmem_fs.h>
@@ -48,6 +50,13 @@
#include "swap_table.h"
#include "internal.h"
#include "swap.h"
+#define DEF_SWAP_PRIO -1
+
+/*
+ * Above every real swap device: one ahead of xswap would take the page
+ * directly, and taking that slot is what xswap exists to avoid.
+ */
+#define DEF_XSWAP_PRIO SWAP_FLAG_PRIO_MASK
#ifdef CONFIG_XSWAP
/*
@@ -66,7 +75,65 @@ static int xswap_map_clusters(struct swap_info_struct *si,
static void xswap_unmap_clusters(struct swap_info_struct *si,
unsigned long start_idx, unsigned long nr);
static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data);
-#endif
+
+static int xswap_create(int prio);
+
+static ssize_t xswap_create_store(struct kobject *kobj,
+ struct kobj_attribute *attr,
+ const char *buf, size_t count)
+{
+ int prio = DEF_XSWAP_PRIO;
+ int err;
+
+ if (!capable(CAP_SYS_ADMIN))
+ return -EPERM;
+
+ /* "[<prio>]" is optional; a lower one is allowed, not a higher. */
+ if (*skip_spaces(buf)) {
+ err = kstrtoint(buf, 10, &prio);
+ if (err)
+ return err;
+ }
+
+ err = xswap_create(prio);
+ if (err < 0)
+ return err;
+
+ return count;
+}
+
+static struct kobj_attribute xswap_create_attr = __ATTR(create, 0200, NULL,
+ xswap_create_store);
+
+static struct attribute *xswap_attrs[] = {
+ &xswap_create_attr.attr,
+ NULL,
+};
+
+static const struct attribute_group xswap_attr_group = {
+ .attrs = xswap_attrs,
+};
+
+static struct kobject *xswap_kobj;
+
+static void xswap_sysfs_init(void)
+{
+ xswap_kobj = kobject_create_and_add("xswap", mm_kobj);
+ if (!xswap_kobj) {
+ pr_err("xswap: failed to create sysfs kobject\n");
+ return;
+ }
+ if (sysfs_create_group(xswap_kobj, &xswap_attr_group)) {
+ pr_err("xswap: failed to create sysfs group\n");
+ kobject_put(xswap_kobj);
+ xswap_kobj = NULL;
+ }
+}
+#else /* !CONFIG_XSWAP */
+static inline void xswap_sysfs_init(void)
+{
+}
+#endif /* CONFIG_XSWAP */
static void swap_range_alloc(struct swap_info_struct *si,
unsigned int nr_entries);
@@ -94,7 +161,6 @@ atomic_t nr_real_swapfiles;
EXPORT_SYMBOL_GPL(nr_swap_pages);
/* protected with swap_lock. reading in vm_swap_full() doesn't need lock */
long total_swap_pages;
-#define DEF_SWAP_PRIO -1
unsigned long swapfile_maximum_size;
#ifdef CONFIG_MIGRATION
bool swap_migration_ad_supported;
@@ -2275,6 +2341,9 @@ static int __find_hibernation_swap_type(dev_t device, sector_t offset)
if (!(sis->flags & SWP_WRITEOK))
continue;
+ /* xswap has no bdev to match a resume device */
+ if (sis->flags & SWP_XSWAP)
+ continue;
if (device == sis->bdev->bd_dev) {
struct swap_extent *se = first_se(sis);
@@ -2462,6 +2531,8 @@ int find_first_swap(dev_t *device)
if (!(sis->flags & SWP_WRITEOK))
continue;
+ if (sis->flags & SWP_XSWAP)
+ continue;
*device = sis->bdev->bd_dev;
spin_unlock(&swap_lock);
return type;
@@ -3425,7 +3496,8 @@ static void *swap_start(struct seq_file *swap, loff_t *pos)
return SEQ_START_TOKEN;
for (type = 0; (si = swap_type_to_info(type)); type++) {
- if (!(si->swap_file))
+ if (!si->swap_file &&
+ !((si->flags & SWP_XSWAP) && (si->flags & SWP_WRITEOK)))
continue;
if (!--l)
return si;
@@ -3446,7 +3518,8 @@ static void *swap_next(struct seq_file *swap, void *v, loff_t *pos)
++(*pos);
for (; (si = swap_type_to_info(type)); type++) {
- if (!(si->swap_file))
+ if (!si->swap_file &&
+ !((si->flags & SWP_XSWAP) && (si->flags & SWP_WRITEOK)))
continue;
return si;
}
@@ -3488,7 +3561,14 @@ static int swap_show(struct seq_file *swap, void *v)
inuse = K(swap_usage_in_pages(si));
file = si->swap_file;
- len = seq_file_path(swap, file, " \t\n\\");
+ if (file)
+ len = seq_file_path(swap, file, " \t\n\\");
+ else {
+ char name[16];
+
+ len = scnprintf(name, sizeof(name), "xswap%d", si->type);
+ seq_puts(swap, name);
+ }
seq_printf(swap, "%*s%s\t%lu\t%s%lu\t%s%d\n",
len < 40 ? 40 - len : 1, " ",
swap_type_str(si),
@@ -4011,6 +4091,83 @@ static int setup_swap_clusters_info(struct swap_info_struct *si,
return err;
}
+#ifdef CONFIG_XSWAP
+/* Create a file-less xswap device, sized at twice RAM. */
+static int xswap_create(int prio)
+{
+ struct swap_info_struct *si;
+ unsigned long ram, maxpages;
+ int error;
+
+ if (prio != DEF_XSWAP_PRIO &&
+ (prio < 0 || prio > SWAP_FLAG_PRIO_MASK))
+ return -EINVAL;
+
+ si = alloc_swap_info();
+ if (IS_ERR(si))
+ return PTR_ERR(si);
+
+ INIT_WORK(&si->discard_work, swap_discard_work);
+ INIT_WORK(&si->reclaim_work, swap_reclaim_work);
+
+ ram = totalram_pages();
+ maxpages = min_t(unsigned long, ram * 2, swapfile_maximum_size);
+ /* si->max is an unsigned int: don't overflow it. */
+ if (maxpages > UINT_MAX)
+ maxpages = UINT_MAX;
+ /* Cluster-aligned, so no cluster holds a slot past si->max. */
+ if (maxpages > SWAPFILE_CLUSTER)
+ maxpages = rounddown(maxpages, SWAPFILE_CLUSTER);
+ if (maxpages < 2)
+ maxpages = 2;
+
+ si->bdev = NULL;
+ si->flags |= SWP_XSWAP | SWP_SOLIDSTATE;
+ si->max = maxpages;
+ si->pages = maxpages - 1;
+ /*
+ * No backing file: setup_swap_extents() is only reachable from the
+ * file-backed swapon() path, so set ops here. Only ops->flags is
+ * used, by may_enter_fs(); the IO methods are never called because
+ * swap_writeout()/swap_read_folio() short circuit xswap.
+ */
+ si->ops = &swap_bdev_ops;
+
+ error = setup_swap_clusters_info(si, NULL, maxpages);
+ if (error)
+ goto bad_swap;
+
+ error = zswap_swapon(si->type, si->max);
+ if (error)
+ goto bad_swap;
+
+ mutex_lock(&swapon_mutex);
+ si->prio = prio;
+ si->list.prio = -si->prio;
+ si->avail_list.prio = -si->prio;
+ /* si->swap_file stays NULL: this is a file-less device */
+ enable_swap_info(si);
+ pr_info("xswap: adding extendable swap type %d (prio %d, %u pages, max %lu)\n",
+ si->type, prio, si->pages, maxpages);
+ mutex_unlock(&swapon_mutex);
+ atomic_inc(&proc_poll_event);
+ wake_up_interruptible(&proc_poll_wait);
+
+ return si->type;
+
+bad_swap:
+ kfree(si->global_cluster);
+ si->global_cluster = NULL;
+ destroy_swap_extents(si, NULL); /* safe: xswap never sets SWP_ACTIVATED */
+ free_swap_cluster_info(si);
+ si->cluster_info = NULL;
+ spin_lock(&swap_lock);
+ si->flags = 0;
+ spin_unlock(&swap_lock);
+ return error;
+}
+#endif /* CONFIG_XSWAP */
+
SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
{
struct swap_info_struct *si;
@@ -4357,6 +4514,8 @@ static int __init swapfile_init(void)
swap_migration_ad_supported = true;
#endif /* CONFIG_MIGRATION */
+ xswap_sysfs_init();
+
return 0;
}
subsys_initcall(swapfile_init);
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 05/14] mm, swap: add xswap grow trigger on cluster allocation
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
` (3 preceding siblings ...)
2026-10-03 0:31 ` [PATCH v4 04/14] mm, swap: add sysfs create interface for xswap Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-07 6:52 ` KunWu Chan
2026-10-03 0:31 ` [PATCH v4 06/14] mm, swap: add xswap_try_shrink and shrink trigger on cluster free Baoquan He
` (9 subsequent siblings)
14 siblings, 1 reply; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
Grow the mapped range before it runs out rather than after, from
cluster_alloc_swap_entry().
Growing maps pages into the VM_SPARSE area and can sleep. Doing it from
the allocation that would otherwise have failed puts that sleep on the
swap-out path with nothing left to fall back on. Growing at 85% in use
instead leaves the allocation that pays for it a reserve to land in, and
leaves one grow's worth of room in front of the next.
The target is 65%: past the point where the shrink starts looking (50%)
and short of the point that triggers the next grow (85%), so a grow
cannot leave the range in a state the shrink would undo.
Since xswap is always SWP_SOLIDSTATE, global_cluster_lock is never
held on this path.
This makes the xswap cluster space grow transparently as swap usage
increases, without any userspace intervention.
The caller holds local_lock(&percpu_swap_cluster.lock) across the whole
slow path, so drop it around xswap_map_clusters() and take it again
afterwards; it only protects the per-cpu cluster cache, which this path
does not touch.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 93 +++++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 93 insertions(+)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 9c204e455ef9..45050f8a2ed8 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -70,6 +70,19 @@
#define XSWAP_GROW_CLUSTERS \
max_t(unsigned long, PAGE_SIZE / sizeof(struct swap_cluster_info), 16)
+/*
+ * Grow and shrink thresholds, as a percentage of the mapped range in use.
+ * Each one lands inside (SHRINK_WHEN, GROW_WHEN), so neither can leave the
+ * range in a state that wakes the other up.
+ */
+#define XSWAP_SHRINK_WHEN 50
+#define XSWAP_SHRINK_UNTIL 70
+#define XSWAP_GROW_UNTIL 65
+#define XSWAP_GROW_WHEN 85
+#define XSWAP_GROW_CHUNKS 4
+
+static bool xswap_should_grow(struct swap_info_struct *si);
+static unsigned long xswap_grow(struct swap_info_struct *si);
static int xswap_map_clusters(struct swap_info_struct *si,
unsigned long start_idx, unsigned long nr);
static void xswap_unmap_clusters(struct swap_info_struct *si,
@@ -1204,6 +1217,17 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
if (order && !(si->flags & SWP_BLKDEV))
return 0;
+#ifdef CONFIG_XSWAP
+ /*
+ * Top the range up early. Every path below leaves through `done`,
+ * so this has to come first; mapping pages can sleep, and doing it
+ * from the allocation that would otherwise fail leaves nothing to
+ * fall back on.
+ */
+ if ((si->flags & SWP_XSWAP) && xswap_should_grow(si))
+ xswap_grow(si);
+#endif
+
if (!(si->flags & SWP_SOLIDSTATE)) {
/* Serialize HDD SWAP allocation for each device. */
spin_lock(&si->global_cluster_lock);
@@ -1280,6 +1304,12 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
if (found)
goto done;
}
+
+#ifdef CONFIG_XSWAP
+ /* A concurrent free or grow may have added clusters; retry once. */
+ if (!found && (si->flags & SWP_XSWAP))
+ found = alloc_swap_scan_list(si, &si->free_clusters, folio, false);
+#endif
done:
if (!(si->flags & SWP_SOLIDSTATE))
spin_unlock(&si->global_cluster_lock);
@@ -3909,6 +3939,69 @@ static int xswap_map_clusters(struct swap_info_struct *si,
return -ENOMEM;
}
+static bool xswap_should_grow(struct swap_info_struct *si)
+{
+ unsigned long mapped = READ_ONCE(si->nr_clusters_mapped);
+
+ if (mapped >= READ_ONCE(si->nr_clusters_max))
+ return false;
+
+ return swap_usage_in_pages(si) * 100 >=
+ mapped * SWAPFILE_CLUSTER * XSWAP_GROW_WHEN;
+}
+
+/*
+ * Map more of the cluster_info array, up to GROW_UNTIL in use. The caller
+ * holds percpu_swap_cluster.lock, which this drops while it sleeps.
+ */
+static unsigned long xswap_grow(struct swap_info_struct *si)
+{
+ unsigned long mapped = READ_ONCE(si->nr_clusters_mapped);
+ unsigned long nr_new, want, i;
+ int ret;
+
+ /* The unmap below is paired with the caller's lock. */
+ lockdep_assert_held(&percpu_swap_cluster.lock);
+
+ want = DIV_ROUND_UP(swap_usage_in_pages(si) * 100,
+ XSWAP_GROW_UNTIL * SWAPFILE_CLUSTER);
+ if (want <= mapped)
+ return 0;
+
+ nr_new = rounddown(want - mapped, XSWAP_GROW_CLUSTERS);
+ if (!nr_new)
+ nr_new = XSWAP_GROW_CLUSTERS;
+ if (nr_new > XSWAP_GROW_CHUNKS * XSWAP_GROW_CLUSTERS)
+ nr_new = XSWAP_GROW_CHUNKS * XSWAP_GROW_CLUSTERS;
+ nr_new = min(nr_new, READ_ONCE(si->nr_clusters_max) - mapped);
+
+ /*
+ * Mapping pages can sleep. The lock only guards the per-cpu
+ * cluster cache, which this path does not touch.
+ */
+ local_unlock(&percpu_swap_cluster.lock);
+ ret = xswap_map_clusters(si, mapped, nr_new);
+ local_lock(&percpu_swap_cluster.lock);
+ if (ret)
+ return 0;
+
+ for (i = mapped; i < mapped + nr_new; i++) {
+ struct swap_cluster_info *ci = &si->cluster_info[i];
+
+ /*
+ * A concurrent grower may have taken these already;
+ * only add the off-list ones.
+ */
+ spin_lock(&ci->lock);
+ if (ci->flags == CLUSTER_FLAG_NONE)
+ move_cluster(si, ci, &si->free_clusters,
+ CLUSTER_FLAG_FREE);
+ spin_unlock(&ci->lock);
+ }
+
+ return nr_new;
+}
+
static void xswap_unmap_clusters(struct swap_info_struct *si,
unsigned long start_idx, unsigned long nr)
{
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 06/14] mm, swap: add xswap_try_shrink and shrink trigger on cluster free
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
` (4 preceding siblings ...)
2026-10-03 0:31 ` [PATCH v4 05/14] mm, swap: add xswap grow trigger on cluster allocation Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-03 0:31 ` [PATCH v4 07/14] mm, swap: free backing pages in xswap_unmap_clusters Baoquan He
` (8 subsequent siblings)
14 siblings, 0 replies; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
When a cluster becomes entirely free, the mapped range may now have a
free tail that nothing is using. Reclaim it: unmap the tail clusters
from the cluster_info VM_SPARSE area.
Only reclaim once the mapped range is at most 50% in use. Growth is
demand driven, so reclaiming on a smaller dip would only map the same
clusters again, and every unmap costs an RCU grace period.
Stop at 70% in use rather than at the end of the free tail. Unmapping
the whole tail would leave the range all but full, and the grow, which
starts at 85%, would be straight back for it; 70% is far enough below
that the workload has to do real work before the grow can fire. Keep at
least one chunk of free tail, so the next allocation has somewhere to
land.
Unmapping fewer than four chunks is not worth the RCU grace period, so
leave the range alone in that case.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 114 +++++++++++++++++++++++++++++++++++++++++++++++---
1 file changed, 108 insertions(+), 6 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 45050f8a2ed8..b8572f301d31 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -60,12 +60,13 @@
#ifdef CONFIG_XSWAP
/*
- * xswap: dynamically grow the cluster_info array via a VM_SPARSE area.
+ * xswap: dynamically grow and shrink the cluster_info array via a
+ * VM_SPARSE area.
*
- * XSWAP_GROW_CLUSTERS is the number of clusters to map in one grow
- * operation. It is set to the number of cluster_info structs that
- * fit in a single page (at least 16), so that the vmalloc page table
- * overhead is proportional to the number of clusters mapped.
+ * XSWAP_GROW_CLUSTERS is the number of clusters to map/unmap in one
+ * grow/shrink operation: the number of cluster_info structs that fit in
+ * a single page (at least 16), so that the vmalloc page table overhead
+ * is proportional to the number of clusters mapped.
*/
#define XSWAP_GROW_CLUSTERS \
max_t(unsigned long, PAGE_SIZE / sizeof(struct swap_cluster_info), 16)
@@ -88,6 +89,7 @@ static int xswap_map_clusters(struct swap_info_struct *si,
static void xswap_unmap_clusters(struct swap_info_struct *si,
unsigned long start_idx, unsigned long nr);
static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data);
+static void xswap_try_shrink(struct swap_info_struct *si);
static int xswap_create(int prio);
@@ -716,6 +718,9 @@ static void __free_cluster(struct swap_info_struct *si, struct swap_cluster_info
swap_cluster_free_table(ci);
move_cluster(si, ci, &si->free_clusters, CLUSTER_FLAG_FREE);
ci->order = 0;
+#ifdef CONFIG_XSWAP
+ xswap_try_shrink(si);
+#endif
}
/*
@@ -1083,6 +1088,9 @@ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,
lockdep_assert_held(&ci->lock);
VM_WARN_ON(!cluster_is_usable(ci, order));
+ /* ci is used without ci->lock; an xswap unmap waits for this. */
+ rcu_read_lock();
+
if (end < nr_pages || ci->count + nr_pages > SWAPFILE_CLUSTER)
goto out;
@@ -1111,6 +1119,7 @@ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,
out:
relocate_cluster(si, ci);
swap_cluster_unlock(ci);
+ rcu_read_unlock();
if (si->flags & SWP_SOLIDSTATE) {
this_cpu_write(percpu_swap_cluster.offset[order], next);
this_cpu_write(percpu_swap_cluster.si[order], si);
@@ -1154,6 +1163,9 @@ static void swap_reclaim_full_clusters(struct swap_info_struct *si, bool force)
to_scan = swap_usage_in_pages(si) / SWAPFILE_CLUSTER;
while ((ci = isolate_lock_cluster(si, &si->full_clusters))) {
+ /* As in alloc_swap_scan_cluster(). */
+ rcu_read_lock();
+
offset = cluster_offset(si, ci);
end = min(si->max, offset + SWAPFILE_CLUSTER);
to_scan--;
@@ -1178,6 +1190,7 @@ static void swap_reclaim_full_clusters(struct swap_info_struct *si, bool force)
relocate_cluster(si, ci);
swap_cluster_unlock(ci);
+ rcu_read_unlock();
if (to_scan <= 0)
break;
@@ -1502,11 +1515,17 @@ static bool swap_alloc_fast(struct folio *folio)
/*
* Once allocated, swap_info_struct will never be completely freed,
* so checking it's liveness by get_swap_device_info is enough.
+ *
+ * The cached offset indexes si->cluster_info, which xswap can
+ * unmap; cover both the read and the use with RCU.
*/
+ rcu_read_lock();
si = this_cpu_read(percpu_swap_cluster.si[order]);
offset = this_cpu_read(percpu_swap_cluster.offset[order]);
- if (!si || !offset || !get_swap_device_info(si))
+ if (!si || !offset || !get_swap_device_info(si)) {
+ rcu_read_unlock();
return false;
+ }
ci = swap_cluster_lock(si, offset);
if (cluster_is_usable(ci, order)) {
@@ -1518,6 +1537,7 @@ static bool swap_alloc_fast(struct folio *folio)
}
put_swap_device(si);
+ rcu_read_unlock();
return folio_test_swapcache(folio);
}
@@ -2308,8 +2328,11 @@ swp_entry_t swap_alloc_hibernation_slot(int type)
/*
* Try the local cluster first if it matches the device. If
* not, try grab a new cluster and override local cluster.
+ *
+ * Same RCU requirement as swap_alloc_fast().
*/
local_lock(&percpu_swap_cluster.lock);
+ rcu_read_lock();
pcp_si = this_cpu_read(percpu_swap_cluster.si[0]);
pcp_offset = this_cpu_read(percpu_swap_cluster.offset[0]);
if (pcp_si == si && pcp_offset) {
@@ -2319,6 +2342,7 @@ swp_entry_t swap_alloc_hibernation_slot(int type)
else
swap_cluster_unlock(ci);
}
+ rcu_read_unlock();
if (!offset)
offset = cluster_alloc_swap_entry(si, NULL);
local_unlock(&percpu_swap_cluster.lock);
@@ -4019,6 +4043,15 @@ static void xswap_unmap_clusters(struct swap_info_struct *si,
return;
}
+ /*
+ * A per-cpu cluster cache can still hold an offset in this range.
+ * Invalidate those references, then wait out the readers that have
+ * already loaded one, so that nobody can dereference cluster_info
+ * past this point. swapoff() needs the same before it releases.
+ */
+ flush_percpu_swap_cluster(si);
+ synchronize_rcu();
+
vm_area_unmap_pages(si->cluster_vm, vm_start, vm_end);
/* vm_area_unmap_pages() clears PTEs but does not free pages. */
/* TODO: free backing pages via page table walk or tracking bitmap */
@@ -4037,6 +4070,75 @@ static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data)
*mapped_end = addr + PAGE_SIZE;
return 0;
}
+
+/*
+ * Automatic reclaim: leave one chunk of the free tail mapped as slack, so
+ * that the next allocation has somewhere to land, and only unmap once
+ * several chunks can go, so the unmap is worth the RCU grace period it
+ * costs.
+ */
+#define XSWAP_SHRINK_SLACK XSWAP_GROW_CLUSTERS
+#define XSWAP_SHRINK_MIN (XSWAP_GROW_CLUSTERS * 8)
+
+/*
+ * Try to shrink the cluster_info tail: unmap contiguous free clusters
+ * at the end of the mapped range.
+ */
+static void xswap_try_shrink(struct swap_info_struct *si)
+{
+ struct swap_cluster_info *ci;
+ unsigned long nr_mapped, last, keep, idx;
+
+ if (!(si->flags & SWP_XSWAP))
+ return;
+
+ nr_mapped = READ_ONCE(si->nr_clusters_mapped);
+ if (nr_mapped <= 1) /* keep cluster 0 */
+ return;
+
+ /*
+ * Reclaim on our own, but only once the mapped range is at most
+ * SHRINK_WHEN in use: growth is demand driven, so reclaiming on a
+ * smaller dip would only map the same clusters again, and every
+ * unmap costs an RCU grace period.
+ */
+ if (swap_usage_in_pages(si) * 100 >
+ nr_mapped * SWAPFILE_CLUSTER * XSWAP_SHRINK_WHEN)
+ return;
+
+ /* Find the last non-free cluster from the tail */
+ last = nr_mapped;
+ while (last > 1) {
+ idx = last - 1;
+ ci = &si->cluster_info[idx];
+ if (ci->count || ci->flags != CLUSTER_FLAG_FREE)
+ break;
+ last = idx;
+ }
+
+ if (last == nr_mapped)
+ return; /* nothing to shrink */
+
+ /* Below `last` has to stay mapped: the free ones in between are
+ * not part of the tail, and unmapping them orphans what is above.
+ */
+ if (nr_mapped - last < XSWAP_SHRINK_SLACK + XSWAP_SHRINK_MIN)
+ return;
+
+ /*
+ * Stop at SHRINK_UNTIL rather than at the end of the tail, or the
+ * range comes out full enough for the grow to be woken.
+ */
+ keep = DIV_ROUND_UP(swap_usage_in_pages(si) * 100,
+ XSWAP_SHRINK_UNTIL * SWAPFILE_CLUSTER);
+ if (keep < last + XSWAP_SHRINK_SLACK)
+ keep = last + XSWAP_SHRINK_SLACK;
+
+ if (nr_mapped < keep + XSWAP_SHRINK_MIN)
+ return;
+
+ xswap_unmap_clusters(si, keep, nr_mapped - keep);
+}
#endif /* CONFIG_XSWAP */
static int setup_swap_clusters_info(struct swap_info_struct *si,
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 07/14] mm, swap: free backing pages in xswap_unmap_clusters
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
` (5 preceding siblings ...)
2026-10-03 0:31 ` [PATCH v4 06/14] mm, swap: add xswap_try_shrink and shrink trigger on cluster free Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-03 0:31 ` [PATCH v4 08/14] mm, swap: defer xswap shrink to workqueue to avoid lock recursion Baoquan He
` (7 subsequent siblings)
14 siblings, 0 replies; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
vm_area_unmap_pages() does not free the backing pages that
xswap_map_clusters() allocated, so they leaked on every unmap.
Collect the backing pages from the PTEs before unmapping and free them
after the PTEs are cleared. The collection array is allocated with
kvmalloc_array(GFP_KERNEL), which falls back to vmalloc() and may
reclaim; if it still fails, unmap in bounded batches with a stack array
instead. The unmap therefore cannot fail and the callers do not need a
retry loop.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 95 ++++++++++++++++++++++++++++++++++++++++++++++-----
1 file changed, 86 insertions(+), 9 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index b8572f301d31..7c062e772f5b 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -84,6 +84,14 @@
static bool xswap_should_grow(struct swap_info_struct *si);
static unsigned long xswap_grow(struct swap_info_struct *si);
+
+/* Batch size for the fallback unmap; keeps the stack array small. */
+#define XSWAP_UNMAP_BATCH_PAGES 16
+#define XSWAP_UNMAP_BATCH_CLUSTERS \
+ max_t(unsigned long, \
+ ((XSWAP_UNMAP_BATCH_PAGES * PAGE_SIZE) / \
+ sizeof(struct swap_cluster_info)), 16)
+
static int xswap_map_clusters(struct swap_info_struct *si,
unsigned long start_idx, unsigned long nr);
static void xswap_unmap_clusters(struct swap_info_struct *si,
@@ -3341,9 +3349,8 @@ static void free_swap_cluster_info(struct swap_info_struct *si)
}
spin_unlock(&ci->lock);
}
- /* Unmap all mapped clusters and free the VM_SPARSE area */
- if (si->nr_clusters_mapped > 0)
- xswap_unmap_clusters(si, 0, si->nr_clusters_mapped);
+ /* free_vm_area() drops the mapping without freeing the pages. */
+ xswap_unmap_clusters(si, 0, si->nr_clusters_mapped);
free_vm_area(si->cluster_vm);
si->cluster_vm = NULL;
si->cluster_info = NULL;
@@ -4026,21 +4033,65 @@ static unsigned long xswap_grow(struct swap_info_struct *si)
return nr_new;
}
+struct xswap_page_data {
+ struct page **pages;
+ unsigned long nr;
+ unsigned long max;
+};
+
+static int xswap_collect_page(pte_t *pte, unsigned long addr, void *data)
+{
+ struct xswap_page_data *xpd = data;
+ pte_t pteval = ptep_get(pte);
+
+ if (!pte_present(pteval))
+ return 0;
+ if (WARN_ON_ONCE(xpd->nr >= xpd->max))
+ return 0;
+ xpd->pages[xpd->nr++] = pte_page(pteval);
+ return 0;
+}
+
+/*
+ * Free the backing pages mapped in [vm_start, vm_end) and clear the PTEs.
+ * @pages must have room for @max pointers.
+ */
+static void xswap_unmap_range(struct swap_info_struct *si,
+ unsigned long vm_start, unsigned long vm_end,
+ struct page **pages, unsigned long max)
+{
+ struct xswap_page_data xpd = {
+ .pages = pages,
+ .max = max,
+ };
+ unsigned long i;
+
+ apply_to_existing_page_range(&init_mm, vm_start, vm_end - vm_start,
+ xswap_collect_page, &xpd);
+ vm_area_unmap_pages(si->cluster_vm, vm_start, vm_end);
+
+ for (i = 0; i < xpd.nr; i++)
+ __free_page(pages[i]);
+}
+
static void xswap_unmap_clusters(struct swap_info_struct *si,
unsigned long start_idx, unsigned long nr)
{
unsigned long start_addr = (unsigned long)si->cluster_info +
(size_t)start_idx * sizeof(struct swap_cluster_info);
- unsigned long end_addr = start_addr + (size_t)nr * sizeof(struct swap_cluster_info);
+ unsigned long end_addr = start_addr +
+ (size_t)nr * sizeof(struct swap_cluster_info);
unsigned long vm_start = PAGE_ALIGN(start_addr);
unsigned long vm_end = PAGE_ALIGN(end_addr);
+ struct page **pages;
+ unsigned int noreclaim_flags;
+ unsigned long npages, idx;
mutex_lock(&si->xswap_lock);
if (vm_start >= vm_end) {
WRITE_ONCE(si->nr_clusters_mapped, start_idx);
- mutex_unlock(&si->xswap_lock);
- return;
+ goto out_unlock;
}
/*
@@ -4052,11 +4103,37 @@ static void xswap_unmap_clusters(struct swap_info_struct *si,
flush_percpu_swap_cluster(si);
synchronize_rcu();
- vm_area_unmap_pages(si->cluster_vm, vm_start, vm_end);
- /* vm_area_unmap_pages() clears PTEs but does not free pages. */
- /* TODO: free backing pages via page table walk or tracking bitmap */
+ npages = (vm_end - vm_start) >> PAGE_SHIFT;
+
+ /* Reclaim here would re-enter xswap_map_clusters() and its lock. */
+ noreclaim_flags = memalloc_noreclaim_save();
+ pages = kvmalloc_array(npages, sizeof(*pages), GFP_KERNEL);
+ memalloc_noreclaim_restore(noreclaim_flags);
+ if (pages) {
+ xswap_unmap_range(si, vm_start, vm_end, pages, npages);
+ kvfree(pages);
+ WRITE_ONCE(si->nr_clusters_mapped, start_idx);
+ goto out_unlock;
+ }
+
+ /* No memory for the array: unmap in bounded batches instead. */
+ for (idx = start_idx; idx < start_idx + nr; idx += XSWAP_UNMAP_BATCH_CLUSTERS) {
+ unsigned long batch = min(start_idx + nr - idx,
+ XSWAP_UNMAP_BATCH_CLUSTERS);
+ unsigned long addr = (unsigned long)si->cluster_info +
+ (size_t)idx * sizeof(struct swap_cluster_info);
+ unsigned long bstart = PAGE_ALIGN(addr);
+ unsigned long bend = PAGE_ALIGN(addr + (size_t)batch *
+ sizeof(struct swap_cluster_info));
+ struct page *batch_pages[XSWAP_UNMAP_BATCH_PAGES + 1];
+
+ if (bstart < bend)
+ xswap_unmap_range(si, bstart, bend, batch_pages,
+ ARRAY_SIZE(batch_pages));
+ }
WRITE_ONCE(si->nr_clusters_mapped, start_idx);
+out_unlock:
mutex_unlock(&si->xswap_lock);
}
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 08/14] mm, swap: defer xswap shrink to workqueue to avoid lock recursion
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
` (6 preceding siblings ...)
2026-10-03 0:31 ` [PATCH v4 07/14] mm, swap: free backing pages in xswap_unmap_clusters Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-08 9:10 ` KunWu Chan
2026-10-03 0:31 ` [PATCH v4 09/14] mm, swap: refactor swapoff and add xswap_destroy Baoquan He
` (6 subsequent siblings)
14 siblings, 1 reply; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
__free_cluster() called xswap_try_shrink() while holding ci->lock, but
shrinking unmaps the backing pages and the subsequent unlock faults on
the unmapped address. Run the shrink via schedule_work() instead, so no
cluster lock is held. The work is only scheduled for xswap devices and
is cancelled on swapoff.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
include/linux/swap.h | 1 +
mm/swapfile.c | 129 ++++++++++++++++++++++++++++++++-----------
2 files changed, 97 insertions(+), 33 deletions(-)
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 382a578140b5..12cdc00f78f9 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -246,6 +246,7 @@ struct swap_info_struct {
struct vm_struct *cluster_vm; /* VM_SPARSE area for cluster_info */
unsigned long nr_clusters_max;/* total clusters in the xswap address space */
unsigned long nr_clusters_mapped; /* currently mapped cluster count */
+ struct work_struct xswap_shrink_work; /* deferred shrink trigger */
struct mutex xswap_lock; /* serialize map/unmap operations */
#endif
struct list_head free_clusters; /* free clusters list */
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 7c062e772f5b..49703731ffd5 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -727,7 +727,9 @@ static void __free_cluster(struct swap_info_struct *si, struct swap_cluster_info
move_cluster(si, ci, &si->free_clusters, CLUSTER_FLAG_FREE);
ci->order = 0;
#ifdef CONFIG_XSWAP
- xswap_try_shrink(si);
+ /* Only xswap devices, and not while the device is being torn down. */
+ if ((si->flags & SWP_XSWAP) && (si->flags & SWP_WRITEOK))
+ schedule_work(&si->xswap_shrink_work);
#endif
}
@@ -3334,6 +3336,7 @@ static void free_swap_cluster_info(struct swap_info_struct *si)
if (si->flags & SWP_XSWAP) {
unsigned long nr_mapped;
+ cancel_work_sync(&si->xswap_shrink_work);
/*
* Cluster 0 keeps the bad header slot, so it never empties
* and __free_cluster() never frees its table.
@@ -3452,6 +3455,11 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
spin_unlock(&p->lock);
spin_unlock(&swap_lock);
+#ifdef CONFIG_XSWAP
+ if (p->flags & SWP_XSWAP)
+ cancel_work_sync(&p->xswap_shrink_work);
+#endif
+
wait_for_allocation(p);
set_current_oom_origin();
@@ -4074,8 +4082,9 @@ static void xswap_unmap_range(struct swap_info_struct *si,
__free_page(pages[i]);
}
-static void xswap_unmap_clusters(struct swap_info_struct *si,
- unsigned long start_idx, unsigned long nr)
+/* Caller must hold si->xswap_lock. Cannot fail. */
+static void xswap_unmap_clusters_locked(struct swap_info_struct *si,
+ unsigned long start_idx, unsigned long nr)
{
unsigned long start_addr = (unsigned long)si->cluster_info +
(size_t)start_idx * sizeof(struct swap_cluster_info);
@@ -4087,11 +4096,9 @@ static void xswap_unmap_clusters(struct swap_info_struct *si,
unsigned int noreclaim_flags;
unsigned long npages, idx;
- mutex_lock(&si->xswap_lock);
-
if (vm_start >= vm_end) {
WRITE_ONCE(si->nr_clusters_mapped, start_idx);
- goto out_unlock;
+ return;
}
/*
@@ -4113,7 +4120,7 @@ static void xswap_unmap_clusters(struct swap_info_struct *si,
xswap_unmap_range(si, vm_start, vm_end, pages, npages);
kvfree(pages);
WRITE_ONCE(si->nr_clusters_mapped, start_idx);
- goto out_unlock;
+ return;
}
/* No memory for the array: unmap in bounded batches instead. */
@@ -4133,7 +4140,13 @@ static void xswap_unmap_clusters(struct swap_info_struct *si,
}
WRITE_ONCE(si->nr_clusters_mapped, start_idx);
-out_unlock:
+}
+
+static void xswap_unmap_clusters(struct swap_info_struct *si,
+ unsigned long start_idx, unsigned long nr)
+{
+ mutex_lock(&si->xswap_lock);
+ xswap_unmap_clusters_locked(si, start_idx, nr);
mutex_unlock(&si->xswap_lock);
}
@@ -4157,6 +4170,16 @@ static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data)
#define XSWAP_SHRINK_SLACK XSWAP_GROW_CLUSTERS
#define XSWAP_SHRINK_MIN (XSWAP_GROW_CLUSTERS * 8)
+static void xswap_shrink_work_fn(struct work_struct *work)
+{
+ struct swap_info_struct *si = container_of(work,
+ struct swap_info_struct, xswap_shrink_work);
+
+ if (!(READ_ONCE(si->flags) & SWP_WRITEOK))
+ return;
+ xswap_try_shrink(si);
+}
+
/*
* Try to shrink the cluster_info tail: unmap contiguous free clusters
* at the end of the mapped range.
@@ -4164,14 +4187,20 @@ static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data)
static void xswap_try_shrink(struct swap_info_struct *si)
{
struct swap_cluster_info *ci;
- unsigned long nr_mapped, last, keep, idx;
+ unsigned long nr_mapped, nr_tail, keep, nr_unmap, start_idx, i;
if (!(si->flags & SWP_XSWAP))
return;
+ mutex_lock(&si->xswap_lock);
+
+ /* A swapoff raced us and is about to walk this mapping. */
+ if (!(READ_ONCE(si->flags) & SWP_WRITEOK))
+ goto out_unlock;
+
nr_mapped = READ_ONCE(si->nr_clusters_mapped);
- if (nr_mapped <= 1) /* keep cluster 0 */
- return;
+ if (nr_mapped <= 1) /* keep cluster 0 */
+ goto out_unlock;
/*
* Reclaim on our own, but only once the mapped range is at most
@@ -4181,40 +4210,73 @@ static void xswap_try_shrink(struct swap_info_struct *si)
*/
if (swap_usage_in_pages(si) * 100 >
nr_mapped * SWAPFILE_CLUSTER * XSWAP_SHRINK_WHEN)
- return;
+ goto out_unlock;
- /* Find the last non-free cluster from the tail */
- last = nr_mapped;
- while (last > 1) {
- idx = last - 1;
- ci = &si->cluster_info[idx];
- if (ci->count || ci->flags != CLUSTER_FLAG_FREE)
+ /*
+ * Count the free clusters at the tail of the mapped range. Scanned,
+ * not tracked: the count must be exact to size the unmap, and an
+ * incremental count falls behind on out-of-order frees.
+ */
+ nr_tail = 0;
+ while (nr_mapped - nr_tail > 1) {
+ ci = &si->cluster_info[nr_mapped - nr_tail - 1];
+ if (READ_ONCE(ci->count) ||
+ READ_ONCE(ci->flags) != CLUSTER_FLAG_FREE)
break;
- last = idx;
+ nr_tail++;
}
-
- if (last == nr_mapped)
- return; /* nothing to shrink */
-
- /* Below `last` has to stay mapped: the free ones in between are
- * not part of the tail, and unmapping them orphans what is above.
- */
- if (nr_mapped - last < XSWAP_SHRINK_SLACK + XSWAP_SHRINK_MIN)
- return;
+ if (nr_tail < XSWAP_SHRINK_SLACK + XSWAP_SHRINK_MIN)
+ goto out_unlock;
/*
* Stop at SHRINK_UNTIL rather than at the end of the tail, or the
- * range comes out full enough for the grow to be woken.
+ * range comes out full enough for the grow to be woken. Below the
+ * tail everything stays mapped, and one chunk of tail with it.
*/
keep = DIV_ROUND_UP(swap_usage_in_pages(si) * 100,
XSWAP_SHRINK_UNTIL * SWAPFILE_CLUSTER);
- if (keep < last + XSWAP_SHRINK_SLACK)
- keep = last + XSWAP_SHRINK_SLACK;
+ if (keep < nr_mapped - nr_tail + XSWAP_SHRINK_SLACK)
+ keep = nr_mapped - nr_tail + XSWAP_SHRINK_SLACK;
if (nr_mapped < keep + XSWAP_SHRINK_MIN)
- return;
+ goto out_unlock;
+
+ nr_unmap = rounddown(nr_mapped - keep, XSWAP_GROW_CLUSTERS);
+ if (!nr_unmap)
+ goto out_unlock;
+ start_idx = nr_mapped - nr_unmap;
+
+ /*
+ * Only shrink a run that reaches the mapped end; otherwise
+ * truncating nr_clusters_mapped would orphan the active tail.
+ */
+ spin_lock(&si->lock);
+ for (i = start_idx; i < nr_mapped; i++) {
+ ci = &si->cluster_info[i];
+ if (READ_ONCE(ci->flags) != CLUSTER_FLAG_FREE)
+ break;
+ if (!spin_trylock(&ci->lock)) {
+ spin_unlock(&si->lock);
+ goto out_unlock;
+ }
+ spin_unlock(&ci->lock);
+ }
+ if (i != nr_mapped) {
+ spin_unlock(&si->lock);
+ goto out_unlock;
+ }
+
+ for (i = start_idx; i < nr_mapped; i++) {
+ ci = &si->cluster_info[i];
+ list_del_init(&ci->list);
+ WRITE_ONCE(ci->flags, CLUSTER_FLAG_NONE);
+ }
+ spin_unlock(&si->lock);
- xswap_unmap_clusters(si, keep, nr_mapped - keep);
+ xswap_unmap_clusters_locked(si, start_idx, nr_unmap);
+
+out_unlock:
+ mutex_unlock(&si->xswap_lock);
}
#endif /* CONFIG_XSWAP */
@@ -4278,6 +4340,7 @@ static int setup_swap_clusters_info(struct swap_info_struct *si,
}
}
+ INIT_WORK(&si->xswap_shrink_work, xswap_shrink_work_fn);
return 0;
err_unmap:
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 09/14] mm, swap: refactor swapoff and add xswap_destroy
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
` (7 preceding siblings ...)
2026-10-03 0:31 ` [PATCH v4 08/14] mm, swap: defer xswap shrink to workqueue to avoid lock recursion Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-03 0:31 ` [PATCH v4 10/14] mm, swap: require zswap for xswap devices Baoquan He
` (5 subsequent siblings)
14 siblings, 0 replies; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
Extract __swapoff() from sys_swapoff() so the teardown logic can be
shared, and make it work for file-less devices.
Add xswap_destroy() to tear down a file-less xswap device by swap type,
exposed via /sys/kernel/mm/xswap/destroy. Writing a swap type tears down
that device; it requires CAP_SYS_ADMIN.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 133 ++++++++++++++++++++++++++++++++++++++++++--------
1 file changed, 114 insertions(+), 19 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 49703731ffd5..0c8b716ec5ba 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -100,6 +100,7 @@ static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data);
static void xswap_try_shrink(struct swap_info_struct *si);
static int xswap_create(int prio);
+static int xswap_destroy(int type);
static ssize_t xswap_create_store(struct kobject *kobj,
struct kobj_attribute *attr,
@@ -128,8 +129,35 @@ static ssize_t xswap_create_store(struct kobject *kobj,
static struct kobj_attribute xswap_create_attr = __ATTR(create, 0200, NULL,
xswap_create_store);
+static ssize_t xswap_destroy_store(struct kobject *kobj,
+ struct kobj_attribute *attr,
+ const char *buf, size_t count)
+{
+ unsigned long type;
+ int err;
+
+ if (!capable(CAP_SYS_ADMIN))
+ return -EPERM;
+
+ err = kstrtoul(buf, 0, &type);
+ if (err)
+ return err;
+ if (type >= MAX_SWAPFILES)
+ return -EINVAL;
+
+ err = xswap_destroy(type);
+ if (err)
+ return err;
+
+ return count;
+}
+
+static struct kobj_attribute xswap_destroy_attr = __ATTR(destroy, 0200, NULL,
+ xswap_destroy_store);
+
static struct attribute *xswap_attrs[] = {
&xswap_create_attr.attr,
+ &xswap_destroy_attr.attr,
NULL,
};
@@ -3398,13 +3426,14 @@ static void flush_percpu_swap_cluster(struct swap_info_struct *si)
}
}
+static int swap_info_remove(struct swap_info_struct *p);
+static int __swapoff(struct swap_info_struct *p);
SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
{
struct swap_info_struct *p = NULL;
- struct file *swap_file, *victim;
+ struct file *victim;
struct address_space *mapping;
- struct inode *inode;
int err, found = 0;
if (!capable(CAP_SYS_ADMIN))
@@ -3421,7 +3450,7 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
spin_lock(&swap_lock);
plist_for_each_entry(p, &swap_active_head, list) {
if (p->flags & SWP_WRITEOK) {
- if (p->swap_file->f_mapping == mapping) {
+ if (p->swap_file && p->swap_file->f_mapping == mapping) {
found = 1;
break;
}
@@ -3440,24 +3469,58 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
goto out_dput;
}
- if (!security_vm_enough_memory_mm(current->mm, p->pages))
- vm_unacct_memory(p->pages);
- else {
- err = -ENOMEM;
+ err = swap_info_remove(p);
+ if (err) {
spin_unlock(&swap_lock);
goto out_dput;
}
+ spin_unlock(&swap_lock);
+
+ err = __swapoff(p);
+
+out_dput:
+ filp_close(victim, NULL);
+ return err;
+}
+
+/*
+ * Drop @p from the avail and active lists and undo its accounting. The
+ * caller must hold swap_lock and have checked that @p is WRITEOK and not
+ * pinned for hibernation.
+ *
+ * Returns 0, or -ENOMEM with swap_lock still held.
+ */
+static int swap_info_remove(struct swap_info_struct *p)
+{
+ if (!security_vm_enough_memory_mm(current->mm, p->pages))
+ vm_unacct_memory(p->pages);
+ else
+ return -ENOMEM;
+
spin_lock(&p->lock);
del_from_avail_list(p, true);
plist_del(&p->list, &swap_active_head);
atomic_long_sub(p->pages, &nr_swap_pages);
total_swap_pages -= p->pages;
spin_unlock(&p->lock);
- spin_unlock(&swap_lock);
+ return 0;
+}
+
+/* Common swap teardown after list removal; shared by sys_swapoff() and
+ * xswap_destroy().
+ */
+static int __swapoff(struct swap_info_struct *p)
+{
+ struct file *swap_file = NULL;
+ int err;
#ifdef CONFIG_XSWAP
- if (p->flags & SWP_XSWAP)
+ if (p->flags & SWP_XSWAP) {
cancel_work_sync(&p->xswap_shrink_work);
+ /* Wait out a shrink racing us from the sysfs write path. */
+ mutex_lock(&p->xswap_lock);
+ mutex_unlock(&p->xswap_lock);
+ }
#endif
wait_for_allocation(p);
@@ -3469,7 +3532,7 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
if (err) {
/* re-insert swap space back into swap_list */
reinsert_swap_info(p);
- goto out_dput;
+ return err;
}
/*
@@ -3509,15 +3572,21 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
kfree(p->global_cluster);
p->global_cluster = NULL;
free_swap_cluster_info(p);
+ /*
+ * The device is off swap_active_head and no longer WRITEOK, so no
+ * reader can observe these; clearing them here needs no lock.
+ */
p->max = 0;
p->cluster_info = NULL;
- inode = mapping->host;
+ if (swap_file) {
+ struct inode *inode = swap_file->f_mapping->host;
- inode_lock(inode);
- inode->i_flags &= ~S_SWAPFILE;
- inode_unlock(inode);
- filp_close(swap_file, NULL);
+ inode_lock(inode);
+ inode->i_flags &= ~S_SWAPFILE;
+ inode_unlock(inode);
+ filp_close(swap_file, NULL);
+ }
/*
* Clear the SWP_USED flag after all resources are freed so that swapon
@@ -3528,13 +3597,10 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
p->flags = 0;
spin_unlock(&swap_lock);
- err = 0;
atomic_inc(&proc_poll_event);
wake_up_interruptible(&proc_poll_wait);
-out_dput:
- filp_close(victim, NULL);
- return err;
+ return 0;
}
#ifdef CONFIG_PROC_FS
@@ -4501,6 +4567,35 @@ static int xswap_create(int prio)
spin_unlock(&swap_lock);
return error;
}
+
+/* Tear down a file-less xswap device by its swap type. */
+static int xswap_destroy(int type)
+{
+ struct swap_info_struct *p;
+ int err;
+
+ p = swap_type_to_info(type);
+ if (!p)
+ return -EINVAL;
+
+ spin_lock(&swap_lock);
+ if (!(p->flags & SWP_WRITEOK) || !(p->flags & SWP_XSWAP)) {
+ spin_unlock(&swap_lock);
+ return -EINVAL;
+ }
+ /* Refuse swapoff while the device is pinned for hibernation */
+ if (p->flags & SWP_HIBERNATION) {
+ spin_unlock(&swap_lock);
+ return -EBUSY;
+ }
+
+ err = swap_info_remove(p);
+ spin_unlock(&swap_lock);
+ if (err)
+ return err;
+
+ return __swapoff(p);
+}
#endif /* CONFIG_XSWAP */
SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 10/14] mm, swap: require zswap for xswap devices
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
` (8 preceding siblings ...)
2026-10-03 0:31 ` [PATCH v4 09/14] mm, swap: refactor swapoff and add xswap_destroy Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-03 0:31 ` [PATCH v4 11/14] mm, swap: do not give an xswap slot to a cgroup that cannot zswap Baoquan He
` (4 subsequent siblings)
14 siblings, 0 replies; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
xswap pages live only in zswap. Without zswap, swapout cannot free the page,
but still takes a swap entry. So the device consumes swap entries without
freeing memory. Fail to create a device when zswap is unavailable.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
Acked-by: Chris Li <chrisl@kernel.org>
---
mm/swapfile.c | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 0c8b716ec5ba..8c8e99ea403e 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -4504,6 +4504,10 @@ static int xswap_create(int prio)
(prio < 0 || prio > SWAP_FLAG_PRIO_MASK))
return -EINVAL;
+ /* xswap has no backing store, it relies on zswap. */
+ if (!zswap_is_enabled())
+ return -EOPNOTSUPP;
+
si = alloc_swap_info();
if (IS_ERR(si))
return PTR_ERR(si);
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 11/14] mm, swap: do not give an xswap slot to a cgroup that cannot zswap
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
` (9 preceding siblings ...)
2026-10-03 0:31 ` [PATCH v4 10/14] mm, swap: require zswap for xswap devices Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-03 0:31 ` [PATCH v4 12/14] mm, swap: widen swap_info_struct max/pages to unsigned long Baoquan He
` (3 subsequent siblings)
14 siblings, 0 replies; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
An xswap slot holds only a zswap-backed page: it reserves no storage of
its own. A cgroup that cannot take another zswap page still gets an
xswap slot on swapout, but zswap refuses the page, so the slot cannot
hold it.
Skip xswap devices during allocation when the folio's cgroup cannot
zswap, and allocate from a physical swap device instead. That routes the
folio to a device that can store it, and avoids taking an xswap slot
that zswap will not fill.
The test is obj_cgroup_may_zswap(), which is false both when
memory.zswap.max is 0 in the hierarchy and when the hierarchy is already
at that limit. So the routing follows how full the pool is, and not only
how the cgroup is configured; it is stricter than zswap_store(), which
stores anyway if its shrink found nothing to free.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 47 ++++++++++++++++++++++++++++++++++++++++-------
1 file changed, 40 insertions(+), 7 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 8c8e99ea403e..818b5a147b73 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -204,6 +204,8 @@ static DEFINE_SPINLOCK(swap_lock);
static unsigned int nr_swapfiles;
atomic_long_t nr_swap_pages;
atomic_t nr_real_swapfiles;
+/* Active xswap devices, as in nr_real_swapfiles. */
+static atomic_t nr_xswap_files;
/*
* Some modules use swappable objects and may try to swap them out under
* memory pressure (via the shrinker). Before doing so, they may wish to
@@ -1385,7 +1387,9 @@ static void del_from_avail_list(struct swap_info_struct *si, bool swapoff)
lockdep_assert_held(&si->lock);
si->flags &= ~SWP_WRITEOK;
/* Count active devices, not merely those on the avail list. */
- if (!(si->flags & SWP_XSWAP))
+ if (si->flags & SWP_XSWAP)
+ atomic_sub(1, &nr_xswap_files);
+ else
atomic_sub(1, &nr_real_swapfiles);
atomic_long_or(SWAP_USAGE_OFFLIST_BIT, &si->inuse_pages);
} else {
@@ -1444,8 +1448,12 @@ static void add_to_avail_list(struct swap_info_struct *si, bool swapon)
}
plist_add(&si->avail_list, &swap_avail_head);
- if (swapon && !(si->flags & SWP_XSWAP))
- atomic_add(1, &nr_real_swapfiles);
+ if (swapon) {
+ if (si->flags & SWP_XSWAP)
+ atomic_add(1, &nr_xswap_files);
+ else
+ atomic_add(1, &nr_real_swapfiles);
+ }
skip:
spin_unlock(&swap_avail_lock);
@@ -1543,7 +1551,7 @@ static bool get_swap_device_info(struct swap_info_struct *si)
* Fast path try to get swap entries with specified order from current
* CPU's swap entry pool (a cluster).
*/
-static bool swap_alloc_fast(struct folio *folio)
+static bool swap_alloc_fast(struct folio *folio, bool may_zswap)
{
unsigned int order = folio_order(folio);
struct swap_cluster_info *ci;
@@ -1564,6 +1572,11 @@ static bool swap_alloc_fast(struct folio *folio)
rcu_read_unlock();
return false;
}
+ if (!may_zswap && (si->flags & SWP_XSWAP)) {
+ put_swap_device(si);
+ rcu_read_unlock();
+ return false;
+ }
ci = swap_cluster_lock(si, offset);
if (cluster_is_usable(ci, order)) {
@@ -1580,13 +1593,19 @@ static bool swap_alloc_fast(struct folio *folio)
}
/* Rotate the device and switch to a new cluster */
-static void swap_alloc_slow(struct folio *folio)
+static void swap_alloc_slow(struct folio *folio, bool may_zswap)
{
struct swap_info_struct *si, *next;
spin_lock(&swap_avail_lock);
start_over:
plist_for_each_entry_safe(si, next, &swap_avail_head, avail_list) {
+ /*
+ * Do not rotate a skipped device: the walk follows the
+ * rotation and would spin forever.
+ */
+ if (!may_zswap && (si->flags & SWP_XSWAP))
+ continue;
/* Rotate the device and switch to a new cluster */
plist_requeue(&si->avail_list, &swap_avail_head);
spin_unlock(&swap_avail_lock);
@@ -1925,10 +1944,24 @@ int folio_alloc_swap(struct folio *folio)
{
unsigned int order = folio_order(folio);
unsigned int size = 1 << order;
+ struct obj_cgroup *objcg;
+ /* True unless the gate below finds an xswap device to route away from. */
+ bool may_zswap = true;
VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio);
VM_BUG_ON_FOLIO(!folio_test_uptodate(folio), folio);
+ /*
+ * A folio that zswap will not take must not get an xswap slot. With
+ * no xswap device to route away from, skip the lookup.
+ */
+ if (atomic_read(&nr_xswap_files)) {
+ objcg = get_obj_cgroup_from_folio(folio);
+ may_zswap = !objcg || obj_cgroup_may_zswap(objcg);
+ if (objcg)
+ obj_cgroup_put(objcg);
+ }
+
if (order) {
/*
* Reject large allocation when THP_SWAP is disabled. Check below
@@ -1949,8 +1982,8 @@ int folio_alloc_swap(struct folio *folio)
again:
local_lock(&percpu_swap_cluster.lock);
- if (!swap_alloc_fast(folio))
- swap_alloc_slow(folio);
+ if (!swap_alloc_fast(folio, may_zswap))
+ swap_alloc_slow(folio, may_zswap);
local_unlock(&percpu_swap_cluster.lock);
if (!order && unlikely(!folio_test_swapcache(folio))) {
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 12/14] mm, swap: widen swap_info_struct max/pages to unsigned long
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
` (10 preceding siblings ...)
2026-10-03 0:31 ` [PATCH v4 11/14] mm, swap: do not give an xswap slot to a cgroup that cannot zswap Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-03 0:31 ` [PATCH v4 13/14] mm, swap: add debugfs counters for xswap Baoquan He
` (2 subsequent siblings)
14 siblings, 0 replies; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
si->max and si->pages are unsigned int. This limits a swap device to
16 TB with 4 KB pages. read_swap_header() need to clamp the size to that
limit, and xswap_create() has the same clamp.
Change both fields to unsigned long. The clamp in read_swap_header() is
removed. The clamp in xswap_create() becomes a limit on the number of
clusters instead: a cluster index is an unsigned int, so the device can
hold at most UINT_MAX clusters. swapfile_maximum_size still limits how
much a swap area can address.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
include/linux/swap.h | 4 ++--
mm/swapfile.c | 57 +++++++++++++++++++++-----------------------
2 files changed, 29 insertions(+), 32 deletions(-)
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 12cdc00f78f9..c56adbbcfd3b 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -240,7 +240,7 @@ struct swap_info_struct {
signed short prio; /* swap priority of this type */
struct plist_node list; /* entry in swap_active_head */
signed char type; /* strange name for an index */
- unsigned int max; /* size of this swap device */
+ unsigned long max; /* size of this swap device */
struct swap_cluster_info *cluster_info; /* array, one entry per cluster */
#ifdef CONFIG_XSWAP
struct vm_struct *cluster_vm; /* VM_SPARSE area for cluster_info */
@@ -255,7 +255,7 @@ struct swap_info_struct {
/* list of cluster that contains at least one free slot */
struct list_head frag_clusters[SWAP_NR_ORDERS];
/* list of cluster that are fragmented or contented */
- unsigned int pages; /* total of usable pages of swap */
+ unsigned long pages; /* total of usable pages of swap */
atomic_long_t inuse_pages; /* number of those currently in use */
struct swap_sequential_cluster *global_cluster; /* Use one global cluster for rotating device */
spinlock_t global_cluster_lock; /* Serialize usage of global cluster */
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 818b5a147b73..f823cebb7937 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -541,10 +541,10 @@ static inline unsigned int cluster_index(struct swap_info_struct *si,
return ci - si->cluster_info;
}
-static inline unsigned int cluster_offset(struct swap_info_struct *si,
- struct swap_cluster_info *ci)
+static inline unsigned long cluster_offset(struct swap_info_struct *si,
+ struct swap_cluster_info *ci)
{
- return cluster_index(si, ci) * SWAPFILE_CLUSTER;
+ return (unsigned long)cluster_index(si, ci) * SWAPFILE_CLUSTER;
}
static void swap_cluster_free_table_folio_rcu_cb(struct rcu_head *head)
@@ -948,7 +948,7 @@ static int swap_cluster_setup_bad_slot(struct swap_info_struct *si,
/* si->max may got shrunk by swap swap_activate() */
if (offset >= si->max && !mask) {
- pr_debug("Ignoring bad slot %u (max: %u)\n", offset, si->max);
+ pr_debug("Ignoring bad slot %u (max: %lu)\n", offset, si->max);
return 0;
}
/*
@@ -1114,11 +1114,12 @@ static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
}
/* Try use a new cluster for current CPU and allocate from it. */
-static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,
- struct swap_cluster_info *ci,
- struct folio *folio, unsigned long offset)
+static unsigned long alloc_swap_scan_cluster(struct swap_info_struct *si,
+ struct swap_cluster_info *ci,
+ struct folio *folio,
+ unsigned long offset)
{
- unsigned int next = SWAP_ENTRY_INVALID, found = SWAP_ENTRY_INVALID;
+ unsigned long next = SWAP_ENTRY_INVALID, found = SWAP_ENTRY_INVALID;
unsigned long start = ALIGN_DOWN(offset, SWAPFILE_CLUSTER);
unsigned int order = likely(folio) ? folio_order(folio) : 0;
unsigned long end = start + SWAPFILE_CLUSTER;
@@ -1169,12 +1170,12 @@ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,
return found;
}
-static unsigned int alloc_swap_scan_list(struct swap_info_struct *si,
- struct list_head *list,
- struct folio *folio,
- bool scan_all)
+static unsigned long alloc_swap_scan_list(struct swap_info_struct *si,
+ struct list_head *list,
+ struct folio *folio,
+ bool scan_all)
{
- unsigned int found = SWAP_ENTRY_INVALID;
+ unsigned long found = SWAP_ENTRY_INVALID;
do {
struct swap_cluster_info *ci = isolate_lock_cluster(si, list);
@@ -1261,7 +1262,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
{
struct swap_cluster_info *ci;
unsigned int order = likely(folio) ? folio_order(folio) : 0;
- unsigned int offset = SWAP_ENTRY_INVALID, found = SWAP_ENTRY_INVALID;
+ unsigned long offset = SWAP_ENTRY_INVALID, found = SWAP_ENTRY_INVALID;
/*
* Swapfile is not block device so unable
@@ -1556,7 +1557,7 @@ static bool swap_alloc_fast(struct folio *folio, bool may_zswap)
unsigned int order = folio_order(folio);
struct swap_cluster_info *ci;
struct swap_info_struct *si;
- unsigned int offset;
+ unsigned long offset;
/*
* Once allocated, swap_info_struct will never be completely freed,
@@ -3010,8 +3011,8 @@ static int unuse_mm(struct mm_struct *mm, unsigned int type)
* Return 0 if there are no inuse entries after prev till end of
* the map.
*/
-static unsigned int find_next_to_unuse(struct swap_info_struct *si,
- unsigned int prev)
+static unsigned long find_next_to_unuse(struct swap_info_struct *si,
+ unsigned long prev)
{
struct swap_cluster_info *ci;
unsigned long i, cluster_end, end;
@@ -3081,7 +3082,7 @@ static int try_to_unuse(unsigned int type)
struct swap_info_struct *si = swap_info[type];
struct folio *folio;
swp_entry_t entry;
- unsigned int i;
+ unsigned long i;
if (!swap_usage_in_pages(si))
goto success;
@@ -3950,12 +3951,8 @@ static unsigned long read_swap_header(struct swap_info_struct *si,
pr_warn("Truncating oversized swap area, only using %luk out of %luk\n",
K(maxpages), K(last_page));
}
- if (maxpages > last_page) {
+ if (maxpages > last_page)
maxpages = last_page + 1;
- /* p->max is an unsigned int: don't overflow it */
- if ((unsigned int)maxpages == 0)
- maxpages = UINT_MAX;
- }
if (!maxpages)
return 0;
@@ -4550,9 +4547,9 @@ static int xswap_create(int prio)
ram = totalram_pages();
maxpages = min_t(unsigned long, ram * 2, swapfile_maximum_size);
- /* si->max is an unsigned int: don't overflow it. */
- if (maxpages > UINT_MAX)
- maxpages = UINT_MAX;
+ /* A cluster index is an unsigned int, so cap the cluster count. */
+ if (maxpages / SWAPFILE_CLUSTER > UINT_MAX)
+ maxpages = (unsigned long)UINT_MAX * SWAPFILE_CLUSTER;
/* Cluster-aligned, so no cluster holds a slot past si->max. */
if (maxpages > SWAPFILE_CLUSTER)
maxpages = rounddown(maxpages, SWAPFILE_CLUSTER);
@@ -4585,8 +4582,8 @@ static int xswap_create(int prio)
si->avail_list.prio = -si->prio;
/* si->swap_file stays NULL: this is a file-less device */
enable_swap_info(si);
- pr_info("xswap: adding extendable swap type %d (prio %d, %u pages, max %lu)\n",
- si->type, prio, si->pages, maxpages);
+ pr_info("xswap: adding extendable swap type %d (prio %d, %lu pages)\n",
+ si->type, prio, si->pages);
mutex_unlock(&swapon_mutex);
atomic_inc(&proc_poll_event);
wake_up_interruptible(&proc_poll_wait);
@@ -4739,7 +4736,7 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
goto bad_swap_unlock_inode;
}
if (si->pages != si->max - 1) {
- pr_err("swap:%u != (max:%u - 1)\n", si->pages, si->max);
+ pr_err("swap:%lu != (max:%lu - 1)\n", si->pages, si->max);
error = -EINVAL;
goto bad_swap_unlock_inode;
}
@@ -4830,7 +4827,7 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
/* Sets SWP_WRITEOK, resurrect the percpu ref, expose the swap device */
enable_swap_info(si);
- pr_info("Adding %uk swap on %s. Priority:%d extents:%d across:%lluk %s%s%s%s\n",
+ pr_info("Adding %luk swap on %s. Priority:%d extents:%d across:%lluk %s%s%s%s\n",
K(si->pages), name->name, si->prio, nr_extents,
K((unsigned long long)span),
(si->flags & SWP_SOLIDSTATE) ? "SS" : "",
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 13/14] mm, swap: add debugfs counters for xswap
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
` (11 preceding siblings ...)
2026-10-03 0:31 ` [PATCH v4 12/14] mm, swap: widen swap_info_struct max/pages to unsigned long Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-03 0:31 ` [PATCH v4 14/14] mm, swap: let the command line set the xswap device size Baoquan He
2026-10-07 16:28 ` [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Klara Modin
14 siblings, 0 replies; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
The mapped range and the number of grow and shrink operations are not
visible from userspace, which makes the grow and shrink thresholds hard
to check: nothing says whether the range settles, how often it moves, or
where it lands.
Put them in debugfs rather than sysfs. This is not an interface and does
not want to be one, so it does not belong in the ABI:
# cat /sys/kernel/debug/xswap/type0
clusters_max 2038783
clusters_mapped 292
usage_pages 91456
grows 3
shrinks 1
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
include/linux/swap.h | 3 ++
mm/swapfile.c | 80 ++++++++++++++++++++++++++++++++++++++++++++
2 files changed, 83 insertions(+)
diff --git a/include/linux/swap.h b/include/linux/swap.h
index c56adbbcfd3b..ae2e49386443 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -248,6 +248,9 @@ struct swap_info_struct {
unsigned long nr_clusters_mapped; /* currently mapped cluster count */
struct work_struct xswap_shrink_work; /* deferred shrink trigger */
struct mutex xswap_lock; /* serialize map/unmap operations */
+ struct dentry *xswap_debugfs; /* mapped range and counters */
+ atomic_long_t xswap_grows; /* grow operations so far */
+ atomic_long_t xswap_shrinks; /* shrink operations so far */
#endif
struct list_head free_clusters; /* free clusters list */
struct list_head full_clusters; /* full clusters list */
diff --git a/mm/swapfile.c b/mm/swapfile.c
index f823cebb7937..ab040d24bd2f 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -49,6 +49,8 @@
#include <linux/leafops.h>
#include "swap_table.h"
#include "internal.h"
+#include <linux/debugfs.h>
+
#include "swap.h"
#define DEF_SWAP_PRIO -1
@@ -84,6 +86,7 @@
static bool xswap_should_grow(struct swap_info_struct *si);
static unsigned long xswap_grow(struct swap_info_struct *si);
+static void xswap_debugfs_del(struct swap_info_struct *si);
/* Batch size for the fallback unmap; keeps the stack array small. */
#define XSWAP_UNMAP_BATCH_PAGES 16
@@ -3416,6 +3419,7 @@ static void free_swap_cluster_info(struct swap_info_struct *si)
}
/* free_vm_area() drops the mapping without freeing the pages. */
xswap_unmap_clusters(si, 0, si->nr_clusters_mapped);
+ xswap_debugfs_del(si);
free_vm_area(si->cluster_vm);
si->cluster_vm = NULL;
si->cluster_info = NULL;
@@ -4134,6 +4138,7 @@ static unsigned long xswap_grow(struct swap_info_struct *si)
spin_unlock(&ci->lock);
}
+ atomic_long_inc(&si->xswap_grows);
return nr_new;
}
@@ -4370,10 +4375,79 @@ static void xswap_try_shrink(struct swap_info_struct *si)
spin_unlock(&si->lock);
xswap_unmap_clusters_locked(si, start_idx, nr_unmap);
+ atomic_long_inc(&si->xswap_shrinks);
out_unlock:
mutex_unlock(&si->xswap_lock);
}
+
+/* Not an interface: how far the mapped range has moved, for testing. */
+static ssize_t xswap_debugfs_read(struct file *file, char __user *buf,
+ size_t count, loff_t *ppos)
+{
+ struct swap_info_struct *si = file->private_data;
+ struct swap_cluster_info *ci;
+ unsigned long mapped, tail = 0;
+ char tmp[192];
+ int len;
+
+ /*
+ * The same scan the shrink does. xswap_lock is what keeps
+ * cluster_info from being unmapped under it.
+ */
+ mutex_lock(&si->xswap_lock);
+ mapped = READ_ONCE(si->nr_clusters_mapped);
+ if (si->cluster_info) {
+ while (mapped - tail > 1) {
+ ci = &si->cluster_info[mapped - tail - 1];
+ if (READ_ONCE(ci->count) ||
+ READ_ONCE(ci->flags) != CLUSTER_FLAG_FREE)
+ break;
+ tail++;
+ }
+ }
+ len = scnprintf(tmp, sizeof(tmp),
+ "clusters_max %lu\nclusters_mapped %lu\nusage_pages %lu\ntail_free %lu\ngrows %lu\nshrinks %lu\n",
+ READ_ONCE(si->nr_clusters_max), mapped,
+ swap_usage_in_pages(si), tail,
+ atomic_long_read(&si->xswap_grows),
+ atomic_long_read(&si->xswap_shrinks));
+ mutex_unlock(&si->xswap_lock);
+
+ return simple_read_from_buffer(buf, count, ppos, tmp, len);
+}
+
+static const struct file_operations xswap_debugfs_fops = {
+ .read = xswap_debugfs_read,
+ .open = simple_open,
+ .llseek = default_llseek,
+};
+
+static struct dentry *xswap_debugfs_root;
+
+static void xswap_debugfs_add(struct swap_info_struct *si)
+{
+ char name[16];
+
+ if (!xswap_debugfs_root)
+ xswap_debugfs_root = debugfs_create_dir("xswap", NULL);
+ if (IS_ERR_OR_NULL(xswap_debugfs_root))
+ return;
+
+ /* debugfs_create_file() takes the name as it is; no formatting. */
+ snprintf(name, sizeof(name), "type%d", si->type);
+ si->xswap_debugfs = debugfs_create_file(name, 0444,
+ xswap_debugfs_root, si,
+ &xswap_debugfs_fops);
+ if (IS_ERR(si->xswap_debugfs))
+ si->xswap_debugfs = NULL;
+}
+
+static void xswap_debugfs_del(struct swap_info_struct *si)
+{
+ debugfs_remove(si->xswap_debugfs);
+ si->xswap_debugfs = NULL;
+}
#endif /* CONFIG_XSWAP */
static int setup_swap_clusters_info(struct swap_info_struct *si,
@@ -4568,10 +4642,16 @@ static int xswap_create(int prio)
*/
si->ops = &swap_bdev_ops;
+ si->xswap_debugfs = NULL;
+ atomic_long_set(&si->xswap_grows, 0);
+ atomic_long_set(&si->xswap_shrinks, 0);
+
error = setup_swap_clusters_info(si, NULL, maxpages);
if (error)
goto bad_swap;
+ xswap_debugfs_add(si);
+
error = zswap_swapon(si->type, si->max);
if (error)
goto bad_swap;
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* [PATCH v4 14/14] mm, swap: let the command line set the xswap device size
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
` (12 preceding siblings ...)
2026-10-03 0:31 ` [PATCH v4 13/14] mm, swap: add debugfs counters for xswap Baoquan He
@ 2026-10-03 0:31 ` Baoquan He
2026-10-07 16:28 ` [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Klara Modin
14 siblings, 0 replies; 21+ messages in thread
From: Baoquan He @ 2026-10-03 0:31 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
yosry, shikemeng, chengming.zhou, baoquan.he, david, linux-kernel,
kunwu.chan, klarasmodin, Baoquan He
A new xswap device is always sized at twice RAM. Add xswap.max= to pick
the size of devices created after it:
xswap.max=300% three times RAM
xswap.max=8G eight GiB
Twice RAM is right for most users, but a machine with a lot of memory and
a small workingset gets a SwapTotal much larger than needed.
The size is resolved at device creation, not at parameter parsing: a
percentage is relative to RAM, and totalram_pages() is not ready that
early. A size that does not fit, or is too small to hold anything, is
refused, not clamped.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 86 +++++++++++++++++++++++++++++++++++++++++++++------
1 file changed, 76 insertions(+), 10 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index ab040d24bd2f..bf95663c2593 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -164,6 +164,63 @@ static struct attribute *xswap_attrs[] = {
NULL,
};
+/* Size of a new device, from xswap.max=. A percentage unless bytes is set. */
+static unsigned int xswap_size_pct = 200;
+static unsigned long xswap_size_bytes;
+
+static int __init xswap_max_setup(char *str)
+{
+ unsigned int pct;
+ char *end;
+
+ end = strchr(str, '%');
+ if (end) {
+ if (end[1]) {
+ pr_err("xswap: bad xswap.max=%s\n", str);
+ return 1;
+ }
+ *end = '\0'; /* writable copy, see parse_args() */
+ if (kstrtouint(str, 0, &pct) || !pct) {
+ pr_err("xswap: bad xswap.max=%s%%\n", str);
+ return 1;
+ }
+ xswap_size_pct = pct;
+ xswap_size_bytes = 0;
+ return 1;
+ }
+
+ xswap_size_bytes = memparse(str, &end);
+ if (*end || xswap_size_bytes < PAGE_SIZE) {
+ pr_err("xswap: bad xswap.max=%s\n", str);
+ xswap_size_bytes = 0;
+ return 1;
+ }
+ xswap_size_pct = 0;
+ return 1;
+}
+__setup("xswap.max=", xswap_max_setup);
+
+/* The most a device can address, whatever xswap.max= asked for. */
+static unsigned long xswap_size_limit(void)
+{
+ unsigned long limit = swapfile_maximum_size;
+
+ /* A cluster index is an unsigned int. */
+ if (limit / SWAPFILE_CLUSTER > UINT_MAX)
+ limit = (unsigned long)UINT_MAX * SWAPFILE_CLUSTER;
+
+ return limit;
+}
+
+static unsigned long xswap_size_pages(void)
+{
+ if (xswap_size_bytes)
+ return xswap_size_bytes >> PAGE_SHIFT;
+
+ return div_u64((unsigned long long)totalram_pages() * xswap_size_pct,
+ 100);
+}
+
static const struct attribute_group xswap_attr_group = {
.attrs = xswap_attrs,
};
@@ -4601,7 +4658,7 @@ static int setup_swap_clusters_info(struct swap_info_struct *si,
static int xswap_create(int prio)
{
struct swap_info_struct *si;
- unsigned long ram, maxpages;
+ unsigned long maxpages;
int error;
if (prio != DEF_XSWAP_PRIO &&
@@ -4612,6 +4669,23 @@ static int xswap_create(int prio)
if (!zswap_is_enabled())
return -EOPNOTSUPP;
+ /* Refuse rather than clamp: xswap.max= should give what was asked. */
+ maxpages = xswap_size_pages();
+ if (maxpages > xswap_size_limit()) {
+ pr_err("xswap: xswap.max is larger than the device can address\n");
+ return -EINVAL;
+ }
+ /*
+ * Below one cluster the device is not cluster-aligned, so its only
+ * cluster would hold slots past si->max. Nothing masks those, and a
+ * slot allocated there is past the end of the device: it can never be
+ * read back or given up again.
+ */
+ if (maxpages < SWAPFILE_CLUSTER) {
+ pr_err("xswap: xswap.max is smaller than one cluster\n");
+ return -EINVAL;
+ }
+
si = alloc_swap_info();
if (IS_ERR(si))
return PTR_ERR(si);
@@ -4619,16 +4693,8 @@ static int xswap_create(int prio)
INIT_WORK(&si->discard_work, swap_discard_work);
INIT_WORK(&si->reclaim_work, swap_reclaim_work);
- ram = totalram_pages();
- maxpages = min_t(unsigned long, ram * 2, swapfile_maximum_size);
- /* A cluster index is an unsigned int, so cap the cluster count. */
- if (maxpages / SWAPFILE_CLUSTER > UINT_MAX)
- maxpages = (unsigned long)UINT_MAX * SWAPFILE_CLUSTER;
/* Cluster-aligned, so no cluster holds a slot past si->max. */
- if (maxpages > SWAPFILE_CLUSTER)
- maxpages = rounddown(maxpages, SWAPFILE_CLUSTER);
- if (maxpages < 2)
- maxpages = 2;
+ maxpages = rounddown(maxpages, SWAPFILE_CLUSTER);
si->bdev = NULL;
si->flags |= SWP_XSWAP | SWP_SOLIDSTATE;
--
2.54.0
^ permalink raw reply related [flat|nested] 21+ messages in thread
* Re: [PATCH v4 01/14] mm: xswap support for zswap
2026-10-03 0:31 ` [PATCH v4 01/14] mm: xswap support for zswap Baoquan He
@ 2026-10-07 4:22 ` KunWu Chan
0 siblings, 0 replies; 21+ messages in thread
From: KunWu Chan @ 2026-10-07 4:22 UTC (permalink / raw)
To: Baoquan He
Cc: linux-mm, akpm, chrisl, kasong, hannes, nphamcs, baohua,
youngjun.park, yosry, shikemeng, chengming.zhou, baoquan.he,
david, linux-kernel, klarasmodin
On Sat, Oct 3, 2026 at 8:31 AM Baoquan He <hebaoquan@kylinos.cn> wrote:
>
> From: Chris Li <chrisl@kernel.org>
>
> Introduce extendable swap device support - xswap.
>
> An xswap device has no backing storage and no swap data section, so
> it wastes no disk space. Creation is via a sysfs interface added in a
> later patch.
>
> Zswap writeback is gated on whether a real (non-xswap) swap device is
> active. nr_real_swapfiles counts such devices and is maintained at
> swapon/swapoff only, so the gate reflects "a device exists to write
> back to" rather than "a device currently has free slots". This keeps
> writeback working even when the real swap device is full, and avoids
> a double decrement when a full device is swapped off.
>
> An xswap entry is not added to the zswap writeback LRU: with no backing
> store there is nothing to write it back to. zswap_lru_del() tolerates
> that, since an entry that was never added is on no list.
>
> Co-developed-by: Baoquan He <hebaoquan@kylinos.cn>
> Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
> Signed-off-by: Chris Li <chrisl@kernel.org>
> ---
> include/linux/swap.h | 2 ++
> mm/page_io.c | 19 +++++++++++++++++++
> mm/swap_state.c | 9 +++++++++
> mm/swapfile.c | 28 ++++++++++++++++++++++++----
> mm/zswap.c | 11 +++++++++--
> 5 files changed, 63 insertions(+), 6 deletions(-)
>
> diff --git a/include/linux/swap.h b/include/linux/swap.h
> index 4f686709d71f..9073d29377d8 100644
> --- a/include/linux/swap.h
> +++ b/include/linux/swap.h
> @@ -201,6 +201,7 @@ enum {
> SWP_STABLE_WRITES = (1 << 11), /* no overwrite PG_writeback pages */
> SWP_SYNCHRONOUS_IO = (1 << 12), /* synchronous IO is efficient */
> SWP_HIBERNATION = (1 << 13), /* pinned for hibernation */
> + SWP_XSWAP = (1 << 14), /* extendable swap device */
> /* add others here before... */
> };
>
> @@ -378,6 +379,7 @@ void free_folio_and_swap_cache(struct folio *folio);
> void free_pages_and_swap_cache(struct encoded_page **, int);
> /* linux/mm/swapfile.c */
> extern atomic_long_t nr_swap_pages;
> +extern atomic_t nr_real_swapfiles;
> extern long total_swap_pages;
> extern atomic_t nr_rotate_swap;
>
> diff --git a/mm/page_io.c b/mm/page_io.c
> index c6824fcd483e..25fa9b82ed46 100644
> --- a/mm/page_io.c
> +++ b/mm/page_io.c
> @@ -248,6 +248,15 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
> }
> rcu_read_unlock();
>
> + /*
> + * xswap has no backing store: keep the folio. ctx->sis is not set
> + * yet, so look the device up from the entry.
> + */
> + if (unlikely(__swap_entry_to_info(folio->swap)->flags & SWP_XSWAP)) {
> + folio_mark_dirty(folio);
> + return AOP_WRITEPAGE_ACTIVATE;
> + }
> +
> __swap_writeout(ctx, folio);
> return 0;
> out_unlock:
> @@ -486,6 +495,16 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct folio *folio)
> if (zswap_load(folio) != -ENOENT)
> goto finish;
>
> + if (unlikely(sis->flags & SWP_XSWAP)) {
> + /*
> + * An xswap entry only ever lives in zswap, so zswap_load()
> + * must have found it. Unlock and let the caller retry.
> + */
> + WARN_ON_ONCE(1);
> + folio_unlock(folio);
> + goto finish;
> + }
> +
> /* We have to read from slower devices. Increase zswap protection. */
> zswap_folio_swapin(folio);
> swap_add_folio(ctx, folio, READ);
> diff --git a/mm/swap_state.c b/mm/swap_state.c
> index ebef568cd62e..3a4e9e447b0b 100644
> --- a/mm/swap_state.c
> +++ b/mm/swap_state.c
> @@ -884,6 +884,10 @@ struct folio *swap_cluster_readahead(swp_entry_t entry, gfp_t gfp_mask,
> struct blk_plug plug;
> swp_entry_t ra_entry;
>
> + /* xswap entries live only in zswap; readahead does not help. */
> + if (si->flags & SWP_XSWAP)
> + goto skip;
> +
> mask = swapin_nr_pages(offset) - 1;
> if (!mask)
> goto skip;
> @@ -969,6 +973,7 @@ static int swap_vma_ra_win(struct vm_fault *vmf, unsigned long *start,
> static struct folio *swap_vma_readahead(swp_entry_t targ_entry, gfp_t gfp_mask,
> struct mempolicy *mpol, pgoff_t targ_ilx, struct vm_fault *vmf)
> {
> + struct swap_info_struct *si = __swap_entry_to_info(targ_entry);
> struct swap_io_ctx ctx = {};
> struct blk_plug plug;
> struct folio *folio;
> @@ -977,6 +982,10 @@ static struct folio *swap_vma_readahead(swp_entry_t targ_entry, gfp_t gfp_mask,
> unsigned long start, end, addr;
> pgoff_t ilx = targ_ilx;
>
> + /* xswap entries live only in zswap; readahead does not help. */
> + if (si->flags & SWP_XSWAP)
> + goto skip;
> +
> win = swap_vma_ra_win(vmf, &start, &end);
> if (win == 1)
> goto skip;
> diff --git a/mm/swapfile.c b/mm/swapfile.c
> index c3288910b3e3..32c3133e1211 100644
> --- a/mm/swapfile.c
> +++ b/mm/swapfile.c
> @@ -66,6 +66,7 @@ static void move_cluster(struct swap_info_struct *si,
> static DEFINE_SPINLOCK(swap_lock);
> static unsigned int nr_swapfiles;
> atomic_long_t nr_swap_pages;
> +atomic_t nr_real_swapfiles;
> /*
> * Some modules use swappable objects and may try to swap them out under
> * memory pressure (via the shrinker). Before doing so, they may wish to
> @@ -733,7 +734,8 @@ static void free_cluster(struct swap_info_struct *si, struct swap_cluster_info *
> /*
> * If the swap is discardable, prepare discard the cluster
> * instead of free it immediately. The cluster will be freed
> - * after discard.
> + * after discard. xswap has no bdev and never sets
> + * SWP_PAGE_DISCARD, so it always takes the free path below.
> */
> if ((si->flags & (SWP_WRITEOK | SWP_PAGE_DISCARD)) ==
> (SWP_WRITEOK | SWP_PAGE_DISCARD)) {
> @@ -1216,6 +1218,9 @@ static void del_from_avail_list(struct swap_info_struct *si, bool swapoff)
> */
> lockdep_assert_held(&si->lock);
> si->flags &= ~SWP_WRITEOK;
> + /* Count active devices, not merely those on the avail list. */
> + if (!(si->flags & SWP_XSWAP))
> + atomic_sub(1, &nr_real_swapfiles);
> atomic_long_or(SWAP_USAGE_OFFLIST_BIT, &si->inuse_pages);
> } else {
> /*
> @@ -1273,6 +1278,8 @@ static void add_to_avail_list(struct swap_info_struct *si, bool swapon)
> }
>
> plist_add(&si->avail_list, &swap_avail_head);
> + if (swapon && !(si->flags & SWP_XSWAP))
> + atomic_add(1, &nr_real_swapfiles);
>
> skip:
> spin_unlock(&swap_avail_lock);
> @@ -3268,7 +3275,8 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
>
> destroy_swap_extents(p, p->swap_file);
>
> - if (!(p->flags & SWP_SOLIDSTATE))
> + if (!(p->flags & SWP_XSWAP) &&
> + !(p->flags & SWP_SOLIDSTATE))
> atomic_dec(&nr_rotate_swap);
>
> mutex_lock(&swapon_mutex);
> @@ -3378,6 +3386,19 @@ static void swap_stop(struct seq_file *swap, void *v)
> mutex_unlock(&swapon_mutex);
> }
>
> +static const char *swap_type_str(struct swap_info_struct *si)
> +{
> + struct file *file = si->swap_file;
> +
> + if (si->flags & SWP_XSWAP)
> + return "xswap\t";
> +
> + if (S_ISBLK(file_inode(file)->i_mode))
> + return "partition";
> +
> + return "file\t";
> +}
> +
> static int swap_show(struct seq_file *swap, void *v)
> {
> struct swap_info_struct *si = v;
> @@ -3397,8 +3418,7 @@ static int swap_show(struct seq_file *swap, void *v)
> len = seq_file_path(swap, file, " \t\n\\");
> seq_printf(swap, "%*s%s\t%lu\t%s%lu\t%s%d\n",
> len < 40 ? 40 - len : 1, " ",
> - S_ISBLK(file_inode(file)->i_mode) ?
> - "partition" : "file\t",
> + swap_type_str(si),
> bytes, bytes < 10000000 ? "\t" : "",
> inuse, inuse < 10000000 ? "\t" : "",
> si->prio);
> diff --git a/mm/zswap.c b/mm/zswap.c
> index ae19e301fced..7fe25f0b157c 100644
> --- a/mm/zswap.c
> +++ b/mm/zswap.c
> @@ -1018,6 +1018,11 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
> if (IS_ERR_OR_NULL(si))
> return -ENOENT;
>
> + if (si->flags & SWP_XSWAP) {
> + put_swap_device(si);
> + return -EINVAL;
> + }
> +
> mpol = get_task_policy(current);
> folio = __swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol,
> NO_INTERLEAVE_INDEX);
> @@ -1511,7 +1516,9 @@ static bool zswap_store_page(struct folio *folio, long index,
> entry->referenced = true;
> if (entry->length) {
> INIT_LIST_HEAD(&entry->lru);
> - zswap_lru_add(entry);
> + /* No backing store: nothing to write these back to. */
> + if (!(__swap_entry_to_info(page_swpentry)->flags & SWP_XSWAP))
> + zswap_lru_add(entry);
> }
>
> return true;
> @@ -1581,7 +1588,7 @@ bool zswap_store(struct folio *folio)
> zswap_pool_put(pool);
> put_objcg:
> obj_cgroup_put(objcg);
> - if (!ret && zswap_pool_reached_full)
> + if (!ret && zswap_pool_reached_full && atomic_read(&nr_real_swapfiles))
> queue_work(shrink_wq, &zswap_shrink_work);
> check_old:
> /*
> --
> 2.54.0
>
Hi Baoquan,
I checked the xswap handling across the swap I/O, readahead, and
zswap writeback paths.
The SWP_XSWAP handling is consistent with the no-backing-store
invariant: swap_writeout() keeps the folio dirty instead of issuing
swap I/O, swap_read_folio() treats a zswap miss as an invariant
violation rather than falling back to backing-store I/O, and both
readahead paths are bypassed because there is no backing device to
read from.
The nr_real_swapfiles accounting also correctly tracks real swap
devices independently of xswap, so zswap writeback remains enabled
as long as there is a real swap device available. Excluding xswap
entries from the zswap writeback LRU is consistent with the same
invariant.
Reviewed-by: Kunwu Chan <chentao@kylinos.cn>
Thanks,
Kunwu
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [PATCH v4 02/14] mm, swap: refactor free_swap_cluster_info to take swap_info_struct
2026-10-03 0:31 ` [PATCH v4 02/14] mm, swap: refactor free_swap_cluster_info to take swap_info_struct Baoquan He
@ 2026-10-07 4:32 ` KunWu Chan
0 siblings, 0 replies; 21+ messages in thread
From: KunWu Chan @ 2026-10-07 4:32 UTC (permalink / raw)
To: Baoquan He
Cc: linux-mm, akpm, chrisl, kasong, hannes, nphamcs, baohua,
youngjun.park, yosry, shikemeng, chengming.zhou, baoquan.he,
david, linux-kernel, klarasmodin
On Sat, Oct 3, 2026 at 8:31 AM Baoquan He <hebaoquan@kylinos.cn> wrote:
>
> Change free_swap_cluster_info() to take struct swap_info_struct* instead
> of (cluster_info, maxpages). It now extracts the fields from si and
> clears si->cluster_info after freeing to avoid a double free on the
> swapon() error path.
>
> The new parameter also lets the xswap path access si->flags in the
> function.
>
> Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
> Reviewed-by: Chris Li <chrisl@kernel.org>
> ---
> mm/swapfile.c | 26 ++++++++++++++------------
> 1 file changed, 14 insertions(+), 12 deletions(-)
>
> diff --git a/mm/swapfile.c b/mm/swapfile.c
> index 32c3133e1211..13ec5c0f3c2f 100644
> --- a/mm/swapfile.c
> +++ b/mm/swapfile.c
> @@ -3144,14 +3144,17 @@ static void wait_for_allocation(struct swap_info_struct *si)
> }
> }
>
> -static void free_swap_cluster_info(struct swap_cluster_info *cluster_info,
> - unsigned long maxpages)
> +static void free_swap_cluster_info(struct swap_info_struct *si)
> {
> + struct swap_cluster_info *cluster_info = si->cluster_info;
> + unsigned long maxpages = si->max;
> struct swap_cluster_info *ci;
> - int i, nr_clusters = DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER);
> + int i, nr_clusters;
>
> if (!cluster_info)
> return;
> +
> + nr_clusters = DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER);
> for (i = 0; i < nr_clusters; i++) {
> ci = cluster_info + i;
> /* Cluster with bad marks count will have a remaining table */
> @@ -3163,6 +3166,7 @@ static void free_swap_cluster_info(struct swap_cluster_info *cluster_info,
> spin_unlock(&ci->lock);
> }
> kvfree(cluster_info);
> + si->cluster_info = NULL;
> }
>
> /*
> @@ -3190,11 +3194,9 @@ static void flush_percpu_swap_cluster(struct swap_info_struct *si)
> SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
> {
> struct swap_info_struct *p = NULL;
> - struct swap_cluster_info *cluster_info;
> struct file *swap_file, *victim;
> struct address_space *mapping;
> struct inode *inode;
> - unsigned int maxpages;
> int err, found = 0;
>
> if (!capable(CAP_SYS_ADMIN))
> @@ -3286,10 +3288,6 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
>
> swap_file = p->swap_file;
> p->swap_file = NULL;
> - maxpages = p->max;
> - cluster_info = p->cluster_info;
> - p->max = 0;
> - p->cluster_info = NULL;
> spin_unlock(&p->lock);
> spin_unlock(&swap_lock);
> arch_swap_invalidate_area(p->type);
> @@ -3297,7 +3295,9 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
> mutex_unlock(&swapon_mutex);
> kfree(p->global_cluster);
> p->global_cluster = NULL;
> - free_swap_cluster_info(cluster_info, maxpages);
> + free_swap_cluster_info(p);
> + p->max = 0;
> + p->cluster_info = NULL;
>
> inode = mapping->host;
>
> @@ -3664,6 +3664,8 @@ static int setup_swap_clusters_info(struct swap_info_struct *si,
> if (!cluster_info)
> goto err;
>
> + si->cluster_info = cluster_info;
> +
> for (i = 0; i < nr_clusters; i++)
> spin_lock_init(&cluster_info[i].lock);
>
> @@ -3727,7 +3729,7 @@ static int setup_swap_clusters_info(struct swap_info_struct *si,
> si->cluster_info = cluster_info;
> return 0;
> err:
> - free_swap_cluster_info(cluster_info, maxpages);
> + free_swap_cluster_info(si);
> return err;
> }
>
> @@ -3949,7 +3951,7 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
> si->global_cluster = NULL;
> inode = NULL;
> destroy_swap_extents(si, swap_file);
> - free_swap_cluster_info(si->cluster_info, si->max);
> + free_swap_cluster_info(si);
> si->cluster_info = NULL;
> /*
> * Clear the SWP_USED flag after all resources are freed so
> --
> 2.54.0
>
Straightforward refactor. Passing si lets the helper obtain the
cluster metadata and its size from swap_info_struct, while
clearing si->cluster_info after freeing it makes the error-path
cleanup consistent with the normal teardown path.
Reviewed-by: Kunwu Chan <chentao@kylinos.cn>
Thanks,
Kunwu
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [PATCH v4 04/14] mm, swap: add sysfs create interface for xswap
2026-10-03 0:31 ` [PATCH v4 04/14] mm, swap: add sysfs create interface for xswap Baoquan He
@ 2026-10-07 6:32 ` KunWu Chan
0 siblings, 0 replies; 21+ messages in thread
From: KunWu Chan @ 2026-10-07 6:32 UTC (permalink / raw)
To: Baoquan He
Cc: linux-mm, akpm, chrisl, kasong, hannes, nphamcs, baohua,
youngjun.park, yosry, shikemeng, chengming.zhou, baoquan.he,
david, linux-kernel, klarasmodin
On Sat, Oct 3, 2026 at 8:32 AM Baoquan He <hebaoquan@kylinos.cn> wrote:
>
> xswap devices have no backing storage, so there is no file to swapon.
> Add a sysfs interface to create them directly:
>
> /sys/kernel/mm/xswap/create write an optional priority to create
> a device, empty for the default
>
> A new device is created with si->max and nr_clusters_max both equal to
> twice RAM: the cluster_info array is a sparse VM_SPARSE area mapped
> lazily, so an idle device costs nothing. The device shows up in
> /proc/swaps as "xswap<N>".
>
> It takes the highest priority by default. Swapout picks the
> highest-priority device that has room, so a real device ahead of an xswap
> one would take the page directly: zswap would still cache it, but the
> slot on the real device is taken either way, and not taking it is the
> point of xswap. A lower priority can be asked for, not a higher one.
>
> Hibernation device discovery skips xswap devices: they have no bdev
> to carry a resume image.
>
> Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
> ---
> mm/swapfile.c | 169 ++++++++++++++++++++++++++++++++++++++++++++++++--
> 1 file changed, 164 insertions(+), 5 deletions(-)
>
> diff --git a/mm/swapfile.c b/mm/swapfile.c
> index 767d3f46877f..9c204e455ef9 100644
> --- a/mm/swapfile.c
> +++ b/mm/swapfile.c
> @@ -16,6 +16,8 @@
> #include <linux/kernel_stat.h>
> #include <linux/swap.h>
> #include <linux/vmalloc.h>
> +#include <linux/kobject.h>
> +#include <linux/sysfs.h>
> #include <linux/pagemap.h>
> #include <linux/namei.h>
> #include <linux/shmem_fs.h>
> @@ -48,6 +50,13 @@
> #include "swap_table.h"
> #include "internal.h"
> #include "swap.h"
> +#define DEF_SWAP_PRIO -1
> +
> +/*
> + * Above every real swap device: one ahead of xswap would take the page
> + * directly, and taking that slot is what xswap exists to avoid.
> + */
> +#define DEF_XSWAP_PRIO SWAP_FLAG_PRIO_MASK
>
> #ifdef CONFIG_XSWAP
> /*
> @@ -66,7 +75,65 @@ static int xswap_map_clusters(struct swap_info_struct *si,
> static void xswap_unmap_clusters(struct swap_info_struct *si,
> unsigned long start_idx, unsigned long nr);
> static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data);
> -#endif
> +
> +static int xswap_create(int prio);
> +
> +static ssize_t xswap_create_store(struct kobject *kobj,
> + struct kobj_attribute *attr,
> + const char *buf, size_t count)
> +{
> + int prio = DEF_XSWAP_PRIO;
> + int err;
> +
> + if (!capable(CAP_SYS_ADMIN))
> + return -EPERM;
> +
> + /* "[<prio>]" is optional; a lower one is allowed, not a higher. */
> + if (*skip_spaces(buf)) {
> + err = kstrtoint(buf, 10, &prio);
> + if (err)
> + return err;
> + }
> +
> + err = xswap_create(prio);
> + if (err < 0)
> + return err;
> +
> + return count;
> +}
> +
> +static struct kobj_attribute xswap_create_attr = __ATTR(create, 0200, NULL,
> + xswap_create_store);
> +
> +static struct attribute *xswap_attrs[] = {
> + &xswap_create_attr.attr,
> + NULL,
> +};
> +
> +static const struct attribute_group xswap_attr_group = {
> + .attrs = xswap_attrs,
> +};
> +
> +static struct kobject *xswap_kobj;
> +
> +static void xswap_sysfs_init(void)
> +{
> + xswap_kobj = kobject_create_and_add("xswap", mm_kobj);
> + if (!xswap_kobj) {
> + pr_err("xswap: failed to create sysfs kobject\n");
> + return;
> + }
> + if (sysfs_create_group(xswap_kobj, &xswap_attr_group)) {
> + pr_err("xswap: failed to create sysfs group\n");
> + kobject_put(xswap_kobj);
> + xswap_kobj = NULL;
> + }
> +}
> +#else /* !CONFIG_XSWAP */
> +static inline void xswap_sysfs_init(void)
> +{
> +}
> +#endif /* CONFIG_XSWAP */
>
> static void swap_range_alloc(struct swap_info_struct *si,
> unsigned int nr_entries);
> @@ -94,7 +161,6 @@ atomic_t nr_real_swapfiles;
> EXPORT_SYMBOL_GPL(nr_swap_pages);
> /* protected with swap_lock. reading in vm_swap_full() doesn't need lock */
> long total_swap_pages;
> -#define DEF_SWAP_PRIO -1
> unsigned long swapfile_maximum_size;
> #ifdef CONFIG_MIGRATION
> bool swap_migration_ad_supported;
> @@ -2275,6 +2341,9 @@ static int __find_hibernation_swap_type(dev_t device, sector_t offset)
>
> if (!(sis->flags & SWP_WRITEOK))
> continue;
> + /* xswap has no bdev to match a resume device */
> + if (sis->flags & SWP_XSWAP)
> + continue;
>
> if (device == sis->bdev->bd_dev) {
> struct swap_extent *se = first_se(sis);
> @@ -2462,6 +2531,8 @@ int find_first_swap(dev_t *device)
>
> if (!(sis->flags & SWP_WRITEOK))
> continue;
> + if (sis->flags & SWP_XSWAP)
> + continue;
> *device = sis->bdev->bd_dev;
> spin_unlock(&swap_lock);
> return type;
> @@ -3425,7 +3496,8 @@ static void *swap_start(struct seq_file *swap, loff_t *pos)
> return SEQ_START_TOKEN;
>
> for (type = 0; (si = swap_type_to_info(type)); type++) {
> - if (!(si->swap_file))
> + if (!si->swap_file &&
> + !((si->flags & SWP_XSWAP) && (si->flags & SWP_WRITEOK)))
> continue;
> if (!--l)
> return si;
> @@ -3446,7 +3518,8 @@ static void *swap_next(struct seq_file *swap, void *v, loff_t *pos)
>
> ++(*pos);
> for (; (si = swap_type_to_info(type)); type++) {
> - if (!(si->swap_file))
> + if (!si->swap_file &&
> + !((si->flags & SWP_XSWAP) && (si->flags & SWP_WRITEOK)))
> continue;
> return si;
> }
> @@ -3488,7 +3561,14 @@ static int swap_show(struct seq_file *swap, void *v)
> inuse = K(swap_usage_in_pages(si));
>
> file = si->swap_file;
> - len = seq_file_path(swap, file, " \t\n\\");
> + if (file)
> + len = seq_file_path(swap, file, " \t\n\\");
> + else {
> + char name[16];
> +
> + len = scnprintf(name, sizeof(name), "xswap%d", si->type);
> + seq_puts(swap, name);
> + }
> seq_printf(swap, "%*s%s\t%lu\t%s%lu\t%s%d\n",
> len < 40 ? 40 - len : 1, " ",
> swap_type_str(si),
> @@ -4011,6 +4091,83 @@ static int setup_swap_clusters_info(struct swap_info_struct *si,
> return err;
> }
>
> +#ifdef CONFIG_XSWAP
> +/* Create a file-less xswap device, sized at twice RAM. */
> +static int xswap_create(int prio)
> +{
> + struct swap_info_struct *si;
> + unsigned long ram, maxpages;
> + int error;
> +
> + if (prio != DEF_XSWAP_PRIO &&
> + (prio < 0 || prio > SWAP_FLAG_PRIO_MASK))
> + return -EINVAL;
> +
> + si = alloc_swap_info();
> + if (IS_ERR(si))
> + return PTR_ERR(si);
> +
> + INIT_WORK(&si->discard_work, swap_discard_work);
> + INIT_WORK(&si->reclaim_work, swap_reclaim_work);
> +
> + ram = totalram_pages();
> + maxpages = min_t(unsigned long, ram * 2, swapfile_maximum_size);
> + /* si->max is an unsigned int: don't overflow it. */
> + if (maxpages > UINT_MAX)
> + maxpages = UINT_MAX;
> + /* Cluster-aligned, so no cluster holds a slot past si->max. */
> + if (maxpages > SWAPFILE_CLUSTER)
> + maxpages = rounddown(maxpages, SWAPFILE_CLUSTER);
> + if (maxpages < 2)
> + maxpages = 2;
> +
> + si->bdev = NULL;
> + si->flags |= SWP_XSWAP | SWP_SOLIDSTATE;
> + si->max = maxpages;
> + si->pages = maxpages - 1;
> + /*
> + * No backing file: setup_swap_extents() is only reachable from the
> + * file-backed swapon() path, so set ops here. Only ops->flags is
> + * used, by may_enter_fs(); the IO methods are never called because
> + * swap_writeout()/swap_read_folio() short circuit xswap.
> + */
> + si->ops = &swap_bdev_ops;
> +
> + error = setup_swap_clusters_info(si, NULL, maxpages);
> + if (error)
> + goto bad_swap;
> +
> + error = zswap_swapon(si->type, si->max);
> + if (error)
> + goto bad_swap;
> +
> + mutex_lock(&swapon_mutex);
> + si->prio = prio;
> + si->list.prio = -si->prio;
> + si->avail_list.prio = -si->prio;
> + /* si->swap_file stays NULL: this is a file-less device */
> + enable_swap_info(si);
> + pr_info("xswap: adding extendable swap type %d (prio %d, %u pages, max %lu)\n",
> + si->type, prio, si->pages, maxpages);
> + mutex_unlock(&swapon_mutex);
> + atomic_inc(&proc_poll_event);
> + wake_up_interruptible(&proc_poll_wait);
> +
> + return si->type;
> +
> +bad_swap:
> + kfree(si->global_cluster);
> + si->global_cluster = NULL;
> + destroy_swap_extents(si, NULL); /* safe: xswap never sets SWP_ACTIVATED */
> + free_swap_cluster_info(si);
> + si->cluster_info = NULL;
> + spin_lock(&swap_lock);
> + si->flags = 0;
> + spin_unlock(&swap_lock);
> + return error;
> +}
> +#endif /* CONFIG_XSWAP */
> +
> SYSCALL_DEFINE2(swapon, const char __user *, specialfile, int, swap_flags)
> {
> struct swap_info_struct *si;
> @@ -4357,6 +4514,8 @@ static int __init swapfile_init(void)
> swap_migration_ad_supported = true;
> #endif /* CONFIG_MIGRATION */
>
> + xswap_sysfs_init();
> +
> return 0;
> }
> subsys_initcall(swapfile_init);
> --
> 2.54.0
>
The create path and priority handling look consistent with the
existing swap priority rules. In particular, an xswap device has no
`swap_file`, and the changes to `swap_show()` / `swap_start()` /
`swap_next()` account for that when enumerating and displaying swap
devices.
The `VM_SPARSE` mapping avoids allocating the full cluster metadata
backing up front, while allowing the mapped range to grow later.
The overcommit interaction (`si->pages` -> `total_swap_pages` ->
`__vm_enough_memory()`) is an open design item already called out in
the cover letter. I don't see it as a reason to block this phase, but
it should remain explicitly tracked.
Acked-by: Kunwu Chan <chentao@kylinos.cn>
Thanks,
Kunwu
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [PATCH v4 05/14] mm, swap: add xswap grow trigger on cluster allocation
2026-10-03 0:31 ` [PATCH v4 05/14] mm, swap: add xswap grow trigger on cluster allocation Baoquan He
@ 2026-10-07 6:52 ` KunWu Chan
0 siblings, 0 replies; 21+ messages in thread
From: KunWu Chan @ 2026-10-07 6:52 UTC (permalink / raw)
To: Baoquan He
Cc: linux-mm, akpm, chrisl, kasong, hannes, nphamcs, baohua,
youngjun.park, yosry, shikemeng, chengming.zhou, baoquan.he,
david, linux-kernel, klarasmodin
On Sat, Oct 3, 2026 at 8:32 AM Baoquan He <hebaoquan@kylinos.cn> wrote:
>
> Grow the mapped range before it runs out rather than after, from
> cluster_alloc_swap_entry().
>
> Growing maps pages into the VM_SPARSE area and can sleep. Doing it from
> the allocation that would otherwise have failed puts that sleep on the
> swap-out path with nothing left to fall back on. Growing at 85% in use
> instead leaves the allocation that pays for it a reserve to land in, and
> leaves one grow's worth of room in front of the next.
>
> The target is 65%: past the point where the shrink starts looking (50%)
> and short of the point that triggers the next grow (85%), so a grow
> cannot leave the range in a state the shrink would undo.
>
> Since xswap is always SWP_SOLIDSTATE, global_cluster_lock is never
> held on this path.
>
> This makes the xswap cluster space grow transparently as swap usage
> increases, without any userspace intervention.
>
> The caller holds local_lock(&percpu_swap_cluster.lock) across the whole
> slow path, so drop it around xswap_map_clusters() and take it again
> afterwards; it only protects the per-cpu cluster cache, which this path
> does not touch.
>
> Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
> ---
> mm/swapfile.c | 93 +++++++++++++++++++++++++++++++++++++++++++++++++++
> 1 file changed, 93 insertions(+)
>
> diff --git a/mm/swapfile.c b/mm/swapfile.c
> index 9c204e455ef9..45050f8a2ed8 100644
> --- a/mm/swapfile.c
> +++ b/mm/swapfile.c
> @@ -70,6 +70,19 @@
> #define XSWAP_GROW_CLUSTERS \
> max_t(unsigned long, PAGE_SIZE / sizeof(struct swap_cluster_info), 16)
>
> +/*
> + * Grow and shrink thresholds, as a percentage of the mapped range in use.
> + * Each one lands inside (SHRINK_WHEN, GROW_WHEN), so neither can leave the
> + * range in a state that wakes the other up.
> + */
> +#define XSWAP_SHRINK_WHEN 50
> +#define XSWAP_SHRINK_UNTIL 70
> +#define XSWAP_GROW_UNTIL 65
> +#define XSWAP_GROW_WHEN 85
> +#define XSWAP_GROW_CHUNKS 4
> +
> +static bool xswap_should_grow(struct swap_info_struct *si);
> +static unsigned long xswap_grow(struct swap_info_struct *si);
> static int xswap_map_clusters(struct swap_info_struct *si,
> unsigned long start_idx, unsigned long nr);
> static void xswap_unmap_clusters(struct swap_info_struct *si,
> @@ -1204,6 +1217,17 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
> if (order && !(si->flags & SWP_BLKDEV))
> return 0;
>
> +#ifdef CONFIG_XSWAP
> + /*
> + * Top the range up early. Every path below leaves through `done`,
> + * so this has to come first; mapping pages can sleep, and doing it
> + * from the allocation that would otherwise fail leaves nothing to
> + * fall back on.
> + */
> + if ((si->flags & SWP_XSWAP) && xswap_should_grow(si))
> + xswap_grow(si);
> +#endif
> +
> if (!(si->flags & SWP_SOLIDSTATE)) {
> /* Serialize HDD SWAP allocation for each device. */
> spin_lock(&si->global_cluster_lock);
> @@ -1280,6 +1304,12 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
> if (found)
> goto done;
> }
> +
> +#ifdef CONFIG_XSWAP
> + /* A concurrent free or grow may have added clusters; retry once. */
> + if (!found && (si->flags & SWP_XSWAP))
> + found = alloc_swap_scan_list(si, &si->free_clusters, folio, false);
> +#endif
> done:
> if (!(si->flags & SWP_SOLIDSTATE))
> spin_unlock(&si->global_cluster_lock);
> @@ -3909,6 +3939,69 @@ static int xswap_map_clusters(struct swap_info_struct *si,
> return -ENOMEM;
> }
>
> +static bool xswap_should_grow(struct swap_info_struct *si)
> +{
> + unsigned long mapped = READ_ONCE(si->nr_clusters_mapped);
> +
> + if (mapped >= READ_ONCE(si->nr_clusters_max))
> + return false;
> +
> + return swap_usage_in_pages(si) * 100 >=
> + mapped * SWAPFILE_CLUSTER * XSWAP_GROW_WHEN;
> +}
> +
> +/*
> + * Map more of the cluster_info array, up to GROW_UNTIL in use. The caller
> + * holds percpu_swap_cluster.lock, which this drops while it sleeps.
> + */
> +static unsigned long xswap_grow(struct swap_info_struct *si)
> +{
> + unsigned long mapped = READ_ONCE(si->nr_clusters_mapped);
> + unsigned long nr_new, want, i;
> + int ret;
> +
> + /* The unmap below is paired with the caller's lock. */
> + lockdep_assert_held(&percpu_swap_cluster.lock);
> +
> + want = DIV_ROUND_UP(swap_usage_in_pages(si) * 100,
> + XSWAP_GROW_UNTIL * SWAPFILE_CLUSTER);
> + if (want <= mapped)
> + return 0;
> +
> + nr_new = rounddown(want - mapped, XSWAP_GROW_CLUSTERS);
> + if (!nr_new)
> + nr_new = XSWAP_GROW_CLUSTERS;
> + if (nr_new > XSWAP_GROW_CHUNKS * XSWAP_GROW_CLUSTERS)
> + nr_new = XSWAP_GROW_CHUNKS * XSWAP_GROW_CLUSTERS;
> + nr_new = min(nr_new, READ_ONCE(si->nr_clusters_max) - mapped);
> +
> + /*
> + * Mapping pages can sleep. The lock only guards the per-cpu
> + * cluster cache, which this path does not touch.
> + */
> + local_unlock(&percpu_swap_cluster.lock);
> + ret = xswap_map_clusters(si, mapped, nr_new);
> + local_lock(&percpu_swap_cluster.lock);
> + if (ret)
> + return 0;
> +
> + for (i = mapped; i < mapped + nr_new; i++) {
> + struct swap_cluster_info *ci = &si->cluster_info[i];
> +
> + /*
> + * A concurrent grower may have taken these already;
> + * only add the off-list ones.
> + */
> + spin_lock(&ci->lock);
> + if (ci->flags == CLUSTER_FLAG_NONE)
> + move_cluster(si, ci, &si->free_clusters,
> + CLUSTER_FLAG_FREE);
> + spin_unlock(&ci->lock);
> + }
> +
> + return nr_new;
> +}
> +
> static void xswap_unmap_clusters(struct swap_info_struct *si,
> unsigned long start_idx, unsigned long nr)
> {
> --
> 2.54.0
>
Hi Baoquan,
The 85% grow threshold gives the allocation path headroom before
the currently mapped range is exhausted.
I also checked the lock scope in `xswap_grow()`. Dropping the
per-CPU `local_lock` around `xswap_map_clusters()` is appropriate
because the mapping operation may sleep, while this path does not
access the per-CPU cluster cache.
`xswap_should_grow()` uses `READ_ONCE()` for
`nr_clusters_mapped`, so a concurrent grow can make an allocation
miss the grow check for one round. This is benign because a
subsequent allocation will observe the updated mapped range.
Reviewed-by: Kunwu Chan <chentao@kylinos.cn>
Thanks,
Kunwu
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I)
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
` (13 preceding siblings ...)
2026-10-03 0:31 ` [PATCH v4 14/14] mm, swap: let the command line set the xswap device size Baoquan He
@ 2026-10-07 16:28 ` Klara Modin
14 siblings, 0 replies; 21+ messages in thread
From: Klara Modin @ 2026-10-07 16:28 UTC (permalink / raw)
To: Baoquan He
Cc: linux-mm, akpm, chrisl, kasong, hannes, nphamcs, baohua,
youngjun.park, yosry, shikemeng, chengming.zhou, baoquan.he,
david, linux-kernel, kunwu.chan
On 2026-10-03 08:31:22 +0800, Baoquan He wrote:
> A normal swap device has a fixed size, set when it is enabled. This
> does not fit zswap well. With zswap, the swapped data stays in RAM in
> compressed form. So how much swap a machine can really use depends on
> how well the workload compresses. It does not depend on a number we
> pick before we know the workload. If we make the device large enough
> for the worst case, most of it is never used. If we make it just large
> enough for the average case, the machine runs out of swap while zswap
> still has free room.
>
> xswap is a swap device with no backing file. Its swap entries hold pages
> in zswap, and its cluster_info array lives in a VM_SPARSE area that is
> mapped a chunk at a time, so the device covers a large address space
> while costing only the chunks it has actually used. The device grows as
> the workload needs it and gives the tail back when it does not.
>
> This series is the device itself: the sparse cluster array, growth and
> shrink, and the interfaces that create and destroy one. It is phase I
> of three; the physical backend and the swap accounting builds on it.
>
> Design notes
> ------------
>
> The device is only an address space. si->max is set when the device is
> created and does not move for its lifetime; what moves is how much of it
> is mapped. It defaults to twice RAM, and xswap.max= sets another size at
> boot, in bytes or as a percentage of RAM:
>
> xswap.max=8G 8 GiB
> xswap.max=300% three times RAM
>
> That is a boot parameter but not a runtime knob because the address space
> cannot move once the device exists; it is fixed at creation and every
> mapped index stays valid for the device's life. swap_info_struct carries
> nr_clusters_max (the address space) and nr_clusters_mapped (the mapped
> prefix), and the two walkers that touch cluster_info are bounded by the
> latter.
>
> Growth starts when 85% of the mapped range is in use. The call comes from
> cluster_alloc_swap_entry(). Growing maps pages into the VM_SPARSE area.
> Mapping can sleep, so it cannot be done inside the allocation that would
> otherwise fail. If it were, the swap-out path would have nothing to fall
> back on.
>
> Shrink gives the free tail back when the mapped range is at most 50% in
> use. It is triggered when a cluster becomes completely free. Shrinking on
> a smaller drop is not useful. Growth happens on demand, so the same
> clusters would just be mapped again, and each unmap also costs an RCU
> grace period.
>
> A file-less device has no swapon, so /sys/kernel/mm/xswap/create makes
> one and destroy takes its swap type. The teardown is shared with
> sys_swapoff() rather than duplicated.
>
> The only argument create takes is a priority. The size comes from
> xswap.max= at boot, so there is one place to ask for a size and one to
> ask for a priority, and a priority can be given per device while a size
> cannot:
>
> # echo > /sys/kernel/mm/xswap/create default priority
> # echo 100 > /sys/kernel/mm/xswap/create priority 100
>
> The default is the highest priority, SWAP_FLAG_PRIO_MASK. An xswap
> device is then the first one tried, ahead of any ordinary swap device.
> Setting a lower priority is allowed.
Nice, this is now a good default which the user should not have to
change.
>
> The device does not move once it exists, so the address space cannot be
> changed from this interface either. xswap.max= at boot and this together
> are the whole configuration.
>
> The device needs zswap. Without it swapout takes a swap entry and frees
> no memory, so creating one fails.
>
> Testing
> -------
>
> qemu KVM guest, 8G RAM, booted with xswap.max=10G. memhog is a small
> local helper: it faults <total_gb> of anon, fills it with a fixed
> pattern, and holds it.
>
> Set up zswap and create a device:
>
> # echo 1 > /sys/module/zswap/parameters/enabled
> # echo > /sys/kernel/mm/xswap/create
> # awk 'NR == 1 || /xswap/' /proc/swaps
> Filename Type Size Used Priority
> xswap0 xswap 10485756 0 32767
> # cat /sys/kernel/debug/xswap/type0
> clusters_max 5120
> clusters_mapped 73
> usage_pages 0
> tail_free 72
> grows 0
> shrinks 0
>
> Swap 9G of anon through a 2G cgroup:
>
> # mkdir -p /sys/fs/cgroup/xswap_limit
> # echo 2G > /sys/fs/cgroup/xswap_limit/memory.max
> # echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
> # ( echo $BASHPID > /sys/fs/cgroup/xswap_limit/cgroup.procs
> # exec env MEMHOG_FILL=pattern numactl --cpunodebind=0 \
> # ./memhog 9 600 ) &
>
> clusters_mapped grows to cover the pages that go to zswap, while
> SwapTotal stays at 10485756. Throughout, SwapTotal never moves and
> SwapFree stays in [0, SwapTotal]. Pushing past the size leaves the
> cgroup out of room and the OOM killer takes the workload, which is the
> pass signal there.
>
> Growth and shrink, measured with two cgroups so that one keeps its pages
> while the other goes away:
>
> idle mapped=73 occ=0.000 grows=0 shrinks=0
> 2G in a 1G cgroup mapped=803 occ=0.737 grows=8 shrinks=0
> 3G more in a second 1G cgroup mapped=2190 occ=0.777 grows=13 shrinks=0
> kill the second cgroup mapped=876 occ=0.675 grows=13 shrinks=1
>
> The second cgroup is started later, so its clusters sit at the tail, and
> killing it leaves a free tail to take back. The first stays alive, so
> usage falls but not to zero: that is the state the shrink has to land in.
> occ is usage over the mapped range, which is what the grow and shrink
> thresholds are expressed in.
>
> # echo 0 > /sys/kernel/mm/xswap/destroy
>
> Also tested things as below, and nothing crashed or warned:
>
> - a shrink racing a swapoff of the same device, 20 rounds
> - a cgroup that cannot zswap: it must get no xswap slot, and the
> allocator walk must not spin
> - a create/destroy cycle: it must not leak per-cluster state
>
> Open items
> ----------
>
> - xswap_create() sets si->pages to the whole address space (twice RAM by
> default) and _enable_swap_info() adds it to nr_swap_pages and
> total_swap_pages, which is what __vm_enough_memory() uses for
> overcommit. That raises the commit limit by 2xRAM for a device whose
> real capacity is bounded by the zswap pool and by how well the workload
> compresses. The question is which semantics overcommit should follow:
> worst-case backing, or the address space.
>
> - destroy takes a bare swap type. The type is an internal index that
> alloc_swap_info() recycles as soon as SWP_USED clears, so a value read
> earlier can name a different device by the time it is written. The
> swapoff ABI uses a path for this reason. It wants a generation or
> another stable handle.
>
> - The new interfaces (/sys/kernel/mm/xswap/create and destroy, xswap.max=,
> the new type string in /proc/swaps) are not documented yet.
To me it feels a bit weird to have the interface split between a kernel
parameter and sysfs, meaning I have to configure this in two places now.
As far as I gather, there's also nothing that needs the limit to be
fixed at boot since it is applied at creation. If we want to expose this
limit to users (as xswap does) it would be nice to have a 'max'
shorthand for tha largest possible value, i.e. SWAPFILE_CLUSTER *
UNIT_MAX (or whatever it might be in the future). That would I guess in
practice give the same behaviour as vm.overcommit_memory=1 which might
or might not be desirable.
>
> Changelog
> =========
> v3 -> v4:
> - Rebased onto the latest mm-new.
>
> - Add boot parameter xswap.max= .
>
> - The device has one size now: set at creation, fixed, with xswap.max= to
> ask for another one at boot. v3's runtime ceiling, the per-device limit,
> the clamp it wrote and the shrink to that ceiling are all gone as
> Johannes suggested (v3's patches 12, 13 and 14).
>
> - v3's patches 2 and 4 are the new patch 3: neither builds a device
> alone. It also bounds find_next_to_unuse() and wait_for_allocation(),
> which walked up to si->max past the mapped range. Chris suggested
> this.
>
> - New patches 11 to 13: a cgroup that cannot zswap is not given an xswap
> slot, swap_info_struct max and pages are unsigned long, and debugfs
> counters for the mapped range and the grow and shrink counts.
>
> - v3's patches 1 and 3, and 5 to 11, are unchanged; the merge shifts them
> to 1 and 2, and 4 to 10.
>
> - Minor comment and cleanup changes.
>
> v2 -> v3:
> - Rebased onto the latest mm-new.
>
> - The grow path now honors the user-set ceiling (si->nr_clusters) instead
> of growing up to nr_clusters_max, and a ceiling below the mapped range
> is unmapped exactly instead of rounded to a chunk (patches 12 and 14).
>
> - The limit write clamps the ceiling up to the clusters covering the pages
> in use, replacing the earlier WARN_ONCE; si->pages becomes mutable at
> runtime (patch 13).
>
> - Minor comment and cleanup changes.
>
> v1->v2:
> - Patch 1 (mm: zswap: return -ENOENT when the swap device is gone) is not
> part of this series; it was posted separately.
>
> - There is only one size knob now. The runtime ceiling and the debugfs
> per-device limit are gone. All that is left is the optional per-device
> cap, /sys/kernel/mm/xswap/type<N>/limit. Grow and shrink work without
> it.
>
> - The shrink no longer keeps its own count of the free tail. It scans the
> tail instead, and dropping the counter also removes a call from the
> cluster allocation path.
>
> - The priority is no longer a patch of its own. The create attribute
> takes it:
> echo 100 > /sys/kernel/mm/xswap/create
>
> RFC v3 -> RFC v2
> - Add patch 16 to support setting xswap device priority at creation.
> The create sysfs interface (/sys/kernel/mm/xswap/create) previously
> hardcoded every new device's priority to DEF_SWAP_PRIO, it now
> accepts an optional priority:
>
> echo "<percent> [<prio>]" > /sys/kernel/mm/xswap/create
>
> - Bug fix: xswap_lock init ordering. mutex_init(&si->xswap_lock) was called
> after xswap_map_clusters() (which locks it), i.e. locking an uninitialized
> mutex. Init now before the first xswap_map_clusters() call. Thanks to Klara.
>
> - Bug fix: Fixes a compile error in !CONFIG_XSWAP builds. xswap_debugfs_root
> is declared inside CONFIG_XSWAP ifdeffery scope, so the ungarded use
> caused error when CONFIG_XSWAP is off.
>
> RFC v2-> RFC v3:
> - Replace the "header-only swap file + swapon" creation hack with a
> proper file-less device created and destroyed via sysfs
> (/sys/kernel/mm/xswap/{create,destroy}). This required the
> __swapoff() refactor and the free_swap_cluster_info() signature
> change (patches 4, 6, 14).
>
> - Require zswap: refuse to create an xswap device when zswap is
> unavailable (patch 15).
>
> - Split the unrelated zswap -ENOENT fix out of the series into a
> standalone patch (patch 1).
>
> - Fix nr_free_tail over-counting on concurrent grow, shrink leaking
> detached clusters on early bail-out, a re-init race on cluster
> spinlocks in xswap_map_clusters(), the nr_clusters_mapped update
> ordering, and swapoff accessing the shrinker-unmapped cluster tail.
>
> - Minor cleanups (checkpatch, /proc/swaps alignment, commit messages).
>
> RFC v1-> RFC v2:
> - Added __GFP_HIGH | __GFP_NOMEMALLOC to alloc_page() and kmalloc_array()
> in the grow path, plus memalloc_noreclaim_save/restore() wrapping,
> to prevent the grow path from consuming emergency memory reserves
> or recursing into swap under PF_MEMALLOC. This is folded into patch 3.
> This was pointed out by Nhat.
>
> - Folded the mutex serialization fix into the cluster grow patch (patch
> 3). This is suggested by Nhat.
>
> - Fixed coding style issues: corrected indentation of declarations in
> xswap_unmap_clusters(), removed unnecessary block scope around the
> err variable in xswap_map_clusters().
>
> - Rebased onto mm-unstable
>
> Baoquan He (13):
> mm, swap: refactor free_swap_cluster_info to take swap_info_struct
> mm, swap: back the cluster_info array with a sparse VM_SPARSE area
> mm, swap: add sysfs create interface for xswap
> mm, swap: add xswap grow trigger on cluster allocation
> mm, swap: add xswap_try_shrink and shrink trigger on cluster free
> mm, swap: free backing pages in xswap_unmap_clusters
> mm, swap: defer xswap shrink to workqueue to avoid lock recursion
> mm, swap: refactor swapoff and add xswap_destroy
> mm, swap: require zswap for xswap devices
> mm, swap: do not give an xswap slot to a cgroup that cannot zswap
> mm, swap: widen swap_info_struct max/pages to unsigned long
> mm, swap: add debugfs counters for xswap
> mm, swap: let the command line set the xswap device size
>
> Chris Li (1):
> mm: xswap support for zswap
>
> include/linux/swap.h | 16 +-
> mm/Kconfig | 9 +
> mm/page_io.c | 19 +
> mm/swap_state.c | 9 +
> mm/swapfile.c | 1221 +++++++++++++++++++++++++++++++++++++++---
> mm/zswap.c | 11 +-
> 6 files changed, 1205 insertions(+), 80 deletions(-)
>
>
> base-commit: 9a0542b19a541eda3f82934ee747c61ec3f18847
> --
> 2.54.0
>
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [PATCH v4 08/14] mm, swap: defer xswap shrink to workqueue to avoid lock recursion
2026-10-03 0:31 ` [PATCH v4 08/14] mm, swap: defer xswap shrink to workqueue to avoid lock recursion Baoquan He
@ 2026-10-08 9:10 ` KunWu Chan
0 siblings, 0 replies; 21+ messages in thread
From: KunWu Chan @ 2026-10-08 9:10 UTC (permalink / raw)
To: Baoquan He
Cc: linux-mm, akpm, chrisl, kasong, hannes, nphamcs, baohua,
youngjun.park, yosry, shikemeng, chengming.zhou, baoquan.he,
david, linux-kernel, klarasmodin
On Sat, Oct 3, 2026 at 8:32 AM Baoquan He <hebaoquan@kylinos.cn> wrote:
>
> __free_cluster() called xswap_try_shrink() while holding ci->lock, but
> shrinking unmaps the backing pages and the subsequent unlock faults on
> the unmapped address. Run the shrink via schedule_work() instead, so no
> cluster lock is held. The work is only scheduled for xswap devices and
> is cancelled on swapoff.
>
> Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
> ---
> include/linux/swap.h | 1 +
> mm/swapfile.c | 129 ++++++++++++++++++++++++++++++++-----------
> 2 files changed, 97 insertions(+), 33 deletions(-)
>
> diff --git a/include/linux/swap.h b/include/linux/swap.h
> index 382a578140b5..12cdc00f78f9 100644
> --- a/include/linux/swap.h
> +++ b/include/linux/swap.h
> @@ -246,6 +246,7 @@ struct swap_info_struct {
> struct vm_struct *cluster_vm; /* VM_SPARSE area for cluster_info */
> unsigned long nr_clusters_max;/* total clusters in the xswap address space */
> unsigned long nr_clusters_mapped; /* currently mapped cluster count */
> + struct work_struct xswap_shrink_work; /* deferred shrink trigger */
> struct mutex xswap_lock; /* serialize map/unmap operations */
> #endif
> struct list_head free_clusters; /* free clusters list */
> diff --git a/mm/swapfile.c b/mm/swapfile.c
> index 7c062e772f5b..49703731ffd5 100644
> --- a/mm/swapfile.c
> +++ b/mm/swapfile.c
> @@ -727,7 +727,9 @@ static void __free_cluster(struct swap_info_struct *si, struct swap_cluster_info
> move_cluster(si, ci, &si->free_clusters, CLUSTER_FLAG_FREE);
> ci->order = 0;
> #ifdef CONFIG_XSWAP
> - xswap_try_shrink(si);
> + /* Only xswap devices, and not while the device is being torn down. */
> + if ((si->flags & SWP_XSWAP) && (si->flags & SWP_WRITEOK))
> + schedule_work(&si->xswap_shrink_work);
> #endif
> }
>
> @@ -3334,6 +3336,7 @@ static void free_swap_cluster_info(struct swap_info_struct *si)
> if (si->flags & SWP_XSWAP) {
> unsigned long nr_mapped;
>
> + cancel_work_sync(&si->xswap_shrink_work);
> /*
> * Cluster 0 keeps the bad header slot, so it never empties
> * and __free_cluster() never frees its table.
> @@ -3452,6 +3455,11 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specialfile)
> spin_unlock(&p->lock);
> spin_unlock(&swap_lock);
>
> +#ifdef CONFIG_XSWAP
> + if (p->flags & SWP_XSWAP)
> + cancel_work_sync(&p->xswap_shrink_work);
> +#endif
> +
> wait_for_allocation(p);
>
> set_current_oom_origin();
> @@ -4074,8 +4082,9 @@ static void xswap_unmap_range(struct swap_info_struct *si,
> __free_page(pages[i]);
> }
>
> -static void xswap_unmap_clusters(struct swap_info_struct *si,
> - unsigned long start_idx, unsigned long nr)
> +/* Caller must hold si->xswap_lock. Cannot fail. */
> +static void xswap_unmap_clusters_locked(struct swap_info_struct *si,
> + unsigned long start_idx, unsigned long nr)
> {
> unsigned long start_addr = (unsigned long)si->cluster_info +
> (size_t)start_idx * sizeof(struct swap_cluster_info);
> @@ -4087,11 +4096,9 @@ static void xswap_unmap_clusters(struct swap_info_struct *si,
> unsigned int noreclaim_flags;
> unsigned long npages, idx;
>
> - mutex_lock(&si->xswap_lock);
> -
> if (vm_start >= vm_end) {
> WRITE_ONCE(si->nr_clusters_mapped, start_idx);
> - goto out_unlock;
> + return;
> }
>
> /*
> @@ -4113,7 +4120,7 @@ static void xswap_unmap_clusters(struct swap_info_struct *si,
> xswap_unmap_range(si, vm_start, vm_end, pages, npages);
> kvfree(pages);
> WRITE_ONCE(si->nr_clusters_mapped, start_idx);
> - goto out_unlock;
> + return;
> }
>
> /* No memory for the array: unmap in bounded batches instead. */
> @@ -4133,7 +4140,13 @@ static void xswap_unmap_clusters(struct swap_info_struct *si,
> }
>
> WRITE_ONCE(si->nr_clusters_mapped, start_idx);
> -out_unlock:
> +}
> +
> +static void xswap_unmap_clusters(struct swap_info_struct *si,
> + unsigned long start_idx, unsigned long nr)
> +{
> + mutex_lock(&si->xswap_lock);
> + xswap_unmap_clusters_locked(si, start_idx, nr);
> mutex_unlock(&si->xswap_lock);
> }
>
> @@ -4157,6 +4170,16 @@ static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data)
> #define XSWAP_SHRINK_SLACK XSWAP_GROW_CLUSTERS
> #define XSWAP_SHRINK_MIN (XSWAP_GROW_CLUSTERS * 8)
>
> +static void xswap_shrink_work_fn(struct work_struct *work)
> +{
> + struct swap_info_struct *si = container_of(work,
> + struct swap_info_struct, xswap_shrink_work);
> +
> + if (!(READ_ONCE(si->flags) & SWP_WRITEOK))
> + return;
> + xswap_try_shrink(si);
> +}
> +
> /*
> * Try to shrink the cluster_info tail: unmap contiguous free clusters
> * at the end of the mapped range.
> @@ -4164,14 +4187,20 @@ static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data)
> static void xswap_try_shrink(struct swap_info_struct *si)
> {
> struct swap_cluster_info *ci;
> - unsigned long nr_mapped, last, keep, idx;
> + unsigned long nr_mapped, nr_tail, keep, nr_unmap, start_idx, i;
>
> if (!(si->flags & SWP_XSWAP))
> return;
>
> + mutex_lock(&si->xswap_lock);
> +
> + /* A swapoff raced us and is about to walk this mapping. */
> + if (!(READ_ONCE(si->flags) & SWP_WRITEOK))
> + goto out_unlock;
> +
> nr_mapped = READ_ONCE(si->nr_clusters_mapped);
> - if (nr_mapped <= 1) /* keep cluster 0 */
> - return;
> + if (nr_mapped <= 1) /* keep cluster 0 */
> + goto out_unlock;
>
> /*
> * Reclaim on our own, but only once the mapped range is at most
> @@ -4181,40 +4210,73 @@ static void xswap_try_shrink(struct swap_info_struct *si)
> */
> if (swap_usage_in_pages(si) * 100 >
> nr_mapped * SWAPFILE_CLUSTER * XSWAP_SHRINK_WHEN)
> - return;
> + goto out_unlock;
>
> - /* Find the last non-free cluster from the tail */
> - last = nr_mapped;
> - while (last > 1) {
> - idx = last - 1;
> - ci = &si->cluster_info[idx];
> - if (ci->count || ci->flags != CLUSTER_FLAG_FREE)
> + /*
> + * Count the free clusters at the tail of the mapped range. Scanned,
> + * not tracked: the count must be exact to size the unmap, and an
> + * incremental count falls behind on out-of-order frees.
> + */
> + nr_tail = 0;
> + while (nr_mapped - nr_tail > 1) {
> + ci = &si->cluster_info[nr_mapped - nr_tail - 1];
> + if (READ_ONCE(ci->count) ||
> + READ_ONCE(ci->flags) != CLUSTER_FLAG_FREE)
> break;
> - last = idx;
> + nr_tail++;
> }
> -
> - if (last == nr_mapped)
> - return; /* nothing to shrink */
> -
> - /* Below `last` has to stay mapped: the free ones in between are
> - * not part of the tail, and unmapping them orphans what is above.
> - */
> - if (nr_mapped - last < XSWAP_SHRINK_SLACK + XSWAP_SHRINK_MIN)
> - return;
> + if (nr_tail < XSWAP_SHRINK_SLACK + XSWAP_SHRINK_MIN)
> + goto out_unlock;
>
> /*
> * Stop at SHRINK_UNTIL rather than at the end of the tail, or the
> - * range comes out full enough for the grow to be woken.
> + * range comes out full enough for the grow to be woken. Below the
> + * tail everything stays mapped, and one chunk of tail with it.
> */
> keep = DIV_ROUND_UP(swap_usage_in_pages(si) * 100,
> XSWAP_SHRINK_UNTIL * SWAPFILE_CLUSTER);
> - if (keep < last + XSWAP_SHRINK_SLACK)
> - keep = last + XSWAP_SHRINK_SLACK;
> + if (keep < nr_mapped - nr_tail + XSWAP_SHRINK_SLACK)
> + keep = nr_mapped - nr_tail + XSWAP_SHRINK_SLACK;
>
> if (nr_mapped < keep + XSWAP_SHRINK_MIN)
> - return;
> + goto out_unlock;
> +
> + nr_unmap = rounddown(nr_mapped - keep, XSWAP_GROW_CLUSTERS);
> + if (!nr_unmap)
> + goto out_unlock;
> + start_idx = nr_mapped - nr_unmap;
> +
> + /*
> + * Only shrink a run that reaches the mapped end; otherwise
> + * truncating nr_clusters_mapped would orphan the active tail.
> + */
> + spin_lock(&si->lock);
> + for (i = start_idx; i < nr_mapped; i++) {
> + ci = &si->cluster_info[i];
> + if (READ_ONCE(ci->flags) != CLUSTER_FLAG_FREE)
> + break;
> + if (!spin_trylock(&ci->lock)) {
> + spin_unlock(&si->lock);
> + goto out_unlock;
> + }
> + spin_unlock(&ci->lock);
> + }
> + if (i != nr_mapped) {
> + spin_unlock(&si->lock);
> + goto out_unlock;
> + }
> +
> + for (i = start_idx; i < nr_mapped; i++) {
> + ci = &si->cluster_info[i];
> + list_del_init(&ci->list);
> + WRITE_ONCE(ci->flags, CLUSTER_FLAG_NONE);
> + }
> + spin_unlock(&si->lock);
>
> - xswap_unmap_clusters(si, keep, nr_mapped - keep);
> + xswap_unmap_clusters_locked(si, start_idx, nr_unmap);
> +
> +out_unlock:
> + mutex_unlock(&si->xswap_lock);
> }
> #endif /* CONFIG_XSWAP */
>
> @@ -4278,6 +4340,7 @@ static int setup_swap_clusters_info(struct swap_info_struct *si,
> }
> }
>
> + INIT_WORK(&si->xswap_shrink_work, xswap_shrink_work_fn);
> return 0;
>
> err_unmap:
> --
> 2.54.0
>
Hi Baoquan,
While following the shrink path across these patches, I wanted to
clarify the lifetime guarantee for the `cluster_info` backing pages.
P07 flushes the per-CPU swap cluster cache and calls
`synchronize_rcu()` before unmapping the tail. P09 clears
`SWP_WRITEOK` and cancels the shrink work before proceeding with
`wait_for_allocation()` and `try_to_unuse()`.
Could you clarify how the shrink path ensures that, before the tail
is unmapped, no allocation-side reader can still access the tail or
republish a stale per-CPU cache entry that points into it, and no
swapoff-side walker can still access the unmapped range?
Thanks,
Kunwu
^ permalink raw reply [flat|nested] 21+ messages in thread
end of thread, other threads:[~2026-10-08 9:10 UTC | newest]
Thread overview: 21+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-03 0:31 [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Baoquan He
2026-10-03 0:31 ` [PATCH v4 01/14] mm: xswap support for zswap Baoquan He
2026-10-07 4:22 ` KunWu Chan
2026-10-03 0:31 ` [PATCH v4 02/14] mm, swap: refactor free_swap_cluster_info to take swap_info_struct Baoquan He
2026-10-07 4:32 ` KunWu Chan
2026-10-03 0:31 ` [PATCH v4 03/14] mm, swap: back the cluster_info array with a sparse VM_SPARSE area Baoquan He
2026-10-03 0:31 ` [PATCH v4 04/14] mm, swap: add sysfs create interface for xswap Baoquan He
2026-10-07 6:32 ` KunWu Chan
2026-10-03 0:31 ` [PATCH v4 05/14] mm, swap: add xswap grow trigger on cluster allocation Baoquan He
2026-10-07 6:52 ` KunWu Chan
2026-10-03 0:31 ` [PATCH v4 06/14] mm, swap: add xswap_try_shrink and shrink trigger on cluster free Baoquan He
2026-10-03 0:31 ` [PATCH v4 07/14] mm, swap: free backing pages in xswap_unmap_clusters Baoquan He
2026-10-03 0:31 ` [PATCH v4 08/14] mm, swap: defer xswap shrink to workqueue to avoid lock recursion Baoquan He
2026-10-08 9:10 ` KunWu Chan
2026-10-03 0:31 ` [PATCH v4 09/14] mm, swap: refactor swapoff and add xswap_destroy Baoquan He
2026-10-03 0:31 ` [PATCH v4 10/14] mm, swap: require zswap for xswap devices Baoquan He
2026-10-03 0:31 ` [PATCH v4 11/14] mm, swap: do not give an xswap slot to a cgroup that cannot zswap Baoquan He
2026-10-03 0:31 ` [PATCH v4 12/14] mm, swap: widen swap_info_struct max/pages to unsigned long Baoquan He
2026-10-03 0:31 ` [PATCH v4 13/14] mm, swap: add debugfs counters for xswap Baoquan He
2026-10-03 0:31 ` [PATCH v4 14/14] mm, swap: let the command line set the xswap device size Baoquan He
2026-10-07 16:28 ` [PATCH v4 00/14] mm, swap: extendable swap devices (xswap phase I) Klara Modin
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox