* [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend
@ 2026-09-20 7:20 Baoquan He
2026-09-20 7:20 ` [RFC PATCH 01/17] mm, swap: prepare the swap IO path for xswap backends Baoquan He
` (18 more replies)
0 siblings, 19 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
This is the writeback layer for xswap, on top of the base series. Post it
as RFC for discussion.
The base series keeps every swapped-out page in zswap. So the pool must
refuse new pages once it's full. This series gives an xswap slot a
backend to go, and charge it only when it really goes to real disk
space.
Design
------
- Backend: when zswap refuses a page, take a slot on the real swap
device with the highest priority and write the page there. One IO per
slot.
- Ownership: that slot has no swap cache folio, so record its owner in
the swap table entry. 0b100 in the low bits marks a backend slot, and
the xswap type/offset go above. The physical side then does not treat
it as free, and readahead skips it.
- Per-slot record: ci->xs_table[], one unsigned long per slot, allocated
the first time a cluster takes a backend. Zero means no backend.
- Read: a written-out slot has no zswap copy, so swap_read_folio()
follows the pointer and reads from the backend.
- Release: verify the slot still points back to the xswap entry first;
the record can go stale.
- swapoff: try_to_unuse() used to skip these slots. It now reads them
back, marks the folio dirty so reclaim re-stores it in zswap, then
drops the backend.
- Reclaim: an xswap entry's count can reach 0 while its folio is still
in the swap cache; the physical scanner reclaims the redundant slot
through __try_to_reclaim_swap().
- Large folios: a read starts at the first physical slot, so the run is
reserved whole and must be backed contiguously; otherwise the fault
retries at a smaller order. xswap is SWP_SYNCHRONOUS_IO, which also
skips readahead.
- Charging: Record the owner at allocation but do not charge. Take the
charge when the entry gets a physical backend slot, and give it back
with that slot.memory.swap.current then counts only real on-disk swap
usage, and a cgroup with memory.swap.max at 0 can still swap out through
zswap.
Note
----
The base series is sereis of xswap foundation. This one depends on it.
So if anyone wants to apply this patchset, the order is base-commit
as below, then xswap foundation, finally this patchset.
[PATCH v3 00/14] mm, swap: extendable swap devices (xswap)
https://lore.kernel.org/all/20260916101929.149106-1-hebaoquan@kylinos.cn/T/#u
base-commit: baa8de2f3448d1466a888a805c18d01c998fe052
This partially refers to Nhat's vswap-v4 series. E.g patch 1 is
consistent with his patch 3, patch 12/13/14/15 are from his patch 8.
And the last 2 patches are fixing bugs when charging patches are added.
Since the charging of xswap is still under discussion, I dind't merge
them into commits. Will squash them or take them off once decision is
made.
Testing
-------
qemu KVM guest, 8G RAM.
A 4G swap disk /dev/vdb is added as the physical backend, and zswap is
turned off so that every page has to reach a backend slot. Note that it
need create xswap device firstly then disable zswap, so every page has
to reach backend slot.
The workload is memhog, as in the base series. Every page is filled with
a fixed pattern, so the data can be checked after a readback.
# echo 1 > /sys/module/zswap/parameters/enabled
# echo 100 > /sys/kernel/mm/xswap/create
# mkswap /dev/vdb && swapon -p 0 /dev/vdb
# echo 0 > /sys/module/zswap/parameters/enabled
# mkdir -p /sys/fs/cgroup/xswap_limit
# echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
# MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 ./memhog
1. Swapout to the backend
Run the workload with 5G of memory in a cgroup capped at 4G:
# awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
Filename Type Size Used Priority
xswap0 xswap 8155132 2392176 100
/dev/vdb partition 4194300 2392176 0
The two Used values are equal, so the pages reached the backend and
the xswap side accounted for them.
2. Pages are read back from the device
Lifting memory.max and touching the pages again reads them back. The
pages-swapped-in counter went up by 265642, so the read did reach the
device. 1010 of the folios read back were 1M, each read with one IO:
# echo max > /sys/fs/cgroup/xswap_limit/memory.max
# cat /sys/kernel/mm/transparent_hugepage/hugepages-1024kB/stats/swpin
1010
hugepages-2048kB/stats/swpin stays at 0, because the swapin order is
capped below the PMD order.
3. verifies data correctness
After the readback the data is compared byte for byte against the
pattern, one time with zswap on and one time with zswap off, so the
pages come from zswap in one time and from the backend in the other
time. All bytes matched.
4. Destroying a device returns its backend slots
The device was filled, then destroyed with no readback first:
# awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
3157220
# echo 0 > /sys/kernel/mm/xswap/destroy
# awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
0
5. verifies the charge is correct
An xswap entry is charged only when it holds a backend slot, so the
two counters have to show the same number of pages:
# echo $(( $(cat /sys/fs/cgroup/xswap_limit/memory.swap.current) / 4096 ))
526852
# awk '$1 == "/dev/vdb" { print $4 / 4 }' /proc/swaps
526852
Nothing is charged while the pages stay in zswap: with the pool holding
them and the backend untouched, memory.swap.current stays 0.
6. A cgroup with no swap room can still reclaim its anon memory
With memory.swap.max at 0 and no xswap device, the cgroup ran out of
room and the workload was killed. With an xswap device and zswap on,
it swapped out through zswap, was charged nothing, and was not killed.
E.g if memory.swap.max is set at 100M, the cgroup will be killed when
the cap was reached.
Baoquan He (12):
mm, swap: tag a swap table entry with its owning xswap entry
mm, swap: prepare the folio-less allocation path for xswap
mm, swap: add a physical backend for xswap slots
mm, swap: use the xswap physical backend
mm, swap: fall back to disk when zswap refuses an xswap page
mm, swap: support swapoff of an xswap physical backend
mm, swap: reclaim physical slots backing cache-only xswap entries
mm, swap: back a large xswap folio with a contiguous physical run
mm, swap: enable THP swapin for xswap entries
mm, swap: drop swap_folio_sector()
mm, swap: do not retake the cluster lock when uncharging an xswap slot
mm, swap: drop a refused xswap backend run directly
Nhat Pham (5):
mm, swap: prepare the swap IO path for xswap backends
mm, swap: split the swap memcg charge helpers
mm, swap: do not charge zswap-backed xswap entries
mm, swap: charge an xswap entry when it gets physical backing
mm, swap: don't gate xswap on the physical swap free count
.../admin-guide/cgroup-v1/memcg_test.rst | 2 +-
include/linux/memcontrol.h | 6 +
include/linux/swap.h | 69 +-
include/linux/swap_ops.h | 9 +-
mm/memcontrol-v1.c | 10 +-
mm/memcontrol.c | 147 ++--
mm/memory.c | 7 +-
mm/page_io.c | 95 ++-
mm/swap.h | 35 +-
mm/swap_state.c | 12 +-
mm/swap_table.h | 53 ++
mm/swapfile.c | 742 ++++++++++++++++--
mm/zswap.c | 20 +-
13 files changed, 1026 insertions(+), 181 deletions(-)
--
2.54.0
^ permalink raw reply [flat|nested] 25+ messages in thread
* [RFC PATCH 01/17] mm, swap: prepare the swap IO path for xswap backends
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 02/17] mm, swap: tag a swap table entry with its owning xswap entry Baoquan He
` (17 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
From: Nhat Pham <nphamcs@gmail.com>
An xswap folio can be written to a physical device. The folio stays in
the swap cache of its own xswap entry. The IO goes to a slot on that
physical device. So folio->swap no longer describes where the IO goes.
The batched swap IO path takes the bio sector from folio->swap. For
such a folio it resolves the xswap device, not the physical one. An
xswap device has no bdev and owns no extents, so both the device and
the sector come out wrong.
Carry the entry explicitly instead. __swap_writepage(), swap_add_folio()
and the swap_ops->can_merge() all take it as parameter. swap_iocb records
the first entry of each batch. The merge check, the bio sector and the
completion path all address the IO from it.
Signed-off-by: Nhat Pham <nphamcs@gmail.com>
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
include/linux/swap.h | 1 +
include/linux/swap_ops.h | 9 ++++---
mm/page_io.c | 55 ++++++++++++++++++++--------------------
mm/swap.h | 3 ++-
mm/swapfile.c | 13 ++++++++++
mm/zswap.c | 2 +-
6 files changed, 49 insertions(+), 34 deletions(-)
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 804189b4b4eb..b9a9a02007cf 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -419,6 +419,7 @@ extern bool swap_entry_swapped(struct swap_info_struct *si, swp_entry_t entry);
extern int swp_swapcount(swp_entry_t entry);
extern struct swap_info_struct *get_swap_device(swp_entry_t entry);
sector_t swap_folio_sector(struct folio *folio);
+sector_t swap_entry_sector(swp_entry_t entry);
/*
* If there is an existing swap slot reference (swap entry) and the caller
diff --git a/include/linux/swap_ops.h b/include/linux/swap_ops.h
index 57ac6c703f68..198c5738c3c7 100644
--- a/include/linux/swap_ops.h
+++ b/include/linux/swap_ops.h
@@ -10,6 +10,7 @@ struct swap_iocb {
struct bio bio;
};
struct bio_vec bvecs[SWAP_CLUSTER_MAX];
+ swp_entry_t entry; /* First entry of the batch */
int nr_bvecs;
int len;
};
@@ -30,15 +31,15 @@ struct swap_io_ctx {
struct swap_ops {
unsigned int flags;
- bool (*can_merge)(struct folio *folio, struct folio *prev_folio,
- size_t prev_folio_size, int rw);
+ bool (*can_merge)(struct folio *folio, swp_entry_t entry,
+ struct swap_iocb *sio, int rw);
void (*submit_write)(struct swap_io_ctx *ctx);
void (*submit_read)(struct swap_io_ctx *ctx);
};
void swap_fs_prepare_rw(struct swap_io_ctx *ctx, int rw, struct iov_iter *iter);
-bool swap_fs_can_merge(struct folio *folio, struct folio *prev_folio,
- size_t prev_folio_size, int rw);
+bool swap_fs_can_merge(struct folio *folio, swp_entry_t entry,
+ struct swap_iocb *sio, int rw);
int swap_fs_activate(struct swap_info_struct *sis, const struct swap_ops *ops);
#endif /* _MM_SWAP_OPS_H */
diff --git a/mm/page_io.c b/mm/page_io.c
index d685c2e2429a..2b1387c87e51 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -257,7 +257,7 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
return AOP_WRITEPAGE_ACTIVATE;
}
- __swap_writeout(ctx, folio);
+ __swap_writeout(ctx, folio, folio->swap);
return 0;
out_unlock:
folio_unlock(folio);
@@ -328,24 +328,22 @@ int sio_pool_init(void)
}
static bool swap_can_merge(struct swap_io_ctx *ctx, struct folio *folio,
- int rw)
+ swp_entry_t entry, int rw)
{
- struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
- struct bio_vec *last_bv = &ctx->sio->bvecs[ctx->sio->nr_bvecs - 1];
- struct folio *prev_folio = bvec_folio(last_bv);
- size_t prev_folio_size = folio_size(prev_folio);
+ struct swap_info_struct *sis = __swap_entry_to_info(entry);
if (ctx->sis != sis)
return false;
- return sis->ops->can_merge(folio, prev_folio, prev_folio_size, rw);
+ return sis->ops->can_merge(folio, entry, ctx->sio, rw);
}
-static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, int rw)
+static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio,
+ swp_entry_t entry, int rw)
{
- struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
+ struct swap_info_struct *sis = __swap_entry_to_info(entry);
struct swap_iocb *sio = ctx->sio;
- if (sio && !swap_can_merge(ctx, folio, rw)) {
+ if (sio && !swap_can_merge(ctx, folio, entry, rw)) {
if (rw == WRITE)
swap_write_submit(ctx);
else
@@ -358,6 +356,7 @@ static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, int rw)
ctx->sio = sio = mempool_alloc(sio_pool, GFP_NOIO);
sio->nr_bvecs = 0;
sio->len = 0;
+ sio->entry = entry;
}
bvec_set_folio(&sio->bvecs[sio->nr_bvecs], folio, folio_size(folio), 0);
sio->len += folio_size(folio);
@@ -378,7 +377,8 @@ static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, int rw)
}
}
-void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
+void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio,
+ swp_entry_t entry)
{
VM_BUG_ON_FOLIO(!folio_test_swapcache(folio), folio);
@@ -394,7 +394,7 @@ void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
folio_start_writeback(folio);
folio_unlock(folio);
- swap_add_folio(ctx, folio, WRITE);
+ swap_add_folio(ctx, folio, entry, WRITE);
}
/*
@@ -503,7 +503,7 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct folio *folio)
/* We have to read from slower devices. Increase zswap protection. */
zswap_folio_swapin(folio);
- swap_add_folio(ctx, folio, READ);
+ swap_add_folio(ctx, folio, folio->swap, READ);
finish:
if (workingset) {
@@ -535,8 +535,6 @@ static void swap_fs_write_complete(struct kiocb *iocb, long ret)
bool failed = ret != sio->len;
if (failed) {
- struct folio *folio = bvec_folio(&sio->bvecs[0]);
-
/*
* In the case of swap-over-nfs, this can be a temporary failure
* if the system has limited memory for allocating transmit
@@ -544,7 +542,7 @@ static void swap_fs_write_complete(struct kiocb *iocb, long ret)
* folio_rotate_reclaimable but rate-limit the messages.
*/
pr_err_ratelimited("Write error %ld on dio swapfile (%llu)\n",
- ret, swap_dev_pos(folio->swap));
+ ret, swap_dev_pos(sio->entry));
}
swap_write_end(sio, failed);
@@ -616,7 +614,7 @@ static void swap_bdev_submit_write(struct swap_io_ctx *ctx)
bio_init(bio, ctx->sis->bdev, sio->bvecs, ARRAY_SIZE(sio->bvecs),
REQ_OP_WRITE | REQ_SWAP);
bio->bi_iter.bi_size = sio->len;
- bio->bi_iter.bi_sector = swap_folio_sector(bio_first_folio_all(bio));
+ bio->bi_iter.bi_sector = swap_entry_sector(sio->entry);
bio_associate_blkg_from_folio(bio, bio_first_folio_all(bio));
if (ctx->sis->flags & SWP_SYNCHRONOUS_IO) {
@@ -636,7 +634,7 @@ static void swap_bdev_submit_read(struct swap_io_ctx *ctx)
bio_init(bio, ctx->sis->bdev, sio->bvecs, ARRAY_SIZE(sio->bvecs),
REQ_OP_READ);
bio->bi_iter.bi_size = sio->len;
- bio->bi_iter.bi_sector = swap_folio_sector(bio_first_folio_all(bio));
+ bio->bi_iter.bi_sector = swap_entry_sector(sio->entry);
if (ctx->sis->flags & SWP_SYNCHRONOUS_IO) {
/*
@@ -654,13 +652,15 @@ static void swap_bdev_submit_read(struct swap_io_ctx *ctx)
}
}
-static bool swap_bdev_can_merge(struct folio *folio, struct folio *prev_folio,
- size_t prev_folio_size, int rw)
+static bool swap_bdev_can_merge(struct folio *folio, swp_entry_t entry,
+ struct swap_iocb *sio, int rw)
{
- if (swap_folio_sector(folio) !=
- swap_folio_sector(prev_folio) + (prev_folio_size >> SECTOR_SHIFT))
+ if (swap_entry_sector(entry) !=
+ swap_entry_sector(sio->entry) + (sio->len >> SECTOR_SHIFT))
return false;
- if (rw == WRITE && !folio_blkg_can_merge(folio, prev_folio))
+ if (rw == WRITE &&
+ !folio_blkg_can_merge(folio,
+ bvec_folio(&sio->bvecs[sio->nr_bvecs - 1])))
return false;
return true;
}
@@ -676,7 +676,7 @@ void swap_fs_prepare_rw(struct swap_io_ctx *ctx, int rw, struct iov_iter *iter)
struct swap_iocb *sio = ctx->sio;
init_sync_kiocb(&sio->iocb, ctx->sis->swap_file);
- sio->iocb.ki_pos = swap_dev_pos(bvec_folio(&sio->bvecs[0])->swap);
+ sio->iocb.ki_pos = swap_dev_pos(sio->entry);
if (rw == WRITE)
sio->iocb.ki_complete = swap_fs_write_complete;
else
@@ -687,11 +687,10 @@ void swap_fs_prepare_rw(struct swap_io_ctx *ctx, int rw, struct iov_iter *iter)
}
EXPORT_SYMBOL_GPL(swap_fs_prepare_rw);
-bool swap_fs_can_merge(struct folio *folio, struct folio *prev_folio,
- size_t prev_folio_size, int rw)
+bool swap_fs_can_merge(struct folio *folio, swp_entry_t entry,
+ struct swap_iocb *sio, int rw)
{
- return swap_dev_pos(folio->swap) ==
- swap_dev_pos(prev_folio->swap) + prev_folio_size;
+ return swap_dev_pos(entry) == swap_dev_pos(sio->entry) + sio->len;
}
EXPORT_SYMBOL_GPL(swap_fs_can_merge);
diff --git a/mm/swap.h b/mm/swap.h
index b3b54c28929a..14b8c6b9d468 100644
--- a/mm/swap.h
+++ b/mm/swap.h
@@ -258,7 +258,8 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct folio *folio);
void swap_read_submit(struct swap_io_ctx *ctx);
void swap_write_submit(struct swap_io_ctx *ctx);
int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio);
-void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio);
+void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio,
+ swp_entry_t entry);
/* linux/mm/swap_state.c */
extern struct address_space swap_space __read_mostly;
diff --git a/mm/swapfile.c b/mm/swapfile.c
index ab64bfc4e0f7..29118082cd7a 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -452,6 +452,19 @@ sector_t swap_folio_sector(struct folio *folio)
return sector << (PAGE_SHIFT - 9);
}
+sector_t swap_entry_sector(swp_entry_t entry)
+{
+ struct swap_info_struct *sis = __swap_entry_to_info(entry);
+ struct swap_extent *se;
+ sector_t sector;
+ pgoff_t offset;
+
+ offset = swp_offset(entry);
+ se = offset_to_swap_extent(sis, offset);
+ sector = se->start_block + (offset - se->start_page);
+ return sector << (PAGE_SHIFT - 9);
+}
+
/*
* swap allocation tell device that a cluster of swap can now be discarded,
* to allow the swap device to optimize its wear-levelling.
diff --git a/mm/zswap.c b/mm/zswap.c
index 96fb993d18cb..b4ffd6fb83eb 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1073,7 +1073,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
folio_set_reclaim(folio);
/* start writeback */
- __swap_writeout(&ctx, folio);
+ __swap_writeout(&ctx, folio, folio->swap);
swap_write_submit(&ctx);
out:
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 02/17] mm, swap: tag a swap table entry with its owning xswap entry
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
2026-09-20 7:20 ` [RFC PATCH 01/17] mm, swap: prepare the swap IO path for xswap backends Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 03/17] mm, swap: prepare the folio-less allocation path for xswap Baoquan He
` (16 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
An xswap slot can be written out to a physical device. The physical
slot then holds its data. That slot has no folio in the swap cache.
Record the xswap entry in the physical slot's swap table entry, so the
physical side can find it again.
A swap table entry can tell the type by its low three bits. Now 0b100
is left for extension. Use it to point at the owner. The xswap type and
offset go in the bits above.
Two paths hand out or cache a slot. They know only about the other two
kinds. Teach them to skip a pointer entry. cluster_scan_range() must
not hand it out as free. __swap_cache_add_check() must not put a folio
over it. Readahead on a physical device can otherwise walk into it.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swap_state.c | 6 +++++-
mm/swap_table.h | 53 +++++++++++++++++++++++++++++++++++++++++++++++++
mm/swapfile.c | 3 +++
3 files changed, 61 insertions(+), 1 deletion(-)
diff --git a/mm/swap_state.c b/mm/swap_state.c
index 8bba3e533b28..135d9573a9dc 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -179,6 +179,9 @@ static int __swap_cache_add_check(struct swap_cluster_info *ci,
return -ENOENT;
ci_off = swp_cluster_offset(targ_entry);
old_tb = __swap_table_get(ci, ci_off);
+ /* Readahead of a physical device can hit an xswap-backing slot. */
+ if (swp_tb_is_pointer(old_tb))
+ return -ENOENT;
if (swp_tb_is_folio(old_tb))
return -EEXIST;
if (!__swp_tb_get_count(old_tb))
@@ -196,7 +199,8 @@ static int __swap_cache_add_check(struct swap_cluster_info *ci,
ci_end = ci_off + nr;
do {
old_tb = __swap_table_get(ci, ci_off);
- if (unlikely(swp_tb_is_folio(old_tb) ||
+ if (unlikely(swp_tb_is_pointer(old_tb) ||
+ swp_tb_is_folio(old_tb) ||
!__swp_tb_get_count(old_tb) ||
is_zero != __swap_table_test_zero(ci, ci_off) ||
(memcg_id && *memcg_id != __swap_cgroup_get(ci, ci_off))))
diff --git a/mm/swap_table.h b/mm/swap_table.h
index e6613e62f8d0..524d397a995a 100644
--- a/mm/swap_table.h
+++ b/mm/swap_table.h
@@ -81,6 +81,59 @@ struct swap_memcg_table {
/* Bad slot: ends with 0b1000 and rests of bits are all 1 */
#define SWP_TB_BAD ((~0UL) << 3)
+#ifdef CONFIG_XSWAP
+/*
+ * Pointer-tagged swap table entry: the reverse map from a physical slot to
+ * the xswap entry it backs. Layout:
+ *
+ * Pointer: | xswap_type(8) | xswap_offset(53) |100|
+ *
+ * The low three bits identify the entry: 0b000 is a free or bad slot, a
+ * shadow entry ends in 0b01 and a cached folio in 0b10, which leaves
+ * 0b100 as the only marker still free.
+ */
+#define SWP_TB_PTR_MARK 0b100UL
+#define SWP_TB_PTR_OFF_BITS 53
+#define SWP_TB_PTR_TYPE_BITS 8
+#define SWP_TB_PTR_OFF_SHIFT 3
+#define SWP_TB_PTR_TYPE_SHIFT (SWP_TB_PTR_OFF_SHIFT + \
+ SWP_TB_PTR_OFF_BITS)
+#define SWP_TB_PTR_OFF_MASK ((1UL << SWP_TB_PTR_OFF_BITS) - 1)
+#define SWP_TB_PTR_TYPE_MASK ((1UL << SWP_TB_PTR_TYPE_BITS) - 1)
+
+static inline bool swp_tb_is_pointer(unsigned long swp_tb)
+{
+ return (swp_tb & (BIT(3) - 1)) == SWP_TB_PTR_MARK;
+}
+
+static inline unsigned long xswap_entry_to_rmap(swp_entry_t entry)
+{
+ unsigned long type = swp_type(entry);
+ unsigned long off = swp_offset(entry);
+
+ VM_WARN_ON_ONCE(type > SWP_TB_PTR_TYPE_MASK);
+ VM_WARN_ON_ONCE(off > SWP_TB_PTR_OFF_MASK);
+ return (type << SWP_TB_PTR_TYPE_SHIFT) |
+ (off << SWP_TB_PTR_OFF_SHIFT) |
+ SWP_TB_PTR_MARK;
+}
+
+static inline swp_entry_t xswap_rmap_to_entry(unsigned long swp_tb)
+{
+ unsigned long type = (swp_tb >> SWP_TB_PTR_TYPE_SHIFT) &
+ SWP_TB_PTR_TYPE_MASK;
+ unsigned long off = (swp_tb >> SWP_TB_PTR_OFF_SHIFT) &
+ SWP_TB_PTR_OFF_MASK;
+
+ return swp_entry(type, off);
+}
+#else /* !CONFIG_XSWAP */
+static inline bool swp_tb_is_pointer(unsigned long swp_tb)
+{
+ return false;
+}
+#endif /* CONFIG_XSWAP */
+
/* Macro for shadow offset calculation */
#define SWAP_COUNT_SHIFT SWP_TB_FLAGS_BITS
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 29118082cd7a..ad6703aadd3d 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -1032,6 +1032,9 @@ static bool cluster_scan_range(struct swap_info_struct *si,
swp_tb = __swap_table_get(ci, ci_off);
if (swp_tb_is_null(swp_tb))
continue;
+ /* Backs an xswap entry: the slot is not allocatable. */
+ if (swp_tb_is_pointer(swp_tb))
+ return false;
if (swp_tb_is_folio(swp_tb) && !__swp_tb_get_count(swp_tb)) {
if (!vm_swap_full())
return false;
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 03/17] mm, swap: prepare the folio-less allocation path for xswap
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
2026-09-20 7:20 ` [RFC PATCH 01/17] mm, swap: prepare the swap IO path for xswap backends Baoquan He
2026-09-20 7:20 ` [RFC PATCH 02/17] mm, swap: tag a swap table entry with its owning xswap entry Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 04/17] mm, swap: add a physical backend for xswap slots Baoquan He
` (15 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
The xswap backend writes to a slot with no folio of its own; only
hibernation allocated such slots before, so the NULL-folio case must be
allowed more widely. The slot also comes from a lower-priority device,
so keep that cluster out of this CPU's swapout cache, or later swapouts
would follow the backend device. Pack the slot into a used cluster
rather than a free one, which a large order may need.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 35 +++++++++++++++++++++++++----------
1 file changed, 25 insertions(+), 10 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index ad6703aadd3d..69ffa9a7a646 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -1068,8 +1068,10 @@ static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
* Such swap slots starts with count == 0 and will be increased
* upon folio unmap.
*
- * Else, it's a exclusive order 0 allocation for hibernation.
- * The slot starts with count == 1 and never increases.
+ * Else, it's an exclusive order 0 allocation for a slot that has
+ * no folio of its own: hibernation, and the physical slot that
+ * backs an xswap entry. The slot starts with count == 1 and
+ * never increases.
*/
if (likely(folio)) {
order = folio_order(folio);
@@ -1077,16 +1079,12 @@ static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
swap_cluster_assert_empty(ci, ci_off, nr_pages, false);
__swap_cache_add_folio(ci, folio, swp_entry(si->type,
ci_off + cluster_offset(si, ci)));
- } else if (IS_ENABLED(CONFIG_HIBERNATION)) {
+ } else {
order = 0;
nr_pages = 1;
swap_cluster_assert_empty(ci, ci_off, 1, false);
- /* Fake shadow placeholder with no flag, hibernation does not use the zeromap */
+ /* Fake shadow placeholder with no flag; no zeromap is used. */
__swap_table_set(ci, ci_off, __swp_tb_mk_count(shadow_to_swp_tb(NULL, 0), 1));
- } else {
- /* Allocation without folio is only possible with hibernation */
- WARN_ON_ONCE(1);
- return false;
}
/*
@@ -1149,8 +1147,15 @@ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,
swap_cluster_unlock(ci);
rcu_read_unlock();
if (si->flags & SWP_SOLIDSTATE) {
- this_cpu_write(percpu_swap_cluster.offset[order], next);
- this_cpu_write(percpu_swap_cluster.si[order], si);
+ /*
+ * A NULL folio is a backend or hibernation allocation: it must
+ * not displace the swapout cluster cached here, which
+ * swap_alloc_fast() reuses without a device check.
+ */
+ if (folio) {
+ this_cpu_write(percpu_swap_cluster.offset[order], next);
+ this_cpu_write(percpu_swap_cluster.si[order], si);
+ }
} else {
si->global_cluster->next[order] = next;
}
@@ -1279,6 +1284,16 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
}
new_cluster:
+ /*
+ * A slot with no folio packs into a cluster already in use: a free
+ * cluster is the only thing a large order can allocate from.
+ */
+ if (!folio && order < PMD_ORDER) {
+ found = alloc_swap_scan_list(si, &si->frag_clusters[order], folio, false);
+ if (found)
+ goto done;
+ }
+
/*
* If the device need discard, prefer new cluster over nonfull
* to spread out the writes.
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 04/17] mm, swap: add a physical backend for xswap slots
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (2 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 03/17] mm, swap: prepare the folio-less allocation path for xswap Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 05/17] mm, swap: use the xswap physical backend Baoquan He
` (14 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
An xswap slot keeps its data only in the zswap pool. When the pool
drops the entry, the data is gone. So give each xswap slot a physical
backend on the highest-priority real swap device. Each backend is
recorded in ci->xs_table[], which is allocated on first use. Zero
means no backend.
Freeing an xswap slot also releases its physical slot. But the record
can be out of date, so first check that the physical slot still points
back to this xswap entry. It may have been freed and given to another
entry. That entry's data must not be dropped.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swap.h | 24 +++++
mm/swapfile.c | 286 ++++++++++++++++++++++++++++++++++++++++++++++++++
2 files changed, 310 insertions(+)
diff --git a/mm/swap.h b/mm/swap.h
index 14b8c6b9d468..30b67915c459 100644
--- a/mm/swap.h
+++ b/mm/swap.h
@@ -58,6 +58,9 @@ struct swap_cluster_info {
u8 order;
atomic_long_t __rcu *table; /* Swap table entries, see mm/swap_table.h */
unsigned int *extend_table; /* For large swap count, protected by ci->lock */
+#ifdef CONFIG_XSWAP
+ unsigned long *xs_table; /* Physical backend per slot, lazily allocated */
+#endif
#ifdef CONFIG_MEMCG
struct swap_memcg_table *memcg_table; /* Swap table entries' cgroup record */
#endif
@@ -470,4 +473,25 @@ extern const struct swap_ops swap_bdev_ops;
int shmem_writeout(struct swap_io_ctx *ctx, struct folio *folio,
struct list_head *folio_list);
+#ifdef CONFIG_XSWAP
+swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci, unsigned int slot);
+swp_entry_t xswap_backend_alloc(swp_entry_t entry);
+void xswap_backend_free(swp_entry_t entry, swp_entry_t phys);
+#else
+static inline swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci,
+ unsigned int slot)
+{
+ return (swp_entry_t){};
+}
+
+static inline swp_entry_t xswap_backend_alloc(swp_entry_t entry)
+{
+ return (swp_entry_t){};
+}
+
+static inline void xswap_backend_free(swp_entry_t entry, swp_entry_t phys)
+{
+}
+#endif /* CONFIG_XSWAP */
+
#endif /* _MM_SWAP_H */
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 69ffa9a7a646..7b44458472a1 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -73,6 +73,11 @@ static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data);
static void xswap_try_shrink(struct swap_info_struct *si);
static int xswap_dev_kobj_add(struct swap_info_struct *si);
static void xswap_dev_kobj_del(struct swap_info_struct *si);
+static void xswap_free_phys_slot(struct swap_info_struct *si,
+ swp_entry_t entry, swp_entry_t phys);
+static void xswap_release_slot_backend(struct swap_info_struct *si,
+ struct swap_cluster_info *ci,
+ unsigned int slot);
static int xswap_create(int prio);
static int xswap_destroy(int type);
@@ -2164,6 +2169,11 @@ void __swap_cluster_free_entries(struct swap_info_struct *si,
*/
VM_WARN_ON(!swp_tb_is_shadow(old_tb) || __swp_tb_get_count(old_tb) > 1);
+#ifdef CONFIG_XSWAP
+ if (ci->xs_table && ci->xs_table[ci_off])
+ xswap_release_slot_backend(si, ci, ci_off);
+#endif
+
/* Resetting the slot to NULL also clears the inline flags. */
__swap_table_set(ci, ci_off, null_to_swp_tb());
if (!SWAP_TABLE_HAS_ZEROFLAG)
@@ -3066,6 +3076,69 @@ static unsigned int find_next_to_unuse(struct swap_info_struct *si,
return 0;
}
+
+#ifdef CONFIG_XSWAP
+/*
+ * Free the physical slot @phys and clear its reverse mapping. @entry is the
+ * xswap entry that recorded @phys. The slot is only freed while it still
+ * says it belongs to @entry: our record can go stale, and @phys may by then
+ * have been handed to another xswap entry, whose data must not be dropped.
+ *
+ * free_cluster()/partial_free_cluster() require the cluster lock.
+ */
+static void xswap_free_phys_slot(struct swap_info_struct *si,
+ swp_entry_t entry, swp_entry_t phys)
+{
+ struct swap_cluster_info *ci;
+ unsigned long swp_tb;
+ unsigned int offset = swp_offset(phys);
+
+ ci = swap_cluster_lock(si, offset);
+ if (!ci)
+ return;
+
+ swp_tb = __swap_table_get(ci, offset % SWAPFILE_CLUSTER);
+ if (!swp_tb_is_pointer(swp_tb) ||
+ xswap_rmap_to_entry(swp_tb).val != entry.val) {
+ /* Already freed, or recycled for another xswap entry. */
+ VM_WARN_ON_ONCE(!swp_tb_is_null(swp_tb) &&
+ !swp_tb_is_pointer(swp_tb));
+ swap_cluster_unlock(ci);
+ return;
+ }
+
+ __swap_table_set(ci, offset % SWAPFILE_CLUSTER, null_to_swp_tb());
+ ci->count--;
+ if (!ci->count)
+ free_cluster(si, ci);
+ else
+ partial_free_cluster(si, ci);
+ swap_cluster_unlock(ci);
+
+ swap_range_free(si, offset, 1);
+}
+
+/*
+ * An xswap slot being freed may still own a physical slot. Drop it, or the
+ * physical space and its reverse mapping leak. Caller holds the xswap
+ * cluster lock; the physical cluster lock is only ever taken below it.
+ */
+static void xswap_release_slot_backend(struct swap_info_struct *si,
+ struct swap_cluster_info *ci,
+ unsigned int slot)
+{
+ swp_entry_t entry = swp_entry(si->type, cluster_offset(si, ci) + slot);
+ struct swap_info_struct *psi;
+ swp_entry_t phys = { .val = ci->xs_table[slot] };
+
+ WRITE_ONCE(ci->xs_table[slot], 0);
+
+ psi = swap_type_to_info(swp_type(phys));
+ if (psi)
+ xswap_free_phys_slot(psi, entry, phys);
+}
+
+#endif /* CONFIG_XSWAP */
static int try_to_unuse(unsigned int type)
{
struct mm_struct *prev_mm;
@@ -3973,6 +4046,211 @@ static unsigned long read_swap_header(struct swap_info_struct *si,
}
#ifdef CONFIG_XSWAP
+/* The physical slot backing @slot, or zero. */
+swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci, unsigned int slot)
+{
+ unsigned long *xs = READ_ONCE(ci->xs_table);
+
+ if (!xs || !xs[slot])
+ return (swp_entry_t){};
+
+ return (swp_entry_t){ .val = xs[slot] };
+}
+
+/* Cluster lock held. xswap never unmaps a cluster below the mapped prefix. */
+static unsigned long *xswap_xs_table_locked(struct swap_cluster_info *ci,
+ gfp_t gfp)
+{
+ unsigned long *xs = ci->xs_table;
+
+ if (xs)
+ return xs;
+
+ if (gfp & __GFP_DIRECT_RECLAIM) {
+ spin_unlock(&ci->lock);
+ xs = kcalloc(SWAPFILE_CLUSTER, sizeof(*xs), gfp);
+ spin_lock(&ci->lock);
+ } else {
+ xs = kcalloc(SWAPFILE_CLUSTER, sizeof(*xs), gfp);
+ }
+
+ if (!xs)
+ return NULL;
+ if (cmpxchg(&ci->xs_table, NULL, xs)) {
+ kfree(xs);
+ xs = ci->xs_table;
+ }
+ return xs;
+}
+
+/*
+ * 4K per cluster, and this runs in reclaim, so it has to be able to sleep.
+ */
+static bool xswap_slot_set_backend(struct swap_cluster_info *ci,
+ unsigned int slot, swp_entry_t phys)
+{
+ unsigned long *xs = xswap_xs_table_locked(ci, GFP_KERNEL | __GFP_HIGH |
+ __GFP_NOMEMALLOC);
+
+ if (!xs)
+ return false;
+ WRITE_ONCE(xs[slot], phys.val);
+ return true;
+}
+
+/*
+ * Drop the table of cluster @idx and every backend it still records. The
+ * cluster is going away, so nothing else will: the physical slots would stay
+ * allocated and their reverse mappings would keep pointing at an xswap entry
+ * whose cluster_info is about to be unmapped. Caller holds ci->lock.
+ */
+static void xswap_free_xs_table(struct swap_info_struct *si, unsigned long idx)
+{
+ struct swap_cluster_info *ci = &si->cluster_info[idx];
+ unsigned long *xs = ci->xs_table;
+ unsigned int slot;
+
+ if (!xs)
+ return;
+
+ for (slot = 0; slot < SWAPFILE_CLUSTER; slot++) {
+ if (xs[slot])
+ xswap_release_slot_backend(si, ci, slot);
+ }
+
+ kfree(xs);
+ ci->xs_table = NULL;
+}
+
+/* Order-0 slot from @si, no folio attached; goes through the per-CPU cache. */
+static unsigned long xswap_alloc_slot_local(struct swap_info_struct *si)
+{
+ struct swap_cluster_info *ci;
+ unsigned long pcp_offset, offset = SWAP_ENTRY_INVALID;
+
+ local_lock(&percpu_swap_cluster.lock);
+ if (this_cpu_read(percpu_swap_cluster.si[0]) == si) {
+ pcp_offset = this_cpu_read(percpu_swap_cluster.offset[0]);
+ if (pcp_offset) {
+ ci = swap_cluster_lock(si, pcp_offset);
+ if (cluster_is_usable(ci, 0))
+ offset = alloc_swap_scan_cluster(si, ci, NULL,
+ pcp_offset);
+ else
+ swap_cluster_unlock(ci);
+ }
+ }
+ if (offset == SWAP_ENTRY_INVALID)
+ offset = cluster_alloc_swap_entry(si, NULL);
+ local_unlock(&percpu_swap_cluster.lock);
+
+ return offset;
+}
+
+static swp_entry_t xswap_alloc_phys_slot(void)
+{
+ struct swap_info_struct *devs[MAX_SWAPFILES], *si;
+ int nr = 0, i;
+
+ /*
+ * Snapshot the candidate devices in priority order. swap_info
+ * structs are never freed, so the pointers stay valid; liveness is
+ * re-checked below.
+ */
+ spin_lock(&swap_avail_lock);
+ plist_for_each_entry(si, &swap_avail_head, avail_list) {
+ if (si->flags & SWP_XSWAP)
+ continue;
+ if (nr == MAX_SWAPFILES)
+ break;
+ devs[nr++] = si;
+ }
+ spin_unlock(&swap_avail_lock);
+
+ for (i = 0; i < nr; i++) {
+ unsigned long offset;
+
+ if (!get_swap_device_info(devs[i]))
+ continue;
+ offset = xswap_alloc_slot_local(devs[i]);
+ put_swap_device(devs[i]);
+ if (offset && offset != SWAP_ENTRY_INVALID)
+ return swp_entry(devs[i]->type, offset);
+ }
+
+ return (swp_entry_t){};
+}
+
+/* Reverse mapping: record @entry as the owner of physical slot @phys. */
+static void xswap_install_rmap(swp_entry_t phys, swp_entry_t entry)
+{
+ struct swap_info_struct *si = swap_type_to_info(swp_type(phys));
+ struct swap_cluster_info *ci;
+ unsigned long off = swp_offset(phys);
+
+ if (!si)
+ return;
+ ci = swap_cluster_lock(si, off);
+ if (ci) {
+ __swap_table_set(ci, off % SWAPFILE_CLUSTER,
+ xswap_entry_to_rmap(entry));
+ swap_cluster_unlock(ci);
+ }
+}
+
+swp_entry_t xswap_backend_alloc(swp_entry_t entry)
+{
+ struct swap_info_struct *si = __swap_entry_to_info(entry);
+ struct swap_cluster_info *ci;
+ unsigned long off = swp_offset(entry);
+ swp_entry_t phys;
+ bool recorded = false;
+
+ phys = xswap_alloc_phys_slot();
+ if (!phys.val)
+ return phys;
+
+ xswap_install_rmap(phys, entry);
+
+ ci = swap_cluster_lock(si, off);
+ if (ci) {
+ recorded = xswap_slot_set_backend(ci, off % SWAPFILE_CLUSTER, phys);
+ swap_cluster_unlock(ci);
+ }
+
+ if (!recorded) {
+ /* The backend is unrecorded, so the data would be unreachable. */
+ xswap_backend_free(entry, phys);
+ return (swp_entry_t){};
+ }
+
+ return phys;
+}
+
+/* Undo xswap_backend_alloc(): the zswap copy is still in place. */
+void xswap_backend_free(swp_entry_t entry, swp_entry_t phys)
+{
+ struct swap_info_struct *si = __swap_entry_to_info(entry);
+ struct swap_info_struct *psi;
+ struct swap_cluster_info *ci;
+ unsigned long off = swp_offset(entry);
+
+ ci = swap_cluster_lock(si, off);
+ if (ci) {
+ unsigned int slot = off % SWAPFILE_CLUSTER;
+ unsigned long *xs = ci->xs_table;
+
+ if (xs && xs[slot] == phys.val)
+ WRITE_ONCE(xs[slot], 0);
+ swap_cluster_unlock(ci);
+ }
+
+ /* xswap_free_phys_slot() clears the reverse mapping. */
+ psi = swap_type_to_info(swp_type(phys));
+ if (psi)
+ xswap_free_phys_slot(psi, entry, phys);
+}
+
static int xswap_map_clusters(struct swap_info_struct *si,
unsigned long start_idx, unsigned long nr)
{
@@ -4135,6 +4413,14 @@ static int xswap_unmap_clusters_locked(struct swap_info_struct *si,
unsigned int noreclaim_flags;
int i;
+ for (unsigned long idx = start_idx; idx < start_idx + nr; idx++) {
+ struct swap_cluster_info *ci = &si->cluster_info[idx];
+
+ spin_lock(&ci->lock);
+ xswap_free_xs_table(si, idx);
+ spin_unlock(&ci->lock);
+ }
+
if (vm_start >= vm_end) {
WRITE_ONCE(si->nr_clusters_mapped, start_idx);
return 0;
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 05/17] mm, swap: use the xswap physical backend
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (3 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 04/17] mm, swap: add a physical backend for xswap slots Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 06/17] mm, swap: fall back to disk when zswap refuses an xswap page Baoquan He
` (13 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
An xswap entry that zswap evicts had nowhere to go, so
zswap_writeback_entry() refused to handle it. Send it to the physical
backend instead. Reserve the backend before decompressing the entry,
so a failed allocation wastes no work.
A slot that was written out no longer has a zswap copy, so
swap_read_folio() forwards its read to the physical entry rather than
dropping it.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/page_io.c | 28 ++++++++++++++++++++--------
mm/zswap.c | 20 +++++++++++++++-----
2 files changed, 35 insertions(+), 13 deletions(-)
diff --git a/mm/page_io.c b/mm/page_io.c
index 2b1387c87e51..32e426779dce 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -463,6 +463,7 @@ static bool swap_read_folio_zeromap(struct folio *folio)
void swap_read_folio(struct swap_io_ctx *ctx, struct folio *folio)
{
struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
+ swp_entry_t entry = folio->swap;
bool synchronous = sis->flags & SWP_SYNCHRONOUS_IO;
bool workingset = folio_test_workingset(folio);
unsigned long pflags;
@@ -492,18 +493,29 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct folio *folio)
goto finish;
if (unlikely(sis->flags & SWP_XSWAP)) {
- /*
- * An xswap entry only ever lives in zswap, so zswap_load()
- * must have found it. Unlock and let the caller retry.
- */
- WARN_ON_ONCE(1);
- folio_unlock(folio);
- goto finish;
+ struct swap_cluster_info *ci;
+ swp_entry_t phys = {};
+ unsigned long offset = swp_offset(folio->swap);
+
+ /* May have been written back; reuse the physical read path. */
+ ci = __swap_offset_to_cluster(sis, offset);
+ if (ci) {
+ spin_lock(&ci->lock);
+ phys = xswap_slot_backend(ci, offset % SWAPFILE_CLUSTER);
+ spin_unlock(&ci->lock);
+ }
+ if (!phys.val) {
+ /* No backend and no zswap entry: data is gone. */
+ folio_unlock(folio);
+ goto finish;
+ }
+
+ entry = phys;
}
/* We have to read from slower devices. Increase zswap protection. */
zswap_folio_swapin(folio);
- swap_add_folio(ctx, folio, folio->swap, READ);
+ swap_add_folio(ctx, folio, entry, READ);
finish:
if (workingset) {
diff --git a/mm/zswap.c b/mm/zswap.c
index b4ffd6fb83eb..24904b4c4dce 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1011,6 +1011,8 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
struct mempolicy *mpol;
struct swap_info_struct *si;
struct swap_io_ctx ctx = {};
+ swp_entry_t phys = {};
+ bool is_xswap;
int ret = 0;
/* try to allocate swap cache folio */
@@ -1018,10 +1020,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
if (IS_ERR_OR_NULL(si))
return -ENOENT;
- if (si->flags & SWP_XSWAP) {
- put_swap_device(si);
- return -EINVAL;
- }
+ is_xswap = !!(si->flags & SWP_XSWAP);
mpol = get_task_policy(current);
folio = swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol,
@@ -1053,6 +1052,15 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
goto out;
}
+ /* Reserve the destination before dropping the zswap copy. */
+ if (is_xswap) {
+ phys = xswap_backend_alloc(swpentry);
+ if (!phys.val) {
+ ret = -ENOMEM;
+ goto out;
+ }
+ }
+
if (!zswap_decompress(entry, folio)) {
ret = -EIO;
goto out;
@@ -1073,11 +1081,13 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
folio_set_reclaim(folio);
/* start writeback */
- __swap_writeout(&ctx, folio, folio->swap);
+ __swap_writeout(&ctx, folio, is_xswap ? phys : folio->swap);
swap_write_submit(&ctx);
out:
if (ret) {
+ if (is_xswap && phys.val)
+ xswap_backend_free(swpentry, phys);
swap_cache_del_folio(folio);
folio_unlock(folio);
}
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 06/17] mm, swap: fall back to disk when zswap refuses an xswap page
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (4 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 05/17] mm, swap: use the xswap physical backend Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 07/17] mm, swap: support swapoff of an xswap physical backend Baoquan He
` (12 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
zswap_store() can refuse a page if the pool may be at its limit, or
the page may not compress well enough. For an xswap device there was
no way to store it, so swap_writeout() put the page back on the LRU.
Under memory pressure that turns into a livelock: reclaim keeps picking
the same page and the pool keeps refusing it.
Now that xswap slots have a physical backend, write the page out
instead. If no backend can be reserved, keep the old behaviour.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/page_io.c | 14 ++++++++++++--
1 file changed, 12 insertions(+), 2 deletions(-)
diff --git a/mm/page_io.c b/mm/page_io.c
index 32e426779dce..d521ea335df1 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -253,8 +253,18 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
* yet, so look the device up from the entry.
*/
if (unlikely(__swap_entry_to_info(folio->swap)->flags & SWP_XSWAP)) {
- folio_mark_dirty(folio);
- return AOP_WRITEPAGE_ACTIVATE;
+ swp_entry_t entry = folio->swap;
+ swp_entry_t phys;
+
+ /* The pool refused it: fall back to a real swap device. */
+ phys = xswap_backend_alloc(entry);
+ if (!phys.val) {
+ folio_mark_dirty(folio);
+ return AOP_WRITEPAGE_ACTIVATE;
+ }
+
+ __swap_writeout(ctx, folio, phys);
+ return 0;
}
__swap_writeout(ctx, folio, folio->swap);
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 07/17] mm, swap: support swapoff of an xswap physical backend
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (5 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 06/17] mm, swap: fall back to disk when zswap refuses an xswap page Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 08/17] mm, swap: reclaim physical slots backing cache-only xswap entries Baoquan He
` (11 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
swapoff of a real swap device iterates its slots through try_to_unuse()
and reads each one back into memory. A slot that was written back
from an xswap entry has no folio in the swap cache, so the loop
skips it and the slot is freed without ever reading the data. While
the xswap entry will keep pointing at a slot that no longer holds
anything.
Now change try_to_unuse() to analyze such a slot through its reverse
mapping: bring the folio back, mark it dirty so that reclaim stores it
into zswap again, and only then drop the backend.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 97 ++++++++++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 96 insertions(+), 1 deletion(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 7b44458472a1..9afccdc03845 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -44,6 +44,7 @@
#include <linux/suspend.h>
#include <linux/zswap.h>
#include <linux/plist.h>
+#include <linux/swap_ops.h>
#include <asm/tlbflush.h>
#include <linux/leafops.h>
@@ -3138,6 +3139,96 @@ static void xswap_release_slot_backend(struct swap_info_struct *si,
xswap_free_phys_slot(psi, entry, phys);
}
+/*
+ * A written-back slot has no folio of its own here, so the unuse loop would
+ * skip it. Resolve it through the reverse mapping and free both sides.
+ */
+static int xswap_unuse_rmap(struct swap_info_struct *si, unsigned long offset)
+{
+ struct swap_cluster_info *pci, *xci;
+ struct swap_info_struct *xsi;
+ struct swap_io_ctx ctx = {};
+ struct mempolicy *mpol;
+ struct folio *folio;
+ unsigned long swp_tb, xoff;
+ swp_entry_t xentry, phys;
+
+ phys = swp_entry(si->type, offset);
+
+ pci = swap_cluster_lock(si, offset);
+ if (!pci)
+ return 0;
+ swp_tb = __swap_table_get(pci, offset % SWAPFILE_CLUSTER);
+ if (!swp_tb_is_pointer(swp_tb)) {
+ swap_cluster_unlock(pci);
+ return 0;
+ }
+ xentry = xswap_rmap_to_entry(swp_tb);
+ swap_cluster_unlock(pci);
+
+ xoff = swp_offset(xentry);
+
+ /*
+ * Pin the owner before reading its cluster_info:
+ * free_swap_cluster_info() only unmaps that after killing si->users.
+ * No owner left means the slot is dropped below.
+ */
+ xsi = swap_type_to_info(swp_type(xentry));
+ if (!xsi || !(xsi->flags & SWP_XSWAP) ||
+ xoff >= READ_ONCE(xsi->nr_clusters_mapped) * SWAPFILE_CLUSTER ||
+ !get_swap_device_info(xsi))
+ goto out_free;
+
+ /*
+ * Bring the data back before the slot goes. A dirty folio is what
+ * reclaim stores into zswap next time.
+ */
+ folio = swap_cache_get_folio(xentry);
+ if (!folio) {
+ mpol = get_task_policy(current);
+ folio = swap_cache_alloc_folio(xentry, GFP_HIGHUSER_MOVABLE,
+ BIT(0), NULL, mpol,
+ NO_INTERLEAVE_INDEX);
+ if (IS_ERR(folio)) {
+ put_swap_device(xsi);
+ return PTR_ERR(folio) == -ENOMEM ? -ENOMEM : 0;
+ }
+ swap_read_folio(&ctx, folio);
+ swap_read_submit(&ctx);
+ folio_lock(folio);
+ } else {
+ folio_lock(folio);
+ }
+
+ if (folio_matches_swap_entry(folio, xentry)) {
+ folio_wait_writeback(folio);
+ if (unlikely(!folio_test_uptodate(folio)))
+ swap_cache_del_folio(folio);
+ else
+ folio_mark_dirty(folio);
+ }
+ folio_unlock(folio);
+ folio_put(folio);
+
+ /* The slot is going away, so our record of it is stale either way. */
+ xci = swap_cluster_lock(xsi, xoff);
+ if (xci) {
+ if (xswap_slot_backend(xci, xoff % SWAPFILE_CLUSTER).val == phys.val)
+ WRITE_ONCE(xci->xs_table[xoff % SWAPFILE_CLUSTER], 0);
+ swap_cluster_unlock(xci);
+ }
+ put_swap_device(xsi);
+out_free:
+ /* Leaving the mapping in place would make the unuse loop spin. */
+ xswap_free_phys_slot(si, xentry, phys);
+ return 0;
+}
+#else
+static inline int xswap_unuse_rmap(struct swap_info_struct *si,
+ unsigned long offset)
+{
+ return 0;
+}
#endif /* CONFIG_XSWAP */
static int try_to_unuse(unsigned int type)
{
@@ -3197,8 +3288,12 @@ static int try_to_unuse(unsigned int type)
entry = swp_entry(type, i);
folio = swap_cache_get_folio(entry);
- if (!folio)
+ if (!folio) {
+ retval = xswap_unuse_rmap(si, i);
+ if (retval)
+ return retval;
continue;
+ }
/*
* It is conceivable that a racing task removed this folio from
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 08/17] mm, swap: reclaim physical slots backing cache-only xswap entries
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (6 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 07/17] mm, swap: support swapoff of an xswap physical backend Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 09/17] mm, swap: back a large xswap folio with a contiguous physical run Baoquan He
` (10 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
An xswap entry's swap count can reach 0 while its folio is still in the
swap cache. Then the entry is cache-only, and its physical slot is
redundant. But the slot is only freed when the entry itself goes away.
A folio can stay in the swap cache for a long time, so the slot stays
pinned.
So reclaim these slots from the physical scanner. It already walks
full clusters when swap is more than half full. A slot with the
reverse mapping belongs to an xswap entry. The scanner can hand that
entry to __try_to_reclaim_swap(). That frees the folio from the swap
cache once no page table reference is left. The physical slot is
released with it.
So __try_to_reclaim_swap() now takes the entry, not a device and an
offset. The entry names its own device. For existing callers the two
are the same.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 42 +++++++++++++++++++++++++++++++++---------
1 file changed, 33 insertions(+), 9 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 9afccdc03845..875b40f2e362 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -311,13 +311,17 @@ static bool swap_only_has_cache(struct swap_cluster_info *ci,
* returns number of pages in the folio that backs the swap entry. If positive,
* the folio was reclaimed. If negative, the folio was not reclaimed. If 0, no
* folio was associated with the swap entry.
+ *
+ * @entry names its own device, which may differ from the device being scanned:
+ * a physical slot is backed by an xswap entry, and it is that entry's folio
+ * and cluster the reclaim works on.
*/
-static int __try_to_reclaim_swap(struct swap_info_struct *si,
- unsigned long offset, unsigned long flags)
+static int __try_to_reclaim_swap(swp_entry_t entry, unsigned long flags)
{
- const swp_entry_t entry = swp_entry(si->type, offset);
+ struct swap_info_struct *si = __swap_entry_to_info(entry);
struct swap_cluster_info *ci;
struct folio *folio;
+ unsigned long offset;
int ret, nr_pages;
bool need_reclaim;
@@ -993,7 +997,8 @@ static bool cluster_reclaim_range(struct swap_info_struct *si,
if (swp_tb_get_count(swp_tb))
break;
if (swp_tb_is_folio(swp_tb))
- if (__try_to_reclaim_swap(si, offset, TTRS_ANYWAY) < 0)
+ if (__try_to_reclaim_swap(swp_entry(si->type, offset),
+ TTRS_ANYWAY) < 0)
break;
} while (++offset < end);
spin_lock(&ci->lock);
@@ -1213,13 +1218,31 @@ static void swap_reclaim_full_clusters(struct swap_info_struct *si, bool force)
swp_tb = swap_table_get(ci, offset % SWAPFILE_CLUSTER);
if (swp_tb_is_folio(swp_tb) && !__swp_tb_get_count(swp_tb)) {
spin_unlock(&ci->lock);
- nr_reclaim = __try_to_reclaim_swap(si, offset,
- TTRS_ANYWAY);
+ nr_reclaim = __try_to_reclaim_swap(
+ swp_entry(si->type, offset),
+ TTRS_ANYWAY);
spin_lock(&ci->lock);
if (nr_reclaim) {
offset += abs(nr_reclaim);
continue;
}
+#ifdef CONFIG_XSWAP
+ } else if (swp_tb_is_pointer(swp_tb)) {
+ /*
+ * The slot is backed by an xswap entry, and
+ * that entry's swap count decides whether the
+ * slot can go.
+ */
+ spin_unlock(&ci->lock);
+ nr_reclaim = __try_to_reclaim_swap(
+ xswap_rmap_to_entry(swp_tb),
+ TTRS_ANYWAY);
+ spin_lock(&ci->lock);
+ if (nr_reclaim) {
+ offset += abs(nr_reclaim);
+ continue;
+ }
+#endif
}
offset++;
}
@@ -1854,8 +1877,9 @@ static void swap_put_entries_cluster(struct swap_info_struct *si,
return;
do {
- nr_reclaimed = __try_to_reclaim_swap(si, offset,
- TTRS_UNMAPPED | TTRS_FULL);
+ nr_reclaimed = __try_to_reclaim_swap(
+ swp_entry(si->type, offset),
+ TTRS_UNMAPPED | TTRS_FULL);
offset++;
if (nr_reclaimed)
offset = round_up(offset, abs(nr_reclaimed));
@@ -2459,7 +2483,7 @@ void swap_free_hibernation_slot(swp_entry_t entry)
swap_cluster_unlock(ci);
/* In theory readahead might add it to the swap cache by accident */
- __try_to_reclaim_swap(si, offset, TTRS_ANYWAY);
+ __try_to_reclaim_swap(swp_entry(si->type, offset), TTRS_ANYWAY);
}
static int __find_hibernation_swap_type(dev_t device, sector_t offset)
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 09/17] mm, swap: back a large xswap folio with a contiguous physical run
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (7 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 08/17] mm, swap: reclaim physical slots backing cache-only xswap entries Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 10/17] mm, swap: enable THP swapin for xswap entries Baoquan He
` (9 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
When zswap refuses a folio, the folio goes to a physical swap device.
A small folio takes one slot there. A large folio is different as it must
sit in a contiguous run of slots. This is because swap_read_folio()
reads it back with one IO that starts at its first physical slot.
So the whole run has to be reserved at once. The problem is that the
reservation path has no folio, so it only knows order 0. Pass the order
down through cluster_alloc_swap_entry(), and let it fill a range of
shadow entries.
The xswap device must accept large orders too. Otherwise a large anon
folio cannot be swapped out to it at all.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/page_io.c | 2 +-
mm/swap.h | 8 +--
mm/swapfile.c | 164 +++++++++++++++++++++++++++-----------------------
mm/zswap.c | 4 +-
4 files changed, 95 insertions(+), 83 deletions(-)
diff --git a/mm/page_io.c b/mm/page_io.c
index d521ea335df1..dda68d7e1d94 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -257,7 +257,7 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
swp_entry_t phys;
/* The pool refused it: fall back to a real swap device. */
- phys = xswap_backend_alloc(entry);
+ phys = xswap_backend_alloc(entry, folio_order(folio));
if (!phys.val) {
folio_mark_dirty(folio);
return AOP_WRITEPAGE_ACTIVATE;
diff --git a/mm/swap.h b/mm/swap.h
index 30b67915c459..4fcbb21a583d 100644
--- a/mm/swap.h
+++ b/mm/swap.h
@@ -475,8 +475,8 @@ int shmem_writeout(struct swap_io_ctx *ctx, struct folio *folio,
#ifdef CONFIG_XSWAP
swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci, unsigned int slot);
-swp_entry_t xswap_backend_alloc(swp_entry_t entry);
-void xswap_backend_free(swp_entry_t entry, swp_entry_t phys);
+swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order);
+void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order);
#else
static inline swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci,
unsigned int slot)
@@ -484,12 +484,12 @@ static inline swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci,
return (swp_entry_t){};
}
-static inline swp_entry_t xswap_backend_alloc(swp_entry_t entry)
+static inline swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order)
{
return (swp_entry_t){};
}
-static inline void xswap_backend_free(swp_entry_t entry, swp_entry_t phys)
+static inline void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order)
{
}
#endif /* CONFIG_XSWAP */
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 875b40f2e362..4d72c854c7b3 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -1063,10 +1063,10 @@ static bool cluster_scan_range(struct swap_info_struct *si,
static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
struct swap_cluster_info *ci,
struct folio *folio,
- unsigned int ci_off)
+ unsigned int ci_off, unsigned int order)
{
- unsigned int order;
- unsigned long nr_pages;
+ unsigned long nr_pages = 1UL << order;
+ unsigned long i;
lockdep_assert_held(&ci->lock);
@@ -1075,33 +1075,24 @@ static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
/*
* All mm swap allocation starts with a folio (folio_alloc_swap),
- * it's also the only allocation path for large orders allocation.
- * Such swap slots starts with count == 0 and will be increased
- * upon folio unmap.
+ * it's also the allocation path for large orders. Such swap slots
+ * starts with count == 0 and will be increased upon folio unmap.
*
- * Else, it's an exclusive order 0 allocation for a slot that has
- * no folio of its own: hibernation, and the physical slot that
- * backs an xswap entry. The slot starts with count == 1 and
- * never increases.
+ * Else, it's an exclusive allocation for a range of slots that has no
+ * folio of its own: hibernation, and the range that backs an xswap
+ * folio. Each slot starts with count == 1 and never increases.
*/
+ swap_cluster_assert_empty(ci, ci_off, nr_pages, false);
if (likely(folio)) {
- order = folio_order(folio);
- nr_pages = 1 << order;
- swap_cluster_assert_empty(ci, ci_off, nr_pages, false);
__swap_cache_add_folio(ci, folio, swp_entry(si->type,
ci_off + cluster_offset(si, ci)));
} else {
- order = 0;
- nr_pages = 1;
- swap_cluster_assert_empty(ci, ci_off, 1, false);
- /* Fake shadow placeholder with no flag; no zeromap is used. */
- __swap_table_set(ci, ci_off, __swp_tb_mk_count(shadow_to_swp_tb(NULL, 0), 1));
+ /* Fake shadow placeholders with no flag; no zeromap is used. */
+ for (i = 0; i < nr_pages; i++)
+ __swap_table_set(ci, ci_off + i,
+ __swp_tb_mk_count(shadow_to_swp_tb(NULL, 0), 1));
}
- /*
- * The first allocation in a cluster makes the
- * cluster exclusive to this order
- */
if (cluster_is_empty(ci))
ci->order = order;
ci->count += nr_pages;
@@ -1109,17 +1100,15 @@ static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
return true;
}
-
-/* Try use a new cluster for current CPU and allocate from it. */
static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,
struct swap_cluster_info *ci,
- struct folio *folio, unsigned long offset)
+ struct folio *folio, unsigned long offset,
+ unsigned int order)
{
unsigned int next = SWAP_ENTRY_INVALID, found = SWAP_ENTRY_INVALID;
unsigned long start = ALIGN_DOWN(offset, SWAPFILE_CLUSTER);
- unsigned int order = likely(folio) ? folio_order(folio) : 0;
unsigned long end = start + SWAPFILE_CLUSTER;
- unsigned int nr_pages = 1 << order;
+ unsigned long nr_pages = 1UL << order;
bool need_reclaim, ret, usable;
lockdep_assert_held(&ci->lock);
@@ -1145,7 +1134,7 @@ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,
if (!ret)
continue;
}
- if (!__swap_cluster_alloc_entries(si, ci, folio, offset % SWAPFILE_CLUSTER))
+ if (!__swap_cluster_alloc_entries(si, ci, folio, offset % SWAPFILE_CLUSTER, order))
break;
found = offset;
offset += nr_pages;
@@ -1176,7 +1165,7 @@ static unsigned int alloc_swap_scan_cluster(struct swap_info_struct *si,
static unsigned int alloc_swap_scan_list(struct swap_info_struct *si,
struct list_head *list,
struct folio *folio,
- bool scan_all)
+ unsigned int order, bool scan_all)
{
unsigned int found = SWAP_ENTRY_INVALID;
@@ -1187,7 +1176,7 @@ static unsigned int alloc_swap_scan_list(struct swap_info_struct *si,
if (!ci)
break;
offset = cluster_offset(si, ci);
- found = alloc_swap_scan_cluster(si, ci, folio, offset);
+ found = alloc_swap_scan_cluster(si, ci, folio, offset, order);
if (found)
break;
} while (scan_all);
@@ -1279,17 +1268,16 @@ static void swap_reclaim_work(struct work_struct *work)
* cluster for current CPU too.
*/
static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
- struct folio *folio)
+ struct folio *folio, unsigned int order)
{
struct swap_cluster_info *ci;
- unsigned int order = likely(folio) ? folio_order(folio) : 0;
unsigned int offset = SWAP_ENTRY_INVALID, found = SWAP_ENTRY_INVALID;
/*
* Swapfile is not block device so unable
- * to allocate large entries.
+ * to allocate large entries. xswap has no backing file, so it can.
*/
- if (order && !(si->flags & SWP_BLKDEV))
+ if (order && !(si->flags & SWP_BLKDEV) && !(si->flags & SWP_XSWAP))
return 0;
if (!(si->flags & SWP_SOLIDSTATE)) {
@@ -1304,7 +1292,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
if (cluster_is_usable(ci, order)) {
if (cluster_is_empty(ci))
offset = cluster_offset(si, ci);
- found = alloc_swap_scan_cluster(si, ci, folio, offset);
+ found = alloc_swap_scan_cluster(si, ci, folio, offset, order);
} else {
swap_cluster_unlock(ci);
}
@@ -1318,7 +1306,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
* cluster is the only thing a large order can allocate from.
*/
if (!folio && order < PMD_ORDER) {
- found = alloc_swap_scan_list(si, &si->frag_clusters[order], folio, false);
+ found = alloc_swap_scan_list(si, &si->frag_clusters[order], folio, order, false);
if (found)
goto done;
}
@@ -1328,19 +1316,19 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
* to spread out the writes.
*/
if (si->flags & SWP_PAGE_DISCARD) {
- found = alloc_swap_scan_list(si, &si->free_clusters, folio, false);
+ found = alloc_swap_scan_list(si, &si->free_clusters, folio, order, false);
if (found)
goto done;
}
if (order < PMD_ORDER) {
- found = alloc_swap_scan_list(si, &si->nonfull_clusters[order], folio, true);
+ found = alloc_swap_scan_list(si, &si->nonfull_clusters[order], folio, order, true);
if (found)
goto done;
}
if (!(si->flags & SWP_PAGE_DISCARD)) {
- found = alloc_swap_scan_list(si, &si->free_clusters, folio, false);
+ found = alloc_swap_scan_list(si, &si->free_clusters, folio, order, false);
if (found)
goto done;
}
@@ -1356,7 +1344,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
* failure is not critical. Scanning one cluster still
* keeps the list rotated and reclaimed (for clean swap cache).
*/
- found = alloc_swap_scan_list(si, &si->frag_clusters[order], folio, false);
+ found = alloc_swap_scan_list(si, &si->frag_clusters[order], folio, order, false);
if (found)
goto done;
}
@@ -1370,11 +1358,11 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
* Clusters here have at least one usable slots and can't fail order 0
* allocation, but reclaim may drop si->lock and race with another user.
*/
- found = alloc_swap_scan_list(si, &si->frag_clusters[o], folio, true);
+ found = alloc_swap_scan_list(si, &si->frag_clusters[o], folio, order, true);
if (found)
goto done;
- found = alloc_swap_scan_list(si, &si->nonfull_clusters[o], folio, true);
+ found = alloc_swap_scan_list(si, &si->nonfull_clusters[o], folio, order, true);
if (found)
goto done;
}
@@ -1417,7 +1405,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
}
found = alloc_swap_scan_list(si, &si->free_clusters,
- folio, false);
+ folio, order, false);
}
}
#endif
@@ -1630,7 +1618,7 @@ static bool swap_alloc_fast(struct folio *folio)
if (cluster_is_usable(ci, order)) {
if (cluster_is_empty(ci))
offset = cluster_offset(si, ci);
- alloc_swap_scan_cluster(si, ci, folio, offset);
+ alloc_swap_scan_cluster(si, ci, folio, offset, order);
} else {
swap_cluster_unlock(ci);
}
@@ -1652,7 +1640,7 @@ static void swap_alloc_slow(struct folio *folio)
plist_requeue(&si->avail_list, &swap_avail_head);
spin_unlock(&swap_avail_lock);
if (get_swap_device_info(si)) {
- cluster_alloc_swap_entry(si, folio);
+ cluster_alloc_swap_entry(si, folio, folio_order(folio));
put_swap_device(si);
if (folio_test_swapcache(folio))
return;
@@ -2443,13 +2431,13 @@ swp_entry_t swap_alloc_hibernation_slot(int type)
if (pcp_si == si && pcp_offset) {
ci = swap_cluster_lock(si, pcp_offset);
if (cluster_is_usable(ci, 0))
- offset = alloc_swap_scan_cluster(si, ci, NULL, pcp_offset);
+ offset = alloc_swap_scan_cluster(si, ci, NULL, pcp_offset, 0);
else
swap_cluster_unlock(ci);
}
rcu_read_unlock();
if (!offset)
- offset = cluster_alloc_swap_entry(si, NULL);
+ offset = cluster_alloc_swap_entry(si, NULL, 0);
local_unlock(&percpu_swap_cluster.lock);
if (offset)
entry = swp_entry(si->type, offset);
@@ -4241,32 +4229,33 @@ static void xswap_free_xs_table(struct swap_info_struct *si, unsigned long idx)
ci->xs_table = NULL;
}
-/* Order-0 slot from @si, no folio attached; goes through the per-CPU cache. */
-static unsigned long xswap_alloc_slot_local(struct swap_info_struct *si)
+/* A range of 1 << @order slots, no folio attached. */
+static unsigned long xswap_alloc_slot_local(struct swap_info_struct *si,
+ unsigned int order)
{
struct swap_cluster_info *ci;
unsigned long pcp_offset, offset = SWAP_ENTRY_INVALID;
local_lock(&percpu_swap_cluster.lock);
- if (this_cpu_read(percpu_swap_cluster.si[0]) == si) {
- pcp_offset = this_cpu_read(percpu_swap_cluster.offset[0]);
+ if (this_cpu_read(percpu_swap_cluster.si[order]) == si) {
+ pcp_offset = this_cpu_read(percpu_swap_cluster.offset[order]);
if (pcp_offset) {
ci = swap_cluster_lock(si, pcp_offset);
- if (cluster_is_usable(ci, 0))
+ if (cluster_is_usable(ci, order))
offset = alloc_swap_scan_cluster(si, ci, NULL,
- pcp_offset);
+ pcp_offset, order);
else
swap_cluster_unlock(ci);
}
}
if (offset == SWAP_ENTRY_INVALID)
- offset = cluster_alloc_swap_entry(si, NULL);
+ offset = cluster_alloc_swap_entry(si, NULL, order);
local_unlock(&percpu_swap_cluster.lock);
return offset;
}
-static swp_entry_t xswap_alloc_phys_slot(void)
+static swp_entry_t xswap_alloc_phys_slot(unsigned int order)
{
struct swap_info_struct *devs[MAX_SWAPFILES], *si;
int nr = 0, i;
@@ -4291,7 +4280,7 @@ static swp_entry_t xswap_alloc_phys_slot(void)
if (!get_swap_device_info(devs[i]))
continue;
- offset = xswap_alloc_slot_local(devs[i]);
+ offset = xswap_alloc_slot_local(devs[i], order);
put_swap_device(devs[i]);
if (offset && offset != SWAP_ENTRY_INVALID)
return swp_entry(devs[i]->type, offset);
@@ -4317,29 +4306,46 @@ static void xswap_install_rmap(swp_entry_t phys, swp_entry_t entry)
}
}
-swp_entry_t xswap_backend_alloc(swp_entry_t entry)
+/* Back the range of 1 << @order slots at @entry with contiguous physical slots. */
+swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order)
{
struct swap_info_struct *si = __swap_entry_to_info(entry);
struct swap_cluster_info *ci;
+ unsigned long nr = 1UL << order;
unsigned long off = swp_offset(entry);
swp_entry_t phys;
+ unsigned int i;
bool recorded = false;
- phys = xswap_alloc_phys_slot();
+ phys = xswap_alloc_phys_slot(order);
if (!phys.val)
return phys;
- xswap_install_rmap(phys, entry);
+ /* One IO reads a large folio, so its range has to be contiguous. */
+ VM_WARN_ON_ONCE(swp_offset(phys) % nr);
+ VM_WARN_ON_ONCE(off % SWAPFILE_CLUSTER + nr > SWAPFILE_CLUSTER);
+
+ for (i = 0; i < nr; i++)
+ xswap_install_rmap(swp_entry(swp_type(phys), swp_offset(phys) + i),
+ swp_entry(si->type, off + i));
ci = swap_cluster_lock(si, off);
if (ci) {
- recorded = xswap_slot_set_backend(ci, off % SWAPFILE_CLUSTER, phys);
+ recorded = true;
+ for (i = 0; i < nr; i++) {
+ if (!xswap_slot_set_backend(ci, (off + i) % SWAPFILE_CLUSTER,
+ swp_entry(swp_type(phys),
+ swp_offset(phys) + i))) {
+ recorded = false;
+ break;
+ }
+ }
swap_cluster_unlock(ci);
}
if (!recorded) {
/* The backend is unrecorded, so the data would be unreachable. */
- xswap_backend_free(entry, phys);
+ xswap_backend_free(entry, phys, order);
return (swp_entry_t){};
}
@@ -4347,27 +4353,33 @@ swp_entry_t xswap_backend_alloc(swp_entry_t entry)
}
/* Undo xswap_backend_alloc(): the zswap copy is still in place. */
-void xswap_backend_free(swp_entry_t entry, swp_entry_t phys)
+void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order)
{
struct swap_info_struct *si = __swap_entry_to_info(entry);
- struct swap_info_struct *psi;
- struct swap_cluster_info *ci;
+ struct swap_info_struct *psi = swap_type_to_info(swp_type(phys));
+ unsigned long nr = 1UL << order;
unsigned long off = swp_offset(entry);
+ unsigned long poff = swp_offset(phys);
+ unsigned int i;
- ci = swap_cluster_lock(si, off);
- if (ci) {
- unsigned int slot = off % SWAPFILE_CLUSTER;
- unsigned long *xs = ci->xs_table;
+ for (i = 0; i < nr; i++) {
+ struct swap_cluster_info *ci;
- if (xs && xs[slot] == phys.val)
- WRITE_ONCE(xs[slot], 0);
- swap_cluster_unlock(ci);
- }
+ ci = swap_cluster_lock(si, off + i);
+ if (ci) {
+ unsigned int slot = (off + i) % SWAPFILE_CLUSTER;
+ unsigned long *xs = ci->xs_table;
- /* xswap_free_phys_slot() clears the reverse mapping. */
- psi = swap_type_to_info(swp_type(phys));
- if (psi)
- xswap_free_phys_slot(psi, entry, phys);
+ if (xs && xs[slot] == phys.val + i)
+ WRITE_ONCE(xs[slot], 0);
+ swap_cluster_unlock(ci);
+ }
+
+ /* xswap_free_phys_slot() clears the reverse mapping. */
+ if (psi)
+ xswap_free_phys_slot(psi, swp_entry(si->type, off + i),
+ swp_entry(swp_type(phys), poff + i));
+ }
}
static int xswap_map_clusters(struct swap_info_struct *si,
diff --git a/mm/zswap.c b/mm/zswap.c
index 24904b4c4dce..c79c527e6345 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1054,7 +1054,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
/* Reserve the destination before dropping the zswap copy. */
if (is_xswap) {
- phys = xswap_backend_alloc(swpentry);
+ phys = xswap_backend_alloc(swpentry, 0);
if (!phys.val) {
ret = -ENOMEM;
goto out;
@@ -1087,7 +1087,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
out:
if (ret) {
if (is_xswap && phys.val)
- xswap_backend_free(swpentry, phys);
+ xswap_backend_free(swpentry, phys, 0);
swap_cache_del_folio(folio);
folio_unlock(folio);
}
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 10/17] mm, swap: enable THP swapin for xswap entries
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (8 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 09/17] mm, swap: back a large xswap folio with a contiguous physical run Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 11/17] mm, swap: drop swap_folio_sector() Baoquan He
` (8 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
Swap a large anon folio back in as one unit, instead of falling back to
order-0 faults. This needs all of the folio's xswap entries to be
backed contiguously by one physical swap device.
Otherwise the folio is refused and read page by page. Its data may
still be in zswap, or the entries may span more than one backend
device. zswap cannot load a large folio in either case. The fault
then retries at a smaller order.
An xswap device is marked SWP_SYNCHRONOUS_IO. Its reads come from
memory, so readahead is skipped. Skipping readahead also avoids
pinning swap cache slots that xswap wants to shrink.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/memory.c | 7 +++++--
mm/swap.h | 8 ++++++++
mm/swap_state.c | 6 ++++++
mm/swapfile.c | 34 +++++++++++++++++++++++++++++++++-
4 files changed, 52 insertions(+), 3 deletions(-)
diff --git a/mm/memory.c b/mm/memory.c
index 926276d41920..c08ef143b8f1 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -4772,15 +4772,18 @@ static unsigned long thp_swapin_suitable_orders(struct vm_fault *vmf)
if (unlikely(userfaultfd_armed(vma)))
return 0;
+ entry = softleaf_from_pte(vmf->orig_pte);
+
/*
* A large swapped out folio could be partially or fully in zswap. We
* lack handling for such cases, so fallback to swapping in order-0
* folio.
+ * An xswap entry is deferred to __swap_cache_add_check() instead.
*/
- if (!zswap_never_enabled())
+ if (!(__swap_entry_to_info(entry)->flags & SWP_XSWAP) &&
+ !zswap_never_enabled())
return 0;
- entry = softleaf_from_pte(vmf->orig_pte);
/*
* Get a list of all the (large) orders below PMD_ORDER that are enabled
* and suitable for swapping THP.
diff --git a/mm/swap.h b/mm/swap.h
index 4fcbb21a583d..990accefd3b6 100644
--- a/mm/swap.h
+++ b/mm/swap.h
@@ -477,6 +477,8 @@ int shmem_writeout(struct swap_io_ctx *ctx, struct folio *folio,
swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci, unsigned int slot);
swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order);
void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order);
+int xswap_check_backing(struct swap_cluster_info *ci, unsigned int off,
+ unsigned int nr);
#else
static inline swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci,
unsigned int slot)
@@ -492,6 +494,12 @@ static inline swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int or
static inline void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order)
{
}
+
+static inline int xswap_check_backing(struct swap_cluster_info *ci,
+ unsigned int off, unsigned int nr)
+{
+ return nr;
+}
#endif /* CONFIG_XSWAP */
#endif /* _MM_SWAP_H */
diff --git a/mm/swap_state.c b/mm/swap_state.c
index 135d9573a9dc..112ab952fa0c 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -194,6 +194,12 @@ static int __swap_cache_add_check(struct swap_cluster_info *ci,
if (nr == 1)
return 0;
+ /* A zswap-backed or mixed batch falls back to a smaller order. */
+ if (__swap_entry_to_info(targ_entry)->flags & SWP_XSWAP) {
+ if (xswap_check_backing(ci, round_down(ci_off, nr), nr) != nr)
+ return -EBUSY;
+ }
+
is_zero = __swap_table_test_zero(ci, ci_off);
ci_off = round_down(ci_off, nr);
ci_end = ci_off + nr;
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 4d72c854c7b3..852f10642b99 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -4164,6 +4164,34 @@ swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci, unsigned int slot)
return (swp_entry_t){ .val = xs[slot] };
}
+/*
+ * A large xswap folio is read back as one IO only when the whole range is
+ * backed by one device, contiguously. Returns how many leading slots are.
+ */
+int xswap_check_backing(struct swap_cluster_info *ci, unsigned int off,
+ unsigned int nr)
+{
+ swp_entry_t first;
+ unsigned int i;
+
+ lockdep_assert_held(&ci->lock);
+
+ /* In zswap, which cannot load a large folio. */
+ if (!ci->xs_table || !ci->xs_table[off])
+ return 1;
+
+ first.val = ci->xs_table[off];
+ for (i = 1; i < nr; i++) {
+ swp_entry_t cur = { .val = ci->xs_table[off + i] };
+
+ if (swp_type(cur) != swp_type(first) ||
+ swp_offset(cur) != swp_offset(first) + i)
+ break;
+ }
+
+ return i;
+}
+
/* Cluster lock held. xswap never unmaps a cluster below the mapped prefix. */
static unsigned long *xswap_xs_table_locked(struct swap_cluster_info *ci,
gfp_t gfp)
@@ -5066,7 +5094,11 @@ static int xswap_create(int prio)
nr_clusters = DIV_ROUND_UP(ram, SWAPFILE_CLUSTER);
si->bdev = NULL;
- si->flags |= SWP_XSWAP | SWP_SOLIDSTATE;
+ /*
+ * Reads are served from zswap, so no readahead. Only a synchronous
+ * device can swap a large folio back in as a unit.
+ */
+ si->flags |= SWP_XSWAP | SWP_SOLIDSTATE | SWP_SYNCHRONOUS_IO;
si->max = maxpages;
si->pages = min_t(unsigned long, nr_clusters * SWAPFILE_CLUSTER,
si->max) - 1;
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 11/17] mm, swap: drop swap_folio_sector()
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (9 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 10/17] mm, swap: enable THP swapin for xswap entries Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 12/17] mm, swap: split the swap memcg charge helpers Baoquan He
` (7 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
No one uses it now.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
include/linux/swap.h | 1 -
mm/swapfile.c | 13 -------------
2 files changed, 14 deletions(-)
diff --git a/include/linux/swap.h b/include/linux/swap.h
index b9a9a02007cf..188821b24e51 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -418,7 +418,6 @@ extern int __swap_count(swp_entry_t entry);
extern bool swap_entry_swapped(struct swap_info_struct *si, swp_entry_t entry);
extern int swp_swapcount(swp_entry_t entry);
extern struct swap_info_struct *get_swap_device(swp_entry_t entry);
-sector_t swap_folio_sector(struct folio *folio);
sector_t swap_entry_sector(swp_entry_t entry);
/*
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 852f10642b99..70ec544acd6b 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -449,19 +449,6 @@ offset_to_swap_extent(struct swap_info_struct *sis, unsigned long offset)
BUG();
}
-sector_t swap_folio_sector(struct folio *folio)
-{
- struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
- struct swap_extent *se;
- sector_t sector;
- pgoff_t offset;
-
- offset = swp_offset(folio->swap);
- se = offset_to_swap_extent(sis, offset);
- sector = se->start_block + (offset - se->start_page);
- return sector << (PAGE_SHIFT - 9);
-}
-
sector_t swap_entry_sector(swp_entry_t entry)
{
struct swap_info_struct *sis = __swap_entry_to_info(entry);
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 12/17] mm, swap: split the swap memcg charge helpers
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (10 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 11/17] mm, swap: drop swap_folio_sector() Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 13/17] mm, swap: do not charge zswap-backed xswap entries Baoquan He
` (6 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
From: Nhat Pham <nphamcs@gmail.com>
Split __mem_cgroup_try_charge_swap() into separate get, charge, record,
uncharge and put helpers, and factor mem_cgroup_may_zswap() out of
obj_cgroup_may_zswap(). Recording the owner of a swap slot and charging
it no longer have to happen together, so a later patch can charge swap
only once it gets physical backing.
No functional change.
Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Nhat Pham <nphamcs@gmail.com>
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
.../admin-guide/cgroup-v1/memcg_test.rst | 2 +-
include/linux/memcontrol.h | 6 +
include/linux/swap.h | 61 ++++++--
mm/memcontrol-v1.c | 10 +-
mm/memcontrol.c | 133 +++++++++++-------
mm/swapfile.c | 34 ++++-
6 files changed, 181 insertions(+), 65 deletions(-)
diff --git a/Documentation/admin-guide/cgroup-v1/memcg_test.rst b/Documentation/admin-guide/cgroup-v1/memcg_test.rst
index d9951c319ef5..cd565626c435 100644
--- a/Documentation/admin-guide/cgroup-v1/memcg_test.rst
+++ b/Documentation/admin-guide/cgroup-v1/memcg_test.rst
@@ -43,7 +43,7 @@ Please note that implementation details can be changed.
mem_cgroup_uncharge()
Called when a page's refcount goes down to 0.
- mem_cgroup_uncharge_swap()
+ mem_cgroup_swap_uncharge()
Called when swp_entry's refcnt goes down to 0. A charge against swap
disappears.
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index 46bf724cae7a..4add06affefa 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -1933,6 +1933,7 @@ static inline void mem_cgroup_calculate_protection_path(struct mem_cgroup *root,
#if defined(CONFIG_MEMCG) && defined(CONFIG_ZSWAP)
bool obj_cgroup_may_zswap(struct obj_cgroup *objcg);
+bool mem_cgroup_may_zswap(struct mem_cgroup *memcg, bool may_flush);
void obj_cgroup_charge_zswap(struct obj_cgroup *objcg, size_t size);
void obj_cgroup_uncharge_zswap(struct obj_cgroup *objcg, size_t size);
bool mem_cgroup_zswap_writeback_enabled(struct mem_cgroup *memcg);
@@ -1941,6 +1942,11 @@ static inline bool obj_cgroup_may_zswap(struct obj_cgroup *objcg)
{
return true;
}
+
+static inline bool mem_cgroup_may_zswap(struct mem_cgroup *memcg, bool may_flush)
+{
+ return true;
+}
static inline void obj_cgroup_charge_zswap(struct obj_cgroup *objcg,
size_t size)
{
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 188821b24e51..7ceac868a885 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -529,36 +529,81 @@ static inline void folio_throttle_swaprate(struct folio *folio, gfp_t gfp)
#endif
#if defined(CONFIG_MEMCG) && defined(CONFIG_SWAP)
-int __mem_cgroup_try_charge_swap(struct folio *folio);
-static inline int mem_cgroup_try_charge_swap(struct folio *folio)
+struct mem_cgroup *__mem_cgroup_swap_get(struct folio *folio);
+static inline struct mem_cgroup *mem_cgroup_swap_get(struct folio *folio)
+{
+ if (mem_cgroup_disabled())
+ return NULL;
+ return __mem_cgroup_swap_get(folio);
+}
+
+int __mem_cgroup_swap_charge(struct mem_cgroup *memcg, unsigned int nr_pages);
+static inline int mem_cgroup_swap_charge(struct mem_cgroup *memcg,
+ unsigned int nr_pages)
{
if (mem_cgroup_disabled())
return 0;
- return __mem_cgroup_try_charge_swap(folio);
+ return __mem_cgroup_swap_charge(memcg, nr_pages);
}
-extern void __mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_pages);
-static inline void mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_pages)
+void __mem_cgroup_swap_record(struct folio *folio, struct mem_cgroup *memcg);
+static inline void mem_cgroup_swap_record(struct folio *folio,
+ struct mem_cgroup *memcg)
{
if (mem_cgroup_disabled())
return;
- __mem_cgroup_uncharge_swap(id, nr_pages);
+ __mem_cgroup_swap_record(folio, memcg);
+}
+
+void __mem_cgroup_swap_uncharge(struct mem_cgroup *memcg,
+ unsigned int nr_pages);
+static inline void mem_cgroup_swap_uncharge(struct mem_cgroup *memcg,
+ unsigned int nr_pages)
+{
+ if (mem_cgroup_disabled())
+ return;
+ __mem_cgroup_swap_uncharge(memcg, nr_pages);
+}
+
+void __mem_cgroup_swap_put(struct mem_cgroup *memcg, unsigned int nr_pages);
+static inline void mem_cgroup_swap_put(struct mem_cgroup *memcg,
+ unsigned int nr_pages)
+{
+ if (mem_cgroup_disabled())
+ return;
+ __mem_cgroup_swap_put(memcg, nr_pages);
}
long mem_cgroup_get_folio_swap_margin(struct folio *folio);
extern long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg);
extern bool mem_cgroup_swap_full(struct folio *folio);
#else
-static inline int mem_cgroup_try_charge_swap(struct folio *folio)
+static inline struct mem_cgroup *mem_cgroup_swap_get(struct folio *folio)
+{
+ return NULL;
+}
+
+static inline int mem_cgroup_swap_charge(struct mem_cgroup *memcg,
+ unsigned int nr_pages)
{
return 0;
}
-static inline void mem_cgroup_uncharge_swap(unsigned short id,
+static inline void mem_cgroup_swap_record(struct folio *folio,
+ struct mem_cgroup *memcg)
+{
+}
+
+static inline void mem_cgroup_swap_uncharge(struct mem_cgroup *memcg,
unsigned int nr_pages)
{
}
+static inline void mem_cgroup_swap_put(struct mem_cgroup *memcg,
+ unsigned int nr_pages)
+{
+}
+
static inline long mem_cgroup_get_folio_swap_margin(struct folio *folio)
{
return PAGE_COUNTER_MAX;
diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c
index bf2c7d53b01b..3e06a8bdf46e 100644
--- a/mm/memcontrol-v1.c
+++ b/mm/memcontrol-v1.c
@@ -341,6 +341,7 @@ void __memcg1_swapout(struct folio *folio, struct swap_cluster_info *ci)
void memcg1_swapin(struct folio *folio)
{
struct swap_cluster_info *ci;
+ struct mem_cgroup *memcg;
unsigned long nr_pages;
unsigned short id;
@@ -372,7 +373,14 @@ void memcg1_swapin(struct folio *folio)
id = __swap_cgroup_clear(ci, swp_cluster_offset(folio->swap),
nr_pages);
swap_cluster_unlock(ci);
- mem_cgroup_uncharge_swap(id, nr_pages);
+
+ rcu_read_lock();
+ memcg = mem_cgroup_from_private_id(id);
+ if (memcg) {
+ mem_cgroup_swap_uncharge(memcg, nr_pages);
+ mem_cgroup_swap_put(memcg, nr_pages);
+ }
+ rcu_read_unlock();
}
#endif
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 1460cba53588..1d35b7ae8d70 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -5926,80 +5926,111 @@ int __init mem_cgroup_init(void)
#ifdef CONFIG_SWAP
/**
- * __mem_cgroup_try_charge_swap - try charging swap space for a folio
+ * __mem_cgroup_swap_get - pin the memcg to account a folio's swap slots to
* @folio: folio being added to swap
*
- * Try to charge @folio's memcg for the swap space at folio->swap.
+ * Pins one private ID ref per page of @folio on its memcg, or on its closest
+ * online ancestor if it has been offlined. The caller charges and records
+ * against whichever memcg is returned, so both land on the same one.
*
- * Returns 0 on success, -ENOMEM on failure.
+ * Return: the pinned memcg, or NULL if there is nothing to account. Drop the
+ * pins with __mem_cgroup_swap_put().
*/
-int __mem_cgroup_try_charge_swap(struct folio *folio)
+struct mem_cgroup *__mem_cgroup_swap_get(struct folio *folio)
{
unsigned int nr_pages = folio_nr_pages(folio);
- struct swap_cluster_info *ci;
- struct page_counter *counter;
struct mem_cgroup *memcg;
struct obj_cgroup *objcg;
if (do_memsw_account())
- return 0;
+ return NULL;
objcg = folio_objcg(folio);
VM_WARN_ON_ONCE_FOLIO(!objcg, folio);
if (!objcg)
- return 0;
+ return NULL;
rcu_read_lock();
memcg = obj_cgroup_memcg(objcg);
if (!folio_test_swapcache(folio)) {
memcg_memory_event(memcg, MEMCG_SWAP_FAIL);
rcu_read_unlock();
- return 0;
+ return NULL;
}
memcg = mem_cgroup_private_id_get_online(memcg, nr_pages);
/* memcg is pined by memcg ID. */
rcu_read_unlock();
+ return memcg;
+}
+
+/**
+ * __mem_cgroup_swap_charge - charge physical swap space
+ * @memcg: the mem_cgroup to charge (may be NULL)
+ * @nr_pages: the amount of swap space to charge
+ *
+ * Return: 0 on success, -ENOMEM if memory.swap.max is exceeded.
+ */
+int __mem_cgroup_swap_charge(struct mem_cgroup *memcg, unsigned int nr_pages)
+{
+ struct page_counter *counter;
+
+ if (do_memsw_account() || !memcg)
+ return 0;
+
if (!mem_cgroup_is_root(memcg) &&
!page_counter_try_charge(&memcg->swap, nr_pages, &counter)) {
memcg_memory_event(memcg, MEMCG_SWAP_MAX);
memcg_memory_event(memcg, MEMCG_SWAP_FAIL);
- mem_cgroup_private_id_put(memcg, nr_pages);
return -ENOMEM;
}
mod_memcg_state(memcg, MEMCG_SWAP, nr_pages);
+ return 0;
+}
+
+/**
+ * __mem_cgroup_swap_record - record the owner of a folio's swap slots
+ * @folio: folio being added to swap
+ * @memcg: the memcg pinned by __mem_cgroup_swap_get()
+ */
+void __mem_cgroup_swap_record(struct folio *folio, struct mem_cgroup *memcg)
+{
+ struct swap_cluster_info *ci;
ci = swap_cluster_get_and_lock(folio);
- __swap_cgroup_set(ci, swp_cluster_offset(folio->swap), nr_pages,
- mem_cgroup_private_id(memcg));
+ __swap_cgroup_set(ci, swp_cluster_offset(folio->swap),
+ folio_nr_pages(folio), mem_cgroup_private_id(memcg));
swap_cluster_unlock(ci);
-
- return 0;
}
/**
- * __mem_cgroup_uncharge_swap - uncharge swap space
- * @id: cgroup id to uncharge
+ * __mem_cgroup_swap_uncharge - uncharge physical swap space
+ * @memcg: the mem_cgroup to uncharge (may be NULL)
* @nr_pages: the amount of swap space to uncharge
*/
-void __mem_cgroup_uncharge_swap(unsigned short id, unsigned int nr_pages)
+void __mem_cgroup_swap_uncharge(struct mem_cgroup *memcg, unsigned int nr_pages)
{
- struct mem_cgroup *memcg;
+ if (!memcg)
+ return;
- rcu_read_lock();
- memcg = mem_cgroup_from_private_id(id);
- if (memcg) {
- if (!mem_cgroup_is_root(memcg)) {
- if (do_memsw_account())
- page_counter_uncharge(&memcg->memsw, nr_pages);
- else
- page_counter_uncharge(&memcg->swap, nr_pages);
- }
- mod_memcg_state(memcg, MEMCG_SWAP, -nr_pages);
- mem_cgroup_private_id_put(memcg, nr_pages);
+ if (!mem_cgroup_is_root(memcg)) {
+ if (do_memsw_account())
+ page_counter_uncharge(&memcg->memsw, nr_pages);
+ else
+ page_counter_uncharge(&memcg->swap, nr_pages);
}
- rcu_read_unlock();
+ mod_memcg_state(memcg, MEMCG_SWAP, -nr_pages);
+}
+
+/**
+ * __mem_cgroup_swap_put - drop the private ID refs taken for swap slots
+ * @memcg: the pinned mem_cgroup
+ * @nr_pages: number of refs to drop
+ */
+void __mem_cgroup_swap_put(struct mem_cgroup *memcg, unsigned int nr_pages)
+{
+ mem_cgroup_private_id_put(memcg, nr_pages);
}
long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
@@ -6198,8 +6229,10 @@ static struct cftype swap_files[] = {
#ifdef CONFIG_ZSWAP
/**
- * obj_cgroup_may_zswap - check if this cgroup can zswap
- * @objcg: the object cgroup
+ * mem_cgroup_may_zswap - check if this cgroup can zswap
+ * @memcg: the memcg to query
+ * @may_flush: force-flush stats for an accurate check (sleeps). Pass false
+ * from atomic contexts; the check is then best-effort.
*
* Check if the hierarchical zswap limit has been reached.
*
@@ -6209,36 +6242,38 @@ static struct cftype swap_files[] = {
* spending cycles on compression when there is already no room left
* or zswap is disabled altogether somewhere in the hierarchy.
*/
-bool obj_cgroup_may_zswap(struct obj_cgroup *objcg)
+bool mem_cgroup_may_zswap(struct mem_cgroup *memcg, bool may_flush)
{
- struct mem_cgroup *memcg, *original_memcg;
- bool ret = true;
-
if (!cgroup_subsys_on_dfl(memory_cgrp_subsys))
return true;
- original_memcg = get_mem_cgroup_from_objcg(objcg);
- for (memcg = original_memcg; !mem_cgroup_is_root(memcg);
- memcg = parent_mem_cgroup(memcg)) {
+ for (; !mem_cgroup_is_root(memcg); memcg = parent_mem_cgroup(memcg)) {
unsigned long max = READ_ONCE(memcg->zswap_max);
unsigned long pages;
if (max == PAGE_COUNTER_MAX)
continue;
- if (max == 0) {
- ret = false;
- break;
- }
+ if (max == 0)
+ return false;
/* Force flush to get accurate stats for charging */
- __mem_cgroup_flush_stats(memcg, true);
+ if (may_flush)
+ __mem_cgroup_flush_stats(memcg, true);
pages = memcg_page_state(memcg, MEMCG_ZSWAP_B) / PAGE_SIZE;
- if (pages < max)
- continue;
- ret = false;
- break;
+ if (pages >= max)
+ return false;
}
- mem_cgroup_put(original_memcg);
+ return true;
+}
+
+bool obj_cgroup_may_zswap(struct obj_cgroup *objcg)
+{
+ struct mem_cgroup *memcg;
+ bool ret;
+
+ memcg = get_mem_cgroup_from_objcg(objcg);
+ ret = mem_cgroup_may_zswap(memcg, true);
+ mem_cgroup_put(memcg);
return ret;
}
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 70ec544acd6b..7dcb4b48645d 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -1962,6 +1962,7 @@ int folio_alloc_swap(struct folio *folio)
{
unsigned int order = folio_order(folio);
unsigned int size = 1 << order;
+ struct mem_cgroup *memcg;
VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio);
VM_BUG_ON_FOLIO(!folio_test_uptodate(folio), folio);
@@ -1995,10 +1996,18 @@ int folio_alloc_swap(struct folio *folio)
goto again;
}
- /* Need to call this even if allocation failed, for MEMCG_SWAP_FAIL. */
- if (unlikely(mem_cgroup_try_charge_swap(folio))) {
- swap_cache_del_folio(folio);
- goto failed;
+ /*
+ * Need to call this even if allocation failed, for MEMCG_SWAP_FAIL.
+ * The memcg is pinned here, then charged and recorded.
+ */
+ memcg = mem_cgroup_swap_get(folio);
+ if (memcg) {
+ if (unlikely(mem_cgroup_swap_charge(memcg, size))) {
+ mem_cgroup_swap_put(memcg, size);
+ swap_cache_del_folio(folio);
+ goto failed;
+ }
+ mem_cgroup_swap_record(folio, memcg);
}
if (unlikely(!folio_test_swapcache(folio)))
@@ -2144,6 +2153,19 @@ struct swap_info_struct *get_swap_device(swp_entry_t entry)
return ERR_PTR(-EIO);
}
+static void memcg_swap_free(unsigned short id, unsigned int nr)
+{
+ struct mem_cgroup *memcg;
+
+ rcu_read_lock();
+ memcg = mem_cgroup_from_private_id(id);
+ if (memcg) {
+ mem_cgroup_swap_uncharge(memcg, nr);
+ mem_cgroup_swap_put(memcg, nr);
+ }
+ rcu_read_unlock();
+}
+
/*
* Free a set of swap slots after their swap count dropped to zero, or will be
* zero after putting the last ref (saves one __swap_cluster_put_entry call).
@@ -2186,14 +2208,14 @@ void __swap_cluster_free_entries(struct swap_info_struct *si,
id_cur = __swap_cgroup_clear(ci, ci_off, 1);
if (batch_id != id_cur) {
if (batch_id)
- mem_cgroup_uncharge_swap(batch_id, ci_off - batch_off);
+ memcg_swap_free(batch_id, ci_off - batch_off);
batch_id = id_cur;
batch_off = ci_off;
}
} while (++ci_off < ci_end);
if (batch_id)
- mem_cgroup_uncharge_swap(batch_id, ci_off - batch_off);
+ memcg_swap_free(batch_id, ci_off - batch_off);
swap_range_free(si, ci_head + ci_start, nr_pages);
swap_cluster_assert_empty(ci, ci_start, nr_pages, false);
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 13/17] mm, swap: do not charge zswap-backed xswap entries
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (11 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 12/17] mm, swap: split the swap memcg charge helpers Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-21 11:54 ` Chris Li
2026-09-20 7:20 ` [RFC PATCH 14/17] mm, swap: charge an xswap entry when it gets physical backing Baoquan He
` (5 subsequent siblings)
18 siblings, 1 reply; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
From: Nhat Pham <nphamcs@gmail.com>
An xswap entry occupies no physical swap space: its data lives in zswap
until a backend is taken. It was nevertheless charged against
memcg->swap at allocation, as if it did.
Record the memcg at allocation but do not charge. The charge is taken
when the entry acquires physical backing, and released with it.
memory.swap.current therefore counts only on-disk swap usage, not
zswap-backed xswap entries. A cgroup can reclaim its anon memory with
memory.swap.max set to 0, provided zswap is allowed for it.
Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Nhat Pham <nphamcs@gmail.com>
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 21 ++++++++++++++++-----
1 file changed, 16 insertions(+), 5 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 7dcb4b48645d..627f2cab8077 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -2002,7 +2002,12 @@ int folio_alloc_swap(struct folio *folio)
*/
memcg = mem_cgroup_swap_get(folio);
if (memcg) {
- if (unlikely(mem_cgroup_swap_charge(memcg, size))) {
+ /*
+ * An xswap entry has no physical swap yet, so only record the
+ * memcg here. Its backing is charged when one is taken.
+ */
+ if (!(__swap_entry_to_info(folio->swap)->flags & SWP_XSWAP) &&
+ unlikely(mem_cgroup_swap_charge(memcg, size))) {
mem_cgroup_swap_put(memcg, size);
swap_cache_del_folio(folio);
goto failed;
@@ -2153,14 +2158,19 @@ struct swap_info_struct *get_swap_device(swp_entry_t entry)
return ERR_PTR(-EIO);
}
-static void memcg_swap_free(unsigned short id, unsigned int nr)
+static void memcg_swap_free(unsigned short id, unsigned int nr, bool is_xswap)
{
struct mem_cgroup *memcg;
rcu_read_lock();
memcg = mem_cgroup_from_private_id(id);
if (memcg) {
- mem_cgroup_swap_uncharge(memcg, nr);
+ /*
+ * An xswap entry was not charged here; its backing, if any,
+ * was uncharged when it was released.
+ */
+ if (!is_xswap)
+ mem_cgroup_swap_uncharge(memcg, nr);
mem_cgroup_swap_put(memcg, nr);
}
rcu_read_unlock();
@@ -2179,6 +2189,7 @@ void __swap_cluster_free_entries(struct swap_info_struct *si,
unsigned int ci_off = ci_start, ci_end = ci_start + nr_pages;
unsigned long ci_head = cluster_offset(si, ci);
unsigned int batch_off = ci_off;
+ bool is_xswap = si->flags & SWP_XSWAP;
VM_WARN_ON(ci->count < nr_pages);
@@ -2208,14 +2219,14 @@ void __swap_cluster_free_entries(struct swap_info_struct *si,
id_cur = __swap_cgroup_clear(ci, ci_off, 1);
if (batch_id != id_cur) {
if (batch_id)
- memcg_swap_free(batch_id, ci_off - batch_off);
+ memcg_swap_free(batch_id, ci_off - batch_off, is_xswap);
batch_id = id_cur;
batch_off = ci_off;
}
} while (++ci_off < ci_end);
if (batch_id)
- memcg_swap_free(batch_id, ci_off - batch_off);
+ memcg_swap_free(batch_id, ci_off - batch_off, is_xswap);
swap_range_free(si, ci_head + ci_start, nr_pages);
swap_cluster_assert_empty(ci, ci_start, nr_pages, false);
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 14/17] mm, swap: charge an xswap entry when it gets physical backing
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (12 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 13/17] mm, swap: do not charge zswap-backed xswap entries Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 15/17] mm, swap: don't gate xswap on the physical swap free count Baoquan He
` (4 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
From: Nhat Pham <nphamcs@gmail.com>
An xswap entry is recorded but not charged at allocation. Charge
memcg->swap when the entry takes a physical backend slot, in
xswap_backend_alloc(), and uncharge it when the backend is released.
memory.swap.current therefore still counts the pages that do reach the
disk, and memory.swap.max is still enforced on them.
Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Nhat Pham <nphamcs@gmail.com>
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 60 +++++++++++++++++++++++++++++++++++++++++++++++++--
1 file changed, 58 insertions(+), 2 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 627f2cab8077..f848c6edc6e7 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -79,6 +79,11 @@ static void xswap_free_phys_slot(struct swap_info_struct *si,
static void xswap_release_slot_backend(struct swap_info_struct *si,
struct swap_cluster_info *ci,
unsigned int slot);
+static struct mem_cgroup *xswap_entry_memcg(struct swap_cluster_info *ci,
+ unsigned long off);
+static void xswap_backend_uncharge(swp_entry_t entry, unsigned int nr);
+static void __xswap_backend_unlink(swp_entry_t entry, swp_entry_t phys,
+ unsigned int order);
static int xswap_create(int prio);
static int xswap_destroy(int type);
@@ -3169,6 +3174,8 @@ static void xswap_release_slot_backend(struct swap_info_struct *si,
psi = swap_type_to_info(swp_type(phys));
if (psi)
xswap_free_phys_slot(psi, entry, phys);
+
+ xswap_backend_uncharge(entry, 1);
}
/*
@@ -4359,6 +4366,7 @@ swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order)
{
struct swap_info_struct *si = __swap_entry_to_info(entry);
struct swap_cluster_info *ci;
+ struct mem_cgroup *memcg = NULL;
unsigned long nr = 1UL << order;
unsigned long off = swp_offset(entry);
swp_entry_t phys;
@@ -4369,6 +4377,21 @@ swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order)
if (!phys.val)
return phys;
+ /*
+ * The xswap entry was only recorded at allocation, so charge the
+ * physical swap here. On failure, drop the physical run without
+ * uncharging it.
+ */
+ ci = swap_cluster_lock(si, off);
+ if (ci) {
+ memcg = xswap_entry_memcg(ci, off);
+ swap_cluster_unlock(ci);
+ }
+ if (mem_cgroup_swap_charge(memcg, nr)) {
+ __xswap_backend_unlink(entry, phys, order);
+ return (swp_entry_t){};
+ }
+
/* One IO reads a large folio, so its range has to be contiguous. */
VM_WARN_ON_ONCE(swp_offset(phys) % nr);
VM_WARN_ON_ONCE(off % SWAPFILE_CLUSTER + nr > SWAPFILE_CLUSTER);
@@ -4400,8 +4423,34 @@ swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order)
return phys;
}
-/* Undo xswap_backend_alloc(): the zswap copy is still in place. */
-void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order)
+/* The memcg that owns an xswap entry, pinned by its recorded ID. */
+static struct mem_cgroup *xswap_entry_memcg(struct swap_cluster_info *ci,
+ unsigned long off)
+{
+ unsigned short id = __swap_cgroup_get(ci, off % SWAPFILE_CLUSTER);
+
+ return id ? mem_cgroup_from_private_id(id) : NULL;
+}
+
+static void xswap_backend_uncharge(swp_entry_t entry, unsigned int nr)
+{
+ struct swap_info_struct *si = __swap_entry_to_info(entry);
+ struct swap_cluster_info *ci;
+ struct mem_cgroup *memcg = NULL;
+ unsigned long off = swp_offset(entry);
+
+ ci = swap_cluster_lock(si, off);
+ if (ci) {
+ memcg = xswap_entry_memcg(ci, off);
+ swap_cluster_unlock(ci);
+ }
+ if (memcg)
+ mem_cgroup_swap_uncharge(memcg, nr);
+}
+
+/* Drop the physical run and the reverse mapping, without any charging. */
+static void __xswap_backend_unlink(swp_entry_t entry, swp_entry_t phys,
+ unsigned int order)
{
struct swap_info_struct *si = __swap_entry_to_info(entry);
struct swap_info_struct *psi = swap_type_to_info(swp_type(phys));
@@ -4430,6 +4479,13 @@ void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order)
}
}
+/* Undo xswap_backend_alloc(): the zswap copy is still in place. */
+void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order)
+{
+ __xswap_backend_unlink(entry, phys, order);
+ xswap_backend_uncharge(entry, 1UL << order);
+}
+
static int xswap_map_clusters(struct swap_info_struct *si,
unsigned long start_idx, unsigned long nr)
{
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 15/17] mm, swap: don't gate xswap on the physical swap free count
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (13 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 14/17] mm, swap: charge an xswap entry when it gets physical backing Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 16/17] mm, swap: do not retake the cluster lock when uncharging an xswap slot Baoquan He
` (3 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
From: Nhat Pham <nphamcs@gmail.com>
An xswap entry is charged only when it gets physical backing, so a
zswap-capable memcg can keep swapping out through xswap even when its
physical swap margin is zero. mem_cgroup_get_nr_swap_pages() would
otherwise starve anon reclaim with memory.swap.max set to 0.
Return PAGE_COUNTER_MAX when an xswap device is active, zswap is on,
and the memcg allows zswap. Track active xswap devices in nr_xswap_files,
mirroring nr_real_swapfiles.
Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Nhat Pham <nphamcs@gmail.com>
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
include/linux/swap.h | 6 ++++++
mm/memcontrol.c | 14 +++++++++++++-
mm/swapfile.c | 5 +++++
3 files changed, 24 insertions(+), 1 deletion(-)
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 7ceac868a885..540a48f203e4 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -391,6 +391,12 @@ void free_pages_and_swap_cache(struct encoded_page **, int);
/* linux/mm/swapfile.c */
extern atomic_long_t nr_swap_pages;
extern atomic_t nr_real_swapfiles;
+extern atomic_t nr_xswap_files;
+
+static inline bool xswap_enabled(void)
+{
+ return atomic_read(&nr_xswap_files) > 0;
+}
extern long total_swap_pages;
extern atomic_t nr_rotate_swap;
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 1d35b7ae8d70..7646e66bc5df 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -6035,8 +6035,20 @@ void __mem_cgroup_swap_put(struct mem_cgroup *memcg, unsigned int nr_pages)
long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
{
- long nr_swap_pages = get_nr_swap_pages();
+ long nr_swap_pages;
+ /*
+ * xswap charges physical backing, not allocation, so virtual swap is
+ * unbounded for a zswap-capable memcg and the swap.max walk below
+ * would starve anon reclaim. swap.max is still enforced when the
+ * backing is charged.
+ */
+ if (xswap_enabled() && zswap_is_enabled() &&
+ (mem_cgroup_disabled() || do_memsw_account() ||
+ mem_cgroup_may_zswap(memcg, false)))
+ return PAGE_COUNTER_MAX;
+
+ nr_swap_pages = get_nr_swap_pages();
if (!mem_cgroup_disabled() && !do_memsw_account())
nr_swap_pages = min(nr_swap_pages, page_counter_margin(&memcg->swap));
diff --git a/mm/swapfile.c b/mm/swapfile.c
index f848c6edc6e7..f59dc74f574e 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -189,6 +189,7 @@ static DEFINE_SPINLOCK(swap_lock);
static unsigned int nr_swapfiles;
atomic_long_t nr_swap_pages;
atomic_t nr_real_swapfiles;
+atomic_t nr_xswap_files;
/*
* Some modules use swappable objects and may try to swap them out under
* memory pressure (via the shrinker). Before doing so, they may wish to
@@ -1425,6 +1426,8 @@ static void del_from_avail_list(struct swap_info_struct *si, bool swapoff)
/* Count active devices, not merely those on the avail list. */
if (!(si->flags & SWP_XSWAP))
atomic_sub(1, &nr_real_swapfiles);
+ else
+ atomic_sub(1, &nr_xswap_files);
atomic_long_or(SWAP_USAGE_OFFLIST_BIT, &si->inuse_pages);
} else {
/*
@@ -1486,6 +1489,8 @@ static void add_to_avail_list(struct swap_info_struct *si, bool swapon)
plist_add(&si->avail_list, &swap_avail_head);
if (swapon && !(si->flags & SWP_XSWAP))
atomic_add(1, &nr_real_swapfiles);
+ else if (swapon)
+ atomic_add(1, &nr_xswap_files);
skip:
spin_unlock(&swap_avail_lock);
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 16/17] mm, swap: do not retake the cluster lock when uncharging an xswap slot
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (14 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 15/17] mm, swap: don't gate xswap on the physical swap free count Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-20 7:20 ` [RFC PATCH 17/17] mm, swap: drop a refused xswap backend run directly Baoquan He
` (2 subsequent siblings)
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
xswap_release_slot_backend() runs with ci->lock held. Its callers free
the slots of a cluster they have locked. It called
xswap_backend_uncharge() to give the swap charge back, and that takes
the same cluster lock again to read the owner out of the swap table.
A spinlock is not recursive, so it spins forever. The swapin path hangs
on it:
do_swap_page
folio_free_swap
swap_cache_del_folio
__swap_cluster_free_entries
xswap_release_slot_backend
xswap_backend_uncharge <- spins here
Read the owner from ci directly instead. The lock is already held, and
the cgroup id is still in the swap table at this point: the loop clears
it after this call.
xswap_backend_uncharge() stays for xswap_backend_free(), which runs
without the cluster lock.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 10 +++++++++-
1 file changed, 9 insertions(+), 1 deletion(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index f59dc74f574e..7a67c2ff4e89 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -3172,6 +3172,7 @@ static void xswap_release_slot_backend(struct swap_info_struct *si,
{
swp_entry_t entry = swp_entry(si->type, cluster_offset(si, ci) + slot);
struct swap_info_struct *psi;
+ struct mem_cgroup *memcg;
swp_entry_t phys = { .val = ci->xs_table[slot] };
WRITE_ONCE(ci->xs_table[slot], 0);
@@ -3180,7 +3181,14 @@ static void xswap_release_slot_backend(struct swap_info_struct *si,
if (psi)
xswap_free_phys_slot(psi, entry, phys);
- xswap_backend_uncharge(entry, 1);
+ /*
+ * ci->lock is held here, so read the owner out of the swap table
+ * directly. xswap_backend_uncharge() would take the same lock
+ * again, and a spinlock is not recursive.
+ */
+ memcg = xswap_entry_memcg(ci, slot);
+ if (memcg)
+ mem_cgroup_swap_uncharge(memcg, 1);
}
/*
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 17/17] mm, swap: drop a refused xswap backend run directly
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (15 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 16/17] mm, swap: do not retake the cluster lock when uncharging an xswap slot Baoquan He
@ 2026-09-20 7:20 ` Baoquan He
2026-09-22 14:22 ` [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Chris Li
2026-09-24 12:17 ` Klara Modin
18 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20 7:20 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt, Baoquan He
On a refused charge, xswap_backend_alloc() gave the physical run back
with __xswap_backend_unlink(), which finds the run through the reverse
mapping. That mapping is not installed yet at this point, so it
recognised nothing and freed nothing: every refused charge leaked the
run and tripped a VM_WARN_ON_ONCE().
Drop the run directly instead; the caller holds it and nothing points
at it yet.
Seen with a cgroup at memory.swap.max: the backend held 48663 pages
while memory.swap.current was 0.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 35 ++++++++++++++++++++++++++++++++++-
1 file changed, 34 insertions(+), 1 deletion(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 7a67c2ff4e89..2308d99b5e7c 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -3161,6 +3161,39 @@ static void xswap_free_phys_slot(struct swap_info_struct *si,
swap_range_free(si, offset, 1);
}
+/*
+ * Give a run of backend slots back. Nothing points at them yet, so there is
+ * no reverse mapping to look up and no record to clear.
+ */
+static void xswap_drop_phys_slots(swp_entry_t phys, unsigned int nr)
+{
+ struct swap_info_struct *psi = swap_type_to_info(swp_type(phys));
+ unsigned long poff = swp_offset(phys);
+ struct swap_cluster_info *pci;
+ unsigned int i;
+
+ if (!psi)
+ return;
+
+ pci = swap_cluster_lock(psi, poff);
+ if (!pci)
+ return;
+
+ VM_WARN_ON(pci->count < nr);
+ pci->count -= nr;
+ for (i = 0; i < nr; i++)
+ __swap_table_set(pci, (poff + i) % SWAPFILE_CLUSTER,
+ null_to_swp_tb());
+
+ if (!pci->count)
+ free_cluster(psi, pci);
+ else
+ partial_free_cluster(psi, pci);
+ swap_cluster_unlock(pci);
+
+ swap_range_free(psi, poff, nr);
+}
+
/*
* An xswap slot being freed may still own a physical slot. Drop it, or the
* physical space and its reverse mapping leak. Caller holds the xswap
@@ -4401,7 +4434,7 @@ swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order)
swap_cluster_unlock(ci);
}
if (mem_cgroup_swap_charge(memcg, nr)) {
- __xswap_backend_unlink(entry, phys, order);
+ xswap_drop_phys_slots(phys, nr);
return (swp_entry_t){};
}
--
2.54.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 13/17] mm, swap: do not charge zswap-backed xswap entries
2026-09-20 7:20 ` [RFC PATCH 13/17] mm, swap: do not charge zswap-backed xswap entries Baoquan He
@ 2026-09-21 11:54 ` Chris Li
0 siblings, 0 replies; 25+ messages in thread
From: Chris Li @ 2026-09-21 11:54 UTC (permalink / raw)
To: Baoquan He
Cc: linux-mm, akpm, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt
On Sat, Sep 19, 2026 at 9:22 PM Baoquan He <hebaoquan@kylinos.cn> wrote:
>
> From: Nhat Pham <nphamcs@gmail.com>
>
> An xswap entry occupies no physical swap space: its data lives in zswap
> until a backend is taken. It was nevertheless charged against
> memcg->swap at allocation, as if it did.
>
> Record the memcg at allocation but do not charge. The charge is taken
> when the entry acquires physical backing, and released with it.
See my other email thread discussion. This is a user-visible behavior
change that breaks existing deployment.
Please remove this patch from your series.
Chris
>
> memory.swap.current therefore counts only on-disk swap usage, not
> zswap-backed xswap entries. A cgroup can reclaim its anon memory with
> memory.swap.max set to 0, provided zswap is allowed for it.
>
> Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
> Signed-off-by: Nhat Pham <nphamcs@gmail.com>
> Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
> ---
> mm/swapfile.c | 21 ++++++++++++++++-----
> 1 file changed, 16 insertions(+), 5 deletions(-)
>
> diff --git a/mm/swapfile.c b/mm/swapfile.c
> index 7dcb4b48645d..627f2cab8077 100644
> --- a/mm/swapfile.c
> +++ b/mm/swapfile.c
> @@ -2002,7 +2002,12 @@ int folio_alloc_swap(struct folio *folio)
> */
> memcg = mem_cgroup_swap_get(folio);
> if (memcg) {
> - if (unlikely(mem_cgroup_swap_charge(memcg, size))) {
> + /*
> + * An xswap entry has no physical swap yet, so only record the
> + * memcg here. Its backing is charged when one is taken.
> + */
> + if (!(__swap_entry_to_info(folio->swap)->flags & SWP_XSWAP) &&
> + unlikely(mem_cgroup_swap_charge(memcg, size))) {
> mem_cgroup_swap_put(memcg, size);
> swap_cache_del_folio(folio);
> goto failed;
> @@ -2153,14 +2158,19 @@ struct swap_info_struct *get_swap_device(swp_entry_t entry)
> return ERR_PTR(-EIO);
> }
>
> -static void memcg_swap_free(unsigned short id, unsigned int nr)
> +static void memcg_swap_free(unsigned short id, unsigned int nr, bool is_xswap)
> {
> struct mem_cgroup *memcg;
>
> rcu_read_lock();
> memcg = mem_cgroup_from_private_id(id);
> if (memcg) {
> - mem_cgroup_swap_uncharge(memcg, nr);
> + /*
> + * An xswap entry was not charged here; its backing, if any,
> + * was uncharged when it was released.
> + */
> + if (!is_xswap)
> + mem_cgroup_swap_uncharge(memcg, nr);
> mem_cgroup_swap_put(memcg, nr);
> }
> rcu_read_unlock();
> @@ -2179,6 +2189,7 @@ void __swap_cluster_free_entries(struct swap_info_struct *si,
> unsigned int ci_off = ci_start, ci_end = ci_start + nr_pages;
> unsigned long ci_head = cluster_offset(si, ci);
> unsigned int batch_off = ci_off;
> + bool is_xswap = si->flags & SWP_XSWAP;
>
> VM_WARN_ON(ci->count < nr_pages);
>
> @@ -2208,14 +2219,14 @@ void __swap_cluster_free_entries(struct swap_info_struct *si,
> id_cur = __swap_cgroup_clear(ci, ci_off, 1);
> if (batch_id != id_cur) {
> if (batch_id)
> - memcg_swap_free(batch_id, ci_off - batch_off);
> + memcg_swap_free(batch_id, ci_off - batch_off, is_xswap);
> batch_id = id_cur;
> batch_off = ci_off;
> }
> } while (++ci_off < ci_end);
>
> if (batch_id)
> - memcg_swap_free(batch_id, ci_off - batch_off);
> + memcg_swap_free(batch_id, ci_off - batch_off, is_xswap);
>
> swap_range_free(si, ci_head + ci_start, nr_pages);
> swap_cluster_assert_empty(ci, ci_start, nr_pages, false);
> --
> 2.54.0
>
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (16 preceding siblings ...)
2026-09-20 7:20 ` [RFC PATCH 17/17] mm, swap: drop a refused xswap backend run directly Baoquan He
@ 2026-09-22 14:22 ` Chris Li
2026-09-22 14:30 ` Johannes Weiner
2026-09-24 1:56 ` KunWu Chan
2026-09-24 12:17 ` Klara Modin
18 siblings, 2 replies; 25+ messages in thread
From: Chris Li @ 2026-09-22 14:22 UTC (permalink / raw)
To: Baoquan He
Cc: linux-mm, akpm, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt
Hi Baoquan,
First of all, thank you so much for the great work. That closes a huge
gap in the xswap stack. It is a huge step forward.
On Sat, Sep 19, 2026 at 9:20 PM Baoquan He <hebaoquan@kylinos.cn> wrote:
>
> This is the writeback layer for xswap, on top of the base series. Post it
> as RFC for discussion.
>
> The base series keeps every swapped-out page in zswap. So the pool must
> refuse new pages once it's full. This series gives an xswap slot a
> backend to go, and charge it only when it really goes to real disk
> space.
We should also preserve the charge to the swap counter when swapping
out to xswap. I don't mind having an additional counter track
compression memory vs. disk usage, etc. The swap counter charge should
preserve the current behavior.
> Design
> ------
> - Backend: when zswap refuses a page, take a slot on the real swap
> device with the highest priority and write the page there. One IO per
> slot.
More than that, it should also happen when zswap needs to do the
writeback. Zswap has some write back hooks when the system is under
pressure. Those hooks should be introduced into the core swap stack so
other swap types can benefit from them as well.
I would like to point out that the allocation strategy should be
different when allocating for the swap backend. We might want to
introduce an allocation flag for that. The backend doesn't care about
fragmentation as much. Fragmentation is a performance optimization
here. The backend can tolerate fragmentation. In the upper layer, swap
entries are 2^N aligned and sized. It is a hard requirement.
> - Ownership: that slot has no swap cache folio, so record its owner in
Clarify: by "that slot," you mean the backend swap entry has no swap
cache folio.
> the swap table entry. 0b100 in the low bits marks a backend slot, and
> the xswap type/offset go above. The physical side then does not treat
> it as free, and readahead skips it.
I don't want to get ahead of ourselves. You have only considered
writeback from xswap to disk so far. A stretch goal for you is to
consider how we can mark the swap table so that it can work with other
swap types as well. e.g. e.g., writing data from SSD to HDD. We don't
have to do it in this series. But it is good to start considering this
kind of thing so that we don't lock ourself out when we want more
generic writeback between swap tiers.
Stretch 2 is considering how to deal with data moving from xswap to
SSD, then SSD to HDD. The front swap only needs to point to the HDD
and skip the SSD completely. It shouldn't require a forward chain.
> - Per-slot record: ci->xs_table[], one unsigned long per slot, allocated
> the first time a cluster takes a backend. Zero means no backend.
There might be some micro optimization we can do if the whole cluster
does not point to any back ends, we might be able to design the array
access so that we don't need to allocate the back end pointer array
for it.
> - Read: a written-out slot has no zswap copy, so swap_read_folio()
> follows the pointer and reads from the backend.
We need to pay attention to the race condition here when moving data
from xswap to the backend. I am not saying you have a bug here, I
haven't read the code yet. I imagine there will be some very tricky
swap cache interaction going on here.
> - Release: verify the slot still points back to the xswap entry first;
> the record can go stale.
What do you mean the record can go stale here?
> - swapoff: try_to_unuse() used to skip these slots. It now reads them
> back, marks the folio dirty so reclaim re-stores it in zswap, then
> drops the backend.
It seems data can remain in zswap past swapoff? That seems wrong to
me. In my mind, it should just read into the swap cache or promote to
the upper swap tier.
> - Reclaim: an xswap entry's count can reach 0 while its folio is still
> in the swap cache; the physical scanner reclaims the redundant slot
> through __try_to_reclaim_swap().
Clarify: I assume you mean swap cache reclaim not the memory reclaim.
> - Large folios: a read starts at the first physical slot, so the run is
> reserved whole and must be backed contiguously; otherwise the fault
> retries at a smaller order. xswap is SWP_SYNCHRONOUS_IO, which also
> skips readahead.
I imagine the front swap needs to be contiguous, but the back swap doesn't?
> - Charging: Record the owner at allocation but do not charge. Take the
> charge when the entry gets a physical backend slot, and give it back
> with that slot.memory.swap.current then counts only real on-disk swap
> usage, and a cgroup with memory.swap.max at 0 can still swap out through
> zswap.
I want the old swap charging behavior back. Please don't change that
behavior. That behavior change should be a separate discussion and
split out from this series. BTW, I am against that change.
Great work.
Chris
>
> Note
> ----
> The base series is sereis of xswap foundation. This one depends on it.
Do you have a git repo I can fetch from? I will try applying your
patches as well. I assume mm-unstable should be the base?
> So if anyone wants to apply this patchset, the order is base-commit
> as below, then xswap foundation, finally this patchset.
>
> [PATCH v3 00/14] mm, swap: extendable swap devices (xswap)
> https://lore.kernel.org/all/20260916101929.149106-1-hebaoquan@kylinos.cn/T/#u
>
> base-commit: baa8de2f3448d1466a888a805c18d01c998fe052
>
> This partially refers to Nhat's vswap-v4 series. E.g patch 1 is
> consistent with his patch 3, patch 12/13/14/15 are from his patch 8.
>
> And the last 2 patches are fixing bugs when charging patches are added.
> Since the charging of xswap is still under discussion, I dind't merge
> them into commits. Will squash them or take them off once decision is
> made.
>
> Testing
> -------
> qemu KVM guest, 8G RAM.
>
> A 4G swap disk /dev/vdb is added as the physical backend, and zswap is
> turned off so that every page has to reach a backend slot. Note that it
> need create xswap device firstly then disable zswap, so every page has
> to reach backend slot.
>
> The workload is memhog, as in the base series. Every page is filled with
> a fixed pattern, so the data can be checked after a readback.
>
> # echo 1 > /sys/module/zswap/parameters/enabled
> # echo 100 > /sys/kernel/mm/xswap/create
> # mkswap /dev/vdb && swapon -p 0 /dev/vdb
> # echo 0 > /sys/module/zswap/parameters/enabled
> # mkdir -p /sys/fs/cgroup/xswap_limit
> # echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
> # MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 ./memhog
>
> 1. Swapout to the backend
>
> Run the workload with 5G of memory in a cgroup capped at 4G:
>
> # awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
> Filename Type Size Used Priority
> xswap0 xswap 8155132 2392176 100
> /dev/vdb partition 4194300 2392176 0
>
> The two Used values are equal, so the pages reached the backend and
> the xswap side accounted for them.
>
> 2. Pages are read back from the device
>
> Lifting memory.max and touching the pages again reads them back. The
> pages-swapped-in counter went up by 265642, so the read did reach the
> device. 1010 of the folios read back were 1M, each read with one IO:
>
> # echo max > /sys/fs/cgroup/xswap_limit/memory.max
> # cat /sys/kernel/mm/transparent_hugepage/hugepages-1024kB/stats/swpin
> 1010
>
> hugepages-2048kB/stats/swpin stays at 0, because the swapin order is
> capped below the PMD order.
>
> 3. verifies data correctness
>
> After the readback the data is compared byte for byte against the
> pattern, one time with zswap on and one time with zswap off, so the
> pages come from zswap in one time and from the backend in the other
> time. All bytes matched.
>
> 4. Destroying a device returns its backend slots
>
> The device was filled, then destroyed with no readback first:
>
> # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
> 3157220
> # echo 0 > /sys/kernel/mm/xswap/destroy
> # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
> 0
>
> 5. verifies the charge is correct
>
> An xswap entry is charged only when it holds a backend slot, so the
> two counters have to show the same number of pages:
>
> # echo $(( $(cat /sys/fs/cgroup/xswap_limit/memory.swap.current) / 4096 ))
> 526852
> # awk '$1 == "/dev/vdb" { print $4 / 4 }' /proc/swaps
> 526852
>
> Nothing is charged while the pages stay in zswap: with the pool holding
> them and the backend untouched, memory.swap.current stays 0.
>
> 6. A cgroup with no swap room can still reclaim its anon memory
>
> With memory.swap.max at 0 and no xswap device, the cgroup ran out of
> room and the workload was killed. With an xswap device and zswap on,
> it swapped out through zswap, was charged nothing, and was not killed.
>
> E.g if memory.swap.max is set at 100M, the cgroup will be killed when
> the cap was reached.
>
> Baoquan He (12):
> mm, swap: tag a swap table entry with its owning xswap entry
> mm, swap: prepare the folio-less allocation path for xswap
> mm, swap: add a physical backend for xswap slots
> mm, swap: use the xswap physical backend
> mm, swap: fall back to disk when zswap refuses an xswap page
> mm, swap: support swapoff of an xswap physical backend
> mm, swap: reclaim physical slots backing cache-only xswap entries
> mm, swap: back a large xswap folio with a contiguous physical run
> mm, swap: enable THP swapin for xswap entries
> mm, swap: drop swap_folio_sector()
> mm, swap: do not retake the cluster lock when uncharging an xswap slot
> mm, swap: drop a refused xswap backend run directly
>
> Nhat Pham (5):
> mm, swap: prepare the swap IO path for xswap backends
> mm, swap: split the swap memcg charge helpers
> mm, swap: do not charge zswap-backed xswap entries
> mm, swap: charge an xswap entry when it gets physical backing
> mm, swap: don't gate xswap on the physical swap free count
>
> .../admin-guide/cgroup-v1/memcg_test.rst | 2 +-
> include/linux/memcontrol.h | 6 +
> include/linux/swap.h | 69 +-
> include/linux/swap_ops.h | 9 +-
> mm/memcontrol-v1.c | 10 +-
> mm/memcontrol.c | 147 ++--
> mm/memory.c | 7 +-
> mm/page_io.c | 95 ++-
> mm/swap.h | 35 +-
> mm/swap_state.c | 12 +-
> mm/swap_table.h | 53 ++
> mm/swapfile.c | 742 ++++++++++++++++--
> mm/zswap.c | 20 +-
> 13 files changed, 1026 insertions(+), 181 deletions(-)
>
> --
> 2.54.0
>
>
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend
2026-09-22 14:22 ` [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Chris Li
@ 2026-09-22 14:30 ` Johannes Weiner
2026-09-24 1:56 ` KunWu Chan
1 sibling, 0 replies; 25+ messages in thread
From: Johannes Weiner @ 2026-09-22 14:30 UTC (permalink / raw)
To: Chris Li
Cc: Baoquan He, linux-mm, akpm, kasong, nphamcs, baohua,
youngjun.park, david, kunwu.chan, baoquan.he, gourry, riel,
mhocko, roman.gushchin, shakeel.butt
On Tue, Sep 22, 2026 at 04:22:51AM -1000, Chris Li wrote:
> On Sat, Sep 19, 2026 at 9:20 PM Baoquan He <hebaoquan@kylinos.cn> wrote:
> > - Charging: Record the owner at allocation but do not charge. Take the
> > charge when the entry gets a physical backend slot, and give it back
> > with that slot.memory.swap.current then counts only real on-disk swap
> > usage, and a cgroup with memory.swap.max at 0 can still swap out through
> > zswap.
>
> I want the old swap charging behavior back. Please don't change that
> behavior. That behavior change should be a separate discussion and
> split out from this series. BTW, I am against that change.
Charging the swap counter is a hard NAK from the cgroup side.
The reasons have been spelled out over and over. Feel free to engage
with those. But what you're doing here is not productive or in good
faith.
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend
2026-09-22 14:22 ` [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Chris Li
2026-09-22 14:30 ` Johannes Weiner
@ 2026-09-24 1:56 ` KunWu Chan
2026-09-24 11:36 ` Chris Li
1 sibling, 1 reply; 25+ messages in thread
From: KunWu Chan @ 2026-09-24 1:56 UTC (permalink / raw)
To: Chris Li
Cc: Baoquan He, linux-mm, akpm, kasong, hannes, nphamcs, baohua,
youngjun.park, david, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt
On Tue, Sep 22, 2026 at 10:23 PM Chris Li <chrisl@kernel.org> wrote:
>
> Hi Baoquan,
>
> First of all, thank you so much for the great work. That closes a huge
> gap in the xswap stack. It is a huge step forward.
>
> On Sat, Sep 19, 2026 at 9:20 PM Baoquan He <hebaoquan@kylinos.cn> wrote:
> >
> > This is the writeback layer for xswap, on top of the base series. Post it
> > as RFC for discussion.
> >
> > The base series keeps every swapped-out page in zswap. So the pool must
> > refuse new pages once it's full. This series gives an xswap slot a
> > backend to go, and charge it only when it really goes to real disk
> > space.
>
> We should also preserve the charge to the swap counter when swapping
> out to xswap. I don't mind having an additional counter track
> compression memory vs. disk usage, etc. The swap counter charge should
> preserve the current behavior.
>
> > Design
> > ------
> > - Backend: when zswap refuses a page, take a slot on the real swap
> > device with the highest priority and write the page there. One IO per
> > slot.
>
> More than that, it should also happen when zswap needs to do the
> writeback. Zswap has some write back hooks when the system is under
> pressure. Those hooks should be introduced into the core swap stack so
> other swap types can benefit from them as well.
>
> I would like to point out that the allocation strategy should be
> different when allocating for the swap backend. We might want to
> introduce an allocation flag for that. The backend doesn't care about
> fragmentation as much. Fragmentation is a performance optimization
> here. The backend can tolerate fragmentation. In the upper layer, swap
> entries are 2^N aligned and sized. It is a hard requirement.
>
> > - Ownership: that slot has no swap cache folio, so record its owner in
>
> Clarify: by "that slot," you mean the backend swap entry has no swap
> cache folio.
>
> > the swap table entry. 0b100 in the low bits marks a backend slot, and
> > the xswap type/offset go above. The physical side then does not treat
> > it as free, and readahead skips it.
>
> I don't want to get ahead of ourselves. You have only considered
> writeback from xswap to disk so far. A stretch goal for you is to
> consider how we can mark the swap table so that it can work with other
> swap types as well. e.g. e.g., writing data from SSD to HDD. We don't
> have to do it in this series. But it is good to start considering this
> kind of thing so that we don't lock ourself out when we want more
> generic writeback between swap tiers.
>
> Stretch 2 is considering how to deal with data moving from xswap to
> SSD, then SSD to HDD. The front swap only needs to point to the HDD
> and skip the SSD completely. It shouldn't require a forward chain.
>
> > - Per-slot record: ci->xs_table[], one unsigned long per slot, allocated
> > the first time a cluster takes a backend. Zero means no backend.
>
> There might be some micro optimization we can do if the whole cluster
> does not point to any back ends, we might be able to design the array
> access so that we don't need to allocate the back end pointer array
> for it.
>
> > - Read: a written-out slot has no zswap copy, so swap_read_folio()
> > follows the pointer and reads from the backend.
>
> We need to pay attention to the race condition here when moving data
> from xswap to the backend. I am not saying you have a bug here, I
> haven't read the code yet. I imagine there will be some very tricky
> swap cache interaction going on here.
>
Hi Chris,
> > - Release: verify the slot still points back to the xswap entry first;
> > the record can go stale.
>
> What do you mean the record can go stale here?
>
I read the "stale record" here as the xswap-side xs_table[] record
of the physical slot in patch 04/17. xswap_backend_alloc() sets it
with xswap_slot_set_backend(), while the physical slot also has a
reverse mapping installed by xswap_install_rmap().
The two records are maintained separately. In xswap_backend_free(),
the forward record is cleared only if it still matches, and
xswap_free_phys_slot() then checks the reverse mapping in the
physical slot's swap table entry against the xswap entry being freed.
If the physical slot has already been recycled for another xswap
entry, that check fails and the slot is left untouched.
So my reading is that xs_table[] can still record a physical slot
after that slot has been recycled, while the reverse mapping in the
physical slot's swap table entry identifies a different xswap owner.
Thanks,
Kunwu
> > - swapoff: try_to_unuse() used to skip these slots. It now reads them
> > back, marks the folio dirty so reclaim re-stores it in zswap, then
> > drops the backend.
>
> It seems data can remain in zswap past swapoff? That seems wrong to
> me. In my mind, it should just read into the swap cache or promote to
> the upper swap tier.
>
> > - Reclaim: an xswap entry's count can reach 0 while its folio is still
> > in the swap cache; the physical scanner reclaims the redundant slot
> > through __try_to_reclaim_swap().
>
> Clarify: I assume you mean swap cache reclaim not the memory reclaim.
>
> > - Large folios: a read starts at the first physical slot, so the run is
> > reserved whole and must be backed contiguously; otherwise the fault
> > retries at a smaller order. xswap is SWP_SYNCHRONOUS_IO, which also
> > skips readahead.
>
> I imagine the front swap needs to be contiguous, but the back swap doesn't?
>
> > - Charging: Record the owner at allocation but do not charge. Take the
> > charge when the entry gets a physical backend slot, and give it back
> > with that slot.memory.swap.current then counts only real on-disk swap
> > usage, and a cgroup with memory.swap.max at 0 can still swap out through
> > zswap.
>
> I want the old swap charging behavior back. Please don't change that
> behavior. That behavior change should be a separate discussion and
> split out from this series. BTW, I am against that change.
>
> Great work.
>
> Chris
>
> >
> > Note
> > ----
> > The base series is sereis of xswap foundation. This one depends on it.
>
> Do you have a git repo I can fetch from? I will try applying your
> patches as well. I assume mm-unstable should be the base?
>
> > So if anyone wants to apply this patchset, the order is base-commit
> > as below, then xswap foundation, finally this patchset.
> >
> > [PATCH v3 00/14] mm, swap: extendable swap devices (xswap)
> > https://lore.kernel.org/all/20260916101929.149106-1-hebaoquan@kylinos.cn/T/#u
> >
> > base-commit: baa8de2f3448d1466a888a805c18d01c998fe052
> >
> > This partially refers to Nhat's vswap-v4 series. E.g patch 1 is
> > consistent with his patch 3, patch 12/13/14/15 are from his patch 8.
> >
> > And the last 2 patches are fixing bugs when charging patches are added.
> > Since the charging of xswap is still under discussion, I dind't merge
> > them into commits. Will squash them or take them off once decision is
> > made.
> >
> > Testing
> > -------
> > qemu KVM guest, 8G RAM.
> >
> > A 4G swap disk /dev/vdb is added as the physical backend, and zswap is
> > turned off so that every page has to reach a backend slot. Note that it
> > need create xswap device firstly then disable zswap, so every page has
> > to reach backend slot.
> >
> > The workload is memhog, as in the base series. Every page is filled with
> > a fixed pattern, so the data can be checked after a readback.
> >
> > # echo 1 > /sys/module/zswap/parameters/enabled
> > # echo 100 > /sys/kernel/mm/xswap/create
> > # mkswap /dev/vdb && swapon -p 0 /dev/vdb
> > # echo 0 > /sys/module/zswap/parameters/enabled
> > # mkdir -p /sys/fs/cgroup/xswap_limit
> > # echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
> > # MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 ./memhog
> >
> > 1. Swapout to the backend
> >
> > Run the workload with 5G of memory in a cgroup capped at 4G:
> >
> > # awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
> > Filename Type Size Used Priority
> > xswap0 xswap 8155132 2392176 100
> > /dev/vdb partition 4194300 2392176 0
> >
> > The two Used values are equal, so the pages reached the backend and
> > the xswap side accounted for them.
> >
> > 2. Pages are read back from the device
> >
> > Lifting memory.max and touching the pages again reads them back. The
> > pages-swapped-in counter went up by 265642, so the read did reach the
> > device. 1010 of the folios read back were 1M, each read with one IO:
> >
> > # echo max > /sys/fs/cgroup/xswap_limit/memory.max
> > # cat /sys/kernel/mm/transparent_hugepage/hugepages-1024kB/stats/swpin
> > 1010
> >
> > hugepages-2048kB/stats/swpin stays at 0, because the swapin order is
> > capped below the PMD order.
> >
> > 3. verifies data correctness
> >
> > After the readback the data is compared byte for byte against the
> > pattern, one time with zswap on and one time with zswap off, so the
> > pages come from zswap in one time and from the backend in the other
> > time. All bytes matched.
> >
> > 4. Destroying a device returns its backend slots
> >
> > The device was filled, then destroyed with no readback first:
> >
> > # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
> > 3157220
> > # echo 0 > /sys/kernel/mm/xswap/destroy
> > # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
> > 0
> >
> > 5. verifies the charge is correct
> >
> > An xswap entry is charged only when it holds a backend slot, so the
> > two counters have to show the same number of pages:
> >
> > # echo $(( $(cat /sys/fs/cgroup/xswap_limit/memory.swap.current) / 4096 ))
> > 526852
> > # awk '$1 == "/dev/vdb" { print $4 / 4 }' /proc/swaps
> > 526852
> >
> > Nothing is charged while the pages stay in zswap: with the pool holding
> > them and the backend untouched, memory.swap.current stays 0.
> >
> > 6. A cgroup with no swap room can still reclaim its anon memory
> >
> > With memory.swap.max at 0 and no xswap device, the cgroup ran out of
> > room and the workload was killed. With an xswap device and zswap on,
> > it swapped out through zswap, was charged nothing, and was not killed.
> >
> > E.g if memory.swap.max is set at 100M, the cgroup will be killed when
> > the cap was reached.
> >
> > Baoquan He (12):
> > mm, swap: tag a swap table entry with its owning xswap entry
> > mm, swap: prepare the folio-less allocation path for xswap
> > mm, swap: add a physical backend for xswap slots
> > mm, swap: use the xswap physical backend
> > mm, swap: fall back to disk when zswap refuses an xswap page
> > mm, swap: support swapoff of an xswap physical backend
> > mm, swap: reclaim physical slots backing cache-only xswap entries
> > mm, swap: back a large xswap folio with a contiguous physical run
> > mm, swap: enable THP swapin for xswap entries
> > mm, swap: drop swap_folio_sector()
> > mm, swap: do not retake the cluster lock when uncharging an xswap slot
> > mm, swap: drop a refused xswap backend run directly
> >
> > Nhat Pham (5):
> > mm, swap: prepare the swap IO path for xswap backends
> > mm, swap: split the swap memcg charge helpers
> > mm, swap: do not charge zswap-backed xswap entries
> > mm, swap: charge an xswap entry when it gets physical backing
> > mm, swap: don't gate xswap on the physical swap free count
> >
> > .../admin-guide/cgroup-v1/memcg_test.rst | 2 +-
> > include/linux/memcontrol.h | 6 +
> > include/linux/swap.h | 69 +-
> > include/linux/swap_ops.h | 9 +-
> > mm/memcontrol-v1.c | 10 +-
> > mm/memcontrol.c | 147 ++--
> > mm/memory.c | 7 +-
> > mm/page_io.c | 95 ++-
> > mm/swap.h | 35 +-
> > mm/swap_state.c | 12 +-
> > mm/swap_table.h | 53 ++
> > mm/swapfile.c | 742 ++++++++++++++++--
> > mm/zswap.c | 20 +-
> > 13 files changed, 1026 insertions(+), 181 deletions(-)
> >
> > --
> > 2.54.0
> >
> >
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend
2026-09-24 1:56 ` KunWu Chan
@ 2026-09-24 11:36 ` Chris Li
0 siblings, 0 replies; 25+ messages in thread
From: Chris Li @ 2026-09-24 11:36 UTC (permalink / raw)
To: KunWu Chan
Cc: Baoquan He, linux-mm, akpm, kasong, hannes, nphamcs, baohua,
youngjun.park, david, baoquan.he, gourry, riel, mhocko,
roman.gushchin, shakeel.butt
On Wed, Sep 23, 2026 at 3:56 PM KunWu Chan <kunwu.chan@gmail.com> wrote:
>
> Hi Chris,
>
> > > - Release: verify the slot still points back to the xswap entry first;
> > > the record can go stale.
> >
> > What do you mean the record can go stale here?
> >
>
> I read the "stale record" here as the xswap-side xs_table[] record
> of the physical slot in patch 04/17. xswap_backend_alloc() sets it
> with xswap_slot_set_backend(), while the physical slot also has a
> reverse mapping installed by xswap_install_rmap().
>
> The two records are maintained separately. In xswap_backend_free(),
> the forward record is cleared only if it still matches, and
> xswap_free_phys_slot() then checks the reverse mapping in the
> physical slot's swap table entry against the xswap entry being freed.
> If the physical slot has already been recycled for another xswap
> entry, that check fails and the slot is left untouched.
>
> So my reading is that xs_table[] can still record a physical slot
> after that slot has been recycled, while the reverse mapping in the
> physical slot's swap table entry identifies a different xswap owner.
I see, thanks for the explanation. So the original sentence can be written as:
"the record can go stale and point to a different xswap entry".
Thanks again
Chris
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
` (17 preceding siblings ...)
2026-09-22 14:22 ` [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Chris Li
@ 2026-09-24 12:17 ` Klara Modin
2026-09-28 3:30 ` Baoquan He
18 siblings, 1 reply; 25+ messages in thread
From: Klara Modin @ 2026-09-24 12:17 UTC (permalink / raw)
To: Baoquan He
Cc: linux-mm, akpm, chrisl, kasong, hannes, nphamcs, baohua,
youngjun.park, david, kunwu.chan, baoquan.he, gourry, riel,
mhocko, roman.gushchin, shakeel.butt
On 2026-09-20 15:20:26 +0800, Baoquan He wrote:
> This is the writeback layer for xswap, on top of the base series. Post it
> as RFC for discussion.
>
> The base series keeps every swapped-out page in zswap. So the pool must
> refuse new pages once it's full. This series gives an xswap slot a
> backend to go, and charge it only when it really goes to real disk
> space.
There is some weird trailing whitespace in this email for some sections.
>
> Design
> ------
> - Backend: when zswap refuses a page, take a slot on the real swap
> device with the highest priority and write the page there. One IO per
> slot.
> - Ownership: that slot has no swap cache folio, so record its owner in
> the swap table entry. 0b100 in the low bits marks a backend slot, and
> the xswap type/offset go above. The physical side then does not treat
> it as free, and readahead skips it.
> - Per-slot record: ci->xs_table[], one unsigned long per slot, allocated
> the first time a cluster takes a backend. Zero means no backend.
> - Read: a written-out slot has no zswap copy, so swap_read_folio()
> follows the pointer and reads from the backend.
> - Release: verify the slot still points back to the xswap entry first;
> the record can go stale.
> - swapoff: try_to_unuse() used to skip these slots. It now reads them
> back, marks the folio dirty so reclaim re-stores it in zswap, then
> drops the backend.
> - Reclaim: an xswap entry's count can reach 0 while its folio is still
> in the swap cache; the physical scanner reclaims the redundant slot
> through __try_to_reclaim_swap().
> - Large folios: a read starts at the first physical slot, so the run is
> reserved whole and must be backed contiguously; otherwise the fault
> retries at a smaller order. xswap is SWP_SYNCHRONOUS_IO, which also
> skips readahead.
> - Charging: Record the owner at allocation but do not charge. Take the
> charge when the entry gets a physical backend slot, and give it back
> with that slot.memory.swap.current then counts only real on-disk swap
> usage, and a cgroup with memory.swap.max at 0 can still swap out through
> zswap.
>
> Note
> ----
> The base series is sereis of xswap foundation. This one depends on it.
> So if anyone wants to apply this patchset, the order is base-commit
> as below, then xswap foundation, finally this patchset.
>
> [PATCH v3 00/14] mm, swap: extendable swap devices (xswap)
> https://lore.kernel.org/all/20260916101929.149106-1-hebaoquan@kylinos.cn/T/#u
>
> base-commit: baa8de2f3448d1466a888a805c18d01c998fe052
>
> This partially refers to Nhat's vswap-v4 series. E.g patch 1 is
> consistent with his patch 3, patch 12/13/14/15 are from his patch 8.
>
> And the last 2 patches are fixing bugs when charging patches are added.
> Since the charging of xswap is still under discussion, I dind't merge
> them into commits. Will squash them or take them off once decision is
> made.
>
> Testing
> -------
> qemu KVM guest, 8G RAM.
>
> A 4G swap disk /dev/vdb is added as the physical backend, and zswap is
> turned off so that every page has to reach a backend slot. Note that it
> need create xswap device firstly then disable zswap, so every page has
> to reach backend slot.
>
> The workload is memhog, as in the base series. Every page is filled with
> a fixed pattern, so the data can be checked after a readback.
>
> # echo 1 > /sys/module/zswap/parameters/enabled
> # echo 100 > /sys/kernel/mm/xswap/create
> # mkswap /dev/vdb && swapon -p 0 /dev/vdb
> # echo 0 > /sys/module/zswap/parameters/enabled
> # mkdir -p /sys/fs/cgroup/xswap_limit
> # echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
> # MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 ./memhog
You did not include e.g. cgexec here, so is memhog really running in the
cgroup you created?
>
> 1. Swapout to the backend
>
> Run the workload with 5G of memory in a cgroup capped at 4G:
>
> # awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
> Filename Type Size Used Priority
> xswap0 xswap 8155132 2392176 100
> /dev/vdb partition 4194300 2392176 0
>
> The two Used values are equal, so the pages reached the backend and
> the xswap side accounted for them.
This is a bit confusing to me. If I now run free, the written back data
will be represented twice in the sum? Since the written back data is
still presented as used in the xswap device, does it still consume space
there, i.e. does it prevent that amount of new data being stored in the
xswap device? If this is the case, it would make it even harder to
dimension the xswap size since one would also have to take into account
all the other swap devices which are present in the system.
>
> 2. Pages are read back from the device
>
> Lifting memory.max and touching the pages again reads them back. The
> pages-swapped-in counter went up by 265642, so the read did reach the
> device. 1010 of the folios read back were 1M, each read with one IO:
>
> # echo max > /sys/fs/cgroup/xswap_limit/memory.max
> # cat /sys/kernel/mm/transparent_hugepage/hugepages-1024kB/stats/swpin
> 1010
>
> hugepages-2048kB/stats/swpin stays at 0, because the swapin order is
> capped below the PMD order.
>
> 3. verifies data correctness
>
> After the readback the data is compared byte for byte against the
> pattern, one time with zswap on and one time with zswap off, so the
> pages come from zswap in one time and from the backend in the other
> time. All bytes matched.
>
> 4. Destroying a device returns its backend slots
>
> The device was filled, then destroyed with no readback first:
>
> # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
> 3157220
> # echo 0 > /sys/kernel/mm/xswap/destroy
> # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
> 0
>
> 5. verifies the charge is correct
>
> An xswap entry is charged only when it holds a backend slot, so the
> two counters have to show the same number of pages:
>
> # echo $(( $(cat /sys/fs/cgroup/xswap_limit/memory.swap.current) / 4096 ))
> 526852
> # awk '$1 == "/dev/vdb" { print $4 / 4 }' /proc/swaps
> 526852
>
> Nothing is charged while the pages stay in zswap: with the pool holding
> them and the backend untouched, memory.swap.current stays 0.
>
> 6. A cgroup with no swap room can still reclaim its anon memory
>
> With memory.swap.max at 0 and no xswap device, the cgroup ran out of
> room and the workload was killed. With an xswap device and zswap on,
> it swapped out through zswap, was charged nothing, and was not killed.
>
> E.g if memory.swap.max is set at 100M, the cgroup will be killed when
> the cap was reached.
Disregarding my comments and opinons on the interface, it seems to be
working so far and I can see writeback happen during a GCC 17 build on
my BPI-F3, but the entire thing takes roughly 18 hours and is not done
yet.
>
> Baoquan He (12):
> mm, swap: tag a swap table entry with its owning xswap entry
> mm, swap: prepare the folio-less allocation path for xswap
> mm, swap: add a physical backend for xswap slots
> mm, swap: use the xswap physical backend
> mm, swap: fall back to disk when zswap refuses an xswap page
> mm, swap: support swapoff of an xswap physical backend
> mm, swap: reclaim physical slots backing cache-only xswap entries
> mm, swap: back a large xswap folio with a contiguous physical run
> mm, swap: enable THP swapin for xswap entries
> mm, swap: drop swap_folio_sector()
> mm, swap: do not retake the cluster lock when uncharging an xswap slot
> mm, swap: drop a refused xswap backend run directly
>
> Nhat Pham (5):
> mm, swap: prepare the swap IO path for xswap backends
> mm, swap: split the swap memcg charge helpers
> mm, swap: do not charge zswap-backed xswap entries
> mm, swap: charge an xswap entry when it gets physical backing
> mm, swap: don't gate xswap on the physical swap free count
>
> .../admin-guide/cgroup-v1/memcg_test.rst | 2 +-
> include/linux/memcontrol.h | 6 +
> include/linux/swap.h | 69 +-
> include/linux/swap_ops.h | 9 +-
> mm/memcontrol-v1.c | 10 +-
> mm/memcontrol.c | 147 ++--
> mm/memory.c | 7 +-
> mm/page_io.c | 95 ++-
> mm/swap.h | 35 +-
> mm/swap_state.c | 12 +-
> mm/swap_table.h | 53 ++
> mm/swapfile.c | 742 ++++++++++++++++--
> mm/zswap.c | 20 +-
> 13 files changed, 1026 insertions(+), 181 deletions(-)
>
> --
> 2.54.0
>
>
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend
2026-09-24 12:17 ` Klara Modin
@ 2026-09-28 3:30 ` Baoquan He
0 siblings, 0 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-28 3:30 UTC (permalink / raw)
To: Klara Modin
Cc: Baoquan He, linux-mm, akpm, chrisl, kasong, hannes, nphamcs,
baohua, youngjun.park, david, kunwu.chan, gourry, riel, mhocko,
roman.gushchin, shakeel.butt
On 09/24/26 at 02:17pm, Klara Modin wrote:
> On 2026-09-20 15:20:26 +0800, Baoquan He wrote:
> > This is the writeback layer for xswap, on top of the base series. Post it
> > as RFC for discussion.
> >
> > The base series keeps every swapped-out page in zswap. So the pool must
> > refuse new pages once it's full. This series gives an xswap slot a
> > backend to go, and charge it only when it really goes to real disk
> > space.
>
> There is some weird trailing whitespace in this email for some sections.
Thanks for careful checking, will clean this up before sending out.
...snip...
> > Testing
> > -------
> > qemu KVM guest, 8G RAM.
> >
> > A 4G swap disk /dev/vdb is added as the physical backend, and zswap is
> > turned off so that every page has to reach a backend slot. Note that it
> > need create xswap device firstly then disable zswap, so every page has
> > to reach backend slot.
> >
> > The workload is memhog, as in the base series. Every page is filled with
> > a fixed pattern, so the data can be checked after a readback.
> >
> > # echo 1 > /sys/module/zswap/parameters/enabled
> > # echo 100 > /sys/kernel/mm/xswap/create
> > # mkswap /dev/vdb && swapon -p 0 /dev/vdb
> > # echo 0 > /sys/module/zswap/parameters/enabled
> > # mkdir -p /sys/fs/cgroup/xswap_limit
> > # echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
> > # MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 ./memhog
>
> You did not include e.g. cgexec here, so is memhog really running in the
> cgroup you created?
You are right, I didn't write out the full command, just list the steps.
And here memhog is not a standard one from installed rpm package
numactl, it's a small local helper. It faults <total_GB> of anon, holds
it, and can drop the mapping or read it back. And the helper keep a known
byte pattern, so the data can be checked byte for byte after it is read
back.
# echo 1 > /sys/module/zswap/parameters/enabled
# echo 100 > /sys/kernel/mm/xswap/create
# mkswap /dev/vdb && swapon -p 0 /dev/vdb
# echo 0 > /sys/module/zswap/parameters/enabled
# mkdir -p /sys/fs/cgroup/xswap_limit
# echo 4G > /sys/fs/cgroup/xswap_limit/memory.max
# echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
# ( echo $BASHPID > /sys/fs/cgroup/xswap_limit/cgroup.procs
# exec env MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 \
# /mnt/kernel_src/testing_kernel_hebq/xswap/memhog 5 600 ) &
I will add the complete block in next version, sorry for the
inconvenience.
>
> >
> > 1. Swapout to the backend
> >
> > Run the workload with 5G of memory in a cgroup capped at 4G:
> >
> > # awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
> > Filename Type Size Used Priority
> > xswap0 xswap 8155132 2392176 100
> > /dev/vdb partition 4194300 2392176 0
> >
> > The two Used values are equal, so the pages reached the backend and
> > the xswap side accounted for them.
>
> This is a bit confusing to me. If I now run free, the written back data
> will be represented twice in the sum? Since the written back data is
> still presented as used in the xswap device, does it still consume space
> there, i.e. does it prevent that amount of new data being stored in the
> xswap device? If this is the case, it would make it even harder to
> dimension the xswap size since one would also have to take into account
> all the other swap devices which are present in the system.
I have to admit you found out an issue, thanks. I noticed it too while
I havne't thought of a satisfying way to fix it. What I can think of
is to mark the usage size of writeback on disk and list it separately.
# awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
Filename Type Size Used Priority
xswap0 xswap 8155132 2392176 100
/dev/vdb partition 4194300 2392176 0
2392176 (xswap)
If xswap is disabled and full, /dev/vdb got swapped out data by its own,
then the amount of used slot will be reflected as before.
# awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
Filename Type Size Used Priority
xswap0 xswap 8155132 2392176 100
/dev/vdb partition 4194300 3000000 0
2392176 (xswap)
I think this can be fixed later, I see it as a non-blocker for now.
>
> >
> > 2. Pages are read back from the device
> >
> > Lifting memory.max and touching the pages again reads them back. The
> > pages-swapped-in counter went up by 265642, so the read did reach the
> > device. 1010 of the folios read back were 1M, each read with one IO:
> >
> > # echo max > /sys/fs/cgroup/xswap_limit/memory.max
> > # cat /sys/kernel/mm/transparent_hugepage/hugepages-1024kB/stats/swpin
> > 1010
> >
> > hugepages-2048kB/stats/swpin stays at 0, because the swapin order is
> > capped below the PMD order.
> >
> > 3. verifies data correctness
> >
> > After the readback the data is compared byte for byte against the
> > pattern, one time with zswap on and one time with zswap off, so the
> > pages come from zswap in one time and from the backend in the other
> > time. All bytes matched.
> >
> > 4. Destroying a device returns its backend slots
> >
> > The device was filled, then destroyed with no readback first:
> >
> > # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
> > 3157220
> > # echo 0 > /sys/kernel/mm/xswap/destroy
> > # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
> > 0
> >
> > 5. verifies the charge is correct
> >
> > An xswap entry is charged only when it holds a backend slot, so the
> > two counters have to show the same number of pages:
> >
> > # echo $(( $(cat /sys/fs/cgroup/xswap_limit/memory.swap.current) / 4096 ))
> > 526852
> > # awk '$1 == "/dev/vdb" { print $4 / 4 }' /proc/swaps
> > 526852
> >
> > Nothing is charged while the pages stay in zswap: with the pool holding
> > them and the backend untouched, memory.swap.current stays 0.
> >
> > 6. A cgroup with no swap room can still reclaim its anon memory
> >
> > With memory.swap.max at 0 and no xswap device, the cgroup ran out of
> > room and the workload was killed. With an xswap device and zswap on,
> > it swapped out through zswap, was charged nothing, and was not killed.
> >
> > E.g if memory.swap.max is set at 100M, the cgroup will be killed when
> > the cap was reached.
>
> Disregarding my comments and opinons on the interface, it seems to be
> working so far and I can see writeback happen during a GCC 17 build on
> my BPI-F3, but the entire thing takes roughly 18 hours and is not done
> yet.
Thanks again for your careful checking, testing.
Thanks
Baoquan
>
> >
> > Baoquan He (12):
> > mm, swap: tag a swap table entry with its owning xswap entry
> > mm, swap: prepare the folio-less allocation path for xswap
> > mm, swap: add a physical backend for xswap slots
> > mm, swap: use the xswap physical backend
> > mm, swap: fall back to disk when zswap refuses an xswap page
> > mm, swap: support swapoff of an xswap physical backend
> > mm, swap: reclaim physical slots backing cache-only xswap entries
> > mm, swap: back a large xswap folio with a contiguous physical run
> > mm, swap: enable THP swapin for xswap entries
> > mm, swap: drop swap_folio_sector()
> > mm, swap: do not retake the cluster lock when uncharging an xswap slot
> > mm, swap: drop a refused xswap backend run directly
> >
> > Nhat Pham (5):
> > mm, swap: prepare the swap IO path for xswap backends
> > mm, swap: split the swap memcg charge helpers
> > mm, swap: do not charge zswap-backed xswap entries
> > mm, swap: charge an xswap entry when it gets physical backing
> > mm, swap: don't gate xswap on the physical swap free count
> >
> > .../admin-guide/cgroup-v1/memcg_test.rst | 2 +-
> > include/linux/memcontrol.h | 6 +
> > include/linux/swap.h | 69 +-
> > include/linux/swap_ops.h | 9 +-
> > mm/memcontrol-v1.c | 10 +-
> > mm/memcontrol.c | 147 ++--
> > mm/memory.c | 7 +-
> > mm/page_io.c | 95 ++-
> > mm/swap.h | 35 +-
> > mm/swap_state.c | 12 +-
> > mm/swap_table.h | 53 ++
> > mm/swapfile.c | 742 ++++++++++++++++--
> > mm/zswap.c | 20 +-
> > 13 files changed, 1026 insertions(+), 181 deletions(-)
> >
> > --
> > 2.54.0
> >
> >
^ permalink raw reply [flat|nested] 25+ messages in thread
end of thread, other threads:[~2026-09-28 3:31 UTC | newest]
Thread overview: 25+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
2026-09-20 7:20 ` [RFC PATCH 01/17] mm, swap: prepare the swap IO path for xswap backends Baoquan He
2026-09-20 7:20 ` [RFC PATCH 02/17] mm, swap: tag a swap table entry with its owning xswap entry Baoquan He
2026-09-20 7:20 ` [RFC PATCH 03/17] mm, swap: prepare the folio-less allocation path for xswap Baoquan He
2026-09-20 7:20 ` [RFC PATCH 04/17] mm, swap: add a physical backend for xswap slots Baoquan He
2026-09-20 7:20 ` [RFC PATCH 05/17] mm, swap: use the xswap physical backend Baoquan He
2026-09-20 7:20 ` [RFC PATCH 06/17] mm, swap: fall back to disk when zswap refuses an xswap page Baoquan He
2026-09-20 7:20 ` [RFC PATCH 07/17] mm, swap: support swapoff of an xswap physical backend Baoquan He
2026-09-20 7:20 ` [RFC PATCH 08/17] mm, swap: reclaim physical slots backing cache-only xswap entries Baoquan He
2026-09-20 7:20 ` [RFC PATCH 09/17] mm, swap: back a large xswap folio with a contiguous physical run Baoquan He
2026-09-20 7:20 ` [RFC PATCH 10/17] mm, swap: enable THP swapin for xswap entries Baoquan He
2026-09-20 7:20 ` [RFC PATCH 11/17] mm, swap: drop swap_folio_sector() Baoquan He
2026-09-20 7:20 ` [RFC PATCH 12/17] mm, swap: split the swap memcg charge helpers Baoquan He
2026-09-20 7:20 ` [RFC PATCH 13/17] mm, swap: do not charge zswap-backed xswap entries Baoquan He
2026-09-21 11:54 ` Chris Li
2026-09-20 7:20 ` [RFC PATCH 14/17] mm, swap: charge an xswap entry when it gets physical backing Baoquan He
2026-09-20 7:20 ` [RFC PATCH 15/17] mm, swap: don't gate xswap on the physical swap free count Baoquan He
2026-09-20 7:20 ` [RFC PATCH 16/17] mm, swap: do not retake the cluster lock when uncharging an xswap slot Baoquan He
2026-09-20 7:20 ` [RFC PATCH 17/17] mm, swap: drop a refused xswap backend run directly Baoquan He
2026-09-22 14:22 ` [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Chris Li
2026-09-22 14:30 ` Johannes Weiner
2026-09-24 1:56 ` KunWu Chan
2026-09-24 11:36 ` Chris Li
2026-09-24 12:17 ` Klara Modin
2026-09-28 3:30 ` Baoquan He
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox