* [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II)
@ 2026-10-03 0:58 Baoquan He
2026-10-03 0:58 ` [PATCH 01/12] mm, swap: prepare the swap IO path for xswap backends Baoquan He
` (12 more replies)
0 siblings, 13 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:58 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
An xswap entry holds a page in zswap. That works until zswap will not
take the page. Then the page has nowhere to go.
This series gives an xswap entry a second home. When zswap refuses a page,
the entry takes a slot on a real swap device and the data is written
there. The xswap entry does not change when the data moves. It is what the
process's PTE names, and the physical slot is recorded behind it.
Phase II of three. Phase I is the device itself, 14 patches.
Design notes
------------
An xswap slot and a physical slot are different things with different
owners, and the physical side has to find its xswap entry again. A swap
table entry already carries its type in the low three bits, and 0b100 was
unused. It now points at the owner, with the xswap type and offset above.
The places that hand out or cache a swap table entry learn to skip a
pointer entry.
folio->swap no longer says where the IO goes: a folio whose data is on the
backend still sits in the swap cache of its own xswap entry. So the
batched IO path carries the entry explicitly, instead of deriving the
device and sector from folio->swap. A large folio is read back with one
IO starting at its first slot, so its slots have to be contiguous; the
reserved-run allocation provides that.
An xswap entry's swap count can reach zero while its folio is still in
the swap cache. A page faulted back in during its write to the backend
ends up there: do_swap_page() does not wait for writeback, folio_free_swap()
refuses a folio under writeback, and a write to a real device does not drop
the cache when it completes. The physical slot is then redundant, but still
held. A single bit in the slot's swap table entry marks this, and the
physical reclaim scanner frees it. The scanner reads the bit directly;
taking the xswap cluster lock would invert the lock order.
An xswap entry that is still in zswap costs no swap space, so it is not
charged. Charging is phase III. Here the entry is only recorded.
Testing
-------
qemu KVM guest, 8G RAM. memhog is a small local helper: it faults
<total_gb> of anon, fills it with a fixed pattern, and holds it.
A 4G swap disk /dev/vdb is the backend. The xswap device has to exist
before zswap is turned off, because creating one requires zswap. The
number written to create is the device's swap priority; 100 puts it above
the backend at 0.
# echo 1 > /sys/module/zswap/parameters/enabled
# echo 100 > /sys/kernel/mm/xswap/create
# mkswap /dev/vdb && swapon -p 0 /dev/vdb
# echo 0 > /sys/module/zswap/parameters/enabled
# mkdir -p /sys/fs/cgroup/xswap_limit
# echo 4G > /sys/fs/cgroup/xswap_limit/memory.max
# echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
# ( echo $BASHPID > /sys/fs/cgroup/xswap_limit/cgroup.procs
# exec env MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 \
# ./memhog 5 600 ) &
The pages reach the backend, and the xswap side accounts for them:
# awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
Filename Type Size Used Priority
xswap0 xswap 8155132 2392176 100
/dev/vdb partition 4194300 2392176 0
The two Used values are equal. An xswap slot and its backend slot each
hold one reference to the same page while it is swapped out. It is one
page of data, not two, and the backend slot is given back when the xswap
entry goes.
With the workload still alive, destroying the device returns every
backend slot:
# echo 0 > /sys/kernel/mm/xswap/destroy
# sleep 5; awk 'NR == 1 || $1 ~ /vdb/' /proc/swaps
Filename Type Size Used Priority
/dev/vdb partition 4194300 0 0
Also tested these, and nothing crashed or warned:
- ten create/fill/destroy cycles, each returning /dev/vdb to 0
- a shrink racing a swapoff of the same device, 20 rounds
- a swapoff of a device that has pages on the backend
- THP swapin: a large folio read back as one IO, and byte for byte
what was written
Changelog
=========
RFC -> v1:
- Charging is phase III, so its six patches are out. An entry is
recorded here and not charged, and the RFC's two charge tests go with
them.
- New patch 12: skip a NOFS backend when reclaim cannot enter the fs.
may_enter_fs() sees the xswap device, not the backend.
- Cache-only reclaim marks the slot with SWP_RMAP_CACHE_ONLY instead of
handing it to __try_to_reclaim_swap().
- A large folio is only written to a block device backend, and the
backend is recorded before the reverse mapping that finds it.
- An xswap entry goes back on the zswap writeback LRU.
- swapoff reads an entry back under memalloc_noreclaim_save() and
xswap_lock.
- Patches 1 to 11 keep their subjects and order, rebased onto the new
phase I, with rewritten commit messages.
Baoquan He (11):
mm, swap: tag a swap table entry with its owning xswap entry
mm, swap: prepare the folio-less allocation path for xswap
mm, swap: add a physical backend for xswap slots
mm, swap: use the xswap physical backend
mm, swap: fall back to disk when zswap refuses an xswap page
mm, swap: support swapoff of an xswap physical backend
mm, swap: reclaim physical slots backing cache-only xswap entries
mm, swap: back a large xswap folio with a contiguous physical run
mm, swap: enable THP swapin for xswap entries
mm, swap: drop swap_folio_sector()
mm, swap: skip a NOFS backend when reclaim cannot enter the fs
Nhat Pham (1):
mm, swap: prepare the swap IO path for xswap backends
include/linux/swap.h | 2 +-
include/linux/swap_ops.h | 10 +-
include/linux/zswap.h | 4 +-
mm/memory.c | 7 +-
mm/page_io.c | 102 +++--
mm/swap.h | 37 +-
mm/swap_state.c | 12 +-
mm/swap_table.h | 92 ++++-
mm/swapfile.c | 783 ++++++++++++++++++++++++++++++++++++---
mm/vmscan.c | 4 +-
mm/zswap.c | 61 ++-
11 files changed, 989 insertions(+), 125 deletions(-)
base-commit: 128a882559e489eea535bf31b7293058708d351b
--
2.54.0
^ permalink raw reply [flat|nested] 16+ messages in thread
* [PATCH 01/12] mm, swap: prepare the swap IO path for xswap backends
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
@ 2026-10-03 0:58 ` Baoquan He
2026-10-03 0:58 ` [PATCH 02/12] mm, swap: tag a swap table entry with its owning xswap entry Baoquan He
` (11 subsequent siblings)
12 siblings, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:58 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
From: Nhat Pham <nphamcs@gmail.com>
An xswap folio is written to a physical device, but stays in the swap
cache of its own xswap entry, so folio->swap no longer says where the IO
goes.
The batched swap IO path takes the bio sector from folio->swap, which for
such a folio resolves the xswap device. That device has no bdev and no
extents, so the IO is set up wrong.
Pass the entry explicitly instead: __swap_writeout(), swap_add_folio()
and swap_ops->can_merge() take it as a parameter, and swap_iocb records
the first entry of each batch. The merge check, the bio sector and the
completion path all use it.
Signed-off-by: Nhat Pham <nphamcs@gmail.com>
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
include/linux/swap.h | 1 +
include/linux/swap_ops.h | 9 ++++---
mm/page_io.c | 55 ++++++++++++++++++++--------------------
mm/swap.h | 3 ++-
mm/swapfile.c | 13 ++++++++++
mm/zswap.c | 2 +-
6 files changed, 49 insertions(+), 34 deletions(-)
diff --git a/include/linux/swap.h b/include/linux/swap.h
index ae2e49386443..2080540c6e39 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -418,6 +418,7 @@ extern bool swap_entry_swapped(struct swap_info_struct *si, swp_entry_t entry);
extern int swp_swapcount(swp_entry_t entry);
extern struct swap_info_struct *get_swap_device(swp_entry_t entry);
sector_t swap_folio_sector(struct folio *folio);
+sector_t swap_entry_sector(swp_entry_t entry);
/*
* If there is an existing swap slot reference (swap entry) and the caller
diff --git a/include/linux/swap_ops.h b/include/linux/swap_ops.h
index 57ac6c703f68..198c5738c3c7 100644
--- a/include/linux/swap_ops.h
+++ b/include/linux/swap_ops.h
@@ -10,6 +10,7 @@ struct swap_iocb {
struct bio bio;
};
struct bio_vec bvecs[SWAP_CLUSTER_MAX];
+ swp_entry_t entry; /* First entry of the batch */
int nr_bvecs;
int len;
};
@@ -30,15 +31,15 @@ struct swap_io_ctx {
struct swap_ops {
unsigned int flags;
- bool (*can_merge)(struct folio *folio, struct folio *prev_folio,
- size_t prev_folio_size, int rw);
+ bool (*can_merge)(struct folio *folio, swp_entry_t entry,
+ struct swap_iocb *sio, int rw);
void (*submit_write)(struct swap_io_ctx *ctx);
void (*submit_read)(struct swap_io_ctx *ctx);
};
void swap_fs_prepare_rw(struct swap_io_ctx *ctx, int rw, struct iov_iter *iter);
-bool swap_fs_can_merge(struct folio *folio, struct folio *prev_folio,
- size_t prev_folio_size, int rw);
+bool swap_fs_can_merge(struct folio *folio, swp_entry_t entry,
+ struct swap_iocb *sio, int rw);
int swap_fs_activate(struct swap_info_struct *sis, const struct swap_ops *ops);
#endif /* _MM_SWAP_OPS_H */
diff --git a/mm/page_io.c b/mm/page_io.c
index 25fa9b82ed46..16ae84e6b785 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -257,7 +257,7 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
return AOP_WRITEPAGE_ACTIVATE;
}
- __swap_writeout(ctx, folio);
+ __swap_writeout(ctx, folio, folio->swap);
return 0;
out_unlock:
folio_unlock(folio);
@@ -332,24 +332,22 @@ int sio_pool_init(void)
}
static bool swap_can_merge(struct swap_io_ctx *ctx, struct folio *folio,
- int rw)
+ swp_entry_t entry, int rw)
{
- struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
- struct bio_vec *last_bv = &ctx->sio->bvecs[ctx->sio->nr_bvecs - 1];
- struct folio *prev_folio = bvec_folio(last_bv);
- size_t prev_folio_size = folio_size(prev_folio);
+ struct swap_info_struct *sis = __swap_entry_to_info(entry);
if (ctx->sis != sis)
return false;
- return sis->ops->can_merge(folio, prev_folio, prev_folio_size, rw);
+ return sis->ops->can_merge(folio, entry, ctx->sio, rw);
}
-static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, int rw)
+static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio,
+ swp_entry_t entry, int rw)
{
- struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
+ struct swap_info_struct *sis = __swap_entry_to_info(entry);
struct swap_iocb *sio = ctx->sio;
- if (sio && !swap_can_merge(ctx, folio, rw)) {
+ if (sio && !swap_can_merge(ctx, folio, entry, rw)) {
if (rw == WRITE)
swap_write_submit(ctx);
else
@@ -362,6 +360,7 @@ static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, int rw)
ctx->sio = sio = mempool_alloc(sio_pool, GFP_NOIO);
sio->nr_bvecs = 0;
sio->len = 0;
+ sio->entry = entry;
}
bvec_set_folio(&sio->bvecs[sio->nr_bvecs], folio, folio_size(folio), 0);
sio->len += folio_size(folio);
@@ -382,7 +381,8 @@ static void swap_add_folio(struct swap_io_ctx *ctx, struct folio *folio, int rw)
}
}
-void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
+void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio,
+ swp_entry_t entry)
{
VM_BUG_ON_FOLIO(!folio_test_swapcache(folio), folio);
@@ -398,7 +398,7 @@ void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
folio_start_writeback(folio);
folio_unlock(folio);
- swap_add_folio(ctx, folio, WRITE);
+ swap_add_folio(ctx, folio, entry, WRITE);
}
/*
@@ -507,7 +507,7 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct folio *folio)
/* We have to read from slower devices. Increase zswap protection. */
zswap_folio_swapin(folio);
- swap_add_folio(ctx, folio, READ);
+ swap_add_folio(ctx, folio, folio->swap, READ);
finish:
if (workingset) {
@@ -539,8 +539,6 @@ static void swap_fs_write_complete(struct kiocb *iocb, long ret)
bool failed = ret != sio->len;
if (failed) {
- struct folio *folio = bvec_folio(&sio->bvecs[0]);
-
/*
* In the case of swap-over-nfs, this can be a temporary failure
* if the system has limited memory for allocating transmit
@@ -548,7 +546,7 @@ static void swap_fs_write_complete(struct kiocb *iocb, long ret)
* folio_rotate_reclaimable but rate-limit the messages.
*/
pr_err_ratelimited("Write error %ld on dio swapfile (%llu)\n",
- ret, swap_dev_pos(folio->swap));
+ ret, swap_dev_pos(sio->entry));
}
swap_write_end(sio, failed);
@@ -620,7 +618,7 @@ static void swap_bdev_submit_write(struct swap_io_ctx *ctx)
bio_init(bio, ctx->sis->bdev, sio->bvecs, ARRAY_SIZE(sio->bvecs),
REQ_OP_WRITE | REQ_SWAP);
bio->bi_iter.bi_size = sio->len;
- bio->bi_iter.bi_sector = swap_folio_sector(bio_first_folio_all(bio));
+ bio->bi_iter.bi_sector = swap_entry_sector(sio->entry);
bio_associate_blkg_from_folio(bio, bio_first_folio_all(bio));
if (ctx->sis->flags & SWP_SYNCHRONOUS_IO) {
@@ -649,7 +647,7 @@ static void swap_bdev_submit_read(struct swap_io_ctx *ctx)
bio_init(bio, ctx->sis->bdev, sio->bvecs, ARRAY_SIZE(sio->bvecs),
REQ_OP_READ);
bio->bi_iter.bi_size = sio->len;
- bio->bi_iter.bi_sector = swap_folio_sector(bio_first_folio_all(bio));
+ bio->bi_iter.bi_sector = swap_entry_sector(sio->entry);
if (ctx->sis->flags & SWP_SYNCHRONOUS_IO) {
/*
@@ -667,13 +665,15 @@ static void swap_bdev_submit_read(struct swap_io_ctx *ctx)
}
}
-static bool swap_bdev_can_merge(struct folio *folio, struct folio *prev_folio,
- size_t prev_folio_size, int rw)
+static bool swap_bdev_can_merge(struct folio *folio, swp_entry_t entry,
+ struct swap_iocb *sio, int rw)
{
- if (swap_folio_sector(folio) !=
- swap_folio_sector(prev_folio) + (prev_folio_size >> SECTOR_SHIFT))
+ if (swap_entry_sector(entry) !=
+ swap_entry_sector(sio->entry) + (sio->len >> SECTOR_SHIFT))
return false;
- if (rw == WRITE && !folio_blkg_can_merge(folio, prev_folio))
+ if (rw == WRITE &&
+ !folio_blkg_can_merge(folio,
+ bvec_folio(&sio->bvecs[sio->nr_bvecs - 1])))
return false;
return true;
}
@@ -689,7 +689,7 @@ void swap_fs_prepare_rw(struct swap_io_ctx *ctx, int rw, struct iov_iter *iter)
struct swap_iocb *sio = ctx->sio;
init_sync_kiocb(&sio->iocb, ctx->sis->swap_file);
- sio->iocb.ki_pos = swap_dev_pos(bvec_folio(&sio->bvecs[0])->swap);
+ sio->iocb.ki_pos = swap_dev_pos(sio->entry);
if (rw == WRITE)
sio->iocb.ki_complete = swap_fs_write_complete;
else
@@ -700,11 +700,10 @@ void swap_fs_prepare_rw(struct swap_io_ctx *ctx, int rw, struct iov_iter *iter)
}
EXPORT_SYMBOL_GPL(swap_fs_prepare_rw);
-bool swap_fs_can_merge(struct folio *folio, struct folio *prev_folio,
- size_t prev_folio_size, int rw)
+bool swap_fs_can_merge(struct folio *folio, swp_entry_t entry,
+ struct swap_iocb *sio, int rw)
{
- return swap_dev_pos(folio->swap) ==
- swap_dev_pos(prev_folio->swap) + prev_folio_size;
+ return swap_dev_pos(entry) == swap_dev_pos(sio->entry) + sio->len;
}
EXPORT_SYMBOL_GPL(swap_fs_can_merge);
diff --git a/mm/swap.h b/mm/swap.h
index d5bf21f517dc..fcde284b16b4 100644
--- a/mm/swap.h
+++ b/mm/swap.h
@@ -258,7 +258,8 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct folio *folio);
void swap_read_submit(struct swap_io_ctx *ctx);
void swap_write_submit(struct swap_io_ctx *ctx);
int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio);
-void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio);
+void __swap_writeout(struct swap_io_ctx *ctx, struct folio *folio,
+ swp_entry_t entry);
/* linux/mm/swap_state.c */
extern struct address_space swap_space __read_mostly;
diff --git a/mm/swapfile.c b/mm/swapfile.c
index bf95663c2593..0d1288d36452 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -540,6 +540,19 @@ sector_t swap_folio_sector(struct folio *folio)
return sector << (PAGE_SHIFT - 9);
}
+sector_t swap_entry_sector(swp_entry_t entry)
+{
+ struct swap_info_struct *sis = __swap_entry_to_info(entry);
+ struct swap_extent *se;
+ sector_t sector;
+ pgoff_t offset;
+
+ offset = swp_offset(entry);
+ se = offset_to_swap_extent(sis, offset);
+ sector = se->start_block + (offset - se->start_page);
+ return sector << (PAGE_SHIFT - 9);
+}
+
/*
* swap allocation tell device that a cluster of swap can now be discarded,
* to allow the swap device to optimize its wear-levelling.
diff --git a/mm/zswap.c b/mm/zswap.c
index 7fe25f0b157c..cdba35e0fb5a 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1087,7 +1087,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
folio_put(folio);
/* start writeback */
- __swap_writeout(&ctx, folio);
+ __swap_writeout(&ctx, folio, folio->swap);
swap_write_submit(&ctx);
return 0;
--
2.54.0
^ permalink raw reply related [flat|nested] 16+ messages in thread
* [PATCH 02/12] mm, swap: tag a swap table entry with its owning xswap entry
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
2026-10-03 0:58 ` [PATCH 01/12] mm, swap: prepare the swap IO path for xswap backends Baoquan He
@ 2026-10-03 0:58 ` Baoquan He
2026-10-03 0:58 ` [PATCH 03/12] mm, swap: prepare the folio-less allocation path for xswap Baoquan He
` (10 subsequent siblings)
12 siblings, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:58 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
An xswap folio can be written out to a physical device. The data then
lives in a slot on that device, with no folio in the swap cache. Store
the xswap entry in that slot's swap table entry, so the owner of a
physical slot can be found again.
The low three bits of a swap table entry tell its type; 0b100 is free.
Use it to mark an entry that points at its owner, with the xswap type and
offset in the bits above.
Two paths hand out or cache a slot and know only the other two types.
Make them skip a pointer entry: cluster_scan_range() must not hand it out
as free, and __swap_cache_add_check() must not put a folio over it.
Otherwise readahead on a physical device can walk into it.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swap_state.c | 6 ++++-
mm/swap_table.h | 61 +++++++++++++++++++++++++++++++++++++++++++++----
mm/swapfile.c | 3 +++
3 files changed, 65 insertions(+), 5 deletions(-)
diff --git a/mm/swap_state.c b/mm/swap_state.c
index 3a4e9e447b0b..c5c182489d6a 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -179,6 +179,9 @@ static int __swap_cache_add_check(struct swap_cluster_info *ci,
return -ENOENT;
ci_off = swp_cluster_offset(targ_entry);
old_tb = __swap_table_get(ci, ci_off);
+ /* Readahead of a physical device can hit an xswap-backing slot. */
+ if (swp_tb_is_pointer(old_tb))
+ return -ENOENT;
if (swp_tb_is_folio(old_tb))
return -EEXIST;
if (!__swp_tb_get_count(old_tb))
@@ -196,7 +199,8 @@ static int __swap_cache_add_check(struct swap_cluster_info *ci,
ci_end = ci_off + nr;
do {
old_tb = __swap_table_get(ci, ci_off);
- if (unlikely(swp_tb_is_folio(old_tb) ||
+ if (unlikely(swp_tb_is_pointer(old_tb) ||
+ swp_tb_is_folio(old_tb) ||
!__swp_tb_get_count(old_tb) ||
is_zero != __swap_table_test_zero(ci, ci_off) ||
(memcg_id && *memcg_id != __swap_cgroup_get(ci, ci_off))))
diff --git a/mm/swap_table.h b/mm/swap_table.h
index e6613e62f8d0..22322519c3fe 100644
--- a/mm/swap_table.h
+++ b/mm/swap_table.h
@@ -28,7 +28,7 @@ struct swap_memcg_table {
* NULL: |---------------- 0 ---------------| - Free slot
* Shadow: |SWAP_COUNT|Z|---- SHADOW_VAL ---|1| - Swapped out slot
* PFN: |SWAP_COUNT|Z|------ PFN -------|10| - Cached slot
- * Pointer: |----------- Pointer ----------|100| - (Unused)
+ * Pointer: |C|-- xswap type/offset --|100| - Backend slot
* Bad: |------------- 1 -------------|1000| - Bad slot
*
* COUNT is `SWP_TB_COUNT_BITS` long, Z is the `SWP_TB_ZERO_FLAG` bit,
@@ -49,9 +49,9 @@ struct swap_memcg_table {
* - PFN: Swap slot is in use, and cached. Memcg info is recorded on the page
* struct.
*
- * - Pointer: Unused yet. `0b100` is reserved for potential pointer usage
- * because only the lower three bits can be used as a marker for 8 bytes
- * aligned pointers.
+ * - Pointer: A physical swap slot backed by an xswap entry. `0b100` marks it
+ * and the xswap type and offset go above it; C is SWP_RMAP_CACHE_ONLY. See
+ * the layout in the CONFIG_XSWAP block below.
*
* - Bad: Swap slot is reserved, protects swap header or holes on swap devices.
*/
@@ -81,6 +81,59 @@ struct swap_memcg_table {
/* Bad slot: ends with 0b1000 and rests of bits are all 1 */
#define SWP_TB_BAD ((~0UL) << 3)
+#ifdef CONFIG_XSWAP
+/*
+ * Pointer-tagged swap table entry: the reverse map from a physical slot to
+ * the xswap entry it backs. Layout:
+ *
+ * Pointer: | xswap_type(8) | xswap_offset(53) |100|
+ *
+ * The low three bits identify the entry: 0b000 is a free or bad slot, a
+ * shadow entry ends in 0b01 and a cached folio in 0b10, which leaves
+ * 0b100 as the only marker still free.
+ */
+#define SWP_TB_PTR_MARK 0b100UL
+#define SWP_TB_PTR_OFF_BITS 53
+#define SWP_TB_PTR_TYPE_BITS 8
+#define SWP_TB_PTR_OFF_SHIFT 3
+#define SWP_TB_PTR_TYPE_SHIFT (SWP_TB_PTR_OFF_SHIFT + \
+ SWP_TB_PTR_OFF_BITS)
+#define SWP_TB_PTR_OFF_MASK ((1UL << SWP_TB_PTR_OFF_BITS) - 1)
+#define SWP_TB_PTR_TYPE_MASK ((1UL << SWP_TB_PTR_TYPE_BITS) - 1)
+
+static inline bool swp_tb_is_pointer(unsigned long swp_tb)
+{
+ return (swp_tb & (BIT(3) - 1)) == SWP_TB_PTR_MARK;
+}
+
+static inline unsigned long xswap_entry_to_rmap(swp_entry_t entry)
+{
+ unsigned long type = swp_type(entry);
+ unsigned long off = swp_offset(entry);
+
+ VM_WARN_ON_ONCE(type > SWP_TB_PTR_TYPE_MASK);
+ VM_WARN_ON_ONCE(off > SWP_TB_PTR_OFF_MASK);
+ return (type << SWP_TB_PTR_TYPE_SHIFT) |
+ (off << SWP_TB_PTR_OFF_SHIFT) |
+ SWP_TB_PTR_MARK;
+}
+
+static inline swp_entry_t xswap_rmap_to_entry(unsigned long swp_tb)
+{
+ unsigned long type = (swp_tb >> SWP_TB_PTR_TYPE_SHIFT) &
+ SWP_TB_PTR_TYPE_MASK;
+ unsigned long off = (swp_tb >> SWP_TB_PTR_OFF_SHIFT) &
+ SWP_TB_PTR_OFF_MASK;
+
+ return swp_entry(type, off);
+}
+#else /* !CONFIG_XSWAP */
+static inline bool swp_tb_is_pointer(unsigned long swp_tb)
+{
+ return false;
+}
+#endif /* CONFIG_XSWAP */
+
/* Macro for shadow offset calculation */
#define SWAP_COUNT_SHIFT SWP_TB_FLAGS_BITS
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 0d1288d36452..af7a7524fc4d 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -1120,6 +1120,9 @@ static bool cluster_scan_range(struct swap_info_struct *si,
swp_tb = __swap_table_get(ci, ci_off);
if (swp_tb_is_null(swp_tb))
continue;
+ /* Backs an xswap entry: the slot is not allocatable. */
+ if (swp_tb_is_pointer(swp_tb))
+ return false;
if (swp_tb_is_folio(swp_tb) && !__swp_tb_get_count(swp_tb)) {
if (!vm_swap_full())
return false;
--
2.54.0
^ permalink raw reply related [flat|nested] 16+ messages in thread
* [PATCH 03/12] mm, swap: prepare the folio-less allocation path for xswap
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
2026-10-03 0:58 ` [PATCH 01/12] mm, swap: prepare the swap IO path for xswap backends Baoquan He
2026-10-03 0:58 ` [PATCH 02/12] mm, swap: tag a swap table entry with its owning xswap entry Baoquan He
@ 2026-10-03 0:58 ` Baoquan He
2026-10-03 0:58 ` [PATCH 04/12] mm, swap: add a physical backend for xswap slots Baoquan He
` (9 subsequent siblings)
12 siblings, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:58 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
The xswap backend writes data to a slot on a real swap device. That slot
has no folio of its own, a case only hibernation used before, so the NULL
folio path is no longer limited to hibernation.
The backend takes its slots from a lower-priority device. The cluster it
fills must not go into this CPU's swapout cache: swap_alloc_fast() reuses
that cache without checking the device, so one backend allocation would
send every later swapout there and xswap would never be picked again.
A folio-less slot packs into a cluster already in use instead of taking a
free cluster. Only a free cluster can serve a large order, so the free
ones are kept for large folios.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 35 +++++++++++++++++++++++++----------
1 file changed, 25 insertions(+), 10 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index af7a7524fc4d..dd18aad7cfd9 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -1156,8 +1156,10 @@ static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
* Such swap slots starts with count == 0 and will be increased
* upon folio unmap.
*
- * Else, it's a exclusive order 0 allocation for hibernation.
- * The slot starts with count == 1 and never increases.
+ * Else, it's an exclusive order 0 allocation for a slot that has
+ * no folio of its own: hibernation, and the physical slot that
+ * backs an xswap entry. The slot starts with count == 1 and
+ * never increases.
*/
if (likely(folio)) {
order = folio_order(folio);
@@ -1165,16 +1167,12 @@ static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
swap_cluster_assert_empty(ci, ci_off, nr_pages, false);
__swap_cache_add_folio(ci, folio, swp_entry(si->type,
ci_off + cluster_offset(si, ci)));
- } else if (IS_ENABLED(CONFIG_HIBERNATION)) {
+ } else {
order = 0;
nr_pages = 1;
swap_cluster_assert_empty(ci, ci_off, 1, false);
- /* Fake shadow placeholder with no flag, hibernation does not use the zeromap */
+ /* Fake shadow placeholder with no flag; no zeromap is used. */
__swap_table_set(ci, ci_off, __swp_tb_mk_count(shadow_to_swp_tb(NULL, 0), 1));
- } else {
- /* Allocation without folio is only possible with hibernation */
- WARN_ON_ONCE(1);
- return false;
}
/*
@@ -1238,8 +1236,15 @@ static unsigned long alloc_swap_scan_cluster(struct swap_info_struct *si,
swap_cluster_unlock(ci);
rcu_read_unlock();
if (si->flags & SWP_SOLIDSTATE) {
- this_cpu_write(percpu_swap_cluster.offset[order], next);
- this_cpu_write(percpu_swap_cluster.si[order], si);
+ /*
+ * A NULL folio is a backend or hibernation allocation: it must
+ * not displace the swapout cluster cached here, which
+ * swap_alloc_fast() reuses without picking a device again.
+ */
+ if (folio) {
+ this_cpu_write(percpu_swap_cluster.offset[order], next);
+ this_cpu_write(percpu_swap_cluster.si[order], si);
+ }
} else {
si->global_cluster->next[order] = next;
}
@@ -1379,6 +1384,16 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
}
new_cluster:
+ /*
+ * A slot with no folio packs into a cluster already in use: a free
+ * cluster is the only thing a large order can allocate from.
+ */
+ if (!folio && order < PMD_ORDER) {
+ found = alloc_swap_scan_list(si, &si->frag_clusters[order], folio, false);
+ if (found)
+ goto done;
+ }
+
/*
* If the device need discard, prefer new cluster over nonfull
* to spread out the writes.
--
2.54.0
^ permalink raw reply related [flat|nested] 16+ messages in thread
* [PATCH 04/12] mm, swap: add a physical backend for xswap slots
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
` (2 preceding siblings ...)
2026-10-03 0:58 ` [PATCH 03/12] mm, swap: prepare the folio-less allocation path for xswap Baoquan He
@ 2026-10-03 0:58 ` Baoquan He
2026-10-03 0:58 ` [PATCH 05/12] mm, swap: use the xswap physical backend Baoquan He
` (8 subsequent siblings)
12 siblings, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:58 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
An xswap slot holds its data only in the zswap pool. There is no
second copy anywhere, so once the pool drops the entry the data is
gone. Give xswap slots a physical backend they can be written out to:
a slot on the highest-priority real swap device.
The backend is recorded per slot in ci->xs_table[], an array of
SWAPFILE_CLUSTER unsigned longs that is allocated on first use. If a
cluster only ever holds zswap-only slots, xs_table[] is never
allocated. Zero means "no backend".
Freeing an xswap slot releases its physical slot. The release checks
that the physical slot still belongs to this xswap entry. If it belongs
to another one, the record is stale. That slot was freed and handed out
again. Then the slot may by belong to another entry, whose data must not
be dropped.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swap.h | 24 ++++
mm/swapfile.c | 317 +++++++++++++++++++++++++++++++++++++++++++++++++-
2 files changed, 339 insertions(+), 2 deletions(-)
diff --git a/mm/swap.h b/mm/swap.h
index fcde284b16b4..838dac4282f6 100644
--- a/mm/swap.h
+++ b/mm/swap.h
@@ -58,6 +58,9 @@ struct swap_cluster_info {
u8 order;
atomic_long_t __rcu *table; /* Swap table entries, see mm/swap_table.h */
unsigned int *extend_table; /* For large swap count, protected by ci->lock */
+#ifdef CONFIG_XSWAP
+ unsigned long *xs_table; /* Physical backend per slot, lazily allocated */
+#endif
#ifdef CONFIG_MEMCG
struct swap_memcg_table *memcg_table; /* Swap table entries' cgroup record */
#endif
@@ -470,4 +473,25 @@ extern const struct swap_ops swap_bdev_ops;
int shmem_writeout(struct swap_io_ctx *ctx, struct folio *folio,
struct list_head *folio_list);
+#ifdef CONFIG_XSWAP
+swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci, unsigned int slot);
+swp_entry_t xswap_backend_alloc(swp_entry_t entry);
+void xswap_backend_free(swp_entry_t entry, swp_entry_t phys);
+#else
+static inline swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci,
+ unsigned int slot)
+{
+ return (swp_entry_t){};
+}
+
+static inline swp_entry_t xswap_backend_alloc(swp_entry_t entry)
+{
+ return (swp_entry_t){};
+}
+
+static inline void xswap_backend_free(swp_entry_t entry, swp_entry_t phys)
+{
+}
+#endif /* CONFIG_XSWAP */
+
#endif /* _MM_SWAP_H */
diff --git a/mm/swapfile.c b/mm/swapfile.c
index dd18aad7cfd9..20a9192c39de 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -101,6 +101,13 @@ static void xswap_unmap_clusters(struct swap_info_struct *si,
unsigned long start_idx, unsigned long nr);
static int xswap_mapped_end(pte_t *pte, unsigned long addr, void *data);
static void xswap_try_shrink(struct swap_info_struct *si);
+static void xswap_free_phys_slot(struct swap_info_struct *si,
+ swp_entry_t entry, swp_entry_t phys);
+static void xswap_release_slot_backend(struct swap_info_struct *si,
+ struct swap_cluster_info *ci,
+ unsigned int slot);
+static void xswap_free_xs_table(struct swap_info_struct *si,
+ struct swap_cluster_info *ci);
static int xswap_create(int prio);
static int xswap_destroy(int type);
@@ -826,6 +833,14 @@ static void swap_cluster_schedule_discard(struct swap_info_struct *si,
static void __free_cluster(struct swap_info_struct *si, struct swap_cluster_info *ci)
{
swap_cluster_assert_empty(ci, 0, SWAPFILE_CLUSTER, false);
+#ifdef CONFIG_XSWAP
+ /*
+ * An empty cluster can still hold records. Give them back first:
+ * swap_cluster_free_table() frees the table the owner ID comes from.
+ */
+ if (si->flags & SWP_XSWAP)
+ xswap_free_xs_table(si, ci);
+#endif
swap_cluster_free_table(ci);
move_cluster(si, ci, &si->free_clusters, CLUSTER_FLAG_FREE);
ci->order = 0;
@@ -1072,6 +1087,12 @@ static bool cluster_reclaim_range(struct swap_info_struct *si,
spin_unlock(&ci->lock);
do {
swp_tb = swap_table_get(ci, offset % SWAPFILE_CLUSTER);
+ /*
+ * A pointer-tagged slot backs an xswap entry, not a cached
+ * folio, so it can't be reclaimed here. Reject the range.
+ */
+ if (swp_tb_is_pointer(swp_tb))
+ break;
if (swp_tb_get_count(swp_tb))
break;
if (swp_tb_is_folio(swp_tb))
@@ -1836,6 +1857,14 @@ static void __swap_cluster_put_entry(struct swap_cluster_info *ci,
lockdep_assert_held(&ci->lock);
swp_tb = __swap_table_get(ci, ci_off);
+ /*
+ * A pointer-tagged slot has no count of its own. Putting one would
+ * corrupt the reverse mapping, so never touch it here.
+ */
+ if (swp_tb_is_pointer(swp_tb)) {
+ VM_WARN_ON_ONCE(1);
+ return;
+ }
count = __swp_tb_get_count(swp_tb);
VM_WARN_ON_ONCE(count <= 0);
@@ -1944,8 +1973,8 @@ static int __swap_cluster_dup_entry(struct swap_cluster_info *ci,
lockdep_assert_held(&ci->lock);
swp_tb = __swap_table_get(ci, ci_off);
- /* Bad or special slots can't be handled */
- if (WARN_ON_ONCE(swp_tb_is_bad(swp_tb)))
+ /* Bad, pointer-tagged and other special slots can't be handled */
+ if (WARN_ON_ONCE(swp_tb_is_bad(swp_tb) || swp_tb_is_pointer(swp_tb)))
return -EINVAL;
count = __swp_tb_get_count(swp_tb);
/* Must be either cached or have a count already */
@@ -2257,6 +2286,11 @@ void __swap_cluster_free_entries(struct swap_info_struct *si,
*/
VM_WARN_ON(!swp_tb_is_shadow(old_tb) || __swp_tb_get_count(old_tb) > 1);
+#ifdef CONFIG_XSWAP
+ if (ci->xs_table && ci->xs_table[ci_off])
+ xswap_release_slot_backend(si, ci, ci_off);
+#endif
+
/* Resetting the slot to NULL also clears the inline flags. */
__swap_table_set(ci, ci_off, null_to_swp_tb());
if (!SWAP_TABLE_HAS_ZEROFLAG)
@@ -3164,6 +3198,71 @@ static unsigned long find_next_to_unuse(struct swap_info_struct *si,
return 0;
}
+
+#ifdef CONFIG_XSWAP
+/*
+ * Free the physical slot @phys and clear its reverse mapping. @entry is the
+ * xswap entry that recorded @phys. The slot is only freed while it still
+ * says it belongs to @entry: our record can go stale, and @phys may by then
+ * have been handed to another xswap entry, whose data must not be dropped.
+ *
+ * free_cluster()/partial_free_cluster() require the cluster lock.
+ */
+static void xswap_free_phys_slot(struct swap_info_struct *si,
+ swp_entry_t entry, swp_entry_t phys)
+{
+ struct swap_cluster_info *ci;
+ unsigned long swp_tb;
+ unsigned int offset = swp_offset(phys);
+
+ ci = swap_cluster_lock(si, offset);
+ if (!ci)
+ return;
+
+ swp_tb = __swap_table_get(ci, offset % SWAPFILE_CLUSTER);
+ if (!swp_tb_is_pointer(swp_tb) ||
+ xswap_rmap_to_entry(swp_tb).val != entry.val) {
+ /* Already freed, or recycled for another xswap entry. */
+ VM_WARN_ON_ONCE(!swp_tb_is_null(swp_tb) &&
+ !swp_tb_is_pointer(swp_tb));
+ swap_cluster_unlock(ci);
+ return;
+ }
+
+ __swap_table_set(ci, offset % SWAPFILE_CLUSTER, null_to_swp_tb());
+ ci->count--;
+
+ /* The slot must not be allocatable before its invalidate hooks run. */
+ swap_range_free(si, offset, 1);
+
+ if (!ci->count)
+ free_cluster(si, ci);
+ else
+ partial_free_cluster(si, ci);
+ swap_cluster_unlock(ci);
+}
+
+/*
+ * An xswap slot being freed may still own a physical slot. Drop it, or the
+ * physical space and its reverse mapping leak. Caller holds the xswap
+ * cluster lock; the physical cluster lock is only ever taken below it.
+ */
+static void xswap_release_slot_backend(struct swap_info_struct *si,
+ struct swap_cluster_info *ci,
+ unsigned int slot)
+{
+ swp_entry_t entry = swp_entry(si->type, cluster_offset(si, ci) + slot);
+ struct swap_info_struct *psi;
+ swp_entry_t phys = { .val = ci->xs_table[slot] };
+
+ WRITE_ONCE(ci->xs_table[slot], 0);
+
+ psi = swap_type_to_info(swp_type(phys));
+ if (psi)
+ xswap_free_phys_slot(psi, entry, phys);
+}
+
+#endif /* CONFIG_XSWAP */
static int try_to_unuse(unsigned int type)
{
struct mm_struct *prev_mm;
@@ -4062,6 +4161,212 @@ static unsigned long read_swap_header(struct swap_info_struct *si,
}
#ifdef CONFIG_XSWAP
+/* The physical slot backing @slot, or zero. */
+swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci, unsigned int slot)
+{
+ unsigned long *xs = READ_ONCE(ci->xs_table);
+
+ if (!xs || !xs[slot])
+ return (swp_entry_t){};
+
+ return (swp_entry_t){ .val = xs[slot] };
+}
+
+/* Cluster lock held. xswap never unmaps a cluster below the mapped prefix. */
+static unsigned long *xswap_xs_table_locked(struct swap_cluster_info *ci,
+ gfp_t gfp)
+{
+ unsigned long *xs = ci->xs_table;
+
+ if (xs)
+ return xs;
+
+ /* The owning entry is pinned, so the cluster survives the relock. */
+ if (gfp & __GFP_DIRECT_RECLAIM) {
+ spin_unlock(&ci->lock);
+ xs = kcalloc(SWAPFILE_CLUSTER, sizeof(*xs), gfp);
+ spin_lock(&ci->lock);
+ } else {
+ xs = kcalloc(SWAPFILE_CLUSTER, sizeof(*xs), gfp);
+ }
+
+ if (!xs)
+ return NULL;
+ if (cmpxchg(&ci->xs_table, NULL, xs)) {
+ kfree(xs);
+ xs = ci->xs_table;
+ }
+ return xs;
+}
+
+/*
+ * 4K per cluster, and this runs in reclaim, so it has to be able to sleep.
+ */
+static bool xswap_slot_set_backend(struct swap_cluster_info *ci,
+ unsigned int slot, swp_entry_t phys)
+{
+ unsigned long *xs = xswap_xs_table_locked(ci, GFP_KERNEL | __GFP_HIGH |
+ __GFP_NOMEMALLOC);
+
+ if (!xs)
+ return false;
+ WRITE_ONCE(xs[slot], phys.val);
+ return true;
+}
+
+/*
+ * Drop the table of cluster @idx and every backend it still records. The
+ * cluster is going away, so nothing else will: the physical slots would stay
+ * allocated and their reverse mappings would keep pointing at an xswap entry
+ * whose cluster_info is about to be unmapped. Caller holds ci->lock.
+ */
+static void xswap_free_xs_table(struct swap_info_struct *si,
+ struct swap_cluster_info *ci)
+{
+ unsigned long *xs = ci->xs_table;
+ unsigned int slot;
+
+ if (!xs)
+ return;
+
+ for (slot = 0; slot < SWAPFILE_CLUSTER; slot++) {
+ if (xs[slot])
+ xswap_release_slot_backend(si, ci, slot);
+ }
+
+ kfree(xs);
+ ci->xs_table = NULL;
+}
+
+/* Order-0 slot from @si, no folio attached; goes through the per-CPU cache. */
+static unsigned long xswap_alloc_slot_local(struct swap_info_struct *si)
+{
+ struct swap_cluster_info *ci;
+ unsigned long pcp_offset, offset = SWAP_ENTRY_INVALID;
+
+ local_lock(&percpu_swap_cluster.lock);
+ if (this_cpu_read(percpu_swap_cluster.si[0]) == si) {
+ pcp_offset = this_cpu_read(percpu_swap_cluster.offset[0]);
+ if (pcp_offset) {
+ ci = swap_cluster_lock(si, pcp_offset);
+ if (cluster_is_usable(ci, 0))
+ offset = alloc_swap_scan_cluster(si, ci, NULL,
+ pcp_offset);
+ else
+ swap_cluster_unlock(ci);
+ }
+ }
+ if (offset == SWAP_ENTRY_INVALID)
+ offset = cluster_alloc_swap_entry(si, NULL);
+ local_unlock(&percpu_swap_cluster.lock);
+
+ return offset;
+}
+
+static swp_entry_t xswap_alloc_phys_slot(void)
+{
+ struct swap_info_struct *devs[MAX_SWAPFILES], *si;
+ int nr = 0, i;
+
+ /*
+ * Snapshot the candidate devices in priority order. swap_info
+ * structs are never freed, so the pointers stay valid; liveness is
+ * re-checked below.
+ */
+ spin_lock(&swap_avail_lock);
+ plist_for_each_entry(si, &swap_avail_head, avail_list) {
+ if (si->flags & SWP_XSWAP)
+ continue;
+ if (nr == MAX_SWAPFILES)
+ break;
+ devs[nr++] = si;
+ }
+ spin_unlock(&swap_avail_lock);
+
+ for (i = 0; i < nr; i++) {
+ unsigned long offset;
+
+ if (!get_swap_device_info(devs[i]))
+ continue;
+ offset = xswap_alloc_slot_local(devs[i]);
+ put_swap_device(devs[i]);
+ if (offset && offset != SWAP_ENTRY_INVALID)
+ return swp_entry(devs[i]->type, offset);
+ }
+
+ return (swp_entry_t){};
+}
+
+/* Reverse mapping: record @entry as the owner of physical slot @phys. */
+static void xswap_install_rmap(swp_entry_t phys, swp_entry_t entry)
+{
+ struct swap_info_struct *si = swap_type_to_info(swp_type(phys));
+ struct swap_cluster_info *ci;
+ unsigned long off = swp_offset(phys);
+
+ if (!si)
+ return;
+ ci = swap_cluster_lock(si, off);
+ if (ci) {
+ __swap_table_set(ci, off % SWAPFILE_CLUSTER,
+ xswap_entry_to_rmap(entry));
+ swap_cluster_unlock(ci);
+ }
+}
+
+swp_entry_t xswap_backend_alloc(swp_entry_t entry)
+{
+ struct swap_info_struct *si = __swap_entry_to_info(entry);
+ struct swap_cluster_info *ci;
+ unsigned long off = swp_offset(entry);
+ swp_entry_t phys;
+ bool recorded = false;
+
+ phys = xswap_alloc_phys_slot();
+ if (!phys.val)
+ return phys;
+
+ xswap_install_rmap(phys, entry);
+
+ ci = swap_cluster_lock(si, off);
+ if (ci) {
+ recorded = xswap_slot_set_backend(ci, off % SWAPFILE_CLUSTER, phys);
+ swap_cluster_unlock(ci);
+ }
+
+ if (!recorded) {
+ /* The backend is unrecorded, so the data would be unreachable. */
+ xswap_backend_free(entry, phys);
+ return (swp_entry_t){};
+ }
+
+ return phys;
+}
+
+/* Undo xswap_backend_alloc(): the zswap copy is still in place. */
+void xswap_backend_free(swp_entry_t entry, swp_entry_t phys)
+{
+ struct swap_info_struct *si = __swap_entry_to_info(entry);
+ struct swap_info_struct *psi;
+ struct swap_cluster_info *ci;
+ unsigned long off = swp_offset(entry);
+
+ ci = swap_cluster_lock(si, off);
+ if (ci) {
+ unsigned int slot = off % SWAPFILE_CLUSTER;
+ unsigned long *xs = ci->xs_table;
+
+ if (xs && xs[slot] == phys.val)
+ WRITE_ONCE(xs[slot], 0);
+ swap_cluster_unlock(ci);
+ }
+
+ /* xswap_free_phys_slot() clears the reverse mapping. */
+ psi = swap_type_to_info(swp_type(phys));
+ if (psi)
+ xswap_free_phys_slot(psi, entry, phys);
+}
+
static int xswap_map_clusters(struct swap_info_struct *si,
unsigned long start_idx, unsigned long nr)
{
@@ -4285,6 +4590,14 @@ static void xswap_unmap_clusters_locked(struct swap_info_struct *si,
unsigned int noreclaim_flags;
unsigned long npages, idx;
+ for (unsigned long idx = start_idx; idx < start_idx + nr; idx++) {
+ struct swap_cluster_info *ci = &si->cluster_info[idx];
+
+ spin_lock(&ci->lock);
+ xswap_free_xs_table(si, ci);
+ spin_unlock(&ci->lock);
+ }
+
if (vm_start >= vm_end) {
WRITE_ONCE(si->nr_clusters_mapped, start_idx);
return;
--
2.54.0
^ permalink raw reply related [flat|nested] 16+ messages in thread
* [PATCH 05/12] mm, swap: use the xswap physical backend
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
` (3 preceding siblings ...)
2026-10-03 0:58 ` [PATCH 04/12] mm, swap: add a physical backend for xswap slots Baoquan He
@ 2026-10-03 0:58 ` Baoquan He
2026-10-03 0:58 ` [PATCH 06/12] mm, swap: fall back to disk when zswap refuses an xswap page Baoquan He
` (7 subsequent siblings)
12 siblings, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:58 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
An xswap entry that zswap evicts had nowhere to go, so
zswap_writeback_entry() refused to handle it. Send it to the physical
backend instead. Reserve the backend before decompressing the entry,
so a failed allocation wastes no work.
Now that an xswap entry can be written back, put it on the zswap
writeback LRU again. The base series kept it off because there was
nothing to write it back to.
A slot that was written out no longer has a zswap copy, so
swap_read_folio() forwards its read to the physical entry rather than
dropping it.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/page_io.c | 32 ++++++++++++++++++++++++--------
mm/zswap.c | 28 +++++++++++++++++++++-------
2 files changed, 45 insertions(+), 15 deletions(-)
diff --git a/mm/page_io.c b/mm/page_io.c
index 16ae84e6b785..6dc90182f761 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -467,6 +467,7 @@ static bool swap_read_folio_zeromap(struct folio *folio)
void swap_read_folio(struct swap_io_ctx *ctx, struct folio *folio)
{
struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
+ swp_entry_t entry = folio->swap;
bool synchronous = sis->flags & SWP_SYNCHRONOUS_IO;
bool workingset = folio_test_workingset(folio);
unsigned long pflags;
@@ -496,18 +497,33 @@ void swap_read_folio(struct swap_io_ctx *ctx, struct folio *folio)
goto finish;
if (unlikely(sis->flags & SWP_XSWAP)) {
- /*
- * An xswap entry only ever lives in zswap, so zswap_load()
- * must have found it. Unlock and let the caller retry.
- */
- WARN_ON_ONCE(1);
- folio_unlock(folio);
- goto finish;
+ struct swap_cluster_info *ci;
+ swp_entry_t phys = {};
+ unsigned long offset = swp_offset(folio->swap);
+
+ /* May have been written back; reuse the physical read path. */
+ ci = __swap_offset_to_cluster(sis, offset);
+ if (ci) {
+ spin_lock(&ci->lock);
+ phys = xswap_slot_backend(ci, offset % SWAPFILE_CLUSTER);
+ spin_unlock(&ci->lock);
+ }
+ if (!phys.val) {
+ /*
+ * No folio_mark_uptodate(), so do_swap_page() sees
+ * this as a failed read and SIGBUSes silently.
+ */
+ WARN_ON_ONCE(1);
+ folio_unlock(folio);
+ goto finish;
+ }
+
+ entry = phys;
}
/* We have to read from slower devices. Increase zswap protection. */
zswap_folio_swapin(folio);
- swap_add_folio(ctx, folio, folio->swap, READ);
+ swap_add_folio(ctx, folio, entry, READ);
finish:
if (workingset) {
diff --git a/mm/zswap.c b/mm/zswap.c
index cdba35e0fb5a..e614c04697f2 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1011,6 +1011,8 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
struct mempolicy *mpol;
struct swap_info_struct *si;
struct swap_io_ctx ctx = {};
+ swp_entry_t phys = {};
+ bool is_xswap;
int ret = 0;
/* try to allocate swap cache folio */
@@ -1018,10 +1020,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
if (IS_ERR_OR_NULL(si))
return -ENOENT;
- if (si->flags & SWP_XSWAP) {
- put_swap_device(si);
- return -EINVAL;
- }
+ is_xswap = !!(si->flags & SWP_XSWAP);
mpol = get_task_policy(current);
folio = __swap_cache_alloc_folio(swpentry, GFP_KERNEL, BIT(0), NULL, mpol,
@@ -1055,6 +1054,15 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
goto err;
}
+ /* Reserve the destination before dropping the zswap copy. */
+ if (is_xswap) {
+ phys = xswap_backend_alloc(swpentry);
+ if (!phys.val) {
+ ret = -ENOMEM;
+ goto err;
+ }
+ }
+
if (!zswap_decompress(entry, folio)) {
ret = -EIO;
goto err;
@@ -1087,12 +1095,14 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
folio_put(folio);
/* start writeback */
- __swap_writeout(&ctx, folio, folio->swap);
+ __swap_writeout(&ctx, folio, is_xswap ? phys : folio->swap);
swap_write_submit(&ctx);
return 0;
err:
+ if (is_xswap && phys.val)
+ xswap_backend_free(swpentry, phys);
swap_cache_del_folio(folio);
folio_unlock(folio);
folio_put(folio);
@@ -1516,8 +1526,12 @@ static bool zswap_store_page(struct folio *folio, long index,
entry->referenced = true;
if (entry->length) {
INIT_LIST_HEAD(&entry->lru);
- /* No backing store: nothing to write these back to. */
- if (!(__swap_entry_to_info(page_swpentry)->flags & SWP_XSWAP))
+ /*
+ * An xswap entry with no real swap device cannot be written
+ * back, so keep it off the shrinker's LRU.
+ */
+ if (!(__swap_entry_to_info(page_swpentry)->flags & SWP_XSWAP) ||
+ atomic_read(&nr_real_swapfiles))
zswap_lru_add(entry);
}
--
2.54.0
^ permalink raw reply related [flat|nested] 16+ messages in thread
* [PATCH 06/12] mm, swap: fall back to disk when zswap refuses an xswap page
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
` (4 preceding siblings ...)
2026-10-03 0:58 ` [PATCH 05/12] mm, swap: use the xswap physical backend Baoquan He
@ 2026-10-03 0:58 ` Baoquan He
2026-10-03 0:58 ` [PATCH 07/12] mm, swap: support swapoff of an xswap physical backend Baoquan He
` (6 subsequent siblings)
12 siblings, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:58 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
zswap_store() can refuse a page if the pool may be at its limit, or
the page may not compress well enough. For an xswap device there was
no way to store it, so swap_writeout() put the page back on the LRU.
Under memory pressure that turns into a livelock: reclaim keeps picking
the same page and the pool keeps refusing it.
Now that xswap slots have a physical backend, write the page out
instead. If no backend can be reserved, keep the old behaviour.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/page_io.c | 14 ++++++++++++--
1 file changed, 12 insertions(+), 2 deletions(-)
diff --git a/mm/page_io.c b/mm/page_io.c
index 6dc90182f761..7a7b7eecb2a8 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -253,8 +253,18 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
* yet, so look the device up from the entry.
*/
if (unlikely(__swap_entry_to_info(folio->swap)->flags & SWP_XSWAP)) {
- folio_mark_dirty(folio);
- return AOP_WRITEPAGE_ACTIVATE;
+ swp_entry_t entry = folio->swap;
+ swp_entry_t phys;
+
+ /* The pool refused it: fall back to a real swap device. */
+ phys = xswap_backend_alloc(entry);
+ if (!phys.val) {
+ folio_mark_dirty(folio);
+ return AOP_WRITEPAGE_ACTIVATE;
+ }
+
+ __swap_writeout(ctx, folio, phys);
+ return 0;
}
__swap_writeout(ctx, folio, folio->swap);
--
2.54.0
^ permalink raw reply related [flat|nested] 16+ messages in thread
* [PATCH 07/12] mm, swap: support swapoff of an xswap physical backend
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
` (5 preceding siblings ...)
2026-10-03 0:58 ` [PATCH 06/12] mm, swap: fall back to disk when zswap refuses an xswap page Baoquan He
@ 2026-10-03 0:58 ` Baoquan He
2026-10-03 0:58 ` [PATCH 08/12] mm, swap: reclaim physical slots backing cache-only xswap entries Baoquan He
` (5 subsequent siblings)
12 siblings, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:58 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
swapoff of a real swap device iterates its slots through try_to_unuse()
and reads each one back into memory. A slot that was written back
from an xswap entry has no folio in the swap cache, so the loop
skips it and the slot is freed without ever reading the data. While
the xswap entry will keep pointing at a slot that no longer holds
anything.
Now change try_to_unuse() to analyze such a slot through its reverse
mapping: bring the folio back, mark it dirty so that reclaim stores it
into zswap again, and only then drop the backend.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swapfile.c | 147 +++++++++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 146 insertions(+), 1 deletion(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 20a9192c39de..45a871eb7bc2 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -44,6 +44,7 @@
#include <linux/suspend.h>
#include <linux/zswap.h>
#include <linux/plist.h>
+#include <linux/swap_ops.h>
#include <asm/tlbflush.h>
#include <linux/leafops.h>
@@ -3262,6 +3263,138 @@ static void xswap_release_slot_backend(struct swap_info_struct *si,
xswap_free_phys_slot(psi, entry, phys);
}
+/*
+ * A written-back slot has no folio of its own here, so the unuse loop would
+ * skip it. Resolve it through the reverse mapping and free both sides.
+ */
+static int xswap_unuse_rmap(struct swap_info_struct *si, unsigned long offset)
+{
+ struct swap_cluster_info *pci, *xci;
+ struct swap_info_struct *xsi;
+ struct swap_io_ctx ctx = {};
+ struct mempolicy *mpol;
+ struct folio *folio;
+ unsigned int noreclaim_flags;
+ unsigned long swp_tb, xoff;
+ swp_entry_t xentry, phys;
+ bool fresh = false;
+ bool drop = false;
+ int ret = 0;
+
+ phys = swp_entry(si->type, offset);
+
+ pci = swap_cluster_lock(si, offset);
+ if (!pci)
+ return 0;
+ /*
+ * The cluster can have emptied since find_next_to_unuse() read its
+ * count without the lock, and freeing a cluster clears its table.
+ */
+ if (!cluster_table_is_alloced(pci)) {
+ swap_cluster_unlock(pci);
+ return 0;
+ }
+ swp_tb = __swap_table_get(pci, offset % SWAPFILE_CLUSTER);
+ if (!swp_tb_is_pointer(swp_tb)) {
+ swap_cluster_unlock(pci);
+ return 0;
+ }
+ xentry = xswap_rmap_to_entry(swp_tb);
+ swap_cluster_unlock(pci);
+
+ xoff = swp_offset(xentry);
+
+ /*
+ * Pin the owner before reading its cluster_info:
+ * free_swap_cluster_info() only unmaps that after killing si->users.
+ * No owner left means the slot is dropped below.
+ */
+ xsi = swap_type_to_info(swp_type(xentry));
+ if (!xsi || !(xsi->flags & SWP_XSWAP) || !get_swap_device_info(xsi))
+ goto out_free;
+
+ /* Keeps the shrink, and a reclaim below, off xsi->cluster_info. */
+ noreclaim_flags = memalloc_noreclaim_save();
+ mutex_lock(&xsi->xswap_lock);
+
+ /* The shrink lowers this under the lock, so check it here. */
+ if (xoff >= READ_ONCE(xsi->nr_clusters_mapped) * SWAPFILE_CLUSTER) {
+ /* Owner gone, so the slot is an orphan and has to go. */
+ drop = true;
+ goto out_unlock;
+ }
+
+ /*
+ * Bring the data back before the slot goes. A dirty folio is what
+ * reclaim stores into zswap next time.
+ */
+ folio = swap_cache_get_folio(xentry);
+ if (!folio) {
+ mpol = get_task_policy(current);
+ folio = __swap_cache_alloc_folio(xentry, GFP_HIGHUSER_MOVABLE,
+ BIT(0), NULL, mpol,
+ NO_INTERLEAVE_INDEX);
+ if (IS_ERR(folio)) {
+ ret = PTR_ERR(folio) == -ENOMEM ? -ENOMEM : -EAGAIN;
+ goto out_unlock;
+ }
+ swap_read_folio(&ctx, folio);
+ swap_read_submit(&ctx);
+ folio_lock(folio);
+ fresh = true;
+ } else {
+ folio_lock(folio);
+ }
+
+ if (folio_matches_swap_entry(folio, xentry)) {
+ folio_wait_writeback(folio);
+ if (unlikely(!folio_test_uptodate(folio))) {
+ /* Drop the folio so a later read can try again. */
+ swap_cache_del_folio(folio);
+ ret = -EAGAIN;
+ } else {
+ folio_mark_dirty(folio);
+ /* Reclaim cannot find a folio that is not on the LRU. */
+ if (fresh)
+ folio_add_lru(folio);
+ drop = true;
+ }
+ } else {
+ /* Taken from under us, so nothing was read back either. */
+ VM_WARN_ON_ONCE(1);
+ ret = -EAGAIN;
+ }
+ folio_unlock(folio);
+ folio_put(folio);
+
+ /* The slot is going away, so drop the record that names it. */
+ if (drop) {
+ xci = swap_cluster_lock(xsi, xoff);
+ if (xci) {
+ if (xswap_slot_backend(xci, xoff % SWAPFILE_CLUSTER).val ==
+ phys.val)
+ WRITE_ONCE(xci->xs_table[xoff % SWAPFILE_CLUSTER],
+ 0);
+ swap_cluster_unlock(xci);
+ }
+ }
+out_unlock:
+ mutex_unlock(&xsi->xswap_lock);
+ memalloc_noreclaim_restore(noreclaim_flags);
+ put_swap_device(xsi);
+ if (!drop)
+ return ret;
+out_free:
+ /* Leaving the mapping in place would make the unuse loop spin. */
+ xswap_free_phys_slot(si, xentry, phys);
+ return 0;
+}
+#else
+static inline int xswap_unuse_rmap(struct swap_info_struct *si,
+ unsigned long offset)
+{
+ return 0;
+}
#endif /* CONFIG_XSWAP */
static int try_to_unuse(unsigned int type)
{
@@ -3321,8 +3454,20 @@ static int try_to_unuse(unsigned int type)
entry = swp_entry(type, i);
folio = swap_cache_get_folio(entry);
- if (!folio)
+ if (!folio) {
+ /*
+ * Only a physical backend holds pointer entries. An
+ * xswap device's slots are shadows, and its
+ * cluster_info is what find_next_to_unuse() reads
+ * without the lock.
+ */
+ if (!(si->flags & SWP_XSWAP)) {
+ retval = xswap_unuse_rmap(si, i);
+ if (retval)
+ return retval;
+ }
continue;
+ }
/*
* It is conceivable that a racing task removed this folio from
--
2.54.0
^ permalink raw reply related [flat|nested] 16+ messages in thread
* [PATCH 08/12] mm, swap: reclaim physical slots backing cache-only xswap entries
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
` (6 preceding siblings ...)
2026-10-03 0:58 ` [PATCH 07/12] mm, swap: support swapoff of an xswap physical backend Baoquan He
@ 2026-10-03 0:58 ` Baoquan He
2026-10-03 0:58 ` [PATCH 09/12] mm, swap: back a large xswap folio with a contiguous physical run Baoquan He
` (4 subsequent siblings)
12 siblings, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:58 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
The swap count of an xswap entry can drop to 0 while its folio is still
in the swap cache, which leaves the physical slot redundant but occupied
until the folio leaves the cache. This happens when the page is faulted
back in during its write to the backend: do_swap_page() does not wait for
writeback, folio_free_swap() refuses a folio under writeback, and a
real-device write does not drop the cache when it completes.
Mark the slot with a cache-only bit in the swap table entry that points
back to the owner, set when the count drops to 0 and cleared when the
slot is reused. The physical reclaim scanner already walks full clusters;
make it free the slots with the bit set. It reads the bit directly,
without the xswap cluster lock, which would invert the lock order.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/swap_table.h | 37 +++++++++++++--
mm/swapfile.c | 122 ++++++++++++++++++++++++++++++++++++++++++++++++
2 files changed, 156 insertions(+), 3 deletions(-)
diff --git a/mm/swap_table.h b/mm/swap_table.h
index 22322519c3fe..40bab9890670 100644
--- a/mm/swap_table.h
+++ b/mm/swap_table.h
@@ -86,21 +86,52 @@ struct swap_memcg_table {
* Pointer-tagged swap table entry: the reverse map from a physical slot to
* the xswap entry it backs. Layout:
*
- * Pointer: | xswap_type(8) | xswap_offset(53) |100|
+ * Pointer: |C| xswap_type(7) | xswap_offset(53) |100|
*
* The low three bits identify the entry: 0b000 is a free or bad slot, a
* shadow entry ends in 0b01 and a cached folio in 0b10, which leaves
- * 0b100 as the only marker still free.
+ * 0b100 as the only marker still free. C is SWP_RMAP_CACHE_ONLY, the one
+ * bit left at the top.
*/
#define SWP_TB_PTR_MARK 0b100UL
#define SWP_TB_PTR_OFF_BITS 53
-#define SWP_TB_PTR_TYPE_BITS 8
+#define SWP_TB_PTR_TYPE_BITS 7
#define SWP_TB_PTR_OFF_SHIFT 3
#define SWP_TB_PTR_TYPE_SHIFT (SWP_TB_PTR_OFF_SHIFT + \
SWP_TB_PTR_OFF_BITS)
#define SWP_TB_PTR_OFF_MASK ((1UL << SWP_TB_PTR_OFF_BITS) - 1)
#define SWP_TB_PTR_TYPE_MASK ((1UL << SWP_TB_PTR_TYPE_BITS) - 1)
+/*
+ * Set when the xswap entry owning a physical slot has swap count 0 but its
+ * folio is still in the swap cache. The slot is redundant then, and the
+ * physical reclaim scanner may free it. The bit sits on the physical slot
+ * so that scanner can read it without the xswap cluster lock.
+ */
+#define SWP_RMAP_CACHE_ONLY (1UL << (BITS_PER_LONG - 1))
+
+/* swp_type() must fit, and SWP_RMAP_CACHE_ONLY must own the top bit. */
+static_assert(BITS_PER_LONG - SWP_TYPE_SHIFT <= SWP_TB_PTR_TYPE_BITS,
+ "xswap rmap type field is too narrow");
+static_assert(SWP_TB_PTR_TYPE_SHIFT + SWP_TB_PTR_TYPE_BITS == BITS_PER_LONG - 1,
+ "xswap rmap fields do not pack");
+
+static inline void swap_rmap_mark_cache_only(struct swap_cluster_info *ci,
+ unsigned int off)
+{
+ atomic_long_t *table = rcu_dereference_check(ci->table, true);
+
+ atomic_long_or(SWP_RMAP_CACHE_ONLY, &table[off]);
+}
+
+static inline void swap_rmap_clear_cache_only(struct swap_cluster_info *ci,
+ unsigned int off)
+{
+ atomic_long_t *table = rcu_dereference_check(ci->table, true);
+
+ atomic_long_and(~SWP_RMAP_CACHE_ONLY, &table[off]);
+}
+
static inline bool swp_tb_is_pointer(unsigned long swp_tb)
{
return (swp_tb & (BIT(3) - 1)) == SWP_TB_PTR_MARK;
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 45a871eb7bc2..5b33b2c78ee4 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -109,6 +109,12 @@ static void xswap_release_slot_backend(struct swap_info_struct *si,
unsigned int slot);
static void xswap_free_xs_table(struct swap_info_struct *si,
struct swap_cluster_info *ci);
+static void xswap_mark_cache_only(struct swap_cluster_info *ci,
+ unsigned int slot);
+static void xswap_clear_cache_only(struct swap_cluster_info *ci,
+ unsigned int start, unsigned int nr);
+static int xswap_reclaim_backing(struct swap_info_struct *si,
+ unsigned long offset, swp_entry_t xentry);
static int xswap_create(int prio);
static int xswap_destroy(int type);
@@ -1325,6 +1331,18 @@ static void swap_reclaim_full_clusters(struct swap_info_struct *si, bool force)
offset += abs(nr_reclaim);
continue;
}
+#ifdef CONFIG_XSWAP
+ } else if (swp_tb_is_pointer(swp_tb) &&
+ (swp_tb & SWP_RMAP_CACHE_ONLY)) {
+ spin_unlock(&ci->lock);
+ nr_reclaim = xswap_reclaim_backing(si, offset,
+ xswap_rmap_to_entry(swp_tb));
+ spin_lock(&ci->lock);
+ if (nr_reclaim) {
+ offset += abs(nr_reclaim);
+ continue;
+ }
+#endif
}
offset++;
}
@@ -1935,6 +1953,10 @@ static void swap_put_entries_cluster(struct swap_info_struct *si,
}
/* count will be 0 after put, slot can be reclaimed */
need_reclaim = true;
+#ifdef CONFIG_XSWAP
+ if (ci->xs_table)
+ xswap_mark_cache_only(ci, ci_off);
+#endif
}
/*
* A count != 1 or cached slot can't be freed. Put its swap
@@ -2041,6 +2063,10 @@ static int swap_dup_entries_cluster(struct swap_info_struct *si,
goto failed;
}
} while (++ci_off < ci_end);
+#ifdef CONFIG_XSWAP
+ if (ci->xs_table)
+ xswap_clear_cache_only(ci, ci_start, nr);
+#endif
swap_cluster_unlock(ci);
return 0;
failed:
@@ -3263,6 +3289,102 @@ static void xswap_release_slot_backend(struct swap_info_struct *si,
xswap_free_phys_slot(psi, entry, phys);
}
+/*
+ * Mark the physical slot backing xswap slot @slot as cache-only. Its folio
+ * is in the swap cache, so the physical copy is redundant. The physical
+ * cluster lock is needed: xswap_free_phys_slot() clears the rmap under it.
+ */
+static void xswap_mark_cache_only(struct swap_cluster_info *ci,
+ unsigned int slot)
+{
+ struct swap_cluster_info *pci;
+ swp_entry_t phys = { .val = ci->xs_table[slot] };
+
+ if (!phys.val)
+ return; /* still in zswap, no backend */
+ pci = __swap_entry_to_cluster(phys);
+ spin_lock(&pci->lock);
+ swap_rmap_mark_cache_only(pci, swp_cluster_offset(phys));
+ spin_unlock(&pci->lock);
+}
+
+/*
+ * Clear the cache-only mark of slots that were re-referenced. A slot that
+ * was cache-only had count 0, so count 1 is exactly the one to clear.
+ */
+static void xswap_clear_cache_only(struct swap_cluster_info *ci,
+ unsigned int start, unsigned int nr)
+{
+ unsigned int slot;
+
+ for (slot = start; slot < start + nr; slot++) {
+ struct swap_cluster_info *pci;
+ unsigned long swp_tb;
+ swp_entry_t phys;
+
+ swp_tb = __swap_table_get(ci, slot);
+ if (!swp_tb_is_folio(swp_tb) || swp_tb_get_count(swp_tb) != 1)
+ continue;
+ phys.val = ci->xs_table[slot];
+ if (!phys.val)
+ continue;
+ pci = __swap_entry_to_cluster(phys);
+ spin_lock(&pci->lock);
+ swap_rmap_clear_cache_only(pci, swp_cluster_offset(phys));
+ spin_unlock(&pci->lock);
+ }
+}
+
+/*
+ * Try to reclaim the physical slot backing cache-only @xentry. The physical
+ * cluster lock must not be held. Returns the folio size, negated if the free
+ * failed, or 0 if @offset turned out not to be the folio's slot.
+ */
+static int xswap_reclaim_backing(struct swap_info_struct *si,
+ unsigned long offset, swp_entry_t xentry)
+{
+ struct swap_info_struct *xsi = __swap_entry_to_info(xentry);
+ struct swap_cluster_info *xci;
+ swp_entry_t first;
+ struct folio *folio;
+ unsigned long xoff, i;
+ int ret = 0;
+
+ folio = swap_cache_get_folio(xentry);
+ if (!folio)
+ return 0;
+ if (!folio_trylock(folio)) {
+ folio_put(folio);
+ return 0;
+ }
+
+ /*
+ * The folio must own @xentry, and @offset must be the slot this device
+ * holds for that page of the folio. Otherwise the rmap went stale.
+ */
+ if (!folio_matches_swap_entry(folio, xentry))
+ goto out;
+ i = xentry.val - folio->swap.val;
+ xoff = swp_offset(folio->swap);
+ xci = swap_cluster_lock(xsi, xoff);
+ if (!xci)
+ goto out;
+ first = xswap_slot_backend(xci, xoff % SWAPFILE_CLUSTER);
+ swap_cluster_unlock(xci);
+ if (!first.val || swp_type(first) != si->type ||
+ swp_offset(first) + i != offset)
+ goto out;
+
+ /* The run is ours: skip it all, whether or not the free succeeds. */
+ ret = folio_nr_pages(folio);
+ if (!folio_free_swap(folio))
+ ret = -ret;
+out:
+ folio_unlock(folio);
+ folio_put(folio);
+ return ret;
+}
+
/*
* A written-back slot has no folio of its own here, so the unuse loop would
* skip it. Resolve it through the reverse mapping and free both sides.
--
2.54.0
^ permalink raw reply related [flat|nested] 16+ messages in thread
* [PATCH 09/12] mm, swap: back a large xswap folio with a contiguous physical run
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
` (7 preceding siblings ...)
2026-10-03 0:58 ` [PATCH 08/12] mm, swap: reclaim physical slots backing cache-only xswap entries Baoquan He
@ 2026-10-03 0:58 ` Baoquan He
2026-10-03 0:58 ` [PATCH 10/12] mm, swap: enable THP swapin for xswap entries Baoquan He
` (3 subsequent siblings)
12 siblings, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:58 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
A folio that zswap refuses goes to a physical swap device, one slot per
entry. A large folio is read back with a single IO from its first physical
slot, so its slots must be contiguous. Reserve the whole range at once by
passing the order down the allocation path.
The contiguous range must come from a block device backend: if a real
device exists but none is a block device, refuse the large order, so the
folio is split and swapped one page at a time. Record the backend in
xs_table before the reverse map that swapoff uses, so a slot is never
visible without its owner.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/page_io.c | 2 +-
mm/swap.h | 8 +-
mm/swapfile.c | 199 ++++++++++++++++++++++++++++++--------------------
mm/zswap.c | 4 +-
4 files changed, 125 insertions(+), 88 deletions(-)
diff --git a/mm/page_io.c b/mm/page_io.c
index 7a7b7eecb2a8..7e8e8150bdcf 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -257,7 +257,7 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
swp_entry_t phys;
/* The pool refused it: fall back to a real swap device. */
- phys = xswap_backend_alloc(entry);
+ phys = xswap_backend_alloc(entry, folio_order(folio));
if (!phys.val) {
folio_mark_dirty(folio);
return AOP_WRITEPAGE_ACTIVATE;
diff --git a/mm/swap.h b/mm/swap.h
index 838dac4282f6..9f04c8486b98 100644
--- a/mm/swap.h
+++ b/mm/swap.h
@@ -475,8 +475,8 @@ int shmem_writeout(struct swap_io_ctx *ctx, struct folio *folio,
#ifdef CONFIG_XSWAP
swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci, unsigned int slot);
-swp_entry_t xswap_backend_alloc(swp_entry_t entry);
-void xswap_backend_free(swp_entry_t entry, swp_entry_t phys);
+swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order);
+void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order);
#else
static inline swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci,
unsigned int slot)
@@ -484,12 +484,12 @@ static inline swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci,
return (swp_entry_t){};
}
-static inline swp_entry_t xswap_backend_alloc(swp_entry_t entry)
+static inline swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order)
{
return (swp_entry_t){};
}
-static inline void xswap_backend_free(swp_entry_t entry, swp_entry_t phys)
+static inline void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order)
{
}
#endif /* CONFIG_XSWAP */
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 5b33b2c78ee4..8e005e8d2250 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -1168,10 +1168,10 @@ static bool cluster_scan_range(struct swap_info_struct *si,
static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
struct swap_cluster_info *ci,
struct folio *folio,
- unsigned int ci_off)
+ unsigned int ci_off, unsigned int order)
{
- unsigned int order;
- unsigned long nr_pages;
+ unsigned long nr_pages = 1UL << order;
+ unsigned long i;
lockdep_assert_held(&ci->lock);
@@ -1180,33 +1180,24 @@ static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
/*
* All mm swap allocation starts with a folio (folio_alloc_swap),
- * it's also the only allocation path for large orders allocation.
- * Such swap slots starts with count == 0 and will be increased
- * upon folio unmap.
+ * it's also the allocation path for large orders. Such swap slots
+ * starts with count == 0 and will be increased upon folio unmap.
*
- * Else, it's an exclusive order 0 allocation for a slot that has
- * no folio of its own: hibernation, and the physical slot that
- * backs an xswap entry. The slot starts with count == 1 and
- * never increases.
+ * Else, it's an exclusive allocation for a range of slots that has no
+ * folio of its own: hibernation, and the range that backs an xswap
+ * folio. Each slot starts with count == 1 and never increases.
*/
+ swap_cluster_assert_empty(ci, ci_off, nr_pages, false);
if (likely(folio)) {
- order = folio_order(folio);
- nr_pages = 1 << order;
- swap_cluster_assert_empty(ci, ci_off, nr_pages, false);
__swap_cache_add_folio(ci, folio, swp_entry(si->type,
ci_off + cluster_offset(si, ci)));
} else {
- order = 0;
- nr_pages = 1;
- swap_cluster_assert_empty(ci, ci_off, 1, false);
- /* Fake shadow placeholder with no flag; no zeromap is used. */
- __swap_table_set(ci, ci_off, __swp_tb_mk_count(shadow_to_swp_tb(NULL, 0), 1));
+ /* Fake shadow placeholders with no flag; no zeromap is used. */
+ for (i = 0; i < nr_pages; i++)
+ __swap_table_set(ci, ci_off + i,
+ __swp_tb_mk_count(shadow_to_swp_tb(NULL, 0), 1));
}
- /*
- * The first allocation in a cluster makes the
- * cluster exclusive to this order
- */
if (cluster_is_empty(ci))
ci->order = order;
ci->count += nr_pages;
@@ -1219,13 +1210,13 @@ static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
static unsigned long alloc_swap_scan_cluster(struct swap_info_struct *si,
struct swap_cluster_info *ci,
struct folio *folio,
- unsigned long offset)
+ unsigned long offset,
+ unsigned int order)
{
unsigned long next = SWAP_ENTRY_INVALID, found = SWAP_ENTRY_INVALID;
unsigned long start = ALIGN_DOWN(offset, SWAPFILE_CLUSTER);
- unsigned int order = likely(folio) ? folio_order(folio) : 0;
unsigned long end = start + SWAPFILE_CLUSTER;
- unsigned int nr_pages = 1 << order;
+ unsigned long nr_pages = 1UL << order;
bool need_reclaim, ret, usable;
lockdep_assert_held(&ci->lock);
@@ -1251,7 +1242,7 @@ static unsigned long alloc_swap_scan_cluster(struct swap_info_struct *si,
if (!ret)
continue;
}
- if (!__swap_cluster_alloc_entries(si, ci, folio, offset % SWAPFILE_CLUSTER))
+ if (!__swap_cluster_alloc_entries(si, ci, folio, offset % SWAPFILE_CLUSTER, order))
break;
found = offset;
offset += nr_pages;
@@ -1282,7 +1273,7 @@ static unsigned long alloc_swap_scan_cluster(struct swap_info_struct *si,
static unsigned long alloc_swap_scan_list(struct swap_info_struct *si,
struct list_head *list,
struct folio *folio,
- bool scan_all)
+ unsigned int order, bool scan_all)
{
unsigned long found = SWAP_ENTRY_INVALID;
@@ -1293,7 +1284,7 @@ static unsigned long alloc_swap_scan_list(struct swap_info_struct *si,
if (!ci)
break;
offset = cluster_offset(si, ci);
- found = alloc_swap_scan_cluster(si, ci, folio, offset);
+ found = alloc_swap_scan_cluster(si, ci, folio, offset, order);
if (found)
break;
} while (scan_all);
@@ -1374,25 +1365,49 @@ static void swap_reclaim_work(struct work_struct *work)
swap_reclaim_full_clusters(si, true);
}
+#ifdef CONFIG_XSWAP
+/* Whether a block device that can back a large xswap folio is available. */
+static bool xswap_has_large_backend(void)
+{
+ struct swap_info_struct *si;
+ bool found = false;
+
+ spin_lock(&swap_avail_lock);
+ plist_for_each_entry(si, &swap_avail_head, avail_list) {
+ if (!(si->flags & SWP_XSWAP) && (si->flags & SWP_BLKDEV)) {
+ found = true;
+ break;
+ }
+ }
+ spin_unlock(&swap_avail_lock);
+
+ return found;
+}
+#endif /* CONFIG_XSWAP */
+
/*
* Try to allocate swap entries with specified order and try set a new
* cluster for current CPU too.
*/
static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
- struct folio *folio)
+ struct folio *folio, unsigned int order)
{
struct swap_cluster_info *ci;
- unsigned int order = likely(folio) ? folio_order(folio) : 0;
unsigned long offset = SWAP_ENTRY_INVALID, found = SWAP_ENTRY_INVALID;
/*
* Swapfile is not block device so unable
- * to allocate large entries.
+ * to allocate large entries. xswap has no backing file, so it can.
*/
- if (order && !(si->flags & SWP_BLKDEV))
+ if (order && !(si->flags & SWP_BLKDEV) && !(si->flags & SWP_XSWAP))
return 0;
#ifdef CONFIG_XSWAP
+ /* A large xswap folio needs a contiguous block device run behind it. */
+ if (order && (si->flags & SWP_XSWAP) &&
+ atomic_read(&nr_real_swapfiles) && !xswap_has_large_backend())
+ return 0;
+
/*
* Top the range up early. Every path below leaves through `done`,
* so this has to come first; mapping pages can sleep, and doing it
@@ -1415,7 +1430,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
if (cluster_is_usable(ci, order)) {
if (cluster_is_empty(ci))
offset = cluster_offset(si, ci);
- found = alloc_swap_scan_cluster(si, ci, folio, offset);
+ found = alloc_swap_scan_cluster(si, ci, folio, offset, order);
} else {
swap_cluster_unlock(ci);
}
@@ -1429,7 +1444,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
* cluster is the only thing a large order can allocate from.
*/
if (!folio && order < PMD_ORDER) {
- found = alloc_swap_scan_list(si, &si->frag_clusters[order], folio, false);
+ found = alloc_swap_scan_list(si, &si->frag_clusters[order], folio, order, false);
if (found)
goto done;
}
@@ -1439,19 +1454,19 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
* to spread out the writes.
*/
if (si->flags & SWP_PAGE_DISCARD) {
- found = alloc_swap_scan_list(si, &si->free_clusters, folio, false);
+ found = alloc_swap_scan_list(si, &si->free_clusters, folio, order, false);
if (found)
goto done;
}
if (order < PMD_ORDER) {
- found = alloc_swap_scan_list(si, &si->nonfull_clusters[order], folio, true);
+ found = alloc_swap_scan_list(si, &si->nonfull_clusters[order], folio, order, true);
if (found)
goto done;
}
if (!(si->flags & SWP_PAGE_DISCARD)) {
- found = alloc_swap_scan_list(si, &si->free_clusters, folio, false);
+ found = alloc_swap_scan_list(si, &si->free_clusters, folio, order, false);
if (found)
goto done;
}
@@ -1467,7 +1482,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
* failure is not critical. Scanning one cluster still
* keeps the list rotated and reclaimed (for clean swap cache).
*/
- found = alloc_swap_scan_list(si, &si->frag_clusters[order], folio, false);
+ found = alloc_swap_scan_list(si, &si->frag_clusters[order], folio, order, false);
if (found)
goto done;
}
@@ -1481,11 +1496,11 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
* Clusters here have at least one usable slots and can't fail order 0
* allocation, but reclaim may drop si->lock and race with another user.
*/
- found = alloc_swap_scan_list(si, &si->frag_clusters[o], folio, true);
+ found = alloc_swap_scan_list(si, &si->frag_clusters[o], folio, order, true);
if (found)
goto done;
- found = alloc_swap_scan_list(si, &si->nonfull_clusters[o], folio, true);
+ found = alloc_swap_scan_list(si, &si->nonfull_clusters[o], folio, order, true);
if (found)
goto done;
}
@@ -1493,7 +1508,7 @@ static unsigned long cluster_alloc_swap_entry(struct swap_info_struct *si,
#ifdef CONFIG_XSWAP
/* A concurrent free or grow may have added clusters; retry once. */
if (!found && (si->flags & SWP_XSWAP))
- found = alloc_swap_scan_list(si, &si->free_clusters, folio, false);
+ found = alloc_swap_scan_list(si, &si->free_clusters, folio, order, false);
#endif
done:
if (!(si->flags & SWP_SOLIDSTATE))
@@ -1714,7 +1729,7 @@ static bool swap_alloc_fast(struct folio *folio, bool may_zswap)
if (cluster_is_usable(ci, order)) {
if (cluster_is_empty(ci))
offset = cluster_offset(si, ci);
- alloc_swap_scan_cluster(si, ci, folio, offset);
+ alloc_swap_scan_cluster(si, ci, folio, offset, order);
} else {
swap_cluster_unlock(ci);
}
@@ -1742,7 +1757,7 @@ static void swap_alloc_slow(struct folio *folio, bool may_zswap)
plist_requeue(&si->avail_list, &swap_avail_head);
spin_unlock(&swap_avail_lock);
if (get_swap_device_info(si)) {
- cluster_alloc_swap_entry(si, folio);
+ cluster_alloc_swap_entry(si, folio, folio_order(folio));
put_swap_device(si);
if (folio_test_swapcache(folio))
return;
@@ -2562,13 +2577,13 @@ swp_entry_t swap_alloc_hibernation_slot(int type)
if (pcp_si == si && pcp_offset) {
ci = swap_cluster_lock(si, pcp_offset);
if (cluster_is_usable(ci, 0))
- offset = alloc_swap_scan_cluster(si, ci, NULL, pcp_offset);
+ offset = alloc_swap_scan_cluster(si, ci, NULL, pcp_offset, 0);
else
swap_cluster_unlock(ci);
}
rcu_read_unlock();
if (!offset)
- offset = cluster_alloc_swap_entry(si, NULL);
+ offset = cluster_alloc_swap_entry(si, NULL, 0);
local_unlock(&percpu_swap_cluster.lock);
if (offset)
entry = swp_entry(si->type, offset);
@@ -4505,32 +4520,33 @@ static void xswap_free_xs_table(struct swap_info_struct *si,
ci->xs_table = NULL;
}
-/* Order-0 slot from @si, no folio attached; goes through the per-CPU cache. */
-static unsigned long xswap_alloc_slot_local(struct swap_info_struct *si)
+/* A range of 1 << @order slots, no folio attached. */
+static unsigned long xswap_alloc_slot_local(struct swap_info_struct *si,
+ unsigned int order)
{
struct swap_cluster_info *ci;
unsigned long pcp_offset, offset = SWAP_ENTRY_INVALID;
local_lock(&percpu_swap_cluster.lock);
- if (this_cpu_read(percpu_swap_cluster.si[0]) == si) {
- pcp_offset = this_cpu_read(percpu_swap_cluster.offset[0]);
+ if (this_cpu_read(percpu_swap_cluster.si[order]) == si) {
+ pcp_offset = this_cpu_read(percpu_swap_cluster.offset[order]);
if (pcp_offset) {
ci = swap_cluster_lock(si, pcp_offset);
- if (cluster_is_usable(ci, 0))
+ if (cluster_is_usable(ci, order))
offset = alloc_swap_scan_cluster(si, ci, NULL,
- pcp_offset);
+ pcp_offset, order);
else
swap_cluster_unlock(ci);
}
}
if (offset == SWAP_ENTRY_INVALID)
- offset = cluster_alloc_swap_entry(si, NULL);
+ offset = cluster_alloc_swap_entry(si, NULL, order);
local_unlock(&percpu_swap_cluster.lock);
return offset;
}
-static swp_entry_t xswap_alloc_phys_slot(void)
+static swp_entry_t xswap_alloc_phys_slot(unsigned int order)
{
struct swap_info_struct *devs[MAX_SWAPFILES], *si;
int nr = 0, i;
@@ -4555,7 +4571,7 @@ static swp_entry_t xswap_alloc_phys_slot(void)
if (!get_swap_device_info(devs[i]))
continue;
- offset = xswap_alloc_slot_local(devs[i]);
+ offset = xswap_alloc_slot_local(devs[i], order);
put_swap_device(devs[i]);
if (offset && offset != SWAP_ENTRY_INVALID)
return swp_entry(devs[i]->type, offset);
@@ -4581,57 +4597,78 @@ static void xswap_install_rmap(swp_entry_t phys, swp_entry_t entry)
}
}
-swp_entry_t xswap_backend_alloc(swp_entry_t entry)
+/* Back the range of 1 << @order slots at @entry with contiguous physical slots. */
+swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order)
{
struct swap_info_struct *si = __swap_entry_to_info(entry);
struct swap_cluster_info *ci;
+ unsigned long nr = 1UL << order;
unsigned long off = swp_offset(entry);
swp_entry_t phys;
- bool recorded = false;
+ unsigned int i;
+
+ /*
+ * Get the xswap side ready before a physical slot exists, so the
+ * reverse map installed below cannot expose a slot with no owner.
+ */
+ ci = swap_cluster_lock(si, off);
+ if (!ci)
+ return (swp_entry_t){};
+ if (!xswap_xs_table_locked(ci, GFP_KERNEL | __GFP_HIGH | __GFP_NOMEMALLOC)) {
+ swap_cluster_unlock(ci);
+ return (swp_entry_t){};
+ }
+ swap_cluster_unlock(ci);
- phys = xswap_alloc_phys_slot();
+ phys = xswap_alloc_phys_slot(order);
if (!phys.val)
return phys;
- xswap_install_rmap(phys, entry);
+ /* One IO reads a large folio, so its range has to be contiguous. */
+ VM_WARN_ON_ONCE(swp_offset(phys) % nr);
+ VM_WARN_ON_ONCE(off % SWAPFILE_CLUSTER + nr > SWAPFILE_CLUSTER);
+ /* Owner first, then the reverse map, under one hold of the lock. */
ci = swap_cluster_lock(si, off);
- if (ci) {
- recorded = xswap_slot_set_backend(ci, off % SWAPFILE_CLUSTER, phys);
- swap_cluster_unlock(ci);
- }
+ for (i = 0; i < nr; i++) {
+ swp_entry_t cur = swp_entry(swp_type(phys), swp_offset(phys) + i);
- if (!recorded) {
- /* The backend is unrecorded, so the data would be unreachable. */
- xswap_backend_free(entry, phys);
- return (swp_entry_t){};
+ xswap_slot_set_backend(ci, (off + i) % SWAPFILE_CLUSTER, cur);
+ xswap_install_rmap(cur, swp_entry(si->type, off + i));
}
+ swap_cluster_unlock(ci);
return phys;
}
/* Undo xswap_backend_alloc(): the zswap copy is still in place. */
-void xswap_backend_free(swp_entry_t entry, swp_entry_t phys)
+void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order)
{
struct swap_info_struct *si = __swap_entry_to_info(entry);
- struct swap_info_struct *psi;
- struct swap_cluster_info *ci;
+ struct swap_info_struct *psi = swap_type_to_info(swp_type(phys));
+ unsigned long nr = 1UL << order;
unsigned long off = swp_offset(entry);
+ unsigned long poff = swp_offset(phys);
+ unsigned int i;
- ci = swap_cluster_lock(si, off);
- if (ci) {
- unsigned int slot = off % SWAPFILE_CLUSTER;
- unsigned long *xs = ci->xs_table;
+ for (i = 0; i < nr; i++) {
+ struct swap_cluster_info *ci;
- if (xs && xs[slot] == phys.val)
- WRITE_ONCE(xs[slot], 0);
- swap_cluster_unlock(ci);
- }
+ ci = swap_cluster_lock(si, off + i);
+ if (ci) {
+ unsigned int slot = (off + i) % SWAPFILE_CLUSTER;
+ unsigned long *xs = ci->xs_table;
- /* xswap_free_phys_slot() clears the reverse mapping. */
- psi = swap_type_to_info(swp_type(phys));
- if (psi)
- xswap_free_phys_slot(psi, entry, phys);
+ if (xs && xs[slot] == phys.val + i)
+ WRITE_ONCE(xs[slot], 0);
+ swap_cluster_unlock(ci);
+ }
+
+ /* xswap_free_phys_slot() clears the reverse mapping. */
+ if (psi)
+ xswap_free_phys_slot(psi, swp_entry(si->type, off + i),
+ swp_entry(swp_type(phys), poff + i));
+ }
}
static int xswap_map_clusters(struct swap_info_struct *si,
diff --git a/mm/zswap.c b/mm/zswap.c
index e614c04697f2..03171e4f2852 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1056,7 +1056,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
/* Reserve the destination before dropping the zswap copy. */
if (is_xswap) {
- phys = xswap_backend_alloc(swpentry);
+ phys = xswap_backend_alloc(swpentry, 0);
if (!phys.val) {
ret = -ENOMEM;
goto err;
@@ -1102,7 +1102,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
err:
if (is_xswap && phys.val)
- xswap_backend_free(swpentry, phys);
+ xswap_backend_free(swpentry, phys, 0);
swap_cache_del_folio(folio);
folio_unlock(folio);
folio_put(folio);
--
2.54.0
^ permalink raw reply related [flat|nested] 16+ messages in thread
* [PATCH 10/12] mm, swap: enable THP swapin for xswap entries
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
` (8 preceding siblings ...)
2026-10-03 0:58 ` [PATCH 09/12] mm, swap: back a large xswap folio with a contiguous physical run Baoquan He
@ 2026-10-03 0:58 ` Baoquan He
2026-10-03 0:58 ` [PATCH 11/12] mm, swap: drop swap_folio_sector() Baoquan He
` (2 subsequent siblings)
12 siblings, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:58 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
Read a large anon folio back as one unit when all of its xswap entries are
contiguous on one physical swap device, instead of falling back to order-0
faults. The device is marked SWP_SYNCHRONOUS_IO, which is required for THP
swapin and also skips readahead, so no swap cache slot is pinned while
xswap tries to shrink.
A batched read is refused when its data is still in zswap or spans more
than one device, since zswap cannot load a large folio. It is then read
page by page and the fault retries at a smaller order.
xswap_check_backing() decides this under the cluster lock, before the folio
is allocated, so zswap_load() takes the physical path instead of warning.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
mm/memory.c | 7 +++++--
mm/swap.h | 8 ++++++++
mm/swap_state.c | 6 ++++++
mm/swapfile.c | 36 +++++++++++++++++++++++++++++++++++-
4 files changed, 54 insertions(+), 3 deletions(-)
diff --git a/mm/memory.c b/mm/memory.c
index 1f83a26f8733..6fd9258772ea 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -4833,15 +4833,18 @@ static unsigned long thp_swapin_suitable_orders(struct vm_fault *vmf)
if (unlikely(userfaultfd_armed(vma)))
return 0;
+ entry = softleaf_from_pte(vmf->orig_pte);
+
/*
* A large swapped out folio could be partially or fully in zswap. We
* lack handling for such cases, so fallback to swapping in order-0
* folio.
+ * An xswap entry is deferred to __swap_cache_add_check() instead.
*/
- if (!zswap_never_enabled())
+ if (!(__swap_entry_to_info(entry)->flags & SWP_XSWAP) &&
+ !zswap_never_enabled())
return 0;
- entry = softleaf_from_pte(vmf->orig_pte);
/*
* Get a list of all the (large) orders below PMD_ORDER that are enabled
* and suitable for swapping THP.
diff --git a/mm/swap.h b/mm/swap.h
index 9f04c8486b98..628395885919 100644
--- a/mm/swap.h
+++ b/mm/swap.h
@@ -477,6 +477,8 @@ int shmem_writeout(struct swap_io_ctx *ctx, struct folio *folio,
swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci, unsigned int slot);
swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order);
void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order);
+int xswap_check_backing(struct swap_cluster_info *ci, unsigned int off,
+ unsigned int nr);
#else
static inline swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci,
unsigned int slot)
@@ -492,6 +494,12 @@ static inline swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int or
static inline void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order)
{
}
+
+static inline int xswap_check_backing(struct swap_cluster_info *ci,
+ unsigned int off, unsigned int nr)
+{
+ return nr;
+}
#endif /* CONFIG_XSWAP */
#endif /* _MM_SWAP_H */
diff --git a/mm/swap_state.c b/mm/swap_state.c
index c5c182489d6a..56e4879abc93 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -194,6 +194,12 @@ static int __swap_cache_add_check(struct swap_cluster_info *ci,
if (nr == 1)
return 0;
+ /* A zswap-backed or mixed batch falls back to a smaller order. */
+ if (__swap_entry_to_info(targ_entry)->flags & SWP_XSWAP) {
+ if (xswap_check_backing(ci, round_down(ci_off, nr), nr) != nr)
+ return -EBUSY;
+ }
+
is_zero = __swap_table_test_zero(ci, ci_off);
ci_off = round_down(ci_off, nr);
ci_end = ci_off + nr;
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 8e005e8d2250..8e6b0cb899e5 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -4454,6 +4454,36 @@ swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci, unsigned int slot)
return (swp_entry_t){ .val = xs[slot] };
}
+/*
+ * A large xswap folio is read back as one IO only when the whole range is
+ * backed by one device, contiguously. Returns how many leading slots are.
+ */
+int xswap_check_backing(struct swap_cluster_info *ci, unsigned int off,
+ unsigned int nr)
+{
+ swp_entry_t first;
+ unsigned int i;
+
+ lockdep_assert_held(&ci->lock);
+ /* A folio's slots never straddle a cluster, or the walk runs off. */
+ VM_WARN_ON_ONCE(off + nr > SWAPFILE_CLUSTER);
+
+ /* In zswap, which cannot load a large folio. */
+ if (!ci->xs_table || !ci->xs_table[off])
+ return 1;
+
+ first.val = ci->xs_table[off];
+ for (i = 1; i < nr; i++) {
+ swp_entry_t cur = { .val = ci->xs_table[off + i] };
+
+ if (swp_type(cur) != swp_type(first) ||
+ swp_offset(cur) != swp_offset(first) + i)
+ break;
+ }
+
+ return i;
+}
+
/* Cluster lock held. xswap never unmaps a cluster below the mapped prefix. */
static unsigned long *xswap_xs_table_locked(struct swap_cluster_info *ci,
gfp_t gfp)
@@ -5345,7 +5375,11 @@ static int xswap_create(int prio)
maxpages = rounddown(maxpages, SWAPFILE_CLUSTER);
si->bdev = NULL;
- si->flags |= SWP_XSWAP | SWP_SOLIDSTATE;
+ /*
+ * Reads are served from zswap, so no readahead. Only a synchronous
+ * device can swap a large folio back in as a unit.
+ */
+ si->flags |= SWP_XSWAP | SWP_SOLIDSTATE | SWP_SYNCHRONOUS_IO;
si->max = maxpages;
si->pages = maxpages - 1;
/*
--
2.54.0
^ permalink raw reply related [flat|nested] 16+ messages in thread
* [PATCH 11/12] mm, swap: drop swap_folio_sector()
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
` (9 preceding siblings ...)
2026-10-03 0:58 ` [PATCH 10/12] mm, swap: enable THP swapin for xswap entries Baoquan He
@ 2026-10-03 0:58 ` Baoquan He
2026-10-03 0:59 ` [PATCH 12/12] mm, swap: skip a NOFS backend when reclaim cannot enter the fs Baoquan He
2026-10-03 9:35 ` [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Nhat Pham
12 siblings, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:58 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
No one uses it now.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
include/linux/swap.h | 1 -
mm/swapfile.c | 13 -------------
2 files changed, 14 deletions(-)
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 2080540c6e39..e78717327126 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -417,7 +417,6 @@ extern int __swap_count(swp_entry_t entry);
extern bool swap_entry_swapped(struct swap_info_struct *si, swp_entry_t entry);
extern int swp_swapcount(swp_entry_t entry);
extern struct swap_info_struct *get_swap_device(swp_entry_t entry);
-sector_t swap_folio_sector(struct folio *folio);
sector_t swap_entry_sector(swp_entry_t entry);
/*
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 8e6b0cb899e5..c63f1c641035 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -541,19 +541,6 @@ offset_to_swap_extent(struct swap_info_struct *sis, unsigned long offset)
BUG();
}
-sector_t swap_folio_sector(struct folio *folio)
-{
- struct swap_info_struct *sis = __swap_entry_to_info(folio->swap);
- struct swap_extent *se;
- sector_t sector;
- pgoff_t offset;
-
- offset = swp_offset(folio->swap);
- se = offset_to_swap_extent(sis, offset);
- sector = se->start_block + (offset - se->start_page);
- return sector << (PAGE_SHIFT - 9);
-}
-
sector_t swap_entry_sector(swp_entry_t entry)
{
struct swap_info_struct *sis = __swap_entry_to_info(entry);
--
2.54.0
^ permalink raw reply related [flat|nested] 16+ messages in thread
* [PATCH 12/12] mm, swap: skip a NOFS backend when reclaim cannot enter the fs
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
` (10 preceding siblings ...)
2026-10-03 0:58 ` [PATCH 11/12] mm, swap: drop swap_folio_sector() Baoquan He
@ 2026-10-03 0:59 ` Baoquan He
2026-10-03 9:35 ` [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Nhat Pham
12 siblings, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03 0:59 UTC (permalink / raw)
To: linux-mm
Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
Baoquan He
An xswap folio's swap entry points at the xswap device, but its data may
be written to a real backend. may_enter_fs() reads the device out of that
entry, so it checks the xswap device and not the backend: a context that
cannot enter the fs is still allowed to write, and
xswap_alloc_phys_slot() may pick NFS or SMB. Writing there re-enters the
filesystem.
Marking the xswap device REQUIRE_NOFS would block zswap too, and zswap
needs no filesystem. Pass may_enter_fs down in swap_io_ctx instead, taken
from sc->gfp_mask. xswap_alloc_phys_slot() skips a NOFS backend when it
is false. With none left, AOP_WRITEPAGE_ACTIVATE keeps the folio.
zswap reaches the allocator from two places. Its shrinker runs only when
count_objects() saw __GFP_FS, and zswap_store() on the swapout path
carries the reclaim's mask down through shrink_memcg(). The other
swap_io_ctx users default to false, as shmem_write_folio() does, reached
that way from i915 and TTM.
Signed-off-by: Baoquan He <hebaoquan@kylinos.cn>
---
include/linux/swap_ops.h | 1 +
include/linux/zswap.h | 4 ++--
mm/page_io.c | 5 +++--
mm/swap.h | 6 ++++--
mm/swapfile.c | 10 +++++++---
mm/vmscan.c | 4 +++-
mm/zswap.c | 35 +++++++++++++++++++++++++----------
7 files changed, 45 insertions(+), 20 deletions(-)
diff --git a/include/linux/swap_ops.h b/include/linux/swap_ops.h
index 198c5738c3c7..ed0017be12fb 100644
--- a/include/linux/swap_ops.h
+++ b/include/linux/swap_ops.h
@@ -18,6 +18,7 @@ struct swap_iocb {
struct swap_io_ctx {
struct swap_iocb *sio;
struct swap_info_struct *sis;
+ bool may_enter_fs; /* this context may enter the fs */
};
/*
diff --git a/include/linux/zswap.h b/include/linux/zswap.h
index df6cafbe95dc..809b8b839f3b 100644
--- a/include/linux/zswap.h
+++ b/include/linux/zswap.h
@@ -25,7 +25,7 @@ struct zswap_lruvec_state {
};
unsigned long zswap_total_pages(void);
-bool zswap_store(struct folio *folio);
+bool zswap_store(struct folio *folio, bool may_enter_fs);
int zswap_load(struct folio *folio);
void zswap_invalidate(int type, pgoff_t offset, unsigned long nr_entries);
int zswap_swapon(int type, unsigned long nr_pages);
@@ -39,7 +39,7 @@ bool zswap_never_enabled(void);
struct zswap_lruvec_state {};
-static inline bool zswap_store(struct folio *folio)
+static inline bool zswap_store(struct folio *folio, bool may_enter_fs)
{
return false;
}
diff --git a/mm/page_io.c b/mm/page_io.c
index 7e8e8150bdcf..fb32b9db1df7 100644
--- a/mm/page_io.c
+++ b/mm/page_io.c
@@ -235,7 +235,7 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
*/
swap_zeromap_folio_clear(folio);
- if (zswap_store(folio)) {
+ if (zswap_store(folio, ctx->may_enter_fs)) {
count_mthp_stat(folio_order(folio), MTHP_STAT_ZSWPOUT);
goto out_unlock;
}
@@ -257,7 +257,8 @@ int swap_writeout(struct swap_io_ctx *ctx, struct folio *folio)
swp_entry_t phys;
/* The pool refused it: fall back to a real swap device. */
- phys = xswap_backend_alloc(entry, folio_order(folio));
+ phys = xswap_backend_alloc(entry, folio_order(folio),
+ ctx->may_enter_fs);
if (!phys.val) {
folio_mark_dirty(folio);
return AOP_WRITEPAGE_ACTIVATE;
diff --git a/mm/swap.h b/mm/swap.h
index 628395885919..f7c832b7fdc4 100644
--- a/mm/swap.h
+++ b/mm/swap.h
@@ -475,7 +475,8 @@ int shmem_writeout(struct swap_io_ctx *ctx, struct folio *folio,
#ifdef CONFIG_XSWAP
swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci, unsigned int slot);
-swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order);
+swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order,
+ bool may_enter_fs);
void xswap_backend_free(swp_entry_t entry, swp_entry_t phys, unsigned int order);
int xswap_check_backing(struct swap_cluster_info *ci, unsigned int off,
unsigned int nr);
@@ -486,7 +487,8 @@ static inline swp_entry_t xswap_slot_backend(struct swap_cluster_info *ci,
return (swp_entry_t){};
}
-static inline swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order)
+static inline swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order,
+ bool may_enter_fs)
{
return (swp_entry_t){};
}
diff --git a/mm/swapfile.c b/mm/swapfile.c
index c63f1c641035..180cf538f9ab 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -4563,7 +4563,7 @@ static unsigned long xswap_alloc_slot_local(struct swap_info_struct *si,
return offset;
}
-static swp_entry_t xswap_alloc_phys_slot(unsigned int order)
+static swp_entry_t xswap_alloc_phys_slot(unsigned int order, bool may_enter_fs)
{
struct swap_info_struct *devs[MAX_SWAPFILES], *si;
int nr = 0, i;
@@ -4577,6 +4577,9 @@ static swp_entry_t xswap_alloc_phys_slot(unsigned int order)
plist_for_each_entry(si, &swap_avail_head, avail_list) {
if (si->flags & SWP_XSWAP)
continue;
+ /* Writing there would re-enter the fs this context is in. */
+ if (!may_enter_fs && (si->ops->flags & SWAP_OPS_F_REQUIRE_NOFS))
+ continue;
if (nr == MAX_SWAPFILES)
break;
devs[nr++] = si;
@@ -4615,7 +4618,8 @@ static void xswap_install_rmap(swp_entry_t phys, swp_entry_t entry)
}
/* Back the range of 1 << @order slots at @entry with contiguous physical slots. */
-swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order)
+swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order,
+ bool may_enter_fs)
{
struct swap_info_struct *si = __swap_entry_to_info(entry);
struct swap_cluster_info *ci;
@@ -4637,7 +4641,7 @@ swp_entry_t xswap_backend_alloc(swp_entry_t entry, unsigned int order)
}
swap_cluster_unlock(ci);
- phys = xswap_alloc_phys_slot(order);
+ phys = xswap_alloc_phys_slot(order, may_enter_fs);
if (!phys.val)
return phys;
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 91295070ca33..fc4754779d76 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -1169,7 +1169,9 @@ static unsigned int shrink_folio_list(struct list_head *folio_list,
unsigned int nr_reclaimed = 0, nr_demoted = 0;
unsigned int pgactivate = 0;
bool do_demote_pass;
- struct swap_io_ctx ctx = {};
+ struct swap_io_ctx ctx = {
+ .may_enter_fs = !!(sc->gfp_mask & __GFP_FS),
+ };
folio_batch_init(&free_folios);
memset(stat, 0, sizeof(*stat));
diff --git a/mm/zswap.c b/mm/zswap.c
index 03171e4f2852..789730c2294f 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -1003,7 +1003,7 @@ static bool zswap_decompress(struct zswap_entry *entry, struct folio *folio)
* freed.
*/
static int zswap_writeback_entry(struct zswap_entry *entry,
- swp_entry_t swpentry)
+ swp_entry_t swpentry, bool may_enter_fs)
{
struct xarray *tree;
pgoff_t offset = swp_offset(swpentry);
@@ -1056,7 +1056,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
/* Reserve the destination before dropping the zswap copy. */
if (is_xswap) {
- phys = xswap_backend_alloc(swpentry, 0);
+ phys = xswap_backend_alloc(swpentry, 0, may_enter_fs);
if (!phys.val) {
ret = -ENOMEM;
goto err;
@@ -1134,11 +1134,18 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
* can expect from writeback. We scale down the number of objects available
* for reclaim by this ratio.
*/
+
+struct zswap_shrink_arg {
+ bool *swapcache_hit; /* page was already in the swap cache */
+ bool may_enter_fs;
+};
+
static enum lru_status shrink_memcg_cb(struct list_head *item, struct list_lru_one *l,
void *arg)
{
+ struct zswap_shrink_arg *shrinking = arg;
struct zswap_entry *entry = container_of(item, struct zswap_entry, lru);
- bool *encountered_page_in_swapcache = (bool *)arg;
+ bool *encountered_page_in_swapcache = shrinking->swapcache_hit;
swp_entry_t swpentry;
enum lru_status ret = LRU_REMOVED_RETRY;
int writeback_result;
@@ -1193,7 +1200,8 @@ static enum lru_status shrink_memcg_cb(struct list_head *item, struct list_lru_o
*/
spin_unlock(&l->lock);
- writeback_result = zswap_writeback_entry(entry, swpentry);
+ writeback_result = zswap_writeback_entry(entry, swpentry,
+ shrinking->may_enter_fs);
if (writeback_result) {
zswap_reject_reclaim_fail++;
@@ -1220,6 +1228,7 @@ static unsigned long zswap_shrinker_scan(struct shrinker *shrinker,
{
unsigned long shrink_ret;
bool encountered_page_in_swapcache = false;
+ struct zswap_shrink_arg shrinking = {};
if (!zswap_shrinker_enabled ||
!mem_cgroup_zswap_writeback_enabled(sc->memcg)) {
@@ -1227,8 +1236,10 @@ static unsigned long zswap_shrinker_scan(struct shrinker *shrinker,
return SHRINK_STOP;
}
+ shrinking.swapcache_hit = &encountered_page_in_swapcache;
+ shrinking.may_enter_fs = !!(sc->gfp_mask & __GFP_FS);
shrink_ret = list_lru_shrink_walk(&zswap_list_lru, sc, &shrink_memcg_cb,
- &encountered_page_in_swapcache);
+ &shrinking);
if (encountered_page_in_swapcache)
return SHRINK_STOP;
@@ -1332,8 +1343,11 @@ static struct shrinker *zswap_alloc_shrinker(void)
* were scanned but none could be written back, or -ENOENT if @memcg has
* writeback disabled, is a zombie cgroup, or has empty zswap LRUs.
*/
-static int shrink_memcg(struct mem_cgroup *memcg)
+static int shrink_memcg(struct mem_cgroup *memcg, bool may_enter_fs)
{
+ struct zswap_shrink_arg shrinking = {
+ .may_enter_fs = may_enter_fs,
+ };
int nid, shrunk = 0, scanned = 0;
if (!mem_cgroup_zswap_writeback_enabled(memcg))
@@ -1350,7 +1364,8 @@ static int shrink_memcg(struct mem_cgroup *memcg)
unsigned long nr_to_walk = SWAP_CLUSTER_MAX;
shrunk += list_lru_walk_one(&zswap_list_lru, nid, memcg,
- &shrink_memcg_cb, NULL, &nr_to_walk);
+ &shrink_memcg_cb, &shrinking,
+ &nr_to_walk);
scanned += SWAP_CLUSTER_MAX - nr_to_walk;
}
@@ -1427,7 +1442,7 @@ static void shrink_worker(struct work_struct *w)
goto resched;
}
- ret = shrink_memcg(memcg);
+ ret = shrink_memcg(memcg, true);
/* drop the extra reference */
mem_cgroup_put(memcg);
@@ -1544,7 +1559,7 @@ static bool zswap_store_page(struct folio *folio, long index,
return false;
}
-bool zswap_store(struct folio *folio)
+bool zswap_store(struct folio *folio, bool may_enter_fs)
{
long nr_pages = folio_nr_pages(folio);
swp_entry_t swp = folio->swap;
@@ -1563,7 +1578,7 @@ bool zswap_store(struct folio *folio)
objcg = get_obj_cgroup_from_folio(folio);
if (objcg && !obj_cgroup_may_zswap(objcg)) {
memcg = get_mem_cgroup_from_objcg(objcg);
- if (shrink_memcg(memcg)) {
+ if (shrink_memcg(memcg, may_enter_fs)) {
mem_cgroup_put(memcg);
goto put_objcg;
}
--
2.54.0
^ permalink raw reply related [flat|nested] 16+ messages in thread
* Re: [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II)
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
` (11 preceding siblings ...)
2026-10-03 0:59 ` [PATCH 12/12] mm, swap: skip a NOFS backend when reclaim cannot enter the fs Baoquan He
@ 2026-10-03 9:35 ` Nhat Pham
2026-10-03 10:07 ` Nhat Pham
2026-10-05 7:05 ` Baoquan He
12 siblings, 2 replies; 16+ messages in thread
From: Nhat Pham @ 2026-10-03 9:35 UTC (permalink / raw)
To: Baoquan He
Cc: linux-mm, akpm, chrisl, kasong, hannes, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin
On Sat, Oct 3, 2026 at 1:59 AM Baoquan He <hebaoquan@kylinos.cn> wrote:
>
> An xswap entry holds a page in zswap. That works until zswap will not
> take the page. Then the page has nowhere to go.
>
> This series gives an xswap entry a second home. When zswap refuses a page,
> the entry takes a slot on a real swap device and the data is written
> there. The xswap entry does not change when the data moves. It is what the
> process's PTE names, and the physical slot is recorded behind it.
>
> Phase II of three. Phase I is the device itself, 14 patches.
It's not cool to take my idea and code, modify a bit to fit your
design, then send it out without any proper attribution.
Adding a per-cluster array to store backend (in place of the existing
zswap tree) was my idea. The only difference is you put it in the
existing struct swap cluster, and I separate that array out into a new
vswap-only struct to separate the vswap-only metadata:
https://lore.kernel.org/all/20261003005900.909710-5-hebaoquan@kylinos.cn/
which, ironically, was also my proposal:
https://lore.kernel.org/all/CAKEwX=OvR7GbU_9f2h_MtU4m0g6s-esHmNQKYNhJz610M0P3Sw@mail.gmail.com/
And that's not the only one. I mean, you're too lazy to even rename
the macros of another mechanism I proposed:
https://lore.kernel.org/all/?q=SWP_RMAP_CACHE_ONLY
Yet, my name and my work is not mentioned or acknowledged at all in
this patch series, other than a small prep patch.
I have been very polite and deferential so far, trying my best to
accommodate your use case and innovation as best as I can, even go so
far as testing and debugging for you, because I want the best swap
virtualization scheme for the kernel:
https://lore.kernel.org/all/CAKEwX=Pe+qMZd2xhnU-PAGQtgXkp56c-JwYCbt2Lux9htgB67Q@mail.gmail.com/
But I believe this has crossed a line. This is a very unprofessional
and unethical. Be better.
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II)
2026-10-03 9:35 ` [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Nhat Pham
@ 2026-10-03 10:07 ` Nhat Pham
2026-10-05 7:05 ` Baoquan He
1 sibling, 0 replies; 16+ messages in thread
From: Nhat Pham @ 2026-10-03 10:07 UTC (permalink / raw)
To: Baoquan He
Cc: linux-mm, akpm, chrisl, kasong, hannes, baohua, youngjun.park,
david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin
On Sat, Oct 3, 2026 at 10:35 AM Nhat Pham <nphamcs@gmail.com> wrote:
>
> On Sat, Oct 3, 2026 at 1:59 AM Baoquan He <hebaoquan@kylinos.cn> wrote:
> >
> > An xswap entry holds a page in zswap. That works until zswap will not
> > take the page. Then the page has nowhere to go.
> >
> > This series gives an xswap entry a second home. When zswap refuses a page,
> > the entry takes a slot on a real swap device and the data is written
> > there. The xswap entry does not change when the data moves. It is what the
> > process's PTE names, and the physical slot is recorded behind it.
> >
> > Phase II of three. Phase I is the device itself, 14 patches.
>
> It's not cool to take my idea and code, modify a bit to fit your
> design, then send it out without any proper attribution.
>
> Adding a per-cluster array to store backend (in place of the existing
> zswap tree) was my idea. The only difference is you put it in the
Just for the record, here's me doing it since February of this year:
https://lore.kernel.org/all/20260208215839.87595-10-nphamcs@gmail.com/
https://lore.kernel.org/all/20260208215839.87595-15-nphamcs@gmail.com/
Look at struct vswap_cluster's and swp_desc's definitions. This design
predated the current swap table edition :)
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II)
2026-10-03 9:35 ` [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Nhat Pham
2026-10-03 10:07 ` Nhat Pham
@ 2026-10-05 7:05 ` Baoquan He
1 sibling, 0 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-05 7:05 UTC (permalink / raw)
To: Nhat Pham
Cc: Baoquan He, linux-mm, akpm, chrisl, kasong, hannes, baohua,
youngjun.park, david, kunwu.chan, gourry, riel, klarasmodin
On 10/03/26 at 10:35am, Nhat Pham wrote:
> On Sat, Oct 3, 2026 at 1:59 AM Baoquan He <hebaoquan@kylinos.cn> wrote:
> >
> > An xswap entry holds a page in zswap. That works until zswap will not
> > take the page. Then the page has nowhere to go.
> >
> > This series gives an xswap entry a second home. When zswap refuses a page,
> > the entry takes a slot on a real swap device and the data is written
> > there. The xswap entry does not change when the data moves. It is what the
> > process's PTE names, and the physical slot is recorded behind it.
> >
> > Phase II of three. Phase I is the device itself, 14 patches.
>
> It's not cool to take my idea and code, modify a bit to fit your
> design, then send it out without any proper attribution.
Nhat,
You have made this a public character judgement — "unprofessional and
unethical", "too lazy" — on top of a factual claim about attribution.
I will answer your factual claims completely, and I will correct the
things I did wrong. I will not reply with insults. But I will not leave
it on the record that I "took my idea and code, modify a bit, and sent
it out without attribution", because that is not what the postings show.
The v1 cover dropped the paragraph the RFC had. The RFC carried this
line:
"This partially refers to Nhat's vswap-v4 series. E.g patch 1 is
consistent with his patch 3, patch 12/13/14/15 are from his patch 8."
When I rewrote the cover for v1, I had just spent more than a week debugging
that series, and then another day on one intermittent failure that I could
not reproduce. I sent it out that morning and planned to keep testing after.
I simply forgot to copy that paragraph into the new cover. It was not a
deliberate attempt to remove your name. And I am happy now xswap phase II
should satisfy kashiko now.
Second, how xswap started. One time Kairui shared a presentation in public
about swap table and TODO list we can do. I got and idea that we can use swap
table pointer entry to point at zswap entry directly to improve efficiency.
After checking with Kairui I posted the Pointer-entry RFC. With that,
zswap got up to ~20% lower latency and ~10% higher throughput.
- https://lore.kernel.org/all/20260707073215.72183-1-baoquan.he@linux.dev/
Then I posted the ghost-swapfile RFC, which used lazy VM_SPARSE mapping for
the cluster_info[] array from the very beginning. Its cover says: "a sparse
64MB virtual address space is reserved at swapon time ... physical pages for
cluster_info[] structures are allocated only on demand". When I was working
on this part, your vswap was only RFC v1 of 5 patches, more than 2000 lines
of code adding, and RFC v2 of 7 patches, more than 2200 lines of code adding.
And your rfc v1 was not cc-ed to me.
- https://lore.kernel.org/all/20260707082614.95030-1-baoquan.he@linux.dev/
Third, the collaboration. When I posted ghost swapfile RFC, you came and
added comment. After discussion, You said let's solve one piece at a time,
and I took that seriously. After a week of careful consideration, I laid
out a bottom-up plan — the backend pointer in swap_cluster_info,
dynamic growth first, writeback and rmap after, swap tiering long term. I
did not get a specific technical objection from you.
- https://lore.kernel.org/all/al8ohWshSSZ64AtT@MiWiFi-R3L-srv/
Then I posted the dynamic cluster management foundation on July 27, and its
cover said explicitly:
"once the dynamic cluster management lands, Nhat's VS series can build
the per-slot backend pointer, writeback, and rmap on top of it without
an extra xarray lookup in the cluster access path."
- https://lore.kernel.org/all/20260727135029.1059441-1-baoquan.he@linux.dev/
I didn't touch writeback part and had been waiting for your. I left writeback
alone because the split was that you would do it. I asked Kairui to review
the foundation. He said he was busy with MGLRU and swap core at first, and
when he did look at it he called it "a really smart idea, really good job",
and gave detailed improvement suggestions. Other developers contacted me
privately, saying they were interested in xswap and wanted to help test,
improve or something, I welcomed all of it.
- Kairui: https://lore.kernel.org/all/apVNB2FRTFMDtvGP@KASONG-MC4/
On September 1 Kairui reviewed the xswap v1 series and praised the VM_SPARSE
foundation. On September 2 you replied to that same thread twice in public,
and then wrote to me off-list the same morning.
- Kairui: https://lore.kernel.org/all/apVNB2FRTFMDtvGP@KASONG-MC4/
- your first public reply:
https://lore.kernel.org/all/CAKEwX=MU9uXVenEu7he+h-Pq7K3BmCGvvKS-KiPK61Xd0ox_mQ@mail.gmail.com/
- your second public reply:
https://lore.kernel.org/all/CAKEwX=NNzX=sXdqPATrWw7cw7sdG-vOO5LqjH=JnV1YiZLO5MA@mail.gmail.com/
In that message you said you wanted to discuss how we could work together on
xswap and vswap, and you asked for a video call. I agreed. I replied with the
phase split: phase I the VM_SPARSE foundation, phase II writeback/rmap/zero-
page, phase III the zswap xarray removal, and my proposal that you take
phase II while I test and review it. And you can take phase III too if you like.
I also told you I did not like the xarray-based cluster array (struct
swap_cluster_info_dynamic) and the large, intrusive changes it made to the
swap core. You followed up in a friendly way. You said that struct was someone
else's code, that you thought the interface and the use case matter more than
the code design. I was working on an answer to that when, the next day, you
suddenly posted "Path forward for Virtualized Swap?" on the list and made
the whole thing public. To this day, I still do not understand why you reached
me to make a private conversation and got a friendly feedback, then immediately
raised a public mail to put me in the middle of a storm, with a lot of
your friends coming at me.
What happened next is the part I disliked most. The "Path forward for
Virtualized Swap?" thread filled up with comments on xswap's interface and
use cases — things that are often a one-line change — and almost nothing
on the core design and implementation.
Fourth, writeback. I had already implemented it by then; I held it back
because the split was yours. When the "no writeback" complaint kept coming, I
had to post the writeback RFC on September 20. Johannes then required the memcg
charging changes, so I added them. When you wrote in the v3 thread that
you tried very hard but found building writeback on the xswap foundation was
a lot of work, I told you not to worry, the RFC was posted, and you could
take over the writeback series. I never got a reply from you.
- my reply: https://lore.kernel.org/all/arDSdRDF9CnHjkE0@fedora/
So I did it myself: over the following week I fixed 11 regressions and could
not reproduce one intermittent failure in two days of continuous testing,
and then posted foundation v4, writeback v1 and memcg charging.
- writeback v1:
https://lore.kernel.org/all/20261003005900.909710-1-hebaoquan@kylinos.cn/
What happens next. I will restore the cover paragraph. If you want to drive
phase II, take it — adjust or rewrite the shared parts, including re-authoring
them. But I am not going to stay quiet while my name is called unprofessional
and unethical over a missing claim in a series that carries your own patch
under your Signed-off-by and whose RFC credited you. "Too lazy to rename a
macro" is not a technical review either.
If your point is that I missed attribution, I hear you, and I am fixing it.
If your point is that I am unethical, I do not accept it, and I want you to
stop saying it.
Baoquan
>
> Adding a per-cluster array to store backend (in place of the existing
> zswap tree) was my idea. The only difference is you put it in the
> existing struct swap cluster, and I separate that array out into a new
> vswap-only struct to separate the vswap-only metadata:
>
> https://lore.kernel.org/all/20261003005900.909710-5-hebaoquan@kylinos.cn/
>
> which, ironically, was also my proposal:
>
> https://lore.kernel.org/all/CAKEwX=OvR7GbU_9f2h_MtU4m0g6s-esHmNQKYNhJz610M0P3Sw@mail.gmail.com/
>
> And that's not the only one. I mean, you're too lazy to even rename
> the macros of another mechanism I proposed:
>
> https://lore.kernel.org/all/?q=SWP_RMAP_CACHE_ONLY
>
> Yet, my name and my work is not mentioned or acknowledged at all in
> this patch series, other than a small prep patch.
>
> I have been very polite and deferential so far, trying my best to
> accommodate your use case and innovation as best as I can, even go so
> far as testing and debugging for you, because I want the best swap
> virtualization scheme for the kernel:
>
> https://lore.kernel.org/all/CAKEwX=Pe+qMZd2xhnU-PAGQtgXkp56c-JwYCbt2Lux9htgB67Q@mail.gmail.com/
>
> But I believe this has crossed a line. This is a very unprofessional
> and unethical. Be better.
^ permalink raw reply [flat|nested] 16+ messages in thread
end of thread, other threads:[~2026-10-05 7:05 UTC | newest]
Thread overview: 16+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-03 0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
2026-10-03 0:58 ` [PATCH 01/12] mm, swap: prepare the swap IO path for xswap backends Baoquan He
2026-10-03 0:58 ` [PATCH 02/12] mm, swap: tag a swap table entry with its owning xswap entry Baoquan He
2026-10-03 0:58 ` [PATCH 03/12] mm, swap: prepare the folio-less allocation path for xswap Baoquan He
2026-10-03 0:58 ` [PATCH 04/12] mm, swap: add a physical backend for xswap slots Baoquan He
2026-10-03 0:58 ` [PATCH 05/12] mm, swap: use the xswap physical backend Baoquan He
2026-10-03 0:58 ` [PATCH 06/12] mm, swap: fall back to disk when zswap refuses an xswap page Baoquan He
2026-10-03 0:58 ` [PATCH 07/12] mm, swap: support swapoff of an xswap physical backend Baoquan He
2026-10-03 0:58 ` [PATCH 08/12] mm, swap: reclaim physical slots backing cache-only xswap entries Baoquan He
2026-10-03 0:58 ` [PATCH 09/12] mm, swap: back a large xswap folio with a contiguous physical run Baoquan He
2026-10-03 0:58 ` [PATCH 10/12] mm, swap: enable THP swapin for xswap entries Baoquan He
2026-10-03 0:58 ` [PATCH 11/12] mm, swap: drop swap_folio_sector() Baoquan He
2026-10-03 0:59 ` [PATCH 12/12] mm, swap: skip a NOFS backend when reclaim cannot enter the fs Baoquan He
2026-10-03 9:35 ` [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Nhat Pham
2026-10-03 10:07 ` Nhat Pham
2026-10-05 7:05 ` Baoquan He
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox