Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II)
@ 2026-10-03  0:58 Baoquan He
  2026-10-03  0:58 ` [PATCH 01/12] mm, swap: prepare the swap IO path for xswap backends Baoquan He
                   ` (12 more replies)
  0 siblings, 13 replies; 16+ messages in thread
From: Baoquan He @ 2026-10-03  0:58 UTC (permalink / raw)
  To: linux-mm
  Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
	david, kunwu.chan, baoquan.he, gourry, riel, klarasmodin,
	Baoquan He

An xswap entry holds a page in zswap. That works until zswap will not
take the page. Then the page has nowhere to go. 

This series gives an xswap entry a second home. When zswap refuses a page,
the entry takes a slot on a real swap device and the data is written
there. The xswap entry does not change when the data moves. It is what the
process's PTE names, and the physical slot is recorded behind it.

Phase II of three. Phase I is the device itself, 14 patches.

Design notes
------------

An xswap slot and a physical slot are different things with different
owners, and the physical side has to find its xswap entry again. A swap
table entry already carries its type in the low three bits, and 0b100 was
unused. It now points at the owner, with the xswap type and offset above.
The places that hand out or cache a swap table entry learn to skip a
pointer entry.

folio->swap no longer says where the IO goes: a folio whose data is on the
backend still sits in the swap cache of its own xswap entry. So the
batched IO path carries the entry explicitly, instead of deriving the
device and sector from folio->swap. A large folio is read back with one
IO starting at its first slot, so its slots have to be contiguous; the
reserved-run allocation provides that.

An xswap entry's swap count can reach zero while its folio is still in
the swap cache. A page faulted back in during its write to the backend
ends up there: do_swap_page() does not wait for writeback, folio_free_swap()
refuses a folio under writeback, and a write to a real device does not drop
the cache when it completes. The physical slot is then redundant, but still
held. A single bit in the slot's swap table entry marks this, and the
physical reclaim scanner frees it. The scanner reads the bit directly;
taking the xswap cluster lock would invert the lock order.

An xswap entry that is still in zswap costs no swap space, so it is not
charged. Charging is phase III. Here the entry is only recorded.

Testing
-------

qemu KVM guest, 8G RAM. memhog is a small local helper: it faults
<total_gb> of anon, fills it with a fixed pattern, and holds it.

A 4G swap disk /dev/vdb is the backend. The xswap device has to exist
before zswap is turned off, because creating one requires zswap. The
number written to create is the device's swap priority; 100 puts it above
the backend at 0.

  # echo 1 > /sys/module/zswap/parameters/enabled
  # echo 100 > /sys/kernel/mm/xswap/create
  # mkswap /dev/vdb && swapon -p 0 /dev/vdb
  # echo 0 > /sys/module/zswap/parameters/enabled
  # mkdir -p /sys/fs/cgroup/xswap_limit
  # echo 4G > /sys/fs/cgroup/xswap_limit/memory.max
  # echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
  # ( echo $BASHPID > /sys/fs/cgroup/xswap_limit/cgroup.procs
  #   exec env MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 \
  #       ./memhog 5 600 ) &

The pages reach the backend, and the xswap side accounts for them:

  # awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
  Filename                Type        Size      Used      Priority
  xswap0                  xswap       8155132   2392176   100
  /dev/vdb                partition   4194300   2392176   0

The two Used values are equal. An xswap slot and its backend slot each
hold one reference to the same page while it is swapped out. It is one
page of data, not two, and the backend slot is given back when the xswap
entry goes.

With the workload still alive, destroying the device returns every
backend slot:

  # echo 0 > /sys/kernel/mm/xswap/destroy
  # sleep 5; awk 'NR == 1 || $1 ~ /vdb/' /proc/swaps
  Filename                Type        Size      Used      Priority
  /dev/vdb                partition   4194300   0         0

Also tested these, and nothing crashed or warned:

  - ten create/fill/destroy cycles, each returning /dev/vdb to 0
  - a shrink racing a swapoff of the same device, 20 rounds
  - a swapoff of a device that has pages on the backend
  - THP swapin: a large folio read back as one IO, and byte for byte
    what was written

Changelog
=========
RFC -> v1:

- Charging is phase III, so its six patches are out. An entry is
  recorded here and not charged, and the RFC's two charge tests go with
  them.

- New patch 12: skip a NOFS backend when reclaim cannot enter the fs.
  may_enter_fs() sees the xswap device, not the backend.

- Cache-only reclaim marks the slot with SWP_RMAP_CACHE_ONLY instead of
  handing it to __try_to_reclaim_swap().

- A large folio is only written to a block device backend, and the
  backend is recorded before the reverse mapping that finds it.

- An xswap entry goes back on the zswap writeback LRU.

- swapoff reads an entry back under memalloc_noreclaim_save() and
  xswap_lock.

- Patches 1 to 11 keep their subjects and order, rebased onto the new
  phase I, with rewritten commit messages.

Baoquan He (11):
  mm, swap: tag a swap table entry with its owning xswap entry
  mm, swap: prepare the folio-less allocation path for xswap
  mm, swap: add a physical backend for xswap slots
  mm, swap: use the xswap physical backend
  mm, swap: fall back to disk when zswap refuses an xswap page
  mm, swap: support swapoff of an xswap physical backend
  mm, swap: reclaim physical slots backing cache-only xswap entries
  mm, swap: back a large xswap folio with a contiguous physical run
  mm, swap: enable THP swapin for xswap entries
  mm, swap: drop swap_folio_sector()
  mm, swap: skip a NOFS backend when reclaim cannot enter the fs

Nhat Pham (1):
  mm, swap: prepare the swap IO path for xswap backends

 include/linux/swap.h     |   2 +-
 include/linux/swap_ops.h |  10 +-
 include/linux/zswap.h    |   4 +-
 mm/memory.c              |   7 +-
 mm/page_io.c             | 102 +++--
 mm/swap.h                |  37 +-
 mm/swap_state.c          |  12 +-
 mm/swap_table.h          |  92 ++++-
 mm/swapfile.c            | 783 ++++++++++++++++++++++++++++++++++++---
 mm/vmscan.c              |   4 +-
 mm/zswap.c               |  61 ++-
 11 files changed, 989 insertions(+), 125 deletions(-)


base-commit: 128a882559e489eea535bf31b7293058708d351b
-- 
2.54.0



^ permalink raw reply	[flat|nested] 16+ messages in thread

end of thread, other threads:[~2026-10-05  7:05 UTC | newest]

Thread overview: 16+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-03  0:58 [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Baoquan He
2026-10-03  0:58 ` [PATCH 01/12] mm, swap: prepare the swap IO path for xswap backends Baoquan He
2026-10-03  0:58 ` [PATCH 02/12] mm, swap: tag a swap table entry with its owning xswap entry Baoquan He
2026-10-03  0:58 ` [PATCH 03/12] mm, swap: prepare the folio-less allocation path for xswap Baoquan He
2026-10-03  0:58 ` [PATCH 04/12] mm, swap: add a physical backend for xswap slots Baoquan He
2026-10-03  0:58 ` [PATCH 05/12] mm, swap: use the xswap physical backend Baoquan He
2026-10-03  0:58 ` [PATCH 06/12] mm, swap: fall back to disk when zswap refuses an xswap page Baoquan He
2026-10-03  0:58 ` [PATCH 07/12] mm, swap: support swapoff of an xswap physical backend Baoquan He
2026-10-03  0:58 ` [PATCH 08/12] mm, swap: reclaim physical slots backing cache-only xswap entries Baoquan He
2026-10-03  0:58 ` [PATCH 09/12] mm, swap: back a large xswap folio with a contiguous physical run Baoquan He
2026-10-03  0:58 ` [PATCH 10/12] mm, swap: enable THP swapin for xswap entries Baoquan He
2026-10-03  0:58 ` [PATCH 11/12] mm, swap: drop swap_folio_sector() Baoquan He
2026-10-03  0:59 ` [PATCH 12/12] mm, swap: skip a NOFS backend when reclaim cannot enter the fs Baoquan He
2026-10-03  9:35 ` [PATCH 00/12] mm, swap: give xswap a physical backend (xswap phase II) Nhat Pham
2026-10-03 10:07   ` Nhat Pham
2026-10-05  7:05   ` Baoquan He

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox