Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend
@ 2026-09-20  7:20 Baoquan He
  2026-09-20  7:20 ` [RFC PATCH 01/17] mm, swap: prepare the swap IO path for xswap backends Baoquan He
                   ` (18 more replies)
  0 siblings, 19 replies; 25+ messages in thread
From: Baoquan He @ 2026-09-20  7:20 UTC (permalink / raw)
  To: linux-mm
  Cc: akpm, chrisl, kasong, hannes, nphamcs, baohua, youngjun.park,
	david, kunwu.chan, baoquan.he, gourry, riel, mhocko,
	roman.gushchin, shakeel.butt, Baoquan He

This is the writeback layer for xswap, on top of the base series. Post it
as RFC for discussion.

The base series keeps every swapped-out page in zswap. So the pool must
refuse new pages once it's full. This series gives an xswap slot a
backend to go, and charge it only when it really goes to real disk
space.

Design                                                                                                
------                                                                                                
- Backend: when zswap refuses a page, take a slot on the real swap                                    
  device with the highest priority and write the page there. One IO per                               
  slot.                                                                                               
- Ownership: that slot has no swap cache folio, so record its owner in                                
  the swap table entry. 0b100 in the low bits marks a backend slot, and                               
  the xswap type/offset go above. The physical side then does not treat                               
  it as free, and readahead skips it.                                                                 
- Per-slot record: ci->xs_table[], one unsigned long per slot, allocated                              
  the first time a cluster takes a backend. Zero means no backend.                                    
- Read: a written-out slot has no zswap copy, so swap_read_folio()                                    
  follows the pointer and reads from the backend.                                                     
- Release: verify the slot still points back to the xswap entry first;                                
  the record can go stale.                                                                            
- swapoff: try_to_unuse() used to skip these slots. It now reads them                                 
  back, marks the folio dirty so reclaim re-stores it in zswap, then                                  
  drops the backend.                                                                                  
- Reclaim: an xswap entry's count can reach 0 while its folio is still                                
  in the swap cache; the physical scanner reclaims the redundant slot                                 
  through __try_to_reclaim_swap().                                                                    
- Large folios: a read starts at the first physical slot, so the run is                               
  reserved whole and must be backed contiguously; otherwise the fault                                 
  retries at a smaller order. xswap is SWP_SYNCHRONOUS_IO, which also                                 
  skips readahead.                                                                                    
- Charging: Record the owner at allocation but do not charge. Take the
  charge when the entry gets a physical backend slot, and give it back
  with that slot.memory.swap.current then counts only real on-disk swap
  usage, and a cgroup with memory.swap.max at 0 can still swap out through
  zswap.

Note                                                                                                  
----                                                                                                  
The base series is sereis of xswap foundation. This one depends on it.
So if anyone wants to apply this patchset, the order is base-commit
as below, then xswap foundation, finally this patchset.

[PATCH v3 00/14] mm, swap: extendable swap devices (xswap)
https://lore.kernel.org/all/20260916101929.149106-1-hebaoquan@kylinos.cn/T/#u

base-commit: baa8de2f3448d1466a888a805c18d01c998fe052

This partially refers to Nhat's vswap-v4 series. E.g patch 1 is
consistent with his patch 3, patch 12/13/14/15 are from his patch 8.

And the last 2 patches are fixing bugs when charging patches are added.
Since the charging of xswap is still under discussion, I dind't merge
them into commits. Will squash them or take them off once decision is
made.

Testing
-------
qemu KVM guest, 8G RAM.    
   
A 4G swap disk /dev/vdb is added as the physical backend, and zswap is
turned off so that every page has to reach a backend slot. Note that it
need create xswap device firstly then disable zswap, so every page has
to reach backend slot.

The workload is memhog, as in the base series. Every page is filled with
a fixed pattern, so the data can be checked after a readback.

  # echo 1 > /sys/module/zswap/parameters/enabled                                                     
  # echo 100 > /sys/kernel/mm/xswap/create                                                            
  # mkswap /dev/vdb && swapon -p 0 /dev/vdb                                                           
  # echo 0 > /sys/module/zswap/parameters/enabled                                                     
  # mkdir -p /sys/fs/cgroup/xswap_limit                                                               
  # echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max                                             
  # MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 ./memhog

1. Swapout to the backend

   Run the workload with 5G of memory in a cgroup capped at 4G:

   # awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
   Filename                Type        Size      Used      Priority
   xswap0                  xswap       8155132   2392176   100
   /dev/vdb                partition   4194300   2392176   0

   The two Used values are equal, so the pages reached the backend and
   the xswap side accounted for them.

2. Pages are read back from the device

   Lifting memory.max and touching the pages again reads them back. The
   pages-swapped-in counter went up by 265642, so the read did reach the
   device. 1010 of the folios read back were 1M, each read with one IO:

   # echo max > /sys/fs/cgroup/xswap_limit/memory.max
   # cat /sys/kernel/mm/transparent_hugepage/hugepages-1024kB/stats/swpin
   1010

   hugepages-2048kB/stats/swpin stays at 0, because the swapin order is
   capped below the PMD order.

3. verifies data correctness

   After the readback the data is compared byte for byte against the
   pattern, one time with zswap on and one time with zswap off, so the
   pages come from zswap in one time and from the backend in the other
   time.  All bytes matched.

4. Destroying a device returns its backend slots

   The device was filled, then destroyed with no readback first:

   # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
   3157220
   # echo 0 > /sys/kernel/mm/xswap/destroy
   # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
   0

5. verifies the charge is correct

   An xswap entry is charged only when it holds a backend slot, so the
   two counters have to show the same number of pages:

   # echo $(( $(cat /sys/fs/cgroup/xswap_limit/memory.swap.current) / 4096 ))
   526852
   # awk '$1 == "/dev/vdb" { print $4 / 4 }' /proc/swaps
   526852

   Nothing is charged while the pages stay in zswap: with the pool holding
   them and the backend untouched, memory.swap.current stays 0.

6. A cgroup with no swap room can still reclaim its anon memory

   With memory.swap.max at 0 and no xswap device, the cgroup ran out of
   room and the workload was killed. With an xswap device and zswap on,
   it swapped out through zswap, was charged nothing, and was not killed.

   E.g if memory.swap.max is set at 100M, the cgroup will be killed when
   the cap was reached.

Baoquan He (12):
  mm, swap: tag a swap table entry with its owning xswap entry
  mm, swap: prepare the folio-less allocation path for xswap
  mm, swap: add a physical backend for xswap slots
  mm, swap: use the xswap physical backend
  mm, swap: fall back to disk when zswap refuses an xswap page
  mm, swap: support swapoff of an xswap physical backend
  mm, swap: reclaim physical slots backing cache-only xswap entries
  mm, swap: back a large xswap folio with a contiguous physical run
  mm, swap: enable THP swapin for xswap entries
  mm, swap: drop swap_folio_sector()
  mm, swap: do not retake the cluster lock when uncharging an xswap slot
  mm, swap: drop a refused xswap backend run directly

Nhat Pham (5):
  mm, swap: prepare the swap IO path for xswap backends
  mm, swap: split the swap memcg charge helpers
  mm, swap: do not charge zswap-backed xswap entries
  mm, swap: charge an xswap entry when it gets physical backing
  mm, swap: don't gate xswap on the physical swap free count

 .../admin-guide/cgroup-v1/memcg_test.rst      |   2 +-
 include/linux/memcontrol.h                    |   6 +
 include/linux/swap.h                          |  69 +-
 include/linux/swap_ops.h                      |   9 +-
 mm/memcontrol-v1.c                            |  10 +-
 mm/memcontrol.c                               | 147 ++--
 mm/memory.c                                   |   7 +-
 mm/page_io.c                                  |  95 ++-
 mm/swap.h                                     |  35 +-
 mm/swap_state.c                               |  12 +-
 mm/swap_table.h                               |  53 ++
 mm/swapfile.c                                 | 742 ++++++++++++++++--
 mm/zswap.c                                    |  20 +-
 13 files changed, 1026 insertions(+), 181 deletions(-)

-- 
2.54.0



^ permalink raw reply	[flat|nested] 25+ messages in thread

end of thread, other threads:[~2026-09-28  3:31 UTC | newest]

Thread overview: 25+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-20  7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
2026-09-20  7:20 ` [RFC PATCH 01/17] mm, swap: prepare the swap IO path for xswap backends Baoquan He
2026-09-20  7:20 ` [RFC PATCH 02/17] mm, swap: tag a swap table entry with its owning xswap entry Baoquan He
2026-09-20  7:20 ` [RFC PATCH 03/17] mm, swap: prepare the folio-less allocation path for xswap Baoquan He
2026-09-20  7:20 ` [RFC PATCH 04/17] mm, swap: add a physical backend for xswap slots Baoquan He
2026-09-20  7:20 ` [RFC PATCH 05/17] mm, swap: use the xswap physical backend Baoquan He
2026-09-20  7:20 ` [RFC PATCH 06/17] mm, swap: fall back to disk when zswap refuses an xswap page Baoquan He
2026-09-20  7:20 ` [RFC PATCH 07/17] mm, swap: support swapoff of an xswap physical backend Baoquan He
2026-09-20  7:20 ` [RFC PATCH 08/17] mm, swap: reclaim physical slots backing cache-only xswap entries Baoquan He
2026-09-20  7:20 ` [RFC PATCH 09/17] mm, swap: back a large xswap folio with a contiguous physical run Baoquan He
2026-09-20  7:20 ` [RFC PATCH 10/17] mm, swap: enable THP swapin for xswap entries Baoquan He
2026-09-20  7:20 ` [RFC PATCH 11/17] mm, swap: drop swap_folio_sector() Baoquan He
2026-09-20  7:20 ` [RFC PATCH 12/17] mm, swap: split the swap memcg charge helpers Baoquan He
2026-09-20  7:20 ` [RFC PATCH 13/17] mm, swap: do not charge zswap-backed xswap entries Baoquan He
2026-09-21 11:54   ` Chris Li
2026-09-20  7:20 ` [RFC PATCH 14/17] mm, swap: charge an xswap entry when it gets physical backing Baoquan He
2026-09-20  7:20 ` [RFC PATCH 15/17] mm, swap: don't gate xswap on the physical swap free count Baoquan He
2026-09-20  7:20 ` [RFC PATCH 16/17] mm, swap: do not retake the cluster lock when uncharging an xswap slot Baoquan He
2026-09-20  7:20 ` [RFC PATCH 17/17] mm, swap: drop a refused xswap backend run directly Baoquan He
2026-09-22 14:22 ` [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Chris Li
2026-09-22 14:30   ` Johannes Weiner
2026-09-24  1:56   ` KunWu Chan
2026-09-24 11:36     ` Chris Li
2026-09-24 12:17 ` Klara Modin
2026-09-28  3:30   ` Baoquan He

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox