Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Klara Modin <klarasmodin@gmail.com>
To: Baoquan He <hebaoquan@kylinos.cn>
Cc: linux-mm@kvack.org, akpm@linux-foundation.org, chrisl@kernel.org,
	 kasong@tencent.com, hannes@cmpxchg.org, nphamcs@gmail.com,
	baohua@kernel.org,  youngjun.park@lge.com, david@kernel.org,
	kunwu.chan@gmail.com, baoquan.he@linux.dev,  gourry@gourry.net,
	riel@surriel.com, mhocko@kernel.org, roman.gushchin@linux.dev,
	 shakeel.butt@linux.dev
Subject: Re: [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend
Date: Thu, 24 Sep 2026 14:17:10 +0200	[thread overview]
Message-ID: <arUKSOVFzPF-iHxU@parmesan.int.kasm.eu> (raw)
In-Reply-To: <20260920072043.430390-1-hebaoquan@kylinos.cn>

On 2026-09-20 15:20:26 +0800, Baoquan He wrote:
> This is the writeback layer for xswap, on top of the base series. Post it
> as RFC for discussion.
> 
> The base series keeps every swapped-out page in zswap. So the pool must
> refuse new pages once it's full. This series gives an xswap slot a
> backend to go, and charge it only when it really goes to real disk
> space.

There is some weird trailing whitespace in this email for some sections.

> 
> Design                                                                                                
> ------                                                                                                
> - Backend: when zswap refuses a page, take a slot on the real swap                                    
>   device with the highest priority and write the page there. One IO per                               
>   slot.                                                                                               
> - Ownership: that slot has no swap cache folio, so record its owner in                                
>   the swap table entry. 0b100 in the low bits marks a backend slot, and                               
>   the xswap type/offset go above. The physical side then does not treat                               
>   it as free, and readahead skips it.                                                                 
> - Per-slot record: ci->xs_table[], one unsigned long per slot, allocated                              
>   the first time a cluster takes a backend. Zero means no backend.                                    
> - Read: a written-out slot has no zswap copy, so swap_read_folio()                                    
>   follows the pointer and reads from the backend.                                                     
> - Release: verify the slot still points back to the xswap entry first;                                
>   the record can go stale.                                                                            
> - swapoff: try_to_unuse() used to skip these slots. It now reads them                                 
>   back, marks the folio dirty so reclaim re-stores it in zswap, then                                  
>   drops the backend.                                                                                  
> - Reclaim: an xswap entry's count can reach 0 while its folio is still                                
>   in the swap cache; the physical scanner reclaims the redundant slot                                 
>   through __try_to_reclaim_swap().                                                                    
> - Large folios: a read starts at the first physical slot, so the run is                               
>   reserved whole and must be backed contiguously; otherwise the fault                                 
>   retries at a smaller order. xswap is SWP_SYNCHRONOUS_IO, which also                                 
>   skips readahead.                                                                                    
> - Charging: Record the owner at allocation but do not charge. Take the
>   charge when the entry gets a physical backend slot, and give it back
>   with that slot.memory.swap.current then counts only real on-disk swap
>   usage, and a cgroup with memory.swap.max at 0 can still swap out through
>   zswap.

> 
> Note                                                                                                  
> ----                                                                                                  
> The base series is sereis of xswap foundation. This one depends on it.
> So if anyone wants to apply this patchset, the order is base-commit
> as below, then xswap foundation, finally this patchset.
> 
> [PATCH v3 00/14] mm, swap: extendable swap devices (xswap)
> https://lore.kernel.org/all/20260916101929.149106-1-hebaoquan@kylinos.cn/T/#u
> 
> base-commit: baa8de2f3448d1466a888a805c18d01c998fe052
> 
> This partially refers to Nhat's vswap-v4 series. E.g patch 1 is
> consistent with his patch 3, patch 12/13/14/15 are from his patch 8.
> 
> And the last 2 patches are fixing bugs when charging patches are added.
> Since the charging of xswap is still under discussion, I dind't merge
> them into commits. Will squash them or take them off once decision is
> made.
> 
> Testing
> -------
> qemu KVM guest, 8G RAM.    
>    
> A 4G swap disk /dev/vdb is added as the physical backend, and zswap is
> turned off so that every page has to reach a backend slot. Note that it
> need create xswap device firstly then disable zswap, so every page has
> to reach backend slot.
> 
> The workload is memhog, as in the base series. Every page is filled with
> a fixed pattern, so the data can be checked after a readback.
> 
>   # echo 1 > /sys/module/zswap/parameters/enabled                                                     
>   # echo 100 > /sys/kernel/mm/xswap/create                                                            
>   # mkswap /dev/vdb && swapon -p 0 /dev/vdb                                                           
>   # echo 0 > /sys/module/zswap/parameters/enabled                                                     
>   # mkdir -p /sys/fs/cgroup/xswap_limit                                                               
>   # echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max                                             
>   # MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 ./memhog

You did not include e.g. cgexec here, so is memhog really running in the
cgroup you created?

> 
> 1. Swapout to the backend
> 
>    Run the workload with 5G of memory in a cgroup capped at 4G:
> 
>    # awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
>    Filename                Type        Size      Used      Priority
>    xswap0                  xswap       8155132   2392176   100
>    /dev/vdb                partition   4194300   2392176   0
> 
>    The two Used values are equal, so the pages reached the backend and
>    the xswap side accounted for them.

This is a bit confusing to me. If I now run free, the written back data
will be represented twice in the sum? Since the written back data is
still presented as used in the xswap device, does it still consume space
there, i.e. does it prevent that amount of new data being stored in the
xswap device? If this is the case, it would make it even harder to
dimension the xswap size since one would also have to take into account
all the other swap devices which are present in the system.

> 
> 2. Pages are read back from the device
> 
>    Lifting memory.max and touching the pages again reads them back. The
>    pages-swapped-in counter went up by 265642, so the read did reach the
>    device. 1010 of the folios read back were 1M, each read with one IO:
> 
>    # echo max > /sys/fs/cgroup/xswap_limit/memory.max
>    # cat /sys/kernel/mm/transparent_hugepage/hugepages-1024kB/stats/swpin
>    1010
> 
>    hugepages-2048kB/stats/swpin stays at 0, because the swapin order is
>    capped below the PMD order.
> 
> 3. verifies data correctness
> 
>    After the readback the data is compared byte for byte against the
>    pattern, one time with zswap on and one time with zswap off, so the
>    pages come from zswap in one time and from the backend in the other
>    time.  All bytes matched.
> 
> 4. Destroying a device returns its backend slots
> 
>    The device was filled, then destroyed with no readback first:
> 
>    # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
>    3157220
>    # echo 0 > /sys/kernel/mm/xswap/destroy
>    # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
>    0
> 
> 5. verifies the charge is correct
> 
>    An xswap entry is charged only when it holds a backend slot, so the
>    two counters have to show the same number of pages:
> 
>    # echo $(( $(cat /sys/fs/cgroup/xswap_limit/memory.swap.current) / 4096 ))
>    526852
>    # awk '$1 == "/dev/vdb" { print $4 / 4 }' /proc/swaps
>    526852
> 
>    Nothing is charged while the pages stay in zswap: with the pool holding
>    them and the backend untouched, memory.swap.current stays 0.
> 
> 6. A cgroup with no swap room can still reclaim its anon memory
> 
>    With memory.swap.max at 0 and no xswap device, the cgroup ran out of
>    room and the workload was killed. With an xswap device and zswap on,
>    it swapped out through zswap, was charged nothing, and was not killed.
> 
>    E.g if memory.swap.max is set at 100M, the cgroup will be killed when
>    the cap was reached.

Disregarding my comments and opinons on the interface, it seems to be
working so far and I can see writeback happen during a GCC 17 build on
my BPI-F3, but the entire thing takes roughly 18 hours and is not done
yet.

> 
> Baoquan He (12):
>   mm, swap: tag a swap table entry with its owning xswap entry
>   mm, swap: prepare the folio-less allocation path for xswap
>   mm, swap: add a physical backend for xswap slots
>   mm, swap: use the xswap physical backend
>   mm, swap: fall back to disk when zswap refuses an xswap page
>   mm, swap: support swapoff of an xswap physical backend
>   mm, swap: reclaim physical slots backing cache-only xswap entries
>   mm, swap: back a large xswap folio with a contiguous physical run
>   mm, swap: enable THP swapin for xswap entries
>   mm, swap: drop swap_folio_sector()
>   mm, swap: do not retake the cluster lock when uncharging an xswap slot
>   mm, swap: drop a refused xswap backend run directly
> 
> Nhat Pham (5):
>   mm, swap: prepare the swap IO path for xswap backends
>   mm, swap: split the swap memcg charge helpers
>   mm, swap: do not charge zswap-backed xswap entries
>   mm, swap: charge an xswap entry when it gets physical backing
>   mm, swap: don't gate xswap on the physical swap free count
> 
>  .../admin-guide/cgroup-v1/memcg_test.rst      |   2 +-
>  include/linux/memcontrol.h                    |   6 +
>  include/linux/swap.h                          |  69 +-
>  include/linux/swap_ops.h                      |   9 +-
>  mm/memcontrol-v1.c                            |  10 +-
>  mm/memcontrol.c                               | 147 ++--
>  mm/memory.c                                   |   7 +-
>  mm/page_io.c                                  |  95 ++-
>  mm/swap.h                                     |  35 +-
>  mm/swap_state.c                               |  12 +-
>  mm/swap_table.h                               |  53 ++
>  mm/swapfile.c                                 | 742 ++++++++++++++++--
>  mm/zswap.c                                    |  20 +-
>  13 files changed, 1026 insertions(+), 181 deletions(-)
> 
> -- 
> 2.54.0
> 
> 


  parent reply	other threads:[~2026-09-24 12:17 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-20  7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
2026-09-20  7:20 ` [RFC PATCH 01/17] mm, swap: prepare the swap IO path for xswap backends Baoquan He
2026-09-20  7:20 ` [RFC PATCH 02/17] mm, swap: tag a swap table entry with its owning xswap entry Baoquan He
2026-09-20  7:20 ` [RFC PATCH 03/17] mm, swap: prepare the folio-less allocation path for xswap Baoquan He
2026-09-20  7:20 ` [RFC PATCH 04/17] mm, swap: add a physical backend for xswap slots Baoquan He
2026-09-20  7:20 ` [RFC PATCH 05/17] mm, swap: use the xswap physical backend Baoquan He
2026-09-20  7:20 ` [RFC PATCH 06/17] mm, swap: fall back to disk when zswap refuses an xswap page Baoquan He
2026-09-20  7:20 ` [RFC PATCH 07/17] mm, swap: support swapoff of an xswap physical backend Baoquan He
2026-09-20  7:20 ` [RFC PATCH 08/17] mm, swap: reclaim physical slots backing cache-only xswap entries Baoquan He
2026-09-20  7:20 ` [RFC PATCH 09/17] mm, swap: back a large xswap folio with a contiguous physical run Baoquan He
2026-09-20  7:20 ` [RFC PATCH 10/17] mm, swap: enable THP swapin for xswap entries Baoquan He
2026-09-20  7:20 ` [RFC PATCH 11/17] mm, swap: drop swap_folio_sector() Baoquan He
2026-09-20  7:20 ` [RFC PATCH 12/17] mm, swap: split the swap memcg charge helpers Baoquan He
2026-09-20  7:20 ` [RFC PATCH 13/17] mm, swap: do not charge zswap-backed xswap entries Baoquan He
2026-09-21 11:54   ` Chris Li
2026-09-20  7:20 ` [RFC PATCH 14/17] mm, swap: charge an xswap entry when it gets physical backing Baoquan He
2026-09-20  7:20 ` [RFC PATCH 15/17] mm, swap: don't gate xswap on the physical swap free count Baoquan He
2026-09-20  7:20 ` [RFC PATCH 16/17] mm, swap: do not retake the cluster lock when uncharging an xswap slot Baoquan He
2026-09-20  7:20 ` [RFC PATCH 17/17] mm, swap: drop a refused xswap backend run directly Baoquan He
2026-09-22 14:22 ` [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Chris Li
2026-09-22 14:30   ` Johannes Weiner
2026-09-24  1:56   ` KunWu Chan
2026-09-24 11:36     ` Chris Li
2026-09-24 12:17 ` Klara Modin [this message]
2026-09-28  3:30   ` Baoquan He

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arUKSOVFzPF-iHxU@parmesan.int.kasm.eu \
    --to=klarasmodin@gmail.com \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baoquan.he@linux.dev \
    --cc=chrisl@kernel.org \
    --cc=david@kernel.org \
    --cc=gourry@gourry.net \
    --cc=hannes@cmpxchg.org \
    --cc=hebaoquan@kylinos.cn \
    --cc=kasong@tencent.com \
    --cc=kunwu.chan@gmail.com \
    --cc=linux-mm@kvack.org \
    --cc=mhocko@kernel.org \
    --cc=nphamcs@gmail.com \
    --cc=riel@surriel.com \
    --cc=roman.gushchin@linux.dev \
    --cc=shakeel.butt@linux.dev \
    --cc=youngjun.park@lge.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox