Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Baoquan He <baoquan.he@linux.dev>
To: Klara Modin <klarasmodin@gmail.com>
Cc: Baoquan He <hebaoquan@kylinos.cn>,
	linux-mm@kvack.org, akpm@linux-foundation.org, chrisl@kernel.org,
	kasong@tencent.com, hannes@cmpxchg.org, nphamcs@gmail.com,
	baohua@kernel.org, youngjun.park@lge.com, david@kernel.org,
	kunwu.chan@gmail.com, gourry@gourry.net, riel@surriel.com,
	mhocko@kernel.org, roman.gushchin@linux.dev,
	shakeel.butt@linux.dev
Subject: Re: [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend
Date: Mon, 28 Sep 2026 11:30:59 +0800	[thread overview]
Message-ID: <arnfc8oXEdCJrWEw@MiWiFi-R3L-srv> (raw)
In-Reply-To: <arUKSOVFzPF-iHxU@parmesan.int.kasm.eu>

On 09/24/26 at 02:17pm, Klara Modin wrote:
> On 2026-09-20 15:20:26 +0800, Baoquan He wrote:
> > This is the writeback layer for xswap, on top of the base series. Post it
> > as RFC for discussion.
> > 
> > The base series keeps every swapped-out page in zswap. So the pool must
> > refuse new pages once it's full. This series gives an xswap slot a
> > backend to go, and charge it only when it really goes to real disk
> > space.
> 
> There is some weird trailing whitespace in this email for some sections.

Thanks for careful checking, will clean this up before sending out.

...snip... 
> > Testing
> > -------
> > qemu KVM guest, 8G RAM.    
> >    
> > A 4G swap disk /dev/vdb is added as the physical backend, and zswap is
> > turned off so that every page has to reach a backend slot. Note that it
> > need create xswap device firstly then disable zswap, so every page has
> > to reach backend slot.
> > 
> > The workload is memhog, as in the base series. Every page is filled with
> > a fixed pattern, so the data can be checked after a readback.
> > 
> >   # echo 1 > /sys/module/zswap/parameters/enabled                                                     
> >   # echo 100 > /sys/kernel/mm/xswap/create                                                            
> >   # mkswap /dev/vdb && swapon -p 0 /dev/vdb                                                           
> >   # echo 0 > /sys/module/zswap/parameters/enabled                                                     
> >   # mkdir -p /sys/fs/cgroup/xswap_limit                                                               
> >   # echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max                                             
> >   # MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 ./memhog
> 
> You did not include e.g. cgexec here, so is memhog really running in the
> cgroup you created?

You are right, I didn't write out the full command, just list the steps.
And here memhog is not a standard one from installed rpm package
numactl, it's a small local helper. It faults <total_GB> of anon, holds
it, and can drop the mapping or read it back. And the helper keep a known
byte pattern, so the data can be checked byte for byte after it is read
back.

# echo 1 > /sys/module/zswap/parameters/enabled
# echo 100 > /sys/kernel/mm/xswap/create
# mkswap /dev/vdb && swapon -p 0 /dev/vdb
# echo 0 > /sys/module/zswap/parameters/enabled
# mkdir -p /sys/fs/cgroup/xswap_limit
# echo 4G > /sys/fs/cgroup/xswap_limit/memory.max
# echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
# ( echo $BASHPID > /sys/fs/cgroup/xswap_limit/cgroup.procs
#   exec env MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 \
#       /mnt/kernel_src/testing_kernel_hebq/xswap/memhog 5 600 ) &

I will add the complete block in next version, sorry for the
inconvenience.

> 
> > 
> > 1. Swapout to the backend
> > 
> >    Run the workload with 5G of memory in a cgroup capped at 4G:
> > 
> >    # awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
> >    Filename                Type        Size      Used      Priority
> >    xswap0                  xswap       8155132   2392176   100
> >    /dev/vdb                partition   4194300   2392176   0
> > 
> >    The two Used values are equal, so the pages reached the backend and
> >    the xswap side accounted for them.
> 
> This is a bit confusing to me. If I now run free, the written back data
> will be represented twice in the sum? Since the written back data is
> still presented as used in the xswap device, does it still consume space
> there, i.e. does it prevent that amount of new data being stored in the
> xswap device? If this is the case, it would make it even harder to
> dimension the xswap size since one would also have to take into account
> all the other swap devices which are present in the system.

I have to admit you found out an issue, thanks. I noticed it too while
I havne't thought of a satisfying way to fix it. What I can think of
is to mark the usage size of writeback on disk and list it separately. 

# awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
Filename                Type        Size      Used            Priority
xswap0                  xswap       8155132   2392176         100
/dev/vdb                partition   4194300   2392176         0
                                              2392176 (xswap) 

If xswap is disabled and full, /dev/vdb got swapped out data by its own,
then the amount of used slot will be reflected as before.

# awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
Filename                Type        Size      Used            Priority
xswap0                  xswap       8155132   2392176         100
/dev/vdb                partition   4194300   3000000         0
                                              2392176 (xswap) 
I think this can be fixed later, I see it as a non-blocker for now.
	
> 
> > 
> > 2. Pages are read back from the device
> > 
> >    Lifting memory.max and touching the pages again reads them back. The
> >    pages-swapped-in counter went up by 265642, so the read did reach the
> >    device. 1010 of the folios read back were 1M, each read with one IO:
> > 
> >    # echo max > /sys/fs/cgroup/xswap_limit/memory.max
> >    # cat /sys/kernel/mm/transparent_hugepage/hugepages-1024kB/stats/swpin
> >    1010
> > 
> >    hugepages-2048kB/stats/swpin stays at 0, because the swapin order is
> >    capped below the PMD order.
> > 
> > 3. verifies data correctness
> > 
> >    After the readback the data is compared byte for byte against the
> >    pattern, one time with zswap on and one time with zswap off, so the
> >    pages come from zswap in one time and from the backend in the other
> >    time.  All bytes matched.
> > 
> > 4. Destroying a device returns its backend slots
> > 
> >    The device was filled, then destroyed with no readback first:
> > 
> >    # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
> >    3157220
> >    # echo 0 > /sys/kernel/mm/xswap/destroy
> >    # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
> >    0
> > 
> > 5. verifies the charge is correct
> > 
> >    An xswap entry is charged only when it holds a backend slot, so the
> >    two counters have to show the same number of pages:
> > 
> >    # echo $(( $(cat /sys/fs/cgroup/xswap_limit/memory.swap.current) / 4096 ))
> >    526852
> >    # awk '$1 == "/dev/vdb" { print $4 / 4 }' /proc/swaps
> >    526852
> > 
> >    Nothing is charged while the pages stay in zswap: with the pool holding
> >    them and the backend untouched, memory.swap.current stays 0.
> > 
> > 6. A cgroup with no swap room can still reclaim its anon memory
> > 
> >    With memory.swap.max at 0 and no xswap device, the cgroup ran out of
> >    room and the workload was killed. With an xswap device and zswap on,
> >    it swapped out through zswap, was charged nothing, and was not killed.
> > 
> >    E.g if memory.swap.max is set at 100M, the cgroup will be killed when
> >    the cap was reached.
> 
> Disregarding my comments and opinons on the interface, it seems to be
> working so far and I can see writeback happen during a GCC 17 build on
> my BPI-F3, but the entire thing takes roughly 18 hours and is not done
> yet.

Thanks again for your careful checking, testing.

Thanks
Baoquan

> 
> > 
> > Baoquan He (12):
> >   mm, swap: tag a swap table entry with its owning xswap entry
> >   mm, swap: prepare the folio-less allocation path for xswap
> >   mm, swap: add a physical backend for xswap slots
> >   mm, swap: use the xswap physical backend
> >   mm, swap: fall back to disk when zswap refuses an xswap page
> >   mm, swap: support swapoff of an xswap physical backend
> >   mm, swap: reclaim physical slots backing cache-only xswap entries
> >   mm, swap: back a large xswap folio with a contiguous physical run
> >   mm, swap: enable THP swapin for xswap entries
> >   mm, swap: drop swap_folio_sector()
> >   mm, swap: do not retake the cluster lock when uncharging an xswap slot
> >   mm, swap: drop a refused xswap backend run directly
> > 
> > Nhat Pham (5):
> >   mm, swap: prepare the swap IO path for xswap backends
> >   mm, swap: split the swap memcg charge helpers
> >   mm, swap: do not charge zswap-backed xswap entries
> >   mm, swap: charge an xswap entry when it gets physical backing
> >   mm, swap: don't gate xswap on the physical swap free count
> > 
> >  .../admin-guide/cgroup-v1/memcg_test.rst      |   2 +-
> >  include/linux/memcontrol.h                    |   6 +
> >  include/linux/swap.h                          |  69 +-
> >  include/linux/swap_ops.h                      |   9 +-
> >  mm/memcontrol-v1.c                            |  10 +-
> >  mm/memcontrol.c                               | 147 ++--
> >  mm/memory.c                                   |   7 +-
> >  mm/page_io.c                                  |  95 ++-
> >  mm/swap.h                                     |  35 +-
> >  mm/swap_state.c                               |  12 +-
> >  mm/swap_table.h                               |  53 ++
> >  mm/swapfile.c                                 | 742 ++++++++++++++++--
> >  mm/zswap.c                                    |  20 +-
> >  13 files changed, 1026 insertions(+), 181 deletions(-)
> > 
> > -- 
> > 2.54.0
> > 
> > 


      reply	other threads:[~2026-09-28  3:31 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-20  7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
2026-09-20  7:20 ` [RFC PATCH 01/17] mm, swap: prepare the swap IO path for xswap backends Baoquan He
2026-09-20  7:20 ` [RFC PATCH 02/17] mm, swap: tag a swap table entry with its owning xswap entry Baoquan He
2026-09-20  7:20 ` [RFC PATCH 03/17] mm, swap: prepare the folio-less allocation path for xswap Baoquan He
2026-09-20  7:20 ` [RFC PATCH 04/17] mm, swap: add a physical backend for xswap slots Baoquan He
2026-09-20  7:20 ` [RFC PATCH 05/17] mm, swap: use the xswap physical backend Baoquan He
2026-09-20  7:20 ` [RFC PATCH 06/17] mm, swap: fall back to disk when zswap refuses an xswap page Baoquan He
2026-09-20  7:20 ` [RFC PATCH 07/17] mm, swap: support swapoff of an xswap physical backend Baoquan He
2026-09-20  7:20 ` [RFC PATCH 08/17] mm, swap: reclaim physical slots backing cache-only xswap entries Baoquan He
2026-09-20  7:20 ` [RFC PATCH 09/17] mm, swap: back a large xswap folio with a contiguous physical run Baoquan He
2026-09-20  7:20 ` [RFC PATCH 10/17] mm, swap: enable THP swapin for xswap entries Baoquan He
2026-09-20  7:20 ` [RFC PATCH 11/17] mm, swap: drop swap_folio_sector() Baoquan He
2026-09-20  7:20 ` [RFC PATCH 12/17] mm, swap: split the swap memcg charge helpers Baoquan He
2026-09-20  7:20 ` [RFC PATCH 13/17] mm, swap: do not charge zswap-backed xswap entries Baoquan He
2026-09-21 11:54   ` Chris Li
2026-09-20  7:20 ` [RFC PATCH 14/17] mm, swap: charge an xswap entry when it gets physical backing Baoquan He
2026-09-20  7:20 ` [RFC PATCH 15/17] mm, swap: don't gate xswap on the physical swap free count Baoquan He
2026-09-20  7:20 ` [RFC PATCH 16/17] mm, swap: do not retake the cluster lock when uncharging an xswap slot Baoquan He
2026-09-20  7:20 ` [RFC PATCH 17/17] mm, swap: drop a refused xswap backend run directly Baoquan He
2026-09-22 14:22 ` [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Chris Li
2026-09-22 14:30   ` Johannes Weiner
2026-09-24  1:56   ` KunWu Chan
2026-09-24 11:36     ` Chris Li
2026-09-24 12:17 ` Klara Modin
2026-09-28  3:30   ` Baoquan He [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arnfc8oXEdCJrWEw@MiWiFi-R3L-srv \
    --to=baoquan.he@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=chrisl@kernel.org \
    --cc=david@kernel.org \
    --cc=gourry@gourry.net \
    --cc=hannes@cmpxchg.org \
    --cc=hebaoquan@kylinos.cn \
    --cc=kasong@tencent.com \
    --cc=klarasmodin@gmail.com \
    --cc=kunwu.chan@gmail.com \
    --cc=linux-mm@kvack.org \
    --cc=mhocko@kernel.org \
    --cc=nphamcs@gmail.com \
    --cc=riel@surriel.com \
    --cc=roman.gushchin@linux.dev \
    --cc=shakeel.butt@linux.dev \
    --cc=youngjun.park@lge.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox