From: Baoquan He <baoquan.he@linux.dev>
To: Klara Modin <klarasmodin@gmail.com>
Cc: Baoquan He <hebaoquan@kylinos.cn>,
linux-mm@kvack.org, akpm@linux-foundation.org, chrisl@kernel.org,
kasong@tencent.com, hannes@cmpxchg.org, nphamcs@gmail.com,
baohua@kernel.org, youngjun.park@lge.com, david@kernel.org,
kunwu.chan@gmail.com, gourry@gourry.net, riel@surriel.com,
mhocko@kernel.org, roman.gushchin@linux.dev,
shakeel.butt@linux.dev
Subject: Re: [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend
Date: Mon, 28 Sep 2026 11:30:59 +0800 [thread overview]
Message-ID: <arnfc8oXEdCJrWEw@MiWiFi-R3L-srv> (raw)
In-Reply-To: <arUKSOVFzPF-iHxU@parmesan.int.kasm.eu>
On 09/24/26 at 02:17pm, Klara Modin wrote:
> On 2026-09-20 15:20:26 +0800, Baoquan He wrote:
> > This is the writeback layer for xswap, on top of the base series. Post it
> > as RFC for discussion.
> >
> > The base series keeps every swapped-out page in zswap. So the pool must
> > refuse new pages once it's full. This series gives an xswap slot a
> > backend to go, and charge it only when it really goes to real disk
> > space.
>
> There is some weird trailing whitespace in this email for some sections.
Thanks for careful checking, will clean this up before sending out.
...snip...
> > Testing
> > -------
> > qemu KVM guest, 8G RAM.
> >
> > A 4G swap disk /dev/vdb is added as the physical backend, and zswap is
> > turned off so that every page has to reach a backend slot. Note that it
> > need create xswap device firstly then disable zswap, so every page has
> > to reach backend slot.
> >
> > The workload is memhog, as in the base series. Every page is filled with
> > a fixed pattern, so the data can be checked after a readback.
> >
> > # echo 1 > /sys/module/zswap/parameters/enabled
> > # echo 100 > /sys/kernel/mm/xswap/create
> > # mkswap /dev/vdb && swapon -p 0 /dev/vdb
> > # echo 0 > /sys/module/zswap/parameters/enabled
> > # mkdir -p /sys/fs/cgroup/xswap_limit
> > # echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
> > # MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 ./memhog
>
> You did not include e.g. cgexec here, so is memhog really running in the
> cgroup you created?
You are right, I didn't write out the full command, just list the steps.
And here memhog is not a standard one from installed rpm package
numactl, it's a small local helper. It faults <total_GB> of anon, holds
it, and can drop the mapping or read it back. And the helper keep a known
byte pattern, so the data can be checked byte for byte after it is read
back.
# echo 1 > /sys/module/zswap/parameters/enabled
# echo 100 > /sys/kernel/mm/xswap/create
# mkswap /dev/vdb && swapon -p 0 /dev/vdb
# echo 0 > /sys/module/zswap/parameters/enabled
# mkdir -p /sys/fs/cgroup/xswap_limit
# echo 4G > /sys/fs/cgroup/xswap_limit/memory.max
# echo max > /sys/fs/cgroup/xswap_limit/memory.swap.max
# ( echo $BASHPID > /sys/fs/cgroup/xswap_limit/cgroup.procs
# exec env MEMHOG_FILL=pattern numactl --cpunodebind=0 --membind=0 \
# /mnt/kernel_src/testing_kernel_hebq/xswap/memhog 5 600 ) &
I will add the complete block in next version, sorry for the
inconvenience.
>
> >
> > 1. Swapout to the backend
> >
> > Run the workload with 5G of memory in a cgroup capped at 4G:
> >
> > # awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
> > Filename Type Size Used Priority
> > xswap0 xswap 8155132 2392176 100
> > /dev/vdb partition 4194300 2392176 0
> >
> > The two Used values are equal, so the pages reached the backend and
> > the xswap side accounted for them.
>
> This is a bit confusing to me. If I now run free, the written back data
> will be represented twice in the sum? Since the written back data is
> still presented as used in the xswap device, does it still consume space
> there, i.e. does it prevent that amount of new data being stored in the
> xswap device? If this is the case, it would make it even harder to
> dimension the xswap size since one would also have to take into account
> all the other swap devices which are present in the system.
I have to admit you found out an issue, thanks. I noticed it too while
I havne't thought of a satisfying way to fix it. What I can think of
is to mark the usage size of writeback on disk and list it separately.
# awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
Filename Type Size Used Priority
xswap0 xswap 8155132 2392176 100
/dev/vdb partition 4194300 2392176 0
2392176 (xswap)
If xswap is disabled and full, /dev/vdb got swapped out data by its own,
then the amount of used slot will be reflected as before.
# awk 'NR == 1 || $1 ~ /xswap|vdb/' /proc/swaps
Filename Type Size Used Priority
xswap0 xswap 8155132 2392176 100
/dev/vdb partition 4194300 3000000 0
2392176 (xswap)
I think this can be fixed later, I see it as a non-blocker for now.
>
> >
> > 2. Pages are read back from the device
> >
> > Lifting memory.max and touching the pages again reads them back. The
> > pages-swapped-in counter went up by 265642, so the read did reach the
> > device. 1010 of the folios read back were 1M, each read with one IO:
> >
> > # echo max > /sys/fs/cgroup/xswap_limit/memory.max
> > # cat /sys/kernel/mm/transparent_hugepage/hugepages-1024kB/stats/swpin
> > 1010
> >
> > hugepages-2048kB/stats/swpin stays at 0, because the swapin order is
> > capped below the PMD order.
> >
> > 3. verifies data correctness
> >
> > After the readback the data is compared byte for byte against the
> > pattern, one time with zswap on and one time with zswap off, so the
> > pages come from zswap in one time and from the backend in the other
> > time. All bytes matched.
> >
> > 4. Destroying a device returns its backend slots
> >
> > The device was filled, then destroyed with no readback first:
> >
> > # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
> > 3157220
> > # echo 0 > /sys/kernel/mm/xswap/destroy
> > # awk '$1 == "/dev/vdb" { print $4 }' /proc/swaps
> > 0
> >
> > 5. verifies the charge is correct
> >
> > An xswap entry is charged only when it holds a backend slot, so the
> > two counters have to show the same number of pages:
> >
> > # echo $(( $(cat /sys/fs/cgroup/xswap_limit/memory.swap.current) / 4096 ))
> > 526852
> > # awk '$1 == "/dev/vdb" { print $4 / 4 }' /proc/swaps
> > 526852
> >
> > Nothing is charged while the pages stay in zswap: with the pool holding
> > them and the backend untouched, memory.swap.current stays 0.
> >
> > 6. A cgroup with no swap room can still reclaim its anon memory
> >
> > With memory.swap.max at 0 and no xswap device, the cgroup ran out of
> > room and the workload was killed. With an xswap device and zswap on,
> > it swapped out through zswap, was charged nothing, and was not killed.
> >
> > E.g if memory.swap.max is set at 100M, the cgroup will be killed when
> > the cap was reached.
>
> Disregarding my comments and opinons on the interface, it seems to be
> working so far and I can see writeback happen during a GCC 17 build on
> my BPI-F3, but the entire thing takes roughly 18 hours and is not done
> yet.
Thanks again for your careful checking, testing.
Thanks
Baoquan
>
> >
> > Baoquan He (12):
> > mm, swap: tag a swap table entry with its owning xswap entry
> > mm, swap: prepare the folio-less allocation path for xswap
> > mm, swap: add a physical backend for xswap slots
> > mm, swap: use the xswap physical backend
> > mm, swap: fall back to disk when zswap refuses an xswap page
> > mm, swap: support swapoff of an xswap physical backend
> > mm, swap: reclaim physical slots backing cache-only xswap entries
> > mm, swap: back a large xswap folio with a contiguous physical run
> > mm, swap: enable THP swapin for xswap entries
> > mm, swap: drop swap_folio_sector()
> > mm, swap: do not retake the cluster lock when uncharging an xswap slot
> > mm, swap: drop a refused xswap backend run directly
> >
> > Nhat Pham (5):
> > mm, swap: prepare the swap IO path for xswap backends
> > mm, swap: split the swap memcg charge helpers
> > mm, swap: do not charge zswap-backed xswap entries
> > mm, swap: charge an xswap entry when it gets physical backing
> > mm, swap: don't gate xswap on the physical swap free count
> >
> > .../admin-guide/cgroup-v1/memcg_test.rst | 2 +-
> > include/linux/memcontrol.h | 6 +
> > include/linux/swap.h | 69 +-
> > include/linux/swap_ops.h | 9 +-
> > mm/memcontrol-v1.c | 10 +-
> > mm/memcontrol.c | 147 ++--
> > mm/memory.c | 7 +-
> > mm/page_io.c | 95 ++-
> > mm/swap.h | 35 +-
> > mm/swap_state.c | 12 +-
> > mm/swap_table.h | 53 ++
> > mm/swapfile.c | 742 ++++++++++++++++--
> > mm/zswap.c | 20 +-
> > 13 files changed, 1026 insertions(+), 181 deletions(-)
> >
> > --
> > 2.54.0
> >
> >
prev parent reply other threads:[~2026-09-28 3:31 UTC|newest]
Thread overview: 25+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-20 7:20 [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Baoquan He
2026-09-20 7:20 ` [RFC PATCH 01/17] mm, swap: prepare the swap IO path for xswap backends Baoquan He
2026-09-20 7:20 ` [RFC PATCH 02/17] mm, swap: tag a swap table entry with its owning xswap entry Baoquan He
2026-09-20 7:20 ` [RFC PATCH 03/17] mm, swap: prepare the folio-less allocation path for xswap Baoquan He
2026-09-20 7:20 ` [RFC PATCH 04/17] mm, swap: add a physical backend for xswap slots Baoquan He
2026-09-20 7:20 ` [RFC PATCH 05/17] mm, swap: use the xswap physical backend Baoquan He
2026-09-20 7:20 ` [RFC PATCH 06/17] mm, swap: fall back to disk when zswap refuses an xswap page Baoquan He
2026-09-20 7:20 ` [RFC PATCH 07/17] mm, swap: support swapoff of an xswap physical backend Baoquan He
2026-09-20 7:20 ` [RFC PATCH 08/17] mm, swap: reclaim physical slots backing cache-only xswap entries Baoquan He
2026-09-20 7:20 ` [RFC PATCH 09/17] mm, swap: back a large xswap folio with a contiguous physical run Baoquan He
2026-09-20 7:20 ` [RFC PATCH 10/17] mm, swap: enable THP swapin for xswap entries Baoquan He
2026-09-20 7:20 ` [RFC PATCH 11/17] mm, swap: drop swap_folio_sector() Baoquan He
2026-09-20 7:20 ` [RFC PATCH 12/17] mm, swap: split the swap memcg charge helpers Baoquan He
2026-09-20 7:20 ` [RFC PATCH 13/17] mm, swap: do not charge zswap-backed xswap entries Baoquan He
2026-09-21 11:54 ` Chris Li
2026-09-20 7:20 ` [RFC PATCH 14/17] mm, swap: charge an xswap entry when it gets physical backing Baoquan He
2026-09-20 7:20 ` [RFC PATCH 15/17] mm, swap: don't gate xswap on the physical swap free count Baoquan He
2026-09-20 7:20 ` [RFC PATCH 16/17] mm, swap: do not retake the cluster lock when uncharging an xswap slot Baoquan He
2026-09-20 7:20 ` [RFC PATCH 17/17] mm, swap: drop a refused xswap backend run directly Baoquan He
2026-09-22 14:22 ` [RFC PATCH 00/17] mm, swap: xswap writeback to a physical backend Chris Li
2026-09-22 14:30 ` Johannes Weiner
2026-09-24 1:56 ` KunWu Chan
2026-09-24 11:36 ` Chris Li
2026-09-24 12:17 ` Klara Modin
2026-09-28 3:30 ` Baoquan He [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arnfc8oXEdCJrWEw@MiWiFi-R3L-srv \
--to=baoquan.he@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=chrisl@kernel.org \
--cc=david@kernel.org \
--cc=gourry@gourry.net \
--cc=hannes@cmpxchg.org \
--cc=hebaoquan@kylinos.cn \
--cc=kasong@tencent.com \
--cc=klarasmodin@gmail.com \
--cc=kunwu.chan@gmail.com \
--cc=linux-mm@kvack.org \
--cc=mhocko@kernel.org \
--cc=nphamcs@gmail.com \
--cc=riel@surriel.com \
--cc=roman.gushchin@linux.dev \
--cc=shakeel.butt@linux.dev \
--cc=youngjun.park@lge.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox