From: Baoquan He <baoquan.he@linux.dev>
To: Nhat Pham <nphamcs@gmail.com>
Cc: Kairui Song <ryncsn@gmail.com>, Baoquan He <hebaoquan@kylinos.cn>,
linux-mm@kvack.org, akpm@linux-foundation.org, chrisl@kernel.org,
kasong@tencent.com, baohua@kernel.org, youngjun.park@lge.com,
hannes@cmpxchg.org, yosry@kernel.org, shikemeng@huaweicloud.com,
chengming.zhou@linux.dev, david@kernel.org,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH 00/16] xswap: extendable swap device backed by zswap
Date: Fri, 4 Sep 2026 17:42:20 +0800 [thread overview]
Message-ID: <apqSfFGyJpUMWCY9@fedora> (raw)
In-Reply-To: <CAKEwX=MU9uXVenEu7he+h-Pq7K3BmCGvvKS-KiPK61Xd0ox_mQ@mail.gmail.com>
On 09/02/26 at 10:10am, Nhat Pham wrote:
> On Tue, Sep 1, 2026 at 7:03 AM Baoquan He <baoquan.he@linux.dev> wrote:
> >
> > On 09/01/26 at 01:54am, Kairui Song wrote:
> > > On Thu, Aug 27, 2026 at 05:44:50PM +0800, Baoquan He wrote:
> > > > xswap is an extendable swap device with no backing storage. Swapped-out
> > > > pages live only in zswap, so the device wastes no disk space and its
> > > > size is independent of any physical device.
> > > >
> > > > xswap decouples PTE swap entries from physical backing storage. The
> > > > cluster_info array is backed by a sparse vmalloc (VM_SPARSE) area that is
> > > > grown and shrunk on demand:
> > > >
> > > > - Grow: when cluster allocation runs out of free clusters and the device
> > > > is below its ceiling, more physical pages are mapped into the VM_SPARSE
> > > > area and their clusters are added to the free list.
> > > >
> > > > - Shrink: when contiguous free clusters accumulate at the tail of the
> > > > mapped range (tracked in O(1) via nr_free_tail), they are unmapped and
> > > > the backing pages freed. Shrink is deferred to a workqueue to avoid
> > > > lock recursion.
> > > >
> > >
> > > Hi Baoquan,
> > >
> > > I didn't check too many details on how the implementation in previous
> > > RFC until now, After looking at it, using VM_SPARSE to setup the cluster
> > > info area is a really smart idea, really good job!
> > >
> > > I think many info are missing in the cover letter though so I wasn't
> > > sure how this grow and shrink works from the description, after
> > > checking the code, it looks much cleaner to me now, correct me
> > > if I'm wrong:
> > >
> > > Every xswap device will have a huge and fixed "hard limit"
> > > (si->max and si->nr_clusters_max), and practically can be considered
> > > large enough to hold any workload, and won't change once swapon
> > > is done.
> > >
> > > The actually data (si->cluster_info) of xswap device is completely
> > > sparse and dynamic using VM_SPARSE, and so we don't need to change
> > > any existing routine. It grow/alloc and shrink/free automatically by
> > > the kernel, limited or driven by a "soft limit" (si->nr_clusters
> > > and si->pages) which you can modify using the interface below.
> >
> > Thanks a lot for careful checking, and you are quite right about the
> > mechanism and details.
> >
> > >
> > > Once concern is that the "hard limit" is now the total RAM size. Isn't
> > > that actually a bit small? Will be better if that one is tunable too?
> > > With a parameter, and before swap on, as the hard limit is hard to
> > > adjust once swapon is done. Any thing limiting this?
> >
> > Chris and I talked about this, we both think the total RAM size is a
> > good hard limit. Because xswap is similar with zswap/zram in essence by
> > compressing memory content to save memory. So the real limit is the
> > zswap pool, not the slot count. In fact it's never able to utilize the
> > total system RAM, right? Making it larger than system RAM is
> > meaningless.
>
> No. This is not quite right. The size of this device is the size of
> the "swapped out" data, which is multiple times the post-compression
> size (i.e zswap pool size). The multiple here depends on how well the
> data is compressed.
>
> This is not to consider the other swap backends:
>
> 1. zero-filled swap pages have effectively 0 memory footprint.
>
> 2. disk swap pages (I know this is not currently supported yet, but
> it's a consideration for the overall design).
You are right, system RAM may not be a good hard limit. Then 2x system
RAM?
>
> >
> > Memory hotplug is a case in which system RAM can be enlarged during
> > system running, while that can be taken into account later as a enhanced
> > feature if it's really wanted.
>
> A lot of these problems are self-inflicted. If we design a fully
> dynamic swap device, then it's not in consideration.
As initial version, I'd like to make it not fully dynamic. The current
mechanism can make the change of xswap being full dynamic very easy.
I am glad to see people can post patch to change it later with
justification. I don't think we need to make everything perfect and anyone
satisfied at the beginning.
>
> >
> > >
> > > And I think these details better be mentioned bit more too.
> >
> > Sure, I can put these thoughts into cover letter or patch log for
> > reference.
> >
> > >
> > > > A per-device ceiling (nr_clusters) bounds growth and is adjustable at
> > > > runtime via debugfs.
> > > >
> > > > Interface:
> > > >
> > > > /sys/kernel/mm/xswap/create write "<percent> [<prio>]" to
> > > > create a device; percent is a
> > > > percent of RAM (0 for the default),
> > > > prio is an optional swap priority
> > > > (default DEF_SWAP_PRIO)
> > >
> > > With what I have read so far, the mandatory percent limit here is kind of
> > > strange, even with 0 as default. Why not make both args optional and just
> > > let it grow without any limit by default? It looks more "fully dynamic"
> > > that way.
> >
> > I'd like to clarify why we default to a soft limit rather than "no limit".
> >
> > The soft limit is the administrator's deliberate size choice, similar
> > to how zram requires an explicit size. On a multi-TB system the
> > cluster_info array is not free, so planning how much of it to allow is a
> > real decision. The current behavior is: grow up to the soft limit as usage
> > demands, then stay there. We do not shrink on idle, and shrink only happens
> > when the admin lowers the limit. So there is no grow/shrink oscillation in
> > normal operation.
> >
> > A default of "no limit / fully dynamic" will instead let the device grow
> > without restriction under memory pressure. While allocating cluster_info
> > pages exactly when memory is scarce, relying on shrink to reclaim afterwards,
> > which is the oscillation we want to avoid. So we'll make both create arguments
> > optional, but the default will be a sensible ceiling rather than unbounded.
> >
> > >
> > > > /sys/kernel/mm/xswap/destroy write a swap type to tear down
> > > > a device
> > > > /sys/kernel/debug/xswap/type<N>_cluster_limit
> > > > read/write the per-device
> > > > cluster ceiling
> > >
>
> What's the point of having multiple xswap devices if it's going to be
> dynamic, cluster-based anyway?
next prev parent reply other threads:[~2026-09-04 9:42 UTC|newest]
Thread overview: 36+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-27 9:44 [PATCH 00/16] xswap: extendable swap device backed by zswap Baoquan He
2026-08-27 9:44 ` [PATCH 01/16] mm: zswap: return -ENOENT when the swap device is gone Baoquan He
2026-09-02 14:53 ` Nhat Pham
2026-09-03 7:54 ` Baoquan He
2026-08-27 9:44 ` [PATCH 02/16] mm: xswap support for zswap Baoquan He
2026-08-27 9:44 ` [PATCH 03/16] mm, swap: add CONFIG_XSWAP and xswap fields to swap_info_struct Baoquan He
2026-08-27 9:44 ` [PATCH 04/16] mm, swap: refactor free_swap_cluster_info to take swap_info_struct Baoquan He
2026-08-27 9:44 ` [PATCH 05/16] mm, swap: add xswap cluster grow via VM_SPARSE vmalloc Baoquan He
2026-08-27 9:44 ` [PATCH 06/16] mm, swap: add sysfs create interface for xswap Baoquan He
2026-08-27 9:44 ` [PATCH 07/16] mm, swap: add xswap grow trigger on cluster allocation Baoquan He
2026-09-02 14:15 ` Nhat Pham
2026-09-03 8:24 ` Baoquan He
2026-08-27 9:44 ` [PATCH 08/16] mm, swap: add xswap_try_shrink and shrink trigger on cluster free Baoquan He
2026-08-27 9:44 ` [PATCH 09/16] mm, swap: free backing pages in xswap_unmap_clusters Baoquan He
2026-08-27 9:45 ` [PATCH 10/16] mm, swap: add nr_free_tail for O(1) xswap shrink detection Baoquan He
2026-08-27 9:45 ` [PATCH 11/16] mm, swap: add adjustable runtime ceiling (nr_clusters) for xswap Baoquan He
2026-08-27 9:45 ` [PATCH 12/16] mm, swap: add debugfs knob for xswap per-device cluster limit Baoquan He
2026-08-27 9:45 ` [PATCH 13/16] mm, swap: defer xswap shrink to workqueue to avoid lock recursion Baoquan He
2026-09-02 14:50 ` Nhat Pham
2026-09-03 9:17 ` Baoquan He
2026-08-27 9:45 ` [PATCH 14/16] mm, swap: refactor swapoff + add xswap_destroy Baoquan He
2026-09-03 6:59 ` Youngjun Park
2026-09-04 5:33 ` Baoquan He
2026-08-27 9:45 ` [PATCH 15/16] mm, swap: require zswap for xswap devices Baoquan He
2026-09-03 6:52 ` Youngjun Park
2026-09-04 7:57 ` Baoquan He
2026-08-27 9:45 ` [PATCH 16/16] mm, swap: allow setting xswap device priority at creation Baoquan He
2026-08-27 13:59 ` [syzbot ci] Re: xswap: extendable swap device backed by zswap syzbot ci
2026-08-31 8:35 ` Baoquan He
2026-08-31 17:54 ` [PATCH 00/16] " Kairui Song
2026-09-01 11:03 ` Baoquan He
2026-09-02 14:10 ` Nhat Pham
2026-09-04 9:42 ` Baoquan He [this message]
2026-09-02 14:33 ` Nhat Pham
2026-09-03 7:35 ` Youngjun Park
2026-09-04 3:33 ` Baoquan He
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=apqSfFGyJpUMWCY9@fedora \
--to=baoquan.he@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=chengming.zhou@linux.dev \
--cc=chrisl@kernel.org \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=hebaoquan@kylinos.cn \
--cc=kasong@tencent.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=nphamcs@gmail.com \
--cc=ryncsn@gmail.com \
--cc=shikemeng@huaweicloud.com \
--cc=yosry@kernel.org \
--cc=youngjun.park@lge.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox