The Linux Kernel Mailing List
 help / color / mirror / Atom feed
From: Baoquan He <baoquan.he@linux.dev>
To: Youngjun Park <youngjun.park@lge.com>
Cc: Johannes Weiner <hannes@cmpxchg.org>,
	linux-mm@kvack.org, chrisl@kernel.org, nphamcs@gmail.com,
	kasong@tencent.com, baohua@kernel.org, yosry@kernel.org,
	david@kernel.org, shikemeng@huaweicloud.com,
	chengming.zhou@linux.dev, linux-kernel@vger.kernel.org
Subject: Re: [RFC PATCH v2 01/10] mm: xswap support for zswap
Date: Mon, 10 Aug 2026 15:25:07 +0800	[thread overview]
Message-ID: <anl80zDN2t4Yex8K@MiWiFi-R3L-srv> (raw)
In-Reply-To: <ankuMRz8BjHomhY9@yjaykim-PowerEdge-T330>

On 08/10/26 at 10:49am, Youngjun Park wrote:
> On Fri, Aug 07, 2026 at 05:11:05PM +0800, Baoquan He wrote:
> > Hi Johannes,
> > 
> > On 08/05/26 at 10:17am, Johannes Weiner wrote:
> > > On Wed, Aug 05, 2026 at 03:53:24PM +0800, Baoquan He wrote:
> > > > From: Chris Li <chrisl@kernel.org>
> > > > 
> > > > Introduce extendable (virtual) swap device support ??? xswap.
> > > > 
> > > > The current zswap requires a backing swapfile. The swap slot used
> > > > by zswap is not able to be used by the swapfile, wasting swapfile
> > > > space.
> > > > 
> > > > An xswap device is a swapfile that only contains the swap header,
> > > > with the header indicating the size of the virtual swap space. There
> > > > is no swap data section, therefore no waste of swapfile space. Any
> > > > write to an xswap device will fail. To prevent accidental read or
> > > > write, bdev of swap_info_struct is set to NULL. Xswap devices set
> > > > the SSD flag because there is no rotational disk access when using
> > > > zswap.
> > > > 
> > > > Zswap writeback is disabled if all swapfiles in the system are
> > > > xswap devices (tracked via nr_real_swapfiles).
> > > > 
> > > > How to create an xswap device:
> > > >   touch swap.1G
> > > >   truncate -s 1G swap.1G
> > > >   mkswap swap.1G
> > > >   dd if=swap.1G of=xswap.1G bs=4096 count=1
> > > >   # xswap.1G is 4K on disk but reports 1G capacity
> > > >   swapon xswap.1G
> > > 
> > > Sigh.
> > > 
> > > Why does the user have to go through this dance?
> > > 
> > > Why does the user have to decide in advance what size the space needs
> > > to be?
> > > 
> > > You point out no inherent limit to how much can be compressed, so
> > > there is no reason to make userspace decide on an arbitrary one.
> > > 
> > > There is no reason to tie an address space that can be managed
> > > transparently inside the kernel to TWO named files on disk.
> > 
> > Thanks for looking into this.
> > 
> > The file-based creation dance is there only because this is RFC —
> > I wanted to reuse the existing swapon path so the core grow/shrink
> > machinery could be measured and tested without also designing a new
> > userspace interface. I agree it's not the right final interface.
> > 
> > The direction I'm thinking for the next revision:
> > 
> > - Drop the file requirement entirely.  An xswap device has no backing
> >   store, so there is no reason it needs a file.
> > 
> > - Use totalram_pages as the initial per-device size.  Chris suggested
> >   this, and it's a natural bound: if all anonymous memory is swapped
> >   out, that is the maximum number of swap entries zswap will ever need,
> >   assuming a reasonable compression ratio. The hard upper limit could
> >   be 2 times of system RAM, or the max system RAM memory hotplug can
> >   add to.
> > 
> >   Doing this because we need consider swap.tier support. A single global
> >   xswap device in swap.tier would mean all memcgs compress into the
> >   same device — there is only one swap entry namespace. With per-device
> >   xswap instances, swap.tier can bind different memcgs to different xswap
> >   devices, giving each its own swap slot namespace. Total isolation on slot
> >   usage, no cross-memcg interference.
> 
> Hello Baoquan :)

Thanks for the thoughtful input, Youngjun.

> 
> Is there concrete user scenario isolation is needed?

Hmm, not yet, it just comes from my personal understanding and thinking.
I will exlain later.

> 
> >    -----
> >    Hi Chris, Joungjun,
> >    Please correct me if I misunderstood the swap.tier concept and xswap
> >    use case in there.)
> 
> >    -----
> 
> Anysway, if we want to use xswap isolation like below,
> 
> 	xswap t1  xswap t2     tier1     tier2
>        | x1 |     | x2 |     | dev1 |  | dev2 |
> 
> then each memcg may have its own xswap front-end and backing tie
> 
>         memcg1: xswap t1  + tier1
>         memcg2: xswap t2 + tier2
> 
> However, with the current tier design, the root cgroup needs to see the
> whole tier layout. In that case, I think the root view may become unclear
> if there are multiple xswap instances. From the root cgroup point of
> view, it may be better to see xswap as one logical tier, not as two
> separate tiers.
> 
> For example, the layout could be like this:
> 
>               xswap tier        tier1     tier2
>         | xswap1 xswap2 |     | dev1 |  | dev2 |

Exactly, this is what I am thinking about. Unless we assign each xswap
different memory tier or memory type in the future. Otherwise, all xswap
instances should share the same tier.

> 
> Then each memcg can have its own mapping or policy:
> 
>         memcg1: xswap tier + tier1  (xswap1 + dev1)
>         memcg2: xswap tier + tier2  (xswap2 + dev2)
> 
> With this model, the root cgroup can keep one simple global view of the
> xswap tier. At the same time, each memcg can still use a specific xswap
> area and a specific backing swap tier.

Exactly. This is what per xswap device is trying to do.

> 
> P.s I am thinking about multiple xswap usecase on tier. 
> (this is just mind map. I don't know whether it is right or not)
> 
> Another possible layout may be to use xswap as a RAM buffer for each
> tier
> 
>               tier1                 tier2
>         | xswap + fast dev |   | xswap + slow dev |
> 
> We can use xswap as a simple buffering layer?
> In that case, we would need a clear policy
> for how xswap is assigned to each tier, and how the backing swap device is
> selected for each tier. 

I may not get your point. Do you mean kernel will decide how to assign
each xswap to tier and how to select back device via policy, but not
user? 

> 
> So why I am saying this is that, 
> if xswap is managed as part of swap tiers, it would
> be helpful to define a more concrete layout and policy for xswap
> assignment.
> 
> This would make the isolation use case much clearer

Here let me explain why I think per-device xswap is the right starting
point.

The primary scenario I have in mind is cloud environments where different
containers need different swap tier chains:

memcg A (latency-sensitive service):
    xswap-A → NVMe tier (fast writeback, small capacity)

memcg B (batch processing): 
    xswap-B → HDD tier (slow writeback, large capacity) 

memcg C (best-effort):
    xswap-C → no backend (compression only, never write to disk)

With a single global xswap, all three memcgs share the same backend. If
memcg B triggers heavy writeback to HDD, memcg A's swap-in path is unaffected
in the IO sense, but the shared slot namespace means A and B compete for
swap entries. Per-device xswap gives each its own slot pool and its own
backing device — full isolation.

Beyond IO and slot isolation, per-device xswap also makes per-memcg swap
statistics straightforward: slot usage, compressed page count, and writeback
throughput are naturally scoped to each device.

Futhermore, the existing swap subsystem is already per-device — each swapon
allocates an independent swap_info_struct with its own cluster table, slot
namespace, statistics, and teardown path. Per-device xswap inherits all of
this for free:

- Slot allocation via the existing per-CPU cluster allocator (no new locking)
- Per-device statistics already exposed in /proc/swaps and sysfs
- Device lifecycle via swapon/swapoff with no new management interface
- swap.tier naturally binds tiers to per-device instances

Thanks
Baoquan

  reply	other threads:[~2026-08-10  7:25 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-05  7:53 [RFC PATCH v2 00/10] mm, swap: dynamic cluster management for xswap devices Baoquan He
2026-08-05  7:53 ` [RFC PATCH v2 01/10] mm: xswap support for zswap Baoquan He
2026-08-05 14:17   ` Johannes Weiner
2026-08-07  9:11     ` Baoquan He
2026-08-10  1:49       ` Youngjun Park
2026-08-10  7:25         ` Baoquan He [this message]
2026-08-05  7:53 ` [RFC PATCH v2 02/10] mm, swap: add CONFIG_XSWAP and xswap fields to swap_info_struct Baoquan He
2026-08-05  7:53 ` [RFC PATCH v2 03/10] mm, swap: add xswap cluster grow via VM_SPARSE vmalloc Baoquan He
2026-08-05  7:53 ` [RFC PATCH v2 04/10] mm, swap: add xswap grow trigger on cluster allocation Baoquan He
2026-08-05  7:53 ` [RFC PATCH v2 05/10] mm, swap: add xswap_try_shrink and shrink trigger on cluster free Baoquan He
2026-08-05  7:53 ` [RFC PATCH v2 06/10] mm, swap: free backing pages in xswap_unmap_clusters Baoquan He
2026-08-05  7:53 ` [RFC PATCH v2 07/10] mm, swap: add nr_free_tail for O(1) xswap shrink detection Baoquan He
2026-08-05  7:53 ` [RFC PATCH v2 08/10] mm, swap: add adjustable runtime ceiling (nr_clusters) for xswap Baoquan He
2026-08-05  7:53 ` [RFC PATCH v2 09/10] mm, swap: add debugfs knob for xswap per-device cluster limit Baoquan He
2026-08-05  7:53 ` [RFC PATCH v2 10/10] mm, swap: defer xswap shrink to workqueue to avoid lock recursion Baoquan He

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anl80zDN2t4Yex8K@MiWiFi-R3L-srv \
    --to=baoquan.he@linux.dev \
    --cc=baohua@kernel.org \
    --cc=chengming.zhou@linux.dev \
    --cc=chrisl@kernel.org \
    --cc=david@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=kasong@tencent.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=nphamcs@gmail.com \
    --cc=shikemeng@huaweicloud.com \
    --cc=yosry@kernel.org \
    --cc=youngjun.park@lge.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox