From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from lgeamrelo13.lge.com (lgeamrelo13.lge.com [156.147.23.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0D3F621770B for ; Mon, 10 Aug 2026 01:49:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=156.147.23.53 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786326590; cv=none; b=i6wJTWqv4AIsvwte7eMc0GfyhteoNnV+k2gSNA+SEb4LhLpPdVa11pbnUit85PeKqBOg4PJBYD8CE7b5kbfe05bSNlAhRpxZ0TpRG7JJupWncoLlFa3M4QNrIJvDfBMMJWSeBQRHmicnss59BMRpIgl0gM0oo8EIuVVNh5UE2qQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786326590; c=relaxed/simple; bh=2zkH3SPoZeWlnoU18mPCYzeB5XEBEjKST2yQDYgQoWg=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=FskZwBDjfkZPoDx24+S9+DCSv+20tfFESUnoO2J4SFa97KMZgxeke4ZF/ULzOF7WDI3OWyN2XhtStCw4tBEufIRjNsabkYU4g3Pj69geHJcKF9cqclFmbGiVe+u5qSM1MptCzHZEamBioopGFijlM1/BNF6soCfdfGXndtWuDQc= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=lge.com; spf=pass smtp.mailfrom=lge.com; arc=none smtp.client-ip=156.147.23.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=lge.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=lge.com Received: from unknown (HELO lgeamrelo01.lge.com) (156.147.1.125) by 156.147.23.53 with ESMTP; 10 Aug 2026 10:49:37 +0900 X-Original-SENDERIP: 156.147.1.125 X-Original-MAILFROM: youngjun.park@lge.com Received: from unknown (HELO yjaykim-PowerEdge-T330) (10.177.112.156) by 156.147.1.125 with ESMTP; 10 Aug 2026 10:49:37 +0900 X-Original-SENDERIP: 10.177.112.156 X-Original-MAILFROM: youngjun.park@lge.com Date: Mon, 10 Aug 2026 10:49:37 +0900 From: Youngjun Park To: Baoquan He Cc: Johannes Weiner , linux-mm@kvack.org, chrisl@kernel.org, nphamcs@gmail.com, kasong@tencent.com, baohua@kernel.org, yosry@kernel.org, david@kernel.org, shikemeng@huaweicloud.com, chengming.zhou@linux.dev, linux-kernel@vger.kernel.org Subject: Re: [RFC PATCH v2 01/10] mm: xswap support for zswap Message-ID: References: <20260805075336.3579395-1-baoquan.he@linux.dev> <20260805075336.3579395-2-baoquan.he@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: On Fri, Aug 07, 2026 at 05:11:05PM +0800, Baoquan He wrote: > Hi Johannes, > > On 08/05/26 at 10:17am, Johannes Weiner wrote: > > On Wed, Aug 05, 2026 at 03:53:24PM +0800, Baoquan He wrote: > > > From: Chris Li > > > > > > Introduce extendable (virtual) swap device support ??? xswap. > > > > > > The current zswap requires a backing swapfile. The swap slot used > > > by zswap is not able to be used by the swapfile, wasting swapfile > > > space. > > > > > > An xswap device is a swapfile that only contains the swap header, > > > with the header indicating the size of the virtual swap space. There > > > is no swap data section, therefore no waste of swapfile space. Any > > > write to an xswap device will fail. To prevent accidental read or > > > write, bdev of swap_info_struct is set to NULL. Xswap devices set > > > the SSD flag because there is no rotational disk access when using > > > zswap. > > > > > > Zswap writeback is disabled if all swapfiles in the system are > > > xswap devices (tracked via nr_real_swapfiles). > > > > > > How to create an xswap device: > > > touch swap.1G > > > truncate -s 1G swap.1G > > > mkswap swap.1G > > > dd if=swap.1G of=xswap.1G bs=4096 count=1 > > > # xswap.1G is 4K on disk but reports 1G capacity > > > swapon xswap.1G > > > > Sigh. > > > > Why does the user have to go through this dance? > > > > Why does the user have to decide in advance what size the space needs > > to be? > > > > You point out no inherent limit to how much can be compressed, so > > there is no reason to make userspace decide on an arbitrary one. > > > > There is no reason to tie an address space that can be managed > > transparently inside the kernel to TWO named files on disk. > > Thanks for looking into this. > > The file-based creation dance is there only because this is RFC — > I wanted to reuse the existing swapon path so the core grow/shrink > machinery could be measured and tested without also designing a new > userspace interface. I agree it's not the right final interface. > > The direction I'm thinking for the next revision: > > - Drop the file requirement entirely. An xswap device has no backing > store, so there is no reason it needs a file. > > - Use totalram_pages as the initial per-device size. Chris suggested > this, and it's a natural bound: if all anonymous memory is swapped > out, that is the maximum number of swap entries zswap will ever need, > assuming a reasonable compression ratio. The hard upper limit could > be 2 times of system RAM, or the max system RAM memory hotplug can > add to. > > Doing this because we need consider swap.tier support. A single global > xswap device in swap.tier would mean all memcgs compress into the > same device — there is only one swap entry namespace. With per-device > xswap instances, swap.tier can bind different memcgs to different xswap > devices, giving each its own swap slot namespace. Total isolation on slot > usage, no cross-memcg interference. Hello Baoquan :) Is there concrete user scenario isolation is needed? > ----- > Hi Chris, Joungjun, > Please correct me if I misunderstood the swap.tier concept and xswap > use case in there.) > ----- Anysway, if we want to use xswap isolation like below, xswap t1 xswap t2 tier1 tier2 | x1 | | x2 | | dev1 | | dev2 | then each memcg may have its own xswap front-end and backing tie memcg1: xswap t1 + tier1 memcg2: xswap t2 + tier2 However, with the current tier design, the root cgroup needs to see the whole tier layout. In that case, I think the root view may become unclear if there are multiple xswap instances. From the root cgroup point of view, it may be better to see xswap as one logical tier, not as two separate tiers. For example, the layout could be like this: xswap tier tier1 tier2 | xswap1 xswap2 | | dev1 | | dev2 | Then each memcg can have its own mapping or policy: memcg1: xswap tier + tier1 (xswap1 + dev1) memcg2: xswap tier + tier2 (xswap2 + dev2) With this model, the root cgroup can keep one simple global view of the xswap tier. At the same time, each memcg can still use a specific xswap area and a specific backing swap tier. P.s I am thinking about multiple xswap usecase on tier. (this is just mind map. I don't know whether it is right or not) Another possible layout may be to use xswap as a RAM buffer for each tier tier1 tier2 | xswap + fast dev | | xswap + slow dev | We can use xswap as a simple buffering layer? In that case, we would need a clear policy for how xswap is assigned to each tier, and how the backing swap device is selected for each tier. So why I am saying this is that, if xswap is managed as part of swap tiers, it would be helpful to define a more concrete layout and policy for xswap assignment. This would make the isolation use case much clearer Thanks! Youngjun