From: Usama Arif <usama.arif@linux.dev>
To: Yosry Ahmed <yosry@kernel.org>
Cc: Alexandre Ghiti <alex@ghiti.fr>,
Andrew Morton <akpm@linux-foundation.org>,
david@kernel.org, chrisl@kernel.org, kasong@tencent.com,
ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org,
ying.huang@linux.alibaba.com, Baoquan He <baoquan.he@linux.dev>,
willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org,
riel@surriel.com, shakeel.butt@linux.dev, kas@kernel.org,
baohua@kernel.org, dev.jain@arm.com,
baolin.wang@linux.alibaba.com, Nico Pache <nico.pache@linux.dev>,
"Liam R. Howlett" <liam@infradead.org>,
ryan.roberts@arm.com, Vlastimil Babka <vbabka@kernel.org>,
lance.yang@linux.dev, linux-kernel@vger.kernel.org,
nphamcs@gmail.com, shikemeng@huaweicloud.com,
kernel-team@meta.com, Alexandre Ghiti <alexghiti@fb.com>
Subject: Re: [PATCH v5 04/11] mm: zswap: add range lookup for large-folio swapin
Date: Fri, 24 Jul 2026 10:59:40 +0100 [thread overview]
Message-ID: <c40f4a8b-1f60-45ea-962d-0597c413ec3d@linux.dev> (raw)
In-Reply-To: <CAO9r8zMzjCf620EHwzfosngc94fu=vAb=ukrtOi9Qf3rA3vv7w@mail.gmail.com>
On 23/07/2026 17:42, Yosry Ahmed wrote:
>>>> @@ -1595,13 +1611,19 @@ int zswap_load(struct folio *folio)
>>>> return -ENOENT;
>>>>
>>>> /*
>>>> - * Large folios should not be swapped in while zswap is being used, as
>>>> - * they are not properly handled. Zswap does not properly load large
>>>> - * folios, and a large folio may only be partially in zswap.
>>>> + * A large folio reaches zswap_load() only when its whole range is
>>>> + * expected to be on disk: PMD swap-entry consumers split before
>>>> + * calling into PMD-order swapin whenever any slot is still in zswap.
>>>> + * Confirm the range is entirely absent from zswap and return -ENOENT
>>>> + * so the caller reads it from disk; if a slot is unexpectedly still in
>>>> + * zswap, fail the read rather than return partially-initialized data.
>>>> */
>>>> - if (WARN_ON_ONCE(folio_test_large(folio))) {
>>>> - folio_unlock(folio);
>>>> - return -EINVAL;
>>>> + if (folio_test_large(folio)) {
>>>> + if (zswap_is_present(swp, folio_nr_pages(folio))) {
>>>
>>> Is dropping the warning here intentional (for the folio_test_large() &&
>>> zswap_is_present() case)?
>>>
>>
>>
>> Yes, so we can end up in a race, which should be handled gracefully.
>>
>> For example, lets say we have zswap writeback enabled, which means we
>> can end up in a state where we have a PMD swap entry and 511 of the 512
>> slots have been written to disk, but 1 slot (slot X) is still in zswap.
>>
>> We can then have the following race:
>>
>>
>> CPU A: PMD swap-in CPU B: split-PTE swap-in
>> ------------------ ------------------------
>>
>> Checks swap cache: empty
>>
>> Faults slot X
>> Adds order-0 folio F
>> zswap_load(F):
>> removes X from zswap
>> marks F dirty
>>
>> Checks zswap range: empty
>> A is preempted
>>
>> Unmaps/reclaims F
>> zswap_store(F):
>> puts X back in zswap
>> Removes F from swap cache
>>
>> A resumes
>> Allocates large swap-cache folio G
>> zswap_load(G) finds X in zswap
>
> Why don't we check the zswap range after allocating a folio in the
> swap cache? I am assuming at this point we have the folio locked and
> the result should be stable?
>
Yes that makes sense. The folio is locked and the result will be stable.
How about something like below?
diff --git a/mm/swap_state.c b/mm/swap_state.c
index 1f08fb522036..41f168377cb1 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -12,6 +12,7 @@
#include <linux/kernel_stat.h>
#include <linux/mempolicy.h>
#include <linux/swap.h>
+#include <linux/zswap.h>
#include <linux/leafops.h>
#include <linux/init.h>
#include <linux/pagemap.h>
@@ -500,6 +501,17 @@ static struct folio *__swap_cache_alloc(struct swap_cluster_info *ci,
__folio_set_locked(folio);
__folio_set_swapbacked(folio);
__swap_cache_do_add_folio(ci, folio, entry);
+ /*
+ * Reject mixed zswap/disk backing before starting the read.
+ */
+ if (order && zswap_is_present(entry, nr_pages)) {
+ __swap_cache_do_del_folio(ci, folio, entry, shadow);
+ spin_unlock(&ci->lock);
+ folio_unlock(folio);
+ /* nr_pages refs from swap cache, 1 from allocation */
+ folio_put_refs(folio, nr_pages + 1);
+ return ERR_PTR(-EBUSY);
+ }
spin_unlock(&ci->lock);
if (mem_cgroup_swapin_charge_folio(folio, memcg_id,
>>
>> The correct action would be to reject the PMD order read and fall back
>> to per-page loading.
>>
>> I am bit torn about what to do for zswap here. The series is quite big
>> already. Alexandre is looking at adding support for PMD swap, but will
>> send patches once this series gets merged. My initial versions 1
>> and 2, basically stopped installing PMD swap entry if zswap was ever enabled.
>> From v3, I used Alexandre's suggestion to do zswap_is_present() test
>> so that PMD swap entry can keep on working.
>>
>> Both of these paths are temporary till Alex sends his series. I
>> do feel my initial version was simpler but will basically stop
>> working if zswap is ever enabled. Do you have any suggestions on
>> what your preference is for zswap?
>
> Even with PMD zswap load support, we still have to deal with the case
> where some order-0 pages are in zswap and some are on disk, right? I
> don't see how that would eliminate the case you mentioned above.
Ack, yeah writeback will still mess with things. I will keep the current
implementation.
next prev parent reply other threads:[~2026-07-24 9:59 UTC|newest]
Thread overview: 23+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-22 15:19 [PATCH v5 00/11] mm: PMD-level swap entries for anonymous THPs Usama Arif
2026-07-22 15:19 ` [PATCH v5 01/11] mm: add PMD swap entry detection support Usama Arif
2026-07-24 6:10 ` Dev Jain
2026-07-24 10:00 ` Usama Arif
2026-07-22 15:19 ` [PATCH v5 02/11] mm: add PMD swap entry splitting support Usama Arif
2026-07-24 6:52 ` Dev Jain
2026-07-24 10:05 ` Usama Arif
2026-07-22 15:19 ` [PATCH v5 03/11] mm: handle PMD swap entries in fork path Usama Arif
2026-07-22 15:19 ` [PATCH v5 04/11] mm: zswap: add range lookup for large-folio swapin Usama Arif
2026-07-23 0:01 ` Yosry Ahmed
2026-07-23 12:45 ` Usama Arif
2026-07-23 16:42 ` Yosry Ahmed
2026-07-23 17:15 ` Nhat Pham
2026-07-24 9:59 ` Usama Arif [this message]
2026-07-22 15:19 ` [PATCH v5 05/11] mm: swap in PMD swap entries as whole THPs during swapoff Usama Arif
2026-07-22 15:19 ` [PATCH v5 06/11] mm: handle PMD swap entries in non-present PMD walkers Usama Arif
2026-07-22 15:19 ` [PATCH v5 07/11] mm: handle PMD swap entries in MADV_WILLNEED Usama Arif
2026-07-22 15:19 ` [PATCH v5 08/11] mm: handle PMD swap entries in UFFDIO_MOVE Usama Arif
2026-07-22 15:19 ` [PATCH v5 09/11] mm: handle PMD swap entry faults on swap-in Usama Arif
2026-07-22 15:19 ` [PATCH v5 10/11] mm: install PMD swap entries on swap-out Usama Arif
2026-07-23 19:16 ` Matthew Wilcox
2026-07-24 10:27 ` Usama Arif
2026-07-22 15:19 ` [PATCH v5 11/11] selftests/mm: add PMD swap entry tests Usama Arif
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=c40f4a8b-1f60-45ea-962d-0597c413ec3d@linux.dev \
--to=usama.arif@linux.dev \
--cc=akpm@linux-foundation.org \
--cc=alex@ghiti.fr \
--cc=alexghiti@fb.com \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=baoquan.he@linux.dev \
--cc=chrisl@kernel.org \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=hannes@cmpxchg.org \
--cc=kas@kernel.org \
--cc=kasong@tencent.com \
--cc=kernel-team@meta.com \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=nico.pache@linux.dev \
--cc=nphamcs@gmail.com \
--cc=riel@surriel.com \
--cc=ryan.roberts@arm.com \
--cc=shakeel.butt@linux.dev \
--cc=shikemeng@huaweicloud.com \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
--cc=ying.huang@linux.alibaba.com \
--cc=yosry@kernel.org \
--cc=youngjun.park@lge.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.