From: Youngjun Park <youngjun.park@lge.com>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: Chris Li <chrisl@kernel.org>, Kairui Song <kasong@tencent.com>,
Kemeng Shi <shikemeng@huaweicloud.com>,
Nhat Pham <nphamcs@gmail.com>, Baoquan He <baoquan.he@linux.dev>,
Barry Song <baohua@kernel.org>,
Jianyue Wu <wujianyue000@gmail.com>,
her0gyugyu@gmail.com, linux-mm@kvack.org,
linux-kernel@vger.kernel.org,
Youngjun Park <youngjun.park@lge.com>
Subject: [PATCH v3 2/4] mm, swap: only allow swapped-out slots into the swap cache
Date: Tue, 11 Aug 2026 22:22:07 +0900 [thread overview]
Message-ID: <20260811132209.2862708-3-youngjun.park@lge.com> (raw)
In-Reply-To: <20260811132209.2862708-1-youngjun.park@lge.com>
__swap_cache_add_check() turns away folio entries and slots with no count
and lets everything else in. That is safe only when the caller owns the
slot. Cluster readahead owns nothing, it walks a raw page_cluster sized
window of offsets around the faulting entry, so it can land on any slot.
A bad slot gets in. The check reads the count with __swp_tb_get_count(),
which shifts the count bits out without looking at the type, and
SWP_TB_BAD has all of them set, so the slot reads as SWP_TB_COUNT_MAX.
Readahead then allocates a folio and reads the offset off the device for a
slot nothing will ever swap in, and the folio entry that replaces it drops
the bad marker.
Readahead used to be guarded by swap_entry_swapped(), which goes through
swp_tb_get_count() and gets -EINVAL for a bad slot. That call went away
when the swap cache checks moved into __swap_cache_add_check(), and the
raw accessor there does not do the same type test.
Require a shadow entry instead. A slot dropped from the swap cache always
gets one, empty if there is no workingset value. The type test runs first,
so the count is only read off a countable entry, and the check as a whole
runs before the folio allocation in __swap_cache_alloc().
The large folio walk in the same function does the same raw reads. A bad
slot cannot be in its range, but the range is not pinned, so a slot freed
and then taken by hibernation still trips the countable assertion there.
Give the walk the same shadow test, the folio check folds into it.
Reproduced with a badpages list written into the swap header by hand.
Readahead took over four bad slots before this patch and none after. It
needs a crafted header, so a normal setup will not hit it.
Fixes: e1e6750df3b4 ("mm, swap: add support for stable large allocation in swap cache directly")
Acked-by: Kairui Song <kasong@tencent.com>
Signed-off-by: Youngjun Park <youngjun.park@lge.com>
---
mm/swap_state.c | 11 ++++++++---
1 file changed, 8 insertions(+), 3 deletions(-)
diff --git a/mm/swap_state.c b/mm/swap_state.c
index b76eb3d876fd..6341f1bfffa2 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -181,9 +181,14 @@ static int __swap_cache_add_check(struct swap_cluster_info *ci,
old_tb = __swap_table_get(ci, ci_off);
if (swp_tb_is_folio(old_tb))
return -EEXIST;
- if (!__swp_tb_get_count(old_tb))
+ /*
+ * Only a swapped-out slot may be brought into the swap cache.
+ * Cluster readahead walks raw offset ranges, so it can land on
+ * slots that are free, bad, or owned by hibernation.
+ */
+ if (!swp_tb_is_shadow(old_tb) || !__swp_tb_get_count(old_tb))
return -ENOENT;
- if (shadowp && swp_tb_is_shadow(old_tb))
+ if (shadowp)
*shadowp = swp_tb_to_shadow(old_tb);
if (memcg_id)
*memcg_id = __swap_cgroup_get(ci, ci_off);
@@ -196,7 +201,7 @@ static int __swap_cache_add_check(struct swap_cluster_info *ci,
ci_end = ci_off + nr;
do {
old_tb = __swap_table_get(ci, ci_off);
- if (unlikely(swp_tb_is_folio(old_tb) ||
+ if (unlikely(!swp_tb_is_shadow(old_tb) ||
!__swp_tb_get_count(old_tb) ||
is_zero != __swap_table_test_zero(ci, ci_off) ||
(memcg_id && *memcg_id != __swap_cgroup_get(ci, ci_off))))
--
2.48.1
next prev parent reply other threads:[~2026-08-11 13:22 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-11 13:22 [PATCH v3 0/4] mm, swap: keep hibernation swap slots out of the swap cache Youngjun Park
2026-08-11 13:22 ` [PATCH v3 1/4] mm, swap: don't free a hibernation slot that is in " Youngjun Park
2026-08-11 13:22 ` Youngjun Park [this message]
2026-08-11 13:22 ` [PATCH v3 3/4] mm, swap: give hibernation swap slots their own swap table entry type Youngjun Park
2026-08-11 16:48 ` Kairui Song
2026-08-11 13:22 ` [PATCH v3 4/4] mm, swap: drop the swap cache guard and reclaim in swap_free_hibernation_slot() Youngjun Park
2026-08-11 17:09 ` Kairui Song
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260811132209.2862708-3-youngjun.park@lge.com \
--to=youngjun.park@lge.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baoquan.he@linux.dev \
--cc=chrisl@kernel.org \
--cc=her0gyugyu@gmail.com \
--cc=kasong@tencent.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=nphamcs@gmail.com \
--cc=shikemeng@huaweicloud.com \
--cc=wujianyue000@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox