* [PATCH v2 1/4] mm, swap: don't free a hibernation slot that is in the swap cache
2026-08-09 14:45 [PATCH v2 0/4] mm, swap: keep hibernation swap slots out of the swap cache Youngjun Park
@ 2026-08-09 14:45 ` Youngjun Park
2026-08-09 14:45 ` [PATCH v2 2/4] mm, swap: only allow swapped-out slots into " Youngjun Park
` (2 subsequent siblings)
3 siblings, 0 replies; 5+ messages in thread
From: Youngjun Park @ 2026-08-09 14:45 UTC (permalink / raw)
To: Andrew Morton
Cc: Chris Li, Kairui Song, Kemeng Shi, Nhat Pham, Baoquan He,
Barry Song, Jianyue Wu, Youngjun Park, her0gyugyu, linux-mm,
linux-kernel, stable
A slot with a folio in the swap cache is freed when the folio leaves the
cache, not when its count drops. swap_put_entries_cluster() follows that
rule. swap_free_hibernation_slot() does not, it calls
__swap_cluster_free_entries() whether or not a folio sits on the slot.
Cluster readahead can put one there. It walks a raw page_cluster sized
window of offsets around the faulting entry, and a hibernation slot passes
__swap_cache_add_check() because it is not a folio and its count is not
zero. Freeing the slot then clears the entry under that folio.
The folio is now unreachable from the swap table, and the offset goes back
to the allocator. The folio is still on the LRU though, so reclaim can
pick it up later. It then takes the old offset out of folio->swap and
overwrites the table entry there, which by then may belong to someone else.
Check for a cached folio before freeing. The slot is then left in the
ordinary state where only the swap cache holds it, and it is freed when the
folio leaves the cache, either through the reclaim below or through normal
reclaim later.
Fixes: 0d6af9bcf383 ("mm, swap: use the swap table to track the swap count")
Cc: <stable@vger.kernel.org>
Acked-by: Kairui Song <kasong@tencent.com>
Signed-off-by: Youngjun Park <youngjun.park@lge.com>
---
mm/swapfile.c | 9 ++++++++-
1 file changed, 8 insertions(+), 1 deletion(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index dea2d3b36e06..f5dfc7e59191 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -2202,7 +2202,14 @@ void swap_free_hibernation_slot(swp_entry_t entry)
ci = swap_cluster_lock(si, offset);
__swap_cluster_put_entry(ci, offset % SWAPFILE_CLUSTER);
- __swap_cluster_free_entries(si, ci, offset % SWAPFILE_CLUSTER, 1);
+ /*
+ * A slot with a folio in the swap cache is freed when the folio
+ * leaves the cache, the same rule swap_put_entries_cluster() follows.
+ * Readahead can put a folio here, and freeing the slot now would
+ * leave that folio with no entry behind it.
+ */
+ if (!swp_tb_is_folio(__swap_table_get(ci, offset % SWAPFILE_CLUSTER)))
+ __swap_cluster_free_entries(si, ci, offset % SWAPFILE_CLUSTER, 1);
swap_cluster_unlock(ci);
/* In theory readahead might add it to the swap cache by accident */
--
2.48.1
^ permalink raw reply related [flat|nested] 5+ messages in thread* [PATCH v2 2/4] mm, swap: only allow swapped-out slots into the swap cache
2026-08-09 14:45 [PATCH v2 0/4] mm, swap: keep hibernation swap slots out of the swap cache Youngjun Park
2026-08-09 14:45 ` [PATCH v2 1/4] mm, swap: don't free a hibernation slot that is in " Youngjun Park
@ 2026-08-09 14:45 ` Youngjun Park
2026-08-09 14:45 ` [PATCH v2 3/4] mm, swap: give hibernation swap slots their own swap table entry type Youngjun Park
2026-08-09 14:45 ` [PATCH v2 4/4] mm, swap: drop the swap cache guard and reclaim in swap_free_hibernation_slot() Youngjun Park
3 siblings, 0 replies; 5+ messages in thread
From: Youngjun Park @ 2026-08-09 14:45 UTC (permalink / raw)
To: Andrew Morton
Cc: Chris Li, Kairui Song, Kemeng Shi, Nhat Pham, Baoquan He,
Barry Song, Jianyue Wu, Youngjun Park, her0gyugyu, linux-mm,
linux-kernel
__swap_cache_add_check() turns away folio entries and slots with no count
and lets everything else in. That is safe only when the caller owns the
slot. Cluster readahead owns nothing, it walks a raw page_cluster sized
window of offsets around the faulting entry, so it can land on any slot.
A bad slot gets in. The check reads the count with __swp_tb_get_count(),
which shifts the count bits out without looking at the type, and
SWP_TB_BAD has all of them set, so the slot reads as SWP_TB_COUNT_MAX.
Readahead then allocates a folio and reads the offset off the device for a
slot nothing will ever swap in, and the folio entry that replaces it drops
the bad marker.
Readahead used to be guarded by swap_entry_swapped(), which goes through
swp_tb_get_count() and gets -EINVAL for a bad slot. That call went away
when the swap cache checks moved into __swap_cache_add_check(), and the
raw accessor there does not do the same type test.
Require a shadow entry instead. A slot dropped from the swap cache always
gets one, empty if there is no workingset value. The type test runs first,
so the count is only read off a countable entry, and the check as a whole
runs before the folio allocation in __swap_cache_alloc().
Reproduced with a badpages list written into the swap header by hand.
Readahead took over four bad slots before this patch and none after. It
needs a crafted header, so a normal setup will not hit it.
Fixes: e1e6750df3b4 ("mm, swap: add support for stable large allocation in swap cache directly")
Signed-off-by: Youngjun Park <youngjun.park@lge.com>
---
mm/swap_state.c | 9 +++++++--
1 file changed, 7 insertions(+), 2 deletions(-)
diff --git a/mm/swap_state.c b/mm/swap_state.c
index 5be825911e64..9f2cc5918713 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -180,9 +180,14 @@ static int __swap_cache_add_check(struct swap_cluster_info *ci,
old_tb = __swap_table_get(ci, ci_off);
if (swp_tb_is_folio(old_tb))
return -EEXIST;
- if (!__swp_tb_get_count(old_tb))
+ /*
+ * Only a swapped-out slot may be brought into the swap cache.
+ * Cluster readahead walks raw offset ranges, so it can land on
+ * slots that are free, bad, or owned by hibernation.
+ */
+ if (!swp_tb_is_shadow(old_tb) || !__swp_tb_get_count(old_tb))
return -ENOENT;
- if (shadowp && swp_tb_is_shadow(old_tb))
+ if (shadowp)
*shadowp = swp_tb_to_shadow(old_tb);
if (memcg_id)
*memcg_id = __swap_cgroup_get(ci, ci_off);
--
2.48.1
^ permalink raw reply related [flat|nested] 5+ messages in thread* [PATCH v2 3/4] mm, swap: give hibernation swap slots their own swap table entry type
2026-08-09 14:45 [PATCH v2 0/4] mm, swap: keep hibernation swap slots out of the swap cache Youngjun Park
2026-08-09 14:45 ` [PATCH v2 1/4] mm, swap: don't free a hibernation slot that is in " Youngjun Park
2026-08-09 14:45 ` [PATCH v2 2/4] mm, swap: only allow swapped-out slots into " Youngjun Park
@ 2026-08-09 14:45 ` Youngjun Park
2026-08-09 14:45 ` [PATCH v2 4/4] mm, swap: drop the swap cache guard and reclaim in swap_free_hibernation_slot() Youngjun Park
3 siblings, 0 replies; 5+ messages in thread
From: Youngjun Park @ 2026-08-09 14:45 UTC (permalink / raw)
To: Andrew Morton
Cc: Chris Li, Kairui Song, Kemeng Shi, Nhat Pham, Baoquan He,
Barry Song, Jianyue Wu, Youngjun Park, her0gyugyu, linux-mm,
linux-kernel
swap_alloc_hibernation_slot() stores a fake shadow in the slot it hands
out. An anon slot swapped out with no workingset shadow looks exactly the
same, so nothing in mm can tell the two apart.
Give hibernation slots their own type. Bit 4 and every bit above it are
set, except the count field, which stays 0. Bits 0 to 3 are taken by the
shadow, PFN, pointer and bad marks, so bit 4 is the first free one. The
type holds no data, so the value alone says what it is.
The entry has no swap count. Hibernation only allocates and frees a slot,
so a count would never change. swap_free_hibernation_slot() frees the slot
directly, there is no count to put first.
The slot is no longer a shadow, so the previous patch keeps it out of the
swap cache. The count field stays 0 as well, so code that reads the count
without checking the type sees an unused slot instead of one at
SWP_TB_COUNT_MAX, and a wrong put is caught by the existing underflow
check.
Suggested-by: Kairui Song <kasong@tencent.com>
Link: https://lore.kernel.org/linux-mm/abp7aDgYLrxF3Me8@KASONG-MC4/
Signed-off-by: Youngjun Park <youngjun.park@lge.com>
---
mm/swap_table.h | 13 +++++++++++++
mm/swapfile.c | 13 +++++++------
2 files changed, 20 insertions(+), 6 deletions(-)
diff --git a/mm/swap_table.h b/mm/swap_table.h
index e6613e62f8d0..b916a6493521 100644
--- a/mm/swap_table.h
+++ b/mm/swap_table.h
@@ -30,6 +30,7 @@ struct swap_memcg_table {
* PFN: |SWAP_COUNT|Z|------ PFN -------|10| - Cached slot
* Pointer: |----------- Pointer ----------|100| - (Unused)
* Bad: |------------- 1 -------------|1000| - Bad slot
+ * Hibern: | 0 |------- 1 -------|10000| - Hibernation slot
*
* COUNT is `SWP_TB_COUNT_BITS` long, Z is the `SWP_TB_ZERO_FLAG` bit,
* and together they form the `SWP_TB_FLAGS_BITS` wide flags field.
@@ -54,6 +55,10 @@ struct swap_memcg_table {
* aligned pointers.
*
* - Bad: Swap slot is reserved, protects swap header or holes on swap devices.
+ *
+ * - Hibern: Swap slot is reserved by hibernation for the suspend image, and
+ * must never enter the swap cache. The count field is kept 0 so it never
+ * reads as a slot in use.
*/
/* NULL Entry, all 0 */
@@ -81,6 +86,9 @@ struct swap_memcg_table {
/* Bad slot: ends with 0b1000 and rests of bits are all 1 */
#define SWP_TB_BAD ((~0UL) << 3)
+/* Hibernation slot: ends with 0b10000, no count, rests of bits are all 1 */
+#define SWP_TB_HIB (((~0UL) << 4) & ~SWP_TB_COUNT_MASK)
+
/* Macro for shadow offset calculation */
#define SWAP_COUNT_SHIFT SWP_TB_FLAGS_BITS
@@ -166,6 +174,11 @@ static inline bool swp_tb_is_bad(unsigned long swp_tb)
return swp_tb == SWP_TB_BAD;
}
+static inline bool swp_tb_is_hibernation(unsigned long swp_tb)
+{
+ return swp_tb == SWP_TB_HIB;
+}
+
static inline bool swp_tb_is_countable(unsigned long swp_tb)
{
return (swp_tb_is_shadow(swp_tb) || swp_tb_is_folio(swp_tb) ||
diff --git a/mm/swapfile.c b/mm/swapfile.c
index f5dfc7e59191..a337387f7431 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -928,7 +928,7 @@ static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
* upon folio unmap.
*
* Else, it's a exclusive order 0 allocation for hibernation.
- * The slot starts with count == 1 and never increases.
+ * The slot carries no swap count and is freed by offset.
*/
if (likely(folio)) {
order = folio_order(folio);
@@ -940,8 +940,8 @@ static bool __swap_cluster_alloc_entries(struct swap_info_struct *si,
order = 0;
nr_pages = 1;
swap_cluster_assert_empty(ci, ci_off, 1, false);
- /* Fake shadow placeholder with no flag, hibernation does not use the zeromap */
- __swap_table_set(ci, ci_off, __swp_tb_mk_count(shadow_to_swp_tb(NULL, 0), 1));
+ /* Exclusively owned by hibernation, must never enter the swap cache */
+ __swap_table_set(ci, ci_off, SWP_TB_HIB);
} else {
/* Allocation without folio is only possible with hibernation */
WARN_ON_ONCE(1);
@@ -1929,9 +1929,11 @@ void __swap_cluster_free_entries(struct swap_info_struct *si,
old_tb = __swap_table_get(ci, ci_off);
/*
* Freeing is done after release of the last swap count
- * ref, or after swap cache is dropped
+ * ref, or after swap cache is dropped. A hibernation slot
+ * has no count and is freed directly by its owner.
*/
- VM_WARN_ON(!swp_tb_is_shadow(old_tb) || __swp_tb_get_count(old_tb) > 1);
+ VM_WARN_ON(!swp_tb_is_hibernation(old_tb) &&
+ (!swp_tb_is_shadow(old_tb) || __swp_tb_get_count(old_tb) > 1));
/* Resetting the slot to NULL also clears the inline flags. */
__swap_table_set(ci, ci_off, null_to_swp_tb());
@@ -2201,7 +2203,6 @@ void swap_free_hibernation_slot(swp_entry_t entry)
pgoff_t offset = swp_offset(entry);
ci = swap_cluster_lock(si, offset);
- __swap_cluster_put_entry(ci, offset % SWAPFILE_CLUSTER);
/*
* A slot with a folio in the swap cache is freed when the folio
* leaves the cache, the same rule swap_put_entries_cluster() follows.
--
2.48.1
^ permalink raw reply related [flat|nested] 5+ messages in thread* [PATCH v2 4/4] mm, swap: drop the swap cache guard and reclaim in swap_free_hibernation_slot()
2026-08-09 14:45 [PATCH v2 0/4] mm, swap: keep hibernation swap slots out of the swap cache Youngjun Park
` (2 preceding siblings ...)
2026-08-09 14:45 ` [PATCH v2 3/4] mm, swap: give hibernation swap slots their own swap table entry type Youngjun Park
@ 2026-08-09 14:45 ` Youngjun Park
3 siblings, 0 replies; 5+ messages in thread
From: Youngjun Park @ 2026-08-09 14:45 UTC (permalink / raw)
To: Andrew Morton
Cc: Chris Li, Kairui Song, Kemeng Shi, Nhat Pham, Baoquan He,
Barry Song, Jianyue Wu, Youngjun Park, her0gyugyu, linux-mm,
linux-kernel
Both are there for a folio that readahead might have put on the slot. A
hibernation entry is not a shadow and has no swap count, so the swap cache
turns it away and no such folio can exist.
Signed-off-by: Youngjun Park <youngjun.park@lge.com>
---
mm/swapfile.c | 12 +-----------
1 file changed, 1 insertion(+), 11 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index a337387f7431..4fba1770ce61 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -2203,18 +2203,8 @@ void swap_free_hibernation_slot(swp_entry_t entry)
pgoff_t offset = swp_offset(entry);
ci = swap_cluster_lock(si, offset);
- /*
- * A slot with a folio in the swap cache is freed when the folio
- * leaves the cache, the same rule swap_put_entries_cluster() follows.
- * Readahead can put a folio here, and freeing the slot now would
- * leave that folio with no entry behind it.
- */
- if (!swp_tb_is_folio(__swap_table_get(ci, offset % SWAPFILE_CLUSTER)))
- __swap_cluster_free_entries(si, ci, offset % SWAPFILE_CLUSTER, 1);
+ __swap_cluster_free_entries(si, ci, offset % SWAPFILE_CLUSTER, 1);
swap_cluster_unlock(ci);
-
- /* In theory readahead might add it to the swap cache by accident */
- __try_to_reclaim_swap(si, offset, TTRS_ANYWAY);
}
static int __find_hibernation_swap_type(dev_t device, sector_t offset)
--
2.48.1
^ permalink raw reply related [flat|nested] 5+ messages in thread