* [PATCH v2 1/3] mm, swap: ratelimit bad swap entry reports
2026-08-13 10:02 [PATCH v2 0/3] mm, swap: don't spin or flood the console on a bad swap entry Breno Leitao
@ 2026-08-13 10:02 ` Breno Leitao
2026-08-13 10:02 ` [PATCH v2 2/3] mm, swap: distinguish a malformed swap entry from a dying device Breno Leitao
2026-08-13 10:02 ` [PATCH v2 3/3] mm: fail the fault on a malformed swap entry instead of retrying it Breno Leitao
2 siblings, 0 replies; 4+ messages in thread
From: Breno Leitao @ 2026-08-13 10:02 UTC (permalink / raw)
To: Andrew Morton, Chris Li, Kairui Song, Kemeng Shi, Nhat Pham,
Baoquan He, Barry Song, Youngjun Park, David Hildenbrand,
Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Michal Hocko, Jann Horn, Pedro Falcato,
Hugh Dickins, Baolin Wang, Peter Xu, Johannes Weiner, Yosry Ahmed,
Chengming Zhou
Cc: linux-mm, linux-kernel, kernel-team, Breno Leitao
A corrupt page table hands the same bogus entry to get_swap_device() on
every access to the mapping, and every rejection is logged. One machine
logged 6185620 copies of the same line in a few hours.
swap_dup_entry_direct() prints the same message from the fork path, once
per call: the WARN_ON_ONCE() guarding it warns once, the pr_err() inside
does not.
Rate limit all three prints.
Signed-off-by: Breno Leitao <leitao@debian.org>
---
mm/swapfile.c | 6 +++---
1 file changed, 3 insertions(+), 3 deletions(-)
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 4d4e3e3059f6b..31c8a340606bb 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -1899,11 +1899,11 @@ struct swap_info_struct *get_swap_device(swp_entry_t entry)
return si;
bad_nofile:
- pr_err("%s: %s%08lx\n", __func__, Bad_file, entry.val);
+ pr_err_ratelimited("%s: %s%08lx\n", __func__, Bad_file, entry.val);
out:
return NULL;
put_out:
- pr_err("%s: %s%08lx\n", __func__, Bad_offset, entry.val);
+ pr_err_ratelimited("%s: %s%08lx\n", __func__, Bad_offset, entry.val);
percpu_ref_put(&si->users);
return NULL;
}
@@ -3876,7 +3876,7 @@ int swap_dup_entry_direct(swp_entry_t entry)
si = swap_entry_to_info(entry);
if (WARN_ON_ONCE(!si)) {
- pr_err("%s%08lx\n", Bad_file, entry.val);
+ pr_err_ratelimited("%s%08lx\n", Bad_file, entry.val);
return -EINVAL;
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 4+ messages in thread* [PATCH v2 2/3] mm, swap: distinguish a malformed swap entry from a dying device
2026-08-13 10:02 [PATCH v2 0/3] mm, swap: don't spin or flood the console on a bad swap entry Breno Leitao
2026-08-13 10:02 ` [PATCH v2 1/3] mm, swap: ratelimit bad swap entry reports Breno Leitao
@ 2026-08-13 10:02 ` Breno Leitao
2026-08-13 10:02 ` [PATCH v2 3/3] mm: fail the fault on a malformed swap entry instead of retrying it Breno Leitao
2 siblings, 0 replies; 4+ messages in thread
From: Breno Leitao @ 2026-08-13 10:02 UTC (permalink / raw)
To: Andrew Morton, Chris Li, Kairui Song, Kemeng Shi, Nhat Pham,
Baoquan He, Barry Song, Youngjun Park, David Hildenbrand,
Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Michal Hocko, Jann Horn, Pedro Falcato,
Hugh Dickins, Baolin Wang, Peter Xu, Johannes Weiner, Yosry Ahmed,
Chengming Zhou
Cc: linux-mm, linux-kernel, kernel-team, Breno Leitao
get_swap_device() returns NULL for two different things: an entry whose
type names no swap device or whose offset is past the end of one, and a
device that swapoff is taking away. The first never becomes valid, the
second does, and callers cannot tell them apart.
Return ERR_PTR(-EIO) for the two malformed cases and keep NULL for
swapoff. copy_nonpresent_pte() already reports -EIO for the same
corruption on the fork path.
Callers bail out on failure either way, so switch them to
IS_ERR_OR_NULL() and clear si where the cleanup path would otherwise
put an ERR_PTR. No functional change.
Signed-off-by: Breno Leitao <leitao@debian.org>
---
mm/memory.c | 6 ++++--
mm/mincore.c | 2 +-
mm/shmem.c | 2 +-
mm/swap_state.c | 4 ++--
mm/swapfile.c | 14 +++++++++-----
mm/userfaultfd.c | 3 ++-
mm/zswap.c | 2 +-
7 files changed, 20 insertions(+), 13 deletions(-)
diff --git a/mm/memory.c b/mm/memory.c
index d9cf941967cf0..7201e848129a7 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -4954,10 +4954,12 @@ vm_fault_t do_swap_page(struct vm_fault *vmf)
goto out;
}
- /* Prevent swapoff from happening to us. */
+ /* Prevent swapoff from happening to us, and reject a bad entry. */
si = get_swap_device(entry);
- if (unlikely(!si))
+ if (IS_ERR_OR_NULL(si)) {
+ si = NULL;
goto out;
+ }
folio = swap_cache_get_folio(entry);
if (folio)
diff --git a/mm/mincore.c b/mm/mincore.c
index ff4ac82817683..c086836bc4bcc 100644
--- a/mm/mincore.c
+++ b/mm/mincore.c
@@ -71,7 +71,7 @@ static unsigned char mincore_swap(swp_entry_t entry, bool shmem)
*/
if (shmem) {
si = get_swap_device(entry);
- if (!si)
+ if (IS_ERR_OR_NULL(si))
return 0;
}
folio = swap_cache_get_folio(entry);
diff --git a/mm/shmem.c b/mm/shmem.c
index 65572cbf1bd3c..d0a9f52bfed71 100644
--- a/mm/shmem.c
+++ b/mm/shmem.c
@@ -2276,7 +2276,7 @@ static int shmem_swapin_folio(struct inode *inode, pgoff_t index,
si = get_swap_device(index_entry);
order = shmem_confirm_swap(mapping, index, index_entry);
- if (unlikely(!si)) {
+ if (IS_ERR_OR_NULL(si)) {
if (order < 0)
return -EEXIST;
else
diff --git a/mm/swap_state.c b/mm/swap_state.c
index 4b7a3303c463b..f2e86d6626ecc 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -715,7 +715,7 @@ struct folio *read_swap_cache_async(struct swap_io_ctx *ctx, swp_entry_t entry,
struct folio *folio;
si = get_swap_device(entry);
- if (!si)
+ if (IS_ERR_OR_NULL(si))
return NULL;
mpol = get_vma_policy(vma, addr, 0, &ilx);
@@ -951,7 +951,7 @@ static struct folio *swap_vma_readahead(swp_entry_t targ_entry, gfp_t gfp_mask,
*/
if (swp_type(entry) != swp_type(targ_entry)) {
si = get_swap_device(entry);
- if (!si)
+ if (IS_ERR_OR_NULL(si))
continue;
}
folio = swap_cache_read_folio(&ctx, entry, gfp_mask, mpol, ilx,
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 31c8a340606bb..b96bc89815935 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -1504,7 +1504,7 @@ int swap_retry_table_alloc(swp_entry_t entry, gfp_t gfp)
unsigned long offset = swp_offset(entry);
si = get_swap_device(entry);
- if (!si)
+ if (IS_ERR_OR_NULL(si))
return 0;
ci = __swap_offset_to_cluster(si, offset);
@@ -1859,7 +1859,10 @@ void folio_put_swap(struct folio *folio, struct page *page)
* Check whether swap entry is valid in the swap device. If so,
* return pointer to swap_info_struct, and keep the swap entry valid
* via preventing the swap device from being swapoff, until
- * put_swap_device() is called. Otherwise return NULL.
+ * put_swap_device() is called. Return NULL for an empty entry or a
+ * device that is going away, and ERR_PTR(-EIO) if the entry's type
+ * names no swap device or its offset is past the end of one. These EIOs
+ * are preceded by pr_err().
*
* Notice that swapoff or swapoff+swapon can still happen before the
* percpu_ref_tryget_live() in get_swap_device() or after the
@@ -1900,12 +1903,13 @@ struct swap_info_struct *get_swap_device(swp_entry_t entry)
return si;
bad_nofile:
pr_err_ratelimited("%s: %s%08lx\n", __func__, Bad_file, entry.val);
+ return ERR_PTR(-EIO);
out:
return NULL;
put_out:
pr_err_ratelimited("%s: %s%08lx\n", __func__, Bad_offset, entry.val);
percpu_ref_put(&si->users);
- return NULL;
+ return ERR_PTR(-EIO);
}
/*
@@ -2001,7 +2005,7 @@ int swp_swapcount(swp_entry_t entry)
int count;
si = get_swap_device(entry);
- if (!si)
+ if (IS_ERR_OR_NULL(si))
return 0;
ci = swap_cluster_lock(si, swp_offset(entry));
@@ -2127,7 +2131,7 @@ void swap_put_entries_direct(swp_entry_t entry, int nr)
struct swap_info_struct *si;
si = get_swap_device(entry);
- if (WARN_ON_ONCE(!si))
+ if (WARN_ON_ONCE(IS_ERR_OR_NULL(si)))
return;
if (WARN_ON_ONCE(end_offset > si->max))
goto out;
diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c
index 24a4d92ffa3c2..bf7bc7fb1aa0f 100644
--- a/mm/userfaultfd.c
+++ b/mm/userfaultfd.c
@@ -1700,7 +1700,8 @@ static long move_pages_ptes(struct mm_struct *mm, pmd_t *dst_pmd, pmd_t *src_pmd
}
si = get_swap_device(entry);
- if (unlikely(!si)) {
+ if (IS_ERR_OR_NULL(si)) {
+ si = NULL;
ret = -EAGAIN;
goto out;
}
diff --git a/mm/zswap.c b/mm/zswap.c
index f7c9c89f6449c..bc9b931d6f447 100644
--- a/mm/zswap.c
+++ b/mm/zswap.c
@@ -997,7 +997,7 @@ static int zswap_writeback_entry(struct zswap_entry *entry,
/* try to allocate swap cache folio */
si = get_swap_device(swpentry);
- if (!si)
+ if (IS_ERR_OR_NULL(si))
return -EEXIST;
mpol = get_task_policy(current);
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 4+ messages in thread* [PATCH v2 3/3] mm: fail the fault on a malformed swap entry instead of retrying it
2026-08-13 10:02 [PATCH v2 0/3] mm, swap: don't spin or flood the console on a bad swap entry Breno Leitao
2026-08-13 10:02 ` [PATCH v2 1/3] mm, swap: ratelimit bad swap entry reports Breno Leitao
2026-08-13 10:02 ` [PATCH v2 2/3] mm, swap: distinguish a malformed swap entry from a dying device Breno Leitao
@ 2026-08-13 10:02 ` Breno Leitao
2 siblings, 0 replies; 4+ messages in thread
From: Breno Leitao @ 2026-08-13 10:02 UTC (permalink / raw)
To: Andrew Morton, Chris Li, Kairui Song, Kemeng Shi, Nhat Pham,
Baoquan He, Barry Song, Youngjun Park, David Hildenbrand,
Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Michal Hocko, Jann Horn, Pedro Falcato,
Hugh Dickins, Baolin Wang, Peter Xu, Johannes Weiner, Yosry Ahmed,
Chengming Zhou
Cc: linux-mm, linux-kernel, kernel-team, Breno Leitao
do_swap_page() returns 0 when get_swap_device() fails, which the fault
handler reads as "handled". For an entry that can never become valid
the retry takes the same fault again, so the thread spins forever,
retrying on the same fault.
Return VM_FAULT_SIGBUS (Bad access) for a malformed entry (pr_err() was
called at get_swap_device()).
Signed-off-by: Breno Leitao <leitao@debian.org>
---
mm/memory.c | 3 +++
1 file changed, 3 insertions(+)
diff --git a/mm/memory.c b/mm/memory.c
index 7201e848129a7..fa2b3d2ad3202 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -4957,6 +4957,9 @@ vm_fault_t do_swap_page(struct vm_fault *vmf)
/* Prevent swapoff from happening to us, and reject a bad entry. */
si = get_swap_device(entry);
if (IS_ERR_OR_NULL(si)) {
+ /* A malformed entry never becomes valid, so don't retry it. */
+ if (IS_ERR(si))
+ ret = VM_FAULT_SIGBUS;
si = NULL;
goto out;
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 4+ messages in thread