linux-mm.kvack.org archive mirror
 help / color / mirror / Atom feed
* [PATCH 0/3] mm, swap: don't spin or flood the console on a bad swap entry
@ 2026-08-10 16:26 Breno Leitao
  2026-08-10 16:26 ` [PATCH 1/3] mm, swap: ratelimit bad swap entry reports in get_swap_device() Breno Leitao
                   ` (3 more replies)
  0 siblings, 4 replies; 5+ messages in thread
From: Breno Leitao @ 2026-08-10 16:26 UTC (permalink / raw)
  To: Andrew Morton, Chris Li, Kairui Song, Kemeng Shi, Nhat Pham,
	Baoquan He, Barry Song, Youngjun Park, David Hildenbrand,
	Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
	Suren Baghdasaryan, Michal Hocko, Jann Horn, Pedro Falcato,
	Hugh Dickins, Baolin Wang, Peter Xu, Johannes Weiner, Yosry Ahmed,
	Chengming Zhou
  Cc: linux-mm, linux-kernel, kernel-team, Breno Leitao

I've seen some machines at Meta flete that show the following type of
problem:

1) It gets some weird warning:

  BUG: Bad page map in process khugepaged  pte:f000eef300000017 pmd:00000067
  addr:00007f57c0a01000 vm_flags:20200073 anon_vma:ffff88829af7c340 mapping:0000000000000000 index:7f57c0a01

The corruption is most likely the collapse/PT_RECLAIM race fixed by
commit 366a4532d96f ("mm: fix the race between collapse and PT_RECLAIM
under per-vma lock"). But this series is not about tha.

2) Then it floods all the monitoring of the fleet, sending the same
   message in the loop, crashing the our fleet kernel monitoring
   subsystem (which is the part that I am interested in protecting)

  get_swap_device: Bad swap offset entry 3ffffffc043c5

For instance, in a host today it logged 6M in a few hours, and it is still
going forever. Two things go wrong.

1) get_swap_device() prints unconditionally, unlike print_bad_pte() next
   door which suppresses itself with is_bad_page_map_ratelimited().

1) do_swap_page() returns 0 when get_swap_device() fails, so the
   fault is retried, reads the same entry and faults again.
   Nothing in the round trip changes the PTE.

Trying to fix it in a naive way:

Patch 1 is super simple, and rate limits the two prints.

Patch 2 makes get_swap_device() return ERR_PTR(-EINVAL) for an entry
that can never name a slot on any device, keeping NULL for a device
swapoff is taking away, and converts the callers. No functional change
expected.

Patch 3 uses that to return VM_FAULT_SIGBUS instead of retrying.

---
Breno Leitao (3):
      mm, swap: ratelimit bad swap entry reports in get_swap_device()
      mm, swap: distinguish a malformed swap entry from a dying device
      mm: fail the fault on a malformed swap entry instead of retrying it

 mm/memory.c      |  7 ++++++-
 mm/mincore.c     |  2 +-
 mm/shmem.c       |  2 +-
 mm/swap_state.c  |  4 ++--
 mm/swapfile.c    | 17 ++++++++++-------
 mm/userfaultfd.c |  3 ++-
 mm/zswap.c       |  2 +-
 7 files changed, 23 insertions(+), 14 deletions(-)
---
base-commit: 6b8c8af514d739d0335f5579b585e02babe8a727
change-id: 20260810-swap-25420f9c8ba9

Best regards,
--  
Breno Leitao <leitao@debian.org>



^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-08-10 17:20 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-10 16:26 [PATCH 0/3] mm, swap: don't spin or flood the console on a bad swap entry Breno Leitao
2026-08-10 16:26 ` [PATCH 1/3] mm, swap: ratelimit bad swap entry reports in get_swap_device() Breno Leitao
2026-08-10 16:26 ` [PATCH 2/3] mm, swap: distinguish a malformed swap entry from a dying device Breno Leitao
2026-08-10 16:26 ` [PATCH 3/3] mm: fail the fault on a malformed swap entry instead of retrying it Breno Leitao
2026-08-10 17:20 ` [PATCH 0/3] mm, swap: don't spin or flood the console on a bad swap entry Andrew Morton

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).