From: Breno Leitao <leitao@debian.org>
To: Andrew Morton <akpm@linux-foundation.org>,
Chris Li <chrisl@kernel.org>, Kairui Song <kasong@tencent.com>,
Kemeng Shi <shikemeng@huaweicloud.com>,
Nhat Pham <nphamcs@gmail.com>, Baoquan He <baoquan.he@linux.dev>,
Barry Song <baohua@kernel.org>,
Youngjun Park <youngjun.park@lge.com>,
David Hildenbrand <david@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>,
"Liam R. Howlett" <liam@infradead.org>,
Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>,
Suren Baghdasaryan <surenb@google.com>,
Michal Hocko <mhocko@suse.com>, Jann Horn <jannh@google.com>,
Pedro Falcato <pfalcato@suse.de>,
Hugh Dickins <hughd@google.com>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
Peter Xu <peterx@redhat.com>,
Johannes Weiner <hannes@cmpxchg.org>,
Yosry Ahmed <yosry@kernel.org>,
Chengming Zhou <chengming.zhou@linux.dev>
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org,
kernel-team@meta.com, Breno Leitao <leitao@debian.org>
Subject: [PATCH v2 0/3] mm, swap: don't spin or flood the console on a bad swap entry
Date: Thu, 13 Aug 2026 03:02:19 -0700 [thread overview]
Message-ID: <20260813-swap-v2-0-4a625ccabdae@debian.org> (raw)
I've seen some machines at Meta fleet that show the following type of
problem:
1) It gets some weird warning:
BUG: Bad page map in process khugepaged pte:f000eef300000017 pmd:00000067
addr:00007f57c0a01000 vm_flags:20200073 anon_vma:ffff88829af7c340 mapping:0000000000000000 index:7f57c0a01
The corruption is most likely the collapse/PT_RECLAIM race fixed by
commit 366a4532d96f ("mm: fix the race between collapse and PT_RECLAIM
under per-vma lock"). But this series is not about this one.
2) Then it floods all the monitoring of the fleet, sending the same
message in the loop, crashing the our fleet kernel monitoring
subsystem (which is the part that I am interested in protecting)
get_swap_device: Bad swap offset entry 3ffffffc043c5
For instance, in a host today it logged 6M in a few hours, and it is still
going forever. Two things go wrong.
1) get_swap_device() prints unconditionally, unlike print_bad_pte() next
door which suppresses itself with is_bad_page_map_ratelimited().
1) do_swap_page() returns 0 when get_swap_device() fails, so the
fault is retried, reads the same entry and faults again.
Nothing in the round trip changes the PTE.
Trying to fix it in a naive way:
Patch 1 is super simple, and rate limits the two prints.
Patch 2 makes get_swap_device() return ERR_PTR(-EINVAL) for an entry
that can never name a slot on any device, keeping NULL for a device
swapoff is taking away, and converts the callers. No functional change
expected.
Patch 3 uses that to return VM_FAULT_SIGBUS instead of retrying.
PS: Sashiko flagged several pre-existing issues, and get_swap_device()
returning an error opens the door to fixing some of them. For this
series, I am focused in landing the basic cases first and build on top,
if needed.
---
Changes in v2:
- Rate limit swap_dup_entry_direct()'s print too (Andrew)
- Drop "in get_swap_device()" from patch 1's subject, it now covers all
three prints
- Return ERR_PTR(-EIO) rather than ERR_PTR(-EINVAL) for a malformed
entry; -EINVAL is too soft for a corrupt page table (David)
- Document the malformed entry case in get_swap_device()'s kerneldoc,
in patch 2 instead of patch 3 (David)
- Reword patch 2's changelog, "an entry that can never name a slot on
any device" was unclear (David)
- Link to v1: https://patch.msgid.link/20260810-swap-v1-0-375ef0767206@debian.org
To: Andrew Morton <akpm@linux-foundation.org>
To: Chris Li <chrisl@kernel.org>
To: Kairui Song <kasong@tencent.com>
To: Kemeng Shi <shikemeng@huaweicloud.com>
To: Nhat Pham <nphamcs@gmail.com>
To: Baoquan He <baoquan.he@linux.dev>
To: Barry Song <baohua@kernel.org>
To: Youngjun Park <youngjun.park@lge.com>
To: David Hildenbrand <david@kernel.org>
To: Lorenzo Stoakes <ljs@kernel.org>
To: "Liam R. Howlett" <liam@infradead.org>
To: Vlastimil Babka <vbabka@kernel.org>
To: Mike Rapoport <rppt@kernel.org>
To: Suren Baghdasaryan <surenb@google.com>
To: Michal Hocko <mhocko@suse.com>
To: Jann Horn <jannh@google.com>
To: Pedro Falcato <pfalcato@suse.de>
To: Hugh Dickins <hughd@google.com>
To: Baolin Wang <baolin.wang@linux.alibaba.com>
To: Peter Xu <peterx@redhat.com>
To: Johannes Weiner <hannes@cmpxchg.org>
To: Yosry Ahmed <yosry@kernel.org>
To: Chengming Zhou <chengming.zhou@linux.dev>
Cc: linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org
---
Breno Leitao (3):
mm, swap: ratelimit bad swap entry reports
mm, swap: distinguish a malformed swap entry from a dying device
mm: fail the fault on a malformed swap entry instead of retrying it
mm/memory.c | 9 +++++++--
mm/mincore.c | 2 +-
mm/shmem.c | 2 +-
mm/swap_state.c | 4 ++--
mm/swapfile.c | 20 ++++++++++++--------
mm/userfaultfd.c | 3 ++-
mm/zswap.c | 2 +-
7 files changed, 26 insertions(+), 16 deletions(-)
---
base-commit: 6b8c8af514d739d0335f5579b585e02babe8a727
change-id: 20260810-swap-25420f9c8ba9
Best regards,
--
Breno Leitao <leitao@debian.org>
next reply other threads:[~2026-08-13 10:03 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-13 10:02 Breno Leitao [this message]
2026-08-13 10:02 ` [PATCH v2 1/3] mm, swap: ratelimit bad swap entry reports Breno Leitao
2026-08-13 10:02 ` [PATCH v2 2/3] mm, swap: distinguish a malformed swap entry from a dying device Breno Leitao
2026-08-13 10:02 ` [PATCH v2 3/3] mm: fail the fault on a malformed swap entry instead of retrying it Breno Leitao
2026-08-13 20:34 ` [PATCH v2 0/3] mm, swap: don't spin or flood the console on a bad swap entry Andrew Morton
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260813-swap-v2-0-4a625ccabdae@debian.org \
--to=leitao@debian.org \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=baoquan.he@linux.dev \
--cc=chengming.zhou@linux.dev \
--cc=chrisl@kernel.org \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=hughd@google.com \
--cc=jannh@google.com \
--cc=kasong@tencent.com \
--cc=kernel-team@meta.com \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=nphamcs@gmail.com \
--cc=peterx@redhat.com \
--cc=pfalcato@suse.de \
--cc=rppt@kernel.org \
--cc=shikemeng@huaweicloud.com \
--cc=surenb@google.com \
--cc=vbabka@kernel.org \
--cc=yosry@kernel.org \
--cc=youngjun.park@lge.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.