Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Andrew Morton <akpm@linux-foundation.org>
To: Breno Leitao <leitao@debian.org>
Cc: Chris Li <chrisl@kernel.org>, Kairui Song <kasong@tencent.com>,
	Kemeng Shi <shikemeng@huaweicloud.com>,
	Nhat Pham <nphamcs@gmail.com>, Baoquan He <baoquan.he@linux.dev>,
	Barry Song <baohua@kernel.org>,
	Youngjun Park <youngjun.park@lge.com>,
	David Hildenbrand <david@kernel.org>,
	Lorenzo Stoakes <ljs@kernel.org>,
	"Liam R. Howlett" <liam@infradead.org>,
	Vlastimil Babka <vbabka@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	Michal Hocko <mhocko@suse.com>, Jann Horn <jannh@google.com>,
	Pedro Falcato <pfalcato@suse.de>, Hugh Dickins <hughd@google.com>,
	Baolin Wang <baolin.wang@linux.alibaba.com>,
	Peter Xu <peterx@redhat.com>,
	Johannes Weiner <hannes@cmpxchg.org>,
	Yosry Ahmed <yosry@kernel.org>,
	Chengming Zhou <chengming.zhou@linux.dev>,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	kernel-team@meta.com
Subject: Re: [PATCH v3 0/2] mm, swap: don't spin on a bad swap entry
Date: Sat, 29 Aug 2026 17:42:49 -0700	[thread overview]
Message-ID: <20260829174249.d723a5f0c29b9eee8d59826b@linux-foundation.org> (raw)
In-Reply-To: <20260818-swap-v3-0-d3fa52598a59@debian.org>

On Tue, 18 Aug 2026 03:06:22 -0700 Breno Leitao <leitao@debian.org> wrote:

> I've seen some machines at Meta fleet that show the following type of
> problem:
> 
> 1) It gets some weird warning:
> 
>   BUG: Bad page map in process khugepaged  pte:f000eef300000017 pmd:00000067
>   addr:00007f57c0a01000 vm_flags:20200073 anon_vma:ffff88829af7c340 mapping:0000000000000000 index:7f57c0a01
> 
> The corruption is most likely the collapse/PT_RECLAIM race fixed by
> commit 366a4532d96f ("mm: fix the race between collapse and PT_RECLAIM
> under per-vma lock"). But this series is not about this one.
> 
> 2) Then the fault never makes progress. do_swap_page() returns 0 when
>    get_swap_device() fails, so the fault is retried, reads the same
>    entry and faults again. Nothing in the round trip changes the PTE,
>    and the same line comes out on every pass:
> 
>   get_swap_device: Bad swap offset entry 3ffffffc043c5
> 
> Patch 1 makes get_swap_device() return ERR_PTR(-EIO) for a malformed
> entry, keeping NULL for a device swapoff is taking away, and converts
> the callers. No functional change expected.
> 
> Patch 2 uses that to return VM_FAULT_SIGBUS instead of retrying.
> 
> The rate limiting patch that used to open this series was split out and
> posted on its own as a backportable hotfix [1], per Andrew's request. It
> should land first: patch 1 here touches the lines next to it in
> get_swap_device(). This patch will probably conflict with [1], but the
> merge should be trivial, given the only change in [1] is the
> addition of the __ratelimited() suffix. 
> 
> 	pr_err("%s: %s%08lx\n", __func__, Bad_file, entry.val);
> 	pr_err_ratelimited("%s: %s%08lx\n", __func__, Bad_file, entry.val);
> 

Thanks.  The patches have bitrotted a little - get_swap_device() was an
easy fixup but please double-check that I didn't miss anything.

Or perhaps just refresh-retest-resend if there's any doubt.


> The rate limiting patch that used to open this series was split out and
> posted on its own as a backportable hotfix [1], per Andrew's request. It
> should land first: patch 1 here touches the lines next to it in
> get_swap_device(). This patch will probably conflict with [1], but the
> merge should be trivial, given the only change in [1] is the
> addition of the __ratelimited() suffix. 
> 
> 	pr_err("%s: %s%08lx\n", __func__, Bad_file, entry.val);
> 		pr_err_ratelimited("%s: %s%08lx\n", __func__, Bad_file, entry.val);

oh, you already said that.

As usual when inspecting our error-path code, Sashiko said "you all suck":

	https://sashiko.dev/#/patchset/20260818-swap-v3-0-d3fa52598a59@debian.org

These things do seem on-topic for the changes you're proposing here, so
please take a look?


      parent reply	other threads:[~2026-08-30  0:42 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-18 10:06 [PATCH v3 0/2] mm, swap: don't spin on a bad swap entry Breno Leitao
2026-08-18 10:06 ` [PATCH v3 1/2] mm, swap: distinguish a malformed swap entry from a dying device Breno Leitao
2026-08-18 16:18   ` Nhat Pham
2026-08-18 18:17   ` David Hildenbrand (Arm)
2026-08-18 10:06 ` [PATCH v3 2/2] mm: fail the fault on a malformed swap entry instead of retrying it Breno Leitao
2026-08-18 16:18   ` Nhat Pham
2026-08-18 18:18   ` David Hildenbrand (Arm)
2026-08-30  0:42 ` Andrew Morton [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260829174249.d723a5f0c29b9eee8d59826b@linux-foundation.org \
    --to=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=baoquan.he@linux.dev \
    --cc=chengming.zhou@linux.dev \
    --cc=chrisl@kernel.org \
    --cc=david@kernel.org \
    --cc=hannes@cmpxchg.org \
    --cc=hughd@google.com \
    --cc=jannh@google.com \
    --cc=kasong@tencent.com \
    --cc=kernel-team@meta.com \
    --cc=leitao@debian.org \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=nphamcs@gmail.com \
    --cc=peterx@redhat.com \
    --cc=pfalcato@suse.de \
    --cc=rppt@kernel.org \
    --cc=shikemeng@huaweicloud.com \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    --cc=yosry@kernel.org \
    --cc=youngjun.park@lge.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox