From: Matthew Wilcox <willy@infradead.org>
To: "David Hildenbrand (Arm)" <david@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
Jane Chu <jane.chu@oracle.com>,
linux-mm@kvack.org, Muchun Song <muchun.song@linux.dev>,
Oscar Salvador <osalvador@suse.de>,
Miaohe Lin <linmiaohe@huawei.com>,
Naoya Horiguchi <nao.horiguchi@gmail.com>,
Jan Kara <jack@suse.cz>,
linux-fsdevel@vger.kernel.org,
Christian Brauner <christian@brauner.io>,
Jiaqi Yan <jiaqiyan@google.com>,
"Gregory Price (Meta)" <gourry@gourry.net>
Subject: Re: [PATCH v9 08/15] hugetlb: Use the has_hwpoisoned flag
Date: Mon, 21 Sep 2026 22:32:39 +0100 [thread overview]
Message-ID: <arGidwZQrrWVPAoj@casper.infradead.org> (raw)
In-Reply-To: <7608c80e-645b-4d07-bba3-ef45ee36fba4@kernel.org>
On Fri, Sep 18, 2026 at 03:44:12PM +0200, David Hildenbrand (Arm) wrote:
> I'm a terrible person and it took me way too long to get back to this :(
>
> >
> > You make it sound so easy ;-)
>
> Heh, as long as you hold a folio reference it really is :)
>
> >
> > If we don't hold a reference on the folio, then the folio (whether it's
> > hugetlb or not) can be split. And then folio_test_hwpoison() can hit
> > the assertion that it's now a tail page.
>
> Right. I am/was missing the connection to "generic_file_read_iter() in hugetlbfs".
We might be able to split this into two series at this point. Is that
worth me spending time on doing? Originally I was going to use
is_page_hwpoison() in filemap, then it became obvious that this was
inefficient when we had a reference to the folio, and so we ended up with
is_ref_page_hwpoison().
> > Before this patch series, we don't always hold the hugetlb lock when
> > splitting a hugetlb folio that contains hwpoison. And we don't want
> > to have to grab the hugetlb lock if we can avoid it -- unprivileged
> > userspace can hammer these paths hard, and I don't want to see this
> > lock be contended.
>
> Right. But doesn't at least the pagecache always hold a folio reference when
> testing for hwpoison?
>
> I mean, there is other code that might need care (PFN walkers like memory
> offlining), but I was surprised to see pagecache code require a rework of
> lockless hwpoison checking.
>
> I'm sure I am missing something.
It's just how everything evolved.
> > Oh, I'm not happy about it. But we need an atomic way to determine if
> > a page belongs to a hugetlb folio with hwpoison detected. Short of a
> > complete rearchitecture of how we handle hwpoison, this is the best I've
> > come up with.
>
> (did I express how much I hate the hwpoison infrastructure and how it's racy
> left and right? :) )
Yes, and you'll have the chance to do it again on Wednesday ;-)
> >> (what on earth is "huge_poison" is this supposed to be "hugetlb_poison" ? Why
> >> "poison" and not "hwpoison"? Really odd)
> >
> > We're inconsistent in our naming on both of these things. It doesn't
> > help that somebody decided to reuse the term "poison" to mean
> > "uninitialised struct page".
>
> Right, but let's be consistent with hugetlb and with hwpoison. ;)
I'll take another look ...
> >> I'd expect that we actually get rid of folio_test_hwpoison entirely and
> >> exclusively work on per-page state and has_hwpoison. But IIUC, now it's some
> >> mixture of folio checks, page checks, folio_has, mixed with some hugetlb oddity.
> >
> > If we could get rid of HVO, we could do it entirely on struct page.
>
> That's what I hope we will achieve at some point.
I'm not sure that's a realistic hope in the next few years. Even if
we get struct page down to 8 bytes, that's 2MB of memory per 1GiB allocation
(assuming 4KiB pages). Always a tempting target for someone looking for
memory savings.
I would prefer an out-of-line tree that doesn't rely on a bit in struct
page. Maybe a bit in struct folio wuld be fine (which tells you whether
it's OK to skip the tree search).
> > Or, as above, entirely rearchitecture hwpoison handling to not be based
> > around pages or folios any more.
> >
> >> I am not quite clear whether the change you propose here is actually required
> >> for the remainder of this series?
> >>
> >> [PATCH v9 00/15] Use generic_file_read_iter() in hugetlbfs
> >>
> >> IOW, do we really need all this hugetlb hwpoison handling just to accomplish
> >> that, or could some of that (bigger hwpoison rework) be done separately?
> >
> > Blame Sashiko! I'm going to ignore it in future, but it's really good
> > at nerd-sniping "hey all of this is already broken and you could fix it
> > as part of this series".
>
> Right, and this is only the tip of the iceberg, because the entire hwpoison
> infrastructure is a complete hacked-on racy piece of ... engineering excellence.
>
> To summarize my question: is it possible to separate for this series the
> pagecache part (Use generic_file_read_iter() in hugetlbfs) from all the hugetlb
> hwpoison rework?
>
> OTOH, if it's really not avoidable, please let me know.
I think Sashiko will whine and complain incessantly if we drop the
race elimination patches off the front. I can try it, but is it worth
doing? Eliminating the races seems like good engineering anyway.
> > try_to_unmap_one() is already 200 lines and contains some tricky code flow.
> > I'm trying to at least make the problem not-worse. I think we should
> > pull out even more code into helpers, but not in this patch series.
>
> Note that this code now was partly reworked such that this helper will soon no
> longer be required.
Rebasing it on current Linus simplified this a lot.
next prev parent reply other threads:[~2026-09-21 21:32 UTC|newest]
Thread overview: 54+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-05 21:05 [PATCH v9 00/15] Use generic_file_read_iter() in hugetlbfs Matthew Wilcox (Oracle)
2026-08-05 21:05 ` [PATCH v9 01/15] memory-failure: Fix hardware poison check in unpoison_memory() again Matthew Wilcox (Oracle)
2026-08-05 21:05 ` [PATCH v9 02/15] memory-failure: Prevent hugetlb freeing during unpoisoning Matthew Wilcox (Oracle)
2026-09-04 3:38 ` Miaohe Lin
2026-08-05 21:05 ` [PATCH v9 03/15] mm: Rename folio_contain_hwpoison_page() to folio_has_hwpoison_page() Matthew Wilcox (Oracle)
2026-08-13 7:20 ` David Hildenbrand (Arm)
2026-08-05 21:05 ` [PATCH v9 04/15] hugetlb: Mark some function arguments as const Matthew Wilcox (Oracle)
2026-08-05 21:05 ` [PATCH v9 05/15] guest_memfd: Use folio_has_hwpoisoned_page() Matthew Wilcox (Oracle)
2026-08-13 7:20 ` David Hildenbrand (Arm)
2026-08-13 11:50 ` Matthew Wilcox
2026-08-25 18:41 ` David Hildenbrand (Arm)
2026-08-05 21:05 ` [PATCH v9 06/15] kpageflags: Use is_page_hwpoison() to set KPF_HWPOISON Matthew Wilcox (Oracle)
2026-08-13 7:25 ` David Hildenbrand (Arm)
2026-08-13 9:01 ` David Hildenbrand (Arm)
2026-09-23 18:59 ` Matthew Wilcox
2026-08-05 21:05 ` [PATCH v9 07/15] hugetlb: Move poison to pages before clearing hugetlb page type Matthew Wilcox (Oracle)
2026-08-05 21:05 ` [PATCH v9 08/15] hugetlb: Use the has_hwpoisoned flag Matthew Wilcox (Oracle)
2026-08-13 8:29 ` David Hildenbrand (Arm)
2026-08-13 16:17 ` Matthew Wilcox
2026-09-18 13:44 ` David Hildenbrand (Arm)
2026-09-21 21:32 ` Matthew Wilcox [this message]
2026-09-23 11:40 ` David Hildenbrand (Arm)
2026-09-23 19:10 ` Matthew Wilcox
2026-09-24 10:32 ` David Hildenbrand (Arm)
2026-09-24 18:53 ` Matthew Wilcox
2026-09-24 19:21 ` David Hildenbrand (Arm)
2026-09-24 21:28 ` Andrew Morton
2026-08-13 11:12 ` Pedro Falcato
2026-09-04 3:46 ` Miaohe Lin
2026-09-23 19:33 ` Matthew Wilcox
2026-08-05 21:05 ` [PATCH v9 09/15] mm: Remove locking mf_mutex in is_raw_hwpoison_page_in_hugepage() Matthew Wilcox (Oracle)
2026-09-04 6:35 ` Miaohe Lin
2026-09-23 19:34 ` Matthew Wilcox
2026-08-05 21:05 ` [PATCH v9 10/15] mm: Check individual hugetlb pages for poison Matthew Wilcox (Oracle)
2026-09-04 6:37 ` Miaohe Lin
2026-08-05 21:05 ` [PATCH v9 11/15] filemap: Add hwpoison handling to filemap_read() Matthew Wilcox (Oracle)
2026-08-13 10:55 ` Pedro Falcato
2026-08-13 16:33 ` Matthew Wilcox
2026-09-18 13:47 ` David Hildenbrand (Arm)
2026-09-21 20:08 ` Matthew Wilcox
2026-08-05 21:05 ` [PATCH v9 12/15] filemap: Remove checks in mapping_set_folio_order_range() Matthew Wilcox (Oracle)
2026-08-13 11:05 ` Pedro Falcato
2026-08-05 21:05 ` [PATCH v9 13/15] hugetlb: Set mapping folio order Matthew Wilcox (Oracle)
2026-08-13 11:06 ` Pedro Falcato
2026-09-18 13:47 ` David Hildenbrand (Arm)
2026-08-05 21:05 ` [PATCH v9 14/15] filemap: Add support for authoritative mappings Matthew Wilcox (Oracle)
2026-08-13 11:10 ` Pedro Falcato
2026-09-18 13:51 ` David Hildenbrand (Arm)
2026-09-21 21:19 ` Matthew Wilcox
2026-09-23 13:55 ` David Hildenbrand (Arm)
2026-08-05 21:05 ` [PATCH v9 15/15] hugetlb: replace hugetlbfs_read_iter() with generic_file_read_iter() Matthew Wilcox (Oracle)
2026-08-13 11:10 ` Pedro Falcato
2026-08-06 19:47 ` [PATCH v9 00/15] Use generic_file_read_iter() in hugetlbfs jane.chu
2026-08-13 8:37 ` Lorenzo Stoakes (ARM)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arGidwZQrrWVPAoj@casper.infradead.org \
--to=willy@infradead.org \
--cc=akpm@linux-foundation.org \
--cc=christian@brauner.io \
--cc=david@kernel.org \
--cc=gourry@gourry.net \
--cc=jack@suse.cz \
--cc=jane.chu@oracle.com \
--cc=jiaqiyan@google.com \
--cc=linmiaohe@huawei.com \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=muchun.song@linux.dev \
--cc=nao.horiguchi@gmail.com \
--cc=osalvador@suse.de \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox