Linux filesystem development
 help / color / mirror / Atom feed
From: Matthew Wilcox <willy@infradead.org>
To: "David Hildenbrand (Arm)" <david@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
	Jane Chu <jane.chu@oracle.com>,
	linux-mm@kvack.org, Muchun Song <muchun.song@linux.dev>,
	Oscar Salvador <osalvador@suse.de>,
	Miaohe Lin <linmiaohe@huawei.com>,
	Naoya Horiguchi <nao.horiguchi@gmail.com>,
	Jan Kara <jack@suse.cz>,
	linux-fsdevel@vger.kernel.org,
	Christian Brauner <christian@brauner.io>,
	Jiaqi Yan <jiaqiyan@google.com>,
	"Gregory Price (Meta)" <gourry@gourry.net>
Subject: Re: [PATCH v9 08/15] hugetlb: Use the has_hwpoisoned flag
Date: Mon, 21 Sep 2026 22:32:39 +0100	[thread overview]
Message-ID: <arGidwZQrrWVPAoj@casper.infradead.org> (raw)
In-Reply-To: <7608c80e-645b-4d07-bba3-ef45ee36fba4@kernel.org>

On Fri, Sep 18, 2026 at 03:44:12PM +0200, David Hildenbrand (Arm) wrote:
> I'm a terrible person and it took me way too long to get back to this :(
> 
> > 
> > You make it sound so easy ;-)
> 
> Heh, as long as you hold a folio reference it really is :)
> 
> > 
> > If we don't hold a reference on the folio, then the folio (whether it's
> > hugetlb or not) can be split.  And then folio_test_hwpoison() can hit
> > the assertion that it's now a tail page.
> 
> Right. I am/was missing the connection to "generic_file_read_iter() in hugetlbfs".

We might be able to split this into two series at this point.  Is that
worth me spending time on doing?  Originally I was going to use
is_page_hwpoison() in filemap, then it became obvious that this was
inefficient when we had a reference to the folio, and so we ended up with
is_ref_page_hwpoison().

> > Before this patch series, we don't always hold the hugetlb lock when
> > splitting a hugetlb folio that contains hwpoison.  And we don't want
> > to have to grab the hugetlb lock if we can avoid it -- unprivileged
> > userspace can hammer these paths hard, and I don't want to see this
> > lock be contended.
> 
> Right. But doesn't at least the pagecache always hold a folio reference when
> testing for hwpoison?
> 
> I mean, there is other code that might need care (PFN walkers like memory
> offlining), but I was surprised to see pagecache code require a rework of
> lockless hwpoison checking.
> 
> I'm sure I am missing something.

It's just how everything evolved.

> > Oh, I'm not happy about it.  But we need an atomic way to determine if
> > a page belongs to a hugetlb folio with hwpoison detected.  Short of a
> > complete rearchitecture of how we handle hwpoison, this is the best I've
> > come up with.
> 
> (did I express how much I hate the hwpoison infrastructure and how it's racy
> left and right? :) )

Yes, and you'll have the chance to do it again on Wednesday ;-)

> >> (what on earth is "huge_poison" is this supposed to be "hugetlb_poison" ? Why
> >> "poison" and not "hwpoison"? Really odd)
> > 
> > We're inconsistent in our naming on both of these things.  It doesn't
> > help that somebody decided to reuse the term "poison" to mean
> > "uninitialised struct page".
> 
> Right, but let's be consistent with hugetlb and with hwpoison. ;)

I'll take another look ...

> >> I'd expect that we actually get rid of folio_test_hwpoison entirely and
> >> exclusively work on per-page state and has_hwpoison. But IIUC, now it's some
> >> mixture of folio checks, page checks, folio_has, mixed with some hugetlb oddity.
> > 
> > If we could get rid of HVO, we could do it entirely on struct page.
> 
> That's what I hope we will achieve at some point.

I'm not sure that's a realistic hope in the next few years.  Even if
we get struct page down to 8 bytes, that's 2MB of memory per 1GiB allocation
(assuming 4KiB pages).  Always a tempting target for someone looking for
memory savings.

I would prefer an out-of-line tree that doesn't rely on a bit in struct
page.  Maybe a bit in struct folio wuld be fine (which tells you whether
it's OK to skip the tree search).

> > Or, as above, entirely rearchitecture hwpoison handling to not be based
> > around pages or folios any more.
> > 
> >> I am not quite clear whether the change you propose here is actually required
> >> for the remainder of this series?
> >>
> >> 	[PATCH v9 00/15] Use generic_file_read_iter() in hugetlbfs
> >>
> >> IOW, do we really need all this hugetlb hwpoison handling just to accomplish
> >> that, or could some of that (bigger hwpoison rework) be done separately?
> > 
> > Blame Sashiko!  I'm going to ignore it in future, but it's really good
> > at nerd-sniping "hey all of this is already broken and you could fix it
> > as part of this series".
> 
> Right, and this is only the tip of the iceberg, because the entire hwpoison
> infrastructure is a complete hacked-on racy piece of ... engineering excellence.
> 
> To summarize my question: is it possible to separate for this series the
> pagecache part (Use generic_file_read_iter() in hugetlbfs) from all the hugetlb
> hwpoison rework?
> 
> OTOH, if it's really not avoidable, please let me know.

I think Sashiko will whine and complain incessantly if we drop the
race elimination patches off the front.  I can try it, but is it worth
doing?  Eliminating the races seems like good engineering anyway.

> > try_to_unmap_one() is already 200 lines and contains some tricky code flow.
> > I'm trying to at least make the problem not-worse.  I think we should
> > pull out even more code into helpers, but not in this patch series.
> 
> Note that this code now was partly reworked such that this helper will soon no
> longer be required.

Rebasing it on current Linus simplified this a lot.


  reply	other threads:[~2026-09-21 21:32 UTC|newest]

Thread overview: 54+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-05 21:05 [PATCH v9 00/15] Use generic_file_read_iter() in hugetlbfs Matthew Wilcox (Oracle)
2026-08-05 21:05 ` [PATCH v9 01/15] memory-failure: Fix hardware poison check in unpoison_memory() again Matthew Wilcox (Oracle)
2026-08-05 21:05 ` [PATCH v9 02/15] memory-failure: Prevent hugetlb freeing during unpoisoning Matthew Wilcox (Oracle)
2026-09-04  3:38   ` Miaohe Lin
2026-08-05 21:05 ` [PATCH v9 03/15] mm: Rename folio_contain_hwpoison_page() to folio_has_hwpoison_page() Matthew Wilcox (Oracle)
2026-08-13  7:20   ` David Hildenbrand (Arm)
2026-08-05 21:05 ` [PATCH v9 04/15] hugetlb: Mark some function arguments as const Matthew Wilcox (Oracle)
2026-08-05 21:05 ` [PATCH v9 05/15] guest_memfd: Use folio_has_hwpoisoned_page() Matthew Wilcox (Oracle)
2026-08-13  7:20   ` David Hildenbrand (Arm)
2026-08-13 11:50     ` Matthew Wilcox
2026-08-25 18:41       ` David Hildenbrand (Arm)
2026-08-05 21:05 ` [PATCH v9 06/15] kpageflags: Use is_page_hwpoison() to set KPF_HWPOISON Matthew Wilcox (Oracle)
2026-08-13  7:25   ` David Hildenbrand (Arm)
2026-08-13  9:01     ` David Hildenbrand (Arm)
2026-09-23 18:59       ` Matthew Wilcox
2026-08-05 21:05 ` [PATCH v9 07/15] hugetlb: Move poison to pages before clearing hugetlb page type Matthew Wilcox (Oracle)
2026-08-05 21:05 ` [PATCH v9 08/15] hugetlb: Use the has_hwpoisoned flag Matthew Wilcox (Oracle)
2026-08-13  8:29   ` David Hildenbrand (Arm)
2026-08-13 16:17     ` Matthew Wilcox
2026-09-18 13:44       ` David Hildenbrand (Arm)
2026-09-21 21:32         ` Matthew Wilcox [this message]
2026-09-23 11:40           ` David Hildenbrand (Arm)
2026-09-23 19:10             ` Matthew Wilcox
2026-09-24 10:32               ` David Hildenbrand (Arm)
2026-09-24 18:53                 ` Matthew Wilcox
2026-09-24 19:21                   ` David Hildenbrand (Arm)
2026-09-24 21:28                   ` Andrew Morton
2026-08-13 11:12   ` Pedro Falcato
2026-09-04  3:46   ` Miaohe Lin
2026-09-23 19:33     ` Matthew Wilcox
2026-08-05 21:05 ` [PATCH v9 09/15] mm: Remove locking mf_mutex in is_raw_hwpoison_page_in_hugepage() Matthew Wilcox (Oracle)
2026-09-04  6:35   ` Miaohe Lin
2026-09-23 19:34     ` Matthew Wilcox
2026-08-05 21:05 ` [PATCH v9 10/15] mm: Check individual hugetlb pages for poison Matthew Wilcox (Oracle)
2026-09-04  6:37   ` Miaohe Lin
2026-08-05 21:05 ` [PATCH v9 11/15] filemap: Add hwpoison handling to filemap_read() Matthew Wilcox (Oracle)
2026-08-13 10:55   ` Pedro Falcato
2026-08-13 16:33     ` Matthew Wilcox
2026-09-18 13:47       ` David Hildenbrand (Arm)
2026-09-21 20:08         ` Matthew Wilcox
2026-08-05 21:05 ` [PATCH v9 12/15] filemap: Remove checks in mapping_set_folio_order_range() Matthew Wilcox (Oracle)
2026-08-13 11:05   ` Pedro Falcato
2026-08-05 21:05 ` [PATCH v9 13/15] hugetlb: Set mapping folio order Matthew Wilcox (Oracle)
2026-08-13 11:06   ` Pedro Falcato
2026-09-18 13:47   ` David Hildenbrand (Arm)
2026-08-05 21:05 ` [PATCH v9 14/15] filemap: Add support for authoritative mappings Matthew Wilcox (Oracle)
2026-08-13 11:10   ` Pedro Falcato
2026-09-18 13:51   ` David Hildenbrand (Arm)
2026-09-21 21:19     ` Matthew Wilcox
2026-09-23 13:55       ` David Hildenbrand (Arm)
2026-08-05 21:05 ` [PATCH v9 15/15] hugetlb: replace hugetlbfs_read_iter() with generic_file_read_iter() Matthew Wilcox (Oracle)
2026-08-13 11:10   ` Pedro Falcato
2026-08-06 19:47 ` [PATCH v9 00/15] Use generic_file_read_iter() in hugetlbfs jane.chu
2026-08-13  8:37 ` Lorenzo Stoakes (ARM)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arGidwZQrrWVPAoj@casper.infradead.org \
    --to=willy@infradead.org \
    --cc=akpm@linux-foundation.org \
    --cc=christian@brauner.io \
    --cc=david@kernel.org \
    --cc=gourry@gourry.net \
    --cc=jack@suse.cz \
    --cc=jane.chu@oracle.com \
    --cc=jiaqiyan@google.com \
    --cc=linmiaohe@huawei.com \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=muchun.song@linux.dev \
    --cc=nao.horiguchi@gmail.com \
    --cc=osalvador@suse.de \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox