From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Matthew Wilcox <willy@infradead.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
Jane Chu <jane.chu@oracle.com>,
linux-mm@kvack.org, Muchun Song <muchun.song@linux.dev>,
Oscar Salvador <osalvador@suse.de>,
Miaohe Lin <linmiaohe@huawei.com>,
Naoya Horiguchi <nao.horiguchi@gmail.com>,
Jan Kara <jack@suse.cz>,
linux-fsdevel@vger.kernel.org,
Christian Brauner <christian@brauner.io>,
Jiaqi Yan <jiaqiyan@google.com>,
"Gregory Price (Meta)" <gourry@gourry.net>
Subject: Re: [PATCH v9 08/15] hugetlb: Use the has_hwpoisoned flag
Date: Wed, 23 Sep 2026 13:40:24 +0200 [thread overview]
Message-ID: <7a23511f-365f-436f-b934-6a520510409c@kernel.org> (raw)
In-Reply-To: <arGidwZQrrWVPAoj@casper.infradead.org>
On 9/21/26 23:32, Matthew Wilcox wrote:
> On Fri, Sep 18, 2026 at 03:44:12PM +0200, David Hildenbrand (Arm) wrote:
>> I'm a terrible person and it took me way too long to get back to this :(
>>
>>>
>>> You make it sound so easy ;-)
>>
>> Heh, as long as you hold a folio reference it really is :)
>>
>>>
>>> If we don't hold a reference on the folio, then the folio (whether it's
>>> hugetlb or not) can be split. And then folio_test_hwpoison() can hit
>>> the assertion that it's now a tail page.
>>
>> Right. I am/was missing the connection to "generic_file_read_iter() in hugetlbfs".
>
> We might be able to split this into two series at this point.
That would be good, because I suspect the "generic_file_read_iter() in
hugetlbfs" part is less controversial than the hugetlb cleanup and we can just
merge that easily without touching too much other code.
[...]
>>> If we could get rid of HVO, we could do it entirely on struct page.
>>
>> That's what I hope we will achieve at some point.
>
> I'm not sure that's a realistic hope in the next few years.
:(
> Even if
> we get struct page down to 8 bytes, that's 2MB of memory per 1GiB allocation
> (assuming 4KiB pages). Always a tempting target for someone looking for
> memory savings.
Ack
>
> I would prefer an out-of-line tree that doesn't rely on a bit in struct
> page. Maybe a bit in struct folio wuld be fine (which tells you whether
> it's OK to skip the tree search).
Yes, that would likely be better. A single bit would likely just remove any
overhead of checking in the common case (no hwpoison).
Not sure what to do with non-folio things.
>
>>> Or, as above, entirely rearchitecture hwpoison handling to not be based
>>> around pages or folios any more.
>>>
>>>
>>> Blame Sashiko! I'm going to ignore it in future, but it's really good
>>> at nerd-sniping "hey all of this is already broken and you could fix it
>>> as part of this series".
>>
>> Right, and this is only the tip of the iceberg, because the entire hwpoison
>> infrastructure is a complete hacked-on racy piece of ... engineering excellence.
>>
>> To summarize my question: is it possible to separate for this series the
>> pagecache part (Use generic_file_read_iter() in hugetlbfs) from all the hugetlb
>> hwpoison rework?
>>
>> OTOH, if it's really not avoidable, please let me know.
>
> I think Sashiko will whine and complain incessantly if we drop the
> race elimination patches off the front. I can try it, but is it worth
> doing?
Note that Sashiko will apparently no longer complain about preexisting stuff.
I've been told that these reports are now collected elsewhere.
So I wouldn't worry about Sashiko crying about hwpoison stuff that was broken
for ever. I'd rather have fixes for that handled separately (after we properly
discuss a better long-term solution to handle most these races and where to
store hwpoison information long-term).
--
Cheers,
David
next prev parent reply other threads:[~2026-09-23 11:40 UTC|newest]
Thread overview: 54+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-05 21:05 [PATCH v9 00/15] Use generic_file_read_iter() in hugetlbfs Matthew Wilcox (Oracle)
2026-08-05 21:05 ` [PATCH v9 01/15] memory-failure: Fix hardware poison check in unpoison_memory() again Matthew Wilcox (Oracle)
2026-08-05 21:05 ` [PATCH v9 02/15] memory-failure: Prevent hugetlb freeing during unpoisoning Matthew Wilcox (Oracle)
2026-09-04 3:38 ` Miaohe Lin
2026-08-05 21:05 ` [PATCH v9 03/15] mm: Rename folio_contain_hwpoison_page() to folio_has_hwpoison_page() Matthew Wilcox (Oracle)
2026-08-13 7:20 ` David Hildenbrand (Arm)
2026-08-05 21:05 ` [PATCH v9 04/15] hugetlb: Mark some function arguments as const Matthew Wilcox (Oracle)
2026-08-05 21:05 ` [PATCH v9 05/15] guest_memfd: Use folio_has_hwpoisoned_page() Matthew Wilcox (Oracle)
2026-08-13 7:20 ` David Hildenbrand (Arm)
2026-08-13 11:50 ` Matthew Wilcox
2026-08-25 18:41 ` David Hildenbrand (Arm)
2026-08-05 21:05 ` [PATCH v9 06/15] kpageflags: Use is_page_hwpoison() to set KPF_HWPOISON Matthew Wilcox (Oracle)
2026-08-13 7:25 ` David Hildenbrand (Arm)
2026-08-13 9:01 ` David Hildenbrand (Arm)
2026-09-23 18:59 ` Matthew Wilcox
2026-08-05 21:05 ` [PATCH v9 07/15] hugetlb: Move poison to pages before clearing hugetlb page type Matthew Wilcox (Oracle)
2026-08-05 21:05 ` [PATCH v9 08/15] hugetlb: Use the has_hwpoisoned flag Matthew Wilcox (Oracle)
2026-08-13 8:29 ` David Hildenbrand (Arm)
2026-08-13 16:17 ` Matthew Wilcox
2026-09-18 13:44 ` David Hildenbrand (Arm)
2026-09-21 21:32 ` Matthew Wilcox
2026-09-23 11:40 ` David Hildenbrand (Arm) [this message]
2026-09-23 19:10 ` Matthew Wilcox
2026-09-24 10:32 ` David Hildenbrand (Arm)
2026-09-24 18:53 ` Matthew Wilcox
2026-09-24 19:21 ` David Hildenbrand (Arm)
2026-09-24 21:28 ` Andrew Morton
2026-08-13 11:12 ` Pedro Falcato
2026-09-04 3:46 ` Miaohe Lin
2026-09-23 19:33 ` Matthew Wilcox
2026-08-05 21:05 ` [PATCH v9 09/15] mm: Remove locking mf_mutex in is_raw_hwpoison_page_in_hugepage() Matthew Wilcox (Oracle)
2026-09-04 6:35 ` Miaohe Lin
2026-09-23 19:34 ` Matthew Wilcox
2026-08-05 21:05 ` [PATCH v9 10/15] mm: Check individual hugetlb pages for poison Matthew Wilcox (Oracle)
2026-09-04 6:37 ` Miaohe Lin
2026-08-05 21:05 ` [PATCH v9 11/15] filemap: Add hwpoison handling to filemap_read() Matthew Wilcox (Oracle)
2026-08-13 10:55 ` Pedro Falcato
2026-08-13 16:33 ` Matthew Wilcox
2026-09-18 13:47 ` David Hildenbrand (Arm)
2026-09-21 20:08 ` Matthew Wilcox
2026-08-05 21:05 ` [PATCH v9 12/15] filemap: Remove checks in mapping_set_folio_order_range() Matthew Wilcox (Oracle)
2026-08-13 11:05 ` Pedro Falcato
2026-08-05 21:05 ` [PATCH v9 13/15] hugetlb: Set mapping folio order Matthew Wilcox (Oracle)
2026-08-13 11:06 ` Pedro Falcato
2026-09-18 13:47 ` David Hildenbrand (Arm)
2026-08-05 21:05 ` [PATCH v9 14/15] filemap: Add support for authoritative mappings Matthew Wilcox (Oracle)
2026-08-13 11:10 ` Pedro Falcato
2026-09-18 13:51 ` David Hildenbrand (Arm)
2026-09-21 21:19 ` Matthew Wilcox
2026-09-23 13:55 ` David Hildenbrand (Arm)
2026-08-05 21:05 ` [PATCH v9 15/15] hugetlb: replace hugetlbfs_read_iter() with generic_file_read_iter() Matthew Wilcox (Oracle)
2026-08-13 11:10 ` Pedro Falcato
2026-08-06 19:47 ` [PATCH v9 00/15] Use generic_file_read_iter() in hugetlbfs jane.chu
2026-08-13 8:37 ` Lorenzo Stoakes (ARM)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=7a23511f-365f-436f-b934-6a520510409c@kernel.org \
--to=david@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=christian@brauner.io \
--cc=gourry@gourry.net \
--cc=jack@suse.cz \
--cc=jane.chu@oracle.com \
--cc=jiaqiyan@google.com \
--cc=linmiaohe@huawei.com \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=muchun.song@linux.dev \
--cc=nao.horiguchi@gmail.com \
--cc=osalvador@suse.de \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox