From: "Vlastimil Babka (SUSE)" <vbabka@kernel.org>
To: Jiaqi Yan <jiaqiyan@google.com>,
linmiaohe@huawei.com, ljs@kernel.org, ziy@nvidia.com
Cc: osalvador@kernel.org, harry.yoo@oracle.com, willy@infradead.org,
osalvador@suse.de, jackmanb@google.com, hannes@cmpxchg.org,
nao.horiguchi@gmail.com, david@kernel.org,
william.roche@oracle.com, tony.luck@intel.com,
wangkefeng.wang@huawei.com, jane.chu@oracle.com,
akpm@linux-foundation.org, muchun.song@linux.dev,
liam@infradead.org, rientjes@google.com, duenwen@google.com,
jthoughton@google.com, linux-mm@kvack.org,
linux-kernel@vger.kernel.org, rppt@kernel.org, shuah@kernel.org,
surenb@google.com, mhocko@suse.com, boudewijn@delta-utec.com
Subject: Re: [PATCH v6 2/5] mm/page_alloc: only free healthy pages in high-order has_hwpoisoned folio
Date: Wed, 22 Jul 2026 11:40:29 +0200 [thread overview]
Message-ID: <a77cde29-20e5-415f-be47-7c1009db4f67@kernel.org> (raw)
In-Reply-To: <20260705180714.3708947-3-jiaqiyan@google.com>
On 7/5/26 20:07, Jiaqi Yan wrote:
> At the end of dissolve_free_hugetlb_folio(), a free HugeTLB folio
> becomes non-HugeTLB, and it is released to buddy allocator
> as a high-order folio, e.g. a folio that contains 262144 pages
> if the folio was a 1G HugeTLB hugepage.
>
> This is problematic if the HugeTLB hugepage contained HWPoison
> subpages. In that case, since buddy allocator does not check
> HWPoison for non-zero-order folio, the raw HWPoison page can
> be given out with its buddy page and be re-used by either
> kernel or userspace.
>
> Memory failure recovery (MFR) in kernel does attempt to take
> raw HWPoison page off buddy allocator after
> dissolve_free_hugetlb_folio(). However, there is always a time
> window between dissolve_free_hugetlb_folio() frees a HWPoison
> high-order folio to buddy allocator and MFR takes HWPoison
> raw page off buddy allocator.
>
> Another similar situation is when a transparent huge page (THP)
> runs into memory failure but splitting failed. Such THP will
> eventually be released to buddy allocator when owning userspace
> processes are gone, but with certain subpages having HWPoison.
>
> One obvious way to avoid both problems is to add page sanity
> checks in page allocate or free path. However, it is against
> the past efforts to reduce sanity check overhead [1,2,3].
>
> Introduce free_has_hwpoisoned() to only free the healthy pages
> and to exclude the HWPoison ones in the high-order folio.
> The idea is to iterate through the sub-pages of the folio to
> identify contiguous ranges of healthy pages.
>
> free_has_hwpoisoned() is added at the end of __free_pages_prepare()
> as a shortcut and only if PG_has_hwpoisoned indicates HWPoison page
> exists and after checks and preparations in __free_pages_prepare()
> all succeeded. It then use __free_prepared_contig_range() to
> decompose healthy range into the largest possible chunks of
> different orders, then freed via __free_frozen_pages().
>
> free_has_hwpoisoned() has linear time complexity wrt the number
> of pages in the folio. While the power-of-two decomposition
> ensures that the number of calls to the buddy allocator is
> logarithmic for each contiguous healthy range, the mandatory
> linear scan of pages to identify PageHWPoison() defines the
> overall time complexity. For a 1G hugepage having 8 HWPoison
> pages, free_has_hwpoisoned() takes around 1ms on average on
> a system having 56 Intel Skylake physical cores. This is
> 15x to the case of freeing no HWPoison page. The cost is far
> from triggering soft lockup, and fair for handling exceptional
> hardware memory errors.
>
> [1] https://lore.kernel.org/linux-mm/1460711275-1130-15-git-send-email-mgorman@techsingularity.net
> [2] https://lore.kernel.org/linux-mm/1460711275-1130-16-git-send-email-mgorman@techsingularity.net
> [3] https://lore.kernel.org/all/20230216095131.17336-1-vbabka@suse.cz
>
> Signed-off-by: Jiaqi Yan <jiaqiyan@google.com>
Reviewed-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
One thing below:
> @@ -6956,6 +7016,61 @@ void __free_contig_range(unsigned long pfn, unsigned long nr_pages)
> __free_contig_range_common(pfn, nr_pages, /* is_frozen= */ false);
> }
>
> +/*
> + * Given some contiguous pages that have certain number of HWPoison page(s),
> + * free only the healthy ones.
> + *
> + * Used at the end of __free_pages_prepare(). Even if having HWPoison pages,
> + * breaking down compound page and clearing metadata (e.g. page owner, alloc
> + * tag) can be done together during __free_pages_prepare(), which simplifies
> + * the splitting here: unlike __split_unmapped_folio(), there is no need to
> + * turn split pages into a compound page or to carry metadata.
> + *
> + * It scans every raw page of the compound page and causes nontrivial overhead.
> + * So only use this when the compound page contains HWPoison page(s).
> + *
> + * It also works when order == 0, regardless of PageHWPoison() or not.
> + *
> + * This implementation needs rework in memdesc world.
> + */
> +static void free_has_hwpoisoned(struct page *page, unsigned int order,
> + fpi_t fpi_flags)
> +{
> + unsigned long curr = page_to_pfn(page);
> + unsigned long end_pfn = curr + (1 << order);
> + unsigned long next;
> + unsigned long total_freed = 0;
> + unsigned long total_hwp = 0;
> +
> + while (curr < end_pfn) {
> + next = curr;
> +
> + while (next < end_pfn && !PageHWPoison(pfn_to_page(next)))
> + ++next;
> +
> + if (next != end_pfn) {
> + /*
> + * Avoid accounting error when the page is freed
> + * by unpoison_memory().
> + */
> + clear_page_tag_ref(pfn_to_page(next));
> + ++total_hwp;
> + }
> +
> + __free_prepared_contig_range(pfn_to_page(curr), next - curr,
> + fpi_flags);
> + total_freed += next - curr;
> +
> + if (next == end_pfn)
> + break;
> +
> + curr = next + 1;
> + }
> +
> + pr_info("Freed %#lx pages, excluded %#lx HWPoison pages\n",
> + total_freed, total_hwp);
Should we really print this? Maybe just pr_debug() or not at all?
> +}
> +
> #ifdef CONFIG_CONTIG_ALLOC
> /* Usage: See admin-guide/dynamic-debug-howto.rst */
> static void alloc_contig_dump_pages(struct list_head *page_list)
next prev parent reply other threads:[~2026-07-22 9:40 UTC|newest]
Thread overview: 38+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-05 18:07 [PATCH v6 0/5] Only free healthy pages in high-order has_hwpoisoned folio Jiaqi Yan
2026-07-05 18:07 ` [PATCH v6 1/5] mm/page_alloc: introduce __free_prepared_contig_range() with fpi_t Jiaqi Yan
2026-07-17 7:17 ` Miaohe Lin
2026-07-22 9:38 ` Vlastimil Babka (SUSE)
2026-07-25 3:54 ` Matthew Wilcox
2026-08-17 0:28 ` Jiaqi Yan
2026-07-05 18:07 ` [PATCH v6 2/5] mm/page_alloc: only free healthy pages in high-order has_hwpoisoned folio Jiaqi Yan
2026-07-17 7:19 ` Miaohe Lin
2026-07-22 9:40 ` Vlastimil Babka (SUSE) [this message]
2026-07-05 18:07 ` [PATCH v6 3/5] mm/memory-failure: set has_hwpoisoned flags on dissolved HugeTLB folio Jiaqi Yan
2026-07-25 3:05 ` Jiaqi Yan
2026-07-05 18:07 ` [PATCH v6 4/5] mm/memory-failure: skip take_page_off_buddy after dissolving HWPoison HugeTLB page Jiaqi Yan
2026-07-17 7:37 ` Miaohe Lin
2026-08-17 0:29 ` Jiaqi Yan
2026-08-17 7:23 ` Miaohe Lin
2026-08-18 3:30 ` Jiaqi Yan
2026-08-18 9:06 ` Miaohe Lin
2026-07-05 18:07 ` [PATCH v6 5/5] selftests/mm: add hard memory failure anonymous HugeTLB test Jiaqi Yan
2026-07-17 7:37 ` Miaohe Lin
2026-07-25 3:05 ` Jiaqi Yan
2026-07-05 18:50 ` [PATCH v6 0/5] Only free healthy pages in high-order has_hwpoisoned folio Andrew Morton
2026-07-06 9:03 ` David Hildenbrand (Arm)
2026-07-25 3:06 ` Jiaqi Yan
2026-07-17 10:18 ` David Hildenbrand (Arm)
2026-07-17 13:06 ` William Roche
2026-07-22 8:27 ` Vlastimil Babka (SUSE)
2026-07-22 9:03 ` Vlastimil Babka (SUSE)
2026-07-25 6:00 ` Jiaqi Yan
2026-07-25 6:52 ` Jiaqi Yan
2026-07-27 14:20 ` David Hildenbrand (Arm)
2026-07-27 18:20 ` Matthew Wilcox
2026-07-28 17:18 ` Vlastimil Babka (SUSE)
2026-08-04 19:51 ` David Hildenbrand (Arm)
2026-08-05 8:58 ` Vlastimil Babka (SUSE)
2026-08-05 9:09 ` David Hildenbrand (Arm)
2026-08-28 23:36 ` Linux MM Alignment Session on hwpoison (was: Re: [PATCH v6 0/5] Only free healthy pages in high-order has_hwpoisoned folio) David Rientjes
2026-08-31 19:26 ` David Rientjes
2026-09-07 13:46 ` Linux MM Alignment Session on hwpoison David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=a77cde29-20e5-415f-be47-7c1009db4f67@kernel.org \
--to=vbabka@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=boudewijn@delta-utec.com \
--cc=david@kernel.org \
--cc=duenwen@google.com \
--cc=hannes@cmpxchg.org \
--cc=harry.yoo@oracle.com \
--cc=jackmanb@google.com \
--cc=jane.chu@oracle.com \
--cc=jiaqiyan@google.com \
--cc=jthoughton@google.com \
--cc=liam@infradead.org \
--cc=linmiaohe@huawei.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=muchun.song@linux.dev \
--cc=nao.horiguchi@gmail.com \
--cc=osalvador@kernel.org \
--cc=osalvador@suse.de \
--cc=rientjes@google.com \
--cc=rppt@kernel.org \
--cc=shuah@kernel.org \
--cc=surenb@google.com \
--cc=tony.luck@intel.com \
--cc=wangkefeng.wang@huawei.com \
--cc=william.roche@oracle.com \
--cc=willy@infradead.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.