* [PATCH v3 0/2] mm/page_isolation: fix UBSAN shift-out-of-bounds in isolate_single_pageblock
@ 2026-08-25 12:05 Qi Xi
2026-08-25 12:05 ` [PATCH v3 1/2] mm/page_isolation: fix UBSAN shift-out-of-bounds warning Qi Xi
2026-08-25 12:05 ` [PATCH v3 2/2] mm/page_isolation: guard compound_order() against racing Qi Xi
0 siblings, 2 replies; 5+ messages in thread
From: Qi Xi @ 2026-08-25 12:05 UTC (permalink / raw)
To: Andrew Morton, Vlastimil Babka
Cc: Suren Baghdasaryan, Michal Hocko, Brendan Jackman,
Johannes Weiner, Zi Yan, linux-mm, linux-kernel, sunnanyong,
wangkefeng.wang, xiqi2
Patch 1 fixes a UBSAN shift-out-of-bounds warning in the PageBuddy branch
of isolate_single_pageblock() triggered by concurrent buddy allocation.
Patch 2 addresses the same class of issue in the PageCompound branch, where
racy compound_order() and compound_head() reads could lead to out-of-range
shifts or incorrect page skipping.
Changes since v2:
- Patch 1: add a comment clarifying pageblock_isolate_and_move_free_pages()
splits cross-boundary PageBuddy. No code change.
Changes since v1:
- Patch 1: drop VM_WARN_ON_ONCE, bail out with -EBUSY instead
- Patch 2: new patch addressing PageCompound branch race
Qi Xi (2):
mm/page_isolation: fix UBSAN shift-out-of-bounds warning
mm/page_isolation: guard compound_order() against racing
mm/page_isolation.c | 40 ++++++++++++++++++++++++++++++++--------
1 file changed, 32 insertions(+), 8 deletions(-)
--
2.33.0
^ permalink raw reply [flat|nested] 5+ messages in thread* [PATCH v3 1/2] mm/page_isolation: fix UBSAN shift-out-of-bounds warning 2026-08-25 12:05 [PATCH v3 0/2] mm/page_isolation: fix UBSAN shift-out-of-bounds in isolate_single_pageblock Qi Xi @ 2026-08-25 12:05 ` Qi Xi 2026-08-25 12:05 ` [PATCH v3 2/2] mm/page_isolation: guard compound_order() against racing Qi Xi 1 sibling, 0 replies; 5+ messages in thread From: Qi Xi @ 2026-08-25 12:05 UTC (permalink / raw) To: Andrew Morton, Vlastimil Babka Cc: Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Johannes Weiner, Zi Yan, linux-mm, linux-kernel, sunnanyong, wangkefeng.wang, xiqi2 A contig-range allocation racing with buddy allocation on the adjacent pageblock can trigger: UBSAN: shift-out-of-bounds in mm/page_isolation.c:393:15 shift exponent -749042176 is negative Call trace: isolate_single_pageblock start_isolate_page_range alloc_contig_frozen_range_noprof alloc_contig_range_noprof isolate_single_pageblock() first calls set_migratetype_isolate() with zone->lock held, which marks the pageblock MIGRATE_ISOLATE and moves any free page straddling the boundary out of the way. Once the lock is dropped, it scans the MAX_ORDER_NR_PAGES-aligned window [start_pfn, boundary_pfn) locklessly, only to skip the free pages already handled above and to detect in-use pages straddling the boundary. Since this scan only reads page state to decide how far to skip and returns -EBUSY on a straddling in-use page, it does not take the lock. The window also covers the adjacent pageblock, whose free pages stay on the normal movable/CMA freelist and can be allocated concurrently. So after the scan observes PageBuddy(page), another CPU can allocate the page, leaving a stale value in page->private that makes "1 << order" shift out of range. Use buddy_order_unsafe() with READ_ONCE to read the order, and validate it is within MAX_PAGE_ORDER before shifting to prevent UBSAN warnings. Since pageblock_isolate_and_move_free_pages() already handles free pages straddling boundary_pfn under zone->lock, bail out with -EBUSY instead of VM_WARN_ON_ONCE() when a PageBuddy page appears to cross the boundary during the lockless scan. Fixes: b2c9e2fbba32 ("mm: make alloc_contig_range work at pageblock granularity") Cc: stable@vger.kernel.org Reviewed-by: Zi Yan <ziy@nvidia.com> Signed-off-by: Qi Xi <xiqi2@huawei.com> --- mm/page_isolation.c | 18 ++++++++++++------ 1 file changed, 12 insertions(+), 6 deletions(-) diff --git a/mm/page_isolation.c b/mm/page_isolation.c index 32ce8a7d9df3..61efc03500ef 100644 --- a/mm/page_isolation.c +++ b/mm/page_isolation.c @@ -387,13 +387,19 @@ static int isolate_single_pageblock(unsigned long boundary_pfn, } if (PageBuddy(page)) { - int order = buddy_order(page); + unsigned int order = buddy_order_unsafe(page); - /* pageblock_isolate_and_move_free_pages() handled this */ - VM_WARN_ON_ONCE(pfn + (1 << order) > boundary_pfn); - - pfn += 1UL << order; - continue; + /* buddy_order_unsafe() is racy. Validate the order before shifting. */ + if (order <= MAX_PAGE_ORDER && + /* + * pageblock_isolate_and_move_free_pages() splits + * cross-boundary PageBuddy, verify it. + */ + pfn + (1UL << order) <= boundary_pfn) { + pfn += 1UL << order; + continue; + } + goto failed; } /* -- 2.33.0 ^ permalink raw reply related [flat|nested] 5+ messages in thread
* [PATCH v3 2/2] mm/page_isolation: guard compound_order() against racing 2026-08-25 12:05 [PATCH v3 0/2] mm/page_isolation: fix UBSAN shift-out-of-bounds in isolate_single_pageblock Qi Xi 2026-08-25 12:05 ` [PATCH v3 1/2] mm/page_isolation: fix UBSAN shift-out-of-bounds warning Qi Xi @ 2026-08-25 12:05 ` Qi Xi 2026-08-28 3:47 ` Andrew Morton 1 sibling, 1 reply; 5+ messages in thread From: Qi Xi @ 2026-08-25 12:05 UTC (permalink / raw) To: Andrew Morton, Vlastimil Babka Cc: Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Johannes Weiner, Zi Yan, linux-mm, linux-kernel, sunnanyong, wangkefeng.wang, xiqi2 The PageCompound branch reads compound_head() without holding a reference. A racing split or free can cause compound_head() to return a stale pointer, and compound_nr() reads the order from that stale head, leading to out-of-range shifts and making the skip distance meaningless. Read the order explicitly with compound_order() and validate it is within MAX_FOLIO_ORDER before shifting. Also verify the derived head_pfn against the legitimate pfn: the head must not be past pfn, must be aligned to nr_pages, and pfn must fall within the compound page. Bail out with -EBUSY if any check fails. Fixes: b2c9e2fbba32 ("mm: make alloc_contig_range work at pageblock granularity") Cc: stable@vger.kernel.org Suggested-by: Zi Yan <ziy@nvidia.com> Reviewed-by: Zi Yan <ziy@nvidia.com> Signed-off-by: Qi Xi <xiqi2@huawei.com> --- mm/page_isolation.c | 22 ++++++++++++++++++++-- 1 file changed, 20 insertions(+), 2 deletions(-) diff --git a/mm/page_isolation.c b/mm/page_isolation.c index 61efc03500ef..eca6fb78f73a 100644 --- a/mm/page_isolation.c +++ b/mm/page_isolation.c @@ -418,10 +418,28 @@ static int isolate_single_pageblock(unsigned long boundary_pfn, if (PageCompound(page)) { struct page *head = compound_head(page); unsigned long head_pfn = page_to_pfn(head); - unsigned long nr_pages = compound_nr(head); + unsigned int order = compound_order(head); + unsigned long nr_pages; + + /* compound_order() is racy. Cap it at MAX_FOLIO_ORDER. */ + if (order > MAX_FOLIO_ORDER) + goto failed; + + nr_pages = 1UL << order; + + /* + * compound_head() is also racy, so the derived head_pfn + * needs additional checks to make sure it is valid. + * Otherwise, just fail the check. pfn comes from + * __first_valid_page() as a legitimate PFN, so use it to + * check head_pfn. + */ + if (head_pfn > pfn || !IS_ALIGNED(head_pfn, nr_pages) || + pfn - head_pfn >= nr_pages) + goto failed; if (head_pfn + nr_pages <= boundary_pfn || - PageHuge(page)) { + PageHuge(head)) { pfn = head_pfn + nr_pages; continue; } -- 2.33.0 ^ permalink raw reply related [flat|nested] 5+ messages in thread
* Re: [PATCH v3 2/2] mm/page_isolation: guard compound_order() against racing 2026-08-25 12:05 ` [PATCH v3 2/2] mm/page_isolation: guard compound_order() against racing Qi Xi @ 2026-08-28 3:47 ` Andrew Morton 2026-08-29 7:16 ` Qi Xi 0 siblings, 1 reply; 5+ messages in thread From: Andrew Morton @ 2026-08-28 3:47 UTC (permalink / raw) To: Qi Xi Cc: Vlastimil Babka, Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Johannes Weiner, Zi Yan, linux-mm, linux-kernel, sunnanyong, wangkefeng.wang On Tue, 25 Aug 2026 20:05:49 +0800 Qi Xi <xiqi2@huawei.com> wrote: > The PageCompound branch reads compound_head() without holding a reference. > A racing split or free can cause compound_head() to return a stale pointer, > and compound_nr() reads the order from that stale head, leading to > out-of-range shifts and making the skip distance meaningless. > > Read the order explicitly with compound_order() and validate it is within > MAX_FOLIO_ORDER before shifting. Also verify the derived head_pfn against > the legitimate pfn: the head must not be past pfn, must be aligned to > nr_pages, and pfn must fall within the compound page. Bail out with > -EBUSY if any check fails. > > ... > > --- a/mm/page_isolation.c > +++ b/mm/page_isolation.c > @@ -418,10 +418,28 @@ static int isolate_single_pageblock(unsigned long boundary_pfn, > if (PageCompound(page)) { > struct page *head = compound_head(page); > unsigned long head_pfn = page_to_pfn(head); > - unsigned long nr_pages = compound_nr(head); > + unsigned int order = compound_order(head); > + unsigned long nr_pages; > + > + /* compound_order() is racy. Cap it at MAX_FOLIO_ORDER. */ > + if (order > MAX_FOLIO_ORDER) > + goto failed; > + > + nr_pages = 1UL << order; > + > + /* > + * compound_head() is also racy, so the derived head_pfn > + * needs additional checks to make sure it is valid. > + * Otherwise, just fail the check. pfn comes from > + * __first_valid_page() as a legitimate PFN, so use it to > + * check head_pfn. > + */ > + if (head_pfn > pfn || !IS_ALIGNED(head_pfn, nr_pages) || > + pfn - head_pfn >= nr_pages) > + goto failed; > > if (head_pfn + nr_pages <= boundary_pfn || > - PageHuge(page)) { > + PageHuge(head)) { Sashiko suggests that this PageHuge() test suffers the same issue? https://sashiko.dev/#/patchset/20260825120549.966271-1-xiqi2@huawei.com > pfn = head_pfn + nr_pages; > continue; > } > -- > 2.33.0 > ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v3 2/2] mm/page_isolation: guard compound_order() against racing 2026-08-28 3:47 ` Andrew Morton @ 2026-08-29 7:16 ` Qi Xi 0 siblings, 0 replies; 5+ messages in thread From: Qi Xi @ 2026-08-29 7:16 UTC (permalink / raw) To: Andrew Morton Cc: Vlastimil Babka, Suren Baghdasaryan, Michal Hocko, Brendan Jackman, Johannes Weiner, Zi Yan, linux-mm, linux-kernel, sunnanyong, wangkefeng.wang Thanks for raising this. I checked the path: PageHuge(head) resolves to folio_test_hugetlb(), which reads page_type via data_race(). I don't think it's the same issue this series targets: - It never touches page->flags or PF_SECOND — folio_test_hugetlb() reads page_type on the head page via FOLIO_TYPE_OPS, a separate "page type" mechanism, so neither folio_flags() nor its VM_BUG_ON_PGFLAGS assertions are on its path. - The only shift is a fixed >> 24 on a u32, not a variable 1 << order, so it can't produce the shift-out-of-bounds UBSAN this series fixes. So no extra stabilization seems needed for PageHuge() itself. Happy to add it if you see a case I missed. Qi On 28/08/2026 11:47, Andrew Morton wrote: > On Tue, 25 Aug 2026 20:05:49 +0800 Qi Xi <xiqi2@huawei.com> wrote: > >> The PageCompound branch reads compound_head() without holding a reference. >> A racing split or free can cause compound_head() to return a stale pointer, >> and compound_nr() reads the order from that stale head, leading to >> out-of-range shifts and making the skip distance meaningless. >> >> Read the order explicitly with compound_order() and validate it is within >> MAX_FOLIO_ORDER before shifting. Also verify the derived head_pfn against >> the legitimate pfn: the head must not be past pfn, must be aligned to >> nr_pages, and pfn must fall within the compound page. Bail out with >> -EBUSY if any check fails. >> >> ... >> >> --- a/mm/page_isolation.c >> +++ b/mm/page_isolation.c >> @@ -418,10 +418,28 @@ static int isolate_single_pageblock(unsigned long boundary_pfn, >> if (PageCompound(page)) { >> struct page *head = compound_head(page); >> unsigned long head_pfn = page_to_pfn(head); >> - unsigned long nr_pages = compound_nr(head); >> + unsigned int order = compound_order(head); >> + unsigned long nr_pages; >> + >> + /* compound_order() is racy. Cap it at MAX_FOLIO_ORDER. */ >> + if (order > MAX_FOLIO_ORDER) >> + goto failed; >> + >> + nr_pages = 1UL << order; >> + >> + /* >> + * compound_head() is also racy, so the derived head_pfn >> + * needs additional checks to make sure it is valid. >> + * Otherwise, just fail the check. pfn comes from >> + * __first_valid_page() as a legitimate PFN, so use it to >> + * check head_pfn. >> + */ >> + if (head_pfn > pfn || !IS_ALIGNED(head_pfn, nr_pages) || >> + pfn - head_pfn >= nr_pages) >> + goto failed; >> >> if (head_pfn + nr_pages <= boundary_pfn || >> - PageHuge(page)) { >> + PageHuge(head)) { > Sashiko suggests that this PageHuge() test suffers the same issue? > > https://sashiko.dev/#/patchset/20260825120549.966271-1-xiqi2@huawei.com > >> pfn = head_pfn + nr_pages; >> continue; >> } >> -- >> 2.33.0 >> ^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-08-29 7:17 UTC | newest] Thread overview: 5+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-08-25 12:05 [PATCH v3 0/2] mm/page_isolation: fix UBSAN shift-out-of-bounds in isolate_single_pageblock Qi Xi 2026-08-25 12:05 ` [PATCH v3 1/2] mm/page_isolation: fix UBSAN shift-out-of-bounds warning Qi Xi 2026-08-25 12:05 ` [PATCH v3 2/2] mm/page_isolation: guard compound_order() against racing Qi Xi 2026-08-28 3:47 ` Andrew Morton 2026-08-29 7:16 ` Qi Xi
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox