From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: "David Hildenbrand (Arm)" <david@kernel.org>
Cc: Yunhui Cui <cuiyunhui@bytedance.com>,
akpm@linux-foundation.org, liam@infradead.org,
vbabka@kernel.org, jannh@google.com,
00moses.alexander00@gmail.com, linux-mm@kvack.org,
linux-kernel@vger.kernel.org, stable@vger.kernel.org
Subject: Re: [PATCH v2] mm/madvise: avoid skipping pages after splitting large folios
Date: Thu, 6 Aug 2026 10:13:45 +0100 [thread overview]
Message-ID: <anRPzJiraQUnsHa5@lucifer> (raw)
In-Reply-To: <699fc1b0-7af8-42ac-9882-c01225379586@kernel.org>
On Thu, Aug 06, 2026 at 10:41:06AM +0200, David Hildenbrand (Arm) wrote:
> On 8/6/26 07:55, Yunhui Cui wrote:
> > madvise_inject_error() advances through the requested range using the
> > size of the page returned by get_user_pages_fast(). Saving the size
> > before error injection is required for hugetlb pages because successful
> > soft offlining can dissolve the source huge page.
> >
> > That stride is incorrect for non-hugetlb large folios in system memory.
> > The memory failure handlers split such a folio and handle only the base
> > page for the supplied PFN. Advancing by the pre-split folio size then
> > skips the remaining pages in the requested range while madvise() still
> > reports success.
> >
> > Advance by PAGE_SIZE for non-hugetlb folios in system memory. Retain
> > folio_size() for hugetlb and ZONE_DEVICE folios, as compound Device DAX
> > folios are handled as a whole.
> >
> > Fixes: 19bfbe22f59a ("mm, hugetlb, soft_offline: save compound page order before page migration")
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Yunhui Cui <cuiyunhui@bytedance.com>
> > ---
> > mm/madvise.c | 13 +++++++++----
> > 1 file changed, 9 insertions(+), 4 deletions(-)
> >
> > diff --git a/mm/madvise.c b/mm/madvise.c
> > index 5a09cc24f04a0..e9d4c3bbc5290 100644
> > --- a/mm/madvise.c
> > +++ b/mm/madvise.c
> > @@ -1455,20 +1455,25 @@ static int madvise_inject_error(struct madvise_behavior *madv_behavior)
> >
> > for (; start < end; start += size) {
> > unsigned long pfn;
> > + struct folio *folio;
> > struct page *page;
> > int ret;
> >
> > ret = get_user_pages_fast(start, 1, 0, &page);
> > if (ret != 1)
> > return ret;
> > + folio = page_folio(page);
> > pfn = page_to_pfn(page);
> >
> > /*
> > - * When soft offlining hugepages, after migrating the page
> > - * we dissolve it, therefore in the second loop "page" will
> > - * no longer be a compound page.
> > + * Non-hugetlb large folios in system memory are split and only
> > + * the addressed base page is handled. Hugetlb folios may be
> > + * dissolved and ZONE_DEVICE folios may be handled as a whole,
> > + * so save their size before error injection.
> > */
> > - size = page_size(compound_head(page));
> > + size = PAGE_SIZE;
> > + if (folio_test_hugetlb(folio) || folio_is_zone_device(folio))
> > + size = folio_size(folio);
>
>
> We should never ever try deferring "how much has been mapped" from a single PTE.
Inferring? :)
>
> While this currently works for hugetlb, it's just an anti-pattern to throw
> hugetlb checks and similar around.
>
> So this is not the way to fix it.
All of this code is disgusting. Also what about an address range that is
partially inside a folio...? Then presumably the whole folio is discarded? Or
does it figure it out somehow and does the split for the rest of the range?
And what if end < the end of the hugetlb size?
Ugh god I hate all of this, it's so so bad. And hugetlb being the special
snowflake is the cherry on the s*** cake...
>
> --
> Cheers,
>
> David
--
Cheers, Lorenzo
next prev parent reply other threads:[~2026-08-06 9:14 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-06 5:55 [PATCH v2] mm/madvise: avoid skipping pages after splitting large folios Yunhui Cui
2026-08-06 6:29 ` Andrew Morton
2026-08-06 8:19 ` [External] " yunhui cui
2026-08-06 8:41 ` David Hildenbrand (Arm)
2026-08-06 9:13 ` Lorenzo Stoakes (ARM) [this message]
2026-08-06 9:11 ` Lorenzo Stoakes (ARM)
2026-08-06 11:35 ` David Hildenbrand (Arm)
2026-08-06 14:34 ` Lorenzo Stoakes (ARM)
2026-08-06 14:46 ` David Hildenbrand (Arm)
2026-08-06 15:40 ` Lorenzo Stoakes (ARM)
2026-08-11 2:31 ` [External] " yunhui cui
2026-08-11 7:52 ` Lorenzo Stoakes (ARM)
2026-08-11 14:49 ` David Hildenbrand (Arm)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anRPzJiraQUnsHa5@lucifer \
--to=ljs@kernel.org \
--cc=00moses.alexander00@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=cuiyunhui@bytedance.com \
--cc=david@kernel.org \
--cc=jannh@google.com \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=stable@vger.kernel.org \
--cc=vbabka@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.