From: Baolin Wang <baolin.wang@linux.alibaba.com>
To: Yeoreum Yun <yeoreum.yun@arm.com>, Zi Yan <ziy@nvidia.com>,
"Liam R. Howlett" <liam@infradead.org>,
Nico Pache <nico.pache@linux.dev>,
Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
Barry Song <baohua@kernel.org>, Lance Yang <lance.yang@linux.dev>,
Usama Arif <usama.arif@linux.dev>,
Kiryl Shutsemau <kas@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>,
Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>,
Suren Baghdasaryan <surenb@google.com>,
Michal Hocko <mhocko@suse.com>,
linux-mm@kvack.org, linux-kselftest@vger.kernel.org,
linux-kernel@vger.kernel.org
Cc: Andrew Morton <akpm@linux-foundation.org>,
David Hildenbrand <david@kernel.org>,
Shuah Khan <shuah@kernel.org>
Subject: Re: [PATCH 2/2] kselftest: mm: fix intermittent failure khugepaged test
Date: Wed, 16 Sep 2026 10:56:38 +0800 [thread overview]
Message-ID: <24690e82-3aab-4f2d-95a7-3bba332ca5bc@linux.alibaba.com> (raw)
In-Reply-To: <20260915-fix_khugepagd_fail-v1-2-bb6f04c8759f@arm.com>
On 9/15/26 5:21 PM, Yeoreum Yun wrote:
> There are intermittent failures in collapse_max_ptes_swap() and
> collapse_max_ptes_shared() when using the khugepaged_context:
>
> // while running ./khugepaged -s 2
>
> # Run test: collapse_max_ptes_shared (khugepaged:anon)
> # Allocate huge page... OK
> # Share huge page over fork()... OK
> # Trigger CoW on page 1023 of 2048... OK
> # Maybe collapse with max_ptes_shared exceeded.... OK
> # Trigger CoW on page 1024 of 2048... Fail
> Bail out! Unexpected huge page
> # Planned tests != run tests (26 != 23)
> # Totals: pass:23 fail:0 xfail:0 xpass:0 skip:0 error:0
>
> # Run test: collapse_max_ptes_swap (khugepaged:anon)
> # Swapout 257 of 2048 pages... OK
> # Maybe collapse with max_ptes_swap exceeded.... OK
> # Swapout 256 of 2048 pages... OK
> Bail out! Unexpected huge page
> # Planned tests != run tests (26 != 17)
> # Totals: pass:17 fail:0 xfail:0 xpass:0 skip:0 error:0
>
> This happens because khugepaged may collapse the pages before wait_for_scan()
> is called, causing a sanity check that expects uncollapsed pages to fail.
>
> For example, in collapse_max_ptes_swap(), after faulting the pages back in
> and paging out up to max_ptes_swap pages, khugepaged may collapse them again
> before c->collapse() is called.
>
> To prevent this, change the khugepaged setting from ALWAYS to MADVICE for
> the affected tests, and mark the VMA with MADV_NOHUGEPAGE after it has been
> collapsed by wait_for_scan(). This prevents khugepaged from collapsing it
> again before c->collapse() is called.
>
> This failure was observed on NVIDIA Spark with 16KB page.
>
> Signed-off-by: Yeoreum Yun <yeoreum.yun@arm.com>
> ---
> tools/testing/selftests/mm/khugepaged.c | 15 +++++++++++++++
> 1 file changed, 15 insertions(+)
>
> diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selftests/mm/khugepaged.c
> index c32244b565658..83e9386bbc842 100644
> --- a/tools/testing/selftests/mm/khugepaged.c
> +++ b/tools/testing/selftests/mm/khugepaged.c
> @@ -578,6 +578,8 @@ static bool wait_for_scan(const char *msg, char *p, size_t len,
> usleep(TICK);
> }
>
> + madvise(p, len, MADV_NOHUGEPAGE);
This looks incorrect to me and would reintroduce the previous problem.
Please see commit 7962e05a835f ("selftests: khugepaged: fix the shmem
collapse failure").
> +
> return timeout == -1;
> }
>
> @@ -839,6 +841,7 @@ static void collapse_swapin_single_pte(struct collapse_context *c, struct mem_op
>
> static void collapse_max_ptes_swap(struct collapse_context *c, struct mem_ops *ops)
> {
> + struct thp_settings settings = *thp_current_settings();
> int max_ptes_swap = thp_read_num("khugepaged/max_ptes_swap");
> void *p;
>
> @@ -860,6 +863,9 @@ static void collapse_max_ptes_swap(struct collapse_context *c, struct mem_ops *o
> validate_memory(p, 0, hpage_pmd_size);
>
> if (c->enforce_pte_scan_limits) {
> + settings.hugepages[collapse_order].enabled = THP_MADVISE;
> + thp_push_settings(&settings);
I'm not sure why the collapse_order setting needs to be changed here. In
your test case, you did not use the '-c' parameter to specify the
collapse order.
> +
> ops->fault(p, 0, hpage_pmd_size);
> ksft_print_msg("Swapout %d of %d pages...", max_ptes_swap,
> hpage_pmd_nr);
> @@ -869,12 +875,15 @@ static void collapse_max_ptes_swap(struct collapse_context *c, struct mem_ops *o
> success("OK");
> } else {
> fail("Fail");
> + thp_pop_settings();
> goto out;
> }
>
> c->collapse("Collapse with max_ptes_swap pages swapped out", p,
> 1, ops, true);
> validate_memory(p, 0, hpage_pmd_size);
> +
> + thp_pop_settings();
> }
> out:
> ops->cleanup_area(p, hpage_pmd_size);
> @@ -1075,6 +1084,7 @@ static void collapse_fork_compound(struct collapse_context *c, struct mem_ops *o
>
> static void collapse_max_ptes_shared(struct collapse_context *c, struct mem_ops *ops)
> {
> + struct thp_settings settings = *thp_current_settings();
> int max_ptes_shared = thp_read_num("khugepaged/max_ptes_shared");
> int wstatus;
> void *p;
> @@ -1100,6 +1110,9 @@ static void collapse_max_ptes_shared(struct collapse_context *c, struct mem_ops
> 1, ops, !c->enforce_pte_scan_limits);
>
> if (c->enforce_pte_scan_limits) {
> + settings.hugepages[collapse_order].enabled = THP_MADVISE;
> + thp_push_settings(&settings);
Ditto.
> +
> ksft_print_msg("Trigger CoW on page %d of %d...",
> hpage_pmd_nr - max_ptes_shared, hpage_pmd_nr);
> ops->fault(p, 0, (hpage_pmd_nr - max_ptes_shared) *
> @@ -1111,6 +1124,8 @@ static void collapse_max_ptes_shared(struct collapse_context *c, struct mem_ops
>
> c->collapse("Collapse with max_ptes_shared PTEs shared",
> p, 1, ops, true);
> +
> + thp_pop_settings();
> }
>
> validate_memory(p, 0, hpage_pmd_size);
>
next prev parent reply other threads:[~2026-09-16 2:56 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-15 9:21 [PATCH 0/2] kselftest: mm: fix intermittent failure khugepaged test Yeoreum Yun
2026-09-15 9:21 ` [PATCH 1/2] kselftest: mm: return fail when child test result is fail in khugepaged Yeoreum Yun
2026-09-15 9:21 ` [PATCH 2/2] kselftest: mm: fix intermittent failure khugepaged test Yeoreum Yun
2026-09-16 2:56 ` Baolin Wang [this message]
2026-09-16 3:47 ` Yeoreum Yun
2026-09-16 1:15 ` [PATCH 0/2] " Andrew Morton
2026-09-16 2:21 ` Yeoreum Yun
2026-09-16 6:41 ` David Hildenbrand (Arm)
2026-09-16 7:06 ` Yeoreum Yun
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=24690e82-3aab-4f2d-95a7-3bba332ca5bc@linux.alibaba.com \
--to=baolin.wang@linux.alibaba.com \
--cc=akpm@linux-foundation.org \
--cc=baohua@kernel.org \
--cc=david@kernel.org \
--cc=dev.jain@arm.com \
--cc=kas@kernel.org \
--cc=lance.yang@linux.dev \
--cc=liam@infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=mhocko@suse.com \
--cc=nico.pache@linux.dev \
--cc=rppt@kernel.org \
--cc=ryan.roberts@arm.com \
--cc=shuah@kernel.org \
--cc=surenb@google.com \
--cc=usama.arif@linux.dev \
--cc=vbabka@kernel.org \
--cc=yeoreum.yun@arm.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.