All of lore.kernel.org
 help / color / mirror / Atom feed
From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: Kiryl Shutsemau <kirill@shutemov.name>
Cc: akpm@linux-foundation.org, david@kernel.org,
	nico.pache@linux.dev,  baolin.wang@linux.alibaba.com,
	baohua@kernel.org, dev.jain@arm.com, hughd@google.com,
	 lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com,
	rppt@kernel.org,  ryan.roberts@arm.com, shuah@kernel.org,
	surenb@google.com, usama.arif@linux.dev,  vbabka@kernel.org,
	ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com,
	 linux-mm@kvack.org, linux-kselftest@vger.kernel.org,
	linux-kernel@vger.kernel.org,  kas@kernel.org
Subject: Re: [PATCH v4 11/19] selftests/mm: add order-parameterized khugepaged collapse cases
Date: Tue, 18 Aug 2026 11:38:31 +0100	[thread overview]
Message-ID: <aoQ11Zo40gJw7qI8@lucifer> (raw)
In-Reply-To: <20260815015901.1236937-12-kirill@shutemov.name>

On Sat, Aug 15, 2026 at 02:58:53AM +0100, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> The mthp_khugepaged context runs the generic cases at a sub-PMD order,
> which answers how many folios of that order a range ends up with.  It
> cannot say which window they are in, so "the populated window collapsed
> and its neighbour did not" and "one window collapsed twice" look alike.
>
> Add four cases that check each aligned window on its own, with the
> folio-order helpers in vm_util:
>
> - collapse_order_single_window(): only the populated window collapses;
> - collapse_order_partial_window(): the default max_ptes_none lets a window
>   with one present PTE collapse;
> - collapse_order_max_ptes_none(): with max_ptes_none=0 a full window
>   collapses and one missing a page does not;
> - collapse_order_mixed_sources(): sources that are already large folios of
>   a smaller order collapse to the target.
>
> Each case faults its region before MADV_HUGEPAGE with only the target
> order enabled, so the sources are order 0 and the result can only come
> from khugepaged.  They wait for a full pass rather than for the result to
> appear: without a completed pass, "not collapsed" and "not scanned yet"
> are the same thing.
>
> Assisted-by: Claude-Code:claude-opus-5
> Tested-by: Muhammad Usama Anjum <usama.anjum@arm.com>
> Signed-off-by: Kiryl Shutsemau (Meta) <kas@kernel.org>

Same comments on comments as other patches - far too dense, read like a
discussion and not a terse description of something the code doesn't make
clear.

Please write comments yourself, LLMs are terrible at it (I mean I feel the
same goes for code also).

> ---
>  tools/testing/selftests/mm/khugepaged.c | 230 ++++++++++++++++++++++++
>  1 file changed, 230 insertions(+)
>
> diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selftests/mm/khugepaged.c
> index 0008862e7cbc..0489967d6ee0 100644
> --- a/tools/testing/selftests/mm/khugepaged.c
> +++ b/tools/testing/selftests/mm/khugepaged.c
> @@ -31,6 +31,8 @@ static unsigned long page_size;
>  static int hpage_pmd_nr;
>  static int anon_order;
>  static int collapse_order;
> +static int pagemap_fd = -1;
> +static int kpageflags_fd = -1;
>
>  #define PID_SMAPS "/proc/self/smaps"
>  #define TEST_FILE "collapse_test_file"
> @@ -1227,6 +1229,216 @@ static void madvise_retracted_page_tables(struct collapse_context *c,
>  	ksft_test_result_report(exit_status, "%s\n", __func__);
>  }
>
> +/* Smallest order khugepaged will consider for mTHP collapse. */
> +#define MIN_MTHP_ORDER 2
> +
> +/*
> + * Order-parameterized collapse cases for the mthp_khugepaged context.  What
> + * they add over the generic cases run under that context is per-window
> + * detection: which aligned window collapsed, and which of its neighbours did
> + * not.  check_huge() answers how many folios of the order the range holds,
> + * which cannot tell one window from another.
> + *
> + * The region is faulted before MADV_HUGEPAGE, and the target order is only
> + * enabled for madvise, so the sources are always order 0 and the collapse
> + * product can only have come from khugepaged.
> + */
> +static size_t mthp_window_size(void)
> +{
> +	return page_size << collapse_order;
> +}
> +
> +static void mthp_push_target_order(void)
> +{
> +	struct thp_settings settings = *thp_current_settings();
> +	int i;
> +
> +	/*
> +	 * The target order, for madvise only, and nothing else enabled: the
> +	 * cases fault their region before MADV_HUGEPAGE, so the sources are
> +	 * order 0 whatever -s asked the fault path for.  That matters for the
> +	 * cases built around a hole -- a large source folio would fill it in
> +	 * and the window would collapse after all.
> +	 * collapse_order_mixed_sources enables the source order it wants on
> +	 * top of this.
> +	 */
> +	settings.thp_enabled = THP_NEVER;
> +	for (i = 0; i < NR_ORDERS; i++)
> +		settings.hugepages[i].enabled = THP_NEVER;
> +	settings.hugepages[collapse_order].enabled = THP_MADVISE;
> +	thp_push_settings(&settings);
> +}
> +
> +static bool window_collapsed(void *p, size_t len)
> +{
> +	return is_range_backed_by_folio_orders(p, len, collapse_order,
> +					       pagemap_fd, kpageflags_fd);
> +}
> +
> +/* No aligned window in [p, p + len) is backed at the target order. */
> +static bool window_not_collapsed(void *p, size_t len)
> +{
> +	size_t window = mthp_window_size();
> +	char *addr = p;
> +
> +	for (; len >= window; addr += window, len -= window) {
> +		if (window_collapsed(addr, window))
> +			return false;
> +	}
> +	return true;
> +}
> +
> +static bool khugepaged_wait_full_pass(void)
> +{
> +	/* Wait up to 30 seconds for the pass to complete. */
> +	return khugepaged_full_pass(30);
> +}
> +
> +static void collapse_order_single_window(struct collapse_context *c,
> +					 struct mem_ops *ops)
> +{
> +	size_t window = mthp_window_size();
> +	void *p;
> +
> +	mthp_push_target_order();
> +
> +	p = ops->setup_area(1);
> +	ops->fault(p, window, 2 * window);
> +	if (!window_not_collapsed(p, hpage_pmd_size))
> +		ksft_exit_fail_msg("Unexpected large folio after fault\n");
> +
> +	madvise(p, hpage_pmd_size, MADV_HUGEPAGE);
> +	ksft_print_msg("Collapse one fully populated window...");
> +	if (!khugepaged_wait_full_pass())
> +		fail("Timeout");
> +	else if (window_collapsed(p + window, window) &&
> +		 window_not_collapsed(p, window) &&
> +		 window_not_collapsed(p + 2 * window,
> +				      hpage_pmd_size - 2 * window))
> +		success("OK");
> +	else
> +		fail("Fail");
> +
> +	validate_memory(p, window, 2 * window);
> +	ops->cleanup_area(p, hpage_pmd_size);
> +	thp_pop_settings();
> +	ksft_test_result_report(exit_status, "%s\n", __func__);
> +}
> +
> +static void collapse_order_partial_window(struct collapse_context *c,
> +					  struct mem_ops *ops)
> +{
> +	void *p;
> +
> +	mthp_push_target_order();
> +
> +	p = ops->setup_area(1);
> +	ops->fault(p, 0, page_size);
> +	if (!window_not_collapsed(p, hpage_pmd_size))
> +		ksft_exit_fail_msg("Unexpected large folio after fault\n");
> +
> +	madvise(p, hpage_pmd_size, MADV_HUGEPAGE);
> +	ksft_print_msg("Collapse window with single PTE entry present...");
> +	if (!khugepaged_wait_full_pass())
> +		fail("Timeout");
> +	else if (window_collapsed(p, mthp_window_size()))
> +		success("OK");
> +	else
> +		fail("Fail");
> +
> +	validate_memory(p, 0, page_size);
> +	ops->cleanup_area(p, hpage_pmd_size);
> +	thp_pop_settings();
> +	ksft_test_result_report(exit_status, "%s\n", __func__);
> +}
> +
> +static void collapse_order_max_ptes_none(struct collapse_context *c,
> +					 struct mem_ops *ops)
> +{
> +	struct thp_settings settings;
> +	size_t window = mthp_window_size();
> +	void *p;
> +
> +	mthp_push_target_order();
> +	settings = *thp_current_settings();
> +	settings.khugepaged.max_ptes_none = 0;
> +	thp_push_settings(&settings);
> +
> +	p = ops->setup_area(1);
> +	ops->fault(p, 0, 2 * window - page_size);
> +	if (!window_not_collapsed(p, hpage_pmd_size))
> +		ksft_exit_fail_msg("Unexpected large folio after fault\n");
> +
> +	madvise(p, hpage_pmd_size, MADV_HUGEPAGE);
> +	ksft_print_msg("Collapse full window, not the one missing a page...");
> +	if (!khugepaged_wait_full_pass())
> +		fail("Timeout");
> +	else if (window_collapsed(p, window) &&
> +		 window_not_collapsed(p + window, window))
> +		success("OK");
> +	else
> +		fail("Fail");
> +
> +	validate_memory(p, 0, 2 * window - page_size);
> +	ops->cleanup_area(p, hpage_pmd_size);
> +	thp_pop_settings();
> +	thp_pop_settings();
> +	ksft_test_result_report(exit_status, "%s\n", __func__);
> +}
> +
> +static void collapse_order_mixed_sources(struct collapse_context *c,
> +					 struct mem_ops *ops)
> +{
> +	struct thp_settings settings;
> +	void *p;
> +
> +	if (collapse_order <= MIN_MTHP_ORDER) {
> +		ksft_test_result_skip("%s: no source order below target\n",
> +				      __func__);
> +		return;
> +	}
> +
> +	mthp_push_target_order();
> +
> +	/* Fault the whole region as order-MIN_MTHP_ORDER folios. */
> +	settings = *thp_current_settings();
> +	settings.hugepages[MIN_MTHP_ORDER].enabled = THP_ALWAYS;
> +	thp_push_settings(&settings);
> +	p = ops->setup_area(1);
> +	ops->fault(p, 0, hpage_pmd_size);
> +	thp_pop_settings();
> +
> +	/*
> +	 * The order is enabled, but the allocator can still fall back under
> +	 * fragmentation.  That leaves nothing to collapse from, which is the
> +	 * machine's answer rather than a reason to end the run.
> +	 */
> +	if (!is_range_backed_by_folio_orders(p, hpage_pmd_size, MIN_MTHP_ORDER,
> +					     pagemap_fd, kpageflags_fd)) {
> +		ksft_print_msg("No order-%d sources to collapse...",
> +			       MIN_MTHP_ORDER);
> +		skip("Skip");
> +		ops->cleanup_area(p, hpage_pmd_size);
> +		thp_pop_settings();
> +		ksft_test_result_report(exit_status, "%s\n", __func__);
> +		return;
> +	}
> +
> +	madvise(p, hpage_pmd_size, MADV_HUGEPAGE);
> +	ksft_print_msg("Collapse region backed by smaller large folios...");
> +	if (!khugepaged_wait_full_pass())
> +		fail("Timeout");
> +	else if (window_collapsed(p, hpage_pmd_size))
> +		success("OK");
> +	else
> +		fail("Fail");
> +
> +	validate_memory(p, 0, hpage_pmd_size);
> +	ops->cleanup_area(p, hpage_pmd_size);
> +	thp_pop_settings();
> +	ksft_test_result_report(exit_status, "%s\n", __func__);
> +}
> +
>  static void usage(void)
>  {
>  	fprintf(stderr, "\nUsage: ./khugepaged [OPTIONS] <test type> [dir]\n\n");
> @@ -1395,6 +1607,20 @@ int main(int argc, char **argv)
>
>  	parse_test_type(argc, argv);
>
> +	if (mthp_khugepaged_context &&
> +	    !(thp_supported_orders() & (1UL << collapse_order)))
> +		ksft_exit_skip("Order %d is not a supported anon THP order\n",
> +			       collapse_order);
> +
> +	if (mthp_khugepaged_context) {
> +		pagemap_fd = open("/proc/self/pagemap", O_RDONLY);
> +		if (pagemap_fd < 0)
> +			ksft_exit_fail_perror("open(/proc/self/pagemap)");
> +		kpageflags_fd = open("/proc/kpageflags", O_RDONLY);
> +		if (kpageflags_fd < 0)
> +			ksft_exit_fail_perror("open(/proc/kpageflags)");
> +	}
> +
>  	setbuf(stdout, NULL);
>
>  	/*
> @@ -1450,6 +1676,10 @@ int main(int argc, char **argv)
>  	TEST(collapse_empty, madvise_context, anon_ops);
>
>  	TEST(collapse_single_mthp, mthp_khugepaged_context, anon_ops);
> +	TEST(collapse_order_single_window, mthp_khugepaged_context, anon_ops);
> +	TEST(collapse_order_partial_window, mthp_khugepaged_context, anon_ops);
> +	TEST(collapse_order_max_ptes_none, mthp_khugepaged_context, anon_ops);
> +	TEST(collapse_order_mixed_sources, mthp_khugepaged_context, anon_ops);
>
>  	TEST(collapse_single_pte_entry, khugepaged_context, anon_ops);
>  	TEST(collapse_single_pte_entry, khugepaged_context, read_only_file_ops);
> --
> 2.54.0
>

--
Cheers, Lorenzo

  reply	other threads:[~2026-08-18 10:38 UTC|newest]

Thread overview: 46+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-15  1:58 [PATCH v4 00/19] selftests/mm: improve khugepaged coverage Kiryl Shutsemau
2026-08-15  1:58 ` [PATCH v4 01/19] selftests/mm: raise the khugepaged test-case cap Kiryl Shutsemau
2026-08-18  9:11   ` Lorenzo Stoakes (ARM)
2026-08-15  1:58 ` [PATCH v4 02/19] selftests/mm: skip collapse_compound_extreme() where the PMD is too large Kiryl Shutsemau
2026-08-18  9:22   ` Lorenzo Stoakes (ARM)
2026-08-15  1:58 ` [PATCH v4 03/19] selftests/mm: scale khugepaged's collapse wait with the PMD size Kiryl Shutsemau
2026-08-18 10:04   ` Lorenzo Stoakes (ARM)
2026-08-23 19:17     ` Kiryl Shutsemau
2026-08-15  1:58 ` [PATCH v4 04/19] selftests/mm: skip khugepaged page cache cases without a PMD folio Kiryl Shutsemau
2026-08-18 10:07   ` Lorenzo Stoakes (ARM)
2026-08-23 19:44     ` Kiryl Shutsemau
2026-08-20  8:40   ` Mike Rapoport
2026-08-15  1:58 ` [PATCH v4 05/19] selftests/mm: make the swap cases' swapout reliable Kiryl Shutsemau
2026-08-18 10:11   ` Lorenzo Stoakes (ARM)
2026-08-20  8:40   ` Mike Rapoport
2026-08-23 19:49     ` Kiryl Shutsemau
2026-08-15  1:58 ` [PATCH v4 06/19] selftests/mm: stop khugepaged during the MADV_COLLAPSE cases Kiryl Shutsemau
2026-08-18 10:12   ` Lorenzo Stoakes (ARM)
2026-08-20  8:40   ` Mike Rapoport
2026-08-15  1:58 ` [PATCH v4 07/19] selftests/mm: move is_backed_by_folio() into vm_util Kiryl Shutsemau
2026-08-18 10:14   ` Lorenzo Stoakes (ARM)
2026-08-15  1:58 ` [PATCH v4 08/19] selftests/mm: add folio-order check for address ranges Kiryl Shutsemau
2026-08-18 10:25   ` Lorenzo Stoakes (ARM)
2026-08-20  8:40   ` Mike Rapoport
2026-08-15  1:58 ` [PATCH v4 09/19] selftests/mm: add folio-order detection self-check Kiryl Shutsemau
2026-08-18 10:30   ` Lorenzo Stoakes (ARM)
2026-08-20  8:40   ` Mike Rapoport
2026-08-15  1:58 ` [PATCH v4 10/19] selftests/mm: add khugepaged completion barrier helper Kiryl Shutsemau
2026-08-18 10:34   ` Lorenzo Stoakes (ARM)
2026-08-20  8:40   ` Mike Rapoport
2026-08-15  1:58 ` [PATCH v4 11/19] selftests/mm: add order-parameterized khugepaged collapse cases Kiryl Shutsemau
2026-08-18 10:38   ` Lorenzo Stoakes (ARM) [this message]
2026-08-20  8:40   ` Mike Rapoport
2026-08-15  1:58 ` [PATCH v4 12/19] selftests/mm: parameterize the mixed-source collapse case by source order Kiryl Shutsemau
2026-08-18 10:47   ` Lorenzo Stoakes (ARM)
2026-08-20  8:40   ` Mike Rapoport
2026-08-15  1:58 ` [PATCH v4 13/19] selftests/mm: cover a shared-source collapse write race Kiryl Shutsemau
2026-08-20  8:40   ` Mike Rapoport
2026-08-15  1:58 ` [PATCH v4 14/19] selftests/mm: run every supported collapse order by default Kiryl Shutsemau
2026-08-20  8:40   ` Mike Rapoport
2026-08-15  1:58 ` [PATCH v4 15/19] selftests/mm: check that one khugepaged pass collapses one window Kiryl Shutsemau
2026-08-15  1:58 ` [PATCH v4 16/19] selftests/mm: add khugepaged race harness Kiryl Shutsemau
2026-08-15  1:58 ` [PATCH v4 17/19] selftests/mm: race collapse of windows with holes Kiryl Shutsemau
2026-08-15  1:59 ` [PATCH v4 18/19] selftests/mm: add memory-pressure threads to the khugepaged race harness Kiryl Shutsemau
2026-08-15  1:59 ` [PATCH v4 19/19] selftests/mm: zap whole PTE tables in " Kiryl Shutsemau
2026-08-20  8:40 ` [PATCH v4 00/19] selftests/mm: improve khugepaged coverage Mike Rapoport

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aoQ11Zo40gJw7qI8@lucifer \
    --to=ljs@kernel.org \
    --cc=agordeev@linux.ibm.com \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=hughd@google.com \
    --cc=kas@kernel.org \
    --cc=kirill@shutemov.name \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=mhocko@suse.com \
    --cc=nico.pache@linux.dev \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=shuah@kernel.org \
    --cc=surenb@google.com \
    --cc=usama.anjum@arm.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.