Linux Documentation
 help / color / mirror / Atom feed
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: "Nico Pache (Red Hat)" <nico.pache@linux.dev>,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	linux-doc@vger.kernel.org
Cc: Andrew Morton <akpm@linux-foundation.org>,
	Lorenzo Stoakes <ljs@kernel.org>, Zi Yan <ziy@nvidia.com>,
	Baolin Wang <baolin.wang@linux.alibaba.com>,
	"Liam R. Howlett" <liam@infradead.org>,
	Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
	Barry Song <baohua@kernel.org>, Lance Yang <lance.yang@linux.dev>,
	Usama Arif <usama.arif@linux.dev>,
	Vlastimil Babka <vbabka@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	Michal Hocko <mhocko@suse.com>, Jonathan Corbet <corbet@lwn.net>,
	Shuah Khan <skhan@linuxfoundation.org>
Subject: Re: [PATCH v4 5/7] mm/khugepaged: Refactor the PTE state checks into a helper
Date: Wed, 12 Aug 2026 10:40:08 +0200	[thread overview]
Message-ID: <f1371b4d-6e98-4699-8c1b-7612f6d70f02@kernel.org> (raw)
In-Reply-To: <20260811-khugepaged_pte_refactor-v4-5-ddac39d61c4a@linux.dev>

On 8/11/26 14:48, Nico Pache (Red Hat) wrote:
> For anonymous collapse, the collapse_scan_pmd() and
> __collapse_huge_page_isolate() functions share a large portion of their
> logic. These functions both check the state of the PTEs and verify the
> following:
> 	- max_pte_* values are not exceeded
> 	- uffd is not active
> 	- lazyfree properties
> 	- non-anonymous
> 
> Merge these checks into a helper collapse_check_pte() to reduce code
> duplication. We also add a helper struct for this function called
> pte_check_context which allows us to pass the required parameters in a
> clean and elegant manner.
> 
> A helper function is also introduced pte_check_fail() to provide a clean
> interface to set the pte_check_context failure results and return
> PTE_CHECK_FAIL state. This helps reduce code duplications across the new
> collapse_check_pte function.
> 
> Two slight modifications are done to the original functionality. We now
> warn (instead of crash) if the anon test fails, and we leverage the
> vm_normal_folio function instead of page->folio, this should be
> functionally equivalent.
> 
> No other functional changes intended.
> 
> This patch is heavily based off work done by Lance Yang, but modified to
> deal with conflicts and feedback received during the review cycle [1].
> 

TL;DR, I think this patch here needs some more work, and we should not fast
track it at this point.

@Andrew, can we delay this patch here for this merge window? Removing it from
mm-unstable shouldn't conflict with any other patch in this series.


> [1] https://lore.kernel.org/linux-mm/20251008043748.45554-1-lance.yang@linux.dev/
> Suggested-by: David Hildenbrand <david@kernel.org>
> Signed-off-by: Nico Pache (Red Hat) <nico.pache@linux.dev>
> ---
>  mm/khugepaged.c | 298 +++++++++++++++++++++++++++++---------------------------
>  1 file changed, 157 insertions(+), 141 deletions(-)
> 
> diff --git a/mm/khugepaged.c b/mm/khugepaged.c
> index 90d6e595d282..b7372aba4417 100644
> --- a/mm/khugepaged.c
> +++ b/mm/khugepaged.c
> @@ -65,6 +65,12 @@ enum scan_result {
>  	SCAN_PAGE_DIRTY_OR_WRITEBACK,
>  };
>  
> +enum pte_check_result {
> +	PTE_CHECK_SUCCEED,
> +	PTE_CHECK_FAIL,
> +	PTE_CHECK_CONTINUE,
> +};

I don't love this. "pte_check_result" is a bit too generic for my taste. What is
the difference between "success" and "continue"? Unclear.

Likely, "continue" should actually be something like "skip". But it sounds like
we are mixing two things that shouldn't be mixed (a check that can do more than
just succeed or fail).

Not sure if this was suggested during earlier review, the cover letter doesn't
spell it out. Ideally we'd avoid this completely and just rely on existing error
codes. Like scan_result.


> +
>  #define CREATE_TRACE_POINTS
>  #include <trace/events/huge_memory.h>
>  
> @@ -119,6 +125,20 @@ struct collapse_control {
>  	DECLARE_BITMAP(mthp_present_ptes, MAX_PTRS_PER_PTE);
>  };
>  
> +struct pte_check_context {
> +	struct collapse_control *cc;
> +	struct vm_area_struct *vma;
> +	unsigned int order;
> +	struct folio *folio;
> +	int none_or_zero;
> +	int shared;
> +	int unmapped;
> +	enum scan_result result;
> +	unsigned int max_ptes_none;
> +	unsigned int max_ptes_swap;
> +	unsigned int max_ptes_shared;
> +};
> +
>  /**
>   * struct khugepaged_scan - cursor for scanning
>   * @mm_head: the head of the mm list to scan
> @@ -696,74 +716,131 @@ static void count_collapse_event(unsigned int order, enum vm_event_item vm_event
>  	count_mthp_stat(order, mthp_event);
>  }
>  
> +/*
> + * pte_check_fail() - A simple helper to set the pte_check_context result and
> + * return PTE_CHECK_FAIL.
> + */
> +static enum pte_check_result pte_check_fail(struct pte_check_context *ctx,
> +		enum scan_result result)
> +{
> +	ctx->result = result;
> +	return PTE_CHECK_FAIL;
> +}

Looks a bit over-engineered and the function doc is just unnecessary.

> +
> +/*
> + * collapse_check_pte() - Check if a PTE is suitable for collapse
> + *
> + * Check if a PTE is suitable for collapse based on the following criteria:
> + * - max_pte_* values are not exceeded
> + * - uffd is not active
> + * - lazyfree properties are not present
> + * - only anonymous pages are present
> + *
> + * a helper struct pte_check_context is used to pass and store relevant
> + * information between the collapse_check_pte() function and the caller.
> + *
> + * Return: PTE_CHECK_SUCCEED if the PTE is suitable for collapse,
> + *         PTE_CHECK_FAIL if the PTE is not suitable for collapse,
> + *         PTE_CHECK_CONTINUE if the scan should continue to check the next PTE.
> + */

Why is this doc required?

> +static enum pte_check_result collapse_check_pte(pte_t pteval,
> +		unsigned long addr, struct pte_check_context *ctx)
> +{
> +	if (pte_none_or_zero(pteval)) {
> +		if (++ctx->none_or_zero > ctx->max_ptes_none) {
> +			count_collapse_event(ctx->order, THP_SCAN_EXCEED_NONE_PTE,
> +					     MTHP_STAT_COLLAPSE_EXCEED_NONE);
> +			return pte_check_fail(ctx, SCAN_EXCEED_NONE_PTE);
> +		}
> +		return PTE_CHECK_CONTINUE;
> +	}
> +	if (!pte_present(pteval)) {
> +		if (ctx->unmapped == -1)
> +			return pte_check_fail(ctx, SCAN_PTE_NON_PRESENT);
> +		if (++ctx->unmapped > ctx->max_ptes_swap) {
> +			count_collapse_event(ctx->order, THP_SCAN_EXCEED_SWAP_PTE,
> +					     MTHP_STAT_COLLAPSE_EXCEED_SWAP);
> +			return pte_check_fail(ctx, SCAN_EXCEED_SWAP_PTE);
> +		}
> +		if (pte_swp_uffd_any(pteval))
> +			return pte_check_fail(ctx, SCAN_PTE_UFFD);
> +		return PTE_CHECK_CONTINUE;
> +	}
> +	/*
> +	 * Don't collapse if any of the small PTEs are armed with uffd
> +	 * write protection. Marking the new huge pmd as write protected
> +	 * could bring userfault messages that fall outside of the
> +	 * registered range.
> +	 */
> +	if (pte_uffd(pteval))
> +		return pte_check_fail(ctx, SCAN_PTE_UFFD);
> +
> +	ctx->folio = vm_normal_folio(ctx->vma, addr, pteval);
> +	if (unlikely(!ctx->folio) || unlikely(folio_is_zone_device(ctx->folio)))
> +		return pte_check_fail(ctx, SCAN_PAGE_NULL);
> +
> +	/*
> +	 * If the vma has the VM_DROPPABLE flag, the collapse will
> +	 * preserve the lazyfree property without needing to skip.
> +	 */
> +	if (ctx->cc->is_khugepaged && !(ctx->vma->vm_flags & VM_DROPPABLE) &&
> +	    folio_test_lazyfree(ctx->folio) && !pte_dirty(pteval))
> +		return pte_check_fail(ctx, SCAN_PAGE_LAZYFREE);
> +
> +	if (!folio_test_anon(ctx->folio)) {
> +		VM_WARN_ON_FOLIO(!folio_test_anon(ctx->folio), ctx->folio);

Huh, that looks odd.

That should just be a VM_WARN_ON_FOLIO(true, ..) or sth like that.

But in collapse_scan_pmd() that warning never existed? So this raises eyebrows.

[...]

I'll play with it to see if we can do better and will reply here later.

-- 
Cheers,

David

  parent reply	other threads:[~2026-08-12  8:40 UTC|newest]

Thread overview: 30+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-11 12:48 [PATCH v4 0/7] mm/khugepaged: several cleanups Nico Pache (Red Hat)
2026-08-11 12:48 ` [PATCH v4 1/7] mm/khugepaged: refactor per-scan state clearing into collapse_control_init_scan() Nico Pache (Red Hat)
2026-08-12  9:23   ` Pedro Falcato
2026-08-11 12:48 ` [PATCH v4 2/7] mm/khugepaged: extract reference check into folio_pte_referenced() helper Nico Pache (Red Hat)
2026-08-11 15:47   ` David Hildenbrand (Arm)
2026-08-11 20:45   ` Zi Yan
2026-08-12  9:19   ` Baolin Wang
2026-08-12  9:25   ` Pedro Falcato
2026-08-11 12:48 ` [PATCH v4 3/7] mm/khugepaged: introduce a count_collapse_event() helper Nico Pache (Red Hat)
2026-08-11 20:45   ` Zi Yan
2026-08-12  9:36   ` Pedro Falcato
2026-08-11 12:48 ` [PATCH v4 4/7] mm/khugepaged: fix outdated comments Nico Pache (Red Hat)
2026-08-11 20:48   ` Zi Yan
2026-08-12  9:38   ` Pedro Falcato
2026-08-11 12:48 ` [PATCH v4 5/7] mm/khugepaged: Refactor the PTE state checks into a helper Nico Pache (Red Hat)
2026-08-12  2:04   ` Zi Yan
2026-08-12  8:40   ` David Hildenbrand (Arm) [this message]
2026-08-12  9:51     ` David Hildenbrand (Arm)
2026-08-12 10:06       ` David Hildenbrand (Arm)
2026-08-12 19:39     ` Andrew Morton
2026-08-12 10:50   ` Pedro Falcato
2026-08-11 12:48 ` [PATCH v4 6/7] mm/khugepaged: unmap pte before releasing vma write lock Nico Pache (Red Hat)
2026-08-11 20:54   ` Zi Yan
2026-08-12  9:21   ` Baolin Wang
2026-08-12 10:51   ` Pedro Falcato
2026-08-11 12:48 ` [PATCH v4 7/7] mm: Documentation: clarify where the mTHP stats live Nico Pache (Red Hat)
2026-08-11 20:54   ` Zi Yan
2026-08-12 10:52   ` Pedro Falcato
2026-08-11 18:23 ` [PATCH v4 0/7] mm/khugepaged: several cleanups Andrew Morton
2026-08-11 18:55   ` David Hildenbrand (Arm)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=f1371b4d-6e98-4699-8c1b-7612f6d70f02@kernel.org \
    --to=david@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=corbet@lwn.net \
    --cc=dev.jain@arm.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=nico.pache@linux.dev \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=skhan@linuxfoundation.org \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox