All of lore.kernel.org
 help / color / mirror / Atom feed
From: Lance Yang <lance.yang@linux.dev>
To: david@kernel.org, baolin.wang@linux.alibaba.com,
	pedrodemargomes@gmail.com
Cc: akpm@linux-foundation.org, ljs@kernel.org, ziy@nvidia.com,
	liam@infradead.org, nico.pache@linux.dev, ryan.roberts@arm.com,
	dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev,
	usama.arif@linux.dev, linux-mm@kvack.org,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH mm-stable] mm/khugepaged: avoid unnecessary checking for swap entries when collapsing a mTHP
Date: Thu, 27 Aug 2026 20:11:51 +0800	[thread overview]
Message-ID: <20260827121151.7837-1-lance.yang@linux.dev> (raw)
In-Reply-To: <d02405a5-1d0d-49d5-9056-1a34ccdeef10@kernel.org>


On Thu, Aug 27, 2026 at 01:02:14PM +0200, David Hildenbrand (Arm) wrote:
>On 8/27/26 10:06, Baolin Wang wrote:
>> 
>> 
>> On 8/26/26 3:24 AM, Pedro Demarchi Gomes wrote:
>>> mthp_collapse() tries to swap in PTEs when collapsing a mTHP if there are any
>>> swap PTEs in the PMD range, even if none of those swap PTEs are
>>> actually part of the mTHP's range.
>> 
>> Are you sure? I wonder how you tested your patch? Because we never swapin PTEs
>> for mTHP collapse, see the code in __collapse_huge_page_swapin():
>
>I'm confused as well, this doesn't really make sense.

Well ... the change does remove a redundant PTE walk, IIUC ...

I think Pedro's wording is causing the confusion :)

Yeah, Baolin is right that __collapse_huge_page_swapin() never reaches
do_swap_page() for a mTHP. For an otherwise eligible lower-order
candidate, current code still calls it and walks the candidate's PTE range
when an unrelated swap PTE exists elsewhere in the same PMD :)

Assume an otherwise eligible lower-order candidate has no swap PTE, while
another subrange in the PMD has one. collapse_scan_pmd() counts unmapped
over the full PMD and passes that PMD-wide value into mthp_collapse():

static enum scan_result collapse_scan_pmd(struct mm_struct *mm,
		struct vm_area_struct *vma, unsigned long start_addr,
		bool *lock_dropped, struct collapse_control *cc)
{
...
	int node = NUMA_NO_NODE, unmapped = 0;
...
	for (i = 0; i < HPAGE_PMD_NR; i++) {
		_pte = pte + i;
		addr = start_addr + i * PAGE_SIZE;
		pteval = ptep_get(_pte);
...
		if (pte_none_or_zero(pteval)) {
...
			continue;
		}
		if (!pte_present(pteval)) {
			if (++unmapped > max_ptes_swap) {
...
			}
...
			if (pte_swp_uffd_any(pteval)) {
				result = SCAN_PTE_UFFD;
				goto out_unmap;
			}
			continue;
		}
...
	}
	if (cc->is_khugepaged &&
		   (!referenced ||
		    (unmapped && referenced < HPAGE_PMD_NR / 2))) {
		result = SCAN_LACK_REFERENCED_PAGE;
	} else {
		result = SCAN_SUCCEED;
	}
...
	if (result == SCAN_SUCCEED) {
...
		result = mthp_collapse(mm, start_addr, referenced,
				       unmapped, cc, enabled_orders);
...
	}
...
	return result;
}

unmapped is PMD-wide above. mthp_collapse() then passes the same value to
every attempted candidate:

static enum scan_result mthp_collapse(struct mm_struct *mm,
		unsigned long address, int referenced, int unmapped,
		struct collapse_control *cc, unsigned long enabled_orders)
{
	unsigned int nr_occupied_ptes, nr_ptes, max_ptes_none;
...
	unsigned int order = HPAGE_PMD_ORDER;

	while (offset < HPAGE_PMD_NR) {
		nr_ptes = 1UL << order;

		if (!test_bit(order, &enabled_orders))
			goto next_order;

		max_ptes_none = collapse_max_ptes_none(cc, NULL, order);
		nr_occupied_ptes = bitmap_weight_from(cc->mthp_present_ptes, offset,
						      offset + nr_ptes);

		/*
		 * Swap PTEs accepted during the scan are counted in @unmapped,
		 * not in the present-PTE bitmap. Account them for the PMD-order
		 * candidate.
		 */
		if (is_pmd_order(order))
			nr_occupied_ptes += unmapped;

		if (nr_occupied_ptes >= nr_ptes - max_ptes_none) {
			enum scan_result ret;

			collapse_address = address + offset * PAGE_SIZE;
			ret = collapse_huge_page(mm, collapse_address, referenced,
						 unmapped, cc, order);
...
		}

next_order:
...
		if (order > KHUGEPAGED_MIN_MTHP_ORDER &&
			(enabled_orders & GENMASK(order - 1, 0))) {
			order--;
			continue;
		}
next_offset:
...
		offset += nr_ptes;
		order = max_order_from_offset(offset);
	}
...
}

Once it reaches the lower-order candidate with no swap PTE, unmapped is
still nonzero and collapse_huge_page() calls the swapin helper:

static enum scan_result collapse_huge_page(struct mm_struct *mm, unsigned long start_addr,
		int referenced, int unmapped, struct collapse_control *cc,
		unsigned int order)
{
...
	if (unmapped) {
...
		result = __collapse_huge_page_swapin(mm, vma, start_addr, pmd,
						     referenced, order);
...
	}
...
}

__collapse_huge_page_swapin() only calls do_swap_page() after finding a
non-present, non-none PTE, and lower orders return before that call:

static enum scan_result __collapse_huge_page_swapin(struct mm_struct *mm,
		struct vm_area_struct *vma, unsigned long start_addr,
		pmd_t *pmd, int referenced, unsigned int order)
{
...
	unsigned long addr, end = start_addr + (PAGE_SIZE << order);
	enum scan_result result;
	pte_t *pte = NULL;
	spinlock_t *ptl;

	for (addr = start_addr; addr < end; addr += PAGE_SIZE) {
...
		vmf.orig_pte = ptep_get_lockless(pte);
		if (pte_none(vmf.orig_pte) ||
		    pte_present(vmf.orig_pte))
			continue;
...
		if (!is_pmd_order(order)) {
...
			result = SCAN_EXCEED_SWAP_PTE;
			goto out;
		}

		vmf.pte = pte;
		vmf.ptl = ptl;
		ret = do_swap_page(&vmf);
...
	}
...
	result = SCAN_SUCCEED;
out:
...
	return result;
}

For the candidate above, the loop only reads its PTEs and returns
SCAN_SUCCEED. Pedro's bitmap lets mthp_collapse() determine that upfront
and skip the walk.

Emm ... as I asked before[1], any numbers showing how much the extra scan
costs?

[1] https://lore.kernel.org/lkml/20260825192433.3185880-1-pedrodemargomes@gmail.com/

Cheers, Lance


  reply	other threads:[~2026-08-27 12:12 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-25 19:24 [PATCH mm-stable] mm/khugepaged: avoid unnecessary checking for swap entries when collapsing a mTHP Pedro Demarchi Gomes
2026-08-26  4:43 ` Lance Yang
2026-08-27  8:06 ` Baolin Wang
2026-08-27 11:02   ` David Hildenbrand (Arm)
2026-08-27 12:11     ` Lance Yang [this message]
2026-08-27 14:07       ` Kiryl Shutsemau
2026-08-31 18:53         ` Pedro Demarchi Gomes
2026-09-02 10:08           ` Kiryl Shutsemau
2026-08-28  3:24       ` Pedro Demarchi Gomes

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260827121151.7837-1-lance.yang@linux.dev \
    --to=lance.yang@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=nico.pache@linux.dev \
    --cc=pedrodemargomes@gmail.com \
    --cc=ryan.roberts@arm.com \
    --cc=usama.arif@linux.dev \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.