All of lore.kernel.org
 help / color / mirror / Atom feed
From: Yeoreum Yun <yeoreum.yun@arm.com>
To: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Yeoreum Yun <yeoreum.yun@arm.com>, Zi Yan <ziy@nvidia.com>,
	Andrew Morton <akpm@linux-foundation.org>,
	David Hildenbrand <david@kernel.org>,
	Lorenzo Stoakes <ljs@kernel.org>,
	"Liam R. Howlett" <liam@infradead.org>,
	Nico Pache <nico.pache@linux.dev>,
	Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
	Barry Song <baohua@kernel.org>, Lance Yang <lance.yang@linux.dev>,
	Usama Arif <usama.arif@linux.dev>,
	Vlastimil Babka <vbabka@kernel.org>,
	Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	Michal Hocko <mhocko@suse.com>, Shuah Khan <shuah@kernel.org>,
	Kevin Brodsky <kevin.brodsky@arm.com>,
	linux-mm@kvack.org, linux-kselftest@vger.kernel.org,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH v2 2/2] kselftest: mm: replace usage of /proc/self/smaps for check_huge_xxx() helper
Date: Fri, 28 Aug 2026 08:44:38 +0100	[thread overview]
Message-ID: <apE8ZvWg219b9aW_@e129823.arm.com> (raw)
In-Reply-To: <31e649d3-7653-49c0-b73e-933ba2a86f4a@linux.alibaba.com>

> 
> 
> On 8/28/26 12:54 AM, Yeoreum Yun wrote:
> > On Thu, Aug 27, 2026 at 11:03:17AM -0400, Zi Yan wrote:
> > > On Thu Aug 27, 2026 at 6:44 AM EDT, Yeoreum Yun wrote:
> > > > Hi Baolin,
> > > > 
> > > > > 
> > > > > 
> > > > > On 8/26/26 8:24 PM, Yeoreum Yun wrote:
> > > > > > Since glibc commit 321e1fc73f (“malloc: Enable 2MB THP by default on AArch64”),
> > > > > > glibc may call madvise(MADV_HUGEPAGE) for sufficiently large allocations
> > > > > > made by memalign().
> > > > > > 
> > > > > > The underlying VMA may start at a different address from the aligned
> > > > > > address returned by memalign(). Furthermore, a subsequent
> > > > > > madvise(MADV_HUGEPAGE) call does not split the VMA because the flag is
> > > > > > already set.
> > > > > > 
> > > > > > This causes split_huge_page_test to fail because the check_huge_xxx()
> > > > > > helpers incorrectly require the address returned by memalign() to
> > > > > > match the VMA start address reported in /proc/self/smaps.
> > > > > > 
> > > > > > Fix this by using /proc/self/pagemap and /proc/kpageflags instead of
> > > > > > /proc/self/smaps to detect huge pages.
> > > > > > 
> > > > > > Reported-by: David Hildenbrand (Arm) <david@kernel.org>
> > > > > > Signed-off-by: Yeoreum Yun <yeoreum.yun@arm.com>
> > > > > > ---
> > > > > >    tools/testing/selftests/mm/vm_util.c | 130 ++++++++++++++++++++---------------
> > > > > >    tools/testing/selftests/mm/vm_util.h |   1 +
> > > > > >    2 files changed, 77 insertions(+), 54 deletions(-)
> > > > > > 
> > > > > > diff --git a/tools/testing/selftests/mm/vm_util.c b/tools/testing/selftests/mm/vm_util.c
> > > > > > index 4821a3563036..1d0959b3b9e8 100644
> > > > > > --- a/tools/testing/selftests/mm/vm_util.c
> > > > > > +++ b/tools/testing/selftests/mm/vm_util.c
> > > > > > @@ -351,31 +351,13 @@ char *__get_smap_entry(void *addr, const char *pattern, char *buf, size_t len)
> > > > > >    	return entry;
> > > > > >    }
> > > > > > -static bool __check_pmd_huge(void *addr, char *pattern, int nr_hpages,
> > > > > > -		  uint64_t hpage_size)
> > > > > > -{
> > > > > > -	char buffer[MAX_LINE_LENGTH];
> > > > > > -	uint64_t thp = -1;
> > > > > > -	char *entry;
> > > > > > -
> > > > > > -	entry = __get_smap_entry(addr, pattern, buffer, sizeof(buffer));
> > > > > > -	if (!entry)
> > > > > > -		goto err_out;
> > > > > > -
> > > > > > -	if (sscanf(entry, "%9" SCNu64 " kB", &thp) != 1)
> > > > > > -		ksft_exit_fail_msg("Reading smap error\n");
> > > > > > -
> > > > > > -err_out:
> > > > > > -	return thp == (nr_hpages * (hpage_size >> 10));
> > > > > > -}
> > > > > > -
> > > > > > -static bool check_large_folios(void *addr, size_t len, int nr_hpages,
> > > > > > -		uint64_t hpage_size)
> > > > > > +static bool check_large_folios(int pagemap_fd, int kpageflags_fd,
> > > > > > +			       void *addr, size_t len, int nr_hpages,
> > > > > > +			       uint64_t hpage_size)
> > > > > >    {
> > > > > >    	int order = 0, pagesize = getpagesize();
> > > > > >    	unsigned int nr_pages = hpage_size / pagesize;
> > > > > >    	int orders[MAX_NR_ORDERS], status;
> > > > > > -	int pagemap_fd, kpageflags_fd;
> > > > > >    	bool ret = false;
> > > > > >    	if (!nr_pages)
> > > > > > @@ -386,15 +368,6 @@ static bool check_large_folios(void *addr, size_t len, int nr_hpages,
> > > > > >    		ksft_exit_fail_msg("invalid order\n");
> > > > > >    	memset(orders, 0, sizeof(int) * MAX_NR_ORDERS);
> > > > > > -	pagemap_fd = open(PAGEMAP_PATH, O_RDONLY);
> > > > > > -	if (pagemap_fd == -1)
> > > > > > -		ksft_exit_fail_msg("read pagemap fail\n");
> > > > > > -
> > > > > > -	kpageflags_fd = open(KPAGEFLAGS_PATH, O_RDONLY);
> > > > > > -	if (kpageflags_fd == -1) {
> > > > > > -		close(pagemap_fd);
> > > > > > -		ksft_exit_fail_msg("read kpageflags fail\n");
> > > > > > -	}
> > > > > >    	status = gather_folio_orders(addr, len, pagemap_fd,
> > > > > >    			kpageflags_fd, orders, MAX_NR_ORDERS);
> > > > > > @@ -405,48 +378,97 @@ static bool check_large_folios(void *addr, size_t len, int nr_hpages,
> > > > > >    		ret = true;
> > > > > >    out:
> > > > > > -	close(pagemap_fd);
> > > > > > -	close(kpageflags_fd);
> > > > > >    	return ret;
> > > > > >    }
> > > > > > -bool check_huge_anon(void *addr, size_t len, int nr_hpages, uint64_t hpage_size)
> > > > > > +enum check_huge_type {
> > > > > > +	CHECK_HUGE_ANON,
> > > > > > +	CHECK_HUGE_FILE,
> > > > > > +	CHECK_HUGE_SHMEM,
> > > > > > +};
> > > > > > +
> > > > > > +static bool __check_pmd_huge(void *addr, size_t len, int nr_hpages,
> > > > > > +			     uint64_t hpage_size, enum check_huge_type type)
> > > > > 
> > > > > The original __check_pmd_huge() is only for PMD-sized large folios, but now
> > > > > it not only checks PMD-sized large folios but also mTHP large folios, which
> > > > > I find confusing. Please keep its original semantics, and only check
> > > > > PMD-sized large folios.
> > > > 
> > > > But, It seems to valuable to check other page-flags than checking
> > > > the large-folio only.
> > > > 
> > > > > 
> > > > > >    {
> > > > > > -	uint64_t pmd_pagesize = read_pmd_pagesize();
> > > > > > +	int pagemap_fd, kpageflags_fd;
> > > > > > +	uint64_t pmd_pagesize, granule;
> > > > > > +	uint64_t categories, kpf;
> > > > > > +	unsigned long pfn;
> > > > > > +	bool check_large, huge_mapped;
> > > > > > +	char *start = addr;
> > > > > > +	char *end = start + len;
> > > > > > +	pmd_pagesize = read_pmd_pagesize();
> > > > > >    	if (!pmd_pagesize)
> > > > > >    		ksft_exit_fail_msg("reading PMD pagesize failed\n");
> > > > > > -	if (hpage_size == pmd_pagesize)
> > > > > > -		return __check_pmd_huge(addr, "AnonHugePages: ", nr_hpages, hpage_size);
> > > > > > +	if (nr_hpages > 0) {
> > > > > > +		check_large = true;
> > > > > > +		granule = hpage_size;
> > > > > > +	} else {
> > > > > > +		check_large = false;
> > > > > > +		granule = psize();
> > > > > > +	}
> > > > > 
> > > > > This is incorrect for the mTHP large folio check. I already hit a selftest
> > > > > failure. Please test your patches before sending them out.
> > > > 
> > > > Since for a split case, large folio can be as-is but only remove
> > > > the PMD mapping only, skipping the large_folio checking seems valid
> > > > when nr_hpage is 0.
> > > 
> > > What split care are you referring to? split_huge_page_test() always
> > > splits the folio.
> > 
> > Not for split_huge_page_test case but for khugepage testcase like
> > collapse_full_of_compound() test.
> > 
> > > 
> > > > 
> > > > And the failure of test seems because of unmapped area after
> > > > changinng the collapse-order. Therefore, it seems to fine with below
> > > > patch:
> > > > 
> > > > ---------------&<----------------------
> > > > 
> > > > diff --git a/tools/testing/selftests/mm/vm_util.c b/tools/testing/selftests/mm/vm_util.c
> > > > index 1d0959b3b9e8..f174a76d2310 100644
> > > > --- a/tools/testing/selftests/mm/vm_util.c
> > > > +++ b/tools/testing/selftests/mm/vm_util.c
> > > > @@ -387,14 +387,14 @@ enum check_huge_type {
> > > >          CHECK_HUGE_SHMEM,
> > > >   };
> > > > 
> > > > -static bool __check_pmd_huge(void *addr, size_t len, int nr_hpages,
> > > > -                            uint64_t hpage_size, enum check_huge_type type)
> > > > +static bool __check_huge(void *addr, size_t len, int nr_hpages,
> > > > +                        uint64_t hpage_size, enum check_huge_type type)
> > > >   {
> > > >          int pagemap_fd, kpageflags_fd;
> > > >          uint64_t pmd_pagesize, granule;
> > > >          uint64_t categories, kpf;
> > > >          unsigned long pfn;
> > > > -       bool check_large, huge_mapped;
> > > > +       bool check_large, check_huge_mapped, allow_nomap;
> > > >          char *start = addr;
> > > >          char *end = start + len;
> > > > 
> > > > @@ -405,11 +405,20 @@ static bool __check_pmd_huge(void *addr, size_t len, int nr_hpages,
> > > >          if (nr_hpages > 0) {
> > > >                  check_large = true;
> > > >                  granule = hpage_size;
> > > > +               if (granule == pmd_pagesize)
> > > > +                       check_huge_mapped = true;
> > > > +               else
> > > > +                       check_huge_mapped = false;
> > > >          } else {
> > > >                  check_large = false;
> > > >                  granule = psize();
> > > >          }
> > > 
> > > The else is for nr_hpages == 0? But it looks like we allow negative
> > > nr_hpages. Maybe add a bool expect_huge = nr_hpages > 0 to make it
> > > explicit.
> > 
> > Fair enough. I'll change.
> > 
> > > 
> > > granule is an optimization for PAGE_IS_HUGE scanning? When we expect a
> > > PMD mapping, we just scan at pmd_pagesize granularity, otherwise we
> > > check every single page?
> > > 
> > > At the high level, the function looks good to me, the rules are:
> > > 
> > > 1. if hpage_size == pmd_pagesize, we need to check PAGE_IS_HUGE and
> > > check_large_folio() can be skipped, since we only care about mappings.
> > > This checks for PMD mappings.
> > > 
> > > 2. in other cases, check_large_folios() is always needed. This is for
> > > mTHP checks.
> 
> Yes, this is also what I thought. And the following code looks good to me.
> Thanks.
> 
> > Exactly, but for some testcase where use this function, doesn't
> > trigger split of large _folio but only check the HUGE MAP is removed
> > (nr_hpage == 0), it skips the chekc_large_folio().
> > 
> > > 
> > > I think the ifs at the beginning is confusing. Can we do something like
> > > below to get rid of the ifs? I also moved KPF_* checks in a separate
> > > function. Feel free to make changes if you find any issue there.
> > > 
> > > static bool check_huge_type(uint64_t categories, uint64_t kpageflags,
> > > 			    enum check_huge_type type)
> > > {
> > > 	bool file = categories & PAGE_IS_FILE;
> > > 	bool swapbacked = kpageflags & KPF_SWAPBACKED;
> > > 
> > > 	switch (type) {
> > > 	case CHECK_HUGE_ANON:
> > > 		return !file;
> > > 	case CHECK_HUGE_FILE:
> > > 		return file && !swapbacked;
> > > 	case CHECK_HUGE_SHMEM:
> > > 		return file && swapbacked;
> > > 	}
> > > 
> > > 	return false;
> > > }
> > > 
> > > static bool __check_huge(void *addr, size_t len, int nr_hpages,
> > > 			 uint64_t hpage_size, enum check_huge_type type)
> > > {
> > > 	int pagemap_fd, kpageflags_fd;
> > > 	int nr_pmd_mappings = 0;
> > > 	uint64_t pmd_pagesize, scan_mapping_size;
> > > 	uint64_t categories, kpf;
> > > 	unsigned long pfn;
> > > 	bool check_pmd_mapping;
> > > 	bool allow_nonpresent;
> > > 	bool ret = false;
> > > 	char *start = addr;
> > > 	char *end = start + len;
> > > 
> > > 	pmd_pagesize = read_pmd_pagesize();
> > > 	if (!pmd_pagesize)
> > > 		ksft_exit_fail_msg("reading PMD pagesize failed\n");
> > > 
> > > 	check_pmd_mapping = hpage_size == pmd_pagesize;
> > > 	scan_mapping_size = nr_hpages > 0 ? hpage_size : psize();
> > > 	/* Some mTHP tests check a partially populated PMD-sized range. */
> > > 	allow_nonpresent = (uint64_t)nr_hpages * hpage_size < len;
> > > 
> > > 	pagemap_fd = open(PAGEMAP_PATH, O_RDONLY);
> > > 	if (pagemap_fd < 0)
> > > 		ksft_exit_fail_msg("open pagemap fail\n");
> > > 
> > > 	kpageflags_fd = open(KPAGEFLAGS_PATH, O_RDONLY);
> > > 	if (kpageflags_fd < 0) {
> > > 		close(pagemap_fd);
> > > 		ksft_exit_fail_msg("open kpageflags fail\n");
> > > 	}
> > > 
> > > 	/* PTE-mapped large folios cannot be identified by PAGE_IS_HUGE. */
> > > 	if (!check_pmd_mapping &&
> > > 	    !check_large_folios(pagemap_fd, kpageflags_fd, addr, len,
> > > 				 nr_hpages, hpage_size))
> > > 		goto out;
> > > 
> > > 	for (; start < end; start += scan_mapping_size) {
> > > 		categories = pagemap_scan_get_categories(pagemap_fd, start);
> > > 		pfn = pagemap_get_pfn(pagemap_fd, start);
> > > 		if (pfn == -1UL) {
> > > 			if (!allow_nonpresent)
> > > 				goto out;
> > > 			continue;
> > > 		}
> > > 		if (pageflags_get(pfn, kpageflags_fd, &kpf))
> > > 			ksft_exit_fail_msg("read kpageflags: %s\n", strerror(errno));
> > > 		if (check_pmd_mapping && (categories & PAGE_IS_HUGE))
> > > 			nr_pmd_mappings++;
> > > 		if (kpf & KPF_COMPOUND_TAIL)
> > > 			continue;
> > > 		if (!check_huge_type(categories, kpf, type))
> > > 			goto out;
> > > 	}
> > > 	if (check_pmd_mapping && nr_pmd_mappings != nr_hpages)
> > > 		goto out;
> > 
> > Again, because of collapse_full_of_compound() testcase,
> > it would be failed for this. so, it would better to skip
> > check_large_folioes() when nr_hpages is 0.
> 
> IIUC, when nr_hpages is 0, we should also call check_large_folios() to
> verify that there are no large folios within this address range, which is
> also what the previous code did.
> 
> Also, I tried Zi's code, and it passed all my khugepaged selftests. I'm not
> sure why you saw the collapse_full_of_compound() test case fail.

Ah, Sorry to my misunderstding I've misread the code on condition :(.
(!check_pmd_mapping as check_pmd_mapping)...

Yeap. it seems fine to check PAGE_IS_HUGE instead of large folio for
pmd_hugepage case.

Sorry to make noise.

-- 
Sincerely,
Yeoreum Yun


  reply	other threads:[~2026-08-28  7:44 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-26 12:24 [PATCH v2 0/2] kselftest: mm: fix some failure of split_huge_page_test Yeoreum Yun
2026-08-26 12:24 ` [PATCH v2 1/2] kselftest: mm: prevent random failure of huge page split for khugepaged Yeoreum Yun
2026-08-27 10:59   ` David Hildenbrand (Arm)
2026-08-27 11:11     ` Yeoreum Yun
2026-08-26 12:24 ` [PATCH v2 2/2] kselftest: mm: replace usage of /proc/self/smaps for check_huge_xxx() helper Yeoreum Yun
2026-08-27  8:28   ` Baolin Wang
2026-08-27  9:11     ` Yeoreum Yun
2026-08-27 10:44     ` Yeoreum Yun
2026-08-27 15:03       ` Zi Yan
2026-08-27 16:54         ` Yeoreum Yun
2026-08-28  1:00           ` Baolin Wang
2026-08-28  7:44             ` Yeoreum Yun [this message]
2026-08-27 10:56     ` David Hildenbrand (Arm)
2026-08-27 11:07       ` Yeoreum Yun

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=apE8ZvWg219b9aW_@e129823.arm.com \
    --to=yeoreum.yun@arm.com \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=kevin.brodsky@arm.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=mhocko@suse.com \
    --cc=nico.pache@linux.dev \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=shuah@kernel.org \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.