From: Nathan Gao <zcgao@amazon.com>
To: <sj@kernel.org>
Cc: <akpm@linux-foundation.org>, <baolin.wang@linux.alibaba.com>,
<damon@lists.linux.dev>, <linux-kernel@vger.kernel.org>,
<linux-mm@kvack.org>, <stable@vger.kernel.org>,
<zcgao@amazon.com>
Subject: Re: [PATCH] mm/damon: use a page-aligned sampling address
Date: Mon, 31 Aug 2026 15:50:32 -0700 [thread overview]
Message-ID: <20260831225032.58045-1-zcgao@amazon.com> (raw)
In-Reply-To: <20260829015511.74221-1-sj@kernel.org>
On Fri, 28 Aug 2026 18:55:11 -0700 SJ Park <sj@kernel.org> wrote:
> On Fri, 28 Aug 2026 18:04:55 -0700 Nathan Gao <zcgao@amazon.com> wrote:
>
> > Hi SJ,
> >
> > Thanks for your review!
> >
> > On Thu, 27 Aug 2026 17:22:11 -0700 SJ Park <sj@kernel.org> wrote:
> >
> > > Hello Nathan,
> > >
> > > On Thu, 27 Aug 2026 12:38:21 -0700 Nathan Gao <zcgao@amazon.com> wrote:
> > >
> > > > __damon_va_prepare_access_check() picks a random byte address within the
> > > > region and stores it in r->sampling_addr. There are two users of
> > > > r->sampling_addr in vaddr.c that pass it into a page table walk, and
> > > > both use it as the address of a page.
> > > >
> > > > damon_va_mkold(mm, r->sampling_addr)
> > > > damon_va_walk_page_range(mm, addr, addr + 1)
> > > > damon_mkold_pmd_entry()
> > > > damon_ptep_mkold(pte, vma, addr)
> > > > ptep_test_and_clear_young(vma, addr, pte)
> > > > mmu_notifier_clear_young(mm, addr, addr + PAGE_SIZE)
> > > >
> > > > damon_va_young(mm, r->sampling_addr, &folio_sz)
> > > > damon_va_walk_page_range(mm, addr, addr + 1)
> > > > damon_young_pmd_entry()
> > > > ptep_get(pte)
> > > > mmu_notifier_test_young(walk->mm, addr)
> > > >
> > > > test_and_clear_young_ptes(), which backs ptep_test_and_clear_young() on
> > > > arm64, documents @addr as "Address the first page is mapped at".
> > > >
> > > > For arm64, before commit 6f0e1142173a ("arm64: mm: support batch
> > > > clearing of the young flag for large folios"),
> > >
> > > The @addr documentation is also introduced by this commit. This commit is
> > > authored at 2026-02-09.
>
> I was wrong. The documentation was introduced by commit 6d7237dda44f ("mm: add
> a batched helper to clear the young flag for large folios"), which was authored
> by Baolin on 2026-03-06.
>
> > >
> > > > the contpte helper walked
> > > > exactly CONT_PTES entries from the aligned-down page table pointer and
> > > > used @addr only to pass down to each entry, so an unaligned value was
> > > > harmless:
> > > >
> > > > ptep = contpte_align_down(ptep);
> > > > addr = ALIGN_DOWN(addr, CONT_PTE_SIZE);
> > > > for (i = 0; i < CONT_PTES; i++, ptep++, addr += PAGE_SIZE)
> > >
> > > So, there was no issue before the commit.
> > >
> >
> > Right. Before 6f0e1142173a, unaligned addresses were tolerated but I don't
> > think this is guaranteed.
> >
> > > >
> > > > Align the sampled address down to a page boundary. It is the address of
> > > > the page to sample, so this matches its intended meaning and fixes both
> > > > users in vaddr.c.
> > >
> > > This indeed sounds like can fix the issue to me. However, was it a clear rule
> > > that we should pass only contepte-aligned addrss to
> > > ptep_test_and_clear_young()? And is DAMON the only ptep_test_and_clear_young()
> > > caller that is mistakenly passing the unaligned address?
> > >
> >
> > It is not spelled out as an explicit rule, but the documented "Address
> > the first page is mapped at" implies it,
>
> You mean the comment on test_and_clear_young_ptes(), right? But as I mentioned
> above, the comment was introduced by Baolin's patch that was authored on
> 2026-03-06. I'd still appreciate Baolin's opinion.
>
I think you are right. I shouldn't reference this and I dropped it in v2 commit
message.
> > and these callers are using
> > aligned addresses:
> >
> > mm/page_idle.c: page_idle_clear_pte_refs_one()
> > fs/proc/task_mmu.c: clear_refs_pte_range()
> >
> > > If not, it might make sense to make contpte_test_and_clear_young_ptes() support
> > > unaligned adress again in my opinion. May I ask your opinion, Baolin?
> > >
> > > >
> > > > Fixes: 3f49584b262c ("mm/damon: implement primitives for the virtual memory address spaces")
> > >
> > > I think 6f0e1142173a ("arm64: mm: support batch clearing of the young flag for
> > > large folios") would be mroe correct 'Fixes:', if there was no issue before the
> > > commit.
> > >
> >
> > Will use that in v2.
> >
> > > > - r->sampling_addr = damon_rand(ctx, r->ar.start, r->ar.end);
> > > > + r->sampling_addr = PAGE_ALIGN_DOWN(damon_rand(ctx, r->ar.start,
> > > > + r->ar.end));
> > >
> > > If we need to have the fix in DAMON, this kind of change would be needed.
> > >
> > > However, what happens if the address is backed by large folios?
> > >
> > > Before the commit 6f0e1142173a, also, it was aligning to CONT_PTE_SIZE. Should
> > > we do same?
> >
> > Passing a page-aligned address restores the pre-6f0e1142173a behavior.
> > Before the change, the helper aligned ptep down to the block start and
> > walked a fixed CONT_PTES entries, regardless of addr. After the
> > change, the walk covers [ALIGN_DOWN(addr, CONT_PTE_SIZE),
> > ALIGN(addr + nr * PAGE_SIZE, CONT_PTE_SIZE)). With a sub-page offset,
> > addr + PAGE_SIZE lands just past the block boundary, so the round-up
> > extends the walk a whole block further. With a page-aligned addr,
> > addr + PAGE_SIZE is at most the block end, so the round-up lands
> > exactly on the block end and the walk covers the same CONT_PTES
> > entries as before the commit.
> >
> > We also can't align to CONT_PTE_SIZE in DAMON since it's defined only under
> > arch/arm64/:
> >
> > #define CONT_PTES (1 << (CONT_PTE_SHIFT - PAGE_SHIFT))
> > #define CONT_PTE_SIZE (CONT_PTES * PAGE_SIZE)
>
> DAMON cares only exactly the byte of the address, so I agree this would work
> for DAMON and be safe. But, still the behavior is not exactly same to
> pre-6f0e1142173a, isn't it? I'm not really sure if this is really the correct
> use of the function. Again, I'd appreciate Baolin's comment.
>
I'd appreciate Baolin's clarification as well here.
> >
> >
> > > Also, I think we should pass aligned address to only the functions that
> > > require alignement. Making the alignment to the sampling address in general
> > > sounds too much to me. Particularly, we are working on supporting new page
> > > access check primitives other than PTE Accessed bit, like AMd IBS. In the
> > > case, we might support <PAGE_SIZE granularity monitoring. Aligning sampling
> > > address in general will make it more complicated.
> >
> > Makes sense. I will keep sampling_addr as is and align inside damon_va_mkold()
> > and damon_va_young() in v2.
>
> Regardless of Baolin's comment, let's fix this issue. So the v2 would be
> appreciated. In the v2, could you also add more details about how the issue
> can be reproduced, and the user impact? You mentioned you found memory
> corruption from DAMON selftets. It would be nice if you could make it more
> detailed, such as what selftest reproduces the issue and what symptoms it
> showed you.
Added the crash stack trace and the impact to the v2 commit message:
https://lore.kernel.org/all/20260831221151.50561-1-zcgao@amazon.com/
>
> Nevertheless I'm also wondering if supporting unaligned address again, like
> below also works.
>
> '''
> --- a/arch/arm64/mm/contpte.c
> +++ b/arch/arm64/mm/contpte.c
> @@ -519,9 +519,12 @@ bool contpte_test_and_clear_young_ptes(struct vm_area_struct *vma,
> * of the same large folio in a single VMA and a single page table.
> */
>
> - unsigned long end = addr + nr * PAGE_SIZE;
> + unsigned long end;
> bool young = false;
>
> + ptep = contpte_align_down(ptep);
> + addr = ALIGN_DOWN(addr, CONT_PTE_SIZE);
> + end = addr + nr * PAGE_SIZE;
> ptep = contpte_align_addr_ptep(&addr, &end, ptep, nr);
> for (; addr != end; ptep++, addr += PAGE_SIZE)
> young |= __ptep_test_and_clear_young(vma, addr, ptep);
> '''
>
> Nathan, what do you think? If it makes sense to you, could you also test this?
I think this may introduce an issue for the case that nr > 1. If the
address is in the middle of the block, an align like this would move it
to the beginning of the block, so the walk would cover the wrong entries.
>
>
> Thanks,
> SJ
>
> [...]
Thanks,
Nathan
Sent using hkml (https://github.com/sjp38/hackermail)
prev parent reply other threads:[~2026-08-31 22:50 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-27 19:38 [PATCH] mm/damon: use a page-aligned sampling address Nathan Gao
2026-08-27 19:50 ` sashiko-bot
2026-08-28 0:28 ` SJ Park
2026-08-28 0:22 ` SJ Park
2026-08-29 1:04 ` Nathan Gao
2026-08-29 1:55 ` SJ Park
2026-08-31 22:50 ` Nathan Gao [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260831225032.58045-1-zcgao@amazon.com \
--to=zcgao@amazon.com \
--cc=akpm@linux-foundation.org \
--cc=baolin.wang@linux.alibaba.com \
--cc=damon@lists.linux.dev \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=sj@kernel.org \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.