All of lore.kernel.org
 help / color / mirror / Atom feed
From: Nathan Gao <zcgao@amazon.com>
To: SJ Park <sj@kernel.org>
Cc: Nathan Gao <zcgao@amazon.com>, <akpm@linux-foundation.org>,
	<damon@lists.linux.dev>, <linux-mm@kvack.org>,
	<linux-kernel@vger.kernel.org>, <baolin.wang@linux.alibaba.com>,
	<david@kernel.org>, <ryan.roberts@arm.com>
Subject: Re: [PATCH v2] mm/damon/vaddr: use a page-aligned address for the sampling walks
Date: Tue, 1 Sep 2026 13:25:36 -0700	[thread overview]
Message-ID: <20260901202556.39515-1-zcgao@amazon.com> (raw)
In-Reply-To: <20260901013551.91246-1-sj@kernel.org>

On Mon, 31 Aug 2026 18:35:50 -0700 SJ Park <sj@kernel.org> wrote:

> On Mon, 31 Aug 2026 15:11:51 -0700 Nathan Gao <zcgao@amazon.com> wrote:
> 
> > __damon_va_prepare_access_check() picks a random byte address within the
> > region and stores it in r->sampling_addr. There are two users of
> > r->sampling_addr in vaddr.c that pass it into a page table walk, and
> > both use it as the address of a page.
> > 
> >   damon_va_mkold(mm, r->sampling_addr)
> >     damon_va_walk_page_range(mm, addr, addr + 1)
> >       damon_mkold_pmd_entry()
> >         damon_ptep_mkold(pte, vma, addr)
> >           ptep_test_and_clear_young(vma, addr, pte)
> >           mmu_notifier_clear_young(mm, addr, addr + PAGE_SIZE)
> > 
> >   damon_va_young(mm, r->sampling_addr, &folio_sz)
> >     damon_va_walk_page_range(mm, addr, addr + 1)
> >       damon_young_pmd_entry()
> >         ptep_get(pte)
> >         mmu_notifier_test_young(walk->mm, addr)
> > 
> > For arm64, before commit 6f0e1142173a ("arm64: mm: support batch
> > clearing of the young flag for large folios"), the contpte helper walked
> > exactly CONT_PTES entries from the aligned-down page table pointer and
> > used @addr only to pass down to each entry, so an unaligned value was
> > harmless:
> > 
> >         ptep = contpte_align_down(ptep);
> >         addr = ALIGN_DOWN(addr, CONT_PTE_SIZE);
> >         for (i = 0; i < CONT_PTES; i++, ptep++, addr += PAGE_SIZE)
> > 
> > Now the range to walk is derived from @addr instead: end = addr +
> > nr * PAGE_SIZE, rounded up to CONT_PTE_SIZE. For a sample in the last
> > page of a contpte block, the sub-page offset puts end just past the
> > block boundary, so the round-up lands a whole block further and the
> > walk clears PTE_AF in CONT_PTES entries beyond the sampled block. For
> > the last block in a page table page, those entries are past the end of
> > that page, so the walk writes into the page that follows.
> > 
> > Triggered by the full 7.1/7.2 kernel selftest suite on arm64 (EC2
> > c/m6g.4xlarge). The kernel sometimes crashes at or shortly after the
> > DAMON test.
> > 
> > What the overrun does depends on the page that happens to follow the
> > page table, so there is no single signature. If that page is read-only,
> > the write faults in the sampling path itself:
> > 
> >   Unable to handle kernel write to read-only memory at virtual address ffff0003c5d2d000
> >     FSC = 0x0f: level 3 permission fault
> >     CM = 0, WnR = 1, TnD = 0, TagAccess = 0
> >   CPU: 10 UID: 0 PID: 3487 Comm: kdamond.2
> >   pc : contpte_test_and_clear_young_ptes+0x70/0xc0
> >   lr : damon_ptep_mkold+0x1e8/0x1f8
> >   Call trace:
> >    contpte_test_and_clear_young_ptes+0x70/0xc0 (P)
> >    damon_mkold_pmd_entry+0x150/0x170
> >    walk_pmd_range+0x110/0x2b0
> >    walk_pud_range+0x10c/0x208
> >    walk_pgd_range+0x134/0x258
> >    __walk_page_range+0x98/0x1b0
> >    walk_page_range_vma_unsafe+0x90/0x148
> >    walk_page_range_vma+0x28/0x40
> >    damon_va_walk_page_range+0x114/0x2b8
> >    damon_va_prepare_access_checks+0xec/0x1a8
> >    kdamond_fn+0x534/0x770
> >    kthread+0x128/0x138
> >    ret_from_fork+0x10/0x20
> > 
> > Otherwise the page is writable, the PTE_AF clearing succeeds silently
> > and the damage only surfaces later, in whatever happened to own the
> > page, so the backtrace is unrelated to DAMON and differs between runs.
> 
> Urgh, this must have been a painful debugging.  Sorry about that, and
> appreciate your great work on this!
> 

No problem at all!  Had fun digging into this. 

> > 
> > Align the address down to a page boundary in damon_va_mkold() and
> > damon_va_young(), the two users that pass it into a page table walk. It
> > is the address of the page to sample, so this matches its intended
> > meaning. r->sampling_addr itself is left as is, so the sampling and
> > region bookkeeping semantics are unchanged.
> 
> I'm still wondering if it makes sense to restore unaligned address support in
> contpte_test_and_clear_young_ptes() as a long term fix.
> 

I'd appreciate Baolin's thoughts here.  If it turns out callers are
expected to do the alignment, it would be better to have a WARN to expose
the issue.

> > 
> > Fixes: 6f0e1142173a ("arm64: mm: support batch clearing of the young flag for large folios")
> > Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
> > Cc: David Hildenbrand (Arm) <david@kernel.org>
> > Cc: Ryan Roberts <ryan.roberts@arm.com>
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Nathan Gao <zcgao@amazon.com>
> > ---
> > V1 -> V2:
> > - Align inside damon_va_mkold() and damon_va_young(), the two users that
> >   pass the address into a page table walk, rather than aligning
> >   r->sampling_addr itself, so that sub-page sampling addresses remain
> >   possible for future non-PTE access check primitives (SJ)
> > - Point Fixes: at 6f0e1142173a instead of 3f49584b262c, since the
> >   unaligned address was harmless before that commit (SJ)
> > - Describe how the issue was noticed and what it does to the kernel (SJ)
> > - Reword the subject to match the narrower change
> > 
> > v1: https://lore.kernel.org/all/20260827193821.46115-1-zcgao@amazon.com/
> > 
> >  mm/damon/vaddr.c | 6 ++++++
> >  1 file changed, 6 insertions(+)
> > 
> > diff --git a/mm/damon/vaddr.c b/mm/damon/vaddr.c
> > index 2c1c1952c008d..e7aa18200088f 100644
> > --- a/mm/damon/vaddr.c
> > +++ b/mm/damon/vaddr.c
> > @@ -349,6 +349,9 @@ static void damon_va_mkold(struct mm_struct *mm, unsigned long addr)
> >  		.hugetlb_entry = damon_mkold_hugetlb_entry,
> >  	};
> >  
> > +	/* Arch helpers can derive a page range from @addr; align it down. */
> > +	addr = PAGE_ALIGN_DOWN(addr);
> > +
> >  	damon_va_walk_page_range(mm, addr, addr + 1, &damon_mkold_ops, NULL);
> 
> Could we do the alignment just before passing the addr to
> contpte_test_and_clear_young_ptes(), which is the exact function that disallows
> the unaligned address?
> 

Sent a v3.  It moves the alignment into damon_ptep_mkold(), the only place
in DAMON that reaches contpte_test_and_clear_young_ptes(), so it is as
close to that function as DAMON can get:

https://lore.kernel.org/all/20260901201001.33271-1-zcgao@amazon.com/

> >  }
> >  
> > @@ -482,6 +485,9 @@ static bool damon_va_young(struct mm_struct *mm, unsigned long addr,
> >  		.hugetlb_entry = damon_young_hugetlb_entry,
> >  	};
> >  
> > +	/* Arch helpers can derive a page range from @addr; align it down. */
> > +	addr = PAGE_ALIGN_DOWN(addr);
> > +
> >  	damon_va_walk_page_range(mm, addr, addr + 1, &damon_young_ops, &arg);
> 
> Ditto.
> 
> >  	return arg.young;
> >  }
> > -- 
> > 2.50.1
> > 
> > 
> 
> 
> Thanks,
> SJ

Thanks,
Nathan

Sent using hkml (https://github.com/sjp38/hackermail)


  reply	other threads:[~2026-09-01 20:26 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-31 22:11 [PATCH v2] mm/damon/vaddr: use a page-aligned address for the sampling walks Nathan Gao
2026-08-31 22:39 ` sashiko-bot
2026-09-01  1:27   ` SJ Park
2026-09-01  1:35 ` SJ Park
2026-09-01 20:25   ` Nathan Gao [this message]
2026-09-02  0:16     ` SJ Park

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260901202556.39515-1-zcgao@amazon.com \
    --to=zcgao@amazon.com \
    --cc=akpm@linux-foundation.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=damon@lists.linux.dev \
    --cc=david@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ryan.roberts@arm.com \
    --cc=sj@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.