From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: Kefeng Wang <wangkefeng.wang@huawei.com>
Cc: "David Hildenbrand (Arm)" <david@kernel.org>,
Andrew Morton <akpm@linux-foundation.org>,
linux-mm@kvack.org, "Liam R. Howlett" <liam@infradead.org>,
Vlastimil Babka <vbabka@kernel.org>,
Jann Horn <jannh@google.com>, Pedro Falcato <pfalcato@suse.de>,
Zi Yan <ziy@nvidia.com>
Subject: Re: [PATCH mm-new v3] mm: mincore: use per-vma lock during page table walk
Date: Wed, 16 Sep 2026 12:19:50 +0100 [thread overview]
Message-ID: <aqp4Uvem1MWQOUZA@gremlin> (raw)
In-Reply-To: <1ddf7304-70d9-4e35-935a-9d622b022e4f@huawei.com>
On Wed, Sep 16, 2026 at 02:50:17PM +0800, Kefeng Wang wrote:
>
>
> On 9/16/2026 2:33 PM, David Hildenbrand (Arm) wrote:
> > On 9/16/26 06:31, Kefeng Wang wrote:
> > > do_mincore() performs a read-only, per-VMA residency query,
> > > making it a good candidate for per-VMA locking. Convert it
> > > to acquire the per-VMA lock, thereby reducing contention on
> > > the per-MM mmap_lock.
> > >
> > > Reviewed-by: Pedro Falcato <pfalcato@suse.de>
> > > Signed-off-by: Kefeng Wang <wangkefeng.wang@huawei.com>
LGTM so:
Reviewed-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
> > > ---
> > > v3:
> > > - add vma_assert_locked, update changelog/comment, per David
> >
> > It likely was Lorenzo :)
> >
>
> Oh, I'm completely blind :)
>
> > > - Add RB
> > > v2, (RESEND):
> > > - using new vma_start_read_unlocked() API, suggestted by Pedro Falcato
> > > v1:
> > > - https://lore.kernel.org/linux-mm/20260701144047.3786939-2-wangkefeng.wang@huawei.com/
> > >
> > > mm/mincore.c | 29 +++++++++++++++++------------
> > > 1 file changed, 17 insertions(+), 12 deletions(-)
> > >
> > > diff --git a/mm/mincore.c b/mm/mincore.c
> > > index c086836bc4bc..0fe50f8a7e62 100644
> > > --- a/mm/mincore.c
> > > +++ b/mm/mincore.c
> > > @@ -235,24 +235,22 @@ static const struct mm_walk_ops mincore_walk_ops = {
> > > .pmd_entry = mincore_pte_range,
> > > .pte_hole = mincore_unmapped_range,
> > > .hugetlb_entry = mincore_hugetlb,
> > > - .walk_lock = PGWALK_RDLOCK,
> > > + .walk_lock = PGWALK_VMA_RDLOCK_VERIFY,
> > > };
> > > /*
> > > * Do a chunk of "sys_mincore()". We've already checked
> > > - * all the arguments, we hold the mmap semaphore: we should
> > > + * all the arguments, we hold the VMA read lock: we should
> > > * just return the amount of info we're asked for.
> > > */
> > > -static long do_mincore(unsigned long addr, unsigned long pages, unsigned char *vec)
> > > +static long do_mincore(struct vm_area_struct *vma, unsigned long addr,
> > > + unsigned long pages, unsigned char *vec)
> > > {
> > > - struct vm_area_struct *vma;
> > > - unsigned long end;
> > > + unsigned long end = min(vma->vm_end, addr + (pages << PAGE_SHIFT));
> > > int err;
> > > - vma = vma_lookup(current->mm, addr);
> > > - if (!vma)
> > > - return -ENOMEM;
> > > - end = min(vma->vm_end, addr + (pages << PAGE_SHIFT));
> > > + vma_assert_locked(vma);
> >
> > I think Lorenzo asked whether we should do that. But the walk_page_vma() further
> > below would already verify that due to PGWALK_VMA_RDLOCK_VERIFY (see
> > process_vma_walk_lock) so not sure if that's really required here.
That happens after you do actions which require the lock.
>
> I misunderstood Lorenzo's intention, which has already been verified through
> the PGWALK_VMA_RDLOCK_VERIFY check. I don't think it has much value, but
> adding it doesn't do any harm.
...! I mean, the polite thing might be to ask? :)
It's not critical, since you literally take it immediately prior to calling
do_mincore().
But I felt it'd be a nice, self-contained, self-documenting way of
establishing the invariant given you just changed the function.
However you're also commenting that so it's not vital.
>
> >
> > AFAIKS, everything we do in mincore_pte_range() should be compatible with the
> > VMA lock, including the swap and pagecache handling.
It'd be pretty broken if the VMA lock provided less guarantees than the
mmap read lock in VMA-specific operations.
> >
> > Acked-by: David Hildenbrand (Arm) <david@kernel.org>
> >
>
--
Cheers, Lorenzo
next prev parent reply other threads:[~2026-09-16 11:20 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 4:31 [PATCH mm-new v3] mm: mincore: use per-vma lock during page table walk Kefeng Wang
2026-09-16 6:33 ` David Hildenbrand (Arm)
2026-09-16 6:50 ` Kefeng Wang
2026-09-16 11:19 ` Lorenzo Stoakes (ARM) [this message]
2026-09-16 12:30 ` Pedro Falcato
2026-09-16 12:40 ` Lorenzo Stoakes (ARM)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqp4Uvem1MWQOUZA@gremlin \
--to=ljs@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=david@kernel.org \
--cc=jannh@google.com \
--cc=liam@infradead.org \
--cc=linux-mm@kvack.org \
--cc=pfalcato@suse.de \
--cc=vbabka@kernel.org \
--cc=wangkefeng.wang@huawei.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox