From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: "David Hildenbrand (Arm)" <david@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>,
akpm@linux-foundation.org, liam@infradead.org,
vbabka@kernel.org, willy@infradead.org, jannh@google.com,
paulmck@kernel.org, pfalcato@suse.de, linux-mm@kvack.org,
linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org
Subject: Re: [PATCH v2 3/5] proc/task_mmu: remove special-casing of smap_gather_stats() start parameter
Date: Thu, 10 Sep 2026 17:27:11 +0100 [thread overview]
Message-ID: <aqLUqsoMDfVtjRuN@gremlin> (raw)
In-Reply-To: <8e643f51-a47f-428e-9443-678326288fa6@kernel.org>
On Wed, Sep 09, 2026 at 09:16:23PM +0200, David Hildenbrand (Arm) wrote:
> On 9/9/26 20:28, Suren Baghdasaryan wrote:
> > On Wed, Sep 9, 2026 at 10:16 AM David Hildenbrand (Arm)
> > <david@kernel.org> wrote:
> >>
> >> On 9/7/26 08:39, Suren Baghdasaryan wrote:
> >>> smap_gather_stats() interprets its start parameter to mean vma->vm_start
> >>> when it's set to 0. Eliminate this special interpretation and pass
> >>> vma->vm_start explicitly when needed.
> >>>
> >>> Since smap_gather_stats() operates within a single VMA, we can replace
> >>> walk_page_vma()/walk_page_range() calls with walk_page_range_vma()
> >>> which is simpler and also can be called while holding per-VMA lock.
> >>>
> >>> No functional change intended.
> >>>
> >>> Suggested by: Lorenzo Stoakes <ljs@kernel.org>
> >>> Signed-off-by: Suren Baghdasaryan <surenb@google.com>
> >>> ---
> >>> fs/proc/task_mmu.c | 40 ++++++++++++++++++++++------------------
> >>> 1 file changed, 22 insertions(+), 18 deletions(-)
> >>>
> >>> diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
> >>> index 9908ba32f180..3351decd1172 100644
> >>> --- a/fs/proc/task_mmu.c
> >>> +++ b/fs/proc/task_mmu.c
> >>> @@ -1246,20 +1246,27 @@ get_smaps_shmem_walk_ops(struct proc_maps_private *priv)
> >>> return &smaps_shmem_walk_vma_lock_ops;
> >>> }
> >>>
> >>> -/*
> >>> - * Gather mem stats from @vma with the indicated beginning
> >>> - * address @start, and keep them in @mss.
> >>> +/**
> >>> + * smap_gather_stats() - Gather mem stats from @vma.
> >>> + * @priv: proc maps private state.
> >>> + * @vma: The VMA to gather stats for.
> >>> + * @mss: The accumulated stats.
> >>> + * @start: The address from which to start.
> >>> *
> >>> - * Use vm_start of @vma as the beginning address if @start is 0.
> >>> + * This gathers stats for the whole of the VMA unless the lock was dropped
> >>> + * and VMA grew or got merged and we found it again, in which case we only
> >>> + * gather stats for the remainder of the VMA range.
> >>> */
> >>> static void smap_gather_stats(struct proc_maps_private *priv,
> >>> struct vm_area_struct *vma,
> >>> - struct mem_size_stats *mss, unsigned long start)
> >>> + struct mem_size_stats *mss,
> >>> + unsigned long start)
> >>> {
> >>> const struct mm_walk_ops *ops = get_smaps_walk_ops(priv);
> >>> + const bool is_partial = start > vma->vm_start;
> >>>
> >>> /* Invalid start */
> >>> - if (start >= vma->vm_end)
> >>> + if (start < vma->vm_start || start >= vma->vm_end)
> >>> return;
> >>>
> >>> if (vma == get_gate_vma(priv->lock_ctx.mm))
> >>> @@ -1279,20 +1286,17 @@ static void smap_gather_stats(struct proc_maps_private *priv,
> >>> * Unless we know that the shmem object (or the part mapped by
> >>> * our VMA) has no swapped out pages at all.
> >>> */
> >>> - unsigned long shmem_swapped = shmem_swap_usage(vma);
> >>> + const unsigned long shmem_swapped = shmem_swap_usage(vma);
> >>> + const bool shared_or_ro = vma_test(vma, VMA_SHARED_BIT) ||
> >>> + !vma_test(vma, VMA_WRITE_BIT);
> >>>
> >>> - if (!start && (!shmem_swapped || (vma->vm_flags & VM_SHARED) ||
> >>> - !(vma->vm_flags & VM_WRITE))) {
> >>> + if (!is_partial && (!shmem_swapped || shared_or_ro))
> >>> mss->swap += shmem_swapped;
> >>> - } else {
> >>> + else
> >>> ops = get_smaps_shmem_walk_ops(priv);
> >>> - }
> >>
> >> Horrible, horrible code, really. But not your fault :)
> >>
> >> I think we can just make the shared_or_ro less odd by just checking for cow
> >> mappings (as described in the comment).
> >>
> >> const bool is_cow = vma_is_cow_mapping(vma);
> >>
> >> ...
> >>
> >> if (is_partial || (shmem_swapped && is_cow))
> >> ops = get_smaps_shmem_walk_ops(priv);
> >> else
> >> mss->swap += shmem_swapped;
> >>
> >> That's almost in a form that I could understand what's happening.
> >
> > Hmm. So, are you saying that !is_cow always implies shared_or_ro? Or
> > maybe you are stating that vma_is_cow_mapping() was the actual intent
> > here?
>
> So the comment says:
>
> "For private writable mappings, we might have COW pages that .."
>
> Which translates to:
>
> private writable == vma_is_cow_mapping()
>
> >
> > shared_or_ro = VMA_SHARED_BIT || !VMA_WRITE_BIT
> >
> > is_cow = !VMA_SHARED_BIT && VMA_MAYWRITE_BIT
> > !is_cow = VMA_SHARED_BIT || !VMA_MAYWRITE_BIT
> >
> > so, !is_cow would impy shared_or_ro only if !VMA_MAYWRITE_BIT always
> > implies !VMA_WRITE_BIT. But I think it's possible to have a VMA that
> > has VMA_WRITE_BIT but not VMA_MAYWRITE_BIT, right?
>
> VMA_WRITE should imply VMA_MAYWRITE
Yup it's illegal to set VMA_WRITE_BIT without VMA_MAYWRITE_BIT.
>
> (in sanitize_fault_flags() we even disallow write faults entirely if VM_MAYWRITE
> is missing)
>
> For example, a
> > driver can create such a VMA to allow writing to the VMA but to lock
> > its content once mprotect(PROT_READ) gets called.
I'm not even sure how a driver would achieve that but drivers in general are not
permitted to alter VMA flags after map time.
>
> I don't think that would be valid for a driver to do. But it wouldn't matter
> here because
>
> shmem_mapping(vma->vm_file->f_mapping)
Yup :)
>
>
> I think we could simplify the comment as well to:
>
> diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
> index e671b4fd8dedd..4b7e7089cafa7 100644
> --- a/fs/proc/task_mmu.c
> +++ b/fs/proc/task_mmu.c
> @@ -1303,12 +1303,10 @@ static void smap_gather_stats(struct proc_maps_private
> *priv,
>
> if (vma->vm_file && shmem_mapping(vma->vm_file->f_mapping)) {
> /*
> - * For shared or readonly shmem mappings we know that all
> - * swapped out pages belong to the shmem object, and we can
> - * obtain the swap value much more efficiently. For private
> - * writable mappings, we might have COW pages that are
> - * not affected by the parent swapped out pages of the shmem
> - * object, so we have to distinguish them during the page walk.
> + * In CoW mappings, we might have anon folios that are
> + * independent of the shmem object. So fallback to the less
> + * efficient mechanism in such mappings.
> + *
Maybe tweak to 'CoW mappings might map anon folios that do not belong to shmem,
so perform a less efficient page table walk in this situation' or something like
that?
> * Unless we know that the shmem object (or the part mapped by
> * our VMA) has no swapped out pages at all.
> */
>
> --
> Cheers,
>
> David
--
Cheers, Lorenzo
next prev parent reply other threads:[~2026-09-10 16:27 UTC|newest]
Thread overview: 43+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-07 6:39 [PATCH v2 0/5] read proc/pid/smaps_rollup under per-vma lock Suren Baghdasaryan
2026-09-07 6:39 ` [PATCH v2 1/5] proc/task_mmu: remove unnecessary helpers Suren Baghdasaryan
2026-09-07 16:44 ` Usama Arif
2026-09-08 17:58 ` Liam R. Howlett
2026-09-09 17:06 ` David Hildenbrand (Arm)
2026-09-09 17:13 ` Suren Baghdasaryan
2026-09-09 17:17 ` David Hildenbrand (Arm)
2026-09-09 18:29 ` Suren Baghdasaryan
2026-09-10 15:43 ` Lorenzo Stoakes (ARM)
2026-09-07 6:39 ` [PATCH v2 2/5] proc/task_mmu: remove unnecessary inlines in function definitions Suren Baghdasaryan
2026-09-07 16:49 ` Usama Arif
2026-09-10 15:35 ` Suren Baghdasaryan
2026-09-08 18:01 ` Liam R. Howlett
2026-09-09 17:07 ` David Hildenbrand (Arm)
2026-09-09 17:15 ` Suren Baghdasaryan
2026-09-10 15:55 ` Lorenzo Stoakes (ARM)
2026-09-07 6:39 ` [PATCH v2 3/5] proc/task_mmu: remove special-casing of smap_gather_stats() start parameter Suren Baghdasaryan
2026-09-08 18:07 ` Liam R. Howlett
2026-09-09 17:16 ` David Hildenbrand (Arm)
2026-09-09 18:28 ` Suren Baghdasaryan
2026-09-09 19:16 ` David Hildenbrand (Arm)
2026-09-09 21:51 ` Suren Baghdasaryan
2026-09-10 7:41 ` David Hildenbrand (Arm)
2026-09-10 15:45 ` Suren Baghdasaryan
2026-09-10 16:01 ` David Hildenbrand (Arm)
2026-09-10 16:09 ` Suren Baghdasaryan
2026-09-10 16:21 ` Suren Baghdasaryan
2026-09-10 16:33 ` David Hildenbrand (Arm)
2026-09-10 17:02 ` Suren Baghdasaryan
2026-09-10 16:27 ` Lorenzo Stoakes (ARM) [this message]
2026-09-10 23:30 ` Suren Baghdasaryan
2026-09-07 6:39 ` [PATCH v2 4/5] proc/task_mmu: read proc/pid/smaps_rollup under per-vma lock Suren Baghdasaryan
2026-09-08 18:17 ` Liam R. Howlett
2026-09-09 14:25 ` Usama Arif
2026-09-09 16:13 ` Suren Baghdasaryan
2026-09-09 17:23 ` David Hildenbrand (Arm)
2026-09-09 17:58 ` Suren Baghdasaryan
2026-09-10 7:44 ` David Hildenbrand (Arm)
2026-09-10 21:20 ` Suren Baghdasaryan
2026-09-07 6:39 ` [PATCH v2 5/5] selftests/proc: add /proc/pid/smaps_rollup tearing tests Suren Baghdasaryan
2026-09-08 18:18 ` Liam R. Howlett
2026-09-08 16:04 ` [PATCH v2 0/5] read proc/pid/smaps_rollup under per-vma lock Xueyuan Chen
2026-09-08 16:08 ` Suren Baghdasaryan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aqLUqsoMDfVtjRuN@gremlin \
--to=ljs@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=david@kernel.org \
--cc=jannh@google.com \
--cc=liam@infradead.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=paulmck@kernel.org \
--cc=pfalcato@suse.de \
--cc=surenb@google.com \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.