Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: "David Hildenbrand (Arm)" <david@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>,
	akpm@linux-foundation.org,  liam@infradead.org,
	vbabka@kernel.org, willy@infradead.org, jannh@google.com,
	 paulmck@kernel.org, pfalcato@suse.de, linux-mm@kvack.org,
	 linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org
Subject: Re: [PATCH v2 3/5] proc/task_mmu: remove special-casing of smap_gather_stats() start parameter
Date: Thu, 10 Sep 2026 17:27:11 +0100	[thread overview]
Message-ID: <aqLUqsoMDfVtjRuN@gremlin> (raw)
In-Reply-To: <8e643f51-a47f-428e-9443-678326288fa6@kernel.org>

On Wed, Sep 09, 2026 at 09:16:23PM +0200, David Hildenbrand (Arm) wrote:
> On 9/9/26 20:28, Suren Baghdasaryan wrote:
> > On Wed, Sep 9, 2026 at 10:16 AM David Hildenbrand (Arm)
> > <david@kernel.org> wrote:
> >>
> >> On 9/7/26 08:39, Suren Baghdasaryan wrote:
> >>> smap_gather_stats() interprets its start parameter to mean vma->vm_start
> >>> when it's set to 0. Eliminate this special interpretation and pass
> >>> vma->vm_start explicitly when needed.
> >>>
> >>> Since smap_gather_stats() operates within a single VMA, we can replace
> >>> walk_page_vma()/walk_page_range() calls with walk_page_range_vma()
> >>> which is simpler and also can be called while holding per-VMA lock.
> >>>
> >>> No functional change intended.
> >>>
> >>> Suggested by: Lorenzo Stoakes <ljs@kernel.org>
> >>> Signed-off-by: Suren Baghdasaryan <surenb@google.com>
> >>> ---
> >>>  fs/proc/task_mmu.c | 40 ++++++++++++++++++++++------------------
> >>>  1 file changed, 22 insertions(+), 18 deletions(-)
> >>>
> >>> diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
> >>> index 9908ba32f180..3351decd1172 100644
> >>> --- a/fs/proc/task_mmu.c
> >>> +++ b/fs/proc/task_mmu.c
> >>> @@ -1246,20 +1246,27 @@ get_smaps_shmem_walk_ops(struct proc_maps_private *priv)
> >>>       return &smaps_shmem_walk_vma_lock_ops;
> >>>  }
> >>>
> >>> -/*
> >>> - * Gather mem stats from @vma with the indicated beginning
> >>> - * address @start, and keep them in @mss.
> >>> +/**
> >>> + * smap_gather_stats() - Gather mem stats from @vma.
> >>> + * @priv: proc maps private state.
> >>> + * @vma: The VMA to gather stats for.
> >>> + * @mss: The accumulated stats.
> >>> + * @start: The address from which to start.
> >>>   *
> >>> - * Use vm_start of @vma as the beginning address if @start is 0.
> >>> + * This gathers stats for the whole of the VMA unless the lock was dropped
> >>> + * and VMA grew or got merged and we found it again, in which case we only
> >>> + * gather stats for the remainder of the VMA range.
> >>>   */
> >>>  static void smap_gather_stats(struct proc_maps_private *priv,
> >>>                             struct vm_area_struct *vma,
> >>> -                           struct mem_size_stats *mss, unsigned long start)
> >>> +                           struct mem_size_stats *mss,
> >>> +                           unsigned long start)
> >>>  {
> >>>       const struct mm_walk_ops *ops = get_smaps_walk_ops(priv);
> >>> +     const bool is_partial = start > vma->vm_start;
> >>>
> >>>       /* Invalid start */
> >>> -     if (start >= vma->vm_end)
> >>> +     if (start < vma->vm_start || start >= vma->vm_end)
> >>>               return;
> >>>
> >>>       if (vma == get_gate_vma(priv->lock_ctx.mm))
> >>> @@ -1279,20 +1286,17 @@ static void smap_gather_stats(struct proc_maps_private *priv,
> >>>                * Unless we know that the shmem object (or the part mapped by
> >>>                * our VMA) has no swapped out pages at all.
> >>>                */
> >>> -             unsigned long shmem_swapped = shmem_swap_usage(vma);
> >>> +             const unsigned long shmem_swapped = shmem_swap_usage(vma);
> >>> +             const bool shared_or_ro = vma_test(vma, VMA_SHARED_BIT) ||
> >>> +                                       !vma_test(vma, VMA_WRITE_BIT);
> >>>
> >>> -             if (!start && (!shmem_swapped || (vma->vm_flags & VM_SHARED) ||
> >>> -                                     !(vma->vm_flags & VM_WRITE))) {
> >>> +             if (!is_partial && (!shmem_swapped || shared_or_ro))
> >>>                       mss->swap += shmem_swapped;
> >>> -             } else {
> >>> +             else
> >>>                       ops = get_smaps_shmem_walk_ops(priv);
> >>> -             }
> >>
> >> Horrible, horrible code, really. But not your fault :)
> >>
> >> I think we can just make the shared_or_ro less odd by just checking for cow
> >> mappings (as described in the comment).
> >>
> >>         const bool is_cow = vma_is_cow_mapping(vma);
> >>
> >> ...
> >>
> >>         if (is_partial || (shmem_swapped && is_cow))
> >>                 ops = get_smaps_shmem_walk_ops(priv);
> >>         else
> >>                 mss->swap += shmem_swapped;
> >>
> >> That's almost in a form that I could understand what's happening.
> >
> > Hmm. So, are you saying that !is_cow always implies shared_or_ro? Or
> > maybe you are stating that vma_is_cow_mapping() was the actual intent
> > here?
>
> So the comment says:
>
> "For private writable mappings, we might have COW pages that  .."
>
> Which translates to:
>
> 	private writable == vma_is_cow_mapping()
>
> >
> > shared_or_ro = VMA_SHARED_BIT || !VMA_WRITE_BIT
> >
> > is_cow = !VMA_SHARED_BIT && VMA_MAYWRITE_BIT
> > !is_cow = VMA_SHARED_BIT || !VMA_MAYWRITE_BIT
> >
> > so, !is_cow would impy shared_or_ro only if !VMA_MAYWRITE_BIT always
> > implies !VMA_WRITE_BIT. But I think it's possible to have a VMA that
> > has VMA_WRITE_BIT but not VMA_MAYWRITE_BIT, right?
>
> VMA_WRITE should imply VMA_MAYWRITE

Yup it's illegal to set VMA_WRITE_BIT without VMA_MAYWRITE_BIT.

>
> (in sanitize_fault_flags() we even disallow write faults entirely if VM_MAYWRITE
> is missing)
>
> For example, a
> > driver can create such a VMA to allow writing to the VMA but to lock
> > its content once mprotect(PROT_READ) gets called.

I'm not even sure how a driver would achieve that but drivers in general are not
permitted to alter VMA flags after map time.

>
> I don't think that would be valid for a driver to do. But it wouldn't matter
> here because
>
> 	shmem_mapping(vma->vm_file->f_mapping)

Yup :)

>
>
> I think we could simplify the comment as well to:
>
> diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
> index e671b4fd8dedd..4b7e7089cafa7 100644
> --- a/fs/proc/task_mmu.c
> +++ b/fs/proc/task_mmu.c
> @@ -1303,12 +1303,10 @@ static void smap_gather_stats(struct proc_maps_private
> *priv,
>
>         if (vma->vm_file && shmem_mapping(vma->vm_file->f_mapping)) {
>                 /*
> -                * For shared or readonly shmem mappings we know that all
> -                * swapped out pages belong to the shmem object, and we can
> -                * obtain the swap value much more efficiently. For private
> -                * writable mappings, we might have COW pages that are
> -                * not affected by the parent swapped out pages of the shmem
> -                * object, so we have to distinguish them during the page walk.
> +                * In CoW mappings, we might have anon folios that are
> +                * independent of the shmem object. So fallback to the less
> +                * efficient mechanism in such mappings.
> +                *

Maybe tweak to 'CoW mappings might map anon folios that do not belong to shmem,
so perform a less efficient page table walk in this situation' or something like
that?


>                  * Unless we know that the shmem object (or the part mapped by
>                  * our VMA) has no swapped out pages at all.
>                  */



>
> --
> Cheers,
>
> David

--
Cheers, Lorenzo


  parent reply	other threads:[~2026-09-10 16:27 UTC|newest]

Thread overview: 43+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-07  6:39 [PATCH v2 0/5] read proc/pid/smaps_rollup under per-vma lock Suren Baghdasaryan
2026-09-07  6:39 ` [PATCH v2 1/5] proc/task_mmu: remove unnecessary helpers Suren Baghdasaryan
2026-09-07 16:44   ` Usama Arif
2026-09-08 17:58   ` Liam R. Howlett
2026-09-09 17:06   ` David Hildenbrand (Arm)
2026-09-09 17:13     ` Suren Baghdasaryan
2026-09-09 17:17       ` David Hildenbrand (Arm)
2026-09-09 18:29         ` Suren Baghdasaryan
2026-09-10 15:43   ` Lorenzo Stoakes (ARM)
2026-09-07  6:39 ` [PATCH v2 2/5] proc/task_mmu: remove unnecessary inlines in function definitions Suren Baghdasaryan
2026-09-07 16:49   ` Usama Arif
2026-09-10 15:35     ` Suren Baghdasaryan
2026-09-08 18:01   ` Liam R. Howlett
2026-09-09 17:07   ` David Hildenbrand (Arm)
2026-09-09 17:15     ` Suren Baghdasaryan
2026-09-10 15:55   ` Lorenzo Stoakes (ARM)
2026-09-07  6:39 ` [PATCH v2 3/5] proc/task_mmu: remove special-casing of smap_gather_stats() start parameter Suren Baghdasaryan
2026-09-08 18:07   ` Liam R. Howlett
2026-09-09 17:16   ` David Hildenbrand (Arm)
2026-09-09 18:28     ` Suren Baghdasaryan
2026-09-09 19:16       ` David Hildenbrand (Arm)
2026-09-09 21:51         ` Suren Baghdasaryan
2026-09-10  7:41           ` David Hildenbrand (Arm)
2026-09-10 15:45             ` Suren Baghdasaryan
2026-09-10 16:01               ` David Hildenbrand (Arm)
2026-09-10 16:09                 ` Suren Baghdasaryan
2026-09-10 16:21                   ` Suren Baghdasaryan
2026-09-10 16:33                     ` David Hildenbrand (Arm)
2026-09-10 17:02                       ` Suren Baghdasaryan
2026-09-10 16:27         ` Lorenzo Stoakes (ARM) [this message]
2026-09-10 23:30           ` Suren Baghdasaryan
2026-09-07  6:39 ` [PATCH v2 4/5] proc/task_mmu: read proc/pid/smaps_rollup under per-vma lock Suren Baghdasaryan
2026-09-08 18:17   ` Liam R. Howlett
2026-09-09 14:25   ` Usama Arif
2026-09-09 16:13     ` Suren Baghdasaryan
2026-09-09 17:23   ` David Hildenbrand (Arm)
2026-09-09 17:58     ` Suren Baghdasaryan
2026-09-10  7:44       ` David Hildenbrand (Arm)
2026-09-10 21:20         ` Suren Baghdasaryan
2026-09-07  6:39 ` [PATCH v2 5/5] selftests/proc: add /proc/pid/smaps_rollup tearing tests Suren Baghdasaryan
2026-09-08 18:18   ` Liam R. Howlett
2026-09-08 16:04 ` [PATCH v2 0/5] read proc/pid/smaps_rollup under per-vma lock Xueyuan Chen
2026-09-08 16:08   ` Suren Baghdasaryan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aqLUqsoMDfVtjRuN@gremlin \
    --to=ljs@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=david@kernel.org \
    --cc=jannh@google.com \
    --cc=liam@infradead.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=paulmck@kernel.org \
    --cc=pfalcato@suse.de \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox