From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Suren Baghdasaryan <surenb@google.com>
Cc: akpm@linux-foundation.org, liam@infradead.org, ljs@kernel.org,
vbabka@kernel.org, willy@infradead.org, jannh@google.com,
paulmck@kernel.org, pfalcato@suse.de, linux-mm@kvack.org,
linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org
Subject: Re: [PATCH v2 3/5] proc/task_mmu: remove special-casing of smap_gather_stats() start parameter
Date: Wed, 9 Sep 2026 21:16:23 +0200 [thread overview]
Message-ID: <8e643f51-a47f-428e-9443-678326288fa6@kernel.org> (raw)
In-Reply-To: <CAJuCfpFL_g7tELjT8p8ExM3tX_v2gNDM9rRKwwefPV_KW=fHfA@mail.gmail.com>
On 9/9/26 20:28, Suren Baghdasaryan wrote:
> On Wed, Sep 9, 2026 at 10:16 AM David Hildenbrand (Arm)
> <david@kernel.org> wrote:
>>
>> On 9/7/26 08:39, Suren Baghdasaryan wrote:
>>> smap_gather_stats() interprets its start parameter to mean vma->vm_start
>>> when it's set to 0. Eliminate this special interpretation and pass
>>> vma->vm_start explicitly when needed.
>>>
>>> Since smap_gather_stats() operates within a single VMA, we can replace
>>> walk_page_vma()/walk_page_range() calls with walk_page_range_vma()
>>> which is simpler and also can be called while holding per-VMA lock.
>>>
>>> No functional change intended.
>>>
>>> Suggested by: Lorenzo Stoakes <ljs@kernel.org>
>>> Signed-off-by: Suren Baghdasaryan <surenb@google.com>
>>> ---
>>> fs/proc/task_mmu.c | 40 ++++++++++++++++++++++------------------
>>> 1 file changed, 22 insertions(+), 18 deletions(-)
>>>
>>> diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
>>> index 9908ba32f180..3351decd1172 100644
>>> --- a/fs/proc/task_mmu.c
>>> +++ b/fs/proc/task_mmu.c
>>> @@ -1246,20 +1246,27 @@ get_smaps_shmem_walk_ops(struct proc_maps_private *priv)
>>> return &smaps_shmem_walk_vma_lock_ops;
>>> }
>>>
>>> -/*
>>> - * Gather mem stats from @vma with the indicated beginning
>>> - * address @start, and keep them in @mss.
>>> +/**
>>> + * smap_gather_stats() - Gather mem stats from @vma.
>>> + * @priv: proc maps private state.
>>> + * @vma: The VMA to gather stats for.
>>> + * @mss: The accumulated stats.
>>> + * @start: The address from which to start.
>>> *
>>> - * Use vm_start of @vma as the beginning address if @start is 0.
>>> + * This gathers stats for the whole of the VMA unless the lock was dropped
>>> + * and VMA grew or got merged and we found it again, in which case we only
>>> + * gather stats for the remainder of the VMA range.
>>> */
>>> static void smap_gather_stats(struct proc_maps_private *priv,
>>> struct vm_area_struct *vma,
>>> - struct mem_size_stats *mss, unsigned long start)
>>> + struct mem_size_stats *mss,
>>> + unsigned long start)
>>> {
>>> const struct mm_walk_ops *ops = get_smaps_walk_ops(priv);
>>> + const bool is_partial = start > vma->vm_start;
>>>
>>> /* Invalid start */
>>> - if (start >= vma->vm_end)
>>> + if (start < vma->vm_start || start >= vma->vm_end)
>>> return;
>>>
>>> if (vma == get_gate_vma(priv->lock_ctx.mm))
>>> @@ -1279,20 +1286,17 @@ static void smap_gather_stats(struct proc_maps_private *priv,
>>> * Unless we know that the shmem object (or the part mapped by
>>> * our VMA) has no swapped out pages at all.
>>> */
>>> - unsigned long shmem_swapped = shmem_swap_usage(vma);
>>> + const unsigned long shmem_swapped = shmem_swap_usage(vma);
>>> + const bool shared_or_ro = vma_test(vma, VMA_SHARED_BIT) ||
>>> + !vma_test(vma, VMA_WRITE_BIT);
>>>
>>> - if (!start && (!shmem_swapped || (vma->vm_flags & VM_SHARED) ||
>>> - !(vma->vm_flags & VM_WRITE))) {
>>> + if (!is_partial && (!shmem_swapped || shared_or_ro))
>>> mss->swap += shmem_swapped;
>>> - } else {
>>> + else
>>> ops = get_smaps_shmem_walk_ops(priv);
>>> - }
>>
>> Horrible, horrible code, really. But not your fault :)
>>
>> I think we can just make the shared_or_ro less odd by just checking for cow
>> mappings (as described in the comment).
>>
>> const bool is_cow = vma_is_cow_mapping(vma);
>>
>> ...
>>
>> if (is_partial || (shmem_swapped && is_cow))
>> ops = get_smaps_shmem_walk_ops(priv);
>> else
>> mss->swap += shmem_swapped;
>>
>> That's almost in a form that I could understand what's happening.
>
> Hmm. So, are you saying that !is_cow always implies shared_or_ro? Or
> maybe you are stating that vma_is_cow_mapping() was the actual intent
> here?
So the comment says:
"For private writable mappings, we might have COW pages that .."
Which translates to:
private writable == vma_is_cow_mapping()
>
> shared_or_ro = VMA_SHARED_BIT || !VMA_WRITE_BIT
>
> is_cow = !VMA_SHARED_BIT && VMA_MAYWRITE_BIT
> !is_cow = VMA_SHARED_BIT || !VMA_MAYWRITE_BIT
>
> so, !is_cow would impy shared_or_ro only if !VMA_MAYWRITE_BIT always
> implies !VMA_WRITE_BIT. But I think it's possible to have a VMA that
> has VMA_WRITE_BIT but not VMA_MAYWRITE_BIT, right?
VMA_WRITE should imply VMA_MAYWRITE
(in sanitize_fault_flags() we even disallow write faults entirely if VM_MAYWRITE
is missing)
For example, a
> driver can create such a VMA to allow writing to the VMA but to lock
> its content once mprotect(PROT_READ) gets called.
I don't think that would be valid for a driver to do. But it wouldn't matter
here because
shmem_mapping(vma->vm_file->f_mapping)
I think we could simplify the comment as well to:
diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
index e671b4fd8dedd..4b7e7089cafa7 100644
--- a/fs/proc/task_mmu.c
+++ b/fs/proc/task_mmu.c
@@ -1303,12 +1303,10 @@ static void smap_gather_stats(struct proc_maps_private
*priv,
if (vma->vm_file && shmem_mapping(vma->vm_file->f_mapping)) {
/*
- * For shared or readonly shmem mappings we know that all
- * swapped out pages belong to the shmem object, and we can
- * obtain the swap value much more efficiently. For private
- * writable mappings, we might have COW pages that are
- * not affected by the parent swapped out pages of the shmem
- * object, so we have to distinguish them during the page walk.
+ * In CoW mappings, we might have anon folios that are
+ * independent of the shmem object. So fallback to the less
+ * efficient mechanism in such mappings.
+ *
* Unless we know that the shmem object (or the part mapped by
* our VMA) has no swapped out pages at all.
*/
--
Cheers,
David
next prev parent reply other threads:[~2026-09-09 19:16 UTC|newest]
Thread overview: 31+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-07 6:39 [PATCH v2 0/5] read proc/pid/smaps_rollup under per-vma lock Suren Baghdasaryan
2026-09-07 6:39 ` [PATCH v2 1/5] proc/task_mmu: remove unnecessary helpers Suren Baghdasaryan
2026-09-07 16:44 ` Usama Arif
2026-09-08 17:58 ` Liam R. Howlett
2026-09-09 17:06 ` David Hildenbrand (Arm)
2026-09-09 17:13 ` Suren Baghdasaryan
2026-09-09 17:17 ` David Hildenbrand (Arm)
2026-09-09 18:29 ` Suren Baghdasaryan
2026-09-07 6:39 ` [PATCH v2 2/5] proc/task_mmu: remove unnecessary inlines in function definitions Suren Baghdasaryan
2026-09-07 16:49 ` Usama Arif
2026-09-08 18:01 ` Liam R. Howlett
2026-09-09 17:07 ` David Hildenbrand (Arm)
2026-09-09 17:15 ` Suren Baghdasaryan
2026-09-07 6:39 ` [PATCH v2 3/5] proc/task_mmu: remove special-casing of smap_gather_stats() start parameter Suren Baghdasaryan
2026-09-08 18:07 ` Liam R. Howlett
2026-09-09 17:16 ` David Hildenbrand (Arm)
2026-09-09 18:28 ` Suren Baghdasaryan
2026-09-09 19:16 ` David Hildenbrand (Arm) [this message]
2026-09-09 21:51 ` Suren Baghdasaryan
2026-09-10 7:41 ` David Hildenbrand (Arm)
2026-09-07 6:39 ` [PATCH v2 4/5] proc/task_mmu: read proc/pid/smaps_rollup under per-vma lock Suren Baghdasaryan
2026-09-08 18:17 ` Liam R. Howlett
2026-09-09 14:25 ` Usama Arif
2026-09-09 16:13 ` Suren Baghdasaryan
2026-09-09 17:23 ` David Hildenbrand (Arm)
2026-09-09 17:58 ` Suren Baghdasaryan
2026-09-10 7:44 ` David Hildenbrand (Arm)
2026-09-07 6:39 ` [PATCH v2 5/5] selftests/proc: add /proc/pid/smaps_rollup tearing tests Suren Baghdasaryan
2026-09-08 18:18 ` Liam R. Howlett
2026-09-08 16:04 ` [PATCH v2 0/5] read proc/pid/smaps_rollup under per-vma lock Xueyuan Chen
2026-09-08 16:08 ` Suren Baghdasaryan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=8e643f51-a47f-428e-9443-678326288fa6@kernel.org \
--to=david@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=jannh@google.com \
--cc=liam@infradead.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=paulmck@kernel.org \
--cc=pfalcato@suse.de \
--cc=surenb@google.com \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox