Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Usama Arif <usama.arif@linux.dev>
To: Suren Baghdasaryan <surenb@google.com>
Cc: Usama Arif <usama.arif@linux.dev>,
	akpm@linux-foundation.org, liam@infradead.org, ljs@kernel.org,
	vbabka@kernel.org, david@redhat.com, willy@infradead.org,
	jannh@google.com, paulmck@kernel.org, pfalcato@suse.de,
	linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	linux-fsdevel@vger.kernel.org
Subject: Re: [PATCH v2 4/5] proc/task_mmu: read proc/pid/smaps_rollup under per-vma lock
Date: Wed,  9 Sep 2026 07:25:26 -0700	[thread overview]
Message-ID: <20260909142527.1601051-1-usama.arif@linux.dev> (raw)
In-Reply-To: <20260907063918.3432401-5-surenb@google.com>

On Sun,  6 Sep 2026 23:39:17 -0700 Suren Baghdasaryan <surenb@google.com> wrote:

> proc/pid/smaps_rollup can be read using the combination of RCU and
> VMA read locks, similar to proc/pid/{maps|smaps|numa_maps}. RCU is
> required to safely traverse the VMA tree and VMA lock stabilizes the
> VMA being processed and the pagetable walk.
> Note that we have to keep the logic to drop mmap_lock on contention
> because even when using per-VMA locks we might have to fall back to
> holding the mmap_lock.
> 
> Running Paul's contention benchmark [1] shows considerable improvement
> both in median and in the worst case latencies:
> 
> Execution command: run-proc-vs-map.sh --nsamples 20 --rawdata -- \
> --busyduration 2 --procfile smaps_rollup
> 
> Baseline:
>    Median   Minimum   Maximum
>     0.174     0.161     2.553
>     0.174     0.164     2.663
>     0.174     0.165     2.664
>     0.174     0.166     2.679
>     0.174     0.167     2.691
>     0.174     0.168     2.704
>     0.174     0.169     2.729
>     0.174     0.172     2.741
>     0.174     0.174     2.745
>     0.174     0.174     2.755
>     0.174     0.175     2.790
>     0.174     0.177     2.809
>     0.174     0.179     3.096
>     0.174     0.183     3.144
>     0.174     0.184     3.158
>     0.174     0.185     3.175
>     0.174     0.185     4.568
>     0.174     0.198     4.821
>     0.174     0.214     5.143
>     0.174     0.251     5.220
> 
> Patched:
>    Median   Minimum   Maximum
>     0.007     0.007     1.952
>     0.007     0.007     1.955
>     0.007     0.007     1.955
>     0.007     0.007     1.955
>     0.007     0.007     1.957
>     0.007     0.007     1.969
>     0.007     0.007     2.065
>     0.007     0.007     2.075
>     0.007     0.007     2.146
>     0.007     0.007     2.195
>     0.007     0.007     2.223
>     0.007     0.007     2.259
>     0.007     0.007     2.488
>     0.007     0.007     2.562
>     0.007     0.007     2.599
>     0.007     0.007     2.697
>     0.007     0.007     3.030
>     0.007     0.007     3.075
>     0.007     0.007     3.145
>     0.007     0.007     3.225
> 
> Remove now unused lock_ctx_mm() and move unlock_ctx_vma() next to
> unlock_ctx_mm() as they are logically related.
> 
> Remove a long comment about 4 cases that we handle when dropping the
> mmap lock in the middle of VMA walk due to contention. The first 3
> cases explained there are handled naturally and only case 4 needs to
> be handled in a special way, which is done in smap_gather_stats() by
> gathering stats from the portion of the VMA that has not yet been
> processed.
> For posterity, moving this comment here:
> 
> After dropping the lock, there are four cases to
> consider. See the following example for explanation.
> 
>   +------+------+-----------+
>   | VMA1 | VMA2 | VMA3      |
>   +------+------+-----------+
>   |      |      |           |
>  4k     8k     16k         400k
> 
> Suppose we drop the lock after reading VMA2 due to
> contention, then we get:
> 
> 	last_vma_end = 16k
> 
> 1) VMA2 is freed, but VMA3 exists:
> 
>    vma_next(vmi) will return VMA3.
>    In this case, just continue from VMA3.
> 
> 2) VMA2 still exists:
> 
>    vma_next(vmi) will return VMA3.
>    In this case, just continue from VMA3.
> 
> 3) No more VMAs can be found:
> 
>    vma_next(vmi) will return NULL.
>    No more things to do, just break.
> 
> 4) (last_vma_end - 1) is the middle of a vma (VMA'):
> 
>    vma_next(vmi) will return VMA' whose range
>    contains last_vma_end.
>    Iterate VMA' from last_vma_end.
> 
> [1] https://github.com/paulmckrcu/proc-mmap_sem-test
> 
> Signed-off-by: Suren Baghdasaryan <surenb@google.com>
> ---
>  fs/proc/task_mmu.c | 153 ++++++++++++++++++---------------------------
>  1 file changed, 60 insertions(+), 93 deletions(-)
> 
> diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
> index 3351decd1172..641a155b0c61 100644
> --- a/fs/proc/task_mmu.c
> +++ b/fs/proc/task_mmu.c
> @@ -130,28 +130,12 @@ static void release_task_mempolicy(struct proc_maps_private *priv)
>  }
>  #endif
>  
> -static int lock_ctx_mm(struct proc_maps_locking_ctx *lock_ctx)
> -{
> -	int ret = mmap_read_lock_killable(lock_ctx->mm);
> -
> -	if (!ret)
> -		lock_ctx->mmap_locked = true;
> -
> -	return ret;
> -}
> -
>  static void unlock_ctx_mm(struct proc_maps_locking_ctx *lock_ctx)
>  {
>  	mmap_read_unlock(lock_ctx->mm);
>  	lock_ctx->mmap_locked = false;
>  }
>  
> -static void reset_lock_ctx(struct proc_maps_locking_ctx *lock_ctx)
> -{
> -	lock_ctx->locked_vma = NULL;
> -	lock_ctx->mmap_locked = false;
> -}
> -
>  static void unlock_ctx_vma(struct proc_maps_locking_ctx *lock_ctx)
>  {
>  	if (lock_ctx->locked_vma) {
> @@ -160,6 +144,12 @@ static void unlock_ctx_vma(struct proc_maps_locking_ctx *lock_ctx)
>  	}
>  }
>  
> +static void reset_lock_ctx(struct proc_maps_locking_ctx *lock_ctx)
> +{
> +	lock_ctx->locked_vma = NULL;
> +	lock_ctx->mmap_locked = false;
> +}
> +
>  static struct vm_area_struct *get_next_vma(struct proc_maps_private *priv,
>  					   loff_t last_pos)
>  {
> @@ -1376,12 +1366,14 @@ static int show_smap(struct seq_file *m, void *v)
>  static int show_smaps_rollup(struct seq_file *m, void *v)
>  {
>  	struct proc_maps_private *priv = m->private;
> +	struct proc_maps_locking_ctx *lock_ctx = &priv->lock_ctx;
> +	struct mm_struct *mm = lock_ctx->mm;
>  	struct mem_size_stats mss = {};
> -	struct mm_struct *mm = priv->lock_ctx.mm;
> +	unsigned long last_vma_end = 0;
> +	unsigned long vma_start = 0;
>  	struct vm_area_struct *vma;
> -	unsigned long vma_start = 0, last_vma_end = 0;
> +	loff_t pos = 0;
>  	int ret = 0;
> -	VMA_ITERATOR(vmi, mm, 0);
>  
>  	priv->task = get_proc_task(priv->inode);
>  	if (!priv->task)
> @@ -1392,89 +1384,60 @@ static int show_smaps_rollup(struct seq_file *m, void *v)
>  		goto out_put_task;
>  	}
>  
> -	ret = lock_ctx_mm(&priv->lock_ctx);
> -	if (ret)
> -		goto out_put_mm;
> -
>  	hold_task_mempolicy(priv);
> -	vma = vma_next(&vmi);
> +	rcu_read_lock();
> +	reset_lock_ctx(lock_ctx);
>  
> +	vma_iter_init(&priv->iter, mm, 0);
> +	vma = proc_get_vma(m, &pos);
>  	if (unlikely(!vma))
>  		goto empty_set;
>  
> -	vma_start = vma->vm_start;
> -	do {
> -		smap_gather_stats(priv, vma, &mss, vma->vm_start);
> -		last_vma_end = vma->vm_end;
> +	if (!IS_ERR(vma) && vma != get_gate_vma(lock_ctx->mm))
> +		vma_start = vma->vm_start;
> +
> +	while (vma) {
> +		if (IS_ERR(vma)) {
> +			ret = PTR_ERR(vma);
> +			goto out_unlock;
> +		}
> +
> +		if (vma == get_gate_vma(lock_ctx->mm))
> +			break;
>  
>  		/*
> -		 * Release mmap_lock temporarily if someone wants to
> -		 * access it for write request.
> +		 * If after retaking the lock, already reported VMA grew or
> +		 * merged with the next one, smap_gather_stats() will gather
> +		 * stats for the remaining portion by starting at last_vma_end.
>  		 */
> -		if (mmap_lock_is_contended(mm)) {
> -			vma_iter_invalidate(&vmi);
> -			unlock_ctx_mm(&priv->lock_ctx);
> -			ret = lock_ctx_mm(&priv->lock_ctx);
> -			if (ret) {
> -				release_task_mempolicy(priv);
> -				goto out_put_mm;
> -			}
> +		smap_gather_stats(priv, vma, &mss, last_vma_end);

Patch 3 made smap_gather_stats() reject starts below the VMA, while this
function initializes last_vma_end to zero.

The first ordinary VMA is therefore skipped. lock_next_vma() can also
return a VMA beginning after the requested position, so the first VMA
after every unmapped gap is skipped as well. This causes
smaps_rollup to underreport RSS, PSS, swap, and the other accumulated
values.

I think you need:
  unsigned long start = max(last_vma_end, vma->vm_start);
  smap_gather_stats(priv, vma, &mss, start);



> +		last_vma_end = vma->vm_end;
>  
> +		/*
> +		 * If the VMA lock is not taken, we hold the often contended
> +		 * mmap lock. This can happen if we had to fall back to the
> +		 * mmap lock.
> +		 *
> +		 * To relieve pressure, check if it is indeed contended, then
> +		 * temporarily release it.
> +		 */
> +		if (lock_ctx->mmap_locked &&
> +		    mmap_lock_is_contended(lock_ctx->mm)) {
> +			unlock_ctx_mm(lock_ctx);
>  			/*
> -			 * After dropping the lock, there are four cases to
> -			 * consider. See the following example for explanation.
> -			 *
> -			 *   +------+------+-----------+
> -			 *   | VMA1 | VMA2 | VMA3      |
> -			 *   +------+------+-----------+
> -			 *   |      |      |           |
> -			 *  4k     8k     16k         400k
> -			 *
> -			 * Suppose we drop the lock after reading VMA2 due to
> -			 * contention, then we get:
> -			 *
> -			 *	last_vma_end = 16k
> -			 *
> -			 * 1) VMA2 is freed, but VMA3 exists:
> -			 *
> -			 *    vma_next(vmi) will return VMA3.
> -			 *    In this case, just continue from VMA3.
> -			 *
> -			 * 2) VMA2 still exists:
> -			 *
> -			 *    vma_next(vmi) will return VMA3.
> -			 *    In this case, just continue from VMA3.
> -			 *
> -			 * 3) No more VMAs can be found:
> -			 *
> -			 *    vma_next(vmi) will return NULL.
> -			 *    No more things to do, just break.
> -			 *
> -			 * 4) (last_vma_end - 1) is the middle of a vma (VMA'):
> -			 *
> -			 *    vma_next(vmi) will return VMA' whose range
> -			 *    contains last_vma_end.
> -			 *    Iterate VMA' from last_vma_end.
> +			 * Even though we previously fell back to mmap lock,
> +			 * we try taking VMA lock for the next VMA, since it
> +			 * might not be under modification. In the worst case
> +			 * we will fall back to mmap lock again.
>  			 */
> -			vma = vma_next(&vmi);
> -			/* Case 3 above */
> -			if (!vma)
> -				break;
> -
> -			/* Case 1 and 2 above */
> -			if (vma->vm_start >= last_vma_end) {
> -				smap_gather_stats(priv, vma, &mss, vma->vm_start);
> -				last_vma_end = vma->vm_end;
> -				continue;
> -			}
> -
> -			/* Case 4 above */
> -			if (vma->vm_end > last_vma_end) {
> -				smap_gather_stats(priv, vma, &mss, last_vma_end);
> -				last_vma_end = vma->vm_end;
> -			}
> +			rcu_read_lock();
> +			reset_lock_ctx(lock_ctx);
> +			/* Resume from the last position. */
> +			pos = last_vma_end;
> +			vma_iter_init(&priv->iter, mm, pos);
>  		}
> -	} for_each_vma(vmi, vma);
> +		vma = proc_get_vma(m, &pos);
> +	}
>  
>  empty_set:
>  	show_vma_header_prefix(m, vma_start, last_vma_end, 0, 0, 0, 0);
> @@ -1483,10 +1446,14 @@ static int show_smaps_rollup(struct seq_file *m, void *v)
>  
>  	__show_smap(m, &mss, true);
>  
> +out_unlock:
> +	if (lock_ctx->mmap_locked) {
> +		unlock_ctx_mm(lock_ctx);
> +	} else {
> +		unlock_ctx_vma(lock_ctx);
> +		rcu_read_unlock();
> +	}
>  	release_task_mempolicy(priv);
> -	unlock_ctx_mm(&priv->lock_ctx);
> -
> -out_put_mm:
>  	mmput(mm);
>  out_put_task:
>  	put_task_struct(priv->task);
> -- 
> 2.55.0.979.g7e5102b832-goog
> 
> 


  parent reply	other threads:[~2026-09-09 14:25 UTC|newest]

Thread overview: 31+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-07  6:39 [PATCH v2 0/5] read proc/pid/smaps_rollup under per-vma lock Suren Baghdasaryan
2026-09-07  6:39 ` [PATCH v2 1/5] proc/task_mmu: remove unnecessary helpers Suren Baghdasaryan
2026-09-07 16:44   ` Usama Arif
2026-09-08 17:58   ` Liam R. Howlett
2026-09-09 17:06   ` David Hildenbrand (Arm)
2026-09-09 17:13     ` Suren Baghdasaryan
2026-09-09 17:17       ` David Hildenbrand (Arm)
2026-09-09 18:29         ` Suren Baghdasaryan
2026-09-07  6:39 ` [PATCH v2 2/5] proc/task_mmu: remove unnecessary inlines in function definitions Suren Baghdasaryan
2026-09-07 16:49   ` Usama Arif
2026-09-08 18:01   ` Liam R. Howlett
2026-09-09 17:07   ` David Hildenbrand (Arm)
2026-09-09 17:15     ` Suren Baghdasaryan
2026-09-07  6:39 ` [PATCH v2 3/5] proc/task_mmu: remove special-casing of smap_gather_stats() start parameter Suren Baghdasaryan
2026-09-08 18:07   ` Liam R. Howlett
2026-09-09 17:16   ` David Hildenbrand (Arm)
2026-09-09 18:28     ` Suren Baghdasaryan
2026-09-09 19:16       ` David Hildenbrand (Arm)
2026-09-09 21:51         ` Suren Baghdasaryan
2026-09-10  7:41           ` David Hildenbrand (Arm)
2026-09-07  6:39 ` [PATCH v2 4/5] proc/task_mmu: read proc/pid/smaps_rollup under per-vma lock Suren Baghdasaryan
2026-09-08 18:17   ` Liam R. Howlett
2026-09-09 14:25   ` Usama Arif [this message]
2026-09-09 16:13     ` Suren Baghdasaryan
2026-09-09 17:23   ` David Hildenbrand (Arm)
2026-09-09 17:58     ` Suren Baghdasaryan
2026-09-10  7:44       ` David Hildenbrand (Arm)
2026-09-07  6:39 ` [PATCH v2 5/5] selftests/proc: add /proc/pid/smaps_rollup tearing tests Suren Baghdasaryan
2026-09-08 18:18   ` Liam R. Howlett
2026-09-08 16:04 ` [PATCH v2 0/5] read proc/pid/smaps_rollup under per-vma lock Xueyuan Chen
2026-09-08 16:08   ` Suren Baghdasaryan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260909142527.1601051-1-usama.arif@linux.dev \
    --to=usama.arif@linux.dev \
    --cc=akpm@linux-foundation.org \
    --cc=david@redhat.com \
    --cc=jannh@google.com \
    --cc=liam@infradead.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=paulmck@kernel.org \
    --cc=pfalcato@suse.de \
    --cc=surenb@google.com \
    --cc=vbabka@kernel.org \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox