From: Suren Baghdasaryan <surenb@google.com>
To: akpm@linux-foundation.org
Cc: dave.hansen@linux.intel.com, Liam.Howlett@oracle.com,
ljs@kernel.org, david@kernel.org, willy@infradead.org,
shakeel.butt@linux.dev, vbabka@kernel.org, jannh@google.com,
aliceryhl@google.com, arve@android.com, cmllamas@google.com,
christian@brauner.io, tkjos@android.com, dsahern@kernel.org,
davem@davemloft.net, gregkh@linuxfoundation.org,
linux-kernel@vger.kernel.org, linux-mm@kvack.org,
netdev@vger.kernel.org, surenb@google.com
Subject: [PATCH v4 3/5] mm: Add RCU-based VMA lookup helper that waits for writers
Date: Thu, 6 Aug 2026 13:05:46 -0700 [thread overview]
Message-ID: <20260806200548.3124802-4-surenb@google.com> (raw)
In-Reply-To: <20260806200548.3124802-1-surenb@google.com>
From: Dave Hansen <dave.hansen@linux.intel.com>
== Background ==
There are basically two parallel ways to look up a VMA: the
traditional way, which is protected by mmap_read_lock, and the RCU-based
per-VMA lock way which is based on RCU and refcounts.
== Problem ==
The mmap_lock one is more straightforward to use but it has a big
disadvantage in that it can not be mixed with page faults since those
can take mmap_lock for read, which can deadlock when mixed with nested
page faults and parallel writers.
For example:
mmap_read_lock(mm);
// Another thread does mmap_write_lock().
// New mmap_lock readers are blocked.
vma = vma_lookup(mm, address);
// This deadlocks on mmap_read_lock() if it faults:
copy_from_user(address);
mmap_read_unlock(mm);
The per-VMA lock can be mixed with faults, but they can fail and need to
be able to fall back to the traditional way.
== Solution ==
Add vma_start_read_unlocked() - a variant of the RCU-based lookup that
waits for writers. This is basically the same as the existing RCU-based
lookup, but on a failure to lock it temporarily takes mmap_lock for read
and waits for writers to finish before locking the VMA, dropping the
mmap_lock and returning the locked VMA. This has some advantages:
1. Callers do not need to have a fallback path for when they
collide with writers.
2. It can be used in contexts where page faults can happen because
it can take the mmap_lock for read but never *holds* it.
3. Its fast path does not require taking mmap_lock for read.
Basically, when applied correctly, this approach results in faster
*and* simpler code.
While at it, fix the comments for vma_start_read_locked(),
vma_start_read_locked_nested(), and uffd_lock_vma().
Suggested-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Signed-off-by: Suren Baghdasaryan <surenb@google.com>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: "Liam R. Howlett" <Liam.Howlett@oracle.com>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: linux-mm@kvack.org
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Arve Hjønnevåg <arve@android.com>
Cc: Todd Kjos <tkjos@android.com>
Cc: Christian Brauner <christian@brauner.io>
Cc: Carlos Llamas <cmllamas@google.com>
Cc: Alice Ryhl <aliceryhl@google.com>
Cc: "David S. Miller" <davem@davemloft.net>
Cc: David Ahern <dsahern@kernel.org>
Cc: netdev@vger.kernel.org
---
include/linux/mmap_lock.h | 19 +++++++++++++++----
mm/mmap_lock.c | 35 +++++++++++++++++++++++++++++++++++
mm/userfaultfd.c | 6 ++++--
3 files changed, 54 insertions(+), 6 deletions(-)
diff --git a/include/linux/mmap_lock.h b/include/linux/mmap_lock.h
index 7b2bbb09a952..a23fe6cbe301 100644
--- a/include/linux/mmap_lock.h
+++ b/include/linux/mmap_lock.h
@@ -228,10 +228,14 @@ static inline void vma_refcount_put(struct vm_area_struct *vma)
}
/*
- * Use only while holding mmap read lock which guarantees that locking will not
- * fail (nobody can concurrently write-lock the vma). vma_start_read() should
+ * Use only while holding mmap read lock which guarantees that vma lock is not
+ * contended (nobody can concurrently write-lock the vma). vma_start_read() should
* not be used in such cases because it might fail due to mm_lock_seq overflow.
* This functionality is used to obtain vma read lock and drop the mmap read lock.
+ *
+ * VMA can't be detached while we are holding mmap lock, therefore in practice this
+ * function can fail only when there are so many readers that vm_refcnt overflows.
+ * The failure case is very unlikely and is already annotated as such internally.
*/
static inline bool vma_start_read_locked_nested(struct vm_area_struct *vma, int subclass)
{
@@ -247,16 +251,23 @@ static inline bool vma_start_read_locked_nested(struct vm_area_struct *vma, int
}
/*
- * Use only while holding mmap read lock which guarantees that locking will not
- * fail (nobody can concurrently write-lock the vma). vma_start_read() should
+ * Use only while holding mmap read lock which guarantees that vma lock is not
+ * contended (nobody can concurrently write-lock the vma). vma_start_read() should
* not be used in such cases because it might fail due to mm_lock_seq overflow.
* This functionality is used to obtain vma read lock and drop the mmap read lock.
+ *
+ * VMA can't be detached while we are holding mmap lock, therefore in practice this
+ * function can fail only when there are so many readers that vm_refcnt overflows.
+ * The failure case is very unlikely and is already annotated as such internally.
*/
static inline bool vma_start_read_locked(struct vm_area_struct *vma)
{
return vma_start_read_locked_nested(vma, 0);
}
+struct vm_area_struct *vma_start_read_unlocked(struct mm_struct *mm,
+ unsigned long address);
+
static inline void vma_end_read(struct vm_area_struct *vma)
{
vma_refcount_put(vma);
diff --git a/mm/mmap_lock.c b/mm/mmap_lock.c
index e20d01e8d38f..1c4902131e98 100644
--- a/mm/mmap_lock.c
+++ b/mm/mmap_lock.c
@@ -338,6 +338,41 @@ struct vm_area_struct *lock_vma_under_rcu(struct mm_struct *mm,
return NULL;
}
+/**
+ * vma_start_read_unlocked() - Find the VMA covering 'address' and read-lock it.
+ * @mm: the mm_struct of the address space to search
+ * @address: address that the vma should contain
+ *
+ * The fast path does not take mmap_lock. Waits for writers to finish if the
+ * VMA is being modified by taking mmap_lock.
+ * Use when mmap_lock is not held, otherwise use vma_start_read_locked().
+ * Nothing prevents VMAs being unmapped/mapped before or after the VMA is
+ * looked up, if a stronger guarantee is required, take an mmap_lock.
+ *
+ * Return: If a VMA exists which spans @address, return that VMA, read-locked.
+ * If no VMA is mapped there or, very unlikely, a reference count overflow
+ * occurred, return NULL.
+ */
+struct vm_area_struct *vma_start_read_unlocked(struct mm_struct *mm,
+ unsigned long address)
+{
+ struct vm_area_struct *vma;
+
+ /* Fast path: return stable VMA covering 'address': */
+ vma = lock_vma_under_rcu(mm, address);
+ if (vma)
+ return vma;
+
+ /* Slow path: preclude VMA writers by temporarily getting mmap read lock. */
+ mmap_read_lock(mm);
+ vma = vma_lookup(mm, address);
+ if (vma && !vma_start_read_locked(vma))
+ vma = NULL;
+ mmap_read_unlock(mm);
+
+ return vma;
+}
+
static struct vm_area_struct *lock_next_vma_under_mmap_lock(struct mm_struct *mm,
struct vma_iterator *vmi,
unsigned long from_addr)
diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c
index edd90892f8cc..c3a0c38a3dc3 100644
--- a/mm/userfaultfd.c
+++ b/mm/userfaultfd.c
@@ -129,8 +129,10 @@ struct vm_area_struct *find_vma_and_prepare_anon(struct mm_struct *mm,
*
* Should be called without holding mmap_lock.
*
- * Return: A locked vma containing @address, -ENOENT if no vma is found, or
- * -ENOMEM if anon_vma couldn't be allocated.
+ * Return: A locked vma containing @address, -ENOENT if no vma is found,
+ * -ENOMEM if anon_vma couldn't be allocated, or -EAGAIN if vma refcount
+ * overflow happened due to high number of readers and the caller should
+ * retry later.
*/
static struct vm_area_struct *uffd_lock_vma(struct mm_struct *mm,
unsigned long address)
--
2.55.0.654.g21b8a5bc05-goog
next prev parent reply other threads:[~2026-08-06 20:06 UTC|newest]
Thread overview: 36+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-06 20:05 [PATCH v4 0/5] mm: Unconditional per-VMA locks and cleanups Suren Baghdasaryan
2026-08-06 20:05 ` [PATCH v4 1/5] mm: Make per-VMA locks available universally Suren Baghdasaryan
2026-08-07 15:37 ` Vlastimil Babka (SUSE)
2026-08-08 1:12 ` Matthew Wilcox
2026-08-08 6:08 ` Suren Baghdasaryan
2026-08-10 8:52 ` Lorenzo Stoakes (ARM)
2026-08-10 9:42 ` Lorenzo Stoakes (ARM)
2026-08-10 16:40 ` Suren Baghdasaryan
2026-08-10 10:17 ` Lorenzo Stoakes (ARM)
2026-08-10 15:22 ` Dave Hansen
2026-08-10 15:56 ` Matthew Wilcox
2026-08-10 17:04 ` Suren Baghdasaryan
2026-08-06 20:05 ` [PATCH v4 2/5] binder: Make shrinker rely solely on per-VMA lock Suren Baghdasaryan
2026-08-07 14:26 ` Alice Ryhl
2026-08-10 10:22 ` Lorenzo Stoakes (ARM)
2026-08-10 18:32 ` Carlos Llamas
2026-08-10 18:54 ` Suren Baghdasaryan
2026-08-10 19:16 ` Carlos Llamas
2026-08-10 21:00 ` Suren Baghdasaryan
2026-08-11 2:50 ` Carlos Llamas
2026-08-10 18:57 ` Carlos Llamas
2026-08-06 20:05 ` Suren Baghdasaryan [this message]
2026-08-07 15:39 ` [PATCH v4 3/5] mm: Add RCU-based VMA lookup helper that waits for writers Vlastimil Babka (SUSE)
2026-08-08 7:24 ` Matthew Wilcox
2026-08-09 1:07 ` Suren Baghdasaryan
2026-08-10 10:24 ` Lorenzo Stoakes (ARM)
2026-08-10 15:25 ` Dave Hansen
2026-08-08 7:37 ` Matthew Wilcox
2026-08-10 10:41 ` Lorenzo Stoakes (ARM)
2026-08-10 17:07 ` Suren Baghdasaryan
2026-08-06 20:05 ` [PATCH v4 4/5] binder: Remove mmap_lock fallback Suren Baghdasaryan
2026-08-06 20:05 ` [PATCH v4 5/5] tcp: Remove mmap_lock fallback path Suren Baghdasaryan
2026-08-08 7:26 ` Matthew Wilcox
2026-08-09 0:51 ` Suren Baghdasaryan
2026-08-10 15:26 ` Dave Hansen
2026-08-10 17:09 ` Suren Baghdasaryan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260806200548.3124802-4-surenb@google.com \
--to=surenb@google.com \
--cc=Liam.Howlett@oracle.com \
--cc=akpm@linux-foundation.org \
--cc=aliceryhl@google.com \
--cc=arve@android.com \
--cc=christian@brauner.io \
--cc=cmllamas@google.com \
--cc=dave.hansen@linux.intel.com \
--cc=davem@davemloft.net \
--cc=david@kernel.org \
--cc=dsahern@kernel.org \
--cc=gregkh@linuxfoundation.org \
--cc=jannh@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ljs@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=shakeel.butt@linux.dev \
--cc=tkjos@android.com \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.