From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: Suren Baghdasaryan <surenb@google.com>
Cc: akpm@linux-foundation.org, dave.hansen@linux.intel.com,
Liam.Howlett@oracle.com, david@redhat.com, willy@infradead.org,
shakeel.butt@linux.dev, vbabka@kernel.org, jannh@google.com,
aliceryhl@google.com, arve@android.com, cmllamas@google.com,
christian@brauner.io, tkjos@android.com, dsahern@kernel.org,
davem@davemloft.net, gregkh@linuxfoundation.org,
linux-kernel@vger.kernel.org, linux-mm@kvack.org,
netdev@vger.kernel.org
Subject: Re: [PATCH v3 2/5] binder: Make shrinker rely solely on per-VMA lock
Date: Tue, 4 Aug 2026 10:08:29 +0100 [thread overview]
Message-ID: <anGpvW6XwgWH1t6D@lucifer> (raw)
In-Reply-To: <CAJuCfpHPcE0=7825xiusBN0MmCp1hrnsDuYrfTKwGM43ozmmOQ@mail.gmail.com>
On Mon, Aug 03, 2026 at 11:31:14AM -0700, Suren Baghdasaryan wrote:
> On Mon, Aug 3, 2026 at 4:10 AM Lorenzo Stoakes (ARM) <ljs@kernel.org> wrote:
> >
> > On Sun, Aug 02, 2026 at 02:54:56PM -0700, Suren Baghdasaryan wrote:
> > > From: Dave Hansen <dave.hansen@linux.intel.com>
> > >
> > > tl;dr: lock_vma_under_rcu() is already a trylock. No need to do both
> > > it and mmap_read_trylock().
> > >
> > > Long Version:
> > >
> > > == Background ==
> > >
> > > Historically, binder used an mmap_read_trylock() in its shrinker code.
> > > This ensures that reclaim is not blocked on an mmap_lock. Commit
> > > 95bc2d4a9020 ("binder: use per-vma lock in page reclaiming") added
> > > support for the per-VMA lock, but left mmap_read_trylock() as a
> > > fallback.
> > >
> > > This was presumably because the per-VMA locking can fail for several
> > > reasons and most (all?) lock_vma_under_rcu() callers have a fallback
> > > to mmap_read_trylock().
> > >
> > > == Problem ==
> > >
> > > The fallback is not worth the complexity here. lock_vma_under_rcu() is
> > > essentially already a non-blocking trylock. The main reason it fails
> > > is also the reason mmap_read_trylock() fails: something is holding
> > > mmap_write_lock().
> > >
> > > The only remedy for a collision with mmap_write_lock() is to wait,
> > > which this code can not do. So the "fallback" after
> > > lock_vma_under_rcu() failure is not really a fallback: it is really
> > > likely to just be retrying in vain. That retry in an of itself isn't
> > > horrible. But it adds complexity.
> > >
> > > == Solution ==
> > >
> > > Now that per-VMA locks are universally available, lock_vma_under_rcu()
> > > will not persistently fail. Rely on it alone and simplify the code.
> > >
> > > Full disclosure: I originally tried to do this with
> > > lock_vma_under_rcu_wait(), but it did not fit well with the mmap_lock
> > > trylock semantics. Claude caught this in a review and suggested the
> > > approach in this path. It seemed sane to me. So, Suggesed-by: Claude,
> > > I guess.
> > >
> > > Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
> > > Signed-off-by: Suren Baghdasaryan <surenb@google.com>
> > > Cc: Andrew Morton <akpm@linux-foundation.org>
> > > Cc: "Liam R. Howlett" <Liam.Howlett@oracle.com>
> > > Cc: Vlastimil Babka <vbabka@kernel.org>
> > > Cc: Shakeel Butt <shakeel.butt@linux.dev>
> > > Cc: linux-mm@kvack.org
> > > Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
> > > Cc: Arve Hjønnevåg <arve@android.com>
> > > Cc: Todd Kjos <tkjos@android.com>
> > > Cc: Christian Brauner <christian@brauner.io>
> > > Cc: Carlos Llamas <cmllamas@google.com>
> > > Cc: Alice Ryhl <aliceryhl@google.com>
> > > Cc: "David S. Miller" <davem@davemloft.net>
> > > Cc: David Ahern <dsahern@kernel.org>
> > > Cc: netdev@vger.kernel.org
> > > ---
> > > drivers/android/binder_alloc.c | 29 +++++++++++++++--------------
> > > 1 file changed, 15 insertions(+), 14 deletions(-)
> > >
> > > diff --git a/drivers/android/binder_alloc.c b/drivers/android/binder_alloc.c
> > > index e4488ad86a65..84104ba04e30 100644
> > > --- a/drivers/android/binder_alloc.c
> > > +++ b/drivers/android/binder_alloc.c
> > > @@ -1142,7 +1142,6 @@ enum lru_status binder_alloc_free_page(struct list_head *item,
> > > struct vm_area_struct *vma;
> > > struct page *page_to_free;
> > > unsigned long page_addr;
> > > - int mm_locked = 0;
> > > size_t index;
> > >
> > > if (!mmget_not_zero(mm))
> > > @@ -1151,14 +1150,20 @@ enum lru_status binder_alloc_free_page(struct list_head *item,
> > > index = mdata->page_index;
> > > page_addr = alloc->vm_start + index * PAGE_SIZE;
> > >
> > > - /* attempt per-vma lock first */
> > > + /*
> > > + * Attempt per-vma lock. This is essentially a
> > > + * "trylock". It can fail even if the VMA exists
> > > + * for 'page_addr'.
> > > + */
> >
> > This makes me wonder whether lock_vma_under_rcu() should really become
> > vma_trylock() at some point in time? :)
> >
> > Or at least have 'trylock' in the name.
>
> Makes sense. I think I'll postpone renames until after the series are
> merged. Don't want to mix too many changes together.
Yeah absolutely :) this isn't for this series, just thinking out loud!
>
> >
> > > vma = lock_vma_under_rcu(mm, page_addr);
> > > if (!vma) {
> > > - /* fall back to mmap_lock */
> > > - if (!mmap_read_trylock(mm))
> > > - goto err_mmap_read_lock_failed;
> > > - mm_locked = 1;
> > > - vma = vma_lookup(mm, page_addr);
> > > + /*
> > > + * If the vma exists, we can't continue because we cannot
> > > + * remove the page from the vma. However, if the vma was
> > > + * unmapped, it's okay to continue.
> > > + */
> > > + if (binder_alloc_is_mapped(alloc))
> > > + goto err_vma_lock_failed;
> >
> > Hmm, it seems a bit odd to me that you also have:
> >
> > if (vma && !binder_alloc_is_mapped(alloc))
> > goto err_invalid_vma;
> >
> > Below?
> >
> > So you have:
> >
> > Before:
> >
> > |binder_alloc_is_mapped()?
> > |yes no
> > --------|-----------------
> > vma is mapped? yes |OK abort
> > no |OK OK
> >
> > Now:
> >
> > |binder_alloc_is_mapped()?
> > |yes no
> > --------|-----------------
> > vma is mapped? maybe |abort OK
> > yes |OK abort
> > no |OK OK
> >
> > The 'maybe' is because the VMA trylock failed.
>
> The case you are considering is lock_vma_under_rcu() failed for some
> reason other than lock contention (say seqno overflow). In that case
Well it could also be due to lock contention right?
> we get vma==NULL and binder_alloc_is_mapped() is called without any
> lock (VMA or mmap lock) being held. I'm not sure if this is a real
> problem since I see other places calling binder_alloc_is_mapped()
> without locking. Alice, Carlos, is this a problem?
There seems to be a contradiction here though in that - the
binder_alloc_is_mapped() call below is predicated on vma != NULL.
But here lock contention could mean the VMA is mapped, but then
binder_alloc_is_mapped() returns false but you still proceed.
Anyway I don't really understand the semantics here so will leave it to you guys
as to whether this is actually an issue :)
>
> >
> > So the issue is you might have a case where the VMA _is_ mapped but
> > !binder_alloc_is_mapped(), which previously aborted because of the vma &&
> > !binder_alloc_is_mapped() check.
> >
> > It seems like:
> >
> > /*
> > * Since a binder_alloc can only be mapped once, we ensure
> > * the vma corresponds to this mapping by checking whether
> > * the binder_alloc is still mapped.
> > */
> > if (vma && !binder_alloc_is_mapped(alloc))
> > goto err_invalid_vma;
> >
> > Is testing for a specific scenario 'we found a VMA but it turns out it's
> > invalid' and aborting if so.
> >
> > So either this check should be removed or you should uncondtionally abort if
> > !vma I think?
>
> I think this check is fine because it basically checks
> binder_alloc_is_mapped() after stabilizing the address range.
> Unconditionally aborting if !vma would prevent us from freeing the
> page if the VMA was already unmapped. The ultimate question is whether
> we can rely on binder_alloc_is_mapped() alone when freeing that page.
> IOW, VMA might still be in the VMA tree but
> binder_alloc_is_mapped()==false, can we free the page?
>
Yeah.
> >
> >
> > > }
> > >
> > > if (!mutex_trylock(&alloc->mutex))
> > > @@ -1191,9 +1196,7 @@ enum lru_status binder_alloc_free_page(struct list_head *item,
> > > }
> > >
> > > mutex_unlock(&alloc->mutex);
> > > - if (mm_locked)
> > > - mmap_read_unlock(mm);
> > > - else
> > > + if (vma)
> > > vma_end_read(vma);
> > > mmput_async(mm);
> > > binder_free_page(page_to_free);
> > > @@ -1203,11 +1206,9 @@ enum lru_status binder_alloc_free_page(struct list_head *item,
> > > err_invalid_vma:
> > > mutex_unlock(&alloc->mutex);
> > > err_get_alloc_mutex_failed:
> > > - if (mm_locked)
> > > - mmap_read_unlock(mm);
> > > - else
> > > + if (vma)
> > > vma_end_read(vma);
> > > -err_mmap_read_lock_failed:
> > > +err_vma_lock_failed:
> > > mmput_async(mm);
> > > err_mmget:
> > > return LRU_SKIP;
> > > --
> > > 2.55.0.508.g3f0d502094-goog
> > >
> >
> > --
> > Cheers, Lorenzo
--
Cheers, Lorenzo
next prev parent reply other threads:[~2026-08-04 9:08 UTC|newest]
Thread overview: 48+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-02 21:54 [PATCH v3 0/5] mm: Unconditional per-VMA locks and cleanups Suren Baghdasaryan
2026-08-02 21:54 ` [PATCH v3 1/5] mm: Make per-VMA locks available universally Suren Baghdasaryan
2026-08-03 10:49 ` Lorenzo Stoakes (ARM)
2026-08-03 14:01 ` Vlastimil Babka (SUSE)
2026-08-03 17:45 ` Suren Baghdasaryan
2026-08-03 15:24 ` Suren Baghdasaryan
2026-08-03 16:08 ` Lorenzo Stoakes (ARM)
2026-08-03 17:41 ` Suren Baghdasaryan
2026-08-03 17:45 ` Suren Baghdasaryan
2026-08-03 21:12 ` Jann Horn
2026-08-04 8:56 ` Lorenzo Stoakes (ARM)
2026-08-04 14:59 ` Suren Baghdasaryan
2026-08-03 19:33 ` Jann Horn
2026-08-03 19:43 ` Suren Baghdasaryan
2026-08-02 21:54 ` [PATCH v3 2/5] binder: Make shrinker rely solely on per-VMA lock Suren Baghdasaryan
2026-08-03 9:48 ` Alice Ryhl
2026-08-03 10:50 ` Lorenzo Stoakes (ARM)
2026-08-03 11:11 ` Lorenzo Stoakes (ARM)
2026-08-03 11:33 ` Lorenzo Stoakes (ARM)
2026-08-03 18:02 ` Suren Baghdasaryan
2026-08-04 9:15 ` Lorenzo Stoakes (ARM)
2026-08-03 11:10 ` Lorenzo Stoakes (ARM)
2026-08-03 18:31 ` Suren Baghdasaryan
2026-08-04 9:08 ` Lorenzo Stoakes (ARM) [this message]
2026-08-04 9:04 ` Alice Ryhl
2026-08-04 9:11 ` Lorenzo Stoakes (ARM)
2026-08-04 14:54 ` Suren Baghdasaryan
2026-08-02 21:54 ` [PATCH v3 3/5] mm: Add RCU-based VMA lookup helper that waits for writers Suren Baghdasaryan
2026-08-03 11:28 ` Lorenzo Stoakes (ARM)
2026-08-03 19:01 ` Suren Baghdasaryan
2026-08-04 8:47 ` Lorenzo Stoakes (ARM)
2026-08-04 15:00 ` Suren Baghdasaryan
2026-08-03 14:55 ` Vlastimil Babka (SUSE)
2026-08-03 15:00 ` Lorenzo Stoakes (ARM)
2026-08-03 16:24 ` Vlastimil Babka (SUSE)
2026-08-03 16:43 ` Lorenzo Stoakes (ARM)
2026-08-03 19:13 ` Suren Baghdasaryan
2026-08-04 7:59 ` Vlastimil Babka (SUSE)
2026-08-04 8:44 ` Lorenzo Stoakes (ARM)
2026-08-02 21:54 ` [PATCH v3 4/5] binder: Remove mmap_lock fallback Suren Baghdasaryan
2026-08-03 10:34 ` Alice Ryhl
2026-08-03 19:14 ` Suren Baghdasaryan
2026-08-03 11:33 ` Lorenzo Stoakes (ARM)
2026-08-03 19:16 ` Suren Baghdasaryan
2026-08-02 21:54 ` [PATCH v3 5/5] tcp: Remove mmap_lock fallback path Suren Baghdasaryan
2026-08-03 2:11 ` [PATCH v3 0/5] mm: Unconditional per-VMA locks and cleanups Barry Song
2026-08-03 17:51 ` Suren Baghdasaryan
2026-08-04 9:27 ` Lorenzo Stoakes (ARM)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anGpvW6XwgWH1t6D@lucifer \
--to=ljs@kernel.org \
--cc=Liam.Howlett@oracle.com \
--cc=akpm@linux-foundation.org \
--cc=aliceryhl@google.com \
--cc=arve@android.com \
--cc=christian@brauner.io \
--cc=cmllamas@google.com \
--cc=dave.hansen@linux.intel.com \
--cc=davem@davemloft.net \
--cc=david@redhat.com \
--cc=dsahern@kernel.org \
--cc=gregkh@linuxfoundation.org \
--cc=jannh@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=netdev@vger.kernel.org \
--cc=shakeel.butt@linux.dev \
--cc=surenb@google.com \
--cc=tkjos@android.com \
--cc=vbabka@kernel.org \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.