From: Uladzislau Rezki <urezki@gmail.com>
To: Dev Jain <dev.jain@arm.com>
Cc: Uladzislau Rezki <urezki@gmail.com>,
Andrew Morton <akpm@linux-foundation.org>,
Ye Liu <ye.liu@linux.dev>, Ye Liu <liuye@kylinos.cn>,
linux-mm@kvack.org, linux-kernel@vger.kernel.org
Subject: Re: [PATCH v2] mm: vmalloc: fix vmap_purge_lock livelock under memory pressure
Date: Thu, 3 Sep 2026 11:10:57 +0200 [thread overview]
Message-ID: <apk5oYoBQ1EDc0u9@milan> (raw)
In-Reply-To: <c88d035a-cf81-4431-bc28-bb09878e09d7@arm.com>
On Wed, Sep 02, 2026 at 10:00:57AM +0530, Dev Jain wrote:
>
>
> On 01/09/26 10:29 pm, Uladzislau Rezki wrote:
> > On Tue, Sep 01, 2026 at 06:43:43PM +0200, Uladzislau Rezki wrote:
> >> On Tue, Sep 01, 2026 at 05:07:30PM +0800, Ye Liu wrote:
> >>>
> >>>
> >>> 在 2026/9/1 14:22, Dev Jain 写道:
> >>>>
> >>>>
> >>>> On 31/08/26 3:24 pm, Uladzislau Rezki wrote:
> >>>>> On Mon, Aug 31, 2026 at 11:39:14AM +0530, Dev Jain wrote:
> >>>>>>
> >>>>>>
> >>>>>> On 28/08/26 11:35 pm, Andrew Morton wrote:
> >>>>>>> On Fri, 28 Aug 2026 17:17:53 +0800 Ye Liu <ye.liu@linux.dev> wrote:
> >>>>>>>
> >>>>>>>> From: Ye Liu <liuye@kylinos.cn>
> >>>>>>>>
> >>>>>>>> The vmap_purge_lock mutex can be held for an extended period by
> >>>>>>>> __purge_vmap_area_lazy() which calls flush_work() to wait for
> >>>>>>>> purge_vmap_node workers while holding the lock. Under memory
> >>>>>>>> pressure, those workers may themselves be blocked in direct
> >>>>>>>> reclaim trying to acquire the same lock via the
> >>>>>>>> vmap_node_shrink_scan() shrinker callback, creating a circular
> >>>>>>>> dependency that deadlocks the entire system.
> >>>>>>>>
> >>>>>>>> Two places acquire vmap_purge_lock from paths that can be reached
> >>>>>>>> during direct reclaim:
> >>>>>>>>
> >>>>>>>> 1. vmap_node_shrink_scan(): replace blocking guard(mutex) with
> >>>>>>>> mutex_trylock(). This is a shrinker that only decays the vmap
> >>>>>>>> pool and returns SHRINK_STOP without freeing memory; skipping a
> >>>>>>>> decay cycle when the lock is contended is harmless and prevents
> >>>>>>>> tasks from piling up on the mutex in the direct reclaim path.
> >>>>>>>>
> >>>>>>>> 2. reclaim_and_purge_vmap_areas(): replace mutex_lock() with
> >>>>>>>> mutex_trylock(). This is called from the vmalloc allocation
> >>>>>>>> overflow path; if trylock fails, another thread is already
> >>>>>>>> purging and the allocator's retry will find freed space. The
> >>>>>>>> notifier chain provides a fallback if the retry still fails.
> >>>>>>>>
> >>>>>>>> Both trylock failures break the circular dependency: the lock
> >>>>>>>> holder's flush_work() can complete because workers are no longer
> >>>>>>>> blocked on vmap_purge_lock in the direct reclaim path.
> >>>>>>>
> >>>>>>> Thanks. AI review expressed a couple of concerns:
> >>>>>>> https://sashiko.dev/#/patchset/20260828091753.299295-1-ye.liu@linux.dev
> >>>>>>
> >>>>>>
> >>>>>> Sounds legit to me. Now there is no guarantee of purge being successful, and we
> >>>>>> can get a spurious failure.
> >>>>>>
> >>>>>> How about using a WQ_RECLAIM workqueue:
> >>>>>>
> >>>>>> vmap_purge_wq = alloc_workqueue("vmap_purge",
> >>>>>> WQ_MEM_RECLAIM | WQ_PERCPU, 0);
> >>>>>>
> >>>>>> I see the same pattern in __lru_add_drain_all() and kmem_cache_init_late().
> >>>>>>
> >>>>> WQ_MEM_RECLAIM makes sense but this is another patch.
> >>>>>
> >>>>> I copied here AI comment:
> >>>>>>
> >>>>>> Does replacing this blocking lock with a trylock break synchronization for
> >>>>>> the callers?
> >>>>>>
> >>>>> No it does not. If someone is doing reclaim we do not wait and do not try
> >>>>> to do it again thus fail allocation.
> >>>>>
> >>>>>> When vmalloc space is exhausted, alloc_vmap_area() calls
> >>>>>> reclaim_and_purge_vmap_areas() and relies on its blocking behavior to ensure
> >>>>>> that free space has actually been reclaimed before looping back to retry:
> >>>>>> mm/vmalloc.c:alloc_vmap_area() {
> >>>>>> ...
> >>>>>> overflow:
> >>>>>> if (!purged) {
> >>>>>> reclaim_and_purge_vmap_areas();
> >>>>>> purged = 1;
> >>>>>> goto retry;
> >>>>>> }
> >>>>>> ...
> >>>>>> }
> >>>>>> With this patch, if another thread holds vmap_purge_lock, mutex_trylock()
> >>>>>> fails and the function returns immediately.
> >>>>>>
> >>>>> If reclaim is in progress and trylock fails a caller repeats only one
> >>>>> time to retry an allocation. There is no any infinite loop.
> >>>>>
> >>>>>>
> >>>>>> The allocator then retries
> >>>>>> instantly without waiting for the concurrent purge to complete.
> >>>>>> Because the retry fails and purged is already 1, could this cause the
> >>>>>> allocation to abort and return a spurious vmalloc allocation failure
> >>>>>> (-EBUSY or -ENOMEM)?
> >>>>>>
> >>>>> vmap space can be fragmented and not avail for 32-bit systems. For
> >>>>> 64-bit system it is likely impossible.
> >>>>>
> >>>>> But, i think we can overt mutex_lock() into mutex_trylock() just only
> >>>>> in the:
> >>>>>
> >>>>> static unsigned long
> >>>>> vmap_node_shrink_scan(struct shrinker *shrink, struct shrink_control *sc)
> >>>>> {
> >>>>> struct vmap_node *vn;
> >>>>>
> >>>>> guard(mutex)(&vmap_purge_lock);
> >>>>> for_each_vmap_node(vn)
> >>>>> decay_va_pool_node(vn, true);
> >>>>>
> >>>>> return SHRINK_STOP;
> >>>>> }
> >>>>>
> >>>>> so the reclaim path is not blocked. It should also address an issue
> >>>>> reported by the Ye Liu <ye.liu@linux.dev>.
> >>>>
> >>>> IIUC you are suggesting mutex_trylock() only in the shrinker path. But
> >>>> then, the following is possible no: take purge lock, try to get a
> >>>> worker thread, worker thread is stuck in vmalloc -> alloc_vmap_area
> >>>> -> reclaim_and_purge_vmap_areas -> take purge lock?
> >>>>
> >>>>
> >>>
> >>> Personally, I lean toward the WQ_MEM_RECLAIM workqueue solution.
> >>> After taking a closer look at the code, I noticed a subtle but
> >>> potentially problematic scenario:
> >>>
> >>> When drain_vmap_area_work acquires vmap_purge_lock and calls into
> >>> __purge_vmap_area_lazy, it may subsequently invoke queue_work/queue_work_on
> >>> on the same CPU's system_wq. If the newly queued work ends up waiting
> >>> for an available worker on that same CPU, while the current worker is
> >>> blocked waiting for that very work to complete (via flush_work),
> >>> we could end up with a self-deadlock on a single CPU.
> >>>
> >>> Theoretically, this seems possible. I suspect the reason we don't
> >>> see widespread reports of such deadlocks is that the nr_purge_helpers
> >>> logic limits the number of asynchronous workers; when resources are tight,
> >>> it falls back to synchronous execution (purge_vmap_node directly),
> >>> which avoids queuing additional work.
> >>>
> >>> Using a dedicated workqueue with WQ_MEM_RECLAIM would provide a clean,
> >>> explicit isolation—ensuring forward progress under memory pressure and
> >>> eliminating the risk of interfering with other subsystems' workqueues.
> >>> I believe this approach is more robust in the long run.
> >>>
> >>> Perhaps like the code below:
> >>> the dedicated queue eliminates the self‑deadlock risk, and the trylock
> >>> in the shrinker prevents recursive lock attempts from reclaim contexts.
> >>>
> >>> diff --git a/mm/vmalloc.c b/mm/vmalloc.c
> >>> index bea9f76ed7e7..68fc1f5acb2f 100644
> >>> --- a/mm/vmalloc.c
> >>> +++ b/mm/vmalloc.c
> >>> @@ -2218,6 +2218,9 @@ static unsigned long lazy_max_pages(void)
> >>> */
> >>> static DEFINE_MUTEX(vmap_purge_lock);
> >>>
> >>> +/* Workqueue for lazy vmap purging; WQ_MEM_RECLAIM guarantees progress. */
> >>> +static struct workqueue_struct *vmap_purge_wq;
> >>> +
> >>> /* for per-CPU blocks */
> >>> static void purge_fragmented_blocks_allcpus(void);
> >>>
> >>> @@ -2408,9 +2411,9 @@ static bool __purge_vmap_area_lazy(unsigned long start, unsigned long end,
> >>> INIT_WORK(&vn->purge_work, purge_vmap_node);
> >>>
> >>> if (cpumask_test_cpu(i, cpu_online_mask))
> >>> - schedule_work_on(i, &vn->purge_work);
> >>> + queue_work_on(i, vmap_purge_wq, &vn->purge_work);
> >>> else
> >>> - schedule_work(&vn->purge_work);
> >>> + queue_work(vmap_purge_wq, &vn->purge_work);
> >>>
> >>> nr_purge_helpers--;
> >>> } else {
> >>> @@ -5519,10 +5522,14 @@ vmap_node_shrink_scan(struct shrinker *shrink, struct shrink_control *sc)
> >>> {
> >>> struct vmap_node *vn;
> >>>
> >>> - guard(mutex)(&vmap_purge_lock);
> >>> + if (!mutex_trylock(&vmap_purge_lock))
> >>> + return SHRINK_STOP;
> >>> +
> >>> for_each_vmap_node(vn)
> >>> decay_va_pool_node(vn, true);
> >>>
> >>> + mutex_unlock(&vmap_purge_lock);
> >>> +
> >>> return SHRINK_STOP;
> >>> }
> >>>
> >>> @@ -5575,6 +5582,17 @@ void __init vmalloc_init(void)
> >>> * Now we can initialize a free vmap space.
> >>> */
> >>> vmap_init_free_space();
> >>> +
> >>> + /*
> >>> + * A dedicated workqueue for lazy vmap purging. WQ_MEM_RECLAIM
> >>> + * reserves a rescue worker so queued purge work items are executed
> >>> + * even under memory pressure, when workers of the system workqueue
> >>> + * may be stuck in direct reclaim.
> >>> + */
> >>> + vmap_purge_wq = alloc_workqueue("vmap_purge",
> >>> + WQ_MEM_RECLAIM | WQ_PERCPU, 0);
> >>> + WARN_ON(!vmap_purge_wq);
> >>> +
> >>> vmap_initialized = true;
> >>>
> >> I agree. We should have it and it should be as separate patch, i.e.
> >> split vmap_node_shrink_scan() and dedicated per-cpu WQs per vmap drain.
> >>
> > And i sent out already the WQ_UNBOUND | WQ_MEM_RECLAIM and separate WQ
> > for vmap drain logic. It looks like Andrew/me forgot about it:
> >
> > https://lore.kernel.org/all/20260331202352.879718-1-urezki@gmail.com/
>
> Great! Perhaps resend it and we can review it?
>
I will and add you to Cc.
--
Uladzislau Rezki
next prev parent reply other threads:[~2026-09-03 9:11 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-28 9:17 [PATCH v2] mm: vmalloc: fix vmap_purge_lock livelock under memory pressure Ye Liu
2026-08-28 16:54 ` Uladzislau Rezki
2026-08-28 18:05 ` Andrew Morton
2026-08-31 6:09 ` Dev Jain
2026-08-31 9:54 ` Uladzislau Rezki
2026-09-01 1:42 ` Andrew Morton
2026-09-01 16:36 ` Uladzislau Rezki
2026-09-01 6:22 ` Dev Jain
2026-09-01 9:07 ` Ye Liu
2026-09-01 16:43 ` Uladzislau Rezki
2026-09-01 16:59 ` Uladzislau Rezki
2026-09-02 4:30 ` Dev Jain
2026-09-03 9:10 ` Uladzislau Rezki [this message]
2026-09-01 16:40 ` Uladzislau Rezki
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=apk5oYoBQ1EDc0u9@milan \
--to=urezki@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=dev.jain@arm.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=liuye@kylinos.cn \
--cc=ye.liu@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox