From: vitaly prosyak <vprosyak@amd.com>
To: Prike Liang <Prike.Liang@amd.com>, amd-gfx@lists.freedesktop.org
Cc: Alexander.Deucher@amd.com, Christian.Koenig@amd.com,
Vitaly.Prosyak@amd.com
Subject: Re: [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock
Date: Mon, 28 Sep 2026 18:49:25 -0400 [thread overview]
Message-ID: <123fa18d-1bf7-48e7-bace-cdf429148128@amd.com> (raw)
In-Reply-To: <20260902125001.621629-12-Prike.Liang@amd.com>
Reviewed-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
I have tested this patch on nv31. It successfully resolves the lockdep deadlock warning triggered during the reverse
ordering of &reset_domain->sem and &userq_mgr->userq_mutex. Please note that this patch requires a rebase to apply
cleanly onto the current target branch. Once rebased, the lock inversion issue is gone, and the trylock-and-retry logic
works cleanly to unblock recovery execution path interactions. We can safely enable this on the CI.
On 2026-09-02 08:49, Prike Liang wrote:
> The offending edge was code that did a blocking
> down_read(&adev->reset_domain->sem) as following, while
> holding userq_mutex. Since GPU recovery takes reset_domain
> ->sem for write and then transitively acquires userq_mutex,
> the reverse ordering could deadlock.
>
> .569196]
> other info that might help us debug this:
>
> [ 307.569516] Chain exists of:
> &adev->firmware.mutex --> &userq_mgr->userq_mutex --> &reset_domain->sem
>
> [ 307.570011] Possible unsafe locking scenario:
>
> [ 307.570250] CPU0 CPU1
> [ 307.570438] ---- ----
> [ 307.570624] lock(&reset_domain->sem);
> [ 307.570785] lock(&userq_mgr->userq_mutex);
> [ 307.571061] lock(&reset_domain->sem);
> [ 307.571320] lock(&adev->firmware.mutex);
> [ 307.571491]
> *** DEADLOCK ***
>
> Signed-off-by: Prike Liang <Prike.Liang@amd.com>
> ---
> drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 51 ++++++++++++++++++++---
> 1 file changed, 45 insertions(+), 6 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> index 2534e4a1a530..a7d5ca741a3b 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> @@ -422,9 +422,12 @@ static void amdgpu_userq_detach_doorbell(struct amdgpu_usermode_queue *queue)
> {
> struct amdgpu_device *adev = queue->userq_mgr->adev;
>
> - down_read(&adev->reset_domain->sem);
> + /*
> + * The caller serializes doorbell removal against an in-progress GPU
> + * reset by holding adev->reset_domain->sem for read.
> + */
> + lockdep_assert_held_read(&adev->reset_domain->sem);
> xa_erase_irq(&adev->userq_doorbell_xa, queue->doorbell_index);
> - up_read(&adev->reset_domain->sem);
> }
>
> /**
> @@ -544,11 +547,34 @@ amdgpu_userq_destroy(struct amdgpu_userq_mgr *uq_mgr, struct amdgpu_usermode_que
>
> cancel_delayed_work_sync(&uq_mgr->resume_work);
>
> + /*
> + * Cancel hang detection before serializing against a GPU reset. Hang
> + * detection triggers recovery, which takes reset_domain->sem for write,
> + * so it must not be canceled while that semaphore is held for read.
> + * A reset IRQ can restart hang detection, so this is repeated on retry.
> + */
> + cancel_delayed_work_sync(&queue->hang_detect_work);
> +retry:
> mutex_lock(&uq_mgr->userq_mutex);
> amdgpu_userq_wait_for_last_fence(queue);
>
> + /*
> + * Serialize queue teardown (doorbell detach and MES unmap) against an
> + * in-progress GPU reset. Do not block on the reset semaphore while
> + * holding userq_mutex: recovery takes the semaphore for write and then
> + * (transitively) userq_mutex, so blocking here would invert that order
> + * and deadlock. If the trylock fails, drop userq_mutex, wait for
> + * recovery to finish, and retry.
> + */
> + if (!down_read_trylock(&adev->reset_domain->sem)) {
> + mutex_unlock(&uq_mgr->userq_mutex);
> +
> + down_read(&adev->reset_domain->sem);
> + up_read(&adev->reset_domain->sem);
> + goto retry;
> + }
> +
> amdgpu_userq_detach_doorbell(queue);
> - cancel_delayed_work_sync(&queue->hang_detect_work);
>
> #if defined(CONFIG_DEBUG_FS)
> debugfs_remove_recursive(queue->debugfs_queue);
> @@ -557,6 +583,7 @@ amdgpu_userq_destroy(struct amdgpu_userq_mgr *uq_mgr, struct amdgpu_usermode_que
> atomic_dec(&uq_mgr->userq_count[queue->queue_type]);
> amdgpu_userq_fence_driver_free(queue);
> queue->fence_drv = NULL;
> + up_read(&adev->reset_domain->sem);
> mutex_unlock(&uq_mgr->userq_mutex);
>
> /*
> @@ -734,16 +761,28 @@ amdgpu_userq_create(struct drm_file *filp, union drm_amdgpu_userq *args)
> if (r)
> goto clean_mqd;
>
> +map_retry:
> amdgpu_userq_ensure_ev_fence(&fpriv->userq_mgr, &fpriv->evf_mgr);
>
> /* don't map the queue if scheduling is halted */
> if (!adev->userq_halt_for_enforce_isolation ||
> ((queue->queue_type != AMDGPU_HW_IP_GFX) &&
> (queue->queue_type != AMDGPU_HW_IP_COMPUTE))) {
> - /* Serialize the map against an in-progress GPU reset (MES is
> - * unresponsive during recovery), matching amdgpu_userq_detach_doorbell().
> + /*
> + * Serialize the map against an in-progress GPU reset (MES is
> + * unresponsive during recovery). Do not block on the reset
> + * semaphore while holding userq_mutex: recovery takes the
> + * semaphore for write and then (transitively) userq_mutex, so
> + * blocking here would invert that order and deadlock. If the
> + * trylock fails, drop userq_mutex, wait for recovery, and retry.
> */
> - down_read(&adev->reset_domain->sem);
> + if (!down_read_trylock(&adev->reset_domain->sem)) {
> + mutex_unlock(&uq_mgr->userq_mutex);
> +
> + down_read(&adev->reset_domain->sem);
> + up_read(&adev->reset_domain->sem);
> + goto map_retry;
> + }
> r = amdgpu_userq_map_helper(queue);
> up_read(&adev->reset_domain->sem);
> if (r) {
next prev parent reply other threads:[~2026-09-28 22:49 UTC|newest]
Thread overview: 28+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
2026-09-02 12:49 ` [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check Prike Liang
2026-09-03 19:29 ` Alex Deucher
2026-09-02 12:49 ` [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery Prike Liang
2026-09-13 20:48 ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 04/18] drm/amdgpu: move userq fence wait out of signalling section Prike Liang
2026-09-15 1:49 ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 05/18] drm/amdgpu: defer userq reset after eviction failure Prike Liang
2026-09-02 12:49 ` [PATCH 06/18] drm/amdgpu: skip DRM internal suspend/resume for reseting VKMS Prike Liang
2026-09-02 12:49 ` [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset Prike Liang
2026-09-13 20:53 ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 08/18] drm/amdgpu: skip VMHUB HW access in unaccessiable device Prike Liang
2026-09-02 12:49 ` [PATCH 09/18] drm/amdgpu/userq: complete the hang userq fence Prike Liang
2026-09-02 12:49 ` [PATCH 10/18] drm/amdgpu/mes: put the mes context BO allocation in mes sw_int Prike Liang
2026-09-02 12:49 ` [PATCH 11/18] drm/amdgpu: depart ring scheduler after resumming IP blocks Prike Liang
2026-09-02 12:49 ` [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock Prike Liang
2026-09-28 22:49 ` vitaly prosyak [this message]
2026-09-29 3:44 ` Liang, Prike
2026-09-29 8:56 ` Christian König
2026-09-02 12:49 ` [PATCH 13/18] drm/amdgpu: allocate dma_fence slot explicitly for rearming eviction fence Prike Liang
2026-09-02 12:49 ` [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset Prike Liang
2026-09-03 19:25 ` Alex Deucher
2026-09-02 12:49 ` [PATCH 15/18] drm/amdgpu/jpeg: skip scheduling jpeg/vcn idle_work " Prike Liang
2026-09-02 12:49 ` [PATCH 16/18] drm/amdgpu/mes: skip userq_notify_unmap during gpu reset Prike Liang
2026-09-02 12:50 ` [PATCH 17/18] drm/amdgpu/userq: complete userq eviction fence in pre_reset Prike Liang
2026-09-02 12:50 ` [PATCH 18/18] drm/amdgpu: stop the userq submission prior to removing userq Prike Liang
2026-09-03 19:28 ` [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Alex Deucher
2026-09-07 6:48 ` Liang, Prike
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=123fa18d-1bf7-48e7-bace-cdf429148128@amd.com \
--to=vprosyak@amd.com \
--cc=Alexander.Deucher@amd.com \
--cc=Christian.Koenig@amd.com \
--cc=Prike.Liang@amd.com \
--cc=Vitaly.Prosyak@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox