From: Prike Liang <Prike.Liang@amd.com>
To: <amd-gfx@lists.freedesktop.org>
Cc: <Alexander.Deucher@amd.com>, <Christian.Koenig@amd.com>,
<Vitaly.Prosyak@amd.com>, Prike Liang <Prike.Liang@amd.com>
Subject: [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock
Date: Wed, 2 Sep 2026 20:49:55 +0800 [thread overview]
Message-ID: <20260902125001.621629-12-Prike.Liang@amd.com> (raw)
In-Reply-To: <20260902125001.621629-1-Prike.Liang@amd.com>
The offending edge was code that did a blocking
down_read(&adev->reset_domain->sem) as following, while
holding userq_mutex. Since GPU recovery takes reset_domain
->sem for write and then transitively acquires userq_mutex,
the reverse ordering could deadlock.
.569196]
other info that might help us debug this:
[ 307.569516] Chain exists of:
&adev->firmware.mutex --> &userq_mgr->userq_mutex --> &reset_domain->sem
[ 307.570011] Possible unsafe locking scenario:
[ 307.570250] CPU0 CPU1
[ 307.570438] ---- ----
[ 307.570624] lock(&reset_domain->sem);
[ 307.570785] lock(&userq_mgr->userq_mutex);
[ 307.571061] lock(&reset_domain->sem);
[ 307.571320] lock(&adev->firmware.mutex);
[ 307.571491]
*** DEADLOCK ***
Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 51 ++++++++++++++++++++---
1 file changed, 45 insertions(+), 6 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
index 2534e4a1a530..a7d5ca741a3b 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
@@ -422,9 +422,12 @@ static void amdgpu_userq_detach_doorbell(struct amdgpu_usermode_queue *queue)
{
struct amdgpu_device *adev = queue->userq_mgr->adev;
- down_read(&adev->reset_domain->sem);
+ /*
+ * The caller serializes doorbell removal against an in-progress GPU
+ * reset by holding adev->reset_domain->sem for read.
+ */
+ lockdep_assert_held_read(&adev->reset_domain->sem);
xa_erase_irq(&adev->userq_doorbell_xa, queue->doorbell_index);
- up_read(&adev->reset_domain->sem);
}
/**
@@ -544,11 +547,34 @@ amdgpu_userq_destroy(struct amdgpu_userq_mgr *uq_mgr, struct amdgpu_usermode_que
cancel_delayed_work_sync(&uq_mgr->resume_work);
+ /*
+ * Cancel hang detection before serializing against a GPU reset. Hang
+ * detection triggers recovery, which takes reset_domain->sem for write,
+ * so it must not be canceled while that semaphore is held for read.
+ * A reset IRQ can restart hang detection, so this is repeated on retry.
+ */
+ cancel_delayed_work_sync(&queue->hang_detect_work);
+retry:
mutex_lock(&uq_mgr->userq_mutex);
amdgpu_userq_wait_for_last_fence(queue);
+ /*
+ * Serialize queue teardown (doorbell detach and MES unmap) against an
+ * in-progress GPU reset. Do not block on the reset semaphore while
+ * holding userq_mutex: recovery takes the semaphore for write and then
+ * (transitively) userq_mutex, so blocking here would invert that order
+ * and deadlock. If the trylock fails, drop userq_mutex, wait for
+ * recovery to finish, and retry.
+ */
+ if (!down_read_trylock(&adev->reset_domain->sem)) {
+ mutex_unlock(&uq_mgr->userq_mutex);
+
+ down_read(&adev->reset_domain->sem);
+ up_read(&adev->reset_domain->sem);
+ goto retry;
+ }
+
amdgpu_userq_detach_doorbell(queue);
- cancel_delayed_work_sync(&queue->hang_detect_work);
#if defined(CONFIG_DEBUG_FS)
debugfs_remove_recursive(queue->debugfs_queue);
@@ -557,6 +583,7 @@ amdgpu_userq_destroy(struct amdgpu_userq_mgr *uq_mgr, struct amdgpu_usermode_que
atomic_dec(&uq_mgr->userq_count[queue->queue_type]);
amdgpu_userq_fence_driver_free(queue);
queue->fence_drv = NULL;
+ up_read(&adev->reset_domain->sem);
mutex_unlock(&uq_mgr->userq_mutex);
/*
@@ -734,16 +761,28 @@ amdgpu_userq_create(struct drm_file *filp, union drm_amdgpu_userq *args)
if (r)
goto clean_mqd;
+map_retry:
amdgpu_userq_ensure_ev_fence(&fpriv->userq_mgr, &fpriv->evf_mgr);
/* don't map the queue if scheduling is halted */
if (!adev->userq_halt_for_enforce_isolation ||
((queue->queue_type != AMDGPU_HW_IP_GFX) &&
(queue->queue_type != AMDGPU_HW_IP_COMPUTE))) {
- /* Serialize the map against an in-progress GPU reset (MES is
- * unresponsive during recovery), matching amdgpu_userq_detach_doorbell().
+ /*
+ * Serialize the map against an in-progress GPU reset (MES is
+ * unresponsive during recovery). Do not block on the reset
+ * semaphore while holding userq_mutex: recovery takes the
+ * semaphore for write and then (transitively) userq_mutex, so
+ * blocking here would invert that order and deadlock. If the
+ * trylock fails, drop userq_mutex, wait for recovery, and retry.
*/
- down_read(&adev->reset_domain->sem);
+ if (!down_read_trylock(&adev->reset_domain->sem)) {
+ mutex_unlock(&uq_mgr->userq_mutex);
+
+ down_read(&adev->reset_domain->sem);
+ up_read(&adev->reset_domain->sem);
+ goto map_retry;
+ }
r = amdgpu_userq_map_helper(queue);
up_read(&adev->reset_domain->sem);
if (r) {
--
2.34.1
next prev parent reply other threads:[~2026-09-02 12:50 UTC|newest]
Thread overview: 28+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
2026-09-02 12:49 ` [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check Prike Liang
2026-09-03 19:29 ` Alex Deucher
2026-09-02 12:49 ` [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery Prike Liang
2026-09-13 20:48 ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 04/18] drm/amdgpu: move userq fence wait out of signalling section Prike Liang
2026-09-15 1:49 ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 05/18] drm/amdgpu: defer userq reset after eviction failure Prike Liang
2026-09-02 12:49 ` [PATCH 06/18] drm/amdgpu: skip DRM internal suspend/resume for reseting VKMS Prike Liang
2026-09-02 12:49 ` [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset Prike Liang
2026-09-13 20:53 ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 08/18] drm/amdgpu: skip VMHUB HW access in unaccessiable device Prike Liang
2026-09-02 12:49 ` [PATCH 09/18] drm/amdgpu/userq: complete the hang userq fence Prike Liang
2026-09-02 12:49 ` [PATCH 10/18] drm/amdgpu/mes: put the mes context BO allocation in mes sw_int Prike Liang
2026-09-02 12:49 ` [PATCH 11/18] drm/amdgpu: depart ring scheduler after resumming IP blocks Prike Liang
2026-09-02 12:49 ` Prike Liang [this message]
2026-09-28 22:49 ` [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock vitaly prosyak
2026-09-29 3:44 ` Liang, Prike
2026-09-29 8:56 ` Christian König
2026-09-02 12:49 ` [PATCH 13/18] drm/amdgpu: allocate dma_fence slot explicitly for rearming eviction fence Prike Liang
2026-09-02 12:49 ` [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset Prike Liang
2026-09-03 19:25 ` Alex Deucher
2026-09-02 12:49 ` [PATCH 15/18] drm/amdgpu/jpeg: skip scheduling jpeg/vcn idle_work " Prike Liang
2026-09-02 12:49 ` [PATCH 16/18] drm/amdgpu/mes: skip userq_notify_unmap during gpu reset Prike Liang
2026-09-02 12:50 ` [PATCH 17/18] drm/amdgpu/userq: complete userq eviction fence in pre_reset Prike Liang
2026-09-02 12:50 ` [PATCH 18/18] drm/amdgpu: stop the userq submission prior to removing userq Prike Liang
2026-09-03 19:28 ` [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Alex Deucher
2026-09-07 6:48 ` Liang, Prike
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260902125001.621629-12-Prike.Liang@amd.com \
--to=prike.liang@amd.com \
--cc=Alexander.Deucher@amd.com \
--cc=Christian.Koenig@amd.com \
--cc=Vitaly.Prosyak@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.