All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset
@ 2026-09-02 12:49 Prike Liang
  2026-09-02 12:49 ` [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check Prike Liang
                   ` (17 more replies)
  0 siblings, 18 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

amdgpu_mes_detect_and_reset_hung_queues() already detects
the guilty compute user queue and resets it through
mes_userq_reset_queue(). The additional reset via
mes_userq_reset() is unnecessary, so remove it to unify
the compute userq reset.

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c    | 5 -----
 drivers/gpu/drm/amd/amdgpu/mes_userqueue.c | 2 --
 2 files changed, 7 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c
index a6f95ff47d24..5c3be851ac84 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c
@@ -2384,11 +2384,6 @@ int amdgpu_gfx_reset_mes_compute(struct amdgpu_device *adev,
 		deferred_end[n_deferred].fence = guilty_fence;
 		n_deferred++;
 	}
-	if (uq) {
-		r = mes_userq_reset(uq);
-		if (r)
-			goto out;
-	}
 	for (i = 0; i < num_hung; i++) {
 		struct amdgpu_ring *hr = NULL;
 		struct amdgpu_fence *hf = NULL;
diff --git a/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c b/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
index 7f334f718cd8..83a438d8e117 100644
--- a/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
+++ b/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
@@ -249,8 +249,6 @@ int mes_userq_reset_queue(struct amdgpu_device *adev,
 
 	xa_for_each(&adev->userq_doorbell_xa, uq_id, uq) {
 		if (uq->queue_type == queue_type) {
-			if (uq == guilty_uq)
-				continue;
 			if (uq->doorbell_index == db) {
 				uq->state = AMDGPU_USERQ_STATE_HUNG;
 				if (use_mmio)
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-03 19:29   ` Alex Deucher
  2026-09-02 12:49 ` [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery Prike Liang
                   ` (16 subsequent siblings)
  17 siblings, 1 reply; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

If the userq doesn't support in a system. then there's no
valid userq_doorbell_xa entry to walk over and then has a
no-op.

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 8 --------
 1 file changed, 8 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
index 0a816b3c5ff9..ed329041a648 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
@@ -1373,15 +1373,11 @@ void amdgpu_userq_mgr_fini(struct amdgpu_userq_mgr *userq_mgr)
 
 int amdgpu_userq_suspend(struct amdgpu_device *adev)
 {
-	u32 ip_mask = amdgpu_userq_get_supported_ip_mask(adev);
 	struct amdgpu_usermode_queue *queue;
 	struct amdgpu_userq_mgr *uqm;
 	unsigned long queue_id;
 	int r;
 
-	if (!ip_mask)
-		return 0;
-
 	xa_for_each(&adev->userq_doorbell_xa, queue_id, queue) {
 		uqm = queue->userq_mgr;
 		cancel_delayed_work_sync(&uqm->resume_work);
@@ -1398,15 +1394,11 @@ int amdgpu_userq_suspend(struct amdgpu_device *adev)
 
 int amdgpu_userq_resume(struct amdgpu_device *adev)
 {
-	u32 ip_mask = amdgpu_userq_get_supported_ip_mask(adev);
 	struct amdgpu_usermode_queue *queue;
 	struct amdgpu_userq_mgr *uqm;
 	unsigned long queue_id;
 	int r;
 
-	if (!ip_mask)
-		return 0;
-
 	xa_for_each(&adev->userq_doorbell_xa, queue_id, queue) {
 		uqm = queue->userq_mgr;
 		guard(mutex)(&uqm->userq_mutex);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
  2026-09-02 12:49 ` [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-13 20:48   ` vitaly prosyak
  2026-09-02 12:49 ` [PATCH 04/18] drm/amdgpu: move userq fence wait out of signalling section Prike Liang
                   ` (15 subsequent siblings)
  17 siblings, 1 reply; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

Suspend the drm internal clients has a deadlock risk as acquiring
it while holding the reset domain lock inverts the ordering
established elsewhere (clientlist_mutex -> ... -> reset_domain->sem).

Reset AMDGPU can prevent the user space clients further accessing by
using the reset semaphore, so removing the drm_client_dev_suspend() |
resume() in the reset path.

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_device.c | 4 ----
 1 file changed, 4 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
index d7640da9f6de..bd4eb97336b1 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
@@ -5210,8 +5210,6 @@ int amdgpu_device_reinit_after_reset(struct amdgpu_reset_context *reset_context)
 				if (r)
 					goto out;
 
-				drm_client_dev_resume(adev_to_drm(tmp_adev));
-
 				/*
 				 * The GPU enters bad state once faulty pages
 				 * by ECC has reached the threshold, and ras
@@ -5544,8 +5542,6 @@ static void amdgpu_device_halt_activities(struct amdgpu_device *adev,
 		 */
 		amdgpu_unregister_gpu_instance(tmp_adev);
 
-		drm_client_dev_suspend(adev_to_drm(tmp_adev));
-
 		/* disable ras on ALL IPs */
 		if (!need_emergency_restart && !amdgpu_reset_in_dpc(adev))
 			amdgpu_ras_suspend(tmp_adev);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 04/18] drm/amdgpu: move userq fence wait out of signalling section
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
  2026-09-02 12:49 ` [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check Prike Liang
  2026-09-02 12:49 ` [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-02 12:49 ` [PATCH 05/18] drm/amdgpu: defer userq reset after eviction failure Prike Liang
                   ` (14 subsequent siblings)
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

The eviction fence suspend worker waits for every pending userq fence
from inside a dma_fence_begin_signalling() critical section. Waiting on
another DMA fence while responsible for signalling one violates the
cross-driver fence contract and is reported by lockdep as a
dma_fence_map dependency.

Move the wait before dma_fence_begin_signalling(). Keep userq_mutex held
so queue lifetime remains stable while inspecting last_fence.

Fixes: fc61df151617 ("drm/amdgpu: annotate eviction fence signaling path")
Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c | 3 +++
 drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c          | 4 +---
 drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h          | 1 +
 3 files changed, 5 insertions(+), 3 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
index 4c5e38dea4c2..a0802012de49 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
@@ -68,6 +68,9 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct *work)
 
 	mutex_lock(&uq_mgr->userq_mutex);
 
+	/* Fence waits are not allowed in a fence signalling critical section. */
+	amdgpu_userq_wait_for_signal(uq_mgr);
+
 	/*
 	 * This is intentionally after taking the userq_mutex since we do
 	 * allocate memory while holding this lock, but only after ensuring that
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
index ed329041a648..a7b68fd2360e 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
@@ -1272,7 +1272,7 @@ amdgpu_userq_evict_all(struct amdgpu_userq_mgr *uq_mgr)
 	return ret;
 }
 
-static void
+void
 amdgpu_userq_wait_for_signal(struct amdgpu_userq_mgr *uq_mgr)
 {
 	struct amdgpu_usermode_queue *queue;
@@ -1291,8 +1291,6 @@ amdgpu_userq_wait_for_signal(struct amdgpu_userq_mgr *uq_mgr)
 void
 amdgpu_userq_evict(struct amdgpu_userq_mgr *uq_mgr)
 {
-	/* Wait for any pending userqueue fence work to finish */
-	amdgpu_userq_wait_for_signal(uq_mgr);
 	amdgpu_userq_evict_all(uq_mgr);
 }
 
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
index 6412a7f7b6ef..488dc21d7c81 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
@@ -162,6 +162,7 @@ void amdgpu_userq_mgr_cancel_reset_work(struct amdgpu_device *adev);
 void amdgpu_userq_mgr_cancel_resume(struct amdgpu_userq_mgr *userq_mgr);
 void amdgpu_userq_mgr_fini(struct amdgpu_userq_mgr *userq_mgr);
 
+void amdgpu_userq_wait_for_signal(struct amdgpu_userq_mgr *uq_mgr);
 void amdgpu_userq_evict(struct amdgpu_userq_mgr *uq_mgr);
 
 void amdgpu_userq_ensure_ev_fence(struct amdgpu_userq_mgr *userq_mgr,
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 05/18] drm/amdgpu: defer userq reset after eviction failure
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (2 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 04/18] drm/amdgpu: move userq fence wait out of signalling section Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-02 12:49 ` [PATCH 06/18] drm/amdgpu: skip DRM internal suspend/resume for reseting VKMS Prike Liang
                   ` (13 subsequent siblings)
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

When userq eviction fails, the eviction fence suspend worker flushes
reset_work while holding userq_mutex and from inside a DMA-fence
signalling critical section.

The reset worker runs GPU recovery, which calls amdgpu_userq_suspend()
and tries to acquire the same userq_mutex. The suspend worker therefore
waits for reset_work while reset_work waits for the suspend worker to
release the mutex.

Propagate the eviction error to the suspend worker, set the error on the
eviction fence and do not schedule queue restore. Signal the failed
fence, release userq_mutex, and then run reset recovery synchronously.
This preserves teardown ordering without holding userq_mutex across the
reset.

Fixes: c8ed2de0f2ee ("drm/amdgpu: rework userq reset work handling")
Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c | 13 +++++++++++--
 drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c          |  7 ++-----
 drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h          |  2 +-
 3 files changed, 14 insertions(+), 8 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
index a0802012de49..2ea8553c82f0 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
@@ -65,6 +65,7 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct *work)
 	struct amdgpu_userq_mgr *uq_mgr = &fpriv->userq_mgr;
 	struct dma_fence *ev_fence;
 	bool cookie;
+	int r;
 
 	mutex_lock(&uq_mgr->userq_mutex);
 
@@ -79,7 +80,9 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct *work)
 	cookie = dma_fence_begin_signalling();
 
 	ev_fence = amdgpu_evf_mgr_get_fence(evf_mgr);
-	amdgpu_userq_evict(uq_mgr);
+	r = amdgpu_userq_evict(uq_mgr);
+	if (r)
+		dma_fence_set_error(ev_fence, r);
 
 	/*
 	 * Signaling the eviction fence must be done while holding the
@@ -90,10 +93,16 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct *work)
 	dma_fence_end_signalling(cookie);
 	dma_fence_put(ev_fence);
 
-	if (!evf_mgr->shutdown)
+	if (!r && !evf_mgr->shutdown)
 		schedule_delayed_work(&uq_mgr->resume_work, 0);
 
 	mutex_unlock(&uq_mgr->userq_mutex);
+
+	if (r) {
+		amdgpu_reset_domain_schedule(uq_mgr->adev->reset_domain,
+					     &uq_mgr->reset_work);
+		flush_work(&uq_mgr->reset_work);
+	}
 }
 
 int amdgpu_evf_mgr_attach_fence(struct amdgpu_eviction_fence_mgr *evf_mgr,
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
index a7b68fd2360e..2534e4a1a530 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
@@ -1265,9 +1265,6 @@ amdgpu_userq_evict_all(struct amdgpu_userq_mgr *uq_mgr)
 	if (ret) {
 		drm_file_err(uq_mgr->file,
 			     "Couldn't unmap all the queues, eviction failed ret=%d\n", ret);
-		amdgpu_reset_domain_schedule(uq_mgr->adev->reset_domain,
-					     &uq_mgr->reset_work);
-		flush_work(&uq_mgr->reset_work);
 	}
 	return ret;
 }
@@ -1288,10 +1285,10 @@ amdgpu_userq_wait_for_signal(struct amdgpu_userq_mgr *uq_mgr)
 	}
 }
 
-void
+int
 amdgpu_userq_evict(struct amdgpu_userq_mgr *uq_mgr)
 {
-	amdgpu_userq_evict_all(uq_mgr);
+	return amdgpu_userq_evict_all(uq_mgr);
 }
 
 int amdgpu_userq_mgr_init(struct amdgpu_userq_mgr *userq_mgr, struct drm_file *file_priv,
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
index 488dc21d7c81..4dcf6151de6a 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h
@@ -163,7 +163,7 @@ void amdgpu_userq_mgr_cancel_resume(struct amdgpu_userq_mgr *userq_mgr);
 void amdgpu_userq_mgr_fini(struct amdgpu_userq_mgr *userq_mgr);
 
 void amdgpu_userq_wait_for_signal(struct amdgpu_userq_mgr *uq_mgr);
-void amdgpu_userq_evict(struct amdgpu_userq_mgr *uq_mgr);
+int amdgpu_userq_evict(struct amdgpu_userq_mgr *uq_mgr);
 
 void amdgpu_userq_ensure_ev_fence(struct amdgpu_userq_mgr *userq_mgr,
 				  struct amdgpu_eviction_fence_mgr *evf_mgr);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 06/18] drm/amdgpu: skip DRM internal suspend/resume for reseting VKMS
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (3 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 05/18] drm/amdgpu: defer userq reset after eviction failure Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-02 12:49 ` [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset Prike Liang
                   ` (12 subsequent siblings)
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

Skip duplicate VKMS mode-config suspend/resume while an internal
reset is in progress. This removes the remaining reset_domain->sem
-> clientlist_mutex edge without changing the lock order documented
by amdgpu_lockdep_init().

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_vkms.c | 6 ++++++
 1 file changed, 6 insertions(+)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_vkms.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_vkms.c
index d6b51b3f216b..a49c7dda4c4e 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_vkms.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_vkms.c
@@ -491,6 +491,9 @@ static int amdgpu_vkms_suspend(struct amdgpu_ip_block *ip_block)
 	struct amdgpu_device *adev = ip_block->adev;
 	int r;
 
+	if (amdgpu_in_reset(adev))
+		return 0;
+
 	r = drm_mode_config_helper_suspend(adev_to_drm(adev));
 	if (r)
 		return r;
@@ -505,6 +508,9 @@ static int amdgpu_vkms_resume(struct amdgpu_ip_block *ip_block)
 	r = amdgpu_vkms_hw_init(ip_block);
 	if (r)
 		return r;
+	if (amdgpu_in_reset(ip_block->adev))
+		return 0;
+
 	return drm_mode_config_helper_resume(adev_to_drm(ip_block->adev));
 }
 
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (4 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 06/18] drm/amdgpu: skip DRM internal suspend/resume for reseting VKMS Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-13 20:53   ` vitaly prosyak
  2026-09-02 12:49 ` [PATCH 08/18] drm/amdgpu: skip VMHUB HW access in unaccessiable device Prike Liang
                   ` (11 subsequent siblings)
  17 siblings, 1 reply; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

The eviction fence suspend worker can submit MES REMOVE_QUEUE packets
without holding the reset-domain semaphore. If GPU recovery starts while
the worker is running, both paths can access the hardware concurrently.
This triggers the hardware-access lockdep assertion and can submit a MES
packet while recovery is resetting the device.

Try to take the reset-domain semaphore for read around userq eviction. Do
not block on it while holding userq_mutex because recovery takes the reset
semaphore for write before acquiring buffer reservations and userq_mutex.
Instead, drop userq_mutex, wait for recovery without holding any other
lock, and retry the queue-state checks after recovery completes.

This makes recovery wait for an in-flight MES eviction, while an eviction
which starts after recovery waits without introducing the reverse lock
dependency that caused the reported circular-lock warning.

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c | 9 +++++++++
 1 file changed, 9 insertions(+)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
index 2ea8553c82f0..93307cbf55dc 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
@@ -67,10 +67,18 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct *work)
 	bool cookie;
 	int r;
 
+retry:
 	mutex_lock(&uq_mgr->userq_mutex);
 
 	/* Fence waits are not allowed in a fence signalling critical section. */
 	amdgpu_userq_wait_for_signal(uq_mgr);
+	if (!down_read_trylock(&uq_mgr->adev->reset_domain->sem)) {
+		mutex_unlock(&uq_mgr->userq_mutex);
+
+		down_read(&uq_mgr->adev->reset_domain->sem);
+		up_read(&uq_mgr->adev->reset_domain->sem);
+		goto retry;
+	}
 
 	/*
 	 * This is intentionally after taking the userq_mutex since we do
@@ -81,6 +89,7 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct *work)
 
 	ev_fence = amdgpu_evf_mgr_get_fence(evf_mgr);
 	r = amdgpu_userq_evict(uq_mgr);
+	up_read(&uq_mgr->adev->reset_domain->sem);
 	if (r)
 		dma_fence_set_error(ev_fence, r);
 
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 08/18] drm/amdgpu: skip VMHUB HW access in unaccessiable device
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (5 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-02 12:49 ` [PATCH 09/18] drm/amdgpu/userq: complete the hang userq fence Prike Liang
                   ` (10 subsequent siblings)
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

Only try to access the VMHUB HW and log the page faut detail
when the VMHUB not is an intermediate HW state.

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c | 6 ++++++
 1 file changed, 6 insertions(+)

diff --git a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
index f454aff831b0..c4448a6a86bc 100644
--- a/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/gmc_v11_0.c
@@ -113,6 +113,12 @@ static int gmc_v11_0_process_interrupt(struct amdgpu_device *adev,
 	addr = (u64)entry->src_data[0] << 12;
 	addr |= ((u64)entry->src_data[1] & 0xf) << 44;
 
+	if (amdgpu_in_reset(adev) || adev->no_hw_access) {
+		dev_err(adev->dev, " [inaccessiable device] page fault address 0x%016llx from client %d\n",
+			addr, entry->client_id);
+		return 0;
+	}
+
 	if (retry_fault) {
 		int ret = amdgpu_gmc_handle_retry_fault(adev, entry, addr, 0, 0,
 							write_fault);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 09/18] drm/amdgpu/userq: complete the hang userq fence
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (6 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 08/18] drm/amdgpu: skip VMHUB HW access in unaccessiable device Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-02 12:49 ` [PATCH 10/18] drm/amdgpu/mes: put the mes context BO allocation in mes sw_int Prike Liang
                   ` (9 subsequent siblings)
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

Complete the hang user queue fence unconditionally, including
the case where queue reset fails. Without this fix, a failed
reset leaves the fence unsignaled, causing waiters to block
indefinitely.

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/mes_userqueue.c | 8 +++-----
 1 file changed, 3 insertions(+), 5 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c b/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
index 83a438d8e117..2bb5f231f133 100644
--- a/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
+++ b/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
@@ -255,12 +255,10 @@ int mes_userq_reset_queue(struct amdgpu_device *adev,
 					r = amdgpu_mes_reset_queue_mmio(adev, queue_type, 0, 1, pipe, queue, 0);
 				else
 					r = amdgpu_mes_reset_user_queue(adev, queue_type, db, 0);
-				if (r)
-					return r;
-				r = mes_userq_unmap(uq);
-				if (r)
-					return r;
+				if (!r)
+					r = mes_userq_unmap(uq);
 				amdgpu_userq_fence_driver_force_completion(uq);
+				return r;
 				break;
 			}
 		}
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 10/18] drm/amdgpu/mes: put the mes context BO allocation in mes sw_int
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (7 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 09/18] drm/amdgpu/userq: complete the hang userq fence Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-02 12:49 ` [PATCH 11/18] drm/amdgpu: depart ring scheduler after resumming IP blocks Prike Liang
                   ` (8 subsequent siblings)
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

It's more sense to put the mes context BO allocation in the mes
sw_init, and this also can resolve the risk deadlock issue between
bo allocation reservation_ww_class_mutex and mutex lock voliation.

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/mes_v11_0.c | 20 ++++++++++----------
 1 file changed, 10 insertions(+), 10 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/mes_v11_0.c b/drivers/gpu/drm/amd/amdgpu/mes_v11_0.c
index 33ff1afd7c4c..b755d9239888 100644
--- a/drivers/gpu/drm/amd/amdgpu/mes_v11_0.c
+++ b/drivers/gpu/drm/amd/amdgpu/mes_v11_0.c
@@ -1734,6 +1734,7 @@ static int mes_v11_0_sw_init(struct amdgpu_ip_block *ip_block)
 	struct amdgpu_device *adev = ip_block->adev;
 	int pipe, r, bo_size;
 
+	adev->mes.use_rs64mem = false;
 	adev->mes.funcs = &mes_v11_0_funcs;
 	adev->mes.kiq_hw_init = &mes_v11_0_kiq_hw_init;
 	adev->mes.kiq_hw_fini = &mes_v11_0_kiq_hw_fini;
@@ -1784,6 +1785,15 @@ static int mes_v11_0_sw_init(struct amdgpu_ip_block *ip_block)
 		return r;
 	}
 
+	/* Allocate GPU buffer for array size query results */
+	r = amdgpu_mes_rs64mem_init(&adev->mes);
+	if (r) {
+		dev_warn(adev->dev,
+			"RS64 local memory init failed (%d),"
+			"falling back to system memory path\n", r);
+		adev->mes.use_rs64mem = false;
+	}
+
 	return 0;
 }
 
@@ -1961,8 +1971,6 @@ static int mes_v11_0_hw_init(struct amdgpu_ip_block *ip_block)
 	if (adev->mes.ring[0].sched.ready)
 		goto out;
 
-	adev->mes.use_rs64mem = false;
-
 	if (!adev->enable_mes_kiq) {
 		if (adev->firmware.load_type == AMDGPU_FW_LOAD_DIRECT) {
 			r = mes_v11_0_load_microcode(adev,
@@ -1980,14 +1988,6 @@ static int mes_v11_0_hw_init(struct amdgpu_ip_block *ip_block)
 	if (r)
 		goto failure;
 
-	/* Allocate GPU buffer for array size query results */
-	r = amdgpu_mes_rs64mem_init(&adev->mes);
-	if (r) {
-		dev_warn(adev->dev,
-			 "RS64 local memory init failed (%d),"
-			 "falling back to system memory path\n", r);
-		adev->mes.use_rs64mem = false;
-	}
 	r = mes_v11_0_set_hw_resources(&adev->mes);
 
 	if (r)
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 11/18] drm/amdgpu: depart ring scheduler after resumming IP blocks
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (8 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 10/18] drm/amdgpu/mes: put the mes context BO allocation in mes sw_int Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-02 12:49 ` [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock Prike Liang
                   ` (7 subsequent siblings)
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

During GPU reset, the ring scheduler should be parked
until all IP blocks have been resumed.

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_device.c | 13 +++++++------
 1 file changed, 7 insertions(+), 6 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
index bd4eb97336b1..a23950f3f634 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
@@ -5793,8 +5793,10 @@ int amdgpu_device_gpu_recover(struct amdgpu_device *adev,
 
 	amdgpu_device_halt_activities(adev, job, reset_context, &device_list,
 				      hive, need_emergency_restart);
-	if (need_emergency_restart)
-		goto skip_sched_resume;
+	if (need_emergency_restart) {
+		amdgpu_device_gpu_resume(adev, &device_list, need_emergency_restart);
+		goto reset_unlock;
+	}
 	/*
 	 * Must check guilty signal here since after this point all old
 	 * HW fences are force signaled.
@@ -5804,18 +5806,17 @@ int amdgpu_device_gpu_recover(struct amdgpu_device *adev,
 	if (job && (dma_fence_get_status(&job->hw_fence->base) > 0)) {
 		job_signaled = true;
 		dev_info(adev->dev, "Guilty job already signaled, skipping HW reset");
-		goto skip_hw_reset;
+		goto post_resume;
 	}
 
 	r = amdgpu_device_asic_reset(adev, &device_list, reset_context);
 	if (r)
 		goto reset_unlock;
-skip_hw_reset:
+post_resume:
+	amdgpu_device_gpu_resume(adev, &device_list, need_emergency_restart);
 	r = amdgpu_device_sched_resume(&device_list, reset_context, job_signaled);
 	if (r)
 		goto reset_unlock;
-skip_sched_resume:
-	amdgpu_device_gpu_resume(adev, &device_list, need_emergency_restart);
 reset_unlock:
 	amdgpu_device_recovery_put_reset_lock(adev, &device_list);
 	amdgpu_ras_post_reset(adev, &device_list);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (9 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 11/18] drm/amdgpu: depart ring scheduler after resumming IP blocks Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-02 12:49 ` [PATCH 13/18] drm/amdgpu: allocate dma_fence slot explicitly for rearming eviction fence Prike Liang
                   ` (6 subsequent siblings)
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

The offending edge was code that did a blocking
down_read(&adev->reset_domain->sem) as following, while
holding userq_mutex. Since GPU recovery takes reset_domain
->sem for write and then transitively acquires userq_mutex,
the reverse ordering could deadlock.

.569196]
               other info that might help us debug this:

[  307.569516] Chain exists of:
                 &adev->firmware.mutex --> &userq_mgr->userq_mutex --> &reset_domain->sem

[  307.570011]  Possible unsafe locking scenario:

[  307.570250]        CPU0                    CPU1
[  307.570438]        ----                    ----
[  307.570624]   lock(&reset_domain->sem);
[  307.570785]                                lock(&userq_mgr->userq_mutex);
[  307.571061]                                lock(&reset_domain->sem);
[  307.571320]   lock(&adev->firmware.mutex);
[  307.571491]
                *** DEADLOCK ***

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 51 ++++++++++++++++++++---
 1 file changed, 45 insertions(+), 6 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
index 2534e4a1a530..a7d5ca741a3b 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
@@ -422,9 +422,12 @@ static void amdgpu_userq_detach_doorbell(struct amdgpu_usermode_queue *queue)
 {
 	struct amdgpu_device *adev = queue->userq_mgr->adev;
 
-	down_read(&adev->reset_domain->sem);
+	/*
+	 * The caller serializes doorbell removal against an in-progress GPU
+	 * reset by holding adev->reset_domain->sem for read.
+	 */
+	lockdep_assert_held_read(&adev->reset_domain->sem);
 	xa_erase_irq(&adev->userq_doorbell_xa, queue->doorbell_index);
-	up_read(&adev->reset_domain->sem);
 }
 
 /**
@@ -544,11 +547,34 @@ amdgpu_userq_destroy(struct amdgpu_userq_mgr *uq_mgr, struct amdgpu_usermode_que
 
 	cancel_delayed_work_sync(&uq_mgr->resume_work);
 
+	/*
+	 * Cancel hang detection before serializing against a GPU reset. Hang
+	 * detection triggers recovery, which takes reset_domain->sem for write,
+	 * so it must not be canceled while that semaphore is held for read.
+	 * A reset IRQ can restart hang detection, so this is repeated on retry.
+	 */
+	cancel_delayed_work_sync(&queue->hang_detect_work);
+retry:
 	mutex_lock(&uq_mgr->userq_mutex);
 	amdgpu_userq_wait_for_last_fence(queue);
 
+	/*
+	 * Serialize queue teardown (doorbell detach and MES unmap) against an
+	 * in-progress GPU reset. Do not block on the reset semaphore while
+	 * holding userq_mutex: recovery takes the semaphore for write and then
+	 * (transitively) userq_mutex, so blocking here would invert that order
+	 * and deadlock. If the trylock fails, drop userq_mutex, wait for
+	 * recovery to finish, and retry.
+	 */
+	if (!down_read_trylock(&adev->reset_domain->sem)) {
+		mutex_unlock(&uq_mgr->userq_mutex);
+
+		down_read(&adev->reset_domain->sem);
+		up_read(&adev->reset_domain->sem);
+		goto retry;
+	}
+
 	amdgpu_userq_detach_doorbell(queue);
-	cancel_delayed_work_sync(&queue->hang_detect_work);
 
 #if defined(CONFIG_DEBUG_FS)
 	debugfs_remove_recursive(queue->debugfs_queue);
@@ -557,6 +583,7 @@ amdgpu_userq_destroy(struct amdgpu_userq_mgr *uq_mgr, struct amdgpu_usermode_que
 	atomic_dec(&uq_mgr->userq_count[queue->queue_type]);
 	amdgpu_userq_fence_driver_free(queue);
 	queue->fence_drv = NULL;
+	up_read(&adev->reset_domain->sem);
 	mutex_unlock(&uq_mgr->userq_mutex);
 
 	/*
@@ -734,16 +761,28 @@ amdgpu_userq_create(struct drm_file *filp, union drm_amdgpu_userq *args)
 	if (r)
 		goto clean_mqd;
 
+map_retry:
 	amdgpu_userq_ensure_ev_fence(&fpriv->userq_mgr, &fpriv->evf_mgr);
 
 	/* don't map the queue if scheduling is halted */
 	if (!adev->userq_halt_for_enforce_isolation ||
 	    ((queue->queue_type != AMDGPU_HW_IP_GFX) &&
 	     (queue->queue_type != AMDGPU_HW_IP_COMPUTE))) {
-		/* Serialize the map against an in-progress GPU reset (MES is
-		 * unresponsive during recovery), matching amdgpu_userq_detach_doorbell().
+		/*
+		 * Serialize the map against an in-progress GPU reset (MES is
+		 * unresponsive during recovery). Do not block on the reset
+		 * semaphore while holding userq_mutex: recovery takes the
+		 * semaphore for write and then (transitively) userq_mutex, so
+		 * blocking here would invert that order and deadlock. If the
+		 * trylock fails, drop userq_mutex, wait for recovery, and retry.
 		 */
-		down_read(&adev->reset_domain->sem);
+		if (!down_read_trylock(&adev->reset_domain->sem)) {
+			mutex_unlock(&uq_mgr->userq_mutex);
+
+			down_read(&adev->reset_domain->sem);
+			up_read(&adev->reset_domain->sem);
+			goto map_retry;
+		}
 		r = amdgpu_userq_map_helper(queue);
 		up_read(&adev->reset_domain->sem);
 		if (r) {
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 13/18] drm/amdgpu: allocate dma_fence slot explicitly for rearming eviction fence
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (10 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-02 12:49 ` [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset Prike Liang
                   ` (5 subsequent siblings)
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

It needn't to explicitly reserve the dma_resv slot here, the slot
should ideally be reserved at the call site (e.g., amdgpu_userq_vm_
validate_and_restore_queue()). This patch is needed as a workaround
in the meantime

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c | 10 +++++++---
 1 file changed, 7 insertions(+), 3 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
index 93307cbf55dc..11ae2558f3fe 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
@@ -126,9 +126,13 @@ int amdgpu_evf_mgr_attach_fence(struct amdgpu_eviction_fence_mgr *evf_mgr,
 
 		amdgpu_bo_placement_from_domain(bo, bo->allowed_domains);
 		ret = ttm_bo_validate(&bo->tbo, &bo->placement, &ctx);
-		if (!ret)
-			dma_resv_add_fence(resv, ev_fence,
-					   DMA_RESV_USAGE_BOOKKEEP);
+		if (!ret) {
+			/*TODO: Figure out where the reserve fence slots are missing. */
+			ret = dma_resv_reserve_fences(resv, 1);
+			if (!ret)
+				dma_resv_add_fence(resv, ev_fence,
+						   DMA_RESV_USAGE_BOOKKEEP);
+		}
 	} else {
 		ret = 0;
 	}
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (11 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 13/18] drm/amdgpu: allocate dma_fence slot explicitly for rearming eviction fence Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-03 19:25   ` Alex Deucher
  2026-09-02 12:49 ` [PATCH 15/18] drm/amdgpu/jpeg: skip scheduling jpeg/vcn idle_work " Prike Liang
                   ` (4 subsequent siblings)
  17 siblings, 1 reply; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

During resume from GPU reset, the gfx idle work may invoke switch_power_profile
before the reset completes. This causes the following assert error because the
register access occurs without first releasing the GPU reset semaphore:

[ 1576.768935] CR2: 0000559ea133ead0 CR3: 00000002e6c42000 CR4: 0000000000350ef0
[ 1576.768940] Call Trace:
[ 1576.768944]  <TASK>
[ 1576.768953]  amdgpu_device_rreg+0x21/0x50 [amdgpu]
[ 1576.769158]  smu_msg_v1_send_msg+0x1a4/0x6e0 [amdgpu]
[ 1576.769437]  smu_cmn_send_smc_msg_with_params_ext+0xba/0x120 [amdgpu]
[ 1576.769721]  smu_cmn_send_smc_msg_with_param+0x33/0x40 [amdgpu]
[ 1576.769993]  smu_v13_0_0_set_power_profile_mode+0x192/0x2b0 [amdgpu]
[ 1576.770267]  smu_bump_power_profile_mode+0x5d/0x80 [amdgpu]
[ 1576.770538]  smu_switch_power_profile+0xa4/0xf0 [amdgpu]
[ 1576.770839]  amdgpu_dpm_switch_power_profile+0x6f/0x90 [amdgpu]
[ 1576.771210]  amdgpu_gfx_profile_idle_work_handler+0xe9/0x130 [amdgpu]
[ 1576.771460]  process_one_work+0x23e/0x6f0
[ 1576.771491]  worker_thread+0x1c4/0x380
[ 1576.771506]  kthread+0x10c/0x150
[ 1576.771512]  ? __pfx_worker_thread+0x10/0x10
[ 1576.771518]  ? __pfx_kthread+0x10/0x10
[ 1576.771530]  ret_from_fork+0x314/0x390
[ 1576.771537]  ? __pfx_kthread+0x10/0x10
[ 1576.771546]  ret_from_fork_asm+0x1a/0x30

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/pm/amdgpu_dpm.c | 6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

diff --git a/drivers/gpu/drm/amd/pm/amdgpu_dpm.c b/drivers/gpu/drm/amd/pm/amdgpu_dpm.c
index ce526db4d24a..808be6c425bf 100644
--- a/drivers/gpu/drm/amd/pm/amdgpu_dpm.c
+++ b/drivers/gpu/drm/amd/pm/amdgpu_dpm.c
@@ -348,7 +348,8 @@ int amdgpu_dpm_switch_power_profile(struct amdgpu_device *adev,
 	const struct amd_pm_funcs *pp_funcs = adev->powerplay.pp_funcs;
 	int ret = 0;
 
-	if (amdgpu_sriov_vf(adev))
+	if (amdgpu_sriov_vf(adev) ||
+		amdgpu_in_reset(adev))
 		return 0;
 
 	if (pp_funcs && pp_funcs->switch_power_profile) {
@@ -367,7 +368,8 @@ int amdgpu_dpm_pause_power_profile(struct amdgpu_device *adev,
 	const struct amd_pm_funcs *pp_funcs = adev->powerplay.pp_funcs;
 	int ret = 0;
 
-	if (amdgpu_sriov_vf(adev))
+	if (amdgpu_sriov_vf(adev) ||
+		amdgpu_in_reset(adev))
 		return 0;
 
 	if (pp_funcs && pp_funcs->pause_power_profile) {
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 15/18] drm/amdgpu/jpeg: skip scheduling jpeg/vcn idle_work during GPU reset
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (12 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-02 12:49 ` [PATCH 16/18] drm/amdgpu/mes: skip userq_notify_unmap during gpu reset Prike Liang
                   ` (3 subsequent siblings)
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

Scheduling jpeg idle_work for JPEG/VCN power gating during GPU reset can
trigger a register access assert before the GPU reset semaphore is
released, causing the following error during reset resume:

210] Workqueue: events amdgpu_jpeg_idle_work_handler [amdgpu]
[ 1576.787453] RIP: 0010:amdgpu_device_skip_hw_access+0x73/0x90 [amdgpu]
[ 1576.787656] Code: 85 c0 75 2a 8b 05 71 56 97 f0 85 c0 74 d3 48 8b bb d0 e7 07 00 be ff ff ff ff 48 81 c7 88 00 00 00 e8 f1 50 39 ef 85 c0 75 b7 <0f> 0b eb b3 48 8b bb d0 e7 07 00 48 83 c7 18 e8 39 fe 36 ee eb a1
[ 1576.787661] RSP: 0018:ffffccf700e6bcf0 EFLAGS: 00010246
[ 1576.787668] RAX: 0000000000000000 RBX: ffff89cf52d80000 RCX: 0000000000000002
[ 1576.787673] RDX: 0000000000000000 RSI: ffff89cf04f56bc8 RDI: ffff89cf02fe8f98
[ 1576.787677] RBP: ffffccf700e6bd00 R08: 0000000000000000 R09: 0000000000000001
[ 1576.787681] R10: ffffccf700e6bdb0 R11: ffffffffc1635c96 R12: 0000000000000000
[ 1576.787686] R13: 00000000000084d2 R14: 000000000003ff01 R15: 0000000000000000
[ 1576.787690] FS:  0000000000000000(0000) GS:ffff89d29a37d000(0000) knlGS:0000000000000000
[ 1576.787695] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 1576.787700] CR2: 00007fea0cc1ca50 CR3: 00000002e6c42000 CR4: 0000000000350ef0
[ 1576.787704] Call Trace:
[ 1576.787709]  <TASK>
[ 1576.787717]  amdgpu_device_wreg+0x26/0x50 [amdgpu]
[ 1576.787925]  jpeg_v4_0_stop+0x47/0x140 [amdgpu]
[ 1576.788170]  jpeg_v4_0_set_powergating_state+0x53/0x70 [amdgpu]
[ 1576.788410]  amdgpu_device_ip_set_powergating_state+0x67/0xc0 [amdgpu]
[ 1576.788642]  amdgpu_jpeg_idle_work_handler+0x105/0x120 [amdgpu]
[ 1576.788887]  process_one_work+0x23e/0x6f0
[ 1576.788917]  worker_thread+0x1c4/0x380
[ 1576.788931]  kthread+0x10c/0x150
[ 1576.788937]  ? __pfx_worker_thread+0x10/0x10
[ 1576.788943]  ? __pfx_kthread+0x10/0x10
[ 1576.788954]  ret_from_fork+0x314/0x390
[ 1576.788960]  ? __pfx_kthread+0x10/0x10
[ 1576.788969]  ret_from_fork_asm+0x1a/0x30
[ 1576.789003]  </TASK>

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_jpeg.c | 5 ++++-
 drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c  | 3 ++-
 2 files changed, 6 insertions(+), 2 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_jpeg.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_jpeg.c
index 208566ffe898..a66da05cc3f7 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_jpeg.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_jpeg.c
@@ -145,7 +145,10 @@ void amdgpu_jpeg_ring_begin_use(struct amdgpu_ring *ring)
 
 void amdgpu_jpeg_ring_end_use(struct amdgpu_ring *ring)
 {
-	if (atomic_dec_and_test(&ring->adev->jpeg.total_submission_cnt))
+	struct amdgpu_device *adev = ring->adev;
+
+	if (atomic_dec_and_test(&ring->adev->jpeg.total_submission_cnt) &&
+	   !amdgpu_in_reset(adev))
 		schedule_delayed_work(&ring->adev->jpeg.idle_work,
 				      JPEG_IDLE_TIMEOUT);
 }
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c
index 17db7264269e..6cf08e35dc9b 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c
@@ -549,7 +549,8 @@ void amdgpu_vcn_ring_end_use(struct amdgpu_ring *ring)
 	    !adev->vcn.inst[ring->me].using_unified_queue)
 		atomic_dec(&ring->adev->vcn.inst[ring->me].dpg_enc_submission_cnt);
 
-	if (atomic_dec_and_test(&ring->adev->vcn.inst[ring->me].total_submission_cnt))
+	if (atomic_dec_and_test(&ring->adev->vcn.inst[ring->me].total_submission_cnt) &&
+	    !amdgpu_in_reset(adev))
 		schedule_delayed_work(&ring->adev->vcn.inst[ring->me].idle_work,
 				      VCN_IDLE_TIMEOUT);
 }
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 16/18] drm/amdgpu/mes: skip userq_notify_unmap during gpu reset
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (13 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 15/18] drm/amdgpu/jpeg: skip scheduling jpeg/vcn idle_work " Prike Liang
@ 2026-09-02 12:49 ` Prike Liang
  2026-09-02 12:50 ` [PATCH 17/18] drm/amdgpu/userq: complete userq eviction fence in pre_reset Prike Liang
                   ` (2 subsequent siblings)
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:49 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

To avoid the following issue, it requires skipping
userq_notify_unmap during gpu reset.

[  356.185682] CPU: 13 UID: 0 PID: 599 Comm: kworker/13:2 Tainted: G        W  OE       7.1.0-custom #40 PREEMPT(lazy)
[  356.185690] Tainted: [W]=WARN, [O]=OOT_MODULE, [E]=UNSIGNED_MODULE
[  356.185694] Hardware name: AMD Majolica-RN/Majolica-RN, BIOS RMJ1009A 06/13/2021
[  356.185699] Workqueue: events amdgpu_mes_userq_notify_unmap_work_handler [amdgpu]
[  356.185952] RIP: 0010:amdgpu_device_skip_hw_access+0x73/0x90 [amdgpu]
[  356.186156] Code: 85 c0 75 2a 8b 05 71 56 d7 dc 85 c0 74 d3 48 8b bb d0 e7 07 00 be ff ff ff ff 48 81 c7 88 00 00 00 e8 f1 50 79 db 85 c0 75 b7 <0f> 0b eb b3 48 8b bb d0 e7 07 00 48 83 c7 18 e8 39 fe 76 da eb a1
[  356.186161] RSP: 0018:ffffd39142dbfa58 EFLAGS: 00010046
[  356.186168] RAX: 0000000000000000 RBX: ffff8bfc09d00000 RCX: 0000000000000003
[  356.186172] RDX: 0000000000000000 RSI: ffff8bfc04e00c88 RDI: ffff8bfc0b93df50
[  356.186176] RBP: ffffd39142dbfa68 R08: 0000000000000000 R09: 0000000000000000
[  356.186180] R10: 0000000000000000 R11: 0000000000000000 R12: 0000000000000000
[  356.186184] R13: 0000000000000680 R14: ffff8bfc09d67288 R15: 0000000000000000
[  356.186188] FS:  0000000000000000(0000) GS:ffff8bffb057d000(0000) knlGS:0000000000000000
[  356.186192] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  356.186197] CR2: 00007fbd3797c000 CR3: 0000000150642000 CR4: 0000000000350ef0
[  356.186201] Call Trace:
[  356.186205]  <TASK>
[  356.186214]  amdgpu_mm_wdoorbell64+0x20/0x70 [amdgpu]
[  356.186425]  mes_v11_0_ring_set_wptr+0x71/0x80 [amdgpu]
[  356.186670]  amdgpu_ring_commit+0x5b/0xa0 [amdgpu]
[  356.186885]  mes_v11_0_submit_pkt_and_poll_completion.constprop.0+0x207/0x4a0 [amdgpu]
[  356.187181]  mes_v11_0_misc_op+0x8b/0x240 [amdgpu]
[  356.187462]  amdgpu_mes_notify_unmap_queue+0xa3/0x110 [amdgpu]
[  356.187713]  amdgpu_mes_userq_notify_unmap_work_handler+0x24/0x70 [amdgpu]
[  356.187954]  process_one_work+0x23e/0x6f0
[  356.187984]  worker_thread+0x1c4/0x380
[  356.187998]  kthread+0x10c/0x150
[  356.188004]  ? __pfx_worker_thread+0x10/0x10
[  356.188010]  ? __pfx_kthread+0x10/0x10
[  356.188021]  ret_from_fork+0x314/0x390
[  356.188027]  ? __pfx_kthread+0x10/0x10
[  356.188036]  ret_from_fork_asm+0x1a/0x30
[  356.188069]  </TASK>

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_mes.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_mes.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_mes.c
index e0e38d6bcafc..dbf535adb1b7 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_mes.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_mes.c
@@ -1192,6 +1192,9 @@ static void amdgpu_mes_userq_notify_unmap_work_handler(struct work_struct *work)
 					       userq_notify_unmap_work.work);
 	struct amdgpu_device *adev = mes->adev;
 
+	if (amdgpu_in_reset(adev))
+		return;
+
 	amdgpu_mes_notify_unmap_queue(adev);
 
 	/* Re-arm if still oversubscribed */
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 17/18] drm/amdgpu/userq: complete userq eviction fence in pre_reset
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (14 preceding siblings ...)
  2026-09-02 12:49 ` [PATCH 16/18] drm/amdgpu/mes: skip userq_notify_unmap during gpu reset Prike Liang
@ 2026-09-02 12:50 ` Prike Liang
  2026-09-02 12:50 ` [PATCH 18/18] drm/amdgpu: stop the userq submission prior to removing userq Prike Liang
  2026-09-03 19:28 ` [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Alex Deucher
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:50 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

Complete the userq eviction fence before gpu reset.

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_userq_fence.c | 9 +++++++++
 1 file changed, 9 insertions(+)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq_fence.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq_fence.c
index b556fbb1c31e..49a938436b08 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq_fence.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq_fence.c
@@ -443,6 +443,9 @@ void
 amdgpu_userq_fence_driver_force_completion(struct amdgpu_usermode_queue *userq)
 {
 	struct dma_fence *f = userq->last_fence;
+	struct amdgpu_fpriv *fpriv = userq->userq_mgr->file->driver_priv;
+	struct amdgpu_eviction_fence_mgr *evf_mgr = &fpriv->evf_mgr;
+	struct dma_fence *ev_fence;
 
 	if (f) {
 		struct amdgpu_userq_fence *fence = to_amdgpu_userq_fence(f);
@@ -452,8 +455,14 @@ amdgpu_userq_fence_driver_force_completion(struct amdgpu_usermode_queue *userq)
 		amdgpu_userq_fence_driver_set_error(fence, -ECANCELED);
 		amdgpu_userq_fence_write(fence_drv, wptr);
 		amdgpu_userq_fence_driver_process(fence_drv);
+	}
 
+	ev_fence = amdgpu_evf_mgr_get_fence(evf_mgr);
+	if (!dma_fence_is_signaled(ev_fence)) {
+		dma_fence_set_error(ev_fence, -ECANCELED);
+		dma_fence_signal(ev_fence);
 	}
+	dma_fence_put(ev_fence);
 }
 
 int amdgpu_userq_signal_ioctl(struct drm_device *dev, void *data,
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* [PATCH 18/18] drm/amdgpu: stop the userq submission prior to removing userq
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (15 preceding siblings ...)
  2026-09-02 12:50 ` [PATCH 17/18] drm/amdgpu/userq: complete userq eviction fence in pre_reset Prike Liang
@ 2026-09-02 12:50 ` Prike Liang
  2026-09-03 19:28 ` [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Alex Deucher
  17 siblings, 0 replies; 24+ messages in thread
From: Prike Liang @ 2026-09-02 12:50 UTC (permalink / raw)
  To: amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak, Prike Liang

Stop the userq submission prior to removing userq in the
amdgpu_userq_pre_reset().

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 4 +++-
 1 file changed, 3 insertions(+), 1 deletion(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
index a7d5ca741a3b..c60365bbd7f6 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
@@ -1540,6 +1540,8 @@ void amdgpu_userq_pre_reset(struct amdgpu_device *adev)
 	struct amdgpu_usermode_queue *queue;
 	unsigned long queue_id;
 
+	amdgpu_mes_suspend(adev, 0);
+
 	/* TODO: We probably need a new lock for the queue state */
 	xa_for_each(&adev->userq_doorbell_xa, queue_id, queue) {
 		if (queue->state == AMDGPU_USERQ_STATE_MAPPED) {
@@ -1587,6 +1589,6 @@ int amdgpu_userq_post_reset(struct amdgpu_device *adev, bool vram_lost)
 			queue->state = AMDGPU_USERQ_STATE_MAPPED;
 		}
 	}
-
+	amdgpu_mes_resume(adev, 0);
 	return r;
 }
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 24+ messages in thread

* Re: [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset
  2026-09-02 12:49 ` [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset Prike Liang
@ 2026-09-03 19:25   ` Alex Deucher
  0 siblings, 0 replies; 24+ messages in thread
From: Alex Deucher @ 2026-09-03 19:25 UTC (permalink / raw)
  To: Prike Liang; +Cc: amd-gfx, Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak

Reviewed-by: Alex Deucher <alexander.deucher@amd.com>

On Wed, Sep 2, 2026 at 8:50 AM Prike Liang <Prike.Liang@amd.com> wrote:
>
> During resume from GPU reset, the gfx idle work may invoke switch_power_profile
> before the reset completes. This causes the following assert error because the
> register access occurs without first releasing the GPU reset semaphore:
>
> [ 1576.768935] CR2: 0000559ea133ead0 CR3: 00000002e6c42000 CR4: 0000000000350ef0
> [ 1576.768940] Call Trace:
> [ 1576.768944]  <TASK>
> [ 1576.768953]  amdgpu_device_rreg+0x21/0x50 [amdgpu]
> [ 1576.769158]  smu_msg_v1_send_msg+0x1a4/0x6e0 [amdgpu]
> [ 1576.769437]  smu_cmn_send_smc_msg_with_params_ext+0xba/0x120 [amdgpu]
> [ 1576.769721]  smu_cmn_send_smc_msg_with_param+0x33/0x40 [amdgpu]
> [ 1576.769993]  smu_v13_0_0_set_power_profile_mode+0x192/0x2b0 [amdgpu]
> [ 1576.770267]  smu_bump_power_profile_mode+0x5d/0x80 [amdgpu]
> [ 1576.770538]  smu_switch_power_profile+0xa4/0xf0 [amdgpu]
> [ 1576.770839]  amdgpu_dpm_switch_power_profile+0x6f/0x90 [amdgpu]
> [ 1576.771210]  amdgpu_gfx_profile_idle_work_handler+0xe9/0x130 [amdgpu]
> [ 1576.771460]  process_one_work+0x23e/0x6f0
> [ 1576.771491]  worker_thread+0x1c4/0x380
> [ 1576.771506]  kthread+0x10c/0x150
> [ 1576.771512]  ? __pfx_worker_thread+0x10/0x10
> [ 1576.771518]  ? __pfx_kthread+0x10/0x10
> [ 1576.771530]  ret_from_fork+0x314/0x390
> [ 1576.771537]  ? __pfx_kthread+0x10/0x10
> [ 1576.771546]  ret_from_fork_asm+0x1a/0x30
>
> Signed-off-by: Prike Liang <Prike.Liang@amd.com>
> ---
>  drivers/gpu/drm/amd/pm/amdgpu_dpm.c | 6 ++++--
>  1 file changed, 4 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/pm/amdgpu_dpm.c b/drivers/gpu/drm/amd/pm/amdgpu_dpm.c
> index ce526db4d24a..808be6c425bf 100644
> --- a/drivers/gpu/drm/amd/pm/amdgpu_dpm.c
> +++ b/drivers/gpu/drm/amd/pm/amdgpu_dpm.c
> @@ -348,7 +348,8 @@ int amdgpu_dpm_switch_power_profile(struct amdgpu_device *adev,
>         const struct amd_pm_funcs *pp_funcs = adev->powerplay.pp_funcs;
>         int ret = 0;
>
> -       if (amdgpu_sriov_vf(adev))
> +       if (amdgpu_sriov_vf(adev) ||
> +               amdgpu_in_reset(adev))
>                 return 0;
>
>         if (pp_funcs && pp_funcs->switch_power_profile) {
> @@ -367,7 +368,8 @@ int amdgpu_dpm_pause_power_profile(struct amdgpu_device *adev,
>         const struct amd_pm_funcs *pp_funcs = adev->powerplay.pp_funcs;
>         int ret = 0;
>
> -       if (amdgpu_sriov_vf(adev))
> +       if (amdgpu_sriov_vf(adev) ||
> +               amdgpu_in_reset(adev))
>                 return 0;
>
>         if (pp_funcs && pp_funcs->pause_power_profile) {
> --
> 2.34.1
>

^ permalink raw reply	[flat|nested] 24+ messages in thread

* Re: [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset
  2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
                   ` (16 preceding siblings ...)
  2026-09-02 12:50 ` [PATCH 18/18] drm/amdgpu: stop the userq submission prior to removing userq Prike Liang
@ 2026-09-03 19:28 ` Alex Deucher
  2026-09-07  6:48   ` Liang, Prike
  17 siblings, 1 reply; 24+ messages in thread
From: Alex Deucher @ 2026-09-03 19:28 UTC (permalink / raw)
  To: Prike Liang; +Cc: amd-gfx, Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak

On Wed, Sep 2, 2026 at 8:50 AM Prike Liang <Prike.Liang@amd.com> wrote:
>
> amdgpu_mes_detect_and_reset_hung_queues() already detects
> the guilty compute user queue and resets it through
> mes_userq_reset_queue(). The additional reset via
> mes_userq_reset() is unnecessary, so remove it to unify
> the compute userq reset.

The problem is that amdgpu_mes_detect_and_reset_hung_queues() won't
reset the queue in some cases.  Detect_and_reset() attempts to preempt
the queues and if they fail to preempt they are considered hung,
however, there are queues which can be preempted which are in a state
which won't make progress so the protected fence will never signal.
That's why we have this special case.

Alex

>
> Signed-off-by: Prike Liang <Prike.Liang@amd.com>
> ---
>  drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c    | 5 -----
>  drivers/gpu/drm/amd/amdgpu/mes_userqueue.c | 2 --
>  2 files changed, 7 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c
> index a6f95ff47d24..5c3be851ac84 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c
> @@ -2384,11 +2384,6 @@ int amdgpu_gfx_reset_mes_compute(struct amdgpu_device *adev,
>                 deferred_end[n_deferred].fence = guilty_fence;
>                 n_deferred++;
>         }
> -       if (uq) {
> -               r = mes_userq_reset(uq);
> -               if (r)
> -                       goto out;
> -       }
>         for (i = 0; i < num_hung; i++) {
>                 struct amdgpu_ring *hr = NULL;
>                 struct amdgpu_fence *hf = NULL;
> diff --git a/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c b/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
> index 7f334f718cd8..83a438d8e117 100644
> --- a/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
> +++ b/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
> @@ -249,8 +249,6 @@ int mes_userq_reset_queue(struct amdgpu_device *adev,
>
>         xa_for_each(&adev->userq_doorbell_xa, uq_id, uq) {
>                 if (uq->queue_type == queue_type) {
> -                       if (uq == guilty_uq)
> -                               continue;
>                         if (uq->doorbell_index == db) {
>                                 uq->state = AMDGPU_USERQ_STATE_HUNG;
>                                 if (use_mmio)
> --
> 2.34.1
>

^ permalink raw reply	[flat|nested] 24+ messages in thread

* Re: [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check
  2026-09-02 12:49 ` [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check Prike Liang
@ 2026-09-03 19:29   ` Alex Deucher
  0 siblings, 0 replies; 24+ messages in thread
From: Alex Deucher @ 2026-09-03 19:29 UTC (permalink / raw)
  To: Prike Liang; +Cc: amd-gfx, Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak

Reviewed-by: Alex Deucher <alexander.deucher@amd.com>

On Wed, Sep 2, 2026 at 11:25 AM Prike Liang <Prike.Liang@amd.com> wrote:
>
> If the userq doesn't support in a system. then there's no
> valid userq_doorbell_xa entry to walk over and then has a
> no-op.
>
> Signed-off-by: Prike Liang <Prike.Liang@amd.com>
> ---
>  drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 8 --------
>  1 file changed, 8 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> index 0a816b3c5ff9..ed329041a648 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c
> @@ -1373,15 +1373,11 @@ void amdgpu_userq_mgr_fini(struct amdgpu_userq_mgr *userq_mgr)
>
>  int amdgpu_userq_suspend(struct amdgpu_device *adev)
>  {
> -       u32 ip_mask = amdgpu_userq_get_supported_ip_mask(adev);
>         struct amdgpu_usermode_queue *queue;
>         struct amdgpu_userq_mgr *uqm;
>         unsigned long queue_id;
>         int r;
>
> -       if (!ip_mask)
> -               return 0;
> -
>         xa_for_each(&adev->userq_doorbell_xa, queue_id, queue) {
>                 uqm = queue->userq_mgr;
>                 cancel_delayed_work_sync(&uqm->resume_work);
> @@ -1398,15 +1394,11 @@ int amdgpu_userq_suspend(struct amdgpu_device *adev)
>
>  int amdgpu_userq_resume(struct amdgpu_device *adev)
>  {
> -       u32 ip_mask = amdgpu_userq_get_supported_ip_mask(adev);
>         struct amdgpu_usermode_queue *queue;
>         struct amdgpu_userq_mgr *uqm;
>         unsigned long queue_id;
>         int r;
>
> -       if (!ip_mask)
> -               return 0;
> -
>         xa_for_each(&adev->userq_doorbell_xa, queue_id, queue) {
>                 uqm = queue->userq_mgr;
>                 guard(mutex)(&uqm->userq_mutex);
> --
> 2.34.1
>

^ permalink raw reply	[flat|nested] 24+ messages in thread

* RE: [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset
  2026-09-03 19:28 ` [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Alex Deucher
@ 2026-09-07  6:48   ` Liang, Prike
  0 siblings, 0 replies; 24+ messages in thread
From: Liang, Prike @ 2026-09-07  6:48 UTC (permalink / raw)
  To: Alex Deucher
  Cc: amd-gfx@lists.freedesktop.org, Deucher,  Alexander,
	Koenig, Christian, Prosyak, Vitaly

AMD General

Regards,
      Prike

> -----Original Message-----
> From: Alex Deucher <alexdeucher@gmail.com>
> Sent: Friday, September 4, 2026 3:28 AM
> To: Liang, Prike <Prike.Liang@amd.com>
> Cc: amd-gfx@lists.freedesktop.org; Deucher, Alexander
> <Alexander.Deucher@amd.com>; Koenig, Christian <Christian.Koenig@amd.com>;
> Prosyak, Vitaly <Vitaly.Prosyak@amd.com>
> Subject: Re: [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq
> reset
>
> On Wed, Sep 2, 2026 at 8:50 AM Prike Liang <Prike.Liang@amd.com> wrote:
> >
> > amdgpu_mes_detect_and_reset_hung_queues() already detects the guilty
> > compute user queue and resets it through mes_userq_reset_queue(). The
> > additional reset via
> > mes_userq_reset() is unnecessary, so remove it to unify the compute
> > userq reset.
>
> The problem is that amdgpu_mes_detect_and_reset_hung_queues() won't reset the
> queue in some cases.  Detect_and_reset() attempts to preempt the queues and if
> they fail to preempt they are considered hung, however, there are queues which can
> be preempted which are in a state which won't make progress so the protected fence
> will never signal.
> That's why we have this special case.

Thank you for the detailed background. Regarding amdgpu_mes_detect_and_reset_hung_queues() in amdgpu_gfx_reset_mes_compute(), this function is only responsible for detecting hung queues, not resetting them. On the MES firmware side, a hung queue is identified by reading and comparing the MQD status over a query time span. Because of this, there is a possibility that a spurious hang detection gets scheduled, triggered by a userq signaling timeout. Such timeout cases may originate from a long-running submission or a slow, heavy queue execution. For these false positive timeout scenarios, we could consider preempting and restoring the userq rather than blindly issuing a reset or alternatively, having userspace skip emitting a tracked fence, as you proposed in a separate review thread.

I will investigate further and draft a proper solution for handling these false timeout cases. Until that is finalized, I will drop this patch.

Regards,
    Prike

> Alex
>
> >
> > Signed-off-by: Prike Liang <Prike.Liang@amd.com>
> > ---
> >  drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c    | 5 -----
> >  drivers/gpu/drm/amd/amdgpu/mes_userqueue.c | 2 --
> >  2 files changed, 7 deletions(-)
> >
> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c
> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c
> > index a6f95ff47d24..5c3be851ac84 100644
> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c
> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gfx.c
> > @@ -2384,11 +2384,6 @@ int amdgpu_gfx_reset_mes_compute(struct
> amdgpu_device *adev,
> >                 deferred_end[n_deferred].fence = guilty_fence;
> >                 n_deferred++;
> >         }
> > -       if (uq) {
> > -               r = mes_userq_reset(uq);
> > -               if (r)
> > -                       goto out;
> > -       }
> >         for (i = 0; i < num_hung; i++) {
> >                 struct amdgpu_ring *hr = NULL;
> >                 struct amdgpu_fence *hf = NULL; diff --git
> > a/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
> > b/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
> > index 7f334f718cd8..83a438d8e117 100644
> > --- a/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
> > +++ b/drivers/gpu/drm/amd/amdgpu/mes_userqueue.c
> > @@ -249,8 +249,6 @@ int mes_userq_reset_queue(struct amdgpu_device
> > *adev,
> >
> >         xa_for_each(&adev->userq_doorbell_xa, uq_id, uq) {
> >                 if (uq->queue_type == queue_type) {
> > -                       if (uq == guilty_uq)
> > -                               continue;
> >                         if (uq->doorbell_index == db) {
> >                                 uq->state = AMDGPU_USERQ_STATE_HUNG;
> >                                 if (use_mmio)
> > --
> > 2.34.1
> >

^ permalink raw reply	[flat|nested] 24+ messages in thread

* Re: [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery
  2026-09-02 12:49 ` [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery Prike Liang
@ 2026-09-13 20:48   ` vitaly prosyak
  0 siblings, 0 replies; 24+ messages in thread
From: vitaly prosyak @ 2026-09-13 20:48 UTC (permalink / raw)
  To: Prike Liang, amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak


On 2026-09-02 08:49, Prike Liang wrote:
> Suspend the drm internal clients has a deadlock risk as acquiring
> it while holding the reset domain lock inverts the ordering
> established elsewhere (clientlist_mutex -> ... -> reset_domain->sem).
>
> Reset AMDGPU can prevent the user space clients further accessing by
> using the reset semaphore, so removing the drm_client_dev_suspend() |
> resume() in the reset path.
That is only true for ioctl paths. amdgpu_userq_restore_worker is a
workqueue, not an ioctl. It never takes reset_domain->sem. The removed

suspend/resume call was the only thing blocking it during reset.

Thanks, Vitaly

> Signed-off-by: Prike Liang <Prike.Liang@amd.com>
> ---
>  drivers/gpu/drm/amd/amdgpu/amdgpu_device.c | 4 ----
>  1 file changed, 4 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> index d7640da9f6de..bd4eb97336b1 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> @@ -5210,8 +5210,6 @@ int amdgpu_device_reinit_after_reset(struct amdgpu_reset_context *reset_context)
>  				if (r)
>  					goto out;
>  
> -				drm_client_dev_resume(adev_to_drm(tmp_adev));
> -
>  				/*
>  				 * The GPU enters bad state once faulty pages
>  				 * by ECC has reached the threshold, and ras
> @@ -5544,8 +5542,6 @@ static void amdgpu_device_halt_activities(struct amdgpu_device *adev,
>  		 */
>  		amdgpu_unregister_gpu_instance(tmp_adev);
>  
> -		drm_client_dev_suspend(adev_to_drm(tmp_adev));
> -
>  		/* disable ras on ALL IPs */
>  		if (!need_emergency_restart && !amdgpu_reset_in_dpc(adev))
>  			amdgpu_ras_suspend(tmp_adev);

^ permalink raw reply	[flat|nested] 24+ messages in thread

* Re: [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset
  2026-09-02 12:49 ` [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset Prike Liang
@ 2026-09-13 20:53   ` vitaly prosyak
  0 siblings, 0 replies; 24+ messages in thread
From: vitaly prosyak @ 2026-09-13 20:53 UTC (permalink / raw)
  To: Prike Liang, amd-gfx; +Cc: Alexander.Deucher, Christian.Koenig, Vitaly.Prosyak


On 2026-09-02 08:49, Prike Liang wrote:
> The eviction fence suspend worker can submit MES REMOVE_QUEUE packets
> without holding the reset-domain semaphore. If GPU recovery starts while
> the worker is running, both paths can access the hardware concurrently.
> This triggers the hardware-access lockdep assertion and can submit a MES
> packet while recovery is resetting the device.
>
> Try to take the reset-domain semaphore for read around userq eviction. Do
> not block on it while holding userq_mutex because recovery takes the reset
> semaphore for write before acquiring buffer reservations and userq_mutex.
> Instead, drop userq_mutex, wait for recovery without holding any other
> lock, and retry the queue-state checks after recovery completes.
>
> This makes recovery wait for an in-flight MES eviction, while an eviction
> which starts after recovery waits without introducing the reverse lock
> dependency that caused the reported circular-lock warning.

It only changes amdgpu_eviction_fence_suspend_worker(). It adds
reset_domain->sem around the evict path only.

It does not touch amdgpu_userq_vm_validate_and_restore_queue() or
amdgpu_evf_mgr_rearm().

Eviction is now serialized against reset, but restore/rearm
is not. Same worker, still unprotected.

Thanks, Vitaly

> Signed-off-by: Prike Liang <Prike.Liang@amd.com>
> ---
>  drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c | 9 +++++++++
>  1 file changed, 9 insertions(+)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
> index 2ea8553c82f0..93307cbf55dc 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
> @@ -67,10 +67,18 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct *work)
>  	bool cookie;
>  	int r;
>  
> +retry:
>  	mutex_lock(&uq_mgr->userq_mutex);
>  
>  	/* Fence waits are not allowed in a fence signalling critical section. */
>  	amdgpu_userq_wait_for_signal(uq_mgr);
> +	if (!down_read_trylock(&uq_mgr->adev->reset_domain->sem)) {
> +		mutex_unlock(&uq_mgr->userq_mutex);
> +
> +		down_read(&uq_mgr->adev->reset_domain->sem);
> +		up_read(&uq_mgr->adev->reset_domain->sem);
> +		goto retry;
> +	}
>  
>  	/*
>  	 * This is intentionally after taking the userq_mutex since we do
> @@ -81,6 +89,7 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct *work)
>  
>  	ev_fence = amdgpu_evf_mgr_get_fence(evf_mgr);
>  	r = amdgpu_userq_evict(uq_mgr);
> +	up_read(&uq_mgr->adev->reset_domain->sem);
>  	if (r)
>  		dma_fence_set_error(ev_fence, r);
>  

^ permalink raw reply	[flat|nested] 24+ messages in thread

end of thread, other threads:[~2026-09-13 20:53 UTC | newest]

Thread overview: 24+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
2026-09-02 12:49 ` [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check Prike Liang
2026-09-03 19:29   ` Alex Deucher
2026-09-02 12:49 ` [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery Prike Liang
2026-09-13 20:48   ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 04/18] drm/amdgpu: move userq fence wait out of signalling section Prike Liang
2026-09-02 12:49 ` [PATCH 05/18] drm/amdgpu: defer userq reset after eviction failure Prike Liang
2026-09-02 12:49 ` [PATCH 06/18] drm/amdgpu: skip DRM internal suspend/resume for reseting VKMS Prike Liang
2026-09-02 12:49 ` [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset Prike Liang
2026-09-13 20:53   ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 08/18] drm/amdgpu: skip VMHUB HW access in unaccessiable device Prike Liang
2026-09-02 12:49 ` [PATCH 09/18] drm/amdgpu/userq: complete the hang userq fence Prike Liang
2026-09-02 12:49 ` [PATCH 10/18] drm/amdgpu/mes: put the mes context BO allocation in mes sw_int Prike Liang
2026-09-02 12:49 ` [PATCH 11/18] drm/amdgpu: depart ring scheduler after resumming IP blocks Prike Liang
2026-09-02 12:49 ` [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock Prike Liang
2026-09-02 12:49 ` [PATCH 13/18] drm/amdgpu: allocate dma_fence slot explicitly for rearming eviction fence Prike Liang
2026-09-02 12:49 ` [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset Prike Liang
2026-09-03 19:25   ` Alex Deucher
2026-09-02 12:49 ` [PATCH 15/18] drm/amdgpu/jpeg: skip scheduling jpeg/vcn idle_work " Prike Liang
2026-09-02 12:49 ` [PATCH 16/18] drm/amdgpu/mes: skip userq_notify_unmap during gpu reset Prike Liang
2026-09-02 12:50 ` [PATCH 17/18] drm/amdgpu/userq: complete userq eviction fence in pre_reset Prike Liang
2026-09-02 12:50 ` [PATCH 18/18] drm/amdgpu: stop the userq submission prior to removing userq Prike Liang
2026-09-03 19:28 ` [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Alex Deucher
2026-09-07  6:48   ` Liang, Prike

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.