From: vitaly prosyak <vprosyak@amd.com>
To: Prike Liang <Prike.Liang@amd.com>, amd-gfx@lists.freedesktop.org
Cc: Alexander.Deucher@amd.com, Christian.Koenig@amd.com,
Vitaly.Prosyak@amd.com
Subject: Re: [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset
Date: Sun, 13 Sep 2026 16:53:09 -0400 [thread overview]
Message-ID: <b8bb7c32-6045-4793-a059-2e6175db0f11@amd.com> (raw)
In-Reply-To: <20260902125001.621629-7-Prike.Liang@amd.com>
On 2026-09-02 08:49, Prike Liang wrote:
> The eviction fence suspend worker can submit MES REMOVE_QUEUE packets
> without holding the reset-domain semaphore. If GPU recovery starts while
> the worker is running, both paths can access the hardware concurrently.
> This triggers the hardware-access lockdep assertion and can submit a MES
> packet while recovery is resetting the device.
>
> Try to take the reset-domain semaphore for read around userq eviction. Do
> not block on it while holding userq_mutex because recovery takes the reset
> semaphore for write before acquiring buffer reservations and userq_mutex.
> Instead, drop userq_mutex, wait for recovery without holding any other
> lock, and retry the queue-state checks after recovery completes.
>
> This makes recovery wait for an in-flight MES eviction, while an eviction
> which starts after recovery waits without introducing the reverse lock
> dependency that caused the reported circular-lock warning.
It only changes amdgpu_eviction_fence_suspend_worker(). It adds
reset_domain->sem around the evict path only.
It does not touch amdgpu_userq_vm_validate_and_restore_queue() or
amdgpu_evf_mgr_rearm().
Eviction is now serialized against reset, but restore/rearm
is not. Same worker, still unprotected.
Thanks, Vitaly
> Signed-off-by: Prike Liang <Prike.Liang@amd.com>
> ---
> drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c | 9 +++++++++
> 1 file changed, 9 insertions(+)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
> index 2ea8553c82f0..93307cbf55dc 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c
> @@ -67,10 +67,18 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct *work)
> bool cookie;
> int r;
>
> +retry:
> mutex_lock(&uq_mgr->userq_mutex);
>
> /* Fence waits are not allowed in a fence signalling critical section. */
> amdgpu_userq_wait_for_signal(uq_mgr);
> + if (!down_read_trylock(&uq_mgr->adev->reset_domain->sem)) {
> + mutex_unlock(&uq_mgr->userq_mutex);
> +
> + down_read(&uq_mgr->adev->reset_domain->sem);
> + up_read(&uq_mgr->adev->reset_domain->sem);
> + goto retry;
> + }
>
> /*
> * This is intentionally after taking the userq_mutex since we do
> @@ -81,6 +89,7 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct *work)
>
> ev_fence = amdgpu_evf_mgr_get_fence(evf_mgr);
> r = amdgpu_userq_evict(uq_mgr);
> + up_read(&uq_mgr->adev->reset_domain->sem);
> if (r)
> dma_fence_set_error(ev_fence, r);
>
next prev parent reply other threads:[~2026-09-13 20:53 UTC|newest]
Thread overview: 24+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
2026-09-02 12:49 ` [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check Prike Liang
2026-09-03 19:29 ` Alex Deucher
2026-09-02 12:49 ` [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery Prike Liang
2026-09-13 20:48 ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 04/18] drm/amdgpu: move userq fence wait out of signalling section Prike Liang
2026-09-02 12:49 ` [PATCH 05/18] drm/amdgpu: defer userq reset after eviction failure Prike Liang
2026-09-02 12:49 ` [PATCH 06/18] drm/amdgpu: skip DRM internal suspend/resume for reseting VKMS Prike Liang
2026-09-02 12:49 ` [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset Prike Liang
2026-09-13 20:53 ` vitaly prosyak [this message]
2026-09-02 12:49 ` [PATCH 08/18] drm/amdgpu: skip VMHUB HW access in unaccessiable device Prike Liang
2026-09-02 12:49 ` [PATCH 09/18] drm/amdgpu/userq: complete the hang userq fence Prike Liang
2026-09-02 12:49 ` [PATCH 10/18] drm/amdgpu/mes: put the mes context BO allocation in mes sw_int Prike Liang
2026-09-02 12:49 ` [PATCH 11/18] drm/amdgpu: depart ring scheduler after resumming IP blocks Prike Liang
2026-09-02 12:49 ` [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock Prike Liang
2026-09-02 12:49 ` [PATCH 13/18] drm/amdgpu: allocate dma_fence slot explicitly for rearming eviction fence Prike Liang
2026-09-02 12:49 ` [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset Prike Liang
2026-09-03 19:25 ` Alex Deucher
2026-09-02 12:49 ` [PATCH 15/18] drm/amdgpu/jpeg: skip scheduling jpeg/vcn idle_work " Prike Liang
2026-09-02 12:49 ` [PATCH 16/18] drm/amdgpu/mes: skip userq_notify_unmap during gpu reset Prike Liang
2026-09-02 12:50 ` [PATCH 17/18] drm/amdgpu/userq: complete userq eviction fence in pre_reset Prike Liang
2026-09-02 12:50 ` [PATCH 18/18] drm/amdgpu: stop the userq submission prior to removing userq Prike Liang
2026-09-03 19:28 ` [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Alex Deucher
2026-09-07 6:48 ` Liang, Prike
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=b8bb7c32-6045-4793-a059-2e6175db0f11@amd.com \
--to=vprosyak@amd.com \
--cc=Alexander.Deucher@amd.com \
--cc=Christian.Koenig@amd.com \
--cc=Prike.Liang@amd.com \
--cc=Vitaly.Prosyak@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.