From: vitaly prosyak <vprosyak@amd.com>
To: Prike Liang <Prike.Liang@amd.com>, amd-gfx@lists.freedesktop.org
Cc: Alexander.Deucher@amd.com, Christian.Koenig@amd.com,
Vitaly.Prosyak@amd.com
Subject: Re: [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery
Date: Sun, 13 Sep 2026 16:48:10 -0400 [thread overview]
Message-ID: <046b6af2-147d-494f-a61e-e84da46de022@amd.com> (raw)
In-Reply-To: <20260902125001.621629-3-Prike.Liang@amd.com>
On 2026-09-02 08:49, Prike Liang wrote:
> Suspend the drm internal clients has a deadlock risk as acquiring
> it while holding the reset domain lock inverts the ordering
> established elsewhere (clientlist_mutex -> ... -> reset_domain->sem).
>
> Reset AMDGPU can prevent the user space clients further accessing by
> using the reset semaphore, so removing the drm_client_dev_suspend() |
> resume() in the reset path.
That is only true for ioctl paths. amdgpu_userq_restore_worker is a
workqueue, not an ioctl. It never takes reset_domain->sem. The removed
suspend/resume call was the only thing blocking it during reset.
Thanks, Vitaly
> Signed-off-by: Prike Liang <Prike.Liang@amd.com>
> ---
> drivers/gpu/drm/amd/amdgpu/amdgpu_device.c | 4 ----
> 1 file changed, 4 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> index d7640da9f6de..bd4eb97336b1 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> @@ -5210,8 +5210,6 @@ int amdgpu_device_reinit_after_reset(struct amdgpu_reset_context *reset_context)
> if (r)
> goto out;
>
> - drm_client_dev_resume(adev_to_drm(tmp_adev));
> -
> /*
> * The GPU enters bad state once faulty pages
> * by ECC has reached the threshold, and ras
> @@ -5544,8 +5542,6 @@ static void amdgpu_device_halt_activities(struct amdgpu_device *adev,
> */
> amdgpu_unregister_gpu_instance(tmp_adev);
>
> - drm_client_dev_suspend(adev_to_drm(tmp_adev));
> -
> /* disable ras on ALL IPs */
> if (!need_emergency_restart && !amdgpu_reset_in_dpc(adev))
> amdgpu_ras_suspend(tmp_adev);
next prev parent reply other threads:[~2026-09-13 20:48 UTC|newest]
Thread overview: 28+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
2026-09-02 12:49 ` [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check Prike Liang
2026-09-03 19:29 ` Alex Deucher
2026-09-02 12:49 ` [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery Prike Liang
2026-09-13 20:48 ` vitaly prosyak [this message]
2026-09-02 12:49 ` [PATCH 04/18] drm/amdgpu: move userq fence wait out of signalling section Prike Liang
2026-09-15 1:49 ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 05/18] drm/amdgpu: defer userq reset after eviction failure Prike Liang
2026-09-02 12:49 ` [PATCH 06/18] drm/amdgpu: skip DRM internal suspend/resume for reseting VKMS Prike Liang
2026-09-02 12:49 ` [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset Prike Liang
2026-09-13 20:53 ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 08/18] drm/amdgpu: skip VMHUB HW access in unaccessiable device Prike Liang
2026-09-02 12:49 ` [PATCH 09/18] drm/amdgpu/userq: complete the hang userq fence Prike Liang
2026-09-02 12:49 ` [PATCH 10/18] drm/amdgpu/mes: put the mes context BO allocation in mes sw_int Prike Liang
2026-09-02 12:49 ` [PATCH 11/18] drm/amdgpu: depart ring scheduler after resumming IP blocks Prike Liang
2026-09-02 12:49 ` [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock Prike Liang
2026-09-28 22:49 ` vitaly prosyak
2026-09-29 3:44 ` Liang, Prike
2026-09-29 8:56 ` Christian König
2026-09-02 12:49 ` [PATCH 13/18] drm/amdgpu: allocate dma_fence slot explicitly for rearming eviction fence Prike Liang
2026-09-02 12:49 ` [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset Prike Liang
2026-09-03 19:25 ` Alex Deucher
2026-09-02 12:49 ` [PATCH 15/18] drm/amdgpu/jpeg: skip scheduling jpeg/vcn idle_work " Prike Liang
2026-09-02 12:49 ` [PATCH 16/18] drm/amdgpu/mes: skip userq_notify_unmap during gpu reset Prike Liang
2026-09-02 12:50 ` [PATCH 17/18] drm/amdgpu/userq: complete userq eviction fence in pre_reset Prike Liang
2026-09-02 12:50 ` [PATCH 18/18] drm/amdgpu: stop the userq submission prior to removing userq Prike Liang
2026-09-03 19:28 ` [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Alex Deucher
2026-09-07 6:48 ` Liang, Prike
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=046b6af2-147d-494f-a61e-e84da46de022@amd.com \
--to=vprosyak@amd.com \
--cc=Alexander.Deucher@amd.com \
--cc=Christian.Koenig@amd.com \
--cc=Prike.Liang@amd.com \
--cc=Vitaly.Prosyak@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox