AMD-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: vitaly prosyak <vprosyak@amd.com>
To: Prike Liang <Prike.Liang@amd.com>, amd-gfx@lists.freedesktop.org
Cc: Alexander.Deucher@amd.com, Christian.Koenig@amd.com,
	Vitaly.Prosyak@amd.com
Subject: Re: [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery
Date: Sun, 13 Sep 2026 16:48:10 -0400	[thread overview]
Message-ID: <046b6af2-147d-494f-a61e-e84da46de022@amd.com> (raw)
In-Reply-To: <20260902125001.621629-3-Prike.Liang@amd.com>


On 2026-09-02 08:49, Prike Liang wrote:
> Suspend the drm internal clients has a deadlock risk as acquiring
> it while holding the reset domain lock inverts the ordering
> established elsewhere (clientlist_mutex -> ... -> reset_domain->sem).
>
> Reset AMDGPU can prevent the user space clients further accessing by
> using the reset semaphore, so removing the drm_client_dev_suspend() |
> resume() in the reset path.
That is only true for ioctl paths. amdgpu_userq_restore_worker is a
workqueue, not an ioctl. It never takes reset_domain->sem. The removed

suspend/resume call was the only thing blocking it during reset.

Thanks, Vitaly

> Signed-off-by: Prike Liang <Prike.Liang@amd.com>
> ---
>  drivers/gpu/drm/amd/amdgpu/amdgpu_device.c | 4 ----
>  1 file changed, 4 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> index d7640da9f6de..bd4eb97336b1 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> @@ -5210,8 +5210,6 @@ int amdgpu_device_reinit_after_reset(struct amdgpu_reset_context *reset_context)
>  				if (r)
>  					goto out;
>  
> -				drm_client_dev_resume(adev_to_drm(tmp_adev));
> -
>  				/*
>  				 * The GPU enters bad state once faulty pages
>  				 * by ECC has reached the threshold, and ras
> @@ -5544,8 +5542,6 @@ static void amdgpu_device_halt_activities(struct amdgpu_device *adev,
>  		 */
>  		amdgpu_unregister_gpu_instance(tmp_adev);
>  
> -		drm_client_dev_suspend(adev_to_drm(tmp_adev));
> -
>  		/* disable ras on ALL IPs */
>  		if (!need_emergency_restart && !amdgpu_reset_in_dpc(adev))
>  			amdgpu_ras_suspend(tmp_adev);

  reply	other threads:[~2026-09-13 20:48 UTC|newest]

Thread overview: 28+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
2026-09-02 12:49 ` [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check Prike Liang
2026-09-03 19:29   ` Alex Deucher
2026-09-02 12:49 ` [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery Prike Liang
2026-09-13 20:48   ` vitaly prosyak [this message]
2026-09-02 12:49 ` [PATCH 04/18] drm/amdgpu: move userq fence wait out of signalling section Prike Liang
2026-09-15  1:49   ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 05/18] drm/amdgpu: defer userq reset after eviction failure Prike Liang
2026-09-02 12:49 ` [PATCH 06/18] drm/amdgpu: skip DRM internal suspend/resume for reseting VKMS Prike Liang
2026-09-02 12:49 ` [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset Prike Liang
2026-09-13 20:53   ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 08/18] drm/amdgpu: skip VMHUB HW access in unaccessiable device Prike Liang
2026-09-02 12:49 ` [PATCH 09/18] drm/amdgpu/userq: complete the hang userq fence Prike Liang
2026-09-02 12:49 ` [PATCH 10/18] drm/amdgpu/mes: put the mes context BO allocation in mes sw_int Prike Liang
2026-09-02 12:49 ` [PATCH 11/18] drm/amdgpu: depart ring scheduler after resumming IP blocks Prike Liang
2026-09-02 12:49 ` [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock Prike Liang
2026-09-28 22:49   ` vitaly prosyak
2026-09-29  3:44     ` Liang, Prike
2026-09-29  8:56     ` Christian König
2026-09-02 12:49 ` [PATCH 13/18] drm/amdgpu: allocate dma_fence slot explicitly for rearming eviction fence Prike Liang
2026-09-02 12:49 ` [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset Prike Liang
2026-09-03 19:25   ` Alex Deucher
2026-09-02 12:49 ` [PATCH 15/18] drm/amdgpu/jpeg: skip scheduling jpeg/vcn idle_work " Prike Liang
2026-09-02 12:49 ` [PATCH 16/18] drm/amdgpu/mes: skip userq_notify_unmap during gpu reset Prike Liang
2026-09-02 12:50 ` [PATCH 17/18] drm/amdgpu/userq: complete userq eviction fence in pre_reset Prike Liang
2026-09-02 12:50 ` [PATCH 18/18] drm/amdgpu: stop the userq submission prior to removing userq Prike Liang
2026-09-03 19:28 ` [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Alex Deucher
2026-09-07  6:48   ` Liang, Prike

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=046b6af2-147d-494f-a61e-e84da46de022@amd.com \
    --to=vprosyak@amd.com \
    --cc=Alexander.Deucher@amd.com \
    --cc=Christian.Koenig@amd.com \
    --cc=Prike.Liang@amd.com \
    --cc=Vitaly.Prosyak@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox