All of lore.kernel.org
 help / color / mirror / Atom feed
From: vitaly prosyak <vprosyak@amd.com>
To: Prike Liang <Prike.Liang@amd.com>, amd-gfx@lists.freedesktop.org
Cc: Alexander.Deucher@amd.com, Christian.Koenig@amd.com,
	Vitaly.Prosyak@amd.com
Subject: Re: [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery
Date: Sun, 13 Sep 2026 16:48:10 -0400	[thread overview]
Message-ID: <046b6af2-147d-494f-a61e-e84da46de022@amd.com> (raw)
In-Reply-To: <20260902125001.621629-3-Prike.Liang@amd.com>


On 2026-09-02 08:49, Prike Liang wrote:
> Suspend the drm internal clients has a deadlock risk as acquiring
> it while holding the reset domain lock inverts the ordering
> established elsewhere (clientlist_mutex -> ... -> reset_domain->sem).
>
> Reset AMDGPU can prevent the user space clients further accessing by
> using the reset semaphore, so removing the drm_client_dev_suspend() |
> resume() in the reset path.
That is only true for ioctl paths. amdgpu_userq_restore_worker is a
workqueue, not an ioctl. It never takes reset_domain->sem. The removed

suspend/resume call was the only thing blocking it during reset.

Thanks, Vitaly

> Signed-off-by: Prike Liang <Prike.Liang@amd.com>
> ---
>  drivers/gpu/drm/amd/amdgpu/amdgpu_device.c | 4 ----
>  1 file changed, 4 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> index d7640da9f6de..bd4eb97336b1 100644
> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
> @@ -5210,8 +5210,6 @@ int amdgpu_device_reinit_after_reset(struct amdgpu_reset_context *reset_context)
>  				if (r)
>  					goto out;
>  
> -				drm_client_dev_resume(adev_to_drm(tmp_adev));
> -
>  				/*
>  				 * The GPU enters bad state once faulty pages
>  				 * by ECC has reached the threshold, and ras
> @@ -5544,8 +5542,6 @@ static void amdgpu_device_halt_activities(struct amdgpu_device *adev,
>  		 */
>  		amdgpu_unregister_gpu_instance(tmp_adev);
>  
> -		drm_client_dev_suspend(adev_to_drm(tmp_adev));
> -
>  		/* disable ras on ALL IPs */
>  		if (!need_emergency_restart && !amdgpu_reset_in_dpc(adev))
>  			amdgpu_ras_suspend(tmp_adev);

  reply	other threads:[~2026-09-13 20:48 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
2026-09-02 12:49 ` [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check Prike Liang
2026-09-03 19:29   ` Alex Deucher
2026-09-02 12:49 ` [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery Prike Liang
2026-09-13 20:48   ` vitaly prosyak [this message]
2026-09-02 12:49 ` [PATCH 04/18] drm/amdgpu: move userq fence wait out of signalling section Prike Liang
2026-09-02 12:49 ` [PATCH 05/18] drm/amdgpu: defer userq reset after eviction failure Prike Liang
2026-09-02 12:49 ` [PATCH 06/18] drm/amdgpu: skip DRM internal suspend/resume for reseting VKMS Prike Liang
2026-09-02 12:49 ` [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset Prike Liang
2026-09-13 20:53   ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 08/18] drm/amdgpu: skip VMHUB HW access in unaccessiable device Prike Liang
2026-09-02 12:49 ` [PATCH 09/18] drm/amdgpu/userq: complete the hang userq fence Prike Liang
2026-09-02 12:49 ` [PATCH 10/18] drm/amdgpu/mes: put the mes context BO allocation in mes sw_int Prike Liang
2026-09-02 12:49 ` [PATCH 11/18] drm/amdgpu: depart ring scheduler after resumming IP blocks Prike Liang
2026-09-02 12:49 ` [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock Prike Liang
2026-09-02 12:49 ` [PATCH 13/18] drm/amdgpu: allocate dma_fence slot explicitly for rearming eviction fence Prike Liang
2026-09-02 12:49 ` [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset Prike Liang
2026-09-03 19:25   ` Alex Deucher
2026-09-02 12:49 ` [PATCH 15/18] drm/amdgpu/jpeg: skip scheduling jpeg/vcn idle_work " Prike Liang
2026-09-02 12:49 ` [PATCH 16/18] drm/amdgpu/mes: skip userq_notify_unmap during gpu reset Prike Liang
2026-09-02 12:50 ` [PATCH 17/18] drm/amdgpu/userq: complete userq eviction fence in pre_reset Prike Liang
2026-09-02 12:50 ` [PATCH 18/18] drm/amdgpu: stop the userq submission prior to removing userq Prike Liang
2026-09-03 19:28 ` [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Alex Deucher
2026-09-07  6:48   ` Liang, Prike

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=046b6af2-147d-494f-a61e-e84da46de022@amd.com \
    --to=vprosyak@amd.com \
    --cc=Alexander.Deucher@amd.com \
    --cc=Christian.Koenig@amd.com \
    --cc=Prike.Liang@amd.com \
    --cc=Vitaly.Prosyak@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.