All of lore.kernel.org
 help / color / mirror / Atom feed
From: Alex Deucher <alexander.deucher@amd.com>
To: <amd-gfx@lists.freedesktop.org>, <christian.koenig@amd.com>
Cc: "Christian König" <ckoenig.leichtzumerken@gmail.com>,
	"Alex Deucher" <alexander.deucher@amd.com>
Subject: [PATCH 06/31] drm/amdgpu: rework queue reset scheduler interaction
Date: Wed, 4 Jun 2025 21:45:36 -0400	[thread overview]
Message-ID: <20250605014602.5915-7-alexander.deucher@amd.com> (raw)
In-Reply-To: <20250605014602.5915-1-alexander.deucher@amd.com>

From: Christian König <ckoenig.leichtzumerken@gmail.com>

Stopping the scheduler for queue reset is generally a good idea because
it prevents any worker from touching the ring buffer.

But using amdgpu_fence_driver_force_completion() before restarting it was
a really bad idea because it marked fences as failed while the work was
potentially still running.

Stop doing that and cleanup the comment a bit.

v2: keep amdgpu_fence_driver_force_completion() for non-gfx rings
v3: drop amdgpu_fence_driver_force_completion() for compute ring
v4: avoid a warning when setting an error on the fence

Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
---
 drivers/gpu/drm/amd/amdgpu/amdgpu_job.c | 37 +++++++++++++++----------
 1 file changed, 22 insertions(+), 15 deletions(-)

diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c
index ddb9d3269357c..821f88b64f3f6 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c
@@ -91,8 +91,8 @@ static enum drm_gpu_sched_stat amdgpu_job_timedout(struct drm_sched_job *s_job)
 	struct amdgpu_job *job = to_amdgpu_job(s_job);
 	struct amdgpu_task_info *ti;
 	struct amdgpu_device *adev = ring->adev;
-	int idx;
-	int r;
+	bool set_error = false;
+	int idx, r;
 
 	if (!drm_dev_enter(adev_to_drm(adev), &idx)) {
 		dev_info(adev->dev, "%s - device unplugged skipping recovery on scheduler:%s",
@@ -136,10 +136,12 @@ static enum drm_gpu_sched_stat amdgpu_job_timedout(struct drm_sched_job *s_job)
 	} else if (amdgpu_gpu_recovery && ring->funcs->reset) {
 		bool is_guilty;
 
-		dev_err(adev->dev, "Starting %s ring reset\n", s_job->sched->name);
-		/* stop the scheduler, but don't mess with the
-		 * bad job yet because if ring reset fails
-		 * we'll fall back to full GPU reset.
+		dev_err(adev->dev, "Starting %s ring reset\n",
+			s_job->sched->name);
+
+		/*
+		 * Stop the scheduler to prevent anybody else from touching the
+		 * ring buffer.
 		 */
 		drm_sched_wqueue_stop(&ring->sched);
 
@@ -154,24 +156,29 @@ static enum drm_gpu_sched_stat amdgpu_job_timedout(struct drm_sched_job *s_job)
 
 		if (is_guilty)
 			dma_fence_set_error(&s_job->s_fence->finished, -ETIME);
+			set_error = true;
+		}
 
 		r = amdgpu_ring_reset(ring, job->vmid);
 		if (!r) {
-			if (amdgpu_ring_sched_ready(ring))
-				drm_sched_stop(&ring->sched, s_job);
 			if (is_guilty) {
 				atomic_inc(&ring->adev->gpu_reset_counter);
-				amdgpu_fence_driver_force_completion(ring);
+				if ((ring->funcs->type != AMDGPU_RING_TYPE_GFX) &&
+				    (ring->funcs->type != AMDGPU_RING_TYPE_COMPUTE))
+					amdgpu_fence_driver_force_completion(ring);
 			}
-			if (amdgpu_ring_sched_ready(ring))
-				drm_sched_start(&ring->sched, 0);
-			dev_err(adev->dev, "Ring %s reset succeeded\n", ring->sched.name);
-			drm_dev_wedged_event(adev_to_drm(adev), DRM_WEDGE_RECOVERY_NONE);
+			drm_sched_wqueue_start(&ring->sched);
+			dev_err(adev->dev, "Ring %s reset succeeded\n",
+				ring->sched.name);
+			drm_dev_wedged_event(adev_to_drm(adev),
+					     DRM_WEDGE_RECOVERY_NONE);
 			goto exit;
 		}
-		dev_err(adev->dev, "Ring %s reset failure\n", ring->sched.name);
+		dev_err(adev->dev, "Ring %s reset failed\n", ring->sched.name);
 	}
-	dma_fence_set_error(&s_job->s_fence->finished, -ETIME);
+
+	if (!set_error)
+		dma_fence_set_error(&s_job->s_fence->finished, -ETIME);
 
 	if (amdgpu_device_should_recover_gpu(ring->adev)) {
 		struct amdgpu_reset_context reset_context;
-- 
2.49.0


  parent reply	other threads:[~2025-06-05  1:46 UTC|newest]

Thread overview: 35+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2025-06-05  1:45 [PATCH V6 00/31] Reset improvements for GC10+ Alex Deucher
2025-06-05  1:45 ` [PATCH 01/31] drm/amdgpu: enable legacy enforce isolation by default Alex Deucher
2025-06-05  1:45 ` [PATCH 02/31] drm/amdgpu/gfx7: drop reset_kgq Alex Deucher
2025-06-05  1:45 ` [PATCH 03/31] drm/amdgpu/gfx8: " Alex Deucher
2025-06-05  1:45 ` [PATCH 04/31] drm/amdgpu/gfx9: " Alex Deucher
2025-06-05  1:45 ` [PATCH 05/31] drm/amdgpu: switch job hw_fence to amdgpu_fence Alex Deucher
2025-06-05  1:45 ` Alex Deucher [this message]
2025-06-05  1:45 ` [PATCH 07/31] drm/amdgpu: move force completion into ring resets Alex Deucher
2025-06-05  1:45 ` [PATCH 08/31] drm/amdgpu: track ring state associated with a job Alex Deucher
2025-06-05 12:11   ` Christian König
2025-06-05 13:21     ` Alex Deucher
2025-06-05 13:50       ` Christian König
2025-06-05  1:45 ` [PATCH 09/31] drm/amdgpu: optimize amdgpu_ring_reemit_unprocessed_jobs() Alex Deucher
2025-06-05  1:45 ` [PATCH 10/31] drm/amdgpu/gfx10: re-emit unprocessed state on ring reset Alex Deucher
2025-06-05  1:45 ` [PATCH 11/31] drm/amdgpu/gfx11: " Alex Deucher
2025-06-05  1:45 ` [PATCH 12/31] drm/amdgpu/gfx12: " Alex Deucher
2025-06-05  1:45 ` [PATCH 13/31] drm/amdgpu/gfx9: re-emit unprocessed state on kcq reset Alex Deucher
2025-06-05  1:45 ` [PATCH 14/31] drm/amdgpu/gfx9.4.3: " Alex Deucher
2025-06-05  1:45 ` [PATCH 15/31] drm/amdgpu/sdma4.4.2: re-emit unprocessed state on ring reset Alex Deucher
2025-06-05  1:45 ` [PATCH 16/31] drm/amdgpu/sdma5: " Alex Deucher
2025-06-05  1:45 ` [PATCH 17/31] drm/amdgpu/sdma5.2: " Alex Deucher
2025-06-05  1:45 ` [PATCH 18/31] drm/amdgpu/sdma6: " Alex Deucher
2025-06-05  1:45 ` [PATCH 19/31] drm/amdgpu/sdma7: " Alex Deucher
2025-06-05  1:45 ` [PATCH 20/31] drm/amdgpu/jpeg2: " Alex Deucher
2025-06-05  1:45 ` [PATCH 21/31] drm/amdgpu/jpeg2.5: " Alex Deucher
2025-06-05  1:45 ` [PATCH 22/31] drm/amdgpu/jpeg3: " Alex Deucher
2025-06-05  1:45 ` [PATCH 23/31] drm/amdgpu/jpeg4: " Alex Deucher
2025-06-05  1:45 ` [PATCH 24/31] drm/amdgpu/jpeg4.0.3: " Alex Deucher
2025-06-05  1:45 ` [PATCH 25/31] drm/amdgpu/jpeg5.0.0: add queue reset Alex Deucher
2025-06-05  1:45 ` [PATCH 26/31] drm/amdgpu/jpeg5: re-emit unprocessed state on ring reset Alex Deucher
2025-06-05  1:45 ` [PATCH 27/31] drm/amdgpu/jpeg5.0.1: " Alex Deucher
2025-06-05  1:45 ` [PATCH 28/31] drm/amdgpu/vcn4: " Alex Deucher
2025-06-05  1:45 ` [PATCH 29/31] drm/amdgpu/vcn4.0.3: " Alex Deucher
2025-06-05  1:46 ` [PATCH 30/31] drm/amdgpu/vcn4.0.5: " Alex Deucher
2025-06-05  1:46 ` [PATCH 31/31] drm/amdgpu/vcn5: " Alex Deucher

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20250605014602.5915-7-alexander.deucher@amd.com \
    --to=alexander.deucher@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    --cc=christian.koenig@amd.com \
    --cc=ckoenig.leichtzumerken@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.