From: Pierre-Eric Pelloux-Prayer <pierre-eric@damsy.net>
To: Alex Deucher <alexdeucher@gmail.com>
Cc: Alex Deucher <alexander.deucher@amd.com>,
amd-gfx@lists.freedesktop.org, christian.koenig@amd.com
Subject: Re: [PATCH 05/12] drm/amdgpu: don't call drm_sched_stop/start() in asic reset
Date: Thu, 5 Feb 2026 16:21:01 +0100 [thread overview]
Message-ID: <32a311db-e50c-49cc-a9da-95ae36ab0126@damsy.net> (raw)
In-Reply-To: <CADnq5_OoDPEy2PM5YUmOWU8k8rLk9UBD88oU5rCndh=Hovcu_Q@mail.gmail.com>
Le 05/02/2026 à 15:26, Alex Deucher a écrit :
> On Thu, Feb 5, 2026 at 9:22 AM Pierre-Eric Pelloux-Prayer
> <pierre-eric@damsy.net> wrote:
>>
>>
>>
>> Le 30/01/2026 à 18:30, Alex Deucher a écrit :
>>> We only want to stop the work queues, not mess with the
>>> fences, etc.
>>>
>>> v2: add the job back to the pending list.
>>> v3: return the proper job status so scheduler adds the
>>> job back to the pending list
>>>
>>> Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
>>> ---
>>> drivers/gpu/drm/amd/amdgpu/amdgpu_device.c | 4 ++--
>>> drivers/gpu/drm/amd/amdgpu/amdgpu_job.c | 6 ++----
>>> 2 files changed, 4 insertions(+), 6 deletions(-)
>>>
>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
>>> index e69ab8a923e31..a5b43d57c7b05 100644
>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
>>> @@ -6313,7 +6313,7 @@ static void amdgpu_device_halt_activities(struct amdgpu_device *adev,
>>> if (!amdgpu_ring_sched_ready(ring))
>>> continue;
>>>
>>> - drm_sched_stop(&ring->sched, job ? &job->base : NULL);
>>> + drm_sched_wqueue_stop(&ring->sched);
>>>
>>> if (need_emergency_restart)
>>> amdgpu_job_stop_all_jobs_on_sched(&ring->sched);
>>> @@ -6397,7 +6397,7 @@ static int amdgpu_device_sched_resume(struct list_head *device_list,
>>> if (!amdgpu_ring_sched_ready(ring))
>>> continue;
>>>
>>> - drm_sched_start(&ring->sched, 0);
>>> + drm_sched_wqueue_start(&ring->sched);
>>> }
>>>
>>> if (!drm_drv_uses_atomic_modeset(adev_to_drm(tmp_adev)) && !job_signaled)
>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c
>>> index df06a271bdf99..cd0707737a29b 100644
>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c
>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c
>>> @@ -92,7 +92,6 @@ static enum drm_gpu_sched_stat amdgpu_job_timedout(struct drm_sched_job *s_job)
>>> struct drm_wedge_task_info *info = NULL;
>>> struct amdgpu_task_info *ti = NULL;
>>> struct amdgpu_device *adev = ring->adev;
>>> - enum drm_gpu_sched_stat status = DRM_GPU_SCHED_STAT_RESET;
>>> int idx, r;
>>>
>>> if (!drm_dev_enter(adev_to_drm(adev), &idx)) {
>>> @@ -147,8 +146,6 @@ static enum drm_gpu_sched_stat amdgpu_job_timedout(struct drm_sched_job *s_job)
>>> ring->sched.name);
>>> drm_dev_wedged_event(adev_to_drm(adev),
>>> DRM_WEDGE_RECOVERY_NONE, info);
>>> - /* This is needed to add the job back to the pending list */
>>> - status = DRM_GPU_SCHED_STAT_NO_HANG;
>>> goto exit;
>>> }
>>> dev_err(adev->dev, "Ring %s reset failed\n", ring->sched.name);
>>> @@ -184,7 +181,8 @@ static enum drm_gpu_sched_stat amdgpu_job_timedout(struct drm_sched_job *s_job)
>>> exit:
>>> amdgpu_vm_put_task_info(ti);
>>> drm_dev_exit(idx);
>>> - return status;
>>> + /* This is needed to add the job back to the pending list */
>>> + return DRM_GPU_SCHED_STAT_NO_HANG;
>>
>> This part seems unrelated to the patch and is overwriting what was done
>> in patch 1/12.
>
> Patch 1 fixes the pending list handling for per queue resets. This
> patch reworks the adapter reset path to match the behavior of the per
> queue reset path. After this patch they match so we can safely return
> DRM_GPU_SCHED_STAT_NO_HANG in both cases. Previously the adapter
> reset path called drm_sched_wqueue_stop()/start() which handles
> re-adding the job to the pending list. Since it no longer does, we
> need to return DRM_GPU_SCHED_STAT_NO_HANG for both cases.
I looked a bit more in the patchset adding DRM_GPU_SCHED_STAT_NO_HANG
and your changes make sense. This patch is:
Acked-by: Pierre-Eric Pelloux-Prayer <pierre-eric.pelloux-prayer@amd.com>
next prev parent reply other threads:[~2026-02-05 15:21 UTC|newest]
Thread overview: 25+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-01-30 17:30 [PATCH 00/12] Improvements for IB handling V8 Alex Deucher
2026-01-30 17:30 ` [PATCH 01/12] drm/amdgpu: re-add the bad job to the pending list for ring resets Alex Deucher
2026-02-05 13:34 ` Pierre-Eric Pelloux-Prayer
2026-01-30 17:30 ` [PATCH 02/12] drm/amdgpu/job: use GFP_ATOMIC while in gpu reset Alex Deucher
2026-01-30 17:30 ` [PATCH 03/12] drm/amdgpu: switch all IPs to using job for IBs Alex Deucher
2026-02-02 17:16 ` Alex Deucher
2026-02-05 13:32 ` Pierre-Eric Pelloux-Prayer
2026-02-05 14:20 ` Alex Deucher
2026-02-05 14:43 ` Pierre-Eric Pelloux-Prayer
2026-02-05 15:16 ` Tvrtko Ursulin
2026-01-30 17:30 ` [PATCH 04/12] drm/amdgpu: require a job to schedule an IB Alex Deucher
2026-01-30 17:30 ` [PATCH 05/12] drm/amdgpu: don't call drm_sched_stop/start() in asic reset Alex Deucher
2026-02-05 14:02 ` Pierre-Eric Pelloux-Prayer
2026-02-05 14:26 ` Alex Deucher
2026-02-05 15:21 ` Pierre-Eric Pelloux-Prayer [this message]
2026-01-30 17:30 ` [PATCH 06/12] drm/amdgpu/cs: return -ETIME for guilty contexts Alex Deucher
2026-01-30 17:30 ` [PATCH 07/12] drm/amdgpu: plumb timedout fence through to force completion Alex Deucher
2026-01-30 17:30 ` [PATCH 08/12] drm/amdgpu: simplify VCN reset helper Alex Deucher
2026-01-30 17:30 ` [PATCH 09/12] drm/amdgpu: Call drm_sched_increase_karma() for ring resets Alex Deucher
2026-01-30 17:30 ` [PATCH 10/12] drm/amdgpu: reorder IB schedule sequence Alex Deucher
2026-02-05 14:16 ` Pierre-Eric Pelloux-Prayer
2026-01-30 17:30 ` [PATCH 11/12] drm/amdgpu: add a helper to calculate ring distance Alex Deucher
2026-02-05 14:18 ` Pierre-Eric Pelloux-Prayer
2026-01-30 17:30 ` [PATCH 12/12] drm/amdgpu: rework ring reset backup and reemit v8 Alex Deucher
-- strict thread matches above, loose matches on Subject: below --
2026-01-29 20:37 [PATCH 00/12] Improvements for IB handling V7 Alex Deucher
2026-01-29 20:37 ` [PATCH 05/12] drm/amdgpu: don't call drm_sched_stop/start() in asic reset Alex Deucher
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=32a311db-e50c-49cc-a9da-95ae36ab0126@damsy.net \
--to=pierre-eric@damsy.net \
--cc=alexander.deucher@amd.com \
--cc=alexdeucher@gmail.com \
--cc=amd-gfx@lists.freedesktop.org \
--cc=christian.koenig@amd.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox