AMD-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Pierre-Eric Pelloux-Prayer <pierre-eric@damsy.net>
To: Alex Deucher <alexdeucher@gmail.com>
Cc: Alex Deucher <alexander.deucher@amd.com>,
	amd-gfx@lists.freedesktop.org, christian.koenig@amd.com
Subject: Re: [PATCH 05/12] drm/amdgpu: don't call drm_sched_stop/start() in asic reset
Date: Thu, 5 Feb 2026 16:21:01 +0100	[thread overview]
Message-ID: <32a311db-e50c-49cc-a9da-95ae36ab0126@damsy.net> (raw)
In-Reply-To: <CADnq5_OoDPEy2PM5YUmOWU8k8rLk9UBD88oU5rCndh=Hovcu_Q@mail.gmail.com>



Le 05/02/2026 à 15:26, Alex Deucher a écrit :
> On Thu, Feb 5, 2026 at 9:22 AM Pierre-Eric Pelloux-Prayer
> <pierre-eric@damsy.net> wrote:
>>
>>
>>
>> Le 30/01/2026 à 18:30, Alex Deucher a écrit :
>>> We only want to stop the work queues, not mess with the
>>> fences, etc.
>>>
>>> v2: add the job back to the pending list.
>>> v3: return the proper job status so scheduler adds the
>>>       job back to the pending list
>>>
>>> Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
>>> ---
>>>    drivers/gpu/drm/amd/amdgpu/amdgpu_device.c | 4 ++--
>>>    drivers/gpu/drm/amd/amdgpu/amdgpu_job.c    | 6 ++----
>>>    2 files changed, 4 insertions(+), 6 deletions(-)
>>>
>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
>>> index e69ab8a923e31..a5b43d57c7b05 100644
>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
>>> @@ -6313,7 +6313,7 @@ static void amdgpu_device_halt_activities(struct amdgpu_device *adev,
>>>                        if (!amdgpu_ring_sched_ready(ring))
>>>                                continue;
>>>
>>> -                     drm_sched_stop(&ring->sched, job ? &job->base : NULL);
>>> +                     drm_sched_wqueue_stop(&ring->sched);
>>>
>>>                        if (need_emergency_restart)
>>>                                amdgpu_job_stop_all_jobs_on_sched(&ring->sched);
>>> @@ -6397,7 +6397,7 @@ static int amdgpu_device_sched_resume(struct list_head *device_list,
>>>                        if (!amdgpu_ring_sched_ready(ring))
>>>                                continue;
>>>
>>> -                     drm_sched_start(&ring->sched, 0);
>>> +                     drm_sched_wqueue_start(&ring->sched);
>>>                }
>>>
>>>                if (!drm_drv_uses_atomic_modeset(adev_to_drm(tmp_adev)) && !job_signaled)
>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c
>>> index df06a271bdf99..cd0707737a29b 100644
>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c
>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_job.c
>>> @@ -92,7 +92,6 @@ static enum drm_gpu_sched_stat amdgpu_job_timedout(struct drm_sched_job *s_job)
>>>        struct drm_wedge_task_info *info = NULL;
>>>        struct amdgpu_task_info *ti = NULL;
>>>        struct amdgpu_device *adev = ring->adev;
>>> -     enum drm_gpu_sched_stat status = DRM_GPU_SCHED_STAT_RESET;
>>>        int idx, r;
>>>
>>>        if (!drm_dev_enter(adev_to_drm(adev), &idx)) {
>>> @@ -147,8 +146,6 @@ static enum drm_gpu_sched_stat amdgpu_job_timedout(struct drm_sched_job *s_job)
>>>                                ring->sched.name);
>>>                        drm_dev_wedged_event(adev_to_drm(adev),
>>>                                             DRM_WEDGE_RECOVERY_NONE, info);
>>> -                     /* This is needed to add the job back to the pending list */
>>> -                     status = DRM_GPU_SCHED_STAT_NO_HANG;
>>>                        goto exit;
>>>                }
>>>                dev_err(adev->dev, "Ring %s reset failed\n", ring->sched.name);
>>> @@ -184,7 +181,8 @@ static enum drm_gpu_sched_stat amdgpu_job_timedout(struct drm_sched_job *s_job)
>>>    exit:
>>>        amdgpu_vm_put_task_info(ti);
>>>        drm_dev_exit(idx);
>>> -     return status;
>>> +     /* This is needed to add the job back to the pending list */
>>> +     return DRM_GPU_SCHED_STAT_NO_HANG;
>>
>> This part seems unrelated to the patch and is overwriting what was done
>> in patch 1/12.
> 
> Patch 1 fixes the pending list handling for per queue resets.  This
> patch reworks the adapter reset path to match the behavior of the per
> queue reset path.  After this patch they match so we can safely return
> DRM_GPU_SCHED_STAT_NO_HANG in both cases.  Previously the adapter
> reset path called drm_sched_wqueue_stop()/start() which handles
> re-adding the job to the pending list.  Since it no longer does, we
> need to return DRM_GPU_SCHED_STAT_NO_HANG for both cases.

I looked a bit more in the patchset adding DRM_GPU_SCHED_STAT_NO_HANG 
and your changes make sense. This patch is:

Acked-by: Pierre-Eric Pelloux-Prayer <pierre-eric.pelloux-prayer@amd.com>


  reply	other threads:[~2026-02-05 15:21 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-01-30 17:30 [PATCH 00/12] Improvements for IB handling V8 Alex Deucher
2026-01-30 17:30 ` [PATCH 01/12] drm/amdgpu: re-add the bad job to the pending list for ring resets Alex Deucher
2026-02-05 13:34   ` Pierre-Eric Pelloux-Prayer
2026-01-30 17:30 ` [PATCH 02/12] drm/amdgpu/job: use GFP_ATOMIC while in gpu reset Alex Deucher
2026-01-30 17:30 ` [PATCH 03/12] drm/amdgpu: switch all IPs to using job for IBs Alex Deucher
2026-02-02 17:16   ` Alex Deucher
2026-02-05 13:32   ` Pierre-Eric Pelloux-Prayer
2026-02-05 14:20     ` Alex Deucher
2026-02-05 14:43       ` Pierre-Eric Pelloux-Prayer
2026-02-05 15:16       ` Tvrtko Ursulin
2026-01-30 17:30 ` [PATCH 04/12] drm/amdgpu: require a job to schedule an IB Alex Deucher
2026-01-30 17:30 ` [PATCH 05/12] drm/amdgpu: don't call drm_sched_stop/start() in asic reset Alex Deucher
2026-02-05 14:02   ` Pierre-Eric Pelloux-Prayer
2026-02-05 14:26     ` Alex Deucher
2026-02-05 15:21       ` Pierre-Eric Pelloux-Prayer [this message]
2026-01-30 17:30 ` [PATCH 06/12] drm/amdgpu/cs: return -ETIME for guilty contexts Alex Deucher
2026-01-30 17:30 ` [PATCH 07/12] drm/amdgpu: plumb timedout fence through to force completion Alex Deucher
2026-01-30 17:30 ` [PATCH 08/12] drm/amdgpu: simplify VCN reset helper Alex Deucher
2026-01-30 17:30 ` [PATCH 09/12] drm/amdgpu: Call drm_sched_increase_karma() for ring resets Alex Deucher
2026-01-30 17:30 ` [PATCH 10/12] drm/amdgpu: reorder IB schedule sequence Alex Deucher
2026-02-05 14:16   ` Pierre-Eric Pelloux-Prayer
2026-01-30 17:30 ` [PATCH 11/12] drm/amdgpu: add a helper to calculate ring distance Alex Deucher
2026-02-05 14:18   ` Pierre-Eric Pelloux-Prayer
2026-01-30 17:30 ` [PATCH 12/12] drm/amdgpu: rework ring reset backup and reemit v8 Alex Deucher
  -- strict thread matches above, loose matches on Subject: below --
2026-01-29 20:37 [PATCH 00/12] Improvements for IB handling V7 Alex Deucher
2026-01-29 20:37 ` [PATCH 05/12] drm/amdgpu: don't call drm_sched_stop/start() in asic reset Alex Deucher

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=32a311db-e50c-49cc-a9da-95ae36ab0126@damsy.net \
    --to=pierre-eric@damsy.net \
    --cc=alexander.deucher@amd.com \
    --cc=alexdeucher@gmail.com \
    --cc=amd-gfx@lists.freedesktop.org \
    --cc=christian.koenig@amd.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox