From: "Timur Kristóf" <timur.kristof@gmail.com>
To: amd-gfx@lists.freedesktop.org, Alexander.Deucher@amd.com,
"Christian König" <christian.koenig@amd.com>,
"Natalie Vock" <natalie.vock@gmx.de>,
"Tvrtko Ursulin" <tursulin@ursulin.net>,
"Felix Kuehling" <Felix.Kuehling@amd.com>,
"Lazar, Lijo" <lijo.lazar@amd.com>
Subject: Re: [PATCH 7/8] drm/amdgpu/sdma: Always handle kernel queues in amdgpu_sdma_reset_engine()
Date: Mon, 07 Sep 2026 20:34:48 +0200 [thread overview]
Message-ID: <ZBE1nOSpRoOewoLNHFTFIA@gmail.com> (raw)
In-Reply-To: <e6270791-f276-487e-abc7-4ee1a529f357@amd.com>
On 2026. szeptember 7., hétfő 9:17:12 közép-európai nyári idő Lazar, Lijo
wrote:
> On 04-Sep-26 12:58 PM, Timur Kristóf wrote:
> > Remove the caller_handles_kernel_queues argument from
> > the amdgpu_sdma_reset_engine() function and make it
> > always handle kernel queues.
> >
> > Now the SDMA recovery sequence is more consistent
> > between callers for the KFD as follows.
> > Before recovery: first the KFD is suspended,
> > then the SDMA queue contents are backed up.
> > After recovery: first the SDMA queue contents
> > are restored, then the KFD is resumed.
> >
> > Signed-off-by: Timur Kristóf <timur.kristof@gmail.com>
> > ---
> >
> > drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c | 68 ++++++++++---------
> > drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h | 3 +-
> > drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c | 2 +-
> > .../drm/amd/amdkfd/kfd_device_queue_manager.c | 2 +-
> > 4 files changed, 38 insertions(+), 37 deletions(-)
> >
> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c
> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c index
> > 9eebd8380834..07aac5b3ea92 100644
> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c
> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c
> > @@ -542,16 +542,14 @@ static int amdgpu_sdma_soft_reset(struct
> > amdgpu_device *adev, u32 instance_id)>
> > }
> >
> > /**
> >
> > - * amdgpu_sdma_reset_engine - Reset a specific SDMA engine
> > + * amdgpu_sdma_reset_engine() - Reset a specific SDMA engine instance.
> > + *
> >
> > * @adev: Pointer to the AMDGPU device
> > * @instance_id: Logical ID of the SDMA engine instance to reset
> >
> > - * @caller_handles_kernel_queues: Skip kernel queue processing. Caller
> > - * will handle it.
> >
> > *
> > * Returns: 0 on success, or a negative error code on failure.
> > */
> >
> > -int amdgpu_sdma_reset_engine(struct amdgpu_device *adev, uint32_t
> > instance_id, - bool
caller_handles_kernel_queues)
> > +int amdgpu_sdma_reset_engine(struct amdgpu_device *adev, uint32_t
> > instance_id)>
> > {
> >
> > struct amdgpu_sdma_instance *sdma_instance =
> > &adev->sdma.instance[instance_id]; struct amdgpu_ring *gfx_ring =
> > &sdma_instance->ring;
> >
> > @@ -564,20 +562,23 @@ int amdgpu_sdma_reset_engine(struct amdgpu_device
> > *adev, uint32_t instance_id,>
> > mutex_lock(&sdma_instance->engine_reset_mutex);
> >
> > - if (!caller_handles_kernel_queues) {
> > - /* Stop the scheduler's work queue for the GFX and page
rings if they
> > are running. - * This ensures that no new tasks are
submitted to the
> > queues while - * the reset is in progress.
> > - */
> > + /*
> > + * Stop the scheduler's work queue for the GFX and page rings if
they
> > are running. + * This ensures that no new tasks are submitted to
the
> > queues while + * the reset is in progress.
> > + */
> > + if (amdgpu_ring_sched_ready(gfx_ring) &&
> > !drm_sched_is_stopped(&gfx_ring->sched))>
> > drm_sched_wqueue_stop(&gfx_ring->sched);
> >
> > - gfx_fence = amdgpu_ring_find_guilty_fence(gfx_ring);
> > - amdgpu_ring_reset_helper_begin(gfx_ring, gfx_fence);
> >
> > - if (adev->sdma.has_page_queue) {
> > + gfx_fence = amdgpu_ring_find_guilty_fence(gfx_ring);
> > + amdgpu_ring_reset_helper_begin(gfx_ring, gfx_fence);
> > +
> > + if (adev->sdma.has_page_queue) {
> > + if (amdgpu_ring_sched_ready(page_ring) &&
> > !drm_sched_is_stopped(&page_ring->sched))>
> > drm_sched_wqueue_stop(&page_ring->sched);
> >
> > - page_fence =
amdgpu_ring_find_guilty_fence(page_ring);
> > - amdgpu_ring_reset_helper_begin(page_ring,
page_fence);
> > - }
> > +
> > + page_fence = amdgpu_ring_find_guilty_fence(page_ring);
> > + amdgpu_ring_reset_helper_begin(page_ring, page_fence);
> >
> > }
>
> Since this resets the engine, a different way may be to have something
> like below (similar to amdgpu_multi_ring_reset_helper_begin) which takes
> care of all rings in the engine instance.
>
> amdgpu_ring_engine_reset_helper_begin(guilty_ring, guilty_fence);
>
> amdgpu_ring_engine_reset_helper_end(guilty_ring, guilty_fence);
>
> ring_type = guilty_ring->funcs->type;
> eng_instance = guilty_ring->me
>
> Thanks,
> Lijo
I am planning to do exactly that, but I want to keep this series short and
focused on unifying the code paths for SDMA v4.4.2 and v5.x.
Is it OK if I do that in a follow-up series?
Thanks,
Timur
>
> > if (sdma_instance->funcs->stop_kernel_queue) {
> >
> > @@ -612,22 +613,25 @@ int amdgpu_sdma_reset_engine(struct amdgpu_device
> > *adev, uint32_t instance_id,>
> > }
> >
> > exit:
> > - if (!caller_handles_kernel_queues) {
> > - /* Restart the scheduler's work queue for the GFX and
page rings
> > - * if they were stopped by this function. This allows
new tasks
> > - * to be submitted to the queues after the reset is
complete.
> > - */
> > - if (!ret) {
> > - ret = amdgpu_ring_reset_helper_end(gfx_ring,
gfx_fence);
> > + /* Restart the scheduler's work queue for the GFX and page rings
> > + * if they were stopped by this function. This allows new tasks
> > + * to be submitted to the queues after the reset is complete.
> > + */
> > + if (!ret) {
> > + ret = amdgpu_ring_reset_helper_end(gfx_ring,
gfx_fence);
> > + if (ret)
> > + goto unlock;
> > +
> > + if (amdgpu_ring_sched_ready(gfx_ring))
> > + drm_sched_wqueue_start(&gfx_ring->sched);
> > +
> > + if (adev->sdma.has_page_queue) {
> > + ret =
amdgpu_ring_reset_helper_end(page_ring, page_fence);
> >
> > if (ret)
> >
> > goto unlock;
> >
> > - drm_sched_wqueue_start(&gfx_ring->sched);
> > - if (adev->sdma.has_page_queue) {
> > - ret =
amdgpu_ring_reset_helper_end(page_ring, page_fence);
> > - if (ret)
> > - goto unlock;
> > +
> > + if (amdgpu_ring_sched_ready(page_ring))
> >
> >
drm_sched_wqueue_start(&page_ring->sched);
> >
> > - }
> >
> > }
> >
> > }
> >
> > unlock:
> > @@ -662,13 +666,11 @@ int amdgpu_sdma_reset_queue_legacy(struct
> > amdgpu_ring *ring,>
> > return -EINVAL;
> >
> > }
> >
> > - amdgpu_ring_reset_helper_begin(ring, timedout_fence);
> > -
> >
> > amdgpu_amdkfd_suspend(adev, true);
> >
> > - r = amdgpu_sdma_reset_engine(adev, ring->me, true);
> > + r = amdgpu_sdma_reset_engine(adev, ring->me);
> >
> > amdgpu_amdkfd_resume(adev, true);
> > if (r)
> >
> > return r;
> >
> > - return amdgpu_ring_reset_helper_end(ring, timedout_fence);
> > + return 0;
> >
> > }
> >
> > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h
> > b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h index
> > cb41453c1a19..5709d438e824 100644
> > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h
> > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h
> > @@ -153,8 +153,7 @@ struct amdgpu_buffer_funcs {
> >
> > uint32_t byte_count);
> >
> > };
> >
> > -int amdgpu_sdma_reset_engine(struct amdgpu_device *adev, uint32_t
> > instance_id, - bool
caller_handles_kernel_queues);
> > +int amdgpu_sdma_reset_engine(struct amdgpu_device *adev, uint32_t
> > instance_id);>
> > int amdgpu_sdma_reset_queue_legacy(struct amdgpu_ring *ring,
> >
> > unsigned int vmid,
> >
> > diff --git a/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c
> > b/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c index
> > 77f385b9ef53..796ea9f74763 100644
> > --- a/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c
> > +++ b/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c
> > @@ -1583,7 +1583,7 @@ static int sdma_v4_4_2_reset_queue(struct
> > amdgpu_ring *ring,>
> > int r;
> >
> > amdgpu_amdkfd_suspend(adev, true);
> >
> > - r = amdgpu_sdma_reset_engine(adev, id, false);
> > + r = amdgpu_sdma_reset_engine(adev, id);
> >
> > amdgpu_amdkfd_resume(adev, true);
> > return r;
> >
> > }
> >
> > diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
> > b/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c index
> > a23384571193..f9fbc5628fb4 100644
> > --- a/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
> > +++ b/drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c
> > @@ -2576,7 +2576,7 @@ static int reset_hung_queues_sdma(struct
> > device_queue_manager *dqm)>
> > continue;
> >
> > /* Reset engine and check. */
> >
> > - if (amdgpu_sdma_reset_engine(dqm->dev->adev,
i, false) ||
> > + if (amdgpu_sdma_reset_engine(dqm->dev->adev,
i) ||
> >
> > dqm->dev->kfd2kgd-
>hqd_sdma_get_doorbell(dqm->dev->adev, i, j) ||
> > !set_sdma_queue_as_reset(dqm,
doorbell_off)) {
> >
> > r = -ENOTRECOVERABLE;
next prev parent reply other threads:[~2026-09-07 18:34 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-04 7:28 [PATCH 0/8] drm/amdgpu/sdma: Improve existing SDMA queue resets (v3) Timur Kristóf
2026-09-04 7:28 ` [PATCH 1/8] drm/amdgpu/sdma: Remove unimplemented soft_reset() for SDMA and SI DMA Timur Kristóf
2026-09-04 7:28 ` [PATCH 2/8] drm/amdgpu/sdma: Remove superfluous rlc_resume and rlc_stop functions Timur Kristóf
2026-09-04 7:28 ` [PATCH 3/8] drm/amdgpu/sdma: Fix executing duplicate commands after recovery on SDMA v4.4.2 Timur Kristóf
2026-09-04 7:28 ` [PATCH 4/8] drm/amdgpu/sdma: Remove unnecessary guilty tracking of SDMA queues Timur Kristóf
2026-09-04 7:28 ` [PATCH 5/8] drm/amdgpu/sdma: Clear SDMA rings after reset before starting them Timur Kristóf
2026-09-07 7:04 ` Lazar, Lijo
2026-09-07 18:31 ` Timur Kristóf
2026-09-04 7:28 ` [PATCH 6/8] drm/amdgpu/sdma: Move SDMA v5.x queue reset to common code Timur Kristóf
2026-09-04 7:28 ` [PATCH 7/8] drm/amdgpu/sdma: Always handle kernel queues in amdgpu_sdma_reset_engine() Timur Kristóf
2026-09-07 7:17 ` Lazar, Lijo
2026-09-07 18:34 ` Timur Kristóf [this message]
2026-09-08 3:25 ` Lazar, Lijo
2026-09-04 7:28 ` [PATCH 8/8] drm/amdgpu/sdma: Use common SDMA legacy queue reset on SDMA v4.4.2 Timur Kristóf
-- strict thread matches above, loose matches on Subject: below --
2026-09-01 8:41 [PATCH 0/8] drm/amdgpu/sdma: Improve existing SDMA queue resets (v2) Timur Kristóf
2026-09-01 8:41 ` [PATCH 7/8] drm/amdgpu/sdma: Always handle kernel queues in amdgpu_sdma_reset_engine() Timur Kristóf
2026-09-03 20:43 ` Alex Deucher
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ZBE1nOSpRoOewoLNHFTFIA@gmail.com \
--to=timur.kristof@gmail.com \
--cc=Alexander.Deucher@amd.com \
--cc=Felix.Kuehling@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
--cc=christian.koenig@amd.com \
--cc=lijo.lazar@amd.com \
--cc=natalie.vock@gmx.de \
--cc=tursulin@ursulin.net \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.