From: Prike Liang <Prike.Liang@amd.com>
To: <amd-gfx@lists.freedesktop.org>
Cc: <Alexander.Deucher@amd.com>, <Christian.Koenig@amd.com>,
<Vitaly.Prosyak@amd.com>, Prike Liang <Prike.Liang@amd.com>
Subject: [PATCH 15/18] drm/amdgpu/jpeg: skip scheduling jpeg/vcn idle_work during GPU reset
Date: Wed, 2 Sep 2026 20:49:58 +0800 [thread overview]
Message-ID: <20260902125001.621629-15-Prike.Liang@amd.com> (raw)
In-Reply-To: <20260902125001.621629-1-Prike.Liang@amd.com>
Scheduling jpeg idle_work for JPEG/VCN power gating during GPU reset can
trigger a register access assert before the GPU reset semaphore is
released, causing the following error during reset resume:
210] Workqueue: events amdgpu_jpeg_idle_work_handler [amdgpu]
[ 1576.787453] RIP: 0010:amdgpu_device_skip_hw_access+0x73/0x90 [amdgpu]
[ 1576.787656] Code: 85 c0 75 2a 8b 05 71 56 97 f0 85 c0 74 d3 48 8b bb d0 e7 07 00 be ff ff ff ff 48 81 c7 88 00 00 00 e8 f1 50 39 ef 85 c0 75 b7 <0f> 0b eb b3 48 8b bb d0 e7 07 00 48 83 c7 18 e8 39 fe 36 ee eb a1
[ 1576.787661] RSP: 0018:ffffccf700e6bcf0 EFLAGS: 00010246
[ 1576.787668] RAX: 0000000000000000 RBX: ffff89cf52d80000 RCX: 0000000000000002
[ 1576.787673] RDX: 0000000000000000 RSI: ffff89cf04f56bc8 RDI: ffff89cf02fe8f98
[ 1576.787677] RBP: ffffccf700e6bd00 R08: 0000000000000000 R09: 0000000000000001
[ 1576.787681] R10: ffffccf700e6bdb0 R11: ffffffffc1635c96 R12: 0000000000000000
[ 1576.787686] R13: 00000000000084d2 R14: 000000000003ff01 R15: 0000000000000000
[ 1576.787690] FS: 0000000000000000(0000) GS:ffff89d29a37d000(0000) knlGS:0000000000000000
[ 1576.787695] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 1576.787700] CR2: 00007fea0cc1ca50 CR3: 00000002e6c42000 CR4: 0000000000350ef0
[ 1576.787704] Call Trace:
[ 1576.787709] <TASK>
[ 1576.787717] amdgpu_device_wreg+0x26/0x50 [amdgpu]
[ 1576.787925] jpeg_v4_0_stop+0x47/0x140 [amdgpu]
[ 1576.788170] jpeg_v4_0_set_powergating_state+0x53/0x70 [amdgpu]
[ 1576.788410] amdgpu_device_ip_set_powergating_state+0x67/0xc0 [amdgpu]
[ 1576.788642] amdgpu_jpeg_idle_work_handler+0x105/0x120 [amdgpu]
[ 1576.788887] process_one_work+0x23e/0x6f0
[ 1576.788917] worker_thread+0x1c4/0x380
[ 1576.788931] kthread+0x10c/0x150
[ 1576.788937] ? __pfx_worker_thread+0x10/0x10
[ 1576.788943] ? __pfx_kthread+0x10/0x10
[ 1576.788954] ret_from_fork+0x314/0x390
[ 1576.788960] ? __pfx_kthread+0x10/0x10
[ 1576.788969] ret_from_fork_asm+0x1a/0x30
[ 1576.789003] </TASK>
Signed-off-by: Prike Liang <Prike.Liang@amd.com>
---
drivers/gpu/drm/amd/amdgpu/amdgpu_jpeg.c | 5 ++++-
drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c | 3 ++-
2 files changed, 6 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_jpeg.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_jpeg.c
index 208566ffe898..a66da05cc3f7 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_jpeg.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_jpeg.c
@@ -145,7 +145,10 @@ void amdgpu_jpeg_ring_begin_use(struct amdgpu_ring *ring)
void amdgpu_jpeg_ring_end_use(struct amdgpu_ring *ring)
{
- if (atomic_dec_and_test(&ring->adev->jpeg.total_submission_cnt))
+ struct amdgpu_device *adev = ring->adev;
+
+ if (atomic_dec_and_test(&ring->adev->jpeg.total_submission_cnt) &&
+ !amdgpu_in_reset(adev))
schedule_delayed_work(&ring->adev->jpeg.idle_work,
JPEG_IDLE_TIMEOUT);
}
diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c
index 17db7264269e..6cf08e35dc9b 100644
--- a/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c
+++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c
@@ -549,7 +549,8 @@ void amdgpu_vcn_ring_end_use(struct amdgpu_ring *ring)
!adev->vcn.inst[ring->me].using_unified_queue)
atomic_dec(&ring->adev->vcn.inst[ring->me].dpg_enc_submission_cnt);
- if (atomic_dec_and_test(&ring->adev->vcn.inst[ring->me].total_submission_cnt))
+ if (atomic_dec_and_test(&ring->adev->vcn.inst[ring->me].total_submission_cnt) &&
+ !amdgpu_in_reset(adev))
schedule_delayed_work(&ring->adev->vcn.inst[ring->me].idle_work,
VCN_IDLE_TIMEOUT);
}
--
2.34.1
next prev parent reply other threads:[~2026-09-02 12:50 UTC|newest]
Thread overview: 28+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 12:49 [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Prike Liang
2026-09-02 12:49 ` [PATCH 02/18] drm/amdgpu: clean up the userq support redundant check Prike Liang
2026-09-03 19:29 ` Alex Deucher
2026-09-02 12:49 ` [PATCH 03/18] drm/amdgpu: remove drm_client suspend-resume in the gpu recovery Prike Liang
2026-09-13 20:48 ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 04/18] drm/amdgpu: move userq fence wait out of signalling section Prike Liang
2026-09-15 1:49 ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 05/18] drm/amdgpu: defer userq reset after eviction failure Prike Liang
2026-09-02 12:49 ` [PATCH 06/18] drm/amdgpu: skip DRM internal suspend/resume for reseting VKMS Prike Liang
2026-09-02 12:49 ` [PATCH 07/18] drm/amdgpu: serialize userq eviction with GPU reset Prike Liang
2026-09-13 20:53 ` vitaly prosyak
2026-09-02 12:49 ` [PATCH 08/18] drm/amdgpu: skip VMHUB HW access in unaccessiable device Prike Liang
2026-09-02 12:49 ` [PATCH 09/18] drm/amdgpu/userq: complete the hang userq fence Prike Liang
2026-09-02 12:49 ` [PATCH 10/18] drm/amdgpu/mes: put the mes context BO allocation in mes sw_int Prike Liang
2026-09-02 12:49 ` [PATCH 11/18] drm/amdgpu: depart ring scheduler after resumming IP blocks Prike Liang
2026-09-02 12:49 ` [PATCH 12/18] drm/amdgpu: don't block wait gpu reset whthin userq lock Prike Liang
2026-09-28 22:49 ` vitaly prosyak
2026-09-29 3:44 ` Liang, Prike
2026-09-29 8:56 ` Christian König
2026-09-02 12:49 ` [PATCH 13/18] drm/amdgpu: allocate dma_fence slot explicitly for rearming eviction fence Prike Liang
2026-09-02 12:49 ` [PATCH 14/18] drm/amdgpu: skip gfx switch_power_profile during GPU reset Prike Liang
2026-09-03 19:25 ` Alex Deucher
2026-09-02 12:49 ` Prike Liang [this message]
2026-09-02 12:49 ` [PATCH 16/18] drm/amdgpu/mes: skip userq_notify_unmap during gpu reset Prike Liang
2026-09-02 12:50 ` [PATCH 17/18] drm/amdgpu/userq: complete userq eviction fence in pre_reset Prike Liang
2026-09-02 12:50 ` [PATCH 18/18] drm/amdgpu: stop the userq submission prior to removing userq Prike Liang
2026-09-03 19:28 ` [PATCH 01/18] drm/amdgpu: Remove separate guilty compute userq reset Alex Deucher
2026-09-07 6:48 ` Liang, Prike
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260902125001.621629-15-Prike.Liang@amd.com \
--to=prike.liang@amd.com \
--cc=Alexander.Deucher@amd.com \
--cc=Christian.Koenig@amd.com \
--cc=Vitaly.Prosyak@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.