From: Jagmeet Randhawa <jagmeet.randhawa@intel.com>
To: intel-xe@lists.freedesktop.org
Cc: niranjana.vishwanathapura@intel.com,
Jagmeet Randhawa <jagmeet.randhawa@intel.com>
Subject: [PATCH] drm/xe/guc: Fix race around q->guc->suspend_pending access
Date: Fri, 14 Aug 2026 04:18:48 +0800 [thread overview]
Message-ID: <20260813201847.201867-2-jagmeet.randhawa@intel.com> (raw)
q->guc->suspend_pending is accessed without any common lock.
__suspend_fence_signal(), called from guc_exec_queue_kill() and the
suspend-timeout ban path, clears the flag asynchronously. Meanwhile
handle_sched_done(), guc_exec_queue_stop() and
__guc_exec_queue_process_msg_suspend() check the flag and then call
suspend_fence_signal(), which asserts that it is still set.
As the check and suspend_fence_signal() are not atomic, the clear can
land in between and trip the xe_gt_assert(q->guc->suspend_pending).
The flag is already set and read under the per-queue msg_lock
(xe_sched_msg_lock()) on the suspend and resume paths. Extend that same
lock to the clear paths (kill and ban) and to the three check-then-act
sites so the check and the signal are atomic with respect to the clear.
In __guc_exec_queue_process_msg_suspend() only the non-sleeping branch is
wrapped, since the other branch waits. In handle_sched_done() the flag is
snapshotted under the lock and deregister_exec_queue() is kept outside it.
Signed-off-by: Jagmeet Randhawa <jagmeet.randhawa@intel.com>
---
drivers/gpu/drm/xe/xe_guc_submit.c | 29 ++++++++++++++++++++++++-----
1 file changed, 24 insertions(+), 5 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
index 9036f89dff7d..c565c1d32d3a 100644
--- a/drivers/gpu/drm/xe/xe_guc_submit.c
+++ b/drivers/gpu/drm/xe/xe_guc_submit.c
@@ -1928,9 +1928,13 @@ static void __guc_exec_queue_process_msg_suspend(struct xe_sched_msg *msg)
set_exec_queue_suspended(q);
disable_scheduling(q, false);
}
- } else if (q->guc->suspend_pending) {
- set_exec_queue_suspended(q);
- suspend_fence_signal(q);
+ } else {
+ xe_sched_msg_lock(&q->guc->sched);
+ if (q->guc->suspend_pending) {
+ set_exec_queue_suspended(q);
+ suspend_fence_signal(q);
+ }
+ xe_sched_msg_unlock(&q->guc->sched);
}
}
@@ -2130,7 +2134,9 @@ static void guc_exec_queue_kill(struct xe_exec_queue *q)
{
trace_xe_exec_queue_kill(q);
set_exec_queue_killed(q);
+ xe_sched_msg_lock(&q->guc->sched);
__suspend_fence_signal(q);
+ xe_sched_msg_unlock(&q->guc->sched);
xe_guc_exec_queue_trigger_cleanup(q);
}
@@ -2392,11 +2398,15 @@ static void guc_exec_queue_suspend_timeout_ban(struct xe_exec_queue *q)
*/
if (xe_exec_queue_is_multi_queue(q)) {
set_exec_queue_group_banned(q);
+ xe_sched_msg_lock(&q->guc->sched);
__suspend_fence_signal(q);
+ xe_sched_msg_unlock(&q->guc->sched);
xe_guc_exec_queue_group_trigger_cleanup(q);
} else {
set_exec_queue_banned(q);
+ xe_sched_msg_lock(&q->guc->sched);
__suspend_fence_signal(q);
+ xe_sched_msg_unlock(&q->guc->sched);
xe_guc_exec_queue_trigger_cleanup(q);
}
}
@@ -2614,10 +2624,12 @@ static void guc_exec_queue_stop(struct xe_guc *guc, struct xe_exec_queue *q)
if (exec_queue_destroyed(q))
do_destroy = true;
}
+ xe_sched_msg_lock(sched);
if (q->guc->suspend_pending) {
set_exec_queue_suspended(q);
suspend_fence_signal(q);
}
+ xe_sched_msg_unlock(sched);
atomic_and(EXEC_QUEUE_STATE_WEDGED | EXEC_QUEUE_STATE_BANNED |
EXEC_QUEUE_STATE_KILLED | EXEC_QUEUE_STATE_DESTROYED |
EXEC_QUEUE_STATE_SUSPENDED,
@@ -3222,13 +3234,20 @@ static void handle_sched_done(struct xe_guc *guc, struct xe_exec_queue *q,
smp_wmb();
wake_up_all(&guc->ct.wq);
} else {
+ bool was_pending;
+
xe_gt_assert(guc_to_gt(guc), runnable_state == 0);
xe_gt_assert(guc_to_gt(guc), exec_queue_pending_disable(q));
- if (q->guc->suspend_pending) {
+ xe_sched_msg_lock(&q->guc->sched);
+ was_pending = q->guc->suspend_pending;
+ if (was_pending) {
clear_exec_queue_pending_disable(q);
suspend_fence_signal(q);
- } else {
+ }
+ xe_sched_msg_unlock(&q->guc->sched);
+
+ if (!was_pending) {
if (exec_queue_banned(q)) {
smp_wmb();
wake_up_all(&guc->ct.wq);
--
2.53.0
next reply other threads:[~2026-08-13 20:19 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-13 20:18 Jagmeet Randhawa [this message]
2026-08-13 20:28 ` ✓ CI.KUnit: success for drm/xe/guc: Fix race around q->guc->suspend_pending access Patchwork
2026-08-13 21:32 ` ✓ Xe.CI.BAT: " Patchwork
2026-08-14 1:01 ` [PATCH] " sashiko-bot
2026-08-14 1:22 ` ✓ Xe.CI.FULL: success for " Patchwork
2026-08-14 3:51 ` [PATCH] " Niranjana Vishwanathapura
2026-08-17 17:29 ` Randhawa, Jagmeet
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260813201847.201867-2-jagmeet.randhawa@intel.com \
--to=jagmeet.randhawa@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=niranjana.vishwanathapura@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox