Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH] drm/xe: Skip schedule disable for an already reset exec queue
@ 2026-09-17 19:38 Jagmeet Randhawa
  2026-09-17 20:10 ` Matthew Brost
                   ` (4 more replies)
  0 siblings, 5 replies; 7+ messages in thread
From: Jagmeet Randhawa @ 2026-09-17 19:38 UTC (permalink / raw)
  To: intel-xe
  Cc: matthew.brost, stuart.summers, niranjana.vishwanathapura,
	Jagmeet Randhawa

A memory CAT error causes GuC to reset the offending context and notify
the driver. xe_guc_exec_queue_memory_cat_error_handler() treats this the
same as an engine reset and calls xe_guc_exec_queue_reset_trigger_cleanup(),
which sets EXEC_QUEUE_STATE_RESET.

Roughly 200us later the TDR runs for the same queue. It reads
exec_queue_reset(), uses the result only to set err = -EIO, and then sends
a schedule disable (H2G SCHED_CONTEXT_MODE_SET) anyway. GuC has already
destroyed the context, so no SCHED_DONE is ever returned. The TDR then
waits out its HZ*5 timeout, reports "Schedule disable failed to respond"
and calls xe_gt_reset_async(), escalating a per-context failure into a
full GT reset that kills every context on the tile and reloads GuC.

The failure is intermittent because the enclosing guard is
(exec_queue_enabled || exec_queue_pending_disable): it is a race between
the CAT error cleanup and the TDR. When cleanup wins the block is skipped
and recovery is clean; when the TDR wins the disable is sent and the
timeout fires.

Skip the handshake entirely when the context has already been reset. The
send, the HZ*5 wait and the GT reset escalation all live in the same block,
so a single condition removes all three. err = -EIO is preserved so
userspace still receives the correct error, and execution falls through to
the normal cleanup path, which performs the deregistration.

This matches existing behaviour elsewhere in the driver:
xe_guc_exec_queue_reset_handler() calls the same
xe_guc_exec_queue_reset_trigger_cleanup() and performs no disable
handshake at all. The TDR path was simply inconsistent with it.

Evidence, captured with existing ftrace events only (no instrumentation):

  Healthy acknowledgements take ~350us:
    4806.244755 xe_exec_queue_scheduling_enable  guc_id=2 guc_state=0x7
    4806.245126 xe_exec_queue_scheduling_done    guc_id=2 guc_state=0x7

  The failure:
    4806.245395 xe_exec_queue_memory_cat_error   guc_id=2 guc_state=0x3
    4806.245592 xe_exec_queue_scheduling_disable guc_id=2 guc_state=0x249
                <5.18s, no scheduling_done for any guc_id>
    4811.425872 xe_sched_job_timedout            guc_id=2 guc_state=0x200

  guc_state=0x249 is REGISTERED|PENDING_DISABLE|RESET|BANNED. The RESET bit
  confirms the exec_queue_reset() branch was taken and the disable was sent
  regardless. scheduling_done never fires, so handle_sched_done() never runs
  and the acknowledgement genuinely never arrives.

  With the fix, scheduling_disable is traced at guc_state=0x2d9
  (adds DESTROYED|KILLED), i.e. it now originates from
  disable_scheduling_deregister() on the normal cleanup path, and GuC
  acknowledges it every time.

Testing:
  igt@xe_exec_reset@multi-queue-cancel               0/10 failures
                                                     (baseline 10/10 failures)
  igt@xe_exec_reset@multi-queue-cancel-on-secondary  passing
  Subtest runtime drops from ~5.2s to ~85ms.

Signed-off-by: Jagmeet Randhawa <jagmeet.randhawa@intel.com>
---
 drivers/gpu/drm/xe/xe_guc_submit.c | 14 ++++++++++----
 1 file changed, 10 insertions(+), 4 deletions(-)

diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
index f3ba8abfc228..03450590042a 100644
--- a/drivers/gpu/drm/xe/xe_guc_submit.c
+++ b/drivers/gpu/drm/xe/xe_guc_submit.c
@@ -1609,14 +1609,20 @@ guc_exec_queue_timedout_job(struct drm_sched_job *drm_job)
 		atomic_or(DRM_XE_EXEC_QUEUE_BAN_REASON_GPU_HANG, &q->ban_reason);
 	set_exec_queue_banned(q);
 
+	/*
+	 * GuC already reset this context (the CAT error handler treats a CAT
+	 * error as an engine reset), so there is nothing left to disable and
+	 * SCHED_DONE will never arrive. Skip the handshake rather than waiting
+	 * HZ*5 and escalating to a full GT reset.
+	 */
+	if (exec_queue_reset(primary))
+		err = -EIO;
+
 	/* Kick job / queue off hardware */
-	if (!xe_device_is_in_reset(xe) && !wedged &&
+	if (!xe_device_is_in_reset(xe) && !wedged && !exec_queue_reset(primary) &&
 	    (exec_queue_enabled(primary) || exec_queue_pending_disable(primary))) {
 		int ret;
 
-		if (exec_queue_reset(primary))
-			err = -EIO;
-
 		if (xe_uc_fw_is_running(&guc->fw)) {
 			/*
 			 * Wait for any pending G2H to flush out before
-- 
2.53.0


^ permalink raw reply related	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-09-18 18:31 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-17 19:38 [PATCH] drm/xe: Skip schedule disable for an already reset exec queue Jagmeet Randhawa
2026-09-17 20:10 ` Matthew Brost
2026-09-18 18:31   ` Randhawa, Jagmeet
2026-09-17 21:21 ` ✗ CI.checkpatch: warning for " Patchwork
2026-09-17 21:23 ` ✓ CI.KUnit: success " Patchwork
2026-09-17 22:29 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-18  2:42 ` ✓ Xe.CI.FULL: " Patchwork

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox