Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 0/3] drm/xe: Fix ULLS chained job loss and GT reset replay
@ 2026-09-30  9:40 Varun Gupta
  2026-09-30  9:40 ` [PATCH 1/3] drm/xe: Park on the ULLS semaphore before publishing the next job's tail Varun Gupta
                   ` (6 more replies)
  0 siblings, 7 replies; 8+ messages in thread
From: Varun Gupta @ 2026-09-30  9:40 UTC (permalink / raw)
  To: intel-xe; +Cc: matthew.brost, thomas.hellstrom, himal.prasad.ghimiray

A chained ULLS migration job can be lost: its predecessor's postamble
publishes the ring tail over the job's slot before parking on the
semaphore, so the CS is free to fetch that slot while it still holds
padding. When the job is later written and the semaphore signalled,
the CS runs what it already fetched, drains to the tail and idles. A
lost ULLS_EXIT goes unnoticed, a lost ULLS_ACTIVE hangs the kernel
migration queue, and because the GT reset replay of a chained ULLS job
does not work either, the second timeout wedges the device.

Patch 1 is the fix: park first, publish the tail after the wait, so the
next slot stays beyond RING_TAIL until the job is in place. This puts
the non-posted tail write on the wake-up path, which the original
ordering was chosen to avoid. I could not find a way to keep the tail
ahead of the wait without the CS being able to fetch the slot early.

Patches 2 and 3 make the GT reset replay of a chained ULLS job work,
so a future ULLS hang degrades to a single recoverable reset rather
than a wedge. Patch 3 covers a state patch 1 eliminates and is defence
in depth.

A couple of related items I have left alone and would appreciate a
view on:

 - The SR-IOV VF pause/unpause replay only routes the last_replay job
   through the tail write, so a chained last job publishes nothing
   there either. Adding "|| job->last_replay" to the patch 2 condition
   looks right but I have no VF setup to test it for now.

 - At replay, pending chained jobs still have their semaphore slot
   signalled from before the reset, so the first re-emitted postamble
   passes its wait immediately. The resubmit loop writes every job
   before GuC processes the enable, so this has not been observed;
   clearing the slots in guc_exec_queue_start() would close it.

Varun Gupta (3):
  drm/xe: Park on the ULLS semaphore before publishing the next job's
    tail
  drm/xe/guc: Publish the ring tail when replaying a chained ULLS job
  drm/xe/guc: Rewind the LRC ring head when replaying a ULLS job

 drivers/gpu/drm/xe/xe_guc_submit.c | 19 +++++++++++++++++--
 drivers/gpu/drm/xe/xe_migrate.c    | 24 +++++++++++++-----------
 drivers/gpu/drm/xe/xe_ring_ops.c   | 15 +++++++++++----
 3 files changed, 41 insertions(+), 17 deletions(-)

-- 
2.43.0


^ permalink raw reply	[flat|nested] 8+ messages in thread

end of thread, other threads:[~2026-09-30 17:24 UTC | newest]

Thread overview: 8+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30  9:40 [PATCH 0/3] drm/xe: Fix ULLS chained job loss and GT reset replay Varun Gupta
2026-09-30  9:40 ` [PATCH 1/3] drm/xe: Park on the ULLS semaphore before publishing the next job's tail Varun Gupta
2026-09-30  9:40 ` [PATCH 2/3] drm/xe/guc: Publish the ring tail when replaying a chained ULLS job Varun Gupta
2026-09-30  9:40 ` [PATCH 3/3] drm/xe/guc: Rewind the LRC ring head when replaying a " Varun Gupta
2026-09-30  9:48 ` ✓ CI.KUnit: success for drm/xe: Fix ULLS chained job loss and GT reset replay Patchwork
2026-09-30 11:03 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-30 14:16 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-09-30 17:24 ` [PATCH 0/3] " Matthew Brost

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox