From: Varun Gupta <varun.gupta@intel.com>
To: intel-xe@lists.freedesktop.org
Cc: matthew.brost@intel.com, thomas.hellstrom@linux.intel.com,
himal.prasad.ghimiray@intel.com
Subject: [PATCH 1/3] drm/xe: Park on the ULLS semaphore before publishing the next job's tail
Date: Wed, 30 Sep 2026 15:10:33 +0530 [thread overview]
Message-ID: <20260930094031.3365707-6-varun.gupta@intel.com> (raw)
In-Reply-To: <20260930094031.3365707-5-varun.gupta@intel.com>
The ULLS postamble publishes the next job's ring tail and only then
parks on that job's semaphore. The CS fetches ring contents up to the
tail while parked, so it can fetch the next job's slot while it still
holds padding NOOPs. When the CPU later writes the job and signals the
semaphore, the CS executes the stale fetch, drains to the tail and
idles. The job never runs and its fence never signals.
This is hit whenever the CS parks before the next chained job is
written, e.g. when there is a pause between entering ULLS and the next
migration job. A lost ULLS_EXIT is silent as the next job's seqno write
covers its fence, but a lost ULLS_ACTIVE hangs the migration queue and
the subsequent kernel job timeout wedges the device.
Park first and publish the tail after the wait. The CS never fetches
beyond RING_TAIL, so the next slot cannot be fetched until the job is
in place. This puts the non-posted tail write on the wake-up path,
which the original order was chosen to avoid, but that order is not
safe.
Fixes: 6ec0b87160be ("drm/xe: Add ULLS migration job support to ring ops")
Cc: Matthew Brost <matthew.brost@intel.com>
Assisted-by: LLM
Signed-off-by: Varun Gupta <varun.gupta@intel.com>
---
drivers/gpu/drm/xe/xe_migrate.c | 32 +++++++++++++++++---------------
drivers/gpu/drm/xe/xe_ring_ops.c | 14 ++++++++++----
2 files changed, 27 insertions(+), 19 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_migrate.c b/drivers/gpu/drm/xe/xe_migrate.c
index 0dfc54ba3b8f..fffb3eb6c76e 100644
--- a/drivers/gpu/drm/xe/xe_migrate.c
+++ b/drivers/gpu/drm/xe/xe_migrate.c
@@ -116,19 +116,19 @@
* <copy timestamp, start seqno store>
* <batch buffer start(s)> (skipped on first/last job)
* <seqno write + user interrupt>
- * postamble: SDI saved ring tail = end of next job
+ * postamble: wait on semaphore[seqno + 1]
+ * SDI saved ring tail = end of next job
* LRI RING_TAIL = end of next job
- * wait on semaphore[seqno + 1]
* (skipped on the last job)
* pad: MI_NOOP up to ULLS_JOB_SIZE_DW
*
* The preamble clears the current job's semaphore so it can be reused once
* the seqno space wraps. The postamble is what keeps the engine busy: it
- * advances the ring tail over the next job and then blocks on that job's
- * semaphore, which is only signaled when the job is actually submitted. It
- * advances the saved tail as well as the tail register, keeping the two in
- * step without any help from the CPU, so a context save and restore can not
- * rewind the tail behind work which has already been published.
+ * blocks on the next job's semaphore, which is only signaled when that job is
+ * actually submitted, and then advances the ring tail over it. It advances
+ * the saved tail as well as the tail register, keeping the two in step
+ * without any help from the CPU, so a context save and restore can not rewind
+ * the tail behind work which has already been published.
*
* The tail register write must be non-posted, i.e. it must not carry
* MI_LRI_FORCE_POSTED. Posted, the new tail is free to land after the command
@@ -137,10 +137,12 @@
* A parked context can be switched off the hardware, and the fast path below
* has no H2G with which to ask GuC to bring it back.
*
- * The tail is published ahead of the semaphore wait rather than after it so
- * that the non-posted write drains while the engine is parked anyway, keeping
- * a register round trip off the path between the semaphore being signaled and
- * the next job running.
+ * The tail must be published after the semaphore wait, not before it. The
+ * command streamer fetches ring contents up to the tail while it is parked,
+ * so a tail published ahead of the wait lets it fetch the next job's slot
+ * before the CPU has written the job there. Once released it then executes
+ * the MI_NOOPs it fetched instead of the job, that job's fence never signals,
+ * and the engine drains and idles with nothing left to wake it.
*
* Submission fast path
* --------------------
@@ -151,10 +153,10 @@
* xe_lrc_set_ulls_semaphore(lrc, seqno); release previous job
*
* The XE_GUC_ACTION_SCHED_CONTEXT H2G is suppressed, and so is the write of
- * the saved ring tail: the previous job's postamble has already published
- * this job's tail both in the tail register and in the context image, so the
- * semaphore signal is all that is left. The previous job's semaphore wait is
- * satisfied and the engine walks straight into this job.
+ * the saved ring tail: the previous job's postamble is parked on this job's
+ * semaphore and publishes this job's tail, both in the tail register and in
+ * the context image, as soon as it is released. The semaphore signal is all
+ * that is left, and the engine walks straight into this job.
*
* This does assume the context stays resident for as long as ULLS mode is
* active. Nothing else is scheduled on the reserved engine, so the only ways
diff --git a/drivers/gpu/drm/xe/xe_ring_ops.c b/drivers/gpu/drm/xe/xe_ring_ops.c
index bc4dea606b38..ad63181938e3 100644
--- a/drivers/gpu/drm/xe/xe_ring_ops.c
+++ b/drivers/gpu/drm/xe/xe_ring_ops.c
@@ -537,12 +537,18 @@ static int emit_ulls_ring_tail(struct xe_gt *gt, struct xe_lrc *lrc, u32 *dw,
return i;
}
-/* Publish the next job's tail, then park the engine on its semaphore */
+/*
+ * Park the engine on the next job's semaphore, then publish its tail.
+ *
+ * The tail must not be published before the wait. The command streamer
+ * fetches ring contents up to the tail while parked, so it would fetch the
+ * next job's slot while it still holds MI_NOOPs and execute those once
+ * released, rather than the job the CPU writes there later. Publishing after
+ * the wait keeps the slot beyond the tail until the job is in place.
+ */
static int emit_ulls_postamble(struct xe_gt *gt, struct xe_lrc *lrc, u32 *dw,
int i, u32 seqno, u32 head)
{
- i = emit_ulls_ring_tail(gt, lrc, dw, i, head);
-
dw[i++] = MI_SEMAPHORE_WAIT |
MI_SEMW_GGTT |
MI_SEMW_POLL |
@@ -552,7 +558,7 @@ static int emit_ulls_postamble(struct xe_gt *gt, struct xe_lrc *lrc, u32 *dw,
dw[i++] = 0;
dw[i++] = 0;
- return i;
+ return emit_ulls_ring_tail(gt, lrc, dw, i, head);
}
/* Pad out to the fixed ULLS job size */
--
2.43.0
next prev parent reply other threads:[~2026-09-30 9:40 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 9:40 [PATCH 0/3] drm/xe: Fix ULLS chained job loss and GT reset replay Varun Gupta
2026-09-30 9:40 ` Varun Gupta [this message]
2026-09-30 9:40 ` [PATCH 2/3] drm/xe/guc: Publish the ring tail when replaying a chained ULLS job Varun Gupta
2026-09-30 9:40 ` [PATCH 3/3] drm/xe/guc: Rewind the LRC ring head when replaying a " Varun Gupta
2026-09-30 9:48 ` ✓ CI.KUnit: success for drm/xe: Fix ULLS chained job loss and GT reset replay Patchwork
2026-09-30 11:03 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-30 14:16 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-09-30 17:24 ` [PATCH 0/3] " Matthew Brost
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260930094031.3365707-6-varun.gupta@intel.com \
--to=varun.gupta@intel.com \
--cc=himal.prasad.ghimiray@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=matthew.brost@intel.com \
--cc=thomas.hellstrom@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).