* [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933)
@ 2026-09-22 14:46 Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
` (3 more replies)
0 siblings, 4 replies; 6+ messages in thread
From: Tales A. Mendonça @ 2026-09-22 14:46 UTC (permalink / raw)
To: intel-xe
Cc: matthew.brost, daniele.ceraolospurio, stuart.summers,
julia.filipchuk, thomas.hellstrom, rodrigo.vivi, jani.nikula,
navonjohnlukose, dri-devel, Tales A. Mendonça
Hi,
v6 of the TLB invalidation ack stall fix for ARL. Tracked in:
https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
The whole series now carries Matthew Brost's Reviewed-by - thank you for
working through it, including cross-checking patch 3 against the i915
implementation.
Recap: on the standalone media GT of MTL/ARL the CPU reads stale cache
lines for data the GuC has already written. The visible symptom is TLB
invalidation acks appearing to stall for a near-constant ~2.3s. i915
works around this as Wa_22016122933; xe never inherited it. Patch 3
implements it, applying XE_BO_FLAG_NEEDS_UC to the GuC-shared
allocations (CTBs, log, ADS, SLPC, engine activity) on the standalone
media GT, scoped by a new OOB rule (22016122933 MEDIA_VERSION(1300)).
The one code change since v5 is in patch 1, from a second issue Sashiko
raised: the capture could run without a runtime PM reference. Every
pending invalidation fence holds one, taken in
xe_tlb_inval_fence_init(), and xe_tlb_inval_fence_signal() drops it via
xe_tlb_inval_fence_fini(). The timeout loop signals the expired fences
and only then calls xe_devcoredump_gt(), whose forcewake acquisition has
always relied on the caller holding a PM reference. If those were the
last references the device could begin autosuspending before the
snapshot touched the hardware. v6 takes a reference while the pending
fences still guarantee the device is awake, and releases it after the
capture. Matt confirmed the analysis and the fix on the list.
I am carrying Matt's tag on patch 1 across that change since he reviewed
the fix itself, but flagging it here so it is not silently inherited.
Two open points from earlier revisions, both now settled:
- LRC coverage: i915 also marks the LRC UC on non-dGPU
(__lrc_alloc_state()). That does not match the erratum's direction -
LRC writes come from hardware context save and the GuC reads it -
and Matt's guidance was to leave i915 alone and treat it as out of
scope for xe.
- SIGID: the TLB logging in patch 2 will be converted once a TLB
component exists in DEFINE_XE_LOG_COMPONENTS(); a colleague of
Matt's volunteered to do that as a follow-up on top of this series.
Two notes carried over, still open to either answer:
1. CPU mapping: keeping XE_BO_FLAG_NEEDS_UC (uncached on both sides).
It is the tested configuration and no throughput difference against
the CPU-WC variant was measurable. Matching i915's exact CPU-WC +
GGTT-UC combination needs either a new BO flag or decoupling the
GGTT cache-mode selection from XE_BO_FLAG_NEEDS_UC; happy to add
that plumbing if parity is preferred.
2. Fixes:/Cc: stable are left out, since MTL/ARL is require_force_probe
in xe. Also happy to add them.
Validation of patch 3 is six weeks on two ARL machines (7d51 and 7dd1),
across kernels 7.1.6, 7.1.8 and 7.2, with over 10M TLB invalidations
processed and zero ack stalls. Before the fix both machines reproduced
20-60 stalls/day, every day, on two GuC firmware versions. The 7dd1
machine, which could not survive a day of media workloads on xe without
a platform freeze, has been running xe full time since 11 August with
zero incidents.
checkpatch is clean, except for one --strict CHECK about macro argument
reuse in the xe_devcoredump() wrapper in patch 1, which is intentional:
the macro only exists to forward (_q)->gt alongside _q.
v5 -> v6:
- Rebased on today's drm-tip; builds clean, no conflicts.
- Patch 1: hold a runtime PM reference across the devcoredump capture
(second issue reported by Sashiko, confirmed by Matt).
- Patches 1-3: collected Reviewed-by from Matthew Brost.
- Patches 2-3: otherwise unchanged.
Thanks,
Tales
Tales A. Mendonça (3):
drm/xe: Capture devcoredump on TLB invalidation timeout
drm/xe: Log when a timed out TLB invalidation ack finally arrives
drm/xe: Implement Wa_22016122933
drivers/gpu/drm/xe/xe_devcoredump.c | 46 ++++++++++--------
drivers/gpu/drm/xe/xe_devcoredump.h | 15 ++++--
drivers/gpu/drm/xe/xe_guc.c | 16 ++++++
drivers/gpu/drm/xe/xe_guc.h | 2 +
drivers/gpu/drm/xe/xe_guc_ads.c | 3 +-
drivers/gpu/drm/xe/xe_guc_ct.c | 6 ++-
drivers/gpu/drm/xe/xe_guc_engine_activity.c | 6 ++-
drivers/gpu/drm/xe/xe_guc_log.c | 7 ++-
drivers/gpu/drm/xe/xe_guc_pc.c | 3 +-
drivers/gpu/drm/xe/xe_tlb_inval.c | 54 +++++++++++++++++++++
drivers/gpu/drm/xe/xe_tlb_inval_types.h | 17 +++++++
drivers/gpu/drm/xe/xe_wa_oob.rules | 1 +
12 files changed, 143 insertions(+), 33 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH v6 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout
2026-09-22 14:46 [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Tales A. Mendonça
@ 2026-09-22 14:46 ` Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
` (2 subsequent siblings)
3 siblings, 0 replies; 6+ messages in thread
From: Tales A. Mendonça @ 2026-09-22 14:46 UTC (permalink / raw)
To: intel-xe
Cc: matthew.brost, daniele.ceraolospurio, stuart.summers,
julia.filipchuk, thomas.hellstrom, rodrigo.vivi, jani.nikula,
navonjohnlukose, dri-devel, Tales A. Mendonça
TLB invalidation timeouts currently leave no record of the firmware
state behind: there is no exec queue or job to blame, so nothing calls
xe_devcoredump() and the GuC log content at the time of the hang is
lost.
Add xe_devcoredump_gt(), a variant of xe_devcoredump() for hangs that
are not tied to an exec queue or job. It captures the GuC log and CT
state of the affected GT, reusing the existing snapshot machinery and
the "only first snapshot" policy, and hook it up to the TLB invalidation
timeout path.
This was instrumental in diagnosing GuC TLB invalidation ack stalls on
ARL (see Link), where the invalidation request is consumed from the H2G
CTB immediately but the ack G2H only arrives ~2.3s later, after the
timeout has already fired.
Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
Signed-off-by: Tales A. Mendonça <talesam@gmail.com>
Reviewed-by: Matthew Brost <matthew.brost@intel.com>
---
drivers/gpu/drm/xe/xe_devcoredump.c | 46 ++++++++++++++++-------------
drivers/gpu/drm/xe/xe_devcoredump.h | 15 +++++++---
drivers/gpu/drm/xe/xe_tlb_inval.c | 35 ++++++++++++++++++++++
3 files changed, 71 insertions(+), 25 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_devcoredump.c b/drivers/gpu/drm/xe/xe_devcoredump.c
index b918f046ae1..ab681b14886 100644
--- a/drivers/gpu/drm/xe/xe_devcoredump.c
+++ b/drivers/gpu/drm/xe/xe_devcoredump.c
@@ -74,11 +74,6 @@ static struct xe_device *coredump_to_xe(const struct xe_devcoredump *coredump)
return container_of(coredump, struct xe_device, devcoredump);
}
-static struct xe_guc *exec_queue_to_guc(struct xe_exec_queue *q)
-{
- return &q->gt->uc.guc;
-}
-
static ssize_t __xe_devcoredump_read(char *buffer, ssize_t count,
ssize_t start,
struct xe_devcoredump *coredump)
@@ -325,40 +320,44 @@ static void xe_devcoredump_deferred_snap_work(struct work_struct *work)
}
static void devcoredump_snapshot(struct xe_devcoredump *coredump,
+ struct xe_gt *gt,
struct xe_exec_queue *q,
struct xe_sched_job *job)
{
struct xe_devcoredump_snapshot *ss = &coredump->snapshot;
- struct xe_guc *guc = exec_queue_to_guc(q);
+ struct xe_guc *guc = >->uc.guc;
const char *process_name = "no process";
bool cookie;
ss->snapshot_time = ktime_get_real();
ss->boot_time = ktime_get_boottime();
- if (q->vm && q->vm->xef) {
+ if (q && q->vm && q->vm->xef) {
process_name = q->vm->xef->process_name;
ss->pid = q->vm->xef->pid;
}
strscpy(ss->process_name, process_name);
- ss->gt = q->gt;
+ ss->gt = gt;
INIT_WORK(&ss->work, xe_devcoredump_deferred_snap_work);
/* keep going if fw fails as we still want to save the memory and SW data */
- CLASS(xe_force_wake, fw_ref)(gt_to_fw(q->gt), XE_FORCEWAKE_ALL);
+ CLASS(xe_force_wake, fw_ref)(gt_to_fw(gt), XE_FORCEWAKE_ALL);
cookie = dma_fence_begin_signalling();
ss->guc.log = xe_guc_log_snapshot_capture(&guc->log, true);
ss->guc.ct = xe_guc_ct_snapshot_capture(&guc->ct);
- ss->ge = xe_guc_exec_queue_snapshot_capture(q);
- if (job)
- ss->job = xe_sched_job_snapshot_capture(job);
- ss->vm = xe_vm_snapshot_capture(q->vm);
- xe_engine_snapshot_capture_for_queue(q);
+ if (q) {
+ ss->ge = xe_guc_exec_queue_snapshot_capture(q);
+ if (job)
+ ss->job = xe_sched_job_snapshot_capture(job);
+ ss->vm = xe_vm_snapshot_capture(q->vm);
+
+ xe_engine_snapshot_capture_for_queue(q);
+ }
queue_work(system_dfl_wq, &ss->work);
@@ -366,19 +365,24 @@ static void devcoredump_snapshot(struct xe_devcoredump *coredump,
}
/**
- * xe_devcoredump - Take the required snapshots and initialize coredump device.
- * @q: The faulty xe_exec_queue, where the issue was detected.
- * @job: The faulty xe_sched_job, where the issue was detected.
+ * __xe_devcoredump - Take the required snapshots and initialize coredump device.
+ * @gt: The GT where the issue was detected.
+ * @q: The faulty xe_exec_queue, where the issue was detected, may be NULL for
+ * hangs that are not tied to an exec queue (e.g. TLB invalidation
+ * timeouts); in that case only the GT-level state (GuC log and CT state)
+ * is captured.
+ * @job: The faulty xe_sched_job, where the issue was detected, may be NULL.
* @fmt: Printf format + args to describe the reason for the core dump
*
* This function should be called at the crash time within the serialized
* gt_reset. It is skipped if we still have the core dump device available
* with the information of the 'first' snapshot.
*/
-__printf(3, 4)
-void xe_devcoredump(struct xe_exec_queue *q, struct xe_sched_job *job, const char *fmt, ...)
+__printf(4, 5)
+void __xe_devcoredump(struct xe_gt *gt, struct xe_exec_queue *q,
+ struct xe_sched_job *job, const char *fmt, ...)
{
- struct xe_device *xe = gt_to_xe(q->gt);
+ struct xe_device *xe = gt_to_xe(gt);
struct xe_devcoredump *coredump = &xe->devcoredump;
va_list varg;
@@ -396,7 +400,7 @@ void xe_devcoredump(struct xe_exec_queue *q, struct xe_sched_job *job, const cha
coredump->snapshot.reason = kvasprintf(GFP_ATOMIC, fmt, varg);
va_end(varg);
- devcoredump_snapshot(coredump, q, job);
+ devcoredump_snapshot(coredump, gt, q, job);
drm_info(&xe->drm, "Xe device coredump has been created\n");
drm_info(&xe->drm, "Check your /sys/class/drm/card%d/device/devcoredump/data\n",
diff --git a/drivers/gpu/drm/xe/xe_devcoredump.h b/drivers/gpu/drm/xe/xe_devcoredump.h
index 5391a80a4d1..bc4800b3a3f 100644
--- a/drivers/gpu/drm/xe/xe_devcoredump.h
+++ b/drivers/gpu/drm/xe/xe_devcoredump.h
@@ -11,15 +11,17 @@
struct drm_printer;
struct xe_device;
struct xe_exec_queue;
+struct xe_gt;
struct xe_sched_job;
#ifdef CONFIG_DEV_COREDUMP
-void xe_devcoredump(struct xe_exec_queue *q, struct xe_sched_job *job, const char *fmt, ...);
+void __xe_devcoredump(struct xe_gt *gt, struct xe_exec_queue *q,
+ struct xe_sched_job *job, const char *fmt, ...);
int xe_devcoredump_init(struct xe_device *xe);
#else
-static inline void xe_devcoredump(struct xe_exec_queue *q,
- struct xe_sched_job *job,
- const char *fmt, ...)
+static inline void __xe_devcoredump(struct xe_gt *gt, struct xe_exec_queue *q,
+ struct xe_sched_job *job,
+ const char *fmt, ...)
{
}
@@ -29,6 +31,11 @@ static inline int xe_devcoredump_init(struct xe_device *xe)
}
#endif
+#define xe_devcoredump(_q, _job, _fmt, ...) \
+ __xe_devcoredump((_q)->gt, _q, _job, _fmt, ##__VA_ARGS__)
+#define xe_devcoredump_gt(_gt, _fmt, ...) \
+ __xe_devcoredump(_gt, NULL, NULL, _fmt, ##__VA_ARGS__)
+
void xe_print_blob_ascii85(struct drm_printer *p, const char *prefix, char suffix,
const void *blob, size_t offset, size_t size);
diff --git a/drivers/gpu/drm/xe/xe_tlb_inval.c b/drivers/gpu/drm/xe/xe_tlb_inval.c
index dd1030de9ab..f7d6640d345 100644
--- a/drivers/gpu/drm/xe/xe_tlb_inval.c
+++ b/drivers/gpu/drm/xe/xe_tlb_inval.c
@@ -5,6 +5,7 @@
#include <drm/drm_managed.h>
+#include "xe_devcoredump.h"
#include "xe_device_types.h"
#include "xe_force_wake.h"
#include "xe_gt_stats.h"
@@ -29,6 +30,12 @@
#define FENCE_STACK_BIT DMA_FENCE_FLAG_USER_BITS
+/* The frontend is only ever embedded in a GT */
+static struct xe_gt *tlb_inval_to_gt(struct xe_tlb_inval *tlb_inval)
+{
+ return container_of(tlb_inval, struct xe_gt, tlb_inval);
+}
+
static void xe_tlb_inval_fence_fini(struct xe_tlb_inval_fence *fence)
{
if (WARN_ON_ONCE(!fence->tlb_inval))
@@ -73,6 +80,7 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
struct xe_device *xe = tlb_inval->xe;
struct xe_tlb_inval_fence *fence, *next;
long timeout_delay = tlb_inval->ops->timeout_delay(tlb_inval);
+ int timedout_seqno = 0, seqno_recv = 0;
tlb_inval->ops->flush(tlb_inval);
@@ -90,13 +98,40 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
"TLB invalidation fence timeout, seqno=%d recv=%d",
fence->seqno, tlb_inval->seqno_recv);
+ if (!timedout_seqno) {
+ /*
+ * Hold a PM reference across the capture below. Every
+ * pending fence holds one, so the device is awake
+ * here, but signalling them may drop the last
+ * reference and let it autosuspend before the
+ * snapshot touches the hardware.
+ */
+ xe_pm_runtime_get_noresume(xe);
+ }
+
+ timedout_seqno = fence->seqno;
+
fence->base.error = -ETIME;
xe_tlb_inval_fence_signal(fence);
}
if (!list_empty(&tlb_inval->pending_fences))
queue_delayed_work(tlb_inval->timeout_wq, &tlb_inval->fence_tdr,
timeout_delay);
+ seqno_recv = tlb_inval->seqno_recv;
spin_unlock_irq(&tlb_inval->pending_lock);
+
+ /*
+ * Capture the GuC log and CT state so the firmware side of the hang
+ * can be inspected; there is no queue or job to blame here. Must be
+ * outside pending_lock as the capture takes sleeping locks, hence
+ * @seqno_recv is sampled above while the lock is still held.
+ */
+ if (timedout_seqno) {
+ xe_devcoredump_gt(tlb_inval_to_gt(tlb_inval),
+ "TLB invalidation fence timeout, seqno=%d recv=%d",
+ timedout_seqno, seqno_recv);
+ xe_pm_runtime_put(xe);
+ }
}
/**
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread
* [PATCH v6 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives
2026-09-22 14:46 [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
@ 2026-09-22 14:46 ` Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 3/3] drm/xe: Implement Wa_22016122933 Tales A. Mendonça
2026-09-24 22:05 ` [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Matthew Brost
3 siblings, 0 replies; 6+ messages in thread
From: Tales A. Mendonça @ 2026-09-22 14:46 UTC (permalink / raw)
To: intel-xe
Cc: matthew.brost, daniele.ceraolospurio, stuart.summers,
julia.filipchuk, thomas.hellstrom, rodrigo.vivi, jani.nikula,
navonjohnlukose, dri-devel, Tales A. Mendonça
When a TLB invalidation fence times out we log the timeout, but if the
ack for that seqno later shows up there is no record of it, making it
impossible to tell from logs whether the ack was lost forever or merely
(very) late.
Track the most recent timed out seqno and log how late its ack arrives,
relative to both the original request and the moment the fence was
signaled with -ETIME.
On ARL with GuC 70.53.0 this shows the acks are never lost: they
consistently arrive ~2.3s after the request, tens of milliseconds after
the TDR has already signaled the fence:
TLB invalidation fence timeout, seqno=10992 recv=10991
TLB invalidation late ack: seqno=10992 recv=10992,
request-to-ack=2314ms, timeout-to-ack=45ms
Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
Signed-off-by: Tales A. Mendonça <talesam@gmail.com>
Reviewed-by: Matthew Brost <matthew.brost@intel.com>
---
drivers/gpu/drm/xe/xe_tlb_inval.c | 19 +++++++++++++++++++
drivers/gpu/drm/xe/xe_tlb_inval_types.h | 17 +++++++++++++++++
2 files changed, 36 insertions(+)
diff --git a/drivers/gpu/drm/xe/xe_tlb_inval.c b/drivers/gpu/drm/xe/xe_tlb_inval.c
index f7d6640d345..c9cc5842d65 100644
--- a/drivers/gpu/drm/xe/xe_tlb_inval.c
+++ b/drivers/gpu/drm/xe/xe_tlb_inval.c
@@ -14,6 +14,7 @@
#include "xe_guc_tlb_inval.h"
#include "xe_mmio.h"
#include "xe_pm.h"
+#include "xe_printk.h"
#include "xe_tlb_inval.h"
#include "xe_trace.h"
@@ -110,6 +111,11 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
}
timedout_seqno = fence->seqno;
+ if (!tlb_inval->timedout_seqno) {
+ tlb_inval->timedout_seqno = fence->seqno;
+ tlb_inval->timedout_inval_time = fence->inval_time;
+ tlb_inval->timedout_time = ktime_get();
+ }
fence->base.error = -ETIME;
xe_tlb_inval_fence_signal(fence);
@@ -242,6 +248,7 @@ void xe_tlb_inval_reset(struct xe_tlb_inval *tlb_inval)
else
pending_seqno = tlb_inval->seqno - 1;
WRITE_ONCE(tlb_inval->seqno_recv, pending_seqno);
+ tlb_inval->timedout_seqno = 0;
list_for_each_entry_safe(fence, next,
&tlb_inval->pending_fences, link)
@@ -470,6 +477,18 @@ void xe_tlb_inval_done_handler(struct xe_tlb_inval *tlb_inval, int seqno)
WRITE_ONCE(tlb_inval->seqno_recv, seqno);
+ if (tlb_inval->timedout_seqno &&
+ xe_tlb_inval_seqno_past(tlb_inval, tlb_inval->timedout_seqno)) {
+ ktime_t now = ktime_get();
+
+ xe_warn(xe,
+ "TLB invalidation late ack: seqno=%d recv=%d, request-to-ack=%lldms, timeout-to-ack=%lldms",
+ tlb_inval->timedout_seqno, seqno,
+ ktime_ms_delta(now, tlb_inval->timedout_inval_time),
+ ktime_ms_delta(now, tlb_inval->timedout_time));
+ tlb_inval->timedout_seqno = 0;
+ }
+
list_for_each_entry_safe(fence, next,
&tlb_inval->pending_fences, link) {
trace_xe_tlb_inval_fence_recv(xe, fence);
diff --git a/drivers/gpu/drm/xe/xe_tlb_inval_types.h b/drivers/gpu/drm/xe/xe_tlb_inval_types.h
index d77be1aedc9..80d2019fa20 100644
--- a/drivers/gpu/drm/xe/xe_tlb_inval_types.h
+++ b/drivers/gpu/drm/xe/xe_tlb_inval_types.h
@@ -112,6 +112,23 @@ struct xe_tlb_inval {
* @pending_lock: protects @pending_fences and updating @seqno_recv.
*/
spinlock_t pending_lock;
+ /**
+ * @timedout_seqno: seqno of the most recent timed out TLB
+ * invalidation, 0 if none. Used to measure how late the ack for a
+ * timed out invalidation actually arrives. Protected by
+ * @pending_lock.
+ */
+ int timedout_seqno;
+ /**
+ * @timedout_inval_time: request time of @timedout_seqno. Protected by
+ * @pending_lock.
+ */
+ ktime_t timedout_inval_time;
+ /**
+ * @timedout_time: time @timedout_seqno was signaled with -ETIME.
+ * Protected by @pending_lock.
+ */
+ ktime_t timedout_time;
/**
* @fence_tdr: schedules a delayed call to xe_tlb_fence_timeout after
* the timeout interval is over.
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread
* [PATCH v6 3/3] drm/xe: Implement Wa_22016122933
2026-09-22 14:46 [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
@ 2026-09-22 14:46 ` Tales A. Mendonça
2026-09-24 22:05 ` [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Matthew Brost
3 siblings, 0 replies; 6+ messages in thread
From: Tales A. Mendonça @ 2026-09-22 14:46 UTC (permalink / raw)
To: intel-xe
Cc: matthew.brost, daniele.ceraolospurio, stuart.summers,
julia.filipchuk, thomas.hellstrom, rodrigo.vivi, jani.nikula,
navonjohnlukose, dri-devel, Tales A. Mendonça
On platforms with a standalone media GT and media version 13.00
(MTL/ARL), memory shared between the CPU and the media GT's GuC must
not be mapped cached on the CPU side: the CPU can otherwise read stale
cache lines for data the GuC has already written.
i915 implements this as Wa_22016122933 (see
intel_gt_needs_wa_22016122933(), used by intel_guc_allocate_vma() and
intel_gt_coherent_map_type()); xe never inherited it.
The visible symptom on ARL is TLB invalidation acks stalling for a
near-constant ~2.3s: the GuC writes the G2H ack in time, but the CPU
keeps reading a stale (empty) view of the G2H CTB until the line is
naturally evicted, so the fence timeout at 2.25s fires first. GuC log
decode confirmed all invalidations were handled promptly by the
firmware, and only the media GT was affected. See Link for the full
investigation (three machines affected: 7d51, 7dd1, Arc Pro 130T).
Making the mapping coherent instead of uncached does not help: with a
GGTT PAT entry repurposed to WB|COH_2WAY the driver comes up and the
media GT GuC runs, but the stalls remain (request-to-ack 2290ms).
Uncached really is required here. Note that 2-way coherency is not
normally reachable from a GGTT PTE (only 2 PAT bits), so that
experiment needed a modified PAT table and may not reflect a supported
configuration.
Add the OOB workaround scoped like i915 (media version 13.00, media GT
only - MEDIA_VERSION() OOB rules only match the media GT on standalone
media platforms) and apply XE_BO_FLAG_NEEDS_UC to the GuC-shared
allocations the CPU reads from: the CTBs, the GuC log, ADS, the SLPC
shared data and the engine activity buffers. hwconfig and the G2G
buffer are allocated on the primary GT only, where the workaround does
not apply.
Note that XE_BO_FLAG_NEEDS_UC drives both the CPU mapping (uncached)
and the GGTT cache mode (XE_CACHE_NONE instead of XE_CACHE_WB), which is
stricter than i915: i915 documents the workaround as WC on the CPU side
and UC on the GPU side. A CPU-WC variant with the GGTT side kept at
XE_CACHE_NONE was measured to be equally effective, but expressing it
would require either a new BO flag or decoupling the GGTT cache-mode
selection from XE_BO_FLAG_NEEDS_UC - simply swapping in
XE_BO_FLAG_FORCE_WC silently relaxes the GPU side back to WB and the
stalls return. The stricter mapping is kept here since it is the tested
configuration and no throughput difference between the two was
measurable; the extra plumbing can be added later if parity with i915 is
preferred.
Scope is limited to the GuC-shared allocations, which is where the
failures were observed. i915 additionally covers media-GT LRC/ring
state; extending xe to match can be done as a follow-up if wanted.
Validation on two ARL machines (7d51 and 7dd1): before, 20-60 TLB
invalidation ack stalls per day, every day, for weeks, on every kernel
and on two GuC firmware versions (70.53.0 and 70.72.1). After: zero
stalls in six weeks of combined runtime, over 10M TLB invalidations
processed under the same workloads, across kernels 7.1.6, 7.1.8 and
7.2. The 7dd1 machine, which could not survive a day of media
workloads on xe without a platform freeze, has been running xe full
time since 11 August, including days with heavy video transcoding,
with zero stalls and zero freezes.
The coherency experiment, the CPU-WC measurements and the engine
activity coverage gap were found by Navon John Lukose while A/B
testing v2 on an ARL 7d51.
Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
Suggested-by: Daniele Ceraolo Spurio <daniele.ceraolospurio@intel.com>
Signed-off-by: Tales A. Mendonça <talesam@gmail.com>
Tested-by: Navon John Lukose <navonjohnlukose@gmail.com>
Reviewed-by: Matthew Brost <matthew.brost@intel.com>
---
drivers/gpu/drm/xe/xe_guc.c | 16 ++++++++++++++++
drivers/gpu/drm/xe/xe_guc.h | 2 ++
drivers/gpu/drm/xe/xe_guc_ads.c | 3 ++-
drivers/gpu/drm/xe/xe_guc_ct.c | 6 ++++--
drivers/gpu/drm/xe/xe_guc_engine_activity.c | 6 ++++--
drivers/gpu/drm/xe/xe_guc_log.c | 7 +++++--
drivers/gpu/drm/xe/xe_guc_pc.c | 3 ++-
drivers/gpu/drm/xe/xe_wa_oob.rules | 1 +
8 files changed, 36 insertions(+), 8 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_guc.c b/drivers/gpu/drm/xe/xe_guc.c
index c7f8bbd4cb9..3ab4cb9e496 100644
--- a/drivers/gpu/drm/xe/xe_guc.c
+++ b/drivers/gpu/drm/xe/xe_guc.c
@@ -1469,6 +1469,22 @@ int xe_guc_suspend(struct xe_guc *guc)
return 0;
}
+/**
+ * xe_guc_bo_wa_flags - Extra BO flags for memory shared with the GuC
+ * @gt: the &xe_gt whose GuC the buffer will be shared with
+ *
+ * Wa_22016122933: on the standalone media GT, memory shared between the
+ * CPU and the GuC must not be mapped cached on the CPU side, otherwise
+ * the CPU can read stale data written by the GuC (e.g. G2H CTB writes)
+ * for multiple seconds.
+ *
+ * Return: additional XE_BO_FLAG_* to use when allocating GuC-shared memory
+ */
+u32 xe_guc_bo_wa_flags(struct xe_gt *gt)
+{
+ return XE_GT_WA(gt, 22016122933) ? XE_BO_FLAG_NEEDS_UC : 0;
+}
+
void xe_guc_notify(struct xe_guc *guc)
{
struct xe_gt *gt = guc_to_gt(guc);
diff --git a/drivers/gpu/drm/xe/xe_guc.h b/drivers/gpu/drm/xe/xe_guc.h
index 61e3ee19a59..c4eca40d69c 100644
--- a/drivers/gpu/drm/xe/xe_guc.h
+++ b/drivers/gpu/drm/xe/xe_guc.h
@@ -30,6 +30,7 @@
xe_guc_fw_version_at_least((guc), MAKE_GUC_VER_ARGS(ver))
struct drm_printer;
+struct xe_gt;
void xe_guc_comm_init_early(struct xe_guc *guc);
int xe_guc_init_noalloc(struct xe_guc *guc);
@@ -45,6 +46,7 @@ void xe_guc_runtime_suspend(struct xe_guc *guc);
void xe_guc_runtime_resume(struct xe_guc *guc);
int xe_guc_suspend(struct xe_guc *guc);
int xe_guc_softreset(struct xe_guc *guc);
+u32 xe_guc_bo_wa_flags(struct xe_gt *gt);
void xe_guc_notify(struct xe_guc *guc);
int xe_guc_auth_huc(struct xe_guc *guc, u32 rsa_addr);
int xe_guc_mmio_send(struct xe_guc *guc, const u32 *request, u32 len);
diff --git a/drivers/gpu/drm/xe/xe_guc_ads.c b/drivers/gpu/drm/xe/xe_guc_ads.c
index ff8eee3831a..abc7266fc6f 100644
--- a/drivers/gpu/drm/xe/xe_guc_ads.c
+++ b/drivers/gpu/drm/xe/xe_guc_ads.c
@@ -435,7 +435,8 @@ int xe_guc_ads_init(struct xe_guc_ads *ads)
XE_BO_FLAG_SYSTEM |
XE_BO_FLAG_GGTT |
XE_BO_FLAG_GGTT_INVALIDATE |
- XE_BO_FLAG_PINNED_NORESTORE);
+ XE_BO_FLAG_PINNED_NORESTORE |
+ xe_guc_bo_wa_flags(gt));
if (IS_ERR(bo))
return PTR_ERR(bo);
diff --git a/drivers/gpu/drm/xe/xe_guc_ct.c b/drivers/gpu/drm/xe/xe_guc_ct.c
index 5c4733da385..5c393aa29de 100644
--- a/drivers/gpu/drm/xe/xe_guc_ct.c
+++ b/drivers/gpu/drm/xe/xe_guc_ct.c
@@ -376,7 +376,8 @@ int xe_guc_ct_init(struct xe_guc_ct *ct)
XE_BO_FLAG_SYSTEM |
XE_BO_FLAG_GGTT |
XE_BO_FLAG_GGTT_INVALIDATE |
- XE_BO_FLAG_PINNED_NORESTORE);
+ XE_BO_FLAG_PINNED_NORESTORE |
+ xe_guc_bo_wa_flags(gt));
if (IS_ERR(bo))
return PTR_ERR(bo);
@@ -386,7 +387,8 @@ int xe_guc_ct_init(struct xe_guc_ct *ct)
XE_BO_FLAG_SYSTEM |
XE_BO_FLAG_GGTT |
XE_BO_FLAG_GGTT_INVALIDATE |
- XE_BO_FLAG_PINNED_NORESTORE);
+ XE_BO_FLAG_PINNED_NORESTORE |
+ xe_guc_bo_wa_flags(gt));
if (IS_ERR(bo))
return PTR_ERR(bo);
diff --git a/drivers/gpu/drm/xe/xe_guc_engine_activity.c b/drivers/gpu/drm/xe/xe_guc_engine_activity.c
index a782be57caa..729ce8ac140 100644
--- a/drivers/gpu/drm/xe/xe_guc_engine_activity.c
+++ b/drivers/gpu/drm/xe/xe_guc_engine_activity.c
@@ -97,7 +97,8 @@ static int allocate_engine_activity_buffers(struct xe_guc *guc,
metadata_bo = xe_bo_create_pin_map_novm(gt_to_xe(gt), tile, PAGE_ALIGN(metadata_size),
ttm_bo_type_kernel, XE_BO_FLAG_SYSTEM |
- XE_BO_FLAG_GGTT | XE_BO_FLAG_GGTT_INVALIDATE,
+ XE_BO_FLAG_GGTT | XE_BO_FLAG_GGTT_INVALIDATE |
+ xe_guc_bo_wa_flags(gt),
false);
if (IS_ERR(metadata_bo))
@@ -105,7 +106,8 @@ static int allocate_engine_activity_buffers(struct xe_guc *guc,
bo = xe_bo_create_pin_map_novm(gt_to_xe(gt), tile, PAGE_ALIGN(size),
ttm_bo_type_kernel, XE_BO_FLAG_VRAM_IF_DGFX(tile) |
- XE_BO_FLAG_GGTT | XE_BO_FLAG_GGTT_INVALIDATE, false);
+ XE_BO_FLAG_GGTT | XE_BO_FLAG_GGTT_INVALIDATE |
+ xe_guc_bo_wa_flags(gt), false);
if (IS_ERR(bo)) {
xe_bo_unpin_map_no_vm(metadata_bo);
diff --git a/drivers/gpu/drm/xe/xe_guc_log.c b/drivers/gpu/drm/xe/xe_guc_log.c
index 538d4df0f7a..7d006268ce9 100644
--- a/drivers/gpu/drm/xe/xe_guc_log.c
+++ b/drivers/gpu/drm/xe/xe_guc_log.c
@@ -17,6 +17,7 @@
#include "xe_force_wake.h"
#include "xe_gt_printk.h"
#include "xe_gt_types.h"
+#include "xe_guc.h"
#include "xe_map.h"
#include "xe_mmio.h"
#include "xe_module.h"
@@ -624,14 +625,16 @@ void xe_guc_log_print_lfd(struct xe_guc_log *log, struct drm_printer *p)
int xe_guc_log_init(struct xe_guc_log *log)
{
struct xe_device *xe = log_to_xe(log);
- struct xe_tile *tile = gt_to_tile(log_to_gt(log));
+ struct xe_gt *gt = log_to_gt(log);
+ struct xe_tile *tile = gt_to_tile(gt);
struct xe_bo *bo;
bo = xe_managed_bo_create_pin_map(xe, tile, GUC_LOG_SIZE,
XE_BO_FLAG_SYSTEM |
XE_BO_FLAG_GGTT |
XE_BO_FLAG_GGTT_INVALIDATE |
- XE_BO_FLAG_PINNED_NORESTORE);
+ XE_BO_FLAG_PINNED_NORESTORE |
+ xe_guc_bo_wa_flags(gt));
if (IS_ERR(bo))
return PTR_ERR(bo);
diff --git a/drivers/gpu/drm/xe/xe_guc_pc.c b/drivers/gpu/drm/xe/xe_guc_pc.c
index 097b075bd89..e0105222a2c 100644
--- a/drivers/gpu/drm/xe/xe_guc_pc.c
+++ b/drivers/gpu/drm/xe/xe_guc_pc.c
@@ -1391,7 +1391,8 @@ int xe_guc_pc_init(struct xe_guc_pc *pc)
XE_BO_FLAG_VRAM_IF_DGFX(tile) |
XE_BO_FLAG_GGTT |
XE_BO_FLAG_GGTT_INVALIDATE |
- XE_BO_FLAG_PINNED_NORESTORE);
+ XE_BO_FLAG_PINNED_NORESTORE |
+ xe_guc_bo_wa_flags(gt));
if (IS_ERR(bo))
return PTR_ERR(bo);
diff --git a/drivers/gpu/drm/xe/xe_wa_oob.rules b/drivers/gpu/drm/xe/xe_wa_oob.rules
index dd69ad07f7a..30958b26a1d 100644
--- a/drivers/gpu/drm/xe/xe_wa_oob.rules
+++ b/drivers/gpu/drm/xe/xe_wa_oob.rules
@@ -14,6 +14,7 @@
16017236439 PLATFORM(PVC)
14019821291 MEDIA_VERSION_RANGE(1300, 2000)
14015076503 MEDIA_VERSION(1300)
+22016122933 MEDIA_VERSION(1300)
14018913170 GRAPHICS_VERSION_RANGE(1270, 1274)
MEDIA_VERSION(1300)
PLATFORM(DG2)
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread
* Re: [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933)
2026-09-22 14:46 [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Tales A. Mendonça
` (2 preceding siblings ...)
2026-09-22 14:46 ` [PATCH v6 3/3] drm/xe: Implement Wa_22016122933 Tales A. Mendonça
@ 2026-09-24 22:05 ` Matthew Brost
2026-09-25 1:06 ` Tales A. Mendonça
3 siblings, 1 reply; 6+ messages in thread
From: Matthew Brost @ 2026-09-24 22:05 UTC (permalink / raw)
To: Tales A. Mendonça
Cc: intel-xe, daniele.ceraolospurio, stuart.summers, julia.filipchuk,
thomas.hellstrom, rodrigo.vivi, jani.nikula, navonjohnlukose,
dri-devel
On Tue, Sep 22, 2026 at 11:46:31AM -0300, Tales A. Mendonça wrote:
Thanks for the patches, I've merge to drm-xe-next.
Matt
> Hi,
>
> v6 of the TLB invalidation ack stall fix for ARL. Tracked in:
>
> https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
>
> The whole series now carries Matthew Brost's Reviewed-by - thank you for
> working through it, including cross-checking patch 3 against the i915
> implementation.
>
> Recap: on the standalone media GT of MTL/ARL the CPU reads stale cache
> lines for data the GuC has already written. The visible symptom is TLB
> invalidation acks appearing to stall for a near-constant ~2.3s. i915
> works around this as Wa_22016122933; xe never inherited it. Patch 3
> implements it, applying XE_BO_FLAG_NEEDS_UC to the GuC-shared
> allocations (CTBs, log, ADS, SLPC, engine activity) on the standalone
> media GT, scoped by a new OOB rule (22016122933 MEDIA_VERSION(1300)).
>
> The one code change since v5 is in patch 1, from a second issue Sashiko
> raised: the capture could run without a runtime PM reference. Every
> pending invalidation fence holds one, taken in
> xe_tlb_inval_fence_init(), and xe_tlb_inval_fence_signal() drops it via
> xe_tlb_inval_fence_fini(). The timeout loop signals the expired fences
> and only then calls xe_devcoredump_gt(), whose forcewake acquisition has
> always relied on the caller holding a PM reference. If those were the
> last references the device could begin autosuspending before the
> snapshot touched the hardware. v6 takes a reference while the pending
> fences still guarantee the device is awake, and releases it after the
> capture. Matt confirmed the analysis and the fix on the list.
>
> I am carrying Matt's tag on patch 1 across that change since he reviewed
> the fix itself, but flagging it here so it is not silently inherited.
>
> Two open points from earlier revisions, both now settled:
>
> - LRC coverage: i915 also marks the LRC UC on non-dGPU
> (__lrc_alloc_state()). That does not match the erratum's direction -
> LRC writes come from hardware context save and the GuC reads it -
> and Matt's guidance was to leave i915 alone and treat it as out of
> scope for xe.
>
> - SIGID: the TLB logging in patch 2 will be converted once a TLB
> component exists in DEFINE_XE_LOG_COMPONENTS(); a colleague of
> Matt's volunteered to do that as a follow-up on top of this series.
>
> Two notes carried over, still open to either answer:
>
> 1. CPU mapping: keeping XE_BO_FLAG_NEEDS_UC (uncached on both sides).
> It is the tested configuration and no throughput difference against
> the CPU-WC variant was measurable. Matching i915's exact CPU-WC +
> GGTT-UC combination needs either a new BO flag or decoupling the
> GGTT cache-mode selection from XE_BO_FLAG_NEEDS_UC; happy to add
> that plumbing if parity is preferred.
>
> 2. Fixes:/Cc: stable are left out, since MTL/ARL is require_force_probe
> in xe. Also happy to add them.
>
> Validation of patch 3 is six weeks on two ARL machines (7d51 and 7dd1),
> across kernels 7.1.6, 7.1.8 and 7.2, with over 10M TLB invalidations
> processed and zero ack stalls. Before the fix both machines reproduced
> 20-60 stalls/day, every day, on two GuC firmware versions. The 7dd1
> machine, which could not survive a day of media workloads on xe without
> a platform freeze, has been running xe full time since 11 August with
> zero incidents.
>
> checkpatch is clean, except for one --strict CHECK about macro argument
> reuse in the xe_devcoredump() wrapper in patch 1, which is intentional:
> the macro only exists to forward (_q)->gt alongside _q.
>
> v5 -> v6:
> - Rebased on today's drm-tip; builds clean, no conflicts.
> - Patch 1: hold a runtime PM reference across the devcoredump capture
> (second issue reported by Sashiko, confirmed by Matt).
> - Patches 1-3: collected Reviewed-by from Matthew Brost.
> - Patches 2-3: otherwise unchanged.
>
> Thanks,
> Tales
>
> Tales A. Mendonça (3):
> drm/xe: Capture devcoredump on TLB invalidation timeout
> drm/xe: Log when a timed out TLB invalidation ack finally arrives
> drm/xe: Implement Wa_22016122933
>
> drivers/gpu/drm/xe/xe_devcoredump.c | 46 ++++++++++--------
> drivers/gpu/drm/xe/xe_devcoredump.h | 15 ++++--
> drivers/gpu/drm/xe/xe_guc.c | 16 ++++++
> drivers/gpu/drm/xe/xe_guc.h | 2 +
> drivers/gpu/drm/xe/xe_guc_ads.c | 3 +-
> drivers/gpu/drm/xe/xe_guc_ct.c | 6 ++-
> drivers/gpu/drm/xe/xe_guc_engine_activity.c | 6 ++-
> drivers/gpu/drm/xe/xe_guc_log.c | 7 ++-
> drivers/gpu/drm/xe/xe_guc_pc.c | 3 +-
> drivers/gpu/drm/xe/xe_tlb_inval.c | 54 +++++++++++++++++++++
> drivers/gpu/drm/xe/xe_tlb_inval_types.h | 17 +++++++
> drivers/gpu/drm/xe/xe_wa_oob.rules | 1 +
> 12 files changed, 143 insertions(+), 33 deletions(-)
>
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933)
2026-09-24 22:05 ` [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Matthew Brost
@ 2026-09-25 1:06 ` Tales A. Mendonça
0 siblings, 0 replies; 6+ messages in thread
From: Tales A. Mendonça @ 2026-09-25 1:06 UTC (permalink / raw)
To: Matthew Brost
Cc: intel-xe, daniele.ceraolospurio, stuart.summers, julia.filipchuk,
thomas.hellstrom, rodrigo.vivi, jani.nikula, navonjohnlukose,
dri-devel
Thanks Matt, and thanks everyone for the reviews and the testing.
Tales
--
Tales A. Mendonça
talesam.org | communitybig.org
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-09-25 7:32 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-22 14:46 [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 3/3] drm/xe: Implement Wa_22016122933 Tales A. Mendonça
2026-09-24 22:05 ` [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Matthew Brost
2026-09-25 1:06 ` Tales A. Mendonça
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox