* [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
@ 2026-08-04 2:14 Tales A. Mendonça
2026-08-04 2:14 ` [RFC PATCH 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
` (9 more replies)
0 siblings, 10 replies; 25+ messages in thread
From: Tales A. Mendonça @ 2026-08-04 2:14 UTC (permalink / raw)
To: intel-xe
Cc: matthew.brost, thomas.hellstrom, rodrigo.vivi, dri-devel,
Tales A. Mendonça
Hi,
This series is a follow-up to the TLB invalidation ack stall I have
been debugging on ARL, tracked in:
https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51 and
7dd1 machines here, plus an independent Arc Pro 130T report on the
issue above), TLB invalidation acks intermittently stall for ~2.3s.
The H2G request is consumed from the CTB immediately and the G2H CTB
is empty the whole time - the firmware simply does not send the ack
until much later. The fence timeout fires at 2.25s and the ack lands
tens of ms after it. Userspace blocked on the invalidation (compositor
buffer unmaps etc.) hitches for the full window.
Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has no
exec queue or job to blame - leaves a devcoredump with the GuC log and
CT state behind (Matt suggested capturing devcoredumps when we
discussed the issue; devcoredumps from both machines are attached to
the issue above).
Patch 2 logs when the ack for a timed out invalidation finally
arrives. This is what established that the acks are late rather than
lost.
Patch 3 is the RFC part: a delayed work that pokes the GuC (status
register read, CT flush, doorbell ring) every 250ms while an ack is
overdue. On my machines this converts the guaranteed 2.3s stall into a
sub-500ms hiccup for the majority of occurrences; a minority of severe
episodes ignore 8-9 consecutive doorbells, which points at the GuC
firmware being internally blocked for the whole window. Full data on
the issue. I am happy to rework the approach (different delay,
tying it to the G2H handler, dropping the status read, etc.) - mainly
I would like the firmware side investigated, since no host-side poke
can fix the severe cases.
Based on drm-tip. Tested for several days on both ARL machines under
desktop and VM-heavy workloads.
Thanks,
Tales
Tales A. Mendonça (3):
drm/xe: Capture devcoredump on TLB invalidation timeout
drm/xe: Log when a timed out TLB invalidation ack finally arrives
drm/xe: Kick GuC while TLB invalidation acks are overdue
drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
drivers/gpu/drm/xe/xe_tlb_inval.c | 131 +++++++++++++++++++++++-
drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
4 files changed, 243 insertions(+), 4 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 25+ messages in thread
* [RFC PATCH 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout
2026-08-04 2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
@ 2026-08-04 2:14 ` Tales A. Mendonça
2026-08-04 22:05 ` Matthew Brost
2026-08-04 2:14 ` [RFC PATCH 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
` (8 subsequent siblings)
9 siblings, 1 reply; 25+ messages in thread
From: Tales A. Mendonça @ 2026-08-04 2:14 UTC (permalink / raw)
To: intel-xe
Cc: matthew.brost, thomas.hellstrom, rodrigo.vivi, dri-devel,
Tales A. Mendonça
TLB invalidation timeouts currently leave no record of the firmware
state behind: there is no exec queue or job to blame, so nothing calls
xe_devcoredump() and the GuC log content at the time of the hang is
lost.
Add xe_devcoredump_gt(), a variant of xe_devcoredump() for hangs that
are not tied to an exec queue or job. It captures the GuC log and CT
state of the affected GT, reusing the existing snapshot machinery and
the "only first snapshot" policy, and hook it up to the TLB invalidation
timeout path.
This was instrumental in diagnosing GuC TLB invalidation ack stalls on
ARL (see Link), where the invalidation request is consumed from the H2G
CTB immediately but the ack G2H only arrives ~2.3s later, after the
timeout has already fired.
Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
Signed-off-by: Tales A. Mendonça <talesam@gmail.com>
---
drivers/gpu/drm/xe/xe_devcoredump.c | 68 +++++++++++++++++++++++++++++
drivers/gpu/drm/xe/xe_devcoredump.h | 6 +++
drivers/gpu/drm/xe/xe_tlb_inval.c | 20 +++++++++
3 files changed, 94 insertions(+)
diff --git a/drivers/gpu/drm/xe/xe_devcoredump.c b/drivers/gpu/drm/xe/xe_devcoredump.c
index 5f2b90b18f9..0ccaed176a4 100644
--- a/drivers/gpu/drm/xe/xe_devcoredump.c
+++ b/drivers/gpu/drm/xe/xe_devcoredump.c
@@ -403,6 +403,74 @@ void xe_devcoredump(struct xe_exec_queue *q, struct xe_sched_job *job, const cha
mutex_unlock(&coredump->lock);
}
+static void devcoredump_snapshot_gt(struct xe_devcoredump *coredump,
+ struct xe_gt *gt)
+{
+ struct xe_devcoredump_snapshot *ss = &coredump->snapshot;
+ struct xe_guc *guc = >->uc.guc;
+ bool cookie;
+
+ ss->snapshot_time = ktime_get_real();
+ ss->boot_time = ktime_get_boottime();
+
+ strscpy(ss->process_name, "no process");
+
+ ss->gt = gt;
+ INIT_WORK(&ss->work, xe_devcoredump_deferred_snap_work);
+
+ /* keep going if fw fails as we still want to save the SW data */
+ CLASS(xe_force_wake, fw_ref)(gt_to_fw(gt), XE_FORCEWAKE_ALL);
+
+ cookie = dma_fence_begin_signalling();
+
+ ss->guc.log = xe_guc_log_snapshot_capture(&guc->log, true);
+ ss->guc.ct = xe_guc_ct_snapshot_capture(&guc->ct);
+
+ queue_work(system_dfl_wq, &ss->work);
+
+ dma_fence_end_signalling(cookie);
+}
+
+/**
+ * xe_devcoredump_gt - Take GT-level snapshots and initialize coredump device.
+ * @gt: The GT where the issue was detected.
+ * @fmt: Printf format + args to describe the reason for the core dump
+ *
+ * Variant of xe_devcoredump() for hangs that are not tied to an exec queue
+ * or job, e.g. TLB invalidation timeouts. Captures the GuC log and CT state
+ * of @gt so the firmware side of the hang can be inspected. Skipped if a
+ * coredump is already captured, same as xe_devcoredump().
+ */
+__printf(2, 3)
+void xe_devcoredump_gt(struct xe_gt *gt, const char *fmt, ...)
+{
+ struct xe_device *xe = gt_to_xe(gt);
+ struct xe_devcoredump *coredump = &xe->devcoredump;
+ va_list varg;
+
+ mutex_lock(&coredump->lock);
+
+ if (coredump->captured) {
+ drm_dbg(&xe->drm, "Multiple hangs are occurring, but only the first snapshot was taken\n");
+ mutex_unlock(&coredump->lock);
+ return;
+ }
+
+ coredump->captured = true;
+
+ va_start(varg, fmt);
+ coredump->snapshot.reason = kvasprintf(GFP_ATOMIC, fmt, varg);
+ va_end(varg);
+
+ devcoredump_snapshot_gt(coredump, gt);
+
+ drm_info(&xe->drm, "Xe device coredump has been created\n");
+ drm_info(&xe->drm, "Check your /sys/class/drm/card%d/device/devcoredump/data\n",
+ xe->drm.primary->index);
+
+ mutex_unlock(&coredump->lock);
+}
+
static void xe_driver_devcoredump_fini(void *arg)
{
struct drm_device *drm = arg;
diff --git a/drivers/gpu/drm/xe/xe_devcoredump.h b/drivers/gpu/drm/xe/xe_devcoredump.h
index 5391a80a4d1..f071bd11f24 100644
--- a/drivers/gpu/drm/xe/xe_devcoredump.h
+++ b/drivers/gpu/drm/xe/xe_devcoredump.h
@@ -11,10 +11,12 @@
struct drm_printer;
struct xe_device;
struct xe_exec_queue;
+struct xe_gt;
struct xe_sched_job;
#ifdef CONFIG_DEV_COREDUMP
void xe_devcoredump(struct xe_exec_queue *q, struct xe_sched_job *job, const char *fmt, ...);
+void xe_devcoredump_gt(struct xe_gt *gt, const char *fmt, ...);
int xe_devcoredump_init(struct xe_device *xe);
#else
static inline void xe_devcoredump(struct xe_exec_queue *q,
@@ -23,6 +25,10 @@ static inline void xe_devcoredump(struct xe_exec_queue *q,
{
}
+static inline void xe_devcoredump_gt(struct xe_gt *gt, const char *fmt, ...)
+{
+}
+
static inline int xe_devcoredump_init(struct xe_device *xe)
{
return 0;
diff --git a/drivers/gpu/drm/xe/xe_tlb_inval.c b/drivers/gpu/drm/xe/xe_tlb_inval.c
index bbd21d39306..833fb92cd3e 100644
--- a/drivers/gpu/drm/xe/xe_tlb_inval.c
+++ b/drivers/gpu/drm/xe/xe_tlb_inval.c
@@ -5,6 +5,7 @@
#include <drm/drm_managed.h>
+#include "xe_devcoredump.h"
#include "xe_device_types.h"
#include "xe_force_wake.h"
#include "xe_gt_stats.h"
@@ -29,6 +30,12 @@
#define FENCE_STACK_BIT DMA_FENCE_FLAG_USER_BITS
+/* The frontend is only ever embedded in a GT */
+static struct xe_gt *tlb_inval_to_gt(struct xe_tlb_inval *tlb_inval)
+{
+ return container_of(tlb_inval, struct xe_gt, tlb_inval);
+}
+
static void xe_tlb_inval_fence_fini(struct xe_tlb_inval_fence *fence)
{
if (WARN_ON_ONCE(!fence->tlb_inval))
@@ -73,6 +80,7 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
struct xe_device *xe = tlb_inval->xe;
struct xe_tlb_inval_fence *fence, *next;
long timeout_delay = tlb_inval->ops->timeout_delay(tlb_inval);
+ int timedout_seqno = 0;
tlb_inval->ops->flush(tlb_inval);
@@ -90,6 +98,8 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
"TLB invalidation fence timeout, seqno=%d recv=%d",
fence->seqno, tlb_inval->seqno_recv);
+ timedout_seqno = fence->seqno;
+
fence->base.error = -ETIME;
xe_tlb_inval_fence_signal(fence);
}
@@ -97,6 +107,16 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
queue_delayed_work(tlb_inval->timeout_wq, &tlb_inval->fence_tdr,
timeout_delay);
spin_unlock_irq(&tlb_inval->pending_lock);
+
+ /*
+ * Capture the GuC log and CT state so the firmware side of the hang
+ * can be inspected; there is no queue or job to blame here. Must be
+ * outside pending_lock as the capture takes sleeping locks.
+ */
+ if (timedout_seqno)
+ xe_devcoredump_gt(tlb_inval_to_gt(tlb_inval),
+ "TLB invalidation fence timeout, seqno=%d recv=%d",
+ timedout_seqno, tlb_inval->seqno_recv);
}
/**
--
2.55.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives
2026-08-04 2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
2026-08-04 2:14 ` [RFC PATCH 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
@ 2026-08-04 2:14 ` Tales A. Mendonça
2026-08-04 22:17 ` Matthew Brost
2026-08-04 2:14 ` [RFC PATCH 3/3] drm/xe: Kick GuC while TLB invalidation acks are overdue Tales A. Mendonça
` (7 subsequent siblings)
9 siblings, 1 reply; 25+ messages in thread
From: Tales A. Mendonça @ 2026-08-04 2:14 UTC (permalink / raw)
To: intel-xe
Cc: matthew.brost, thomas.hellstrom, rodrigo.vivi, dri-devel,
Tales A. Mendonça
When a TLB invalidation fence times out we log the timeout, but if the
ack for that seqno later shows up there is no record of it, making it
impossible to tell from logs whether the ack was lost forever or merely
(very) late.
Track the most recent timed out seqno and log how late its ack arrives,
relative to both the original request and the moment the fence was
signaled with -ETIME.
On ARL with GuC 70.53.0 this shows the acks are never lost: they
consistently arrive ~2.3s after the request, tens of milliseconds after
the TDR has already signaled the fence:
TLB invalidation fence timeout, seqno=10992 recv=10991
TLB invalidation late ack: seqno=10992 recv=10992, request-to-ack=2314ms, timeout-to-ack=45ms
Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
Signed-off-by: Tales A. Mendonça <talesam@gmail.com>
---
drivers/gpu/drm/xe/xe_tlb_inval.c | 16 ++++++++++++++++
drivers/gpu/drm/xe/xe_tlb_inval_types.h | 17 +++++++++++++++++
2 files changed, 33 insertions(+)
diff --git a/drivers/gpu/drm/xe/xe_tlb_inval.c b/drivers/gpu/drm/xe/xe_tlb_inval.c
index 833fb92cd3e..9dd04d5bc4c 100644
--- a/drivers/gpu/drm/xe/xe_tlb_inval.c
+++ b/drivers/gpu/drm/xe/xe_tlb_inval.c
@@ -99,6 +99,9 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
fence->seqno, tlb_inval->seqno_recv);
timedout_seqno = fence->seqno;
+ tlb_inval->timedout_seqno = fence->seqno;
+ tlb_inval->timedout_inval_time = fence->inval_time;
+ tlb_inval->timedout_time = ktime_get();
fence->base.error = -ETIME;
xe_tlb_inval_fence_signal(fence);
@@ -227,6 +230,7 @@ void xe_tlb_inval_reset(struct xe_tlb_inval *tlb_inval)
else
pending_seqno = tlb_inval->seqno - 1;
WRITE_ONCE(tlb_inval->seqno_recv, pending_seqno);
+ tlb_inval->timedout_seqno = 0;
list_for_each_entry_safe(fence, next,
&tlb_inval->pending_fences, link)
@@ -424,6 +428,18 @@ void xe_tlb_inval_done_handler(struct xe_tlb_inval *tlb_inval, int seqno)
WRITE_ONCE(tlb_inval->seqno_recv, seqno);
+ if (tlb_inval->timedout_seqno &&
+ xe_tlb_inval_seqno_past(tlb_inval, tlb_inval->timedout_seqno)) {
+ ktime_t now = ktime_get();
+
+ drm_warn(&xe->drm,
+ "TLB invalidation late ack: seqno=%d recv=%d, request-to-ack=%lldms, timeout-to-ack=%lldms",
+ tlb_inval->timedout_seqno, seqno,
+ ktime_ms_delta(now, tlb_inval->timedout_inval_time),
+ ktime_ms_delta(now, tlb_inval->timedout_time));
+ tlb_inval->timedout_seqno = 0;
+ }
+
list_for_each_entry_safe(fence, next,
&tlb_inval->pending_fences, link) {
trace_xe_tlb_inval_fence_recv(xe, fence);
diff --git a/drivers/gpu/drm/xe/xe_tlb_inval_types.h b/drivers/gpu/drm/xe/xe_tlb_inval_types.h
index 3d1797d186f..38288966254 100644
--- a/drivers/gpu/drm/xe/xe_tlb_inval_types.h
+++ b/drivers/gpu/drm/xe/xe_tlb_inval_types.h
@@ -102,6 +102,23 @@ struct xe_tlb_inval {
* @pending_lock: protects @pending_fences and updating @seqno_recv.
*/
spinlock_t pending_lock;
+ /**
+ * @timedout_seqno: seqno of the most recent timed out TLB
+ * invalidation, 0 if none. Used to measure how late the ack for a
+ * timed out invalidation actually arrives. Protected by
+ * @pending_lock.
+ */
+ int timedout_seqno;
+ /**
+ * @timedout_inval_time: request time of @timedout_seqno. Protected by
+ * @pending_lock.
+ */
+ ktime_t timedout_inval_time;
+ /**
+ * @timedout_time: time @timedout_seqno was signaled with -ETIME.
+ * Protected by @pending_lock.
+ */
+ ktime_t timedout_time;
/**
* @fence_tdr: schedules a delayed call to xe_tlb_fence_timeout after
* the timeout interval is over.
--
2.55.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC PATCH 3/3] drm/xe: Kick GuC while TLB invalidation acks are overdue
2026-08-04 2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
2026-08-04 2:14 ` [RFC PATCH 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
2026-08-04 2:14 ` [RFC PATCH 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
@ 2026-08-04 2:14 ` Tales A. Mendonça
2026-08-04 2:15 ` ✗ LGCI.VerificationFailed: failure for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Patchwork
` (6 subsequent siblings)
9 siblings, 0 replies; 25+ messages in thread
From: Tales A. Mendonça @ 2026-08-04 2:14 UTC (permalink / raw)
To: intel-xe
Cc: matthew.brost, thomas.hellstrom, rodrigo.vivi, dri-devel,
Tales A. Mendonça
On ARL with GuC 70.53.0, TLB invalidation acks intermittently stall:
the H2G request is consumed from the CTB immediately, but the G2H ack
only arrives ~2.3s later, tens of milliseconds after the fence timeout
has fired. During the stall the GPU keeps rendering; only work blocked
on the invalidation (e.g. Wayland compositors performing buffer
unmaps) hitches for the full 2.3s. Observed on three machines so far
(7d51, 7dd1, plus an Arc Pro 130T report), see Link.
Experiments ruled out the obvious suspects:
* GT C6 parking: holding forcewake across the whole GT (C6 residency
pinned at 0ms) still produced 9 timeouts in a row.
* Lost interrupt/CT processing on the host: the G2H CTB is empty at
timeout time; the ack genuinely has not been sent by the firmware.
What does help is poking the GuC while the ack is overdue. Add a
delayed work that fires XE_TLB_INVAL_KICK_DELAY_MS after an
invalidation is issued and, while any ack is pending, re-reads the GuC
status register, flushes the CT fast-path and rings the GuC doorbell
(xe_guc_notify()), re-arming itself until the ack arrives; the
existing TDR still bounds the total wait.
Instrumented results from two ARL machines over several days:
* Without kicks: every stall lasts the full ~2.3s and is reported as
a fence timeout (-ETIME), ~1/hour on a desktop workload.
* With kicks: the majority of stalls resolve 15-276ms after one of
the kicks, e.g.:
TLB invalidation ack after kick: seqno=36405 recv=36405, request-to-ack=528ms, last-kick-to-ack=15ms, kicks=2
* A minority of severe episodes ignore 8-9 consecutive doorbells and
still run to the timeout (worst observed: ack 5996ms after request,
3.7s after the last kick), suggesting the firmware is internally
blocked for the whole window rather than missing a wake event.
This is a workaround, not a fix - the root cause looks like a GuC
firmware issue - but it turns a guaranteed 2.3s stall into a sub-500ms
hiccup for most occurrences, and the "ack after kick" log documents
the firmware behavior for further debugging.
Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
Signed-off-by: Tales A. Mendonça <talesam@gmail.com>
---
drivers/gpu/drm/xe/xe_tlb_inval.c | 95 +++++++++++++++++++++++--
drivers/gpu/drm/xe/xe_tlb_inval_types.h | 25 +++++++
2 files changed, 116 insertions(+), 4 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_tlb_inval.c b/drivers/gpu/drm/xe/xe_tlb_inval.c
index 9dd04d5bc4c..16c32d5669f 100644
--- a/drivers/gpu/drm/xe/xe_tlb_inval.c
+++ b/drivers/gpu/drm/xe/xe_tlb_inval.c
@@ -10,7 +10,9 @@
#include "xe_force_wake.h"
#include "xe_gt_stats.h"
#include "xe_gt_types.h"
+#include "xe_guc.h"
#include "xe_guc_ct.h"
+#include "xe_guc_pc.h"
#include "xe_guc_tlb_inval.h"
#include "xe_mmio.h"
#include "xe_pm.h"
@@ -30,6 +32,15 @@
#define FENCE_STACK_BIT DMA_FENCE_FLAG_USER_BITS
+/*
+ * Delay before poking the GuC about a pending invalidation that has not been
+ * acked yet. Acks normally arrive in microseconds; when the GuC stalls they
+ * only show up seconds later, after the timeout has already fired.
+ */
+#define XE_TLB_INVAL_KICK_DELAY_MS 250
+
+static void xe_tlb_inval_kick(struct work_struct *work);
+
/* The frontend is only ever embedded in a GT */
static struct xe_gt *tlb_inval_to_gt(struct xe_tlb_inval *tlb_inval)
{
@@ -54,8 +65,10 @@ xe_tlb_inval_fence_signal(struct xe_tlb_inval_fence *fence)
lockdep_assert_held(&fence->tlb_inval->pending_lock);
list_del(&fence->link);
- if (list_empty(&tlb_inval->pending_fences))
+ if (list_empty(&tlb_inval->pending_fences)) {
cancel_delayed_work(&tlb_inval->fence_tdr);
+ cancel_delayed_work(&tlb_inval->kick_work);
+ }
trace_xe_tlb_inval_fence_signal(fence->tlb_inval->xe, fence);
xe_tlb_inval_fence_fini(fence);
dma_fence_signal(&fence->base);
@@ -168,6 +181,7 @@ int xe_gt_tlb_inval_init_early(struct xe_gt *gt)
spin_lock_init(&tlb_inval->pending_lock);
spin_lock_init(&tlb_inval->lock);
INIT_DELAYED_WORK(&tlb_inval->fence_tdr, xe_tlb_inval_fence_timeout);
+ INIT_DELAYED_WORK(&tlb_inval->kick_work, xe_tlb_inval_kick);
err = drmm_mutex_init(&xe->drm, &tlb_inval->seqno_lock);
if (err)
@@ -218,6 +232,7 @@ void xe_tlb_inval_reset(struct xe_tlb_inval *tlb_inval)
mutex_lock(&tlb_inval->seqno_lock);
spin_lock_irq(&tlb_inval->pending_lock);
cancel_delayed_work(&tlb_inval->fence_tdr);
+ cancel_delayed_work(&tlb_inval->kick_work);
/*
* We might have various kworkers waiting for TLB flushes to complete
* which are not tracked with an explicit TLB fence, however at this
@@ -231,6 +246,8 @@ void xe_tlb_inval_reset(struct xe_tlb_inval *tlb_inval)
pending_seqno = tlb_inval->seqno - 1;
WRITE_ONCE(tlb_inval->seqno_recv, pending_seqno);
tlb_inval->timedout_seqno = 0;
+ tlb_inval->kicked_seqno = 0;
+ tlb_inval->kicked_count = 0;
list_for_each_entry_safe(fence, next,
&tlb_inval->pending_fences, link)
@@ -268,6 +285,55 @@ static bool xe_tlb_inval_seqno_past(struct xe_tlb_inval *tlb_inval, int seqno)
return seqno_recv >= seqno;
}
+static void xe_tlb_inval_kick(struct work_struct *work)
+{
+ struct xe_tlb_inval *tlb_inval = container_of(work, struct xe_tlb_inval,
+ kick_work.work);
+ struct xe_gt *gt = tlb_inval_to_gt(tlb_inval);
+ struct xe_tlb_inval_fence *fence;
+ ktime_t inval_time = 0;
+ int seqno = 0;
+
+ spin_lock_irq(&tlb_inval->pending_lock);
+ fence = list_first_entry_or_null(&tlb_inval->pending_fences,
+ struct xe_tlb_inval_fence, link);
+ if (fence) {
+ seqno = fence->seqno;
+ inval_time = fence->inval_time;
+ }
+ spin_unlock_irq(&tlb_inval->pending_lock);
+
+ if (!seqno)
+ return;
+
+ /*
+ * Poke the GuC: read its status register, flush the CT fast-path and
+ * ring the doorbell. On ARL with GuC 70.53.0 the ack for a pending
+ * invalidation sometimes only arrives seconds after the request even
+ * though the H2G was consumed immediately; a doorbell ring while the
+ * ack is overdue usually unsticks it within ~250ms (see Link in the
+ * commit message). Keep kicking every interval until the ack shows
+ * up; the TDR bounds how long this can go on.
+ */
+ if (gt->gtidle.idle_residency)
+ xe_guc_pc_c_status(>->uc.guc.pc);
+ tlb_inval->ops->flush(tlb_inval);
+ xe_guc_notify(>->uc.guc);
+
+ spin_lock_irq(&tlb_inval->pending_lock);
+ if (!xe_tlb_inval_seqno_past(tlb_inval, seqno)) {
+ if (tlb_inval->kicked_seqno != seqno)
+ tlb_inval->kicked_count = 0;
+ tlb_inval->kicked_seqno = seqno;
+ tlb_inval->kicked_inval_time = inval_time;
+ tlb_inval->kicked_time = ktime_get();
+ tlb_inval->kicked_count++;
+ queue_delayed_work(tlb_inval->timeout_wq, &tlb_inval->kick_work,
+ msecs_to_jiffies(XE_TLB_INVAL_KICK_DELAY_MS));
+ }
+ spin_unlock_irq(&tlb_inval->pending_lock);
+}
+
static void xe_tlb_inval_fence_prep(struct xe_tlb_inval_fence *fence)
{
struct xe_tlb_inval *tlb_inval = fence->tlb_inval;
@@ -279,9 +345,12 @@ static void xe_tlb_inval_fence_prep(struct xe_tlb_inval_fence *fence)
fence->inval_time = ktime_get();
list_add_tail(&fence->link, &tlb_inval->pending_fences);
- if (list_is_singular(&tlb_inval->pending_fences))
+ if (list_is_singular(&tlb_inval->pending_fences)) {
queue_delayed_work(tlb_inval->timeout_wq, &tlb_inval->fence_tdr,
tlb_inval->ops->timeout_delay(tlb_inval));
+ queue_delayed_work(tlb_inval->timeout_wq, &tlb_inval->kick_work,
+ msecs_to_jiffies(XE_TLB_INVAL_KICK_DELAY_MS));
+ }
spin_unlock_irq(&tlb_inval->pending_lock);
tlb_inval->seqno = (tlb_inval->seqno + 1) %
@@ -440,6 +509,20 @@ void xe_tlb_inval_done_handler(struct xe_tlb_inval *tlb_inval, int seqno)
tlb_inval->timedout_seqno = 0;
}
+ if (tlb_inval->kicked_seqno &&
+ xe_tlb_inval_seqno_past(tlb_inval, tlb_inval->kicked_seqno)) {
+ ktime_t now = ktime_get();
+
+ drm_warn(&xe->drm,
+ "TLB invalidation ack after kick: seqno=%d recv=%d, request-to-ack=%lldms, last-kick-to-ack=%lldms, kicks=%d",
+ tlb_inval->kicked_seqno, seqno,
+ ktime_ms_delta(now, tlb_inval->kicked_inval_time),
+ ktime_ms_delta(now, tlb_inval->kicked_time),
+ tlb_inval->kicked_count);
+ tlb_inval->kicked_seqno = 0;
+ tlb_inval->kicked_count = 0;
+ }
+
list_for_each_entry_safe(fence, next,
&tlb_inval->pending_fences, link) {
trace_xe_tlb_inval_fence_recv(xe, fence);
@@ -450,12 +533,16 @@ void xe_tlb_inval_done_handler(struct xe_tlb_inval *tlb_inval, int seqno)
xe_tlb_inval_fence_signal(fence);
}
- if (!list_empty(&tlb_inval->pending_fences))
+ if (!list_empty(&tlb_inval->pending_fences)) {
mod_delayed_work(tlb_inval->timeout_wq,
&tlb_inval->fence_tdr,
tlb_inval->ops->timeout_delay(tlb_inval));
- else
+ mod_delayed_work(tlb_inval->timeout_wq, &tlb_inval->kick_work,
+ msecs_to_jiffies(XE_TLB_INVAL_KICK_DELAY_MS));
+ } else {
cancel_delayed_work(&tlb_inval->fence_tdr);
+ cancel_delayed_work(&tlb_inval->kick_work);
+ }
spin_unlock_irqrestore(&tlb_inval->pending_lock, flags);
}
diff --git a/drivers/gpu/drm/xe/xe_tlb_inval_types.h b/drivers/gpu/drm/xe/xe_tlb_inval_types.h
index 38288966254..f8ae540dc4b 100644
--- a/drivers/gpu/drm/xe/xe_tlb_inval_types.h
+++ b/drivers/gpu/drm/xe/xe_tlb_inval_types.h
@@ -124,6 +124,31 @@ struct xe_tlb_inval {
* the timeout interval is over.
*/
struct delayed_work fence_tdr;
+ /**
+ * @kick_work: pokes the GuC while an invalidation ack is overdue,
+ * bounding ack stalls on GuC firmware that misses CT notifications.
+ */
+ struct delayed_work kick_work;
+ /**
+ * @kicked_seqno: seqno the last kick was issued for, 0 if none.
+ * Protected by @pending_lock.
+ */
+ int kicked_seqno;
+ /**
+ * @kicked_inval_time: request time of @kicked_seqno. Protected by
+ * @pending_lock.
+ */
+ ktime_t kicked_inval_time;
+ /**
+ * @kicked_time: time the last kick for @kicked_seqno ran. Protected
+ * by @pending_lock.
+ */
+ ktime_t kicked_time;
+ /**
+ * @kicked_count: number of kicks issued for @kicked_seqno. Protected
+ * by @pending_lock.
+ */
+ int kicked_count;
/** @job_wq: schedules TLB invalidation jobs */
struct workqueue_struct *job_wq;
/** @tlb_inval.lock: protects TLB invalidation fences */
--
2.55.0
^ permalink raw reply related [flat|nested] 25+ messages in thread
* ✗ LGCI.VerificationFailed: failure for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-04 2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
` (2 preceding siblings ...)
2026-08-04 2:14 ` [RFC PATCH 3/3] drm/xe: Kick GuC while TLB invalidation acks are overdue Tales A. Mendonça
@ 2026-08-04 2:15 ` Patchwork
2026-08-04 16:34 ` [RFC PATCH 0/3] " Tales A. Mendonça
` (5 subsequent siblings)
9 siblings, 0 replies; 25+ messages in thread
From: Patchwork @ 2026-08-04 2:15 UTC (permalink / raw)
To: Tales A. Mendonça; +Cc: intel-xe
== Series Details ==
Series: drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
URL : https://patchwork.freedesktop.org/series/171539/
State : failure
== Summary ==
Series author address 'talesam@gmail.com' is not on the allowlist, which prevents CI from being automatically triggered.
If you want CI to run for this series, ask Patchwork project owners to click 'retest' on the series in Patchwork.
Exception occurred during validation, bailing out!
Build URL: http://intel-gfx-ci-public.igk.intel.com:8080/job/xe_pw_trigger/1237486/ (on master)
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-04 2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
` (3 preceding siblings ...)
2026-08-04 2:15 ` ✗ LGCI.VerificationFailed: failure for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Patchwork
@ 2026-08-04 16:34 ` Tales A. Mendonça
2026-08-04 21:29 ` Matthew Brost
2026-08-04 21:02 ` Summers, Stuart
` (4 subsequent siblings)
9 siblings, 1 reply; 25+ messages in thread
From: Tales A. Mendonça @ 2026-08-04 16:34 UTC (permalink / raw)
To: intel-xe; +Cc: matthew.brost, thomas.hellstrom, rodrigo.vivi, dri-devel
Patchwork reports my address is not on the CI allowlist, so CI was not
triggered for this series:
Series author address 'talesam@gmail.com' is not on the allowlist,
which prevents CI from being automatically triggered.
Could one of the project owners click 'retest' on the series (and/or
add me to the allowlist)? Series URL:
https://patchwork.freedesktop.org/series/171539/
Thanks!
Tales
Em seg., 3 de ago. de 2026 às 23:14, Tales A. Mendonça
<talesam@gmail.com> escreveu:
>
> Hi,
>
> This series is a follow-up to the TLB invalidation ack stall I have
> been debugging on ARL, tracked in:
>
> https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
>
> Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51 and
> 7dd1 machines here, plus an independent Arc Pro 130T report on the
> issue above), TLB invalidation acks intermittently stall for ~2.3s.
> The H2G request is consumed from the CTB immediately and the G2H CTB
> is empty the whole time - the firmware simply does not send the ack
> until much later. The fence timeout fires at 2.25s and the ack lands
> tens of ms after it. Userspace blocked on the invalidation (compositor
> buffer unmaps etc.) hitches for the full window.
>
> Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has no
> exec queue or job to blame - leaves a devcoredump with the GuC log and
> CT state behind (Matt suggested capturing devcoredumps when we
> discussed the issue; devcoredumps from both machines are attached to
> the issue above).
>
> Patch 2 logs when the ack for a timed out invalidation finally
> arrives. This is what established that the acks are late rather than
> lost.
>
> Patch 3 is the RFC part: a delayed work that pokes the GuC (status
> register read, CT flush, doorbell ring) every 250ms while an ack is
> overdue. On my machines this converts the guaranteed 2.3s stall into a
> sub-500ms hiccup for the majority of occurrences; a minority of severe
> episodes ignore 8-9 consecutive doorbells, which points at the GuC
> firmware being internally blocked for the whole window. Full data on
> the issue. I am happy to rework the approach (different delay,
> tying it to the G2H handler, dropping the status read, etc.) - mainly
> I would like the firmware side investigated, since no host-side poke
> can fix the severe cases.
>
> Based on drm-tip. Tested for several days on both ARL machines under
> desktop and VM-heavy workloads.
>
> Thanks,
> Tales
>
> Tales A. Mendonça (3):
> drm/xe: Capture devcoredump on TLB invalidation timeout
> drm/xe: Log when a timed out TLB invalidation ack finally arrives
> drm/xe: Kick GuC while TLB invalidation acks are overdue
>
> drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
> drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
> drivers/gpu/drm/xe/xe_tlb_inval.c | 131 +++++++++++++++++++++++-
> drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
> 4 files changed, 243 insertions(+), 4 deletions(-)
>
> --
> 2.55.0
>
--
Com os cumprimentos,
Tales A. Mendonça
talesam.org
communitybig.org
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-04 2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
` (4 preceding siblings ...)
2026-08-04 16:34 ` [RFC PATCH 0/3] " Tales A. Mendonça
@ 2026-08-04 21:02 ` Summers, Stuart
2026-08-04 21:27 ` Matthew Brost
2026-08-05 12:32 ` ✗ CI.checkpatch: warning for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2) Patchwork
` (3 subsequent siblings)
9 siblings, 1 reply; 25+ messages in thread
From: Summers, Stuart @ 2026-08-04 21:02 UTC (permalink / raw)
To: intel-xe@lists.freedesktop.org, Ceraolo Spurio, Daniele,
talesam@gmail.com
Cc: dri-devel@lists.freedesktop.org, Brost, Matthew, Vivi, Rodrigo,
thomas.hellstrom@linux.intel.com
On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
> Hi,
>
> This series is a follow-up to the TLB invalidation ack stall I have
> been debugging on ARL, tracked in:
>
> https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
>
> Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51 and
> 7dd1 machines here, plus an independent Arc Pro 130T report on the
> issue above), TLB invalidation acks intermittently stall for ~2.3s.
> The H2G request is consumed from the CTB immediately and the G2H CTB
> is empty the whole time - the firmware simply does not send the ack
> until much later. The fence timeout fires at 2.25s and the ack lands
> tens of ms after it. Userspace blocked on the invalidation
> (compositor
> buffer unmaps etc.) hitches for the full window.
Firstly, thanks for the patch!
I haven't looked in to all the details of the sighting you were
debugging, but we have had similar issues that were fixed in a later
GuC version. I think around 70.60.0? It might be worth trying on
something later than that to see if that helps... (+Daniele)
>
> Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has no
Is there a reason we don't just re-use the main xe_devcoredump()?
> exec queue or job to blame - leaves a devcoredump with the GuC log
> and
> CT state behind (Matt suggested capturing devcoredumps when we
> discussed the issue; devcoredumps from both machines are attached to
> the issue above).
>
> Patch 2 logs when the ack for a timed out invalidation finally
> arrives. This is what established that the acks are late rather than
> lost.
>
> Patch 3 is the RFC part: a delayed work that pokes the GuC (status
> register read, CT flush, doorbell ring) every 250ms while an ack is
> overdue. On my machines this converts the guaranteed 2.3s stall into
I'm a little worried we're just papering over something here that needs
to be addressed in GuC, particularly around GT going to sleep or
something around the time we're expecting a response, so the pings on
registers might be prematurely waking things up which is something we'd
want to happen in GuC, not the KMD.
Thanks,
Stuart
> a
> sub-500ms hiccup for the majority of occurrences; a minority of
> severe
> episodes ignore 8-9 consecutive doorbells, which points at the GuC
> firmware being internally blocked for the whole window. Full data on
> the issue. I am happy to rework the approach (different delay,
> tying it to the G2H handler, dropping the status read, etc.) - mainly
> I would like the firmware side investigated, since no host-side poke
> can fix the severe cases.
>
> Based on drm-tip. Tested for several days on both ARL machines under
> desktop and VM-heavy workloads.
>
> Thanks,
> Tales
>
> Tales A. Mendonça (3):
> drm/xe: Capture devcoredump on TLB invalidation timeout
> drm/xe: Log when a timed out TLB invalidation ack finally arrives
> drm/xe: Kick GuC while TLB invalidation acks are overdue
>
> drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
> drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
> drivers/gpu/drm/xe/xe_tlb_inval.c | 131
> +++++++++++++++++++++++-
> drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
> 4 files changed, 243 insertions(+), 4 deletions(-)
>
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-04 21:02 ` Summers, Stuart
@ 2026-08-04 21:27 ` Matthew Brost
2026-08-04 21:33 ` Summers, Stuart
0 siblings, 1 reply; 25+ messages in thread
From: Matthew Brost @ 2026-08-04 21:27 UTC (permalink / raw)
To: Summers, Stuart
Cc: intel-xe@lists.freedesktop.org, Ceraolo Spurio, Daniele,
talesam@gmail.com, dri-devel@lists.freedesktop.org, Vivi, Rodrigo,
thomas.hellstrom@linux.intel.com, julia.filipchuk
On Tue, Aug 04, 2026 at 03:02:44PM -0600, Summers, Stuart wrote:
> On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
> > Hi,
> >
> > This series is a follow-up to the TLB invalidation ack stall I have
> > been debugging on ARL, tracked in:
> >
> > https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> >
> > Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51 and
> > 7dd1 machines here, plus an independent Arc Pro 130T report on the
> > issue above), TLB invalidation acks intermittently stall for ~2.3s.
> > The H2G request is consumed from the CTB immediately and the G2H CTB
> > is empty the whole time - the firmware simply does not send the ack
> > until much later. The fence timeout fires at 2.25s and the ack lands
> > tens of ms after it. Userspace blocked on the invalidation
> > (compositor
> > buffer unmaps etc.) hitches for the full window.
>
> Firstly, thanks for the patch!
>
> I haven't looked in to all the details of the sighting you were
> debugging, but we have had similar issues that were fixed in a later
> GuC version. I think around 70.60.0? It might be worth trying on
> something later than that to see if that helps... (+Daniele)
>
I think this would require an AR on our end to make a new firmware
version available.
The upstream repo only has 70.53.0 available for ARL [1] (iirc, ARL
aliases to MTL for firmware). (+Julia too).
Presumably, the GuC changelogs should indicate whether an issue related
this has been fixed. If so, we need to update all GuC versions across
both i915 and Xe.
[1] https://gitlab.com/kernel-firmware/linux-firmware/-/blob/main/i915/mtl_guc_70.bin?ref_type=heads
> >
> > Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has no
>
> Is there a reason we don't just re-use the main xe_devcoredump()?
>
This is my suggestion: the main devcoredump infrastructure is job-based,
so it cannot be used for hangs that are not associated with a job.
In my opinion, this is a gap on our end. Introducing something like
`xe_devcoredump_gt()`, which can be used for non-job-based hangs (e.g.,
TLB invalidation timeouts like those addressed in this series, or more
generally any GuC protocol hang), makes sense to me.
I haven't looked at the patch yet, but at a high level, adding
`xe_devcoredump_gt()` seems like a reasonable approach.
> > exec queue or job to blame - leaves a devcoredump with the GuC log
> > and
> > CT state behind (Matt suggested capturing devcoredumps when we
> > discussed the issue; devcoredumps from both machines are attached to
> > the issue above).
> >
> > Patch 2 logs when the ack for a timed out invalidation finally
> > arrives. This is what established that the acks are late rather than
> > lost.
> >
> > Patch 3 is the RFC part: a delayed work that pokes the GuC (status
> > register read, CT flush, doorbell ring) every 250ms while an ack is
> > overdue. On my machines this converts the guaranteed 2.3s stall into
>
> I'm a little worried we're just papering over something here that needs
> to be addressed in GuC, particularly around GT going to sleep or
> something around the time we're expecting a response, so the pings on
> registers might be prematurely waking things up which is something we'd
> want to happen in GuC, not the KMD.
>
In general, I agree with this. We should avoid papering over the issue
and instead fix it properly in the GuC. That said, this workaround
provides a pretty strong data point, since it appears to get the TLB
invalidation unstuck.
Matt
> Thanks,
> Stuart
>
> > a
> > sub-500ms hiccup for the majority of occurrences; a minority of
> > severe
> > episodes ignore 8-9 consecutive doorbells, which points at the GuC
> > firmware being internally blocked for the whole window. Full data on
> > the issue. I am happy to rework the approach (different delay,
> > tying it to the G2H handler, dropping the status read, etc.) - mainly
> > I would like the firmware side investigated, since no host-side poke
> > can fix the severe cases.
> >
> > Based on drm-tip. Tested for several days on both ARL machines under
> > desktop and VM-heavy workloads.
> >
> > Thanks,
> > Tales
> >
> > Tales A. Mendonça (3):
> > drm/xe: Capture devcoredump on TLB invalidation timeout
> > drm/xe: Log when a timed out TLB invalidation ack finally arrives
> > drm/xe: Kick GuC while TLB invalidation acks are overdue
> >
> > drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
> > drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
> > drivers/gpu/drm/xe/xe_tlb_inval.c | 131
> > +++++++++++++++++++++++-
> > drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
> > 4 files changed, 243 insertions(+), 4 deletions(-)
> >
>
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-04 16:34 ` [RFC PATCH 0/3] " Tales A. Mendonça
@ 2026-08-04 21:29 ` Matthew Brost
2026-08-05 20:24 ` Matthew Brost
0 siblings, 1 reply; 25+ messages in thread
From: Matthew Brost @ 2026-08-04 21:29 UTC (permalink / raw)
To: Tales A. Mendonça
Cc: intel-xe, thomas.hellstrom, rodrigo.vivi, dri-devel
On Tue, Aug 04, 2026 at 01:34:19PM -0300, Tales A. Mendonça wrote:
> Patchwork reports my address is not on the CI allowlist, so CI was not
> triggered for this series:
>
> Series author address 'talesam@gmail.com' is not on the allowlist,
> which prevents CI from being automatically triggered.
>
> Could one of the project owners click 'retest' on the series (and/or
> add me to the allowlist)? Series URL:
>
> https://patchwork.freedesktop.org/series/171539/
>
We'd have to resend this ourselves. I can do this, but I've requested
for you to be on our allow list as well. I'll ping here once that goes
through.
Matt
> Thanks!
> Tales
>
>
> Em seg., 3 de ago. de 2026 às 23:14, Tales A. Mendonça
> <talesam@gmail.com> escreveu:
> >
> > Hi,
> >
> > This series is a follow-up to the TLB invalidation ack stall I have
> > been debugging on ARL, tracked in:
> >
> > https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> >
> > Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51 and
> > 7dd1 machines here, plus an independent Arc Pro 130T report on the
> > issue above), TLB invalidation acks intermittently stall for ~2.3s.
> > The H2G request is consumed from the CTB immediately and the G2H CTB
> > is empty the whole time - the firmware simply does not send the ack
> > until much later. The fence timeout fires at 2.25s and the ack lands
> > tens of ms after it. Userspace blocked on the invalidation (compositor
> > buffer unmaps etc.) hitches for the full window.
> >
> > Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has no
> > exec queue or job to blame - leaves a devcoredump with the GuC log and
> > CT state behind (Matt suggested capturing devcoredumps when we
> > discussed the issue; devcoredumps from both machines are attached to
> > the issue above).
> >
> > Patch 2 logs when the ack for a timed out invalidation finally
> > arrives. This is what established that the acks are late rather than
> > lost.
> >
> > Patch 3 is the RFC part: a delayed work that pokes the GuC (status
> > register read, CT flush, doorbell ring) every 250ms while an ack is
> > overdue. On my machines this converts the guaranteed 2.3s stall into a
> > sub-500ms hiccup for the majority of occurrences; a minority of severe
> > episodes ignore 8-9 consecutive doorbells, which points at the GuC
> > firmware being internally blocked for the whole window. Full data on
> > the issue. I am happy to rework the approach (different delay,
> > tying it to the G2H handler, dropping the status read, etc.) - mainly
> > I would like the firmware side investigated, since no host-side poke
> > can fix the severe cases.
> >
> > Based on drm-tip. Tested for several days on both ARL machines under
> > desktop and VM-heavy workloads.
> >
> > Thanks,
> > Tales
> >
> > Tales A. Mendonça (3):
> > drm/xe: Capture devcoredump on TLB invalidation timeout
> > drm/xe: Log when a timed out TLB invalidation ack finally arrives
> > drm/xe: Kick GuC while TLB invalidation acks are overdue
> >
> > drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
> > drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
> > drivers/gpu/drm/xe/xe_tlb_inval.c | 131 +++++++++++++++++++++++-
> > drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
> > 4 files changed, 243 insertions(+), 4 deletions(-)
> >
> > --
> > 2.55.0
> >
>
>
> --
> Com os cumprimentos,
>
> Tales A. Mendonça
> talesam.org
> communitybig.org
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-04 21:27 ` Matthew Brost
@ 2026-08-04 21:33 ` Summers, Stuart
2026-08-04 22:08 ` Daniele Ceraolo Spurio
0 siblings, 1 reply; 25+ messages in thread
From: Summers, Stuart @ 2026-08-04 21:33 UTC (permalink / raw)
To: Brost, Matthew
Cc: intel-xe@lists.freedesktop.org, dri-devel@lists.freedesktop.org,
Vivi, Rodrigo, Ceraolo Spurio, Daniele, talesam@gmail.com,
thomas.hellstrom@linux.intel.com, Filipchuk, Julia
On Tue, 2026-08-04 at 14:27 -0700, Matthew Brost wrote:
> On Tue, Aug 04, 2026 at 03:02:44PM -0600, Summers, Stuart wrote:
> > On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
> > > Hi,
> > >
> > > This series is a follow-up to the TLB invalidation ack stall I
> > > have
> > > been debugging on ARL, tracked in:
> > >
> > > https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> > >
> > > Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51
> > > and
> > > 7dd1 machines here, plus an independent Arc Pro 130T report on
> > > the
> > > issue above), TLB invalidation acks intermittently stall for
> > > ~2.3s.
> > > The H2G request is consumed from the CTB immediately and the G2H
> > > CTB
> > > is empty the whole time - the firmware simply does not send the
> > > ack
> > > until much later. The fence timeout fires at 2.25s and the ack
> > > lands
> > > tens of ms after it. Userspace blocked on the invalidation
> > > (compositor
> > > buffer unmaps etc.) hitches for the full window.
> >
> > Firstly, thanks for the patch!
> >
> > I haven't looked in to all the details of the sighting you were
> > debugging, but we have had similar issues that were fixed in a
> > later
> > GuC version. I think around 70.60.0? It might be worth trying on
> > something later than that to see if that helps... (+Daniele)
> >
>
> I think this would require an AR on our end to make a new firmware
> version available.
>
> The upstream repo only has 70.53.0 available for ARL [1] (iirc, ARL
> aliases to MTL for firmware). (+Julia too).
>
> Presumably, the GuC changelogs should indicate whether an issue
> related
> this has been fixed. If so, we need to update all GuC versions across
> both i915 and Xe.
Right... I guess I'd still like to see if we can test this in GuC (or
get confirmation we can't for some reason) before committing something.
My worry is we will prevent bug reports like this by working around it
and miss critical bugs that need to be fixed in the right component.
>
> [1]
> https://gitlab.com/kernel-firmware/linux-firmware/-/blob/main/i915/mtl_guc_70.bin?ref_type=heads
>
> > >
> > > Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has
> > > no
> >
> > Is there a reason we don't just re-use the main xe_devcoredump()?
> >
>
> This is my suggestion: the main devcoredump infrastructure is job-
> based,
> so it cannot be used for hangs that are not associated with a job.
>
> In my opinion, this is a gap on our end. Introducing something like
> `xe_devcoredump_gt()`, which can be used for non-job-based hangs
> (e.g.,
> TLB invalidation timeouts like those addressed in this series, or
> more
> generally any GuC protocol hang), makes sense to me.
Ok makes sense. We can do that here. It would be nice to have a more
inclusive implementation that lets us call this from anywhere so we
aren't duplicating things around for different use cases. But not a
blocker here.
>
> I haven't looked at the patch yet, but at a high level, adding
> `xe_devcoredump_gt()` seems like a reasonable approach.
>
> > > exec queue or job to blame - leaves a devcoredump with the GuC
> > > log
> > > and
> > > CT state behind (Matt suggested capturing devcoredumps when we
> > > discussed the issue; devcoredumps from both machines are attached
> > > to
> > > the issue above).
> > >
> > > Patch 2 logs when the ack for a timed out invalidation finally
> > > arrives. This is what established that the acks are late rather
> > > than
> > > lost.
> > >
> > > Patch 3 is the RFC part: a delayed work that pokes the GuC
> > > (status
> > > register read, CT flush, doorbell ring) every 250ms while an ack
> > > is
> > > overdue. On my machines this converts the guaranteed 2.3s stall
> > > into
> >
> > I'm a little worried we're just papering over something here that
> > needs
> > to be addressed in GuC, particularly around GT going to sleep or
> > something around the time we're expecting a response, so the pings
> > on
> > registers might be prematurely waking things up which is something
> > we'd
> > want to happen in GuC, not the KMD.
> >
>
> In general, I agree with this. We should avoid papering over the
> issue
> and instead fix it properly in the GuC. That said, this workaround
> provides a pretty strong data point, since it appears to get the TLB
> invalidation unstuck.
So if we hit this issue I guess we're already going to have some
performance degredation and the workaround makes that better. I need to
look at the implementation, but we could be potentially introducing
performance penalties in other areas doing these pings.
Again, I'd like to see if we can fix this in the right place before
implementing a workaround for it. Hopefully Daniele or Julia can give
some direction there.
Thanks,
Stuart
>
> Matt
>
> > Thanks,
> > Stuart
> >
> > > a
> > > sub-500ms hiccup for the majority of occurrences; a minority of
> > > severe
> > > episodes ignore 8-9 consecutive doorbells, which points at the
> > > GuC
> > > firmware being internally blocked for the whole window. Full data
> > > on
> > > the issue. I am happy to rework the approach (different delay,
> > > tying it to the G2H handler, dropping the status read, etc.) -
> > > mainly
> > > I would like the firmware side investigated, since no host-side
> > > poke
> > > can fix the severe cases.
> > >
> > > Based on drm-tip. Tested for several days on both ARL machines
> > > under
> > > desktop and VM-heavy workloads.
> > >
> > > Thanks,
> > > Tales
> > >
> > > Tales A. Mendonça (3):
> > > drm/xe: Capture devcoredump on TLB invalidation timeout
> > > drm/xe: Log when a timed out TLB invalidation ack finally
> > > arrives
> > > drm/xe: Kick GuC while TLB invalidation acks are overdue
> > >
> > > drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
> > > drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
> > > drivers/gpu/drm/xe/xe_tlb_inval.c | 131
> > > +++++++++++++++++++++++-
> > > drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
> > > 4 files changed, 243 insertions(+), 4 deletions(-)
> > >
> >
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout
2026-08-04 2:14 ` [RFC PATCH 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
@ 2026-08-04 22:05 ` Matthew Brost
0 siblings, 0 replies; 25+ messages in thread
From: Matthew Brost @ 2026-08-04 22:05 UTC (permalink / raw)
To: Tales A. Mendonça
Cc: intel-xe, thomas.hellstrom, rodrigo.vivi, dri-devel
On Mon, Aug 03, 2026 at 11:14:39PM -0300, Tales A. Mendonça wrote:
> TLB invalidation timeouts currently leave no record of the firmware
> state behind: there is no exec queue or job to blame, so nothing calls
> xe_devcoredump() and the GuC log content at the time of the hang is
> lost.
>
> Add xe_devcoredump_gt(), a variant of xe_devcoredump() for hangs that
> are not tied to an exec queue or job. It captures the GuC log and CT
> state of the affected GT, reusing the existing snapshot machinery and
> the "only first snapshot" policy, and hook it up to the TLB invalidation
> timeout path.
>
> This was instrumental in diagnosing GuC TLB invalidation ack stalls on
> ARL (see Link), where the invalidation request is consumed from the H2G
> CTB immediately but the ack G2H only arrives ~2.3s later, after the
> timeout has already fired.
>
Thanks for doing this. A couple suggestions.
> Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> Signed-off-by: Tales A. Mendonça <talesam@gmail.com>
> ---
> drivers/gpu/drm/xe/xe_devcoredump.c | 68 +++++++++++++++++++++++++++++
> drivers/gpu/drm/xe/xe_devcoredump.h | 6 +++
> drivers/gpu/drm/xe/xe_tlb_inval.c | 20 +++++++++
> 3 files changed, 94 insertions(+)
>
> diff --git a/drivers/gpu/drm/xe/xe_devcoredump.c b/drivers/gpu/drm/xe/xe_devcoredump.c
> index 5f2b90b18f9..0ccaed176a4 100644
> --- a/drivers/gpu/drm/xe/xe_devcoredump.c
> +++ b/drivers/gpu/drm/xe/xe_devcoredump.c
> @@ -403,6 +403,74 @@ void xe_devcoredump(struct xe_exec_queue *q, struct xe_sched_job *job, const cha
> mutex_unlock(&coredump->lock);
> }
>
> +static void devcoredump_snapshot_gt(struct xe_devcoredump *coredump,
> + struct xe_gt *gt)
> +{
> + struct xe_devcoredump_snapshot *ss = &coredump->snapshot;
> + struct xe_guc *guc = >->uc.guc;
> + bool cookie;
> +
> + ss->snapshot_time = ktime_get_real();
> + ss->boot_time = ktime_get_boottime();
> +
> + strscpy(ss->process_name, "no process");
> +
> + ss->gt = gt;
> + INIT_WORK(&ss->work, xe_devcoredump_deferred_snap_work);
> +
> + /* keep going if fw fails as we still want to save the SW data */
> + CLASS(xe_force_wake, fw_ref)(gt_to_fw(gt), XE_FORCEWAKE_ALL);
> +
> + cookie = dma_fence_begin_signalling();
> +
> + ss->guc.log = xe_guc_log_snapshot_capture(&guc->log, true);
> + ss->guc.ct = xe_guc_ct_snapshot_capture(&guc->ct);
> +
> + queue_work(system_dfl_wq, &ss->work);
> +
> + dma_fence_end_signalling(cookie);
> +}
> +
> +/**
> + * xe_devcoredump_gt - Take GT-level snapshots and initialize coredump device.
> + * @gt: The GT where the issue was detected.
> + * @fmt: Printf format + args to describe the reason for the core dump
> + *
> + * Variant of xe_devcoredump() for hangs that are not tied to an exec queue
> + * or job, e.g. TLB invalidation timeouts. Captures the GuC log and CT state
> + * of @gt so the firmware side of the hang can be inspected. Skipped if a
> + * coredump is already captured, same as xe_devcoredump().
> + */
> +__printf(2, 3)
> +void xe_devcoredump_gt(struct xe_gt *gt, const char *fmt, ...)
I'd unify these functions with the existing
xe_devcoredump/devcoredump_snapshot by adding a GT argument to both and
teaching those functions that 'q' can be NULL. Then replace
s/xe_devcoredump/__xe_devcoredump/ and add wrapper macros
xe_devcoredump and xe_devcoredump_gt in the header file.
More below.
> +{
> + struct xe_device *xe = gt_to_xe(gt);
> + struct xe_devcoredump *coredump = &xe->devcoredump;
> + va_list varg;
> +
> + mutex_lock(&coredump->lock);
> +
> + if (coredump->captured) {
> + drm_dbg(&xe->drm, "Multiple hangs are occurring, but only the first snapshot was taken\n");
> + mutex_unlock(&coredump->lock);
> + return;
> + }
> +
> + coredump->captured = true;
> +
> + va_start(varg, fmt);
> + coredump->snapshot.reason = kvasprintf(GFP_ATOMIC, fmt, varg);
> + va_end(varg);
> +
> + devcoredump_snapshot_gt(coredump, gt);
> +
> + drm_info(&xe->drm, "Xe device coredump has been created\n");
> + drm_info(&xe->drm, "Check your /sys/class/drm/card%d/device/devcoredump/data\n",
> + xe->drm.primary->index);
> +
> + mutex_unlock(&coredump->lock);
> +}
> +
> static void xe_driver_devcoredump_fini(void *arg)
> {
> struct drm_device *drm = arg;
> diff --git a/drivers/gpu/drm/xe/xe_devcoredump.h b/drivers/gpu/drm/xe/xe_devcoredump.h
> index 5391a80a4d1..f071bd11f24 100644
> --- a/drivers/gpu/drm/xe/xe_devcoredump.h
> +++ b/drivers/gpu/drm/xe/xe_devcoredump.h
> @@ -11,10 +11,12 @@
> struct drm_printer;
> struct xe_device;
> struct xe_exec_queue;
> +struct xe_gt;
> struct xe_sched_job;
>
> #ifdef CONFIG_DEV_COREDUMP
> void xe_devcoredump(struct xe_exec_queue *q, struct xe_sched_job *job, const char *fmt, ...);
> +void xe_devcoredump_gt(struct xe_gt *gt, const char *fmt, ...);
This what I'm suggesting for a header...
void __xe_devcoredump(struct xe_gt *gt, struct xe_exec_queue *q,
struct xe_sched_job *job, const char *fmt, ...);
#define xe_devcoredump(_q, _job, _fmt, ...) \
__xe_devcoredump((_q)->gt, _q, _job, _fmt, ##__VA_ARGS__)
#define xe_devcoredump_gt(_gt, _fmt, ...) \
__xe_devcoredump(_gt, NULL, NULL, _fmt, ##__VA_ARGS__)
I think we need wrapper macros rather than inline wrappers because of
how .../##__VA_ARGS__ work.
Matt
> int xe_devcoredump_init(struct xe_device *xe);
> #else
> static inline void xe_devcoredump(struct xe_exec_queue *q,
> @@ -23,6 +25,10 @@ static inline void xe_devcoredump(struct xe_exec_queue *q,
> {
> }
>
> +static inline void xe_devcoredump_gt(struct xe_gt *gt, const char *fmt, ...)
> +{
> +}
> +
> static inline int xe_devcoredump_init(struct xe_device *xe)
> {
> return 0;
> diff --git a/drivers/gpu/drm/xe/xe_tlb_inval.c b/drivers/gpu/drm/xe/xe_tlb_inval.c
> index bbd21d39306..833fb92cd3e 100644
> --- a/drivers/gpu/drm/xe/xe_tlb_inval.c
> +++ b/drivers/gpu/drm/xe/xe_tlb_inval.c
> @@ -5,6 +5,7 @@
>
> #include <drm/drm_managed.h>
>
> +#include "xe_devcoredump.h"
> #include "xe_device_types.h"
> #include "xe_force_wake.h"
> #include "xe_gt_stats.h"
> @@ -29,6 +30,12 @@
>
> #define FENCE_STACK_BIT DMA_FENCE_FLAG_USER_BITS
>
> +/* The frontend is only ever embedded in a GT */
> +static struct xe_gt *tlb_inval_to_gt(struct xe_tlb_inval *tlb_inval)
> +{
> + return container_of(tlb_inval, struct xe_gt, tlb_inval);
> +}
> +
> static void xe_tlb_inval_fence_fini(struct xe_tlb_inval_fence *fence)
> {
> if (WARN_ON_ONCE(!fence->tlb_inval))
> @@ -73,6 +80,7 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
> struct xe_device *xe = tlb_inval->xe;
> struct xe_tlb_inval_fence *fence, *next;
> long timeout_delay = tlb_inval->ops->timeout_delay(tlb_inval);
> + int timedout_seqno = 0;
>
> tlb_inval->ops->flush(tlb_inval);
>
> @@ -90,6 +98,8 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
> "TLB invalidation fence timeout, seqno=%d recv=%d",
> fence->seqno, tlb_inval->seqno_recv);
>
> + timedout_seqno = fence->seqno;
> +
> fence->base.error = -ETIME;
> xe_tlb_inval_fence_signal(fence);
> }
> @@ -97,6 +107,16 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
> queue_delayed_work(tlb_inval->timeout_wq, &tlb_inval->fence_tdr,
> timeout_delay);
> spin_unlock_irq(&tlb_inval->pending_lock);
> +
> + /*
> + * Capture the GuC log and CT state so the firmware side of the hang
> + * can be inspected; there is no queue or job to blame here. Must be
> + * outside pending_lock as the capture takes sleeping locks.
> + */
> + if (timedout_seqno)
> + xe_devcoredump_gt(tlb_inval_to_gt(tlb_inval),
> + "TLB invalidation fence timeout, seqno=%d recv=%d",
> + timedout_seqno, tlb_inval->seqno_recv);
> }
>
> /**
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-04 21:33 ` Summers, Stuart
@ 2026-08-04 22:08 ` Daniele Ceraolo Spurio
2026-08-04 23:00 ` Tales A. Mendonça
0 siblings, 1 reply; 25+ messages in thread
From: Daniele Ceraolo Spurio @ 2026-08-04 22:08 UTC (permalink / raw)
To: Summers, Stuart, Brost, Matthew
Cc: intel-xe@lists.freedesktop.org, dri-devel@lists.freedesktop.org,
Vivi, Rodrigo, talesam@gmail.com,
thomas.hellstrom@linux.intel.com, Filipchuk, Julia
On 8/4/2026 2:33 PM, Summers, Stuart wrote:
> On Tue, 2026-08-04 at 14:27 -0700, Matthew Brost wrote:
>> On Tue, Aug 04, 2026 at 03:02:44PM -0600, Summers, Stuart wrote:
>>> On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
>>>> Hi,
>>>>
>>>> This series is a follow-up to the TLB invalidation ack stall I
>>>> have
>>>> been debugging on ARL, tracked in:
>>>>
>>>> https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
>>>>
>>>> Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51
>>>> and
>>>> 7dd1 machines here, plus an independent Arc Pro 130T report on
>>>> the
>>>> issue above), TLB invalidation acks intermittently stall for
>>>> ~2.3s.
>>>> The H2G request is consumed from the CTB immediately and the G2H
>>>> CTB
>>>> is empty the whole time - the firmware simply does not send the
>>>> ack
>>>> until much later. The fence timeout fires at 2.25s and the ack
>>>> lands
>>>> tens of ms after it. Userspace blocked on the invalidation
>>>> (compositor
>>>> buffer unmaps etc.) hitches for the full window.
>>> Firstly, thanks for the patch!
>>>
>>> I haven't looked in to all the details of the sighting you were
>>> debugging, but we have had similar issues that were fixed in a
>>> later
>>> GuC version. I think around 70.60.0? It might be worth trying on
>>> something later than that to see if that helps... (+Daniele)
>>>
>> I think this would require an AR on our end to make a new firmware
>> version available.
>>
>> The upstream repo only has 70.53.0 available for ARL [1] (iirc, ARL
>> aliases to MTL for firmware). (+Julia too).
>>
>> Presumably, the GuC changelogs should indicate whether an issue
>> related
>> this has been fixed. If so, we need to update all GuC versions across
>> both i915 and Xe.
> Right... I guess I'd still like to see if we can test this in GuC (or
> get confirmation we can't for some reason) before committing something.
> My worry is we will prevent bug reports like this by working around it
> and miss critical bugs that need to be fixed in the right component.
>
>> [1]
>> https://gitlab.com/kernel-firmware/linux-firmware/-/blob/main/i915/mtl_guc_70.bin?ref_type=heads
>>
>>>> Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has
>>>> no
>>> Is there a reason we don't just re-use the main xe_devcoredump()?
>>>
>> This is my suggestion: the main devcoredump infrastructure is job-
>> based,
>> so it cannot be used for hangs that are not associated with a job.
>>
>> In my opinion, this is a gap on our end. Introducing something like
>> `xe_devcoredump_gt()`, which can be used for non-job-based hangs
>> (e.g.,
>> TLB invalidation timeouts like those addressed in this series, or
>> more
>> generally any GuC protocol hang), makes sense to me.
> Ok makes sense. We can do that here. It would be nice to have a more
> inclusive implementation that lets us call this from anywhere so we
> aren't duplicating things around for different use cases. But not a
> blocker here.
>
>> I haven't looked at the patch yet, but at a high level, adding
>> `xe_devcoredump_gt()` seems like a reasonable approach.
>>
>>>> exec queue or job to blame - leaves a devcoredump with the GuC
>>>> log
>>>> and
>>>> CT state behind (Matt suggested capturing devcoredumps when we
>>>> discussed the issue; devcoredumps from both machines are attached
>>>> to
>>>> the issue above).
>>>>
>>>> Patch 2 logs when the ack for a timed out invalidation finally
>>>> arrives. This is what established that the acks are late rather
>>>> than
>>>> lost.
>>>>
>>>> Patch 3 is the RFC part: a delayed work that pokes the GuC
>>>> (status
>>>> register read, CT flush, doorbell ring) every 250ms while an ack
>>>> is
>>>> overdue. On my machines this converts the guaranteed 2.3s stall
>>>> into
>>> I'm a little worried we're just papering over something here that
>>> needs
>>> to be addressed in GuC, particularly around GT going to sleep or
>>> something around the time we're expecting a response, so the pings
>>> on
>>> registers might be prematurely waking things up which is something
>>> we'd
>>> want to happen in GuC, not the KMD.
>>>
>> In general, I agree with this. We should avoid papering over the
>> issue
>> and instead fix it properly in the GuC. That said, this workaround
>> provides a pretty strong data point, since it appears to get the TLB
>> invalidation unstuck.
> So if we hit this issue I guess we're already going to have some
> performance degredation and the workaround makes that better. I need to
> look at the implementation, but we could be potentially introducing
> performance penalties in other areas doing these pings.
>
> Again, I'd like to see if we can fix this in the right place before
> implementing a workaround for it. Hopefully Daniele or Julia can give
> some direction there.
Are we seeing this on i915 at all? Given that Xe does not officially
support MTL/ARL and is missing several critical WAs for those platforms,
the approach so far has been to only update the GuC FW if it is required
for i915.
Looking at the GuC release notes, there have been a couple of
TLB-related fixes after 70.53, but they're both marked as only affecting
PVC and Xe2+ platforms, so no fixes seem to be available for ARL (or at
least they're not listed in the release notes).
Daniele
>
> Thanks,
> Stuart
>
>> Matt
>>
>>> Thanks,
>>> Stuart
>>>
>>>> a
>>>> sub-500ms hiccup for the majority of occurrences; a minority of
>>>> severe
>>>> episodes ignore 8-9 consecutive doorbells, which points at the
>>>> GuC
>>>> firmware being internally blocked for the whole window. Full data
>>>> on
>>>> the issue. I am happy to rework the approach (different delay,
>>>> tying it to the G2H handler, dropping the status read, etc.) -
>>>> mainly
>>>> I would like the firmware side investigated, since no host-side
>>>> poke
>>>> can fix the severe cases.
>>>>
>>>> Based on drm-tip. Tested for several days on both ARL machines
>>>> under
>>>> desktop and VM-heavy workloads.
>>>>
>>>> Thanks,
>>>> Tales
>>>>
>>>> Tales A. Mendonça (3):
>>>> drm/xe: Capture devcoredump on TLB invalidation timeout
>>>> drm/xe: Log when a timed out TLB invalidation ack finally
>>>> arrives
>>>> drm/xe: Kick GuC while TLB invalidation acks are overdue
>>>>
>>>> drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
>>>> drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
>>>> drivers/gpu/drm/xe/xe_tlb_inval.c | 131
>>>> +++++++++++++++++++++++-
>>>> drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
>>>> 4 files changed, 243 insertions(+), 4 deletions(-)
>>>>
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives
2026-08-04 2:14 ` [RFC PATCH 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
@ 2026-08-04 22:17 ` Matthew Brost
0 siblings, 0 replies; 25+ messages in thread
From: Matthew Brost @ 2026-08-04 22:17 UTC (permalink / raw)
To: Tales A. Mendonça
Cc: intel-xe, thomas.hellstrom, rodrigo.vivi, dri-devel
On Mon, Aug 03, 2026 at 11:14:40PM -0300, Tales A. Mendonça wrote:
> When a TLB invalidation fence times out we log the timeout, but if the
> ack for that seqno later shows up there is no record of it, making it
> impossible to tell from logs whether the ack was lost forever or merely
> (very) late.
>
> Track the most recent timed out seqno and log how late its ack arrives,
> relative to both the original request and the moment the fence was
> signaled with -ETIME.
>
> On ARL with GuC 70.53.0 this shows the acks are never lost: they
> consistently arrive ~2.3s after the request, tens of milliseconds after
> the TDR has already signaled the fence:
>
> TLB invalidation fence timeout, seqno=10992 recv=10991
> TLB invalidation late ack: seqno=10992 recv=10992, request-to-ack=2314ms, timeout-to-ack=45ms
>
> Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> Signed-off-by: Tales A. Mendonça <talesam@gmail.com>
> ---
> drivers/gpu/drm/xe/xe_tlb_inval.c | 16 ++++++++++++++++
> drivers/gpu/drm/xe/xe_tlb_inval_types.h | 17 +++++++++++++++++
> 2 files changed, 33 insertions(+)
>
> diff --git a/drivers/gpu/drm/xe/xe_tlb_inval.c b/drivers/gpu/drm/xe/xe_tlb_inval.c
> index 833fb92cd3e..9dd04d5bc4c 100644
> --- a/drivers/gpu/drm/xe/xe_tlb_inval.c
> +++ b/drivers/gpu/drm/xe/xe_tlb_inval.c
> @@ -99,6 +99,9 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
> fence->seqno, tlb_inval->seqno_recv);
>
> timedout_seqno = fence->seqno;
Should this be:
if (!tlb_inval->timedout_seqno) {
tlb_inval->timedout_seqno = fence->seqno;
tlb_inval->timedout_inval_time = fence->inval_time;
tlb_inval->timedout_time = ktime_get();
}
To record the very first timeout seqno? I think this makes more sense.
> + tlb_inval->timedout_seqno = fence->seqno;
> + tlb_inval->timedout_inval_time = fence->inval_time;
> + tlb_inval->timedout_time = ktime_get();
>
> fence->base.error = -ETIME;
> xe_tlb_inval_fence_signal(fence);
> @@ -227,6 +230,7 @@ void xe_tlb_inval_reset(struct xe_tlb_inval *tlb_inval)
> else
> pending_seqno = tlb_inval->seqno - 1;
> WRITE_ONCE(tlb_inval->seqno_recv, pending_seqno);
> + tlb_inval->timedout_seqno = 0;
>
> list_for_each_entry_safe(fence, next,
> &tlb_inval->pending_fences, link)
> @@ -424,6 +428,18 @@ void xe_tlb_inval_done_handler(struct xe_tlb_inval *tlb_inval, int seqno)
>
> WRITE_ONCE(tlb_inval->seqno_recv, seqno);
>
> + if (tlb_inval->timedout_seqno &&
> + xe_tlb_inval_seqno_past(tlb_inval, tlb_inval->timedout_seqno)) {
> + ktime_t now = ktime_get();
> +
> + drm_warn(&xe->drm,
I think xe_warn is the preference here.
Matt
> + "TLB invalidation late ack: seqno=%d recv=%d, request-to-ack=%lldms, timeout-to-ack=%lldms",
> + tlb_inval->timedout_seqno, seqno,
> + ktime_ms_delta(now, tlb_inval->timedout_inval_time),
> + ktime_ms_delta(now, tlb_inval->timedout_time));
> + tlb_inval->timedout_seqno = 0;
> + }
> +
> list_for_each_entry_safe(fence, next,
> &tlb_inval->pending_fences, link) {
> trace_xe_tlb_inval_fence_recv(xe, fence);
> diff --git a/drivers/gpu/drm/xe/xe_tlb_inval_types.h b/drivers/gpu/drm/xe/xe_tlb_inval_types.h
> index 3d1797d186f..38288966254 100644
> --- a/drivers/gpu/drm/xe/xe_tlb_inval_types.h
> +++ b/drivers/gpu/drm/xe/xe_tlb_inval_types.h
> @@ -102,6 +102,23 @@ struct xe_tlb_inval {
> * @pending_lock: protects @pending_fences and updating @seqno_recv.
> */
> spinlock_t pending_lock;
> + /**
> + * @timedout_seqno: seqno of the most recent timed out TLB
> + * invalidation, 0 if none. Used to measure how late the ack for a
> + * timed out invalidation actually arrives. Protected by
> + * @pending_lock.
> + */
> + int timedout_seqno;
> + /**
> + * @timedout_inval_time: request time of @timedout_seqno. Protected by
> + * @pending_lock.
> + */
> + ktime_t timedout_inval_time;
> + /**
> + * @timedout_time: time @timedout_seqno was signaled with -ETIME.
> + * Protected by @pending_lock.
> + */
> + ktime_t timedout_time;
> /**
> * @fence_tdr: schedules a delayed call to xe_tlb_fence_timeout after
> * the timeout interval is over.
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-04 22:08 ` Daniele Ceraolo Spurio
@ 2026-08-04 23:00 ` Tales A. Mendonça
2026-08-04 23:50 ` Daniele Ceraolo Spurio
0 siblings, 1 reply; 25+ messages in thread
From: Tales A. Mendonça @ 2026-08-04 23:00 UTC (permalink / raw)
To: Daniele Ceraolo Spurio
Cc: Summers, Stuart, Brost, Matthew, intel-xe@lists.freedesktop.org,
dri-devel@lists.freedesktop.org, Vivi, Rodrigo,
thomas.hellstrom@linux.intel.com, Filipchuk, Julia
> Are we seeing this on i915?
I have not tested i915 on the affected machines yet - the
instrumentation that measured the stalls (late-ack logging, kick
results) is xe-only, so I have no comparable i915 data. I can boot one
of the ARL machines with i915 for a few days and watch for TLB
invalidation timeouts there, if that data helps.
More generally: both machines here reproduce reliably (~1 stall/hour
on a desktop workload, much more under memory pressure), so I am happy
to test anything on them - including any GuC build the firmware team
would like data on.
On Stuart's masking concern: fully agreed, that is why patch 3 is
marked RFC. Patches 1-2 are pure diagnostics and stand on their own; I
am fine holding patch 3 until the firmware side has been looked at.
The data point it adds is that a doorbell ring unblocks the ack in the
majority of episodes, while the severe ones ignore 8-9 consecutive
rings - hopefully that narrows where to look inside the GuC.
I will send a v2 addressing Matt's review comments (the
__xe_devcoredump unification and the fixes on patch 2).
Thanks,
Tales
Em ter., 4 de ago. de 2026 às 19:08, Daniele Ceraolo Spurio
<daniele.ceraolospurio@intel.com> escreveu:
>
>
>
> On 8/4/2026 2:33 PM, Summers, Stuart wrote:
> > On Tue, 2026-08-04 at 14:27 -0700, Matthew Brost wrote:
> >> On Tue, Aug 04, 2026 at 03:02:44PM -0600, Summers, Stuart wrote:
> >>> On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
> >>>> Hi,
> >>>>
> >>>> This series is a follow-up to the TLB invalidation ack stall I
> >>>> have
> >>>> been debugging on ARL, tracked in:
> >>>>
> >>>> https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> >>>>
> >>>> Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51
> >>>> and
> >>>> 7dd1 machines here, plus an independent Arc Pro 130T report on
> >>>> the
> >>>> issue above), TLB invalidation acks intermittently stall for
> >>>> ~2.3s.
> >>>> The H2G request is consumed from the CTB immediately and the G2H
> >>>> CTB
> >>>> is empty the whole time - the firmware simply does not send the
> >>>> ack
> >>>> until much later. The fence timeout fires at 2.25s and the ack
> >>>> lands
> >>>> tens of ms after it. Userspace blocked on the invalidation
> >>>> (compositor
> >>>> buffer unmaps etc.) hitches for the full window.
> >>> Firstly, thanks for the patch!
> >>>
> >>> I haven't looked in to all the details of the sighting you were
> >>> debugging, but we have had similar issues that were fixed in a
> >>> later
> >>> GuC version. I think around 70.60.0? It might be worth trying on
> >>> something later than that to see if that helps... (+Daniele)
> >>>
> >> I think this would require an AR on our end to make a new firmware
> >> version available.
> >>
> >> The upstream repo only has 70.53.0 available for ARL [1] (iirc, ARL
> >> aliases to MTL for firmware). (+Julia too).
> >>
> >> Presumably, the GuC changelogs should indicate whether an issue
> >> related
> >> this has been fixed. If so, we need to update all GuC versions across
> >> both i915 and Xe.
> > Right... I guess I'd still like to see if we can test this in GuC (or
> > get confirmation we can't for some reason) before committing something.
> > My worry is we will prevent bug reports like this by working around it
> > and miss critical bugs that need to be fixed in the right component.
> >
> >> [1]
> >> https://gitlab.com/kernel-firmware/linux-firmware/-/blob/main/i915/mtl_guc_70.bin?ref_type=heads
> >>
> >>>> Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has
> >>>> no
> >>> Is there a reason we don't just re-use the main xe_devcoredump()?
> >>>
> >> This is my suggestion: the main devcoredump infrastructure is job-
> >> based,
> >> so it cannot be used for hangs that are not associated with a job.
> >>
> >> In my opinion, this is a gap on our end. Introducing something like
> >> `xe_devcoredump_gt()`, which can be used for non-job-based hangs
> >> (e.g.,
> >> TLB invalidation timeouts like those addressed in this series, or
> >> more
> >> generally any GuC protocol hang), makes sense to me.
> > Ok makes sense. We can do that here. It would be nice to have a more
> > inclusive implementation that lets us call this from anywhere so we
> > aren't duplicating things around for different use cases. But not a
> > blocker here.
> >
> >> I haven't looked at the patch yet, but at a high level, adding
> >> `xe_devcoredump_gt()` seems like a reasonable approach.
> >>
> >>>> exec queue or job to blame - leaves a devcoredump with the GuC
> >>>> log
> >>>> and
> >>>> CT state behind (Matt suggested capturing devcoredumps when we
> >>>> discussed the issue; devcoredumps from both machines are attached
> >>>> to
> >>>> the issue above).
> >>>>
> >>>> Patch 2 logs when the ack for a timed out invalidation finally
> >>>> arrives. This is what established that the acks are late rather
> >>>> than
> >>>> lost.
> >>>>
> >>>> Patch 3 is the RFC part: a delayed work that pokes the GuC
> >>>> (status
> >>>> register read, CT flush, doorbell ring) every 250ms while an ack
> >>>> is
> >>>> overdue. On my machines this converts the guaranteed 2.3s stall
> >>>> into
> >>> I'm a little worried we're just papering over something here that
> >>> needs
> >>> to be addressed in GuC, particularly around GT going to sleep or
> >>> something around the time we're expecting a response, so the pings
> >>> on
> >>> registers might be prematurely waking things up which is something
> >>> we'd
> >>> want to happen in GuC, not the KMD.
> >>>
> >> In general, I agree with this. We should avoid papering over the
> >> issue
> >> and instead fix it properly in the GuC. That said, this workaround
> >> provides a pretty strong data point, since it appears to get the TLB
> >> invalidation unstuck.
> > So if we hit this issue I guess we're already going to have some
> > performance degredation and the workaround makes that better. I need to
> > look at the implementation, but we could be potentially introducing
> > performance penalties in other areas doing these pings.
> >
> > Again, I'd like to see if we can fix this in the right place before
> > implementing a workaround for it. Hopefully Daniele or Julia can give
> > some direction there.
>
> Are we seeing this on i915 at all? Given that Xe does not officially
> support MTL/ARL and is missing several critical WAs for those platforms,
> the approach so far has been to only update the GuC FW if it is required
> for i915.
> Looking at the GuC release notes, there have been a couple of
> TLB-related fixes after 70.53, but they're both marked as only affecting
> PVC and Xe2+ platforms, so no fixes seem to be available for ARL (or at
> least they're not listed in the release notes).
>
> Daniele
>
> >
> > Thanks,
> > Stuart
> >
> >> Matt
> >>
> >>> Thanks,
> >>> Stuart
> >>>
> >>>> a
> >>>> sub-500ms hiccup for the majority of occurrences; a minority of
> >>>> severe
> >>>> episodes ignore 8-9 consecutive doorbells, which points at the
> >>>> GuC
> >>>> firmware being internally blocked for the whole window. Full data
> >>>> on
> >>>> the issue. I am happy to rework the approach (different delay,
> >>>> tying it to the G2H handler, dropping the status read, etc.) -
> >>>> mainly
> >>>> I would like the firmware side investigated, since no host-side
> >>>> poke
> >>>> can fix the severe cases.
> >>>>
> >>>> Based on drm-tip. Tested for several days on both ARL machines
> >>>> under
> >>>> desktop and VM-heavy workloads.
> >>>>
> >>>> Thanks,
> >>>> Tales
> >>>>
> >>>> Tales A. Mendonça (3):
> >>>> drm/xe: Capture devcoredump on TLB invalidation timeout
> >>>> drm/xe: Log when a timed out TLB invalidation ack finally
> >>>> arrives
> >>>> drm/xe: Kick GuC while TLB invalidation acks are overdue
> >>>>
> >>>> drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
> >>>> drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
> >>>> drivers/gpu/drm/xe/xe_tlb_inval.c | 131
> >>>> +++++++++++++++++++++++-
> >>>> drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
> >>>> 4 files changed, 243 insertions(+), 4 deletions(-)
> >>>>
>
--
Com os cumprimentos,
Tales A. Mendonça
talesam.org
communitybig.org
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-04 23:00 ` Tales A. Mendonça
@ 2026-08-04 23:50 ` Daniele Ceraolo Spurio
2026-08-06 17:36 ` Tales A. Mendonça
0 siblings, 1 reply; 25+ messages in thread
From: Daniele Ceraolo Spurio @ 2026-08-04 23:50 UTC (permalink / raw)
To: Tales A. Mendonça
Cc: Summers, Stuart, Brost, Matthew, intel-xe@lists.freedesktop.org,
dri-devel@lists.freedesktop.org, Vivi, Rodrigo,
thomas.hellstrom@linux.intel.com, Filipchuk, Julia
On 8/4/2026 4:00 PM, Tales A. Mendonça wrote:
>> Are we seeing this on i915?
> I have not tested i915 on the affected machines yet - the
> instrumentation that measured the stalls (late-ack logging, kick
> results) is xe-only, so I have no comparable i915 data. I can boot one
> of the ARL machines with i915 for a few days and watch for TLB
> invalidation timeouts there, if that data helps.
That would definitely help, because if the issue does not happen on i915
it likely means that we're missing a WA or something like that in Xe.
>
> More generally: both machines here reproduce reliably (~1 stall/hour
> on a desktop workload, much more under memory pressure), so I am happy
> to test anything on them - including any GuC build the firmware team
> would like data on.
>
> On Stuart's masking concern: fully agreed, that is why patch 3 is
> marked RFC. Patches 1-2 are pure diagnostics and stand on their own; I
> am fine holding patch 3 until the firmware side has been looked at.
> The data point it adds is that a doorbell ring unblocks the ack in the
> majority of episodes, while the severe ones ignore 8-9 consecutive
> rings - hopefully that narrows where to look inside the GuC.
Just a bit of a terminology update here, to make sure we're on the same
page: we usually refer to the notification you're sending to the GuC as
an H2G interrupt and not a doorbell. I'm making this clarification
because the GuC supports a separate per-context notification mechanism
that is referred to as doorbell and which we currently do not implement
in neither i915 nor Xe.
When receiving the H2G interrupt, the only thing that the GuC does is
look into the CTB and process anything in there; however, you've said
that the contents of the H2G CTB are processed immediately, so the
follow up interrupt should result in the GuC just bailing out and doing
nothing because there is no data to process. It feels like when the
issue occurs something is stuck in HW rather than GuC FW and triggering
the interrupt causes the HW to get unstuck.
I have pushed the latest GuC FW for MTL here in case you want to give it
a go:
https://gitlab.com/dceraolo/drm-firmware/-/blob/f783b931555be057dafc2400b7eb4d445c953fec/i915/mtl_guc_70.72.1.bin
. You can override the GuC firmware used by the driver via the
xe.guc_firmware_path modparam; the path is relative to /lib/firmware/
and the firmware needs to be in initramfs for the driver to find it at
boot. Note that we haven't tested this image on MTL, so it might have
unexpected results.
Also, would you be able to capture the GuC logs when the issue occurs?
The default guc log size is relatively small, so you'd have to capture
right when the issue happens. However, you can make them bigger by
building the kernel with CONFIG_DRM_XE_DEBUG or by simply modifying the
xe_guc_log.h file to pick the bigger size by default. If you go with the
latter, please also set xe.guc_log_level=3 on the command line (this is
automatically added by the kconfig).
Thanks,
Daniele
>
> I will send a v2 addressing Matt's review comments (the
> __xe_devcoredump unification and the fixes on patch 2).
>
> Thanks,
> Tales
>
> Em ter., 4 de ago. de 2026 às 19:08, Daniele Ceraolo Spurio
> <daniele.ceraolospurio@intel.com> escreveu:
>>
>>
>> On 8/4/2026 2:33 PM, Summers, Stuart wrote:
>>> On Tue, 2026-08-04 at 14:27 -0700, Matthew Brost wrote:
>>>> On Tue, Aug 04, 2026 at 03:02:44PM -0600, Summers, Stuart wrote:
>>>>> On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
>>>>>> Hi,
>>>>>>
>>>>>> This series is a follow-up to the TLB invalidation ack stall I
>>>>>> have
>>>>>> been debugging on ARL, tracked in:
>>>>>>
>>>>>> https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
>>>>>>
>>>>>> Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51
>>>>>> and
>>>>>> 7dd1 machines here, plus an independent Arc Pro 130T report on
>>>>>> the
>>>>>> issue above), TLB invalidation acks intermittently stall for
>>>>>> ~2.3s.
>>>>>> The H2G request is consumed from the CTB immediately and the G2H
>>>>>> CTB
>>>>>> is empty the whole time - the firmware simply does not send the
>>>>>> ack
>>>>>> until much later. The fence timeout fires at 2.25s and the ack
>>>>>> lands
>>>>>> tens of ms after it. Userspace blocked on the invalidation
>>>>>> (compositor
>>>>>> buffer unmaps etc.) hitches for the full window.
>>>>> Firstly, thanks for the patch!
>>>>>
>>>>> I haven't looked in to all the details of the sighting you were
>>>>> debugging, but we have had similar issues that were fixed in a
>>>>> later
>>>>> GuC version. I think around 70.60.0? It might be worth trying on
>>>>> something later than that to see if that helps... (+Daniele)
>>>>>
>>>> I think this would require an AR on our end to make a new firmware
>>>> version available.
>>>>
>>>> The upstream repo only has 70.53.0 available for ARL [1] (iirc, ARL
>>>> aliases to MTL for firmware). (+Julia too).
>>>>
>>>> Presumably, the GuC changelogs should indicate whether an issue
>>>> related
>>>> this has been fixed. If so, we need to update all GuC versions across
>>>> both i915 and Xe.
>>> Right... I guess I'd still like to see if we can test this in GuC (or
>>> get confirmation we can't for some reason) before committing something.
>>> My worry is we will prevent bug reports like this by working around it
>>> and miss critical bugs that need to be fixed in the right component.
>>>
>>>> [1]
>>>> https://gitlab.com/kernel-firmware/linux-firmware/-/blob/main/i915/mtl_guc_70.bin?ref_type=heads
>>>>
>>>>>> Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has
>>>>>> no
>>>>> Is there a reason we don't just re-use the main xe_devcoredump()?
>>>>>
>>>> This is my suggestion: the main devcoredump infrastructure is job-
>>>> based,
>>>> so it cannot be used for hangs that are not associated with a job.
>>>>
>>>> In my opinion, this is a gap on our end. Introducing something like
>>>> `xe_devcoredump_gt()`, which can be used for non-job-based hangs
>>>> (e.g.,
>>>> TLB invalidation timeouts like those addressed in this series, or
>>>> more
>>>> generally any GuC protocol hang), makes sense to me.
>>> Ok makes sense. We can do that here. It would be nice to have a more
>>> inclusive implementation that lets us call this from anywhere so we
>>> aren't duplicating things around for different use cases. But not a
>>> blocker here.
>>>
>>>> I haven't looked at the patch yet, but at a high level, adding
>>>> `xe_devcoredump_gt()` seems like a reasonable approach.
>>>>
>>>>>> exec queue or job to blame - leaves a devcoredump with the GuC
>>>>>> log
>>>>>> and
>>>>>> CT state behind (Matt suggested capturing devcoredumps when we
>>>>>> discussed the issue; devcoredumps from both machines are attached
>>>>>> to
>>>>>> the issue above).
>>>>>>
>>>>>> Patch 2 logs when the ack for a timed out invalidation finally
>>>>>> arrives. This is what established that the acks are late rather
>>>>>> than
>>>>>> lost.
>>>>>>
>>>>>> Patch 3 is the RFC part: a delayed work that pokes the GuC
>>>>>> (status
>>>>>> register read, CT flush, doorbell ring) every 250ms while an ack
>>>>>> is
>>>>>> overdue. On my machines this converts the guaranteed 2.3s stall
>>>>>> into
>>>>> I'm a little worried we're just papering over something here that
>>>>> needs
>>>>> to be addressed in GuC, particularly around GT going to sleep or
>>>>> something around the time we're expecting a response, so the pings
>>>>> on
>>>>> registers might be prematurely waking things up which is something
>>>>> we'd
>>>>> want to happen in GuC, not the KMD.
>>>>>
>>>> In general, I agree with this. We should avoid papering over the
>>>> issue
>>>> and instead fix it properly in the GuC. That said, this workaround
>>>> provides a pretty strong data point, since it appears to get the TLB
>>>> invalidation unstuck.
>>> So if we hit this issue I guess we're already going to have some
>>> performance degredation and the workaround makes that better. I need to
>>> look at the implementation, but we could be potentially introducing
>>> performance penalties in other areas doing these pings.
>>>
>>> Again, I'd like to see if we can fix this in the right place before
>>> implementing a workaround for it. Hopefully Daniele or Julia can give
>>> some direction there.
>> Are we seeing this on i915 at all? Given that Xe does not officially
>> support MTL/ARL and is missing several critical WAs for those platforms,
>> the approach so far has been to only update the GuC FW if it is required
>> for i915.
>> Looking at the GuC release notes, there have been a couple of
>> TLB-related fixes after 70.53, but they're both marked as only affecting
>> PVC and Xe2+ platforms, so no fixes seem to be available for ARL (or at
>> least they're not listed in the release notes).
>>
>> Daniele
>>
>>> Thanks,
>>> Stuart
>>>
>>>> Matt
>>>>
>>>>> Thanks,
>>>>> Stuart
>>>>>
>>>>>> a
>>>>>> sub-500ms hiccup for the majority of occurrences; a minority of
>>>>>> severe
>>>>>> episodes ignore 8-9 consecutive doorbells, which points at the
>>>>>> GuC
>>>>>> firmware being internally blocked for the whole window. Full data
>>>>>> on
>>>>>> the issue. I am happy to rework the approach (different delay,
>>>>>> tying it to the G2H handler, dropping the status read, etc.) -
>>>>>> mainly
>>>>>> I would like the firmware side investigated, since no host-side
>>>>>> poke
>>>>>> can fix the severe cases.
>>>>>>
>>>>>> Based on drm-tip. Tested for several days on both ARL machines
>>>>>> under
>>>>>> desktop and VM-heavy workloads.
>>>>>>
>>>>>> Thanks,
>>>>>> Tales
>>>>>>
>>>>>> Tales A. Mendonça (3):
>>>>>> drm/xe: Capture devcoredump on TLB invalidation timeout
>>>>>> drm/xe: Log when a timed out TLB invalidation ack finally
>>>>>> arrives
>>>>>> drm/xe: Kick GuC while TLB invalidation acks are overdue
>>>>>>
>>>>>> drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
>>>>>> drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
>>>>>> drivers/gpu/drm/xe/xe_tlb_inval.c | 131
>>>>>> +++++++++++++++++++++++-
>>>>>> drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
>>>>>> 4 files changed, 243 insertions(+), 4 deletions(-)
>>>>>>
>
^ permalink raw reply [flat|nested] 25+ messages in thread
* ✗ CI.checkpatch: warning for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2)
2026-08-04 2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
` (5 preceding siblings ...)
2026-08-04 21:02 ` Summers, Stuart
@ 2026-08-05 12:32 ` Patchwork
2026-08-05 12:34 ` ✓ CI.KUnit: success " Patchwork
` (2 subsequent siblings)
9 siblings, 0 replies; 25+ messages in thread
From: Patchwork @ 2026-08-05 12:32 UTC (permalink / raw)
To: Tales A. Mendonça; +Cc: intel-xe
== Series Details ==
Series: drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2)
URL : https://patchwork.freedesktop.org/series/171539/
State : warning
== Summary ==
+ KERNEL=/kernel
+ git clone https://gitlab.freedesktop.org/drm/maintainer-tools mt
Cloning into 'mt'...
warning: redirecting to https://gitlab.freedesktop.org/drm/maintainer-tools.git/
+ git -C mt rev-list -n1 origin/master
061140b9bc586ae7f40abc1249c97e1cc72d1b9d
+ cd /kernel
+ git config --global --add safe.directory /kernel
+ git log -n1
commit 34edb36f8f11bba53549e0a81459f317a6fac286
Author: Tales A. Mendonça <talesam@gmail.com>
Date: Mon Aug 3 23:14:41 2026 -0300
drm/xe: Kick GuC while TLB invalidation acks are overdue
On ARL with GuC 70.53.0, TLB invalidation acks intermittently stall:
the H2G request is consumed from the CTB immediately, but the G2H ack
only arrives ~2.3s later, tens of milliseconds after the fence timeout
has fired. During the stall the GPU keeps rendering; only work blocked
on the invalidation (e.g. Wayland compositors performing buffer
unmaps) hitches for the full 2.3s. Observed on three machines so far
(7d51, 7dd1, plus an Arc Pro 130T report), see Link.
Experiments ruled out the obvious suspects:
* GT C6 parking: holding forcewake across the whole GT (C6 residency
pinned at 0ms) still produced 9 timeouts in a row.
* Lost interrupt/CT processing on the host: the G2H CTB is empty at
timeout time; the ack genuinely has not been sent by the firmware.
What does help is poking the GuC while the ack is overdue. Add a
delayed work that fires XE_TLB_INVAL_KICK_DELAY_MS after an
invalidation is issued and, while any ack is pending, re-reads the GuC
status register, flushes the CT fast-path and rings the GuC doorbell
(xe_guc_notify()), re-arming itself until the ack arrives; the
existing TDR still bounds the total wait.
Instrumented results from two ARL machines over several days:
* Without kicks: every stall lasts the full ~2.3s and is reported as
a fence timeout (-ETIME), ~1/hour on a desktop workload.
* With kicks: the majority of stalls resolve 15-276ms after one of
the kicks, e.g.:
TLB invalidation ack after kick: seqno=36405 recv=36405, request-to-ack=528ms, last-kick-to-ack=15ms, kicks=2
* A minority of severe episodes ignore 8-9 consecutive doorbells and
still run to the timeout (worst observed: ack 5996ms after request,
3.7s after the last kick), suggesting the firmware is internally
blocked for the whole window rather than missing a wake event.
This is a workaround, not a fix - the root cause looks like a GuC
firmware issue - but it turns a guaranteed 2.3s stall into a sub-500ms
hiccup for most occurrences, and the "ack after kick" log documents
the firmware behavior for further debugging.
Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
Signed-off-by: Tales A. Mendonça <talesam@gmail.com>
+ /mt/dim checkpatch 4c13a67db118ee0019231dbc1362bcb8115ac660 drm-intel
cc6ee7dbdd9a drm/xe: Capture devcoredump on TLB invalidation timeout
d5341199c3eb drm/xe: Log when a timed out TLB invalidation ack finally arrives
-:24: WARNING:COMMIT_LOG_LONG_LINE: Prefer a maximum 75 chars per line (possible unwrapped commit description?)
#24:
TLB invalidation late ack: seqno=10992 recv=10992, request-to-ack=2314ms, timeout-to-ack=45ms
total: 0 errors, 1 warnings, 0 checks, 57 lines checked
34edb36f8f11 drm/xe: Kick GuC while TLB invalidation acks are overdue
-:38: WARNING:COMMIT_LOG_LONG_LINE: Prefer a maximum 75 chars per line (possible unwrapped commit description?)
#38:
TLB invalidation ack after kick: seqno=36405 recv=36405, request-to-ack=528ms, last-kick-to-ack=15ms, kicks=2
total: 0 errors, 1 warnings, 0 checks, 194 lines checked
^ permalink raw reply [flat|nested] 25+ messages in thread
* ✓ CI.KUnit: success for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2)
2026-08-04 2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
` (6 preceding siblings ...)
2026-08-05 12:32 ` ✗ CI.checkpatch: warning for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2) Patchwork
@ 2026-08-05 12:34 ` Patchwork
2026-08-05 13:11 ` ✗ Xe.CI.BAT: failure " Patchwork
2026-08-05 23:39 ` ✗ Xe.CI.FULL: " Patchwork
9 siblings, 0 replies; 25+ messages in thread
From: Patchwork @ 2026-08-05 12:34 UTC (permalink / raw)
To: Tales A. Mendonça; +Cc: intel-xe
== Series Details ==
Series: drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2)
URL : https://patchwork.freedesktop.org/series/171539/
State : success
== Summary ==
+ trap cleanup EXIT
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/drm/xe/.kunitconfig
[12:32:46] Configuring KUnit Kernel ...
Generating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[12:32:51] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[12:33:22] Starting KUnit Kernel (1/1)...
[12:33:22] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[12:33:22] ================== guc_buf (11 subtests) ===================
[12:33:22] [PASSED] test_smallest
[12:33:22] [PASSED] test_largest
[12:33:22] [PASSED] test_granular
[12:33:22] [PASSED] test_unique
[12:33:22] [PASSED] test_overlap
[12:33:22] [PASSED] test_reusable
[12:33:22] [PASSED] test_too_big
[12:33:22] [PASSED] test_flush
[12:33:22] [PASSED] test_lookup
[12:33:22] [PASSED] test_data
[12:33:22] [PASSED] test_class
[12:33:22] ===================== [PASSED] guc_buf =====================
[12:33:22] =================== guc_dbm (7 subtests) ===================
[12:33:22] [PASSED] test_empty
[12:33:22] [PASSED] test_default
[12:33:22] ======================== test_size ========================
[12:33:22] [PASSED] 4
[12:33:23] [PASSED] 8
[12:33:23] [PASSED] 32
[12:33:23] [PASSED] 256
[12:33:23] ==================== [PASSED] test_size ====================
[12:33:23] ======================= test_reuse ========================
[12:33:23] [PASSED] 4
[12:33:23] [PASSED] 8
[12:33:23] [PASSED] 32
[12:33:23] [PASSED] 256
[12:33:23] =================== [PASSED] test_reuse ====================
[12:33:23] =================== test_range_overlap ====================
[12:33:23] [PASSED] 4
[12:33:23] [PASSED] 8
[12:33:23] [PASSED] 32
[12:33:23] [PASSED] 256
[12:33:23] =============== [PASSED] test_range_overlap ================
[12:33:23] =================== test_range_compact ====================
[12:33:23] [PASSED] 4
[12:33:23] [PASSED] 8
[12:33:23] [PASSED] 32
[12:33:23] [PASSED] 256
[12:33:23] =============== [PASSED] test_range_compact ================
[12:33:23] ==================== test_range_spare =====================
[12:33:23] [PASSED] 4
[12:33:23] [PASSED] 8
[12:33:23] [PASSED] 32
[12:33:23] [PASSED] 256
[12:33:23] ================ [PASSED] test_range_spare =================
[12:33:23] ===================== [PASSED] guc_dbm =====================
[12:33:23] =================== guc_idm (6 subtests) ===================
[12:33:23] [PASSED] bad_init
[12:33:23] [PASSED] no_init
[12:33:23] [PASSED] init_fini
[12:33:23] [PASSED] check_used
[12:33:23] [PASSED] check_quota
[12:33:23] [PASSED] check_all
[12:33:23] ===================== [PASSED] guc_idm =====================
[12:33:23] =============== guc_klv_helpers (9 subtests) ===============
[12:33:23] [PASSED] test_count
[12:33:23] [PASSED] test_encode_u32
[12:33:23] [PASSED] test_encode_u64
[12:33:23] [PASSED] test_encode_string
[12:33:23] [PASSED] test_encode_object_raw
[12:33:23] [PASSED] test_encode_object_klv
[12:33:23] [PASSED] test_encode_object_nested
[12:33:23] [PASSED] test_encode_object_basic
[12:33:23] [PASSED] test_print
[12:33:23] ================= [PASSED] guc_klv_helpers =================
[12:33:23] ================== no_relay (3 subtests) ===================
[12:33:23] [PASSED] xe_drops_guc2pf_if_not_ready
[12:33:23] [PASSED] xe_drops_guc2vf_if_not_ready
[12:33:23] [PASSED] xe_rejects_send_if_not_ready
[12:33:23] ==================== [PASSED] no_relay =====================
[12:33:23] ================== pf_relay (14 subtests) ==================
[12:33:23] [PASSED] pf_rejects_guc2pf_too_short
[12:33:23] [PASSED] pf_rejects_guc2pf_too_long
[12:33:23] [PASSED] pf_rejects_guc2pf_no_payload
[12:33:23] [PASSED] pf_fails_no_payload
[12:33:23] [PASSED] pf_fails_bad_origin
[12:33:23] [PASSED] pf_fails_bad_type
[12:33:23] [PASSED] pf_txn_reports_error
[12:33:23] [PASSED] pf_txn_sends_pf2guc
[12:33:23] [PASSED] pf_sends_pf2guc
[12:33:23] [SKIPPED] pf_loopback_nop (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[12:33:23] [SKIPPED] pf_loopback_echo (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[12:33:23] [SKIPPED] pf_loopback_fail (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[12:33:23] [SKIPPED] pf_loopback_busy (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[12:33:23] [SKIPPED] pf_loopback_retry (requires CONFIG_DRM_XE_DEBUG_SRIOV)
[12:33:23] ==================== [PASSED] pf_relay =====================
[12:33:23] ================== vf_relay (3 subtests) ===================
[12:33:23] [PASSED] vf_rejects_guc2vf_too_short
[12:33:23] [PASSED] vf_rejects_guc2vf_too_long
[12:33:23] [PASSED] vf_rejects_guc2vf_no_payload
[12:33:23] ==================== [PASSED] vf_relay =====================
[12:33:23] ================ pf_gt_config (9 subtests) =================
[12:33:23] [PASSED] fair_contexts_1vf
[12:33:23] [PASSED] fair_doorbells_1vf
[12:33:23] [PASSED] fair_ggtt_1vf
[12:33:23] ====================== fair_vram_1vf ======================
[12:33:23] [PASSED] 3.50 GiB
[12:33:23] [PASSED] 11.5 GiB
[12:33:23] [PASSED] 15.5 GiB
[12:33:23] [PASSED] 31.5 GiB
[12:33:23] [PASSED] 63.5 GiB
[12:33:23] [PASSED] 1.91 GiB
[12:33:23] ================== [PASSED] fair_vram_1vf ==================
[12:33:23] ================ fair_vram_1vf_admin_only =================
[12:33:23] [PASSED] 3.50 GiB
[12:33:23] [PASSED] 11.5 GiB
[12:33:23] [PASSED] 15.5 GiB
[12:33:23] [PASSED] 31.5 GiB
[12:33:23] [PASSED] 63.5 GiB
[12:33:23] [PASSED] 1.91 GiB
[12:33:23] ============ [PASSED] fair_vram_1vf_admin_only =============
[12:33:23] ====================== fair_contexts ======================
[12:33:23] [PASSED] 1 VF
[12:33:23] [PASSED] 2 VFs
[12:33:23] [PASSED] 3 VFs
[12:33:23] [PASSED] 4 VFs
[12:33:23] [PASSED] 5 VFs
[12:33:23] [PASSED] 6 VFs
[12:33:23] [PASSED] 7 VFs
[12:33:23] [PASSED] 8 VFs
[12:33:23] [PASSED] 9 VFs
[12:33:23] [PASSED] 10 VFs
[12:33:23] [PASSED] 11 VFs
[12:33:23] [PASSED] 12 VFs
[12:33:23] [PASSED] 13 VFs
[12:33:23] [PASSED] 14 VFs
[12:33:23] [PASSED] 15 VFs
[12:33:23] [PASSED] 16 VFs
[12:33:23] [PASSED] 17 VFs
[12:33:23] [PASSED] 18 VFs
[12:33:23] [PASSED] 19 VFs
[12:33:23] [PASSED] 20 VFs
[12:33:23] [PASSED] 21 VFs
[12:33:23] [PASSED] 22 VFs
[12:33:23] [PASSED] 23 VFs
[12:33:23] [PASSED] 24 VFs
[12:33:23] [PASSED] 25 VFs
[12:33:23] [PASSED] 26 VFs
[12:33:23] [PASSED] 27 VFs
[12:33:23] [PASSED] 28 VFs
[12:33:23] [PASSED] 29 VFs
[12:33:23] [PASSED] 30 VFs
[12:33:23] [PASSED] 31 VFs
[12:33:23] [PASSED] 32 VFs
[12:33:23] [PASSED] 33 VFs
[12:33:23] [PASSED] 34 VFs
[12:33:23] [PASSED] 35 VFs
[12:33:23] [PASSED] 36 VFs
[12:33:23] [PASSED] 37 VFs
[12:33:23] [PASSED] 38 VFs
[12:33:23] [PASSED] 39 VFs
[12:33:23] [PASSED] 40 VFs
[12:33:23] [PASSED] 41 VFs
[12:33:23] [PASSED] 42 VFs
[12:33:23] [PASSED] 43 VFs
[12:33:23] [PASSED] 44 VFs
[12:33:23] [PASSED] 45 VFs
[12:33:23] [PASSED] 46 VFs
[12:33:23] [PASSED] 47 VFs
[12:33:23] [PASSED] 48 VFs
[12:33:23] [PASSED] 49 VFs
[12:33:23] [PASSED] 50 VFs
[12:33:23] [PASSED] 51 VFs
[12:33:23] [PASSED] 52 VFs
[12:33:23] [PASSED] 53 VFs
[12:33:23] [PASSED] 54 VFs
[12:33:23] [PASSED] 55 VFs
[12:33:23] [PASSED] 56 VFs
[12:33:23] [PASSED] 57 VFs
[12:33:23] [PASSED] 58 VFs
[12:33:23] [PASSED] 59 VFs
[12:33:23] [PASSED] 60 VFs
[12:33:23] [PASSED] 61 VFs
[12:33:23] [PASSED] 62 VFs
[12:33:23] [PASSED] 63 VFs
[12:33:23] ================== [PASSED] fair_contexts ==================
[12:33:23] ===================== fair_doorbells ======================
[12:33:23] [PASSED] 1 VF
[12:33:23] [PASSED] 2 VFs
[12:33:23] [PASSED] 3 VFs
[12:33:23] [PASSED] 4 VFs
[12:33:23] [PASSED] 5 VFs
[12:33:23] [PASSED] 6 VFs
[12:33:23] [PASSED] 7 VFs
[12:33:23] [PASSED] 8 VFs
[12:33:23] [PASSED] 9 VFs
[12:33:23] [PASSED] 10 VFs
[12:33:23] [PASSED] 11 VFs
[12:33:23] [PASSED] 12 VFs
[12:33:23] [PASSED] 13 VFs
[12:33:23] [PASSED] 14 VFs
[12:33:23] [PASSED] 15 VFs
[12:33:23] [PASSED] 16 VFs
[12:33:23] [PASSED] 17 VFs
[12:33:23] [PASSED] 18 VFs
[12:33:23] [PASSED] 19 VFs
[12:33:23] [PASSED] 20 VFs
[12:33:23] [PASSED] 21 VFs
[12:33:23] [PASSED] 22 VFs
[12:33:23] [PASSED] 23 VFs
[12:33:23] [PASSED] 24 VFs
[12:33:23] [PASSED] 25 VFs
[12:33:23] [PASSED] 26 VFs
[12:33:23] [PASSED] 27 VFs
[12:33:23] [PASSED] 28 VFs
[12:33:23] [PASSED] 29 VFs
[12:33:23] [PASSED] 30 VFs
[12:33:23] [PASSED] 31 VFs
[12:33:23] [PASSED] 32 VFs
[12:33:23] [PASSED] 33 VFs
[12:33:23] [PASSED] 34 VFs
[12:33:23] [PASSED] 35 VFs
[12:33:23] [PASSED] 36 VFs
[12:33:23] [PASSED] 37 VFs
[12:33:23] [PASSED] 38 VFs
[12:33:23] [PASSED] 39 VFs
[12:33:23] [PASSED] 40 VFs
[12:33:23] [PASSED] 41 VFs
[12:33:23] [PASSED] 42 VFs
[12:33:23] [PASSED] 43 VFs
[12:33:23] [PASSED] 44 VFs
[12:33:23] [PASSED] 45 VFs
[12:33:23] [PASSED] 46 VFs
[12:33:23] [PASSED] 47 VFs
[12:33:23] [PASSED] 48 VFs
[12:33:23] [PASSED] 49 VFs
[12:33:23] [PASSED] 50 VFs
[12:33:23] [PASSED] 51 VFs
[12:33:23] [PASSED] 52 VFs
[12:33:23] [PASSED] 53 VFs
[12:33:23] [PASSED] 54 VFs
[12:33:23] [PASSED] 55 VFs
[12:33:23] [PASSED] 56 VFs
[12:33:23] [PASSED] 57 VFs
[12:33:23] [PASSED] 58 VFs
[12:33:23] [PASSED] 59 VFs
[12:33:23] [PASSED] 60 VFs
[12:33:23] [PASSED] 61 VFs
[12:33:23] [PASSED] 62 VFs
[12:33:23] [PASSED] 63 VFs
[12:33:23] ================= [PASSED] fair_doorbells ==================
[12:33:23] ======================== fair_ggtt ========================
[12:33:23] [PASSED] 1 VF
[12:33:23] [PASSED] 2 VFs
[12:33:23] [PASSED] 3 VFs
[12:33:23] [PASSED] 4 VFs
[12:33:23] [PASSED] 5 VFs
[12:33:23] [PASSED] 6 VFs
[12:33:23] [PASSED] 7 VFs
[12:33:23] [PASSED] 8 VFs
[12:33:23] [PASSED] 9 VFs
[12:33:23] [PASSED] 10 VFs
[12:33:23] [PASSED] 11 VFs
[12:33:23] [PASSED] 12 VFs
[12:33:23] [PASSED] 13 VFs
[12:33:23] [PASSED] 14 VFs
[12:33:23] [PASSED] 15 VFs
[12:33:23] [PASSED] 16 VFs
[12:33:23] [PASSED] 17 VFs
[12:33:23] [PASSED] 18 VFs
[12:33:23] [PASSED] 19 VFs
[12:33:23] [PASSED] 20 VFs
[12:33:23] [PASSED] 21 VFs
[12:33:23] [PASSED] 22 VFs
[12:33:23] [PASSED] 23 VFs
[12:33:23] [PASSED] 24 VFs
[12:33:23] [PASSED] 25 VFs
[12:33:23] [PASSED] 26 VFs
[12:33:23] [PASSED] 27 VFs
[12:33:23] [PASSED] 28 VFs
[12:33:23] [PASSED] 29 VFs
[12:33:23] [PASSED] 30 VFs
[12:33:23] [PASSED] 31 VFs
[12:33:23] [PASSED] 32 VFs
[12:33:23] [PASSED] 33 VFs
[12:33:23] [PASSED] 34 VFs
[12:33:23] [PASSED] 35 VFs
[12:33:23] [PASSED] 36 VFs
[12:33:23] [PASSED] 37 VFs
[12:33:23] [PASSED] 38 VFs
[12:33:23] [PASSED] 39 VFs
[12:33:23] [PASSED] 40 VFs
[12:33:23] [PASSED] 41 VFs
[12:33:23] [PASSED] 42 VFs
[12:33:23] [PASSED] 43 VFs
[12:33:23] [PASSED] 44 VFs
[12:33:23] [PASSED] 45 VFs
[12:33:23] [PASSED] 46 VFs
[12:33:23] [PASSED] 47 VFs
[12:33:23] [PASSED] 48 VFs
[12:33:23] [PASSED] 49 VFs
[12:33:23] [PASSED] 50 VFs
[12:33:23] [PASSED] 51 VFs
[12:33:23] [PASSED] 52 VFs
[12:33:23] [PASSED] 53 VFs
[12:33:23] [PASSED] 54 VFs
[12:33:23] [PASSED] 55 VFs
[12:33:23] [PASSED] 56 VFs
[12:33:23] [PASSED] 57 VFs
[12:33:23] [PASSED] 58 VFs
[12:33:23] [PASSED] 59 VFs
[12:33:23] [PASSED] 60 VFs
[12:33:23] [PASSED] 61 VFs
[12:33:23] [PASSED] 62 VFs
[12:33:23] [PASSED] 63 VFs
[12:33:23] ==================== [PASSED] fair_ggtt ====================
[12:33:23] ======================== fair_vram ========================
[12:33:23] [PASSED] 1 VF
[12:33:23] [PASSED] 2 VFs
[12:33:23] [PASSED] 3 VFs
[12:33:23] [PASSED] 4 VFs
[12:33:23] [PASSED] 5 VFs
[12:33:23] [PASSED] 6 VFs
[12:33:23] [PASSED] 7 VFs
[12:33:23] [PASSED] 8 VFs
[12:33:23] [PASSED] 9 VFs
[12:33:23] [PASSED] 10 VFs
[12:33:23] [PASSED] 11 VFs
[12:33:23] [PASSED] 12 VFs
[12:33:23] [PASSED] 13 VFs
[12:33:23] [PASSED] 14 VFs
[12:33:23] [PASSED] 15 VFs
[12:33:23] [PASSED] 16 VFs
[12:33:23] [PASSED] 17 VFs
[12:33:23] [PASSED] 18 VFs
[12:33:23] [PASSED] 19 VFs
[12:33:23] [PASSED] 20 VFs
[12:33:23] [PASSED] 21 VFs
[12:33:23] [PASSED] 22 VFs
[12:33:23] [PASSED] 23 VFs
[12:33:23] [PASSED] 24 VFs
[12:33:23] [PASSED] 25 VFs
[12:33:23] [PASSED] 26 VFs
[12:33:23] [PASSED] 27 VFs
[12:33:23] [PASSED] 28 VFs
[12:33:23] [PASSED] 29 VFs
[12:33:23] [PASSED] 30 VFs
[12:33:23] [PASSED] 31 VFs
[12:33:23] [PASSED] 32 VFs
[12:33:23] [PASSED] 33 VFs
[12:33:23] [PASSED] 34 VFs
[12:33:23] [PASSED] 35 VFs
[12:33:23] [PASSED] 36 VFs
[12:33:23] [PASSED] 37 VFs
[12:33:23] [PASSED] 38 VFs
[12:33:23] [PASSED] 39 VFs
[12:33:23] [PASSED] 40 VFs
[12:33:23] [PASSED] 41 VFs
[12:33:23] [PASSED] 42 VFs
[12:33:23] [PASSED] 43 VFs
[12:33:23] [PASSED] 44 VFs
[12:33:23] [PASSED] 45 VFs
[12:33:23] [PASSED] 46 VFs
[12:33:23] [PASSED] 47 VFs
[12:33:23] [PASSED] 48 VFs
[12:33:23] [PASSED] 49 VFs
[12:33:23] [PASSED] 50 VFs
[12:33:23] [PASSED] 51 VFs
[12:33:23] [PASSED] 52 VFs
[12:33:23] [PASSED] 53 VFs
[12:33:23] [PASSED] 54 VFs
[12:33:23] [PASSED] 55 VFs
[12:33:23] [PASSED] 56 VFs
[12:33:23] [PASSED] 57 VFs
[12:33:23] [PASSED] 58 VFs
[12:33:23] [PASSED] 59 VFs
[12:33:23] [PASSED] 60 VFs
[12:33:23] [PASSED] 61 VFs
[12:33:23] [PASSED] 62 VFs
[12:33:23] [PASSED] 63 VFs
[12:33:23] ==================== [PASSED] fair_vram ====================
[12:33:23] ================== [PASSED] pf_gt_config ===================
[12:33:23] ===================== lmtt (1 subtest) =====================
[12:33:23] ======================== test_ops =========================
[12:33:23] [PASSED] 2-level
[12:33:23] [PASSED] multi-level
[12:33:23] ==================== [PASSED] test_ops =====================
[12:33:23] ====================== [PASSED] lmtt =======================
[12:33:23] ================= sriov_packet (1 subtest) =================
[12:33:23] [PASSED] test_descriptor_init
[12:33:23] ================== [PASSED] sriov_packet ===================
[12:33:23] ================= pf_service (11 subtests) =================
[12:33:23] [PASSED] pf_negotiate_any
[12:33:23] [PASSED] pf_negotiate_base_match
[12:33:23] [PASSED] pf_negotiate_base_newer
[12:33:23] [PASSED] pf_negotiate_base_next
[12:33:23] [SKIPPED] pf_negotiate_base_older (no older minor)
[12:33:23] [PASSED] pf_negotiate_base_prev
[12:33:23] [PASSED] pf_negotiate_latest_match
[12:33:23] [PASSED] pf_negotiate_latest_newer
[12:33:23] [PASSED] pf_negotiate_latest_next
[12:33:23] [SKIPPED] pf_negotiate_latest_older (no older minor)
[12:33:23] [SKIPPED] pf_negotiate_latest_prev (no prev major)
[12:33:23] =================== [PASSED] pf_service ====================
[12:33:23] ================= xe_guc_g2g (2 subtests) ==================
[12:33:23] ============== xe_live_guc_g2g_kunit_default ==============
[12:33:23] ========= [SKIPPED] xe_live_guc_g2g_kunit_default ==========
[12:33:23] ============== xe_live_guc_g2g_kunit_allmem ===============
[12:33:23] ========== [SKIPPED] xe_live_guc_g2g_kunit_allmem ==========
[12:33:23] =================== [SKIPPED] xe_guc_g2g ===================
[12:33:23] =================== xe_mocs (2 subtests) ===================
[12:33:23] ================ xe_live_mocs_kernel_kunit ================
[12:33:23] =========== [SKIPPED] xe_live_mocs_kernel_kunit ============
[12:33:23] ================ xe_live_mocs_reset_kunit =================
[12:33:23] ============ [SKIPPED] xe_live_mocs_reset_kunit ============
[12:33:23] ==================== [SKIPPED] xe_mocs =====================
[12:33:23] ================= xe_migrate (2 subtests) ==================
[12:33:23] ================= xe_migrate_sanity_kunit =================
[12:33:23] ============ [SKIPPED] xe_migrate_sanity_kunit =============
[12:33:23] ================== xe_validate_ccs_kunit ==================
[12:33:23] ============= [SKIPPED] xe_validate_ccs_kunit ==============
[12:33:23] =================== [SKIPPED] xe_migrate ===================
[12:33:23] ================== xe_dma_buf (1 subtest) ==================
[12:33:23] ==================== xe_dma_buf_kunit =====================
[12:33:23] ================ [SKIPPED] xe_dma_buf_kunit ================
[12:33:23] =================== [SKIPPED] xe_dma_buf ===================
[12:33:23] ================= xe_bo_shrink (1 subtest) =================
[12:33:23] =================== xe_bo_shrink_kunit ====================
[12:33:23] =============== [SKIPPED] xe_bo_shrink_kunit ===============
[12:33:23] ================== [SKIPPED] xe_bo_shrink ==================
[12:33:23] ==================== xe_bo (2 subtests) ====================
[12:33:23] ================== xe_ccs_migrate_kunit ===================
[12:33:23] ============== [SKIPPED] xe_ccs_migrate_kunit ==============
[12:33:23] ==================== xe_bo_evict_kunit ====================
[12:33:23] =============== [SKIPPED] xe_bo_evict_kunit ================
[12:33:23] ===================== [SKIPPED] xe_bo ======================
[12:33:23] ==================== args (13 subtests) ====================
[12:33:23] [PASSED] count_args_test
[12:33:23] [PASSED] call_args_example
[12:33:23] [PASSED] call_args_test
[12:33:23] [PASSED] drop_first_arg_example
[12:33:23] [PASSED] drop_first_arg_test
[12:33:23] [PASSED] first_arg_example
[12:33:23] [PASSED] first_arg_test
[12:33:23] [PASSED] last_arg_example
[12:33:23] [PASSED] last_arg_test
[12:33:23] [PASSED] pick_arg_example
[12:33:23] [PASSED] if_args_example
[12:33:23] [PASSED] if_args_test
[12:33:23] [PASSED] sep_comma_example
[12:33:23] ====================== [PASSED] args =======================
[12:33:23] =================== xe_pci (3 subtests) ====================
[12:33:23] ==================== check_graphics_ip ====================
[12:33:23] [PASSED] 12.00 Xe_LP
[12:33:23] [PASSED] 12.10 Xe_LP+
[12:33:23] [PASSED] 12.55 Xe_HPG
[12:33:23] [PASSED] 12.60 Xe_HPC
[12:33:23] [PASSED] 12.70 Xe_LPG
[12:33:23] [PASSED] 12.71 Xe_LPG
[12:33:23] [PASSED] 12.74 Xe_LPG+
[12:33:23] [PASSED] 20.01 Xe2_HPG
[12:33:23] [PASSED] 20.02 Xe2_HPG
[12:33:23] [PASSED] 20.04 Xe2_LPG
[12:33:23] [PASSED] 30.00 Xe3_LPG
[12:33:23] [PASSED] 30.01 Xe3_LPG
[12:33:23] [PASSED] 30.03 Xe3_LPG
[12:33:23] [PASSED] 30.04 Xe3_LPG
[12:33:23] [PASSED] 30.05 Xe3_LPG
[12:33:23] [PASSED] 35.10 Xe3p_LPG
[12:33:23] [PASSED] 35.11 Xe3p_XPC
[12:33:23] ================ [PASSED] check_graphics_ip ================
[12:33:23] ===================== check_media_ip ======================
[12:33:23] [PASSED] 12.00 Xe_M
[12:33:23] [PASSED] 12.55 Xe_HPM
[12:33:23] [PASSED] 13.00 Xe_LPM+
[12:33:23] [PASSED] 13.01 Xe2_HPM
[12:33:23] [PASSED] 20.00 Xe2_LPM
[12:33:23] [PASSED] 30.00 Xe3_LPM
[12:33:23] [PASSED] 30.02 Xe3_LPM
[12:33:23] [PASSED] 35.00 Xe3p_LPM
[12:33:23] [PASSED] 35.03 Xe3p_HPM
[12:33:23] ================= [PASSED] check_media_ip ==================
[12:33:23] =================== check_platform_desc ===================
[12:33:23] [PASSED] 0x9A60 (TIGERLAKE)
[12:33:23] [PASSED] 0x9A68 (TIGERLAKE)
[12:33:23] [PASSED] 0x9A70 (TIGERLAKE)
[12:33:23] [PASSED] 0x9A40 (TIGERLAKE)
[12:33:23] [PASSED] 0x9A49 (TIGERLAKE)
[12:33:23] [PASSED] 0x9A59 (TIGERLAKE)
[12:33:23] [PASSED] 0x9A78 (TIGERLAKE)
[12:33:23] [PASSED] 0x9AC0 (TIGERLAKE)
[12:33:23] [PASSED] 0x9AC9 (TIGERLAKE)
[12:33:23] [PASSED] 0x9AD9 (TIGERLAKE)
[12:33:23] [PASSED] 0x9AF8 (TIGERLAKE)
[12:33:23] [PASSED] 0x4C80 (ROCKETLAKE)
[12:33:23] [PASSED] 0x4C8A (ROCKETLAKE)
[12:33:23] [PASSED] 0x4C8B (ROCKETLAKE)
[12:33:23] [PASSED] 0x4C8C (ROCKETLAKE)
[12:33:23] [PASSED] 0x4C90 (ROCKETLAKE)
[12:33:23] [PASSED] 0x4C9A (ROCKETLAKE)
[12:33:23] [PASSED] 0x4680 (ALDERLAKE_S)
[12:33:23] [PASSED] 0x4682 (ALDERLAKE_S)
[12:33:23] [PASSED] 0x4688 (ALDERLAKE_S)
[12:33:23] [PASSED] 0x468A (ALDERLAKE_S)
[12:33:23] [PASSED] 0x468B (ALDERLAKE_S)
[12:33:23] [PASSED] 0x4690 (ALDERLAKE_S)
[12:33:23] [PASSED] 0x4692 (ALDERLAKE_S)
[12:33:23] [PASSED] 0x4693 (ALDERLAKE_S)
[12:33:23] [PASSED] 0x46A0 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46A1 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46A2 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46A3 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46A6 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46A8 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46AA (ALDERLAKE_P)
[12:33:23] [PASSED] 0x462A (ALDERLAKE_P)
[12:33:23] [PASSED] 0x4626 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x4628 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46B0 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46B1 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46B2 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46B3 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46C0 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46C1 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46C2 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46C3 (ALDERLAKE_P)
[12:33:23] [PASSED] 0x46D0 (ALDERLAKE_N)
[12:33:23] [PASSED] 0x46D1 (ALDERLAKE_N)
[12:33:23] [PASSED] 0x46D2 (ALDERLAKE_N)
[12:33:23] [PASSED] 0x46D3 (ALDERLAKE_N)
[12:33:23] [PASSED] 0x46D4 (ALDERLAKE_N)
[12:33:23] [PASSED] 0xA721 (ALDERLAKE_P)
[12:33:23] [PASSED] 0xA7A1 (ALDERLAKE_P)
[12:33:23] [PASSED] 0xA7A9 (ALDERLAKE_P)
[12:33:23] [PASSED] 0xA7AC (ALDERLAKE_P)
[12:33:23] [PASSED] 0xA7AD (ALDERLAKE_P)
[12:33:23] [PASSED] 0xA720 (ALDERLAKE_P)
[12:33:23] [PASSED] 0xA7A0 (ALDERLAKE_P)
[12:33:23] [PASSED] 0xA7A8 (ALDERLAKE_P)
[12:33:23] [PASSED] 0xA7AA (ALDERLAKE_P)
[12:33:23] [PASSED] 0xA7AB (ALDERLAKE_P)
[12:33:23] [PASSED] 0xA780 (ALDERLAKE_S)
[12:33:23] [PASSED] 0xA781 (ALDERLAKE_S)
[12:33:23] [PASSED] 0xA782 (ALDERLAKE_S)
[12:33:23] [PASSED] 0xA783 (ALDERLAKE_S)
[12:33:23] [PASSED] 0xA788 (ALDERLAKE_S)
[12:33:23] [PASSED] 0xA789 (ALDERLAKE_S)
[12:33:23] [PASSED] 0xA78A (ALDERLAKE_S)
[12:33:23] [PASSED] 0xA78B (ALDERLAKE_S)
[12:33:23] [PASSED] 0x4905 (DG1)
[12:33:23] [PASSED] 0x4906 (DG1)
[12:33:23] [PASSED] 0x4907 (DG1)
[12:33:23] [PASSED] 0x4908 (DG1)
[12:33:23] [PASSED] 0x4909 (DG1)
[12:33:23] [PASSED] 0x56C0 (DG2)
[12:33:23] [PASSED] 0x56C2 (DG2)
[12:33:23] [PASSED] 0x56C1 (DG2)
[12:33:23] [PASSED] 0x7D51 (METEORLAKE)
[12:33:23] [PASSED] 0x7DD1 (METEORLAKE)
[12:33:23] [PASSED] 0x7D41 (METEORLAKE)
[12:33:23] [PASSED] 0x7D67 (METEORLAKE)
[12:33:23] [PASSED] 0xB640 (METEORLAKE)
[12:33:23] [PASSED] 0x56A0 (DG2)
[12:33:23] [PASSED] 0x56A1 (DG2)
[12:33:23] [PASSED] 0x56A2 (DG2)
[12:33:23] [PASSED] 0x56BE (DG2)
[12:33:23] [PASSED] 0x56BF (DG2)
[12:33:23] [PASSED] 0x5690 (DG2)
[12:33:23] [PASSED] 0x5691 (DG2)
[12:33:23] [PASSED] 0x5692 (DG2)
[12:33:23] [PASSED] 0x56A5 (DG2)
[12:33:23] [PASSED] 0x56A6 (DG2)
[12:33:23] [PASSED] 0x56B0 (DG2)
[12:33:23] [PASSED] 0x56B1 (DG2)
[12:33:23] [PASSED] 0x56BA (DG2)
[12:33:23] [PASSED] 0x56BB (DG2)
[12:33:23] [PASSED] 0x56BC (DG2)
[12:33:23] [PASSED] 0x56BD (DG2)
[12:33:23] [PASSED] 0x5693 (DG2)
[12:33:23] [PASSED] 0x5694 (DG2)
[12:33:23] [PASSED] 0x5695 (DG2)
[12:33:23] [PASSED] 0x56A3 (DG2)
[12:33:23] [PASSED] 0x56A4 (DG2)
[12:33:23] [PASSED] 0x56B2 (DG2)
[12:33:23] [PASSED] 0x56B3 (DG2)
[12:33:23] [PASSED] 0x5696 (DG2)
[12:33:23] [PASSED] 0x5697 (DG2)
[12:33:23] [PASSED] 0xB69 (PVC)
[12:33:23] [PASSED] 0xB6E (PVC)
[12:33:23] [PASSED] 0xBD4 (PVC)
[12:33:23] [PASSED] 0xBD5 (PVC)
[12:33:23] [PASSED] 0xBD6 (PVC)
[12:33:23] [PASSED] 0xBD7 (PVC)
[12:33:23] [PASSED] 0xBD8 (PVC)
[12:33:23] [PASSED] 0xBD9 (PVC)
[12:33:23] [PASSED] 0xBDA (PVC)
[12:33:23] [PASSED] 0xBDB (PVC)
[12:33:23] [PASSED] 0xBE0 (PVC)
[12:33:23] [PASSED] 0xBE1 (PVC)
[12:33:23] [PASSED] 0xBE5 (PVC)
[12:33:23] [PASSED] 0x7D40 (METEORLAKE)
[12:33:23] [PASSED] 0x7D45 (METEORLAKE)
[12:33:23] [PASSED] 0x7D55 (METEORLAKE)
[12:33:23] [PASSED] 0x7D60 (METEORLAKE)
[12:33:23] [PASSED] 0x7DD5 (METEORLAKE)
[12:33:23] [PASSED] 0x6420 (LUNARLAKE)
[12:33:23] [PASSED] 0x64A0 (LUNARLAKE)
[12:33:23] [PASSED] 0x64B0 (LUNARLAKE)
[12:33:23] [PASSED] 0xE202 (BATTLEMAGE)
[12:33:23] [PASSED] 0xE209 (BATTLEMAGE)
[12:33:23] [PASSED] 0xE20B (BATTLEMAGE)
[12:33:23] [PASSED] 0xE20C (BATTLEMAGE)
[12:33:23] [PASSED] 0xE20D (BATTLEMAGE)
[12:33:23] [PASSED] 0xE210 (BATTLEMAGE)
[12:33:23] [PASSED] 0xE211 (BATTLEMAGE)
[12:33:23] [PASSED] 0xE212 (BATTLEMAGE)
[12:33:23] [PASSED] 0xE216 (BATTLEMAGE)
[12:33:23] [PASSED] 0xE220 (BATTLEMAGE)
[12:33:23] [PASSED] 0xE221 (BATTLEMAGE)
[12:33:23] [PASSED] 0xE222 (BATTLEMAGE)
[12:33:23] [PASSED] 0xE223 (BATTLEMAGE)
[12:33:23] [PASSED] 0xB080 (PANTHERLAKE)
[12:33:23] [PASSED] 0xB081 (PANTHERLAKE)
[12:33:23] [PASSED] 0xB082 (PANTHERLAKE)
[12:33:23] [PASSED] 0xB083 (PANTHERLAKE)
[12:33:23] [PASSED] 0xB084 (PANTHERLAKE)
[12:33:23] [PASSED] 0xB085 (PANTHERLAKE)
[12:33:23] [PASSED] 0xB086 (PANTHERLAKE)
[12:33:23] [PASSED] 0xB087 (PANTHERLAKE)
[12:33:23] [PASSED] 0xB08F (PANTHERLAKE)
[12:33:23] [PASSED] 0xB090 (PANTHERLAKE)
[12:33:23] [PASSED] 0xB0A0 (PANTHERLAKE)
[12:33:23] [PASSED] 0xB0B0 (PANTHERLAKE)
[12:33:23] [PASSED] 0xFD80 (PANTHERLAKE)
[12:33:23] [PASSED] 0xFD81 (PANTHERLAKE)
[12:33:23] [PASSED] 0xD740 (NOVALAKE_S)
[12:33:23] [PASSED] 0xD741 (NOVALAKE_S)
[12:33:23] [PASSED] 0xD742 (NOVALAKE_S)
[12:33:23] [PASSED] 0xD743 (NOVALAKE_S)
[12:33:23] [PASSED] 0xD745 (NOVALAKE_S)
[12:33:23] [PASSED] 0xD74A (NOVALAKE_S)
[12:33:23] [PASSED] 0xD74B (NOVALAKE_S)
[12:33:23] [PASSED] 0x674C (CRESCENTISLAND)
[12:33:23] [PASSED] 0x674D (CRESCENTISLAND)
[12:33:23] [PASSED] 0x674E (CRESCENTISLAND)
[12:33:23] [PASSED] 0x674F (CRESCENTISLAND)
[12:33:23] [PASSED] 0x6750 (CRESCENTISLAND)
[12:33:23] [PASSED] 0xD750 (NOVALAKE_P)
[12:33:23] [PASSED] 0xD751 (NOVALAKE_P)
[12:33:23] [PASSED] 0xD752 (NOVALAKE_P)
[12:33:23] [PASSED] 0xD753 (NOVALAKE_P)
[12:33:23] [PASSED] 0xD754 (NOVALAKE_P)
[12:33:23] [PASSED] 0xD755 (NOVALAKE_P)
[12:33:23] [PASSED] 0xD756 (NOVALAKE_P)
[12:33:23] [PASSED] 0xD757 (NOVALAKE_P)
[12:33:23] [PASSED] 0xD75F (NOVALAKE_P)
[12:33:23] =============== [PASSED] check_platform_desc ===============
[12:33:23] ===================== [PASSED] xe_pci ======================
[12:33:23] ============= xe_rtp_tables_test (5 subtests) ==============
[12:33:23] ================== xe_rtp_table_gt_test ===================
[12:33:23] [PASSED] gt_was/14011060649
[12:33:23] [PASSED] gt_was/14011059788
[12:33:23] [PASSED] gt_was/14015795083
[12:33:23] [PASSED] gt_was/16021867713
[12:33:23] [PASSED] gt_was/14019449301
[12:33:23] [PASSED] gt_was/16028005424
[12:33:23] [PASSED] gt_was/14026578760
[12:33:23] [PASSED] gt_was/1409420604
[12:33:23] [PASSED] gt_was/1408615072
[12:33:23] [PASSED] gt_was/22010523718
[12:33:23] [PASSED] gt_was/14011006942
[12:33:23] [PASSED] gt_was/14014830051
[12:33:23] [PASSED] gt_was/18018781329
[12:33:23] [PASSED] gt_was/1509235366
[12:33:23] [PASSED] gt_was/18018781329
[12:33:23] [PASSED] gt_was/16016694945
[12:33:23] [PASSED] gt_was/14018575942
[12:33:23] [PASSED] gt_was/22016670082
[12:33:23] [PASSED] gt_was/22016670082
[12:33:23] [PASSED] gt_was/14017421178
[12:33:23] [PASSED] gt_was/16025250150
[12:33:23] [PASSED] gt_was/14021871409
[12:33:23] [PASSED] gt_was/16021865536
[12:33:23] [PASSED] gt_was/14021486841
[12:33:23] [PASSED] gt_was/14025160223
[12:33:23] [PASSED] gt_was/14026144927, 16029437861, 14026127056
[12:33:23] [PASSED] gt_was/14025635424
[12:33:23] [PASSED] gt_was/16028005424
[12:33:23] ============== [PASSED] xe_rtp_table_gt_test ===============
[12:33:23] ================== xe_rtp_table_gt_test ===================
[12:33:23] [PASSED] gt_tunings/Tuning: Blend Fill Caching Optimization Disable
[12:33:23] [PASSED] gt_tunings/Tuning: 32B Access Enable
[12:33:23] [PASSED] gt_tunings/Tuning: L3 cache
[12:33:23] [PASSED] gt_tunings/Tuning: L3 cache - media
[12:33:23] [PASSED] gt_tunings/Tuning: Compression Overfetch
[12:33:23] [PASSED] gt_tunings/Tuning: Compression Overfetch - media
[12:33:23] [PASSED] gt_tunings/Tuning: Enable compressible partial write overfetch in L3
[12:33:23] [PASSED] gt_tunings/Tuning: Enable compressible partial write overfetch in L3 - media
[12:33:23] [PASSED] gt_tunings/Tuning: L2 Overfetch Compressible Only
[12:33:23] [PASSED] gt_tunings/Tuning: L2 Overfetch Compressible Only - media
[12:33:23] [PASSED] gt_tunings/Tuning: Stateless compression control
[12:33:23] [PASSED] gt_tunings/Tuning: Stateless compression control - media
[12:33:23] [PASSED] gt_tunings/Tuning: L3 RW flush all Cache
[12:33:23] [PASSED] gt_tunings/Tuning: L3 RW flush all cache - media
[12:33:23] [PASSED] gt_tunings/Tuning: Set STLB Bank Hash Mode to 4KB
[12:33:23] ============== [PASSED] xe_rtp_table_gt_test ===============
[12:33:23] ================== xe_rtp_table_oob_test ==================
[12:33:23] [PASSED] oob_was/1607983814
[12:33:23] [PASSED] oob_was/16010904313
[12:33:23] [PASSED] oob_was/18022495364
[12:33:23] [PASSED] oob_was/22012773006
[12:33:23] [PASSED] oob_was/14014475959
[12:33:23] [PASSED] oob_was/22011391025
[12:33:23] [PASSED] oob_was/22012727170
[12:33:23] [PASSED] oob_was/22012727685
[12:33:23] [PASSED] oob_was/22016596838
[12:33:23] [PASSED] oob_was/18020744125
[12:33:23] [PASSED] oob_was/1409600907
[12:33:23] [PASSED] oob_was/22014953428
[12:33:23] [PASSED] oob_was/16017236439
[12:33:23] [PASSED] oob_was/14019821291
[12:33:23] [PASSED] oob_was/14015076503
[12:33:23] [PASSED] oob_was/14018913170
[12:33:23] [PASSED] oob_was/14018094691
[12:33:23] [PASSED] oob_was/18024947630
[12:33:23] [PASSED] oob_was/16022287689
[12:33:23] [PASSED] oob_was/13011645652
[12:33:23] [PASSED] oob_was/14022293748
[12:33:23] [PASSED] oob_was/22019794406
[12:33:23] [PASSED] oob_was/22019338487
[12:33:23] [PASSED] oob_was/16023588340
[12:33:23] [PASSED] oob_was/14019789679
[12:33:23] [PASSED] oob_was/14022866841
[12:33:23] [PASSED] oob_was/16021333562
[12:33:23] [PASSED] oob_was/14016712196
[12:33:23] [PASSED] oob_was/14015568240
[12:33:23] [PASSED] oob_was/18013179988
[12:33:23] [PASSED] oob_was/1508761755
[12:33:23] [PASSED] oob_was/16023105232
[12:33:23] [PASSED] oob_was/16026508708
[12:33:23] [PASSED] oob_was/14020001231
[12:33:23] [PASSED] oob_was/16023683509
[12:33:23] [PASSED] oob_was/14025515070
[12:33:23] [PASSED] oob_was/15015404425_disable
[12:33:23] [PASSED] oob_was/16026007364
[12:33:23] [PASSED] oob_was/14020316580
[12:33:23] [PASSED] oob_was/14025883347
[12:33:23] [PASSED] oob_was/16029380221
[12:33:23] [PASSED] oob_was/22022079272
[12:33:23] [PASSED] oob_was/16029897822
[12:33:23] [PASSED] oob_was/14027054324
[12:33:23] ============== [PASSED] xe_rtp_table_oob_test ==============
[12:33:23] ================ xe_rtp_table_dev_oob_test ================
[12:33:23] [PASSED] device_oob_was/22010954014
[12:33:23] [PASSED] device_oob_was/15015404425
[12:33:23] [PASSED] device_oob_was/22019338487_display
[12:33:23] [PASSED] device_oob_was/14022085890
[12:33:23] [PASSED] device_oob_was/14026539277
[12:33:23] [PASSED] device_oob_was/14026633728
[12:33:23] [PASSED] device_oob_was/14026746987
[12:33:23] [PASSED] device_oob_was/14026779378
[12:33:23] ============ [PASSED] xe_rtp_table_dev_oob_test ============
[12:33:23] ========== xe_rtp_table_missing_upper_bound_test ==========
[12:33:23] [PASSED] register_whitelist/WaAllowPMDepthAndInvocationCountAccessFromUMD, 1408556865
[12:33:23] [PASSED] register_whitelist/1508744258, 14012131227, 1808121037
[12:33:23] [PASSED] register_whitelist/1806527549
[12:33:23] [PASSED] register_whitelist/allow_read_ctx_timestamp
[12:33:23] [PASSED] register_whitelist/allow_read_queue_timestamp
[12:33:23] [PASSED] register_whitelist/16014440446
[12:33:23] [PASSED] register_whitelist/16017236439
[12:33:23] [PASSED] register_whitelist/16020183090
[12:33:23] [PASSED] register_whitelist/14024997852
[12:33:23] [PASSED] register_whitelist/14024997852
[12:33:23] ====== [PASSED] xe_rtp_table_missing_upper_bound_test ======
[12:33:23] =============== [PASSED] xe_rtp_tables_test ================
[12:33:23] =================== xe_rtp (3 subtests) ====================
[12:33:23] =================== xe_rtp_rules_tests ====================
[12:33:23] [PASSED] no
[12:33:23] [PASSED] yes
[12:33:23] [PASSED] no-and-no
[12:33:23] [PASSED] no-and-yes
[12:33:23] [PASSED] yes-and-no
[12:33:23] [PASSED] yes-and-yes
[12:33:23] [PASSED] no-or-no
[12:33:23] [PASSED] no-or-yes
[12:33:23] [PASSED] yes-or-no
[12:33:23] [PASSED] yes-or-yes
[12:33:23] [PASSED] no-yes-or-yes-no
[12:33:23] [PASSED] no-yes-or-yes-yes
[12:33:23] [PASSED] yes-yes-or-no-yes
[12:33:23] [PASSED] yes-yes-or-yes-yes
[12:33:23] [PASSED] no-no-or-yes-or-no
[12:33:23] [PASSED] or
[12:33:23] [PASSED] or-yes
[12:33:23] [PASSED] or-no
[12:33:23] [PASSED] yes-or
[12:33:23] [PASSED] no-or
[12:33:23] [PASSED] no-or-or-yes
[12:33:23] [PASSED] yes-or-or-no
[12:33:23] [PASSED] no-or-or-no
[12:33:23] [PASSED] missing-context-engine-class
[12:33:23] [PASSED] missing-context-engine-class-or-yes
[12:33:23] [PASSED] missing-context-engine-class-or-or-yes
[12:33:23] =============== [PASSED] xe_rtp_rules_tests ================
[12:33:23] =============== xe_rtp_process_to_sr_tests ================
[12:33:23] [PASSED] coalesce-same-reg
[12:33:23] [PASSED] coalesce-same-reg-literal-and-func
[12:33:23] [PASSED] no-match-no-add
[12:33:23] [PASSED] two-regs-two-entries
[12:33:23] [PASSED] clr-one-set-other
[12:33:23] [PASSED] set-field
[12:33:23] [PASSED] conflict-duplicate
[12:33:23] [PASSED] conflict-not-disjoint
[12:33:23] [PASSED] conflict-not-disjoint-literal-and-func
[12:33:23] [PASSED] conflict-reg-type
[12:33:23] [PASSED] bad-mcr-reg-forced-to-regular
[12:33:23] [PASSED] bad-regular-reg-forced-to-mcr
[12:33:23] =========== [PASSED] xe_rtp_process_to_sr_tests ============
[12:33:23] ================== xe_rtp_process_tests ===================
[12:33:23] [PASSED] active1
[12:33:23] [PASSED] active2
[12:33:23] [PASSED] active-inactive
[12:33:23] [PASSED] inactive-active
[12:33:23] [PASSED] inactive-active-inactive
[12:33:23] [PASSED] inactive-inactive-inactive
[12:33:23] ============== [PASSED] xe_rtp_process_tests ===============
[12:33:23] ===================== [PASSED] xe_rtp ======================
[12:33:23] ==================== xe_wa (1 subtest) =====================
[12:33:23] ======================== xe_wa_gt =========================
[12:33:23] [PASSED] TIGERLAKE B0
[12:33:23] [PASSED] DG1 A0
[12:33:23] [PASSED] DG1 B0
[12:33:23] [PASSED] ALDERLAKE_S A0
[12:33:23] [PASSED] ALDERLAKE_S B0
[12:33:23] [PASSED] ALDERLAKE_S C0
[12:33:23] [PASSED] ALDERLAKE_S D0
[12:33:23] [PASSED] ALDERLAKE_P A0
[12:33:23] [PASSED] ALDERLAKE_P B0
[12:33:23] [PASSED] ALDERLAKE_P C0
[12:33:23] [PASSED] ALDERLAKE_S RPLS D0
[12:33:23] [PASSED] ALDERLAKE_P RPLU E0
[12:33:23] [PASSED] DG2 G10 C0
[12:33:23] [PASSED] DG2 G11 B1
[12:33:23] [PASSED] DG2 G12 A1
[12:33:23] [PASSED] METEORLAKE 12.70(Xe_LPG) A0 13.00(Xe_LPM+) A0
[12:33:23] [PASSED] METEORLAKE 12.71(Xe_LPG) A0 13.00(Xe_LPM+) A0
[12:33:23] [PASSED] METEORLAKE 12.74(Xe_LPG+) A0 13.00(Xe_LPM+) A0
[12:33:23] [PASSED] LUNARLAKE 20.04(Xe2_LPG) A0 20.00(Xe2_LPM) A0
[12:33:23] [PASSED] LUNARLAKE 20.04(Xe2_LPG) B0 20.00(Xe2_LPM) A0
[12:33:23] [PASSED] BATTLEMAGE 20.01(Xe2_HPG) A0 13.01(Xe2_HPM) A1
[12:33:23] [PASSED] PANTHERLAKE 30.00(Xe3_LPG) A0 30.00(Xe3_LPM) A0
[12:33:23] ==================== [PASSED] xe_wa_gt =====================
[12:33:23] ====================== [PASSED] xe_wa ======================
[12:33:23] ============================================================
[12:33:23] Testing complete. Ran 742 tests: passed: 724, skipped: 18
[12:33:23] Elapsed time: 36.657s total, 4.413s configuring, 31.578s building, 0.624s running
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/drm/tests/.kunitconfig
[12:33:23] Configuring KUnit Kernel ...
Regenerating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[12:33:25] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[12:33:49] Starting KUnit Kernel (1/1)...
[12:33:49] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[12:33:50] ============ drm_test_pick_cmdline (2 subtests) ============
[12:33:50] [PASSED] drm_test_pick_cmdline_res_1920_1080_60
[12:33:50] =============== drm_test_pick_cmdline_named ===============
[12:33:50] [PASSED] NTSC
[12:33:50] [PASSED] NTSC-J
[12:33:50] [PASSED] PAL
[12:33:50] [PASSED] PAL-M
[12:33:50] =========== [PASSED] drm_test_pick_cmdline_named ===========
[12:33:50] ============== [PASSED] drm_test_pick_cmdline ==============
[12:33:50] == drm_test_atomic_get_connector_for_encoder (1 subtest) ===
[12:33:50] [PASSED] drm_test_drm_atomic_get_connector_for_encoder
[12:33:50] ==== [PASSED] drm_test_atomic_get_connector_for_encoder ====
[12:33:50] =========== drm_validate_clone_mode (2 subtests) ===========
[12:33:50] ============== drm_test_check_in_clone_mode ===============
[12:33:50] [PASSED] in_clone_mode
[12:33:50] [PASSED] not_in_clone_mode
[12:33:50] ========== [PASSED] drm_test_check_in_clone_mode ===========
[12:33:50] =============== drm_test_check_valid_clones ===============
[12:33:50] [PASSED] not_in_clone_mode
[12:33:50] [PASSED] valid_clone
[12:33:50] [PASSED] invalid_clone
[12:33:50] =========== [PASSED] drm_test_check_valid_clones ===========
[12:33:50] ============= [PASSED] drm_validate_clone_mode =============
[12:33:50] ============= drm_validate_modeset (1 subtest) =============
[12:33:50] [PASSED] drm_test_check_connector_changed_modeset
[12:33:50] ============== [PASSED] drm_validate_modeset ===============
[12:33:50] ====== drm_test_bridge_get_current_state (1 subtest) =======
[12:33:50] [PASSED] drm_test_drm_bridge_get_current_state_atomic
[12:33:50] ======== [PASSED] drm_test_bridge_get_current_state ========
[12:33:50] ====== drm_test_bridge_helper_reset_crtc (3 subtests) ======
[12:33:50] [PASSED] drm_test_drm_bridge_helper_reset_crtc_atomic
[12:33:50] [PASSED] drm_test_drm_bridge_helper_reset_crtc_atomic_disabled
[12:33:50] [PASSED] drm_test_drm_bridge_helper_hdmi_output_bus_fmts
[12:33:50] ======== [PASSED] drm_test_bridge_helper_reset_crtc ========
[12:33:50] ============== drm_bridge_alloc (2 subtests) ===============
[12:33:50] [PASSED] drm_test_drm_bridge_alloc_basic
[12:33:50] [PASSED] drm_test_drm_bridge_alloc_get_put
[12:33:50] ================ [PASSED] drm_bridge_alloc =================
[12:33:50] ============= drm_bridge_bus_fmt (5 subtests) ==============
[12:33:50] [PASSED] drm_test_bridge_rgb_yuv_rgb
[12:33:50] [PASSED] drm_test_bridge_must_convert_to_yuv444
[12:33:50] [PASSED] drm_test_bridge_hdmi_auto_rgb
[12:33:50] [PASSED] drm_test_bridge_auto_first
[12:33:50] [PASSED] drm_test_bridge_rgb_yuv_no_path
[12:33:50] =============== [PASSED] drm_bridge_bus_fmt ================
[12:33:50] ============= drm_cmdline_parser (40 subtests) =============
[12:33:50] [PASSED] drm_test_cmdline_force_d_only
[12:33:50] [PASSED] drm_test_cmdline_force_D_only_dvi
[12:33:50] [PASSED] drm_test_cmdline_force_D_only_hdmi
[12:33:50] [PASSED] drm_test_cmdline_force_D_only_not_digital
[12:33:50] [PASSED] drm_test_cmdline_force_e_only
[12:33:50] [PASSED] drm_test_cmdline_res
[12:33:50] [PASSED] drm_test_cmdline_res_vesa
[12:33:50] [PASSED] drm_test_cmdline_res_vesa_rblank
[12:33:50] [PASSED] drm_test_cmdline_res_rblank
[12:33:50] [PASSED] drm_test_cmdline_res_bpp
[12:33:50] [PASSED] drm_test_cmdline_res_refresh
[12:33:50] [PASSED] drm_test_cmdline_res_bpp_refresh
[12:33:50] [PASSED] drm_test_cmdline_res_bpp_refresh_interlaced
[12:33:50] [PASSED] drm_test_cmdline_res_bpp_refresh_margins
[12:33:50] [PASSED] drm_test_cmdline_res_bpp_refresh_force_off
[12:33:50] [PASSED] drm_test_cmdline_res_bpp_refresh_force_on
[12:33:50] [PASSED] drm_test_cmdline_res_bpp_refresh_force_on_analog
[12:33:50] [PASSED] drm_test_cmdline_res_bpp_refresh_force_on_digital
[12:33:50] [PASSED] drm_test_cmdline_res_bpp_refresh_interlaced_margins_force_on
[12:33:50] [PASSED] drm_test_cmdline_res_margins_force_on
[12:33:50] [PASSED] drm_test_cmdline_res_vesa_margins
[12:33:50] [PASSED] drm_test_cmdline_name
[12:33:50] [PASSED] drm_test_cmdline_name_bpp
[12:33:50] [PASSED] drm_test_cmdline_name_option
[12:33:50] [PASSED] drm_test_cmdline_name_bpp_option
[12:33:50] [PASSED] drm_test_cmdline_rotate_0
[12:33:50] [PASSED] drm_test_cmdline_rotate_90
[12:33:50] [PASSED] drm_test_cmdline_rotate_180
[12:33:50] [PASSED] drm_test_cmdline_rotate_270
[12:33:50] [PASSED] drm_test_cmdline_hmirror
[12:33:50] [PASSED] drm_test_cmdline_vmirror
[12:33:50] [PASSED] drm_test_cmdline_margin_options
[12:33:50] [PASSED] drm_test_cmdline_multiple_options
[12:33:50] [PASSED] drm_test_cmdline_bpp_extra_and_option
[12:33:50] [PASSED] drm_test_cmdline_extra_and_option
[12:33:50] [PASSED] drm_test_cmdline_freestanding_options
[12:33:50] [PASSED] drm_test_cmdline_freestanding_force_e_and_options
[12:33:50] [PASSED] drm_test_cmdline_panel_orientation
[12:33:50] ================ drm_test_cmdline_invalid =================
[12:33:50] [PASSED] margin_only
[12:33:50] [PASSED] interlace_only
[12:33:50] [PASSED] res_missing_x
[12:33:50] [PASSED] res_missing_y
[12:33:50] [PASSED] res_bad_y
[12:33:50] [PASSED] res_missing_y_bpp
[12:33:50] [PASSED] res_bad_bpp
[12:33:50] [PASSED] res_bad_refresh
[12:33:50] [PASSED] res_bpp_refresh_force_on_off
[12:33:50] [PASSED] res_invalid_mode
[12:33:50] [PASSED] res_bpp_wrong_place_mode
[12:33:50] [PASSED] name_bpp_refresh
[12:33:50] [PASSED] name_refresh
[12:33:50] [PASSED] name_refresh_wrong_mode
[12:33:50] [PASSED] name_refresh_invalid_mode
[12:33:50] [PASSED] rotate_multiple
[12:33:50] [PASSED] rotate_invalid_val
[12:33:50] [PASSED] rotate_truncated
[12:33:50] [PASSED] invalid_option
[12:33:50] [PASSED] invalid_tv_option
[12:33:50] [PASSED] truncated_tv_option
[12:33:50] ============ [PASSED] drm_test_cmdline_invalid =============
[12:33:50] =============== drm_test_cmdline_tv_options ===============
[12:33:50] [PASSED] NTSC
[12:33:50] [PASSED] NTSC_443
[12:33:50] [PASSED] NTSC_J
[12:33:50] [PASSED] PAL
[12:33:50] [PASSED] PAL_M
[12:33:50] [PASSED] PAL_N
[12:33:50] [PASSED] SECAM
[12:33:50] [PASSED] MONO_525
[12:33:50] [PASSED] MONO_625
[12:33:50] =========== [PASSED] drm_test_cmdline_tv_options ===========
[12:33:50] =============== [PASSED] drm_cmdline_parser ================
[12:33:50] ========== drmm_connector_hdmi_init (20 subtests) ==========
[12:33:50] [PASSED] drm_test_connector_hdmi_init_valid
[12:33:50] [PASSED] drm_test_connector_hdmi_init_bpc_8
[12:33:50] [PASSED] drm_test_connector_hdmi_init_bpc_10
[12:33:50] [PASSED] drm_test_connector_hdmi_init_bpc_12
[12:33:50] [PASSED] drm_test_connector_hdmi_init_bpc_invalid
[12:33:50] [PASSED] drm_test_connector_hdmi_init_bpc_null
[12:33:50] [PASSED] drm_test_connector_hdmi_init_formats_empty
[12:33:50] [PASSED] drm_test_connector_hdmi_init_formats_no_rgb
[12:33:50] === drm_test_connector_hdmi_init_formats_yuv420_allowed ===
[12:33:50] [PASSED] supported_formats=0x9 yuv420_allowed=1
[12:33:50] [PASSED] supported_formats=0x9 yuv420_allowed=0
[12:33:50] [PASSED] supported_formats=0x5 yuv420_allowed=1
[12:33:50] [PASSED] supported_formats=0x5 yuv420_allowed=0
[12:33:50] === [PASSED] drm_test_connector_hdmi_init_formats_yuv420_allowed ===
[12:33:50] [PASSED] drm_test_connector_hdmi_init_null_ddc
[12:33:50] [PASSED] drm_test_connector_hdmi_init_null_product
[12:33:50] [PASSED] drm_test_connector_hdmi_init_null_vendor
[12:33:50] [PASSED] drm_test_connector_hdmi_init_product_length_exact
[12:33:50] [PASSED] drm_test_connector_hdmi_init_product_length_too_long
[12:33:50] [PASSED] drm_test_connector_hdmi_init_product_valid
[12:33:50] [PASSED] drm_test_connector_hdmi_init_vendor_length_exact
[12:33:50] [PASSED] drm_test_connector_hdmi_init_vendor_length_too_long
[12:33:50] [PASSED] drm_test_connector_hdmi_init_vendor_valid
[12:33:50] ========= drm_test_connector_hdmi_init_type_valid =========
[12:33:50] [PASSED] HDMI-A
[12:33:50] [PASSED] HDMI-B
[12:33:50] ===== [PASSED] drm_test_connector_hdmi_init_type_valid =====
[12:33:50] ======== drm_test_connector_hdmi_init_type_invalid ========
[12:33:50] [PASSED] Unknown
[12:33:50] [PASSED] VGA
[12:33:50] [PASSED] DVI-I
[12:33:50] [PASSED] DVI-D
[12:33:50] [PASSED] DVI-A
[12:33:50] [PASSED] Composite
[12:33:50] [PASSED] SVIDEO
[12:33:50] [PASSED] LVDS
[12:33:50] [PASSED] Component
[12:33:50] [PASSED] DIN
[12:33:50] [PASSED] DP
[12:33:50] [PASSED] TV
[12:33:50] [PASSED] eDP
[12:33:50] [PASSED] Virtual
[12:33:50] [PASSED] DSI
[12:33:50] [PASSED] DPI
[12:33:50] [PASSED] Writeback
[12:33:50] [PASSED] SPI
[12:33:50] [PASSED] USB
[12:33:50] ==== [PASSED] drm_test_connector_hdmi_init_type_invalid ====
[12:33:50] ============ [PASSED] drmm_connector_hdmi_init =============
[12:33:50] ============= drmm_connector_init (3 subtests) =============
[12:33:50] [PASSED] drm_test_drmm_connector_init
[12:33:50] [PASSED] drm_test_drmm_connector_init_null_ddc
[12:33:50] ========= drm_test_drmm_connector_init_type_valid =========
[12:33:50] [PASSED] Unknown
[12:33:50] [PASSED] VGA
[12:33:50] [PASSED] DVI-I
[12:33:50] [PASSED] DVI-D
[12:33:50] [PASSED] DVI-A
[12:33:50] [PASSED] Composite
[12:33:50] [PASSED] SVIDEO
[12:33:50] [PASSED] LVDS
[12:33:50] [PASSED] Component
[12:33:50] [PASSED] DIN
[12:33:50] [PASSED] DP
[12:33:50] [PASSED] HDMI-A
[12:33:50] [PASSED] HDMI-B
[12:33:50] [PASSED] TV
[12:33:50] [PASSED] eDP
[12:33:50] [PASSED] Virtual
[12:33:50] [PASSED] DSI
[12:33:50] [PASSED] DPI
[12:33:50] [PASSED] Writeback
[12:33:50] [PASSED] SPI
[12:33:50] [PASSED] USB
[12:33:50] ===== [PASSED] drm_test_drmm_connector_init_type_valid =====
[12:33:50] =============== [PASSED] drmm_connector_init ===============
[12:33:50] ========= drm_connector_dynamic_init (6 subtests) ==========
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_init
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_init_null_ddc
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_init_not_added
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_init_properties
[12:33:50] ===== drm_test_drm_connector_dynamic_init_type_valid ======
[12:33:50] [PASSED] Unknown
[12:33:50] [PASSED] VGA
[12:33:50] [PASSED] DVI-I
[12:33:50] [PASSED] DVI-D
[12:33:50] [PASSED] DVI-A
[12:33:50] [PASSED] Composite
[12:33:50] [PASSED] SVIDEO
[12:33:50] [PASSED] LVDS
[12:33:50] [PASSED] Component
[12:33:50] [PASSED] DIN
[12:33:50] [PASSED] DP
[12:33:50] [PASSED] HDMI-A
[12:33:50] [PASSED] HDMI-B
[12:33:50] [PASSED] TV
[12:33:50] [PASSED] eDP
[12:33:50] [PASSED] Virtual
[12:33:50] [PASSED] DSI
[12:33:50] [PASSED] DPI
[12:33:50] [PASSED] Writeback
[12:33:50] [PASSED] SPI
[12:33:50] [PASSED] USB
[12:33:50] = [PASSED] drm_test_drm_connector_dynamic_init_type_valid ==
[12:33:50] ======== drm_test_drm_connector_dynamic_init_name =========
[12:33:50] [PASSED] Unknown
[12:33:50] [PASSED] VGA
[12:33:50] [PASSED] DVI-I
[12:33:50] [PASSED] DVI-D
[12:33:50] [PASSED] DVI-A
[12:33:50] [PASSED] Composite
[12:33:50] [PASSED] SVIDEO
[12:33:50] [PASSED] LVDS
[12:33:50] [PASSED] Component
[12:33:50] [PASSED] DIN
[12:33:50] [PASSED] DP
[12:33:50] [PASSED] HDMI-A
[12:33:50] [PASSED] HDMI-B
[12:33:50] [PASSED] TV
[12:33:50] [PASSED] eDP
[12:33:50] [PASSED] Virtual
[12:33:50] [PASSED] DSI
[12:33:50] [PASSED] DPI
[12:33:50] [PASSED] Writeback
[12:33:50] [PASSED] SPI
[12:33:50] [PASSED] USB
[12:33:50] ==== [PASSED] drm_test_drm_connector_dynamic_init_name =====
[12:33:50] =========== [PASSED] drm_connector_dynamic_init ============
[12:33:50] ==== drm_connector_dynamic_register_early (4 subtests) =====
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_register_early_on_list
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_register_early_defer
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_register_early_no_init
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_register_early_no_mode_object
[12:33:50] ====== [PASSED] drm_connector_dynamic_register_early =======
[12:33:50] ======= drm_connector_dynamic_register (7 subtests) ========
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_register_on_list
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_register_no_defer
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_register_no_init
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_register_mode_object
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_register_sysfs
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_register_sysfs_name
[12:33:50] [PASSED] drm_test_drm_connector_dynamic_register_debugfs
[12:33:50] ========= [PASSED] drm_connector_dynamic_register ==========
[12:33:50] = drm_connector_attach_broadcast_rgb_property (2 subtests) =
[12:33:50] [PASSED] drm_test_drm_connector_attach_broadcast_rgb_property
[12:33:50] [PASSED] drm_test_drm_connector_attach_broadcast_rgb_property_hdmi_connector
[12:33:50] === [PASSED] drm_connector_attach_broadcast_rgb_property ===
[12:33:50] ========== drm_get_tv_mode_from_name (2 subtests) ==========
[12:33:50] ========== drm_test_get_tv_mode_from_name_valid ===========
[12:33:50] [PASSED] NTSC
[12:33:50] [PASSED] NTSC-443
[12:33:50] [PASSED] NTSC-J
[12:33:50] [PASSED] PAL
[12:33:50] [PASSED] PAL-M
[12:33:50] [PASSED] PAL-N
[12:33:50] [PASSED] SECAM
[12:33:50] [PASSED] Mono
[12:33:50] ====== [PASSED] drm_test_get_tv_mode_from_name_valid =======
[12:33:50] [PASSED] drm_test_get_tv_mode_from_name_truncated
[12:33:50] ============ [PASSED] drm_get_tv_mode_from_name ============
[12:33:50] = drm_test_connector_hdmi_compute_mode_clock (12 subtests) =
[12:33:50] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb
[12:33:50] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_10bpc
[12:33:50] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_10bpc_vic_1
[12:33:50] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_12bpc
[12:33:50] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_12bpc_vic_1
[12:33:50] [PASSED] drm_test_drm_hdmi_compute_mode_clock_rgb_double
[12:33:50] = drm_test_connector_hdmi_compute_mode_clock_yuv420_valid =
[12:33:50] [PASSED] VIC 96
[12:33:50] [PASSED] VIC 97
[12:33:50] [PASSED] VIC 101
[12:33:50] [PASSED] VIC 102
[12:33:50] [PASSED] VIC 106
[12:33:50] [PASSED] VIC 107
[12:33:50] === [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv420_valid ===
[12:33:50] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv420_10_bpc
[12:33:50] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv420_12_bpc
[12:33:50] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv422_8_bpc
[12:33:50] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv422_10_bpc
[12:33:50] [PASSED] drm_test_connector_hdmi_compute_mode_clock_yuv422_12_bpc
[12:33:50] === [PASSED] drm_test_connector_hdmi_compute_mode_clock ====
[12:33:50] == drm_hdmi_connector_get_broadcast_rgb_name (2 subtests) ==
[12:33:50] === drm_test_drm_hdmi_connector_get_broadcast_rgb_name ====
[12:33:50] [PASSED] Automatic
[12:33:50] [PASSED] Full
[12:33:50] [PASSED] Limited 16:235
[12:33:50] === [PASSED] drm_test_drm_hdmi_connector_get_broadcast_rgb_name ===
[12:33:50] [PASSED] drm_test_drm_hdmi_connector_get_broadcast_rgb_name_invalid
[12:33:50] ==== [PASSED] drm_hdmi_connector_get_broadcast_rgb_name ====
[12:33:50] == drm_hdmi_connector_get_output_format_name (2 subtests) ==
[12:33:50] === drm_test_drm_hdmi_connector_get_output_format_name ====
[12:33:50] [PASSED] RGB
[12:33:50] [PASSED] YUV 4:2:0
[12:33:50] [PASSED] YUV 4:2:2
[12:33:50] [PASSED] YUV 4:4:4
[12:33:50] === [PASSED] drm_test_drm_hdmi_connector_get_output_format_name ===
[12:33:50] [PASSED] drm_test_drm_hdmi_connector_get_output_format_name_invalid
[12:33:50] ==== [PASSED] drm_hdmi_connector_get_output_format_name ====
[12:33:50] ============= drm_damage_helper (21 subtests) ==============
[12:33:50] [PASSED] drm_test_damage_iter_no_damage
[12:33:50] [PASSED] drm_test_damage_iter_no_damage_fractional_src
[12:33:50] [PASSED] drm_test_damage_iter_no_damage_src_moved
[12:33:50] [PASSED] drm_test_damage_iter_no_damage_fractional_src_moved
[12:33:50] [PASSED] drm_test_damage_iter_no_damage_not_visible
[12:33:50] [PASSED] drm_test_damage_iter_no_damage_no_crtc
[12:33:50] [PASSED] drm_test_damage_iter_no_damage_no_fb
[12:33:50] [PASSED] drm_test_damage_iter_simple_damage
[12:33:50] [PASSED] drm_test_damage_iter_single_damage
[12:33:50] [PASSED] drm_test_damage_iter_single_damage_intersect_src
[12:33:50] [PASSED] drm_test_damage_iter_single_damage_outside_src
[12:33:50] [PASSED] drm_test_damage_iter_single_damage_fractional_src
[12:33:50] [PASSED] drm_test_damage_iter_single_damage_intersect_fractional_src
[12:33:50] [PASSED] drm_test_damage_iter_single_damage_outside_fractional_src
[12:33:50] [PASSED] drm_test_damage_iter_single_damage_src_moved
[12:33:50] [PASSED] drm_test_damage_iter_single_damage_fractional_src_moved
[12:33:50] [PASSED] drm_test_damage_iter_damage
[12:33:50] [PASSED] drm_test_damage_iter_damage_one_intersect
[12:33:50] [PASSED] drm_test_damage_iter_damage_one_outside
[12:33:50] [PASSED] drm_test_damage_iter_damage_src_moved
[12:33:50] [PASSED] drm_test_damage_iter_damage_not_visible
[12:33:50] ================ [PASSED] drm_damage_helper ================
[12:33:50] ============== drm_dp_mst_helper (3 subtests) ==============
[12:33:50] ============== drm_test_dp_mst_calc_pbn_mode ==============
[12:33:50] [PASSED] Clock 154000 BPP 30 DSC disabled
[12:33:50] [PASSED] Clock 234000 BPP 30 DSC disabled
[12:33:50] [PASSED] Clock 297000 BPP 24 DSC disabled
[12:33:50] [PASSED] Clock 332880 BPP 24 DSC enabled
[12:33:50] [PASSED] Clock 324540 BPP 24 DSC enabled
[12:33:50] ========== [PASSED] drm_test_dp_mst_calc_pbn_mode ==========
[12:33:50] ============== drm_test_dp_mst_calc_pbn_div ===============
[12:33:50] [PASSED] Link rate 2000000 lane count 4
[12:33:50] [PASSED] Link rate 2000000 lane count 2
[12:33:50] [PASSED] Link rate 2000000 lane count 1
[12:33:50] [PASSED] Link rate 1350000 lane count 4
[12:33:50] [PASSED] Link rate 1350000 lane count 2
[12:33:50] [PASSED] Link rate 1350000 lane count 1
[12:33:50] [PASSED] Link rate 1000000 lane count 4
[12:33:50] [PASSED] Link rate 1000000 lane count 2
[12:33:50] [PASSED] Link rate 1000000 lane count 1
[12:33:50] [PASSED] Link rate 810000 lane count 4
[12:33:50] [PASSED] Link rate 810000 lane count 2
[12:33:50] [PASSED] Link rate 810000 lane count 1
[12:33:50] [PASSED] Link rate 540000 lane count 4
[12:33:50] [PASSED] Link rate 540000 lane count 2
[12:33:50] [PASSED] Link rate 540000 lane count 1
[12:33:50] [PASSED] Link rate 270000 lane count 4
[12:33:50] [PASSED] Link rate 270000 lane count 2
[12:33:50] [PASSED] Link rate 270000 lane count 1
[12:33:50] [PASSED] Link rate 162000 lane count 4
[12:33:50] [PASSED] Link rate 162000 lane count 2
[12:33:50] [PASSED] Link rate 162000 lane count 1
[12:33:50] ========== [PASSED] drm_test_dp_mst_calc_pbn_div ===========
[12:33:50] ========= drm_test_dp_mst_sideband_msg_req_decode =========
[12:33:50] [PASSED] DP_ENUM_PATH_RESOURCES with port number
[12:33:50] [PASSED] DP_POWER_UP_PHY with port number
[12:33:50] [PASSED] DP_POWER_DOWN_PHY with port number
[12:33:50] [PASSED] DP_ALLOCATE_PAYLOAD with SDP stream sinks
[12:33:50] [PASSED] DP_ALLOCATE_PAYLOAD with port number
[12:33:50] [PASSED] DP_ALLOCATE_PAYLOAD with VCPI
[12:33:50] [PASSED] DP_ALLOCATE_PAYLOAD with PBN
[12:33:50] [PASSED] DP_QUERY_PAYLOAD with port number
[12:33:50] [PASSED] DP_QUERY_PAYLOAD with VCPI
[12:33:50] [PASSED] DP_REMOTE_DPCD_READ with port number
[12:33:50] [PASSED] DP_REMOTE_DPCD_READ with DPCD address
[12:33:50] [PASSED] DP_REMOTE_DPCD_READ with max number of bytes
[12:33:50] [PASSED] DP_REMOTE_DPCD_WRITE with port number
[12:33:50] [PASSED] DP_REMOTE_DPCD_WRITE with DPCD address
[12:33:50] [PASSED] DP_REMOTE_DPCD_WRITE with data array
[12:33:50] [PASSED] DP_REMOTE_I2C_READ with port number
[12:33:50] [PASSED] DP_REMOTE_I2C_READ with I2C device ID
[12:33:50] [PASSED] DP_REMOTE_I2C_READ with transactions array
[12:33:50] [PASSED] DP_REMOTE_I2C_WRITE with port number
[12:33:50] [PASSED] DP_REMOTE_I2C_WRITE with I2C device ID
[12:33:50] [PASSED] DP_REMOTE_I2C_WRITE with data array
[12:33:50] [PASSED] DP_QUERY_STREAM_ENC_STATUS with stream ID
[12:33:50] [PASSED] DP_QUERY_STREAM_ENC_STATUS with client ID
[12:33:50] [PASSED] DP_QUERY_STREAM_ENC_STATUS with stream event
[12:33:50] [PASSED] DP_QUERY_STREAM_ENC_STATUS with valid stream event
[12:33:50] [PASSED] DP_QUERY_STREAM_ENC_STATUS with stream behavior
[12:33:50] [PASSED] DP_QUERY_STREAM_ENC_STATUS with a valid stream behavior
[12:33:50] ===== [PASSED] drm_test_dp_mst_sideband_msg_req_decode =====
[12:33:50] ================ [PASSED] drm_dp_mst_helper ================
[12:33:50] ================== drm_exec (7 subtests) ===================
[12:33:50] [PASSED] sanitycheck
[12:33:50] [PASSED] test_lock
[12:33:50] [PASSED] test_lock_unlock
[12:33:50] [PASSED] test_duplicates
[12:33:50] [PASSED] test_prepare
[12:33:50] [PASSED] test_prepare_array
[12:33:50] [PASSED] test_multiple_loops
[12:33:50] ==================== [PASSED] drm_exec =====================
[12:33:50] =========== drm_format_helper_test (17 subtests) ===========
[12:33:50] ============== drm_test_fb_xrgb8888_to_gray8 ==============
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ========== [PASSED] drm_test_fb_xrgb8888_to_gray8 ==========
[12:33:50] ============= drm_test_fb_xrgb8888_to_rgb332 ==============
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ========= [PASSED] drm_test_fb_xrgb8888_to_rgb332 ==========
[12:33:50] ============= drm_test_fb_xrgb8888_to_rgb565 ==============
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ========= [PASSED] drm_test_fb_xrgb8888_to_rgb565 ==========
[12:33:50] ============ drm_test_fb_xrgb8888_to_xrgb1555 =============
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ======== [PASSED] drm_test_fb_xrgb8888_to_xrgb1555 =========
[12:33:50] ============ drm_test_fb_xrgb8888_to_argb1555 =============
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ======== [PASSED] drm_test_fb_xrgb8888_to_argb1555 =========
[12:33:50] ============ drm_test_fb_xrgb8888_to_rgba5551 =============
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ======== [PASSED] drm_test_fb_xrgb8888_to_rgba5551 =========
[12:33:50] ============= drm_test_fb_xrgb8888_to_rgb888 ==============
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ========= [PASSED] drm_test_fb_xrgb8888_to_rgb888 ==========
[12:33:50] ============= drm_test_fb_xrgb8888_to_bgr888 ==============
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ========= [PASSED] drm_test_fb_xrgb8888_to_bgr888 ==========
[12:33:50] ============ drm_test_fb_xrgb8888_to_argb8888 =============
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ======== [PASSED] drm_test_fb_xrgb8888_to_argb8888 =========
[12:33:50] =========== drm_test_fb_xrgb8888_to_xrgb2101010 ===========
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ======= [PASSED] drm_test_fb_xrgb8888_to_xrgb2101010 =======
[12:33:50] =========== drm_test_fb_xrgb8888_to_argb2101010 ===========
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ======= [PASSED] drm_test_fb_xrgb8888_to_argb2101010 =======
[12:33:50] ============== drm_test_fb_xrgb8888_to_mono ===============
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ========== [PASSED] drm_test_fb_xrgb8888_to_mono ===========
[12:33:50] ==================== drm_test_fb_swab =====================
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ================ [PASSED] drm_test_fb_swab =================
[12:33:50] ============ drm_test_fb_xrgb8888_to_xbgr8888 =============
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ======== [PASSED] drm_test_fb_xrgb8888_to_xbgr8888 =========
[12:33:50] ============ drm_test_fb_xrgb8888_to_abgr8888 =============
[12:33:50] [PASSED] single_pixel_source_buffer
[12:33:50] [PASSED] single_pixel_clip_rectangle
[12:33:50] [PASSED] well_known_colors
[12:33:50] [PASSED] destination_pitch
[12:33:50] ======== [PASSED] drm_test_fb_xrgb8888_to_abgr8888 =========
[12:33:50] ================= drm_test_fb_clip_offset =================
[12:33:50] [PASSED] pass through
[12:33:50] [PASSED] horizontal offset
[12:33:50] [PASSED] vertical offset
[12:33:50] [PASSED] horizontal and vertical offset
[12:33:50] [PASSED] horizontal offset (custom pitch)
[12:33:50] [PASSED] vertical offset (custom pitch)
[12:33:50] [PASSED] horizontal and vertical offset (custom pitch)
[12:33:50] ============= [PASSED] drm_test_fb_clip_offset =============
[12:33:50] =================== drm_test_fb_memcpy ====================
[12:33:50] [PASSED] single_pixel_source_buffer: XR24 little-endian (0x34325258)
[12:33:50] [PASSED] single_pixel_source_buffer: XRA8 little-endian (0x38415258)
[12:33:50] [PASSED] single_pixel_source_buffer: YU24 little-endian (0x34325559)
[12:33:50] [PASSED] single_pixel_clip_rectangle: XB24 little-endian (0x34324258)
[12:33:50] [PASSED] single_pixel_clip_rectangle: XRA8 little-endian (0x38415258)
[12:33:50] [PASSED] single_pixel_clip_rectangle: YU24 little-endian (0x34325559)
[12:33:50] [PASSED] well_known_colors: XB24 little-endian (0x34324258)
[12:33:50] [PASSED] well_known_colors: XRA8 little-endian (0x38415258)
[12:33:50] [PASSED] well_known_colors: YU24 little-endian (0x34325559)
[12:33:50] [PASSED] destination_pitch: XB24 little-endian (0x34324258)
[12:33:50] [PASSED] destination_pitch: XRA8 little-endian (0x38415258)
[12:33:50] [PASSED] destination_pitch: YU24 little-endian (0x34325559)
[12:33:50] =============== [PASSED] drm_test_fb_memcpy ================
[12:33:50] ============= [PASSED] drm_format_helper_test ==============
[12:33:50] ================= drm_format (18 subtests) =================
[12:33:50] [PASSED] drm_test_format_block_width_invalid
[12:33:50] [PASSED] drm_test_format_block_width_one_plane
[12:33:50] [PASSED] drm_test_format_block_width_two_plane
[12:33:50] [PASSED] drm_test_format_block_width_three_plane
[12:33:50] [PASSED] drm_test_format_block_width_tiled
[12:33:50] [PASSED] drm_test_format_block_height_invalid
[12:33:50] [PASSED] drm_test_format_block_height_one_plane
[12:33:50] [PASSED] drm_test_format_block_height_two_plane
[12:33:50] [PASSED] drm_test_format_block_height_three_plane
[12:33:50] [PASSED] drm_test_format_block_height_tiled
[12:33:50] [PASSED] drm_test_format_min_pitch_invalid
[12:33:50] [PASSED] drm_test_format_min_pitch_one_plane_8bpp
[12:33:50] [PASSED] drm_test_format_min_pitch_one_plane_16bpp
[12:33:50] [PASSED] drm_test_format_min_pitch_one_plane_24bpp
[12:33:50] [PASSED] drm_test_format_min_pitch_one_plane_32bpp
[12:33:50] [PASSED] drm_test_format_min_pitch_two_plane
[12:33:50] [PASSED] drm_test_format_min_pitch_three_plane_8bpp
[12:33:50] [PASSED] drm_test_format_min_pitch_tiled
[12:33:50] =================== [PASSED] drm_format ====================
[12:33:50] ============== drm_framebuffer (10 subtests) ===============
[12:33:50] ========== drm_test_framebuffer_check_src_coords ==========
[12:33:50] [PASSED] Success: source fits into fb
[12:33:50] [PASSED] Fail: overflowing fb with x-axis coordinate
[12:33:50] [PASSED] Fail: overflowing fb with y-axis coordinate
[12:33:50] [PASSED] Fail: overflowing fb with source width
[12:33:50] [PASSED] Fail: overflowing fb with source height
[12:33:50] ====== [PASSED] drm_test_framebuffer_check_src_coords ======
[12:33:50] [PASSED] drm_test_framebuffer_cleanup
[12:33:50] =============== drm_test_framebuffer_create ===============
[12:33:50] [PASSED] ABGR8888 normal sizes
[12:33:50] [PASSED] ABGR8888 max sizes
[12:33:50] [PASSED] ABGR8888 pitch greater than min required
[12:33:50] [PASSED] ABGR8888 pitch less than min required
[12:33:50] [PASSED] ABGR8888 Invalid width
[12:33:50] [PASSED] ABGR8888 Invalid buffer handle
[12:33:50] [PASSED] No pixel format
[12:33:50] [PASSED] ABGR8888 Width 0
[12:33:50] [PASSED] ABGR8888 Height 0
[12:33:50] [PASSED] ABGR8888 Out of bound height * pitch combination
[12:33:50] [PASSED] ABGR8888 Large buffer offset
[12:33:50] [PASSED] ABGR8888 Buffer offset for inexistent plane
[12:33:50] [PASSED] ABGR8888 Invalid flag
[12:33:50] [PASSED] ABGR8888 Set DRM_MODE_FB_MODIFIERS without modifiers
[12:33:50] [PASSED] ABGR8888 Valid buffer modifier
[12:33:50] [PASSED] ABGR8888 Invalid buffer modifier(DRM_FORMAT_MOD_SAMSUNG_64_32_TILE)
[12:33:50] [PASSED] ABGR8888 Extra pitches without DRM_MODE_FB_MODIFIERS
[12:33:50] [PASSED] ABGR8888 Extra pitches with DRM_MODE_FB_MODIFIERS
[12:33:50] [PASSED] NV12 Normal sizes
[12:33:50] [PASSED] NV12 Max sizes
[12:33:50] [PASSED] NV12 Invalid pitch
[12:33:50] [PASSED] NV12 Invalid modifier/missing DRM_MODE_FB_MODIFIERS flag
[12:33:50] [PASSED] NV12 different modifier per-plane
[12:33:50] [PASSED] NV12 with DRM_FORMAT_MOD_SAMSUNG_64_32_TILE
[12:33:50] [PASSED] NV12 Valid modifiers without DRM_MODE_FB_MODIFIERS
[12:33:50] [PASSED] NV12 Modifier for inexistent plane
[12:33:50] [PASSED] NV12 Handle for inexistent plane
[12:33:50] [PASSED] NV12 Handle for inexistent plane without DRM_MODE_FB_MODIFIERS
[12:33:50] [PASSED] YVU420 DRM_MODE_FB_MODIFIERS set without modifier
[12:33:50] [PASSED] YVU420 Normal sizes
[12:33:50] [PASSED] YVU420 Max sizes
[12:33:50] [PASSED] YVU420 Invalid pitch
[12:33:50] [PASSED] YVU420 Different pitches
[12:33:50] [PASSED] YVU420 Different buffer offsets/pitches
[12:33:50] [PASSED] YVU420 Modifier set just for plane 0, without DRM_MODE_FB_MODIFIERS
[12:33:50] [PASSED] YVU420 Modifier set just for planes 0, 1, without DRM_MODE_FB_MODIFIERS
[12:33:50] [PASSED] YVU420 Modifier set just for plane 0, 1, with DRM_MODE_FB_MODIFIERS
[12:33:50] [PASSED] YVU420 Valid modifier
[12:33:50] [PASSED] YVU420 Different modifiers per plane
[12:33:50] [PASSED] YVU420 Modifier for inexistent plane
[12:33:50] [PASSED] YUV420_10BIT Invalid modifier(DRM_FORMAT_MOD_LINEAR)
[12:33:50] [PASSED] X0L2 Normal sizes
[12:33:50] [PASSED] X0L2 Max sizes
[12:33:50] [PASSED] X0L2 Invalid pitch
[12:33:50] [PASSED] X0L2 Pitch greater than minimum required
[12:33:50] [PASSED] X0L2 Handle for inexistent plane
[12:33:50] [PASSED] X0L2 Offset for inexistent plane, without DRM_MODE_FB_MODIFIERS set
[12:33:50] [PASSED] X0L2 Modifier without DRM_MODE_FB_MODIFIERS set
[12:33:50] [PASSED] X0L2 Valid modifier
[12:33:50] [PASSED] X0L2 Modifier for inexistent plane
[12:33:50] =========== [PASSED] drm_test_framebuffer_create ===========
[12:33:50] [PASSED] drm_test_framebuffer_free
[12:33:50] [PASSED] drm_test_framebuffer_init
[12:33:50] [PASSED] drm_test_framebuffer_init_bad_format
[12:33:50] [PASSED] drm_test_framebuffer_init_dev_mismatch
[12:33:50] [PASSED] drm_test_framebuffer_lookup
[12:33:50] [PASSED] drm_test_framebuffer_lookup_inexistent
[12:33:50] [PASSED] drm_test_framebuffer_modifiers_not_supported
[12:33:50] ================= [PASSED] drm_framebuffer =================
[12:33:50] ================ drm_gem_shmem (8 subtests) ================
[12:33:50] [PASSED] drm_gem_shmem_test_obj_create
[12:33:50] [PASSED] drm_gem_shmem_test_obj_create_private
[12:33:50] [PASSED] drm_gem_shmem_test_pin_pages
[12:33:50] [PASSED] drm_gem_shmem_test_vmap
[12:33:50] [PASSED] drm_gem_shmem_test_get_sg_table
[12:33:50] [PASSED] drm_gem_shmem_test_get_pages_sgt
[12:33:50] [PASSED] drm_gem_shmem_test_madvise
[12:33:50] [PASSED] drm_gem_shmem_test_purge
[12:33:50] ================== [PASSED] drm_gem_shmem ==================
[12:33:50] === drm_atomic_helper_connector_hdmi_check (29 subtests) ===
[12:33:50] [PASSED] drm_test_check_broadcast_rgb_auto_cea_mode
[12:33:50] [PASSED] drm_test_check_broadcast_rgb_auto_cea_mode_vic_1
[12:33:50] [PASSED] drm_test_check_broadcast_rgb_full_cea_mode
[12:33:50] [PASSED] drm_test_check_broadcast_rgb_full_cea_mode_vic_1
[12:33:50] [PASSED] drm_test_check_broadcast_rgb_limited_cea_mode
[12:33:50] [PASSED] drm_test_check_broadcast_rgb_limited_cea_mode_vic_1
[12:33:50] ====== drm_test_check_broadcast_rgb_cea_mode_yuv420 =======
[12:33:50] [PASSED] Automatic
[12:33:50] [PASSED] Full
[12:33:50] [PASSED] Limited 16:235
[12:33:50] == [PASSED] drm_test_check_broadcast_rgb_cea_mode_yuv420 ===
[12:33:50] [PASSED] drm_test_check_broadcast_rgb_crtc_mode_changed
[12:33:50] [PASSED] drm_test_check_broadcast_rgb_crtc_mode_not_changed
[12:33:50] [PASSED] drm_test_check_disable_connector
[12:33:50] [PASSED] drm_test_check_hdmi_funcs_reject_rate
[12:33:50] [PASSED] drm_test_check_max_tmds_rate_bpc_fallback_rgb
[12:33:50] [PASSED] drm_test_check_max_tmds_rate_bpc_fallback_yuv420
[12:33:50] [PASSED] drm_test_check_max_tmds_rate_bpc_fallback_ignore_yuv422
[12:33:50] [PASSED] drm_test_check_max_tmds_rate_bpc_fallback_ignore_yuv420
[12:33:50] [PASSED] drm_test_check_driver_unsupported_fallback_yuv420
[12:33:50] [PASSED] drm_test_check_output_bpc_crtc_mode_changed
[12:33:50] [PASSED] drm_test_check_output_bpc_crtc_mode_not_changed
[12:33:50] [PASSED] drm_test_check_output_bpc_dvi
[12:33:50] [PASSED] drm_test_check_output_bpc_format_vic_1
[12:33:50] [PASSED] drm_test_check_output_bpc_format_display_8bpc_only
[12:33:50] [PASSED] drm_test_check_output_bpc_format_display_rgb_only
[12:33:50] [PASSED] drm_test_check_output_bpc_format_driver_8bpc_only
[12:33:50] [PASSED] drm_test_check_output_bpc_format_driver_rgb_only
[12:33:50] [PASSED] drm_test_check_tmds_char_rate_rgb_8bpc
[12:33:50] [PASSED] drm_test_check_tmds_char_rate_rgb_10bpc
[12:33:50] [PASSED] drm_test_check_tmds_char_rate_rgb_12bpc
[12:33:50] ============ drm_test_check_hdmi_color_format =============
[12:33:50] [PASSED] AUTO -> RGB
[12:33:50] [PASSED] YCBCR422 -> YUV422
[12:33:50] [PASSED] YCBCR420 -> YUV420
[12:33:50] [PASSED] YCBCR444 -> YUV444
[12:33:50] [PASSED] RGB -> RGB
[12:33:50] ======== [PASSED] drm_test_check_hdmi_color_format =========
[12:33:50] ======== drm_test_check_hdmi_color_format_420_only ========
[12:33:50] [PASSED] RGB should fail
[12:33:50] [PASSED] YUV444 should fail
[12:33:50] [PASSED] YUV422 should fail
[12:33:50] [PASSED] YUV420 should work
[12:33:50] ==== [PASSED] drm_test_check_hdmi_color_format_420_only ====
[12:33:50] ===== [PASSED] drm_atomic_helper_connector_hdmi_check ======
[12:33:50] === drm_atomic_helper_connector_hdmi_reset (6 subtests) ====
[12:33:50] [PASSED] drm_test_check_broadcast_rgb_value
[12:33:50] [PASSED] drm_test_check_bpc_8_value
[12:33:50] [PASSED] drm_test_check_bpc_10_value
[12:33:50] [PASSED] drm_test_check_bpc_12_value
[12:33:50] [PASSED] drm_test_check_format_value
[12:33:50] [PASSED] drm_test_check_tmds_char_value
[12:33:50] ===== [PASSED] drm_atomic_helper_connector_hdmi_reset ======
[12:33:50] = drm_atomic_helper_connector_hdmi_mode_valid (7 subtests) =
[12:33:50] [PASSED] drm_test_check_mode_valid
[12:33:50] [PASSED] drm_test_check_mode_valid_reject
[12:33:50] [PASSED] drm_test_check_mode_valid_reject_rate
[12:33:50] [PASSED] drm_test_check_mode_valid_reject_max_clock
[12:33:50] [PASSED] drm_test_check_mode_valid_yuv420_only_max_clock
[12:33:50] [PASSED] drm_test_check_mode_valid_reject_yuv420_only_connector
[12:33:50] [PASSED] drm_test_check_mode_valid_accept_yuv420_also_connector_rgb
[12:33:50] === [PASSED] drm_atomic_helper_connector_hdmi_mode_valid ===
[12:33:50] = drm_atomic_helper_connector_hdmi_infoframes (5 subtests) =
[12:33:50] [PASSED] drm_test_check_infoframes
[12:33:50] [PASSED] drm_test_check_reject_avi_infoframe
[12:33:50] [PASSED] drm_test_check_reject_hdr_infoframe_bpc_8
[12:33:50] [PASSED] drm_test_check_reject_hdr_infoframe_bpc_10
[12:33:50] [PASSED] drm_test_check_reject_audio_infoframe
[12:33:50] === [PASSED] drm_atomic_helper_connector_hdmi_infoframes ===
[12:33:50] ================= drm_managed (2 subtests) =================
[12:33:50] [PASSED] drm_test_managed_release_action
[12:33:50] [PASSED] drm_test_managed_run_action
[12:33:50] =================== [PASSED] drm_managed ===================
[12:33:50] =================== drm_mm (6 subtests) ====================
[12:33:50] [PASSED] drm_test_mm_init
[12:33:50] [PASSED] drm_test_mm_debug
[12:33:50] [PASSED] drm_test_mm_align32
[12:33:50] [PASSED] drm_test_mm_align64
[12:33:50] [PASSED] drm_test_mm_lowest
[12:33:50] [PASSED] drm_test_mm_highest
[12:33:50] ===================== [PASSED] drm_mm ======================
[12:33:50] ============= drm_modes_analog_tv (5 subtests) =============
[12:33:50] [PASSED] drm_test_modes_analog_tv_mono_576i
[12:33:50] [PASSED] drm_test_modes_analog_tv_ntsc_480i
[12:33:50] [PASSED] drm_test_modes_analog_tv_ntsc_480i_inlined
[12:33:50] [PASSED] drm_test_modes_analog_tv_pal_576i
[12:33:50] [PASSED] drm_test_modes_analog_tv_pal_576i_inlined
[12:33:50] =============== [PASSED] drm_modes_analog_tv ===============
[12:33:50] ============== drm_plane_helper (2 subtests) ===============
[12:33:50] =============== drm_test_check_plane_state ================
[12:33:50] [PASSED] clipping_simple
[12:33:50] [PASSED] clipping_rotate_reflect
[12:33:50] [PASSED] positioning_simple
[12:33:50] [PASSED] upscaling
[12:33:50] [PASSED] downscaling
[12:33:50] [PASSED] rounding1
[12:33:50] [PASSED] rounding2
[12:33:50] [PASSED] rounding3
[12:33:50] [PASSED] rounding4
[12:33:50] =========== [PASSED] drm_test_check_plane_state ============
[12:33:50] =========== drm_test_check_invalid_plane_state ============
[12:33:50] [PASSED] positioning_invalid
[12:33:50] [PASSED] upscaling_invalid
[12:33:50] [PASSED] downscaling_invalid
[12:33:50] ======= [PASSED] drm_test_check_invalid_plane_state ========
[12:33:50] ================ [PASSED] drm_plane_helper =================
[12:33:50] ====== drm_connector_helper_tv_get_modes (1 subtest) =======
[12:33:50] ====== drm_test_connector_helper_tv_get_modes_check =======
[12:33:50] [PASSED] None
[12:33:50] [PASSED] PAL
[12:33:50] [PASSED] NTSC
[12:33:50] [PASSED] Both, NTSC Default
[12:33:50] [PASSED] Both, PAL Default
[12:33:50] [PASSED] Both, NTSC Default, with PAL on command-line
[12:33:50] [PASSED] Both, PAL Default, with NTSC on command-line
[12:33:50] == [PASSED] drm_test_connector_helper_tv_get_modes_check ===
[12:33:50] ======== [PASSED] drm_connector_helper_tv_get_modes ========
[12:33:50] ================== drm_rect (9 subtests) ===================
[12:33:50] [PASSED] drm_test_rect_clip_scaled_div_by_zero
[12:33:50] [PASSED] drm_test_rect_clip_scaled_not_clipped
[12:33:50] [PASSED] drm_test_rect_clip_scaled_clipped
[12:33:50] [PASSED] drm_test_rect_clip_scaled_signed_vs_unsigned
[12:33:50] ================= drm_test_rect_intersect =================
[12:33:50] [PASSED] top-left x bottom-right: 2x2+1+1 x 2x2+0+0
[12:33:50] [PASSED] top-right x bottom-left: 2x2+0+0 x 2x2+1-1
[12:33:50] [PASSED] bottom-left x top-right: 2x2+1-1 x 2x2+0+0
[12:33:50] [PASSED] bottom-right x top-left: 2x2+0+0 x 2x2+1+1
[12:33:50] [PASSED] right x left: 2x1+0+0 x 3x1+1+0
[12:33:50] [PASSED] left x right: 3x1+1+0 x 2x1+0+0
[12:33:50] [PASSED] up x bottom: 1x2+0+0 x 1x3+0-1
[12:33:50] [PASSED] bottom x up: 1x3+0-1 x 1x2+0+0
[12:33:50] [PASSED] touching corner: 1x1+0+0 x 2x2+1+1
[12:33:50] [PASSED] touching side: 1x1+0+0 x 1x1+1+0
[12:33:50] [PASSED] equal rects: 2x2+0+0 x 2x2+0+0
[12:33:50] [PASSED] inside another: 2x2+0+0 x 1x1+1+1
[12:33:50] [PASSED] far away: 1x1+0+0 x 1x1+3+6
[12:33:50] [PASSED] points intersecting: 0x0+5+10 x 0x0+5+10
[12:33:50] [PASSED] points not intersecting: 0x0+0+0 x 0x0+5+10
[12:33:50] ============= [PASSED] drm_test_rect_intersect =============
[12:33:50] ================ drm_test_rect_calc_hscale ================
[12:33:50] [PASSED] normal use
[12:33:50] [PASSED] out of max range
[12:33:50] [PASSED] out of min range
[12:33:50] [PASSED] zero dst
[12:33:50] [PASSED] negative src
[12:33:50] [PASSED] negative dst
[12:33:50] ============ [PASSED] drm_test_rect_calc_hscale ============
[12:33:50] ================ drm_test_rect_calc_vscale ================
[12:33:50] [PASSED] normal use
[12:33:50] [PASSED] out of max range
[12:33:50] [PASSED] out of min range
[12:33:50] [PASSED] zero dst
[12:33:50] [PASSED] negative src
[12:33:50] [PASSED] negative dst
[12:33:50] ============ [PASSED] drm_test_rect_calc_vscale ============
[12:33:50] ================== drm_test_rect_rotate ===================
[12:33:50] [PASSED] reflect-x
[12:33:50] [PASSED] reflect-y
[12:33:50] [PASSED] rotate-0
[12:33:50] [PASSED] rotate-90
[12:33:50] [PASSED] rotate-180
[12:33:50] [PASSED] rotate-270
[12:33:50] ============== [PASSED] drm_test_rect_rotate ===============
[12:33:50] ================ drm_test_rect_rotate_inv =================
[12:33:50] [PASSED] reflect-x
[12:33:50] [PASSED] reflect-y
[12:33:50] [PASSED] rotate-0
[12:33:50] [PASSED] rotate-90
[12:33:50] [PASSED] rotate-180
[12:33:50] [PASSED] rotate-270
[12:33:50] ============ [PASSED] drm_test_rect_rotate_inv =============
[12:33:50] ==================== [PASSED] drm_rect =====================
[12:33:50] ============ drm_sysfb_modeset_test (1 subtest) ============
[12:33:50] ============ drm_test_sysfb_build_fourcc_list =============
[12:33:50] [PASSED] no native formats
[12:33:50] [PASSED] XRGB8888 as native format
[12:33:50] [PASSED] remove duplicates
[12:33:50] [PASSED] convert alpha formats
[12:33:50] [PASSED] random formats
[12:33:50] ======== [PASSED] drm_test_sysfb_build_fourcc_list =========
[12:33:50] ============= [PASSED] drm_sysfb_modeset_test ==============
[12:33:50] ================== drm_fixp (2 subtests) ===================
[12:33:50] [PASSED] drm_test_int2fixp
[12:33:50] [PASSED] drm_test_sm2fixp
[12:33:50] ==================== [PASSED] drm_fixp =====================
[12:33:50] ============================================================
[12:33:50] Testing complete. Ran 637 tests: passed: 637
[12:33:50] Elapsed time: 26.605s total, 1.817s configuring, 24.568s building, 0.203s running
+ /kernel/tools/testing/kunit/kunit.py run --kunitconfig /kernel/drivers/gpu/drm/ttm/tests/.kunitconfig
[12:33:50] Configuring KUnit Kernel ...
Regenerating .config ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
[12:33:52] Building KUnit Kernel ...
Populating config with:
$ make ARCH=um O=.kunit olddefconfig
Building with:
$ make all compile_commands.json scripts_gdb ARCH=um O=.kunit --jobs=48
[12:34:02] Starting KUnit Kernel (1/1)...
[12:34:02] ============================================================
Running tests with:
$ .kunit/linux kunit.enable=1 mem=1G console=tty kunit_shutdown=halt
[12:34:02] ================= ttm_device (5 subtests) ==================
[12:34:02] [PASSED] ttm_device_init_basic
[12:34:02] [PASSED] ttm_device_init_multiple
[12:34:02] [PASSED] ttm_device_fini_basic
[12:34:02] [PASSED] ttm_device_init_no_vma_man
[12:34:02] ================== ttm_device_init_pools ==================
[12:34:02] [PASSED] No DMA allocations, no DMA32 required
[12:34:02] [PASSED] DMA allocations, DMA32 required
[12:34:02] [PASSED] No DMA allocations, DMA32 required
[12:34:02] [PASSED] DMA allocations, no DMA32 required
[12:34:02] ============== [PASSED] ttm_device_init_pools ==============
[12:34:02] =================== [PASSED] ttm_device ====================
[12:34:02] ================== ttm_pool (8 subtests) ===================
[12:34:02] ================== ttm_pool_alloc_basic ===================
[12:34:02] [PASSED] One page
[12:34:02] [PASSED] More than one page
[12:34:02] [PASSED] Above the allocation limit
[12:34:02] [PASSED] One page, with coherent DMA mappings enabled
[12:34:02] [PASSED] Above the allocation limit, with coherent DMA mappings enabled
[12:34:02] ============== [PASSED] ttm_pool_alloc_basic ===============
[12:34:02] ============== ttm_pool_alloc_basic_dma_addr ==============
[12:34:02] [PASSED] One page
[12:34:02] [PASSED] More than one page
[12:34:02] [PASSED] Above the allocation limit
[12:34:02] [PASSED] One page, with coherent DMA mappings enabled
[12:34:02] [PASSED] Above the allocation limit, with coherent DMA mappings enabled
[12:34:02] ========== [PASSED] ttm_pool_alloc_basic_dma_addr ==========
[12:34:02] [PASSED] ttm_pool_alloc_order_caching_match
[12:34:02] [PASSED] ttm_pool_alloc_caching_mismatch
[12:34:02] [PASSED] ttm_pool_alloc_order_mismatch
[12:34:02] [PASSED] ttm_pool_free_dma_alloc
[12:34:02] [PASSED] ttm_pool_free_no_dma_alloc
[12:34:02] [PASSED] ttm_pool_fini_basic
[12:34:02] ==================== [PASSED] ttm_pool =====================
[12:34:02] ================ ttm_resource (8 subtests) =================
[12:34:02] ================= ttm_resource_init_basic =================
[12:34:02] [PASSED] Init resource in TTM_PL_SYSTEM
[12:34:02] [PASSED] Init resource in TTM_PL_VRAM
[12:34:02] [PASSED] Init resource in a private placement
[12:34:02] [PASSED] Init resource in TTM_PL_SYSTEM, set placement flags
[12:34:02] ============= [PASSED] ttm_resource_init_basic =============
[12:34:02] [PASSED] ttm_resource_init_pinned
[12:34:02] [PASSED] ttm_resource_fini_basic
[12:34:02] [PASSED] ttm_resource_manager_init_basic
[12:34:02] [PASSED] ttm_resource_manager_usage_basic
[12:34:02] [PASSED] ttm_resource_manager_set_used_basic
[12:34:02] [PASSED] ttm_sys_man_alloc_basic
[12:34:02] [PASSED] ttm_sys_man_free_basic
[12:34:02] ================== [PASSED] ttm_resource ===================
[12:34:02] =================== ttm_tt (15 subtests) ===================
[12:34:02] ==================== ttm_tt_init_basic ====================
[12:34:02] [PASSED] Page-aligned size
[12:34:02] [PASSED] Extra pages requested
[12:34:02] ================ [PASSED] ttm_tt_init_basic ================
[12:34:02] [PASSED] ttm_tt_init_misaligned
[12:34:02] [PASSED] ttm_tt_fini_basic
[12:34:02] [PASSED] ttm_tt_fini_sg
[12:34:02] [PASSED] ttm_tt_fini_shmem
[12:34:02] [PASSED] ttm_tt_create_basic
[12:34:02] [PASSED] ttm_tt_create_invalid_bo_type
[12:34:02] [PASSED] ttm_tt_create_ttm_exists
[12:34:02] [PASSED] ttm_tt_create_failed
[12:34:02] [PASSED] ttm_tt_destroy_basic
[12:34:02] [PASSED] ttm_tt_populate_null_ttm
[12:34:02] [PASSED] ttm_tt_populate_populated_ttm
[12:34:02] [PASSED] ttm_tt_unpopulate_basic
[12:34:02] [PASSED] ttm_tt_unpopulate_empty_ttm
[12:34:02] [PASSED] ttm_tt_swapin_basic
[12:34:02] ===================== [PASSED] ttm_tt ======================
[12:34:02] =================== ttm_bo (14 subtests) ===================
[12:34:02] =========== ttm_bo_reserve_optimistic_no_ticket ===========
[12:34:02] [PASSED] Cannot be interrupted and sleeps
[12:34:02] [PASSED] Cannot be interrupted, locks straight away
[12:34:02] [PASSED] Can be interrupted, sleeps
[12:34:02] ======= [PASSED] ttm_bo_reserve_optimistic_no_ticket =======
[12:34:02] [PASSED] ttm_bo_reserve_locked_no_sleep
[12:34:02] [PASSED] ttm_bo_reserve_no_wait_ticket
[12:34:02] [PASSED] ttm_bo_reserve_double_resv
[12:34:02] [PASSED] ttm_bo_reserve_interrupted
[12:34:02] [PASSED] ttm_bo_reserve_deadlock
[12:34:02] [PASSED] ttm_bo_unreserve_basic
[12:34:02] [PASSED] ttm_bo_unreserve_pinned
[12:34:02] [PASSED] ttm_bo_unreserve_bulk
[12:34:02] [PASSED] ttm_bo_fini_basic
[12:34:02] [PASSED] ttm_bo_fini_shared_resv
[12:34:02] [PASSED] ttm_bo_pin_basic
[12:34:02] [PASSED] ttm_bo_pin_unpin_resource
[12:34:02] [PASSED] ttm_bo_multiple_pin_one_unpin
[12:34:02] ===================== [PASSED] ttm_bo ======================
[12:34:02] ============== ttm_bo_validate (22 subtests) ===============
[12:34:02] ============== ttm_bo_init_reserved_sys_man ===============
[12:34:02] [PASSED] Buffer object for userspace
[12:34:02] [PASSED] Kernel buffer object
[12:34:02] [PASSED] Shared buffer object
[12:34:02] ========== [PASSED] ttm_bo_init_reserved_sys_man ===========
[12:34:02] ============== ttm_bo_init_reserved_mock_man ==============
[12:34:02] [PASSED] Buffer object for userspace
[12:34:02] [PASSED] Kernel buffer object
[12:34:02] [PASSED] Shared buffer object
[12:34:02] ========== [PASSED] ttm_bo_init_reserved_mock_man ==========
[12:34:02] [PASSED] ttm_bo_init_reserved_resv
[12:34:02] ================== ttm_bo_validate_basic ==================
[12:34:02] [PASSED] Buffer object for userspace
[12:34:02] [PASSED] Kernel buffer object
[12:34:02] [PASSED] Shared buffer object
[12:34:02] ============== [PASSED] ttm_bo_validate_basic ==============
[12:34:02] [PASSED] ttm_bo_validate_invalid_placement
[12:34:02] ============= ttm_bo_validate_same_placement ==============
[12:34:02] [PASSED] System manager
[12:34:02] [PASSED] VRAM manager
[12:34:02] ========= [PASSED] ttm_bo_validate_same_placement ==========
[12:34:02] [PASSED] ttm_bo_validate_failed_alloc
[12:34:02] [PASSED] ttm_bo_validate_pinned
[12:34:02] [PASSED] ttm_bo_validate_busy_placement
[12:34:02] ================ ttm_bo_validate_multihop =================
[12:34:02] [PASSED] Buffer object for userspace
[12:34:02] [PASSED] Kernel buffer object
[12:34:02] [PASSED] Shared buffer object
[12:34:02] ============ [PASSED] ttm_bo_validate_multihop =============
[12:34:02] ========== ttm_bo_validate_no_placement_signaled ==========
[12:34:02] [PASSED] Buffer object in system domain, no page vector
[12:34:02] [PASSED] Buffer object in system domain with an existing page vector
[12:34:02] ====== [PASSED] ttm_bo_validate_no_placement_signaled ======
[12:34:02] ======== ttm_bo_validate_no_placement_not_signaled ========
[12:34:02] [PASSED] Buffer object for userspace
[12:34:02] [PASSED] Kernel buffer object
[12:34:02] [PASSED] Shared buffer object
[12:34:02] ==== [PASSED] ttm_bo_validate_no_placement_not_signaled ====
[12:34:02] [PASSED] ttm_bo_validate_move_fence_signaled
[12:34:02] ========= ttm_bo_validate_move_fence_not_signaled =========
[12:34:02] [PASSED] Waits for GPU
[12:34:02] [PASSED] Tries to lock straight away
[12:34:02] ===== [PASSED] ttm_bo_validate_move_fence_not_signaled =====
[12:34:02] [PASSED] ttm_bo_validate_swapout
[12:34:02] [PASSED] ttm_bo_validate_happy_evict
[12:34:02] [PASSED] ttm_bo_validate_all_pinned_evict
[12:34:02] [PASSED] ttm_bo_validate_allowed_only_evict
[12:34:02] [PASSED] ttm_bo_validate_deleted_evict
[12:34:02] [PASSED] ttm_bo_validate_busy_domain_evict
[12:34:02] [PASSED] ttm_bo_validate_evict_gutting
[12:34:02] [PASSED] ttm_bo_validate_recrusive_evict
[12:34:02] ================= [PASSED] ttm_bo_validate =================
[12:34:02] ============================================================
[12:34:02] Testing complete. Ran 102 tests: passed: 102
[12:34:02] Elapsed time: 11.969s total, 1.791s configuring, 9.963s building, 0.184s running
+ cleanup
++ stat -c %u:%g /kernel
+ chown -R 1003:1003 /kernel
^ permalink raw reply [flat|nested] 25+ messages in thread
* ✗ Xe.CI.BAT: failure for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2)
2026-08-04 2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
` (7 preceding siblings ...)
2026-08-05 12:34 ` ✓ CI.KUnit: success " Patchwork
@ 2026-08-05 13:11 ` Patchwork
2026-08-05 23:39 ` ✗ Xe.CI.FULL: " Patchwork
9 siblings, 0 replies; 25+ messages in thread
From: Patchwork @ 2026-08-05 13:11 UTC (permalink / raw)
To: Tales A. Mendonça; +Cc: intel-xe
[-- Attachment #1: Type: text/plain, Size: 5833 bytes --]
== Series Details ==
Series: drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2)
URL : https://patchwork.freedesktop.org/series/171539/
State : failure
== Summary ==
CI Bug Log - changes from xe-5545-4c13a67db118ee0019231dbc1362bcb8115ac660_BAT -> xe-pw-171539v2_BAT
====================================================
Summary
-------
**WARNING**
Minor unknown changes coming with xe-pw-171539v2_BAT need to be verified
manually.
If you think the reported changes have nothing to do with the changes
introduced in xe-pw-171539v2_BAT, please notify your bug team (I915-ci-infra@lists.freedesktop.org) to allow them
to document this new failure mode, which will reduce false positives in CI.
Participating hosts (14 -> 15)
------------------------------
Additional (1): bat-bmg-2
Possible new issues
-------------------
Here are the unknown changes that may have been introduced in xe-pw-171539v2_BAT:
### IGT changes ###
#### Warnings ####
* igt@kms_frontbuffer_tracking@basic:
- bat-nvls-2: [SKIP][1] ([Intel XE#7779]) -> [SKIP][2]
[1]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-5545-4c13a67db118ee0019231dbc1362bcb8115ac660/bat-nvls-2/igt@kms_frontbuffer_tracking@basic.html
[2]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/bat-nvls-2/igt@kms_frontbuffer_tracking@basic.html
Known issues
------------
Here are the changes found in xe-pw-171539v2_BAT that come from known issues:
### IGT changes ###
#### Issues hit ####
* igt@fbdev@write:
- bat-bmg-2: NOTRUN -> [SKIP][3] ([Intel XE#2134]) +4 other tests skip
[3]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/bat-bmg-2/igt@fbdev@write.html
* igt@kms_addfb_basic@addfb25-y-tiled-small-legacy:
- bat-bmg-2: NOTRUN -> [SKIP][4] ([Intel XE#2233])
[4]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/bat-bmg-2/igt@kms_addfb_basic@addfb25-y-tiled-small-legacy.html
* igt@kms_cursor_legacy@basic-flip-after-cursor-legacy:
- bat-bmg-2: NOTRUN -> [SKIP][5] ([Intel XE#2489] / [Intel XE#3419]) +13 other tests skip
[5]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/bat-bmg-2/igt@kms_cursor_legacy@basic-flip-after-cursor-legacy.html
* igt@kms_flip@basic-flip-vs-modeset:
- bat-bmg-2: NOTRUN -> [SKIP][6] ([Intel XE#2482]) +3 other tests skip
[6]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/bat-bmg-2/igt@kms_flip@basic-flip-vs-modeset.html
* igt@kms_frontbuffer_tracking@basic:
- bat-bmg-2: NOTRUN -> [SKIP][7] ([Intel XE#2434] / [Intel XE#2548] / [Intel XE#6314])
[7]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/bat-bmg-2/igt@kms_frontbuffer_tracking@basic.html
* igt@kms_psr@psr-sprite-plane-onoff:
- bat-bmg-2: NOTRUN -> [SKIP][8] ([Intel XE#2234] / [Intel XE#2850]) +2 other tests skip
[8]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/bat-bmg-2/igt@kms_psr@psr-sprite-plane-onoff.html
* igt@xe_exec_multi_queue@priority:
- bat-bmg-2: NOTRUN -> [SKIP][9] ([Intel XE#8364]) +13 other tests skip
[9]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/bat-bmg-2/igt@xe_exec_multi_queue@priority.html
* igt@xe_live_ktest@xe_bo@xe_ccs_migrate_kunit:
- bat-bmg-2: NOTRUN -> [SKIP][10] ([Intel XE#2229])
[10]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/bat-bmg-2/igt@xe_live_ktest@xe_bo@xe_ccs_migrate_kunit.html
* igt@xe_pat@pat-index-xehpc:
- bat-bmg-2: NOTRUN -> [SKIP][11] ([Intel XE#1420] / [Intel XE#7590])
[11]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/bat-bmg-2/igt@xe_pat@pat-index-xehpc.html
* igt@xe_pat@pat-index-xelp:
- bat-bmg-2: NOTRUN -> [SKIP][12] ([Intel XE#2245] / [Intel XE#7590])
[12]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/bat-bmg-2/igt@xe_pat@pat-index-xelp.html
* igt@xe_pat@pat-index-xelpg:
- bat-bmg-2: NOTRUN -> [SKIP][13] ([Intel XE#2236] / [Intel XE#7590])
[13]: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/bat-bmg-2/igt@xe_pat@pat-index-xelpg.html
[Intel XE#1420]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/1420
[Intel XE#2134]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2134
[Intel XE#2229]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2229
[Intel XE#2233]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2233
[Intel XE#2234]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2234
[Intel XE#2236]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2236
[Intel XE#2245]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2245
[Intel XE#2434]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2434
[Intel XE#2482]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2482
[Intel XE#2489]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2489
[Intel XE#2548]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2548
[Intel XE#2850]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/2850
[Intel XE#3419]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/3419
[Intel XE#6314]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/6314
[Intel XE#7590]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7590
[Intel XE#7779]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/7779
[Intel XE#8364]: https://gitlab.freedesktop.org/drm/xe/kernel/issues/8364
Build changes
-------------
* Linux: xe-5545-4c13a67db118ee0019231dbc1362bcb8115ac660 -> xe-pw-171539v2
IGT_9038: 9038
xe-5545-4c13a67db118ee0019231dbc1362bcb8115ac660: 4c13a67db118ee0019231dbc1362bcb8115ac660
xe-pw-171539v2: 171539v2
== Logs ==
For more details see: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/index.html
[-- Attachment #2: Type: text/html, Size: 6775 bytes --]
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-04 21:29 ` Matthew Brost
@ 2026-08-05 20:24 ` Matthew Brost
2026-08-06 17:37 ` Tales A. Mendonça
0 siblings, 1 reply; 25+ messages in thread
From: Matthew Brost @ 2026-08-05 20:24 UTC (permalink / raw)
To: Tales A. Mendonça
Cc: intel-xe, thomas.hellstrom, rodrigo.vivi, dri-devel
On Tue, Aug 04, 2026 at 02:29:27PM -0700, Matthew Brost wrote:
> On Tue, Aug 04, 2026 at 01:34:19PM -0300, Tales A. Mendonça wrote:
> > Patchwork reports my address is not on the CI allowlist, so CI was not
> > triggered for this series:
> >
> > Series author address 'talesam@gmail.com' is not on the allowlist,
> > which prevents CI from being automatically triggered.
> >
> > Could one of the project owners click 'retest' on the series (and/or
> > add me to the allowlist)? Series URL:
> >
> > https://patchwork.freedesktop.org/series/171539/
> >
>
> We'd have to resend this ourselves. I can do this, but I've requested
> for you to be on our allow list as well. I'll ping here once that goes
> through.
>
You are approved on our CI for future patches.
Also I came across this issue in the i915 for MTL [1] which seems to
indicate the same issue (ARL and MTL are very close and share same GuC
firmware), the suggestion there is turn off rc6. Unfortunately Xe
doesn't have a knob to do this but according to [1] you can turn off rc6
in the BIOS. Might be worth a try.
Matt
[1] https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/14469
> Matt
>
> > Thanks!
> > Tales
> >
> >
> > Em seg., 3 de ago. de 2026 às 23:14, Tales A. Mendonça
> > <talesam@gmail.com> escreveu:
> > >
> > > Hi,
> > >
> > > This series is a follow-up to the TLB invalidation ack stall I have
> > > been debugging on ARL, tracked in:
> > >
> > > https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> > >
> > > Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51 and
> > > 7dd1 machines here, plus an independent Arc Pro 130T report on the
> > > issue above), TLB invalidation acks intermittently stall for ~2.3s.
> > > The H2G request is consumed from the CTB immediately and the G2H CTB
> > > is empty the whole time - the firmware simply does not send the ack
> > > until much later. The fence timeout fires at 2.25s and the ack lands
> > > tens of ms after it. Userspace blocked on the invalidation (compositor
> > > buffer unmaps etc.) hitches for the full window.
> > >
> > > Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has no
> > > exec queue or job to blame - leaves a devcoredump with the GuC log and
> > > CT state behind (Matt suggested capturing devcoredumps when we
> > > discussed the issue; devcoredumps from both machines are attached to
> > > the issue above).
> > >
> > > Patch 2 logs when the ack for a timed out invalidation finally
> > > arrives. This is what established that the acks are late rather than
> > > lost.
> > >
> > > Patch 3 is the RFC part: a delayed work that pokes the GuC (status
> > > register read, CT flush, doorbell ring) every 250ms while an ack is
> > > overdue. On my machines this converts the guaranteed 2.3s stall into a
> > > sub-500ms hiccup for the majority of occurrences; a minority of severe
> > > episodes ignore 8-9 consecutive doorbells, which points at the GuC
> > > firmware being internally blocked for the whole window. Full data on
> > > the issue. I am happy to rework the approach (different delay,
> > > tying it to the G2H handler, dropping the status read, etc.) - mainly
> > > I would like the firmware side investigated, since no host-side poke
> > > can fix the severe cases.
> > >
> > > Based on drm-tip. Tested for several days on both ARL machines under
> > > desktop and VM-heavy workloads.
> > >
> > > Thanks,
> > > Tales
> > >
> > > Tales A. Mendonça (3):
> > > drm/xe: Capture devcoredump on TLB invalidation timeout
> > > drm/xe: Log when a timed out TLB invalidation ack finally arrives
> > > drm/xe: Kick GuC while TLB invalidation acks are overdue
> > >
> > > drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
> > > drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
> > > drivers/gpu/drm/xe/xe_tlb_inval.c | 131 +++++++++++++++++++++++-
> > > drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
> > > 4 files changed, 243 insertions(+), 4 deletions(-)
> > >
> > > --
> > > 2.55.0
> > >
> >
> >
> > --
> > Com os cumprimentos,
> >
> > Tales A. Mendonça
> > talesam.org
> > communitybig.org
^ permalink raw reply [flat|nested] 25+ messages in thread
* ✗ Xe.CI.FULL: failure for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2)
2026-08-04 2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
` (8 preceding siblings ...)
2026-08-05 13:11 ` ✗ Xe.CI.BAT: failure " Patchwork
@ 2026-08-05 23:39 ` Patchwork
9 siblings, 0 replies; 25+ messages in thread
From: Patchwork @ 2026-08-05 23:39 UTC (permalink / raw)
To: Tales A. Mendonça; +Cc: intel-xe
[-- Attachment #1: Type: text/plain, Size: 415 bytes --]
== Series Details ==
Series: drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2)
URL : https://patchwork.freedesktop.org/series/171539/
State : failure
== Summary ==
ERROR: The runconfig 'xe-5545-4c13a67db118ee0019231dbc1362bcb8115ac660_FULL' does not exist in the database
== Logs ==
For more details see: https://intel-gfx-ci.01.org/tree/intel-xe/xe-pw-171539v2/index.html
[-- Attachment #2: Type: text/html, Size: 980 bytes --]
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-04 23:50 ` Daniele Ceraolo Spurio
@ 2026-08-06 17:36 ` Tales A. Mendonça
2026-08-06 21:13 ` Daniele Ceraolo Spurio
0 siblings, 1 reply; 25+ messages in thread
From: Tales A. Mendonça @ 2026-08-06 17:36 UTC (permalink / raw)
To: Daniele Ceraolo Spurio
Cc: Summers, Stuart, Brost, Matthew, intel-xe@lists.freedesktop.org,
dri-devel@lists.freedesktop.org, Vivi, Rodrigo,
thomas.hellstrom@linux.intel.com, Filipchuk, Julia
> That would definitely help, because if the issue does not happen on i915
> it likely means that we're missing a WA or something like that in Xe.
Two data points on that, pulling in different directions:
* My second ARL machine (7dd1) has been running i915 for 7 days now,
with the same GuC 70.53.0, and there is not a single TLB invalidation
timeout or GuC error in its logs. Caveat: i915's TLB invalidation
timeout is longer than xe's 2.25s, so short stalls could be silent
there.
* However, Matt just pointed at i915 MTL issue 14469, which looks like
the same problem on i915 - so it may not be xe-specific after all.
If it helps I can also boot i915 on the primary machine (7d51), where I
can compare against days of xe statistics on identical workloads.
> Just a bit of a terminology update here, to make sure we're on the same
> page: we usually refer to the notification you're sending to the GuC as
> an H2G interrupt and not a doorbell.
Thanks for the correction - I will use H2G interrupt from here on and
fix the terminology in v2.
> It feels like when the
> issue occurs something is stuck in HW rather than GuC FW and triggering
> the interrupt causes the HW to get unstuck.
That fits a pattern I can now see clearly with more data. Since
enabling the bigger GuC logs (~1.5 days, 35 stalls): 12 stalls were
unstuck by one of the H2G interrupts within 0.3-1.5s, but 23 ignored
8-9 consecutive interrupts and ran to the end. And in those severe
cases the request-to-ack time is nearly constant: 2.28-2.34s, every
single time. It does not look like congestion - it looks like a fixed
internal timeout expiring somewhere and releasing things.
Related: an A/B experiment I ran earlier (holding forcewake across the
whole GT, C6 residency pinned at 0ms for the whole window) still hit 9
timeouts in a row, so GT-level RC6 avoidance alone does not prevent it.
> Also, would you be able to capture the GuC logs when the issue occurs?
Done. I rebuilt with the debug-sized log buffers (8M event data / 1M
crash dump / 1M state capture) and xe.guc_log_level=3, and the series'
patch 1 (devcoredump on TLB invalidation timeout) captures the GuC log
at the exact moment the timeout fires. I attached three devcoredumps
(4-8.7MB each, containing the full GuC log around severe stalls that
ignored 8-9 H2G interrupts) to the gitlab issue:
https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
I have more captures if useful (9 so far).
> I have pushed the latest GuC FW for MTL here in case you want to give it
> a go:
Downloaded and staged - I will switch to 70.72.1 via
xe.guc_firmware_path today and report back with a few days of data
(this machine currently reproduces 20-35 stalls/day under my normal
workload, so the signal should be quick).
Thanks,
Tales
Em ter., 4 de ago. de 2026 às 20:50, Daniele Ceraolo Spurio
<daniele.ceraolospurio@intel.com> escreveu:
>
>
>
> On 8/4/2026 4:00 PM, Tales A. Mendonça wrote:
> >> Are we seeing this on i915?
> > I have not tested i915 on the affected machines yet - the
> > instrumentation that measured the stalls (late-ack logging, kick
> > results) is xe-only, so I have no comparable i915 data. I can boot one
> > of the ARL machines with i915 for a few days and watch for TLB
> > invalidation timeouts there, if that data helps.
>
> That would definitely help, because if the issue does not happen on i915
> it likely means that we're missing a WA or something like that in Xe.
>
> >
> > More generally: both machines here reproduce reliably (~1 stall/hour
> > on a desktop workload, much more under memory pressure), so I am happy
> > to test anything on them - including any GuC build the firmware team
> > would like data on.
> >
> > On Stuart's masking concern: fully agreed, that is why patch 3 is
> > marked RFC. Patches 1-2 are pure diagnostics and stand on their own; I
> > am fine holding patch 3 until the firmware side has been looked at.
> > The data point it adds is that a doorbell ring unblocks the ack in the
> > majority of episodes, while the severe ones ignore 8-9 consecutive
> > rings - hopefully that narrows where to look inside the GuC.
>
> Just a bit of a terminology update here, to make sure we're on the same
> page: we usually refer to the notification you're sending to the GuC as
> an H2G interrupt and not a doorbell. I'm making this clarification
> because the GuC supports a separate per-context notification mechanism
> that is referred to as doorbell and which we currently do not implement
> in neither i915 nor Xe.
>
> When receiving the H2G interrupt, the only thing that the GuC does is
> look into the CTB and process anything in there; however, you've said
> that the contents of the H2G CTB are processed immediately, so the
> follow up interrupt should result in the GuC just bailing out and doing
> nothing because there is no data to process. It feels like when the
> issue occurs something is stuck in HW rather than GuC FW and triggering
> the interrupt causes the HW to get unstuck.
>
> I have pushed the latest GuC FW for MTL here in case you want to give it
> a go:
> https://gitlab.com/dceraolo/drm-firmware/-/blob/f783b931555be057dafc2400b7eb4d445c953fec/i915/mtl_guc_70.72.1.bin
> . You can override the GuC firmware used by the driver via the
> xe.guc_firmware_path modparam; the path is relative to /lib/firmware/
> and the firmware needs to be in initramfs for the driver to find it at
> boot. Note that we haven't tested this image on MTL, so it might have
> unexpected results.
>
> Also, would you be able to capture the GuC logs when the issue occurs?
> The default guc log size is relatively small, so you'd have to capture
> right when the issue happens. However, you can make them bigger by
> building the kernel with CONFIG_DRM_XE_DEBUG or by simply modifying the
> xe_guc_log.h file to pick the bigger size by default. If you go with the
> latter, please also set xe.guc_log_level=3 on the command line (this is
> automatically added by the kconfig).
>
> Thanks,
> Daniele
>
> >
> > I will send a v2 addressing Matt's review comments (the
> > __xe_devcoredump unification and the fixes on patch 2).
> >
> > Thanks,
> > Tales
> >
> > Em ter., 4 de ago. de 2026 às 19:08, Daniele Ceraolo Spurio
> > <daniele.ceraolospurio@intel.com> escreveu:
> >>
> >>
> >> On 8/4/2026 2:33 PM, Summers, Stuart wrote:
> >>> On Tue, 2026-08-04 at 14:27 -0700, Matthew Brost wrote:
> >>>> On Tue, Aug 04, 2026 at 03:02:44PM -0600, Summers, Stuart wrote:
> >>>>> On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
> >>>>>> Hi,
> >>>>>>
> >>>>>> This series is a follow-up to the TLB invalidation ack stall I
> >>>>>> have
> >>>>>> been debugging on ARL, tracked in:
> >>>>>>
> >>>>>> https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> >>>>>>
> >>>>>> Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51
> >>>>>> and
> >>>>>> 7dd1 machines here, plus an independent Arc Pro 130T report on
> >>>>>> the
> >>>>>> issue above), TLB invalidation acks intermittently stall for
> >>>>>> ~2.3s.
> >>>>>> The H2G request is consumed from the CTB immediately and the G2H
> >>>>>> CTB
> >>>>>> is empty the whole time - the firmware simply does not send the
> >>>>>> ack
> >>>>>> until much later. The fence timeout fires at 2.25s and the ack
> >>>>>> lands
> >>>>>> tens of ms after it. Userspace blocked on the invalidation
> >>>>>> (compositor
> >>>>>> buffer unmaps etc.) hitches for the full window.
> >>>>> Firstly, thanks for the patch!
> >>>>>
> >>>>> I haven't looked in to all the details of the sighting you were
> >>>>> debugging, but we have had similar issues that were fixed in a
> >>>>> later
> >>>>> GuC version. I think around 70.60.0? It might be worth trying on
> >>>>> something later than that to see if that helps... (+Daniele)
> >>>>>
> >>>> I think this would require an AR on our end to make a new firmware
> >>>> version available.
> >>>>
> >>>> The upstream repo only has 70.53.0 available for ARL [1] (iirc, ARL
> >>>> aliases to MTL for firmware). (+Julia too).
> >>>>
> >>>> Presumably, the GuC changelogs should indicate whether an issue
> >>>> related
> >>>> this has been fixed. If so, we need to update all GuC versions across
> >>>> both i915 and Xe.
> >>> Right... I guess I'd still like to see if we can test this in GuC (or
> >>> get confirmation we can't for some reason) before committing something.
> >>> My worry is we will prevent bug reports like this by working around it
> >>> and miss critical bugs that need to be fixed in the right component.
> >>>
> >>>> [1]
> >>>> https://gitlab.com/kernel-firmware/linux-firmware/-/blob/main/i915/mtl_guc_70.bin?ref_type=heads
> >>>>
> >>>>>> Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has
> >>>>>> no
> >>>>> Is there a reason we don't just re-use the main xe_devcoredump()?
> >>>>>
> >>>> This is my suggestion: the main devcoredump infrastructure is job-
> >>>> based,
> >>>> so it cannot be used for hangs that are not associated with a job.
> >>>>
> >>>> In my opinion, this is a gap on our end. Introducing something like
> >>>> `xe_devcoredump_gt()`, which can be used for non-job-based hangs
> >>>> (e.g.,
> >>>> TLB invalidation timeouts like those addressed in this series, or
> >>>> more
> >>>> generally any GuC protocol hang), makes sense to me.
> >>> Ok makes sense. We can do that here. It would be nice to have a more
> >>> inclusive implementation that lets us call this from anywhere so we
> >>> aren't duplicating things around for different use cases. But not a
> >>> blocker here.
> >>>
> >>>> I haven't looked at the patch yet, but at a high level, adding
> >>>> `xe_devcoredump_gt()` seems like a reasonable approach.
> >>>>
> >>>>>> exec queue or job to blame - leaves a devcoredump with the GuC
> >>>>>> log
> >>>>>> and
> >>>>>> CT state behind (Matt suggested capturing devcoredumps when we
> >>>>>> discussed the issue; devcoredumps from both machines are attached
> >>>>>> to
> >>>>>> the issue above).
> >>>>>>
> >>>>>> Patch 2 logs when the ack for a timed out invalidation finally
> >>>>>> arrives. This is what established that the acks are late rather
> >>>>>> than
> >>>>>> lost.
> >>>>>>
> >>>>>> Patch 3 is the RFC part: a delayed work that pokes the GuC
> >>>>>> (status
> >>>>>> register read, CT flush, doorbell ring) every 250ms while an ack
> >>>>>> is
> >>>>>> overdue. On my machines this converts the guaranteed 2.3s stall
> >>>>>> into
> >>>>> I'm a little worried we're just papering over something here that
> >>>>> needs
> >>>>> to be addressed in GuC, particularly around GT going to sleep or
> >>>>> something around the time we're expecting a response, so the pings
> >>>>> on
> >>>>> registers might be prematurely waking things up which is something
> >>>>> we'd
> >>>>> want to happen in GuC, not the KMD.
> >>>>>
> >>>> In general, I agree with this. We should avoid papering over the
> >>>> issue
> >>>> and instead fix it properly in the GuC. That said, this workaround
> >>>> provides a pretty strong data point, since it appears to get the TLB
> >>>> invalidation unstuck.
> >>> So if we hit this issue I guess we're already going to have some
> >>> performance degredation and the workaround makes that better. I need to
> >>> look at the implementation, but we could be potentially introducing
> >>> performance penalties in other areas doing these pings.
> >>>
> >>> Again, I'd like to see if we can fix this in the right place before
> >>> implementing a workaround for it. Hopefully Daniele or Julia can give
> >>> some direction there.
> >> Are we seeing this on i915 at all? Given that Xe does not officially
> >> support MTL/ARL and is missing several critical WAs for those platforms,
> >> the approach so far has been to only update the GuC FW if it is required
> >> for i915.
> >> Looking at the GuC release notes, there have been a couple of
> >> TLB-related fixes after 70.53, but they're both marked as only affecting
> >> PVC and Xe2+ platforms, so no fixes seem to be available for ARL (or at
> >> least they're not listed in the release notes).
> >>
> >> Daniele
> >>
> >>> Thanks,
> >>> Stuart
> >>>
> >>>> Matt
> >>>>
> >>>>> Thanks,
> >>>>> Stuart
> >>>>>
> >>>>>> a
> >>>>>> sub-500ms hiccup for the majority of occurrences; a minority of
> >>>>>> severe
> >>>>>> episodes ignore 8-9 consecutive doorbells, which points at the
> >>>>>> GuC
> >>>>>> firmware being internally blocked for the whole window. Full data
> >>>>>> on
> >>>>>> the issue. I am happy to rework the approach (different delay,
> >>>>>> tying it to the G2H handler, dropping the status read, etc.) -
> >>>>>> mainly
> >>>>>> I would like the firmware side investigated, since no host-side
> >>>>>> poke
> >>>>>> can fix the severe cases.
> >>>>>>
> >>>>>> Based on drm-tip. Tested for several days on both ARL machines
> >>>>>> under
> >>>>>> desktop and VM-heavy workloads.
> >>>>>>
> >>>>>> Thanks,
> >>>>>> Tales
> >>>>>>
> >>>>>> Tales A. Mendonça (3):
> >>>>>> drm/xe: Capture devcoredump on TLB invalidation timeout
> >>>>>> drm/xe: Log when a timed out TLB invalidation ack finally
> >>>>>> arrives
> >>>>>> drm/xe: Kick GuC while TLB invalidation acks are overdue
> >>>>>>
> >>>>>> drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
> >>>>>> drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
> >>>>>> drivers/gpu/drm/xe/xe_tlb_inval.c | 131
> >>>>>> +++++++++++++++++++++++-
> >>>>>> drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
> >>>>>> 4 files changed, 243 insertions(+), 4 deletions(-)
> >>>>>>
> >
>
--
Com os cumprimentos,
Tales A. Mendonça
talesam.org
communitybig.org
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-05 20:24 ` Matthew Brost
@ 2026-08-06 17:37 ` Tales A. Mendonça
2026-08-06 17:56 ` Daniele Ceraolo Spurio
0 siblings, 1 reply; 25+ messages in thread
From: Tales A. Mendonça @ 2026-08-06 17:37 UTC (permalink / raw)
To: Matthew Brost; +Cc: intel-xe, thomas.hellstrom, rodrigo.vivi, dri-devel
> You've been approved in our CI for future revs.
Thanks!
> Also I came across this issue on i915 for MTL [1] which appears to
> indicate the same issue (ARL and MTL are very close and share the same GuC
> firmware), the suggestion is to disable rc6. Unfortunately Xe
> doesn't have a knob for this, but per [1] you can turn off rc6
> in the BIOS. Maybe worth trying.
Good find - that does look like the same ~2.3s signature, and it being
visible on i915/MTL too is an important data point (Daniele asked
exactly that in the other subthread).
On the rc6 angle: I ran an A/B experiment earlier that should be
equivalent to disabling rc6 at the GT level - holding forcewake across
the whole GT for hours (C6 residency pinned at 0ms for the entire
window, verified) - and still hit 9 timeouts in a row, with the same
~2.3s request-to-ack. So at least keeping the GT out of RC6 does not
avoid the stall here. If the BIOS suggestion covers more than GT RC6
(e.g. package C-states), that would be a different experiment - my
consumer ASUS BIOS does not expose an rc6 knob, but I can look for
C-state options if you think it is worth isolating.
Next on my side: switching to the GuC 70.72.1 build Daniele posted and
reporting back, plus GuC logs from severe stalls are now attached to
the gitlab issue.
Thanks,
Tales
Em qua., 5 de ago. de 2026 às 17:25, Matthew Brost
<matthew.brost@intel.com> escreveu:
>
> On Tue, Aug 04, 2026 at 02:29:27PM -0700, Matthew Brost wrote:
> > On Tue, Aug 04, 2026 at 01:34:19PM -0300, Tales A. Mendonça wrote:
> > > Patchwork reports my address is not on the CI allowlist, so CI was not
> > > triggered for this series:
> > >
> > > Series author address 'talesam@gmail.com' is not on the allowlist,
> > > which prevents CI from being automatically triggered.
> > >
> > > Could one of the project owners click 'retest' on the series (and/or
> > > add me to the allowlist)? Series URL:
> > >
> > > https://patchwork.freedesktop.org/series/171539/
> > >
> >
> > We'd have to resend this ourselves. I can do this, but I've requested
> > for you to be on our allow list as well. I'll ping here once that goes
> > through.
> >
>
> You are approved on our CI for future patches.
>
> Also I came across this issue in the i915 for MTL [1] which seems to
> indicate the same issue (ARL and MTL are very close and share same GuC
> firmware), the suggestion there is turn off rc6. Unfortunately Xe
> doesn't have a knob to do this but according to [1] you can turn off rc6
> in the BIOS. Might be worth a try.
>
> Matt
>
> [1] https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/14469
>
> > Matt
> >
> > > Thanks!
> > > Tales
> > >
> > >
> > > Em seg., 3 de ago. de 2026 às 23:14, Tales A. Mendonça
> > > <talesam@gmail.com> escreveu:
> > > >
> > > > Hi,
> > > >
> > > > This series is a follow-up to the TLB invalidation ack stall I have
> > > > been debugging on ARL, tracked in:
> > > >
> > > > https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> > > >
> > > > Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51 and
> > > > 7dd1 machines here, plus an independent Arc Pro 130T report on the
> > > > issue above), TLB invalidation acks intermittently stall for ~2.3s.
> > > > The H2G request is consumed from the CTB immediately and the G2H CTB
> > > > is empty the whole time - the firmware simply does not send the ack
> > > > until much later. The fence timeout fires at 2.25s and the ack lands
> > > > tens of ms after it. Userspace blocked on the invalidation (compositor
> > > > buffer unmaps etc.) hitches for the full window.
> > > >
> > > > Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has no
> > > > exec queue or job to blame - leaves a devcoredump with the GuC log and
> > > > CT state behind (Matt suggested capturing devcoredumps when we
> > > > discussed the issue; devcoredumps from both machines are attached to
> > > > the issue above).
> > > >
> > > > Patch 2 logs when the ack for a timed out invalidation finally
> > > > arrives. This is what established that the acks are late rather than
> > > > lost.
> > > >
> > > > Patch 3 is the RFC part: a delayed work that pokes the GuC (status
> > > > register read, CT flush, doorbell ring) every 250ms while an ack is
> > > > overdue. On my machines this converts the guaranteed 2.3s stall into a
> > > > sub-500ms hiccup for the majority of occurrences; a minority of severe
> > > > episodes ignore 8-9 consecutive doorbells, which points at the GuC
> > > > firmware being internally blocked for the whole window. Full data on
> > > > the issue. I am happy to rework the approach (different delay,
> > > > tying it to the G2H handler, dropping the status read, etc.) - mainly
> > > > I would like the firmware side investigated, since no host-side poke
> > > > can fix the severe cases.
> > > >
> > > > Based on drm-tip. Tested for several days on both ARL machines under
> > > > desktop and VM-heavy workloads.
> > > >
> > > > Thanks,
> > > > Tales
> > > >
> > > > Tales A. Mendonça (3):
> > > > drm/xe: Capture devcoredump on TLB invalidation timeout
> > > > drm/xe: Log when a timed out TLB invalidation ack finally arrives
> > > > drm/xe: Kick GuC while TLB invalidation acks are overdue
> > > >
> > > > drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
> > > > drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
> > > > drivers/gpu/drm/xe/xe_tlb_inval.c | 131 +++++++++++++++++++++++-
> > > > drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
> > > > 4 files changed, 243 insertions(+), 4 deletions(-)
> > > >
> > > > --
> > > > 2.55.0
> > > >
> > >
> > >
> > > --
> > > Com os cumprimentos,
> > >
> > > Tales A. Mendonça
> > > talesam.org
> > > communitybig.org
--
Com os cumprimentos,
Tales A. Mendonça
talesam.org
communitybig.org
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-06 17:37 ` Tales A. Mendonça
@ 2026-08-06 17:56 ` Daniele Ceraolo Spurio
0 siblings, 0 replies; 25+ messages in thread
From: Daniele Ceraolo Spurio @ 2026-08-06 17:56 UTC (permalink / raw)
To: Tales A. Mendonça, Matthew Brost
Cc: intel-xe, thomas.hellstrom, rodrigo.vivi, dri-devel
On 8/6/2026 10:37 AM, Tales A. Mendonça wrote:
>> You've been approved in our CI for future revs.
> Thanks!
>
>> Also I came across this issue on i915 for MTL [1] which appears to
>> indicate the same issue (ARL and MTL are very close and share the same GuC
>> firmware), the suggestion is to disable rc6. Unfortunately Xe
>> doesn't have a knob for this, but per [1] you can turn off rc6
>> in the BIOS. Maybe worth trying.
> Good find - that does look like the same ~2.3s signature, and it being
> visible on i915/MTL too is an important data point (Daniele asked
> exactly that in the other subthread).
I am not sure if 14469 is the same issue. In that one it seems like the
GuC just stops processing incoming messages (i.e., the H2G CTB has
unprocessed data in it), while AFAIU in your case the GuC is still
processing new commands and it is just being slow. It might be different
manifestations of the same underlying issue, but it might also be
completely separate bugs.
>
> On the rc6 angle: I ran an A/B experiment earlier that should be
> equivalent to disabling rc6 at the GT level - holding forcewake across
> the whole GT for hours (C6 residency pinned at 0ms for the entire
> window, verified) - and still hit 9 timeouts in a row, with the same
> ~2.3s request-to-ack. So at least keeping the GT out of RC6 does not
> avoid the stall here. If the BIOS suggestion covers more than GT RC6
> (e.g. package C-states), that would be a different experiment - my
> consumer ASUS BIOS does not expose an rc6 knob, but I can look for
> C-state options if you think it is worth isolating.
The fact that keeping the GT awake didn't help also indicates that this
likely isn't the same 14469.
Daniele
>
> Next on my side: switching to the GuC 70.72.1 build Daniele posted and
> reporting back, plus GuC logs from severe stalls are now attached to
> the gitlab issue.
>
> Thanks,
> Tales
>
> Em qua., 5 de ago. de 2026 às 17:25, Matthew Brost
> <matthew.brost@intel.com> escreveu:
>> On Tue, Aug 04, 2026 at 02:29:27PM -0700, Matthew Brost wrote:
>>> On Tue, Aug 04, 2026 at 01:34:19PM -0300, Tales A. Mendonça wrote:
>>>> Patchwork reports my address is not on the CI allowlist, so CI was not
>>>> triggered for this series:
>>>>
>>>> Series author address 'talesam@gmail.com' is not on the allowlist,
>>>> which prevents CI from being automatically triggered.
>>>>
>>>> Could one of the project owners click 'retest' on the series (and/or
>>>> add me to the allowlist)? Series URL:
>>>>
>>>> https://patchwork.freedesktop.org/series/171539/
>>>>
>>> We'd have to resend this ourselves. I can do this, but I've requested
>>> for you to be on our allow list as well. I'll ping here once that goes
>>> through.
>>>
>> You are approved on our CI for future patches.
>>
>> Also I came across this issue in the i915 for MTL [1] which seems to
>> indicate the same issue (ARL and MTL are very close and share same GuC
>> firmware), the suggestion there is turn off rc6. Unfortunately Xe
>> doesn't have a knob to do this but according to [1] you can turn off rc6
>> in the BIOS. Might be worth a try.
>>
>> Matt
>>
>> [1] https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/14469
>>
>>> Matt
>>>
>>>> Thanks!
>>>> Tales
>>>>
>>>>
>>>> Em seg., 3 de ago. de 2026 às 23:14, Tales A. Mendonça
>>>> <talesam@gmail.com> escreveu:
>>>>> Hi,
>>>>>
>>>>> This series is a follow-up to the TLB invalidation ack stall I have
>>>>> been debugging on ARL, tracked in:
>>>>>
>>>>> https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
>>>>>
>>>>> Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51 and
>>>>> 7dd1 machines here, plus an independent Arc Pro 130T report on the
>>>>> issue above), TLB invalidation acks intermittently stall for ~2.3s.
>>>>> The H2G request is consumed from the CTB immediately and the G2H CTB
>>>>> is empty the whole time - the firmware simply does not send the ack
>>>>> until much later. The fence timeout fires at 2.25s and the ack lands
>>>>> tens of ms after it. Userspace blocked on the invalidation (compositor
>>>>> buffer unmaps etc.) hitches for the full window.
>>>>>
>>>>> Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has no
>>>>> exec queue or job to blame - leaves a devcoredump with the GuC log and
>>>>> CT state behind (Matt suggested capturing devcoredumps when we
>>>>> discussed the issue; devcoredumps from both machines are attached to
>>>>> the issue above).
>>>>>
>>>>> Patch 2 logs when the ack for a timed out invalidation finally
>>>>> arrives. This is what established that the acks are late rather than
>>>>> lost.
>>>>>
>>>>> Patch 3 is the RFC part: a delayed work that pokes the GuC (status
>>>>> register read, CT flush, doorbell ring) every 250ms while an ack is
>>>>> overdue. On my machines this converts the guaranteed 2.3s stall into a
>>>>> sub-500ms hiccup for the majority of occurrences; a minority of severe
>>>>> episodes ignore 8-9 consecutive doorbells, which points at the GuC
>>>>> firmware being internally blocked for the whole window. Full data on
>>>>> the issue. I am happy to rework the approach (different delay,
>>>>> tying it to the G2H handler, dropping the status read, etc.) - mainly
>>>>> I would like the firmware side investigated, since no host-side poke
>>>>> can fix the severe cases.
>>>>>
>>>>> Based on drm-tip. Tested for several days on both ARL machines under
>>>>> desktop and VM-heavy workloads.
>>>>>
>>>>> Thanks,
>>>>> Tales
>>>>>
>>>>> Tales A. Mendonça (3):
>>>>> drm/xe: Capture devcoredump on TLB invalidation timeout
>>>>> drm/xe: Log when a timed out TLB invalidation ack finally arrives
>>>>> drm/xe: Kick GuC while TLB invalidation acks are overdue
>>>>>
>>>>> drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
>>>>> drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
>>>>> drivers/gpu/drm/xe/xe_tlb_inval.c | 131 +++++++++++++++++++++++-
>>>>> drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
>>>>> 4 files changed, 243 insertions(+), 4 deletions(-)
>>>>>
>>>>> --
>>>>> 2.55.0
>>>>>
>>>>
>>>> --
>>>> Com os cumprimentos,
>>>>
>>>> Tales A. Mendonça
>>>> talesam.org
>>>> communitybig.org
>
>
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-06 17:36 ` Tales A. Mendonça
@ 2026-08-06 21:13 ` Daniele Ceraolo Spurio
2026-08-08 0:20 ` Tales A. Mendonça
0 siblings, 1 reply; 25+ messages in thread
From: Daniele Ceraolo Spurio @ 2026-08-06 21:13 UTC (permalink / raw)
To: Tales A. Mendonça
Cc: Summers, Stuart, Brost, Matthew, intel-xe@lists.freedesktop.org,
dri-devel@lists.freedesktop.org, Vivi, Rodrigo,
thomas.hellstrom@linux.intel.com, Filipchuk, Julia
On 8/6/2026 10:36 AM, Tales A. Mendonça wrote:
>> That would definitely help, because if the issue does not happen on i915
>> it likely means that we're missing a WA or something like that in Xe.
> Two data points on that, pulling in different directions:
>
> * My second ARL machine (7dd1) has been running i915 for 7 days now,
> with the same GuC 70.53.0, and there is not a single TLB invalidation
> timeout or GuC error in its logs. Caveat: i915's TLB invalidation
> timeout is longer than xe's 2.25s, so short stalls could be silent
> there.
>
> * However, Matt just pointed at i915 MTL issue 14469, which looks like
> the same problem on i915 - so it may not be xe-specific after all.
>
> If it helps I can also boot i915 on the primary machine (7d51), where I
> can compare against days of xe statistics on identical workloads.
>
>> Just a bit of a terminology update here, to make sure we're on the same
>> page: we usually refer to the notification you're sending to the GuC as
>> an H2G interrupt and not a doorbell.
> Thanks for the correction - I will use H2G interrupt from here on and
> fix the terminology in v2.
>
>> It feels like when the
>> issue occurs something is stuck in HW rather than GuC FW and triggering
>> the interrupt causes the HW to get unstuck.
> That fits a pattern I can now see clearly with more data. Since
> enabling the bigger GuC logs (~1.5 days, 35 stalls): 12 stalls were
> unstuck by one of the H2G interrupts within 0.3-1.5s, but 23 ignored
> 8-9 consecutive interrupts and ran to the end. And in those severe
> cases the request-to-ack time is nearly constant: 2.28-2.34s, every
> single time. It does not look like congestion - it looks like a fixed
> internal timeout expiring somewhere and releasing things.
>
> Related: an A/B experiment I ran earlier (holding forcewake across the
> whole GT, C6 residency pinned at 0ms for the whole window) still hit 9
> timeouts in a row, so GT-level RC6 avoidance alone does not prevent it.
>
>> Also, would you be able to capture the GuC logs when the issue occurs?
> Done. I rebuilt with the debug-sized log buffers (8M event data / 1M
> crash dump / 1M state capture) and xe.guc_log_level=3, and the series'
> patch 1 (devcoredump on TLB invalidation timeout) captures the GuC log
> at the exact moment the timeout fires. I attached three devcoredumps
> (4-8.7MB each, containing the full GuC log around severe stalls that
> ignored 8-9 H2G interrupts) to the gitlab issue:
>
> https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
>
> I have more captures if useful (9 so far).
Do you happen to have the matching dmesg for those? I decoded the logs
but there are thousands of invalidation calls in them; I looked at the
last few but from the GuC POV they've all been handled quickly. Dmesg
error should log exactly which invalidations were delayed from the
driver POV so I can just look at what happens around those.
Also, I have noticed that the failures seems to all be on GT1. On MTL,
there is a caching bug on GT1 and we do not implement the WA for that in
Xe (Wa_22016122933). Not sure if this is the actual root cause, but the
fact that it only happens on GT1 makes me suspicious (i.e., it is
possible that the GuC is replying in time but CPU doesn't see the reply
because the cache-line is not correctly updated, or vice versa). Not
sure if there is an easy way to implement this in Xe to test, that's not
really my field of expertise; maybe adding XE_BO_FLAG_FORCE_WC or
XE_BO_FLAG_NEEDS_UC to the CTB allocation could work as a quick hack?
But I'd like either Matt or Thomas to confirm.
Daniele
>
>> I have pushed the latest GuC FW for MTL here in case you want to give it
>> a go:
> Downloaded and staged - I will switch to 70.72.1 via
> xe.guc_firmware_path today and report back with a few days of data
> (this machine currently reproduces 20-35 stalls/day under my normal
> workload, so the signal should be quick).
>
> Thanks,
> Tales
>
> Em ter., 4 de ago. de 2026 às 20:50, Daniele Ceraolo Spurio
> <daniele.ceraolospurio@intel.com> escreveu:
>>
>>
>> On 8/4/2026 4:00 PM, Tales A. Mendonça wrote:
>>>> Are we seeing this on i915?
>>> I have not tested i915 on the affected machines yet - the
>>> instrumentation that measured the stalls (late-ack logging, kick
>>> results) is xe-only, so I have no comparable i915 data. I can boot one
>>> of the ARL machines with i915 for a few days and watch for TLB
>>> invalidation timeouts there, if that data helps.
>> That would definitely help, because if the issue does not happen on i915
>> it likely means that we're missing a WA or something like that in Xe.
>>
>>> More generally: both machines here reproduce reliably (~1 stall/hour
>>> on a desktop workload, much more under memory pressure), so I am happy
>>> to test anything on them - including any GuC build the firmware team
>>> would like data on.
>>>
>>> On Stuart's masking concern: fully agreed, that is why patch 3 is
>>> marked RFC. Patches 1-2 are pure diagnostics and stand on their own; I
>>> am fine holding patch 3 until the firmware side has been looked at.
>>> The data point it adds is that a doorbell ring unblocks the ack in the
>>> majority of episodes, while the severe ones ignore 8-9 consecutive
>>> rings - hopefully that narrows where to look inside the GuC.
>> Just a bit of a terminology update here, to make sure we're on the same
>> page: we usually refer to the notification you're sending to the GuC as
>> an H2G interrupt and not a doorbell. I'm making this clarification
>> because the GuC supports a separate per-context notification mechanism
>> that is referred to as doorbell and which we currently do not implement
>> in neither i915 nor Xe.
>>
>> When receiving the H2G interrupt, the only thing that the GuC does is
>> look into the CTB and process anything in there; however, you've said
>> that the contents of the H2G CTB are processed immediately, so the
>> follow up interrupt should result in the GuC just bailing out and doing
>> nothing because there is no data to process. It feels like when the
>> issue occurs something is stuck in HW rather than GuC FW and triggering
>> the interrupt causes the HW to get unstuck.
>>
>> I have pushed the latest GuC FW for MTL here in case you want to give it
>> a go:
>> https://gitlab.com/dceraolo/drm-firmware/-/blob/f783b931555be057dafc2400b7eb4d445c953fec/i915/mtl_guc_70.72.1.bin
>> . You can override the GuC firmware used by the driver via the
>> xe.guc_firmware_path modparam; the path is relative to /lib/firmware/
>> and the firmware needs to be in initramfs for the driver to find it at
>> boot. Note that we haven't tested this image on MTL, so it might have
>> unexpected results.
>>
>> Also, would you be able to capture the GuC logs when the issue occurs?
>> The default guc log size is relatively small, so you'd have to capture
>> right when the issue happens. However, you can make them bigger by
>> building the kernel with CONFIG_DRM_XE_DEBUG or by simply modifying the
>> xe_guc_log.h file to pick the bigger size by default. If you go with the
>> latter, please also set xe.guc_log_level=3 on the command line (this is
>> automatically added by the kconfig).
>>
>> Thanks,
>> Daniele
>>
>>> I will send a v2 addressing Matt's review comments (the
>>> __xe_devcoredump unification and the fixes on patch 2).
>>>
>>> Thanks,
>>> Tales
>>>
>>> Em ter., 4 de ago. de 2026 às 19:08, Daniele Ceraolo Spurio
>>> <daniele.ceraolospurio@intel.com> escreveu:
>>>>
>>>> On 8/4/2026 2:33 PM, Summers, Stuart wrote:
>>>>> On Tue, 2026-08-04 at 14:27 -0700, Matthew Brost wrote:
>>>>>> On Tue, Aug 04, 2026 at 03:02:44PM -0600, Summers, Stuart wrote:
>>>>>>> On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
>>>>>>>> Hi,
>>>>>>>>
>>>>>>>> This series is a follow-up to the TLB invalidation ack stall I
>>>>>>>> have
>>>>>>>> been debugging on ARL, tracked in:
>>>>>>>>
>>>>>>>> https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
>>>>>>>>
>>>>>>>> Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51
>>>>>>>> and
>>>>>>>> 7dd1 machines here, plus an independent Arc Pro 130T report on
>>>>>>>> the
>>>>>>>> issue above), TLB invalidation acks intermittently stall for
>>>>>>>> ~2.3s.
>>>>>>>> The H2G request is consumed from the CTB immediately and the G2H
>>>>>>>> CTB
>>>>>>>> is empty the whole time - the firmware simply does not send the
>>>>>>>> ack
>>>>>>>> until much later. The fence timeout fires at 2.25s and the ack
>>>>>>>> lands
>>>>>>>> tens of ms after it. Userspace blocked on the invalidation
>>>>>>>> (compositor
>>>>>>>> buffer unmaps etc.) hitches for the full window.
>>>>>>> Firstly, thanks for the patch!
>>>>>>>
>>>>>>> I haven't looked in to all the details of the sighting you were
>>>>>>> debugging, but we have had similar issues that were fixed in a
>>>>>>> later
>>>>>>> GuC version. I think around 70.60.0? It might be worth trying on
>>>>>>> something later than that to see if that helps... (+Daniele)
>>>>>>>
>>>>>> I think this would require an AR on our end to make a new firmware
>>>>>> version available.
>>>>>>
>>>>>> The upstream repo only has 70.53.0 available for ARL [1] (iirc, ARL
>>>>>> aliases to MTL for firmware). (+Julia too).
>>>>>>
>>>>>> Presumably, the GuC changelogs should indicate whether an issue
>>>>>> related
>>>>>> this has been fixed. If so, we need to update all GuC versions across
>>>>>> both i915 and Xe.
>>>>> Right... I guess I'd still like to see if we can test this in GuC (or
>>>>> get confirmation we can't for some reason) before committing something.
>>>>> My worry is we will prevent bug reports like this by working around it
>>>>> and miss critical bugs that need to be fixed in the right component.
>>>>>
>>>>>> [1]
>>>>>> https://gitlab.com/kernel-firmware/linux-firmware/-/blob/main/i915/mtl_guc_70.bin?ref_type=heads
>>>>>>
>>>>>>>> Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has
>>>>>>>> no
>>>>>>> Is there a reason we don't just re-use the main xe_devcoredump()?
>>>>>>>
>>>>>> This is my suggestion: the main devcoredump infrastructure is job-
>>>>>> based,
>>>>>> so it cannot be used for hangs that are not associated with a job.
>>>>>>
>>>>>> In my opinion, this is a gap on our end. Introducing something like
>>>>>> `xe_devcoredump_gt()`, which can be used for non-job-based hangs
>>>>>> (e.g.,
>>>>>> TLB invalidation timeouts like those addressed in this series, or
>>>>>> more
>>>>>> generally any GuC protocol hang), makes sense to me.
>>>>> Ok makes sense. We can do that here. It would be nice to have a more
>>>>> inclusive implementation that lets us call this from anywhere so we
>>>>> aren't duplicating things around for different use cases. But not a
>>>>> blocker here.
>>>>>
>>>>>> I haven't looked at the patch yet, but at a high level, adding
>>>>>> `xe_devcoredump_gt()` seems like a reasonable approach.
>>>>>>
>>>>>>>> exec queue or job to blame - leaves a devcoredump with the GuC
>>>>>>>> log
>>>>>>>> and
>>>>>>>> CT state behind (Matt suggested capturing devcoredumps when we
>>>>>>>> discussed the issue; devcoredumps from both machines are attached
>>>>>>>> to
>>>>>>>> the issue above).
>>>>>>>>
>>>>>>>> Patch 2 logs when the ack for a timed out invalidation finally
>>>>>>>> arrives. This is what established that the acks are late rather
>>>>>>>> than
>>>>>>>> lost.
>>>>>>>>
>>>>>>>> Patch 3 is the RFC part: a delayed work that pokes the GuC
>>>>>>>> (status
>>>>>>>> register read, CT flush, doorbell ring) every 250ms while an ack
>>>>>>>> is
>>>>>>>> overdue. On my machines this converts the guaranteed 2.3s stall
>>>>>>>> into
>>>>>>> I'm a little worried we're just papering over something here that
>>>>>>> needs
>>>>>>> to be addressed in GuC, particularly around GT going to sleep or
>>>>>>> something around the time we're expecting a response, so the pings
>>>>>>> on
>>>>>>> registers might be prematurely waking things up which is something
>>>>>>> we'd
>>>>>>> want to happen in GuC, not the KMD.
>>>>>>>
>>>>>> In general, I agree with this. We should avoid papering over the
>>>>>> issue
>>>>>> and instead fix it properly in the GuC. That said, this workaround
>>>>>> provides a pretty strong data point, since it appears to get the TLB
>>>>>> invalidation unstuck.
>>>>> So if we hit this issue I guess we're already going to have some
>>>>> performance degredation and the workaround makes that better. I need to
>>>>> look at the implementation, but we could be potentially introducing
>>>>> performance penalties in other areas doing these pings.
>>>>>
>>>>> Again, I'd like to see if we can fix this in the right place before
>>>>> implementing a workaround for it. Hopefully Daniele or Julia can give
>>>>> some direction there.
>>>> Are we seeing this on i915 at all? Given that Xe does not officially
>>>> support MTL/ARL and is missing several critical WAs for those platforms,
>>>> the approach so far has been to only update the GuC FW if it is required
>>>> for i915.
>>>> Looking at the GuC release notes, there have been a couple of
>>>> TLB-related fixes after 70.53, but they're both marked as only affecting
>>>> PVC and Xe2+ platforms, so no fixes seem to be available for ARL (or at
>>>> least they're not listed in the release notes).
>>>>
>>>> Daniele
>>>>
>>>>> Thanks,
>>>>> Stuart
>>>>>
>>>>>> Matt
>>>>>>
>>>>>>> Thanks,
>>>>>>> Stuart
>>>>>>>
>>>>>>>> a
>>>>>>>> sub-500ms hiccup for the majority of occurrences; a minority of
>>>>>>>> severe
>>>>>>>> episodes ignore 8-9 consecutive doorbells, which points at the
>>>>>>>> GuC
>>>>>>>> firmware being internally blocked for the whole window. Full data
>>>>>>>> on
>>>>>>>> the issue. I am happy to rework the approach (different delay,
>>>>>>>> tying it to the G2H handler, dropping the status read, etc.) -
>>>>>>>> mainly
>>>>>>>> I would like the firmware side investigated, since no host-side
>>>>>>>> poke
>>>>>>>> can fix the severe cases.
>>>>>>>>
>>>>>>>> Based on drm-tip. Tested for several days on both ARL machines
>>>>>>>> under
>>>>>>>> desktop and VM-heavy workloads.
>>>>>>>>
>>>>>>>> Thanks,
>>>>>>>> Tales
>>>>>>>>
>>>>>>>> Tales A. Mendonça (3):
>>>>>>>> drm/xe: Capture devcoredump on TLB invalidation timeout
>>>>>>>> drm/xe: Log when a timed out TLB invalidation ack finally
>>>>>>>> arrives
>>>>>>>> drm/xe: Kick GuC while TLB invalidation acks are overdue
>>>>>>>>
>>>>>>>> drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
>>>>>>>> drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
>>>>>>>> drivers/gpu/drm/xe/xe_tlb_inval.c | 131
>>>>>>>> +++++++++++++++++++++++-
>>>>>>>> drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
>>>>>>>> 4 files changed, 243 insertions(+), 4 deletions(-)
>>>>>>>>
>
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
2026-08-06 21:13 ` Daniele Ceraolo Spurio
@ 2026-08-08 0:20 ` Tales A. Mendonça
0 siblings, 0 replies; 25+ messages in thread
From: Tales A. Mendonça @ 2026-08-08 0:20 UTC (permalink / raw)
To: Daniele Ceraolo Spurio
Cc: Summers, Stuart, Brost, Matthew, intel-xe@lists.freedesktop.org,
dri-devel@lists.freedesktop.org, Vivi, Rodrigo,
thomas.hellstrom@linux.intel.com, Filipchuk, Julia
Results from a full day on GuC 70.72.1 under my normal workload:
59 stall events / 20 fence timeouts in ~20.6h, worst
request-to-ack=2346ms, severe episodes still ignoring 8-9
consecutive H2G interrupts.
Statistically identical to 70.53.0 - the firmware version changes
nothing, which matches your decode (GuC handling everything in time).
I attached to the gitlab issue the full dmesg of the day plus a fresh
devcoredump taken on 70.72.1 (same debug-sized GuC log buffers), in
case you want to confirm the decode on the new firmware too.
That leaves the GT1 cache-line theory (Wa_22016122933) as the main
suspect. I have a kernel built with XE_BO_FLAG_NEEDS_UC added to both
CTB allocations as you suggested (blunt version: both directions, all
GTs - just to test causality) and I am booting it tonight, with the
firmware reverted to stock 70.53.0 so only one variable changes.
Results tomorrow.
If the stalls disappear, I would be happy to help turn this into a
proper Wa_22016122933 implementation for xe, mirroring what i915 does
(scoped to the media GT on the affected platforms) - with guidance
from Matt/Thomas on the preferred shape.
Thanks,
Tales
Em qui., 6 de ago. de 2026 às 18:13, Daniele Ceraolo Spurio
<daniele.ceraolospurio@intel.com> escreveu:
>
>
>
> On 8/6/2026 10:36 AM, Tales A. Mendonça wrote:
> >> That would definitely help, because if the issue does not happen on i915
> >> it likely means that we're missing a WA or something like that in Xe.
> > Two data points on that, pulling in different directions:
> >
> > * My second ARL machine (7dd1) has been running i915 for 7 days now,
> > with the same GuC 70.53.0, and there is not a single TLB invalidation
> > timeout or GuC error in its logs. Caveat: i915's TLB invalidation
> > timeout is longer than xe's 2.25s, so short stalls could be silent
> > there.
> >
> > * However, Matt just pointed at i915 MTL issue 14469, which looks like
> > the same problem on i915 - so it may not be xe-specific after all.
> >
> > If it helps I can also boot i915 on the primary machine (7d51), where I
> > can compare against days of xe statistics on identical workloads.
> >
> >> Just a bit of a terminology update here, to make sure we're on the same
> >> page: we usually refer to the notification you're sending to the GuC as
> >> an H2G interrupt and not a doorbell.
> > Thanks for the correction - I will use H2G interrupt from here on and
> > fix the terminology in v2.
> >
> >> It feels like when the
> >> issue occurs something is stuck in HW rather than GuC FW and triggering
> >> the interrupt causes the HW to get unstuck.
> > That fits a pattern I can now see clearly with more data. Since
> > enabling the bigger GuC logs (~1.5 days, 35 stalls): 12 stalls were
> > unstuck by one of the H2G interrupts within 0.3-1.5s, but 23 ignored
> > 8-9 consecutive interrupts and ran to the end. And in those severe
> > cases the request-to-ack time is nearly constant: 2.28-2.34s, every
> > single time. It does not look like congestion - it looks like a fixed
> > internal timeout expiring somewhere and releasing things.
> >
> > Related: an A/B experiment I ran earlier (holding forcewake across the
> > whole GT, C6 residency pinned at 0ms for the whole window) still hit 9
> > timeouts in a row, so GT-level RC6 avoidance alone does not prevent it.
> >
> >> Also, would you be able to capture the GuC logs when the issue occurs?
> > Done. I rebuilt with the debug-sized log buffers (8M event data / 1M
> > crash dump / 1M state capture) and xe.guc_log_level=3, and the series'
> > patch 1 (devcoredump on TLB invalidation timeout) captures the GuC log
> > at the exact moment the timeout fires. I attached three devcoredumps
> > (4-8.7MB each, containing the full GuC log around severe stalls that
> > ignored 8-9 H2G interrupts) to the gitlab issue:
> >
> > https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> >
> > I have more captures if useful (9 so far).
>
> Do you happen to have the matching dmesg for those? I decoded the logs
> but there are thousands of invalidation calls in them; I looked at the
> last few but from the GuC POV they've all been handled quickly. Dmesg
> error should log exactly which invalidations were delayed from the
> driver POV so I can just look at what happens around those.
> Also, I have noticed that the failures seems to all be on GT1. On MTL,
> there is a caching bug on GT1 and we do not implement the WA for that in
> Xe (Wa_22016122933). Not sure if this is the actual root cause, but the
> fact that it only happens on GT1 makes me suspicious (i.e., it is
> possible that the GuC is replying in time but CPU doesn't see the reply
> because the cache-line is not correctly updated, or vice versa). Not
> sure if there is an easy way to implement this in Xe to test, that's not
> really my field of expertise; maybe adding XE_BO_FLAG_FORCE_WC or
> XE_BO_FLAG_NEEDS_UC to the CTB allocation could work as a quick hack?
> But I'd like either Matt or Thomas to confirm.
>
> Daniele
>
> >
> >> I have pushed the latest GuC FW for MTL here in case you want to give it
> >> a go:
> > Downloaded and staged - I will switch to 70.72.1 via
> > xe.guc_firmware_path today and report back with a few days of data
> > (this machine currently reproduces 20-35 stalls/day under my normal
> > workload, so the signal should be quick).
> >
> > Thanks,
> > Tales
> >
> > Em ter., 4 de ago. de 2026 às 20:50, Daniele Ceraolo Spurio
> > <daniele.ceraolospurio@intel.com> escreveu:
> >>
> >>
> >> On 8/4/2026 4:00 PM, Tales A. Mendonça wrote:
> >>>> Are we seeing this on i915?
> >>> I have not tested i915 on the affected machines yet - the
> >>> instrumentation that measured the stalls (late-ack logging, kick
> >>> results) is xe-only, so I have no comparable i915 data. I can boot one
> >>> of the ARL machines with i915 for a few days and watch for TLB
> >>> invalidation timeouts there, if that data helps.
> >> That would definitely help, because if the issue does not happen on i915
> >> it likely means that we're missing a WA or something like that in Xe.
> >>
> >>> More generally: both machines here reproduce reliably (~1 stall/hour
> >>> on a desktop workload, much more under memory pressure), so I am happy
> >>> to test anything on them - including any GuC build the firmware team
> >>> would like data on.
> >>>
> >>> On Stuart's masking concern: fully agreed, that is why patch 3 is
> >>> marked RFC. Patches 1-2 are pure diagnostics and stand on their own; I
> >>> am fine holding patch 3 until the firmware side has been looked at.
> >>> The data point it adds is that a doorbell ring unblocks the ack in the
> >>> majority of episodes, while the severe ones ignore 8-9 consecutive
> >>> rings - hopefully that narrows where to look inside the GuC.
> >> Just a bit of a terminology update here, to make sure we're on the same
> >> page: we usually refer to the notification you're sending to the GuC as
> >> an H2G interrupt and not a doorbell. I'm making this clarification
> >> because the GuC supports a separate per-context notification mechanism
> >> that is referred to as doorbell and which we currently do not implement
> >> in neither i915 nor Xe.
> >>
> >> When receiving the H2G interrupt, the only thing that the GuC does is
> >> look into the CTB and process anything in there; however, you've said
> >> that the contents of the H2G CTB are processed immediately, so the
> >> follow up interrupt should result in the GuC just bailing out and doing
> >> nothing because there is no data to process. It feels like when the
> >> issue occurs something is stuck in HW rather than GuC FW and triggering
> >> the interrupt causes the HW to get unstuck.
> >>
> >> I have pushed the latest GuC FW for MTL here in case you want to give it
> >> a go:
> >> https://gitlab.com/dceraolo/drm-firmware/-/blob/f783b931555be057dafc2400b7eb4d445c953fec/i915/mtl_guc_70.72.1.bin
> >> . You can override the GuC firmware used by the driver via the
> >> xe.guc_firmware_path modparam; the path is relative to /lib/firmware/
> >> and the firmware needs to be in initramfs for the driver to find it at
> >> boot. Note that we haven't tested this image on MTL, so it might have
> >> unexpected results.
> >>
> >> Also, would you be able to capture the GuC logs when the issue occurs?
> >> The default guc log size is relatively small, so you'd have to capture
> >> right when the issue happens. However, you can make them bigger by
> >> building the kernel with CONFIG_DRM_XE_DEBUG or by simply modifying the
> >> xe_guc_log.h file to pick the bigger size by default. If you go with the
> >> latter, please also set xe.guc_log_level=3 on the command line (this is
> >> automatically added by the kconfig).
> >>
> >> Thanks,
> >> Daniele
> >>
> >>> I will send a v2 addressing Matt's review comments (the
> >>> __xe_devcoredump unification and the fixes on patch 2).
> >>>
> >>> Thanks,
> >>> Tales
> >>>
> >>> Em ter., 4 de ago. de 2026 às 19:08, Daniele Ceraolo Spurio
> >>> <daniele.ceraolospurio@intel.com> escreveu:
> >>>>
> >>>> On 8/4/2026 2:33 PM, Summers, Stuart wrote:
> >>>>> On Tue, 2026-08-04 at 14:27 -0700, Matthew Brost wrote:
> >>>>>> On Tue, Aug 04, 2026 at 03:02:44PM -0600, Summers, Stuart wrote:
> >>>>>>> On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
> >>>>>>>> Hi,
> >>>>>>>>
> >>>>>>>> This series is a follow-up to the TLB invalidation ack stall I
> >>>>>>>> have
> >>>>>>>> been debugging on ARL, tracked in:
> >>>>>>>>
> >>>>>>>> https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> >>>>>>>>
> >>>>>>>> Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51
> >>>>>>>> and
> >>>>>>>> 7dd1 machines here, plus an independent Arc Pro 130T report on
> >>>>>>>> the
> >>>>>>>> issue above), TLB invalidation acks intermittently stall for
> >>>>>>>> ~2.3s.
> >>>>>>>> The H2G request is consumed from the CTB immediately and the G2H
> >>>>>>>> CTB
> >>>>>>>> is empty the whole time - the firmware simply does not send the
> >>>>>>>> ack
> >>>>>>>> until much later. The fence timeout fires at 2.25s and the ack
> >>>>>>>> lands
> >>>>>>>> tens of ms after it. Userspace blocked on the invalidation
> >>>>>>>> (compositor
> >>>>>>>> buffer unmaps etc.) hitches for the full window.
> >>>>>>> Firstly, thanks for the patch!
> >>>>>>>
> >>>>>>> I haven't looked in to all the details of the sighting you were
> >>>>>>> debugging, but we have had similar issues that were fixed in a
> >>>>>>> later
> >>>>>>> GuC version. I think around 70.60.0? It might be worth trying on
> >>>>>>> something later than that to see if that helps... (+Daniele)
> >>>>>>>
> >>>>>> I think this would require an AR on our end to make a new firmware
> >>>>>> version available.
> >>>>>>
> >>>>>> The upstream repo only has 70.53.0 available for ARL [1] (iirc, ARL
> >>>>>> aliases to MTL for firmware). (+Julia too).
> >>>>>>
> >>>>>> Presumably, the GuC changelogs should indicate whether an issue
> >>>>>> related
> >>>>>> this has been fixed. If so, we need to update all GuC versions across
> >>>>>> both i915 and Xe.
> >>>>> Right... I guess I'd still like to see if we can test this in GuC (or
> >>>>> get confirmation we can't for some reason) before committing something.
> >>>>> My worry is we will prevent bug reports like this by working around it
> >>>>> and miss critical bugs that need to be fixed in the right component.
> >>>>>
> >>>>>> [1]
> >>>>>> https://gitlab.com/kernel-firmware/linux-firmware/-/blob/main/i915/mtl_guc_70.bin?ref_type=heads
> >>>>>>
> >>>>>>>> Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has
> >>>>>>>> no
> >>>>>>> Is there a reason we don't just re-use the main xe_devcoredump()?
> >>>>>>>
> >>>>>> This is my suggestion: the main devcoredump infrastructure is job-
> >>>>>> based,
> >>>>>> so it cannot be used for hangs that are not associated with a job.
> >>>>>>
> >>>>>> In my opinion, this is a gap on our end. Introducing something like
> >>>>>> `xe_devcoredump_gt()`, which can be used for non-job-based hangs
> >>>>>> (e.g.,
> >>>>>> TLB invalidation timeouts like those addressed in this series, or
> >>>>>> more
> >>>>>> generally any GuC protocol hang), makes sense to me.
> >>>>> Ok makes sense. We can do that here. It would be nice to have a more
> >>>>> inclusive implementation that lets us call this from anywhere so we
> >>>>> aren't duplicating things around for different use cases. But not a
> >>>>> blocker here.
> >>>>>
> >>>>>> I haven't looked at the patch yet, but at a high level, adding
> >>>>>> `xe_devcoredump_gt()` seems like a reasonable approach.
> >>>>>>
> >>>>>>>> exec queue or job to blame - leaves a devcoredump with the GuC
> >>>>>>>> log
> >>>>>>>> and
> >>>>>>>> CT state behind (Matt suggested capturing devcoredumps when we
> >>>>>>>> discussed the issue; devcoredumps from both machines are attached
> >>>>>>>> to
> >>>>>>>> the issue above).
> >>>>>>>>
> >>>>>>>> Patch 2 logs when the ack for a timed out invalidation finally
> >>>>>>>> arrives. This is what established that the acks are late rather
> >>>>>>>> than
> >>>>>>>> lost.
> >>>>>>>>
> >>>>>>>> Patch 3 is the RFC part: a delayed work that pokes the GuC
> >>>>>>>> (status
> >>>>>>>> register read, CT flush, doorbell ring) every 250ms while an ack
> >>>>>>>> is
> >>>>>>>> overdue. On my machines this converts the guaranteed 2.3s stall
> >>>>>>>> into
> >>>>>>> I'm a little worried we're just papering over something here that
> >>>>>>> needs
> >>>>>>> to be addressed in GuC, particularly around GT going to sleep or
> >>>>>>> something around the time we're expecting a response, so the pings
> >>>>>>> on
> >>>>>>> registers might be prematurely waking things up which is something
> >>>>>>> we'd
> >>>>>>> want to happen in GuC, not the KMD.
> >>>>>>>
> >>>>>> In general, I agree with this. We should avoid papering over the
> >>>>>> issue
> >>>>>> and instead fix it properly in the GuC. That said, this workaround
> >>>>>> provides a pretty strong data point, since it appears to get the TLB
> >>>>>> invalidation unstuck.
> >>>>> So if we hit this issue I guess we're already going to have some
> >>>>> performance degredation and the workaround makes that better. I need to
> >>>>> look at the implementation, but we could be potentially introducing
> >>>>> performance penalties in other areas doing these pings.
> >>>>>
> >>>>> Again, I'd like to see if we can fix this in the right place before
> >>>>> implementing a workaround for it. Hopefully Daniele or Julia can give
> >>>>> some direction there.
> >>>> Are we seeing this on i915 at all? Given that Xe does not officially
> >>>> support MTL/ARL and is missing several critical WAs for those platforms,
> >>>> the approach so far has been to only update the GuC FW if it is required
> >>>> for i915.
> >>>> Looking at the GuC release notes, there have been a couple of
> >>>> TLB-related fixes after 70.53, but they're both marked as only affecting
> >>>> PVC and Xe2+ platforms, so no fixes seem to be available for ARL (or at
> >>>> least they're not listed in the release notes).
> >>>>
> >>>> Daniele
> >>>>
> >>>>> Thanks,
> >>>>> Stuart
> >>>>>
> >>>>>> Matt
> >>>>>>
> >>>>>>> Thanks,
> >>>>>>> Stuart
> >>>>>>>
> >>>>>>>> a
> >>>>>>>> sub-500ms hiccup for the majority of occurrences; a minority of
> >>>>>>>> severe
> >>>>>>>> episodes ignore 8-9 consecutive doorbells, which points at the
> >>>>>>>> GuC
> >>>>>>>> firmware being internally blocked for the whole window. Full data
> >>>>>>>> on
> >>>>>>>> the issue. I am happy to rework the approach (different delay,
> >>>>>>>> tying it to the G2H handler, dropping the status read, etc.) -
> >>>>>>>> mainly
> >>>>>>>> I would like the firmware side investigated, since no host-side
> >>>>>>>> poke
> >>>>>>>> can fix the severe cases.
> >>>>>>>>
> >>>>>>>> Based on drm-tip. Tested for several days on both ARL machines
> >>>>>>>> under
> >>>>>>>> desktop and VM-heavy workloads.
> >>>>>>>>
> >>>>>>>> Thanks,
> >>>>>>>> Tales
> >>>>>>>>
> >>>>>>>> Tales A. Mendonça (3):
> >>>>>>>> drm/xe: Capture devcoredump on TLB invalidation timeout
> >>>>>>>> drm/xe: Log when a timed out TLB invalidation ack finally
> >>>>>>>> arrives
> >>>>>>>> drm/xe: Kick GuC while TLB invalidation acks are overdue
> >>>>>>>>
> >>>>>>>> drivers/gpu/drm/xe/xe_devcoredump.c | 68 ++++++++++++
> >>>>>>>> drivers/gpu/drm/xe/xe_devcoredump.h | 6 ++
> >>>>>>>> drivers/gpu/drm/xe/xe_tlb_inval.c | 131
> >>>>>>>> +++++++++++++++++++++++-
> >>>>>>>> drivers/gpu/drm/xe/xe_tlb_inval_types.h | 42 ++++++++
> >>>>>>>> 4 files changed, 243 insertions(+), 4 deletions(-)
> >>>>>>>>
> >
>
--
Com os cumprimentos,
Tales A. Mendonça
talesam.org
communitybig.org
^ permalink raw reply [flat|nested] 25+ messages in thread
end of thread, other threads:[~2026-08-08 0:20 UTC | newest]
Thread overview: 25+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-04 2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
2026-08-04 2:14 ` [RFC PATCH 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
2026-08-04 22:05 ` Matthew Brost
2026-08-04 2:14 ` [RFC PATCH 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
2026-08-04 22:17 ` Matthew Brost
2026-08-04 2:14 ` [RFC PATCH 3/3] drm/xe: Kick GuC while TLB invalidation acks are overdue Tales A. Mendonça
2026-08-04 2:15 ` ✗ LGCI.VerificationFailed: failure for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Patchwork
2026-08-04 16:34 ` [RFC PATCH 0/3] " Tales A. Mendonça
2026-08-04 21:29 ` Matthew Brost
2026-08-05 20:24 ` Matthew Brost
2026-08-06 17:37 ` Tales A. Mendonça
2026-08-06 17:56 ` Daniele Ceraolo Spurio
2026-08-04 21:02 ` Summers, Stuart
2026-08-04 21:27 ` Matthew Brost
2026-08-04 21:33 ` Summers, Stuart
2026-08-04 22:08 ` Daniele Ceraolo Spurio
2026-08-04 23:00 ` Tales A. Mendonça
2026-08-04 23:50 ` Daniele Ceraolo Spurio
2026-08-06 17:36 ` Tales A. Mendonça
2026-08-06 21:13 ` Daniele Ceraolo Spurio
2026-08-08 0:20 ` Tales A. Mendonça
2026-08-05 12:32 ` ✗ CI.checkpatch: warning for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2) Patchwork
2026-08-05 12:34 ` ✓ CI.KUnit: success " Patchwork
2026-08-05 13:11 ` ✗ Xe.CI.BAT: failure " Patchwork
2026-08-05 23:39 ` ✗ Xe.CI.FULL: " Patchwork
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox