Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Matthew Brost <matthew.brost@intel.com>
To: "Summers, Stuart" <stuart.summers@intel.com>
Cc: "intel-xe@lists.freedesktop.org" <intel-xe@lists.freedesktop.org>,
	"Ceraolo Spurio, Daniele" <daniele.ceraolospurio@intel.com>,
	"talesam@gmail.com" <talesam@gmail.com>,
	"dri-devel@lists.freedesktop.org"
	<dri-devel@lists.freedesktop.org>,
	"Vivi, Rodrigo" <rodrigo.vivi@intel.com>,
	"thomas.hellstrom@linux.intel.com"
	<thomas.hellstrom@linux.intel.com>, <julia.filipchuk@intel.com>
Subject: Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
Date: Tue, 4 Aug 2026 14:27:36 -0700	[thread overview]
Message-ID: <anJZSDD+lmzbaE7b@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <c0c50dff0f9b64a2db8e1470d75e450ba35cccde.camel@intel.com>

On Tue, Aug 04, 2026 at 03:02:44PM -0600, Summers, Stuart wrote:
> On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
> > Hi,
> > 
> > This series is a follow-up to the TLB invalidation ack stall I have
> > been debugging on ARL, tracked in:
> > 
> >   https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> > 
> > Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51 and
> > 7dd1 machines here, plus an independent Arc Pro 130T report on the
> > issue above), TLB invalidation acks intermittently stall for ~2.3s.
> > The H2G request is consumed from the CTB immediately and the G2H CTB
> > is empty the whole time - the firmware simply does not send the ack
> > until much later. The fence timeout fires at 2.25s and the ack lands
> > tens of ms after it. Userspace blocked on the invalidation
> > (compositor
> > buffer unmaps etc.) hitches for the full window.
> 
> Firstly, thanks for the patch!
> 
> I haven't looked in to all the details of the sighting you were
> debugging, but we have had similar issues that were fixed in a later
> GuC version. I think around 70.60.0? It might be worth trying on
> something later than that to see if that helps... (+Daniele)
> 

I think this would require an AR on our end to make a new firmware
version available.

The upstream repo only has 70.53.0 available for ARL [1] (iirc, ARL
aliases to MTL for firmware). (+Julia too).

Presumably, the GuC changelogs should indicate whether an issue related
this has been fixed. If so, we need to update all GuC versions across
both i915 and Xe.

[1] https://gitlab.com/kernel-firmware/linux-firmware/-/blob/main/i915/mtl_guc_70.bin?ref_type=heads

> > 
> > Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has no
> 
> Is there a reason we don't just re-use the main xe_devcoredump()?
>

This is my suggestion: the main devcoredump infrastructure is job-based,
so it cannot be used for hangs that are not associated with a job.

In my opinion, this is a gap on our end. Introducing something like
`xe_devcoredump_gt()`, which can be used for non-job-based hangs (e.g.,
TLB invalidation timeouts like those addressed in this series, or more
generally any GuC protocol hang), makes sense to me.

I haven't looked at the patch yet, but at a high level, adding
`xe_devcoredump_gt()` seems like a reasonable approach.

> > exec queue or job to blame - leaves a devcoredump with the GuC log
> > and
> > CT state behind (Matt suggested capturing devcoredumps when we
> > discussed the issue; devcoredumps from both machines are attached to
> > the issue above).
> > 
> > Patch 2 logs when the ack for a timed out invalidation finally
> > arrives. This is what established that the acks are late rather than
> > lost.
> > 
> > Patch 3 is the RFC part: a delayed work that pokes the GuC (status
> > register read, CT flush, doorbell ring) every 250ms while an ack is
> > overdue. On my machines this converts the guaranteed 2.3s stall into
> 
> I'm a little worried we're just papering over something here that needs
> to be addressed in GuC, particularly around GT going to sleep or
> something around the time we're expecting a response, so the pings on
> registers might be prematurely waking things up which is something we'd
> want to happen in GuC, not the KMD.
> 

In general, I agree with this. We should avoid papering over the issue
and instead fix it properly in the GuC. That said, this workaround
provides a pretty strong data point, since it appears to get the TLB
invalidation unstuck.

Matt

> Thanks,
> Stuart
> 
> > a
> > sub-500ms hiccup for the majority of occurrences; a minority of
> > severe
> > episodes ignore 8-9 consecutive doorbells, which points at the GuC
> > firmware being internally blocked for the whole window. Full data on
> > the issue. I am happy to rework the approach (different delay,
> > tying it to the G2H handler, dropping the status read, etc.) - mainly
> > I would like the firmware side investigated, since no host-side poke
> > can fix the severe cases.
> > 
> > Based on drm-tip. Tested for several days on both ARL machines under
> > desktop and VM-heavy workloads.
> > 
> > Thanks,
> > Tales
> > 
> > Tales A. Mendonça (3):
> >   drm/xe: Capture devcoredump on TLB invalidation timeout
> >   drm/xe: Log when a timed out TLB invalidation ack finally arrives
> >   drm/xe: Kick GuC while TLB invalidation acks are overdue
> > 
> >  drivers/gpu/drm/xe/xe_devcoredump.c     |  68 ++++++++++++
> >  drivers/gpu/drm/xe/xe_devcoredump.h     |   6 ++
> >  drivers/gpu/drm/xe/xe_tlb_inval.c       | 131
> > +++++++++++++++++++++++-
> >  drivers/gpu/drm/xe/xe_tlb_inval_types.h |  42 ++++++++
> >  4 files changed, 243 insertions(+), 4 deletions(-)
> > 
> 

  reply	other threads:[~2026-08-04 21:27 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-04  2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
2026-08-04  2:14 ` [RFC PATCH 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
2026-08-04 22:05   ` Matthew Brost
2026-08-04  2:14 ` [RFC PATCH 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
2026-08-04 22:17   ` Matthew Brost
2026-08-04  2:14 ` [RFC PATCH 3/3] drm/xe: Kick GuC while TLB invalidation acks are overdue Tales A. Mendonça
2026-08-04  2:15 ` ✗ LGCI.VerificationFailed: failure for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Patchwork
2026-08-04 16:34 ` [RFC PATCH 0/3] " Tales A. Mendonça
2026-08-04 21:29   ` Matthew Brost
2026-08-05 20:24     ` Matthew Brost
2026-08-06 17:37       ` Tales A. Mendonça
2026-08-06 17:56         ` Daniele Ceraolo Spurio
2026-08-04 21:02 ` Summers, Stuart
2026-08-04 21:27   ` Matthew Brost [this message]
2026-08-04 21:33     ` Summers, Stuart
2026-08-04 22:08       ` Daniele Ceraolo Spurio
2026-08-04 23:00         ` Tales A. Mendonça
2026-08-04 23:50           ` Daniele Ceraolo Spurio
2026-08-06 17:36             ` Tales A. Mendonça
2026-08-06 21:13               ` Daniele Ceraolo Spurio
2026-08-08  0:20                 ` Tales A. Mendonça
2026-08-05 12:32 ` ✗ CI.checkpatch: warning for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2) Patchwork
2026-08-05 12:34 ` ✓ CI.KUnit: success " Patchwork
2026-08-05 13:11 ` ✗ Xe.CI.BAT: failure " Patchwork
2026-08-05 23:39 ` ✗ Xe.CI.FULL: " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anJZSDD+lmzbaE7b@gsse-cloud1.jf.intel.com \
    --to=matthew.brost@intel.com \
    --cc=daniele.ceraolospurio@intel.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=julia.filipchuk@intel.com \
    --cc=rodrigo.vivi@intel.com \
    --cc=stuart.summers@intel.com \
    --cc=talesam@gmail.com \
    --cc=thomas.hellstrom@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox