All of lore.kernel.org
 help / color / mirror / Atom feed
From: Matthew Brost <matthew.brost@intel.com>
To: "Summers, Stuart" <stuart.summers@intel.com>
Cc: "intel-xe@lists.freedesktop.org" <intel-xe@lists.freedesktop.org>,
	"Ceraolo Spurio, Daniele" <daniele.ceraolospurio@intel.com>,
	"talesam@gmail.com" <talesam@gmail.com>,
	"dri-devel@lists.freedesktop.org"
	<dri-devel@lists.freedesktop.org>,
	"Vivi, Rodrigo" <rodrigo.vivi@intel.com>,
	"thomas.hellstrom@linux.intel.com"
	<thomas.hellstrom@linux.intel.com>, <julia.filipchuk@intel.com>
Subject: Re: [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL
Date: Tue, 4 Aug 2026 14:27:36 -0700	[thread overview]
Message-ID: <anJZSDD+lmzbaE7b@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <c0c50dff0f9b64a2db8e1470d75e450ba35cccde.camel@intel.com>

On Tue, Aug 04, 2026 at 03:02:44PM -0600, Summers, Stuart wrote:
> On Mon, 2026-08-03 at 23:14 -0300, Tales A. Mendonça wrote:
> > Hi,
> > 
> > This series is a follow-up to the TLB invalidation ack stall I have
> > been debugging on ARL, tracked in:
> > 
> >   https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> > 
> > Summary of the issue: with GuC 70.53.0 on ARL (reproduced on 7d51 and
> > 7dd1 machines here, plus an independent Arc Pro 130T report on the
> > issue above), TLB invalidation acks intermittently stall for ~2.3s.
> > The H2G request is consumed from the CTB immediately and the G2H CTB
> > is empty the whole time - the firmware simply does not send the ack
> > until much later. The fence timeout fires at 2.25s and the ack lands
> > tens of ms after it. Userspace blocked on the invalidation
> > (compositor
> > buffer unmaps etc.) hitches for the full window.
> 
> Firstly, thanks for the patch!
> 
> I haven't looked in to all the details of the sighting you were
> debugging, but we have had similar issues that were fixed in a later
> GuC version. I think around 70.60.0? It might be worth trying on
> something later than that to see if that helps... (+Daniele)
> 

I think this would require an AR on our end to make a new firmware
version available.

The upstream repo only has 70.53.0 available for ARL [1] (iirc, ARL
aliases to MTL for firmware). (+Julia too).

Presumably, the GuC changelogs should indicate whether an issue related
this has been fixed. If so, we need to update all GuC versions across
both i915 and Xe.

[1] https://gitlab.com/kernel-firmware/linux-firmware/-/blob/main/i915/mtl_guc_70.bin?ref_type=heads

> > 
> > Patch 1 adds xe_devcoredump_gt() so this kind of hang - which has no
> 
> Is there a reason we don't just re-use the main xe_devcoredump()?
>

This is my suggestion: the main devcoredump infrastructure is job-based,
so it cannot be used for hangs that are not associated with a job.

In my opinion, this is a gap on our end. Introducing something like
`xe_devcoredump_gt()`, which can be used for non-job-based hangs (e.g.,
TLB invalidation timeouts like those addressed in this series, or more
generally any GuC protocol hang), makes sense to me.

I haven't looked at the patch yet, but at a high level, adding
`xe_devcoredump_gt()` seems like a reasonable approach.

> > exec queue or job to blame - leaves a devcoredump with the GuC log
> > and
> > CT state behind (Matt suggested capturing devcoredumps when we
> > discussed the issue; devcoredumps from both machines are attached to
> > the issue above).
> > 
> > Patch 2 logs when the ack for a timed out invalidation finally
> > arrives. This is what established that the acks are late rather than
> > lost.
> > 
> > Patch 3 is the RFC part: a delayed work that pokes the GuC (status
> > register read, CT flush, doorbell ring) every 250ms while an ack is
> > overdue. On my machines this converts the guaranteed 2.3s stall into
> 
> I'm a little worried we're just papering over something here that needs
> to be addressed in GuC, particularly around GT going to sleep or
> something around the time we're expecting a response, so the pings on
> registers might be prematurely waking things up which is something we'd
> want to happen in GuC, not the KMD.
> 

In general, I agree with this. We should avoid papering over the issue
and instead fix it properly in the GuC. That said, this workaround
provides a pretty strong data point, since it appears to get the TLB
invalidation unstuck.

Matt

> Thanks,
> Stuart
> 
> > a
> > sub-500ms hiccup for the majority of occurrences; a minority of
> > severe
> > episodes ignore 8-9 consecutive doorbells, which points at the GuC
> > firmware being internally blocked for the whole window. Full data on
> > the issue. I am happy to rework the approach (different delay,
> > tying it to the G2H handler, dropping the status read, etc.) - mainly
> > I would like the firmware side investigated, since no host-side poke
> > can fix the severe cases.
> > 
> > Based on drm-tip. Tested for several days on both ARL machines under
> > desktop and VM-heavy workloads.
> > 
> > Thanks,
> > Tales
> > 
> > Tales A. Mendonça (3):
> >   drm/xe: Capture devcoredump on TLB invalidation timeout
> >   drm/xe: Log when a timed out TLB invalidation ack finally arrives
> >   drm/xe: Kick GuC while TLB invalidation acks are overdue
> > 
> >  drivers/gpu/drm/xe/xe_devcoredump.c     |  68 ++++++++++++
> >  drivers/gpu/drm/xe/xe_devcoredump.h     |   6 ++
> >  drivers/gpu/drm/xe/xe_tlb_inval.c       | 131
> > +++++++++++++++++++++++-
> >  drivers/gpu/drm/xe/xe_tlb_inval_types.h |  42 ++++++++
> >  4 files changed, 243 insertions(+), 4 deletions(-)
> > 
> 

  reply	other threads:[~2026-08-04 21:27 UTC|newest]

Thread overview: 29+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-04  2:14 [RFC PATCH 0/3] drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Tales A. Mendonça
2026-08-04  2:14 ` [RFC PATCH 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
2026-08-04  2:38   ` sashiko-bot
2026-08-04 22:05   ` Matthew Brost
2026-08-04  2:14 ` [RFC PATCH 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
2026-08-04 22:17   ` Matthew Brost
2026-08-04  2:14 ` [RFC PATCH 3/3] drm/xe: Kick GuC while TLB invalidation acks are overdue Tales A. Mendonça
2026-08-04  2:38   ` sashiko-bot
2026-08-04  2:15 ` ✗ LGCI.VerificationFailed: failure for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL Patchwork
2026-08-04 16:34 ` [RFC PATCH 0/3] " Tales A. Mendonça
2026-08-04 21:29   ` Matthew Brost
2026-08-05 20:24     ` Matthew Brost
2026-08-06 17:37       ` Tales A. Mendonça
2026-08-06 17:56         ` Daniele Ceraolo Spurio
2026-08-04 21:02 ` Summers, Stuart
2026-08-04 21:27   ` Matthew Brost [this message]
2026-08-04 21:33     ` Summers, Stuart
2026-08-04 22:08       ` Daniele Ceraolo Spurio
2026-08-04 23:00         ` Tales A. Mendonça
2026-08-04 23:50           ` Daniele Ceraolo Spurio
2026-08-06 17:36             ` Tales A. Mendonça
2026-08-06 21:13               ` Daniele Ceraolo Spurio
2026-08-08  0:20                 ` Tales A. Mendonça
2026-08-08 21:04                 ` Tales A. Mendonça
2026-08-12  0:52                   ` Matthew Brost
2026-08-05 12:32 ` ✗ CI.checkpatch: warning for drm/xe: diagnostics and workaround for GuC TLB invalidation ack stalls on ARL (rev2) Patchwork
2026-08-05 12:34 ` ✓ CI.KUnit: success " Patchwork
2026-08-05 13:11 ` ✗ Xe.CI.BAT: failure " Patchwork
2026-08-05 23:39 ` ✗ Xe.CI.FULL: " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anJZSDD+lmzbaE7b@gsse-cloud1.jf.intel.com \
    --to=matthew.brost@intel.com \
    --cc=daniele.ceraolospurio@intel.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=julia.filipchuk@intel.com \
    --cc=rodrigo.vivi@intel.com \
    --cc=stuart.summers@intel.com \
    --cc=talesam@gmail.com \
    --cc=thomas.hellstrom@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.