All of lore.kernel.org
 help / color / mirror / Atom feed
From: Matthew Brost <matthew.brost@intel.com>
To: "Tales A. Mendonça" <talesam@gmail.com>
Cc: <intel-xe@lists.freedesktop.org>,
	<dri-devel@lists.freedesktop.org>,
	<thomas.hellstrom@linux.intel.com>, <rodrigo.vivi@intel.com>
Subject: Re: [PATCH v1 3/4] drm/xe/guc/ct: Queue G2H worker before flushing it in timeout paths
Date: Thu, 23 Jul 2026 16:19:11 -0700	[thread overview]
Message-ID: <amKhbyYjQqcG1Asc@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <CAHBRX4GTHiv+3YYiNoDFy0L8KAyb7ruqDZR1rV9MNGwHBmLhnA@mail.gmail.com>

On Thu, Jul 23, 2026 at 07:31:10PM -0300, Tales A. Mendonça wrote:
> > What about the H2G CT?
> 
> Hi Matt,
> 
> Thanks for the series and for the H2G question — I extended my
> instrumentation to sample the H2G CTB descriptor as well, and I now
> have results from two ARL machines running your series 170939.
> 
> Setup:
> - Machine A: ARL-H (PCI 7d51, Core Ultra 5 225H)
> - Machine B: ARL-P (PCI 7dd1, Core Ultra 7 255H)
> - Both: GuC firmware 70.53.0, kernel 7.1.3 with your two patches
> backported, plus debug-only instrumentation that samples both CTB
> descriptors and the GT C-state at TDR time and logs late acks. My
> earlier G2H-flush patch (3/4 of my series) is NOT applied, so your CPU
> flush WA is the only recovery path in place, with the stock
> invalidation timeout.
> 
> Results, 7 timeout events so far (6 on machine A, 1 on machine B),
> identical signature on every single one:
> 
>   TLB invalidation timeout gtidle: seqno=89335, gt_c_state request=C0
>     timeout=C6, c6_residency request-to-timeout=2278ms, waited=2279ms
>   TLB invalidation timeout g2h: seqno=89335, ctb pending=0 dw,
>     head=23932, tail=23932, outstanding=1
>   TLB invalidation timeout h2g: seqno=89335, ctb pending=0 dw,
>     head=792, tail=792
>   TLB invalidation late ack: seqno=89335 recv=89335,
>     request-to-ack=2325ms, timeout-to-ack=45ms
> 
> To answer your question directly: the H2G CTB is fully drained at TDR
> time — head == tail as read from the shared descriptor, i.e. the GuC's
> own view of its consumption progress. So the invalidation request is
> not sitting unconsumed in the H2G ring. Combined with the empty G2H
> ring (with one outstanding G2H credit) this means the GuC consumed the
> request but did not send the ack until 2-45ms after the TDR fired,
> roughly 2.3s after the request was posted.
> 
> Also worth noting: your "CPU flush WA resolved %u pending TLB inval
> fence(s)" warning has not fired once across many hours on either
> machine, while 7 timeouts went through the recovery path. So at least
> on ARL the LNL flush WA does not appear to be the mechanism — there is
> no G2H sitting unprocessed on the host side; the ack genuinely arrives
> late from the firmware.
> 
> Everything points at a GuC 70.53.0 stall: request consumed, ack

Yes, this looks like something is going on in the GuC.

> delayed by seconds, firmware "wakes up" right around the TDR. If it
> would help I can file a gitlab issue with the full logs from both
> machines, and I am happy to keep collecting events (the two machines
> produce a steady trickle).

I'd file a gitlab.

> 
> I would also like to give xe_devcoredump_gt() a try as a separate
> series, as you suggested, if no one else is already on it.
> 

If you could wire up `xe_devcoredump_gt` and attach the devcoredump so
we can inspect the GuC log, it would probably be worth us taking a quick
look.
 
We do support ARL on i915, and the firmware is the same. In general, the
GuC firmware operates similarly across most platforms, so it's quite
possible that whatever is happening here could also be an issue on other
platforms.

Matt

> Thanks,
> Tales

  reply	other threads:[~2026-07-23 23:19 UTC|newest]

Thread overview: 22+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-22  0:46 [PATCH v1 0/4] drm/xe: MCR semaphore and TLB invalidation timeout fixes for ARL Tales A. Mendonça
2026-07-22  0:46 ` [PATCH v1 1/4] drm/xe/mcr: Keep GT forcewake during MCR steering Tales A. Mendonça
2026-07-22 14:02   ` sashiko-bot
2026-07-22 17:45   ` Tales A. Mendonça
2026-07-22 18:10   ` Matt Roper
2026-07-22 18:39     ` Tales A. Mendonça
2026-07-22  0:46 ` [PATCH v1 2/4] drm/xe/mcr: Sanitize steering semaphore on GT resume Tales A. Mendonça
2026-07-22 13:57   ` sashiko-bot
2026-07-22 17:47   ` Tales A. Mendonça
2026-10-03 19:45   ` Navon John Lukose
2026-07-22  0:46 ` [PATCH v1 3/4] drm/xe/guc/ct: Queue G2H worker before flushing it in timeout paths Tales A. Mendonça
2026-07-22 14:07   ` sashiko-bot
2026-07-22 17:48   ` Tales A. Mendonça
2026-07-22 21:10   ` Matthew Brost
2026-07-22 22:03     ` Matthew Brost
2026-07-22 23:28       ` Tales A. Mendonça
2026-07-23  0:26         ` Matthew Brost
2026-07-23  2:08           ` Tales A. Mendonça
2026-07-23 22:31           ` Tales A. Mendonça
2026-07-23 23:19             ` Matthew Brost [this message]
2026-07-22  0:46 ` [PATCH v1 4/4] drm/xe: Raise hw_tlb_timeout to cover observed GuC ack latency Tales A. Mendonça
2026-07-22 17:27 ` ✗ LGCI.VerificationFailed: failure for drm/xe: MCR semaphore and TLB invalidation timeout fixes for ARL Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=amKhbyYjQqcG1Asc@gsse-cloud1.jf.intel.com \
    --to=matthew.brost@intel.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=rodrigo.vivi@intel.com \
    --cc=talesam@gmail.com \
    --cc=thomas.hellstrom@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.