From: Matthew Brost <matthew.brost@intel.com>
To: "Tales A. Mendonça" <talesam@gmail.com>
Cc: <intel-xe@lists.freedesktop.org>,
<daniele.ceraolospurio@intel.com>, <stuart.summers@intel.com>,
<julia.filipchuk@intel.com>, <thomas.hellstrom@linux.intel.com>,
<rodrigo.vivi@intel.com>, <jani.nikula@intel.com>,
<navonjohnlukose@gmail.com>, <dri-devel@lists.freedesktop.org>
Subject: Re: [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933)
Date: Thu, 24 Sep 2026 15:05:18 -0700 [thread overview]
Message-ID: <arWenmbBkCZ4m7I8@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <20260922144634.55130-1-talesam@gmail.com>
On Tue, Sep 22, 2026 at 11:46:31AM -0300, Tales A. Mendonça wrote:
Thanks for the patches, I've merge to drm-xe-next.
Matt
> Hi,
>
> v6 of the TLB invalidation ack stall fix for ARL. Tracked in:
>
> https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
>
> The whole series now carries Matthew Brost's Reviewed-by - thank you for
> working through it, including cross-checking patch 3 against the i915
> implementation.
>
> Recap: on the standalone media GT of MTL/ARL the CPU reads stale cache
> lines for data the GuC has already written. The visible symptom is TLB
> invalidation acks appearing to stall for a near-constant ~2.3s. i915
> works around this as Wa_22016122933; xe never inherited it. Patch 3
> implements it, applying XE_BO_FLAG_NEEDS_UC to the GuC-shared
> allocations (CTBs, log, ADS, SLPC, engine activity) on the standalone
> media GT, scoped by a new OOB rule (22016122933 MEDIA_VERSION(1300)).
>
> The one code change since v5 is in patch 1, from a second issue Sashiko
> raised: the capture could run without a runtime PM reference. Every
> pending invalidation fence holds one, taken in
> xe_tlb_inval_fence_init(), and xe_tlb_inval_fence_signal() drops it via
> xe_tlb_inval_fence_fini(). The timeout loop signals the expired fences
> and only then calls xe_devcoredump_gt(), whose forcewake acquisition has
> always relied on the caller holding a PM reference. If those were the
> last references the device could begin autosuspending before the
> snapshot touched the hardware. v6 takes a reference while the pending
> fences still guarantee the device is awake, and releases it after the
> capture. Matt confirmed the analysis and the fix on the list.
>
> I am carrying Matt's tag on patch 1 across that change since he reviewed
> the fix itself, but flagging it here so it is not silently inherited.
>
> Two open points from earlier revisions, both now settled:
>
> - LRC coverage: i915 also marks the LRC UC on non-dGPU
> (__lrc_alloc_state()). That does not match the erratum's direction -
> LRC writes come from hardware context save and the GuC reads it -
> and Matt's guidance was to leave i915 alone and treat it as out of
> scope for xe.
>
> - SIGID: the TLB logging in patch 2 will be converted once a TLB
> component exists in DEFINE_XE_LOG_COMPONENTS(); a colleague of
> Matt's volunteered to do that as a follow-up on top of this series.
>
> Two notes carried over, still open to either answer:
>
> 1. CPU mapping: keeping XE_BO_FLAG_NEEDS_UC (uncached on both sides).
> It is the tested configuration and no throughput difference against
> the CPU-WC variant was measurable. Matching i915's exact CPU-WC +
> GGTT-UC combination needs either a new BO flag or decoupling the
> GGTT cache-mode selection from XE_BO_FLAG_NEEDS_UC; happy to add
> that plumbing if parity is preferred.
>
> 2. Fixes:/Cc: stable are left out, since MTL/ARL is require_force_probe
> in xe. Also happy to add them.
>
> Validation of patch 3 is six weeks on two ARL machines (7d51 and 7dd1),
> across kernels 7.1.6, 7.1.8 and 7.2, with over 10M TLB invalidations
> processed and zero ack stalls. Before the fix both machines reproduced
> 20-60 stalls/day, every day, on two GuC firmware versions. The 7dd1
> machine, which could not survive a day of media workloads on xe without
> a platform freeze, has been running xe full time since 11 August with
> zero incidents.
>
> checkpatch is clean, except for one --strict CHECK about macro argument
> reuse in the xe_devcoredump() wrapper in patch 1, which is intentional:
> the macro only exists to forward (_q)->gt alongside _q.
>
> v5 -> v6:
> - Rebased on today's drm-tip; builds clean, no conflicts.
> - Patch 1: hold a runtime PM reference across the devcoredump capture
> (second issue reported by Sashiko, confirmed by Matt).
> - Patches 1-3: collected Reviewed-by from Matthew Brost.
> - Patches 2-3: otherwise unchanged.
>
> Thanks,
> Tales
>
> Tales A. Mendonça (3):
> drm/xe: Capture devcoredump on TLB invalidation timeout
> drm/xe: Log when a timed out TLB invalidation ack finally arrives
> drm/xe: Implement Wa_22016122933
>
> drivers/gpu/drm/xe/xe_devcoredump.c | 46 ++++++++++--------
> drivers/gpu/drm/xe/xe_devcoredump.h | 15 ++++--
> drivers/gpu/drm/xe/xe_guc.c | 16 ++++++
> drivers/gpu/drm/xe/xe_guc.h | 2 +
> drivers/gpu/drm/xe/xe_guc_ads.c | 3 +-
> drivers/gpu/drm/xe/xe_guc_ct.c | 6 ++-
> drivers/gpu/drm/xe/xe_guc_engine_activity.c | 6 ++-
> drivers/gpu/drm/xe/xe_guc_log.c | 7 ++-
> drivers/gpu/drm/xe/xe_guc_pc.c | 3 +-
> drivers/gpu/drm/xe/xe_tlb_inval.c | 54 +++++++++++++++++++++
> drivers/gpu/drm/xe/xe_tlb_inval_types.h | 17 +++++++
> drivers/gpu/drm/xe/xe_wa_oob.rules | 1 +
> 12 files changed, 143 insertions(+), 33 deletions(-)
>
> --
> 2.55.0
>
next prev parent reply other threads:[~2026-09-24 22:05 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-22 14:46 [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 3/3] drm/xe: Implement Wa_22016122933 Tales A. Mendonça
2026-09-22 15:14 ` ✗ CI.checkpatch: warning for drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) (rev5) Patchwork
2026-09-22 15:16 ` ✓ CI.KUnit: success " Patchwork
2026-09-22 17:01 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-23 2:55 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-09-24 22:05 ` Matthew Brost [this message]
2026-09-25 1:06 ` [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Tales A. Mendonça
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arWenmbBkCZ4m7I8@gsse-cloud1.jf.intel.com \
--to=matthew.brost@intel.com \
--cc=daniele.ceraolospurio@intel.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=intel-xe@lists.freedesktop.org \
--cc=jani.nikula@intel.com \
--cc=julia.filipchuk@intel.com \
--cc=navonjohnlukose@gmail.com \
--cc=rodrigo.vivi@intel.com \
--cc=stuart.summers@intel.com \
--cc=talesam@gmail.com \
--cc=thomas.hellstrom@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox