Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Matthew Brost <matthew.brost@intel.com>
To: "Tales A. Mendonça" <talesam@gmail.com>
Cc: <intel-xe@lists.freedesktop.org>,
	<daniele.ceraolospurio@intel.com>, <stuart.summers@intel.com>,
	<julia.filipchuk@intel.com>, <thomas.hellstrom@linux.intel.com>,
	<rodrigo.vivi@intel.com>, <jani.nikula@intel.com>,
	<navonjohnlukose@gmail.com>, <dri-devel@lists.freedesktop.org>
Subject: Re: [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933)
Date: Thu, 24 Sep 2026 15:05:18 -0700	[thread overview]
Message-ID: <arWenmbBkCZ4m7I8@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <20260922144634.55130-1-talesam@gmail.com>

On Tue, Sep 22, 2026 at 11:46:31AM -0300, Tales A. Mendonça wrote:

Thanks for the patches, I've merge to drm-xe-next.

Matt

> Hi,
> 
> v6 of the TLB invalidation ack stall fix for ARL. Tracked in:
> 
>   https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> 
> The whole series now carries Matthew Brost's Reviewed-by - thank you for
> working through it, including cross-checking patch 3 against the i915
> implementation.
> 
> Recap: on the standalone media GT of MTL/ARL the CPU reads stale cache
> lines for data the GuC has already written. The visible symptom is TLB
> invalidation acks appearing to stall for a near-constant ~2.3s. i915
> works around this as Wa_22016122933; xe never inherited it. Patch 3
> implements it, applying XE_BO_FLAG_NEEDS_UC to the GuC-shared
> allocations (CTBs, log, ADS, SLPC, engine activity) on the standalone
> media GT, scoped by a new OOB rule (22016122933 MEDIA_VERSION(1300)).
> 
> The one code change since v5 is in patch 1, from a second issue Sashiko
> raised: the capture could run without a runtime PM reference. Every
> pending invalidation fence holds one, taken in
> xe_tlb_inval_fence_init(), and xe_tlb_inval_fence_signal() drops it via
> xe_tlb_inval_fence_fini(). The timeout loop signals the expired fences
> and only then calls xe_devcoredump_gt(), whose forcewake acquisition has
> always relied on the caller holding a PM reference. If those were the
> last references the device could begin autosuspending before the
> snapshot touched the hardware. v6 takes a reference while the pending
> fences still guarantee the device is awake, and releases it after the
> capture. Matt confirmed the analysis and the fix on the list.
> 
> I am carrying Matt's tag on patch 1 across that change since he reviewed
> the fix itself, but flagging it here so it is not silently inherited.
> 
> Two open points from earlier revisions, both now settled:
> 
>  - LRC coverage: i915 also marks the LRC UC on non-dGPU
>    (__lrc_alloc_state()). That does not match the erratum's direction -
>    LRC writes come from hardware context save and the GuC reads it -
>    and Matt's guidance was to leave i915 alone and treat it as out of
>    scope for xe.
> 
>  - SIGID: the TLB logging in patch 2 will be converted once a TLB
>    component exists in DEFINE_XE_LOG_COMPONENTS(); a colleague of
>    Matt's volunteered to do that as a follow-up on top of this series.
> 
> Two notes carried over, still open to either answer:
> 
>  1. CPU mapping: keeping XE_BO_FLAG_NEEDS_UC (uncached on both sides).
>     It is the tested configuration and no throughput difference against
>     the CPU-WC variant was measurable. Matching i915's exact CPU-WC +
>     GGTT-UC combination needs either a new BO flag or decoupling the
>     GGTT cache-mode selection from XE_BO_FLAG_NEEDS_UC; happy to add
>     that plumbing if parity is preferred.
> 
>  2. Fixes:/Cc: stable are left out, since MTL/ARL is require_force_probe
>     in xe. Also happy to add them.
> 
> Validation of patch 3 is six weeks on two ARL machines (7d51 and 7dd1),
> across kernels 7.1.6, 7.1.8 and 7.2, with over 10M TLB invalidations
> processed and zero ack stalls. Before the fix both machines reproduced
> 20-60 stalls/day, every day, on two GuC firmware versions. The 7dd1
> machine, which could not survive a day of media workloads on xe without
> a platform freeze, has been running xe full time since 11 August with
> zero incidents.
> 
> checkpatch is clean, except for one --strict CHECK about macro argument
> reuse in the xe_devcoredump() wrapper in patch 1, which is intentional:
> the macro only exists to forward (_q)->gt alongside _q.
> 
> v5 -> v6:
> - Rebased on today's drm-tip; builds clean, no conflicts.
> - Patch 1: hold a runtime PM reference across the devcoredump capture
>   (second issue reported by Sashiko, confirmed by Matt).
> - Patches 1-3: collected Reviewed-by from Matthew Brost.
> - Patches 2-3: otherwise unchanged.
> 
> Thanks,
> Tales
> 
> Tales A. Mendonça (3):
>   drm/xe: Capture devcoredump on TLB invalidation timeout
>   drm/xe: Log when a timed out TLB invalidation ack finally arrives
>   drm/xe: Implement Wa_22016122933
> 
>  drivers/gpu/drm/xe/xe_devcoredump.c         | 46 ++++++++++--------
>  drivers/gpu/drm/xe/xe_devcoredump.h         | 15 ++++--
>  drivers/gpu/drm/xe/xe_guc.c                 | 16 ++++++
>  drivers/gpu/drm/xe/xe_guc.h                 |  2 +
>  drivers/gpu/drm/xe/xe_guc_ads.c             |  3 +-
>  drivers/gpu/drm/xe/xe_guc_ct.c              |  6 ++-
>  drivers/gpu/drm/xe/xe_guc_engine_activity.c |  6 ++-
>  drivers/gpu/drm/xe/xe_guc_log.c             |  7 ++-
>  drivers/gpu/drm/xe/xe_guc_pc.c              |  3 +-
>  drivers/gpu/drm/xe/xe_tlb_inval.c           | 54 +++++++++++++++++++++
>  drivers/gpu/drm/xe/xe_tlb_inval_types.h     | 17 +++++++
>  drivers/gpu/drm/xe/xe_wa_oob.rules          |  1 +
>  12 files changed, 143 insertions(+), 33 deletions(-)
> 
> -- 
> 2.55.0
> 

  parent reply	other threads:[~2026-09-24 22:05 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 14:46 [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 3/3] drm/xe: Implement Wa_22016122933 Tales A. Mendonça
2026-09-22 15:14 ` ✗ CI.checkpatch: warning for drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) (rev5) Patchwork
2026-09-22 15:16 ` ✓ CI.KUnit: success " Patchwork
2026-09-22 17:01 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-23  2:55 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-09-24 22:05 ` Matthew Brost [this message]
2026-09-25  1:06   ` [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Tales A. Mendonça

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arWenmbBkCZ4m7I8@gsse-cloud1.jf.intel.com \
    --to=matthew.brost@intel.com \
    --cc=daniele.ceraolospurio@intel.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=jani.nikula@intel.com \
    --cc=julia.filipchuk@intel.com \
    --cc=navonjohnlukose@gmail.com \
    --cc=rodrigo.vivi@intel.com \
    --cc=stuart.summers@intel.com \
    --cc=talesam@gmail.com \
    --cc=thomas.hellstrom@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox