Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Tales A. Mendonça" <talesam@gmail.com>
To: intel-xe@lists.freedesktop.org
Cc: matthew.brost@intel.com, daniele.ceraolospurio@intel.com,
	stuart.summers@intel.com, julia.filipchuk@intel.com,
	thomas.hellstrom@linux.intel.com, rodrigo.vivi@intel.com,
	jani.nikula@intel.com, navonjohnlukose@gmail.com,
	dri-devel@lists.freedesktop.org,
	"Tales A. Mendonça" <talesam@gmail.com>
Subject: [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933)
Date: Tue, 22 Sep 2026 11:46:31 -0300	[thread overview]
Message-ID: <20260922144634.55130-1-talesam@gmail.com> (raw)

Hi,

v6 of the TLB invalidation ack stall fix for ARL. Tracked in:

  https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678

The whole series now carries Matthew Brost's Reviewed-by - thank you for
working through it, including cross-checking patch 3 against the i915
implementation.

Recap: on the standalone media GT of MTL/ARL the CPU reads stale cache
lines for data the GuC has already written. The visible symptom is TLB
invalidation acks appearing to stall for a near-constant ~2.3s. i915
works around this as Wa_22016122933; xe never inherited it. Patch 3
implements it, applying XE_BO_FLAG_NEEDS_UC to the GuC-shared
allocations (CTBs, log, ADS, SLPC, engine activity) on the standalone
media GT, scoped by a new OOB rule (22016122933 MEDIA_VERSION(1300)).

The one code change since v5 is in patch 1, from a second issue Sashiko
raised: the capture could run without a runtime PM reference. Every
pending invalidation fence holds one, taken in
xe_tlb_inval_fence_init(), and xe_tlb_inval_fence_signal() drops it via
xe_tlb_inval_fence_fini(). The timeout loop signals the expired fences
and only then calls xe_devcoredump_gt(), whose forcewake acquisition has
always relied on the caller holding a PM reference. If those were the
last references the device could begin autosuspending before the
snapshot touched the hardware. v6 takes a reference while the pending
fences still guarantee the device is awake, and releases it after the
capture. Matt confirmed the analysis and the fix on the list.

I am carrying Matt's tag on patch 1 across that change since he reviewed
the fix itself, but flagging it here so it is not silently inherited.

Two open points from earlier revisions, both now settled:

 - LRC coverage: i915 also marks the LRC UC on non-dGPU
   (__lrc_alloc_state()). That does not match the erratum's direction -
   LRC writes come from hardware context save and the GuC reads it -
   and Matt's guidance was to leave i915 alone and treat it as out of
   scope for xe.

 - SIGID: the TLB logging in patch 2 will be converted once a TLB
   component exists in DEFINE_XE_LOG_COMPONENTS(); a colleague of
   Matt's volunteered to do that as a follow-up on top of this series.

Two notes carried over, still open to either answer:

 1. CPU mapping: keeping XE_BO_FLAG_NEEDS_UC (uncached on both sides).
    It is the tested configuration and no throughput difference against
    the CPU-WC variant was measurable. Matching i915's exact CPU-WC +
    GGTT-UC combination needs either a new BO flag or decoupling the
    GGTT cache-mode selection from XE_BO_FLAG_NEEDS_UC; happy to add
    that plumbing if parity is preferred.

 2. Fixes:/Cc: stable are left out, since MTL/ARL is require_force_probe
    in xe. Also happy to add them.

Validation of patch 3 is six weeks on two ARL machines (7d51 and 7dd1),
across kernels 7.1.6, 7.1.8 and 7.2, with over 10M TLB invalidations
processed and zero ack stalls. Before the fix both machines reproduced
20-60 stalls/day, every day, on two GuC firmware versions. The 7dd1
machine, which could not survive a day of media workloads on xe without
a platform freeze, has been running xe full time since 11 August with
zero incidents.

checkpatch is clean, except for one --strict CHECK about macro argument
reuse in the xe_devcoredump() wrapper in patch 1, which is intentional:
the macro only exists to forward (_q)->gt alongside _q.

v5 -> v6:
- Rebased on today's drm-tip; builds clean, no conflicts.
- Patch 1: hold a runtime PM reference across the devcoredump capture
  (second issue reported by Sashiko, confirmed by Matt).
- Patches 1-3: collected Reviewed-by from Matthew Brost.
- Patches 2-3: otherwise unchanged.

Thanks,
Tales

Tales A. Mendonça (3):
  drm/xe: Capture devcoredump on TLB invalidation timeout
  drm/xe: Log when a timed out TLB invalidation ack finally arrives
  drm/xe: Implement Wa_22016122933

 drivers/gpu/drm/xe/xe_devcoredump.c         | 46 ++++++++++--------
 drivers/gpu/drm/xe/xe_devcoredump.h         | 15 ++++--
 drivers/gpu/drm/xe/xe_guc.c                 | 16 ++++++
 drivers/gpu/drm/xe/xe_guc.h                 |  2 +
 drivers/gpu/drm/xe/xe_guc_ads.c             |  3 +-
 drivers/gpu/drm/xe/xe_guc_ct.c              |  6 ++-
 drivers/gpu/drm/xe/xe_guc_engine_activity.c |  6 ++-
 drivers/gpu/drm/xe/xe_guc_log.c             |  7 ++-
 drivers/gpu/drm/xe/xe_guc_pc.c              |  3 +-
 drivers/gpu/drm/xe/xe_tlb_inval.c           | 54 +++++++++++++++++++++
 drivers/gpu/drm/xe/xe_tlb_inval_types.h     | 17 +++++++
 drivers/gpu/drm/xe/xe_wa_oob.rules          |  1 +
 12 files changed, 143 insertions(+), 33 deletions(-)

-- 
2.55.0


             reply	other threads:[~2026-09-22 14:46 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 14:46 Tales A. Mendonça [this message]
2026-09-22 14:46 ` [PATCH v6 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
2026-09-22 14:46 ` [PATCH v6 3/3] drm/xe: Implement Wa_22016122933 Tales A. Mendonça
2026-09-22 15:14 ` ✗ CI.checkpatch: warning for drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) (rev5) Patchwork
2026-09-22 15:16 ` ✓ CI.KUnit: success " Patchwork
2026-09-22 17:01 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-23  2:55 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-09-24 22:05 ` [PATCH v6 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Matthew Brost
2026-09-25  1:06   ` Tales A. Mendonça

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260922144634.55130-1-talesam@gmail.com \
    --to=talesam@gmail.com \
    --cc=daniele.ceraolospurio@intel.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=jani.nikula@intel.com \
    --cc=julia.filipchuk@intel.com \
    --cc=matthew.brost@intel.com \
    --cc=navonjohnlukose@gmail.com \
    --cc=rodrigo.vivi@intel.com \
    --cc=stuart.summers@intel.com \
    --cc=thomas.hellstrom@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox