From: "Tales A. Mendonça" <talesam@gmail.com>
To: intel-xe@lists.freedesktop.org
Cc: dri-devel@lists.freedesktop.org, matthew.brost@intel.com,
thomas.hellstrom@linux.intel.com, rodrigo.vivi@intel.com,
"Tales A. Mendonça" <talesam@gmail.com>
Subject: [PATCH v1 0/4] drm/xe: MCR semaphore and TLB invalidation timeout fixes for ARL
Date: Tue, 21 Jul 2026 21:46:50 -0300 [thread overview]
Message-ID: <20260722004654.744249-1-talesam@gmail.com> (raw)
This series addresses two issues observed on Arrow Lake laptops (two
ASUS Vivobook S 14 S3407CA units, Core Ultra 5 225H, iGPU 7d51/ARL-H
and 7dd1/ARL-P, GuC firmware 70.53.0), reproduced and bisected over
several weeks of daily use with instrumented kernels.
1) MCR steering semaphore held across resume (patches 1-2)
On s2idle resume the first MCR access finds STEER_SEMAPHORE still held
by firmware. The current code waits only 10us while spinning inside
the mcr_lock spinlock and then fires a WARN with a full backtrace:
drm_WARN_ON_ONCE(ret == -110)
WARNING: drivers/gpu/drm/xe/xe_gt_mcr.c:697 mcr_lock
Patch 1 ports the i915 MCR locking model: wait for the hardware
semaphore (up to 100ms, sleeping) before taking the software spinlock,
hold GT forcewake over the lock/steer/unlock cycle (Wa_22018931422),
and demote the WARN to a rate-limited GT error. Patch 2 ports
intel_gt_mcr_lock_sanitize() from i915, releasing an orphaned
semaphore at GT resume before the first MCR access.
Testing: before the fix the semaphore WARN fired on nearly every
s2idle resume. With patches 1-2, zero occurrences across dozens of
s2idle cycles (including several of 10+ hours) on both machines.
2) TLB invalidation fence timeouts / GuC ack latency (patches 3-4)
Under CPU load, TLB invalidation fences intermittently time out:
TLB invalidation fence timeout, seqno=N recv=N-1
Instrumented kernels (logging the GT C-state at request and timeout
time, the G2H CTB occupancy at timeout time, and the ack arrival time)
show, across 26+ events on both machines:
- the ack is not sitting unprocessed in the G2H CTB at timeout time
(CTB head == tail, zero pending dwords), and
- the GT C-state does not matter (events observed both with the GT
parked in C6 the whole wait and, with forcewake artificially held,
in C0 the whole wait), and
- the ack arrives 5-475ms (median ~22ms) after the timeout handler
runs, i.e. right after the handler's forced G2H drain and register
traffic.
This looks like a GuC-side stall posting the ack (a separate report
with the full data is being prepared), but the host side can still be
made more robust: patch 3 fixes the timeout-recovery paths, which
flush a G2H worker that was possibly never queued (flush_work() on an
idle work item is a no-op, so a lost/coalesced GUC2HOST interrupt
turns the flush into a no-op too). Queueing the worker before flushing
guarantees the CTB is actually drained, and with it the fence-timeout
message becomes a reliable indicator of a genuine GuC-side stall.
Patch 4 (RFC) raises the driver-internal TLB invalidation deadline
from HZ/4 to HZ/2. All observed request-to-ack latencies cluster just
above the current 250ms deadline (2.5s only because the fence TDR
rounds up); HZ/2 keeps the WARN meaningful while not firing on
latencies the hardware demonstrably recovers from. This one is a
band-aid for the firmware behaviour above, hence RFC — happy to drop
it if the GuC stall gets fixed at the source.
Series based on drm-tip, compile-tested there; runtime testing was
done on 7.1.3 plus backports of these patches (kernel package
"linux71-comm" of the BigCommunity distribution, in daily use on both
machines).
Tales A. Mendonça (4):
drm/xe/mcr: Keep GT forcewake during MCR steering
drm/xe/mcr: Sanitize steering semaphore on GT resume
drm/xe/guc/ct: Queue G2H worker before flushing it in timeout paths
drm/xe: Raise hw_tlb_timeout to cover observed GuC ack latency
drivers/gpu/drm/xe/xe_gt.c | 7 +++
drivers/gpu/drm/xe/xe_gt_mcr.c | 90 ++++++++++++++++++++++-----
drivers/gpu/drm/xe/xe_gt_mcr.h | 1 +
drivers/gpu/drm/xe/xe_guc_ct.c | 23 ++++++-
drivers/gpu/drm/xe/xe_guc_ct.h | 1 +
drivers/gpu/drm/xe/xe_guc_tlb_inval.c | 4 +-
6 files changed, 106 insertions(+), 20 deletions(-)
--
2.55.0
next reply other threads:[~2026-07-22 13:42 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-22 0:46 Tales A. Mendonça [this message]
2026-07-22 0:46 ` [PATCH v1 1/4] drm/xe/mcr: Keep GT forcewake during MCR steering Tales A. Mendonça
2026-07-22 17:45 ` Tales A. Mendonça
2026-07-22 18:10 ` Matt Roper
2026-07-22 18:39 ` Tales A. Mendonça
2026-07-22 0:46 ` [PATCH v1 2/4] drm/xe/mcr: Sanitize steering semaphore on GT resume Tales A. Mendonça
2026-07-22 17:47 ` Tales A. Mendonça
2026-07-22 0:46 ` [PATCH v1 3/4] drm/xe/guc/ct: Queue G2H worker before flushing it in timeout paths Tales A. Mendonça
2026-07-22 17:48 ` Tales A. Mendonça
2026-07-22 21:10 ` Matthew Brost
2026-07-22 22:03 ` Matthew Brost
2026-07-22 23:28 ` Tales A. Mendonça
2026-07-23 0:26 ` Matthew Brost
2026-07-23 2:08 ` Tales A. Mendonça
2026-07-22 0:46 ` [PATCH v1 4/4] drm/xe: Raise hw_tlb_timeout to cover observed GuC ack latency Tales A. Mendonça
2026-07-22 17:27 ` ✗ LGCI.VerificationFailed: failure for drm/xe: MCR semaphore and TLB invalidation timeout fixes for ARL Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260722004654.744249-1-talesam@gmail.com \
--to=talesam@gmail.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=intel-xe@lists.freedesktop.org \
--cc=matthew.brost@intel.com \
--cc=rodrigo.vivi@intel.com \
--cc=thomas.hellstrom@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox