From: Matthew Brost <matthew.brost@intel.com>
To: "Tales A. Mendonça" <talesam@gmail.com>
Cc: <sashiko-reviews@lists.linux.dev>,
<dri-devel@lists.freedesktop.org>,
<intel-xe@lists.freedesktop.org>
Subject: Re: [PATCH v5 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout
Date: Mon, 21 Sep 2026 19:19:45 -0700 [thread overview]
Message-ID: <arHlwf7vmajXgR3l@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <CAHBRX4GqGrA-JhPct0K1GO+fO-uwPm+jJDiQcAQ9FBvBENcOjw@mail.gmail.com>
On Mon, Sep 21, 2026 at 05:28:22PM -0300, Tales A. Mendonça wrote:
> On Mon, Sep 21, 2026 at 03:33:00PM -0300, sashiko-bot@kernel.org wrote:
> > [High] drm/xe: MMIO access without PM reference in devcoredump on TLB timeout
>
> Correct on both counts, and it is my patch that introduces it.
>
> Tracing it: every pending invalidation fence holds a runtime PM
> reference, taken by xe_pm_runtime_get_noresume() in
> xe_tlb_inval_fence_init(). xe_tlb_inval_fence_signal() calls
> xe_tlb_inval_fence_fini(), which drops it. The timeout loop signals all
> the expired fences and only then, after dropping pending_lock, calls
> xe_devcoredump_gt(). devcoredump_snapshot() acquires forcewake without
> taking a PM reference of its own - it has always relied on the caller
> holding one, which the existing job-timeout caller does. The new
> GT-level caller may have just released the last one, so the device can
> begin autosuspending before the snapshot touches the hardware.
>
> Fixed for v6 by holding a reference across the capture, taken while the
> pending fences still guarantee the device is awake:
>
> if (!timedout_seqno) {
> /*
> * Hold a PM reference across the capture below. Every
> * pending fence holds one, so the device is awake
> * here, but signalling them may drop the last
> * reference and let it autosuspend before the
> * snapshot touches the hardware.
> */
> xe_pm_runtime_get_noresume(xe);
> }
>
> and released after xe_devcoredump_gt() returns.
>
> Taking it inside pending_lock is safe: xe_pm_runtime_get_noresume() goes
> through pm_runtime_get_if_in_use(), and its "Missing outer runtime PM
> protection" warning cannot fire here, because the fence reference is
> still held at that point.
>
This is correct.
Matt
> Tales
>
> Em seg., 21 de set. de 2026 às 15:33, <sashiko-bot@kernel.org> escreveu:
> >
> > Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> > - [High] drm/xe: MMIO access without PM reference in devcoredump on TLB timeout
> > --
> >
> > commit a463de937c804fe9be7ed86bfbfa297e4f35c466
> > Author: Tales A. Mendonça <talesam@gmail.com>
> >
> > drm/xe: Capture devcoredump on TLB invalidation timeout
> >
> > This commit adds xe_devcoredump_gt() to capture the GuC log and CT state
> > for GT-level hangs that are not tied to a specific queue or job. It then
> > hooks this capture mechanism into the TLB invalidation timeout path. This
> > helps in diagnosing firmware stalls where the invalidation request was
> > consumed but the ack was delayed.
> >
> > Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8678
> >
> > > diff --git a/drivers/gpu/drm/xe/xe_devcoredump.c b/drivers/gpu/drm/xe/xe_devcoredump.c
> > > --- a/drivers/gpu/drm/xe/xe_devcoredump.c
> > > +++ b/drivers/gpu/drm/xe/xe_devcoredump.c
> > [ ... ]
> > > @@ -323,40 +318,44 @@ static void devcoredump_snapshot(struct xe_devcoredump *coredump,
> > > struct xe_gt *gt,
> > > struct xe_exec_queue *q,
> > > struct xe_sched_job *job)
> > > {
> > [ ... ]
> > > /* keep going if fw fails as we still want to save the memory and SW data */
> > > - CLASS(xe_force_wake, fw_ref)(gt_to_fw(q->gt), XE_FORCEWAKE_ALL);
> > > + CLASS(xe_force_wake, fw_ref)(gt_to_fw(gt), XE_FORCEWAKE_ALL);
> >
> > [Severity: High]
> > If this capture path is reached without holding a PM runtime reference, does
> > acquiring the forcewake here attempt MMIO accesses while the device is in a
> > low-power state? Accessing MMIO registers when the device has entered
> > autosuspend (D3cold) typically causes unhandled PCIe errors or system
> > crashes.
> >
> > [ ... ]
> > > diff --git a/drivers/gpu/drm/xe/xe_tlb_inval.c b/drivers/gpu/drm/xe/xe_tlb_inval.c
> > > --- a/drivers/gpu/drm/xe/xe_tlb_inval.c
> > > +++ b/drivers/gpu/drm/xe/xe_tlb_inval.c
> > [ ... ]
> > > @@ -90,13 +98,27 @@ static void xe_tlb_inval_fence_timeout(struct work_struct *work)
> > > "TLB invalidation fence timeout, seqno=%d recv=%d",
> > > fence->seqno, tlb_inval->seqno_recv);
> > >
> > > + timedout_seqno = fence->seqno;
> > > +
> > > fence->base.error = -ETIME;
> > > xe_tlb_inval_fence_signal(fence);
> >
> > [Severity: High]
> > Can this sequence drop the last PM runtime reference before the devcoredump
> > is captured?
> >
> > Looking at xe_tlb_inval_fence_signal(), it calls xe_tlb_inval_fence_fini()
> > which drops the PM runtime reference (xe_pm_runtime_put). If this was the
> > last active reference, the device could begin autosuspending.
> >
> > > }
> > > if (!list_empty(&tlb_inval->pending_fences))
> > > queue_delayed_work(tlb_inval->timeout_wq, &tlb_inval->fence_tdr,
> > > timeout_delay);
> > > + seqno_recv = tlb_inval->seqno_recv;
> > > spin_unlock_irq(&tlb_inval->pending_lock);
> > > +
> > > + /*
> > > + * Capture the GuC log and CT state so the firmware side of the hang
> > > + * can be inspected; there is no queue or job to blame here. Must be
> > > + * outside pending_lock as the capture takes sleeping locks, hence
> > > + * @seqno_recv is sampled above while the lock is still held.
> > > + */
> > > + if (timedout_seqno)
> > > + xe_devcoredump_gt(tlb_inval_to_gt(tlb_inval),
> > > + "TLB invalidation fence timeout, seqno=%d recv=%d",
> > > + timedout_seqno, seqno_recv);
> >
> > Since the PM reference was dropped earlier in the loop, would this call to
> > xe_devcoredump_gt() proceed without holding one? If the thread gets
> > preempted long enough for the autosuspend delay to expire, the forcewake
> > operation inside devcoredump_snapshot() might hit the hardware while it is
> > unpowered.
> >
> > --
> > Sashiko AI review · https://sashiko.dev/#/patchset/20260921182121.308217-1-talesam@gmail.com?part=1
>
>
>
> --
> Com os cumprimentos,
>
> Tales A. Mendonça
> talesam.org
> communitybig.org
next prev parent reply other threads:[~2026-09-22 2:19 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-21 18:21 [PATCH v5 0/3] drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) Tales A. Mendonça
2026-09-21 18:21 ` [PATCH v5 1/3] drm/xe: Capture devcoredump on TLB invalidation timeout Tales A. Mendonça
2026-09-21 18:33 ` sashiko-bot
2026-09-21 20:28 ` Tales A. Mendonça
2026-09-22 2:19 ` Matthew Brost [this message]
2026-09-21 18:21 ` [PATCH v5 2/3] drm/xe: Log when a timed out TLB invalidation ack finally arrives Tales A. Mendonça
2026-09-21 18:21 ` [PATCH v5 3/3] drm/xe: Implement Wa_22016122933 Tales A. Mendonça
2026-09-21 18:48 ` Matthew Brost
2026-09-21 20:18 ` Tales A. Mendonça
2026-09-22 2:22 ` Matthew Brost
2026-09-21 18:27 ` ✗ CI.checkpatch: warning for drm/xe: fix GuC TLB invalidation ack stalls on ARL (Wa_22016122933) (rev4) Patchwork
2026-09-21 18:29 ` ✓ CI.KUnit: success " Patchwork
2026-09-21 19:28 ` ✗ Xe.CI.BAT: failure " Patchwork
2026-09-22 4:38 ` ✗ Xe.CI.FULL: " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arHlwf7vmajXgR3l@gsse-cloud1.jf.intel.com \
--to=matthew.brost@intel.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=intel-xe@lists.freedesktop.org \
--cc=sashiko-reviews@lists.linux.dev \
--cc=talesam@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox