From: Matthew Brost <matthew.brost@intel.com>
To: Arvind Yadav <arvind.yadav@intel.com>
Cc: <intel-xe@lists.freedesktop.org>,
<dri-devel@lists.freedesktop.org>, <rodrigo.vivi@intel.com>,
<himal.prasad.ghimiray@intel.com>,
<thomas.hellstrom@linux.intel.com>
Subject: Re: [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery
Date: Fri, 18 Sep 2026 15:31:44 -0700 [thread overview]
Message-ID: <aq270L/+biIRJHkX@gsse-cloud1.jf.intel.com> (raw)
In-Reply-To: <20260916095337.3104891-1-arvind.yadav@intel.com>
On Wed, Sep 16, 2026 at 03:23:32PM +0530, Arvind Yadav wrote:
> On BMG, terminating a process with Ctrl+C while GPU work is running under
> VRAM pressure can cause an RCS page fault followed by a migration queue
> timeout on BCS8 (guc_id 0). The driver resets the GT, but the pending
> migration job can time out again and eventually wedge the device.
>
If we can just avoid faults on ctrl-c or segfault will this series be
needed (or at least only a subset of this series)?
I suggested a way to kill the exec queue first + sync wait on those here
[1] before clobbering the GPU page tables? I don't really care who works
on [1].
We still need root cause why a RCS fault affects the BCS engine - it
shouldn't and need investigation.
Matt
[1] https://patchwork.freedesktop.org/patch/752845/?series=173889&rev=1#comment_1391707
> BCS8 is reserved for paging and runs migration and VM bind work.
> The exact hardware link between the RCS fault and the BCS8 stall
> is still under investigation. This series addresses the teardown
> and recovery problems found while debugging this failure.
>
> During file close, exec queue cleanup starts asynchronously, but the VM
> mappings can be removed before that cleanup finishes. A missed wakeup in
> the GuC disable-completion handler can also turn a completed operation
> into a five-second timeout and an unnecessary GT reset.
>
> The reset replay path rewinds the software ring tail to the oldest pending
> job, but leaves the LRC head at its saved position. This leaves different
> starting positions for replay.
>
> The four patches address these paths:
> 1. Clear pending-disable state before waking waiters, so a completed disable
> does not appear to time out.
>
> 2. Mark VMs as closing before queue cleanup. Reject new work and page faults
> on closing VMs, while allowing existing SVM invalidation to drain mappings.
>
> 3. Keep VM mappings alive until queue cleanup completes. File close uses one
> five-second queue-wait budget across all VMs, then defers any remaining teardown.
> VM destroy defers without waiting. Device references protect deferred close and
> final VM destruction on the module-lifetime destroy workqueue.
>
> 4. Set the software tail and LRC head and tail to the oldest pending job before
> resubmitting jobs after a GT reset.
>
> The five-second budget applies only to the new queue-cleanup wait.
> Existing teardown waits are unchanged.
>
> Arvind Yadav (5):
> drm/xe: Hold a device reference across deferred VM destruction
> drm/xe/guc: Wake disable waiters after clearing pending state
> drm/xe: Mark VMs as closing before queue cleanup
> drm/xe: Defer VM teardown until exec queue cleanup completes
> drm/xe/guc: Reset LRC ring pointers before replay
>
> drivers/gpu/drm/xe/xe_device.c | 25 +++-
> drivers/gpu/drm/xe/xe_exec_queue.c | 22 +++-
> drivers/gpu/drm/xe/xe_guc_submit.c | 60 +++++----
> drivers/gpu/drm/xe/xe_module.c | 6 +-
> drivers/gpu/drm/xe/xe_pagefault.c | 2 +-
> drivers/gpu/drm/xe/xe_svm.c | 3 +-
> drivers/gpu/drm/xe/xe_vm.c | 205 +++++++++++++++++++++++++----
> drivers/gpu/drm/xe/xe_vm.h | 14 +-
> drivers/gpu/drm/xe/xe_vm_types.h | 16 +++
> 9 files changed, 289 insertions(+), 64 deletions(-)
>
> --
> 2.43.0
>
next prev parent reply other threads:[~2026-09-18 22:31 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 9:53 [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Arvind Yadav
2026-09-16 9:53 ` [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction Arvind Yadav
2026-09-18 3:22 ` Matthew Brost
2026-09-18 7:22 ` Thomas Hellström
2026-09-18 22:41 ` Matthew Brost
2026-09-21 7:06 ` Thomas Hellström
2026-09-16 9:53 ` [PATCH 2/5] drm/xe/guc: Wake disable waiters after clearing pending state Arvind Yadav
2026-09-18 22:36 ` Matthew Brost
2026-09-21 6:46 ` Yadav, Arvind
2026-09-16 9:53 ` [PATCH 3/5] drm/xe: Mark VMs as closing before queue cleanup Arvind Yadav
2026-09-16 9:53 ` [PATCH 4/5] drm/xe: Defer VM teardown until exec queue cleanup completes Arvind Yadav
2026-09-16 9:53 ` [PATCH 5/5] drm/xe/guc: Reset LRC ring pointers before replay Arvind Yadav
2026-09-16 10:01 ` ✓ CI.KUnit: success for drm/xe: Fix VM teardown and migration queue recovery Patchwork
2026-09-16 10:59 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-16 12:11 ` ✓ Xe.CI.FULL: " Patchwork
2026-09-18 22:31 ` Matthew Brost [this message]
2026-09-24 10:06 ` [PATCH 0/5] " Yadav, Arvind
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aq270L/+biIRJHkX@gsse-cloud1.jf.intel.com \
--to=matthew.brost@intel.com \
--cc=arvind.yadav@intel.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=himal.prasad.ghimiray@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=rodrigo.vivi@intel.com \
--cc=thomas.hellstrom@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox