Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Arvind Yadav <arvind.yadav@intel.com>
To: intel-xe@lists.freedesktop.org, dri-devel@lists.freedesktop.org
Cc: rodrigo.vivi@intel.com, matthew.brost@intel.com,
	himal.prasad.ghimiray@intel.com,
	thomas.hellstrom@linux.intel.com
Subject: [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery
Date: Wed, 16 Sep 2026 15:23:32 +0530	[thread overview]
Message-ID: <20260916095337.3104891-1-arvind.yadav@intel.com> (raw)

On BMG, terminating a process with Ctrl+C while GPU work is running under
VRAM pressure can cause an RCS page fault followed by a migration queue
timeout on BCS8 (guc_id 0). The driver resets the GT, but the pending
migration job can time out again and eventually wedge the device.

BCS8 is reserved for paging and runs migration and VM bind work.
The exact hardware link between the RCS fault and the BCS8 stall
is still under investigation. This series addresses the teardown
and recovery problems found while debugging this failure.

During file close, exec queue cleanup starts asynchronously, but the VM
mappings can be removed before that cleanup finishes. A missed wakeup in
the GuC disable-completion handler can also turn a completed operation
into a five-second timeout and an unnecessary GT reset.

The reset replay path rewinds the software ring tail to the oldest pending
job, but leaves the LRC head at its saved position. This leaves different
starting positions for replay.

The four patches address these paths:
1. Clear pending-disable state before waking waiters, so a completed disable
does not appear to time out.

2. Mark VMs as closing before queue cleanup. Reject new work and page faults
on closing VMs, while allowing existing SVM invalidation to drain mappings.

3. Keep VM mappings alive until queue cleanup completes. File close uses one
five-second queue-wait budget across all VMs, then defers any remaining teardown.
VM destroy defers without waiting. Device references protect deferred close and
final VM destruction on the module-lifetime destroy workqueue.

4. Set the software tail and LRC head and tail to the oldest pending job before
resubmitting jobs after a GT reset.

The five-second budget applies only to the new queue-cleanup wait.
Existing teardown waits are unchanged.

Arvind Yadav (5):
  drm/xe: Hold a device reference across deferred VM destruction
  drm/xe/guc: Wake disable waiters after clearing pending state
  drm/xe: Mark VMs as closing before queue cleanup
  drm/xe: Defer VM teardown until exec queue cleanup completes
  drm/xe/guc: Reset LRC ring pointers before replay

 drivers/gpu/drm/xe/xe_device.c     |  25 +++-
 drivers/gpu/drm/xe/xe_exec_queue.c |  22 +++-
 drivers/gpu/drm/xe/xe_guc_submit.c |  60 +++++----
 drivers/gpu/drm/xe/xe_module.c     |   6 +-
 drivers/gpu/drm/xe/xe_pagefault.c  |   2 +-
 drivers/gpu/drm/xe/xe_svm.c        |   3 +-
 drivers/gpu/drm/xe/xe_vm.c         | 205 +++++++++++++++++++++++++----
 drivers/gpu/drm/xe/xe_vm.h         |  14 +-
 drivers/gpu/drm/xe/xe_vm_types.h   |  16 +++
 9 files changed, 289 insertions(+), 64 deletions(-)

-- 
2.43.0


             reply	other threads:[~2026-09-16  9:53 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16  9:53 Arvind Yadav [this message]
2026-09-16  9:53 ` [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction Arvind Yadav
2026-09-18  3:22   ` Matthew Brost
2026-09-18  7:22     ` Thomas Hellström
2026-09-18 22:41       ` Matthew Brost
2026-09-21  7:06         ` Thomas Hellström
2026-09-16  9:53 ` [PATCH 2/5] drm/xe/guc: Wake disable waiters after clearing pending state Arvind Yadav
2026-09-18 22:36   ` Matthew Brost
2026-09-21  6:46     ` Yadav, Arvind
2026-09-16  9:53 ` [PATCH 3/5] drm/xe: Mark VMs as closing before queue cleanup Arvind Yadav
2026-09-16  9:53 ` [PATCH 4/5] drm/xe: Defer VM teardown until exec queue cleanup completes Arvind Yadav
2026-09-16  9:53 ` [PATCH 5/5] drm/xe/guc: Reset LRC ring pointers before replay Arvind Yadav
2026-09-16 10:01 ` ✓ CI.KUnit: success for drm/xe: Fix VM teardown and migration queue recovery Patchwork
2026-09-16 10:59 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-16 12:11 ` ✓ Xe.CI.FULL: " Patchwork
2026-09-18 22:31 ` [PATCH 0/5] " Matthew Brost
2026-09-24 10:06   ` Yadav, Arvind

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260916095337.3104891-1-arvind.yadav@intel.com \
    --to=arvind.yadav@intel.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=himal.prasad.ghimiray@intel.com \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=matthew.brost@intel.com \
    --cc=rodrigo.vivi@intel.com \
    --cc=thomas.hellstrom@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox