* [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery
@ 2026-09-16 9:53 Arvind Yadav
2026-09-16 9:53 ` [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction Arvind Yadav
` (5 more replies)
0 siblings, 6 replies; 14+ messages in thread
From: Arvind Yadav @ 2026-09-16 9:53 UTC (permalink / raw)
To: intel-xe, dri-devel
Cc: rodrigo.vivi, matthew.brost, himal.prasad.ghimiray,
thomas.hellstrom
On BMG, terminating a process with Ctrl+C while GPU work is running under
VRAM pressure can cause an RCS page fault followed by a migration queue
timeout on BCS8 (guc_id 0). The driver resets the GT, but the pending
migration job can time out again and eventually wedge the device.
BCS8 is reserved for paging and runs migration and VM bind work.
The exact hardware link between the RCS fault and the BCS8 stall
is still under investigation. This series addresses the teardown
and recovery problems found while debugging this failure.
During file close, exec queue cleanup starts asynchronously, but the VM
mappings can be removed before that cleanup finishes. A missed wakeup in
the GuC disable-completion handler can also turn a completed operation
into a five-second timeout and an unnecessary GT reset.
The reset replay path rewinds the software ring tail to the oldest pending
job, but leaves the LRC head at its saved position. This leaves different
starting positions for replay.
The four patches address these paths:
1. Clear pending-disable state before waking waiters, so a completed disable
does not appear to time out.
2. Mark VMs as closing before queue cleanup. Reject new work and page faults
on closing VMs, while allowing existing SVM invalidation to drain mappings.
3. Keep VM mappings alive until queue cleanup completes. File close uses one
five-second queue-wait budget across all VMs, then defers any remaining teardown.
VM destroy defers without waiting. Device references protect deferred close and
final VM destruction on the module-lifetime destroy workqueue.
4. Set the software tail and LRC head and tail to the oldest pending job before
resubmitting jobs after a GT reset.
The five-second budget applies only to the new queue-cleanup wait.
Existing teardown waits are unchanged.
Arvind Yadav (5):
drm/xe: Hold a device reference across deferred VM destruction
drm/xe/guc: Wake disable waiters after clearing pending state
drm/xe: Mark VMs as closing before queue cleanup
drm/xe: Defer VM teardown until exec queue cleanup completes
drm/xe/guc: Reset LRC ring pointers before replay
drivers/gpu/drm/xe/xe_device.c | 25 +++-
drivers/gpu/drm/xe/xe_exec_queue.c | 22 +++-
drivers/gpu/drm/xe/xe_guc_submit.c | 60 +++++----
drivers/gpu/drm/xe/xe_module.c | 6 +-
drivers/gpu/drm/xe/xe_pagefault.c | 2 +-
drivers/gpu/drm/xe/xe_svm.c | 3 +-
drivers/gpu/drm/xe/xe_vm.c | 205 +++++++++++++++++++++++++----
drivers/gpu/drm/xe/xe_vm.h | 14 +-
drivers/gpu/drm/xe/xe_vm_types.h | 16 +++
9 files changed, 289 insertions(+), 64 deletions(-)
--
2.43.0
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction
2026-09-16 9:53 [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Arvind Yadav
@ 2026-09-16 9:53 ` Arvind Yadav
2026-09-18 3:22 ` Matthew Brost
2026-09-16 9:53 ` [PATCH 2/5] drm/xe/guc: Wake disable waiters after clearing pending state Arvind Yadav
` (4 subsequent siblings)
5 siblings, 1 reply; 14+ messages in thread
From: Arvind Yadav @ 2026-09-16 9:53 UTC (permalink / raw)
To: intel-xe, dri-devel
Cc: rodrigo.vivi, matthew.brost, himal.prasad.ghimiray,
thomas.hellstrom
xe_vm_free() is the drm_gpuvm vm_free callback. It hands the final
teardown to vm_destroy_work_func() on a workqueue and returns.
drm_gpuvm_free() drops its device reference immediately after the
callback returns:
gpuvm->ops->vm_free(gpuvm);
drm_dev_put(drm);
vm_destroy_work_func() then keeps using device state: xe_pm_runtime_put()
for an LR mode VM, ttm_lru_bulk_move_fini() on xe->ttm, and the tile
iteration. If the freed VM held the last device reference, the work runs
against a released xe_device.
Take a device reference in xe_vm_free() and drop it once
vm_destroy_work_func() has finished using the device.
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
---
drivers/gpu/drm/xe/xe_vm.c | 9 +++++++++
1 file changed, 9 insertions(+)
diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
index efa5ff6cc823..264bdab75de2 100644
--- a/drivers/gpu/drm/xe/xe_vm.c
+++ b/drivers/gpu/drm/xe/xe_vm.c
@@ -2054,12 +2054,21 @@ static void vm_destroy_work_func(struct work_struct *w)
xe_file_put(vm->xef);
kfree(vm);
+
+ drm_dev_put(&xe->drm);
}
static void xe_vm_free(struct drm_gpuvm *gpuvm)
{
struct xe_vm *vm = container_of(gpuvm, struct xe_vm, gpuvm);
+ /*
+ * drm_gpuvm drops its device reference as soon as this callback
+ * returns, but vm_destroy_work_func() still uses device state. Hold a
+ * reference across the deferred work.
+ */
+ drm_dev_get(&vm->xe->drm);
+
/* To destroy the VM we need to be able to sleep */
queue_work(system_dfl_wq, &vm->destroy_work);
}
--
2.43.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH 2/5] drm/xe/guc: Wake disable waiters after clearing pending state
2026-09-16 9:53 [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Arvind Yadav
2026-09-16 9:53 ` [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction Arvind Yadav
@ 2026-09-16 9:53 ` Arvind Yadav
2026-09-18 22:36 ` Matthew Brost
2026-09-16 9:53 ` [PATCH 3/5] drm/xe: Mark VMs as closing before queue cleanup Arvind Yadav
` (3 subsequent siblings)
5 siblings, 1 reply; 14+ messages in thread
From: Arvind Yadav @ 2026-09-16 9:53 UTC (permalink / raw)
To: intel-xe, dri-devel
Cc: rodrigo.vivi, matthew.brost, himal.prasad.ghimiray,
thomas.hellstrom
disable_scheduling_deregister() waits for any pending scheduling
operation to complete before destroying an exec queue.
For a disable completion, handle_sched_done() can wake the waiter before
clearing pending_disable, or not wake it at all. The waiter can then miss
a completed operation and expire after five seconds. This causes a
spurious GT reset and immediate TDR.
Clear pending_disable before waking the waitqueue on every completion
path. Sample the destroyed state before clearing the pending state, as
required by the existing destroy protocol.
Also wake CT waiters after a suspend completion.
The queue remains alive until deregistration completes, so waking
waiters before sending the deregister request does not release it.
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
---
drivers/gpu/drm/xe/xe_guc_submit.c | 42 +++++++++++++++++-------------
1 file changed, 24 insertions(+), 18 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
index f3ba8abfc228..ca24a77dfb26 100644
--- a/drivers/gpu/drm/xe/xe_guc_submit.c
+++ b/drivers/gpu/drm/xe/xe_guc_submit.c
@@ -3265,26 +3265,32 @@ static void handle_sched_done(struct xe_guc *guc, struct xe_exec_queue *q,
if (q->guc->suspend_pending) {
clear_exec_queue_pending_disable(q);
suspend_fence_signal(q);
+
+ /*
+ * Publish the cleared state before waking waiters.
+ */
+ smp_wmb();
+ wake_up_all(&guc->ct.wq);
} else {
- if (exec_queue_banned(q)) {
- smp_wmb();
- wake_up_all(&guc->ct.wq);
- }
- if (exec_queue_destroyed(q)) {
- /*
- * Make sure to clear the pending_disable only
- * after sampling the destroyed state. We want
- * to ensure we don't trigger the unregister too
- * early with something intending to only
- * disable scheduling. The caller doing the
- * destroy must wait for an ongoing
- * pending_disable before marking as destroyed.
- */
- clear_exec_queue_pending_disable(q);
+ bool destroyed = exec_queue_destroyed(q);
+
+ /*
+ * Make sure to clear pending_disable only after sampling
+ * the destroyed state. The caller doing the destroy must
+ * wait for an ongoing disable before marking the queue
+ * destroyed.
+ */
+ clear_exec_queue_pending_disable(q);
+
+ /*
+ * Publish the cleared state before waking waiters.
+ */
+ smp_wmb();
+ wake_up_all(&guc->ct.wq);
+
+ /* The queue remains alive until DEREGISTER_DONE. */
+ if (destroyed)
deregister_exec_queue(guc, q);
- } else {
- clear_exec_queue_pending_disable(q);
- }
}
}
}
--
2.43.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH 3/5] drm/xe: Mark VMs as closing before queue cleanup
2026-09-16 9:53 [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Arvind Yadav
2026-09-16 9:53 ` [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction Arvind Yadav
2026-09-16 9:53 ` [PATCH 2/5] drm/xe/guc: Wake disable waiters after clearing pending state Arvind Yadav
@ 2026-09-16 9:53 ` Arvind Yadav
2026-09-16 9:53 ` [PATCH 4/5] drm/xe: Defer VM teardown until exec queue cleanup completes Arvind Yadav
` (2 subsequent siblings)
5 siblings, 0 replies; 14+ messages in thread
From: Arvind Yadav @ 2026-09-16 9:53 UTC (permalink / raw)
To: intel-xe, dri-devel
Cc: rodrigo.vivi, matthew.brost, himal.prasad.ghimiray,
thomas.hellstrom
Queue cleanup is asynchronous. New VM work can therefore race with file
close or VM destroy while cleanup is starting.
Add a closing state and set it before starting queue cleanup. Treat a
closing VM as unavailable for new exec, bind, rebind and page-fault work.
Keep vm->size valid while the VM is closing so SVM invalidation can still
drain existing mappings.
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
---
drivers/gpu/drm/xe/xe_device.c | 4 ++++
drivers/gpu/drm/xe/xe_pagefault.c | 2 +-
drivers/gpu/drm/xe/xe_svm.c | 3 ++-
drivers/gpu/drm/xe/xe_vm.c | 18 +++++++++++++++++-
drivers/gpu/drm/xe/xe_vm.h | 9 ++++++++-
drivers/gpu/drm/xe/xe_vm_types.h | 1 +
6 files changed, 33 insertions(+), 4 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c
index 205cb4e7f9e8..954f0965deb8 100644
--- a/drivers/gpu/drm/xe/xe_device.c
+++ b/drivers/gpu/drm/xe/xe_device.c
@@ -179,6 +179,10 @@ static void xe_file_close(struct drm_device *dev, struct drm_file *file)
guard(xe_pm_runtime)(xe);
+ /* Block new VM work before starting asynchronous queue teardown. */
+ xa_for_each(&xef->vm.xa, idx, vm)
+ xe_vm_close_start(vm);
+
/*
* No need for exec_queue.lock here as there is no contention for it
* when FD is closing as IOCTLs presumably can't be modifying the
diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
index aeb56ff5d58e..4bdd714b9f88 100644
--- a/drivers/gpu/drm/xe/xe_pagefault.c
+++ b/drivers/gpu/drm/xe/xe_pagefault.c
@@ -265,7 +265,7 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
down_read(&vm->lock);
- if (xe_vm_is_closed(vm)) {
+ if (xe_vm_is_closed_or_banned(vm)) {
err = -ENOENT;
goto unlock_vm;
}
diff --git a/drivers/gpu/drm/xe/xe_svm.c b/drivers/gpu/drm/xe/xe_svm.c
index 6c3033fc4db7..c2a9dd98f363 100644
--- a/drivers/gpu/drm/xe/xe_svm.c
+++ b/drivers/gpu/drm/xe/xe_svm.c
@@ -218,7 +218,8 @@ xe_svm_range_notifier_event_end(struct xe_vm *vm, struct drm_gpusvm_range *r,
drm_gpusvm_unmap_pages(&vm->svm.gpusvm, &(to_xe_range(r)->pages),
drm_gpusvm_range_size(r) >> PAGE_SHIFT, &ctx);
- if (!xe_vm_is_closed(vm) && mmu_range->event == MMU_NOTIFY_UNMAP)
+ if (!xe_vm_is_closed(vm) && !xe_vm_is_closing(vm) &&
+ mmu_range->event == MMU_NOTIFY_UNMAP)
xe_svm_garbage_collector_add_range(vm, to_xe_range(r),
mmu_range);
}
diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
index 264bdab75de2..5d0b616a27b3 100644
--- a/drivers/gpu/drm/xe/xe_vm.c
+++ b/drivers/gpu/drm/xe/xe_vm.c
@@ -1913,6 +1913,20 @@ static void xe_vm_close(struct xe_vm *vm)
drm_dev_exit(idx);
}
+void xe_vm_close_start(struct xe_vm *vm)
+{
+ down_write(&vm->lock);
+ if (xe_vm_in_fault_mode(vm))
+ xe_svm_notifier_lock(vm);
+
+ /* Keep size valid so SVM invalidation still performs its full drain. */
+ vm->flags |= XE_VM_FLAG_CLOSING;
+
+ if (xe_vm_in_fault_mode(vm))
+ xe_svm_notifier_unlock(vm);
+ up_write(&vm->lock);
+}
+
void xe_vm_close_and_put(struct xe_vm *vm)
{
LIST_HEAD(contested);
@@ -2214,8 +2228,10 @@ int xe_vm_destroy_ioctl(struct drm_device *dev, void *data,
xa_erase(&xef->vm.xa, args->vm_id);
mutex_unlock(&xef->vm.lock);
- if (!err)
+ if (!err) {
+ xe_vm_close_start(vm);
xe_vm_close_and_put(vm);
+ }
return err;
}
diff --git a/drivers/gpu/drm/xe/xe_vm.h b/drivers/gpu/drm/xe/xe_vm.h
index c5b900f38ded..70456ffe27e2 100644
--- a/drivers/gpu/drm/xe/xe_vm.h
+++ b/drivers/gpu/drm/xe/xe_vm.h
@@ -59,6 +59,11 @@ static inline bool xe_vm_is_closed(struct xe_vm *vm)
return !vm->size;
}
+static inline bool xe_vm_is_closing(struct xe_vm *vm)
+{
+ return vm->flags & XE_VM_FLAG_CLOSING;
+}
+
static inline bool xe_vm_is_banned(struct xe_vm *vm)
{
return vm->flags & XE_VM_FLAG_BANNED;
@@ -67,7 +72,8 @@ static inline bool xe_vm_is_banned(struct xe_vm *vm)
static inline bool xe_vm_is_closed_or_banned(struct xe_vm *vm)
{
lockdep_assert_held(&vm->lock);
- return xe_vm_is_closed(vm) || xe_vm_is_banned(vm);
+ return xe_vm_is_closed(vm) || xe_vm_is_banned(vm) ||
+ xe_vm_is_closing(vm);
}
struct xe_vma *
@@ -214,6 +220,7 @@ int xe_vm_get_property_ioctl(struct drm_device *dev, void *data,
struct drm_file *file);
void xe_vm_close_and_put(struct xe_vm *vm);
+void xe_vm_close_start(struct xe_vm *vm);
static inline bool xe_vm_in_fault_mode(struct xe_vm *vm)
{
diff --git a/drivers/gpu/drm/xe/xe_vm_types.h b/drivers/gpu/drm/xe/xe_vm_types.h
index 648031e64145..16b051a0cf10 100644
--- a/drivers/gpu/drm/xe/xe_vm_types.h
+++ b/drivers/gpu/drm/xe/xe_vm_types.h
@@ -286,6 +286,7 @@ struct xe_vm {
#define XE_VM_FLAG_SET_TILE_ID(tile) FIELD_PREP(GENMASK(7, 6), (tile)->id)
#define XE_VM_FLAG_GSC BIT(8)
#define XE_VM_FLAG_NO_VM_OVERCOMMIT BIT(9)
+#define XE_VM_FLAG_CLOSING BIT(10)
unsigned long flags;
/**
--
2.43.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH 4/5] drm/xe: Defer VM teardown until exec queue cleanup completes
2026-09-16 9:53 [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Arvind Yadav
` (2 preceding siblings ...)
2026-09-16 9:53 ` [PATCH 3/5] drm/xe: Mark VMs as closing before queue cleanup Arvind Yadav
@ 2026-09-16 9:53 ` Arvind Yadav
2026-09-16 9:53 ` [PATCH 5/5] drm/xe/guc: Reset LRC ring pointers before replay Arvind Yadav
2026-09-18 22:31 ` [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Matthew Brost
5 siblings, 0 replies; 14+ messages in thread
From: Arvind Yadav @ 2026-09-16 9:53 UTC (permalink / raw)
To: intel-xe, dri-devel
Cc: rodrigo.vivi, matthew.brost, himal.prasad.ghimiray,
thomas.hellstrom
Exec queue destruction is asynchronous. Closing a VM immediately after
killing its queues can remove mappings before GuC cleanup has finished.
Track the queues using each VM and keep the bind queue's user_vm
reference until final queue destruction. Start queue cleanup before
removing the VM mappings.
The count includes VM-owned bind queues. Release the VM's references
to these queues before waiting for their asynchronous cleanup.
During file close, use one five-second wait budget across all VMs.
If queues remain, defer VM close until the last queue is removed.
The existing runtime-PM reference remains held during this wait.
VM destroy does not wait for queue cleanup. If user exec queues still
reference the VM, its mappings remain after the ioctl returns until
those queues finish cleanup. Userspace must destroy the remaining queue
handles or close the DRM file to release them.
The five-second budget limits the new queue wait, not the lifetime of
the retained mappings. Recovery continues through the existing GuC
reset and teardown paths.
Run deferred close and final VM destruction on the module-lifetime
destroy workqueue. Hold an extra VM reference until the close worker
has released its PM and unplug guards. Final VM destruction retains
the device reference until cleanup finishes.
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
---
drivers/gpu/drm/xe/xe_device.c | 21 +++-
drivers/gpu/drm/xe/xe_exec_queue.c | 22 ++--
drivers/gpu/drm/xe/xe_module.c | 6 +-
drivers/gpu/drm/xe/xe_vm.c | 178 +++++++++++++++++++++++++----
drivers/gpu/drm/xe/xe_vm.h | 5 +
drivers/gpu/drm/xe/xe_vm_types.h | 15 +++
6 files changed, 214 insertions(+), 33 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c
index 954f0965deb8..8966a3bd6eac 100644
--- a/drivers/gpu/drm/xe/xe_device.c
+++ b/drivers/gpu/drm/xe/xe_device.c
@@ -8,6 +8,7 @@
#include <linux/aperture.h>
#include <linux/delay.h>
#include <linux/fault-inject.h>
+#include <linux/jiffies.h>
#include <linux/units.h>
#include <drm/drm_client.h>
@@ -175,6 +176,8 @@ static void xe_file_close(struct drm_device *dev, struct drm_file *file)
struct xe_file *xef = file->driver_priv;
struct xe_vm *vm;
struct xe_exec_queue *q;
+ bool close_deferred = false;
+ unsigned long deadline;
unsigned long idx;
guard(xe_pm_runtime)(xe);
@@ -195,8 +198,24 @@ static void xe_file_close(struct drm_device *dev, struct drm_file *file)
xe_exec_queue_kill(q);
xe_exec_queue_put(q);
}
+
+ /* Start all bind queue teardown before spending the shared wait budget. */
xa_for_each(&xef->vm.xa, idx, vm)
- xe_vm_close_and_put(vm);
+ xe_vm_kill_bind_queues(vm);
+
+ deadline = jiffies + HZ * 5;
+ xa_for_each(&xef->vm.xa, idx, vm) {
+ unsigned long now = jiffies;
+ unsigned long timeout =
+ time_before(now, deadline) ? deadline - now : 0;
+
+ close_deferred |=
+ xe_vm_close_and_put_deferred(vm, timeout);
+ }
+
+ if (close_deferred)
+ drm_dbg(&xe->drm,
+ "VM teardown deferred while exec queues are being stopped\n");
scoped_guard(mutex, &xef->mmio_gem.lock) {
if (xef->mmio_gem.pci_barrier) {
diff --git a/drivers/gpu/drm/xe/xe_exec_queue.c b/drivers/gpu/drm/xe/xe_exec_queue.c
index e63559a2f582..82407dddfbbe 100644
--- a/drivers/gpu/drm/xe/xe_exec_queue.c
+++ b/drivers/gpu/drm/xe/xe_exec_queue.c
@@ -154,9 +154,16 @@ static void __xe_exec_queue_free(struct xe_exec_queue *q)
if (q->vm) {
xe_vm_remove_exec_queue(q->vm, q);
+ if (q->vm->xef)
+ xe_vm_remove_close_queue(q->vm);
xe_vm_put(q->vm);
}
+ if (q->user_vm) {
+ xe_vm_remove_close_queue(q->user_vm);
+ xe_vm_put(q->user_vm);
+ }
+
if (q->xef)
xe_file_put(q->xef);
@@ -250,8 +257,12 @@ static struct xe_exec_queue *__xe_exec_queue_alloc(struct xe_device *xe,
}
}
- if (vm)
+ if (vm) {
q->vm = xe_vm_get(vm);
+ /* vm->xef stays unchanged until final VM destruction. */
+ if (vm->xef)
+ xe_vm_add_close_queue(vm);
+ }
if (extensions) {
/*
@@ -617,8 +628,10 @@ struct xe_exec_queue *xe_exec_queue_create_bind(struct xe_device *xe,
return ERR_PTR(err);
}
- if (user_vm)
+ if (user_vm) {
q->user_vm = xe_vm_get(user_vm);
+ xe_vm_add_close_queue(user_vm);
+ }
}
return q;
@@ -657,11 +670,6 @@ void xe_exec_queue_destroy(struct kref *ref)
xe_exec_queue_put(eq);
}
- if (q->user_vm) {
- xe_vm_put(q->user_vm);
- q->user_vm = NULL;
- }
-
q->ops->destroy(q);
}
diff --git a/drivers/gpu/drm/xe/xe_module.c b/drivers/gpu/drm/xe/xe_module.c
index 4bc28dfc1992..4e500156ed9b 100644
--- a/drivers/gpu/drm/xe/xe_module.c
+++ b/drivers/gpu/drm/xe/xe_module.c
@@ -114,9 +114,9 @@ static void xe_destroy_wq_module_exit(void)
* xe_destroy_wq_queue() - Queue work on the destroy workqueue
* @work: work item to queue
*
- * The destroy workqueue has module lifetime and is used for GuC exec queue
- * teardown that can outlive a single xe_device. SVM pagemap destroy uses the
- * per-device xe->destroy_wq instead.
+ * Queue and VM cleanup can drop the last DRM device reference. This workqueue
+ * has module lifetime so device teardown cannot destroy a worker's own queue.
+ * SVM pagemap destroy uses the per-device xe->destroy_wq instead.
*
* Return: %true if @work was queued, %false if it was already pending.
*/
diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
index 5d0b616a27b3..152540a4741b 100644
--- a/drivers/gpu/drm/xe/xe_vm.c
+++ b/drivers/gpu/drm/xe/xe_vm.c
@@ -29,6 +29,7 @@
#include "xe_exec_queue.h"
#include "xe_gt.h"
#include "xe_migrate.h"
+#include "xe_module.h"
#include "xe_pat.h"
#include "xe_pm.h"
#include "xe_preempt_fence.h"
@@ -1558,6 +1559,7 @@ static const struct xe_pt_ops xelp_pt_ops = {
.pde_encode_bo = xelp_pde_encode_bo,
};
+static void vm_close_work_func(struct work_struct *w);
static void vm_destroy_work_func(struct work_struct *w);
/**
@@ -1701,6 +1703,10 @@ struct xe_vm *xe_vm_create(struct xe_device *xe, u32 flags, struct xe_file *xef)
ttm_lru_bulk_move_init(&vm->lru_bulk_move);
INIT_WORK(&vm->destroy_work, vm_destroy_work_func);
+ atomic_set(&vm->close.num_exec_queues, 0);
+ init_waitqueue_head(&vm->close.wq);
+ spin_lock_init(&vm->close.lock);
+ INIT_WORK(&vm->close.work, vm_close_work_func);
INIT_LIST_HEAD(&vm->preempt.exec_queues);
for (id = 0; id < XE_MAX_TILES_PER_DEVICE * XE_MAX_GT_PER_TILE; ++id)
@@ -1948,25 +1954,7 @@ void xe_vm_close_and_put(struct xe_vm *vm)
if (xe_vm_in_fault_mode(vm))
xe_svm_close(vm);
- down_write(&vm->lock);
- for_each_tile(tile, xe, id) {
- if (vm->q[id]) {
- int i;
-
- xe_exec_queue_last_fence_put(vm->q[id], vm);
- for_each_tlb_inval(i)
- xe_exec_queue_tlb_inval_last_fence_put(vm->q[id], vm, i);
- }
- }
- up_write(&vm->lock);
-
- for_each_tile(tile, xe, id) {
- if (vm->q[id]) {
- xe_exec_queue_kill(vm->q[id]);
- xe_exec_queue_put(vm->q[id]);
- vm->q[id] = NULL;
- }
- }
+ xe_vm_kill_bind_queues(vm);
down_write(&vm->lock);
xe_vm_lock(vm, false);
@@ -2038,6 +2026,148 @@ void xe_vm_close_and_put(struct xe_vm *vm)
xe_vm_put(vm);
}
+static void vm_close_work_func(struct work_struct *w)
+{
+ struct xe_vm *vm = container_of(w, struct xe_vm, close.work);
+ struct xe_device *xe = vm->xe;
+ int idx;
+
+ /* Keep the VM and device alive until the PM and unplug guards unwind. */
+ xe_vm_get(vm);
+
+ if (drm_dev_enter(&xe->drm, &idx)) {
+ xe_pm_runtime_get(xe);
+ xe_vm_close_and_put(vm);
+ xe_pm_runtime_put(xe);
+ drm_dev_exit(idx);
+ } else {
+ /*
+ * The device is gone. Drop the VM without a runtime PM reference.
+ * xe_vm_close() guards its own register access.
+ */
+ xe_vm_close_and_put(vm);
+ }
+
+ xe_vm_put(vm);
+}
+
+/**
+ * xe_vm_kill_bind_queues() - Kill and release the VM bind queues
+ * @vm: The VM whose bind queues should be stopped
+ *
+ * Callers must serialize teardown of @vm. Later calls are harmless because
+ * the first call clears the VM bind queue pointers.
+ */
+void xe_vm_kill_bind_queues(struct xe_vm *vm)
+{
+ struct xe_device *xe = vm->xe;
+ struct xe_tile *tile;
+ u8 id;
+
+ down_write(&vm->lock);
+ for_each_tile(tile, xe, id) {
+ if (vm->q[id]) {
+ int i;
+
+ xe_exec_queue_last_fence_put(vm->q[id], vm);
+ for_each_tlb_inval(i)
+ xe_exec_queue_tlb_inval_last_fence_put(vm->q[id],
+ vm, i);
+ }
+ }
+ up_write(&vm->lock);
+
+ for_each_tile(tile, xe, id) {
+ if (vm->q[id]) {
+ xe_exec_queue_kill(vm->q[id]);
+ xe_exec_queue_put(vm->q[id]);
+ vm->q[id] = NULL;
+ }
+ }
+}
+
+/**
+ * xe_vm_add_close_queue() - Track an exec queue using the VM
+ * @vm: The VM used by the exec queue
+ *
+ * The queue is tracked until its asynchronous destruction completes.
+ */
+void xe_vm_add_close_queue(struct xe_vm *vm)
+{
+ atomic_inc(&vm->close.num_exec_queues);
+}
+
+/**
+ * xe_vm_remove_close_queue() - Stop tracking an exec queue
+ * @vm: The VM used by the exec queue
+ *
+ * If this is the last tracked queue and VM close was deferred, schedule
+ * the deferred VM close.
+ */
+void xe_vm_remove_close_queue(struct xe_vm *vm)
+{
+ bool queue_close = false;
+
+ if (!atomic_dec_and_test(&vm->close.num_exec_queues))
+ return;
+
+ wake_up_all(&vm->close.wq);
+
+ spin_lock(&vm->close.lock);
+ if (vm->close.deferred) {
+ vm->close.deferred = false;
+ queue_close = true;
+ }
+ spin_unlock(&vm->close.lock);
+
+ if (queue_close)
+ xe_destroy_wq_queue(&vm->close.work);
+}
+
+/**
+ * xe_vm_close_and_put_deferred() - Close a VM immediately or defer the close
+ * @vm: The VM reference to consume
+ * @timeout: Maximum time to wait for VM queues, in jiffies
+ *
+ * The VM must already be marked closing, and callers must serialize teardown.
+ * Release the VM-owned bind queues before waiting for queue cleanup. Callers
+ * may release them earlier to start cleanup for several VMs before waiting.
+ *
+ * If queues remain after @timeout, keep the mappings and finish VM close
+ * after the last tracked queue is freed. This consumes the caller's VM
+ * reference on both the immediate and deferred paths.
+ *
+ * Return: %true if VM close was deferred, or %false if it completed now.
+ */
+bool xe_vm_close_and_put_deferred(struct xe_vm *vm, unsigned long timeout)
+{
+ bool queue_close = false;
+
+ /* Drop VM-owned queue references before waiting for their final free. */
+ xe_vm_kill_bind_queues(vm);
+
+ if (!atomic_read(&vm->close.num_exec_queues) ||
+ wait_event_timeout(vm->close.wq,
+ !atomic_read(&vm->close.num_exec_queues),
+ timeout)) {
+ xe_vm_close_and_put(vm);
+ return false;
+ }
+
+ spin_lock(&vm->close.lock);
+ vm->close.deferred = true;
+ if (!atomic_read(&vm->close.num_exec_queues)) {
+ vm->close.deferred = false;
+ queue_close = true;
+ }
+ spin_unlock(&vm->close.lock);
+
+ if (queue_close)
+ xe_destroy_wq_queue(&vm->close.work);
+
+ return true;
+}
+
static void vm_destroy_work_func(struct work_struct *w)
{
struct xe_vm *vm =
@@ -2083,8 +2213,7 @@ static void xe_vm_free(struct drm_gpuvm *gpuvm)
*/
drm_dev_get(&vm->xe->drm);
- /* To destroy the VM we need to be able to sleep */
- queue_work(system_dfl_wq, &vm->destroy_work);
+ xe_destroy_wq_queue(&vm->destroy_work);
}
struct xe_vm *xe_vm_lookup(struct xe_file *xef, u32 id)
@@ -2230,7 +2359,12 @@ int xe_vm_destroy_ioctl(struct drm_device *dev, void *data,
if (!err) {
xe_vm_close_start(vm);
- xe_vm_close_and_put(vm);
+
+ /*
+ * User exec queues can outlive the VM handle. Do not wait for
+ * queues which this ioctl does not destroy.
+ */
+ xe_vm_close_and_put_deferred(vm, 0);
}
return err;
diff --git a/drivers/gpu/drm/xe/xe_vm.h b/drivers/gpu/drm/xe/xe_vm.h
index 70456ffe27e2..efa427651c47 100644
--- a/drivers/gpu/drm/xe/xe_vm.h
+++ b/drivers/gpu/drm/xe/xe_vm.h
@@ -221,6 +221,11 @@ int xe_vm_get_property_ioctl(struct drm_device *dev, void *data,
void xe_vm_close_and_put(struct xe_vm *vm);
void xe_vm_close_start(struct xe_vm *vm);
+bool xe_vm_close_and_put_deferred(struct xe_vm *vm, unsigned long timeout);
+void xe_vm_kill_bind_queues(struct xe_vm *vm);
+
+void xe_vm_add_close_queue(struct xe_vm *vm);
+void xe_vm_remove_close_queue(struct xe_vm *vm);
static inline bool xe_vm_in_fault_mode(struct xe_vm *vm)
{
diff --git a/drivers/gpu/drm/xe/xe_vm_types.h b/drivers/gpu/drm/xe/xe_vm_types.h
index 16b051a0cf10..a3802f4fd00b 100644
--- a/drivers/gpu/drm/xe/xe_vm_types.h
+++ b/drivers/gpu/drm/xe/xe_vm_types.h
@@ -14,6 +14,7 @@
#include <linux/kref.h>
#include <linux/mmu_notifier.h>
#include <linux/scatterlist.h>
+#include <linux/wait.h>
#include "xe_device_types.h"
#include "xe_pt_types.h"
@@ -314,6 +315,20 @@ struct xe_vm {
*/
struct work_struct destroy_work;
+ /** @close: State used to defer VM teardown until exec queues are gone. */
+ struct {
+ /** @close.num_exec_queues: Queues which can access this VM. */
+ atomic_t num_exec_queues;
+ /** @close.wq: Waitqueue for exec queue teardown. */
+ wait_queue_head_t wq;
+ /** @close.lock: Protects deferred and work scheduling. */
+ spinlock_t lock;
+ /** @close.deferred: VM close is waiting for queue teardown. */
+ bool deferred;
+ /** @close.work: Completes a deferred VM close. */
+ struct work_struct work;
+ } close;
+
/**
* @rftree: range fence tree to track updates to page table structure.
* Used to implement conflict tracking between independent bind engines.
--
2.43.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* [PATCH 5/5] drm/xe/guc: Reset LRC ring pointers before replay
2026-09-16 9:53 [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Arvind Yadav
` (3 preceding siblings ...)
2026-09-16 9:53 ` [PATCH 4/5] drm/xe: Defer VM teardown until exec queue cleanup completes Arvind Yadav
@ 2026-09-16 9:53 ` Arvind Yadav
2026-09-18 22:31 ` [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Matthew Brost
5 siblings, 0 replies; 14+ messages in thread
From: Arvind Yadav @ 2026-09-16 9:53 UTC (permalink / raw)
To: intel-xe, dri-devel
Cc: rodrigo.vivi, matthew.brost, himal.prasad.ghimiray,
thomas.hellstrom
A GT reset can stop a context after the LRC head has advanced past the
start of the oldest pending job.
The replay path rewinds the software ring tail to the job's recorded
start so the pending jobs are written again. However, it leaves the LRC
head at its saved later position and initializes the LRC tail from that
position.
The hardware and software ring positions can therefore use different
replay starting points.
Before resubmitting the queue, set the software tail and both LRC ring
pointers to the start of the oldest pending job. Update the pointers
while the context is unregistered, before resubmitting pending jobs.
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
---
drivers/gpu/drm/xe/xe_guc_submit.c | 18 +++++++++---------
1 file changed, 9 insertions(+), 9 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
index ca24a77dfb26..878ba94d22e0 100644
--- a/drivers/gpu/drm/xe/xe_guc_submit.c
+++ b/drivers/gpu/drm/xe/xe_guc_submit.c
@@ -2984,17 +2984,17 @@ static void guc_exec_queue_start(struct xe_exec_queue *q)
trace_xe_exec_queue_resubmit(q);
if (job) {
for (i = 0; i < q->width; ++i) {
+ u32 replay_head = job->ptrs[i].head;
+
/*
- * The GuC context is unregistered at this point
- * time, adjusting software ring tail ensures
- * jobs are rewritten in original placement,
- * adjusting LRC tail ensures the newly loaded
- * GuC / contexts only view the LRC tail
- * increasing as jobs are written out.
+ * A started job may have advanced the saved LRC
+ * head past its original ring position. Rewind
+ * both head and tail before rewriting and
+ * replaying the pending jobs.
*/
- q->lrc[i]->ring.tail = job->ptrs[i].head;
- xe_lrc_set_ring_tail(q->lrc[i],
- xe_lrc_ring_head(q->lrc[i]));
+ q->lrc[i]->ring.tail = replay_head;
+ xe_lrc_set_ring_head(q->lrc[i], replay_head);
+ xe_lrc_set_ring_tail(q->lrc[i], replay_head);
}
}
xe_sched_resubmit_jobs(sched);
--
2.43.0
^ permalink raw reply related [flat|nested] 14+ messages in thread
* Re: [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction
2026-09-16 9:53 ` [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction Arvind Yadav
@ 2026-09-18 3:22 ` Matthew Brost
2026-09-18 7:22 ` Thomas Hellström
0 siblings, 1 reply; 14+ messages in thread
From: Matthew Brost @ 2026-09-18 3:22 UTC (permalink / raw)
To: Arvind Yadav
Cc: intel-xe, dri-devel, rodrigo.vivi, himal.prasad.ghimiray,
thomas.hellstrom
On Wed, Sep 16, 2026 at 03:23:33PM +0530, Arvind Yadav wrote:
> xe_vm_free() is the drm_gpuvm vm_free callback. It hands the final
> teardown to vm_destroy_work_func() on a workqueue and returns.
>
> drm_gpuvm_free() drops its device reference immediately after the
> callback returns:
>
> gpuvm->ops->vm_free(gpuvm);
> drm_dev_put(drm);
>
> vm_destroy_work_func() then keeps using device state: xe_pm_runtime_put()
> for an LR mode VM, ttm_lru_bulk_move_fini() on xe->ttm, and the tile
> iteration. If the freed VM held the last device reference, the work runs
> against a released xe_device.
>
> Take a device reference in xe_vm_free() and drop it once
> vm_destroy_work_func() has finished using the device.
>
> Cc: Matthew Brost <matthew.brost@intel.com>
This is a fix, IMO. Ideally, we should probably push the delayed-destroy
semantics into gpuvm if they are really needed. I'm also questioning
whether the VM destroy worker is actually required. This dates back to
the very early days of Xe, and I doubt we've ever revisited whether it
is necessary.
Let's follow up with one of the following:
- Introduce async destroy in gpuvm and have it own the drm_dev_get/put.
- Drop delayed destroy entirely in Xe.
As a temporary fix that can be backported, this looks good to me, so
with a Fixes tag:
Reviewed-by: Matthew Brost <matthew.brost@intel.com>
> Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
> Assisted-by: Claude:claude-opus-4-8
> Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
> ---
> drivers/gpu/drm/xe/xe_vm.c | 9 +++++++++
> 1 file changed, 9 insertions(+)
>
> diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c
> index efa5ff6cc823..264bdab75de2 100644
> --- a/drivers/gpu/drm/xe/xe_vm.c
> +++ b/drivers/gpu/drm/xe/xe_vm.c
> @@ -2054,12 +2054,21 @@ static void vm_destroy_work_func(struct work_struct *w)
> xe_file_put(vm->xef);
>
> kfree(vm);
> +
> + drm_dev_put(&xe->drm);
> }
>
> static void xe_vm_free(struct drm_gpuvm *gpuvm)
> {
> struct xe_vm *vm = container_of(gpuvm, struct xe_vm, gpuvm);
>
> + /*
> + * drm_gpuvm drops its device reference as soon as this callback
> + * returns, but vm_destroy_work_func() still uses device state. Hold a
> + * reference across the deferred work.
> + */
> + drm_dev_get(&vm->xe->drm);
> +
> /* To destroy the VM we need to be able to sleep */
> queue_work(system_dfl_wq, &vm->destroy_work);
> }
> --
> 2.43.0
>
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction
2026-09-18 3:22 ` Matthew Brost
@ 2026-09-18 7:22 ` Thomas Hellström
2026-09-18 22:41 ` Matthew Brost
0 siblings, 1 reply; 14+ messages in thread
From: Thomas Hellström @ 2026-09-18 7:22 UTC (permalink / raw)
To: Matthew Brost, Arvind Yadav
Cc: intel-xe, dri-devel, rodrigo.vivi, himal.prasad.ghimiray
On Thu, 2026-09-17 at 20:22 -0700, Matthew Brost wrote:
> On Wed, Sep 16, 2026 at 03:23:33PM +0530, Arvind Yadav wrote:
> > xe_vm_free() is the drm_gpuvm vm_free callback. It hands the final
> > teardown to vm_destroy_work_func() on a workqueue and returns.
> >
> > drm_gpuvm_free() drops its device reference immediately after the
> > callback returns:
> >
> > gpuvm->ops->vm_free(gpuvm);
> > drm_dev_put(drm);
> >
> > vm_destroy_work_func() then keeps using device state:
> > xe_pm_runtime_put()
> > for an LR mode VM, ttm_lru_bulk_move_fini() on xe->ttm, and the
> > tile
> > iteration. If the freed VM held the last device reference, the work
> > runs
> > against a released xe_device.
> >
> > Take a device reference in xe_vm_free() and drop it once
> > vm_destroy_work_func() has finished using the device.
> >
> > Cc: Matthew Brost <matthew.brost@intel.com>
>
> This is a fix, IMO. Ideally, we should probably push the delayed-
> destroy
> semantics into gpuvm if they are really needed. I'm also questioning
> whether the VM destroy worker is actually required. This dates back
> to
> the very early days of Xe, and I doubt we've ever revisited whether
> it
> is necessary.
>
> Let's follow up with one of the following:
> - Introduce async destroy in gpuvm and have it own the
> drm_dev_get/put.
> - Drop delayed destroy entirely in Xe.
We need to keep in mind that the drm file keeps a reference on the Xe
module. So once the last close() callback has executed, the module can
typically be unloaded, causing execution UAF. It's therefore not really
recommended to keep file-related structures around with a refcount
after close.
Device references however typically don't necessarily keep the module
pinned. I had a series to fix this for xe only, (Got stalled) [1], but
in general we should be careful about leaking that assumption into DRM
code.
[1] https://patchwork.freedesktop.org/series/163298/
Thanks,
Thomas
>
> As a temporary fix that can be backported, this looks good to me, so
> with a Fixes tag:
>
> Reviewed-by: Matthew Brost <matthew.brost@intel.com>
>
> > Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> > Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> > Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
> > Assisted-by: Claude:claude-opus-4-8
> > Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
> > ---
> > drivers/gpu/drm/xe/xe_vm.c | 9 +++++++++
> > 1 file changed, 9 insertions(+)
> >
> > diff --git a/drivers/gpu/drm/xe/xe_vm.c
> > b/drivers/gpu/drm/xe/xe_vm.c
> > index efa5ff6cc823..264bdab75de2 100644
> > --- a/drivers/gpu/drm/xe/xe_vm.c
> > +++ b/drivers/gpu/drm/xe/xe_vm.c
> > @@ -2054,12 +2054,21 @@ static void vm_destroy_work_func(struct
> > work_struct *w)
> > xe_file_put(vm->xef);
> >
> > kfree(vm);
> > +
> > + drm_dev_put(&xe->drm);
> > }
> >
> > static void xe_vm_free(struct drm_gpuvm *gpuvm)
> > {
> > struct xe_vm *vm = container_of(gpuvm, struct xe_vm,
> > gpuvm);
> >
> > + /*
> > + * drm_gpuvm drops its device reference as soon as this
> > callback
> > + * returns, but vm_destroy_work_func() still uses device
> > state. Hold a
> > + * reference across the deferred work.
> > + */
> > + drm_dev_get(&vm->xe->drm);
> > +
> > /* To destroy the VM we need to be able to sleep */
> > queue_work(system_dfl_wq, &vm->destroy_work);
> > }
> > --
> > 2.43.0
> >
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery
2026-09-16 9:53 [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Arvind Yadav
` (4 preceding siblings ...)
2026-09-16 9:53 ` [PATCH 5/5] drm/xe/guc: Reset LRC ring pointers before replay Arvind Yadav
@ 2026-09-18 22:31 ` Matthew Brost
2026-09-24 10:06 ` Yadav, Arvind
5 siblings, 1 reply; 14+ messages in thread
From: Matthew Brost @ 2026-09-18 22:31 UTC (permalink / raw)
To: Arvind Yadav
Cc: intel-xe, dri-devel, rodrigo.vivi, himal.prasad.ghimiray,
thomas.hellstrom
On Wed, Sep 16, 2026 at 03:23:32PM +0530, Arvind Yadav wrote:
> On BMG, terminating a process with Ctrl+C while GPU work is running under
> VRAM pressure can cause an RCS page fault followed by a migration queue
> timeout on BCS8 (guc_id 0). The driver resets the GT, but the pending
> migration job can time out again and eventually wedge the device.
>
If we can just avoid faults on ctrl-c or segfault will this series be
needed (or at least only a subset of this series)?
I suggested a way to kill the exec queue first + sync wait on those here
[1] before clobbering the GPU page tables? I don't really care who works
on [1].
We still need root cause why a RCS fault affects the BCS engine - it
shouldn't and need investigation.
Matt
[1] https://patchwork.freedesktop.org/patch/752845/?series=173889&rev=1#comment_1391707
> BCS8 is reserved for paging and runs migration and VM bind work.
> The exact hardware link between the RCS fault and the BCS8 stall
> is still under investigation. This series addresses the teardown
> and recovery problems found while debugging this failure.
>
> During file close, exec queue cleanup starts asynchronously, but the VM
> mappings can be removed before that cleanup finishes. A missed wakeup in
> the GuC disable-completion handler can also turn a completed operation
> into a five-second timeout and an unnecessary GT reset.
>
> The reset replay path rewinds the software ring tail to the oldest pending
> job, but leaves the LRC head at its saved position. This leaves different
> starting positions for replay.
>
> The four patches address these paths:
> 1. Clear pending-disable state before waking waiters, so a completed disable
> does not appear to time out.
>
> 2. Mark VMs as closing before queue cleanup. Reject new work and page faults
> on closing VMs, while allowing existing SVM invalidation to drain mappings.
>
> 3. Keep VM mappings alive until queue cleanup completes. File close uses one
> five-second queue-wait budget across all VMs, then defers any remaining teardown.
> VM destroy defers without waiting. Device references protect deferred close and
> final VM destruction on the module-lifetime destroy workqueue.
>
> 4. Set the software tail and LRC head and tail to the oldest pending job before
> resubmitting jobs after a GT reset.
>
> The five-second budget applies only to the new queue-cleanup wait.
> Existing teardown waits are unchanged.
>
> Arvind Yadav (5):
> drm/xe: Hold a device reference across deferred VM destruction
> drm/xe/guc: Wake disable waiters after clearing pending state
> drm/xe: Mark VMs as closing before queue cleanup
> drm/xe: Defer VM teardown until exec queue cleanup completes
> drm/xe/guc: Reset LRC ring pointers before replay
>
> drivers/gpu/drm/xe/xe_device.c | 25 +++-
> drivers/gpu/drm/xe/xe_exec_queue.c | 22 +++-
> drivers/gpu/drm/xe/xe_guc_submit.c | 60 +++++----
> drivers/gpu/drm/xe/xe_module.c | 6 +-
> drivers/gpu/drm/xe/xe_pagefault.c | 2 +-
> drivers/gpu/drm/xe/xe_svm.c | 3 +-
> drivers/gpu/drm/xe/xe_vm.c | 205 +++++++++++++++++++++++++----
> drivers/gpu/drm/xe/xe_vm.h | 14 +-
> drivers/gpu/drm/xe/xe_vm_types.h | 16 +++
> 9 files changed, 289 insertions(+), 64 deletions(-)
>
> --
> 2.43.0
>
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH 2/5] drm/xe/guc: Wake disable waiters after clearing pending state
2026-09-16 9:53 ` [PATCH 2/5] drm/xe/guc: Wake disable waiters after clearing pending state Arvind Yadav
@ 2026-09-18 22:36 ` Matthew Brost
2026-09-21 6:46 ` Yadav, Arvind
0 siblings, 1 reply; 14+ messages in thread
From: Matthew Brost @ 2026-09-18 22:36 UTC (permalink / raw)
To: Arvind Yadav
Cc: intel-xe, dri-devel, rodrigo.vivi, himal.prasad.ghimiray,
thomas.hellstrom
On Wed, Sep 16, 2026 at 03:23:34PM +0530, Arvind Yadav wrote:
> disable_scheduling_deregister() waits for any pending scheduling
> operation to complete before destroying an exec queue.
>
> For a disable completion, handle_sched_done() can wake the waiter before
> clearing pending_disable, or not wake it at all. The waiter can then miss
> a completed operation and expire after five seconds. This causes a
> spurious GT reset and immediate TDR.
>
> Clear pending_disable before waking the waitqueue on every completion
> path. Sample the destroyed state before clearing the pending state, as
> required by the existing destroy protocol.
>
> Also wake CT waiters after a suspend completion.
> The queue remains alive until deregistration completes, so waking
> waiters before sending the deregister request does not release it.
>
> Cc: Matthew Brost <matthew.brost@intel.com>
> Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
I feel like this is an independent fix (like patch 1) that can be sent
on its own with a Fixes tag and CC'd to stable. This also looks very
similar to [1].
Can you coordinate with the author of [1]? Alternatively, if that patch
looks good to you, could you give it an RB?
Matt
[1] https://patchwork.freedesktop.org/patch/753791/?series=174199&rev=1
> Assisted-by: Claude:claude-opus-4-8
> Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
> ---
> drivers/gpu/drm/xe/xe_guc_submit.c | 42 +++++++++++++++++-------------
> 1 file changed, 24 insertions(+), 18 deletions(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
> index f3ba8abfc228..ca24a77dfb26 100644
> --- a/drivers/gpu/drm/xe/xe_guc_submit.c
> +++ b/drivers/gpu/drm/xe/xe_guc_submit.c
> @@ -3265,26 +3265,32 @@ static void handle_sched_done(struct xe_guc *guc, struct xe_exec_queue *q,
> if (q->guc->suspend_pending) {
> clear_exec_queue_pending_disable(q);
> suspend_fence_signal(q);
> +
> + /*
> + * Publish the cleared state before waking waiters.
> + */
> + smp_wmb();
> + wake_up_all(&guc->ct.wq);
> } else {
> - if (exec_queue_banned(q)) {
> - smp_wmb();
> - wake_up_all(&guc->ct.wq);
> - }
> - if (exec_queue_destroyed(q)) {
> - /*
> - * Make sure to clear the pending_disable only
> - * after sampling the destroyed state. We want
> - * to ensure we don't trigger the unregister too
> - * early with something intending to only
> - * disable scheduling. The caller doing the
> - * destroy must wait for an ongoing
> - * pending_disable before marking as destroyed.
> - */
> - clear_exec_queue_pending_disable(q);
> + bool destroyed = exec_queue_destroyed(q);
> +
> + /*
> + * Make sure to clear pending_disable only after sampling
> + * the destroyed state. The caller doing the destroy must
> + * wait for an ongoing disable before marking the queue
> + * destroyed.
> + */
> + clear_exec_queue_pending_disable(q);
> +
> + /*
> + * Publish the cleared state before waking waiters.
> + */
> + smp_wmb();
> + wake_up_all(&guc->ct.wq);
> +
> + /* The queue remains alive until DEREGISTER_DONE. */
> + if (destroyed)
> deregister_exec_queue(guc, q);
> - } else {
> - clear_exec_queue_pending_disable(q);
> - }
> }
> }
> }
> --
> 2.43.0
>
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction
2026-09-18 7:22 ` Thomas Hellström
@ 2026-09-18 22:41 ` Matthew Brost
2026-09-21 7:06 ` Thomas Hellström
0 siblings, 1 reply; 14+ messages in thread
From: Matthew Brost @ 2026-09-18 22:41 UTC (permalink / raw)
To: Thomas Hellström
Cc: Arvind Yadav, intel-xe, dri-devel, rodrigo.vivi,
himal.prasad.ghimiray
On Fri, Sep 18, 2026 at 09:22:53AM +0200, Thomas Hellström wrote:
> On Thu, 2026-09-17 at 20:22 -0700, Matthew Brost wrote:
> > On Wed, Sep 16, 2026 at 03:23:33PM +0530, Arvind Yadav wrote:
> > > xe_vm_free() is the drm_gpuvm vm_free callback. It hands the final
> > > teardown to vm_destroy_work_func() on a workqueue and returns.
> > >
> > > drm_gpuvm_free() drops its device reference immediately after the
> > > callback returns:
> > >
> > > gpuvm->ops->vm_free(gpuvm);
> > > drm_dev_put(drm);
> > >
> > > vm_destroy_work_func() then keeps using device state:
> > > xe_pm_runtime_put()
> > > for an LR mode VM, ttm_lru_bulk_move_fini() on xe->ttm, and the
> > > tile
> > > iteration. If the freed VM held the last device reference, the work
> > > runs
> > > against a released xe_device.
> > >
> > > Take a device reference in xe_vm_free() and drop it once
> > > vm_destroy_work_func() has finished using the device.
> > >
> > > Cc: Matthew Brost <matthew.brost@intel.com>
> >
> > This is a fix, IMO. Ideally, we should probably push the delayed-
> > destroy
> > semantics into gpuvm if they are really needed. I'm also questioning
> > whether the VM destroy worker is actually required. This dates back
> > to
> > the very early days of Xe, and I doubt we've ever revisited whether
> > it
> > is necessary.
> >
> > Let's follow up with one of the following:
> > - Introduce async destroy in gpuvm and have it own the
> > drm_dev_get/put.
> > - Drop delayed destroy entirely in Xe.
>
> We need to keep in mind that the drm file keeps a reference on the Xe
> module. So once the last close() callback has executed, the module can
Yikes, so drm_dev_get won't prevent the module from unloading? Different
issue, right?
> typically be unloaded, causing execution UAF. It's therefore not really
> recommended to keep file-related structures around with a refcount
> after close.
>
> Device references however typically don't necessarily keep the module
> pinned. I had a series to fix this for xe only, (Got stalled) [1], but
> in general we should be careful about leaking that assumption into DRM
> code.
>
So what is the fix here? I believe gpuvm has a drm_dev_ ref and it async
teardown which breaks drm_dev_ assumption. I don't see an issue with
this patch or my suggest follow up.
> [1] https://patchwork.freedesktop.org/series/163298/
Should we push on this? I don't really see an issue with series as long
as 'rmmod xe' works if display isn't holding a module ref.
Matt
>
> Thanks,
> Thomas
>
>
> >
> > As a temporary fix that can be backported, this looks good to me, so
> > with a Fixes tag:
> >
> > Reviewed-by: Matthew Brost <matthew.brost@intel.com>
> >
> > > Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> > > Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> > > Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
> > > Assisted-by: Claude:claude-opus-4-8
> > > Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
> > > ---
> > > drivers/gpu/drm/xe/xe_vm.c | 9 +++++++++
> > > 1 file changed, 9 insertions(+)
> > >
> > > diff --git a/drivers/gpu/drm/xe/xe_vm.c
> > > b/drivers/gpu/drm/xe/xe_vm.c
> > > index efa5ff6cc823..264bdab75de2 100644
> > > --- a/drivers/gpu/drm/xe/xe_vm.c
> > > +++ b/drivers/gpu/drm/xe/xe_vm.c
> > > @@ -2054,12 +2054,21 @@ static void vm_destroy_work_func(struct
> > > work_struct *w)
> > > xe_file_put(vm->xef);
> > >
> > > kfree(vm);
> > > +
> > > + drm_dev_put(&xe->drm);
> > > }
> > >
> > > static void xe_vm_free(struct drm_gpuvm *gpuvm)
> > > {
> > > struct xe_vm *vm = container_of(gpuvm, struct xe_vm,
> > > gpuvm);
> > >
> > > + /*
> > > + * drm_gpuvm drops its device reference as soon as this
> > > callback
> > > + * returns, but vm_destroy_work_func() still uses device
> > > state. Hold a
> > > + * reference across the deferred work.
> > > + */
> > > + drm_dev_get(&vm->xe->drm);
> > > +
> > > /* To destroy the VM we need to be able to sleep */
> > > queue_work(system_dfl_wq, &vm->destroy_work);
> > > }
> > > --
> > > 2.43.0
> > >
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH 2/5] drm/xe/guc: Wake disable waiters after clearing pending state
2026-09-18 22:36 ` Matthew Brost
@ 2026-09-21 6:46 ` Yadav, Arvind
0 siblings, 0 replies; 14+ messages in thread
From: Yadav, Arvind @ 2026-09-21 6:46 UTC (permalink / raw)
To: Matthew Brost
Cc: intel-xe, dri-devel, rodrigo.vivi, himal.prasad.ghimiray,
thomas.hellstrom
On 19-09-2026 04:06, Matthew Brost wrote:
> On Wed, Sep 16, 2026 at 03:23:34PM +0530, Arvind Yadav wrote:
>> disable_scheduling_deregister() waits for any pending scheduling
>> operation to complete before destroying an exec queue.
>>
>> For a disable completion, handle_sched_done() can wake the waiter before
>> clearing pending_disable, or not wake it at all. The waiter can then miss
>> a completed operation and expire after five seconds. This causes a
>> spurious GT reset and immediate TDR.
>>
>> Clear pending_disable before waking the waitqueue on every completion
>> path. Sample the destroyed state before clearing the pending state, as
>> required by the existing destroy protocol.
>>
>> Also wake CT waiters after a suspend completion.
>> The queue remains alive until deregistration completes, so waking
>> waiters before sending the deregister request does not release it.
>>
>> Cc: Matthew Brost <matthew.brost@intel.com>
>> Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
>> Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
>> Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
> I feel like this is an independent fix (like patch 1) that can be sent
> on its own with a Fixes tag and CC'd to stable. This also looks very
> similar to [1].
>
> Can you coordinate with the author of [1]? Alternatively, if that patch
> looks good to you, could you give it an RB?
Thank you for review.
Noted,
Arvind
>
> Matt
>
> [1] https://patchwork.freedesktop.org/patch/753791/?series=174199&rev=1
>
>> Assisted-by: Claude:claude-opus-4-8
>> Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
>> ---
>> drivers/gpu/drm/xe/xe_guc_submit.c | 42 +++++++++++++++++-------------
>> 1 file changed, 24 insertions(+), 18 deletions(-)
>>
>> diff --git a/drivers/gpu/drm/xe/xe_guc_submit.c b/drivers/gpu/drm/xe/xe_guc_submit.c
>> index f3ba8abfc228..ca24a77dfb26 100644
>> --- a/drivers/gpu/drm/xe/xe_guc_submit.c
>> +++ b/drivers/gpu/drm/xe/xe_guc_submit.c
>> @@ -3265,26 +3265,32 @@ static void handle_sched_done(struct xe_guc *guc, struct xe_exec_queue *q,
>> if (q->guc->suspend_pending) {
>> clear_exec_queue_pending_disable(q);
>> suspend_fence_signal(q);
>> +
>> + /*
>> + * Publish the cleared state before waking waiters.
>> + */
>> + smp_wmb();
>> + wake_up_all(&guc->ct.wq);
>> } else {
>> - if (exec_queue_banned(q)) {
>> - smp_wmb();
>> - wake_up_all(&guc->ct.wq);
>> - }
>> - if (exec_queue_destroyed(q)) {
>> - /*
>> - * Make sure to clear the pending_disable only
>> - * after sampling the destroyed state. We want
>> - * to ensure we don't trigger the unregister too
>> - * early with something intending to only
>> - * disable scheduling. The caller doing the
>> - * destroy must wait for an ongoing
>> - * pending_disable before marking as destroyed.
>> - */
>> - clear_exec_queue_pending_disable(q);
>> + bool destroyed = exec_queue_destroyed(q);
>> +
>> + /*
>> + * Make sure to clear pending_disable only after sampling
>> + * the destroyed state. The caller doing the destroy must
>> + * wait for an ongoing disable before marking the queue
>> + * destroyed.
>> + */
>> + clear_exec_queue_pending_disable(q);
>> +
>> + /*
>> + * Publish the cleared state before waking waiters.
>> + */
>> + smp_wmb();
>> + wake_up_all(&guc->ct.wq);
>> +
>> + /* The queue remains alive until DEREGISTER_DONE. */
>> + if (destroyed)
>> deregister_exec_queue(guc, q);
>> - } else {
>> - clear_exec_queue_pending_disable(q);
>> - }
>> }
>> }
>> }
>> --
>> 2.43.0
>>
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction
2026-09-18 22:41 ` Matthew Brost
@ 2026-09-21 7:06 ` Thomas Hellström
0 siblings, 0 replies; 14+ messages in thread
From: Thomas Hellström @ 2026-09-21 7:06 UTC (permalink / raw)
To: Matthew Brost
Cc: Arvind Yadav, intel-xe, dri-devel, rodrigo.vivi,
himal.prasad.ghimiray
On Fri, 2026-09-18 at 15:41 -0700, Matthew Brost wrote:
> On Fri, Sep 18, 2026 at 09:22:53AM +0200, Thomas Hellström wrote:
> > On Thu, 2026-09-17 at 20:22 -0700, Matthew Brost wrote:
> > > On Wed, Sep 16, 2026 at 03:23:33PM +0530, Arvind Yadav wrote:
> > > > xe_vm_free() is the drm_gpuvm vm_free callback. It hands the
> > > > final
> > > > teardown to vm_destroy_work_func() on a workqueue and returns.
> > > >
> > > > drm_gpuvm_free() drops its device reference immediately after
> > > > the
> > > > callback returns:
> > > >
> > > > gpuvm->ops->vm_free(gpuvm);
> > > > drm_dev_put(drm);
> > > >
> > > > vm_destroy_work_func() then keeps using device state:
> > > > xe_pm_runtime_put()
> > > > for an LR mode VM, ttm_lru_bulk_move_fini() on xe->ttm, and the
> > > > tile
> > > > iteration. If the freed VM held the last device reference, the
> > > > work
> > > > runs
> > > > against a released xe_device.
> > > >
> > > > Take a device reference in xe_vm_free() and drop it once
> > > > vm_destroy_work_func() has finished using the device.
> > > >
> > > > Cc: Matthew Brost <matthew.brost@intel.com>
> > >
> > > This is a fix, IMO. Ideally, we should probably push the delayed-
> > > destroy
> > > semantics into gpuvm if they are really needed. I'm also
> > > questioning
> > > whether the VM destroy worker is actually required. This dates
> > > back
> > > to
> > > the very early days of Xe, and I doubt we've ever revisited
> > > whether
> > > it
> > > is necessary.
> > >
> > > Let's follow up with one of the following:
> > > - Introduce async destroy in gpuvm and have it own the
> > > drm_dev_get/put.
> > > - Drop delayed destroy entirely in Xe.
> >
> > We need to keep in mind that the drm file keeps a reference on the
> > Xe
> > module. So once the last close() callback has executed, the module
> > can
>
> Yikes, so drm_dev_get won't prevent the module from unloading?
Exactly.
> Different
> issue, right?
Yes, but if we move into GPUVM it would affect other drivers. I believe
the issue is fixable on a per-driver basis.
>
> > typically be unloaded, causing execution UAF. It's therefore not
> > really
> > recommended to keep file-related structures around with a refcount
> > after close.
> >
> > Device references however typically don't necessarily keep the
> > module
> > pinned. I had a series to fix this for xe only, (Got stalled) [1],
> > but
> > in general we should be careful about leaking that assumption into
> > DRM
> > code.
> >
>
> So what is the fix here? I believe gpuvm has a drm_dev_ ref and it
> async
> teardown which breaks drm_dev_ assumption. I don't see an issue with
> this patch or my suggest follow up.
While the fix for xe is the below series, And I will respin that, If we
move more device refcounts into GPUVM, that will affect other drivers
that aren't yet fixed. Note drm_pagemap handles this with the dev
unhold work, not sure if that is applicable also with GPUVM.
Either way, I don't want to block the series due to a fix missing. Just
want us to keep this in mind since it increases the probability of the
real bug hitting.
>
> > [1] https://patchwork.freedesktop.org/series/163298/
>
> Should we push on this? I don't really see an issue with series as
> long
> as 'rmmod xe' works if display isn't holding a module ref.
Yes, I'll respin.
Thanks,
Thomas
>
> Matt
>
> >
> > Thanks,
> > Thomas
> >
> >
> > >
> > > As a temporary fix that can be backported, this looks good to me,
> > > so
> > > with a Fixes tag:
> > >
> > > Reviewed-by: Matthew Brost <matthew.brost@intel.com>
> > >
> > > > Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
> > > > Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
> > > > Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
> > > > Assisted-by: Claude:claude-opus-4-8
> > > > Signed-off-by: Arvind Yadav <arvind.yadav@intel.com>
> > > > ---
> > > > drivers/gpu/drm/xe/xe_vm.c | 9 +++++++++
> > > > 1 file changed, 9 insertions(+)
> > > >
> > > > diff --git a/drivers/gpu/drm/xe/xe_vm.c
> > > > b/drivers/gpu/drm/xe/xe_vm.c
> > > > index efa5ff6cc823..264bdab75de2 100644
> > > > --- a/drivers/gpu/drm/xe/xe_vm.c
> > > > +++ b/drivers/gpu/drm/xe/xe_vm.c
> > > > @@ -2054,12 +2054,21 @@ static void vm_destroy_work_func(struct
> > > > work_struct *w)
> > > > xe_file_put(vm->xef);
> > > >
> > > > kfree(vm);
> > > > +
> > > > + drm_dev_put(&xe->drm);
> > > > }
> > > >
> > > > static void xe_vm_free(struct drm_gpuvm *gpuvm)
> > > > {
> > > > struct xe_vm *vm = container_of(gpuvm, struct xe_vm,
> > > > gpuvm);
> > > >
> > > > + /*
> > > > + * drm_gpuvm drops its device reference as soon as
> > > > this
> > > > callback
> > > > + * returns, but vm_destroy_work_func() still uses
> > > > device
> > > > state. Hold a
> > > > + * reference across the deferred work.
> > > > + */
> > > > + drm_dev_get(&vm->xe->drm);
> > > > +
> > > > /* To destroy the VM we need to be able to sleep */
> > > > queue_work(system_dfl_wq, &vm->destroy_work);
> > > > }
> > > > --
> > > > 2.43.0
> > > >
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery
2026-09-18 22:31 ` [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Matthew Brost
@ 2026-09-24 10:06 ` Yadav, Arvind
0 siblings, 0 replies; 14+ messages in thread
From: Yadav, Arvind @ 2026-09-24 10:06 UTC (permalink / raw)
To: Matthew Brost
Cc: intel-xe, dri-devel, rodrigo.vivi, himal.prasad.ghimiray,
thomas.hellstrom
On 19-09-2026 04:01, Matthew Brost wrote:
> On Wed, Sep 16, 2026 at 03:23:32PM +0530, Arvind Yadav wrote:
>> On BMG, terminating a process with Ctrl+C while GPU work is running under
>> VRAM pressure can cause an RCS page fault followed by a migration queue
>> timeout on BCS8 (guc_id 0). The driver resets the GT, but the pending
>> migration job can time out again and eventually wedge the device.
>>
> If we can just avoid faults on ctrl-c or segfault will this series be
> needed (or at least only a subset of this series)?
>
> I suggested a way to kill the exec queue first + sync wait on those here
> [1] before clobbering the GPU page tables? I don't really care who works
> on [1].
Thankyou Matt for you review comments and suggestions.
My series waits for final queue cleanup before removing the VM mappings.
Your suggestion is to disable all queues first, then wait for them to
stop before calling xe_vm_close().
I will build on your v6 patch and add the synchronous wait from the
earlier discussion. I will test that approach and check which parts of
my series are still needed.
>
> We still need root cause why a RCS fault affects the BCS engine - it
> shouldn't and need investigation.
I will also continue the GuC investigation into the first BCS8 failure.
Avoiding the teardown fault may prevent this reproducer, but it does not
explain why the RCS fault affects BCS8.
I am discussing this issue with the GuC team, and once we identify the
root cause, I will provide an update.
Thanks,
Arvind
>
> Matt
>
> [1] https://patchwork.freedesktop.org/patch/752845/?series=173889&rev=1#comment_1391707
>
>> BCS8 is reserved for paging and runs migration and VM bind work.
>> The exact hardware link between the RCS fault and the BCS8 stall
>> is still under investigation. This series addresses the teardown
>> and recovery problems found while debugging this failure.
>>
>> During file close, exec queue cleanup starts asynchronously, but the VM
>> mappings can be removed before that cleanup finishes. A missed wakeup in
>> the GuC disable-completion handler can also turn a completed operation
>> into a five-second timeout and an unnecessary GT reset.
>>
>> The reset replay path rewinds the software ring tail to the oldest pending
>> job, but leaves the LRC head at its saved position. This leaves different
>> starting positions for replay.
>>
>> The four patches address these paths:
>> 1. Clear pending-disable state before waking waiters, so a completed disable
>> does not appear to time out.
>>
>> 2. Mark VMs as closing before queue cleanup. Reject new work and page faults
>> on closing VMs, while allowing existing SVM invalidation to drain mappings.
>>
>> 3. Keep VM mappings alive until queue cleanup completes. File close uses one
>> five-second queue-wait budget across all VMs, then defers any remaining teardown.
>> VM destroy defers without waiting. Device references protect deferred close and
>> final VM destruction on the module-lifetime destroy workqueue.
>>
>> 4. Set the software tail and LRC head and tail to the oldest pending job before
>> resubmitting jobs after a GT reset.
>>
>> The five-second budget applies only to the new queue-cleanup wait.
>> Existing teardown waits are unchanged.
>>
>> Arvind Yadav (5):
>> drm/xe: Hold a device reference across deferred VM destruction
>> drm/xe/guc: Wake disable waiters after clearing pending state
>> drm/xe: Mark VMs as closing before queue cleanup
>> drm/xe: Defer VM teardown until exec queue cleanup completes
>> drm/xe/guc: Reset LRC ring pointers before replay
>>
>> drivers/gpu/drm/xe/xe_device.c | 25 +++-
>> drivers/gpu/drm/xe/xe_exec_queue.c | 22 +++-
>> drivers/gpu/drm/xe/xe_guc_submit.c | 60 +++++----
>> drivers/gpu/drm/xe/xe_module.c | 6 +-
>> drivers/gpu/drm/xe/xe_pagefault.c | 2 +-
>> drivers/gpu/drm/xe/xe_svm.c | 3 +-
>> drivers/gpu/drm/xe/xe_vm.c | 205 +++++++++++++++++++++++++----
>> drivers/gpu/drm/xe/xe_vm.h | 14 +-
>> drivers/gpu/drm/xe/xe_vm_types.h | 16 +++
>> 9 files changed, 289 insertions(+), 64 deletions(-)
>>
>> --
>> 2.43.0
>>
^ permalink raw reply [flat|nested] 14+ messages in thread
end of thread, other threads:[~2026-09-24 10:06 UTC | newest]
Thread overview: 14+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-16 9:53 [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Arvind Yadav
2026-09-16 9:53 ` [PATCH 1/5] drm/xe: Hold a device reference across deferred VM destruction Arvind Yadav
2026-09-18 3:22 ` Matthew Brost
2026-09-18 7:22 ` Thomas Hellström
2026-09-18 22:41 ` Matthew Brost
2026-09-21 7:06 ` Thomas Hellström
2026-09-16 9:53 ` [PATCH 2/5] drm/xe/guc: Wake disable waiters after clearing pending state Arvind Yadav
2026-09-18 22:36 ` Matthew Brost
2026-09-21 6:46 ` Yadav, Arvind
2026-09-16 9:53 ` [PATCH 3/5] drm/xe: Mark VMs as closing before queue cleanup Arvind Yadav
2026-09-16 9:53 ` [PATCH 4/5] drm/xe: Defer VM teardown until exec queue cleanup completes Arvind Yadav
2026-09-16 9:53 ` [PATCH 5/5] drm/xe/guc: Reset LRC ring pointers before replay Arvind Yadav
2026-09-18 22:31 ` [PATCH 0/5] drm/xe: Fix VM teardown and migration queue recovery Matthew Brost
2026-09-24 10:06 ` Yadav, Arvind
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox