All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH] drm/xe: Don't wedge shared engine on stale faults from torn-down VMs
@ 2026-09-11 19:23 Sanjay Yadav
  2026-09-11 19:39 ` sashiko-bot
                   ` (4 more replies)
  0 siblings, 5 replies; 6+ messages in thread
From: Sanjay Yadav @ 2026-09-11 19:23 UTC (permalink / raw)
  To: intel-xe; +Cc: matthew.brost, Stuart Summers

A recoverable page fault generated by a process's completed migrate/BCS
work can arrive after that process's VM has been closed and its ASID
removed (abnormal exit while GPU work was in flight).
xe_pagefault_service() returned an error for such faults, which is
reported to the GuC/HW as an unsuccessful response. Repeated
unsuccessful responses eventually escalate to an engine memory CAT
error. Because the migrate/BCS engine is shared, that CAT error
wedges it for unrelated processes, which then hang on their own
copies.

Detect stale faults whose owning VM is closed or whose ASID no longer
maps to a fault-capable VM, and drain them instead of replying
unsuccessful.

Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Stuart Summers <stuart.summers@intel.com>
Assisted-by: GitHub Copilot:claude-opus-4.8
Signed-off-by: Sanjay Yadav <sanjay.kumar.yadav@intel.com>
---
 drivers/gpu/drm/xe/xe_pagefault.c | 22 +++++++++++++++++++---
 1 file changed, 19 insertions(+), 3 deletions(-)

diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
index c82b8bc8bc70..2827652ab8a3 100644
--- a/drivers/gpu/drm/xe/xe_pagefault.c
+++ b/drivers/gpu/drm/xe/xe_pagefault.c
@@ -238,8 +238,10 @@ static struct xe_vm *xe_pagefault_asid_to_vm(struct xe_device *xe, u32 asid)
 	vm = xa_load(&xe->usm.asid_to_vm, asid);
 	if (vm && xe_vm_in_fault_mode(vm))
 		xe_vm_get(vm);
-	else
+	else if (vm)
 		vm = ERR_PTR(-EINVAL);
+	else
+		vm = ERR_PTR(-ENOENT);
 	up_read(&xe->usm.lock);
 
 	return vm;
@@ -260,13 +262,27 @@ static int xe_pagefault_service(struct xe_pagefault *pf)
 		return -EFAULT;
 
 	vm = xe_pagefault_asid_to_vm(xe, asid);
-	if (IS_ERR(vm))
+	if (IS_ERR(vm)) {
+		if (PTR_ERR(vm) == -ENOENT) {
+			drm_info(&xe->drm,
+				 "xe_pf_debug: drain stale fault (no VM) asid=%u addr=0x%llx\n",
+				 asid, pf->consumer.page_addr);
+			xe_pagefault_set_start_addr(pf, pf->consumer.page_addr);
+			xe_pagefault_set_end_addr(pf, pf->consumer.page_addr);
+			return 0;
+		}
 		return PTR_ERR(vm);
+	}
 
 	down_read(&vm->lock);
 
 	if (xe_vm_is_closed(vm)) {
-		err = -ENOENT;
+		drm_info(&xe->drm,
+			 "xe_pf_debug: drain stale fault (closed VM) asid=%u addr=0x%llx\n",
+			 asid, pf->consumer.page_addr);
+		xe_pagefault_set_start_addr(pf, pf->consumer.page_addr);
+		xe_pagefault_set_end_addr(pf, pf->consumer.page_addr);
+		err = 0;
 		goto unlock_vm;
 	}
 
-- 
2.52.0


^ permalink raw reply related	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-09-12  5:42 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-11 19:23 [PATCH] drm/xe: Don't wedge shared engine on stale faults from torn-down VMs Sanjay Yadav
2026-09-11 19:39 ` sashiko-bot
2026-09-11 19:51 ` ✓ CI.KUnit: success for " Patchwork
2026-09-11 20:43 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-11 20:45 ` [PATCH] " Matthew Brost
2026-09-12  5:42 ` ✓ Xe.CI.FULL: success for " Patchwork

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.