Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v2 00/15] drm/xe: Isolate wedged devices from hardware access
@ 2026-09-22 10:16 Arvind Yadav
  2026-09-22 10:16 ` [PATCH v2 01/15] drm/xe/irq: Always free requested IRQs on uninstall Arvind Yadav
                   ` (17 more replies)
  0 siblings, 18 replies; 23+ messages in thread
From: Arvind Yadav @ 2026-09-22 10:16 UTC (permalink / raw)
  To: intel-xe, dri-devel
  Cc: matthew.brost, himal.prasad.ghimiray, thomas.hellstrom,
	rodrigo.vivi

A permanently wedged device may still be accessed by pending work,
background workers and existing VRAM CPU mappings. This can leave DMA,
interrupt or MMIO activity running after userspace has been asked to
recover the device.

Add a common device I/O gate based on the DRM SRCU domain. Hardware
users enter the gate before checking the device state and leave it after
their hardware access has finished. PCI error recovery and permanent
wedge handling first publish the blocked state, signal pending GT fences
and then drain existing readers. This closes the race between checking
the device state and accessing hardware.

Keep recoverable PCI error handling separate from permanent wedging.
AER uses a temporary reset state and returns -ECANCELED for cancelled
GuC requests. A permanent wedge remains terminal and returns
-ENOTRECOVERABLE.

AER recovery recreates the Xe device. Existing DRM file descriptors
remain attached to the unplugged instance and are not recovered.
Userspace must close them and reopen the recovered device. This series
does not add transparent reset recovery.

After a permanent wedge, the driver:
 - Drops queued page faults.
 - Blocks VM, SVM and GuC hardware access.
 - Waits for active device I/O users.
 - Shuts down and unregisters the display.
 - Suspends interrupts and clears PCI bus mastering.
 - invalidates existing VRAM CPU mappings
 - maps later CPU faults to one dummy page per BO
 - Rejects new VRAM allocations.
 - Notifies userspace after isolation completes.

System suspend waits for wedge isolation to finish. PCI disable and
D3Cold operations remain paired even when the device wedges during
suspend. Resume balances the PCI state but skips driver resume for a
permanently wedged device.

The series also fixes IRQ uninstall so requested IRQ handlers are freed
after interrupts have already been suspended.

v2:
 - Add a common device I/O gate on the DRM SRCU domain. Hardware users
   enter the gate before checking device state, and blocked paths drain
   existing readers.
 - Return -ECANCELED for GuC sends blocked by AER recovery, keeping
  -ENOTRECOVERABLE for permanent wedges.
 - Update reported_method only when drm_dev_wedged_event() succeeds.
   Requeue only when the recovery method changes to avoid repeated
   retries after notification failure. (Andi)
 - Map only the faulting address to the per-BO dummy page instead of
   prefaulting the entire VMA.
 - Keep PCI enable/disable and D3Cold operations paired, including when
   the device wedges during suspend. Skip only driver-level PM work
   for wedged devices.
 - Use explicit runtime-PM get/put calls in xe_bo_vm_access() instead
   of mixing scope-based cleanup with goto-based error handling.
 - Protect VM and SVM hardware access with the common SRCU gate.
 - Hold the gate across page-fault, migration and rebind work.
 - Remove the racy device state checks.
 - Simplify the TLB invalidation warning check.
 - report the wedged event from a worker after isolation.
 - invalidate existing VRAM CPU mappings after draining active faults.

Arvind Yadav (14):
  drm/xe/irq: Always free requested IRQs on uninstall
  drm/xe: Separate AER reset state from device wedging
  drm/xe: Protect device I/O with DRM device SRCU
  drm/xe: Drop queued page faults when device I/O is blocked
  drm/xe: Stop VM work when device I/O is blocked
  drm/xe: Send wedged notification from a worker
  drm/xe: Reuse one dummy page per BO after wedge
  drm/xe: Invalidate existing VRAM mappings on wedge
  drm/xe/irq: Protect IRQ state during wedge isolation
  drm/xe: Isolate a wedged device before notifying userspace
  drm/xe/ttm: Reject VRAM allocations on wedged devices
  drm/xe/guc: Skip timeout recovery on a wedged device
  drm/xe: Skip PM notifier preparation when device I/O is blocked
  drm/xe: Block BO VM access when device I/O is unavailable

Thomas Hellström (1):
  drm/drv: Export drm_dev_srcu_synchronize()

 drivers/gpu/drm/drm_drv.c               |  17 ++-
 drivers/gpu/drm/xe/display/xe_display.c |  17 ++-
 drivers/gpu/drm/xe/xe_bo.c              |  72 ++++++++++-
 drivers/gpu/drm/xe/xe_bo.h              |   1 +
 drivers/gpu/drm/xe/xe_bo_types.h        |   4 +
 drivers/gpu/drm/xe/xe_device.c          | 155 ++++++++++++++++++++----
 drivers/gpu/drm/xe/xe_device.h          |  41 +++++++
 drivers/gpu/drm/xe/xe_device_types.h    |  21 ++++
 drivers/gpu/drm/xe/xe_guc_ct.c          |  24 +++-
 drivers/gpu/drm/xe/xe_guc_pc.c          |  10 +-
 drivers/gpu/drm/xe/xe_guc_rc.c          |   4 +-
 drivers/gpu/drm/xe/xe_guc_submit.c      |  13 +-
 drivers/gpu/drm/xe/xe_guc_tlb_inval.c   |  15 ++-
 drivers/gpu/drm/xe/xe_irq.c             |  29 +++--
 drivers/gpu/drm/xe/xe_pagefault.c       |  20 ++-
 drivers/gpu/drm/xe/xe_pci.c             |  51 ++++++--
 drivers/gpu/drm/xe/xe_pci_error.c       |  24 ++--
 drivers/gpu/drm/xe/xe_pm.c              |  12 ++
 drivers/gpu/drm/xe/xe_sriov_pf.c        |   2 +-
 drivers/gpu/drm/xe/xe_svm.c             |  29 ++++-
 drivers/gpu/drm/xe/xe_ttm_vram_mgr.c    |   4 +
 drivers/gpu/drm/xe/xe_vm.c              |  11 +-
 include/drm/drm_drv.h                   |   1 +
 23 files changed, 487 insertions(+), 90 deletions(-)

-- 
2.43.0


^ permalink raw reply	[flat|nested] 23+ messages in thread

end of thread, other threads:[~2026-09-22 20:50 UTC | newest]

Thread overview: 23+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-22 10:16 [PATCH v2 00/15] drm/xe: Isolate wedged devices from hardware access Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 01/15] drm/xe/irq: Always free requested IRQs on uninstall Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 02/15] drm/drv: Export drm_dev_srcu_synchronize() Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 03/15] drm/xe: Separate AER reset state from device wedging Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 04/15] drm/xe: Protect device I/O with DRM device SRCU Arvind Yadav
2026-09-22 10:28   ` sashiko-bot
2026-09-22 10:16 ` [PATCH v2 05/15] drm/xe: Drop queued page faults when device I/O is blocked Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 06/15] drm/xe: Stop VM work " Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 07/15] drm/xe: Send wedged notification from a worker Arvind Yadav
2026-09-22 10:27   ` sashiko-bot
2026-09-22 10:16 ` [PATCH v2 08/15] drm/xe: Reuse one dummy page per BO after wedge Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 09/15] drm/xe: Invalidate existing VRAM mappings on wedge Arvind Yadav
2026-09-22 10:30   ` sashiko-bot
2026-09-22 10:16 ` [PATCH v2 10/15] drm/xe/irq: Protect IRQ state during wedge isolation Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 11/15] drm/xe: Isolate a wedged device before notifying userspace Arvind Yadav
2026-09-22 10:31   ` sashiko-bot
2026-09-22 10:16 ` [PATCH v2 12/15] drm/xe/ttm: Reject VRAM allocations on wedged devices Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 13/15] drm/xe/guc: Skip timeout recovery on a wedged device Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 14/15] drm/xe: Skip PM notifier preparation when device I/O is blocked Arvind Yadav
2026-09-22 10:17 ` [PATCH v2 15/15] drm/xe: Block BO VM access when device I/O is unavailable Arvind Yadav
2026-09-22 10:27 ` ✓ CI.KUnit: success for drm/xe: Isolate wedged devices from hardware access (rev2) Patchwork
2026-09-22 12:26 ` ✗ Xe.CI.BAT: failure " Patchwork
2026-09-22 20:50 ` ✗ Xe.CI.FULL: " Patchwork

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox