dri-devel Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v12 0/4] Introduce cold reset recovery method
@ 2026-07-24 10:03 Mallesh Koujalagi
  2026-07-24 10:03 ` [PATCH v12 1/4] drm: Add DRM_WEDGE_RECOVERY_COLD_RESET " Mallesh Koujalagi
                   ` (3 more replies)
  0 siblings, 4 replies; 11+ messages in thread
From: Mallesh Koujalagi @ 2026-07-24 10:03 UTC (permalink / raw)
  To: intel-xe, dri-devel, rodrigo.vivi
  Cc: andrealmeid, christian.koenig, airlied, simona.vetter, mripard,
	maarten.lankhorst, tzimmermann, anshuman.gupta, badal.nilawar,
	riana.tauro, karthik.poosa, sk.anirban, raag.jadav,
	Mallesh Koujalagi

Add support for handling errors that require a complete
device power cycle (cold reset) to recover.

Certain error conditions leave the device in a persistent hardware
error state that cannot be cleared through existing recovery mechanisms
such as driver reload or PCIe reset. In these cases, functionality can
only be restored by performing a cold reset.

To support this, the series introduces a new DRM wedging recovery
method, DRM_WEDGE_RECOVERY_COLD_RESET (BIT(4)). When a device is wedged
with this method, the DRM core notifies userspace via a uevent that a cold
reset is required. This allows userspace to take appropriate action to
power-cycle the device.

Example uevent received:
  SUBSYSTEM=drm
  WEDGED=cold-reset
  DEVPATH=/devices/.../drm/card0

v2:
- Add use case: Handling errors from power management unit,
  which requires a complete power cycle to
  recover. (Christian)
- Add several instead of number to avoid update. (Jani)

v3:
- Update any scenario that requires cold-reset. (Riana)
- Update document with generic scenario. (Riana)
- Consistent with terminology. (Raag)
- Remove already covered information.
- Use PUNIT instead of PMU. (Riana)
- Use consistent wordingi.
- Remove log. (Raag)

v4:
- Rename cold reset to power cyclce. (Raag)
- Update doc. (Raag/Riana)
- Change commit message. (Raag)
- Make function static. (Raag)

v5:
- Make it consistent with consumer expectations. (Raag)
- Update commit message.
- Remove unbind.
- Simplify cold-reset script.
- Remove kdoc for static function.
- Remove xe_ prefix for static function.

v6:
- Drop "last resort" wording. (Riana)
- Look up the hotplug slot in DEVPATH instead of scanning
  every PCI slot on the system. (Raag)
- Drop arbitrary sleep values from the example script.
- Expand commit message to explain why SUR_DN is masked. (Raag/Riana)
- Check Slot Implemented bit before reading Slot Capabilities, per
  PCIe spec. (Riana)
- Add debug log.

v7:
- Update recovery script. (Raag)
- Handle surprise link down event properly. (Aravind/Riana)
- Update commit message. (Riana)
- Correct log message.

v8:
- Add rescan instead of reset. (Raag)
- Use find_usp_dev() in punit_error_handler() function.

v9:
- Remove unwanted header. (Sashiko)
- Removed #ifdef CONFIG_PCIEAER. (Riana)
- Used pci_find_ext_capability() instead of usp->aer_cap.
- Clear the PCI_ERR_UNC_SURPDN status bit (W1C) after
  reset complete. (Lukas Wunner)
- Use pci_clear_and_set_config_dword() helper.

v10:
- Rebase.
- Fix column width. (Sashiko)

v11:
- Make udev rules in single line. (Sashiko)

v12:
- Trigger punit handler using fault-inject.

Cc: André Almeida <andrealmeid@igalia.com>
Cc: Christian König <christian.koenig@amd.com>
Cc: David Airlie <airlied@gmail.com>
Cc: Simona Vetter <simona.vetter@ffwll.ch>
Cc: Maxime Ripard <mripard@kernel.org>
Cc: Maarten Lankhorst <maarten.lankhorst@linux.intel.com>
Cc: Thomas Zimmermann <tzimmermann@suse.de>

Mallesh Koujalagi (4):
  drm: Add DRM_WEDGE_RECOVERY_COLD_RESET recovery method
  drm/doc: Document DRM_WEDGE_RECOVERY_COLD_RESET recovery method
  drm/xe: Handle PUNIT errors by requesting cold-reset recovery
  drm/xe/ras: Use fault-inject to trigger punit error handler

 Documentation/gpu/drm-uapi.rst  | 93 +++++++++++++++++++++++++++++++--
 drivers/gpu/drm/drm_drv.c       |  2 +
 drivers/gpu/drm/xe/xe_debugfs.c |  4 ++
 drivers/gpu/drm/xe/xe_debugfs.h |  2 +
 drivers/gpu/drm/xe/xe_ras.c     | 25 ++++++++-
 include/drm/drm_device.h        |  1 +
 6 files changed, 121 insertions(+), 6 deletions(-)

-- 
2.48.1


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH v12 1/4] drm: Add DRM_WEDGE_RECOVERY_COLD_RESET recovery method
  2026-07-24 10:03 [PATCH v12 0/4] Introduce cold reset recovery method Mallesh Koujalagi
@ 2026-07-24 10:03 ` Mallesh Koujalagi
  2026-07-24 10:16   ` sashiko-bot
  2026-07-24 10:03 ` [PATCH v12 2/4] drm/doc: Document " Mallesh Koujalagi
                   ` (2 subsequent siblings)
  3 siblings, 1 reply; 11+ messages in thread
From: Mallesh Koujalagi @ 2026-07-24 10:03 UTC (permalink / raw)
  To: intel-xe, dri-devel, rodrigo.vivi
  Cc: andrealmeid, christian.koenig, airlied, simona.vetter, mripard,
	maarten.lankhorst, tzimmermann, anshuman.gupta, badal.nilawar,
	riana.tauro, karthik.poosa, sk.anirban, raag.jadav,
	Mallesh Koujalagi

Introduce DRM_WEDGE_RECOVERY_COLD_RESET (BIT(4)) recovery method to handle
scenarios requiring device power cycle.

This method addresses cases where other recovery mechanisms
(driver reload, PCIe reset, etc.) are insufficient to restore device
functionality. When set, it indicates to userspace that only device power
cycle can recover the device from its current error state.

Signed-off-by: Mallesh Koujalagi <mallesh.koujalagi@intel.com>
Reviewed-by: Raag Jadav <raag.jadav@intel.com>
---
v3:
- Update any scenario that requires cold-reset. (Riana)

v4:
- Rename cold reset to power cycle. (Raag)

v5:
- Make it consistent with consumer expectations. (Raag)

v6:
- Drop "last resort" wording. (Riana)
---
 drivers/gpu/drm/drm_drv.c | 2 ++
 include/drm/drm_device.h  | 1 +
 2 files changed, 3 insertions(+)

diff --git a/drivers/gpu/drm/drm_drv.c b/drivers/gpu/drm/drm_drv.c
index e51ed959da89..8519e97ef5d3 100644
--- a/drivers/gpu/drm/drm_drv.c
+++ b/drivers/gpu/drm/drm_drv.c
@@ -545,6 +545,8 @@ static const char *drm_get_wedge_recovery(unsigned int opt)
 		return "bus-reset";
 	case DRM_WEDGE_RECOVERY_VENDOR:
 		return "vendor-specific";
+	case DRM_WEDGE_RECOVERY_COLD_RESET:
+		return "cold-reset";
 	default:
 		return NULL;
 	}
diff --git a/include/drm/drm_device.h b/include/drm/drm_device.h
index 768a8dae83c5..75f030d027ee 100644
--- a/include/drm/drm_device.h
+++ b/include/drm/drm_device.h
@@ -37,6 +37,7 @@ struct pci_controller;
 #define DRM_WEDGE_RECOVERY_REBIND	BIT(1)	/* unbind + bind driver */
 #define DRM_WEDGE_RECOVERY_BUS_RESET	BIT(2)	/* unbind + reset bus device + bind */
 #define DRM_WEDGE_RECOVERY_VENDOR	BIT(3)	/* vendor specific recovery method */
+#define DRM_WEDGE_RECOVERY_COLD_RESET	BIT(4)	/* remove device + slot power cycle + rescan */
 
 /**
  * struct drm_wedge_task_info - information about the guilty task of a wedge dev
-- 
2.48.1


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH v12 2/4] drm/doc: Document DRM_WEDGE_RECOVERY_COLD_RESET recovery method
  2026-07-24 10:03 [PATCH v12 0/4] Introduce cold reset recovery method Mallesh Koujalagi
  2026-07-24 10:03 ` [PATCH v12 1/4] drm: Add DRM_WEDGE_RECOVERY_COLD_RESET " Mallesh Koujalagi
@ 2026-07-24 10:03 ` Mallesh Koujalagi
  2026-07-24 10:13   ` sashiko-bot
  2026-07-24 10:03 ` [PATCH v12 3/4] drm/xe: Handle PUNIT errors by requesting cold-reset recovery Mallesh Koujalagi
  2026-07-24 10:03 ` [PATCH v12 4/4] drm/xe/ras: Use fault-inject to trigger punit error handler Mallesh Koujalagi
  3 siblings, 1 reply; 11+ messages in thread
From: Mallesh Koujalagi @ 2026-07-24 10:03 UTC (permalink / raw)
  To: intel-xe, dri-devel, rodrigo.vivi
  Cc: andrealmeid, christian.koenig, airlied, simona.vetter, mripard,
	maarten.lankhorst, tzimmermann, anshuman.gupta, badal.nilawar,
	riana.tauro, karthik.poosa, sk.anirban, raag.jadav,
	Mallesh Koujalagi

When ``WEDGED=cold-reset`` is sent, it indicates that the device has
encountered an error condition that cannot be resolved through other
recovery methods such as driver rebind or bus reset, and requires a
complete device power cycle to restore functionality.

Signed-off-by: Mallesh Koujalagi <mallesh.koujalagi@intel.com>
Reviewed-by: Raag Jadav <raag.jadav@intel.com>
---
v2:
- Add several instead of number to avoid update. (Jani)

v3:
- Update document with generic scenario. (Riana)
- Consistent with terminology. (Raag)
- Remove already covered information.

v4:
- Update doc. (Raag/Riana)
- Change commit message.

v5:
- Update commit message. (Raag)
- Remove unbind.
- Simplify cold-reset script.

v6:
- Look up the hotplug slot in DEVPATH instead of scanning
  every PCI slot on the system. (Raag)
- Drop arbitrary sleep values from the example script.

v7:
- Update recovery script. (Raag)

v8:
- Add rescan instead of reset. (Raag)

v10:
- Fix column width. (Sashiko)

v11:
- Make udev rules in single line. (Sashiko)
---
 Documentation/gpu/drm-uapi.rst | 93 ++++++++++++++++++++++++++++++++--
 1 file changed, 88 insertions(+), 5 deletions(-)

diff --git a/Documentation/gpu/drm-uapi.rst b/Documentation/gpu/drm-uapi.rst
index 93df92c4ac8c..52255247a6db 100644
--- a/Documentation/gpu/drm-uapi.rst
+++ b/Documentation/gpu/drm-uapi.rst
@@ -424,7 +424,7 @@ needed.
 Recovery
 --------
 
-Current implementation defines four recovery methods, out of which, drivers
+Current implementation defines several recovery methods, out of which, drivers
 can use any one, multiple or none. Method(s) of choice will be sent in the
 uevent environment as ``WEDGED=<method1>[,..,<methodN>]`` in order of less to
 more side-effects. See the section `Vendor Specific Recovery`_
@@ -434,15 +434,16 @@ method is unknown, ``WEDGED=unknown`` will be sent instead.
 Userspace consumers can parse this event and attempt recovery as per the
 following expectations.
 
-    =============== ========================================
+    =============== =========================================
     Recovery method Consumer expectations
-    =============== ========================================
+    =============== =========================================
     none            optional telemetry collection
     rebind          unbind + bind driver
     bus-reset       unbind + bus reset/re-enumeration + bind
     vendor-specific vendor specific recovery method
+    cold-reset      remove device + slot power cycle + rescan
     unknown         consumer policy
-    =============== ========================================
+    =============== =========================================
 
 No Recovery
 -----------
@@ -453,6 +454,17 @@ debug purpose in order to root cause the hang. This is useful because the first
 hang is usually the most critical one which can result in consequential hangs
 or complete wedging.
 
+Cold Reset Recovery
+-------------------
+
+When ``WEDGED=cold-reset`` is sent, it indicates that the device has
+encountered an error condition that cannot be resolved through other
+recovery methods such as driver rebind or bus reset, and requires a complete
+device power cycle to restore functionality.
+
+This method is used by devices that are plugged directly into the PCIe slot
+which supports removing the power.
+
 Vendor Specific Recovery
 ------------------------
 
@@ -516,7 +528,7 @@ Example - rebind
 
 Udev rule::
 
-    SUBSYSTEM=="drm", ENV{WEDGED}=="rebind", DEVPATH=="*/drm/card[0-9]",
+    SUBSYSTEM=="drm", ENV{WEDGED}=="rebind", DEVPATH=="*/drm/card[0-9]", \
     RUN+="/path/to/rebind.sh $env{DEVPATH}"
 
 Recovery script::
@@ -530,6 +542,77 @@ Recovery script::
     echo -n $DEVICE > $DRIVER/unbind
     echo -n $DEVICE > $DRIVER/bind
 
+Example - cold-reset
+--------------------
+
+Udev rule::
+
+    SUBSYSTEM=="drm", ENV{WEDGED}=="cold-reset", DEVPATH=="*/drm/card[0-9]", \
+    RUN+="/path/to/cold-reset.sh $env{DEVPATH}"
+
+Recovery script::
+
+    #!/bin/sh
+    die() { echo "ERROR: $*" >&2; exit 1; }
+
+    [ -n "$1" ] || die "Usage: $0 <device-path>"
+
+    PCI_DEVS=/sys/bus/pci/devices
+    PCI_SLOTS=/sys/bus/pci/slots
+
+    syspath=$(readlink -f "/sys/$1/device" 2>/dev/null || readlink -f "/sys/$1" 2>/dev/null)
+    [ -n "$syspath" ] || die "cannot resolve sysfs path for: $1"
+
+    dev=$(basename "$syspath")
+    [ -e "$PCI_DEVS/$dev" ] || die "not a PCI device: $dev"
+    echo "device : $dev"
+
+    slot=""
+    walk=$(dirname "$(readlink -f "$PCI_DEVS/$dev")")
+
+    while true; do
+        ancestor=$(basename "$walk")
+        case "$ancestor" in pci*) break ;; esac  # reached the virtual bus root
+
+        ancestor_nofn=${ancestor%.*}  # strip function: 0000:03:01.0 -> 0000:03:01
+
+        for f in "$PCI_SLOTS"/*/address; do
+            [ -f "$f" ] || continue
+            addr=$(cat "$f")
+            case "$ancestor_nofn" in
+                *"$addr") slot=$(basename "$(dirname "$f")"); break ;;
+            esac
+        done
+
+        if [ -n "$slot" ] && [ -e "$PCI_SLOTS/$slot/power" ]; then
+            echo "slot   : $slot (port $ancestor)"
+            break
+        fi
+        slot=""
+        walk=$(dirname "$walk")
+    done
+
+    [ -n "$slot" ] || die "no hotplug slot with power control found in PCIe topology"
+
+    # Cold reset: remove the device, cut slot power, restore power, rescan.
+    echo "Removing $dev..."
+    [ -e "$PCI_DEVS/$dev" ] && echo 1 > "$PCI_DEVS/$dev/remove"
+
+    echo "Powering off slot $slot..."
+    echo 0 > "$PCI_SLOTS/$slot/power"
+
+    echo "Powering on slot $slot..."
+    echo 1 > "$PCI_SLOTS/$slot/power"
+
+    echo "Rescanning PCI bus..."
+    echo 1 > /sys/bus/pci/rescan
+
+    if [ -e "$PCI_DEVS/$dev" ]; then
+        echo "Done: $dev is back online."
+    else
+        echo "WARNING: $dev did not re-appear after rescan."
+    fi
+
 Customization
 -------------
 
-- 
2.48.1


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH v12 3/4] drm/xe: Handle PUNIT errors by requesting cold-reset recovery
  2026-07-24 10:03 [PATCH v12 0/4] Introduce cold reset recovery method Mallesh Koujalagi
  2026-07-24 10:03 ` [PATCH v12 1/4] drm: Add DRM_WEDGE_RECOVERY_COLD_RESET " Mallesh Koujalagi
  2026-07-24 10:03 ` [PATCH v12 2/4] drm/doc: Document " Mallesh Koujalagi
@ 2026-07-24 10:03 ` Mallesh Koujalagi
  2026-07-24 10:29   ` sashiko-bot
  2026-07-24 10:03 ` [PATCH v12 4/4] drm/xe/ras: Use fault-inject to trigger punit error handler Mallesh Koujalagi
  3 siblings, 1 reply; 11+ messages in thread
From: Mallesh Koujalagi @ 2026-07-24 10:03 UTC (permalink / raw)
  To: intel-xe, dri-devel, rodrigo.vivi
  Cc: andrealmeid, christian.koenig, airlied, simona.vetter, mripard,
	maarten.lankhorst, tzimmermann, anshuman.gupta, badal.nilawar,
	riana.tauro, karthik.poosa, sk.anirban, raag.jadav,
	Mallesh Koujalagi

When PUNIT (power management unit) errors are detected that persist across
warm resets, mark the device as wedged with DRM_WEDGE_RECOVERY_COLD_RESET
and notify userspace that a complete device power cycle is required to
restore normal operation.

Signed-off-by: Mallesh Koujalagi <mallesh.koujalagi@intel.com>
Reviewed-by: Raag Jadav <raag.jadav@intel.com>
---
v3:
- Use PUNIT instead of PMU. (Riana)
- Use consistent wording.
- Remove log. (Raag)

v4:
- Make function static. (Raag)

v5:
- Remove kdoc for static function. (Raag)
- Remove xe_ prefix for static function.

v9:
- Remove unwanted header. (Sashiko)
---
 drivers/gpu/drm/xe/xe_ras.c | 8 +++++++-
 1 file changed, 7 insertions(+), 1 deletion(-)

diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
index a31e06b8aa67..92b4181026cb 100644
--- a/drivers/gpu/drm/xe/xe_ras.c
+++ b/drivers/gpu/drm/xe/xe_ras.c
@@ -236,6 +236,12 @@ static u8 handle_core_compute_errors(struct xe_ras_error_array *arr)
 	return XE_RAS_RECOVERY_ACTION_RECOVERED;
 }
 
+static void punit_error_handler(struct xe_device *xe)
+{
+	xe_device_set_wedged_method(xe, DRM_WEDGE_RECOVERY_COLD_RESET);
+	xe_device_declare_wedged(xe);
+}
+
 static u8 handle_soc_internal_errors(struct xe_device *xe, struct xe_ras_error_array *arr)
 {
 	struct xe_ras_soc_error *info = (void *)arr->details;
@@ -267,7 +273,7 @@ static u8 handle_soc_internal_errors(struct xe_device *xe, struct xe_ras_error_a
 			xe_err(xe, "[RAS]: PUNIT %s detected: 0x%x\n",
 			       sev_to_str(counter->common.severity),
 			       ieh_error->global_error_status);
-			/* TODO: Add PUNIT error handling */
+			punit_error_handler(xe);
 			return XE_RAS_RECOVERY_ACTION_DISCONNECT;
 		}
 	}
-- 
2.48.1


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH v12 4/4] drm/xe/ras: Use fault-inject to trigger punit error handler
  2026-07-24 10:03 [PATCH v12 0/4] Introduce cold reset recovery method Mallesh Koujalagi
                   ` (2 preceding siblings ...)
  2026-07-24 10:03 ` [PATCH v12 3/4] drm/xe: Handle PUNIT errors by requesting cold-reset recovery Mallesh Koujalagi
@ 2026-07-24 10:03 ` Mallesh Koujalagi
  2026-07-24 10:20   ` sashiko-bot
                     ` (2 more replies)
  3 siblings, 3 replies; 11+ messages in thread
From: Mallesh Koujalagi @ 2026-07-24 10:03 UTC (permalink / raw)
  To: intel-xe, dri-devel, rodrigo.vivi
  Cc: andrealmeid, christian.koenig, airlied, simona.vetter, mripard,
	maarten.lankhorst, tzimmermann, anshuman.gupta, badal.nilawar,
	riana.tauro, karthik.poosa, sk.anirban, raag.jadav,
	Mallesh Koujalagi

Use fault-inject framework to trigger punit_error_handler()
for testing.

Usage:
  echo 100 > .../inject_punit_error/probability
  echo 1   > .../inject_punit_error/times

Signed-off-by: Mallesh Koujalagi <mallesh.koujalagi@intel.com>
---
 drivers/gpu/drm/xe/xe_debugfs.c |  4 ++++
 drivers/gpu/drm/xe/xe_debugfs.h |  2 ++
 drivers/gpu/drm/xe/xe_ras.c     | 17 +++++++++++++++++
 3 files changed, 23 insertions(+)

diff --git a/drivers/gpu/drm/xe/xe_debugfs.c b/drivers/gpu/drm/xe/xe_debugfs.c
index 5a3877fcb0f0..5599d493feca 100644
--- a/drivers/gpu/drm/xe/xe_debugfs.c
+++ b/drivers/gpu/drm/xe/xe_debugfs.c
@@ -42,6 +42,7 @@
 
 DECLARE_FAULT_ATTR(gt_reset_failure);
 DECLARE_FAULT_ATTR(inject_csc_hw_error);
+DECLARE_FAULT_ATTR(inject_punit_error);
 
 static bool csc_hw_error_available(struct xe_device *xe)
 {
@@ -62,6 +63,8 @@ static struct {
 	{ .name = "inject_csc_hw_error",
 	  .attr = &inject_csc_hw_error,
 	  .is_visible = csc_hw_error_available },
+	{ .name = "inject_punit_error",
+	  .attr = &inject_punit_error },
 };
 
 /*
@@ -76,6 +79,7 @@ bool xe_fault_##name(void)				\
 
 FAULT_ACTION(gt_reset, gt_reset_failure)
 FAULT_ACTION(csc_hw_error, inject_csc_hw_error)
+FAULT_ACTION(punit_error, inject_punit_error)
 
 static void xe_fault_inject_debugfs_register(struct xe_device *xe,
 					     struct dentry *root)
diff --git a/drivers/gpu/drm/xe/xe_debugfs.h b/drivers/gpu/drm/xe/xe_debugfs.h
index cd56f7442b99..dd57914dd4f2 100644
--- a/drivers/gpu/drm/xe/xe_debugfs.h
+++ b/drivers/gpu/drm/xe/xe_debugfs.h
@@ -13,10 +13,12 @@ struct xe_device;
 #ifdef CONFIG_DEBUG_FS
 bool xe_fault_gt_reset(void);
 bool xe_fault_csc_hw_error(void);
+bool xe_fault_punit_error(void);
 void xe_debugfs_register(struct xe_device *xe);
 #else
 static inline bool xe_fault_gt_reset(void) { return false; }
 static inline bool xe_fault_csc_hw_error(void) { return false; }
+static inline bool xe_fault_punit_error(void) { return false; }
 static inline void xe_debugfs_register(struct xe_device *xe) { }
 #endif
 
diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
index 92b4181026cb..da06c16cf7ed 100644
--- a/drivers/gpu/drm/xe/xe_ras.c
+++ b/drivers/gpu/drm/xe/xe_ras.c
@@ -3,6 +3,7 @@
  * Copyright © 2026 Intel Corporation
  */
 
+#include "xe_debugfs.h"
 #include "xe_device.h"
 #include "xe_drm_ras.h"
 #include "xe_pm.h"
@@ -391,6 +392,22 @@ enum xe_ras_recovery_action xe_ras_process_errors(struct xe_device *xe)
 	size_t rlen;
 	int ret;
 
+	/*
+	 * Only allow the injected PUNIT error once the DRM device is registered.
+	 * xe_ras_process_errors() also runs during probe (via xe_ras_init()), before
+	 * drm_dev_register() calls device_add() on the DRM minor's kdev. At that
+	 * point kdev->kobj.parent is not yet linked into the sysfs hierarchy, so
+	 * kobject_get_path() returns "/card0" instead of the real sysfs path.
+	 * drm_dev_wedged_event() would then emit a KOBJ_CHANGE uevent with a wrong
+	 * DEVPATH that udev cannot resolve, silently dropping the event and leaving
+	 * the cold-reset recovery broken.
+	 */
+	if (xe->drm.registered && xe_fault_punit_error()) {
+		xe_err(xe, "[RAS]: PUNIT error injected\n");
+		punit_error_handler(xe);
+		return XE_RAS_RECOVERY_ACTION_DISCONNECT;
+	}
+
 	if (!xe->info.has_sysctrl)
 		return XE_RAS_RECOVERY_ACTION_RESET;
 
-- 
2.48.1


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* Re: [PATCH v12 2/4] drm/doc: Document DRM_WEDGE_RECOVERY_COLD_RESET recovery method
  2026-07-24 10:03 ` [PATCH v12 2/4] drm/doc: Document " Mallesh Koujalagi
@ 2026-07-24 10:13   ` sashiko-bot
  0 siblings, 0 replies; 11+ messages in thread
From: sashiko-bot @ 2026-07-24 10:13 UTC (permalink / raw)
  To: Mallesh Koujalagi; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] The cold-reset recovery script fails to find the hotplug slot for standard single-endpoint PCIe GPUs because it starts its search at the parent bridge, skipping the device itself.
--

--- Patch [2]: [PATCH v12 2/4] drm/doc: Document DRM_WEDGE_RECOVERY_COLD_RESET recovery method ---
commit 40a35f3de596a4e78fa7c7c9f9433ec0c0f1b456
Author: Mallesh Koujalagi <mallesh.koujalagi@intel.com>
drm/doc: Document DRM_WEDGE_RECOVERY_COLD_RESET recovery method

This commit adds documentation for the new "cold-reset" wedge recovery
method. It updates the expectations table and provides an example udev
rule and recovery script to perform a cold reset by removing the device,
toggling slot power, and rescanning the PCI bus.

> diff --git a/Documentation/gpu/drm-uapi.rst b/Documentation/gpu/drm-uapi.rst
> index 93df92c4ac8cf..52255247a6db9 100644
> --- a/Documentation/gpu/drm-uapi.rst
> +++ b/Documentation/gpu/drm-uapi.rst
[ ... ]
> @@ -530,6 +542,77 @@ Recovery script::
>      echo -n $DEVICE > $DRIVER/unbind
>      echo -n $DEVICE > $DRIVER/bind
>  
> +Example - cold-reset
> +--------------------
> +
> +Udev rule::
> +
> +    SUBSYSTEM=="drm", ENV{WEDGED}=="cold-reset", DEVPATH=="*/drm/card[0-9]", \
> +    RUN+="/path/to/cold-reset.sh $env{DEVPATH}"
> +
> +Recovery script::
> +
> +    #!/bin/sh
> +    die() { echo "ERROR: $*" >&2; exit 1; }
> +
> +    [ -n "$1" ] || die "Usage: $0 <device-path>"
> +
> +    PCI_DEVS=/sys/bus/pci/devices
> +    PCI_SLOTS=/sys/bus/pci/slots
> +
> +    syspath=$(readlink -f "/sys/$1/device" 2>/dev/null || readlink -f "/sys/$1" 2>/dev/null)
> +    [ -n "$syspath" ] || die "cannot resolve sysfs path for: $1"
> +
> +    dev=$(basename "$syspath")
> +    [ -e "$PCI_DEVS/$dev" ] || die "not a PCI device: $dev"
> +    echo "device : $dev"
> +
> +    slot=""
> +    walk=$(dirname "$(readlink -f "$PCI_DEVS/$dev")")

[Severity: Medium]
Will this fail to find the hotplug slot for standard single-endpoint PCIe GPUs?

By initializing walk with dirname here, the script appears to start its search
at the parent bridge, skipping the device itself. Since the device's own
address typically matches the hotplug slot address, this skips checking the
device's slot and will result in no hotplug slot being found for single
endpoint devices.

> +
> +    while true; do
> +        ancestor=$(basename "$walk")
> +        case "$ancestor" in pci*) break ;; esac  # reached the virtual bus root

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260724100302.706685-6-mallesh.koujalagi@intel.com?part=2

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v12 1/4] drm: Add DRM_WEDGE_RECOVERY_COLD_RESET recovery method
  2026-07-24 10:03 ` [PATCH v12 1/4] drm: Add DRM_WEDGE_RECOVERY_COLD_RESET " Mallesh Koujalagi
@ 2026-07-24 10:16   ` sashiko-bot
  0 siblings, 0 replies; 11+ messages in thread
From: sashiko-bot @ 2026-07-24 10:16 UTC (permalink / raw)
  To: Mallesh Koujalagi; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] The newly added `DRM_WEDGE_RECOVERY_COLD_RESET` recovery method is not documented in the UAPI documentation.
--

--- Patch [1]: [PATCH v12 1/4] drm: Add DRM_WEDGE_RECOVERY_COLD_RESET recovery method ---
commit 1acdd09ef42a2d7b1fde40c873937c144a20b544
Author: Mallesh Koujalagi <mallesh.koujalagi@intel.com>

drm: Add DRM_WEDGE_RECOVERY_COLD_RESET recovery method

This commit introduces a new recovery method, DRM_WEDGE_RECOVERY_COLD_RESET,
to handle scenarios requiring a device power cycle. It adds the "cold-reset"
string mapping and the corresponding bit flag.

> diff --git a/drivers/gpu/drm/drm_drv.c b/drivers/gpu/drm/drm_drv.c
> index 1ff0bf7cba6aa..94abe7a811a52 100644
> --- a/drivers/gpu/drm/drm_drv.c
> +++ b/drivers/gpu/drm/drm_drv.c
> @@ -537,6 +537,8 @@ static const char *drm_get_wedge_recovery(unsigned int opt)
>  		return "bus-reset";
>  	case DRM_WEDGE_RECOVERY_VENDOR:
>  		return "vendor-specific";
> +	case DRM_WEDGE_RECOVERY_COLD_RESET:
> +		return "cold-reset";

[Severity: Medium]
Should this new recovery method be documented in Documentation/gpu/drm-uapi.rst?

Looking at Documentation/gpu/drm-uapi.rst, there is a table mapping recovery
methods to consumer expectations:

    =============== ========================================
    Recovery method Consumer expectations
    =============== ========================================
    none            optional telemetry collection
    rebind          unbind + bind driver
    bus-reset       unbind + bus reset/re-enumeration + bind
    vendor-specific vendor specific recovery method
    unknown         consumer policy
    =============== ========================================

Could "cold-reset" be missing from this UAPI documentation, leaving userspace
developers without a canonical reference for its expected consumer policies?

>  	default:
>  		return NULL;
>  	}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260724100302.706685-6-mallesh.koujalagi@intel.com?part=1

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v12 4/4] drm/xe/ras: Use fault-inject to trigger punit error handler
  2026-07-24 10:03 ` [PATCH v12 4/4] drm/xe/ras: Use fault-inject to trigger punit error handler Mallesh Koujalagi
@ 2026-07-24 10:20   ` sashiko-bot
  2026-07-28  8:32   ` Tauro, Riana
  2026-07-28 19:25   ` Rodrigo Vivi
  2 siblings, 0 replies; 11+ messages in thread
From: sashiko-bot @ 2026-07-24 10:20 UTC (permalink / raw)
  To: Mallesh Koujalagi; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Real hardware errors (PUNIT/CSC) detected during probe trigger a broken KOBJ_CHANGE uevent with an unresolvable DEVPATH, dropping the event silently and breaking cold-reset recovery. The patch works around this exclusively for injected errors, leaving the real bug unhandled.
--

--- Patch [4]: [PATCH v12 4/4] drm/xe/ras: Use fault-inject to trigger punit error handler ---
commit 7bec129b8ce98871b6643acdf5450a3fe7c276cd
Author: Mallesh Koujalagi <mallesh.koujalagi@intel.com>

drm/xe/ras: Use fault-inject to trigger punit error handler

Use fault-inject framework to trigger punit_error_handler()
for testing.

> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index 92b4181026cbe..da06c16cf7ed9 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -391,6 +392,22 @@ enum xe_ras_recovery_action xe_ras_process_errors(struct xe_device *xe)
>  	size_t rlen;
>  	int ret;
>  
> +	/*
> +	 * Only allow the injected PUNIT error once the DRM device is registered.
> +	 * xe_ras_process_errors() also runs during probe (via xe_ras_init()), before
> +	 * drm_dev_register() calls device_add() on the DRM minor's kdev. At that
> +	 * point kdev->kobj.parent is not yet linked into the sysfs hierarchy, so
> +	 * kobject_get_path() returns "/card0" instead of the real sysfs path.
> +	 * drm_dev_wedged_event() would then emit a KOBJ_CHANGE uevent with a wrong
> +	 * DEVPATH that udev cannot resolve, silently dropping the event and leaving
> +	 * the cold-reset recovery broken.
> +	 */
> +	if (xe->drm.registered && xe_fault_punit_error()) {
> +		xe_err(xe, "[RAS]: PUNIT error injected\n");
> +		punit_error_handler(xe);
> +		return XE_RAS_RECOVERY_ACTION_DISCONNECT;
> +	}

[Severity: High]
Does this same broken uevent issue also apply to real hardware errors detected
during probe?

The comment correctly identifies that calling punit_error_handler() before 
the DRM device is registered emits a KOBJ_CHANGE uevent with a wrong DEVPATH.
However, if a real PUNIT error is detected during probe, the code bypasses
this guard and proceeds to handle_soc_internal_errors():

drivers/gpu/drm/xe/xe_ras.c:xe_ras_process_errors() {
	...
	case XE_RAS_COMP_SOC_INTERNAL:
		action = handle_soc_internal_errors(xe, arr);
		break;
	...
}

Which unconditionally calls punit_error_handler() without the
xe->drm.registered check:

drivers/gpu/drm/xe/xe_ras.c:handle_soc_internal_errors() {
	...
	if (ieh_error->global_error_status & XE_RAS_SOC_IEH_PUNIT) {
		xe_err(xe, "[RAS]: PUNIT %s detected: 0x%x\n",
		       sev_to_str(counter->common.severity),
		       ieh_error->global_error_status);
		punit_error_handler(xe);
		return XE_RAS_RECOVERY_ACTION_DISCONNECT;
	}
	...
}

Should the underlying uevent emission in punit_error_handler() be delayed
or guarded until after registration, rather than only working around it for
injected errors?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260724100302.706685-6-mallesh.koujalagi@intel.com?part=4

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v12 3/4] drm/xe: Handle PUNIT errors by requesting cold-reset recovery
  2026-07-24 10:03 ` [PATCH v12 3/4] drm/xe: Handle PUNIT errors by requesting cold-reset recovery Mallesh Koujalagi
@ 2026-07-24 10:29   ` sashiko-bot
  0 siblings, 0 replies; 11+ messages in thread
From: sashiko-bot @ 2026-07-24 10:29 UTC (permalink / raw)
  To: Mallesh Koujalagi; +Cc: dri-devel

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] Calling `xe_device_declare_wedged()` during driver probe silently drops the `WEDGED` uevent because the DRM device is not yet registered in sysfs, and `xe_ras_init()` ignores the disconnect return value, leaving userspace unaware of a permanently wedged device.
- [High] Unprotected writes to `xe->wedged.method` cause a data race with concurrent wedge events, which can overwrite the critical `DRM_WEDGE_RECOVERY_COLD_RESET` recovery method, and lack of state checking leads to redundant uevent flooding.
--

--- Patch [3]: [PATCH v12 3/4] drm/xe: Handle PUNIT errors by requesting cold-reset recovery ---
commit 340a6fc7bdb22fd2c21615ec29270c2b330f42ea
Author: Mallesh Koujalagi <mallesh.koujalagi@intel.com>

drm/xe: Handle PUNIT errors by requesting cold-reset recovery

This commit handles PUNIT errors by setting the wedge recovery method to
cold reset and declaring the device wedged. This ensures that userspace is
notified that a complete device power cycle is required to restore normal
operation.

> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index a31e06b8aa67a..92b4181026cbe 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -236,6 +236,12 @@ static u8 handle_core_compute_errors(struct xe_ras_error_array *arr)
>  	return XE_RAS_RECOVERY_ACTION_RECOVERED;
>  }
>  
> +static void punit_error_handler(struct xe_device *xe)
> +{
> +	xe_device_set_wedged_method(xe, DRM_WEDGE_RECOVERY_COLD_RESET);

[Severity: High]
Does this unprotected write to xe->wedged.method create a data race?

If another thread (e.g., PCI AER or GT reset) concurrently calls
xe_device_declare_wedged(), it could evaluate if (!xe->wedged.method) before
this assignment, and overwrite DRM_WEDGE_RECOVERY_COLD_RESET with the default
DRM_WEDGE_RECOVERY_REBIND | DRM_WEDGE_RECOVERY_BUS_RESET.

Also, since xe_ras_process_errors() processes multiple errors in a loop:

drivers/gpu/drm/xe/xe_ras.c:xe_ras_process_errors() {
    ...
    for (int i = 0; i < response.num_errors && i < XE_RAS_NUM_ERROR_ARR; i++) {
        // error handler called
    ...
}

Will this cause a flood of redundant uevents and wedge method assignments if
multiple PUNIT errors are returned in a single response?

> +	xe_device_declare_wedged(xe);

[Severity: High]
Can this call silently drop the WEDGED uevent if it occurs during driver probe?

If a PUNIT error is detected during early probe via:
xe_device_probe() -> xe_ras_init() -> xe_ras_process_errors() -> 
handle_soc_internal_errors() -> punit_error_handler()

The uevent sent by drm_dev_wedged_event() (called from
xe_device_declare_wedged()) will fail with -EINVAL because the DRM device is
not yet registered in sysfs (kobj->state_in_sysfs is false).

Furthermore, since xe_ras_init() ignores the XE_RAS_RECOVERY_ACTION_DISCONNECT
return value, will the system boot with a wedged GPU while userspace is
completely unaware?

> +}

[ ... ]

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260724100302.706685-6-mallesh.koujalagi@intel.com?part=3

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v12 4/4] drm/xe/ras: Use fault-inject to trigger punit error handler
  2026-07-24 10:03 ` [PATCH v12 4/4] drm/xe/ras: Use fault-inject to trigger punit error handler Mallesh Koujalagi
  2026-07-24 10:20   ` sashiko-bot
@ 2026-07-28  8:32   ` Tauro, Riana
  2026-07-28 19:25   ` Rodrigo Vivi
  2 siblings, 0 replies; 11+ messages in thread
From: Tauro, Riana @ 2026-07-28  8:32 UTC (permalink / raw)
  To: Mallesh Koujalagi, intel-xe, dri-devel, rodrigo.vivi
  Cc: andrealmeid, christian.koenig, airlied, simona.vetter, mripard,
	maarten.lankhorst, tzimmermann, anshuman.gupta, badal.nilawar,
	karthik.poosa, sk.anirban, raag.jadav

Hi Mallesh

On 24-07-2026 15:33, Mallesh Koujalagi wrote:
> Use fault-inject framework to trigger punit_error_handler()
> for testing.
>
> Usage:
>    echo 100 > .../inject_punit_error/probability
>    echo 1   > .../inject_punit_error/times

I think it would be better to have this as wedged fault injection 
instead of punit,
This fault injection is only sending a wedged uevent.

We could combine both wedged (cold_reset for punit and vendor_specific 
for csc) into one
if there is a better way to do it.  But that could be seperate patch.

Thanks
Riana

>
> Signed-off-by: Mallesh Koujalagi <mallesh.koujalagi@intel.com>
> ---
>   drivers/gpu/drm/xe/xe_debugfs.c |  4 ++++
>   drivers/gpu/drm/xe/xe_debugfs.h |  2 ++
>   drivers/gpu/drm/xe/xe_ras.c     | 17 +++++++++++++++++
>   3 files changed, 23 insertions(+)
>
> diff --git a/drivers/gpu/drm/xe/xe_debugfs.c b/drivers/gpu/drm/xe/xe_debugfs.c
> index 5a3877fcb0f0..5599d493feca 100644
> --- a/drivers/gpu/drm/xe/xe_debugfs.c
> +++ b/drivers/gpu/drm/xe/xe_debugfs.c
> @@ -42,6 +42,7 @@
>   
>   DECLARE_FAULT_ATTR(gt_reset_failure);
>   DECLARE_FAULT_ATTR(inject_csc_hw_error);
> +DECLARE_FAULT_ATTR(inject_punit_error);
>   
>   static bool csc_hw_error_available(struct xe_device *xe)
>   {
> @@ -62,6 +63,8 @@ static struct {
>   	{ .name = "inject_csc_hw_error",
>   	  .attr = &inject_csc_hw_error,
>   	  .is_visible = csc_hw_error_available },
> +	{ .name = "inject_punit_error",
> +	  .attr = &inject_punit_error },
>   };
>   
>   /*
> @@ -76,6 +79,7 @@ bool xe_fault_##name(void)				\
>   
>   FAULT_ACTION(gt_reset, gt_reset_failure)
>   FAULT_ACTION(csc_hw_error, inject_csc_hw_error)
> +FAULT_ACTION(punit_error, inject_punit_error)
>   
>   static void xe_fault_inject_debugfs_register(struct xe_device *xe,
>   					     struct dentry *root)
> diff --git a/drivers/gpu/drm/xe/xe_debugfs.h b/drivers/gpu/drm/xe/xe_debugfs.h
> index cd56f7442b99..dd57914dd4f2 100644
> --- a/drivers/gpu/drm/xe/xe_debugfs.h
> +++ b/drivers/gpu/drm/xe/xe_debugfs.h
> @@ -13,10 +13,12 @@ struct xe_device;
>   #ifdef CONFIG_DEBUG_FS
>   bool xe_fault_gt_reset(void);
>   bool xe_fault_csc_hw_error(void);
> +bool xe_fault_punit_error(void);
>   void xe_debugfs_register(struct xe_device *xe);
>   #else
>   static inline bool xe_fault_gt_reset(void) { return false; }
>   static inline bool xe_fault_csc_hw_error(void) { return false; }
> +static inline bool xe_fault_punit_error(void) { return false; }
>   static inline void xe_debugfs_register(struct xe_device *xe) { }
>   #endif
>   
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index 92b4181026cb..da06c16cf7ed 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -3,6 +3,7 @@
>    * Copyright © 2026 Intel Corporation
>    */
>   
> +#include "xe_debugfs.h"
>   #include "xe_device.h"
>   #include "xe_drm_ras.h"
>   #include "xe_pm.h"
> @@ -391,6 +392,22 @@ enum xe_ras_recovery_action xe_ras_process_errors(struct xe_device *xe)
>   	size_t rlen;
>   	int ret;
>   
> +	/*
> +	 * Only allow the injected PUNIT error once the DRM device is registered.
> +	 * xe_ras_process_errors() also runs during probe (via xe_ras_init()), before
> +	 * drm_dev_register() calls device_add() on the DRM minor's kdev. At that
> +	 * point kdev->kobj.parent is not yet linked into the sysfs hierarchy, so
> +	 * kobject_get_path() returns "/card0" instead of the real sysfs path.
> +	 * drm_dev_wedged_event() would then emit a KOBJ_CHANGE uevent with a wrong
> +	 * DEVPATH that udev cannot resolve, silently dropping the event and leaving
> +	 * the cold-reset recovery broken.
> +	 */
> +	if (xe->drm.registered && xe_fault_punit_error()) {
> +		xe_err(xe, "[RAS]: PUNIT error injected\n");
> +		punit_error_handler(xe);
> +		return XE_RAS_RECOVERY_ACTION_DISCONNECT;
> +	}
> +
>   	if (!xe->info.has_sysctrl)
>   		return XE_RAS_RECOVERY_ACTION_RESET;
>   

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v12 4/4] drm/xe/ras: Use fault-inject to trigger punit error handler
  2026-07-24 10:03 ` [PATCH v12 4/4] drm/xe/ras: Use fault-inject to trigger punit error handler Mallesh Koujalagi
  2026-07-24 10:20   ` sashiko-bot
  2026-07-28  8:32   ` Tauro, Riana
@ 2026-07-28 19:25   ` Rodrigo Vivi
  2 siblings, 0 replies; 11+ messages in thread
From: Rodrigo Vivi @ 2026-07-28 19:25 UTC (permalink / raw)
  To: Mallesh Koujalagi
  Cc: intel-xe, dri-devel, andrealmeid, christian.koenig, airlied,
	simona.vetter, mripard, maarten.lankhorst, tzimmermann,
	anshuman.gupta, badal.nilawar, riana.tauro, karthik.poosa,
	sk.anirban, raag.jadav

On Fri, Jul 24, 2026 at 03:33:07PM +0530, Mallesh Koujalagi wrote:
> Use fault-inject framework to trigger punit_error_handler()
> for testing.
> 
> Usage:
>   echo 100 > .../inject_punit_error/probability
>   echo 1   > .../inject_punit_error/times
> 
> Signed-off-by: Mallesh Koujalagi <mallesh.koujalagi@intel.com>
> ---
>  drivers/gpu/drm/xe/xe_debugfs.c |  4 ++++
>  drivers/gpu/drm/xe/xe_debugfs.h |  2 ++
>  drivers/gpu/drm/xe/xe_ras.c     | 17 +++++++++++++++++
>  3 files changed, 23 insertions(+)
> 
> diff --git a/drivers/gpu/drm/xe/xe_debugfs.c b/drivers/gpu/drm/xe/xe_debugfs.c
> index 5a3877fcb0f0..5599d493feca 100644
> --- a/drivers/gpu/drm/xe/xe_debugfs.c
> +++ b/drivers/gpu/drm/xe/xe_debugfs.c
> @@ -42,6 +42,7 @@
>  
>  DECLARE_FAULT_ATTR(gt_reset_failure);
>  DECLARE_FAULT_ATTR(inject_csc_hw_error);
> +DECLARE_FAULT_ATTR(inject_punit_error);
>  
>  static bool csc_hw_error_available(struct xe_device *xe)
>  {
> @@ -62,6 +63,8 @@ static struct {
>  	{ .name = "inject_csc_hw_error",
>  	  .attr = &inject_csc_hw_error,
>  	  .is_visible = csc_hw_error_available },
> +	{ .name = "inject_punit_error",
> +	  .attr = &inject_punit_error },
>  };
>  
>  /*
> @@ -76,6 +79,7 @@ bool xe_fault_##name(void)				\
>  
>  FAULT_ACTION(gt_reset, gt_reset_failure)
>  FAULT_ACTION(csc_hw_error, inject_csc_hw_error)
> +FAULT_ACTION(punit_error, inject_punit_error)
>  
>  static void xe_fault_inject_debugfs_register(struct xe_device *xe,
>  					     struct dentry *root)
> diff --git a/drivers/gpu/drm/xe/xe_debugfs.h b/drivers/gpu/drm/xe/xe_debugfs.h
> index cd56f7442b99..dd57914dd4f2 100644
> --- a/drivers/gpu/drm/xe/xe_debugfs.h
> +++ b/drivers/gpu/drm/xe/xe_debugfs.h
> @@ -13,10 +13,12 @@ struct xe_device;
>  #ifdef CONFIG_DEBUG_FS
>  bool xe_fault_gt_reset(void);
>  bool xe_fault_csc_hw_error(void);
> +bool xe_fault_punit_error(void);
>  void xe_debugfs_register(struct xe_device *xe);
>  #else
>  static inline bool xe_fault_gt_reset(void) { return false; }
>  static inline bool xe_fault_csc_hw_error(void) { return false; }
> +static inline bool xe_fault_punit_error(void) { return false; }
>  static inline void xe_debugfs_register(struct xe_device *xe) { }
>  #endif
>  
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index 92b4181026cb..da06c16cf7ed 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -3,6 +3,7 @@
>   * Copyright © 2026 Intel Corporation
>   */
>  
> +#include "xe_debugfs.h"
>  #include "xe_device.h"
>  #include "xe_drm_ras.h"
>  #include "xe_pm.h"
> @@ -391,6 +392,22 @@ enum xe_ras_recovery_action xe_ras_process_errors(struct xe_device *xe)
>  	size_t rlen;
>  	int ret;
>  
> +	/*
> +	 * Only allow the injected PUNIT error once the DRM device is registered.

That entirely defeats the purpose! :)

Sashiko identified this corner case so I asked you to ensure that this is covered
and to prove that you should provide a fault-inject with test case.

But then you are skipping exactly the corner case. That makes absolutely no sense.

> +	 * xe_ras_process_errors() also runs during probe (via xe_ras_init()), before
> +	 * drm_dev_register() calls device_add() on the DRM minor's kdev. At that
> +	 * point kdev->kobj.parent is not yet linked into the sysfs hierarchy, so
> +	 * kobject_get_path() returns "/card0" instead of the real sysfs path.
> +	 * drm_dev_wedged_event() would then emit a KOBJ_CHANGE uevent with a wrong
> +	 * DEVPATH that udev cannot resolve, silently dropping the event and leaving
> +	 * the cold-reset recovery broken.
> +	 */
> +	if (xe->drm.registered && xe_fault_punit_error()) {
> +		xe_err(xe, "[RAS]: PUNIT error injected\n");
> +		punit_error_handler(xe);
> +		return XE_RAS_RECOVERY_ACTION_DISCONNECT;
> +	}
> +
>  	if (!xe->info.has_sysctrl)
>  		return XE_RAS_RECOVERY_ACTION_RESET;
>  
> -- 
> 2.48.1
> 

^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-07-28 19:25 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-24 10:03 [PATCH v12 0/4] Introduce cold reset recovery method Mallesh Koujalagi
2026-07-24 10:03 ` [PATCH v12 1/4] drm: Add DRM_WEDGE_RECOVERY_COLD_RESET " Mallesh Koujalagi
2026-07-24 10:16   ` sashiko-bot
2026-07-24 10:03 ` [PATCH v12 2/4] drm/doc: Document " Mallesh Koujalagi
2026-07-24 10:13   ` sashiko-bot
2026-07-24 10:03 ` [PATCH v12 3/4] drm/xe: Handle PUNIT errors by requesting cold-reset recovery Mallesh Koujalagi
2026-07-24 10:29   ` sashiko-bot
2026-07-24 10:03 ` [PATCH v12 4/4] drm/xe/ras: Use fault-inject to trigger punit error handler Mallesh Koujalagi
2026-07-24 10:20   ` sashiko-bot
2026-07-28  8:32   ` Tauro, Riana
2026-07-28 19:25   ` Rodrigo Vivi

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox