From: Rodrigo Vivi <rodrigo.vivi@intel.com>
To: Mallesh Koujalagi <mallesh.koujalagi@intel.com>
Cc: <intel-xe@lists.freedesktop.org>,
<dri-devel@lists.freedesktop.org>, <andrealmeid@igalia.com>,
<christian.koenig@amd.com>, <airlied@gmail.com>,
<simona.vetter@ffwll.ch>, <mripard@kernel.org>,
<maarten.lankhorst@linux.intel.com>, <tzimmermann@suse.de>,
<anshuman.gupta@intel.com>, <badal.nilawar@intel.com>,
<riana.tauro@intel.com>, <karthik.poosa@intel.com>,
<sk.anirban@intel.com>, <raag.jadav@intel.com>
Subject: Re: [PATCH v13 0/4] Introduce cold reset recovery method
Date: Thu, 13 Aug 2026 09:44:22 -0400 [thread overview]
Message-ID: <an3KNgWyfgxy_8Yd@intel.com> (raw)
In-Reply-To: <20260805071152.1225416-6-mallesh.koujalagi@intel.com>
On Wed, Aug 05, 2026 at 12:41:53PM +0530, Mallesh Koujalagi wrote:
> Add support for handling errors that require a complete
> device power cycle (cold reset) to recover.
>
> Certain error conditions leave the device in a persistent hardware
> error state that cannot be cleared through existing recovery mechanisms
> such as driver reload or PCIe reset. In these cases, functionality can
> only be restored by performing a cold reset.
>
> To support this, the series introduces a new DRM wedging recovery
> method, DRM_WEDGE_RECOVERY_COLD_RESET (BIT(4)). When a device is wedged
> with this method, the DRM core notifies userspace via a uevent that a cold
> reset is required. This allows userspace to take appropriate action to
> power-cycle the device.
>
> Example uevent received:
> SUBSYSTEM=drm
> WEDGED=cold-reset
> DEVPATH=/devices/.../drm/card0
>
> v2:
> - Add use case: Handling errors from power management unit,
> which requires a complete power cycle to
> recover. (Christian)
> - Add several instead of number to avoid update. (Jani)
>
> v3:
> - Update any scenario that requires cold-reset. (Riana)
> - Update document with generic scenario. (Riana)
> - Consistent with terminology. (Raag)
> - Remove already covered information.
> - Use PUNIT instead of PMU. (Riana)
> - Use consistent wordingi.
> - Remove log. (Raag)
>
> v4:
> - Rename cold reset to power cyclce. (Raag)
> - Update doc. (Raag/Riana)
> - Change commit message. (Raag)
> - Make function static. (Raag)
>
> v5:
> - Make it consistent with consumer expectations. (Raag)
> - Update commit message.
> - Remove unbind.
> - Simplify cold-reset script.
> - Remove kdoc for static function.
> - Remove xe_ prefix for static function.
>
> v6:
> - Drop "last resort" wording. (Riana)
> - Look up the hotplug slot in DEVPATH instead of scanning
> every PCI slot on the system. (Raag)
> - Drop arbitrary sleep values from the example script.
> - Expand commit message to explain why SUR_DN is masked. (Raag/Riana)
> - Check Slot Implemented bit before reading Slot Capabilities, per
> PCIe spec. (Riana)
> - Add debug log.
>
> v7:
> - Update recovery script. (Raag)
> - Handle surprise link down event properly. (Aravind/Riana)
> - Update commit message. (Riana)
> - Correct log message.
>
> v8:
> - Add rescan instead of reset. (Raag)
> - Use find_usp_dev() in punit_error_handler() function.
>
> v9:
> - Remove unwanted header. (Sashiko)
> - Removed #ifdef CONFIG_PCIEAER. (Riana)
> - Used pci_find_ext_capability() instead of usp->aer_cap.
> - Clear the PCI_ERR_UNC_SURPDN status bit (W1C) after
> reset complete. (Lukas Wunner)
> - Use pci_clear_and_set_config_dword() helper.
>
> v10:
> - Rebase.
> - Fix column width. (Sashiko)
>
> v11:
> - Make udev rules in single line. (Sashiko)
>
> v12:
> - Trigger punit handler using fault-inject.
>
> v13:
> - Rebase.
> - Rename inject_punit_error to wedge_cold_reset. (Riana)
> - Sashiko corner case issue addressed with
> commit 20bc4883c7c0 ("drm/xe/ras: Fix boot-time ras error processing").
pushed to drm-xe-next, thanks for the patch, reviews and acks
>
> Cc: André Almeida <andrealmeid@igalia.com>
> Cc: Christian König <christian.koenig@amd.com>
> Cc: David Airlie <airlied@gmail.com>
> Cc: Simona Vetter <simona.vetter@ffwll.ch>
> Cc: Maxime Ripard <mripard@kernel.org>
> Cc: Maarten Lankhorst <maarten.lankhorst@linux.intel.com>
> Cc: Thomas Zimmermann <tzimmermann@suse.de>
>
> Mallesh Koujalagi (4):
> drm: Add DRM_WEDGE_RECOVERY_COLD_RESET recovery method
> drm/doc: Document DRM_WEDGE_RECOVERY_COLD_RESET recovery method
> drm/xe: Handle PUNIT errors by requesting cold-reset recovery
> drm/xe/ras: Use fault-inject to trigger cold-reset wedge
>
> Documentation/gpu/drm-uapi.rst | 93 +++++++++++++++++++++++++++++++--
> drivers/gpu/drm/drm_drv.c | 2 +
> drivers/gpu/drm/xe/xe_debugfs.c | 4 ++
> drivers/gpu/drm/xe/xe_debugfs.h | 2 +
> drivers/gpu/drm/xe/xe_ras.c | 15 +++++-
> include/drm/drm_device.h | 1 +
> 6 files changed, 111 insertions(+), 6 deletions(-)
>
> --
> 2.48.1
>
prev parent reply other threads:[~2026-08-13 13:44 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-05 7:11 [PATCH v13 0/4] Introduce cold reset recovery method Mallesh Koujalagi
2026-08-05 7:11 ` [PATCH v13 1/4] drm: Add DRM_WEDGE_RECOVERY_COLD_RESET " Mallesh Koujalagi
2026-08-05 7:19 ` sashiko-bot
2026-08-05 7:11 ` [PATCH v13 2/4] drm/doc: Document " Mallesh Koujalagi
2026-08-05 7:20 ` sashiko-bot
2026-08-05 7:11 ` [PATCH v13 3/4] drm/xe: Handle PUNIT errors by requesting cold-reset recovery Mallesh Koujalagi
2026-08-05 7:30 ` sashiko-bot
2026-08-05 7:11 ` [PATCH v13 4/4] drm/xe/ras: Use fault-inject to trigger cold-reset wedge Mallesh Koujalagi
2026-08-05 7:24 ` sashiko-bot
2026-08-05 9:13 ` Raag Jadav
2026-08-05 10:01 ` Mallesh, Koujalagi
2026-08-05 11:50 ` ✓ CI.KUnit: success for Introduce cold reset recovery method (rev13) Patchwork
2026-08-05 12:31 ` ✗ Xe.CI.BAT: failure " Patchwork
2026-08-05 22:59 ` ✗ Xe.CI.FULL: " Patchwork
2026-08-06 16:41 ` ✓ CI.KUnit: success for Introduce cold reset recovery method (rev14) Patchwork
2026-08-06 17:22 ` ✓ Xe.CI.BAT: " Patchwork
2026-08-07 5:53 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-08-13 13:44 ` Rodrigo Vivi [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=an3KNgWyfgxy_8Yd@intel.com \
--to=rodrigo.vivi@intel.com \
--cc=airlied@gmail.com \
--cc=andrealmeid@igalia.com \
--cc=anshuman.gupta@intel.com \
--cc=badal.nilawar@intel.com \
--cc=christian.koenig@amd.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=intel-xe@lists.freedesktop.org \
--cc=karthik.poosa@intel.com \
--cc=maarten.lankhorst@linux.intel.com \
--cc=mallesh.koujalagi@intel.com \
--cc=mripard@kernel.org \
--cc=raag.jadav@intel.com \
--cc=riana.tauro@intel.com \
--cc=simona.vetter@ffwll.ch \
--cc=sk.anirban@intel.com \
--cc=tzimmermann@suse.de \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.