Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Arvind Yadav" <arvind.yadav@intel.com>
Cc: intel-xe@lists.freedesktop.org, dri-devel@lists.freedesktop.org
Subject: Re: [PATCH v2 11/15] drm/xe: Isolate a wedged device before notifying userspace
Date: Tue, 22 Sep 2026 10:31:26 +0000	[thread overview]
Message-ID: <20260922103127.3FDD61F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20260922101721.1583542-12-arvind.yadav@intel.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] drm/xe: TOCTOU race in xe_display_register leaves display active without IRQs
- [High] drm/xe: double suspend of display hardware due to wedge during PM
--

commit f1762c4030fdcf19e70d9970658fdb7e79d5019c
Author: Arvind Yadav <arvind.yadav@intel.com>

drm/xe: Isolate a wedged device before notifying userspace

A permanently wedged device must stop interrupt and DMA activity before
userspace is notified that recovery is required. The commit runs the
sleepable isolation steps from the wedge worker.

> diff --git a/drivers/gpu/drm/xe/display/xe_display.c b/drivers/gpu/drm/xe/display/xe_display.c
> --- a/drivers/gpu/drm/xe/display/xe_display.c
> +++ b/drivers/gpu/drm/xe/display/xe_display.c
> @@ -150,21 +150,27 @@ void xe_display_register(struct xe_device *xe)
>  {
>  	struct intel_display *display = xe->display;
>  
> -	if (!xe->info.probe_display)
> +	if (!xe->info.probe_display || xe->display_registered)
>  		return;
>  
>  	intel_display_driver_register(display);
>  	intel_display_driver_runtime_pm_enable(display);
> +
> +	xe->display_registered = true;

[Severity: High]
Could the async wedge worker race with the display probe thread here?

The display is registered without any locking to prevent concurrent
isolation by the async wedge worker.

>  }
>  
>  void xe_display_unregister(struct xe_device *xe)
>  {
>  	struct intel_display *display = xe->display;
>  
> -	if (!xe->info.probe_display)
> +	if (!xe->info.probe_display || !xe->display_registered)
>  		return;

[Severity: High]
The wedge worker reads xe->display_registered unprotected and returns early
if registration is still in progress in the probe thread.

[ ... ]
> @@ -172,7 +178,8 @@ void xe_display_shutdown(struct xe_device *xe)
>  {
>  	struct intel_display *display = xe->display;
>  
> -	if (!xe->info.probe_display)
> +	if (!xe->info.probe_display || !xe->display_registered ||
> +	    xe->display_shutdown)
>  		return;
>  
>  	intel_display_driver_shutdown(display);
> @@ -180,6 +187,8 @@ void xe_display_shutdown(struct xe_device *xe)
>  	intel_opregion_suspend(display, PCI_D3cold);
>  
>  	intel_dmc_suspend(display);
> +
> +	xe->display_shutdown = true;
>  }

[Severity: High]
When a synchronous wedge worker calls xe_device_wedged_isolate() during
suspend, it executes xe_display_shutdown().

Since xe->display_shutdown is false, it proceeds to call
intel_dmc_suspend(), intel_opregion_suspend(), and
intel_display_driver_shutdown() a second time. Could this cause a wakeref
underflow and unbalanced runtime PM state for the display hardware?

[ ... ]
> diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c
> --- a/drivers/gpu/drm/xe/xe_device.c
> +++ b/drivers/gpu/drm/xe/xe_device.c
[ ... ]
> +static void xe_device_wedged_isolate(struct xe_device *xe)
> +{
[ ... ]
> +	/* Shut down display hardware before stopping device access. */
> +	xe_display_shutdown(xe);
> +	xe_display_unregister(xe);
> +
> +	/* Stop interrupt and DMA activity before changing mappings. */
> +	xe_irq_suspend(xe);
> +	pci_clear_master(to_pci_dev(xe->drm.dev));

[Severity: High]
If the device wedges asynchronously during device probe, could the wedge
worker miss the display unregistration because xe->display_registered is not
yet true, but then continue to disable hardware access?

If the probe thread then proceeds to complete xe_display_register(), the
display remains registered and exposed to userspace, but the device has its
interrupts and DMA (PCI master) permanently disabled. This could cause
userspace modesets and page flips to time out and hang since VBLANK
interrupts will never fire.

[ ... ]
> diff --git a/drivers/gpu/drm/xe/xe_pci.c b/drivers/gpu/drm/xe/xe_pci.c
> --- a/drivers/gpu/drm/xe/xe_pci.c
> +++ b/drivers/gpu/drm/xe/xe_pci.c
> @@ -1325,19 +1325,38 @@ static int xe_pci_suspend(struct device *dev)
>  {
[ ... ]
> +	} else {
> +		err = xe_pm_suspend(xe);
> +		xe_device_io_put(io_idx);
> +
> +		if (xe_device_wedged(xe)) {
> +			wait_for_completion(&xe->wedged.prepared);
> +			flush_work(&xe->wedged.work);
> +		} else if (err) {
> +			return err;
> +		}
> +	}

[Severity: High]
Can a wedge during suspend cause a double suspend of the display hardware?

If a wedge event occurs during xe_pm_suspend(), the display has already been
suspended via xe_display_pm_suspend(). Since xe_pm_suspend() does not set
the xe->display_shutdown flag, this flush_work() will execute the wedge
worker synchronously, resulting in xe_display_shutdown() suspending the
hardware again.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260922101721.1583542-1-arvind.yadav@intel.com?part=11

  reply	other threads:[~2026-09-22 10:31 UTC|newest]

Thread overview: 23+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 10:16 [PATCH v2 00/15] drm/xe: Isolate wedged devices from hardware access Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 01/15] drm/xe/irq: Always free requested IRQs on uninstall Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 02/15] drm/drv: Export drm_dev_srcu_synchronize() Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 03/15] drm/xe: Separate AER reset state from device wedging Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 04/15] drm/xe: Protect device I/O with DRM device SRCU Arvind Yadav
2026-09-22 10:28   ` sashiko-bot
2026-09-22 10:16 ` [PATCH v2 05/15] drm/xe: Drop queued page faults when device I/O is blocked Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 06/15] drm/xe: Stop VM work " Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 07/15] drm/xe: Send wedged notification from a worker Arvind Yadav
2026-09-22 10:27   ` sashiko-bot
2026-09-22 10:16 ` [PATCH v2 08/15] drm/xe: Reuse one dummy page per BO after wedge Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 09/15] drm/xe: Invalidate existing VRAM mappings on wedge Arvind Yadav
2026-09-22 10:30   ` sashiko-bot
2026-09-22 10:16 ` [PATCH v2 10/15] drm/xe/irq: Protect IRQ state during wedge isolation Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 11/15] drm/xe: Isolate a wedged device before notifying userspace Arvind Yadav
2026-09-22 10:31   ` sashiko-bot [this message]
2026-09-22 10:16 ` [PATCH v2 12/15] drm/xe/ttm: Reject VRAM allocations on wedged devices Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 13/15] drm/xe/guc: Skip timeout recovery on a wedged device Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 14/15] drm/xe: Skip PM notifier preparation when device I/O is blocked Arvind Yadav
2026-09-22 10:17 ` [PATCH v2 15/15] drm/xe: Block BO VM access when device I/O is unavailable Arvind Yadav
2026-09-22 10:27 ` ✓ CI.KUnit: success for drm/xe: Isolate wedged devices from hardware access (rev2) Patchwork
2026-09-22 12:26 ` ✗ Xe.CI.BAT: failure " Patchwork
2026-09-22 20:50 ` ✗ Xe.CI.FULL: " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260922103127.3FDD61F000FF@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=arvind.yadav@intel.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox