dri-devel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Arvind Yadav" <arvind.yadav@intel.com>
Cc: dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v2 07/15] drm/xe: Send wedged notification from a worker
Date: Tue, 22 Sep 2026 10:27:20 +0000	[thread overview]
Message-ID: <20260922102721.5870D1F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20260922101721.1583542-8-arvind.yadav@intel.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] xe: TOCTOU race in xe_device_declare_wedged() overwrites recovery method
--

commit 8a75aea67f604c5a996849e643c64617844ddad3
Author: Arvind Yadav <arvind.yadav@intel.com>
Subject: drm/xe: Send wedged notification from a worker

This commit moves the wedged event to a sleepable device worker to allow
isolation work to complete before userspace is notified. It updates the
wedge bookkeeping to rescan the GT state on later declarations and tracks
the last reported method to avoid lost updates.

> diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c
> --- a/drivers/gpu/drm/xe/xe_device.c
> +++ b/drivers/gpu/drm/xe/xe_device.c

[ ... ]

> @@ -1435,14 +1482,22 @@ void xe_device_declare_wedged(struct xe_device *xe)
>  		 */
>  		if (xe->wedged.mode == XE_WEDGED_MODE_UPON_ANY_HANG_NO_RESET)
>  			xe_device_set_wedged_method(xe, DRM_WEDGE_RECOVERY_NONE);
>  	}
> +
> +	/* Re-scan GT submission state on every declaration. */
> +	for_each_gt(gt, xe, id)
> +		xe_gt_declare_wedged(gt);
> +
> +	/* If no wedge recovery method is set, use default */
> +	if (!READ_ONCE(xe->wedged.method))
> +		xe_device_set_wedged_method(xe, DRM_WEDGE_RECOVERY_REBIND |
> +					    DRM_WEDGE_RECOVERY_BUS_RESET);

[Severity: Medium]
Is there a Time-of-Check to Time-of-Use (TOCTOU) race condition here in
xe_device_declare_wedged()?

If multiple threads invoke xe_device_declare_wedged() concurrently (for
example, during simultaneous GT hangs), could Thread A evaluate the
!READ_ONCE(xe->wedged.method) condition as true, get preempted, and then
allow Thread B to successfully set a specific method like
DRM_WEDGE_RECOVERY_NONE?

If Thread A then resumes, it seems it would unconditionally call
xe_device_set_wedged_method() with the default REBIND | BUS_RESET flags,
silently overwriting Thread B's legitimately set method.

Does this check-and-set sequence need atomic protection, such as cmpxchg,
to prevent concurrent threads from overwriting a newly set recovery method
with the default fallback?

> +
> +	if (!atomic_read(&xe->wedged.stopping) &&
> +	    (first ||
> +	     READ_ONCE(xe->wedged.method) !=
> +	     READ_ONCE(xe->wedged.reported_method)))
> +		queue_work(xe->unordered_wq, &xe->wedged.work);
>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260922101721.1583542-1-arvind.yadav@intel.com?part=7

  reply	other threads:[~2026-09-22 10:27 UTC|newest]

Thread overview: 20+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 10:16 [PATCH v2 00/15] drm/xe: Isolate wedged devices from hardware access Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 01/15] drm/xe/irq: Always free requested IRQs on uninstall Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 02/15] drm/drv: Export drm_dev_srcu_synchronize() Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 03/15] drm/xe: Separate AER reset state from device wedging Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 04/15] drm/xe: Protect device I/O with DRM device SRCU Arvind Yadav
2026-09-22 10:28   ` sashiko-bot
2026-09-22 10:16 ` [PATCH v2 05/15] drm/xe: Drop queued page faults when device I/O is blocked Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 06/15] drm/xe: Stop VM work " Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 07/15] drm/xe: Send wedged notification from a worker Arvind Yadav
2026-09-22 10:27   ` sashiko-bot [this message]
2026-09-22 10:16 ` [PATCH v2 08/15] drm/xe: Reuse one dummy page per BO after wedge Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 09/15] drm/xe: Invalidate existing VRAM mappings on wedge Arvind Yadav
2026-09-22 10:30   ` sashiko-bot
2026-09-22 10:16 ` [PATCH v2 10/15] drm/xe/irq: Protect IRQ state during wedge isolation Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 11/15] drm/xe: Isolate a wedged device before notifying userspace Arvind Yadav
2026-09-22 10:31   ` sashiko-bot
2026-09-22 10:16 ` [PATCH v2 12/15] drm/xe/ttm: Reject VRAM allocations on wedged devices Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 13/15] drm/xe/guc: Skip timeout recovery on a wedged device Arvind Yadav
2026-09-22 10:16 ` [PATCH v2 14/15] drm/xe: Skip PM notifier preparation when device I/O is blocked Arvind Yadav
2026-09-22 10:17 ` [PATCH v2 15/15] drm/xe: Block BO VM access when device I/O is unavailable Arvind Yadav

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260922102721.5870D1F000FF@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=arvind.yadav@intel.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox