From: sashiko-bot@kernel.org
To: "Raag Jadav" <raag.jadav@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v1] drm/xe/guc: Allow GuC CT for wedged device
Date: Fri, 21 Aug 2026 12:42:54 +0000 [thread overview]
Message-ID: <20260821124255.5C09C1F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260821122114.567725-1-raag.jadav@intel.com>
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] Removing the `xe_device_wedged` check from `__guc_ct_send_locked` causes a full 1-second delay in `xe_device_declare_wedged` when a GuC hang occurs.
- [Low] The kernel-doc comment for `xe_guc_ct_send()` is stale and incorrectly claims that non-blocking send helpers return `-ENOTRECOVERABLE` when the device is wedged.
--
--- Patch [1]: [PATCH v1] drm/xe/guc: Allow GuC CT for wedged device ---
commit baa99b44631241d9eb913ee6996cb7221b0a82f2
Author: Raag Jadav <raag.jadav@intel.com>
drm/xe/guc: Allow GuC CT for wedged device
Commit 50fa9acac26f ("drm/xe/guc: distinguish wedged from recoverable
cancellation") introduced distinguishable error codes for g2h failure
cases, but also blocked GuC CT for wedged device. This is problematic
in cases where we want to prevent user from accessing the device but
also keep GuC CT functioning on temporarily wedged device. First user
of such requirement is PCIe FLR handling where we require uC firmware
loading while the device is temporarily wedged.
Fixes: 50fa9acac26f ("drm/xe/guc: distinguish wedged from recoverable cancellation")
Signed-off-by: Raag Jadav <raag.jadav@intel.com>
> diff --git a/drivers/gpu/drm/xe/xe_guc_ct.c b/drivers/gpu/drm/xe/xe_guc_ct.c
> index fe70c0fd85c52..541c74180bcaf 100644
> --- a/drivers/gpu/drm/xe/xe_guc_ct.c
> +++ b/drivers/gpu/drm/xe/xe_guc_ct.c
> @@ -1065,11 +1065,6 @@ static int __guc_ct_send_locked(struct xe_guc_ct *ct, const u32 *action,
> xe_gt_assert(gt, g2h_len || !num_g2h);
> lockdep_assert_held(&ct->lock);
>
> - if (xe_device_wedged(ct_to_xe(ct))) {
> - ret = -ENOTRECOVERABLE;
> - goto out;
> - }
> -
[Severity: Medium]
Will removing this check cause a 1-second delay during device wedging?
If guc_ct_send_locked() is executing when a GuC hang occurs and the CTB is full,
it enters a retry loop holding ct->lock. If xe_device_declare_wedged() executes
concurrently, it attempts to call xe_guc_ct_stop() and blocks waiting for
ct->lock.
Without this xe_device_wedged() check, the retry loop relies entirely on the
ct->state == XE_GUC_CT_STATE_STOPPED condition to abort early. However, because
the wedging thread is blocked on ct->lock, the state cannot be updated. This
causes the loop in guc_ct_send_locked() to retry until the 1-second timeout
expires.
[Severity: Low]
Should the kernel-doc for xe_guc_ct_send() be updated to reflect this change?
The documentation currently claims that non-blocking send helpers return
-ENOTRECOVERABLE when the device is wedged:
* * -ENOTRECOVERABLE: the xe device is wedged. Stop submitting new GuC work; the
* request cannot make progress until the device is recovered.
Since the wedged check was removed from the send path, this error code is
no longer returned in that scenario.
> if (unlikely(ct->ctbs.h2g.info.broken)) {
> ret = -EPIPE;
> goto out;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260821122114.567725-1-raag.jadav@intel.com?part=1
next prev parent reply other threads:[~2026-08-21 12:42 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-21 12:21 [PATCH v1] drm/xe/guc: Allow GuC CT for wedged device Raag Jadav
2026-08-21 12:28 ` ✓ CI.KUnit: success for " Patchwork
2026-08-21 12:42 ` sashiko-bot [this message]
2026-08-21 13:31 ` ✓ Xe.CI.BAT: " Patchwork
2026-08-21 17:27 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-08-25 9:39 ` [PATCH v1] " Laguna, Lukasz
2026-08-25 10:04 ` Raag Jadav
2026-08-25 10:53 ` Laguna, Lukasz
2026-08-25 12:49 ` Raag Jadav
2026-08-25 13:40 ` Laguna, Lukasz
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260821124255.5C09C1F000E9@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=intel-xe@lists.freedesktop.org \
--cc=raag.jadav@intel.com \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.