From: sashiko-bot@kernel.org
To: "Umesh Nerlige Ramappa" <umesh.nerlige.ramappa@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v4 4/7] drm/xe/guc: Cleanup error codes and handling for CT errors
Date: Fri, 18 Sep 2026 00:17:27 +0000 [thread overview]
Message-ID: <20260918001727.E63451F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20260917235923.1521112-13-umesh.nerlige.ramappa@intel.com>
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] drm/xe/guc: retry_failure() not updated for -EPERM breaks no_fail requests
- [Low] drm/xe/guc: build failure due to undeclared CT_DEAD_ reason codes
- [Low] drm/xe/guc: kernel-doc mismatch for xe_guc_ct_send return value
--
commit fcb132fd1e8cabdb5fd371a7a87d1bb78381dd54
Author: Umesh Nerlige Ramappa <umesh.nerlige.ramappa@intel.com>
Subject: drm/xe/guc: Cleanup error codes and handling for CT errors
This commit updates GuC CT error handling to return more specific error
codes instead of uniformly returning -EPROTO. It retains -EPROTO for
genuine protocol violations while adjusting others to reflect the actual
failure, and simplifies the reset handling logic.
> diff --git a/drivers/gpu/drm/xe/xe_guc_ct.c b/drivers/gpu/drm/xe/xe_guc_ct.c
> index 3353515aa8747..31ecddab30578 100644
> --- a/drivers/gpu/drm/xe/xe_guc_ct.c
> +++ b/drivers/gpu/drm/xe/xe_guc_ct.c
[ ... ]
> @@ -953,23 +1035,22 @@ static int h2g_write(struct xe_guc_ct *ct, const u32 *action, u32 len,
> u32 desc_status;
>
> desc_status = desc_read(xe, h2g, status);
> - if (desc_status) {
> - xe_gt_err(gt, "CT write: non-zero status: %u\n", desc_status);
> - goto corrupted;
> - }
> + if (desc_status)
> + return ct_corrupted(ct, &ct->ctbs.h2g, ct_id(H2G_WRITE),
> + "write: non-zero status: %u\n", desc_status);
[Severity: Low]
Does this code break the build when CONFIG_DRM_XE_DEBUG is disabled?
The macro ct_id() evaluates to CT_DEAD_##reason_code. If those enum
values are conditionally defined globally (such as inside an
#if IS_ENABLED(CONFIG_DRM_XE_DEBUG) block), using them here in h2g_write()
outside of any #if guards will cause undeclared identifier compilation
errors on non-debug builds.
[ ... ]
> @@ -1060,11 +1137,6 @@ static int __guc_ct_send_locked(struct xe_guc_ct *ct, const u32 *action,
> goto out;
> }
>
> - if (unlikely(ct->ctbs.h2g.info.broken)) {
> - ret = -EPIPE;
> - goto out;
> - }
> -
> if (ct->state == XE_GUC_CT_STATE_DISABLED) {
> ret = -ENODEV;
> goto out;
> @@ -1075,6 +1147,11 @@ static int __guc_ct_send_locked(struct xe_guc_ct *ct, const u32 *action,
> goto out;
> }
>
> + if (unlikely(ct->ctbs.h2g.info.broken)) {
> + ret = -EPERM;
> + goto out;
> + }
> +
[Severity: High]
Does this return code change break the retry logic for no-fail requests?
When the H2G CTB is broken, __guc_ct_send_locked() now returns -EPERM
instead of the legacy -EPIPE. However, retry_failure() still only checks
for -EPIPE:
if (!(ret == -EDEADLK || ret == -EPIPE || ret == -ENODEV))
return false;
Without checking for -EPERM, retry_failure() will return false, causing
no-fail messages (such as critical TLB invalidations) to fail immediately
instead of waiting for a GT reset and retrying.
[Severity: Low]
Also, does the kernel-doc for xe_guc_ct_send() need an update to reflect
this change?
The documentation for xe_guc_ct_send() still lists -EPIPE for a broken
CTB:
* -EPIPE: the H2G CTB is marked broken. The channel stays unusable until the
* CT is restarted, which clears the broken flag.
Since this patch changes the return value for a broken CTB to -EPERM,
this documentation might need to be updated to match.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260917235923.1521112-9-umesh.nerlige.ramappa@intel.com?part=4
next prev parent reply other threads:[~2026-09-18 0:17 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 23:59 [PATCH v4 0/7] Use SIG_ID logs for GuC component Umesh Nerlige Ramappa
2026-09-17 23:59 ` [PATCH v4 1/7] drm/xe/guc: Use different error codes for GuC load errors Umesh Nerlige Ramappa
2026-09-17 23:59 ` [PATCH v4 2/7] drm/xe/guc: Handle CRASH and EXCEPTION G2H with separate helpers Umesh Nerlige Ramappa
2026-09-18 0:08 ` sashiko-bot
2026-09-17 23:59 ` [PATCH v4 3/7] drm/xe/guc: Make ct_dead_capture available on non-debug config Umesh Nerlige Ramappa
2026-09-18 0:06 ` sashiko-bot
2026-09-17 23:59 ` [PATCH v4 4/7] drm/xe/guc: Cleanup error codes and handling for CT errors Umesh Nerlige Ramappa
2026-09-18 0:17 ` sashiko-bot [this message]
2026-09-17 23:59 ` [PATCH v4 5/7] drm/xe/uc: Report DMA failure using SIGID Umesh Nerlige Ramappa
2026-09-18 0:07 ` sashiko-bot
2026-09-17 23:59 ` [PATCH v4 6/7] drm/xe/guc: Report major GuC failures " Umesh Nerlige Ramappa
2026-09-17 23:59 ` [PATCH v4 7/7] drm/xe/guc: Report errors that cause a CT shutdown " Umesh Nerlige Ramappa
2026-09-18 0:08 ` sashiko-bot
2026-09-18 0:58 ` ✗ CI.KUnit: failure for Use SIG_ID logs for GuC component (rev3) Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260918001727.E63451F000FF@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=intel-xe@lists.freedesktop.org \
--cc=sashiko-reviews@lists.linux.dev \
--cc=umesh.nerlige.ramappa@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox