From: Umesh Nerlige Ramappa <umesh.nerlige.ramappa@intel.com>
To: intel-xe@lists.freedesktop.org
Cc: daniele.ceraolospurio@intel.com, michal.wajdeczko@intel.com,
aravind.iddamsetty@intel.com, mallesh.koujalagi@intel.com,
alan.previn.teres.alexis@intel.com, julia.filipchuk@intel.com
Subject: [PATCH v5 4/7] drm/xe/guc: Cleanup error codes and handling for CT errors
Date: Fri, 18 Sep 2026 22:01:00 -0700 [thread overview]
Message-ID: <20260919050055.2092579-13-umesh.nerlige.ramappa@intel.com> (raw)
In-Reply-To: <20260919050055.2092579-9-umesh.nerlige.ramappa@intel.com>
Instead of always returning -EPROTO for most errors, use different error
codes based on the error type. -EPROTO is retained for the cases that
are genuine protocol violations by the GuC. The remaining cases now
report what actually went wrong.
receive_g2h() used to escalate to CT_DEAD + kick_reset() by matching the
two error codes that dequeue_one_g2h() could return on a fatal error.
Invert the check to simplify reset handling.
Signed-off-by: Umesh Nerlige Ramappa <umesh.nerlige.ramappa@intel.com>
Assisted-by: Claude:claude-opus-5
---
v2: (Michal)
- Sync order of errors in g2h_read and __guc_ct_send_locked
- For CT errors use EPIPE and for HXG errors use EPROTO
- Convert the non-fatal EPIPE to EPERM in the helper
v3: (Michal)
- s/EPIPE/EPERM/ in __guc_ct_send_locked
- Add ct_corrupted helper
- Add error documentation
- Remove "reset required" string from logs
v4: (Sashiko)
- s/EPIPE/EPERM/ in retry_failure and xe_guc_ct_send doc
---
drivers/gpu/drm/xe/xe_guc_ct.c | 194 ++++++++++++++++++++++++---------
1 file changed, 143 insertions(+), 51 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_guc_ct.c b/drivers/gpu/drm/xe/xe_guc_ct.c
index 5dbda0fe20f3..572a52773c56 100644
--- a/drivers/gpu/drm/xe/xe_guc_ct.c
+++ b/drivers/gpu/drm/xe/xe_guc_ct.c
@@ -37,6 +37,71 @@
#include "xe_sriov_vf.h"
#include "xe_trace_guc.h"
+/**
+ * DOC: GuC CT Error Codes
+ *
+ * Send path - xe_guc_ct_send(), xe_guc_ct_send_locked(), h2g_write()
+ * ------------------------------------------------------------------
+ *
+ * Channel-state rejections - the message was never queued, nothing is corrupt:
+ *
+ * ``-ENOTRECOVERABLE``: the device is wedged. No CT traffic can make progress
+ * until the device is recovered.
+ * ``-ENODEV``: the CT channel is disabled. Don't retry until it is enabled.
+ * ``-ECANCELED``: the CT channel is stopped or a GT recovery is pending, so the
+ * message was dropped. Often benign; retry after the reset/recovery completes.
+ * ``-EPERM``: the H2G CTB is flagged broken and stays unusable until the CT is
+ * restarted, which clears the flag.
+ *
+ * Internal flow control - never seen by callers, handled by the send helpers:
+ *
+ * ``-EBUSY``: not enough H2G room or G2H credits. Triggers a flush, a wait and
+ * a retry.
+ * ``-EAGAIN``: the message would wrap the H2G ring, so the tail was NOP-filled
+ * and reset to 0. The write is reissued immediately.
+ *
+ * Hard failures:
+ *
+ * ``-EPIPE``: H2G descriptor corruption found before writing - non-zero status,
+ * or head/tail out of range. Marks the CTB broken.
+ * ``-EDEADLK``: H2G credits never recovered after waiting ~1s. An async GT reset
+ * is requested before this is returned.
+ * ``-EPROTO``: the MMIO CT enable/disable handshake got an unexpected reply.
+ * ``-ENOMEM``: internal allocation failure (fence lookup under GFP_ATOMIC, or the
+ * G2H workqueue at init). The blocking path retries the allocation with
+ * GFP_KERNEL.
+ *
+ * Blocking send only (xe_guc_ct_send_recv()) - the H2G went out, the reply is the
+ * problem:
+ *
+ * ``-ETIME``: no G2H response within 1s, even after flushing the G2H worker.
+ * ``-EIO``: the GuC replied RESPONSE_FAILURE; error and hint codes are logged.
+ * ``-ELOOP``: the GuC kept replying NO_RESPONSE_RETRY until the request hit
+ * GUC_SEND_RETRY_LIMIT. (NO_RESPONSE_BUSY is not an error - the waiter re-arms.)
+ *
+ * Receive path - g2h_read(), parse_g2h_msg(), dequeue_one_g2h()
+ * -------------------------------------------------------------
+ *
+ * Channel-state rejections - the same four codes as the send path, checked against
+ * the G2H CTB. g2h_err_is_fatal() treats exactly these as benign; anything else
+ * declares the CT dead and kicks a reset:
+ *
+ * ``-ENOTRECOVERABLE``: device wedged.
+ * ``-ENODEV``: CT disabled.
+ * ``-ECANCELED``: CT stopped.
+ * ``-EPERM``: G2H CTB already flagged broken.
+ *
+ * Hard failures:
+ *
+ * ``-EPIPE``: G2H descriptor corruption - non-zero status, head/tail out of
+ * range, or a declared message length larger than the data actually available.
+ * ``-EPROTO``: HXG protocol violation. Either the G2H origin field is not the
+ * GuC, or a FAST_REQUEST H2G came back rejected. FAST_REQUEST fences aren't
+ * tracked, so there is no caller to report to and the failure is escalated to a
+ * reset instead.
+ * ``-EOPNOTSUPP``: G2H carried an HXG type this driver does not handle.
+ */
+
static void receive_g2h(struct xe_guc_ct *ct);
static void g2h_worker_func(struct work_struct *w);
static void safe_mode_worker_func(struct work_struct *w);
@@ -929,6 +994,23 @@ static bool vf_action_can_safely_fail(struct xe_device *xe, u32 action)
action == XE_GUC_ACTION_REGISTER_CONTEXT);
}
+static int ct_corrupted(struct xe_guc_ct *ct, struct guc_ctb *ctb,
+ u32 reason_code, const char *msg, ...)
+{
+ struct va_format vaf;
+ va_list va_args;
+ int err = -EPIPE;
+
+ va_start(va_args, msg);
+ vaf.fmt = msg;
+ vaf.va = &va_args;
+ xe_gt_err(ct_to_gt(ct), "GUC: CT: %pV", &vaf);
+ va_end(va_args);
+
+ ct_dead_capture(ct, ctb, reason_code);
+ return err;
+}
+
#define H2G_CT_HEADERS (GUC_CTB_HDR_LEN + 1) /* one DW CTB header and one DW HxG header */
static int h2g_write(struct xe_guc_ct *ct, const u32 *action, u32 len,
@@ -953,23 +1035,22 @@ static int h2g_write(struct xe_guc_ct *ct, const u32 *action, u32 len,
u32 desc_status;
desc_status = desc_read(xe, h2g, status);
- if (desc_status) {
- xe_gt_err(gt, "CT write: non-zero status: %u\n", desc_status);
- goto corrupted;
- }
+ if (desc_status)
+ return ct_corrupted(ct, &ct->ctbs.h2g, ct_id(H2G_WRITE),
+ "write: non-zero status: %u\n", desc_status);
if (tail > h2g->info.size) {
desc_write(xe, h2g, status, desc_status | GUC_CTB_STATUS_OVERFLOW);
- xe_gt_err(gt, "CT write: tail out of range: %u vs %u\n",
- tail, h2g->info.size);
- goto corrupted;
+ return ct_corrupted(ct, &ct->ctbs.h2g, ct_id(H2G_WRITE),
+ "write: tail out of range: %u vs %u\n",
+ tail, h2g->info.size);
}
if (desc_head >= h2g->info.size) {
desc_write(xe, h2g, status, desc_status | GUC_CTB_STATUS_OVERFLOW);
- xe_gt_err(gt, "CT write: invalid head offset %u >= %u)\n",
- desc_head, h2g->info.size);
- goto corrupted;
+ return ct_corrupted(ct, &ct->ctbs.h2g, ct_id(H2G_WRITE),
+ "write: invalid head offset %u >= %u)\n",
+ desc_head, h2g->info.size);
}
}
@@ -1034,10 +1115,6 @@ static int h2g_write(struct xe_guc_ct *ct, const u32 *action, u32 len,
desc_read(xe, h2g, head), h2g->info.tail);
return 0;
-
-corrupted:
- CT_DEAD(ct, &ct->ctbs.h2g, H2G_WRITE);
- return -EPIPE;
}
static int __guc_ct_send_locked(struct xe_guc_ct *ct, const u32 *action,
@@ -1060,11 +1137,6 @@ static int __guc_ct_send_locked(struct xe_guc_ct *ct, const u32 *action,
goto out;
}
- if (unlikely(ct->ctbs.h2g.info.broken)) {
- ret = -EPIPE;
- goto out;
- }
-
if (ct->state == XE_GUC_CT_STATE_DISABLED) {
ret = -ENODEV;
goto out;
@@ -1075,6 +1147,11 @@ static int __guc_ct_send_locked(struct xe_guc_ct *ct, const u32 *action,
goto out;
}
+ if (unlikely(ct->ctbs.h2g.info.broken)) {
+ ret = -EPERM;
+ goto out;
+ }
+
xe_gt_assert(gt, xe_guc_ct_enabled(ct));
if (g2h_fence) {
@@ -1213,7 +1290,7 @@ static int guc_ct_send_locked(struct xe_guc_ct *ct, const u32 *action, u32 len,
return ret;
broken:
- xe_gt_err(gt, "No forward process on H2G, reset required\n");
+ xe_gt_err(gt, "No forward progress on H2G\n");
CT_DEAD(ct, &ct->ctbs.h2g, DEADLOCK);
return -EDEADLK;
@@ -1246,7 +1323,7 @@ static int guc_ct_send(struct xe_guc_ct *ct, const u32 *action, u32 len,
*
* * -ENOTRECOVERABLE: the xe device is wedged. Stop submitting new GuC work; the
* request cannot make progress until the device is recovered.
- * * -EPIPE: the H2G CTB is marked broken. The channel stays unusable until the
+ * * -EPERM: the H2G CTB is marked broken. The channel stays unusable until the
* CT is restarted, which clears the broken flag.
* * -ENODEV: the CT channel is disabled, messages not expected in this state.
* Don't retry until it is enabled again.
@@ -1358,7 +1435,7 @@ int xe_guc_ct_send_g2h_handler(struct xe_guc_ct *ct, const u32 *action, u32 len)
*/
static bool retry_failure(struct xe_guc_ct *ct, int ret)
{
- if (!(ret == -EDEADLK || ret == -EPIPE || ret == -ENODEV))
+ if (!(ret == -EDEADLK || ret == -EPERM || ret == -ENODEV))
return false;
#define ct_alive(ct) \
@@ -1667,8 +1744,7 @@ static int parse_g2h_msg(struct xe_guc_ct *ct, u32 *msg, u32 len)
origin = FIELD_GET(GUC_HXG_MSG_0_ORIGIN, hxg[0]);
if (unlikely(origin != GUC_HXG_ORIGIN_GUC)) {
- xe_gt_err(gt, "G2H channel broken on read, origin=%u, reset required\n",
- origin);
+ xe_gt_err(gt, "Invalid G2H origin=%u\n", origin);
CT_DEAD(ct, &ct->ctbs.g2h, PARSE_G2H_ORIGIN);
return -EPROTO;
@@ -1686,8 +1762,7 @@ static int parse_g2h_msg(struct xe_guc_ct *ct, u32 *msg, u32 len)
ret = parse_g2h_response(ct, msg, len);
break;
default:
- xe_gt_err(gt, "G2H channel broken on read, type=%u, reset required\n",
- type);
+ xe_gt_err(gt, "Unexpected G2H message type=%u\n", type);
CT_DEAD(ct, &ct->ctbs.g2h, PARSE_G2H_TYPE);
ret = -EOPNOTSUPP;
@@ -1808,6 +1883,7 @@ static int g2h_read(struct xe_guc_ct *ct, u32 *msg, bool fast_path)
xe_gt_assert(gt, xe_guc_ct_initialized(ct));
lockdep_assert_held(&ct->fast_lock);
+ /* Keep in sync with g2h_err_is_fatal() */
if (xe_device_wedged(xe))
return -ENOTRECOVERABLE;
@@ -1818,7 +1894,7 @@ static int g2h_read(struct xe_guc_ct *ct, u32 *msg, bool fast_path)
return -ECANCELED;
if (g2h->info.broken)
- return -EPIPE;
+ return -EPERM;
xe_gt_assert(gt, xe_guc_ct_enabled(ct));
@@ -1834,10 +1910,9 @@ static int g2h_read(struct xe_guc_ct *ct, u32 *msg, bool fast_path)
desc_status &= ~GUC_CTB_STATUS_DISABLED;
}
- if (desc_status) {
- xe_gt_err(gt, "CT read: non-zero status: %u\n", desc_status);
- goto corrupted;
- }
+ if (desc_status)
+ return ct_corrupted(ct, &ct->ctbs.g2h, ct_id(G2H_READ),
+ "read: non-zero status: %u\n", desc_status);
}
if (IS_ENABLED(CONFIG_DRM_XE_DEBUG)) {
@@ -1858,24 +1933,24 @@ static int g2h_read(struct xe_guc_ct *ct, u32 *msg, bool fast_path)
*
if (g2h->info.head != desc_head) {
desc_write(xe, g2h, status, desc_status | GUC_CTB_STATUS_MISMATCH);
- xe_gt_err(gt, "CT read: head was modified %u != %u\n",
- desc_head, g2h->info.head);
- goto corrupted;
+ return ct_corrupted(ct, &ct->ctbs.g2h, ct_id(G2H_READ),
+ "read: head was modified %u != %u\n",
+ desc_head, g2h->info.head);
}
*/
if (g2h->info.head > g2h->info.size) {
desc_write(xe, g2h, status, desc_status | GUC_CTB_STATUS_OVERFLOW);
- xe_gt_err(gt, "CT read: head out of range: %u vs %u\n",
- g2h->info.head, g2h->info.size);
- goto corrupted;
+ return ct_corrupted(ct, &ct->ctbs.g2h, ct_id(G2H_READ),
+ "read: head out of range: %u vs %u\n",
+ g2h->info.head, g2h->info.size);
}
if (desc_tail >= g2h->info.size) {
desc_write(xe, g2h, status, desc_status | GUC_CTB_STATUS_OVERFLOW);
- xe_gt_err(gt, "CT read: invalid tail offset %u >= %u)\n",
- desc_tail, g2h->info.size);
- goto corrupted;
+ return ct_corrupted(ct, &ct->ctbs.g2h, ct_id(G2H_READ),
+ "read: invalid tail offset %u >= %u)\n",
+ desc_tail, g2h->info.size);
}
}
@@ -1892,11 +1967,10 @@ static int g2h_read(struct xe_guc_ct *ct, u32 *msg, bool fast_path)
xe_map_memcpy_from(xe, msg, &g2h->cmds, sizeof(u32) * g2h->info.head,
sizeof(u32));
len = FIELD_GET(GUC_CTB_MSG_0_NUM_DWORDS, msg[0]) + GUC_CTB_MSG_MIN_LEN;
- if (len > avail) {
- xe_gt_err(gt, "G2H channel broken on read, avail=%d, len=%d, reset required\n",
- avail, len);
- goto corrupted;
- }
+ if (len > avail)
+ return ct_corrupted(ct, &ct->ctbs.g2h, ct_id(G2H_READ),
+ "G2H channel broken on read, avail=%d, len=%d\n",
+ avail, len);
head = (g2h->info.head + 1) % g2h->info.size;
avail = len - 1;
@@ -1942,10 +2016,6 @@ static int g2h_read(struct xe_guc_ct *ct, u32 *msg, bool fast_path)
action, len, g2h->info.head, tail);
return len;
-
-corrupted:
- CT_DEAD(ct, &ct->ctbs.g2h, G2H_READ);
- return -EPROTO;
}
static void g2h_fast_path(struct xe_guc_ct *ct, u32 *msg, u32 len)
@@ -2037,6 +2107,27 @@ static int dequeue_one_g2h(struct xe_guc_ct *ct)
return 1;
}
+/*
+ * Errors reported by dequeue_one_g2h() come in two flavours: either the channel
+ * is simply not available right now, which is expected and handled gracefully,
+ * or the channel state or the message itself is inconsistent, in which case the
+ * only way forward is to declare the CT dead and reset the GuC. Since the
+ * former is a short and well known list, check against that and treat anything
+ * else as fatal, so that new error codes don't silently escape the escalation.
+ */
+static bool g2h_err_is_fatal(int err)
+{
+ switch (err) {
+ case -ENOTRECOVERABLE: /* device wedged */
+ case -ENODEV: /* CT disabled */
+ case -ECANCELED: /* CT stopped */
+ case -EPERM: /* CT already declared broken */
+ return false;
+ default:
+ return true;
+ }
+}
+
static void receive_g2h(struct xe_guc_ct *ct)
{
bool ongoing;
@@ -2074,8 +2165,9 @@ static void receive_g2h(struct xe_guc_ct *ct)
ret = dequeue_one_g2h(ct);
mutex_unlock(&ct->lock);
- if (unlikely(ret == -EPROTO || ret == -EOPNOTSUPP)) {
- xe_gt_err(ct_to_gt(ct), "CT dequeue failed: %d\n", ret);
+ if (unlikely(ret < 0 && g2h_err_is_fatal(ret))) {
+ xe_gt_err(ct_to_gt(ct), "CT dequeue failed, forcing GT reset(%pe)\n",
+ ERR_PTR(ret));
CT_DEAD(ct, NULL, G2H_RECV);
kick_reset(ct);
}
--
2.55.0
next prev parent reply other threads:[~2026-09-19 5:01 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-19 5:00 [PATCH v5 0/7] Use SIG_ID logs for GuC component Umesh Nerlige Ramappa
2026-09-19 5:00 ` [PATCH v5 1/7] drm/xe/guc: Use different error codes for GuC load errors Umesh Nerlige Ramappa
2026-09-19 5:00 ` [PATCH v5 2/7] drm/xe/guc: Handle CRASH and EXCEPTION G2H with separate helpers Umesh Nerlige Ramappa
2026-09-19 5:09 ` sashiko-bot
2026-09-19 5:00 ` [PATCH v5 3/7] drm/xe/guc: Make ct_dead_capture available on non-debug config Umesh Nerlige Ramappa
2026-09-19 5:01 ` Umesh Nerlige Ramappa [this message]
2026-09-19 5:14 ` [PATCH v5 4/7] drm/xe/guc: Cleanup error codes and handling for CT errors sashiko-bot
2026-09-19 5:01 ` [PATCH v5 5/7] drm/xe/uc: Report DMA failure using SIGID Umesh Nerlige Ramappa
2026-09-19 5:01 ` [PATCH v5 6/7] drm/xe/guc: Report major GuC failures " Umesh Nerlige Ramappa
2026-09-19 5:01 ` [PATCH v5 7/7] drm/xe/guc: Report errors that cause a CT shutdown " Umesh Nerlige Ramappa
2026-09-19 5:07 ` ✗ CI.checkpatch: warning for Use SIG_ID logs for GuC component (rev4) Patchwork
2026-09-19 5:09 ` ✓ CI.KUnit: success " Patchwork
2026-09-19 5:48 ` ✗ Xe.CI.BAT: failure " Patchwork
2026-09-19 8:33 ` ✗ Xe.CI.FULL: " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260919050055.2092579-13-umesh.nerlige.ramappa@intel.com \
--to=umesh.nerlige.ramappa@intel.com \
--cc=alan.previn.teres.alexis@intel.com \
--cc=aravind.iddamsetty@intel.com \
--cc=daniele.ceraolospurio@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=julia.filipchuk@intel.com \
--cc=mallesh.koujalagi@intel.com \
--cc=michal.wajdeczko@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.