From: Michal Wajdeczko <michal.wajdeczko@intel.com>
To: Umesh Nerlige Ramappa <umesh.nerlige.ramappa@intel.com>,
<intel-xe@lists.freedesktop.org>,
Rodrigo Vivi <rodrigo.vivi@intel.com>
Cc: <daniele.ceraolospurio@intel.com>, <aravind.iddamsetty@intel.com>,
<mallesh.koujalagi@intel.com>,
<alan.previn.teres.alexis@intel.com>, <julia.filipchuk@intel.com>
Subject: Re: [PATCH v3 5/5] drm/xe/guc: Report errors that cause a CT shutdown using SIGID
Date: Tue, 15 Sep 2026 12:24:34 +0200 [thread overview]
Message-ID: <7e5691ad-d3ed-48e8-90b0-81ddf737a6f9@intel.com> (raw)
In-Reply-To: <20260903233958.475162-12-umesh.nerlige.ramappa@intel.com>
On 9/4/2026 1:40 AM, Umesh Nerlige Ramappa wrote:
> From: Daniele Ceraolo Spurio <daniele.ceraolospurio@intel.com>
>
> Convert any errors that can cause the CT to be declared as dead to
> use the xe_log_err() helper. Errors that are escalated to the callers
> are left for the caller to report with SIGID if needed.
> While at it, update some of the error messages to make what went wrong
> clearer.
>
> v2:
> - use different error codes and better messages (Michal)
> v3:
> - Drop GuC from log messages (Michal)
> - Clean up log messages
move under ---
>
> Signed-off-by: Daniele Ceraolo Spurio <daniele.ceraolospurio@intel.com>
> Cc: Michal Wajdeczko <michal.wajdeczko@intel.com>
> Cc: Aravind Iddamsetty <aravind.iddamsetty@intel.com>
> Cc: Mallesh Koujalagi <mallesh.koujalagi@intel.com>
> Cc: Alan Previn Teres Alexis <alan.previn.teres.alexis@intel.com>
> Cc: Julia Filipchuk <julia.filipchuk@intel.com>
> Assisted-by: Claude:claude-opus-5
missing your s-o-b
> ---
> drivers/gpu/drm/xe/xe_guc_ct.c | 91 +++++++++++++++++++---------------
> 1 file changed, 51 insertions(+), 40 deletions(-)
>
> diff --git a/drivers/gpu/drm/xe/xe_guc_ct.c b/drivers/gpu/drm/xe/xe_guc_ct.c
> index c5a417fef913..623b8c1c6944 100644
> --- a/drivers/gpu/drm/xe/xe_guc_ct.c
> +++ b/drivers/gpu/drm/xe/xe_guc_ct.c
> @@ -30,6 +30,7 @@
> #include "xe_guc_relay.h"
> #include "xe_guc_submit.h"
> #include "xe_guc_tlb_inval.h"
> +#include "xe_log.h"
> #include "xe_map.h"
> #include "xe_page_reclaim.h"
> #include "xe_pm.h"
> @@ -679,7 +680,7 @@ static int __xe_guc_ct_start(struct xe_guc_ct *ct, bool needs_register)
> return 0;
>
> err_out:
> - xe_gt_err(gt, "Failed to enable GuC CT (%pe)\n", ERR_PTR(err));
> + xe_log_err(gt, GUC, err, "CT: Failed to enable\n");
> CT_DEAD(ct, NULL, SETUP);
>
> return err;
> @@ -803,8 +804,9 @@ static bool h2g_has_room(struct xe_guc_ct *ct, u32 cmd_len)
>
> desc_write(xe, h2g, status, desc_status | GUC_CTB_STATUS_OVERFLOW);
>
> - xe_gt_err(ct_to_gt(ct), "CT: invalid head offset %u >= %u)\n",
> - h2g->info.head, h2g->info.size);
> + xe_log_err(ct_to_gt(ct), GUC, -EPIPE,
> + "CT: invalid head offset %u >= %u)\n",
> + h2g->info.head, h2g->info.size);
> CT_DEAD(ct, h2g, H2G_HAS_ROOM);
> return false;
> }
> @@ -873,12 +875,13 @@ static void __g2h_release_space(struct xe_guc_ct *ct, u32 g2h_len)
> bad |= !ct->g2h_outstanding;
>
> if (bad) {
> - xe_gt_err(ct_to_gt(ct), "Invalid G2H release: %d + %d vs %d - %d -> %d vs %d, outstanding = %d!\n",
> - ct->ctbs.g2h.info.space, g2h_len,
> - ct->ctbs.g2h.info.size, ct->ctbs.g2h.info.resv_space,
> - ct->ctbs.g2h.info.space + g2h_len,
> - ct->ctbs.g2h.info.size - ct->ctbs.g2h.info.resv_space,
> - ct->g2h_outstanding);
> + xe_log_err(ct_to_gt(ct), GUC, -ETOOMANYREFS,
> + "CT: Invalid G2H release: %d + %d vs %d - %d -> %d vs %d, outstanding = %d!\n",
> + ct->ctbs.g2h.info.space, g2h_len,
> + ct->ctbs.g2h.info.size, ct->ctbs.g2h.info.resv_space,
> + ct->ctbs.g2h.info.space + g2h_len,
> + ct->ctbs.g2h.info.size - ct->ctbs.g2h.info.resv_space,
> + ct->g2h_outstanding);
> CT_DEAD(ct, &ct->ctbs.g2h, G2H_RELEASE);
> return;
> }
> @@ -963,23 +966,26 @@ static int h2g_write(struct xe_guc_ct *ct, const u32 *action, u32 len,
> desc_status = desc_read(xe, h2g, status);
> if (desc_status) {
> err = -EPIPE;
> - xe_gt_err(gt, "CT write: non-zero status: %u\n", desc_status);
> + xe_log_err(gt, GUC, err,
> + "CT: write: non-zero status: %u\n", desc_status);
> goto corrupted;
> }
>
> if (tail > h2g->info.size) {
> desc_write(xe, h2g, status, desc_status | GUC_CTB_STATUS_OVERFLOW);
> err = -EPIPE;
> - xe_gt_err(gt, "CT write: tail out of range: %u vs %u\n",
> - tail, h2g->info.size);
> + xe_log_err(gt, GUC, err,
> + "CT: write: tail out of range: %u vs %u\n",
> + tail, h2g->info.size);
> goto corrupted;
> }
>
> if (desc_head >= h2g->info.size) {
> desc_write(xe, h2g, status, desc_status | GUC_CTB_STATUS_OVERFLOW);
> err = -EPIPE;
> - xe_gt_err(gt, "CT write: invalid head offset %u >= %u)\n",
> - desc_head, h2g->info.size);
> + xe_log_err(gt, GUC, err,
> + "CT: write: invalid head offset %u >= %u)\n",
> + desc_head, h2g->info.size);
> goto corrupted;
hmm, all 3 above are under DEBUG config, should we really convert them to xe_log?
> }
> }
> @@ -1224,7 +1230,7 @@ static int guc_ct_send_locked(struct xe_guc_ct *ct, const u32 *action, u32 len,
> return ret;
>
> broken:
> - xe_gt_err(gt, "No forward process on H2G, reset required\n");
> + xe_log_err(gt, GUC, -EDEADLK, "CT: No forward progress on H2G, reset required\n");
> CT_DEAD(ct, &ct->ctbs.h2g, DEADLOCK);
>
> return -EDEADLK;
> @@ -1562,11 +1568,12 @@ static int guc_crash_process_msg(struct xe_guc_ct *ct, u32 action)
> struct xe_gt *gt = ct_to_gt(ct);
>
> if (action == XE_GUC_ACTION_NOTIFY_CRASH_DUMP_POSTED)
> - xe_gt_err(gt, "GuC Crash dump notification\n");
> + xe_log_err(gt, GUC, -EHOSTDOWN, "CT: Crash dump notification\n");
drop "CT:", it is a FW crash, and CT is just a comm channel where we get that notif
> else if (action == XE_GUC_ACTION_NOTIFY_EXCEPTION)
> - xe_gt_err(gt, "GuC Exception notification\n");
> + xe_log_err(gt, GUC, -EHOSTDOWN, "CT: Exception notification\n");
ditto
> else
> - xe_gt_err(gt, "Unknown GuC crash notification: 0x%04X\n", action);
> + xe_log_err(gt, GUC, -EHOSTDOWN,
> + "CT: Unknown crash notification: 0x%04X\n", action);
hmm, this looks like our programming error
we call guc_crash_process_msg() only for 2 crash messages
why should we care about something else?
maybe this should be coded outsize xe_guc_ct.c as:
int xe_guc_handle_crash_dump_msg(guc, action[], len)
{
if (len != XE_GUC_ACTION_NOTIFY_CRASH_DUMP_POSTED_MSG_LEN)
return -EPROTO;
xe_log_err(gt, GUC, -EHOSTDOWN, "Crash dump notification\n");
return -EHOSTDOWN;
// or just EPIPE as this should call DEAD_CT
// and kick_reset
}
int xe_guc_handle_exception_msg(guc, action[], len)
{
if (len != XE_GUC_ACTION_NOTIFY_EXCEPTION_MSG_LEN)
return -EPROTO;
xe_log_err(gt, GUC, -EHOSTDOWN, "Exception notification\n");
return -EHOSTDOWN;
// or just EPIPE as this should call DEAD_CT
// and kick_reset
}
>
> CT_DEAD(ct, NULL, CRASH);
>
> @@ -1596,13 +1603,15 @@ static int parse_g2h_response(struct xe_guc_ct *ct, u32 *msg, u32 len)
> */
> if (fence & CT_SEQNO_UNTRACKED) {
> if (type == GUC_HXG_TYPE_RESPONSE_FAILURE)
> - xe_gt_err(gt, "FAST_REQ H2G fence 0x%x failed! e=0x%x, h=%u\n",
> - fence,
> - FIELD_GET(GUC_HXG_FAILURE_MSG_0_ERROR, hxg[0]),
> - FIELD_GET(GUC_HXG_FAILURE_MSG_0_HINT, hxg[0]));
> + xe_log_err(gt, GUC, -EPROTO,
> + "CT: FAST_REQ H2G fence 0x%x failed! e=0x%x, h=%u\n",
> + fence,
> + FIELD_GET(GUC_HXG_FAILURE_MSG_0_ERROR, hxg[0]),
> + FIELD_GET(GUC_HXG_FAILURE_MSG_0_HINT, hxg[0]));
hmm, while FAILURE is unexpected under normal operation,
it is still a valid message according to CTB ABI, likely
caused by host driver mis-programming, so I'm not sure
the -EPROTO is the right error code here, maybe -EINVAL?
> else
> - xe_gt_err(gt, "unexpected response %u for FAST_REQ H2G fence 0x%x!\n",
> - type, fence);
> + xe_log_err(gt, GUC, -EPROTO,
> + "CT: unexpected response %u for FAST_REQ H2G fence 0x%x!\n",
> + type, fence);
OTOH, this looks fine, as there should no other responses
>
> fast_req_report(ct, fence);
>
> @@ -1678,8 +1687,9 @@ static int parse_g2h_msg(struct xe_guc_ct *ct, u32 *msg, u32 len)
>
> origin = FIELD_GET(GUC_HXG_MSG_0_ORIGIN, hxg[0]);
> if (unlikely(origin != GUC_HXG_ORIGIN_GUC)) {
> - xe_gt_err(gt, "G2H channel broken on read, origin=%u, reset required\n",
> - origin);
> + xe_log_err(gt, GUC, -EBADMSG,
> + "CT: Invalid G2H origin=%u, reset required\n",
nit: can we drop this "reset required" phrase?
in other cases (like broken CTB.desc -EPIPE) we don't print that
and we should have some follow up messages saying that we've triggered
a RESET due to something bad in CTB, of from the RESET flow itself:
./xe_gt.c:940: xe_log_info(gt, GT, "reset started\n");
./xe_gt.c:979: xe_log_info(gt, GT, "reset done\n");
.
> + origin);
> CT_DEAD(ct, &ct->ctbs.g2h, PARSE_G2H_ORIGIN);
>
> return -EPROTO;
> @@ -1697,8 +1707,9 @@ static int parse_g2h_msg(struct xe_guc_ct *ct, u32 *msg, u32 len)
> ret = parse_g2h_response(ct, msg, len);
> break;
> default:
> - xe_gt_err(gt, "G2H channel broken on read, type=%u, reset required\n",
> - type);
> + xe_log_err(gt, GUC, -EOPNOTSUPP,
> + "CT: Unexpected G2H message type %u, reset required\n",
> + type);
> CT_DEAD(ct, &ct->ctbs.g2h, PARSE_G2H_TYPE);
>
> ret = -EOPNOTSUPP;
> @@ -1796,8 +1807,8 @@ static int process_g2h_msg(struct xe_guc_ct *ct, u32 *msg, u32 len)
> }
>
> if (ret) {
> - xe_gt_err(gt, "G2H action %#04x failed (%pe) len %u msg %*ph\n",
> - action, ERR_PTR(ret), hxg_len, (int)sizeof(u32) * hxg_len, hxg);
> + xe_log_err(gt, GUC, ret, "CT: G2H action %#04x failed len %u msg %*ph\n",
> + action, hxg_len, (int)sizeof(u32) * hxg_len, hxg);
> CT_DEAD(ct, NULL, PROCESS_FAILED);
> }
>
> @@ -1847,7 +1858,7 @@ static int g2h_read(struct xe_guc_ct *ct, u32 *msg, bool fast_path)
>
> if (desc_status) {
> err = -EPIPE;
> - xe_gt_err(gt, "CT read: non-zero status: %u\n", desc_status);
> + xe_log_err(gt, GUC, err, "CT: read: non-zero status: %u\n", desc_status);
> goto corrupted;
> }
> }
> @@ -1879,16 +1890,16 @@ static int g2h_read(struct xe_guc_ct *ct, u32 *msg, bool fast_path)
> if (g2h->info.head > g2h->info.size) {
> desc_write(xe, g2h, status, desc_status | GUC_CTB_STATUS_OVERFLOW);
> err = -EPIPE;
> - xe_gt_err(gt, "CT read: head out of range: %u vs %u\n",
> - g2h->info.head, g2h->info.size);
> + xe_log_err(gt, GUC, err, "CT: read: head out of range: %u vs %u\n",
> + g2h->info.head, g2h->info.size);
> goto corrupted;
> }
>
> if (desc_tail >= g2h->info.size) {
> desc_write(xe, g2h, status, desc_status | GUC_CTB_STATUS_OVERFLOW);
> err = -EPIPE;
> - xe_gt_err(gt, "CT read: invalid tail offset %u >= %u)\n",
> - desc_tail, g2h->info.size);
> + xe_log_err(gt, GUC, err, "CT: read: invalid tail offset %u >= %u)\n",
> + desc_tail, g2h->info.size);
> goto corrupted;
hmm, again this is under CONFIG_DRM_XE_DEBUG, shall we still use xe_log?
> }
> }
> @@ -1908,8 +1919,9 @@ static int g2h_read(struct xe_guc_ct *ct, u32 *msg, bool fast_path)
> len = FIELD_GET(GUC_CTB_MSG_0_NUM_DWORDS, msg[0]) + GUC_CTB_MSG_MIN_LEN;
> if (len > avail) {
> err = -EPIPE;
> - xe_gt_err(gt, "G2H channel broken on read, avail=%d, len=%d, reset required\n",
> - avail, len);
> + xe_log_err(gt, GUC, err,
> + "CT: G2H channel broken on read, avail=%d, len=%d, reset required\n",
> + avail, len);
> goto corrupted;
> }
>
> @@ -1991,8 +2003,7 @@ static void g2h_fast_path(struct xe_guc_ct *ct, u32 *msg, u32 len)
> }
>
> if (ret) {
> - xe_gt_err(gt, "G2H action 0x%04x failed (%pe)\n",
> - action, ERR_PTR(ret));
> + xe_log_err(gt, GUC, ret, "CT: G2H action 0x%04x failed\n", action);
> CT_DEAD(ct, NULL, FAST_G2H);
> }
> }
> @@ -2111,7 +2122,7 @@ static void receive_g2h(struct xe_guc_ct *ct)
> mutex_unlock(&ct->lock);
>
> if (unlikely(ret < 0 && g2h_err_is_fatal(ret))) {
> - xe_gt_err(ct_to_gt(ct), "CT dequeue failed (%pe)\n", ERR_PTR(ret));
> + xe_log_err(ct_to_gt(ct), GUC, ret, "CT: dequeue failed, forcing GT reset\n");
> CT_DEAD(ct, NULL, G2H_RECV);
> kick_reset(ct);
> }
next prev parent reply other threads:[~2026-09-15 10:24 UTC|newest]
Thread overview: 19+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 23:39 [PATCH v2 0/5] Use SIG_ID logs for GuC component Umesh Nerlige Ramappa
2026-09-03 23:40 ` [PATCH v3 1/5] drm/xe/guc: Use different error codes for GuC load errors Umesh Nerlige Ramappa
2026-09-15 8:58 ` Michal Wajdeczko
2026-09-15 19:19 ` Umesh Nerlige Ramappa
2026-09-03 23:40 ` [PATCH v3 2/5] drm/xe/guc: Use different error codes for CT errors Umesh Nerlige Ramappa
2026-09-03 23:42 ` Umesh Nerlige Ramappa
2026-09-15 9:49 ` Michal Wajdeczko
2026-09-15 22:59 ` Umesh Nerlige Ramappa
2026-09-15 23:10 ` Umesh Nerlige Ramappa
2026-09-18 20:28 ` Umesh Nerlige Ramappa
2026-09-03 23:40 ` [PATCH v3 3/5] drm/xe/uc: Report DMA failure using SIGID Umesh Nerlige Ramappa
2026-09-03 23:40 ` [PATCH v3 4/5] drm/xe/guc: Report major GuC failures " Umesh Nerlige Ramappa
2026-09-15 9:56 ` Michal Wajdeczko
2026-09-03 23:40 ` [PATCH v3 5/5] drm/xe/guc: Report errors that cause a CT shutdown " Umesh Nerlige Ramappa
2026-09-15 10:24 ` Michal Wajdeczko [this message]
2026-09-17 22:25 ` Umesh Nerlige Ramappa
2026-09-04 0:09 ` ✓ CI.KUnit: success for Use SIG_ID logs for GuC component (rev2) Patchwork
2026-09-04 0:48 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-04 14:00 ` ✓ Xe.CI.FULL: " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=7e5691ad-d3ed-48e8-90b0-81ddf737a6f9@intel.com \
--to=michal.wajdeczko@intel.com \
--cc=alan.previn.teres.alexis@intel.com \
--cc=aravind.iddamsetty@intel.com \
--cc=daniele.ceraolospurio@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=julia.filipchuk@intel.com \
--cc=mallesh.koujalagi@intel.com \
--cc=rodrigo.vivi@intel.com \
--cc=umesh.nerlige.ramappa@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.