From: sashiko-bot@kernel.org
To: "Umesh Nerlige Ramappa" <umesh.nerlige.ramappa@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v4 7/7] drm/xe/guc: Report errors that cause a CT shutdown using SIGID
Date: Fri, 18 Sep 2026 00:08:43 +0000 [thread overview]
Message-ID: <20260918000844.05A9A1F00893@smtp.kernel.org> (raw)
In-Reply-To: <20260917235923.1521112-16-umesh.nerlige.ramappa@intel.com>
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] xe: missed xe_log_err conversion in guc crash handlers
- [Medium] xe: mismatched error codes returned in parse_g2h_msg and parse_g2h_response
--
commit 3134241c4666d211fb1d7e98e3ce95c50000c2e9
Author: Daniele Ceraolo Spurio <daniele.ceraolospurio@intel.com>
Subject: drm/xe/guc: Report errors that cause a CT shutdown using SIGID
This patch converts errors that result in a CT being declared dead to use
the xe_log_err helper, which supports SIGID reporting. Escalate-to-caller
errors remain unchanged so callers can apply SIGID if needed. Error
messages are also updated to be clearer about the failure.
> diff --git a/drivers/gpu/drm/xe/xe_guc_ct.c b/drivers/gpu/drm/xe/xe_guc_ct.c
> index 31ecddab30578..63d79987c5d28 100644
> --- a/drivers/gpu/drm/xe/xe_guc_ct.c
> +++ b/drivers/gpu/drm/xe/xe_guc_ct.c
[ ... ]
> @@ -1662,13 +1665,15 @@ static int parse_g2h_response(struct xe_guc_ct *ct, u32 *msg, u32 len)
> */
> if (fence & CT_SEQNO_UNTRACKED) {
> if (type == GUC_HXG_TYPE_RESPONSE_FAILURE)
> - xe_gt_err(gt, "FAST_REQ H2G fence 0x%x failed! e=0x%x, h=%u\n",
> - fence,
> - FIELD_GET(GUC_HXG_FAILURE_MSG_0_ERROR, hxg[0]),
> - FIELD_GET(GUC_HXG_FAILURE_MSG_0_HINT, hxg[0]));
> + xe_log_err(gt, GUC, -EINVAL,
> + "CT: FAST_REQ H2G fence 0x%x failed! e=0x%x, h=%u\n",
> + fence,
> + FIELD_GET(GUC_HXG_FAILURE_MSG_0_ERROR, hxg[0]),
> + FIELD_GET(GUC_HXG_FAILURE_MSG_0_HINT, hxg[0]));
> else
> - xe_gt_err(gt, "unexpected response %u for FAST_REQ H2G fence 0x%x!\n",
> - type, fence);
> + xe_log_err(gt, GUC, -EPROTO,
> + "CT: unexpected response %u for FAST_REQ H2G fence 0x%x!\n",
> + type, fence);
>
> fast_req_report(ct, fence);
>
> /* FIXME: W/A race in the GuC, will get in firmware soon */
> if (xe_gt_recovery_pending(gt))
> return 0;
>
> CT_DEAD(ct, NULL, PARSE_G2H_RESPONSE);
>
> return -EPROTO;
[Severity: Medium]
Does this cause misleading double-logging when receive_g2h() processes the
return value?
Here in parse_g2h_response(), we log the specific -EINVAL error code, but
return -EPROTO to the caller. When the caller receive_g2h() receives this,
it will log the failure a second time as a generic -EPROTO error since the
error codes don't match.
> }
>
[ ... ]
> @@ -1744,7 +1749,7 @@ static int parse_g2h_msg(struct xe_guc_ct *ct, u32 *msg, u32 len)
>
> origin = FIELD_GET(GUC_HXG_MSG_0_ORIGIN, hxg[0]);
> if (unlikely(origin != GUC_HXG_ORIGIN_GUC)) {
> - xe_gt_err(gt, "Invalid G2H origin=%u\n", origin);
> + xe_log_err(gt, GUC, -EBADMSG, "CT: Invalid G2H origin=%u\n", origin);
> CT_DEAD(ct, &ct->ctbs.g2h, PARSE_G2H_ORIGIN);
>
> return -EPROTO;
[Severity: Medium]
Similar to the check in parse_g2h_response(), does logging -EBADMSG here
but returning -EPROTO result in contradictory double-logging in the caller?
> }
>
[ ... ]
> @@ -1862,8 +1867,8 @@ static int process_g2h_msg(struct xe_guc_ct *ct, u32 *msg, u32 len)
> }
>
> if (ret) {
> - xe_gt_err(gt, "G2H action %#04x failed (%pe) len %u msg %*ph\n",
> - action, ERR_PTR(ret), hxg_len, (int)sizeof(u32) * hxg_len, hxg);
> + xe_log_err(gt, GUC, ret, "CT: G2H action %#04x failed len %u msg %*ph\n",
> + action, hxg_len, (int)sizeof(u32) * hxg_len, hxg);
> CT_DEAD(ct, NULL, PROCESS_FAILED);
> }
[Severity: Medium]
The commit message notes this converts any errors causing a CT shutdown to use
xe_log_err, but did it miss the GuC crash and exception handlers?
When a GuC crash or exception occurs, process_g2h_msg() calls
guc_handle_crash_dump_msg() and guc_handle_exception_msg(). Both of these
handlers call CT_DEAD directly and use the older xe_gt_err for logging:
guc_handle_crash_dump_msg()
xe_gt_err(ct_to_gt(ct), "GuC Crash dump notification\n");
CT_DEAD(ct, NULL, CRASH);
Should these be converted to xe_log_err as well to ensure these GuC crashes
receive a SIGID? As it stands, they are untagged, while the fallback generic
PROCESS_FAILED error above does get a SIGID.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260917235923.1521112-9-umesh.nerlige.ramappa@intel.com?part=7
next prev parent reply other threads:[~2026-09-18 0:08 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 23:59 [PATCH v4 0/7] Use SIG_ID logs for GuC component Umesh Nerlige Ramappa
2026-09-17 23:59 ` [PATCH v4 1/7] drm/xe/guc: Use different error codes for GuC load errors Umesh Nerlige Ramappa
2026-09-17 23:59 ` [PATCH v4 2/7] drm/xe/guc: Handle CRASH and EXCEPTION G2H with separate helpers Umesh Nerlige Ramappa
2026-09-18 0:08 ` sashiko-bot
2026-09-17 23:59 ` [PATCH v4 3/7] drm/xe/guc: Make ct_dead_capture available on non-debug config Umesh Nerlige Ramappa
2026-09-18 0:06 ` sashiko-bot
2026-09-17 23:59 ` [PATCH v4 4/7] drm/xe/guc: Cleanup error codes and handling for CT errors Umesh Nerlige Ramappa
2026-09-18 0:17 ` sashiko-bot
2026-09-17 23:59 ` [PATCH v4 5/7] drm/xe/uc: Report DMA failure using SIGID Umesh Nerlige Ramappa
2026-09-18 0:07 ` sashiko-bot
2026-09-17 23:59 ` [PATCH v4 6/7] drm/xe/guc: Report major GuC failures " Umesh Nerlige Ramappa
2026-09-17 23:59 ` [PATCH v4 7/7] drm/xe/guc: Report errors that cause a CT shutdown " Umesh Nerlige Ramappa
2026-09-18 0:08 ` sashiko-bot [this message]
2026-09-18 0:58 ` ✗ CI.KUnit: failure for Use SIG_ID logs for GuC component (rev3) Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260918000844.05A9A1F00893@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=intel-xe@lists.freedesktop.org \
--cc=sashiko-reviews@lists.linux.dev \
--cc=umesh.nerlige.ramappa@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox