Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Umesh Nerlige Ramappa" <umesh.nerlige.ramappa@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v4 7/7] drm/xe/guc: Report errors that cause a CT shutdown using SIGID
Date: Fri, 18 Sep 2026 00:08:43 +0000	[thread overview]
Message-ID: <20260918000844.05A9A1F00893@smtp.kernel.org> (raw)
In-Reply-To: <20260917235923.1521112-16-umesh.nerlige.ramappa@intel.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] xe: missed xe_log_err conversion in guc crash handlers
- [Medium] xe: mismatched error codes returned in parse_g2h_msg and parse_g2h_response
--

commit 3134241c4666d211fb1d7e98e3ce95c50000c2e9
Author: Daniele Ceraolo Spurio <daniele.ceraolospurio@intel.com>
Subject: drm/xe/guc: Report errors that cause a CT shutdown using SIGID

This patch converts errors that result in a CT being declared dead to use
the xe_log_err helper, which supports SIGID reporting. Escalate-to-caller
errors remain unchanged so callers can apply SIGID if needed. Error
messages are also updated to be clearer about the failure.

> diff --git a/drivers/gpu/drm/xe/xe_guc_ct.c b/drivers/gpu/drm/xe/xe_guc_ct.c
> index 31ecddab30578..63d79987c5d28 100644
> --- a/drivers/gpu/drm/xe/xe_guc_ct.c
> +++ b/drivers/gpu/drm/xe/xe_guc_ct.c
[ ... ]
> @@ -1662,13 +1665,15 @@ static int parse_g2h_response(struct xe_guc_ct *ct, u32 *msg, u32 len)
>  	 */
>  	if (fence & CT_SEQNO_UNTRACKED) {
>  		if (type == GUC_HXG_TYPE_RESPONSE_FAILURE)
> -			xe_gt_err(gt, "FAST_REQ H2G fence 0x%x failed! e=0x%x, h=%u\n",
> -				  fence,
> -				  FIELD_GET(GUC_HXG_FAILURE_MSG_0_ERROR, hxg[0]),
> -				  FIELD_GET(GUC_HXG_FAILURE_MSG_0_HINT, hxg[0]));
> +			xe_log_err(gt, GUC, -EINVAL,
> +				   "CT: FAST_REQ H2G fence 0x%x failed! e=0x%x, h=%u\n",
> +				   fence,
> +				   FIELD_GET(GUC_HXG_FAILURE_MSG_0_ERROR, hxg[0]),
> +				   FIELD_GET(GUC_HXG_FAILURE_MSG_0_HINT, hxg[0]));
>  		else
> -			xe_gt_err(gt, "unexpected response %u for FAST_REQ H2G fence 0x%x!\n",
> -				  type, fence);
> +			xe_log_err(gt, GUC, -EPROTO,
> +				   "CT: unexpected response %u for FAST_REQ H2G fence 0x%x!\n",
> +				    type, fence);
>  
>  		fast_req_report(ct, fence);
>  
>  		/* FIXME: W/A race in the GuC, will get in firmware soon */
>  		if (xe_gt_recovery_pending(gt))
>  			return 0;
>  
>  		CT_DEAD(ct, NULL, PARSE_G2H_RESPONSE);
>  
>  		return -EPROTO;

[Severity: Medium]
Does this cause misleading double-logging when receive_g2h() processes the
return value?

Here in parse_g2h_response(), we log the specific -EINVAL error code, but
return -EPROTO to the caller. When the caller receive_g2h() receives this,
it will log the failure a second time as a generic -EPROTO error since the
error codes don't match.

>  	}
>  
[ ... ]
> @@ -1744,7 +1749,7 @@ static int parse_g2h_msg(struct xe_guc_ct *ct, u32 *msg, u32 len)
>  
>  	origin = FIELD_GET(GUC_HXG_MSG_0_ORIGIN, hxg[0]);
>  	if (unlikely(origin != GUC_HXG_ORIGIN_GUC)) {
> -		xe_gt_err(gt, "Invalid G2H origin=%u\n", origin);
> +		xe_log_err(gt, GUC, -EBADMSG, "CT: Invalid G2H origin=%u\n", origin);
>  		CT_DEAD(ct, &ct->ctbs.g2h, PARSE_G2H_ORIGIN);
>  
>  		return -EPROTO;

[Severity: Medium]
Similar to the check in parse_g2h_response(), does logging -EBADMSG here
but returning -EPROTO result in contradictory double-logging in the caller?

>  	}
>  
[ ... ]
> @@ -1862,8 +1867,8 @@ static int process_g2h_msg(struct xe_guc_ct *ct, u32 *msg, u32 len)
>  	}
>  
>  	if (ret) {
> -		xe_gt_err(gt, "G2H action %#04x failed (%pe) len %u msg %*ph\n",
> -			  action, ERR_PTR(ret), hxg_len, (int)sizeof(u32) * hxg_len, hxg);
> +		xe_log_err(gt, GUC, ret, "CT: G2H action %#04x failed len %u msg %*ph\n",
> +			   action, hxg_len, (int)sizeof(u32) * hxg_len, hxg);
>  		CT_DEAD(ct, NULL, PROCESS_FAILED);
>  	}

[Severity: Medium]
The commit message notes this converts any errors causing a CT shutdown to use
xe_log_err, but did it miss the GuC crash and exception handlers?

When a GuC crash or exception occurs, process_g2h_msg() calls
guc_handle_crash_dump_msg() and guc_handle_exception_msg(). Both of these
handlers call CT_DEAD directly and use the older xe_gt_err for logging:

guc_handle_crash_dump_msg()
    xe_gt_err(ct_to_gt(ct), "GuC Crash dump notification\n");
    CT_DEAD(ct, NULL, CRASH);

Should these be converted to xe_log_err as well to ensure these GuC crashes
receive a SIGID? As it stands, they are untagged, while the fallback generic
PROCESS_FAILED error above does get a SIGID.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260917235923.1521112-9-umesh.nerlige.ramappa@intel.com?part=7

  reply	other threads:[~2026-09-18  0:08 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-17 23:59 [PATCH v4 0/7] Use SIG_ID logs for GuC component Umesh Nerlige Ramappa
2026-09-17 23:59 ` [PATCH v4 1/7] drm/xe/guc: Use different error codes for GuC load errors Umesh Nerlige Ramappa
2026-09-17 23:59 ` [PATCH v4 2/7] drm/xe/guc: Handle CRASH and EXCEPTION G2H with separate helpers Umesh Nerlige Ramappa
2026-09-18  0:08   ` sashiko-bot
2026-09-17 23:59 ` [PATCH v4 3/7] drm/xe/guc: Make ct_dead_capture available on non-debug config Umesh Nerlige Ramappa
2026-09-18  0:06   ` sashiko-bot
2026-09-17 23:59 ` [PATCH v4 4/7] drm/xe/guc: Cleanup error codes and handling for CT errors Umesh Nerlige Ramappa
2026-09-18  0:17   ` sashiko-bot
2026-09-17 23:59 ` [PATCH v4 5/7] drm/xe/uc: Report DMA failure using SIGID Umesh Nerlige Ramappa
2026-09-18  0:07   ` sashiko-bot
2026-09-17 23:59 ` [PATCH v4 6/7] drm/xe/guc: Report major GuC failures " Umesh Nerlige Ramappa
2026-09-17 23:59 ` [PATCH v4 7/7] drm/xe/guc: Report errors that cause a CT shutdown " Umesh Nerlige Ramappa
2026-09-18  0:08   ` sashiko-bot [this message]
2026-09-18  0:58 ` ✗ CI.KUnit: failure for Use SIG_ID logs for GuC component (rev3) Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260918000844.05A9A1F00893@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=umesh.nerlige.ramappa@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox