Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Badal Nilawar" <badal.nilawar@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v3 08/12] drm/xe/xe_ras: Report device memory errors using SIGID
Date: Sun, 06 Sep 2026 17:27:41 +0000	[thread overview]
Message-ID: <20260906172742.488C11F00A3A@smtp.kernel.org> (raw)
In-Reply-To: <20260906172604.2215987-22-badal.nilawar@intel.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:

New issues:
- [Medium] Double-bit ECC memory errors are incorrectly reported with CPER_SEV_INFORMATIONAL severity despite being severe enough to trigger a device reset.

Pre-existing issues:
- [High] Hardware CPER records are silently discarded due to an inverted SIGID check.
--

commit 0369a03d8670385fb852c68b7ab519615f2c4e05
Author: Badal Nilawar <badal.nilawar@intel.com>

drm/xe/xe_ras: Report device memory errors using SIGID

This commit updates device memory error reporting to use the
xe_log_comp_info helper. It maps memory errors like poison, data parity,
and double-bit ECC to this new telemetry interface.

> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index 288dbc0942f58..3a7f80370827e 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c

[ ... ]

> @@ -387,14 +387,17 @@ static u8 handle_device_memory_errors(struct xe_device *xe, struct xe_ras_error_
>  	 */
>  	switch (info->category) {
>  	case XE_RAS_MEMORY_POISON:
> -		xe_info(xe, "[RAS]: Poison error detected\n");
> +		xe_log_comp_info(xe, DEVICE_MEMORY, &arr->counter, sizeof(arr->counter),
> +				 "Poison error detected\n");
>  		break;
>  	case XE_RAS_MEMORY_DATA_PARITY:
> -		xe_info(xe, "[RAS]: Data parity error detected\n");
> +		xe_log_comp_info(xe, DEVICE_MEMORY, &arr->counter, sizeof(arr->counter),
> +				 "Data parity error detected\n");
>  		break;
>  	case XE_RAS_MEMORY_DB_ECC:
> -		xe_info(xe, "[RAS]: Double-bit ECC error detected at sw address 0x%llx\n",
> -			info->sw_address);
> +		xe_log_comp_info(xe, DEVICE_MEMORY, &arr->counter, sizeof(arr->counter),
> +				 "Double-bit ECC error detected at sw address 0x%llx\n",
> +				 info->sw_address);

[Severity: Medium]
Will this incorrectly report double-bit ECC memory errors with
CPER_SEV_INFORMATIONAL severity despite being severe enough to trigger a
device reset?

The xe_log_comp_info macro hardcodes the CPER severity to
CPER_SEV_INFORMATIONAL. After logging, the switch statement falls through
to the default case and returns XE_RAS_RECOVERY_ACTION_RESET. 

Could this break automated fleet management that acts on CPER severities
because a severe uncorrectable double-bit ECC error is logged as
informational while the driver actually resets the GPU?

>  		/* TODO: Add page offlining for Double-bit ECC error */
>  		fallthrough;
>  	default:
>  		return XE_RAS_RECOVERY_ACTION_RESET;
>  	}

[ ... ]


This isn't a bug introduced by this patch, but looking at the broader
telemetry handling in drivers/gpu/drm/xe/xe_cper.c:

> void xe_emit_hardware_error_cper(struct pci_dev *pdev, int cper_sev, enum xe_sigid sigid,
> 				 struct xe_ras_error_class *counter,
> 				 struct xe_ras_get_counter_response *response)
> {
> 	struct xe_device *xe = pdev_to_xe_device(pdev);
>       ...
> 
> 	if ((int)sigid >= INTEL_SIGID_GPU_XE_HARDWARE_START)
> 		return;
>       ...

[Severity: High]
Are hardware CPER records being silently discarded here due to an inverted
SIGID check?

When xe_emit_hardware_error_cper() is called to emit a record for a hardware
error, the sigid parameter is greater than or equal to
INTEL_SIGID_GPU_XE_HARDWARE_START, which is the base value for all hardware
SIGIDs. 

This condition evaluates to true, causing the function to immediately return
before any CPER record is emitted, rendering the logging API calls
completely ineffective for hardware faults.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260906172604.2215987-14-badal.nilawar@intel.com?part=8

  reply	other threads:[~2026-09-06 17:27 UTC|newest]

Thread overview: 45+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-06 17:26 [PATCH v3 00/12] Add CPER logging support for CRI Badal Nilawar
2026-09-06 17:16 ` ✗ CI.checkpatch: warning for Add CPER logging support for CRI (rev3) Patchwork
2026-09-06 17:18 ` ✓ CI.KUnit: success " Patchwork
2026-09-06 17:26 ` [PATCH v3 01/12] drm/xe/cper: Hardware error CPER reporting from xe_log Badal Nilawar
2026-09-06 17:21   ` sashiko-bot
2026-09-07 12:38   ` Michal Wajdeczko
2026-09-10 11:39     ` Nilawar, Badal
2026-09-08 10:12   ` Raag Jadav
2026-09-10 12:33     ` Nilawar, Badal
2026-09-06 17:26 ` [PATCH v3 02/12] drm/xe/cper: Retrieve the error counter record for CPER reporting Badal Nilawar
2026-09-06 17:23   ` sashiko-bot
2026-09-08 10:16   ` Raag Jadav
2026-09-09  6:12     ` Raag Jadav
2026-09-10 12:59       ` Nilawar, Badal
2026-09-10 13:19         ` Raag Jadav
2026-09-06 17:26 ` [PATCH v3 03/12] drm/xe/cper: Add Intel specific CPER structures Badal Nilawar
2026-09-07 13:13   ` Michal Wajdeczko
2026-09-10 11:57     ` Nilawar, Badal
2026-09-08 10:18   ` Raag Jadav
2026-09-10 13:36     ` Nilawar, Badal
2026-09-06 17:26 ` [PATCH v3 04/12] drm/xe/cper: Prepare CPER record Badal Nilawar
2026-09-06 17:27   ` sashiko-bot
2026-09-08 10:20   ` Raag Jadav
2026-09-06 17:26 ` [PATCH v3 05/12] drm/xe/xe_ras: Add support to retrieve info queue data for CRI Badal Nilawar
2026-09-06 17:17   ` sashiko-bot
2026-09-09  8:03   ` Raag Jadav
2026-09-06 17:26 ` [PATCH v3 06/12] drm/xe/cper: Prepare Intel CPER error info records Badal Nilawar
2026-09-06 17:30   ` sashiko-bot
2026-09-09 11:58   ` Raag Jadav
2026-09-06 17:26 ` [PATCH v3 07/12] drm/xe/cper: Log CPER records for aggregate counter retrival Badal Nilawar
2026-09-06 17:23   ` sashiko-bot
2026-09-10  6:27   ` Raag Jadav
2026-09-10 22:29     ` Rodrigo Vivi
2026-09-06 17:26 ` [PATCH v3 08/12] drm/xe/xe_ras: Report device memory errors using SIGID Badal Nilawar
2026-09-06 17:27   ` sashiko-bot [this message]
2026-09-06 17:26 ` [PATCH v3 09/12] drm/xe/xe_ras: Report core compute " Badal Nilawar
2026-09-06 17:21   ` sashiko-bot
2026-09-06 17:26 ` [PATCH v3 10/12] drm/xe/xe_ras: Report soc internal " Badal Nilawar
2026-09-06 17:26 ` [PATCH v3 11/12] drm/xe/xe_ras: Report correctable " Badal Nilawar
2026-09-06 17:27   ` sashiko-bot
2026-09-06 17:26 ` [PATCH v3 12/12] drm/xe/cper: Emit cper record to trace buf Badal Nilawar
2026-09-06 17:28   ` sashiko-bot
2026-09-10  7:58   ` Raag Jadav
2026-09-06 17:55 ` ✓ Xe.CI.BAT: success for Add CPER logging support for CRI (rev3) Patchwork
2026-09-06 19:02 ` ✗ Xe.CI.FULL: failure " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260906172742.488C11F00A3A@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=badal.nilawar@intel.com \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox