From: sashiko-bot@kernel.org
To: "Badal Nilawar" <badal.nilawar@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v2 08/11] drm/xe/ras: Report device memory errors using SIGID
Date: Tue, 25 Aug 2026 17:58:56 +0000 [thread overview]
Message-ID: <20260825175856.CF7071F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260825175916.1103841-21-badal.nilawar@intel.com>
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] The `xe_log_err` family of macros passes `ERR_PTR(err)` to the logging framework, which crashes when used with hardware component tags due to an unchecked cast.
--
commit 3a41849b5bfff0854c8490dfdd5d178d31641b09
Author: Badal Nilawar <badal.nilawar@intel.com>
drm/xe/ras: Report device memory errors using SIGID
This commit updates the device memory error handler to use the
xe_log_comp_info() helper instead of xe_info(). This allows reporting
these errors with the DEVICE_MEMORY hardware component tag.
> diff --git a/drivers/gpu/drm/xe/xe_ras.c b/drivers/gpu/drm/xe/xe_ras.c
> index b4cdb5ec64915..172653be1b829 100644
> --- a/drivers/gpu/drm/xe/xe_ras.c
> +++ b/drivers/gpu/drm/xe/xe_ras.c
> @@ -409,14 +409,17 @@ static u8 handle_device_memory_errors(struct xe_device *xe, struct xe_ras_error_
> */
> switch (info->category) {
> case XE_RAS_MEMORY_POISON:
> - xe_info(xe, "[RAS]: Poison error detected\n");
> + xe_log_comp_info(xe, DEVICE_MEMORY, &arr->counter, sizeof(arr->counter),
> + "Poison error detected\n");
[Severity: Medium]
Does introducing the DEVICE_MEMORY hardware tag expose a latent API trap in
the logging framework?
While this specific call passes a valid pointer (&arr->counter), the logging
macro family is documented to allow passing ERR_PTR(err) as the data
parameter.
If a developer later calls xe_log_err(any, DEVICE_MEMORY, err, ...), it
passes ERR_PTR(err) as the data parameter. This gets processed in
log_emit_cper():
drivers/gpu/drm/xe/xe_log.c:log_emit_cper() {
...
if (is_hw_sigid(sigid))
xe_emit_hardware_error_cper(pdev, cper_sev, sigid,
(struct xe_ras_error_class *)data);
...
}
This blindly casts the data parameter to struct xe_ras_error_class *.
Then xe_emit_hardware_error_cper() checks if it is non-NULL but fails to check
for an error pointer:
drivers/gpu/drm/xe/xe_ras.c:xe_emit_hardware_error_cper() {
...
if (counter && !ras_counter_is_valid(xe, counter))
...
}
This leads to an invalid pointer dereference when ras_counter_is_valid()
accesses the severity field:
drivers/gpu/drm/xe/xe_ras.c:ras_counter_is_valid() {
...
u8 severity = counter->common.severity;
...
}
Should an IS_ERR() check be added in log_emit_cper() or
xe_emit_hardware_error_cper() before dereferencing the pointer to prevent a
potential kernel crash?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260825175916.1103841-13-badal.nilawar@intel.com?part=8
next prev parent reply other threads:[~2026-08-25 17:58 UTC|newest]
Thread overview: 35+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-25 17:59 [PATCH v2 00/11] Add CPER logging support for CRI Badal Nilawar
2026-08-25 17:59 ` [PATCH v2 01/11] drm/xe/xe_ras: Add support to retrieve info queue data " Badal Nilawar
2026-08-25 17:53 ` sashiko-bot
2026-08-26 0:54 ` Rodrigo Vivi
2026-08-25 20:48 ` Michal Wajdeczko
2026-08-25 17:59 ` [PATCH v2 02/11] drm/xe/xe_ras: Refactor get_counter() to return response structure Badal Nilawar
2026-08-25 17:59 ` [PATCH v2 03/11] drm/xe/cper: Add CPER structures and trace event Badal Nilawar
2026-08-25 17:51 ` sashiko-bot
2026-08-28 15:23 ` Rodrigo Vivi
2026-08-25 17:59 ` [PATCH v2 04/11] drm/xe/cper: APIs to prepare and log CPER record Badal Nilawar
2026-08-25 18:02 ` sashiko-bot
2026-08-26 0:59 ` Rodrigo Vivi
2026-08-25 17:59 ` [PATCH v2 05/11] drm/xe/cper: Prepare Intel CPER error info from info queue Badal Nilawar
2026-08-25 17:54 ` sashiko-bot
2026-08-25 17:59 ` [PATCH v2 06/11] drm/xe/cper: Log CPER records for aggregate counter retrival Badal Nilawar
2026-08-25 17:55 ` sashiko-bot
2026-08-26 1:01 ` Rodrigo Vivi
2026-08-25 17:59 ` [PATCH v2 07/11] drm/xe/cper: Allow hardware error CPER reporting from xe_log Badal Nilawar
2026-08-25 17:54 ` sashiko-bot
2026-08-27 21:27 ` Michal Wajdeczko
2026-08-25 17:59 ` [PATCH v2 08/11] drm/xe/ras: Report device memory errors using SIGID Badal Nilawar
2026-08-25 17:58 ` sashiko-bot [this message]
2026-08-27 20:25 ` Michal Wajdeczko
2026-08-25 17:59 ` [PATCH v2 09/11] drm/xe/ras: Report core compute " Badal Nilawar
2026-08-25 17:55 ` sashiko-bot
2026-08-25 17:59 ` [PATCH v2 10/11] drm/xe/ras: Report soc internal " Badal Nilawar
2026-08-28 15:20 ` Rodrigo Vivi
2026-08-25 17:59 ` [PATCH v2 11/11] drm/xe/ras: Report correctable " Badal Nilawar
2026-08-25 18:03 ` sashiko-bot
2026-08-25 18:29 ` ✗ CI.checkpatch: warning for Add CPER logging support for CRI (rev2) Patchwork
2026-08-25 18:31 ` ✓ CI.KUnit: success " Patchwork
2026-08-25 19:25 ` ✓ Xe.CI.BAT: " Patchwork
2026-08-25 22:06 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-08-26 19:50 ` [PATCH v2 00/11] Add CPER logging support for CRI Matt Roper
2026-08-27 20:12 ` Rodrigo Vivi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260825175856.CF7071F000E9@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=badal.nilawar@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.