From: "Nilawar, Badal" <badal.nilawar@intel.com>
To: <sashiko-reviews@lists.linux.dev>
Cc: <intel-xe@lists.freedesktop.org>
Subject: Re: [PATCH v3 02/12] drm/xe/cper: Retrieve the error counter record for CPER reporting
Date: Mon, 5 Oct 2026 23:07:58 +0530 [thread overview]
Message-ID: <7c73a849-5994-4664-952a-682823a14b55@intel.com> (raw)
In-Reply-To: <20260906172343.2A77A1F00A3A@smtp.kernel.org>
On 06-09-2026 22:53, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
>
> New issues:
> - [High] Fetching error counters from the generic logging emission path introduces a might_sleep() restriction into generic logging macros.
> - [High] xe_emit_hardware_error_cper reads MMIO registers via sysctrl command without ensuring the device is awake, leading to potential system crashes or PCIe AER errors if called when the device is in D3cold.
> - [Medium] The patch introduces redundant, synchronous sysctrl mailbox commands to fetch counter values that the driver either already possesses or fetches again immediately.
>
> Pre-existing issues:
> - [High] Preexisting logic bug causes xe_emit_hardware_error_cper to immediately abort for hardware SIGIDs, making the newly added code unreachable.
> --
>
> commit ea598ea85d76d6f40d20fe1e15edf6a22c8e705c
> Author: Badal Nilawar <badal.nilawar@intel.com>
>
> drm/xe/cper: Retrieve the error counter record for CPER reporting
>
> Retrieve the error counter record required to build a CPER record when
> one is not provided by the caller.
>
>> diff --git a/drivers/gpu/drm/xe/xe_cper.c b/drivers/gpu/drm/xe/xe_cper.c
>> index e8017e3ee3a09..f04a91223a43c 100644
>> --- a/drivers/gpu/drm/xe/xe_cper.c
>> +++ b/drivers/gpu/drm/xe/xe_cper.c
> [ ... ]
>> @@ -26,6 +28,8 @@ void xe_emit_hardware_error_cper(struct pci_dev *pdev, int cper_sev, enum xe_sig
>> struct xe_ras_get_counter_response *response)
>> {
>> struct xe_device *xe = pdev_to_xe_device(pdev);
>> + struct xe_ras_get_counter_response local_resp = {};
>> + struct xe_ras_get_counter_response *counter_response = response;
>>
>> if (!xe)
>> return;
>> @@ -33,5 +37,16 @@ void xe_emit_hardware_error_cper(struct pci_dev *pdev, int cper_sev, enum xe_sig
>> if ((int)sigid >= INTEL_SIGID_GPU_XE_HARDWARE_START)
>> return;
> [Severity: High]
> This isn't a bug introduced by this patch, but does this condition in
> xe_emit_hardware_error_cper() prevent the newly added code from executing?
>
> Since hardware SIGIDs are strictly greater than or equal to
> INTEL_SIGID_GPU_XE_HARDWARE_START, checking if sigid is greater than or
> equal to this constant seems to cause an immediate return for all hardware
> errors, silently disabling the CPER record emission.
Will fix this.
>
>> + if (!counter || !xe_ras_counter_is_valid(xe, counter))
>> + return;
>> +
>> + if (!counter_response) {
>> + counter_response = &local_resp;
>> + if (xe_ras_get_counter_response(xe, counter, counter_response)) {
> [Severity: High]
> Does fetching error counters from the generic logging emission path
> introduce a sleep restriction into logging macros?
>
> If a generic macro like xe_log_comp_fatal() is called from an atomic context,
> this synchronous fetch via xe_ras_get_counter_response() will invoke
> xe_sysctrl_send_command(). The underlying sysctrl command asserts might_sleep()
> and waits on a mutex, which could cause a bug or deadlock if triggered from an
> interrupt handler or spinlock.
xe_log_comp_fatal() and related logging macros can reach
xe_ras_get_counter_response(), which sends a sysctrl command and may sleep.
Only hardware SIGIDs associated with a valid xe_ras_error_class take
this path. The callers are xe_ras_process_errors() handlers and
xe_ras_counter_threshold_crossed().
These paths already issue sysctrl commands and therefore execute in a
context where sleeping is allowed.
Will add might_sleep() to enforce the sleep-able-context requirement.
>
> [Severity: High]
> Also, does this hardware access ensure the device is runtime-resumed?
>
> The sysctrl command path uses guard(xe_pm_runtime_noresume)(xe), which
> increments the usage counter but avoids waking a suspended device. If this
> logging path executes while the device is in D3cold, will the subsequent
> MMIO reads hit powered-off hardware and trigger unhandled PCIe AER errors?
Callers are expected to hold outer runtime pm reference. Will add
xe_device_assert_mem_access().
>
> [Severity: Medium]
> Finally, does this force a new mailbox command to fetch data that the caller
> might already possess?
>
> For example, in xe_ras_process_errors(), the driver already executes a
> mailbox command that returns the counter values. However, it passes only a
> pointer to the generic logging function, forcing
> xe_emit_hardware_error_cper() to issue an entirely new, redundant synchronous
> mailbox command to fetch the same value. Can this be optimized to reuse the
> existing data and avoid extra latency during error handling?
There is no redundant mailbox command. Call inside xe_ras_process_errors
fetch the aggregate counter for netlink event whereas
xe_emit_hardware_error_cper() fetches detail error counter required for
CPER, which also contains info queue data.
Thanks,
Badal
>
>> + xe_err(xe, "[RAS]: CPER: failed to get counter, skipping record\n");
>> + return;
>> + }
>> + }
>> +
>> /* TODO */
>> }
next prev parent reply other threads:[~2026-10-05 17:38 UTC|newest]
Thread overview: 48+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-06 17:26 [PATCH v3 00/12] Add CPER logging support for CRI Badal Nilawar
2026-09-06 17:16 ` ✗ CI.checkpatch: warning for Add CPER logging support for CRI (rev3) Patchwork
2026-09-06 17:18 ` ✓ CI.KUnit: success " Patchwork
2026-09-06 17:26 ` [PATCH v3 01/12] drm/xe/cper: Hardware error CPER reporting from xe_log Badal Nilawar
2026-09-06 17:21 ` sashiko-bot
2026-09-07 12:38 ` Michal Wajdeczko
2026-09-10 11:39 ` Nilawar, Badal
2026-09-08 10:12 ` Raag Jadav
2026-09-10 12:33 ` Nilawar, Badal
2026-09-06 17:26 ` [PATCH v3 02/12] drm/xe/cper: Retrieve the error counter record for CPER reporting Badal Nilawar
2026-09-06 17:23 ` sashiko-bot
2026-10-05 17:37 ` Nilawar, Badal [this message]
2026-09-08 10:16 ` Raag Jadav
2026-09-09 6:12 ` Raag Jadav
2026-09-10 12:59 ` Nilawar, Badal
2026-09-10 13:19 ` Raag Jadav
2026-09-06 17:26 ` [PATCH v3 03/12] drm/xe/cper: Add Intel specific CPER structures Badal Nilawar
2026-09-07 13:13 ` Michal Wajdeczko
2026-09-10 11:57 ` Nilawar, Badal
2026-09-08 10:18 ` Raag Jadav
2026-09-10 13:36 ` Nilawar, Badal
2026-09-06 17:26 ` [PATCH v3 04/12] drm/xe/cper: Prepare CPER record Badal Nilawar
2026-09-06 17:27 ` sashiko-bot
2026-09-08 10:20 ` Raag Jadav
2026-09-06 17:26 ` [PATCH v3 05/12] drm/xe/xe_ras: Add support to retrieve info queue data for CRI Badal Nilawar
2026-09-06 17:17 ` sashiko-bot
2026-09-09 8:03 ` Raag Jadav
2026-09-06 17:26 ` [PATCH v3 06/12] drm/xe/cper: Prepare Intel CPER error info records Badal Nilawar
2026-09-06 17:30 ` sashiko-bot
2026-09-09 11:58 ` Raag Jadav
2026-09-06 17:26 ` [PATCH v3 07/12] drm/xe/cper: Log CPER records for aggregate counter retrival Badal Nilawar
2026-09-06 17:23 ` sashiko-bot
2026-10-08 13:19 ` Nilawar, Badal
2026-09-10 6:27 ` Raag Jadav
2026-09-10 22:29 ` Rodrigo Vivi
2026-09-06 17:26 ` [PATCH v3 08/12] drm/xe/xe_ras: Report device memory errors using SIGID Badal Nilawar
2026-09-06 17:27 ` sashiko-bot
2026-10-08 13:40 ` Nilawar, Badal
2026-09-06 17:26 ` [PATCH v3 09/12] drm/xe/xe_ras: Report core compute " Badal Nilawar
2026-09-06 17:21 ` sashiko-bot
2026-09-06 17:26 ` [PATCH v3 10/12] drm/xe/xe_ras: Report soc internal " Badal Nilawar
2026-09-06 17:26 ` [PATCH v3 11/12] drm/xe/xe_ras: Report correctable " Badal Nilawar
2026-09-06 17:27 ` sashiko-bot
2026-09-06 17:26 ` [PATCH v3 12/12] drm/xe/cper: Emit cper record to trace buf Badal Nilawar
2026-09-06 17:28 ` sashiko-bot
2026-09-10 7:58 ` Raag Jadav
2026-09-06 17:55 ` ✓ Xe.CI.BAT: success for Add CPER logging support for CRI (rev3) Patchwork
2026-09-06 19:02 ` ✗ Xe.CI.FULL: failure " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=7c73a849-5994-4664-952a-682823a14b55@intel.com \
--to=badal.nilawar@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.