All of lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Riana Tauro" <riana.tauro@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH 3/3] drm/xe/xe_hw_error: Handle non-fatal errors routed via direct IRQ
Date: Sun, 30 Aug 2026 07:09:15 +0000	[thread overview]
Message-ID: <20260830070915.8F2CA1F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260830065603.419055-8-riana.tauro@intel.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] `hw_error_work` accesses system controller MMIO registers without holding an active runtime PM reference, risking MMIO faults if the device is suspended.
- [Medium] `hw_error_work` unconditionally overwrites the device wedge method with `DRM_WEDGE_RECOVERY_BUS_RESET`, clobbering previously set recovery methods and sending duplicate wedge events.
--

commit 7bcb91bb7afbf8e7efbb4e0783cf8098239c131e
Author: Riana Tauro <riana.tauro@intel.com>

drm/xe/xe_hw_error: Handle non-fatal errors routed via direct IRQ
    
This commit modifies xe_hw_error to process, log, and recover non-fatal errors
routed via direct IRQ. For errors that cannot be recovered locally, the device
is wedged using a bus-reset as the recovery action.

> diff --git a/drivers/gpu/drm/xe/xe_hw_error.c b/drivers/gpu/drm/xe/xe_hw_error.c
> index 5a57b7eaebc3b..a895436f2ce7e 100644
> --- a/drivers/gpu/drm/xe/xe_hw_error.c
> +++ b/drivers/gpu/drm/xe/xe_hw_error.c

[ ... ]

>  static void hw_error_work(struct work_struct *work)
>  {
>  	struct xe_tile *tile = container_of(work, typeof(*tile), hw_error_work);
>  	struct xe_device *xe = tile_to_xe(tile);
>  
> +	if (xe->info.has_sysctrl) {
> +		int ret;
> +
> +		ret = xe_ras_process_errors(xe);

[Severity: High]
Does this code access system controller MMIO registers without holding an
active runtime PM reference?

When hw_error_source_handler() schedules hw_error_work() asynchronously, it
doesn't take a runtime PM reference. Once this work executes, it calls
xe_ras_process_errors() which eventually calls xe_sysctrl_send_command().

While xe_sysctrl_send_command() uses guard(xe_pm_runtime_noresume)(xe), this
only increments the usage count and asserts a reference is held; it does not
actually wake a suspended device. If the device transitions to runtime suspend
between the hardware error IRQ firing and the workqueue executing, the
subsequent MMIO reads in sysctrl_read_frame() will access powered-down
hardware, potentially causing PCI target aborts, 0xFFFFFFFF reads, or system
hangs.

> +		/* For any non-fatal errors that do not return recovered, declare wedged */
> +		if (ret) {
> +			xe_device_set_wedged_method(xe, DRM_WEDGE_RECOVERY_BUS_RESET);
> +			xe_device_declare_wedged(xe);

[Severity: Medium]
Can this unconditionally overwrite the device wedge method, clobbering
previously set recovery methods and sending duplicate wedge events?

During xe_ras_process_errors(), functions like handle_soc_internal_errors()
may call xe_survivability_mode_runtime_enable(xe). That function sets the
wedge method to DRM_WEDGE_RECOVERY_VENDOR and declares the device wedged,
before returning XE_RAS_RECOVERY_ACTION_DISCONNECT:

drivers/gpu/drm/xe/xe_ras.c:handle_soc_internal_errors() {
	...
	xe_survivability_mode_runtime_enable(xe);
	return XE_RAS_RECOVERY_ACTION_DISCONNECT;
}

Here, ret (the returned action) is treated as a boolean. By unconditionally
calling xe_device_set_wedged_method(xe, DRM_WEDGE_RECOVERY_BUS_RESET) and
xe_device_declare_wedged(xe) again, the device is presented to userspace with
an incorrect recovery method (BUS_RESET) instead of the true requirement
(VENDOR), and duplicate uevents are emitted.

> +		}
> +
> +		return;
> +	}
> +
>  	xe_survivability_mode_runtime_enable(xe);
>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260830065603.419055-5-riana.tauro@intel.com?part=3

  reply	other threads:[~2026-08-30  7:09 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-30  6:56 [PATCH 0/3] Handle non-fatal errors routed via IRQ Riana Tauro
2026-08-30  6:56 ` [PATCH 1/3] drm/xe/hw_error: Reuse CSC worker for generic error handling Riana Tauro
2026-08-30  6:56 ` [PATCH 2/3] drm/xe/hw_error: Initialize hw_error_work per tile Riana Tauro
2026-08-30  7:14   ` sashiko-bot
2026-08-30  6:56 ` [PATCH 3/3] drm/xe/xe_hw_error: Handle non-fatal errors routed via direct IRQ Riana Tauro
2026-08-30  7:09   ` sashiko-bot [this message]
2026-08-30  7:03 ` ✓ CI.KUnit: success for Handle non-fatal errors routed via IRQ Patchwork
2026-08-30  7:46 ` ✓ Xe.CI.BAT: " Patchwork
2026-08-30  8:58 ` ✓ Xe.CI.FULL: " Patchwork

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260830070915.8F2CA1F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=riana.tauro@intel.com \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.