Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Michal Wajdeczko" <michal.wajdeczko@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v5 22/33] drm/xe/survivability: Report 'boot status' on sysfs failure
Date: Fri, 14 Aug 2026 17:39:27 +0000	[thread overview]
Message-ID: <20260814173928.925011F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260814171801.13347-23-michal.wajdeczko@intel.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:

Pre-existing issues:
- [High] The `csc_hw_error_work` workqueue item is never canceled during driver teardown, leading to a potential Use-After-Free.
- [High] Hardware-controlled infinite loop while traversing the auxiliary info register linked list.
--

--- Patch [22]: [PATCH v5 22/33] drm/xe/survivability: Report 'boot status' on sysfs failure ---
commit 24540c4b13c7579fcae729221e696b114366060f
Author: Michal Wajdeczko <michal.wajdeczko@intel.com>

drm/xe/survivability: Report 'boot status' on sysfs failure

If we fail to create any of the survivability sysfs files, the user
can't view boot status registers any more. Report them into dmesg
instead using our existing SIGID helper.

> diff --git a/drivers/gpu/drm/xe/xe_survivability_mode.c b/drivers/gpu/drm/xe/xe_survivability_mode.c
> index 8b74bbc963903..e14553295171d 100644
> --- a/drivers/gpu/drm/xe/xe_survivability_mode.c
> +++ b/drivers/gpu/drm/xe/xe_survivability_mode.c
[ ... ]
> @@ -329,6 +330,9 @@ static int create_survivability_sysfs(struct pci_dev *pdev)
>  
>  failed:
>  	xe_err(xe, "Failed to create survivability sysfs files: %pe\n", ERR_PTR(ret));
> +	/* no sysfs, dump Survivability info to dmesg instead */
> +	if (needs_info)
> +		log_survivability_info(xe);
>  	return ret;
>  }

[Severity: High]
This is a pre-existing issue, but can this loop run infinitely if the hardware
returns corrupted or circular data?

drivers/gpu/drm/xe/xe_survivability_mode.c:populate_survivability_info() {
    ...
		for (id = REG_FIELD_GET(AUXINFO_REG_OFFSET, reg_value);
		     id >= AUX_INFO0 && id < MAX_SCRATCH_REG;
		     id =  REG_FIELD_GET(AUXINFO_HISTORY_OFFSET, info[id]))
			set_survivability_info(mmio, info, id);
    ...
}

Since survivability mode handles wedged or failing firmware, the hardware might
return an AUXINFO_HISTORY_OFFSET that points to the current or a previously
visited index.

Does this traversal need a bound or cycle detection to prevent hanging a kernel
thread?

[Severity: High]
This is a pre-existing issue, but is there a potential use-after-free during
driver teardown if this work item is not canceled?

drivers/gpu/drm/xe/xe_hw_error.c:xe_hw_error_init() {
    ...
	INIT_WORK(&tile->csc_hw_error_work, csc_hw_error_work);
    ...
}

The work item is initialized without a corresponding teardown action
(like devm_add_action_or_reset) to cancel it on driver unbind.

If the driver is unloaded while hardware error processing is queued, the worker
could execute after the tile and xe structures are freed:

drivers/gpu/drm/xe/xe_hw_error.c:csc_hw_error_work() {
	struct xe_tile *tile = container_of(work, typeof(*tile), csc_hw_error_work);

	xe_hw_error_print(tile, CSC_HW_ERROR);
}

Does the driver need a cancel_work_sync() during teardown to prevent this?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260814171801.13347-1-michal.wajdeczko@intel.com?part=22

  reply	other threads:[~2026-08-14 17:39 UTC|newest]

Thread overview: 51+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-14 17:17 [PATCH v5 00/33] drm/xe: Add structured SIGID error logging infrastructure Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 01/33] drm/xe: Introduce xe_any helpers Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 02/33] drm/xe/log: Add structured SIGID error logging infrastructure Michal Wajdeczko
2026-08-14 17:31   ` Nilawar, Badal
2026-08-14 17:17 ` [PATCH v5 03/33] drm/xe/log: Introduce structured component/location identifiers Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 04/33] drm/xe/log: Add component/location decorations to dmesg Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 05/33] drm/xe/log: Add SIGID log helpers for severity Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 06/33] drm/xe/log: Add SIGID log helpers for location Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 07/33] drm/xe/log: Add SIGID log helpers for location & severity Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 08/33] drm/xe/log: Add SIGID log helpers for components Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 09/33] drm/xe/log: Add SIGID log helpers for component & severity Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 10/33] drm/xe/log: Add SIGID log helpers for errno-only Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 11/33] drm/xe/log: Index all SIGID printk messages Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 12/33] drm/xe/log: Add hardware error signatures Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 13/33] drm/xe/log: Extend components list with hardware items Michal Wajdeczko
2026-08-14 18:39   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 14/33] drm/xe/ras: Check RAS and LOG component definitions Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 15/33] drm/xe/kunit: Setup driver data in the test device Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 16/33] drm/xe/tests: Add Kunit tests for xe_log Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 17/33] drm/xe/tests: Add kunit tests for xe_any Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 18/33] drm/xe: Report 'probe blocked' status using SIGID Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 19/33] drm/xe: Report all probe errors " Michal Wajdeczko
2026-08-14 17:31   ` sashiko-bot
2026-08-14 17:17 ` [PATCH v5 20/33] drm/xe/survivability: Report 'boot status' " Michal Wajdeczko
2026-08-14 17:26   ` sashiko-bot
2026-08-14 18:41   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 21/33] drm/xe/survivability: Report sysfs failure in one place Michal Wajdeczko
2026-08-14 18:43   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 22/33] drm/xe/survivability: Report 'boot status' on sysfs failure Michal Wajdeczko
2026-08-14 17:39   ` sashiko-bot [this message]
2026-08-14 18:45   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 23/33] drm/xe/survivability: Report 'Boot Mode enabled' status using SIGID Michal Wajdeczko
2026-08-14 18:47   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 24/33] drm/xe/survivability: Report 'Runtime " Michal Wajdeczko
2026-08-14 18:49   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 25/33] drm/xe: Report 'device wedged' errors " Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 26/33] drm/xe/pcode: Report 'Mailbox failed' error " Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 27/33] drm/xe/pcode: Report 'timeout, retrying' " Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 28/33] drm/xe/pcode: Report 'initialization timedout' " Michal Wajdeczko
2026-08-14 17:33   ` sashiko-bot
2026-08-14 17:17 ` [PATCH v5 29/33] drm/xe/guc: Report 'GuC mmio' errors " Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 30/33] drm/xe/gt: Report 'reset failed' " Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 31/33] drm/xe/gt: Report 'Fault response' pagefault error " Michal Wajdeczko
2026-08-14 18:50   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 32/33] drm/xe/gt: Report 'Queue full' " Michal Wajdeczko
2026-08-14 17:34   ` sashiko-bot
2026-08-14 17:17 ` [PATCH v5 33/33] drm/xe/pci: Report 'cannot re-enable' " Michal Wajdeczko
2026-08-14 17:25 ` ✗ CI.checkpatch: warning for drm/xe: Add structured SIGID error logging infrastructure (rev5) Patchwork
2026-08-14 17:27 ` ✓ CI.KUnit: success " Patchwork
2026-08-14 18:13 ` ✓ Xe.CI.BAT: " Patchwork
2026-08-14 18:54 ` [PATCH v5 00/33] drm/xe: Add structured SIGID error logging infrastructure Rodrigo Vivi

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260814173928.925011F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=michal.wajdeczko@intel.com \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox