Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Michal Wajdeczko" <michal.wajdeczko@intel.com>
Cc: intel-xe@lists.freedesktop.org
Subject: Re: [PATCH v5 28/33] drm/xe/pcode: Report 'initialization timedout' error using SIGID
Date: Fri, 14 Aug 2026 17:33:46 +0000	[thread overview]
Message-ID: <20260814173346.9E0F51F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260814171801.13347-29-michal.wajdeczko@intel.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:

Pre-existing issues:
- [High] The polling loop in `pcode_try_request` uses unbounded exponential backoff, causing massive delay intervals, dangerous `udelay()` values, and incorrect timeout calculations.
--

--- Patch [28]: [PATCH v5 28/33] drm/xe/pcode: Report 'initialization timedout' error using SIGID ---
commit 37b63ecd447fce41ca2633f291eae720f6ea1ba1
Author: Michal Wajdeczko <michal.wajdeczko@intel.com>

drm/xe/pcode: Report 'initialization timedout' error using SIGID

This commit updates xe_pcode_ready() to report initialization timeouts
using the xe_log_err helper with a SIGID. It also replaces the hardcoded
integer timeout with a calculated value using USEC_PER_SEC.

> diff --git a/drivers/gpu/drm/xe/xe_pcode.c b/drivers/gpu/drm/xe/xe_pcode.c
> index 14b13a11d1d6d..e1b8062541a91 100644
> --- a/drivers/gpu/drm/xe/xe_pcode.c
> +++ b/drivers/gpu/drm/xe/xe_pcode.c
> @@ -297,7 +297,7 @@ int xe_pcode_ready(struct xe_device *xe, bool locked)
>  {
>  	u32 status, request = DGFX_GET_INIT_STATUS;
>  	struct xe_tile *tile = xe_device_get_root_tile(xe);
> -	int timeout_us = 180000000; /* 3 min */
> +	long timeout_us = 3 * 60 * USEC_PER_SEC; /* 3 min */

[Severity: High]
This is a pre-existing issue, but does the polling loop in
pcode_try_request() have an uncapped exponential backoff that causes
dangerous wait intervals here?

With timeout_us set to 3 minutes (180,000,000 us), the backoff loop
in pcode_try_request() doubles the wait variable repeatedly:

	for (slept = 0; slept < timeout_us; slept += wait) {
...
		if (atomic)
			udelay(wait);
		else
			usleep_range(wait, wait << 1);
		wait <<= 1;
	}

Because there is no cap on wait, it will grow exponentially until it
reaches ~83 seconds, leading to a massive sleep interval that could
break hardware responsiveness and stall the kernel thread.

Additionally, when xe_pcode_request() polls with preemption disabled
(atomic = true), wait can grow up to ~32ms, leading to udelay(32000),
which exceeds architecture limits for udelay() and risks soft lockups.

Also, since slept adds wait before the actual sleep duration is completed,
and wait is then doubled, does this miscalculate the elapsed time and
effectively halve the actual timeout?

[ ... ]

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260814171801.13347-1-michal.wajdeczko@intel.com?part=28

  reply	other threads:[~2026-08-14 17:33 UTC|newest]

Thread overview: 51+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-14 17:17 [PATCH v5 00/33] drm/xe: Add structured SIGID error logging infrastructure Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 01/33] drm/xe: Introduce xe_any helpers Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 02/33] drm/xe/log: Add structured SIGID error logging infrastructure Michal Wajdeczko
2026-08-14 17:31   ` Nilawar, Badal
2026-08-14 17:17 ` [PATCH v5 03/33] drm/xe/log: Introduce structured component/location identifiers Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 04/33] drm/xe/log: Add component/location decorations to dmesg Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 05/33] drm/xe/log: Add SIGID log helpers for severity Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 06/33] drm/xe/log: Add SIGID log helpers for location Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 07/33] drm/xe/log: Add SIGID log helpers for location & severity Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 08/33] drm/xe/log: Add SIGID log helpers for components Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 09/33] drm/xe/log: Add SIGID log helpers for component & severity Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 10/33] drm/xe/log: Add SIGID log helpers for errno-only Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 11/33] drm/xe/log: Index all SIGID printk messages Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 12/33] drm/xe/log: Add hardware error signatures Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 13/33] drm/xe/log: Extend components list with hardware items Michal Wajdeczko
2026-08-14 18:39   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 14/33] drm/xe/ras: Check RAS and LOG component definitions Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 15/33] drm/xe/kunit: Setup driver data in the test device Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 16/33] drm/xe/tests: Add Kunit tests for xe_log Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 17/33] drm/xe/tests: Add kunit tests for xe_any Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 18/33] drm/xe: Report 'probe blocked' status using SIGID Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 19/33] drm/xe: Report all probe errors " Michal Wajdeczko
2026-08-14 17:31   ` sashiko-bot
2026-08-14 17:17 ` [PATCH v5 20/33] drm/xe/survivability: Report 'boot status' " Michal Wajdeczko
2026-08-14 17:26   ` sashiko-bot
2026-08-14 18:41   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 21/33] drm/xe/survivability: Report sysfs failure in one place Michal Wajdeczko
2026-08-14 18:43   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 22/33] drm/xe/survivability: Report 'boot status' on sysfs failure Michal Wajdeczko
2026-08-14 17:39   ` sashiko-bot
2026-08-14 18:45   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 23/33] drm/xe/survivability: Report 'Boot Mode enabled' status using SIGID Michal Wajdeczko
2026-08-14 18:47   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 24/33] drm/xe/survivability: Report 'Runtime " Michal Wajdeczko
2026-08-14 18:49   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 25/33] drm/xe: Report 'device wedged' errors " Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 26/33] drm/xe/pcode: Report 'Mailbox failed' error " Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 27/33] drm/xe/pcode: Report 'timeout, retrying' " Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 28/33] drm/xe/pcode: Report 'initialization timedout' " Michal Wajdeczko
2026-08-14 17:33   ` sashiko-bot [this message]
2026-08-14 17:17 ` [PATCH v5 29/33] drm/xe/guc: Report 'GuC mmio' errors " Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 30/33] drm/xe/gt: Report 'reset failed' " Michal Wajdeczko
2026-08-14 17:17 ` [PATCH v5 31/33] drm/xe/gt: Report 'Fault response' pagefault error " Michal Wajdeczko
2026-08-14 18:50   ` Rodrigo Vivi
2026-08-14 17:17 ` [PATCH v5 32/33] drm/xe/gt: Report 'Queue full' " Michal Wajdeczko
2026-08-14 17:34   ` sashiko-bot
2026-08-14 17:17 ` [PATCH v5 33/33] drm/xe/pci: Report 'cannot re-enable' " Michal Wajdeczko
2026-08-14 17:25 ` ✗ CI.checkpatch: warning for drm/xe: Add structured SIGID error logging infrastructure (rev5) Patchwork
2026-08-14 17:27 ` ✓ CI.KUnit: success " Patchwork
2026-08-14 18:13 ` ✓ Xe.CI.BAT: " Patchwork
2026-08-14 18:54 ` [PATCH v5 00/33] drm/xe: Add structured SIGID error logging infrastructure Rodrigo Vivi

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260814173346.9E0F51F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=michal.wajdeczko@intel.com \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox