All of lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Adrián Larumbe" <adrian.larumbe@collabora.com>
Cc: dri-devel@lists.freedesktop.org
Subject: Re: [PATCH v4 09/13] drm/panfrost: Add debugfs knob for manually triggering a GPU reset
Date: Wed, 29 Jul 2026 03:07:28 +0000	[thread overview]
Message-ID: <20260729030728.B3F6E1F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260729-claude-fixes-v4-9-01968f2ec77a@collabora.com>

Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:

New issues:
- [High] Debugfs `reset_set` drops the PM runtime reference on wait timeout or signal before the reset work completes, leading to hardware access while the device is suspended.
- [Medium] Incorrect return value logic in `reset_set` masks signal interruptions and suspended PM states by unconditionally returning `-ETIMEDOUT`.

Pre-existing issues:
- [High] The GPU reset sequence drops PM runtime references from stuck jobs before accessing hardware registers, risking a crash if the device suspends concurrently.
--

--- Patch [9]: [PATCH v4 09/13] drm/panfrost: Add debugfs knob for manually triggering a GPU reset ---
commit f19b1f8bf36d92ef2c4bcc1bdcf953c6bd32d4fd
Author: Adrián Larumbe <adrian.larumbe@collabora.com>

drm/panfrost: Add debugfs knob for manually triggering a GPU reset

This commit adds a debugfs file to manually trigger a GPU reset. This is
intended to help test potential race conditions between the GPU reset
sequence and other parts of the code accessing hardware registers.

> diff --git a/drivers/gpu/drm/panfrost/panfrost_device.c b/drivers/gpu/drm/panfrost/panfrost_device.c
> index a51f1ef8563ce..c0334b72c70b3 100644
> --- a/drivers/gpu/drm/panfrost/panfrost_device.c
> +++ b/drivers/gpu/drm/panfrost/panfrost_device.c

[ ... ]

> @@ -629,3 +630,43 @@ void panfrost_device_reset(struct panfrost_device *pfdev, bool enable_job_int)
>  	if (!panfrost_device_started(pfdev) && !ret)
>  		pfdev->gpu_started = true;
>  }
> +
> +#ifdef CONFIG_DEBUG_FS
> +static int reset_get(void *data, u64 *val)
> +{

[ ... ]

> +static int reset_set(void *data, u64 val)
> +{
> +	struct panfrost_device *pfdev =
> +		container_of(data, struct panfrost_device, base);
> +	int ret;
> +
> +	ret = pm_runtime_get_if_in_use(pfdev->base.dev);
> +
> +	if (ret > 0) {
> +		panfrost_device_schedule_reset(pfdev);
> +		ret = wait_event_interruptible_timeout(pfdev->reset.wait,
> +						       !atomic_read(&pfdev->reset.pending),
> +						       msecs_to_jiffies(60));
> +		pm_runtime_put(pfdev->base.dev);

[Severity: High]
If the wait is interrupted by a signal (like Ctrl-C) or times out, does
this immediately drop the PM usage count to 0 and queue a suspend request?

Since the scheduled reset work continues executing asynchronously without a
PM reference, could it access the GPU registers via panfrost_device_reset()
while the device is already powered down, leading to a system crash?

> +	}
> +
> +	return (ret > 0) ? 0 : -ETIMEDOUT;

[Severity: Medium]
Does this ternary return statement mask signal interruptions and PM runtime
errors?

If wait_event_interruptible_timeout() is interrupted, it returns -ERESTARTSYS.
Converting this negative error code to -ETIMEDOUT prevents the kernel from
correctly restarting the syscall or returning -EINTR to user space. It also
masks legitimate errors if pm_runtime_get_if_in_use() returned 0 initially.

> +}
> +
> +DEFINE_DEBUGFS_ATTRIBUTE(panfrost_reset_debugfs_fops,
> +			 reset_get, reset_set,
> +			 "0x%08llx\n");

[ ... ]

[Severity: High]
This isn't a bug introduced by this patch, but while reviewing the reset
sequences, I noticed a potential race condition in the existing
panfrost_reset() handler in drivers/gpu/drm/panfrost/panfrost_job.c:

panfrost_job.c:panfrost_reset() {
    ...
    for (i = 0; i < NUM_JOB_SLOTS; i++) {
        if (pfdev->jobs[i][0])
            pm_runtime_put_noidle(pfdev->base.dev);
        ...
    }
    memset(pfdev->jobs, 0, sizeof(pfdev->jobs));
    ...
    panfrost_device_reset(pfdev, true);
}

Does dropping the PM runtime references for all jobs before executing the
hardware reset risk a crash?

By dropping the usage count and clearing the jobs array, panfrost_jm_is_idle()
will return true. If a concurrent thread had triggered a PM autosuspend timer,
could the device be suspended before panfrost_device_reset() accesses the GPU
registers? Should panfrost_reset() acquire its own PM reference before
dropping the job references?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260729-claude-fixes-v4-0-01968f2ec77a@collabora.com?part=9

  reply	other threads:[~2026-07-29  3:07 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-29  2:54 [PATCH v4 00/13] RPM, perfcnt and other minor fixes for Panfrost Adrián Larumbe
2026-07-29  2:54 ` [PATCH v4 01/13] drm/panfrost: Check another bo field for cache option query Adrián Larumbe
2026-07-29  2:54 ` [PATCH v4 02/13] drm/panfrost: Prevent division by 0 Adrián Larumbe
2026-07-29  3:00   ` sashiko-bot
2026-07-29  2:54 ` [PATCH v4 03/13] drm/panfrost: Remove unnecessary header file include Adrián Larumbe
2026-07-29  2:54 ` [PATCH v4 04/13] drm/panfrost: Move shrinker initialization and unplug one level down Adrián Larumbe
2026-07-29  2:54 ` [PATCH v4 05/13] drm/panfrost: Move all device power up and down into RPM callbacks Adrián Larumbe
2026-07-29  3:08   ` sashiko-bot
2026-07-29  8:37   ` Philipp Zabel
2026-07-29  2:54 ` [PATCH v4 06/13] drm/panfrost: Explicitly enable MMU interrupts at device init Adrián Larumbe
2026-07-29  2:54 ` [PATCH v4 07/13] drm/panfrost: Sync with IRQ before MMU disable and reset Adrián Larumbe
2026-07-29  3:12   ` sashiko-bot
2026-07-29  2:54 ` [PATCH v4 08/13] drm/panfrost: Rewire reset sequence to avoid concurrent attempts Adrián Larumbe
2026-07-29  3:19   ` sashiko-bot
2026-07-29  2:54 ` [PATCH v4 09/13] drm/panfrost: Add debugfs knob for manually triggering a GPU reset Adrián Larumbe
2026-07-29  3:07   ` sashiko-bot [this message]
2026-07-29  2:54 ` [PATCH v4 10/13] drm/panfrost: Move perfcnt GPU disable sequence into a helper Adrián Larumbe
2026-07-29  3:03   ` sashiko-bot
2026-07-29  2:54 ` [PATCH v4 11/13] drm/panfrost: Introduce a reset lock Adrián Larumbe
2026-07-29  3:08   ` sashiko-bot
2026-07-29  2:54 ` [PATCH v4 12/13] drm/panfrost: Fix races between perfcnt and reset sequence Adrián Larumbe
2026-07-29  3:06   ` sashiko-bot
2026-07-29  2:54 ` [PATCH v4 13/13] drm/panfrost: Bump driver minor to reflect new DUMP IOCTL req field Adrián Larumbe
2026-07-29  3:08   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260729030728.B3F6E1F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=adrian.larumbe@collabora.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.