All of lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Adrián Larumbe" <adrian.larumbe@collabora.com>
Cc: dri-devel@lists.freedesktop.org
Subject: Re: [PATCH v4 08/13] drm/panfrost: Rewire reset sequence to avoid concurrent attempts
Date: Wed, 29 Jul 2026 03:19:33 +0000	[thread overview]
Message-ID: <20260729031933.CD7361F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260729-claude-fixes-v4-8-01968f2ec77a@collabora.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] Removing the synchronous `atomic_set(&pfdev->reset.pending, 1)` from `panfrost_job_handle_err` opens a vulnerable window where new jobs are erroneously submitted to faulted hardware.
- [Medium] The introduced `pfdev->reset.wait` wait queue is never waited on, rendering it dead code and breaking the intended synchronization.
--

--- Patch [8]: [PATCH v4 08/13] drm/panfrost: Rewire reset sequence to avoid concurrent attempts ---
commit c66abdaa510df5e2a3ec11be0069267b5ffde774
Author: Adrián Larumbe <adrian.larumbe@collabora.com>

drm/panfrost: Rewire reset sequence to avoid concurrent attempts

This rewires the reset sequence to use a single reset thread and a wait queue
for other threads wanting to trigger a reset.

> diff --git a/drivers/gpu/drm/panfrost/panfrost_device.h b/drivers/gpu/drm/panfrost/panfrost_device.h
> index eeca57efcbec0..f0acfc634cd47 100644
> --- a/drivers/gpu/drm/panfrost/panfrost_device.h
> +++ b/drivers/gpu/drm/panfrost/panfrost_device.h
[ ... ]
> @@ -341,8 +342,8 @@ bool panfrost_exception_needs_reset(const struct panfrost_device *pfdev,
>  static inline void
>  panfrost_device_schedule_reset(struct panfrost_device *pfdev)
>  {
> -	atomic_set(&pfdev->reset.pending, 1);
> -	queue_work(pfdev->reset.wq, &pfdev->reset.work);
> +	if (!atomic_cmpxchg(&pfdev->reset.pending, 0, 1))
> +		queue_work(pfdev->reset.wq, &pfdev->reset.work);
>  }

[Severity: Medium]
The commit message states that other threads will wait on the ongoing reset
to finish, and a wait queue is added and woken up in panfrost_reset().
However, there do not appear to be any wait_event() calls added for
pfdev->reset.wait in the codebase.

When the atomic_cmpxchg() fails here, the caller returns immediately instead
of blocking on the ongoing reset. Are we missing the wait logic for
concurrent reset requests?

> diff --git a/drivers/gpu/drm/panfrost/panfrost_job.c b/drivers/gpu/drm/panfrost/panfrost_job.c
> index bda1494eb430c..ed8a1fabd47b4 100644
> --- a/drivers/gpu/drm/panfrost/panfrost_job.c
> +++ b/drivers/gpu/drm/panfrost/panfrost_job.c
[ ... ]
> @@ -499,10 +499,8 @@ static void panfrost_job_handle_err(struct panfrost_device *pfdev,
>  
>  	pm_runtime_put_autosuspend(pfdev->base.dev);
>  
> -	if (panfrost_exception_needs_reset(pfdev, js_status)) {
> -		atomic_set(&pfdev->reset.pending, 1);
> +	if (panfrost_exception_needs_reset(pfdev, js_status))
>  		drm_sched_fault(&pfdev->js->queue[js].sched);
> -	}
>  }

[Severity: High]
Does removing the synchronous atomic_set() here create a vulnerability window?

By delegating this to the async timeout worker via drm_sched_fault(),
reset.pending remains 0 when panfrost_job_handle_err() returns.

If reset.pending is still 0, will panfrost_jm_handle_irq() incorrectly
requeue jobs onto the faulted hardware?

panfrost_jm_handle_irq() {
    ...
    } else if (!atomic_read(&pfdev->reset.pending)) {
        /* Requeue the job we removed if no reset is pending */
        job_write(pfdev, JS_COMMAND_NEXT(j), JS_COMMAND_START);
    }
}

Similarly, concurrent job submissions from the scheduler in
panfrost_job_hw_submit() might also incorrectly proceed because they check
if (!atomic_read(&pfdev->reset.pending)).

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260729-claude-fixes-v4-0-01968f2ec77a@collabora.com?part=8

  reply	other threads:[~2026-07-29  3:19 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-29  2:54 [PATCH v4 00/13] RPM, perfcnt and other minor fixes for Panfrost Adrián Larumbe
2026-07-29  2:54 ` [PATCH v4 01/13] drm/panfrost: Check another bo field for cache option query Adrián Larumbe
2026-07-29  2:54 ` [PATCH v4 02/13] drm/panfrost: Prevent division by 0 Adrián Larumbe
2026-07-29  3:00   ` sashiko-bot
2026-07-29  2:54 ` [PATCH v4 03/13] drm/panfrost: Remove unnecessary header file include Adrián Larumbe
2026-07-29  2:54 ` [PATCH v4 04/13] drm/panfrost: Move shrinker initialization and unplug one level down Adrián Larumbe
2026-07-29  2:54 ` [PATCH v4 05/13] drm/panfrost: Move all device power up and down into RPM callbacks Adrián Larumbe
2026-07-29  3:08   ` sashiko-bot
2026-07-29  8:37   ` Philipp Zabel
2026-07-29  2:54 ` [PATCH v4 06/13] drm/panfrost: Explicitly enable MMU interrupts at device init Adrián Larumbe
2026-07-29  2:54 ` [PATCH v4 07/13] drm/panfrost: Sync with IRQ before MMU disable and reset Adrián Larumbe
2026-07-29  3:12   ` sashiko-bot
2026-07-29  2:54 ` [PATCH v4 08/13] drm/panfrost: Rewire reset sequence to avoid concurrent attempts Adrián Larumbe
2026-07-29  3:19   ` sashiko-bot [this message]
2026-07-29  2:54 ` [PATCH v4 09/13] drm/panfrost: Add debugfs knob for manually triggering a GPU reset Adrián Larumbe
2026-07-29  3:07   ` sashiko-bot
2026-07-29  2:54 ` [PATCH v4 10/13] drm/panfrost: Move perfcnt GPU disable sequence into a helper Adrián Larumbe
2026-07-29  3:03   ` sashiko-bot
2026-07-29  2:54 ` [PATCH v4 11/13] drm/panfrost: Introduce a reset lock Adrián Larumbe
2026-07-29  3:08   ` sashiko-bot
2026-07-29  2:54 ` [PATCH v4 12/13] drm/panfrost: Fix races between perfcnt and reset sequence Adrián Larumbe
2026-07-29  3:06   ` sashiko-bot
2026-07-29  2:54 ` [PATCH v4 13/13] drm/panfrost: Bump driver minor to reflect new DUMP IOCTL req field Adrián Larumbe
2026-07-29  3:08   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260729031933.CD7361F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=adrian.larumbe@collabora.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.