dri-devel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Adrián Larumbe" <adrian.larumbe@collabora.com>
Cc: dri-devel@lists.freedesktop.org
Subject: Re: [PATCH v6 13/16] drm/panfrost: Avoid cache flush after perfcnt sample in fully coherent systems
Date: Wed, 26 Aug 2026 21:36:58 +0000	[thread overview]
Message-ID: <20260826213659.445501F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260826-claude-fixes-v6-13-78119af409e7@collabora.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:

Pre-existing issues:
- [High] Missing cancellation and synchronization of asynchronous GPU tasks after a signal or timeout leads to software race conditions, hardware state corruption, and GPU page faults.
--

commit 5c3d775056a5b2e3ebdba73f1d3b9858fe3716ba
Author: Adrián Larumbe <adrian.larumbe@collabora.com>

drm/panfrost: Avoid cache flush after perfcnt sample in fully coherent systems

This commit avoids flushing the GPU caches after a perfcnt sample is made
available if the GPU and CPU are fully coherent.

> diff --git a/drivers/gpu/drm/panfrost/panfrost_perfcnt.c b/drivers/gpu/drm/panfrost/panfrost_perfcnt.c
> index ffc77121070e2..b3f71d7fd82a9 100644
> --- a/drivers/gpu/drm/panfrost/panfrost_perfcnt.c
> +++ b/drivers/gpu/drm/panfrost/panfrost_perfcnt.c
> @@ -52,7 +52,10 @@ void panfrost_perfcnt_clean_cache_done(struct panfrost_device *pfdev)
>  
>  void panfrost_perfcnt_sample_done(struct panfrost_device *pfdev)
>  {
> -	gpu_write(pfdev, GPU_CMD, GPU_CMD_CLEAN_CACHES);
> +	if (pfdev->features.selected_coherency != COHERENCY_ACE)
> +		gpu_write(pfdev, GPU_CMD, GPU_CMD_CLEAN_CACHES);
> +	else
> +		complete(&pfdev->perfcnt->dump_comp);
>  }
>  
>  static int panfrost_perfcnt_dump_locked(struct panfrost_device *pfdev)

[Severity: High]
This is a pre-existing issue, but does this interruptible wait in
panfrost_perfcnt_dump_locked() leave the GPU hardware in a busy state if it
aborts early?

When an unprivileged user issues the DRM_IOCTL_PANFROST_PERFCNT_DUMP ioctl,
the kernel waits for the sample to complete:

panfrost_perfcnt_dump_locked() {
    ...
    ret = wait_for_completion_interruptible_timeout(&pfdev->perfcnt->dump_comp,
                                                    msecs_to_jiffies(1000));
    ...
}

If the user process is interrupted by a signal, this wait returns early,
leaving the GPU_CMD_PERFCNT_SAMPLE command executing asynchronously on the
hardware.

If the syscall is restarted by the kernel, the driver re-issues the
GPU_CMD_PERFCNT_SAMPLE command:

panfrost_perfcnt_dump_locked() {
    ...
    reinit_completion(&pfdev->perfcnt->dump_comp);
    ...
    gpu_write(pfdev, GPU_CMD, GPU_CMD_PERFCNT_SAMPLE);
    ...
}

This issues a new command while the previous one is still in-flight, which
can corrupt the hardware state machine.

Alternatively, if the process exits instead,
panfrost_perfcnt_disable_locked() unmaps the GPU memory without waiting for
the in-flight sample to finish:

panfrost_perfcnt_disable_locked() {
    ...
    drm_gem_vunmap(&perfcnt->mapping->obj->base.base, &map);
    perfcnt->buf = NULL;
    panfrost_gem_close(&perfcnt->mapping->obj->base.base, file_priv);
    ...
}

When the GPU finally writes the sample data, it hits unmapped memory,
causing a GPU page fault and forcing a GPU hard reset. Could proper
cancellation and synchronization of these asynchronous GPU tasks be added
after a signal or timeout?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260826-claude-fixes-v6-0-78119af409e7@collabora.com?part=13

  reply	other threads:[~2026-08-26 21:37 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-26 21:18 [PATCH v6 00/16] Collection of fixes for Panfrost: Perfcnt, RPM, refactorings Adrián Larumbe
2026-08-26 21:18 ` [PATCH v6 01/16] drm/panfrost: Move shrinker initialization and unplug one level down Adrián Larumbe
2026-08-26 21:34   ` sashiko-bot
2026-08-26 21:18 ` [PATCH v6 02/16] drm/panfrost: Move all DRM device initialisation into device_init() Adrián Larumbe
2026-08-26 21:18 ` [PATCH v6 03/16] drm/panfrost: Move lock and modparam initialisations into their subsystems Adrián Larumbe
2026-08-26 21:18 ` [PATCH v6 04/16] drm/panfrost: Move debugfs initialisation to relevant subsystems Adrián Larumbe
2026-08-26 21:18 ` [PATCH v6 05/16] drm/panfrost: Skip NULL checks for clock enable/disabling Adrián Larumbe
2026-08-26 21:18 ` [PATCH v6 06/16] drm/panfrost: Consolidate device clock management and reset Adrián Larumbe
2026-08-26 21:18 ` [PATCH v6 07/16] drm/panfrost: Split subsystem init/reset from interrupt enablement Adrián Larumbe
2026-08-26 21:34   ` sashiko-bot
2026-08-26 21:18 ` [PATCH v6 08/16] drm/panfrost: Fix PM refcnt and autosuspend issues at device probe/remove Adrián Larumbe
2026-08-26 21:31   ` sashiko-bot
2026-08-26 21:18 ` [PATCH v6 09/16] drm/panfrost: Add warning messages to fatal error conditions Adrián Larumbe
2026-08-26 21:27   ` sashiko-bot
2026-08-26 21:18 ` [PATCH v6 10/16] drm/panfrost: Add debugfs knob for manually triggering a GPU reset Adrián Larumbe
2026-08-26 21:34   ` sashiko-bot
2026-08-26 21:18 ` [PATCH v6 11/16] drm/panfrost: Move perfcnt GPU disable sequence into a helper Adrián Larumbe
2026-08-26 21:18 ` [PATCH v6 12/16] drm/panfrost: Skip cache flush/invalidate when enabling perfcnt Adrián Larumbe
2026-08-26 21:18 ` [PATCH v6 13/16] drm/panfrost: Avoid cache flush after perfcnt sample in fully coherent systems Adrián Larumbe
2026-08-26 21:36   ` sashiko-bot [this message]
2026-08-26 21:18 ` [PATCH v6 14/16] drm/panfrost: Introduce a reset lock Adrián Larumbe
2026-08-26 21:18 ` [PATCH v6 15/16] drm/panfrost: Fix races between perfcnt and reset sequence Adrián Larumbe
2026-08-26 21:35   ` sashiko-bot
2026-08-26 21:18 ` [PATCH v6 16/16] drm/panfrost: Bump driver minor to reflect new DUMP IOCTL req field Adrián Larumbe
2026-08-26 21:37   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260826213659.445501F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=adrian.larumbe@collabora.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox