dri-devel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Boris Brezillon <boris.brezillon@collabora.com>
To: "Adrián Larumbe" <adrian.larumbe@collabora.com>
Cc: Rob Herring <robh@kernel.org>,
	Steven Price <steven.price@arm.com>,
	Maarten Lankhorst <maarten.lankhorst@linux.intel.com>,
	Maxime Ripard <mripard@kernel.org>,
	Thomas Zimmermann <tzimmermann@suse.de>,
	David Airlie <airlied@gmail.com>, Simona Vetter <simona@ffwll.ch>,
	Faith Ekstrand <faith.ekstrand@collabora.com>,
	"Marty E. Plummer" <hanetzer@startmail.com>,
	Tomeu Vizoso <tomeu@tomeuvizoso.net>,
	Eric Anholt <eric@anholt.net>,
	Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com>,
	Robin Murphy <robin.murphy@arm.com>,
	Philipp Zabel <p.zabel@pengutronix.de>,
	dri-devel@lists.freedesktop.org, linux-kernel@vger.kernel.org,
	Collabora Kernel Team <kernel@collabora.com>,
	Neil Armstrong <neil.armstrong@linaro.org>
Subject: Re: [PATCH v7 08/17] drm/panfrost: Split subsystem init/reset from interrupt enablement
Date: Wed, 2 Sep 2026 18:05:06 +0200	[thread overview]
Message-ID: <20260902180506.66fd5d00@fedora-21.home> (raw)
In-Reply-To: <aphDm1vDtk2Pei3i@sobremesa>

On Wed, 2 Sep 2026 16:41:40 +0100
Adrián Larumbe <adrian.larumbe@collabora.com> wrote:

> On 01.09.2026 15:08, Boris Brezillon wrote:
> > On Fri, 28 Aug 2026 21:56:48 +0100
> > Adrián Larumbe <adrian.larumbe@collabora.com> wrote:
> >   
> > > Because MMU interrupts are only enabled when the device is reset, it
> > > happened that after DRM device registration, the very first job targeting
> > > the tiler heap BO would always time out. The reason is the reset sequence
> > > is only part of PM runtime resume, which is not called explicitly at driver
> > > probe time, and an actual reset work item manually triggered after a HW
> > > error.
> > > 
> > > I have attempted a somewhat drastic solution, which is completely
> > > decoupling GPU/MMU/JM subsystem initialisation and reset from interrupt
> > > enablement, so that we can handle IRQ toggling a bit more flexibly.
> > > 
> > > To this end:
> > > - Ensure every subsystem with its own IRQ has an 'enable interrupts'
> > > method, and that it doesn't enable them anywhere else.
> > > - Force IRQ masking at MMU reset time. Up until, now, panfrost_mmu_reset()
> > > was clearing the MMU IRQ suspension bit, but at no point that is set during
> > > the reset sequence.
> > > 
> > > Then manually enable all interrupts when the device is fully initialised at
> > > probe time, right before DRM device registration, or after the reset
> > > sequence is complete. Also disable all interrupts at device remove time,
> > > so that their IRQs can be sync'ed right before tearing the device down.
> > > 
> > > Fixes: 635430797d3f ("drm/panfrost: Rework runtime PM initialization")
> > > Fixes: 876b15d2c88d ("drm/panfrost: Fix module unload")
> > > Signed-off-by: Adrián Larumbe <adrian.larumbe@collabora.com>
> > > ---
> > >  drivers/gpu/drm/panfrost/panfrost_device.c | 40 ++++++++++++++++++++++--------
> > >  drivers/gpu/drm/panfrost/panfrost_device.h |  3 ++-
> > >  drivers/gpu/drm/panfrost/panfrost_gpu.c    | 19 ++++++++------
> > >  drivers/gpu/drm/panfrost/panfrost_gpu.h    |  2 ++
> > >  drivers/gpu/drm/panfrost/panfrost_job.c    |  7 +++---
> > >  drivers/gpu/drm/panfrost/panfrost_mmu.c    |  9 +++++--
> > >  drivers/gpu/drm/panfrost/panfrost_mmu.h    |  2 ++
> > >  7 files changed, 56 insertions(+), 26 deletions(-)
> > > 
> > > diff --git a/drivers/gpu/drm/panfrost/panfrost_device.c b/drivers/gpu/drm/panfrost/panfrost_device.c
> > > index 9e02fb5f73c8..99f7da2180f9 100644
> > > --- a/drivers/gpu/drm/panfrost/panfrost_device.c
> > > +++ b/drivers/gpu/drm/panfrost/panfrost_device.c
> > > @@ -226,6 +226,27 @@ static int panfrost_pm_domain_init(struct panfrost_device *pfdev)
> > >  	return err;
> > >  }
> > >  
> > > +void panfrost_device_enable_int(struct panfrost_device *pfdev)
> > > +{
> > > +	panfrost_gpu_enable_interrupts(pfdev);
> > > +	panfrost_mmu_enable_interrupts(pfdev);
> > > +	panfrost_jm_enable_interrupts(pfdev);
> > > +}
> > > +
> > > +static void panfrost_device_enable_hw(struct panfrost_device *pfdev)
> > > +{
> > > +	panfrost_device_enable_int(pfdev);
> > > +	panfrost_devfreq_resume(pfdev);
> > > +}
> > > +
> > > +static void panfrost_device_disable_hw(struct panfrost_device *pfdev)
> > > +{
> > > +	panfrost_devfreq_suspend(pfdev);
> > > +	panfrost_jm_suspend_irq(pfdev);
> > > +	panfrost_mmu_suspend_irq(pfdev);
> > > +	panfrost_gpu_suspend_irq(pfdev);  
> > 
> > Hm, I think I'd prefer if those suspend/resume_irq() were hidden in
> > some subcomponent panfrost_<subcomp>_suspend,resume() helpers. And
> > then we just have to resume/suspend component in the right order
> > instead of treating IRQs as a standalone object (enabling/disabling
> > only makes sense if the subcomponent handling those interrupts is
> > resumed/suspended).  
> 
> I thought it would only make sense to enable interupts for a given subsystem
> when all the other subsystems are also resumed or initialised. This was prompted
> by Sashiko warning of the possibility of spurious interrupts causing a handler
> to be run when one of the subsystems it touches on hasn't yet been initialised.

Well, in practice things tend to be well isolated, for instance, an
MMU IRQ should be processed entirely inside panfrost_mmu.c, with no
particular interaction with the other subsystems. So, if an MMU
interrupt fires before, say, the JM subsystem is up and running, that
shouldn't be a problem. In panthor, we have a few cases where events
get propagated between subsystems, and for those we have some
is_initialized checks. I'm not sure this applies to panfrost though.

The other advantage with this approach is that it's one step towards a
better subsystem isolation like we have in panthor, where subsystems
only see their internal state/data plus the general state exposed by
panthor_device, instead of having everything in panfrost_device, and
everyone having the ability to modify/check the state of other
subsystems. panfrost_device.c then just acts as a glue layer that knows
about the order things should be executed in, but doesn't have all the
internal details about subsystem initialization/teardown steps.

  reply	other threads:[~2026-09-02 16:05 UTC|newest]

Thread overview: 58+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-28 20:56 [PATCH v7 00/17] Collection of fixes for Panfrost: Perfcnt, RPM, refactorings Adrián Larumbe
2026-08-28 20:56 ` [PATCH v7 01/17] drm/panfrost: Move shrinker initialization and unplug one level down Adrián Larumbe
2026-09-01 11:28   ` Boris Brezillon
2026-09-02 15:36     ` Adrián Larumbe
2026-08-28 20:56 ` [PATCH v7 02/17] drm/panfrost: Move all DRM device initialisation into device_init() Adrián Larumbe
2026-09-01 11:49   ` Boris Brezillon
2026-09-02 15:38     ` Adrián Larumbe
2026-09-02 15:50       ` Boris Brezillon
2026-08-28 20:56 ` [PATCH v7 03/17] drm/panfrost: Move lock and modparam initialisations into their subsystems Adrián Larumbe
2026-08-28 21:14   ` sashiko-bot
2026-09-01 12:10   ` Boris Brezillon
2026-08-28 20:56 ` [PATCH v7 04/17] drm/panfrost: Move debugfs initialisation to relevant subsystems Adrián Larumbe
2026-09-01 12:30   ` Boris Brezillon
2026-09-02 15:40     ` Adrián Larumbe
2026-08-28 20:56 ` [PATCH v7 05/17] drm/panfrost: Skip NULL checks for clock enable/disabling Adrián Larumbe
2026-09-01 12:31   ` Boris Brezillon
2026-08-28 20:56 ` [PATCH v7 06/17] drm/panfrost: Consolidate device clock management and reset Adrián Larumbe
2026-09-01 12:38   ` Boris Brezillon
2026-09-02 15:41     ` Adrián Larumbe
2026-08-28 20:56 ` [PATCH v7 07/17] drm/panfrost: Stop all jobs before commencing device teardown Adrián Larumbe
2026-08-28 21:16   ` sashiko-bot
2026-09-01 12:58   ` Boris Brezillon
2026-08-28 20:56 ` [PATCH v7 08/17] drm/panfrost: Split subsystem init/reset from interrupt enablement Adrián Larumbe
2026-08-28 21:11   ` sashiko-bot
2026-09-01 13:08   ` Boris Brezillon
2026-09-02 15:41     ` Adrián Larumbe
2026-09-02 16:05       ` Boris Brezillon [this message]
2026-08-28 20:56 ` [PATCH v7 09/17] drm/panfrost: Fix PM refcnt and autosuspend issues at device probe/remove Adrián Larumbe
2026-08-28 21:09   ` sashiko-bot
2026-09-01 13:18   ` Boris Brezillon
2026-09-02 15:42     ` Adrián Larumbe
2026-09-02 16:14       ` Boris Brezillon
2026-08-28 20:56 ` [PATCH v7 10/17] drm/panfrost: Add warning messages to fatal error conditions Adrián Larumbe
2026-08-28 21:10   ` sashiko-bot
2026-09-01 13:20   ` Boris Brezillon
2026-08-28 20:56 ` [PATCH v7 11/17] drm/panfrost: Add debugfs knob for manually triggering a GPU reset Adrián Larumbe
2026-08-28 21:12   ` sashiko-bot
2026-09-01 13:27   ` Boris Brezillon
2026-09-02 15:42     ` Adrián Larumbe
2026-09-02 16:23       ` Boris Brezillon
2026-08-28 20:56 ` [PATCH v7 12/17] drm/panfrost: Move perfcnt GPU disable sequence into a helper Adrián Larumbe
2026-08-28 20:56 ` [PATCH v7 13/17] drm/panfrost: Skip cache flush/invalidate when enabling perfcnt Adrián Larumbe
2026-09-01 13:32   ` Boris Brezillon
2026-09-02 15:43     ` Adrián Larumbe
2026-09-02 16:29       ` Boris Brezillon
2026-08-28 20:56 ` [PATCH v7 14/17] drm/panfrost: Avoid cache flush after perfcnt sample in fully coherent systems Adrián Larumbe
2026-08-28 21:14   ` sashiko-bot
2026-09-01 13:37   ` Boris Brezillon
2026-09-02 15:44     ` Adrián Larumbe
2026-09-02 16:33       ` Boris Brezillon
2026-09-02 16:34   ` Boris Brezillon
2026-08-28 20:56 ` [PATCH v7 15/17] drm/panfrost: Introduce a reset lock Adrián Larumbe
2026-08-28 20:56 ` [PATCH v7 16/17] drm/panfrost: Fix races between perfcnt and reset sequence Adrián Larumbe
2026-08-28 21:17   ` sashiko-bot
2026-09-01 14:03   ` Boris Brezillon
2026-09-02 15:45     ` Adrián Larumbe
2026-09-02 16:51       ` Boris Brezillon
2026-08-28 20:56 ` [PATCH v7 17/17] drm/panfrost: Bump driver minor to reflect new DUMP IOCTL req field Adrián Larumbe

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260902180506.66fd5d00@fedora-21.home \
    --to=boris.brezillon@collabora.com \
    --cc=adrian.larumbe@collabora.com \
    --cc=airlied@gmail.com \
    --cc=alyssa.rosenzweig@collabora.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=eric@anholt.net \
    --cc=faith.ekstrand@collabora.com \
    --cc=hanetzer@startmail.com \
    --cc=kernel@collabora.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=maarten.lankhorst@linux.intel.com \
    --cc=mripard@kernel.org \
    --cc=neil.armstrong@linaro.org \
    --cc=p.zabel@pengutronix.de \
    --cc=robh@kernel.org \
    --cc=robin.murphy@arm.com \
    --cc=simona@ffwll.ch \
    --cc=steven.price@arm.com \
    --cc=tomeu@tomeuvizoso.net \
    --cc=tzimmermann@suse.de \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox