* [PATCH 0/2] drm/v3d: Skip perfmon serialization when it cannot be observed
@ 2026-09-21 21:24 Maíra Canal
2026-09-21 21:24 ` [PATCH 1/2] drm/v3d: Skip perfmon serialization while the device has no perfmon Maíra Canal
` (2 more replies)
0 siblings, 3 replies; 7+ messages in thread
From: Maíra Canal @ 2026-09-21 21:24 UTC (permalink / raw)
To: Melissa Wen, Iago Toral, David Airlie, Simona Vetter
Cc: kernel-dev, dri-devel, Maíra Canal
The V3D core has a single set of performance counters. So that a perfmon
only counts the jobs it is attached to, a job carrying a non-global
perfmon waits for every job still in flight, and while it is in flight
every later job waits for it. Knowing what is in flight requires every
job to merge its finished fence into a per-queue accumulator.
That accumulator is maintained unconditionally, so a client pays a fence
merge on every job even when no dependency can ever be built out of it.
This series turns the bookkeeping off in the two cases where that holds:
1. No perfmon is alive on the device, which is the common case since
perfmons only exist while userspace runs a performance query;
2. A global perfmon is set and the job carries no perfmon of its own.
A global perfmon counts concurrent activity from every job, so
nothing has to wait. Jobs carrying their own perfmon stay tracked,
since the global perfmon may be cleared before they run and only
one of them can be programmed in the HW at a time.
Both share the same trade-off: a job that goes unrecorded never carries a
perfmon, so the only cost is that the first job measured after tracking
resumes may overlap it. Considering the use cases of performance monitors
and that it will only affect the first job, it's a reasonable compromise
to make for the average case.
Best regards,
- Maíra
---
Maíra Canal (2):
drm/v3d: Skip perfmon serialization while the device has no perfmon
drm/v3d: Skip perfmon serialization while a global perfmon is set
drivers/gpu/drm/v3d/v3d_drv.h | 7 +++++++
drivers/gpu/drm/v3d/v3d_perfmon.c | 8 ++++++--
drivers/gpu/drm/v3d/v3d_submit.c | 19 ++++++++++++-------
3 files changed, 25 insertions(+), 9 deletions(-)
---
base-commit: 402491eefc5ea9a40e19b08dccf0c045aec4efc2
change-id: 20260921-v3d-quick-exit-perfmon-serialize-9e9e8ceac587
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH 1/2] drm/v3d: Skip perfmon serialization while the device has no perfmon
2026-09-21 21:24 [PATCH 0/2] drm/v3d: Skip perfmon serialization when it cannot be observed Maíra Canal
@ 2026-09-21 21:24 ` Maíra Canal
2026-09-22 6:35 ` Iago Toral
2026-09-21 21:24 ` [PATCH 2/2] drm/v3d: Skip perfmon serialization while a global perfmon is set Maíra Canal
2026-09-22 7:11 ` [PATCH 0/2] drm/v3d: Skip perfmon serialization when it cannot be observed Iago Toral
2 siblings, 1 reply; 7+ messages in thread
From: Maíra Canal @ 2026-09-21 21:24 UTC (permalink / raw)
To: Melissa Wen, Iago Toral, David Airlie, Simona Vetter
Cc: kernel-dev, dri-devel, Maíra Canal
Every job merges its finished fence into the per-queue accumulator so
that a job carrying a perfmon can later depend on everything still in
flight. This is a (relatively high) cost that every job pays before
it is submitted.
Perfmons only exist while userspace runs a performance query, so the
common case is a client paying that on every job for a dependency that is
never actually added, increasing the submission latency.
Address this situation by counting the number of perfmons alive on the
device and returning early while that count is zero. This means that if
no perfmon exists, v3d_serialize_for_perfmon() will bail out immediately.
Jobs submitted while no perfmon is alive stay unaccounted and may overlap
the first measured job. The count returns to zero whenever the last
perfmon is destroyed, so that window reopens on every measurement cycle.
Closing it would mean merging a fence on every submission for the
lifetime of the device, so it is a reasonable compromise for the average
use case.
Signed-off-by: Maíra Canal <mcanal@igalia.com>
---
drivers/gpu/drm/v3d/v3d_drv.h | 7 +++++++
drivers/gpu/drm/v3d/v3d_perfmon.c | 8 ++++++--
drivers/gpu/drm/v3d/v3d_submit.c | 7 +++++++
3 files changed, 20 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/v3d/v3d_drv.h b/drivers/gpu/drm/v3d/v3d_drv.h
index 79ef20c89ff2..cd544af8d8c8 100644
--- a/drivers/gpu/drm/v3d/v3d_drv.h
+++ b/drivers/gpu/drm/v3d/v3d_drv.h
@@ -84,6 +84,8 @@ struct v3d_queue_state {
* This way, only events related to a specific submission will be counted.
*/
struct v3d_perfmon {
+ struct v3d_dev *v3d;
+
/* Tracks the number of users of the perfmon, when this counter reaches
* zero the perfmon is destroyed.
*/
@@ -184,6 +186,11 @@ struct v3d_dev {
/* Perfmon currently programmed in HW (or NULL if none). */
struct v3d_perfmon *active;
+ /* Number of perfmons alive on this device. Jobs are not
+ * serialized if the number is zero.
+ */
+ atomic_t nperfmons;
+
/* Finished fence of the most recently submitted job that
* opened a serialization window (i.e. a job with a non-global
* perfmon attached).
diff --git a/drivers/gpu/drm/v3d/v3d_perfmon.c b/drivers/gpu/drm/v3d/v3d_perfmon.c
index 07dab7fb3060..359334616bd8 100644
--- a/drivers/gpu/drm/v3d/v3d_perfmon.c
+++ b/drivers/gpu/drm/v3d/v3d_perfmon.c
@@ -217,8 +217,10 @@ void v3d_perfmon_get(struct v3d_perfmon *perfmon)
void v3d_perfmon_put(struct v3d_perfmon *perfmon)
{
- if (perfmon && refcount_dec_and_test(&perfmon->refcnt))
+ if (perfmon && refcount_dec_and_test(&perfmon->refcnt)) {
+ atomic_dec(&perfmon->v3d->perfmon_state.nperfmons);
kfree(perfmon);
+ }
}
static void v3d_perfmon_hw_start(struct v3d_dev *v3d, struct v3d_perfmon *perfmon)
@@ -434,13 +436,15 @@ int v3d_perfmon_create_ioctl(struct drm_device *dev, void *data,
perfmon->counters[i] = req->counters[i];
perfmon->ncounters = req->ncounters;
+ perfmon->v3d = v3d;
refcount_set(&perfmon->refcnt, 1);
+ atomic_inc(&v3d->perfmon_state.nperfmons);
ret = xa_alloc(&v3d_priv->perfmons, &id, perfmon, xa_limit_32b,
GFP_KERNEL);
if (ret < 0) {
- kfree(perfmon);
+ v3d_perfmon_put(perfmon);
return ret;
}
diff --git a/drivers/gpu/drm/v3d/v3d_submit.c b/drivers/gpu/drm/v3d/v3d_submit.c
index 834d52030979..bc3c43fd4fd9 100644
--- a/drivers/gpu/drm/v3d/v3d_submit.c
+++ b/drivers/gpu/drm/v3d/v3d_submit.c
@@ -357,6 +357,10 @@ v3d_attach_perfmon_to_jobs(struct v3d_submit *submit, u32 perfmon_id)
*
* We don't serialize the jobs when using a global perfmon as it's expected to
* track concurrent activity from all jobs.
+ *
+ * Keeping track of the in-flight jobs costs a fence merge per job, so it is
+ * only done while at least one perfmon is alive. Jobs submitted while no
+ * perfmon exists go untracked and may overlap the first measured job.
*/
static int
v3d_serialize_for_perfmon(struct v3d_job *job)
@@ -368,6 +372,9 @@ v3d_serialize_for_perfmon(struct v3d_job *job)
lockdep_assert_held(&v3d->sched_lock);
+ if (!atomic_read(&v3d->perfmon_state.nperfmons))
+ return 0;
+
scoped_guard(spinlock_irqsave, &v3d->perfmon_state.lock)
is_global_perfmon = !!v3d->global_perfmon;
--
2.55.0
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH 2/2] drm/v3d: Skip perfmon serialization while a global perfmon is set
2026-09-21 21:24 [PATCH 0/2] drm/v3d: Skip perfmon serialization when it cannot be observed Maíra Canal
2026-09-21 21:24 ` [PATCH 1/2] drm/v3d: Skip perfmon serialization while the device has no perfmon Maíra Canal
@ 2026-09-21 21:24 ` Maíra Canal
2026-09-21 21:35 ` sashiko-bot
2026-09-22 7:11 ` Iago Toral
2026-09-22 7:11 ` [PATCH 0/2] drm/v3d: Skip perfmon serialization when it cannot be observed Iago Toral
2 siblings, 2 replies; 7+ messages in thread
From: Maíra Canal @ 2026-09-21 21:24 UTC (permalink / raw)
To: Melissa Wen, Iago Toral, David Airlie, Simona Vetter
Cc: kernel-dev, dri-devel, Maíra Canal
A global perfmon counts concurrent activity from every job, so no
dependency is added while one is set. Skip the accumulator for the jobs
carrying no perfmon while a global perfmon is enabled, which is every job
in the common case.
Jobs carrying their own perfmon stay tracked. The global perfmon may be
cleared before such a job runs, and the next job carrying a perfmon has
to wait for it, otherwise both are in flight at once and only one gets
programmed in the HW.
Signed-off-by: Maíra Canal <mcanal@igalia.com>
---
drivers/gpu/drm/v3d/v3d_submit.c | 18 ++++++++----------
1 file changed, 8 insertions(+), 10 deletions(-)
diff --git a/drivers/gpu/drm/v3d/v3d_submit.c b/drivers/gpu/drm/v3d/v3d_submit.c
index bc3c43fd4fd9..7548faab67ba 100644
--- a/drivers/gpu/drm/v3d/v3d_submit.c
+++ b/drivers/gpu/drm/v3d/v3d_submit.c
@@ -359,15 +359,15 @@ v3d_attach_perfmon_to_jobs(struct v3d_submit *submit, u32 perfmon_id)
* track concurrent activity from all jobs.
*
* Keeping track of the in-flight jobs costs a fence merge per job, so it is
- * only done while at least one perfmon is alive. Jobs submitted while no
- * perfmon exists go untracked and may overlap the first measured job.
+ * only done while at least one perfmon is alive, and is skipped for the jobs
+ * carrying no perfmon while a global perfmon is set. Jobs submitted while
+ * tracking is off go untracked and may overlap the first measured job.
*/
static int
v3d_serialize_for_perfmon(struct v3d_job *job)
{
struct v3d_dev *v3d = job->v3d;
struct dma_fence *merged;
- bool is_global_perfmon;
int ret;
lockdep_assert_held(&v3d->sched_lock);
@@ -375,11 +375,10 @@ v3d_serialize_for_perfmon(struct v3d_job *job)
if (!atomic_read(&v3d->perfmon_state.nperfmons))
return 0;
- scoped_guard(spinlock_irqsave, &v3d->perfmon_state.lock)
- is_global_perfmon = !!v3d->global_perfmon;
-
- if (is_global_perfmon)
- goto publish;
+ scoped_guard(spinlock_irqsave, &v3d->perfmon_state.lock) {
+ if (!job->perfmon && v3d->global_perfmon)
+ return 0;
+ }
if (job->perfmon) {
for (enum v3d_queue q = 0; q < V3D_MAX_QUEUES; q++) {
@@ -400,7 +399,6 @@ v3d_serialize_for_perfmon(struct v3d_job *job)
return ret;
}
-publish:
/*
* Accumulate every in-flight job on this queue into one merged fence.
* A HW queue is fed by several scheduler entities (one per-fd), so jobs
@@ -414,7 +412,7 @@ v3d_serialize_for_perfmon(struct v3d_job *job)
dma_fence_put(v3d->perfmon_state.last_hw_fence[job->queue]);
v3d->perfmon_state.last_hw_fence[job->queue] = merged;
- if (job->perfmon && !is_global_perfmon) {
+ if (job->perfmon) {
dma_fence_put(v3d->perfmon_state.fence);
v3d->perfmon_state.fence = dma_fence_get(job->done_fence);
}
--
2.55.0
^ permalink raw reply related [flat|nested] 7+ messages in thread
* Re: [PATCH 2/2] drm/v3d: Skip perfmon serialization while a global perfmon is set
2026-09-21 21:24 ` [PATCH 2/2] drm/v3d: Skip perfmon serialization while a global perfmon is set Maíra Canal
@ 2026-09-21 21:35 ` sashiko-bot
2026-09-22 7:11 ` Iago Toral
1 sibling, 0 replies; 7+ messages in thread
From: sashiko-bot @ 2026-09-21 21:35 UTC (permalink / raw)
To: Maíra Canal; +Cc: dri-devel
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] drm/v3d: counter pollution due to skipped job tracking with global perfmon
--
commit 16d01b295e0df230e5bdde84096f44906f1411fe
Author: Maíra Canal <mcanal@igalia.com>
drm/v3d: Skip perfmon serialization while a global perfmon is set
A global perfmon counts concurrent activity from every job, so no dependency
is added while one is set. This patch skips the accumulator for the jobs
carrying no perfmon while a global perfmon is enabled.
> diff --git a/drivers/gpu/drm/v3d/v3d_submit.c b/drivers/gpu/drm/v3d/v3d_submit.c
> index bc3c43fd4fd9e..7548faab67baa 100644
> --- a/drivers/gpu/drm/v3d/v3d_submit.c
> +++ b/drivers/gpu/drm/v3d/v3d_submit.c
[ ... ]
> @@ -375,11 +375,10 @@ v3d_serialize_for_perfmon(struct v3d_job *job)
> if (!atomic_read(&v3d->perfmon_state.nperfmons))
> return 0;
>
> - scoped_guard(spinlock_irqsave, &v3d->perfmon_state.lock)
> - is_global_perfmon = !!v3d->global_perfmon;
> -
> - if (is_global_perfmon)
> - goto publish;
> + scoped_guard(spinlock_irqsave, &v3d->perfmon_state.lock) {
> + if (!job->perfmon && v3d->global_perfmon)
> + return 0;
[Severity: Medium]
Does this early return cause subsequent local perfmon jobs to overlap with
untracked ones?
If a user sets a global perfmon and a job without a local perfmon is
submitted, v3d_serialize_for_perfmon() returns early here. By skipping the
publish step, the job is never added to last_hw_fence.
If the global perfmon is subsequently cleared and a job with a local perfmon
is submitted, it will check last_hw_fence but won't find a dependency for
the previously untracked job.
Could this result in both jobs executing concurrently and polluting the
performance counters for the local perfmon?
> + }
>
> if (job->perfmon) {
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260921-v3d-quick-exit-perfmon-serialize-v1-0-e58aea1075e2@igalia.com?part=2
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 1/2] drm/v3d: Skip perfmon serialization while the device has no perfmon
2026-09-21 21:24 ` [PATCH 1/2] drm/v3d: Skip perfmon serialization while the device has no perfmon Maíra Canal
@ 2026-09-22 6:35 ` Iago Toral
0 siblings, 0 replies; 7+ messages in thread
From: Iago Toral @ 2026-09-22 6:35 UTC (permalink / raw)
To: Maíra Canal, Melissa Wen, David Airlie, Simona Vetter
Cc: kernel-dev, dri-devel
El lun, 21-09-2026 a las 18:24 -0300, Maíra Canal escribió:
> Every job merges its finished fence into the per-queue accumulator so
> that a job carrying a perfmon can later depend on everything still in
> flight. This is a (relatively high) cost that every job pays before
> it is submitted.
>
> Perfmons only exist while userspace runs a performance query, so the
> common case is a client paying that on every job for a dependency
> that is
> never actually added, increasing the submission latency.
>
> Address this situation by counting the number of perfmons alive on
> the
> device and returning early while that count is zero. This means that
> if
> no perfmon exists, v3d_serialize_for_perfmon() will bail out
> immediately.
>
> Jobs submitted while no perfmon is alive stay unaccounted and may
> overlap
> the first measured job. The count returns to zero whenever the last
> perfmon is destroyed, so that window reopens on every measurement
> cycle.
> Closing it would mean merging a fence on every submission for the
> lifetime of the device, so it is a reasonable compromise for the
> average
> use case.
>
> Signed-off-by: Maíra Canal <mcanal@igalia.com>
> ---
> drivers/gpu/drm/v3d/v3d_drv.h | 7 +++++++
> drivers/gpu/drm/v3d/v3d_perfmon.c | 8 ++++++--
> drivers/gpu/drm/v3d/v3d_submit.c | 7 +++++++
> 3 files changed, 20 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/gpu/drm/v3d/v3d_drv.h
> b/drivers/gpu/drm/v3d/v3d_drv.h
> index 79ef20c89ff2..cd544af8d8c8 100644
> --- a/drivers/gpu/drm/v3d/v3d_drv.h
> +++ b/drivers/gpu/drm/v3d/v3d_drv.h
> @@ -84,6 +84,8 @@ struct v3d_queue_state {
> * This way, only events related to a specific submission will be
> counted.
> */
> struct v3d_perfmon {
> + struct v3d_dev *v3d;
> +
> /* Tracks the number of users of the perfmon, when this
> counter reaches
> * zero the perfmon is destroyed.
> */
> @@ -184,6 +186,11 @@ struct v3d_dev {
> /* Perfmon currently programmed in HW (or NULL if
> none). */
> struct v3d_perfmon *active;
>
> + /* Number of perfmons alive on this device. Jobs are
> not
> + * serialized if the number is zero.
> + */
> + atomic_t nperfmons;
> +
Maybe we should also amend the comment for the last_hw_fence field, to
make it explicit that last fence tracking is only accurate while
nperfmons > 0?
> /* Finished fence of the most recently submitted job
> that
> * opened a serialization window (i.e. a job with a
> non-global
> * perfmon attached).
> diff --git a/drivers/gpu/drm/v3d/v3d_perfmon.c
> b/drivers/gpu/drm/v3d/v3d_perfmon.c
> index 07dab7fb3060..359334616bd8 100644
> --- a/drivers/gpu/drm/v3d/v3d_perfmon.c
> +++ b/drivers/gpu/drm/v3d/v3d_perfmon.c
> @@ -217,8 +217,10 @@ void v3d_perfmon_get(struct v3d_perfmon
> *perfmon)
>
> void v3d_perfmon_put(struct v3d_perfmon *perfmon)
> {
> - if (perfmon && refcount_dec_and_test(&perfmon->refcnt))
> + if (perfmon && refcount_dec_and_test(&perfmon->refcnt)) {
> + atomic_dec(&perfmon->v3d->perfmon_state.nperfmons);
> kfree(perfmon);
> + }
> }
>
> static void v3d_perfmon_hw_start(struct v3d_dev *v3d, struct
> v3d_perfmon *perfmon)
> @@ -434,13 +436,15 @@ int v3d_perfmon_create_ioctl(struct drm_device
> *dev, void *data,
> perfmon->counters[i] = req->counters[i];
>
> perfmon->ncounters = req->ncounters;
> + perfmon->v3d = v3d;
>
> refcount_set(&perfmon->refcnt, 1);
> + atomic_inc(&v3d->perfmon_state.nperfmons);
>
> ret = xa_alloc(&v3d_priv->perfmons, &id, perfmon,
> xa_limit_32b,
> GFP_KERNEL);
> if (ret < 0) {
> - kfree(perfmon);
> + v3d_perfmon_put(perfmon);
> return ret;
> }
>
> diff --git a/drivers/gpu/drm/v3d/v3d_submit.c
> b/drivers/gpu/drm/v3d/v3d_submit.c
> index 834d52030979..bc3c43fd4fd9 100644
> --- a/drivers/gpu/drm/v3d/v3d_submit.c
> +++ b/drivers/gpu/drm/v3d/v3d_submit.c
> @@ -357,6 +357,10 @@ v3d_attach_perfmon_to_jobs(struct v3d_submit
> *submit, u32 perfmon_id)
> *
> * We don't serialize the jobs when using a global perfmon as it's
> expected to
> * track concurrent activity from all jobs.
> + *
> + * Keeping track of the in-flight jobs costs a fence merge per job,
> so it is
> + * only done while at least one perfmon is alive. Jobs submitted
> while no
> + * perfmon exists go untracked and may overlap the first measured
> job.
> */
> static int
> v3d_serialize_for_perfmon(struct v3d_job *job)
> @@ -368,6 +372,9 @@ v3d_serialize_for_perfmon(struct v3d_job *job)
>
> lockdep_assert_held(&v3d->sched_lock);
>
> + if (!atomic_read(&v3d->perfmon_state.nperfmons))
> + return 0;
> +
> scoped_guard(spinlock_irqsave, &v3d->perfmon_state.lock)
> is_global_perfmon = !!v3d->global_perfmon;
>
>
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 2/2] drm/v3d: Skip perfmon serialization while a global perfmon is set
2026-09-21 21:24 ` [PATCH 2/2] drm/v3d: Skip perfmon serialization while a global perfmon is set Maíra Canal
2026-09-21 21:35 ` sashiko-bot
@ 2026-09-22 7:11 ` Iago Toral
1 sibling, 0 replies; 7+ messages in thread
From: Iago Toral @ 2026-09-22 7:11 UTC (permalink / raw)
To: Maíra Canal, Melissa Wen, David Airlie, Simona Vetter
Cc: kernel-dev, dri-devel
El lun, 21-09-2026 a las 18:24 -0300, Maíra Canal escribió:
> A global perfmon counts concurrent activity from every job, so no
> dependency is added while one is set. Skip the accumulator for the
> jobs
> carrying no perfmon while a global perfmon is enabled, which is every
> job
> in the common case.
>
> Jobs carrying their own perfmon stay tracked. The global perfmon may
> be
> cleared before such a job runs, and the next job carrying a perfmon
> has
> to wait for it, otherwise both are in flight at once and only one
> gets
> programmed in the HW.
>
> Signed-off-by: Maíra Canal <mcanal@igalia.com>
> ---
> drivers/gpu/drm/v3d/v3d_submit.c | 18 ++++++++----------
> 1 file changed, 8 insertions(+), 10 deletions(-)
>
> diff --git a/drivers/gpu/drm/v3d/v3d_submit.c
> b/drivers/gpu/drm/v3d/v3d_submit.c
> index bc3c43fd4fd9..7548faab67ba 100644
> --- a/drivers/gpu/drm/v3d/v3d_submit.c
> +++ b/drivers/gpu/drm/v3d/v3d_submit.c
> @@ -359,15 +359,15 @@ v3d_attach_perfmon_to_jobs(struct v3d_submit
> *submit, u32 perfmon_id)
> * track concurrent activity from all jobs.
> *
> * Keeping track of the in-flight jobs costs a fence merge per job,
> so it is
> - * only done while at least one perfmon is alive. Jobs submitted
> while no
> - * perfmon exists go untracked and may overlap the first measured
> job.
> + * only done while at least one perfmon is alive, and is skipped for
> the jobs
> + * carrying no perfmon while a global perfmon is set. Jobs submitted
> while
> + * tracking is off go untracked and may overlap the first measured
> job.
> */
> static int
> v3d_serialize_for_perfmon(struct v3d_job *job)
> {
> struct v3d_dev *v3d = job->v3d;
> struct dma_fence *merged;
> - bool is_global_perfmon;
> int ret;
>
> lockdep_assert_held(&v3d->sched_lock);
> @@ -375,11 +375,10 @@ v3d_serialize_for_perfmon(struct v3d_job *job)
> if (!atomic_read(&v3d->perfmon_state.nperfmons))
> return 0;
>
> - scoped_guard(spinlock_irqsave, &v3d->perfmon_state.lock)
> - is_global_perfmon = !!v3d->global_perfmon;
> -
> - if (is_global_perfmon)
> - goto publish;
> + scoped_guard(spinlock_irqsave, &v3d->perfmon_state.lock) {
> + if (!job->perfmon && v3d->global_perfmon)
> + return 0;
Same suggestion as in previous patch about updating docs for
last_hw_fence. This change means we don't update it in this case
either, to we should probably reflect that the field documentation.
> + }
>
> if (job->perfmon) {
> for (enum v3d_queue q = 0; q < V3D_MAX_QUEUES; q++)
> {
> @@ -400,7 +399,6 @@ v3d_serialize_for_perfmon(struct v3d_job *job)
> return ret;
> }
>
> -publish:
> /*
> * Accumulate every in-flight job on this queue into one
> merged fence.
> * A HW queue is fed by several scheduler entities (one per-
> fd), so jobs
> @@ -414,7 +412,7 @@ v3d_serialize_for_perfmon(struct v3d_job *job)
> dma_fence_put(v3d->perfmon_state.last_hw_fence[job->queue]);
> v3d->perfmon_state.last_hw_fence[job->queue] = merged;
>
> - if (job->perfmon && !is_global_perfmon) {
> + if (job->perfmon) {
> dma_fence_put(v3d->perfmon_state.fence);
> v3d->perfmon_state.fence = dma_fence_get(job-
> >done_fence);
> }
>
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 0/2] drm/v3d: Skip perfmon serialization when it cannot be observed
2026-09-21 21:24 [PATCH 0/2] drm/v3d: Skip perfmon serialization when it cannot be observed Maíra Canal
2026-09-21 21:24 ` [PATCH 1/2] drm/v3d: Skip perfmon serialization while the device has no perfmon Maíra Canal
2026-09-21 21:24 ` [PATCH 2/2] drm/v3d: Skip perfmon serialization while a global perfmon is set Maíra Canal
@ 2026-09-22 7:11 ` Iago Toral
2 siblings, 0 replies; 7+ messages in thread
From: Iago Toral @ 2026-09-22 7:11 UTC (permalink / raw)
To: Maíra Canal, Melissa Wen, David Airlie, Simona Vetter
Cc: kernel-dev, dri-devel
Hi Maíra,
this looks good to me. I suggested to amend the docs for last_hw_fence
to make more explicit how this MR affects it. Otherwise this is:
Reviewed-by: Iago Toral Quiroga <itoral@igalia.com>
El lun, 21-09-2026 a las 18:24 -0300, Maíra Canal escribió:
> The V3D core has a single set of performance counters. So that a
> perfmon
> only counts the jobs it is attached to, a job carrying a non-global
> perfmon waits for every job still in flight, and while it is in
> flight
> every later job waits for it. Knowing what is in flight requires
> every
> job to merge its finished fence into a per-queue accumulator.
>
> That accumulator is maintained unconditionally, so a client pays a
> fence
> merge on every job even when no dependency can ever be built out of
> it.
> This series turns the bookkeeping off in the two cases where that
> holds:
>
> 1. No perfmon is alive on the device, which is the common case
> since
> perfmons only exist while userspace runs a performance query;
>
> 2. A global perfmon is set and the job carries no perfmon of its
> own.
> A global perfmon counts concurrent activity from every job, so
> nothing has to wait. Jobs carrying their own perfmon stay
> tracked,
> since the global perfmon may be cleared before they run and only
> one of them can be programmed in the HW at a time.
>
> Both share the same trade-off: a job that goes unrecorded never
> carries a
> perfmon, so the only cost is that the first job measured after
> tracking
> resumes may overlap it. Considering the use cases of performance
> monitors
> and that it will only affect the first job, it's a reasonable
> compromise
> to make for the average case.
>
> Best regards,
> - Maíra
>
> ---
> Maíra Canal (2):
> drm/v3d: Skip perfmon serialization while the device has no
> perfmon
> drm/v3d: Skip perfmon serialization while a global perfmon is
> set
>
> drivers/gpu/drm/v3d/v3d_drv.h | 7 +++++++
> drivers/gpu/drm/v3d/v3d_perfmon.c | 8 ++++++--
> drivers/gpu/drm/v3d/v3d_submit.c | 19 ++++++++++++-------
> 3 files changed, 25 insertions(+), 9 deletions(-)
> ---
> base-commit: 402491eefc5ea9a40e19b08dccf0c045aec4efc2
> change-id: 20260921-v3d-quick-exit-perfmon-serialize-9e9e8ceac587
>
>
^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-09-22 7:11 UTC | newest]
Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-21 21:24 [PATCH 0/2] drm/v3d: Skip perfmon serialization when it cannot be observed Maíra Canal
2026-09-21 21:24 ` [PATCH 1/2] drm/v3d: Skip perfmon serialization while the device has no perfmon Maíra Canal
2026-09-22 6:35 ` Iago Toral
2026-09-21 21:24 ` [PATCH 2/2] drm/v3d: Skip perfmon serialization while a global perfmon is set Maíra Canal
2026-09-21 21:35 ` sashiko-bot
2026-09-22 7:11 ` Iago Toral
2026-09-22 7:11 ` [PATCH 0/2] drm/v3d: Skip perfmon serialization when it cannot be observed Iago Toral
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox