* [PATCH 1/8] drm/sched: Add locking to drm_sched_entity_modify_sched
[not found] <20240913160559.49054-1-tursulin@igalia.com>
@ 2024-09-13 16:05 ` Tvrtko Ursulin
2024-09-13 16:05 ` [PATCH 2/8] drm/sched: Always wake up correct scheduler in drm_sched_entity_push_job Tvrtko Ursulin
2024-09-13 16:05 ` [PATCH 3/8] drm/sched: Always increment correct scheduler score Tvrtko Ursulin
2 siblings, 0 replies; 10+ messages in thread
From: Tvrtko Ursulin @ 2024-09-13 16:05 UTC (permalink / raw)
To: amd-gfx, dri-devel
Cc: kernel-dev, Tvrtko Ursulin, Christian König, Alex Deucher,
Luben Tuikov, Matthew Brost, David Airlie, Daniel Vetter,
Philipp Stanner, stable
From: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Without the locking amdgpu currently can race between
amdgpu_ctx_set_entity_priority() (via drm_sched_entity_modify_sched()) and
drm_sched_job_arm(), leading to the latter accesing potentially
inconsitent entity->sched_list and entity->num_sched_list pair.
v2:
* Improve commit message. (Philipp)
Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Fixes: b37aced31eb0 ("drm/scheduler: implement a function to modify sched list")
Cc: Christian König <christian.koenig@amd.com>
Cc: Alex Deucher <alexander.deucher@amd.com>
Cc: Luben Tuikov <ltuikov89@gmail.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: David Airlie <airlied@gmail.com>
Cc: Daniel Vetter <daniel@ffwll.ch>
Cc: dri-devel@lists.freedesktop.org
Cc: Philipp Stanner <pstanner@redhat.com>
Cc: <stable@vger.kernel.org> # v5.7+
Reviewed-by: Christian König <christian.koenig@amd.com>
---
drivers/gpu/drm/scheduler/sched_entity.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/drivers/gpu/drm/scheduler/sched_entity.c b/drivers/gpu/drm/scheduler/sched_entity.c
index 58c8161289fe..ae8be30472cd 100644
--- a/drivers/gpu/drm/scheduler/sched_entity.c
+++ b/drivers/gpu/drm/scheduler/sched_entity.c
@@ -133,8 +133,10 @@ void drm_sched_entity_modify_sched(struct drm_sched_entity *entity,
{
WARN_ON(!num_sched_list || !sched_list);
+ spin_lock(&entity->rq_lock);
entity->sched_list = sched_list;
entity->num_sched_list = num_sched_list;
+ spin_unlock(&entity->rq_lock);
}
EXPORT_SYMBOL(drm_sched_entity_modify_sched);
--
2.46.0
^ permalink raw reply related [flat|nested] 10+ messages in thread* [PATCH 2/8] drm/sched: Always wake up correct scheduler in drm_sched_entity_push_job
[not found] <20240913160559.49054-1-tursulin@igalia.com>
2024-09-13 16:05 ` [PATCH 1/8] drm/sched: Add locking to drm_sched_entity_modify_sched Tvrtko Ursulin
@ 2024-09-13 16:05 ` Tvrtko Ursulin
2024-09-13 16:05 ` [PATCH 3/8] drm/sched: Always increment correct scheduler score Tvrtko Ursulin
2 siblings, 0 replies; 10+ messages in thread
From: Tvrtko Ursulin @ 2024-09-13 16:05 UTC (permalink / raw)
To: amd-gfx, dri-devel
Cc: kernel-dev, Tvrtko Ursulin, Christian König, Alex Deucher,
Luben Tuikov, Matthew Brost, David Airlie, Daniel Vetter,
Philipp Stanner, stable
From: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Since drm_sched_entity_modify_sched() can modify the entities run queue,
lets make sure to only dereference the pointer once so both adding and
waking up are guaranteed to be consistent.
Alternative of moving the spin_unlock to after the wake up would for now
be more problematic since the same lock is taken inside
drm_sched_rq_update_fifo().
v2:
* Improve commit message. (Philipp)
* Cache the scheduler pointer directly. (Christian)
Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Fixes: b37aced31eb0 ("drm/scheduler: implement a function to modify sched list")
Cc: Christian König <christian.koenig@amd.com>
Cc: Alex Deucher <alexander.deucher@amd.com>
Cc: Luben Tuikov <ltuikov89@gmail.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: David Airlie <airlied@gmail.com>
Cc: Daniel Vetter <daniel@ffwll.ch>
Cc: Philipp Stanner <pstanner@redhat.com>
Cc: dri-devel@lists.freedesktop.org
Cc: <stable@vger.kernel.org> # v5.7+
Reviewed-by: Christian König <christian.koenig@amd.com>
---
drivers/gpu/drm/scheduler/sched_entity.c | 10 ++++++++--
1 file changed, 8 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/scheduler/sched_entity.c b/drivers/gpu/drm/scheduler/sched_entity.c
index ae8be30472cd..76e422548d40 100644
--- a/drivers/gpu/drm/scheduler/sched_entity.c
+++ b/drivers/gpu/drm/scheduler/sched_entity.c
@@ -599,6 +599,9 @@ void drm_sched_entity_push_job(struct drm_sched_job *sched_job)
/* first job wakes up scheduler */
if (first) {
+ struct drm_gpu_scheduler *sched;
+ struct drm_sched_rq *rq;
+
/* Add the entity to the run queue */
spin_lock(&entity->rq_lock);
if (entity->stopped) {
@@ -608,13 +611,16 @@ void drm_sched_entity_push_job(struct drm_sched_job *sched_job)
return;
}
- drm_sched_rq_add_entity(entity->rq, entity);
+ rq = entity->rq;
+ sched = rq->sched;
+
+ drm_sched_rq_add_entity(rq, entity);
spin_unlock(&entity->rq_lock);
if (drm_sched_policy == DRM_SCHED_POLICY_FIFO)
drm_sched_rq_update_fifo(entity, submit_ts);
- drm_sched_wakeup(entity->rq->sched, entity);
+ drm_sched_wakeup(sched, entity);
}
}
EXPORT_SYMBOL(drm_sched_entity_push_job);
--
2.46.0
^ permalink raw reply related [flat|nested] 10+ messages in thread* [PATCH 3/8] drm/sched: Always increment correct scheduler score
[not found] <20240913160559.49054-1-tursulin@igalia.com>
2024-09-13 16:05 ` [PATCH 1/8] drm/sched: Add locking to drm_sched_entity_modify_sched Tvrtko Ursulin
2024-09-13 16:05 ` [PATCH 2/8] drm/sched: Always wake up correct scheduler in drm_sched_entity_push_job Tvrtko Ursulin
@ 2024-09-13 16:05 ` Tvrtko Ursulin
2024-09-30 13:01 ` Tvrtko Ursulin
2 siblings, 1 reply; 10+ messages in thread
From: Tvrtko Ursulin @ 2024-09-13 16:05 UTC (permalink / raw)
To: amd-gfx, dri-devel
Cc: kernel-dev, Tvrtko Ursulin, Nirmoy Das, Christian König,
Luben Tuikov, Matthew Brost, David Airlie, Daniel Vetter, stable,
Nirmoy Das
From: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Entities run queue can change during drm_sched_entity_push_job() so make
sure to update the score consistently.
Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Fixes: d41a39dda140 ("drm/scheduler: improve job distribution with multiple queues")
Cc: Nirmoy Das <nirmoy.das@amd.com>
Cc: Christian König <christian.koenig@amd.com>
Cc: Luben Tuikov <ltuikov89@gmail.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: David Airlie <airlied@gmail.com>
Cc: Daniel Vetter <daniel@ffwll.ch>
Cc: dri-devel@lists.freedesktop.org
Cc: <stable@vger.kernel.org> # v5.9+
Reviewed-by: Christian König <christian.koenig@amd.com>
Reviewed-by: Nirmoy Das <nirmoy.das@intel.com>
---
drivers/gpu/drm/scheduler/sched_entity.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/gpu/drm/scheduler/sched_entity.c b/drivers/gpu/drm/scheduler/sched_entity.c
index 76e422548d40..6645a8524699 100644
--- a/drivers/gpu/drm/scheduler/sched_entity.c
+++ b/drivers/gpu/drm/scheduler/sched_entity.c
@@ -586,7 +586,6 @@ void drm_sched_entity_push_job(struct drm_sched_job *sched_job)
ktime_t submit_ts;
trace_drm_sched_job(sched_job, entity);
- atomic_inc(entity->rq->sched->score);
WRITE_ONCE(entity->last_user, current->group_leader);
/*
@@ -614,6 +613,7 @@ void drm_sched_entity_push_job(struct drm_sched_job *sched_job)
rq = entity->rq;
sched = rq->sched;
+ atomic_inc(sched->score);
drm_sched_rq_add_entity(rq, entity);
spin_unlock(&entity->rq_lock);
--
2.46.0
^ permalink raw reply related [flat|nested] 10+ messages in thread* Re: [PATCH 3/8] drm/sched: Always increment correct scheduler score
2024-09-13 16:05 ` [PATCH 3/8] drm/sched: Always increment correct scheduler score Tvrtko Ursulin
@ 2024-09-30 13:01 ` Tvrtko Ursulin
2024-09-30 13:07 ` Christian König
0 siblings, 1 reply; 10+ messages in thread
From: Tvrtko Ursulin @ 2024-09-30 13:01 UTC (permalink / raw)
To: Tvrtko Ursulin, amd-gfx, dri-devel
Cc: kernel-dev, Christian König, Luben Tuikov, Matthew Brost,
David Airlie, Daniel Vetter, stable, Nirmoy Das, Nirmoy Das
On 13/09/2024 17:05, Tvrtko Ursulin wrote:
> From: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
>
> Entities run queue can change during drm_sched_entity_push_job() so make
> sure to update the score consistently.
>
> Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
> Fixes: d41a39dda140 ("drm/scheduler: improve job distribution with multiple queues")
> Cc: Nirmoy Das <nirmoy.das@amd.com>
> Cc: Christian König <christian.koenig@amd.com>
> Cc: Luben Tuikov <ltuikov89@gmail.com>
> Cc: Matthew Brost <matthew.brost@intel.com>
> Cc: David Airlie <airlied@gmail.com>
> Cc: Daniel Vetter <daniel@ffwll.ch>
> Cc: dri-devel@lists.freedesktop.org
> Cc: <stable@vger.kernel.org> # v5.9+
> Reviewed-by: Christian König <christian.koenig@amd.com>
> Reviewed-by: Nirmoy Das <nirmoy.das@intel.com>
> ---
> drivers/gpu/drm/scheduler/sched_entity.c | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/drivers/gpu/drm/scheduler/sched_entity.c b/drivers/gpu/drm/scheduler/sched_entity.c
> index 76e422548d40..6645a8524699 100644
> --- a/drivers/gpu/drm/scheduler/sched_entity.c
> +++ b/drivers/gpu/drm/scheduler/sched_entity.c
> @@ -586,7 +586,6 @@ void drm_sched_entity_push_job(struct drm_sched_job *sched_job)
> ktime_t submit_ts;
>
> trace_drm_sched_job(sched_job, entity);
> - atomic_inc(entity->rq->sched->score);
> WRITE_ONCE(entity->last_user, current->group_leader);
>
> /*
> @@ -614,6 +613,7 @@ void drm_sched_entity_push_job(struct drm_sched_job *sched_job)
> rq = entity->rq;
> sched = rq->sched;
>
> + atomic_inc(sched->score);
Ugh this is wrong. :(
I was working on some further consolidation and realised this.
It will create an imbalance in score since score is currently supposed
to be accounted twice:
1. +/- 1 for each entity (de-)queued
2. +/- 1 for each job queued/completed
By moving it into the "if (first) branch" it unbalances it.
But it is still true the original placement is racy. It looks like what
is required is an unconditional entity->lock section after
spsc_queue_push. AFAICT that's the only way to be sure entity->rq is set
for the submission at hand.
Question also is, why +/- score in entity add/remove and not just for jobs?
In the meantime patch will need to get reverted.
Regards,
Tvrtko
> drm_sched_rq_add_entity(rq, entity);
> spin_unlock(&entity->rq_lock);
>
^ permalink raw reply [flat|nested] 10+ messages in thread* Re: [PATCH 3/8] drm/sched: Always increment correct scheduler score
2024-09-30 13:01 ` Tvrtko Ursulin
@ 2024-09-30 13:07 ` Christian König
2024-09-30 13:22 ` Tvrtko Ursulin
0 siblings, 1 reply; 10+ messages in thread
From: Christian König @ 2024-09-30 13:07 UTC (permalink / raw)
To: Tvrtko Ursulin, Tvrtko Ursulin, amd-gfx, dri-devel
Cc: kernel-dev, Luben Tuikov, Matthew Brost, David Airlie,
Daniel Vetter, stable, Nirmoy Das
Am 30.09.24 um 15:01 schrieb Tvrtko Ursulin:
>
> On 13/09/2024 17:05, Tvrtko Ursulin wrote:
>> From: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
>>
>> Entities run queue can change during drm_sched_entity_push_job() so make
>> sure to update the score consistently.
>>
>> Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
>> Fixes: d41a39dda140 ("drm/scheduler: improve job distribution with
>> multiple queues")
>> Cc: Nirmoy Das <nirmoy.das@amd.com>
>> Cc: Christian König <christian.koenig@amd.com>
>> Cc: Luben Tuikov <ltuikov89@gmail.com>
>> Cc: Matthew Brost <matthew.brost@intel.com>
>> Cc: David Airlie <airlied@gmail.com>
>> Cc: Daniel Vetter <daniel@ffwll.ch>
>> Cc: dri-devel@lists.freedesktop.org
>> Cc: <stable@vger.kernel.org> # v5.9+
>> Reviewed-by: Christian König <christian.koenig@amd.com>
>> Reviewed-by: Nirmoy Das <nirmoy.das@intel.com>
>> ---
>> drivers/gpu/drm/scheduler/sched_entity.c | 2 +-
>> 1 file changed, 1 insertion(+), 1 deletion(-)
>>
>> diff --git a/drivers/gpu/drm/scheduler/sched_entity.c
>> b/drivers/gpu/drm/scheduler/sched_entity.c
>> index 76e422548d40..6645a8524699 100644
>> --- a/drivers/gpu/drm/scheduler/sched_entity.c
>> +++ b/drivers/gpu/drm/scheduler/sched_entity.c
>> @@ -586,7 +586,6 @@ void drm_sched_entity_push_job(struct
>> drm_sched_job *sched_job)
>> ktime_t submit_ts;
>> trace_drm_sched_job(sched_job, entity);
>> - atomic_inc(entity->rq->sched->score);
>> WRITE_ONCE(entity->last_user, current->group_leader);
>> /*
>> @@ -614,6 +613,7 @@ void drm_sched_entity_push_job(struct
>> drm_sched_job *sched_job)
>> rq = entity->rq;
>> sched = rq->sched;
>> + atomic_inc(sched->score);
>
> Ugh this is wrong. :(
>
> I was working on some further consolidation and realised this.
>
> It will create an imbalance in score since score is currently supposed
> to be accounted twice:
>
> 1. +/- 1 for each entity (de-)queued
> 2. +/- 1 for each job queued/completed
>
> By moving it into the "if (first) branch" it unbalances it.
>
> But it is still true the original placement is racy. It looks like
> what is required is an unconditional entity->lock section after
> spsc_queue_push. AFAICT that's the only way to be sure entity->rq is
> set for the submission at hand.
>
> Question also is, why +/- score in entity add/remove and not just for
> jobs?
>
> In the meantime patch will need to get reverted.
Ok going to revert that.
I also just realized that we don't need to change anything. The rq can't
change as soon as there is a job armed for it.
So having the increment right before pushing the armed job to the entity
was actually correct in the first place.
Regards,
Christian.
>
> Regards,
>
> Tvrtko
>
>> drm_sched_rq_add_entity(rq, entity);
>> spin_unlock(&entity->rq_lock);
^ permalink raw reply [flat|nested] 10+ messages in thread* Re: [PATCH 3/8] drm/sched: Always increment correct scheduler score
2024-09-30 13:07 ` Christian König
@ 2024-09-30 13:22 ` Tvrtko Ursulin
2024-09-30 13:27 ` Christian König
0 siblings, 1 reply; 10+ messages in thread
From: Tvrtko Ursulin @ 2024-09-30 13:22 UTC (permalink / raw)
To: Christian König, Tvrtko Ursulin, amd-gfx, dri-devel
Cc: kernel-dev, Luben Tuikov, Matthew Brost, David Airlie,
Daniel Vetter, stable, Nirmoy Das
On 30/09/2024 14:07, Christian König wrote:
> Am 30.09.24 um 15:01 schrieb Tvrtko Ursulin:
>>
>> On 13/09/2024 17:05, Tvrtko Ursulin wrote:
>>> From: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
>>>
>>> Entities run queue can change during drm_sched_entity_push_job() so make
>>> sure to update the score consistently.
>>>
>>> Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
>>> Fixes: d41a39dda140 ("drm/scheduler: improve job distribution with
>>> multiple queues")
>>> Cc: Nirmoy Das <nirmoy.das@amd.com>
>>> Cc: Christian König <christian.koenig@amd.com>
>>> Cc: Luben Tuikov <ltuikov89@gmail.com>
>>> Cc: Matthew Brost <matthew.brost@intel.com>
>>> Cc: David Airlie <airlied@gmail.com>
>>> Cc: Daniel Vetter <daniel@ffwll.ch>
>>> Cc: dri-devel@lists.freedesktop.org
>>> Cc: <stable@vger.kernel.org> # v5.9+
>>> Reviewed-by: Christian König <christian.koenig@amd.com>
>>> Reviewed-by: Nirmoy Das <nirmoy.das@intel.com>
>>> ---
>>> drivers/gpu/drm/scheduler/sched_entity.c | 2 +-
>>> 1 file changed, 1 insertion(+), 1 deletion(-)
>>>
>>> diff --git a/drivers/gpu/drm/scheduler/sched_entity.c
>>> b/drivers/gpu/drm/scheduler/sched_entity.c
>>> index 76e422548d40..6645a8524699 100644
>>> --- a/drivers/gpu/drm/scheduler/sched_entity.c
>>> +++ b/drivers/gpu/drm/scheduler/sched_entity.c
>>> @@ -586,7 +586,6 @@ void drm_sched_entity_push_job(struct
>>> drm_sched_job *sched_job)
>>> ktime_t submit_ts;
>>> trace_drm_sched_job(sched_job, entity);
>>> - atomic_inc(entity->rq->sched->score);
>>> WRITE_ONCE(entity->last_user, current->group_leader);
>>> /*
>>> @@ -614,6 +613,7 @@ void drm_sched_entity_push_job(struct
>>> drm_sched_job *sched_job)
>>> rq = entity->rq;
>>> sched = rq->sched;
>>> + atomic_inc(sched->score);
>>
>> Ugh this is wrong. :(
>>
>> I was working on some further consolidation and realised this.
>>
>> It will create an imbalance in score since score is currently supposed
>> to be accounted twice:
>>
>> 1. +/- 1 for each entity (de-)queued
>> 2. +/- 1 for each job queued/completed
>>
>> By moving it into the "if (first) branch" it unbalances it.
>>
>> But it is still true the original placement is racy. It looks like
>> what is required is an unconditional entity->lock section after
>> spsc_queue_push. AFAICT that's the only way to be sure entity->rq is
>> set for the submission at hand.
>>
>> Question also is, why +/- score in entity add/remove and not just for
>> jobs?
>>
>> In the meantime patch will need to get reverted.
>
> Ok going to revert that.
Thank you, and sorry for the trouble!
> I also just realized that we don't need to change anything. The rq can't
> change as soon as there is a job armed for it.
>
> So having the increment right before pushing the armed job to the entity
> was actually correct in the first place.
Are you sure? Two threads racing to arm and push on the same entity?
T1 T2
arm job
rq1 selected
..
push job arm job
inc score rq1
spsc_queue_count check passes
--- just before T1 spsc_queue_push ---
changed to rq2
spsc_queue_push
if (first)
resamples entity->rq
queues rq2
Where rq1 and rq2 belong to different schedulers.
Regards,
Tvrtko
> Regards,
> Christian.
>
>>
>> Regards,
>>
>> Tvrtko
>>
>>> drm_sched_rq_add_entity(rq, entity);
>>> spin_unlock(&entity->rq_lock);
>
^ permalink raw reply [flat|nested] 10+ messages in thread* Re: [PATCH 3/8] drm/sched: Always increment correct scheduler score
2024-09-30 13:22 ` Tvrtko Ursulin
@ 2024-09-30 13:27 ` Christian König
0 siblings, 0 replies; 10+ messages in thread
From: Christian König @ 2024-09-30 13:27 UTC (permalink / raw)
To: Tvrtko Ursulin, Tvrtko Ursulin, amd-gfx, dri-devel
Cc: kernel-dev, Luben Tuikov, Matthew Brost, David Airlie,
Daniel Vetter, stable, Nirmoy Das
Am 30.09.24 um 15:22 schrieb Tvrtko Ursulin:
>
> On 30/09/2024 14:07, Christian König wrote:
>> Am 30.09.24 um 15:01 schrieb Tvrtko Ursulin:
>>>
>>> On 13/09/2024 17:05, Tvrtko Ursulin wrote:
>>>> From: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
>>>>
>>>> Entities run queue can change during drm_sched_entity_push_job() so
>>>> make
>>>> sure to update the score consistently.
>>>>
>>>> Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
>>>> Fixes: d41a39dda140 ("drm/scheduler: improve job distribution with
>>>> multiple queues")
>>>> Cc: Nirmoy Das <nirmoy.das@amd.com>
>>>> Cc: Christian König <christian.koenig@amd.com>
>>>> Cc: Luben Tuikov <ltuikov89@gmail.com>
>>>> Cc: Matthew Brost <matthew.brost@intel.com>
>>>> Cc: David Airlie <airlied@gmail.com>
>>>> Cc: Daniel Vetter <daniel@ffwll.ch>
>>>> Cc: dri-devel@lists.freedesktop.org
>>>> Cc: <stable@vger.kernel.org> # v5.9+
>>>> Reviewed-by: Christian König <christian.koenig@amd.com>
>>>> Reviewed-by: Nirmoy Das <nirmoy.das@intel.com>
>>>> ---
>>>> drivers/gpu/drm/scheduler/sched_entity.c | 2 +-
>>>> 1 file changed, 1 insertion(+), 1 deletion(-)
>>>>
>>>> diff --git a/drivers/gpu/drm/scheduler/sched_entity.c
>>>> b/drivers/gpu/drm/scheduler/sched_entity.c
>>>> index 76e422548d40..6645a8524699 100644
>>>> --- a/drivers/gpu/drm/scheduler/sched_entity.c
>>>> +++ b/drivers/gpu/drm/scheduler/sched_entity.c
>>>> @@ -586,7 +586,6 @@ void drm_sched_entity_push_job(struct
>>>> drm_sched_job *sched_job)
>>>> ktime_t submit_ts;
>>>> trace_drm_sched_job(sched_job, entity);
>>>> - atomic_inc(entity->rq->sched->score);
>>>> WRITE_ONCE(entity->last_user, current->group_leader);
>>>> /*
>>>> @@ -614,6 +613,7 @@ void drm_sched_entity_push_job(struct
>>>> drm_sched_job *sched_job)
>>>> rq = entity->rq;
>>>> sched = rq->sched;
>>>> + atomic_inc(sched->score);
>>>
>>> Ugh this is wrong. :(
>>>
>>> I was working on some further consolidation and realised this.
>>>
>>> It will create an imbalance in score since score is currently
>>> supposed to be accounted twice:
>>>
>>> 1. +/- 1 for each entity (de-)queued
>>> 2. +/- 1 for each job queued/completed
>>>
>>> By moving it into the "if (first) branch" it unbalances it.
>>>
>>> But it is still true the original placement is racy. It looks like
>>> what is required is an unconditional entity->lock section after
>>> spsc_queue_push. AFAICT that's the only way to be sure entity->rq is
>>> set for the submission at hand.
>>>
>>> Question also is, why +/- score in entity add/remove and not just
>>> for jobs?
>>>
>>> In the meantime patch will need to get reverted.
>>
>> Ok going to revert that.
>
> Thank you, and sorry for the trouble!
>
>> I also just realized that we don't need to change anything. The rq
>> can't change as soon as there is a job armed for it.
>>
>> So having the increment right before pushing the armed job to the
>> entity was actually correct in the first place.
>
> Are you sure? Two threads racing to arm and push on the same entity?
>
>
> T1 T2
>
> arm job
> rq1 selected
> ..
> push job arm job
> inc score rq1
> spsc_queue_count check passes
> --- just before T1 spsc_queue_push ---
> changed to rq2
> spsc_queue_push
> if (first)
> resamples entity->rq
> queues rq2
>
> Where rq1 and rq2 belong to different schedulers.
arm/push must be protected by an external lock preventing two threads
pushing into the same entity at the same time.
That's what this misleading comment from Sima we already discussed
should have meant.
Regards,
Christian.
>
> Regards,
>
> Tvrtko
>
>
>> Regards,
>> Christian.
>>
>>>
>>> Regards,
>>>
>>> Tvrtko
>>>
>>>> drm_sched_rq_add_entity(rq, entity);
>>>> spin_unlock(&entity->rq_lock);
>>
^ permalink raw reply [flat|nested] 10+ messages in thread