Linux kernel -stable discussions
 help / color / mirror / Atom feed
From: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
To: "Christian König" <christian.koenig@amd.com>,
	"Tvrtko Ursulin" <tursulin@igalia.com>,
	amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org
Cc: kernel-dev@igalia.com, Luben Tuikov <ltuikov89@gmail.com>,
	Matthew Brost <matthew.brost@intel.com>,
	David Airlie <airlied@gmail.com>, Daniel Vetter <daniel@ffwll.ch>,
	stable@vger.kernel.org, Nirmoy Das <nirmoy.das@intel.com>
Subject: Re: [PATCH 3/8] drm/sched: Always increment correct scheduler score
Date: Mon, 30 Sep 2024 14:22:28 +0100	[thread overview]
Message-ID: <2b0860a2-5ef0-496f-9283-d5056433af58@igalia.com> (raw)
In-Reply-To: <cf135523-92ca-4d41-9acf-e979c9769ad9@amd.com>


On 30/09/2024 14:07, Christian König wrote:
> Am 30.09.24 um 15:01 schrieb Tvrtko Ursulin:
>>
>> On 13/09/2024 17:05, Tvrtko Ursulin wrote:
>>> From: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
>>>
>>> Entities run queue can change during drm_sched_entity_push_job() so make
>>> sure to update the score consistently.
>>>
>>> Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
>>> Fixes: d41a39dda140 ("drm/scheduler: improve job distribution with 
>>> multiple queues")
>>> Cc: Nirmoy Das <nirmoy.das@amd.com>
>>> Cc: Christian König <christian.koenig@amd.com>
>>> Cc: Luben Tuikov <ltuikov89@gmail.com>
>>> Cc: Matthew Brost <matthew.brost@intel.com>
>>> Cc: David Airlie <airlied@gmail.com>
>>> Cc: Daniel Vetter <daniel@ffwll.ch>
>>> Cc: dri-devel@lists.freedesktop.org
>>> Cc: <stable@vger.kernel.org> # v5.9+
>>> Reviewed-by: Christian König <christian.koenig@amd.com>
>>> Reviewed-by: Nirmoy Das <nirmoy.das@intel.com>
>>> ---
>>>   drivers/gpu/drm/scheduler/sched_entity.c | 2 +-
>>>   1 file changed, 1 insertion(+), 1 deletion(-)
>>>
>>> diff --git a/drivers/gpu/drm/scheduler/sched_entity.c 
>>> b/drivers/gpu/drm/scheduler/sched_entity.c
>>> index 76e422548d40..6645a8524699 100644
>>> --- a/drivers/gpu/drm/scheduler/sched_entity.c
>>> +++ b/drivers/gpu/drm/scheduler/sched_entity.c
>>> @@ -586,7 +586,6 @@ void drm_sched_entity_push_job(struct 
>>> drm_sched_job *sched_job)
>>>       ktime_t submit_ts;
>>>         trace_drm_sched_job(sched_job, entity);
>>> -    atomic_inc(entity->rq->sched->score);
>>>       WRITE_ONCE(entity->last_user, current->group_leader);
>>>         /*
>>> @@ -614,6 +613,7 @@ void drm_sched_entity_push_job(struct 
>>> drm_sched_job *sched_job)
>>>           rq = entity->rq;
>>>           sched = rq->sched;
>>>   +        atomic_inc(sched->score);
>>
>> Ugh this is wrong. :(
>>
>> I was working on some further consolidation and realised this.
>>
>> It will create an imbalance in score since score is currently supposed 
>> to be accounted twice:
>>
>>  1. +/- 1 for each entity (de-)queued
>>  2. +/- 1 for each job queued/completed
>>
>> By moving it into the "if (first) branch" it unbalances it.
>>
>> But it is still true the original placement is racy. It looks like 
>> what is required is an unconditional entity->lock section after 
>> spsc_queue_push. AFAICT that's the only way to be sure entity->rq is 
>> set for the submission at hand.
>>
>> Question also is, why +/- score in entity add/remove and not just for 
>> jobs?
>>
>> In the meantime patch will need to get reverted.
> 
> Ok going to revert that.

Thank you, and sorry for the trouble!

> I also just realized that we don't need to change anything. The rq can't 
> change as soon as there is a job armed for it.
> 
> So having the increment right before pushing the armed job to the entity 
> was actually correct in the first place.

Are you sure? Two threads racing to arm and push on the same entity?

	
	T1		T2

	arm job
	rq1 selected
	..
	push job	arm job
	inc score rq1
			spsc_queue_count check passes
	 ---  just before T1 spsc_queue_push ---
			changed to rq2
	spsc_queue_push
	if (first)
	  resamples entity->rq
	  queues rq2

Where rq1 and rq2 belong to different schedulers.	

Regards,

Tvrtko


> Regards,
> Christian.
> 
>>
>> Regards,
>>
>> Tvrtko
>>
>>>           drm_sched_rq_add_entity(rq, entity);
>>>           spin_unlock(&entity->rq_lock);
> 

  reply	other threads:[~2024-09-30 13:22 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <20240913160559.49054-1-tursulin@igalia.com>
2024-09-13 16:05 ` [PATCH 1/8] drm/sched: Add locking to drm_sched_entity_modify_sched Tvrtko Ursulin
2024-09-13 16:05 ` [PATCH 2/8] drm/sched: Always wake up correct scheduler in drm_sched_entity_push_job Tvrtko Ursulin
2024-09-13 16:05 ` [PATCH 3/8] drm/sched: Always increment correct scheduler score Tvrtko Ursulin
2024-09-30 13:01   ` Tvrtko Ursulin
2024-09-30 13:07     ` Christian König
2024-09-30 13:22       ` Tvrtko Ursulin [this message]
2024-09-30 13:27         ` Christian König
     [not found] <20240924101914.2713-1-tursulin@igalia.com>
2024-09-24 10:19 ` Tvrtko Ursulin
     [not found] <20240909171937.51550-1-tursulin@igalia.com>
2024-09-09 17:19 ` Tvrtko Ursulin

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=2b0860a2-5ef0-496f-9283-d5056433af58@igalia.com \
    --to=tvrtko.ursulin@igalia.com \
    --cc=airlied@gmail.com \
    --cc=amd-gfx@lists.freedesktop.org \
    --cc=christian.koenig@amd.com \
    --cc=daniel@ffwll.ch \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=kernel-dev@igalia.com \
    --cc=ltuikov89@gmail.com \
    --cc=matthew.brost@intel.com \
    --cc=nirmoy.das@intel.com \
    --cc=stable@vger.kernel.org \
    --cc=tursulin@igalia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox