dri-devel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Luben Tuikov <ltuikov89@gmail.com>
To: Danilo Krummrich <dakr@redhat.com>, tvrtko.ursulin@linux.intel.com
Cc: matthew.brost@intel.com, robdclark@chromium.org,
	sarah.walker@imgtec.com, ketil.johnsen@arm.com,
	lina@asahilina.net, mcanal@igalia.com, Liviu.Dudau@arm.com,
	dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org,
	boris.brezillon@collabora.com, donald.robson@imgtec.com,
	christian.koenig@amd.com, faith.ekstrand@collabora.com
Subject: Re: [PATCH] drm/sched: Don't disturb the entity when in RR-mode scheduling
Date: Thu, 9 Nov 2023 18:49:14 -0500	[thread overview]
Message-ID: <6273fadf-267a-4965-82ab-89c5b3f28cf2@gmail.com> (raw)
In-Reply-To: <6170dbc1-e8ab-41e9-916b-ccdc2be7ac6b@redhat.com>


[-- Attachment #1.1.1: Type: text/plain, Size: 5011 bytes --]

On 2023-11-09 18:41, Danilo Krummrich wrote:
> On 11/9/23 20:24, Danilo Krummrich wrote:
>> On 11/9/23 07:52, Luben Tuikov wrote:
>>> Hi,
>>>
>>> On 2023-11-07 19:41, Danilo Krummrich wrote:
>>>> On 11/7/23 05:10, Luben Tuikov wrote:
>>>>> Don't call drm_sched_select_entity() in drm_sched_run_job_queue().  In fact,
>>>>> rename __drm_sched_run_job_queue() to just drm_sched_run_job_queue(), and let
>>>>> it do just that, schedule the work item for execution.
>>>>>
>>>>> The problem is that drm_sched_run_job_queue() calls drm_sched_select_entity()
>>>>> to determine if the scheduler has an entity ready in one of its run-queues,
>>>>> and in the case of the Round-Robin (RR) scheduling, the function
>>>>> drm_sched_rq_select_entity_rr() does just that, selects the _next_ entity
>>>>> which is ready, sets up the run-queue and completion and returns that
>>>>> entity. The FIFO scheduling algorithm is unaffected.
>>>>>
>>>>> Now, since drm_sched_run_job_work() also calls drm_sched_select_entity(), then
>>>>> in the case of RR scheduling, that would result in drm_sched_select_entity()
>>>>> having been called twice, which may result in skipping a ready entity if more
>>>>> than one entity is ready. This commit fixes this by eliminating the call to
>>>>> drm_sched_select_entity() from drm_sched_run_job_queue(), and leaves it only
>>>>> in drm_sched_run_job_work().
>>>>>
>>>>> v2: Rebased on top of Tvrtko's renames series of patches. (Luben)
>>>>>       Add fixes-tag. (Tvrtko)
>>>>>
>>>>> Signed-off-by: Luben Tuikov <ltuikov89@gmail.com>
>>>>> Fixes: f7fe64ad0f22ff ("drm/sched: Split free_job into own work item")
>>>>> ---
>>>>>    drivers/gpu/drm/scheduler/sched_main.c | 16 +++-------------
>>>>>    1 file changed, 3 insertions(+), 13 deletions(-)
>>>>>
>>>>> diff --git a/drivers/gpu/drm/scheduler/sched_main.c b/drivers/gpu/drm/scheduler/sched_main.c
>>>>> index 27843e37d9b769..cd0dc3f81d05f0 100644
>>>>> --- a/drivers/gpu/drm/scheduler/sched_main.c
>>>>> +++ b/drivers/gpu/drm/scheduler/sched_main.c
>>>>> @@ -256,10 +256,10 @@ drm_sched_rq_select_entity_fifo(struct drm_sched_rq *rq)
>>>>>    }
>>>>>    /**
>>>>> - * __drm_sched_run_job_queue - enqueue run-job work
>>>>> + * drm_sched_run_job_queue - enqueue run-job work
>>>>>     * @sched: scheduler instance
>>>>>     */
>>>>> -static void __drm_sched_run_job_queue(struct drm_gpu_scheduler *sched)
>>>>> +static void drm_sched_run_job_queue(struct drm_gpu_scheduler *sched)
>>>>>    {
>>>>>        if (!READ_ONCE(sched->pause_submit))
>>>>>            queue_work(sched->submit_wq, &sched->work_run_job);
>>>>> @@ -928,7 +928,7 @@ static bool drm_sched_can_queue(struct drm_gpu_scheduler *sched)
>>>>>    void drm_sched_wakeup(struct drm_gpu_scheduler *sched)
>>>>>    {
>>>>>        if (drm_sched_can_queue(sched))
>>>>> -        __drm_sched_run_job_queue(sched);
>>>>> +        drm_sched_run_job_queue(sched);
>>>>>    }
>>>>>    /**
>>>>> @@ -1040,16 +1040,6 @@ drm_sched_pick_best(struct drm_gpu_scheduler **sched_list,
>>>>>    }
>>>>>    EXPORT_SYMBOL(drm_sched_pick_best);
>>>>> -/**
>>>>> - * drm_sched_run_job_queue - enqueue run-job work if there are ready entities
>>>>> - * @sched: scheduler instance
>>>>> - */
>>>>> -static void drm_sched_run_job_queue(struct drm_gpu_scheduler *sched)
>>>>> -{
>>>>> -    if (drm_sched_select_entity(sched))
>>>>
>>>> Hm, now that I rebase my patch to implement dynamic job-flow control I recognize that
>>>> we probably need the peek semantics here. If we do not select an entity here, we also
>>>> do not check whether the corresponding job fits on the ring.
>>>>
>>>> Alternatively, we simply can't do this check in drm_sched_wakeup(). The consequence would
>>>> be that we don't detect that we need to wait for credits to free up before the run work is
>>>> already executing and the run work selects an entity.
>>>
>>> So I rebased v5 on top of the latest drm-misc-next, and looked around and found out that
>>> drm_sched_wakeup() is missing drm_sched_entity_is_ready(). It should look like the following,
>>
>> Yeah, but that's just the consequence of re-basing it onto Tvrtko's patch.
>>
>> My point is that by removing drm_sched_select_entity() from drm_sched_run_job_queue() we do not
>> only loose the check whether the selected entity is ready, but also whether we have enough
>> credits to actually run a new job. This can lead to queuing up work that does nothing but calling
>> drm_sched_select_entity() and return.
> 
> Ok, I see it now.  We don't need to peek, we know the entity at drm_sched_wakeup().
> 
> However, the missing drm_sched_entity_is_ready() check should have been added already when
> drm_sched_select_entity() was removed. Gonna send a fix for that as well.

Let me do that, since I added it to your patch.
Then you can rebase your credits patch onto mine.

-- 
Regards,
Luben

[-- Attachment #1.1.2: OpenPGP public key --]
[-- Type: application/pgp-keys, Size: 677 bytes --]

[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 236 bytes --]

  reply	other threads:[~2023-11-09 23:49 UTC|newest]

Thread overview: 28+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2023-10-31  3:24 [PATCH v8 0/5] DRM scheduler changes for Xe Matthew Brost
2023-10-31  3:24 ` [PATCH v8 1/5] drm/sched: Add drm_sched_wqueue_* helpers Matthew Brost
2023-10-31  3:24 ` [PATCH v8 2/5] drm/sched: Convert drm scheduler to use a work queue rather than kthread Matthew Brost
2023-10-31  3:24 ` [PATCH v8 3/5] drm/sched: Split free_job into own work item Matthew Brost
2023-11-01 22:13   ` Luben Tuikov
2023-11-02 11:13   ` Tvrtko Ursulin
2023-11-02 22:46     ` [PATCH] drm/sched: Eliminate drm_sched_run_job_queue_if_ready() Luben Tuikov
2023-11-03 10:39       ` Tvrtko Ursulin
2023-11-04  0:25         ` Luben Tuikov
2023-11-06 12:54           ` Tvrtko Ursulin
2023-11-03 15:13       ` Matthew Brost
2023-11-04  0:24         ` Luben Tuikov
2023-11-02 22:58     ` [PATCH v8 3/5] drm/sched: Split free_job into own work item Luben Tuikov
2023-11-07  4:10     ` [PATCH] drm/sched: Don't disturb the entity when in RR-mode scheduling Luben Tuikov
2023-11-07 11:48       ` Matthew Brost
2023-11-08  3:28         ` Luben Tuikov
2023-11-07 17:53       ` Danilo Krummrich
2023-11-08  3:29         ` Luben Tuikov
2023-11-08  0:41       ` Danilo Krummrich
2023-11-09  6:52         ` Luben Tuikov
2023-11-09 19:24           ` Danilo Krummrich
2023-11-09 23:41             ` Danilo Krummrich
2023-11-09 23:49               ` Luben Tuikov [this message]
2023-11-27 13:30                 ` [PATCH] Revert "drm/sched: Qualify drm_sched_wakeup() by drm_sched_entity_is_ready()" Bert Karwatzki
2023-11-27 15:14                   ` Luben Tuikov
2023-10-31  3:24 ` [PATCH v8 4/5] drm/sched: Add drm_sched_start_timeout_unlocked helper Matthew Brost
2023-10-31  3:24 ` [PATCH v8 5/5] drm/sched: Add a helper to queue TDR immediately Matthew Brost
2023-11-01 22:16 ` [PATCH v8 0/5] DRM scheduler changes for Xe Luben Tuikov

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=6273fadf-267a-4965-82ab-89c5b3f28cf2@gmail.com \
    --to=ltuikov89@gmail.com \
    --cc=Liviu.Dudau@arm.com \
    --cc=boris.brezillon@collabora.com \
    --cc=christian.koenig@amd.com \
    --cc=dakr@redhat.com \
    --cc=donald.robson@imgtec.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=faith.ekstrand@collabora.com \
    --cc=intel-xe@lists.freedesktop.org \
    --cc=ketil.johnsen@arm.com \
    --cc=lina@asahilina.net \
    --cc=matthew.brost@intel.com \
    --cc=mcanal@igalia.com \
    --cc=robdclark@chromium.org \
    --cc=sarah.walker@imgtec.com \
    --cc=tvrtko.ursulin@linux.intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox