AMD-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Andrey Grodzovsky <andrey.grodzovsky@amd.com>
To: "Chen, JingWen" <JingWen.Chen2@amd.com>
Cc: "Christian König" <ckoenig.leichtzumerken@gmail.com>,
	"Chen, Horace" <Horace.Chen@amd.com>,
	"Lazar, Lijo" <Lijo.Lazar@amd.com>,
	"amd-gfx@lists.freedesktop.org" <amd-gfx@lists.freedesktop.org>,
	"daniel@ffwll.ch" <daniel@ffwll.ch>,
	"Deucher, Alexander" <Alexander.Deucher@amd.com>,
	"Koenig, Christian" <Christian.Koenig@amd.com>,
	"Liu, Monk" <Monk.Liu@amd.com>
Subject: Re: [RFC v4 02/11] drm/amdgpu: Move scheduler init to after XGMI is ready
Date: Wed, 2 Mar 2022 22:16:47 -0500	[thread overview]
Message-ID: <7e9dc4ea-3665-7632-280f-9e8ed8948b45@amd.com> (raw)
In-Reply-To: <C3715E75-B013-409C-B2A3-E10CD79FD027@amd.com>

OK, i will do quick smoke test tomorrow and push all of it it then.

Andrey

On 2022-03-02 21:59, Chen, JingWen wrote:
> Hi Andrey,
>
> I don't have the bare mental environment, I can only test the SRIOV cases.
>
> Best Regards,
> JingWen Chen
>
>
>
>> On Mar 3, 2022, at 01:55, Grodzovsky, Andrey <Andrey.Grodzovsky@amd.com> wrote:
>>
>> The patch is acked-by: Andrey Grodzovsky <andrey.grodzovsky@amd.com>
>>
>> If you also smoked tested bare metal feel free to apply all the patches, if no let me know.
>>
>> Andrey
>>
>> On 2022-03-02 04:51, JingWen Chen wrote:
>>> Hi Andrey,
>>>
>>> Most part of the patches are OK, but the code will introduce a ib test fail on the disabled vcn of sienna_cichlid.
>>>
>>> In SRIOV use case we will disable one vcn on sienna_cichlid, I have attached a patch to fix this issue, please check the attachment.
>>>
>>> Best Regards,
>>>
>>> Jingwen Chen
>>>
>>>
>>> On 2/26/22 5:22 AM, Andrey Grodzovsky wrote:
>>>> Hey, patches attached - i applied the patches and resolved merge conflicts but weren't able to test as my on board's network card doesn't work with 5.16 kernel (it does with 5.17, maybe it's Kconfig issue and i need to check more).
>>>> The patches are on top of 'cababde192b2 Yifan Zhang         2 days ago     drm/amd/pm: fix mode2 reset fail for smu 13.0.5 ' commit.
>>>>
>>>> Please test and let me know. Maybe by Monday I will be able to resolve the connectivity issue on 5.16.
>>>>
>>>> Andrey
>>>>
>>>> On 2022-02-24 22:13, JingWen Chen wrote:
>>>>> Hi Andrey,
>>>>>
>>>>> Sorry for the misleading, I mean the whole patch series. We are depending on this patch series to fix the concurrency issue within SRIOV TDR sequence.
>>>>>
>>>>>
>>>>>
>>>>> On 2/25/22 1:26 AM, Andrey Grodzovsky wrote:
>>>>>> No problem if so but before I do,
>>>>>>
>>>>>>
>>>>>> JingWen - why you think this patch is needed as a standalone now ? It has no use without the
>>>>>> entire feature together with it. Is it some changes you want to do on top of that code ?
>>>>>>
>>>>>>
>>>>>> Andrey
>>>>>>
>>>>>>
>>>>>> On 2022-02-24 12:12, Deucher, Alexander wrote:
>>>>>>> [Public]
>>>>>>>
>>>>>>>
>>>>>>> If it applies cleanly, feel free to drop it in.  I'll drop those patches for drm-next since they are already in drm-misc.
>>>>>>>
>>>>>>> Alex
>>>>>>>
>>>>>>> ------------------------------------------------------------------------
>>>>>>> *From:* amd-gfx <amd-gfx-bounces@lists.freedesktop.org> on behalf of Andrey Grodzovsky <andrey.grodzovsky@amd.com>
>>>>>>> *Sent:* Thursday, February 24, 2022 11:24 AM
>>>>>>> *To:* Chen, JingWen <JingWen.Chen2@amd.com>; Christian König <ckoenig.leichtzumerken@gmail.com>; dri-devel@lists.freedesktop.org <dri-devel@lists.freedesktop.org>; amd-gfx@lists.freedesktop.org <amd-gfx@lists.freedesktop.org>
>>>>>>> *Cc:* Liu, Monk <Monk.Liu@amd.com>; Chen, Horace <Horace.Chen@amd.com>; Lazar, Lijo <Lijo.Lazar@amd.com>; Koenig, Christian <Christian.Koenig@amd.com>; daniel@ffwll.ch <daniel@ffwll.ch>
>>>>>>> *Subject:* Re: [RFC v4 02/11] drm/amdgpu: Move scheduler init to after XGMI is ready
>>>>>>> No because all the patch-set including this patch was landed into
>>>>>>> drm-misc-next and will reach amd-staging-drm-next on the next upstream
>>>>>>> rebase i guess.
>>>>>>>
>>>>>>> Andrey
>>>>>>>
>>>>>>> On 2022-02-24 01:47, JingWen Chen wrote:
>>>>>>>> Hi Andrey,
>>>>>>>>
>>>>>>>> Will you port this patch into amd-staging-drm-next?
>>>>>>>>
>>>>>>>> on 2/10/22 2:06 AM, Andrey Grodzovsky wrote:
>>>>>>>>> All comments are fixed and code pushed. Thanks for everyone
>>>>>>>>> who helped reviewing.
>>>>>>>>>
>>>>>>>>> Andrey
>>>>>>>>>
>>>>>>>>> On 2022-02-09 02:53, Christian König wrote:
>>>>>>>>>> Am 09.02.22 um 01:23 schrieb Andrey Grodzovsky:
>>>>>>>>>>> Before we initialize schedulers we must know which reset
>>>>>>>>>>> domain are we in - for single device there iis a single
>>>>>>>>>>> domain per device and so single wq per device. For XGMI
>>>>>>>>>>> the reset domain spans the entire XGMI hive and so the
>>>>>>>>>>> reset wq is per hive.
>>>>>>>>>>>
>>>>>>>>>>> Signed-off-by: Andrey Grodzovsky <andrey.grodzovsky@amd.com>
>>>>>>>>>> One more comment below, with that fixed Reviewed-by: Christian König <christian.koenig@amd.com>.
>>>>>>>>>>
>>>>>>>>>>> ---
>>>>>>>>>>> drivers/gpu/drm/amd/amdgpu/amdgpu_device.c | 45 ++++++++++++++++++++++
>>>>>>>>>>> drivers/gpu/drm/amd/amdgpu/amdgpu_fence.c  | 34 ++--------------
>>>>>>>>>>> drivers/gpu/drm/amd/amdgpu/amdgpu_ring.h   |  2 +
>>>>>>>>>>>       3 files changed, 51 insertions(+), 30 deletions(-)
>>>>>>>>>>>
>>>>>>>>>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
>>>>>>>>>>> index 9704b0e1fd82..00123b0013d3 100644
>>>>>>>>>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
>>>>>>>>>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_device.c
>>>>>>>>>>> @@ -2287,6 +2287,47 @@ static int amdgpu_device_fw_loading(struct amdgpu_device *adev)
>>>>>>>>>>>           return r;
>>>>>>>>>>>       }
>>>>>>>>>>>       +static int amdgpu_device_init_schedulers(struct amdgpu_device *adev)
>>>>>>>>>>> +{
>>>>>>>>>>> +    long timeout;
>>>>>>>>>>> +    int r, i;
>>>>>>>>>>> +
>>>>>>>>>>> +    for (i = 0; i < AMDGPU_MAX_RINGS; ++i) {
>>>>>>>>>>> +        struct amdgpu_ring *ring = adev->rings[i];
>>>>>>>>>>> +
>>>>>>>>>>> +        /* No need to setup the GPU scheduler for rings that don't need it */
>>>>>>>>>>> +        if (!ring || ring->no_scheduler)
>>>>>>>>>>> +            continue;
>>>>>>>>>>> +
>>>>>>>>>>> +        switch (ring->funcs->type) {
>>>>>>>>>>> +        case AMDGPU_RING_TYPE_GFX:
>>>>>>>>>>> +            timeout = adev->gfx_timeout;
>>>>>>>>>>> +            break;
>>>>>>>>>>> +        case AMDGPU_RING_TYPE_COMPUTE:
>>>>>>>>>>> +            timeout = adev->compute_timeout;
>>>>>>>>>>> +            break;
>>>>>>>>>>> +        case AMDGPU_RING_TYPE_SDMA:
>>>>>>>>>>> +            timeout = adev->sdma_timeout;
>>>>>>>>>>> +            break;
>>>>>>>>>>> +        default:
>>>>>>>>>>> +            timeout = adev->video_timeout;
>>>>>>>>>>> +            break;
>>>>>>>>>>> +        }
>>>>>>>>>>> +
>>>>>>>>>>> +        r = drm_sched_init(&ring->sched, &amdgpu_sched_ops,
>>>>>>>>>>> + ring->num_hw_submission, amdgpu_job_hang_limit,
>>>>>>>>>>> +                   timeout, adev->reset_domain.wq, ring->sched_score, ring->name);
>>>>>>>>>>> +        if (r) {
>>>>>>>>>>> +            DRM_ERROR("Failed to create scheduler on ring %s.\n",
>>>>>>>>>>> +                  ring->name);
>>>>>>>>>>> +            return r;
>>>>>>>>>>> +        }
>>>>>>>>>>> +    }
>>>>>>>>>>> +
>>>>>>>>>>> +    return 0;
>>>>>>>>>>> +}
>>>>>>>>>>> +
>>>>>>>>>>> +
>>>>>>>>>>>       /**
>>>>>>>>>>>        * amdgpu_device_ip_init - run init for hardware IPs
>>>>>>>>>>>        *
>>>>>>>>>>> @@ -2419,6 +2460,10 @@ static int amdgpu_device_ip_init(struct amdgpu_device *adev)
>>>>>>>>>>>               }
>>>>>>>>>>>           }
>>>>>>>>>>>       +    r = amdgpu_device_init_schedulers(adev);
>>>>>>>>>>> +    if (r)
>>>>>>>>>>> +        goto init_failed;
>>>>>>>>>>> +
>>>>>>>>>>>           /* Don't init kfd if whole hive need to be reset during init */
>>>>>>>>>>>           if (!adev->gmc.xgmi.pending_reset)
>>>>>>>>>>> amdgpu_amdkfd_device_init(adev);
>>>>>>>>>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_fence.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_fence.c
>>>>>>>>>>> index 45977a72b5dd..fa302540c69a 100644
>>>>>>>>>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_fence.c
>>>>>>>>>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_fence.c
>>>>>>>>>>> @@ -457,8 +457,6 @@ int amdgpu_fence_driver_init_ring(struct amdgpu_ring *ring,
>>>>>>>>>>>                         atomic_t *sched_score)
>>>>>>>>>>>       {
>>>>>>>>>>>           struct amdgpu_device *adev = ring->adev;
>>>>>>>>>>> -    long timeout;
>>>>>>>>>>> -    int r;
>>>>>>>>>>>             if (!adev)
>>>>>>>>>>>               return -EINVAL;
>>>>>>>>>>> @@ -478,36 +476,12 @@ int amdgpu_fence_driver_init_ring(struct amdgpu_ring *ring,
>>>>>>>>>>> spin_lock_init(&ring->fence_drv.lock);
>>>>>>>>>>>           ring->fence_drv.fences = kcalloc(num_hw_submission * 2, sizeof(void *),
>>>>>>>>>>>                            GFP_KERNEL);
>>>>>>>>>>> -    if (!ring->fence_drv.fences)
>>>>>>>>>>> -        return -ENOMEM;
>>>>>>>>>>>       -    /* No need to setup the GPU scheduler for rings that don't need it */
>>>>>>>>>>> -    if (ring->no_scheduler)
>>>>>>>>>>> -        return 0;
>>>>>>>>>>> +    ring->num_hw_submission = num_hw_submission;
>>>>>>>>>>> +    ring->sched_score = sched_score;
>>>>>>>>>> Let's move this into the caller and then use ring->num_hw_submission in the fence code as well.
>>>>>>>>>>
>>>>>>>>>> The maximum number of jobs on the ring is not really fence specific.
>>>>>>>>>>
>>>>>>>>>> Regards,
>>>>>>>>>> Christian.
>>>>>>>>>>
>>>>>>>>>>>       -    switch (ring->funcs->type) {
>>>>>>>>>>> -    case AMDGPU_RING_TYPE_GFX:
>>>>>>>>>>> -        timeout = adev->gfx_timeout;
>>>>>>>>>>> -        break;
>>>>>>>>>>> -    case AMDGPU_RING_TYPE_COMPUTE:
>>>>>>>>>>> -        timeout = adev->compute_timeout;
>>>>>>>>>>> -        break;
>>>>>>>>>>> -    case AMDGPU_RING_TYPE_SDMA:
>>>>>>>>>>> -        timeout = adev->sdma_timeout;
>>>>>>>>>>> -        break;
>>>>>>>>>>> -    default:
>>>>>>>>>>> -        timeout = adev->video_timeout;
>>>>>>>>>>> -        break;
>>>>>>>>>>> -    }
>>>>>>>>>>> -
>>>>>>>>>>> -    r = drm_sched_init(&ring->sched, &amdgpu_sched_ops,
>>>>>>>>>>> -               num_hw_submission, amdgpu_job_hang_limit,
>>>>>>>>>>> -               timeout, NULL, sched_score, ring->name);
>>>>>>>>>>> -    if (r) {
>>>>>>>>>>> -        DRM_ERROR("Failed to create scheduler on ring %s.\n",
>>>>>>>>>>> -              ring->name);
>>>>>>>>>>> -        return r;
>>>>>>>>>>> -    }
>>>>>>>>>>> +    if (!ring->fence_drv.fences)
>>>>>>>>>>> +        return -ENOMEM;
>>>>>>>>>>>             return 0;
>>>>>>>>>>>       }
>>>>>>>>>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ring.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_ring.h
>>>>>>>>>>> index fae7d185ad0d..7f20ce73a243 100644
>>>>>>>>>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ring.h
>>>>>>>>>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ring.h
>>>>>>>>>>> @@ -251,6 +251,8 @@ struct amdgpu_ring {
>>>>>>>>>>>           bool has_compute_vm_bug;
>>>>>>>>>>>           bool            no_scheduler;
>>>>>>>>>>>           int            hw_prio;
>>>>>>>>>>> +    unsigned num_hw_submission;
>>>>>>>>>>> +    atomic_t        *sched_score;
>>>>>>>>>>>       };
>>>>>>>>>>>         #define amdgpu_ring_parse_cs(r, p, ib) ((r)->funcs->parse_cs((p), (ib)))

  reply	other threads:[~2022-03-03  3:16 UTC|newest]

Thread overview: 34+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2022-02-09  0:23 [RFC v4 00/11] Define and use reset domain for GPU recovery in amdgpu Andrey Grodzovsky
2022-02-09  0:23 ` [RFC v4 01/11] drm/amdgpu: Introduce reset domain Andrey Grodzovsky
2022-02-09  0:23 ` [RFC v4 02/11] drm/amdgpu: Move scheduler init to after XGMI is ready Andrey Grodzovsky
2022-02-09  7:53   ` Christian König
2022-02-09 18:06     ` Andrey Grodzovsky
2022-02-24  6:47       ` JingWen Chen
2022-02-24 16:24         ` Andrey Grodzovsky
2022-02-24 17:12           ` Deucher, Alexander
2022-02-24 17:26             ` Andrey Grodzovsky
2022-02-25  3:13               ` JingWen Chen
2022-02-25 21:22                 ` Andrey Grodzovsky
2022-03-02  9:51                   ` JingWen Chen
2022-03-02 17:55                     ` Andrey Grodzovsky
2022-03-03  2:59                       ` Chen, JingWen
2022-03-03  3:16                         ` Andrey Grodzovsky [this message]
2022-03-03 16:36                           ` Andrey Grodzovsky
2022-03-04  6:35                             ` Chen, JingWen
2022-03-04 10:18                     ` Christian König
2022-02-09  0:23 ` [RFC v4 03/11] drm/amdgpu: Serialize non TDR gpu recovery with TDRs Andrey Grodzovsky
2022-02-09  0:23 ` [RFC v4 04/11] drm/amd/virt: For SRIOV send GPU reset directly to TDR queue Andrey Grodzovsky
2022-02-09  0:49   ` Liu, Shaoyun
2022-02-09  7:54   ` Christian König
2022-02-09  0:23 ` [RFC v4 05/11] drm/amdgpu: Drop hive->in_reset Andrey Grodzovsky
2022-02-09  0:23 ` [RFC v4 06/11] drm/amdgpu: Drop concurrent GPU reset protection for device Andrey Grodzovsky
2022-02-09  0:23 ` [RFC v4 07/11] drm/amdgpu: Rework reset domain to be refcounted Andrey Grodzovsky
2022-02-09  7:57   ` Christian König
2022-02-09  0:23 ` [RFC v4 08/11] drm/amdgpu: Move reset sem into reset_domain Andrey Grodzovsky
2022-02-09  7:59   ` Christian König
2022-02-09  0:23 ` [RFC v4 09/11] drm/amdgpu: Move in_gpu_reset " Andrey Grodzovsky
2022-02-09  8:00   ` Christian König
2022-02-09  0:23 ` [RFC v4 10/11] drm/amdgpu: Rework amdgpu_device_lock_adev Andrey Grodzovsky
2022-02-09  8:04   ` Christian König
2022-02-09  0:23 ` [RFC v4 11/11] Revert 'drm/amdgpu: annotate a false positive recursive locking' Andrey Grodzovsky
2022-02-09  8:06   ` Christian König

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=7e9dc4ea-3665-7632-280f-9e8ed8948b45@amd.com \
    --to=andrey.grodzovsky@amd.com \
    --cc=Alexander.Deucher@amd.com \
    --cc=Christian.Koenig@amd.com \
    --cc=Horace.Chen@amd.com \
    --cc=JingWen.Chen2@amd.com \
    --cc=Lijo.Lazar@amd.com \
    --cc=Monk.Liu@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    --cc=ckoenig.leichtzumerken@gmail.com \
    --cc=daniel@ffwll.ch \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox