AMD-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: James Zhu <jamesz@amd.com>
To: "Christian König" <christian.koenig@amd.com>,
	"James Zhu" <James.Zhu@amd.com>,
	amd-gfx@lists.freedesktop.org
Subject: Re: [PATCH 1/4] drm/amdgpu/vcn: fix race condition issue for vcn start
Date: Wed, 4 Mar 2020 09:57:15 -0500	[thread overview]
Message-ID: <6294bcbb-e663-79e5-371b-28eb3a47dbfb@amd.com> (raw)
In-Reply-To: <11274ae4-da5e-5c7d-ba1a-877f22daea24@amd.com>


On 2020-03-04 3:53 a.m., Christian König wrote:
> Am 03.03.20 um 23:48 schrieb James Zhu:
>>
>> On 2020-03-03 2:03 p.m., James Zhu wrote:
>>>
>>> On 2020-03-03 1:44 p.m., Christian König wrote:
>>>> Am 03.03.20 um 19:16 schrieb James Zhu:
>>>>> Fix race condition issue when multiple vcn starts are called.
>>>>>
>>>>> Signed-off-by: James Zhu <James.Zhu@amd.com>
>>>>> ---
>>>>>   drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c | 4 ++++
>>>>>   drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.h | 1 +
>>>>>   2 files changed, 5 insertions(+)
>>>>>
>>>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c 
>>>>> b/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c
>>>>> index f96464e..aa7663f 100644
>>>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c
>>>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.c
>>>>> @@ -63,6 +63,7 @@ int amdgpu_vcn_sw_init(struct amdgpu_device *adev)
>>>>>       int i, r;
>>>>>         INIT_DELAYED_WORK(&adev->vcn.idle_work, 
>>>>> amdgpu_vcn_idle_work_handler);
>>>>> +    mutex_init(&adev->vcn.vcn_pg_lock);
>>>>>         switch (adev->asic_type) {
>>>>>       case CHIP_RAVEN:
>>>>> @@ -210,6 +211,7 @@ int amdgpu_vcn_sw_fini(struct amdgpu_device 
>>>>> *adev)
>>>>>       }
>>>>>         release_firmware(adev->vcn.fw);
>>>>> +    mutex_destroy(&adev->vcn.vcn_pg_lock);
>>>>>         return 0;
>>>>>   }
>>>>> @@ -321,6 +323,7 @@ void amdgpu_vcn_ring_begin_use(struct 
>>>>> amdgpu_ring *ring)
>>>>>       struct amdgpu_device *adev = ring->adev;
>>>>>       bool set_clocks = 
>>>>> !cancel_delayed_work_sync(&adev->vcn.idle_work);
>>>>>   +    mutex_lock(&adev->vcn.vcn_pg_lock);
>>>>
>>>> That still won't work correctly here.
>>>>
>>>> The whole idea of the cancel_delayed_work_sync() and 
>>>> schedule_delayed_work() dance is that you have exactly one user of 
>>>> that. If you have multiple rings that whole thing won't work 
>>>> correctly.
>>>>
>>>> To fix this you need to call mutex_lock() before 
>>>> cancel_delayed_work_sync() and schedule_delayed_work() before 
>>>> mutex_unlock().
>>>
>>> Big lock definitely works. I am trying to use as smaller lock as 
>>> possible here. the share resource which needs protect here are power 
>>> gate process and dpg mode switch process.
>>>
>>> if we move mutex_unlock() before schedule_delayed_work(. I am 
>>> wondering what are the other necessary resources which need protect.
>>
>> By the way, cancel_delayed_work_sync() supports multiple thread 
>> itself, so I didn't put it into protection area.
>
> Yeah, but that's correct but it still won't working correctly :)
>
> See the problem is that only for the first caller 
> cancel_delayed_work_sync() returns true because it canceled the 
> delayed work.

if the 1st caller gets true. the 2nd caller unfortunately may miss this 
pending status, so it will ungate the power which is unexpected.

But in power gate/ungate function, a power state is maintained, so this 
miss won't be really triggered to ungate the power.

So I think cancel_delayed_work_sync() / schedule_delayed_work() are not 
necessary be protected here.

Best Regards!

James

>
> For all others it returns false and those would then think that they 
> need to ungate the power.
>
> The only solution I see is to either put both the 
> cancel_delayed_work_sync() and schedule_delayed_work() under the same 
> mutex protection or start to use an atomic or other counter to note 
> concurrent processing.
>
>> power gate is shared by all VCN IP instances and different rings , so 
>> it  needs be put into protection area.
>>
>> each ring's job itself is serialized by scheduler. it doesn't need 
>> be  put into this protection area.
>
> Yes, those should work as expected.
>
> Regards,
> Christian.
>
>>
>>>
>>> Thanks!
>>>
>>> James
>>>
>>>>
>>>> Regards,
>>>> Christian.
>>>>
>>>>>       if (set_clocks) {
>>>>>           amdgpu_gfx_off_ctrl(adev, false);
>>>>>           amdgpu_device_ip_set_powergating_state(adev, 
>>>>> AMD_IP_BLOCK_TYPE_VCN,
>>>>> @@ -345,6 +348,7 @@ void amdgpu_vcn_ring_begin_use(struct 
>>>>> amdgpu_ring *ring)
>>>>>             adev->vcn.pause_dpg_mode(adev, ring->me, &new_state);
>>>>>       }
>>>>> +    mutex_unlock(&adev->vcn.vcn_pg_lock);
>>>>>   }
>>>>>     void amdgpu_vcn_ring_end_use(struct amdgpu_ring *ring)
>>>>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.h 
>>>>> b/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.h
>>>>> index 6fe0573..2ae110d 100644
>>>>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.h
>>>>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_vcn.h
>>>>> @@ -200,6 +200,7 @@ struct amdgpu_vcn {
>>>>>       struct drm_gpu_scheduler 
>>>>> *vcn_dec_sched[AMDGPU_MAX_VCN_INSTANCES];
>>>>>       uint32_t         num_vcn_enc_sched;
>>>>>       uint32_t         num_vcn_dec_sched;
>>>>> +    struct mutex         vcn_pg_lock;
>>>>>         unsigned    harvest_config;
>>>>>       int (*pause_dpg_mode)(struct amdgpu_device *adev,
>>>>
>
_______________________________________________
amd-gfx mailing list
amd-gfx@lists.freedesktop.org
https://lists.freedesktop.org/mailman/listinfo/amd-gfx

  reply	other threads:[~2020-03-04 14:57 UTC|newest]

Thread overview: 27+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2020-03-03 18:16 [PATCH 1/4] drm/amdgpu/vcn: fix race condition issue for vcn start James Zhu
2020-03-03 18:16 ` [PATCH 2/4] drm/amdgpu/vcn: fix race condition issue for dpg unpause mode switch James Zhu
2020-03-09 16:58   ` [PATCH v3 " James Zhu
2020-03-11 14:16   ` [PATCH v4 " James Zhu
2020-03-03 18:16 ` [PATCH 3/4] drm/amdgpu/vcn2.0: stall DPG when WPTR/RPTR reset James Zhu
2020-03-03 18:16 ` [PATCH 4/4] drm/amdgpu/vcn2.5: " James Zhu
2020-03-10 19:58   ` [PATCH v2 4/4] drm/amdgpu/vcn2.5: add sync " James Zhu
2020-03-10 20:03     ` Leo Liu
2020-03-03 18:44 ` [PATCH 1/4] drm/amdgpu/vcn: fix race condition issue for vcn start Christian König
2020-03-03 19:03   ` James Zhu
2020-03-03 22:48     ` James Zhu
2020-03-04  8:53       ` Christian König
2020-03-04 14:57         ` James Zhu [this message]
2020-03-04 15:03           ` Christian König
2020-03-04 15:09             ` James Zhu
2020-03-04 16:34 ` [PATCH v2 " James Zhu
2020-03-05 11:25   ` Christian König
2020-03-05 11:27     ` Christian König
2020-03-05 14:34       ` James Zhu
2020-03-09 16:57 ` [PATCH v3 " James Zhu
2020-03-11 11:30   ` Zhu, James
2020-03-11 11:38     ` Christian König
2020-03-11 14:15       ` James Zhu
2020-03-11 14:15 ` [PATCH v4 " James Zhu
2020-03-11 14:39   ` Christian König
2020-03-11 15:04 ` [PATCH v5 " James Zhu
2020-03-11 15:13   ` Christian König

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=6294bcbb-e663-79e5-371b-28eb3a47dbfb@amd.com \
    --to=jamesz@amd.com \
    --cc=James.Zhu@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    --cc=christian.koenig@amd.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox