AMD-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Lazar, Lijo" <lijo.lazar@amd.com>
To: "Zhou1, Tao" <Tao.Zhou1@amd.com>,
	"amd-gfx@lists.freedesktop.org" <amd-gfx@lists.freedesktop.org>,
	"Zhang, Hawking" <Hawking.Zhang@amd.com>,
	"Clements, John" <John.Clements@amd.com>,
	"Yang, Stanley" <Stanley.Yang@amd.com>,
	"Quan, Evan" <Evan.Quan@amd.com>
Subject: Re: [PATCH] drm/amdgpu: support new mode-1 reset interface
Date: Tue, 16 Nov 2021 14:26:36 +0530	[thread overview]
Message-ID: <33df0fe4-95db-5663-eef0-0ad23b9cb149@amd.com> (raw)
In-Reply-To: <DM6PR12MB46504CA4EDBDF823266985DCB0999@DM6PR12MB4650.namprd12.prod.outlook.com>



On 11/16/2021 2:17 PM, Zhou1, Tao wrote:
> [AMD Official Use Only]
> 
> Hi Lijo,
> 
> Your concern is reasonable, but in fact smu_v13_0_mode1_reset is used only by ALDEBARAN currently. I assume the PMFW of new smu v13 ASIC in the future will follow this design, otherwise we could move the implementation into xxx_ppt.c.
> 

Actually, this is meant to be a common logic for SMU13 based ASICs. The 
version check in a common file is not maintainable. I see there is a 
version check before also, even that is not proper :)

It is better to do it properly when support is added rather than 
thinking of refactoring with future ASICs.

Thanks,
Lijo

> Regards,
> Tao
> 
>> -----Original Message-----
>> From: Lazar, Lijo <Lijo.Lazar@amd.com>
>> Sent: Tuesday, November 16, 2021 3:44 PM
>> To: Zhou1, Tao <Tao.Zhou1@amd.com>; amd-gfx@lists.freedesktop.org; Zhang,
>> Hawking <Hawking.Zhang@amd.com>; Clements, John
>> <John.Clements@amd.com>; Yang, Stanley <Stanley.Yang@amd.com>; Quan,
>> Evan <Evan.Quan@amd.com>
>> Subject: Re: [PATCH] drm/amdgpu: support new mode-1 reset interface
>>
>>
>>
>> On 11/16/2021 12:53 PM, Tao Zhou wrote:
>>> If gpu reset is triggered by ras fatal error, tell it to smu in mode-1
>>> reset message.
>>>
>>> Signed-off-by: Tao Zhou <tao.zhou1@amd.com>
>>> ---
>>>    .../gpu/drm/amd/pm/swsmu/smu13/smu_v13_0.c    | 21
>> ++++++++++++++++---
>>>    1 file changed, 18 insertions(+), 3 deletions(-)
>>>
>>> diff --git a/drivers/gpu/drm/amd/pm/swsmu/smu13/smu_v13_0.c
>>> b/drivers/gpu/drm/amd/pm/swsmu/smu13/smu_v13_0.c
>>> index 35145db6eedf..6f3d064a8232 100644
>>> --- a/drivers/gpu/drm/amd/pm/swsmu/smu13/smu_v13_0.c
>>> +++ b/drivers/gpu/drm/amd/pm/swsmu/smu13/smu_v13_0.c
>>> @@ -1426,16 +1426,31 @@ int smu_v13_0_set_azalia_d3_pme(struct
>>> smu_context *smu)
>>>
>>>    int smu_v13_0_mode1_reset(struct smu_context *smu)
>>>    {
>>> -   u32 smu_version;
>>> +   u32 smu_version, fatal_err, param;
>>>      int ret = 0;
>>> +   struct amdgpu_device *adev = smu->adev;
>>> +   struct amdgpu_ras *ras = amdgpu_ras_get_context(adev);
>>> +
>>> +   fatal_err = 0;
>>> +   param = SMU_RESET_MODE_1;
>>> +
>>>      /*
>>>      * PM FW support SMU_MSG_GfxDeviceDriverReset from 68.07
>>>      */
>>>      smu_cmn_get_smc_version(smu, NULL, &smu_version);
>>>      if (smu_version < 0x00440700)
>>>              ret = smu_cmn_send_smc_msg(smu, SMU_MSG_Mode1Reset,
>> NULL);
>>> -   else
>>> -           ret = smu_cmn_send_smc_msg_with_param(smu,
>> SMU_MSG_GfxDeviceDriverReset, SMU_RESET_MODE_1, NULL);
>>> +   else {
>>> +           /* fatal error triggered by ras, PMFW supports the flag
>>> +              from 68.44.0 */
>>> +           if ((smu_version >= 0x00442c00) && ras &&
>>> +               atomic_read(&ras->in_recovery))
>>> +                   fatal_err = 1;
>>> +
>>
>>   From PMFW version, this looks specific to aldebaran. Since there is version
>> check as well, the implementation needs to be moved to aldebaran_ppt.c
>>
>> Thanks,
>> Lijo
>>
>>> +           param |= (fatal_err << 16);
>>> +           ret = smu_cmn_send_smc_msg_with_param(smu,
>>> +                                   SMU_MSG_GfxDeviceDriverReset,
>> param, NULL);
>>> +   }
>>>
>>>      if (!ret)
>>>              msleep(SMU13_MODE1_RESET_WAIT_TIME_IN_MS);
>>>

      reply	other threads:[~2021-11-16  8:56 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2021-11-16  7:23 [PATCH] drm/amdgpu: support new mode-1 reset interface Tao Zhou
2021-11-16  7:40 ` Zhang, Hawking
2021-11-16  7:44 ` Lazar, Lijo
2021-11-16  8:47   ` Zhou1, Tao
2021-11-16  8:56     ` Lazar, Lijo [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=33df0fe4-95db-5663-eef0-0ad23b9cb149@amd.com \
    --to=lijo.lazar@amd.com \
    --cc=Evan.Quan@amd.com \
    --cc=Hawking.Zhang@amd.com \
    --cc=John.Clements@amd.com \
    --cc=Stanley.Yang@amd.com \
    --cc=Tao.Zhou1@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox