AMD-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs
@ 2026-09-28  6:11 Chi Hong
  2026-09-29  7:25 ` Hong, Chi
  0 siblings, 1 reply; 6+ messages in thread
From: Chi Hong @ 2026-09-28  6:11 UTC (permalink / raw)
  To: amd-gfx; +Cc: Chi Hong

SR-IOV VFs use amdgpu_virt_ras_hw_init() which skips ras_core_hw_init(),
leaving fw_gpu_mem unallocated. ras_psp_reload_firmwares() then hits the
NULL buffer and returns -ENOMEM (-12) for every VF on resume, aborting
ip_resume_phase2 before amdgpu_fence_driver_hw_init() runs. The next
hibernate double-puts IRQ refcounts causing a WARN storm.

Skip the reload for VFs — there is nothing local to reload. The
post-reset path (amdgpu_ras_mgr_resume_after_reset) is unaffected.

Fixes: 5e5be56a5b76 ("drm/amdgpu/ras: reload RAS TA from ras mgr resume")
Signed-off-by: Chi Hong <Chi.Hong@amd.com>
---
 drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
index a02167397ce2..5bac792554f5 100644
--- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
+++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
@@ -529,6 +529,9 @@ static int amdgpu_ras_mgr_resume(struct amdgpu_ip_block *ip_block)
 	if (!ras_mgr->ras_is_ready)
 		return 0;
 
+	if (amdgpu_sriov_vf(adev))
+		return 0;
+
 	ret = ras_psp_reload_firmwares(ras_mgr->ras_core, 0);
 	if (ret)
 		RAS_DEV_ERR(adev,
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 6+ messages in thread

* RE: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs
  2026-09-28  6:11 [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs Chi Hong
@ 2026-09-29  7:25 ` Hong, Chi
  2026-09-30  6:32   ` Zhang, Tiantian (Celine)
  2026-09-30 11:25   ` Lazar, Lijo
  0 siblings, 2 replies; 6+ messages in thread
From: Hong, Chi @ 2026-09-29  7:25 UTC (permalink / raw)
  To: Hong, Chi, amd-gfx@lists.freedesktop.org
  Cc: Lazar, Lijo, Zhang, Hawking, Zhang, Tiantian (Celine)

AMD General

Ping

-----Original Message-----
From: Chi Hong <Chi.Hong@amd.com>
Sent: Monday, September 28, 2026 2:12 PM
To: amd-gfx@lists.freedesktop.org
Cc: Hong, Chi <Chi.Hong@amd.com>
Subject: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs

SR-IOV VFs use amdgpu_virt_ras_hw_init() which skips ras_core_hw_init(), leaving fw_gpu_mem unallocated. ras_psp_reload_firmwares() then hits the NULL buffer and returns -ENOMEM (-12) for every VF on resume, aborting
ip_resume_phase2 before amdgpu_fence_driver_hw_init() runs. The next hibernate double-puts IRQ refcounts causing a WARN storm.

Skip the reload for VFs - there is nothing local to reload. The post-reset path (amdgpu_ras_mgr_resume_after_reset) is unaffected.

Fixes: 5e5be56a5b76 ("drm/amdgpu/ras: reload RAS TA from ras mgr resume")
Signed-off-by: Chi Hong <Chi.Hong@amd.com>
---
 drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
index a02167397ce2..5bac792554f5 100644
--- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
+++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
@@ -529,6 +529,9 @@ static int amdgpu_ras_mgr_resume(struct amdgpu_ip_block *ip_block)
        if (!ras_mgr->ras_is_ready)
                return 0;

+       if (amdgpu_sriov_vf(adev))
+               return 0;
+
        ret = ras_psp_reload_firmwares(ras_mgr->ras_core, 0);
        if (ret)
                RAS_DEV_ERR(adev,
--
2.43.0


^ permalink raw reply related	[flat|nested] 6+ messages in thread

* RE: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs
  2026-09-29  7:25 ` Hong, Chi
@ 2026-09-30  6:32   ` Zhang, Tiantian (Celine)
  2026-09-30 11:25   ` Lazar, Lijo
  1 sibling, 0 replies; 6+ messages in thread
From: Zhang, Tiantian (Celine) @ 2026-09-30  6:32 UTC (permalink / raw)
  To: Hong, Chi, amd-gfx@lists.freedesktop.org, Lazar, Lijo; +Cc: Zhang, Hawking

AMD General

Hi @Lazar, Lijo,

Could you please help to review the patch Chi provided?
You can contact Chi directly if you have any questions, thanks.


Best Regards,
Celine Zhang
-----Original Message-----
From: Hong, Chi <Chi.Hong@amd.com>
Sent: Tuesday, September 29, 2026 3:25 PM
To: Hong, Chi <Chi.Hong@amd.com>; amd-gfx@lists.freedesktop.org
Cc: Lazar, Lijo <Lijo.Lazar@amd.com>; Zhang, Hawking <Hawking.Zhang@amd.com>; Zhang, Tiantian (Celine) <Tiantian.Zhang@amd.com>
Subject: RE: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs

AMD General

Ping

-----Original Message-----
From: Chi Hong <Chi.Hong@amd.com>
Sent: Monday, September 28, 2026 2:12 PM
To: amd-gfx@lists.freedesktop.org
Cc: Hong, Chi <Chi.Hong@amd.com>
Subject: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs

SR-IOV VFs use amdgpu_virt_ras_hw_init() which skips ras_core_hw_init(), leaving fw_gpu_mem unallocated. ras_psp_reload_firmwares() then hits the NULL buffer and returns -ENOMEM (-12) for every VF on resume, aborting
ip_resume_phase2 before amdgpu_fence_driver_hw_init() runs. The next hibernate double-puts IRQ refcounts causing a WARN storm.

Skip the reload for VFs - there is nothing local to reload. The post-reset path (amdgpu_ras_mgr_resume_after_reset) is unaffected.

Fixes: 5e5be56a5b76 ("drm/amdgpu/ras: reload RAS TA from ras mgr resume")
Signed-off-by: Chi Hong <Chi.Hong@amd.com>
---
 drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
index a02167397ce2..5bac792554f5 100644
--- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
+++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
@@ -529,6 +529,9 @@ static int amdgpu_ras_mgr_resume(struct amdgpu_ip_block *ip_block)
        if (!ras_mgr->ras_is_ready)
                return 0;

+       if (amdgpu_sriov_vf(adev))
+               return 0;
+
        ret = ras_psp_reload_firmwares(ras_mgr->ras_core, 0);
        if (ret)
                RAS_DEV_ERR(adev,
--
2.43.0



^ permalink raw reply related	[flat|nested] 6+ messages in thread

* Re: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs
  2026-09-29  7:25 ` Hong, Chi
  2026-09-30  6:32   ` Zhang, Tiantian (Celine)
@ 2026-09-30 11:25   ` Lazar, Lijo
  2026-10-08  2:00     ` Hong, Chi
  1 sibling, 1 reply; 6+ messages in thread
From: Lazar, Lijo @ 2026-09-30 11:25 UTC (permalink / raw)
  To: Hong, Chi, amd-gfx@lists.freedesktop.org
  Cc: Zhang, Hawking, Zhang, Tiantian (Celine), Zhou1, Tao, YiPeng Chai

cc: Tao/Thomas to take a look.

Thanks,
Lijo

On 29-Sep-26 12:55 PM, Hong, Chi wrote:
> AMD General
> 
> Ping
> 
> -----Original Message-----
> From: Chi Hong <Chi.Hong@amd.com>
> Sent: Monday, September 28, 2026 2:12 PM
> To: amd-gfx@lists.freedesktop.org
> Cc: Hong, Chi <Chi.Hong@amd.com>
> Subject: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs
> 
> SR-IOV VFs use amdgpu_virt_ras_hw_init() which skips ras_core_hw_init(), leaving fw_gpu_mem unallocated. ras_psp_reload_firmwares() then hits the NULL buffer and returns -ENOMEM (-12) for every VF on resume, aborting
> ip_resume_phase2 before amdgpu_fence_driver_hw_init() runs. The next hibernate double-puts IRQ refcounts causing a WARN storm.
> 
> Skip the reload for VFs - there is nothing local to reload. The post-reset path (amdgpu_ras_mgr_resume_after_reset) is unaffected.
> 
> Fixes: 5e5be56a5b76 ("drm/amdgpu/ras: reload RAS TA from ras mgr resume")
> Signed-off-by: Chi Hong <Chi.Hong@amd.com>
> ---
>   drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c | 3 +++
>   1 file changed, 3 insertions(+)
> 
> diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
> index a02167397ce2..5bac792554f5 100644
> --- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
> +++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
> @@ -529,6 +529,9 @@ static int amdgpu_ras_mgr_resume(struct amdgpu_ip_block *ip_block)
>          if (!ras_mgr->ras_is_ready)
>                  return 0;
> 
> +       if (amdgpu_sriov_vf(adev))
> +               return 0;
> +
>          ret = ras_psp_reload_firmwares(ras_mgr->ras_core, 0);
>          if (ret)
>                  RAS_DEV_ERR(adev,
> --
> 2.43.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* RE: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs
  2026-09-30 11:25   ` Lazar, Lijo
@ 2026-10-08  2:00     ` Hong, Chi
  2026-10-08  2:35       ` Chai, Thomas
  0 siblings, 1 reply; 6+ messages in thread
From: Hong, Chi @ 2026-10-08  2:00 UTC (permalink / raw)
  To: Lazar, Lijo, amd-gfx@lists.freedesktop.org, Zhou1, Tao,
	Chai, Thomas
  Cc: Zhang, Hawking, Zhang, Tiantian (Celine)

AMD General

Hi

@Chai, Thomas@Zhou1, Tao. Please take a look at this patch.

Thanks,
Chi

-----Original Message-----
From: Lazar, Lijo <Lijo.Lazar@amd.com>
Sent: Wednesday, September 30, 2026 7:26 PM
To: Hong, Chi <Chi.Hong@amd.com>; amd-gfx@lists.freedesktop.org
Cc: Zhang, Hawking <Hawking.Zhang@amd.com>; Zhang, Tiantian (Celine) <Tiantian.Zhang@amd.com>; Zhou1, Tao <Tao.Zhou1@amd.com>; Chai, Thomas <YiPeng.Chai@amd.com>
Subject: Re: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs

cc: Tao/Thomas to take a look.

Thanks,
Lijo

On 29-Sep-26 12:55 PM, Hong, Chi wrote:
> AMD General
>
> Ping
>
> -----Original Message-----
> From: Chi Hong <Chi.Hong@amd.com>
> Sent: Monday, September 28, 2026 2:12 PM
> To: amd-gfx@lists.freedesktop.org
> Cc: Hong, Chi <Chi.Hong@amd.com>
> Subject: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs
>
> SR-IOV VFs use amdgpu_virt_ras_hw_init() which skips ras_core_hw_init(), leaving fw_gpu_mem unallocated. ras_psp_reload_firmwares() then hits the NULL buffer and returns -ENOMEM (-12) for every VF on resume, aborting
> ip_resume_phase2 before amdgpu_fence_driver_hw_init() runs. The next hibernate double-puts IRQ refcounts causing a WARN storm.
>
> Skip the reload for VFs - there is nothing local to reload. The post-reset path (amdgpu_ras_mgr_resume_after_reset) is unaffected.
>
> Fixes: 5e5be56a5b76 ("drm/amdgpu/ras: reload RAS TA from ras mgr resume")
> Signed-off-by: Chi Hong <Chi.Hong@amd.com>
> ---
>   drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c | 3 +++
>   1 file changed, 3 insertions(+)
>
> diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
> index a02167397ce2..5bac792554f5 100644
> --- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
> +++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
> @@ -529,6 +529,9 @@ static int amdgpu_ras_mgr_resume(struct amdgpu_ip_block *ip_block)
>          if (!ras_mgr->ras_is_ready)
>                  return 0;
>
> +       if (amdgpu_sriov_vf(adev))
> +               return 0;
> +
>          ret = ras_psp_reload_firmwares(ras_mgr->ras_core, 0);
>          if (ret)
>                  RAS_DEV_ERR(adev,
> --
> 2.43.0


^ permalink raw reply	[flat|nested] 6+ messages in thread

* RE: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs
  2026-10-08  2:00     ` Hong, Chi
@ 2026-10-08  2:35       ` Chai, Thomas
  0 siblings, 0 replies; 6+ messages in thread
From: Chai, Thomas @ 2026-10-08  2:35 UTC (permalink / raw)
  To: Hong, Chi, Lazar, Lijo, amd-gfx@lists.freedesktop.org, Zhou1, Tao
  Cc: Zhang, Hawking, Zhang, Tiantian (Celine)

AMD General

Reviewed-by: YiPeng Chai <YiPeng.Chai@amd.com>

Best Regards,
Thomas
-----Original Message-----
From: Hong, Chi <Chi.Hong@amd.com>
Sent: Thursday, October 8, 2026 10:01 AM
To: Lazar, Lijo <Lijo.Lazar@amd.com>; amd-gfx@lists.freedesktop.org; Zhou1, Tao <Tao.Zhou1@amd.com>; Chai, Thomas <YiPeng.Chai@amd.com>
Cc: Zhang, Hawking <Hawking.Zhang@amd.com>; Zhang, Tiantian (Celine) <Tiantian.Zhang@amd.com>
Subject: RE: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs

AMD General

Hi

@Chai, Thomas@Zhou1, Tao. Please take a look at this patch.

Thanks,
Chi

-----Original Message-----
From: Lazar, Lijo <Lijo.Lazar@amd.com>
Sent: Wednesday, September 30, 2026 7:26 PM
To: Hong, Chi <Chi.Hong@amd.com>; amd-gfx@lists.freedesktop.org
Cc: Zhang, Hawking <Hawking.Zhang@amd.com>; Zhang, Tiantian (Celine) <Tiantian.Zhang@amd.com>; Zhou1, Tao <Tao.Zhou1@amd.com>; Chai, Thomas <YiPeng.Chai@amd.com>
Subject: Re: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs

cc: Tao/Thomas to take a look.

Thanks,
Lijo

On 29-Sep-26 12:55 PM, Hong, Chi wrote:
> AMD General
>
> Ping
>
> -----Original Message-----
> From: Chi Hong <Chi.Hong@amd.com>
> Sent: Monday, September 28, 2026 2:12 PM
> To: amd-gfx@lists.freedesktop.org
> Cc: Hong, Chi <Chi.Hong@amd.com>
> Subject: [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs
>
> SR-IOV VFs use amdgpu_virt_ras_hw_init() which skips ras_core_hw_init(), leaving fw_gpu_mem unallocated. ras_psp_reload_firmwares() then hits the NULL buffer and returns -ENOMEM (-12) for every VF on resume, aborting
> ip_resume_phase2 before amdgpu_fence_driver_hw_init() runs. The next hibernate double-puts IRQ refcounts causing a WARN storm.
>
> Skip the reload for VFs - there is nothing local to reload. The post-reset path (amdgpu_ras_mgr_resume_after_reset) is unaffected.
>
> Fixes: 5e5be56a5b76 ("drm/amdgpu/ras: reload RAS TA from ras mgr resume")
> Signed-off-by: Chi Hong <Chi.Hong@amd.com>
> ---
>   drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c | 3 +++
>   1 file changed, 3 insertions(+)
>
> diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
> index a02167397ce2..5bac792554f5 100644
> --- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
> +++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_mgr.c
> @@ -529,6 +529,9 @@ static int amdgpu_ras_mgr_resume(struct amdgpu_ip_block *ip_block)
>          if (!ras_mgr->ras_is_ready)
>                  return 0;
>
> +       if (amdgpu_sriov_vf(adev))
> +               return 0;
> +
>          ret = ras_psp_reload_firmwares(ras_mgr->ras_core, 0);
>          if (ret)
>                  RAS_DEV_ERR(adev,
> --
> 2.43.0



^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-10-08  2:35 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28  6:11 [PATCH] drm/amd/ras: skip RAS TA firmware reload on resume for SR-IOV VFs Chi Hong
2026-09-29  7:25 ` Hong, Chi
2026-09-30  6:32   ` Zhang, Tiantian (Celine)
2026-09-30 11:25   ` Lazar, Lijo
2026-10-08  2:00     ` Hong, Chi
2026-10-08  2:35       ` Chai, Thomas

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox