* [PATCH 1/2] drm/amd/ras: tell KFD the reset came from an ECC error
@ 2026-08-24 10:21 Xiang Liu
2026-08-24 10:21 ` [PATCH 2/2] drm/amd/ras: record the fatal state on every device of the hive Xiang Liu
2026-08-24 11:01 ` [PATCH 1/2] drm/amd/ras: tell KFD the reset came from an ECC error Zhang, Hawking
0 siblings, 2 replies; 3+ messages in thread
From: Xiang Liu @ 2026-08-24 10:21 UTC (permalink / raw)
To: amd-gfx; +Cc: Hawking.Zhang, Tao.Zhou1, Stanley.Yang, YiPeng.Chai, Xiang Liu
kfd_signal_reset_event() picks between KFD_HW_EXCEPTION_ECC and
KFD_HW_EXCEPTION_GPU_HANG from the SRAM ECC flag, and only delivers the
memory exception event for the former. Nothing raises that flag on the
RAS module paths, so a reset caused by an uncorrectable or a consumed
poison error is reported to every process on the device as a plain hang
and the runtime carries on instead of tearing the workload down.
Raise it the way the per IP callbacks used to.
Signed-off-by: Xiang Liu <xiang.liu@amd.com>
---
drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c | 1 +
drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c | 2 ++
2 files changed, 3 insertions(+)
diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c
index 39452a900615..9dd44fb5b885 100644
--- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c
+++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c
@@ -89,6 +89,7 @@ int amdgpu_ras_process_handle_umc_interrupt(struct amdgpu_device *adev, void *da
int amdgpu_ras_process_handle_unexpected_interrupt(struct amdgpu_device *adev, void *data)
{
+ kgd2kfd_set_sram_ecc_flag(adev->kfd.dev);
amdgpu_ras_set_fed(adev, true);
return amdgpu_ras_mgr_reset_gpu(adev, AMDGPU_RAS_GPU_RESET_MODE1_RESET);
}
diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
index afb539f068c2..081516c46cf8 100644
--- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
+++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
@@ -54,6 +54,8 @@ static int amdgpu_ras_sys_poison_consumption_event(struct ras_core_context *ras_
if (!req)
return -EINVAL;
+ kgd2kfd_set_sram_ecc_flag(adev->kfd.dev);
+
if (req->pasid_fn) {
pasid_fn = (pasid_notify)req->pasid_fn;
pasid_fn(adev, req->pasid, req->data);
--
2.34.1
^ permalink raw reply related [flat|nested] 3+ messages in thread
* [PATCH 2/2] drm/amd/ras: record the fatal state on every device of the hive
2026-08-24 10:21 [PATCH 1/2] drm/amd/ras: tell KFD the reset came from an ECC error Xiang Liu
@ 2026-08-24 10:21 ` Xiang Liu
2026-08-24 11:01 ` [PATCH 1/2] drm/amd/ras: tell KFD the reset came from an ECC error Zhang, Hawking
1 sibling, 0 replies; 3+ messages in thread
From: Xiang Liu @ 2026-08-24 10:21 UTC (permalink / raw)
To: amd-gfx; +Cc: Hawking.Zhang, Tao.Zhou1, Stanley.Yang, YiPeng.Chai, Xiang Liu
The fatal error interrupt is broadcast to every device of the hive and
they all race for amdgpu_ras_global_ras_isr(), which hands -EBUSY to
everyone but the winner. Treating that as a failure returns before the
device is marked, so seven devices out of eight are left without their
fatal and SRAM ECC state, the one that actually logged the error among
them. KFD then tells the processes on those devices that the reset was
a plain hang.
-EBUSY only means the reset has already been asked for. Record the
state anyway and leave the request to the winner.
Signed-off-by: Xiang Liu <xiang.liu@amd.com>
---
drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c | 10 ++++++++--
1 file changed, 8 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
index 081516c46cf8..291b2a96cbb5 100644
--- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
+++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
@@ -33,8 +33,14 @@ static int amdgpu_ras_sys_detect_fatal_event(struct ras_core_context *ras_core,
uint64_t seq_no;
ret = amdgpu_ras_global_ras_isr(adev);
- if (ret)
- return ret;
+ if (ret) {
+ /* Another device of the hive already asked for the reset, this
+ * one still has to record that it saw the error.
+ */
+ kgd2kfd_set_sram_ecc_flag(adev->kfd.dev);
+ amdgpu_ras_set_fed(adev, true);
+ return ret == -EBUSY ? 0 : ret;
+ }
seq_no = amdgpu_ras_mgr_gen_ras_event_seqno(adev, RAS_SEQNO_TYPE_UE);
RAS_DEV_INFO(adev,
--
2.34.1
^ permalink raw reply related [flat|nested] 3+ messages in thread
* RE: [PATCH 1/2] drm/amd/ras: tell KFD the reset came from an ECC error
2026-08-24 10:21 [PATCH 1/2] drm/amd/ras: tell KFD the reset came from an ECC error Xiang Liu
2026-08-24 10:21 ` [PATCH 2/2] drm/amd/ras: record the fatal state on every device of the hive Xiang Liu
@ 2026-08-24 11:01 ` Zhang, Hawking
1 sibling, 0 replies; 3+ messages in thread
From: Zhang, Hawking @ 2026-08-24 11:01 UTC (permalink / raw)
To: Liu, Xiang(Dean), amd-gfx@lists.freedesktop.org
Cc: Zhou1, Tao, Yang, Stanley, Chai, Thomas
AMD General
Series is
Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
Regards,
Hawking
-----Original Message-----
From: Liu, Xiang(Dean) <Xiang.Liu@amd.com>
Sent: Monday, August 24, 2026 6:21 PM
To: amd-gfx@lists.freedesktop.org
Cc: Zhang, Hawking <Hawking.Zhang@amd.com>; Zhou1, Tao <Tao.Zhou1@amd.com>; Yang, Stanley <Stanley.Yang@amd.com>; Chai, Thomas <YiPeng.Chai@amd.com>; Liu, Xiang(Dean) <Xiang.Liu@amd.com>
Subject: [PATCH 1/2] drm/amd/ras: tell KFD the reset came from an ECC error
kfd_signal_reset_event() picks between KFD_HW_EXCEPTION_ECC and KFD_HW_EXCEPTION_GPU_HANG from the SRAM ECC flag, and only delivers the memory exception event for the former. Nothing raises that flag on the RAS module paths, so a reset caused by an uncorrectable or a consumed poison error is reported to every process on the device as a plain hang and the runtime carries on instead of tearing the workload down.
Raise it the way the per IP callbacks used to.
Signed-off-by: Xiang Liu <xiang.liu@amd.com>
---
drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c | 1 +
drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c | 2 ++
2 files changed, 3 insertions(+)
diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c
index 39452a900615..9dd44fb5b885 100644
--- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c
+++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c
@@ -89,6 +89,7 @@ int amdgpu_ras_process_handle_umc_interrupt(struct amdgpu_device *adev, void *da
int amdgpu_ras_process_handle_unexpected_interrupt(struct amdgpu_device *adev, void *data) {
+ kgd2kfd_set_sram_ecc_flag(adev->kfd.dev);
amdgpu_ras_set_fed(adev, true);
return amdgpu_ras_mgr_reset_gpu(adev, AMDGPU_RAS_GPU_RESET_MODE1_RESET);
}
diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
index afb539f068c2..081516c46cf8 100644
--- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
+++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
@@ -54,6 +54,8 @@ static int amdgpu_ras_sys_poison_consumption_event(struct ras_core_context *ras_
if (!req)
return -EINVAL;
+ kgd2kfd_set_sram_ecc_flag(adev->kfd.dev);
+
if (req->pasid_fn) {
pasid_fn = (pasid_notify)req->pasid_fn;
pasid_fn(adev, req->pasid, req->data);
--
2.34.1
^ permalink raw reply related [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-08-24 11:02 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-24 10:21 [PATCH 1/2] drm/amd/ras: tell KFD the reset came from an ECC error Xiang Liu
2026-08-24 10:21 ` [PATCH 2/2] drm/amd/ras: record the fatal state on every device of the hive Xiang Liu
2026-08-24 11:01 ` [PATCH 1/2] drm/amd/ras: tell KFD the reset came from an ECC error Zhang, Hawking
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.