All of lore.kernel.org
 help / color / mirror / Atom feed
From: Xiang Liu <xiang.liu@amd.com>
To: <amd-gfx@lists.freedesktop.org>
Cc: <Hawking.Zhang@amd.com>, <Tao.Zhou1@amd.com>,
	<Stanley.Yang@amd.com>, <YiPeng.Chai@amd.com>,
	Xiang Liu <xiang.liu@amd.com>
Subject: [PATCH 1/2] drm/amd/ras: tell KFD the reset came from an ECC error
Date: Mon, 24 Aug 2026 18:21:10 +0800	[thread overview]
Message-ID: <20260824102111.264161-1-xiang.liu@amd.com> (raw)

kfd_signal_reset_event() picks between KFD_HW_EXCEPTION_ECC and
KFD_HW_EXCEPTION_GPU_HANG from the SRAM ECC flag, and only delivers the
memory exception event for the former. Nothing raises that flag on the
RAS module paths, so a reset caused by an uncorrectable or a consumed
poison error is reported to every process on the device as a plain hang
and the runtime carries on instead of tearing the workload down.

Raise it the way the per IP callbacks used to.

Signed-off-by: Xiang Liu <xiang.liu@amd.com>
---
 drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c | 1 +
 drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c     | 2 ++
 2 files changed, 3 insertions(+)

diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c
index 39452a900615..9dd44fb5b885 100644
--- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c
+++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_process.c
@@ -89,6 +89,7 @@ int amdgpu_ras_process_handle_umc_interrupt(struct amdgpu_device *adev, void *da
 
 int amdgpu_ras_process_handle_unexpected_interrupt(struct amdgpu_device *adev, void *data)
 {
+	kgd2kfd_set_sram_ecc_flag(adev->kfd.dev);
 	amdgpu_ras_set_fed(adev, true);
 	return amdgpu_ras_mgr_reset_gpu(adev, AMDGPU_RAS_GPU_RESET_MODE1_RESET);
 }
diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
index afb539f068c2..081516c46cf8 100644
--- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
+++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
@@ -54,6 +54,8 @@ static int amdgpu_ras_sys_poison_consumption_event(struct ras_core_context *ras_
 	if (!req)
 		return -EINVAL;
 
+	kgd2kfd_set_sram_ecc_flag(adev->kfd.dev);
+
 	if (req->pasid_fn) {
 		pasid_fn = (pasid_notify)req->pasid_fn;
 		pasid_fn(adev, req->pasid, req->data);
-- 
2.34.1


             reply	other threads:[~2026-08-24 10:22 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-24 10:21 Xiang Liu [this message]
2026-08-24 10:21 ` [PATCH 2/2] drm/amd/ras: record the fatal state on every device of the hive Xiang Liu
2026-08-24 11:01 ` [PATCH 1/2] drm/amd/ras: tell KFD the reset came from an ECC error Zhang, Hawking

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260824102111.264161-1-xiang.liu@amd.com \
    --to=xiang.liu@amd.com \
    --cc=Hawking.Zhang@amd.com \
    --cc=Stanley.Yang@amd.com \
    --cc=Tao.Zhou1@amd.com \
    --cc=YiPeng.Chai@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.