All of lore.kernel.org
 help / color / mirror / Atom feed
From: Xiang Liu <xiang.liu@amd.com>
To: <amd-gfx@lists.freedesktop.org>
Cc: <Hawking.Zhang@amd.com>, <Tao.Zhou1@amd.com>,
	<Stanley.Yang@amd.com>, <YiPeng.Chai@amd.com>,
	Xiang Liu <xiang.liu@amd.com>
Subject: [PATCH 2/2] drm/amd/ras: record the fatal state on every device of the hive
Date: Mon, 24 Aug 2026 18:21:11 +0800	[thread overview]
Message-ID: <20260824102111.264161-2-xiang.liu@amd.com> (raw)
In-Reply-To: <20260824102111.264161-1-xiang.liu@amd.com>

The fatal error interrupt is broadcast to every device of the hive and
they all race for amdgpu_ras_global_ras_isr(), which hands -EBUSY to
everyone but the winner. Treating that as a failure returns before the
device is marked, so seven devices out of eight are left without their
fatal and SRAM ECC state, the one that actually logged the error among
them. KFD then tells the processes on those devices that the reset was
a plain hang.

-EBUSY only means the reset has already been asked for. Record the
state anyway and leave the request to the winner.

Signed-off-by: Xiang Liu <xiang.liu@amd.com>
---
 drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c | 10 ++++++++--
 1 file changed, 8 insertions(+), 2 deletions(-)

diff --git a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
index 081516c46cf8..291b2a96cbb5 100644
--- a/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
+++ b/drivers/gpu/drm/amd/ras/ras_mgr/amdgpu_ras_sys.c
@@ -33,8 +33,14 @@ static int amdgpu_ras_sys_detect_fatal_event(struct ras_core_context *ras_core,
 	uint64_t seq_no;
 
 	ret = amdgpu_ras_global_ras_isr(adev);
-	if (ret)
-		return ret;
+	if (ret) {
+		/* Another device of the hive already asked for the reset, this
+		 * one still has to record that it saw the error.
+		 */
+		kgd2kfd_set_sram_ecc_flag(adev->kfd.dev);
+		amdgpu_ras_set_fed(adev, true);
+		return ret == -EBUSY ? 0 : ret;
+	}
 
 	seq_no = amdgpu_ras_mgr_gen_ras_event_seqno(adev, RAS_SEQNO_TYPE_UE);
 	RAS_DEV_INFO(adev,
-- 
2.34.1


  reply	other threads:[~2026-08-24 10:22 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-24 10:21 [PATCH 1/2] drm/amd/ras: tell KFD the reset came from an ECC error Xiang Liu
2026-08-24 10:21 ` Xiang Liu [this message]
2026-08-24 11:01 ` Zhang, Hawking

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260824102111.264161-2-xiang.liu@amd.com \
    --to=xiang.liu@amd.com \
    --cc=Hawking.Zhang@amd.com \
    --cc=Stanley.Yang@amd.com \
    --cc=Tao.Zhou1@amd.com \
    --cc=YiPeng.Chai@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.