From: Soham Purkait <soham.purkait@intel.com>
To: igt-dev@lists.freedesktop.org, riana.tauro@intel.com,
badal.nilawar@intel.com, kamil.konieczny@intel.com,
sk.anirban@intel.com, raag.jadav@intel.com
Cc: anshuman.gupta@intel.com, soham.purkait@intel.com
Subject: [PATCH i-g-t v6] tests/intel/xe_ras: Add test for GPU health indicator
Date: Mon, 5 Oct 2026 23:17:02 +0530 [thread overview]
Message-ID: <20261005174702.2574483-2-soham.purkait@intel.com> (raw)
Add a new Xe RAS test exercising the gpu_health sysfs attribute
exposed by the Xe driver on platforms that provide the system
controller. The attribute reports and allows updating the GPU
health state.
The gpu-health subtest validates each valid state by writing it
and reading the value back, and checks that invalid writes are
rejected with -EINVAL, guarding against regressions.
v1:
- Add platform-conditional skip and exit handler. (Anirban)
v2:
- Add dynamic subtests for each case. (Riana)
v3:
- Add health state description.
- Store original health state as an index instead of a char pointer. (Riana)
Signed-off-by: Soham Purkait <soham.purkait@intel.com>
Reviewed-by: Sk Anirban <sk.anirban@intel.com>
---
tests/intel/xe_ras.c | 158 +++++++++++++++++++++++++++++++++++++++++++
tests/meson.build | 1 +
2 files changed, 159 insertions(+)
create mode 100644 tests/intel/xe_ras.c
diff --git a/tests/intel/xe_ras.c b/tests/intel/xe_ras.c
new file mode 100644
index 000000000..c98a10f00
--- /dev/null
+++ b/tests/intel/xe_ras.c
@@ -0,0 +1,158 @@
+// SPDX-License-Identifier: MIT
+/*
+ * Copyright © 2026 Intel Corporation
+ */
+
+#include <errno.h>
+#include <string.h>
+#include <unistd.h>
+
+#include "igt.h"
+#include "igt_sysfs.h"
+
+/**
+ * TEST: Test Xe RAS (Reliability, Availability, Serviceability) functionality
+ * Category: Core
+ * Mega feature: RAS
+ * Sub-category: RAS tests
+ * Functionality: ras
+ * Test category: Functional tests
+ *
+ * SUBTEST: gpu-health
+ * Description: Verify the gpu_health sysfs attribute accepts each valid state
+ * and rejects invalid writes with -EINVAL.
+ * The valid states are:
+ * ok - The gpu is healthy and operating within normal
+ * parameters.
+ * warning - The gpu is experiencing minor issues but remains
+ * operational.
+ * critical - The gpu is in a critical state and may not be
+ * operational.
+ */
+
+IGT_TEST_DESCRIPTION("Tests for Xe RAS");
+
+#define GPU_HEALTH_ATTR "device/gpu_health"
+
+static const char * const gpu_health_states[] = {
+ "ok",
+ "warning",
+ "critical",
+ "invalid-input",
+};
+
+#define GPU_HEALTH_VALID_STATE_COUNT (ARRAY_SIZE(gpu_health_states) - 1)
+
+static struct {
+ int sys_fd;
+ int orig_health;
+} gpu_health_ctx = { .sys_fd = -1, .orig_health = -1 };
+
+static void restore_gpu_health(int sig)
+{
+ if (gpu_health_ctx.sys_fd < 0)
+ return;
+
+ if (gpu_health_ctx.orig_health >= 0) {
+ igt_sysfs_set(gpu_health_ctx.sys_fd, GPU_HEALTH_ATTR,
+ gpu_health_states[gpu_health_ctx.orig_health]);
+ igt_info("Restored initial gpu_health to '%s'\n",
+ gpu_health_states[gpu_health_ctx.orig_health]);
+ gpu_health_ctx.orig_health = -1;
+ }
+ close(gpu_health_ctx.sys_fd);
+ gpu_health_ctx.sys_fd = -1;
+}
+
+static bool valid_gpu_health(const char *s)
+{
+ int i;
+
+ if (!s)
+ return false;
+
+ for (i = 0; i < GPU_HEALTH_VALID_STATE_COUNT; i++)
+ if (!strcmp(s, gpu_health_states[i]))
+ return true;
+
+ return false;
+}
+
+static void setup_gpu_health(int xe)
+{
+ char *health;
+ int i;
+
+ gpu_health_ctx.sys_fd = igt_sysfs_open(xe);
+ igt_assert(gpu_health_ctx.sys_fd >= 0);
+
+ igt_skip_on_f(!igt_sysfs_has_attr(gpu_health_ctx.sys_fd, GPU_HEALTH_ATTR),
+ "gpu_health sysfs attribute not exposed by driver\n");
+
+ health = igt_sysfs_get(gpu_health_ctx.sys_fd, GPU_HEALTH_ATTR);
+ igt_assert_f(valid_gpu_health(health),
+ "Unexpected initial gpu_health value: '%s'\n", health);
+
+ for (i = 0; i < GPU_HEALTH_VALID_STATE_COUNT; i++)
+ if (!strcmp(health, gpu_health_states[i]))
+ gpu_health_ctx.orig_health = i;
+
+ igt_info("Initial gpu_health: %s\n", health);
+ free(health);
+}
+
+static void test_gpu_health_state(const char *state)
+{
+ char *health;
+ int ret;
+
+ if (!valid_gpu_health(state)) {
+ ret = igt_sysfs_write(gpu_health_ctx.sys_fd, GPU_HEALTH_ATTR,
+ state, strlen(state));
+ igt_assert_f(ret == -EINVAL,
+ "Writing invalid value to %s returned %d, expected -EINVAL\n",
+ GPU_HEALTH_ATTR, ret);
+ return;
+ }
+
+ igt_info("Setting gpu_health to '%s'\n", state);
+
+ igt_assert_f(igt_sysfs_set(gpu_health_ctx.sys_fd, GPU_HEALTH_ATTR, state),
+ "Failed to write '%s' to %s\n",
+ state, GPU_HEALTH_ATTR);
+
+ health = igt_sysfs_get(gpu_health_ctx.sys_fd, GPU_HEALTH_ATTR);
+ igt_assert(health);
+ igt_assert_f(!strcmp(health, state),
+ "gpu_health readback mismatch: wrote '%s', read '%s'\n",
+ state, health);
+ igt_info("Verified gpu_health state readback for '%s' successfully\n",
+ state);
+ free(health);
+}
+
+int igt_main()
+{
+ int xe;
+ int state;
+
+ igt_fixture() {
+ xe = drm_open_driver(DRIVER_XE);
+ }
+
+ igt_subtest_with_dynamic("gpu-health") {
+ igt_install_exit_handler(restore_gpu_health);
+ setup_gpu_health(xe);
+
+ for (state = 0; state < ARRAY_SIZE(gpu_health_states); state++) {
+ igt_dynamic_f("%s", gpu_health_states[state]) {
+ test_gpu_health_state(gpu_health_states[state]);
+ }
+ }
+
+ restore_gpu_health(0);
+ }
+
+ igt_fixture()
+ drm_close_driver(xe);
+}
diff --git a/tests/meson.build b/tests/meson.build
index 1ac89bab7..3c2d4fdda 100644
--- a/tests/meson.build
+++ b/tests/meson.build
@@ -333,6 +333,7 @@ intel_xe_progs = [
'xe_prime_self_import',
'xe_pxp',
'xe_query',
+ 'xe_ras',
'xe_render_copy',
'xe_vm',
'xe_userptr_pressure',
--
2.43.0
next reply other threads:[~2026-10-05 17:47 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-05 17:47 Soham Purkait [this message]
2026-10-05 18:42 ` ✓ i915.CI.BAT: success for tests/intel/xe_ras: Add test for GPU health indicator (rev7) Patchwork
2026-10-05 18:59 ` ✓ Xe.CI.BAT: " Patchwork
2026-10-05 23:24 ` ✗ i915.CI.Full: failure " Patchwork
2026-10-06 1:46 ` ✗ Xe.CI.FULL: " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261005174702.2574483-2-soham.purkait@intel.com \
--to=soham.purkait@intel.com \
--cc=anshuman.gupta@intel.com \
--cc=badal.nilawar@intel.com \
--cc=igt-dev@lists.freedesktop.org \
--cc=kamil.konieczny@intel.com \
--cc=raag.jadav@intel.com \
--cc=riana.tauro@intel.com \
--cc=sk.anirban@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.