From: Nitin Gote <nitin.r.gote@intel.com>
To: igt-dev@lists.freedesktop.org,
janusz.krzysztofik@linux.intel.com,
kamil.konieczny@linux.intel.com
Cc: matthew.auld@intel.com, nitin.r.gote@intel.com
Subject: [PATCH] tests/core_hotunplug: add *-with-load subtests exercising GPU workload during hotplug
Date: Mon, 18 May 2026 17:31:34 +0530 [thread overview]
Message-ID: <20260518120133.5651-2-nitin.r.gote@intel.com> (raw)
The six fd-holding hot* subtests keep a DRM fd open across unbind/unplug
but leave the GPU idle, so the kernel hotplug path is never exercised
against an active VM, exec queue and BOs.
Add six new dedicated *-with-load subtests that run a GPU spinner across
the unbind/unplug sequence so the device is removed while a workload is
actively using it.
The workload helpers use an explicit chipset dispatch block in
workload_start()/stop() so that other vendors can contribute
support by adding an else-if branch. New subtests are gated on
workload_available(), which currently covers DRIVER_XE and DRIVER_INTEL,
and returns false (SKIP) on any other driver.
igt_spin_free() is avoided in workload_stop() since the device may be
gone by then — igt_spin_end() writes the stop condition via a userspace
mmap and lets the DRM fd close reclaim kernel-side resources.
Validated on BMG and DG2 with a KASAN-enabled kernel:
all six new subtests pass with no KASAN reports.
v2: Pass the hotunplug priv struct to workload_start()/stop()
and dispatch via an explicit if/else-if chipset block to invite
other vendors to contribute. (Janusz)
Cc: Matthew Auld <matthew.auld@intel.com>
Cc: Janusz Krzysztofik <janusz.krzysztofik@linux.intel.com>
Cc: Kamil Konieczny <kamil.konieczny@linux.intel.com>
Assisted-by: Copilot:claude-opus-4.7
Signed-off-by: Nitin Gote <nitin.r.gote@intel.com>
---
tests/core_hotunplug.c | 290 +++++++++++++++++++++++++++++++++++++++++
1 file changed, 290 insertions(+)
diff --git a/tests/core_hotunplug.c b/tests/core_hotunplug.c
index 7c9dae1bf..38dcbb97d 100644
--- a/tests/core_hotunplug.c
+++ b/tests/core_hotunplug.c
@@ -80,6 +80,34 @@
* SUBTEST: unplug-rescan
* Description: Check if a device believed to be closed can be cleanly
* unplugged, then restored
+ *
+ * SUBTEST: hotunbind-rebind-with-load
+ * Description: Check if the driver can be cleanly unbound from an open device
+ * with a background GPU workload in flight, then released and
+ * rebound
+ *
+ * SUBTEST: hotunplug-rescan-with-load
+ * Description: Check if an open device with a background GPU workload in
+ * flight can be cleanly unplugged, then released and restored
+ *
+ * SUBTEST: hotrebind-with-load
+ * Description: Check if the driver can be cleanly rebound to a device with a
+ * still open hot unbound driver instance while a background GPU
+ * workload is in flight
+ *
+ * SUBTEST: hotreplug-with-load
+ * Description: Check if a hot unplugged and still open device can be cleanly
+ * restored while a background GPU workload is in flight
+ *
+ * SUBTEST: hotrebind-lateclose-with-load
+ * Description: Check if a hot unbound driver instance still open after hot
+ * rebind with a background GPU workload in flight can be cleanly
+ * released
+ *
+ * SUBTEST: hotreplug-lateclose-with-load
+ * Description: Check if an instance of a still open while hot replugged
+ * device with a background GPU workload in flight can be cleanly
+ * released
*/
IGT_TEST_DESCRIPTION("Examine behavior of a driver on device hot unplug");
@@ -565,6 +593,52 @@ static void set_filter_from_device(int fd)
igt_assert_eq(igt_device_filter_add(filter), 1);
}
+/* GPU workload helpers */
+
+struct gpu_workload {
+ igt_spin_t *spin;
+ uint64_t ahnd;
+};
+
+static bool workload_available(int chipset)
+{
+ return chipset == DRIVER_XE || chipset == DRIVER_INTEL;
+}
+
+static void workload_start(struct hotunplug *priv, int fd,
+ struct gpu_workload *w)
+{
+ if (priv->chipset == DRIVER_XE || priv->chipset == DRIVER_INTEL) {
+ local_debug("%s\n", "starting GPU spinner");
+ priv->failure = "GPU workload start failure!";
+ w->ahnd = intel_allocator_open(fd, 0, INTEL_ALLOCATOR_RELOC);
+ w->spin = igt_spin_new(fd, .ahnd = w->ahnd);
+ priv->failure = NULL;
+ }
+}
+
+static void workload_stop(struct hotunplug *priv, struct gpu_workload *w)
+{
+ if (priv->chipset == DRIVER_XE || priv->chipset == DRIVER_INTEL) {
+ if (!w->spin)
+ return;
+
+ /*
+ * The device may be gone (unbind/unplug), so avoid
+ * igt_spin_free(): on Xe it would wait forever on a syncobj
+ * that never signals; on i915 it would touch the dead fd via
+ * gem_munmap()/gem_close(). Instead, end the spinner via a
+ * userspace mmap write and let the DRM fd close reclaim the
+ * kernel-side resources.
+ */
+ local_debug("%s\n", "stopping GPU spinner");
+ igt_spin_end(w->spin);
+ put_ahnd(w->ahnd);
+ w->spin = NULL;
+ w->ahnd = 0;
+ }
+}
+
/* Subtests */
static void unbind_rebind(struct hotunplug *priv)
@@ -679,6 +753,150 @@ static void hotreplug_lateclose(struct hotunplug *priv)
igt_assert_f(healthcheck(priv, false), "%s\n", priv->failure);
}
+static void hotunbind_rebind_with_load(struct hotunplug *priv)
+{
+ struct gpu_workload w = { 0 };
+
+ pre_check(priv);
+
+ igt_require_f(workload_available(priv->chipset),
+ "No GPU workload support for this driver\n");
+
+ priv->fd.drm = local_drm_open_driver(false, "", " for hot unbind");
+
+ workload_start(priv, priv->fd.drm, &w);
+
+ driver_unbind(priv, "hot ", 0);
+
+ workload_stop(priv, &w);
+
+ priv->fd.drm = close_device(priv->fd.drm, "late ", "unbound ");
+ igt_assert_eq(priv->fd.drm, -1);
+
+ driver_bind(priv, 0);
+
+ igt_assert_f(healthcheck(priv, false), "%s\n", priv->failure);
+}
+
+static void hotunplug_rescan_with_load(struct hotunplug *priv)
+{
+ struct gpu_workload w = { 0 };
+
+ pre_check(priv);
+
+ igt_require_f(workload_available(priv->chipset),
+ "No GPU workload support for this driver\n");
+
+ priv->fd.drm = local_drm_open_driver(false, "", " for hot unplug");
+
+ workload_start(priv, priv->fd.drm, &w);
+
+ device_unplug(priv, "hot ", 0);
+
+ workload_stop(priv, &w);
+
+ priv->fd.drm = close_device(priv->fd.drm, "late ", "removed ");
+ igt_assert_eq(priv->fd.drm, -1);
+
+ bus_rescan(priv, 0);
+
+ igt_assert_f(healthcheck(priv, false), "%s\n", priv->failure);
+}
+
+static void hotrebind_with_load(struct hotunplug *priv)
+{
+ struct gpu_workload w = { 0 };
+
+ pre_check(priv);
+
+ igt_require_f(workload_available(priv->chipset),
+ "No GPU workload support for this driver\n");
+
+ priv->fd.drm = local_drm_open_driver(false, "", " for hot rebind");
+
+ workload_start(priv, priv->fd.drm, &w);
+
+ driver_unbind(priv, "hot ", 60);
+
+ driver_bind(priv, 0);
+
+ workload_stop(priv, &w);
+
+ igt_assert_f(healthcheck(priv, false), "%s\n", priv->failure);
+}
+
+static void hotreplug_with_load(struct hotunplug *priv)
+{
+ struct gpu_workload w = { 0 };
+
+ pre_check(priv);
+
+ igt_require_f(workload_available(priv->chipset),
+ "No GPU workload support for this driver\n");
+
+ priv->fd.drm = local_drm_open_driver(false, "", " for hot replug");
+
+ workload_start(priv, priv->fd.drm, &w);
+
+ device_unplug(priv, "hot ", 60);
+
+ bus_rescan(priv, 0);
+
+ workload_stop(priv, &w);
+
+ igt_assert_f(healthcheck(priv, false), "%s\n", priv->failure);
+}
+
+static void hotrebind_lateclose_with_load(struct hotunplug *priv)
+{
+ struct gpu_workload w = { 0 };
+
+ pre_check(priv);
+
+ igt_require_f(workload_available(priv->chipset),
+ "No GPU workload support for this driver\n");
+
+ priv->fd.drm = local_drm_open_driver(false, "", " for hot rebind");
+
+ workload_start(priv, priv->fd.drm, &w);
+
+ driver_unbind(priv, "hot ", 60);
+
+ driver_bind(priv, 0);
+
+ workload_stop(priv, &w);
+
+ priv->fd.drm = close_device(priv->fd.drm, "late ", "unbound ");
+ igt_assert_eq(priv->fd.drm, -1);
+
+ igt_assert_f(healthcheck(priv, false), "%s\n", priv->failure);
+}
+
+static void hotreplug_lateclose_with_load(struct hotunplug *priv)
+{
+ struct gpu_workload w = { 0 };
+
+ pre_check(priv);
+
+ igt_require_f(workload_available(priv->chipset),
+ "No GPU workload support for this driver\n");
+
+ priv->fd.drm = local_drm_open_driver(false, "", " for hot replug");
+
+ workload_start(priv, priv->fd.drm, &w);
+
+ device_unplug(priv, "hot ", 60);
+
+ bus_rescan(priv, 0);
+
+ workload_stop(priv, &w);
+
+ priv->fd.drm = close_device(priv->fd.drm, "late ", "removed ");
+ igt_assert_eq(priv->fd.drm, -1);
+
+ igt_assert_f(healthcheck(priv, false), "%s\n", priv->failure);
+}
+
/* Main */
int igt_main()
@@ -817,6 +1035,78 @@ int igt_main()
recover(&priv);
}
+ igt_fixture()
+ post_healthcheck(&priv);
+
+ igt_subtest_group() {
+ igt_describe("Check if the driver can be cleanly unbound from an open device with a background GPU workload in flight, then released and rebound");
+ igt_subtest("hotunbind-rebind-with-load")
+ hotunbind_rebind_with_load(&priv);
+
+ igt_fixture()
+ recover(&priv);
+ }
+
+ igt_fixture()
+ post_healthcheck(&priv);
+
+ igt_subtest_group() {
+ igt_describe("Check if an open device with a background GPU workload in flight can be cleanly unplugged, then released and restored");
+ igt_subtest("hotunplug-rescan-with-load")
+ hotunplug_rescan_with_load(&priv);
+
+ igt_fixture()
+ recover(&priv);
+ }
+
+ igt_fixture()
+ post_healthcheck(&priv);
+
+ igt_subtest_group() {
+ igt_describe("Check if the driver can be cleanly rebound to a device with a still open hot unbound driver instance while a background GPU workload is in flight");
+ igt_subtest("hotrebind-with-load")
+ hotrebind_with_load(&priv);
+
+ igt_fixture()
+ recover(&priv);
+ }
+
+ igt_fixture()
+ post_healthcheck(&priv);
+
+ igt_subtest_group() {
+ igt_describe("Check if a hot unplugged and still open device can be cleanly restored while a background GPU workload is in flight");
+ igt_subtest("hotreplug-with-load")
+ hotreplug_with_load(&priv);
+
+ igt_fixture()
+ recover(&priv);
+ }
+
+ igt_fixture()
+ post_healthcheck(&priv);
+
+ igt_subtest_group() {
+ igt_describe("Check if a hot unbound driver instance still open after hot rebind with a background GPU workload in flight can be cleanly released");
+ igt_subtest("hotrebind-lateclose-with-load")
+ hotrebind_lateclose_with_load(&priv);
+
+ igt_fixture()
+ recover(&priv);
+ }
+
+ igt_fixture()
+ post_healthcheck(&priv);
+
+ igt_subtest_group() {
+ igt_describe("Check if an instance of a still open while hot replugged device with a background GPU workload in flight can be cleanly released");
+ igt_subtest("hotreplug-lateclose-with-load")
+ hotreplug_lateclose_with_load(&priv);
+
+ igt_fixture()
+ recover(&priv);
+ }
+
igt_fixture() {
post_healthcheck(&priv);
--
2.50.1
next reply other threads:[~2026-05-18 11:26 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-05-18 12:01 Nitin Gote [this message]
2026-05-18 12:53 ` [PATCH] tests/core_hotunplug: add *-with-load subtests exercising GPU workload during hotplug Janusz Krzysztofik
-- strict thread matches above, loose matches on Subject: below --
2026-08-21 6:58 Nitin Gote
2026-08-21 9:43 ` Janusz Krzysztofik
2026-08-21 14:36 ` Gote, Nitin R
2026-05-18 15:48 Nitin Gote
2026-05-20 13:38 ` Janusz Krzysztofik
2026-05-21 12:46 ` Kamil Konieczny
2026-05-18 8:27 Nitin Gote
2026-05-18 8:48 ` Janusz Krzysztofik
2026-05-18 9:59 ` Gote, Nitin R
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260518120133.5651-2-nitin.r.gote@intel.com \
--to=nitin.r.gote@intel.com \
--cc=igt-dev@lists.freedesktop.org \
--cc=janusz.krzysztofik@linux.intel.com \
--cc=kamil.konieczny@linux.intel.com \
--cc=matthew.auld@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox