All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH] drm/xe/guc: Guard page-fault ack with runtime PM check
@ 2026-09-03  6:27 Varun Gupta
  2026-09-03  6:33 ` ✗ CI.checkpatch: warning for " Patchwork
                   ` (4 more replies)
  0 siblings, 5 replies; 9+ messages in thread
From: Varun Gupta @ 2026-09-03  6:27 UTC (permalink / raw)
  To: intel-xe; +Cc: stuart.summers, matthew.brost, szymon.markiewicz

During VM teardown, the VM's runtime PM reference is dropped
asynchronously, allowing the device to autosuspend while stale page
faults belonging to the now-dead VM are still queued. When the
page-fault worker later tries to ack one of these, it calls into
guc_ct_send_locked() on an already-suspended device, tripping:

  Assertion `!xe_pm_runtime_suspended(xe)` failed!
  WARNING at xe_device.c:1267 xe_device_assert_mem_access+0x11c/0x140 [xe]

A live VM/exec queue always holds a PM reference while it has
outstanding work, so if the device is suspended at ack time, the
owning context is already gone and the fault is stale.

Guard both ack sites (guc_ack_fault() and the batched flush in
guc_ack_fault_end()) with xe_pm_runtime_get_if_active() instead of
forcing a resume. get_if_active() is non-blocking, so it's safe to
call under guc->ct.lock, unlike a blocking get() which would risk
lock-ordering issues against the suspend path and needlessly wake
the device to deliver an ack nobody is waiting for.

Reported-by: Szymon Markiewicz <szymon.markiewicz@intel.com>
Signed-off-by: Varun Gupta <varun.gupta@intel.com>
---
 drivers/gpu/drm/xe/xe_guc_pagefault.c | 18 ++++++++++++++++--
 1 file changed, 16 insertions(+), 2 deletions(-)

diff --git a/drivers/gpu/drm/xe/xe_guc_pagefault.c b/drivers/gpu/drm/xe/xe_guc_pagefault.c
index 8f8210a732e9..a18de20c8245 100644
--- a/drivers/gpu/drm/xe/xe_guc_pagefault.c
+++ b/drivers/gpu/drm/xe/xe_guc_pagefault.c
@@ -9,6 +9,7 @@
 #include "xe_guc_pagefault.h"
 #include "xe_pagefault.h"
 #include "xe_pagefault_types.h"
+#include "xe_pm.h"
 
 #define XE_GUC_PAGEFAULT_FLUSH_PERIOD	BIT(4)	/* Sixteen */
 
@@ -52,19 +53,32 @@ static void guc_ack_fault(struct xe_pagefault *pf, int err)
 		FIELD_PREP(PFR_PDATA, pdata),
 	};
 	struct xe_guc *guc = pf->producer.private;
-	bool write_only = guc->pagefault_ack_counter++ &
+	struct xe_device *xe = guc_to_xe(guc);
+	bool write_only;
+
+	if (!xe_pm_runtime_get_if_active(xe))
+		return;
+
+	write_only = guc->pagefault_ack_counter++ &
 		(XE_GUC_PAGEFAULT_FLUSH_PERIOD - 1);
 
 	xe_guc_ct_send_locked(&guc->ct, action, ARRAY_SIZE(action),
 			      write_only);
+
+	xe_pm_runtime_put(xe);
 }
 
 static void guc_ack_fault_end(void *private)
 {
 	struct xe_guc *guc = private;
+	struct xe_device *xe = guc_to_xe(guc);
 
-	if ((guc->pagefault_ack_counter & (XE_GUC_PAGEFAULT_FLUSH_PERIOD - 1)) != 1)
+	if ((guc->pagefault_ack_counter &
+	     (XE_GUC_PAGEFAULT_FLUSH_PERIOD - 1)) != 1 &&
+	    xe_pm_runtime_get_if_active(xe)) {
 		xe_guc_ct_send_flush(&guc->ct);
+		xe_pm_runtime_put(xe);
+	}
 	xe_guc_ct_unlock(&guc->ct);
 }
 
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread
* [PATCH v3] drm/xe: Guard page-fault worker with runtime PM check
@ 2026-09-07  5:00 Varun Gupta
  2026-09-07  5:22 ` Upadhyay, Tejas
  0 siblings, 1 reply; 9+ messages in thread
From: Varun Gupta @ 2026-09-07  5:00 UTC (permalink / raw)
  To: intel-xe

During VM teardown, the VM's runtime PM reference is dropped
asynchronously, allowing the device to autosuspend while stale page
faults belonging to the now-dead VM are still queued. When the
page-fault worker later tries to ack one of these, it calls into
guc_ct_send_locked() on an already-suspended device, tripping:

  Assertion `!xe_pm_runtime_suspended(xe)` failed!
  WARNING at xe_device.c:1267 xe_device_assert_mem_access+0x11c/0x140 [xe]

A live VM/exec queue always holds a PM reference while it has
outstanding work, so if the device is suspended at ack time, the
owning context is already gone and the fault is stale.

Take a runtime PM reference across the entire pagefault
queue worker to safely deliver acks for torn-down VMs.

v3:
 - Move PM ref to the generic xe_pagefault_queue_work using
   guard(xe_pm_runtime)(xe) instead of tracking it in the GuC
   backend(Matt Brost).

v2:
 - Hold PM ref across the entire batch (begin/end) instead of per-ack.
   This prevents the device from autosuspending mid-batch, which would
   leave write_only acks written but the end flush skipped, and skip
   counter++, desyncing the cadence check.(Himal)
 - Add a comment explaining stale faults.(Himal)

Fixes: f289f7807119 ("drm/xe: Add xe_guc_pagefault layer")
Reported-by: Szymon Markiewicz <szymon.markiewicz@intel.com>
Signed-off-by: Varun Gupta <varun.gupta@intel.com>
Reviewed-by: Matthew Brost <matthew.brost@intel.com>
---
 drivers/gpu/drm/xe/xe_pagefault.c | 8 ++++++++
 1 file changed, 8 insertions(+)

diff --git a/drivers/gpu/drm/xe/xe_pagefault.c b/drivers/gpu/drm/xe/xe_pagefault.c
index 2e415995f067..3e8e8f7f5f40 100644
--- a/drivers/gpu/drm/xe/xe_pagefault.c
+++ b/drivers/gpu/drm/xe/xe_pagefault.c
@@ -17,6 +17,7 @@
 #include "xe_log.h"
 #include "xe_pagefault.h"
 #include "xe_pagefault_types.h"
+#include "xe_pm.h"
 #include "xe_svm.h"
 #include "xe_trace_bo.h"
 #include "xe_vm.h"
@@ -604,6 +605,13 @@ static void xe_pagefault_queue_work(struct work_struct *w)
 	u64 cache_start = XE_PAGEFAULT_CACHE_START_INVALID, cache_end = 0;
 	u32 cache_asid = 0;
 
+	/*
+	 * A live VM holds a PM reference, but a torn-down VM does not.
+	 * Guard the entire worker loop to safely drain stale faults and
+	 * prevent autosuspends from desyncing batched CT flushes.
+	 */
+	guard(xe_pm_runtime)(xe);
+
 #define USM_QUEUE_MAX_RUNTIME_MS      20
 	threshold = jiffies + msecs_to_jiffies(USM_QUEUE_MAX_RUNTIME_MS);
 
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-09-07  5:22 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-03  6:27 [PATCH] drm/xe/guc: Guard page-fault ack with runtime PM check Varun Gupta
2026-09-03  6:33 ` ✗ CI.checkpatch: warning for " Patchwork
2026-09-03  6:35 ` ✓ CI.KUnit: success " Patchwork
2026-09-03  7:17 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-03 18:53 ` ✗ Xe.CI.FULL: failure " Patchwork
2026-09-04  9:16 ` [PATCH v3] drm/xe: Guard page-fault worker " Varun Gupta
2026-09-04 18:00   ` Matthew Brost
  -- strict thread matches above, loose matches on Subject: below --
2026-09-07  5:00 Varun Gupta
2026-09-07  5:22 ` Upadhyay, Tejas

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.