From: Raag Jadav <raag.jadav@intel.com>
To: intel-xe@lists.freedesktop.org
Cc: riana.tauro@intel.com, michal.wajdeczko@intel.com,
lukasz.laguna@intel.com, matthew.d.roper@intel.com,
matthew.brost@intel.com, rodrigo.vivi@intel.com,
Raag Jadav <raag.jadav@intel.com>
Subject: [PATCH v2 1/5] drm/xe/gt: Use GT ordered workqueue for wedging
Date: Mon, 31 Aug 2026 09:55:49 +0530 [thread overview]
Message-ID: <20260831042633.1760474-2-raag.jadav@intel.com> (raw)
In-Reply-To: <20260831042633.1760474-1-raag.jadav@intel.com>
Currently, xe_gt_declare_wedged() stops GuC CT and initiates exec queue
teardown synchronously. This is problematic in cases where the jobs are
still in-flight when the device is declated wedged.
Introduce a worker for GT specific wedge handling and queue the teardown
on GT ordered workqueue, so we don't disrupt the scheduler while jobs
are still in-flight.
Fixes: c9474b726b93 ("drm/xe: Wedge the entire device")
Signed-off-by: Raag Jadav <raag.jadav@intel.com>
---
drivers/gpu/drm/xe/xe_gt.c | 15 +++++++++++++--
drivers/gpu/drm/xe/xe_gt_types.h | 9 +++++++++
2 files changed, 22 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_gt.c b/drivers/gpu/drm/xe/xe_gt.c
index 478e047031f4..5c201327a747 100644
--- a/drivers/gpu/drm/xe/xe_gt.c
+++ b/drivers/gpu/drm/xe/xe_gt.c
@@ -173,6 +173,7 @@ static void xe_gt_enable_comp_1wcoh(struct xe_gt *gt)
}
static void gt_reset_worker(struct work_struct *w);
+static void gt_wedge_worker(struct work_struct *w);
static int emit_job_sync(struct xe_exec_queue *q, struct xe_bb *bb,
long timeout_jiffies, bool force_reset)
@@ -704,6 +705,8 @@ static void xe_gt_fini(void *arg)
struct xe_gt *gt = arg;
int i;
+ disable_work_sync(>->wedge.worker);
+
if (disable_work_sync(>->reset.worker))
/*
* If gt_reset_worker was halted from executing, take care of
@@ -723,6 +726,7 @@ int xe_gt_init(struct xe_gt *gt)
int i;
INIT_WORK(>->reset.worker, gt_reset_worker);
+ INIT_WORK(>->wedge.worker, gt_wedge_worker);
for (i = 0; i < XE_ENGINE_CLASS_MAX; ++i) {
gt->ring_ops[i] = xe_ring_ops_get(gt, i);
@@ -1005,6 +1009,14 @@ void xe_gt_reset_async(struct xe_gt *gt)
xe_pm_runtime_put(xe);
}
+static void gt_wedge_worker(struct work_struct *w)
+{
+ struct xe_gt *gt = container_of(w, typeof(*gt), wedge.worker);
+
+ xe_uc_declare_wedged(>->uc);
+ xe_tlb_inval_reset(>->tlb_inval);
+}
+
void xe_gt_suspend_prepare(struct xe_gt *gt)
{
xe_uc_suspend_prepare(>->uc);
@@ -1195,6 +1207,5 @@ void xe_gt_declare_wedged(struct xe_gt *gt)
{
xe_gt_assert(gt, gt_to_xe(gt)->wedged.mode);
- xe_uc_declare_wedged(>->uc);
- xe_tlb_inval_reset(>->tlb_inval);
+ queue_work(gt->ordered_wq, >->wedge.worker);
}
diff --git a/drivers/gpu/drm/xe/xe_gt_types.h b/drivers/gpu/drm/xe/xe_gt_types.h
index 628911346455..ca1e2f82470d 100644
--- a/drivers/gpu/drm/xe/xe_gt_types.h
+++ b/drivers/gpu/drm/xe/xe_gt_types.h
@@ -233,6 +233,15 @@ struct xe_gt {
struct work_struct worker;
} reset;
+ /** @wedge: state for GT wedge */
+ struct {
+ /**
+ * @wedge.worker: work so GT wedge to be done async allowing the wedge
+ * code to safely flush all code paths
+ */
+ struct work_struct worker;
+ } wedge;
+
/** @tlb_inval: TLB invalidation state */
struct xe_tlb_inval tlb_inval;
--
2.43.0
next prev parent reply other threads:[~2026-08-31 4:26 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-31 4:25 [PATCH v2 0/5] Introduce xe_wedge Raag Jadav
2026-08-31 4:25 ` Raag Jadav [this message]
2026-08-31 4:25 ` [PATCH v2 2/5] drm/xe: Make xe_device_declare_wedged() IRQ safe Raag Jadav
2026-08-31 4:40 ` sashiko-bot
2026-08-31 4:25 ` [PATCH v2 3/5] drm/xe: Introduce xe_wedge Raag Jadav
2026-09-02 17:56 ` Rodrigo Vivi
2026-08-31 4:25 ` [PATCH v2 4/5] drm/xe/debugfs: Consolidate wedged_mode debt into xe_wedge Raag Jadav
2026-09-07 7:24 ` Laguna, Lukasz
2026-09-07 7:53 ` Raag Jadav
2026-09-07 8:40 ` Laguna, Lukasz
2026-08-31 4:25 ` [PATCH v2 5/5] drm/xe/wedge: Update naming to match with xe_wedge Raag Jadav
2026-08-31 4:35 ` sashiko-bot
2026-08-31 4:33 ` ✗ CI.checkpatch: warning for Introduce xe_wedge Patchwork
2026-08-31 4:34 ` ✗ CI.KUnit: failure " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260831042633.1760474-2-raag.jadav@intel.com \
--to=raag.jadav@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=lukasz.laguna@intel.com \
--cc=matthew.brost@intel.com \
--cc=matthew.d.roper@intel.com \
--cc=michal.wajdeczko@intel.com \
--cc=riana.tauro@intel.com \
--cc=rodrigo.vivi@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.