From: Raag Jadav <raag.jadav@intel.com>
To: intel-xe@lists.freedesktop.org
Cc: riana.tauro@intel.com, michal.wajdeczko@intel.com,
lukasz.laguna@intel.com, matthew.d.roper@intel.com,
matthew.brost@intel.com, rodrigo.vivi@intel.com,
Raag Jadav <raag.jadav@intel.com>
Subject: [PATCH v3 1/5] drm/xe/gt: Use GT ordered workqueue for wedging
Date: Mon, 7 Sep 2026 14:04:36 +0530 [thread overview]
Message-ID: <20260907083541.2194747-2-raag.jadav@intel.com> (raw)
In-Reply-To: <20260907083541.2194747-1-raag.jadav@intel.com>
Currently, xe_gt_declare_wedged() stops GuC CT and initiates exec queue
teardown synchronously. This can result in tdr timeout in cases where
the device is declared wedged while jobs are still in-flight.
Introduce a worker for GT specific wedge handling and queue the teardown
on GT ordered workqueue, so we don't disrupt the scheduler while jobs
are still in-flight.
Fixes: c9474b726b93 ("drm/xe: Wedge the entire device")
Signed-off-by: Raag Jadav <raag.jadav@intel.com>
---
v2: Split fixes into separate patches (Rodrigo, Michal)
---
drivers/gpu/drm/xe/xe_gt.c | 17 +++++++++++++++--
drivers/gpu/drm/xe/xe_gt_types.h | 6 ++++++
2 files changed, 21 insertions(+), 2 deletions(-)
diff --git a/drivers/gpu/drm/xe/xe_gt.c b/drivers/gpu/drm/xe/xe_gt.c
index 478e047031f4..200e34331c19 100644
--- a/drivers/gpu/drm/xe/xe_gt.c
+++ b/drivers/gpu/drm/xe/xe_gt.c
@@ -173,6 +173,7 @@ static void xe_gt_enable_comp_1wcoh(struct xe_gt *gt)
}
static void gt_reset_worker(struct work_struct *w);
+static void gt_wedge_worker(struct work_struct *w);
static int emit_job_sync(struct xe_exec_queue *q, struct xe_bb *bb,
long timeout_jiffies, bool force_reset)
@@ -704,6 +705,8 @@ static void xe_gt_fini(void *arg)
struct xe_gt *gt = arg;
int i;
+ disable_work_sync(>->wedge.worker);
+
if (disable_work_sync(>->reset.worker))
/*
* If gt_reset_worker was halted from executing, take care of
@@ -723,6 +726,7 @@ int xe_gt_init(struct xe_gt *gt)
int i;
INIT_WORK(>->reset.worker, gt_reset_worker);
+ INIT_WORK(>->wedge.worker, gt_wedge_worker);
for (i = 0; i < XE_ENGINE_CLASS_MAX; ++i) {
gt->ring_ops[i] = xe_ring_ops_get(gt, i);
@@ -1005,6 +1009,14 @@ void xe_gt_reset_async(struct xe_gt *gt)
xe_pm_runtime_put(xe);
}
+static void gt_wedge_worker(struct work_struct *w)
+{
+ struct xe_gt *gt = container_of(w, typeof(*gt), wedge.worker);
+
+ xe_uc_declare_wedged(>->uc);
+ xe_tlb_inval_reset(>->tlb_inval);
+}
+
void xe_gt_suspend_prepare(struct xe_gt *gt)
{
xe_uc_suspend_prepare(>->uc);
@@ -1195,6 +1207,7 @@ void xe_gt_declare_wedged(struct xe_gt *gt)
{
xe_gt_assert(gt, gt_to_xe(gt)->wedged.mode);
- xe_uc_declare_wedged(>->uc);
- xe_tlb_inval_reset(>->tlb_inval);
+ /* Prevent queueing uninitialized worker if hit during probe */
+ if (gt->wedge.worker.func)
+ queue_work(gt->ordered_wq, >->wedge.worker);
}
diff --git a/drivers/gpu/drm/xe/xe_gt_types.h b/drivers/gpu/drm/xe/xe_gt_types.h
index 628911346455..b11ac52b0ab7 100644
--- a/drivers/gpu/drm/xe/xe_gt_types.h
+++ b/drivers/gpu/drm/xe/xe_gt_types.h
@@ -233,6 +233,12 @@ struct xe_gt {
struct work_struct worker;
} reset;
+ /** @wedge: state for GT wedge */
+ struct {
+ /** @wedge.worker: worker for GT wedge to be done async */
+ struct work_struct worker;
+ } wedge;
+
/** @tlb_inval: TLB invalidation state */
struct xe_tlb_inval tlb_inval;
--
2.43.0
next prev parent reply other threads:[~2026-09-07 8:36 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-07 8:34 [PATCH v3 0/5] Introduce xe_wedge Raag Jadav
2026-09-07 8:34 ` Raag Jadav [this message]
2026-09-07 8:50 ` [PATCH v3 1/5] drm/xe/gt: Use GT ordered workqueue for wedging sashiko-bot
2026-09-07 12:02 ` Michal Wajdeczko
2026-09-07 13:06 ` Raag Jadav
2026-09-07 8:34 ` [PATCH v3 2/5] drm/xe: Move xe_device_wedged_fini() registration to xe_device_probe_early() Raag Jadav
2026-09-07 8:54 ` sashiko-bot
2026-09-07 8:34 ` [PATCH v3 3/5] drm/xe: Make xe_device_declare_wedged() IRQ safe Raag Jadav
2026-09-07 8:34 ` [PATCH v3 4/5] drm/xe: Introduce xe_wedge Raag Jadav
2026-09-07 8:34 ` [PATCH v3 5/5] drm/xe/ras: Move xe_ras_process_errors() to xe_ras Raag Jadav
2026-09-07 8:55 ` sashiko-bot
2026-09-07 9:35 ` ✗ CI.checkpatch: warning for Introduce xe_wedge (rev2) Patchwork
2026-09-07 9:37 ` ✓ CI.KUnit: success " Patchwork
2026-09-07 10:26 ` ✓ Xe.CI.BAT: " Patchwork
2026-09-07 11:39 ` ✓ Xe.CI.FULL: " Patchwork
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260907083541.2194747-2-raag.jadav@intel.com \
--to=raag.jadav@intel.com \
--cc=intel-xe@lists.freedesktop.org \
--cc=lukasz.laguna@intel.com \
--cc=matthew.brost@intel.com \
--cc=matthew.d.roper@intel.com \
--cc=michal.wajdeczko@intel.com \
--cc=riana.tauro@intel.com \
--cc=rodrigo.vivi@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.